Serasa Platform

Serasa Experian · Platform transformation · 2021

From tightly coupled architecture to a platform built for scale

Leadership of a transformation combining performance, reliability, architecture, and operating model for products including Limpa Nome, eCred, and Premium.

43 miregistered users
260 mimonthly accesses
20 mil/speak sustained for 10 minutes
0–1incidents/month after the change

The challenge

Strategic products grew rapidly, but critical integrations, mainframe dependencies, and duplicated capabilities created a tightly coupled architecture. Marketing campaigns frequently caused incidents, and some journeys took almost four minutes.

Before

  • Reactive and expensive scaling during incidents
  • Unclear ownership between platform and vertical teams
  • Heavy operations in the synchronous customer journey
  • Recurring campaign failures
→

Target state

  • Centralized and reusable shared capabilities
  • Clear end-to-end ownership
  • Observable asynchronous processing
  • Capacity planned before campaigns

What I personally led

Shared platform

Direct leadership of eight engineers responsible for the foundation used by Limpa Nome, eCred, and Premium.

Cross-functional coordination

Alignment across SRE, DevOps, infrastructure, product, security, and business-unit leaders.

Direction and priorities

Platform strategy, critical journeys, and requirements for LGPD, fraud prevention, access, performance, and reliability.

The key decision

Stop treating every failure as an isolated incident. We mapped journeys end to end and prioritized bottlenecks by customer impact, operational risk, scalability, and reuse potential.
1

Map

Latency, mainframe, integrations, duplication, and ownership.

2

Prioritize

Impact, risk, scale, and reuse potential.

3

Decouple

Caching, Go, Redis, queues, retries, and observability.

4

Operate

Capacity planning and contingency before campaigns.

Execution across two main workstreams

1 · Positive Credit Registry

Scalable caching, Redis-based enrichment, and cloud services in Go reduced dependency on remote processing in São Paulo and the mainframe.

Delivery: approximately 1 month

2 · Limpa Nome

Heavy operations were removed from the synchronous path and replaced with asynchronous processing, safe retries, better failure handling, and improved observability.

Transformation: approximately 2 months

3 · Campaigns

Product, SRE, DevOps, and infrastructure began reviewing traffic, scale, monitoring, and contingency plans in advance.

Model: continuous planning

Measurable outcomes

Critical journey latency≈ 4 min → 200 ms
Operational incidents≈ 20/month → 0–1/month
Campaign capacity2× previous traffic

Scale proven

Peaks of approximately 20,000 accesses per second for ten minutes.

Stable campaigns

Production stopped failing during marketing campaigns.

Cost avoided

Eliminated reliance on emergency scaling of up to ten additional machines for ten hours.

Approximate values based on project records.

Lesson de liderança

The problem was not only performance. It also involved architecture, ownership, and the operating model. Centralizing shared capabilities, clarifying responsibilities, and replacing reactive scaling with planning enabled growth with greater confidence, lower cost, and higher reliability.
Engineering and platform leadership case studySerasa Experian · 2021
Scroll to Top