IT outsourcing
Platform engineering and reliability for a Latin American fintech
At a glance
- Industry
- Financial services and fintech
- Client geography
- Latin American fintech, multi-country
- Client size
- Scale-up, multi-market digital financial platform
- Service line
- IT outsourcing, platform engineering, DevOps, reliability
- Primary language
- English and Spanish
- Delivery site
- Bogotá
- Engagement duration
- 16 months, ongoing
- Team size
- 20 engineers, 2 leads, 7 senior, 9 mid, 2 junior
Client profile
The client is a Latin American fintech operating a digital financial platform across several countries, headquartered in the region with a distributed engineering organisation. Its platform handles payments and financial transactions at scale, where availability and reliability are existential rather than merely important.
The challenge
The fintech's platform availability stood at 97.2%, which for a payments platform represents an unacceptable amount of downtime. Every minute the platform is unavailable is transactions that fail and trust that erodes, and the company's rapid growth was outpacing the reliability of its infrastructure. The engineering organisation was consumed by incident response and had no capacity to improve the platform. Every reliability initiative in the preceding year had been abandoned partway when incident load reclaimed the team, a familiar and corrosive pattern. Deployment was slow and risky. Without progressive delivery infrastructure, every release carried full-platform risk, which made the team release infrequently, which made each release larger and riskier, a self-reinforcing problem. The fintech had tried to hire senior platform and reliability engineers across its markets and found the profile scarce and expensive, and its own engineers were being pulled toward feature work by product pressure.
Why Corpshore Colombia
Corpshore Colombia proposed a two-track engagement, a run track stabilising the platform and a build track constructing the reliability and delivery infrastructure, with the build track contractually ring-fenced from incident response so it could not be consumed by firefighting, which had killed the fintech's prior internal attempts.
Bogotá placed the team in the country's financial and engineering centre, on the client's clock, and Corpshore's experience with regulated financial platforms and its security posture mattered for a payments environment.
The engagement
Twenty engineers in Bogotá: two technical leads, seven senior, nine mid-level and two junior, split across a run track and a ringfenced build track. Run coverage is 24 hours for critical incidents; the build track works standard hours and is contractually protected from incident response. Stack: Kubernetes, Terraform, AWS, Go, Python, Prometheus, Grafana, with the client's existing platform.
Approach and methodology
Stabilise by cause, not by ticket. Incident response moved from firefighting to cause elimination, with every critical incident producing a root cause record and a preventive action. The availability curve reflects cause elimination rather than faster recovery. Ring-fenced reliability build. The build track constructed observability, progressive delivery and self-service infrastructure, protected from incident load, which is the only reason it survived to deliver where the fintech's prior attempts had not. Progressive delivery. Feature flags, canary deployments and automated rollback decoupled release from risk, allowing frequent, safe deployment. Documented knowledge transfer. The infrastructure and runbooks are built to a standard the client's own engineers operate, with defined transfer milestones.
Core platform availability
Results
Core platform availability rose from 97.2% to 99.87% by month 12, reducing monthly downtime from roughly twenty hours to under one, on a payments platform where that difference is existential. Mean time to resolution on critical incidents fell from over eleven hours to 0.7 hours, and critical incident frequency fell as cause elimination took hold. The fintech's own engineers were returned to feature work as the platform stabilised, and deployment frequency rose to weekly or better per team.
Enduring value
The observability and progressive delivery infrastructure are client-owned and operated by the fintech's engineers. The ring-fenced two-track model has become the client's template for infrastructure work. Corpshore Colombia has extended the engagement to security engineering for the platform.
Key indicators
| Metric | Baseline | Month 12 | Change |
|---|---|---|---|
| Core platform availability | 97.2% | 99.87% | +2.67 pts |
| Monthly downtime | ≈ 20 hours | < 1 hour | -95% |
| Mean time to resolution, critical | 11 h 10 m | 42 min | -94% |
| Critical incidents per month | 13 | 3 | -77% |
| Deployment frequency | Infrequent | Weekly+ | Structural change |
| Change failure rate | 22% | 7% | -68% |
| Reliability build initiatives completed | 0 of prior attempts | Delivered | New capability |
Where a metric disclosed only a change, the absolute figures are shown as a dash. Figures are client-reported or jointly measured.
Every reliability programme we had run was eaten by incidents. Ring-fencing the build team contractually is the only reason this one survived to deliver anything. On a payments platform, that reliability is the business.
Related topics