Re-architecting checkout across seven microservices
How we moved Urban Company's multi-category checkout to a new money-flow stack with zero downtime — and what distributed systems lessons stuck.
At Urban Company, checkout was not one service — it was a graph. Cart, pricing, coupons, payments, fulfilment, notifications: each owned by a different team, each with its own release cadence. Adding multi-category support meant every hop in that graph had to understand categories without breaking existing flows.
The constraint: zero downtime
We could not flip a feature flag and pray. Home services run 24/7; a broken checkout is revenue stopped. The migration had to be incremental: dual-write, shadow reads, cutover per category, rollback paths at every stage.
What we changed
- Money-flow stack redesign — decoupled payment capture from the user-facing journey so retries and failures did not block the UI.
- Redis distributed locks — multi-device coupon fraud was real; locks on redemption paths saved real money.
- Consolidated payment-summary library — 10+ ad-hoc implementations became one Node.js package every service imported.
Numbers that mattered
- 7+ microservices touched, zero breakage on cutover
- 1M+ daily events through the pipeline at 98% accuracy
- Gifts micro-frontend: 40% latency cut, +2% repeat conversion
Lessons
Ownership beats abstraction. We did not build a “checkout platform” in year one. We fixed the money path first, then extracted shared libraries once patterns were proven.
Observability is not optional. Grafana and ELK were how we knew shadow reads matched production before we cut traffic over.
Teams ship at different speeds. The architecture had to allow Service A to migrate while Service B still sent the old payload shape — versioned contracts, not big-bang deploys.
That project is why I still reach for event-driven designs and explicit failure modes before I reach for a new framework.