Products›💳 Payments›Stripe
💳

Stripe

Global financial infrastructure processing hundreds of billions

99.9999%
SLA
10
SERVICES
10
NODES
REQ/SEC0
LATENCY0ms
ERROR RATE0%
CACHE HIT95%
ACTIVE CONNS0
QUEUE DEPTH0

Stripe Architecture Blueprint

Click ▶ RUN to animate active particle streams across microservices

client
gateway
service
database
cache
queue
cdn
storage
SYSTEM ARCHITECTURE WALKTHROUGH

How Traffic Flows Through Stripe

EDGE TIER01

1. Ingress & Edge Routing

User requests arrive at the edge network. Global CDNs cache static assets and media. API Gateways terminate TLS, validate JWT authentication tokens, enforce token-bucket rate limits, and scrub malicious bot traffic before forwarding to internal services.

CustomerMerchant ServerAPI Gateway
APPLICATION TIER02

2. Microservice Processing

Stateless domain services execute core business logic. Microservices communicate via high-performance internal gRPC/REST APIs and autoscaling worker pods, ensuring that high load on one domain never exhausts compute resources of another.

Payment ServiceRadar Fraud MLWebhook Service
DATA TIER03

3. In-Memory Caching & Storage

Read-heavy traffic is served from in-memory Redis clusters with sub-millisecond latencies, protecting primary databases. Persistent databases (PostgreSQL, Cassandra, DynamoDB) maintain ACID consistency for financial ledgers, user accounts, and immutable state records.

Ledger ServicePayment DBImmutable Audit Log
MESSAGING TIER04

4. Asynchronous Event Streams

Heavy operations (notifications, audit logging, analytics, ML training, fan-out delivery) are decoupled into durable event logs like Kafka and SQS. This prevents user-facing requests from blocking on slow external networks.

Card Network / Bank
📖 SYSTEM DESIGN WHITE PAPERS & LOW-LEVEL SPECIFICATIONS

Study Stripe's database schemas, capacity math & production contracts

Beyond the visual blueprint, explore the exhaustive 7-section engineering whitepaper with real DDL schemas, API endpoints, failure mitigation matrices, and 45-minute FAANG interview scripts.

🏆 #1 HARDEST SYSTEM CHALLENGE

Zero Double-Charges & Distributed Financial Idempotency

⚠️The Engineering Bottleneck

In distributed internet payments, network connections drop, merchant servers timeout, and users click "Pay" multiple times. Charging a customer twice causes catastrophic fraud, customer fury, and merchant chargeback penalties.

💡The Winning Architectural Solution

Stripe requires client-generated idempotency keys on all mutating operations. The API gateway locks the key in an in-memory Redis cluster. If a retry arrives while processing, it waits; if it arrives after completion, Stripe immediately returns the cached original response without contacting the bank network.

SCALE & PRODUCTION METRICS:Zero duplicate charges across $1 Trillion+ lifetime volume, 99.9999% ledger accuracy.

⚖️ Architectural Trade-Offs & Decisions

Why the engineering team chose this specific stack over competing alternatives

Why Double-Entry Accounting instead of a simple balance column in MySQL?
CHOSEN:✓ Immutable Double-Entry Ledgervs Single balance column in MySQL, Eventual consistency DynamoDB table

Simple balance columns can be corrupted by race conditions, silent network partitions, or software bugs. In double-entry accounting, money cannot move without offsetting debit and credit journal entries. The sum of all accounts must balance to zero, providing automated mathematical proof of ledger integrity.

Why Webhooks with Exponential Backoff instead of Long Polling?
CHOSEN:✓ Webhooks with Exponential Backoff + Jittervs Long Polling, Server-Sent Events, Manual Merchant Polling

Card authorizations, 3D-Secure biometric challenges, and bank transfers can take seconds to hours to settle. Holding open HTTP connections is impossible across millions of merchants. Webhooks decouple processing and guarantee event delivery with exponential retry policies.

🚨 REAL-WORLD POST-MORTEM

The Webhook Thundering Herd Retry Storm

The Incident

A major e-commerce platform's servers crashed, returning 500 errors to Stripe webhooks. Stripe's automated retry engine began retrying millions of webhook events at identical fixed intervals, accidentally creating a self-inflicted DDoS on internal message queues.

Root Cause Analysis

Exponential backoff algorithms lacked randomized jitter (noise), causing millions of retried webhook requests to synchronize into periodic high-volume spikes.

How They Re-Architected It

Stripe introduced "Full Jitter" to all exponential backoff retry algorithms, ensuring retry attempts are spread uniformly across time windows, and implemented per-merchant fair-share rate limiters.

📋 Complete Microservice Specifications

Every service in the Stripe ecosystem with production tech stacks and failure impact

ComponentTier / LayerTech StackProduction FunctionStatus / Chaos
CustomerCLIENT
Stripe.jsElements
Consumer entering payment credentials on web checkout
Merchant ServerCLIENT
RubyPythonNodeGo
E-commerce merchant backend initiating charges via Stripe API
API GatewayGATEWAY
EnvoyGoRate Limiter
API gateway handling TLS termination, API key verification, and rate limiting
Payment ServiceSERVICE
GoRubySaga
Core orchestration engine coordinating fraud checks, card networks, and ledger
Radar Fraud MLSERVICE
PythonMLC++
Real-time machine learning engine evaluating fraud risk on every single transaction
Card Network / BankEXTERNAL
ISO 8583Card Networks
Visa, Mastercard, American Express, and acquiring banks
Webhook ServiceSERVICE
GoSQSRedis
Guaranteed webhook delivery engine dispatching payment events to merchant servers
Ledger ServiceDATABASE
PostgreSQLLedger Engine
Double-entry immutable accounting ledger where total debits always equal total credits
Payment DBDATABASE
PostgreSQLSpanner
Sharded transactional database storing immutable charge records and idempotency locks
Immutable Audit LogDATABASE
ClickHouseKafka
Append-only cryptographic event log for financial regulatory compliance