Products›🚗 Ride & Transport›Uber
🚗

Uber

Real-time ride matching at global planetary scale

99.99%
SLA
12
SERVICES
13
NODES
REQ/SEC0
LATENCY0ms
ERROR RATE0%
CACHE HIT95%
ACTIVE CONNS0
QUEUE DEPTH0

Uber Architecture Blueprint

Click ▶ RUN to animate active particle streams across microservices

client
gateway
service
database
cache
queue
cdn
storage
SYSTEM ARCHITECTURE WALKTHROUGH

How Traffic Flows Through Uber

EDGE TIER01

1. Ingress & Edge Routing

User requests arrive at the edge network. Global CDNs cache static assets and media. API Gateways terminate TLS, validate JWT authentication tokens, enforce token-bucket rate limits, and scrub malicious bot traffic before forwarding to internal services.

Rider AppDriver AppAPI GatewayLoad Balancer
APPLICATION TIER02

2. Microservice Processing

Stateless domain services execute core business logic. Microservices communicate via high-performance internal gRPC/REST APIs and autoscaling worker pods, ensuring that high load on one domain never exhausts compute resources of another.

Location ServiceMatching ServiceTrip ServicePricing / Surge
DATA TIER03

3. In-Memory Caching & Storage

Read-heavy traffic is served from in-memory Redis clusters with sub-millisecond latencies, protecting primary databases. Persistent databases (PostgreSQL, Cassandra, DynamoDB) maintain ACID consistency for financial ledgers, user accounts, and immutable state records.

Redis Geo CachePostgreSQL DB
MESSAGING TIER04

4. Asynchronous Event Streams

Heavy operations (notifications, audit logging, analytics, ML training, fan-out delivery) are decoupled into durable event logs like Kafka and SQS. This prevents user-facing requests from blocking on slow external networks.

Kafka Event Bus
📖 SYSTEM DESIGN WHITE PAPERS & LOW-LEVEL SPECIFICATIONS

Study Uber's database schemas, capacity math & production contracts

Beyond the visual blueprint, explore the exhaustive 7-section engineering whitepaper with real DDL schemas, API endpoints, failure mitigation matrices, and 45-minute FAANG interview scripts.

🏆 #1 HARDEST SYSTEM CHALLENGE

The Hexagonal Geospatial Spatial Index (Uber H3)

⚠️The Engineering Bottleneck

Traditional SQL spatial queries (`ST_DWithin`) compute Euclidean distances between millions of moving drivers every 4 seconds. At 5M+ active drivers, this creates catastrophic lock contention, quadratic O(N×M) CPU bottlenecks, and database crashes.

💡The Winning Architectural Solution

Uber engineered and open-sourced H3, dividing the globe into a hierarchical hexagonal grid. Every latitude/longitude maps directly to a 64-bit integer cell ID in O(1) time. Finding nearby drivers is simply querying the driver’s hexagon and its immediate 1-ring neighbor hexagons from in-memory Redis Sorted Sets in sub-millisecond time.

SCALE & PRODUCTION METRICS:Sub-15ms P99 matching latency, 250,000 GPS writes/second, 99.99% system availability.

⚖️ Architectural Trade-Offs & Decisions

Why the engineering team chose this specific stack over competing alternatives

Why Redis Sorted Sets instead of PostgreSQL PostGIS for driver coordinates?
CHOSEN:✓ Redis Sorted Sets with H3vs PostgreSQL PostGIS, MongoDB 2dsphere, Cassandra

Driver coordinates expire every 60 seconds and are updated every 4 seconds. Writing 250,000 updates/second to a relational database creates massive WAL disk I/O write amplification. Redis keeps data entirely in RAM, executes GEOADD in O(log N) time, and automatically garbage-collects disconnected drivers via TTLs.

Why Kafka instead of direct HTTP RPC between Location and Matching?
CHOSEN:✓ Apache Kafkavs RabbitMQ, AWS SQS, Direct gRPC calls

During massive traffic spikes (like New Year's Eve or sudden rainstorms), direct HTTP calls would overwhelm the matching engine with cascading 504 timeouts. Kafka acts as an elastic shock-absorber that buffers millions of events and enables multiple downstream microservices (Surge, Analytics, Heatmaps) to consume the same stream independently.

🚨 REAL-WORLD POST-MORTEM

The Monolithic PostgreSQL Connection Pool Exhaustion

The Incident

During a severe thunderstorm in New York City, ride requests surged 12x in 5 minutes. Uber's early monolithic PostgreSQL database was swamped by open connection requests, knocking down ride matching globally.

Root Cause Analysis

Each application server instance opened dedicated TCP connections to PostgreSQL. As autoscaling spawned hundreds of backend workers, connection pool limits were exceeded, locking the database in connection thrashing.

How They Re-Architected It

Uber decoupled real-time location storage into Redis, adopted connection multiplexing proxies (PgBouncer/Envoy), built Schemaless on top of MySQL shards, and broke the monolith into independently scaling microservices.

📋 Complete Microservice Specifications

Every service in the Uber ecosystem with production tech stacks and failure impact

ComponentTier / LayerTech StackProduction FunctionStatus / Chaos
Rider AppCLIENT
iOSAndroidWebSockets
iOS/Android app sending pickup requests and polling driver ETA
Driver AppCLIENT
iOSAndroidgRPC
Driver mobile app broadcasting GPS location every 4 seconds
API GatewayGATEWAY
KongOAuth2WAF
Kong/Nginx reverse proxy for SSL termination, JWT auth, and DDoS protection
Load BalancerLB
EnvoyALB
Layer 7 Envoy load balancer distributing traffic across healthy instances
Location ServiceSERVICE
GoUber H3Redis
Geospatial indexing engine using Uber H3 hexagonal spatial cells
Matching ServiceSERVICE
PythonC++ML
Two-phase driver matching engine: proximity filter followed by ML ranking
Trip ServiceSERVICE
JavaPostgreSQLSaga
Finite State Machine managing trip lifecycle (REQUESTED -> ACCEPTED -> ACTIVE -> COMPLETED)
Pricing / SurgeSERVICE
PythonSparkRay
Real-time dynamic pricing model calculating surge multipliers per H3 hexagon
Payment ServiceSERVICE
JavaStripeBraintree
Payment orchestration using idempotency keys to guarantee exactly-once charging
Notification ServiceSERVICE
GoAPNsFCM
Real-time push delivery service dispatching trip requests to driver apps
Redis Geo CacheCACHE
RedisSorted Sets
In-memory Redis Sorted Sets storing driver coordinates with 60s expiration
Kafka Event BusQUEUE
KafkaSchema Registry
Distributed event log decoupling trip events from analytics, surge, and ML pipelines
PostgreSQL DBDATABASE
PostgreSQLSchemaless
Sharded relational database storing immutable trip records, users, and audit journals