⚖️
Ingress & TrafficBeginner
Load Balancing & Elastic Distribution
Distributes incoming client requests across a pool of backend servers to prevent single-node bottlenecks and maximize throughput.
KEY TAKEAWAY: Layer 4 (TCP/UDP) operates on IP/Port; Layer 7 (HTTP) operates on URLs, headers, and cookies.
PRO:Eliminates single point of failure; enables horizontal autoscaling.
CON:Requires health checks and session persistence (sticky sessions) if stateful.
🛡️
Ingress & TrafficBeginner
Reverse Proxy (NGINX / Envoy)
Sits in front of internal servers to terminate TLS, compress responses (Gzip/Brotli), cache static assets, and shield internal IP addresses.
KEY TAKEAWAY: Reverse proxies protect internal network topology from public exposure.
PRO:Centralizes SSL certificates and DDoS scrubbing.
CON:Additional network hop adds 1-2ms to request processing.
🚪
Ingress & TrafficIntermediate
API Gateway Pattern
Unified entry point for microservices handling authentication (JWT), token bucket rate limiting, protocol translation (REST to gRPC), and request routing.
KEY TAKEAWAY: Prevents client mobile apps from making 20 individual microservice requests.
PRO:Centralized security, rate limiting, and observability.
CON:Can become a monolithic bottleneck if domain business logic leaks into it.
🌐
Ingress & TrafficBeginner
DNS Resolution & Anycast Routing
Translates human-readable domain names into IP addresses using hierarchical nameservers (Root -> TLD -> Authoritative). Anycast routes packets to the topologically closest PoP.
KEY TAKEAWAY: GeoDNS and Anycast route clients to geographically nearest data centers.
PRO:Globally distributed, heavily cached at ISP resolvers.
CON:TTL caching delays DNS propagation during emergency disaster failovers.
📈
Ingress & TrafficBeginner
Horizontal vs Vertical Scaling
Vertical scaling (Scale Up) upgrades CPU/RAM on a single box. Horizontal scaling (Scale Out) adds more stateless worker nodes behind a load balancer.
KEY TAKEAWAY: Stateless application servers scale horizontally indefinitely; databases require sharding.
PRO:Horizontal scaling has no hardware ceiling and offers linear cost efficiency.
CON:Requires stateless code and distributed session stores (Redis).
🔀
Data & StorageAdvanced
Database Sharding & Consistent Hashing
Partitions large database tables horizontally across independent physical servers using a shard key (e.g., hash(user_id) % N or Consistent Hash Ring).
KEY TAKEAWAY: Consistent hashing bounds key redistribution to K/N when nodes join or leave.
PRO:Breaks write throughput limits of single database master nodes.
CON:Cross-shard SQL joins and distributed transactions become prohibitively expensive.
👯
Data & StorageIntermediate
Database Replication (Leader-Follower)
Master node accepts write transactions and replicates Write-Ahead Logs (WAL) to multiple read replicas. Read replicas serve high-volume read traffic.
KEY TAKEAWAY: Async replication causes replication lag: clients might read stale data after writing.
PRO:Massively increases read QPS capacity and provides automated failover.
CON:Replication lag can violate read-your-own-writes consistency.
🌲
Data & StorageAdvanced
B-Tree vs LSM-Tree Storage Engines
B-Trees (PostgreSQL/MySQL) organize data into balanced tree pages on disk, optimizing for fast in-place reads. LSM-Trees (Cassandra/RocksDB) append writes sequentially to MemTable and SSTables, maximizing write throughput.
KEY TAKEAWAY: LSM-Trees eliminate random disk writes by buffering writes in memory and flushing sequentially.
PRO:LSM has 10x higher write throughput; B-Tree has predictable sub-millisecond point reads.
CON:LSM background compaction consumes background CPU and disk I/O.
🏛️
Data & StorageAdvanced
ACID Transactions & Isolation Levels
Guarantees Atomicity, Consistency, Isolation, and Durability. Isolation levels (Read Committed -> Repeatable Read -> Serializable) trade concurrency for phantom read protection.
KEY TAKEAWAY: Serializable isolation eliminates race conditions via 2-Phase Locking or SSI.
PRO:Guarantees financial ledger integrity and prevents double-spending.
CON:Serializable locking introduces lock contention and degrades write throughput.
⚡
Caching & MemoryBeginner
Cache-Aside Pattern (Lazy Loading)
Application code queries the cache first. On cache hit, data returns in <2ms. On cache miss, application queries database and writes data into cache for future requests.
KEY TAKEAWAY: Only requested data is cached, keeping memory footprints compact.
PRO:Cache node failure does not crash the system (graceful database fallback).
CON:Cache miss latency penalty on cold start; potential stale data if TTL is long.
📝
Caching & MemoryIntermediate
Write-Through vs Write-Behind Caching
Write-Through writes to cache and DB simultaneously. Write-Behind (Write-Back) writes immediately to cache and acknowledges the client, flushing asynchronously to DB in batches.
KEY TAKEAWAY: Write-Behind delivers lightning-fast write latency but risks data loss if cache crashes.
PRO:Write-Behind absorbs devastating write spikes (e.g. video view counters).
CON:Risk of data loss on power failure before asynchronous flush commits.
🦬
Caching & MemoryAdvanced
Cache Stampede & Thundering Herd Defense
When a popular cached key expires, thousands of concurrent requests miss simultaneously and flood the primary database, crashing storage servers.
KEY TAKEAWAY: Defend using probabilistic early expiration (XFetch algorithm) or distributed mutex locks.
PRO:Guarantees only one single worker thread rebuilds the cache key.
CON:Requires distributed locking primitives (Redlock) or algorithm tuning.
🌸
Caching & MemoryIntermediate
Bloom Filters (Probabilistic Sets)
A space-efficient bit array that tests whether an element is definitely NOT in a set or possibly in a set in O(k) time without disk I/O.
KEY TAKEAWAY: Zero false negatives: if Bloom filter returns false, key definitely does not exist on disk.
PRO:Prevents expensive disk seeks for non-existent database keys.
CON:Elements cannot be deleted from standard bloom filters; small false positive rate.
🔌
Resilience & StabilityIntermediate
Circuit Breaker Pattern
Monitors downstream service failure rates. If errors exceed a threshold (e.g. 50%), the breaker trips to OPEN, instantly short-circuiting calls and preventing thread pool exhaustion.
KEY TAKEAWAY: Transitions: CLOSED (normal) -> OPEN (fast fail) -> HALF-OPEN (probe recovery).
PRO:Prevents cascading failures from bringing down the entire microservice ecosystem.
CON:Requires fallback degraded responses (e.g. cached static data).
🚦
Resilience & StabilityIntermediate
Rate Limiting (Token Bucket / Sliding Window)
Controls request velocity per IP/User to prevent DDoS attacks, credential stuffing, and noisy-neighbor compute starvation. Drops excess calls with HTTP 429.
KEY TAKEAWAY: Token Bucket allows bursts; Leaky Bucket guarantees smooth outbound rate.
PRO:Protects backend infrastructure from being overwhelmed.
CON:Requires centralized Redis state to enforce limits across distributed gateway nodes.
🔁
Resilience & StabilityIntermediate
Distributed Idempotency Keys
Clients attach a unique UUID in the Idempotency-Key header. Servers lock the key in Redis and cache the completed response, guaranteeing zero duplicate payments on retry.
KEY TAKEAWAY: Making mutating POST requests idempotent enables safe automatic client retries.
PRO:Eliminates double-charges in e-commerce and financial payment systems.
CON:Requires distributed lock management and response payload caching with TTLs.
🚢
Resilience & StabilityIntermediate
Bulkhead Isolation Pattern
Isolates resources (thread pools, connection pools, memory) into distinct compartments. An outage in an analytics service cannot exhaust threads allocated for user checkout.
KEY TAKEAWAY: Named after naval ship bulkheads: a leak in one hull compartment does not sink the ship.
PRO:Guarantees critical revenue-generating paths remain operational during partial outages.
CON:Can lead to resource under-utilization if thread pools are sized rigidly.
📜
Transactions & MessagingAdvanced
Saga Pattern (Orchestration vs Choreography)
Manages distributed transactions across microservices without 2-Phase Locking. Each service executes a local transaction; if a step fails, compensating transactions undo earlier changes.
KEY TAKEAWAY: Compensating transactions reverse partial work when distributed operations fail.
PRO:Supports long-running business workflows across decoupled microservices.
CON:Eventual consistency: intermediate states are visible to other queries.
🤝
Transactions & MessagingAdvanced
Two-Phase Commit (2PC)
Atomic consensus across multiple database nodes. Phase 1 (Prepare): coordinator asks all nodes if they can commit. Phase 2 (Commit): coordinator issues final commit or rollback.
KEY TAKEAWAY: Guarantees strict ACID across shards, but coordinator failure blocks all nodes.
PRO:Guarantees strong atomic consistency across distributed databases.
CON:Blocking protocol: holding locks across network roundtrips severely hurts throughput.
📋
Transactions & MessagingAdvanced
Event Sourcing & CQRS Pattern
Stores an immutable log of events instead of current state. CQRS separates Command (mutating write models) from Query (read-optimized materialized views).
KEY TAKEAWAY: Current state is calculated by replaying past events from genesis.
PRO:Complete auditable history; zero data loss; separate read/write scaling.
CON:High architectural complexity; eventual consistency between write and read models.
📮
Transactions & MessagingAdvanced
Transactional Outbox & CDC (Debezium)
Guarantees database update and message publishing commit together. Writes event to an `outbox` table in the same DB transaction; Debezium tails the WAL to publish to Kafka.
KEY TAKEAWAY: Eliminates dual-write race conditions between databases and message brokers.
PRO:Guaranteed at-least-once message delivery without distributed two-phase commit.
CON:Requires message consumers to implement idempotency handling.
🔌
Real-Time & ProtocolsBeginner
WebSockets vs Server-Sent Events (SSE)
WebSockets provide full-duplex bidirectional TCP communication. SSE provides lightweight unidirectional streaming from server to browser over standard HTTP/2.
KEY TAKEAWAY: Use WebSockets for interactive chat; use SSE for AI token generation and stock feeds.
PRO:SSE works over HTTP/2, supports auto-reconnect, and traverses corporate proxies easily.
CON:WebSockets require custom connection management and load balancer sticky sessions.
🕸️
Real-Time & ProtocolsIntermediate
GraphQL Federation & Gateway Mesh
Combines multiple microservice GraphQL schemas into a single unified supergraph. Apollo Gateway coordinates sub-queries and resolves entities across independent services.
KEY TAKEAWAY: Clients request exact fields needed in 1 roundtrip, eliminating over/under-fetching.
PRO:Frontends query one unified graph endpoint across 50+ backend microservices.
CON:Complex N+1 query performance traps if DataLoader batching is not configured.
🤖
GenAI & LLM SystemsIntermediate
RAG Pipeline & Hybrid Vector Search
Retrieval-Augmented Generation enriches prompts with private data. Text chunks are embedded into dense vectors, searched via HNSW Cosine Similarity, and injected into LLM context.
KEY TAKEAWAY: Hybrid Search combines dense vector semantic search with sparse BM25 keyword matching.
PRO:Eliminates LLM hallucinations and provides grounded citations without fine-tuning.
CON:Chunking strategies, embedding latency, and vector database operational overhead.
🧠
GenAI & LLM SystemsAdvanced
PagedAttention & KV-Cache Virtual Paging (vLLM)
Manages Transformer Key-Value activations in GPU VRAM like OS virtual memory pages. Allocates non-contiguous memory blocks on-demand, cutting VRAM waste from 70% to <4%.
KEY TAKEAWAY: PagedAttention increases LLM serving concurrency and throughput by up to 4x.
PRO:Massive serving cost reduction and support for long context windows.
CON:Requires custom CUDA kernels and complex memory block manager orchestration.
🛡️
GenAI & LLM SystemsIntermediate
AI Gateway (Semantic Caching & Guardrails)
Reverse proxy for LLM endpoints providing semantic caching (sub-20ms cache hit for similar queries), automatic model fallback, rate limiting, and prompt injection filters.
KEY TAKEAWAY: Semantic caching matches prompt embeddings using cosine similarity thresholds (>0.92).
PRO:Reduces LLM API costs by up to 60% and slashes Time To First Token (TTFT).
CON:Slight risk of serving obsolete cached answers for time-sensitive prompts.
🔄
GenAI & LLM SystemsAdvanced
Agent ReAct Loop (Thought -> Action -> Observation)
Autonomous AI agent execution pattern combining reasoning traces and action calls. Agents iteratively reason, invoke tools (web search, SQL, Python interpreter), and observe outputs.
KEY TAKEAWAY: ReAct loops enable LLMs to interact dynamically with external APIs and databases.
PRO:Solves complex multi-step reasoning problems beyond single-shot prompt capabilities.
CON:Can enter infinite loops or make expensive recursive tool calls without max-iteration bounds.