All 48 articles, sorted alphabetically
System Design: Rate Limiting — Token Bucket Algorithm — Belgavi.AI Lab
A comprehensive guide to rate limiting: token bucket mechanics, distributed Redis-backed implementations, sliding window and leaky bucket algorithms, …
Read article →System Design: Load Balancing - AICassindra
Deep dive into load balancing: algorithms (round-robin, least connections, IP hash, consistent hashing), L4 vs L7 decisions, health checks, session st…
Read article →Airbnb Smart Pricing Architecture
The ML system that suggests nightly prices to millions of Airbnb hosts. Covers feature extraction from location and history, demand modeling with tree…
Read article →Amazon Shopping Cart Architecture, in depth: the always-writable cart, replication and merge semantics, guest carts, price revalidation and the checkout handoff
How to design an Amazon-scale shopping cart from first principles: why the cart must accept writes during failures, what the public Dynamo paper says …
Read article →Designing a Real-Time Chat System (1M concurrent users)
End-to-end architecture for a WhatsApp-scale chat system, from connection layer to message storage.
Read article →Designing an Event-Driven Order System, in depth: state machine, event contracts, reservations, payments and failure walk-throughs
An end-to-end design for an event-driven e-commerce order system: the order state machine, event envelopes, per-order ordering by key, inventory reser…
Read article →Designing a Feature Flag System, in depth: requirements, data model, evaluation semantics, propagation and failure modes
A system design walk-through of a feature flag platform: requirements and capacity estimates, the flag data model and versioning, a precise evaluation…
Read article →Designing a File Upload Service (S3-style)
Designing an upload service: control plane versus data plane, multipart and resumable uploads, integrity checks, content-addressed dedup, and lifecycl…
Read article →Designing a Metrics Aggregation System
Designing a metrics platform: series cardinality, push versus pull ingest, sharding by series, the TSDB write path, rollups, and per-tenant limits.
Read article →Designing a Payment System, in depth: a marketplace pay-in and pay-out design from requirements to reconciliation
An end-to-end payment system design for a marketplace: requirements and capacity, the data model, a guarded payment state machine, safe PSP calls, pay…
Read article →Designing a Quota System, in depth: allocation, rate and budget quotas, tenant hierarchies, leases and reconciliation
How to design a multi-tenant quota system: the difference between rate limits and quotas, the tenant hierarchy and data model, transactional reservati…
Read article →Designing Search at Scale, in depth: sizing shards and replicas, taming fan-out tail latency, fresh indexing and zero-downtime reindexing
How to take a search system from one cluster to production scale: the three independent scale axes, sizing shards from measured bytes per document, si…
Read article →Designing Uber Dispatch System
Ride dispatch as an assignment problem: location ingestion, supply indexing, batched matching, offer state machines, idempotent dispatch and regional …
Read article →Email Delivery Architecture, in depth: SMTP relay, SPF, DKIM and DMARC alignment, retry queues, bounces, and staying out of the spam folder
How a production email sending system works end to end: submission and relay over SMTP, envelope versus header sender, SPF, DKIM and DMARC with alignm…
Read article →Google Drive + Docs Real-Time Collaboration Architecture, in depth: the ordering point, the op log, sessions, permissions and offline edits
A first-principles reference architecture for Google Docs-style real-time collaboration on top of a Drive-style file service: central-server operation…
Read article →Meta Threads Architecture
How Meta launched Threads in 5 days by leveraging Instagram&am…
Read article →Rate Limiter Architecture, in depth: GCRA in one Redis key, composite limits, local token leasing, failure policy and client signalling
How to build a production rate limiter rather than pick an algorithm: the decision contract, GCRA implemented as an atomic Redis script with server ti…
Read article →API gateway architecture
Deep-dive on API gateway design: data plane vs control plane, route matching and filter chains, JWKS-cached auth, distributed rate limiting, retry bud…
Read article →Backpressure Architecture in Depth: Bounded Queues, Credit-Based Flow Control and Propagating Overload Upstream
How backpressure keeps systems stable when producers outrun consumers: bounded queues and Little&a…
Read article →Content Delivery Network Architecture in Depth
A 2500-word walkthrough of CDN architecture: client, edge PoP, origin shield, origin, cache rules, purge, optimization, TLS, edge compute, analytics, …
Read article →CQRS architecture
Deep-dive on CQRS (Command Query Responsibility Segregation): splitting the write model that validates commands and mutates a normalized source of tru…
Read article →Dead-letter queue architecture
Deep-dive on the dead-letter queue: the delivery counter and retry policy, broker-level vs application-level dead-lettering, the failure-metadata cont…
Read article →Distributed lock architecture
Deep-dive on distributed locks: consensus-backed lock services (etcd, ZooKeeper), lease TTLs and sessions, fencing tokens checked at the resource, wai…
Read article →Event sourcing architecture
Deep-dive on event sourcing: commands and aggregates that emit events, the append-only event store as source of truth, rehydrating state by replay, sn…
Read article →Geo-distributed systems -- serving the world with low latency
Deep-dive on geo-distributed systems: the latency/availability/law drivers, multi-region replicas, data placement, the speed-of-light consistency-vs-l…
Read article →Gossip protocol architecture - epidemic membership, failure detection, and anti-entropy
Deep-dive on gossip protocols: the SYN/ACK/ACK2 digest exchange and version merge, phi-accrual failure detection and suspicion lifecycle, anti-entropy…
Read article →Hot-key mitigation architecture
Deep-dive on surviving hot keys in sharded systems: why hashing concentrates load, approximate hotspot detection, edge caching for read-hot keys, key …
Read article →Idempotency architecture
Deep-dive on idempotency architecture: keys, dedupe store, retry policy, conflict handling, cross-service propagation, TTL, and audit.
Read article →Leader election architecture
Deep-dive on leader election: quorum-based campaigning, terms/epochs, leases and heartbeats for bounded failover, fencing tokens that neutralize a sta…
Read article →Load shedding -- dropping work to survive overload
Deep-dive on load shedding: the overload-collapse problem, rejecting excess load, serve-some-well-rather-than-all-badly, fail fast (early rejection), …
Read article →Notification System Architecture in Depth
A 2500-word walkthrough of a notification system: producers, ingest queue, notification service, preferences, dedup + batching, channels, delivery tra…
Read article →Transactional outbox architecture
Deep-dive on the transactional outbox pattern: why commit-then-publish dual writes lose or fabricate events, writing the event to an outbox table in t…
Read article →Pub/sub system design architecture
Deep-dive on pub/sub design: topics, partitions, consumers, delivery guarantees, retention, DLQ, schema registry, and metrics.
Read article →Distributed Rate Limiter Architecture in Depth
A 2500-word walkthrough of a production distributed rate limiter: edge, gateway, local token bucket, global sliding window, Redis + Lua, fallback, and…
Read article →Rate limiting architecture
Deep-dive on rate-limiting architecture: local buckets + shared counters + sliding-window smoothing + policy pipeline for safe changes.
Read article →Read-replica routing architecture
Deep-dive on routing reads to replicas safely: read classification, LSN/GTID write watermarks, session stores, byte vs time lag, replica ejection, pri…
Read article →Saga pattern architecture
Deep-dive on the saga pattern: orchestration vs choreography, the durable saga log, compensating transactions and pivot placement, idempotency keys an…
Read article →Scaling Patterns in Depth: Finding the Bottleneck, Doing the Capacity Math, and Choosing the Next Move
A bottleneck-first method for scaling a system: name the saturated resource, size it with Little&a…
Read article →Search System Architecture in Depth
A 2500-word walkthrough of search architecture: docs, indexer, inverted index, query, hybrid retrieval, reranking, personalization, analytics, shardin…
Read article →Service discovery architecture
Deep-dive on service discovery: the health-checked registry, client-side vs server-side discovery, DNS and mesh-sidecar packagings, watches and cachin…
Read article →Database sharding architecture
Deep-dive on database sharding: choosing the shard key, range versus hash partitioning, the routing layer and shard directory, online rebalancing, per…
Read article →Signal Protocol Architecture, in depth: PQXDH key agreement, the Double Ratchet, and a server that learns almost nothing
How the Signal Protocol delivers end-to-end encryption to devices that are usually offline: the system architecture and what the server stores, identi…
Read article →Backend for Frontend (BFF), in depth: per-client backends, aggregation, token handling and when not to build one
The Backend for Frontend pattern from first principles: why one general-purpose API serves many clients badly, how a BFF differs from an API gateway a…
Read article →Caching
An orienting map of caching: the tiers a byte can live in, cache-aside versus write-through, why invalidation deletes rather than updates, the fill ra…
Read article →Load Balancing
How load balancers distribute traffic across backend servers: L4 vs L7, round-robin vs least-connections, and health-check design.
Read article →Message Queues
The main message queue types (log-based Kafka vs broker-based RabbitMQ vs managed SQS), semantics (at-least-once vs at-most-once vs exactly-once), and…
Read article →Twitter Timeline Generation Architecture, in depth: the Following cache, the For You ranking pipeline, and how both are served, repaired and kept fast
How Twitter-style home timelines are generated: the precomputed Following timeline cache and its fan-out workers, read-time merging and visibility fil…
Read article →Vector Search at Scale, in depth: segments, capacity math, sharding, filtering, re-embedding and recall in production
How to design a vector search service for hundreds of millions of embeddings: separate write and read paths, growing and sealed segments, worked memor…
Read article →