⚙️ How Big Tech Systems Handle Millions of Requests per Second

Architecture Lessons from Amazon, Flipkart, Google & Apple — Deep Dive into Scalability, Caching, Event-Driven Systems & Resilience

🔍 11,000+ words deep research⚡ Real-world case studies🏆 Production-proven patterns

📌 Core Pillars of Scalable Architecture

🌐 API Gateway → Load Balancer → Services → Data LayerEnd-to-end request flow in <50ms
📈 Horizontal Scaling & Auto-scaling Groups100k+ instances in 20 minutes
🗄️ Sharding + Replication10T+ requests daily
🚀 Caching: Redis, CDN, EdgeSub-millisecond latency
📡 Event-Driven: Kafka, Pub/Sub1M+ messages/second
🛡️ Failure Handling & Circuit Breakers99.999% availability
🔍 Observability: Tracing & Metrics1T+ spans daily
🔄

🌐 Request Lifecycle at Scale

API Gateway → Load Balancer → Services → Data Layer

When millions of requests hit Amazon, Flipkart, Google, or Apple every second, the journey begins at an API Gateway (like Amazon API Gateway, Envoy, or Zuul). The gateway handles authentication, rate limiting, and request routing. Next, a Layer 4/7 Load Balancer (e.g., AWS ELB, Google Cloud Load Bal...

2,000,000+requests/sec<50msp99 latency
⚙️

📦 Horizontal Scaling Strategies

Stateless services & Auto-scaling at Global Scale

Big Tech avoids vertical scaling (bigger machines) due to cost and physical limits. Instead, they add more instances — horizontal scaling. Amazon uses Auto Scaling Groups with dynamic policies based on CPU, memory, or custom metrics like request queue depth. Google's Borg/Kubernetes schedules pods a...

100,000+instances30 secondscold start
💾

🗃️ Data Partitioning: Sharding & Replication

DynamoDB, Cassandra, Spanner, CockroachDB

To handle massive write/read loads, databases must split data across nodes. Sharding (horizontal partitioning) distributes rows by a shard key (e.g., user_id, region, or timestamp). Amazon DynamoDB uses consistent hashing with virtual nodes to evenly distribute data across partitions, handling 10+ t...

10+ Trilliondaily requests100,000+QPS
🧠

⚡ Caching Layers: Redis, CDN, Edge Computing

Multi-tier caching for sub-millisecond responses

Caching is the silver bullet for latency. Browser/CDN (CloudFront, Fastly, Cloudflare) caches static assets with 500+ PoPs globally, serving 30% of all internet traffic. Edge caching (Cloudflare Workers, Lambda@Edge, Fly.io) runs logic closer to users, reducing latency by 80%. Next, in-memory caches...

100+ Millionops/second<1mslatency
📨

📡 Event-Driven Systems (Kafka, Pub/Sub)

Decoupling producers & consumers at Scale

When an order is placed on Amazon or a like on Instagram, the system emits events. Apache Kafka (used by LinkedIn, Uber, Apple, and Netflix) handles 1T+ messages/day with 1M+ messages/sec throughput per cluster. Kafka's partitioning enables parallel processing — each partition can handle 100MB/sec. ...

1+ Trillionmessages/day1,000,000+msg/sec
🔌

🛡️ Failure Handling: Timeouts, Retries, Circuit Breakers

Resilience patterns from Netflix, Amazon & Google

At scale, failures are inevitable. Amazon's architecture uses timeout budgets (50ms for internal calls), exponential backoff retries with jitter (base 100ms, max 10s), and circuit breakers (like Hystrix or Resilience4j). When a downstream service degrades, the circuit opens and fails fast for 30 sec...

99.999%availability30 secondsrecovery
📊

🔭 Observability: Logs, Metrics, Tracing

Prometheus, Jaeger, OpenTelemetry, Datadog

To debug millions of requests, you need deep visibility. Google's Dapper pioneered distributed tracing with 1 in 1000 sampling and 95% accuracy. Today, OpenTelemetry standardizes traces, metrics, and logs with automatic instrumentation. Amazon CloudWatch + X-Ray trace requests across services, proce...

1+ Trillionspans daily100+ Millionmetrics/sec

🔬 Deep Research & Real-World Examples

🏢

🏢 Amazon's Cell-Based Architecture

Amazon's 'cells' architecture limits blast radius to 1000 customers per cell. Each cell has independent API gateway, load balancers, services, and databases. Cells can fail independently without affecting others. This enables 99.999% availability with rapid deployment. Amazon runs 5000+ cells globally, each handling 100,000+ requests/second. Cell-based architecture reduces deployment risk and enables canary testing at scale.

🌍

🌍 Google's Global Load Balancing

Google's Andromeda network virtualization stack powers Maglev, their consistent hashing L4 load balancer. Maglev handles 10M+ RPS with 10-second configuration propagation. Their 'global VIP' routes users to nearest region using anycast IP and BGP routing. Google's network carries 40% of global internet traffic daily. Maglev's consistent hashing minimizes connection disruptions during backend changes.

🇮🇳

🇮🇳 Flipkart's Big Billion Day Prep

Flipkart conducts 100+ chaos experiments before Big Billion Day. They use in-memory data grids (Hazelcast) with 100GB+ cache, read replicas across 3 availability zones, and sharded Redis clusters with 95% hit rates. Their auto-scaling triggers 15 minutes before predicted spikes, scaling from 10k to 200k instances within 20 minutes. They simulate 3x peak traffic during load testing.

🍎

🍎 Apple's Privacy-First Observability

Apple's 'Differential Privacy' adds statistical noise to metrics while preserving insights. They leverage FoundationDB for distributed ACID transactions across 1000+ nodes. iCloud uses 'CloudCache' with 500+ edge locations and 95% hit rates. Apple processes 100B+ iCloud requests daily with <50ms latency. Privacy-preserving analytics enable feature improvement without user tracking.

📺

📺 Netflix's API Gateway Evolution

Netflix's Zuul API gateway handles 2M+ concurrent requests, 10B+ daily API calls. It evolved from Zuul 1 (sync) to Zuul 2 (async non-blocking), reducing latency by 40%. Zuul includes circuit breakers, rate limiting, request routing, and canary analysis for gradual rollouts. Zuul processes 1 trillion API calls monthly across 100+ regions.

🚗

🚗 Uber's Event-Driven Architecture

Uber's Kafka-based event bus processes 500M+ events daily for trip tracking, surge pricing, and driver matching. Their 'Chaperone' service validates 100% of events for schema compliance. Uber uses 100+ Kafka clusters with 10PB of storage, processing 1M+ events/second at peak. Event streaming powers real-time ETA calculations and fraud detection.

⛓️ Circuit Breaker Pattern in Action

Netflix Hystrix / Resilience4j — Used by Amazon, Flipkart, Google, Apple

CLOSED ✅ Requests passing normally
→ failure threshold (50%) →
OPEN ❌ Fails fast for 30s
← timeout (30s) ←
HALF-OPEN 🔄 Trial requests allowed

Amazon, Flipkart implement adaptive retries with jitter. Google's 'client-side backoff' reduces thundering herds. Apple uses exponential backoff for iCloud sync. Netflix's Chaos Monkey tests resilience daily.

📊 Throughput Legends

  • 🏢 Amazon: 1.5M+ RPS during Prime Day (API Gateway + DynamoDB)
  • 🌍 Google Search: ~5M RPS across global datacenters
  • 🇮🇳 Flipkart: 2M+ RPS at peak Big Billion Day
  • 🍎 Apple: 500k+ RPS for App Store & iCloud
  • 📺 Netflix: 2M+ concurrent streams, 10B+ API calls daily
  • 🚗 Uber: 500M+ events daily, 1M+ trips per day

🧠 Edge Caching Leaders

  • 🌐 CloudFront (Amazon) – 400+ PoPs globally, 30M+ RPS
  • 🔵 Google Cloud CDN – anycast routing, 130+ edge locations
  • 🍎 Apple's Edge Cache – powers iOS updates, 50TB/s throughput
  • 🟠 Flipkart's Dynamic CDN – product images cached at edge
  • ⚡ Cloudflare – 250+ cities, 30M+ RPS globally
  • 🚀 Fastly – 100+ PoPs, 1T+ requests daily
📡 Apache Kafka at Scale: 1 Trillion+ messages/day

Event-driven architecture powers real-time fraud detection at Flipkart, order tracking at Amazon, and Apple Maps traffic updates. Pub/Sub decouples producers from consumers, enabling asynchronous processing and resilience. LinkedIn processes 1T+ Kafka messages daily across 100+ clusters. Uber's event bus handles 500M+ events daily for real-time pricing and dispatch.

📈

Metrics (Prometheus + Grafana)

Flipkart monitors RED (Rate, Errors, Duration) for every microservice. Amazon CloudWatch custom metrics at 1-second granularity. 100M+ metrics/sec processed globally.

🔗

Distributed Tracing (Jaeger/Zipkin)

Google Dapper header propagation: X-Request-Id. Apple's internal tracing correlates billions of spans daily. Uber's Jaeger handles 1B+ spans per day with 1 in 1000 sampling.

📋

Centralized Logging (ELK/Loki)

Amazon OpenSearch Service processes petabyte-scale logs. Structured logging (JSON) enables fast debugging. Netflix's logging pipeline handles 10PB+ daily.

Explore More

The Death of Framework Loyalty

Apr 4, 2026

The Death of Framework Loyalty

READ →
AI in Disease Detection: Revolutionizing Healthcare

Apr 13, 2026

AI in Disease Detection: Revolutionizing Healthcare

READ →
Why CRUD Apps Are Dead

Apr 4, 2026

Why CRUD Apps Are Dead

READ →