Products›🤖 AI Products›ChatGPT / AI Assistant
🤖

ChatGPT / AI Assistant

LLM inference at global scale — token streaming & RAG architecture

99.9%
SLA
12
SERVICES
12
NODES
REQ/SEC0
LATENCY0ms
ERROR RATE0%
CACHE HIT95%
ACTIVE CONNS0
QUEUE DEPTH0

ChatGPT / AI Assistant Architecture Blueprint

Click ▶ RUN to animate active particle streams across microservices

client
gateway
service
database
cache
queue
cdn
storage
SYSTEM ARCHITECTURE WALKTHROUGH

How Traffic Flows Through ChatGPT / AI Assistant

EDGE TIER01

1. Ingress & Edge Routing

User requests arrive at the edge network. Global CDNs cache static assets and media. API Gateways terminate TLS, validate JWT authentication tokens, enforce token-bucket rate limits, and scrub malicious bot traffic before forwarding to internal services.

User ClientAPI GatewayAI Routing Gateway
APPLICATION TIER02

2. Microservice Processing

Stateless domain services execute core business logic. Microservices communicate via high-performance internal gRPC/REST APIs and autoscaling worker pods, ensuring that high load on one domain never exhausts compute resources of another.

Context & History ServiceInference DispatcherGPU Inference ClusterTool & Sandbox Executor
DATA TIER03

3. In-Memory Caching & Storage

Read-heavy traffic is served from in-memory Redis clusters with sub-millisecond latencies, protecting primary databases. Persistent databases (PostgreSQL, Cassandra, DynamoDB) maintain ACID consistency for financial ledgers, user accounts, and immutable state records.

Vector DB (RAG)Semantic CacheConversation DB
MESSAGING TIER04

4. Asynchronous Event Streams

Heavy operations (notifications, audit logging, analytics, ML training, fan-out delivery) are decoupled into durable event logs like Kafka and SQS. This prevents user-facing requests from blocking on slow external networks.

📖 SYSTEM DESIGN WHITE PAPERS & LOW-LEVEL SPECIFICATIONS

Study ChatGPT / AI Assistant's database schemas, capacity math & production contracts

Beyond the visual blueprint, explore the exhaustive 7-section engineering whitepaper with real DDL schemas, API endpoints, failure mitigation matrices, and 45-minute FAANG interview scripts.

🏆 #1 HARDEST SYSTEM CHALLENGE

PagedAttention & GPU Memory Management (vLLM)

⚠️The Engineering Bottleneck

During LLM generation, Key-Value (KV) cache grows unpredictably as tokens are generated. Traditional systems allocate contiguous memory blocks, resulting in 60-80% of scarce GPU High Bandwidth Memory (HBM) being wasted due to memory fragmentation.

💡The Winning Architectural Solution

OpenAI and vLLM developed PagedAttention, which manages the KV cache like virtual memory in an operating system. KV cache blocks are stored in non-contiguous physical memory pages. This allows almost 100% of GPU memory to be utilized, boosting concurrent request throughput by up to 24x per server.

SCALE & PRODUCTION METRICS:Sub-20ms Time-To-First-Token, 24x higher throughput per GPU, 200M+ active users.

⚖️ Architectural Trade-Offs & Decisions

Why the engineering team chose this specific stack over competing alternatives

Why Server-Sent Events (SSE) instead of WebSockets for token streaming?
CHOSEN:✓ Server-Sent Events (SSE)vs WebSockets, gRPC-Web, HTTP Polling

WebSockets are bidirectional and require custom protocol framing and connection management. LLM completions are strictly unidirectional (server pushes tokens to client). SSE runs over standard HTTP/2, automatically supports reconnection, passes through corporate firewalls, and works seamlessly with standard CDN caching proxies.

Why Semantic Vector Caching instead of exact-string Redis cache?
CHOSEN:✓ Semantic Vector Cachingvs Exact-String Redis Cache, No Caching

Users rarely type identical prompts ("How does Uber matching work?" vs "Explain Uber's ride matching algorithm"). Exact-string matching yields <2% cache hit rates. Embedding queries into vector embeddings and checking cosine similarity allows caching semantically identical queries, saving thousands of GPU compute hours.

🚨 REAL-WORLD POST-MORTEM

The Redis Cluster Cache Leak Incident (March 2023)

The Incident

A brief bug in ChatGPT's conversation caching layer allowed some users to see titles of active chat histories belonging to other users.

Root Cause Analysis

An asynchronous Redis client library bug in the Python backend caused connection pool cancellations to misalign request/response pairs under high concurrency, returning another user's cached session data.

How They Re-Architected It

The backend was patched with strict request-scoped connection isolation, extensive automated concurrency fuzz testing was implemented, and customer data encryption keys were strictly bound to session tokens.

📋 Complete Microservice Specifications

Every service in the ChatGPT / AI Assistant ecosystem with production tech stacks and failure impact

ComponentTier / LayerTech StackProduction FunctionStatus / Chaos
User ClientCLIENT
ReactMobileAPI
Web browser, mobile app, and API consumers submitting prompts and receiving tokens
API GatewayGATEWAY
KongGoRate Limiter
Gateway managing JWT auth, user tiers (Free vs Plus), and token rate limits
AI Routing GatewayGATEWAY
GoCustom
Intelligent router selecting optimal model (GPT-4o, o1, mini) based on prompt complexity
Context & History ServiceSERVICE
PythonRedisTiktoken
Retrieves conversation memory, truncates sliding windows, and counts tokens
Vector DB (RAG)DATABASE
PineconeMilvusQdrant
High-dimensional vector database for semantic search and document retrieval
Inference DispatcherSERVICE
GogRPCvLLM
Load balances prompts across distributed GPU worker nodes running vLLM
GPU Inference ClusterSERVICE
NVIDIA H100CUDAvLLM
Fleet of NVIDIA H100 GPUs executing transformer forward passes with PagedAttention
Tool & Sandbox ExecutorSERVICE
PythonDockergVisor
Isolated Docker sandbox executing web searches, code execution, and third-party APIs
SSE Streaming EngineSERVICE
GoNettySSE
Server-Sent Events server streaming generated tokens to the client connection
Semantic CacheCACHE
RedisVector Index
Caches embeddings and responses for identical or semantically similar queries
Conversation DBDATABASE
PostgreSQL
PostgreSQL database storing conversation histories, user settings, and billing logs
Safety & Guardrail FilterSERVICE
PythonTriton
Moderation model screening inputs for prompt injections, malware, and harmful content