Introduction
Deploying generative AI in production requires careful planning around GPU memory, compute throughput, and serving architecture. This guide covers the full spectrum — from sizing a single inference server for a 7B model to designing a multi-node cluster capable of serving GPT-4-class models at enterprise throughput.
The core constraint in LLM deployment is GPU VRAM. A 7B parameter model in FP16 precision requires approximately 14GB of VRAM just to load the weights — before accounting for the KV cache, activations, and batch overhead. A 70B model requires 140GB minimum, necessitating multi-GPU configurations.
Beyond raw memory, production serving requires attention to throughput, latency, and cost efficiency. A single A100 80GB GPU can serve a 7B model at roughly 1,000-2,000 tokens per second in batch mode, but achieving that throughput requires careful configuration of batching, KV cache sizing, and request scheduling.
Architecture overview
A production LLM serving stack has seven distinct layers, each with its own scaling characteristics and failure modes.
LLM Serving Architecture
Monitoring
Metrics, tracing, alerting
Model Weights Storage
Persistent weight storage and loading
GPU Memory Pool
VRAM allocation across model layers
KV Cache Manager
Attention key-value cache allocation
Inference API Gateway
Auth, throttling, request queuing
Load Balancer
Request routing and rate limiting
User Requests
REST / WebSocket / gRPC clients
Technical deep dive
GPU VRAM requirements scale linearly with parameter count. The table below shows minimum and recommended configurations for common model sizes.
LLM infrastructure requirements by model size
| Model Class | Parameters | Min VRAM | Recommended GPU | Nodes | Throughput (tok/s) | P50 Latency |
|---|---|---|---|---|---|---|
| LLaMA-7B | 7B | 14 GB | 1x A10G 24GB | 1 | 800–1,200 | 80 ms |
| LLaMA-70B | 70B | 140 GB | 2x A100 80GB | 1 | 400–600 | 200 ms |
| LLaMA-405B | 405B | 810 GB | 8x H100 80GB | 1–2 | 150–250 | 500 ms |
| GPT-4 class | ~1.8T MoE | ~1.6 TB | 16x H100 80GB | 2–4 | 80–120 | 800 ms |
| Gemini Ultra class | ~1T+ | ~2 TB | 32x H100 80GB | 4–8 | 60–100 | 1,200 ms |
The 2x memory rule
KV cache sizing
Implementation guide
Follow these steps to design and deploy a production LLM serving infrastructure.
- 1
Determine model size and precision
Select the model based on quality requirements. Calculate VRAM: parameters × 2 bytes (FP16) or × 1 byte (INT8). Add 25% headroom for KV cache and activations. This determines your minimum GPU configuration.
- 2
Choose serving framework
vLLM is the production standard for most use cases, offering PagedAttention for efficient KV cache management. Text Generation Inference (TGI) from Hugging Face is a strong alternative. NVIDIA Triton is preferred for multi-model serving environments.
- 3
Configure tensor parallelism
For models requiring multiple GPUs, configure tensor parallelism to split model layers across GPUs. vLLM supports this natively. Use NVLink-connected GPUs for tensor parallel configurations to minimize inter-GPU communication overhead.
- 4
Size the KV cache
Configure KV cache to use 80–90% of remaining VRAM after model weights are loaded. Monitor cache hit rates — a hit rate below 60% indicates insufficient cache size or poor request batching.
- 5
Configure dynamic batching
Enable continuous batching (iteration-level scheduling) in your serving framework. This allows new requests to join in-flight batches, improving GPU utilization from 20–30% to 70–85%.
- 6
Deploy load balancing and autoscaling
Place a load balancer in front of multiple inference replicas. Configure autoscaling based on GPU utilization and request queue depth. Target 70–80% GPU utilization for cost efficiency.
- 7
Implement observability
Instrument with Prometheus metrics: tokens per second, GPU utilization, KV cache hit rate, request queue depth, and P50/P95/P99 latency. Set alerts on queue depth (>100 requests) and GPU utilization (>90% sustained).
Business benefits and ROI
Cost reduction vs managed API
Self-hosted vs OpenAI at scale
Throughput improvement
With continuous batching enabled
GPU utilization achievable
With dynamic batching
Latency reduction
Speculative decoding benefit
LLM infrastructure cost calculator
Estimate GPU requirements and monthly costs for your LLM serving workload.
Estimated results
GPUs required
Monthly GPU cost
Cost per 1M tokens
OpenAI API equivalent
Monthly savings
Common mistakes
Undersizing KV cache
Ignoring tensor parallelism overhead
No request queuing under load
Vendor considerations
LLM serving framework comparison
| Framework | Best For | Batching | Multi-GPU | Quantization | Maturity |
|---|---|---|---|---|---|
| vLLM | High-throughput serving | Continuous | Tensor + Pipeline | GPTQ, AWQ, INT8 | Production |
| TGI (Hugging Face) | HF model ecosystem | Continuous | Tensor Parallel | GPTQ, AWQ | Production |
| NVIDIA Triton | Multi-model serving | Dynamic | Full support | TensorRT | Production |
| Ollama | Local / dev use | Sequential | Limited | GGUF | Development |
| LMDeploy | Efficiency-focused | Continuous | Tensor Parallel | W4A16, W8A8 | Production |
Reference architecture
Enterprise LLM Serving Reference Architecture
Model Registry
Versioned model weight storage
GPU Cluster
A100/H100 nodes with NVLink
Inference Fleet
Autoscaling vLLM / TGI replicas
Request Queue
Durable queue for traffic spike absorption
API Gateway + Auth
Rate limiting, authentication, request routing
Client Layer
Applications consuming the LLM API
Future trends
The LLM infrastructure landscape is evolving rapidly. Disaggregated prefill-decode architectures are emerging as the next major efficiency improvement, with early benchmarks showing 2-4x throughput gains.
Mixture-of-Experts (MoE) architectures are becoming dominant for frontier models, activating only a fraction of parameters per token. This changes the infrastructure calculus significantly.
Custom silicon (Google TPUs, AWS Trainium/Inferentia, Groq LPUs) is maturing rapidly and will challenge NVIDIA's dominance for specific workloads.
Frequently asked questions
How much GPU memory does an LLM require?
The minimum VRAM equals the parameter count multiplied by 2 (for FP16 precision). A 7B model needs 14 GB minimum, a 70B model needs 140 GB. Add 20–30% headroom for KV cache and activations. INT8 quantization halves these requirements with minimal quality loss.
How do you serve LLMs at scale?
Use a production serving framework like vLLM or TGI with continuous batching enabled. Place a load balancer in front of multiple inference replicas. Implement a request queue for traffic spike absorption. Autoscale based on GPU utilization and queue depth.
What is KV cache and why does it matter?
The KV (key-value) cache stores intermediate attention computations for tokens already processed in a sequence. Without it, every new token would require recomputing attention over the entire context. KV cache enables efficient autoregressive generation but consumes significant VRAM — often as much as the model weights themselves under load.
How do you reduce LLM inference costs?
The highest-impact optimizations are: INT8/INT4 quantization to reduce memory and increase throughput, continuous batching to improve GPU utilization, semantic caching to avoid redundant inference, and model routing to use smaller models for simpler queries. Combined, these can reduce costs by 70–90%.
What is speculative decoding?
Speculative decoding uses a small draft model to predict multiple tokens ahead, then verifies them in parallel with the large model. When predictions are correct (typically 60–80% of the time), multiple tokens are accepted in a single forward pass, reducing latency by 2–3x without changing output quality.