Introduction

Deploying generative AI in production requires careful planning around GPU memory, compute throughput, and serving architecture. This guide covers the full spectrum — from sizing a single inference server for a 7B model to designing a multi-node cluster capable of serving GPT-4-class models at enterprise throughput.

The core constraint in LLM deployment is GPU VRAM. A 7B parameter model in FP16 precision requires approximately 14GB of VRAM just to load the weights — before accounting for the KV cache, activations, and batch overhead. A 70B model requires 140GB minimum, necessitating multi-GPU configurations.

Beyond raw memory, production serving requires attention to throughput, latency, and cost efficiency. A single A100 80GB GPU can serve a 7B model at roughly 1,000-2,000 tokens per second in batch mode, but achieving that throughput requires careful configuration of batching, KV cache sizing, and request scheduling.

Architecture overview

A production LLM serving stack has seven distinct layers, each with its own scaling characteristics and failure modes.

LLM Serving Architecture

Monitoring

Metrics, tracing, alerting

PrometheusGrafanaOpenTelemetry

Model Weights Storage

Persistent weight storage and loading

NVMe SSDObject StorageNFS

GPU Memory Pool

VRAM allocation across model layers

Tensor ParallelPipeline Parallel

KV Cache Manager

Attention key-value cache allocation

PagedAttentionRadixAttention

Inference API Gateway

Auth, throttling, request queuing

vLLMTGITriton

Load Balancer

Request routing and rate limiting

NGINXEnvoyAWS ALB

User Requests

REST / WebSocket / gRPC clients

Web AppsMobileAPI Consumers
Stack layers — top to bottom: highest to lowest abstraction

Technical deep dive

GPU VRAM requirements scale linearly with parameter count. The table below shows minimum and recommended configurations for common model sizes.

LLM infrastructure requirements by model size

Model ClassParametersMin VRAMRecommended GPUNodesThroughput (tok/s)P50 Latency
LLaMA-7B7B14 GB1x A10G 24GB1800–1,20080 ms
LLaMA-70B70B140 GB2x A100 80GB1400–600200 ms
LLaMA-405B405B810 GB8x H100 80GB1–2150–250500 ms
GPT-4 class~1.8T MoE~1.6 TB16x H100 80GB2–480–120800 ms
Gemini Ultra class~1T+~2 TB32x H100 80GB4–860–1001,200 ms

The 2x memory rule

Multiply the model parameter count by 2 to get the minimum VRAM in gigabytes for FP16 inference. A 13B model needs 26 GB minimum. Add 20–30% headroom for KV cache and activations in production.

KV cache sizing

The KV cache grows with sequence length and batch size. For a 7B model with 4K context and batch size 32, the KV cache alone can consume 8–16 GB of VRAM. Size your GPU memory budget with KV cache as a first-class concern, not an afterthought.

Implementation guide

Follow these steps to design and deploy a production LLM serving infrastructure.

  1. 1

    Determine model size and precision

    Select the model based on quality requirements. Calculate VRAM: parameters × 2 bytes (FP16) or × 1 byte (INT8). Add 25% headroom for KV cache and activations. This determines your minimum GPU configuration.

  2. 2

    Choose serving framework

    vLLM is the production standard for most use cases, offering PagedAttention for efficient KV cache management. Text Generation Inference (TGI) from Hugging Face is a strong alternative. NVIDIA Triton is preferred for multi-model serving environments.

  3. 3

    Configure tensor parallelism

    For models requiring multiple GPUs, configure tensor parallelism to split model layers across GPUs. vLLM supports this natively. Use NVLink-connected GPUs for tensor parallel configurations to minimize inter-GPU communication overhead.

  4. 4

    Size the KV cache

    Configure KV cache to use 80–90% of remaining VRAM after model weights are loaded. Monitor cache hit rates — a hit rate below 60% indicates insufficient cache size or poor request batching.

  5. 5

    Configure dynamic batching

    Enable continuous batching (iteration-level scheduling) in your serving framework. This allows new requests to join in-flight batches, improving GPU utilization from 20–30% to 70–85%.

  6. 6

    Deploy load balancing and autoscaling

    Place a load balancer in front of multiple inference replicas. Configure autoscaling based on GPU utilization and request queue depth. Target 70–80% GPU utilization for cost efficiency.

  7. 7

    Implement observability

    Instrument with Prometheus metrics: tokens per second, GPU utilization, KV cache hit rate, request queue depth, and P50/P95/P99 latency. Set alerts on queue depth (>100 requests) and GPU utilization (>90% sustained).

Business benefits and ROI

70%

Cost reduction vs managed API

Self-hosted vs OpenAI at scale

3x

Throughput improvement

With continuous batching enabled

80%

GPU utilization achievable

With dynamic batching

2–3x

Latency reduction

Speculative decoding benefit

LLM infrastructure cost calculator

Estimate GPU requirements and monthly costs for your LLM serving workload.

10,000 users
1001,000,000
1,000 tokens
1004,000
500 ms
1002,000
8 $/hr
230

Estimated results

1

GPUs required

$5,760

Monthly GPU cost

$6.40

Cost per 1M tokens

$13,500

OpenAI API equivalent

$7,740

Monthly savings

Common mistakes

Undersizing KV cache

The most common mistake is allocating insufficient VRAM for the KV cache. Teams calculate model weight memory correctly but forget that the KV cache can consume as much VRAM as the model itself under load. Always benchmark with realistic batch sizes and sequence lengths before finalizing GPU selection.

Ignoring tensor parallelism overhead

Splitting a model across GPUs with tensor parallelism introduces communication overhead. For models that fit on a single GPU, tensor parallelism often reduces throughput. Only use tensor parallelism when the model genuinely does not fit on a single GPU.

No request queuing under load

Without a proper request queue, traffic spikes cause inference servers to OOM or return errors. Implement a durable queue (Redis, RabbitMQ) between the API gateway and inference workers. Set queue depth alerts and autoscaling triggers before going to production.

Vendor considerations

LLM serving framework comparison

FrameworkBest ForBatchingMulti-GPUQuantizationMaturity
vLLMHigh-throughput servingContinuousTensor + PipelineGPTQ, AWQ, INT8Production
TGI (Hugging Face)HF model ecosystemContinuousTensor ParallelGPTQ, AWQProduction
NVIDIA TritonMulti-model servingDynamicFull supportTensorRTProduction
OllamaLocal / dev useSequentialLimitedGGUFDevelopment
LMDeployEfficiency-focusedContinuousTensor ParallelW4A16, W8A8Production

Reference architecture

Enterprise LLM Serving Reference Architecture

Model Registry

Versioned model weight storage

MLflowS3NFS Cache

GPU Cluster

A100/H100 nodes with NVLink

8x A100 80GBNVLink 3.0InfiniBand HDR

Inference Fleet

Autoscaling vLLM / TGI replicas

vLLM Replica 1vLLM Replica 2vLLM Replica N

Request Queue

Durable queue for traffic spike absorption

Redis StreamsRabbitMQSQS

API Gateway + Auth

Rate limiting, authentication, request routing

KongAWS API GWCustom Gateway

Client Layer

Applications consuming the LLM API

Web AppMobile AppInternal ToolsAutomation
Stack layers — top to bottom: highest to lowest abstraction

Frequently asked questions

How much GPU memory does an LLM require?

The minimum VRAM equals the parameter count multiplied by 2 (for FP16 precision). A 7B model needs 14 GB minimum, a 70B model needs 140 GB. Add 20–30% headroom for KV cache and activations. INT8 quantization halves these requirements with minimal quality loss.

How do you serve LLMs at scale?

Use a production serving framework like vLLM or TGI with continuous batching enabled. Place a load balancer in front of multiple inference replicas. Implement a request queue for traffic spike absorption. Autoscale based on GPU utilization and queue depth.

What is KV cache and why does it matter?

The KV (key-value) cache stores intermediate attention computations for tokens already processed in a sequence. Without it, every new token would require recomputing attention over the entire context. KV cache enables efficient autoregressive generation but consumes significant VRAM — often as much as the model weights themselves under load.

How do you reduce LLM inference costs?

The highest-impact optimizations are: INT8/INT4 quantization to reduce memory and increase throughput, continuous batching to improve GPU utilization, semantic caching to avoid redundant inference, and model routing to use smaller models for simpler queries. Combined, these can reduce costs by 70–90%.

What is speculative decoding?

Speculative decoding uses a small draft model to predict multiple tokens ahead, then verifies them in parallel with the large model. When predictions are correct (typically 60–80% of the time), multiple tokens are accepted in a single forward pass, reducing latency by 2–3x without changing output quality.