Skip to content
ansezz.
← Back to blog
Architecture May 24, 2026 9 min read 1,603 words

GPU-aware load balancing for AI inference

Round-robin breaks when LLM requests span 50 to 50,000 tokens. GPU-aware load balancing with prefill/decode disaggregation and the metrics that cut P99.

Anass Ez-zouaine

Backend · Architect · AI

▸ Share

GPU-aware load balancer routing AI inference traffic across a fleet
▸ On this page (6)

Your load balancer says CPU is at 40% and everything is fine. Your GPUs are maxed out and your P99 latency is in the gutter.

You just scaled your RAG application to a hundred concurrent users. Suddenly, your latency spikes.

Some users get their answers in two seconds, while others stare at a loading spinner for thirty. This is the problem GPU-aware load balancing exists to solve.

Models are not web servers

You are treating your AI models like traditional web servers. Sending a 4,000-token prompt to the same GPU that is generating a 50-token summary is a recipe for disaster.

Round-robin routing is a relic of the past for LLM inference. If you don’t account for how GPUs handle compute and memory, you waste money and hurt your user experience at the same time.

More GPUs won’t fix it alone

The fix is a load balancer that understands what is happening inside the model. That means GPU-aware routing, prefill vs decode disaggregation, and treating your KV cache as one of the most valuable assets in your stack.

Why round-robin is a trap for LLMs

In traditional software, a request is a request. Whether it’s a GET /users or a POST /orders, the variance in resource use is usually small and predictable.

Standard load balancers like Nginx or HAProxy work great here. They look at basic health checks and send traffic to the next available worker.

Requests have very different weights

AI is different. One user might ask “what is 2+2?” while another uploads a 50-page PDF and asks for a deep analysis.

If your load balancer sends both to the same GPU, the heavy request hogs the compute, and the light request waits in a queue.

CPU metrics tell you nothing

A GPU can be at 100% utilization while doing very different types of work. Some work is compute-bound and needs raw processing power.

Other work is memory-bound, limited by how fast data moves in and out of VRAM. To route well, you have to look inside the inference lifecycle.

Prefill vs decode: the performance gap

LLM inference happens in two distinct phases. Understanding the difference is the “aha!” moment for GPU load balancing.

Prefill: read the whole prompt

In prefill, the model reads your entire prompt and processes all the tokens at once. It is a heavy, compute-intensive task that builds the KV cache (key-value cache).

Prefill loves big batches and fast tensor cores. It is where the heavy lifting happens.

Decode: one token at a time

In decode, the model generates the response one token at a time. Each new token only needs the previous tokens and the KV cache.

This phase is light on compute but very heavy on memory bandwidth. It is slow and long-lived.

Why mixing them hurts

When you mix the two on one GPU without a smart scheduler, the prefill of a new request often pauses the decode of existing requests. That causes the jittery, stuttering text users hate.

GPU-aware load balancing lets you treat these phases differently across your fleet.

PhaseWhat it doesBound byMetric to watch
PrefillReads the full prompt at onceComputeTime to first token (TTFT)
DecodeWrites the answer token by tokenMemory bandwidthInter-token latency (ITL)
Comic panel of a strong robot tearing through a paper stack while a slim robot typing one letter at a time gets pushed aside
Prefill and decode need different things, so mixing them on one GPU makes answers stutter.

Metrics for the real world

To build a better router, stop looking at CPU and start looking at these four metrics:

  1. Token queue depth: how many tokens are waiting to be processed? This shows “load” far better than request counts.
  2. KV cache utilization: GPUs have limited VRAM, and the KV cache holds the “memory” of ongoing conversations. If a GPU’s VRAM is 90% full of KV cache, it cannot accept a large new prompt, even if it looks “idle.”
  3. Time to first token (TTFT): the latency of the prefill phase. If TTFT is climbing, your prefill pool is congested.
  4. Inter-token latency (ITL): the speed of the decode phase. If this is high, your GPUs are likely limited by memory bandwidth.

Get the numbers from your engine

I often recommend tools like vLLM because they expose these metrics out of the box.

You can pipe them into a custom gateway that routes on real-time VRAM availability instead of a simple “is the server up?” check.

Prefix-aware routing: reuse the KV cache

The most expensive part of a RAG request is often re-processing the same system prompt or long context again and again.

If you send five follow-up questions about the same document to five different GPUs, each GPU runs prefill for that document from scratch. On hosted APIs, prompt caching attacks the same cost, and my LLM cost calculator shows the effect on your monthly bill.

Send the prompt where the cache is warm

Prefix-aware routing (also called KV-cache-aware routing) fixes this. Your load balancer tokenizes the start of the prompt and looks for a GPU that already has that content in its KV cache.

Matching the prefix lets you skip prefill for the shared part and reuse the cached keys and values directly. That can turn hundreds of milliseconds of TTFT into a near-instant first token.

Every repeated prefill is a GPU redoing work it already finished somewhere else.

A common RAG blind spot

This is one of the most effective optimizations in a production RAG system. I’ve written before about common RAG mistakes, and ignoring cache locality is one of them.

For the layer above this (caching whole responses and embeddings), see Redis and semantic caching for RAG.

Splitting the fleet into specialized pools

As you scale, stop treating every GPU as a generalist. Build disaggregated inference fleets instead.

I like to split my GPUs into two pools:

  • The prefill pool: compute-heavy GPUs (like H100s or B200s) tuned to chew through large prompts fast. These nodes process the initial context, then hand off the KV state.
  • The decode pool: GPUs sized for memory bandwidth and capacity rather than raw FLOPs (think A100s or H200s) that focus on producing tokens for existing requests.

Scale the pool your users stress

If your users upload huge documents but ask for short summaries, scale your prefill pool. If they have long, chatty conversations, scale your decode pool.

This is the same logic used in modern DevOps with Coolify. You wouldn’t put your heavy database on the same tiny instance as your frontend, so why mix heavy prefill work with light decode work?

One mixed pool

Every GPU does prefill and decode. New long prompts pause ongoing answers, and you can only scale everything at once.

Split pools

Prefill GPUs read prompts and hand off the KV state. Decode GPUs keep streaming tokens, and you scale each pool for the load it actually gets.
Comic panel of a traffic robot sending big crates to heavy-lifter robots and small cubes to fast typist robots
Separate prefill and decode pools let you scale each for the load it actually gets.

Implementing your first GPU-aware router

You don’t need to build a custom engine from scratch. Here is the practical path I follow when setting this up for a new SaaS:

  1. Centralize your metrics: use Prometheus to scrape vLLM or TGI metrics from every GPU node.
  2. Use a smart gateway: write a middleware in Go or Rust (or even a heavy-duty Lua script in OpenResty) that checks these metrics before choosing a target.
  3. Prioritize KV cache: route on prefix overlap when you can. A cheap first approximation is session stickiness: if a node served this conversation_id recently and isn’t at 100% KV utilization, send the follow-up there so the warm cache gets reused.
  4. Set hard limits: if a GPU reaches 85% VRAM usage, take it out of the rotation for new prompts until some sessions finish.

From black box to context-aware

Managing AI compute means moving from “black box” infrastructure to “context-aware” infrastructure. When your load balancer knows the difference between a 10-token greeting and a 10,000-token context window, your costs go down and your users stay happy.

It’s easy to get lost in the hype of “agentic systems” and context-aware agents, but none of that matters if your infrastructure is buckling under unoptimized routing.

Key takeaways

  • Round-robin fails for LLMs because request weight varies from a few tokens to whole documents.
  • Prefill is compute-bound, decode is memory-bound. Mixing them on one GPU causes stutter.
  • Route on token queue depth, KV cache use, TTFT and ITL, not CPU.
  • Prefix-aware routing reuses the KV cache and cuts time to first token.
  • Split prefill and decode pools so you can scale each for your real workload.

If round-robin is hurting your inference latency, here’s how I help teams build GPU-aware routing for LLM serving, or drop a note via contact.

If you are still using round-robin for your AI models, what is the main bottleneck you are seeing in your P99 latency right now?

▸ Made it to the end? Send it around.

▸ Share

▸ Comments