de_DE Deutsch |
Modern enterprise server room with GPU servers for AI inference

Inference Servers: How vLLM Makes Self-Hosted AI Pay Off

Continuous batching, PagedAttention, quantization: how inference servers saturate the GPU and make self-hosted AI economical.

A company buys a GPU card for 15,000 euros, loads an open-weight model onto it, and discovers something awkward: the answers arrive, but the card is bored. Utilization sits at 12 percent; the rest evaporates. This is where it gets decided whether self-hosted AI stays an expensive hobby or actually pays off. The model is not the lever. The server that feeds it is.

Most discussions circle around models: which Llama, which Qwen, which Mistral. But the economics are settled one layer down, in the inference server. It determines how many requests a single card handles at once, how much memory goes to waste, and how quickly your team can scale when demand spikes. Ignore it and you pay twice — once for hardware, once for idle compute.

Why loading a model isn’t enough

A language model produces text token by token. Each step runs a pass through several billion parameters. If a naive server loads the model and answers one request after another, the GPU computes for a single user and then waits. In a company with a hundred parallel requests from support, accounting, and sales, that is ruinous.

The trick behind modern inference servers is to push many requests through the same card simultaneously. A GPU is a parallel machine with thousands of compute units. Feeding it properly doubles or tenfolds throughput without buying a second card. The gap between a naive script and vLLM is often the gap between ten and a hundred concurrent users.

Continuous batching: never let the card idle

Classic systems bundle requests into fixed groups. Five requests arrive, and the server waits until all five finish before starting the next group. If one answer completes early, the freed slot stays blocked until the slowest neighbor is done. With wildly varying answer lengths — a yes-no question next to a three-page summary — that wastes an enormous amount of compute.

Continuous batching breaks this rigid pattern. The moment one request finishes, the next slides into the free slot. The card works without interruption instead of waiting for stragglers. vLLM popularized the technique; SGLang and Hugging Face’s Text Generation Inference use it too. For your operation it means more requests per second on the same hardware, and you feel that directly on the electricity bill.

There’s a second effect on latency. Because no user is stuck behind a slow batch neighbor anymore, response times stay stable even under load. Anyone embedding AI into a customer portal notices the difference in abandonment rates.

GPU accelerator card inside an inference server
Not the model, but how well the GPU is utilized decides the cost of self-hosted AI. · AI-Designed

PagedAttention: the memory nobody sees

To keep track of a conversation, a model builds a so-called KV cache — a kind of short-term memory for each active request. Conventional servers reserve one large, contiguous block of memory per request, sized for the maximum possible length. If a request uses only a fraction of that, the rest lies fallow. In practice, 60 to 80 percent of expensive GPU memory can vanish this way.

PagedAttention, vLLM’s core innovation, borrows an old idea from operating systems and applies it to the GPU: memory is split into small pages and handed out only when actually needed. The model accesses scattered pages through a table as if they were contiguous. Freed memory goes straight to new requests. Concretely, that means many more users run in parallel on the same card, because the memory is no longer wasted.

Quantization: the same model on smaller hardware

Models usually store their parameters as 16-bit numbers. Quantization lowers that precision to 8 or even 4 bits. A 70-billion-parameter model that occupies two large cards in full precision often fits on a single card once quantized. With well-chosen methods, answer quality barely suffers, while the hardware requirement drops sharply.

For budget planning, this is the single biggest lever. Instead of investing in a second GPU, you run the model quantized on existing hardware and gain throughput on top, because less data moves around. Methods like AWQ and GPTQ are built directly into vLLM and SGLang. One caveat: test quality on your own tasks, not on someone else’s benchmarks. A legal department forgives inaccuracy differently than an internal FAQ.

IT team monitoring throughput and latency of an AI inference server
Measure throughput and latency under realistic load before committing to an inference server. · AI-Designed

vLLM, SGLang, TGI, TensorRT-LLM: the candidates

There is no universal winner. vLLM has become the de facto standard, broadly proven and backed by a huge community — the sensible starting point for most shops. SGLang shines with structured outputs and very high concurrency. Hugging Face’s Text Generation Inference is a mature, well-documented all-rounder. TensorRT-LLM with the Triton server squeezes the last bit of performance out of NVIDIA cards, at the cost of more setup effort.

The honest advice: benchmark with your real requests. A server that dazzles on short chat messages can collapse on long documents. Measure throughput and latency under realistic load before you commit. Two afternoons of comparison save weeks of rebuilding later.

Cold starts and autoscaling in daily operation

One practical problem surfaces at the latest when you scale: model weights are huge, from a few gigabytes to well over a hundred. When a new server node starts, it has to load those weights first — the notorious cold start, which can take minutes. Anyone trying to absorb load spikes automatically trips over exactly this.

The proven remedies are unspectacular but effective: download the weights once and mount them as a shared volume on every node, keep container images lean, and speed up loading via streaming. That shrinks the multi-minute cold start to a few seconds, and autoscaling by requests per second becomes viable in the first place. Combined with an AI gateway in front that distributes requests, you get a system that breathes with the load instead of suffocating under it.

What this means for your budget

Don’t reckon in model prices; reckon in requests per second per card. That single metric decides how many users an investment supports. Continuous batching and PagedAttention raise it several times over, while quantization lowers the hardware you need. Together, these three techniques push the break-even point for running your own stack versus the cloud noticeably forward.

Building it takes care, not wizardry. The components are open source and mature, the pitfalls well understood. Know them, and you run your own AI at costs that were unthinkable two years ago — while keeping the data in your own house.

Thinking about putting an open-weight model into production in your own data center, and want the serving layer set up right from day one? We help you choose, benchmark, and operate your self-hosted AI. Talk to us at ai-designers.eu.

Images: AI-Designed

Leave a Reply

Your email address will not be published. Required fields are marked *