de_DE Deutsch |
Multi-LoRA serving: many AI adapters on one shared base model and a single GPU

Multi-LoRA: Run Dozens of Fine-Tuned AI Models on One GPU

One fine-tuned AI per team, but only one GPU? Multi-LoRA serving makes dozens of model variants affordable on a single card.

Legal wants a model that summarises contracts in their register. Support asks for one that sorts tickets into the categories they actually use. Marketing would like an AI that gets product names and claims right. Three teams, three fine-tuned models — and one quick question that follows: do we now need three GPUs? Or twelve?

This is exactly where multi-LoRA serving comes in. Instead of loading a full model into GPU memory for every variant, all the variants share a single base model. The differences live in tiny add-on files that can be swapped per request. It sounds like a detail for infrastructure teams. In practice it decides whether self-hosted AI stays affordable in your organisation or dies on the hardware bill.

What LoRA actually does — briefly, no fog

LoRA stands for Low-Rank Adaptation. Classic fine-tuning changes the weights of the entire model — for a 7-billion-parameter model that means several gigabytes you have to store, load and maintain. LoRA takes a different route. The base model stays frozen, untouched. You train only two small matrices that attach to the model and nudge its behaviour in the direction you want.

The size gap is dramatic. A LoRA adapter for Mistral-7B weighs around 13.6 megabytes. The base model itself tips the scales at 14.48 gigabytes. So the adapter is smaller than one thousandth of the model. In practice you budget for roughly one percent of extra memory per adapter. That lever is what makes the whole approach worthwhile.

One thing matters for understanding this: an adapter is not a standalone model. Without the base model it is useless. Think of it as a lens you screw onto the same camera — the sensor stays the same, the picture changes.

LoRA adapters attach to a shared AI base model
One base model, many lightweight adapters: LoRA bolts the specialisation onto the shared model. · AI-Designed

One base model, many adapters — how it runs

The heart of multi-LoRA serving: you load the heavy base model into GPU memory exactly once. On top of it you stack dozens of lightweight adapters. When a request arrives, the server picks the right adapter by an identifier — often simply an adapter_id — and answers in that adapter’s character. The very next request can use a different one.

Memory usage stays surprisingly low. In one documented setup, 30 adapters loaded at the same time added just three percent more GPU memory. So you get 30 model personalities for a little more than the price of one. Frameworks like vLLM, Text Generation Inference and LoRAX now handle this pattern in production; the optimised kernels behind them come from projects such as Punica.

Throughput does not have to suffer either. On a comparatively cheap Nvidia L4 card, one such setup reached 75 requests per second at an average of 450 input and 234 output tokens. That is plenty for many internal applications — and it runs on hardware that does not cost a fortune.

What it really saves in the data centre

Let us do the maths, with a real example from the community. A provider had trained one LoRA adapter per business customer on a shared Llama-3.1-8B — 40 of them in the end. Running each as a separate deployment would have cost around 24,000 US dollars a month in mostly idle GPU time. With multi-LoRA, all 40 adapters fit on two A100 cards.

The real win is not only the hardware purchase. It is utilisation. Individual fine-tuned models spend most of their time waiting for requests while burning expensive memory. Bundle them onto one base model and they share the baseline load. The card works instead of waiting.

There is a second effect that is harder to put a number on but adds up fast in operations: fewer moving parts. You update one base model once. A new use case means a new adapter of a few megabytes, not a new deployment with its own monitoring, its own memory and its own update chain.

One base model serves several departments with their own adapters
Legal, support, procurement: each team gets its own adapter on a shared base. · AI-Designed

Where multi-LoRA shines day to day

The clearest case is multi-tenant operation. If you offer software as a service and every customer expects an AI tuned to their data, a model per customer would be ruinous. An adapter per customer on a shared base, by contrast, scales cleanly — from the tenth to the hundredth tenant, the underlying infrastructure barely changes.

The departmental case inside your own walls is just as convincing. Legal, support, procurement, HR — each team speaks its own language and works from its own templates. Instead of one generic model that half-fits everyone, you give each team an adapter that hits their exact terms, formats and tone. All on one card, all maintained centrally.

And finally, development itself. New adapters are cheap to train and cheap to test. You can run a revised version next to the old one, split traffic and compare, without spinning up a second machine. That lowers the bar for experiments noticeably — and good AI rarely arrives on the first try.

Limits and pitfalls

Multi-LoRA is no cure-all. Every adapter hangs off the same base model. If two use cases need different base models — say one language model for German and one with special code understanding — they will not share a GPU this way. You plan per base model instead.

Quality has limits too. LoRA shifts behaviour, but it does not conjure knowledge the base model simply lacks. If your task demands deep domain expertise, you will not get around clean, well-prepared data and often retrieval as well. The adapter shapes the answer — the facts have to come from somewhere else.

And one operational detail: many adapters on one card also means that if the card fails, many services go down at once. Part of what you save on hardware should go back into redundancy and monitoring. A second node that steps in belongs to any serious setup.

How to start pragmatically

Start small. Pick a solid open base model that fits your tasks and train two or three adapters for concrete, well-bounded use cases. Choose a serving framework that supports multi-LoRA natively, so per-request switching comes for free.

Measure early and honestly: throughput, response time, memory per additional adapter. Those numbers tell you how far one card carries before you need the next. And document which adapter belongs to which task — with 30 variants you lose track faster than you would like.

If you want to find out how several fine-tuned AI models can be bundled on affordable hardware in your organisation, talk to us. At AI-Designers we guide companies from model selection through fine-tuning to the stable operation of self-hosted AI — hands-on, measurable, and with your data sovereignty in mind.

Images: AI-Designed

Leave a Reply

Your email address will not be published. Required fields are marked *