AI-Designers
One big model that teaches. A small one that works.
Plenty of companies are stuck on the same problem. The best language model for their task has hundreds of billions of parameters, runs only on expensive hardware or in someone else’s cloud, and burns compute on every single call. The small model that would happily run on their own GPU doesn’t understand the domain language and guesses wrong too often. Neither end is fit for daily work.
Model distillation resolves exactly that tension. A large, capable model takes on the role of the teacher. A small model learns from it until it handles the same task almost as well — with a fraction of the parameters. The idea isn’t new; Geoffrey Hinton and colleagues described it back in 2015. But only now, with strong open models freely available, has it become practical for ordinary companies.

What actually happens during distillation
The trick lies in what the teacher gives away. A normal training example only says: this invoice is an incoming invoice. The teacher says more. It reveals that the document is 88 percent an incoming invoice, 9 percent a credit note, and 3 percent a reminder. Those soft probabilities carry knowledge about similarities and edge cases that a hard yes/no label throws away entirely.
So the small model learns not just the right answer, but the way the big one thinks. On tasks that need reasoning it goes further still. The teacher writes out its intermediate steps, and the small model practices walking the same path — not merely parroting the result. Specialists call this rationale distillation.
In practice it runs in three steps. You collect typical inputs from your operation, let the teacher answer them, and train the small model on those pairs. The teacher only has to work once. After that you’re left with a dataset you can reuse as often as you like.
Why this pays off for mid-sized firms
A distilled model with a few billion parameters runs on a single GPU, often on hardware that already sits in the rack. No per-query licence, no data leaving the building, no waiting on an overloaded API. Anyone classifying tens of thousands of documents a day or pre-sorting email feels the difference immediately — on the power bill and in response time.
Then there’s speed. A small model answers in milliseconds instead of seconds. That opens applications that weren’t possible before: suggestions right in the input field, checks in the middle of a workflow step, analysis on the production line second by second. Latency isn’t a comfort feature. It often decides whether a function gets used at all.
And operations become predictable. A model you own and run locally doesn’t change overnight because a vendor rolls out a new version. You test once, and the behaviour stays stable — an argument that carries real weight in regulated industries.

Where distillation shines — and where it doesn’t
The method is strongest on narrow, clearly defined tasks. Assigning invoices to the right accounts. Sorting support tickets by urgency. Generating product descriptions from master data. Routing customer enquiries to the right department. Wherever the task recurs and the field is bounded, a small model nearly catches up to its big teacher.
Open, broad tasks are a different story. An assistant meant to chat about anything and reason through complexity loses noticeable substance when it shrinks. Here the large model stays ahead. The honest rule: the sharper you scope the task, the smaller the model can be. Try to do everything at once, and you distill the strength right out of it.
Data quality sets limits too. The small model inherits the teacher’s mistakes. If the big one guesses wrong, the small one learns the same error — only faster. That’s why every project needs a sample a human checks before the dataset goes into training.
The path to your own small model
Start with a single task that costs a lot of manual effort today and is clearly measurable. Define up front what success means: a hit rate, a processing time, an error rate. Without that yardstick you can’t tell later whether the small model is good enough.
Then pick the teacher. It can be the most expensive, best model you can reach — you pay for it only once, to generate the training set, not in daily operation. As the student, take an open model of the right size and train it on the generated pairs. The tools for this are mature and well documented.
Finally, compare both models on real, unseen cases. If the student meets your yardstick, it goes into production. If it falls just short, more or better training material usually helps more than a bigger model. This loop — measure, retrain, measure again — is the actual work, and it’s worth it, because the result belongs to you for good.
A building block, not a cure-all
Distillation replaces neither clean data nor clear thinking about the process. It’s a lever for transferring existing strength from large models into small, affordable ones. Combined with quantization and a lean inference server, it produces AI that runs in your own building, answers fast, and stays predictable.
For many companies this is the more realistic route to their own AI than trying to run a giant. You borrow the big model’s cleverness once and keep a small model that does exactly what the business needs.
Wondering which of your recurring tasks would suit your own distilled model? At AI-Designers we analyse your use case, choose the right teacher, and deliver a lean model that runs locally on your hardware. Get in touch — we’ll show you, on a concrete example, what’s possible.
Images: AI-Designed
Deutsch

