The core idea is old: train a small 'student' model not just on the data the large model saw, but on the large model's outputs — its probability distributions, its reasoning traces, its judgments. The student learns the teacher's decision function rather than rediscovering it from raw data, achieving much of the capability at a fraction of the size.
Modern practice refines this considerably. Reasoning traces from frontier models — the chain-of-thought a large model produces while solving a problem — have proven especially valuable as training material, teaching small models not just answers but the process that generates them. On-policy variants train the student on its own generated outputs scored by the teacher, closing gaps that offline data leaves behind.
Why it matters
Inference is where AI's costs live. A frontier model served at scale costs orders of magnitude more per query than a distilled variant, which is why every product team's roadmap runs through distillation: flagship model for the hard cases, distilled model for the volume.
On-device AI raises the stakes. Phones, cars and laptops cannot host frontier-class models, so the features users experience — summarization, translation, photo understanding, voice assistants — are almost universally distilled. When a phone maker advertises on-device AI, it is advertising a distillation program, whether or not the word appears.
How it works in practice
The pipeline: generate outputs from the teacher model across the task distribution, filter for quality, train the student on the outputs (and sometimes the reasoning), then evaluate against a benchmark suite. The teacher's logs and traces are the asset; labs that keep them hold a compounding advantage.
Evaluation is the hard part. Distillation quality varies unpredictably across tasks — a student may match its teacher on summarization while lagging badly on multi-step reasoning. Production teams therefore benchmark distillations per use case, not per model, and increasingly route queries: cheap model first, escalate to the frontier model when confidence is low.
Evidence
The pattern is visible in every major model family: OpenAI, Google and Meta all ship small variants positioned as distilled or efficiency-tuned, with published benchmarks showing capability retention far above what model size alone would predict. Open-weight ecosystems show the same dynamic, with community distillations of frontier models dominating the small-model leaderboards.
Licensing has become the contested edge: some open-weight model licenses explicitly permit distillation, others restrict using outputs to train competing models. The legal fights over whether model outputs can be used as training data at all are ongoing, and the answers will shape how freely the technique propagates.
The competing read
Distillation skeptics note a ceiling: students inherit their teachers' blind spots and cap out below the frontier, so sustained progress still requires expensive frontier training. If distillation were the whole story, capability would plateau at the current frontier's level plus modest compression gains.
Pragmatists respond that this is exactly the point: distillation diffuses capability, it doesn't create it. The frontier lab competes on new capability; everyone else competes on efficiently serving what exists. Both dynamics can be true, and the current market clearly rewards the second more than most commentary acknowledges.
What happens next
Watch for stricter distillation terms in next-generation model licenses, watch on-device model benchmarks as phone makers' offerings mature, and watch the routing infrastructure layer — the systems that decide which model handles which query — which is quietly becoming one of the most economically consequential parts of the stack.
