AI Model Distillation Explained: How Smaller Models Learn From Larger Rivals

AI distillation illustration.

AI model distillation is a widely used technique for building smaller, cheaper AI systems, but allegations of unauthorized copying have turned it into a growing intellectual property and geopolitical dispute. Image: Generated via ChatGPT

Aug 3, 2026
5 minute read
eWeek Le contenu et les recommandations de produits sont indépendants de la rédaction. Nous pouvons gagner de l'argent lorsque vous cliquez sur des liens vers nos partenaires. En savoir plus

AI companies are accusing their rivals of copying their homework — millions of answers at a time.

At the center of the dispute is model distillation, a standard training technique in which a large, expensive AI system helps teach a smaller and cheaper one. The method powers efficient models used in phones, private networks, and enterprise software, but allegations that companies are extracting capabilities from closed U.S. systems have turned it into a geopolitical and intellectual property fight.

The difficult question is not whether distillation should exist. It is where legitimate learning ends and unauthorized copying begins.

It's not a new idea. Google's Chief Scientist Jeff Dean said on a podcast in February that his team stumbled upon distillation years ago while trying to boost performance without relying on a single giant image recognition model. 

Distillation overlaps with synthetic-data training and model extraction, but the terms are not interchangeable. Traditional distillation usually involves deliberately transferring behavior from a teacher model to a student, while extraction often refers to reproducing a system’s capabilities through repeated queries without the owner’s cooperation.

Why companies distill AI models

Frontier models can require enormous investments in chips, data centers, energy, and engineering. Distillation offers a cheaper route: companies can transfer some of a larger model’s capabilities into a system designed for a narrower job.

That tradeoff is why distilled models show up in phones, factory equipment, cars, and private company networks — places where a bulky frontier model would never fit. 

In 2019, Hugging Face introduced DistilBERT, one of the best-known distilled versions of Google's BERT language model. It reduced the parameter count by nearly half, ran about 60% faster, and retained roughly 97% of BERT's performance on standard benchmarks.

Why reasoning outputs matter

Older distillation mostly copied final answers. Newer systems can also produce extended reasoning outputs: written intermediate steps showing how they appear to work through a problem. Those outputs are not necessarily a faithful window into a model’s internal computation, but they can provide much richer training material than a final answer alone.

Florian Tramèr, an assistant professor at ETH Zurich who studies machine-learning security, compared it to solving math homework. Learning from a book that only shows final answers is much harder than learning from one that walks through every step, he said. Because reasoning traces reveal more about how a model works, companies now treat access to them as more sensitive than plain outputs.

Advertisement

Where the line gets crossed

Distillation itself isn't illegal or unusual. Stanford's Alpaca project and Microsoft's Orca research both trained smaller models using outputs from bigger ones, and Chinese researchers have published similar work building Chinese-language instruction models. The dispute is about permission and scale, not the technique.

Open-weight models make their parameters available for inspection or modification, but the rights attached to them depend on the license. They generally give developers more room to study and adapt a system than closed models do, although commercial and redistribution restrictions may still apply.

Anthropic has accused Chinese firms DeepSeek, Moonshot AI and MiniMax of running large-scale campaigns to pull capabilities out of Claude, reportedly using around 24,000 fake accounts to generate 16 million exchanges targeting software engineering and reasoning skills. 

The dispute escalated further after Moonshot released its Kimi K3 model, which tested as competitive with top U.S. systems. Director of the White House Office of Science and Technology Policy (OSTP) Michael Kratsios alleged on X that Moonshot "distilled Anthropic's Fable for the development of its K3 model" using "a sophisticated internal platform to conduct large-scale distillation against U.S. models" designed to dodge detection. 

Moonshot has rejected the claim. 

OpenAI has said it has also spotted Chinese attempts to distill its models

US tech companies are divided

More than 60 companies — including Nvidia, Microsoft, Meta and Palantir — signed a letter warning policymakers against "premature restrictions" on open-weight AI that would "stifle competition or drive innovation overseas." On distillation specifically, the group wrote that it "is a widely used technique for model improvement, evaluation, and validation." Nvidia used the method to help train its own Llama Nemotron models.

Notably, Anthropic did not sign that letter; one of the companies with the most to lose from unauthorized distillation of their own systems sat it out, while firms selling chips, cloud infrastructure and open-weight tools backed looser rules.

Pukar Hamal, founder of AI security firm SecurityPal, described unauthorized distillation bluntly: "It's almost like someone went to the lectures, read the textbook, and did all the hard work of doing the homework. Then some other student is like, 'Hey, I didn't do that. Can I just copy your work?'" 

Advertisement

Yet Hamal also said he'd run a Chinese open-weight model like Kimi K3 at his own company if it saved money, so long as it was checked for backdoors and hosted on his own infrastructure, according to CNBC.

What businesses should watch

Companies deploying AI have practical reasons to care beyond the geopolitics. If a business fine-tunes a model on its own proprietary data, that fine-tuned version is a business asset that can itself be distilled by outsiders without permission. Basic protections include rate-limiting API access, logging queries for patterns that look like systematic extraction, writing anti-distillation clauses into terms of service, and keeping the most sensitive models off public-facing endpoints entirely.

Buyers evaluating any model — American or Chinese, closed or open-weight — also have reason to ask where its training data actually came from, since that provenance question is now tied up in export policy discussions in Washington.

The uncomfortable part

Washington's push to treat unauthorized distillation as IP theft runs into an awkward complication: both OpenAI and Anthropic built their own frontier models partly on other people's copyrighted material, and both are currently being sued over it

Max Pritt, an attorney at Boies Schiller Flexner who represents authors in copyright litigation against AI firms, said the government's focus has been lopsided. "The administration, at least publicly, has focused its efforts on the protection of technology companies' intellectual property, while remaining silent in large part about creators and individuals' intellectual property that was used without authorization," Pritt said, per CNBC.

That double standard is likely to shape how any distillation rules get written. Any policy response will therefore face a credibility problem: lawmakers would be protecting AI companies from unauthorized extraction while courts are still deciding whether those same companies lawfully used copyrighted human work to build their models.

Advertisement

The bottom line

Distillation will remain central to AI because it helps turn expensive research systems into models that businesses and consumers can actually run. The unresolved issue is not whether one model may learn from another, but what permission, compensation, and disclosure should be required when it does.

As models become more costly to build and their outputs become more useful as training data, that question will move beyond obscure terms of service. Courts, regulators, AI developers, and their customers will increasingly have to decide when learning from a rival becomes copying one.

Other News: OpenAI has slashed API prices for its GPT-5.6 Terra and Luna models by up to 80% just weeks after launch, aiming to deliver more AI performance per dollar as competition in the enterprise AI market intensifies.

Aminu Abdullahi

Aminu Abdullahi is a B2C and B2B technology and finance writer with more than six years of experience covering enterprise IT, cybersecurity, cloud computing, artificial intelligence, fintech, business software, and emerging technologies. His work has appeared in publications including TechRepublic, eWEEK, Channel Insider, Geekflare, Enterprise Networking Planet, eSecurity Planet, CIO Insight, and Webopedia. With a technical background in computer science, he specializes in translating complex technology topics into clear, accessible content for business leaders and decision-makers.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Propriété de TechnologyAdvice. © 2026 TechnologyAdvice. Tous droits réservés

Divulgation publicitaire : Certains des produits qui apparaissent sur ce site proviennent d'entreprises dont TechnologyAdvice reçoit une compensation. Cette compensation peut influencer la façon dont les produits apparaissent sur ce site, notamment l'ordre dans lequel ils apparaissent. TechnologyAdvice n'inclut pas toutes les entreprises ou tous les types de produits disponibles sur le marché.