7 October 2026International edition
Vol. I · No.
7 October 2026
AI in Fashion
DAILY
The daily briefing on AI in the fashion business
Where fashion meets artificial intelligence.
Glossary

What is model distillation in AI?

A technique that trains a smaller AI model to imitate the outputs of a larger one, making it cheaper and faster to run with similar quality on a task.

In short

Model distillation is a technique in which a smaller "student" model is trained to reproduce the outputs of a larger "teacher" model. The result is a model that is cheaper and faster to run, while keeping much of the larger model's quality for a specific task.

How does it work in practice?

First, a large foundation model is used to produce answers for many examples, such as product attribute labels for thousands of images or descriptions for a range of styles. These outputs, sometimes including the teacher's probability scores, become training data for a smaller model. The student learns to give similar answers. After testing, the small model replaces the large one in production for that task, often at a fraction of the inference cost.

Why does it matter for fashion businesses?

Fashion catalogues contain many thousands of products and variants, and tasks such as tagging, translation or search must run constantly. Using the largest models for every request can become expensive and slow. A distilled model can handle high-volume, repeatable tasks cheaply, while the large model is kept for complex cases. It can also run on a company's own servers, which helps when product or customer data must not leave controlled systems.

How does AI use it?

AI providers regularly release smaller model versions created partly through distillation. Companies apply the same idea internally: they use a powerful model to label data or generate examples, then train a small language model or vision model for their own narrow use case. Distillation is often combined with fine-tuning on company-specific data.

Common pitfalls

  • Expecting the student model to handle tasks outside the training examples.
  • Copying the teacher's mistakes and biases into the smaller model.
  • Ignoring licence terms that may restrict using one model's outputs to train another.
  • Skipping evaluation against real business cases before switching production.

Frequently asked questions

What is the difference between distillation and fine-tuning?

Fine-tuning adapts an existing model with new training data. Distillation trains a smaller model specifically to imitate a larger one, and the training data often comes from the larger model itself.

Why would a company use a distilled model?

Distilled models are cheaper, faster and easier to host. They suit high-volume, narrow tasks where the full capability of a large model is not needed.

All terms