Machine learning basics for merchandisers: models, features and training data
What a model, a feature, a label and training data actually are, explained with fashion merchandising examples, and the questions a merchandiser should ask before trusting a forecast.
KEY TAKEAWAYS Summary by the editors
- In machine learning, features are the input variables a model uses, the label is the value it predicts, and training is the process of adjusting the model's weights to reduce the gap between predicted and actual values.
- For fashion demand forecasting, useful features typically include price and discount, stock availability, product attributes such as brand and category, and seasonality.
- Zalando researchers reported in 2023 that a deep learning demand forecasting model forecasting 26 weeks ahead at weekly resolution had been in use at Zalando in different versions since 2019.
- Overfitting means a model learns its training data so closely that it fails on new data, which is why models must be evaluated on data they have not seen.
- Forecasting models should be tested on later periods than the ones they were trained on, a method known as time series cross-validation or evaluation on a rolling forecasting origin.
Machine learning is a way of building software that learns patterns from historical data instead of following hand-written rules. For merchandisers, the most common application is forecasting: a model looks at past sales together with information such as price, stock and product attributes, and estimates future demand per style, colour or size. Understanding a handful of terms (model, feature, label, training data and validation) is enough to ask the right questions about any forecast a system produces.
What is a machine learning model?
A model is a mathematical function that turns inputs into a prediction. Google's Machine Learning Crash Course uses linear regression as its starting example: a simple equation in which each input is multiplied by a weight and added up to give a predicted value. During training, the model calculates the weights and bias that produce the best predictions, by reducing the difference between predicted and actual values.
More complex models, such as gradient-boosted trees or deep neural networks, follow the same principle with many more parameters. They can capture interactions, for example that a discount works differently for a basic T-shirt than for a seasonal dress, but they also need more data and are harder to explain.
Most merchandising applications use supervised learning, where the model learns from examples for which the correct answer is known, such as past weeks with recorded sales. Other approaches exist: clustering groups similar stores or customers without predefined labels, which can support assortment planning, and recommendation models learn which products tend to be bought together. Whatever the method, the merchandiser's knowledge of the business is what makes the inputs meaningful, for instance knowing that a week of low sales was caused by a late delivery rather than weak demand.
What are features and labels?
In Google's terminology, a feature is an input variable used to make a prediction, and the label is the value the model predicts. In merchandising, the label is usually units sold or demand in a given week, store or channel. The features are everything that might explain it. The table shows typical examples.
| Feature group | Examples | Why it matters | Common data problem |
|---|---|---|---|
| Price | Full price, discount, markdown timing | Demand reacts strongly to price | Promotions not recorded consistently |
| Availability | Stock on hand, sizes in stock, deliveries | Sales stop when stock runs out | Lost sales hidden as low demand |
| Product attributes | Brand, category, colour, fit, fabric | Lets new styles borrow from similar ones | Inconsistent or missing attributes |
| Time | Week of year, season, holidays | Captures seasonal patterns | Shifting calendars across markets |
| Channel and location | Store, region, online, wholesale account | Demand differs by outlet | Different systems per channel |
| Marketing | Campaigns, newsletter placement | Explains spikes | Often not tracked at article level |

How do fashion retailers use machine learning for forecasting?
A well-documented example comes from Zalando. In a 2023 paper, Zalando researchers described a deep learning model that forecasts demand 26 weeks ahead at weekly resolution and, according to the paper, has been in use at Zalando in different versions since 2019. Its inputs include discounts, the recommended retail price, stock levels, historical sales, brand and commodity group, and an encoding of the annual seasonal cycle. The authors identify the challenges that make fashion forecasting hard: large data volumes, irregular demand, high catalogue turnover, and the need to model carefully how price affects demand.
The lesson for merchandisers is that the model is only one part of the work. Most effort goes into deciding which features to include, cleaning them and making sure the relationship between price, stock and demand is represented correctly.
What is training data, and why does it decide quality?
Training data is the historical record the model learns from. Its quality sets the ceiling for forecast quality. Three issues are typical in fashion:
- Censored demand: when a size sells out, recorded sales understate true demand. Without stock data the model learns that the item was less popular than it was.
- Short histories: most styles live for one season, so the model must learn from similar products via their attributes rather than from the item's own past.
- Changing conditions: unusual seasons, new channels or a change in pricing strategy mean the past is a weaker guide to the future.
What is overfitting, and how are models tested?
Google describes overfitting as a model learning the training data so closely that it fails to make correct predictions on new data: it works in the lab but not in the real world. The defence is to keep data back. Models are trained on one part of the data, tuned on a validation set and judged on a test set they have never seen.
For forecasts, the split must respect time. Hyndman and Athanasopoulos's textbook Forecasting: Principles and Practice describes time series cross-validation, also called evaluation on a rolling forecasting origin: each test uses only observations that came before it, and accuracy is averaged across many such tests. In practice, a merchandiser should ask to see how a model would have performed on last season, forecast from data available before that season started.

How should merchandisers work with machine learning models?
- Agree on the label: what exactly is forecast, at what level (style, colour, size, store) and for what horizon.
- Audit the features: make sure prices, promotions, stock and attributes are recorded consistently.
- Demand a fair benchmark: compare the model with the existing method on the same, unseen periods.
- Keep judgement in the loop: models do not know about a new campaign or a supplier delay unless someone tells them.
- Monitor after launch: accuracy drifts as assortments and customers change, so models need regular review and retraining.
Frequently asked questions
What is the difference between a feature and a label in machine learning?
A feature is an input variable the model uses to make a prediction, such as price, discount or product category. The label is the value the model predicts, such as units sold in a given week. Training adjusts the model so that its predicted labels match the actual ones as closely as possible.
How much data do you need for a fashion demand forecast?
There is no fixed amount, but the model needs consistent sales, price, stock and attribute history across several seasons. Because most fashion styles are short-lived, the breadth and quality of product attributes often matter more than the length of any single item's history.
What is overfitting in simple terms?
Overfitting happens when a model memorises its training data instead of learning general patterns, so it performs well on the past but poorly on new data. It is detected by testing the model on data it has not seen during training.
How do I know if an AI forecast is better than our current planning?
Compare both methods on the same past periods that the model did not see during training, using forecasts made only with data available at the time. This approach, known as time series cross-validation, gives a fair comparison of accuracy before the model is used for buying decisions.
One edition every weekday morning. Read in five minutes. Free for industry professionals.
SOURCES
- Google for Developers: Machine Learning Crash Course, Linear regression
- Google for Developers: Machine Learning Crash Course, Overfitting
- arXiv: Deep Learning based Forecasting, a case study from the online fashion industry (Kunz et al.)
- OTexts: Forecasting: Principles and Practice (3rd ed), 5.10 Time series cross-validation



