How do you measure forecast accuracy in fashion? MAPE, WAPE, bias and MASE
Forecast accuracy decides whether an AI planning tool is worth its cost, yet the most popular metric is badly suited to fashion. A practical guide to MAPE, WAPE, bias and MASE, and what good looks like.
KEY TAKEAWAYS Summary by the editors
- Forecast accuracy in fashion should be measured with at least two numbers: an error size metric such as WAPE or MASE, and a bias metric that shows whether forecasts systematically run high or low.
- MAPE is undefined when actual sales are zero and becomes extreme when sales are close to zero, which makes it unreliable for slow-selling sizes, colours and stores.
- MASE, proposed by Hyndman and Koehler, scales errors against a simple naive forecast, so a value below one means the forecast beats that naive benchmark.
- Accuracy must be measured at the level where decisions are taken, such as style-colour by week or store, because aggregate accuracy hides errors that cause markdowns and stockouts.
- There is no universal good accuracy figure for fashion: the meaningful test is improvement against the current method on the same products, horizon and level.
Forecast accuracy in fashion is best measured with an error metric that copes with low and zero sales, such as WAPE or MASE, plus a separate bias measure. MAPE, the most quoted metric, breaks down on slow sellers and should not be the only number on the dashboard. What counts as good depends on the level, horizon and product type, so accuracy should always be judged against a benchmark on the same data.
Why does forecast accuracy matter so much in fashion?
Every buy, allocation and re-order decision rests on a forecast. Over-forecasting becomes excess stock and markdowns; under-forecasting becomes stockouts and lost full-price sales. When a brand evaluates an AI forecasting tool, accuracy is also the main evidence for its return on investment. If the measurement is wrong, the business case is wrong too.
What are the main forecast accuracy metrics?
The textbook Forecasting: Principles and Practice by Hyndman and Athanasopoulos groups accuracy measures into scale-dependent errors, percentage errors and scaled errors. The table summarises the metrics most often used in retail planning.
| Metric | What it measures | Strength | Weakness in fashion |
|---|---|---|---|
| MAE (mean absolute error) | Average size of error in units | Easy to explain | Cannot compare items with different volumes |
| RMSE (root mean squared error) | Error with large misses weighted more | Highlights big misses | Sensitive to outliers; unit-dependent |
| MAPE (mean absolute percentage error) | Average error as a percentage of actuals | Unit-free, intuitive | Undefined at zero sales; extreme near zero |
| WAPE (weighted absolute percentage error) | Total absolute error divided by total actual sales | Robust to low-volume items | Dominated by high-volume items |
| MASE (mean absolute scaled error) | Error relative to a naive forecast | Comparable across items; works with zeros | Less intuitive for business users |
| Bias | Average signed error (over or under) | Shows systematic direction | Says nothing about error size |

What is wrong with MAPE for fashion forecasts?
MAPE divides each error by the actual value. Hyndman and Athanasopoulos note that percentage measures are infinite or undefined when an actual value is zero, and give extreme results when actuals are close to zero. In fashion, zero and near-zero sales are everywhere: a size 34 in one store, a niche colourway, the first or last week of a product's life.
In their paper Another look at measures of forecast accuracy, Hyndman and Koehler go further: for intermittent demand with small counts, frequent zeros make percentage measures impossible to use. They also point out that the M3 forecasting competition avoided the issue by excluding such data, a fix that is not available in practice. The textbook adds a second problem: MAPE penalises negative errors more heavily than positive ones, which can nudge teams towards systematically low forecasts.
What should fashion teams use instead?
Two alternatives are widely used. WAPE sums absolute errors across items and divides by total sales, so a missed size on a slow item does not blow up the result. It is the metric reported, for example, in the VISUELLE research on new fast-fashion products by Skenderi and colleagues. Its drawback is that bestsellers dominate the figure, so it should be read alongside a breakdown by product segment.
MASE, proposed by Hyndman and Koehler, scales each error by the error of a simple naive forecast on historical data. A value below one means the forecast beats that naive benchmark; above one means it does worse. Because it remains defined when there are zeros, it suits granular fashion data, and it builds a benchmark into the metric itself.
Why measure bias separately?
Error metrics measure size, not direction. A forecast can have acceptable WAPE while consistently over-forecasting by a few percent, which quietly builds excess stock season after season. Bias is the average signed error, or total forecast minus total actuals as a share of actuals. It should be tracked by category, channel and planner override, because systematic optimism often enters through manual adjustments rather than the model.
The choice of metric also shapes what a model optimises. The textbook notes that a method minimising MAE produces forecasts of the median, while minimising RMSE produces forecasts of the mean. For skewed fashion demand those two can differ noticeably, so the metric in a vendor contract effectively defines what kind of forecast is delivered.
At what level should accuracy be measured?
Accuracy improves as data is aggregated, so the same forecast can look excellent at brand-month level and poor at style-colour-size-store-week level. Measure at the level where the decision is taken.
- Buying: style or style-colour, total season, at the lead time when the buy is committed.
- Allocation: style-colour by store or cluster, by week.
- Replenishment and re-order: SKU by location, over the replenishment lead time.
- Financial planning: category by month, where aggregate accuracy is appropriate.
Sales should also be corrected for stockouts before measuring. A forecast that was right about demand looks wrong if the product sold out and recorded sales were capped.
What does good forecast accuracy look like?
There is no reliable industry-wide target for fashion, and any vendor quoting a universal percentage should be asked how it was measured. Published research gives a sense of scale: in the Fisher and Vaidyanathan study of new-product forecasting in retail categories, errors on sales shares ranged from 16.2% to 28.7% MAPE, at an aggregate level and outside fashion. Granular fashion forecasts will typically show much higher errors.
A practical evaluation routine follows from this.
- Fix the decision level and lead time for each use case before testing.
- Hold out a past season the model has not seen and forecast it as if live.
- Report WAPE or MASE, plus bias, by category, channel and new versus continuing products.
- Compare with the current planner forecast and with a simple naive benchmark.
- Translate error reduction into stock and margin effects, so finance can judge the return on investment.
Frequently asked questions
What is a good MAPE for fashion forecasting?
There is no universal good MAPE for fashion, because accuracy depends heavily on level, horizon and product type. MAPE is also unreliable for low-volume items because it is undefined at zero sales. It is better to compare WAPE or MASE against your current method on the same data.
What is the difference between MAPE and WAPE?
MAPE averages the percentage error of each item, so a small item with a large percentage miss can distort the result. WAPE divides total absolute error by total actual sales, which makes it robust to low-volume items but weights bestsellers more heavily.
What is forecast bias?
Forecast bias is the tendency of forecasts to be systematically too high or too low. It is measured as the average signed error, and it matters in fashion because a persistent positive bias builds excess stock and markdowns even when error size looks acceptable.
What does a MASE below 1 mean?
A MASE below 1 means the forecast's average error is smaller than that of a simple naive forecast on the historical data. It is a quick check that a model, or an AI tool, adds value beyond a basic benchmark.
One edition every weekday morning. Read in five minutes. Free for industry professionals.
SOURCES
- Hyndman and Athanasopoulos, Forecasting: Principles and Practice (3rd ed.): Evaluating point forecast accuracy
- Rob J. Hyndman: Another look at measures of forecast accuracy (Hyndman and Koehler)
- arXiv: Well Googled is Half Done, multimodal forecasting of new fashion product sales with image-based Google Trends (Skenderi et al.)
- IDEAS/RePEc: A Demand Estimation Procedure for Retail Assortment Optimization with Results from Implementations (Fisher and Vaidyanathan, Management Science 2014)

