Forecast value added: proving a model beats the planner
Forecast value added compares each step of a forecasting process with a simple benchmark, so a team can see whether models and manual overrides improve accuracy or only add effort.

KEY TAKEAWAYS Summary by the editors
- Forecast value added (FVA) is the change in a forecast performance metric attributable to a particular step or participant in the forecasting process, introduced by Mike Gilliland in Foresight in spring 2013.
- FVA starts from a naive benchmark, usually the last observed value or a seasonal naive forecast, and compares each later step (statistical model, planner override, final consensus) with the step before it.
- In a case cited in Foresight, moving from naive to a statistical forecast added about 5% accuracy at Newell Rubbermaid, while judgmental overrides reduced accuracy by about 2%.
- A critique published in Foresight in 2024 argues FVA shows correlation, not causation, ignores uncertainty and financial cost, and can be gamed, but concedes it is useful as a one-off reality check.
- A fair test needs the same items, the same periods, the statistical forecast, the final forecast and the actual values, measured on data frozen at the time of the forecast.
To prove that a model beats a planner, compare both with the same simple benchmark on the same items and periods, using forecasts that were frozen before the actual sales were known. The method for doing this is called forecast value added, or FVA: it measures whether each step in the forecasting process makes the forecast better or worse than the step before.
This how-to explains the method, the data you need, a step-by-step procedure, how to read the results, and the critiques that should shape how you use it.
What is forecast value added?
According to a Foresight article, FVA is the change in a performance metric, usually a MAPE variant, that is attributable to a particular step or participant in the forecasting process. It was introduced by Mike Gilliland, a forecasting practitioner at SAS, in Foresight: The International Journal of Applied Forecasting, Spring 2013, issue 29. The article lists the typical process as: sales history, forecasting model, statistical forecast, management override, final forecast.
Each step is compared with the step before it. A positive FVA means the step improved the forecast and a negative FVA means it made the forecast worse. The result is often shown as a stairstep report with a row per step. In the article's own illustrative example, a naive forecast has a MAPE of 50%, the statistical forecast 40% and the final forecast 42%, so the statistical step adds value and the override step destroys some of it.
Why start with a naive benchmark?
The naive forecast is a cheap, honest baseline. The Foresight article defines it as the random walk or no-change model, where the last observed value becomes the forecast for all future periods. For seasonal data the textbook by Hyndman and Athanasopoulos describes the seasonal naive forecast, where each forecast copies the value from the same season of the last year. The textbook states that any new forecasting method should be compared against these simple methods, and that if it cannot beat them it is not worth considering.
In fashion, a naive benchmark needs thought. Styles are short-lived, so last year's value for the same style may not exist. Common substitutes are the same period last year at category level, or a moving average of the first weeks of sales. Whichever is chosen, document it and keep it fixed.

How do you run an FVA analysis step by step?
- Map your process: list each step from history to final consensus forecast and who touches it.
- Capture the forecasts: for each item, location and period, store the statistical forecast, the final forecast and the actual. Freeze them at the time of the forecast.
- Choose the metric: use a measure that suits sparse data, for example weighted absolute percentage error at an aggregated level, plus a bias measure.
- Choose the benchmark: naive or seasonal naive, or a simple moving average when history is short.
- Compute the metric for each step and the difference between steps, by item group and overall.
- Look at the distribution, not just the average: a positive average can hide many items where the model did worse.
That last point comes from the Foresight article's own case. At Newell Rubbermaid, as reported by Schubert and Rickard in 2011, the move from naive to statistical forecast added about 5% accuracy, while judgmental overrides reduced accuracy by about 2%. Histograms of item-level FVA revealed many items where the statistical forecast did worse, even though the average was positive. The article also describes Tempur-Pedic using FVA to make a collaborative process visible, challenging salespeople to beat the statistical forecast, which reduced unnecessary adjustments. These are cases from consumer goods companies, not fashion, and results will differ by category.
| Comparison | Question | If FVA is positive | If FVA is negative |
|---|---|---|---|
| Statistical forecast vs naive | Does the model beat a trivial rule? | The model is worth its cost | Simplify or fix the model and data |
| Final forecast vs statistical forecast | Do planner overrides help? | Keep targeted overrides | Restrict overrides to defined situations |
| Machine learning model vs statistical model | Is the extra complexity justified? | Consider adopting it for that item group | Keep the simpler method |
| Consensus with sales vs planner forecast | Does sales input add information? | Keep the step, document why | Drop or limit the step |
How do you interpret the results fairly?
- Segment by item type: core, fashion and new items behave differently, and an override that helps for new styles may hurt for core lines.
- Check enough periods: a few weeks can favour either side by luck.
- Keep overrides labelled with a reason, so you can see whether reasons such as a known campaign or a supply constraint add value.
- Test with an out-of-sample logic: forecasts must be frozen before actuals, otherwise the comparison is contaminated.
- Report forecast bias alongside accuracy: planners often add optimism, which accuracy measures can hide.
Metric choice matters. The Hyndman and Athanasopoulos textbook lists drawbacks of MAPE: it is undefined when an actual value is zero, unstable near zero, and penalises negative errors more heavily than positive ones. Scaled errors such as MASE divide by the error of a naive forecast, so a value below 1 beats the naive benchmark. For fashion data with many zeros, an aggregated weighted error or a scaled error is usually safer than a plain MAPE.
What are the criticisms of FVA?
A 2024 Foresight article by Conor Doherty of Lokad argues that FVA has limited value as an ongoing tool. His main points are that FVA shows correlation, not causation, so a positive override may be luck; that it relies on point forecasts and ignores uncertainty; that accuracy is not profit; that it diagnoses but does not explain; and that once accuracy becomes a performance measure, people may change their behaviour to look good. He also notes that FVA can add bureaucracy. Lokad sells forecasting software, so readers should weigh the argument with that interest in mind.
The critique does concede that FVA can be useful as a one-off reality check that shows overconfident forecasters how unreliable manual overrides can be. A balanced practice follows from both sources: use FVA to test claims, such as whether a model beats the planner, and complement it with financial measures such as markdown cost and lost sales, so that accuracy gains are tied to money.

How do you turn the result into a decision?
If the model beats the naive benchmark but the final forecast is worse than the model, restrict overrides to defined situations, such as confirmed campaigns, and require a reason. If the model does not beat the benchmark, look first at data quality, the hierarchy level at which you forecast, and the handling of stockouts before replacing the model. If planners beat the model for particular groups, such as new styles, treat that as information about where the model lacks inputs, not as a verdict on the whole approach.
Frequently asked questions
What is forecast value added (FVA)?
It is the change in a forecast performance metric attributable to a step or participant in the forecasting process, such as a statistical model or a manual override. It compares each step with the one before, starting from a naive benchmark.
How do I test whether AI beats my demand planner?
Freeze both forecasts before actuals are known, use the same items and periods, add a naive benchmark, and compare errors by item group. Report bias as well as accuracy, and check enough periods to avoid results driven by luck.
Do manual overrides improve forecasts?
It depends. In a case reported in Foresight, judgmental overrides reduced accuracy by about 2% at Newell Rubbermaid, while statistical forecasting added about 5% over naive. Other settings differ, so each company should measure its own overrides.
What are the limits of FVA?
Critics argue it shows correlation rather than causation, ignores uncertainty and cost, and can be gamed. It is best used as a periodic reality check alongside financial measures such as markdown cost and lost sales.
One edition every weekday morning. Read in five minutes. Free for industry professionals.
SOURCES
- Foresight: Forecast Value Added: A Reality Check on Forecast Accuracy
- Doherty (Lokad): A Critical Evaluation of Forecast Value Added, Foresight 2024
- Hyndman and Athanasopoulos: Forecasting: Principles and Practice, Some simple forecasting methods
- Hyndman and Athanasopoulos: Forecasting: Principles and Practice, Evaluating forecast accuracy




