Deep diveRetail
Measuring retail demand forecast accuracy at store-SKU level
Retail demand forecasting accuracy should be judged at the level where a decision is made, with measures that survive zero sales and expose bias, and ultimately by what the forecast does to availability, waste and stock once replenishment rules act on it. A lower error score that leaves shelves emptier or bins fuller is not an improvement. This page explains the measures, the data traps and an evaluation method that links them to outcomes.
On this page
- Why one MAPE figure misleads at store-SKU-day level
- Accuracy measures and terms retail planners use
- Matching forecast grain, horizon and measure to each decision
- From raw POS history to a forecast you can score honestly
- Slow movers, empty shelves, promotions and new items
- Choosing an evaluation approach for each item segment
- A bakery category where lower WAPE meant more waste
- Evaluation mistakes that flatter a retail forecast
- Questions and answers
- Sources
Why one MAPE figure misleads at store-SKU-day level
Most store-SKU-day series are sparse. A typical store sells many items only a few times a week, so a large share of days record zero sales. Mean absolute percentage error divides by the actual value, which makes it undefined on zero days and explosive on days with one or two units1. Teams then drop the zeros or cap the error, and the metric quietly stops describing the assortment.
Percentage errors are also asymmetric: they punish over-forecasts and under-forecasts differently1, which pushes models toward under-forecasting. In grocery or fashion, a habit of under-forecasting shows up as empty shelves, the most expensive error a retailer can make.
The fix is not a better single number. It is a small set of measures, each answering one question, computed at the level where someone acts on the forecast.
Accuracy measures and terms retail planners use
- WAPE
- Weighted absolute percentage error: total absolute error divided by total actual sales across a set of items and days. It weights fast sellers more, tolerates zero days and reads as a share of volume.
- Bias
- The signed sum of errors over a period divided by actual sales. A positive value means persistent over-forecasting (excess stock); negative means under-forecasting (lost sales). Track it separately from WAPE, which hides direction.
- MASE
- Mean absolute scaled error: the model's error divided by the in-sample error of a naive or seasonal naive forecast1. Below one means the model beats that simple benchmark; it works on series with zeros.
- Quantile (pinball) loss
- Scores a forecast of a specific quantile, such as the level demand stays under nine days in ten2. It is the right measure when the forecast feeds an order-up-to level.
- Censored demand
- Recorded sales capped by available stock. When the shelf is empty, sales of zero do not mean demand of zero.
- Forecast value added
- The change in error contributed by each step of the process, such as a planner override or a promotion uplift, measured against the step before it.
Matching forecast grain, horizon and measure to each decision
One forecast rarely serves every decision. Score each use at its own level.
| Decision | Grain | Horizon | Primary measure | Watch for |
|---|---|---|---|---|
| Store replenishment | Store-SKU-day | Lead time plus review period | Quantile loss at the target service level | Censored history from past stockouts |
| Fresh and short-life ordering | Store-SKU-day, sometimes intraday | One to a few days | Quantile loss, plus waste in simulation | Over-forecast cost as high as under-forecast |
| Distribution center allocation | DC-SKU-week | Supplier lead time | WAPE and bias | Store-level noise averaged away |
| Promotion planning | SKU-region-event | Weeks before the event | Bias on promoted lines | Cannibalization of related items |
| Store labor scheduling | Store-hour, transactions or units | One to several weeks | WAPE by hour band | Click-and-collect picking hidden in units |
Horizons are illustrative ranges; set yours from actual supplier lead times and review cycles.
From raw POS history to a forecast you can score honestly
- POS and stock history
Daily sales, on-hand inventory, deliveries and returns by store and item.
- Flag and correct stockouts
Mark days where stock hit zero or availability was low and estimate the demand that went unrecorded.
- Add demand drivers
Price, promotion mechanics, display, holidays, local events and weather as explicit features.
- Probabilistic forecast
Predict a distribution or several quantiles rather than a single point.
- Reconcile the hierarchy
Make store, region, category and total forecasts add up coherently.
- Simulate replenishment
Replay ordering rules on held-out weeks to estimate availability, waste and stock.
Slow movers, empty shelves, promotions and new items
Intermittent demand. For items that sell irregularly, Croston-type methods forecast the size of a sale and the interval between sales separately; the original method is known to be biased and gives no prediction intervals3. Many retailers now forecast a full distribution for these items and set stock from a quantile, which treats the zeros as information rather than noise.
Censored demand. If an item sold out by midday, the recorded sales are a floor. Training on that history teaches the model to expect low demand and keeps the shelf thin. Use on-hand and delivery data to flag stockout days, then either exclude them from scoring or estimate lost demand from intraday sales patterns and similar stores.
Promotions. Promoted weeks need their own evaluation. Track the uplift on the promoted item, cannibalization of substitutes, halo on complements and the post-promotion dip. A model that nails promoted volume but ignores the dip will overstock the week after.
New products. With no history, forecast from analogues: items with similar attributes, price points and launch support. Score these launches separately for their first weeks, because blending them with mature items hides how poorly the cold start performs.
Hierarchy. Forecasts built at different levels rarely add up. Reconciliation methods, from bottom-up and top-down to optimal combination across levels, make them coherent and often improve accuracy at both ends4.
Choosing an evaluation approach for each item segment
- If
Fast-moving staples that sell most days in most stores.
ThenTrack WAPE and bias weekly by category, with quantile loss at the replenishment service level.
Volume makes errors stable and comparable, and bias drives most excess stock and lost sales here.
- If
Slow or intermittent items, including long-tail ecommerce lines.
ThenUse MASE and quantile loss; judge the stocking quantile, not the mean.
Percentage measures break on zeros, and the business question is how much to hold, not the expected sale.
- If
Fresh, bakery or short-life lines.
ThenScore by simulated waste and availability together, using the item's actual order cycle.
Error measures treat over- and under-forecasting alike, while the costs of waste and empty shelves differ by item.
- If
Promoted lines and new launches.
ThenEvaluate on their own, event by event, and report bias before error.
Blended with base weeks, their large errors are averaged away and repeat in the next event.
A bakery category where lower WAPE meant more waste
Evaluation mistakes that flatter a retail forecast
Scoring on data the model has seen
Early signalBack-test error far below live error in the first weeks.
MitigationUse rolling-origin evaluation: train up to a date, forecast the next horizon, roll forward and repeat.
Leaking future information
Early signalFeatures such as final promotion volumes or end-of-day stock appear in training rows.
MitigationBuild every feature as it was known at forecast time, including the promotion plan as it stood then.
Aggregating away the problem
Early signalCategory accuracy looks good while store complaints about gaps rise.
MitigationReport at the decision grain and segment by store size, region and velocity band.
Counting overrides as model error
Early signalNobody can say whether planner adjustments help.
MitigationStore the statistical forecast and each override separately and measure forecast value added.
Questions and answers
What forecast accuracy is good for a retailer?
There is no universal figure, because achievable accuracy depends on grain, horizon and how intermittent demand is. Compare against a naive or seasonal naive benchmark on the same items, which is what MASE does, and against your current method. Then ask whether the new forecast improves availability, waste or stock in simulation. A benchmark from another retailer with a different assortment tells you little.
Should retail forecasts be probabilistic?
For replenishment, usually yes. Ordering systems set stock to cover demand at a chosen service level, which is a quantile of the demand distribution, not its average. A probabilistic forecast supplies that quantile directly and expresses how uncertain each item is, which a point forecast plus a fixed safety factor approximates poorly for slow movers and promoted lines.
How often should retail forecasting models retrain?
Separate retraining from re-forecasting. Forecasts should refresh as often as ordering runs, often daily, using the latest sales, stock and promotion plans. Model parameters can retrain less often, for example weekly or monthly, with a scheduled check for drift after assortment changes, new store formats or shifts in shopping patterns. Retrain sooner when bias on recent weeks moves beyond an agreed tolerance.
Can we score a forecast on days when the item was out of stock?
Not at face value. On stockout days recorded sales understate demand, so scoring against them rewards forecasts that predict low sales. Either exclude flagged stockout days from the evaluation set or score against an estimate of unconstrained demand, and report how many days were affected so readers can judge the result.
Sources
- Forecasting: Principles and Practice (3rd ed), section 5.8: Evaluating point forecast accuracy — OTexts (Hyndman and Athanasopoulos) · checked 10 October 2026
- Forecasting: Principles and Practice (3rd ed), section 5.9: Evaluating distributional forecast accuracy — OTexts (Hyndman and Athanasopoulos) · checked 10 October 2026
- Forecasting: Principles and Practice (3rd ed), section 13.2: Time series of counts — OTexts (Hyndman and Athanasopoulos) · checked 10 October 2026
- Forecasting: Principles and Practice (3rd ed), chapter 11: Forecasting hierarchical and grouped time series — OTexts (Hyndman and Athanasopoulos) · checked 10 October 2026