Deep diveRetail

Measuring retail demand forecast accuracy at store-SKU level

Retail demand forecasting accuracy should be judged at the level where a decision is made, with measures that survive zero sales and expose bias, and ultimately by what the forecast does to availability, waste and stock once replenishment rules act on it. A lower error score that leaves shelves emptier or bins fuller is not an improvement. This page explains the measures, the data traps and an evaluation method that links them to outcomes.

Reviewed 7 min read

On this page
  1. Why one MAPE figure misleads at store-SKU-day level
  2. Accuracy measures and terms retail planners use
  3. Matching forecast grain, horizon and measure to each decision
  4. From raw POS history to a forecast you can score honestly
  5. Slow movers, empty shelves, promotions and new items
  6. Choosing an evaluation approach for each item segment
  7. A bakery category where lower WAPE meant more waste
  8. Evaluation mistakes that flatter a retail forecast
  9. Questions and answers
  10. Sources

Why one MAPE figure misleads at store-SKU-day level

Most store-SKU-day series are sparse. A typical store sells many items only a few times a week, so a large share of days record zero sales. Mean absolute percentage error divides by the actual value, which makes it undefined on zero days and explosive on days with one or two units1. Teams then drop the zeros or cap the error, and the metric quietly stops describing the assortment.

Percentage errors are also asymmetric: they punish over-forecasts and under-forecasts differently1, which pushes models toward under-forecasting. In grocery or fashion, a habit of under-forecasting shows up as empty shelves, the most expensive error a retailer can make.

The fix is not a better single number. It is a small set of measures, each answering one question, computed at the level where someone acts on the forecast.

Accuracy measures and terms retail planners use

WAPE
Weighted absolute percentage error: total absolute error divided by total actual sales across a set of items and days. It weights fast sellers more, tolerates zero days and reads as a share of volume.
Bias
The signed sum of errors over a period divided by actual sales. A positive value means persistent over-forecasting (excess stock); negative means under-forecasting (lost sales). Track it separately from WAPE, which hides direction.
MASE
Mean absolute scaled error: the model's error divided by the in-sample error of a naive or seasonal naive forecast1. Below one means the model beats that simple benchmark; it works on series with zeros.
Quantile (pinball) loss
Scores a forecast of a specific quantile, such as the level demand stays under nine days in ten2. It is the right measure when the forecast feeds an order-up-to level.
Censored demand
Recorded sales capped by available stock. When the shelf is empty, sales of zero do not mean demand of zero.
Forecast value added
The change in error contributed by each step of the process, such as a planner override or a promotion uplift, measured against the step before it.

Matching forecast grain, horizon and measure to each decision

One forecast rarely serves every decision. Score each use at its own level.

DecisionGrainHorizonPrimary measureWatch for
Store replenishmentStore-SKU-dayLead time plus review periodQuantile loss at the target service levelCensored history from past stockouts
Fresh and short-life orderingStore-SKU-day, sometimes intradayOne to a few daysQuantile loss, plus waste in simulationOver-forecast cost as high as under-forecast
Distribution center allocationDC-SKU-weekSupplier lead timeWAPE and biasStore-level noise averaged away
Promotion planningSKU-region-eventWeeks before the eventBias on promoted linesCannibalization of related items
Store labor schedulingStore-hour, transactions or unitsOne to several weeksWAPE by hour bandClick-and-collect picking hidden in units

Horizons are illustrative ranges; set yours from actual supplier lead times and review cycles.

From raw POS history to a forecast you can score honestly

01POS and stock history02Flag and correctstockouts03Add demand drivers04Probabilistic forecast05Reconcile the hierarchy06Simulate replenishment
  1. POS and stock history

    Daily sales, on-hand inventory, deliveries and returns by store and item.

  2. Flag and correct stockouts

    Mark days where stock hit zero or availability was low and estimate the demand that went unrecorded.

  3. Add demand drivers

    Price, promotion mechanics, display, holidays, local events and weather as explicit features.

  4. Probabilistic forecast

    Predict a distribution or several quantiles rather than a single point.

  5. Reconcile the hierarchy

    Make store, region, category and total forecasts add up coherently.

  6. Simulate replenishment

    Replay ordering rules on held-out weeks to estimate availability, waste and stock.

Conceptual evaluation pipeline for store-level forecasts. It shows the order of steps, not a specific product or a measured result.

Slow movers, empty shelves, promotions and new items

Intermittent demand. For items that sell irregularly, Croston-type methods forecast the size of a sale and the interval between sales separately; the original method is known to be biased and gives no prediction intervals3. Many retailers now forecast a full distribution for these items and set stock from a quantile, which treats the zeros as information rather than noise.

Censored demand. If an item sold out by midday, the recorded sales are a floor. Training on that history teaches the model to expect low demand and keeps the shelf thin. Use on-hand and delivery data to flag stockout days, then either exclude them from scoring or estimate lost demand from intraday sales patterns and similar stores.

Promotions. Promoted weeks need their own evaluation. Track the uplift on the promoted item, cannibalization of substitutes, halo on complements and the post-promotion dip. A model that nails promoted volume but ignores the dip will overstock the week after.

New products. With no history, forecast from analogues: items with similar attributes, price points and launch support. Score these launches separately for their first weeks, because blending them with mature items hides how poorly the cold start performs.

Hierarchy. Forecasts built at different levels rarely add up. Reconciliation methods, from bottom-up and top-down to optimal combination across levels, make them coherent and often improve accuracy at both ends4.

Choosing an evaluation approach for each item segment

  • If

    Fast-moving staples that sell most days in most stores.

    Then

    Track WAPE and bias weekly by category, with quantile loss at the replenishment service level.

    Volume makes errors stable and comparable, and bias drives most excess stock and lost sales here.

  • If

    Slow or intermittent items, including long-tail ecommerce lines.

    Then

    Use MASE and quantile loss; judge the stocking quantile, not the mean.

    Percentage measures break on zeros, and the business question is how much to hold, not the expected sale.

  • If

    Fresh, bakery or short-life lines.

    Then

    Score by simulated waste and availability together, using the item's actual order cycle.

    Error measures treat over- and under-forecasting alike, while the costs of waste and empty shelves differ by item.

  • If

    Promoted lines and new launches.

    Then

    Evaluate on their own, event by event, and report bias before error.

    Blended with base weeks, their large errors are averaged away and repeat in the next event.

A bakery category where lower WAPE meant more waste

Evaluation mistakes that flatter a retail forecast

Scoring on data the model has seen

Early signalBack-test error far below live error in the first weeks.

MitigationUse rolling-origin evaluation: train up to a date, forecast the next horizon, roll forward and repeat.

Leaking future information

Early signalFeatures such as final promotion volumes or end-of-day stock appear in training rows.

MitigationBuild every feature as it was known at forecast time, including the promotion plan as it stood then.

Aggregating away the problem

Early signalCategory accuracy looks good while store complaints about gaps rise.

MitigationReport at the decision grain and segment by store size, region and velocity band.

Counting overrides as model error

Early signalNobody can say whether planner adjustments help.

MitigationStore the statistical forecast and each override separately and measure forecast value added.

Questions and answers

What forecast accuracy is good for a retailer?

There is no universal figure, because achievable accuracy depends on grain, horizon and how intermittent demand is. Compare against a naive or seasonal naive benchmark on the same items, which is what MASE does, and against your current method. Then ask whether the new forecast improves availability, waste or stock in simulation. A benchmark from another retailer with a different assortment tells you little.

Should retail forecasts be probabilistic?

For replenishment, usually yes. Ordering systems set stock to cover demand at a chosen service level, which is a quantile of the demand distribution, not its average. A probabilistic forecast supplies that quantile directly and expresses how uncertain each item is, which a point forecast plus a fixed safety factor approximates poorly for slow movers and promoted lines.

How often should retail forecasting models retrain?

Separate retraining from re-forecasting. Forecasts should refresh as often as ordering runs, often daily, using the latest sales, stock and promotion plans. Model parameters can retrain less often, for example weekly or monthly, with a scheduled check for drift after assortment changes, new store formats or shifts in shopping patterns. Retrain sooner when bias on recent weeks moves beyond an agreed tolerance.

Can we score a forecast on days when the item was out of stock?

Not at face value. On stockout days recorded sales understate demand, so scoring against them rewards forecasts that predict low sales. Either exclude flagged stockout days from the evaluation set or score against an estimate of unconstrained demand, and report how many days were affected so readers can judge the result.

Sources

  1. Forecasting: Principles and Practice (3rd ed), section 5.8: Evaluating point forecast accuracy — OTexts (Hyndman and Athanasopoulos) · checked 10 October 2026
  2. Forecasting: Principles and Practice (3rd ed), section 5.9: Evaluating distributional forecast accuracy — OTexts (Hyndman and Athanasopoulos) · checked 10 October 2026
  3. Forecasting: Principles and Practice (3rd ed), section 13.2: Time series of counts — OTexts (Hyndman and Athanasopoulos) · checked 10 October 2026
  4. Forecasting: Principles and Practice (3rd ed), chapter 11: Forecasting hierarchical and grouped time series — OTexts (Hyndman and Athanasopoulos) · checked 10 October 2026

More in Retail

Back to Retail

Next step

Get a second opinion on how your forecasts are scored

Send your current accuracy report and a description of how replenishment uses the forecast. We will reply with the measures and simulation we would add, and where stockout or promotion data may be distorting the picture.

Discuss forecast evaluation