Accuracy & evaluation
What is “good” forecast accuracy? Benchmarks you can actually use
Good is beating a naive baseline on the same grain, horizon, and folds — by enough to change a decision. Industry round numbers are gossip until they match your SKU mix and your clock.
Updated Aug 16, 2026·8 min read
Beat the baseline first
Good is beating a naive rule on the same grain, horizon, and folds — by enough that a human will change a decision. If you cannot beat last-week-this-weekday, you are formatting, not forecasting. MASE makes that visible. So does putting the naive on the same chart.
Vendor numbers without a baseline, grain, or fold description are marketing. Ask for those four: grain, horizon, folds, baseline. If they stall, you already have the forecast of the relationship.
Typical WAPE ranges (gossip until your grain matches)
Grain changes everything
Chain-week totals are smoother. SKU-store-day is a forest of zeros and spikes. Quoting “we are at 8%” without the grain is how two teams think they agree. Energy hourly load can sit in the low single digits and still be a serious system. Grocery daily SKU can sit at 30%+ and still beat last year. Both can be “good.”
Ranges, not trophies
The bands in the figure are gossip with a source, not a contract. Published competitions and production write-ups cluster there. Your mix (promos, intermittence, data delay) will sit somewhere else. Use them to sanity-check a slide, not to write a bonus plan.
Good enough to decide
A 14% WAPE that buyers trust and refresh weekly beats an 8% WAPE that lives in a deck. Adoption is part of accuracy. So is bias: a slightly worse WAPE that is unbiased can cost less than a pretty score that is always heavy on the perishable aisle.
FAQ
- What is a good WAPE for retail demand?
- Weekly category or chain totals often land in the high single digits to low teens. Daily SKU-store can sit at 30%+ and still beat last-year-this-week. Quote the grain when you quote the number.
- Why do vendor benchmarks look better than ours?
- They may score aggregates, drop hard SKUs, or report in-sample fit. Ask for the grain, the horizon, the folds, and the baseline. If those four are missing, the number is marketing.
- When is a 'worse' forecast still worth using?
- When it is less wrong than the process it replaces, and when people will actually consume it. A 14% WAPE that buyers trust beats an 8% WAPE sitting in a slide deck.
Keep going
Guide
WAPE, explained
WAPE (weighted absolute percentage error) is total absolute error divided by total actuals. It behaves when some periods are zero, and it does not let tiny SKUs dominate the score the way MAPE does.
Guide
Choosing a forecast accuracy metric
Pick the metric that matches the cost of being wrong. WAPE for volume, MAE when units are the pain, RMSE when big misses hurt more than small ones, bias when you always run heavy or light.
Guide
How to backtest a forecasting model
Backtesting asks the model to forecast a past window it was not trained on, then scores the miss. Walk the window forward so you see many 'futures,' not one lucky test month.
Use case
Retail demand forecasting
Anticipate demand by product, store, channel, and region before buying or allocation decisions are locked.
Use case
Energy load forecasting
Forecast demand by interval, feeder, zone, and customer class with weather-driven uncertainty.