Skip to content
Predict.ai

Accuracy & evaluation

What is “good” forecast accuracy? Benchmarks you can actually use

Good is beating a naive baseline on the same grain, horizon, and folds — by enough to change a decision. Industry round numbers are gossip until they match your SKU mix and your clock.

Updated Aug 16, 2026·8 min read

Beat the baseline first

Good is beating a naive rule on the same grain, horizon, and folds — by enough that a human will change a decision. If you cannot beat last-week-this-weekday, you are formatting, not forecasting. MASE makes that visible. So does putting the naive on the same chart.

Vendor numbers without a baseline, grain, or fold description are marketing. Ask for those four: grain, horizon, folds, baseline. If they stall, you already have the forecast of the relationship.

Typical WAPE ranges (gossip until your grain matches)

Typical WAPE ranges by grainChain · weekly612%Category · weekly816%SKU-store · daily2245%Grid load · hourly26%Web traffic · daily718%
These are order-of-magnitude bands from published bake-offs and production teams, not trophies. Always print vs a naive baseline on the same folds.

Grain changes everything

Chain-week totals are smoother. SKU-store-day is a forest of zeros and spikes. Quoting “we are at 8%” without the grain is how two teams think they agree. Energy hourly load can sit in the low single digits and still be a serious system. Grocery daily SKU can sit at 30%+ and still beat last year. Both can be “good.”

Ranges, not trophies

The bands in the figure are gossip with a source, not a contract. Published competitions and production write-ups cluster there. Your mix (promos, intermittence, data delay) will sit somewhere else. Use them to sanity-check a slide, not to write a bonus plan.

Good enough to decide

A 14% WAPE that buyers trust and refresh weekly beats an 8% WAPE that lives in a deck. Adoption is part of accuracy. So is bias: a slightly worse WAPE that is unbiased can cost less than a pretty score that is always heavy on the perishable aisle.

FAQ

What is a good WAPE for retail demand?
Weekly category or chain totals often land in the high single digits to low teens. Daily SKU-store can sit at 30%+ and still beat last-year-this-week. Quote the grain when you quote the number.
Why do vendor benchmarks look better than ours?
They may score aggregates, drop hard SKUs, or report in-sample fit. Ask for the grain, the horizon, the folds, and the baseline. If those four are missing, the number is marketing.
When is a 'worse' forecast still worth using?
When it is less wrong than the process it replaces, and when people will actually consume it. A 14% WAPE that buyers trust beats an 8% WAPE sitting in a slide deck.

Keep going

Try it on your data.

Connect an outcome and its history. Predict.ai finds the drivers, races the models, and keeps the forecast live — in the workspace, over the API, and through agents.