usecaseinai

Use cases / Supply chain / CPG

20% less inventory: how AB InBev makes 85% of its demand plans touchless

Everyone agrees forecasting is useful. Almost nobody can name who's actually leveraged it, or what it returned. AB InBev can — and the interesting part isn't the model, it's where they drew the line between "trust it" and "call a human."

17 Sept 20265 mindemand-forecasting · supply-chain · cpgSource: usecaseinai analysis. Operational figures are o9 Solutions' AB InBev case study (vendor-reported, marked in-text). Independent ranges are from McKinsey operations research. (Swap in the Substack long-form URL once live.)

The shape of the problem

Ask any supply-chain professional whether demand forecasting is useful and you'll get a yes in one second. Ask them who has actually leveraged it in production and what business value it returned, and the room goes quiet. That gap is the whole problem. Forecasting is the most agreed-upon, least-quantified use case in the industry — and worse, the people buying it are usually buying the wrong promise. They want accuracy. What a forecast actually sells is a bounded, probabilistic estimate of an uncertain future. Those are not the same product, and confusing them is what turns a working system into a disappointed client.

What the system does

INGEST

sales history, POS, distributor sell-through, promotions, price, weather, calendar events

ENRICH

correct censored demand, build lag/seasonality/promo features

FORECAST

ensemble → probabilistic forecast at SKU-location-week

GATE

auto-commit only above the confidence threshold

ROUTE

everything else goes to a human demand planner

A demand signal comes in, gets cleaned and enriched, and an ensemble produces not a single number but a distribution — expected demand plus a range. A confidence threshold decides what's safe to commit automatically. Everything the system isn't sure about stays on a planner's desk. AB InBev reports that roughly 85% of its US demand plans now run touchless this way (vendor-reported, via o9). The pattern is identical to confidence-calibrated email triage: don't trust the model everywhere — only where it's earned it.

It's a range, not a promise

This is the line most buyers never internalize. A forecast that says "+5%" is not a commitment to 5%. It's "roughly +2% to +5%, at this confidence," and the job of a decision-maker is to take that band and adjust it with domain knowledge and a read on the trend. The model's entire purpose is to quantify uncertainty, not remove it. When a client or an executive demands a single accurate number from a statistical model, they're asking the tool to be something it never claimed to be — and when reality lands inside the range but off the point estimate, they call it a failure. It wasn't. It did exactly what it said.

The gating eval: touchless only where it's earned

The number that matters here isn't average accuracy, it's how much of the plan you can safely stop touching. AB InBev pushed forecast accuracy past 87% and inventory down about 20% (both vendor-reported), but the operational unlock was the confidence gate: the stable, high-signal SKUs commit themselves, and planners spend their week on the exceptions where human judgment actually adds value. The discipline that keeps this honest is Forecast Value Added — you measure whether each human override actually beat the model. If a touchpoint doesn't add value, you remove it. Most planning teams have never measured this, which is why most planning teams override far more than they should.

Where it silently breaks

Three failure modes account for most of the disappointment, and none of them are about the model being "dumb."

  • Censored demand. A stock-out gets recorded as zero demand. The model learns nobody wanted the item, under-forecasts, and stocks you out again next cycle. The forecast wasn't wrong — the data lied to it. This is the single most common quiet killer, and only people who've shipped one think to look for it.
  • Silent drift. A model trained on last year's patterns doesn't break loudly when the world shifts — it goes stale. It keeps producing confident, plausible, wrong numbers. Quarterly retraining feels responsible right up until conditions change in a week.
  • The override trap. Studies of tens of thousands of real forecasts found that roughly half of human overrides made the forecast worse, and in some organizations planners override the system up to 80% of the time. So the fix isn't "trust the humans" — but it isn't "trust the machine" either. It's Forecast Value Added: keep the human only where the human demonstrably wins.

What to actually monitor

  • WMAPE and bias, per segment — not one global accuracy number; new products and promos behave nothing like stable base demand.
  • Forecast Value Added, per touchpoint — is each override, and each model, beating a naive baseline?
  • Touchless rate — the share of the plan committed without a human. This is the real efficiency metric.
  • Drift — distribution shift on inputs and errors; trigger retraining on change, not on the calendar.
  • Business outcomes — service level and inventory turns, because MAPE is not what finance signs off on.

Why this isn't really a modeling story

AB InBev's edge isn't a smarter algorithm — the algorithms are largely commodity. The edge is architectural and organizational: clean data that isn't lying about censored demand, a confidence line that decides what's safe to automate, Forecast Value Added to keep humans only where they help, and — most underrated — expectation management. The independent benchmark backs the upside: McKinsey finds AI-driven forecasting cuts errors 20–50%, reduces lost sales and unavailability by up to 65%, and pulls inventory costs down 10–15%. Those numbers are real and repeatable. They just aren't a promise of certainty. The companies that win here didn't buy a better crystal ball. They understood exactly what they'd bought: a statistical model that delivers based on the data you feed it — and nothing more.