Believable isn't accurate. Anyone can render a synthetic shopper that looks plausible — the hard part is proving it predicts the real sale.
These are our working notes on how we do it: fit the simulation to real transactions, then test it on sales it never saw, and report where it stops holding.
Internal working notes and reproducible benchmark reports — the case for calibration, and the evidence behind it.
Why synthetic-shopper simulation needs a calibration layer before anyone should trust it with real money. A market that looks right can still land nowhere near the truth — and looking right is exactly what makes that dangerous.
Request the report →Fitting the agent simulation to real transactions cuts market-share error from a raw ~26× miss down to near-exact. On sales the model never saw, held-out R²=0.64 — a blind test, not a curve fit to the past.
Request the report →How we recover true market shares on sales the model was never shown. Validated on two independent public datasets — Nevo (ready-to-eat cereal) and dunnhumby (frozen pizza) — with share recovery moving from −0.35 to +0.44.
Request the report →Why calibration to real transactions compounds into a moat that persona-scale alone can't buy. More real sales tighten the fit; a tighter fit earns more trust and more data. The advantage is the loop, not the model.
Request the report →Raw agent simulation looks believable but lands far from the truth. The calibration layer is what closes the gap between “looks right” and “is right” — and we measure it on held-out data.
Bar length = how close each lands to the true market share. Watch calibration snap onto reality.
Every number here can be checked. The full proof — the tests, the honest error ranges, and the categories where it stops working — is in the report. Request the report →
On held-out markets from three real supermarket datasets, an actual LLM multi-agent marketplace looks plausible but barely recovers real market shares (correlation r ≈ 0.07–0.22). Our calibration lands closest to reality every time (r = 0.44–0.65) — cutting the distance to real by 1.6–5.2×.
Cereal (Nevo) · frozen pizza & oral care (dunnhumby) · held-out / out-of-sample markets · LLM agents = gpt-4o-mini · distance-to-real = Jensen–Shannon divergence.
Four steps between “a market that looks real” and “a market you can bet on.”
Thousands of simulated shoppers weigh brand, price and features and pick a product — a believable market, but on its own only a starting guess.
A BLP-style demand fit tunes the simulation to real transactions, so its shares and price sensitivities match the actual market — not just a plausible one.
We score the calibrated model on sales it never saw and recover true market shares — the blind test that separates accuracy from overfitting.
Every answer ships with honest error ranges and the categories where it stops holding. Directionally decision-grade — never dollar-exact, and we say so.
Working notes and reproducible benchmark reports from the Avanti team. Full report available under NDA on request.
We'll send the reproducible benchmark — the held-out tests, the honest error ranges, and the categories where the model stops working — and run the live simulator on your own category. Available under NDA on request.
Request the full benchmark report →