00Research

The science of calibrated simulation.

Believable isn't accurate. Anyone can render a synthetic shopper that looks plausible — the hard part is proving it predicts the real sale.

These are our working notes on how we do it: fit the simulation to real transactions, then test it on sales it never saw, and report where it stops holding.

02Proof · believable ≠ accurate
We tested it on real sales it had never seen.

Raw agent simulation looks believable but lands far from the truth. The calibration layer is what closes the gap between “looks right” and “is right” — and we measure it on held-out data.

◇ Simulation vs real sales (held-out)

Bar length = how close each lands to the true market share. Watch calibration snap onto reality.

Raw sim
off by ~26×
Calibrated
≈ real
Real sales
truth
Raw agent simulation looks believable but lands far from the truth. Adding our calibration layer moves share-recovery from −0.35 to +0.44 — the gap between “looks right” and “is right.”
Held-out accuracy · R² (0–1)
0.64
How well we predict real market shares on sales the model never saw. 1 = perfect; strong for a blind test — and something no “synthetic shopper” tool publishes.
~26×
raw simulation error before the calibration layer
3
independent real supermarket datasets validated on
4 of 5×
directional calls we get right on real price moves
NEW  Benchmark · three real datasets
We ran a real LLM marketplace against reality — three times.

On held-out markets from three real supermarket datasets, an actual LLM multi-agent marketplace looks plausible but barely recovers real market shares (correlation r ≈ 0.07–0.22). Our calibration lands closest to reality every time (r = 0.44–0.65) — cutting the distance to real by 1.6–5.2×.

Across cereal (Nevo), frozen pizza and oral care (dunnhumby), raw LLM simulation lands far from real market shares while Avanti's calibration is closest, on held-out markets.

Cereal (Nevo) · frozen pizza & oral care (dunnhumby) · held-out / out-of-sample markets · LLM agents = gpt-4o-mini · distance-to-real = Jensen–Shannon divergence.

03Methodology
How a plausible simulation becomes an accurate one.

Four steps between “a market that looks real” and “a market you can bet on.”

01
Agent choice layer

Thousands of simulated shoppers weigh brand, price and features and pick a product — a believable market, but on its own only a starting guess.

02
Structural calibration

A BLP-style demand fit tunes the simulation to real transactions, so its shares and price sensitivities match the actual market — not just a plausible one.

03
Held-out validation

We score the calibrated model on sales it never saw and recover true market shares — the blind test that separates accuracy from overfitting.

04
Confidence & failure modes

Every answer ships with honest error ranges and the categories where it stops holding. Directionally decision-grade — never dollar-exact, and we say so.

04Research index
The full list.

Working notes and reproducible benchmark reports from the Avanti team. Full report available under NDA on request.

Position
Believable ≠ accurateWhy synthetic-shopper simulation needs a calibration layer to be trusted with real money.
Request
Benchmark
26× error reduction on held-out salesFitting the agent simulation to real transactions cuts share error from raw ~26× to near-exact; held-out R²=0.64.
Request
Method
Held-out validation on public supermarket dataRecovering true market shares on sales the model never saw; validated on Nevo (cereal) and dunnhumby (pizza); share recovery −0.35 → +0.44.
Request
Position
The data flywheelWhy calibration to real transactions compounds into a moat that persona-scale alone can't.
Request
05Get the evidence

Request the full benchmark report

We'll send the reproducible benchmark — the held-out tests, the honest error ranges, and the categories where the model stops working — and run the live simulator on your own category. Available under NDA on request.

Request the full benchmark report →