Accuracy · run 153 · 2026-09-11

How close we get to real people

Scored against 64 held out questions with real answers. Every number on this page comes from the run named above.

Methodology
NDAM · our stricter measure
33.2%
distribution match, all option counts
1-MAE · industry standard
87.8%
inflates with more options, shown for comparison
Human ceiling
82.3%
split half agreement between real reviewers
Ungrounded LLM
37.9%
same model, no audience, no reviews
Holdout: time split. test window from 2026-06-13. 205,256 held out reviews. 204,987 in the training window. run 153. 64 of 64 questions scored. grounded model claude-haiku-4-5 (fast tier). cost $2.99 over 2586 LLM calls. 3 samples per persona. corpus prior NDAM 79.9%. uniform baseline NDAM 58.1%. Every question counts equally. NDAM is 1 minus total variation distance, 1-MAE is 1 minus mean absolute error over option shares.
Time split: grounding retrieval is restricted to the train window, but the personas in use were built on all reviews (run with rebuild_personas for the leakage-safe number).
issue_mix first run: panel vs corpus prior on the extracted complaint mix

By question family

NDAM in bold
FamilyQuestions1-MAENDAMCeilingUngrounded
issue mix6487.8%33.2%82.3%37.9%
How we test
01
Split
Reviews up to a cut off date build the audiences. Everything written after it is held out.
02
Ask
Each held out question is put to the audience the way a survey would ask it, with no sight of the test window.
03
Compare
The predicted distribution is scored against what real reviewers actually said, next to every baseline.
04
Publish
Every run is stored and this page updates from it. Nothing is cherry picked and no number is typed in by hand.

By audience

NDAM in bold
AudienceQuestions1-MAENDAMCeilingUngrounded
pub:airbnb185.1%18.1%80.5%28.7%
pub:alltrails184.0%12.2%79.5%33.4%
pub:bbc184.4%14.1%86.1%16.8%
pub:booking_com191.0%50.6%83.8%15.9%
pub:brave182.7%4.7%86.7%44.6%
pub:bumble_dating_app188.6%37.4%86.1%45.7%
pub:cash_app193.3%63.2%83.1%28.6%
pub:chatgpt184.8%16.2%66.5%55.7%
pub:chime193.6%64.7%83.5%20.3%
pub:chrome183.0%6.6%91.2%32.8%
pub:claude185.4%19.6%79.7%48.7%
pub:cnn187.9%33.5%70.1%20.4%
pub:coinbase191.8%54.9%81.8%38.8%
pub:deliveroo186.9%28.1%85.5%32.9%
pub:discord183.5%9.2%85.2%40.6%
pub:disney182.9%6.1%87.6%29.4%
pub:doordash188.7%37.9%85.0%32.8%
pub:duckduckgo185.7%21.6%78.7%24.8%
pub:edge185.6%20.9%83.0%52.1%
pub:evernote182.7%4.9%77.1%24.2%
pub:expedia183.8%11.0%81.5%26.0%
pub:firefox189.9%44.5%80.6%23.7%
pub:fitbit188.8%38.1%89.5%60.5%
pub:fox191.1%51.3%88.4%51.3%
pub:garmin_connect189.2%40.8%85.3%51.4%
pub:grubhub188.9%38.8%85.8%22.9%
pub:guardian183.0%6.7%67.0%50.5%
pub:hinge_dating_app185.5%20.2%85.0%46.3%
pub:hopper189.8%43.7%71.1%13.0%
pub:hulu192.6%59.3%86.2%65.3%
pub:instacart192.9%60.9%83.3%24.0%
pub:just_eat192.5%58.7%85.1%18.9%
pub:lyft188.0%34.0%83.8%35.2%
pub:match184.8%16.4%86.2%45.2%
pub:max186.0%22.7%84.0%45.7%
pub:messenger189.5%42.0%84.1%33.1%
pub:microsoft_teams187.5%31.5%85.5%54.5%
pub:myfitnesspal191.3%52.2%87.4%55.6%
pub:netflix191.1%51.0%85.0%32.0%
pub:nike_run_club184.7%15.8%84.7%38.7%
pub:notion186.5%25.8%81.4%24.5%
pub:nytimes184.4%14.2%82.1%55.0%
pub:okcupid_dating183.5%9.5%84.4%51.5%
pub:opera187.8%32.7%84.1%55.0%
pub:perplexity190.0%45.0%67.2%47.9%
pub:plenty_of_fish189.0%39.4%85.3%9.0%
pub:robinhood188.3%35.8%79.7%45.9%
pub:runna187.4%30.6%76.4%60.5%
pub:samsung_internet191.0%50.4%84.9%40.9%
pub:signal184.6%15.3%80.8%38.9%
pub:slack193.8%66.0%77.2%71.0%
pub:spotify189.9%44.2%82.9%18.0%
pub:strava187.7%32.2%84.6%59.6%
pub:telegram_messenger185.8%22.1%86.3%12.9%
pub:tinder_dating_app190.0%45.1%85.6%37.3%
pub:todoist190.1%45.6%64.6%46.6%
pub:uber188.4%36.3%82.7%32.5%
pub:uber_eats190.6%48.2%87.3%24.2%
pub:venmo191.5%53.3%85.8%19.4%
pub:viber182.3%2.5%86.5%36.3%
pub:vivaldi190.0%45.2%81.2%47.8%
pub:whatsapp_messenger186.9%28.1%86.2%30.1%
pub:ynab193.8%66.0%74.7%59.2%
pub:youtube_music186.8%27.2%84.7%39.1%
Where it does not work
Thin audiences. Below our minimum cluster size we refuse to build personas and say so.
Claims that need real people. Regulated, legal or published claims stay with human respondents.
The general public. Reviewers are self selected, so read differences and movements rather than absolute levels.
Why this page exists

Published, reproducible, and honest about the gaps.

A synthetic audience that cannot be checked is a rumour with a chart on it. This page explains exactly what we measure, how we hold data back so the measurement means something, what the baselines are, and where the method does not work.

Metrics

Two numbers, one forgiving and one not.

1-MAE

One minus mean absolute error

For a question with options, we compare the share the synthetic audience gives each option with the share the held out real data gives it, take the mean absolute error across options, and subtract it from one. Higher is better. It is the number most widely reported in this space, which makes it the comparable one, but it is forgiving: averaging over many options dilutes a single bad miss.

NDAM

Normalised deviation from actual maximum

A stricter read. It scores the error against the worst error that was possible for that question, so a miss on a question where the options were close is penalised far harder than a miss on a question where almost any answer was nearly right. We report it alongside 1-MAE precisely because it is less flattering.

Hold out method

How we make the measurement mean something.

01 · Method

Split the data by time, not at random

Personas are built only from reviews up to a cut off date. The test set is reviews written after it. A random split leaks the answer, because the same reviewer's vocabulary appears on both sides.

02 · Method

Turn held out reviews into questions with known answers

From the held out window we derive distributions we can actually check, for example the share of reviews raising a given theme, or the rating distribution for a given topic. The synthetic audience is asked the same thing without seeing that window.

03 · Method

Add external published surveys where they exist

Where a credible published survey overlaps one of our categories, we use its reported distribution as a second test set. This guards against the failure mode where we are only ever good at predicting more reviews.

04 · Method

Score against baselines, not against nothing

A number on its own is unreadable. We report a uniform baseline, a corpus prior baseline (guess the overall distribution every time), and an ungrounded LLM baseline (the same model asked the same question with no persona and no review grounding). The gap between userken and the ungrounded baseline is the part that our work is responsible for.

The human ceiling

The ceiling is not 100%

Real humans do not agree with themselves. We estimate the ceiling by split half agreement: split the held out real data in two at random, and score one half against the other with the same metric. Whatever that comes out as is the best score any method could get on this task, because it is the level at which the truth disagrees with itself. Scores should be read as a share of that ceiling, not as a percentage of perfect.

Where it does not work

The honest list.

Every method has a shape of problem it is wrong about. Here is ours, so you can avoid using this where it will let you down.

Thin audiences

A persona cluster built from a handful of reviews is a guess with a name on it. Below our minimum cluster size we refuse to build the audience rather than ship a confident sounding fiction. Narrow filters are the most common way to get a thin audience by accident.

Reviewers are not the general public

People who write app reviews are self selected: more motivated, more annoyed, and more likely to have hit a specific problem than the average user. A userken audience represents the reviewing population of a category, which is useful and is not the same thing as a representative population sample. Read differences between audiences and movements over time rather than absolute levels.

Regulated and high stakes claims

Do not use this for medical, financial, legal or safety claims, for anything that will be published as a substantiated statistic, or for any claim that needs a defensible methodology in front of a regulator. Use a real sample.

Genuinely new things

An audience built from text about what exists cannot reliably tell you about a category that does not exist yet. The further a concept is from anything the corpus has opinions about, the more the answer is the model's prior rather than the audience's view.

Small differences

If two options come back two points apart, that is a tie. Treat the method as a way to find large effects and eliminate bad options, not as a way to pick between close ones.

Languages and markets we have not indexed

Coverage is what it is, and /sources lists it. An audience cannot represent a market whose reviews are not in the index.

Methodology

The long version.

The full write up covers the corpus, the cut off dates, how questions are derived from held out data, the exact baseline definitions, and the split half ceiling calculation.

Hold us to this page.

If the numbers here are not good enough for your decision, do not use a synthetic audience for it. That is the point of publishing them.