How close we get to real people
Scored against 64 held out questions with real answers. Every number on this page comes from the run named above.
By question family
NDAM in bold| Family | Questions | 1-MAE | NDAM | Ceiling | Ungrounded |
|---|---|---|---|---|---|
| issue mix | 64 | 87.8% | 33.2% | 82.3% | 37.9% |
By audience
NDAM in bold| Audience | Questions | 1-MAE | NDAM | Ceiling | Ungrounded |
|---|---|---|---|---|---|
| pub:airbnb | 1 | 85.1% | 18.1% | 80.5% | 28.7% |
| pub:alltrails | 1 | 84.0% | 12.2% | 79.5% | 33.4% |
| pub:bbc | 1 | 84.4% | 14.1% | 86.1% | 16.8% |
| pub:booking_com | 1 | 91.0% | 50.6% | 83.8% | 15.9% |
| pub:brave | 1 | 82.7% | 4.7% | 86.7% | 44.6% |
| pub:bumble_dating_app | 1 | 88.6% | 37.4% | 86.1% | 45.7% |
| pub:cash_app | 1 | 93.3% | 63.2% | 83.1% | 28.6% |
| pub:chatgpt | 1 | 84.8% | 16.2% | 66.5% | 55.7% |
| pub:chime | 1 | 93.6% | 64.7% | 83.5% | 20.3% |
| pub:chrome | 1 | 83.0% | 6.6% | 91.2% | 32.8% |
| pub:claude | 1 | 85.4% | 19.6% | 79.7% | 48.7% |
| pub:cnn | 1 | 87.9% | 33.5% | 70.1% | 20.4% |
| pub:coinbase | 1 | 91.8% | 54.9% | 81.8% | 38.8% |
| pub:deliveroo | 1 | 86.9% | 28.1% | 85.5% | 32.9% |
| pub:discord | 1 | 83.5% | 9.2% | 85.2% | 40.6% |
| pub:disney | 1 | 82.9% | 6.1% | 87.6% | 29.4% |
| pub:doordash | 1 | 88.7% | 37.9% | 85.0% | 32.8% |
| pub:duckduckgo | 1 | 85.7% | 21.6% | 78.7% | 24.8% |
| pub:edge | 1 | 85.6% | 20.9% | 83.0% | 52.1% |
| pub:evernote | 1 | 82.7% | 4.9% | 77.1% | 24.2% |
| pub:expedia | 1 | 83.8% | 11.0% | 81.5% | 26.0% |
| pub:firefox | 1 | 89.9% | 44.5% | 80.6% | 23.7% |
| pub:fitbit | 1 | 88.8% | 38.1% | 89.5% | 60.5% |
| pub:fox | 1 | 91.1% | 51.3% | 88.4% | 51.3% |
| pub:garmin_connect | 1 | 89.2% | 40.8% | 85.3% | 51.4% |
| pub:grubhub | 1 | 88.9% | 38.8% | 85.8% | 22.9% |
| pub:guardian | 1 | 83.0% | 6.7% | 67.0% | 50.5% |
| pub:hinge_dating_app | 1 | 85.5% | 20.2% | 85.0% | 46.3% |
| pub:hopper | 1 | 89.8% | 43.7% | 71.1% | 13.0% |
| pub:hulu | 1 | 92.6% | 59.3% | 86.2% | 65.3% |
| pub:instacart | 1 | 92.9% | 60.9% | 83.3% | 24.0% |
| pub:just_eat | 1 | 92.5% | 58.7% | 85.1% | 18.9% |
| pub:lyft | 1 | 88.0% | 34.0% | 83.8% | 35.2% |
| pub:match | 1 | 84.8% | 16.4% | 86.2% | 45.2% |
| pub:max | 1 | 86.0% | 22.7% | 84.0% | 45.7% |
| pub:messenger | 1 | 89.5% | 42.0% | 84.1% | 33.1% |
| pub:microsoft_teams | 1 | 87.5% | 31.5% | 85.5% | 54.5% |
| pub:myfitnesspal | 1 | 91.3% | 52.2% | 87.4% | 55.6% |
| pub:netflix | 1 | 91.1% | 51.0% | 85.0% | 32.0% |
| pub:nike_run_club | 1 | 84.7% | 15.8% | 84.7% | 38.7% |
| pub:notion | 1 | 86.5% | 25.8% | 81.4% | 24.5% |
| pub:nytimes | 1 | 84.4% | 14.2% | 82.1% | 55.0% |
| pub:okcupid_dating | 1 | 83.5% | 9.5% | 84.4% | 51.5% |
| pub:opera | 1 | 87.8% | 32.7% | 84.1% | 55.0% |
| pub:perplexity | 1 | 90.0% | 45.0% | 67.2% | 47.9% |
| pub:plenty_of_fish | 1 | 89.0% | 39.4% | 85.3% | 9.0% |
| pub:robinhood | 1 | 88.3% | 35.8% | 79.7% | 45.9% |
| pub:runna | 1 | 87.4% | 30.6% | 76.4% | 60.5% |
| pub:samsung_internet | 1 | 91.0% | 50.4% | 84.9% | 40.9% |
| pub:signal | 1 | 84.6% | 15.3% | 80.8% | 38.9% |
| pub:slack | 1 | 93.8% | 66.0% | 77.2% | 71.0% |
| pub:spotify | 1 | 89.9% | 44.2% | 82.9% | 18.0% |
| pub:strava | 1 | 87.7% | 32.2% | 84.6% | 59.6% |
| pub:telegram_messenger | 1 | 85.8% | 22.1% | 86.3% | 12.9% |
| pub:tinder_dating_app | 1 | 90.0% | 45.1% | 85.6% | 37.3% |
| pub:todoist | 1 | 90.1% | 45.6% | 64.6% | 46.6% |
| pub:uber | 1 | 88.4% | 36.3% | 82.7% | 32.5% |
| pub:uber_eats | 1 | 90.6% | 48.2% | 87.3% | 24.2% |
| pub:venmo | 1 | 91.5% | 53.3% | 85.8% | 19.4% |
| pub:viber | 1 | 82.3% | 2.5% | 86.5% | 36.3% |
| pub:vivaldi | 1 | 90.0% | 45.2% | 81.2% | 47.8% |
| pub:whatsapp_messenger | 1 | 86.9% | 28.1% | 86.2% | 30.1% |
| pub:ynab | 1 | 93.8% | 66.0% | 74.7% | 59.2% |
| pub:youtube_music | 1 | 86.8% | 27.2% | 84.7% | 39.1% |
Published, reproducible, and honest about the gaps.
A synthetic audience that cannot be checked is a rumour with a chart on it. This page explains exactly what we measure, how we hold data back so the measurement means something, what the baselines are, and where the method does not work.
Two numbers, one forgiving and one not.
One minus mean absolute error
For a question with options, we compare the share the synthetic audience gives each option with the share the held out real data gives it, take the mean absolute error across options, and subtract it from one. Higher is better. It is the number most widely reported in this space, which makes it the comparable one, but it is forgiving: averaging over many options dilutes a single bad miss.
Normalised deviation from actual maximum
A stricter read. It scores the error against the worst error that was possible for that question, so a miss on a question where the options were close is penalised far harder than a miss on a question where almost any answer was nearly right. We report it alongside 1-MAE precisely because it is less flattering.
How we make the measurement mean something.
Split the data by time, not at random
Personas are built only from reviews up to a cut off date. The test set is reviews written after it. A random split leaks the answer, because the same reviewer's vocabulary appears on both sides.
Turn held out reviews into questions with known answers
From the held out window we derive distributions we can actually check, for example the share of reviews raising a given theme, or the rating distribution for a given topic. The synthetic audience is asked the same thing without seeing that window.
Add external published surveys where they exist
Where a credible published survey overlaps one of our categories, we use its reported distribution as a second test set. This guards against the failure mode where we are only ever good at predicting more reviews.
Score against baselines, not against nothing
A number on its own is unreadable. We report a uniform baseline, a corpus prior baseline (guess the overall distribution every time), and an ungrounded LLM baseline (the same model asked the same question with no persona and no review grounding). The gap between userken and the ungrounded baseline is the part that our work is responsible for.
The ceiling is not 100%
Real humans do not agree with themselves. We estimate the ceiling by split half agreement: split the held out real data in two at random, and score one half against the other with the same metric. Whatever that comes out as is the best score any method could get on this task, because it is the level at which the truth disagrees with itself. Scores should be read as a share of that ceiling, not as a percentage of perfect.
The honest list.
Every method has a shape of problem it is wrong about. Here is ours, so you can avoid using this where it will let you down.
Thin audiences
A persona cluster built from a handful of reviews is a guess with a name on it. Below our minimum cluster size we refuse to build the audience rather than ship a confident sounding fiction. Narrow filters are the most common way to get a thin audience by accident.
Reviewers are not the general public
People who write app reviews are self selected: more motivated, more annoyed, and more likely to have hit a specific problem than the average user. A userken audience represents the reviewing population of a category, which is useful and is not the same thing as a representative population sample. Read differences between audiences and movements over time rather than absolute levels.
Regulated and high stakes claims
Do not use this for medical, financial, legal or safety claims, for anything that will be published as a substantiated statistic, or for any claim that needs a defensible methodology in front of a regulator. Use a real sample.
Genuinely new things
An audience built from text about what exists cannot reliably tell you about a category that does not exist yet. The further a concept is from anything the corpus has opinions about, the more the answer is the model's prior rather than the audience's view.
Small differences
If two options come back two points apart, that is a tie. Treat the method as a way to find large effects and eliminate bad options, not as a way to pick between close ones.
Languages and markets we have not indexed
Coverage is what it is, and /sources lists it. An audience cannot represent a market whose reviews are not in the index.
The long version.
The full write up covers the corpus, the cut off dates, how questions are derived from held out data, the exact baseline definitions, and the split half ceiling calculation.
Hold us to this page.
If the numbers here are not good enough for your decision, do not use a synthetic audience for it. That is the point of publishing them.