No Instrument Is Valid Until It Is Joined
Four ways of knowing a senior — synthetic agents, stated intent, a learned in-call model, a grounding label — scored against the same settled money
Comprenda Research — working paper. First posted 2026-08-20; revised 2026-08-21 (two-judge say–do audit, robustness section); revised 2026-09-08 (four-instrument reframe, methods section, say–do section rewritten around the 17-conclusion result, ceiling corrected for label grain). Contact: addie@comprenda.ai.
Superseded title, kept findable: “Simulating Seniors: The Behavioral Validity of Synthetic Agents Against Real Economic Outcomes.” The scope changed rather than the evidence: the paper now reports two of our own instruments failing beside the two we did not build.
Abstract
Synthetic respondents — generative agents standing in for human participants — are validated almost entirely against what people say. The strongest published result reports 0.83–0.85 normalized accuracy on survey answers, and on incentivized behavior reports a normalized correlation of 0.66 on continuous lab-game amounts — a different metric on a different kind of outcome, and one on which that paper's own ANOVA finds no significant difference between its interview-built agents and a plain demographic baseline. We evaluate the behavioral validity of four instruments — synthetic agents, stated intent, a learned in-call model of ours and a grounding label of ours — where it is hardest and most consequential: real purchase decisions by U.S. seniors, resolved against settled money. Two of the four are ours and two of the four fail, which is why this is a result rather than advocacy. Given a senior's demographics, four frontier models (GPT-5, Gemini 2.5 Pro, Claude Opus 4.8, Fable 5) predict who actually enrolls no better than chance (AUC 0.47–0.53) — though that is one leg of a two-leg contrast, and handed the same transcript the frontier arm draws with ours (0.708 against 0.743, not significant), so the demographics-only leg is the whole of the gap. Asked the easier question of what the senior says NEXT, given the real conversation so far, a frontier persona scores 2.179 log-loss against a population base rate of 1.600 — substantially worse than knowing nothing about the individual at all. Reading only the first five caller turns, an outcome-grounded model reaches AUC 0.7544 (0.729–0.778) on 1,821 held-out calls carrying 466 settled sales, coverage-clean fold c11287d535b4005f, with positive Brier skill against the base rate (+0.1384, 0.1036–0.1689) on the calibrated probabilities, or 68.3% balanced accuracy (65.8–70.7) at the pre-specified class-prior threshold. The two-sided read, which takes the agent's speech as well as the senior's, requires a refit on this fold and is not reported here. We then measure the ceiling those numbers run against, on seniors observed deciding more than once. On connected pairs at a 30-day separation a senior agrees with their own settled purchase decision 59.7% of the time against 58.8% expected by chance — kappa +0.022, at chance with themselves. The higher figures this paper previously reported (69.5%, and 72.1% at three days) are superseded: they were computed on a person-month label without excluding within-month pairs, so they measure the billing period as much as the person, and the correction moves the ceiling down rather than up. The individual propensity that synthetic-respondent methods set out to recover does not determine the outcome, which bounds that entire class of method well below where it is assumed to sit. Handed those same conversations, the frontier models and the outcome-grounded model track what the senior said about equally well; the grounded model's edge is a training label they do not have, worth 26 more settled sales per 1,821 calls than GPT-5 at 30% of dial capacity (95% CI +2 to +46) on the one fold measured, which is the only point in the capacity table that clears zero; at 20% and at 5% of capacity the difference is within noise. Frontier models read medication count as a buying signal; on settled outcomes it carries none (enrolment flat at 23 to 28% across 0 to 9+ medications). On the stated side, of 17 business conclusions frozen from what these seniors said before any outcome was opened, 5 survive, 4 reverse and 8 disappear once settled money is joined. On one senior pharmacy book, about half of the seniors an AI reads as positively engaged do not settle (49–56%). Two of our own instruments were tested the same way and did not survive: a turn-10 in-call model whose margin is reproduced at 104% by a one-column forecast of eventual call duration, and a grounding label that is reliable at kappa 0.7205 and predictively null, beaten by a character count. None of this refutes simulation; it bounds it — synthetic agents may generate, explore and prioritise, and may not adjudicate. The gap between simulated saying and real doing is not closed by scale or prompting; it closes only with the settled outcome, which someone must hold.
Four instruments, one endpoint
The thesis is not that synthetic agents fail. It is the more general and more defensible one, and it is the claim this paper's own results were allowed to falsify: you cannot know the behavioural validity of any instrument — synthetic, stated, or learned — until it is joined to what people actually do, on the same identified person, against money that settled.
Four instruments are tested here against the same endpoint. All four are plausible a priori. Two of them are ours, which is the reason this is a result rather than advocacy: a paper that only reports other people's methods failing is a sales document.
| instrument | what it promised | what settled money said |
|---|---|---|
| synthetic agents (four frontier labs) | a persona built from who the person is stands in for the person | given a senior's demographics, AUC 0.47–0.53 on who enrols. That is one leg of a two-leg contrast and the other leg removes the gap: handed the same transcript the frontier arm draws with ours (0.708 against 0.743, not significant; 0.7434 against 0.7075 at equal n, z 0.7). Asked the easier question of what the senior says next, a frontier persona scores 2.179 log-loss against a 1.600 population base rate — worse than knowing nothing about the individual |
| stated intent | the person tells you what they will do | of 17 business conclusions frozen from the stated side alone, 5 survive, 4 reverse, 8 disappear |
| a learned in-call model — ours | the conversation at a landmark predicts the outcome | withdrawn at turn 10. A one-column forecast of eventual call duration, built from turn-10 quantities only, reproduces 104% of the margin that was published (excess −0.0022, closer-clustered [−0.0155, +0.0121]) |
| a grounding label — ours | evidence that the senior understood predicts the decision | two frontier instruments agree at kappa 0.7205 [0.6372, 0.7937] and the label is predictively null: out of fold the seven label dummies rank settlement at 0.381 and 0.483, at or below chance, while a character count through the same turn ranks it at 0.6255 |
Read the table for what it does and does not say. Two of the four carry real information at some level — stated intent is a genuine ordered signal over part of its range, and the caller channel is not redundant with the agent's — and the grounding label is reliable as an instrument even where it is null as a predictor. The claim is not that instruments are worthless. It is that their behavioural validity is unknowable without the join, and that the join is what separated the two that carry information from the two that carry length, register, or nothing.
The methods contribution
Two procedures came out of this work, and they may outlast every finding in the paper. Both are stated as rules because both were found by breaking our own results with them.
A length-so-far baseline at a landmark is not a length control
At landmark k — score the call at the senior's k-th turn, on the risk set of calls still active there — turns-so-far is constant = k by construction. A baseline fitted on how much has been said by turn k therefore controls length inside the window and says nothing about how long the call will eventually run. On this book, how long the call eventually runs is very nearly the outcome itself. Settled rate moves from 0.018 to 0.893 across quintiles of eventual duration, a 50-fold spread, and actual eventual duration ranks settlement at AUC 0.9428 at turn 10. At the turn-10 landmark of a second study on the same book, frame settle rate runs 0.0628 below 450 s, 0.7072 at 451–900 s and 0.9175 above 900 s.
Eventual duration is post-treatment, so it cannot be used as a covariate. The move is decomposition rather than control: fit the same turn-1..k text to log(eventual duration), collapse it to one scalar, and rank the outcome with that scalar alone. If the scalar reproduces the published margin, the published margin was a duration forecast wearing a state finding's clothes.
It was. At the exact holdout our own turn-10 margin was quoted at, the one-column duration forecast reads 0.6711 against the full model's 0.6689 — an excess of −0.0022, closer-clustered [−0.0155, +0.0121], reproducing 104% of the published margin — and length-so-far plus that one column beats the full model outright at 0.6790. The pre-declared kill criterion (excess ≤ 0.02 with an interval crossing zero) was met, and the turn-10 result is withdrawn as a state finding. The turn-20 landmark survives at about a quarter of what was published: excess +0.0246 [+0.0080, +0.0390] closer-held-out and +0.0234 [+0.0149, +0.0313] person-held-out. A corollary went with it: the function-word "register" placebo was a weaker proxy for the same duration channel and is superseded.
This killed our own headline the day it was written. It generalises to any survival-entangled prediction on unfolding interaction data — a call, a chat, a session, a visit — and it is the most transferable thing in this paper.
Outcome-correlated missingness, and what a re-run is for
The second rule arrived with a measured example. 5,232 of 20,861 billed transfers (25.1%) carried no usable transcript text, and the missingness was not ignorable: on billed transfers with a settled read, calls that had text settle at 0.2302 and calls that did not settle at 0.6124 — the unreadable calls are where the money is, at roughly two and a half times the rate. The mechanism is duration, and it is partly a counterparty's hard 450-second exclusion rather than an analytic choice.
First, 85.1% of the missing calls were missing only from the reader. 4,443 of the 5,218 with audio had already been transcribed and were sitting on disk, outside the table every reader keys on: the merge moved that cache from 44,818 rows to 49,261. Recovery cost $0 and made no model or network call. A store being full is not the same as the reader seeing it, and that distinction is worth a quarter of this corpus.
Second, the re-run is the contribution, not the repair. Every load-bearing result was re-run on the repaired corpus and the three did not land the same way.
| result | on the repaired corpus |
|---|---|
| the 17-conclusion say-to-do proof | reproduces bit-identical — 1,379 result leaves compared, 0 differ; it reads the opener leg and joins settled money by call record, never by text, and the cap is on the buyer leg |
| a hesitation contrast, +0.105 | reproduces on a disjoint stratum under a second transcription engine — published +0.10534, recovered +0.10546 |
| an engagement contrast, −0.141 | magnitude withdrawn — recovered −0.05075 on comparable n and comparable closer counts, and the two intervals do not overlap. It is a property of the short-call half of the book and must not be quoted as a size |
And a result that runs against the easy version of this argument: duration-decile inverse-probability weighting removed 85% of the bias in the marginal rate (naive 0.2302, 26.4% low; IPW 0.3007, 3.9% low; after recovery 0.3080, 1.6% low; truth 0.3129). IPW was not useless. The narrower and stronger claim is that IPW can reweight a rate and cannot reweight a feature that does not exist — every load-bearing claim here is a within-call textual contrast, and a call with no text has nothing to weight at any price.
The prior work
Park, Liang, Bernstein et al. built a generative agent for each of about a thousand people from two-hour interviews, and validated the agents against:
- the participants' own survey answers — normalized accuracy 0.83 for interview-built agents
(raw 65.67% divided by participants' own two-week self-consistency of 79.53%), rising to ~0.85 for survey-plus-interview agents;
- the participants' behavior in five incentivized economic games (dictator, both trust roles,
public goods, prisoner's dilemma) — a normalized correlation of 0.66. Their words: "Since the outcomes in these economic games are continuous measures, we calculated correlation coefficients and normalized correlations."
These two numbers are not on the same scale, and the difference between them is not a percentage-point gap. One is an accuracy on categorical survey items; the other is a correlation on continuous amounts standardized to 0–1. Subtracting them is not meaningful, and neither is comparable to a balanced accuracy on a binary purchase. We therefore make no claim of beating or trailing 0.66, and no comparable figure exists anywhere in this paper.
The finding in that paper that does bear on ours is a different one, and it is theirs, not our reading of it. On the survey side, interview-built agents clearly beat a demographic baseline (0.83 against 0.74). On the behavioral side, the advantage disappears: "A one-way ANOVA comparing correlations across all five agent types found no significant differences in economic-game performance (F(4, 5255) = 1.63, p = 0.16), and no pairwise contrast was significant." Interview agents scored 0.66, persona agents 0.57, demographic agents 0.48, and those differences do not separate. Two hours of interview buys a great deal on stated attitudes and, on this evidence, nothing statistically distinguishable on real-stakes behavior.
That is measured on ~$10 one-shot lab stakes in a nationally representative sample. It is not a result about seniors, and no senior-specific behavioral figure has been published. That is what this study measures.
What stated intent is worth once you open the behaviour side
The question a stated-preference instrument has to answer is not whether people sometimes change their minds. It is whether the conclusions you would have drawn from what they said survive contact with what they did.
So we froze them. A stated-preference product was given only what it could plausibly elicit before any outcome exists — who the person is (age band, sex, state), which list they came from, and what they said — and was given no call duration, no transfer, no billable status and no settled outcome. From that side alone two kinds of conclusion were written down and frozen: prioritise segment A over segment B, for every segment pair the stated side reliably separates, and concern X matters more than concern Y. Only then was the behaviour leg opened.
Of 17 business conclusions a stated-preference study would have reached from this population, 5 survive as decisions worth acting on, 4 reverse outright, and 8 evaporate. Zero point the right way but are too small to justify the reallocation they recommend.
| verdict | n | share | what it means for a buyer |
|---|---|---|---|
| HOLDS | 5 | 29% | the stated-preference study was right and the money agrees |
| REVERSES | 4 | 24% | acting on the stated-preference study moves budget the wrong way |
| DISAPPEARS | 8 | 47% | the difference is real in what people say and absent in what they do |
| UNECONOMIC | 0 | 0% | directionally right, too small to pay for the reallocation |
The four inversions, named, because a reversal is the expensive kind of error:
| a stated-preference study would have said | on what basis | behaviour says |
|---|---|---|
| prioritise stated-intent level 3 over level 2 | the study's own intent ordering | 17.4% enrol against 30.2% — 12.8% the wrong way [−0.193, −0.074] |
| prioritise not verbally polite over verbally polite | 81.8% express willingness against 76.1% | 19.0% against 31.2% — 12.2% the wrong way [−0.144, −0.098] |
| prioritise stated-intent level 4 over level 2 | the study's own intent ordering | 18.0% against 30.2% — 12.2% the wrong way [−0.178, −0.062] |
| prioritise under 65 over 75–84 | 84.1% express willingness against 79.5% | 15.2% against 25.2% — 10.0% the wrong way [−0.147, −0.057] |
Scope: 11,980 calls across 11,462 people; segments under 100 calls were not eligible to produce a conclusion; UNECONOMIC is a declared threshold in the script rather than a judgement made after the fact; every interval is a person-clustered bootstrap. The settled feed covers one buyer, so this is a description of that book rather than a population quantity, and the stated side is a reconstruction from sales conversations rather than a fielded survey. Nothing here says surveys are useless. It says one thing: without the join, the validity of the instrument is unknown, and unknown in a direction you cannot guess.
The per-call gap
On one senior pharmacy book, about half of the seniors an AI reads as positively engaged do not settle (49–56%). "Engaged" is a model-coded read of the call; "settle" is the settled purchase, joined per call.
Against the literature
The intention–behavior gap is among the most replicated findings in social psychology, which gives a benchmark. Sheeran and Webb put it plainly: "The intention–behavior gap is large — current evidence suggests that intentions get translated into action approximately one-half of the time." The underlying meta-analysis of meta-analyses spans 422 studies and finds a sample-weighted intention–behavior correlation of r = 0.53 across more than 80,000 participants (Sheeran 2002).
On the same book, stated intent predicts settlement at 0.58–0.67, below that benchmark (0.81 AUC-equivalent). The value of measuring the gap here is that it can be measured per conversation, against money that actually settled.
Why this is not compared to the synthetic-agent literature
It is tempting to place this beside the generative-agent results and it cannot be done. Those numbers describe how well a simulated agent reproduces a real person's survey answers versus that same person's behavior in economic games — simulation fidelity across two instruments. Ours is the proportion of real people who did what they said. They are different objects, and any spread computed between them is meaningless. We make no such comparison anywhere in this paper.
The label was audited, and it mattered
An outcome-blind review of 150 balanced calls (75 settled, 75 not), by a language model kept blind to both the outcome and the original label, found the "wanted to enroll" label 89% accurate on calls that settled and only 28% accurate on calls that did not. It systematically over-attributed intent to non-committal, information-gathering calls. Our own model-coded "say" was confounded, and only checking it against the settled outcome caught it. That is this paper's thesis applied to its own instrument.
Three limits, stated because they bound the claim. Every tier above is a machine reading of the call, the started-an-application row included — the feeds this paper reads carry no field for a senior's intent, so no version of this measurement is free of a model. The auditor was also a language model, not a human panel. And the audit is a single judge at n=150; a larger two-judge version exists in working notes but its per-call grader output is not committed, so nothing from it is quoted here.
What real economic outcomes add (and what our approach honestly cannot match)
We hold the one thing that makes their weakest number look generous: real, high-stakes economic behavior at scale. A senior on a live sales call deciding whether to buy — with real money, a real product, in the wild — is a vastly stronger behavioral ground truth than a lab game.
| dimension | Park et al. | This study |
|---|---|---|
| behavioral ground truth | ~$10 incentivized lab games | real purchase, real money (settled $) |
| ecological validity | artificial, one-shot | live sales conversation in the field |
| N (behavioral outcomes) | ~1,000 | 400 in the demographics-only LLM head-to-head |
| refresh | frozen snapshot | settles daily, continuously |
| population | general US | seniors — least AI-represented, most economically valuable |
| models tested | GPT-4-class | GPT-5, Gemini 2.5 Pro, Claude Opus 4.8, Fable 5 |
| the loop | measures the gap | measures AND closes it (grounding lifts prediction) |
Honest limit: Park built person-specific twins from 2-hour interviews; our unit is the population/segment + the conversation, not a deep individual twin. So the claim is NOT "we redid their study with better twins." It is: their method raised a question it could not answer — on real high-stakes behavior, how do synthetic agents do, and can the gap be closed? — and we answer it on the ground truth they lacked.
The task (matched to theirs)
Given a synthetic representation of a senior, predict that senior's real economic decision (settle / no-settle) on a held-out call. This is their behavioral-prediction task, ported from a lab game to a real purchase.
Preliminary result (already run; us_vs_models, n=400 held-out, 74 positives)
Predicting the real senior behavioral outcome:
| arm | AUC | 95% CI |
|---|---|---|
| GPT-5 — ungrounded synthetic senior | 0.466 | [0.391, 0.538] |
| Gemini 2.5 Pro — ungrounded | 0.483 | [0.413, 0.553] |
| Claude Opus 4.8 — ungrounded | 0.505 | [0.437, 0.572] |
| Fable 5 — ungrounded | 0.528 | [0.460, 0.596] |
| Ours — profile-grounded, NO conversation | 0.566 | [0.497, 0.635] |
Given only a synthetic representation of a senior, every ungrounded frontier arm sits at chance. Our own profile-only arm does no better, which is the point: the signal is not in the profile.
Our standing conversation result. Reading only the first five caller turns, before any decision is on the table, we rank who enrolls at AUC 0.7544 (95% CI 0.729–0.778) on 1,821 held-out calls carrying 466 settled sales, person-disjoint and forward in time, coverage-clean fold c11287d535b4005f. Calibrated on an inner slice and scored once on the test fold, the same model carries positive Brier skill against that population's base rate (+0.1384, 95% CI +0.1036 to +0.1689; mean predicted 0.2622 against an observed base rate of 0.2559), so the model prices as well as ranks. Ranking and pricing are read off two arms of one model and we name which is which: the 0.7544 is the uncalibrated ranking score, and the calibrated probabilities that carry the Brier skill and the threshold rule below rank at 0.7472. Reading both sides of the conversation is a separate model that has not been refit on this fold; the Park-units comparison below reports the caller-only leg and marks the two-sided leg pending.
Every ungrounded frontier model's CI includes 0.50 — statistically a coin flip. The best two (GPT-5, Gemini) point below chance. Complementary evidence: the confirmation-machine test — six flagship models across four labs, against a sealed pre-registered answer key, asserted senior-persuasion tactics our settled-outcome data does not support; Gemini endorsed all six at 92–95% confidence.
Our accuracy in behavioral-match units
Park's headline is a match rate: how often the agent's predicted behavior matches the real behavior, on a roughly balanced binary choice. Balanced accuracy is the statistic that ports across base rates, so that is what we report — on the real settled outcome, person-disjoint held-out, with the threshold fixed before the test fold is read.
| what the model was given | behavioral-match (balanced accuracy) | |
|---|---|---|
| Ungrounded frontier LLMs | a senior's demographics | ~50% (chance; AUC ≈ 0.50) |
| Comprenda, caller speech only | the first five things the senior said on a cold call | 68.3% (65.8–70.7) |
| Comprenda, both sides of the call | the conversation through the senior's fourth turn | pending a refit on this fold |
Reading only what the senior says we land at 68.3%, on held-out calls, person-disjoint, with the threshold fixed before the test fold was read. The two-sided model, meaning what the senior said and what the agent said back, is a separate fit and this fold does not yet carry it, so the row above is marked pending rather than filled from a different population. Neither figure is a comparison to the prior work's 0.66, which is a correlation on continuous lab-game amounts and does not convert into these units. What follows is our number on our task.
It should still be read against what the model was handed. Park's agent had a two-hour interview with the individual whose behavior it then predicted, and predicted a one-shot choice worth about ten dollars in a lab, made by a participant who had consented to be studied. Ours had the first five caller turns of a cold sales call, no prior contact with the person, and predicted whether they would actually enrol in something that takes money from their account every month, scored against the settled record.
Two honest caveats on the 68.3%. It is an operating point, not a ranking: it is read off the calibrated arm at the class-prior threshold, and the arm that produces the headline AUC is the uncalibrated one, so the two figures describe the same model at two settings and are quoted as such. And a balanced accuracy in this band is what the two-sided model reached on its own earlier fold, so the caller-only and two-sided legs can no longer be presented as a ladder; whether reading the agent's speech adds anything on this fold is exactly the open question the pending refit answers, and until it lands we do not claim it does.
Superseded, kept findable. The figure the caller-only leg replaces is AUC 0.726 [0.699, 0.7507] on 1,660 calls carrying 472 enrolments, K5, Brier skill +0.1195 — canon from 2026-09-04 until the coverage-clean fold replaced it on 2026-09-06 — and the 64.1% balanced accuracy this paper carried until 2026-09-08 was that fold's operating point, retired with it. The earlier caller-only and two-sided Brier figures of +0.120 and +0.144 belong to the same superseded fold.
The result barely depends on the threshold rule. On the calibrated probabilities, thresholding at the class prior gives 68.3% and flagging the top prior-fraction of the list gives 67.2%, a gap of about a point. On the uncalibrated ranking arm the two rules diverge, 65.5% by fraction against 60.1% by threshold, which is the expected behaviour of a score that is not a probability, and the reason the operating point is reported on the calibrated arm and not on that one.
The threshold is pre-specified. Balanced accuracy weights both classes equally, so for a calibrated posterior the balanced-accuracy-optimal rule is to predict positive when the probability exceeds the class prior. We use the training base rate, fixed before the test fold is touched — no search, no sweep, nothing read off the test set. Both models are reported at that rule, and both are also reported under a threshold selected on an inner validation slice; the two rules agree to within a point on every arm, so neither number depends on the choice.
What the number is, and is not
It is a position on a curve, not a ceiling. Accuracy here is a function of how many settled outcomes we hold, and the curve has not bent:
| outcome-linked calls in training | AUC on a fixed held-out fold |
|---|---|
| 250 | 0.559 |
| 500 | 0.610 |
| 1,000 | 0.658 |
| 2,000 | 0.687 |
| 4,000 | 0.710 |
| 6,754 | 0.720 |
That is +0.019 AUC per doubling, still rising at full n. Depth moves it too: reading twelve caller turns instead of five reaches AUC 0.708 with the leakage gate passing un-exempted at every depth. Park's method has no equivalent axis. A generative agent built from an interview does not improve as purchases settle, because nothing in that pipeline ever observes a purchase. Ours improves at the rate real people make real decisions, which is slow, unglamorous, and not purchasable.
The models' error is not a direction. It is a missing label. This is the part a match rate hides. On the same calls and the same settled outcomes, the outcome-grounded model and three frontier models track what the senior said about equally well. The grounded model's edge is a training label they do not have, and the one place it clears zero is at 30% of dial capacity, where it is worth 26 more settled sales per 1,821 calls than GPT-5 (95% CI +2 to +46) on the one fold measured. The spread across all three frontier models at that same 30% point is 26 to 39 sales and carries no interval. At 20% and at 5% of capacity the difference is within noise. Where the missing label shows most plainly is an inference the corpus never corrects: frontier models read medication count as a buying signal, and on settled outcomes it carries none.
We read that as a property of the training medium rather than of any lab's choices. A corpus of language records what people said they would do in enormous volume, and what they then did almost never, because the record of what they did sits in a private billing ledger that appears in no corpus. A model trained on that medium learns the plausible inference, that more medication means more need means more likely to buy, and never observes the purchase that would correct it. Whether a message is persuasive is legible in text. Whether it converts is not. The implication for scale is the uncomfortable one: more language does not supply the label.
The medication table
The per-band figures behind the paragraph above, on real senior conversations graded against the settled enrolment (demographics-known fold, n=1,020, 246 sales):
| medications | n | actually enrol | Claude Opus | GPT-5 | Gemini 2.5 Pro | Comprenda |
|---|---|---|---|---|---|---|
| 0–2 | 63 | 25.4% | 24.5% | 20.9% | 14.0% | 21.1% |
| 3–5 | 494 | 24.5% | 34.9% | 41.3% | 41.1% | 23.9% |
| 6–8 | 269 | 27.5% | 35.4% | 44.0% | 46.9% | 21.8% |
| 9+ | 56 | 23.2% | 35.2% | 47.2% | 50.5% | 21.6% |
The real rate never leaves 23 to 28%. The frontier predictions climb with medication count; the grounded model stays flat with the truth. The adjusted effect of medication count on the corrected label is +0.0035 logit per medication, 95% CI (−0.053, +0.057) — a null.
Superseded, kept findable. Until 2026-09-06 this section reported per-model over-prediction multiples on the legacy phone-grain label; they are withdrawn (canon row B6b). On the corrected per-call label the frontier arms and the grounded arm track each other, which is why this section now reports a missing label rather than a direction.
What simulation is for, and what it cannot decide
Nothing above refutes simulation. It bounds it, and the boundary is worth stating explicitly rather than leaving a reader to infer a blanket dismissal we do not hold.
What the measurements license is a claim about one job: synthetic agents cannot adjudicate real treatment effects on this population. Given a senior's demographics they rank the settled decision at chance; given the same transcript they draw with an outcome-grounded model rather than losing to it; and asked the easier question of the senior's very next turn they score worse than knowing nothing about the individual. On the endpoint that decides money, they are not an instrument you can settle a question with.
That is one job of several, and the others are not touched by any result here. Simulation can generate — hypotheses, wordings, edge cases, adversarial scripts. It can explore failure modes cheaply, in volume, before anything is dialled. It can prioritise what deserves an expensive validation, which is exactly what this estate uses it for: the screen decides what gets tested, and the field decides what is true. And its failures are informative in their own right — a model that reads medication count as a buying signal is telling you something true about the medium it was trained on.
The rule we operate by, stated plainly: simulation may propose; only the settled outcome may dispose. A synthetic result is a candidate, never a finding, and it is promoted only when a real outcome that nobody could have read in advance agrees with it.
The headline
The ceiling on predicting a person
Every synthetic-respondent method rests on an assumption that is almost never stated, let alone measured: that there is a stable individual propensity to recover. Build a good enough model of the person — from demographics, from a persona, from a two-hour interview — and their behavior follows.
That assumption is testable wherever the same person can be observed deciding more than once. Ours can. Among seniors who had more than one real conversation with a settled outcome on each, we asked how well one decision predicts the same senior's next one, taking a single observation per senior-day and scoring in the same balanced units the model is scored in.
On the measurement as it was first run — and read the correction two subsections below before quoting any of it — a senior reproduces their own purchase decision 69.5% of the time (63.2–75.3, 307 pairs from 261 seniors, resampling seniors rather than pairs).
The number that matters is not the pooled one, though. It is the left edge:
| gap between decisions | pairs | seniors | reproduces own decision |
|---|---|---|---|
| 0–3 days | 50 | 49 | 72.1% (55.0–85.1) |
| 4–7 days | 47 | 41 | 76.3% (62.3–88.3) |
| 8–14 days | 59 | 55 | 70.5% (56.8–82.8) |
| 15–21 days | 36 | 32 | 75.0% (56.2–91.1) |
| 22–45 days | 75 | 73 | 67.6% (56.6–78.0) |
| 46+ days | 40 | 39 | 49.8% (35.9–65.6) |
Within three days, the same senior lands differently more than a quarter of the time. Three days is too short for much genuine change: they have not aged, moved, lost coverage, or reconsidered their finances. Whatever is producing that disagreement is not the person being different. It is that the decision was never a fixed property of the person to begin with.
That is the ceiling. A model that perfectly recovered a senior's stable propensity would still be wrong about a quarter of their decisions, because the propensity does not determine the outcome. It is a bound on the entire category of methods that work by simulating an individual, and it is measured rather than assumed.
The correction the ceiling needs: the settled label is a person-month stamp
The argument stands; the arithmetic above needs a correction that makes it stronger, and the correction is that the settled purchase label is not a per-call receipt. It is a person-month stamp. Of 708 person-months holding more than one settled-labelled call, 649 (91.7%) carry an identical label on every call in that month, and of 336 adjacent settled-to-settled pairs, 284 (84.5%) carry the identical sale date — one enrolment stamped onto two or more calls.
What that does to any repeat-decision number is not subtle:
| pairs | n | agreement (kappa) |
|---|---|---|
| all adjacent pairs | 1,223 | +0.684 [+0.640, +0.726] |
| same calendar month | 838 | +0.850 [+0.809, +0.888] |
| different month | 385 | +0.271 [+0.165, +0.366] |
| gap > 30 days | 211 | +0.087 [−0.046, +0.217] |
A naive read says a senior's purchase behaviour is highly repeatable. The clean read says it is indistinguishable from zero, which reproduces this estate's standing figure of kappa +0.022 on connected pairs at 30 days. Any repeat-decision number computed without excluding within-month pairs is measuring the billing period, not the person.
That applies to the table above, and it is the sharper of two defects in it. The repeat-decision run read the phone-grain settled store rather than the corrected per-call label, and it did not exclude within-calendar-month pairs — calendar month entered that script only as a descriptive field. The 0–3 day bucket is by construction almost entirely inside one calendar month, so its 72.1% is substantially the billing period agreeing with itself, and so is the pooled 69.5%.
The direction matters. The grain biases toward agreement, so the true short-interval consistency of a senior with themselves is lower than the figures printed here, and the ceiling those figures describe is lower than it looks. The bound on the synthetic-respondent class is strengthened by the correction, not weakened. The superseding measurement, on connected pairs at a 30-day separation, is that 191 pairs agree 59.7% of the time against 58.8% expected by chance on those same pairs — kappa +0.022. On the clean grain a senior is at chance with themselves about enrolling.
Superseded, kept findable. 69.5% (63.2–75.3, 307 pairs from 261 seniors) and the gap-bucket table are retained above because the shape of the decay is what the permutation test was run on, and because a retired number that disappears gets re-derived. They are not quotable as properties of a person: they are properties of a person and a billing period. Where this paper needs a ceiling, the number is kappa +0.022 at 30 days.
Three independent results line up behind it. Given a senior's full demographics, four frontier models predict who enrolls at chance. Given a two-hour interview with the specific person, the prior work's own ANOVA finds no significant advantage over a demographic baseline on real-stakes behavior. And given the conversation, we reach 0.7544. The information that predicts the purchase is in the encounter, not in the participant.
What we can and cannot attribute
Consistency does fall with the gap, and that fall is a real function of the gap rather than an artifact of how pairs are formed: shuffling the gaps across pairs reproduces a decline this steep in 0.55% of 2,000 permutations.
We are not claiming it shows a person changing. Two things prevent it. The classical test–retest literature is explicit that a longer retest interval mixes genuine change into what looks like unreliability, so a decline at six weeks is the expected result rather than a discovery. And our own attempt to hold the person fixed fails on sample size: only two seniors contribute pairs at both a short and a long horizon, which cannot separate person-change from cohort composition. The long bucket is also structurally drawn from earlier arrivals, since a 46-day gap requires a first contact early in the window, and it carries a lower settle rate (0.225 against 0.27–0.35 elsewhere).
So the decay is reported as consistent with what the retest literature predicts, and the short-interval number is the one carrying weight. Forty pairs is the binding constraint on the long arm, and every interval here is wide.
For the same reason we make no claim from the comparison with the prior work's 79.53% two-week survey self-consistency. At a matched interval ours is 70.5%, and the interval spans 56.8 to 82.8.
Calibration, not just accuracy
A match rate says whether the ranking is right. It says nothing about whether the numbers mean anything, and for a decision that is the part that matters: a model that says 40% must be right about 40% of the time, or the decision built on it is wrong even when the ranking is perfect.
Our probabilities carry positive Brier skill against the population's own base rate: +0.1384 (95% CI +0.1036 to +0.1689) for the caller-only model on the coverage-clean fold, with mean predicted 0.2622 against an observed base rate of 0.2559. The model prices as well as it ranks. The two-sided model's Brier skill is not restated here, because the two-sided arm has not been refit on this fold.
The frontier models are not scored on calibration in this edition. The per-model over-prediction multiples an earlier draft carried here rested on a label since corrected and are withdrawn; the level result that survives on the corrected label is the medication contrast above, which an accuracy-only evaluation cannot see. The earlier caller-only and two-sided Brier figures of +0.120 and +0.144 belong to the superseded fold and are retired with it.
What our own false-alarm rate is
Every result above rests on a scorer deciding whether two things differ. That scorer has an error rate, and asserting it is not the same as measuring it. We ran ours against 150 independent null pairs — comparisons where, by construction, there is nothing to find — and it called a difference 5.9% of the time (upper bound 10.5%) at a stated alpha of 5%.
Whether it transfers
The results above are on prescription-benefit calls. Whether they transfer to other senior verticals is not reported in this edition.
What this is measured on
101,023 recorded hours of senior conversation. They are two legs and are never summed into a rate denominator: the closer book at 40,406 hours, and the opener leg at 60,617 hours counted on the 4,500,102 sessions where the senior actually spoke. It is the decision itself rather than an interview about the person.
Those conversations are in text where text exists: 102,547 transcribed closer-leg calls (medical alert 41,597, prescription delivery 24,423, final expense 10,833) and 778,377 opener-leg conversations across five partner feeds, counting only sessions where the senior took three or more turns.
Three distinctions matter and are easy to blur. These are conversations, not calls — the opener and the closer are different agents having different exchanges with the same senior, and roughly half the opener legs sit on a call that also has a closer leg. They are not people: the same senior is dialled repeatedly, and the warehouse join that carries an exact date of birth covers 364,934 rows rather than 364,934 individuals. And recorded is not transcribed, and neither is labelled — the population carrying a settled enrolment label, which is what every accuracy figure in this paper is scored against, is a small fraction of either. The dial counts are larger still and are not conversations at all.
That gap between what is recorded and what carries a settled outcome is the honest shape of this asset, and it is why the constraint on the work is settled outcomes rather than conversation volume.
A frontier persona predicts a senior's next turn worse than the base rate
Everything above asks a model to predict a purchase. That is a hard question, and failing it is forgivable. So we asked the easy one instead.
Give a model the real conversation up to a point — the senior's actual words, the agent's actual last line — and ask only what the senior says next. No outcome, no money, no thirty-day window. Just the next turn, scored as a probability over the fourteen caller states our tagger recognises.
| predictor | log-loss | 95% CI |
|---|---|---|
| frontier persona (Claude Sonnet 5) | 2.179 | [2.127, 2.235] |
| the population base rate | 1.600 | [1.572, 1.627] |
| transition table fitted on our conversations | 1.559 | [1.531, 1.589] |
The persona is substantially worse than the base rate. Knowing nothing whatever about the individual — just how seniors in general distribute across the next turn — beats a frontier model that has read the whole conversation. Paired difference against our table: −0.620 [−0.666, −0.575], on all 1,644 held-out persons × 4 turns = 6,576 predictions, person-disjoint, pre-registered, bootstrapped over callers.
It is not model size. A stronger persona (Claude Opus 5) narrows the gap to −0.250 [−0.333, −0.175] and still loses. The extension was pre-registered with a softening rule — if the stronger model tied, the claim retracted to "matches a frontier persona" — and the rule was not triggered. At 6.6× the original sample the gap widened rather than closed.
We are not claiming to be good at this. Our table beats the base rate by +0.040 [+0.031, +0.051]. That is real and it is small. The −0.620 is almost entirely the model failing, not the table succeeding, and the honest statement is the first one: a frontier model predicts a senior's next turn worse than the population average does.
How it fails is the informative part. The stronger persona wins top-1 accuracy — 0.432 against our table's 0.394 — while losing log-loss. It picks the single likeliest reply slightly better and is badly calibrated about everything else. And where the senior actually went quiet, its log-loss is 4.40 against the table's 2.45. The model imagines a senior who assents and asks articulate questions; real seniors are quiet or non-committal in 45% of turns.
This measures turn fidelity and nothing else. It is not a claim about predicting the settled sale.
Robustness
The grounded model's signal survives every identity control we can throw at it:
| control | result |
|---|---|
| Person-disjoint, walk-forward split | held throughout (no person appears in train and test) |
| Source-grouped CV (model never sees the test supply source) | −0.009 against its own baseline |
| Brand-masked (all brand/agent-name tokens removed) | −0.003 against its own baseline |
| Trivial-detector gate at every turn depth (length, closing lexicon, consent, carrier names, qualification counts, PII verification, buyer identity) | passes at turn 3, 5, 8 and 12 with no exemption claimed |
| Leave-one-vertical-out, zero-shot on settled dollars | 4 of 4 senior verticals above chance (AUC 0.60–0.68; every bootstrap CI excludes 0.50; permutation p < 0.001) |
The source-grouped and brand-masked rows report deltas against each model's own baseline, which is the robustness quantity; the turn-depth gate rows are on the standing model. The signal is what is said on the call, not who is calling, which list they came from, or which product is being sold. A well-known failure mode in adjacent content models is that a large share of apparent "content" signal turns out to be source identity; these controls rule that out here.
Method and rigor
Every arm is scored on the same blind, person-disjoint held-out set of real seniors, and the settled-outcome labels are fixed before any model is scored — no arm can be tuned to the answer after the fact. The comparison is reported as rank accuracy (AUC), which is invariant to the base rate and therefore not gameable by a model that simply predicts the majority class.
What was fixed in advance, and what was not
Stated per decision, because "held out" covers a lot of choices that are not.
| decision | when it was made |
|---|---|
| Train/test split | person-disjoint and forward in time; no senior appears on both sides |
| Settled-outcome labels | fixed before any model was scored |
| Turn depth reported as the shipping depth | chosen by a written rule on the inner validation slice, before the test fold was read |
| Classification threshold | the class prior, and separately an inner-validation choice; neither reads the test set |
| Leakage detectors and their 0.60 bar | fixed before the fold was cut; no exemption is claimed anywhere in this paper |
| Depths 3, 5, 8 and 12 | all reported, so the curve is not a selected point |
| Choice of AUC and balanced accuracy | both reported for every arm |
Exploratory rather than pre-registered: the cross-vertical transfer result (archetype-specific; pooling hurt RX), the self-consistency measurement, and the medication-count contrast. They are reported as findings, not as tests of a registered hypothesis, and each is a single measurement rather than a selected best.
What is quotable, and at what grain
Not every number in this paper is the same kind of claim, and the difference is the independence cluster rather than the sample size.
independence.py implements one rule: two observations are in the same independence cluster when they share a cell — workspace × vertical × direction with dates stripped — or share any script version, closed transitively. Run on the closer book these results are measured on, it returns one cluster: one workspace, one vertical, one direction, and no script version recorded anywhere. A cluster bootstrap over one cluster is an undefined interval rather than a wide one, because every resample is the same resample and the width comes out at zero, which errs toward the analyst.
So, per claim:
| claim | what it is |
|---|---|
| the caller-only fold — AUC 0.7544, Brier skill +0.1384, 68.3% balanced accuracy | a held-out measurement on one fold of one buyer's settled book, person-disjoint and forward in time, with phone-clustered intervals. Quotable as what this model does on this book, not as a population quantity: the cluster rule above was run on the closer book, and the intervals here are quoted at the phone grain they were computed at |
| the 17-conclusion result, 5 survive / 4 reverse / 8 disappear | a description of one book, person-clustered. The procedure generalises; the count does not |
| the grounding null and its reliability (kappa 0.7205) | a description of the only book it was measured on. A null needs no cluster-level generalization to be a fair description of that book. Not quotable: any ordering of the seven labels, any per-label effect size as a population quantity, or any claim that the layer would also fail on a second operator or vertical — that needs a second cluster, and there is exactly one |
| the turn-10 withdrawal and the 104% reproduction | a decomposition on the same single-cluster book. The rule it produced is general; the margin it killed was never a population quantity to begin with |
| the repeat-decision ceiling | a description of repeat-contacted seniors, who are dialler-selected, and now additionally bounded by the person-month label grain above |
Where a number is a description of one book, this paper says so beside the number rather than in a footnote, and no interval is quoted at a grain the module refuses.
The gate is a fixed instrument, not a judgement call
Before any arm is quoted, its text is scored by six trivial detectors — length, closing lexicon, consent language, carrier names, qualification counts, and PII verification — plus a buyer-identity detector and a survival check on how many turns were available. If any of them predicts the outcome better than 0.60 on its own, the arm does not ship. The detectors run at every reported depth, the bar is fixed in advance, and passing produces a receipt hash recorded with the result.
The bar binds. An arm that does not clear it is not reported, whatever its accuracy, and the gate runs before a number is quoted rather than after it is questioned.