Comprenda Research
Comprenda Research studies the validity of the tools used to understand human behaviour, including our own. Comprenda Research asks what each source is actually entitled to tell us.
A census statistic does not have to predict enrolment to say how common broadband access is. A published hearing study does not have to predict conversion to make auditory load worth considering. Validity belongs to a use, a population, a criterion and a unit of decision, and no source gets to borrow another source’s authority.
Research principle
Where possible we join each instrument to an external criterion, usually an observed business outcome, and report what survives.
Working paper · September 2026
Outcome-linked validation of synthetic respondents, stated intent and learned conversational signals in a senior-targeted acquisition setting. The result is not that every instrument fails. It is that validity depends on the use, the criterion and the unit of decision.
Stated evidence
17 conclusions frozen before the outcome was opened: 5 held, 4 reversed, 8 disappeared.
A grounding label
Reached inter-model agreement but did not establish outcome validity, so it is no longer described as a comprehension measure.
A retired result
A previously reported turn-10 prediction was withdrawn after eventual call duration reproduced the same signal.
The result we did not have to publish
The paper reports two of our own instruments failing beside the two we did not build. A learned conversational signal was withdrawn once eventual call duration reproduced it, and a grounding label that agreed with itself across model readers turned out not to establish outcome validity. So it is no longer described as a comprehension measure.
A validation system that only ever indicts other people’s instruments is not a validation system. This is the check on us.
Outcome linkage does not make an instrument valid. It makes invalidity observable.
What the gap is made of
On the same calls and the same settled outcomes, Comprenda’s ranking and three frontier models track what the senior said about equally well. Comprenda’s edge is a training label they do not have, and it shows between roughly 20% and 30% of dial capacity, where it is worth 26 to 39 more settled sales per 1,821 calls on the one fold measured. At tighter capacity the difference is within noise.
Corrected-label canonical AUC: 0.7544 on the coverage-clean fold (n=1,821, 466 sales), quoted beside sales-at-capacity, never as an accuracy headline.
Comprenda is calibrated to real outcomes at the population level. Below the level where it was calibrated it greys out rather than fabricates. A system that will answer any question at any resolution is telling you something about its confidence, not about your customers.
The same discipline applied to speech: the senior voice eval scores every engine, ours included, on the older adult’s channel of the same sealed calls.
The reason this needed building
Underrepresented
People over 65 are poorly represented in many of the datasets and instruments used to model consumers. In one audit of 92 face datasets, five documented anyone over 65; a separate audit of opinion simulation names 65+ among the worst-represented groups in every model it tested.
Demographics carry little
Given only demographics, four frontier models predict who enrolls at a coin flip (0.47–0.53); handed the same transcript, they rank about as well as ours (0.708 against 0.743, not significant). The demographics carry almost no signal for anyone: on the corrected label they separate seniors by 3.4 points.
Misleading when asked
We froze 17 business conclusions using what people said, then opened what those same people actually did. Twelve changed.
So we built the record everyone else was missing.
Three shortcuts, tested
A profile is not an outcome. An opinion is not an outcome. A model’s confidence is not an outcome. That is why the decision has to be connected to what happened next.
Proof 01 · The person
On a common holdout, a person-level information set, age, sex, geography, contact history, predicted the outcome at chance in both splits. The persona was empty.
AUC 0.513 · 0.505
Proof 02 · The instrument
Acting on what people said improved the outcome on two decision families and made it worse than doing nothing on two others, and nothing on the stated side told us which case we were in. The damage grew with conviction.
6 of 12 allocation cells worse than status quo
Proof 03 · The model
Given blinded descriptions of real experiments, three frontier models scored at chance. Two exactly matching change nothing. They invented winners between identical scripts, and every one was more confident when wrong.
9/17 · 9/17 · 11/17 | chance = 50%
Scoped to blinded experiment descriptions and demographics. Handed a transcript, the same models rank with ours; what they cannot do is price the outcome. See the gap above.
Finding an idea is easy. Surviving the evidence is hard.
Prove
A script test changes everything and guesses afterward what caused the result. The state definition is part of the treatment: if the moment is mis-detected, the experiment is aimed at the wrong population.
In one senior-pharmacy mine, 141 apparent patterns entered the funnel. Seven survived the statistical guards, agent and independence effects, call-length artifacts, matched placebo. Four earned a test.
The last gate is the one most methods do not have. Three of the seven survivors belonged to a single decision moment. And when that moment’s detector was hand-audited it was right only a quarter of the time. All three were held out before a single call was assigned.
Surviving the data is not enough. The moment itself has to be real. A move randomized at a mis-detected moment is a different intervention from the one on the protocol, and no amount of sample fixes it. So before anything goes live, each moment gets a passport: definition, positive and negative examples, the ambiguous class and its abstain rule, a hand-labelled precision and recall audit, safety exclusions, permitted moves, the current move distribution, assignment rule, delivery verification, authoritative outcome, sample requirement and kill criterion.
Try it
All four looked real at first. Each shows the raw association we measured on one book, before the guards ran.
The outcome decides. Not our opinion.
Research standard
Science becomes marketing when only the surviving result is visible. Every current research artifact carries its version, population, outcome, analysis date, what changed from prior versions, and known limitations.
Earlier values that no longer belong to the current evidence base stay in version history rather than being silently overwritten. The current revision carries explicit supersessions, among them the retirement of an earlier headline AUC that turned out to be reading transcript length and a consent phrase rather than predicting the outcome.
Every result should be able to lose. If a method only produces wins, it is not a validation system.
The corpus
The research base includes 101,026 recorded hours of older-adult conversation across two stages of the same journey, 4.5 million conversations in which the older adult spoke, out of 24.2 million contacts; two-thirds of those are a single turn. That is a scale fact about recorded audio, and it is not the same thing as a joined asset. Beside the calls sits a bank of 56 published senior-survey instruments and standards and 243 distinct research sources across two literature registers (250 cited rows, deduplicated 2026-09-10; the registers include standards and regulations, so "sources" rather than "studies").
Recorded is not the same as transcribed.
Transcribed is not the same as outcome-labelled.
Sessions are not people.
Different analyses use different populations.
So every result reports its own denominator, and the two legs are never summed into one for any rate.