Comprenda Research

Test what we think we know
about older adults.

Comprenda Research studies the validity of the tools used to understand human behaviour, including our own. Comprenda Research asks what each source is actually entitled to tell us.

A census statistic does not have to predict enrolment to say how common broadband access is. A published hearing study does not have to predict conversion to make auditory load worth considering. Validity belongs to a use, a population, a criterion and a unit of decision, and no source gets to borrow another source’s authority.

Research principle

No instrument validates itself.

A survey
cannot establish its own behavioural validity.
A synthetic respondent
cannot prove that its simulation matches real behaviour.
A model
cannot use its own confidence as evidence that it is correct.
A conversational score
cannot be assumed to matter economically because it sounds psychologically plausible.

Where possible we join each instrument to an external criterion, usually an observed business outcome, and report what survives.

Working paper · September 2026

No Instrument Is Valid
Until It Is Joined

Outcome-linked validation of synthetic respondents, stated intent and learned conversational signals in a senior-targeted acquisition setting. The result is not that every instrument fails. It is that validity depends on the use, the criterion and the unit of decision.

Stated evidence

17 conclusions frozen before the outcome was opened: 5 held, 4 reversed, 8 disappeared.

A grounding label

Reached inter-model agreement but did not establish outcome validity, so it is no longer described as a comprehension measure.

A retired result

A previously reported turn-10 prediction was withdrawn after eventual call duration reproduced the same signal.

The result we did not have to publish

The paper reports two of our own instruments failing beside the two we did not build. A learned conversational signal was withdrawn once eventual call duration reproduced it, and a grounding label that agreed with itself across model readers turned out not to establish outcome validity. So it is no longer described as a comprehension measure.

A validation system that only ever indicts other people’s instruments is not a validation system. This is the check on us.

Outcome linkage does not make an instrument valid. It makes invalidity observable.

What the gap is made of

The edge is the label,
not the model.

On the same calls and the same settled outcomes, Comprenda’s ranking and three frontier models track what the senior said about equally well. Comprenda’s edge is a training label they do not have, and it shows between roughly 20% and 30% of dial capacity, where it is worth 26 to 39 more settled sales per 1,821 calls on the one fold measured. At tighter capacity the difference is within noise.

Corrected-label canonical AUC: 0.7544 on the coverage-clean fold (n=1,821, 466 sales), quoted beside sales-at-capacity, never as an accuracy headline.


Where the answer stops, the product says so.

Comprenda is calibrated to real outcomes at the population level. Below the level where it was calibrated it greys out rather than fabricates. A system that will answer any question at any resolution is telling you something about its confidence, not about your customers.

The same discipline applied to speech: the senior voice eval scores every engine, ours included, on the older adult’s channel of the same sealed calls.

The reason this needed building

Why older customers are
unusually easy to get wrong.

Underrepresented

People over 65 are poorly represented in many of the datasets and instruments used to model consumers. In one audit of 92 face datasets, five documented anyone over 65; a separate audit of opinion simulation names 65+ among the worst-represented groups in every model it tested.

Demographics carry little

Given only demographics, four frontier models predict who enrolls at a coin flip (0.47–0.53); handed the same transcript, they rank about as well as ours (0.708 against 0.743, not significant). The demographics carry almost no signal for anyone: on the corrected label they separate seniors by 3.4 points.

Misleading when asked

We froze 17 business conclusions using what people said, then opened what those same people actually did. Twelve changed.

So we built the record everyone else was missing.

Three shortcuts, tested

Where the answer isn’t.

A profile is not an outcome. An opinion is not an outcome. A model’s confidence is not an outcome. That is why the decision has to be connected to what happened next.

Proof 01 · The person

Who they are told us almost nothing.

On a common holdout, a person-level information set, age, sex, geography, contact history, predicted the outcome at chance in both splits. The persona was empty.

AUC 0.513 · 0.505

Proof 02 · The instrument

Stated evidence helped half the time.

Acting on what people said improved the outcome on two decision families and made it worse than doing nothing on two others, and nothing on the stated side told us which case we were in. The damage grew with conviction.

6 of 12 allocation cells worse than status quo

Proof 03 · The model

Frontier models did not beat chance.

Given blinded descriptions of real experiments, three frontier models scored at chance. Two exactly matching change nothing. They invented winners between identical scripts, and every one was more confident when wrong.

9/17 · 9/17 · 11/17  |  chance = 50%

Scoped to blinded experiment descriptions and demographics. Handed a transcript, the same models rank with ours; what they cannot do is price the outcome. See the gap above.

Finding an idea is easy. Surviving the evidence is hard.

Prove

We do not A/B a whole script.
We randomize the next move at the moment.

A script test changes everything and guesses afterward what caused the result. The state definition is part of the treatment: if the moment is mis-detected, the experiment is aimed at the wrong population.

Freeze
Define the question before the answer exists.
Assign
Create a meaningful comparison.
Verify
Confirm what was actually delivered.
Grade
Use the authoritative outcome.
Decide
Deploy, don’t deploy, or conclude that more evidence is needed.

Built to kill bad answers.

In one senior-pharmacy mine, 141 apparent patterns entered the funnel. Seven survived the statistical guards, agent and independence effects, call-length artifacts, matched placebo. Four earned a test.

141Raw patterns
→
35Independence
→
12Duration
→
7Placebo
→
4Moment audit

The last gate is the one most methods do not have. Three of the seven survivors belonged to a single decision moment. And when that moment’s detector was hand-audited it was right only a quarter of the time. All three were held out before a single call was assigned.

Surviving the data is not enough. The moment itself has to be real. A move randomized at a mis-detected moment is a different intervention from the one on the protocol, and no amount of sample fixes it. So before anything goes live, each moment gets a passport: definition, positive and negative examples, the ambiguous class and its abstain rule, a hand-labelled precision and recall audit, safety exclusions, permitted moves, the current move distribution, assignment rule, delivery verification, authoritative outcome, sample requirement and kill criterion.


Try it

Four patterns we found. Guess which one earned a test.

All four looked real at first. Each shows the raw association we measured on one book, before the guards ran.

The outcome decides. Not our opinion.

Research standard

We publish the corrections too.

Science becomes marketing when only the surviving result is visible. Every current research artifact carries its version, population, outcome, analysis date, what changed from prior versions, and known limitations.

Earlier values that no longer belong to the current evidence base stay in version history rather than being silently overwritten. The current revision carries explicit supersessions, among them the retirement of an earlier headline AUC that turned out to be reading transcript length and a consent phrase rather than predicting the outcome.

Every result should be able to lose. If a method only produces wins, it is not a validation system.

The corpus

A large record, with the
denominators kept honest.

The research base includes 101,026 recorded hours of older-adult conversation across two stages of the same journey, 4.5 million conversations in which the older adult spoke, out of 24.2 million contacts; two-thirds of those are a single turn. That is a scale fact about recorded audio, and it is not the same thing as a joined asset. Beside the calls sits a bank of 56 published senior-survey instruments and standards and 243 distinct research sources across two literature registers (250 cited rows, deduplicated 2026-09-10; the registers include standards and regulations, so "sources" rather than "studies").

Recorded is not the same as transcribed.

Transcribed is not the same as outcome-labelled.

Sessions are not people.

Different analyses use different populations.

So every result reports its own denominator, and the two legs are never summed into one for any rate.