Senior voice eval
Every model in a voice stack consumes a transcript. On a two-channel call the older adult and the agent can be scored separately by the same engine, on the same line, about the same topic. That isolates the speaker from the channel. The older adult’s words are the ones that decide every downstream label, and they are the ones engines disagree on most.
Reference status
Human-corrected reference transcripts on file: 0. Every word-error row on this page is engine agreement with another engine on the same calls, not accuracy. It says which engine another engine agrees with. Accuracy rows appear when human-corrected references land, and they replace these rather than sit beside them.
The gap, on the same calls
One production telephone engine, scored on the older adult’s channel and the agent’s channel of the same recorded calls. Confidence is the engine’s own per-word self-report. It is a difficulty signal that needs no reference and scales to every call; it is not accuracy.
| side of the same calls | channels | mean word confidence | words under 0.5 confidence | words per minute | turns of three words or fewer |
|---|---|---|---|---|---|
| older adult | 2,364 | 0.916 | 3.6% | 37 | 45% |
| agent | 2,375 | 0.963 | 1.1% | 82 | 24% |
Low-confidence share, older adult over agent: 3.4x. The agent is the within-call comparison: a speaker on a headset, same line, same subject. What replicates is caller-on-a-phone against agent-on-a-headset inside one call. On the sealed set below the same gap reads +3.1 points of low-confidence words [+2.5, +3.7] on 202 calls. By age band the difficulty runs flat to falling, so this is not evidence that older adults are harder to transcribe. It is evidence that the caller channel, where the older adult sits, is where the error lives, and that a whole-call number hides it.
Published leaderboards score whole calls, read speech, or younger speakers on headsets. None score the older adult on the phone, by age band, with the disfluencies kept.
The metric
Senior-only
Only the older adult’s words are scored. Agent words are excluded from numerator and denominator. Whole-call error is reported as a secondary row so vendors can compare with their own leaderboards.
Disfluency-preserving
Fillers, restarts, repeats and repair tokens such as “pardon?” are kept in the reference and count as words. An engine that drops them is penalised, because they are the signal that the person did not follow.
Same calls, same bar
Every engine is scored on the same sealed calls, the same references, the same normaliser and the same bootstrap by phone number, and its row is published beside the others. Comprenda’s own engines are rows in the table, not the bar.
Results are reported by age band, with intervals from resampling people rather than calls, so an older adult with several calls is one draw. Age is banded at source; the benchmark never holds a date of birth.
What is on file today
Twenty-eight calls, scored against a provisional reference produced by a frontier audio model. These rows measure agreement with that reference, not accuracy, and they are here so the shape of the gap is visible before the human references arrive.
| engine | whole-call word error, median | senior-only word error, median | senior-only over whole-call |
|---|---|---|---|
| production telephone vendor | 5.3% | 11.4% | 2.2x |
| fused (vendor spine plus a four-engine vote) | 7.8% | 20.6% | 2.6x |
| second telephone engine | 12.1% | 26.1% | 2.2x |
The row that matters is the middle column. On every engine the older adult’s words carry roughly twice the error of the call as a whole, which is what a whole-call leaderboard hides.
The sealed set
Frozen on 14 September 2026: one call per phone number, two-channel audio, the older adult’s
channel present with at least ten words, calls of one to thirty minutes in senior verticals, the oldest band
oversampled. A benchmark of 300 calls, split into a public third vendors can be scored on and a sealed two thirds
that is never released, plus a reserve of 700. The manifest holds no date of birth, no raw phone number and no
text. Its fingerprint is b36619a8d1b34838b35a56d72752bc0fe564f522c42ecb22bd2a2e642fe8ee63.
| age band | benchmark, public | benchmark, sealed | reserve | eligible phones |
|---|---|---|---|---|
| under 65 | 20 | 40 | 120 | 852 |
| 65 to 74 | 20 | 40 | 120 | 1,967 |
| 75 to 84 | 20 | 40 | 120 | 941 |
| 85 and over | 20 | 40 | 90 | 150 |
| age unknown | 20 | 40 | 250 | 10,944 |
| all | 100 | 200 | 700 | 14,854 |
The frame, stated before the number: the set was drawn from the 25,617 calls that already carry a channel-aware engine output, which earlier studies chose for their own reasons. It is not a uniform draw of the recording estate. Seventeen organisations and three product verticals are represented; none is named. The 85-and-over band took every eligible phone. Age is banded at source and reported as unknown where the join was empty, never dropped.
What makes this set different from a speech corpus is the join. Each call sits beside whether that person’s enrolment settled, so a transcription error can be priced in the outcome it changed, not only counted as a word.
Who this is for
Voice-agent platforms entering Medicare, pharmacy, insurance or medical-alert outreach, where the caller is over 65 and the regulator listens. Speech vendors whose published error rates were measured on younger speakers. Contact-centre suites adding an AI agent to a floor that already serves this population. Submit engine output for the public split; Comprenda runs the harness and publishes the row, yours beside every other, including ours.