Senior voice eval

Can your speech stack hear a 78-year-old
on an ordinary phone line?

Every model in a voice stack consumes a transcript. On a two-channel call the older adult and the agent can be scored separately by the same engine, on the same line, about the same topic. That isolates the speaker from the channel. The older adult’s words are the ones that decide every downstream label, and they are the ones engines disagree on most.

Reference status

Human-corrected reference transcripts on file: 0. Every word-error row on this page is engine agreement with another engine on the same calls, not accuracy. It says which engine another engine agrees with. Accuracy rows appear when human-corrected references land, and they replace these rather than sit beside them.

The gap, on the same calls

Same line. Same topic. One speaker is over 65.

One production telephone engine, scored on the older adult’s channel and the agent’s channel of the same recorded calls. Confidence is the engine’s own per-word self-report. It is a difficulty signal that needs no reference and scales to every call; it is not accuracy.

side of the same callschannelsmean word confidencewords under 0.5 confidencewords per minuteturns of three words or fewer
older adult2,3640.9163.6%3745%
agent2,3750.9631.1%8224%

Low-confidence share, older adult over agent: 3.4x. The agent is the within-call comparison: a speaker on a headset, same line, same subject. What replicates is caller-on-a-phone against agent-on-a-headset inside one call. On the sealed set below the same gap reads +3.1 points of low-confidence words [+2.5, +3.7] on 202 calls. By age band the difficulty runs flat to falling, so this is not evidence that older adults are harder to transcribe. It is evidence that the caller channel, where the older adult sits, is where the error lives, and that a whole-call number hides it.

Published leaderboards score whole calls, read speech, or younger speakers on headsets. None score the older adult on the phone, by age band, with the disfluencies kept.

The metric

Senior-only, disfluency-preserving word error,
against human-corrected references.

Senior-only

Only the older adult’s words are scored. Agent words are excluded from numerator and denominator. Whole-call error is reported as a secondary row so vendors can compare with their own leaderboards.

Disfluency-preserving

Fillers, restarts, repeats and repair tokens such as “pardon?” are kept in the reference and count as words. An engine that drops them is penalised, because they are the signal that the person did not follow.

Same calls, same bar

Every engine is scored on the same sealed calls, the same references, the same normaliser and the same bootstrap by phone number, and its row is published beside the others. Comprenda’s own engines are rows in the table, not the bar.

Results are reported by age band, with intervals from resampling people rather than calls, so an older adult with several calls is one draw. Age is banded at source; the benchmark never holds a date of birth.

What is on file today

Provisional rows, labelled as such.

Twenty-eight calls, scored against a provisional reference produced by a frontier audio model. These rows measure agreement with that reference, not accuracy, and they are here so the shape of the gap is visible before the human references arrive.

enginewhole-call word error, mediansenior-only word error, mediansenior-only over whole-call
production telephone vendor5.3%11.4%2.2x
fused (vendor spine plus a four-engine vote)7.8%20.6%2.6x
second telephone engine12.1%26.1%2.2x

The row that matters is the middle column. On every engine the older adult’s words carry roughly twice the error of the call as a whole, which is what a whole-call leaderboard hides.

The sealed set

One call per person, two channels,
joined to what settled.

Frozen on 14 September 2026: one call per phone number, two-channel audio, the older adult’s channel present with at least ten words, calls of one to thirty minutes in senior verticals, the oldest band oversampled. A benchmark of 300 calls, split into a public third vendors can be scored on and a sealed two thirds that is never released, plus a reserve of 700. The manifest holds no date of birth, no raw phone number and no text. Its fingerprint is b36619a8d1b34838b35a56d72752bc0fe564f522c42ecb22bd2a2e642fe8ee63.

age bandbenchmark, publicbenchmark, sealedreserveeligible phones
under 652040120852
65 to 7420401201,967
75 to 842040120941
85 and over204090150
age unknown204025010,944
all10020070014,854

The frame, stated before the number: the set was drawn from the 25,617 calls that already carry a channel-aware engine output, which earlier studies chose for their own reasons. It is not a uniform draw of the recording estate. Seventeen organisations and three product verticals are represented; none is named. The 85-and-over band took every eligible phone. Age is banded at source and reported as unknown where the join was empty, never dropped.

What makes this set different from a speech corpus is the join. Each call sits beside whether that person’s enrolment settled, so a transcription error can be priced in the outcome it changed, not only counted as a word.

Who this is for

Anyone whose model talks to older adults by phone.

Voice-agent platforms entering Medicare, pharmacy, insurance or medical-alert outreach, where the caller is over 65 and the regulator listens. Speech vendors whose published error rates were measured on younger speakers. Contact-centre suites adding an AI agent to a floor that already serves this population. Submit engine output for the public split; Comprenda runs the harness and publishes the row, yours beside every other, including ours.