Five minutes

Two instruments for AI that talks
to older adults by phone.

A sealed benchmark of how well any speech engine hears the older adult on an ordinary phone line, and a certified judge that recovers what actually happened on a call from its transcript, graded against settled money. Five chapters, one minute each, with a recorded narration. Press play to hear it, or read at your own pace. Every number carries the population it was measured on, and the last chapter says what none of this claims.

What is live on this page

The benchmark’s seal and the judge’s certification are fetched from this site while you read, so the fingerprints you see are the ones the server holds now. The four transcripts in chapter three were written for this page, belong to no client, and were scored by the judge on the day the page was built; their digests are checked against the live certification in front of you.

0:00

0:00 · chapter one

The gap, on the same calls

Same line. Same topic. One speaker is over 65.

On a two-channel recording the older adult and the agent can be scored separately by the same engine, on the same line, about the same subject. That isolates the speaker from the channel. One production telephone engine, its own per-word confidence, both sides of the same calls.

side of the same callschannelsmean word confidencewords under 0.5 confidencewords per minuteturns of three words or fewer
older adult2,3640.9163.6%3745%
agent2,3750.9631.1%8224%

Low-confidence share, older adult over agent: 3.4x. On the sealed set below, one call per person, the same paired gap reads +3.1 points of low-confidence words [+2.5, +3.7] on 202 calls. Confidence is a difficulty signal that needs no reference; it is not accuracy.

Two things this does not say. It does not say older adults are harder to transcribe: by age band the difficulty runs flat to falling on every engine on file. And it does not rank any vendor. What replicates is caller-on-a-phone against agent-on-a-headset inside one call, and the older adult sits on the caller side, where the words that decide every downstream label are the ones engines disagree on most.

Measured  Full tables, frame and intervals on the senior voice eval page.

1:00 · chapter two

The sealed set

One call per person, two channels,
frozen and fingerprinted.

A benchmark of 300 calls, a public third that vendors can be scored on and a sealed two thirds that is never released, plus a reserve of 700. The manifest holds no date of birth, no raw phone number and no text. The seal below is read from this server as you look at it.

Live: GET /api/voice-eval/seal

fetching…
age bandbenchmark, publicbenchmark, sealedreserveeligible phones
under 652040120852
65 to 7420401201,967
75 to 842040120941
85 and over204090150
age unknown204025010,944
all10020070014,854

How big is the question the benchmark answers? On the sealed calls where the two strongest engines on file both produced a senior-channel transcript, they disagree on about 31 senior words in 100 [27, 36] across 63 calls. That is engine agreement, not accuracy. Human-corrected references on file today: 0. When they land, accuracy rows replace the agreement rows rather than sit beside them, scored senior-only, disfluencies kept, every engine on the same calls at the same bar, with Comprenda’s own engines as rows in the table and not the bar.

Sealed  Frame: drawn from the 25,617 calls that already carried a channel-aware engine output, not a uniform draw of the recording estate. Seventeen organisations, three senior verticals, none named. The 85-and-over band took every eligible phone.

2:00 · chapter three

The outcome judge

Read a transcript. Say what happened.
Abstain when you cannot.

The judge recovers the sale or enrolment outcome from the call’s own words. It is outcome recovery, not forecasting. It answers SALE above a probability of 0.80, NO_SALE below 0.15, and abstains in between rather than guess. Four illustrative calls, written for this page in the shape of the close it was certified on. Reveal each verdict.

Live certification check

The live certification check is offline while a public version without client or vendor names is prepared.

Illustrative  These four calls belong to no client and were never dialled. The judge is certified on real settled calls, not on these. They show the interface, the band and the provenance; chapter four shows the certification.

3:00 · chapter four

Where it holds, and where it does not

Certified on one population, out of sample by time,
on the calls where the sales actually live.

The certification is a document the API returns with every verdict, so a consumer always sees the bar the judge was held to. The headline is one prescription-delivery buyer’s settled calls, after the head’s training cut, scored on the transcript engine it was certified with.

2,901settled calls, held out by time
0.973area under the curve
0.961SALE precision
0.949SALE recall
0.977NO_SALE precision
92%gradeable (not abstained)
0.972agreement with the settled label
32.1%base rate on this population

Where it is weaker, stated with the number. On calls of 450 seconds or less, where 3% of calls settle, recall is 0.068: the judge misses most short-call sales, and 88% of settled sales sit above 450 seconds, where recall is 0.997. On 57 gradeable long calls that never enrolled, it said SALE 31 times: the words said enrolled and the money said no, which is a cancellation, a decline at the pharmacy, or a missed postback, and it is the reason the judge is reconciled against settled money by buyer every day rather than trusted.

Where it does not hold at all. The same judge, pointed at three other buyers’ calls with human-certified sales, recognised 0 of 27. The fourth transcript in chapter three is that finding in one call: a sale written in a vocabulary the certified population never uses reads as NO_SALE with high confidence, not as an abstain. So a certificate is issued per buyer, per vertical, per call direction and per transcript engine, and never extended: of 63 population rows in the table, 8 are certified today. The bar for a row is at least 150 settled sales and 150 settled non-sales, recall and precision at 0.90 with a lower bound of 0.85, 80% gradeable, and an area under the curve of 0.95.

Certified Null on other buyers  Source documents: the certification the API serves, and the ledger rows of 6 and 7 September on the buyer-holdout test.

4:00 · chapter five

What it is for

Anyone whose model talks to older adults by phone.

Voice-agent platforms entering Medicare, pharmacy, insurance or medical-alert outreach, where the caller is over 65 and the regulator listens: submit engine output for the public split and get a row beside every other, including ours. Speech vendors whose published error rates were measured on younger speakers on headsets. Contact-centre suites and labs that need a settled-outcome judge certified on their own population, not a generic one: the bar in chapter four is the bar, and a population without a row is uncertified by construction.

# open: the benchmark seal (fingerprint, seed, counts, funnel)
curl -s https://comprenda.ai/api/voice-eval/seal

# key-gated: a verdict on your own transcript; the response never carries a word of it
curl -s -X POST https://comprenda.ai/api/judge/outcome \
  -H "X-Comprenda-Key: $KEY" -H "Content-Type: application/json" \
  -d '{"vertical":"rx","turns":[{"speaker":"agent","text":"..."},{"speaker":"caller","text":"..."}]}'
Not claimed
That any engine is accurate or inaccurate on older adults. There is no human reference yet, and the page says so before any row.
Not claimed
That older adults are harder to transcribe. The gap that replicates is the caller channel against the agent channel inside one call.
Not claimed
That the judge transfers. It is certified per population and returned 0 of 27 on buyers it never saw.
Not claimed
That a verdict predicts anything. The judge recovers what happened; forecasting is a different instrument with its own page.
Not yet
Vendor rows on the public split. They appear when human-corrected references exist, and the release conditions for the public split are a decision still open.