Five minutes
A sealed benchmark of how well any speech engine hears the older adult on an ordinary phone line, and a certified judge that recovers what actually happened on a call from its transcript, graded against settled money. Five chapters, one minute each, with a recorded narration. Press play to hear it, or read at your own pace. Every number carries the population it was measured on, and the last chapter says what none of this claims.
What is live on this page
The benchmark’s seal and the judge’s certification are fetched from this site while you read, so the fingerprints you see are the ones the server holds now. The four transcripts in chapter three were written for this page, belong to no client, and were scored by the judge on the day the page was built; their digests are checked against the live certification in front of you.
0:00 · chapter one
The gap, on the same calls
On a two-channel recording the older adult and the agent can be scored separately by the same engine, on the same line, about the same subject. That isolates the speaker from the channel. One production telephone engine, its own per-word confidence, both sides of the same calls.
| side of the same calls | channels | mean word confidence | words under 0.5 confidence | words per minute | turns of three words or fewer |
|---|---|---|---|---|---|
| older adult | 2,364 | 0.916 | 3.6% | 37 | 45% |
| agent | 2,375 | 0.963 | 1.1% | 82 | 24% |
Low-confidence share, older adult over agent: 3.4x. On the sealed set below, one call per person, the same paired gap reads +3.1 points of low-confidence words [+2.5, +3.7] on 202 calls. Confidence is a difficulty signal that needs no reference; it is not accuracy.
Two things this does not say. It does not say older adults are harder to transcribe: by age band the difficulty runs flat to falling on every engine on file. And it does not rank any vendor. What replicates is caller-on-a-phone against agent-on-a-headset inside one call, and the older adult sits on the caller side, where the words that decide every downstream label are the ones engines disagree on most.
Measured Full tables, frame and intervals on the senior voice eval page.
1:00 · chapter two
The sealed set
A benchmark of 300 calls, a public third that vendors can be scored on and a sealed two thirds that is never released, plus a reserve of 700. The manifest holds no date of birth, no raw phone number and no text. The seal below is read from this server as you look at it.
Live: GET /api/voice-eval/seal
| age band | benchmark, public | benchmark, sealed | reserve | eligible phones |
|---|---|---|---|---|
| under 65 | 20 | 40 | 120 | 852 |
| 65 to 74 | 20 | 40 | 120 | 1,967 |
| 75 to 84 | 20 | 40 | 120 | 941 |
| 85 and over | 20 | 40 | 90 | 150 |
| age unknown | 20 | 40 | 250 | 10,944 |
| all | 100 | 200 | 700 | 14,854 |
How big is the question the benchmark answers? On the sealed calls where the two strongest engines on file both produced a senior-channel transcript, they disagree on about 31 senior words in 100 [27, 36] across 63 calls. That is engine agreement, not accuracy. Human-corrected references on file today: 0. When they land, accuracy rows replace the agreement rows rather than sit beside them, scored senior-only, disfluencies kept, every engine on the same calls at the same bar, with Comprenda’s own engines as rows in the table and not the bar.
Sealed Frame: drawn from the 25,617 calls that already carried a channel-aware engine output, not a uniform draw of the recording estate. Seventeen organisations, three senior verticals, none named. The 85-and-over band took every eligible phone.
2:00 · chapter three
The outcome judge
The judge recovers the sale or enrolment outcome from the call’s own words. It is outcome recovery, not forecasting. It answers SALE above a probability of 0.80, NO_SALE below 0.15, and abstains in between rather than guess. Four illustrative calls, written for this page in the shape of the close it was certified on. Reveal each verdict.
Live certification check
Illustrative These four calls belong to no client and were never dialled. The judge is certified on real settled calls, not on these. They show the interface, the band and the provenance; chapter four shows the certification.
3:00 · chapter four
Where it holds, and where it does not
The certification is a document the API returns with every verdict, so a consumer always sees the bar the judge was held to. The headline is one prescription-delivery buyer’s settled calls, after the head’s training cut, scored on the transcript engine it was certified with.
Where it is weaker, stated with the number. On calls of 450 seconds or less, where 3% of calls settle, recall is 0.068: the judge misses most short-call sales, and 88% of settled sales sit above 450 seconds, where recall is 0.997. On 57 gradeable long calls that never enrolled, it said SALE 31 times: the words said enrolled and the money said no, which is a cancellation, a decline at the pharmacy, or a missed postback, and it is the reason the judge is reconciled against settled money by buyer every day rather than trusted.
Where it does not hold at all. The same judge, pointed at three other buyers’ calls with human-certified sales, recognised 0 of 27. The fourth transcript in chapter three is that finding in one call: a sale written in a vocabulary the certified population never uses reads as NO_SALE with high confidence, not as an abstain. So a certificate is issued per buyer, per vertical, per call direction and per transcript engine, and never extended: of 63 population rows in the table, 8 are certified today. The bar for a row is at least 150 settled sales and 150 settled non-sales, recall and precision at 0.90 with a lower bound of 0.85, 80% gradeable, and an area under the curve of 0.95.
Certified Null on other buyers Source documents: the certification the API serves, and the ledger rows of 6 and 7 September on the buyer-holdout test.
4:00 · chapter five
What it is for
Voice-agent platforms entering Medicare, pharmacy, insurance or medical-alert outreach, where the caller is over 65 and the regulator listens: submit engine output for the public split and get a row beside every other, including ours. Speech vendors whose published error rates were measured on younger speakers on headsets. Contact-centre suites and labs that need a settled-outcome judge certified on their own population, not a generic one: the bar in chapter four is the bar, and a population without a row is uncertified by construction.
# open: the benchmark seal (fingerprint, seed, counts, funnel)
curl -s https://comprenda.ai/api/voice-eval/seal
# key-gated: a verdict on your own transcript; the response never carries a word of it
curl -s -X POST https://comprenda.ai/api/judge/outcome \
-H "X-Comprenda-Key: $KEY" -H "Content-Type: application/json" \
-d '{"vertical":"rx","turns":[{"speaker":"agent","text":"..."},{"speaker":"caller","text":"..."}]}'