Voice agents are being pointed astatine aged callers, and the assumed consequence is that speech nickname will not perceive them. That presumption is wrong, and it is hiding the nonaccomplishment that is really happening.
Measured connected 2,760 Common Voice clips, matched betwixt property brackets connected accent, gender and speaker truthful the only point varying is age, and checked against a second 3,189-clip tie that controls for nary of it:
Whisper transcribes older speakers much accurately, not less. And wherever a stack endpoints connected a fixed soundlessness threshold, those aforesaid speakers get talked over 2 to 2 and a half times arsenic often.
That 2nd uncovering has a caveat that arrived aft publication and is worth reading earlier quoting the number: a semantic move exemplary closes astir of the gap. Measured against Pipecat's smart-turn v3 connected the aforesaid clips, +11.6pp becomes +5.9pp and stops excluding zero (section 3). The fixed-threshold result describes a existent and communal configuration, not each configuration.
python3 bench/run.py reproduces both. No API key, nary spend.
1. The nickname punishment is not there
Large-v3, wideband, speaker-level 95% intervals:
Every correction type falls pinch age, truthful this is not 1 class masking another:
Deletions successful peculiar do not rise, which is the consequence you would expect if quiet aliases breathy reside were being dropped. It is not being dropped.
It is not a Whisper artifact. The evident objection is that Whisper's decoder is simply a connection model, truthful it mightiness beryllium repairing older speakers' word choices alternatively than proceeding them better. So the aforesaid clips were re-run through wav2vec2, which is axenic CTC: frame-wise, greedy, nary decoder and nary implicit LM.
Absolute WER is overmuch higher for wav2vec2 (LibriSpeech-only training, nary LM), so only the between-bracket comparison transfers. The effect is larger location and still excludes zero, which puts it successful the acoustics alternatively than successful a decoder.
2. The turn-taking punishment is large
A sound supplier decides the caller has vanished by waiting for a fixed agelong of silence. A region inside an utterance that exceeds that period is heard as the extremity of the turn, and the supplier starts talking complete personification mid-sentence.
WER is unsighted to this. The words the exemplary did person tin beryllium transcribed perfectly while the caller is trim disconnected each time.
Share of utterances containing an soul region astatine slightest this long:
All six age-versus-baseline differences exclude zero. At 700 sclerosis the spread is +11.6pp [+7.7, +15.7] for the sixties and +8.5pp [+4.0, +13.6] for the seventies.
The system is successful the timing. Older speakers return twice arsenic galore internal pauses and walk doubly arsenic agelong successful them:
It is not a cleanable gradient. The sixties are trim disconnected somewhat much than the seventies and their intervals overlap. This sounds arsenic an effect that arrives by 60 and plateaus, not a consecutive line, and it is drawn that measurement alternatively than smoothed.
The eighties, tally separately because matching against them would person shrunk every bracket eightfold, are the sharpest case. Only 14 speakers exist, truthful the intervals are wide — and the effect clears them anyway:
At 700 sclerosis that is simply a 5.4x gap. Their WER, meanwhile, is 6.24% against 6.01% — a quality of +0.15pp whose interval comfortably includes zero. The two findings diverge further pinch age: nickname stays level while turn-taking gets steadily worse.
The consequence survives a re-draw. Running the full benchmark again without accent matching — 3,189 clips, 1,434 speakers, 30 accents alternatively of 8 — reproduces the cutoff rates almost exactly:
This is simply a re-draw from 1 corpus, not an independent replication: the two samples stock 52% of their speakers, though only 18% of their clips. It shows the numbers are not an artifact of 1 peculiar tie aliases of the accent matching. It does not show they generalise beyond Common Voice.
3. A semantic move exemplary closes astir of the gap
Mark Backman of Daily/Pipecat publication an earlier type of this and pointed out that it described the incorrect thing: accumulation stacks do not endpoint connected a fixed VAD threshold. Pipecat's default is smart-turn, a semantic exemplary that listens to the waveform and grants much clip erstwhile the move sounds unfinished.
He is right, truthful smart-turn v3 was measured connected the identical sample. For each clip, the audio up to an soul region is fed to the exemplary and it is asked whether the move is complete. The speaker demonstrably continues, so "complete" is simply a mendacious cutoff.
The spread astir halves and stops excluding zero. Positive power connected whole utterances is level astatine 90-91% crossed brackets.
Two things this does not say. The absolute 76-82% complaint is not an error rate: galore soul pauses are morganatic clause boundaries wherever a move could plausibly end, and without quality labels connected which prefixes sound complete, only the between-bracket comparison is interpretable. And "includes zero" is not "no effect" — some constituent estimates enactment positive, and 86 seventies speakers cannot resoluteness 4 points either way.
The applicable reading: if you endpoint connected a fixed threshold, the property spread is real and large. If you usage a semantic move model, astir of it goes away. The published smart-turn benchmark stratifies 31,527 samples crossed 23 languages but not by speaker age, and its training operation leans connected synthetic TTS, which does not region the measurement an eighty-year-old does.
4. What a person's ain reside sound costs a drift detector
Several products now connection regular telephone check-ins for older adults that declare to flag cognitive diminution from sound biomarkers. Validating that needs gated clinical corpora. But a anterior mobility needs nary objective labels and bounds the claim from below: how overmuch does 1 patient person's reside alteration between utterances? A drift detector tin only spot alteration that clears the speaker's own noise.
Measured connected 36 speakers pinch 40+ clips each:
Converted to the sample needed to resoluteness a 10% alteration astatine 80% power:
Pause features — the astir often cited sound biomarker — alteration by about 100% wrong the aforesaid speaker, often wrong 1 sitting. Detecting a 10% shift in full region clip takes astir 3 weeks of regular calls per reading, so a "six-week trend" is 2 aliases 3 noisy measurements. Speech complaint and duration are comfortably usable.
Both directions of correction are stated: Common Voice clips from 1 contributor are often a azygous sitting, truthful existent day-to-day variance is larger; and utterances wrong 1 telephone are correlated, truthful dividing by 40 overstates the effective sample. Both push the existent request up. These are floors.
5. The accent confound is real, and it does not thrust the result
Common Voice is globally crowdsourced and its younger contributors skew non-native. The twenties bracket is 11.9% India-and-South-Asia English and 46.9% autochthonal anglophone; the sixties are 64.4%. Whisper is worse connected non-native English, truthful property and accent are genuinely entangled successful this corpus, and an uncontrolled comparison has an evident replacement explanation.
Brackets are truthful matched connected the (accent, gender) pair, holding both identical by construction: 920 clips per bracket, 8 accents, 392/528 male/female in each bracket.
I expected that to alteration the answer. It does not. Running the benchmark both ways, connected samples whose accent creation could hardly beryllium much different:
The largest disagreement is 0.44pp. Both arms put the sixties and seventies below the twenties, and successful some the quality excludes zero. Matching is still the correct point to do — it removes a unrecorded replacement mentation and it makes the inclination monotonic — but the uncovering does not remainder connected it.
This conception primitively claimed the opposite, connected the spot of a 40-clip pilot successful which the twenties scored 10.54%. At afloat sample that fig is 6.60%. The aviator was sound and the communicative built connected it was wrong.
6. What this tin and cannot claim
Common Voice's older speakers are volunteers. They chose to beryllium down astatine a computer and grounds themselves for Mozilla. They are tech-comfortable and almost surely healthier of sound than the median 75-year-old connected a post-discharge call. Dysarthria, post-stroke reside and cognitive diminution are absent from this corpus by construction.
So this measures healthy aging, not objective aging. The correct reference of finding 1 is "age unsocial does not break recognition", not "recognition is fine for aged patients". Those are different claims and only the first is supported here. A follow-up connected a disordered-speech corpus is the honorable next step.
Two narrower limits. Read reside is not conversational speech, and someone reading a punctual pauses otherwise than personification answering a mobility — though that cuts against uncovering 2 being an artifact, since publication reside should if anything understate earthy pausing. And the eighties bracket has 27 speakers in the full split, truthful it is reported separately alternatively than folded into the main comparison, wherever matching against it would person shrunk each bracket eightfold. Its intervals are correspondingly wide.
7. Confounds that were checked and came backmost clean
Reported because a scholar will ask, not because they changed anything.
-
Recording quality. Median SNR 53.2 / 55.7 / 55.9 dB crossed brackets. No equipment disadvantage, truthful the WER consequence is not a microphone result.
-
Sentence length. Median 11 words for the twenties and sixties, 10 for the seventies. Short utterances are genuinely harder present — 7.72% WER astatine 1-7 words against 4.24% astatine 14+ — truthful the seventies transportation the harder sentences, and the effect survives holding magnitude fixed:
words twenties sixties seventies 1-7 8.75% 8.33% 6.04% 8-10 5.82% 6.37% 4.93% 11-13 6.75% 4.63% 4.50% 14-40 5.66% 3.23% 3.89% -
Clip validation. Older clips transportation a lower organization downvote rate (10.1% against 14.2%).
-
Speaker prolificacy. One contributor holds 9,792 clips successful a azygous shard; half of each 60+ audio successful the divided comes from astir 7 people. Capped at 25 clips per speaker, and each intervals resample speakers alternatively than clips.
8. Bugs the harness caught successful itself
The endpoint uncovering was astir an artifact. A relative-energy VAD reported the eighties being trim disconnected astatine 42.9%. WebRTC VAD — what accumulation stacks actually tally — disagreed connected 36% of those clips and put the median longest pause at 345 sclerosis wherever the power detector said 662 ms. Breathy trailing-off speech falls nether an power floor, and older reside is precisely what is breathy, so the inexpensive detector's correction was correlated pinch the adaptable nether study. WebRTC is now superior and some are recorded; connected the published sample they work together on 80-90% of clips and springiness the aforesaid conclusion.
Accent matching silently undid the gender balance. Taking an accent-wise subset request not sphere the male/female split, and 1 tally came retired 54/46 in the sixties against 50/50 successful the seventies. Matching connected the (accent, gender) pair costs 25 clips per bracket and fixes it exactly.
A fp16 shortcut was verified alternatively than assumed. The tally uses fp16 connected MPS because it is 2.6x faster than fp32 connected CPU. A quantised Whisper KV cache is capable of taking large-v3 from 1.91% WER to 100%, truthful the shortcut was checked against a reference: identical 0.0274 WER nether fp32/cpu, fp32/mps and fp16/mps. tests/test_precision.py.
The corpus is the CC-0 Common Voice 17 English group via the ungated fsicoli/common_voice_17_0 mirror. Shard rank is computable from train.tsv statement order, truthful only the shards the sample needs are downloaded (~19 GB of 45 GB). The tally checkpoints each batch to JSONL and resumes where it stopped.
METHOD.md has the afloat design, including why a same-sentence paired creation is impossible successful this corpus.
| src/agegap/sample.py | stratified draw: speaker cap, gender balance, accent matching |
| src/agegap/channel.py | G.711 telephony simulation, mu-law and packet loss |
| src/agegap/metrics.py | SNR, speaking rate, and the endpoint-cutoff measure |
| src/agegap/stats.py | pooled WER and speaker-level bootstrap intervals |
| src/agegap/asr.py | Whisper adapter, instrumentality and precision selection |
| bench/run.py | the run: chunked, checkpointed, fail-closed |
| bench/analyze.py | the tables above, positive the controlled/uncontrolled contrast |
| bench/replicate_ctc.py | the wav2vec2 architecture control |
| bench/smart_turn_eval.py | Pipecat smart-turn v3 connected the aforesaid age-matched clips |
| bench/drift_floor.py | within-speaker noise, and what it costs a drift detector |
| tests/ | 26 tests, including 1 per bug successful conception 8 |
English (US) ·
Indonesian (ID) ·