Show HN: Whisper transcribes 70-year-olds more accurately than 20-year-olds
Technical analysis of OpenAI's Whisper model reveals that it transcribes speech from elderly individuals more accurately than from younger adults. However, the study notes that elderly speakers are frequently interrupted by systems using fixed silence thresholds for turn-taking.
Why it matters
This highlights a critical design flaw in voice-activated systems that negatively impacts accessibility for older users.
Voice agents are being pointed at elderly callers, and the assumed risk is that speech recognition will not hear them. That assumption is wrong, and it is hiding the failure that is actually happening.
Measured on 2,760 Common Voice clips, matched between age brackets on accent, gender and speaker so the only thing varying is age, and checked against a second 3,189-clip draw that controls for none of it:
word error rate premature cutoff (700 ms) twenties n=920 6.53% 8.0% sixties n=920 5.23% -1.31pp * 19.7% +11.6pp * seventies n=920 4.67% -1.86pp * 16.6% +8.5pp * * speaker-bootstrapped 95% interval excludes zero Whisper transcribes older speakers more accurately, not less. And where a stack endpoints on a fixed silence threshold, those same speakers get talked over two to two and a half times as often.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in