Noesa
← The Casebook

Global · 2021–2025

Speech recognition understood a third of them. Then it was trained on their own voices.

Off-the-shelf models worked for 35% of speakers with impaired speech. After about eighteen minutes of their own recordings, 79% — though for severe impairment, still fewer than half.

Published 2 September 2026

What happened

Automatic speech recognition is quoted at accuracies around 95%, and for millions of people it simply does not work. Speech affected by ALS, Parkinson's disease, cerebral palsy, stroke, Down syndrome or vocal cord paralysis falls outside what these systems were trained on, and the failure is not marginal: across 432 speakers with disordered speech, two speaker-independent commercial systems recorded median word error rates of 31.5% and 29.4% [1].

That is the population that would benefit most from voice control, because speech impairments frequently accompany mobility impairments — the people for whom talking to a device is not a convenience but the interface [2].

Google's Project Euphonia set out to collect the data that did not exist. By February 2025 the corpus held over 1.5 million utterances from around 3,000 speakers, recorded remotely on the participants' own phones, tablets and computers, in their own homes, with consent obtained for research and product use [1][3].

Trained on a speaker's own recordings, the same underlying system changed character. Across those 432 speakers the personalized models reached a median word error rate of 4.6%, against 31.5% and 29.4% for the off-the-shelf systems [1].

On a subset of 116 speakers, the researchers also compared the models against three speech-language pathologists experienced with dysarthria, transcribing 30 randomly selected phrases. The personalized models beat the expert human listeners, with median and maximum recognition accuracy gains of 9% and 80% [1].

A separate study put a number on how much speech that takes. For 195 speakers, an unadapted model hit the target error rate for 35% of them. Personalizing on 250 recorded utterances — about 18 minutes on average — raised that to 79%. Even 50 utterances, roughly three and a half minutes, reached 63% [2].

The limits are in the same papers. Broken down by severity, personalization on 250 utterances reached the target for 96% of speakers with mild impairment, 81% with moderate, and 47% with severe; using every recording a speaker had made lifted severe speakers only to 69% [2]. In the larger cohort the severe group's personalized median error rate stayed at 13%, against 3.8% pooled across typical, mild and moderate speech, and 7% of all speakers remained above a 24% error rate [1].

Two other findings complicate the good news. What predicted a poor result was often not the impairment but the recording — low signal-to-noise ratio and fewer training utterances went with high error rates even among mildly impaired speakers [1]. And the results are for short phrases of two to four words in a home-automation vocabulary; the authors say plainly that longer utterances and spontaneous conversation need further evaluation [1].

The work has since been extended beyond English, with data collection in Spanish, French, Japanese and Hindi — 132 speakers so far, unevenly distributed, and the authors note that their evaluation covered only two of those languages and that rater confidence was lower where their clinicians were not fluent [3].

Where it helped

This is a genuine and unglamorous good. Nobody invented a new architecture; they collected data from people the industry had not collected data from, then fitted the model to individuals. The result cleared not just the commercial baseline but expert human listeners — clinicians who work with dysarthria daily — on short phrases [1]. And the research is honest in the way that makes it usable: it publishes the severity breakdown where the method works least well, names recording quality rather than disability as a leading cause of failure, and states that conversational speech is not yet evaluated [1][2].

Where it can still burn

"It can be personalized" is not "it works", and the gap falls exactly where the need is greatest: fewer than half of speakers with severe impairment reached the target on 250 utterances, and barely two thirds did on everything they had recorded [2]. There is a bill attached, too, and it is presented to the person with the impairment — recording enough phrases is a real investment of time and effort for someone who may also have motor or cognitive impairments [2]. A procurement decision that reads "98% of our users are served" from a headline figure, or a product that ships personalization as a checkbox and leaves the recording to the customer, has taken the encouraging half of this record and dropped the rest.

The tell

When a tool fails for you, find out whether it can be trained on you before you accept the verdict — and ask how much of your own data that takes. "This does not work for people like me" and "nobody has fitted this to me yet" are different problems with different fixes.

Most people read a bad result as a fact about the tool, or worse, about themselves, and stop. Often it is a fact about the default — a model fitted to an average that you are not near. The second half of the question is the one that decides anything: if adaptation needs three minutes of your data, that is a Tuesday; if it needs six hours of it, someone has moved the work onto you and called it a feature. Ask both, in that order, and ask who is paying for the data. The same two questions apply to a dictation tool that mangles your accent, a model that misreads your team's documents, and any vendor whose demo was not run on anything resembling your input.

Share this case

The image has the link printed on it, so it still leads back here.

The check is a habit, and habits are trained. Using AI, day by day is about directing AI at your actual work rather than accepting its defaults — including when to adapt a tool to your material instead of abandoning it.

Sources

Every source below was opened and read. Last verified 2 September 2026.

  1. [1] Automatic Speech Recognition of Disordered Speech: Personalized models outperforming human listeners on short phrasesGreen et al., Proc. Interspeech 2021 (ISCA), 30 August 2021
  2. [2] Personalized Automatic Speech Recognition Trained on Small Disordered Speech DatasetsTobin & Tomanek, Google Research (arXiv:2110.04612), 9 October 2021
  3. [3] Project Euphonia: advancing inclusive speech recognition through expanded data collection and evaluationMartin et al., Frontiers in Language Sciences, 20 June 2025