Project Vaani steps up inclusivity with dataset for atypical speech in three languages

Researchers at the Indian Institute of Science (IISc) are developing a dataset for atypical speech in Indian languages under Project Vaani. This initiative aims to make Automatic Speech Recognition (ASR) systems more inclusive for individuals with neurological or motor conditions.
Why it matters
Improving speech recognition for atypical patterns is a significant step toward digital accessibility and inclusivity for people with disabilities.
While speech technology for Indian languages has moved beyond the infancy stage, it still trails English and struggles with regional dialects and cultural contexts. The challenge becomes even more acute when it comes to atypical speech (speech that deviates from standard speech patterns) in Indian languages, as most Automatic Speech Recognition (ASR) systems are designed for standard speech.
Researchers at the Indian Institute of Science (IISc), Bengaluru, have been trying to address this by building a dataset of atypical speech in Indic Languages.
The Vaani Atypical Speech Corpus, a pilot initiative under Project Vaani by IISc and ARTPARK, has so far collected approximately 10 hours of atypical speech from around 40 speakers across three languages and aims to expand the corpus in the coming days.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in