tiny-audio, nanoGPT for speech-to-text
A new open-source project called 'tiny-audio' allows users to train a speech-to-text system for approximately $25. The model connects a pretrained speech encoder to an LLM, offering high accuracy and features like speaker diarization and word-level timestamps.
Why it matters
This democratizes access to high-performance speech recognition technology, allowing developers to run sophisticated AI models on consumer-grade hardware.
A speech-to-text system you can train for $25.
Tiny Audio connects a frozen, pretrained speech encoder to a pretrained LLM with a small trainable projector. The model published from this repo gets 1.8% WER on LibriSpeech test-clean and 7.4% across 12 benchmarks (11,822 samples pooled) while training only ~80M parameters. The codebase is small enough to read in an afternoon, and you can run a training loop on your laptop in about five minutes.
No install: open the live demo , record yourself or upload a file, and get a transcript.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in