Can gzip be a language model?
This article explores the theoretical connection between data compression algorithms and language modeling. It demonstrates that gzip, by identifying patterns to compress data, inherently functions as a basic prediction model capable of generating text.
Why it matters
It provides a unique perspective on information theory and the fundamental mechanics behind how AI models predict and generate sequences.
A while back I wrote about language modeling without neural networks , where I generated Shakespeare with an unbounded n-gram model: no weights, no training, just counting. Fortuitously, I came across the paper Language Modeling is Compression , which mentioned the compression–prediction equivalence :
every prediction model is inherently a compressor, and all compression algorithms are prediction models .
This led to the natural question: can gzip do language modeling? 1 No neural network, no learned parameters, nothing. Just the compressor that ships with your operating system. You prime it with a corpus, give it a normal text prompt, and it continues that prompt by searching for the byte sequences that compress best. Here's some real, unedited output after priming it on tiny Shakespeare:
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in