These startups are chasing the next big thing in LLMs

The article explores the limitations of transformer-based neural networks, which currently power most large language models. It highlights a new wave of startups attempting to innovate beyond these architectures to address fundamental flaws in how current AI models process information.
Why it matters
As the foundation of modern AI, the potential transition away from transformers could disrupt the current AI industry landscape and lead to more efficient, capable models.
Meet the new kids nipping at the heels of the AI giants.
MIT Technology Review ’s What’s Next series looks across industries, trends, and technologies to give you a first look at the future. You can read the rest of them here .
Way back in the summer of 2017, AI researchers at Google put out a paper called “Attention Is All You Need,” in which they described a new type of neural network called a transformer. It proved to be very good at processing long sequences of data, especially text.
Nine years on, transformers are the engines inside every major large language model on the market. “The entire AI industry is built on transformers,” says Justin Dangel, cofounder and CEO of the AI startup Subquadratic. “They are one of the most important innovations in the history of computer science, and they’ve changed the world.”
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in