The Download: the next big thing in LLMs and how AI academic research is shifting

MIT Technology Review explores the limitations of transformer-based neural networks, which have served as the foundation for modern LLMs. The article discusses emerging research aimed at creating more efficient and capable architectures to overcome current scaling bottlenecks.
Why it matters
As transformer models reach their performance limits, the next generation of AI architecture will determine the future speed and efficiency of global AI development.
Plus: Nvidia has secured $500 billion from Wall Street for AI infrastructure.
This is today's edition of The Download , our weekday newsletter that provides a daily dose of what's going on in the world of technology.
Nine years after Google researchers introduced the transformer, this family of neural networks has become the engine inside every major large language model. But transformers are starting to show their age.
As LLMs get bigger and better, transformers have become a bottleneck. Their dense attention mechanism becomes increasingly expensive as the amount of text grows, and they’re not great at keeping track of a lot of information at once.
Here are four new ideas for how to solve the transformer problem —innovations that could change LLMs for good, making them faster, far more efficient, and (maybe) even smarter.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in