Dust: Pretraining Transformers Without Backpropagation
Researchers have introduced 'Dust,' a new method for pretraining transformer language models that avoids the traditional backpropagation algorithm. The study suggests that this zeroth-order method is computationally efficient and could potentially outperform backpropagation in compute-rich environments.
Why it matters
If successful, this could fundamentally change how large-scale AI models are trained, potentially bypassing the limitations of current gradient-based optimization.
Samip Dahal, Bishwas Mandal, Serdar Gülbahar, Akshay Vegesna
October 2026 Correspondence to s@qlabs.sh · Code · Cite
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in