LeMario: Training a JEPA World Model on Super Mario Bros

A developer details their attempt to train a Joint-Embedding Predictive Architecture (JEPA) model on Super Mario Bros to learn world dynamics. While the model successfully predicted short-term game frames, it failed to learn long-term planning or goal-oriented navigation, serving as a postmortem on AI world modeling.
Why it matters
This highlights the current limitations of predictive AI models in mastering complex, long-horizon tasks despite their ability to simulate short-term dynamics.
I wanted to reproduce LeWorldModel , a small Joint-Embedding Predictive Architecture (JEPA) that learns world dynamics from pixels and actions. The original paper used it for reward-free planning in Push-T. But, since I loved video games, and at the same time wanted to learn more deeply about LeCun's JEPA architecture, I decided to write the whole architecture from scratch and train it on Super Mario Bros.
The model passed every test I initially thought mattered. It generalized to held-out episodes, used the actions, and predicted five-step futures better than strong baselines. Raw reward-free planning could move Mario toward nearby image goals and finish within two and five pixels of the targets. :D
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in