Livenerf: Has Opus 5.5 been nerfed yet?
The 'livenerf' project is a new, deterministic benchmark designed to track whether AI models like Claude Opus 5.5 degrade in performance over time after their initial release. By using a standardized, automated testing protocol, it aims to move beyond anecdotal 'vibes' to provide statistical evidence of model drift.
Why it matters
As AI models become integral to infrastructure, verifying that their performance remains consistent post-launch is critical for reliability and trust.
A long-running, deterministic-as-possible benchmark for detecting whether a frontier model gets quietly worse after launch.
📋 The plan · 📊 Results · 🔬 How it works · 🧪 Pre-registration
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in