Are AI Labs Pelicanmaxxing?

An analysis of AI model performance using the 'pelican riding a bicycle' SVG benchmark suggests that AI labs may be optimizing models specifically to pass informal community tests. The study tested seven frontier models to determine if they were 'benchmaxxing' on this specific prompt.
Why it matters
This highlights the growing tension between standardized AI benchmarks and the informal, often meme-driven evaluation methods used by the developer community.
For the past few years, Simon Willison has tested every major LLM release with the same prompt: “Generate an SVG of a pelican riding a bicycle”.
What began as a tongue-in-cheek benchmark has become one of the most famous informal benchmarks in AI. Simon’s pelican-on-a-bicycle results are often among the most upvoted comments on Hacker News threads announcing new releases from AI labs.
The benchmark is now famous enough that there’s plenty of discussion about its usefulness and about whether AI labs might be benchmaxxing 1 on it. When billions or even trillions of dollars are at stake, and a strong result could help persuade users, wouldn’t it be tempting to pelicanmaxx your model just a bit?
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in