I tested 10 model/harness combinations on the same Three.js task
A developer conducted a comparative test of 10 different AI model and harness combinations to determine which produces the best Three.js code for a sci-fi hangar project. The analysis tracks output tokens, reasoning capabilities, and tool error rates.
Why it matters
As AI coding assistants become more prevalent, benchmarking their performance on specific technical tasks is essential for developers to optimize their workflows.
I've been testing a simple prompt with different model and harness combinations to work out which one produces best results. I do this in /goal mode.
Prompt: Build a single-page Three.js sci-fi hangar with hovering drones, animated warning lights, emissive runway strips, and subtle volumetric-style fog planes. Include drone formation toggle and cinematic camera path. Output one self-contained HTML file with inline JavaScript.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in