Boris Cherny on Trying to Get Claude Code to Rewrite the Claude App

Boris Cherny, head of Claude Code at Anthropic, discusses the shift in AI development from prompt engineering to task verification. He shares an experiment where he tasked an AI agent with rewriting a desktop application by verifying its own code output against visual benchmarks.
Why it matters
This demonstrates the evolving capability of AI agents to perform complex, multi-step software engineering tasks with minimal human intervention.
Stop searching the App Store. Start telling your phone what to build.
Boris Cherny, head of Claude Code at Anthropic, was interviewed by Diana Hu on stage at Y Combinator’s Startup School 2026 last week. Starting around 20:30 in the video , he briefly discussed directing Claude to perform difficult tasks:
Cherny: I think the skill nowadays is less about prompt engineering and more about figuring out how do you give Claude a hard task that seems a little bit too hard. Then how do you make it possible for Claude to verify its work along the way? The verification is probably the single most important thing that people do not get right, largely.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in