Goodfire says its new ‘inside-out’ monitors catch rogue AI agents at a fraction of the cost

Startup Goodfire has launched a new monitoring system for AI agents that analyzes internal model activations rather than just output text. This 'inside-out' approach is designed to be more cost-effective and efficient at detecting rogue behavior compared to traditional secondary AI monitors.
Why it matters
As AI agents become more autonomous, cost-effective safety and interpretability tools are essential for preventing security breaches and unintended model behavior.
The standard way to keep an AI agent in line is to have a second AI read over its shoulder . It’s been the default approach, but it can get expensive fast when agents run for hours and process the equivalent of several novels’ worth of text.
Goodfire, a startup focused on interpretability (figuring out how AI models work internally), launched a cheaper option on Thursday: monitors that watch what’s happening inside an AI model as it works, rather than just reading what it writes. The monitors are available to customers of Baseten, which hosts and runs AI models for other companies.
Baseten’s Base Labs announced a safety partnership with Goodfire and the AI platform Hugging Face last month.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in