Show HN: Cactus Hybrid: We taught Gemma 4 to know when it's wrong
Cactus has released a new hybrid AI model, Gemma 4 E2B, which uses internal confidence scoring to decide whether to answer a query locally or route it to a larger, more powerful model. This approach aims to balance the speed and privacy of on-device processing with the accuracy of larger cloud-based systems.
Why it matters
This represents a significant shift in AI architecture, moving toward cost-effective, efficient, and privacy-conscious hybrid models that optimize computational resources.
A small, on-device model is fast and private, but sometimes wrong. At Cactus we post-train models to know when they are wrong : we ship probes inside the checkpoint that score every answer with a confidence between 0 and 1, returned as structured data (never parsed out of the answer text). Answer on-device when confidence is high; you can re-route to a bigger model when it's low:
if confidence < 0.85 : answer = ask_a_bigger_model ( prompt ) We start the rollout with Gemma 4 E2B Hybrid , all builds live in the Cactus Hybrid collection on Hugging Face.
Gemma 4 E2B hybrid, the smallest Gemma model, matches Gemini 3.1 Flash-Lite on most benchmarks by routing only 15–35% of queries to the Gemini 3.1 Flash-Lite and running the remnant itself.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in