Show HN: Jevstiller – Distill Jev into a local model, with a disagreement bound

Jevstiller is a tool designed to optimize AI agent loops by distilling large model responses into smaller, faster local models with a defined disagreement bound. It allows developers to maintain high accuracy while significantly reducing latency and costs.
Why it matters
As AI agents become more prevalent, reducing the latency of network calls is critical for real-time performance and operational efficiency.
September 2026. Every number here is from the benchmarks , and bash experiments/bench.sh --no-record reruns them without an API key.
If you classify text with Jev , every answer is a network call to one vendor and comes back in about 300 ms, at any load. For a batch job that is fine. For an agent loop that decides, acts, and decides again, or a game tick, or anything that classifies then acts, 300 ms per step is the whole budget.
Jevstiller sits in front of that call, learns a small local model from Jev’s own answers, and lets it answer what it is sure about in about 15 ms on a CPU. The interesting part is not the small model. It is the contract:
Set one number, say 98%. Jevstiller returns the label Jev would have returned on at least that share of requests.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in