Jev-Driven SRE Diagnosis: What Worked and What Failed
Researchers have developed a Jev-driven SRE diagnosis pipeline that automates incident investigation by analyzing cluster evidence without the need for an LLM agent. The system successfully diagnosed 76% of tested faults by programmatically collecting data and using Jev to make classification decisions.
Why it matters
It highlights the potential for specialized decision models to outperform or replace LLMs in specific, high-stakes operational tasks like SRE diagnosis.
In our first study , we experimented with Jev as a decision aid for an LLM agent. The agent diagnosed and repaired incidents; Jev helped rank the agent's proposed tests and reviewed the evidence before submission.
That post ended with a more ambitious idea: giving Jev a broad view of the cluster and letting its fast, cheap judgments guide the investigation.
In this post, we present a Jev-driven diagnosis pipeline without any LLM agent . The pipeline programmatically collects and organizes cluster evidence, then feeds it to Jev. Jev selects a likely root cause and supporting observations, and the pipeline uses them to assemble a diagnosis report.
Across 21 SREGym-Lite faults, the Jev-driven pipeline passes 80 of 105 diagnoses (76.2%) , with a median diagnosis time of 14.6 seconds .
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in