Show HN: AI SRE Arena, an Open Benchmark for AI SRE Agents on Kubernetes
The AI SRE Arena is a new open-source benchmark designed to evaluate AI agents in Kubernetes environments. It allows users to deploy clusters, inject faults, and score the performance of various AI models against standardized incident scenarios.
Why it matters
As AI agents are increasingly deployed for IT operations and site reliability engineering, standardized benchmarks are essential for measuring their effectiveness and reliability.
A vendor-neutral starter for deploying a disposable Kubernetes fixture, injecting faults, saving investigation records from any product, and scoring completed investigations with a configurable judge. Python 3.10+ is the only Python dependency. Kubernetes operations additionally need kubectl ; local cluster creation needs Docker and kind .
Choose local kind or AWS EKS for your cluster, then select the 21-scenario full suite or six-scenario smoke fixture. Cluster type and scenario suite are separate choices.
flowchart LR A[Deploy cluster and application] --> B[Connect your product] B --> C[Inject and verify a fault] C --> D[Let your product investigate] D --> E[Save investigation records] E --> F[Score with your chosen judge] F --> G[Compare results] Loading Run commands from the Project Arena directory. Cluster setup, product integration, scenario execution, and scoring are separate steps; follow the sections below in order.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in