Which AI SRE tool diagnoses Kubernetes faults correctly?
Five tools, the same 50 broken clusters, the same question and judge. On faults that announce themselves they score within a few of each other. On the 14 faults where every pod reports Ready, Claude Code + Radar MCP got 12 and AWS DevOps Agent got 5. Claude Code + Radar MCP answered in an average of 79s* at $0.27 a fault. We build Radar, so every number here comes from the dataset below, which you can download.
* Newer Radar versions improved this: Radar 1.15 averaged 64s on the same 50 faults.
- RadarClaude Code + Radar MCP44/50
- kubectl-ai41/50
- HolmesGPT38/50
- OpenSRE38/50
- AWS DevOps AgentAWS-managed model33/50 · 5 refused
Claude Code + Radar MCP was the most accurate of the five: clearly ahead of HolmesGPT, OpenSRE and AWS DevOps Agent, and slightly ahead of kubectl-ai. Each fault ran once, so a gap of a fault or two can move between runs. The time and cost gaps are much larger.
- RadarClaude Code + Radar MCP79s avg55s median
- kubectl-ai288s avg184s median
- HolmesGPT176s avg116s median
- OpenSRE151s avg137s median
- AWS DevOps AgentAWS-managed model123s avg99s median
- RadarClaude Code + Radar MCP$0.27 avg$0.23 median
- kubectl-ai$1.67 avg$0.63 median
- HolmesGPT$0.37 avg$0.26 median
- OpenSRE$2.07 avg$1.21 median
- AWS DevOps Agentagent-time, not tokens$1.27 avg$1.02 median
How did each AI SRE tool score on accuracy, time and cost?
| Product | Correct | Average answer | Median answer | Cost per fault | Model |
|---|---|---|---|---|---|
| Claude Code + Radar MCP 1.13.1 | 44 / 50 | 79 s | 55 s | $0.27 | claude-sonnet-5 on Amazon Bedrock |
| kubectl-ai 0.0.31 | 41 / 50 | 288 s | 184 s | $1.67 no prompt caching, upper bound | claude-sonnet-5 on Amazon Bedrock |
| HolmesGPT 0.41.0 | 38 / 50 | 176 s | 116 s | $0.37 | claude-sonnet-5 on Amazon Bedrock |
| OpenSRE 0.1.2026.9.12 | 38 / 50 | 151 s | 137 s | $2.07 | claude-sonnet-5 on Amazon Bedrock |
| AWS DevOps Agent GA (generally available 2026-03-31) | 33 / 50 | 123 s | 99 s | $1.27 agent-time | undisclosed AWS-managed model, default "Balanced" tier - not selectable |
Time is each product’s own answer time, question sent to answer returned, over every trial whether right or wrong - the wait a user actually sees.
All five handle crash loops. Radar pulls ahead on faults where every pod reports Ready
Split the 50 faults by what makes them hard and the ranking above stops being one number. On faults that announce themselves, all five land within a few of each other. On faults where every pod reports Ready, the spread is 7 of 14.
The fault announces itself
28 faultsThe workload is visibly unhealthy before anyone investigates: a crash loop, a Pending pod, a failing probe, an image that will not pull.
spread: 3 of 28
- Radar27/28
- kubectl-ai26/28
- HolmesGPT25/28
- OpenSRE26/28
- AWS DevOps Agent24/28
Ready, but broken
14 faultsEvery pod reports Ready and nothing is restarting, but the application is wrong. The evidence is in a configuration object; pod status looks healthy.
spread: 7 of 14
- Radar12/14
- kubectl-ai9/14
- HolmesGPT7/14
- OpenSRE6/14
- AWS DevOps Agent5/14
Controllers, jobs, webhooks
8 faultsThe fault lives in a controller, Job, webhook, HPA or admission path rather than in the workload's own spec.
spread: 2 of 8
- Radar5/8
- kubectl-ai6/8
- HolmesGPT6/8
- OpenSRE6/8
- AWS DevOps Agent4/8
Our grouping, not SREGym's, and assigned after the runs completed. Every scenario's class ships in this file so the cut can be checked or redrawn.
How we ran the AI SRE benchmark
- Benchmark
- SREGym (Microsoft + UIUC), diagnosis-only faults
- Cluster
- 3-node Amazon EKS, live application, injected fault
- Judge
- claude-opus-5 on SREGym's own rubric, judged on three questions: where the fault is, what broke, and how far it spread
- Pass rule
- SREGym composite ≥ 0.7
- Timing
- product answer time: question sent → final answer returned. A trial still running at 1100s is stopped by the harness and its last output graded; those are marked capped and enter the timing aggregates at the cap, which understates how long that product would have taken.
- Sample
- 1 counted trial per fault, 250 trials in all. A trial was re-run only when a product or our harness failed to run, never for scoring low.
What we held constant, and what we could not
Same harness, same fault injections, same question, same judge and rubric, same live cluster, read-only for everyone. Each product brought its own agent loop, its own instructions and its own way of looking at the cluster, because that is the product.
We changed only what a product needed to run here. kubectl-ai reached Bedrock through a local proxy, because its own Claude providers are broken in that version, and we raised its iteration cap from 20 to 100. HolmesGPT also got a higher step cap. AWS DevOps Agent got AWS's own read-only EKS access policies. Each product's full setup is in the dataset.
Four arms ran Claude Sonnet 5 on Amazon Bedrock for every model role. AWS DevOps Agent is a managed service running an undisclosed model at a default tier that cannot be selected, so its result is a product result and not a model comparison.
AWS DevOps Agent also sees the cluster unfiltered, including the chaos-engineering namespaces the other four arms never see - extra noise that counts against it. And every fault is injected minutes after the application deploys, which flatters a tool that correlates recent change and penalises one that reads logs first.
One trial per fault. We have watched individual faults flip between draws, so a difference of a fault or two can move between runs. That makes the accuracy lead over kubectl-ai a small one; the time and cost gaps are much larger.
Dataset
Every score, answer time, tool call count, cost and outcome for all 250 trials, as JSON. The numbers on this page are derived from this file at build time, so they cannot drift from it.
Download the datasetPublished · Last updated .
AI SRE benchmark questions
Which AI SRE tool is most accurate on Kubernetes faults?
What separates the tools, if they all handle crashing pods?
Did every tool run on the same model?
Why did AWS DevOps Agent refuse to answer on some faults?
How is cost measured when the tools bill differently?
Can I reproduce this benchmark?
Related
Run Radar on your own cluster
Radar is open source, one binary. Run it on your laptop or in your cluster.
$curl -fsSL https://get.radarhq.io | sh && kubectl radarApache 2.0 · No account for local use · Run Radar OSS forever