SREGym · 50 live faults · 5 products

Which AI SRE tool diagnoses Kubernetes faults correctly?

Five tools, the same 50 broken clusters, the same question and judge. On faults that announce themselves they score within a few of each other. On the 14 faults where every pod reports Ready, Claude Code + Radar MCP got 12 and AWS DevOps Agent got 5. Claude Code + Radar MCP answered in an average of 79s* at $0.27 a fault. We build Radar, so every number here comes from the dataset below, which you can download.

* Newer Radar versions improved this: Radar 1.15 averaged 64s on the same 50 faults.

Faults diagnosed correctly, of 50

Claude Code + Radar MCP was the most accurate of the five: clearly ahead of HolmesGPT, OpenSRE and AWS DevOps Agent, and slightly ahead of kubectl-ai. Each fault ran once, so a gap of a fault or two can move between runs. The time and cost gaps are much larger.

Time to an answer - average and median, all 50 faults
Cost per fault - average and median, all 50 faults

How did each AI SRE tool score on accuracy, time and cost?

ProductCorrectAverage answerMedian answerCost per faultModel
Claude Code + Radar MCP
1.13.1
44 / 5079 s55 s$0.27claude-sonnet-5 on Amazon Bedrock
kubectl-ai
0.0.31
41 / 50288 s184 s$1.67
no prompt caching, upper bound
claude-sonnet-5 on Amazon Bedrock
HolmesGPT
0.41.0
38 / 50176 s116 s$0.37claude-sonnet-5 on Amazon Bedrock
OpenSRE
0.1.2026.9.12
38 / 50151 s137 s$2.07claude-sonnet-5 on Amazon Bedrock
AWS DevOps Agent
GA (generally available 2026-03-31)
33 / 50123 s99 s$1.27
agent-time
undisclosed AWS-managed model, default "Balanced" tier - not selectable

Time is each product’s own answer time, question sent to answer returned, over every trial whether right or wrong - the wait a user actually sees.

The finding

All five handle crash loops. Radar pulls ahead on faults where every pod reports Ready

Split the 50 faults by what makes them hard and the ranking above stops being one number. On faults that announce themselves, all five land within a few of each other. On faults where every pod reports Ready, the spread is 7 of 14.

The fault announces itself

28 faults

The workload is visibly unhealthy before anyone investigates: a crash loop, a Pending pod, a failing probe, an image that will not pull.

spread: 3 of 28

  1. Radar27/28
  2. kubectl-ai26/28
  3. HolmesGPT25/28
  4. OpenSRE26/28
  5. AWS DevOps Agent24/28

Ready, but broken

14 faults

Every pod reports Ready and nothing is restarting, but the application is wrong. The evidence is in a configuration object; pod status looks healthy.

spread: 7 of 14

  1. Radar12/14
  2. kubectl-ai9/14
  3. HolmesGPT7/14
  4. OpenSRE6/14
  5. AWS DevOps Agent5/14

Controllers, jobs, webhooks

8 faults

The fault lives in a controller, Job, webhook, HPA or admission path rather than in the workload's own spec.

spread: 2 of 8

  1. Radar5/8
  2. kubectl-ai6/8
  3. HolmesGPT6/8
  4. OpenSRE6/8
  5. AWS DevOps Agent4/8

Our grouping, not SREGym's, and assigned after the runs completed. Every scenario's class ships in this file so the cut can be checked or redrawn.

Methodology

How we ran the AI SRE benchmark

Benchmark
SREGym (Microsoft + UIUC), diagnosis-only faults
Cluster
3-node Amazon EKS, live application, injected fault
Judge
claude-opus-5 on SREGym's own rubric, judged on three questions: where the fault is, what broke, and how far it spread
Pass rule
SREGym composite ≥ 0.7
Timing
product answer time: question sent → final answer returned. A trial still running at 1100s is stopped by the harness and its last output graded; those are marked capped and enter the timing aggregates at the cap, which understates how long that product would have taken.
Sample
1 counted trial per fault, 250 trials in all. A trial was re-run only when a product or our harness failed to run, never for scoring low.

What we held constant, and what we could not

Same harness, same fault injections, same question, same judge and rubric, same live cluster, read-only for everyone. Each product brought its own agent loop, its own instructions and its own way of looking at the cluster, because that is the product.

We changed only what a product needed to run here. kubectl-ai reached Bedrock through a local proxy, because its own Claude providers are broken in that version, and we raised its iteration cap from 20 to 100. HolmesGPT also got a higher step cap. AWS DevOps Agent got AWS's own read-only EKS access policies. Each product's full setup is in the dataset.

Four arms ran Claude Sonnet 5 on Amazon Bedrock for every model role. AWS DevOps Agent is a managed service running an undisclosed model at a default tier that cannot be selected, so its result is a product result and not a model comparison.

AWS DevOps Agent also sees the cluster unfiltered, including the chaos-engineering namespaces the other four arms never see - extra noise that counts against it. And every fault is injected minutes after the application deploys, which flatters a tool that correlates recent change and penalises one that reads logs first.

One trial per fault. We have watched individual faults flip between draws, so a difference of a fault or two can move between runs. That makes the accuracy lead over kubectl-ai a small one; the time and cost gaps are much larger.

Dataset

Every score, answer time, tool call count, cost and outcome for all 250 trials, as JSON. The numbers on this page are derived from this file at build time, so they cannot drift from it.

Download the dataset

Published · Last updated .

FAQ

AI SRE benchmark questions

Which AI SRE tool is most accurate on Kubernetes faults?
Claude Code + Radar MCP was the most accurate of the five: clearly ahead of HolmesGPT, OpenSRE and AWS DevOps Agent, and slightly ahead of kubectl-ai. On this 50-fault set it diagnosed 44 correctly, kubectl-ai 41, HolmesGPT 38, OpenSRE 38 and AWS DevOps Agent 33. Each fault ran once, so a gap of a fault or two can move between runs. The time and cost gaps are much larger.
What separates the tools, if they all handle crashing pods?
Fault class. On the 28 faults that announce themselves - a crash loop, a Pending pod, a failing probe - every tool scored between 24 and 27 of 28. On the 14 faults where every pod reports Ready and the evidence is a configuration object, the range is 5 to 12 of 14. That is where the products differ.
Did every tool run on the same model?
Four of the five did: Claude Sonnet 5 on Amazon Bedrock, for every model role. AWS DevOps Agent is a managed service that runs an undisclosed AWS model at its default tier and offers no way to select one, so that arm's result is a product result and cannot be attributed to the harness alone.
Why did AWS DevOps Agent refuse to answer on some faults?
On 5 of 50 faults its own content safety filter replaced a completed investigation with a refusal to answer. The question text was identical on every trial and the refusals span three applications, so we could not determine what triggers it. They count as misses because a customer would get the same thing. Excluding them, the score is 33 of 45.
How is cost measured when the tools bill differently?
Four tools are billed for LLM tokens, measured from Amazon Bedrock's own CloudWatch metrics over each trial and priced at on-demand list rates. AWS DevOps Agent is billed for agent-seconds, so its per-fault figure is the account's metered total for the run allocated across trials in proportion to answer time. The two are comparable as money spent per fault diagnosed. They are not comparable as token counts, and AWS Support customers receive monthly credits that make list price an upper bound.
Can I reproduce this benchmark?
The full per-fault dataset is downloadable from this page as JSON, including every score, answer time, tool call count, cost and outcome. The benchmark itself is SREGym, an open harness from Microsoft and UIUC. We build Radar, so read the methodology and check the numbers against the data.

Run Radar on your own cluster

Radar is open source, one binary. Run it on your laptop or in your cluster.

Quick install
$curl -fsSL https://get.radarhq.io | sh && kubectl radar

Apache 2.0 · No account for local use · Run Radar OSS forever