Benchmark · SREGym · 50 live faults

One broken Kubernetes cluster, two agents: raw kubectl vs Radar's MCP. Which finds the root cause first?

We gave the same model the same fault twice: once with raw kubectl, once reading Radar's live model of the cluster with kubectl blocked. 50 paired fault scenarios on a live EKS cluster. Pick any of them below and watch both agents troubleshoot it, call by call.

3× faster
to a diagnosis, on average across all 50 faults
64s through Radar vs 191s on kubectl
Slightly higher
accuracy
mean judge score 0.936 vs 0.905
86% vs 60%
right within two minutes
faults diagnosed correctly inside an incident's first two minutes

The gap is widest on the slowest faults: nine in ten Radar diagnoses landed within 94s. For kubectl, the same mark was 632s, while the kubectl agent went hunting.

Replay any of the 50 Kubernetes faults

Nothing here is hand-picked. Every paired scenario from the run is in the dropdown, including the ones Radar got wrong. Each solid mark is one tool call, drawn at the moment it was issued and as wide as it blocked for; the hatched ground between marks is the model thinking. Hover any mark to see the call.

Correct diagnoses within a time budget

By the final 30 min time budget the two arms land close together: 92% vs 90%. But nobody lets an agent grind for half an hour while production is down. Cap it at a budget you would really grant, and the gap opens up, because Radar has already finished.

Time budgetRadar MCPkubectl
2 min86%60%
5 min90%78%
10 min92%82%
15 min92%90%
30 min92%90%

Not every Kubernetes MCP server would do this

MCP is how an agent reaches a cluster now, and it is the right interface. But the protocol is a connector, and Kubernetes MCP servers differ enormously in what they send back through it. One that proxies kubectl hands the model the same wall of YAML with an extra hop in front of it. We benchmarked four Kubernetes MCP servers side by side: across 25 faults, Radar was 2.3× faster on average than the next-best one. What moved this benchmark is what came back: a model of the cluster Radar keeps current, so the agent asks one question instead of rebuilding the resource graph from a dozen serial commands. Time spent waiting on the cluster fell from 80s to 8s on average.

Radar MCP64s average to diagnosis
55s
8s blocked · 55s model

13% of the incident blocked on the cluster

kubectl191s average to diagnosis
80s
111s
80s blocked · 111s model

42% of the incident blocked on the cluster

blocked on a toolmodel workingAverages across 50 paired scenarios, both bars on one scale.

With raw kubectl, the agent spends 42% of the incident waiting on the cluster: rollout status, log tails, sleep-and-retry loops, re-listing resources to rebuild context it already lost. Radar answers from a model of the cluster it is already maintaining, so it spends 13% of its time blocked.

That flips diagnosis from I/O-bound to model-bound. Radar's floor is how fast the model thinks; kubectl's floor is how fast a cluster answers serial commands. That is why we expect faster models to widen this gap rather than close it - a prediction, not something this benchmark measured.

Which is why “does a Kubernetes MCP server help?” is the wrong question to ask of the category. The useful question is what a given server sends back: raw resources the model still has to assemble, or state that has already been correlated. The five-surface Kubernetes MCP benchmark measures that category question on a separate 25-fault sample; this page is the deeper 50-fault, two-arm evidence for one server, measured rather than asserted.

Every scenario, including the ones kubectl won

One row per paired fault, sorted by how long kubectl took. The question isn't whether Radar was faster; it's how much, and where kubectl blew past ten minutes.

Where Radar's MCP server loses to kubectl

Radar fails four of the 50 faults, and every one is in the dropdown. Only one is a fault kubectl got right and Radar didn't; the other three beat both arms.

Right component, wrong mechanism. On a Valkey memory disruption Radar found the failing cache immediately and then described it as a statically undersized memory limit rather than an injected disruption, pulling a second workload in as a co-suspect. It localized correctly and lost on characterization - the one fault here where kubectl did better.

A policy object rewrote the workload. In two faults a cluster-scoped object changed the pod at admission - a mutating webhook forcing a 16Mi memory limit, a PriorityClass mis-scoped as a global default. Both agents read the mutated pod accurately and neither reliably named the object that wrote it, because nothing in the pod points back to it. Radar now compares a live pod against its ReplicaSet's template and names the admission candidates that could explain the difference, which moved it from never finding the webhook to finding it on some runs.

A convincing decoy. On a DNS-resolution fault both agents found a real, unrelated misconfiguration elsewhere in the application, explained it carefully, and submitted that instead.

How we ran the Kubernetes agent benchmark

  • Benchmark: SREGym (Microsoft + UIUC), faults injected into OpenTelemetry Astronomy Shop, DeathStarBench Hotel Reservation, and Social Network.
  • Cluster: live 3-node EKS, us-east-1. Same cluster, same faults for both arms.
  • Agent: Claude Code on claude-sonnet-5, both arms.
  • Arms are disjoint: the Radar arm cannot fall back to a shell - kubectl is blocked outright. That models how Radar actually deploys, and it attributes results to the tool surface rather than to whichever tool the agent felt like reaching for.
  • Grading: SREGym's own LLM judge at temperature 0. The graded artifact is the first diagnosis each agent submits.
  • Timing: agent start to submitted diagnosis, read from the raw transcripts. Not session wall-clock, which also counts the separate remediation stage the benchmark runs afterwards.
FAQ

AI agents on Kubernetes: common questions

Should I give an AI agent kubectl access to my Kubernetes cluster?
Not unrestricted access. On 50 live faults the same model averaged 191 seconds to a diagnosis with raw kubectl against 64 through Radar's MCP server, and a shell hands the agent everything its identity can do - one fault here prints a production password into the context window. If an agent needs cluster access, give it a least-privilege, RBAC-scoped surface rather than a shell.
Does a Kubernetes MCP server actually make an AI agent better, or is it hype?
MCP itself is plumbing: a server that just proxies kubectl returns the same YAML with an extra hop. What moved this benchmark was what came back - a correlated model of the cluster, so the agent spent 13% of the incident waiting on the cluster instead of 42%. Accuracy was slightly higher; time to an answer was 3x shorter.
How much faster is an AI agent with an MCP server than with kubectl?
3x on average - 64 seconds against 191 across all 50 faults - and 2x at the median. The gap is widest in the slow tail, where kubectl's 90th percentile was 632 seconds against Radar's 94, and it holds at 2.7x on the 44 faults both arms got right.
How does Radar compare with AI SRE tools?
This study compares one agent with and without Radar. For Radar against HolmesGPT, kubectl-ai, OpenSRE and AWS DevOps Agent on the same 50 faults, see the AI SRE benchmark.
Is a Kubernetes MCP server that wraps kubectl any better than kubectl?
We did not measure that head to head, but our studies suggest the gain is marginal at best. MCP is a simple, generic interface for agents, and agents already know kubectl well from training, so a wrapper adds little. The gain we measured comes from answering out of a continuously maintained model of the cluster, with tools that do more of the work before the agent has to reason.
Where does Radar's MCP server lose to raw kubectl?
It fails 4 of the 50 faults. Only one is a fault kubectl got right: a Valkey memory disruption Radar localized correctly and then mis-described. The rest beat both arms - two where a cluster-scoped policy object rewrote the pod at admission, one where both chased a convincing decoy - and all of them are in the replay.
How was this Kubernetes MCP benchmark run?
SREGym, the fault-injection benchmark from Microsoft and UIUC, on a live 3-node EKS cluster. Both arms ran claude-sonnet-5 on Amazon Bedrock against the same faults, with kubectl blocked outright in the MCP arm so the result reflects the tool surface. SREGym's own LLM judge graded the first diagnosis each agent submitted, and per-trial data is downloadable.

More on AI agents and Kubernetes

  • All Kubernetes AI agent benchmarks - the three studies side by side, with the shared setup and how to compare them.
  • AI SRE benchmark - Radar against kubectl-ai, HolmesGPT, OpenSRE and AWS DevOps Agent on the same 50 faults, with cost per fault.
  • Kubernetes MCP server benchmark - Radar, two general Kubernetes MCP servers, k8sgpt, and raw kubectl on a separate 25-fault sample.
  • Best Kubernetes MCP servers, ranked - the broader product decision across speed, safety, writes, coverage, setup, multi-cluster reach, and governance.
  • The full write-up - methodology, per-scenario analysis, and what we changed in Radar because of it.
  • Radar's MCP server - the 23 read tools and 7 annotated write tools the Radar arm was driving, and how to point your own agent at them.
  • AI investigations - the same correlation engine the benchmark measured, driven from Radar's own UI.

Run Radar on your own cluster

Radar is open source, one binary. Run it on your laptop or in your cluster.

Quick install
$curl -fsSL https://get.radarhq.io | sh && kubectl radar

Apache 2.0 · No account for local use · Run Radar OSS forever