One broken Kubernetes cluster, two agents: raw kubectl vs Radar's MCP. Which finds the root cause first?
We gave the same model the same fault twice: once with raw kubectl, once reading Radar's live model of the cluster with kubectl blocked. 50 paired fault scenarios on a live EKS cluster. Pick any of them below and watch both agents troubleshoot it, call by call.
The gap is widest on the slowest faults: nine in ten Radar diagnoses landed within 94s. For kubectl, the same mark was 632s, while the kubectl agent went hunting.
Replay any of the 50 Kubernetes faults
Nothing here is hand-picked. Every paired scenario from the run is in the dropdown, including the ones Radar got wrong. Each solid mark is one tool call, drawn at the moment it was issued and as wide as it blocked for; the hatched ground between marks is the model thinking. Hover any mark to see the call.
Correct diagnoses within a time budget
By the final 30 min time budget the two arms land close together: 92% vs 90%. But nobody lets an agent grind for half an hour while production is down. Cap it at a budget you would really grant, and the gap opens up, because Radar has already finished.
| Time budget | Radar MCP | kubectl |
|---|---|---|
| 2 min | 86% | 60% |
| 5 min | 90% | 78% |
| 10 min | 92% | 82% |
| 15 min | 92% | 90% |
| 30 min | 92% | 90% |
Not every Kubernetes MCP server would do this
MCP is how an agent reaches a cluster now, and it is the right interface. But the protocol is a connector, and Kubernetes MCP servers differ enormously in what they send back through it. One that proxies kubectl hands the model the same wall of YAML with an extra hop in front of it. We benchmarked four Kubernetes MCP servers side by side: across 25 faults, Radar was 2.3× faster on average than the next-best one. What moved this benchmark is what came back: a model of the cluster Radar keeps current, so the agent asks one question instead of rebuilding the resource graph from a dozen serial commands. Time spent waiting on the cluster fell from 80s to 8s on average.
13% of the incident blocked on the cluster
42% of the incident blocked on the cluster
With raw kubectl, the agent spends 42% of the incident waiting on the cluster: rollout status, log tails, sleep-and-retry loops, re-listing resources to rebuild context it already lost. Radar answers from a model of the cluster it is already maintaining, so it spends 13% of its time blocked.
That flips diagnosis from I/O-bound to model-bound. Radar's floor is how fast the model thinks; kubectl's floor is how fast a cluster answers serial commands. That is why we expect faster models to widen this gap rather than close it - a prediction, not something this benchmark measured.
Which is why “does a Kubernetes MCP server help?” is the wrong question to ask of the category. The useful question is what a given server sends back: raw resources the model still has to assemble, or state that has already been correlated. The five-surface Kubernetes MCP benchmark measures that category question on a separate 25-fault sample; this page is the deeper 50-fault, two-arm evidence for one server, measured rather than asserted.
Every scenario, including the ones kubectl won
One row per paired fault, sorted by how long kubectl took. The question isn't whether Radar was faster; it's how much, and where kubectl blew past ten minutes.
Where Radar's MCP server loses to kubectl
Radar fails four of the 50 faults, and every one is in the dropdown. Only one is a fault kubectl got right and Radar didn't; the other three beat both arms.
Right component, wrong mechanism. On a Valkey memory disruption Radar found the failing cache immediately and then described it as a statically undersized memory limit rather than an injected disruption, pulling a second workload in as a co-suspect. It localized correctly and lost on characterization - the one fault here where kubectl did better.
A policy object rewrote the workload. In two faults a cluster-scoped object changed the pod at admission - a mutating webhook forcing a 16Mi memory limit, a PriorityClass mis-scoped as a global default. Both agents read the mutated pod accurately and neither reliably named the object that wrote it, because nothing in the pod points back to it. Radar now compares a live pod against its ReplicaSet's template and names the admission candidates that could explain the difference, which moved it from never finding the webhook to finding it on some runs.
A convincing decoy. On a DNS-resolution fault both agents found a real, unrelated misconfiguration elsewhere in the application, explained it carefully, and submitted that instead.
How we ran the Kubernetes agent benchmark
- Benchmark: SREGym (Microsoft + UIUC), faults injected into OpenTelemetry Astronomy Shop, DeathStarBench Hotel Reservation, and Social Network.
- Cluster: live 3-node EKS, us-east-1. Same cluster, same faults for both arms.
- Agent: Claude Code on claude-sonnet-5, both arms.
- Arms are disjoint: the Radar arm cannot fall back to a shell - kubectl is blocked outright. That models how Radar actually deploys, and it attributes results to the tool surface rather than to whichever tool the agent felt like reaching for.
- Grading: SREGym's own LLM judge at temperature 0. The graded artifact is the first diagnosis each agent submits.
- Timing: agent start to submitted diagnosis, read from the raw transcripts. Not session wall-clock, which also counts the separate remediation stage the benchmark runs afterwards.
AI agents on Kubernetes: common questions
Should I give an AI agent kubectl access to my Kubernetes cluster?
Does a Kubernetes MCP server actually make an AI agent better, or is it hype?
How much faster is an AI agent with an MCP server than with kubectl?
How does Radar compare with AI SRE tools?
Is a Kubernetes MCP server that wraps kubectl any better than kubectl?
Where does Radar's MCP server lose to raw kubectl?
How was this Kubernetes MCP benchmark run?
More on AI agents and Kubernetes
- All Kubernetes AI agent benchmarks - the three studies side by side, with the shared setup and how to compare them.
- AI SRE benchmark - Radar against kubectl-ai, HolmesGPT, OpenSRE and AWS DevOps Agent on the same 50 faults, with cost per fault.
- Kubernetes MCP server benchmark - Radar, two general Kubernetes MCP servers, k8sgpt, and raw kubectl on a separate 25-fault sample.
- Best Kubernetes MCP servers, ranked - the broader product decision across speed, safety, writes, coverage, setup, multi-cluster reach, and governance.
- The full write-up - methodology, per-scenario analysis, and what we changed in Radar because of it.
- Radar's MCP server - the 23 read tools and 7 annotated write tools the Radar arm was driving, and how to point your own agent at them.
- AI investigations - the same correlation engine the benchmark measured, driven from Radar's own UI.
Run Radar on your own cluster
Radar is open source, one binary. Run it on your laptop or in your cluster.
$curl -fsSL https://get.radarhq.io | sh && kubectl radarApache 2.0 · No account for local use · Run Radar OSS forever