Radar MCP vs kubectl: A Kubernetes Agent Benchmark on 50 Faults
The same model diagnosed 50 live Kubernetes faults 3x faster on average through Radar's MCP server than through raw kubectl, with slightly higher accuracy. What we measured, why it happens, and where Radar still loses.

Updated September 27, 2026. This post now reports the third run of this benchmark: 50 faults, both arms on
claude-sonnet-5through Amazon Bedrock. The first version (July) usedclaude-sonnet-4-6and timed whole sessions rather than the diagnosis. The August run, on 54 faults, measured a larger gap, about four times at the median; the shell arm was unusually slow in it, while Radar's times reproduced within seconds across every run. Every number below is from the current run.
I wanted to know one thing: when an AI agent debugs a broken Kubernetes cluster, does it matter what tools we hand it?
Most vendors hand-wave this - "our MCP makes your agent smarter." That is not the bar. If you're giving Claude, Cursor, or an in-house agent access to Kubernetes, the tool surface has to beat raw kubectl on the things you feel during an incident: how long you wait, whether the answer is right, and whether you can let it near production. So I ran the same Kubernetes benchmark with the same model on two tool surfaces. One arm drove raw kubectl through a shell. The other drove Radar's Kubernetes MCP server.
TL;DR
On 50 real Kubernetes fault-injection scenarios, an agent driving Radar's Kubernetes MCP server reached a diagnosis in 64 seconds on average against 191 seconds for the same agent on raw kubectl - three times faster. Accuracy was slightly higher (a mean judge score of 0.936 vs 0.905). The bigger difference is how long you wait for it: within two minutes, Radar had the right answer on 86% of faults against 60%.
Worth being precise about what won, because it isn't MCP. MCP is a connector; an MCP server that just proxies kubectl returns the same wall of YAML with an extra hop in front of it. We benchmarked four Kubernetes MCP servers side by side, and Radar was more than twice as fast on average as the next-best one. What won is what came back through the connector: a model of the cluster Radar keeps current. The kubectl agent spends 42% of the incident blocked, waiting on the cluster to answer. Through Radar it's 13%.
What I tested
The setup was deliberately simple: same benchmark suite, same model, same goal, two tool surfaces.
- Scenarios: SREGym (Microsoft + UIUC) - 50 paired fault-injection scenarios on a live cluster. Bad DNS policies, wrong selectors, missing env vars, broken probes, auth disruptions, stale secrets, admission webhooks, slow image pulls. Real faults, not puzzles.
- Cluster: EKS, 3 nodes, us-east-1.
- Model:
claude-sonnet-5on Amazon Bedrock, both arms. - Tool surfaces: Arm A got raw
kubectlvia Bash, includingexec. Arm B got Radar's MCP tools withkubectlblocked outright, so it could not fall back to a shell. - Grading: SREGym's own LLM judge on the first diagnosis each agent submits. Time is measured from agent start to that submission.
The numbers
| Metric | kubectl | Radar MCP |
|---|---|---|
| Average time to a diagnosis (all 50 faults) | 191s | 64s |
| Median time to a diagnosis | 87s | 44s |
| Mean judge score (0-1) | 0.905 | 0.936 |
| Answered sooner, of the 44 both got right | 7 | 37 |
| Slowest correct diagnosis | 757s (13 min) | 358s (6 min) |
| Share of the incident blocked on the cluster | 42% | 13% |
I lead with the average rather than the median on purpose. The median hides the tail, and the tail is the story: the kubectl agent went past ten minutes on six faults, and got two of those wrong anyway. Radar never went past seven minutes.
Correct diagnoses within a time budget
Reporting accuracy and speed separately understates what is happening, because nobody lets an agent grind for thirty minutes while production is down. Put the two together - how many faults each arm has correctly diagnosed inside a given time budget - and the tie on accuracy turns into a clear gap:
| Agent time budget | Radar MCP | kubectl |
|---|---|---|
| 2 minutes | 86% | 60% |
| 5 minutes | 90% | 78% |
| 10 minutes | 92% | 82% |
| 30 minutes | 92% | 90% |
Radar has nearly every correct answer it will ever produce inside two minutes. kubectl gets to the same place eventually, if you give it a quarter of an hour.
Why this Kubernetes MCP server beats raw kubectl - and why not all of them would
This is the part that explains the numbers instead of just reporting them.
Radar's MCP is not a kubectl wrapper - that distinction is the whole mechanism. Give an agent raw kubectl and a single kubectl get pod -o yaml hands back hundreds of lines: managedFields, full status blocks, annotations, the entire object as the API server stores it. The one useful signal - an OOMKilled condition, a failing readiness probe - is buried in the wall of text. And the agent re-reads its accumulated context every turn, so by call 30 it's dragging a huge transcript of YAML it already parsed.
To trace a cross-cutting fault - "this Service can't reach its pods" - a kubectl agent has to reconstruct the graph by hand: list the Service, read its selector, list pods, check labels, read endpoints, check the network policy, pull events, cross-reference timestamps. Worse, a lot of that is waiting. Rollout status, log tails, sleep-and-retry loops, and throwaway pods launched just to run curl from inside the cluster. That is where the 42% comes from.
Radar hands over the already-computed answer instead. The MCP returns minified, enriched, secret-redacted data: pre-built topology graphs, problem-correlated timelines, deduplicated events, error-filtered logs. It calls the issues API and gets the causal chain - what broke, what it's connected to, what changed - instead of the raw materials to derive it. Most of those calls come back in under a second, because the expensive work happened continuously, before the incident.
That flips diagnosis from I/O-bound to model-bound. Radar's floor is how fast the model thinks; kubectl's floor is how fast a cluster answers serial commands.
This is exactly the argument Roy made in "Agents, UIs, and CLIs: The False Choice in Kubernetes Operations" - a cluster is a live graph, and raw YAML is a poor way for anything to make sense of it, human or LLM. This benchmark is that argument with a number attached.
Where Radar pulls ahead, and where it loses
Radar's biggest wins are the faults where nothing crashes. A rotated Secret the pod never picked up, a missing environment variable, a changed Valkey password. Every pod reports Ready, describe looks clean, and there is no event to find. That is where most of the kubectl agent's ten-minute hunts happened. "What changed, and does it correlate with the symptom" is precisely what Radar computes continuously.
Radar missed four of the 50, and they are all in the replay if you want to watch them go wrong. Only one is a fault kubectl got right: on a Valkey memory disruption, Radar found the failing cache immediately and then described it as an undersized memory limit rather than an injected disruption. Right component, wrong mechanism.
The other three beat both arms. In two, a cluster-scoped object rewrote the pod at admission - a mutating webhook forcing a 16Mi memory limit, a PriorityClass mis-scoped as the global default. Both agents read the mutated pod accurately and neither reliably named the object that wrote it, because nothing in the pod points back to it. Radar now compares a live pod against its ReplicaSet's template and names the admission candidates that could explain the difference, which is a gap this benchmark found. In the third, both agents found a real but unrelated misconfiguration elsewhere in the app and submitted that instead of the DNS fault.
You don't want an agent loose on raw kubectl
Speed and accuracy aside, raw kubectl through a shell is the ungoverned option. The agent inherits whatever that shell can do, every call is a potential mutation, and you get no record of what it touched. This benchmark contains a fault whose most natural diagnostic command prints a production password in plaintext, straight into an LLM context window.
Radar's MCP is the opposite by design: read tools are read-only and secret-redacted, writes are RBAC-enforced and gated so the client can prompt before anything changes, and Radar Cloud adds the governance layer on top - an inventory of every agent and token with cluster access, the same Kubernetes impersonation a human gets, and audited tool activity for agent actions.
"Faster, and slightly more accurate" is the benchmark result. "And you can actually let it near production" is why teams route the agent through Radar in the first place.
Why this is the argument for an agent surface
Step back and the result isn't really about AI - it's about data shape. Agents reason over Kubernetes far faster when the data is structured for them. It's the same reason we build dashboards instead of reading raw YAML: eyes need the graph rendered, and it turns out agents need the same graph modeled.
So Radar is one engine that watches your clusters, with two surfaces over it. Humans consume it as a UI - still the surface most of our users live in. Agents consume it as MCP. Both need the same thing: the cluster made sense of, not just dumped. This benchmark was about the second surface, but the human-first one is still the point - we build Radar for the person staring at a cluster at 2am, and increasingly some of that work is done by an agent on their behalf.
Watch it, then run it yourself
You don't have to take my word for any of this. Every one of the 50 scenarios is replayable side by side, call by call, including the ones Radar got wrong:
Radar is open source with the MCP server on by default, so you can point your own agent at it and break something on purpose to see how it reasons.
curl -fsSL https://get.radarhq.io | sh && kubectl radarThat gets you the UI and the MCP server in one go. Aim your agent at the MCP endpoint and watch it work a fault on a cluster you trust.
This study asks whether Radar's correlated MCP surface beats raw kubectl. The newer five-surface Kubernetes MCP benchmark compares Radar with other MCP servers on 25 live faults, the AI SRE benchmark puts Radar against kubectl-ai, HolmesGPT, OpenSRE, AWS DevOps Agent and kagent on the same 50 faults, and the Kubernetes MCP server comparison covers operations, safety, and fleet governance. All three studies are on the benchmarks page.
Get the next issue in your inbox
Kubernetes deep-dives, new Radar releases, and what we learned shipping them.
About one email a month. No sales sequences. Unsubscribe anytime.
Keep reading
Give Your Agent Context Before You Buy an AI SRE
We ran the same agent and model on 50 live Kubernetes faults. With read-only cluster context it reached the answer about three times faster, and was slightly more accurate. Here's what it did without that context, and how to run the same test before you buy an AI SRE.
Read-Only Is Not a Safety Boundary
Two sentences get said in every meeting where a team wires an AI agent into a cluster: it's read-only, and just use the service account. Both are wrong later.
Radar on Cloud Native FM: Kubernetes as a Graph, and What Agents Change
Nadav Erell took Radar through a live broken cluster on Cloud Native FM. The short version: why a cluster is a graph, and what MCP actually changes.