Live-cluster benchmark · 25 injected faults

Which Kubernetes MCP server is fastest?

The same agent diagnosed the same 25 live faults five times: through Radar, two general Kubernetes MCP servers, k8sgpt, and raw kubectl. Radar took 57 seconds on average to submit a diagnosis, 2.3× faster than the next-fastest MCP server and 4.3× faster than raw kubectl. At the median it was 38s against 62s (1.6×).

57s
average to a diagnosis
all 25 faults
2.3×
faster on average
than the next-fastest MCP server
8
median tool calls
vs 14–18 for the other MCP arms

Average time to first submitted diagnosis

Bars include all 25 faults. Each line below shows the stricter check on the same 16 faults every arm diagnosed correctly. Lower is faster.

agent start → first submitted diagnosis

  1. 01Radarfastest

    Correlated cluster model · median 38s · 8 median calls

    57s
    same 16 correct faults47s average
  2. 02Flux159/mcp-server-kubernetes

    kubectl/Helm command tools · median 62s · 14 median calls

    135s
    same 16 correct faults93s average
  3. 03containers/kubernetes-mcp-server

    Native Kubernetes API client · median 64s · 14 median calls

    149s
    same 16 correct faults103s average
  4. 04k8sgpt

    Analyzer plus resource reads · median 119s · 18 median calls

    174s
    same 16 correct faults115s average
  5. 05kubectl

    Shell and kubectl · median 72s · 8 median calls

    249s
    same 16 correct faults128s average

Safety, write support, coverage, setup and fleet support are compared on the Kubernetes MCP server comparison.

Per fault

All 25 faults, all five arms

Pick a fault to see each arm's time to a diagnosis and its judge score.

Pick an injected fault

The same agent saw the same fault through each tool surface.

Astronomy Shop

Readiness probe misconfiguration

bar = timescore = judge
  1. Radarcorrect
    34.2s · score 1.00
  2. containers/kubernetes-mcp-servercorrect
    64.3s · score 1.00
  3. Flux159/mcp-server-kubernetescorrect
    91.7s · score 1.00
  4. kubectlcorrect
    328.3s · score 1.00
  5. k8sgptcorrect
    82.3s · score 1.00

A correct badge is the judge's verdict; the number keeps partial credit.

What we measured

One agent, five ways to reach the cluster. Everything else stayed fixed.

The faults
25 fault scenarios from SREGym (Microsoft and UIUC), injected into real applications on a 3-node EKS, us-east-1 cluster.
The agent
claude-sonnet-5 with the same task on every arm. Each arm reached the cluster only through its own tools; the kubectl arm had a shell.
The clock
From agent start to its first submitted diagnosis.
The grade
SREGym's judge (claude-opus-5) on that first submission.
Arm
Radar MCP1.9.2Built-in MCP server; diagnosis task only
containers/kubernetes-mcp-serverv0.0.66--disable-destructive
Flux159/mcp-server-kubernetes4.1.2ALLOW_ONLY_NON_DESTRUCTIVE_TOOLS=true
Raw kubectlEKS cluster clientDirect shell access
k8sgpt MCP0.4.36k8sgpt serve --mcp --mcp-http
Accuracy. Radar had the highest mean judge score (0.94). Radar and the two general MCP servers got 22–23 of 25 right; raw kubectl got 20 and k8sgpt 19. The large gap is speed, and Radar keeps it on the 16 faults every arm got right.

Download the dataset

25 scenarios, with every score, verdict, timing, and tool-call count in one JSON file.

Download JSON

Published · Last updated . SREGym scenarios, configs, and public data verified on the date above.

Why Radar is faster

A Kubernetes cluster is a graph that changes over time. Deployments own ReplicaSets, which own Pods. Services select Pods by label, Ingresses route to Services, Pods mount ConfigMaps and Secrets. Most faults sit on an edge of that graph, or in a recent change to it.

routes toselects app=webownsownsmountsreadsIngressshopServicewebPodweb-7d9f…ReplicaSetweb-7d9fDeploymentwebConfigMapweb-configSecretweb-dbtimeline10:02 Deployment web: revision 310:04 ConfigMap web-config: DB_URL removed10:05 Pod web-7d9f: CrashLoopBackOffIssue: Pod web-7d9f is crash-looping. Likely cause: ConfigMap web-config lost DB_URL at 10:04, two minutes before the first restart.
What Radar keeps for every cluster, before anyone asks: the resources joined into a graph, every change on a timeline, and detectors that connect a symptom to its likely cause. Illustrative example.

14–18

median tool calls

Most MCP servers wrap the Kubernetes API

Their tools map to Kubernetes API calls: list, get, describe, logs. The agent rebuilds the graph itself: list the Pods, read the Service selector, match labels, pull events, line up timestamps. Each step is a round trip that returns raw YAML, and the next step waits for the model to read it.

8

median tool calls

Radar returns the joined graph

Radar watches every resource and keeps the joins current: owners, selectors, routes, config references. It records every change and event on a timeline and runs deterministic detectors over the result. The agent's first call returns what's broken, what it's connected to and what changed.

This work is deterministic: joins, a graph and detectors, kept current from the Kubernetes watch stream. Most of the engineering in Radar goes into it. The model is left to read the evidence and decide. The lead holds on the 16 faults every arm got right: 47s on average against 93s for the next MCP server.

FAQ

Kubernetes MCP benchmark questions

What is the fastest Kubernetes MCP server?
Radar was fastest in this benchmark. Across all 25 faults, Radar took 57s on average to submit a diagnosis. The next-fastest MCP server averaged 135s (2.3x slower) and raw kubectl 249s. Radar also led on the 16 faults every arm got right: 47s against 93s for the next MCP server.
Does this benchmark prove which Kubernetes MCP server is best?
It measures how fast each server gets an agent to a correct diagnosis. Radar was the fastest, had the highest average judge score, and used the fewest tool calls of the MCP servers. Choosing a server also depends on safety controls, write support, resource coverage, installation, multi-cluster reach, governance, and how much cluster context the server returns. The companion comparison page evaluates those product dimensions separately.
How were the Kubernetes MCP servers compared fairly?
Every arm used the same claude-sonnet-5 agent, SREGym fault scenarios, 3-node EKS, us-east-1, and claude-opus-5 judge. Time runs from agent start to its first submitted diagnosis. The page shows all 25 faults, then repeats the comparison on the 16 faults every arm diagnosed correctly so quick misses cannot explain the result. Radar leads in both views.
Which Kubernetes MCP server was most accurate?
This benchmark does not establish a meaningful accuracy winner. Most tools scored highly and the gaps are too small for a confident ranking. The public dataset includes every score; the clearer result here is diagnosis speed.
How does this differ from Radar's 50-fault MCP vs kubectl benchmark?
It is a separate, deeper two-arm study - Radar against raw kubectl on 50 faults, where Radar was 3x faster on average with slightly higher accuracy. The samples and runs differ, so compare each benchmark on its own.

Point your agent at Radar's MCP server.

It ships in the open source binary. Run it on your laptop or in your cluster; no account needed.

Quick install
$curl -fsSL https://get.radarhq.io | sh && kubectl radar

Apache 2.0 · No account for local use · Run Radar OSS forever