AI investigations

Root cause for a broken cluster, in about a minute.

An investigation is only as good as the cluster context the agent reads. Ask your own agent, click Investigate in Radar, or let Radar Cloud run it on every alert. The same cluster model sits underneath all three.

1.6–3.6×faster30–85%cheaperthan the AI SRE tools we tested, on live faults

Three ways to run one

Claude Code · Radar MCPbenchmark run · kubectl took 11.7 min
  1. 4.6sissues namespace=astronomy-shop
  2. 4.9sget_changes namespace=astronomy-shop
  3. 11.8sget_resource deployment/frontend
  4. 12.3sget_workload_logs frontend
  5. 20.4sdiagnose deployment/product-catalog
  6. 34.8sfrontend lost CART_ADDR (was cart:8080) at 22:15:29, so checkout can't reach the cart service. Cart itself is healthy.

Open source

Your agent, over MCP

Ask Claude Code, Codex, Cursor, Copilot or any MCP agent from your terminal or IDE. Radar's MCP server answers with the cluster already correlated.

The MCP server

Open source

Your agent CLI, run from Radar

Click Investigate on anything broken. Radar runs Claude Code, Codex, Cursor or OpenCode read-only on your machine and lays out the findings, the evidence and a fix you approve.

What you see
  1. 1

    Alert fires

    A rule matches an issue on any connected cluster.

  2. 2

    Radar Cloud investigates

    Its agent runs read-only against the same cluster model.

  3. 3

    Root cause in Slack

    With the evidence, and a link to the full investigation.

Radar Cloud

Radar Cloud's own agent

Runs automatically when an alert fires, or on demand, and sends the root cause to Slack with the evidence behind it. Nothing to install on anyone's laptop.

Radar Cloud

Why Radar is faster than dedicated AI SRE tools

The agent starts from a cluster model that is already joined, time-indexed and checked for failures.

  • Failures are already detected

    Radar's Issues engine ranks failures before anyone asks. In our benchmark it was the agent's first call on all 50 faults.

  • Changes are on a timeline

    Every spec change and event sits on a timeline, so "what changed before this broke" is one call.

  • Resources come back joined

    Owners, topology and error-filtered logs come back correlated and secret-redacted in one response, so the agent starts with the connections already made.

routes toselects app=webownsownsmountsreadsIngressshopServicewebPodweb-7d9f…ReplicaSetweb-7d9fDeploymentwebConfigMapweb-configSecretweb-dbtimeline10:02 Deployment web: revision 310:04 ConfigMap web-config: DB_URL removed10:05 Pod web-7d9f: CrashLoopBackOffIssue: Pod web-7d9f is crash-looping. Likely cause: ConfigMap web-config lost DB_URL at 10:04, two minutes before the first restart.
What Radar keeps for every cluster, before anyone asks: the resources joined into a graph, every change on a timeline, and detectors that connect a symptom to its likely cause. Illustrative example.

Measured in the AI SRE benchmark, the Kubernetes MCP server benchmark and the MCP vs kubectl benchmark.

What the investigation workspace shows

From Radar's UI or Radar Cloud, an investigation shows its findings, the evidence behind each one and a proposed fix.

  • Starts from the issue

    Investigate sits on the issue itself. There's no prompt to write and no context to paste.

  • Evidence for each claim

    Evidence cards carry the log lines, events and resources the agent relied on. One click opens the real object.

  • Stated confidence

    Established, Likely, or Still open, with the questions it couldn't settle listed plainly.

  • Fixes need approval

    Approve a fix and it runs as a separate session. Radar then checks whether it actually worked.

Investigate is on every issue.
The Activity tab shows every call the agent made.

Safety limits

  • In Radar's UI and Radar Cloud, investigations are read-only: write tools aren't loaded. Over MCP, your agent's write access is your choice, within your RBAC.
  • An approved fix runs alone, in its own session.
  • Argo CD, Flux and Helm-managed changes ask first.
  • Open source runs on your machine and your model account.

Why the limits live in the tool layer: Read-Only Is Not a Safety Boundary.

FAQ

AI investigation questions

What is an AI investigation in Radar?
An AI agent works out why something in your Kubernetes cluster is broken, using Radar's live model of the cluster, and returns the root cause, the evidence behind it and a proposed fix. You can run one from your own agent over MCP, from Radar's UI with your agent CLI, or let Radar Cloud run one when an alert fires.
Do I need to buy an AI SRE tool?
Probably not. You likely already have the agent. What it needs is the cluster context: an investigation is only as good as the cluster context the agent reads. Radar gives that context to the agent you already use: Claude Code, Codex, Cursor, anything that speaks MCP. In our benchmark of five AI SRE products on 50 live Kubernetes faults, Claude Code on Radar's tools was the most accurate of the five: clearly ahead of HolmesGPT, OpenSRE and AWS DevOps Agent, and slightly ahead of kubectl-ai. It answered in an average of 79 seconds at $0.27 a fault, the fastest and cheapest of the five. If you'd rather not run it yourself, Radar Cloud runs the same setup on every alert: Claude on Radar's tools, the combination we benchmarked.
How does Radar compare with HolmesGPT, kubectl-ai and AWS DevOps Agent?
On faults that announce themselves, like a crash loop or a failed image pull, all five products scored between 24 and 27 of 28. They separate on the 14 faults where every pod reports Ready: Radar 12, kubectl-ai 9, HolmesGPT 7, OpenSRE 6 and AWS DevOps Agent 5. Every result is downloadable from the benchmark page.
Which agents work with Radar?
Over MCP, any client that speaks it, including Claude Code, Codex, Cursor, GitHub Copilot, OpenCode and Google Antigravity. From Radar's UI, Claude Code, Codex, Cursor and OpenCode; Radar detects the ones on your PATH. Radar Cloud brings its own agent.
Where does my cluster data go?
With open source Radar, investigations run on your machine and what the agent reads goes to your own model provider, never to us. In Radar Cloud, each investigation runs as a sandboxed job with read-only access, and an owner accepts that once for the organization.
Can an investigation change my cluster?
It depends on how you run it. Radar's UI and Radar Cloud run investigations read-only: write tools aren't loaded. When you connect your own agent over MCP, you decide: Radar's write tools are there if you enable them, and your RBAC limits what they can do. When you approve a fix in Radar, it runs in a separate session bound to that one change, and changes to resources managed by Argo CD, Flux or Helm ask for an explicit acknowledgment because the next sync would revert them.

Connect your agent to Radar

Radar is open source, one binary. Run it on your laptop or in your cluster.

Quick install
$curl -fsSL https://get.radarhq.io | sh && kubectl radar

Apache 2.0 · No account for local use · Run Radar OSS forever