AI Kubernetes Troubleshooting Agent

An AI platform that investigates Kubernetes failures, finds the root cause, and suggests fixes through a web dashboard.

AI Kubernetes Troubleshooting Agent

What changed

  • Turned manual Kubernetes debugging into a one-click investigation that ends with a root cause, a concrete fix and a confidence score.
  • Saved every investigation, so the team can look back at past incidents.

What I worked on

  • I built the investigation layer with five inspectors that collect pods, logs, events, deployments, and network state.
  • I designed the AI agent flow that turns raw cluster evidence into a root cause, a fix, and a confidence score.
  • I wired the InsForge backend for auth, investigation history, and realtime progress updates in the dashboard.

When a pod keeps failing, the answer is usually spread across its status, its logs, the cluster events and the deployment config. This agent collects all of it in one click and hands it to an AI to reason over.

Example

On a test cluster with three broken apps, it found each problem separately: a missing DATABASE_URL crashing the payment service, an image tag that doesn't exist, and an analytics worker in a crash loop with empty logs. It gave kubectl commands for the first two, and said its confidence was lower for the third because the evidence was thin.

Each step of the investigation ticks off live in the dashboard

How it's built

Five inspectors talk to the Kubernetes API and gather their findings into one evidence report. A prompt builder passes the report to an LLM through OpenRouter (Claude, GPT or DeepSeek), and the agent returns the root cause, a fix such as a kubectl command or a YAML change, and a confidence score. InsForge handles auth, investigation history and the live progress updates in the dashboard.

Past investigations are saved with their root cause and confidence