Your cluster, healing itself, with your approval
Dorgu watches your cluster, diagnoses failures with AI, proposes a reviewable fix, and applies it when you approve. Runs in your own cluster. Apache-2.0.

An AI SRE for teams without an SRE
Kubernetes restarts a crash-looping pod forever and never asks why. Dorgu asks, then shows you the fix and waits for your call.
AI Self-Healing
Detect, diagnose, propose, approve, heal, remember. Dorgu spots OOMKills, crash loops, saturation, and node or control-plane trouble, works out the root cause, and writes an ordered plan. Every step carries its rationale, risk level, and a YAML diff you can read before anything happens.
the loopHuman-in-the-loop by default
Every remediation is approval-gated. Resource changes are capped at 2× blast radius, limited to 5 per app per hour, and kube-system is always excluded. If health regresses after a fix, Dorgu rolls it back automatically.
approval requiredIncident memory
IncidentMemory and RemediationAction CRDs keep the signal, the root cause, the confidence, the plan, and the outcome as first-class cluster objects. Organizational memory that outlives the Slack thread and feeds the next diagnosis.
IncidentMemory CRDYour cluster, your keys
Apache-2.0 and self-hosted. AI is optional and bring-your-own Anthropic key. Detection, diagnosis, and remediation all work rule-based with no key at all. Your incidents stay as CRDs in your cluster. No lock-in.
Apache 2.0Kubernetes Operator
Validate deployments against personas. Advisory or enforcing webhooks, Prometheus-based resource learning, ArgoCD sync tracking. It never creates or modifies your workloads, only the persona and incident records.
OperatorApplication Personas
Give your apps identity. ApplicationPersona CRDs capture what your app needs (resources, scaling, health, dependencies, ownership) and give every signal something to correlate to.
CRDCluster Personas
Give your cluster a soul. ClusterPersona CRDs auto-discover nodes, addons, capacity, and state: the cluster context the AI plans against.
ClusterPersona CRDAI Manifest Generation
Getting started from scratch? Point dorgu at your Dockerfile or Compose file for production-ready Deployments, Services, Ingress, HPA, ArgoCD config, CI/CD workflows, and a matching persona.
dorgu generateCluster Setup Wizard
Bootstrap a production stack in minutes. cert-manager, ingress-nginx, CloudNativePG, OpenObserve, Argo CD, External Secrets, with an educational wizard that teaches as it installs.
Blessed StackGitOps Native
Generates ArgoCD Applications, scaffolds App-of-Apps directories, respects your GitOps workflows. Approve a fix with --no-heal and apply it through your own pipeline.
ArgoCDPlatform Dashboard
A live view of your cluster: nodes, capacity, addons, and ClusterPersona state over WebSockets. Incidents and remediations are reviewed from the CLI today.
dorgu platform serve
From failure to fix, in six steps
Detect, diagnose, propose, approve, heal, remember.
Code detects. AI explains. A human approves. Nothing touches your workloads until you say so. Every command is readable, and every record stays in your cluster.
Install
One Helm command, in your own cluster. Add an Anthropic key if you want AI diagnosis and AI-written plans. Everything works rule-based without one.
Dorgu detects
The health-check reconciler watches for OOMKills, crash loops, image-pull failures, CPU and memory saturation, and node or control-plane trouble, every 60s by default or 30s for a tight loop. Each signal opens an IncidentMemory.
AI diagnoses
Deterministic rules produce a root cause and a confidence score. With a key configured, Claude enhances that with cluster context. Any AI failure degrades to the rules and never blocks the loop.
It proposes a fix
An ordered, reviewable plan lands as a RemediationAction, every step carrying its rationale, risk level, and a YAML diff. Capped at 2× blast radius, 5 remediations per app per hour, kube-system excluded.
You approve, it heals
Nothing is applied until you say so. The operator patches the persona's desired state; the CLI patches the Deployment with your credentials. If health regresses during the verification window, Dorgu rolls it back.
It remembers
The signal, the root cause, the plan, and the outcome persist as CRDs in your cluster, and become context the next proposal is written against.
Simple, transparent pricing
The entire self-healing loop is open source and free forever. Paid tiers are on the roadmap, and nothing in them ships today.
Join the Waitlist
Help us understand your needs. Takes less than 3 minutes.