Action Agent
The action agent executes Kubernetes remediation actions in response to decisions from the analysis agent. It only processes decisions marked AutoApprove: true.
Actions
Six Kubernetes actions are supported. Scale/Restart/Cordon are only ever triggered automatically, from a consumed analysis.decisions message — there's no standalone HTTP endpoint to trigger them manually. Drain/Rollback/Uncordon have both an automated path and a dedicated POST endpoint for a human or an AI agent to trigger directly.
Scale Deployment
Patches the target Deployment's replica count via a JSON merge patch. If an explicit target replica count is given, it scales to exactly that value; otherwise it defaults to the current count plus one. There's no HPA maxReplicas cap applied here — that's an HPA-owned concern, not something this action checks.
// JSON merge patch
{ "spec": { "replicas": target } }
Scaling down is genuinely supported, not human-only: before a scale-down, the executor checks every matching PodDisruptionBudget in the namespace and refuses the patch if it would violate minAvailable or a maxUnavailable: 0 PDB.
Canary verification on scale-up (not gated behind any config beyond being on by default): after a successful scale-up, a background check waits CANARY_SETTLE_SECONDS (default 60s), then compares k8s-monitor's health score before and after. If it dropped by more than CANARY_ROLLBACK_HEALTH_DROP points (default 10.0), the scale-up is automatically reverted back to the pre-action replica count. This entire check can be disabled with CANARY_ENABLED=false.
Restart Pod
Deletes the target pod. The owning controller (Deployment, StatefulSet, DaemonSet) automatically recreates it with a fresh container:
client.CoreV1().Pods(namespace).Delete(ctx, podName, metav1.DeleteOptions{})
This is the standard Kubernetes restart mechanism — equivalent to kubectl rollout restart but targeted at a specific pod.
Cordon Node
Patches the node to set spec.unschedulable: true. New pods will not be scheduled on this node, but existing pods are not evicted:
// JSON merge patch
{ "spec": { "unschedulable": true } }
Cordoning a node does not drain it. Existing pods continue running. Use Drain (below) to also evict pods.
Uncordon Node
Reverses a previous cordon by patching spec.unschedulable: false. The scheduler resumes placing pods on the node.
Drain Node
Gracefully drains a node for maintenance or decommissioning:
- Cordons the node (
spec.unschedulable: true) - Lists all non-DaemonSet pods scheduled on the node
- Evicts each pod, one at a time, via the Kubernetes Eviction API (
policy/v1) with a 30-second grace period
DaemonSet pods are skipped, since they'd be immediately recreated on the same node. Eviction genuinely respects PodDisruptionBudgets — the API itself returns 429 when an eviction would violate a PDB, and each eviction is retried with exponential backoff (5s, doubling up to a 60s cap, 5 attempts total) rather than failing immediately. Because pods are evicted sequentially rather than in parallel, a node with several PDB-constrained pods can genuinely take a few minutes to fully drain — there's no separate fixed overall timeout on top of that.
POST /api/v1/actions/drain { "node_name": "ip-10-0-1-45.ec2.internal" }
Rollback Deployment
Triggers a rolling restart of a Deployment by patching the kubectl.kubernetes.io/restartedAt annotation with the current timestamp. Kubernetes treats this as a configuration change and performs a zero-downtime rolling restart, effectively rolling back to a known-good container state — this isn't a real revision rollback to a prior ReplicaSet, it's a fresh rollout of the pod template as it stands today.
POST /api/v1/actions/rollback { "namespace": "production", "deployment_name": "payments-api" }
Action Execution
The action agent consumes analysis.decisions from RabbitMQ. For each decision:
- Checks
decision.AutoApprove == true— skips if false - Executes the appropriate Kubernetes action in a goroutine
- Logs an
ActionRecordto PostgreSQL (status: executing) - On completion, updates
ActionRecord.statustosuccessorfailed - Publishes an
ActionOutcometo theaction.outcomesexchange
Decisions with AutoApprove: false show up instead in the pending-approval queue below, for a human to act on.
Human approval workflow
For decisions that aren't auto-approved, there's a real, separate REST-based approve/reject flow — GET /api/v1/actions/pending lists everything awaiting a decision, and POST /api/v1/actions/{id}/approve / .../reject resolve it. This is a plain authenticated REST call, not an AI-agent tool; no agent-runtime tool currently wraps either endpoint.
Kubernetes Auth
The action agent uses rest.InClusterConfig() when running inside Kubernetes, falling back to the standard ~/.kube/config path for local development (this fallback is a hardcoded path, not driven by a KUBECONFIG environment variable). It requires a ServiceAccount with the following permissions:
rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["list", "get", "delete"]
- apiGroups: [""]
resources: ["nodes"]
verbs: ["list", "get", "patch"]
- apiGroups: ["apps"]
resources: ["deployments"]
verbs: ["list", "get", "patch", "update"]
- apiGroups: ["policy"]
resources: ["pods/eviction", "poddisruptionbudgets"]
verbs: ["create", "list"]
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/actions | Action log (?cluster_id=&limit=) |
GET | /api/v1/actions/pending | Actions awaiting human approval |
GET | /api/v1/actions/{id} | Action detail |
POST | /api/v1/actions/{id}/approve | Approve a pending action |
POST | /api/v1/actions/{id}/reject | Reject a pending action |
POST | /api/v1/actions/drain | Drain a node ({ "node_name": "..." }) |
POST | /api/v1/actions/rollback | Rollback a deployment ({ "namespace": "...", "deployment_name": "..." }) |
POST | /api/v1/actions/uncordon | Uncordon a node ({ "node_name": "..." }) |
GET | /healthz | Health check |
Environment Variables
| Variable | Default | Description |
|---|---|---|
DATABASE_URL | — | PostgreSQL connection |
RABBITMQ_URL | — | RabbitMQ connection |
AUTH_JWT_ACCESS_SECRET | — | Required. Validates bearer tokens on the authenticated REST routes |
K8S_MONITOR_URL | http://k8s-monitor-srv:8085 | Used for the canary health-score check on scale-up |
CANARY_ENABLED | true | Toggles the scale-up canary check |
CANARY_SETTLE_SECONDS | 60 | How long to wait before comparing pre/post health scores |
CANARY_ROLLBACK_HEALTH_DROP | 10.0 | Health-score drop that triggers an automatic scale-up rollback |
PORT | 8094 | HTTP port |