Skip to main content
Version: 1.0

Action Agent

The action agent executes Kubernetes remediation actions in response to decisions from the analysis agent. It only processes decisions marked AutoApprove: true.

Actions​

Six Kubernetes actions are supported. Scale/Restart/Cordon are only ever triggered automatically, from a consumed analysis.decisions message — there's no standalone HTTP endpoint to trigger them manually. Drain/Rollback/Uncordon have both an automated path and a dedicated POST endpoint for a human or an AI agent to trigger directly.

Scale Deployment​

Patches the target Deployment's replica count via a JSON merge patch. If an explicit target replica count is given, it scales to exactly that value; otherwise it defaults to the current count plus one. There's no HPA maxReplicas cap applied here — that's an HPA-owned concern, not something this action checks.

// JSON merge patch
{ "spec": { "replicas": target } }

Scaling down is genuinely supported, not human-only: before a scale-down, the executor checks every matching PodDisruptionBudget in the namespace and refuses the patch if it would violate minAvailable or a maxUnavailable: 0 PDB.

Canary verification on scale-up (not gated behind any config beyond being on by default): after a successful scale-up, a background check waits CANARY_SETTLE_SECONDS (default 60s), then compares k8s-monitor's health score before and after. If it dropped by more than CANARY_ROLLBACK_HEALTH_DROP points (default 10.0), the scale-up is automatically reverted back to the pre-action replica count. This entire check can be disabled with CANARY_ENABLED=false.

Restart Pod​

Deletes the target pod. The owning controller (Deployment, StatefulSet, DaemonSet) automatically recreates it with a fresh container:

client.CoreV1().Pods(namespace).Delete(ctx, podName, metav1.DeleteOptions{})

This is the standard Kubernetes restart mechanism — equivalent to kubectl rollout restart but targeted at a specific pod.

Cordon Node​

Patches the node to set spec.unschedulable: true. New pods will not be scheduled on this node, but existing pods are not evicted:

// JSON merge patch
{ "spec": { "unschedulable": true } }
warning

Cordoning a node does not drain it. Existing pods continue running. Use Drain (below) to also evict pods.

Uncordon Node​

Reverses a previous cordon by patching spec.unschedulable: false. The scheduler resumes placing pods on the node.

Drain Node​

Gracefully drains a node for maintenance or decommissioning:

  1. Cordons the node (spec.unschedulable: true)
  2. Lists all non-DaemonSet pods scheduled on the node
  3. Evicts each pod, one at a time, via the Kubernetes Eviction API (policy/v1) with a 30-second grace period

DaemonSet pods are skipped, since they'd be immediately recreated on the same node. Eviction genuinely respects PodDisruptionBudgets — the API itself returns 429 when an eviction would violate a PDB, and each eviction is retried with exponential backoff (5s, doubling up to a 60s cap, 5 attempts total) rather than failing immediately. Because pods are evicted sequentially rather than in parallel, a node with several PDB-constrained pods can genuinely take a few minutes to fully drain — there's no separate fixed overall timeout on top of that.

POST /api/v1/actions/drain  { "node_name": "ip-10-0-1-45.ec2.internal" }

Rollback Deployment​

Triggers a rolling restart of a Deployment by patching the kubectl.kubernetes.io/restartedAt annotation with the current timestamp. Kubernetes treats this as a configuration change and performs a zero-downtime rolling restart, effectively rolling back to a known-good container state — this isn't a real revision rollback to a prior ReplicaSet, it's a fresh rollout of the pod template as it stands today.

POST /api/v1/actions/rollback  { "namespace": "production", "deployment_name": "payments-api" }

Action Execution​

The action agent consumes analysis.decisions from RabbitMQ. For each decision:

  1. Checks decision.AutoApprove == true — skips if false
  2. Executes the appropriate Kubernetes action in a goroutine
  3. Logs an ActionRecord to PostgreSQL (status: executing)
  4. On completion, updates ActionRecord.status to success or failed
  5. Publishes an ActionOutcome to the action.outcomes exchange

Decisions with AutoApprove: false show up instead in the pending-approval queue below, for a human to act on.

Human approval workflow​

For decisions that aren't auto-approved, there's a real, separate REST-based approve/reject flow — GET /api/v1/actions/pending lists everything awaiting a decision, and POST /api/v1/actions/{id}/approve / .../reject resolve it. This is a plain authenticated REST call, not an AI-agent tool; no agent-runtime tool currently wraps either endpoint.

Kubernetes Auth​

The action agent uses rest.InClusterConfig() when running inside Kubernetes, falling back to the standard ~/.kube/config path for local development (this fallback is a hardcoded path, not driven by a KUBECONFIG environment variable). It requires a ServiceAccount with the following permissions:

rules:
- apiGroups: [""]
resources: ["pods"]
verbs: ["list", "get", "delete"]
- apiGroups: [""]
resources: ["nodes"]
verbs: ["list", "get", "patch"]
- apiGroups: ["apps"]
resources: ["deployments"]
verbs: ["list", "get", "patch", "update"]
- apiGroups: ["policy"]
resources: ["pods/eviction", "poddisruptionbudgets"]
verbs: ["create", "list"]

REST API​

MethodPathDescription
GET/api/v1/actionsAction log (?cluster_id=&limit=)
GET/api/v1/actions/pendingActions awaiting human approval
GET/api/v1/actions/{id}Action detail
POST/api/v1/actions/{id}/approveApprove a pending action
POST/api/v1/actions/{id}/rejectReject a pending action
POST/api/v1/actions/drainDrain a node ({ "node_name": "..." })
POST/api/v1/actions/rollbackRollback a deployment ({ "namespace": "...", "deployment_name": "..." })
POST/api/v1/actions/uncordonUncordon a node ({ "node_name": "..." })
GET/healthzHealth check

Environment Variables​

VariableDefaultDescription
DATABASE_URL—PostgreSQL connection
RABBITMQ_URL—RabbitMQ connection
AUTH_JWT_ACCESS_SECRET—Required. Validates bearer tokens on the authenticated REST routes
K8S_MONITOR_URLhttp://k8s-monitor-srv:8085Used for the canary health-score check on scale-up
CANARY_ENABLEDtrueToggles the scale-up canary check
CANARY_SETTLE_SECONDS60How long to wait before comparing pre/post health scores
CANARY_ROLLBACK_HEALTH_DROP10.0Health-score drop that triggers an automatic scale-up rollback
PORT8094HTTP port