Skip to main content
Version: 1.0

Reactive AI Pipeline

The reactive pipeline detects problems, takes action, and measures results — detection, analysis, and action form a real, working chain, running automatically with no human intervention needed for auto-approved decisions. The one piece that isn't real yet is the "learns" part its name implies: the final loop, where feedback about an action's outcome tunes the system's own detection thresholds, has the logic built on both ends but isn't wired together — see below for exactly where.

Event Flow​

Reactive AI Pipeline — Event Flow

Autonomous detection → action → feedback. The loop back to detection isn't wired up yet.

MetricPoint
observability.telemetry
Decisions
analysis.decisions
action.outcomes
feedback.signals · no consumer
Insights
analysis.insights
Select any service to see its role in the event pipeline.

Adaptive Thresholds — designed, not yet running​

The analysis agent maintains per-cluster, per-metric Z-score thresholds in memory, and has real, tested logic to adjust them:

Feedback typeMeaningThreshold change
reinforcementThe triggering action executed successfully-0.05 (more sensitive)
correctionThe triggering action failed or errored+0.10 (less sensitive)

If this loop were actually running, a cluster whose auto-approved actions consistently succeed would gradually receive faster alerts (lower threshold), while one where actions keep failing would self-quieten. In practice, analysis-agent-srv never starts a consumer for the feedback.signals exchange this table's inputs are published to, so this adjustment code has never executed outside of tests. Worth being precise about what "the action improved the metric" means, too: feedback-agent-srv derives reinforcement/correction purely from whether the Kubernetes action itself succeeded or failed — there's no follow-up telemetry comparison checking whether cluster health actually got better. See Analysis Agent and Feedback Agent for the full detail on both halves.

Telemetry Message Schema​

type TelemetryPublishMessage struct {
SnapshotID string
ClusterID string
Timestamp time.Time
Snapshot TelemetrySnapshot
}

type TelemetrySnapshot struct {
ClusterID string
Source string // "k8s-monitor" | "security-api" | "cicd-gateway"
HealthScore float64
CPUUsagePct float64
MemoryUsagePct float64
ReadyNodeCount int
NodeCount int
PodCount int
FailedPodCount int
CrashLoopCount int
APIServerLatencyMs float64
NetworkRxBytesPS float64
NetworkTxBytesPS float64
SecurityPostureScore float64 // host-cluster targets only
}

See Observability Agent for how ClusterID gets populated — it's a real cluster ID for host-cluster targets, but a tenant's vCluster namespace for the per-tenant targets this service also discovers and collects from.

Decision Types​

The analysis agent produces decisions of the following types — see Analysis Agent for the exact metric/severity conditions and which decisions are auto-approved by default:

TypeDefault action (via action-agent-srv)
scaleScale the target Deployment's replicas
restartDelete the target Pod
cordonPatch the target Node unschedulable
notifyLog only; no Kubernetes action taken

Configuring Alert Rules​

This section covers a separate, related mechanism — anomaly-detector's own AlertRule system, not the analysis-agent Z-score pipeline described above. anomaly-detector consumes its own anomaly events from the k8s.anomalies exchange (a different exchange from analysis-agent's analysis.decisions/analysis.insights) and evaluates them against configured rules:

POST /api/anomalies/api/v1/alert-rules
Content-Type: application/json

{
"cluster_id": "prod-us-east",
"name": "High CPU on API pods",
"metric": "cpu_usage_pct",
"condition": "z_score_gt",
"threshold": 2.5,
"severity": "high",
"action": "scale_up",
"enabled": true,
"auto_approve": false
}

Set auto_approve: true to enable fully automated remediation. When false, the anomaly is logged and visible in the UI but no Kubernetes action is taken.