Nodes Manager
Repo: o-apps/nodes-manager · Port: 8115 · DB Schema: nodes
nodes-manager owns node-lifecycle reconciliation for the cluster. It is worth being precise about what it actually does today versus what it's designed to eventually do, because the two are further apart here than for most services in this platform: a substantial amount of intelligent scheduling and Karpenter integration code exists and is exposed over its REST API, but is not yet wired into the service's own background loop, and Karpenter itself is not yet installed in any environment. What genuinely runs continuously today is node state synchronization and an auto-uncordon reconciler described below.
What actually runs continuously
Every scan interval (SCAN_INTERVAL_SECONDS, no fixed default — set per environment), nodes-manager does two things:
- Syncs every Kubernetes node's state into Postgres (
nodes.nodes, upserted), so the REST API below has something real to read. - Runs the auto-uncordon reconciler. AWS EC2 issues a Spot rebalance recommendation when a Spot instance is at elevated risk of reclamation.
aws-node-termination-handler, running as a DaemonSet, reacts to that recommendation by cordoning the affected node — but a rebalance recommendation is only a risk signal, not a guaranteed interruption, and most of the time no actual reclaim follows. Left alone, that node would sitReadybut permanently unschedulable, since nothing else in the cluster ever reverses the cordon. nodes-manager watches for exactly this pattern — a node cordoned by nothing but the plainnode.kubernetes.io/unschedulabletaint (never a manual maintenance cordon or, once Karpenter is installed, one of its own disruption taints), continuouslyReadythroughout, and past a configurable grace period (AUTO_UNCORDON_GRACE_PERIOD_MINUTES, default 15 minutes — comfortably longer than a real Spot interruption's roughly two-minute final notice) — and uncordons it automatically, recording the action as ascaling_eventsrow and, when RabbitMQ is configured, anodes.eventsmessage.
A heartbeat is published to the nodes.events RabbitMQ topic exchange alongside these events when RABBITMQ_URL is set.
Built, exposed over the API, but not yet wired into the running loop
The remaining capabilities below are real, tested code — reachable through the REST endpoints in the table further down — but the background scheduling logic they represent (ScaleUpUseCase, ScaleDownUseCase, OptimizeUseCase) is constructed at startup and explicitly not invoked from the management loop yet. Calling their endpoints directly still works; nothing currently triggers them on a schedule.
- AI-scored workload placement. Each pod is classified into a
WorkloadClass(latency_sensitive,gpu,batch,stateful, orstateless, by QoS class, GPU requests, owning controller, and volume mounts) and scored across five weighted dimensions — resource fit (35%), topology spread across zones (25%), cost efficiency of spot versus on-demand (20%), spot interruption risk (15%), and data locality to existing volumes (5%) — to recommend a placement, with a plain-language rationale attached to each decision. - Karpenter NodePool management.
NodePoolandNodeClaimare accessed through the Kubernetes dynamic client, so no Karpenter Go library is required. If the CRDs aren't installed — the case in every environment as of this writing — every Karpenter-related endpoint degrades gracefully to an empty list with a log warning rather than an error, rather than failing. - Spot market advisor. Queries EC2's
DescribeSpotPriceHistory(5-minute cache per region) for current spot pricing and a rough interruption-frequency estimate per instance type and zone. - Consolidation planning. Identifies underutilized nodes that would be safe to drain given PodDisruptionBudget constraints, and estimates the resulting savings.
- On-demand AWS provisioning. A thin wrapper over
RunInstances/DescribeInstanceTypes, for provisioning outside of Karpenter.
REST API
| Method | Path | Description |
|---|---|---|
GET | /api/v1/nodes | List nodes with state, resources, and pod counts |
GET | /api/v1/nodes/{id} | Node detail with running pods |
POST | /api/v1/nodes/{id}/drain | Gracefully drain a node |
GET | /api/v1/decisions | AI scheduling decision log (populated once the scheduler is wired in) |
GET | /api/v1/optimize/plan | Consolidation plan with savings and PDB impact |
POST | /api/v1/optimize/trigger | Manually trigger an optimisation cycle |
POST | /api/v1/simulate | Simulate scheduling N pods → projected node additions |
GET | /api/v1/nodepools | List Karpenter NodePools (empty until Karpenter is installed) |
GET | /api/v1/nodepools/{name} | NodePool detail |
POST | /api/v1/nodepools | Create NodePool |
PATCH | /api/v1/nodepools/{name} | Update NodePool limits or disruption policy |
DELETE | /api/v1/nodepools/{name} | Delete NodePool |
GET | /api/v1/nodepools/{name}/nodeclaims | List NodeClaims for a pool |
POST | /api/v1/nodepools/{name}/consolidate | Trigger Karpenter consolidation |
GET | /api/v1/spot/market | Spot pricing and interruption data by zone |
GET | /api/v1/workloads/placement | All workloads with node, class, score, and rationale |
GET | /healthz | Health check |
Environment Variables
| Variable | Default | Description |
|---|---|---|
DATABASE_URL | — | PostgreSQL connection |
AUTH_JWT_ACCESS_SECRET | — | Required — shared JWT signing secret for the HTTP API |
RABBITMQ_URL | — | Optional — enables publishing to nodes.events |
AUTO_UNCORDON_ENABLED | true | Enables the reconciler described above |
AUTO_UNCORDON_GRACE_PERIOD_MINUTES | 15 | How long a node must have been cordoned, with nothing but the bare unschedulable taint, before it's reversed |
KUBECONFIG | in-cluster | Path to kubeconfig (local development only) |
AWS_REGION | — | Region for spot pricing queries |
AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY | — | AWS credentials for the on-demand provisioner |
APP_PORT | 8115 | HTTP port |
A note on the service's history
This service has, for most of its deployment, been silently non-functional: its Docker image never actually shipped its own database migrations, so every startup ran goose against an empty directory, logged a warning, and continued — meaning nodes.nodes and the other tables never existed, and every read or write against them failed. A second, related bug (missing type:jsonb annotations on the ORM's JSON-holding columns) was masked entirely by the first, and only surfaced once the migration itself was fixed. Both are now fixed, and the auto-uncordon reconciler described above has been confirmed live: during its own rollout verification, it found several real nodes that had sat cordoned for 33 minutes to over an hour from an actual Spot rebalance event, and correctly uncordoned all of them with no manual intervention.