<#1009 Alerting system Stage 5: node-inactive / no...
# osctrl
g
#1009 Alerting system Stage 5: node-inactive / node-recovered watcher Pull request opened by javuto Alerting system Stage 5: node-inactive / node-recovered watcher Problem Stages 1–4 cover alerting on log content — matches in result, status, and distributed-query logs — but not on absence of traffic. "Alert when a node has gone silent" was one of the two original alert classes and cannot be observed from an event stream: only a clock can detect that a node stopped reporting. This PR adds the sweep-based watcher for
node_inactive
and
node_recovered
rules. Change Watcher — new
pkg/alerts/inactive.go
• Periodic sweep with state-transition semantics: an active→inactive transition fires
node_inactive
once per node; an inactive→active transition fires
node_recovered
once (only when a
node_recovered
rule exists — recovery alerts are opt-in); an unchanged sweep is silent • Redis set (
osctrl:alert:inactive
) records flagged UUIDs so a restart does not re-alert the whole fleet; the set is replaced wholesale each pass so recovered nodes drop out • State access is behind a small
inactiveState
interface (Redis adapter in production, in-memory map in tests); load/save failures fail open — same posture as the cooldown gate (alert availability beats perfect dedupe on a Redis blip) • Zero-cost guards: no node-state rules in the snapshot → the sweep returns before any source query; the active-node query runs only when recovery rules exist •
Run(stop, interval)
ticker loop, 5-minute default cadence Snapshot extension —
pkg/alerts/rules.go
,
manager.go
•
RuleSet
gains
nodeInactive
/
nodeRecovered
buckets;
LoadSnapshot
routes the sources so hot-reload (Stage 4's
reload-alerts
command) covers them too • Env scoping identical to the log matchers: global rules (env 0) fire everywhere, scoped rules only in their environment — enforced per node via
ruleApplies(envID)
osquery-facing source — new
cmd/tls/tls_alerts.go
•
tlsNodeSource
adapts the nodes manager to the watcher's
NodeSource
• Queries are batched per environment so each environment's
inactive_hours
setting applies correctly (global setting then the 72h default as fallbacks) — a single global query would use the wrong cutoff for environments with custom thresholds • One environment's query failure does not drop the others Wiring —
cmd/tls/main.go
• Watcher constructed and started behind
--alerts-enabled
(5-minute interval), stopped on shutdown/restart via the reworked
stopAlerts
(now closes the watcher loop before draining the dispatch queue) Validation •
go build ./...
,
go vet
,
gofmt
— clean •
golangci-lint run ./pkg/alerts/... ./cmd/tls/...
— 0 issues • New tests, all passing: • Full lifecycle with in-memory state + synthetic clock: inactive fires once → unchanged second sweep silent → recovery fires once → not repeated → re-inactive fires again • Env scoping: an env-7 rule hits only the prod node, not the dev node • No-rules no-work (no hits, no source queries); source errors no-op •
humanSince
rendering table (never / 30m / 5h / 4d) • Live-Redis variant asserts the same lifecycle through the real Redis set (auto-skipped without
REDIS_URL
) • Full affected suites green:
pkg/alerts
,
cmd/tls
,
cmd/tls/handlers
,
cmd/api/handlers
Performance • The sweep never touches the ingest path; it runs on its own ticker • Per-env batched queries (one
GetByEnv
per environment per pass) — no per-node lookups; the Stage-0 design target of "batched SQL, not per-node queries" is met • No-rules deployments pay zero queries; recovery detection costs one extra query per env only when recovery rules exist configured Roadmap position Stage 5 of 7 (
alerts-inactive
). Both original alert classes are now live: log-content matching (Stages 1–4) and node liveness (this stage), both manageable via the Stage-4 API/CLI and hot-reloadable. Remaining: Stage 6 (SPA "Alerts" section), Stage 7 (load hardening + alert-history retention sweep). jmpsec/osctrl