GitHub
08/31/2026, 6:16 PMnode_inactive and
node_recovered rules.
Change
Watcher — new pkg/alerts/inactive.go
• Periodic sweep with state-transition semantics: an active→inactive
transition fires node_inactive once per node; an inactive→active
transition fires node_recovered once (only when a node_recovered rule
exists — recovery alerts are opt-in); an unchanged sweep is silent
• Redis set (osctrl:alert:inactive) records flagged UUIDs so a restart
does not re-alert the whole fleet; the set is replaced wholesale each pass
so recovered nodes drop out
• State access is behind a small inactiveState interface (Redis adapter in
production, in-memory map in tests); load/save failures fail open —
same posture as the cooldown gate (alert availability beats perfect
dedupe on a Redis blip)
• Zero-cost guards: no node-state rules in the snapshot → the sweep returns
before any source query; the active-node query runs only when recovery
rules exist
• Run(stop, interval) ticker loop, 5-minute default cadence
Snapshot extension — pkg/alerts/rules.go, manager.go
• RuleSet gains nodeInactive / nodeRecovered buckets; LoadSnapshot
routes the sources so hot-reload (Stage 4's reload-alerts command)
covers them too
• Env scoping identical to the log matchers: global rules (env 0) fire
everywhere, scoped rules only in their environment — enforced per node
via ruleApplies(envID)
osquery-facing source — new cmd/tls/tls_alerts.go
• tlsNodeSource adapts the nodes manager to the watcher's NodeSource
• Queries are batched per environment so each environment's
inactive_hours setting applies correctly (global setting then the
72h default as fallbacks) — a single global query would use the wrong
cutoff for environments with custom thresholds
• One environment's query failure does not drop the others
Wiring — cmd/tls/main.go
• Watcher constructed and started behind --alerts-enabled (5-minute
interval), stopped on shutdown/restart via the reworked stopAlerts
(now closes the watcher loop before draining the dispatch queue)
Validation
• go build ./..., go vet, gofmt — clean
• golangci-lint run ./pkg/alerts/... ./cmd/tls/... — 0 issues
• New tests, all passing:
• Full lifecycle with in-memory state + synthetic clock: inactive fires
once → unchanged second sweep silent → recovery fires once → not
repeated → re-inactive fires again
• Env scoping: an env-7 rule hits only the prod node, not the dev node
• No-rules no-work (no hits, no source queries); source errors no-op
• humanSince rendering table (never / 30m / 5h / 4d)
• Live-Redis variant asserts the same lifecycle through the real Redis
set (auto-skipped without REDIS_URL)
• Full affected suites green: pkg/alerts, cmd/tls, cmd/tls/handlers,
cmd/api/handlers
Performance
• The sweep never touches the ingest path; it runs on its own ticker
• Per-env batched queries (one GetByEnv per environment per pass) — no
per-node lookups; the Stage-0 design target of "batched SQL, not
per-node queries" is met
• No-rules deployments pay zero queries; recovery detection costs one
extra query per env only when recovery rules exist configured
Roadmap position
Stage 5 of 7 (alerts-inactive). Both original alert classes are now
live: log-content matching (Stages 1–4) and node liveness (this stage),
both manageable via the Stage-4 API/CLI and hot-reloadable. Remaining:
Stage 6 (SPA "Alerts" section), Stage 7 (load hardening + alert-history
retention sweep).
jmpsec/osctrlGitHub
08/31/2026, 7:00 PM