<#1011 Alerting system Stage 7: retention sweep + ...
# osctrl
g
#1011 Alerting system Stage 7: retention sweep + ingest load hardening Pull request opened by javuto Alerting system Stage 7: retention sweep + ingest load hardening Problem The alerting pipeline (Stages 1–6) wrote a history row per dispatched alert with nothing bounding the
alert_history
table, and the Stage-0 performance contract — matcher attachment must not meaningfully regress log ingest — had only been validated at the matcher level, never on the full ingest path. This stage closes both gaps and is the final one of the alerting roadmap. Change Alert-history retention •
pkg/settings
— new
alert_history_retention_days
setting with an
AlertHistoryRetentionDays()
getter; *default 30 days*; zero/negative values fall back to the default so the table can never grow unbounded by accident •
pkg/alerts
—
PruneHistoryWithRetention(days, now)
computes the cutoff from the window with a defensive zero-guard (a direct caller cannot accidentally delete everything) •
cmd/tls
—
watchAlertRetention
prunes once at boot (catch-up after downtime) then daily, honoring the setting live; never touches the ingest path; stopped on shutdown via the extended
stopAlerts
Load validation — and a real optimization it surfaced • New `pkg/logging/bench_alerts_test.go`: benchmarks the full
ProcessLogs
path
over realistic 10-entry batches, with and without a matcher attached, for both result and status logs — the CI-runnable form of the Stage-0 regression gate • First status-path numbers showed +50% on the ProcessLogs slice: the Stage-2 hook had decoded the status batch a second time for the matcher • Fixed:
ProcessLogs
now decodes the status batch once into the full schema when a matcher is attached and derives the generic metadata view from it — the matcher consumes the same slice, no second unmarshal anywhere on the hot path. The feature-off path keeps the historical cheap generic-only decode, so the no-alerts baseline is byte-for-byte unchanged Performance results | Path (10-entry batch) | baseline | + matcher | delta | | --------------------- | -------- | --------- | ------------------- | | Result logs | 51.7 µs | 51.7 µs | ~0% | | Status logs | 37.9 µs | 45.3 µs | +19% of ProcessLogs | Endpoint-level context: the added ~0.75 µs/entry sits inside a
/log
request dominated by envelope parse and DB writes (≥1 ms typical), so the request-level regression is < 1% — comfortably within the Stage-0 "< 5%" contract. The matcher-side gates from Stage 1 still hold: 3.0 ms for 20 rules × 500 rows, 2.1 ns / 0 allocations when the feature is off or no rules exist. Validation • New tests:
TestPruneHistoryWithRetention
(40-day row pruned, 5-day and 29-day rows survive, zero-retention is a no-op) and
TestAlertHistoryRetentionSetting
(absent → default, zero → default, 90-day override honored) • All affected suites pass with
-race
:
pkg/alerts
,
pkg/logging
,
pkg/settings
,
cmd/tls
,
cmd/tls/handlers
,
cmd/api
,
cmd/api/handlers
•
go build ./...
,
go vet
,
gofmt
clean;
golangci-lint
— 0 issues • Full benchmark report captured above for the PR record Roadmap complete This closes the 7-stage alerting plan delivered across the previous PRs: 1. Engine — lock-free rule snapshots, pure matchers, Redis cooldown gate #1005 2. Ingest hook — non-blocking queue + worker pool, Prometheus metrics #1006 3. Channels — webhook (HMAC, SSRF guard, retries) + email, fan-out isolation #1007 4. API + CLI — full CRUD, secret redaction,
reload-alerts
hot-reload #1008 5. Node liveness — transition-based inactive/recovered sweep, Redis state #1009 6. SPA — Alerts section (rules/channels/history tabs, registry-driven forms) #1010 7. Hardening — this PR The subsystem remains fully inert behind `--alerts-enabled`: no tables, no hooks, no routes, no UI when the flag is off. jmpsec/osctrl