<#1034 Health system status for `osctrl-api` and `...
# osctrl
g
#1034 Health system status for `osctrl-api` and `osctrl-tls` Pull request opened by javuto Health / System Status Adds an admin page answering "is my deployment healthy?" — Database, Redis, osctrl-api, osctrl-tls and the TLS workers, with Go runtime detail per service and upgrade status. Behind
--health-enabled
, off by default. Why The signals existed but were invisible.
pkg/backend.DBHealth
tracked database degradation and only ever surfaced as a 204 on the unauthenticated
/health
probe. Redis was checked once at boot. osctrl-tls ran the ingest path, the alerts dispatcher, the batch writer and the activity writer entirely in its own process — a wedged worker looked exactly like a healthy one. The upgrade check already ran at boot and was thrown into a log line. Version skew between osctrl-api and osctrl-tls was observable nowhere. Costs nothing when off
--health-enabled
/
SERVICE_HEALTH_ENABLED
, default false, declared in
initServiceFlags
so both services take it. With the flag unset: the manager is never constructed, so
AutoMigrate
never runs and
service_status
is never
*created*; no route is registered; no heartbeat is written; no goroutine starts;
features.health
is false and the SPA hides the section. Residual cost is one bool in the flag struct. The final review traced all five of those through both
main.go
files. How it collects Nothing expensive on a timer, nothing cross-process on a request: | Source | When | | ------------------------------------------------------------------ | --------------------------- | | DB degraded flag + live ping, Redis PING, osctrl-api's own runtime | on request | | osctrl-tls: uptime, goroutines, worker counters, runtime stats | heartbeat, every 60s | | Upgrade check (external HTTP) | cached, refreshed every 24h |
runtime.ReadMemStats
stops the world, so it runs only on the API request path. osctrl-tls samples the same numbers through
runtime/metrics
, which reads the same counters without a global pause. The heartbeat rides the existing 5s
service_commands
poll loop, firing every 12th tick — no new goroutine, no new ticker: ~1,440 writes/day against the 17,280 selects that loop already issues, into a table that holds one upserted row per service forever. Endpoint and page
GET /api/v1/health/status
— admin only, 503 when the flag is off, one response carrying every component so the page's drill-downs expand rather than refetch. Statuses are `operational | degraded | down | stale | unknown`; a missing heartbeat is unknown, never down ("not reporting — enable
--health-enabled
on osctrl-tls"), since a service that was never enabled is not a dead one. Version skew between services surfaces as
degraded
, naming both sides. The page (
/_app/health
, nav gated on
features.health
) mirrors the three sections: component list with expandable details, upgrade status, and a runtime card per reporting service. Each card is labelled live (osctrl-api, read while serving) or as of HHMMSS (osctrl-tls, up to a heartbeat old) so nobody debugs a memory spike against stale numbers. Two defects the whole-branch review caught • The endpoint could hang during the outage it exists to report. Its DB and Redis checks had no deadline —
r.Context()
carries none — so a saturated pool blocked instead of returning
down
. Now bounded by a 3s budget threaded through all three checks, via new context-aware
Manager.GetContext
and
RedisManager.CheckContext
(existing callers unchanged). •
WorkersComponent
blamed osctrl-tls for database outages.
Manager.Get
returns
ErrNotReporting
only for a missing row; any other DB error rendered as "osctrl-tls is not reporting", pointing operators at a healthy service. It now mirrors `ServiceComponent`'s branch and surfaces the real error. Known limitations (documented in ARCHITECTURE.md) • Needs the flag on both services for a complete picture; api-only shows tls as not reporting. • Multiple osctrl-tls replicas share one row (last writer wins). Liveness stays correct; per-replica detail is lost. Upgrade path is a
(service, instance_id)
key plus pruning. •
LastGC
and
Lookups
have no
runtime/metrics
equivalent and stay zero for osctrl-tls;
PauseTotalNs
there is a histogram approximation. All documented at the sampler. • Snapshot only — no history. Files 37 files, +1,976/−6. New:
pkg/health
(model/manager, two runtime samplers, status derivation, version cache),
cmd/tls/health.go
,
cmd/api/handlers/health_status.go
,
frontend/src/{api/health.ts,features/health/}
. Modified: both `main.go`s,
pkg/config
(flag),
pkg/cache
+
pkg/alerts
(context/accessor additions), SPA router/nav/features, README, ARCHITECTURE, sample configs, regenerated
osctrl-api.yaml
. Verification
go build ./...
clean; full
go test ./...
green; frontend 329 tests pass;
tsc --noEmit
clean;
make openapi-check
exits 0. Every task was spec-and-quality reviewed, plus a whole-branch review whose two findings are fixed above. Not verified: the runtime smoke test — nobody booted the services against a fresh database to confirm the table is absent with the flag off, or watched a heartbeat flip to
stale
after stopping osctrl-tls. Both properties were traced statically and hold. jmpsec/osctrl