GitHub
09/07/2026, 6:04 AM--health-enabled, off by default.
Why
The signals existed but were invisible. pkg/backend.DBHealth tracked database
degradation and only ever surfaced as a 204 on the unauthenticated /health
probe. Redis was checked once at boot. osctrl-tls ran the ingest path, the
alerts dispatcher, the batch writer and the activity writer entirely in its own
process — a wedged worker looked exactly like a healthy one. The upgrade check
already ran at boot and was thrown into a log line. Version skew between
osctrl-api and osctrl-tls was observable nowhere.
Costs nothing when off
--health-enabled / SERVICE_HEALTH_ENABLED, default false, declared in
initServiceFlags so both services take it. With the flag unset: the manager is
never constructed, so AutoMigrate never runs and service_status is never
*created*; no route is registered; no heartbeat is written; no goroutine starts;
features.health is false and the SPA hides the section. Residual cost is one
bool in the flag struct. The final review traced all five of those through both
main.go files.
How it collects
Nothing expensive on a timer, nothing cross-process on a request:
| Source | When |
| ------------------------------------------------------------------ | --------------------------- |
| DB degraded flag + live ping, Redis PING, osctrl-api's own runtime | on request |
| osctrl-tls: uptime, goroutines, worker counters, runtime stats | heartbeat, every 60s |
| Upgrade check (external HTTP) | cached, refreshed every 24h |
runtime.ReadMemStats stops the world, so it runs only on the API request path.
osctrl-tls samples the same numbers through runtime/metrics, which reads
the same counters without a global pause. The heartbeat rides the existing 5s
service_commands poll loop, firing every 12th tick — no new goroutine, no
new ticker: ~1,440 writes/day against the 17,280 selects that loop already
issues, into a table that holds one upserted row per service forever.
Endpoint and page
GET /api/v1/health/status — admin only, 503 when the flag is off, one response
carrying every component so the page's drill-downs expand rather than refetch.
Statuses are `operational | degraded | down | stale | unknown`; a missing
heartbeat is unknown, never down ("not reporting — enable --health-enabled
on osctrl-tls"), since a service that was never enabled is not a dead one.
Version skew between services surfaces as degraded, naming both sides.
The page (/_app/health, nav gated on features.health) mirrors the three
sections: component list with expandable details, upgrade status, and a runtime
card per reporting service. Each card is labelled live (osctrl-api, read
while serving) or as of HHMMSS (osctrl-tls, up to a heartbeat old) so
nobody debugs a memory spike against stale numbers.
Two defects the whole-branch review caught
• The endpoint could hang during the outage it exists to report. Its DB and
Redis checks had no deadline — r.Context() carries none — so a saturated
pool blocked instead of returning down. Now bounded by a 3s budget threaded
through all three checks, via new context-aware Manager.GetContext and
RedisManager.CheckContext (existing callers unchanged).
• WorkersComponent blamed osctrl-tls for database outages. Manager.Get
returns ErrNotReporting only for a missing row; any other DB error rendered
as "osctrl-tls is not reporting", pointing operators at a healthy service. It
now mirrors `ServiceComponent`'s branch and surfaces the real error.
Known limitations (documented in ARCHITECTURE.md)
• Needs the flag on both services for a complete picture; api-only shows tls
as not reporting.
• Multiple osctrl-tls replicas share one row (last writer wins). Liveness stays
correct; per-replica detail is lost. Upgrade path is a (service, instance_id)
key plus pruning.
• LastGC and Lookups have no runtime/metrics equivalent and stay zero for
osctrl-tls; PauseTotalNs there is a histogram approximation. All documented
at the sampler.
• Snapshot only — no history.
Files
37 files, +1,976/−6. New: pkg/health (model/manager, two runtime samplers,
status derivation, version cache), cmd/tls/health.go, cmd/api/handlers/health_status.go,
frontend/src/{api/health.ts,features/health/}. Modified: both `main.go`s,
pkg/config (flag), pkg/cache + pkg/alerts (context/accessor additions),
SPA router/nav/features, README, ARCHITECTURE, sample configs, regenerated
osctrl-api.yaml.
Verification
go build ./... clean; full go test ./... green; frontend 329 tests pass;
tsc --noEmit clean; make openapi-check exits 0. Every task was
spec-and-quality reviewed, plus a whole-branch review whose two findings are
fixed above.
Not verified: the runtime smoke test — nobody booted the services against a
fresh database to confirm the table is absent with the flag off, or watched a
heartbeat flip to stale after stopping osctrl-tls. Both properties were traced
statically and hold.
jmpsec/osctrlGitHub
09/07/2026, 6:35 AM