Argus
Distributed multi-region service monitor: three checkers on three continents vote on whether your service is actually down before anything pages you.
- Required 2 of 3 regions to agree over a sliding 90-second window, so one checker on a bad network path cannot page anyone: a checker losing a third of its requests raised 0 false alarms.
- Serialized per-monitor evaluation with Postgres advisory locks inside a single transaction, chosen non-blocking so contention skips rather than queues, and the lock releases on commit rather than needing cleanup.
- Inserted a four-state model (up / degraded / down / recovering) with configurable time thresholds between raw verdicts and notifications, absorbing 199 brief status flips into 0 pages while still catching every real outage.
- Detected latency anomalies online with per-checker EWMA baselines and z-scores, distinguishing a service-wide slowdown (two of three paths) from one slow region, so a bad transit route is recorded rather than paged.
- Reconstructed uptime from an append-only transition log with half-open interval arithmetic, excluding scheduled maintenance and gaps where monitoring itself was down.
- Made alert delivery crash-safe with a transactional outbox and a pg-boss retry queue, keeping network I/O out of the locked transaction entirely.
- Load-tested the write path at 506 req/s with 0 errors and 81 ms p99, behind ~210 tests including integration tests against real Postgres via Testcontainers.
