A page goes out once
The database refuses a second copy of the same page to the same person for the same step. A job that runs twice finds the page already queued and moves on.
Reliability
A pager is only worth anything when your own systems are failing. So Transmit keeps its timers in the database, runs somewhere your apps probably don’t, and is watched from outside its own cloud.
Every escalation step and every page is a job written in the same database transaction as the change that caused it. If the alert is stored, its first page is stored with it. Nothing waits only in memory, so a crash or a deploy loses nothing.
The database refuses a second copy of the same page to the same person for the same step. A job that runs twice finds the page already queued and moves on.
A step that is due runs immediately rather than waiting for a scheduler pass; a test fails the build if that ever changes.
The worker that fires escalation timers runs on an instance that is never scaled down, so a step due at 3:02 pages at 3:02, not whenever the next request wakes something up.
A small program on a different provider from Transmit itself checks the API, the worker and the dashboard every minute and emails the team if any is down. Transmit watches the watchdog back with one of its own heartbeats. Uptime checks also run from four regions.
Anything that does arithmetic on time (rotations, restrictions, daylight-saving changes, escalation offsets, heartbeat deadlines) is checked with property tests at ten thousand cases each, before every deploy. Everything between an alert and the email provider runs for real against Postgres in tests.
We would rather say so here than have you find out at 3am.
Transmit is in early access. Tell us about your on-call setup, and we will help you import it and check it against Opsgenie before you switch.