On-call arrangements tend to be designed for coverage and not for the people providing it. The result is predictable: the most experienced engineers quietly disengage, and the knowledge needed to handle incidents leaves with them.

Measure the actual load

Count how many times someone was woken, how long each incident took, and how many were genuinely urgent. Most organisations have never measured this, and the people carrying it are not in a strong position to complain about it.

Once visible, the number usually makes its own case for fixing the underlying noise.

Every page must be actionable

If the response to an alert at three in the morning is to acknowledge it and go back to sleep, it should not have paged. Move it to a dashboard or a next-morning queue.

Alerting on customer-visible symptoms rather than on internal conditions removes most of this noise in one change.

Give people a real runbook

A tired person at three in the morning should not be reasoning from first principles. A runbook per alert — what it means, how to confirm, what to try, when to escalate — is the difference between a fifteen-minute incident and a two-hour one.

Write it when the alert is created, not after the first bad night.

Compensate it and bound it

On-call is work, whether or not anything happens, because it constrains where someone can be and what they can do. Pay for it or give time back, and be explicit about the expectation.

Cap consecutive weeks, guarantee recovery time after a bad night, and make it genuinely acceptable to hand over when someone is exhausted.

Feed incidents back into the work

If the same alert fires every rotation and never gets fixed, the rotation is absorbing a problem rather than surfacing it. Reserve capacity in normal planning for reliability work generated by on-call.

Without that loop, on-call becomes a permanent tax rather than a temporary safety net.

A sustainable rotation and a reliable system are the same project. Anything that makes on-call quieter has already made the system better.

Written by the Global IT Solutions engineering team. Have a project this touches on?

Start a conversation