Designing on-call that people will accept
It is rarely the frequency of pages that makes people leave. It is being woken for something they are not permitted to fix, and then having no time in the following weeks to remove the cause, so the same page arrives again the next month. A quiet rota with no route to improvement wears people down more slowly but just as completely as a loud one, and both are visible to candidates within two questions.
What actually drives people off a rota?
Futility, then unpredictability, then volume, in that order. Being paged for a condition you cannot resolve, because the fix belongs to another team or requires access you do not have, is the fastest way to make on-call intolerable. The engineer is left awake, accountable and powerless, which is a combination nobody tolerates for long regardless of what it pays.
Unpredictability comes next. A rota where the shift dates move, where cover is arranged by asking in a group chat, or where the secondary is theoretical, means the person cannot plan a week of their life. That is a bigger imposition than the pages themselves, and it is entirely a management problem rather than a technical one.
Volume matters, but it is the one people will tolerate if the first two are handled, and only if the volume is falling. An engineer who is paged often but can see the count declining because their follow-up work is landing will stay. One paged rarely by the same recurring fault that nobody is allowed to fix will not.
How small can a rota be before it stops working?
Smaller than most guidance suggests and larger than most companies have. The constraint is not only the rotation interval but what happens when somebody is ill, on holiday or has left. A rota that becomes unworkable when one person is unavailable is not a rota, it is an individual with a backup plan.
| Rota shape | Realistic minimum | Works when | Fails when |
|---|---|---|---|
| Single primary, weekly rotation | Six people | Pages are rare and the service is well understood | Anyone leaves, and the interval collapses to every fourth week |
| Primary plus secondary, weekly | Eight people, across both tiers | Escalation is genuinely needed and the secondary is real | The secondary is nominal and the primary knows it |
| Follow the sun across time zones | Two teams of four in separated regions | You already have engineers in those regions for other reasons | Handover is verbal, so context is lost twice a day |
| Service teams on call for their own service | Four per team, with a platform escalation | Teams can deploy and roll back their own service unaided | Teams cannot access production, so every page escalates anyway |
| Business hours only, best effort overnight | Three people | Nothing genuinely breaks overnight and customers accept it | The informal expectation exists but is neither paid nor scheduled |
Should on-call be paid, and how?
Yes, separately from salary, and the mechanism matters more than the amount. Paying a standby allowance for carrying the pager plus a call-out payment for being woken creates the incentive you want: the organisation notices the cost of a noisy service, and the person bearing it is compensated for the specific imposition rather than told it is part of the job.
Time off in lieu is the common alternative and it works only if it is actually taken. In practice compensatory rest that has to be requested rarely gets used, because the person who was awake at three in the morning is the same person with a deadline that afternoon. If you offer it, schedule it automatically rather than making it a claim.
The arrangement to avoid is folding on-call into salary as an unstated expectation. It removes the signal that a noisy service costs money, it makes the burden invisible in planning, and it is the single most common thing candidates in this field are trying to escape when they start looking.
How do you make the follow-up actually happen?
Give the person who was paged the ticket, the authority to close it however they judge best, and protected time in the next sprint. The last part is the one that gets dropped, and dropping it converts every incident review into a list of things that will happen when there is time, which there never is.
Cap it so it is affordable. A standing allocation of a fixed share of the following sprint for incident follow-up is small enough to survive a planning meeting and large enough to remove several recurring pages a quarter. The alternative, where follow-up competes with feature work item by item, loses every time, because the feature has a stakeholder in the room and the alert does not.
Include deletion as a valid outcome. The correct resolution for a large share of pages is that the alert should not have existed, and the person who was woken by it is the best-placed person to say so. A review process that only accepts fixes as outcomes accumulates alerts forever, which is how rotas become intolerable without anyone deciding they should.
What should escalate and what should wait until morning?
Route by whether a human can act now and whether a user is affected now. Anything failing for customers with a known intervention goes to the pager. Anything that will become urgent but is not yet, such as a disk projected to fill in a week, goes to a ticket. Anything that has already been automatically remediated goes to a report that someone reads on Monday.
This requires a second, quieter destination that most setups lack. Without a low-urgency channel that is genuinely reviewed, every signal ends up either paging or being lost, and the pressure to keep things visible pushes them into the pager. A weekly triage of the low-urgency queue, with an owner, is what makes it safe to route things there.
The mechanism to insist on is that the person on call may downgrade an alert during their shift. Not silence it permanently, but move it out of the pager with a reason, for review afterwards. Giving that authority to the person experiencing the noise is what stops noise accumulating, and withholding it is why so many rotas get worse over time.
What will candidates ask, and which answer loses them?
They will ask how often the pager goes off, whether it is compensated, what happened to the last recurring incident, and whether they will be able to fix the causes. Experienced candidates ask all four, and they are assessing your engineering culture more accurately through these than through any technical discussion.
The answer that loses people is that it is quiet, honestly. Said without a number or an example it reads as either a lack of measurement or an evasion, and both are worse than a candid figure. A specific answer, including a bad month and what was done about it, is more persuasive than a reassuring one, because it demonstrates that somebody is counting.
The other losing answer is that on-call is shared informally and rarely used. Anyone who has done this work hears an unpaid expectation with no rota, no escalation path and no cover when someone is ill. If that is genuinely your situation, say so and say what you intend to change, because the alternative is that they discover it in month two.
Common questions
- What makes on-call unbearable for engineers?
- Being paged for problems they are not permitted or able to fix, more than the number of pages. An engineer woken for a condition owned by another team is awake, accountable and powerless, which nobody tolerates for long. Unpredictable scheduling comes second, and raw volume third. Volume is tolerable when it is falling, because that shows follow-up work is being done and the rota is improving.
- How many people do you need for an on-call rota?
- Around six for a single weekly primary rotation, and about eight if you want a genuine primary and secondary. The binding constraint is not the rotation interval but resilience: a rota that becomes unworkable when one person is ill, on holiday or leaves is not a rota. Smaller teams can run business hours cover with an honest, stated position on what happens overnight.
- Should on-call be paid separately?
- Yes, and paying a standby allowance for carrying the pager plus a call-out payment for actual interruptions works best. It compensates the specific imposition and makes the cost of a noisy service visible in planning. Time off in lieu only works if it is scheduled automatically rather than claimed, because the person who was awake overnight usually has a deadline the same day and never takes it.
- How do you stop the same incident recurring?
- Give the ticket to whoever was paged, give them the authority to decide the fix, and protect a fixed share of the next sprint for it. Follow-up that competes item by item with feature work loses every time, because the feature has someone in the room arguing for it. Accept alert deletion as a legitimate outcome, since a large share of pages should not have existed.
- What should page someone and what should wait?
- Page when a user is affected now and a human can act now. Route conditions that will become urgent later, such as a disk projected to fill next week, into a ticket queue that is genuinely reviewed on a schedule with a named owner. Anything already remediated automatically belongs in a weekly report. Without a reviewed low-urgency destination, everything drifts back into the pager.
- What do candidates ask about on-call in interviews?
- How often the pager fires, whether it is compensated, what happened after the last recurring incident, and whether they will have time to fix causes. Answering that it is quiet, without a figure or an example, reads as an evasion or as evidence that nobody is measuring. A specific answer including a bad month and the response to it is more convincing than a reassuring vague one.