Keep the lights on (KTLO)
How a paused product stays supported without the pause dissolving. STOP only works if the squads doing remediation are actually left alone — and clients only accept the pause if production stays solid. KTLO is the small, named, ring-fenced capacity that makes both true at once.
The capacity rule of thumb
- Ring-fence roughly 10–15% of squad capacity — in practice, one named engineer pair per product (or small product cluster), rotating on a sensible cadence so knowledge spreads and nobody burns out on interrupts.
- Named, not notional — "someone will pick it up" is how STOP squads get raided one ticket at a time. The KTLO pair is on the rota; everyone else is off-limits.
- Capacity is a budget — if KTLO demand persistently exceeds the ring-fence, that's a signal to Leadership at the weekly review (the product may be less stable than assumed), never a licence to quietly pull remediation engineers across.
What interrupts the pause — and what queues
Severity per the rapid-fix SLAs decides:
- S1/S2 interrupt — regulatory-reporting failures, claims-flow outages, a broken quote-and-buy journey with no workaround. The KTLO pair drops queued work and responds to SLA; fixes still go through the normal gates.
- Regulatory-dated work counts as interrupt-class — an FCA-dated obligation doesn't queue because we're pausing; it's ring-fenced via the commitment triage honour disposition.
- S3/S4 queue — minor defects and cosmetic issues are logged, batched, and worked by the KTLO pair as capacity allows, or held for the resumed delivery flow. The client-facing line: logged, prioritised, not lost.
- Feature and enhancement requests never interrupt — they are not defects at any severity; they route to the client-commitment register as new entries for triage.
Routing
- One front door — everything arrives via support triage; nothing lands directly on an engineer, however friendly the relationship.
- Triage classifies first — defect (severity per the SLA matrix) or request (to commitment triage). That single fork is the whole control.
- KTLO pair only — classified defects go to the named pair; STOP squads are not in the routing table at all.
- Escalation is explicit — if an S1 genuinely exceeds the pair, pulling a STOP engineer is a logged Leadership decision with a return date, not a hallway grab.
The classic failure mode: config requests un-pausing the pause
The pause rarely dies by announcement — it dies by config. A rate-table refresh here, an MTA wording tweak there, a scheme toggle "that's just config, five minutes" — each one small, each one reasonable, and three months later the pause exists in name only while remediation crawls.
- The rule — if it changes what the product does for a client, it's a commitment, not a ticket. Config requests route through the client-commitment register and get a disposition like everything else; the support side door is for defects only.
- The tell — KTLO throughput trending up while the defect count doesn't. The weekly review watches for exactly this: it means requests are leaking in through the ticket queue relabelled as fixes.
- The exception is honest — some config is honoured during the pause (regulatory dates, binding commitments). The difference is that it arrives via a scored disposition with Leadership visibility, not by osmosis.
Standards referenced: ITIL 4.