Every platform that carries revenue eventually forces one question: who is responsible when it fails at 03:00? Most companies answer it by default rather than by decision. The engineers who built the system keep it running, an on-call rotation forms informally, and operations becomes something the product team absorbs. That arrangement holds until it becomes expensive.
This article compares the two deliberate answers: building an internal operations team, or contracting managed IT services under an SLA. We sell the second, so read with that in mind. We have tried to keep the comparison honest — there are cases where internal is clearly the right call, and we start with those.
When an internal team is the right call
If operations is part of what you sell, own it. A company whose product is infrastructure — a hosting platform, a payments processor, a trading system — competes on operational capability, and outsourcing it would mean outsourcing a differentiator. The same applies when operational knowledge is genuinely proprietary: specialised hardware, hard latency constraints, or regulated environments where day-to-day control must stay in-house.
Internal also wins when you have the scale to do it properly. A team of eight or more senior engineers can staff a sustainable rotation, build its own tooling, and keep incident knowledge inside the company. If you can hire and retain those people, an internal team gives you the shortest possible loop between product decisions and operational reality. No contract matches that.
The honest test is not whether you could build an operations team. It is whether the money and management attention it consumes would produce more value there than on your product roadmap.
The real cost of an internal on-call rotation
The visible cost is headcount. Sustainable 24/7 coverage needs five to six engineers at minimum; with fewer, people are on call every second or third week, which is a resignation letter on a timer. At Western European senior rates, that rotation is a substantial six-figure annual commitment before anyone writes a line of product code.
The less visible costs accumulate around it:
None of this appears as a line item called operations. It is spread across salaries, recruiting fees, delayed releases and exit interviews — which is exactly why it is underestimated.
- On-call compensation, time off in lieu, and scheduling around holidays and sick leave
- Monitoring, paging and incident-management tooling, plus the time to maintain it
- Daytime productivity lost to interrupted sleep and context switching
- Attrition: engineers hired to build products resent nights spent firefighting, and they leave
- Re-hiring: replacing a senior operations engineer takes months, and incident knowledge leaves with them
What an SLA actually transfers
A managed operations contract transfers a defined set of obligations, in writing. In our contracts that means 24/7 monitoring, four severity levels (P1–P4) with response and restoration targets per level, security patching, CI/CD upkeep and bug fixing — with financial consequences for us when we miss a target. We answer P1 incidents within the hour because the contract says we must, not because someone happens to be awake.
The economic effect is the conversion of an unpredictable internal cost — staffing, tooling, attrition, 03:00 escalations — into a fixed monthly amount with a defined scope. The rotation math, the paging discipline and the post-mortem process become the provider’s problem. That is transfer of responsibility in a precise sense: measurable obligations, backed by penalties, held by a party whose business depends on meeting them.
What an SLA does not transfer
Accountability to your customers stays with you. When your platform is down, your clients call you, not your provider. An SLA can guarantee response and restoration; it cannot make your architecture decisions for you, and it does not absorb your regulatory position — under GDPR you typically remain the data controller, with the provider acting as processor under a data processing agreement.
A serious contract also demands things from you. The provider needs real access, current documentation, a named technical counterpart and the authority to change what it operates. An SLA over infrastructure nobody is allowed to touch is theatre. If your organisation cannot delegate that authority, the contract will underperform — and that failure will be yours as much as the provider’s.
Decision criteria
There is no formula, but five questions settle most cases:
If operations is a differentiator and you have the scale, build internal. If it is a dependency and the honest answers to the other questions are uncomfortable, transfer it — under a contract that makes every obligation explicit.
- Scale: can you staff five to six engineers on a rotation without starving the roadmap?
- Differentiation: do your customers pay you for operational capability, or is it a dependency?
- Cost of downtime: what does one hour of unplanned outage cost, in revenue and contract penalties?
- Hiring reality: can you attract and retain senior operations engineers in your market this year?
- Velocity: how much of your engineers’ week currently disappears into firefighting?
Where the line usually falls
Most companies we work with land on a division of labour rather than a wholesale choice. Their engineers own the product, the architecture direction and the decisions that differentiate the business. The provider owns the run: monitoring, incident response, patching and deployment discipline, governed by the SLA. The boundary is written down, reviewed regularly, and adjustable as the platform evolves.
Whichever way you decide, decide deliberately — and keep it reversible. Require documentation as a deliverable, an exit clause and a defined handover procedure from any provider, including us. The one model we advise against is the default: operations absorbed informally by the product team, unbudgeted and unowned. It is the most expensive option of all, and its costs are the hardest to see.