Strategy

It Broke at 2am: Who Do You Actually Call?

SKIMBOX Team

Most maintenance contracts describe what gets done, not what happens when something stops working at the worst possible moment. Here is how to define severity, response and escalation before you need them.

It Broke at 2am: Who Do You Actually Call?

At two in the morning your booking system stops taking payments. Somebody notices at seven, when the first complaint arrives. By nine you have found a phone number, left a message, and sent an email to an address that was set up three years ago.

The support gets sorted eventually. What made the night expensive was not the fault. It was that nobody had ever agreed what happens when there is one.

Most maintenance contracts describe what work gets done: updates, backups, patches, monitoring. Our guides on website maintenance and app maintenance cover that scope and what it costs. This article is about the other half, which is frequently absent from the same agreement: what happens when something breaks unexpectedly, and who does what.

The four questions to settle before you need them

If you take nothing else from this article, settle these four with your supplier while everything is working.

Who, on what channel, in what hours? Not a company name. A route that reaches a person, and the hours during which that route is actually staffed. A surprising number of businesses discover during their first serious incident that the only real path to their supplier is one individual's mobile number, and that individual is on a flight.

What is committed, at what severity? A response time attached to each severity level, in writing. If the honest answer is that nothing is committed outside working hours, that is a legitimate arrangement and you should know you are in it rather than assume otherwise.

Who has the access needed to fix it? This one catches people. If the only person with production credentials is on leave, the response time in your contract is irrelevant. Ask how many people on their side can actually act, and make sure the answer is more than one.

What happens if we cannot reach you? The escalation path, with names and wait times.

None of these requires negotiation or additional spend to establish. They are questions with answers that either exist or do not, and finding out which is the entire point. A supplier who cannot answer all four in an email has told you the shape of your next incident.

Severity, agreed in advance

Without an agreed way to say how bad something is, every report you make is urgent to you and routine to them, and you end up negotiating priority during the incident itself. That is the worst possible time.

The useful framing comes from how incidents get prioritised in public sector guidance. US federal incident handling guidance recommends prioritising by the functional impact on business operations, the impact on information, and how recoverable the situation is [1].

Translated for a business, that means four questions. What has actually stopped working? How many customers or staff are affected? Is money or data at risk? And is it getting worse?

Notice that none of those asks about the technical cause. Severity is about business impact, and defining it that way avoids the conversation where a supplier explains that the underlying issue is minor while your customers cannot pay you.

Three or four levels is enough:

The service is down, or money cannot be taken. Nothing else matters until this is resolved.

A major function is broken with no workaround. People cannot do their jobs, but the business is still trading.

Something is broken and there is a workaround. Annoying, costly in aggregate, survivable for a day.

Everything else. Handled in the normal queue.

Resist adding more levels. Every additional boundary is something somebody will argue about at exactly the wrong moment.

Agree also that the initial classification is yours. A supplier deciding that your incident is not urgent is an argument you cannot afford to be having while it is happening, and genuine misclassification can be corrected calmly afterwards.

Response is not resolution

These get conflated in contracts and they are different commitments.

Response means somebody has acknowledged it and started work. Resolution means it is fixed.

Suppliers commit to response times because they control them. They are cautious about resolution times because nobody can know in advance how difficult a fault will be. That caution is fair, and you should not read it as evasion.

What you can reasonably ask for instead is updates at agreed intervals until it is resolved. Every thirty minutes on a severe incident, hourly on a moderate one. That is the thing that actually reduces the pain of an outage, because most of the distress in an incident comes from not knowing what is happening.

Be prepared for an honest answer about what your current arrangement includes. For a great many businesses the truthful response is that nothing is committed outside working hours, because nothing was ever paid for.

Work out what an hour costs

This single number decides most of the questions in this article, and almost nobody calculates it.

Take the revenue that flows through the affected system in a typical period and divide it down to an hourly figure. Add the cost of staff who cannot work. Add anything you owe customers when you fail them.

It will be rough. Rough is enough, because you are choosing between paying for standby cover and not, and those two options are usually far apart in cost.

For a business whose systems are used only in office hours, paying for overnight cover is generally waste. For anything taking orders around the clock, the arithmetic flips fast. The mistake is not choosing wrong. It is never doing the calculation and then being surprised by both the outage and the invoice.

Out-of-hours cover costs more because it requires somebody available who would otherwise be asleep. When buying it, ask specifically whether you are getting a person on standby or best-effort availability. The two read almost identically in a contract and behave completely differently at 2am on a public holiday.

The escalation path

The common failure is not a supplier refusing to help. It is nobody answering and no agreed next step.

So write the sequence down: the support channel, then a named account manager, then a named director, with a wait time at each step. Thirty minutes at the first level for a severe incident, two hours for a moderate one.

Agree the wait times in advance rather than judging in the moment, because judgement under pressure reliably errs toward waiting too long. When the trigger is defined, escalating becomes a procedure you are following rather than a decision you have to justify to somebody who might be asleep.

For severe incidents, get a phone number and confirm it is answered by a person rather than a mailbox. Email and ticket systems are right for most requests and are checked on a schedule. An incident costing money every hour needs a channel that interrupts somebody.

The access problem nobody checks

A response time is worthless if the person who responds cannot actually fix anything.

This is the most common hidden failure in support arrangements and it takes one question to surface: how many people on your supplier's side hold the credentials needed to act on a production problem? If the answer is one, your effective response time is that individual's availability, whatever the contract says. Holidays, illness and resignations all become your outage.

The mirror of that question applies to you. Who on your side can authorise emergency spending at 2am? If the only person who can approve an unplanned invoice is unreachable overnight, a supplier willing to work may end up waiting for permission. Agree a threshold in advance, something like: for a severe incident, work proceeds and the commercial conversation happens afterwards.

There is a third version worth checking, which is access to the accounts themselves. If the hosting account, the domain registrar or the payment gateway is only reachable through credentials your supplier holds, then an incident where you and that supplier are out of contact becomes unrecoverable rather than merely slow. Our guide on the accounts your business must own covers what should sit in your name and why.

All three are questions with concrete answers, and all three are much cheaper to ask on a quiet Tuesday than to discover at two in the morning.

Test it, because nobody does

Once a quarter, use the emergency route outside normal hours with a genuine but low-severity issue. Tell the supplier afterwards that it was a test.

You will find out whether the number still works, whether the named contact still works there, whether the inbox is monitored, and how long the whole thing actually takes. Every one of those has a way of degrading quietly between incidents, and finding out during a real one is the expensive version.

This is the single cheapest useful thing in this article and virtually nobody does it.

Your side of the arrangement

Half of a good incident response is internal, and it is entirely within your control.

Write one page covering who to contact, on which number, in which hours, who to escalate to and after how long, and what your team should gather before calling.

Store it somewhere that does not depend on the system being up. That detail sounds obvious and is learned the hard way with impressive regularity: the incident procedure lives on the intranet, and the intranet is what is down.

On what to gather first: what is broken in plain terms, when it started, whether anything changed recently, how many people are affected, what error appears and where, and whether it affects everyone or only some users. Five minutes assembling that saves considerably more, because it lets a supplier start working rather than start questioning.

That "whether anything changed recently" item is worth emphasising. Recent change is the first thing any competent responder checks, because it is the most common cause by a wide margin: a deployment, a configuration change, a certificate expiring, a third-party update. Keeping even a simple record of what changed and when pays for itself the first time you have an incident.

Find out before your customers do

Basic uptime monitoring that checks your site every few minutes and alerts somebody costs very little and is among the highest-value things you can add.

Being told by a customer that your site is down means it has probably been down for a while. The first question any incident review asks is how long it took anybody to notice, and "a customer told us" is a poor answer.

One refinement matters more than the rest: monitor the transaction, not just the page. A homepage returning a healthy response while checkout silently fails is the most common blind spot in basic monitoring. Watch the actual business action, a completed test order or a successful login, because a server answering tells you remarkably little about whether the business is working.

Backups you have never restored

A backup that has never been restored is an untested assumption sitting in a file.

Ask three questions: when was a restore last actually performed, how long did it take, and how much data would have been lost? If nobody can answer, that is your finding, and it is far better found now.

Note also the difference between a backup and a recovery plan. A backup is a copy of data. A recovery plan is a documented sequence somebody can follow to get a working system back: where things are, in what order they start, who holds which credentials. Most businesses have the first and not the second, which is why recovery so often takes days rather than hours. Our guide on business continuity covers that planning properly.

Afterwards

Hold a short review once things are stable. Half an hour, half a page.

What happened, why did it take as long as it did to notice and to fix, and what would prevent a repeat. Keep it about the system rather than about the person, because a review that assigns blame reliably produces less information the next time, and less information is the opposite of what you want.

The output should be two or three specific changes with owners. Monitoring on the thing that failed silently. A documented step that was missing. A dependency that needs upgrading. If a review produces no changes at all, either the incident was genuinely unpreventable or nobody asked hard enough questions, and the second is much more likely.

If it is a hosted product

You still need the answers even though you cannot change most of them.

What does the provider commit to? Where is their status page, found now rather than during an outage? How do you raise something urgent? And what does your business do while they fix it?

That last question is the one worth thinking about in advance, because it is the only part you control. Knowing that the platform itself is down is genuinely useful information, since it tells you to communicate with customers rather than keep diagnosing.

Where two suppliers each point at the other, our guide on suppliers blaming each other covers how to break that standoff with evidence.

What to do this week

Write the page. Who to call, on what number, in what hours, who to escalate to and after how long, what to gather first. Store it somewhere independent of the systems it covers.

Twenty minutes, no cost, and it is the highest-return preparation available to most businesses.

If you want help beyond that, defining severity levels, response expectations, an escalation path and a one-page procedure for your team starts from around AED 1,500 with us. Reviewing an arrangement you already have, including actually testing whether the routes work, sits in the same range. Final pricing depends on scope, and these are our own figures rather than a market survey.

References

  1. NIST Special Publication 800-61, incident response recommendations and considerations for cybersecurity risk management
  2. NIST, computer security incident handling guidance
  3. SKIMBOX, website maintenance and AMC in Dubai
  4. SKIMBOX, mobile app maintenance cost in Dubai
  5. SKIMBOX, business continuity and disaster recovery in the UAE
  6. SKIMBOX, when two suppliers blame each other

NIST guidance addresses cybersecurity incident response for US federal agencies. It is cited here for its prioritisation approach, which adapts well to ordinary operational incidents, rather than as a standard binding on private businesses in the UAE.

Frequently asked questions

  • What is the first thing to establish about support?

    Who to contact, through which channel, and what hours that channel is actually staffed. A surprising number of businesses discover during their first serious incident that the only route to their supplier is one person's mobile number, and that person is on a flight. Establish all of that while nothing is wrong, and write it somewhere your own team can find at 2am without needing access to the system that might be down.

  • Why does my maintenance contract not cover this?

    Because most maintenance agreements describe what work gets done, updates, backups, patches and monitoring, rather than what happens when something stops working unexpectedly. Those are different things. Our maintenance guides cover the scope and cost of ongoing work. This article covers incident response, which is a different thing and frequently absent from the same agreement entirely, because nobody thought to ask for it.

  • What is incident severity and why does it matter?

    It is an agreed way of saying how bad something is, so that both sides respond proportionately without arguing about it during a crisis. Without it, everything you report is urgent to you and routine to them. With it, a checkout being down and a typo on a page get treated differently by prior agreement rather than according to whoever happens to be most insistent at the time.

  • How should severity be defined?

    By business impact rather than by technical cause. US federal incident guidance recommends prioritising by the functional impact on business operations, the impact on information, and how recoverable the situation is. Adapted for an ordinary business, that means asking four things: what has actually stopped working, how many customers or staff are affected, whether money or data is at risk, and whether the situation is getting worse.

  • What severity levels should I use?

    Three or four is enough for most businesses. Something like: the service is down or money cannot be taken, a major function is broken with no workaround, something is broken but there is a workaround, and everything else. Resist the urge to add more levels than that, because every additional one creates another boundary that somebody will want to argue about at exactly the wrong moment.

  • Who decides the severity of an incident?

    You raise it and the supplier can propose a change, with the definitions settling most disagreements. Agree in advance that the initial classification is yours, because a supplier deciding your incident is not urgent is exactly the argument you cannot afford to be having at the time. Genuine misclassification can be corrected calmly after the event, when both sides can look at what actually happened without the pressure of an active outage.

  • What is a realistic response time?

    It depends entirely on what you are paying for, which is the point most businesses miss. Anything genuinely out of hours costs money because it requires somebody on standby. Ask what response time is actually included at your current price rather than assuming one exists, and be prepared for the honest answer to be that nothing at all is committed outside working hours.

  • What is the difference between response and resolution?

    Response is somebody acknowledging and beginning work. Resolution is the problem being fixed. Suppliers commit to response times because they control them, and are cautious about resolution times because they cannot know in advance how hard a fault will be. That distinction is fair rather than evasive. What you can reasonably ask for instead is updates at agreed intervals until it is resolved, because most of the distress in an outage comes from not knowing what is happening.

  • Should I pay for out-of-hours support?

    It depends on what an hour of downtime costs you. For a business whose systems are only used in office hours, paying for overnight cover is usually waste. For anything taking orders or serving customers around the clock, the arithmetic changes quickly. Work out the hourly cost of being down before deciding either way, because in most cases it makes the answer immediately obvious and removes the argument entirely.

  • How do I work out what downtime costs?

    Take your revenue through the affected system for a typical period and divide it down to an hourly figure, then add the staff cost of people who cannot work and any obligation you owe customers. It will be rough. Rough is entirely sufficient, because you are choosing between paying for standby cover and not paying for it, and those two options are usually very far apart in cost.

  • What is an escalation path?

    The named sequence of people to contact when the first person does not respond, with a stated wait at each step. First the support channel, then the account manager, then a director, with a time attached to each. It exists because the real failure mode is almost never a supplier refusing to help. It is nobody answering, and nobody having agreed what the next step should be.

  • How long should I wait before escalating?

    Agree it in advance rather than judging in the moment, because judgement under pressure tends toward waiting too long. Something like thirty minutes at the first level for a severe incident and two hours for a moderate one. The value lies in having a defined trigger, so that escalating becomes a procedure you are following rather than a decision you have to justify to somebody senior who might be asleep.

  • Should I have a phone number, not just email?

    For anything severe, yes, and confirm it is a number somebody actually answers rather than a mailbox. Email and ticket systems are entirely appropriate for most requests and are checked on a schedule rather than continuously. An incident costing money every hour needs a channel that actually interrupts somebody, and that channel should be tested periodically rather than assumed to work.

  • Should I test the support arrangement?

    Yes, and almost nobody does. Once a quarter, use the emergency channel outside normal hours with a low-severity issue, tell the supplier afterwards that it was a test, and see what happened. That single exercise reliably reveals the dead phone number, the contact who left the company and the unmonitored inbox, before a real incident reveals them for you at much greater cost.

  • What should my own team know?

    Who to contact, what counts as severe enough to use the emergency route, what information to gather first, and who internally can authorise spending to fix it. Put all of it on one page that does not live inside the system which might itself be down. That last detail sounds obvious and gets learned the hard way with impressive regularity.

  • What information should we gather before calling?

    What is broken in plain terms, when it started, whether anything changed recently, how many people are affected, what error appears and where, and whether it is happening for everyone or only some users. Five minutes spent assembling that information saves considerably more than five minutes of diagnosis later, because it lets a supplier begin working on the problem rather than begin questioning you about it.

  • Does a deployment usually cause it?

    Recent change is the first thing any competent responder checks, because it is the most common cause by a wide margin. Whether that is a deployment, a configuration change, a certificate expiring or a third-party update, knowing what changed in the previous day or two shortens diagnosis dramatically. Keeping even a simple written record of what changed and when pays for itself the very first time you have an incident, and costs almost nothing to maintain.

  • What if my supplier says it is not their fault?

    That may be true, and it does not tell you who fixes it. Separate the two questions: who resolves it now, and who pays for it afterwards. Where two suppliers each point at the other, our guide on suppliers blaming each other covers how to break that standoff with evidence rather than with argument, and how to separate the fix from the invoice.

  • Who is responsible if the hosting provider goes down?

    Commercially, that depends on your contracts, and practically your customers hold you responsible regardless. Large providers publish status pages and post-incident reports, so establish where yours is now rather than searching for it during an outage. Knowing that the platform itself is down is genuinely useful information, because it tells you to start communicating with customers rather than continue diagnosing something you cannot fix.

  • Should I have monitoring so I find out before customers do?

    Yes, and it is one of the cheapest useful things available. Basic uptime monitoring that checks a page every few minutes and alerts you costs very little. Being told by a customer that your site is down means it has probably been down for some time already, and the first question any incident review asks is how long it took anybody to notice.

  • What should be monitored beyond whether the site loads?

    The transaction that actually matters to the business. A homepage returning a healthy response while checkout silently fails is by far the most common blind spot in basic monitoring, and it is the one that costs real money. Monitor the actual business action, a completed test order or a successful login, rather than only whether the server responds, because a server answering tells you remarkably little about whether the business itself is working.

  • What is a post-incident review?

    A short session after things are stable, looking at what happened, why it took as long as it did to notice and fix, and what would prevent a repeat. Keep it about the system rather than about the person, because a review that assigns blame reliably produces less information next time. Half an hour and half a page is usually enough.

  • What should the review actually produce?

    Two or three specific changes with owners, not a narrative. Better monitoring on the thing that failed silently. A documented step that was missing. A dependency that needs upgrading. If a review produces no changes at all, either the incident was genuinely unpreventable or nobody asked hard enough questions, and in our experience the second explanation is considerably more likely.

  • How long should recovery from backups take?

    Whatever your supplier has actually demonstrated, which for most businesses is an unknown. A backup that has never been restored is an untested assumption. Ask when a restore was last performed, how long it took, and how much data would have been lost. If nobody can answer any of the three, that is itself the finding, and it is far better discovered now than during the incident where it actually matters.

  • What is the difference between a backup and a recovery plan?

    A backup is a copy of data. A recovery plan is a documented sequence somebody can follow to get a working system back, including where things are, in what order they start, and who has the credentials. Most businesses have the first and not the second, which is precisely why recovery so often takes days rather than the hours everybody assumed it would.

  • Do I need this if my system is a hosted product?

    You still need to know the answers even though you cannot change most of them. What does the provider commit to, where is their status page, how do you raise something urgent, and what happens to your business while they fix it. Your customers will hold you responsible for the outage regardless of who operates the platform, so the part worth planning is what your business does while somebody else fixes it.

  • What does out-of-hours cover typically involve?

    Somebody being reachable and available to work outside normal hours, which is why it costs more than daytime support. What varies is whether that means a person on standby or best-effort availability. Ask specifically which of those you are buying, because the two read almost identically in a contract and behave completely differently at two in the morning on a public holiday.

  • Is a retainer better than paying per incident?

    A retainer buys availability and priority, which is the thing you actually need in an incident, and you pay for it whether or not you use it. Paying per incident is cheaper when nothing happens and slower when something does, because you are negotiating while the system is down. Match the choice to the cost of an hour of downtime.

  • Can you help us set this up?

    We can. Defining severity levels, response expectations, an escalation path and a one-page incident procedure for your own team starts from around AED 1,500 with us. Reviewing an arrangement you already have, including actually testing whether the routes work, sits in the same range. Final pricing depends on scope, and these are our own figures rather than a market survey.

  • What should I do this week?

    Write one page: who to call, on which number, in what hours, who to escalate to and after how long, and what your team should gather before calling. Then store it somewhere that does not depend on the systems it covers being available. That page takes about twenty minutes to write, costs nothing, and is the highest-return preparation available to most businesses.

SKIMBOX Team

Tech Consultancy

Get fresh writing in your inbox

One email a fortnight. No filler.

By subscribing, you agree to our privacy policy.

Want us to build something?

We work with teams across MENA, UK, USA, and India to build products, run programs, and grow.

Get in touch

Continue reading