How to Write a Website Incident Response Runbook for Website Failures
When a website failure happens, the first few minutes can feel disproportionately difficult. The homepage may be unavailable, checkout may be rejecting orders, a lead form may have stopped working or a third-party service may be returning errors. Meanwhile, the person who normally knows what to do is in a meeting, travelling or outside working hours.
A website incident response runbook gives your team a practical route through that uncertainty. It does not need to be a large technical manual or a promise of 24-hour coverage. It should be a clear, current document that explains how to identify the problem, assess its impact, contact the right people, protect customers and confirm recovery.
This guide is for UK business owners, ecommerce teams, digital leads and operations managers who rely on their website but do not have a dedicated round-the-clock operations team.
What a website incident response runbook should achieve
A runbook is an operational instruction set for a known situation. In the context of website failures, it should help the team answer five questions quickly:
- What appears to be wrong?
- How serious is the impact?
- Who owns the next decision?
- Who needs to be informed?
- What proves that the website has recovered?
The runbook is not a replacement for monitoring, technical documentation or a support agreement. It connects those things. A monitor may tell you that a checkout journey has failed, while the runbook tells you whether to investigate immediately, contact a supplier, pause advertising or use an agreed fallback.
Start with the failures that matter commercially
Do not begin by documenting every possible error code or hosting scenario. Start with the website journeys that affect revenue, enquiries, customer service or operational continuity.
Typical priority journeys include:
- Homepage and core navigation
- Product, category or service pages
- Search and filtering
- Basket and checkout
- Payment-provider handoff
- Lead, quote or contact forms
- Customer login or trade portal access
- Order confirmation and internal notifications
For each journey, write down what failure looks like from the customer’s perspective. A checkout problem might mean an error after payment, not only a page that fails to load. A form incident might mean the visitor sees a success message but no lead reaches the CRM or inbox.
Define simple incident severity levels
A small team needs a severity model that can be understood without a meeting. Three or four levels are normally enough. The labels matter less than the decision attached to each one.
Critical incident
A critical incident affects a core trading or customer journey for a significant proportion of users. Examples include checkout failure, payment errors, widespread 5xx responses, missing order creation or a serious data-handling concern.
The response should normally include immediate ownership, technical investigation, stakeholder notification and a decision about whether to pause campaigns, disable a feature or use a rollback route.
High-priority incident
A high-priority incident affects an important journey or customer group but may have a workaround or limited scope. Examples include a trade portal issue affecting one account tier, a major product feed problem or a lead form failing on mobile.
Standard incident
A standard incident is a genuine problem with limited commercial or operational impact. It may affect a low-traffic page, a non-critical integration or a small group of users. It should be recorded and assigned, but may wait for normal working hours.
Monitoring signal or investigation
This is an alert or report that needs checking but has not yet been confirmed as a customer-impacting failure. Avoid treating every failed automated check as an outage. Confirm the result using another browser, device, location or test route where appropriate.
VERIFY: severity thresholds should reflect your traffic, trading hours, support arrangements and the cost of a missed order or enquiry. Do not copy another business’s response targets without checking that they fit your operation.
Give every incident one clear owner
Shared responsibility often means nobody is sure who should act next. Each incident should have one named incident owner, even if several people contribute to the investigation.
The incident owner is responsible for keeping the response moving. They do not need to fix the technical fault themselves. Their responsibilities may include:
- Confirming the incident and recording the start time
- Assigning the severity level
- Contacting the technical owner or support partner
- Keeping internal stakeholders updated
- Recording decisions, evidence and next actions
- Confirming recovery before closing the incident
In a small business, the incident owner might be the ecommerce manager during working hours and a director or duty contact outside those hours. The important point is that the fallback route is written down rather than assumed.
Document the website support escalation route
Your runbook should show exactly how to escalate a website failure. Avoid relying on a general team inbox if a serious issue could sit there unnoticed.
Record the following for each important supplier, agency or technical contact:
- Organisation or team name
- Primary contact route
- Secondary contact route
- Systems or services they support
- Information they need when an incident is raised
- Expected acknowledgement or response terms, where agreed
- Internal owner for the relationship
VERIFY: supplier support hours, response targets and escalation terms should be checked against your current contracts. Do not describe a service as 24-hour support unless that is explicitly provided and understood.
A useful escalation message includes the affected URL or journey, the time first observed, the symptoms, the number of failed tests, the business impact and any relevant order, form or monitoring reference. This gives the recipient enough context to begin investigating without asking the same basic questions first.
Include a first-response incident checklist
The first response should be short enough to follow under pressure. A practical incident response checklist might be:
- Record the time the issue was reported or detected.
- Run a controlled test of the affected journey.
- Check whether the issue affects one page, one device, one account or the wider site.
- Assign a severity level and incident owner.
- Capture a screenshot, error message, URL and relevant reference.
- Check recent releases, configuration changes and supplier incidents.
- Escalate to the technical owner or support contact.
- Decide whether any traffic, promotion or feature should be paused.
- Set the next update time, even if there is no resolution yet.
Do not make multiple unrelated changes while diagnosing the issue. If possible, record each action and its result. Otherwise, the team may lose the evidence needed to identify the original cause.
Define customer and stakeholder communications
Communication is part of the website outage process, not an optional extra. Customers, sales teams, fulfilment staff and marketing managers may all need different information.
Customer-facing communication
Use clear, factual wording. Explain what customers should do next without speculating about the technical cause. For example, if orders cannot be completed, advise customers to retry later or contact the business through an available route. If a form is unavailable, provide a monitored email address or telephone number if one exists.
Internal communication
Internal updates should state what is affected, what is not yet known, who owns the response and when the next update will happen. This prevents sales, customer service and marketing teams from working from different versions of the situation.
Marketing and paid traffic decisions
If a failure affects a campaign landing page, checkout or lead route, decide whether Google Ads, social campaigns or email activity should be paused. The decision should be based on the affected journey, not only on whether the homepage is available.
HOFK’s guide to scoring competing Google Ads landing page variants covers pre-launch readiness. During an incident, the same principle applies in reverse: do not continue sending paid traffic to a route that cannot complete its intended action.
Write recovery checks for each critical journey
An incident is not resolved simply because the page loads again. Recovery checks should prove that the customer journey and the internal handoff are working.
For an ecommerce incident, check:
- Product or service pages load correctly
- The relevant item can be added to the basket
- Prices, stock and delivery information are plausible
- Checkout progresses to the payment step
- The payment or order handoff returns correctly
- The order appears in the expected back-office system
- Confirmation messages or emails are generated
For a lead-generation incident, submit a controlled test form and confirm the success state, CRM or inbox delivery, source information and follow-up notification. For a B2B portal, test login, account pricing and the relevant ordering or quote route.
HOFK’s article on synthetic uptime monitoring for critical website journeys explains why testing the complete journey is more useful than checking whether one URL responds.
Record what happened and review it afterwards
Once the incident is stable, capture a short record while the details are fresh. You do not need a lengthy post-mortem for every minor issue, but repeated incidents deserve more than a ticket marked closed.
Record:
- Incident title and severity
- Start and recovery times
- Affected journeys and users
- Detection source
- Key actions and decisions
- Technical or supplier cause, if confirmed
- Customer or commercial impact, where known
- Follow-up work and owner
Ask three practical questions: what detected the problem, what slowed the response, and what would make the next response easier? The outcome may be a code fix, a better monitor, clearer documentation, a support escalation change or a small process improvement.
HOFK’s article on prioritising a website support retainer backlog provides a useful approach for deciding which incident follow-ups deserve attention first.
Keep the runbook current
A website incident response runbook becomes unreliable when systems, contacts and responsibilities change without updating it. Review it after every significant incident, platform change, agency change or new critical journey.
At a minimum, review:
- Contact details and escalation routes
- Critical journeys and recovery tests
- Severity definitions
- Monitoring links and alert destinations
- Known workarounds and rollback instructions
- Internal and customer communication templates
Store the current version somewhere the relevant people can reach during an incident. Avoid keeping the only copy inside the website or application that may be unavailable.
Website incident response runbook template
A concise runbook can use this structure:
- Purpose: what the runbook covers.
- Critical journeys: what must be tested first.
- Severity levels: what makes an incident critical, high or standard.
- Roles: incident owner, technical owner, business owner and communicator.
- Escalation contacts: primary and fallback routes.
- First-response checklist: the immediate actions to take.
- Communication templates: internal and customer-facing wording.
- Recovery checks: evidence required before closure.
- Incident record: fields for timing, impact, actions and follow-up.
- Review date: when the runbook itself will be checked.
Where HOFK can help
Website incidents often cross responsive websites, ecommerce systems, integrations, forms, analytics, hosting and operational workflows. HOFK can help review critical journeys, improve monitoring, investigate recurring failures or support the full stack development behind a more maintainable website.
Relevant work may include synthetic checks, checkout and form monitoring, incident documentation, escalation design, ecommerce support, responsive improvements or automation around operational alerts. The aim is not to create unnecessary process. It is to give a small team enough visibility and structure to respond calmly when the website fails.
Conclusion
A website incident response runbook gives a business a shared method for handling website failures. Start with the journeys that matter commercially, define simple severity levels, name one incident owner, document website support escalation and include clear communication and recovery checks.
The runbook should be short enough to use under pressure and specific enough to prevent guesswork. Review it after incidents and platform changes so it reflects the website you operate today, not the one you had when the document was first written.
Frequently asked questions
What is a website incident response runbook?
It is a practical document that explains how to identify, assess, escalate, communicate and recover from website failures.
What should a website outage process include?
It should include severity levels, named owners, technical and supplier contacts, first-response actions, communications, recovery checks and post-incident follow-up.
Do small businesses need a 24-hour incident response team?
No. A small business can still benefit from a clear runbook, agreed escalation contacts and realistic support arrangements without claiming round-the-clock coverage.
How often should an incident response checklist be reviewed?
Review it after significant incidents, platform or supplier changes, changes to critical customer journeys and at a regular operational interval that suits the business.
When should website support escalation happen?
Escalate when a failure affects a critical journey, cannot be explained by a local test issue, exceeds the agreed severity threshold or requires access held by a technical supplier.