

Website Maintenance SLA: Severity, Evidence, and Responsibility
Author
A website maintenance SLA is a contractual agreement that defines the expected level of service for keeping a website operational, secure, and up-to-date. It includes severity-based incident response, maintenance windows, backup and recovery commitments, escalation paths, exclusions, and monthly evidence reporting.
What Is a Website Maintenance SLA?
A Website Maintenance SLA (Service Level Agreement) is a formal contract between a website owner and a maintenance provider that defines the expected level of service for keeping the website operational, secure, and up-to-date.
It specifies response times, resolution targets, maintenance windows, backup schedules, and reporting obligations.
Unlike a generic support SLA, a website maintenance SLA focuses on the unique needs of a web presence: uptime, security patches, content updates, and recovery from failures.
The core components include: response and resolution times for incidents, scheduled maintenance windows, backup frequency and retention, and reporting requirements.
Illustrative adjustable assumption: Each component is tied to a measurable target, such as "respond to a critical incident within 30 minutes" or "perform daily backups with 30-day retention."
These targets are agreed upon based on the business impact of downtime and data loss.
A website maintenance SLA differs from a hosting SLA. A hosting SLA typically covers server uptime and network availability, while a maintenance SLA covers the application layer: the CMS, plugins, themes, and custom code.
For example, a hosting SLA might guarantee 99.9% server uptime, but a maintenance SLA would cover how quickly a plugin conflict is resolved after an update.
The Incident Severity Matrix: From P1 to P4
The incident severity matrix is a classification system used to prioritize issues based on their impact on the business. It typically has four levels: P1 (critical), P2 (high), P3 (medium), and P4 (low).
Each level has defined response and resolution targets, which are agreed upon in the SLA.
A P1 incident is a complete outage or a security breach that compromises sensitive data. For example, if the website is completely down, or if a hacker gains access to the admin panel, it is a P1.
Illustrative adjustable assumption: The SLA might require a response within 15 minutes and a resolution within 4 hours. These numbers are illustrative; the actual targets should be adjusted based on the business’s tolerance for downtime.
A P2 incident is a major functionality failure that affects a significant portion of users but does not take the site completely down. For instance, a broken checkout process on an e-commerce site would be a P2.
Illustrative adjustable assumption: The response time might be 30 minutes, with a resolution within 8 hours.
A P3 incident is a minor issue that affects a small number of users or has a workaround. An example is a misaligned image on a landing page. Illustrative adjustable assumption: The response time might be 4 hours, with a resolution within 2 business days.
A P4 incident is a cosmetic or low-priority issue, such as a typo or a broken link. The response time might be 1 business day, with a resolution within the next scheduled maintenance window.
The severity matrix is not static; it should be reviewed and updated as the website evolves. The SLA should also define who assigns the severity level and how disputes are resolved.
Maintenance Windows and Scheduled Downtime
Maintenance windows are predefined periods during which the provider can perform scheduled tasks that may cause downtime. These windows are typically chosen to minimize impact on users, such as late at night or on weekends.
The SLA should specify the exact hours and days for maintenance windows, as well as the expected duration.
Scheduled downtime is not counted against the uptime guarantee, but it must be communicated in advance.
Illustrative adjustable assumption: The SLA should require the provider to notify the website owner at least 24 hours before a maintenance window, unless it is an emergency.
This notification should include the expected duration and the reason for the maintenance.
During a maintenance window, the provider may perform tasks such as applying security patches, updating the CMS, or migrating the server. These tasks are necessary for the long-term health of the website, but they can cause temporary disruption.
The SLA should define what happens if the maintenance exceeds the scheduled window, such as extending the window or providing a credit.
Maintenance windows are not the same as incident response. If an incident occurs during a maintenance window, the severity matrix still applies.
For example, if a P1 incident happens during scheduled maintenance, the provider must respond immediately, not wait until the window ends.
The SLA should also define exclusions: what is not covered by the maintenance service. For example, changes made by the website owner without prior approval, or third-party integrations that are not supported, may be excluded.
Backup and Recovery Commitments
Backup and recovery commitments define how often backups are taken, how long they are retained, and how quickly the website can be restored after a failure. These commitments are essential for minimizing data loss and downtime.
Backup frequency is typically daily, but it can be more frequent for sites with frequent content updates. The SLA should specify the backup schedule, such as "daily at 2:00 AM UTC."
It should also specify the retention period, such as "30 days of daily backups, plus 12 monthly backups." These are illustrative; the actual numbers should be based on the business’s data recovery needs.
Recovery Time Objective (RTO) is the maximum acceptable time to restore the website after a failure. For example, an RTO of 4 hours means the website must be back online within 4 hours of a disaster.
Illustrative adjustable assumption: Recovery Point Objective (RPO) is the maximum acceptable data loss, such as "no more than 24 hours of data." These objectives are defined in the SLA and tested regularly.
The SLA should also define the backup storage location, such as off-site or cloud-based, and the security measures for protecting backups. It should specify who is responsible for testing backups and how often.
Regular backup testing is essential to ensure that backups are actually recoverable.
Evidence of backup and recovery is provided through reports. The SLA should require the provider to deliver monthly reports that include backup logs, test results, and any incidents.
Escalation Paths and Dependencies
Escalation paths determine how an incident moves up the chain when initial response does not resolve it.
A typical path starts with the first-line support engineer, then moves to a senior engineer, and finally to the provider’s management or the vendor’s support team. Each step should have a defined trigger, such as elapsed time or unresolved status.
For example, if a critical incident is not resolved within the agreed recovery target, it escalates to the next tier automatically.
Dependencies are third-party services that the website relies on, such as hosting providers, CDNs, domain registrars, or payment gateways.
The SLA must state that the provider’s responsibility is limited to coordinating with these parties, not guaranteeing their performance.
For instance, if the hosting provider has an outage, the maintenance provider is responsible for monitoring, communicating, and escalating to the host, but cannot be held liable for the host’s downtime.
The SLA should include a dependency matrix listing each third-party, the contact point, and the expected response time from that party.
Escalation also involves communication. The SLA should specify who is notified at each stage, via what channel (email, phone, ticket system), and within what timeframe.
For example, a critical incident may require a phone call to the client’s designated contact within 15 minutes, while a low-severity issue may only require a ticket update.
Exclusions and Limitations
Exclusions define what the SLA does not cover. Common exclusions include issues caused by user-generated content, such as a site owner uploading a corrupted file or a plugin conflict introduced by the client’s own changes.
Also excluded are force majeure events like natural disasters, war, or internet backbone failures. The SLA should list these explicitly to avoid ambiguity.
Another limitation is scheduled maintenance windows. The SLA typically excludes downtime during agreed maintenance periods, such as a weekly two-hour window for updates.
If the provider performs maintenance during this window and the site is temporarily unavailable, it does not count as an incident. However, the provider must notify the client in advance and ensure the maintenance is completed within the window.
Other exclusions may include issues arising from third-party services that are not under the provider’s control, such as a payment gateway outage. Also, security breaches caused by the client’s negligence, like sharing admin credentials, may be excluded.
The SLA should also state that it does not cover changes requested by the client that are outside the agreed scope, such as new feature development.
Monthly Evidence and Reporting
Monthly evidence is the proof that the SLA was met. The provider must collect data on incident response times, resolution times, uptime, and maintenance activities. This data is compiled into a monthly report that shows compliance against the agreed targets.
The report should include a summary of all incidents, their severity, response and resolution times, and any escalations that occurred.
Evidence collection methods include ticketing system logs, monitoring tools, and server logs. For example, the ticketing system records when an incident was reported and when the first response was made.
Monitoring tools track uptime and can generate automated reports. The SLA should specify which tools are used and how the data is stored. The provider should also keep records of all communications related to incidents, such as emails and phone call logs.
The monthly report should be delivered by a specific date, such as the fifth business day of the following month. It should be in a format that is easy to understand, with charts and tables.
The report should also include a section on any exclusions that were applied, such as scheduled maintenance downtime. If the provider failed to meet any targets, the report should explain why and what corrective actions are being taken.
Example: A Real-World SLA Clause
Here is a concrete example of an SLA clause that includes a severity matrix and evidence requirements. This is an illustrative example; adjust the numbers to fit your specific needs.
**Severity Matrix**
| Severity | Definition | Response Time | Resolution Time |
|———-|————|—————|—————–|
| Critical | Website is down or major functionality is broken, affecting all users. Illustrative adjustable assumption: | 15 minutes | 4 hours |
| High | Significant functionality is impaired, but the site is accessible. Illustrative adjustable assumption: | 30 minutes | 8 hours |
| Medium | Minor functionality issue, no major impact. Illustrative adjustable assumption: | 2 hours | 24 hours |
| Low | Cosmetic issue or question. Illustrative adjustable assumption: | 4 hours | 48 hours |
**Escalation**
Illustrative adjustable assumption: – If a Critical incident is not resolved within 2 hours, it escalates to the senior engineer. Illustrative adjustable assumption: – If not resolved within 4 hours, it escalates to the provider’s management.
Illustrative adjustable assumption: – The provider will notify the client’s designated contact via phone within 15 minutes of a Critical incident.
**Dependencies**
The provider is not responsible for downtime caused by the hosting provider, CDN, or domain registrar. The provider will coordinate with these parties and keep the client informed.
**Exclusions**
This SLA does not cover issues caused by user-generated content, client-initiated changes, or force majeure events. Scheduled maintenance windows are excluded from uptime calculations.
**Evidence and Reporting**
The provider will use a ticketing system and monitoring tools to record all incidents. A monthly report will be delivered by the fifth business day of the following month, including incident logs, response times, resolution times, and any escalations.
The report will also list any exclusions applied.
This clause provides a clear framework for accountability and performance measurement. By defining severity, escalation, exclusions, and evidence, both parties know what to expect and how to verify compliance.
Next step
Ready to define your own maintenance SLA? Contact us to discuss a tailored agreement that fits your website’s needs.
Related services and further reading
Official references and sources
Comments (0)
No comments yet. Be the first!