The client says, "It's urgent."
The technician thinks, "Urgent for whom?"
The clock is running… and nobody agreed on what a response means.
That's the problem a useful SLA can prevent.
A service-level agreement isn't there to promise that everything will be fixed in minutes. It defines what you support, when the clock starts, who responds, how work escalates, and what evidence the client receives.
This guide gives a small MSP an operational starting point. It doesn't replace legal review or the commercial terms in your contract; use it to structure the service before turning it into a legal commitment.
1) Define what the SLA controls
An SLA isn't a phrase like "fast support."
You can't measure that.
The NIST glossary describes an SLA as a commitment between a provider and customer that can address responsibilities, service type, expected performance, response times, reporting, resolution, and termination.
For an MSP, that becomes six questions:
- which services and endpoints are covered?;
- during what hours does the commitment run?;
- which channel creates a valid request?;
- how is priority assigned?;
- when should the client receive a response and updates?;
- what is excluded or quoted separately?
If one of those answers exists only in your head, you don't have an SLA yet. You have a risky expectation.
2) Separate response, update, and resolution
Three clocks. Three different things.
First-response time: when a person acknowledges the request, assigns an owner, and communicates the next step.
Update cadence: how often you communicate while the incident remains open.
Restoration or resolution target: when you expect to restore service or close the cause, depending on scope and external dependencies.
PeopleCert's official ITIL Service Level Management practice explains that targets should be business-based and that agreed service levels should be monitored, reported, and improved.
That's why you shouldn't promise "resolution in 30 minutes" when you might depend on an ISP, replacement hardware, Microsoft, a vendor, or client approval.
3) Assign priority with impact and urgency
Everything is urgent when there's no matrix.
A useful priority combines:
- impact: how many people, locations, or business processes are affected;
- urgency: how long the operation can wait before the consequence becomes serious.
Start with this table, then adjust it to your real capacity:
| Priority | Example | First response | Updates | Starting target |
|---|---|---|---|---|
| P1 Critical | core operation stopped, multiple users unable to work, or suspected active security incident | 30 minutes during coverage | every 60 minutes | work continuously until restored or contained |
| P2 High | important function degraded, several users affected, limited workaround available | 2 business hours | at least once per business day | restore within 1 business day when controlled by the MSP |
| P3 Normal | individual incident with a workable alternative | 4 business hours | when status changes | resolve within 2 business days as a target |
| P4 Request | access, change, question, or improvement without interruption | 1 business day | when the schedule is agreed | schedule within 3 to 5 business days |
P1 Critical
core operation stopped, multiple users unable to work, or suspected active security incident
30 minutes during coverage
every 60 minutes
work continuously until restored or contained
P2 High
important function degraded, several users affected, limited workaround available
2 business hours
at least once per business day
restore within 1 business day when controlled by the MSP
P3 Normal
individual incident with a workable alternative
4 business hours
when status changes
resolve within 2 business days as a target
P4 Request
access, change, question, or improvement without interruption
1 business day
when the schedule is agreed
schedule within 3 to 5 business days
Don't copy these targets if you can't sustain them. Measure your ticket volume, technician capacity, and sold coverage first.
4) Write when the clock starts, pauses, and stops
The client sends a WhatsApp message at 11:48 p.m.
Has the SLA started?
The answer needs to be written before that happens.
Define at least:
- time zone and support hours;
- holidays and maintenance windows;
- the channel that creates a valid ticket;
- minimum data: client, user, endpoint, symptom, and impact;
- whether P1 time runs outside normal hours;
- when time pauses for access, approval, information, or client response;
- what counts as restored, resolved, or closed;
- how many contact attempts happen before closing for no response.
A pause shouldn't become a trick for hiding bad metrics. It needs a visible reason, timestamp, and evidence.
5) Define channels, ownership, and escalation
A strong SLA doesn't only say "how fast."
It also says "who's next."
Simple example:
| Moment | Owner | Action |
|---|---|---|
| Intake | service desk or on-call technician | validate client, scope, impact, and priority |
| Assignment | responsible technician | acknowledge and communicate the next action |
| Technical escalation | lead or specialist | step in when access, knowledge, or technical authority is missing |
| Commercial escalation | account owner | resolve scope, authorization, cost, or expectation |
| Executive communication | MSP owner or service lead | communicate P1 incidents and required decisions |
Intake
service desk or on-call technician
validate client, scope, impact, and priority
Assignment
responsible technician
acknowledge and communicate the next action
Technical escalation
lead or specialist
step in when access, knowledge, or technical authority is missing
Commercial escalation
account owner
resolve scope, authorization, cost, or expectation
Executive communication
MSP owner or service lead
communicate P1 incidents and required decisions
For every P1, also define whom the MSP calls on the client's side. A critical incident without an authorized contact gets stuck right when it matters most.
6) Protect scope with exclusions and responsibilities
The SLA shouldn't turn your monthly package into infinite work.
Clarify what is included and what needs a separate quote, approval, or schedule:
- unregistered endpoints, users, and locations;
- projects, migrations, cabling, and major changes;
- after-hours support;
- third-party failures, even if you coordinate follow-up;
- hardware, licenses, and replacement parts;
- recovery when no valid backup exists;
- incidents caused by unauthorized access or changes;
- onsite work and travel expenses.
The client also needs responsibilities: keep contacts current, provide authorized access, answer information requests, approve changes, and report through the agreed channel.
7) Use this template as a starting point
Adapt this block for your proposal or operational appendix:
```text SERVICE-LEVEL AGREEMENT — OPERATIONAL BASELINE
Coverage:
- Included services: [list]
- Endpoints/users/locations: [quantity and scope]
- Hours: [days, times, and time zone]
- Valid channel: [portal or email]
Priorities:
- P1 Critical: [definition] — response [time] — update [frequency]
- P2 High: [definition] — response [time] — update [frequency]
- P3 Normal: [definition] — response [time] — target [time]
- P4 Request: [definition] — response [time] — scheduling [time]
Clock rules:
- Starts when [condition].
- Pauses when [conditions and evidence].
- Stops when [restoration, resolution, or closure].
Escalation:
- Initial owner: [role]
- Technical escalation: [role and condition]
- Commercial escalation: [role and condition]
- Authorized client contact: [role]
Exclusions:
- [projects, third parties, after-hours, hardware, and others]
Monthly reporting:
- first-response compliance;
- tickets by priority;
- restoration/resolution times;
- reopened or aging tickets;
- client decisions still pending.
```
Before signing it, review capacity, dependencies, remedies, liability limits, and applicable law with qualified professional advice.
8) Measure the SLA and review it every month
An SLA sitting in a PDF doesn't improve service.
A monthly review can.
Measure at least:
- percentage of tickets with first response inside target;
- median first-response time by priority;
- restoration or resolution time;
- open tickets by age;
- reopened tickets;
- recurring incidents;
- time waiting for a client decision or response;
- leading causes of missed targets.
CISA recommends that even small teams document and practice an incident response plan. For an MSP, a simple P1 rehearsal can verify contacts, escalation, and communication before the real incident.
Don't use the report to hide failures. Use it to decide whether to adjust capacity, redefine scope, automate a task, or correct a promise.
FAQ: SLAs for MSPs
Does an SLA mean guaranteed resolution time?
Not necessarily. It can cover first response, updates, availability, and restoration or resolution targets. Separate every commitment and its dependencies.
Does a small MSP need four priorities?
Four is a useful starting point, but three can work for a simple operation. What matters is giving each level clear criteria and examples.
Should I offer 24/7 P1 support?
Only if you have the coverage, rotation, and price to sustain it. Otherwise, define business hours and the after-hours process without selling availability that doesn't exist.
Does this template replace a contract?
No. It's an operational baseline. A legal professional should adapt obligations, remedies, privacy, liability, and jurisdiction.
How does Lunixar RMM help with an SLA?
An RMM helps centralize operational context such as endpoints, alerts, patches, tickets, remote support, and reports. Compliance still depends on your scope, people, process, and capacity.
An SLA should protect the client and the MSP
The client needs to know what to expect.
Your team needs to know what to do.
And your MSP needs commitments it can actually deliver.
Lunixar RMM can support that operation by connecting visibility, alerts, support, and evidence in one console. Explore Lunixar RMM for MSPs or start a 2-week free trial with no credit card, up to 5 devices, and full access.
Keep reading:












