
Writing an IT SLA that works
Most IT service level agreements measure things that are easy to measure rather than things that matter. The result is a monthly report showing every target met, alongside a business that is not happy with the service. When that gap appears, the agreement is measuring the wrong things.
A useful SLA measures what the customer of the service actually experiences, and it is honest about what is outside the provider's control.
The metrics that look good and mean little
Response time alone. An automated acknowledgement within fifteen minutes satisfies a response target and does nothing for the person waiting. Response time is worth measuring only alongside a resolution measure; on its own it rewards acknowledgement rather than progress.
Availability as a single percentage. Ninety-nine point nine percent uptime sounds precise. It permits roughly forty-three minutes of downtime a month, and it matters enormously whether that lands at three in the morning or during month-end close. A single number cannot distinguish them.
Ticket volume. Falling volume can mean the estate is more stable, or it can mean people have stopped reporting problems because reporting them achieves nothing. It measures the wrong direction of causation and is easily influenced.
First-call resolution, unqualified. Genuinely useful when it reflects competent first-line staff. Easily inflated by closing tickets that were not resolved, or by categorising anything complicated as a new request.
Average resolution time. Averages hide the cases that damage the relationship. One ticket open for three weeks disappears into an average with two hundred quick ones. Percentiles are more honest.
The metrics worth agreeing
Time to resolution, by priority, measured at a percentile
Not the average. Something like "ninety percent of priority two incidents resolved within eight working hours". The percentile makes the long tail visible, which is where dissatisfaction actually comes from.
Agree the priority definitions carefully, because everything depends on them. Priority should be a function of business impact and how many people are affected, defined with examples, not left to whoever raises the ticket.
Availability of named services, in business hours
Measure the services the business cares about (the finance system, email, the network at a specific site) rather than infrastructure components. Weight the measurement to when it matters. Downtime during business hours should count differently from downtime at two on a Sunday morning.
Time to restore after a major incident
The number the business actually cares about after something significant breaks. Worth stating separately because it is a different capability from routine ticket resolution.
Backlog age
The oldest open ticket, and the count of tickets older than an agreed threshold. This is the single most effective measure for surfacing the work that quietly never gets done. Resolution-time targets can be met consistently while a backlog of difficult tickets grows behind them.
Change success rate
The proportion of changes completed without causing an incident or requiring rollback. This measures operational discipline better than almost anything else, and it is hard to manipulate.
Something the users say
A short satisfaction measure attached to closed tickets. It is imperfect and it responds to things outside the provider's control. It is also the only metric that reflects the experience rather than the process, and a persistent divergence between good numbers and poor satisfaction is worth investigating rather than explaining away.
Write the exclusions honestly
An agreement without clear exclusions produces arguments at every review. Be specific about:
- Third-party dependencies. If resolution requires a software vendor or a telecom provider, the clock should reflect that. Agree how it is measured and, importantly, that the provider is still responsible for chasing.
- Scheduled maintenance, with how much notice is required.
- Anything out of scope: unsupported software, personal devices, systems the customer administers themselves.
- Customer-caused delays. Waiting for information, waiting for approval, waiting for site access. The clock should pause, and the report should show how much time was spent paused, because a large number there is itself a finding.
That last one is frequently abused in both directions. Making the paused time visible keeps it honest.
Set targets you can meet
Targets set too aggressively during a sales process produce one of two outcomes: consistent failure, which poisons the relationship, or gaming, which is worse because it is invisible.
Better practice is to measure for a period before committing. Run the service, gather three months of data, then set targets at a level that is demanding but achievable, with an agreed improvement trajectory.
A provider who proposes this rather than accepting whatever targets you ask for is usually one worth having.
Make the review the point
The agreement is not the important artefact. The monthly or quarterly review is.
A useful review spends most of its time on exceptions rather than on the metrics that were met. What breached, why, what is being done. What is in the backlog and why it is stuck. What changed in the estate. What the provider needs from the customer that they are not getting.
That last item matters. Most service failures have a shared cause: access that was not granted, information that was not provided, a decision that was not made. A review where only one side is accountable produces a defensive relationship and worse service.
The test of a good agreement
The honest test is whether the report and the business's opinion of the service move together.
When every target is green and the business is unhappy, the agreement is measuring the wrong things and should be changed. That conversation is uncomfortable and it is far more productive than continuing to report against metrics that nobody believes.
Want this looked at in your own environment?
Talk to an expert →Keep reading


