GuideManaged Services
Writing SLAs for business operations that AI helps run
An SLA for AI-assisted operations has to measure more than speed. When software does part of the work, the main risk shifts from slow output to confident, wrong output at volume. This guide covers choosing indicators per process, evidencing quality with sampled human review, objectives and error budgets, credits, and the clauses that govern model changes and human escalation.
On this page
- What an SLA can and cannot promise for AI-assisted work
- Indicators that matter for each family of operational process
- Designing the quality sample and the human QA behind it
- Setting objectives and error budgets for operational work
- Credits, earn-back and exclusions that change behavior
- Contract clauses for model, prompt and threshold changes
- Human-review commitments and escalation times
- Hypothetical SLA schedule for a finance-operations service
- Reporting cadence and the forums that read the reports
- Questions and answers
- Sources
What an SLA can and cannot promise for AI-assisted work
A service level agreement is a contractual promise about outcomes both parties can measure. In AI-assisted operations, many of those outcomes are decided by software: whether an invoice lands on the right ledger account, whether a claim reaches the right handler, whether a reply states the policy correctly. The agreement has to describe the quality of decisions, not only the time they take.
An SLA can reasonably promise processing times, limits on backlog age, the time a person takes to pick up an item the software has set aside, quality on a defined sample, notice before material changes and named escalation contacts. It cannot promise that no single output is ever wrong, that a model behaves identically after its vendor updates it, or that targets hold when inputs the client controls arrive late.
ColdAI describes the Operate phase of its managed services as continuous delivery by dedicated teams with SLA-backed performance1. This page explains how such an agreement should be shaped; it does not state ColdAI's own targets, which are set per process from a measured baseline.
Indicators that matter for each family of operational process
Choose two or three service level indicators per process: one for flow, one for quality and, where people review the software's work, one for the hand-off.
| Process family | Speed and flow | Quality | Exceptions and people |
|---|---|---|---|
| Finance operations (payables, reconciliations) | Cycle time from receipt to posting; age of the oldest unprocessed item | First-time-right coding and matching, checked on a stratified sample | Share of items held for a person; time to resolve a held invoice |
| Customer service | Time to first response and to resolution, by channel | Accuracy and policy compliance of answers, scored against a rubric | Escalations to human agents; reopened cases |
| Insurance claims support | Time to triage and to a settlement recommendation | Correct coverage and reserve recommendations on sampled files | Time for a licensed handler to pick up a referral; audit findings |
| Banking onboarding and KYC | Time to complete a case; age of the queue | Correct risk rating and a complete evidence pack on sampled cases | Cases passed to compliance; time to clear screening alerts |
| Trust and safety | Time from report to action | Precision of enforcement decisions on samples; appeals overturned | Items held for specialist review; limits on reviewer exposure |
Keep further measures in operational reporting rather than the contract; a long list dilutes attention and complicates credits.
Designing the quality sample and the human QA behind it
Quality is the indicator buyers most often leave vague. These steps turn it into something both sides can check.
Define what counts as right
Write a scoring rubric per output type with error classes: critical (a wrong payee, a wrong coverage decision, a reply that breaks policy), major and minor, with agreed examples of each.
Stratify the population
Sample separately from items the software completed alone, items a person touched and high-value or high-risk items. A healthy average can hide a weak segment, and the automated one changes fastest.
Set sample sizes and frequency
Size samples from volume and the precision you need, with larger samples for new processes and after changes. Acceptance-sampling schemes such as those in the ISO 2859 series are one established reference for this kind of plan3.
Keep reviewers independent and calibrated
Reviewers should not grade their own team's work, and the client should be able to draw its own blind sample. Calibrate reviewers on a shared set of items and record disagreements.
Log causes, not just counts
For every sampled error, record its class and root cause: the model, a rule, the input data or a person. Root causes decide the fix; a raw count only says something went wrong.
Agree what a breach triggers
Write down what a breached threshold triggers: expanded sampling, a root-cause review within an agreed time and, for critical errors, pausing the automated path until a fix is verified.
Setting objectives and error budgets for operational work
A service level objective is the internal target the provider runs to; the SLA is the commitment it is paid against. Running to a stricter objective gives the provider room to correct a decline before it becomes a breach.
An error budget turns that gap into a working rule. Site reliability engineering treats the budget as the unreliability a service may spend in a period: while budget remains, changes ship; when it is spent, effort moves to reliability2. In operations, express it as tolerated late items or critical errors per month. When it runs out, planned automation changes pause until quality recovers.
Set objectives from the baseline measured during transition rather than from a proposal, and revisit them after a full business cycle that includes month-end or the seasonal peak.
Credits, earn-back and exclusions that change behavior
- If
A missed level reaches your customers, suppliers or regulators directly, such as late supplier payments or missed complaint deadlines
ThenAttach service credits to that indicator and require a written remediation plan with dates
Credits signal priority; the plan is what changes the outcome.
- If
Misses are occasional and the provider recovers quickly
ThenAllow earn-back, so credits are returned when the level holds for an agreed run of following months
It rewards sustained recovery over punishing one bad month.
- If
A miss is caused by your own inputs, such as late source files, unannounced policy changes or outages in your systems
ThenExclude it, but only where the dependency is written into the SLA with its own measure
Loose exclusions become the explanation for every miss.
- If
Speed holds while sampled quality falls
ThenMake quality a gate: speed performance cannot earn back credits in a month when the quality threshold is breached
Otherwise fast but wrong output looks like success.
- If
A critical indicator is missed repeatedly
ThenDefine a persistent-failure threshold that gives you step-in, re-scoping or termination rights
Without it, credits can settle into a routine cost for the provider.
Contract clauses for model, prompt and threshold changes
Software-run work can change without anyone touching the process map. These clauses keep such changes visible and reversible.
Human-review commitments and escalation times
List the outputs that always go to a person before they take effect, such as payment releases above a limit, claim declines, account closures or replies to regulatory complaints. Then measure the time from the software's hand-off to a qualified person picking the item up. An overloaded human in the loop becomes a rubber stamp, and this measure is the earliest sign of it.
Commit to roles and competence rather than headcount: which decisions need a licensed or authorized reviewer, how cover works outside normal hours, and an escalation matrix naming contacts and response times at each level.
Hypothetical SLA schedule for a finance-operations service
Reporting cadence and the forums that read the reports
A weekly view covers queues, holds and incidents; a monthly pack covers each indicator, sampling results with root causes, credits and changes released; a quarterly review asks whether the targets still describe what the business needs. How those forums run is covered in the managed service governance checklist.
Questions and answers
Should an SLA for AI-assisted work include accuracy targets?
Yes, but expressed as quality on a defined sample rather than a blanket accuracy claim. Agree the rubric, the error classes, the sampling plan and who reviews before go-live, then set the target from a measured baseline. A speed target without a quality target rewards fast but wrong output, which is the specific failure mode of automated work.
Who should do the quality sampling, the provider or the client?
Usually the provider runs routine sampling with an independent QA function, and the client keeps the right to draw and grade its own blind sample. Calibrating both groups on a shared set of items keeps their scores comparable. For regulated decisions, the client's own second line may also need to review samples directly.
Can a provider exclude errors caused by a third-party model update?
It can ask to, but a broad exclusion leaves the client carrying a risk it cannot control. A better balance is that the provider must monitor for upstream model changes, re-validate before relying on a new version and roll back when results degrade. Errors that occur despite that process can then be handled through remediation rather than credits.
Sources
- Managed Services: operations areas and six-phase methodology — ColdAI
- Site Reliability Engineering, Chapter 3: Embracing Risk — Google · checked 10 October 2026
- ISO 2859-1 Sampling procedures for inspection by attributes, Part 1: Sampling schemes indexed by acceptance quality limit (AQL) for lot-by-lot inspection — International Organization for Standardization · checked 10 October 2026