GuideManaged Services

Writing SLAs for business operations that AI helps run

An SLA for AI-assisted operations has to measure more than speed. When software does part of the work, the main risk shifts from slow output to confident, wrong output at volume. This guide covers choosing indicators per process, evidencing quality with sampled human review, objectives and error budgets, credits, and the clauses that govern model changes and human escalation.

Reviewed 8 min read

On this page
  1. What an SLA can and cannot promise for AI-assisted work
  2. Indicators that matter for each family of operational process
  3. Designing the quality sample and the human QA behind it
  4. Setting objectives and error budgets for operational work
  5. Credits, earn-back and exclusions that change behavior
  6. Contract clauses for model, prompt and threshold changes
  7. Human-review commitments and escalation times
  8. Hypothetical SLA schedule for a finance-operations service
  9. Reporting cadence and the forums that read the reports
  10. Questions and answers
  11. Sources

What an SLA can and cannot promise for AI-assisted work

A service level agreement is a contractual promise about outcomes both parties can measure. In AI-assisted operations, many of those outcomes are decided by software: whether an invoice lands on the right ledger account, whether a claim reaches the right handler, whether a reply states the policy correctly. The agreement has to describe the quality of decisions, not only the time they take.

An SLA can reasonably promise processing times, limits on backlog age, the time a person takes to pick up an item the software has set aside, quality on a defined sample, notice before material changes and named escalation contacts. It cannot promise that no single output is ever wrong, that a model behaves identically after its vendor updates it, or that targets hold when inputs the client controls arrive late.

ColdAI describes the Operate phase of its managed services as continuous delivery by dedicated teams with SLA-backed performance1. This page explains how such an agreement should be shaped; it does not state ColdAI's own targets, which are set per process from a measured baseline.

Indicators that matter for each family of operational process

Choose two or three service level indicators per process: one for flow, one for quality and, where people review the software's work, one for the hand-off.

Process familySpeed and flowQualityExceptions and people
Finance operations (payables, reconciliations)Cycle time from receipt to posting; age of the oldest unprocessed itemFirst-time-right coding and matching, checked on a stratified sampleShare of items held for a person; time to resolve a held invoice
Customer serviceTime to first response and to resolution, by channelAccuracy and policy compliance of answers, scored against a rubricEscalations to human agents; reopened cases
Insurance claims supportTime to triage and to a settlement recommendationCorrect coverage and reserve recommendations on sampled filesTime for a licensed handler to pick up a referral; audit findings
Banking onboarding and KYCTime to complete a case; age of the queueCorrect risk rating and a complete evidence pack on sampled casesCases passed to compliance; time to clear screening alerts
Trust and safetyTime from report to actionPrecision of enforcement decisions on samples; appeals overturnedItems held for specialist review; limits on reviewer exposure

Keep further measures in operational reporting rather than the contract; a long list dilutes attention and complicates credits.

Designing the quality sample and the human QA behind it

Quality is the indicator buyers most often leave vague. These steps turn it into something both sides can check.

  1. Define what counts as right

    Write a scoring rubric per output type with error classes: critical (a wrong payee, a wrong coverage decision, a reply that breaks policy), major and minor, with agreed examples of each.

    Output
    Rubric with error classes and worked examples
    Owner
    Client process owner with the provider's QA lead
  2. Stratify the population

    Sample separately from items the software completed alone, items a person touched and high-value or high-risk items. A healthy average can hide a weak segment, and the automated one changes fastest.

    Output
    Sampling frame by segment
  3. Set sample sizes and frequency

    Size samples from volume and the precision you need, with larger samples for new processes and after changes. Acceptance-sampling schemes such as those in the ISO 2859 series are one established reference for this kind of plan3.

    Output
    Sampling schedule
  4. Keep reviewers independent and calibrated

    Reviewers should not grade their own team's work, and the client should be able to draw its own blind sample. Calibrate reviewers on a shared set of items and record disagreements.

    Output
    Calibration record
  5. Log causes, not just counts

    For every sampled error, record its class and root cause: the model, a rule, the input data or a person. Root causes decide the fix; a raw count only says something went wrong.

    Output
    Quality log feeding the monthly service pack
  6. Agree what a breach triggers

    Write down what a breached threshold triggers: expanded sampling, a root-cause review within an agreed time and, for critical errors, pausing the automated path until a fix is verified.

    Output
    Remediation playbook annexed to the SLA

Setting objectives and error budgets for operational work

A service level objective is the internal target the provider runs to; the SLA is the commitment it is paid against. Running to a stricter objective gives the provider room to correct a decline before it becomes a breach.

An error budget turns that gap into a working rule. Site reliability engineering treats the budget as the unreliability a service may spend in a period: while budget remains, changes ship; when it is spent, effort moves to reliability2. In operations, express it as tolerated late items or critical errors per month. When it runs out, planned automation changes pause until quality recovers.

Set objectives from the baseline measured during transition rather than from a proposal, and revisit them after a full business cycle that includes month-end or the seasonal peak.

Credits, earn-back and exclusions that change behavior

  • If

    A missed level reaches your customers, suppliers or regulators directly, such as late supplier payments or missed complaint deadlines

    Then

    Attach service credits to that indicator and require a written remediation plan with dates

    Credits signal priority; the plan is what changes the outcome.

  • If

    Misses are occasional and the provider recovers quickly

    Then

    Allow earn-back, so credits are returned when the level holds for an agreed run of following months

    It rewards sustained recovery over punishing one bad month.

  • If

    A miss is caused by your own inputs, such as late source files, unannounced policy changes or outages in your systems

    Then

    Exclude it, but only where the dependency is written into the SLA with its own measure

    Loose exclusions become the explanation for every miss.

  • If

    Speed holds while sampled quality falls

    Then

    Make quality a gate: speed performance cannot earn back credits in a month when the quality threshold is breached

    Otherwise fast but wrong output looks like success.

  • If

    A critical indicator is missed repeatedly

    Then

    Define a persistent-failure threshold that gives you step-in, re-scoping or termination rights

    Without it, credits can settle into a routine cost for the provider.

Contract clauses for model, prompt and threshold changes

Software-run work can change without anyone touching the process map. These clauses keep such changes visible and reversible.

0 of 6 checked

Human-review commitments and escalation times

List the outputs that always go to a person before they take effect, such as payment releases above a limit, claim declines, account closures or replies to regulatory complaints. Then measure the time from the software's hand-off to a qualified person picking the item up. An overloaded human in the loop becomes a rubber stamp, and this measure is the earliest sign of it.

Commit to roles and competence rather than headcount: which decisions need a licensed or authorized reviewer, how cover works outside normal hours, and an escalation matrix naming contacts and response times at each level.

Hypothetical SLA schedule for a finance-operations service

Reporting cadence and the forums that read the reports

A weekly view covers queues, holds and incidents; a monthly pack covers each indicator, sampling results with root causes, credits and changes released; a quarterly review asks whether the targets still describe what the business needs. How those forums run is covered in the managed service governance checklist.

Questions and answers

Should an SLA for AI-assisted work include accuracy targets?

Yes, but expressed as quality on a defined sample rather than a blanket accuracy claim. Agree the rubric, the error classes, the sampling plan and who reviews before go-live, then set the target from a measured baseline. A speed target without a quality target rewards fast but wrong output, which is the specific failure mode of automated work.

Who should do the quality sampling, the provider or the client?

Usually the provider runs routine sampling with an independent QA function, and the client keeps the right to draw and grade its own blind sample. Calibrating both groups on a shared set of items keeps their scores comparable. For regulated decisions, the client's own second line may also need to review samples directly.

Can a provider exclude errors caused by a third-party model update?

It can ask to, but a broad exclusion leaves the client carrying a risk it cannot control. A better balance is that the provider must monitor for upstream model changes, re-validate before relying on a new version and roll back when results degrade. Errors that occur despite that process can then be handled through remediation rather than credits.

Sources

  1. Managed Services: operations areas and six-phase methodology — ColdAI
  2. Site Reliability Engineering, Chapter 3: Embracing Risk — Google · checked 10 October 2026
  3. ISO 2859-1 Sampling procedures for inspection by attributes, Part 1: Sampling schemes indexed by acceptance quality limit (AQL) for lot-by-lot inspection — International Organization for Standardization · checked 10 October 2026

More in Managed Services

Back to Managed Services

Next step

Send us the process you want measured and the data you hold on it

Describe the process, its current volumes and any service levels you already report. We will suggest indicators and a sampling approach, and say where a baseline has to be measured before targets are set.

Discuss service levels