GuideTechnology

How to set SLOs and error budgets for platforms and AI services

To set SLOs and error budgets, start from the few user journeys that matter, measure each as the share of good events among valid events, set a target users would actually notice missing, and agree in writing what the team does when the budget runs out. Alert on how fast the budget is burning rather than on raw errors. Services that call language models need extra indicators for streaming latency, tool calls and answer quality.

Reviewed 7 min read

On this page
  1. Indicator, objective, agreement: the order to define them in
  2. Five steps from a user journey to a working alert
  3. Request-based or window-based indicators: which to use where
  4. Indicators that make sense for LLM-backed services
  5. Counting third-party model APIs against your budget
  6. Responding as the error budget is consumed
  7. SLO anti-patterns that empty error budgets of meaning
  8. Questions and answers
  9. Sources

Indicator, objective, agreement: the order to define them in

Google's SRE book separates three terms that are often blurred. A service level indicator is a carefully defined measurement of some aspect of the service; an objective is the target value for that indicator; an agreement is a contract with users that carries consequences when objectives are missed1. If there is no consequence, what you have is an objective, not an agreement.

Define them in that order. Teams that start from a contractual number tend to choose indicators that make the number easy to hit rather than indicators that reflect user pain. Internal objectives should be stricter than anything promised externally, so that the team has room to react before a contract is breached.

ColdAI's technology practice lists SRE practices and observability-first design as part of operating the platforms it builds2. The method below is the one we apply; nothing in it requires outside help.

Five steps from a user journey to a working alert

  1. Choose the critical user journeys

    List the handful of things users come to the service to do: sign in, search, submit an order, get an answer from an assistant. Rank them by the cost of failure and keep the top three to five. A journey nobody would escalate is not worth an objective.

    Output
    Ranked journey list
    Owner
    Product owner with the service team
  2. Specify the indicator, then decide how to measure it

    The SRE Workbook distinguishes the indicator specification, the outcome that matters to users, from its implementation, the method used to measure it3. Express each indicator as good events divided by valid events. Measuring at the load balancer is cheap but misses client failures; synthetic probes and client telemetry cost more and see more.

    Output
    Indicator specification and measurement point
    Owner
    Service team
  3. Set the target from evidence and user tolerance

    Look at recent performance, then ask where users start to complain or abandon. The Workbook warns against simply adopting current performance, which can lock you into a target that is stricter than needed; use it as a starting point and revise3. Pick a rolling window, commonly four weeks, so the budget recovers gradually.

    Output
    Target and window per indicator
    Owner
    Service owner, agreed with product
  4. Write and sign the error-budget policy

    The budget is the share of events allowed to fail. The policy states what the team does as it is consumed and when it is exhausted, who can grant exceptions and who settles disputes. Google's example policy halts changes other than urgent fixes and security fixes until the service is back within objective, and escalates disagreements to the CTO4.

    Output
    Signed error-budget policy
    Owner
    Engineering and product leadership
  5. Alert on burn rate across two windows

    Burn rate is how fast the budget is being spent relative to the objective; a rate of one spends exactly the whole budget over the window5. Fire an alert only when both a long window and a short window exceed the threshold: the long window proves the spend is significant, the short window proves it is still happening. Route fast burns to a pager and slow burns to a ticket queue.

    Output
    Paging and ticketing alert rules
    Owner
    On-call lead

Request-based or window-based indicators: which to use where

QuestionRequest-basedWindow-based
What it countsGood requests over valid requestsGood time slices over all time slices
Low or bursty trafficNoisy: one failure moves the ratio sharplySteadier, but a quiet hour counts as fully good
Fairness to heavy usersWeights every request equallyHides failures concentrated in busy minutes
Typical fitCustomer-facing APIs and checkout flowsBatch jobs, pipelines and low-volume internal tools
Budget meaningNumber of requests that may failMinutes or hours of degraded service allowed

Most teams use request-based indicators for interactive services and window-based ones where traffic is too thin for a ratio to mean anything.

Indicators that make sense for LLM-backed services

Classic availability and latency still apply. These indicators cover what users experience in a streaming, tool-using assistant.

IndicatorWhat counts as goodWhere to measureWatch out for
Time to first tokenFirst streamed token within the threshold users tolerateYour gateway, not the provider's dashboardLong prompts and retrieval add hidden delay
Full response latencyComplete answer within an agreed ceilingClient or gateway timingPenalizes useful long answers; set per journey
Tool-call successCall is well-formed, authorized and returns a usable resultAgent runtime logsRetries can hide failures; count the first attempt
Grounded-answer pass rateAnswer passes an automated evaluation on a fixed setScheduled evaluation runsEvaluation sets go stale; refresh them deliberately
Safe completionResponse is neither blocked wrongly nor unsafeGuardrail and review logsOver-blocking is a reliability failure too

Evaluation-based indicators are sampled, not measured on every request, so treat their budgets as slower-moving and alert on them through tickets rather than pages.

Counting third-party model APIs against your budget

If your assistant depends on a hosted model, its outages are your outages from the user's point of view, so they count against your objective. Excluding them makes the dashboard look healthy while users wait. What you can change is the design: a fallback model, cached answers for common questions, or a graceful message that hands the user to a person.

Track the dependency separately as well. A per-provider indicator tells you whether failures come from the provider or from your own retrieval and orchestration, and gives evidence for contract conversations. Your own objective cannot sensibly be stricter than the combined reliability of the dependencies on the critical path unless you add redundancy.

Responding as the error budget is consumed

  • If

    Less than half the budget is spent and the trend is flat.

    Then

    Ship normally and consider using the spare budget for riskier releases or planned maintenance.

    An unspent budget usually means the target is stricter than users need, or the team is moving too cautiously.

  • If

    A fast-burn alert fires during a release.

    Then

    Roll back first and investigate afterwards.

    At high burn rates the budget can be gone within hours, and diagnosis rarely beats a rollback.

  • If

    The budget is exhausted for the window.

    Then

    Apply the policy: pause feature changes, prioritize reliability work and hold a review of the incidents that spent it.

    Without a pre-agreed consequence the objective becomes a number people watch rather than one that changes behavior.

  • If

    The budget is repeatedly exhausted but users are not complaining.

    Then

    Revisit the indicator or the target before demanding more engineering effort.

    The objective may be measuring something users do not feel, or set above what they need.

SLO anti-patterns that empty error budgets of meaning

An objective for every metric

Early signalDashboards list dozens of objectives per service and nobody can name the important ones.

MitigationKeep a few per critical journey and demote the rest to diagnostic metrics.

Aspirational targets

Early signalThe target is chosen because it sounds good and is missed every window.

MitigationBase it on evidence of user tolerance and tighten only when users need it.

Alerting on raw error rates

Early signalOn-call engineers are paged for brief spikes that never threaten the budget.

MitigationReplace threshold alerts with multiwindow burn-rate rules.

No review cadence

Early signalObjectives set at launch are never revisited as traffic and features change.

MitigationReview indicators and targets each quarter with product, and after every major incident.

Questions and answers

How many SLOs should one service have?

Usually one to three per critical user journey, and rarely more than a handful per service. Each additional objective dilutes attention and adds alert rules to maintain. If an indicator never changes a decision, keep it as a diagnostic metric on a dashboard instead of giving it a target and a budget.

Do internal developer platforms need SLOs?

Yes. A platform's users are engineers, and their journeys are things like running a pipeline, provisioning an environment or deploying a service. Objectives for those journeys tell the platform team where reliability work matters and give product teams an honest picture of what they can depend on.

What is the difference between an SLO and an SLA in a contract?

An SLA is a promise to a customer with consequences, often service credits, if it is missed. An SLO is an internal target that drives engineering decisions. Set SLOs stricter than any SLA so that burning the internal budget warns you well before a contractual breach occurs.

How do you set an SLO for a brand-new service with no history?

Start from the user journey and a reasoned guess about tolerance, label the objective provisional, and measure for a full window before enforcing the error-budget policy. Collect data on where users abandon or complain, then set the first enforced target from that evidence rather than from the launch estimate.

Sources

  1. Site Reliability Engineering, Chapter 4: Service Level Objectives — Google · checked 10 October 2026
  2. Technology capability: Operate & Evolve — ColdAI
  3. The Site Reliability Workbook, Chapter 2: Implementing SLOs — Google · checked 10 October 2026
  4. The Site Reliability Workbook, Appendix B: Example Error Budget Policy — Google · checked 10 October 2026
  5. The Site Reliability Workbook, Chapter 5: Alerting on SLOs — Google · checked 10 October 2026

More in Technology

Back to Technology

Next step

Get a second opinion on your first SLOs and alert rules

Send one service's user journeys, current dashboards and alert rules. We will return suggested indicators, targets to test and a draft error-budget policy you can adapt.

Send a service to review