GuideTechnology
How to set SLOs and error budgets for platforms and AI services
To set SLOs and error budgets, start from the few user journeys that matter, measure each as the share of good events among valid events, set a target users would actually notice missing, and agree in writing what the team does when the budget runs out. Alert on how fast the budget is burning rather than on raw errors. Services that call language models need extra indicators for streaming latency, tool calls and answer quality.
On this page
- Indicator, objective, agreement: the order to define them in
- Five steps from a user journey to a working alert
- Request-based or window-based indicators: which to use where
- Indicators that make sense for LLM-backed services
- Counting third-party model APIs against your budget
- Responding as the error budget is consumed
- SLO anti-patterns that empty error budgets of meaning
- Questions and answers
- Sources
Indicator, objective, agreement: the order to define them in
Google's SRE book separates three terms that are often blurred. A service level indicator is a carefully defined measurement of some aspect of the service; an objective is the target value for that indicator; an agreement is a contract with users that carries consequences when objectives are missed1. If there is no consequence, what you have is an objective, not an agreement.
Define them in that order. Teams that start from a contractual number tend to choose indicators that make the number easy to hit rather than indicators that reflect user pain. Internal objectives should be stricter than anything promised externally, so that the team has room to react before a contract is breached.
ColdAI's technology practice lists SRE practices and observability-first design as part of operating the platforms it builds2. The method below is the one we apply; nothing in it requires outside help.
Five steps from a user journey to a working alert
Choose the critical user journeys
List the handful of things users come to the service to do: sign in, search, submit an order, get an answer from an assistant. Rank them by the cost of failure and keep the top three to five. A journey nobody would escalate is not worth an objective.
Specify the indicator, then decide how to measure it
The SRE Workbook distinguishes the indicator specification, the outcome that matters to users, from its implementation, the method used to measure it3. Express each indicator as good events divided by valid events. Measuring at the load balancer is cheap but misses client failures; synthetic probes and client telemetry cost more and see more.
Set the target from evidence and user tolerance
Look at recent performance, then ask where users start to complain or abandon. The Workbook warns against simply adopting current performance, which can lock you into a target that is stricter than needed; use it as a starting point and revise3. Pick a rolling window, commonly four weeks, so the budget recovers gradually.
Write and sign the error-budget policy
The budget is the share of events allowed to fail. The policy states what the team does as it is consumed and when it is exhausted, who can grant exceptions and who settles disputes. Google's example policy halts changes other than urgent fixes and security fixes until the service is back within objective, and escalates disagreements to the CTO4.
Alert on burn rate across two windows
Burn rate is how fast the budget is being spent relative to the objective; a rate of one spends exactly the whole budget over the window5. Fire an alert only when both a long window and a short window exceed the threshold: the long window proves the spend is significant, the short window proves it is still happening. Route fast burns to a pager and slow burns to a ticket queue.
Request-based or window-based indicators: which to use where
| Question | Request-based | Window-based |
|---|---|---|
| What it counts | Good requests over valid requests | Good time slices over all time slices |
| Low or bursty traffic | Noisy: one failure moves the ratio sharply | Steadier, but a quiet hour counts as fully good |
| Fairness to heavy users | Weights every request equally | Hides failures concentrated in busy minutes |
| Typical fit | Customer-facing APIs and checkout flows | Batch jobs, pipelines and low-volume internal tools |
| Budget meaning | Number of requests that may fail | Minutes or hours of degraded service allowed |
Most teams use request-based indicators for interactive services and window-based ones where traffic is too thin for a ratio to mean anything.
Indicators that make sense for LLM-backed services
Classic availability and latency still apply. These indicators cover what users experience in a streaming, tool-using assistant.
| Indicator | What counts as good | Where to measure | Watch out for |
|---|---|---|---|
| Time to first token | First streamed token within the threshold users tolerate | Your gateway, not the provider's dashboard | Long prompts and retrieval add hidden delay |
| Full response latency | Complete answer within an agreed ceiling | Client or gateway timing | Penalizes useful long answers; set per journey |
| Tool-call success | Call is well-formed, authorized and returns a usable result | Agent runtime logs | Retries can hide failures; count the first attempt |
| Grounded-answer pass rate | Answer passes an automated evaluation on a fixed set | Scheduled evaluation runs | Evaluation sets go stale; refresh them deliberately |
| Safe completion | Response is neither blocked wrongly nor unsafe | Guardrail and review logs | Over-blocking is a reliability failure too |
Evaluation-based indicators are sampled, not measured on every request, so treat their budgets as slower-moving and alert on them through tickets rather than pages.
Counting third-party model APIs against your budget
If your assistant depends on a hosted model, its outages are your outages from the user's point of view, so they count against your objective. Excluding them makes the dashboard look healthy while users wait. What you can change is the design: a fallback model, cached answers for common questions, or a graceful message that hands the user to a person.
Track the dependency separately as well. A per-provider indicator tells you whether failures come from the provider or from your own retrieval and orchestration, and gives evidence for contract conversations. Your own objective cannot sensibly be stricter than the combined reliability of the dependencies on the critical path unless you add redundancy.
Responding as the error budget is consumed
- If
Less than half the budget is spent and the trend is flat.
ThenShip normally and consider using the spare budget for riskier releases or planned maintenance.
An unspent budget usually means the target is stricter than users need, or the team is moving too cautiously.
- If
A fast-burn alert fires during a release.
ThenRoll back first and investigate afterwards.
At high burn rates the budget can be gone within hours, and diagnosis rarely beats a rollback.
- If
The budget is exhausted for the window.
ThenApply the policy: pause feature changes, prioritize reliability work and hold a review of the incidents that spent it.
Without a pre-agreed consequence the objective becomes a number people watch rather than one that changes behavior.
- If
The budget is repeatedly exhausted but users are not complaining.
ThenRevisit the indicator or the target before demanding more engineering effort.
The objective may be measuring something users do not feel, or set above what they need.
SLO anti-patterns that empty error budgets of meaning
An objective for every metric
Early signalDashboards list dozens of objectives per service and nobody can name the important ones.
MitigationKeep a few per critical journey and demote the rest to diagnostic metrics.
Aspirational targets
Early signalThe target is chosen because it sounds good and is missed every window.
MitigationBase it on evidence of user tolerance and tighten only when users need it.
Alerting on raw error rates
Early signalOn-call engineers are paged for brief spikes that never threaten the budget.
MitigationReplace threshold alerts with multiwindow burn-rate rules.
No review cadence
Early signalObjectives set at launch are never revisited as traffic and features change.
MitigationReview indicators and targets each quarter with product, and after every major incident.
Questions and answers
How many SLOs should one service have?
Usually one to three per critical user journey, and rarely more than a handful per service. Each additional objective dilutes attention and adds alert rules to maintain. If an indicator never changes a decision, keep it as a diagnostic metric on a dashboard instead of giving it a target and a budget.
Do internal developer platforms need SLOs?
Yes. A platform's users are engineers, and their journeys are things like running a pipeline, provisioning an environment or deploying a service. Objectives for those journeys tell the platform team where reliability work matters and give product teams an honest picture of what they can depend on.
What is the difference between an SLO and an SLA in a contract?
An SLA is a promise to a customer with consequences, often service credits, if it is missed. An SLO is an internal target that drives engineering decisions. Set SLOs stricter than any SLA so that burning the internal budget warns you well before a contractual breach occurs.
How do you set an SLO for a brand-new service with no history?
Start from the user journey and a reasoned guess about tolerance, label the objective provisional, and measure for a full window before enforcing the error-budget policy. Collect data on where users abandon or complain, then set the first enforced target from that evidence rather than from the launch estimate.
Sources
- Site Reliability Engineering, Chapter 4: Service Level Objectives — Google · checked 10 October 2026
- Technology capability: Operate & Evolve — ColdAI
- The Site Reliability Workbook, Chapter 2: Implementing SLOs — Google · checked 10 October 2026
- The Site Reliability Workbook, Appendix B: Example Error Budget Policy — Google · checked 10 October 2026
- The Site Reliability Workbook, Chapter 5: Alerting on SLOs — Google · checked 10 October 2026