Buyer's guideArtificial Intelligence
Estimating the total cost of an AI system, component by component
The price per million tokens is the easiest number to find and often a minor part of what an AI system costs. Evaluation, human review, integration, monitoring, re-indexing and model migrations recur for as long as the system runs. This guide sets out a component-based cost model, a worksheet built from variables rather than prices, levers that cut cost without hurting quality and questions for vendors.
On this page
- Why the per-token price is a poor budget anchor
- The cost components around one AI workload
- What drives each cost line and when it recurs
- Laying out the AI cost worksheet
- A hypothetical claims-intake assistant on the worksheet
- Build, buy or managed service: where the costs land
- Cost levers that do not cost you quality
- Comparing AI providers with source-attributed evidence
- Questions to ask an AI vendor before signing
- Questions and answers
- Sources
Why the per-token price is a poor budget anchor
Provider price lists quote rates for input and output tokens, and some offer lower rates for cached input or for requests processed in batches123. Those rates matter, but they describe a single ingredient. A request that pulls several long passages into the prompt, retries after a malformed response and is then checked by a person can cost far more than its token count suggests, and most of the difference never appears on the model invoice.
Budget surprises tend to come from three places: production volumes and request lengths that differ from the pilot; recurring work treated as a one-off, such as re-indexing, evaluation and model migrations; and people's time spent reviewing outputs and handling exceptions. Naming those components up front turns the conversation with finance into a discussion of assumptions rather than one number to defend.
The cost components around one AI workload
- The AI workload
One application serving a defined workflow at an expected volume.
- Inference
Model calls priced by tokens, time or reserved capacity, multiplied by retries.
- Retrieval and storage
Embedding, indexing and search storage, plus re-indexing whenever sources change.
- Evaluation runs
Test-set runs after every change, including judge-model calls and reviewer time.
- Human review
Approvals, exception handling and quality sampling, counted in people's hours.
- Integration and change
Connectors, identity, workflow redesign and training for the people affected.
- Operations
Monitoring, log retention, incident response, security reviews and on-call cover.
- Model migrations
Moving to new model versions or providers, with re-evaluation and prompt rework.
What drives each cost line and when it recurs
| Component | Main drivers | When it recurs | Why it gets missed |
|---|---|---|---|
| Inference | Requests per period, input and output length, model tier and retry rate | Continuously, in line with volume | Pilot inputs are short and clean; production inputs are longer and messier |
| Retrieval and storage | Corpus size, chunking, embedding model and how often sources change | Whenever sources change or the embedding model is replaced | Re-embedding a whole corpus is treated as a one-time setup task |
| Evaluation | Set size, judge-model calls, reviewer hours and release frequency | After every prompt, model or source change | It sits in engineering time rather than on a vendor invoice |
| Human review | Share of items reviewed, minutes per item and reviewer cost | For every reviewed item, for the life of the system | Business cases assume review falls away faster than evidence allows |
| Integration and change | Number of systems touched, identity and permissions work, training | With each new workflow, system upgrade or policy change | It is booked as project cost and left out of run-rate estimates |
| Operations and migrations | Monitoring tools, log retention, incidents and model deprecations | Ongoing, plus each time a provider retires a model version | Model versions are retired on the provider's timetable, not yours |
Laying out the AI cost worksheet
Use variables you can defend, fill them from pilot logs and keep provider rates in one table.
Split traffic into request classes
Group requests by shape, such as short classifications, document questions, long-form drafting and multi-step agent tasks, and record the expected volume per period for each class.
Measure tokens and calls per class
From pilot logs, record average input and output tokens, retrieved passages added to each prompt, tool calls and the retry rate, rather than figures from a vendor demonstration.
Apply rates by route
Assign each class to the model and processing route it will use, such as a smaller model for classification or batch processing for overnight jobs, and apply current rates from the provider's own pricing page, for example Amazon Bedrock or Azure OpenAI78.
Add the people lines
Multiply reviewed items by minutes per item and a loaded hourly cost, then add exception handling, source owners' maintenance time and evaluation reviewer hours per release.
Add platform and operations lines
Include search or vector storage, re-indexing runs, monitoring, log retention and an allowance for incidents and security reviews.
Divide by completed outcomes
Express the total as cost per completed task, such as a resolved request or a processed document, and set it beside the cost of doing the same work today.
Stress the assumptions
Rerun the worksheet with higher volumes, longer inputs, a larger review share and a model migration, so the budget shows a range rather than a point.
A hypothetical claims-intake assistant on the worksheet
Build, buy or managed service: where the costs land
| Cost line | Build in-house | Buy a packaged product | Managed service |
|---|---|---|---|
| Upfront spend | Engineering time to design, integrate and evaluate | License and configuration; integration is still needed | Onboarding and transition, with design work shared with the provider |
| Run-cost visibility | Full: you see every model call and storage line | Bundled into seat or usage pricing | Defined by the contract's usage metering |
| Model and provider choice | Yours, including routing between providers | The vendor's, sometimes with limited options | Shared, according to the service agreement |
| Migration effort | Yours to plan and fund | Absorbed by the vendor, on the vendor's timetable | Usually handled by the provider within agreed change terms |
Cost levers that do not cost you quality
Pull a lever only after the evaluation set confirms that quality holds.
- If
Many requests share the same long instructions or reference text
- If
Some steps are simple classification or routing
ThenRoute them to a smaller, cheaper model and confirm on the evaluation set that accuracy holds.
- If
Results are not needed immediately, as with overnight document processing or evaluation runs
- If
Usage is high and steady, or the data must stay in your own environment
ThenModel reserved-capacity or self-hosted options against pay-per-use over the same period; see enterprise AI for private deployment.
Comparing AI providers with source-attributed evidence
Rates, context limits and benchmark claims change often, and comparisons copied into slide decks go stale quickly. Link the worksheet's rate table to primary sources, record the date each rate was checked and keep published benchmark scores separate from results on your own evaluation set.
ColdAI built SFAIE, an AI model and provider intelligence exchange, for this kind of research. It brings models, inference providers, pricing and source-attributed benchmarks into one interface, with a calculator that turns expected volume, token counts and cache assumptions into indicative costs6. Treat any such estimate as a starting point for the worksheet, not a substitute for usage measured in your own pilot.
Questions to ask an AI vendor before signing
Questions and answers
How should an AI pilot be budgeted differently from production?
Budget a pilot for learning: integration with real systems, building the evaluation set and reviewer time are the main lines, and inference is often small. Production adds volume, monitoring, on-call cover, log retention and migrations. Use the pilot to measure what the production worksheet needs, such as tokens per request class and review minutes per item.
Is cost per outcome a better measure than cost per call?
For budget decisions, yes. Cost per call rewards cheap calls even when they fail and need a retry or a person to fix the result. Cost per completed outcome, such as a resolved request or a processed claim, includes retries, review and exceptions, and can be compared directly with what the same work costs today.
When does self-hosting a model pay for itself?
It can when usage is high and steady, when data must stay in your environment or when you need control over model versions. It also brings hardware or reserved-capacity costs, operations staff and your own upgrade cycle. Compare both options over the same period at the same quality bar; our self-hosted LLM reference architecture covers the design.
Why do evaluation runs belong in the cost model?
Every prompt, model or source change should be followed by an evaluation rerun, and each run consumes model calls and reviewer time. Teams that leave this out either under-budget or start skipping evaluations, which moves the cost into production incidents. Our guide to building an LLM evaluation set explains what a run involves.
Sources
- Claude pricing — Anthropic · checked 10 October 2026
- Batch processing with the Message Batches API — Anthropic · checked 10 October 2026
- Batch API guide — OpenAI · checked 10 October 2026
- Prompt caching — Anthropic · checked 10 October 2026
- Prompt caching guide — OpenAI · checked 10 October 2026
- SFAIE: AI model and provider intelligence exchange (case study) — ColdAI
- Amazon Bedrock pricing — Amazon Web Services · checked 10 October 2026
- Azure OpenAI pricing — Microsoft · checked 10 October 2026