Buyer's guideArtificial Intelligence

Estimating the total cost of an AI system, component by component

The price per million tokens is the easiest number to find and often a minor part of what an AI system costs. Evaluation, human review, integration, monitoring, re-indexing and model migrations recur for as long as the system runs. This guide sets out a component-based cost model, a worksheet built from variables rather than prices, levers that cut cost without hurting quality and questions for vendors.

Reviewed 8 min read

On this page
  1. Why the per-token price is a poor budget anchor
  2. The cost components around one AI workload
  3. What drives each cost line and when it recurs
  4. Laying out the AI cost worksheet
  5. A hypothetical claims-intake assistant on the worksheet
  6. Build, buy or managed service: where the costs land
  7. Cost levers that do not cost you quality
  8. Comparing AI providers with source-attributed evidence
  9. Questions to ask an AI vendor before signing
  10. Questions and answers
  11. Sources

Why the per-token price is a poor budget anchor

Provider price lists quote rates for input and output tokens, and some offer lower rates for cached input or for requests processed in batches123. Those rates matter, but they describe a single ingredient. A request that pulls several long passages into the prompt, retries after a malformed response and is then checked by a person can cost far more than its token count suggests, and most of the difference never appears on the model invoice.

Budget surprises tend to come from three places: production volumes and request lengths that differ from the pilot; recurring work treated as a one-off, such as re-indexing, evaluation and model migrations; and people's time spent reviewing outputs and handling exceptions. Naming those components up front turns the conversation with finance into a discussion of assumptions rather than one number to defend.

The cost components around one AI workload

01The AI workload02Inference03Retrieval andstorage04Evaluation runs05Human review06Integration andchange07Operations08Model migrations
  1. The AI workload

    One application serving a defined workflow at an expected volume.

  2. Inference

    Model calls priced by tokens, time or reserved capacity, multiplied by retries.

  3. Retrieval and storage

    Embedding, indexing and search storage, plus re-indexing whenever sources change.

  4. Evaluation runs

    Test-set runs after every change, including judge-model calls and reviewer time.

  5. Human review

    Approvals, exception handling and quality sampling, counted in people's hours.

  6. Integration and change

    Connectors, identity, workflow redesign and training for the people affected.

  7. Operations

    Monitoring, log retention, incident response, security reviews and on-call cover.

  8. Model migrations

    Moving to new model versions or providers, with re-evaluation and prompt rework.

Conceptual map of the cost components of one AI workload; relative sizes vary by design and are not to scale.

What drives each cost line and when it recurs

ComponentMain driversWhen it recursWhy it gets missed
InferenceRequests per period, input and output length, model tier and retry rateContinuously, in line with volumePilot inputs are short and clean; production inputs are longer and messier
Retrieval and storageCorpus size, chunking, embedding model and how often sources changeWhenever sources change or the embedding model is replacedRe-embedding a whole corpus is treated as a one-time setup task
EvaluationSet size, judge-model calls, reviewer hours and release frequencyAfter every prompt, model or source changeIt sits in engineering time rather than on a vendor invoice
Human reviewShare of items reviewed, minutes per item and reviewer costFor every reviewed item, for the life of the systemBusiness cases assume review falls away faster than evidence allows
Integration and changeNumber of systems touched, identity and permissions work, trainingWith each new workflow, system upgrade or policy changeIt is booked as project cost and left out of run-rate estimates
Operations and migrationsMonitoring tools, log retention, incidents and model deprecationsOngoing, plus each time a provider retires a model versionModel versions are retired on the provider's timetable, not yours

Laying out the AI cost worksheet

Use variables you can defend, fill them from pilot logs and keep provider rates in one table.

  1. Split traffic into request classes

    Group requests by shape, such as short classifications, document questions, long-form drafting and multi-step agent tasks, and record the expected volume per period for each class.

    Output
    Request-class table
  2. Measure tokens and calls per class

    From pilot logs, record average input and output tokens, retrieved passages added to each prompt, tool calls and the retry rate, rather than figures from a vendor demonstration.

    Output
    Usage profile per class
  3. Apply rates by route

    Assign each class to the model and processing route it will use, such as a smaller model for classification or batch processing for overnight jobs, and apply current rates from the provider's own pricing page, for example Amazon Bedrock or Azure OpenAI78.

    Output
    Inference cost per class
  4. Add the people lines

    Multiply reviewed items by minutes per item and a loaded hourly cost, then add exception handling, source owners' maintenance time and evaluation reviewer hours per release.

    Output
    People cost per period
  5. Add platform and operations lines

    Include search or vector storage, re-indexing runs, monitoring, log retention and an allowance for incidents and security reviews.

    Output
    Run-rate cost
  6. Divide by completed outcomes

    Express the total as cost per completed task, such as a resolved request or a processed document, and set it beside the cost of doing the same work today.

    Output
    Cost per outcome
  7. Stress the assumptions

    Rerun the worksheet with higher volumes, longer inputs, a larger review share and a model migration, so the budget shows a range rather than a point.

    Output
    Cost range for approval

A hypothetical claims-intake assistant on the worksheet

Build, buy or managed service: where the costs land

Cost lineBuild in-houseBuy a packaged productManaged service
Upfront spendEngineering time to design, integrate and evaluateLicense and configuration; integration is still neededOnboarding and transition, with design work shared with the provider
Run-cost visibilityFull: you see every model call and storage lineBundled into seat or usage pricingDefined by the contract's usage metering
Model and provider choiceYours, including routing between providersThe vendor's, sometimes with limited optionsShared, according to the service agreement
Migration effortYours to plan and fundAbsorbed by the vendor, on the vendor's timetableUsually handled by the provider within agreed change terms

Cost levers that do not cost you quality

Pull a lever only after the evaluation set confirms that quality holds.

  • If

    Many requests share the same long instructions or reference text

    Then

    Use prompt caching where the provider supports it, and keep the stable content at the start of the prompt45.

    Providers that offer caching bill cached input below their standard input rate45.

  • If

    Some steps are simple classification or routing

    Then

    Route them to a smaller, cheaper model and confirm on the evaluation set that accuracy holds.

  • If

    Results are not needed immediately, as with overnight document processing or evaluation runs

    Then

    Use the provider's batch interface, which some providers price below synchronous requests23.

  • If

    Usage is high and steady, or the data must stay in your own environment

    Then

    Model reserved-capacity or self-hosted options against pay-per-use over the same period; see enterprise AI for private deployment.

Comparing AI providers with source-attributed evidence

Rates, context limits and benchmark claims change often, and comparisons copied into slide decks go stale quickly. Link the worksheet's rate table to primary sources, record the date each rate was checked and keep published benchmark scores separate from results on your own evaluation set.

ColdAI built SFAIE, an AI model and provider intelligence exchange, for this kind of research. It brings models, inference providers, pricing and source-attributed benchmarks into one interface, with a calculator that turns expected volume, token counts and cache assumptions into indicative costs6. Treat any such estimate as a starting point for the worksheet, not a substitute for usage measured in your own pilot.

Questions to ask an AI vendor before signing

0 of 6 checked

Questions and answers

How should an AI pilot be budgeted differently from production?

Budget a pilot for learning: integration with real systems, building the evaluation set and reviewer time are the main lines, and inference is often small. Production adds volume, monitoring, on-call cover, log retention and migrations. Use the pilot to measure what the production worksheet needs, such as tokens per request class and review minutes per item.

Is cost per outcome a better measure than cost per call?

For budget decisions, yes. Cost per call rewards cheap calls even when they fail and need a retry or a person to fix the result. Cost per completed outcome, such as a resolved request or a processed claim, includes retries, review and exceptions, and can be compared directly with what the same work costs today.

When does self-hosting a model pay for itself?

It can when usage is high and steady, when data must stay in your environment or when you need control over model versions. It also brings hardware or reserved-capacity costs, operations staff and your own upgrade cycle. Compare both options over the same period at the same quality bar; our self-hosted LLM reference architecture covers the design.

Why do evaluation runs belong in the cost model?

Every prompt, model or source change should be followed by an evaluation rerun, and each run consumes model calls and reviewer time. Teams that leave this out either under-budget or start skipping evaluations, which moves the cost into production incidents. Our guide to building an LLM evaluation set explains what a run involves.

Sources

  1. Claude pricing — Anthropic · checked 10 October 2026
  2. Batch processing with the Message Batches API — Anthropic · checked 10 October 2026
  3. Batch API guide — OpenAI · checked 10 October 2026
  4. Prompt caching — Anthropic · checked 10 October 2026
  5. Prompt caching guide — OpenAI · checked 10 October 2026
  6. SFAIE: AI model and provider intelligence exchange (case study) — ColdAI
  7. Amazon Bedrock pricing — Amazon Web Services · checked 10 October 2026
  8. Azure OpenAI pricing — Microsoft · checked 10 October 2026

More in Artificial Intelligence

Back to Artificial Intelligence

Next step

Get a second opinion on your AI cost model

Share the workflow, expected volumes and any vendor quotes you have. We will point out missing cost lines and the assumptions most worth testing in a pilot.

Review my cost model