Skip to content
SFAI
SFAI / PRACTICAL GUIDE

Input vs output token pricing explained

A useful cost comparison separates the prompt you send from the answer you receive.

Two different quantities

Input includes the material processed for the request; output is the generated response. Token counts are model and tokenizer dependent. A fixed number of words is not an exact token estimate across models, and multimodal inputs may follow different billing rules.

A worked example

For an illustrative workload of two million input tokens and half a million output tokens, at $1 per million input and $4 per million output, the base token charge is $4. This is arithmetic, not a provider quote. Cache, tools, media, reasoning and other charges may need separate treatment.

Compare a repeatable scenario

Use the same input/output assumptions for each candidate. Save the model version, provider, rates and date with the result. A short-answer classification workload and a long-form generation workload can produce different cost rankings even when they use the same models.

Apply the method to current data

Browse current source-qualified model profiles · Review sources and publication limits