ComparisonAWS Bedrock
Choosing how to buy and run Amazon Bedrock inference capacity
Amazon Bedrock sells inference in several ways: on-demand requests in the Standard, Priority or Flex tiers, asynchronous batch jobs, hourly Provisioned Throughput and Reserved capacity. Cross-Region inference profiles and prompt caching change throughput and cost further. This guide compares the options, explains how quotas and throttling work, and maps each option to bursty, steady and offline traffic. Prices change, so treat the AWS pricing page as the reference for rates.
On this page
- Why the capacity mode matters as much as the model
- Bedrock consumption options side by side
- How on-demand quotas and throttling behave
- Batch inference for work nobody is waiting on
- Provisioned Throughput, Reserved capacity and custom models
- Geographic or global cross-Region inference
- Request-level levers to pull before buying capacity
- Matching traffic shape to a capacity mode
- A hypothetical application that splits its traffic three ways
- Questions and answers
- Sources
Why the capacity mode matters as much as the model
Two applications using the same model can have very different bills and reliability depending on how they consume it. An interactive assistant that throttles at peak hours fails its users however good its answers are; an overnight document job running at interactive rates pays for speed nobody needed. Choosing a consumption mode means matching how capacity is bought to the shape of the traffic.
The options differ on four axes: how you pay, whether you commit in advance, how requests are prioritized when capacity is tight, and where requests are processed. The table below compares them; current rates and discounts are published on the Amazon Bedrock pricing page1.
Bedrock consumption options side by side
| Option | How it is consumed | Commitment | Fits | Key limits |
|---|---|---|---|---|
| On-demand, Standard tier | Per token, request by request; the default when no tier is set2 | None | Everyday interactive and background traffic | Shared per-model quotas and throttling at peak |
| Priority tier | Per token at a premium, served ahead of Standard and Flex2 | None | Customer-facing paths that need the fastest responses | Draws on the same on-demand quota2 |
| Flex tier | Per token at a discount, for work that tolerates longer processing2 | None | Evaluations, summarization and background agent work | Slower; check that your models support it |
| Batch inference | Asynchronous jobs with JSONL input in S3 and results written back to S33 | None | Large offline datasets | No tool calling or structured output; not for provisioned models3 |
| Provisioned Throughput | Hourly billing for model units of dedicated throughput4 | None, one month or six months4 | Steady high volume on one model; customized models | Billed until deleted; not usable through inference profiles5 |
| Reserved tier | Reserved tokens-per-minute capacity that overflows to Standard when exceeded2 | One or three months, arranged through the AWS account team2 | Mission-critical traffic that cannot tolerate throttling | Minimum capacity applies; billed monthly |
Tier support varies by model, so check the model card for every model you route to.
How on-demand quotas and throttling behave
On-demand use is limited by per-model quotas on tokens and requests per minute, plus a daily token quota. At the start of each request, Bedrock reserves the input tokens plus the request's max_tokens value against the quota, then adjusts to the actual output when the request finishes6. Setting max_tokens far above typical response length therefore reduces how many requests can run at once, even though you are billed only for the tokens actually used.
For some models, each output token counts several times against the quota through a burndown rate, so a quota sized on billed tokens can throttle earlier than expected6. Throttled requests should be retried with exponential backoff and jitter; persistent throttling is a signal to request a quota increase, spread load across Regions or move part of the traffic to another mode. The Priority, Standard and Flex tiers all share the same on-demand quota2.
Batch inference for work nobody is waiting on
Batch inference suits jobs such as classifying an archive, summarizing a quarter's documents or generating evaluation outputs. Each request becomes a JSONL record in S3 in the InvokeModel or Converse format; you submit a job and collect results from S3 when it completes, with EventBridge notifying you of state changes instead of polling3.
Because each record is processed independently, batch supports neither tool calling nor structured output, and it cannot target provisioned models3. Prompt caching is unavailable for batch jobs too7. Keep batch prompts self-contained and validate the outputs after the job, since there is no interactive retry.
Provisioned Throughput, Reserved capacity and custom models
Provisioned Throughput buys model units, each delivering a set input and output token rate for one model, billed hourly whether or not it is used; longer commitments lower the hourly price4. The Provisioned Throughput documentation states that a customized model needs it4, while the model lifecycle page also describes on-demand deployments for custom models8, so confirm which applies to your base model before planning capacity.
The Reserved tier works differently: you reserve input and output tokens-per-minute capacity separately, and traffic above the reservation overflows to the Standard tier instead of failing2. Both options make sense only once on-demand metrics show sustained demand, and both should be sized from measured peaks, including cache-write tokens, rather than from forecasts.
Geographic or global cross-Region inference
Inference profiles let Bedrock send a request to another Region with spare capacity. The two profile types differ mainly in where data may be processed5.
| Consideration | Geographic profile | Global profile |
|---|---|---|
| Where requests are processed | A Region within one geography, such as the US or the EU | Any supported commercial Region worldwide |
| Fits | Workloads with data-residency requirements | Workloads without geographic restrictions that prioritize cost |
| Pricing | Standard pricing, based on the Region you call from | Described by AWS as the lower-cost option5 |
| Organization policies | Service control policies must allow every destination Region in the profile | Policies must allow requests whose requested Region is unspecified |
| Audit trail | CloudTrail in the source Region records where each request was processed | The same; check the inference Region field in each event |
Inference profiles cannot be combined with Provisioned Throughput5.
Request-level levers to pull before buying capacity
Matching traffic shape to a capacity mode
- If
Interactive traffic is bursty and hard to predict.
ThenStay on on-demand Standard, with a geographic cross-Region profile for headroom and retries with backoff.
You pay only for use, and profiles add headroom without a commitment.
- If
One customer-facing path needs the fastest responses but not round-the-clock reserved capacity.
ThenSend that path's requests in the Priority tier and keep everything else on Standard.
The tier is set per request, so only the latency-sensitive path pays the premium.
- If
Volume on one model is steady and high, and throttling would breach a service level.
ThenMeasure peaks on demand first, then size Provisioned Throughput or a Reserved tier reservation.
A commitment is cheaper and safer only when it is sized from observed demand.
- If
Large datasets are processed overnight or in bulk.
ThenUse batch inference, or the Flex tier where tool calling or structured output is needed.
Nobody is waiting on these jobs, so lower priority costs nothing in user experience.
- If
Data must stay within one jurisdiction.
ThenUse in-Region inference or geographic profiles only, and block global profiles with service control policies.
Global profiles may process requests in any supported commercial Region.
A hypothetical application that splits its traffic three ways
Questions and answers
How do we request a higher Bedrock quota?
Most on-demand quotas can be viewed and raised through the Service Quotas console or API for your account and Region, while some capacity, such as the Reserved tier and details of model units, goes through your AWS account team. Bring CloudWatch evidence of throttling and expected peaks, because requests backed by measured demand are easier to justify and to size.
Can we test Provisioned Throughput before committing to a term?
Yes. Provisioned Throughput can be bought with no commitment term and deleted at any time, which lets you load-test a model unit against your real traffic pattern. Billing is hourly until the provisioned model is deleted, so schedule the test, record throughput and latency, and remove it when you are done4.
Do cached prompt tokens count against our Bedrock quota?
No. Tokens read from the prompt cache do not count toward the tokens-per-minute and daily quotas, while tokens written to the cache do6. That makes caching a throughput lever as well as a cost lever for applications that resend long, stable prefixes such as system instructions, tool definitions or reference documents.
Which metrics show whether our capacity choice is working?
Watch input, output, cache-read and cache-write token counts, invocation latency and throttles per model in CloudWatch, along with the resolved service tier, which shows the tier that actually served each request. Compare peaks with your quota and any reservation, and review the trend monthly so commitments follow demand rather than the original forecast.
Sources
- Amazon Bedrock pricing — Amazon Web Services · checked 10 October 2026
- Service tiers for optimizing performance and cost — Amazon Web Services · checked 10 October 2026
- Process multiple prompts with batch inference — Amazon Web Services · checked 10 October 2026
- Increase model invocation capacity with Provisioned Throughput in Amazon Bedrock — Amazon Web Services · checked 10 October 2026
- Route model inference requests across AWS Regions with cross-Region inference — Amazon Web Services · checked 10 October 2026
- How tokens are counted in Amazon Bedrock — Amazon Web Services · checked 10 October 2026
- Prompt caching for faster model inference — Amazon Web Services · checked 10 October 2026
- Model lifecycle (Amazon Bedrock User Guide) — Amazon Web Services · checked 10 October 2026