ComparisonAWS Bedrock

Choosing how to buy and run Amazon Bedrock inference capacity

Amazon Bedrock sells inference in several ways: on-demand requests in the Standard, Priority or Flex tiers, asynchronous batch jobs, hourly Provisioned Throughput and Reserved capacity. Cross-Region inference profiles and prompt caching change throughput and cost further. This guide compares the options, explains how quotas and throttling work, and maps each option to bursty, steady and offline traffic. Prices change, so treat the AWS pricing page as the reference for rates.

Reviewed 7 min read

On this page
  1. Why the capacity mode matters as much as the model
  2. Bedrock consumption options side by side
  3. How on-demand quotas and throttling behave
  4. Batch inference for work nobody is waiting on
  5. Provisioned Throughput, Reserved capacity and custom models
  6. Geographic or global cross-Region inference
  7. Request-level levers to pull before buying capacity
  8. Matching traffic shape to a capacity mode
  9. A hypothetical application that splits its traffic three ways
  10. Questions and answers
  11. Sources

Why the capacity mode matters as much as the model

Two applications using the same model can have very different bills and reliability depending on how they consume it. An interactive assistant that throttles at peak hours fails its users however good its answers are; an overnight document job running at interactive rates pays for speed nobody needed. Choosing a consumption mode means matching how capacity is bought to the shape of the traffic.

The options differ on four axes: how you pay, whether you commit in advance, how requests are prioritized when capacity is tight, and where requests are processed. The table below compares them; current rates and discounts are published on the Amazon Bedrock pricing page1.

Bedrock consumption options side by side

OptionHow it is consumedCommitmentFitsKey limits
On-demand, Standard tierPer token, request by request; the default when no tier is set2NoneEveryday interactive and background trafficShared per-model quotas and throttling at peak
Priority tierPer token at a premium, served ahead of Standard and Flex2NoneCustomer-facing paths that need the fastest responsesDraws on the same on-demand quota2
Flex tierPer token at a discount, for work that tolerates longer processing2NoneEvaluations, summarization and background agent workSlower; check that your models support it
Batch inferenceAsynchronous jobs with JSONL input in S3 and results written back to S33NoneLarge offline datasetsNo tool calling or structured output; not for provisioned models3
Provisioned ThroughputHourly billing for model units of dedicated throughput4None, one month or six months4Steady high volume on one model; customized modelsBilled until deleted; not usable through inference profiles5
Reserved tierReserved tokens-per-minute capacity that overflows to Standard when exceeded2One or three months, arranged through the AWS account team2Mission-critical traffic that cannot tolerate throttlingMinimum capacity applies; billed monthly

Tier support varies by model, so check the model card for every model you route to.

How on-demand quotas and throttling behave

On-demand use is limited by per-model quotas on tokens and requests per minute, plus a daily token quota. At the start of each request, Bedrock reserves the input tokens plus the request's max_tokens value against the quota, then adjusts to the actual output when the request finishes6. Setting max_tokens far above typical response length therefore reduces how many requests can run at once, even though you are billed only for the tokens actually used.

For some models, each output token counts several times against the quota through a burndown rate, so a quota sized on billed tokens can throttle earlier than expected6. Throttled requests should be retried with exponential backoff and jitter; persistent throttling is a signal to request a quota increase, spread load across Regions or move part of the traffic to another mode. The Priority, Standard and Flex tiers all share the same on-demand quota2.

Batch inference for work nobody is waiting on

Batch inference suits jobs such as classifying an archive, summarizing a quarter's documents or generating evaluation outputs. Each request becomes a JSONL record in S3 in the InvokeModel or Converse format; you submit a job and collect results from S3 when it completes, with EventBridge notifying you of state changes instead of polling3.

Because each record is processed independently, batch supports neither tool calling nor structured output, and it cannot target provisioned models3. Prompt caching is unavailable for batch jobs too7. Keep batch prompts self-contained and validate the outputs after the job, since there is no interactive retry.

Provisioned Throughput, Reserved capacity and custom models

Provisioned Throughput buys model units, each delivering a set input and output token rate for one model, billed hourly whether or not it is used; longer commitments lower the hourly price4. The Provisioned Throughput documentation states that a customized model needs it4, while the model lifecycle page also describes on-demand deployments for custom models8, so confirm which applies to your base model before planning capacity.

The Reserved tier works differently: you reserve input and output tokens-per-minute capacity separately, and traffic above the reservation overflows to the Standard tier instead of failing2. Both options make sense only once on-demand metrics show sustained demand, and both should be sized from measured peaks, including cache-write tokens, rather than from forecasts.

Geographic or global cross-Region inference

Inference profiles let Bedrock send a request to another Region with spare capacity. The two profile types differ mainly in where data may be processed5.

ConsiderationGeographic profileGlobal profile
Where requests are processedA Region within one geography, such as the US or the EUAny supported commercial Region worldwide
FitsWorkloads with data-residency requirementsWorkloads without geographic restrictions that prioritize cost
PricingStandard pricing, based on the Region you call fromDescribed by AWS as the lower-cost option5
Organization policiesService control policies must allow every destination Region in the profilePolicies must allow requests whose requested Region is unspecified
Audit trailCloudTrail in the source Region records where each request was processedThe same; check the inference Region field in each event

Inference profiles cannot be combined with Provisioned Throughput5.

Request-level levers to pull before buying capacity

0 of 7 checked

Matching traffic shape to a capacity mode

  • If

    Interactive traffic is bursty and hard to predict.

    Then

    Stay on on-demand Standard, with a geographic cross-Region profile for headroom and retries with backoff.

    You pay only for use, and profiles add headroom without a commitment.

  • If

    One customer-facing path needs the fastest responses but not round-the-clock reserved capacity.

    Then

    Send that path's requests in the Priority tier and keep everything else on Standard.

    The tier is set per request, so only the latency-sensitive path pays the premium.

  • If

    Volume on one model is steady and high, and throttling would breach a service level.

    Then

    Measure peaks on demand first, then size Provisioned Throughput or a Reserved tier reservation.

    A commitment is cheaper and safer only when it is sized from observed demand.

  • If

    Large datasets are processed overnight or in bulk.

    Then

    Use batch inference, or the Flex tier where tool calling or structured output is needed.

    Nobody is waiting on these jobs, so lower priority costs nothing in user experience.

  • If

    Data must stay within one jurisdiction.

    Then

    Use in-Region inference or geographic profiles only, and block global profiles with service control policies.

    Global profiles may process requests in any supported commercial Region.

A hypothetical application that splits its traffic three ways

Questions and answers

How do we request a higher Bedrock quota?

Most on-demand quotas can be viewed and raised through the Service Quotas console or API for your account and Region, while some capacity, such as the Reserved tier and details of model units, goes through your AWS account team. Bring CloudWatch evidence of throttling and expected peaks, because requests backed by measured demand are easier to justify and to size.

Can we test Provisioned Throughput before committing to a term?

Yes. Provisioned Throughput can be bought with no commitment term and deleted at any time, which lets you load-test a model unit against your real traffic pattern. Billing is hourly until the provisioned model is deleted, so schedule the test, record throughput and latency, and remove it when you are done4.

Do cached prompt tokens count against our Bedrock quota?

No. Tokens read from the prompt cache do not count toward the tokens-per-minute and daily quotas, while tokens written to the cache do6. That makes caching a throughput lever as well as a cost lever for applications that resend long, stable prefixes such as system instructions, tool definitions or reference documents.

Which metrics show whether our capacity choice is working?

Watch input, output, cache-read and cache-write token counts, invocation latency and throttles per model in CloudWatch, along with the resolved service tier, which shows the tier that actually served each request. Compare peaks with your quota and any reservation, and review the trend monthly so commitments follow demand rather than the original forecast.

Sources

  1. Amazon Bedrock pricing — Amazon Web Services · checked 10 October 2026
  2. Service tiers for optimizing performance and cost — Amazon Web Services · checked 10 October 2026
  3. Process multiple prompts with batch inference — Amazon Web Services · checked 10 October 2026
  4. Increase model invocation capacity with Provisioned Throughput in Amazon Bedrock — Amazon Web Services · checked 10 October 2026
  5. Route model inference requests across AWS Regions with cross-Region inference — Amazon Web Services · checked 10 October 2026
  6. How tokens are counted in Amazon Bedrock — Amazon Web Services · checked 10 October 2026
  7. Prompt caching for faster model inference — Amazon Web Services · checked 10 October 2026
  8. Model lifecycle (Amazon Bedrock User Guide) — Amazon Web Services · checked 10 October 2026

More in AWS Bedrock

Back to AWS Bedrock

Next step

Share your Bedrock traffic pattern and get a capacity plan

Send recent CloudWatch token and throttling metrics, or a description of expected traffic. We will suggest which requests belong in which mode and what to measure before any commitment.

Plan Bedrock capacity