DynamoDB burst capacity is a temporary buffer that can help with brief increases in read or write traffic. For tables and indexes using provisioned capacity, DynamoDB reserves some of the unused throughput for a later spike. AWS documents retention of up to five minutes (300 seconds), but the buffer is not guaranteed: DynamoDB may also use it for background maintenance. The five-minute detail below describes unused provisioned throughput. AWS also says an on-demand table can sometimes exceed a configured maximum through burst capacity; that separate case is covered below. Treat burst capacity as extra breathing room, not as capacity you can count on. See AWS documentation on burst and adaptive capacity.

Key takeaways:

  • In provisioned mode, burst capacity uses some recently unused read or write throughput; on-demand tables with a configured maximum can sometimes exceed it through burst capacity.
  • Five minutes is the documented maximum retention window, not a promise that a full five minutes of throughput is available.
  • CloudWatch shows capacity use and throttling, but it does not report a remaining burst-capacity balance.
  • Partition limits, table or index capacity, account quotas, and configured limits can still cause throttling.
  • Burst capacity does not lower the cost of provisioned throughput or replace a suitable baseline.

How DynamoDB burst capacity works

How unused capacity builds up

In provisioned mode, a table or global secondary index has configured read and write capacity. When actual use falls below that capacity, DynamoDB reserves a portion of the unused throughput for possible later bursts. Read and write throughput are separate, so unused read capacity does not provide write capacity, or vice versa.

AWS describes the reserve as a portion of unused capacity. It does not publish a formula that lets you calculate an exact balance from a table's provisioned and consumed capacity. Available burst capacity can also change as DynamoDB performs background work.

The five-minute retention window

DynamoDB retains up to five minutes (300 seconds) of unused read and write capacity. This is an upper bound on how far back unused throughput may be retained, not a guarantee that every idle second adds a fixed number of usable units. A spike can consume the available buffer faster than the configured per-second provisioned rate would suggest.

What happens during a traffic spike

If burst capacity is available, some requests above the provisioned rate may succeed temporarily. As the buffer is used, or if it is unavailable, requests that exceed a limit can be throttled. A short spike may fit inside the available headroom; sustained demand needs enough provisioned capacity or another capacity strategy.

Tables and global secondary indexes have distinct capacity settings in provisioned mode. Check which resource is affected when throttling occurs: for example, insufficient GSI write capacity can also cause writes to the base table to be throttled.

Monitoring capacity and throttling

CloudWatch helps you see usage patterns and identify throttling, but there is no standard CloudWatch metric for the amount of burst capacity remaining. Comparing consumed and provisioned capacity can show whether a resource is running close to its baseline; it cannot reveal an exact burst balance.

Metrics to use

CloudWatch metric What it tells you How to use it
ConsumedReadCapacityUnits and ConsumedWriteCapacityUnits Read and write capacity consumed during the selected period. Compare trends with provisioned capacity to understand normal load and spikes.
ProvisionedReadCapacityUnits and ProvisionedWriteCapacityUnits Configured read and write capacity for a table or index. Use the right dimensions to inspect the table or a specific global secondary index.
ThrottledRequests Requests that contain at least one throttled event. Use it as a request-level signal; one request can include multiple throttled events.
ReadThrottleEvents, WriteThrottleEvents, ReadKeyRangeThroughputThrottleEvents, and WriteKeyRangeThroughputThrottleEvents Read, write, and partition-level throttling events. Use these with the throttling reason to narrow down the resource and limit involved.

Compare these metrics using the same resource dimensions. A Sum for ConsumedReadCapacityUnits or ConsumedWriteCapacityUnits is the total units used during the CloudWatch period, while provisioned capacity is expressed per second. Divide the consumed Sum by the period length in seconds—for example, 60 for a one-minute period—to get the average consumed units per second. When checking a global secondary index, include both TableName and GlobalSecondaryIndexName; the table dimension alone does not include that index. This average can hide a brief spike inside the period. See AWS's DynamoDB metrics and dimensions guide for metric definitions.

For a fuller metric list and their meanings, see AWS's DynamoDB CloudWatch throttling metrics guide. Set alarms around the impact you need to control, such as recurring throttling or a sustained capacity trend. Choose thresholds from your workload and latency goals rather than treating one utilization percentage as a universal rule.

Finding the cause of throttling

Start with the throttling exception and its reason, then check the matching CloudWatch metric for the affected table or index. Table-level capacity can look healthy while a hot partition is throttled. For partition-related throttling, CloudWatch Contributor Insights for DynamoDB can identify frequently accessed or throttled keys. It is optional and has separate CloudWatch charges; if key values may be sensitive, review how Contributor Insights publishes them before enabling it.

Setting up useful monitoring

  1. Graph consumed and provisioned read and write capacity for each table and global secondary index that matters to the workload.
  2. Alarm on throttling signals that affect your application's latency or completion goals. Use the appropriate metric dimensions so the alarm monitors the intended resource.
  3. When throttling occurs, inspect its reason and compare request-level metrics with read, write, or key-range events before changing capacity.

Using burst capacity in capacity planning

Plan for steady and expected load

Set the provisioned baseline to support the traffic your application expects to sustain. Burst capacity may absorb a brief deviation, but it can be drained and cannot protect against a long plateau above provisioned throughput. DynamoDB Auto Scaling can adjust provisioned capacity as load changes, but it takes time to react; configure its target and capacity bounds with enough headroom for your workload.

Provisioned tables are billed for capacity provisioned, not just the capacity consumed. Burst capacity is not a cost-saving pool that makes it safe to provision below sustained demand. If traffic is unpredictable, compare provisioned mode with on-demand capacity mode. On-demand mode removes the need to set provisioned read and write units, but it has its own scaling behavior and can still throttle, including when traffic rises to more than twice its previous peak within 30 minutes or when another limit is reached. You can also set a maximum throughput for an on-demand table or index. AWS treats that maximum as a best-effort target, and burst capacity may let traffic exceed it temporarily; see the on-demand maximum throughput guidance.

Common mistakes to avoid

  • Do not assume all unused provisioned capacity becomes available as burst capacity or that a full five-minute reserve is waiting.
  • Do not infer a remaining burst balance from CloudWatch consumed and provisioned metrics.
  • Do not increase table capacity blindly when the throttling reason points to a hot partition, index, account quota, or configured maximum.
  • Use bounded retries with backoff to handle transient throttling, but do not treat retries as a substitute for addressing sustained demand or a hot key.

For broader causes and mitigations beyond short capacity spikes, see our DynamoDB throttling guide.

Burst capacity limits and alternatives

Partition and table limits still apply

Burst capacity is only one part of DynamoDB's capacity behavior. A hot partition can be throttled even when other parts of the table have unused capacity. DynamoDB adaptive capacity helps with uneven access patterns, but it does not remove the maximum throughput of a partition or let a workload exceed the table's total provisioned capacity. Read and write throttling can also come from an index or account-level limit. See AWS guidance on partition-level throttling and provisioned-throughput throttling.

Choose a capacity mode for the workload

Use provisioned capacity when demand is predictable enough to set a cost-effective baseline, with Auto Scaling where appropriate. Consider on-demand capacity when request volume varies sharply or is difficult to forecast. Neither mode removes partition-level hot spots, account quotas, or configured throughput limits, so diagnose the throttling reason before choosing a remedy.

Summary

DynamoDB burst capacity can help handle brief read or write spikes. In provisioned mode, it uses some recently unused throughput; on-demand tables with a configured maximum can also exceed that target temporarily in some cases. AWS documents retention of up to 300 seconds for unused capacity, but DynamoDB can consume the reserve for background work and does not expose its remaining balance through CloudWatch.

Monitor consumed, provisioned, and throttling metrics to understand what your workload is doing. When throttling appears, identify whether the cause is provisioned capacity, an index, a hot partition, or another limit. Keep enough baseline capacity for sustained demand, and choose Auto Scaling or on-demand capacity based on how your traffic behaves.