DynamoDB CloudWatch Metrics: Capacity, Latency, and Throttling
Published on · Updated on
Corrected metric scope, dimensions, and statistics; replaced fixed capacity rules with cause-specific throttling guidance; clarified alarm behavior, TTL metadata, and Streams/Lambda monitoring.
Amazon DynamoDB publishes service metrics to Amazon CloudWatch in the AWS/DynamoDB namespace. They help answer different questions: how much capacity a workload uses, whether requests are throttled, and how long successful service operations take. For a useful view, match each metric to the capacity mode, table or index, and application objective you need to monitor.
Start with these signals:
- Capacity:
ConsumedReadCapacityUnitsandConsumedWriteCapacityUnits; compare with provisioned capacity only when the resource uses provisioned mode. - Throttling:
ThrottledRequests,ReadThrottleEvents, andWriteThrottleEvents, followed by the cause-specific throttle metrics. - Service latency:
SuccessfulRequestLatency, separated by operation. - TTL and stream consumers:
TimeToLiveDeletedItemCountfor TTL deletions, and Lambda metrics for stream-processing lag and failures.
Choose the right metric and dimensions
A CloudWatch metric is identified by its namespace, name, and complete set of dimensions. Find DynamoDB metrics under AWS/DynamoDB, then select the dimension combination for the resource you mean. A table series and a table-plus-index series are separate metrics; an alarm with the wrong dimensions will not match the intended series. See AWS's DynamoDB metrics and dimensions reference.
Read capacity metrics in context
ConsumedReadCapacityUnits and ConsumedWriteCapacityUnits report capacity units used over a selected period in both provisioned and on-demand modes. For a total over a period, use Sum. For example, divide the one-minute sum by 60 to estimate the average units consumed per second. That average can hide a short spike, so use it for trend and utilization analysis rather than as a precise record of moment-to-moment demand.
In provisioned mode, compare consumption with ProvisionedReadCapacityUnits and ProvisionedWriteCapacityUnits. These metrics describe configured throughput; their valid statistics include Average, Minimum, and Maximum, not a period total. Provisioned capacity metrics do not provide a baseline for on-demand tables. If an on-demand maximum throughput is configured, OnDemandMaxReadRequestUnits and OnDemandMaxWriteRequestUnits show that configured ceiling; they are not measures of consumed requests.
Check index capacity separately. For the base table, the TableName dimension returns table consumption. To inspect a global secondary index (GSI), include both TableName and GlobalSecondaryIndexName. A GSI has its own read and write capacity behavior, and an index write limit can throttle writes to the base table, even when the table itself has capacity. See AWS's guidance on GSI throughput and write back-pressure. Global-table write metrics can also include a Source dimension to distinguish customer writes from replicated writes.
Use capacity mode and cost data together when reviewing efficiency. No single utilization percentage determines whether provisioned or on-demand mode is cheaper: the decision depends on the workload's traffic pattern and current regional pricing. Auto Scaling can adjust provisioned capacity within its configured limits, but its changes take time to apply. It cannot fix a hot partition, a configured on-demand maximum, or an account quota. For the temporary headroom unused provisioned capacity may provide, see DynamoDB burst capacity and its limits.
Diagnose throttling by cause
ThrottledRequests counts requests, while ReadThrottleEvents and WriteThrottleEvents count throttled read or write events against the table or an index. Use the Sum statistic for counts. ThrottledRequests uses TableName and Operation dimensions. For read/write event series, add GlobalSecondaryIndexName with TableName to inspect a GSI; the table-only series does not include that index's separate events. One API request can create several events—for example, a write to a table and its GSIs—so the request metric may increase once while multiple event metrics increase. For batch reads and writes, ThrottledRequests increases only when every operation in the batch is throttled. These metrics are complementary; neither alone gives a complete picture.
After confirming a throttle, narrow down its cause with the matching metrics:
ReadProvisionedThroughputThrottleEventsandWriteProvisionedThroughputThrottleEventspoint to provisioned throughput limits.ReadKeyRangeThroughputThrottleEventsandWriteKeyRangeThroughputThrottleEventspoint to partition-level limits, often caused by a hot key or concentrated key range.ReadMaxOnDemandThroughputThrottleEventsandWriteMaxOnDemandThroughputThrottleEventspoint to a configured on-demand maximum.ReadAccountLimitThrottleEventsandWriteAccountLimitThrottleEventspoint to account-level limits.
Use the exact dimension set published for each cause-specific metric; provisioned, key-range, and on-demand-maximum event metrics can be separated by TableName and GlobalSecondaryIndexName. In application logs, retain the throttling exception's ThrottlingReasons and affected resource ARN. A reason such as IndexWriteKeyRangeThroughputExceeded identifies an index hot-key problem; increasing table capacity alone will not correct it. AWS describes the reason codes and corresponding CloudWatch throttling metrics in its diagnosis guide.
Choose the response from the cause: raise or autoscale provisioned capacity when that resource's configured throughput is the limit; adjust key design or access patterns when one key range is saturated; review the configured on-demand maximum or applicable service quota when those limits are named. On-demand mode and Auto Scaling do not remove every source of throttling. AWS's throttling diagnosis workflow maps each reason to its resolution.
CloudWatch Contributor Insights for DynamoDB can reveal the most accessed or throttled keys when aggregate metrics point to a hot key. It adds a monthly charge per rule and event-based charges; processing every read and write in its accessed-and-throttled mode can cost more than processing throttled events only. The feature publishes partition and sort key values to CloudWatch, where principals with the relevant permissions can view them. Avoid enabling it on tables whose keys contain sensitive data your policies do not allow in CloudWatch. AWS documents its billing and data handling.
Separate service latency from application latency
SuccessfulRequestLatency measures successful operations inside DynamoDB or DynamoDB Streams, in milliseconds. It excludes network time, client work, and time spent retrying a request, and it does not include failed requests. The AWS dimension reference lists TableName, Operation, and StreamLabel; use the published dimension combination for the operation you are monitoring. StreamLabel applies to Streams operations. Use the Average or a percentile such as p50 or p99 to study the latency distribution, and check SampleCount to see how many successful requests support the period.
This service metric is not end-to-end SDK latency. Measure request duration and retries in the application as well. For example, repeated SDK retries can make a call slow to the user while the successful DynamoDB operation remains fast. Compare both views, filtered to the same operation and time window, before attributing a slowdown.
Interpret errors and application conflicts
SystemErrors counts requests that returned HTTP 500 errors. AWS SDKs commonly retry these responses, so an individual error may appear in CloudWatch without reaching the application as a final failure; persistent errors or a rise in client latency deserve investigation. For a per-table alarm, include both the TableName and Operation dimensions.
UserErrors counts many HTTP 400 responses, but it is not a per-table measure of every application failure: it is aggregated by account and Region, and it excludes throttles and failed conditional writes. Use ConditionalCheckFailedRequests for conditional writes and TransactionConflict for rejected item-level operations caused by concurrent transactions. Treat these as signals to investigate against application behavior, not as interchangeable error totals.
Monitor TTL and stream consumers
TimeToLiveDeletedItemCount reports the number of items deleted by TTL during a period. Use its Sum statistic to track deletion volume. TTL is asynchronous: expired items can remain in the table for a few days before the background process deletes them, so this metric is not an immediate count of expired-but-pending items. For approximate inventory, DescribeTable returns ItemCount and TableSizeBytes; DynamoDB refreshes both values about every six hours, so they are not live traffic metrics.
If you process TTL deletions through DynamoDB Streams, enable Streams and choose a stream view that includes the old item when your consumer needs its contents. A keys-only view cannot supply the deleted item's prior attributes. TTL deletions are identifiable as service deletions in the Region where deletion occurs; replicated deletes in other global-table Regions do not carry the same identity marker. See AWS's guide to DynamoDB TTL records in Streams.
For stream API activity, ReturnedRecordsCount reports records returned by DynamoDB Streams GetRecords calls. Use its Sum statistic and the Operation, StreamLabel, and TableName dimensions. Streams operations also use SuccessfulRequestLatency with those dimensions. These DynamoDB metrics describe stream reads, not whether a Lambda consumer is keeping up. In the AWS/Lambda namespace, monitor IteratorAge for record delay and Errors and Throttles for function failures and concurrency limits; filter by event source mapping UUID when needed. You can also opt in to event source mapping metrics such as PolledEventCount, FailedInvokeEventCount, and DroppedEventCount. Counts can include repeated events when Lambda retries a batch. See AWS's Lambda metrics reference.
Set CloudWatch alarms for workload goals
Pick a metric, statistic, and threshold that correspond to an operational response. A utilization alarm can warn that provisioned headroom is shrinking; it does not prove users are affected. For workloads where any throttle matters, alarm on throttle events. Otherwise, set the threshold and evaluation window from the application's latency or completion objective. A fixed 80% utilization rule is not suitable for every table, capacity mode, or traffic shape.
CloudWatch evaluates a metric alarm over periods. For example, a one-minute, two-out-of-three alarm needs two breaching datapoints among the last three; they do not have to be consecutive. Match the period to the metric's publication cadence and the time available to respond. Also account for a DynamoDB-specific behavior: alarms on the AWS/DynamoDB namespace always ignore missing data, regardless of the configured missing-data treatment, and remain in their current state when data is absent. Use a separate signal if missing telemetry itself needs to alert you. For a broader guide to CloudWatch alarm thresholds, periods, and conditions, see the site's alarm guide.
Conclusion
Use DynamoDB CloudWatch metrics to answer a specific question, then verify the result at the right level: capacity mode, table or GSI, operation, throttle cause, and application outcome. Compare request-level throttles with event-level and cause-specific metrics, and compare DynamoDB's successful service latency with latency measured by your application. Alarms become more useful when their thresholds reflect a workload objective and their dimensions identify the affected resource.