Monitor Lambda Errors with CloudWatch
Published on · Updated on
Rebuilt the walkthrough around valid Lambda metrics, structured and plain-text Logs Insights queries, a runnable metric-math error-rate alarm, correct logging permissions, and deliberate missing-data handling.
CloudWatch can show whether a Lambda function is failing, throttled, slow, or building a backlog. Use the AWS/Lambda metrics to detect changes, then inspect the function's CloudWatch Logs to find the invocation or application message behind them. This walkthrough creates an error-rate alarm and runs focused Logs Insights queries.
The examples use one function named orders-api in us-east-1. Replace the function, Region, and notification topic with values from your account. CloudWatch Lambda metrics are emitted automatically; the function's execution role needs logging permissions only if it sends logs to CloudWatch Logs.
Key steps: metrics, logs, and an alarm
- Open the CloudWatch metrics for the function and graph
Errors,Invocations, andThrottlesusing theSumstatistic. - Open
/aws/lambda/orders-apiin CloudWatch Logs and use Logs Insights to inspect recent error records. - Create a metric-math alarm from
Errorsdivided byInvocations, with a threshold and missing-data choice that fit the function's traffic. - Check alarm history and a test notification so the configured action reaches its intended owner.
Set up Lambda metrics and permissions
Open the function's CloudWatch metrics
- In the AWS console, select the Region where the Lambda function runs.
- Open CloudWatch, choose Metrics, then select the
AWS/Lambdanamespace and the view grouped by function name. - Select
Errors,Invocations, andThrottlesfororders-api. AddDurationif execution time is part of the issue. - Use
Sumfor counts. Compare errors and invocations over the same time range; use a period that gives enough datapoints for the traffic you have.
Errors counts invocations that returned a function or runtime error, including exceptions and timeouts. It does not count handled errors when the handler returns normally, and Throttles is a separate metric. Use Invocations as the denominator for an approximate function error rate. The two metrics are published at one-minute intervals, although a long invocation can take several minutes to appear. The FunctionName dimension aggregates versions and aliases; choose a resource or executed-version dimension when an alarm needs to isolate a rollout. See AWS's Lambda metric definitions, its guide to viewing metrics by dimension, or our guide to Lambda metric dimensions and statistics.
Lambda publishes these standard service metrics without extra permissions on the function's execution role. To view them in the console, the IAM identity you use must have permission to read CloudWatch metrics and alarms.
Create an error-rate alarm
A single alarm on Errors is useful when any failed invocation deserves attention. A rate alarm is useful when the proportion of failed calls matters more than the raw count. The following AWS CLI command creates an example rate alarm for orders-api. It evaluates five-minute sums, requires two of the last three periods to reach five percent, and calculates a rate only when a period includes at least 20 invocations.
aws cloudwatch put-metric-alarm \
--region "us-east-1" \
--alarm-name "orders-api-error-rate" \
--evaluation-periods 3 \
--datapoints-to-alarm 2 \
--threshold 5 \
--comparison-operator GreaterThanOrEqualToThreshold \
--treat-missing-data notBreaching \
--alarm-actions "arn:aws:sns:us-east-1:123456789012:service-alerts" \
--metrics '[
{
"Id": "errors",
"MetricStat": {
"Metric": {
"Namespace": "AWS/Lambda",
"MetricName": "Errors",
"Dimensions": [{"Name": "FunctionName", "Value": "orders-api"}]
},
"Period": 300,
"Stat": "Sum"
},
"ReturnData": false
},
{
"Id": "invocations",
"MetricStat": {
"Metric": {
"Namespace": "AWS/Lambda",
"MetricName": "Invocations",
"Dimensions": [{"Name": "FunctionName", "Value": "orders-api"}]
},
"Period": 300,
"Stat": "Sum"
},
"ReturnData": false
},
{
"Id": "error_rate",
"Expression": "IF(invocations >= 20, 100 * errors / invocations)",
"Label": "Error rate (%)",
"ReturnData": true
}
]'
Create an SNS topic and confirm its subscriptions before running the command, then replace the example topic ARN with that topic in the same Region and account. The 20-invocation gate and five-percent threshold are illustrative settings, not recommended defaults. Choose them from the function's traffic and error budget. With fewer than 20 invocations, this expression has no rate datapoint; if any error at low traffic matters, add a separate Errors count alarm.
This example treats missing rate data as not breaching, which avoids treating an idle period as an error-rate incident. It also means missing telemetry will not page through this alarm. If the function should be active, monitor expected invocations or a heartbeat separately. CloudWatch alarms can be in OK, ALARM, or INSUFFICIENT_DATA; check the graph and history after creation, and test the notification path. Learn more about metric-math alarms and alarm periods and missing-data choices.
Find the invocation details in CloudWatch Logs
By default, Lambda sends function logs to a log group named /aws/lambda/<function-name>, such as /aws/lambda/orders-api. The group is created when the function first runs if it does not already exist. In CloudWatch, choose Logs, then Log groups, open the function's group, and select a recent stream. Use the same Region and a time range around the metric increase.
The execution role needs CloudWatch Logs write permissions. The AWS managed policy AWSLambdaBasicExecutionRole includes the common permissions to create a log group and stream and put log events. If a policy is scoped to an existing custom log group, check that it still allows the required log-event writes. These role permissions control log delivery; they do not grant the signed-in operator permission to run Logs Insights queries. See AWS's guidance for Lambda and CloudWatch Logs.
Query errors and slow invocations
Lambda's default log format is plain text. To enable JSON, open the function in the Lambda console, choose Configuration > Monitoring and operations tools, edit Logging configuration, set Log format to JSON, and save. This changes new system logs; application logs become structured when the runtime and logger support the format. For example, supported Node.js runtimes use the built-in console methods and add fields such as level, message, and requestId. This complete test handler logs and rethrows only when an input event sets simulateFailure to true; do not leave the test flag in a production request path. See AWS's log-format and runtime requirements.
export const handler = async (event) => {
try {
if (event.simulateFailure) {
throw new Error("Simulated dependency failure");
}
return { ok: true };
} catch (error) {
console.error("Dependency request failed", error);
throw error;
}
};
Configure the function's log format as JSON, then invoke it once with {"simulateFailure": true} to see the error fields. A caught error that is handled and returned normally will not increment Lambda's Errors metric; rethrow only when the invocation should fail.
Run this Logs Insights query against the function's log group to inspect recent structured Node.js error records:
fields @timestamp, requestId, level, message, errorType, errorMessage
| filter level = "ERROR"
| sort @timestamp desc
| limit 50
For existing plain-text logs, start with a broader search and adjust the terms to the messages your runtime and application actually emit:
fields @timestamp, @logStream, @message
| filter @message like /(?i)error|timed out/
| sort @timestamp desc
| limit 50
With JSON system logs enabled, Lambda writes report events as platform.report records with duration under record.metrics.durationMs. Run this query for a five-minute duration trend:
filter type = "platform.report"
| stats count(*) as invocations,
avg(record.metrics.durationMs) as avgDurationMs,
max(record.metrics.durationMs) as maxDurationMs
by bin(5m) as period
| sort period desc
If the function still uses plain-text system logs, query the standard discovered fields instead:
filter @type = "REPORT"
| stats count(*) as invocations,
avg(@duration) as avgDurationMs,
max(@duration) as maxDurationMs
by bin(5m) as period
| sort period desc
Both duration queries count completed invocation report records; they help analyze execution time but do not replace the Errors metric. Error text varies by runtime and handler. A log query that counts messages can overcount when one invocation writes several errors, and can miss failures that use different text. AWS documents the platform report event schema, Lambda Logs Insights query examples, and structured logging.
Set retention and keep queries focused
CloudWatch Logs keeps events indefinitely by default unless a retention period is set. Choose retention for the function's troubleshooting, audit, and compliance needs rather than applying a fixed duration to every environment. Logs incur ingestion and storage charges, and Logs Insights charges can depend on the volume scanned. Limit the time range and selected log groups, and avoid logging secrets or full event payloads.
Lambda uses /aws/lambda/<function-name> by default. Choose a custom log group only when it supports a real ownership or retention need, and update the execution-role policy if the function is restricted to that group. For current logging destinations and formats, use AWS's Lambda logging guide.
Add Lambda Insights for runtime details
Standard Lambda metrics show invocations, function errors, duration, and concurrency; they do not provide an actual-memory-use time series. CloudWatch Lambda Insights can add execution-environment and resource telemetry, including memory information. It is an optional extension with its own compatibility, IAM, log, and cost considerations. See our guide to enabling Lambda Insights and interpreting its metrics.

Enable it for a function
- Open the function in the Lambda console and choose Configuration.
- Under Monitoring and operations tools, edit the additional monitoring tools and enable CloudWatch Lambda Insights.
- Confirm the execution role has
CloudWatchLambdaInsightsExecutionRolePolicy. The console can attach it if your IAM permissions allow. - Invoke the function, then open CloudWatch Insights > Lambda Insights in the same Region.
Check AWS's current runtime, architecture, and extension requirements before enabling it. A dashboard shows only functions that have the extension configured and have emitted telemetry.
Compare only enabled functions
The multi-function view helps compare telemetry across functions that have Lambda Insights enabled. Use a single-function view to inspect one function's runtime signals, then return to its application log group for error messages. Insights does not identify every business failure and does not replace the function's standard CloudWatch metrics.
Use the signals in context
CloudWatch's Errors, Invocations, and Throttles answer different questions. A handler that catches an exception and returns normally can produce an error log without incrementing Errors. Conversely, a timeout can increment Errors even when the application did not write a final log line. For trigger-specific retries, destinations, and backlog signals, see the operational Lambda monitoring guide.
Review trends and correlate request IDs
Compare the metric window with Logs Insights results for the same function, Region, and time range. Use requestId to follow one invocation through application messages and Lambda platform logs. For an event source mapping, also inspect the source queue or stream: a low function error count does not prove that messages are being delivered or processed on time.
On a release, review the affected function version or alias rather than relying only on an aggregate function graph. Check alarm history before changing its threshold, and confirm whether a rise came from function errors, throttles, or source delivery behavior.
Measure handled business failures explicitly
If a handler catches an exception and intentionally returns success, Lambda's Errors metric will not count that outcome. Emit a custom metric for a business failure that needs an alarm, or log a structured field that you can query. Keep the metric's dimensions bounded, such as operation or error category; adding a unique request ID as a metric dimension creates a separate time series for each value.
Use a rate alarm only when a rate expresses the service risk. Add a count alarm for important low-volume failures, and choose missing-data behavior based on whether idle periods are normal. Avoid fixed rules such as a universal one-percent error rate or memory threshold; set conditions from the function's traffic, service objective, and response plan.
AWS documentation references
- Lambda CloudWatch metric definitions
- CloudWatch Logs Insights sample queries
- CloudWatch alarm missing-data behavior
Conclusion
Start with the function's standard metrics, use Logs Insights to inspect specific invocations, and add an error-rate alarm only when its denominator, threshold, evaluation window, and missing-data behavior match the workload. Keep throttles and source backlogs visible as separate signals, and measure handled application failures explicitly when they matter.