All topic guides

Monitoring and reliability

Build an observation path from metrics and logs to alarms and investigation. Then use the reliability section when your question involves recovery, fault testing, or behavior across services and regions. Workload-specific guides are grouped separately so you can choose an EC2, Lambda, container, database, or DynamoDB entry without reading an unrelated setup guide.

Start here: CloudWatch Standard vs Detailed Monitoring: Key Differences

Set up monitoring and alarms

Use the comparison to orient your monitoring question, then move to custom metrics, alarm conditions, or a dashboard integration. The Config guide concerns resource inventory and change tracking, which complements workload signals but serves a different task.

Investigate a workload

Pick the guide for the service and symptom you are investigating. Metric definitions, recovery alarms, and error-monitoring practices have different scopes; use them together when you need to connect an observed signal to an operational response.

Plan and test resilience

Begin with the recovery-pattern comparison or cross-service resilience guide, then explore fault injection and automated recovery testing. Multi-region active-active design has its own application and data decisions, so keep that architectural question distinct from a single alarm or retry policy.

Explore another topic · Browse all articles by date