Monitoring and reliability
Build an observation path from metrics and logs to alarms and investigation. Then use the reliability section when your question involves recovery, fault testing, or behavior across services and regions. Workload-specific guides are grouped separately so you can choose an EC2, Lambda, container, database, or DynamoDB entry without reading an unrelated setup guide.
Start here: CloudWatch Standard vs Detailed Monitoring: Key Differences
Set up monitoring and alarms
Use the comparison to orient your monitoring question, then move to custom metrics, alarm conditions, or a dashboard integration. The Config guide concerns resource inventory and change tracking, which complements workload signals but serves a different task.
- CloudWatch Standard vs Detailed Monitoring: Key Differences
- Custom CloudWatch Metrics: Setup Guide 2024
- CloudWatch Alarms: Best Practices for Thresholds & Conditions
- AWS CloudWatch + Grafana: Setup Guide
- AWS LogicMonitor Integration: 7 Best Practices
- AWS DevOps Monitoring: Metrics, Tools, Best Practices
- AWS Health Dashboards: Monitoring & Debugging Guide
- AWS Config: Resource Inventory, Change Tracking Guide
Investigate a workload
Pick the guide for the service and symptom you are investigating. Metric definitions, recovery alarms, and error-monitoring practices have different scopes; use them together when you need to connect an observed signal to an operational response.
- How to Monitor EC2 with CloudWatch
- CloudWatch Alarms: Automate EC2 Recovery
- AWS Fargate Metrics with CloudWatch
- AWS Lambda Metrics Explained
- AWS Lambda Insights: Serverless Monitoring Guide
- Monitor Lambda Errors with CloudWatch
- Key CloudWatch Metrics for DynamoDB Performance
- AWS Performance Insights: Monitoring RDS Databases
Plan and test resilience
Begin with the recovery-pattern comparison or cross-service resilience guide, then explore fault injection and automated recovery testing. Multi-region active-active design has its own application and data decisions, so keep that architectural question distinct from a single alarm or retry policy.