HPC Monitoring on AWS: Metrics, Dashboards, and Alarms
Published on · Updated on
Rewrote the guide for current AWS metric coverage, scheduler signals, and cost-aware monitoring; removed unsupported savings claims and thresholds.
HPC monitoring on AWS should answer three questions: is the cluster healthy, are jobs progressing as expected, and what is limiting a representative run? CloudWatch provides service metrics, logs, dashboards, and alarms, but no single default feed shows instance health, scheduler state, GPU use, filesystem pressure, and MPI communication together.
Build visibility in layers. Start with AWS resource metrics, add host and scheduler data where needed, then use application profiling to explain how a job spends its time. Tie each signal to a decision or response so a dashboard stays useful as compute fleets scale up and down.
What CloudWatch shows by default
AWS services publish metrics for their own resources. For example, the AWS/EC2 namespace includes CPU, network, status, and storage I/O measurements; the exact metrics depend on the instance and attached resources. On a subset of accelerated instance types, it also exposes GPUPowerUtilization, a power-use measurement rather than a measure of GPU compute busy time. EC2 basic monitoring generally publishes five-minute datapoints. Enabling detailed monitoring changes supported EC2 metrics to one-minute datapoints and adds cost. Neither setting measures guest memory, filesystem capacity, scheduler queues, or application communication time by itself. Check the EC2 metrics list and CloudWatch monitoring intervals for the resources you use.
Install and configure the CloudWatch agent on nodes when you need in-guest metrics such as memory and filesystem use, or want to ship system and application logs. The agent only publishes the measurements selected in its configuration; it does not turn every host or workload signal on automatically. Choose a collection interval that fits the question, and include cluster, queue, and instance dimensions that let operators find the affected resources.
| Signal | Where to collect it | What it tells you | Scope to remember |
|---|---|---|---|
| Instance CPU, network traffic, status, and supported I/O | EC2 metrics in CloudWatch | Whether an instance or fleet is healthy or busy | Instance-level usage is not the same as job or rank utilization; some accelerators also expose power-use metrics. |
| Guest memory, filesystem use, selected process or log data | CloudWatch agent on Linux or Windows nodes | Whether host resources are constrained and which logs explain node issues | Install and configure the agent on the nodes that need coverage, including short-lived compute nodes. |
| NVIDIA GPU compute utilization and device memory | CloudWatch agent GPU collection on Linux nodes | GPU compute use, memory use, temperature, and related device measurements | GPU collection must be configured and the NVIDIA driver must be installed. EC2's limited GPUPowerUtilization metric on supported types measures power use, not compute utilization. |
| EFA traffic and device counters | CloudWatch agent EFA collection on Linux nodes with EFA devices | Per-device traffic, drops, retransmissions, and RDMA operation counters | These counters describe EFA devices, not per-job MPI time or application latency. |
| Filesystem I/O, capacity, and metadata/server health | FSx for Lustre metrics in AWS/FSx | Whether the shared filesystem or a target is under pressure | Metrics use file-system and, for some measurements, storage-target dimensions; they do not identify the job that caused activity. |
Collect the signals the workload needs
GPU and EFA metrics need explicit collection
The CloudWatch agent can collect supported NVIDIA GPU metrics when an NVIDIA driver is installed and GPU collection is enabled. These include GPU compute and memory measurements that the EC2 GPUPowerUtilization metric does not provide. The agent also collects EFA device metrics when configured on Linux servers with EFA devices. For EFA, counters such as efa_rx_dropped and retransmission events can help identify transport problems. By contrast, EC2's NetworkPacketsOut is an instance traffic count; it does not report EFA packet drops or prove that an MPI collective is slow.
Do not treat device activity as a complete explanation of a distributed job. EFA counters do not say how much time an application spent waiting in MPI calls, and high GPU utilization does not show whether useful work is completing. Pair node metrics with logs and a profile from a representative application run. See the site's HPC workload profiling guide for ways to investigate CPU, GPU, and communication bottlenecks.
Use FSx for Lustre metrics to investigate storage
Amazon FSx for Lustre publishes file-system metrics to CloudWatch's AWS/FSx namespace. The service emits usage metrics every minute, including network I/O, capacity, and metrics for storage and metadata targets. Use the dimensions and statistics documented in the FSx for Lustre metrics reference to compare aggregate and per-target behavior. A high data rate alone does not prove that storage is a bottleneck; compare the observed pattern with the job's read/write phases, filesystem configuration, and application profile.
Monitor scheduler and job progress
AWS ParallelCluster with Slurm
AWS ParallelCluster provides CloudWatch Logs integration, a cluster dashboard, and version-dependent cluster alarms. Its dashboard can show head-node metrics and logs, shared storage metrics, and Slurm cluster-health information. Current ParallelCluster guidance describes default head-node alarms for health, CPU, memory, and root-disk use in supported versions. These defaults are not a complete compute-node or job monitor, and ParallelCluster does not configure alarm actions for you. Check the dashboard, alarm, and monitoring configuration docs for your ParallelCluster version and settings.
For Slurm queue and job-level visibility, use scheduler data such as squeue for current state and sacct for completed-job accounting when Slurm accounting is configured. ParallelCluster supports Slurm accounting through slurmdbd and a MySQL-compatible database; you must configure and operate that accounting path. If operators need queue wait, completion, failure, or resource-use trends in CloudWatch, publish selected aggregates from scheduler/accounting data or your job hooks. Keep job identifiers in logs or an accounting store when they would create too many unique metric dimensions. The ParallelCluster Slurm accounting guide explains the supported setup.
AWS Batch jobs
AWS Batch publishes job lifecycle and duration metrics in the AWS/Batch namespace for jobs submitted to a job queue. They are useful for comparing submitted, runnable, running, succeeded, and failed work over time. AWS describes these metrics as best-effort monitoring signals: a fast job transition can be missed, so do not use them as billing, audit, or exact reconciliation counts. For compute-environment and job CPU, memory, disk, and network measurements, consider CloudWatch Container Insights; AWS bills its metrics as custom metrics. The AWS Batch metrics guide documents the signals and limitations.
Design dashboards around diagnosis
Give each dashboard a clear path from cluster health to the job or component that needs investigation. A practical HPC view can include:
- Fleet health: node state, EC2 status checks, CPU, memory, filesystem use, and provisioning or Slurm health errors.
- Job progress: queue age, running and completed jobs, failures, elapsed time, and resource use by queue or workload class.
- Data path: FSx throughput, operations and capacity alongside EFA traffic, drops, retransmissions, and the job's input/output phases.
- Application behavior: representative run time, useful work completed, and profiler results for the code path under study.
Label units, dimensions, statistics, and time periods so that a node-level metric cannot be mistaken for a job-wide result. Add dashboards for short operational windows and for longer run-to-run comparisons when the workload needs both. Amazon Managed Grafana can query CloudWatch for cross-cluster views; see the site's CloudWatch and Grafana setup guide for data-source and dashboard setup.
Set alarms that lead to an action
Alarm on symptoms that an operator can act on: repeated node-health failures, a sustained resource limit, a growing queue wait, a missing ParallelCluster cluster-management heartbeat where your version publishes one, an FSx capacity or performance condition, or an EFA error counter that departs from the workload's baseline. Set periods and thresholds from normal runs and service objectives; the correct values depend on the application, cluster size, and expected queue behavior. Choose how missing data should be treated, and avoid triggering on a single harmless spike when a sustained condition is what matters.
For each alarm, define its owner and response. Route notifications through the mechanism your team uses, and configure alarm actions explicitly; an alarm without a response path is only a graph state. CloudWatch alarm patterns are covered in the site's AWS monitoring guide and AWS's CloudWatch alarms documentation.
Control telemetry cost and access
Before increasing collection frequency, decide what question the extra resolution will answer. Detailed monitoring, custom metrics, high-resolution data, CloudWatch API reads, and log ingestion and storage can add charges. Use only the measurements and dimensions you need, especially when many nodes or per-job labels could multiply metric counts. Review charges by CloudWatch usage type and update retention settings to match the operational and audit needs for each log group. AWS maintains current cost guidance in Analyzing and reducing CloudWatch costs.
Use AWS Budgets for account, service, or tagged cost notifications, and assign an owner to investigate HPC spend. Budgets are useful for planning and follow-up, not immediate runtime control: AWS updates budget data up to three times per day, and usage-to-billing delays can postpone a notification. Keep that lag in mind when deciding what should stop or limit a job. For ongoing AWS API and configuration audit, use CloudTrail and AWS Config for their respective control-plane and resource-configuration signals; they do not replace host, scheduler, or application telemetry. Restrict access to monitoring data, then set log encryption and retention according to your data classification and audit requirements.
A practical monitoring checklist
- Choose one representative workload and record a baseline for queue wait, run time, completed work, and relevant compute, network, and storage use.
- Confirm which service metrics already exist, then configure the CloudWatch agent only for missing host, GPU, or EFA signals on the nodes that need them.
- Decide where scheduler accounting and job logs live, and preserve the data long enough to diagnose failures after short-lived compute nodes terminate.
- Build dashboards around cluster, scheduler, storage, and application questions; set alarm thresholds from observed workload behavior.
- Review CloudWatch and compute costs after enabling new metrics, logs, or higher collection rates.
A useful HPC monitoring system combines infrastructure health, scheduler outcomes, storage and network evidence, and application profiling. CloudWatch can bring many of those signals together, provided each metric's scope and collection method are clear. Revisit the setup when the scheduler, node types, application, or cluster scale changes.