HPC monitoring on AWS should answer three questions: is the cluster healthy, are jobs progressing as expected, and what is limiting a representative run? CloudWatch provides service metrics, logs, dashboards, and alarms, but no single default feed shows instance health, scheduler state, GPU use, filesystem pressure, and MPI communication together.

Build visibility in layers. Start with AWS resource metrics, add host and scheduler data where needed, then use application profiling to explain how a job spends its time. Tie each signal to a decision or response so a dashboard stays useful as compute fleets scale up and down.

What CloudWatch shows by default

AWS services publish metrics for their own resources. For example, the AWS/EC2 namespace includes CPU, network, status, and storage I/O measurements; the exact metrics depend on the instance and attached resources. On a subset of accelerated instance types, it also exposes GPUPowerUtilization, a power-use measurement rather than a measure of GPU compute busy time. EC2 basic monitoring generally publishes five-minute datapoints. Enabling detailed monitoring changes supported EC2 metrics to one-minute datapoints and adds cost. Neither setting measures guest memory, filesystem capacity, scheduler queues, or application communication time by itself. Check the EC2 metrics list and CloudWatch monitoring intervals for the resources you use.

Install and configure the CloudWatch agent on nodes when you need in-guest metrics such as memory and filesystem use, or want to ship system and application logs. The agent only publishes the measurements selected in its configuration; it does not turn every host or workload signal on automatically. Choose a collection interval that fits the question, and include cluster, queue, and instance dimensions that let operators find the affected resources.

SignalWhere to collect itWhat it tells youScope to remember
Instance CPU, network traffic, status, and supported I/OEC2 metrics in CloudWatchWhether an instance or fleet is healthy or busyInstance-level usage is not the same as job or rank utilization; some accelerators also expose power-use metrics.
Guest memory, filesystem use, selected process or log dataCloudWatch agent on Linux or Windows nodesWhether host resources are constrained and which logs explain node issuesInstall and configure the agent on the nodes that need coverage, including short-lived compute nodes.
NVIDIA GPU compute utilization and device memoryCloudWatch agent GPU collection on Linux nodesGPU compute use, memory use, temperature, and related device measurementsGPU collection must be configured and the NVIDIA driver must be installed. EC2's limited GPUPowerUtilization metric on supported types measures power use, not compute utilization.
EFA traffic and device countersCloudWatch agent EFA collection on Linux nodes with EFA devicesPer-device traffic, drops, retransmissions, and RDMA operation countersThese counters describe EFA devices, not per-job MPI time or application latency.
Filesystem I/O, capacity, and metadata/server healthFSx for Lustre metrics in AWS/FSxWhether the shared filesystem or a target is under pressureMetrics use file-system and, for some measurements, storage-target dimensions; they do not identify the job that caused activity.

Collect the signals the workload needs

GPU and EFA metrics need explicit collection

The CloudWatch agent can collect supported NVIDIA GPU metrics when an NVIDIA driver is installed and GPU collection is enabled. These include GPU compute and memory measurements that the EC2 GPUPowerUtilization metric does not provide. The agent also collects EFA device metrics when configured on Linux servers with EFA devices. For EFA, counters such as efa_rx_dropped and retransmission events can help identify transport problems. By contrast, EC2's NetworkPacketsOut is an instance traffic count; it does not report EFA packet drops or prove that an MPI collective is slow.

Do not treat device activity as a complete explanation of a distributed job. EFA counters do not say how much time an application spent waiting in MPI calls, and high GPU utilization does not show whether useful work is completing. Pair node metrics with logs and a profile from a representative application run. See the site's HPC workload profiling guide for ways to investigate CPU, GPU, and communication bottlenecks.

Use FSx for Lustre metrics to investigate storage

Amazon FSx for Lustre publishes file-system metrics to CloudWatch's AWS/FSx namespace. The service emits usage metrics every minute, including network I/O, capacity, and metrics for storage and metadata targets. Use the dimensions and statistics documented in the FSx for Lustre metrics reference to compare aggregate and per-target behavior. A high data rate alone does not prove that storage is a bottleneck; compare the observed pattern with the job's read/write phases, filesystem configuration, and application profile.

Monitor scheduler and job progress

AWS ParallelCluster with Slurm

AWS ParallelCluster provides CloudWatch Logs integration, a cluster dashboard, and version-dependent cluster alarms. Its dashboard can show head-node metrics and logs, shared storage metrics, and Slurm cluster-health information. Current ParallelCluster guidance describes default head-node alarms for health, CPU, memory, and root-disk use in supported versions. These defaults are not a complete compute-node or job monitor, and ParallelCluster does not configure alarm actions for you. Check the dashboard, alarm, and monitoring configuration docs for your ParallelCluster version and settings.

For Slurm queue and job-level visibility, use scheduler data such as squeue for current state and sacct for completed-job accounting when Slurm accounting is configured. ParallelCluster supports Slurm accounting through slurmdbd and a MySQL-compatible database; you must configure and operate that accounting path. If operators need queue wait, completion, failure, or resource-use trends in CloudWatch, publish selected aggregates from scheduler/accounting data or your job hooks. Keep job identifiers in logs or an accounting store when they would create too many unique metric dimensions. The ParallelCluster Slurm accounting guide explains the supported setup.

AWS Batch jobs

AWS Batch publishes job lifecycle and duration metrics in the AWS/Batch namespace for jobs submitted to a job queue. They are useful for comparing submitted, runnable, running, succeeded, and failed work over time. AWS describes these metrics as best-effort monitoring signals: a fast job transition can be missed, so do not use them as billing, audit, or exact reconciliation counts. For compute-environment and job CPU, memory, disk, and network measurements, consider CloudWatch Container Insights; AWS bills its metrics as custom metrics. The AWS Batch metrics guide documents the signals and limitations.

Design dashboards around diagnosis

Give each dashboard a clear path from cluster health to the job or component that needs investigation. A practical HPC view can include:

  • Fleet health: node state, EC2 status checks, CPU, memory, filesystem use, and provisioning or Slurm health errors.
  • Job progress: queue age, running and completed jobs, failures, elapsed time, and resource use by queue or workload class.
  • Data path: FSx throughput, operations and capacity alongside EFA traffic, drops, retransmissions, and the job's input/output phases.
  • Application behavior: representative run time, useful work completed, and profiler results for the code path under study.

Label units, dimensions, statistics, and time periods so that a node-level metric cannot be mistaken for a job-wide result. Add dashboards for short operational windows and for longer run-to-run comparisons when the workload needs both. Amazon Managed Grafana can query CloudWatch for cross-cluster views; see the site's CloudWatch and Grafana setup guide for data-source and dashboard setup.

Set alarms that lead to an action

Alarm on symptoms that an operator can act on: repeated node-health failures, a sustained resource limit, a growing queue wait, a missing ParallelCluster cluster-management heartbeat where your version publishes one, an FSx capacity or performance condition, or an EFA error counter that departs from the workload's baseline. Set periods and thresholds from normal runs and service objectives; the correct values depend on the application, cluster size, and expected queue behavior. Choose how missing data should be treated, and avoid triggering on a single harmless spike when a sustained condition is what matters.

For each alarm, define its owner and response. Route notifications through the mechanism your team uses, and configure alarm actions explicitly; an alarm without a response path is only a graph state. CloudWatch alarm patterns are covered in the site's AWS monitoring guide and AWS's CloudWatch alarms documentation.

Control telemetry cost and access

Before increasing collection frequency, decide what question the extra resolution will answer. Detailed monitoring, custom metrics, high-resolution data, CloudWatch API reads, and log ingestion and storage can add charges. Use only the measurements and dimensions you need, especially when many nodes or per-job labels could multiply metric counts. Review charges by CloudWatch usage type and update retention settings to match the operational and audit needs for each log group. AWS maintains current cost guidance in Analyzing and reducing CloudWatch costs.

Use AWS Budgets for account, service, or tagged cost notifications, and assign an owner to investigate HPC spend. Budgets are useful for planning and follow-up, not immediate runtime control: AWS updates budget data up to three times per day, and usage-to-billing delays can postpone a notification. Keep that lag in mind when deciding what should stop or limit a job. For ongoing AWS API and configuration audit, use CloudTrail and AWS Config for their respective control-plane and resource-configuration signals; they do not replace host, scheduler, or application telemetry. Restrict access to monitoring data, then set log encryption and retention according to your data classification and audit requirements.

A practical monitoring checklist

  1. Choose one representative workload and record a baseline for queue wait, run time, completed work, and relevant compute, network, and storage use.
  2. Confirm which service metrics already exist, then configure the CloudWatch agent only for missing host, GPU, or EFA signals on the nodes that need them.
  3. Decide where scheduler accounting and job logs live, and preserve the data long enough to diagnose failures after short-lived compute nodes terminate.
  4. Build dashboards around cluster, scheduler, storage, and application questions; set alarm thresholds from observed workload behavior.
  5. Review CloudWatch and compute costs after enabling new metrics, logs, or higher collection rates.

A useful HPC monitoring system combines infrastructure health, scheduler outcomes, storage and network evidence, and application profiling. CloudWatch can bring many of those signals together, provided each metric's scope and collection method are clear. Revisit the setup when the scheduler, node types, application, or cluster scale changes.