AWS Fault Injection Service (FIS): Test Workload Resilience
Published on · Updated on
Reworked FIS actions, targeting, guardrails, recovery, and cost guidance.
AWS Fault Injection Service (AWS FIS) runs controlled experiments against real AWS resources. An experiment tests a specific reliability hypothesis, such as whether a service stays within its latency objective when one test instance stops. The result applies only to the actions, resources, and workload conditions you tested.
For a first experiment, choose one reversible fault, target a single test resource, define a measurable steady state, configure a CloudWatch alarm stop condition, and decide how the resource will recover before you start.
Before you start
Set up the account and test environment
Use an AWS account and Region where you can isolate the experiment from production traffic and data. Prepare the target resource, a test workload, and the monitoring needed to observe user-visible behavior. FIS actions change resources or calls in the target workload, so treat an experiment as a real operational change.
Use separate IAM permissions for running FIS and changing resources
The identity that creates or starts an experiment needs the relevant AWS FIS permissions. When a caller creates or updates a template with an experiment-role ARN, it also needs iam:PassRole permission scoped to that role; AWS's FIS template policy example shows this separate grant. The experiment role is the role AWS FIS assumes to perform selected actions. Give it only the permissions required for the action and target, and restrict its trust policy to fis.amazonaws.com. AWS recommends adding aws:SourceAccount and aws:SourceArn conditions to reduce confused-deputy risk.
AWS FIS also uses the service-linked role AWSServiceRoleForFIS, backed by AmazonFISServiceRolePolicy, for monitoring and resource selection. AWS FIS creates it when an experiment starts; it does not replace the experiment role that authorizes fault actions. The identity used for the first run may need permission to create that service-linked role. See AWS's guides to experiment roles and the FIS service-linked role.
Check action prerequisites
Before choosing an action, check its supported resource type, required permissions, parameters, and any agent or service prerequisites in the current AWS FIS actions reference. An EC2 stop action needs EC2 stop permission and start permission if you configure FIS to restart the instance. Encrypted EBS volumes can also require access to the KMS key used for encryption.
How FIS experiments work

Templates define experiments; runs execute them
An experiment template defines actions, targets, stop conditions, and the experiment role. Starting a template creates an experiment using a snapshot of that template, so later edits do not change a run already in progress. Actions can run in sequence or in parallel. A stopped or failed experiment cannot be resumed; start a new run from the template after reviewing what happened. See AWS's guide to starting an experiment from a template.
Actions are specific to supported services and resources
FIS does not provide one generic switch for every kind of outage. Its action catalog covers specific operations for supported AWS resource types. The API fault-injection actions aws:fis:inject-api-internal-error, aws:fis:inject-api-throttle-error, and aws:fis:inject-api-unavailable-error target an IAM role and can inject those errors into configured EC2 or Kinesis API operations; they do not affect arbitrary application endpoints. Check the action reference before planning around any fault.
Test how the workload responds
Use an experiment to check a prediction: for example, requests continue to meet a latency target after one instance is stopped, alarms detect a dependency failure, or queued work drains after a consumer restarts. FIS supplies controlled actions and run records; your application metrics and recovery checks determine whether the workload behaved acceptably. For related patterns around retries, queues, and fault boundaries, see AWS cross-service resilience patterns.
Plan a bounded, observable experiment
Define scope and success criteria
Write down the question, expected behavior, affected resource, maximum fault duration, and the person or process watching the run. Define the workload's steady state using service-level measures such as request success rate, latency, or retry volume. An infrastructure metric such as CPU can help explain a failure, but may not show whether users can complete the operation.
Start in a contained test environment. If a later production experiment is necessary, review its scope, customer impact, timing, recovery procedure, and operational approval before running it. For broader recovery evidence, pair FIS experiments with a restore or disaster-recovery drill; our AWS disaster recovery testing guide explains how those tests answer different questions.
Define alarms for unacceptable impact
Create a CloudWatch alarm stop condition for a workload signal that means the experiment should end, such as sustained error rate or latency above the service objective. Confirm the alarm evaluates the intended metric before the run, choose its period and evaluation window with the signal's reporting delay in mind, and decide how missing data should be treated.
A stop condition stops the experiment when its alarm triggers; it is not an instantaneous circuit breaker and does not guarantee that every fault is reversed. CloudWatch must first receive and evaluate the metric, so a slow or missing signal can delay or prevent the alarm from reaching ALARM; see the CloudWatch guidance on alarm evaluation and missing data. Configure a supported action's post-action or recovery behavior where available. An irreversible action such as terminating an instance cannot be undone by stopping the experiment. AWS documents which actions support post-actions in its FIS action guidance. FIS also has a regional safety lever that can stop running experiments and prevent new ones in that account and Region; it is a broader emergency control, not a per-experiment rollback.
Build a template and start a run
Choose the console or AWS CLI
Create a template in the AWS FIS console or with aws fis create-experiment-template. The console flow asks you to define actions and targets, select the experiment role, and optionally add stop conditions and logging. The CLI accepts a JSON template.
Select a valid action and its parameters
This example stops one running EC2 instance tagged for the test and asks FIS to start it again after two minutes. It uses the supported action ID aws:ec2:stop-instances; it does not terminate the instance. Before creating the template, tag exactly one disposable test instance with fis-test=true, create a CloudWatch alarm in us-east-1, and replace the example account, role, and alarm values with your own. If you use another Region, change the alarm ARN and each CLI --region value to match the target resources and alarm.
{
"description": "Stop and restart one tagged test instance",
"roleArn": "arn:aws:iam::123456789012:role/FIS-EC2-TestRole",
"stopConditions": [
{
"source": "aws:cloudwatch:alarm",
"value": "arn:aws:cloudwatch:us-east-1:123456789012:alarm:FIS-Test-ErrorRate"
}
],
"targets": {
"testInstance": {
"resourceType": "aws:ec2:instance",
"resourceTags": {
"fis-test": "true"
},
"filters": [
{
"path": "State.Name",
"values": ["running"]
}
],
"selectionMode": "COUNT(1)"
}
},
"actions": {
"stopOneInstance": {
"actionId": "aws:ec2:stop-instances",
"parameters": {
"startInstancesAfterDuration": "PT2M"
},
"targets": {
"Instances": "testInstance"
}
}
}
}
Save the JSON as fis-template.json and create the template in the same Region as the instance and alarm:
aws fis create-experiment-template --region us-east-1 --cli-input-json file://fis-template.json
The target uses COUNT(1), which selects one matching instance at random. Keep the tag limited to the disposable test instance; if more instances match, FIS may stop a different one. The experiment role must allow the EC2 operations used by the action and any required KMS access for encrypted volumes. Consult the action reference for exact permissions and parameters.
Preview and narrow the targets
FIS supports explicit resource IDs and, for supported resource types, tags, filters, or parameters. Selection mode can be ALL, COUNT(n), or PERCENT(n). FIS resolves targets when an experiment starts; an empty result normally causes the experiment to fail. Use the console's target preview to check the intended resources. Preview skips fault actions, does not verify action permissions, and can differ from the eventual run if resources change or targets are sampled randomly.
Add the stop condition and experiment role
Put the CloudWatch alarm ARN in stopConditions and set roleArn to the experiment role trusted by fis.amazonaws.com. Verify that role's permissions cover the selected action. A stopped run cannot be resumed. For multi-account experiments, additional account roles and CloudWatch alarm sharing are required; follow the current AWS instructions rather than reusing this single-account example unchanged.
Start the reviewed template
Inspect the template, alarm, experiment role, and resolved target. In the console, select the template and choose Start experiment. From the CLI, run this in the template's Region and replace the example template ID:
aws fis start-experiment --region us-east-1 --experiment-template-id EXTxxxxxxxxx
Starting a template snapshots its current configuration for that run.
Monitor actions and workload health
Watch the experiment state and action progress in the FIS console or CLI. At the same time, monitor user-facing success and latency, dependencies, and the alarm state. Keep the operator who can stop the run available until the fault ends and recovery is confirmed.
Confirm recovery before ending the exercise
Check that the target returned to a safe state and that the workload recovered. If an action does not have a configured post-action, perform and verify the required recovery manually before another run.
Examples of supported FIS actions
Stop and restart a test EC2 instance
Use aws:ec2:stop-instances with startInstancesAfterDuration to test instance replacement, health checks, or application failover. The restart delay is configured for that action; an EC2 termination action is a separate, destructive operation.
Disrupt a defined network path
Use aws:network:disrupt-connectivity on supported subnet targets to deny a selected traffic scope temporarily. For example, the all scope blocks traffic entering and leaving the subnet but allows traffic within the subnet. FIS temporarily associates a cloned network ACL and restores the original association when the action completes. This tests connectivity loss, not arbitrary network latency.
Apply CPU stress through Systems Manager
FIS can use aws:ssm:send-command with the AWS FIS CPU stress document on an EC2 instance managed by Systems Manager. Confirm the SSM Agent, instance profile, document prerequisites, permissions, and the document's own duration settings before running it; see AWS's guide to using Systems Manager documents with FIS.
Treat storage faults as service-specific tests
For EBS, aws:ebs:pause-volume-io pauses volume I/O, while aws:ebs:volume-io-latency injects read or write latency. Both require target volumes to be in the same Availability Zone and do not support volumes attached to Outpost instances; pause-I/O targets must also use Nitro-based instances. Neither action injects arbitrary disk errors or data corruption. Check the current action reference for all prerequisites and recovery behavior, and use a separate backup and restore test when you need evidence that data can be recovered.
Schedule repeatable experiments and control cost
Repeat after meaningful changes
Rerun an experiment when a change to the workload, scaling policy, dependencies, monitoring, or recovery process could change the result. FIS supports one-time or recurring schedules through EventBridge Scheduler. Before enabling a schedule, verify the target set, alarm behavior, recovery actions, notification path, and charges. In CI/CD, use an explicit controlled stage rather than running disruptive actions on every code change by default.
Budget for FIS action time and for the AWS resources and monitoring used by the test. The current FIS pricing page lists $0.10 per action-minute in most Regions, plus $0.10 per action-minute for each additional target account. In AWS GovCloud (US-East and US-West), the base and each-additional-account rates are $0.12. Action time is measured from start to stop and rounded to the nearest minute. Parallel actions accrue action time separately, and the number of affected resources does not change the FIS action charge. Optional experiment reports currently cost $5 each. Check the current AWS FIS pricing and the pricing of other services in the experiment before estimating a run.
Review results, improve, and troubleshoot
Compare results with the hypothesis
Record the experiment and template IDs, action and target set, alarm state, workload metrics, and the time until the service returned to its expected state. Distinguish the FIS action result from the application result: an action can complete successfully while the workload violates its service objective. For a more durable record, enable experiment logging or configure an experiment report with a CloudWatch dashboard; reports and associated storage or API calls can add charges.
Find gaps in detection and recovery
Look for missed alarms, error spikes, slow failover, capacity limits, requests that did not recover, or effects that persisted after the action ended. Check the FIS experiment history alongside CloudWatch metrics and logs. Assign an owner to each finding, make the smallest change that addresses it, and rerun the same bounded experiment to see whether the result improved.
Check the experiment and action state
Use the FIS experiment details to find the failed action, target, or service response. Confirm that selected resources still exist, match the target filters, and are in a state the action supports. An empty resolved target set can fail the experiment; update the template and preview again before retrying.
Verify the signal and recovery path
If the workload behaves differently than expected, compare service metrics, alarm history, experiment logs, and the action's post-action status. Check whether the CloudWatch alarm evaluated the intended metric and whether missing data or its evaluation window delayed a state change. Confirm the resource is healthy and any manual recovery steps are complete before another run.
Separate caller and role permission failures
Check that the identity starting the experiment can call the needed FIS operations and create the service-linked role if required. Then check that the experiment role trusts AWS FIS and allows the selected action. For encrypted volumes, verify KMS key permissions too. The action reference lists the service permissions for each supported action.
Use the findings to improve the workload or its runbook, then repeat the test when meaningful changes could alter its behavior.