5 Patterns for Resilient Serverless State Management on AWS
Published on · Updated on
Updated AWS service guarantees for DynamoDB consistency and transactions, Lambda retries, Step Functions workflows, queue buffering, event sourcing, and sidecars.
Serverless compute does not remove application state; it changes where durable state belongs. A standard AWS Lambda function should not depend on memory or temporary files surviving its invocation, even though Lambda may reuse an execution environment. Store business data in a durable service and make each request safe to retry. AWS also offers Durable Functions for checkpointed workflows that can run for up to a year. Their checkpoints preserve workflow progress; they do not replace a durable system of record or make an external business operation atomic.
The five patterns below address different boundaries: keeping current state, coordinating work across services, buffering accepted work, rebuilding read models from events, and adding a helper process to a container workload. They can be combined, but each needs a clear source of truth and a defined response to retries, concurrent writes, and partial failure.
- External storage: Keep canonical application state in a database or object store that outlives a function invocation.
- Saga orchestration: Coordinate local transactions across services and define the business action that follows each failure.
- Storage-first processing: Persist or enqueue work before acknowledging it, then process it asynchronously.
- CQRS with event sourcing: Keep an append-only event history and build query models from it when audit, replay, or different read patterns justify the added work.
- Sidecar: Run a closely related helper beside an application container; use it for local support work, not as a durable state store.
Quick Comparison
| Pattern | Use and limit |
|---|---|
| External storage | Keep shared, durable current state. A database commit does not include a later external side effect. |
| Saga | Coordinate independently committed services. Compensation is a new business action that can also fail. |
| Storage-first | Accept work for later processing. Delivery can repeat, and queue acceptance does not mean completion. |
| CQRS with event sourcing | Build read models from durable history. Projections can lag; a short-lived change stream is not an event store. |
| Sidecar | Package a helper with a container workload. It does not provide durable storage or fleet-wide coordination. |
Start by asking what must survive a failed invocation, what the caller is allowed to assume when it receives a response, and which component owns each business update. Those decisions matter more than labeling an architecture “serverless.”
1. External Storage Pattern for Stateful Functions
Use an external store when a value must be available to later invocations or other consumers. A shopping-cart API, for example, can store each cart in DynamoDB and load it when a request arrives. Standard Lambda environments can be reused, but AWS advises treating them as if they existed for one invocation; commit permanent changes to a durable store before the function exits. See AWS guidance on implementing statelessness in Lambda.
Failure handling and consistency
Persisting state protects it from a function restart only after the storage service has accepted the write. A client timeout can leave the outcome uncertain: the write might have committed even though the response was lost. Give a retried business request the same idempotency key, store that key with the result, and use a conditional write to prevent a second order; if the key already exists, return the saved result. If two writers update the same item, a version attribute and condition expression can detect a stale update; the losing writer must reload and resolve the conflict rather than overwrite it blindly. DynamoDB documents this approach as optimistic locking.
Choose consistency deliberately. DynamoDB reads are eventually consistent by default; a strongly consistent read is available for a table or local secondary index, but not for a global secondary index or stream. A global table adds a separate choice: multi-Region eventual consistency (MREC) replicates changes asynchronously, while multi-Region strong consistency (MRSC) synchronously replicates writes and supports strongly consistent reads across replicas. On MREC tables, a transaction is atomic in the Region where it ran, but replicas may temporarily expose only some of its writes. MRSC global tables do not support DynamoDB transaction APIs. These boundaries are described in AWS’s documentation for read consistency and global tables.
A practical implementation
For a cart, use a stable customer or cart key and update only the attributes that changed. For an order, create a pending record with the request key and an initial version, then advance its status with a condition such as “status is still PENDING and version is still 1.” A conditional write protects one item from a stale update. A DynamoDB transaction can atomically change multiple items in one Region, such as an order and its inventory reservation. It cannot include an SQS send, a payment-provider call, or another service’s database operation in that same transaction.
For MREC global tables, two Regions can update the same item concurrently and DynamoDB resolves the conflict on a last-writer-wins basis. A local version condition is not a cross-Region lock. If updates must be serialized, assign each record a write Region or choose a consistency design whose documented guarantees meet that need. MRSC offers stronger multi-Region reads and writes, with higher write latency and no transaction API.
Performance and limits
Every database read or write adds a network call, so cache only data that can safely be stale and make the cache disposable. DynamoDB is suited to structured items with known access patterns; S3 is suited to objects such as uploaded files. Neither choice makes an external action—such as sending an email or charging a card—part of the storage commit. Record the action’s status separately, pass the same idempotency key to a provider that supports one, and reconcile ambiguous outcomes.
Cost and operational burden
Compare the complete request path: reads, writes, indexes, retries, backups, data transfer, and any cache or orchestration layer. A durable database adds latency and request cost, but it avoids relying on a warm Lambda environment for user or business state. Use this pattern for current state; use the event-sourcing pattern below only when keeping and replaying a history has a concrete purpose.
2. Saga Pattern with AWS Step Functions
A saga coordinates a business operation whose services commit their own local transactions. AWS Step Functions can orchestrate those steps, but it does not turn them into one distributed database transaction or invent the compensating actions. Each service owns its update; the workflow records progress and decides what to try next. AWS describes saga orchestration and choreography in its Saga patterns guidance.
Failure handling and compensation
Consider an order workflow that reserves inventory, authorizes payment, and then confirms the order. If authorization fails after inventory was reserved, the workflow can release the reservation. If payment authorization succeeds but the order update fails, it can retry the update or void the authorization. If payment was already captured, the compensating action may be a refund. These are new business operations, not time reversal: they may fail, be delayed, or require manual reconciliation.
Make every retryable step safe to call again. Use a stable operation key such as orderId:reserve-inventory and have the owning service conditionally apply it once. A timeout does not tell the workflow whether a remote service completed its work before the response was lost. A compensation can also time out after it succeeds. Persist each step’s outcome, use idempotency keys supported by external providers, and expose unresolved orders for reconciliation.
Workflow and implementation
In Step Functions, model each local transaction as a Task state, use Retry for transient errors, and Catch to route permanent failures to the appropriate compensation path. Keep validation failures out of broad retry policies; retrying a declined card or invalid address will not make it valid. Test the timeout path as well as the explicit error path, because they can leave different evidence at the participant service. Step Functions supports configurable Retry and Catch behavior.
Workflow type affects duration and execution semantics. Standard workflows can run for up to one year and follow an exactly-once workflow model unless the definition configures retries. Express workflows can run for up to five minutes; asynchronous Express uses at-least-once execution, while synchronous Express uses at-most-once execution. These guarantees describe Step Functions workflow behavior; they do not make several service databases atomic, undo an earlier commit, or remove the need for idempotency when a task is retried. See AWS’s current guide to choosing a workflow type.
Latency and consistency
Orchestration adds transitions and network calls, and the saga is usually eventually consistent while it runs. Return an order identifier and a pending status when the workflow outlasts an API request; let clients check or receive an update rather than holding the connection open. Step Functions keeps workflow state, while each participant’s database remains authoritative for its own transaction.
Cost and limits
Standard workflows are billed by state transitions; Express workflows are billed by executions and duration. Retries and compensation add work and may add transitions or invocations, so include failure paths in cost estimates. Choose based on audit history, duration, integration needs, traffic, and idempotency—not on the assumption that one workflow type makes an entire business transaction “exactly once.”
3. Storage-First Pattern with Event Processing
Use storage-first processing when a caller needs a quick acceptance response but the work can finish later. For example, a webhook endpoint can validate an order event, send it to an SQS queue, and return 202 Accepted after SQS confirms the message was accepted. A Lambda event source mapping can then invoke a worker. The response means “the queue accepted this request,” not “the order completed.”
Durable handoff and retries
SQS retains a message until a consumer deletes it, it expires, or (when configured) a redrive policy moves it to a dead-letter queue. The default retention is four days and the maximum is 14 days. Standard queues provide at-least-once delivery, so duplicate messages and occasional reordering are possible. Lambda’s SQS event source mapping retries failed batches; by default, messages in a failed batch can be made visible again, including messages already processed successfully. Use an idempotent consumer, configure a redrive policy and dead-letter queue for repeated failures, and consider partial batch responses to retry only failed records. See AWS’s guides to SQS standard queues, message retention, and handling SQS event-source errors.
Give each incoming business event a stable identifier. Before applying an effect, the worker can conditionally record that identifier with the state change. If it calls another service, use that service’s idempotency key when available. A DLQ is a place to inspect messages that exceeded a configured receive threshold; it is not proof that the business work succeeded, and replaying it can repeat the effect. If your design uses EventBridge to route business events as well, see our EventBridge guide to retries, dead-letter queues, replay, and idempotent consumers.
The outbox failure window
A common failure window appears when an API first writes an order to a database and then sends an SQS message. The database write can succeed while the send fails, leaving an order that no worker knows about. Reversing the order creates the opposite risk: the message can be sent even though the database write fails. The transactional outbox pattern closes this dual-write gap by committing the business update and an outbox record in the same database transaction. A separate publisher sends outbox records to the queue or event bus.
The publisher can send a message successfully and then fail before marking the outbox record as sent. It will send that record again after recovery, so the message needs a stable ID and the consumer must deduplicate it. The outbox makes the database update and notification intent atomic; it does not make delivery or downstream processing exactly once. AWS explains these trade-offs in its transactional outbox pattern guidance.
Backpressure and processing delay
A queue absorbs bursts and lets consumers process at a sustainable rate, but it can turn overload into a growing backlog. Watch age of oldest message, queue depth, processing errors, and DLQ growth. A successful enqueue does not protect a message forever: retention is finite, and a message can reach a DLQ under its redrive policy. SQS is a buffer, not a substitute for an audit store or a read-after-write status model.
Cost and trade-offs
Budget for queue requests, worker invocations, retries, DLQ inspection, replay, and any S3 storage used for large payloads. If payloads are large, store the object in S3 and put a reference plus a stable event ID in SQS. If a consumer cannot tolerate duplicate work, fix that in the consumer or downstream service; queue type alone does not guarantee a business effect happened once.
4. CQRS with Event Sourcing
CQRS separates the model that accepts commands from the model optimized for queries. Event sourcing is a separate choice: it stores accepted state changes as an append-only event history and derives current or specialized views from that history. The two patterns are often combined, but CQRS does not require event sourcing. Use them when read and write needs differ, or when an auditable history and replay are worth the extra design and operating work. AWS outlines the distinctions in its guides to CQRS and event sourcing.
History, replay, and failure recovery
For an order system, the command side might append OrderPlaced, ItemReserved, and PaymentAuthorized events. A projection worker consumes those events to maintain views such as an order summary or sales report. If a projection is damaged, rebuild it from the durable event store. Replaying events should rebuild views; it should not resend emails, charge cards, or repeat other external effects unless that behavior is explicitly intended and guarded.
Do not confuse a database change stream with the event store. DynamoDB Streams captures item-level changes for up to 24 hours, and a Lambda consumer can process the same record more than once. It is useful for change data capture and projection updates, but its retention is not a permanent business history. Store domain events durably in an event table or another event store, and use streams only as a delivery mechanism when their retention and recovery behavior fit the design. AWS documents DynamoDB Streams retention and duplicate processing with Lambda consumers.
Ordering, concurrency, and event versions
Give each aggregate (such as one order) an ordered sequence or version. When appending an event, check that the current version matches the version the command read, then advance it. This optimistic concurrency check prevents two writers in the same Region from silently overwriting each other; a conflict must be rejected or retried from fresh state. A request ID recorded with the accepted command can prevent a client retry from appending the same business event twice.
If DynamoDB is the event store, a same-Region transaction can conditionally advance the aggregate version and append the event together. MREC global tables replicate items asynchronously, can resolve concurrent cross-Region writes on a last-writer-wins basis, and do not replicate a transaction as one atomic unit. MRSC provides cross-Region strong consistency but does not support transaction APIs. Do not assume a per-Region sequence condition creates a global event order; route an aggregate’s writes through a defined owner or choose a store and consistency model that supports the ordering the domain requires.
Version event schemas and keep consumers able to read older records, or provide explicit upcasting when events evolve. Track projection progress and make projection writes idempotent so a retry does not apply the same event twice. If a projection lags, reads can be stale; return the accepted command result or read from the command side when a caller requires immediate confirmation.
Read performance and consistency
Separate read models let you shape indexes and storage around query patterns, but they add propagation delay. A report can be eventually consistent while an order-status screen may need to show the result returned by the command. Choose per use case instead of assuming one consistency level fits every read.
Cost and trade-offs
Event history grows over time. Account for retention, backups, schema migration, snapshots, projection rebuild time, and the cost of keeping multiple read stores. Event sourcing is useful when the history itself has value; if the application only needs current state, a conventional database with an audit record may be simpler.
5. Sidecars in Container Workloads
A sidecar is a helper container deployed alongside an application container, commonly to provide a proxy, telemetry collector, or other closely related capability. In Amazon ECS, for example, Service Connect adds a proxy sidecar to a task. This is a container deployment pattern, not a serverless persistence pattern. See AWS’s description of the ECS Service Connect sidecar.
Failure boundaries
A sidecar can keep supporting code out of the application process, but it is deployed alongside the application in the same task or pod. In ECS on Fargate, containers in a task share the task’s allocated CPU and memory. A helper crash, bad configuration, or saturated proxy can still interrupt requests that depend on it; process separation alone does not make the application more fault tolerant. AWS describes the resource boundary for Fargate tasks.
When it fits
A telemetry collector or service proxy can be a useful sidecar in ECS or EKS. A small local cache can also sit beside a container, provided the database remains authoritative and the design can handle stale entries, restarts, and concurrent updates from other tasks. A sidecar is not a general-purpose state manager for standard Lambda functions, and it does not coordinate writes across a fleet. Keep durable business state in an external store, use a queue for buffered work, or use a checkpointed workflow when the requirement is workflow progress.
Performance and operations
Local communication can be quick, but every extra process or container adds configuration, health checks, deployment coordination, logs, and resource use. The application and helper may need compatible versions and startup behavior. If only one application needs the capability, a library or a managed service may be easier to operate.
Cost and trade-offs
Include the sidecar’s CPU and memory in task sizing, then compare that cost with the operational work it removes. Use it to package a helper with an application, not to claim isolation from infrastructure failure or to keep the only copy of important state.
Choosing Between the Patterns
- Current state across invocations: Start with external storage, an idempotency key, and conditional updates.
- Work that can finish later: Use a queue with stable event IDs, consumer deduplication, a configured DLQ, and backlog monitoring.
- Multi-service business steps: Use a saga with idempotent actions, explicit compensation, and reconciliation for ambiguous outcomes.
- Historical queries or rebuildable read models: Consider event sourcing with a durable log, aggregate versions, event schema evolution, and idempotent projections.
- A proxy or collector beside a container: Consider a sidecar, with health handling, resource sizing, and an external source of truth.
These choices can work together. An order service might commit a pending order and an outbox record, publish the event to SQS, and start a saga that reserves inventory and authorizes payment. The order database owns current status; the queue buffers work; the saga owns orchestration progress. If the business also needs replayable history or specialized reports, an event store and projections may be justified. A sidecar could support a containerized worker, but it would not replace any of those durable records.
At every handoff, define the success response, the durable record that proves acceptance, the key used to recognize a retry, the version or ordering rule for concurrent changes, and the recovery path if a later step fails. Do not infer an exactly-once business outcome across independent systems from a workflow or queue guarantee alone. Reliability comes from clear ownership, durable intent, safe retries, and a way to find and reconcile work that remains incomplete.
Conclusion
Use external storage for current state, a saga for multi-service business steps, and a queue when a caller can hand off work for later processing. Add CQRS and event sourcing when independent read models or durable history justify their complexity. Use sidecars for container helpers, not as a substitute for a durable data store. Choose each pattern by the failure it handles, then design explicitly for duplicate delivery, concurrent updates, uncertain outcomes, and compensation that may need its own recovery.
FAQs
How do I keep state between Lambda invocations?
For standard Lambda functions, persist application data in a database or object store and load what each invocation needs. Lambda Durable Functions can checkpoint progress for long-running workflows, but a checkpoint is workflow state; keep durable business records in the system that owns them.
What should I consider when using Step Functions for a saga?
Define every forward action and its compensating action, make retryable steps idempotent, and use a workflow type that fits the required duration and execution semantics. A compensation can fail or remain ambiguous, so persist step outcomes and provide reconciliation for unresolved transactions.
What is a sidecar, and does it make a serverless application more fault tolerant?
A sidecar is a helper container that runs alongside an application container, often as a proxy or telemetry collector. It can separate code responsibilities, but is deployed with the application. In ECS on Fargate, containers in a task share allocated CPU and memory. It does not provide durable storage or automatically isolate the application from failures in a helper it needs.