Rewrote the guide to compare Lambda, Fargate, and rehosting; added a staged migration plan and clarified Lambda durable execution, invocation limits, and cost considerations.
A legacy application can move to AWS without becoming a collection of microservices. Modernization is a series of choices about where each workload should run, which behavior should change, and how to prove the change is safe. AWS Lambda can suit bounded, event-driven work and supported checkpointed workflows; Amazon ECS on AWS Fargate can run containerized services and longer processes; rehosting or retaining an application can be the right first step when compatibility or migration risk matters more than redesign.
Use this guide to choose a first workload, test it against real constraints, and cut it over with a recovery path. Serverless can reduce infrastructure management, but availability and savings depend on the design, workload, dependencies, and operating model.
Choose a runtime that fits the workload
“Serverless” means AWS manages more of the underlying compute infrastructure. Your team still configures capacity behavior, networking, permissions, data, deployments, monitoring, and recovery. Compare the options for each workload rather than selecting one target for the whole application.
Path
Good fit when
Trade-offs to check
Retain or rehost on conventional compute
Retain the application where it is while a constraint or business case is unresolved, or rehost it on AWS compute such as EC2 when moving is the priority and it depends on a specific operating system, runtime, local process model, or vendor configuration.
Rehosting can reduce data-center work, but it does not remove application maintenance or prove lower total cost. You still operate the target compute and plan its capacity.
ECS on Fargate
The application fits a container and needs a long-running process, existing web server, background worker, or greater control over CPU and memory.
Fargate removes server management. ECS service scaling, task count, health checks, load balancing, and deployment behavior still need configuration and validation.
Lambda
A bounded piece of work can start from an HTTP request, schedule, queue message, file, or other event and finish within the function’s limits; a multi-step workflow may also fit if it can checkpoint and pause between steps.
Handlers and integrations have different request, timeout, state, and retry behavior. Default Lambda compute scales concurrency with demand; Managed Instances use a different model. Any scaling must account for quotas and downstream capacity.
Lambda limits depend on the compute and invocation model:
Default compute: Standard functions run for at most 15 minutes per invocation, with 128 MB to 10,240 MB of memory. Check the current Lambda quotas.
Lambda Managed Instances: Asynchronous invocations and supported event-source mappings can run for up to 90 minutes. They use managed EC2 capacity with a different scaling and pricing model; see the Managed Instances overview and current quotas.
Durable Functions: An asynchronous, checkpointed workflow can run for up to one year. A synchronous caller still waits no more than 15 minutes. With an event-source mapping, total execution is limited by the function timeout: up to 15 minutes on default capacity or 90 minutes on Managed Instances for supported sources. See AWS's Durable Functions invocation guidance and event-source mapping limits.
Confirm runtime, Region, and invocation-source support in the linked AWS documentation before choosing a long-running path. AWS's Lambda or Fargate decision guide describes the different execution models.
With ECS, automatic scaling changes a service's desired task count only after you configure a scaling policy; see ECS service auto scaling.
For a broader portfolio decision, AWS describes seven migration strategies: retire, retain, rehost, relocate, repurchase, replatform, and refactor. Refactoring is one option, not the default outcome. See AWS Prescriptive Guidance on migration strategies.
Assess the application before choosing a slice
Map the workflows the application actually runs, not just its components on an architecture diagram. Include user-facing requests, batch and scheduled work, databases, file shares, queues, identity systems, third-party APIs, and manual operating procedures. For each workflow, record its business owner, peak and normal load, latency expectations, failure impact, data it reads or writes, and current recovery steps.
Use source inspection, deployment records, logs, traces, and owner interviews together. A code reference does not prove a dependency is active, and a quiet test environment does not prove an infrequent job is unused. Record what you observed, when you observed it, and what remains uncertain. The site's application portfolio analysis guide covers portfolio-level discovery; five strategies for analyzing legacy code gives a deeper code and runtime assessment method.
Before changing a workflow, establish a baseline from representative traffic: response or processing time, throughput, error rate, resource use, database load, and current infrastructure cost. Include peak periods and important failure paths where practical. These measurements give the pilot a comparison point and reveal dependencies that may limit Lambda or container scaling.
Plan one bounded migration
Choose a first workload with a clear owner, observable input and output, manageable dependencies, and a way to limit its impact. A useful pilot tests a real constraint, such as request latency, a database connection ceiling, deployment recovery, or a variable batch workload. Do not choose a slice solely because it is small if it cannot be measured or rolled back safely.
Write down the goal and acceptance measures before implementation. For example: preserve the existing user-visible behavior, meet the agreed latency and error targets at peak load, keep database usage below a safe limit, and stay within a cost range calculated for the measured workload. Set the comparison period and data sources so the pilot can be judged consistently.
Decide where the boundary between old and new behavior will sit. A route, API, or message can direct one capability to its new implementation while the rest stays in the monolith. This incremental pattern is often called the strangler fig pattern; AWS describes its purpose and trade-offs in its Strangler fig pattern guidance. Introduce a routing layer only if the migration needs one; a small application may already have a suitable seam.
Keep the first change narrow enough to understand. One deployment unit does not require one microservice per feature, and a serverless application can call an existing monolith or managed database. Split components when independent scaling, release ownership, failure isolation, or a clear domain boundary justifies the added distributed-system work.
Adapt only the behavior the target requires
For a Lambda candidate, move the selected behavior behind a handler that receives an explicit request or event, validates it, performs one bounded unit of work, and returns or records an outcome. Keep durable state in its existing database or another durable service; an execution environment and its local memory or temporary files are not durable application state, as described in AWS's Lambda application design guidance. Read the Lambda basics guide for the handler, invocation, and permission model.
Before choosing Lambda, check whether the code assumes a continuously running process, local session state, long transactions, large in-memory datasets, or native libraries. A long-running workflow fits Durable Functions only when it can be divided into safe checkpointed steps and waits; continuous processes or long uninterruptible work may call for a container on Fargate or a staged move that leaves the workload in place while surrounding dependencies are changed.
On Lambda's default compute type, execution environments scale with concurrent demand, subject to regional quotas and scaling behavior. A burst can therefore affect a database or external API before that dependency can absorb the extra connections or requests. Measure peak concurrency against downstream limits. Lambda concurrency controls explain the default account pool, reserved concurrency, and provisioned concurrency: reserved concurrency reserves a function's share and caps its maximum, while provisioned concurrency pre-initializes environments for latency-sensitive workloads. Managed Instances use a different scaling and concurrency model.
Preserve clear ownership of data and events
Changing compute does not require changing the database at the same time. Keeping the current relational database can avoid combining a compute migration with a data-model migration. First verify network access, query behavior, connection limits, transaction needs, and recovery requirements. If Lambda's scale can create more database connections than the database can serve, connection reuse or Amazon RDS Proxy may help pool connections for supported RDS and Aurora engines; it does not replace capacity testing. Check current RDS Proxy engine and Region support and see AWS's Lambda and RDS connection guidance.
Give each data domain a clear writer during the transition. Avoid letting the old and new implementations independently write the same records unless you have designed and tested conflict handling. If data must be copied or transformed, validate counts and business-level results, monitor changes during the copy, and rehearse how writes will be handled at cutover. A schema change that remains compatible with both versions is easier to roll back than one that requires an immediate irreversible conversion. AWS's cutover guidance also treats data recovery as a separate part of the rollback plan after new writes reach the target.
Choose synchronous calls when the caller needs an immediate result. Use a queue or event only when buffering, independent processing, or looser timing is useful enough to justify extra delivery and failure handling. Event semantics vary by trigger and invocation type: a failed asynchronous Lambda invocation is retried by default, and queue-backed event source mappings can deliver records more than once. Make side effects idempotent, for example by recording a stable request or event identifier with the result, and define how failed work is inspected and replayed. See AWS on Lambda retry behavior, idempotent function code, and SQS visibility timeouts and dead-letter queues.
Build and release the pilot as a production change
Define the function or container, its event source or route, permissions, network settings, alarms, and environment-specific configuration in infrastructure as code. Keep secrets in a managed secrets store rather than source code or plain-text templates. Grant each runtime identity only the actions and resources it needs; AWS explains this boundary in its Lambda execution role guidance. Review both who can deploy and what the deployed code can access.
Deploy through the same repeatable pipeline used for later releases. Test the handler or container locally where useful, then test the deployed integration with representative events and realistic dependencies. Verify authorization, timeouts, payload boundaries, retries, duplicate events, database behavior, logs, and rollback before exposing normal traffic. A unit test alone cannot establish that cloud permissions, routing, and trigger configuration work together.
Carry security controls through the migration
Map existing controls to the new path before traffic moves: identity, least-privilege access, secret rotation, encryption in transit and at rest, network boundaries, audit records, and applicable data-retention or regulatory requirements. Use separate permissions for deployment and runtime access. Review every event source and API for who can invoke it, and avoid logging credentials, payment details, or other sensitive payload fields.
Instrument the workflow end to end
Collect structured logs with a request or event identifier that can connect activity across the entry point, function or task, database, and external calls. Track user-visible latency, throughput, errors, throttles, concurrency, queue age or depth, and database saturation. Use CloudWatch metrics for Lambda and alarms for service health; add distributed tracing when a request crosses enough components that logs alone cannot explain latency or failure.
For asynchronous work, alert on growing backlog, old messages, repeated failures, and dead-letter queue activity, not just invocation errors. Give the on-call owner a way to find the failed input, decide whether it is safe to retry, and replay or reconcile it without repeating a business side effect.
Validate performance under representative load
Run a staged load test with the same event shapes, data size, concurrency pattern, and important downstream calls expected in production. Compare response or processing-time percentiles, throughput, errors, throttles, database behavior, and resource use against the baseline. Test a sudden burst and a dependency failure if they are realistic risks; increasing Lambda concurrency is not an improvement if it only overloads the database.
For Lambda's default compute type, profile initialization and handler work before configuring provisioned concurrency. Tune memory and timeout with measured runs; memory also affects allocated CPU. For ECS on Fargate, test task startup, health checks, desired task count, and any configured auto-scaling policy. Scale policies need a metric, thresholds, and enough task capacity to meet the recovery time your service requires. The AWS migration benchmarking guide can help structure a before-and-after comparison.
Compare total cost using measured demand
There is no automatic serverless saving. Estimate cost for the same workload and service level in the target region. Include the selected Lambda compute model: requests and execution duration, provisioned concurrency, Durable Functions operations and checkpoint data, or Managed Instances capacity charges, as applicable. Also include API and event services, Fargate vCPU, memory and task uptime, database capacity, storage, data transfer, logs, monitoring, and the remaining engineering and support work. For on-demand Durable Functions, suspended waits do not accrue compute-duration charges, but operations and checkpoint storage still affect cost.
With Lambda's default compute type, reserved concurrency does not lower request or duration prices. It reserves concurrency for a function and sets its ceiling, which can protect a critical function or limit pressure on a dependency; unused reserved capacity can reduce what remains available to other functions. On this compute type, provisioned concurrency can reduce initialization latency, but AWS bills for configured provisioned capacity as well as requests and execution. Fargate resource billing starts when AWS begins downloading the container image and ends when the task terminates, subject to the current minimum billing duration. Include task startup and any continuously running baseline capacity in the estimate. Check current regional rates in the Lambda pricing page and Fargate pricing page, then use the AWS Pricing Calculator and compare actual pilot usage with the estimate.
Variable, intermittent work may benefit from per-invocation billing; steady, continuously running workloads can favor a different compute or pricing model. Include transition costs, duplicate environments during cutover, and database costs in the comparison. Decide from measured workload and current account pricing rather than applying a fixed savings percentage.
Cut over gradually and rehearse recovery
Establish a production-like baseline. Confirm the current behavior, data state, dependencies, and recovery steps. Agree on the metrics and limits that will stop the rollout.
Deploy without sending normal traffic. Verify permissions, health checks, alarms, event handling, and database access. Use test or replay data that cannot create unintended customer-facing effects.
Expose a limited workload. Route a small, controlled share of compatible traffic or events to the new path. Compare behavior and data outcomes with the baseline; watch downstream capacity as concurrency changes.
Increase exposure only when evidence is healthy. Check error rate, latency, backlog, resource use, cost, and data reconciliation at each step. Keep a named owner and a clear decision point.
Rollback by switching routing or disabling the new trigger. Confirm in advance that the old path can still serve requests and understand how writes made by the new path will be read or reconciled. A route switch does not undo incompatible data writes, so define a forward-fix or recovery process for that case.
Retire the old path only after a stable period. Confirm data ownership, restore procedures, operational handoff, and any retention obligations before removing the former implementation.
Operate the new path before expanding the migration
After cutover, review service health, cost, and operational effort against the pilot's acceptance measures over representative peak and normal periods. Update the runbook with the observed failure modes and the tested recovery procedure. Expand only when the current slice has an owner, reliable deployment, clear data ownership, and results that justify the next change. If the evidence favors Fargate, rehosting, or leaving another component in place, keep that option in the migration plan.
A measured path to modernization
Assess the workload, choose a runtime that fits its behavior, and migrate one bounded capability at a time. Protect data ownership, design for the trigger's retry semantics, cap and observe load where dependencies require it, and validate the rollback path before broad cutover. Measure performance and total cost using the actual workload. This produces a modernization decision grounded in how the application runs, rather than an assumption that every legacy system should become serverless.
FAQs
Should a legacy .NET application be rewritten for Lambda?
Not automatically. First check its runtime, framework, native dependencies, process model, and request or job duration against the target's supported environment. A compatible bounded workflow may fit Lambda; a long-running containerized application may fit ECS on Fargate; an application with tightly coupled behavior may be safer to rehost or retain while you modernize a specific seam. Choose from workload evidence and migration goals, then test a pilot.