Serverless Backend for Frontend (BFF) with AWS Lambda
Published on · Updated on
Rewrote the guide to clarify BFF responsibilities, request timeouts, concurrency, downstream limits, and API authorization.
A Backend for Frontend (BFF) is an API layer shaped around the needs of a particular client, such as a browser, mobile app, or partner integration. It can combine calls to existing services, translate their responses into a client-friendly contract, and keep internal service details away from the client. AWS Lambda can run this request-handling code without you managing application servers; Amazon API Gateway can provide the HTTP entry point.
Lambda does not make an API reliable or fast by itself. A synchronous request still depends on every service it calls, each timeout along the request path, available concurrency, and the way failures are handled. This guide shows how to design those boundaries for a practical serverless BFF.
What a BFF does
A BFF serves a client-specific API contract. For example, a mobile screen might need a compact product summary and a cart count, while a browser page might need richer product details and recommendations. The BFF can gather those fields and return one response, reducing client-side orchestration and avoiding data the client does not need.
The BFF should compose and present domain capabilities, not become a second home for their rules. Keep durable business decisions—such as whether an order can be placed—in the service that owns that domain. Otherwise, separate BFFs can develop conflicting versions of the same rule.
Use a BFF when different clients have meaningfully different data or interaction needs, or when a frontend should not coordinate several internal APIs. A single stable API may be simpler when clients share the same contract. Do not create a separate BFF for every screen or release; each API adds code, deployment, access policy, and operational work.
Architecture and request flow
A common synchronous design has a client call API Gateway, which authenticates or authorizes the request and invokes a Lambda function. The function validates the input, calls the required domain APIs or data services, combines only the permitted results, and returns an HTTP response through API Gateway.
| Component | Responsibility | Design question |
|---|---|---|
| Client | Renders the experience and sends a request with its user credentials. | What data does this client need for this operation? |
| API Gateway | Provides routes, request controls, and an HTTP integration with Lambda. | Which API type, authorization method, and request limits fit the API? |
| Lambda BFF | Validates input, applies presentation-level composition, and returns a client contract. | Can the function finish within the end-to-end request deadline? |
| Domain services | Own business rules and their data. | What are their rate, connection, and availability limits? |
Keep routes and contracts explicit even if several routes share one Lambda function. A single catch-all function can be useful for a small API, but it can also concentrate unrelated release schedules, permissions, and traffic. Split functions when the difference improves ownership, access control, scaling, or deployment—not simply because the clients have different names.
When a request is expected to take longer than a user-facing API can wait, separate acceptance from completion. Validate the request, enqueue or start durable work, return an identifier, and expose a way to check status or receive completion. A queue or workflow service is a better fit for long-running tasks than keeping an HTTP request open.
Design the request path and timeout budget
There is no single timeout for a serverless BFF. The client, any CDN or proxy in front of the API, API Gateway's integration, the Lambda function, and each downstream call can have separate deadlines. The request fails from the caller's point of view when the earliest applicable deadline expires. Increasing Lambda's timeout does not extend a shorter API Gateway or client timeout.
Lambda's configurable function timeout can be set up to 900 seconds (15 minutes), but that does not make a normal HTTP integration a 15-minute request. API Gateway REST APIs have a default integration timeout of 29 seconds; Regional and private REST API limits can be increased, sometimes with a lower account-level throttle quota. HTTP APIs have a 30-second maximum integration timeout. Check the current limit for your API type and endpoint in AWS's REST API quotas and HTTP API quotas; a browser, mobile network, or other intermediary can impose a shorter limit. Lambda's execution limit is described in the Lambda timeout documentation.
Set an end-to-end latency objective first. Reserve time for authorization, Lambda startup and processing, response serialization, and network transfer. Then give each downstream call a shorter deadline than the time remaining for the whole request. If one of several parallel calls fails or exhausts its budget, decide whether the endpoint can return a useful partial response, a clear error, or a retryable status. Do not let every dependency consume the full Lambda timeout.
A timeout at the HTTP boundary does not guarantee that Lambda stopped running. AWS notes that a Lambda function can continue after the connection to API Gateway closes because of a timeout. The user may retry while the original function is still performing a write. Make state-changing operations idempotent: use a client-supplied idempotency key or a domain-specific deduplication rule, and make retry behavior explicit. See AWS's notes on API Gateway response timeouts and synchronous Lambda invocations.
Measure initialization before tuning
For Java or other workloads where initialization measurably affects latency, first measure cold and warm invocation behavior. Consider smaller dependencies, suitable memory, and provisioned concurrency only when measurements justify the cost and operational settings. The practical options and their trade-offs are covered in our guide to mitigating AWS Lambda cold starts.
Control concurrency and downstream load
Lambda creates execution environments as concurrent demand grows, subject to account and function limits. Concurrency is the number of requests in flight at a time; it is not an unlimited capacity guarantee. Reaching a quota or a configured concurrency cap can cause throttling. AWS explains the distinction between account concurrency, function scaling, and reserved concurrency in its Lambda concurrency guide.
Autoscaling the BFF can increase pressure on a database or service that scales differently. A single inbound request may fan out to several downstream calls, so a traffic burst can multiply into a larger burst of dependency work. Estimate expected concurrency from request rate and execution duration, then check whether each dependency can support the resulting call rate and open connections.
- Set bounded connection pools and parallel fan-out. Avoid unbounded work for every incoming request.
- Use reserved concurrency when a function needs a cap that protects a downstream system or preserves capacity for other functions. A cap also means excess requests can be throttled, so define the client and retry behavior.
- Apply throttling at the API boundary where appropriate, and use backoff with jitter for retryable dependency failures. Avoid immediate retry loops that amplify an outage.
- Cache only data with an acceptable freshness and authorization model. Never share user-specific results across identities through an unkeyed cache.
- For relational database connections from Lambda, assess whether a connection proxy or a different access pattern is needed. More function concurrency does not create unlimited database connections.
Protect the user-facing path from slow or unavailable dependencies. Set a deadline per call, handle errors deliberately, and return a meaningful error without exposing internal exception details. A circuit breaker or fallback can help when its behavior is safe for the domain; returning stale or incomplete data is not appropriate for every operation.
Separate authentication from authorization
Secure the API and its data
Authentication establishes who made a request. Authorization decides whether that identity may perform the requested action on the specific resource. A valid login is not permission to read another tenant's order, and a client label such as “mobile” is not proof of a user's identity.
Choose an API Gateway authorizer or application-level verification that matches the API type and identity provider. For HTTP APIs, a JWT authorizer can validate token claims and scopes; REST APIs support Cognito user-pool authorizers and other options. In the BFF, use trusted identity and scope information to check the user's permission for each operation and resource. AWS documents JWT authorizers for HTTP APIs and Cognito authorization for REST APIs.
API Gateway API keys are for identifying API consumers and applying usage-plan throttling or quotas. AWS explicitly says not to use them for authentication or authorization. A key embedded in a browser or mobile app can be copied, so it cannot establish the identity or permissions of the person using that app. Use an appropriate user or service identity mechanism; reserve API keys for usage management where they fit. See API Gateway usage plans and API key guidance.
Lambda's execution role grants the function permission to call AWS services. It is separate from the end user's authorization. Grant each function only the AWS actions and resources it needs, and avoid broad policies that turn a BFF into a privileged data access path. Store credentials in a managed secret store rather than source code, and keep tokens, passwords, and sensitive payloads out of logs. AWS describes Lambda execution roles and least privilege.
Attach a function to a VPC when it needs network access to private resources, with routing and security groups configured for the required paths. VPC attachment is not a substitute for authentication, authorization, or least-privilege IAM; it also does not automatically provide internet access. AWS explains the networking requirements in its guide to internet access for VPC-connected Lambda functions.
If the API uses browser cookies, set cookie attributes and cross-origin rules for the application, and consider the relevant CSRF protections for the authentication design. Validate request shape, size, and values before using input in downstream calls. Return only fields the caller is allowed to see; response shaping is part of the BFF's security boundary.
If the BFF manages session state, keep it in an appropriate shared store or use signed, short-lived tokens with a deliberate revocation strategy. Do not depend on a particular Lambda execution environment being reused. See our guide to session management in AWS Lambda for the trade-offs involved.
A practical implementation workflow
- Write the client contract. Identify the page or operation, the fields it needs, freshness expectations, and how errors appear to the user. Make the API independent of the shape of any one downstream response.
- Map each field to its owner. Identify the domain service responsible for the data or decision. Keep ownership and business rules there; use the BFF for aggregation, protocol adaptation, and presentation-specific composition.
- Choose the request and authorization path. Define routes, authentication, resource-level authorization, validation, payload limits, and throttling. Decide which operations need synchronous responses and which should return a job identifier.
- Implement bounded dependency calls. Use explicit connection and request deadlines, bounded parallelism, safe retries, and idempotency for writes. Return only the data in the contract.
- Deploy repeatably. Define API Gateway routes, Lambda settings, IAM policies, secrets, and environment-specific configuration in infrastructure as code. Use versions or aliases and a staged rollout where the release process needs them.
- Test across boundaries. Unit-test mapping and authorization decisions; integration-test the BFF against its dependencies or controlled substitutes; test API responses, timeouts, throttling, malformed input, permission failures, and duplicate requests. Load-test dependencies as well as Lambda, because the slowest limit may be downstream.
- Operate from service signals. Track API Gateway request and integration latency and status codes alongside Lambda duration, errors, throttles, and concurrency. Correlate logs with a request identifier, but do not log secrets or unnecessary personal data. Add traces when they answer a concrete question about calls across services.
Our Lambda unit testing guide covers separating handler code from application logic.
Monitor the complete request path
For production diagnosis, see AWS Lambda metrics explained; standard metrics are useful, but the API boundary and downstream services need their own signals as well.
Benefits and trade-offs
A Lambda BFF can reduce frontend coordination, fit traffic that varies over time, and let teams deploy client-facing API changes without changing every backend service. The managed runtime removes server provisioning from the application team's routine work. You still own the API contract, permissions, dependency behavior, monitoring, deployment process, and cost controls.
Account for the additional boundary
The pattern adds another network hop and another component to debug. Lambda startup, dependency latency, API Gateway limits, concurrency quotas, and downstream capacity all affect the result. Multiple BFFs can also duplicate mapping code or drift in policy. Share stable domain logic through the domain services that own it, and keep client-specific code focused on composition.
Choose it when client needs justify a separate API
Compare the pattern with a direct API, a shared backend, or a container service based on the workload you actually have. A BFF is a useful boundary when client needs differ; it is not a default requirement for every frontend. A long-running process, a high and steady workload, or a service needing persistent connections may fit another compute model better.
Conclusion
A serverless BFF works well when it has a clear client contract and a narrow composition role. API Gateway and Lambda provide a managed HTTP-to-code path, while the application's reliability still depends on aligning timeouts, limiting fan-out, protecting downstream capacity, and enforcing authorization at the resource level.
Start with one user-visible request. Map its data to the services that own it, set a deadline for each dependency within the API's end-to-end budget, and test what happens when a dependency is slow or a client retries. Add more client-specific routes or functions when a real difference in contract, ownership, permissions, or traffic justifies them.
FAQs
How is a BFF different from an API proxy?
A proxy forwards or adapts requests to another endpoint. A BFF is designed around the needs of a client and can compose multiple domain responses into a deliberate client contract. A BFF may use proxy integrations, but forwarding requests alone does not provide client-specific composition or resource authorization.
Does Lambda scaling mean a BFF has unlimited capacity?
No. Lambda concurrency and scaling have quotas, and API Gateway and downstream services have their own limits. Scaling the function can increase load on databases and APIs. Set suitable concurrency and request limits, bound fan-out, and monitor throttles and dependency saturation.
Can I set the Lambda timeout to the API timeout?
They are separate settings. The request can time out at the client or API integration before Lambda reaches its own configured timeout, and the function may continue after the caller stops waiting. For interactive requests, keep Lambda and downstream deadlines inside the shortest request-path deadline. Use asynchronous work for operations that need longer.
Are API Gateway API keys enough to secure a BFF?
No. API keys can identify consumers for usage plans, but AWS advises against using them for authentication or authorization. Authenticate the user or calling service with a suitable mechanism and authorize each operation against the requested resource.