AWS serverless architecture means building a system from managed services that scale with demand and bill per request or per unit of work, so there are no servers to patch, size or keep warm. The usual pieces are API Gateway at the edge, Lambda for compute, DynamoDB, S3 or Aurora Serverless for state, and EventBridge, SQS and Step Functions to connect them. The individual services are well documented; the hard part is the seams between them, where retries, ordering, concurrency limits and cost behave differently from a long-running server.
This article walks one order service end to end: the request path, the event path, how each kind of invocation retries, how to stop a burst from flattening a database, what it costs, and how it fails.
What counts as serverless
A useful definition is operational, not marketing: a component is serverless when you never choose an instance size, it scales out without you, and an idle component costs little or nothing. Lambda, DynamoDB on-demand, S3, SQS, EventBridge, Step Functions and API Gateway fit. Aurora Serverless v2 is closer to autoscaling than to zero cost: it scales in capacity units and, since November 2024, can be configured to pause at 0 ACUs, but resuming takes on the order of 15 seconds, which a user-facing request will notice. Fargate removes servers but still bills for running tasks whether they are busy or not.
The architectural consequence is that compute becomes stateless and short-lived. A Lambda execution environment is reused between invocations when it can be, but you cannot rely on it; anything that must survive goes to a data store, a queue or a workflow.
What counts as serverless
A useful definition is operational, not marketing: a component is serverless when you never choose an instance size, it scales out without you, and an idle component costs little or nothing. Lambda, DynamoDB on-demand, S3, SQS, EventBridge, Step Functions and API Gateway fit. Aurora Serverless v2 is closer to autoscaling than to zero cost: it scales in capacity units and, since November 2024, can be configured to pause at 0 ACUs, but resuming takes on the order of 15 seconds, which a user-facing request will notice. Fargate removes servers but still bills for running tasks whether they are busy or not.
The architectural consequence is that compute becomes stateless and short-lived. A Lambda execution environment is reused between invocations when it can be, but you cannot rely on it; anything that must survive goes to a data store, a queue or a workflow.
The architecture, piece by piece
In the diagram, the synchronous path handles GET /orders/{id} and POST /orders. API Gateway validates the JWT, applies throttling, and invokes the API function, which reads or writes DynamoDB and returns. Nothing else happens in the request. The asynchronous path starts from the table's stream: a publisher function turns change records into domain events such as OrderPlaced and puts them on EventBridge. Rules route each event to its own SQS queue, one per consumer, so a slow email provider cannot hold up fulfilment. Multi-step processes with waits, compensation or human approval run in Step Functions rather than as a Lambda calling Lambdas. Every function has its own IAM role, and the whole stack is one template.
The synchronous path
Keep the synchronous function small and initialise clients outside the handler so warm invocations reuse connections:
import json, os, boto3
table = boto3.resource("dynamodb").Table(os.environ["TABLE"]) # created once per environment
def get_order(event, context):
order_id = event["pathParameters"]["id"]
item = table.get_item(Key={"pk": f"order#{order_id}"}).get("Item")
if item is None:
return {"statusCode": 404, "body": json.dumps({"error": "not found"})}
return {"statusCode": 200, "body": json.dumps(item, default=str)}A warm invocation of this function is dominated by the DynamoDB round trip, single-digit milliseconds. A cold start adds environment creation, runtime start and your initialisation code, from a fraction of a second for small interpreted functions to several seconds for heavy JVM ones without SnapStart. Treat that as a tail-latency issue and measure it; Lambda cold starts covers the mitigations. Synchronous callers also inherit API Gateway's integration timeout, so long work belongs on the asynchronous path.
The event path
The event path deserves as much design as the API. DynamoDB Streams delivers each change at least once, in order for a given item key, and keeps records for 24 hours, so a consumer that is broken for longer than that loses data. The publisher function should translate raw change records into stable domain events rather than leaking the table layout to every consumer; then the table can be remodelled without breaking anyone. Give each event a type, a version, an ID that consumers can use for deduplication, and the correlation ID from the original request.
EventBridge rules match on event content. The fulfilment rule only wants paid orders:
{
"source": ["shop.orders"],
"detail-type": ["OrderPlaced"],
"detail": {
"payment_status": ["PAID"],
"schema_version": [{"numeric": [">=", 2]}]
}
}Routing each rule to its own queue rather than straight to a function buys three things: a buffer when the consumer is slow or down, an independent retry policy and DLQ per consumer, and a place to cap concurrency. EventBridge archives and replay are worth enabling, because replaying a day of events into a fixed consumer is the cleanest recovery after a bug.
Choosing where state lives
Choose state stores by access pattern, not by habit. DynamoDB suits key-value and single-table access with predictable queries, scales without connection limits and pairs naturally with Lambda; ad hoc queries and joins are where it hurts. Aurora Serverless v2 gives SQL and transactions across tables, but connections are a finite resource, so put RDS Proxy in front and keep the minimum capacity above zero for anything user-facing. S3 holds large objects and event archives and can trigger processing on upload. Many systems use DynamoDB for the operational path and stream changes to S3 for analytics with Athena.
Retries, duplicates and idempotency
The single most important serverless fact is that each invocation type retries differently:
| Invocation type | Examples | Who retries | Where failures end up |
|---|---|---|---|
| Synchronous | API Gateway, SDK Invoke | the caller | returned to the caller |
| Asynchronous | EventBridge, S3, SNS | Lambda: 2 retries by default | on-failure destination or DLQ |
| Poll-based | SQS, DynamoDB/Kinesis streams | until success, expiry or redrive | SQS redrive DLQ, stream failure destination |
Two consequences follow. Every consumer will occasionally see the same event twice, so handlers must be idempotent. And for poll-based sources, one bad message can block or re-run a whole batch unless the handler reports which items failed. With ReportBatchItemFailures enabled, only the listed messages return to the queue:
import json, boto3
from botocore.exceptions import ClientError
done = boto3.resource("dynamodb").Table("fulfilment-done")
def fulfil(event, context):
failures = []
for rec in event["Records"]:
msg = json.loads(rec["body"])
key = f"fulfil#{msg['order_id']}"
try:
if "Item" in done.get_item(Key={"pk": key}):
continue # already handled: a duplicate delivery
ship(msg, idempotency_key=key) # downstream API also deduplicates on key
done.put_item(Item={"pk": key},
ConditionExpression="attribute_not_exists(pk)")
except ClientError as e:
if e.response["Error"]["Code"] != "ConditionalCheckFailedException":
failures.append({"itemIdentifier": rec["messageId"]})
except Exception:
failures.append({"itemIdentifier": rec["messageId"]})
return {"batchItemFailures": failures}The record of completion is written after the side effect, and the side effect itself carries an idempotency key, because a crash between the two would otherwise ship twice. Where the downstream cannot deduplicate, the Powertools for AWS Lambda idempotency utility implements the in-progress and completed states for you.
Concurrency and backpressure
Lambda scales by running more concurrent environments, up to a per-region account quota shared by every function, so check the current quota and per-function scaling rate in your account rather than assuming. That elasticity is a weapon pointed at whatever sits downstream. A queue of 200,000 messages can wake hundreds of consumers that each open a database connection.
- Maximum concurrency on the SQS event source caps how many environments that queue may drive, without throttling the function for other callers.
- Reserved concurrency both guarantees and caps a function's share of the account pool; use it for critical functions and to stop one function starving the rest.
- Visibility timeout on the queue should be at least six times the function timeout, as AWS recommends, so messages are not redelivered while a slow batch is still running.
- RDS Proxy pools connections in front of Aurora so a burst of environments shares a bounded set of database connections.
- API Gateway throttling and usage plans shed load at the edge before it costs compute.
Infrastructure as code
Define everything in one template so the wiring is reviewable. In AWS SAM the fulfilment consumer reads:
Transform: AWS::Serverless-2016-10-31
Resources:
FulfilDLQ:
Type: AWS::SQS::Queue
FulfilQueue:
Type: AWS::SQS::Queue
Properties:
VisibilityTimeout: 180 # 6 x function timeout
RedrivePolicy:
deadLetterTargetArn: !GetAtt FulfilDLQ.Arn
maxReceiveCount: 5
Fulfil:
Type: AWS::Serverless::Function
Properties:
Runtime: python3.12
Handler: app.fulfil
Timeout: 30
MemorySize: 512
Events:
Orders:
Type: SQS
Properties:
Queue: !GetAtt FulfilQueue.Arn
BatchSize: 10
FunctionResponseTypes: [ReportBatchItemFailures]
ScalingConfig:
MaximumConcurrency: 20The same file grants the function exactly the table and queue permissions it needs and nothing else; per-function roles are what make a compromised function a small incident.
Worked example: what it costs
Lambda bills requests plus GB-seconds of execution. Work an example with the us-east-1 x86 list prices at the time of writing, $0.20 per million requests and about $0.0000166667 per GB-second, and check the pricing page before relying on them. Suppose the API serves 50 million requests a month, each running 120 ms at 512 MB:
compute = 50e6 requests x 0.120 s x 0.5 GB = 3.0M GB-s x $0.0000166667 = $50
requests = 50 x $0.20 = $10
total Lambda = $60 / monthThat is cheap for bursty traffic. Now hold 1,000 requests per second around the clock, about 2.6 billion a month: the same arithmetic gives roughly $2,600 in compute and $520 in requests, before API Gateway, which also bills per request. A small always-on container fleet sized for that steady load often costs less. The crossover depends on duty cycle: serverless wins when utilisation is spiky or low and loses when it is flat and high. Memory also buys CPU in Lambda, so a larger setting can finish faster and cost the same or less; measure rather than guess.
Failure modes
- Poison messages retried forever; fix with partial batch responses, a redrive policy and an alarm on DLQ depth.
- Duplicate side effects from at-least-once delivery; fix with idempotency keys.
- Recursive loops, such as a function writing to the bucket that triggers it; separate prefixes or buckets, and keep Lambda's recursive loop detection enabled.
- Downstream meltdown when concurrency outruns a database or third-party API; cap it at the event source.
- Throttling surprises when one function consumes the account's concurrency; use reserved concurrency and alarm on throttles.
- Lambda-calls-Lambda chains that pay for waiting and lose errors; move orchestration into Step Functions.
- Hot partitions in DynamoDB from a low-cardinality key; design keys for spread.
Operating it
Alarm on errors, throttles, iterator age for streams, approximate age of oldest message for queues and DLQ depth; those catch nearly every async incident. Use structured JSON logs with a correlation ID passed through events, and X-Ray or OpenTelemetry tracing across API Gateway, Lambda and downstream calls. Deploy with aliases and gradual traffic shifting so a bad version rolls back automatically. For the building blocks see Step Functions, EventBridge, SQS and DynamoDB.
Trade-offs
Serverless removes capacity planning, patching and idle cost and gives fine-grained isolation per function. It costs you control: hard execution time limits, cold-start tail latency, per-request pricing that punishes steady high throughput, and a distributed system whose behaviour lives in configuration as much as in code. Teams that succeed treat event contracts, retry semantics and concurrency limits as first-class design, and move individual hot paths to containers when the numbers say so.
What to do next
- Draw your system and label every arrow synchronous, asynchronous or poll-based.
- Make every asynchronous consumer idempotent and give every queue a DLQ with an alarm.
- Enable partial batch responses on every SQS and stream consumer.
- Cap concurrency at the event source for anything that touches a database or external API.
- Model monthly cost at today's load and at ten times it, and compare with a container baseline.
- Put the whole stack in SAM or CDK, with one IAM role per function.
- Move any function that waits on other functions into Step Functions.