AWS serverless architecture means building a system from managed services that scale with demand and bill per request or per unit of work, so there are no servers to patch, size or keep warm. The usual pieces are API Gateway at the edge, Lambda for compute, DynamoDB, S3 or Aurora Serverless for state, and EventBridge, SQS and Step Functions to connect them. The individual services are well documented; the hard part is the seams between them, where retries, ordering, concurrency limits and cost behave differently from a long-running server.

This article walks one order service end to end: the request path, the event path, how each kind of invocation retries, how to stop a burst from flattening a database, what it costs, and how it fails.

Serverless order service: a synchronous path and an asynchronous pathClientHTTPS + JWTAPI Gatewayauth, throttlingLambda: APIvalidate, read, writeDynamoDBorders tablesynchronous: caller waits, caller retriesDynamoDB Streamchange recordsLambda: publishermap to domain eventEventBridgerules route eventsSQS: fulfilmentvisibility, redriveSQS: emailindependent rateStep Functionsmulti-step workflowLambda: fulfilidempotent, batchedLambda: notifyemail providerAurora Serverless v2via RDS ProxyFulfilment DLQPer-function IAMleast privilegeS3uploads, exportsIaC (SAM / CDK) + CloudWatch, X-Ray
The request path stays short and synchronous; everything else flows from the table's change stream through EventBridge into per-consumer queues and workflows, each with its own retries and dead-letter queue.

What counts as serverless

A useful definition is operational, not marketing: a component is serverless when you never choose an instance size, it scales out without you, and an idle component costs little or nothing. Lambda, DynamoDB on-demand, S3, SQS, EventBridge, Step Functions and API Gateway fit. Aurora Serverless v2 is closer to autoscaling than to zero cost: it scales in capacity units and, since November 2024, can be configured to pause at 0 ACUs, but resuming takes on the order of 15 seconds, which a user-facing request will notice. Fargate removes servers but still bills for running tasks whether they are busy or not.

The architectural consequence is that compute becomes stateless and short-lived. A Lambda execution environment is reused between invocations when it can be, but you cannot rely on it; anything that must survive goes to a data store, a queue or a workflow.

What counts as serverless

A useful definition is operational, not marketing: a component is serverless when you never choose an instance size, it scales out without you, and an idle component costs little or nothing. Lambda, DynamoDB on-demand, S3, SQS, EventBridge, Step Functions and API Gateway fit. Aurora Serverless v2 is closer to autoscaling than to zero cost: it scales in capacity units and, since November 2024, can be configured to pause at 0 ACUs, but resuming takes on the order of 15 seconds, which a user-facing request will notice. Fargate removes servers but still bills for running tasks whether they are busy or not.

The architectural consequence is that compute becomes stateless and short-lived. A Lambda execution environment is reused between invocations when it can be, but you cannot rely on it; anything that must survive goes to a data store, a queue or a workflow.

The architecture, piece by piece

In the diagram, the synchronous path handles GET /orders/{id} and POST /orders. API Gateway validates the JWT, applies throttling, and invokes the API function, which reads or writes DynamoDB and returns. Nothing else happens in the request. The asynchronous path starts from the table's stream: a publisher function turns change records into domain events such as OrderPlaced and puts them on EventBridge. Rules route each event to its own SQS queue, one per consumer, so a slow email provider cannot hold up fulfilment. Multi-step processes with waits, compensation or human approval run in Step Functions rather than as a Lambda calling Lambdas. Every function has its own IAM role, and the whole stack is one template.

The synchronous path

Keep the synchronous function small and initialise clients outside the handler so warm invocations reuse connections:

import json, os, boto3

table = boto3.resource("dynamodb").Table(os.environ["TABLE"])   # created once per environment

def get_order(event, context):
    order_id = event["pathParameters"]["id"]
    item = table.get_item(Key={"pk": f"order#{order_id}"}).get("Item")
    if item is None:
        return {"statusCode": 404, "body": json.dumps({"error": "not found"})}
    return {"statusCode": 200, "body": json.dumps(item, default=str)}

A warm invocation of this function is dominated by the DynamoDB round trip, single-digit milliseconds. A cold start adds environment creation, runtime start and your initialisation code, from a fraction of a second for small interpreted functions to several seconds for heavy JVM ones without SnapStart. Treat that as a tail-latency issue and measure it; Lambda cold starts covers the mitigations. Synchronous callers also inherit API Gateway's integration timeout, so long work belongs on the asynchronous path.

The event path

The event path deserves as much design as the API. DynamoDB Streams delivers each change at least once, in order for a given item key, and keeps records for 24 hours, so a consumer that is broken for longer than that loses data. The publisher function should translate raw change records into stable domain events rather than leaking the table layout to every consumer; then the table can be remodelled without breaking anyone. Give each event a type, a version, an ID that consumers can use for deduplication, and the correlation ID from the original request.

EventBridge rules match on event content. The fulfilment rule only wants paid orders:

{
  "source": ["shop.orders"],
  "detail-type": ["OrderPlaced"],
  "detail": {
    "payment_status": ["PAID"],
    "schema_version": [{"numeric": [">=", 2]}]
  }
}

Routing each rule to its own queue rather than straight to a function buys three things: a buffer when the consumer is slow or down, an independent retry policy and DLQ per consumer, and a place to cap concurrency. EventBridge archives and replay are worth enabling, because replaying a day of events into a fixed consumer is the cleanest recovery after a bug.

Choosing where state lives

Choose state stores by access pattern, not by habit. DynamoDB suits key-value and single-table access with predictable queries, scales without connection limits and pairs naturally with Lambda; ad hoc queries and joins are where it hurts. Aurora Serverless v2 gives SQL and transactions across tables, but connections are a finite resource, so put RDS Proxy in front and keep the minimum capacity above zero for anything user-facing. S3 holds large objects and event archives and can trigger processing on upload. Many systems use DynamoDB for the operational path and stream changes to S3 for analytics with Athena.

Retries, duplicates and idempotency

The single most important serverless fact is that each invocation type retries differently:

Invocation typeExamplesWho retriesWhere failures end up
SynchronousAPI Gateway, SDK Invokethe callerreturned to the caller
AsynchronousEventBridge, S3, SNSLambda: 2 retries by defaulton-failure destination or DLQ
Poll-basedSQS, DynamoDB/Kinesis streamsuntil success, expiry or redriveSQS redrive DLQ, stream failure destination

Two consequences follow. Every consumer will occasionally see the same event twice, so handlers must be idempotent. And for poll-based sources, one bad message can block or re-run a whole batch unless the handler reports which items failed. With ReportBatchItemFailures enabled, only the listed messages return to the queue:

import json, boto3
from botocore.exceptions import ClientError

done = boto3.resource("dynamodb").Table("fulfilment-done")

def fulfil(event, context):
    failures = []
    for rec in event["Records"]:
        msg = json.loads(rec["body"])
        key = f"fulfil#{msg['order_id']}"
        try:
            if "Item" in done.get_item(Key={"pk": key}):
                continue                                  # already handled: a duplicate delivery
            ship(msg, idempotency_key=key)                # downstream API also deduplicates on key
            done.put_item(Item={"pk": key},
                          ConditionExpression="attribute_not_exists(pk)")
        except ClientError as e:
            if e.response["Error"]["Code"] != "ConditionalCheckFailedException":
                failures.append({"itemIdentifier": rec["messageId"]})
        except Exception:
            failures.append({"itemIdentifier": rec["messageId"]})
    return {"batchItemFailures": failures}

The record of completion is written after the side effect, and the side effect itself carries an idempotency key, because a crash between the two would otherwise ship twice. Where the downstream cannot deduplicate, the Powertools for AWS Lambda idempotency utility implements the in-progress and completed states for you.

Concurrency and backpressure

Lambda scales by running more concurrent environments, up to a per-region account quota shared by every function, so check the current quota and per-function scaling rate in your account rather than assuming. That elasticity is a weapon pointed at whatever sits downstream. A queue of 200,000 messages can wake hundreds of consumers that each open a database connection.

  • Maximum concurrency on the SQS event source caps how many environments that queue may drive, without throttling the function for other callers.
  • Reserved concurrency both guarantees and caps a function's share of the account pool; use it for critical functions and to stop one function starving the rest.
  • Visibility timeout on the queue should be at least six times the function timeout, as AWS recommends, so messages are not redelivered while a slow batch is still running.
  • RDS Proxy pools connections in front of Aurora so a burst of environments shares a bounded set of database connections.
  • API Gateway throttling and usage plans shed load at the edge before it costs compute.

Infrastructure as code

Define everything in one template so the wiring is reviewable. In AWS SAM the fulfilment consumer reads:

Transform: AWS::Serverless-2016-10-31
Resources:
  FulfilDLQ:
    Type: AWS::SQS::Queue
  FulfilQueue:
    Type: AWS::SQS::Queue
    Properties:
      VisibilityTimeout: 180            # 6 x function timeout
      RedrivePolicy:
        deadLetterTargetArn: !GetAtt FulfilDLQ.Arn
        maxReceiveCount: 5
  Fulfil:
    Type: AWS::Serverless::Function
    Properties:
      Runtime: python3.12
      Handler: app.fulfil
      Timeout: 30
      MemorySize: 512
      Events:
        Orders:
          Type: SQS
          Properties:
            Queue: !GetAtt FulfilQueue.Arn
            BatchSize: 10
            FunctionResponseTypes: [ReportBatchItemFailures]
            ScalingConfig:
              MaximumConcurrency: 20

The same file grants the function exactly the table and queue permissions it needs and nothing else; per-function roles are what make a compromised function a small incident.

Worked example: what it costs

Lambda bills requests plus GB-seconds of execution. Work an example with the us-east-1 x86 list prices at the time of writing, $0.20 per million requests and about $0.0000166667 per GB-second, and check the pricing page before relying on them. Suppose the API serves 50 million requests a month, each running 120 ms at 512 MB:

compute = 50e6 requests x 0.120 s x 0.5 GB  = 3.0M GB-s  x $0.0000166667 = $50
requests = 50 x $0.20                                                      = $10
total Lambda                                                               = $60 / month

That is cheap for bursty traffic. Now hold 1,000 requests per second around the clock, about 2.6 billion a month: the same arithmetic gives roughly $2,600 in compute and $520 in requests, before API Gateway, which also bills per request. A small always-on container fleet sized for that steady load often costs less. The crossover depends on duty cycle: serverless wins when utilisation is spiky or low and loses when it is flat and high. Memory also buys CPU in Lambda, so a larger setting can finish faster and cost the same or less; measure rather than guess.

Failure modes

  • Poison messages retried forever; fix with partial batch responses, a redrive policy and an alarm on DLQ depth.
  • Duplicate side effects from at-least-once delivery; fix with idempotency keys.
  • Recursive loops, such as a function writing to the bucket that triggers it; separate prefixes or buckets, and keep Lambda's recursive loop detection enabled.
  • Downstream meltdown when concurrency outruns a database or third-party API; cap it at the event source.
  • Throttling surprises when one function consumes the account's concurrency; use reserved concurrency and alarm on throttles.
  • Lambda-calls-Lambda chains that pay for waiting and lose errors; move orchestration into Step Functions.
  • Hot partitions in DynamoDB from a low-cardinality key; design keys for spread.

Operating it

Alarm on errors, throttles, iterator age for streams, approximate age of oldest message for queues and DLQ depth; those catch nearly every async incident. Use structured JSON logs with a correlation ID passed through events, and X-Ray or OpenTelemetry tracing across API Gateway, Lambda and downstream calls. Deploy with aliases and gradual traffic shifting so a bad version rolls back automatically. For the building blocks see Step Functions, EventBridge, SQS and DynamoDB.

Trade-offs

Serverless removes capacity planning, patching and idle cost and gives fine-grained isolation per function. It costs you control: hard execution time limits, cold-start tail latency, per-request pricing that punishes steady high throughput, and a distributed system whose behaviour lives in configuration as much as in code. Teams that succeed treat event contracts, retry semantics and concurrency limits as first-class design, and move individual hot paths to containers when the numbers say so.

What to do next

  1. Draw your system and label every arrow synchronous, asynchronous or poll-based.
  2. Make every asynchronous consumer idempotent and give every queue a DLQ with an alarm.
  3. Enable partial batch responses on every SQS and stream consumer.
  4. Cap concurrency at the event source for anything that touches a database or external API.
  5. Model monthly cost at today's load and at ten times it, and compare with a container baseline.
  6. Put the whole stack in SAM or CDK, with one IAM role per function.
  7. Move any function that waits on other functions into Step Functions.
Key takeaway: A serverless system on AWS is defined by its seams: keep the request path short, push everything else through events and queues, make consumers idempotent, cap concurrency where it meets stateful systems, and check the cost curve against a container baseline as load grows.