A Lambda cold start is the work Lambda does when a request arrives and no idle, initialized execution environment exists for that function version: fetch code, start a microVM, start extensions and the language runtime, and run your initialization code before the handler. AWS says cold starts typically affect under 1% of invocations and range from under 100 ms to over a second, but that 1% lands on the requests at the worst moment, during scale-ups, deploys and the first request after a quiet period.

This article looks at the platform side: what the environment lifecycle is, which limits and billing rules apply to each phase, how many cold starts a traffic pattern will produce, and how provisioned concurrency and SnapStart change the architecture. If you need to find which lines of your initialization code are slow, read the companion guide to measuring and profiling cold starts alongside this one.

Where a Lambda cold start happensCallerAPI, event sourceInvoke frontendauth, limits, routingEnvironment lookupidle warm one free?Warm paththaw, run handlerCold pathnew environmentyesnoWorker host: one Firecracker microVM per execution environmentCode fetchZIP or image chunksExtension initexternal processesRuntime initJVM, Python, NodeFunction inityour top-level codeInvokehandler, then freezeLeversProvisioned concurrency:init done ahead of timeSnapStart:restore a snapshotYour code:make function init cheapInit: 10 s limit on demand, billed since 1 Aug 2025; up to 15 min with PC or SnapStart.
A cold start happens when no idle environment exists for the version; the init steps run inside a new microVM before the handler can start.

Execution environments: the unit of a cold start

Lambda runs each function in execution environments: isolated sandboxes, each inside a Firecracker microVM on a worker host, holding your code, the runtime and any extensions. In the standard model an environment serves one request at a time, so concurrency, the number of requests in flight, equals the number of busy environments. (Lambda Managed Instances, which run on capacity in your account, follow a different model with concurrent invocations per environment and are not covered here.)

When an invocation arrives, the frontend authenticates it, applies concurrency limits and looks for an idle environment for that function version. If one exists, Lambda thaws it and runs the handler; module-level objects such as SDK clients and connection pools are still in memory. If none exists, Lambda creates one, and that is the cold start. Idle environments are kept for a while and then reclaimed, and Lambda also recycles environments every few hours for maintenance even under steady traffic, so a small trickle of cold starts never fully disappears. AWS does not publish the idle timeout, and you should not design around a guessed value.

Each published version and $LATEST has its own environments. A deploy that shifts an alias to a new version starts that version with zero warm environments, which is why the minutes after a deploy show a burst of cold starts.

Execution environments: the unit of a cold start

Lambda runs each function in execution environments: isolated sandboxes, each inside a Firecracker microVM on a worker host, holding your code, the runtime and any extensions. In the standard model an environment serves one request at a time, so concurrency, the number of requests in flight, equals the number of busy environments. (Lambda Managed Instances, which run on capacity in your account, follow a different model with concurrent invocations per environment and are not covered here.)

When an invocation arrives, the frontend authenticates it, applies concurrency limits and looks for an idle environment for that function version. If one exists, Lambda thaws it and runs the handler; module-level objects such as SDK clients and connection pools are still in memory. If none exists, Lambda creates one, and that is the cold start. Idle environments are kept for a while and then reclaimed, and Lambda also recycles environments every few hours for maintenance even under steady traffic, so a small trickle of cold starts never fully disappears. AWS does not publish the idle timeout, and you should not design around a guessed value.

Each published version and $LATEST has its own environments. A deploy that shifts an alias to a new version starts that version with zero warm environments, which is why the minutes after a deploy show a burst of cold starts.

The lifecycle phases and their limits

The lifecycle has three phases, plus a fourth for SnapStart, and each has its own limits.

PhaseWhat runsLimit and billing
InitExtension init, runtime init, function init (your top-level code); before-snapshot hooks with SnapStart10 s on demand; if exceeded, Lambda retries init at the first invoke under the function timeout. Up to 15 min with provisioned concurrency or SnapStart. Billed for all configurations since 1 Aug 2025
Restore (SnapStart only)Resume from the cached snapshot, then after-restore hooksRuntime load plus hooks within 10 s, else SnapStartTimeoutException; hooks are billed
InvokeYour handler, plus extensions until they signal doneFunction timeout, up to 900 s
ShutdownShutdown event to extensions0 ms with no extensions, 500 ms with internal, 2,000 ms with external ones

Three details change how you design. First, until 1 August 2025 the init phase of on-demand ZIP functions on managed runtimes was not billed, which made heavy initialization look free. It is now billed like invoke duration, so a 2-second init costs the same as 2 seconds of handler time. Second, if an invoke crashes or times out, Lambda resets the environment and re-runs init on the next request in that environment, a suppressed init that does not get its own log line; its time appears inside that request's duration. A function that crashes often is therefore also a function with hidden cold starts. Third, extensions run during init and can delay readiness; an observability extension that takes 300 ms to start adds 300 ms to every cold start.

Code loading: ZIP packages and container images

Where your code comes from affects the first step. ZIP packages are limited to 250 MB unzipped including layers; container images can be up to 10 GB. That makes it tempting to assume large images start slowly, but AWS's published design for image functions (the USENIX ATC 2023 paper On-demand Container Loading in AWS Lambda) loads image contents lazily in chunks, deduplicates common chunks across customers, and caches them in several tiers. What costs you is the bytes actually read during startup, not the size of the image. An image that imports a large ML library at init pays for reading it; an image that carries the library but loads it lazily on the first request that needs it does not.

How many cold starts will you get?

Cold starts are a capacity question. By Little's law, concurrency equals request rate times duration: 500 requests per second at 80 ms needs about 40 environments. Cold starts happen whenever concurrency rises above the number of warm environments, whenever environments are recycled, and after every deploy. A steady workload therefore sees few cold starts, while a spiky one sees one per environment added on each spike.

Lambda also limits how fast a single function can scale: each function in each Region can add 1,000 execution environments every 10 seconds, refilled continuously. Requests beyond that, or beyond reserved or account concurrency, are throttled with a 429. Synchronous callers see the error; asynchronous and event source invocations are retried by Lambda or the poller. A traffic step from 0 to 5,000 concurrent requests therefore takes at least 50 seconds to absorb, and every environment created during it is a cold start.

A small simulation makes the trade-off visible before you buy anything. It tracks warm environments, counts cold starts for a given arrival trace, and lets you test a provisioned floor. The idle reclaim time is a parameter precisely because AWS does not publish it; try several values.

import heapq, random

def simulate(arrivals, duration_s, init_s, idle_reclaim_s=600, provisioned=0):
    """arrivals: sorted request timestamps (s). Returns (cold_starts, total)."""
    # idle: heap of (became_idle_at, env_id) for free environments
    busy = []                   # (free_at, env_id) heap of running environments
    pinned = set(range(provisioned))        # pre-initialized, never reclaimed
    idle = [(0.0, env) for env in pinned]
    next_id, cold = provisioned, 0
    for t in arrivals:
        while busy and busy[0][0] <= t:     # finished requests free their environments
            free_at, env = heapq.heappop(busy)
            heapq.heappush(idle, (free_at, env))
        # drop on-demand environments idle longer than the reclaim window
        idle = [(s, e) for (s, e) in idle if e in pinned or t - s < idle_reclaim_s]
        heapq.heapify(idle)
        if idle:
            _, env = heapq.heappop(idle)
            heapq.heappush(busy, (t + duration_s, env))
        else:
            cold += 1
            heapq.heappush(busy, (t + init_s + duration_s, next_id)); next_id += 1
    return cold, len(arrivals)

# 5 req/s baseline with a 60 s burst at 200 req/s
trace, t = [], 0.0
while t < 3600:
    rate = 200 if 1800 <= t < 1860 else 5
    t += random.expovariate(rate)
    trace.append(t)
for pc in (0, 10, 30):
    print(pc, simulate(trace, duration_s=0.08, init_s=1.5, provisioned=pc))

Provisioned concurrency and SnapStart as architecture

Provisioned concurrency keeps a configured number of environments initialized for a version or alias. Init runs when you set the configuration, not on a request, and requests above the provisioned level spill over to normal on-demand environments, which can be cold. You pay for the configured amount for as long as it is set, so schedule it or drive it with Application Auto Scaling rather than holding peak capacity all day. It is the tool for strict latency targets.

SnapStart runs init once when you publish a version, takes an encrypted Firecracker snapshot of the initialized memory and disk state, caches it, and creates new environments by restoring that snapshot. It supports Java 11 and later, Python 3.12 and later and .NET 8 and later, on published versions and aliases only, never $LATEST. It cannot be combined with provisioned concurrency, Amazon EFS, or ephemeral storage above 512 MB, so these are alternatives for a given version, not layers to stack. For Java managed runtimes there is no extra charge; for Python and .NET you pay for snapshot caching (minimum 3 hours, per published version while it stays active) and for each restore.

The catch is that one snapshot becomes the starting state for many environments. Anything unique created during init, such as random seeds, UUIDs, temporary credentials or open connections, is now shared or stale. Runtime hooks let you fix that.

# Python 3.12+ with SnapStart: hooks from the snapshot-restore-py library
import boto3, random
from snapshot_restore_py import register_before_snapshot, register_after_restore

s3 = boto3.client("s3")          # SDK clients usually re-establish connections on their own
db = connect()                   # opened during init, so it would be in the snapshot

@register_before_snapshot
def before_snapshot():
    global db
    db.close()                   # never snapshot a live socket
    db = None

@register_after_restore
def after_restore():
    global db
    random.seed()                # snapshot froze the PRNG state; reseed per environment
    db = connect()               # fresh connection per restored environment

def lambda_handler(event, context):
    return handle(event, db)
# turn SnapStart on and publish, or provision a warm floor on an alias
aws lambda update-function-configuration --function-name orders-api \
    --snap-start ApplyOn=PublishedVersions
aws lambda publish-version --function-name orders-api

aws lambda put-provisioned-concurrency-config --function-name orders-api \
    --qualifier live --provisioned-concurrent-executions 20

Design choices that shape cold starts

  • Memory buys CPU. Lambda allocates CPU in proportion to memory, so init-heavy runtimes such as the JVM often start much faster at higher memory, sometimes at similar total cost. Test, do not assume.
  • Fewer environments, fewer cold starts. For queue consumers, a batch size above 1 on the event source mapping means fewer concurrent environments for the same throughput; see Amazon SQS architecture for batching and partial failures.
  • VPC attachment is no longer per cold start. Since 2019, Lambda creates shared Hyperplane network interfaces when a VPC function is configured, not at each cold start. Subnet IP capacity and security group design still matter.
  • Deploys are scale-from-zero events. Weighted alias shifting warms a new version gradually; with SnapStart, publishing pays the init cost once instead of on user requests.
  • Know when Lambda is the wrong shape. If a service needs many requests per instance or multi-second warm-up, a container platform with request concurrency may fit better; compare Cloud Run concurrency.

Worked example: a morning batch burst

An internal pricing API on Python 3.12 serves 50 requests per second in office hours with an 80 ms handler, so it averages about 4 concurrent environments. Init takes 1.5 s, mostly importing a data science library and loading a 200 MB reference table. Every morning a batch job fires 400 requests per second for two minutes, and those requests carry a 300 ms latency target.

At 400 per second and 80 ms, concurrency rises to about 32, so roughly 28 new environments start in the first second of the burst, well within the 1,000 per 10 seconds scaling rate. Each of those 28 first requests waits about 1.6 s. The fix options compare as follows. Provisioned concurrency of 35 on the alias, scheduled for the batch window, removes the cold starts for the price of 35 warm environments for that window. SnapStart replaces the 1.5 s init with a restore; the table load becomes part of the snapshot, as long as the reference data does not need refreshing per environment. Code changes, such as importing only the needed submodules and reading the table with memory mapping, shrink init for every invoke path and are worth doing first. The team chose code changes plus a scheduled provisioned floor, because the reference table refreshes hourly and would otherwise go stale inside a snapshot, and verified the result in the next window's REPORT lines.

Failure modes

  • Init timeouts on demand. Init over 10 s is retried at invoke under the function timeout; callers see slow first requests and sometimes timeouts. Fix: cut init, or use SnapStart or provisioned concurrency.
  • Crash loops hiding as latency. Each invoke failure resets the environment and triggers a suppressed init. Fix: watch error rate and duration together.
  • Spillover past the provisioned floor. Traffic above the configured level runs on cold on-demand environments. Fix: alert on spillover invocations and autoscale the floor.
  • Snapshot uniqueness bugs. Shared random seeds, duplicate IDs or stale credentials after restore. Fix: generate unique state in after-restore hooks or the handler.
  • Throttling during steps. Bursts beyond the scaling rate or reserved concurrency return 429. Fix: queue in front of Lambda, or raise limits before planned events.

Trade-offs

LeverRemovesCostsDoes not fix
Leaner init codePart of every cold startEngineering timePlatform and runtime start
SnapStartMost init time on restoreCaching and restore fees (not on Java); uniqueness careRestore time; per-environment fresh state
Provisioned concurrencyCold starts up to the floorPaying for idle environmentsSpillover above the floor
More memoryCPU-bound init timeHigher price per msI/O-bound init

For the wider picture of event-driven designs on AWS see AWS serverless architecture in depth.

What to do next

  1. Compute concurrency per function as request rate times duration, and plot it against warm environments over a day.
  2. Pull Init Duration from REPORT lines and find which functions pay the most total init time now that it is billed.
  3. Remove unused imports and lazy-load anything only some requests need; re-measure.
  4. Check that init stays well under 10 s on demand, and that extensions are not adding startup time.
  5. For Java, Python 3.12+ or .NET 8+, trial SnapStart on a published version and add restore hooks for uniqueness.
  6. Where a latency target is strict and traffic is predictable, schedule provisioned concurrency and alert on spillover.
  7. Rehearse a deploy and a traffic step, and confirm the scaling rate and throttling behave as you expect.
Key takeaway: A cold start is one new execution environment: a microVM, extensions, runtime and your init code, billed since August 2025 and capped at 10 s on demand. Their number follows from concurrency changes, deploys and recycling, so model them with Little's law, shrink init first, then pick SnapStart or provisioned concurrency per version, knowing the two cannot be combined.