Learn / AWS Lambda for backend devs / The execution model and real limits

Lesson 1 of 5 8 min

The execution model and real limits

How a Lambda invocation actually runs, and the quotas that shape design decisions - verified against current AWS docs.

The execution model, briefly

A Lambda function is packaged code plus configuration (runtime, memory, timeout, environment variables, IAM role). AWS runs it inside a managed, sandboxed execution environment - you never provision or patch the underlying host. Two states matter for how it behaves:

  • Cold start: no existing execution environment for this function version, so AWS provisions one from scratch (download code, start the runtime, run any top-level initialization in your code) before your handler runs.
  • Warm start: an execution environment from a previous invocation is reused - your handler runs immediately, skipping initialization.

Lesson 3 goes deep on why this distinction matters for latency-sensitive workloads. For now, the key model to hold onto: you write a handler function; AWS decides when and how many execution environments exist to run it, scaling them up and down based on concurrent invocations.

The quotas that actually shape design decisions

Verified against AWS’s Lambda quotas documentation (docs.aws.amazon.com/lambda/latest/dg/gettingstarted-limits.html):

QuotaValueWhy it matters
Max execution timeout900 seconds (15 min)A task that can run longer needs Step Functions or a different service - not “just increase the timeout”
Memory128 MB - 10,240 MB, 1 MB incrementsCPU scales with memory - a memory bump can also be a CPU/performance lever
/tmp ephemeral storage512 MB default, up to 10,240 MB, 1 MB incrementsRelevant for anything that downloads, unpacks, or buffers files mid-execution
Deployment package (zipped)50 MB via API/SDK, 50 MB via consolePush past this with container images or by uploading via S3 first
Deployment package (unzipped, incl. layers)250 MBThe real ceiling for how much code + dependencies you can ship without a container image
Synchronous invoke payload6 MB request and responseMatters for API Gateway-fronted functions returning large bodies
Synchronous invoke, streamed response200 MBResponse streaming raises the ceiling substantially for large payloads
Asynchronous invoke payload1 MBMuch tighter than sync - relevant for event-driven (S3, SNS, EventBridge-triggered) functions

What this rules in and out

  • Batch/ETL jobs longer than 15 minutes: not a single Lambda invocation. Break the work into chunks orchestrated by Step Functions, or move to a container/batch service.
  • A function that unpacks a large archive: check it fits in /tmp at the configured ephemeral storage size (default 512 MB is easy to exceed with, say, an ML model file).
  • A large ML dependency (numpy, pandas, a model file): the 250 MB unzipped limit is the real constraint - lesson 2 covers packaging strategies, including when a container image is the better fit than a zip deployment.
  • A CPU-heavy function that’s slow: before assuming you need to rewrite the algorithm, try raising memory - the proportional CPU increase is often the cheapest fix available.

Key takeaways

  • A Lambda function runs in a managed, ephemeral execution environment - your code doesn't control the underlying host, and the environment can be reused (a 'warm start') or freshly created (a 'cold start', see lesson 3).
  • Max execution timeout is 900 seconds (15 minutes) - anything longer needs Step Functions orchestration or a different compute service, not a bigger Lambda.
  • Memory is configurable from 128 MB to 10,240 MB in 1 MB increments, and CPU is allocated proportionally to memory - more memory also means more CPU, which can make a CPU-bound function faster, not just able to hold more data.
  • /tmp ephemeral storage defaults to 512 MB and is configurable up to 10,240 MB - relevant for anything that downloads or unpacks files during execution.

Quick check

3 questions - see how much stuck.

1. What is the maximum execution timeout for a single Lambda invocation?
2. What's the actual range for a Lambda function's configured memory?
3. Why would increasing a Lambda function's memory setting make a CPU-bound function run faster, even if it doesn't need more RAM?