Learn / AWS Lambda for backend devs / Debugging and observability
Debugging and observability
Where Lambda logs actually go, structured logging, tracing a request across services, and the metrics worth alerting on.
Logs: automatic, but where do they land
Anything your function writes to stdout or stderr is automatically sent to CloudWatch
Logs, in a log group named /aws/lambda/<function-name> - no logging agent to install or
configure for the basics. Each invocation’s output is tagged with a request ID, which is the
thread to pull on when debugging one specific failed invocation.
console.log / print() works, but a raw free-text log line is hard to query at scale:
# Hard to query reliably at scale
print(f"User {user_id} request failed: {error}")
# Structured - queryable by field in CloudWatch Logs Insights
import json
print(json.dumps({"level": "error", "requestId": request_id, "userId": user_id, "error": str(error)}))
CloudWatch Logs Insights lets you run query-language searches across a log group - with
structured fields, you can filter level = "error" and group by userId directly; with
free-text logs, you’re stuck writing fragile regex against however each message happened to be
phrased.
Tracing a request across functions
A single log group shows you one function’s own output. Once a request fans out across multiple functions (an API Gateway function calling another function via SQS, which writes to DynamoDB and triggers a third function) no single log group shows the whole picture. AWS X-Ray (or an OpenTelemetry-based alternative, if your stack already standardizes on that) instruments the request with a trace ID that follows it across service boundaries, producing a timeline that shows exactly where the total time went - which hop was slow, which one failed.
Enable it per function (a small configuration flag), and instrument any outbound calls (HTTP requests, AWS SDK calls) you want visible as their own segments in the trace, not just the top-level invocation.
Metrics worth alerting on
CloudWatch automatically publishes per-function metrics; three matter most for catching problems before a user reports them:
- Errors / error rate - the direct signal something in the code is failing. Alert on rate, not raw count, so it scales with traffic.
- Throttles - invocations rejected because you’ve hit a concurrency limit (either the account-level limit or a function-level reserved concurrency setting). A throttle spike means real traffic is being dropped, not just slow - it’s a distinct failure mode from an error and needs its own alert.
- Duration approaching the configured timeout - a function that’s usually fast but occasionally runs close to its timeout (lesson 1’s 900-second ceiling) is heading toward timeout failures under slightly more load or a slightly slower downstream dependency; catch the trend before it becomes an outage.
A practical debugging flow
- Reproduce (or find) the failing request’s ID from a user report or an alert.
- Pull that request ID’s log lines from CloudWatch Logs (fast, if logs are structured with the request ID as a field).
- If the request crossed multiple functions/services, pull the same request’s X-Ray trace to see the full path and where time/errors occurred.
- Check whether it’s an isolated incident or a rate/duration trend on the relevant CloudWatch metric - that distinction decides whether the fix is “patch this one bug” or “this function needs more memory/concurrency/a timeout increase.”
Key takeaways
- Lambda sends everything written to stdout/stderr to CloudWatch Logs automatically, in a log group named /aws/lambda/<function-name> - no logging agent to install.
- Structured (JSON) logs are far easier to query in CloudWatch Logs Insights than free-text logs - log objects, not just interpolated strings.
- X-Ray (or an OpenTelemetry-based equivalent) traces a request across multiple Lambda functions and downstream services, which plain per-function logs can't show you on their own.
- The metrics worth alerting on: error rate, throttles (means you're hitting a concurrency limit), and duration approaching the configured timeout - each points at a different class of problem.
Quick check
3 questions - see how much stuck.