TIL / A queue consumer needs to catch SIGTERM, not just crash on it
A queue consumer needs to catch SIGTERM, not just crash on it
The problem
A rolling deploy or a KEDA scale-down sends SIGTERM to a pod, waits out
terminationGracePeriodSeconds, then sends SIGKILL. A queue consumer that has no signal
handler just gets killed mid-message: whatever it was processing is lost or, worse, left in a
half-applied state if the work wasn’t idempotent.
The fix
Install a SIGTERM handler that flips a flag, let the current message finish, stop pulling new
ones, and exit cleanly before the grace period runs out.
import signal
shutting_down = False
def handle_sigterm(signum, frame):
global shutting_down
shutting_down = True
signal.signal(signal.SIGTERM, handle_sigterm)
while not shutting_down:
message = queue.receive(wait_seconds=5)
if message:
process(message)
queue.delete(message)
Gotcha
The grace period has to be longer than the slowest single message can realistically take, or
Kubernetes sends SIGKILL anyway before the handler finishes - check
terminationGracePeriodSeconds against your actual p99 processing time, not the average.