TIL / A stale-job reaper stops a crashed worker from wedging a pipeline forever
A stale-job reaper stops a crashed worker from wedging a pipeline forever
The problem
A multi-stage document pipeline marks each job processing when a worker picks it up and
done when it finishes. That’s fine as long as every worker finishes what it starts - but a
worker that gets OOM-killed, loses its pod, or hits an unhandled exception leaves the job
stuck at processing forever. Nothing else in the pipeline will ever touch it again, and it
silently disappears from throughput without ever showing up as a failure.
The fix
A periodic reaper query finds jobs that have been processing past a reasonable timeout and
resets them to pending (or failed, past a retry limit) so a healthy worker can pick them
back up.
UPDATE jobs
SET status = 'pending', attempts = attempts + 1
WHERE status = 'processing'
AND updated_at < now() - interval '15 minutes'
AND attempts < 3;
Gotcha
The timeout has to be longer than the slowest legitimate job, or the reaper starts fighting a worker that’s still genuinely working - pick it from real p99 duration, not a guess, and log every reap so a job that keeps getting reset (rather than completing) shows up as a real alert instead of quietly retrying forever.