Background jobs and task queues
Moving slow work off the request path, task queues and workers, retries with backoff, dead-letter queues, idempotent jobs, priorities, scheduled and delayed jobs, and how to keep a job system healthy.
Reading is half of it. See this used in a real interview: walk through Design a Distributed Job Scheduler →
Sending an email, resizing a photo, generating a report, charging a renewal, transcoding a video: none of these should make a user wait. Almost every design moves some work into background jobs, and interviewers then ask the follow-ups that matter: what happens when a job fails, runs twice, or never runs at all. This guide covers the patterns that answer them.
Why move work off the request path
- Latency: the user gets a response in milliseconds; the slow work happens afterwards.
- Reliability: a failing dependency (an email provider, a payment gateway) no longer fails the user’s request; the job retries later.
- Load smoothing: a burst of 10,000 uploads becomes a queue the workers drain at their own pace instead of a spike that overloads everything.
- Independent scaling: workers scale with queue depth, separately from web servers.
The rule of thumb: if the user does not need the result to see the next screen, it can be a job.
The basic shape
- The request handler validates input, writes what must be durable, and enqueues a job (an id and a few parameters, not a large payload).
- Workers pull jobs, do the work, and acknowledge.
- If the worker crashes or the job fails, the job becomes visible again and is retried.
Common tools: SQS, RabbitMQ, Redis-based queues (Sidekiq, BullMQ, Celery with Redis), Kafka for high-volume streams, and workflow engines (Temporal, Step Functions) for multi-step processes. See message queues and event streams for the differences.
Enqueue reliably. If the handler writes the database and then enqueues, a crash between the two loses the job. Use a transactional outbox: write a job row in the same transaction, and a relay moves it to the queue. Or let the job table itself be the queue for modest volumes. See event sourcing, CQRS and CDC.
Delivery is at least once, so jobs must be idempotent
A worker can finish the work and crash before acknowledging, so the job runs again. Queues therefore deliver at least once, and every job must be safe to run twice:
- Check-then-do with a durable marker: "if invoice 123 is already sent, stop".
- Use idempotency keys when calling external APIs (payment providers support them).
- Make writes upserts keyed by the job’s natural id.
See distributed transactions and idempotency.
Retries, backoff and dead letters
- Retry with exponential backoff and jitter: 1 s, 2 s, 4 s, 8 s… plus randomness, so a dependency outage does not produce synchronized retry storms.
- Cap attempts (say 5 to 10) and distinguish errors: retry timeouts and 5xx responses; do not retry validation errors or 4xx responses that will never succeed.
- Dead-letter queue (DLQ): jobs that exhaust retries go to a separate queue for inspection and replay, instead of retrying forever or vanishing. Alert on DLQ growth.
- Poison messages: a job that crashes the worker every time must not block the queue; the attempt cap plus DLQ handles it.
Visibility timeouts and long jobs
Most queues hide a job while a worker holds it (a visibility timeout or lease). If the worker does not finish or extend the lease in time, the job reappears for another worker. Set the timeout above the normal job duration, and have long jobs heartbeat to extend it. Otherwise a slow-but-alive job runs on two workers at once.
Priorities and fairness
- Use separate queues for different priorities (password reset emails before marketing newsletters) and have workers pull from high priority first, or give each queue dedicated workers so low priority still progresses.
- In multi-tenant systems, one customer’s 1 M-job import must not delay everyone else: per-tenant queues or fair scheduling with per-tenant concurrency limits.
- Rate limit jobs that call external APIs to stay within the provider’s limits.
Scheduled and delayed jobs
- Delayed jobs ("send a reminder in 24 hours"): most queues support a delay or a "run at" time; for long delays, store the job with a
run_atcolumn and have a scheduler enqueue due jobs. - Recurring jobs (cron): a scheduler must fire each occurrence exactly once even with several scheduler instances. Use leader election or a lock with fencing, or a database row per occurrence with a unique constraint. See Design a Distributed Job Scheduler.
- Batching: for huge recurring work ("renew 10 M subscriptions at midnight"), the scheduler enqueues per-item jobs gradually rather than one giant job.
Multi-step work
When a job is really a workflow (charge the card, then reserve stock, then email), chain jobs with state stored between steps, or use a workflow engine that persists each step’s result and handles retries, timers and compensation. This is the saga pattern in practice.
Keeping it healthy
Watch:
- Queue depth and the age of the oldest job (age is the clearest signal of falling behind).
- Processing time per job type, at p50 and p99.
- Failure and retry rates, and DLQ size.
- Worker utilisation, and autoscale workers on queue age or depth. See autoscaling and capacity planning.
In the interview
Say what is synchronous and what is a job, how the job is enqueued reliably, that jobs are idempotent because delivery is at least once, how retries and the DLQ work, and how you would know the queue is falling behind.
Checklist
- Which work leaves the request path.
- Reliable enqueue (outbox or job table).
- Idempotent jobs.
- Retries with backoff and jitter, attempt caps, a DLQ.
- Visibility timeouts and heartbeats for long jobs.
- Priorities and per-tenant fairness.
- Scheduled jobs that fire exactly once.
- Queue age, failure rate and DLQ alerts.