Delayed jobs, timers and distributed cron
Running work at a specific time across a cluster: cron at scale, delayed queues, time-bucketed tables, timing wheels, leader election and partitioned schedulers, exactly-once firing, time zones, missed runs and backfills.
Reading is half of it. See this used in a real interview: walk through Design a Distributed Job Scheduler →
Plenty of work must happen later, not now: send a reminder in 24 hours, retry a payment in 10 minutes, expire a reservation after 15 minutes, run a report every night at 2 a.m., publish a scheduled post. With one server, cron and an in-memory timer do it. Across a cluster, with millions of timers and machines that crash, you need to decide where timers live, who fires them, and how to avoid firing twice or never.
Kinds of scheduled work
- Recurring jobs: cron expressions ("every day at 02:00 UTC").
- One-off delayed jobs: "run this payload at time T", often created by the millions (reminders, expirations, retries).
- Timeouts that are usually cancelled before they fire (a reservation hold that ends when the purchase completes).
Storing timers
| Approach | How | Good for |
|---|---|---|
| Delayed queue | the broker hides a message until its time (SQS delay, RabbitMQ delayed exchange, Redis sorted set) | short delays, simple setups |
| Table indexed by time | rows with run_at, polled by WHERE run_at <= now() | durable, queryable, cancellable timers |
| Time buckets | timers grouped by minute (bucket = run_at rounded) and partitioned | millions of timers; scan one bucket per minute |
| Timing wheel | in-memory circular buffer of slots, each holding timers due then | very many short timers on one node |
| Workflow engine | durable timers inside workflows (Temporal and similar) | timers that are part of multi-step processes |
A Redis sorted set with the due time as score is a popular delayed queue: workers pop items with score at or below now. Persistence and failover need care. See Redis data structures.
Who fires the timers
- One leader scans for due timers and enqueues them for workers. Simple, but the leader must be elected safely, with leases and fencing, and becomes a throughput limit. See distributed locks and leases.
- Partitioned schedulers: timers are partitioned (by hash of job id or by time bucket), and each scheduler node owns some partitions, reassigned on failure. Scales horizontally.
- Claim with conditional updates: any worker can grab due rows with
UPDATE ... SET owner = me WHERE id = ? AND owner IS NULL(orSELECT ... FOR UPDATE SKIP LOCKED), so no row is taken twice.
The scheduler only dispatches: it puts due jobs on a normal work queue, and a worker pool runs them. Separating firing from execution keeps a slow job from delaying other timers. See background jobs.
Exactly once, in practice
Firing involves a crash window: a job may be dispatched and the scheduler crash before recording it. So:
- Dispatch is at least once; each firing has an id (job id plus scheduled time).
- Handlers are idempotent, keyed by that id. See distributed transactions and idempotency.
- Record state transitions (
scheduled,dispatched,running,succeeded,failed) so you can see and retry stuck jobs.
Missed runs and catch-up
If the scheduler is down for an hour, what happens to the jobs due during that hour? Decide per job: run all missed occurrences, run once to catch up, or skip. Nightly reports usually run once; billing runs must all happen. Also bound catch-up so a backlog does not overload downstream systems.
Time and time zones
- Store times in UTC; convert at the edges.
- Recurring jobs in a user's local time ("9 a.m. every day") must handle daylight saving changes: 2:30 a.m. may not exist or may happen twice.
- Machines' clocks drift; the scheduler should be tolerant of a few seconds of skew, and nothing correctness-critical should depend on exact timing. See unique ids, ordering and time.
Thundering herds
"Every day at midnight" means everything fires at once. Add jitter (spread over a window), especially for millions of user-level jobs like reminders, and rate-limit dispatch to protect downstream services. See Design a Notification System.
Observability
Track the lag between scheduled time and actual start (the key SLI for a scheduler), the number of overdue timers, failures and retries per job type, and job durations. Alert when lag grows. See SLIs, SLOs and error budgets.
In the interview
For Design a Job Scheduler: a durable job store indexed by next run time (time-bucketed for scale), partitioned scheduler nodes that claim due jobs with conditional updates and dispatch them to a queue, idempotent workers with retries and state tracking, missed-run policies, UTC and DST handling, jitter, and lag monitoring.
Checklist
- Timers in a durable store indexed or bucketed by due time.
- Partitioned or leader-based firing with safe claiming.
- Dispatch separated from execution.
- At-least-once firing with idempotent handlers.
- Explicit policy for missed runs; bounded catch-up.
- UTC storage, DST-aware recurrences, jitter for popular times.
- Schedule lag as the main metric.