Running untrusted code safely
How online judges, notebooks, CI systems and AI agents execute code they did not write: isolation with containers, gVisor and microVMs, resource limits, network and filesystem restrictions, judging pipelines, warm pools, fairness and abuse prevention.
Reading is half of it. See this used in a real interview: walk through Design LeetCode →
Some products exist to run code written by strangers: coding interview platforms and online judges, cloud notebooks, CI services, serverless platforms, and now AI agents that write and execute code. That code may be buggy (infinite loops, memory leaks), wasteful (crypto mining) or hostile (escaping to the host, attacking the network). "Design LeetCode" lives or dies on how you run submissions safely, fairly and quickly.
Threats
- Resource exhaustion: infinite loops, fork bombs, huge memory or disk use.
- Escape: exploiting the kernel or runtime to reach the host and other users' code.
- Network abuse: attacking other systems, sending spam, mining, exfiltrating test data.
- Information leaks: reading hidden test cases, other submissions or secrets in the environment.
- Noisy neighbours: one job slowing everyone on the same host.
Isolation options
| Technology | Isolation | Startup | Notes |
|---|---|---|---|
| Language sandbox (restricted interpreter) | weak | instant | easy to escape; not enough alone |
| Containers (namespaces, cgroups, seccomp) | moderate; shared kernel | under a second | fast and common; kernel bugs are the risk |
| gVisor-style user-space kernels | stronger; system calls intercepted | fast | some performance overhead and compatibility limits |
| MicroVMs (Firecracker and similar) | strong; separate kernel | about 100 ms or less | used by serverless platforms; good balance |
| Full VMs | strongest | seconds to minutes | heavy; for long-lived environments |
For untrusted code from the internet, prefer microVMs or a gVisor-style layer over plain containers, and always add defence in depth.
Limits and restrictions
Whatever the isolation, apply:
- CPU time and wall-clock limits (kill after the limit).
- Memory limits, with out-of-memory kills reported as such.
- Process count limits to stop fork bombs.
- Disk quotas and a read-only root filesystem, with a small writable temp directory.
- No network by default; if needed, an egress allowlist through a proxy.
- Seccomp filters to block dangerous system calls; drop all capabilities; run as an unprivileged user.
- No secrets in the sandbox environment; test data mounted read-only and only the inputs needed.
A judging pipeline
- The API receives a submission, stores code and metadata, and enqueues a job. See background jobs.
- A worker takes the job and starts (or takes from a warm pool) a sandbox for the language.
- Compile if needed (with its own time limit), then run against each test case with per-test limits.
- Compare output (exact, with tolerance for floats, or with a custom checker).
- Record the verdict (accepted, wrong answer, time limit, memory limit, runtime error, compile error) and per-test stats.
- Destroy the sandbox; never reuse one across users.
- Notify the client by polling or a WebSocket.
See Design LeetCode.
Fairness and consistent timing
Execution time is the verdict, so measurements must be consistent:
- Measure CPU time, not wall time, where possible.
- Pin sandboxes to dedicated cores; avoid overcommitting CPU on judging hosts.
- Run on uniform hardware, or calibrate limits per machine type.
- Rerun borderline results to reduce noise.
Scaling and latency
- Warm pools of pre-started sandboxes per language remove startup time; snapshots of initialised microVMs make this cheap.
- Separate queues for interactive "run" requests and full "submit" judging; contests get dedicated capacity.
- Autoscale workers on queue depth and age. Contest starts produce sudden spikes; pre-scale for scheduled events. See autoscaling.
- Per-user rate limits and concurrent job limits. See multi-tenancy.
AI agents and code interpreters
Agents that write and execute code need the same isolation, plus longer-lived sessions with files and package installs, and careful control over network access and any credentials the agent may use. Treat the model's output as untrusted user code. See LLM systems.
In the interview
For Design LeetCode: a queue of submissions, workers running each in a fresh microVM or hardened container with CPU, memory, process, disk and network limits, warm pools per language, CPU-time measurement on pinned cores, verdicts stored and pushed to the client, and pre-scaling for contests.
Checklist
- Strong isolation (microVM or user-space kernel) plus hardened containers.
- CPU, wall time, memory, process and disk limits; read-only filesystem.
- No network by default; no secrets; read-only test data.
- Fresh sandbox per submission, destroyed after use.
- Consistent timing on pinned, uniform hardware.
- Warm pools, queue-based autoscaling and per-user limits.