Skip to content
All writing
Research3 min read

Notes on running other people's code

What building HuntCode's sandbox taught me: the escape is rarely the first thing that breaks, and every performance shortcut is a reuse of something.

HuntCode executes code written by people whose stated goal, on that exact page, is to find a way out. That is a comfortable threat model to design against, because it removes the temptation to reason about intent. Everything is hostile; the only question is what holds.

The escape is not the first thing that breaks

Before anyone gets near a container escape, they will have found the cheaper attacks: allocate until the host swaps, fork until the pid table is full, write until the disk is gone, print until the log shipper falls over, or simply ask for ten thousand executions at once. None of those require a vulnerability. They require an API.

  • Memory and CPU ceilings, enforced by the runtime rather than by the program.
  • A pid limit, because fork bombs are three characters long.
  • A wall-clock timeout, because a program that never ends looks exactly like one that is slow.
  • An output byte ceiling — unbounded stdout is a denial of service against whatever reads it.
  • Per-account concurrency limits before global ones, so one account cannot spend everyone's capacity.

Egress is the control that pays for itself

An execution container has no outbound network. That single decision removes exfiltration of anything the container learns, removes the platform's usefulness as an anonymising proxy, removes callbacks to attacker infrastructure, and removes the entire class of attacks that begin with pulling a second stage.

Every speed-up is a reuse

Cold container start is slow enough to be felt on every run, and the obvious fix is to keep containers around and reuse them. That is also the one thing the design cannot allow: reuse between players is exactly the boundary the sandbox exists to enforce.

The distinction that made it work is between pre-warming and recycling. A pool of containers is started ahead of demand and kept idle, but each one is used for exactly one execution and then destroyed. The startup cost leaves the critical path; no state crosses between players.

text
pre-warm   start → idle → assign → run → destroy      no state crosses
recycle    start → run → reset? → run → reset? → run   one missed reset away

The output is untrusted data

Program output is streamed to a terminal view in the browser. It is tempting to treat it as text and stop thinking, but it is attacker-chosen bytes arriving at a rendering surface — the same category as any other untrusted input. It gets rendered as text, never as markup, and it gets a length limit on the way.

What I would tell myself at the start

  1. 01Write the attacker's side of the feature first. It is the fastest way to find out which assumptions were decorative.
  2. 02Limits enforced by the thing being limited are not limits.
  3. 03Prefer removing a capability to detecting its misuse. Detection is a maintenance commitment; removal is not.
  4. 04Assume you will be wrong about one of these, and make the container cheap enough to throw away that being wrong costs one execution.