Autograder Trust-Boundary Sandbox

COMP 435 final project · the same submission is graded by a naive autograder and a hardened one, side by side. Everything runs in your browser (Pyodide / WebAssembly); nothing is sent to a server.

The problem both graders test

Implement solve(a, b) so it returns a + b. Four hidden tests: (2,3)→5   (-1,1)→0   (10,20)→30   (0,0)→0

The naive grader (left) execs your code in a namespace that also holds the hidden answer key and the live score object, and checks results with got == expected. The hardened grader (right) runs your code in an isolated namespace with no answers and no score in scope, and compares by canonical JSON plus a strict type check.

Load a submission

booting Python runtime…

Naive grader UNSAFE

—

Hardened grader FIXED

—

What each attack shows (your presentation notes)

Attack 1 — assertion subversion via __eq__

The submission returns an object whose __eq__ always returns True. Because the naive grader checks got == expected, every test "passes." Root cause: the grader trusts an operation the untrusted object controls. Fix: compare canonical representations and check the type, so the object's own __eq__ is never called.

Attack 2 — answer-key disclosure

The naive grader execs the submission in a scope that still contains ANSWER_KEY, so the code reads the hidden answers (watch the leaked output) and returns exact correct values. This defeats even a type/canonical check, because the returned values are genuinely correct. Only scope isolation stops it: the hardened grader never puts the key in the submission's namespace, so the name is undefined.

Attack 3 — grade tampering

The naive grader hands the submission the live results object, so the code sets RESULTS['score'] to full marks and returns nothing. Every test actually fails, yet the score is maxed. Fix: the score is a local variable the submission can never reach. In a real system (e.g. Gradescope's results.json) this is the file the grader writes; if student code shares that filesystem and user, it can overwrite it.

The one lesson

An autograder runs attacker-controlled code inside the same trust zone as the things it protects: the score, the answer key, and the grading harness. Defense means putting a real boundary between them — in scope, in comparison logic, and (in production) at the OS level.

Limitations of this sandbox

This demo fixes the logic of the trust boundary. It runs in one Pyodide interpreter, so it cannot demonstrate OS-level isolation. A production autograder must also run each submission as a separate low-privilege user in a container, mount the answer key read-only, write results on a channel the submission cannot touch, drop network access, and enforce CPU, memory, and wall-clock limits. Those defeat the reverse-shell and container-escape attacks documented in the real-world writeups below.

References