Code Verification Lab
Patchglass is an RL environment I built in Python for code-repair agents: an agent patches a repository in a container, and its reward comes from running every test on the buggy code and on the patched code. This page rebuilds that grading in your browser on two small tasks.
Original built with Python · Docker · SQLite · Pydantic · Typer
What I found
fail-to-pass tests (ones the bug fails and the fix must pass) repaired by the example-only patch. It recognizes the one demonstrated input, and the smoke suite, a quick two-test check, still passes it.
pass-to-pass test broken by the ordering patch while it repairs all 3 fail-to-pass tests. A patch passes only if every test passes (a conjunction, not an average), so repairs cannot cancel a regression.
failing tests fixed by a patch written by an AI model in Patchglass's own run, though it applied cleanly and kept 52/52 existing tests passing. These are results from the original evaluation run; the public repo is a cleaned release without those run artifacts.
Try it
On the two-test smoke suite, 3 of 4 patches pass. On the full suite, 1 does.
The task: Merge overlapping or touching closed intervals, including point intervals. Return ascending disjoint ranges without mutating the input. The bug: The baseline treats equality at an endpoint as a gap.
Each row is a candidate patch; each column a test, run twice, on the buggy function and on the patched one. A fail-to-pass test is one the bug fails and the fix must pass; a pass-to-pass test passes before the fix and must still pass after. The smoke suite is a quick check that runs one test of each kind; the full suite runs all six. Compare the two verdict columns: the smoke suite passes patches the full suite fails. Select a patch to inspect its evidence below, and switch the suite to see which tests each one runs.
- Repaired fail → pass
- Still passes pass → pass
- Still broken fail → fail
- Regressed pass → fail
| Patch | Smoke verdict | Full verdict | Fail-to-passTouching pair | Fail-to-passUnsorted touching chainnot in smoke suite | Fail-to-passPoint on negative endpointnot in smoke suite | Pass-to-passExisting overlap | Pass-to-passSeparated ranges stay ascendingnot in smoke suite | Pass-to-passEmpty inputnot in smoke suite |
|---|---|---|---|---|---|---|---|---|
| Passes | Passes | |||||||
| Passes | Fails | |||||||
| Passes | Fails | |||||||
| Fails | Fails |
Hatched columns are the tests the smoke suite skips: they show what a full run would find, and the smoke verdict ignores them.
Merge overlapping or touching closed intervals, including point intervals. Return ascending disjoint ranges without mutating the input.
Baseline faultThe baseline treats equality at an endpoint as a gap.
Adds a special case for the first demonstration input.
Browser evaluation · deterministic, synthetic fixtures.
Curated functions only; no arbitrary code execution.
The two-test smoke suite is satisfied. This is not full verification.
Touching pair
FailPass- Input
[[1,3],[3,5]]- Expected
[[1,5]]- Buggy actual
[[1,3],[3,5]]- After actual
[[1,5]]
Implementation & replacement diff
Each function as written; the page's tests run this text against the function it executes, so they cannot differ. The diff compares the whole buggy function with the selected one, not a repository patch, and marks the characters that change.
Contents
What a passing test establishes
A passing example is weak evidence for a repair. A patch can recognize that exact input and return its expected answer while staying wrong everywhere else; another can fix the reported bug and break an ordering that callers already depend on. A green badge hides both.
In an RL environment that matters twice over, because the grade is the reward an agent learns from. A reward that checks only the bug's own example pays an agent for memorizing that example, so the environment has to grade every test, on both versions of the code.
So the verifier grades a patch by transitions: how each test's result changes. Every test runs twice, on the buggy function and on the patched one, against the same independently written expected value. A failure that now passes is a repair, a pass that still passes is preserved behavior, a failure that still fails is still broken, and a pass that now fails is a regression. The patch is resolved only when the change is not empty and every selected test passes after it.
What the four patches show
Each task has 3 fail-to-pass tests, which the bug fails and the fix must pass, and 3 pass-to-pass tests, which must pass both before and after. The smoke suite, a quick check, keeps one of each. On it, 3 of the 4 patches pass: every patch that changes the function at all. On the full suite, 1 does.
The example-only patch adds a special case for the demonstrated input. It repairs 1 of 3 fail-to-pass tests and leaves 2 still broken, which the smoke suite never runs. The ordering patch repairs all 3 and reverses the output order, breaking 1 pass-to-pass test. Repair evidence cannot cancel a regression: a patch passes only if every test passes, a conjunction, not an average.
Where this comes from
This page is a reduced rebuild of Patchglass, the container-based repair and test-synthesis harness I built in Python. It packages a repository at a commit, gives a solver a restricted workspace, and grades the patch in a fresh container. In its saved runs, a patch written by an AI model applied cleanly and broke none of 52 existing tests, yet fixed 0 of the 4 failing ones: a patch that applies and breaks nothing can still fix nothing. These come from the same original run as the findings above.
What this page is not
Two small tasks with six tests each and four curated patches, not model output. There is no container or repository here, so nothing on this page measures an agent's ability to find or fix a bug.
For engineers
How a patch is graded, how authored tests are judged, what Patchglass runs that this page does not, and how to rerun it.
How a patch is graded
Baseline polarity first. Before any patch is graded, the engine checks that every fail-to-pass test fails on the buggy function and every pass-to-pass test passes. A fixture that breaks that rule stops the run rather than producing a result.
Curated functions, not uploaded code. The patches are ordinary typed functions called directly; nothing the visitor types is executed as code. Visitors supply bounded input and expected-output data instead: up to 64 items, strings up to 128 characters and 16 authored tests, with malformed JSON, reversed intervals and mismatched shapes rejected.
Evidence and replay. A report records the task, patch, mode, suite, authored tests, every observation and the transition counts. Import recomputes the result from the configuration rather than trusting the saved verdict, so editing a report's verdict changes nothing.
Test synthesis mode
The same holds in test-synthesis mode, where the visitor writes the tests. A synthesized test is useful only if it fails before the reference repair and passes after it, and every authored test must pass on the fixed function. An already-passing assertion reproduces nothing; a wrong expectation stays wrong after the repair and rejects the suite.
Inside Patchglass
A task bundle pins a repository commit, an image digest, test commands, a parser and graded test buckets. The solver gets a stripped workspace without the graded tests; grading happens in a fresh container with no network, dropped capabilities, a seccomp profile and resource limits, and hidden tests are staged only after the patch is applied. Results land in SQLite with the patch, image digest and task hash.
| Saved run | Fail-to-pass | Pass-to-pass | Verdict |
|---|---|---|---|
| Python repository, reference patch | 4/4 | 52/52 | Resolved |
| Same task, successful model patch | 4/4 | 52/52 | Resolved |
| Same task, unsuccessful model patch | 0/4 | 52/52 | Not resolved |
| JavaScript repository, reference patch | 3/3 | 288/288 | Resolved |
| Go repository, reference patch | 2/2 | none declared | Resolved |
Both model patches, written by AI models through Patchglass's model solvers, applied cleanly. One repaired all four failures; the other repaired none and preserved every existing test, which is the point of this page at repository scale. The table is from the original run described above.
What the rebuild leaves out
Passing the tasks does not prove a function correct; authored tests can probe new boundaries, but property tests or a proof would be stronger. Every test ships with the page, so this is an inspection tool, not a hidden-test leaderboard, and the string task uses JavaScript lowercase, not full Unicode case folding. There is no patch application, network policy or test-output parser here.
Reproduce it
The engine, patches and fixtures are in the code for this page. The tests cover baseline polarity, every patch, invalid inputs, contradictory synthesis expectations and replay. From the site's Next.js app:
npx vitest run src/lib/projects/code-verification