Skip to content

Code Verification Lab

Patchglass is an RL environment I built in Python for code-repair agents: an agent patches a repository in a container, and its reward comes from running every test on the buggy code and on the patched code. This page rebuilds that grading in your browser on two small tasks.

Original built with Python · Docker · SQLite · Pydantic · Typer

What I found

1/3

fail-to-pass tests (ones the bug fails and the fix must pass) repaired by the example-only patch. It recognizes the one demonstrated input, and the smoke suite, a quick two-test check, still passes it.

1

pass-to-pass test broken by the ordering patch while it repairs all 3 fail-to-pass tests. A patch passes only if every test passes (a conjunction, not an average), so repairs cannot cancel a regression.

0/4

failing tests fixed by a patch written by an AI model in Patchglass's own run, though it applied cleanly and kept 52/52 existing tests passing. These are results from the original evaluation run; the public repo is a cleaned release without those run artifacts.

Try it

Task
Suite

On the two-test smoke suite, 3 of 4 patches pass. On the full suite, 1 does.

The task: Merge overlapping or touching closed intervals, including point intervals. Return ascending disjoint ranges without mutating the input. The bug: The baseline treats equality at an endpoint as a gap.

Each row is a candidate patch; each column a test, run twice, on the buggy function and on the patched one. A fail-to-pass test is one the bug fails and the fix must pass; a pass-to-pass test passes before the fix and must still pass after. The smoke suite is a quick check that runs one test of each kind; the full suite runs all six. Compare the two verdict columns: the smoke suite passes patches the full suite fails. Select a patch to inspect its evidence below, and switch the suite to see which tests each one runs.

  • Repaired fail → pass
  • Still passes pass → pass
  • Still broken fail → fail
  • Regressed pass → fail
PatchSmoke verdictFull verdictFail-to-passTouching pairFail-to-passUnsorted touching chainnot in smoke suiteFail-to-passPoint on negative endpointnot in smoke suitePass-to-passExisting overlapPass-to-passSeparated ranges stay ascendingnot in smoke suitePass-to-passEmpty inputnot in smoke suite
PassesPasses
PassesFails
PassesFails
FailsFails

Hatched columns are the tests the smoke suite skips: they show what a full run would find, and the smoke verdict ignores them.

Merge overlapping or touching closed intervals, including point intervals. Return ascending disjoint ranges without mutating the input.

Baseline faultThe baseline treats equality at an endpoint as a gap.

Adds a special case for the first demonstration input.

Browser evaluation · deterministic, synthetic fixtures.
Curated functions only; no arbitrary code execution.

Suite satisfied

The two-test smoke suite is satisfied. This is not full verification.

Fail → Pass1Repaired
Pass → Pass1Still passes
Fail → Fail0Still broken
Pass → Fail0Regressed
AssertionBuggyAfter

Touching pair

FailPass
Input
1 … 3
3 … 5
Buggy
1 … 3
3 … 5
After
1 … 5
Expected
1 … 5
15
Closed endpoints share one scale. Each range occupies its own lane.
Input
[[1,3],[3,5]]
Expected
[[1,5]]
Buggy actual
[[1,3],[3,5]]
After actual
[[1,5]]
Implementation & replacement diff
Example-only repair

Each function as written; the page's tests run this text against the function it executes, so they cannot differ. The diff compares the whole buggy function with the selected one, not a repository patch, and marks the characters that change.

Contents

What a passing test establishes

A passing example is weak evidence for a repair. A patch can recognize that exact input and return its expected answer while staying wrong everywhere else; another can fix the reported bug and break an ordering that callers already depend on. A green badge hides both.

In an RL environment that matters twice over, because the grade is the reward an agent learns from. A reward that checks only the bug's own example pays an agent for memorizing that example, so the environment has to grade every test, on both versions of the code.

So the verifier grades a patch by transitions: how each test's result changes. Every test runs twice, on the buggy function and on the patched one, against the same independently written expected value. A failure that now passes is a repair, a pass that still passes is preserved behavior, a failure that still fails is still broken, and a pass that now fails is a regression. The patch is resolved only when the change is not empty and every selected test passes after it.

What the four patches show

Each task has 3 fail-to-pass tests, which the bug fails and the fix must pass, and 3 pass-to-pass tests, which must pass both before and after. The smoke suite, a quick check, keeps one of each. On it, 3 of the 4 patches pass: every patch that changes the function at all. On the full suite, 1 does.

The example-only patch adds a special case for the demonstrated input. It repairs 1 of 3 fail-to-pass tests and leaves 2 still broken, which the smoke suite never runs. The ordering patch repairs all 3 and reverses the output order, breaking 1 pass-to-pass test. Repair evidence cannot cancel a regression: a patch passes only if every test passes, a conjunction, not an average.

Where this comes from

This page is a reduced rebuild of Patchglass, the container-based repair and test-synthesis harness I built in Python. It packages a repository at a commit, gives a solver a restricted workspace, and grades the patch in a fresh container. In its saved runs, a patch written by an AI model applied cleanly and broke none of 52 existing tests, yet fixed 0 of the 4 failing ones: a patch that applies and breaks nothing can still fix nothing. These come from the same original run as the findings above.

What this page is not

Two small tasks with six tests each and four curated patches, not model output. There is no container or repository here, so nothing on this page measures an agent's ability to find or fix a bug.

For engineers

How a patch is graded, how authored tests are judged, what Patchglass runs that this page does not, and how to rerun it.

How a patch is graded

Baseline polarity first. Before any patch is graded, the engine checks that every fail-to-pass test fails on the buggy function and every pass-to-pass test passes. A fixture that breaks that rule stops the run rather than producing a result.

Curated functions, not uploaded code. The patches are ordinary typed functions called directly; nothing the visitor types is executed as code. Visitors supply bounded input and expected-output data instead: up to 64 items, strings up to 128 characters and 16 authored tests, with malformed JSON, reversed intervals and mismatched shapes rejected.

Evidence and replay. A report records the task, patch, mode, suite, authored tests, every observation and the transition counts. Import recomputes the result from the configuration rather than trusting the saved verdict, so editing a report's verdict changes nothing.

Test synthesis mode

The same holds in test-synthesis mode, where the visitor writes the tests. A synthesized test is useful only if it fails before the reference repair and passes after it, and every authored test must pass on the fixed function. An already-passing assertion reproduces nothing; a wrong expectation stays wrong after the repair and rejects the suite.

Inside Patchglass

A task bundle pins a repository commit, an image digest, test commands, a parser and graded test buckets. The solver gets a stripped workspace without the graded tests; grading happens in a fresh container with no network, dropped capabilities, a seccomp profile and resource limits, and hidden tests are staged only after the patch is applied. Results land in SQLite with the patch, image digest and task hash.

Saved runFail-to-passPass-to-passVerdict
Python repository, reference patch4/452/52Resolved
Same task, successful model patch4/452/52Resolved
Same task, unsuccessful model patch0/452/52Not resolved
JavaScript repository, reference patch3/3288/288Resolved
Go repository, reference patch2/2none declaredResolved

Both model patches, written by AI models through Patchglass's model solvers, applied cleanly. One repaired all four failures; the other repaired none and preserved every existing test, which is the point of this page at repository scale. The table is from the original run described above.

What the rebuild leaves out

Passing the tasks does not prove a function correct; authored tests can probe new boundaries, but property tests or a proof would be stronger. Every test ships with the page, so this is an inspection tool, not a hidden-test leaderboard, and the string task uses JavaScript lowercase, not full Unicode case folding. There is no patch application, network policy or test-output parser here.

Reproduce it

The engine, patches and fixtures are in the code for this page. The tests cover baseline polarity, every patch, invalid inputs, contradictory synthesis expectations and replay. From the site's Next.js app:

npx vitest run src/lib/projects/code-verification