Browser Agent Completion Gate
Parallax is a browser agent I built in Python: it plans a task, then clicks and types through websites to carry it out. This page rebuilds its executor on a synthetic booking app and adds a completion gate, a check that calls a task done only when the saved record proves it.
Original built with Python · Playwright · SQLite
What I found
attempts pass Parallax's completion check, the rule my earlier agent uses, which accepts any valid form after a typing step. One saved the booking that was asked for.
attempts end on the success notice "Reservation saved". Two of those notices are wrong.
conditions the completion gate checks on the page, so it never relies on the notice alone: record, title, room, save, dialog, notice.
Try it
6 attempts to book North lab for “Spectral scan”, and only one saved the booking that was asked for. Only the completion gate gets every attempt right.
Each row is one attempt. The saved-as-asked column says whether a record with the requested title and room was committed; the three columns after it are ways of deciding the booking worked. The completion gate is this rebuild's check: it calls a booking done only when a committed record exists and matches the request. Red marks a check that got it wrong. Click an attempt to run it in the booking app below. Try “Notice shown, nothing saved”: the app says “Reservation saved” and its list stays empty.
| Attempt | Saved as askeda committed record with the requested title and room | Success notice"Reservation saved" shown | Parallax's checkmy earlier agent | Completion gatethis rebuild's check |
|---|---|---|---|---|
| yes | done | done | done | |
| none | donewrong | donewrong | not done | |
| South lab, wrong room | donewrong | donewrong | not done | |
| none | not done | donewrong | not done | |
| none | not done | donewrong | not done | |
| none | not done | not done | not done |
Loaded in the app Saves normally
Lab reservations
| Reservation | Room | State |
|---|---|---|
| No reservations | ||
Completion gate: done only when all 6 hold
- ○Committed record exists
- ○Title matches the requested task
- ○Room matches the requested task
- ○Save reached its completed state
- Dialog is closed
- ○Success notice is present
Settings, the action plan and the step-by-step trace
Action plan
- 1Open reservation dialog
button[Reserve slot] - 2Fill task title
textbox[Reservation title]"Spectral scan" - 3Choose a room
combobox[Room]"north" - 4Submit reservation form
button[Save reservation] - 5Verify committed reservation
all completion conditions
Observed trace
No actions observed.
A fixed plan on the app's real controls, with records kept in memory. Nothing leaves the page: no model call, no network request, no screenshot. The trace samples the page's state before and after each action.
Contents
When is a reservation made?
An agent that operates a website has to decide when it is finished. The cheap answer is to trust the site: a success message, a closed dialog, a form that validates. Each of those can be true while the thing the user asked for does not exist. A notice can come from a save that wrote nothing, a correct-looking record can be in the wrong room, and a form validates the moment its fields are filled, before anything is submitted.
So the question here is narrow and testable: what on the page justifies saying the reservation exists? The scheduling app is new and synthetic, built for this page. The executor, the program that does the clicking, operates its real controls: it clicks the button, types the title, picks the room, submits, and then waits on conditions rather than a clock. The completion gate is the check that decides whether it is done.
What each kind of evidence says
All 6 attempts ask for the same thing: Spectral scan, in North lab. The success notice is right when nothing goes wrong and when the save is refused. It is wrong two times: a save that reports success and commits no record, and a save to the wrong room, whose notice is identical to the good one. Three notices, one reservation.
Parallax's own check, the rule this rebuild started from, is weaker. After any step it treats as interactive, such as typing, filling or submitting, it accepts a status or alert element on the page, or a form with no invalid field. The form is valid as soon as the title is typed, so every attempt that gets that far passes: 5 of 6, including the save the server refused and the run that gave up before its save landed. Only the renamed button, where the form never opened, fails it. The same detector counts an error alert as a signal.
The completion gate agrees with the committed record, the reservation the app actually stored, in every row because it reads that record. That is the design, not a discovery: a gate is only as good as its conditions, and these are written for this task. The record exists, its title and room match the request, the save completed, the dialog closed, and the notice is present. The notice is one of the six conditions, never the only one.
Where this comes from
This page rebuilds one narrow piece of Parallax, a browser agent I built in Python. It asks a model for a plan, carries the plan out in a real browser through Playwright, and records what it saw after every step. Reading its source for this page turned up the gap the table shows: its completion check accepts signs that something happened, not proof that the task did. The table runs that rule, ported to TypeScript, over the same attempts.
What this page is not
One task on one synthetic app with a fixed plan: success here says nothing about arbitrary sites or about planning. Parallax itself is not run; its completion rule is ported and applied to this app's states.
For engineers
How the executor and the gate work, why two runs fail, the rest of Parallax's source, the limits in full, and how to rerun it.
How the executor and the gate work
A fixed plan on real controls. The executor runs 5 actions: open reservation dialog, fill task title, choose a room, submit reservation form, verify committed reservation. Each goes through a small adapter scoped to the app: it finds the button by its accessible name, sets the title and room through the native inputs and their events, and clicks the real submit button. Before and after every action it samples the dialog, the inputs, the notice and the table, and records the change in role and name pairs as a Jaccard distance, the role-set distance in the trace. A large distance means the structure changed, not that anything succeeded.
Waiting on conditions. After submitting, the executor polls the page until all 6 gate conditions hold or the deadline passes, and stops early if the save is rejected or a success notice appears without a record. An action budget caps the run, and a run that spends it is reported as exhausted, never as partly complete. Stop, reset and leaving the page cancel pending work, and a cancelled run cannot write into a newer one.
Replay recomputes. Export writes the configuration, the app's events and the observed trace under a schema version. Import rebuilds the states by replaying those events through the same transition function and recomputes completion instead of trusting a saved status. An unknown version, an illegal event order or more than 64 events is refused. Replay is a reconstruction, not a recording: it does not recapture the page or its timing.
Why two executor runs fail
The two executor failures differ in kind. With a fixed 200 ms wait on a 500 ms save, the executor checks before the save lands, calls the run failed and cancels the pending save; the condition wait polls every 20 ms up to its 1500 ms deadline and sees it commit. With exact-name selectors the renamed button is never found and the run stops at its first step; the alias policy, which knows one alternate name, recovers. That is one declared alias, not a planner finding a new strategy.
Parallax, in detail
Parallax's interpreter asks a model provider for a structured plan and validates the plan's structure. Its navigator executes the plan through Playwright, resolving controls by role, test id, selector, text or XPath, with bounded retries, slow retyping when a value does not stick, and an optional vision fallback. Its observer records the URL, roles, dialogs, notices, form validity and loaders after every step, with screenshots at several viewport sizes, and its archivist writes JSONL, SQLite and readable reports beside a Playwright trace.
Parallax decides an interactive task is complete when, after a step whose description mentions typing, filling, submitting or saving, the page has a status or alert element or a form with no invalid field. Those are signs that something happened, not that the task did. The rebuild keeps the separation Parallax's design was reaching for, between the action attempted, the transition observed and the task validated, and makes the last of those read the outcome.
Two more things the source shows. When the local planner's reply cannot be parsed or comes back empty, its fallback is a one-step plan that opens the start page, which can look like progress without planning anything. And screenshot redaction is on by default, but its helper returns the unredacted image if processing fails, so it is not a privacy guarantee.
Its limits in detail
One task, one curated plan and one declared alias: success here says nothing about arbitrary sites or autonomous planning. The executor and the app share a page, so the trace is inspectable but not attested, and this is no isolation boundary against a hostile site. A committed row lives in memory and is gone on reload. The gate checks the conditions written for this task; a wrong or missing condition would give a wrong verdict just as confidently. There is no model, network request or screenshot here, and Parallax itself is not run: its completion rule is ported and applied to this app's states.
Reproduce it
Pick an attempt in the table to run it, or step through one action at a time and compare each action's before and after under the hood. There, set the wait policy to a fixed delay and the save latency below 200 ms, and the early check succeeds. Export a run, reset, import it, and drag the replay to its last frame. The engine, the attempts, the ported Parallax rule and the adapter are in the code for this page; the tests run every executor attempt against the transition model and check each verdict. From the site's Next.js app:
npx vitest run src/lib/projects/workflow-observatory src/components/projects/workflow-observatory