Skip to content

Flight Routing Lab

Gatebound is a reinforcement-learning system I built in Python to get a passenger there by a deadline when flights can be delayed or cancelled. This page rebuilds its simulator and runs four booking policies (rules for which flight to book) over the same 64 worlds, simulated days with fixed delays.

Original built with Python · Gymnasium · Parquet

What I found

0 vs 40

worlds out of 64 on time with the deadline 5 min before the first nonstop, F3, lands: Nonstop first, which books that nonstop, makes 0; Deadline lookahead, which plans every leg ahead, makes 40.

−8

arrivals overall, 54 against 62: the price of those on-time worlds. The connection through Denver is two flights that can each be disrupted, not one.

58 of 64

worlds where Greedy next arrival, which takes the soonest landing, strands the passenger in Chicago, a dead end built into this network on purpose. An average score alone would hide why the rule fails; the run records how every trip ends.

Try it

A passenger has to get from SFO to JFK before a deadline, and any flight can be delayed or cancelled. Each panel runs one policy, a rule for which flight to book next, over the same 64 worlds: simulated days whose delays and cancellations are fixed by a seed, so every policy meets the same ones. Try the other deadline and see whether lookahead still beats the nonstop.

Deadline
Disruptions
Route

Mixed: each flight is cancelled 1 time in 10 and delayed or diverted 1 time in 10.

Deadline lookahead is on time in 40 of 64 worlds, Nonstop first in 0. Lookahead gets 8 fewer passengers there at all (54 against 62).

  • On time
  • Late
  • Never arrived

Each square is one world (seeds 42–105), in the same place in all four panels, so any difference between panels is the policy's alone. Click a square to replay that world below.

Nonstop first

Books the earliest nonstop to the destination.

0/64 on time

62 late · 2 never arrived

Deadline lookahead

Plans every leg ahead and books the flight most likely to make the deadline.

40/64 on time

14 late · 10 never arrived

Greedy next arrival

Takes whichever flight lands soonest, wherever it goes.

0/64 on time

0 late · 64 never arrived

Seeded random

Picks a bookable flight at random, the same pick every time in a given world.

5/64 on time

38 late · 21 never arrived

World 1 of 64

Seed 42 · Deadline lookahead · 2 decisions

Status
Arrived on time
Passenger
At JFK
Clock / deadline
07:50 / 07:55
Attempts
2 of 4 allowed
Reward (out of 1)
0.9510
FlownRing: where the passenger isSFO San Francisco · DEN Denver · ORD Chicago · JFK New York

Flight schedule: thin bars are the timetable, thick bars what actually flew. The dashed line is the deadline, the solid one the passenger's clock.

F1
F2
F3
F4
F5
F6
00:00Deadline 07:5516:00

What happened

  1. F2 SFO → DENLeft 01:30, landed 03:20.
  2. F5 DEN → JFKLeft 04:20, landed 07:50.

Rewind to take another flight in this same world.

On time
0.8000 / 0.80
Arrived
0.1000 / 0.10
Earliness
0.0510 / 0.10
Reward
0.9510

The reward is paid when the trip ends: 0.80 for landing by the 07:55 deadline, 0.10 for landing at all, and up to 0.10 more the earlier the landing before the horizon, 16:00, when the simulated day ends.

Under the hood: world settings and export

First world seed 0–2,147,483,584; deadline 1–1440 and no later than the horizon; buffer 0–1440; 1–6 attempts.

Contents

Getting there on time, not just getting there

A passenger has to reach JFK before a deadline. Landing somewhere soon is not the goal, and neither is landing eventually: the quickest next flight can strand them, and the fastest whole itinerary can expose them to a second cancellation. Gatebound scores a trip with a reward, a number paid once when the trip ends: any on-time arrival outscores any late one, and a late arrival still beats never arriving.

When lookahead wins, and what it costs

Four policies, each a rule for which flight to book next, run over the same 64 seeded worlds: simulated days whose delays and cancellations are fixed by a seed, so every policy meets the same ones. The network has two nonstops to JFK; the deadlines here are set against the first, F3, which lands at 08:00. With a 09:00 deadline, Deadline lookahead books that nonstop as Nonstop first does, and the two match world for world: 55 of 64 on time, 62 arrived. Move the deadline to 07:55, 5 min before the first nonstop, F3, lands, and Nonstop first is on time in 0 worlds while lookahead switches to the connection through Denver and is on time in 40. It pays for them: 54 arrivals instead of 62, because two flights can each be disrupted.

PolicyOn time, 09:00On time, 07:55Arrived, 09:00 / 07:55
Nonstop first55062 / 62
Deadline lookahead554062 / 54
Greedy next arrival000 / 0
Seeded random16543 / 43

Greedy next arrival never reaches JFK. It takes the earliest landing, F1 to Chicago, where this network has no onward flight: 58 worlds strand there and 6 are cancelled first. That is a counterexample to the rule, not evidence about real connections, and it is why the recorded run keeps every way a trip can end rather than only a mean reward, the average score.

The full system

This page is a small rebuild of part of Gatebound, the routing and verification system I built in Python. Its 2020–2024 pipeline run covered all 60 months of US flight-performance records: 33,173,483 raw rows, 2,651,910 kept within a 12-hub scope. In its historical simulator, six fixed 2025 requests at an eight-hour deadline, the deadline planner tied nonstop-first on five; on Newark to Salt Lake City it made 10 of 100 deadline arrivals against none, while eventual arrivals fell from 100 to 90. That is the trade this page shows, measured on real schedules. Those figures are results from the original evaluation run; the public repo is a cleaned release without those run artifacts.

What this page is not

Six invented flights with made-up delays on two mirrored networks. Nothing here is a forecast, a passenger-benefit estimate or a training result.

For engineers

The reward, how lookahead plans, the environment's rules, Gatebound's pipeline, what this rebuild leaves out, and how to reproduce the run.

The reward

The reward is terminal, paid once when the trip ends:

R = 0.80 × on time
  + 0.10 × arrived
  + 0.10 × arrived
    × (horizon − arrival) / horizon

The horizon, 16:00, is where the simulated day ends. The first nonstop, F3, landing on schedule at 08:00 earns 0.95; the same flight landing at 11:00, past the 09:00 deadline, earns 0.13125; a cancellation earns nothing.

How lookahead plans

Lookahead plans on the model, never on the sampled outcome. For each bookable flight it runs through the ten joint outcomes in that flight's pool and recurses to the end of the trip. Under the mixed profile the first nonstop's modelled chance of arriving by 09:00 is 80%, and by 07:55 0%; the Denver connection's are 64% and 64%. These are probabilities under this fixture, not forecasts. More lookahead is not a general win; it is a different trade, and the reward decides when the trade is worth taking.

Like Gatebound's planner, it maximizes only the chance of making the deadline and breaks a tie by the earliest scheduled arrival, then the first listed flight. So when no flight can make the deadline, every choice ties at zero and it books the earliest arrival, which can be a dead end: set an unreachable deadline under the hood and lookahead strands the passenger as greedy does.

The environment

A deadline-aware routing environment with masked actions, joint disruption samples and paired policy evaluation, each of which the paragraphs below spell out.

State and actions. The state is the passenger's airport, clock and attempts so far. The actions are the flights leaving that airport at least the connection buffer after the clock and before the horizon, in a fixed order, padded to eight slots with a binary mask, so a policy can never pick a flight that has already left.

Joint outcomes. Taking a flight reveals one whole outcome row: departure delay, arrival delay, cancellation and diversion together, never sampled field by field, so the simulator cannot invent a cancellation plus an unrelated delay. A cancellation keeps the passenger at the origin. An unresolved diversion ends the trip without moving them to an airport they never confirmed. After a landing, the next action set is built from the actual landing time, not the scheduled one.

Paired worlds. The outcome row is picked by hashing the world's seed with the flight, so a flight has the same outcome in a world whichever policy reaches it and whenever. Comparing policies on shared worlds takes sampling noise out of the comparison, and it makes a rewind in the replay a true counterfactual: the other branch of the same day.

Terminal reasons. Arrived, cancelled, diverted, no onward flight, out of attempts and out of time are kept apart. A stranded passenger and a cancelled one both score zero; the run still tells them apart.

Gatebound’s pipeline

Gatebound ingests monthly US flight-performance archives into partitioned Parquet, builds comparable-flight outcome pools, runs masked Gymnasium environments with a deadline planner and a masked REINFORCE learner, and checks every episode record against its configured schedules and outcome rows, so a forged but plausible record cannot earn reward.

What is left out

Six invented flights, ten joint outcome rows a flight and two mirrored networks. Chicago is a deliberate dead end, and the return route mirrors the outbound one rather than pretending to be a second dataset. The clear and stress profiles are controlled edits, not weather models, and times are minutes from midnight, not local clocks.

Left out next to Gatebound: prebooked itineraries, rebooking after a cancellation, fares, carriers, time zones, historical pool fitting, record authentication and the learned policies. Nothing here is a passenger-benefit estimate or a training result. The outcome hash is FNV-1a with an avalanche step, not Gatebound's cryptographic generator, so seeds do not carry across; and the outcomes ship with the page, so hiding them in the interface is not a security boundary.

Reproduce it

The engine, fixtures and reference cases are in the code for this page. 10 reference cases pin terminal paths (on time, late, cancelled, diverted, stranded, out of attempts, out of time, an invalid action) to exact clocks and rewards, and a test recomputes the 64-world run and compares it with the committed JSON. From the site's Next.js app:

npx vitest run src/lib/projects/flight-routing

Export, under the hood in the replay, writes the trip with its configuration, fixture version, outcome rows, reward terms and the 64-world comparison, so a run can be checked outside the page.