Skip to content

Offline Policy Evaluation Lab

Counterledger is an offline policy evaluator I built in Python: it estimates how a new decision policy (the target) would have done from records of past decisions (the logger), without trying it for real. This page rebuilds its importance-sampling core and puts an apparent gain through four checks.

Original built with Python · NumPy · pandas · scikit-learn · SciPy

What I found

65%

of the new policy’s (the target’s) apparent gain over the past decisions (the logger) also goes to a constant control that acts at the same rate but ignores the state. Most of the gain is how often to act, not when; the part that depends on the state is 0.10, paired 95% interval 0.03 to 0.17. Default settings.

−0.21 to 0.05

the paired 95% interval for the new policy minus the past decisions once the reward stops paying for gain: it crosses zero, and the improvement is gone.

68 of 160

logged trajectories, each one case’s run of 4 decisions, that the estimate effectively uses at the last step (its effective sample size): 42% of the data. Default settings.

Try it

Offline evaluation scores a new policy, the target, a rule for when to act, on records of decisions someone else made: the logging policy, or logger. Nothing is tried for real; each logged decision is reweighted by how likely the new policy was to make it. Here each made-up trajectory passes through 4 decisions; at each, a load level (low, moderate, high) is observed and one action is taken. Try “Gain dropped” or “None at low load” and watch the verdict change.

Reward
How often the logger chose Intensify
Estimator: importance sampling (IS), how the rewards are reweighted

Rare: the logger chose Intensify 2% of the time at low load, 10% at moderate load and 18% at high load. Changing this changes the past decisions themselves, so the logger’s own return moves too.

On its face the target beats the logger by 0.27 (paired 95% interval for target − logger: 0.12 to 0.40), but a control that never reads the state (the load) gets 65% of that gain: most of it is a shift in how often to act, not when.

The part that depends on reading the state, target − control, is 0.10 (paired 95% interval 0.03 to 0.17). The 65% is a ratio of the two point estimates, taken before rounding.

Average discounted return per logged trajectory, higher is better, over 160 trajectories of 4 decisions each, with 95% intervals from 256 paired bootstrap resamples. The dashed line marks the logger. The last row is the paired difference the verdict reads: each resample scores both policies on the same trajectories, so its interval is tighter than the two rows above it suggest.

State-responsive targetthe new policy: acts more as load rises
−0.10−0.39 to 0.13
Constant controlsame rate as the target, ignores the load
−0.19−0.50 to 0.06
Logging policythe past decisions: what actually happened
−0.37−0.59 to −0.17
Target − loggerthe paired difference the verdict reads
0.270.12 to 0.40

In effect, the estimate leans on 68 of the 160 logged trajectories (its effective sample size).

Under the hood: support, effective sample size, estimators, one trajectory's ledger, settings and export

The evaluation underneath

Seed 2026 · 160 trajectories · 4 decisions each

Final-step effective sample size (ESS)
68 / 160
Target mass on thin support
6.3%
Weights above the cap
1.4%
Largest final-step share
5.8%
Who the logger supports how often the logger and the target choose each action, by load
LoadHoldAdjustIntensify
Low load80.0% → 68.8%187 logged18.0% → 13.7%39 logged2.0% → 17.6%3 logged
Moderate load60.0% → 48.8%112 logged30.0% → 21.9%58 logged10.0% → 29.3%26 logged
High load40.0% → 28.7%85 logged42.0% → 30.2%89 logged18.0% → 41.1%41 logged

LoggerTargetLogger under 5%: thin supportLogger never acts there: no support

Trajectories the weights really use effective sample size (ESS) by step, of 160
  1. Step 198 / 116 capped
  2. Step 280 / 95 capped
  3. Step 381 / 86 capped
  4. Step 468 / 77 capped

Raw weightsCapped weightsMore capped ESS is less concentration, not less bias.

The target under each estimator paired against the logger on the same resamples

EstimatorReturn95% intervalVersus loggerPaired interval
Raw IS−0.07−0.32 to 0.170.300.14 to 0.45
Capped IS−0.09−0.34 to 0.120.280.14 to 0.39
Capped self-normalized IS−0.10−0.39 to 0.130.270.12 to 0.40

One trajectory's ledger its weight is the product of the target-to-logger ratios so far

StepLoadLogged actionRewardWeightCapped
1Low loadHold0.180.860.86
2Low loadHold0.320.740.74
3Low loadHold0.180.630.63
4Low loadHold−0.040.550.55

Settings

Target policy
0.55
0.20
0.25
Reward and discount
1.00
1.50
0.95
Logged cohort and estimator

Seed 0–4,294,967,295; trajectories 8–320.

Contents

The questions an estimate has to survive

Offline evaluation scores a new policy, the target, a rule for when to act, on decisions someone else made: the logging policy, or logger. Reweight each logged trajectory, one case's run of 4 decisions, by how much more or less likely the new policy was to take the logged actions, average the rewards, and a number comes out. That reweighting is importance sampling. The number comes out even when the logs barely cover what the new policy does and even when the reward pays for the wrong thing.

So the number has to survive four checks before it means better. Is the gain larger than its uncertainty? Does a policy that ignores the state, here the load level, get the same gain? Does the ranking hold under a different reward? Is there logged evidence at all for what the new policy wants to do?

What the checks show

Under the default reward (gain minus 1.5 × harm) and rarely logged intensification, the state-responsive target, the new policy, scores −0.10 against the logger's −0.37: a gain of 0.27, with a paired 95% interval from 0.12 to 0.40 over 256 bootstrap resamples. That is the number a dashboard would report.

The constant control intervenes at the same average rate but ignores load. It scores −0.19, 65% of the target's gain. Most of the apparent improvement is a shift in how often to act, not in when. That 65% is a ratio of two point estimates, taken before rounding (from the rounded figures here it would read 67%), so it carries no interval of its own. The part of the target's value that depends on reading the state, target minus control on the same resamples, is 0.10, paired 95% interval 0.03 to 0.17: real on this data, and small.

Drop the gain term from the reward and the target falls to −1.01 against −0.93, an interval that crosses zero. Give the logger broad support for intensifying and two things change, because that setting changes the past decisions themselves: the logger becomes a different, stronger policy, its return rising from −0.37 to 0.08, and the target, blended 25% toward the logger, moves from −0.10 to −0.01. Against that logger the target loses outright: −0.17 to −0.01.

Every one of those is the same estimator on the same kind of data. The demo's controls switch between them, and the verdict above the plot is computed from the paired interval each time rather than written in advance.

The full study

This page is a small rebuild of the questions behind Counterledger, the offline policy evaluator I built in Python. Its own study, 3,000 twelve-step trajectories, ended where these checks point: the leading estimate did not survive them, and no state-dependent value was shown. The result was a reproducible evaluation system and a negative finding. Those figures are results from the original evaluation run; the public repo is a cleaned release without those run artifacts.

What this page is not

Made-up load states and abstract actions (Hold, Adjust, Intensify), with logging probabilities known by construction. Nothing here is evidence about any real decision or its effects.

For engineers

The estimators and their formulas, the support check that refuses to estimate, Counterledger's method, what this rebuild leaves out, and how to reproduce the run.

Three estimators, three compromises

w(i,t) = Π[k≤t] π(a | s) / b(a | s)        cumulative ratio
c(i,t) = min(cap, w(i,t))
Raw IS        = Σ γ^t · mean(w · r)
Capped IS     = Σ γ^t · mean(c · r)
Self-norm. IS = Σ γ^t · Σ(c · r) / Σ c
ESS(t)        = (Σ w)² / Σ w²

π is the target, b the logger, w a trajectory's importance weight and γ the discount. The cap applies to the cumulative trajectory ratio at each step, never to each one-step ratio, which would let the bound grow as the cap to the horizon. Capping trades variance for bias. Self-normalizing keeps the estimate inside the range of possible returns but is biased in finite samples too. The effective sample size says how many trajectories the weights really use: under the default target it falls from 98 at the first step to 68 at the last, out of 160.

Intervals come from 256 bootstrap draws of whole trajectories, shared by every policy, so a difference interval is a paired contrast, not two intervals subtracted. Normalizing denominators are recomputed in each draw. If any draw has a zero denominator, the interval is withheld rather than quietly narrowed.

Support, and when to refuse

Importance weighting needs the logger to have taken, at least sometimes, every action the target takes. When it never did, no reweighting recovers the outcome. With no logged intensification at low load, the target still puts 17% on Intensify at low load, so the evaluator reports nothing for it: no value, no interval, no comparison. Capping cannot create evidence that was never logged.

The check runs against every state the generator can reach, not only the states a sample happened to contain, so a small sample that misses low load still refuses. Thin support, logged under 5%, is shown in the support table without blocking: it is a warning about variance, not a missing outcome.

Counterledger’s method

Counterledger fits backward finite-horizon Q functions with gradient boosting, adds a sequential doubly robust correction with capped cumulative ratios, separates subject-level roles for policy development, nuisance fitting and evaluation, and audits every fitted helper rather than trusting its name.

Its study covered 3,000 twelve-step trajectories: 25,200 training records, 5,400 validation records with outcomes and 5,400 outcome-blind test observations. The candidate policy came from a local quantized 14-billion-parameter model scoring restricted action choices, blended 25% with 75% behavior cloning. The leading fitted-Q estimate did not survive: a constant control nearly reproduced it, permuting the state alignment preserved it, and removing one reward component reversed the ranking. Those figures are results from the original evaluation run; the public repo is a cleaned release without those run artifacts.

What is left out

A generated cohort of generic load states and abstract actions, with logging probabilities known by construction. That removes propensity-model error, which real observational data never does. Load evolves independently of actions, so reweighting fixed trajectories is evaluation, not a simulator: nobody can take an action here and see its counterfactual future.

This page runs importance sampling only. It fits no Q function, runs no doubly robust correction and no language model, and its intervals are pointwise, unadjusted for choosing settings interactively.

Reproduce it

The engine, generator and estimators are in the code for this page. Its tests pin the one-step value, the logging-policy identity (blend fully toward the logger and every estimator returns the factual average), cumulative capping, support refusal, reward ablation and deterministic replay. From the site's Next.js app:

npx vitest run src/lib/projects/offline-policy-evaluation

Export, under the hood in the demo, writes the configuration, every logged trajectory with its rewards and weights, the support check and the bootstrap intervals, so evaluate(exported.config) reproduces the run exactly.