Skip to content

Spreadsheet Reasoning Inspector

Formuloom is a pipeline I built in Python for reviewing changes to financial models: it sorts a workbook's changed cells into final outputs, the values someone would report, and intermediates, the working steps. Rebuilt here on one small workbook, with rules voting on each cell.

Original built with Python

What I found

0.970 / 0.920

precision and recall of Formuloom itself on its 24-sheet development corpus: of the cells it marked final, 97% were, and it found 92% of the real ones. These are results from the original evaluation run; the public repo is a cleaned release without those run artifacts. The two cards beside it are this rebuild's counterexamples.

1

false final in this rebuild's workbook under balanced votes, Formuloom's own majority rule: a scratch formula no other cell uses, final only because it is a dependency sink, a cell nothing depends on.

0%

precision of the precision gate, a stricter policy this page adds, on the same workbook: it sends both real finals to review and still keeps the scratch cell final.

Try it

With balanced votes the rules find 2 of 2 final values and wrongly mark 1 scratch cell final. The precision gate, built to be stricter, finds none and still marks the scratch cell.

A final value is one someone would report; an intermediate is a working step. The answer key marks Calc!B4 and Report!B2 final (both “Total margin”); Report!B4, “Scratch estimate”, is not. 7 simple rules each vote final, intermediate or abstain on every cell, and the rule policy turns their votes into a label; a cell the policy cannot settle is sent to review, for a person to decide. Lines run from a value to the cells that use it; on a phone the inspector below lists them. Click a cell to see how each rule voted.

Rule policy
Prune-only review

A prune-only review can remove a final and never add one. The reviewer here can see the answer key, so dropping Report!B4 is not the rules getting better.

Finals foundof the key's finals
2 of 2
False finalsmarked final, not in the key
1
Sent to reviewfinal and intermediate votes tied
1
Precisionmarked finals that are right
67%
F1precision and recall in one score
80%
  • Final
  • Intermediate
  • Review: final and intermediate votes tied
  • Wrong against the key

Inputs support

Calc output

Report output

Report!B4

Scratch estimate: 6 final

Balanced: positive votes must outnumber negative votes

Each rule votes final or intermediate, or abstains when it has no opinion.

  • Dependency sinka formula no other cell usesfinal0 downstream consumers
  • Range aggregationadds up a range with SUMabstainsNo range aggregation
  • Output lexiconits label says total, net or outputabstainsNo output term in label
  • Input literala typed-in value, not a formulaabstainsFormula
  • Pass-throughonly points at one other cellabstainsNot a bare reference
  • Consumed downstreamother cells use itabstains0 downstream consumers
  • Emphasized checkpointemphasized in the sheet, with a total label or a SUMabstainsNo emphasis
Edit the workbook, review a cell, the illustrative ballots, export and replay

11 cells · 10 dependency edges · computed in your browser, no model calls. Edit a cell, then recalculate.

CellRow labelValueLabel
Units120intermediate
Unit price12intermediate
Unit cost5intermediate
Overhead100intermediate
Revenue1,440intermediate
Expense contribution-600intermediate
Total margin840final
After overhead740intermediate
Total margin840final
Working check740review
Scratch estimate6final

Edit Report!B4

A support sheet's cells are always intermediate: it holds working steps, not results.

Review can only remove a proposed final; it never adds one.

Illustrative ballots

Three votes per cell, showing how a majority over repeated votes reads. Authored illustrative ballots, not actual model output. Bound to the unedited baseline workbook. No live inference.

finalintermediateintermediate → majority intermediate

Contents

When is a spreadsheet value final?

Formuloom was built for reviewing changes to financial models, such as discounted cash flow, three-statement and M&A workbooks. One update can change thousands of cells on a sheet, and a reviewer needs the outputs that moved, not every working step that moved with them.

A cell can be numerically right and still the wrong thing to report. A margin subtotal can be a real result for its section and also feed a later calculation; a scratch formula can have no consumers at all and mean nothing. So the problem is not evaluating formulas, it is recovering the scope in which a value is final. A dependency graph, the map of which cells feed which, is evidence for that, not the definition.

What the rules get right and wrong

The workbook here is 11 cells across an inputs sheet and two output sheets, with 2 values final in its answer key. Balanced votes, which call a cell final when more rules vote final than intermediate (abstentions do not count), find both and mark one extra: the scratch estimate Report!B4, final only because nothing uses it. Precision, the share of marked finals that are right, is 67%; recall, the share of real finals found, 100%.

Balanced votes are Formuloom's own rule. The precision gate is not: Formuloom has no such vote rule, and this page adds it as a stricter alternative. It requires a final to have no negative vote. That sounds stricter, and on this workbook it is worse: the margin checkpoint Calc!B4 is used by a later cell, which draws one negative vote, so the gate sends it to review along with the report total, while the scratch cell, with no negative vote at all, stays final. Precision 0%. Rules that agree are not calibrated, and a gate named for a metric does not deliver it.

A prune-only review, a second pass that can remove a final but never add one, takes the balanced policy to precision 100% and F1 100% by dropping the scratch cell. That is a reviewer with the key in view, not a model improvement.

Where this comes from

This page is a reduced rebuild of Formuloom, the workbook classification pipeline I built in Python: given two versions of a workbook, it sorts the changed cells into final outputs and intermediates. On its development corpus of 24 sheets it recorded precision 0.970 and recall 0.920, the original run's figures in the first card above.

What this page is not

One small authored workbook and its key, with restricted arithmetic: not an Excel engine, and no model runs here. The rules are Formuloom's approach, rebuilt, not its Python pipeline.

For engineers

The formula grammar, the rules and policies, Formuloom's full pipeline and evaluation, the limits in full, and how to rerun it.

How a cell is labelled

A bounded formula grammar. Formulas parse into an expression tree; nothing typed is executed as code. Numbers, A1 and absolute references, simple cross-sheet references, arithmetic and SUM over ranges of up to 64 cells. A depth-first evaluation finds cycles; a cycle, a missing reference or a division by zero is an explicit error that blocks its dependents, never a silent zero.

Abstaining rules. 7 labelling functions vote final, intermediate or abstain: Dependency sink, Range aggregation, Output lexicon, Input literal, Pass-through, Consumed downstream, Emphasized checkpoint. For the checkpoint the evidence is 3 final votes against 1. Balanced votes need more final than intermediate votes, abstentions ignored, and send a tie to review; the precision gate also sends to review any final with a vote against; a support sheet routes its cells to intermediate.

Which policy is Formuloom's. Balanced votes follow Formuloom's majority_vote_label (labelmodel.py), which marks a cell final when its final votes outnumber its intermediate votes; Formuloom labels a tie intermediate, where this page sends it to review. The precision gate has no counterpart in its code: Formuloom's V11 precision guards (cli.py) work on whole sheets, dropping every proposed final on a large granular or dense source sheet, not on one cell's votes.

Prune-only review. A reviewer can keep or drop a proposed final and nothing else, so review can never introduce a final the rules did not propose or rescue one they missed. The illustrative ballots under the hood are authored, not model output, and disappear as soon as the workbook changes.

Formuloom, in detail

Formuloom extracts formulas and cached values from paired workbook snapshots, builds cross-sheet dependency features, compresses repeated rows into formula patterns, and classifies changed cells with rules, repeated model votes, a weak-supervision label model and a structural router over precision- and recall-oriented variants, with a prune-only judge as the second pass.

Its development evaluation, five task bundles and 24 sheets, recorded precision 0.970, recall 0.920 and micro F1 0.945 for the structural router (1,543 true positives, 47 false positives, 134 false negatives), with a sheet-clustered bootstrap F1 interval of 0.876 to 0.982 over 1,000 resamples. Like the card above, these come from the original run, and they describe that development corpus, not new workbooks.

The lesson this page keeps: repeated votes suppress unstable decisions, but consensus can reinforce a wrong convention for what final means, and a global prune can remove legitimate outputs. Applying review only where the sheet structure called for it worked better than applying it everywhere.

Its limits in detail

Up to 48 cells and restricted arithmetic: not an Excel engine. No named ranges, dynamic arrays, dates, text, external workbooks or Excel coercion. The key fits one authored workbook; editing a formula, a label or a sheet role may change what is final, so an edited run keeps its computation and drops the comparison rather than scoring against a key that no longer applies. There is no model, no workbook ingestion and no trained label model here.

Reproduce it

The parser, engine, rules and fixtures are in the code for this page. The tests cover arithmetic and precedence, references, ranges, cycles, invalid sessions, the counterexamples above, prune-only membership and replay; an export is a compact session of at most 150 KB, and replay recomputes every value and label instead of trusting the file. From the site's Next.js app:

npx vitest run src/lib/projects/spreadsheet-reasoning