← All benchmarks
Finance benchmark · by Column Tax

TaxCalcBench

US tax return benchmark: Can an AI complete a held-out US tax-return case correctly? TaxCalcBench gives the model a taxpayer’s W-2s, 1099s, and details, then asks it to produce the complete Form 1040. Every line it computes is checked against Column Tax’s reference 1040 for that held-out case, and the score is its by-line accuracy.

Use this board to compare tested agent configurations on these held-out tax returns. It measures by-line accuracy on this task suite, not general model quality or production reliability.

How to read this benchmark

Each held-out return is graded line-by-line against Column Tax’s reference 1040 for that case. Chart cost is the mean recorded cost per return in the published run data.

Last run
Held-out tasks
51
Configurations
5 agent configurations
Chart metric
By-line accuracy on a 0–100% scale

An agent profile is the exact instruction and tool configuration used for a run. Hover a chart label to see its immutable profile ID. 95% Student-t interval across the 51 per-return by-line accuracy scores. Small suites are narrow capability checks, not evidence that one model is generally better.

Read the broader agent-evaluation methodology →
harness (tools)raw model95% CI
By-line accuracy, %
0
20
40
60
80
100
91.4
91.2
90.1
72.8
70.7
gpt-5.5via codexDefault profilen=51mean cost/return $0.18
opusvia Claude CodeDefault profilen=51mean cost/return $0.21
Claude Sonnetvia Claude CodeDefault profilen=51mean cost/return $0.17
haikuvia Claude CodeDefault profilen=51mean cost/return $0.06
gpt-5.1via direct callNo profilen=51mean cost/return $0.03

Same 51 held-out cases for every bar.