TaxCalcBench
US tax return benchmark: Can an AI complete a held-out US tax-return case correctly? TaxCalcBench gives the model a taxpayer’s W-2s, 1099s, and details, then asks it to produce the complete Form 1040. Every line it computes is checked against Column Tax’s reference 1040 for that held-out case, and the score is its by-line accuracy.
Use this board to compare tested agent configurations on these held-out tax returns. It measures by-line accuracy on this task suite, not general model quality or production reliability.
How to read this benchmark
Each held-out return is graded line-by-line against Column Tax’s reference 1040 for that case. Chart cost is the mean recorded cost per return in the published run data.
- Last run
- Held-out tasks
- 51
- Configurations
- 5 agent configurations
- Chart metric
- By-line accuracy on a 0–100% scale
An agent profile is the exact instruction and tool configuration used for a run. Hover a chart label to see its immutable profile ID. 95% Student-t interval across the 51 per-return by-line accuracy scores. Small suites are narrow capability checks, not evidence that one model is generally better.
Read the broader agent-evaluation methodology →
gpt-5.5via codexDefault profilen=51mean cost/return $0.18Same 51 held-out cases for every bar.