Individual tax prep v1
Preparing an individual tax return from client documents to a return ready for practitioner review.
We evaluate AI models on economically valuable tasks drawn from the businesses Thrive Holdings owns and operates, measuring performance alongside execution time and cost.
IT values use the 52-ticket export. Token counts are available for individual tax prep only. Focus a point for its model, effort, score, and cost.
Best result per model + Pareto frontier. Hover or focus a point for details.
AI could improve the quality and lower the cost of services people and businesses rely on every day. Evaluating the work behind those services gives us a tangible way to assess AI model progress. Across the 70+ businesses we own and operate, our engineers work alongside practitioners to turn real assignments, decisions, and corrections into tests that reflect professional standards. This helps our businesses identify where AI can create value.
We're releasing benchmarks for individual tax preparation at our accounting platform Current (opens in a new tab) and helpdesk ticket resolution at our IT services platform Shield (opens in a new tab). Drawing on 150 end-to-end tasks from production workflows, they preserve the context needed to measure task performance, execution time, and cost. We use the results to choose models, identify failures, and improve the systems around them. We're also sharing a small selection of tasks so other builders can examine the work and test their own approaches.
Preparing an individual tax return from client documents to a return ready for practitioner review.
Investigating support requests and autonomously resolving issues to help customers get back to work.
Holdings-IndividualTaxBench
As an AI tax preparer, models review client documents, identify relevant tax information, and map it to the fields and schedules required for a structured federal Form 1040 return.
| Model | Effort | Accuracy | Recall | Precision | Minutes per package | Cost per return (USD) | Tokens per return |
|---|---|---|---|---|---|---|---|
| Astra | high | 79.20% | 79.20% | 78.91% | 3.55 | $3.47 | 1.13M |
| Grok 4.6 | medium | 79.05% | 79.05% | 79.69% | 7.15 | $1.54 | 1.03M |
| Astra | xhigh | 78.93% | 78.93% | 78.25% | 4.00 | $4.05 | 1.43M |
| Opus 5 | medium | 78.72% | 78.72% | 80.34% | 3.75 | $3.06 | 2.94M |
| Grok 4.6 | xhigh | 77.83% | 77.83% | 78.56% | 9.08 | $1.57 | 1.24M |
| Astra | low | 77.75% | 77.75% | 79.56% | 2.14 | $3.09 | 986.79K |
| Sol | high | 77.64% | 77.64% | 79.26% | 4.25 | $2.01 | 1.63M |
| Sol | medium | 76.98% | 76.98% | 80.11% | 2.26 | $1.45 | 972.5K |
| GLM 5.3 | high | 76.51% | 76.51% | 78.57% | 4.14 | $1.06 | 2.7M |
| Opus 5 | low | 76.41% | 76.41% | 79.71% | 2.62 | $2.50 | 2.22M |
| Terra | high | 74.84% | 74.84% | 80.33% | 2.54 | $0.94 | 1.23M |
| Sonnet 5 | xhigh | 74.29% | 74.29% | 78.85% | 10.10 | $1.92 | 3.81M |
| Sonnet 5 | high | 74.28% | 74.28% | 80.94% | 7.02 | $1.59 | 3.46M |
| Kimi K3 | low | 74.02% | 74.02% | 78.92% | 3.60 | $0.93 | 1.48M |
| Gemini 3.7 Flash | low | 73.69% | 73.69% | 81.61% | 5.36 | $0.98 | 7.78M |
| Luna | high | 73.28% | 73.28% | 77.36% | 3.31 | $0.13 | 1.49M |
| Grok 4.6 | low | 73.24% | 73.24% | 79.82% | 3.03 | $0.82 | 671.33K |
| DeepSeek V4 Pro | high | 72.46% | 72.46% | 78.98% | 11.52 | $0.40 | 2.19M |
| Terra | medium | 72.05% | 72.05% | 80.94% | 1.30 | $0.60 | 726.08K |
| Sol | low | 71.62% | 71.62% | 80.13% | 1.29 | $0.99 | 570.67K |
| GLM 5.3 | low | 71.20% | 71.20% | 79.23% | 1.47 | $0.59 | 1.47M |
| Sonnet 5 | medium | 67.98% | 67.98% | 80.80% | 4.32 | $1.22 | 2.61M |
| Terra | low | 67.00% | 67.00% | 81.07% | 1.11 | $0.54 | 638.87K |
| Luna | medium | 65.00% | 65.00% | 78.32% | 1.36 | $0.07 | 764.06K |
| MiniMax M3 | high | 57.77% | 57.77% | 79.92% | 6.40 | $0.37 | 4.13M |
| Sonnet 5 | low | 55.33% | 55.33% | 81.75% | 2.61 | $0.88 | 1.5M |
| Luna | low | 52.67% | 52.67% | 83.34% | 0.98 | $0.05 | 503.79K |
Grok 4.6 medium achieves 79.05% accuracy at $1.54 per return, compared with Astra high at 79.20% and $3.47. Grok costs roughly 56% less but takes about twice as long, averaging 7.15 versus 3.55 minutes per return. For this pair, the practical trade-off is lower cost versus faster completion, with little separation in measured accuracy.
Luna at high reasoning achieves 73.28% accuracy at $0.13 per return, while the open-weight DeepSeek V4 Pro model at high reasoning achieves 72.46% at $0.40. Luna delivers higher measured accuracy at roughly one-third the execution cost. In this comparison, choosing the proprietary model reduces cost rather than adding a premium.
Compared with Astra high, extra-high reasoning scores better on 10 returns, ties on 27, and scores worse on 3. One decline of 18.90 percentage points outweighs the smaller gains elsewhere. Average accuracy falls from 79.20% to 78.93% while cost rises to $4.05 per return. Additional reasoning improves more returns than it hurts, but the size of the regressions determines the overall result. On this cohort, a higher reasoning budget does not produce more reliable performance.
All 27 model and reasoning configurations are evaluated on the same 40 returns, with each return weighted equally in the reported averages. The current complexity labels are 19 simple, 12 moderate, 5 complex, and 1 ultra-complex; the remaining 3 returns have no complexity label.
We use the lightweight Pi harness (opens in a new tab) to evaluate models under consistent conditions, with the same inputs, tools, and task instructions and no model-specific workflow optimizations. Each tax package is handled by one model and one agent. Each run has access to frozen Markdown documents, original Excel workbooks, sandboxed file and Bash tools, and web search. Agents can inspect files, execute commands, write and run scripts, and search the web.
Every configuration receives the same client documents and uses the same tax engine version. Prior-year information and preparation or review notes, when included, are provided as frozen snapshots to keep the task consistent across models.
We convert the model output and the reference return exported from our tax engine into canonical Form 1040 JSON, then compare them using Astra at medium reasoning as an LLM-as-judge. Exact matching alone is insufficient because entity names, formatting, and how information is represented can differ between outputs. The judge matches corresponding entities and accounts for these nuances to distinguish equivalent information from substantive discrepancies.
Our production Tax AI product uses specialized, form-specific extractors, prompts, and model selection, with document processing, context, and tools that vary by workflow. These benchmarks more closely reflect raw model capabilities under a shared evaluation setup than the performance of our specialized production product.
Practitioner corrections become focused evaluations that help us improve Tax AI's prompts, tools, and evaluation system. Repeating the same tasks under consistent conditions lets us measure whether those changes improve performance. Read more in Building self-improving tax agents with Codex (opens in a new tab).
| # | MODEL | ||||||
|---|---|---|---|---|---|---|---|
| 1 | Astrahigh | 79.20% | 78.91% | 3.55 | $3.47 | 1.13M | 12.70 |
| 2 | Grok 4.6medium | 79.05% | 79.69% | 7.15 | $1.54 | 1.03M | 11.75 |
| 3 | Astraxhigh | 78.93% | 78.25% | 4.00 | $4.05 | 1.43M | 14.93 |
| 4 | Opus 5medium | 78.72% | 80.34% | 3.75 | $3.06 | 2.94M | 22.23 |
| 5 | Grok 4.6xhigh | 77.83% | 78.56% | 9.08 | $1.57 | 1.24M | 13.18 |
| 6 | Astralow | 77.75% | 79.56% | 2.14 | $3.09 | 986.79K | 12.58 |
| 7 | Solhigh | 77.64% | 79.26% | 4.25 | $2.01 | 1.63M | 15.35 |
| 8 | Solmedium | 76.98% | 80.11% | 2.26 | $1.45 | 972.5K | 10.88 |
| 9 | GLM 5.3high | 76.51% | 78.57% | 4.14 | $1.06 | 2.7M | 25.48 |
| 10 | Opus 5low | 76.41% | 79.71% | 2.62 | $2.50 | 2.22M | 17.50 |
Package means across the same shared cohort for every model. Costs are in USD per return. LLM-judge scores are not supplied in this release.
Reported aggregate results. Precision is Classic Precision. Time, cost, and tokens are averages per return.
Launched v1 of the individual tax preparation and helpdesk ticket resolution benchmarks.