Individual tax prep v1
Preparing an individual tax return from client documents to a return ready for practitioner review.
We evaluate AI models on economically valuable tasks drawn from the businesses Thrive Holdings owns and operates, measuring performance alongside execution time and cost.
IT values use the 52-ticket export. Token counts are available for individual tax prep only. Focus a point for its model, effort, score, and cost.
Best result per model + Pareto frontier. Hover or focus a point for details.
AI could improve the quality and lower the cost of services people and businesses rely on every day. Evaluating the work behind those services gives us a tangible way to assess AI model progress. Across the 70+ businesses we own and operate, our engineers work alongside practitioners to turn real assignments, decisions, and corrections into tests that reflect professional standards. This helps our businesses identify where AI can create value.
We're releasing benchmarks for individual tax preparation at our accounting platform Current (opens in a new tab) and helpdesk ticket resolution at our IT services platform Shield (opens in a new tab). Drawing on 150 end-to-end tasks from production workflows, they preserve the context needed to measure task performance, execution time, and cost. We use the results to choose models, identify failures, and improve the systems around them. We're also sharing a small selection of tasks so other builders can examine the work and test their own approaches.
Preparing an individual tax return from client documents to a return ready for practitioner review.
Investigating support requests and autonomously resolving issues to help customers get back to work.
Holdings-IndividualTaxBench
As an AI tax preparer, models review client documents, identify relevant tax information, and map it to the fields and schedules required for a structured federal Form 1040 return.
| Model | Effort | Accuracy | Recall | Precision | Minutes per package | Cost per return (USD) | Tokens per return |
|---|---|---|---|---|---|---|---|
| Astra | high | 79.20% | 79.20% | 78.91% | 3.55 | $3.47 | 1.13M |
| Grok 4.6 | medium | 79.05% | 79.05% | 79.69% | 7.15 | $1.54 | 1.03M |
| Astra | xhigh | 78.93% | 78.93% | 78.25% | 4.00 | $4.05 | 1.43M |
| Opus 5 | medium | 78.72% | 78.72% | 80.34% | 3.75 | $3.06 | 2.94M |
| Grok 4.6 | xhigh | 77.83% | 77.83% | 78.56% | 9.08 | $1.57 | 1.24M |
| Astra | low | 77.75% | 77.75% | 79.56% | 2.14 | $3.09 | 986.79K |
| Sol | high | 77.64% | 77.64% | 79.26% | 4.25 | $2.01 | 1.63M |
| Sol | medium | 76.98% | 76.98% | 80.11% | 2.26 | $1.45 | 972.5K |
| GLM 5.3 | high | 76.51% | 76.51% | 78.57% | 4.14 | $1.06 | 2.7M |
| Opus 5 | low | 76.41% | 76.41% | 79.71% | 2.62 | $2.50 | 2.22M |
| Terra | high | 74.84% | 74.84% | 80.33% | 2.54 | $0.94 | 1.23M |
| Sonnet 5 | xhigh | 74.29% | 74.29% | 78.85% | 10.10 | $1.92 | 3.81M |
| Sonnet 5 | high | 74.28% | 74.28% | 80.94% | 7.02 | $1.59 | 3.46M |
| Kimi K3 | low | 74.02% | 74.02% | 78.92% | 3.60 | $0.93 | 1.48M |
| Gemini 3.7 Flash | low | 73.69% | 73.69% | 81.61% | 5.36 | $0.98 | 7.78M |
| Luna | high | 73.28% | 73.28% | 77.36% | 3.31 | $0.13 | 1.49M |
| Grok 4.6 | low | 73.24% | 73.24% | 79.82% | 3.03 | $0.82 | 671.33K |
| DeepSeek V4 Pro | high | 72.46% | 72.46% | 78.98% | 11.52 | $0.40 | 2.19M |
| Terra | medium | 72.05% | 72.05% | 80.94% | 1.30 | $0.60 | 726.08K |
| Sol | low | 71.62% | 71.62% | 80.13% | 1.29 | $0.99 | 570.67K |
| GLM 5.3 | low | 71.20% | 71.20% | 79.23% | 1.47 | $0.59 | 1.47M |
| Sonnet 5 | medium | 67.98% | 67.98% | 80.80% | 4.32 | $1.22 | 2.61M |
| Terra | low | 67.00% | 67.00% | 81.07% | 1.11 | $0.54 | 638.87K |
| Luna | medium | 65.00% | 65.00% | 78.32% | 1.36 | $0.07 | 764.06K |
| MiniMax M3 | high | 57.77% | 57.77% | 79.92% | 6.40 | $0.37 | 4.13M |
| Sonnet 5 | low | 55.33% | 55.33% | 81.75% | 2.61 | $0.88 | 1.5M |
| Luna | low | 52.67% | 52.67% | 83.34% | 0.98 | $0.05 | 503.79K |
Grok 4.6 medium achieves 79.05% accuracy at $1.54 per return, compared with Astra high at 79.20% and $3.47. Grok costs roughly 56% less but takes about twice as long, averaging 7.15 versus 3.55 minutes per return. For this pair, the practical trade-off is lower cost versus faster completion, with little separation in measured accuracy.
Luna at high reasoning achieves 73.28% accuracy at $0.13 per return, while the open-weight DeepSeek V4 Pro model at high reasoning achieves 72.46% at $0.40. Luna delivers higher measured accuracy at roughly one-third the execution cost. In this comparison, choosing the proprietary model reduces cost rather than adding a premium.
Compared with Astra high, extra-high reasoning scores better on 10 returns, ties on 27, and scores worse on 3. One decline of 18.90 percentage points outweighs the smaller gains elsewhere. Average accuracy falls from 79.20% to 78.93% while cost rises to $4.05 per return. Additional reasoning improves more returns than it hurts, but the size of the regressions determines the overall result. On this cohort, a higher reasoning budget does not produce more reliable performance.
The task is to reason over a package of PDFs, spreadsheets, text, and email and produce a structured individual tax return. Inputs range from W-2s, 1099s, and K-1s to brokerage statements, accountant emails, and handwritten or scanned notes. We parse documents into Markdown so the benchmark compares tax reasoning consistently, including for models without vision capabilities. Each agent fills a JSON schema containing hundreds to more than a thousand possible fields for Form 1040 and its supporting schedules, ready to map into tax software.
Some values map directly from a document. Others require calculations and reconciliation across the package—for example, combining business income and expenses for Schedule C or organizing rental activity and separately extracted K-1 information for Schedule E. The model must decide which fields apply while keeping entities, values, and dependencies consistent across schedules.
Knowledge base
PDFs, spreadsheets, text, email
Parsed into Markdown
Agent
One session, end to end
Reasons over the full package
Structured output
Form 1040 JSON
Hundreds to 1,000+ possible fields
Complexity reflects how difficult a package is to prepare, from familiar W-2 and 1099 mappings to consolidated statements, mixed sources, and schedules that require cross-document calculations. Document volume alone does not capture that difficulty: a large stack of similar K-1s can be easier than a few ambiguous workpapers. These 40 packages retain their original labels as an approximate measure of task difficulty.
3 packages are unlabeled.
These tasks are curated from production work across the accounting firms we own. The source material includes client documents, prior-year returns, preparer and reviewer interactions, initial AI passes, and the corrections behind practitioner-reviewed outcomes.
We check that each package's supplied inputs support its expected answers, then freeze those inputs and keep the reference return separate. All 27 model and reasoning configurations are evaluated on the same 40 packages, with each package weighted equally.
We use the lightweight Pi harness (opens in a new tab) to give one agent the full preparation task in a single session. Every configuration receives the same prompt, schemas, frozen client documents, tax-engine version, and sandboxed file and Bash tools. Agents can inspect files, run scripts, and search the web, but they do not use sub-agents or our specialized production extractors. This isolates the model and reasoning effort within a consistent, minimal harness.
We convert both the model output and reference return into canonical Form 1040 JSON. Because entities can appear in different places and names or formatting can vary, Astra at medium reasoning acts as an LLM judge to reconcile corresponding schedules and values. The grader awards partial credit for individual facts rather than passing or failing an entire return. Accuracy therefore measures how much expected information was captured correctly; it is not a whole-return pass rate.
Our production Tax AI product uses specialized classification, splitting, extraction, prompting, and model-selection workflows. This benchmark measures raw model capabilities under a shared evaluation setup, not the performance of that optimized production system. Practitioner corrections become focused evaluations that help improve the production prompts, tools, and grading system. Read more in Building self-improving tax agents with Codex (opens in a new tab).
| # | MODEL | ||||||
|---|---|---|---|---|---|---|---|
| 1 | Astrahigh | 79.20% | 78.91% | 3.55 | $3.47 | 1.13M | 12.70 |
| 2 | Grok 4.6medium | 79.05% | 79.69% | 7.15 | $1.54 | 1.03M | 11.75 |
| 3 | Astraxhigh | 78.93% | 78.25% | 4.00 | $4.05 | 1.43M | 14.93 |
| 4 | Opus 5medium | 78.72% | 80.34% | 3.75 | $3.06 | 2.94M | 22.23 |
| 5 | Grok 4.6xhigh | 77.83% | 78.56% | 9.08 | $1.57 | 1.24M | 13.18 |
| 6 | Astralow | 77.75% | 79.56% | 2.14 | $3.09 | 986.79K | 12.58 |
| 7 | Solhigh | 77.64% | 79.26% | 4.25 | $2.01 | 1.63M | 15.35 |
| 8 | Solmedium | 76.98% | 80.11% | 2.26 | $1.45 | 972.5K | 10.88 |
| 9 | GLM 5.3high | 76.51% | 78.57% | 4.14 | $1.06 | 2.7M | 25.48 |
| 10 | Opus 5low | 76.41% | 79.71% | 2.62 | $2.50 | 2.22M | 17.50 |
Package means across the same shared cohort for every model. Costs are in USD per return. LLM-judge scores are not supplied in this release.
Reported aggregate results. Precision is Classic Precision. Time, cost, and tokens are averages per return.
Launched v1 of the individual tax preparation and helpdesk ticket resolution benchmarks.