Individual tax prep
Preparing an individual tax return from client documents to a return ready for practitioner review.
We evaluate AI models on economically valuable tasks drawn from the businesses Thrive Holdings owns and operates, measuring performance alongside execution time and cost.
IT values use the 52-ticket export. Token counts are available for individual tax prep only. Focus a point for its model, effort, score, and cost.
AI could improve the quality and lower the cost of services people and businesses rely on every day. Evaluating the work behind those services gives us a tangible way to assess AI model progress. Across the 70+ businesses we own and operate, our engineers work alongside practitioners to turn real assignments, decisions, and corrections into tests that reflect professional standards. This helps our businesses identify where AI can create value.
We're releasing benchmarks for individual tax preparation at Current (opens in a new tab) and IT helpdesk ticket resolution at Shield (opens in a new tab). Drawing on 92 tasks from production workflows, they preserve the context needed to measure task performance, execution time, and cost. We use the results to choose models, identify failures, and improve the systems around them. We're also sharing a small selection of tasks so other builders can examine the work and test their own approaches.
Preparing an individual tax return from client documents to a return ready for practitioner review.
Investigating support requests and autonomously resolving issues to help customers get back to work.
Holdings-IndTaxBench
Across 40 tax returns, models work from client documents provided in markdown, identify the relevant tax information, and map it to the fields and schedules required by our tax engine. The task evaluates how well models turn source material into a structured Form 1040 return for practitioner review.
| Model | Effort | Accuracy | Recall | Precision | Minutes per package | Extraction USD per return | Tokens per return |
|---|---|---|---|---|---|---|---|
| Astra | xhigh | 77.63% | 77.63% | 76.77% | 4.10 | $4.09 | 1.46M |
| Astra | high | 77.53% | 77.53% | 76.92% | 3.60 | $3.50 | 1.14M |
| Opus 5 | medium | 75.30% | 75.30% | 77.65% | 3.80 | $3.09 | 2.96M |
| Grok 4.6 | medium | 75.64% | 75.64% | 77.18% | 7.20 | $1.55 | 1.04M |
| Grok 4.6 | xhigh | 75.12% | 75.12% | 76.39% | 9.10 | $1.58 | 1.25M |
| Sol | high | 75.80% | 75.80% | 78.84% | 4.30 | $2.03 | 1.64M |
| GLM 5.3 | high | 75.08% | 75.08% | 77.79% | 4.20 | $1.07 | 2.74M |
| Astra | low | 74.49% | 74.49% | 77.55% | 2.20 | $3.11 | 993.5K |
| Sol | medium | 74.71% | 74.71% | 78.96% | 2.30 | $1.47 | 987K |
| Opus 5 | low | 72.87% | 72.87% | 77.44% | 2.60 | $2.51 | 2.23M |
| Sonnet 5 | high | 73.32% | 73.32% | 80.50% | 7.20 | $1.61 | 3.49M |
| Sonnet 5 | xhigh | 71.46% | 71.46% | 77.06% | 10.30 | $1.94 | 3.85M |
| Kimi K3 | low | 72.02% | 72.02% | 78.07% | 3.60 | $0.94 | 1.49M |
| Luna | high | 71.09% | 71.09% | 75.00% | 3.30 | $0.13 | 1.51M |
| Gemini 3.7 Flash | low | 70.33% | 70.33% | 79.27% | 5.40 | $0.99 | 7.84M |
| Terra | high | 70.33% | 70.33% | 78.88% | 2.60 | $0.95 | 1.24M |
| DeepSeek V4 Pro | high | 70.14% | 70.14% | 77.91% | 11.60 | $0.41 | 2.21M |
| Grok 4.6 | low | 70.33% | 70.33% | 79.71% | 3.00 | $0.82 | 672.5K |
| Terra | medium | 68.58% | 68.58% | 80.95% | 1.30 | $0.60 | 733.25K |
| GLM 5.3 | low | 66.41% | 66.41% | 77.19% | 1.50 | $0.59 | 1.48M |
| Sonnet 5 | medium | 63.64% | 63.64% | 80.22% | 4.40 | $1.24 | 2.63M |
| Sol | low | 63.83% | 63.83% | 78.85% | 1.30 | $1.00 | 583K |
| Terra | low | 60.60% | 60.60% | 80.01% | 1.10 | $0.55 | 643K |
| Luna | medium | 56.91% | 56.91% | 76.18% | 1.40 | $0.07 | 769.75K |
| MiniMax M3 | high | 53.30% | 53.30% | 72.56% | 6.50 | $0.37 | 4.17M |
| Sonnet 5 | low | 53.00% | 53.00% | 80.73% | 2.60 | $0.89 | 1.51M |
| Luna | low | 40.85% | 40.85% | 80.94% | 1.00 | $0.06 | 509K |
27 model and reasoning configurations
| # | MODEL | |||||
|---|---|---|---|---|---|---|
| 1 | Astraxhigh | 77.63% | 76.77% | 4.10 | $4.09 | 1.46M |
| 2 | Astrahigh | 77.53% | 76.92% | 3.60 | $3.50 | 1.14M |
| 3 | Solhigh | 75.80% | 78.84% | 4.30 | $2.03 | 1.64M |
| 4 | Grok 4.6medium | 75.64% | 77.18% | 7.20 | $1.55 | 1.04M |
| 5 | Opus 5medium | 75.30% | 77.65% | 3.80 | $3.09 | 2.96M |
| 6 | Grok 4.6xhigh | 75.12% | 76.39% | 9.10 | $1.58 | 1.25M |
| 7 | GLM 5.3high | 75.08% | 77.79% | 4.20 | $1.07 | 2.74M |
| 8 | Solmedium | 74.71% | 78.96% | 2.30 | $1.47 | 987K |
| 9 | Astralow | 74.49% | 77.55% | 2.20 | $3.11 | 993.5K |
| 10 | Sonnet 5high | 73.32% | 80.50% | 7.20 | $1.61 | 3.49M |
Reported aggregate results. Precision is Classic Precision. Time, cost, and tokens are averages per return.
Our benchmarks draw on work performed inside Current and Shield, with the context and tools needed to complete each assignment. Models run through the Holdings proprietary evaluation system under consistent conditions. Tax runs use the same client documents, prior-year information, and tax engine version. IT runs replay recorded tool responses in an isolated environment, with write actions mocked and the starting state reset between runs.
The overview shows the Pareto frontier: configurations offering the best score for their selected cost, time, or token usage. Detailed charts and tables include all reported configurations. Tax cost and tokens per return are the source cohort totals divided by 40; time is average extraction time. IT reports cost and elapsed time per task; its export does not include token usage. Completion counts, run identifiers, and turn counts are not reported. Computer-use results and human-cost comparisons remain illustrative. Human-cost scenarios add an assumed practitioner-review allowance to execution cost and compare it with a fixed labor baseline, using the assumptions stated beneath each chart. They do not represent measured savings.
Practitioner corrections become focused evaluations that help us improve prompts, tools, and our evaluation system. Repeating tasks under consistent conditions helps us understand whether those changes improve performance. Results describe a particular system and task set, not every job in an industry. We plan to extend these benchmarks as our systems improve and we build in additional industries, helping practitioners spend more time on what customers value most. Read how we use practitioner feedback to improve products like Tax AI in Building self-improving tax agents with Codex (opens in a new tab).
Updated individual tax preparation results from 28 to 40 graded returns per configuration.
Launched two benchmarks: individual tax prep and IT helpdesk ticket resolution.