Thrive Holdings

New York City

  

Measuring AI on real-world work

We evaluate AI models on economically valuable tasks drawn from the businesses Thrive Holdings owns and operates, measuring performance alongside execution time and cost.

Model performance across benchmarks

IT values use the 52-ticket export. Token counts are available for individual tax prep only. Focus a point for its model, effort, score, and cost.

Why we built these benchmarks

AI could improve the quality and lower the cost of services people and businesses rely on every day. Evaluating the work behind those services gives us a tangible way to assess AI model progress. Across the 70+ businesses we own and operate, our engineers work alongside practitioners to turn real assignments, decisions, and corrections into tests that reflect professional standards. This helps our businesses identify where AI can create value.

We're releasing benchmarks for individual tax preparation at Current (opens in a new tab) and IT helpdesk ticket resolution at Shield (opens in a new tab). Drawing on 92 tasks from production workflows, they preserve the context needed to measure task performance, execution time, and cost. We use the results to choose models, identify failures, and improve the systems around them. We're also sharing a small selection of tasks so other builders can examine the work and test their own approaches.

Our benchmarks

Individual tax prep

Preparing an individual tax return from client documents to a return ready for practitioner review.

IT helpdesk ticket resolution

Investigating support requests and autonomously resolving issues to help customers get back to work.

Holdings-IndTaxBench

Individual tax prep

Form 1040 tax preparation

Across 40 tax returns, models work from client documents provided in markdown, identify the relevant tax information, and map it to the fields and schedules required by our tax engine. The task evaluates how well models turn source material into a structured Form 1040 return for practitioner review.

Reasoning effort
  • low
  • medium
  • high
  • xhigh
Lines connect reasoning efforts within the same model family
Models
  • Astra
  • Opus 5
  • Grok 4.6
  • Sol
  • GLM 5.3
  • Sonnet 5
  • Kimi K3
  • Luna
  • Gemini 3.7 Flash
  • Terra
  • DeepSeek V4 Pro
  • MiniMax M3
All individual tax model and reasoning effort results
ModelEffortAccuracyRecallPrecisionMinutes per packageExtraction USD per returnTokens per return
Astraxhigh77.63%77.63%76.77%4.10$4.091.46M
Astrahigh77.53%77.53%76.92%3.60$3.501.14M
Opus 5medium75.30%75.30%77.65%3.80$3.092.96M
Grok 4.6medium75.64%75.64%77.18%7.20$1.551.04M
Grok 4.6xhigh75.12%75.12%76.39%9.10$1.581.25M
Solhigh75.80%75.80%78.84%4.30$2.031.64M
GLM 5.3high75.08%75.08%77.79%4.20$1.072.74M
Astralow74.49%74.49%77.55%2.20$3.11993.5K
Solmedium74.71%74.71%78.96%2.30$1.47987K
Opus 5low72.87%72.87%77.44%2.60$2.512.23M
Sonnet 5high73.32%73.32%80.50%7.20$1.613.49M
Sonnet 5xhigh71.46%71.46%77.06%10.30$1.943.85M
Kimi K3low72.02%72.02%78.07%3.60$0.941.49M
Lunahigh71.09%71.09%75.00%3.30$0.131.51M
Gemini 3.7 Flashlow70.33%70.33%79.27%5.40$0.997.84M
Terrahigh70.33%70.33%78.88%2.60$0.951.24M
DeepSeek V4 Prohigh70.14%70.14%77.91%11.60$0.412.21M
Grok 4.6low70.33%70.33%79.71%3.00$0.82672.5K
Terramedium68.58%68.58%80.95%1.30$0.60733.25K
GLM 5.3low66.41%66.41%77.19%1.50$0.591.48M
Sonnet 5medium63.64%63.64%80.22%4.40$1.242.63M
Sollow63.83%63.83%78.85%1.30$1.00583K
Terralow60.60%60.60%80.01%1.10$0.55643K
Lunamedium56.91%56.91%76.18%1.40$0.07769.75K
MiniMax M3high53.30%53.30%72.56%6.50$0.374.17M
Sonnet 5low53.00%53.00%80.73%2.60$0.891.51M
Lunalow40.85%40.85%80.94%1.00$0.06509K

All individual tax results

27 model and reasoning configurations

Download CSV
Individual tax reported aggregate results. First 10 of 27 model and reasoning configurations, sorted by accuracy descending.
#MODEL
1Astraxhigh77.63%76.77%4.10$4.091.46M
2Astrahigh77.53%76.92%3.60$3.501.14M
3Solhigh75.80%78.84%4.30$2.031.64M
4Grok 4.6medium75.64%77.18%7.20$1.551.04M
5Opus 5medium75.30%77.65%3.80$3.092.96M
6Grok 4.6xhigh75.12%76.39%9.10$1.581.25M
7GLM 5.3high75.08%77.79%4.20$1.072.74M
8Solmedium74.71%78.96%2.30$1.47987K
9Astralow74.49%77.55%2.20$3.11993.5K
10Sonnet 5high73.32%80.50%7.20$1.613.49M

Reported aggregate results. Precision is Classic Precision. Time, cost, and tokens are averages per return.

Methodology

Our benchmarks draw on work performed inside Current and Shield, with the context and tools needed to complete each assignment. Models run through the Holdings proprietary evaluation system under consistent conditions. Tax runs use the same client documents, prior-year information, and tax engine version. IT runs replay recorded tool responses in an isolated environment, with write actions mocked and the starting state reset between runs.

The overview shows the Pareto frontier: configurations offering the best score for their selected cost, time, or token usage. Detailed charts and tables include all reported configurations. Tax cost and tokens per return are the source cohort totals divided by 40; time is average extraction time. IT reports cost and elapsed time per task; its export does not include token usage. Completion counts, run identifiers, and turn counts are not reported. Computer-use results and human-cost comparisons remain illustrative. Human-cost scenarios add an assumed practitioner-review allowance to execution cost and compare it with a fixed labor baseline, using the assumptions stated beneath each chart. They do not represent measured savings.

Practitioner corrections become focused evaluations that help us improve prompts, tools, and our evaluation system. Repeating tasks under consistent conditions helps us understand whether those changes improve performance. Results describe a particular system and task set, not every job in an industry. We plan to extend these benchmarks as our systems improve and we build in additional industries, helping practitioners spend more time on what customers value most. Read how we use practitioner feedback to improve products like Tax AI in Building self-improving tax agents with Codex (opens in a new tab).

Changelog

Updated individual tax preparation results from 28 to 40 graded returns per configuration.

Launched two benchmarks: individual tax prep and IT helpdesk ticket resolution.