Thrive Holdings

New York City

  

Measuring AI on real-world work

We evaluate AI models on economically valuable tasks drawn from the businesses Thrive Holdings owns and operates, measuring performance alongside execution time and cost.

Model performance across benchmarks

IT values use the 52-ticket export. CUA v1 and the tax overview remain illustrative. Focus a point for its model, effort, score, and cost.

Reasoning effort
  • low
  • medium
  • high
  • xhigh

Why we built these benchmarks

AI could improve the quality and lower the cost of services people and businesses rely on every day. Evaluating the work behind those services gives us a tangible way to assess AI model progress. Across the 70+ businesses we own and operate, our engineers work alongside practitioners to turn real assignments, decisions, and corrections into tests that reflect professional standards. This helps our businesses identify where AI can create value.

We're releasing benchmarks for individual tax preparation at Current (opens in a new tab) and IT helpdesk ticket resolution at Shield (opens in a new tab). Drawing on 100 tasks from production workflows, they preserve the context needed to measure task performance, execution time, and cost. We use the results to choose models, identify failures, and improve the systems around them. We're also sharing a small selection of tasks so other builders can examine the work and test their own approaches.

Our benchmarks

Individual tax prep

Preparing an individual tax return from client documents to a return ready for practitioner review.

IT helpdesk ticket resolution

Investigating support requests and autonomously resolving issues to help customers get back to work.

Holdings-IndTaxBench

Individual tax prep

Form 1040 tax preparation

Across 50 tax returns, models work from client documents provided in markdown, identify the relevant tax information, and map it to the fields and schedules required by our tax engine. The task evaluates how well models turn source material into a structured Form 1040 return for practitioner review.

Reasoning effort
  • low
  • medium
  • high
  • xhigh
Pareto frontier · highest accuracy for a given cost
Models
  • Astra
  • Sol
  • Terra
  • Kimi K3
  • Gemini Flash
  • Grok
  • Opus
  • GLM
  • Sonnet
  • Luna
  • DeepSeek
All individual tax model and reasoning effort results
ModelEffortAccuracyRecallPrecisionMinutes per packageExtraction USD per return
Astraxhigh76.69%76.69%75.65%3.65$3.16
Astrahigh76.67%76.67%76.13%3.07$2.69
Astramedium76.51%76.51%75.56%2.08$2.36
Solmedium75.99%75.99%77.66%1.97$1.10
Terraxhigh75.47%75.47%76.85%4.42$0.92
Kimi K3medium75.40%75.40%77.07%6.52$0.97
Terrahigh75.33%75.33%76.86%2.28$0.65
Solxhigh74.82%74.82%76.17%5.45$1.76
Gemini Flashhigh74.78%74.78%76.77%7.18$0.92
Grokmedium74.67%74.67%75.63%6.30$0.92
Kimi K3high74.62%74.62%76.37%7.88$0.99
Opusxhigh74.52%74.52%75.88%5.43$3.07
Astralow74.41%74.41%75.80%1.75$2.28
Opushigh74.28%74.28%76.00%4.05$2.62
Sollow74.19%74.19%77.27%1.20$0.83
Opusmedium74.07%74.07%75.76%3.08$2.33
Grokxhigh74.06%74.06%74.77%8.43$1.11
GLMhigh73.75%73.75%75.17%3.48$0.79
Sonnethigh73.73%73.73%78.05%5.93$1.24
Lunaxhigh73.20%73.20%74.76%4.87$0.10
Opuslow72.94%72.94%74.88%2.10$1.81
Terramedium72.79%72.79%79.26%1.18$0.46
Sonnetxhigh72.33%72.33%75.06%8.62$1.48
Lunahigh71.69%71.69%73.66%2.55$0.07
Kimi K3low71.47%71.47%75.42%3.30$0.71
Sonnetmedium69.29%69.29%78.63%3.55$0.94
DeepSeekhigh68.16%68.16%74.74%10.77$0.34
Lunamedium67.89%67.89%73.76%1.15$0.05
Terralow67.24%67.24%78.65%1.03$0.43

All individual tax results

31 model and reasoning configurations

Download CSV
Individual tax reported aggregate results. First 10 of 31 model and reasoning configurations, sorted by accuracy descending.
#MODEL
1Astraxhigh76.69%75.65%3.65$3.16
2Astrahigh76.67%76.13%3.07$2.69
3Astramedium76.51%75.56%2.08$2.36
4Solmedium75.99%77.66%1.97$1.10
5Terraxhigh75.47%76.85%4.42$0.92
6Kimi K3medium75.40%77.07%6.52$0.97
7Terrahigh75.33%76.86%2.28$0.65
8Solxhigh74.82%76.17%5.45$1.76
9Gemini Flashhigh74.78%76.77%7.18$0.92
10Grokmedium74.67%75.63%6.30$0.92

Reported aggregate results. Precision is Classic Precision. Time is average minutes per return; cost is USD per return.

Methodology

Our benchmarks draw on work performed inside Current and Shield, with the context and tools needed to complete each assignment. Models run through the Holdings proprietary evaluation system under consistent conditions. Tax runs use the same client documents, prior-year information, and tax engine version. IT runs replay recorded tool responses in an isolated environment, with write actions mocked and the starting state reset between runs.

The overview shows the Pareto frontier: configurations offering the best score for their cost. Detailed charts and tables include all reported configurations. Tax cost per return is the source cohort total divided by 28; time is average extraction time. IT reports cost and elapsed time per task. Completion counts, run identifiers, and turn counts are not reported. Computer-use results and human-cost comparisons remain illustrative. Human-cost scenarios add an assumed practitioner-review allowance to execution cost and compare it with a fixed labor baseline, using the assumptions stated beneath each chart. They do not represent measured savings.

Practitioner corrections become focused evaluations that help us improve prompts, tools, and our evaluation system. Repeating tasks under consistent conditions helps us understand whether those changes improve performance. Results describe a particular system and task set, not every job in an industry. We plan to extend these benchmarks as our systems improve and we build in additional industries, helping practitioners spend more time on what customers value most. Read how we use practitioner feedback to improve products like Tax AI in Building self-improving tax agents with Codex (opens in a new tab).

Changelog

Launched two benchmarks: individual tax prep and IT helpdesk ticket resolution.