Thrive Holdings

New York City

  

Measuring AI on real-world work

We evaluate AI models on economically valuable tasks drawn from the businesses Thrive Holdings owns and operates, measuring performance alongside execution time and cost.

Model performance across benchmarks

IT values use the 52-ticket export. Token counts are available for individual tax prep only. Focus a point for its model, effort, score, and cost.

Best result per model + Pareto frontier. Hover or focus a point for details.

Why we built these benchmarks

AI could improve the quality and lower the cost of services people and businesses rely on every day. Evaluating the work behind those services gives us a tangible way to assess AI model progress. Across the 70+ businesses we own and operate, our engineers work alongside practitioners to turn real assignments, decisions, and corrections into tests that reflect professional standards. This helps our businesses identify where AI can create value.

We're releasing benchmarks for individual tax preparation at our accounting platform Current (opens in a new tab) and helpdesk ticket resolution at our IT services platform Shield (opens in a new tab). Drawing on 150 end-to-end tasks from production workflows, they preserve the context needed to measure task performance, execution time, and cost. We use the results to choose models, identify failures, and improve the systems around them. We're also sharing a small selection of tasks so other builders can examine the work and test their own approaches.

Our benchmarks

Individual tax prep v1

Preparing an individual tax return from client documents to a return ready for practitioner review.

Helpdesk ticket resolution v1

Investigating support requests and autonomously resolving issues to help customers get back to work.

Holdings-IndividualTaxBench

Individual tax prep

Federal Form 1040 tax preparation

As an AI tax preparer, models review client documents, identify relevant tax information, and map it to the fields and schedules required for a structured federal Form 1040 return.

Reasoning effort
  • low
  • medium
  • high
  • xhigh
Models
  • Astra
  • Grok 4.6
  • Opus 5
  • Sol
  • GLM 5.3
  • Terra
  • Sonnet 5
  • Kimi K3
  • Gemini 3.7 Flash
  • Luna
  • DeepSeek V4 Pro
  • MiniMax M3
Individual tax results · All complexities
ModelEffortAccuracyRecallPrecisionMinutes per packageCost per return (USD)Tokens per return
Astrahigh79.20%79.20%78.91%3.55$3.471.13M
Grok 4.6medium79.05%79.05%79.69%7.15$1.541.03M
Astraxhigh78.93%78.93%78.25%4.00$4.051.43M
Opus 5medium78.72%78.72%80.34%3.75$3.062.94M
Grok 4.6xhigh77.83%77.83%78.56%9.08$1.571.24M
Astralow77.75%77.75%79.56%2.14$3.09986.79K
Solhigh77.64%77.64%79.26%4.25$2.011.63M
Solmedium76.98%76.98%80.11%2.26$1.45972.5K
GLM 5.3high76.51%76.51%78.57%4.14$1.062.7M
Opus 5low76.41%76.41%79.71%2.62$2.502.22M
Terrahigh74.84%74.84%80.33%2.54$0.941.23M
Sonnet 5xhigh74.29%74.29%78.85%10.10$1.923.81M
Sonnet 5high74.28%74.28%80.94%7.02$1.593.46M
Kimi K3low74.02%74.02%78.92%3.60$0.931.48M
Gemini 3.7 Flashlow73.69%73.69%81.61%5.36$0.987.78M
Lunahigh73.28%73.28%77.36%3.31$0.131.49M
Grok 4.6low73.24%73.24%79.82%3.03$0.82671.33K
DeepSeek V4 Prohigh72.46%72.46%78.98%11.52$0.402.19M
Terramedium72.05%72.05%80.94%1.30$0.60726.08K
Sollow71.62%71.62%80.13%1.29$0.99570.67K
GLM 5.3low71.20%71.20%79.23%1.47$0.591.47M
Sonnet 5medium67.98%67.98%80.80%4.32$1.222.61M
Terralow67.00%67.00%81.07%1.11$0.54638.87K
Lunamedium65.00%65.00%78.32%1.36$0.07764.06K
MiniMax M3high57.77%57.77%79.92%6.40$0.374.13M
Sonnet 5low55.33%55.33%81.75%2.61$0.881.5M
Lunalow52.67%52.67%83.34%0.98$0.05503.79K

Takeaways

Similar accuracy, different economics

Grok 4.6 medium achieves 79.05% accuracy at $1.54 per return, compared with Astra high at 79.20% and $3.47. Grok costs roughly 56% less but takes about twice as long, averaging 7.15 versus 3.55 minutes per return. For this pair, the practical trade-off is lower cost versus faster completion, with little separation in measured accuracy.

Open source does not automatically mean lower cost

Luna at high reasoning achieves 73.28% accuracy at $0.13 per return, while the open-weight DeepSeek V4 Pro model at high reasoning achieves 72.46% at $0.40. Luna delivers higher measured accuracy at roughly one-third the execution cost. In this comparison, choosing the proprietary model reduces cost rather than adding a premium.

More reasoning can introduce regressions

Compared with Astra high, extra-high reasoning scores better on 10 returns, ties on 27, and scores worse on 3. One decline of 18.90 percentage points outweighs the smaller gains elsewhere. Average accuracy falls from 79.20% to 78.93% while cost rises to $4.05 per return. Additional reasoning improves more returns than it hurts, but the size of the regressions determines the overall result. On this cohort, a higher reasoning budget does not produce more reliable performance.

How we built the benchmark

Task selection

All 27 model and reasoning configurations are evaluated on the same 40 returns, with each return weighted equally in the reported averages. The current complexity labels are 19 simple, 12 moderate, 5 complex, and 1 ultra-complex; the remaining 3 returns have no complexity label.

The harness

We use the lightweight Pi harness (opens in a new tab) to evaluate models under consistent conditions, with the same inputs, tools, and task instructions and no model-specific workflow optimizations. Each tax package is handled by one model and one agent. Each run has access to frozen Markdown documents, original Excel workbooks, sandboxed file and Bash tools, and web search. Agents can inspect files, execute commands, write and run scripts, and search the web.

Every configuration receives the same client documents and uses the same tax engine version. Prior-year information and preparation or review notes, when included, are provided as frozen snapshots to keep the task consistent across models.

How we grade

We convert the model output and the reference return exported from our tax engine into canonical Form 1040 JSON, then compare them using Astra at medium reasoning as an LLM-as-judge. Exact matching alone is insufficient because entity names, formatting, and how information is represented can differ between outputs. The judge matches corresponding entities and accounts for these nuances to distinguish equivalent information from substantive discrepancies.

How this compares to our production product

Our production Tax AI product uses specialized, form-specific extractors, prompts, and model selection, with document processing, context, and tools that vary by workflow. These benchmarks more closely reflect raw model capabilities under a shared evaluation setup than the performance of our specialized production product.

Practitioner corrections become focused evaluations that help us improve Tax AI's prompts, tools, and evaluation system. Repeating the same tasks under consistent conditions lets us measure whether those changes improve performance. Read more in Building self-improving tax agents with Codex (opens in a new tab).

All individual tax results

Individual tax reported aggregate results. First 10 of 27 model and reasoning configurations, sorted by accuracy descending.
#MODEL
1Astrahigh79.20%78.91%3.55$3.471.13M12.70
2Grok 4.6medium79.05%79.69%7.15$1.541.03M11.75
3Astraxhigh78.93%78.25%4.00$4.051.43M14.93
4Opus 5medium78.72%80.34%3.75$3.062.94M22.23
5Grok 4.6xhigh77.83%78.56%9.08$1.571.24M13.18
6Astralow77.75%79.56%2.14$3.09986.79K12.58
7Solhigh77.64%79.26%4.25$2.011.63M15.35
8Solmedium76.98%80.11%2.26$1.45972.5K10.88
9GLM 5.3high76.51%78.57%4.14$1.062.7M25.48
10Opus 5low76.41%79.71%2.62$2.502.22M17.50

Package means across the same shared cohort for every model. Costs are in USD per return. LLM-judge scores are not supplied in this release.

Reported aggregate results. Precision is Classic Precision. Time, cost, and tokens are averages per return.

Changelog

Launched v1 of the individual tax preparation and helpdesk ticket resolution benchmarks.