Thrive Holdings

New York City

  

Measuring AI on real-world work

We evaluate AI models on economically valuable tasks drawn from the businesses Thrive Holdings owns and operates, measuring performance alongside execution time and cost.

Model performance across benchmarks

IT values use the 52-ticket export. Token counts are available for individual tax prep only. Focus a point for its model, effort, score, and cost.

Best result per model + Pareto frontier. Hover or focus a point for details.

Why we built these benchmarks

AI could improve the quality and lower the cost of services people and businesses rely on every day. Evaluating the work behind those services gives us a tangible way to assess AI model progress. Across the 70+ businesses we own and operate, our engineers work alongside practitioners to turn real assignments, decisions, and corrections into tests that reflect professional standards. This helps our businesses identify where AI can create value.

We're releasing benchmarks for individual tax preparation at our accounting platform Current (opens in a new tab) and helpdesk ticket resolution at our IT services platform Shield (opens in a new tab). Drawing on 150 end-to-end tasks from production workflows, they preserve the context needed to measure task performance, execution time, and cost. We use the results to choose models, identify failures, and improve the systems around them. We're also sharing a small selection of tasks so other builders can examine the work and test their own approaches.

Our benchmarks

Individual tax prep v1

Preparing an individual tax return from client documents to a return ready for practitioner review.

Helpdesk ticket resolution v1

Investigating support requests and autonomously resolving issues to help customers get back to work.

Holdings-IndividualTaxBench

Individual tax prep

Federal Form 1040 tax preparation

As an AI tax preparer, models review client documents, identify relevant tax information, and map it to the fields and schedules required for a structured federal Form 1040 return.

Reasoning effort
  • low
  • medium
  • high
  • xhigh
Models
  • Astra
  • Grok 4.6
  • Opus 5
  • Sol
  • GLM 5.3
  • Terra
  • Sonnet 5
  • Kimi K3
  • Gemini 3.7 Flash
  • Luna
  • DeepSeek V4 Pro
  • MiniMax M3
Individual tax results · All complexities
ModelEffortAccuracyRecallPrecisionMinutes per packageCost per return (USD)Tokens per return
Astrahigh79.20%79.20%78.91%3.55$3.471.13M
Grok 4.6medium79.05%79.05%79.69%7.15$1.541.03M
Astraxhigh78.93%78.93%78.25%4.00$4.051.43M
Opus 5medium78.72%78.72%80.34%3.75$3.062.94M
Grok 4.6xhigh77.83%77.83%78.56%9.08$1.571.24M
Astralow77.75%77.75%79.56%2.14$3.09986.79K
Solhigh77.64%77.64%79.26%4.25$2.011.63M
Solmedium76.98%76.98%80.11%2.26$1.45972.5K
GLM 5.3high76.51%76.51%78.57%4.14$1.062.7M
Opus 5low76.41%76.41%79.71%2.62$2.502.22M
Terrahigh74.84%74.84%80.33%2.54$0.941.23M
Sonnet 5xhigh74.29%74.29%78.85%10.10$1.923.81M
Sonnet 5high74.28%74.28%80.94%7.02$1.593.46M
Kimi K3low74.02%74.02%78.92%3.60$0.931.48M
Gemini 3.7 Flashlow73.69%73.69%81.61%5.36$0.987.78M
Lunahigh73.28%73.28%77.36%3.31$0.131.49M
Grok 4.6low73.24%73.24%79.82%3.03$0.82671.33K
DeepSeek V4 Prohigh72.46%72.46%78.98%11.52$0.402.19M
Terramedium72.05%72.05%80.94%1.30$0.60726.08K
Sollow71.62%71.62%80.13%1.29$0.99570.67K
GLM 5.3low71.20%71.20%79.23%1.47$0.591.47M
Sonnet 5medium67.98%67.98%80.80%4.32$1.222.61M
Terralow67.00%67.00%81.07%1.11$0.54638.87K
Lunamedium65.00%65.00%78.32%1.36$0.07764.06K
MiniMax M3high57.77%57.77%79.92%6.40$0.374.13M
Sonnet 5low55.33%55.33%81.75%2.61$0.881.5M
Lunalow52.67%52.67%83.34%0.98$0.05503.79K

Takeaways

Similar accuracy, different economics

Grok 4.6 medium achieves 79.05% accuracy at $1.54 per return, compared with Astra high at 79.20% and $3.47. Grok costs roughly 56% less but takes about twice as long, averaging 7.15 versus 3.55 minutes per return. For this pair, the practical trade-off is lower cost versus faster completion, with little separation in measured accuracy.

Open source does not automatically mean lower cost

Luna at high reasoning achieves 73.28% accuracy at $0.13 per return, while the open-weight DeepSeek V4 Pro model at high reasoning achieves 72.46% at $0.40. Luna delivers higher measured accuracy at roughly one-third the execution cost. In this comparison, choosing the proprietary model reduces cost rather than adding a premium.

More reasoning can introduce regressions

Compared with Astra high, extra-high reasoning scores better on 10 returns, ties on 27, and scores worse on 3. One decline of 18.90 percentage points outweighs the smaller gains elsewhere. Average accuracy falls from 79.20% to 78.93% while cost rises to $4.05 per return. Additional reasoning improves more returns than it hurts, but the size of the regressions determines the overall result. On this cohort, a higher reasoning budget does not produce more reliable performance.

Methodology

Task overview

The task is to reason over a package of PDFs, spreadsheets, text, and email and produce a structured individual tax return. Inputs range from W-2s, 1099s, and K-1s to brokerage statements, accountant emails, and handwritten or scanned notes. We parse documents into Markdown so the benchmark compares tax reasoning consistently, including for models without vision capabilities. Each agent fills a JSON schema containing hundreds to more than a thousand possible fields for Form 1040 and its supporting schedules, ready to map into tax software.

Some values map directly from a document. Others require calculations and reconciliation across the package—for example, combining business income and expenses for Schedule C or organizing rental activity and separately extracted K-1 information for Schedule E. The model must decide which fields apply while keeping entities, values, and dependencies consistent across schedules.

A knowledge base of PDFs, spreadsheets, text, and email goes into one agent session, which produces structured Form 1040 JSON.
  1. Knowledge base

    PDFs, spreadsheets, text, email

    Parsed into Markdown

  2. Agent

    One session, end to end

    Reasons over the full package

  3. Structured output

    Form 1040 JSON

    Hundreds to 1,000+ possible fields

Complexity

Complexity reflects how difficult a package is to prepare, from familiar W-2 and 1099 mappings to consolidated statements, mixed sources, and schedules that require cross-document calculations. Document volume alone does not capture that difficulty: a large stack of similar K-1s can be easier than a few ambiguous workpapers. These 40 packages retain their original labels as an approximate measure of task difficulty.

The benchmark groups tax packages into simple, moderate, complex, and ultra-complex tiers, with some packages left unclassified.
  1. Simple19Mostly W-2s and 1099s with direct mappings
  2. Moderate12Consolidated statements and mixed sources, with more to reconcile
  3. Complex5More schedules, K-1s, and cross-document calculations
  4. Ultra-complex1The long tail: many schedules stacked on a large packet

3 packages are unlabeled.

Task mining and curation

These tasks are curated from production work across the accounting firms we own. The source material includes client documents, prior-year returns, preparer and reviewer interactions, initial AI passes, and the corrections behind practitioner-reviewed outcomes.

We check that each package's supplied inputs support its expected answers, then freeze those inputs and keep the reference return separate. All 27 model and reasoning configurations are evaluated on the same 40 packages, with each package weighted equally.

Harness and runtime

We use the lightweight Pi harness (opens in a new tab) to give one agent the full preparation task in a single session. Every configuration receives the same prompt, schemas, frozen client documents, tax-engine version, and sandboxed file and Bash tools. Agents can inspect files, run scripts, and search the web, but they do not use sub-agents or our specialized production extractors. This isolates the model and reasoning effort within a consistent, minimal harness.

Grading

We convert both the model output and reference return into canonical Form 1040 JSON. Because entities can appear in different places and names or formatting can vary, Astra at medium reasoning acts as an LLM judge to reconcile corresponding schedules and values. The grader awards partial credit for individual facts rather than passing or failing an entire return. Accuracy therefore measures how much expected information was captured correctly; it is not a whole-return pass rate.

Relationship to the production product

Our production Tax AI product uses specialized classification, splitting, extraction, prompting, and model-selection workflows. This benchmark measures raw model capabilities under a shared evaluation setup, not the performance of that optimized production system. Practitioner corrections become focused evaluations that help improve the production prompts, tools, and grading system. Read more in Building self-improving tax agents with Codex (opens in a new tab).

All individual tax results

Individual tax reported aggregate results. First 10 of 27 model and reasoning configurations, sorted by accuracy descending.
#MODEL
1Astrahigh79.20%78.91%3.55$3.471.13M12.70
2Grok 4.6medium79.05%79.69%7.15$1.541.03M11.75
3Astraxhigh78.93%78.25%4.00$4.051.43M14.93
4Opus 5medium78.72%80.34%3.75$3.062.94M22.23
5Grok 4.6xhigh77.83%78.56%9.08$1.571.24M13.18
6Astralow77.75%79.56%2.14$3.09986.79K12.58
7Solhigh77.64%79.26%4.25$2.011.63M15.35
8Solmedium76.98%80.11%2.26$1.45972.5K10.88
9GLM 5.3high76.51%78.57%4.14$1.062.7M25.48
10Opus 5low76.41%79.71%2.62$2.502.22M17.50

Package means across the same shared cohort for every model. Costs are in USD per return. LLM-judge scores are not supplied in this release.

Reported aggregate results. Precision is Classic Precision. Time, cost, and tokens are averages per return.

Changelog

Launched v1 of the individual tax preparation and helpdesk ticket resolution benchmarks.