Argo-Bench review: the first public data-agent test that grades business actions

Share





FuturPulse analysis · 4 October 2026

Argo-Bench, released 1 October 2026, is the strongest public test yet for enterprise data agents: its best model fully solved 34.8% of 210 action-scored tasks across a 235-table warehouse containing 7.5 billion rows.

At a glance

  • Argo-Bench simulates a New York City food-delivery company with 81 million orders during 2024.
  • Claude Opus 5.5 led the reported field with a mean score of 59.5 out of 100.
  • Argo-Bench scores filed bans, forecasts, journal entries and payouts, rather than judging SQL text alone.
  • The released warehouse occupies 71 GiB in Parquet form, making local reproduction possible but not lightweight.
  • AgentBench covers eight agent environments, but Argo-Bench concentrates on one large operational data setting.

What does Argo-Bench actually test?

Argo-Bench tests an agent’s full work loop: inspect the warehouse, calculate an answer, then file a business action. The test has 210 tasks across fraud, forecasting, finance, dashboards and commercial analysis. That is a different target from text-to-SQL, where success can mean producing a query whose output matches an answer key.

The authors built a synthetic food-delivery business with 235 connected tables, 3.4 million active customers and 81 million orders. The agent sees warehouse records, but not the simulator’s hidden state. That matters because the grader can check whether a fraud ban prevented the right losses, or whether a forecast matched a month withheld from the agent.

Its practical distinction is the filing. An agent can submit customer bans, payout decisions, general-ledger entries, forecasts or dashboard data sources through a Mission Control interface. The repository says the agent is graded on those filings, not on a persuasive written explanation. A fluent answer therefore cannot compensate for an action that harms legitimate couriers or misses the actual fraud ring.

Why this is harder than “write SQL”: an agent must choose useful tables, interpret business rules, use Python for statistics where needed, and make a decision with consequences. SQL is only one tool in that chain.

How hard are the reported Argo-Bench tasks?

The first reported result is sobering. The paper’s best run, Claude Opus 5.5, achieved a mean score of 59.5 and fully solved 34.8% of tasks. “Solved” means scoring at least 95 out of 100, so it is a high bar rather than merely returning a filing.

The paper reports 95% bootstrap intervals of 52.0 to 67.1 for Opus’s mean score, and 27.3% to 43.3% for its fully solved share. Those wide ranges reflect a benchmark with several task families, not a stable one-number league table. The useful reading is that even the leading system often produced an incomplete or economically wrong action.

One documented forecasting example shows the difference. A reference solution reads six tables and earns 94 points. Repeating April’s $1.44 million value for May misses by $1.41 million and receives minus one point. An agent can therefore perform a familiar calculation and still fail the business task because it ignored a changing pay rule.

The benchmark also measures effort. A run may use up to 500 model turns and one hour of wall-clock time per question. SQL queries can run for five minutes and scan up to 20 GiB each. Those limits make Argo-Bench closer to a bounded analyst workflow than a chatbot question, though they still do not model a company’s full access controls or approval chain.

Which models lead on Argo-Bench?

Claude Opus 5.5 leads the reported set, but the more useful finding is the gap between scores and complete outcomes. GPT-6 Astra and Claude Sonnet 5.5 each average 51.8 points, while their full-solve rates stay below 29%. A buyer should not read a middling mean score as reliable autonomy.

Reported model settingMean task scoreFully solvedAverage model API cost per task
Claude Opus 5.559.5 / 10034.8%$4.71
GPT-6 Astra51.8 / 10027.6%$2.71
Claude Sonnet 5.551.8 / 10028.6%$3.74
GPT-6.1 Sol49.5 / 10024.8%$0.49
GPT-6 Sol36.8 / 10017.6%$0.85

Source: Argo-Bench Table 2. API cost excludes BigQuery warehouse spending.

Higher scores still do not mean dependable business actions: Claude Opus 5.5, GPT-6 Astra, Claude Sonnet 5.5, GPT-6.1 Sol, GPT-6 Sol
Higher scores still do not mean dependable business actions · Source: arxiv.org

The paper separately estimates warehouse-query spending. Across 9,869 graded runs, agents scanned 986 TiB, costing about $6,160 at the stated BigQuery list price, versus $14,502 of model API spending. For the best-scoring model, warehouse scans averaged $0.82 per task. This means deployment cost is not just model tokens, the chunks of text a model reads and bills by.

How does Argo-Bench compare with AgentBench?

AgentBench and Argo-Bench measure different claims. AgentBench, published at ICLR 2024, evaluates general agent reasoning and decisions across eight environments. Its original suite spans operating systems, databases, knowledge graphs, games, browsing, shopping and household-style tasks.

Argo-Bench trades that breadth for depth in one business setting. It asks whether an agent can reconcile an ERP warehouse, apply a policy, then take a measured action. AgentBench’s database environment tests tool-using database interaction; Argo-Bench tests a larger chain that can include SQL, Python analysis, hidden future values and economic consequences.

The comparison should not be collapsed into “which benchmark is harder.” AgentBench answers whether a model transfers between different environments. Argo-Bench answers whether it can survive a particular enterprise-style analytical workflow. AgentBench’s current function-calling version provides containerized versions of five tasks, including database, operating-system and web-shopping tasks, but it does not claim to score the downstream profit, loss or accounting effect of an operational filing.

That makes Argo-Bench more relevant for teams evaluating an analyst agent that could flag fraud, forecast payouts or support a month-end close. It makes AgentBench more useful for teams choosing a general-purpose agent framework. Neither is a direct substitute for an internal evaluation using the company’s schemas, definitions and approval rules.

Which enterprise workflows does Argo-Bench add?

Argo-Bench adds workflows that current general benchmarks mostly leave abstract: fraud triage, courier incentive allocation, back-pay calculations, forecasting after a policy shock, dashboard lineage and financial reconciliation. These are enterprise-scale because several records must agree across order, payment, party and ledger systems.

The warehouse is not merely large in row count. It contains 3,949 columns across 235 tables, including 159 standard Oracle E-Business Suite tables and 76 custom extensions. That design creates a familiar analyst risk: the query may run and return plausible values while joining the wrong business concept.

America’s Next Top Modeler, a separate Theory Ventures challenge, points toward the next missing layer. Its starter kit combines 24 Parquet tables, 19 JSONL logs and PDFs, with 25 practice questions. A post-event account reports a median human-participant score of 23 out of 65 and identifies unstructured documents as a major bottleneck.

Argo-Bench is stronger on action grading than that competition format. The Theory Ventures exercise is broader in data modality. A mature enterprise evaluation should combine both: action consequences from Argo-Bench and messy internal evidence, including documents and logs, that an analyst must turn into usable facts.

What metrics should enterprise buyers add?

Argo-Bench’s outcome score is its best idea, but production teams need diagnostic measures too. Hamel Husain and Shreya Shankar recommend starting with end-to-end task success, then examining tool choice, argument extraction, error handling, context retention and efficiency. Their evaluation guidance also recommends transition-failure matrices that locate the first failed step.

That approach would expose why a model earned 40 points instead of 90. Did it choose the wrong table Did it use the right records but the wrong metric Did it calculate correctly but file without authorization Argo-Bench logs filings, token use, timing and run endings for scored submissions, which gives maintainers useful raw material for this kind of analysis.

For an internal rollout, measure at least four additional things:

  • Approval compliance: whether the agent acted only within the human or policy approval it received.
  • Data-governance compliance: whether it accessed only permitted tables, columns and tenant data.
  • Recovery quality: whether it stops safely after empty results, stale data or tool failures.
  • Cost per accepted action: model, warehouse and human-review cost divided by actions a business owner approves.

Who is making each claim, and why?

The Argo-Bench claims come from Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand and Joseph J Ma. Their paper and TextQLLabs repository have a clear adoption interest: a widely used benchmark validates their task design, harness and leaderboard. That does not invalidate the measurements, but it means enterprises should reproduce relevant tasks and seek external replications.

AgentBench’s maintainers benefit when researchers adopt its broad agent suite and function-calling version. Theory Ventures has a commercial interest as an early-stage venture capital firm that runs internal research and software efforts; its benchmark-adjacent challenge can shape how builders judge agent capability. Its public starter kit nevertheless supplies concrete multimodal scope rather than a marketing claim.

Other names in the search results should not be confused with this benchmark. Argos Myriad sells an annotation, evaluation and review platform with QA and audit trails; it is unrelated to Argo-Bench. HypePaper’s current score of 54.5 out of 100 measures paper momentum, not benchmark validity, and its business is ranking attention around research papers.

What could not verify?

No independent replication of Argo-Bench’s full 14-model result set is yet public. The paper reports the authors’ runs, while external teams would need to submit logs for scoring because the answer keys remain held out. Benchmark maintainers could settle reproducibility questions by publishing independently verified submissions and a public scoring service.

Argo-Bench also does not establish how models behave with real customer data, live permissions, changing warehouse schemas or legally binding financial controls. Enterprises, auditors and deployment vendors could settle those questions only through controlled internal trials with documented human approvals, access boundaries and post-action reviews.

Sources



Maya Chen
Maya Chen
Maya Chen covers AI agents, orchestration frameworks, tool-use, and evaluation. She focuses on what actually works in production—failure modes, safety boundaries, and measurable performance—without the hype.

Read more

Local News