Benchmark report · Updated September 2, 2026
Financial Agent Benchmark
Financial answers should survive inspection. On the Vals Finance Agent public set, Drillr MCP scores 94.06%, against 76.30% for Claude Opus 4.6 with built-in web search. In a controlled run over 200 questions, the same model scored 80.76% with Drillr MCP and 33.86% with web search, using 55% fewer tool calls.
01 · Why Drillr MCP
Search finds pages. Financial research needs facts.
- Where web search breaks
- Semantic search returns passages that look like the question, stripped of the table, period and filing they came from. That loss of context is the root problem.
- What Drillr MCP does
- Drillr parses filings into a knowledge graph: entities, periods, line items, reported and restated values, each tied to its source. One call returns the fact with its citation.
02 · Public benchmark
Public financial AI agent results.
The Drillr MCP and Web Search rows were run with Claude Opus 4.6 on the Finance Agent v1.1 public test set, a snapshot taken at the time of writing. Rows marked Vals baseline quote the official results Vals published from its internal test set. The live Vals leaderboard may differ.
03 · Internal benchmark
Ahead on every question type, with half the tool calls.
| Accuracy | Efficiency | |||||||
|---|---|---|---|---|---|---|---|---|
| Rank | Configuration | Rate | Full | Partial | Zero | Total calls | Avg. calls | Over 20 calls |
| 01 | Sonnet 5 + Drillr MCPCurrent | 80.76% | 138 | 35 | 20 | 1,333 | 6.91 | 7 |
| 02 | Sonnet 5 + Web SearchBaseline | 33.86% | 43 | 76 | 74 | 2,970 | 15.39 | 70 |
Both runs used Sonnet 5 on the same 200 questions, one with Drillr MCP and the control with general web search. Answers were scored against a fixed rubric and every tool call was counted.
Category breakdown
Where each retrieval approach wins.
| Task | Web Search | Drillr MCP |
|---|---|---|
| EPS numerator | 27.14% | 94.29% |
| Cross-entity breadth | 78.33% | 100.00% |
| Company Search | 26.54% | 72.62% |
| Event base rate | 0.00% | 66.67% |
| Footnote nonstandard | 25.00% | 75.00% |
| Keyword collision | 67.29% | 96.91% |
| Cross-family revision | 0.00% | 100.00% |
| Restatement | 26.67% | 84.91% |
| Other public benchmarks | 31.47% | 55.66% |
Other public benchmarks pools the questions drawn from Finance Agent v2, Daloopa, FinSearchComp and BigFinanceBench.
04 · Cases
The score gap, question by question.
What net income figure did Claros Mortgage Trust use as the numerator in its fiscal year 2022 basic EPS calculation?
$112,064 thousand
Stopped at net income attributable to common stockholders—the intermediate value before the two-class allocation.
27 tool calls$109,667 thousand
Found the EPS reconciliation, subtracted $2,397 thousand for participating securities, and tied the result to $0.79 EPS.
3 tool callsWhat net income attributable to parent did Pathward Financial report for the three months ended June 30, 2023?
$45,096 thousand
Returned the superseded figure from the original 10-Q and press release without finding the correcting 10-K/A.
22 tool calls$36,080 thousand · restated
Distinguished the originally reported $45,096 thousand from the corrected figure and cited the amending filing.
4 tool callsIn Intel’s fiscal 2025 statement of stockholders’ equity, what aggregate shares and net proceeds appear on the “Net proceeds from stock issuances and warrants” line?
≈735.3 million shares · ≈$12.706 billion
Reconstructed an estimate from transaction announcements and incorrectly included separately classified escrowed shares.
21 tool calls580 million shares · $11.835 billion
Retrieved the exact roll-forward line instead of combining similarly worded figures from separate disclosures.
5 tool callsThe three cases cover three failure modes of the web-search run: selecting an intermediate value, missing a later restatement, and combining similar figures from separate sources.
05 · Methodology and limitations
How the results were produced, and what they do not establish.
Every question is answered from Drillr’s own datasets: the same tables, filings and derived metrics the product serves. How those datasets are sourced, normalized and refreshed is documented separately.
Numeric answers allow a 3% relative error unless a rubric sets another tolerance. Company Search uses deterministic entity recall.
An LLM judge scores each rubric item. When two passes disagree, a third pass decides by majority. Full, partial and zero outcomes are kept separate so the aggregate does not hide the shape of errors.
The public Vals snapshot and the internal Drillr benchmark use different protocols. Scores from the two should not be compared directly.
Read the Vals methodologyDrillr created and evaluated the internal benchmark. The results are published as product research, not as an independent audit or third-party certification.
Some task groups hold only a few questions. A single answer can move a small category’s percentage sharply.
Model behavior, source availability, and retrieval systems change. This page identifies the August 14, 2026 Drillr MCP run so later results can be separated.
Test the workflow