Benchmark agents on enterprise finance.
From 10-Ks to deal models.

FinanceBench is an agent benchmark built from expert-authored, multi-step tasks grounded in real-world financial documents — 10-Ks, 10-Qs, spreadsheets, and deal materials. Each task requires long-context reasoning, tool use, and structured computation in a sandboxed notebook environment, with outputs scored by deterministic test suites and rubric-based LLM checks.

Multi-step reasoning
Real document grounding
Deterministic + LLM scoring
Tool-use in sandboxed env
Expert-authored tasks
Leaderboard
#ModelScore
Model Comparison — All Tasks

Scoretotal verifier checks passed across all tasks, as a percentage of all verifier checks run. Each model was evaluated on a single trajectory per task.

Select Task
~/workspace
Select a file
Click a file to view its contents

Scoreverifier checks passed for this task, as a percentage of total verifier checks run. Based on a single trajectory per model.

Task Lifecycle
The end-to-end pipeline for creating, validating, and generating evaluation tasks — from raw data procurement through final deliverable review.
Manual Work
QC Work
Automated Work
Procurement
Sourcers
QA
Domain Experts
Expert 1
Expert 2
Expert 3
Reviewers
Reviewer 1
Reviewer 2
1
Procurement of Data
Sourcers
Source Repository + Operational Artifacts
Repo + Artifacts QA
QA
Expert 1
2
Task Creation
if pass@k ≤ 50%
3
Task Iteration & QC
Proceed when verifiers PASS against golden data
4
Verifier Iteration & QC
Proceed when verifiers PASS against ALL golden data
5
Trajectory Generation + QC
Shared Filesystem
Docs MCP.docx
Spreadsheet.xlsx
PDF MCP.pdf
Presentation.ppt
Calendar.ical
Mail MCP.mbox
Chat MCP.json
Filesystem MCP
Container Runtime
Dockerfile + Repo
Tool
Execution