FinanceBench is an agent benchmark built from expert-authored, multi-step tasks grounded in real-world financial documents — 10-Ks, 10-Qs, spreadsheets, and deal materials. Each task requires long-context reasoning, tool use, and structured computation in a sandboxed notebook environment, with outputs scored by deterministic test suites and rubric-based LLM checks.
| # | Model | Score |
|---|
Score — total verifier checks passed across all tasks, as a percentage of all verifier checks run. Each model was evaluated on a single trajectory per task.