Deployment-Aware and Reliable Evaluation of Models as Agents
A workload- and deployment-aware benchmark for evaluating models as agents.
- Institute of Information Engineering, Chinese Academy of Sciences
- School of Cyber Security, University of Chinese Academy of Sciences
- MiLM Plus, Xiaomi Inc.
- Department of Computer Science, Brown University
Motivation
What deployment-oriented agent evaluation must answer
TL;DR
Read the abstract
Overview
From source benchmarks to deployment reports
Workload matrix
Tasks organized by what the agent has to do
Instead of reporting by source benchmark, DAREBench groups tasks along two workload dimensions: input modality (text or multimodal) and execution form (single-step, multi-step, or multi-step with tools). Each source benchmark maps to exactly one group. Select a cell to see its sources, scoring modes, an example prompt and the strongest models.
Source benchmarks
Browse every task file
Metadata only (from the task files' frontmatter). Group and scoring follow the paper's source mapping; prompts, answers and rubrics are not shown.
Leaderboard
No single model dominates all workload groups
Accuracy & cost
Text and multimodal workloads have different frontiers
Tokens and dollars rank models differently
Evidence-based audit
Removing credit the evidence does not support
The audit funnel
Two judges, two model families
Hallucination patterns caught by the audit
Key findings
What the results say about choosing an agent model
Model selection should consider workload profiles, deployment mode, and accuracy–cost trade-offs rather than a single aggregate score.
Where agents fail
Quick start
Evaluate your own model
The commands below are taken verbatim from the repository README. The repository holds the task files, graders and run scripts; by default, runs go through an OpenClaw agent runtime, as in the paper.
Requirements
Anatomy of a task file
Cite
If DAREBench helps your work
Please cite the paper: