Deployment-Aware and Reliable Evaluation of Models as Agents

A workload- and deployment-aware benchmark for evaluating models as agents.

  • Yu Liu,
  • Zhilin Liu,
  • Zhiwei Yang,
  • Shaojie Zhang,
  • Zheyuan Deng,
  • Tingwei Huang,
  • Zhenbo Luo,
  • Lei Jiang,
  • Yanbing Liu,
  • Pei Fu
  1. Institute of Information Engineering, Chinese Academy of Sciences
  2. School of Cyber Security, University of Chinese Academy of Sciences
  3. MiLM Plus, Xiaomi Inc.
  4. Department of Computer Science, Brown University

Cartoon lobster evaluator holding a magnifying glass and a clipboard

    Motivation

    What deployment-oriented agent evaluation must answer

      TL;DR

        Read the abstract

        Overview

        From source benchmarks to deployment reports

          Workload matrix

          Tasks organized by what the agent has to do

          Instead of reporting by source benchmark, DAREBench groups tasks along two workload dimensions: input modality (text or multimodal) and execution form (single-step, multi-step, or multi-step with tools). Each source benchmark maps to exactly one group. Select a cell to see its sources, scoring modes, an example prompt and the strongest models.

          Source benchmarks

          Browse every task file

          Metadata only (from the task files' frontmatter). Group and scoring follow the paper's source mapping; prompts, answers and rubrics are not shown.

          Leaderboard

          No single model dominates all workload groups

          Accuracy & cost

          Text and multimodal workloads have different frontiers

          Scopes every chart and note in this section.

          Tokens and dollars rank models differently

          Evidence-based audit

          Removing credit the evidence does not support

          The audit funnel

            Two judges, two model families

            Hallucination patterns caught by the audit

              Key findings

              What the results say about choosing an agent model

              Model selection should consider workload profiles, deployment mode, and accuracy–cost trade-offs rather than a single aggregate score.

              Where agents fail

              Quick start

              Evaluate your own model

              The commands below are taken verbatim from the repository README. The repository holds the task files, graders and run scripts; by default, runs go through an OpenClaw agent runtime, as in the paper.

              Requirements

                  Anatomy of a task file

                    
                            

                    Cite

                    If DAREBench helps your work

                    Please cite the paper: