← Back to the briefs
Research

EvalResearchBench tests whether agents can design useful evaluations

A new preprint compares agent-built evaluation suites with established benchmarks and finds that matching visible tests does not guarantee matching hidden ones.

A preprint submitted on 3 October introduces EvalResearchBench. Yaolun Zhang and colleagues ask agents to design executable evaluators under fixed budgets.

It compares nine researcher configurations, 13 candidate models and 14 benchmarks across coding, workplace tasks and reasoning. We have not reproduced the experiments.

Agreement with a reference is the test

The authors report that the best evaluators match roughly 75% of reference orderings between model pairs. That measures ordering, not task completion. Disagreements between references set an estimated ceiling near 91% in this experiment.

The leading evaluator on visible targets does not lead on hidden targets. A human-designed task sample remains competitive.

Reported defects include truncated answers, exhausted budgets and scores dominated by a few questions. The results describe this experimental setup, rather than universal evaluator reliability.

Our assessment

If you ask an agent to build your evaluation, treat the evaluator as software that needs checking. Start with cases whose expected outcome you already know. Include correct answers, plausible mistakes and incomplete responses, then inspect how the grader handles each one.

Keep some cases hidden during development. For example, you could let an agent revise tests against one set of document-processing tasks, then assess its finished evaluator on separate documents. Define the scoring rules before showing it that second set. This is an illustrative test plan, not a result from the paper.

Compare the proposed evaluator with a simple baseline that a person can inspect. Record disagreements and examine whether a score changed because of a better answer, different task coverage or a grader defect. A useful ranking alone may not tell you which important errors were missed.

Check the operating limits too. Set explicit time and call budgets, retain failed runs, and look for unanswered tasks before accepting an average score. Budget exhaustion should remain visible in the review, so a shorter or cheaper evaluator does not appear successful merely because it assessed less work.

Source: EvalResearchBench, arXiv preprint, version 1.

Original source

Read the original material

arXiv (opens in a new tab)

The article date belongs to Agentic Horizon. The source publication date is listed separately. Vendor claims remain attributed; we have not independently tested the reported capability.

Search the publication

Find a story by title, topic, or keyword.

Press Escape to close