A preprint by Liu Zhang and Mark Esposito examines reliability ratings for a deployed voice-and-video interviewing agent. Submitted on 4 October, it uses 2,611 scored interviews, overlapping reviewers, deployment records and support tickets.
This is an observational study of one deployment; we have not independently checked its data.
Reviewers can alter the apparent trend
The authors report that two reviewers assessing the same batches differed by 0.79 standard deviations on a combined score. Changes in reviewer composition also affected the apparent trend.
After adjusting for reviewer differences and smoothing batch estimates, the study finds higher evaluated reliability between March and August. The authors also observe shifts around deployments and a related movement in support tickets.
Those associations do not establish that evaluation caused the improvement. The paper calls for comparable raters, checks on how current the evidence remains, and comparison with operational outcomes.
Our assessment
For your own evaluation programme, keep a small set of reference cases that reviewers can assess independently. Discuss disagreements against a shared rubric, and retain some overlap when people join or leave the review group. This gives you a way to notice changes in scoring practice before interpreting a trend as a product improvement.
Record the system version with each result. If you change the model, tools, prompts or task selection, mark that change in the evaluation record. Compare similar cases before drawing conclusions about whether the agent has become more reliable. These are suggested controls, not a replication of the study.
Add an outcome that reflects the task, such as corrections required after a test workflow. Examine that measure alongside the ratings. If one improves and the other worsens, inspect the underlying cases before reporting progress. Avoid presenting an association as proof that a particular evaluation practice caused the change.
Keep the study’s scope in view. Findings from an interviewing agent do not establish reliability for browser automation, coding or document processing, and reliability ratings alone do not establish fairness in hiring. Your application needs evidence for its own tasks and decisions, refreshed as the system changes.
Source: Reliability of AI Agents, arXiv preprint, version 1.
Read the original material
arXiv (opens in a new tab)The article date belongs to Agentic Horizon. The source publication date is listed separately. Vendor claims remain attributed; we have not independently tested the reported capability.


