Evals for Agents: When Benchmarks Lie and When They Work
A new leaderboard can crown an autonomous agent benchmark winner while leaving deployment questions unanswered. Consider a hypothetical rollout: the agent loops for twelve minutes editing irrelevant files, damages a