
HarnessEval-W: Agentifying the Evaluation of Visual Worlds — A HarnessEval Agent Evaluation Guide
HarnessEval-W is an agentified evaluation benchmark that brings the “harness” paradigm from the LLM ecosystem to world-model benchmarking. Instead of computing a fixed rubric over generated rollouts, it interprets each evaluation case, decomposes the question into measurable sub-questions, and spawns specialized sub-agents with tailored context and diagnostic tools. A parent agent then validates the gathered evidence and aggregates it into a final verdict, producing a transparent evidence tree with a complete reasoning chain for every single rollout. ...