
AI Agent Benchmark Exploitation: How Berkeley RDI Broke Every Major Benchmark
Introduction — The Benchmark Illusion For the last two years, the AI industry has been racing to top leaderboards on agent benchmarks like SWE-bench, WebArena, and GAIA. These scores drive funding rounds, product launches, and enterprise procurement decisions. But a landmark study from Berkeley RDI reveals a devastating truth: every single major AI agent benchmark can be exploited to produce near-perfect scores without the agent solving a single real task. The study, led by Hao Wang, Qiuyang Mang, Alvin Cheung, Koushik Sen, and Dawn Song, introduces BenchJack — an automated vulnerability scanner that achieved 100% scores on eight benchmarks using nothing more than environment manipulation, configuration leakage, and broken validation logic. The paper, titled “Do Androids Dream of Breaking the Game?” (arXiv 2605.12673), demonstrates that the current state of AI agent evaluation is fundamentally broken, and the problem is far more urgent than most in the industry realize. ...