Announcement_82
New EvalEval blog post: AI evals are becoming the new compute bottleneck, with Yifan Mai, Georgia Channing, and Leshem Choshen. We dig into how benchmarking frontier systems now routinely costs tens of thousands of dollars per run, why agent evals are especially unpredictable, and what that concentration of validation authority means for the broader research community.