Skip to content
The DecoderJonathan Kemper

Google researchers find a way to keep self-improving AI agents from memorizing their tests

Google 研究人员提出 RRSI,通过约束智能体运行框架的自优化,减少测试任务过拟合并改善未见任务表现。研究在模型 Claude Opus 4.8 保持冻结的情况下测试了 8 个基准,训练任务提升最高 14.1 点,5 个未见基准提升最高 4.7 点,运行时 token 用量比未正则化版本减少约 30%。

Read the original ↗Export Markdown