核心摘要
AI 摘要:In a study covering seven benchmarks, the UK's AI Security Institute shows that standard AI evaluations systematically underestimate agent capabilities by capping the compute budge
为什么重要
这条信息可能影响用户对 AI 产品、公司动态或行业趋势的判断,值得结合后续进展继续观察。
关键信息
- 来源:The Decoder
- 分类:news
- 标签:AI News、Models、Research
- 原文链接:https://the-decoder.com/uks-ai-security-institute-finds-standard-benchmarks-systematically-underestimate-what-ai-agents-can-actually-do/