How Far Behind the Frontier are Leading Open Weight Models on Cyber?
英国 AI Security Institute 首次公开分析领先开源权重模型的网络安全能力差距,测试发现 GLM-5.2 是测试时最强的开源权重模型,在 AISI 窄域网络任务上接近 4 个月前发布的 Opus 4.6,在长时程网络靶场上接近 Opus 4.5,整体落后前沿 4 至 7 个月,窄于 2025 年 1 至 9 月内部评估测得的 6 至 10 个月。