GPTProto

评测基准

125 picksNewest firstLatest article Oct 2, 2026, 8:17 AM
4 stories
3 stories
3 stories
08:00 OpenRouter:Announcements(RSS)AI score 62/100

How to Test Tool-Calling Accuracy in AI Agents

An agent can fail in two places when it uses a tool. It can choose the wrong tool, or it can choose the right one and send the wrong arguments. Those failures tell you different things. If an agent calls lookup_order instead of refund_order, the problem is tool selection. If it calls refund_order with the wrong order_id, it chose the right tool and passed the wrong arguments. This guide covers three ways to test tool-calling behavior. The first is a reference-free large language model (LLM) judg

00:05 X:Arena (@arena)AI score 67/100

LLM judges prefer their own responses 70% more often than humans do

Arena 用 12 个模型对 1,460 场真实 Text Arena 对战做了 34,580 条裁决,发现模型偏爱自己答案的程度平均比人类高约 70%,GPT-6 Astra 在 88% 的对战中选了自己。OpenAI 三款裁判对 OpenAI 模型比人类宽容 37 分,AI 裁判之间一致率 79.4%,但与人类投票者只有 56.9%;模型很少判平局,Sol 在 96% 的对战中强行选出赢家。

5 stories
00:00 Artificial Analysis 完整文章(网页)AI score 69/100

GPT-6.1 Sol replaces GPT-6 Sol after just 7 days, with near-Astra intelligence

See model page GPT-6.1 Sol replaces GPT-6 Sol after just 7 days. It scores 1 point below GPT-6 Astra in the Intelligence Index at less than one quarter of the Cost per Task. Pricing matches GPT-6 Sol at $2/$10 per million input/output tokens, except that the cache read discount rises from 90% to 95%. GPT-6.1 Sol’s overall blended price for agentic workloads is therefore slightly lower than GPT-6 Sol. This represents an additional price cut, following GPT-6 Sol’s original 50% discount from GPT-5.

00:00 Artificial Analysis 完整文章(网页)AI score 64/100

AA-AgentPerf-Local: Benchmarking local AI agents on laptops and workstations

Announcing AA-AgentPerf-Local, our open-source inference testing tool for local AI models - test how fast agentic AI can run on your own laptop or workstation, and browse our list of serving configurations to plan your next agent setup Key points: ➤ We’re open sourcing AA-AgentPerf-Local, which replays real agent trajectories on laptop & workstation hardware to test inference performance ➤ We’re releasing initial results for NVIDIA DGX Spark, NVIDIA GeForce RTX 5090, AMD Ryzen AI Halo, and MacBo

2 stories
1 story
1 story
1 story
4 stories
1 story
1 story
2 stories
1 story
1 story
1 story
1 story
1 story
1 story
1 story
2 stories
1 story
1 story
1 story
1 story
1 story
1 story
2 stories
4 stories
08:00 HuggingFace Daily Papers(社区热门论文)AI score 78/100

What AI Red-Team Evaluations Can and Cannot Prove

一项新研究为AI红队评测划定了可计算的证据上限,即固定测试预算下单一结果能改变信念的最大倍数,并以闭式解定位其边界。研究发现,在可计算的危害率之上,中等规模基准足以按既定证据标准认证某类别,且零失败记录比单次复现的失败更具说服力;低于该阈值则任何可行规模的被动基准都无法提供指定安全证据。对八个评测套件的审计显示,现有基准对高频危害类别充分,但对罕见灾难性危害类别仍差数个数量级。

08:00 HuggingFace Daily Papers(社区热门论文)AI score 73/100

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

腾讯推出 WorkBuddy Bench,一个覆盖 Code、Web、Office、Security 四个工作领域的编码智能体评测套件。每个任务均从真实 commit、PR 或业务场景逆向工程而来,改写为口语化角色扮演请求,从构造上抵抗数据污染。该基准在 CodeBuddy Code 和 Claude Code 上运行,所有任务目录、环境镜像、评分工具和参考方案均开源发布。

2 stories
1 story
2 stories
1 story
1 story
2 stories
1 story
1 story
2 stories
18:16 The Decoder:AI News(RSS)AI score 70/100

Only three AI models finished above starting capital in a 500-day startup survival test

普林斯顿大学推出CEO-Bench基准测试,让AI智能体在模拟环境中运营订阅软件公司NovaMind 500天,起始资金100万美元。14个测试模型中,仅Claude Fable 5(最佳轮次盈利4715万美元)、Claude Opus 4.8(2780万美元)和GPT-5.5(2130万美元)在最佳运行中超过起始资本。一个不调用语言模型的简单规则启发式方法通过固定定价、配额和针对性开发达到1576万美元,超越除上述三款外的所有模型。多数模型无法保持连贯策略,在模拟结束前破产。该测试旨在衡量AI的长期战略决策能力。

08:00 HuggingFace Daily Papers(社区热门论文)AI score 82/100

OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

OSWorld2.0 发布,包含108个长时域计算机使用工作流,覆盖日常与专业任务。每项任务用户中位数约1.6小时完成,Claude Opus 4.7(最大思考)平均需318次工具调用(OSWorld 1.0约30次)。基准聚焦流交互、动态环境、跨源推理、隐式状态推断、视觉空间精度等真实挑战。任务基于真实输入工件和状态化用户档案,附安全报告。500步二元完成指标下,Claude Opus 4.8(最大思考+批量调用)得分最高仅20.6%(部分54.8%);GPT-5.5更省token但约13%。结果表明当前智能体远未达专业级:瓶颈不在基本GUI控制或编码,而是丢失约束、错过中途信息、猜测而非询问、跳过验证,尤其依赖隐藏状态时最差。

1 story
1 story
1 story
2 stories
08:00 HuggingFace Daily Papers(社区热门论文)AI score 70/100

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers?

NatureBench是一个跨学科基准测试,包含90个从Nature系列同行评审论文中提取的任务,用于评估AI编码智能体能否超越复现、实现发现。基准基于NatureGym自动化管线,为每个任务提供标准化容器化环境,解决环境碎片化问题。在严格禁用网络搜索的协议下评估10种前沿智能体配置,最强模型仅在17.8%任务上超过已发表SOTA(g>0.1准则)。分析表明,智能体成功主要依赖方法论翻译,失败主因为方法选择错误和计算预算不足。已发布基准、NatureGym管线及公共排行榜。

08:00 Apple Machine Learning Research(RSS)AI score 68/100

Nine Judges, Two Effective Votes: Correlated Errors Undermine LLM Evaluation Panels

苹果机器学习研究团队发现,LLM-as-a-judge面板因模型间高度相关而严重受限。对7个模型家族的9个前沿大语言模型在3个自然语言推理数据集上的测试表明,9位评委实际仅提供约2个独立投票的信息量,面板准确率比独立投票理想值低8–22个百分点,最佳单一模型的表现已匹敌或超越整个面板。增加评委数量或改进聚合算法收效甚微,即使允许算法获取正确答案也仅能缩小至多11%的差距。该结论在多种提示变体、温度设置及偏好任务中均得到验证,瓶颈在于评委间的相关性而非聚合算法。

2 stories
1 story
2 stories
1 story
2 stories
1 story
1 story
2 stories
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
3 stories
08:00 HuggingFace Daily Papers(社区热门论文)AI score 70/100

Recovering Policy-Induced Errors: Benchmarking and Trajectory Synthesis for Robust GUI Agents

针对GUI智能体缺乏从自身错误中恢复能力的问题,本研究提出了GUI-RobustEval基准测试和RoTS轨迹合成框架。GUI-RobustEval包含1216个可执行测试用例,系统评估智能体在多种错误模式下的恢复能力。RoTS框架通过基于树的流程合成了80万条高质量数据。在此基础上训练的RoTS-7B和RoTS-32B模型,在GUI-RobustEval及传统基准上均获得显著提升。其中RoTS-32B在OSWorld上取得了47.4%的成功率和33.8%的All-Pass@4分数,表明长时程错误恢复能力的增强对鲁棒性和整体性能均有贡献。

08:00 HuggingFace Daily Papers(社区热门论文)AI score 75/100

WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction

针对现有基准无法精确诊断多模态智能体记忆在动态环境中的具体失败阶段,研究提出了“行动-世界交互循环”记忆模型,并构建了WorldMemArena基准。该基准包含400个多会话多模态任务,涵盖“终身进化”和“智能体执行”两类场景,支持对记忆写入、维护、检索和使用的阶段级评估。研究首次对长上下文、RAG等手工设计系统与基于框架的记忆智能体进行直接比较,发现记忆写入与存储质量的提升不直接带来性能改善,且多模态记忆在利用视觉证据及跨领域稳定性上仍存在挑战。

01:20 Hugging Face:Blog(RSS)AI score 70/100

ITBench-AA: Frontier Models Score Below 50% on the First Benchmark for Agentic Enterprise IT Tasks — by Artificial Analysis and IBM

由Artificial Analysis和IBM推出的ITBench-AA SRE基准测试显示,所有前沿大模型得分均未超过50%。Claude Opus 4.7(自适应推理,最大努力)以47%领先,GPT-5.5(xhigh)和Qwen3.7 Max分别得46%和42%。该测试包含59个需要通过Shell命令调查Kubernetes事件快照并提交根因诊断的智能体任务。关键发现是模型推理轮次差异近3倍,但更长的轨迹并不转化为更高准确率,过度调查的模型会因提交误报而受罚。在成本方面,开源模型Gemma 4 31B(Reasoning)以每任务$0.14的成本获得37%得分,优于成本更高但得分更低的闭源模型。

2 stories
1 story
1 story
1 story
1 story
2 stories
2 stories
Showing the latest 100 of 125 stories.