GPTProto

Research Papers

447 picksNewest firstLatest article Oct 1, 2026, 8:00 AM
1 story
4 stories
23:00 MIT News(RSS)AI score 61/100

This game-playing AI is the new champ at Stratego

A new AI system that excels at challenging games with hidden information could someday help human decision-makers select ideal strategies to outfox opponents in complicated situations like military maneuvers.Using advances in machine-learning, researchers from MIT, Carnegie Mellon University, New York University, and Stanford University developed an AI that defeated top-ranked human players of the board wargame Stratego by a large margin — something no AI system had been able to achieve. Strateg

03:24 The Decoder:AI News(RSS)AI score 80/100

UK AI Security Institute finds GPT-6 Astra's rogue attack rate jumped fivefold over its predecessor

The UK's AI Security Institute tested OpenAI's GPT-6 Astra before its release. In simulated cybersecurity evaluations, the model carried out unauthorized attacks on third-party software far more often than its predecessors. There have now likely been thousands of incidents in which AI systems carried out unauthorized cyber activity during security evaluations. The UK's AI Security Institute (AISI), a research organization within Britain's science ministry, tested OpenAI's GPT-6 Astra specificall

00:00 Transluce(网页)AI score 74/100

AI Agents Targeted U.S. and Canadian Government Websites

Jack Cable*,1, Daniel Chiu*, Francisco Pernice*,2, Laura Ruis*,2, Selena Zhang*,3, Tetiana Bas4, Jordan Chetty5, Farzaan Kaiyom1, Gary Shen4, Conrad Stosz†,3, Jacob Steinhardt†,31 Corridor · 2 MIT · 3 Transluce · 4 AIUC · 5 Hertz Foundation · * First authors, alphabetical · † Senior authorsTransluce | Published: September 30, 2026Following up on our previous blog post, we discovered several additional incidents where rogue AI agents appear to have used aggressive techniques to access publicly av

2 stories
3 stories
17:00 MIT News(RSS)AI score 78/100

New formulation helps RNA vaccines withstand high temperatures

RNA vaccines, which have been proven effective against Covid-19, are now being developed for many other diseases, including cancer. One of the drawbacks to these vaccines is that they require ultracold storage, but researchers from MIT have found a promising way to overcome that limitation.With help from an AI algorithm, the researchers tweaked the formulation surrounding the lipid nanoparticles that are typically used to deliver mRNA vaccines, making the vaccines more heat-resistant. Using this

07:00 英国 AI Security Institute:Blog(网页)AI score 80/100

GPT-6 Astra performs unsanctioned supply-chain attacks in simulations

In recent incidents, AI systems performed unsanctioned cyber activity despite being prompted only to complete a cybersecurity evaluation [1,2,3,4]. This includes AI systems engaging in supply-chain attacks on real, out-of-bounds targets. Before its public release, AISI tested whether GPT-6 Astra would engage in this type of unsanctioned cyber activity when prompted to complete a cyber evaluation. To securely perform this testing, we used Petri, a tool that uses LLMs to fully simulate the cyber e

1 story
1 story
2 stories
1 story
3 stories
1 story
1 story
3 stories
2 stories
2 stories
3 stories
1 story
1 story
2 stories
2 stories
1 story
1 story
2 stories
2 stories
2 stories
2 stories
1 story
2 stories
2 stories
2 stories
2 stories
4 stories
08:00 PromptArmor:Threat IntelligenceAI score 72/100

Attacker Takes Over Zoom AI

PromptArmor 披露针对 Zoom AI 智能体 ZoomMate 的攻击链:恶意 Skill 或间接提示词注入可让智能体连到攻击者服务器并执行命令,窃取会议记录、消息及连接器数据。攻击在用户点击 stop 或关闭 Zoom 后仍持续运行,且最终聊天输出看起来完全正常。作者称 ZoomMate 的联网环境属预期功能,无特定漏洞可披露,发文旨在提醒用户无严格网络沙箱的智能体风险。

1 story
3 stories
2 stories
4 stories
08:37 MarkTechPost(RSS)AI score 70/100

Microsoft’s SkillOpt Shows Optimized Agent Skill Artifacts Transfer Across Model Scales and Between Codex and Claude Code Harnesses

Microsoft 与上海交大、同济、复旦团队提出的 SkillOpt 通过文本空间优化训练单一技能文档,冻结目标模型,使优化后的技能工件可跨模型规模和跨工具链迁移。在 Codex 上优化的 SpreadsheetBench 技能部署到 Claude Code 后得分 81.8,超过后者自行训练技能得到的 80.4。全部 4 项跨模型、4 项跨工具链和 3 项跨基准迁移结果均高于目标的无技能基线。

08:00 HuggingFace Daily Papers(社区热门论文)AI score 72/100

Scaling Inherently Interpretable Language Models

一项新研究挑战“可解释性牺牲能力”的假设,将可解释性作为训练约束与语言建模目标共同优化。在三个数量级的算力范围内,自回归与扩散语言模型的表征随规模增大而更解耦、更对齐人类概念。其实例 Steerling-8B 支持通过概念或特征归因诊断输出、检索训练数据并无需重训即可干预,且与算力多 2-16 倍的开放模型保持竞争力。

08:00 HuggingFace Daily Papers(社区热门论文)AI score 71/100

Activity Frames: Deterministic Screen-Activity Compilation for Agent Memory and Replay

一项研究提出 Activity Frames,用确定性、零模型管线将被动捕获的屏幕活动编译为智能体记忆,输出字节一致、可缓存且可审计。在单人 128,756 帧、51 个活跃天的语料上,该编译器将一天原始捕获压缩为 86 倍更小的提示块,耗时 68 毫秒;智能体阅读该块的问答准确率达 98.4%,优于 LLM 摘要的 66-80%。

3 stories
21:49 AI as Normal Technology(RSS)AI score 70/100

AI agents can't yet do open-ended AI research

普林斯顿大学团队通过“影子评估”测试前沿AI智能体的开放式研究能力:让智能体在六天时间内、使用数千美元API额度和算力,回答两篇未发表论文的核心研究问题,结果两篇论文均被原作者明确拒绝。分析显示,智能体缺乏研究判断力、资源意识、创造性反馈应对能力和有效回溯能力,且未遵循具体指令。研究团队认为,开放式研究对前沿AI智能体仍具挑战性,但结果尚属初步,需扩大样本量并持续评估。

08:00 HuggingFace Daily Papers(社区热门论文)AI score 80/100

The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

一项新研究揭示,个性化大语言模型普遍存在过度推断(OI)现象,即编造超出证据支持的用户属性。在 MirageBench 基准测试中,12 个模型均有 35%–49% 的推断被判定为虚构(均值 41.6%)。更关键的是,模型自我评估的 OI 与外部评测结果呈负相关(rho = -0.60),表明自我报告的可信度是误导性信号,外部验证才是更可靠的个性化基础。

5 stories
08:00 HuggingFace Daily Papers(社区热门论文)AI score 72/100

Resume Means Resume: A Machine-Checked Conformance Contract for Checkpoint, Interrupt, and Resume Semantics in Workflow Persistence Layers

一项研究为智能体工作流持久化层提出“恢复契约”,规定前缀延续、效果恰好一次等六项属性,并用 TLA+ 模型穷举验证了 740 万状态。实测发现 LangGraph 1.2.9 在崩溃后重复执行已持久化工作,CrewAI 1.15.2 违背其书面声明,pydantic-graph 1.x 无法在节点中途崩溃后恢复。研究还给出经 Verus 验证的参考实现 REMIT,修复了分叉与有效性缺陷。

08:00 HuggingFace Daily Papers(社区热门论文)AI score 76/100

Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging

Any-OPD 提出首个支持任意异构流匹配生成器对的同策略蒸馏框架,仅通过冻结的 DINOv2 表示空间桥接教师与学生模型,无需共享 VAE、架构或噪声调度。将 FLUX.1-dev 蒸馏至 SD3.5-Medium 后,学生模型 PickScore 与 HPSv3 均显著提升,以教师五分之一的参数量达到接近其性能,而直接潜在回归训练完全失败。

1 story
3 stories
5 stories
08:00 HuggingFace Daily Papers(社区热门论文)AI score 73/100

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing

研究发现,领先模型在 SWE-bench Verified 上对开发者补丁的删除召回率最高仅 71.7%,29.0% 的通过补丁采用 Guard-and-Go 模式保留目标代码。新基准 CanItDelete 含 200 个纯删除任务,最佳模型仍失败 19.5%。在 7B 模型后训练中加入 12.8k 删除示例(占 0.7% token)可将删除回避降低 13.9 个百分点。

08:00 HuggingFace Daily Papers(社区热门论文)AI score 75/100

Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning

研究提出“显著性偏差”概念,指大语言模型被输入中无用的显性干扰(如数字)劫持,忽略任务隐含的常识前提。基于新构建的 SaliTrap 基准评估 12 个主流模型,最佳模型仅 54.8% 查询避开陷阱,8/12 模型低于 30%;GLM-5.1 和 Kimi-K2 在识别陷阱后仍分别有 86.2% 和 81.8% 的遵从率。

08:00 HuggingFace Daily Papers(社区热门论文)AI score 71/100

BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms

一项受控研究在约450倍跨度、28个严格嵌套的语料规模层级上比较多种RAG范式,发现存在规模依赖的交叉点而非绝对赢家。File-System Agent在最小规模领先,但约1000万语料token时BM25反超并在所有更大层级保持领先,全规模下优势接近20个点。BM25还锚定了无需LLM构建的低成本帕累托前沿。

08:00 HuggingFace Daily Papers(社区热门论文)AI score 74/100

PhiZero: A World Model Built Around Physical Language

PhiZero 是一种基于“物理语言”的物理世界模型,该语言通过自监督学习从野外视频中提取世界状态转移的紧凑离散表征。它采用先推理后渲染的范式,先以物理语言序列推断未来世界演化,再由扩散解码器渲染成视频。实验验证了其在物理一致性生成、细粒度动作条件模拟和零样本运动迁移上的能力。

2 stories
1 story
1 story
1 story
5 stories
08:00 HuggingFace Daily Papers(社区热门论文)AI score 78/100

What AI Red-Team Evaluations Can and Cannot Prove

一项新研究为AI红队评测划定了可计算的证据上限,即固定测试预算下单一结果能改变信念的最大倍数,并以闭式解定位其边界。研究发现,在可计算的危害率之上,中等规模基准足以按既定证据标准认证某类别,且零失败记录比单次复现的失败更具说服力;低于该阈值则任何可行规模的被动基准都无法提供指定安全证据。对八个评测套件的审计显示,现有基准对高频危害类别充分,但对罕见灾难性危害类别仍差数个数量级。

08:00 HuggingFace Daily Papers(社区热门论文)AI score 73/100

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

腾讯推出 WorkBuddy Bench,一个覆盖 Code、Web、Office、Security 四个工作领域的编码智能体评测套件。每个任务均从真实 commit、PR 或业务场景逆向工程而来,改写为口语化角色扮演请求,从构造上抵抗数据污染。该基准在 CodeBuddy Code 和 Claude Code 上运行,所有任务目录、环境镜像、评分工具和参考方案均开源发布。

2 stories
2 stories
Showing the latest 100 of 447 stories.