GPTProto

推理能力

450 picksNewest firstLatest article Oct 2, 2026, 12:51 AM
1 story
1 story
3 stories
23:00 MIT News(RSS)AI score 61/100

This game-playing AI is the new champ at Stratego

A new AI system that excels at challenging games with hidden information could someday help human decision-makers select ideal strategies to outfox opponents in complicated situations like military maneuvers.Using advances in machine-learning, researchers from MIT, Carnegie Mellon University, New York University, and Stanford University developed an AI that defeated top-ranked human players of the board wargame Stratego by a large margin — something no AI system had been able to achieve. Strateg

18:30 OpenAI:官网动态(RSS · 排除企业/客户案例)AI score 80/100

Disrupting a coordinated model-distillation campaign

We recently identified and disrupted a coordinated campaign designed to extract protected reasoning from our models, with the earliest observed activity occurring in the first week of July. This activity is consistent with adversarial distillation: the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model. Protected reasoning is the model’s internal record for working through a task; extracting it can reveal information withheld fr

00:00 Artificial Analysis 完整文章(网页)AI score 78/100

Gemini 4 Argon: Google is back as one of the top three labs in intelligence achieved

See model page Google’s new Gemini 4 Argon equals GPT-6 Astra on the Artificial Analysis Intelligence Index at 60% of the Cost per Task with discounted prices Gemini 4 Argon is Google DeepMind’s first proprietary model above the Flash class in over 7 months. With high reasoning (the highest available), it scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53) and 1 point ahead of GPT-6.1 Sol (max, 52), with gains driven by lower hallucinations and stronger agenti

2 stories
1 story
2 stories
3 stories
2 stories
3 stories
1 story
1 story
4 stories
2 stories
3 stories
1 story
1 story
2 stories
9 stories
2 stories
1 story
2 stories
2 stories
2 stories
1 story
4 stories
08:00 Tomer Tunguz 博客(VC 分析)AI score 64/100

Mainframes became personal. So will your data center.

斯坦福大学与Together AI研究显示,本地AI模型在超百万条真实查询中,对89%的日常聊天与推理问题的回答质量已不输云端前沿模型。本地模型对前沿模型的胜/平率从2023年的23.2%升至2025年的71.3%,智能每瓦特效率同期提升5.3倍。相比全云端方案,本地模型加路由器的组合可降低80%能耗、77%算力与74%成本,但云端在多步推理与高难度领域仍具优势。

00:00 LMSYS:Blog(Chatbot Arena 团队)AI score 63/100

Blog Chasing the Batch-1 Floor: Ling-3.0-flash Speculative Decode on Blackwell Batch-1 decode keeps getting more important. Xiaomi MiMo, for example, announced MiMo-V2.5-Pro UltraSpeed in June, claiming 1,000 tok/s decode on a one-trillion-parameter MoE model. Batch 1 gives an ... RadixArk SGLang Team, Ant Ling Infra Team

蚂蚁 Ling Infra 团队与 RadixArk SGLang 团队将 Ling-3.0-flash 混合线性注意力 MoE 模型的单请求解码速度从 288 tok/s 提升至 606 tok/s,平均 TPOT 从 3.33 ms 降至 1.53 ms。

1 story
2 stories
1 story
1 story
2 stories
5 stories
01:00 Google Blog:AI(RSS)AI score 69/100

AMIE, our research medical AI system, demonstrates real-time clinical video consultation capabilities in a first-of-its-kind study.

Google Research 与 Google DeepMind 推进医疗 AI 系统 AMIE,实现实时临床视频问诊,首次在此场景展示专家级 AI 能力。该系统基于 Gemini 和 Project Astra 构建,可解读视觉与听觉线索、引导虚拟体格检查并实时诊断推理。随机研究中,临床评估者对 AMIE 的病史采集、诊断准确性等核心能力给予好评,患者演员也更偏好视频体验。

00:31 Dwarkesh Patel:Podcast & Blog(RSS)AI score 60/100

Ryan Greenblatt – Human level AIs might build runaway superintelligences by 2032

Dwarkesh Patel与Redwood Research首席科学家Ryan Greenblatt探讨递归自我改进(RSI)的可能性:一旦AI达到人类顶级专家水平,可能在一年内实现相当于4-5年的AI进展,Ryan的中位预期是2031年自动化AI研发。双方还讨论了超级智能的对齐对象、奖励黑客行为是否会升级为AI联手接管世界等风险。

00:00 Google Research:Blog(网页)AI score 66/100

Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

Google Research 提出知识画像框架,发现前沿 LLM(如 Gemini3、GPT-5)的事实编码接近饱和,但回忆(recall)能力不足,多数事实错误源于“丢钥匙”而非“空货架”。该框架将事实分为编码失败、回忆失败等五类画像,并配套推出 WikiProfile 基准,含 2,150 条维基百科事实,每条配 10 个问题,用于分别探测编码、回忆与识别能力。

5 stories
21:51 LMSYS:Blog(Chatbot Arena 团队)AI score 72/100

Blog Unified Radix Cache: One Tree for Hybrid Model Prefix Caching Prefix caching reuses KV when requests share the same token prefix. Under full attention, once the KV for a shared prefix is computed, it remains valid as more tokens are appended. A later request wit... Zhangheng Huang, Ke Bao, Yi Zhang, Jialin Ouyang, Sicheng Pan

LMSYS 团队提出 Unified Radix Cache,用单一 token 键控 radix 拓扑统一管理混合模型的 FULL、SWA 和 MAMBA 组件缓存,各组件独立执行路径、滑动窗口和检查点复用语义。

21:51 LMSYS:Blog(Chatbot Arena 团队)AI score 74/100

Blog SGLang Adds Day-0 Support for NVIDIA Nemotron 3.5 Lightning SGLang is excited to announce Day-0 support for NVIDIA Nemotron 3.5 Lightning, a customizable open model built to power always-on agents across local systems, the edge, the datacenter, and the cloud. ... NVIDIA Nemotron Team and SGLang Team

SGLang 宣布对 NVIDIA Nemotron 3.5 Lightning 提供 Day-0 支持,该开源模型为 30B 总参数、3B 激活参数的混合专家架构,支持最长 1M token 上下文,可从 Hugging Face 下载 BF16 和 NVFP4 权重。模型支持 MTP、DFlash、DSpark 三种投机解码技术,并可通过 OpenAI 兼容 API 接入智能体工作流。

4 stories
08:00 OpenRouter:Announcements(RSS)AI score 77/100

Model Routing Powered by Wisdom of the Market

OpenRouter 基于每周超 55T token 的社区消费数据,推出新版 Auto 路由器(openrouter/auto),其模型选择在多数任务和成本档位上优于旧版。新路由器按约 30 种任务类型匹配近 7 天社区实际消费的模型,支持 cost_tier 参数(low 至 max)并遵循账户隐私设置。在 MMLU Pro 等基准上,新默认档位在多数领域以更低成本达到旧版同等性能。

1 story
1 story
1 story
3 stories
1 story
1 story
1 story
3 stories
2 stories
1 story
1 story
1 story
2 stories
2 stories
1 story
1 story
Showing the latest 100 of 450 stories.