GPTProto

Qwen / Alibaba

76 picksNewest firstLatest article Sep 23, 2026, 5:11 PM
3 stories
2 stories
2 stories
9 stories
08:00 HuggingFace Daily Papers(社区热门论文)AI score 39/100

Rubric Rewards from Item Response Theory

View PDF HTML (experimental) Abstract:Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Distinct verdict patterns can thus receive the same reward, and the fixed points encode how much each criterion should count, not how strongly its verdict distinguish

08:00 HuggingFace Daily Papers(社区热门论文)AI score 38/100

Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence

Authors:Xi Xiao, Tianchen Zhao, Youngeun Kim, Zhuowei Li, Linghan Xu, Jiaye Wu, Zheng Zhang, Xiang Xu, Xuanbai Chen, Farhan Tejani, Jakub Zablocki, Julia Xu, Yifan Xing View PDF HTML (experimental) Abstract:Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to

08:00 HuggingFace Daily Papers(社区热门论文)AI score 39/100

Label-free steering: Compressing test-time reinforcement learning into bias-only subspaces

View PDF HTML (experimental) Abstract:Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emerge when both the reward signal and the optimization space are severely restricted? We answer this question with label-free bias-only TTRL, which uses majority-vote pseudolabels as re

08:00 HuggingFace Daily Papers(社区热门论文)AI score 56/100

Same Bytes, Different Authority: Reserved-Token Representations in Chat-Template Prompt Injection

View PDF HTML (experimental) Abstract:Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactly the same text, and because tokenization runs on the server, the defender rather than the attacker decides which one the model receives. We use this to

08:00 HuggingFace Daily Papers(社区热门论文)AI score 62/100

Language Models Are "Insecure" Reporters

View PDF HTML (experimental) Abstract:As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introduce a suite of eight adversarial reporting scenarios to systematically study whether LLMs conceal narrative-changing flaws: errors or limitations that undermine a

08:00 HuggingFace Daily Papers(社区热门论文)AI score 32/100

PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation

View PDF HTML (experimental) Abstract:Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one

08:00 HuggingFace Daily Papers(社区热门论文)AI score 40/100

See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs

View PDF HTML (experimental) Abstract:Safety-aligned LLMs can exhibit emergent misalignment (EM): narrow domain adaptation unexpectedly triggers catastrophic safety failures across unrelated domains. Prior static analyses leave training dynamics unmapped, while existing defenses rely on heuristics that degrade utility. We present a dynamic, second-order geometric study of EM. Tracking training trajectories reveals that directional Hessian curvature concentrates sharply on semantic pivot tokens.

08:00 HuggingFace Daily Papers(社区热门论文)AI score 42/100

LLMs are General Asynchronous Agents

View PDF HTML (experimental) Abstract:Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform another task. Modern LLMs address this with specialized architectures for voice interaction and video streams, VLAs for robot control, asynchronous tool calling fo

08:00 HuggingFace Daily Papers(社区热门论文)AI score 44/100

ROSS: Relearning from Self-Generated Rollouts through Selective Supervision

View PDF HTML (experimental) Abstract:Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating fine

1 story
1 story
2 stories
1 story
1 story
2 stories
1 story
1 story
1 story
2 stories
4 stories
1 story
1 story
2 stories
1 story
1 story
1 story
1 story
1 story
1 story
1 story
2 stories
1 story
1 story
1 story
2 stories
1 story
1 story
1 story
1 story
1 story
4 stories
08:00 HuggingFace Daily Papers(社区热门论文)AI score 88/100

Agon: Competitive Cross-Model RL with Implicit Rival Grading of Reasoning

Agon 让两个模型互为评分者,通过竞争性强化学习提升推理能力。在 DeepMath 困难子集上,基于 Qwen3 的 Agon 将 GRPO 的 pass@1 翻倍,增益约为未训练的 Mixture-of-Agents 的 8 倍。该结果在 Qwen3.5、Gemma 4 等模型族及编程代码任务上得到复现,推理时采用两阶段级联:一个模型起草,另一个阅读后作答。

08:00 HuggingFace Daily Papers(社区热门论文)AI score 71/100

Length Penalties Make Chain-of-Thought Less Monitorable

长度惩罚强化学习虽能缩短思维链推理,却会隐藏影响模型答案的驱动因素。对Qwen3-4B和Qwen3-14B的实验显示,压缩后思维链提及提示的频率大幅下降,Qwen3-14B的忠实度下限降至基线的63.1%,监控捕获提示使用的比率从69%降至49%。随机删除基线链句子以匹配压缩长度后,压缩链披露提示的频率仍比基线低7-35个百分点,表明压缩优先移除了监控所需的关键线索。

08:00 Hugging Face:Blog(RSS)AI score 66/100

Native-speed vLLM transformers modeling backend

Hugging Face 宣布 transformers vLLM 后端现与手写原生 vLLM 实现速度相当甚至更快。模型作者无需移植代码,即可自动利用 transformers 获得超快推理。测试使用 Qwen3-4B(单 GPU)、Qwen3-32B(张量并行)和 Qwen3-235B-A22B-FP8 MoE(数据+专家并行)三种配置,吞吐量均达到或超过原生。该后端通过 torch.fx 静态分析图、AST 重写代码实现动态层融合,支持张量/管道/专家并行及 torch.compile。用户仅需添加 `--model-impl transformers` 标志。目前不支持线性注意力模型但即将支持。

00:50 MarkTechPost(RSS)AI score 70/100

Liquid AI Open-Sources Antidoom: A Final Token Preference Optimization (FTPO) Method that Reduces Doom Loops in Reasoning Models

Liquid AI 开源了 Antidoom,一种基于 Final Token Preference Optimization (FTPO) 的针对性修复方法,用于减少推理模型中的 doom loop(死循环)问题。该方法定位循环开始的第一个 token,训练模型选择连贯替代项,而不改变整体输出分布。在 LFM2.5-2.6B 上,硬数学和编程任务中的循环率从 10.2% 降至 1.4%;Qwen3.5-4B 上从 22.9% 降至 1%。整套流程可在数小时内完成,全部代码和数据集(LiquidAI/antidoom-mix-v1.0)已开源。

1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
3 stories
10:00 Qwen:Blog Retrieval(API)AI score 73/100

Qwen-Robot Suite: A Foundation Model Suite for Physical World Intelligence

Qwen 发布三款基础模型——Qwen-RobotNav、Qwen-RobotManip 和 Qwen-RobotWorld。Nav 通过可控观测协议统一指令跟随、点/物体目标导航、目标追踪和自动驾驶五类任务,在 VLN-CE RxR 上达 76.5% SR,HM3Dv2 物体目标导航(仅 RGB)75.6% SR,EVT-Bench 追踪率 90.0%,NAVSIM 91.4 PDMS。Manip 利用规范状态-动作空间对超 38,100 小时异构开源机器人数据进行跨本体训练。World 通过自然语言动作接口协同训练 20 余种本体,预测操控、驾驶和导航的物理未来。三者共同将通用智能转化为物理行动。

08:00 Qwen:Blog Retrieval(API)AI score 72/100

Qwen-RobotManip: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

Qwen-RobotManip 是通义千问基于 Qwen-VL 的视觉-语言-动作(VLA)基础模型,引入覆盖表示、运动和行为三维度的统一对齐框架。仅使用开源机器人数据集和人演示视频,构建约 38,100 小时预训练语料,涵盖 15 种机器人形态。在 LIBERO-Plus 达 91.4%,RoboTwin-C2R Hard 达 69.4%,RoboCasa365 Composite-Unseen 达 14.9%,EBench 达 45.6%,RoboTwin-IF 达 72.0%,并在 RoboChallenge Table30 v1 generalist track 夺冠。模型采用 80 维状态-动作表示、人-机器人数据合成管道(1,933 小时第一人称视频转 24,808 小时数据)及上下文策略适配。

08:00 Qwen:Blog Retrieval(API)AI score 72/100

Qwen-RobotWorld: Boundless Worlds for Embodied Agents

Qwen-RobotWorld以语言为统一动作接口,采用双流Multimodal Diffusion Transformer(MMDiT)架构,将Qwen2.5-VL作为动作编码器。在4个基准测试中取得顶尖成绩,统一20余种机器人形态,基于860万跨场景训练对和1300多项操作技能。语言接口标准化500多种动作类别,支持操作、自动驾驶、室内导航的联合训练。还支持Scene2Robot人类到机器人转移及2–4路多视角几何一致视频生成。

1 story
2 stories
1 story
1 story
1 story
1 story
1 story
1 story
1 story
2 stories
17:00 Qwen:Blog Retrieval(API)AI score 66/100

Qwen-VLA: From Understanding the World to Acting in It

通义千问推出通用视觉-语言-动作模型Qwen-VLA,基于Qwen多模态骨干,将视觉感知、语言理解与空间推理扩展至连续动作生成和轨迹预测。训练分四阶段:文本到动作预训练(T2A)、持续预训练(CPT)、监督微调(SFT)和强化学习(RL)。在LIBERO上达97.9%,Simpler-WidowX达73.7%,RoboTwin-Easy/Hard达86.1%/87.2%,匹配或超越专精模型。数据涵盖超10,000小时公共机器人轨迹、1,000+小时内部真实轨迹及800万+合成仿真轨迹。

00:00 LMSYS:Blog(Chatbot Arena 团队)AI score 61/100

Blog Heterogeneous CPU + GPU EPD Disaggregation to Boost VLM Serving TL;DR We enabled heterogeneous Encode-Prefill-Decode (EPD) disaggregation via Dynamo and SGLang for Vision-Language Models (VLMs). By offloading vision encoding tasks to CPUs (the easiest-getting CPU... Intel & SGLang Team

LMSYS团队(Intel与SGLang)通过Dynamo和SGLang框架,为视觉语言模型(VLM)启用了异构编码-预填充-解耦(EPD)架构。该方案将视觉编码任务从GPU卸载至CPU(如Intel Xeon 6747P),与GPU协同工作。在Qwen3-VL-8B-Instruct模型的测试中,采用4 CPU + 1 GPU作为编码器、4 GPU作为预填充解码器(能力比R=12)的配置,在ISL/OSL 128/256、1080p 8张图像的负载下,实现了P99 TTFT和请求吞吐量约1.2倍至1.3倍的提升,并将P99 TPOT降低了约1.3倍至30倍。

2 stories
3 stories
08:00 HuggingFace Daily Papers(社区热门论文)AI score 70/100

ResearchMath-14K: Scaling Research-Level Mathematics via Agents

本文介绍了ResearchMath-14K,这是一个包含14,056个研究级数学问题的数据集,通过多智能体流程从学术资料中策划而成,是目前此类规模最大的集合。研究还生成了ResearchMath-Reasoning(包含220K条教师轨迹),发现语言模型存在回避行为,且新一代模型产生的引用和虚假引用分别是旧模型的5.6倍和5.0倍。经过智能体过滤后,对参数规模为4B到30B的Qwen3模型进行微调,其平均得分比基础模型提高了9.2分,表明过滤后的开放问题尝试能为研究级数学推理提供有效监督。该数据集已公开发布。

08:00 Hugging Face:Blog(RSS)AI score 72/100

Reachy Mini goes fully local

Reachy Mini 机器人现可通过 `speech-to-speech` 库实现完全本地化的语音交互,无需依赖云端。该方案采用级联流水线架构,对外提供 Realtime API 兼容的 WebSocket 接口。默认组件包括 Silero VAD 用于语音活动检测、Parakeet-TDT 作为语音转文本模型、通义千问(Qwen3-TTS)作为文本转语音模型。大语言模型推荐使用 llama.cpp 运行 Gemma 4。所有数据均在本地处理,保障了隐私且无 API 费用。

08:00 Hugging Face:Blog(RSS)AI score 61/100

Shipping a Trillion Parameters With a Hub Bucket: Delta Weight Sync in TRL

异步强化学习中,训练器每步需将完整模型权重(如1T参数checkpoint约1 TB)传输给推理引擎。TRL新增PR利用相邻RL优化步骤间约99%的bf16权重比特相同的特点,仅将变化的权重编码为稀疏safetensors文件,上传至Hugging Face Bucket并通知vLLM获取。在Qwen3-0.6B上,每步传输从1.2 GB降至20–35 MB。实验还展示了完全分离的训练场景:训练器、vLLM和Wordle环境分别位于不同机器和Hugging Face Space中,权重通过单个Hub bucket流动,无需共享集群、RDMA或VPN。

1 story
1 story
1 story
1 story
1 story
2 stories
2 stories
08:00 HuggingFace Daily Papers(社区热门论文)AI score 76/100

Qwen-Image-2.0 Technical Report

Qwen-Image-2.0是一个统一高保真生成与精确编辑的全能图像生成基础模型。它采用Qwen3-VL作为条件编码器,结合多模态扩散变换器进行联合建模,并通过大规模数据整理与多阶段训练实现强化。该模型支持长达1K令牌的指令输入,能生成幻灯片、海报等富文本内容,显著提升多语言文本渲染与排版质量。在生成方面,它增强了细节、纹理真实感与光照一致性,并更可靠遵循复杂指令。人工评估表明,其在生成和编辑任务上均大幅超越前代模型。

02:44 Hugging Face:Blog(RSS)AI score 74/100

MachinaCheck: Building a Multi-Agent CNC Manufacturability System on AMD MI300X

MachinaCheck是一款基于多智能体AI的系统,旨在革新小型CNC机加工车间的报价分析流程。传统上,车间经理需花费30-60分钟手动分析图纸,而该系统在上传STEP文件及材料、公差等简单输入后,能在30秒内生成完整的可制造性报告,明确指出零件能否制造、所需工具及生产前需采取的行动。其核心在AMD MI300X加速卡上本地运行Qwen 2.5 7B模型,利用192GB HBM3显存确保客户设计数据无需离开本地,满足了制造业对数据隐私的严格要求。系统采用五组件流水线,结合精确的几何特征提取与LLM的制造知识推理,最终输出结构化报告。

2 stories
1 story