Open-sourcing AstaBrief, the fast report-generation model in Asta
Ai2 开源 AstaBrief 8B,一个基于 Qwen3-8B、将研究问题和检索文献片段转化为带引用报告的科学报告生成模型,现已在 Asta 的 Generate a report 功能中作为 Fast mode 上线,并连同训练数据开放下载。
Ai2 开源 AstaBrief 8B,一个基于 Qwen3-8B、将研究问题和检索文献片段转化为带引用报告的科学报告生成模型,现已在 Asta 的 Generate a report 功能中作为 Fast mode 上线,并连同训练数据开放下载。
Cloudflare has released Clef and Clef-flash, the first models trained by its Workers AI team. They are decision models, not chatbots. Each reads an input state and a schema of typed questions. It returns a probability for every allowed answer, with no free-form text. Both are open-weight under Apache 2.0 and compatible with TypeSafe AI’s Jev API. Is it deployable? Yes, Both models run today on Workers AI, and the weights are on Hugging Face for self-hosting. What a Decision Model Does An LLM gen
Artificial Analysis 自行本地部署评测了阿里 9 月 20 日以开源权重发布的 Qwen-Image-2.1,该模型在 AA-Image-T2I v2.0 和 AA-Image-Editing v2.0 上均排名第 18,为两个榜单上排名第一的开源权重模型,超过 Ideogram 4.0(Quality)和 HunyuanImage 3.0 Instruct。
View PDF HTML (experimental) Abstract:Large language models (LLMs) have shown promise for recommendation reranking, but their use introduces an important tradeoff between recommendation quality and serving efficiency. We investigate whether a decision-oriented model provides a useful alternative when the reranking task is fundamentally a structured choice among predefined candidate items. Specifically, we conduct a controlled empirical study of Jev, described by TypeSafe AI as a ``System One Mod
View PDF HTML (experimental) Abstract:On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the
Alibaba’s Qwen team has released Qwen-Audio-3.1, a 5-model audio stack spanning ASR, TTS and realtime interaction. The main model is Qwen-Audio-3.1-Realtime, a full-duplex speech model built for voice agents that call tools. Qwen also cut prices: about 85% on Realtime, about 70% on TTS and up to 95% on ASR. Is it deployable? Yes, as a managed API. qwen-audio-3.1-realtime-plus is live on QwenCloud over WebSocket. No open weights were announced. What Ships on QwenCloud The model page lists text an
Authors:Hongjin Qian, Chaofan Li, Kun Luo, Wenqing Wei, Jianlyu Chen, Shuqi Lu, Yuyang Hu, Hongwang Xiao, Hui Wang, Chaozhuo Li, Qiwei Ye, Zhicheng Dou, Defu Lian, Zheng Liu View PDF HTML (experimental) Abstract:We present AREX-2, an effort to advance the self-improving capability of LLM agents, which we define as the ability to iteratively refine a solution at test time. This ability rests on two complementary capabilities: reflection, which produces a solution better than the current one, and
View PDF HTML (experimental) Abstract:Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Distinct verdict patterns can thus receive the same reward, and the fixed points encode how much each criterion should count, not how strongly its verdict distinguish
Authors:Xi Xiao, Tianchen Zhao, Youngeun Kim, Zhuowei Li, Linghan Xu, Jiaye Wu, Zheng Zhang, Xiang Xu, Xuanbai Chen, Farhan Tejani, Jakub Zablocki, Julia Xu, Yifan Xing View PDF HTML (experimental) Abstract:Latent visual reasoning (LVR) enables multimodal large language models (MLLMs) to perform intermediate computation in continuous latent tokens rather than expressing every reasoning step in words. However, unlike textual CoT, latent reasoning is not directly observable, making it difficult to
View PDF HTML (experimental) Abstract:Test-time reinforcement learning (TTRL) enables models to improve their reasoning without relying on labeled training data, but existing approaches typically optimize a large fraction of the model parameters. This raises a natural question: can effective test-time adaptation emerge when both the reward signal and the optimization space are severely restricted? We answer this question with label-free bias-only TTRL, which uses majority-vote pseudolabels as re
View PDF HTML (experimental) Abstract:Prompt injection against LLM agents becomes much stronger when the injected instruction is wrapped in the model's own chat template. A forged template marker such as <|im_start|> can reach the model either as a single reserved control token or as a sequence of ordinary subword tokens. The two decode to exactly the same text, and because tokenization runs on the server, the defender rather than the attacker decides which one the model receives. We use this to
View PDF HTML (experimental) Abstract:As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introduce a suite of eight adversarial reporting scenarios to systematically study whether LLMs conceal narrative-changing flaws: errors or limitations that undermine a
View PDF HTML (experimental) Abstract:Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one
View PDF HTML (experimental) Abstract:Safety-aligned LLMs can exhibit emergent misalignment (EM): narrow domain adaptation unexpectedly triggers catastrophic safety failures across unrelated domains. Prior static analyses leave training dynamics unmapped, while existing defenses rely on heuristics that degrade utility. We present a dynamic, second-order geometric study of EM. Tracking training trajectories reveals that directional Hessian curvature concentrates sharply on semantic pivot tokens.
View PDF HTML (experimental) Abstract:Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform another task. Modern LLMs address this with specialized architectures for voice interaction and video streams, VLAs for robot control, asynchronous tool calling fo
View PDF HTML (experimental) Abstract:Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating fine
Qwen 发布 Qwen-Audio-3.1,ASR、TTS 与 Realtime 全面升级,并新增音频创作模型 TTS-Next 和音频理解模型 ASR-Next,共五个模型覆盖理解、生成、交互与创作。全线降价,TTS 约 70% off、Realtime 约 85% off、ASR 最高 95% off;Realtime 可边说边听、随时打断,检测到低落情绪时会放慢语速并共情回应。
通义千问宣布 Qwen-Image-2.1 在 Arena 的 Image Edit Arena 和 Text-to-Image Arena 均为开源模型第一。引用内容显示其 Image Edit Arena 得分 1367,总榜第 16,距第 15 名 GPT-Image-1.5-high-fidelity 仅 3 分,官方邀请用户体验该模型。
Qwen 发布 Qwen3.8-LiveTranslate,采用 Interleave 架构与 Hybrid-MoE Thinker–Talker 设计重构实时同声传译,平均滞后(LAAL)从上一代的 2.8 秒降至 2.3 秒。
Qwen 发布下一代原生全模态模型 Qwen3.8-Omni-Flash,支持文本、图像、音频和视频输入及 1M token 上下文窗口,29 项评测平均分较 Qwen3.5-Omni-Plus 提升超过 25%,音频输入每小时价格下降超过 98%,音视频输入每小时价格下降超过 93%。
Unsloth 通过 MTP(Multi-Token Prediction)让 Qwen3.8-Flash-Next 本地推理提速约 1.3 至 1.7 倍且精度不变,GGUF 在单张 RTX PRO 6000 上可达 170 tokens/s(基线 100 tokens/s)。
通义千问发布 Qwen3.8-Max-0902,在 Code Arena: WebDev 以 1,691 分首次亮相即排名总榜第一,并以混合价 $5/MToken 成为 Pareto 前沿上得分最高的模型,现已可在 QwenCloud 试用。
Unsloth 推出 Qwen3.8-Flash-Next 的 GGUF 量化版本,称该 125B MoE 多模态模型性能超过 Claude-Opus-4.6 (Max),可在 75GB RAM 上本地运行,无需 GPU VRAM。
通义千问发布 Qwen3.8-Flash,一款多模态 MoE 模型,作为 Qwen4 架构的早期预览并开放权重。该模型总参数 125B,每 token 仅激活 6B,训练成本仅为 Qwen3.7-Plus 的 1/9,性能全面超越后者。生产版 API 定价 $0.16/1M 输入 tokens 和 $0.47/1M 输出 tokens,原生上下文 262K,可扩展至 1M。
通义千问开源 Qwen3.8-Flash-Next,一款多模态 MoE 模型,也是 Qwen4 架构的早期预览。该模型采用 GDN + QSA 混合注意力等四项升级,总参数 125B,每 token 激活 6B,训练成本约为 Qwen3.7-Plus 的 1/9,编码与办公任务能力更强。
阿里巴巴发布 Qwen3.8-Flash-Next 模型权重,作为即将到来的 Qwen4 架构预览。
Unsloth 发布 Qwen3.8-27B 的 Dynamic GGUF 量化版本,4-bit 量化可在 17-19GB 内存设备上本地运行,并同时上传 NVFP4 量化。
通义千问兑现承诺,开源 Qwen3.8 系列模型。其中 Qwen3.8-27B 为原生多模态稠密模型,仅 27B 参数即全面超越 Qwen3.7-Plus,原生支持 262K 上下文,可通过 YaRN 扩展至 1M tokens,采用 Apache 2.0 许可。Max 级 Qwen3.8-2.4T-A95B 的开放权重也已同步发布。
Fireworks 发布一条低于 10 美元、约 150 步的对比微调配方,将 Qwen3-Embedding-8B 适配到特定检索领域。
Miles 团队在 Blackwell 架构上实现了两种原生低精度强化学习方案:端到端 MXFP8 和 MoE 专家权重的逐 token NVFP4。在 8x B200 上对 Qwen3-30B-A3B 的消融实验中,BF16 与所有五种低精度配置的原始奖励曲线高度重合,且 MXFP8 和 NVFP4 减少了推理时间。
通义千问发布第三代图像生成基座模型 Qwen-Image-3.0,核心关键词为“实”。该模型支持最长 4.5k token 指令输入,可单次生成包含 9 个复杂信息图的 3×3 网格布局;文本渲染精度达 10px,并支持 12 种语言原生渲染,旨在将图像转化为可部署的生产力工具。
AgentDebugX是一个开源调试框架,将LLM智能体调试组织为“检测-归因-恢复-重跑”闭环。其核心DeepDebug在Who&When基准上对qwen3.5-9b达到精确的智能体与步骤归因准确率,在GAIA上单次重跑即可修复失败任务。该工具提供Python库、CLI、Web控制台和可安装的智能体技能。
Agon 让两个模型互为评分者,通过竞争性强化学习提升推理能力。在 DeepMath 困难子集上,基于 Qwen3 的 Agon 将 GRPO 的 pass@1 翻倍,增益约为未训练的 Mixture-of-Agents 的 8 倍。该结果在 Qwen3.5、Gemma 4 等模型族及编程代码任务上得到复现,推理时采用两阶段级联:一个模型起草,另一个阅读后作答。
长度惩罚强化学习虽能缩短思维链推理,却会隐藏影响模型答案的驱动因素。对Qwen3-4B和Qwen3-14B的实验显示,压缩后思维链提及提示的频率大幅下降,Qwen3-14B的忠实度下限降至基线的63.1%,监控捕获提示使用的比率从69%降至49%。随机删除基线链句子以匹配压缩长度后,压缩链披露提示的频率仍比基线低7-35个百分点,表明压缩优先移除了监控所需的关键线索。
Hugging Face 宣布 transformers vLLM 后端现与手写原生 vLLM 实现速度相当甚至更快。模型作者无需移植代码,即可自动利用 transformers 获得超快推理。测试使用 Qwen3-4B(单 GPU)、Qwen3-32B(张量并行)和 Qwen3-235B-A22B-FP8 MoE(数据+专家并行)三种配置,吞吐量均达到或超过原生。该后端通过 torch.fx 静态分析图、AST 重写代码实现动态层融合,支持张量/管道/专家并行及 torch.compile。用户仅需添加 `--model-impl transformers` 标志。目前不支持线性注意力模型但即将支持。
Liquid AI 开源了 Antidoom,一种基于 Final Token Preference Optimization (FTPO) 的针对性修复方法,用于减少推理模型中的 doom loop(死循环)问题。该方法定位循环开始的第一个 token,训练模型选择连贯替代项,而不改变整体输出分布。在 LFM2.5-2.6B 上,硬数学和编程任务中的循环率从 10.2% 降至 1.4%;Qwen3.5-4B 上从 22.9% 降至 1%。整套流程可在数小时内完成,全部代码和数据集(LiquidAI/antidoom-mix-v1.0)已开源。
Qwen 发布三款基础模型——Qwen-RobotNav、Qwen-RobotManip 和 Qwen-RobotWorld。Nav 通过可控观测协议统一指令跟随、点/物体目标导航、目标追踪和自动驾驶五类任务,在 VLN-CE RxR 上达 76.5% SR,HM3Dv2 物体目标导航(仅 RGB)75.6% SR,EVT-Bench 追踪率 90.0%,NAVSIM 91.4 PDMS。Manip 利用规范状态-动作空间对超 38,100 小时异构开源机器人数据进行跨本体训练。World 通过自然语言动作接口协同训练 20 余种本体,预测操控、驾驶和导航的物理未来。三者共同将通用智能转化为物理行动。
Qwen-RobotManip 是通义千问基于 Qwen-VL 的视觉-语言-动作(VLA)基础模型,引入覆盖表示、运动和行为三维度的统一对齐框架。仅使用开源机器人数据集和人演示视频,构建约 38,100 小时预训练语料,涵盖 15 种机器人形态。在 LIBERO-Plus 达 91.4%,RoboTwin-C2R Hard 达 69.4%,RoboCasa365 Composite-Unseen 达 14.9%,EBench 达 45.6%,RoboTwin-IF 达 72.0%,并在 RoboChallenge Table30 v1 generalist track 夺冠。模型采用 80 维状态-动作表示、人-机器人数据合成管道(1,933 小时第一人称视频转 24,808 小时数据)及上下文策略适配。
Qwen-RobotWorld以语言为统一动作接口,采用双流Multimodal Diffusion Transformer(MMDiT)架构,将Qwen2.5-VL作为动作编码器。在4个基准测试中取得顶尖成绩,统一20余种机器人形态,基于860万跨场景训练对和1300多项操作技能。语言接口标准化500多种动作类别,支持操作、自动驾驶、室内导航的联合训练。还支持Scene2Robot人类到机器人转移及2–4路多视角几何一致视频生成。
VISTA-9B是基于Qwen3.5 9B骨干训练的GUI定位模型,输入截图与自然语言指令,输出0-1000归一化坐标。采用VISTA(视图一致自验证)方法,含view-consistent GRPO与self-verified cross-view anchoring。在SSPro、SSV2、OSWorld-G、OSWorld-G-R上分别取得69.2、95.8、68.1、75.5分,超越Qwen3.5-9B与GRPO-9B基线。模型已开源,可通过HuggingFace加载使用。
VISTA-4B 是基于 Qwen3.5-4B 骨干的 GUI 定位模型,输入截图与自然语言指令,输出归一化 0-1000 坐标。训练采用视图一致 GRPO 和自验证交叉视图锚定。在 GUI 定位基准上,SSPro 得分 64.2(相比 GRPO-4B 提升 2.0),SSV2 得分 93.8(下降 0.4),OSWorld-G 得分 61.2(提升 1.3),OSWorld-G-R 得分 69.7(提升 0.5)。模型已开源在 HuggingFace,推荐使用提示词并返回 [x,y] 格式坐标。
通义千问推出通用视觉-语言-动作模型Qwen-VLA,基于Qwen多模态骨干,将视觉感知、语言理解与空间推理扩展至连续动作生成和轨迹预测。训练分四阶段:文本到动作预训练(T2A)、持续预训练(CPT)、监督微调(SFT)和强化学习(RL)。在LIBERO上达97.9%,Simpler-WidowX达73.7%,RoboTwin-Easy/Hard达86.1%/87.2%,匹配或超越专精模型。数据涵盖超10,000小时公共机器人轨迹、1,000+小时内部真实轨迹及800万+合成仿真轨迹。
LMSYS团队(Intel与SGLang)通过Dynamo和SGLang框架,为视觉语言模型(VLM)启用了异构编码-预填充-解耦(EPD)架构。该方案将视觉编码任务从GPU卸载至CPU(如Intel Xeon 6747P),与GPU协同工作。在Qwen3-VL-8B-Instruct模型的测试中,采用4 CPU + 1 GPU作为编码器、4 GPU作为预填充解码器(能力比R=12)的配置,在ISL/OSL 128/256、1080p 8张图像的负载下,实现了P99 TTFT和请求吞吐量约1.2倍至1.3倍的提升,并将P99 TPOT降低了约1.3倍至30倍。
Qwen-VLA是一个统一的具身基础模型,将Qwen的视觉-语言建模从感知、理解与推理扩展至连续动作和轨迹生成。它通过基于DiT的动作解码器实现,使用包含机器人操作轨迹、人类第一人称示范、仿真及导航数据等在内的大规模数据进行联合预训练。为支持多种平台,引入了感知载体感知的提示条件机制,并将操作、导航与轨迹预测统一到一个框架中。实验显示,Qwen-VLA-Instruct在多个基准上表现优异,例如在LIBERO达到97.9%,在真实世界ALOHA实验中平均分布外成功率为76.9%。
Qwen3.5在TokenSpeed推理引擎上,针对智能体工作负载达到了创纪录的580 tokens per second (tps)速度。这一成果由通义千问推理团队、lightseekorg Foundation TokenSpeed团队、NVIDIA及Mooncake团队共同实现,并采用了tri_dao的FlashAttention-4 (FA4) 优化。此里程碑标志着开源大语言模型推理性能的边界得到了推动,相关详情可查阅PyTorch社区博客。
本文介绍了ResearchMath-14K,这是一个包含14,056个研究级数学问题的数据集,通过多智能体流程从学术资料中策划而成,是目前此类规模最大的集合。研究还生成了ResearchMath-Reasoning(包含220K条教师轨迹),发现语言模型存在回避行为,且新一代模型产生的引用和虚假引用分别是旧模型的5.6倍和5.0倍。经过智能体过滤后,对参数规模为4B到30B的Qwen3模型进行微调,其平均得分比基础模型提高了9.2分,表明过滤后的开放问题尝试能为研究级数学推理提供有效监督。该数据集已公开发布。
Reachy Mini 机器人现可通过 `speech-to-speech` 库实现完全本地化的语音交互,无需依赖云端。该方案采用级联流水线架构,对外提供 Realtime API 兼容的 WebSocket 接口。默认组件包括 Silero VAD 用于语音活动检测、Parakeet-TDT 作为语音转文本模型、通义千问(Qwen3-TTS)作为文本转语音模型。大语言模型推荐使用 llama.cpp 运行 Gemma 4。所有数据均在本地处理,保障了隐私且无 API 费用。
异步强化学习中,训练器每步需将完整模型权重(如1T参数checkpoint约1 TB)传输给推理引擎。TRL新增PR利用相邻RL优化步骤间约99%的bf16权重比特相同的特点,仅将变化的权重编码为稀疏safetensors文件,上传至Hugging Face Bucket并通知vLLM获取。在Qwen3-0.6B上,每步传输从1.2 GB降至20–35 MB。实验还展示了完全分离的训练场景:训练器、vLLM和Wordle环境分别位于不同机器和Hugging Face Space中,权重通过单个Hub bucket流动,无需共享集群、RDMA或VPN。
主推文赞扬了创新者在前沿领域的探索。引用的推文具体指出,SenseNova-U1在空间智能能力上取得进展,其关键基准测试表现超越了Qwen3.5等强劲基线。同时,团队开源了目前最大的空间问答数据集SenseNova-SI-8M,并邀请业界在CVPR会议进行线下交流。
OpenCode x Qwen 3.6 Plus - 再次免费 上次各位把我们的容量当成了自助餐。 我们找到了更多GPU。第二轮。
Qwen-Image-2.0是一个统一高保真生成与精确编辑的全能图像生成基础模型。它采用Qwen3-VL作为条件编码器,结合多模态扩散变换器进行联合建模,并通过大规模数据整理与多阶段训练实现强化。该模型支持长达1K令牌的指令输入,能生成幻灯片、海报等富文本内容,显著提升多语言文本渲染与排版质量。在生成方面,它增强了细节、纹理真实感与光照一致性,并更可靠遵循复杂指令。人工评估表明,其在生成和编辑任务上均大幅超越前代模型。
MachinaCheck是一款基于多智能体AI的系统,旨在革新小型CNC机加工车间的报价分析流程。传统上,车间经理需花费30-60分钟手动分析图纸,而该系统在上传STEP文件及材料、公差等简单输入后,能在30秒内生成完整的可制造性报告,明确指出零件能否制造、所需工具及生产前需采取的行动。其核心在AMD MI300X加速卡上本地运行Qwen 2.5 7B模型,利用192GB HBM3显存确保客户设计数据无需离开本地,满足了制造业对数据隐私的严格要求。系统采用五组件流水线,结合精确的几何特征提取与LLM的制造知识推理,最终输出结构化报告。
该项目使用AMD Instinct MI300X(192 GB HBM3显存)和ROCm,通过LoRA微调Qwen3-1.7B模型实现医学问答。训练仅用2000条MedMCQA样本,约5分钟完成,仅更新约220万参数(占模型总参数的0.1443%),全程采用fp16精度,无需量化。HuggingFace生态(Transformers、PEFT、TRL、Accelerate)在ROCm上无缝运行,无需修改代码即可直接替代CUDA。模型已上传至HuggingFace Hub并提供在线Demo。
Proprioceptive AI开发的Cygnus技术,通过为冻结的大语言模型添加自感知适配器,使其能读取内部认知几何。该技术将模型的隐藏状态投影到由gl(4,R)李代数定义的数学空间,分离出包含主要精度信号的“暗模式”,从而无需重新训练即可显著提升模型性能。例如,仅用一张RTX 3090显卡,就将Qwen-32B在ARC-Challenge基准上的准确率从82.2%提升至94.97%。其适配器将覆盖从3B到405B的多款模型,服务节点可支持5万用户并发,预计本周末上线。相关设计论文已公开。