Open-sourcing AstaBrief, the fast report-generation model in Asta
Ai2 开源 AstaBrief 8B,一个基于 Qwen3-8B、将研究问题和检索文献片段转化为带引用报告的科学报告生成模型,现已在 Asta 的 Generate a report 功能中作为 Fast mode 上线,并连同训练数据开放下载。
Ai2 开源 AstaBrief 8B,一个基于 Qwen3-8B、将研究问题和检索文献片段转化为带引用报告的科学报告生成模型,现已在 Asta 的 Generate a report 功能中作为 Fast mode 上线,并连同训练数据开放下载。
Oct 2, 2026 Black Forest Labs has released Flux 3 Image, the image side of its Flux 3 model family. The model supports multi-step edits without changing other parts of the image, BFL claims, and covers text-to-image, image-to-image, text rendering, and photorealism. Users can compose scenes with bounding boxes, include up to ten reference images, and output up to 4K. A free demo is available here. API access is 50 percent off through October 8. Companies can license commercial weights to run and
AWS Strands Labs releases Strands Decider 2B, an open source decision model. It does not generate text. It reads a state and typed questions, then returns a choice, a yes/no probability, or a score with a calibrated confidence. The model has 1.9 billion parameters and runs locally on a CPU, a consumer GPU, or an Apple silicon Mac. Is it deployable? Yes, for local and self-hosted use. Weights are on Hugging Face under Apache-2.0, and pip install strands-decider gives a CLI and an HTTP server. The
Arena 宣布 Xiaomi MiMo-V2.6-Pro 和 MiMo-V2.6-Flash 登陆 Agent Arena。Pro 在 8.1K+ 真实智能体会话中净提升 +3.17%,列开源模型第5,较 MiMo-V2.5-Pro(第13,-7.23%)提升9个名次;其 Confirmed Success 得分 +7.35%,列开源模型第2。
Oct 1, 2026 Ideogram says its new model Ideogram 4.5 solves one of AI image editing's biggest headaches. When you edit part of an image, the rest shouldn't change. Swap someone's outfit, and their body shape and background should stay intact. Leading models like GPT-Image 2.5 and Nano Banana have gotten much better at this, but they still tend to produce artifacts, especially after multiple edits. Ideogram 4.5 only touches what the user tells it to, the company claims. Target use cases include p
Amazon Web Services released an open source decision model inspired by TypeSafe’s Jev, with AI developers increasingly seeking intelligence that is more suited to computer automation than frontier LLMs. Amazon’s Strands Decider 2B, released the same week OpenAI announced a similar offering, is a high-speed, low-cost way to sort between pre-decided options and deliver a measure of how confident it is in its choice. The model is fully open sourced, available now, and small enough to run locally. A
Today we’re releasing Olmo-core 3, a significant upgrade to our framework for developing large language models featuring a redesigned open mixture-of-experts (MoE) training system.Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency. It’s one of the core systems behind the next generation of Olmo, and part of our ongoing commitment to open up the tools and training infrastructure behind each new model.Training large AI models t
NVIDIA has released Kumo Tabular, a new family of tabular foundation models (TFMs) for classification and regression. If you have followed TabPFN or TabICL, the setup will look familiar. The model takes labeled rows as context and predicts new rows in one forward pass. There is no training, no hyperparameter tuning, and no feature engineering. Kumo Tabular comes in Small, Medium, and Large versions, spanning about 28M to 215M parameters. It runs through NVIDIA’s open-source structured-data-model
Perplexity Research and turbopuffer have released pplx-embed-v2-context-9b-preview, a contextual embedding model for RAG pipelines. Each chunk is embedded with the full document in view. The real change is the training signal. The model learns to retrieve the answer along with the context needed to verify it, not one ‘gold passage.’ PIs it deployable? Yes, as a self-hosted preview. Weights are on Hugging Face under the MIT license. Loading requires transformers>=5.4.0 with trust_remote_code=True
Artificial Analysis 数据显示,GPT-6.1 Sol 的 Cost per Task 约 $0.72,比 GPT-6 Sol($1.05)低约 30%,后者已约为 GPT-5.6 Sol($1.99)的一半。
Sep 30, 2026 | Gemini 4 Argon delivers frontier performance in complex workflows across real-world software engineering, enterprise knowledge work like legal and finance, and cybersecurity defense. In this article Today, we’re announcing our new frontier model, Gemini 4 Argon, which is rolling out to a set of trusted cyber defenders through our Fairwind Program. Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Go
Introducing Cohere Embed 5: our new state-of-the-art family of embeddings models.Get frontier capabilities with Embed 5 Pro or low-latency performance with Embed 5 Fast.
Liquid AI has released d1, a decision model built for structured choices instead of text generation. You give it context and a set of typed questions. It returns calibrated probabilities across a fixed set of outcomes in a single call, with zero generated tokens. The target is the work many teams still send to general LLMs: classification, ticket routing, scoring, moderation, reranking and LLM-as-judge checks. Is it deployable? Yes, today, as a hosted API. d1 runs on the Liquid API under the mod
Near-Astra intelligence for a fifth of the price We’re introducing GPT‑6.1 Sol, an upgrade to GPT‑6 Sol that nearly matches GPT‑6 Astra’s intelligence on agentic coding, computer use, and professional work at one-fifth of Astra’s standard input and output token prices. Cached input costs just $0.10 per million tokens—95% less than standard input pricing and 50% less than GPT‑6 Sol’s cached input pricing—giving developers more room to build and run capable agents that reuse context across request
Korean AI Lab 🇰🇷 Upstage has released Solar Mini 4 which scores 24 on the Artificial Analysis Intelligence Index, but costs ~5x as much per task as GPT-6 Luna (max) despite similar per-token pricesSee model page Upstage has released Solar Mini 4, a new proprietary reasoning model. Upstage reports 35B total and 3B active parameters, setting a new Pareto optimal point on Intelligence Index vs. Active Parameters for models under 3B active parameters. It also scores 16 points higher than Upstage's
See model page Google’s new Gemini 4 Argon equals GPT-6 Astra on the Artificial Analysis Intelligence Index at 60% of the Cost per Task with discounted prices Gemini 4 Argon is Google DeepMind’s first proprietary model above the Flash class in over 7 months. With high reasoning (the highest available), it scores 53 on the Artificial Analysis Intelligence Index, matching GPT-6 Astra (max, 53) and 1 point ahead of GPT-6.1 Sol (max, 52), with gains driven by lower hallucinations and stronger agenti
Highlights (TL;DR) NVIDIA Kumo Tabular, part of the NVIDIA Kumo Structured model collection, is an open foundation model for tabular data now available on Hugging Face. Given a table of labeled rows, it predicts the labels of new rows in a single forward pass, with no training, no tuning, and no feature engineering, for both classification and regression. It was pretrained only on artificial data, comes in three sizes (28M to 215M parameters), runs through our open-source library, and is release
Elevenlabs is releasing Eleven v4, a new speech model that follows direction cues more accurately and keeps voices consistent across long productions. A new model architecture also powers the Turbo variant for real-time voice agents. Eleven v4 generates laughter, whispers, and sounds like slamming doors more reliably than its predecessor. The Turbo variant for voice agents starts producing speech in about 150 milliseconds. Eleven v3, released just over a year ago, already supported these audio t
DevDay 2026 is our biggest yet, with more than 20 major announcements across ChatGPT, Codex, our models, and entirely new forms of working with AI. We believe AI can help bring about a new renaissance of creativity and discovery. It should give people more time for what matters to them, more freedom to pursue their ideas, and the ability to do things they didn’t think were possible. Today, we introduced agents that can take on ongoing responsibilities and new ways for people and AI to work toget
H Company has released Holo4, a family of generalist computer-use models for AI agents. One set of weights clicks and types on screens. It also writes code and calls MCP or API tools. Holo4 ships in 2 sizes: Holo4 27B (dense) and Holo4 35B-A3B (Mixture of Experts, 3B active). Both serve a 256K context on the H Models API. Is it deployable? Yes. Holo4 35B-A3B ships Apache 2.0 weights for commercial self-hosting. Holo4 27B weights are CC BY-NC 4.0, so commercial use of 27B runs through the H Model
Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work.Sonnet 5.5 is a faster, lower-cost complement to Claude Opus 5.5. Where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. It’s also got a sharp eye for design. Claude Haiku 5.5, built fo
A line of text can change significantly depending on how it’s spoken. “I need you to stay calm" should sound different depending on who's saying it, whether that's a doctor delivering it gently to a frightened patient, or a character in a game shouting to his squad before dropping into battle.Today we're launching Eleven v4, our most emotive text-to-speech model yet, and its low-latency variant, Eleven v4 Turbo.Ranked #1 by Artificial Analysis1, and preferred by ~75% of listeners in blind head-t
Claude Sonnet 5.5 is our second model in the Claude 5.5 family after Opus 5.5. It's a clear upgrade over Sonnet 5 and is smarter, more efficient and 30% faster. The per-token price is unchanged and because Sonnet 5.5 typically needs far fewer tokens to do the same work, it costs up to 30% less for most work. FIG ALance Martin's code-to-painting demo: each model writes code that repaints the same photograph. From left: the photograph, Claude Sonnet 5, Claude Sonnet 5.5 and Claude Opus 5.5. Credit
Holo4 is our new series of agentic models. It comes in two sizes: 27B dense and 35B-A3B Mixture of Experts. Both are available on the H Models API. We are also releasing an updated version of Holotron 3: Holotron4 Nano. Holo4 builds on our previous model and interacts with software through any available interface: GUIs, code, MCP and APIs. It scores well on academic benchmarks, but we built it for real business workflows. It was trained through supervised and reinforcement learning on a large se
The Claude Platform release notes list changes to the Claude API, the client SDKs, and the Claude Console, newest first. September 30, 2026 We announced the deprecation of the Claude Sonnet 4.5 model (claude-sonnet-4-5-20250929), with retirement on the Claude API scheduled for November 30, 2026. We recommend migrating to Claude Sonnet 5.5. Read more in Model deprecations. September 28, 2026 We've launched Claude Sonnet 5.5 (claude-sonnet-5-5). It's available on the Claude API, Claude in Amazon B
Meta 发布 Muse Realtime Avatar,将 Muse Realtime Voice 的语音流实时转化为带表情、手势和全身动作的交互化身,可基于参考图像驱动人像、插画、动物或日常物品。
OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna,将 GPT-6 Astra 的训练方法用于更快更便宜的模型,API 价格较 GPT-5.6 促销价下调 50%(Sol 输入 $4→$2、输出 $20→$10;Luna 输入 $0.20→$0.10、输出 $1.20→$0.50,每百万 token)。
Fireworks Research 发布基于 Kimi K3 训练的专用模型 Ember-1,通过学习精简不必要的推理,在保持质量的同时减少约 35-50% 的 reasoning token。
Qwen 发布 Qwen3.8-LiveTranslate,采用 Interleave 架构与 Hybrid-MoE Thinker–Talker 设计重构实时同声传译,平均滞后(LAAL)从上一代的 2.8 秒降至 2.3 秒。
Qwen 发布下一代原生全模态模型 Qwen3.8-Omni-Flash,支持文本、图像、音频和视频输入及 1M token 上下文窗口,29 项评测平均分较 Qwen3.5-Omni-Plus 提升超过 25%,音频输入每小时价格下降超过 98%,音视频输入每小时价格下降超过 93%。
WorkBuddy 宣布 DeepSeek V4.1-Flash 已在其平台上线,免费试用两周。引用的 DeepSeek 公告称 V4.1-Flash 已登陆 DeepSeek API 并支持原生多模态。
DeepSeek AI 发布多模态 MoE 模型 DeepSeek-V4.1-Flash,552B 主干加 196B Engram 参数,上下文窗口 1M,prefill 激活 8B 参数、decode 激活 16B,全局 KV 缓存降至每 token 890 字节,约为 DeepSeek-V4-Flash 的 1/4、DeepSeek-V1 的 1/437。
Suno 发布 v6 模型,可将图片、视频和语音备忘录转化为音乐,并对已创建的歌曲进行精确修改。同时提供 v6-wild 版本供探索更多可能性,官方附有 2 分钟以内的功能演示视频。
Cognition 发布最先进编码模型 SWE-2,基于 2.8T 参数的 Kimi K33 后训练,在 FrontierCode 1.1 Main 取得 50.0%,仅比 Fable 5.1 低不到一分且便宜 64%。
Microsoft AI 发布转录模型 MAI-Transcribe-2,称其在 FLEURS 基准 60 种语言上平均词错率 5.2% 排名第一,并占据 Artificial Analysis 准确率-延迟 Pareto 前沿。
OpenAI 发布 GPT-6 Astra 系统卡,将其在生物化学领域预防性地定为 High 能力,未达 Critical 阈值。
通义千问发布 Qwen3.8-Max-0902,在 Code Arena: WebDev 以 1,691 分首次亮相即排名总榜第一,并以混合价 $5/MToken 成为 Pareto 前沿上得分最高的模型,现已可在 QwenCloud 试用。
Meta 发布 Muse Spark 1.3,在智能体和编码任务上性能提升,max reasoning 版本现已在 Muse Code 和 Meta Model API 开放。
Google DeepMind 发布 Gemini 3.8 Flash,在 Artificial Analysis Intelligence Index 上以 high 推理得分 59,较 Gemini 3.7 Flash 提升 3 分。
OpenAI 宣布 Astra 在其 Preparedness Framework 下达到 Critical 网络安全能力阈值,是首个被评定为该级别的模型,可在少人干预下发现未知漏洞并构建利用链。
上海AI实验室发布 InternLumina-U2,一个 16B 参数(激活 1B)的 MoE 扩散语言模型,用 8 码本全离散视觉表示统一覆盖文本问答、文生图、图像理解与编辑、视频及 3D 理解。
Meta Superintelligence Labs 发布 Muse Voice Transcribe,称其为首个实时音频感知模型,提供流式 ASR、最多 20+ 说话人的 diarization 和 endpointing。
Anthropic 发布 Claude Fable 5.1(claude-fable-5-1),面向长时间运行的智能体编码、知识工作与研究,Claude Mythos 5.1 面向 Project Glasswing 参与者。
Anthropic 发布 Claude Fable 5.1 和 Claude Mythos 5.1,两者为同一模型但安全防护级别不同,Mythos 5.1 仅通过可信访问计划提供给网络安全和生命科学领域的受审机构。
GLM-5.3 现已开放权重。 我们最强大的智能体编码与网络防御模型,现已可供下载、运行和定制。 权重:https://huggingface.co/zai-org/GLM-5.3 技术博客:https://z.ai/blog/glm-5.3
Midjourney 开始向所有用户开放其首个 V8.2 图像编辑模型的测试。该模型支持指令编辑、以图生图(最多同时引用 4 张参考图)、局部重绘与扩画,并兼容个性化、moodboards 和 srefs 功能。用户可通过网页端或 Discord 的 `--edit` 命令使用,官方同步更新了 midjourney.com 与 alpha.midjourney.com 的界面。
Google 推出 Gemini Omni 1.1 Flash,为开发者提供更强的生成式视频控制能力。新模型支持场景扩展(可分析最多 10 秒先前上下文,以 10 秒为增量累计延长至 40 秒)、指定首尾帧生成平滑过渡,以及 4K 高清输出。
Ox Alpha = GLM-5.3 Flash AA = 57,以1/100的前沿价格, 由纯国产芯片驱动。在OpenRouter上实现了近20%的周token份额(第一)。 感谢大家的支持。
Google DeepMind 推出 Gemini 3.5 Transcribe 语音转文本模型,支持流式与非流式两种 API。据 Artificial Analysis 评测,其流式与非流式平均词错率分别为 4.0% 和 2.6%,支持超 85 种语言、自定义词汇及最多三人说话人识别。
Unsloth 推出 Qwen3.8-Flash-Next 的 GGUF 量化版本,称该 125B MoE 多模态模型性能超过 Claude-Opus-4.6 (Max),可在 75GB RAM 上本地运行,无需 GPU VRAM。
通义千问发布 Qwen3.8-Flash,一款多模态 MoE 模型,作为 Qwen4 架构的早期预览并开放权重。该模型总参数 125B,每 token 仅激活 6B,训练成本仅为 Qwen3.7-Plus 的 1/9,性能全面超越后者。生产版 API 定价 $0.16/1M 输入 tokens 和 $0.47/1M 输出 tokens,原生上下文 262K,可扩展至 1M。
通义千问开源 Qwen3.8-Flash-Next,一款多模态 MoE 模型,也是 Qwen4 架构的早期预览。该模型采用 GDN + QSA 混合注意力等四项升级,总参数 125B,每 token 激活 6B,训练成本约为 Qwen3.7-Plus 的 1/9,编码与办公任务能力更强。
面壁智能 OpenBMB 推出 MathForm,一个面向 Lean 4 数学自动形式化的开源框架、数据集与模型。其 FormalVerse 数据集含 367K+ 已验证示例;在匹配 100K 预算下,基于其训练的模型 Consistency Check 达 60.32%,优于 FineLeanCorpus(46.53%)与 NuminaMath-LEAN(41.49%)。
Hugging Face 发布 LFM2.5 系列三款模型的 DSpark 草稿模型检查点,通过投机解码在不改变输出质量的前提下,GPU 吞吐最高提升 3.18 倍,端侧最高 2.87 倍。草稿模型约 300M 参数,LFM2.5-2.6B 函数调用延迟平均降低 57%,已开源支持 llama.cpp 和 SGLang。
Unsloth 发布 Qwen3.8-27B 的 Dynamic GGUF 量化版本,4-bit 量化可在 17-19GB 内存设备上本地运行,并同时上传 NVFP4 量化。
智谱发布 GLM-5.3,为网络安全任务最强模型,在漏洞发现、利用分析和多步安全任务上大幅提升。CyberGym 得分 84.5%(GLM-5.2 为 77.2%),ExploitBench 达 54.4%(此前 24.4%),ExploitGym 两小时完成 105 项任务(此前 29 项)。因双重用途风险,将先由安全伙伴受控评估,再开放 API 和完整权重。
Google DeepMind 发布 Gemini 3.7 Flash,距 3.6 Flash 仅三周,主打编程与智能体任务,输入/输出价格分别为每百万 token $0.75 和 $3.75,为原 3.6 Flash 的一半。
FireRedTeam 发布 FireRedTTS3,包含 Base 和 Instruct 两个变体,基于语义增强连续语音表征实现统一语音生成与编辑。
MiniMax 推出 Music 3.0,新一代音乐生成模型,可根据创意概念和可选歌词一次性完成整首歌的作曲、编曲、演奏与制作,最长支持五分钟。
Grok 4.6 现已推出 🚀🚀🚀 智能、快速,性价比惊人!
vLLM 宣布对 Qwen3.8-2.4T-A95B 提供 Day-0 支持,该模型是基于 Qwen 3.5 架构的 2.4 万亿参数稀疏 MoE 模型,含 512 个专家,92 层混合骨干中每 4 层使用一次 full attention,其余 69 层为线性注意力。
Cursor 与 SpaceXAI 今日发布 Grok 4.6,重点强化长时运行智能体与交互式视觉任务,在多项智能体编程与知识工作基准上达到前沿水平,并在 Artificial Analysis Intelligence Index 上追平 GPT-5.6 Sol。
xAI 今日发布 Grok 4.6,在 Grok 4.5 基础上重点强化长时运行智能体及更复杂的交互式与视觉工作能力。该模型在多项智能体编码与知识工作基准上达到前沿水平,在 Artificial Analysis Intelligence Index(九项基准综合分)上追平 GPT-5.6 Sol。
Liquid AI 发布 LFM2.5-VL-3B 视觉语言模型,屏幕理解 ScreenSpot-v2 平均 80.7,ToolSandbox 从 26.4 升至 59.5,RefCOCO 从 57.1 升至 87.9。
SGLang 宣布对 NVIDIA Nemotron 3.5 Lightning 提供 Day-0 支持,该开源模型为 30B 总参数、3B 激活参数的混合专家架构,支持最长 1M token 上下文,可从 Hugging Face 下载 BF16 和 NVFP4 权重。模型支持 MTP、DFlash、DSpark 三种投机解码技术,并可通过 OpenAI 兼容 API 接入智能体工作流。
NVIDIA 发布 Nemotron 3.5 Lightning,一款可定制的开源 30B 混合专家(MoE)模型,专为常驻智能体设计。相比同类开源模型,其 token 生成速度最高提升 4 倍,任务完成时间缩短 30%。该模型采用开放权重,支持用户微调以匹配特定任务,并可在 RTX PC、DGX Spark 及 Jetson 等设备上运行。
蚂蚁 inclusionAI 开源 Ling-3.0 系列语言基础模型,含 tiny 与 flash 两个尺寸及预训练、中期训练、合并(WSM)等阶段检查点。该系列采用混合线性注意力与稀疏 MoE 架构,总参数 7.9B,每 token 仅激活 1.3B 参数(128 个路由专家中激活 8 个)。
蚂蚁 inclusionAI 开源 Ling-3.0 系列高效语言基础模型,并发布预训练、中期训练及合并(WSM)等阶段的检查点,支持继续预训练、微调与研究。Ling-3.0-tiny-base 采用稀疏 MoE 架构,含 128 个路由专家,每 token 仅激活 1.3B 参数,总参数量 7.9B,并原生混合线性注意力以高效处理长上下文。
蚂蚁 inclusionAI 开源 Ling-3.0 系列语言基座模型,采用高度稀疏(1/64)MoE 架构,含 512 个路由专家,每个 token 仅激活 8 个路由专家和 1 个共享专家,激活参数仅 5.1B(Non-emb)。
蚂蚁 inclusionAI 开源 Ling-3.0 系列语言基础模型,并发布预训练、中期训练及合并(WSM)等多个阶段检查点。该系列采用高度稀疏(1/64)MoE 架构,512 个路由专家中每 token 仅激活 8 个,激活参数仅 5.1B(Non-emb),总参数 124B,并原生结合 KDA 与 Gated MLA 混合线性注意力以高效处理长上下文。
蚂蚁 inclusionAI 开源 Ling-3.0 系列语言基座模型,采用高度稀疏(1/64)MoE 架构,512 个路由专家中每 token 仅激活 8 个,总参数量 124B,激活参数仅 5.1B。该系列原生融合线性注意力,并以加权检查点合并替代传统学习率衰减。此次发布包含预训练、中期训练及合并(WSM)等多个训练阶段的检查点,支持继续预训练与微调,模型采用 MIT 许可证。
蚂蚁 inclusionAI 开源 Ling-3.0 系列高效语言基础模型,并发布预训练、中间训练及合并(WSM)等多个训练阶段检查点,支持继续预训练、微调与研究。
Inception Labs 发布面向搜索管线的 Mercury 2,扩散式并行解码超 1000 tokens/秒,在 WideSearch 各步骤比 Gemini 3.1 Flash Lite 快近 2 倍、比 GPT-5 Mini 快 10 倍,定价 $0.25/M 输入、$0.75/M 输出。
Microsoft 发布 MAI-Image-2.6,在 Arena 文生图排行榜上位列第二,超过 Meta、Google 和 xAI 的模型。该模型相较 MAI-Image-2.5 整体提升 +79 Elo,文本渲染单项提升 +91 Elo,各测量类别均有改进;现已可在 Arena 体验,本周后续登陆 MAI Playground,并将逐步推广到 Microsoft Foundry 等产品。
SGLang 与 Meta Superintelligence Labs 合作,为 30B 参数多模态模型 Muse Glimmer 提供 Day-0 支持,该模型拥有 128k+ token 上下文窗口。
OpenAI 发布网络安全专用模型 GPT-5.6-Cyber,可通过 Daybreak Red 获取,用于授权的漏洞研究、漏洞验证和安全测试。该模型旨在应对网络防御窗口不断收窄的挑战,为安全研究人员提供专门工具。
NVIDIA 发布开源端到端全双工语音对话模型 NemotronLabs VoiceChat 11B,在统一网络中完成流式语音理解与生成,实测轮换延迟 448 毫秒。该模型为首个支持对话中工具调用的开源全双工模型,通过独立输出通道及预置“保持”话术避免 API 执行期间冷场。权重与容器已公开,但仅限研究用途,需单张 80 GB 显存 GPU,目前无托管 API。
Meta 的 Muse Glimmer 30B 稠密模型已在 Fireworks 上提供 serverless 和按需部署,主打多轮工具调用与失败恢复等常驻 Agent 场景。
NVIDIA 发布 Cosmos 3,一个基于混合 Transformer 架构的开放物理 AI 基础全模态模型,整合视觉推理、世界生成与动作预测。
ChatGPT 推出改进版 GPT-5.6 Sol,提升准确性与一致性,同时扩大免费用户访问权限。免费用户还可无限次使用 GPT-5.6 Luna 进行日常对话。
NVIDIA 发布 Alpamayo 2 Super,一款 34B 参数的视觉-语言-动作(VLA)模型,专为自动驾驶长尾事件设计,权重采用 Linux 基金会 OpenMDW-1.1 许可,代码为 Apache 2.0,发布首日即可商用。
@bfl_ai 的 FLUX 3 Video 现已在 OpenRouter 上向所有人开放。 一个统一的视频、音频、图像和动作预测多模态模型家族。严肃、有趣、创意、真实、电影感,随你所需。基于统一架构联合训练。
商汤发布开源模型 SenseNova U1,可在统一流程中同时进行推理与图像生成。其信息图模式可将单条提示词转为结构化幻灯片,交错模式则逐步生成图文内容,如演示六步画龙教程。模型已上线 SenseNova Studio、HuggingFace 及 GitHub。
NVIDIA Alpamayo 2 Super 现已开放商用,基于 Cosmos 3 Super Reasoner 构建,采用强化学习后训练,支持轨迹预测、因果链推理、元动作、自动标注及视觉问答等多任务输出。