GPTProto

多模态

305 picksNewest firstLatest article Sep 29, 2026, 10:00 PM
2 stories
1 story
1 story
3 stories
1 story
2 stories
1 story
3 stories
1 story
1 story
5 stories
2 stories
2 stories
1 story
2 stories
1 story
3 stories
1 story
2 stories
2 stories
2 stories
2 stories
3 stories
19:51 LMSYS:Blog(Chatbot Arena 团队)AI score 72/100

Blog SGLang Adds Day-0 Support for Muse Glimmer, a Multimodal Model Built for Local Agentic Workflows We're excited to partner with Meta Superintelligence Labs to bring Day-0 support for Muse Glimmer to SGLang, with dedicated optimizations tailored for high-performance inference of agentic workflows o... Meta Superintelligence Labs and the SGLang Team

SGLang 与 Meta Superintelligence Labs 合作,为 30B 参数多模态模型 Muse Glimmer 提供 Day-0 支持,该模型拥有 128k+ token 上下文窗口。

07:58 MarkTechPost(RSS)AI score 75/100

NVIDIA Releases NemotronLabs VoiceChat 11B: An Open Full-Duplex Speech-to-Speech Model with ~450 ms Turn-Taking and Live Tool Calling

NVIDIA 发布开源端到端全双工语音对话模型 NemotronLabs VoiceChat 11B,在统一网络中完成流式语音理解与生成,实测轮换延迟 448 毫秒。该模型为首个支持对话中工具调用的开源全双工模型,通过独立输出通道及预置“保持”话术避免 API 执行期间冷场。权重与容器已公开,但仅限研究用途,需单张 80 GB 显存 GPU,目前无托管 API。

2 stories
2 stories
4 stories
5 stories
5 stories
08:00 HuggingFace Daily Papers(社区热门论文)AI score 70/100

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

SwanTale 提出统一的多说话人语音与音频生成模型,同时支持零样本与指令任务。研究配套推出 SwanData-Caption 数据方案,通过清洗、合成覆盖与多级标注解决数据稀缺问题,并引入 SwanVAE、Unified MoE、GRPO 后训练等技术。实验显示,SwanTale 在多项零样本与指令指标上领先,并在两项任务的表达力评分中均取得最佳成绩。

1 story
3 stories
17:59 MiniMax:Blog(网页)AI score 78/100

MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities

MiniMax 正式推出全能多模态生成模型 H3,可联合理解文本、图像、视频和音频,生成最高 2K 分辨率、15 秒时长且带原生立体声的视频。H3 在指令跟随、文字与品牌呈现、V2V 动作迁移上表现突出,2K 下每秒价格低于主流模型三分之一,768p 下低于主流 720p 价格一半。官方计划近日开源模型权重,以支持开源社区并加速硬件兼容。

01:47 X:Thinking Machines (@thinkymachines)AI score 64/100

Today, we are releasing Inkling-Small. Inkling-Small achieves comparable performance to Inkling at ...

今天,我们发布 Inkling-Small。 Inkling-Small 在仅为 Inkling 四分之一规模的情况下,实现了与之相当的性能。它拥有 276B 总参数,12B 激活参数。我们将开放完整权重。 https://thinkingmachines.ai/news/inkling-small/ 现在即可在 Tinker 上对其进行微调,或在 Tinker Playground 中以文本、图像和音频形式与之对话。

3 stories
1 story
1 story
1 story
1 story
3 stories
1 story
1 story
3 stories
2 stories
4 stories
00:51 Google Developers Blog(RSS)AI score 58/100

Unlocking the Next Era of On-Device AI with Google Tensor and Pixel

在 Google I/O Connect India 上,Google 展示了由定制 Tensor SoC 和 TPU 驱动的 Pixel 10 系列所支持的 100% 私有端侧 AI 未来。活动首次推出轻量级 Gemma 4 E2B 模型,该模型原生运行于设备端,可实现完全离线的多模态功能,包括 AI 聊天、实时图像识别和个人智能体任务。开发者即日起可通过新发布的 Tensor SDK beta 及其配套开源资源,开始构建这些安全的边缘应用。

2 stories
1 story
3 stories
15:56 MarkTechPost(RSS)AI score 74/100

Ant Group’s Robbyant Unveils LingBot-VA 2.0: A Causal Video-Action Model Built Natively for Physical AI

蚂蚁集团旗下具身智能团队 Robbyant 发布 LingBot-VA 2.0,首个原生具身基础模型。该模型采用因果 DiT 架构,视频专家约 13.0B 参数(约 1.9B 激活),训练规模约 15.3B 参数,推理时每 token 约 2.5B 激活。模型引入多块预测(MCP)实现 2.3 倍训练加速,并通过前瞻推理将推理延迟降至 142 ms/chunk。在 RoboTwin 2.0 的 50 个任务上,干净与随机演示数据平均成功率分别达 93.8% 和 93.4%。

08:00 HuggingFace Daily Papers(社区热门论文)AI score 70/100

GigaChat Audio: Time-aware Large Audio Language Model

GigaChat Audio 是一种时间感知的音频大语言模型,支持长达120分钟的输入,并能生成带有明确时间戳的问答、片段描述和摘要。该模型通过将周期性时间标记与连续音频 token 交错,并利用级联合成数据管道进行大规模训练,在短时和长时基准上均实现了强时间定位精度。研究团队已开源模型权重及超过1万小时的时序数据集。

2 stories
3 stories
2 stories
Showing the latest 100 of 305 stories.