GPTProto

Speech & Audio

108 picksNewest firstLatest article Oct 2, 2026, 4:35 AM
1 story
1 story
1 story
1 story
3 stories
1 story
1 story
2 stories
1 story
1 story
2 stories
1 story
2 stories
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
2 stories
2 stories
2 stories
08:00 xAI:News(网页)AI score 77/100

Introducing the Voice Agent Builder Jul 1, 2026 # Introducing the Voice Agent Builder Create a personalized voice agent in under 2 minutes without a single line of code. Read More

xAI 推出 Voice Agent Builder 测试版,这是一个基于 Grok Voice 的无代码平台,可在两分钟内创建生产级语音智能体。它集成电话、知识检索、工具、MCP、Guardrails 及可观测性,支持连接现有 SIP 号码、API 和 WebSocket,采用语音到语音路径。在 τ-voice Bench 上,Grok Voice Think Fast 1.0 得分 67.3%,领先 Gemini 3.1 Flash Live(43.8%)和 GPT Realtime 1.5(35.3%)。定价为每分钟音频 0.05 美元、电话费 0.01 美元,提供 80+ 种语音及声音克隆,每个账户附赠一个免费电话号码。

00:59 Apple:Newsroom(RSS)AI score 66/100

Apple Creator Studio gets smarter, faster, and more connected

Apple Creator Studio 推出多项 AI 增强更新。Final Cut Pro 新增 on-device AI 驱动的 Generate Captions(自动转录音频生成字幕)和 Edit Detection(自动检测剪辑点)。Mac 版加入 Auto Mask(自动识别皮肤、天空等主体)、增强的 Match Color 和 Advanced Trimming。支持将帧发送至 Pixelmator Pro 编辑,并在 Keynote、Pages、Numbers 中直接调用 Pixelmator Pro 修改图片。Logic Pro 新增 Grammy 制作人制作的 Producer Project 及 Chord ID 改进。订阅价 $12.99/月或 $129/年,新用户免费试用一个月,教育用户 $2.99/月。

1 story
1 story
1 story
1 story
2 stories
1 story
1 story
1 story
2 stories
20:59 Hugging Face:Blog(RSS)AI score 75/100

How to Fine-Tune Nemotron 3.5 ASR for Your Language, Domain, or Accent

Nemotron 3.5 ASR 是一个 600M 参数的多语言流式语音识别模型,单个检查点覆盖 40 种语言-地区(含英、西、德、法、意、日、韩、中、阿拉伯等)。采用 Cache-Aware FastConformer 编码器与 RNNT 解码器,缓存内部状态避免重复计算,实现低延迟流式转录且不损失精度。模型原生输出带标点和大写的生产级文本,无需后处理。支持指定语言(target_lang=es-ES)或自动语言检测(target_lang=auto)。通过注意力上下文大小(att_context_size)可在推理时直接调节延迟-准确率权衡,范围从 80ms 到 1.12s,无需重新训练。模型以 NeMo 检查点形式发布,可用于微调以适配特定语言、领域或口音。

00:00 LMSYS:Blog(Chatbot Arena 团队)AI score 77/100

Blog Higgs Audio v3 TTS on SGLang-Omni: Real-Time, Controllable Speech for Voice Agents Today we are announcing end-to-end serving for Higgs Audio v3 TTS on SGLang-Omni. Higgs Audio v3 TTS is Boson AI's text-to-speech model for conversational voice agents: it generates natural and expres... Boson AI & SGLang-Omni Team

Boson AI 与 LMSYS 联合推出基于 SGLang-Omni 推理框架的 Higgs Audio v3 TTS 端到端服务。该模型约 4B 参数,基于 Qwen3-4B 骨干,支持 100 种语言(内部评测覆盖 111 种),在 Seed-TTS、CV3、MiniMax-Multilingual 及 Higgs-Multilingual 零样本语音克隆任务中达到单字级 WER/CER。开发者可通过文本内控制标签实时调整情感(20+种)、风格、韵律(语速/音高/停顿)及音效。模型支持流式合成,文本未完整时即可开始生成语音并保持一致性。SGLang-Omni 专为多阶段生成模型设计,统一调度 AR 解码与轻量计算,实现低延迟推理。

1 story
2 stories
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
2 stories
1 story
2 stories
1 story
1 story
1 story
2 stories
2 stories
1 story
1 story
2 stories
1 story
1 story
1 story
6 stories
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
1 story
Showing the latest 100 of 108 stories.