GPTProto

Blog The next generation of speculative decoding: DFlash and Spec V2 Using Modal and Z Lab's DFlash speculative decoding models with SGLang’s newly default Spec V2 engine, you can achieve state-of-the-art latencies for LLM inference serving. Our new, jointly-released D... Z Lab, Modal, and SGLang Teams

LMSYS:Blog(Chatbot Arena 团队)·Jun 15, 2026, 12:00 AM·Qwen / Alibaba

Z Lab、Modal 与 SGLang 团队联合发布 DFlash 投机解码模型和 SGLang 的默认 Spec V2 引擎。DFlash 采用块扩散+KV 注入并行生成整块 draft token,在 Qwen 3.5 397B-A17B(BF16)的 HumanEval 数据集上、并发 1 时吞吐量达到基线的 4.3