Blog Pushing the Limits of Serving DeepSeek-V4-Pro DeepSeek-V4-Pro is a 1.6-trillion-parameter Mixture-of-Experts (MoE) model released with both FP8 and FP4 weights. Models at this scale naturally benefit from accelerators such as NVIDIA Blackwell GPU... Tianyu Zhang, Yusong Gao, Yun Zhang
LMSYS 团队针对 1.6 万亿参数的 MoE 模型 DeepSeek-V4-Pro,在 H20 GPU 上通过场景化服务配置逼近 B300 性能。单节点 H20-141GB 参考实现达 271 output tokens/s,与 B300 的 383.7 tokens/s 性能差距缩小至 1.42×。