Blog Running DeepSeek-V4-Flash and Kimi-K3 on Consumer Hardware with SSD Expert Pack SGLang brings the core idea of SSD-LLaMA to MoE inference: keep routed experts that do not fit in VRAM and host RAM on an NVMe SSD, load only the experts selected by the router, and use Expert Pack la... WiCi AI Team, SGLang Team
SGLang 将 SSD-LLaMA 的思路引入 MoE 推理,把放不进显存和内存的路由专家放在 NVMe SSD 上,通过 Expert Pack 连续布局、O_DIRECT 直接 I/O 和带预算的 GPU LFU/LRU 缓存,只加载 router 选中的专家。