# DeepSeek-V4-Flash-DSpark 修复记录(GitHub Issue #47648) > 本文件记录 vLLM DSpark 在 DeepSeek-V4-Flash 上的启动失败问题、根因、修复 diff 以及验证结果,便于后续直接用于提交 PR。 --- ## 1. 问题概述 在 H200/SM90(以及同样走 NVIDIA 路径的 B200/SM120)上,使用 `--spec-method dspark` 部署 `DeepSeek-V4-Flash-DSpark` 时,模型初始化阶段会失败。主要表现为两类错误: 1. **DSpark draft 权重加载路径错配**:draft 权重(checkpoint 中以 `mtp.{i}.*` 命名)没有被正确加载到模型中。 2. **KV cache shape mismatch**:当启用 `fp8_ds_mla` layout 时,KV cache 的 per-token slot size 被错误计算为 512B,而实际需要 584B,导致 shape 断言失败。 关联 issue:[vllm-project/vllm#47648](https://github.com/vllm-project/vllm/issues/47648) --- ## 2. 环境信息 - vLLM 版本:`0.23.1rc1.dev788+gfa4321de3`(vllm-dspark wheel) - PyTorch:`2.11.0+cu129` - GPU:8× NVIDIA H200(SM90),TP=8 - 模型:`/data/models/DeepSeek-V4-Flash-DSpark` - 启动命令: ```bash vllm serve /data/models/DeepSeek-V4-Flash-DSpark \ --trust-remote-code \ --tensor-parallel-size 8 \ --kv-cache-dtype fp8 \ --block-size 256 \ --max-model-len auto \ --max-num-seqs 256 \ --tokenizer-mode deepseek_v4 \ --reasoning-parser deepseek_v4 \ --spec-method dspark \ --spec-model /data/models/DeepSeek-V4-Flash-DSpark \ --spec-tokens 5 \ --no-disable-hybrid-kv-cache-manager \ --disable-uvicorn-access-log \ --port 30004 ``` --- ## 3. 根因分析 ### 3.1 根因一:DSpark draft 权重加载路径错配 #### 位置 `vllm/models/deepseek_v4/nvidia/dspark.py`,`_remap_dspark_name` 方法。 #### 问题描述 DeepSeek-V4-Flash-DSpark 的 draft 权重是嵌在目标模型 checkpoint 中的,命名格式为 `mtp.{i}.*`,例如: ```text mtp.0.self_attn.wq_b.weight mtp.0.ffn.experts.0.w1.weight mtp.1.self_attn.wq_b.weight ... ``` `_remap_dspark_name` 负责把这些 checkpoint key 映射到 DSpark draft 模型内部的真实参数名。原来的实现: ```python return f"model.layers.{self.num_hidden_layers + stage}.{rest}" ``` 把 `mtp.0.*` 映射成了 `model.layers.{num_hidden_layers + 0}.*`(例如 `model.layers.64.*`)。 但实际上,DSpark draft 的 3 个 `DeepseekV4DecoderLayer` 是放在一个 `nn.ModuleList` 里的,PyTorch 真实注册的参数名只跟 `ModuleList` 的索引有关: ```text model.layers.0.* model.layers.1.* model.layers.2.* ``` `DeepseekV4DecoderLayer` 构造函数里传入的 `prefix=f"layers.{num_hidden_layers + i}"` 只是为了内部计算 `compress_ratio`(让 `layer_id >= num_hidden_layers` 时固定 `compress_ratio=1`),**不是真实的参数名前缀**。 因此,原来的映射导致所有 draft block 权重都找不到对应的模型参数,draft layer 实际上没有被正确加载。 #### 修复 把 `mtp.{i}` 的 block 权重映射到 `model.layers.{i}`: ```python def _remap_dspark_name(self, name: str) -> str | None: """Map a checkpoint ``mtp.{i}.*`` name to this model's parameter path. Returns None for non-mtp weights (owned by the target model). """ m = re.match(r"mtp\.(\d+)\.(.*)", name) if m is None: return None stage = int(m.group(1)) rest = m.group(2) # The confidence head is not wired into inference yet; drop its weights. if rest.startswith("confidence_head."): return None # Head-stack params live at model level (mtp.last), context combiner at # model level (mtp.0); everything else is a per-layer decoder block. head_prefixes = ( "norm.", "hc_head_fn", "hc_head_base", "hc_head_scale", "markov_head.", ) if rest.startswith(("main_proj.", "main_norm.")) or rest.startswith( head_prefixes ): return f"model.{rest}" # Draft layers live in a ModuleList, so their actual parameter names are # model.layers.{stage}.* even though the prefix passed to the decoder # layer is layers.{num_hidden_layers + stage} (for compress_ratio). return f"model.layers.{stage}.{rest}" ``` ### 3.2 根因二:DeepSeek-V4 KV cache shape mismatch #### 位置 `vllm/models/deepseek_v4/attention.py`,`DeepseekV4Attention.get_kv_cache_spec` 方法。 #### 问题描述 `DeepseekV4Attention.get_kv_cache_spec` 返回的 `MLAAttentionSpec` 没有设置 `kv_quant_mode`。在 `gpu_model_runner.py` 中,这个缺失导致 `cache_dtype_str` 被当成 `"auto"` 传给 `DeepseekV4FlashMLABackend.get_kv_cache_shape`,返回的 shape 是 `(num_blocks, block_size, 512)`。 但对于 `fp8_ds_mla` layout(UE8M0 block-scaled fp8,以 `uint8` 打包),每个 token 的 KV slot 实际大小是 584 字节,而不是 512。于是 KV cache 分配的空间不够,触发 shape mismatch / assert。 #### 修复 在 `MLAAttentionSpec` 中加上 `kv_quant_mode=get_kv_quant_mode(self.kv_cache_dtype)`: ```python def get_kv_cache_spec(self, vllm_config: VllmConfig) -> KVCacheSpec | None: if ( self.compress_ratio <= 1 ): # SWA part. Allocated separately as DeepseekV4SWACache. return None # fp8_ds_mla is a UE8M0 block-scaled uint8 layout and needs 576B # alignment; plain bf16 / per-tensor fp8 rows use natural element-size # pages. uses_fp8_ds_mla_layout = self.kv_cache_dtype == "fp8_ds_mla" return MLAAttentionSpec( block_size=vllm_config.cache_config.block_size, num_kv_heads=1, head_size=self.head_dim, dtype=torch.uint8 if uses_fp8_ds_mla_layout else self.kv_cache_torch_dtype, compress_ratio=self.compress_ratio, cache_dtype_str=self.kv_cache_dtype, alignment=576 if uses_fp8_ds_mla_layout else None, model_version="deepseek_v4", kv_quant_mode=get_kv_quant_mode(self.kv_cache_dtype), ) ``` --- ## 4. 修改文件清单 | 文件 | 修改内容 | |---|---| | `vllm/models/deepseek_v4/nvidia/dspark.py` | `_remap_dspark_name`:draft block 权重从 `model.layers.{num_hidden_layers + stage}` 改为 `model.layers.{stage}` | | `vllm/models/deepseek_v4/attention.py` | `DeepseekV4Attention.get_kv_cache_spec`:在 `MLAAttentionSpec` 中补充 `kv_quant_mode=get_kv_quant_mode(self.kv_cache_dtype)` | > 如果上游 vllm-main 仓库与 vllm-dspark 安装包的文件结构一致,也需要同步修改 `vllm-main` 下对应路径的同名文件。 --- ## 5. 完整 diff(面向 PR) ```diff --- a/vllm/models/deepseek_v4/nvidia/dspark.py +++ b/vllm/models/deepseek_v4/nvidia/dspark.py @@ -486,8 +486,8 @@ class DSparkDeepseekV4ForCausalLM(nn.Module): if rest.startswith(("main_proj.", "main_norm.")) or rest.startswith( head_prefixes ): return f"model.{rest}" - # Draft layers live after the target layers in the decoder stack. - return f"model.layers.{self.num_hidden_layers + stage}.{rest}" + # Draft layers live in a ModuleList, so their actual parameter names are + # model.layers.{stage}.* even though the prefix passed to the decoder + # layer is layers.{num_hidden_layers + stage} (for compress_ratio). + return f"model.layers.{stage}.{rest}" ``` ```diff --- a/vllm/models/deepseek_v4/attention.py +++ b/vllm/models/deepseek_v4/attention.py @@ -613,6 +613,7 @@ class DeepseekV4Attention(nn.Module, AttentionLayerBase, ABC): cache_dtype_str=self.kv_cache_dtype, alignment=576 if uses_fp8_ds_mla_layout else None, model_version="deepseek_v4", + kv_quant_mode=get_kv_quant_mode(self.kv_cache_dtype), ) ``` --- ## 6. 验证结果 ### 6.1 Qwen3 DSpark 通路验证 - 目标模型:`/data/models/Qwen3-4B` - Draft 模型:`deepseek-ai/dspark_qwen3_4b_block7` - 结果:服务启动成功,`/v1/completions` 返回结果正常,speculative decoding 工作。 ### 6.2 DeepSeek-V4-Flash-DSpark 服务验证 修复后,使用第 2 节的命令可以成功启动服务,监听 `http://127.0.0.1:30004`。简单 completion 请求返回正常: ```bash curl -s -X POST http://127.0.0.1:30004/v1/completions \ -H "Content-Type: application/json" \ -d '{ "model": "/data/models/DeepSeek-V4-Flash-DSpark", "prompt": "The capital of France is", "max_tokens": 20, "temperature": 0.0 }' # 返回: " Paris." ``` ### 6.3 Serving benchmark 使用 vLLM 内置 `vllm bench serve`: ```bash vllm bench serve \ --host 127.0.0.1 --port 30004 \ --backend openai \ --dataset-name sharegpt \ --dataset-path /data/user1/yy/datasets/ShareGPT_filtered_chat.json \ --sharegpt-output-len 256 \ --num-prompts 50 \ --max-concurrency 16 \ --endpoint /v1/completions \ --model /data/models/DeepSeek-V4-Flash-DSpark \ --seed 42 ``` 结果: | 指标 | 数值 | |---|---| | Request throughput | **2.38 req/s** | | Output token throughput | **610.47 tok/s** | | Acceptance rate | **28.20%** | | Mean acceptance length | **2.41** | | Mean TTFT | 2162 ms | | Mean TPOT | 16.85 ms | | Total input tokens | 16381 | | Total generated tokens | 12800 | 服务器日志中的 SpecDecoding metrics 在稳态下 acceptance rate 落在 25%–36% 区间,与 DSpark 论文预期一致。 --- ## 7. 影响范围 - **受影响**:所有使用 NVIDIA 路径运行 DeepSeek-V4-Flash-DSpark 的 GPU(包括 H200/SM90 的 FlashMLA 路径,以及 B200/SM120 的 FlashInfer SM120 路径)。 - **不受影响**: - Qwen3 DSpark(使用独立的 `qwen3_dspark.py` 实现,不共享 `_remap_dspark_name`)。 - AMD/ROCm 和 XPU 平台(目前 DSpark 仅支持 NVIDIA)。 --- ## 8. 后续 PR 待办 - [ ] 在 vllm-main 上同步应用以上两处修改。 - [ ] 跑通 DeepSeek-V4-Flash-DSpark 的单元测试 / 冒烟测试。 - [ ] 检查 `DeepseekV4IndexerCache.get_kv_cache_spec` 是否也需要补充 `kv_quant_mode`(当前未改动,因为未触发错误)。 - [ ] 补充 DSpark draft 权重加载的回归测试(可选,但建议)。 - [ ] 向 vllm-project/vllm 提交 PR,并在描述中引用 issue #47648。 --- ## 9. 备注 - 本文件中的 token、路径、版本号均基于 2026-07-05 的实际运行环境。 - GitHub issue 评论已发布:https://github.com/vllm-project/vllm/issues/47648#issuecomment-4886275279