feat(pro6000/GLM-5.3): 方案 D/E/F 部署资产入库(TP2PP4 生产配方 / TP8+DFlash2 / PD 分离四角色链)

- scripts: 11 个服务器原样脚本入库(md5 对照表更新至 README);D=60.1 生产原样配方、
  E=v5 DFlash 底稿、F=PD 链四角色部署+launch+双场景压测驱动
- profiles: 新增 6 个 .env(D/E 单机 + F 四角色,均带镜像 digest
  sha256:28e0d260…,对齐 kimi3 PD 多角色先例)
- deploy/PD_CHAIN.md: 方案 F 编排手册(启动顺序 mc-master→prefill→decode→router、
  基础设施依赖表、质量门口径、拆链恢复、÷2 单机等效判决)
- platforms/patches/pro6000/glm53_pd_chain/: sglang 补丁树 vs 镜像原版 11 文件
  unified diff 快照——宿主树无 .git,此为唯一版本记录(DFlash+PP+PD 解锁全集)
- deploy/manifests/: GLM-5.3-NVFP4(47分片)/GLM-5.3-DFlash2(单分片) 权重 md5 清单
- deploy/CURRENT.md: 全集群现役状态页(2026-09-08 八机实测)
- deploy/verify_profile.sh: 防漂移核验工具(digest+参数 token 比对+端口/health,
  已在 60.1 生产容器实测 PASS)
This commit is contained in:
yy-fighting 2026-09-08 16:36:15 +08:00
parent b3165a1d3c
commit 5c749cda03
26 changed files with 1647 additions and 2 deletions

35
deploy/CURRENT.md Normal file
View File

@ -0,0 +1,35 @@
# 现役部署状态页live 核验于 2026-09-08
> 本页记录 174.1.60.x 集群"现在跑的是什么"。改动机器前后先读这页;
> **任何变动(停/起/改配置)完成后必须更新本页**。核验命令:`docker ps` + `nvidia-smi`
## 机器状态2026-09-08 实测)
| 机器 | 在役 | 口径 / 归属 | 对应 profile |
|---|---|---|---|
| 60.1 (6000D-1) | `glm53-pp4`Up8 卡满载,:30000 | **方案 D 生产**。09-08 PD 压测窗口停机 ~2.5h 后已恢复并核验 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp2pp4.env` |
| 60.2 (6000D-2) | 空无容器GPU 全 0 MiB | 方案 F PD 链已于 09-08 拆除。GPU7 曾有外部裸金属任务 main_v2.py现已结束动卡前仍先核实归属 | `profiles/pro6000/glm53_nvfp4_pro6000_pd_{prefill→decode 侧}.env` |
| 60.3 | 空(无容器) | — | — |
| 60.4 | `dsv4_scan` 容器Up 5min+ 裸金属 sglang `DeepSeek-V4-Flash-0731` DSPARK DP2/TP4 :30000 | **并行会话/外部在役,勿动** | — |
| 60.5 | `glm53-nvfp4`Up 29h | **NVFP4 团队生产**deploy_glm53_605.shmd5 fcd9109b。生产机铁律不实验、不重启、不覆盖脚本 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(方案 A 口径) |
| 60.6 | 空(无容器) | — | — |
| 60.7 | 空 | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — |
| 60.8 | 空 | 09-08 已停服清空、router 拆除,不再恢复 | — |
## 方案 A-F 一览GLM-5.3-NVFP4 @ pro60002026-09-08 双场景报告口径)
| 方案 | 一句话 | profile / 脚本 |
|---|---|---|
| A | TP8 + EAGLE基准60.5 生产口径 | `glm53_nvfp4_pro6000_sglang_tp8eagle.env` |
| B | TP4 PP2 | `glm53_nvfp4_pro6000_sglang_tp4pp2.env` |
| C | TP4 PP2 + IndexCache(freq=4) | `glm53_nvfp4_pro6000_sglang_tp4pp2_index.env` |
| D | TP2 PP460.1 生产在役 | `glm53_nvfp4_pro6000_sglang_tp2pp4.env` |
| E | TP8 + DFlash2 投机 | `glm53_nvfp4_pro6000_sglang_tp8dflash2.env` |
| F | PD 分离双机链mc-master→prefill→decode→router | `glm53_nvfp4_pro6000_pd_{master,prefill,decode,router}.env` + `deploy/PD_CHAIN.md` |
压测数据飞书《GLM-5.3-NVFP4 双场景压测报告》(`SZUSdEqY1oRVxGxgILBcHqPJnEc`)。
## 防漂移
每个在役容器用 `deploy/verify_profile.sh <profile.env>` 定期核验(镜像 digest + 启动参数 +
端口),发现不一致 = 容器被人手改过,先查清归属再处理。

96
deploy/PD_CHAIN.md Normal file
View File

@ -0,0 +1,96 @@
# 方案FGLM-5.3-NVFP4 PD 分离完整链(双机 6000D编排手册
> 2026-09-08 实测终态。四角色、两台机、启动顺序强制。吞吐换算口径:**链合计 ÷2 = 单机等效**(与单机方案 A-E 可比)。
> 完整压测数据见飞书《GLM-5.3-NVFP4 双场景压测报告》方案 F 行(文档 `SZUSdEqY1oRVxGxgILBcHqPJnEc` / wiki `NzbMwzmKviYidRkZYRrc8GYQnrf`)。
## 拓扑
| 角色 | 机器 | 容器 | 端口 | profile |
|---|---|---|---|---|
| 1. mc-masterMooncake 元数据) | 174.1.60.1 | `mc-master` | 50051 | `profiles/pro6000/glm53_nvfp4_pro6000_pd_master.env` |
| 2. prefillTP4 PP2 + DFLASH 草稿) | 174.1.60.1 | `glm53-pd-smoke-prefill` | 30000 | `profiles/pro6000/glm53_nvfp4_pro6000_pd_prefill.env` |
| 3. decodeTP8 + DFLASH v5 配方) | 174.1.60.2 | `glm53-s1-decode` | 30000 | `profiles/pro6000/glm53_nvfp4_pro6000_pd_decode.env` |
| 4. routerMiniLB | 174.1.60.2 | `pd-smoke-router` | 31000 | `profiles/pro6000/glm53_nvfp4_pro6000_pd_router.env` |
镜像统一:`lmsysorg/sglang:nightly-dev-20260828-daf63171`
digest `sha256:28e0d26073161e49ca56eba808d264a4223804a212020f1dfe1b2405b9f8a399`)。
## 启动顺序(强制)
```
mc-master (60.1) → prefill (60.1) → decode (60.2) → router (60.2)
```
对应脚本(`experiments/pro6000/glm53_nvfp4_pro6000d_sglang_dual_scenario_bench/scripts/`
```bash
# 60.1注意60.1 日常跑生产容器 glm53-pp4先停它见下节
bash /tmp/deploy_pd_smoke_master.sh # 判据 ss :50051
bash /tmp/deploy_pd_probe.sh p1b_prefill_tp4pp2_cps8k_radix_launch.sh # 判据 :30000/health 200~10-15min
# 60.2
bash /tmp/deploy_s1_decode.sh # 判据 :30000/health 200~10min
bash /tmp/deploy_pd_smoke_router.sh # 判据 ss :31000
```
场景一90% 命中)必须用 **radix 变体** launch`p1b_prefill_tp4pp2_cps8k_radix_launch.sh`
与基线唯一差异 = 无 `--disable-radix-cache`)。场景二 0 命中radix 开销可忽略,两场景共用同一部署。
## 基础设施依赖(缺一不可)
| 依赖 | 位置 | 说明 |
|---|---|---|
| sglang 补丁树 | `/data/sglang_patch_glm53` → 容器 `/sgl-workspace/sglang` | 与镜像原版差 **11 个文件**10 改 + 1 新增 `dflash_pp.py`),清单与 diff 见 `platforms/patches/pro6000/glm53_pd_chain/`。无补丁则 DFlash+PD 冷启动接线缺失decode 首请求 400 |
| mooncake wheel | `/data/flashkda_deploy/wheels/mooncake_transfer_engine_cuda13-0.3.12.post1-cp312-cp312-manylinux_2_28_x86_64.whl` | 容器内 launch 脚本 `pip install --no-deps` 自装 |
| DFlash2 草稿权重 | `/data/hf_models/GLM-5.3-DFlash2` | 两角色都要挂 |
| IB 设备 | mlx5_0mlx5_3 + `--device /dev/infiniband` + `--ulimit memlock=-1` | RDMA KV 传输通道 |
| 环境变量 | `MOONCAKE_MASTER=174.1.60.1:50051` `MOONCAKE_PROTOCOL=rdma` | prefill/decode 容器都要 |
## 质量门(经 router :31000全链路含 KV transfer + 草稿)
60.2 上生成 router 口径的 gate 副本:
```bash
sed 's/PORT=30000/PORT=31000/' /root/quality_gate_605.sh > /tmp/quality_gate_31000.sh
```
判据GSM8K×5 + 中文推理必须过。tool call 预期不过(链配方无 parser与方案 B/D 同口径的已知配置缺口,记录放行)。
2026-09-08 实测 **6/7**核心项全过DFlash 草稿无质量损失)。
## 压测
驱动脚本:`run_pd_s1.sh` / `run_pd_s2.sh`(同 scripts 目录),跑在 60.1(语料本地,
`--url http://174.1.60.2:31000/generate` 走 router`--container glm53-pd-smoke-prefill` 读 prefill 日志核命中率)。
- **场景一**128k/64k × cc1-4nreq8shared-frac 0.9canonical 窗口 9301-9308
- **场景二**cc8/16/32 canonical 9311-9313nreq 16/32/32cc40/64 全新窗口 pool-override
16,384,000 / 17,203,200nreq 40/64 单轮满波)
- 指标读取:命中率读 prefill 容器日志PP2 日志计数 ×2不影响比值accept 从 decode 日志读
router 转发后 usage 字段可能缺失)
- 每点核验ok/failed=0、retractions、s1 hit≈0.90、s2 hit=0
### 2026-09-08 实测判决13/13 干净点)
- **场景一**128k cc1 TTFT 3.91s = 六方案最低(跨机重叠把 prefill 与 RDMA 传输完全藏住),
但 cc≥2 时 decode 侧排队TTFT 堆到 31.0s÷2 输出吞吐 21.2-23.5 tok/s仅为 A 的 30-57%。
根因decode KV 池 214,336 token 按 16k 场景定容 → 128k 仅 1 驻留、64k 仅 3。
- **场景二**÷2 单机等效 74-85 tok/s = A 的 75-96%,但 TTFT 较 A 减半。
判决:**两机买 TTFT、不买吞吐**。decode MRR12 在 cc16+ 饱和DFlash accept 在 12-batch
verify 下掉到 1.8(单机 EAGLE 2.84)。
- **PD 机制验证**:跨机 prefill/decode 重叠成立、mooncake RDMA 传输被吸收——机制无罪,
容量配置decode 池按 16k 定容)是判决主因。
## 拆链与生产恢复(顺序固定)
```bash
# 1. 拆链
# 60.1: docker rm -f mc-master glm53-pd-smoke-prefill
# 60.2: docker rm -f glm53-s1-decode pd-smoke-router
# 2. 等双机显存排空60.1 <2000MiB60.2 GPU7 常有外部裸金属任务 main_v2.py核实归属勿清
# 3. 恢复 60.1 生产bash /tmp/deploy_glm53_pp4.sh → health 200 + docker inspect Args 核对
```
## 60.1 生产停机窗口提示
60.1 日常跑生产 `glm53-pp4`(方案 DTP2PP48 卡满载)。拉 prefill 角色前必须停生产;
`deploy_glm53_pp4.sh` 生成的参数与生产容器 `docker inspect` Args 已核对逐字一致,
恢复即完全复现。全程停机 ~2-2.5h(部署+13 点压测+质量门+恢复)。

View File

@ -0,0 +1,4 @@
302f15789d891aa39620209286252fed ./README.md
ce1962554abf73c6ad953aa0ac3f20d7 ./config.json
5e170425c97cda8f798c74041979a569 ./configuration.json
32c81842f12e56e6ac2a1feaafd5bfa7 ./model.safetensors

View File

@ -0,0 +1,58 @@
434f8692a4914f39714df2c4493c92d7 ./.gitattributes
245254422e0ca504a5f09e37719ca718 ./LICENSE
02e0c207304666abdd93de6a3780c52a ./README.md
e4c9e1000a513f680dd796f76d1097e0 ./chat_template.jinja
d42bba3f99ce9b2705bb19c41930ea0c ./config.json
5e170425c97cda8f798c74041979a569 ./configuration.json
ee2560496cd373851033b3c05ee461fb ./generation_config.json
027f9b7a7bb96f30962b5be7e4e4a5f1 ./hf_quant_config.json
57cdec3085e0805c4ad5875f61c63b68 ./model-00030-of-00047.safetensors
be93abf8f84e890b8dd1d4187ef540c2 ./model-00001-of-00047.safetensors
a4a627a33e69e91491c9a3d3a7c9cba9 ./model-00031-of-00047.safetensors
500a87d1fd1dd6a93400539227361f5d ./model-00002-of-00047.safetensors
c732e4ee2f910af666418abd547d34cb ./model-00032-of-00047.safetensors
68024a4367c3341402766fcc40d383f6 ./model-00003-of-00047.safetensors
94188d8acbc7fd47ee99400507793d6a ./model-00033-of-00047.safetensors
b5870fe5e8704e879895a7b5a3be5136 ./model-00004-of-00047.safetensors
87b17b02643b1c14ad2e90900b3dfa89 ./model-00034-of-00047.safetensors
f2f51d0be2c49ba6569d59d8d1db5236 ./model-00005-of-00047.safetensors
915184c001e70fea709ffbb5242e4a8c ./model-00035-of-00047.safetensors
fe2146b7c71775e51bad2322c52c34d1 ./model-00006-of-00047.safetensors
56d618d4ba615cc19a0898d02d82bb52 ./model-00036-of-00047.safetensors
03c0d5ea0471553cce658b70daed7b47 ./model-00007-of-00047.safetensors
c76ac8dd7dbddbc4a636118ad8a5f713 ./model-00038-of-00047.safetensors
49e246d6faa068641d2858e36b6bb5cd ./model-00008-of-00047.safetensors
8f31262bc0a079c946a9ec1cdeb4051e ./model-00037-of-00047.safetensors
b956d449de03a2db14fee14df7aaddac ./model-00009-of-00047.safetensors
087456e9659f4d7c8c9a4af093b7c828 ./model-00039-of-00047.safetensors
f94083fc155df999ab9ff27f6f67b978 ./model-00010-of-00047.safetensors
bfd9606be27f5efbb34d85dc8fad3cea ./model-00040-of-00047.safetensors
775a242af0663e5c65bc1b41a8930314 ./model-00011-of-00047.safetensors
a4c70502c1470d1cd520f4f2b0f968d8 ./model-00041-of-00047.safetensors
1d9d4ca43e411b977cded63b3447b8b8 ./model-00012-of-00047.safetensors
f00a801a8de9a31b2ff686f6f8924d41 ./model-00042-of-00047.safetensors
5feebc534bb09cee7b7e908aab8e2108 ./model-00013-of-00047.safetensors
b58238047ce49362d39c36b54761c4cf ./model-00043-of-00047.safetensors
c15c8fa6325af6ca302e10425b357199 ./model-00014-of-00047.safetensors
fb652edaa0a450b7ce333829e363ee7d ./model.safetensors.index.json
ce7e000a8d1c076b13346374f378e1f3 ./model-00015-of-00047.safetensors
752f6cd2e6a4a2ea824d1b513530e0b0 ./tokenizer.json
96ab98059044ac21eff43da3d9882689 ./model-00016-of-00047.safetensors
3312b710133454ca7150bccad6381bd8 ./tokenizer_config.json
82429b288fe2a97cb771e78bea60a0bd ./model-00017-of-00047.safetensors
7605255172c381931a496b316faecdb7 ./model-00018-of-00047.safetensors
e7fdbb1bd3e82f97831f4c1c69e0403a ./model-00019-of-00047.safetensors
f5d355f9a737d7aa270d178a1bceeeec ./model-00020-of-00047.safetensors
50b23a8dc4932840e5ae2062ca04b60a ./model-00021-of-00047.safetensors
285172cbc5e900681efc5783f04a346a ./model-00022-of-00047.safetensors
accdb4eb9e0831705bcd2d39e4f3bdcb ./model-00023-of-00047.safetensors
9182f13eb2670db2a5312aaca2f9b1da ./model-00024-of-00047.safetensors
198809340ec4795bdcfafdf7cd3074c7 ./model-00025-of-00047.safetensors
da5e3a1afe95fe14c5d2d1f432f4a933 ./model-00026-of-00047.safetensors
1d15d1696fd6d2cab3ed8726c3b25112 ./model-00027-of-00047.safetensors
6147f7b5acb9a297eab276b16a58efe7 ./model-00028-of-00047.safetensors
529b4048cf2f18d9eeaf8d1ce37b60cd ./model-00029-of-00047.safetensors
ffe813a8385daf74f69f632d9437085c ./model-00044-of-00047.safetensors
f67ddcaf1b68509fd6ab8d0399632a75 ./model-00045-of-00047.safetensors
489dc1a568e671176a98e7cc23519c17 ./model-00046-of-00047.safetensors
3a7abe2df1bd98e749ee8733370ea78b ./model-00047-of-00047.safetensors

View File

@ -0,0 +1,15 @@
# 权重完整性清单md5
2026-09-08 在 174.1.60.1 `/data/hf_models/` 下对两个模型目录逐文件 `md5sum` 的原样输出。
用于:新机器部署前核对权重传输完整性、或怀疑权重被改动时做漂移检测。
校验方法(在权重目录下):
```bash
md5sum -c /path/to/GLM-5.3-NVFP4.md5 # 清单内路径为 ./ 相对路径
```
| 清单 | 模型 | 规模 | 用途 |
|---|---|---|---|
| `GLM-5.3-NVFP4.md5` | GLM-5.3-NVFP4 主模型 | 47 分片 + 配置共 55 文件 | 方案 A-F 全部部署的目标模型 |
| `GLM-5.3-DFlash2.md5` | GLM-5.3-DFlash2 草稿模型 | 单分片 model.safetensors + 配置共 4 文件 | 方案 E 与方案 F 的 DFLASH 投机草稿 |

View File

@ -0,0 +1,52 @@
# GLM-5.3-NVFP4 PD 分离链 - 角色3: decode 节点方案F跑在 6000D-2 = 174.1.60.28 卡)。
# 四角色链之一,启动顺序强制: mc-master -> prefill -> [本角色] -> router。
# 完整链编排见 deploy/PD_CHAIN.md。
# 可执行部署脚本: experiments/pro6000/glm53_nvfp4_pro6000d_sglang_dual_scenario_bench/scripts/deploy_s1_decode.sh
# launch 脚本同目录 s1_decode_launch.sh
#
# 关键点(实测踩坑,勿随意改):
# - 配方 = 方案 E v5TP8 + DFLASH + MRR12 + fa4 + window2048+ PD decode flags
# KV 池 214,336 token比单机 E 的 243,584 少PD 传输缓冲占显存)
# - 容量属性(实测判决的核心): 池按 16k 场景定容 → 128k 仅容 1 个驻留cc4 时 TTFT
# 堆到 31.0s、64k 容 3、16k 容 12MRR12 上限)。长上下文负载该池就是瓶颈
# - 必须挂补丁树 /data/sglang_patch_glm53:/sgl-workspace/sglang唯一改动
# python/sglang/srt/speculative/spec_info.py 接 build_dflash_family_disagg_draft_input
# DFlash PD 冷启动接线)+ 装 mooncake wheel--no-depslaunch 脚本自装)
# - --network host + --device /dev/infiniband + --ulimit memlock=-1IB 设备 mlx5_0-3
# - DFlash accept 在 12-batch verify 下掉到 1.8(单机 EAGLE 2.84——decode 侧并发
# verify 是 DFlash 的弱势区,场景二吞吐上限由此而来
# - 60.2 平时空闲但 GPU7 常有外部裸金属任务main_v2.py动卡前核实归属勿清
#
# 实测成绩: 飞书《GLM-5.3-NVFP4 双场景压测报告》方案 F 行(文档 SZUSdEqY1oRVxGxgILBcHqPJnEc
# 场景二判决: "两机买 TTFT、不买吞吐"÷2 单机等效 74-85 tok/s = A 的 75-96%TTFT 减半)。
PLATFORM=pro6000
EXPERIMENT=glm53_nvfp4_pro6000_pd_decode
MODEL_NAME=GLM-5.3-NVFP4
ENGINE=sglang
RUNTIME=docker
ROLE=pd-decode
NODE=174.1.60.2
DOCKER_IMAGE=lmsysorg/sglang:nightly-dev-20260828-daf63171
DOCKER_IMAGE_DIGEST=sha256:28e0d26073161e49ca56eba808d264a4223804a212020f1dfe1b2405b9f8a399
CONTAINER_NAME=glm53-s1-decode
MODEL_PATH=/data/hf_models/GLM-5.3-NVFP4
DRAFT_MODEL_PATH=/data/hf_models/GLM-5.3-DFlash2
PORT=30000
HEALTH_PATH=/health
HEALTH_WAIT_S=1200
TP=8
MEM_FRACTION_STATIC=0.85
MAX_RUNNING_REQUESTS=12
CHUNKED_PREFILL_SIZE=8192
DEVICE_VARS="CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7"
ENGINE_ENV="MOONCAKE_MASTER=174.1.60.1:50051 MOONCAKE_PROTOCOL=rdma PYTHONUNBUFFERED=1"
DOCKER_FLAGS="--gpus all --network host --ipc=host --shm-size 64g --ulimit memlock=-1 --device /dev/infiniband --restart unless-stopped"
VOLUMES="/data/hf_models:/data/hf_models /data/flashkda_deploy/wheels:/mc_wheels:ro /data/sglang_patch_glm53:/sgl-workspace/sglang"
BOOTSTRAP="pip install /mc_wheels/mooncake_transfer_engine_cuda13-0.3.12.post1-cp312-cp312-manylinux_2_28_x86_64.whl --no-deps -q && python3 -m sglang.launch_server ${LAUNCH_ARGS}"
LAUNCH_ARGS="--model-path ${MODEL_PATH} --tp-size ${TP} --mem-fraction-static ${MEM_FRACTION_STATIC} --max-running-requests ${MAX_RUNNING_REQUESTS} --disable-radix-cache --disable-shared-experts-fusion --moe-runner-backend flashinfer_cutlass --disable-flashinfer-autotune --disable-custom-all-reduce --chunked-prefill-size ${CHUNKED_PREFILL_SIZE} --speculative-algorithm DFLASH --speculative-draft-model-path ${DRAFT_MODEL_PATH} --speculative-draft-attention-backend fa4 --speculative-draft-window-size 2048 --disaggregation-mode decode --disaggregation-transfer-backend mooncake --disaggregation-bootstrap-port 28800 --disaggregation-ib-device mlx5_0,mlx5_1,mlx5_2,mlx5_3 --host 0.0.0.0 --port ${PORT} --json-model-override-args {\"index_topk_freq\": 4}"

View File

@ -0,0 +1,29 @@
# GLM-5.3-NVFP4 PD 分离链 - 角色1: Mooncake 元数据服务 mc-master方案F双机 6000D-1 + 6000D-2
# 四角色链之一,启动顺序强制: mc-master(60.1:50051) -> prefill(60.1:30000) -> decode(60.2:30000) -> router(60.2:31000)。
# 完整链编排见 deploy/PD_CHAIN.md可执行脚本: experiments/pro6000/glm53_nvfp4_pro6000d_sglang_dual_scenario_bench/scripts/deploy_pd_smoke_master.sh
#
# 关键点:
# - 镜像自带的 mooncake_master 二进制(/opt/sglang/bin/mooncake_master无需 GPU
# --network host监听 50051prefill/decode 容器通过 MOONCAKE_MASTER=174.1.60.1:50051 注册
# - 就绪判据: ss -tln | grep :50051脚本 sleep 3 后检查)
# - 拆链时必须先删本容器之外的角色再删它? 否——顺序无依赖,但生产恢复 60.1 时
# 本容器与 glm53-pd-smoke-prefill 都要删干净、等显存排空再拉生产容器
#
# 实测背景2026-09-08 方案F 双场景压测): 13/13 干净点,质量门 6/7。
PLATFORM=pro6000
EXPERIMENT=glm53_nvfp4_pro6000_pd_master
MODEL_NAME=GLM-5.3-NVFP4
ENGINE=sglang-mooncake-master
RUNTIME=docker
ROLE=pd-master
NODE=174.1.60.1
DOCKER_IMAGE=lmsysorg/sglang:nightly-dev-20260828-daf63171
DOCKER_IMAGE_DIGEST=sha256:28e0d26073161e49ca56eba808d264a4223804a212020f1dfe1b2405b9f8a399
CONTAINER_NAME=mc-master
PORT=50051
DOCKER_FLAGS="--network host --restart unless-stopped"
ENTRYPOINT="/opt/sglang/bin/mooncake_master"
READY_CHECK="ss -tln | grep :50051"

View File

@ -0,0 +1,59 @@
# GLM-5.3-NVFP4 PD 分离链 - 角色2: prefill 节点方案F跑在 6000D-1 = 174.1.60.18 卡)。
# 四角色链之一,启动顺序强制: mc-master -> [本角色] -> decode -> router。
# 完整链编排见 deploy/PD_CHAIN.md。
# 可执行部署脚本: experiments/pro6000/glm53_nvfp4_pro6000d_sglang_dual_scenario_bench/scripts/deploy_pd_probe.sh
# (参数 = launch 脚本名launch 脚本同目录 p1b_prefill_tp4pp2_cps8k_launch.sh 及其 radix 变体)
#
# 关键点(实测踩坑,勿随意改):
# - 拓扑 TP4 PP24卡/stage × 2 stage+ DFLASH 草稿(草稿只跑 prefill 侧草稿 KV
# export SGLANG_DFLASH_PD_DRAFT_KV_TRANSFER=0 = 草稿 KV 不跨机传输)
# - mem 0.78 低于单机方案PD 模式下 mooncake 传输缓冲占显存cps 8192非 D 方案的 16384
# - 基线 launch 脚本带 --disable-radix-cache场景一90% 命中)实测用 radix 变体
# p1b_prefill_tp4pp2_cps8k_radix_launch.sh唯一差异 = 删掉该 flag
# - 必须挂补丁树 /data/sglang_patch_glm53:/sgl-workspace/sglang含 1 个文件改动:
# python/sglang/srt/speculative/spec_info.py 接 build_dflash_family_disagg_draft_input
# DFlash PD 冷启动接线;无此补丁 decode 首请求 400
# - 必须装 mooncake wheel /data/flashkda_deploy/wheels/mooncake_transfer_engine_cuda13-
# 0.3.12.post1-cp312-cp312-manylinux_2_28_x86_64.whl--no-deps容器内 launch 脚本自装)
# - --network host + --device /dev/infiniband + --ulimit memlock=-1IB 设备 mlx5_0-3
# - 前置硬检查: mc-master 50051 必须已监听deploy_pd_probe.sh 自带)
# - 60.1 是生产机(日常跑 glm53-pp4拉本角色前须停生产容器测完等显存 <2000MiB
# 再跑 deploy_glm53_pp4.sh 恢复
#
# 实测成绩: 飞书《GLM-5.3-NVFP4 双场景压测报告》方案 F 行(文档 SZUSdEqY1oRVxGxgILBcHqPJnEc
# 128k cc1 TTFT 3.91s = 六方案最低;跨机 prefill/decode 重叠使 RDMA 传输被完全吸收。
PLATFORM=pro6000
EXPERIMENT=glm53_nvfp4_pro6000_pd_prefill
MODEL_NAME=GLM-5.3-NVFP4
ENGINE=sglang
RUNTIME=docker
ROLE=pd-prefill
NODE=174.1.60.1
DOCKER_IMAGE=lmsysorg/sglang:nightly-dev-20260828-daf63171
DOCKER_IMAGE_DIGEST=sha256:28e0d26073161e49ca56eba808d264a4223804a212020f1dfe1b2405b9f8a399
CONTAINER_NAME=glm53-pd-smoke-prefill
MODEL_PATH=/data/hf_models/GLM-5.3-NVFP4
DRAFT_MODEL_PATH=/data/hf_models/GLM-5.3-DFlash2
PORT=30000
HEALTH_PATH=/health
HEALTH_WAIT_S=1500
TP=4
PP=2
MEM_FRACTION_STATIC=0.78
MAX_RUNNING_REQUESTS=48
CHUNKED_PREFILL_SIZE=8192
DEVICE_VARS="CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7"
ENGINE_ENV="MOONCAKE_MASTER=174.1.60.1:50051 MOONCAKE_PROTOCOL=rdma SGLANG_DFLASH_PD_DRAFT_KV_TRANSFER=0 PYTHONUNBUFFERED=1"
DOCKER_FLAGS="--gpus all --network host --ipc=host --shm-size 64g --ulimit memlock=-1 --device /dev/infiniband --restart unless-stopped"
VOLUMES="/data/hf_models:/data/hf_models /data/flashkda_deploy/wheels:/mc_wheels:ro /data/sglang_patch_glm53:/sgl-workspace/sglang"
BOOTSTRAP="pip install /mc_wheels/mooncake_transfer_engine_cuda13-0.3.12.post1-cp312-cp312-manylinux_2_28_x86_64.whl --no-deps -q && python3 -m sglang.launch_server ${LAUNCH_ARGS}"
LAUNCH_ARGS="--model-path ${MODEL_PATH} --tp-size ${TP} --pp-size ${PP} --mem-fraction-static ${MEM_FRACTION_STATIC} --max-running-requests ${MAX_RUNNING_REQUESTS} --disable-shared-experts-fusion --moe-runner-backend flashinfer_cutlass --disable-flashinfer-autotune --disable-custom-all-reduce --chunked-prefill-size ${CHUNKED_PREFILL_SIZE} --speculative-algorithm DFLASH --speculative-draft-model-path ${DRAFT_MODEL_PATH} --speculative-draft-attention-backend fa4 --disaggregation-mode prefill --disaggregation-transfer-backend mooncake --disaggregation-bootstrap-port 28800 --disaggregation-ib-device mlx5_0,mlx5_1,mlx5_2,mlx5_3 --host 0.0.0.0 --port ${PORT} --json-model-override-args {\"index_topk_freq\": 4}"
# 场景一90% 命中)变体: LAUNCH_ARGS 追加 --disable-radix-cache 删除radix 开)。
# 基线(本 profile 口径)= radix off。

View File

@ -0,0 +1,35 @@
# GLM-5.3-NVFP4 PD 分离链 - 角色4: MiniLB 路由方案F跑在 6000D-2 = 174.1.60.2:31000
# 四角色链之一,启动顺序强制: mc-master -> prefill -> decode -> [本角色]。
# 完整链编排见 deploy/PD_CHAIN.md。
# 可执行脚本: experiments/pro6000/glm53_nvfp4_pro6000d_sglang_dual_scenario_bench/scripts/deploy_pd_smoke_router.sh
#
# 关键点:
# - 镜像内自带 sglang_routerpython3 -m sglang_router.launch_router无需 GPU
# - --prefill 参数形态: URL 后跟 prefill 侧 disaggregation-bootstrap-port28800
# 这是 KV 传输握手口,漏掉则路由建立后首请求挂起;--decode 只有 URL
# - bench 与质量门全部经 :31000/generate 打(--url http://174.1.60.2:31000/generate
# 即全链路含 KV transfer + DFlash 草稿
# - usage/accept 指标 router 转发后可能缺失: 命中率读 prefill 容器日志
# docker logs glm53-pd-smoke-prefillPP2 下日志计数 ×2 不影响比值),
# accept 读 decode 容器日志
# - 质量门脚本: 用 quality_gate_605.sh 副本 sed 's/PORT=30000/PORT=31000/' 生成
# - 就绪判据: ss -tln | grep :31000
#
# 实测背景: 全链 13/13 干净点、质量门 6/7tool call 为 parser 配置缺口,与 B/D 同口径)。
PLATFORM=pro6000
EXPERIMENT=glm53_nvfp4_pro6000_pd_router
MODEL_NAME=GLM-5.3-NVFP4
ENGINE=sglang-router
RUNTIME=docker
ROLE=pd-router
NODE=174.1.60.2
DOCKER_IMAGE=lmsysorg/sglang:nightly-dev-20260828-daf63171
DOCKER_IMAGE_DIGEST=sha256:28e0d26073161e49ca56eba808d264a4223804a212020f1dfe1b2405b9f8a399
CONTAINER_NAME=pd-smoke-router
PORT=31000
DOCKER_FLAGS="--network host --restart unless-stopped --entrypoint python3"
LAUNCH_ARGS="-m sglang_router.launch_router --pd-disaggregation --mini-lb --prefill http://174.1.60.1:30000 28800 --decode http://174.1.60.2:30000 --host 0.0.0.0 --port 31000"
READY_CHECK="ss -tln | grep :31000"

View File

@ -0,0 +1,52 @@
# GLM-5.3-NVFP4 SGLang TP=2 PP=4 deployment profile (single RTX 6000D node, 8 GPUs).
# 方案 D2026-09-08 双场景补测6000D-1 现役生产容器 glm53-pp4 的原样配方)。
# 可执行部署脚本experiments/pro6000/glm53_nvfp4_pro6000d_sglang_dual_scenario_bench/scripts/deploy_glm53_pp4.sh
#
# 关键点(实测踩坑,勿随意改):
# - 脚本头注释是早期 TP4PP2 版残留(写"TP4 PP2 + IndexCache"),实际配置 TP=2 PP=4
# 以脚本 docker run 段为准md5 def3c64c5dc19e1d507080e3366c3762
# - 串行 prefill 有效速率 6.5-6.7k tok/s 为各方案最高KV 池 1,040,384 token
# A 的 3.76 倍≈61.6 请求驻留,超过 MRR 48——池在高并发下不构成约束
# - 场景二16k 独立输入 cc8-64成立cc32 起反超 TP4PP2方案Bcc40/64 输出
# 203.8/208.2 tok/sTTFT p50 五档全档低于 B
# - 场景一90% 命中长上下文8/8 全败TP2 长上下文每卡 KV 读量翻倍 + PP4 低并发
# 流水空泡,单请求 decode 仅 16-19 tok/s。生产态 radix off 前缀命中恒 0实际表现
# 比报告 D 行radix-on 最好情况)更差——长上下文/共享前缀负载勿用
# - DSA 实测:上下文长度不影响 TPOT52.8ms 恒定),并发才是驱动
# - 上生产须补 --tool-call-parser glm47 与 --reasoning-parser glm45
# - bench 口径:场景二 cc40/64 用 nreq 80/128与其他方案 40/64 单轮满波不同,
# 已在报告表注声明)
#
# 实测成绩飞书《GLM-5.3-NVFP4 双场景压测报告》方案 D 行(文档 SZUSdEqY1oRVxGxgILBcHqPJnEc
# 质量GSM8K×5 + 中文推理通过6/7tool call 为 parser 配置缺口)。
PLATFORM=pro6000
EXPERIMENT=glm53_nvfp4_pro6000_sglang_tp2pp4
MODEL_NAME=GLM-5.3-NVFP4
ENGINE=sglang
RUNTIME=docker
DOCKER_IMAGE=lmsysorg/sglang:nightly-dev-20260828-daf63171
DOCKER_IMAGE_DIGEST=sha256:28e0d26073161e49ca56eba808d264a4223804a212020f1dfe1b2405b9f8a399
CONTAINER_NAME=glm53-pp4
MODEL_PATH=/data/hf_models/GLM-5.3-NVFP4
SERVED_MODEL_NAME=/data/hf_models/GLM-5.3-NVFP4
PORT=30000
HEALTH_PATH=/health
HEALTH_WAIT_S=600
CONTAINER_PYTHON=python3
TP=2
PP=4
MEM_FRACTION_STATIC=0.85
MAX_RUNNING_REQUESTS=48
CHUNKED_PREFILL_SIZE=16384
DEVICE_VARS="CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7"
ENGINE_ENV="PYTHONUNBUFFERED=1 HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1"
DOCKER_FLAGS="--gpus all --shm-size 64g --ipc=host -p ${PORT}:${PORT}"
VOLUMES="/data/hf_models:/data/hf_models"
BOOTSTRAP="python3 -m sglang.launch_server ${LAUNCH_ARGS}"
LAUNCH_ARGS="--model-path ${MODEL_PATH} --tp-size ${TP} --pp-size ${PP} --mem-fraction-static ${MEM_FRACTION_STATIC} --max-running-requests ${MAX_RUNNING_REQUESTS} --disable-radix-cache --disable-shared-experts-fusion --moe-runner-backend flashinfer_cutlass --disable-flashinfer-autotune --disable-custom-all-reduce --chunked-prefill-size ${CHUNKED_PREFILL_SIZE} --host 0.0.0.0 --port ${PORT} --json-model-override-args {\"index_topk_freq\": 4}"

View File

@ -0,0 +1,52 @@
# GLM-5.3-NVFP4 SGLang TP=8 + DFlash2 speculative decoding profile (single RTX 6000D node, 8 GPUs).
# 方案 E2026-09-08 场景一补测6000D-2
# 可执行部署脚本experiments/pro6000/glm53_nvfp4_pro6000d_sglang_dual_scenario_bench/scripts/deploy_glm53_tp8_dflash2.sh
# v5 终版底稿,含四轮 OOM 战役完整教训注释md5 5bf2b47c9349e5855b963e571e35c096
#
# 关键点(实测踩坑,勿随意改):
# - v5 配方核心MRR 48→12verify CUDA graph 4.04→0.83GB,真正起作用的杠杆)+
# --speculative-draft-window-size 2048 + mem0.85/cps8192。四轮 OOM 根因与推导见脚本头注释
# - DFLASH block-diffusion 草稿 7 tokens/步draft 权重 GLM-5.3-DFlash2fa4 draft
# attentionfa4 会把 draft KV 强制 bf16fp8 需换 flashinfer/triton 后端,仅省 0.35GB 未用)
# - KV 池 243,584 tokenkv fp8_e4m3 由模型配置自动带出(无需显式 flag
# - 底稿(本 profile LAUNCH_ARGSradix/AR 均为禁用;场景一实测变体共四处 delta
# ① 去 --disable-radix-cache90% 命中前提)② 去 --disable-custom-all-reduce
# v1 CAR 与方案 A 一致开启)③ 加 --context-length 270336 ④ 加 --reasoning-parser
# glm45 --tool-call-parser glm47质量门 7/7 的前提)
# - 判决:场景一 8 点全部低于方案 C、7 点低于 A——DFlash accept 低于 EAGLE同语料
# 2.53 vs 2.84)而每步墙钟相当,劣势全在接受率。投机栈选型维持 EAGLE3勿用
# DFlash2 替换(性能问题非质量问题)
#
# 实测成绩飞书《GLM-5.3-NVFP4 双场景压测报告》方案 E 行(文档 SZUSdEqY1oRVxGxgILBcHqPJnEc
# 质量:变体配置下质量门 7/7GSM8K×5、中文推理、tool call 全过DFlash 草稿无质量损失。
PLATFORM=pro6000
EXPERIMENT=glm53_nvfp4_pro6000_sglang_tp8dflash2
MODEL_NAME=GLM-5.3-NVFP4
ENGINE=sglang
RUNTIME=docker
DOCKER_IMAGE=lmsysorg/sglang:nightly-dev-20260828-daf63171
DOCKER_IMAGE_DIGEST=sha256:28e0d26073161e49ca56eba808d264a4223804a212020f1dfe1b2405b9f8a399
CONTAINER_NAME=glm53-tp8-dflash2
MODEL_PATH=/data/hf_models/GLM-5.3-NVFP4
DRAFT_MODEL_PATH=/data/hf_models/GLM-5.3-DFlash2
SERVED_MODEL_NAME=/data/hf_models/GLM-5.3-NVFP4
PORT=30000
HEALTH_PATH=/health
HEALTH_WAIT_S=900
CONTAINER_PYTHON=python3
TP=8
MEM_FRACTION_STATIC=0.85
MAX_RUNNING_REQUESTS=12
CHUNKED_PREFILL_SIZE=8192
DEVICE_VARS="CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7"
ENGINE_ENV="PYTHONUNBUFFERED=1 HF_HUB_OFFLINE=1 TRANSFORMERS_OFFLINE=1"
DOCKER_FLAGS="--gpus all --shm-size 64g --ipc=host -p ${PORT}:${PORT}"
VOLUMES="/data/hf_models:/data/hf_models"
BOOTSTRAP="python3 -m sglang.launch_server ${LAUNCH_ARGS}"
LAUNCH_ARGS="--model-path ${MODEL_PATH} --tp-size ${TP} --mem-fraction-static ${MEM_FRACTION_STATIC} --max-running-requests ${MAX_RUNNING_REQUESTS} --disable-radix-cache --disable-shared-experts-fusion --moe-runner-backend flashinfer_cutlass --disable-flashinfer-autotune --disable-custom-all-reduce --chunked-prefill-size ${CHUNKED_PREFILL_SIZE} --speculative-algorithm DFLASH --speculative-draft-model-path ${DRAFT_MODEL_PATH} --speculative-draft-attention-backend fa4 --speculative-draft-window-size 2048 --host 0.0.0.0 --port ${PORT} --json-model-override-args {\"index_topk_freq\": 4}"

91
deploy/verify_profile.sh Normal file
View File

@ -0,0 +1,91 @@
#!/bin/bash
# verify_profile.sh —— 防漂移核验:运行中容器 vs 仓库 profile 声明
# 在目标服务器上运行(需 docker 读权限,无需 GPU。用法
# bash verify_profile.sh deploy/profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp2pp4.env
# bash verify_profile.sh <profile.env> [container_name] # container_name 缺省取 profile 的 CONTAINER_NAME
# 核验项:①容器在跑 ②镜像 registry digest ③launch 参数token 级比对)④端口监听。
# PD 链角色的参数在挂载的 launch 脚本内docker Args 只有 bash /smoke_launch.sh
# 第③项自动跳过并提示改用库内脚本 md5 对比scripts/ 目录各脚本头注释有 md5
set -uo pipefail
PROFILE="$1"
[ -r "$PROFILE" ] || { echo "FATAL: 无法读取 profile: $PROFILE"; exit 2; }
# profile 是声明式清单BOOTSTRAP 行引用 ${LAUNCH_ARGS}定义在后set -u 下直接
# source 会炸;临时关 -u。source 完成后 LAUNCH_ARGS 已是全部变量展开后的实参串。
set +u
# shellcheck disable=SC1090
source "$PROFILE"
set -u
CONT="${2:-${CONTAINER_NAME:-}}"
fail=0
section() { printf '\n== %s ==\n' "$1"; }
section "容器状态"
if [ -z "$CONT" ]; then echo "FATAL: profile 未定义 CONTAINER_NAME 且未显式传入"; exit 2; fi
if docker ps --format '{{.Names}}' | grep -qx "$CONT"; then
echo "OK $CONT 在跑"
else
echo "FAIL $CONT 未在运行"; docker ps -a --filter "name=$CONT" --format '{{.Names}} {{.Status}}' | head -3
exit 1
fi
section "镜像 digest"
if [ -n "${DOCKER_IMAGE_DIGEST:-}" ]; then
# RepoDigests 在 image 对象上container 先取 image ID再 image inspect
img_id=$(docker inspect --format '{{.Image}}' "$CONT" 2>/dev/null)
repo_digests=$(docker image inspect --format '{{join .RepoDigests "\n"}}' "$img_id" 2>/dev/null)
digest_hex="${DOCKER_IMAGE_DIGEST#sha256:}"
if echo "$repo_digests" | grep -q "$digest_hex"; then
echo "OK registry digest 一致: ${DOCKER_IMAGE_DIGEST}"
elif [ "$img_id" = "sha256:${digest_hex}" ]; then
echo "OK image ID 一致: ${DOCKER_IMAGE_DIGEST}"
else
echo "FAIL digest 不一致"
echo " profile: ${DOCKER_IMAGE_DIGEST}"
echo " 实际 image ID: ${img_id}"
echo " 实际 RepoDigests: $(echo "$repo_digests" | head -2 | tr '\n' ' ')"
fail=1
fi
else
echo "SKIP profile 未定义 DOCKER_IMAGE_DIGEST"
fi
section "启动参数"
args=$(docker inspect --format '{{join .Args " "}}' "$CONT" 2>/dev/null)
if echo "$args" | grep -q launch_server; then
actual=$(echo "$args" | sed 's/.*launch_server //')
expected="${LAUNCH_ARGS:-}"
tr ' ' '\n' <<<"$expected" | sed '/^$/d' | sort > /tmp/vp_expected.$$
tr ' ' '\n' <<<"$actual" | sed '/^$/d' | sort > /tmp/vp_actual.$$
if diff -q /tmp/vp_expected.$$ /tmp/vp_actual.$$ >/dev/null; then
echo "OK 参数一致token 比对,共 $(wc -l < /tmp/vp_expected.$$) 项)"
else
echo "FAIL 参数有漂移(< profile 声明 / > 容器实际):"
diff /tmp/vp_expected.$$ /tmp/vp_actual.$$ | sed 's/^/ /'
fail=1
fi
rm -f /tmp/vp_expected.$$ /tmp/vp_actual.$$
else
echo "SKIP docker Args 为 '$args' —— 参数在挂载的 launch 脚本内"
echo " 改用库内脚本 md5 对比experiments/.../scripts/ 各脚本头注释"
fi
section "端口 ${PORT:-?}"
if [ -n "${PORT:-}" ]; then
if ss -tln | grep -q ":${PORT} "; then
echo "OK :${PORT} 在监听"
else
echo "FAIL :${PORT} 未监听"; fail=1
fi
if [ -n "${HEALTH_PATH:-}" ]; then
code=$(curl -s -o /dev/null -w '%{http_code}' "http://localhost:${PORT}${HEALTH_PATH}" 2>/dev/null)
[ "$code" = "200" ] && echo "OK health ${HEALTH_PATH} -> 200" || { echo "FAIL health ${HEALTH_PATH} -> ${code}"; fail=1; }
fi
else
echo "SKIP profile 未定义 PORT"
fi
echo
[ $fail -eq 0 ] && echo "VERDICT: PASS" || echo "VERDICT: DRIFT DETECTED"
exit $fail

View File

@ -110,6 +110,31 @@ deploy 脚本内置集群踩坑防护:`docker rm -f` 异步滞留 → 轮询
<2000 MiB最长 15 分钟**不等显存归零就重部署会把新容器 KV 池压小**脚本支持环境变量
调参(`MEMFRAC/STEPS/TOPK/DRAFT/CTXLEN/CHUNK/MAXPRE/EXTRA/RESTART`),用法见脚本头注释。
## 方案 A-F 总表2026-09-08 六方案报告口径)
双场景报告(飞书 `SZUSdEqY1oRVxGxgILBcHqPJnEc`wiki `NzbMwzmKviYidRkZYRrc8GYQnrf`)已从
A/B 两配置扩展为六方案对比D/E/F 部署资产 09-08 入库。方案 FPD 分离链)占两台机,
吞吐须 **÷2 换算单机等效**后才与单机方案可比。
| 方案 | 部署脚本scripts/ | profiledeploy/profiles/pro6000/ | 实测机器 | 一句话结论 |
|---|---|---|---|---|
| ATP8+EAGLE | `deploy_glm53_605.sh` | `glm53_nvfp4_pro6000_sglang_tp8eagle.env` | 60.5/60.7 | 基准场景一最优60.5 生产在役 |
| BTP4PP2 | `deploy_glm53_optimal.sh` | `glm53_nvfp4_pro6000_sglang_tp4pp2_indexcache.env` | 60.7 | 场景二最优(吞吐 +41~79% |
| CB+IndexCache | 同 B`index_topk_freq=4` | 同 B | 60.7 | freq=4 为原生默认,与 B 恒等;报告口径保留 |
| DTP2PP4 | `deploy_glm53_pp4.sh` | `glm53_nvfp4_pro6000_sglang_tp2pp4.env` | 60.1 | 串行 prefill 有效速率最高6.5-6.7k);场景二 cc32+ 反超 B、场景一 8/8 败60.1 生产在役 |
| ETP8+DFlash2 | `deploy_glm53_tp8_dflash2.sh` | `glm53_nvfp4_pro6000_sglang_tp8dflash2.env` | 60.2 | 场景一 8 点全败于 A/CDFlash accept 2.53 < EAGLE 2.84劣势全在接受率投机栈维持 EAGLE |
| FPD 分离链 | `deploy_pd_smoke_master.sh`+`deploy_pd_probe.sh`+`deploy_s1_decode.sh`+`deploy_pd_smoke_router.sh`launch 脚本×3 | `glm53_nvfp4_pro6000_pd_{master,prefill,decode,router}.env` | 60.1+60.2 | 128k cc1 TTFT 3.91s 六方案最低÷2 后吞吐仅为 A 的 30-96%——"两机买 TTFT、不买吞吐" |
方案 F 配套资产:
- **链编排手册**(启动顺序 mc-master→prefill→decode→router、基础设施依赖表、质量门、拆链恢复
`deploy/PD_CHAIN.md`
- **sglang 补丁树快照**(宿主树无 .git此 patch 是唯一版本记录):`platforms/patches/pro6000/glm53_pd_chain/`
- **压测驱动**`scripts/run_pd_s1.sh` / `run_pd_s2.sh`(经 router :31000命中率读 prefill 日志、
accept 读 decode 日志)
- **防漂移核验**`bash deploy/verify_profile.sh <profile.env>`2026-09-08 已在 60.1 生产容器实测 PASS
- **现役状态页**`deploy/CURRENT.md`
## 基线数字2026-09-07真实语料60.7 干净机器)
**场景一 · 配置 A**run-id 9301-9308命中率 89.99%/89.94%0 失败 0 回撤):
@ -208,6 +233,22 @@ CAR 补丁(实验性)用法:先在**运行中的原生容器**上生成补
| `scripts/bench_report.sh` | 282f8cdf5873d3a3eb8a950a1564acae | 仅路径适配¹ |
| `scripts/bench_matrix.sh` | b961a56bc7c8f26c519fdc7f5f627ba1 | 仅路径适配¹ |
09-08 新增(方案 D/E/F服务器原版 md5 = 仓库版,逐字保留):
| 文件 | 60.1/60.2 原版 md5 | 说明 |
|---|---|---|
| `scripts/deploy_glm53_pp4.sh` | def3c64c5dc19e1d507080e3366c3762 | 方案 D 部署60.1 生产原样配方;脚本头注释为早期 TP4PP2 版残留,以 docker run 段为准) |
| `scripts/deploy_glm53_tp8_dflash2.sh` | 5bf2b47c9349e5855b963e571e35c096 | 方案 E v5 底稿(含四轮 OOM 教训bench 变体 4 处 delta 见 profile 注释) |
| `scripts/deploy_pd_smoke_master.sh` | 6e406234ba3c1634899dc89950d0139b | 方案 F 角色1mc-master |
| `scripts/deploy_pd_probe.sh` | 0d44fc8bf5b577ca0ea77c99f8f574cd | 方案 F 角色2prefill 通用探针(参数=launch 脚本名) |
| `scripts/p1b_prefill_tp4pp2_cps8k_launch.sh` | 006ffa5851eb188b19dde8c5b5dc296e | prefill launch 基线radix off |
| `scripts/p1b_prefill_tp4pp2_cps8k_radix_launch.sh` | 17822f059a93ecd54d0e309f8ef4ae02 | prefill launch 场景一变体(唯一差异=去 --disable-radix-cache |
| `scripts/deploy_s1_decode.sh` | 10ae7a186eb6eb309f1098c65fd87e1b | 方案 F 角色3decode 部署 |
| `scripts/s1_decode_launch.sh` | 932410b846f91ead3108122585c3cc34 | decode launchv5 配方+PD flags头注释"1 个文件改动"已过时,实为 11 文件,见补丁快照) |
| `scripts/deploy_pd_smoke_router.sh` | 77b3242ec6890402cbb4d28d79113fbb | 方案 F 角色4MiniLB router |
| `scripts/run_pd_s1.sh` | 889adc7aa9a75b58275e32052cebb2b8 | 方案 F 场景一压测驱动(经 router :31000 |
| `scripts/run_pd_s2.sh` | 980b74bb7b6fb3daa3e987d138b2c8ce | 方案 F 场景二压测驱动 |
¹ 仓库规范README 注意事项要求脚本不写绝对路径bench 脚本改按脚本所在目录解析,
语料/日志目录可用 `CORPUS`/`LOG_DIR` 环境变量覆盖,默认仍为 `/root`(服务器原布局,行为不变)。
@ -216,8 +257,11 @@ CAR 补丁(实验性)用法:先在**运行中的原生容器**上生成补
## 关联
- 部署 profile`python -m sskj.deploy` 消费):`deploy/profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`
`deploy/profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp4pp2_indexcache.env`
- 部署 profile方案 A-F 全量见上文"方案 A-F 总表";文件在 `deploy/profiles/pro6000/`
A=`tp8eagle`、B/C=`tp4pp2_indexcache`、D=`tp2pp4`、E=`tp8dflash2`、F=`pd_*` 四角色)
- 方案 F 编排:`deploy/PD_CHAIN.md`;补丁快照:`platforms/patches/pro6000/glm53_pd_chain/`
- 权重完整性清单:`deploy/manifests/GLM-5.3-NVFP4.md5``GLM-5.3-DFlash2.md5`
- 防漂移核验工具:`deploy/verify_profile.sh`;现役状态页:`deploy/CURRENT.md`
- 前序实验:`experiments/pro6000/glm53_nvfp4_pro6000d_sglang_tp4pp2_profile/`profile 与配置级证伪)、
`experiments/pro6000/glm53_nvfp4_pro6000d_sglang_ppmtp_deepdive/`PP+MTP 深挖,均在 hzy 分支)
- 场景一深度优化报告E7b CAR 补丁出处):`D:\sskj\reports\GLM53_NVFP4_60.7_场景一优化实验报告_2026-09-07.md`

View File

@ -0,0 +1,55 @@
#!/bin/bash
# ============================================================
# GLM-5.3 最优部署方案6000D 8卡TP4 PP2 + IndexCache freq=4
# 2026-09-07
#
# - 基线配置TP4 PP2 + cps16k + mem0.85131.6 tok/s 吞吐基线)
# - IndexCacheindex_topk_freq=4层轴索引复用省 75% indexer
# 16K 场景无损失128K 长上下文并发 1.35-1.47× 提速
# - 禁 radix cache禁投机解码PP2 与投机框架不兼容,已实测)
#
# 用法sudo bash deploy_glm53_optimal.sh (在 174.1.60.1 节点上执行)
# ============================================================
set -uo pipefail
CONTAINER="glm53-pp4"
IMAGE="lmsysorg/sglang:nightly-dev-20260828-daf63171"
MODEL="/data/hf_models/GLM-5.3-NVFP4"
PORT=30000
TP=2; PP=4; MEM=0.85; MRR=48; CPS=16384
docker rm -f ${CONTAINER} 2>/dev/null || true
sleep 2
docker run -d --name ${CONTAINER} --gpus all --shm-size 64g --ipc=host \
--restart unless-stopped \
-p ${PORT}:${PORT} \
-v /data/hf_models:/data/hf_models \
${IMAGE} \
python3 -m sglang.launch_server \
--model-path ${MODEL} \
--tp-size ${TP} --pp-size ${PP} \
--mem-fraction-static ${MEM} \
--max-running-requests ${MRR} \
--disable-radix-cache \
--disable-shared-experts-fusion \
--moe-runner-backend flashinfer_cutlass \
--disable-flashinfer-autotune \
--disable-custom-all-reduce \
--chunked-prefill-size ${CPS} \
--host 0.0.0.0 --port ${PORT} \
--json-model-override-args '{"index_topk_freq": 4}'
echo "容器已启动,等待就绪..."
for i in $(seq 1 60); do
code=$(curl -s -o /dev/null -w '%{http_code}' http://localhost:${PORT}/health 2>/dev/null)
if [ "$code" = "200" ]; then
echo "READY after ${i}0s"
docker ps --filter name=${CONTAINER} --format '{{.Names}} {{.Status}}'
# 确认 override 生效
echo "override args: $(docker inspect ${CONTAINER} --format '{{.Config.Cmd}}' | grep -o 'index_topk_freq[^,}]*')"
exit 0
fi
sleep 10
done
echo "TIMEOUT"; exit 1

View File

@ -0,0 +1,65 @@
#!/bin/bash
# GLM-5.3-NVFP4 TP8 + DFlash2 speculative decoding (6000D 8卡)
# 2026-09-07 v5(最终可用配置)
# - TP8 单节点, DFLASH block-diffusion draft (incoai/GLM-5.3-DFlash2, 7 draft tokens/step)
# - 四轮 OOM 战役教训(v1 mem0.85/cps16K → v2 mem0.80/cps16K → v3 mem0.85/cps8192 全崩在
# bench warmup 第一条 16K prompt):
# ① flashinfer_cutlass MoE workspace 随 chunk 尺寸走:16K=3.0GB, 8K=1.54GB,须落在静态池外空闲
# ② mem 与 cps 互相拆台:mem↑ 池大但空闲小, cps↓ workspace 小
# ③ 真正大头 = target verify CUDA graph 按 bs≤MRR 捕获(MRR48 时吃 4.04GB),
# 而 TP8 KV 池容量只支撑 ~14 并发(每 rank 装全部 78 层 KV, 池=TP2PP4 的 1/4)
# - v5 两改: MRR 48→12(verify graph 砍到 bs≤12: 4.04→0.83GB 腾 3.2GB,这是真正起作用的杠杆;
# 池 243K>12×16.9K=203K 无节流无 retract) + --speculative-draft-window-size 2048
# (DFLASH compact draft KV cache;实测不缩池分配——仍镜像目标池 0.70GB,只改运行期窗口化
# 寻址,留用无害) + mem0.85/cps8192 维持。最终空闲 6.77GB vs 8K chunk 峰值需求 ~3.9GB
# - 备选(未用): draft 后端 fa4→flashinfer/triton 可让 draft KV 保持 fp8(kv_cache_dtype.py
# 的 bf16 覆盖是 fa4 专属, 注释 "fp8-capable backends keep the target dtype"),但只省 0.35GB
# 且偏离官方配方(draft 噪声 → 接收率风险),不如 window 开关省得多
set -uo pipefail
CONTAINER="glm53-tp8-dflash2"
IMAGE="lmsysorg/sglang:nightly-dev-20260828-daf63171"
MODEL="/data/hf_models/GLM-5.3-NVFP4"
DRAFT="/data/hf_models/GLM-5.3-DFlash2"
PORT=30000
docker rm -f ${CONTAINER} 2>/dev/null || true
sleep 2
docker run -d --name ${CONTAINER} --gpus all --shm-size 64g --ipc=host \
--restart unless-stopped \
-p ${PORT}:${PORT} \
-v /data/hf_models:/data/hf_models \
${IMAGE} \
python3 -m sglang.launch_server \
--model-path ${MODEL} \
--tp-size 8 \
--mem-fraction-static 0.85 \
--max-running-requests 12 \
--disable-radix-cache \
--disable-shared-experts-fusion \
--moe-runner-backend flashinfer_cutlass \
--disable-flashinfer-autotune \
--disable-custom-all-reduce \
--chunked-prefill-size 8192 \
--speculative-algorithm DFLASH \
--speculative-draft-model-path ${DRAFT} \
--speculative-draft-attention-backend fa4 \
--speculative-draft-window-size 2048 \
--host 0.0.0.0 --port ${PORT} \
--json-model-override-args '{"index_topk_freq": 4}'
echo "容器已启动,等待就绪..."
for i in $(seq 1 90); do
code=$(curl -s -o /dev/null -w '%{http_code}' http://localhost:${PORT}/health 2>/dev/null)
if [ "$code" = "200" ]; then
echo "READY after $((i*10))s"
docker ps --filter name=${CONTAINER} --format '{{.Names}} {{.Status}}'
exit 0
fi
if ! docker ps --format '{{.Names}}' | grep -q "^${CONTAINER}$"; then
echo "CONTAINER DIED before ready"; docker logs ${CONTAINER} 2>&1 | tail -20; exit 1
fi
sleep 10
done
echo "TIMEOUT"; exit 1

View File

@ -0,0 +1,36 @@
#!/bin/bash
# PD prefill 探针通用部署:参数 = launch 脚本名(/tmp 下)。容器名固定 glm53-pd-smoke-prefill。
# 前置:mc-master 已起;decode(-2)+router 由调用方另行启动。
set -uo pipefail
LAUNCH="$1"
CONTAINER="glm53-pd-smoke-prefill"
IMAGE="lmsysorg/sglang:nightly-dev-20260828-daf63171"
if ! curl -s -o /dev/null -m3 http://174.1.60.1:50051/ 2>/dev/null; then
if ! ss -tln | grep :50051 >/dev/null; then echo "FATAL: mc-master 50051 未监听"; exit 1; fi
fi
docker rm -f ${CONTAINER} 2>/dev/null || true
sleep 2
docker run -d --name ${CONTAINER} --gpus all --network host --ipc=host --shm-size 64g \
--ulimit memlock=-1 --device /dev/infiniband \
--restart unless-stopped \
-e MOONCAKE_MASTER=174.1.60.1:50051 \
-e MOONCAKE_PROTOCOL=rdma \
-v /data/hf_models:/data/hf_models \
-v /data/flashkda_deploy/wheels:/mc_wheels:ro \
-v /data/sglang_patch_glm53:/sgl-workspace/sglang \
-v /tmp/${LAUNCH}:/smoke_launch.sh:ro \
${IMAGE} bash /smoke_launch.sh
echo "PD probe(${LAUNCH}) 容器已启动,等待就绪..."
for i in $(seq 1 150); do
code=$(curl -s -o /dev/null -w '%{http_code}' http://localhost:30000/health 2>/dev/null)
if [ "$code" = "200" ]; then echo "PD PROBE READY after $((i*10))s"; exit 0; fi
if ! docker ps --format '{{.Names}}' | grep -q "^${CONTAINER}$"; then
echo "PD PROBE CONTAINER DIED"; docker logs ${CONTAINER} 2>&1 | tail -40; exit 1
fi
sleep 10
done
echo "PD PROBE TIMEOUT"; exit 1

View File

@ -0,0 +1,12 @@
#!/bin/bash
# PD 冒烟第一步:mc-master(Mooncake 元数据服务)on 6000D-1 = 174.1.60.1:50051
# 依据 feishu_docs/pd_separation.md 成功配方:启动顺序 mc-master -> prefill -> decode -> router
set -u
docker rm -f mc-master 2>/dev/null || true
sleep 1
docker run -d --name mc-master --network host \
--restart unless-stopped \
--entrypoint /opt/sglang/bin/mooncake_master \
lmsysorg/sglang:nightly-dev-20260828-daf63171
sleep 3
if ss -tln | grep -q :50051; then echo "MC_MASTER_UP (50051)"; else echo "MC_MASTER_DOWN"; docker logs mc-master 2>&1 | tail -10; fi

View File

@ -0,0 +1,16 @@
#!/bin/bash
# PD 冒烟 router(MiniLB)on 6000D-2:31000。前置:prefill(-1:30000) + decode(-2:30000) 均 READY。
# 命令形态照搬 feishu_docs/pd_separation.md 验证过的写法(prefill URL 后跟 bootstrap 端口)。
set -u
docker rm -f pd-smoke-router 2>/dev/null || true
sleep 1
docker run -d --name pd-smoke-router --network host \
--restart unless-stopped \
--entrypoint python3 \
lmsysorg/sglang:nightly-dev-20260828-daf63171 \
-m sglang_router.launch_router --pd-disaggregation --mini-lb \
--prefill http://174.1.60.1:30000 28800 \
--decode http://174.1.60.2:30000 \
--host 0.0.0.0 --port 31000
sleep 3
if ss -tln | grep -q :31000; then echo "ROUTER_UP (31000)"; else echo "ROUTER_DOWN"; docker logs pd-smoke-router 2>&1 | tail -15; fi

View File

@ -0,0 +1,31 @@
#!/bin/bash
# S1' PD + decode 侧 DFLASH 冷启动实验(-2)。前置:mc-master(-1) 起、prefill(-1) READY、
# 补丁树 /data/sglang_patch_glm53 已含 spec_info.py 接线。
set -uo pipefail
CONTAINER="glm53-s1-decode"
IMAGE="lmsysorg/sglang:nightly-dev-20260828-daf63171"
docker rm -f ${CONTAINER} 2>/dev/null || true
sleep 2
docker run -d --name ${CONTAINER} --gpus all --network host --ipc=host --shm-size 64g \
--ulimit memlock=-1 --device /dev/infiniband \
--restart unless-stopped \
-e MOONCAKE_MASTER=174.1.60.1:50051 \
-e MOONCAKE_PROTOCOL=rdma \
-v /data/hf_models:/data/hf_models \
-v /data/flashkda_deploy/wheels:/mc_wheels:ro \
-v /data/sglang_patch_glm53:/sgl-workspace/sglang \
-v /tmp/s1_decode_launch.sh:/smoke_launch.sh:ro \
${IMAGE} bash /smoke_launch.sh
echo "S1 decode 容器已启动,等待就绪..."
for i in $(seq 1 120); do
code=$(curl -s -o /dev/null -w '%{http_code}' http://localhost:30000/health 2>/dev/null)
if [ "$code" = "200" ]; then echo "S1 DECODE READY after $((i*10))s"; exit 0; fi
if ! docker ps --format '{{.Names}}' | grep -q "^${CONTAINER}$"; then
echo "S1 DECODE CONTAINER DIED"; docker logs ${CONTAINER} 2>&1 | tail -40; exit 1
fi
sleep 10
done
echo "S1 DECODE TIMEOUT"; exit 1

View File

@ -0,0 +1,17 @@
set -e
# P2 探针:PD prefill TP4PP2+DFLASH(拓扑变体:4卡/stage × 2 stage,cps 16384,其余同 P1-B 基线)。
pip install /mc_wheels/mooncake_transfer_engine_cuda13-0.3.12.post1-cp312-cp312-manylinux_2_28_x86_64.whl --no-deps -q
echo WHEEL_OK
python3 -c "import sglang; print('SGLANG_FILE', sglang.__file__)"
export SGLANG_DFLASH_PD_DRAFT_KV_TRANSFER=0
exec python3 -m sglang.launch_server --model-path /data/hf_models/GLM-5.3-NVFP4 --tp-size 4 --pp-size 2 \
--mem-fraction-static 0.78 --max-running-requests 48 --disable-radix-cache \
--disable-shared-experts-fusion --moe-runner-backend flashinfer_cutlass \
--disable-flashinfer-autotune --disable-custom-all-reduce --chunked-prefill-size 8192 \
--speculative-algorithm DFLASH \
--speculative-draft-model-path /data/hf_models/GLM-5.3-DFlash2 \
--speculative-draft-attention-backend fa4 \
--disaggregation-mode prefill --disaggregation-transfer-backend mooncake \
--disaggregation-bootstrap-port 28800 --disaggregation-ib-device mlx5_0,mlx5_1,mlx5_2,mlx5_3 \
--host 0.0.0.0 --port 30000 \
--json-model-override-args '{"index_topk_freq": 4}'

View File

@ -0,0 +1,17 @@
set -e
# P2 探针:PD prefill TP4PP2+DFLASH(拓扑变体:4卡/stage × 2 stage,cps 16384,其余同 P1-B 基线)。
pip install /mc_wheels/mooncake_transfer_engine_cuda13-0.3.12.post1-cp312-cp312-manylinux_2_28_x86_64.whl --no-deps -q
echo WHEEL_OK
python3 -c "import sglang; print('SGLANG_FILE', sglang.__file__)"
export SGLANG_DFLASH_PD_DRAFT_KV_TRANSFER=0
exec python3 -m sglang.launch_server --model-path /data/hf_models/GLM-5.3-NVFP4 --tp-size 4 --pp-size 2 \
--mem-fraction-static 0.78 --max-running-requests 48 \
--disable-shared-experts-fusion --moe-runner-backend flashinfer_cutlass \
--disable-flashinfer-autotune --disable-custom-all-reduce --chunked-prefill-size 8192 \
--speculative-algorithm DFLASH \
--speculative-draft-model-path /data/hf_models/GLM-5.3-DFlash2 \
--speculative-draft-attention-backend fa4 \
--disaggregation-mode prefill --disaggregation-transfer-backend mooncake \
--disaggregation-bootstrap-port 28800 --disaggregation-ib-device mlx5_0,mlx5_1,mlx5_2,mlx5_3 \
--host 0.0.0.0 --port 30000 \
--json-model-override-args '{"index_topk_freq": 4}'

View File

@ -0,0 +1,20 @@
#!/bin/bash
# run_pd_s1.sh — PD-P3 full chain scenario-1 8-point bench via MiniLB router (2026-09-08)
# prefill glm53-pd-smoke-prefill @60.1 (TP4PP2 radix-on) | decode glm53-s1-decode @60.2 (TP8+DFLASH)
# canonical windows 9301-9308 (virgin chain, identical inputs to A-E s1 rows)
# throughput numbers = chain aggregate; per-machine equivalent = aggregate/2
set -u
mkdir -p /root/bench_logs
code=$(curl -s -o /dev/null -w '%{http_code}' http://174.1.60.2:31000/health)
[ "$code" = "200" ] || { echo "router not healthy: $code"; exit 1; }
for spec in "131072 1 9301" "131072 2 9302" "131072 3 9303" "131072 4 9304" "65536 1 9305" "65536 2 9306" "65536 3 9307" "65536 4 9308"; do
set -- $spec; IL=$1; CC=$2; RID=$3
echo "[$(date +%H:%M:%S)] start rid=$RID il=$IL cc=$CC" | tee -a /root/bench_logs/pd_runner.log
python3 /root/bench_corpus.py --corpus /root/corpus_ids.json \
--url http://174.1.60.2:31000/generate \
--input-len $IL --concurrency $CC --num-requests 8 --run-id $RID \
--shared-frac 0.9 --output-len 512 --container glm53-pd-smoke-prefill \
> /root/bench_logs/pd_s1_run${RID}.log 2>&1
echo "[$(date +%H:%M:%S)] done rid=$RID rc=$?" | tee -a /root/bench_logs/pd_runner.log
done
echo ALL_DONE | tee -a /root/bench_logs/pd_runner.log

View File

@ -0,0 +1,27 @@
#!/bin/bash
# run_pd_s2.sh — PD-P3 full chain scenario-2 5-point bench via MiniLB router (2026-09-08)
# cc8/16/32 = canonical 9311/9312/9313 (identical windows to A/D s2 rows)
# cc40/64 = fresh pool-override windows 16384000/17203200, nreq 40/64 (A/B/C convention)
# run-id 9330/9331 label-only under override. Throughput = chain aggregate; /2 = per-machine.
set -u
mkdir -p /root/bench_logs
code=$(curl -s -o /dev/null -w '%{http_code}' http://174.1.60.2:31000/health)
[ "$code" = "200" ] || { echo "router not healthy: $code"; exit 1; }
run_point() {
CC=$1; NR=$2; RID=$3; PO=$4
EXTRA=""
[ -n "$PO" ] && EXTRA="--pool-override $PO"
echo "=== cc=$CC nreq=$NR run=$RID pool=$PO start $(date +%T) ===" | tee -a /root/bench_logs/pd_runner.log
python3 /root/bench_corpus.py --corpus /root/corpus_ids.json \
--url http://174.1.60.2:31000/generate \
--input-len 16384 --concurrency $CC --num-requests $NR --run-id $RID \
--shared-frac 0 --output-len 512 --container glm53-pd-smoke-prefill $EXTRA \
> /root/bench_logs/pd_s2_run${RID}.log 2>&1
echo "=== run=$RID done rc=$? $(date +%T) ===" | tee -a /root/bench_logs/pd_runner.log
}
run_point 8 16 9311 ""
run_point 16 32 9312 ""
run_point 32 32 9313 ""
run_point 40 40 9330 16384000
run_point 64 64 9331 17203200
echo ALL_DONE | tee -a /root/bench_logs/pd_runner.log

View File

@ -0,0 +1,18 @@
set -e
# S1' decode 侧(6000D-2 容器内):参照 v5 DFLASH 参数 + PD decode flags + 挂载补丁树
# 补丁树当前含 1 个文件改动:spec_info.py 接 build_dflash_family_disagg_draft_input(冷启动接线)
pip install /mc_wheels/mooncake_transfer_engine_cuda13-0.3.12.post1-cp312-cp312-manylinux_2_28_x86_64.whl --no-deps -q
echo WHEEL_OK
python3 -c "import sglang; print('SGLANG_FILE', sglang.__file__)"
exec python3 -m sglang.launch_server --model-path /data/hf_models/GLM-5.3-NVFP4 --tp-size 8 \
--mem-fraction-static 0.85 --max-running-requests 12 --disable-radix-cache \
--disable-shared-experts-fusion --moe-runner-backend flashinfer_cutlass \
--disable-flashinfer-autotune --disable-custom-all-reduce --chunked-prefill-size 8192 \
--speculative-algorithm DFLASH \
--speculative-draft-model-path /data/hf_models/GLM-5.3-DFlash2 \
--speculative-draft-attention-backend fa4 \
--speculative-draft-window-size 2048 \
--disaggregation-mode decode --disaggregation-transfer-backend mooncake \
--disaggregation-bootstrap-port 28800 --disaggregation-ib-device mlx5_0,mlx5_1,mlx5_2,mlx5_3 \
--host 0.0.0.0 --port 30000 \
--json-model-override-args '{"index_topk_freq": 4}'

View File

@ -0,0 +1,44 @@
# glm53_pd_chainsglang 补丁树 vs 镜像原版差异2026-09-08 快照)
## 这是什么
生产/实验容器通过 `-v /data/sglang_patch_glm53:/sgl-workspace/sglang` 挂载的补丁源码树
(宿主机 60.1/60.2 均有,无 `.git`)。本目录的 `sglang_patch_vs_image_20260908.patch`
是它与镜像 `lmsysorg/sglang:nightly-dev-20260828-daf63171`digest
`sha256:28e0d26073161e49ca56eba808d264a4223804a212020f1dfe1b2405b9f8a399`)内置原版源码的
统一 diff是这套补丁**唯一的版本记录**(宿主机树上没有 git 历史可查)。
生成方式60.2 上,容器化 diffpristine 在前保证 patch 语义为"打在原版上"
```bash
docker run --rm -v /data/sglang_patch_glm53:/host_patch:ro -v /tmp:/host_tmp \
lmsysorg/sglang:nightly-dev-20260828-daf63171 bash -c \
'diff -ruN --exclude=__pycache__ --exclude="._*" --exclude=.git \
/sgl-workspace/sglang /host_patch > /host_tmp/pd_chain.patch'
```
## 改动清单11 文件10 改 + 1 新增)
主题 = **DFlash 家族DFLASH/DSPARK+ PP 流水 + PD 分离** 三者组合的解锁与修错。
| 文件python/sglang/srt/ 下) | 改动 |
|---|---|
| `arg_groups/speculative_hook.py` | 放开 PP 断言DFLASH/DSPARK 允许 pp>1限 PD prefill/decode server |
| `disaggregation/prefill.py` | 草稿 KV 跨机传输 opt-outenv `SGLANG_DFLASH_PD_DRAFT_KV_TRANSFER=0`PP 下仅最后一个 rank 持活草稿池、仅它注册 draft bufferTP 不匹配时MHA 头分片 → kv_item_lens 不同)会 RDMA 段失败/静默错位——注释里有完整推导 |
| `managers/scheduler_pp_mixin.py` | PP 末 rank 为 DFlash 家族构造 next_draft_input 并经 RelayPayload 中继topk_p/index/hidden_statesEAGLE 草稿中继路径dflash 前后 set/clear pp_proxy_tensors |
| `managers/scheduler.py` | PP 投机分支:仅 pp_size==1 时做 D2H copy_to_cpu末 rank 结果是 device 张量送 rank 0非末 stage next_token_ids=None 直接拷会崩) |
| `model_executor/model_runner_components/layer_setup.py` | MTP 层守卫放行 dflash 家族(独立草稿模型不用目标模型原生 MTP 层) |
| `models/deepseek_v2.py` | DFLASH PP prefill各 stage 原始 aux 捕获打包进 proxy_tensors["dspark_aux_hidden_states"]`set_dflash_layers_to_capture` 改为按 consumer 入口捕获(+1 偏移,各 rank 只留本地 consumer |
| `models/dflash.py` | 新增 `project_target_hidden_partial`:各 PP rank 只对自己捕获的特征列做 fc 部分投影,末 stage 求和并只做一次 hidden_norm数学等价切片 |
| `server_args.py` | PP + 投机断言改写:仅允许 PD prefill server + DFLASH/DSPARK#33863 |
| `speculative/dflash_pp.py` **新增** | `kimi_pp_capture_layer_ids`Kimi 后置层流捕获归属下一 stage 的辅助函数 |
| `speculative/dflash_worker_v2.py` | PD-prefill 非末 PP rank = context-only不做草稿 forward、无草稿 KV只投影本地捕获特征草稿 worker 以 pp_size=1 构建;`_init_pp_context_features` |
| `speculative/spec_info.py` | DFlash PD 冷启动接线disagg 草稿输入走 `build_dflash_family_disagg_draft_input`——无此项 decode 首请求 400 |
## 如何使用
- **部署**:不 apply patch直接整树挂载`deploy/PD_CHAIN.md` 基础设施依赖表)。
- **审阅/重建**:把镜像源码导出后 `patch -d <sglang仓根> -p2 < sglang_patch_vs_image_20260908.patch`
diff 两侧绝对路径去掉前两段后即仓内相对路径)。
- **漂移检测**:重跑上面的容器化 diff与库内 patch 比对;不一致说明宿主机补丁树又被人改过,
需要重新快照入库。

View File

@ -0,0 +1,665 @@
diff -ruN '--exclude=__pycache__' '--exclude=._*' '--exclude=.git' /sgl-workspace/sglang/python/sglang/srt/arg_groups/speculative_hook.py /host_patch/python/sglang/srt/arg_groups/speculative_hook.py
--- /sgl-workspace/sglang/python/sglang/srt/arg_groups/speculative_hook.py 2026-08-28 03:46:53.000000000 +0000
+++ /host_patch/python/sglang/srt/arg_groups/speculative_hook.py 2026-09-07 12:26:53.359744593 +0000
@@ -196,9 +196,9 @@
"Currently DFLASH speculative decoding does not support dp attention."
)
- if cfg.pp_size != 1:
+ if cfg.pp_size != 1 and cfg.disaggregation_mode != "prefill":
raise ValueError(
- "Currently DFLASH speculative decoding only supports pp_size == 1."
+ "DFLASH with pp_size > 1 is only supported on a PD prefill server."
)
if cfg.speculative_draft_model_path is None:
@@ -390,9 +390,13 @@
f"(got {cfg.speculative_moe_a2a_backend!r})."
)
- if cfg.pp_size != 1:
+ if cfg.pp_size != 1 and cfg.disaggregation_mode not in (
+ "prefill",
+ "decode",
+ ):
raise ValueError(
- "Currently DSpark speculative decoding only supports pp_size == 1."
+ "Currently DSpark speculative decoding with pp_size > 1 is only "
+ "supported under PD disaggregation."
)
if cfg.speculative_draft_model_path is None:
diff -ruN '--exclude=__pycache__' '--exclude=._*' '--exclude=.git' /sgl-workspace/sglang/python/sglang/srt/disaggregation/prefill.py /host_patch/python/sglang/srt/disaggregation/prefill.py
--- /sgl-workspace/sglang/python/sglang/srt/disaggregation/prefill.py 2026-08-28 03:46:53.000000000 +0000
+++ /host_patch/python/sglang/srt/disaggregation/prefill.py 2026-09-07 13:36:24.566851061 +0000
@@ -21,6 +21,7 @@
import hashlib
import logging
+import os
from array import array
from collections import deque
from http import HTTPStatus
@@ -202,9 +203,26 @@
)
layer_shard_rank = getattr(self.token_to_kv_pool, "layer_shard_rank", None)
layer_shard_size = getattr(self.token_to_kv_pool, "layer_shard_size", 1)
+ # Under PP, only the last rank owns a live DFLASH draft pool (earlier
+ # ranks hold page-size stubs); registering their buffers would push
+ # draft entries at the wrong layer offset into the decode KV layout.
+ # The draft KV transfer additionally requires the prefill and decode
+ # TP sizes to yield identical per-rank draft cells: DFlash draft
+ # attention is MHA head-sharded (num_kv_heads = total // tp), so e.g.
+ # a TP2 prefill and a TP8 decode over an 8-kv-head draft register
+ # 4x-different kv_item_lens and every draft write lands past the
+ # receiver's slots (RDMA segment failure on long prompts, silent
+ # misalignment on short ones). The target MLA KV is replicated and
+ # unaffected. Opt out per deployment when the TP layouts mismatch.
transfer_draft_cache = (
+ self.pp_size <= 1 or self.pp_rank == self.pp_size - 1
+ ) and (
not layer_shard_enabled or layer_shard_rank == layer_shard_size - 1
)
+ if transfer_draft_cache and os.environ.get(
+ "SGLANG_DFLASH_PD_DRAFT_KV_TRANSFER", "1"
+ ) == "0":
+ transfer_draft_cache = False
kv_args.prefill_start_layer = (
getattr(
self.token_to_kv_pool,
diff -ruN '--exclude=__pycache__' '--exclude=._*' '--exclude=.git' /sgl-workspace/sglang/python/sglang/srt/managers/scheduler_pp_mixin.py /host_patch/python/sglang/srt/managers/scheduler_pp_mixin.py
--- /sgl-workspace/sglang/python/sglang/srt/managers/scheduler_pp_mixin.py 2026-08-28 03:46:53.000000000 +0000
+++ /host_patch/python/sglang/srt/managers/scheduler_pp_mixin.py 2026-09-07 12:26:53.349255446 +0000
@@ -1172,17 +1172,61 @@
logits_output = LogitsProcessorOutput(next_token_logits=None)
logits_output.auxiliary_device_output = auxiliary_output
next_token_ids = pp_outputs["next_token_ids"].to(torch.int64)
+ next_draft_input = None
+ if isinstance(batch, ScheduleBatch) and batch.spec_algorithm.is_dflash_family():
+ next_token_ids = next_token_ids.to(
+ device=batch.device,
+ dtype=torch.int64,
+ non_blocking=True,
+ )
+ from sglang.srt.speculative.dspark_components.dspark_draft import (
+ make_next_draft_input,
+ )
+
+ if batch.spec_algorithm.is_dflash():
+ from sglang.srt.speculative.draft_worker_common import (
+ make_draft_input_v2 as make_next_draft_input,
+ )
+
+ next_draft_input = make_next_draft_input(
+ bonus_tokens=next_token_ids,
+ new_seq_lens=batch.seq_lens,
+ )
+ batch.spec_info = next_draft_input
+ elif "draft_topk_p" in pp_outputs.tensors:
+ from sglang.srt.speculative.eagle_info import EagleDraftInput
+
+ next_draft_input = EagleDraftInput(
+ topk_p=pp_outputs["draft_topk_p"],
+ topk_index=pp_outputs["draft_topk_index"],
+ hidden_states=pp_outputs["draft_hidden_states"],
+ bonus_tokens=next_token_ids,
+ num_tokens_per_req=1,
+ num_tokens_for_logprob_per_req=1,
+ )
+ batch.spec_info = next_draft_input
# PP rank 0 also relays into output_tokens_buf so the next iter's
# resolve_forward_inputs finds these tokens for the decode portion
# of mixed-chunk batches (which gather via mix_running_indices).
self.future_map.stash(
- batch.req_pool_indices, RelayPayload(bonus_tokens=next_token_ids)
+ batch.req_pool_indices,
+ RelayPayload(
+ bonus_tokens=next_token_ids,
+ topk_p=None if next_draft_input is None else next_draft_input.topk_p,
+ topk_index=(
+ None if next_draft_input is None else next_draft_input.topk_index
+ ),
+ hidden_states=(
+ None if next_draft_input is None else next_draft_input.hidden_states
+ ),
+ ),
)
batch.input_ids = None
output_result = GenerationBatchResult(
logits_output=logits_output,
pp_hidden_states_proxy_tensors=None,
next_token_ids=pp_outputs["next_token_ids"],
+ next_draft_input=next_draft_input,
extend_input_len_per_req=extend_input_len_per_req,
extend_logprob_start_len_per_req=extend_logprob_start_len_per_req,
can_run_cuda_graph=mb_metadata.can_run_cuda_graph,
@@ -1315,7 +1359,15 @@
"set_run_batch_cpu_start_time",
trace_only=True,
)
- result = self.run_batch(cur_batch, pp_proxy_tensors)
+ if cur_batch.spec_algorithm.is_dflash_family():
+ self.model_worker.set_pp_proxy_tensors_for_next_forward(
+ pp_proxy_tensors
+ )
+ try:
+ result = self.run_batch(cur_batch, pp_proxy_tensors)
+ finally:
+ if cur_batch.spec_algorithm.is_dflash_family():
+ self.model_worker.set_pp_proxy_tensors_for_next_forward(None)
set_time_batch(
cur_batch.reqs,
"set_run_batch_cpu_end_time",
diff -ruN '--exclude=__pycache__' '--exclude=._*' '--exclude=.git' /sgl-workspace/sglang/python/sglang/srt/managers/scheduler.py /host_patch/python/sglang/srt/managers/scheduler.py
--- /sgl-workspace/sglang/python/sglang/srt/managers/scheduler.py 2026-08-28 03:46:53.000000000 +0000
+++ /host_patch/python/sglang/srt/managers/scheduler.py 2026-09-07 13:22:04.673564095 +0000
@@ -3961,11 +3961,19 @@
batch.input_ids = None # rebuilt next iter from draft_token
self.update_cache_from_scheduler(batch, batch_result)
# Sync D2H so the result processor can read CPU tensors.
- batch_result.copy_done = self.device_module.Event()
- batch_result.copy_to_cpu(
- return_logprob=batch.return_logprob,
- return_hidden_states=batch.return_hidden_states,
- )
+ # PP mode: the last-rank result is packed as on-device tensors
+ # and sent to rank 0, whose pp-mixin path owns the D2H copies
+ # (copy_stream_ctx + d2h_event) — exactly like the non-spec
+ # branch below, which never calls copy_to_cpu in run_batch.
+ # Non-final stages have next_token_ids=None, so the
+ # unconditional copy would crash; the last rank's copy would
+ # also hand a CPU tensor to the device-only pp output send.
+ if get_parallel().pp_size == 1:
+ batch_result.copy_done = self.device_module.Event()
+ batch_result.copy_to_cpu(
+ return_logprob=batch.return_logprob,
+ return_hidden_states=batch.return_hidden_states,
+ )
else:
kwargs = (
{"pp_proxy_tensors": pp_proxy_tensors}
diff -ruN '--exclude=__pycache__' '--exclude=._*' '--exclude=.git' /sgl-workspace/sglang/python/sglang/srt/model_executor/model_runner_components/layer_setup.py /host_patch/python/sglang/srt/model_executor/model_runner_components/layer_setup.py
--- /sgl-workspace/sglang/python/sglang/srt/model_executor/model_runner_components/layer_setup.py 2026-08-28 03:46:53.000000000 +0000
+++ /host_patch/python/sglang/srt/model_executor/model_runner_components/layer_setup.py 2026-09-07 12:42:30.789352374 +0000
@@ -200,6 +200,10 @@
assert (
(not model_has_mtp_layers)
or (spec_algorithm.is_none())
+ # DFlash-family drafts are standalone models; the target's native MTP
+ # layer stays unused, so a PP-split target with DFLASH/DSPARK (#33863)
+ # does not hit the native-MTP-drafting hazard this guard exists for.
+ or (spec_algorithm.is_dflash_family())
or (
(not spec_algorithm.is_none())
and (num_effective_layers == model_num_layers)
diff -ruN '--exclude=__pycache__' '--exclude=._*' '--exclude=.git' /sgl-workspace/sglang/python/sglang/srt/models/deepseek_v2.py /host_patch/python/sglang/srt/models/deepseek_v2.py
--- /sgl-workspace/sglang/python/sglang/srt/models/deepseek_v2.py 2026-08-28 03:46:53.000000000 +0000
+++ /host_patch/python/sglang/srt/models/deepseek_v2.py 2026-09-07 13:04:54.575535893 +0000
@@ -2938,6 +2938,20 @@
(0, get_dsa_index_topk(self.config)), dtype=torch.int32
)
proxy_tensors["topk_indices"] = topk_indices
+ # DFLASH PP prefill: hand this stage's raw aux captures to the next
+ # stage on the wire; the spec worker swaps them for the accumulated
+ # partial projection (dflash_ctx_acc) before the send leaves.
+ # Model-scope capture flag: ForCausalLM owns capture_aux_hidden_states,
+ # the model owns layers_to_capture — this is DeepseekV2Model.forward.
+ if len(self.layers_to_capture) > 0:
+ if len(aux_hidden_states) > 0:
+ proxy_tensors["dspark_aux_hidden_states"] = (
+ aux_hidden_states.finalize()
+ )
+ else:
+ proxy_tensors["dspark_aux_hidden_states"] = hidden_states.new_empty(
+ hidden_states.shape[0], 0
+ )
return PPProxyTensors(proxy_tensors)
else:
if not forward_batch.forward_mode.is_idle():
@@ -3159,11 +3173,12 @@
hidden_states = self.model(
input_ids, positions, forward_batch, input_embeds, pp_proxy_tensors
)
- aux_hidden_states = None
- if self.capture_aux_hidden_states:
- hidden_states, aux_hidden_states = hidden_states
-
if self.pp_group.is_last_rank:
+ # Under PP the model returns PPProxyTensors on non-last ranks; the
+ # (hidden, aux) capture tuple only exists on the last rank (#33863).
+ aux_hidden_states = None
+ if self.capture_aux_hidden_states:
+ hidden_states, aux_hidden_states = hidden_states
return self.logits_processor(
input_ids, hidden_states, self.lm_head, forward_batch, aux_hidden_states
)
@@ -3218,16 +3233,27 @@
self.model.layers_to_capture = list(layer_ids)
def set_dflash_layers_to_capture(self, layer_ids: List[int]):
- if not self.pp_group.is_last_rank:
- return
-
if layer_ids is None:
raise ValueError(
"DFLASH requires explicit layer_ids for aux hidden capture."
)
- self.capture_aux_hidden_states = True
- self.model.layers_to_capture = [val + 1 for val in layer_ids]
+ # Capture at consumer entry: the stream of layer i is captured inside
+ # layer min(i+1, L-1)'s prepare_attn (the +1 shift), so under PP every
+ # capture stays local to the stage that owns the consumer layer and a
+ # PP-boundary capture (consumer == start_layer) reads the incoming
+ # proxy residual. Each rank keeps only its own consumer ids.
+ num_layers = self.config.num_hidden_layers
+ consumer_ids = [min(int(v) + 1, num_layers - 1) for v in layer_ids]
+ local_consumer_ids = sorted(
+ {
+ c
+ for c in consumer_ids
+ if self.model.start_layer <= c < self.model.end_layer
+ }
+ )
+ self.capture_aux_hidden_states = bool(local_consumer_ids)
+ self.model.layers_to_capture = local_consumer_ids
def prepare_context_parallel_metadata_for_dcp(
self,
diff -ruN '--exclude=__pycache__' '--exclude=._*' '--exclude=.git' /sgl-workspace/sglang/python/sglang/srt/models/dflash.py /host_patch/python/sglang/srt/models/dflash.py
--- /sgl-workspace/sglang/python/sglang/srt/models/dflash.py 2026-08-28 03:46:53.000000000 +0000
+++ /host_patch/python/sglang/srt/models/dflash.py 2026-09-07 12:26:53.334773984 +0000
@@ -679,6 +679,45 @@
projected = projected[0]
return self.hidden_norm(projected)
+ def project_target_hidden_partial(
+ self, target_hidden: torch.Tensor, feature_indices: list[int]
+ ) -> torch.Tensor:
+ """Project only this PP rank's captured features through their fc columns.
+
+ Mathematically exact slice of project_target_hidden's fc matmul:
+ concat(h) @ W.T == sum_i (h_i @ W_i.T); the caller sums the per-rank
+ partials and applies hidden_norm exactly once on the final stage.
+ """
+ if not feature_indices:
+ raise ValueError("feature_indices must be non-empty.")
+ feature_indices = [int(i) for i in feature_indices]
+ if (
+ min(feature_indices) < 0
+ or max(feature_indices) >= self.num_context_features
+ ):
+ raise ValueError(
+ "feature_indices out of range for DFLASH context projection: "
+ f"{feature_indices=} {self.num_context_features=}."
+ )
+ hidden_size = int(self.config.hidden_size)
+ expected = len(feature_indices) * hidden_size
+ if target_hidden.ndim != 2 or int(target_hidden.shape[-1]) != expected:
+ raise ValueError(
+ "DFLASH partial target_hidden feature dim mismatch. "
+ f"Expected shape [N, {expected}] for {feature_indices=}, "
+ f"but got shape={tuple(target_hidden.shape)}."
+ )
+
+ cols = []
+ for idx in feature_indices:
+ start = idx * hidden_size
+ cols.extend(range(start, start + hidden_size))
+ index = torch.tensor(cols, dtype=torch.long, device=self.fc.weight.device)
+ weight = self.fc.weight.index_select(1, index)
+ if target_hidden.dtype != weight.dtype:
+ target_hidden = target_hidden.to(weight.dtype)
+ return F.linear(target_hidden, weight)
+
@torch.no_grad()
def forward(
self,
diff -ruN '--exclude=__pycache__' '--exclude=._*' '--exclude=.git' /sgl-workspace/sglang/python/sglang/srt/server_args.py /host_patch/python/sglang/srt/server_args.py
--- /sgl-workspace/sglang/python/sglang/srt/server_args.py 2026-08-28 03:46:53.000000000 +0000
+++ /host_patch/python/sglang/srt/server_args.py 2026-09-07 12:36:00.443055298 +0000
@@ -10308,9 +10308,19 @@
)
if cfg.pp_size > 1:
- assert (
- cfg.disable_overlap_schedule and cfg.speculative_algorithm is None
- ), "Pipeline parallelism is not compatible with overlap schedule, speculative decoding"
+ assert cfg.disable_overlap_schedule, (
+ "Pipeline parallelism is not compatible with overlap schedule."
+ )
+ if cfg.speculative_algorithm is not None:
+ # #33863: PP + speculative is allowed only as a PD prefill
+ # server running a dflash-family draft warmup.
+ assert (
+ cfg.disaggregation_mode == "prefill"
+ and cfg.speculative_algorithm in ("DFLASH", "DSPARK")
+ ), (
+ "Pipeline parallelism with speculative decoding is only "
+ "supported on a PD prefill server with DFLASH/DSPARK."
+ )
assert cfg.min_free_slots_delay is None, (
"--min-free-slots-delay is not supported with pipeline "
"parallelism: allocatable slots per microbatch are bounded by "
diff -ruN '--exclude=__pycache__' '--exclude=._*' '--exclude=.git' /sgl-workspace/sglang/python/sglang/srt/speculative/dflash_pp.py /host_patch/python/sglang/srt/speculative/dflash_pp.py
--- /sgl-workspace/sglang/python/sglang/srt/speculative/dflash_pp.py 1970-01-01 00:00:00.000000000 +0000
+++ /host_patch/python/sglang/srt/speculative/dflash_pp.py 2026-09-07 12:26:53.339334845 +0000
@@ -0,0 +1,19 @@
+def kimi_pp_capture_layer_ids(layer_ids, start_layer, end_layer, num_layers):
+ """Own stream captures where their next consumer's weights are local.
+
+ Kimi's post-layer stream uses the next layer's attention-residual weights.
+ A PP-boundary capture therefore belongs to the next stage, before its first
+ decoder layer, rather than to the stage that just produced the raw stream.
+ """
+ ids = list(layer_ids)
+ if not ids or any(type(i) is not int for i in ids):
+ raise ValueError("DFLASH requires explicit integer capture layer IDs")
+ if len(ids) != len(set(ids)) or min(ids) < 0 or max(ids) >= num_layers:
+ raise ValueError("DFLASH capture layer IDs must be unique and in range")
+ if ids != sorted(ids):
+ raise ValueError("Kimi DFLASH capture layer IDs must follow model layer order")
+ if not 0 <= start_layer < end_layer <= num_layers:
+ raise ValueError("Invalid Kimi PP layer range")
+ return sorted(
+ i for i in ids if start_layer <= min(i + 1, num_layers - 1) < end_layer
+ )
diff -ruN '--exclude=__pycache__' '--exclude=._*' '--exclude=.git' /sgl-workspace/sglang/python/sglang/srt/speculative/dflash_worker_v2.py /host_patch/python/sglang/srt/speculative/dflash_worker_v2.py
--- /sgl-workspace/sglang/python/sglang/srt/speculative/dflash_worker_v2.py 2026-08-28 03:46:53.000000000 +0000
+++ /host_patch/python/sglang/srt/speculative/dflash_worker_v2.py 2026-09-07 12:26:53.343964789 +0000
@@ -34,6 +34,7 @@
compute_position,
)
from sglang.srt.runtime_context import (
+ get_disagg,
get_exec,
get_schedule,
get_spec,
@@ -298,6 +299,16 @@
self.draft_window_size: Optional[int] = get_spec().speculative_draft_window_size
self.use_compact_draft_cache = self.draft_window_size is not None
self.device = target_worker.device
+ # PD-prefill PP ranks before the last hold the draft only to project
+ # their locally captured context features (no draft forward, no draft
+ # KV); see _forward_pp_prefill.
+ self._is_pd_prefill = get_disagg().disaggregation_mode == "prefill"
+ self._is_context_only_pp_prefill_rank = (
+ self._is_pd_prefill and ps.pp_rank < ps.pp_size - 1
+ )
+ self._next_pp_proxy_tensors = None
+ self._pp_context_feature_indices: list = []
+ self._pp_expects_incoming_context = False
self._warned_sampling_fallback = False
self._draft_probs_buf = None
@@ -306,7 +317,7 @@
bundle = build_draft_tp_worker(
server_args=server_args,
gpu_id=gpu_id,
- ps=replace(ps, pp_rank=0),
+ ps=replace(ps, pp_rank=0, pp_size=1),
nccl_port=nccl_port,
target_model_config=target_worker.model_runner.model_config,
algo_label="DFLASH",
@@ -315,6 +326,8 @@
self.draft_model_runner = bundle.draft_model_runner
self._draft_sampler = None
self.draft_model = bundle.draft_model
+ if ps.pp_size > 1:
+ self._init_pp_context_features()
self.selector = self.draft_model.candidate_selector
draft_config = parse_dflash_draft_config(
draft_hf_config=self.draft_model_runner.model_config.hf_config
@@ -427,6 +440,8 @@
def spec_v2_attn_backends(self) -> tuple:
# Every attn backend a spec_v2 forward touches; consumed by
# decide_needs_cpu_seq_lens to gate the seq_lens_cpu D2H.
+ if self._is_context_only_pp_prefill_rank:
+ return (self._target_worker.model_runner.attn_backend,)
return (
self._target_worker.model_runner.attn_backend,
self.draft_model_runner.attn_backend,
@@ -443,6 +458,23 @@
# enabled, the draft worker keeps a private compact req->token table
# over the same global KV index space, so radix-cache/prefix-hit KV
# remains reusable while draft attention sees only the recent window.
+ if memory_pool_config is not None and self._is_context_only_pp_prefill_rank:
+ # Context-only PP prefill ranks never run a draft forward; clamp
+ # the draft pool to a one-page stub.
+ memory_pool_config = replace(
+ memory_pool_config,
+ max_total_num_tokens=self.page_size,
+ full_max_total_num_tokens=(
+ self.page_size
+ if memory_pool_config.full_max_total_num_tokens
+ else memory_pool_config.full_max_total_num_tokens
+ ),
+ swa_max_total_num_tokens=(
+ self.page_size
+ if memory_pool_config.swa_max_total_num_tokens
+ else memory_pool_config.swa_max_total_num_tokens
+ ),
+ )
self._draft_worker.alloc_memory_pool(
memory_pool_config=memory_pool_config,
req_to_token_pool=(
@@ -451,7 +483,49 @@
token_to_kv_pool_allocator=token_to_kv_pool_allocator,
)
+ def legacy_physical_transfer_locs(
+ self, req_pool_idx: int, start: int, end: int
+ ) -> torch.Tensor:
+ """Resolve absolute draft positions in the legacy full physical pool.
+
+ Non-compact DFlash materializes draft K/V at the target allocator's
+ physical indices. Prefill and decode request-slot/allocator choices are
+ role-local, so each peer must gather its own ``req_to_token`` suffix;
+ sending either role's raw indices to the other would address unrelated
+ rows. Slot zero is the padding sentinel and is never a valid transfer
+ destination.
+ """
+ if start < 0 or end < start:
+ raise ValueError(f"invalid transfer range: start={start}, end={end}")
+
+ req_to_token = self.model_runner.req_to_token_pool.req_to_token
+ num_owners, table_width = req_to_token.shape
+ if req_pool_idx <= 0 or req_pool_idx >= int(num_owners):
+ raise ValueError(
+ "invalid DFlash transfer owner: "
+ f"owner={req_pool_idx}, valid=[1,{int(num_owners) - 1}]"
+ )
+ if end > int(table_width):
+ raise ValueError(
+ "DFlash transfer range exceeds req_to_token width: "
+ f"end={end}, width={int(table_width)}"
+ )
+
+ locations = req_to_token[int(req_pool_idx), start:end].to(torch.int64)
+ if int(locations.numel()) != end - start:
+ raise RuntimeError(
+ "DFlash legacy suffix length mismatch: "
+ f"expected={end - start}, actual={int(locations.numel())}"
+ )
+ if bool(torch.any(locations <= 0).item()):
+ raise RuntimeError(
+ "DFlash legacy suffix contains an unallocated/padding KV slot"
+ )
+ return locations
+
def init_attention_backends(self):
+ if self._is_context_only_pp_prefill_rank:
+ return
self._draft_worker.init_attention_backends()
self._need_mamba_verify_commit = mambaish_config(
self.model_runner.model_config
@@ -1657,15 +1731,150 @@
) -> DFlashDraftInputV2:
return make_draft_input_v2(bonus_tokens=bonus_tokens, new_seq_lens=new_seq_lens)
+ def _init_pp_context_features(self):
+ from sglang.srt.speculative.dflash_pp import kimi_pp_capture_layer_ids
+
+ target_model = self.model_runner.model
+ if (
+ not self._is_pd_prefill
+ or not hasattr(target_model, "set_dflash_layers_to_capture")
+ or not hasattr(self.draft_model, "project_target_hidden_partial")
+ ):
+ raise ValueError(
+ "PP DFLASH currently requires a PD prefill target with DFLASH "
+ "aux capture and a draft model with partial context projection"
+ )
+ layer_ids = self.model_runner.spec_aux_config.dflash_target_layer_ids
+ if not layer_ids or len(layer_ids) != self.draft_model.num_context_features:
+ raise ValueError("DFLASH capture count does not match the draft projection")
+ info = self.model_runner.layer_info
+ num_layers = self.model_runner.model_config.num_hidden_layers
+ local_ids = kimi_pp_capture_layer_ids(
+ layer_ids, info.start_layer, info.end_layer, num_layers
+ )
+ self._pp_context_feature_indices = [layer_ids.index(i) for i in local_ids]
+ self._pp_expects_incoming_context = any(
+ min(i + 1, num_layers - 1) < info.start_layer for i in layer_ids
+ )
+ logger.info(
+ "DFLASH PP rank %s: capture layers=%s, projection columns=%s, incoming=%s",
+ self.ps.pp_rank,
+ local_ids,
+ self._pp_context_feature_indices,
+ self._pp_expects_incoming_context,
+ )
+
+ def set_pp_proxy_tensors_for_next_forward(self, pp_proxy_tensors):
+ self._next_pp_proxy_tensors = pp_proxy_tensors
+
+ @torch.no_grad()
+ def _forward_pp_prefill(self, batch, on_publish, pp_proxy_tensors):
+ result = self.target_worker.forward_batch_generation(
+ batch,
+ pp_proxy_tensors=pp_proxy_tensors,
+ capture_hidden_mode=CaptureHiddenMode.FULL,
+ )
+ output = result.pp_hidden_states_proxy_tensors
+ logits = result.logits_output
+ target_hidden = (
+ logits.hidden_states
+ if logits is not None
+ else (
+ output.tensors.get("dspark_aux_hidden_states")
+ if output is not None
+ else None
+ )
+ )
+ incoming = (
+ pp_proxy_tensors.tensors.get("dflash_ctx_acc")
+ if pp_proxy_tensors is not None
+ else None
+ )
+ if (incoming is not None) != self._pp_expects_incoming_context:
+ raise RuntimeError("DFLASH PP context missing or unexpectedly duplicated")
+ if batch.extend_lens is None or batch.prefix_lens is None:
+ raise RuntimeError("DFLASH PP prefill requires extend_lens and prefix_lens")
+ if batch.out_cache_loc is None:
+ raise RuntimeError("DFLASH PP prefill requires out_cache_loc")
+ tokens = sum(batch.extend_lens)
+ shape = (tokens, self.draft_model.config.hidden_size)
+ local = None
+ if self._pp_context_feature_indices:
+ if target_hidden is None or target_hidden.shape[0] != tokens:
+ raise RuntimeError(
+ "DFLASH PP local hidden capture is missing or truncated"
+ )
+ local = self.draft_model.project_target_hidden_partial(
+ target_hidden, self._pp_context_feature_indices
+ )
+ elif logits is None and target_hidden is not None and target_hidden.numel():
+ raise RuntimeError("DFLASH PP captured unassigned layer features")
+ for context in (incoming, local):
+ if context is not None and tuple(context.shape) != shape:
+ raise RuntimeError("DFLASH PP accumulated context shape mismatch")
+ context = incoming
+ if local is not None:
+ context = local if incoming is None else incoming.to(local) + local
+
+ if self.ps.pp_rank < self.ps.pp_size - 1:
+ if output is None:
+ raise RuntimeError(
+ "DFLASH non-final PP stage did not return proxy tensors"
+ )
+ output.tensors.pop("dspark_aux_hidden_states", None)
+ if context is not None:
+ output.tensors["dflash_ctx_acc"] = context
+ else:
+ if output is not None or context is None or result.next_token_ids is None:
+ raise RuntimeError(
+ "DFLASH final PP stage lacks complete context or logits"
+ )
+ prefixes = torch.tensor(
+ batch.prefix_lens, dtype=torch.int32, device=self.device
+ )
+ extends = torch.tensor(
+ batch.extend_lens, dtype=torch.int32, device=self.device
+ )
+ positions, _ = compute_position(
+ self.model_runner.prefill_attention_backend_str,
+ prefixes,
+ extends,
+ tokens,
+ )
+ # The linear partials are summed before applying RMSNorm exactly once.
+ self._append_target_hidden_sequential(
+ ctx_hidden=self.draft_model.hidden_norm(context),
+ ctx_positions=positions.to(dtype=torch.int64),
+ ctx_cache_loc=batch.out_cache_loc.to(dtype=torch.int64),
+ )
+ result.next_draft_input = self._make_next_draft_input_prefill(
+ bonus_tokens=result.next_token_ids, seq_lens=batch.seq_lens
+ )
+ if logits is not None:
+ logits.hidden_states = None
+ result.new_seq_lens = batch.seq_lens
+ if on_publish is not None:
+ on_publish(result.new_seq_lens)
+ return result
+
def forward_batch_generation(
self,
batch: ScheduleBatch,
on_publish=None,
grammar_barrier=None,
+ pp_proxy_tensors=None,
) -> GenerationBatchResult:
+ # PP mode: the scheduler passes the incoming proxy tensors either as
+ # an explicit argument or via set_pp_proxy_tensors_for_next_forward.
+ if pp_proxy_tensors is None:
+ pp_proxy_tensors = self._next_pp_proxy_tensors
+ self._next_pp_proxy_tensors = None
self._validate_phase1_sampling_support(batch)
+ is_extend = batch.forward_mode.is_extend() or batch.is_extend_in_batch
- if batch.forward_mode.is_extend() or batch.is_extend_in_batch:
+ if is_extend:
+ if self.ps.pp_size > 1:
+ return self._forward_pp_prefill(batch, on_publish, pp_proxy_tensors)
# Target prefill: capture DFlash aux hidden states for prompt tokens.
batch_output = self.target_worker.forward_batch_generation(
batch, capture_hidden_mode=CaptureHiddenMode.FULL
diff -ruN '--exclude=__pycache__' '--exclude=._*' '--exclude=.git' /sgl-workspace/sglang/python/sglang/srt/speculative/spec_info.py /host_patch/python/sglang/srt/speculative/spec_info.py
--- /sgl-workspace/sglang/python/sglang/srt/speculative/spec_info.py 2026-08-28 03:46:53.000000000 +0000
+++ /host_patch/python/sglang/srt/speculative/spec_info.py 2026-09-07 12:26:53.324493421 +0000
@@ -190,6 +190,14 @@
return build_dspark_disagg_draft_input(
batch, last_tokens_tensor, future_map
)
+ if self.is_dflash():
+ from sglang.srt.speculative.dflash_disaggregation import (
+ build_dflash_family_disagg_draft_input,
+ )
+
+ return build_dflash_family_disagg_draft_input(
+ batch, last_tokens_tensor, future_map
+ )
return None
def need_topk(self) -> bool: