hit90 scenario bench: TP4PP2-nomtp@0.90 winner, DP-attention/DCP verdicts (60.8 serving, 60.5 v3 delivered, 09-09)

- 冠军 A6 TP4PP2-nomtp@0.90(池 647,040):hit90 out cc1-4=28.8/47.5/63.9/76.2、cc8/16=106.4/125.7(vs TP2PP4 基线 cc4+14%/cc8+15%),并发独立 128k 文档 4 条、512k 单条 ✓、质量门 7/7;60.8 在役(实测实例切 unless-stopped)
- 判决:DP attention 容量负收益判死(非 MoE 权重按 attn 组复制致 dp4 池 96,448/rank + EAGLE×DP 两层崩溃);DCP 对 DSA 静默算错禁用(dsa_backend 零引用无 guard);MTP@128k accept 2.07 判负(vs EAGLE 3.46 存活但池 276k 过不了容量门槛)
- 交付:60.5 v3 脚本(只交未执行)+ 补丁束(60.8:/root,md5 6922e534)+ profile tp4pp2_hicache.env + CURRENT.md 更新
- 归档:70 文件全量日志(含 DP 臂失败记录、A3 容器日志)+ 7 脚本 + README(双口径协议:hit90 主扫 + distinct-doc 容量探针)
This commit is contained in:
yy-fighting 2026-09-09 17:11:04 +08:00
parent 0d929948a5
commit ad1853f49b
12 changed files with 488 additions and 4 deletions

View File

@ -1,9 +1,9 @@
# 现役部署状态页live 核验于 2026-09-08 # 现役部署状态页live 核验于 2026-09-09
> 本页记录 174.1.60.x 集群"现在跑的是什么"。改动机器前后先读这页; > 本页记录 174.1.60.x 集群"现在跑的是什么"。改动机器前后先读这页;
> **任何变动(停/起/改配置)完成后必须更新本页**。核验命令:`docker ps` + `nvidia-smi` > **任何变动(停/起/改配置)完成后必须更新本页**。核验命令:`docker ps` + `nvidia-smi`
## 机器状态2026-09-08 实测) ## 机器状态2026-09-09 实测)
| 机器 | 在役 | 口径 / 归属 | 对应 profile | | 机器 | 在役 | 口径 / 归属 | 对应 profile |
|---|---|---|---| |---|---|---|---|
@ -11,10 +11,10 @@
| 60.2 (6000D-2) | `glm53-nvfp4` 实验容器09-08 晚 TP1PP8 phaserun_phase.sh 实验链进行中) | **他人实验进行中,勿动**(方案 F PD 链已拆除GPU7 曾有外部裸金属任务)。动卡前仍先核实归属 | `profiles/pro6000/glm53_nvfp4_pro6000_pd_{prefill→decode 侧}.env` | | 60.2 (6000D-2) | `glm53-nvfp4` 实验容器09-08 晚 TP1PP8 phaserun_phase.sh 实验链进行中) | **他人实验进行中,勿动**(方案 F PD 链已拆除GPU7 曾有外部裸金属任务)。动卡前仍先核实归属 | `profiles/pro6000/glm53_nvfp4_pro6000_pd_{prefill→decode 侧}.env` |
| 60.3 | 无容器,但 8 卡被外部裸金属训练占用(`/data/mas/larm`09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — | | 60.3 | 无容器,但 8 卡被外部裸金属训练占用(`/data/mas/larm`09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — |
| 60.4 | `glm53-nvfp4`Up:30000restart=unless-stopped | **TP8+EAGLE+custom-AR 1stageE7b 配方)在役**09-09 下午部署EAGLE 4/1/5/mem0.90/MRR16/chunk8192/ctxlen270336/fp8KV+hicache3/decode 图桶 1-8/双 parserCAR 补丁三处全注入、8 rank `SSKJ_CAR_PATCH_ACTIVE` 确认health 200、生成冒烟、质量门 7/7 见 `/root/qg_604_car.log`)。启动 `/root/deploy_glm53_604_exp.sh`=60.7 实验版逐字拷贝,`RESTART=yes CAR_PATCH=1`),补丁 `/root/patches/`md5 与仓库 car_patch 归档一致)。当日早间曾短暂部署 TP2PP4 D 配方复刻deploy_s2_test_604.sh 留盘可切回)后被本方案替换;同日经授权清退外部 vllm 评测流水线tmux `mas` 的 run_multiseed.sh 链,--resume 可续跑) | `experiments/pro6000/glm53_nvfp4_pro6000d_sglang_dual_scenario_bench/scripts/`deploy_glm53_604_exp.sh + car_patch/ 补丁快照) | | 60.4 | `glm53-nvfp4`Up:30000restart=unless-stopped | **TP8+EAGLE+custom-AR 1stageE7b 配方)在役**09-09 下午部署EAGLE 4/1/5/mem0.90/MRR16/chunk8192/ctxlen270336/fp8KV+hicache3/decode 图桶 1-8/双 parserCAR 补丁三处全注入、8 rank `SSKJ_CAR_PATCH_ACTIVE` 确认health 200、生成冒烟、质量门 7/7 见 `/root/qg_604_car.log`)。启动 `/root/deploy_glm53_604_exp.sh`=60.7 实验版逐字拷贝,`RESTART=yes CAR_PATCH=1`),补丁 `/root/patches/`md5 与仓库 car_patch 归档一致)。当日早间曾短暂部署 TP2PP4 D 配方复刻deploy_s2_test_604.sh 留盘可切回)后被本方案替换;同日经授权清退外部 vllm 评测流水线tmux `mas` 的 run_multiseed.sh 链,--resume 可续跑) | `experiments/pro6000/glm53_nvfp4_pro6000d_sglang_dual_scenario_bench/scripts/`deploy_glm53_604_exp.sh + car_patch/ 补丁快照) |
| 60.5 | `glm53-nvfp4`Up 29h | **NVFP4 团队生产**deploy_glm53_605.shmd5 fcd9109b。生产机铁律不实验、不重启、不覆盖脚本。**09-09 已交付容量扩容 v2 脚本TP2PP4-hicache 优胜配置,池 909,632/c6/单条 909k与 60.8 现役同款),未执行,择维护窗口跑 `/root/deploy_glm53_605_v2.sh`;回滚=原脚本** | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(方案 A 口径);待切 `glm53_nvfp4_pro6000_sglang_tp2pp4_hicache.env` | | 60.5 | `glm53-nvfp4`Up 2d09-09 只读核验 | **NVFP4 团队生产**deploy_glm53_605.shmd5 fcd9109b。生产机铁律不实验、不重启、不覆盖脚本。**交付升级路径已更新为 v309-09 hit90 场景优胜 TP4PP2-hicache@0.90,池 647,040/c4 并发独立文档/hit90 cc8 out +34% vs 现役):脚本 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/scripts/deploy_glm53_605_v3.sh` + 补丁束60.8:/root/glm53_r37_patch_bundle_v3.tar.gzmd5 6922e534需 scp 至 60.5+ 需 docker pull cu13 镜像;未执行、未落 60.5 磁盘60.5:/root 仅有原脚本,核验过)。此前 v2TP2PP4-hicache冷缓存口径优胜被 v3 取代,仍留仓库可作"容量优先 6 条文档/512k 单条最快"备选;回滚=原脚本** | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(方案 A 口径);待切 `glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`(容量优先备选 `..._tp2pp4_hicache.env` |
| 60.6 | 无容器,但 8 卡被外部裸金属实验占用(`/data/hzy/sparse-opd-*`09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — | | 60.6 | 无容器,但 8 卡被外部裸金属实验占用(`/data/hzy/sparse-opd-*`09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — |
| 60.7 | 基本空4 卡仍有 `/home/user/dirA_exp` 外部小任务09-08 晚实测) | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — | | 60.7 | 基本空4 卡仍有 `/home/user/dirA_exp` 外部小任务09-08 晚实测) | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — |
| 60.8 | `glm53-nvfp4`Up:30000restart=unless-stopped | **TP2PP4-nomtp 容量扩容口径在役**09-09 按用户决定取代 r37TP2PP4 + radix/hicache 开 + ctx 1048576 + cu13 栈 9 挂载 + de-GLOOmemfrac 0.85/chunk 8192/MRR 16KV 池 **909,632**,启动 `bash /root/deploy_ppmtp_r37.sh '--tp 2 --pp-size 4 --disable-overlap-schedule --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length 1048576' nomtp 8192 0.85`。四臂对拍优胜vs A-mirror/r37/B'i128k/o512 c1-4c4 输入/输出 **5,907/23.1 tok/s**、TTFT p50 43.4s、并发上限 c6、单条 ~909k900k 实跑、512k 单条 TTFT 88.4s、90% 命中 c4 输入 17,625 tok/s质量门 7/7。r37 的 MTP 优势区间是 i16k/cc≤16切回 = deploy_ppmtp_r37.sh mtp 模式(详见 `experiments/pro6000/glm53_nvfp4_128k_capacity_topology/` | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp2pp4_hicache.env`;实验全量 `experiments/pro6000/glm53_nvfp4_128k_capacity_topology/`;前史 r37 `experiments/pro6000/glm53_ppmtp_r37_verify_graph/` | | 60.8 | `glm53-nvfp4`Up:30000restart=unless-stopped | **TP4PP2-nomtp@0.90 hit90 优胜口径在役**09-09 晚接替同日早间的 TP2PP4@0.85:同 r37 栈 cu13+9 挂载+de-GLOO仅改 tp4/pp2 + memfrac 0.90KV 池 **647,040**;实测实例原样保留切 restart 策略,未重建容器。启动 `bash /root/deploy_ppmtp_r37.sh '--tp 4 --pp-size 2 --disable-overlap-schedule --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length 1048576' nomtp 8192 0.90 1 1 0 0`。hit9090% 命中 i128k/o512out cc1/2/3/4/8/16 = **28.8/47.5/63.9/76.2/106.4/125.7**vs TP2PP4 基线 cc4 +14%/cc8 +15%);并发独立 128k 文档 **4 条**cap cc4 98.5 零排队、512k 单条 ✓ TTFT 135.8s、质量门 **7/7**同轮判决DP attention 容量负收益判死、DCP 对 DSA 静默算错禁用、MTP@128k accept 2.07 判负。让步项(知情):容量 4 条TP2PP4 为 6、512k 单条慢 33%。回滚 TP2PP4 = 同命令改 `'--tp 2 --pp-size 4'` + memfrac 0.85 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`;实验全量 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/`;前史 `experiments/pro6000/glm53_nvfp4_128k_capacity_topology/`TP2PP4 口径)+ r37 `experiments/pro6000/glm53_ppmtp_r37_verify_graph/` |
## 方案 A-F 一览GLM-5.3-NVFP4 @ pro60002026-09-08 双场景报告口径) ## 方案 A-F 一览GLM-5.3-NVFP4 @ pro60002026-09-08 双场景报告口径)

View File

@ -0,0 +1,47 @@
# GLM-5.3-NVFP4 SGLang TP=4 PP=2 + hicache deployment profile (single RTX 6000D node, 8 GPUs).
# 2026-09-09 hit90 场景实验优胜配置60.8 现役60.5 交付 deploy_glm53_605_v3.sh 待执行)。
# 可执行部署脚本experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/scripts/deploy_glm53_605_v3.sh
# 60.8 启动口径bash /root/deploy_ppmtp_r37.sh '--tp 4 --pp-size 2 --disable-overlap-schedule
# --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length 1048576' nomtp 8192 0.90 1 1 0 0
#
# 与 TP2PP4-hicacheglm53_nvfp4_pro6000_sglang_tp2pp4_hicache.env的关键差异勿混淆
# - tp2/pp4 → tp4/pp2memfrac 0.85 → 0.90其余cu13+9 挂载、chunk 8192、MRR 16、
# hicache 3、ctx 1048576、双 parser逐项同构
# - 实测 KV 池 647,040fp8 KV 18.52GB/rankPP0 初始化后剩 10.94GB
# - 取胜依据hit90=90% 命中 i128k/o512 主场景out cc1/2/3/4 = 28.8/47.5/63.9/76.2、
# cc8/16 = 106.4/125.7vs TP2PP4cc4 +14%/cc8 +15%/cc16 +1%cap cc2/cc4 = 62.9/98.5
# 零排队;质量门 7/7
# - 让步项(知情选择):并发独立 128k 文档 4 条TP2PP4 为 6512k 单条 151.9s
# TP2PP4 为 114.4sPP4 单条巨请求 prefill 流水更优);无投机解码
# - 判决背景DP attention 对本模型容量负收益、DCP 对 DSA 静默算错、MTP@128k accept 2.07 判负
# (见 experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/README.md
PLATFORM=pro6000
EXPERIMENT=glm53_nvfp4_hit90_dp_dcp_bench
MODEL_NAME=GLM-5.3-NVFP4
ENGINE=sglang
RUNTIME=docker
DOCKER_IMAGE=lmsysorg/sglang:nightly-dev-cu13-20260901-07c8f729
CONTAINER_NAME=glm53-nvfp4
MODEL_PATH=/data/hf_models/GLM-5.3-NVFP4
SERVED_MODEL_NAME=/data/hf_models/GLM-5.3-NVFP4
PORT=30000
HEALTH_PATH=/health
HEALTH_WAIT_S=1800
CONTAINER_PYTHON=python3
TP=4
PP=2
MEM_FRACTION_STATIC=0.90
MAX_RUNNING_REQUESTS=16
CHUNKED_PREFILL_SIZE=8192
CONTEXT_LENGTH=1048576
DEVICE_VARS="CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7"
ENGINE_ENV="PYTHONUNBUFFERED=1 SGLANG_PP_DEGLOO=1 SGLANG_PP_SPEC_FORCE_EAGER_DRAFT=1 SGLANG_PP_FORCE_EAGER_VERIFY=0 SGLANG_PP_SPEC_DEBUG=0"
DOCKER_FLAGS="--gpus all --shm-size 64g --ipc=host --cap-add SYS_PTRACE -p ${PORT}:${PORT}"
VOLUMES="/data/hf_models:/data/hf_models + 9 patch ro-mounts (full list in scripts/deploy_glm53_605_v3.sh; files live in /root + /root/sglang_patch2 on the host, bundle md5 6922e53439991bc13feee72f3760704f)"
BOOTSTRAP="python3 -m sglang.launch_server ${LAUNCH_ARGS}"
LAUNCH_ARGS="--model-path ${MODEL_PATH} --tp-size ${TP} --pp-size ${PP} --mem-fraction-static ${MEM_FRACTION_STATIC} --max-running-requests ${MAX_RUNNING_REQUESTS} --chunked-prefill-size ${CHUNKED_PREFILL_SIZE} --disable-shared-experts-fusion --moe-runner-backend flashinfer_cutlass --disable-flashinfer-autotune --reasoning-parser glm45 --tool-call-parser glm47 --enable-hierarchical-cache --hicache-ratio 3 --disable-overlap-schedule --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length ${CONTEXT_LENGTH} --host 0.0.0.0 --port ${PORT}"

View File

@ -0,0 +1,56 @@
# GLM-5.3-NVFP4 hit90 场景输出吞吐与容量优化实验(含 DP attention / DCP 探索)— 2026-09-09 @ 174.1.60.8
## 一、结论速览
**优胜配置 = A6TP4PP2 nomtp @ memfrac 0.90**cu13 镜像 + r37 9 补丁挂载radix/hicache 开,`--context-length 1048576`。60.8 已在役(实测实例原样保留切 `restart=unless-stopped`)。
场景(用户修正后):**缓存命中 90%** 的 i128k/o512 长上下文重复查询60.5 真实流量形态),指标 = 输出吞吐 + 容量上限,且不限于 PPDP attention / DCP 均需验证)。
| 维度 | A1 在役TP2PP4@0.85 | A3 = 60.5 生产TP8+EAGLE | **A6 优胜TP4PP2@0.90** |
|---|---|---|---|
| KV 池token | 909,632 | 276,864 | **647,040** |
| 并发独立 128k 文档上限 | 6 | 2 | **4** |
| hit90 out cc1/2/3/4 | —(仅 cc4 66.9 | —(仅 cc4 72.0 | **28.8 / 47.5 / 63.9 / 76.2** |
| hit90 out cc8/16 | 92.6 / 124.1 | 79.2 / 82.0 | **106.4 / 125.7** |
| cap cc2 单请求解码 | 18.4 tok/s | **44.7** | 31.5 |
| 512k 单条 | 114.4s ✓ | ✗(池不够) | 151.9s ✓TTFT 135.8s |
| 质量门 | 7/7 | — | **7/7** |
- A6 对 A1hit90 cc4/cc8 **+14%/+15%**cc16 +1%(主指标全胜);容量 6→4 为唯一让步hit90 共享前缀口径下 cc16 仅占 ~33 万 token超出部分 hicache host 层兜底cap cc8 排队场景 96.5 vs A1 的 67.6 仍优雅)。
- A6 对 60.5 生产A16并发独立文档容量 **×2**2→4、hit90 cc8 **+34%**;代价 = EAGLE 低并发单请求极致速度消失44.7→31.5 tok/s/req
- 512k 单条 A6 比 A1 慢 33%PP4 对单条巨请求的 4 级 prefill 流水更优)——判据只要求"可完成",如实记录。
## 二、三条架构级判决(本轮核心增量)
1. **DP attention 对 GLM-5.3 容量是负收益**:非 MoE 权重DSA indexer/shared expert/dense按 attn 组整份复制attn_tp=4 → 58.37GB/rankdp4 池塌到 96,448/rank、dp2 215,232/组,聚合均 < 单机 TP2PP4 909,632叠加 EAGLE×DP 两层独立崩溃mlp-sync padding NoneType 补丁修复SM120 sparse-MLA seq_lens 形状 (8,)vs(6,) cu13 栈未修)→ **DP 臂整体判死**证据`results/20260909/dp_arms_failure_record.md` + arm4/arm5 日志
2. **DCP 对 DSA 是静默死路**`--dcp-size` 只有稠密 MLA 内核消费dsa_backend.py 零引用且无 guard → 配了**静默算错**。点亮需 SM120 sparse-MLA 内核补丁(多天级),列为后续工程项,禁止生产尝试。
3. **MTP 在 128k 长上下文判负**accept 2.0716k 时 3.5+),单请求 14.9 < 无投机 20.7 tok/sDSA verify 成本随上下文暴涨break-even accept2.87EAGLE 深树 accept 3.46@128k 存活但 A16 配方池太小未来投机方向 = "TP4PP2+EAGLE 扩池"未验证
## 三、方法学(双口径协议)
- **hit90 主扫**shared-frac 0.9(共享 117,968 前缀cc∈{1,2,3,4,8,16}nreq=2×cc每点 flush+warm90实测 hit_rate=0.8999。注意共享前缀口径容量被天然美化,不能单独下容量结论。
- **容量探针cap**:独立文档,种子遍(串行 o8 喂 radix→ 计分遍(并发 o512 不 flush**TTFT_max 突增 = 排队 = 容量上限信号**(种子过的文档重复查询 TTFT≈0.5s、admission 给足信用)。
- **判据**:硬门槛 = cap cc4 零排队 + 512k 单条可完成 + 质量门 7/7入围按 hit90 out_tps 定夺。
- 窗口映射全臂同映射PG19 语料 --pool-override 重放jit 18.0M / hit90 cc1-3 17.0/17.3/17.6M / cc4 15.5M / cc8 15.7M / cc16 16.1M / cap cc2-8 18.5/18.8/19.4/20.2M / 512k 10.0M(语料总 21,296,780
- 复现性A1 hit90 cc4=66.9 vs 上轮独立验证点 68.85(偏差 2.9% 预热抖动范围)。
## 四、资产
- `scripts/`run_arm90.sh单臂全套编排、warm90.py共享前缀预热、finals_arm6.sh质量门+512k+cc1-3 决赛、deploy_dp_eagle.sh + patch_fbi.py + launch_arm5.shDP 臂部署与 fbi None 守卫补丁)、**deploy_glm53_605_v3.sh60.5 交付脚本,只交未执行)**。
- `results/20260909/`70 文件全量日志A1/A2/A3/A6 主扫+cap+决赛、DP 臂失败记录、部署日志、A3 容器全量 docker 日志)。
- 复用既有归档deploy_ppmtp_r37.sh`experiments/pro6000/glm53_ppmtp_r37_verify_graph/scripts/`、quality_gate_605.sh`.../glm53_nvfp4_pro6000d_sglang_dual_scenario_bench/scripts/`)。
- 60.5 部署补丁束60.8:/root/glm53_r37_patch_bundle_v3.tar.gz9 文件md5 `6922e53439991bc13feee72f3760704f`9 补丁文本快照另见 r37 实验目录)。
- 完整报告:`D:\sskj\reports\GLM53_NVFP4_60.8_hit90场景输出吞吐与容量优化实验报告_2026-09-09.md`
- Profile`deploy/profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`
## 五、60.8 在役部署口径2026-09-09 起)
```
bash /root/deploy_ppmtp_r37.sh '--tp 4 --pp-size 2 --disable-overlap-schedule \
--max-prefill-tokens 16384 --disable-custom-all-reduce --context-length 1048576' \
nomtp 8192 0.90 1 1 0 0
# 池 647,040容器 glm53-nvfp4 restart=unless-stopped
# 回滚 A1 = 同命令改 '--tp 2 --pp-size 4' + memfrac 0.85
```
60.5 切换scp 补丁束 + `deploy_glm53_605_v3.sh` → 维护窗口执行(停机 ~7 分钟,需 docker pull cu13 镜像);回滚 = 原 deploy_glm53_605.shA16 配方)。

View File

@ -0,0 +1,43 @@
# DP attention 臂失败记录2026-09-0960.8cu13 nightly-dev-cu13-20260901-07c8f729
## A4: TP8 + DP4 (attn_tp=2) + EAGLE 4/1/5 @0.90 —— 容量判死 + 崩溃 #1
启动日志实测(容器 0161ef23已删
- `Load weight end ... mem usage=64.75 GB`TP8 时约 38.5GB → 非 MoE 权重按 attn 组复制attn_tp=2 时每 rank 多吃 ~26GB
- `KV Cache is allocated ... #tokens: 96448, KV size: 5.52 GB`(每 rank
- `Capture target prefill CUDA graph ... num_tokens=[4..2048]` 吃 5.37GB
- 聚合池 = 4 × 96,448 = 385,792 << A1(TP2PP4-nomtp) 909,632单组 96k < 131k一条 128k 都装不下
- 崩溃 #106:06:58全 DP rank
`forward_batch_info.py prepare_mlp_sync_batch → _pad_inputs_to_size → _pad_tensor_to_size
AttributeError: 'NoneType' object has no attribute 'new_zeros'`
机理:`eagle_prepare_for_verify` 仅对非 idle batch 赋 `batch.input_ids`eagle_utils.py `if not
batch.forward_mode.is_idle()`DP attention 下空闲 rank 跑 idle batch 对齐 MoE 集合通信,
mlp-sync padding 无条件 pad input_ids → None 崩溃。
修复fbi_cu13_patched.py/rootmd5 0afeb2bb2f07e9df28198c91601ca490加 None 守卫,验证有效(不再崩此处)。
## A5: TP8 + DP2 (attn_tp=4) + EAGLE 4/1/5 @0.90 + fbi 补丁 + 图削减 —— 容量仍判死 + 崩溃 #2
启动日志实测(容器 e817d310已删
- `Load weight end ... mem usage=58.37 GB`attn_tp=4 复制仍吃 ~20GB
- `KV Cache is allocated ... #tokens: 215232, KV size: 12.32 GB`(每 rank/组)
- 图削减生效prefill graph 0.99GB(原 5.37、verify 0.84、draft 0.25+0.37`max_running_requests=24`48/dp2
- 聚合池 = 2 × 215,232 = 430,464 < A1 909,632单组 215k 装不下两条独立 128k262k独立文档并发上限 2
- 崩溃 #206:25:00 DP0 TP0fbi 补丁后新崩点):
`prefill_cuda_graph_runner.execute → deepseek_v2.forward → mla_bmm_then_unified_attention →
flashinfer trtllm_batch_decode_sparse_mla_v32_sm120 → _normalize_sm120_sparse_v32_topk_length
ValueError: seq_lens for SM120 sparse MLA v32/GLM must be shaped either (6,), (6,), or (6, 1); got (8,)`
机理EAGLE verifydraft 5+1=6 位置)经 prefill(extend) 图 runner 走 SM120 sparse-MLA decode 内核,
期望 seq_lens 形状 (6,),实际收到 (8,)——cu13 栈 EAGLE×GLM(DSA) 的内核级形状缺陷(旧镜像 20260828
的生产 EAGLE 配方从未在 cu13 验证过cu13 上唯一验证过的投机路径是 MTP+r37 九挂载补丁栈)。
## 判决
DP attention 在 GLM-5.3-NVFP4 + 本栈上对"扩池"目标为负:
1. 容量负收益(架构性):非 MoE 权重(含 DSA indexer/shared expert/dense按 attn 组复制,
dp 越大每 rank 权重越重 → dp4 池 38.6万、dp2 池 43万均低于 TP2PP4 的 90.9 万。
"聚合容量 ≈ ×dp_size" 的前提是非 MoE 权重占比小GLM-5.3 不满足。
2. 投机解码两层独立崩溃EAGLE 路径idle-batch mlp-sync padding已补丁+ SM120 sparse-MLA
verify 形状断言(需内核侧工作,超出本轮范围)。
3. MTP×DP 未测——r37 九挂载补丁栈只验证过 PP 路径,与 dp 组合属未验证组合,且容量已被 1 否决。
资产deploy_dp_eagle.shSPEC/FBIPATCH 可调、patch_fbi.py、fbi_cu13_patched.py60.8 /root

View File

@ -0,0 +1,69 @@
#!/bin/bash
# GLM-5.3-NVFP4 TP8 + DP attention (+ EAGLE) on 60.8 (cu13, optional fbi patch mount).
# DP attention: per-dp-group scheduler + KV pool. NOTE (measured 09-09): non-MoE
# weights replicate per attn group (attn_tp = tp/dp), so per-rank weights grow as
# dp grows -> dp4 pool collapsed to 96k/rank. dp2 (attn_tp4) is the sweet spot.
# EAGLE+dp needs the forward_batch_info.py None-guard patch (idle-batch mlp-sync
# padding); pass FBIPATCH=/root/fbi_cu13_patched.py when SPEC=1.
# Usage: DP=2 bash deploy_dp_eagle.sh
# Env: DP (4) SPEC (1) MEMFRAC (0.90) MRR (48) CHUNK (16384) FBIPATCH ("") EXTRA ("")
set -e
DP=${DP:-4}
SPEC=${SPEC:-1}
MEMFRAC=${MEMFRAC:-0.90}
MRR=${MRR:-48}
CHUNK=${CHUNK:-16384}
FBIPATCH=${FBIPATCH:-}
EXTRA=${EXTRA:-}
SPECARGS=""
if [ "$SPEC" = "1" ]; then
SPECARGS="--speculative-algorithm EAGLE --speculative-num-steps 4 --speculative-eagle-topk 1 --speculative-num-draft-tokens 5 --speculative-skip-dp-mlp-sync"
fi
PATCHMOUNT=""
if [ -n "$FBIPATCH" ]; then
PATCHMOUNT="-v $FBIPATCH:/sgl-workspace/sglang/python/sglang/srt/model_executor/forward_batch_info.py:ro"
fi
for i in $(seq 1 20); do
CID=$(docker ps -a --filter name=glm53-nvfp4 -q)
[ -z "$CID" ] && break
docker rm -f glm53-nvfp4 >/dev/null 2>&1 || true
sleep 3
done
[ -z "$(docker ps -a --filter name=glm53-nvfp4 -q)" ] || { echo "ERROR: old container cannot be removed"; exit 1; }
for i in $(seq 1 60); do
used=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | awk '{s+=$1} END {print s}')
[ "$used" -lt 1500 ] && break
sleep 3
done
echo "[deploy] VRAM now: ${used} MiB total; starting TP8 DP${DP} SPEC=${SPEC} memfrac=${MEMFRAC} mrr=${MRR} chunk=${CHUNK} patch='${FBIPATCH}' extra='${EXTRA}'"
docker run -d --name glm53-nvfp4 --gpus all --shm-size 64g --ipc=host --restart no \
-p 30000:30000 -v /data/hf_models:/data/hf_models $PATCHMOUNT \
lmsysorg/sglang:nightly-dev-cu13-20260901-07c8f729 \
python3 -m sglang.launch_server \
--model-path /data/hf_models/GLM-5.3-NVFP4 \
--tp 8 --dp-size $DP --enable-dp-attention \
--mem-fraction-static $MEMFRAC --max-running-requests $MRR \
--chunked-prefill-size $CHUNK --max-prefill-tokens 16384 \
--disable-shared-experts-fusion --moe-runner-backend flashinfer_cutlass --disable-flashinfer-autotune \
$SPECARGS \
--enable-hierarchical-cache --hicache-ratio 3 \
--context-length 1048576 --reasoning-parser glm45 --tool-call-parser glm47 \
--host 0.0.0.0 --port 30000 $EXTRA
echo "[deploy] container started; waiting for health..."
for i in $(seq 10 10 2400); do
code=$(curl -s -o /dev/null -m3 -w '%{http_code}' http://127.0.0.1:30000/health 2>/dev/null)
if [ "$code" = "200" ]; then
echo "healthy after ${i}s"
docker logs glm53-nvfp4 2>&1 | grep -oE "max_total_num_tokens = [0-9]+" | head -4
exit 0
fi
sleep 10
done
echo "ERROR: health timeout after 2400s"
docker logs glm53-nvfp4 2>&1 | tail -40
exit 1

View File

@ -0,0 +1,90 @@
#!/bin/bash
# deploy_glm53_605_v3.sh — GLM-5.3-NVFP4 hit90 实验冠军配方部署脚本60.5 团队自用机)
#
# 冠军配方2026-09-09 hit90 实验60.8 全套实测):
# TP4PP2-nomtp @ memfrac 0.90cu13 镜像 + r37 9 补丁挂载
# 池 647,040 tokensfp8 KV 18.52GB/rank→ 4 条并发 128k 独立文档 + 512k 单条
# hit90 输出吞吐 76.2/106.4/125.7 tok/s @cc4/8/16vs 旧 A16 配方 72.0/79.2/82.0
# 容量 vs 旧配方2 条 → 4 条并发独立文档×2hit90 cc8 +34%
# 质量门 7/7、512k 单条通过(详见实验报告)
#
# 前置条件(操作者手动完成):
# 1) 本脚本与补丁束 glm53_r37_patch_bundle_v3.tar.gz (md5 6922e53439991bc13feee72f3760704f)
# 放在同一目录(补丁束在 60.8:/root/glm53_r37_patch_bundle_v3.tar.gz先 scp 过来)
# 2) docker pull lmsysorg/sglang:nightly-dev-cu13-20260901-07c8f729
# 3) 停机窗口 ~7 分钟(会移除现有 glm53-nvfp4 容器)
#
# 用法: bash deploy_glm53_605_v3.sh
set -euo pipefail
BUNDLE=glm53_r37_patch_bundle_v3.tar.gz
IMG=lmsysorg/sglang:nightly-dev-cu13-20260901-07c8f729
MODEL=/data/hf_models/GLM-5.3-NVFP4
PORT=30000
NAME=glm53-nvfp4
HERE=$(cd "$(dirname "$0")" && pwd)
# ---- 0) 补丁束解包与校验(挂载布局须与 60.8 实测完全一致)----
if [ ! -f /root/eagle_worker_v2_mask.py ] || [ ! -f /root/sglang_patch2/deepseek_v2.py ]; then
[ -f "$HERE/$BUNDLE" ] || { echo "缺补丁束 $BUNDLE(从 60.8:/root/ scp"; exit 1; }
tar xzf "$HERE/$BUNDLE" -C /root/
fi
for f in sglang_patch2/layer_setup.py sglang_patch2/validation_hook.py \
eagle_worker_v2_mask.py sglang_patch2/eagle_worker_common.py \
sglang_patch2/deepseek_nextn.py scheduler_pp_mixin_r35.py \
request_receiver_degloo.py sglang_patch2/deepseek_v2.py \
decode_cuda_graph_runner_fix.py; do
[ -f "/root/$f" ] || { echo "补丁缺失: /root/$f(重新解包 $BUNDLE"; exit 1; }
done
# ---- 1) 镜像 ----
docker images --format '{{.Repository}}:{{.Tag}}' | grep -qx "$IMG" || docker pull "$IMG"
# ---- 2) 旧容器清理(重试 + 等显存归零,防 CUDA OOM----
for i in 1 2 3 4 5; do docker rm -f "$NAME" >/dev/null 2>&1 && break || sleep 5; done
for i in $(seq 1 60); do
m=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | sort -rn | head -1)
[ "$m" -lt 1500 ] && break
sleep 5
done
for i in $(seq 1 15); do ss -ltn 2>/dev/null | grep -q ":$PORT " || break; sleep 2; done
# ---- 3) 启动(= 60.8 实测冠军容器参数 1:1----
docker run -d --name "$NAME" --gpus all --shm-size 64g --ipc=host --cap-add SYS_PTRACE \
-v /root/sglang_patch2/layer_setup.py:/sgl-workspace/sglang/python/sglang/srt/model_executor/model_runner_components/layer_setup.py:ro \
-v /root/sglang_patch2/validation_hook.py:/sgl-workspace/sglang/python/sglang/srt/arg_groups/validation_hook.py:ro \
-v /root/eagle_worker_v2_mask.py:/sgl-workspace/sglang/python/sglang/srt/speculative/eagle_worker_v2.py:ro \
-v /root/sglang_patch2/eagle_worker_common.py:/sgl-workspace/sglang/python/sglang/srt/speculative/eagle_worker_common.py:ro \
-v /root/sglang_patch2/deepseek_nextn.py:/sgl-workspace/sglang/python/sglang/srt/models/deepseek_nextn.py:ro \
-v /root/scheduler_pp_mixin_r35.py:/sgl-workspace/sglang/python/sglang/srt/managers/scheduler_pp_mixin.py:ro \
-v /root/request_receiver_degloo.py:/sgl-workspace/sglang/python/sglang/srt/managers/scheduler_components/request_receiver.py:ro \
-v /root/sglang_patch2/deepseek_v2.py:/sgl-workspace/sglang/python/sglang/srt/models/deepseek_v2.py:ro \
-v /root/decode_cuda_graph_runner_fix.py:/sgl-workspace/sglang/python/sglang/srt/model_executor/runner/decode_cuda_graph_runner.py:ro \
-e SGLANG_PP_SPEC_DEBUG=0 -e SGLANG_PP_DEGLOO=1 \
-e SGLANG_PP_SPEC_FORCE_EAGER_DRAFT=1 -e SGLANG_PP_FORCE_EAGER_VERIFY=0 \
--restart unless-stopped -p $PORT:$PORT \
-v /data/hf_models:/data/hf_models \
"$IMG" \
python3 -m sglang.launch_server \
--model-path "$MODEL" \
--tp 4 --pp-size 2 \
--mem-fraction-static 0.90 \
--max-running-requests 16 \
--chunked-prefill-size 8192 \
--disable-shared-experts-fusion \
--moe-runner-backend flashinfer_cutlass \
--disable-flashinfer-autotune \
--reasoning-parser glm45 --tool-call-parser glm47 \
--enable-hierarchical-cache --hicache-ratio 3 \
--disable-overlap-schedule --max-prefill-tokens 16384 \
--disable-custom-all-reduce --context-length 1048576 \
--host 0.0.0.0 --port "$PORT"
echo "deployed v3 (TP4PP2-nomtp@0.90); waiting for health..."
for i in $(seq 1 240); do
code=$(curl -s -m 2 "http://127.0.0.1:$PORT/health" -o /dev/null -w '%{http_code}' || true)
[ "$code" = "200" ] && { echo "healthy after $((i*10))s"; break; }
sleep 10
done
docker logs "$NAME" 2>&1 | grep -E 'max_total_num_tokens|KV Cache is allocated' | tail -5
echo "预期: max_total_num_tokens=647040±0若显著低于此值检查是否漏挂补丁或镜像代差"

View File

@ -0,0 +1,52 @@
#!/bin/bash
# Finals for arm6 (TP4PP2-nomtp@0.90): quality gate 7/7 + 512k single + hit90 cc1/2/3.
# Usage: nohup bash /root/finals_arm6.sh > /root/bench_logs/hit90exp/arm6_finals_progress.log 2>&1 &
set -u
LOG=/root/bench_logs/hit90exp
URL=http://127.0.0.1:30000
CORPUS=/root/corpus_ids.json
idle_wait() {
for i in $(seq 1 90); do
local last
last=$(docker logs --since 90s glm53-nvfp4 2>&1 | grep 'running-req' | tail -1)
if [ -z "$last" ] || echo "$last" | grep -q 'running-req: 0'; then return 0; fi
sleep 10
done
echo "[warn] idle_wait timeout, continuing"
}
flush() { curl -s -X POST $URL/flush_cache >/dev/null; sleep 3; }
m() { grep -o "\"$1\": [0-9.]*" "$2" | tail -1; }
echo "=== finals arm6 start $(date) ==="
# 1) quality gate 7/7
idle_wait; flush
bash /root/quality_gate_605.sh > "$LOG/arm6_quality_gate.log" 2>&1
echo "[quality] pass=$(grep -c '\[PASS\]' $LOG/arm6_quality_gate.log) fail=$(grep -c '\[FAIL\]' $LOG/arm6_quality_gate.log)"
# 2) 512k single (mirrors last-round arm3_TP2PP4_single512k protocol)
idle_wait; flush
python3 /root/bench_corpus.py --corpus $CORPUS --input-len 523776 --output-len 512 --shared-frac 0.0 \
--concurrency 1 --num-requests 1 --run-id 9601 --pool-override 10000000 \
> "$LOG/arm6_single512k.log" 2>&1
echo "[512k single] wall=$(m wall_s $LOG/arm6_single512k.log) out=$(m output_throughput_tok_s $LOG/arm6_single512k.log) in=$(m input_throughput_tok_s $LOG/arm6_single512k.log) ttft_max=$(grep -A3 '\"ttft_s\"' $LOG/arm6_single512k.log | grep -o '\"max\": [0-9.]*' | tail -1)"
# 3) hit90 cc1/2/3 (user's original c1-4 scenario, low-cc points)
for cc in 1 2 3; do
case $cc in
1) base=17000000;;
2) base=17300000;;
3) base=17600000;;
esac
idle_wait; flush
python3 /root/warm90.py 1 > "$LOG/arm6_warm_cc${cc}.log" 2>&1
python3 /root/bench_corpus.py --corpus $CORPUS --input-len 131072 --output-len 512 --shared-frac 0.9 \
--concurrency "$cc" --num-requests $((cc * 2)) --run-id $((9000 + cc)) --pool-override "$base" \
> "$LOG/arm6_hit90_cc${cc}.log" 2>&1
echo "[hit90 cc=$cc] out=$(m output_throughput_tok_s $LOG/arm6_hit90_cc${cc}.log) in=$(m input_throughput_tok_s $LOG/arm6_hit90_cc${cc}.log) wall=$(m wall_s $LOG/arm6_hit90_cc${cc}.log)"
done
echo "=== finals arm6 ALL DONE $(date) ==="

View File

@ -0,0 +1,5 @@
#!/bin/bash
# A5: TP8 + DP2 (attn_tp4) + EAGLE 4/1/5, fbi None-guard patch, capped graphs.
export DP=2 SPEC=1 FBIPATCH=/root/fbi_cu13_patched.py
export EXTRA="--cuda-graph-max-bs-decode 16 --cuda-graph-bs-decode 1 2 4 6 8 12 16 --cuda-graph-max-bs-prefill 64"
bash /root/deploy_dp_eagle.sh

View File

@ -0,0 +1,20 @@
#!/usr/bin/env python3
"""Patch cu13 forward_batch_info.py: guard input_ids padding against None.
Under DP attention, idle-mode batches (ranks with no requests this step,
running to keep MoE collectives aligned) have input_ids=None. The DP mlp-sync
padding path (_pad_inputs_to_size) unconditionally padded input_ids and
crashed with AttributeError. eagle_prepare_for_verify only assigns
batch.input_ids for non-idle batches, so None here is legitimate.
"""
SRC = "/root/fbi_cu13.py"
DST = "/root/fbi_cu13_patched.py"
src = open(SRC).read()
old = " self.input_ids = self._pad_tensor_to_size(self.input_ids, num_tokens)\n"
new = (" if self.input_ids is not None:\n"
" self.input_ids = self._pad_tensor_to_size(self.input_ids, num_tokens)\n")
n = src.count(old)
assert n == 1, f"expected exactly 1 occurrence, found {n}"
open(DST, "w").write(src.replace(old, new))
print("patched OK ->", DST)

View File

@ -0,0 +1,71 @@
#!/bin/bash
# One full hit90 experiment pass for the CURRENTLY-SERVING arm:
# JIT/dp-priming warm point -> hit90 sweep cc4/8/16 (shared-frac 0.9)
# -> distinct-doc capacity probe cc2/4/6/8 (seed serial o8, then scored cc o512, no flush)
# Corpus windows are FIXED across all arms (pool-override replay of consumed
# region, flush per point = cold-unique fairness):
# jit-warm 18,000,000 | hit90: cc4 15,500,000 / cc8 15,700,000 / cc16 16,100,000
# cap: cc2 18,500,000 / cc4 18,800,000 / cc6 19,400,000 / cc8 20,200,000
# Usage: nohup bash run_arm90.sh <arm_name> <warmn> \
# > /root/bench_logs/hit90exp/<arm>_progress.log 2>&1 &
set -u
ARM=$1
WARMN=${2:-1}
LOG=/root/bench_logs/hit90exp
mkdir -p "$LOG"
CORPUS=/root/corpus_ids.json
URL=http://127.0.0.1:30000
idle_wait() {
for i in $(seq 1 90); do
local last
last=$(docker logs --since 90s glm53-nvfp4 2>&1 | grep 'running-req' | tail -1)
if [ -z "$last" ] || echo "$last" | grep -q 'running-req: 0'; then return 0; fi
sleep 10
done
echo "[warn] idle_wait timeout, continuing"
}
flush() { curl -s -X POST $URL/flush_cache >/dev/null; sleep 3; }
m() { grep -o "\"$1\": [0-9.]*" "$2" | tail -1; }
hit90_point() { # cc base
local cc=$1 base=$2
idle_wait; flush
python3 /root/warm90.py "$WARMN" > "$LOG/${ARM}_warm_cc${cc}.log" 2>&1
python3 /root/bench_corpus.py --corpus $CORPUS --input-len 131072 --output-len 512 --shared-frac 0.9 \
--concurrency "$cc" --num-requests $((cc * 2)) --run-id $((9000 + cc)) --pool-override "$base" \
> "$LOG/${ARM}_hit90_cc${cc}.log" 2>&1
echo "[hit90 cc=$cc] out=$(m output_throughput_tok_s "$LOG/${ARM}_hit90_cc${cc}.log") in=$(m input_throughput_tok_s "$LOG/${ARM}_hit90_cc${cc}.log") wall=$(m wall_s "$LOG/${ARM}_hit90_cc${cc}.log")"
}
cap_point() { # cc base
local cc=$1 base=$2
idle_wait; flush
python3 /root/bench_corpus.py --corpus $CORPUS --input-len 131072 --output-len 8 --shared-frac 0.0 \
--concurrency 1 --num-requests "$cc" --run-id 9993 --pool-override "$base" \
> "$LOG/${ARM}_cap${cc}_seed.log" 2>&1
python3 /root/bench_corpus.py --corpus $CORPUS --input-len 131072 --output-len 512 --shared-frac 0.0 \
--concurrency "$cc" --num-requests "$cc" --run-id $((9100 + cc)) --pool-override "$base" \
> "$LOG/${ARM}_cap${cc}.log" 2>&1
echo "[cap cc=$cc] out=$(m output_throughput_tok_s "$LOG/${ARM}_cap${cc}.log") ttft_max=$(grep -A3 '"ttft_s"' "$LOG/${ARM}_cap${cc}.log" | grep -o '"max": [0-9.]*' | tail -1) wall=$(m wall_s "$LOG/${ARM}_cap${cc}.log")"
}
echo "=== arm=$ARM start $(date) ==="
idle_wait; flush
python3 /root/warm90.py "$WARMN" > "$LOG/${ARM}_warm_jit.log" 2>&1
python3 /root/bench_corpus.py --corpus $CORPUS --input-len 131072 --output-len 512 --shared-frac 0.9 \
--concurrency 2 --num-requests 2 --run-id 9992 --pool-override 18000000 \
> "$LOG/${ARM}_jitwarm.log" 2>&1
echo "[jit-warm done]"
hit90_point 4 15500000
hit90_point 8 15700000
hit90_point 16 16100000
cap_point 2 18500000
cap_point 4 18800000
cap_point 6 19400000
cap_point 8 20200000
echo "=== arm=$ARM ALL DONE $(date) ==="

View File

@ -0,0 +1,31 @@
#!/usr/bin/env python3
"""Prime the hit90 shared prefix (ids[0:117968]) into radix cache.
Sent N times sequentially so every DP-attention group's radix gets a copy
(requests are allocated round-robin across dp ranks). Payload is identical
to bench_corpus.py's internal warmup (shared prefix + 64 spare-region
tokens), so it never pollutes any scored unique-suffix window.
Usage: python3 warm90.py [N]
"""
import json
import sys
import time
import requests
N = int(sys.argv[1]) if len(sys.argv) > 1 else 1
corpus = json.load(open("/root/corpus_ids.json"))
ids = corpus["ids"]
shared = ids[0:117968] # 131072 - 13104 (bench_corpus split_lens @ 0.9)
warm = ids[2071040:2071040 + 64] # SPARE_BASE slice, same as bench warmup
sess = requests.Session()
sess.trust_env = False
payload = {
"input_ids": shared + warm,
"sampling_params": {"max_new_tokens": 8, "temperature": 0.0, "ignore_eos": True},
}
for i in range(N):
t0 = time.perf_counter()
r = sess.post("http://127.0.0.1:30000/generate", json=payload, timeout=1800)
print(f"[warm90] {i + 1}/{N} http={r.status_code} wall={time.perf_counter() - t0:.2f}s", flush=True)