From c5d91ceafefc2c4ec6598b0af7d82e3e3cc4a7f7 Mon Sep 17 00:00:00 2001 From: yy-fighting <2351884576@qq.com> Date: Thu, 10 Sep 2026 18:52:54 +0800 Subject: [PATCH] b300-equivalent matrix: E7b high-concurrency retest (MRR64 + decode-graph buckets 1-64) - graph-drop cliff fixed, report numbers overwritten in place - deploy_glm53_e7b_hicc.sh (md5 165db732): only delta vs 607_exp = MRR 16->64 + cuda-graph-bs-decode 1..64; KV pool 276,480 unchanged, avail 6.11GB after capture - 10 retest points (16K/4.1/4.2 at c8/16/32/64) all OK, hit=0.0 (fresh container = recycled windows virgin again), 0 retraction, QG 7/7 - verdicts: 16K output 92.1/97.6/99.6 (+18~33% vs initial, still TP2PP4-dominated, prefill wall ~100 plateau); 4.1 c32/64 229/272 (gap narrowed to 1.2x); 4.2 402/676/826.5 - E7b wins ALL cc tiers, c64 826.5 tok/s = machine-wide best output (+72% vs TP2PP4 482), TTFT 12.18s / TPOT 85.1ms; c8 anchors within +-3% prove no env drift - REPORT.md + Feishu A7V3wZTQeifCB4krdi6cA834nW9 overwritten in place (user directive: no appended chapter); initial MRR16 run archived as baseline in results/e7b/ - provenance.md: e7b64 VRAM (idle 79.3k, peak 83,627 MiB), second in-service restore verified (fired up/health 200/16K+C4 spot/KV pool 647,040 identical) --- deploy/CURRENT.md | 2 +- .../README.md | 38 ++--- .../REPORT.md | 121 ++++++++-------- .../results/e7b64/all_results.jsonl | 10 ++ .../results/e7b64/md5_assets.txt | 3 + .../results/e7b64/retest_compare.md | 27 ++++ .../results/e7b64/server_facts.txt | 19 +++ .../results/e7b64/status.txt | 10 ++ .../results/provenance.md | 15 +- .../scripts/deploy_glm53_e7b_hicc.sh | 69 +++++++++ .../scripts/gen_retest_compare.py | 108 ++++++++++++++ .../scripts/run_retest_e7b64.sh | 133 ++++++++++++++++++ 12 files changed, 477 insertions(+), 78 deletions(-) create mode 100644 experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/all_results.jsonl create mode 100644 experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/md5_assets.txt create mode 100644 experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/retest_compare.md create mode 100644 experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/server_facts.txt create mode 100644 experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/status.txt create mode 100644 experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/deploy_glm53_e7b_hicc.sh create mode 100644 experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/gen_retest_compare.py create mode 100644 experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/run_retest_e7b64.sh diff --git a/deploy/CURRENT.md b/deploy/CURRENT.md index b377fc8..1eb96fc 100644 --- a/deploy/CURRENT.md +++ b/deploy/CURRENT.md @@ -14,7 +14,7 @@ | 60.5 | `glm53-nvfp4`(Up 2d,09-09 只读核验) | **NVFP4 团队生产**(deploy_glm53_605.sh,md5 fcd9109b)。生产机铁律:不实验、不重启、不覆盖脚本。**交付升级路径已更新为 v3(09-09 hit90 场景优胜 TP4PP2-hicache@0.90,池 647,040/c4 并发独立文档/hit90 cc8 out +34% vs 现役):脚本 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/scripts/deploy_glm53_605_v3.sh` + 补丁束(60.8:/root/glm53_r37_patch_bundle_v3.tar.gz,md5 6922e534,需 scp 至 60.5)+ 需 docker pull cu13 镜像;未执行、未落 60.5 磁盘(60.5:/root 仅有原脚本,核验过)。此前 v2(TP2PP4-hicache,冷缓存口径优胜)被 v3 取代,仍留仓库可作"容量优先 6 条文档/512k 单条最快"备选;回滚=原脚本** | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(方案 A 口径);待切 `glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`(容量优先备选 `..._tp2pp4_hicache.env`) | | 60.6 | 无容器,但 8 卡被外部裸金属实验占用(`/data/hzy/sparse-opd-*`,09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — | | 60.7 | 基本空(4 卡仍有 `/home/user/dirA_exp` 外部小任务,09-08 晚实测) | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — | -| 60.8 | `glm53-nvfp4`(Up,:30000,restart=unless-stopped) | **TP4PP2-nomtp@0.90 hit90 优胜口径在役**(09-09 晚接替同日早间的 TP2PP4@0.85:同 r37 栈 cu13+9 挂载+de-GLOO,仅改 tp4/pp2 + memfrac 0.90;KV 池 **647,040**;实测实例原样保留切 restart 策略,未重建容器。启动 `bash /root/deploy_ppmtp_r37.sh '--tp 4 --pp-size 2 --disable-overlap-schedule --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length 1048576' nomtp 8192 0.90 1 1 0 0`)。hit90(90% 命中 i128k/o512)out cc1/2/3/4/8/16 = **28.8/47.5/63.9/76.2/106.4/125.7**(vs TP2PP4 基线 cc4 +14%/cc8 +15%);并发独立 128k 文档 **4 条**(cap cc4 98.5 零排队)、512k 单条 ✓ TTFT 135.8s、质量门 **7/7**;同轮判决:DP attention 容量负收益判死、DCP 对 DSA 静默算错禁用、MTP@128k accept 2.07 判负。让步项(知情):容量 4 条(TP2PP4 为 6)、512k 单条慢 33%。回滚 TP2PP4 = 同命令改 `'--tp 2 --pp-size 4'` + memfrac 0.85。**09-10 B300 对标战役**:停役(容器 rename 保全 `glm53-nvfp4-insvc`)→ 双臂 B300 场景矩阵(TP2PP4-D 口径 24 点 + E7b 配方 16+1 点,判决=分界 C8/E7b 窗口≤C8/边界仅 TP2PP4 可达/与 B300 绝对差 4-5×,报告飞书 wiki `A7V3wZTQeifCB4krdi6cA834nW9`)→ **原容器恢复并核验**(rename 回 + start,health 200、16K 抽测 ok、显存水位与停役前一致,口径未变) | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`;实验全量 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/`;B300 对标 `experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/`;前史 `experiments/pro6000/glm53_nvfp4_128k_capacity_topology/`(TP2PP4 口径)+ r37 `experiments/pro6000/glm53_ppmtp_r37_verify_graph/` | +| 60.8 | `glm53-nvfp4`(Up,:30000,restart=unless-stopped) | **TP4PP2-nomtp@0.90 hit90 优胜口径在役**(09-09 晚接替同日早间的 TP2PP4@0.85:同 r37 栈 cu13+9 挂载+de-GLOO,仅改 tp4/pp2 + memfrac 0.90;KV 池 **647,040**;实测实例原样保留切 restart 策略,未重建容器。启动 `bash /root/deploy_ppmtp_r37.sh '--tp 4 --pp-size 2 --disable-overlap-schedule --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length 1048576' nomtp 8192 0.90 1 1 0 0`)。hit90(90% 命中 i128k/o512)out cc1/2/3/4/8/16 = **28.8/47.5/63.9/76.2/106.4/125.7**(vs TP2PP4 基线 cc4 +14%/cc8 +15%);并发独立 128k 文档 **4 条**(cap cc4 98.5 零排队)、512k 单条 ✓ TTFT 135.8s、质量门 **7/7**;同轮判决:DP attention 容量负收益判死、DCP 对 DSA 静默算错禁用、MTP@128k accept 2.07 判负。让步项(知情):容量 4 条(TP2PP4 为 6)、512k 单条慢 33%。回滚 TP2PP4 = 同命令改 `'--tp 2 --pp-size 4'` + memfrac 0.85。**09-10 B300 对标战役**:停役(容器 rename 保全 `glm53-nvfp4-insvc`)→ 双臂 B300 场景矩阵(TP2PP4-D 口径 24 点 + E7b 配方)→ E7b 高并发调参复测(初测 MRR16+图1-8 掉图断崖 → MRR64+图桶 1-64,`deploy_glm53_e7b_hicc.sh`,三场景 c8-c64 共 10 点全 OK;报告正文采用复测值:掉图断崖已修复、16K 仍 TP2PP4 占优(E7b 被 prefill 墙封 ~100 tok/s 平台)、decode 密集 1K→4K E7b 全档反超(c64 out 826.5 tok/s 全场最高、超 TP2PP4 72%)、边界仅 TP2PP4 可达、与 B300 绝对差 4-5×;报告飞书 wiki `A7V3wZTQeifCB4krdi6cA834nW9`,已整文更新)→ **原容器恢复并二次核验**(rename 回 + start,fired up/health 200/16K 单发+C4 抽测 ok、KV 池 647,040 与启动口径逐字一致) | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`;实验全量 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/`;B300 对标 `experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/`;前史 `experiments/pro6000/glm53_nvfp4_128k_capacity_topology/`(TP2PP4 口径)+ r37 `experiments/pro6000/glm53_ppmtp_r37_verify_graph/` | ## 方案 A-F 一览(GLM-5.3-NVFP4 @ pro6000,2026-09-08 双场景报告口径) diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/README.md b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/README.md index 60866ec..f4ddc10 100644 --- a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/README.md +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/README.md @@ -3,47 +3,48 @@ 日期:2026-09-10 | 机器:174.1.60.8(8×RTX 6000D,96GB GDDR7,无 NVLink) 模型:GLM-5.3-NVFP4(modelopt)| 镜像:`nightly-dev-20260828-daf63171`(两臂同) 对标基线:飞书《GLM 5.3 | SGLang | Low-latency & High-Throughput 测试结果》(B300 报告,wiki UPB2w4Y5yi65qwkMxJJcZko5nUc) -完整报告:本目录 `REPORT.md`(= 飞书发布版 A7V3wZTQeifCB4krdi6cA834nW9) +完整报告:本目录 `REPORT.md`(= 飞书发布版 A7V3wZTQeifCB4krdi6cA834nW9;E7b 高并发复测后整文更新,正文采用复测值) ## 目标 -在 6000D 上复刻 B300 报告的全部场景(主场景 16K→512、4.1 短输入、4.2 长输出、5.1 长上下文、5.2 边界),对两套在役部署方案各跑一遍完整矩阵,产出对齐 B300 8 章结构的对标报告。测后 60.8 在役服务(TP4PP2@0.90)原容器恢复(已验证:health 200 + 16K 抽测 ok + 显存水位一致)。 +在 6000D 上复刻 B300 报告的全部场景(主场景 16K→512、4.1 短输入、4.2 长输出、5.1 长上下文、5.2 边界),对两套在役部署方案各跑一遍完整矩阵,产出对齐 B300 8 章结构的对标报告。E7b 初测暴露 MRR16 + decode 图 bs1-8 的高并发掉图断崖后,按用户决策调参为 MRR64 + 图桶 1-64(`deploy_glm53_e7b_hicc.sh`,其余配方逐字不变)复测三场景 C=8/16/32/64 共 10 点,**报告正文一律采用复测值**(C=1 与 5.1/5.2 点受池上限约束、与调参无关,沿用初测值;C=8 锚点前后偏差 ≤3% 证两轮环境无漂移)。测后 60.8 在役服务(TP4PP2@0.90)原容器恢复并二次验证。 ## 实验臂 | 臂 | 方案 | 关键配置 | 质量门 | |---|---|---|---| | tp2pp4 | **D 生产口径**(deploy_glm53_pp4.sh) | TP2PP4、mem0.85、MRR48、cps16384、radix 关、KV fp8_e4m3 池 1,040,384、无投机、index_topk_freq=4(=原生默认,恒等)、ctx 1,048,576 | 6/7(仅 tool-call:无 parser,历史已知) | -| e7b | **TP8+EAGLE3+AR**(deploy_glm53_607_exp.sh + CAR 补丁注入) | TP8、EAGLE 4/1/5、mem0.90、MRR16、cps8192、radix 开+hicache×3、KV fp8_e4m3 GPU 池 276,480、decode 图 bs1-8、ctx 270,336、custom-AR 1stage(`SGLANG_CUSTOM_ALLREDUCE_ALGO=1stage`,8 rank `SSKJ_CAR_PATCH_ACTIVE` 验证) | 7/7 | +| e7b | **TP8+EAGLE3+AR 初测**(deploy_glm53_607_exp.sh + CAR 补丁注入) | TP8、EAGLE 4/1/5、mem0.90、MRR16、cps8192、radix 开+hicache×3、KV fp8_e4m3 GPU 池 276,480、decode 图 bs1-8、ctx 270,336、custom-AR 1stage(`SGLANG_CUSTOM_ALLREDUCE_ALGO=1stage`,8 rank `SSKJ_CAR_PATCH_ACTIVE` 验证) | 7/7 | +| e7b64 | **TP8+EAGLE3+AR 高并发复测**(deploy_glm53_e7b_hicc.sh,= e7b 仅改 MRR 16→64 + decode 图桶 1-64) | 其余配方与 e7b 逐字一致;图捕获后 avail 6.11 GB/卡,KV 池 276,480 不变 | 7/7 | ## 场景矩阵与并发档位(用户裁决收敛:16K 封 64、64K/128K 封 8/4) | B300 章节 | 场景 | 并发档位 | TP2PP4 活跃上限 | E7b 活跃上限 | |---|---|---|---|---| -| §3 主场景 | 16K→512 | 1/8/16/32/64 | 48(MRR) | 16(MRR=池贴边) | -| §4.1 | 1K→128 | 1/8/32/64 | 48 | 16 | -| §4.2 | 1K→4K | 1/8/32/64 | 48 | 16(c64 用户中止 rc=143) | +| §3 主场景 | 16K→512 | 1/8/16/32/64 | 48(MRR) | 16(池 276,480÷17.4K;复测臂 MRR64 不再是约束) | +| §4.1 | 1K→128 | 1/8/32/64 | 48 | 64(复测臂 MRR;初测 16) | +| §4.2 | 1K→4K | 1/8/32/64 | 48 | ~52(池;复测臂 c64 有 12 条排队;初测 c64 中止 rc=143,复测已补齐) | | §5.1 | 64K→512 | 1/4/8 | 15(池) | 4(池) | | §5.1 | 128K→512 | 1/2/4 | 7(池) | 2(池) | | §5.2 | 256K→1 | 1/2/3 | 3(池) | 1(池贴边) | | §5.2 | 512K→1 | 1 | 1 | 结构性不可(ctx) | | §5.2 | 896K→1(代 B300"约1M") | 1 | 1 | 结构性不可(ctx) | -测量协议:冷缓存(shared-frac 0)+ 每点 flush + 服务端命中核验 ≤0.01(超限重试一次);nreq=max(8, 2×cc);P95 nearest-rank 对齐 B300;语料耗尽(21.23M/21.30M)下用回收窗口(`--pool-override` 基址映射见 REPORT 附录 A)。有效测量点 41 个(tp2pp4 24 + e7b 16 + e7b 256K C=1 全新文本重测),全点 0 retraction、0 OOM。 +测量协议:冷缓存(shared-frac 0)+ 每点 flush + 服务端命中核验 ≤0.01(超限重试一次);nreq=max(8, 2×cc);P95 nearest-rank 对齐 B300;语料耗尽(21.23M/21.30M)下用回收窗口(`--pool-override` 基址映射见 REPORT 附录 A;复测臂为全新容器实例,同基址文本对 hicache 宿主层重新成为处女文本,命中核验全 0.0)。有效测量点 44 个(tp2pp4 24 + e7b/e7b64 合计 20:初测 10 点沿用 + 复测 10 点采用),全点 0 retraction、0 OOM。 -## 判决速览(详见 REPORT.md) +## 判决速览(详见 REPORT.md,E7b 数值均为复测值) -- **分界 C=8**:E7b 全场景 C=1 占优(16K out 49.6 vs 16.7=3.0×、TPOT 13.7 vs 50.3ms=3.7×;1K→4K out 135 vs 20.3=6.7×);C≥8 TP2PP4 全指标反超并随并发拉大(16K c64 out 211 vs 84=2.5×)。 -- **E7b 可用窗口 ≤C8**:MRR16 + decode 掉图(bs>8)双击,16K c16 TPOT 298ms 断崖。 -- **decode 密集甜点 = E7b c8**:1K→4K out 412 tok/s、TPOT 22.5ms、accept 4.0。 -- **TP2PP4 甜点 c16+**:16K 近线性至 c64(MRR48 未饱和);全场最高输出 4.2 c64 482 tok/s(但 TTFT 338s,仅离线)。 -- **边界只有 TP2PP4 可达**:256K/512K/896K 全测(input 7,050/5,662/4,153 tok/s);E7b ctx 270,336 结构性封顶。 +- **掉图断崖已修复(复测核心判决)**:初测 MRR16 + 图 bs1-8 把 E7b 窗口封死 C≤8(16K c16 TPOT 298ms 断崖);调参 MRR64 + 图桶 1-64 后 C=16 TPOT 降至 228.8ms,16K c16/32/64 输出 92.1/97.6/99.6 tok/s(较初测 +18~33%),窗口扩到 C=64。 +- **分界负载形态化**:prefill 密集(16K 主场景)仍 TP2PP4 占优——c64 out 211 vs 99.6 = 2.1×(自初测 2.5× 收窄),E7b 输出被 prefill 墙(chunk8192+TP8 无 PP 流水)封在 ~100 tok/s 平台;短输入 c32/64 TP2PP4 领先收窄到 1.2×(276/341 vs 229/272),且 E7b c64 TTFT 反超(12.15 vs 19.96s)。 +- **decode 密集(1K→4K)E7b 全档反超**:c8/32/64 = 402/676/**826.5 tok/s**,c64 为全场最高输出吞吐(超 TP2PP4 同点 482 达 72%),TTFT 12.18s、TPOT 85.1ms 双优;初测排队断崖(c32 TTFT 305s)消除为 6.19s。 +- **E7b C=1 优势不变**:16K out 49.6 vs 16.7=3.0×、TPOT 13.7 vs 50.3ms=3.7×;1K→4K out 135 vs 20.3=6.7×、TPOT 8.7ms。 +- **TP2PP4 甜点 c16+**:16K 近线性至 c64(MRR48 未饱和);边界 256K/512K/896K 只有 TP2PP4 可达(input 7,050/5,662/4,153 tok/s);E7b ctx 270,336 结构性封顶。 - **DSA 复现**:C=1 TPOT 对上下文不敏感(50.3/50.0/49.7ms @16/64/128K),并发才是驱动(c8: 70.9→161.8ms)。 -- **vs B300**:定性结构完全复现(LL/HT 分野一致),绝对差 4-5×,边界 prefill 差距收窄至 ~2×;分界点本机更靠前(C8 vs C64-128),原因是容量上限(MRR/池)而非算力。 +- **vs B300**:定性结构复现(LL/HT 分野一致),绝对差 4-5×,边界 prefill 差距收窄至 ~2×;主场景分界本机更靠前(C8 vs C64-128,容量上限而非算力),decode 密集场景调参后 E7b 全档无交叉(B300 报告未呈现该形态)。 ## 关键坑位(复测必读) -1. **hicache 宿主层陷阱**:256K prewarm KV 占池 94.8% 触发宿主层下放,`flush_cache` 清不掉宿主层 → 同文本测量命中 0.9998。冷缓存复测**必须换该实例从未发过的文本**(本战役 E7b 256K C=1 用窗口 9,900,000 重测达标)。 +1. **hicache 宿主层陷阱**:256K prewarm KV 占池 94.8% 触发宿主层下放,`flush_cache` 清不掉宿主层 → 同文本测量命中 0.9998。冷缓存复测**必须换该实例从未发过的文本**(本战役 E7b 256K C=1 用窗口 9,900,000 重测达标;e7b64 复测臂为全新容器实例,沿用同基址窗口即满足该条件,10 点命中全 0.0)。 2. 语料已耗尽:回收窗口复用仅在 flush+命中核验协议下有效;TP2PP4 臂 radix 本来就关,零污染。 3. 在役保全流程:`docker stop` → `docker rename glm53-nvfp4 glm53-nvfp4-insvc`(必须先改名,E7b 部署脚本会 rm -f 同名容器)→ 测毕 `rename` 回 + `start`。docker stop/rm 偶发 "zombie PID" 报错是收尾边界现象,容器终态 exited(137)、显存归零,稍等重试即可。 @@ -53,6 +54,9 @@ 1e34dd8d813ccb7ac8243e8d1573a49e bench_corpus_v2.py (scripts/) 4c126d067d33b5ea27c268f37561634c extract_summary.py (scripts/) 5892b44610b2ce61f533f1721105625e run_b300_matrix.sh (scripts/) +165db732fa3237a80a4d53a963a128e7 deploy_glm53_e7b_hicc.sh (scripts/,E7b 高并发版:MRR64+图桶1-64,其余与 607_exp 逐字一致) +fe46da2eb22f5ae23eb963bb9d467922 run_retest_e7b64.sh (scripts/,10 点复测驱动,run-id 96xx) +376bedfd42fa29ba9de8bfdf0c67c3b2 gen_retest_compare.py (scripts/,复测前后对比表生成器) def3c64c5dc19e1d507080e3366c3762 deploy_glm53_pp4.sh (见 dual_scenario_bench/scripts/,md5 对照一致) 21db641f1d997e26fcf1ddc97163b87e deploy_glm53_607_exp.sh (见 dual_scenario_bench/scripts/,md5 对照一致) a8fc9a504c041acc40831a848b4985d8 patches/custom_all_reduce.py @@ -63,6 +67,6 @@ gen_report_tables.py 为本地表格生成器(md5 未入台账,60.8 侧执 ## 原始数据 -- 本目录 `results/{tp2pp4,e7b}/`:all_results.jsonl(逐点 SUMMARY + 命中核验)、status.txt、server_facts.txt(启动参数+池分配日志摘录)、gpu_inventory_idle/final.csv、vram_timeline.csv(30s 采样全矩阵) -- 60.8 侧:`/root/bench_logs/b300eq_tp2pp4_20260910_1138/`、`/root/bench_logs/b300eq_e7b_20260910_1428/` +- 本目录 `results/{tp2pp4,e7b,e7b64}/`:all_results.jsonl(逐点 SUMMARY + 命中核验)、status.txt、server_facts.txt(启动参数+池分配日志摘录)、gpu_inventory_idle/final.csv、vram_timeline.csv(30s 采样全矩阵;csv/log 属 gitignore 中间件,关键事实折叠于 provenance.md) +- 60.8 侧:`/root/bench_logs/b300eq_tp2pp4_20260910_1138/`、`/root/bench_logs/b300eq_e7b_20260910_1428/`(初测留档基线)、`/root/bench_logs/b300eq_e7b64_20260910_1732/`(复测,正文采用值)+ `/root/bench_logs/retest_compare.md`(前后对比,本目录 results/e7b64/ 有同名镜像) - E7b 256K C=1:all_results.jsonl 同 tag 共 4 条,最后一条为干净重测值(生成器 dict 载入后写覆盖,天然生效) diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/REPORT.md b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/REPORT.md index 78f5703..8a0c152 100644 --- a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/REPORT.md +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/REPORT.md @@ -1,23 +1,23 @@ # GLM-5.3-NVFP4 | RTX 6000D | SGLang 双方案 B300 对标场景压测报告 -- 测试日期:2026-09-10(单日单机完成两臂) +- 测试日期:2026-09-10(单日单机完成两臂;E7b 臂当日调参 MRR 64 + decode 图 ≤64 后完成高并发复测,正文一律采用复测值) - 测试机:174.1.60.8(6000D,8 卡) - 对标基线:飞书《GLM 5.3 | SGLang | Low-latency & High-Throughput 测试结果》(B300 报告,wiki UPB2w4Y5yi65qwkMxJJcZko5nUc) -- 测后状态:60.8 在役服务(TP4PP2@0.90 口径)已原容器恢复并验证(health 200 + 16K 抽测 ok + 显存水位与停役前一致) +- 测后状态:60.8 在役服务(TP4PP2@0.90 口径,KV 池 647,040)已原容器恢复并二次验证(两轮测试后均按 rename→start 流程恢复;本轮复测拆台后 fired up + health 200 + 16K 单发与 C=4 抽测 ok,配置与池逐字一致) ## 1. 结论摘要 本轮在单台 8 卡 RTX 6000D 上,用 GLM-5.3-**NVFP4** 完整复刻 B300 报告的场景矩阵,测试了两套在役部署方案。两套方案代表完整部署形态,不是单参数 A/B:**TP2PP4** 为 D 生产口径(吞吐/长上下文形态),**TP8+EAGLE3+AR** 为 E7b 配方(低延迟形态,含 custom allreduce 1stage 补丁)。 - **低并发优先 E7b(TP8+EAGLE3+AR)**:主场景 `16K→512, C=1` 输出 49.6 tok/s、TPOT 13.7 ms,对 TP2PP4(16.7 tok/s、50.3 ms)分别是 **3.0×** 与 **3.7×**;长输出 `1K→4K, C=1` 输出 135 tok/s、TPOT 8.7 ms,对 TP2PP4(20.3、49.8 ms)是 **6.7×** 与 **5.7×**。 -- **高并发优先 TP2PP4**:主场景 C=64 达到 6,748 input tok/s / 211 output tok/s,对 E7b(2,697 / 84.3)均为 **2.5×**;短输入 C=32/64 输出 276/341 tok/s,对 E7b(100/103)为 **2.7~3.3×**。 -- **E7b 的可用并发窗口比 B300 Low-Latency 窄一个数量级**:MRR=16 且 CUDA graph 仅覆盖 decode bs 1–8,并发 ≥8 即掉图,主场景 C=16 TPOT P95 跳到 297.8 ms(C=8 为 135.1 ms;同点 TP2PP4 仅 98.5 ms)。E7b 的生产甜点上限 = **C≤8**。 -- **TP2PP4 甜点在 C=16 之后**:主场景输出吞吐从 C=8 的 101 近线性爬到 C=64 的 211 tok/s(MRR48 尚未饱和),4.2 长输出 C=64 达全场最高 482 tok/s(但 TTFT P95 338 s,需要排队预算)。 -- **decode 密集低并发的最优解是 E7b C=8**:`1K→4K, C=8` 输出 412 tok/s、TPOT 22.5 ms,对 TP2PP4(101 tok/s、79.4 ms)为 4.1×;EAGLE 实测 accept length 4.0。 +- **prefill 密集场景(16K 输入)高并发优先 TP2PP4**:主场景 C=64 达到 6,748 input tok/s / 211 output tok/s,对 E7b(3,188 / 99.6)为 **2.1×**;短输入 C=32/64 TP2PP4 仍占优(276/341 vs 229/272 tok/s)但差距收窄到 **1.2×**。 +- **E7b 初测的掉图断崖是配置产物,调参后已消除**:初测 MRR=16 + decode 图仅覆盖 bs 1–8,并发 >8 即掉图(主场景 C=16 TPOT P95 298 ms)。按决策调参为 **MRR=64 + decode 图桶 1–64**(其余配方逐字不变)后复测:C=16 TPOT P95 降至 228.8 ms,16K 场景 C=16/32/64 输出 92.1/97.6/99.6 tok/s(较初测 +18~33%),E7b 可用并发窗口从 C≤8 扩到 **C=64**。 +- **TP2PP4 甜点在 C=16 之后**:主场景输出吞吐从 C=8 的 101 近线性爬到 C=64 的 211 tok/s(MRR48 尚未饱和);其 4.2 长输出 C=64 的 482 tok/s 已被 E7b 反超(826.5),但 16K 主场景仍是本机 prefill 吞吐之王。 +- **decode 密集负载(1K 进、长出)E7b 全并发档最优**:`1K→4K` C=8/32/64 输出 402/676/**826.5 tok/s**——C=64 为全场最高输出吞吐(超 TP2PP4 同点 482 达 72%),TTFT P95 12.2 s、TPOT P95 85.1 ms,EAGLE accept 3.8~4.0。 - **长上下文与容量边界只有 TP2PP4 可达**:128K C=1 两方案输出打平(11.8 vs 11.9 tok/s)但 TP2PP4 TTFT 减半(18.2 s vs 37.6 s);256K/512K/896K 边界 E7b 结构性不可测(ctx 270,336 封顶 + KV 池 276,480 贴边),TP2PP4 全部完成(256K C=1/2/3、512K/896K C=1)。 - **DSA 特性在 6000D 复现**:TP2PP4 C=1 的 TPOT 对上下文长度不敏感(16K/64K/128K = 50.3/50.0/49.7 ms 恒定),并发才是 TPOT 驱动因子(16K 行 C=8→C=64:70.9→203.7 ms)。 - **与 B300 的绝对差距约 4~5×**,边界 prefill 差距收窄到约 2×(256K C=1 input 7,050 vs 17,457;896K 4,153 vs 约1M 行 7,897)。硬件与量化口径不同(B300 报告未写明量化方式),绝对值仅量级可比,两份报告的结构性结论一致(见第 9 章)。 -- 全部 41 个有效测量点 **0 回退(retraction)、0 OOM**,冷缓存命中核验全部 ≤0.01(E7b 256K C=1 首测触 hicache 宿主层陷阱,用全新文本重测达标,见 5.2 注记)。 +- 全部 44 个有效测量点 **0 回退(retraction)、0 OOM**,冷缓存命中核验全部 ≤0.01(E7b 256K C=1 首测触 hicache 宿主层陷阱,用全新文本重测达标,见 5.2 注记;E7b 复测 10 点命中核验全部 0.0)。 ## 2. 测试环境与配置 @@ -29,17 +29,19 @@ | 并行 | TP2 × PP4 | TP8 | | 投机解码 | 无 | EAGLE3,num_steps=4,topk=1,draft_tokens=5 | | `mem-fraction-static` | 0.85 | 0.90 | -| 最大活跃请求(MRR) | 48 | 16 | +| 最大活跃请求(MRR) | 48 | 64 | | Chunk Prefill | 16,384 | 8,192 | | KV dtype | fp8_e4m3 | fp8_e4m3 | | KV 池(服务端实测) | **1,040,384 tokens**(12.6~13.4 GB/rank,无宿主层) | **276,480 tokens GPU**(15.8 GB/rank)+ 分层缓存 hicache×3 宿主层(write_through) | | radix cache | 关(`disable_radix_cache=True`) | 开(分层缓存) | | 上下文上限 | 1,048,576(config 原生) | 270,336(显存约束下的部署值) | -| CUDA graph | 常规捕获 | decode 图 bs 1–8(bs>8 掉图) | +| CUDA graph | 常规捕获 | decode/verify 图桶 bs 1–64(1,2,3,4,6,8,12,16,24,32,48,64) | | custom allreduce | — | 1stage 补丁注入(`SGLANG_CUSTOM_ALLREDUCE_ALGO=1stage` 环境强制;8 rank `SSKJ_CAR_PATCH_ACTIVE` 日志验证全出现) | | `index_topk_freq` | 4(override,等于原生默认,恒等) | 原生默认 4 | | 质量门 | 6/7(仅 tool-call 失败:D 口径未配 parser,历史已知;其余全过) | **7/7** | +**E7b 臂配置演进注记**:E7b 初测为 MRR=16 + decode 图桶 bs 1–8,高并发点(C>8)出现 decode 掉图 + MRR 排队双击。按决策将 MRR 调至 64、decode 图桶扩至 1–64(部署脚本 `deploy_glm53_e7b_hicc.sh`,其余配方与在役 E7b 逐字一致),三场景 C=8/16/32/64 共 10 点全部重测,正文一律采用复测值;C=1 各点与 5.1/5.2 长上下文点受 KV 池上限约束(活跃 1~4 条),行为与该调参无关,沿用初测值。C=8 锚点前后偏差 ≤3%(16K 81.4→83.9、1K 152→151、1K→4K 412→402 tok/s),证明两轮环境无漂移。初测原始数据留档于服务器 `b300eq_e7b_20260910_1428/` 与库内 `results/e7b/`。 + 因此,下文比较回答的是"两种部署形态谁更适合该负载",不能把差异单独归因于 EAGLE、PP 流水、chunk、radix 或图覆盖中的某一项(与 B300 报告同款声明)。 **测量协议**(对齐 B300 口径): @@ -54,7 +56,9 @@ | 场景 | TP2PP4 活跃上限 | E7b 活跃上限 | |-|-|-| -| 16K / 1K | 48(MRR) | 16(MRR,=池贴边) | +| 16K | 48(MRR) | 16(池 276,480 ÷ 约17.4K;MRR64 不再是约束) | +| 1K→128 | 48(MRR) | 64(MRR) | +| 1K→4K | 48(MRR) | ~52(池;C=64 有 12 条排队) | | 64K | 15(池) | 4(池) | | 128K | 7(池) | 2(池) | | 256K | 3(池) | 1(池 262K KV / 276K 贴边) | @@ -69,20 +73,20 @@ B300 跑了 C=1/8/32/64/128/256;本机按 MRR 上限收敛为 C=1/8/16/32/64 | 1 | TP2PP4 | 534 | 16.7 | 5.36 s | 50.3 ms | | 1 | TP8+EAGLE3+AR | 1,587 | 49.6 | **3.98 s** | **13.7 ms** | | 8 | TP2PP4 | 3,236 | **101** | **14.79 s** | **70.9 ms** | -| 8 | TP8+EAGLE3+AR | 2,605 | 81.4 | 32.52 s | 135.1 ms | +| 8 | TP8+EAGLE3+AR | 2,684 | 83.9 | 32.47 s | 136.1 ms | | 16 | TP2PP4 | 4,674 | **146** | **25.78 s** | **98.5 ms** | -| 16 | TP8+EAGLE3+AR | 2,217 | 69.3 | 60.86 s | 297.8 ms | +| 16 | TP8+EAGLE3+AR | 2,947 | 92.1 | 59.91 s | 228.8 ms | | 32 | TP2PP4 | **6,185** | **193** | **46.43 s** | **152.0 ms** | -| 32 | TP8+EAGLE3+AR | 2,527 | 79.0 | 169.68 s | 259.9 ms | +| 32 | TP8+EAGLE3+AR | 3,123 | 97.6 | 140.33 s | 213.6 ms | | 64 | TP2PP4 | **6,748** | **211** | **134.78 s** | **203.7 ms** | -| 64 | TP8+EAGLE3+AR | 2,697 | 84.3 | 345.69 s | 231.5 ms | +| 64 | TP8+EAGLE3+AR | 3,188 | 99.6 | 293.21 s | 214.6 ms | 趋势: -- **分界在 C=8**:C=1 E7b 全指标占优;C=8 起 TP2PP4 全指标反超,且输出吞吐差距随并发拉大(101 vs 81 → 211 vs 84)。 -- **E7b 在 C=8→16 输出吞吐倒退**(81.4→69.3 tok/s):MRR=16 开始排队 + decode 掉图(bs>8 无图)双击;TPOT P95 从 135 ms 跳到 298 ms。EAGLE accept length 随并发从 2.14 爬到 2.99,但被掉图抵消。 -- **prefill 墙的差异**:E7b 的 input TPS 几乎不随并发增长(C=8→64:2,605→2,697,+3%),TP2PP4 翻倍(3,236→6,748,+108%)——chunk 8192 + TP8 无 PP 流水的 prefill 瓶颈 vs chunk 16384 + PP4 流水摊满。 -- TP2PP4 到 C=64 仍在爬坡(C=32→64 +9%),MRR48 未饱和;TTFT P95 在 C=64 达 134.8 s,同 B300 一样高并发 TTFT 需要准入控制。 +- **分界仍在 C=8,但差距显著收窄**:C=1 E7b 全指标占优;C=8 起 TP2PP4 四指标反超。输出吞吐差距从初测的 2.5× 收到 2.1×(211 vs 99.6),TPOT P95 在 C=64 已接近(203.7 vs 214.6 ms)。 +- **调参消除掉图断崖**:E7b 输出从 C=8 的 83.9 单调爬到 C=64 的 99.6 tok/s(+19%),C=16 TPOT P95 从初测 298 ms 降到 228.8 ms,C=32/64 稳定在 ~214 ms——decode 图全程覆盖运行批。EAGLE accept 随并发从 2.56 爬到 2.95。 +- **prefill 墙成为 E7b 的输出上限**:其 input TPS 从 C=8 的 2,684 到 C=64 仅 +19%(3,188),TP2PP4 同区间 +108%(3,236→6,748)——chunk 8192 + TP8 无 PP 流水 vs chunk 16384 + PP4 摊满。16K 场景 E7b 输出被 prefill 封死在 ~100 tok/s 平台,并发再高也不突破。 +- **E7b 本场景活跃上限 = KV 池(~16 条)**:C=32/64 为排队观察点(TTFT P95 140/293 s);TP2PP4 到 C=64 仍在爬坡(C=32→64 +9%,MRR48 未饱和),其 TTFT P95 134.8 s 同样需要准入控制。 ## 4. 短输入与长输出 @@ -93,13 +97,13 @@ B300 跑了 C=1/8/32/64/128/256;本机按 MRR 上限收敛为 C=1/8/16/32/64 | 1 | TP2PP4 | 155 | 19.3 | 386 ms | 49.3 ms | | 1 | TP8+EAGLE3+AR | **518** | **64.7** | **326 ms** | **14.3 ms** | | 8 | TP2PP4 | 815 | 102 | **1.69 s** | 72.9 ms | -| 8 | TP8+EAGLE3+AR | **1,219** | **152** | 2.16 s | **53.7 ms** | +| 8 | TP8+EAGLE3+AR | **1,211** | **151** | 2.17 s | **53.7 ms** | | 32 | TP2PP4 | **2,204** | **276** | **4.87 s** | **98.5 ms** | -| 32 | TP8+EAGLE3+AR | 801 | 100 | 25.60 s | 190.5 ms | -| 64 | TP2PP4 | **2,724** | **341** | **19.96 s** | **101.8 ms** | -| 64 | TP8+EAGLE3+AR | 826 | 103 | 65.81 s | 177.7 ms | +| 32 | TP8+EAGLE3+AR | 1,833 | 229 | 7.09 s | 150.6 ms | +| 64 | TP2PP4 | **2,724** | **341** | 19.96 s | **101.8 ms** | +| 64 | TP8+EAGLE3+AR | 2,178 | 272 | **12.15 s** | 292.7 ms | -短输入下 E7b 在 C≤8 显著占优(C=8 输出 152 vs 102,1.5×),C=32 起 TP2PP4 大幅拉开(2.7~3.3×)。E7b 的 input TPS 反而在 C=8 最高(1,219)后回落——MRR16 排队开始挤占 prefill。B300 同场景 Low-Latency 到 C=128 才被反超,本机提前到 C=8~32 之间,同样是容量上限(MRR16/池)而非算力所致。 +短输入下 E7b 在 C≤8 显著占优(C=8 输出 151 vs 102,1.5×);C=32/64 TP2PP4 输出仍领先(276 vs 229、341 vs 272)但差距只有 **1.2×**(初测为 2.7~3.3×),且 E7b 的 C=64 TTFT P95 反而更优(12.15 vs 19.96 s)——代价是 TPOT P95 292.7 ms(64 条同时在飞;TP2PP4 同点 MRR48 只保持 48 活跃,TPOT 101.8 ms)。B300 同场景 Low-Latency 到 C=128 才被反超,本机在 C=8~32 之间,主因是 MRR/池容量而非算力。 ### 4.2 `1K -> 4K` @@ -108,19 +112,17 @@ B300 跑了 C=1/8/32/64/128/256;本机按 MRR 上限收敛为 C=1/8/16/32/64 | 1 | TP2PP4 | 20.3 | 395 ms | 49.8 ms | | 1 | TP8+EAGLE3+AR | **135** | **320 ms** | **8.7 ms** | | 8 | TP2PP4 | 101 | 1.87 s | 79.4 ms | -| 8 | TP8+EAGLE3+AR | **412** | **1.68 s** | **22.5 ms** | -| 32 | TP2PP4 | **367** | **4.14 s** | **90.5 ms** | -| 32 | TP8+EAGLE3+AR | 252 | 304.75 s | 74.6 ms | -| 64 | TP2PP4 | **482** | 337.92 s | 95.5 ms | -| 64 | TP8+EAGLE3+AR | 用户中止*(见注) | — | — | +| 8 | TP8+EAGLE3+AR | **402** | **1.65 s** | **24.3 ms** | +| 32 | TP2PP4 | 367 | **4.14 s** | 90.5 ms | +| 32 | TP8+EAGLE3+AR | **676** | 6.19 s | **54.6 ms** | +| 64 | TP2PP4 | 482 | 337.92 s | 95.5 ms | +| 64 | TP8+EAGLE3+AR | **826.5** | **12.18 s** | **85.1 ms** | -\* E7b C=64 点按用户指示中止("并发 64 太高",rc=143),未获得有效数据;同点 TP2PP4 已完成。E7b 该点 nreq=128 远超 MRR=16,属排队观察点,中止不影响结论完整性。 +长输出放大了两形态的差异——调参后 E7b 在本场景全并发档反超: -长输出放大了两形态的差异: - -- **E7b C=1/C=8 是 decode 密集负载的最优区间**:C=1 输出 135 tok/s、TPOT 8.7 ms(全场最低),C=8 输出 412 tok/s(全场第二),EAGLE accept 长达 3.7~4.0——长输出让草稿模型进入"顺笔"状态,accept 显著高于 4.1 短输出行(2.0~2.1)。 -- **E7b C=32 的 TTFT P95 304.75 s** 是纯排队(nreq=64 / MRR=16,4 波串行),其 TPOT 74.6 ms 与掉图后水平一致。 -- **TP2PP4 C=64 输出 482 tok/s 为全场最高**,但 TTFT P95 338 s 意味着该点只适合离线批处理;交互负载应压在 C=32(367 tok/s、TTFT 4.1 s)。 +- **E7b 全档最优,C=64 = 826.5 tok/s 为全场最高输出吞吐**:C=1/8/32/64 输出 135/402/676/826.5,对 TP2PP4(20.3/101/367/482)为 6.7×/4.0×/1.8×/1.7×;C=64 同时拿下 TTFT(12.18 vs 337.92 s)与 TPOT(85.1 vs 95.5 ms)双优。EAGLE accept 3.7~4.0——长输出让草稿模型进入"顺笔"状态,显著高于 4.1 短输出行(2.0~2.1)。 +- **初测的排队断崖已消除**:C=32 TTFT P95 从初测 304.75 s(nreq=64 / MRR=16 四波串行)降到 6.19 s;C=64 从未完成变为 12.18 s。MRR64 下 1K→4K 的池上限约 52 条活跃,C=64 仅 12 条排队,请求几乎全程满飞。 +- **TP2PP4 本场景全程被压**:其优势场景是 prefill 密集(16K 主场景),1K 短进长出下既无 prefill 墙可摊、也无投机解码加成;C=64 输出 482 tok/s 且 TTFT P95 338 s,只适合离线批处理。 ## 5. 长上下文观察 @@ -144,7 +146,7 @@ B300 跑了 C=1/8/32/64/128/256;本机按 MRR 上限收敛为 C=1/8/16/32/64 | 128K→512 | 4 | TP8+EAGLE3+AR | 3,133 | 12.2 | 159.50 s | 163.4 ms | - 64K C=1 E7b 仍占优(22.5 vs 14.3 tok/s),但 128K C=1 两方案输出打平(11.8 vs 11.9)——prefill 逐渐成为长上下文的主导成本,E7b 的 decode 优势被稀释;其 TTFT 反而慢 2×(37.6 vs 18.2 s)。 -- C≥2 起 TP2PP4 全指标占优;E7b 的 output TPS 在 64K/128K 行几乎不随并发变化(22.5→25.6、11.9→12.2),与主场景同一形态:容量上限 + 掉图封死并发收益。 +- C≥2 起 TP2PP4 全指标占优;E7b 的 output TPS 在 64K/128K 行几乎不随并发变化(22.5→25.6、11.9→12.2)——KV 池把活跃钉在 ~4/~2 条,并发收益被容量封死(与 MRR/图调参无关,故沿用初测值)。 - DSA 的 TPOT 上下文不变性(TP2PP4 C=1:50.0 ms @64K ≈ 49.7 ms @128K ≈ 50.3 ms @16K)与并发驱动性(C=8:70.9 ms @16K → 161.8 ms @64K)在本组完整呈现。 ### 5.2 上下文边界 @@ -167,29 +169,30 @@ B300 跑了 C=1/8/32/64/128/256;本机按 MRR 上限收敛为 C=1/8/16/32/64 ## 6. 显存状态 - **TP2PP4**:服务加载后空载 64.6 GiB/卡,矩阵峰值 **85.0 GiB/卡**(主场景 C=64 时逼近打满,最紧张卡余量约 0.6 GiB)。mem 0.85 下 KV 池按卡容量贴满分配,属预期;继续上调 MRR 或上下文没有余量,扩容前必须先降 mem-fraction。 -- **TP8+EAGLE3+AR**:空载 77.9 GiB/卡(EAGLE 草稿权重 + mem 0.90 大池),矩阵峰值 **83.6 GiB/卡**(余量约 2.1 GiB)。 -- 两臂全矩阵 **0 OOM、0 retraction**(全部 41 点 retractions_total=0)——B300 未披露该指标,本机在自身容量上限内运行无回退。 +- **TP8+EAGLE3+AR**:空载 77.5 GiB/卡(EAGLE 草稿权重 + mem 0.90 大池 + 13 档 decode 图,图捕获完成后余 6.11 GB/卡),复测矩阵峰值 **81.7 GiB/卡**(余量约 3.9 GiB;初测臂含 64K/128K/256K 长上下文的 hicache 传输,峰值 83.6 GiB/卡)。 +- 两臂全矩阵 **0 OOM、0 retraction**(全部 44 点 retractions_total=0)——B300 未披露该指标,本机在自身容量上限内运行无回退。 - 显存时间线逐 30 s 采样留档(vram_timeline.csv),可复核任一时刻的卡间分布。 ## 7. 建议 -1. **低并发交互/agent 长思考(C≤8)用 E7b**:主场景 C=1 TPOT 13.7 ms、长输出 C=8 输出 412 tok/s。生产并发上限建议钉在 ≤8:C=16 起 decode 掉图 + MRR16 排队使其全面劣于 TP2PP4。 -2. **高并发吞吐/长上下文(C≥8 或输入 ≥64K)用 TP2PP4**:主场景 C=64 输出 211 tok/s、128K C=1 TTFT 18.2 s、896K 可达;MRR48 内未饱和,吞吐上限即 MRR。 -3. **负载形态分界线**:prefill 吞吐需求 >2.2K tok/s 或并发 >8 → TP2PP4;decode 为主且并发 ≤8 → E7b。两臂在 C=8 附近的输出吞吐交叉(主场景 101 vs 81、短输入 102 vs 152、长输出 101 vs 412)——按输出长度分布选型,不能只看并发。 -4. **E7b 扩窗口的两个前置**:MRR 16→更高需先扩 KV 池(hicache 宿主层只救命中场景,不增并发容量);decode 图覆盖 bs 8→16/32 才能消掉 C=16 的 298 ms TPOT 断崖。 +1. **低并发交互/agent 长思考(C≤8)用 E7b**:主场景 C=1 TPOT 13.7 ms、长输出 C=8 输出 402 tok/s;该区间对 TP2PP4 的优势最大(3.0~6.7×)。 +2. **prefill 密集/长上下文(16K 主场景、输入 ≥64K)用 TP2PP4**:主场景 C=64 输出 211 tok/s(E7b 99.6)、128K C=1 TTFT 18.2 s、896K 可达;MRR48 内未饱和,吞吐上限即 MRR。 +3. **负载形态分界线(调参后)**:decode 密集(短进长出,1K→4K)任何并发档 E7b 全优(C=64 输出 826 vs 482 tok/s);prefill 密集(16K 主场景)C≥8 TP2PP4 全优(211 vs 100);短输入短输出(1K→128)C≤8 E7b、C≥32 TP2PP4(差距仅 1.2×)。选型看输入/输出长度分布,不能只看并发。 +4. **E7b 剩余瓶颈在 prefill 墙与 KV 池,不再在图**:MRR64 + 图桶 1–64 已消掉掉图断崖(本报告即复测值);16K 主场景输出封顶 ~100 tok/s 是 chunk 8192 + TP8 无 PP 流水所致。扩并发容量的唯一杠杆是 KV 池(hicache 宿主层只救命中场景,不增并发容量)。 5. **边界与超长上下文只有 TP2PP4 口径可服务**:E7b 若要对标 B300 512K/约1M 行,需要 ctx ≥524,288 与池 ≥52 万 tokens 的部署形态,本版(ctx 270,336 / 池 276,480)结构性不可达。 6. **不要把两臂差异单归因 EAGLE**:两臂同时差在并行拓扑、chunk、radix、MRR 与图覆盖;单变量消融未做(与 B300 报告建议 3 同款)。 -7. **生产容量同时设吞吐和延迟 SLO**:TP2PP4 主场景 C=64 输出最高但 TTFT P95 已到 135 s;E7b C=32 长输出 TTFT P95 305 s。只看峰值 TPS 会掩盖排队长尾。 +7. **生产容量同时设吞吐和延迟 SLO**:TP2PP4 主场景 C=64 输出最高但 TTFT P95 已到 135 s;E7b 主场景 C=64 TTFT P95 293 s(池限 16 活跃的排队)。只看峰值 TPS 会掩盖排队长尾。 8. **分层缓存运维**:宿主层缓存不受 `flush_cache` 影响,任何冷缓存测量/复测必须更换输入文本(见 5.2 注记)。 ## 8. 原始结果与复现 - 服务器原始结果(60.8): - TP2PP4 臂:`/root/bench_logs/b300eq_tp2pp4_20260910_1138/`(all_results.jsonl 24 点、status.txt、server_facts.txt、gpu_inventory、vram_timeline.csv) - - E7b 臂:`/root/bench_logs/b300eq_e7b_20260910_1428/`(all_results.jsonl 20 条:16 点 OK + 4.2 C=64 用户中止 + 256K C=1 干净重测覆盖前 3 条污染记录;status.txt 含 4 个结构性跳过与 HIT_FAIL_FINAL 首测记录) + - E7b 臂初测(MRR16/图 1–8,留档基线):`/root/bench_logs/b300eq_e7b_20260910_1428/`(all_results.jsonl 20 条:16 点 OK + 4.2 C=64 用户中止 + 256K C=1 干净重测;status.txt 含 4 个结构性跳过与 HIT_FAIL_FINAL 首测记录) + - E7b 臂复测(MRR64/图 1–64,**正文采用值**):`/root/bench_logs/b300eq_e7b64_20260910_1732/`(all_results.jsonl 10 点全 OK:16K/4.1/4.2 三场景 C=8/16/32/64,命中核验全 0.0、0 retraction);前后对比表 `/root/bench_logs/retest_compare.md` - 资产 md5 台账:`/root/bench_logs/b300eq_md5_ledger.txt` -- 本地镜像:`D:\sskj\b300eq\{tp2pp4,e7b}\`(上述全部文件)、`D:\sskj\b300eq\report_tables.md`(表格生成器输出) -- 部署脚本:`/root/deploy_glm53_pp4.sh`(md5 def3c64c…,与库内 sskj main 副本一致)、`/root/deploy_glm53_607_exp.sh`;CAR 补丁:`/root/patches/custom_all_reduce.py`(a8fc9a50…)+ `custom_all_reduce_utils.py`(65a4d22b…),三处(60.7 原件/本地/库内)md5 一致 +- 本地镜像:`D:\sskj\b300eq\{tp2pp4,e7b,e7b64}\`(上述全部文件)、`D:\sskj\b300eq\report_tables.md`(表格生成器输出)、`D:\sskj\b300eq\retest_compare.md`(复测前后对比) +- 部署脚本:`/root/deploy_glm53_pp4.sh`(md5 def3c64c…,与库内 sskj main 副本一致)、`/root/deploy_glm53_607_exp.sh`(E7b 在役配方)、`/root/deploy_glm53_e7b_hicc.sh`(165db732…,E7b 高并发版:MRR64 + 图桶 1–64,其余与前者逐字一致);复测驱动 `/root/run_retest_e7b64.sh`(fe46da2e…)、对比生成器 `/root/gen_retest_compare.py`;CAR 补丁:`/root/patches/custom_all_reduce.py`(a8fc9a50…)+ `custom_all_reduce_utils.py`(65a4d22b…),三处(60.7 原件/本地/库内)md5 一致 - 测量工具:`/root/bench_corpus_v2.py`(md5 1e34dd8d…,p95 nearest-rank + 逐请求 dump)、`/root/extract_summary.py`(4c126d06…)、`/root/run_b300_matrix.sh`(5892b446…,矩阵驱动:alive/idle_wait/prewarm/flush/命中核验/重试/VRAM 采样) - 复现命令(单点示例): @@ -199,6 +202,7 @@ python3 /root/bench_corpus_v2.py --input-len 262144 --output-len 1 --shared-frac --concurrency 1 --num-requests 1 --run-id 9551 --pool-override 9900000 \ --dump-records $L/b52_256k_c1_v2_records.jsonl # 全矩阵:nohup bash /root/run_b300_matrix.sh > 2>&1 & +# E7b 高并发复测(MRR64/图≤64,10 点):nohup bash /root/run_retest_e7b64.sh > 2>&1 & ``` ## 9. 与 B300 对比观察 @@ -206,17 +210,17 @@ python3 /root/bench_corpus_v2.py --input-len 262144 --output-len 1 --shared-frac > **口径声明**:B300 报告未写明模型量化方式(若为原始 BF16 权重,则与本机 NVFP4 非同模型形态);硬件为 8×B300(288 GB HBM3e)vs 本机 8×RTX 6000D(96 GB GDDR7);镜像 v0.5.18-cu130-dev4 vs nightly-20260828-daf63171。**绝对值仅量级可比,本对比只对结构性结论负责**。 - **定性结构完全复现**:低延迟配方(B300 Low-Latency = TP8+EAGLE vs 本机 E7b = TP8+EAGLE3+AR)在 C=1 占优、吞吐配方(B300 High-Throughput = DP8+DeepEP vs 本机 TP2PP4 = D 生产口径)在高并发占优——两套硬件上"低延迟 vs 高吞吐"的分野方向一致。 -- **分界点本机更靠前**:B300 的交叉点在 C=64~128(主场景 HT C=128 反超 33%);本机在 C=8 附近。原因不是算力而是**容量上限**:本机两臂 MRR/池上限(48/16)远小于 B300 配方的 256/默认,先于算力撞墙。 +- **分界点本机更靠前,且调参后由负载形态决定**:B300 的交叉点在 C=64~128(主场景 HT C=128 反超 33%);本机主场景交叉仍在 C=8 附近(容量上限所致:本臂 MRR64/池贴边 16 活跃 vs B300 配方 256/默认,先于算力撞墙),但 decode 密集场景(1K→4K)E7b 调参后全档占优、无交叉——这一形态差异 B300 报告未呈现。 - **绝对差距 4~5×(主场景)**:C=1 输出 246 vs 49.6 tok/s(5.0×)、input 7,882 vs 1,587(5.0×);吞吐侧峰值 997 vs 211(4.7×)、31,889 vs 6,748(4.7×)。与显存带宽硬件代差量级一致。 - **边界 prefill 差距收窄到 ~2×**:256K C=1 input 7,050 vs 17,457(2.5×)→ 512K 5,662 vs 12,421(2.2×)→ 896K/约1M 4,153 vs 7,897(1.9×)。计算密集的超长 prefill 是 6000D 相对最能打的位置(PP 流水摊满 + 带宽占比下降)。 - **TPOT 差距小于吞吐差距**:B300 LL C=1 4.36 ms vs E7b 13.7 ms(3.1×);高并发侧 B300 HT C=128 165 ms vs TP2PP4 C=64 204 ms(1.2×)——NVFP4 + DSA 把 decode 单步成本压得相对不差,差距主要在吞吐面。 -- **饱和形态不同**:B300 LL 在 C=64 后进入 24K input tok/s 平台、HT 在 C=128 达峰后 C=256 回退 19%;本机 TP2PP4 到 C=64 仍在爬坡(MRR 未饱和),E7b 则被 MRR16+掉图封死在 C=8。本机没有一档出现吞吐回退——"甜点=并发上限"由 MRR 决定而非算力。 +- **饱和形态不同**:B300 LL 在 C=64 后进入 24K input tok/s 平台、HT 在 C=128 达峰后 C=256 回退 19%;本机 TP2PP4 到 C=64 仍在爬坡(MRR 未饱和),E7b 调参后在 16K 场景呈 ~100 tok/s 输出平台(池限 16 活跃 + prefill 墙)、在 1K→4K 场景爬到 826 tok/s 无回退。本机没有一档出现吞吐回退——"甜点=并发上限"由 MRR/池决定而非算力。 - **EAGLE 配方差异**:B300 LL 为 5 steps/6 draft tokens,本机 E7b 为 4 steps/topk1/5 draft tokens;本机实测 accept 2.0~4.0(短输出 2.0、主场景 2.1~3.0、长输出 3.7~4.0,随 decode 深入上升)。B300 未披露 accept,无法直接对比投机效率。 - **容量边界差距最大**:B300 两模式都完成约 1M 输入 C=1/2/4;本机仅 TP2PP4 可达 896K 且 C=1 单条(池 1,040,384 刚容一条),E7b 连 512K 都结构性不可测(ctx 270,336)。96 GB 卡上"上下文边界=显存边界"比 B300 严酷得多。 ## 附录 A:语料窗口映射(回收窗口) -语料总量 21,296,780 tokens,此前场景一/二战役已消费至 21,235,008。冷缓存协议下回收复用:窗口基址 `--pool-override` 显式指定,每点窗口在基址上顺序推进(逐记录 `corpus_window.start/end` 留档),每点 flush + 命中核验 ≤0.01 保证冷。 +语料总量 21,296,780 tokens,此前场景一/二战役已消费至 21,235,008。冷缓存协议下回收复用:窗口基址 `--pool-override` 显式指定,每点窗口在基址上顺序推进(逐记录 `corpus_window.start/end` 留档),每点 flush + 命中核验 ≤0.01 保证冷。E7b 复测臂沿用与初测相同的窗口基址(2,300,000 / 4,500,000 / 4,700,000)——复测为全新容器实例,这些文本对 hicache 宿主层重新成为"处女文本",冷缓存协议成立(复测 10 点命中核验全部 0.0)。 | 场景 | 窗口基址 | 备注 | |-|-|-| @@ -236,28 +240,29 @@ python3 /root/bench_corpus_v2.py --input-len 262144 --output-len 1 --shared-frac | b3_16k_c1 | TP2PP4 | 8/8 | 245.55 | 16.68 | 533.8 | 5.35/5.36/5.36 | 49.6/50.3/50.3 | 0 | None | | b3_16k_c1 | TP8+EAGLE3+AR | 8/8 | 82.58 | 49.6 | 1587.15 | 3.93/3.98/3.98 | 12.5/13.7/13.7 | 0 | 2.138 | | b3_16k_c8 | TP2PP4 | 16/16 | 81.0 | 101.13 | 3236.22 | 10.05/14.79/14.79 | 59.5/70.9/70.9 | 0 | None | -| b3_16k_c8 | TP8+EAGLE3+AR | 16/16 | 100.63 | 81.41 | 2605.13 | 11.75/32.52/32.52 | 71.6/135.1/135.1 | 0 | 2.375 | +| b3_16k_c8 | TP8+EAGLE3+AR | 16/16 | 97.66 | 83.88 | 2684.26 | 12.12/32.47/32.47 | 69.2/136.1/136.1 | 0 | 2.556 | | b3_16k_c16 | TP2PP4 | 32/32 | 112.16 | 146.08 | 4674.5 | 15.53/25.78/25.84 | 79.2/98.5/101.2 | 0 | None | -| b3_16k_c16 | TP8+EAGLE3+AR | 32/32 | 236.44 | 69.29 | 2217.4 | 23.60/60.86/90.39 | 173.3/297.8/305.5 | 0 | 2.517 | +| b3_16k_c16 | TP8+EAGLE3+AR | 32/32 | 177.89 | 92.1 | 2947.27 | 19.76/59.91/63.69 | 129.7/228.8/275.9 | 0 | 2.485 | | b3_16k_c32 | TP2PP4 | 64/64 | 169.54 | 193.27 | 6184.73 | 26.52/46.43/47.89 | 113.7/152.0/157.3 | 0 | None | -| b3_16k_c32 | TP8+EAGLE3+AR | 64/64 | 414.96 | 78.97 | 2526.91 | 100.26/169.68/190.73 | 169.8/259.9/319.1 | 0 | 2.767 | +| b3_16k_c32 | TP8+EAGLE3+AR | 64/64 | 335.79 | 97.58 | 3122.67 | 82.65/140.33/154.66 | 138.1/213.6/277.5 | 0 | 2.777 | | b3_16k_c64 | TP2PP4 | 128/128 | 310.79 | 210.87 | 6747.91 | 63.05/134.78/139.99 | 139.2/203.7/214.3 | 0 | None | -| b3_16k_c64 | TP8+EAGLE3+AR | 128/128 | 777.55 | 84.29 | 2697.14 | 240.10/345.69/380.96 | 168.0/231.5/311.9 | 0 | 2.986 | +| b3_16k_c64 | TP8+EAGLE3+AR | 128/128 | 657.79 | 99.63 | 3188.2 | 204.27/293.21/322.84 | 142.5/214.6/276.6 | 0 | 2.954 | | b41_1k_c1 | TP2PP4 | 8/8 | 52.95 | 19.34 | 154.7 | 0.38/0.39/0.39 | 49.1/49.3/49.3 | 0 | None | | b41_1k_c1 | TP8+EAGLE3+AR | 8/8 | 15.82 | 64.74 | 517.93 | 0.30/0.33/0.33 | 13.2/14.3/14.3 | 0 | 2.028 | | b41_1k_c8 | TP2PP4 | 16/16 | 20.1 | 101.91 | 815.3 | 1.31/1.69/1.69 | 68.6/72.9/72.9 | 0 | None | -| b41_1k_c8 | TP8+EAGLE3+AR | 16/16 | 13.44 | 152.4 | 1219.24 | 1.21/2.16/2.16 | 41.4/53.7/53.7 | 0 | 2.048 | +| b41_1k_c8 | TP8+EAGLE3+AR | 16/16 | 13.53 | 151.33 | 1210.61 | 1.25/2.17/2.17 | 42.1/53.7/53.7 | 0 | 1.992 | | b41_1k_c32 | TP2PP4 | 64/64 | 29.73 | 275.55 | 2204.4 | 3.89/4.87/4.88 | 85.6/98.5/106.8 | 0 | None | -| b41_1k_c32 | TP8+EAGLE3+AR | 64/64 | 81.8 | 100.14 | 801.15 | 17.56/25.60/27.89 | 146.6/190.5/195.2 | 0 | 2.054 | +| b41_1k_c32 | TP8+EAGLE3+AR | 64/64 | 35.75 | 229.18 | 1833.42 | 2.99/7.09/7.10 | 111.9/150.6/179.8 | 0 | 2.035 | | b41_1k_c64 | TP2PP4 | 128/128 | 48.11 | 340.53 | 2724.25 | 8.14/19.96/20.01 | 93.3/101.8/125.1 | 0 | None | -| b41_1k_c64 | TP8+EAGLE3+AR | 128/128 | 158.73 | 103.22 | 825.75 | 46.74/65.81/70.71 | 144.5/177.7/212.5 | 0 | 2.087 | +| b41_1k_c64 | TP8+EAGLE3+AR | 128/128 | 60.18 | 272.24 | 2177.88 | 4.76/12.15/13.48 | 191.1/292.7/331.5 | 0 | 2.094 | | b42_1k4k_c1 | TP2PP4 | 8/8 | 1614.16 | 20.3 | 5.08 | 0.39/0.40/0.40 | 49.2/49.8/49.8 | 0 | None | | b42_1k4k_c1 | TP8+EAGLE3+AR | 8/8 | 242.18 | 135.31 | 33.83 | 0.31/0.32/0.32 | 7.3/8.7/8.7 | 0 | 3.701 | | b42_1k4k_c8 | TP2PP4 | 16/16 | 649.69 | 100.87 | 25.22 | 1.35/1.87/1.87 | 79.0/79.4/79.4 | 0 | None | -| b42_1k4k_c8 | TP8+EAGLE3+AR | 16/16 | 159.24 | 411.54 | 102.89 | 0.96/1.68/1.68 | 18.1/22.5/22.5 | 0 | 4.003 | +| b42_1k4k_c8 | TP8+EAGLE3+AR | 16/16 | 162.84 | 402.47 | 100.62 | 0.95/1.65/1.65 | 18.5/24.3/24.3 | 0 | 3.951 | | b42_1k4k_c32 | TP2PP4 | 64/64 | 714.14 | 367.08 | 91.77 | 1.98/4.14/4.15 | 86.1/90.5/91.1 | 0 | None | -| b42_1k4k_c32 | TP8+EAGLE3+AR | 64/64 | 1040.79 | 251.87 | 62.97 | 195.07/304.75/334.69 | 60.7/74.6/92.3 | 0 | 3.852 | +| b42_1k4k_c32 | TP8+EAGLE3+AR | 64/64 | 387.9 | 675.81 | 168.95 | 2.58/6.19/6.20 | 43.0/54.6/60.4 | 0 | 3.82 | | b42_1k4k_c64 | TP2PP4 | 128/128 | 1088.63 | 481.6 | 120.4 | 80.74/337.92/338.47 | 91.2/95.5/96.3 | 0 | None | +| b42_1k4k_c64 | TP8+EAGLE3+AR | 128/128 | 634.35 | 826.5 | 206.62 | 4.36/12.18/12.50 | 69.9/85.1/101.5 | 0 | 3.878 | | b51_64k_c1 | TP2PP4 | 8/8 | 285.91 | 14.33 | 1833.72 | 10.28/10.32/10.32 | 49.8/50.0/50.0 | 0 | None | | b51_64k_c1 | TP8+EAGLE3+AR | 8/8 | 182.23 | 22.48 | 2877.06 | 17.04/17.17/17.17 | 11.2/13.5/13.5 | 0 | 2.468 | | b51_64k_c4 | TP2PP4 | 8/8 | 122.3 | 33.49 | 4286.94 | 19.55/28.83/28.83 | 81.4/99.5/99.5 | 0 | None | @@ -277,4 +282,4 @@ python3 /root/bench_corpus_v2.py --input-len 262144 --output-len 1 --shared-frac | b52_512k_c1 | TP2PP4 | 1/1 | 92.6 | 0.01 | 5662.03 | 92.60/92.60/92.60 | — | 0 | None | | b52_896k_c1 | TP2PP4 | 1/1 | 220.91 | 0.0 | 4153.24 | 220.91/220.91/220.91 | — | 0 | None | -注:E7b 4.2 C=64 用户中止(rc=143)无 SUMMARY,不在表内;E7b 256K/512K/896K C>1 为结构性跳过;E7b 256K C=1 为全新文本干净重测值(窗口 9,900,000)。 +注:E7b 的 16K/4.1/4.2 场景 C=8~C=64 共 10 点为调参后(MRR64/decode 图 1–64)复测值,即正文采用值;C=1 各点与 5.1/5.2 行沿用初测值(活跃数受池上限约束,与调参无关)。E7b 256K/512K/896K C>1 为结构性跳过;256K C=1 为全新文本干净重测值(窗口 9,900,000)。初测(MRR16/图 1–8)全量原始数据留档于 `results/e7b/`。 diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/all_results.jsonl b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/all_results.jsonl new file mode 100644 index 0000000..c4cfbd9 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/all_results.jsonl @@ -0,0 +1,10 @@ +{"tag": "b3_16k_c8", "arm": "e7b64", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9602, "corpus_window": {"start": 2300000, "end": 2562144}, "ok": 16, "failed": 0, "wall_s": 97.66, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 8192, "output_throughput_tok_s": 83.88, "input_throughput_tok_s": 2684.26, "ttft_s": {"mean": 12.1187, "p50": 8.3735, "p95": 32.4651, "max": 32.4651, "min": 3.9862}, "tpot_s": {"mean": 0.0692, "p50": 0.0688, "p95": 0.1361, "max": 0.1361, "min": 0.0264}, "e2e_s": {"mean": 47.483, "p50": 50.5394, "p95": 78.7413, "max": 78.7413, "min": 18.1949}, "per_req_out_tok_s_e2e": {"mean": 13.2102, "p50": 11.2272, "p95": 28.1397, "max": 28.1397, "min": 6.5023}, "per_req_decode_tok_s": {"mean": 18.5895, "p50": 14.9416, "p95": 37.9333, "max": 37.9333, "min": 7.3633}, "spec_accept_length_mean": 2.556, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 262144, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b3_16k_c16", "arm": "e7b64", "summary": {"concurrency": 16, "num_requests": 32, "run_id": 9603, "corpus_window": {"start": 2300000, "end": 2824288}, "ok": 32, "failed": 0, "wall_s": 177.89, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 16384, "output_throughput_tok_s": 92.1, "input_throughput_tok_s": 2947.27, "ttft_s": {"mean": 19.7632, "p50": 7.8805, "p95": 59.91, "max": 63.6881, "min": 4.062}, "tpot_s": {"mean": 0.1297, "p50": 0.134, "p95": 0.2288, "max": 0.2759, "min": 0.0334}, "e2e_s": {"mean": 86.0176, "p50": 83.7193, "p95": 145.8411, "max": 150.1856, "min": 22.7897}, "per_req_out_tok_s_e2e": {"mean": 7.4324, "p50": 6.5209, "p95": 16.0562, "max": 22.4663, "min": 3.4091}, "per_req_decode_tok_s": {"mean": 10.0542, "p50": 7.5804, "p95": 27.3398, "max": 29.9645, "min": 3.6318}, "spec_accept_length_mean": 2.485, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b3_16k_c32", "arm": "e7b64", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9604, "corpus_window": {"start": 2300000, "end": 3348576}, "ok": 64, "failed": 0, "wall_s": 335.79, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 32768, "output_throughput_tok_s": 97.58, "input_throughput_tok_s": 3122.67, "ttft_s": {"mean": 82.6459, "p50": 87.8342, "p95": 140.3273, "max": 154.661, "min": 5.2339}, "tpot_s": {"mean": 0.1381, "p50": 0.1443, "p95": 0.2136, "max": 0.2775, "min": 0.026}, "e2e_s": {"mean": 153.2127, "p50": 153.6215, "p95": 227.6964, "max": 243.4494, "min": 78.5488}, "per_req_out_tok_s_e2e": {"mean": 3.581, "p50": 3.3906, "p95": 5.2619, "max": 6.5182, "min": 2.1031}, "per_req_decode_tok_s": {"mean": 9.3726, "p50": 7.1484, "p95": 24.3778, "max": 38.5186, "min": 3.6109}, "spec_accept_length_mean": 2.777, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b3_16k_c64", "arm": "e7b64", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9605, "corpus_window": {"start": 2300000, "end": 4397152}, "ok": 128, "failed": 0, "wall_s": 657.79, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 65536, "output_throughput_tok_s": 99.63, "input_throughput_tok_s": 3188.2, "ttft_s": {"mean": 204.2735, "p50": 244.6246, "p95": 293.2111, "max": 322.84, "min": 5.231}, "tpot_s": {"mean": 0.1425, "p50": 0.1465, "p95": 0.2146, "max": 0.2766, "min": 0.0297}, "e2e_s": {"mean": 277.1081, "p50": 309.1898, "p95": 368.9412, "max": 410.789, "min": 77.0243}, "per_req_out_tok_s_e2e": {"mean": 2.133, "p50": 1.6583, "p95": 4.7327, "max": 6.6473, "min": 1.2464}, "per_req_decode_tok_s": {"mean": 8.0518, "p50": 6.8676, "p95": 16.3762, "max": 33.7377, "min": 3.6219}, "spec_accept_length_mean": 2.954, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b41_1k_c8", "arm": "e7b64", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9612, "corpus_window": {"start": 4500000, "end": 4516384}, "ok": 16, "failed": 0, "wall_s": 13.53, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 2048, "output_throughput_tok_s": 151.33, "input_throughput_tok_s": 1210.61, "ttft_s": {"mean": 1.2544, "p50": 1.7591, "p95": 2.1654, "max": 2.1654, "min": 0.3904}, "tpot_s": {"mean": 0.0421, "p50": 0.041, "p95": 0.0537, "max": 0.0537, "min": 0.0336}, "e2e_s": {"mean": 6.6071, "p50": 6.6204, "p95": 8.6618, "max": 8.6618, "min": 4.8636}, "per_req_out_tok_s_e2e": {"mean": 19.981, "p50": 19.5317, "p95": 26.3178, "max": 26.3178, "min": 14.7775}, "per_req_decode_tok_s": {"mean": 24.481, "p50": 24.6734, "p95": 29.9848, "max": 29.9848, "min": 18.7591}, "spec_accept_length_mean": 1.992, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 9, "new_tokens": 16384, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b41_1k_c32", "arm": "e7b64", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9613, "corpus_window": {"start": 4500000, "end": 4565536}, "ok": 64, "failed": 0, "wall_s": 35.75, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 8192, "output_throughput_tok_s": 229.18, "input_throughput_tok_s": 1833.42, "ttft_s": {"mean": 2.9919, "p50": 2.2946, "p95": 7.0943, "max": 7.1032, "min": 0.5556}, "tpot_s": {"mean": 0.1119, "p50": 0.1092, "p95": 0.1506, "max": 0.1798, "min": 0.0728}, "e2e_s": {"mean": 17.1982, "p50": 16.6865, "p95": 23.6063, "max": 25.9273, "min": 9.7993}, "per_req_out_tok_s_e2e": {"mean": 7.8468, "p50": 7.6792, "p95": 11.3781, "max": 13.0622, "min": 4.9369}, "per_req_decode_tok_s": {"mean": 9.4487, "p50": 9.2601, "p95": 12.987, "max": 13.852, "min": 5.6066}, "spec_accept_length_mean": 2.035, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 20, "new_tokens": 65536, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b41_1k_c64", "arm": "e7b64", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9614, "corpus_window": {"start": 4500000, "end": 4631072}, "ok": 128, "failed": 0, "wall_s": 60.18, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 16384, "output_throughput_tok_s": 272.24, "input_throughput_tok_s": 2177.88, "ttft_s": {"mean": 4.7575, "p50": 2.3032, "p95": 12.1513, "max": 13.4791, "min": 0.6534}, "tpot_s": {"mean": 0.1911, "p50": 0.1882, "p95": 0.2927, "max": 0.3315, "min": 0.1063}, "e2e_s": {"mean": 29.0341, "p50": 28.6144, "p95": 42.031, "max": 44.989, "min": 14.7114}, "per_req_out_tok_s_e2e": {"mean": 4.7331, "p50": 4.4822, "p95": 7.5631, "max": 8.7007, "min": 2.8451}, "per_req_decode_tok_s": {"mean": 5.6253, "p50": 5.3567, "p95": 8.1286, "max": 9.481, "min": 3.0404}, "spec_accept_length_mean": 2.094, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 27, "new_tokens": 131072, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b42_1k4k_c8", "arm": "e7b64", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9622, "corpus_window": {"start": 4700000, "end": 4716384}, "ok": 16, "failed": 0, "wall_s": 162.84, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 65536, "output_throughput_tok_s": 402.47, "input_throughput_tok_s": 100.62, "ttft_s": {"mean": 0.9469, "p50": 1.3161, "p95": 1.6549, "max": 1.6549, "min": 0.3801}, "tpot_s": {"mean": 0.0185, "p50": 0.0181, "p95": 0.0243, "max": 0.0243, "min": 0.0149}, "e2e_s": {"mean": 76.5237, "p50": 75.395, "p95": 101.2464, "max": 101.2464, "min": 61.5722}, "per_req_out_tok_s_e2e": {"mean": 54.4322, "p50": 55.1122, "p95": 66.5236, "max": 66.5236, "min": 40.4557}, "per_req_decode_tok_s": {"mean": 55.085, "p50": 56.3616, "p95": 66.9683, "max": 66.9683, "min": 41.1259}, "spec_accept_length_mean": 3.951, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 10, "new_tokens": 16384, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b42_1k4k_c32", "arm": "e7b64", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9623, "corpus_window": {"start": 4700000, "end": 4765536}, "ok": 64, "failed": 0, "wall_s": 387.9, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 262144, "output_throughput_tok_s": 675.81, "input_throughput_tok_s": 168.95, "ttft_s": {"mean": 2.5762, "p50": 1.5202, "p95": 6.1895, "max": 6.2023, "min": 0.4797}, "tpot_s": {"mean": 0.043, "p50": 0.042, "p95": 0.0546, "max": 0.0604, "min": 0.0317}, "e2e_s": {"mean": 178.7057, "p50": 174.8683, "p95": 227.0262, "max": 253.2517, "min": 130.3185}, "per_req_out_tok_s_e2e": {"mean": 23.3694, "p50": 23.4457, "p95": 28.0936, "max": 31.4307, "min": 16.1736}, "per_req_decode_tok_s": {"mean": 23.6957, "p50": 23.8798, "p95": 28.3919, "max": 31.568, "min": 16.5524}, "spec_accept_length_mean": 3.82, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 36, "new_tokens": 65536, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b42_1k4k_c64", "arm": "e7b64", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9624, "corpus_window": {"start": 4700000, "end": 4831072}, "ok": 128, "failed": 0, "wall_s": 634.35, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 524288, "output_throughput_tok_s": 826.5, "input_throughput_tok_s": 206.62, "ttft_s": {"mean": 4.359, "p50": 2.4011, "p95": 12.1834, "max": 12.4977, "min": 0.536}, "tpot_s": {"mean": 0.0699, "p50": 0.0687, "p95": 0.0851, "max": 0.1015, "min": 0.0521}, "e2e_s": {"mean": 290.7916, "p50": 283.6317, "p95": 358.4567, "max": 420.1138, "min": 213.9936}, "per_req_out_tok_s_e2e": {"mean": 14.2904, "p50": 14.4588, "p95": 16.8051, "max": 19.1408, "min": 9.7497}, "per_req_decode_tok_s": {"mean": 14.5029, "p50": 14.5947, "p95": 16.9633, "max": 19.2102, "min": 9.8594}, "spec_accept_length_mean": 3.878, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 69, "new_tokens": 131072, "cached_tokens": 0, "hit_rate": 0.0}}} diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/md5_assets.txt b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/md5_assets.txt new file mode 100644 index 0000000..c154f14 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/md5_assets.txt @@ -0,0 +1,3 @@ +a8f3909211adb92501cf007ab71d3b29 /root/corpus_ids.json +1e34dd8d813ccb7ac8243e8d1573a49e /root/bench_corpus_v2.py +4c126d067d33b5ea27c268f37561634c /root/extract_summary.py diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/retest_compare.md b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/retest_compare.md new file mode 100644 index 0000000..6d03792 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/retest_compare.md @@ -0,0 +1,27 @@ +| 场景 | 并发 | Output TPS 旧→新 | Δ | TPOT P95 旧→新 | TTFT P95 旧→新 | accept 旧→新 | +|---|---|---|---|---|---|---| +| 主场景 16K→512 | 8 | 81.4 → **83.9** | +3% | 135.1 ms → 136.1 ms | 32.52 s → 32.47 s | 2.375 → 2.556 | +| 主场景 16K→512 | 16 | 69.3 → **92.1** | +33% | 297.8 ms → 228.8 ms | 60.86 s → 59.91 s | 2.517 → 2.485 | +| 主场景 16K→512 | 32 | 79.0 → **97.6** | +24% | 259.9 ms → 213.6 ms | 169.68 s → 140.33 s | 2.767 → 2.777 | +| 主场景 16K→512 | 64 | 84.3 → **99.6** | +18% | 231.5 ms → 214.6 ms | 345.69 s → 293.21 s | 2.986 → 2.954 | +| 4.1 短输入 1K→128 | 8 | 152 → **151** | -1% | 53.7 ms → 53.7 ms | 2.16 s → 2.17 s | 2.048 → 1.992 | +| 4.1 短输入 1K→128 | 32 | 100 → **229** | +129% | 190.5 ms → 150.6 ms | 25.60 s → 7.09 s | 2.054 → 2.035 | +| 4.1 短输入 1K→128 | 64 | 103 → **272** | +164% | 177.7 ms → 292.7 ms | 65.81 s → 12.15 s | 2.087 → 2.094 | +| 4.2 长输出 1K→4K | 8 | 412 → **402** | -2% | 22.5 ms → 24.3 ms | 1.68 s → 1.65 s | 4.003 → 3.951 | +| 4.2 长输出 1K→4K | 32 | 252 → **676** | +168% | 74.6 ms → 54.6 ms | 304.75 s → 6.19 s | 3.852 → 3.82 | +| 4.2 长输出 1K→4K | 64 | (无旧数据) | | 826 (新) | 85.1 ms | 3.878 | + +## 附录:E7b64 复测全量指标 + +| 场景点 | ok/nreq | wall(s) | out tok/s | in tok/s | TTFT mean/p95/max(s) | TPOT mean/p95/max(ms) | retractions | accept | hit | +|---|---|---|---|---|---|---|---|---|---| +| b3_16k_c8 | 16/16 | 97.66 | 83.88 | 2684.26 | 12.12/32.47/32.47 | 69.2/136.1/136.1 | 0 | 2.556 | 0.0 | +| b3_16k_c16 | 32/32 | 177.89 | 92.1 | 2947.27 | 19.76/59.91/63.69 | 129.7/228.8/275.9 | 0 | 2.485 | 0.0 | +| b3_16k_c32 | 64/64 | 335.79 | 97.58 | 3122.67 | 82.65/140.33/154.66 | 138.1/213.6/277.5 | 0 | 2.777 | 0.0 | +| b3_16k_c64 | 128/128 | 657.79 | 99.63 | 3188.2 | 204.27/293.21/322.84 | 142.5/214.6/276.6 | 0 | 2.954 | 0.0 | +| b41_1k_c8 | 16/16 | 13.53 | 151.33 | 1210.61 | 1.25/2.17/2.17 | 42.1/53.7/53.7 | 0 | 1.992 | 0.0 | +| b41_1k_c32 | 64/64 | 35.75 | 229.18 | 1833.42 | 2.99/7.09/7.10 | 111.9/150.6/179.8 | 0 | 2.035 | 0.0 | +| b41_1k_c64 | 128/128 | 60.18 | 272.24 | 2177.88 | 4.76/12.15/13.48 | 191.1/292.7/331.5 | 0 | 2.094 | 0.0 | +| b42_1k4k_c8 | 16/16 | 162.84 | 402.47 | 100.62 | 0.95/1.65/1.65 | 18.5/24.3/24.3 | 0 | 3.951 | 0.0 | +| b42_1k4k_c32 | 64/64 | 387.9 | 675.81 | 168.95 | 2.58/6.19/6.20 | 43.0/54.6/60.4 | 0 | 3.82 | 0.0 | +| b42_1k4k_c64 | 128/128 | 634.35 | 826.5 | 206.62 | 4.36/12.18/12.50 | 69.9/85.1/101.5 | 0 | 3.878 | 0.0 | diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/server_facts.txt b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/server_facts.txt new file mode 100644 index 0000000..5201e9c --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/server_facts.txt @@ -0,0 +1,19 @@ +[2026-09-10 09:22:59] server_args={'model_path': '/data/hf_models/GLM-5.3-NVFP4', 'tokenizer_path': '/data/hf_models/GLM-5.3-NVFP4', 'tokenizer_mode': 'auto', 'tokenizer_backend': 'huggingface', 'tokenizer_worker_num': 1, 'detokenizer_worker_num': 1, 'skip_tokenizer_init': False, 'load_format': 'auto', 'model_loader_extra_config': '{}', 'trust_remote_code': False, 'context_length': 270336, 'is_embedding': False, 'enable_multimodal': None, 'revision': None, 'model_impl': 'auto', 'model_config_parser': 'auto', 'json_model_override_args': '{}', 'dtype': 'auto', 'quantization': None, 'quantization_param_path': None, 'kv_cache_dtype': 'fp8_e4m3', 'enable_fp32_lm_head': False, 'modelopt_quant': None, 'modelopt_checkpoint_restore_path': None, 'modelopt_checkpoint_save_path': None, 'modelopt_export_path': None, 'quantize_and_serve': False, 'rl_quant_profile': None, 'enable_tf32_matmul': False, 'mem_fraction_static': 0.9, 'max_running_requests': 64, 'max_queued_requests': None, 'max_total_tokens': None, 'chunked_prefill_size': 8192, 'prefill_decode_interval': 0, 'enable_dynamic_chunking': False, 'max_prefill_tokens': 16384, 'prefill_max_requests': None, 'schedule_policy': 'fcfs', 'enable_priority_scheduling': False, 'disable_priority_preemption': False, 'default_priority_value': None, 'abort_on_priority_when_disabled': False, 'schedule_low_priority_values_first': False, 'priority_scheduling_preemption_threshold': 10, 'retraction_policy': 'length', 'schedule_conservativeness': 1.0, 'page_size': 64, 'c128_page_size': 16, 'swa_full_tokens_ratio': 0.8, 'disable_hybrid_swa_memory': False, 'radix_eviction_policy': 'lru', 'prefill_only_disable_kv_cache': False, 'disable_radix_cache': False, 'enable_page_major_kv_layout': False, 'enable_unified_memory': False, 'disable_chunked_prefix_cache': False, 'disable_overlap_schedule': False, 'num_continuous_decode_steps': 1, 'scheduler_recv_interval': 1, 'enable_mixed_chunk': False, 'nccl_port': None, 'dist_timeout': None, 'dist_init_addr': None, 'gated_launch_port': None, 'nnodes': 1, 'node_rank': 0, 'tp_size': 8, 'dcp_size': 1, 'pp_size': 1, 'pp_max_micro_batch_size': None, 'pp_async_batch_depth': 0, 'dp_size': 1, 'load_balance_method': 'round_robin', 'attn_cp_size': 1, 'moe_dp_size': 1, 'dwdp_size': 1, 'dcp_comm_backend': 'ag_rs', 'dcp_replicate_q_proj': None, 'enable_prefill_cp': False, 'cp_strategy': None, 'enable_dsa_cache_layer_split': False, 'enable_dsa_prefill_context_parallel': False, 'dsa_prefill_cp_mode': 'round-robin-split', 'enable_prefill_context_parallel': False, 'prefill_cp_mode': 'in-seq-split', 'enable_cp_decode_attn_tp': False, 'enable_dp_attention': False, 'enable_dp_attention_local_control_broadcast': False, 'enable_dp_lm_head': False, 'enable_tp_lm_head_all_to_all': False, 'enable_attn_tp_input_scattered': False, 'enable_shared_experts_attn_tp': False, 'enable_dense_mlp_attn_tp': False, 'disable_attn_tp_gather': False, 'enable_p2p_check': False, 'device': 'cuda', 'base_gpu_id': 0, 'gpu_id_step': 1, 'random_seed': 825816159, 'mlx_enable_sampling': False, 'watchdog_timeout': 300, 'soft_watchdog_timeout': None, 'sleep_on_idle': False, 'use_ray': False, 'custom_sigquit_handler': None, 'numa_node': None, 'gc_threshold': None, 'host': '0.0.0.0', 'port': 30000, 'fastapi_root_path': '', 'smg_grpc_mode': False, 'grpc_mode': False, 'grpc_port': None, 'grpc_worker_threads': 4, 'sidecar': None, 'sidecar_args': None, 'skip_server_warmup': False, 'warmups': None, 'enable_http2': False, 'http2_max_concurrent_streams': 200, 'ssl_keyfile': None, 'ssl_certfile': None, 'ssl_ca_certs': None, 'ssl_keyfile_password': None, 'enable_ssl_refresh': False, 'api_key': None, 'admin_api_key': None, 'served_model_name': '/data/hf_models/GLM-5.3-NVFP4', 'weight_version': 'default', 'chat_template': None, 'hf_chat_template_name': None, 'completion_template': None, 'file_storage_path': 'sglang_storage', 'enable_cache_report': False, 'reasoning_parser': 'glm45', 'default_chat_template_kwargs': None, 'strip_thinking_cache': False, 'enable_strict_thinking': False, 'tool_call_parser': 'glm47', 'tool_server': None, 'sampling_defaults': 'model', 'asr_max_buffer_seconds': 60, 'asr_max_concurrent_sessions': 32, 'preferred_sampling_params': None, 'allow_auto_truncate': False, 'stream_interval': 1, 'batch_notify_size': 16, 'stream_response_default_include_usage': False, 'incremental_streaming_output': False, 'enable_streaming_session': False, 'enable_session_radix_cache': False, 'log_level': 'info', 'log_level_http': None, 'log_requests': False, 'log_requests_level': 2, 'log_requests_format': 'text', 'log_requests_target': None, 'uvicorn_access_log_exclude_prefixes': [], 'crash_dump_folder': None, 'show_time_cost': False, 'enable_metrics': False, 'smg_http_sidecar_port': None, 'enable_mfu_metrics': False, 'enable_metrics_for_all_schedulers': False, 'load_snapshot_publish_interval': 15, 'tokenizer_metrics_custom_labels_header': 'x-custom-labels', 'tokenizer_metrics_allowed_custom_labels': None, 'extra_metric_labels': None, 'bucket_time_to_first_token': None, 'bucket_inter_token_latency': None, 'bucket_e2e_request_latency': None, 'prompt_tokens_buckets': None, 'generation_tokens_buckets': None, 'gc_warning_threshold_secs': 0.0, 'decode_log_interval': 40, 'enable_request_time_stats_logging': False, 'kv_events_config': None, 'load_publish_endpoint': None, 'enable_forward_pass_metrics': False, 'forward_pass_metrics_worker_id': '', 'forward_pass_metrics_ipc_name': None, 'enable_trace': False, 'trace_modules': 'request', 'otlp_traces_endpoint': 'localhost:4317', 'export_metrics_to_file': False, 'export_metrics_to_file_dir': None, 'stat_loggers': None, 'constrained_json_whitespace_pattern': None, 'constrained_json_disable_any_whitespace': False, 'attention_backend': 'dsa', 'decode_attention_backend': None, 'prefill_attention_backend': None, 'sampling_backend': 'flashinfer', 'grammar_backend': 'xgrammar', 'radix_cache_backend': None, 'mm_attention_backend': None, 'fp8_gemm_runner_backend': 'auto', 'fp4_gemm_runner_backend': 'auto', 'bf16_gemm_backend': 'auto', 'dsa_prefill_backend': 'flashinfer_sparse_mla', 'dsv4_prefill_backend': 'auto', 'dsa_decode_backend': 'flashinfer_sparse_mla', 'dsa_paged_mqa_logits_backend': 'auto', 'dsa_topk_backend': 'sgl-kernel', 'disable_flashinfer_autotune': True, 'flashinfer_autotune_skip_ops': None, 'mamba_backend': 'triton', 'cuda_graph_config': {'decode': {'backend': 'full', 'max_bs': 64, 'bs': [1, 2, 3, 4, 6, 8, 12, 16, 24, 32, 48, 64], 'tc_compiler': 'eager', 'full_prefill_max_req': None, 'full_prefill_prefix_chunk_tokens': None}, 'prefill': {'backend': 'breakable', 'max_bs': 8, 'bs': [4, 8], 'tc_compiler': 'eager', 'full_prefill_max_req': None, 'full_prefill_prefix_chunk_tokens': None}}, 'cuda_graph_backend_decode': None, 'cuda_graph_backend_prefill': None, 'cuda_graph_max_bs_decode': 64, 'cuda_graph_max_bs_prefill': 8, 'cuda_graph_bs_decode': [1, 2, 3, 4, 6, 8, 12, 16, 24, 32, 48, 64], 'cuda_graph_bs_prefill': None, 'cuda_graph_tc_compiler': None, 'disable_prefill_cuda_graph': False, 'disable_decode_cuda_graph': False, 'disable_cuda_graph': False, 'disable_cuda_graph_padding': False, 'enable_profile_cuda_graph': False, 'enable_cudagraph_gc': False, 'debug_cuda_graph': False, 'enable_layerwise_nvtx_marker': False, 'enable_nccl_nvls': False, 'enable_symm_mem': False, 'triton_attention_reduce_in_fp32': False, 'triton_attention_num_kv_splits': 8, 'triton_attention_split_tile_size': None, 'flashinfer_mla_disable_ragged': False, 'enable_fused_qk_norm_rope': False, 'enable_precise_embedding_interpolation': False, 'enable_fused_moe_sum_all_reduce': False, 'enable_deepseek_v4_fp4_indexer': False, 'disable_custom_all_reduce': False, 'enable_mscclpp': False, 'enable_torch_symm_mem': False, 'enable_scattered_sconv': False, 'pre_warm_nccl': False, 'enable_quant_communications': False, 'enable_flashinfer_allreduce_fusion': False, 'enforce_disable_flashinfer_allreduce_fusion': False, 'flashinfer_allreduce_fusion_backend': None, 'enable_aiter_allreduce_fusion': False, 'enable_torch_compile': False, 'enable_torch_compile_debug_mode': False, 'torch_compile_max_bs': 32, 'speculative_algorithm': 'EAGLE', 'speculative_draft_model_path': '/data/hf_models/GLM-5.3-NVFP4', 'speculative_draft_model_revision': None, 'speculative_draft_load_format': None, 'speculative_num_steps': 4, 'speculative_eagle_topk': 1, 'speculative_num_draft_tokens': 5, 'speculative_dflash_block_size': None, 'speculative_dspark_block_size': None, 'speculative_dspark_sps_table_path': None, 'speculative_dspark_confidence_sts_path': None, 'speculative_dspark_align_verify_tokens_to_graph_tier': False, 'speculative_accept_threshold_single': 1.0, 'speculative_accept_threshold_acc': 1.0, 'speculative_use_rejection_sampling': False, 'speculative_token_map': None, 'speculative_attention_mode': 'prefill', 'speculative_draft_attention_backend': None, 'speculative_dsa_topk_backend': 'sgl-kernel', 'speculative_draft_kv_cache_dtype': None, 'speculative_draft_window_size': None, 'speculative_moe_runner_backend': 'flashinfer_cutlass', 'speculative_moe_a2a_backend': None, 'speculative_draft_model_quantization': None, '_speculative_draft_quantization_explicitly_set': False, 'speculative_skip_dp_mlp_sync': False, 'enable_multi_layer_eagle': False, 'speculative_adaptive': False, 'speculative_adaptive_config': None, 'decoupled_spec_bind_endpoint': None, 'decoupled_spec_connect_endpoints': None, 'decoupled_spec_rank': None, 'decoupled_spec_role': 'null', 'spec_trace_dir': None, 'speculative_ngram_min_bfs_breadth': 1, 'speculative_ngram_max_bfs_breadth': 10, 'speculative_ngram_match_type': 'BFS', 'speculative_ngram_max_trie_depth': 18, 'speculative_ngram_capacity': 10000000, 'speculative_ngram_external_corpus_path': None, 'speculative_ngram_external_sam_budget': 0, 'speculative_ngram_external_corpus_max_tokens': 10000000, 'ep_size': 1, 'moe_a2a_backend': 'none', 'enable_w4a4_mxfp4_megamoe': False, 'deepep_v2_mode': 'direct', 'moe_runner_backend': 'flashinfer_cutlass', 'flashinfer_mxfp4_moe_precision': 'default', 'deepep_mode': 'auto', 'fuseep_mode': 2, 'deepep_dispatcher_output_dtype': 'auto', 'ep_num_redundant_experts': 0, 'ep_dispatch_algorithm': None, 'init_expert_location': 'trivial', 'enable_eplb': False, 'eplb_algorithm': 'auto', 'eplb_rebalance_num_iterations': 1000, 'eplb_rebalance_layers_per_chunk': None, 'eplb_min_rebalancing_utilization_threshold': 1.0, 'expert_distribution_recorder_mode': None, 'expert_distribution_recorder_buffer_size': 1000, 'expert_balancedness_report_mode': 'off', 'deepep_config': None, 'moe_dense_tp_size': None, 'elastic_ep_backend': None, 'enable_elastic_expert_backup': False, 'mooncake_ib_device': None, 'enable_waterfill': False, 'ep_join_mode': None, 'ep_join_rank_offset': 0, 'elastic_ep_initial_size': None, 'max_ep_size': None, 'elastic_ep_scale_timeout': 600, 'elastic_ep_rejoin': False, 'disable_flashinfer_cutlass_moe_fp4_allgather': False, 'disable_shared_experts_fusion': True, 'enforce_shared_experts_fusion': False, 'max_mamba_cache_size': None, 'mamba_ssm_dtype': None, 'mamba_max_states_per_path': -1, 'enable_mamba_cache_stochastic_rounding': False, 'mamba_cache_philox_rounds': 0, 'mamba_full_memory_ratio': 0.9, 'mamba_radix_cache_strategy': 'auto', 'uses_mamba_radix_cache': False, 'mamba_track_interval': 256, 'enable_int8_mamba_checkpoint': False, 'int8_mamba_ckpt_size': None, 'linear_attn_backend': 'triton', 'linear_attn_decode_backend': None, 'linear_attn_prefill_backend': None, 'linear_attn_verify_backend': None, 'enable_linear_replayssm': False, 'linear_replayssm_cache_len': 16, 'enable_linear_replayssm_spec': False, 'enable_hierarchical_cache': True, 'hicache_host_memory_mode': 'cache', 'hicache_ratio': 3.0, 'hicache_size': 0, 'hicache_write_policy': 'write_through', 'hicache_io_backend': 'kernel', 'hicache_mem_layout': 'page_first', 'hicache_storage_backend': None, 'hicache_storage_prefetch_policy': 'timeout', 'hicache_storage_backend_extra_config': None, 'enable_hisparse': False, 'hisparse_config': None, 'enable_broadcast_mm_inputs_process': False, 'enable_prefix_mm_cache': False, 'mm_enable_dp_encoder': False, 'mm_process_config': {}, 'mm_processor_worker_num': 0, 'mm_io_worker_num': 0, 'allowed_media_domains': [], 'media_url_max_file_size_mb': 64, 'mm_preprocess_cache_size_mb': None, 'trust_mm_content_hashes': False, 'limit_mm_data_per_request': None, 'enable_mm_global_cache': False, 'image_processor_backend': 'auto', 'mm_global_cache_backend': 'mooncake', 'disable_fast_image_processor': False, 'mm_feature_transport': 'cpu', 'keep_mm_feature_on_device': False, 'enable_lora': None, 'enable_lora_overlap_loading': None, 'max_lora_rank': None, 'lora_target_modules': None, 'lora_paths': None, 'max_loaded_loras': None, 'max_loras_per_batch': 8, 'lora_eviction_policy': 'lru', 'lora_backend': 'csgmv', 'max_lora_chunk_size': 16, 'experts_shared_outer_loras': None, 'lora_use_virtual_experts': False, 'lora_strict_loading': False, 'lora_drain_wait_threshold': 0.0, 'enable_two_batch_overlap': False, 'enable_single_batch_overlap': False, 'tbo_token_distribution_threshold': 0.48, 'cpu_offload_gb': 0, 'offload_group_size': -1, 'offload_num_in_group': 1, 'offload_prefetch_step': 1, 'offload_mode': 'cpu', 'enable_lmcache': False, 'lmcache_config_file': None, 'enable_flexkv': False, 'flexkv_config_file': None, 'kt_weight_path': None, 'kt_method': 'AMXINT4', 'kt_cpuinfer': None, 'kt_threadpool_count': 2, 'kt_num_gpu_experts': None, 'kt_max_deferred_experts_per_token': None, 'dllm_algorithm': None, 'dllm_algorithm_config': None, 'dllm_fdfo': True, 'disaggregation_mode': 'null', 'disaggregation_transfer_backend': 'mooncake', 'disaggregation_bootstrap_port': 8998, 'disaggregation_ib_device': None, 'disaggregation_decode_enable_radix_cache': False, 'disaggregation_decode_enable_offload_kvcache': False, 'disaggregation_decode_retraction_backup': None, 'num_reserved_decode_tokens': 512, 'disaggregation_decode_extra_slots': None, 'disaggregation_decode_polling_interval': 1, 'optimistic_prefill_attempts': 0, 'encoder_only': False, 'language_only': False, 'language_model_only': False, 'encoder_transfer_backend': 'zmq_to_scheduler', 'encoder_urls': [], 'encoder_bootstrap_port': 8997, 'encoder_register_urls': [], 'enable_adaptive_dispatch_to_encoder': False, 'enable_pdmux': False, 'pdmux_config_path': None, 'sm_group_num': 8, 'startup_weight_load_mode': 'serial', 'custom_weight_loader': [], 'weight_loader_disable_mmap': False, 'weight_loader_prefetch_checkpoints': False, 'weight_loader_prefetch_num_threads': 4, 'weight_loader_drop_cache_after_load': False, 'remote_instance_weight_loader_seed_instance_ip': None, 'remote_instance_weight_loader_seed_instance_service_port': None, 'remote_instance_weight_loader_send_weights_group_ports': None, 'remote_instance_weight_loader_backend': 'nccl', 'remote_instance_weight_loader_start_seed_via_transfer_engine': False, 'engine_info_bootstrap_port': 6789, 'modelexpress_config': None, 'download_dir': None, 'model_checksum': None, 'delete_ckpt_after_loading': False, 'decrypted_config_file': None, 'decrypted_draft_config_file': None, 'checkpoint_engine_wait_weights_before_ready': False, 'enable_prefill_delayer': False, 'prefill_delayer_max_delay_passes': 30, 'prefill_delayer_token_usage_low_watermark': None, 'prefill_delayer_forward_passes_buckets': None, 'prefill_delayer_wait_seconds_buckets': None, 'prefill_delayer_queue_min_ratio': None, 'prefill_delayer_max_delay_ms': None, 'min_free_slots_delay': None, 'enable_deterministic_inference': False, 'rl_on_policy_target': None, 'kv_canary': 'none', 'kv_canary_real_data': 'none', 'kv_canary_sweep_interval': 0, 'enable_dynamic_batch_tokenizer': False, 'dynamic_batch_tokenizer_batch_size': 32, 'dynamic_batch_tokenizer_batch_timeout': 0.002, 'enable_tokenizer_batch_encode': False, 'disable_tokenizer_batch_decode': False, 'debug_tensor_dump_output_folder': None, 'debug_tensor_dump_layers': None, 'debug_tensor_dump_input_file': None, 'enable_memory_saver': False, 'enable_weights_cpu_backup': False, 'enable_draft_weights_cpu_backup': False, 'enable_custom_logit_processor': False, 'enable_return_hidden_states': False, 'return_hidden_states_mode': None, 'enable_return_routed_experts': False, 'enable_return_indexer_topk': False, 'disable_outlines_disk_cache': False, 'enable_mis': False, 'weight_cache_mode': 'off', 'weight_cache_socket': None, 'weight_cache_timeout': 1800, 'forward_hooks': None, 'msprobe_dump_config': None} +[2026-09-10 09:27:16 TP4] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 09:27:16 TP2] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 09:27:16 TP2] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 09:27:16 TP4] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 09:27:16 TP3] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 09:27:16 TP7] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 09:27:16 TP6] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 09:27:16 TP3] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 09:27:16 TP1] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 09:27:16 TP5] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 09:27:16 TP7] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 09:27:16 TP1] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 09:27:16 TP6] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 09:27:16 TP5] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 09:27:16 TP0] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 09:27:16 TP0] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 09:29:42 TP0] max_total_num_tokens=276480, chunked_prefill_size=8192, max_prefill_tokens=16384, max_running_requests=64, context_len=270336, available_gpu_mem=6.11 GB +[2026-09-10 09:30:27] Engine startup timings (s): load_weight=217.07, kv_cache_allocation=0.42, scheduler_e2e=431.88, cuda_graph={prefill=94.48, decode=0.00, target_verify=35.95, draft_prefill=0.00, draft_decode=13.40, draft_extend=1.85}, tokenizer_e2e=447.94 diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/status.txt b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/status.txt new file mode 100644 index 0000000..c59f6b1 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b64/status.txt @@ -0,0 +1,10 @@ +b3_16k_c8 OK +b3_16k_c16 OK +b3_16k_c32 OK +b3_16k_c64 OK +b41_1k_c8 OK +b41_1k_c32 OK +b41_1k_c64 OK +b42_1k4k_c8 OK +b42_1k4k_c32 OK +b42_1k4k_c64 OK diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/provenance.md b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/provenance.md index f386e15..82927a5 100644 --- a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/provenance.md +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/provenance.md @@ -1,7 +1,7 @@ # 数据来源与核验记录(provenance) 原始 csv/log 按仓库惯例不入库(.gitignore `*.csv`/`*.log`),完整文件在 60.8 -`/root/bench_logs/b300eq_{tp2pp4_20260910_1138,e7b_20260910_1428}/` 与本地镜像 +`/root/bench_logs/b300eq_{tp2pp4_20260910_1138,e7b_20260910_1428,e7b64_20260910_1732}/` 与本地镜像 `D:/sskj/b300eq/`。本文件固化其中的关键事实。 ## GPU 清单(nvidia-smi,8×RTX 6000D,总 85,651 MiB/卡) @@ -27,10 +27,20 @@ vram_timeline.csv(30s 采样)全程峰值 **85,013 MiB**(主场景 C=64, 全程峰值 **83,553 MiB**(余量 ~2.1 GiB)。 +### E7b64 复测臂(MRR64 + decode 图桶 1–64,20260910_1732) + +- 部署:`deploy_glm53_e7b_hicc.sh`(md5 165db732…),与 E7b 初测唯一差异 = MRR 16→64 + 图参数 `--cuda-graph-max-bs-decode 64 --cuda-graph-bs-decode "1 2 3 4 6 8 12 16 24 32 48 64"`;KV 池 276,480 不变,server_args 核验 `max_running_requests=64`、图桶 13 档 +- 启动计时(09:30:27):load_weight=217.07 s;cuda_graph={prefill=94.48, target_verify=35.95, draft_decode=13.40, draft_extend=1.85};捕获后 avail_gpu_mem=6.11 GB +- 空载 79,317~79,411 MiB/卡;复测矩阵(16K/1K 场景,vram_timeline 30s 采样 116 帧)全程峰值 **83,627 MiB**(余量 ~3.9 GiB) +- 10 点全部 OK、命中核验全 0.0(全新容器实例=回收窗口重新处女文本)、0 retraction;run-id 9601-9610 +- C=8 锚点 vs 初测偏差:16K 81.4→83.9(+3%)、1K 152→151(−1%)、1K→4K 412→402(−2%)→ 两轮环境无漂移 +- 前后对比表:`retest_compare.md`(本目录镜像 = 60.8 `/root/bench_logs/retest_compare.md`) + ## 质量门判决 - TP2PP4 臂:`PASS=6 FAIL=1`(唯一失败 = tool-call,D 口径无 parser,历史已知;GSM8K×5 + 中文推理全过) - E7b 臂:`PASS=7 FAIL=0`(含 tool-call `get_weather{"city": "北京"}`) +- E7b64 复测臂:`PASS=7 FAIL=0`(部署后以 "The server is fired up" 真就绪信号判定后跑门,7/7) ## 在役容器保全与恢复(60.8,TP4PP2-nomtp@0.90 口径) @@ -38,7 +48,8 @@ vram_timeline.csv(30s 采样)全程峰值 **85,013 MiB**(主场景 C=64, - 镜像:`lmsysorg/sglang:nightly-dev-cu13-20260901-07c8f729`(sha256:eb090e39…) - 停役前显存:82,221~82,395 MiB/卡 - 恢复流程:E7b 测试容器 rm(显存排干 0 MiB)→ `docker rename glm53-nvfp4-insvc glm53-nvfp4 && docker start` -- 恢复核验(09-10):health 200(启动后 ~4 min);16K/16tok 冷抽测 ok=1/1、wall 3.59 s;显存 GPU4-7 与停役前持平(82,2xx MiB)、GPU0-3 低 ~5 GiB(重启后 radix 池未回填,正常);容器口径未变 +- 恢复核验第一轮(09-10 午,初测后):health 200(启动后 ~4 min);16K/16tok 冷抽测 ok=1/1、wall 3.59 s;显存 GPU4-7 与停役前持平(82,2xx MiB)、GPU0-3 低 ~5 GiB(重启后 radix 池未回填,正常);容器口径未变 +- 恢复核验第二轮(09-10 晚,e7b64 复测拆台后,`restore_insvc_e7b64.sh` 自动化):fired up(start 后 ~4 min)→ health 200 → 16K/32tok 抽测 ok + C=4×16K/256tok 抽测 ok;KV 池 647,040 tokens(8 rank 一致)+ server_args 与归档启动命令逐字一致(TP4PP2/mem0.90/MRR16/cps8192/hicache×3/ctx 1,048,576);显存 GPU4-7 82,2xx MiB 持平、GPU0-3 77.2 GiB = mem0.90 静态预算水位(较停役前热稳态多 ~5 GiB 余量,与第一轮同象) - 僵尸 PID 现象记录:`docker stop`/`rm` 偶发 "container PID xxx is zombie and can not be killed",实为收尾边界现象(容器终态 exited 137、显存归零),等待 ~20s 重试即成功 ## 执行资产 md5(60.8 = 本目录 = 60.7 原件,三方一致) diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/deploy_glm53_e7b_hicc.sh b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/deploy_glm53_e7b_hicc.sh new file mode 100644 index 0000000..3877b03 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/deploy_glm53_e7b_hicc.sh @@ -0,0 +1,69 @@ +#!/bin/bash +# deploy_glm53_e7b_hicc.sh — E7b high-concurrency variant for the 60.8 retest +# Base: deploy_glm53_607_exp.sh (md5 21db641f1d997e26fcf1ddc97163b87e) with ONLY: +# 1) MRR env param (default 64) replaces hardcoded --max-running-requests 16 +# 2) decode CUDA graph bucket extended 1..64 (was "1 2 3 4 6 8", max 8); +# prefill graph max stays 8 (single-variable change vs E7b: MRR + decode graph) +# Everything else identical recipe: image nightly-dev-20260828-daf63171, EAGLE 4/1/5, +# mem0.90, chunk8192, hicache3, ctx270336, parsers, CAR 1stage patch flow. +# Env: MEMFRAC(0.90) STEPS(4) TOPK(1) DRAFT(5) CTXLEN(270336) CHUNK(8192) +# MAXPRE(16384) MRR(64) GMAXD(64) GBSD("1 2 3 4 6 8 12 16 24 32 48 64") +# RESTART(no|yes) EXTRA("") CAR_PATCH(1) +set -e +MEMFRAC=${MEMFRAC:-0.90} +STEPS=${STEPS:-4} +TOPK=${TOPK:-1} +DRAFT=${DRAFT:-5} +CTXLEN=${CTXLEN:-270336} +CHUNK=${CHUNK:-8192} +MAXPRE=${MAXPRE:-16384} +MRR=${MRR:-64} +GMAXD=${GMAXD:-64} +GBSD=${GBSD:-"1 2 3 4 6 8 12 16 24 32 48 64"} +EXTRA=${EXTRA:-} +if [ "$RESTART" = "yes" ]; then RP="--restart unless-stopped"; else RP="--restart no"; fi + +echo "[deploy] removing old container (if any, exact-name match only)" +for i in $(seq 1 45); do + CID=$(docker ps -a --filter name=^/glm53-nvfp4$ -q) + [ -z "$CID" ] && break + docker rm -f glm53-nvfp4 >/dev/null 2>&1 || true + sleep 2 +done + +echo "[deploy] waiting for VRAM drain" +for i in $(seq 1 45); do + used=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | awk '{s+=$1} END {print s}') + [ "$used" -lt 2000 ] && break + sleep 2 +done +used=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | awk '{s+=$1} END {print s}') +echo "[deploy] VRAM now: ${used} MiB total" + +echo "[deploy] starting: TP8 EAGLE ${STEPS}/${TOPK}/${DRAFT} memfrac=${MEMFRAC} mrr=${MRR} decode-graph<=${GMAXD} chunk=${CHUNK} extra='${EXTRA}'" +docker run -d --name glm53-nvfp4 --gpus all --shm-size 64g --ipc=host $RP \ + -p 30000:30000 -v /data/hf_models:/data/hf_models \ + lmsysorg/sglang:nightly-dev-20260828-daf63171 \ + python3 -m sglang.launch_server \ + --model-path /data/hf_models/GLM-5.3-NVFP4 --tp 8 \ + --mem-fraction-static $MEMFRAC --max-running-requests $MRR \ + --chunked-prefill-size $CHUNK --max-prefill-tokens $MAXPRE \ + --disable-shared-experts-fusion --moe-runner-backend flashinfer_cutlass --disable-flashinfer-autotune \ + --speculative-algorithm EAGLE --speculative-num-steps $STEPS --speculative-eagle-topk $TOPK --speculative-num-draft-tokens $DRAFT \ + --kv-cache-dtype fp8_e4m3 --enable-hierarchical-cache --hicache-ratio 3 \ + --cuda-graph-max-bs-decode $GMAXD --cuda-graph-bs-decode $GBSD --cuda-graph-max-bs-prefill 8 \ + --context-length $CTXLEN --reasoning-parser glm45 --tool-call-parser glm47 \ + --host 0.0.0.0 --port 30000 $EXTRA + +echo "[deploy] container started; poll: docker logs -f glm53-nvfp4" + +# --- CAR 1stage patch injection (E7b winner) --- +if [ "${CAR_PATCH:-1}" != "0" ] && [ -f /root/patches/custom_all_reduce.py ]; then + echo "[deploy] CAR_PATCH: injecting custom-AR 1stage patch" + docker stop -t 20 glm53-nvfp4 >/dev/null 2>&1 || true + DPATH=/sgl-workspace/sglang/python/sglang/srt/distributed/device_communicators + docker cp /root/patches/custom_all_reduce.py glm53-nvfp4:$DPATH/custom_all_reduce.py + docker cp /root/patches/custom_all_reduce_utils.py glm53-nvfp4:$DPATH/custom_all_reduce_utils.py + docker start glm53-nvfp4 + echo "[deploy] CAR_PATCH injected, container restarted; poll health as usual" +fi diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/gen_retest_compare.py b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/gen_retest_compare.py new file mode 100644 index 0000000..3f6ae3d --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/gen_retest_compare.py @@ -0,0 +1,108 @@ +#!/usr/bin/env python3 +"""gen_retest_compare.py + +Before/after table for the E7b high-concurrency retest: + E7b (MRR16, decode graph bs 1-8) vs E7b64 (MRR64, decode graph bs 1-64) +Emits B300-style markdown per scenario with Output TPS / TTFT P95 / TPOT P95 +and the delta, plus a full-metrics appendix. Also prints TP2PP4 reference +values from the first file's arm field where available (3-column context). +""" +import json +import sys + +POINTS = [ + ("b3_16k_c8", "主场景 16K→512", 8), + ("b3_16k_c16", "主场景 16K→512", 16), + ("b3_16k_c32", "主场景 16K→512", 32), + ("b3_16k_c64", "主场景 16K→512", 64), + ("b41_1k_c8", "4.1 短输入 1K→128", 8), + ("b41_1k_c32", "4.1 短输入 1K→128", 32), + ("b41_1k_c64", "4.1 短输入 1K→128", 64), + ("b42_1k4k_c8", "4.2 长输出 1K→4K", 8), + ("b42_1k4k_c32", "4.2 长输出 1K→4K", 32), + ("b42_1k4k_c64", "4.2 长输出 1K→4K", 64), +] + +def load(path, arm): + rows = {} + with open(path, encoding="utf-8") as f: + for line in f: + rec = json.loads(line) + if rec.get("arm") == arm: + rows[rec["tag"]] = rec["summary"] # later records override + return rows + +def g(s, path): + cur = s + for key in path.split("."): + if cur is None: + return None + cur = cur.get(key) + return cur + +def fmt_tps(v): + if v is None: + return "—" + return f"{v:,.0f}" if v >= 100 else f"{v:,.1f}" + +def fmt_s(v): + if v is None: + return "—" + return f"{v:,.2f} s" if v >= 1 else f"{v*1000:,.0f} ms" + +def fmt_ms(v): + return "—" if v is None else f"{v*1000:,.1f} ms" + +def ratio(new, old): + if new is None or old in (None, 0): + return "" + d = (new - old) / old * 100 + return f"{d:+.0f}%" + +def main(): + old = load(sys.argv[1], "e7b") + new = load(sys.argv[2], "e7b64") + print("| 场景 | 并发 | Output TPS 旧→新 | Δ | TPOT P95 旧→新 | TTFT P95 旧→新 | accept 旧→新 |") + print("|---|---|---|---|---|---|---|") + for tag, scene, cc in POINTS: + o, n = old.get(tag), new.get(tag) + if n is None and o is None: + continue + cells = [scene, str(cc)] + if o is None: + cells += ["(无旧数据)", "", fmt_tps(g(n, "output_throughput_tok_s")) + " (新)", fmt_ms(g(n, "tpot_s.p95")), str(g(n, "spec_accept_length_mean") or "—")] + elif n is None: + cells += [fmt_tps(g(o, "output_throughput_tok_s")) + " (旧)", "", "(无新数据)", "", ""] + else: + cells += [ + f"{fmt_tps(g(o,'output_throughput_tok_s'))} → **{fmt_tps(g(n,'output_throughput_tok_s'))}**", + ratio(g(n, "output_throughput_tok_s"), g(o, "output_throughput_tok_s")), + f"{fmt_ms(g(o,'tpot_s.p95'))} → {fmt_ms(g(n,'tpot_s.p95'))}", + f"{fmt_s(g(o,'ttft_s.p95'))} → {fmt_s(g(n,'ttft_s.p95'))}", + f"{g(o,'spec_accept_length_mean') or '—'} → {g(n,'spec_accept_length_mean') or '—'}", + ] + print("| " + " | ".join(cells) + " |") + + print("\n## 附录:E7b64 复测全量指标\n") + print("| 场景点 | ok/nreq | wall(s) | out tok/s | in tok/s | TTFT mean/p95/max(s) | TPOT mean/p95/max(ms) | retractions | accept | hit |") + print("|---|---|---|---|---|---|---|---|---|---|") + for tag, _, _ in POINTS: + n = new.get(tag) + if n is None: + continue + ttft, tpot = n.get("ttft_s") or {}, n.get("tpot_s") or {} + def trio(d, ms=False): + m, p, x = d.get("mean"), d.get("p95"), d.get("max") + if m is None: + return "—" + if ms: + return f"{m*1000:.1f}/{p*1000:.1f}/{x*1000:.1f}" + return f"{m:.2f}/{p:.2f}/{x:.2f}" + hit = (n.get("cache_hit_from_logs") or {}).get("hit_rate") + print(f"| {tag} | {n.get('ok')}/{n.get('num_requests')} | {n.get('wall_s')} " + f"| {n.get('output_throughput_tok_s')} | {n.get('input_throughput_tok_s')} " + f"| {trio(ttft)} | {trio(tpot, True)} | {n.get('retractions_total')} " + f"| {n.get('spec_accept_length_mean')} | {hit} |") + +if __name__ == "__main__": + main() diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/run_retest_e7b64.sh b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/run_retest_e7b64.sh new file mode 100644 index 0000000..9e9c208 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/run_retest_e7b64.sh @@ -0,0 +1,133 @@ +#!/bin/bash +# run_retest_e7b64.sh — E7b high-concurrency retest (MRR 64 + decode graph <=64) +# +# Retests ONLY the points where the original E7b arm (MRR 16, decode graph +# bs 1-8) lost the CUDA graph at running bs > 8: 16K/4.1/4.2 ladders at +# cc 8/16/32/64 (cc8 = on-graph control anchor; 4.2 c64 was user-aborted in +# the original run, now in scope with the cc ceiling raised to 64). +# Same protocol as run_b300_matrix.sh: cold points, recycled corpus windows +# (fresh container instance = windows virgin again), per-point idle-wait -> +# scenario prewarm (first point) -> flush -> bench -> hit verification. +# run-ids 96xx distinguish from the original 95xx e7b arm. +# +# Usage: nohup bash /root/run_retest_e7b64.sh \ +# > /root/bench_logs/b300eq_e7b64_progress.log 2>&1 & +set -u +ARM=e7b64 +CONTAINER=glm53-nvfp4 + +CORPUS=/root/corpus_ids.json +BENCH=/root/bench_corpus_v2.py +EXTRACT=/root/extract_summary.py +URL=http://127.0.0.1:30000 +STAMP=$(date +%Y%m%d_%H%M) +LOG=/root/bench_logs/b300eq_${ARM}_${STAMP} +mkdir -p "$LOG" +RESULTS="$LOG/all_results.jsonl" +STATUS="$LOG/status.txt" +: > "$STATUS" + +echo "=== [$ARM] retest start $(date) LOGDIR=$LOG ===" + +# ---- server facts snapshot ---- +alive() { [ -n "$(docker ps --filter name=$CONTAINER --filter status=running -q)" ]; } +alive || { echo "=== [$ARM] ABORT: container $CONTAINER not running ==="; exit 1; } +docker inspect "$CONTAINER" --format '{{.Config.Cmd}}' > "$LOG/server_cmd.txt" 2>&1 +docker logs "$CONTAINER" 2>&1 | grep -E "max_total_num_tokens|KV Cache is allocated|context_len|chunked_prefill_size|max_running_request|cuda_graph|speculative_num_steps|Capture cuda graph" | head -40 > "$LOG/server_facts.txt" 2>&1 +nvidia-smi --query-gpu=index,name,memory.total,memory.used --format=csv > "$LOG/gpu_inventory_idle.csv" 2>&1 +md5sum "$CORPUS" "$BENCH" "$EXTRACT" > "$LOG/md5_assets.txt" 2>&1 +echo "--- server_cmd: $(cat "$LOG/server_cmd.txt")" +echo "--- server_facts:"; cat "$LOG/server_facts.txt" + +# ---- per-rank VRAM sampler (whole-arm timeline, 30s cadence) ---- +( while true; do + echo "# $(date +%s) $(date +%T)" + nvidia-smi --query-gpu=index,memory.used,utilization.gpu --format=csv,noheader + sleep 30 + done ) > "$LOG/vram_timeline.csv" 2>&1 & +SAMPLER=$! +trap 'kill $SAMPLER 2>/dev/null' EXIT + +idle_wait() { + for i in $(seq 1 90); do + local last + last=$(docker logs --since 90s "$CONTAINER" 2>&1 | grep 'running-req' | tail -1) + if [ -z "$last" ] || echo "$last" | grep -q 'running-req: 0'; then return 0; fi + sleep 10 + done + echo "[warn] idle_wait timeout, continuing" +} + +flush() { curl -s -m 60 -X POST $URL/flush_cache >/dev/null; sleep 3; } + +prewarm() { # $1=input_len $2=base — one uncounted request at scenario length + timeout 1800 python3 "$BENCH" --corpus "$CORPUS" --input-len "$1" --output-len 16 \ + --shared-frac 0 --concurrency 1 --num-requests 1 --run-id 9699 \ + --pool-override "$2" --url $URL/generate --container "$CONTAINER" \ + > "$LOG/prewarm_$1.log" 2>&1 || true +} + +bench_call() { # $1=log $2=tag $3=isl $4=osl $5=cc $6=nreq $7=base $8=rid + timeout 7200 python3 "$BENCH" --corpus "$CORPUS" --input-len "$3" --output-len "$4" \ + --shared-frac 0 --concurrency "$5" --num-requests "$6" --run-id "$8" \ + --pool-override "$7" --url $URL/generate --container "$CONTAINER" \ + --dump-records "$LOG/${2}_records.jsonl" > "$1" 2>&1 +} + +run_point() { # tag isl osl cc nreq base rid [prewarm=1] + local TAG=$1 ISL=$2 OSL=$3 CC=$4 NREQ=$5 BASE=$6 RID=$7 PW=${8:-0} + alive || { echo "$TAG CONTAINER_DEAD" >> "$STATUS"; echo "=== $TAG ABORT: container dead ==="; exit 1; } + echo "=== [$ARM] $TAG isl=$ISL osl=$OSL cc=$CC nreq=$NREQ base=$BASE start $(date +%T) ===" + idle_wait + [ "$PW" = "1" ] && prewarm "$ISL" "$BASE" + flush + local rc=0 + bench_call "$LOG/${TAG}.log" "$TAG" "$ISL" "$OSL" "$CC" "$NREQ" "$BASE" "$RID" || rc=$? + if [ "$rc" = "0" ]; then + python3 "$EXTRACT" "$LOG/${TAG}.log" "$TAG" "$ARM" "$RESULTS"; rc=$? + fi + if [ "$rc" = "4" ]; then + echo "=== $TAG hit_rate>0.01, retrying once ===" + idle_wait; flush + bench_call "$LOG/${TAG}_retry.log" "${TAG}_retry" "$ISL" "$OSL" "$CC" "$NREQ" "$BASE" "$RID" || rc=$? + if [ "$rc" = "0" ]; then + python3 "$EXTRACT" "$LOG/${TAG}_retry.log" "${TAG}_retry" "$ARM" "$RESULTS"; rc=$? + fi + if [ "$rc" = "4" ]; then + echo "$TAG HIT_FAIL_FINAL" >> "$STATUS" + echo "=== $TAG hit_rate still >0.01 after retry — ABORTING ARM (contamination) ===" + exit 1 + fi + fi + if [ "$rc" = "0" ]; then + echo "$TAG OK" >> "$STATUS" + else + echo "$TAG BENCH_FAIL rc=$rc" >> "$STATUS" + echo "=== $TAG bench rc=$rc (recorded, continuing) ===" + fi + echo "=== $TAG done $(date +%T) ===" +} + +# ---- same recycled window bases as the original e7b arm (fresh instance = +# virgin texts for this server process, hit-check still enforced) ---- +B_16K=2300000; B_1K=4500000; B_1K4=4700000 + +# ===== B300 §3: main scenario 16K -> 512 (cc1 unaffected by graph, skipped) ===== +run_point b3_16k_c8 16384 512 8 16 $B_16K 9602 1 +run_point b3_16k_c16 16384 512 16 32 $B_16K 9603 +run_point b3_16k_c32 16384 512 32 64 $B_16K 9604 +run_point b3_16k_c64 16384 512 64 128 $B_16K 9605 + +# ===== B300 §4.1: short input 1K -> 128 ===== +run_point b41_1k_c8 1024 128 8 16 $B_1K 9612 1 +run_point b41_1k_c32 1024 128 32 64 $B_1K 9613 +run_point b41_1k_c64 1024 128 64 128 $B_1K 9614 + +# ===== B300 §4.2: long output 1K -> 4K (c64 now in scope, ceiling raised to 64) ===== +run_point b42_1k4k_c8 1024 4096 8 16 $B_1K4 9622 1 +run_point b42_1k4k_c32 1024 4096 32 64 $B_1K4 9623 +run_point b42_1k4k_c64 1024 4096 64 128 $B_1K4 9624 + +nvidia-smi --query-gpu=index,name,memory.total,memory.used --format=csv > "$LOG/gpu_inventory_final.csv" 2>&1 +echo "=== [$ARM] RETEST ALL DONE $(date) — $LOG ===" +echo "--- status:"; cat "$STATUS"