diff --git a/deploy/CURRENT.md b/deploy/CURRENT.md index 3f8ec0a..b377fc8 100644 --- a/deploy/CURRENT.md +++ b/deploy/CURRENT.md @@ -1,4 +1,4 @@ -# 现役部署状态页(live 核验于 2026-09-09) +# 现役部署状态页(live 核验于 2026-09-10,60.8 当日核验;其余机器 09-09 口径) > 本页记录 174.1.60.x 集群"现在跑的是什么"。改动机器前后先读这页; > **任何变动(停/起/改配置)完成后必须更新本页**。核验命令:`docker ps` + `nvidia-smi`。 @@ -14,7 +14,7 @@ | 60.5 | `glm53-nvfp4`(Up 2d,09-09 只读核验) | **NVFP4 团队生产**(deploy_glm53_605.sh,md5 fcd9109b)。生产机铁律:不实验、不重启、不覆盖脚本。**交付升级路径已更新为 v3(09-09 hit90 场景优胜 TP4PP2-hicache@0.90,池 647,040/c4 并发独立文档/hit90 cc8 out +34% vs 现役):脚本 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/scripts/deploy_glm53_605_v3.sh` + 补丁束(60.8:/root/glm53_r37_patch_bundle_v3.tar.gz,md5 6922e534,需 scp 至 60.5)+ 需 docker pull cu13 镜像;未执行、未落 60.5 磁盘(60.5:/root 仅有原脚本,核验过)。此前 v2(TP2PP4-hicache,冷缓存口径优胜)被 v3 取代,仍留仓库可作"容量优先 6 条文档/512k 单条最快"备选;回滚=原脚本** | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(方案 A 口径);待切 `glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`(容量优先备选 `..._tp2pp4_hicache.env`) | | 60.6 | 无容器,但 8 卡被外部裸金属实验占用(`/data/hzy/sparse-opd-*`,09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — | | 60.7 | 基本空(4 卡仍有 `/home/user/dirA_exp` 外部小任务,09-08 晚实测) | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — | -| 60.8 | `glm53-nvfp4`(Up,:30000,restart=unless-stopped) | **TP4PP2-nomtp@0.90 hit90 优胜口径在役**(09-09 晚接替同日早间的 TP2PP4@0.85:同 r37 栈 cu13+9 挂载+de-GLOO,仅改 tp4/pp2 + memfrac 0.90;KV 池 **647,040**;实测实例原样保留切 restart 策略,未重建容器。启动 `bash /root/deploy_ppmtp_r37.sh '--tp 4 --pp-size 2 --disable-overlap-schedule --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length 1048576' nomtp 8192 0.90 1 1 0 0`)。hit90(90% 命中 i128k/o512)out cc1/2/3/4/8/16 = **28.8/47.5/63.9/76.2/106.4/125.7**(vs TP2PP4 基线 cc4 +14%/cc8 +15%);并发独立 128k 文档 **4 条**(cap cc4 98.5 零排队)、512k 单条 ✓ TTFT 135.8s、质量门 **7/7**;同轮判决:DP attention 容量负收益判死、DCP 对 DSA 静默算错禁用、MTP@128k accept 2.07 判负。让步项(知情):容量 4 条(TP2PP4 为 6)、512k 单条慢 33%。回滚 TP2PP4 = 同命令改 `'--tp 2 --pp-size 4'` + memfrac 0.85 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`;实验全量 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/`;前史 `experiments/pro6000/glm53_nvfp4_128k_capacity_topology/`(TP2PP4 口径)+ r37 `experiments/pro6000/glm53_ppmtp_r37_verify_graph/` | +| 60.8 | `glm53-nvfp4`(Up,:30000,restart=unless-stopped) | **TP4PP2-nomtp@0.90 hit90 优胜口径在役**(09-09 晚接替同日早间的 TP2PP4@0.85:同 r37 栈 cu13+9 挂载+de-GLOO,仅改 tp4/pp2 + memfrac 0.90;KV 池 **647,040**;实测实例原样保留切 restart 策略,未重建容器。启动 `bash /root/deploy_ppmtp_r37.sh '--tp 4 --pp-size 2 --disable-overlap-schedule --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length 1048576' nomtp 8192 0.90 1 1 0 0`)。hit90(90% 命中 i128k/o512)out cc1/2/3/4/8/16 = **28.8/47.5/63.9/76.2/106.4/125.7**(vs TP2PP4 基线 cc4 +14%/cc8 +15%);并发独立 128k 文档 **4 条**(cap cc4 98.5 零排队)、512k 单条 ✓ TTFT 135.8s、质量门 **7/7**;同轮判决:DP attention 容量负收益判死、DCP 对 DSA 静默算错禁用、MTP@128k accept 2.07 判负。让步项(知情):容量 4 条(TP2PP4 为 6)、512k 单条慢 33%。回滚 TP2PP4 = 同命令改 `'--tp 2 --pp-size 4'` + memfrac 0.85。**09-10 B300 对标战役**:停役(容器 rename 保全 `glm53-nvfp4-insvc`)→ 双臂 B300 场景矩阵(TP2PP4-D 口径 24 点 + E7b 配方 16+1 点,判决=分界 C8/E7b 窗口≤C8/边界仅 TP2PP4 可达/与 B300 绝对差 4-5×,报告飞书 wiki `A7V3wZTQeifCB4krdi6cA834nW9`)→ **原容器恢复并核验**(rename 回 + start,health 200、16K 抽测 ok、显存水位与停役前一致,口径未变) | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`;实验全量 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/`;B300 对标 `experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/`;前史 `experiments/pro6000/glm53_nvfp4_128k_capacity_topology/`(TP2PP4 口径)+ r37 `experiments/pro6000/glm53_ppmtp_r37_verify_graph/` | ## 方案 A-F 一览(GLM-5.3-NVFP4 @ pro6000,2026-09-08 双场景报告口径) diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/README.md b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/README.md new file mode 100644 index 0000000..60866ec --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/README.md @@ -0,0 +1,68 @@ +# GLM-5.3-NVFP4 双方案 B300 对标场景矩阵压测 — 60.8(6000D) + +日期:2026-09-10 | 机器:174.1.60.8(8×RTX 6000D,96GB GDDR7,无 NVLink) +模型:GLM-5.3-NVFP4(modelopt)| 镜像:`nightly-dev-20260828-daf63171`(两臂同) +对标基线:飞书《GLM 5.3 | SGLang | Low-latency & High-Throughput 测试结果》(B300 报告,wiki UPB2w4Y5yi65qwkMxJJcZko5nUc) +完整报告:本目录 `REPORT.md`(= 飞书发布版 A7V3wZTQeifCB4krdi6cA834nW9) + +## 目标 + +在 6000D 上复刻 B300 报告的全部场景(主场景 16K→512、4.1 短输入、4.2 长输出、5.1 长上下文、5.2 边界),对两套在役部署方案各跑一遍完整矩阵,产出对齐 B300 8 章结构的对标报告。测后 60.8 在役服务(TP4PP2@0.90)原容器恢复(已验证:health 200 + 16K 抽测 ok + 显存水位一致)。 + +## 实验臂 + +| 臂 | 方案 | 关键配置 | 质量门 | +|---|---|---|---| +| tp2pp4 | **D 生产口径**(deploy_glm53_pp4.sh) | TP2PP4、mem0.85、MRR48、cps16384、radix 关、KV fp8_e4m3 池 1,040,384、无投机、index_topk_freq=4(=原生默认,恒等)、ctx 1,048,576 | 6/7(仅 tool-call:无 parser,历史已知) | +| e7b | **TP8+EAGLE3+AR**(deploy_glm53_607_exp.sh + CAR 补丁注入) | TP8、EAGLE 4/1/5、mem0.90、MRR16、cps8192、radix 开+hicache×3、KV fp8_e4m3 GPU 池 276,480、decode 图 bs1-8、ctx 270,336、custom-AR 1stage(`SGLANG_CUSTOM_ALLREDUCE_ALGO=1stage`,8 rank `SSKJ_CAR_PATCH_ACTIVE` 验证) | 7/7 | + +## 场景矩阵与并发档位(用户裁决收敛:16K 封 64、64K/128K 封 8/4) + +| B300 章节 | 场景 | 并发档位 | TP2PP4 活跃上限 | E7b 活跃上限 | +|---|---|---|---|---| +| §3 主场景 | 16K→512 | 1/8/16/32/64 | 48(MRR) | 16(MRR=池贴边) | +| §4.1 | 1K→128 | 1/8/32/64 | 48 | 16 | +| §4.2 | 1K→4K | 1/8/32/64 | 48 | 16(c64 用户中止 rc=143) | +| §5.1 | 64K→512 | 1/4/8 | 15(池) | 4(池) | +| §5.1 | 128K→512 | 1/2/4 | 7(池) | 2(池) | +| §5.2 | 256K→1 | 1/2/3 | 3(池) | 1(池贴边) | +| §5.2 | 512K→1 | 1 | 1 | 结构性不可(ctx) | +| §5.2 | 896K→1(代 B300"约1M") | 1 | 1 | 结构性不可(ctx) | + +测量协议:冷缓存(shared-frac 0)+ 每点 flush + 服务端命中核验 ≤0.01(超限重试一次);nreq=max(8, 2×cc);P95 nearest-rank 对齐 B300;语料耗尽(21.23M/21.30M)下用回收窗口(`--pool-override` 基址映射见 REPORT 附录 A)。有效测量点 41 个(tp2pp4 24 + e7b 16 + e7b 256K C=1 全新文本重测),全点 0 retraction、0 OOM。 + +## 判决速览(详见 REPORT.md) + +- **分界 C=8**:E7b 全场景 C=1 占优(16K out 49.6 vs 16.7=3.0×、TPOT 13.7 vs 50.3ms=3.7×;1K→4K out 135 vs 20.3=6.7×);C≥8 TP2PP4 全指标反超并随并发拉大(16K c64 out 211 vs 84=2.5×)。 +- **E7b 可用窗口 ≤C8**:MRR16 + decode 掉图(bs>8)双击,16K c16 TPOT 298ms 断崖。 +- **decode 密集甜点 = E7b c8**:1K→4K out 412 tok/s、TPOT 22.5ms、accept 4.0。 +- **TP2PP4 甜点 c16+**:16K 近线性至 c64(MRR48 未饱和);全场最高输出 4.2 c64 482 tok/s(但 TTFT 338s,仅离线)。 +- **边界只有 TP2PP4 可达**:256K/512K/896K 全测(input 7,050/5,662/4,153 tok/s);E7b ctx 270,336 结构性封顶。 +- **DSA 复现**:C=1 TPOT 对上下文不敏感(50.3/50.0/49.7ms @16/64/128K),并发才是驱动(c8: 70.9→161.8ms)。 +- **vs B300**:定性结构完全复现(LL/HT 分野一致),绝对差 4-5×,边界 prefill 差距收窄至 ~2×;分界点本机更靠前(C8 vs C64-128),原因是容量上限(MRR/池)而非算力。 + +## 关键坑位(复测必读) + +1. **hicache 宿主层陷阱**:256K prewarm KV 占池 94.8% 触发宿主层下放,`flush_cache` 清不掉宿主层 → 同文本测量命中 0.9998。冷缓存复测**必须换该实例从未发过的文本**(本战役 E7b 256K C=1 用窗口 9,900,000 重测达标)。 +2. 语料已耗尽:回收窗口复用仅在 flush+命中核验协议下有效;TP2PP4 臂 radix 本来就关,零污染。 +3. 在役保全流程:`docker stop` → `docker rename glm53-nvfp4 glm53-nvfp4-insvc`(必须先改名,E7b 部署脚本会 rm -f 同名容器)→ 测毕 `rename` 回 + `start`。docker stop/rm 偶发 "zombie PID" 报错是收尾边界现象,容器终态 exited(137)、显存归零,稍等重试即可。 + +## 资产与 md5 台账(60.8 执行件 = 本目录 = 60.7 原件 三方一致) + +``` +1e34dd8d813ccb7ac8243e8d1573a49e bench_corpus_v2.py (scripts/) +4c126d067d33b5ea27c268f37561634c extract_summary.py (scripts/) +5892b44610b2ce61f533f1721105625e run_b300_matrix.sh (scripts/) +def3c64c5dc19e1d507080e3366c3762 deploy_glm53_pp4.sh (见 dual_scenario_bench/scripts/,md5 对照一致) +21db641f1d997e26fcf1ddc97163b87e deploy_glm53_607_exp.sh (见 dual_scenario_bench/scripts/,md5 对照一致) +a8fc9a504c041acc40831a848b4985d8 patches/custom_all_reduce.py +65a4d22bd87ae343e7d68c4879c14f98 patches/custom_all_reduce_utils.py +``` + +gen_report_tables.py 为本地表格生成器(md5 未入台账,60.8 侧执行件同源)。 + +## 原始数据 + +- 本目录 `results/{tp2pp4,e7b}/`:all_results.jsonl(逐点 SUMMARY + 命中核验)、status.txt、server_facts.txt(启动参数+池分配日志摘录)、gpu_inventory_idle/final.csv、vram_timeline.csv(30s 采样全矩阵) +- 60.8 侧:`/root/bench_logs/b300eq_tp2pp4_20260910_1138/`、`/root/bench_logs/b300eq_e7b_20260910_1428/` +- E7b 256K C=1:all_results.jsonl 同 tag 共 4 条,最后一条为干净重测值(生成器 dict 载入后写覆盖,天然生效) diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/REPORT.md b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/REPORT.md new file mode 100644 index 0000000..78f5703 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/REPORT.md @@ -0,0 +1,280 @@ +# GLM-5.3-NVFP4 | RTX 6000D | SGLang 双方案 B300 对标场景压测报告 + +- 测试日期:2026-09-10(单日单机完成两臂) +- 测试机:174.1.60.8(6000D,8 卡) +- 对标基线:飞书《GLM 5.3 | SGLang | Low-latency & High-Throughput 测试结果》(B300 报告,wiki UPB2w4Y5yi65qwkMxJJcZko5nUc) +- 测后状态:60.8 在役服务(TP4PP2@0.90 口径)已原容器恢复并验证(health 200 + 16K 抽测 ok + 显存水位与停役前一致) + +## 1. 结论摘要 + +本轮在单台 8 卡 RTX 6000D 上,用 GLM-5.3-**NVFP4** 完整复刻 B300 报告的场景矩阵,测试了两套在役部署方案。两套方案代表完整部署形态,不是单参数 A/B:**TP2PP4** 为 D 生产口径(吞吐/长上下文形态),**TP8+EAGLE3+AR** 为 E7b 配方(低延迟形态,含 custom allreduce 1stage 补丁)。 + +- **低并发优先 E7b(TP8+EAGLE3+AR)**:主场景 `16K→512, C=1` 输出 49.6 tok/s、TPOT 13.7 ms,对 TP2PP4(16.7 tok/s、50.3 ms)分别是 **3.0×** 与 **3.7×**;长输出 `1K→4K, C=1` 输出 135 tok/s、TPOT 8.7 ms,对 TP2PP4(20.3、49.8 ms)是 **6.7×** 与 **5.7×**。 +- **高并发优先 TP2PP4**:主场景 C=64 达到 6,748 input tok/s / 211 output tok/s,对 E7b(2,697 / 84.3)均为 **2.5×**;短输入 C=32/64 输出 276/341 tok/s,对 E7b(100/103)为 **2.7~3.3×**。 +- **E7b 的可用并发窗口比 B300 Low-Latency 窄一个数量级**:MRR=16 且 CUDA graph 仅覆盖 decode bs 1–8,并发 ≥8 即掉图,主场景 C=16 TPOT P95 跳到 297.8 ms(C=8 为 135.1 ms;同点 TP2PP4 仅 98.5 ms)。E7b 的生产甜点上限 = **C≤8**。 +- **TP2PP4 甜点在 C=16 之后**:主场景输出吞吐从 C=8 的 101 近线性爬到 C=64 的 211 tok/s(MRR48 尚未饱和),4.2 长输出 C=64 达全场最高 482 tok/s(但 TTFT P95 338 s,需要排队预算)。 +- **decode 密集低并发的最优解是 E7b C=8**:`1K→4K, C=8` 输出 412 tok/s、TPOT 22.5 ms,对 TP2PP4(101 tok/s、79.4 ms)为 4.1×;EAGLE 实测 accept length 4.0。 +- **长上下文与容量边界只有 TP2PP4 可达**:128K C=1 两方案输出打平(11.8 vs 11.9 tok/s)但 TP2PP4 TTFT 减半(18.2 s vs 37.6 s);256K/512K/896K 边界 E7b 结构性不可测(ctx 270,336 封顶 + KV 池 276,480 贴边),TP2PP4 全部完成(256K C=1/2/3、512K/896K C=1)。 +- **DSA 特性在 6000D 复现**:TP2PP4 C=1 的 TPOT 对上下文长度不敏感(16K/64K/128K = 50.3/50.0/49.7 ms 恒定),并发才是 TPOT 驱动因子(16K 行 C=8→C=64:70.9→203.7 ms)。 +- **与 B300 的绝对差距约 4~5×**,边界 prefill 差距收窄到约 2×(256K C=1 input 7,050 vs 17,457;896K 4,153 vs 约1M 行 7,897)。硬件与量化口径不同(B300 报告未写明量化方式),绝对值仅量级可比,两份报告的结构性结论一致(见第 9 章)。 +- 全部 41 个有效测量点 **0 回退(retraction)、0 OOM**,冷缓存命中核验全部 ≤0.01(E7b 256K C=1 首测触 hicache 宿主层陷阱,用全新文本重测达标,见 5.2 注记)。 + +## 2. 测试环境与配置 + +| 项目 | TP2PP4(D 生产口径) | TP8+EAGLE3+AR(E7b 配方) | +|-|-|-| +| 硬件 | 单机 8 × NVIDIA RTX 6000D(96 GB GDDR7,nvidia-smi 可见 85,651 MiB/卡) | 同左 | +| 模型 | GLM-5.3-NVFP4(modelopt 量化,/data/hf_models/GLM-5.3-NVFP4) | 同左 | +| 镜像 | `nightly-dev-20260828-daf63171` | 同左 | +| 并行 | TP2 × PP4 | TP8 | +| 投机解码 | 无 | EAGLE3,num_steps=4,topk=1,draft_tokens=5 | +| `mem-fraction-static` | 0.85 | 0.90 | +| 最大活跃请求(MRR) | 48 | 16 | +| Chunk Prefill | 16,384 | 8,192 | +| KV dtype | fp8_e4m3 | fp8_e4m3 | +| KV 池(服务端实测) | **1,040,384 tokens**(12.6~13.4 GB/rank,无宿主层) | **276,480 tokens GPU**(15.8 GB/rank)+ 分层缓存 hicache×3 宿主层(write_through) | +| radix cache | 关(`disable_radix_cache=True`) | 开(分层缓存) | +| 上下文上限 | 1,048,576(config 原生) | 270,336(显存约束下的部署值) | +| CUDA graph | 常规捕获 | decode 图 bs 1–8(bs>8 掉图) | +| custom allreduce | — | 1stage 补丁注入(`SGLANG_CUSTOM_ALLREDUCE_ALGO=1stage` 环境强制;8 rank `SSKJ_CAR_PATCH_ACTIVE` 日志验证全出现) | +| `index_topk_freq` | 4(override,等于原生默认,恒等) | 原生默认 4 | +| 质量门 | 6/7(仅 tool-call 失败:D 口径未配 parser,历史已知;其余全过) | **7/7** | + +因此,下文比较回答的是"两种部署形态谁更适合该负载",不能把差异单独归因于 EAGLE、PP 流水、chunk、radix 或图覆盖中的某一项(与 B300 报告同款声明)。 + +**测量协议**(对齐 B300 口径): + +- 冷缓存:`--shared-frac 0`,每点前 `POST /flush_cache`,服务端 Prefill 日志核算命中率,>0.01 重测一次,仍超停点排查;全矩阵命中核验最终全部达标。 +- 指标:Input TPS / Output TPS / TTFT P95 / TPOT P95(P95 为 nearest-rank;Input TPS = 输入 token / 全程墙钟,与 B300 口径一致)。 +- 负载:PG19 真实语料 token 切片(`corpus_ids.json`),input_ids 直打 `/generate`,temperature=0、ignore_eos、流式;nreq = max(8, 2×并发),边界行 nreq=并发。 +- 并发档位按决策收敛:16K/1K 类封顶 64;64K 封 8、128K 封 4(128K C=2 补一档);超出活跃上限的档位是**排队观察点**(与 B300 C=256 同性质,保留为有效观察)。 +- 语料已耗尽(21.23M/21.30M),冷缓存口径下用**回收窗口**复用(窗口基址见附录 A,逐记录 `corpus_window` 字段留档)。 + +**两臂活跃上限**(MRR 与 KV 池决定,解释各行哪些并发是排队观察点): + +| 场景 | TP2PP4 活跃上限 | E7b 活跃上限 | +|-|-|-| +| 16K / 1K | 48(MRR) | 16(MRR,=池贴边) | +| 64K | 15(池) | 4(池) | +| 128K | 7(池) | 2(池) | +| 256K | 3(池) | 1(池 262K KV / 276K 贴边) | +| 512K / 896K | 1(池) | 结构性不可(ctx 270,336) | + +## 3. 主场景:16K 输入、512 输出 + +B300 跑了 C=1/8/32/64/128/256;本机按 MRR 上限收敛为 C=1/8/16/32/64(C=16 为本机甜点档,B300 无此档;C=128/256 超出两臂 MRR)。 + +| 并发 | 方案 | Input TPS | Output TPS | TTFT P95 | TPOT P95 | +|-|-|-|-|-|-| +| 1 | TP2PP4 | 534 | 16.7 | 5.36 s | 50.3 ms | +| 1 | TP8+EAGLE3+AR | 1,587 | 49.6 | **3.98 s** | **13.7 ms** | +| 8 | TP2PP4 | 3,236 | **101** | **14.79 s** | **70.9 ms** | +| 8 | TP8+EAGLE3+AR | 2,605 | 81.4 | 32.52 s | 135.1 ms | +| 16 | TP2PP4 | 4,674 | **146** | **25.78 s** | **98.5 ms** | +| 16 | TP8+EAGLE3+AR | 2,217 | 69.3 | 60.86 s | 297.8 ms | +| 32 | TP2PP4 | **6,185** | **193** | **46.43 s** | **152.0 ms** | +| 32 | TP8+EAGLE3+AR | 2,527 | 79.0 | 169.68 s | 259.9 ms | +| 64 | TP2PP4 | **6,748** | **211** | **134.78 s** | **203.7 ms** | +| 64 | TP8+EAGLE3+AR | 2,697 | 84.3 | 345.69 s | 231.5 ms | + +趋势: + +- **分界在 C=8**:C=1 E7b 全指标占优;C=8 起 TP2PP4 全指标反超,且输出吞吐差距随并发拉大(101 vs 81 → 211 vs 84)。 +- **E7b 在 C=8→16 输出吞吐倒退**(81.4→69.3 tok/s):MRR=16 开始排队 + decode 掉图(bs>8 无图)双击;TPOT P95 从 135 ms 跳到 298 ms。EAGLE accept length 随并发从 2.14 爬到 2.99,但被掉图抵消。 +- **prefill 墙的差异**:E7b 的 input TPS 几乎不随并发增长(C=8→64:2,605→2,697,+3%),TP2PP4 翻倍(3,236→6,748,+108%)——chunk 8192 + TP8 无 PP 流水的 prefill 瓶颈 vs chunk 16384 + PP4 流水摊满。 +- TP2PP4 到 C=64 仍在爬坡(C=32→64 +9%),MRR48 未饱和;TTFT P95 在 C=64 达 134.8 s,同 B300 一样高并发 TTFT 需要准入控制。 + +## 4. 短输入与长输出 + +### 4.1 `1K -> 128` + +| 并发 | 方案 | Input TPS | Output TPS | TTFT P95 | TPOT P95 | +|-|-|-|-|-|-| +| 1 | TP2PP4 | 155 | 19.3 | 386 ms | 49.3 ms | +| 1 | TP8+EAGLE3+AR | **518** | **64.7** | **326 ms** | **14.3 ms** | +| 8 | TP2PP4 | 815 | 102 | **1.69 s** | 72.9 ms | +| 8 | TP8+EAGLE3+AR | **1,219** | **152** | 2.16 s | **53.7 ms** | +| 32 | TP2PP4 | **2,204** | **276** | **4.87 s** | **98.5 ms** | +| 32 | TP8+EAGLE3+AR | 801 | 100 | 25.60 s | 190.5 ms | +| 64 | TP2PP4 | **2,724** | **341** | **19.96 s** | **101.8 ms** | +| 64 | TP8+EAGLE3+AR | 826 | 103 | 65.81 s | 177.7 ms | + +短输入下 E7b 在 C≤8 显著占优(C=8 输出 152 vs 102,1.5×),C=32 起 TP2PP4 大幅拉开(2.7~3.3×)。E7b 的 input TPS 反而在 C=8 最高(1,219)后回落——MRR16 排队开始挤占 prefill。B300 同场景 Low-Latency 到 C=128 才被反超,本机提前到 C=8~32 之间,同样是容量上限(MRR16/池)而非算力所致。 + +### 4.2 `1K -> 4K` + +| 并发 | 方案 | Output TPS | TTFT P95 | TPOT P95 | +|-|-|-|-|-| +| 1 | TP2PP4 | 20.3 | 395 ms | 49.8 ms | +| 1 | TP8+EAGLE3+AR | **135** | **320 ms** | **8.7 ms** | +| 8 | TP2PP4 | 101 | 1.87 s | 79.4 ms | +| 8 | TP8+EAGLE3+AR | **412** | **1.68 s** | **22.5 ms** | +| 32 | TP2PP4 | **367** | **4.14 s** | **90.5 ms** | +| 32 | TP8+EAGLE3+AR | 252 | 304.75 s | 74.6 ms | +| 64 | TP2PP4 | **482** | 337.92 s | 95.5 ms | +| 64 | TP8+EAGLE3+AR | 用户中止*(见注) | — | — | + +\* E7b C=64 点按用户指示中止("并发 64 太高",rc=143),未获得有效数据;同点 TP2PP4 已完成。E7b 该点 nreq=128 远超 MRR=16,属排队观察点,中止不影响结论完整性。 + +长输出放大了两形态的差异: + +- **E7b C=1/C=8 是 decode 密集负载的最优区间**:C=1 输出 135 tok/s、TPOT 8.7 ms(全场最低),C=8 输出 412 tok/s(全场第二),EAGLE accept 长达 3.7~4.0——长输出让草稿模型进入"顺笔"状态,accept 显著高于 4.1 短输出行(2.0~2.1)。 +- **E7b C=32 的 TTFT P95 304.75 s** 是纯排队(nreq=64 / MRR=16,4 波串行),其 TPOT 74.6 ms 与掉图后水平一致。 +- **TP2PP4 C=64 输出 482 tok/s 为全场最高**,但 TTFT P95 338 s 意味着该点只适合离线批处理;交互负载应压在 C=32(367 tok/s、TTFT 4.1 s)。 + +## 5. 长上下文观察 + +### 5.1 `64K/128K -> 512` + +(并发档位按用户指示收敛:64K 封 8、128K 封 4,另补 128K C=2;B300 同场景为 64K C=8/32/64、128K C=8/32。) + +| 场景 | 并发 | 方案 | Input TPS | Output TPS | TTFT P95 | TPOT P95 | +|-|-|-|-|-|-|-| +| 64K→512 | 1 | TP2PP4 | 1,834 | 14.3 | **10.32 s** | 50.0 ms | +| 64K→512 | 1 | TP8+EAGLE3+AR | **2,877** | **22.5** | 17.17 s | **13.5 ms** | +| 64K→512 | 4 | TP2PP4 | **4,287** | **33.5** | **28.83 s** | **99.5 ms** | +| 64K→512 | 4 | TP8+EAGLE3+AR | 3,292 | 25.7 | 68.87 s | 154.6 ms | +| 64K→512 | 8 | TP2PP4 | **5,650** | **44.1** | **53.38 s** | **161.8 ms** | +| 64K→512 | 8 | TP8+EAGLE3+AR | 3,280 | 25.6 | 149.45 s | 157.5 ms | +| 128K→512 | 1 | TP2PP4 | 3,017 | 11.8 | **18.23 s** | 49.7 ms | +| 128K→512 | 1 | TP8+EAGLE3+AR | **3,054** | **11.9** | 37.60 s | **12.3 ms** | +| 128K→512 | 2 | TP2PP4 | **4,310** | **16.8** | **32.42 s** | **83.7 ms** | +| 128K→512 | 2 | TP8+EAGLE3+AR | 3,145 | 12.3 | 75.81 s | 162.2 ms | +| 128K→512 | 4 | TP2PP4 | **5,626** | **22.0** | **60.35 s** | **146.8 ms** | +| 128K→512 | 4 | TP8+EAGLE3+AR | 3,133 | 12.2 | 159.50 s | 163.4 ms | + +- 64K C=1 E7b 仍占优(22.5 vs 14.3 tok/s),但 128K C=1 两方案输出打平(11.8 vs 11.9)——prefill 逐渐成为长上下文的主导成本,E7b 的 decode 优势被稀释;其 TTFT 反而慢 2×(37.6 vs 18.2 s)。 +- C≥2 起 TP2PP4 全指标占优;E7b 的 output TPS 在 64K/128K 行几乎不随并发变化(22.5→25.6、11.9→12.2),与主场景同一形态:容量上限 + 掉图封死并发收益。 +- DSA 的 TPOT 上下文不变性(TP2PP4 C=1:50.0 ms @64K ≈ 49.7 ms @128K ≈ 50.3 ms @16K)与并发驱动性(C=8:70.9 ms @16K → 161.8 ms @64K)在本组完整呈现。 + +### 5.2 上下文边界 + +(OSL=1,只验证容量与 prefill,不比较 Output TPS/TPOT——与 B300 同声明。B300 完成到约 1M;本机以 896K=917,504 tokens 对应 B300"约 1M"档。) + +| 输入长度 | 方案 | 已完成并发 | C=1 Input TPS | 最高并发 TTFT P95 | +|-|-|-|-|-| +| 256K | TP2PP4 | 1/2/3 | **7,050** | 102.99 s(C=3) | +| 256K | TP8+EAGLE3+AR | 1 | 2,867 | 91.45 s(C=1) | +| 512K | TP2PP4 | 1 | **5,662** | 92.60 s(C=1) | +| 512K | TP8+EAGLE3+AR | 结构性不可(ctx 270,336) | — | — | +| 896K | TP2PP4 | 1 | **4,153** | 220.91 s(C=1) | +| 896K | TP8+EAGLE3+AR | 结构性不可(ctx 270,336) | — | — | + +- 256K C=1 两臂相差 2.5×(7,050 vs 2,867 tok/s):TP2PP4 的 chunk 16384 + PP4 流水对超长 prefill 的摊满优势,在边界长度上比 128K 行(几乎打平)进一步放大;E7b 的 chunk 8192 代价随长度累积。 +- TP2PP4 边界 input TPS 随长度衰减平缓(7,050 → 5,662 → 4,153),896K 单条 220.9 s 完成、池 1,040,384 tokens 单条可容(KV 917,504 + 余量)。 +- **E7b 256K C=1 命中核验注记(hicache 宿主层陷阱)**:首测命中率 0.9998、重试仍超——根因是该点 nreq=1 测量文本与预热完全相同,256K 预热 KV(262,160 tokens)占池 94.8% 触发分层缓存宿主层下放,`flush_cache` 只清 GPU radix 树、清不掉宿主层。改用该服务实例从未发过的文本(窗口基址 9,900,000)无预热重测,命中 0.0,数据干净。此为分层缓存运维要点:**宿主层缓存不受 flush_cache 影响,冷测必须换文本**。 + +## 6. 显存状态 + +- **TP2PP4**:服务加载后空载 64.6 GiB/卡,矩阵峰值 **85.0 GiB/卡**(主场景 C=64 时逼近打满,最紧张卡余量约 0.6 GiB)。mem 0.85 下 KV 池按卡容量贴满分配,属预期;继续上调 MRR 或上下文没有余量,扩容前必须先降 mem-fraction。 +- **TP8+EAGLE3+AR**:空载 77.9 GiB/卡(EAGLE 草稿权重 + mem 0.90 大池),矩阵峰值 **83.6 GiB/卡**(余量约 2.1 GiB)。 +- 两臂全矩阵 **0 OOM、0 retraction**(全部 41 点 retractions_total=0)——B300 未披露该指标,本机在自身容量上限内运行无回退。 +- 显存时间线逐 30 s 采样留档(vram_timeline.csv),可复核任一时刻的卡间分布。 + +## 7. 建议 + +1. **低并发交互/agent 长思考(C≤8)用 E7b**:主场景 C=1 TPOT 13.7 ms、长输出 C=8 输出 412 tok/s。生产并发上限建议钉在 ≤8:C=16 起 decode 掉图 + MRR16 排队使其全面劣于 TP2PP4。 +2. **高并发吞吐/长上下文(C≥8 或输入 ≥64K)用 TP2PP4**:主场景 C=64 输出 211 tok/s、128K C=1 TTFT 18.2 s、896K 可达;MRR48 内未饱和,吞吐上限即 MRR。 +3. **负载形态分界线**:prefill 吞吐需求 >2.2K tok/s 或并发 >8 → TP2PP4;decode 为主且并发 ≤8 → E7b。两臂在 C=8 附近的输出吞吐交叉(主场景 101 vs 81、短输入 102 vs 152、长输出 101 vs 412)——按输出长度分布选型,不能只看并发。 +4. **E7b 扩窗口的两个前置**:MRR 16→更高需先扩 KV 池(hicache 宿主层只救命中场景,不增并发容量);decode 图覆盖 bs 8→16/32 才能消掉 C=16 的 298 ms TPOT 断崖。 +5. **边界与超长上下文只有 TP2PP4 口径可服务**:E7b 若要对标 B300 512K/约1M 行,需要 ctx ≥524,288 与池 ≥52 万 tokens 的部署形态,本版(ctx 270,336 / 池 276,480)结构性不可达。 +6. **不要把两臂差异单归因 EAGLE**:两臂同时差在并行拓扑、chunk、radix、MRR 与图覆盖;单变量消融未做(与 B300 报告建议 3 同款)。 +7. **生产容量同时设吞吐和延迟 SLO**:TP2PP4 主场景 C=64 输出最高但 TTFT P95 已到 135 s;E7b C=32 长输出 TTFT P95 305 s。只看峰值 TPS 会掩盖排队长尾。 +8. **分层缓存运维**:宿主层缓存不受 `flush_cache` 影响,任何冷缓存测量/复测必须更换输入文本(见 5.2 注记)。 + +## 8. 原始结果与复现 + +- 服务器原始结果(60.8): + - TP2PP4 臂:`/root/bench_logs/b300eq_tp2pp4_20260910_1138/`(all_results.jsonl 24 点、status.txt、server_facts.txt、gpu_inventory、vram_timeline.csv) + - E7b 臂:`/root/bench_logs/b300eq_e7b_20260910_1428/`(all_results.jsonl 20 条:16 点 OK + 4.2 C=64 用户中止 + 256K C=1 干净重测覆盖前 3 条污染记录;status.txt 含 4 个结构性跳过与 HIT_FAIL_FINAL 首测记录) + - 资产 md5 台账:`/root/bench_logs/b300eq_md5_ledger.txt` +- 本地镜像:`D:\sskj\b300eq\{tp2pp4,e7b}\`(上述全部文件)、`D:\sskj\b300eq\report_tables.md`(表格生成器输出) +- 部署脚本:`/root/deploy_glm53_pp4.sh`(md5 def3c64c…,与库内 sskj main 副本一致)、`/root/deploy_glm53_607_exp.sh`;CAR 补丁:`/root/patches/custom_all_reduce.py`(a8fc9a50…)+ `custom_all_reduce_utils.py`(65a4d22b…),三处(60.7 原件/本地/库内)md5 一致 +- 测量工具:`/root/bench_corpus_v2.py`(md5 1e34dd8d…,p95 nearest-rank + 逐请求 dump)、`/root/extract_summary.py`(4c126d06…)、`/root/run_b300_matrix.sh`(5892b446…,矩阵驱动:alive/idle_wait/prewarm/flush/命中核验/重试/VRAM 采样) +- 复现命令(单点示例): + +```bash +# 冷缓存压测(E7b 256K C=1 干净版) +python3 /root/bench_corpus_v2.py --input-len 262144 --output-len 1 --shared-frac 0 \ + --concurrency 1 --num-requests 1 --run-id 9551 --pool-override 9900000 \ + --dump-records $L/b52_256k_c1_v2_records.jsonl +# 全矩阵:nohup bash /root/run_b300_matrix.sh > 2>&1 & +``` + +## 9. 与 B300 对比观察 + +> **口径声明**:B300 报告未写明模型量化方式(若为原始 BF16 权重,则与本机 NVFP4 非同模型形态);硬件为 8×B300(288 GB HBM3e)vs 本机 8×RTX 6000D(96 GB GDDR7);镜像 v0.5.18-cu130-dev4 vs nightly-20260828-daf63171。**绝对值仅量级可比,本对比只对结构性结论负责**。 + +- **定性结构完全复现**:低延迟配方(B300 Low-Latency = TP8+EAGLE vs 本机 E7b = TP8+EAGLE3+AR)在 C=1 占优、吞吐配方(B300 High-Throughput = DP8+DeepEP vs 本机 TP2PP4 = D 生产口径)在高并发占优——两套硬件上"低延迟 vs 高吞吐"的分野方向一致。 +- **分界点本机更靠前**:B300 的交叉点在 C=64~128(主场景 HT C=128 反超 33%);本机在 C=8 附近。原因不是算力而是**容量上限**:本机两臂 MRR/池上限(48/16)远小于 B300 配方的 256/默认,先于算力撞墙。 +- **绝对差距 4~5×(主场景)**:C=1 输出 246 vs 49.6 tok/s(5.0×)、input 7,882 vs 1,587(5.0×);吞吐侧峰值 997 vs 211(4.7×)、31,889 vs 6,748(4.7×)。与显存带宽硬件代差量级一致。 +- **边界 prefill 差距收窄到 ~2×**:256K C=1 input 7,050 vs 17,457(2.5×)→ 512K 5,662 vs 12,421(2.2×)→ 896K/约1M 4,153 vs 7,897(1.9×)。计算密集的超长 prefill 是 6000D 相对最能打的位置(PP 流水摊满 + 带宽占比下降)。 +- **TPOT 差距小于吞吐差距**:B300 LL C=1 4.36 ms vs E7b 13.7 ms(3.1×);高并发侧 B300 HT C=128 165 ms vs TP2PP4 C=64 204 ms(1.2×)——NVFP4 + DSA 把 decode 单步成本压得相对不差,差距主要在吞吐面。 +- **饱和形态不同**:B300 LL 在 C=64 后进入 24K input tok/s 平台、HT 在 C=128 达峰后 C=256 回退 19%;本机 TP2PP4 到 C=64 仍在爬坡(MRR 未饱和),E7b 则被 MRR16+掉图封死在 C=8。本机没有一档出现吞吐回退——"甜点=并发上限"由 MRR 决定而非算力。 +- **EAGLE 配方差异**:B300 LL 为 5 steps/6 draft tokens,本机 E7b 为 4 steps/topk1/5 draft tokens;本机实测 accept 2.0~4.0(短输出 2.0、主场景 2.1~3.0、长输出 3.7~4.0,随 decode 深入上升)。B300 未披露 accept,无法直接对比投机效率。 +- **容量边界差距最大**:B300 两模式都完成约 1M 输入 C=1/2/4;本机仅 TP2PP4 可达 896K 且 C=1 单条(池 1,040,384 刚容一条),E7b 连 512K 都结构性不可测(ctx 270,336)。96 GB 卡上"上下文边界=显存边界"比 B300 严酷得多。 + +## 附录 A:语料窗口映射(回收窗口) + +语料总量 21,296,780 tokens,此前场景一/二战役已消费至 21,235,008。冷缓存协议下回收复用:窗口基址 `--pool-override` 显式指定,每点窗口在基址上顺序推进(逐记录 `corpus_window.start/end` 留档),每点 flush + 命中核验 ≤0.01 保证冷。 + +| 场景 | 窗口基址 | 备注 | +|-|-|-| +| 主场景 16K→512 | 2,300,000 | | +| 4.1 短输入 1K→128 | 4,500,000 | | +| 4.2 长输出 1K→4K | 4,700,000 | | +| 5.1 64K→512 | 5,000,000 | | +| 5.1 128K→512 | 6,200,000 | | +| 5.2 256K→1 | 8,400,000 | E7b 干净重测改用 9,900,000(实例首用文本,避 hicache 宿主层残留) | +| 5.2 512K→1 | 9,300,000 | 仅 TP2PP4 | +| 5.2 896K→1 | 9,900,000 | 仅 TP2PP4 | + +## 附录 B:全量指标(含 mean/p95/max、回退、投机接受长度) + +| 场景点 | 方案 | ok/nreq | wall(s) | out tok/s | in tok/s | TTFT mean/p95/max(s) | TPOT mean/p95/max(ms) | retractions | accept | +|---|---|---|---|---|---|---|---|---|---| +| b3_16k_c1 | TP2PP4 | 8/8 | 245.55 | 16.68 | 533.8 | 5.35/5.36/5.36 | 49.6/50.3/50.3 | 0 | None | +| b3_16k_c1 | TP8+EAGLE3+AR | 8/8 | 82.58 | 49.6 | 1587.15 | 3.93/3.98/3.98 | 12.5/13.7/13.7 | 0 | 2.138 | +| b3_16k_c8 | TP2PP4 | 16/16 | 81.0 | 101.13 | 3236.22 | 10.05/14.79/14.79 | 59.5/70.9/70.9 | 0 | None | +| b3_16k_c8 | TP8+EAGLE3+AR | 16/16 | 100.63 | 81.41 | 2605.13 | 11.75/32.52/32.52 | 71.6/135.1/135.1 | 0 | 2.375 | +| b3_16k_c16 | TP2PP4 | 32/32 | 112.16 | 146.08 | 4674.5 | 15.53/25.78/25.84 | 79.2/98.5/101.2 | 0 | None | +| b3_16k_c16 | TP8+EAGLE3+AR | 32/32 | 236.44 | 69.29 | 2217.4 | 23.60/60.86/90.39 | 173.3/297.8/305.5 | 0 | 2.517 | +| b3_16k_c32 | TP2PP4 | 64/64 | 169.54 | 193.27 | 6184.73 | 26.52/46.43/47.89 | 113.7/152.0/157.3 | 0 | None | +| b3_16k_c32 | TP8+EAGLE3+AR | 64/64 | 414.96 | 78.97 | 2526.91 | 100.26/169.68/190.73 | 169.8/259.9/319.1 | 0 | 2.767 | +| b3_16k_c64 | TP2PP4 | 128/128 | 310.79 | 210.87 | 6747.91 | 63.05/134.78/139.99 | 139.2/203.7/214.3 | 0 | None | +| b3_16k_c64 | TP8+EAGLE3+AR | 128/128 | 777.55 | 84.29 | 2697.14 | 240.10/345.69/380.96 | 168.0/231.5/311.9 | 0 | 2.986 | +| b41_1k_c1 | TP2PP4 | 8/8 | 52.95 | 19.34 | 154.7 | 0.38/0.39/0.39 | 49.1/49.3/49.3 | 0 | None | +| b41_1k_c1 | TP8+EAGLE3+AR | 8/8 | 15.82 | 64.74 | 517.93 | 0.30/0.33/0.33 | 13.2/14.3/14.3 | 0 | 2.028 | +| b41_1k_c8 | TP2PP4 | 16/16 | 20.1 | 101.91 | 815.3 | 1.31/1.69/1.69 | 68.6/72.9/72.9 | 0 | None | +| b41_1k_c8 | TP8+EAGLE3+AR | 16/16 | 13.44 | 152.4 | 1219.24 | 1.21/2.16/2.16 | 41.4/53.7/53.7 | 0 | 2.048 | +| b41_1k_c32 | TP2PP4 | 64/64 | 29.73 | 275.55 | 2204.4 | 3.89/4.87/4.88 | 85.6/98.5/106.8 | 0 | None | +| b41_1k_c32 | TP8+EAGLE3+AR | 64/64 | 81.8 | 100.14 | 801.15 | 17.56/25.60/27.89 | 146.6/190.5/195.2 | 0 | 2.054 | +| b41_1k_c64 | TP2PP4 | 128/128 | 48.11 | 340.53 | 2724.25 | 8.14/19.96/20.01 | 93.3/101.8/125.1 | 0 | None | +| b41_1k_c64 | TP8+EAGLE3+AR | 128/128 | 158.73 | 103.22 | 825.75 | 46.74/65.81/70.71 | 144.5/177.7/212.5 | 0 | 2.087 | +| b42_1k4k_c1 | TP2PP4 | 8/8 | 1614.16 | 20.3 | 5.08 | 0.39/0.40/0.40 | 49.2/49.8/49.8 | 0 | None | +| b42_1k4k_c1 | TP8+EAGLE3+AR | 8/8 | 242.18 | 135.31 | 33.83 | 0.31/0.32/0.32 | 7.3/8.7/8.7 | 0 | 3.701 | +| b42_1k4k_c8 | TP2PP4 | 16/16 | 649.69 | 100.87 | 25.22 | 1.35/1.87/1.87 | 79.0/79.4/79.4 | 0 | None | +| b42_1k4k_c8 | TP8+EAGLE3+AR | 16/16 | 159.24 | 411.54 | 102.89 | 0.96/1.68/1.68 | 18.1/22.5/22.5 | 0 | 4.003 | +| b42_1k4k_c32 | TP2PP4 | 64/64 | 714.14 | 367.08 | 91.77 | 1.98/4.14/4.15 | 86.1/90.5/91.1 | 0 | None | +| b42_1k4k_c32 | TP8+EAGLE3+AR | 64/64 | 1040.79 | 251.87 | 62.97 | 195.07/304.75/334.69 | 60.7/74.6/92.3 | 0 | 3.852 | +| b42_1k4k_c64 | TP2PP4 | 128/128 | 1088.63 | 481.6 | 120.4 | 80.74/337.92/338.47 | 91.2/95.5/96.3 | 0 | None | +| b51_64k_c1 | TP2PP4 | 8/8 | 285.91 | 14.33 | 1833.72 | 10.28/10.32/10.32 | 49.8/50.0/50.0 | 0 | None | +| b51_64k_c1 | TP8+EAGLE3+AR | 8/8 | 182.23 | 22.48 | 2877.06 | 17.04/17.17/17.17 | 11.2/13.5/13.5 | 0 | 2.468 | +| b51_64k_c4 | TP2PP4 | 8/8 | 122.3 | 33.49 | 4286.94 | 19.55/28.83/28.83 | 81.4/99.5/99.5 | 0 | None | +| b51_64k_c4 | TP8+EAGLE3+AR | 8/8 | 159.24 | 25.72 | 3292.45 | 30.55/68.87/68.87 | 94.9/154.6/154.6 | 0 | 2.438 | +| b51_64k_c8 | TP2PP4 | 16/16 | 185.59 | 44.14 | 5650.11 | 31.83/53.38/53.38 | 119.2/161.8/161.8 | 0 | None | +| b51_64k_c8 | TP8+EAGLE3+AR | 16/16 | 319.66 | 25.63 | 3280.31 | 90.92/149.45/149.45 | 105.8/157.5/157.5 | 0 | 2.442 | +| b51_128k_c1 | TP2PP4 | 8/8 | 347.5 | 11.79 | 3017.45 | 18.21/18.23/18.23 | 49.4/49.7/49.7 | 0 | None | +| b51_128k_c1 | TP8+EAGLE3+AR | 8/8 | 343.36 | 11.93 | 3053.83 | 37.59/37.60/37.60 | 10.4/12.3/12.3 | 0 | 2.728 | +| b51_128k_c2 | TP2PP4 | 8/8 | 243.3 | 16.84 | 4309.84 | 25.27/32.42/32.42 | 69.6/83.7/83.7 | 0 | None | +| b51_128k_c2 | TP8+EAGLE3+AR | 8/8 | 333.36 | 12.29 | 3145.44 | 48.21/75.81/75.81 | 68.0/162.2/162.2 | 0 | 2.769 | +| b51_128k_c4 | TP2PP4 | 8/8 | 186.37 | 21.98 | 5626.35 | 39.29/60.35/60.35 | 105.4/146.8/146.8 | 0 | None | +| b51_128k_c4 | TP8+EAGLE3+AR | 8/8 | 334.71 | 12.24 | 3132.76 | 110.14/159.50/159.50 | 68.7/163.4/163.4 | 0 | 2.595 | +| b52_256k_c1 | TP2PP4 | 1/1 | 37.18 | 0.03 | 7050.23 | 37.18/37.18/37.18 | — | 0 | None | +| b52_256k_c1 | TP8+EAGLE3+AR | 1/1 | 91.45 | 0.01 | 2866.59 | 91.45/91.45/91.45 | — | 0 | None | +| b52_256k_c2 | TP2PP4 | 2/2 | 70.19 | 0.03 | 7469.22 | 53.69/70.12/70.12 | — | 0 | None | +| b52_256k_c3 | TP2PP4 | 3/3 | 103.12 | 0.03 | 7626.22 | 70.12/102.99/102.99 | — | 0 | None | +| b52_512k_c1 | TP2PP4 | 1/1 | 92.6 | 0.01 | 5662.03 | 92.60/92.60/92.60 | — | 0 | None | +| b52_896k_c1 | TP2PP4 | 1/1 | 220.91 | 0.0 | 4153.24 | 220.91/220.91/220.91 | — | 0 | None | + +注:E7b 4.2 C=64 用户中止(rc=143)无 SUMMARY,不在表内;E7b 256K/512K/896K C>1 为结构性跳过;E7b 256K C=1 为全新文本干净重测值(窗口 9,900,000)。 diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/patches/custom_all_reduce.py b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/patches/custom_all_reduce.py new file mode 100644 index 0000000..4ec6995 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/patches/custom_all_reduce.py @@ -0,0 +1,424 @@ +# SPDX-License-Identifier: Apache-2.0 +# SPDX-FileCopyrightText: Copyright contributors to the vLLM project +# Adapted from https://github.com/vllm-project/vllm/blob/v0.6.4.post1/vllm/distributed/device_communicators/custom_all_reduce.py + +import ctypes +import logging +import os +from contextlib import contextmanager +from functools import partial +from typing import Any, List, Optional, Union + +import torch +import torch.distributed as dist +from torch.distributed import ProcessGroup + +import sglang.srt.distributed.device_communicators.custom_all_reduce_ops as ops +from sglang.srt.distributed.device_communicators.cuda_wrapper import CudaRTLibrary +from sglang.srt.distributed.device_communicators.custom_all_reduce_utils import ( + can_use_custom_all_reduce_with_nvlink, + is_weak_contiguous, +) +from sglang.srt.environ import envs +from sglang.srt.model_executor.runner_backend_utils.tc_piecewise_cuda_graph import ( + is_in_tc_piecewise_cuda_graph, +) +from sglang.srt.utils import ( + get_bool_env_var, + is_cuda, + is_hip, + is_musa, + log_info_on_rank0, +) + +_is_cuda = is_cuda() +_is_hip = is_hip() +_is_musa = is_musa() + +logger = logging.getLogger(__name__) +os.environ.setdefault("SGLANG_CUSTOM_ALLREDUCE_ALGO", "1stage") # SSKJ-PATCH: C++ dispatch no-ops when full_nvlink=False; force kernel launch + + +class CustomAllreduce: + _SUPPORTED_WORLD_SIZES = [2, 4, 6, 8] + _MAX_CAR_SIZE = 8192 * 1024 + if _is_hip: + # crossover is at 16MB buffer size for ROCm + _MAX_CAR_SIZE = 2 * 8192 * 1024 + if _is_musa: + # crossover is at 128MB buffer size for MUSA + _MAX_CAR_SIZE = 16 * 8196 * 1024 + + # max_size: max supported allreduce size + def __init__( + self, + group: ProcessGroup, + device: Union[int, str, torch.device], + max_size=_MAX_CAR_SIZE, + ) -> None: + """ + Args: + group: the process group to work on. If None, it will use the + default process group. + device: the device to bind the CustomAllreduce to. If None, + it will be bind to f"cuda:{local_rank}". + It is the caller's responsibility to make sure each communicator + is bind to a unique device, and all communicators in this group + are in the same node. + """ + self._IS_CAPTURING = False + self.disabled = True # This can be modified in-place by context manager in piecewise cuda graph runner + self.original_disabled = True # To store the original state + self.use_amd_deterministic_impl = _use_amd_deterministic_impl() + + if not ops.IS_CUSTOM_AR_AVAILABLE: + # disable because of missing custom allreduce library + # e.g. in a non-cuda environment + return + + rank = dist.get_rank(group=group) + world_size = dist.get_world_size(group=group) + + if isinstance(device, int): + device = torch.device(f"cuda:{device}") + elif isinstance(device, str): + device = torch.device(device) + # now `device` is a `torch.device` object + assert isinstance(device, torch.device) + self.device = device + full_nvlink = can_use_custom_all_reduce_with_nvlink( + group=group, + device=device, + supported_world_size=self._SUPPORTED_WORLD_SIZES, + cls_name="CustomAllreduce", + ) + if full_nvlink is None: + return # fail to get nvlink status + + self.group = group + self.max_size = max_size + self.rank = rank + self.world_size = world_size + self.full_nvlink = full_nvlink + + if not _is_hip: + # Buffers memory are owned by this Python class and passed to C++. + # Meta data composes of two parts: meta data for synchronization and a + # temporary buffer for storing intermediate allreduce results. + self.meta_ptrs = self.create_shared_buffer( + ops.meta_size() + max_size, group=group + ) + # This is a pre-registered IPC buffer. In eager mode, input tensors + # are first copied into this buffer before allreduce is performed + self.buffer_ptrs = self.create_shared_buffer(max_size, group=group) + # This is a buffer for storing the tuples of pointers pointing to + # IPC buffers from all ranks. Each registered tuple has size of + # 8*world_size bytes where world_size is at most 8. Allocating 8MB + # is enough for 131072 such tuples. The largest model I've seen only + # needs less than 10000 of registered tuples. + self.rank_data = torch.empty( + max_size, dtype=torch.uint8, device=self.device + ) + self._ptr = ops.init_custom_ar( + self.meta_ptrs, self.rank_data, rank, self.full_nvlink + ) + ops.register_buffer(self._ptr, self.buffer_ptrs) + else: + # meta data buffers need to be "uncached" for signal on MI200 + self.meta = ops.allocate_meta_buffer(ops.meta_size() + max_size) + self.buffer = torch.empty(max_size, dtype=torch.uint8, device=self.device) + handle = ops.get_meta_buffer_ipc_handle(self.meta) + shard_data = ( + bytes(handle), # ipc handle to base ptr + 0, # offset of base ptr + ) + handles, offsets = self._gather_ipc_meta(shard_data) + self.rank_data = torch.empty( + max_size, dtype=torch.uint8, device=self.device + ) + self._ptr = ops.init_custom_ar( + self.meta, self.rank_data, handles, offsets, rank, self.full_nvlink + ) + self.register_buffer(self.buffer) + + self.disabled = False + self.original_disabled = False # Ensure original_disabled == disabled + logger.warning(f"SSKJ_CAR_PATCH_ACTIVE ws={self.world_size} full_nvlink={self.full_nvlink}") + self.tms_cudagraph = envs.SGLANG_MEMORY_SAVER_CUDA_GRAPH.get() + + @staticmethod + def create_shared_buffer( + size_in_bytes: int, group: Optional[ProcessGroup] = None + ) -> List[int]: + """ + Creates a shared buffer and returns a list of pointers + representing the buffer on all processes in the group. + """ + lib = CudaRTLibrary() + pointer = lib.cudaMalloc(size_in_bytes) + if _is_musa: + lib.cudaMemset(pointer, 0, size_in_bytes) + handle = lib.cudaIpcGetMemHandle(pointer) + world_size = dist.get_world_size(group=group) + rank = dist.get_rank(group=group) + handles = [None] * world_size + dist.all_gather_object(handles, handle, group=group) + + pointers: List[int] = [] + for i, h in enumerate(handles): + if i == rank: + pointers.append(pointer.value) # type: ignore + else: + pointers.append(lib.cudaIpcOpenMemHandle(h).value) # type: ignore + + return pointers + + @staticmethod + def free_shared_buffer( + pointers: List[int], group: Optional[ProcessGroup] = None + ) -> None: + rank = dist.get_rank(group=group) + lib = CudaRTLibrary() + lib.cudaFree(ctypes.c_void_p(pointers[rank])) + + @contextmanager + def capture(self): + """ + The main responsibility of this context manager is the + `register_graph_buffers` call at the end of the context. + It records all the buffer addresses used in the CUDA graph. + """ + try: + self._IS_CAPTURING = True + yield + finally: + self._IS_CAPTURING = False + if not self.disabled: + self.register_graph_buffers() + + def _get_ipc_meta(self, inp: torch.Tensor): + # _share_cuda_() doesn't accept meta buffer not allocated from + # PyTorch cache allocator, use direct HIP call to get IPC handle + handle = ops.get_meta_buffer_ipc_handle(inp) + shard_data = ( + bytes(handle), # ipc handle to base ptr + 0, # offset of base ptr + ) + return self._gather_ipc_meta(shard_data) + + def _gather_ipc_meta(self, shard_data): + # Note: don't use `[[None]] * self.world_size` here + # because it will create a list of the same reference + all_data: List[Optional[Any]] = [[None] for i in range(self.world_size)] + all_data[self.rank][0] = shard_data + + ranks = dist.get_process_group_ranks(group=self.group) + ranks.sort() + for i, rank in enumerate(ranks): + dist.broadcast_object_list( + all_data[i], src=rank, group=self.group, device="cpu" + ) + + # we cannot directly use `dist.all_gather_object` here + # because it is incompatible with `gloo` backend under inference mode. + # see https://github.com/pytorch/pytorch/issues/126032 for details. + + handles = [] + offsets = [] + for i in range(len(all_data)): + handles.append(all_data[i][0][0]) # type: ignore + offsets.append(all_data[i][0][1]) # type: ignore + return handles, offsets + + def register_buffer(self, inp: torch.Tensor): + handles, offsets = self._get_ipc_meta(inp) + ops.register_buffer(self._ptr, inp, handles, offsets) + + def register_graph_buffers(self): + if _is_hip: + handle, offset = ops.get_graph_buffer_ipc_meta(self._ptr) + handles, offsets = self._gather_ipc_meta((bytes(handle), offset)) + log_info_on_rank0(logger, f"Registering {len(offset)} cuda graph addresses") + ops.register_graph_buffers(self._ptr, handles, offsets) + else: + handle, offset = ops.get_graph_buffer_ipc_meta(self._ptr) + log_info_on_rank0(logger, f"Registering {len(offset)} cuda graph addresses") + # We cannot directly use `dist.all_gather_object` here + # because it is incompatible with `gloo` backend under inference mode. + # see https://github.com/pytorch/pytorch/issues/126032 for details. + all_data = [ + [None, None] for _ in range(dist.get_world_size(group=self.group)) + ] + all_data[self.rank] = [handle, offset] + ranks = sorted(dist.get_process_group_ranks(group=self.group)) + for i, rank in enumerate(ranks): + dist.broadcast_object_list( + all_data[i], src=rank, group=self.group, device="cpu" + ) + # Unpack list of tuples to tuple of lists. + handles = [d[0] for d in all_data] # type: ignore + offsets = [d[1] for d in all_data] # type: ignore + ops.register_graph_buffers(self._ptr, handles, offsets) + + def should_custom_ar(self, inp: torch.Tensor): + if self.disabled: + return False + inp_size = inp.numel() * inp.element_size() + # custom allreduce requires input byte size to be multiples of 16 + if inp_size % 16 != 0: + return False + if not is_weak_contiguous(inp): + return False + # for 4 or more non NVLink-capable GPUs, custom allreduce provides + # little performance improvement over NCCL. + if not _is_hip: + if True: + return inp_size <= self.max_size + return False + + if _is_hip: + if self.use_amd_deterministic_impl: + return True + if self.full_nvlink: + return inp_size <= self.max_size + return False + + return False + + def _all_reduce_impl(self, inp: torch.Tensor, registered: bool): + out = torch.empty_like(inp) + if not _is_hip: # CUDA-like + if registered: + ops.all_reduce(self._ptr, inp, out, 0, 0) + else: + ops.all_reduce( + self._ptr, inp, out, self.buffer_ptrs[self.rank], self.max_size + ) + elif self.use_amd_deterministic_impl: + inp_size = inp.numel() * inp.element_size() + if inp_size < self.max_size: + reg_buffer = self.buffer.view(inp.dtype)[: inp.numel()] + ops.deterministic_all_reduce_unreg(self._ptr, inp, reg_buffer, out) + else: + self.register_buffer(inp) + ops.deterministic_all_reduce_reg(self._ptr, inp, out) + else: # normal AMD ROCm path + if registered: + ops.all_reduce_reg(self._ptr, inp, out) + else: + ops.all_reduce_unreg(self._ptr, inp, self.buffer, out) + return out + + def custom_all_reduce(self, input: torch.Tensor) -> Optional[torch.Tensor]: + """The main allreduce API that provides support for cuda graph.""" + # When custom allreduce is disabled, this will be None. + if self.disabled or not self.should_custom_ar(input): + return None + if self._IS_CAPTURING: + if torch.cuda.is_current_stream_capturing(): + return self._all_reduce_impl(input, registered=not self.tms_cudagraph) + else: + # Could be warmup OR piecewise cuda graph split op execution. + # In piecewise cuda graph, split ops run eagerly outside the graph + # but _IS_CAPTURING is still True. We need to do real all-reduce. + if is_in_tc_piecewise_cuda_graph(): + # Split op execution - do real all-reduce + return self._all_reduce_impl(input, registered=False) + else: + # True warmup - mimic the allocation pattern since custom + # allreduce is out-of-place. + return torch.zeros_like(input) + else: + return self._all_reduce_impl(input, registered=False) + + def close(self): + if not self.disabled and self._ptr: + if ops is not None: + ops.dispose(self._ptr) + if _is_cuda: + self.free_shared_buffer(self.meta_ptrs) + self.free_shared_buffer(self.buffer_ptrs) + self._ptr = 0 + + def __del__(self): + self.close() + + +def dispatch_custom_allreduce( + group: ProcessGroup, + device: torch.device, +): + """Return the CustomAllreduce class to use (aiter on ROCm if enabled). + + On AMD with 1-stage AR enabled, use sglang's CustomAllreduce. + Otherwise use AiterCustomAllreduce if available. + + On CUDA, the JIT-compiled v2 implementation is used by default. + Set SGLANG_OPT_USE_CUSTOM_ALL_REDUCE_V2=0 to fall back to the legacy CustomAllreduce. + Multi-node v2 is admitted only for a single NVLink clique (see + can_use_custom_all_reduce_v2); other cross-node groups fall back to NCCL. + """ + if _is_cuda and envs.SGLANG_OPT_USE_CUSTOM_ALL_REDUCE_V2.get(): + from .custom_all_reduce_v2 import ( + CustomAllReduceV2, + can_use_custom_all_reduce_v2, + ) + + if can_use_custom_all_reduce_v2(group=group, device=device): + logger.debug("[AR] Using CustomAllReduceV2 (JIT-compiled)") + return CustomAllReduceV2 + + if _is_cuda or _is_musa: + return CustomAllreduce + + assert _is_hip + + if envs.SGLANG_USE_1STAGE_ALLREDUCE.is_set(): + if envs.SGLANG_USE_1STAGE_ALLREDUCE.get(): + logger.debug( + "[AR] All-reduce: 1-stage kernel (SGLANG_USE_1STAGE_ALLREDUCE=1)" + ) + else: + logger.debug("[AR] All-reduce: default (SGLANG_USE_1STAGE_ALLREDUCE=0)") + elif envs.SGLANG_ENABLE_DETERMINISTIC_INFERENCE.get(): + logger.debug( + "[AR] All-reduce: 1-stage kernel (deterministic inference enabled)" + ) + else: + logger.debug("[AR] All-reduce: default") + + # On AMD with 1-stage AR, use sglang's CustomAllreduce + # (AiterCustomAllreduce doesn't have deterministic_all_reduce method) + if _use_amd_deterministic_impl(): + return CustomAllreduce + + if get_bool_env_var("SGLANG_USE_AITER_AR", default="true"): + try: + from aiter.dist.device_communicators.custom_all_reduce import ( + CustomAllreduce as AiterCustomAllreduce, + ) + + logger.info("[AR] Using AiterCustomAllreduce (AMD default)") + tms_cudagraph = envs.SGLANG_MEMORY_SAVER_CUDA_GRAPH.get() + return partial( + AiterCustomAllreduce, + enable_register_for_capturing=not tms_cudagraph, + ) + except ImportError as e: + logger.warning( + "[AR] Aiter custom all-reduce not available; " + "falling back to sglang CustomAllreduce. Details: %s", + e, + ) + return CustomAllreduce + + return CustomAllreduce + + +def _use_amd_deterministic_impl() -> bool: + if not _is_hip: # CUDA is always deterministic + return False + if envs.SGLANG_USE_1STAGE_ALLREDUCE.is_set(): + return envs.SGLANG_USE_1STAGE_ALLREDUCE.get() + else: + return envs.SGLANG_ENABLE_DETERMINISTIC_INFERENCE.get() diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/patches/custom_all_reduce_utils.py b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/patches/custom_all_reduce_utils.py new file mode 100644 index 0000000..70a70de --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/patches/custom_all_reduce_utils.py @@ -0,0 +1,519 @@ +# SPDX-License-Identifier: Apache-2.0 +# SPDX-FileCopyrightText: Copyright contributors to the vLLM project +# Adapted from https://github.com/vllm-project/vllm/blob/v0.6.4.post1/vllm/distributed/device_communicators/custom_all_reduce_utils.py + +import ctypes +import json +import logging +import os +import pickle +import subprocess +import sys +import tempfile +from functools import wraps +from itertools import product +from typing import Callable, Dict, List, Optional, Sequence, TypeVar + +import torch +import torch.distributed as dist +import torch.multiprocessing as mp +from typing_extensions import ParamSpec + +from sglang.srt.distributed.device_communicators.cuda_wrapper import CudaRTLibrary +from sglang.srt.distributed.parallel_state import in_the_same_node_as +from sglang.srt.environ import envs as sglang_envs +from sglang.srt.utils import is_cuda, is_hip, is_musa +from sglang.srt.utils.cuda_vmm_utils import _gpu_fabric_clique + +logger = logging.getLogger(__name__) + +_is_cuda = is_cuda() +_is_hip = is_hip() +_is_musa = is_musa() + +if _is_cuda: + try: + import pynvml + except ImportError as e: + logger.warning("Failed to import pynvml with %r", e) + +if _is_musa: + try: + import pymtml as pynvml + except ImportError as e: + logger.warning("Failed to import pymtml with %r", e) + +if _is_hip: + try: + from amdsmi import ( + AmdSmiException, + amdsmi_get_processor_handles, + amdsmi_init, + amdsmi_shut_down, + amdsmi_topo_get_link_type, + ) + except ImportError as e: + logger.warning("Failed to import amdsmi with %r", e) + +_P = ParamSpec("_P") +_R = TypeVar("_R") + + +def update_environment_variables(envs: Dict[str, str]): + for k, v in envs.items(): + if k in os.environ and os.environ[k] != v: + logger.warning( + "Overwriting environment variable %s " "from '%s' to '%s'", + k, + os.environ[k], + v, + ) + os.environ[k] = v + + +def producer( + batch_src: Sequence[int], + producer_queue, + consumer_queue, + result_queue, + cuda_visible_devices: Optional[str] = None, +): + if cuda_visible_devices is not None: + update_environment_variables({"CUDA_VISIBLE_DEVICES": cuda_visible_devices}) + + lib = CudaRTLibrary() + for i in batch_src: + lib.cudaSetDevice(i) + pointer = lib.cudaMalloc(1024) + lib.cudaMemset(pointer, 1, 1024) + lib.cudaDeviceSynchronize() + handle = lib.cudaIpcGetMemHandle(pointer) + producer_queue.put(handle) + open_success = consumer_queue.get() + if open_success: + # use two queues to simulate barrier + producer_queue.put(0) + consumer_queue.get() + # check if the memory is modified + host_data = (ctypes.c_char * 1024)() + lib.cudaMemcpy(host_data, pointer, 1024) # type: ignore + for i in range(1024): + if ord(host_data[i]) != 2: + open_success = False + break + result_queue.put(open_success) + lib.cudaDeviceReset() + + +def consumer( + batch_tgt: Sequence[int], + producer_queue, + consumer_queue, + result_queue, + cuda_visible_devices: Optional[str] = None, +): + if cuda_visible_devices is not None: + update_environment_variables({"CUDA_VISIBLE_DEVICES": cuda_visible_devices}) + + lib = CudaRTLibrary() + for j in batch_tgt: + lib.cudaSetDevice(j) + handle = producer_queue.get() + open_success = False + try: + pointer = lib.cudaIpcOpenMemHandle(handle) # type: ignore + open_success = True + except RuntimeError: + # cannot error out here, because the producer process + # is still waiting for the response. + pass + consumer_queue.put(open_success) + if open_success: + # modify the memory + lib.cudaMemset(pointer, 2, 1024) + lib.cudaDeviceSynchronize() + # use two queues to simulate barrier + producer_queue.get() + consumer_queue.put(0) + # check if the memory is modified + host_data = (ctypes.c_char * 1024)() + lib.cudaMemcpy(host_data, pointer, 1024) # type: ignore + for i in range(1024): + if ord(host_data[i]) != 2: + open_success = False + break + result_queue.put(open_success) + lib.cudaDeviceReset() + + +def can_actually_p2p( + batch_src: Sequence[int], + batch_tgt: Sequence[int], +) -> Sequence[bool]: + """ + Usually, checking if P2P access is enabled can be done by + `torch.cuda.can_device_access_peer(src, tgt)`. However, sometimes + the driver might be broken, and `torch.cuda.can_device_access_peer(src, tgt)` + returns `True` even if P2P access is not actually possible. + See https://github.com/vllm-project/vllm/issues/2728 and + https://forums.developer.nvidia.com/t/direct-gpu-gpu-communication-does-not-seem-to-work-properly/283264/10 + Therefore, we have to perform a real P2P access to check if it is actually + possible. + + Note on p2p and cuda IPC: + Usually, one process uses one GPU: + GPU src --> cuda context src --> tensor src --> process src + + We need to combine p2p and cuda IPC, so that: + GPU src --> cuda context src --> tensor src --> process src + |shared| + GPU tgt --> cuda context tgt --> tensor tgt --> process tgt + That is to say, process src creates a tensor in GPU src, passes IPC handle to + process tgt, and process tgt accesses the tensor in GPU tgt. Any operation on the + tensor in process tgt will be reflected in the tensor in process src, because + they are the same memory segment. + It is important to note that process tgt accesses the tensor in GPU tgt, not + GPU src. That's why we need p2p access. + + The most time-consuming part is the process creation. To avoid creating + processes for every pair of GPUs, we use batched testing. We create two + processes for testing all pairs of GPUs in batch. The trick is to reset + the device after each test (which is not available in PyTorch). + """ # noqa + cuda_visible_devices = os.environ.get("CUDA_VISIBLE_DEVICES", None) + # pass the CUDA_VISIBLE_DEVICES to the child process + # to make sure they see the same set of GPUs + + # make sure the processes are spawned + smp = mp.get_context("spawn") + producer_queue = smp.Queue() + consumer_queue = smp.Queue() + result_queue = smp.Queue() + p_src = smp.Process( + target=producer, + args=( + batch_src, + producer_queue, + consumer_queue, + result_queue, + cuda_visible_devices, + ), + ) + p_tgt = smp.Process( + target=consumer, + args=( + batch_tgt, + producer_queue, + consumer_queue, + result_queue, + cuda_visible_devices, + ), + ) + p_src.start() + p_tgt.start() + p_src.join() + p_tgt.join() + assert p_src.exitcode == 0 and p_tgt.exitcode == 0 + result: List[bool] = [] + for src, tgt in zip(batch_src, batch_tgt): + a = result_queue.get() + b = result_queue.get() + if a != b: + logger.warning( + "Two processes do not agree on the P2P access" + " status on %d -> %d, treat as disabled.", + src, + tgt, + ) + result.append(False) + else: + result.append(a) + return result + + +# why do we need this cache? +# we are testing peer-to-peer (p2p) access between GPUs,across processes. +# if we test it every time, it will be very slow, because we need to create +# N * N * 2 processes, where N is the world size. This is very slow. +# to reduce the time, we use a cache file to store the p2p access status. +# the cache file is generated by the master process if it does not exist. +# then all the processes can read the cache file to check the p2p access status. +# Note that the cache file is suffixed by the CUDA_VISIBLE_DEVICES, so that we +# can have different cache files for different CUDA_VISIBLE_DEVICES settings, +# e.g. used by different vllm engines. The device id in the cache file is a +# **local** device id, i.e. from 0 to num_dev-1, where num_dev is the number +# of visible devices in the vllm engine. +_gpu_p2p_access_cache: Optional[Dict[str, bool]] = None + + +def gpu_p2p_access_check(src: int, tgt: int) -> bool: + """Check if GPU src can access GPU tgt.""" + + # if the cache variable is already calculated, + # read from the cache instead of checking it again + global _gpu_p2p_access_cache + if _gpu_p2p_access_cache is not None: + return _gpu_p2p_access_cache[f"{src}->{tgt}"] + + is_distributed = dist.is_initialized() + + num_dev = torch.cuda.device_count() + cuda_visible_devices = os.environ.get("CUDA_VISIBLE_DEVICES", None) + if cuda_visible_devices is None: + cuda_visible_devices = ",".join(str(i) for i in range(num_dev)) + + # VLLM_CACHE_ROOT -> SGLANG_CACHE_ROOT + # "~/.cache/vllm" -> envs.SGLANG_CACHE_DIR + SGLANG_CACHE_ROOT = os.path.expanduser(sglang_envs.SGLANG_CACHE_DIR.get()) + path = os.path.join( + SGLANG_CACHE_ROOT, f"gpu_p2p_access_cache_for_{cuda_visible_devices}.json" + ) + cache_dir = os.path.dirname(path) + try: + os.makedirs(cache_dir, exist_ok=True) + except (FileExistsError, NotADirectoryError): + if not os.path.isdir(cache_dir): + # Path exists as a file (stale cache/lock). Remove and retry. + try: + os.remove(cache_dir) + except OSError: + pass + os.makedirs(cache_dir, exist_ok=True) + from sglang.srt.distributed.parallel_state import get_world_group + + if (not is_distributed or get_world_group().local_rank == 0) and ( + not os.path.exists(path) + ): + # only the local master process (with local_rank == 0) can + # enter this block to calculate the cache + logger.info("generating GPU P2P access cache in %s", path) + cache: Dict[str, bool] = {} + ids = list(range(num_dev)) + # batch of all pairs of GPUs + batch_src, batch_tgt = zip(*list(product(ids, ids))) + # NOTE: we use `subprocess` rather than `multiprocessing` here + # because the caller might not have `if __name__ == "__main__":`, + # in that case we cannot use spawn method in multiprocessing. + # However, `can_actually_p2p` requires spawn method. + # The fix is, we use `subprocess` to call the function, + # where we have `if __name__ == "__main__":` in this file. + + # use a temporary file to store the result + # we don't use the output of the subprocess directly, + # because the subprocess might produce logging output + with tempfile.NamedTemporaryFile() as output_file: + input_bytes = pickle.dumps((batch_src, batch_tgt, output_file.name)) + returned = subprocess.run( + [sys.executable, __file__], input=input_bytes, capture_output=True + ) + # check if the subprocess is successful + try: + returned.check_returncode() + except Exception as e: + # wrap raised exception to provide more information + raise RuntimeError( + f"Error happened when batch testing " + f"peer-to-peer access from {batch_src} to {batch_tgt}:\n" + f"{returned.stderr.decode()}" + ) from e + with open(output_file.name, "rb") as f: + result = pickle.load(f) + for _i, _j, r in zip(batch_src, batch_tgt, result): + cache[f"{_i}->{_j}"] = r + with open(path, "w") as f: + json.dump(cache, f, indent=4) + if is_distributed: + get_world_group().barrier() + logger.info("reading GPU P2P access cache from %s", path) + with open(path) as f: + cache = json.load(f) + _gpu_p2p_access_cache = cache + return _gpu_p2p_access_cache[f"{src}->{tgt}"] + + +def with_nvml_context(fn: Callable[_P, _R]) -> Callable[_P, _R]: + @wraps(fn) + def wrapper(*args: _P.args, **kwargs: _P.kwargs) -> _R: + if _is_hip: + try: + amdsmi_init() + return fn(*args, **kwargs) + finally: + amdsmi_shut_down() + else: + pynvml.nvmlInit() + try: + return fn(*args, **kwargs) + finally: + pynvml.nvmlShutdown() + + return wrapper + + +@with_nvml_context +def is_full_nvlink(physical_device_ids: List[int], world_size: int) -> bool: + if _is_hip: + """ + query if the set of gpus are fully connected by xgmi (1 hop) + """ + handles = [amdsmi_get_processor_handles()[i] for i in physical_device_ids] + for i, handle in enumerate(handles): + for j, peer_handle in enumerate(handles): + if i < j: + try: + link_type = amdsmi_topo_get_link_type(handle, peer_handle) + # type is 2 for XGMI + if link_type["hops"] != 1 or link_type["type"] != 2: + return False + except AmdSmiException as error: + logger.error("AMD 1 hop XGMI detection failed.", exc_info=error) + return False + return True + else: + """ + query if the set of gpus are fully connected by nvlink (1 hop) + """ + handles = [pynvml.nvmlDeviceGetHandleByIndex(i) for i in physical_device_ids] + for i, handle in enumerate(handles): + for j, peer_handle in enumerate(handles): + if i < j: + try: + p2p_status = pynvml.nvmlDeviceGetP2PStatus( + handle, peer_handle, pynvml.NVML_P2P_CAPS_INDEX_NVLINK + ) + if p2p_status != pynvml.NVML_P2P_STATUS_OK: + return False + except pynvml.NVMLError: + logger.exception( + "NVLink detection failed. This is normal if your" + " machine has no NVLink equipped." + ) + return False + return True + + +@with_nvml_context +def is_one_nvlink_clique( + group: torch.distributed.ProcessGroup, device: torch.device +) -> bool: + """True iff every rank's GPU is in the same NVLink fabric clique (one NVL72 / + MNNVL domain). Such a clique shares a single NVLink address space even across + nodes, so custom-AR v2's symm-mem storage + fabric peer VAs are valid group-wide.""" + if _is_hip: + return False + try: + clique = _gpu_fabric_clique(device) + except Exception as e: + logger.warning( + "GPU fabric clique query failed (%r); custom-AR stays intra-node.", e + ) + clique = None + # Always all-gather (every rank calls it once) so a failed query on any rank + # resolves to a clean False rather than a collective mismatch. + world_size = dist.get_world_size(group=group) + gathered: List[object] = [None] * world_size + dist.all_gather_object(gathered, clique, group=group) + if any(c is None for c in gathered): + return False + return len(set(gathered)) == 1 + + +def is_weak_contiguous(inp: torch.Tensor): + return inp.is_contiguous() or ( + inp.storage().nbytes() - inp.storage_offset() * inp.element_size() + == inp.numel() * inp.element_size() + ) + + +def can_p2p(rank: int, world_size: int) -> bool: + # SGLANG_SKIP_P2P_CHECK can be set to False in sglang + SGLANG_SKIP_P2P_CHECK = os.getenv("SGLANG_SKIP_P2P_CHECK", "0") == "1" + for i in range(world_size): + if i == rank: + continue + if SGLANG_SKIP_P2P_CHECK: + logger.info("Skipping P2P check and trusting the driver's P2P report.") + return torch.cuda.can_device_access_peer(rank, i) + if not gpu_p2p_access_check(rank, i): + return False + return True + + +def can_use_custom_all_reduce_with_nvlink( + group: torch.distributed.ProcessGroup, + device: torch.device, + supported_world_size: List[int], + cls_name: str, +) -> Optional[bool]: # None if fail; otherwise return whether NVLink is available + assert ( + dist.get_backend(group) != dist.Backend.NCCL + ), f"{cls_name} should be attached to a non-NCCL group." + + rank = dist.get_rank(group=group) + world_size = dist.get_world_size(group=group) + + # No need to initialize custom allreduce for single GPU case. + if world_size == 1: + return + + # No need to initialize custom allreduce for multi-node case. + if not all(in_the_same_node_as(group, source_rank=0)): + logger.warning( + f"{cls_name} is disabled because this process group" " spans across nodes." + ) + return + + # For not supported world size, we disable custom allreduce. + if world_size not in supported_world_size: + logger.warning( + f"{cls_name} is disabled due to an unsupported world" + f" size: {world_size}. Supported world sizes: {supported_world_size}. " + "To silence this warning, specify disable_custom_all_reduce=True explicitly.", + ) + return + + cuda_visible_devices = os.environ.get("CUDA_VISIBLE_DEVICES", None) + if cuda_visible_devices: + device_ids = list(map(int, cuda_visible_devices.split(","))) + else: + device_ids = list(range(torch.cuda.device_count())) + physical_device_id = device_ids[device.index] + tensor = torch.tensor([physical_device_id], dtype=torch.int, device="cpu") + gather_list = [ + torch.tensor([0], dtype=torch.int, device="cpu") for _ in range(world_size) + ] + dist.all_gather(gather_list, tensor, group=group) + physical_device_ids = [int(t) for t in gather_list] + full_nvlink = is_full_nvlink(physical_device_ids, world_size) + + # test nvlink first, this will filter out most of the cases + # where custom allreduce is not supported + # this checks hardware and driver support for NVLink + if False: + logger.warning( + f"{cls_name} is disabled because it's not supported on" + " more than two PCIe-only GPUs. To silence this warning, " + "specify disable_custom_all_reduce=True explicitly." + ) + return + + # test P2P capability, this checks software/cudaruntime support + # this is expensive to compute at the first time + # then we cache the result + # On AMD GPU, p2p is always enabled between XGMI connected GPUs + if not _is_hip and not can_p2p(rank, world_size): + logger.warning( + f"{cls_name} is disabled because your platform lacks " + "GPU P2P capability or P2P test failed. To silence this " + "warning, specify disable_custom_all_reduce=True explicitly." + ) + return + + return full_nvlink + + +if __name__ == "__main__": + batch_src, batch_tgt, output_file = pickle.loads(sys.stdin.buffer.read()) + result = can_actually_p2p(batch_src, batch_tgt) + with open(output_file, "wb") as f: + f.write(pickle.dumps(result)) diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b/all_results.jsonl b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b/all_results.jsonl new file mode 100644 index 0000000..174c9f4 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b/all_results.jsonl @@ -0,0 +1,20 @@ +{"tag": "b3_16k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9501, "corpus_window": {"start": 2300000, "end": 2431072}, "ok": 8, "failed": 0, "wall_s": 82.58, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 49.6, "input_throughput_tok_s": 1587.15, "ttft_s": {"mean": 3.9252, "p50": 3.9266, "p95": 3.9787, "max": 3.9787, "min": 3.8899}, "tpot_s": {"mean": 0.0125, "p50": 0.0128, "p95": 0.0137, "max": 0.0137, "min": 0.0112}, "e2e_s": {"mean": 10.3227, "p50": 10.4215, "p95": 10.9871, "max": 10.9871, "min": 9.6445}, "per_req_out_tok_s_e2e": {"mean": 49.6617, "p50": 49.8845, "p95": 53.0871, "max": 53.0871, "min": 46.6001}, "per_req_decode_tok_s": {"mean": 80.2896, "p50": 80.4694, "p95": 89.5483, "max": 89.5483, "min": 72.9492}, "spec_accept_length_mean": 2.138, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 16, "new_tokens": 131072, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b3_16k_c8", "arm": "e7b", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9502, "corpus_window": {"start": 2300000, "end": 2562144}, "ok": 16, "failed": 0, "wall_s": 100.63, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 8192, "output_throughput_tok_s": 81.41, "input_throughput_tok_s": 2605.13, "ttft_s": {"mean": 11.7518, "p50": 6.5296, "p95": 32.5233, "max": 32.5233, "min": 3.9869}, "tpot_s": {"mean": 0.0716, "p50": 0.0731, "p95": 0.1351, "max": 0.1351, "min": 0.0266}, "e2e_s": {"mean": 48.3305, "p50": 50.9027, "p95": 78.2404, "max": 78.2404, "min": 17.5929}, "per_req_out_tok_s_e2e": {"mean": 12.9171, "p50": 10.4857, "p95": 29.1027, "max": 29.1027, "min": 6.5439}, "per_req_decode_tok_s": {"mean": 16.8775, "p50": 15.7396, "p95": 37.6553, "max": 37.6553, "min": 7.4162}, "spec_accept_length_mean": 2.375, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 262144, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b3_16k_c16", "arm": "e7b", "summary": {"concurrency": 16, "num_requests": 32, "run_id": 9503, "corpus_window": {"start": 2300000, "end": 2824288}, "ok": 32, "failed": 0, "wall_s": 236.44, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 16384, "output_throughput_tok_s": 69.29, "input_throughput_tok_s": 2217.4, "ttft_s": {"mean": 23.5964, "p50": 13.173, "p95": 60.864, "max": 90.3879, "min": 4.2969}, "tpot_s": {"mean": 0.1733, "p50": 0.1784, "p95": 0.2978, "max": 0.3055, "min": 0.0668}, "e2e_s": {"mean": 112.1663, "p50": 105.5499, "p95": 170.4271, "max": 183.3467, "min": 52.7259}, "per_req_out_tok_s_e2e": {"mean": 5.0312, "p50": 4.8797, "p95": 8.0542, "max": 9.7106, "min": 2.7925}, "per_req_decode_tok_s": {"mean": 6.7269, "p50": 5.8271, "p95": 12.9785, "max": 15.0103, "min": 3.2802}, "spec_accept_length_mean": 2.517, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b3_16k_c32", "arm": "e7b", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9504, "corpus_window": {"start": 2300000, "end": 3348576}, "ok": 64, "failed": 0, "wall_s": 414.96, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 32768, "output_throughput_tok_s": 78.97, "input_throughput_tok_s": 2526.91, "ttft_s": {"mean": 100.2552, "p50": 109.5841, "p95": 169.6812, "max": 190.7286, "min": 5.2261}, "tpot_s": {"mean": 0.1698, "p50": 0.1728, "p95": 0.2599, "max": 0.3191, "min": 0.0328}, "e2e_s": {"mean": 187.0411, "p50": 190.6599, "p95": 271.5134, "max": 292.7775, "min": 86.5615}, "per_req_out_tok_s_e2e": {"mean": 2.9412, "p50": 2.6906, "p95": 4.5951, "max": 5.9149, "min": 1.7488}, "per_req_decode_tok_s": {"mean": 7.1251, "p50": 5.846, "p95": 15.414, "max": 30.5614, "min": 3.1404}, "spec_accept_length_mean": 2.767, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b3_16k_c64", "arm": "e7b", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9505, "corpus_window": {"start": 2300000, "end": 4397152}, "ok": 128, "failed": 0, "wall_s": 777.55, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 65536, "output_throughput_tok_s": 84.29, "input_throughput_tok_s": 2697.14, "ttft_s": {"mean": 240.0965, "p50": 287.3833, "p95": 345.6933, "max": 380.9581, "min": 5.2288}, "tpot_s": {"mean": 0.168, "p50": 0.1702, "p95": 0.2315, "max": 0.3119, "min": 0.0404}, "e2e_s": {"mean": 325.934, "p50": 364.8118, "p95": 432.439, "max": 474.7226, "min": 83.7008}, "per_req_out_tok_s_e2e": {"mean": 1.8253, "p50": 1.4064, "p95": 4.196, "max": 6.117, "min": 1.0785}, "per_req_decode_tok_s": {"mean": 6.6174, "p50": 5.8876, "p95": 12.5326, "max": 24.7977, "min": 3.2126}, "spec_accept_length_mean": 2.986, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b41_1k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9511, "corpus_window": {"start": 4500000, "end": 4508192}, "ok": 8, "failed": 0, "wall_s": 15.82, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 1024, "output_throughput_tok_s": 64.74, "input_throughput_tok_s": 517.93, "ttft_s": {"mean": 0.3034, "p50": 0.3015, "p95": 0.3258, "max": 0.3258, "min": 0.2959}, "tpot_s": {"mean": 0.0132, "p50": 0.0135, "p95": 0.0143, "max": 0.0143, "min": 0.0119}, "e2e_s": {"mean": 1.9769, "p50": 2.0115, "p95": 2.1177, "max": 2.1177, "min": 1.8091}, "per_req_out_tok_s_e2e": {"mean": 64.9178, "p50": 66.8893, "p95": 70.753, "max": 70.753, "min": 60.442}, "per_req_decode_tok_s": {"mean": 76.7983, "p50": 79.4095, "p95": 84.9204, "max": 84.9204, "min": 70.4319}, "spec_accept_length_mean": 2.028, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 8, "new_tokens": 8192, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b41_1k_c8", "arm": "e7b", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9512, "corpus_window": {"start": 4500000, "end": 4516384}, "ok": 16, "failed": 0, "wall_s": 13.44, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 2048, "output_throughput_tok_s": 152.4, "input_throughput_tok_s": 1219.24, "ttft_s": {"mean": 1.21, "p50": 1.7505, "p95": 2.1601, "max": 2.1601, "min": 0.387}, "tpot_s": {"mean": 0.0414, "p50": 0.042, "p95": 0.0537, "max": 0.0537, "min": 0.0319}, "e2e_s": {"mean": 6.4682, "p50": 6.5412, "p95": 8.9777, "max": 8.9777, "min": 4.4513}, "per_req_out_tok_s_e2e": {"mean": 20.4798, "p50": 20.2955, "p95": 28.7556, "max": 28.7556, "min": 14.2576}, "per_req_decode_tok_s": {"mean": 24.8432, "p50": 24.8024, "p95": 31.5592, "max": 31.5592, "min": 18.7758}, "spec_accept_length_mean": 2.048, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 10, "new_tokens": 16384, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b41_1k_c32", "arm": "e7b", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9513, "corpus_window": {"start": 4500000, "end": 4565536}, "ok": 64, "failed": 0, "wall_s": 81.8, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 8192, "output_throughput_tok_s": 100.14, "input_throughput_tok_s": 801.15, "ttft_s": {"mean": 17.5622, "p50": 21.7567, "p95": 25.6003, "max": 27.8932, "min": 1.5226}, "tpot_s": {"mean": 0.1466, "p50": 0.1501, "p95": 0.1905, "max": 0.1952, "min": 0.088}, "e2e_s": {"mean": 36.1812, "p50": 40.1766, "p95": 46.8967, "max": 47.1947, "min": 18.2249}, "per_req_out_tok_s_e2e": {"mean": 3.8358, "p50": 3.1917, "p95": 6.6145, "max": 7.0234, "min": 2.7122}, "per_req_decode_tok_s": {"mean": 7.1041, "p50": 6.7208, "p95": 10.1791, "max": 11.4499, "min": 5.164}, "spec_accept_length_mean": 2.054, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 38, "new_tokens": 65536, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b41_1k_c64", "arm": "e7b", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9514, "corpus_window": {"start": 4500000, "end": 4631072}, "ok": 128, "failed": 0, "wall_s": 158.73, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 16384, "output_throughput_tok_s": 103.22, "input_throughput_tok_s": 825.75, "ttft_s": {"mean": 46.7369, "p50": 57.65, "p95": 65.8132, "max": 70.7061, "min": 1.5163}, "tpot_s": {"mean": 0.1445, "p50": 0.144, "p95": 0.1777, "max": 0.2125, "min": 0.0759}, "e2e_s": {"mean": 65.0893, "p50": 74.2949, "p95": 84.7332, "max": 89.7485, "min": 18.3747}, "per_req_out_tok_s_e2e": {"mean": 2.4029, "p50": 1.7238, "p95": 6.272, "max": 6.9661, "min": 1.4262}, "per_req_decode_tok_s": {"mean": 7.1513, "p50": 7.0024, "p95": 9.0245, "max": 13.2811, "min": 4.7422}, "spec_accept_length_mean": 2.087, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 85, "new_tokens": 131072, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b42_1k4k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9521, "corpus_window": {"start": 4700000, "end": 4708192}, "ok": 8, "failed": 0, "wall_s": 242.18, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 32768, "output_throughput_tok_s": 135.31, "input_throughput_tok_s": 33.83, "ttft_s": {"mean": 0.3142, "p50": 0.3133, "p95": 0.3204, "max": 0.3204, "min": 0.3097}, "tpot_s": {"mean": 0.0073, "p50": 0.0078, "p95": 0.0087, "max": 0.0087, "min": 0.0059}, "e2e_s": {"mean": 30.2716, "p50": 32.1753, "p95": 35.9141, "max": 35.9141, "min": 24.509}, "per_req_out_tok_s_e2e": {"mean": 137.6058, "p50": 136.5088, "p95": 167.1221, "max": 167.1221, "min": 114.0498}, "per_req_decode_tok_s": {"mean": 139.1062, "p50": 137.9513, "p95": 169.3421, "max": 169.3421, "min": 115.0748}, "spec_accept_length_mean": 3.701, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 8, "new_tokens": 8192, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b42_1k4k_c8", "arm": "e7b", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9522, "corpus_window": {"start": 4700000, "end": 4716384}, "ok": 16, "failed": 0, "wall_s": 159.24, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 65536, "output_throughput_tok_s": 411.54, "input_throughput_tok_s": 102.89, "ttft_s": {"mean": 0.9638, "p50": 1.2709, "p95": 1.6805, "max": 1.6805, "min": 0.378}, "tpot_s": {"mean": 0.0181, "p50": 0.0175, "p95": 0.0225, "max": 0.0225, "min": 0.0153}, "e2e_s": {"mean": 75.1101, "p50": 73.2342, "p95": 93.9711, "max": 93.9711, "min": 64.3187}, "per_req_out_tok_s_e2e": {"mean": 55.3003, "p50": 58.2448, "p95": 63.6829, "max": 63.6829, "min": 43.5879}, "per_req_decode_tok_s": {"mean": 55.9977, "p50": 58.5719, "p95": 65.3811, "max": 65.3811, "min": 44.3769}, "spec_accept_length_mean": 4.003, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 10, "new_tokens": 16384, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b42_1k4k_c32", "arm": "e7b", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9523, "corpus_window": {"start": 4700000, "end": 4765536}, "ok": 64, "failed": 0, "wall_s": 1040.79, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 262144, "output_throughput_tok_s": 251.87, "input_throughput_tok_s": 62.97, "ttft_s": {"mean": 195.07, "p50": 240.7323, "p95": 304.7528, "max": 334.6933, "min": 1.3472}, "tpot_s": {"mean": 0.0607, "p50": 0.0581, "p95": 0.0746, "max": 0.0923, "min": 0.0465}, "e2e_s": {"mean": 443.6795, "p50": 478.2794, "p95": 587.5712, "max": 597.8883, "min": 219.3532}, "per_req_out_tok_s_e2e": {"mean": 10.0455, "p50": 8.6106, "p95": 17.4441, "max": 18.6731, "min": 6.8508}, "per_req_decode_tok_s": {"mean": 16.784, "p50": 17.4276, "p95": 20.0358, "max": 21.5315, "min": 10.8381}, "spec_accept_length_mean": 3.852, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 51, "new_tokens": 65536, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b51_64k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9531, "corpus_window": {"start": 5000000, "end": 5524288}, "ok": 8, "failed": 0, "wall_s": 182.23, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 22.48, "input_throughput_tok_s": 2877.06, "ttft_s": {"mean": 17.0384, "p50": 17.022, "p95": 17.1703, "max": 17.1703, "min": 17.0028}, "tpot_s": {"mean": 0.0112, "p50": 0.0117, "p95": 0.0135, "max": 0.0135, "min": 0.008}, "e2e_s": {"mean": 22.7786, "p50": 23.0017, "p95": 23.9301, "max": 23.9301, "min": 21.118}, "per_req_out_tok_s_e2e": {"mean": 22.5108, "p50": 22.5486, "p95": 24.2447, "max": 24.2447, "min": 21.3957}, "per_req_decode_tok_s": {"mean": 91.4983, "p50": 90.0463, "p95": 125.007, "max": 125.007, "min": 74.0091}, "spec_accept_length_mean": 2.468, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b51_64k_c4", "arm": "e7b", "summary": {"concurrency": 4, "num_requests": 8, "run_id": 9532, "corpus_window": {"start": 5000000, "end": 5524288}, "ok": 8, "failed": 0, "wall_s": 159.24, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 25.72, "input_throughput_tok_s": 3292.45, "ttft_s": {"mean": 30.5492, "p50": 18.5951, "p95": 68.8675, "max": 68.8675, "min": 17.0513}, "tpot_s": {"mean": 0.0949, "p50": 0.1145, "p95": 0.1546, "max": 0.1546, "min": 0.02}, "e2e_s": {"mean": 79.0616, "p50": 80.1169, "p95": 131.9282, "max": 131.9282, "min": 27.2645}, "per_req_out_tok_s_e2e": {"mean": 8.1791, "p50": 6.6417, "p95": 18.779, "max": 18.779, "min": 3.8809}, "per_req_decode_tok_s": {"mean": 16.1959, "p50": 11.2172, "p95": 50.1581, "max": 50.1581, "min": 6.4814}, "spec_accept_length_mean": 2.438, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b51_64k_c8", "arm": "e7b", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9533, "corpus_window": {"start": 5000000, "end": 6048576}, "ok": 16, "failed": 0, "wall_s": 319.66, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 8192, "output_throughput_tok_s": 25.63, "input_throughput_tok_s": 3280.31, "ttft_s": {"mean": 90.9186, "p50": 96.5085, "p95": 149.4465, "max": 149.4465, "min": 18.5948}, "tpot_s": {"mean": 0.1058, "p50": 0.1221, "p95": 0.1575, "max": 0.1575, "min": 0.0172}, "e2e_s": {"mean": 144.9847, "p50": 155.6435, "p95": 211.845, "max": 211.845, "min": 75.7166}, "per_req_out_tok_s_e2e": {"mean": 3.7942, "p50": 3.6883, "p95": 6.7621, "max": 6.7621, "min": 2.4169}, "per_req_decode_tok_s": {"mean": 13.0455, "p50": 8.4248, "p95": 58.0944, "max": 58.0944, "min": 6.3604}, "spec_accept_length_mean": 2.442, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b51_128k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9541, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 343.36, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 11.93, "input_throughput_tok_s": 3053.83, "ttft_s": {"mean": 37.5917, "p50": 37.5968, "p95": 37.6012, "max": 37.6012, "min": 37.5564}, "tpot_s": {"mean": 0.0104, "p50": 0.0104, "p95": 0.0123, "max": 0.0123, "min": 0.009}, "e2e_s": {"mean": 42.9201, "p50": 42.9071, "p95": 43.9018, "max": 43.9018, "min": 42.1336}, "per_req_out_tok_s_e2e": {"mean": 11.9312, "p50": 11.9694, "p95": 12.1518, "max": 12.1518, "min": 11.6624}, "per_req_decode_tok_s": {"mean": 97.1304, "p50": 98.8664, "p95": 111.8659, "max": 111.8659, "min": 81.2653}, "spec_accept_length_mean": 2.728, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b51_128k_c2", "arm": "e7b", "summary": {"concurrency": 2, "num_requests": 8, "run_id": 9542, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 333.36, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 12.29, "input_throughput_tok_s": 3145.44, "ttft_s": {"mean": 48.2081, "p50": 39.1403, "p95": 75.8082, "max": 75.8082, "min": 37.5834}, "tpot_s": {"mean": 0.068, "p50": 0.0858, "p95": 0.1622, "max": 0.1622, "min": 0.0099}, "e2e_s": {"mean": 82.9494, "p50": 84.1258, "p95": 122.0311, "max": 122.0311, "min": 42.6804}, "per_req_out_tok_s_e2e": {"mean": 7.0763, "p50": 6.2562, "p95": 11.9961, "max": 11.9961, "min": 4.1957}, "per_req_decode_tok_s": {"mean": 37.0402, "p50": 12.1918, "p95": 100.963, "max": 100.963, "min": 6.1768}, "spec_accept_length_mean": 2.769, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b51_128k_c4", "arm": "e7b", "summary": {"concurrency": 4, "num_requests": 8, "run_id": 9543, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 334.71, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 12.24, "input_throughput_tok_s": 3132.76, "ttft_s": {"mean": 110.1385, "p50": 120.8044, "p95": 159.4963, "max": 159.4963, "min": 39.1664}, "tpot_s": {"mean": 0.0687, "p50": 0.0152, "p95": 0.1634, "max": 0.1634, "min": 0.0097}, "e2e_s": {"mean": 145.2581, "p50": 130.4964, "p95": 204.1758, "max": 204.1758, "min": 83.563}, "per_req_out_tok_s_e2e": {"mean": 3.8065, "p50": 4.04, "p95": 6.1271, "max": 6.1271, "min": 2.5076}, "per_req_decode_tok_s": {"mean": 54.2853, "p50": 73.1655, "p95": 103.6085, "max": 103.6085, "min": 6.1313}, "spec_accept_length_mean": 2.595, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b52_256k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9551, "corpus_window": {"start": 8400000, "end": 8662144}, "ok": 1, "failed": 0, "wall_s": 0.5, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 2.01, "input_throughput_tok_s": 527883.82, "ttft_s": {"mean": 0.4955, "p50": 0.4955, "p95": 0.4955, "max": 0.4955, "min": 0.4955}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 0.4957, "p50": 0.4957, "p95": 0.4957, "max": 0.4957, "min": 0.4957}, "per_req_out_tok_s_e2e": {"mean": 2.0172, "p50": 2.0172, "p95": 2.0172, "max": 2.0172, "min": 2.0172}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 1, "new_tokens": 64, "cached_tokens": 262080, "hit_rate": 0.9998}}} +{"tag": "b52_256k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9551, "corpus_window": {"start": 9900000, "end": 10162144}, "ok": 1, "failed": 0, "wall_s": 91.45, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 0.01, "input_throughput_tok_s": 2866.59, "ttft_s": {"mean": 91.4465, "p50": 91.4465, "p95": 91.4465, "max": 91.4465, "min": 91.4465}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 91.4468, "p50": 91.4468, "p95": 91.4468, "max": 91.4468, "min": 91.4468}, "per_req_out_tok_s_e2e": {"mean": 0.0109, "p50": 0.0109, "p95": 0.0109, "max": 0.0109, "min": 0.0109}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 262144, "cached_tokens": 0, "hit_rate": 0.0}}} diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b/md5_assets.txt b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b/md5_assets.txt new file mode 100644 index 0000000..c154f14 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b/md5_assets.txt @@ -0,0 +1,3 @@ +a8f3909211adb92501cf007ab71d3b29 /root/corpus_ids.json +1e34dd8d813ccb7ac8243e8d1573a49e /root/bench_corpus_v2.py +4c126d067d33b5ea27c268f37561634c /root/extract_summary.py diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b/server_facts.txt b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b/server_facts.txt new file mode 100644 index 0000000..4c48460 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b/server_facts.txt @@ -0,0 +1,18 @@ +[2026-09-10 06:13:36] server_args={'model_path': '/data/hf_models/GLM-5.3-NVFP4', 'tokenizer_path': '/data/hf_models/GLM-5.3-NVFP4', 'tokenizer_mode': 'auto', 'tokenizer_backend': 'huggingface', 'tokenizer_worker_num': 1, 'detokenizer_worker_num': 1, 'skip_tokenizer_init': False, 'load_format': 'auto', 'model_loader_extra_config': '{}', 'trust_remote_code': False, 'context_length': 270336, 'is_embedding': False, 'enable_multimodal': None, 'revision': None, 'model_impl': 'auto', 'model_config_parser': 'auto', 'json_model_override_args': '{}', 'dtype': 'auto', 'quantization': None, 'quantization_param_path': None, 'kv_cache_dtype': 'fp8_e4m3', 'enable_fp32_lm_head': False, 'modelopt_quant': None, 'modelopt_checkpoint_restore_path': None, 'modelopt_checkpoint_save_path': None, 'modelopt_export_path': None, 'quantize_and_serve': False, 'rl_quant_profile': None, 'enable_tf32_matmul': False, 'mem_fraction_static': 0.9, 'max_running_requests': 16, 'max_queued_requests': None, 'max_total_tokens': None, 'chunked_prefill_size': 8192, 'prefill_decode_interval': 0, 'enable_dynamic_chunking': False, 'max_prefill_tokens': 16384, 'prefill_max_requests': None, 'schedule_policy': 'fcfs', 'enable_priority_scheduling': False, 'disable_priority_preemption': False, 'default_priority_value': None, 'abort_on_priority_when_disabled': False, 'schedule_low_priority_values_first': False, 'priority_scheduling_preemption_threshold': 10, 'retraction_policy': 'length', 'schedule_conservativeness': 1.0, 'page_size': 64, 'c128_page_size': 16, 'swa_full_tokens_ratio': 0.8, 'disable_hybrid_swa_memory': False, 'radix_eviction_policy': 'lru', 'prefill_only_disable_kv_cache': False, 'disable_radix_cache': False, 'enable_page_major_kv_layout': False, 'enable_unified_memory': False, 'disable_chunked_prefix_cache': False, 'disable_overlap_schedule': False, 'num_continuous_decode_steps': 1, 'scheduler_recv_interval': 1, 'enable_mixed_chunk': False, 'nccl_port': None, 'dist_timeout': None, 'dist_init_addr': None, 'gated_launch_port': None, 'nnodes': 1, 'node_rank': 0, 'tp_size': 8, 'dcp_size': 1, 'pp_size': 1, 'pp_max_micro_batch_size': None, 'pp_async_batch_depth': 0, 'dp_size': 1, 'load_balance_method': 'round_robin', 'attn_cp_size': 1, 'moe_dp_size': 1, 'dwdp_size': 1, 'dcp_comm_backend': 'ag_rs', 'dcp_replicate_q_proj': None, 'enable_prefill_cp': False, 'cp_strategy': None, 'enable_dsa_cache_layer_split': False, 'enable_dsa_prefill_context_parallel': False, 'dsa_prefill_cp_mode': 'round-robin-split', 'enable_prefill_context_parallel': False, 'prefill_cp_mode': 'in-seq-split', 'enable_cp_decode_attn_tp': False, 'enable_dp_attention': False, 'enable_dp_attention_local_control_broadcast': False, 'enable_dp_lm_head': False, 'enable_tp_lm_head_all_to_all': False, 'enable_attn_tp_input_scattered': False, 'enable_shared_experts_attn_tp': False, 'enable_dense_mlp_attn_tp': False, 'disable_attn_tp_gather': False, 'enable_p2p_check': False, 'device': 'cuda', 'base_gpu_id': 0, 'gpu_id_step': 1, 'random_seed': 516482816, 'mlx_enable_sampling': False, 'watchdog_timeout': 300, 'soft_watchdog_timeout': None, 'sleep_on_idle': False, 'use_ray': False, 'custom_sigquit_handler': None, 'numa_node': None, 'gc_threshold': None, 'host': '0.0.0.0', 'port': 30000, 'fastapi_root_path': '', 'smg_grpc_mode': False, 'grpc_mode': False, 'grpc_port': None, 'grpc_worker_threads': 4, 'sidecar': None, 'sidecar_args': None, 'skip_server_warmup': False, 'warmups': None, 'enable_http2': False, 'http2_max_concurrent_streams': 200, 'ssl_keyfile': None, 'ssl_certfile': None, 'ssl_ca_certs': None, 'ssl_keyfile_password': None, 'enable_ssl_refresh': False, 'api_key': None, 'admin_api_key': None, 'served_model_name': '/data/hf_models/GLM-5.3-NVFP4', 'weight_version': 'default', 'chat_template': None, 'hf_chat_template_name': None, 'completion_template': None, 'file_storage_path': 'sglang_storage', 'enable_cache_report': False, 'reasoning_parser': 'glm45', 'default_chat_template_kwargs': None, 'strip_thinking_cache': False, 'enable_strict_thinking': False, 'tool_call_parser': 'glm47', 'tool_server': None, 'sampling_defaults': 'model', 'asr_max_buffer_seconds': 60, 'asr_max_concurrent_sessions': 32, 'preferred_sampling_params': None, 'allow_auto_truncate': False, 'stream_interval': 1, 'batch_notify_size': 16, 'stream_response_default_include_usage': False, 'incremental_streaming_output': False, 'enable_streaming_session': False, 'enable_session_radix_cache': False, 'log_level': 'info', 'log_level_http': None, 'log_requests': False, 'log_requests_level': 2, 'log_requests_format': 'text', 'log_requests_target': None, 'uvicorn_access_log_exclude_prefixes': [], 'crash_dump_folder': None, 'show_time_cost': False, 'enable_metrics': False, 'smg_http_sidecar_port': None, 'enable_mfu_metrics': False, 'enable_metrics_for_all_schedulers': False, 'load_snapshot_publish_interval': 15, 'tokenizer_metrics_custom_labels_header': 'x-custom-labels', 'tokenizer_metrics_allowed_custom_labels': None, 'extra_metric_labels': None, 'bucket_time_to_first_token': None, 'bucket_inter_token_latency': None, 'bucket_e2e_request_latency': None, 'prompt_tokens_buckets': None, 'generation_tokens_buckets': None, 'gc_warning_threshold_secs': 0.0, 'decode_log_interval': 40, 'enable_request_time_stats_logging': False, 'kv_events_config': None, 'load_publish_endpoint': None, 'enable_forward_pass_metrics': False, 'forward_pass_metrics_worker_id': '', 'forward_pass_metrics_ipc_name': None, 'enable_trace': False, 'trace_modules': 'request', 'otlp_traces_endpoint': 'localhost:4317', 'export_metrics_to_file': False, 'export_metrics_to_file_dir': None, 'stat_loggers': None, 'constrained_json_whitespace_pattern': None, 'constrained_json_disable_any_whitespace': False, 'attention_backend': 'dsa', 'decode_attention_backend': None, 'prefill_attention_backend': None, 'sampling_backend': 'flashinfer', 'grammar_backend': 'xgrammar', 'radix_cache_backend': None, 'mm_attention_backend': None, 'fp8_gemm_runner_backend': 'auto', 'fp4_gemm_runner_backend': 'auto', 'bf16_gemm_backend': 'auto', 'dsa_prefill_backend': 'flashinfer_sparse_mla', 'dsv4_prefill_backend': 'auto', 'dsa_decode_backend': 'flashinfer_sparse_mla', 'dsa_paged_mqa_logits_backend': 'auto', 'dsa_topk_backend': 'sgl-kernel', 'disable_flashinfer_autotune': True, 'flashinfer_autotune_skip_ops': None, 'mamba_backend': 'triton', 'cuda_graph_config': {'decode': {'backend': 'full', 'max_bs': 8, 'bs': [1, 2, 3, 4, 6, 8], 'tc_compiler': 'eager', 'full_prefill_max_req': None, 'full_prefill_prefix_chunk_tokens': None}, 'prefill': {'backend': 'breakable', 'max_bs': 8, 'bs': [4, 8], 'tc_compiler': 'eager', 'full_prefill_max_req': None, 'full_prefill_prefix_chunk_tokens': None}}, 'cuda_graph_backend_decode': None, 'cuda_graph_backend_prefill': None, 'cuda_graph_max_bs_decode': 8, 'cuda_graph_max_bs_prefill': 8, 'cuda_graph_bs_decode': [1, 2, 3, 4, 6, 8], 'cuda_graph_bs_prefill': None, 'cuda_graph_tc_compiler': None, 'disable_prefill_cuda_graph': False, 'disable_decode_cuda_graph': False, 'disable_cuda_graph': False, 'disable_cuda_graph_padding': False, 'enable_profile_cuda_graph': False, 'enable_cudagraph_gc': False, 'debug_cuda_graph': False, 'enable_layerwise_nvtx_marker': False, 'enable_nccl_nvls': False, 'enable_symm_mem': False, 'triton_attention_reduce_in_fp32': False, 'triton_attention_num_kv_splits': 8, 'triton_attention_split_tile_size': None, 'flashinfer_mla_disable_ragged': False, 'enable_fused_qk_norm_rope': False, 'enable_precise_embedding_interpolation': False, 'enable_fused_moe_sum_all_reduce': False, 'enable_deepseek_v4_fp4_indexer': False, 'disable_custom_all_reduce': False, 'enable_mscclpp': False, 'enable_torch_symm_mem': False, 'enable_scattered_sconv': False, 'pre_warm_nccl': False, 'enable_quant_communications': False, 'enable_flashinfer_allreduce_fusion': False, 'enforce_disable_flashinfer_allreduce_fusion': False, 'flashinfer_allreduce_fusion_backend': None, 'enable_aiter_allreduce_fusion': False, 'enable_torch_compile': False, 'enable_torch_compile_debug_mode': False, 'torch_compile_max_bs': 32, 'speculative_algorithm': 'EAGLE', 'speculative_draft_model_path': '/data/hf_models/GLM-5.3-NVFP4', 'speculative_draft_model_revision': None, 'speculative_draft_load_format': None, 'speculative_num_steps': 4, 'speculative_eagle_topk': 1, 'speculative_num_draft_tokens': 5, 'speculative_dflash_block_size': None, 'speculative_dspark_block_size': None, 'speculative_dspark_sps_table_path': None, 'speculative_dspark_confidence_sts_path': None, 'speculative_dspark_align_verify_tokens_to_graph_tier': False, 'speculative_accept_threshold_single': 1.0, 'speculative_accept_threshold_acc': 1.0, 'speculative_use_rejection_sampling': False, 'speculative_token_map': None, 'speculative_attention_mode': 'prefill', 'speculative_draft_attention_backend': None, 'speculative_dsa_topk_backend': 'sgl-kernel', 'speculative_draft_kv_cache_dtype': None, 'speculative_draft_window_size': None, 'speculative_moe_runner_backend': 'flashinfer_cutlass', 'speculative_moe_a2a_backend': None, 'speculative_draft_model_quantization': None, '_speculative_draft_quantization_explicitly_set': False, 'speculative_skip_dp_mlp_sync': False, 'enable_multi_layer_eagle': False, 'speculative_adaptive': False, 'speculative_adaptive_config': None, 'decoupled_spec_bind_endpoint': None, 'decoupled_spec_connect_endpoints': None, 'decoupled_spec_rank': None, 'decoupled_spec_role': 'null', 'spec_trace_dir': None, 'speculative_ngram_min_bfs_breadth': 1, 'speculative_ngram_max_bfs_breadth': 10, 'speculative_ngram_match_type': 'BFS', 'speculative_ngram_max_trie_depth': 18, 'speculative_ngram_capacity': 10000000, 'speculative_ngram_external_corpus_path': None, 'speculative_ngram_external_sam_budget': 0, 'speculative_ngram_external_corpus_max_tokens': 10000000, 'ep_size': 1, 'moe_a2a_backend': 'none', 'enable_w4a4_mxfp4_megamoe': False, 'deepep_v2_mode': 'direct', 'moe_runner_backend': 'flashinfer_cutlass', 'flashinfer_mxfp4_moe_precision': 'default', 'deepep_mode': 'auto', 'fuseep_mode': 2, 'deepep_dispatcher_output_dtype': 'auto', 'ep_num_redundant_experts': 0, 'ep_dispatch_algorithm': None, 'init_expert_location': 'trivial', 'enable_eplb': False, 'eplb_algorithm': 'auto', 'eplb_rebalance_num_iterations': 1000, 'eplb_rebalance_layers_per_chunk': None, 'eplb_min_rebalancing_utilization_threshold': 1.0, 'expert_distribution_recorder_mode': None, 'expert_distribution_recorder_buffer_size': 1000, 'expert_balancedness_report_mode': 'off', 'deepep_config': None, 'moe_dense_tp_size': None, 'elastic_ep_backend': None, 'enable_elastic_expert_backup': False, 'mooncake_ib_device': None, 'enable_waterfill': False, 'ep_join_mode': None, 'ep_join_rank_offset': 0, 'elastic_ep_initial_size': None, 'max_ep_size': None, 'elastic_ep_scale_timeout': 600, 'elastic_ep_rejoin': False, 'disable_flashinfer_cutlass_moe_fp4_allgather': False, 'disable_shared_experts_fusion': True, 'enforce_shared_experts_fusion': False, 'max_mamba_cache_size': None, 'mamba_ssm_dtype': None, 'mamba_max_states_per_path': -1, 'enable_mamba_cache_stochastic_rounding': False, 'mamba_cache_philox_rounds': 0, 'mamba_full_memory_ratio': 0.9, 'mamba_radix_cache_strategy': 'auto', 'uses_mamba_radix_cache': False, 'mamba_track_interval': 256, 'enable_int8_mamba_checkpoint': False, 'int8_mamba_ckpt_size': None, 'linear_attn_backend': 'triton', 'linear_attn_decode_backend': None, 'linear_attn_prefill_backend': None, 'linear_attn_verify_backend': None, 'enable_linear_replayssm': False, 'linear_replayssm_cache_len': 16, 'enable_linear_replayssm_spec': False, 'enable_hierarchical_cache': True, 'hicache_host_memory_mode': 'cache', 'hicache_ratio': 3.0, 'hicache_size': 0, 'hicache_write_policy': 'write_through', 'hicache_io_backend': 'kernel', 'hicache_mem_layout': 'page_first', 'hicache_storage_backend': None, 'hicache_storage_prefetch_policy': 'timeout', 'hicache_storage_backend_extra_config': None, 'enable_hisparse': False, 'hisparse_config': None, 'enable_broadcast_mm_inputs_process': False, 'enable_prefix_mm_cache': False, 'mm_enable_dp_encoder': False, 'mm_process_config': {}, 'mm_processor_worker_num': 0, 'mm_io_worker_num': 0, 'allowed_media_domains': [], 'media_url_max_file_size_mb': 64, 'mm_preprocess_cache_size_mb': None, 'trust_mm_content_hashes': False, 'limit_mm_data_per_request': None, 'enable_mm_global_cache': False, 'image_processor_backend': 'auto', 'mm_global_cache_backend': 'mooncake', 'disable_fast_image_processor': False, 'mm_feature_transport': 'cpu', 'keep_mm_feature_on_device': False, 'enable_lora': None, 'enable_lora_overlap_loading': None, 'max_lora_rank': None, 'lora_target_modules': None, 'lora_paths': None, 'max_loaded_loras': None, 'max_loras_per_batch': 8, 'lora_eviction_policy': 'lru', 'lora_backend': 'csgmv', 'max_lora_chunk_size': 16, 'experts_shared_outer_loras': None, 'lora_use_virtual_experts': False, 'lora_strict_loading': False, 'lora_drain_wait_threshold': 0.0, 'enable_two_batch_overlap': False, 'enable_single_batch_overlap': False, 'tbo_token_distribution_threshold': 0.48, 'cpu_offload_gb': 0, 'offload_group_size': -1, 'offload_num_in_group': 1, 'offload_prefetch_step': 1, 'offload_mode': 'cpu', 'enable_lmcache': False, 'lmcache_config_file': None, 'enable_flexkv': False, 'flexkv_config_file': None, 'kt_weight_path': None, 'kt_method': 'AMXINT4', 'kt_cpuinfer': None, 'kt_threadpool_count': 2, 'kt_num_gpu_experts': None, 'kt_max_deferred_experts_per_token': None, 'dllm_algorithm': None, 'dllm_algorithm_config': None, 'dllm_fdfo': True, 'disaggregation_mode': 'null', 'disaggregation_transfer_backend': 'mooncake', 'disaggregation_bootstrap_port': 8998, 'disaggregation_ib_device': None, 'disaggregation_decode_enable_radix_cache': False, 'disaggregation_decode_enable_offload_kvcache': False, 'disaggregation_decode_retraction_backup': None, 'num_reserved_decode_tokens': 512, 'disaggregation_decode_extra_slots': None, 'disaggregation_decode_polling_interval': 1, 'optimistic_prefill_attempts': 0, 'encoder_only': False, 'language_only': False, 'language_model_only': False, 'encoder_transfer_backend': 'zmq_to_scheduler', 'encoder_urls': [], 'encoder_bootstrap_port': 8997, 'encoder_register_urls': [], 'enable_adaptive_dispatch_to_encoder': False, 'enable_pdmux': False, 'pdmux_config_path': None, 'sm_group_num': 8, 'startup_weight_load_mode': 'serial', 'custom_weight_loader': [], 'weight_loader_disable_mmap': False, 'weight_loader_prefetch_checkpoints': False, 'weight_loader_prefetch_num_threads': 4, 'weight_loader_drop_cache_after_load': False, 'remote_instance_weight_loader_seed_instance_ip': None, 'remote_instance_weight_loader_seed_instance_service_port': None, 'remote_instance_weight_loader_send_weights_group_ports': None, 'remote_instance_weight_loader_backend': 'nccl', 'remote_instance_weight_loader_start_seed_via_transfer_engine': False, 'engine_info_bootstrap_port': 6789, 'modelexpress_config': None, 'download_dir': None, 'model_checksum': None, 'delete_ckpt_after_loading': False, 'decrypted_config_file': None, 'decrypted_draft_config_file': None, 'checkpoint_engine_wait_weights_before_ready': False, 'enable_prefill_delayer': False, 'prefill_delayer_max_delay_passes': 30, 'prefill_delayer_token_usage_low_watermark': None, 'prefill_delayer_forward_passes_buckets': None, 'prefill_delayer_wait_seconds_buckets': None, 'prefill_delayer_queue_min_ratio': None, 'prefill_delayer_max_delay_ms': None, 'min_free_slots_delay': None, 'enable_deterministic_inference': False, 'rl_on_policy_target': None, 'kv_canary': 'none', 'kv_canary_real_data': 'none', 'kv_canary_sweep_interval': 0, 'enable_dynamic_batch_tokenizer': False, 'dynamic_batch_tokenizer_batch_size': 32, 'dynamic_batch_tokenizer_batch_timeout': 0.002, 'enable_tokenizer_batch_encode': False, 'disable_tokenizer_batch_decode': False, 'debug_tensor_dump_output_folder': None, 'debug_tensor_dump_layers': None, 'debug_tensor_dump_input_file': None, 'enable_memory_saver': False, 'enable_weights_cpu_backup': False, 'enable_draft_weights_cpu_backup': False, 'enable_custom_logit_processor': False, 'enable_return_hidden_states': False, 'return_hidden_states_mode': None, 'enable_return_routed_experts': False, 'enable_return_indexer_topk': False, 'disable_outlines_disk_cache': False, 'enable_mis': False, 'weight_cache_mode': 'off', 'weight_cache_socket': None, 'weight_cache_timeout': 1800, 'forward_hooks': None, 'msprobe_dump_config': None} +[2026-09-10 06:18:05 TP7] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 06:18:05 TP7] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 06:18:05 TP2] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 06:18:05 TP4] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 06:18:05 TP1] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 06:18:05 TP2] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 06:18:05 TP4] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 06:18:05 TP1] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 06:18:05 TP6] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 06:18:05 TP3] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 06:18:05 TP5] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 06:18:05 TP6] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 06:18:05 TP3] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 06:18:05 TP5] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 06:18:05 TP0] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 15.83 GB +[2026-09-10 06:18:05 TP0] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 276480, KV size: 0.20 GB +[2026-09-10 06:20:15 TP0] max_total_num_tokens=276480, chunked_prefill_size=8192, max_prefill_tokens=16384, max_running_requests=16, context_len=270336, available_gpu_mem=7.53 GB diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b/status.txt b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b/status.txt new file mode 100644 index 0000000..f1b1e64 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/e7b/status.txt @@ -0,0 +1,20 @@ +b3_16k_c1 OK +b3_16k_c8 OK +b3_16k_c16 OK +b3_16k_c32 OK +b3_16k_c64 OK +b41_1k_c1 OK +b41_1k_c8 OK +b41_1k_c32 OK +b41_1k_c64 OK +b42_1k4k_c1 OK +b42_1k4k_c8 OK +b42_1k4k_c32 OK +b42_1k4k_c64 BENCH_FAIL rc=143 +b51_64k_c1 OK +b51_64k_c4 OK +b51_64k_c8 OK +b51_128k_c1 OK +b51_128k_c2 OK +b51_128k_c4 OK +b52_256k_c1 HIT_FAIL_FINAL diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/md5_ledger.txt b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/md5_ledger.txt new file mode 100644 index 0000000..72c33f2 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/md5_ledger.txt @@ -0,0 +1,7 @@ +1e34dd8d813ccb7ac8243e8d1573a49e bench_corpus_v2.py +4c126d067d33b5ea27c268f37561634c extract_summary.py +5892b44610b2ce61f533f1721105625e run_b300_matrix.sh +def3c64c5dc19e1d507080e3366c3762 deploy_glm53_pp4.sh +21db641f1d997e26fcf1ddc97163b87e deploy_glm53_607_exp.sh +a8fc9a504c041acc40831a848b4985d8 patches/custom_all_reduce.py +65a4d22bd87ae343e7d68c4879c14f98 patches/custom_all_reduce_utils.py diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/tp2pp4/all_results.jsonl b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/tp2pp4/all_results.jsonl new file mode 100644 index 0000000..11763a0 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/tp2pp4/all_results.jsonl @@ -0,0 +1,24 @@ +{"tag": "b3_16k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9501, "corpus_window": {"start": 2300000, "end": 2431072}, "ok": 8, "failed": 0, "wall_s": 245.55, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 16.68, "input_throughput_tok_s": 533.8, "ttft_s": {"mean": 5.3493, "p50": 5.3474, "p95": 5.3628, "max": 5.3628, "min": 5.3438}, "tpot_s": {"mean": 0.0496, "p50": 0.0495, "p95": 0.0503, "max": 0.0503, "min": 0.0493}, "e2e_s": {"mean": 30.6931, "p50": 30.657, "p95": 31.0364, "max": 31.0364, "min": 30.5593}, "per_req_out_tok_s_e2e": {"mean": 16.6817, "p50": 16.7233, "p95": 16.7543, "max": 16.7543, "min": 16.4968}, "per_req_decode_tok_s": {"mean": 20.2031, "p50": 20.2689, "p95": 20.3081, "max": 20.3081, "min": 19.9299}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b3_16k_c8", "arm": "tp2pp4", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9502, "corpus_window": {"start": 2300000, "end": 2562144}, "ok": 16, "failed": 0, "wall_s": 81.0, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 8192, "output_throughput_tok_s": 101.13, "input_throughput_tok_s": 3236.22, "ttft_s": {"mean": 10.0539, "p50": 10.6041, "p95": 14.7948, "max": 14.7948, "min": 5.2557}, "tpot_s": {"mean": 0.0595, "p50": 0.0602, "p95": 0.0709, "max": 0.0709, "min": 0.0481}, "e2e_s": {"mean": 40.4682, "p50": 41.4728, "p95": 41.6589, "max": 41.6589, "min": 39.2731}, "per_req_out_tok_s_e2e": {"mean": 12.662, "p50": 13.0063, "p95": 13.0369, "max": 13.0369, "min": 12.2903}, "per_req_decode_tok_s": {"mean": 17.0379, "p50": 17.0019, "p95": 20.8395, "max": 20.8395, "min": 14.1226}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b3_16k_c16", "arm": "tp2pp4", "summary": {"concurrency": 16, "num_requests": 32, "run_id": 9503, "corpus_window": {"start": 2300000, "end": 2824288}, "ok": 32, "failed": 0, "wall_s": 112.16, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 16384, "output_throughput_tok_s": 146.08, "input_throughput_tok_s": 4674.5, "ttft_s": {"mean": 15.5283, "p50": 16.1378, "p95": 25.7815, "max": 25.8389, "min": 5.267}, "tpot_s": {"mean": 0.0792, "p50": 0.0797, "p95": 0.0985, "max": 0.1012, "min": 0.0571}, "e2e_s": {"mean": 55.9934, "p50": 56.8023, "p95": 57.0853, "max": 57.1347, "min": 54.9678}, "per_req_out_tok_s_e2e": {"mean": 9.1467, "p50": 9.2952, "p95": 9.3139, "max": 9.3145, "min": 8.9613}, "per_req_decode_tok_s": {"mean": 12.9884, "p50": 12.7194, "p95": 16.7725, "max": 17.5544, "min": 9.9045}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b3_16k_c32", "arm": "tp2pp4", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9504, "corpus_window": {"start": 2300000, "end": 3348576}, "ok": 64, "failed": 0, "wall_s": 169.54, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 32768, "output_throughput_tok_s": 193.27, "input_throughput_tok_s": 6184.73, "ttft_s": {"mean": 26.5164, "p50": 27.1445, "p95": 46.4256, "max": 47.8906, "min": 5.2618}, "tpot_s": {"mean": 0.1137, "p50": 0.114, "p95": 0.152, "max": 0.1573, "min": 0.07}, "e2e_s": {"mean": 84.6344, "p50": 85.322, "p95": 85.7436, "max": 85.8459, "min": 83.6335}, "per_req_out_tok_s_e2e": {"mean": 6.0502, "p50": 6.108, "p95": 6.1205, "max": 6.1219, "min": 5.9642}, "per_req_decode_tok_s": {"mean": 9.2806, "p50": 8.8387, "p95": 13.3026, "max": 14.3169, "min": 6.3708}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b3_16k_c64", "arm": "tp2pp4", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9505, "corpus_window": {"start": 2300000, "end": 4397152}, "ok": 128, "failed": 0, "wall_s": 310.79, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 65536, "output_throughput_tok_s": 210.87, "input_throughput_tok_s": 6747.91, "ttft_s": {"mean": 63.0522, "p50": 49.1362, "p95": 134.7794, "max": 139.9931, "min": 5.3652}, "tpot_s": {"mean": 0.1392, "p50": 0.1358, "p95": 0.2037, "max": 0.2143, "min": 0.0707}, "e2e_s": {"mean": 134.1654, "p50": 114.4018, "p95": 226.099, "max": 226.1544, "min": 84.0644}, "per_req_out_tok_s_e2e": {"mean": 4.1916, "p50": 4.476, "p95": 6.0862, "max": 6.0906, "min": 2.2639}, "per_req_decode_tok_s": {"mean": 7.7966, "p50": 7.3815, "p95": 11.9051, "max": 14.1753, "min": 4.6761}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 512, "new_tokens": 8388608, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b41_1k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9511, "corpus_window": {"start": 4500000, "end": 4508192}, "ok": 8, "failed": 0, "wall_s": 52.95, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 1024, "output_throughput_tok_s": 19.34, "input_throughput_tok_s": 154.7, "ttft_s": {"mean": 0.3848, "p50": 0.3852, "p95": 0.3865, "max": 0.3865, "min": 0.3825}, "tpot_s": {"mean": 0.0491, "p50": 0.0491, "p95": 0.0493, "max": 0.0493, "min": 0.0488}, "e2e_s": {"mean": 6.619, "p50": 6.62, "p95": 6.6447, "max": 6.6447, "min": 6.5863}, "per_req_out_tok_s_e2e": {"mean": 19.3385, "p50": 19.3379, "p95": 19.4343, "max": 19.4343, "min": 19.2635}, "per_req_decode_tok_s": {"mean": 20.5327, "p50": 20.5305, "p95": 20.6346, "max": 20.6346, "min": 20.4442}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 32768, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b41_1k_c8", "arm": "tp2pp4", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9512, "corpus_window": {"start": 4500000, "end": 4516384}, "ok": 16, "failed": 0, "wall_s": 20.1, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 2048, "output_throughput_tok_s": 101.91, "input_throughput_tok_s": 815.3, "ttft_s": {"mean": 1.3131, "p50": 1.2998, "p95": 1.6901, "max": 1.6901, "min": 0.6481}, "tpot_s": {"mean": 0.0686, "p50": 0.0684, "p95": 0.0729, "max": 0.0729, "min": 0.0669}, "e2e_s": {"mean": 10.024, "p50": 10.1096, "p95": 10.2887, "max": 10.2887, "min": 9.7907}, "per_req_out_tok_s_e2e": {"mean": 12.7741, "p50": 12.9718, "p95": 13.0736, "max": 13.0736, "min": 12.4409}, "per_req_decode_tok_s": {"mean": 14.7032, "p50": 14.7329, "p95": 15.0683, "max": 15.0683, "min": 13.8323}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 24, "new_tokens": 65536, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b41_1k_c32", "arm": "tp2pp4", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9513, "corpus_window": {"start": 4500000, "end": 4565536}, "ok": 64, "failed": 0, "wall_s": 29.73, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 8192, "output_throughput_tok_s": 275.55, "input_throughput_tok_s": 2204.4, "ttft_s": {"mean": 3.8882, "p50": 3.9615, "p95": 4.8674, "max": 4.8759, "min": 1.5565}, "tpot_s": {"mean": 0.0856, "p50": 0.0869, "p95": 0.0985, "max": 0.1068, "min": 0.0785}, "e2e_s": {"mean": 14.758, "p50": 14.8384, "p95": 15.0559, "max": 15.1157, "min": 14.5917}, "per_req_out_tok_s_e2e": {"mean": 8.6742, "p50": 8.7413, "p95": 8.7701, "max": 8.7721, "min": 8.468}, "per_req_decode_tok_s": {"mean": 11.8377, "p50": 11.6015, "p95": 12.822, "max": 12.8317, "min": 9.4403}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 36, "new_tokens": 262144, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b41_1k_c64", "arm": "tp2pp4", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9514, "corpus_window": {"start": 4500000, "end": 4631072}, "ok": 128, "failed": 0, "wall_s": 48.11, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 16384, "output_throughput_tok_s": 340.53, "input_throughput_tok_s": 2724.25, "ttft_s": {"mean": 8.1393, "p50": 4.8485, "p95": 19.9553, "max": 20.0096, "min": 0.3836}, "tpot_s": {"mean": 0.0933, "p50": 0.0924, "p95": 0.1018, "max": 0.1251, "min": 0.0818}, "e2e_s": {"mean": 19.9938, "p50": 16.0958, "p95": 32.0105, "max": 32.0464, "min": 15.7372}, "per_req_out_tok_s_e2e": {"mean": 6.9903, "p50": 7.9525, "p95": 8.1268, "max": 8.1336, "min": 3.9942}, "per_req_decode_tok_s": {"mean": 10.8849, "p50": 10.9086, "p95": 12.2874, "max": 12.3245, "min": 8.0552}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 52, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b42_1k4k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9521, "corpus_window": {"start": 4700000, "end": 4708192}, "ok": 8, "failed": 0, "wall_s": 1614.16, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 32768, "output_throughput_tok_s": 20.3, "input_throughput_tok_s": 5.08, "ttft_s": {"mean": 0.3926, "p50": 0.394, "p95": 0.3952, "max": 0.3952, "min": 0.3852}, "tpot_s": {"mean": 0.0492, "p50": 0.0491, "p95": 0.0498, "max": 0.0498, "min": 0.049}, "e2e_s": {"mean": 201.7695, "p50": 201.6506, "p95": 204.4013, "max": 204.4013, "min": 200.9309}, "per_req_out_tok_s_e2e": {"mean": 20.3009, "p50": 20.3439, "p95": 20.3851, "max": 20.3851, "min": 20.039}, "per_req_decode_tok_s": {"mean": 20.3407, "p50": 20.3837, "p95": 20.4254, "max": 20.4254, "min": 20.078}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 32768, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b42_1k4k_c8", "arm": "tp2pp4", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9522, "corpus_window": {"start": 4700000, "end": 4716384}, "ok": 16, "failed": 0, "wall_s": 649.69, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 65536, "output_throughput_tok_s": 100.87, "input_throughput_tok_s": 25.22, "ttft_s": {"mean": 1.352, "p50": 1.4385, "p95": 1.8726, "max": 1.8726, "min": 0.4364}, "tpot_s": {"mean": 0.079, "p50": 0.0791, "p95": 0.0794, "max": 0.0794, "min": 0.0788}, "e2e_s": {"mean": 324.8285, "p50": 325.5653, "p95": 325.6331, "max": 325.6331, "min": 324.0447}, "per_req_out_tok_s_e2e": {"mean": 12.6098, "p50": 12.6389, "p95": 12.6402, "max": 12.6402, "min": 12.5786}, "per_req_decode_tok_s": {"mean": 12.6626, "p50": 12.6734, "p95": 12.7003, "max": 12.7003, "min": 12.5955}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 20, "new_tokens": 65536, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b42_1k4k_c32", "arm": "tp2pp4", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9523, "corpus_window": {"start": 4700000, "end": 4765536}, "ok": 64, "failed": 0, "wall_s": 714.14, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 262144, "output_throughput_tok_s": 367.08, "input_throughput_tok_s": 91.77, "ttft_s": {"mean": 1.9772, "p50": 1.1755, "p95": 4.1378, "max": 4.1467, "min": 0.3867}, "tpot_s": {"mean": 0.0861, "p50": 0.0887, "p95": 0.0905, "max": 0.0911, "min": 0.0805}, "e2e_s": {"mean": 354.393, "p50": 367.2171, "p95": 373.3199, "max": 374.0319, "min": 330.2086}, "per_req_out_tok_s_e2e": {"mean": 11.5833, "p50": 11.8878, "p95": 12.2953, "max": 12.4043, "min": 10.9509}, "per_req_decode_tok_s": {"mean": 11.6457, "p50": 11.9132, "p95": 12.3308, "max": 12.4283, "min": 10.9854}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 262144, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b42_1k4k_c64", "arm": "tp2pp4", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9524, "corpus_window": {"start": 4700000, "end": 4831072}, "ok": 128, "failed": 0, "wall_s": 1088.63, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 524288, "output_throughput_tok_s": 481.6, "input_throughput_tok_s": 120.4, "ttft_s": {"mean": 80.7415, "p50": 4.7569, "p95": 337.9195, "max": 338.4726, "min": 0.3985}, "tpot_s": {"mean": 0.0912, "p50": 0.0914, "p95": 0.0955, "max": 0.0963, "min": 0.0852}, "e2e_s": {"mean": 454.4024, "p50": 385.4538, "p95": 718.7004, "max": 732.2026, "min": 351.072}, "per_req_out_tok_s_e2e": {"mean": 9.6298, "p50": 10.6531, "p95": 11.4339, "max": 11.6671, "min": 5.5941}, "per_req_decode_tok_s": {"mean": 10.9736, "p50": 10.9486, "p95": 11.4764, "max": 11.7362, "min": 10.3887}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 200, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b51_64k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9531, "corpus_window": {"start": 5000000, "end": 5524288}, "ok": 8, "failed": 0, "wall_s": 285.91, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 14.33, "input_throughput_tok_s": 1833.72, "ttft_s": {"mean": 10.2811, "p50": 10.2794, "p95": 10.3216, "max": 10.3216, "min": 10.2634}, "tpot_s": {"mean": 0.0498, "p50": 0.0499, "p95": 0.05, "max": 0.05, "min": 0.0496}, "e2e_s": {"mean": 35.739, "p50": 35.7693, "p95": 35.838, "max": 35.838, "min": 35.6073}, "per_req_out_tok_s_e2e": {"mean": 14.3261, "p50": 14.3345, "p95": 14.3791, "max": 14.3791, "min": 14.2865}, "per_req_decode_tok_s": {"mean": 20.112, "p50": 20.1156, "p95": 20.2096, "max": 20.2096, "min": 20.052}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b51_64k_c4", "arm": "tp2pp4", "summary": {"concurrency": 4, "num_requests": 8, "run_id": 9532, "corpus_window": {"start": 5000000, "end": 5524288}, "ok": 8, "failed": 0, "wall_s": 122.3, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 33.49, "input_throughput_tok_s": 4286.94, "ttft_s": {"mean": 19.5548, "p50": 22.5526, "p95": 28.8262, "max": 28.8262, "min": 10.2823}, "tpot_s": {"mean": 0.0814, "p50": 0.0873, "p95": 0.0995, "max": 0.0995, "min": 0.0632}, "e2e_s": {"mean": 61.1314, "p50": 61.2078, "p95": 61.2548, "max": 61.2548, "min": 61.002}, "per_req_out_tok_s_e2e": {"mean": 8.3754, "p50": 8.3851, "p95": 8.3932, "max": 8.3932, "min": 8.3585}, "per_req_decode_tok_s": {"mean": 12.6679, "p50": 13.2834, "p95": 15.8466, "max": 15.8466, "min": 10.0675}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b51_64k_c8", "arm": "tp2pp4", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9533, "corpus_window": {"start": 5000000, "end": 6048576}, "ok": 16, "failed": 0, "wall_s": 185.59, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 8192, "output_throughput_tok_s": 44.14, "input_throughput_tok_s": 5650.11, "ttft_s": {"mean": 31.833, "p50": 34.8143, "p95": 53.3841, "max": 53.3841, "min": 10.2802}, "tpot_s": {"mean": 0.1192, "p50": 0.1246, "p95": 0.1618, "max": 0.1618, "min": 0.0765}, "e2e_s": {"mean": 92.7517, "p50": 92.8482, "p95": 92.9902, "max": 92.9902, "min": 92.4957}, "per_req_out_tok_s_e2e": {"mean": 5.5201, "p50": 5.5226, "p95": 5.5354, "max": 5.5354, "min": 5.506}, "per_req_decode_tok_s": {"mean": 8.9039, "p50": 8.8159, "p95": 13.0909, "max": 13.0909, "min": 6.1915}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b51_128k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9541, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 347.5, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 11.79, "input_throughput_tok_s": 3017.45, "ttft_s": {"mean": 18.2074, "p50": 18.2043, "p95": 18.2252, "max": 18.2252, "min": 18.1922}, "tpot_s": {"mean": 0.0494, "p50": 0.0496, "p95": 0.0497, "max": 0.0497, "min": 0.0485}, "e2e_s": {"mean": 43.4375, "p50": 43.5633, "p95": 43.6252, "max": 43.6252, "min": 43.0007}, "per_req_out_tok_s_e2e": {"mean": 11.7874, "p50": 11.7541, "p95": 11.9068, "max": 11.9068, "min": 11.7363}, "per_req_decode_tok_s": {"mean": 20.2951, "p50": 20.1994, "p95": 20.6444, "max": 20.6444, "min": 20.1578}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b51_128k_c2", "arm": "tp2pp4", "summary": {"concurrency": 2, "num_requests": 8, "run_id": 9542, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 243.3, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 16.84, "input_throughput_tok_s": 4309.84, "ttft_s": {"mean": 25.2658, "p50": 32.1952, "p95": 32.4188, "max": 32.4188, "min": 18.173}, "tpot_s": {"mean": 0.0696, "p50": 0.0831, "p95": 0.0837, "max": 0.0837, "min": 0.0556}, "e2e_s": {"mean": 60.8183, "p50": 60.7483, "p95": 61.2374, "max": 61.2374, "min": 60.6134}, "per_req_out_tok_s_e2e": {"mean": 8.4186, "p50": 8.4352, "p95": 8.447, "max": 8.447, "min": 8.3609}, "per_req_decode_tok_s": {"mean": 14.9825, "p50": 17.8153, "p95": 18.0089, "max": 18.0089, "min": 11.965}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b51_128k_c4", "arm": "tp2pp4", "summary": {"concurrency": 4, "num_requests": 8, "run_id": 9543, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 186.37, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 21.98, "input_throughput_tok_s": 5626.35, "ttft_s": {"mean": 39.2939, "p50": 46.1927, "p95": 60.3496, "max": 60.3496, "min": 18.2774}, "tpot_s": {"mean": 0.1054, "p50": 0.1187, "p95": 0.1468, "max": 0.1468, "min": 0.0638}, "e2e_s": {"mean": 93.1534, "p50": 93.2237, "p95": 93.3139, "max": 93.3139, "min": 92.97}, "per_req_out_tok_s_e2e": {"mean": 5.4963, "p50": 5.497, "p95": 5.5072, "max": 5.5072, "min": 5.4869}, "per_req_decode_tok_s": {"mean": 10.4415, "p50": 10.8799, "p95": 15.6959, "max": 15.6959, "min": 6.8234}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b52_256k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9551, "corpus_window": {"start": 8400000, "end": 8662144}, "ok": 1, "failed": 0, "wall_s": 37.18, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 0.03, "input_throughput_tok_s": 7050.23, "ttft_s": {"mean": 37.1811, "p50": 37.1811, "p95": 37.1811, "max": 37.1811, "min": 37.1811}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 37.1814, "p50": 37.1814, "p95": 37.1814, "max": 37.1814, "min": 37.1814}, "per_req_out_tok_s_e2e": {"mean": 0.0269, "p50": 0.0269, "p95": 0.0269, "max": 0.0269, "min": 0.0269}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b52_256k_c2", "arm": "tp2pp4", "summary": {"concurrency": 2, "num_requests": 2, "run_id": 9552, "corpus_window": {"start": 8400000, "end": 8924288}, "ok": 2, "failed": 0, "wall_s": 70.19, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 2, "output_throughput_tok_s": 0.03, "input_throughput_tok_s": 7469.22, "ttft_s": {"mean": 53.6945, "p50": 70.1229, "p95": 70.1229, "max": 70.1229, "min": 37.2662}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 53.6948, "p50": 70.1232, "p95": 70.1232, "max": 70.1232, "min": 37.2664}, "per_req_out_tok_s_e2e": {"mean": 0.0205, "p50": 0.0268, "p95": 0.0268, "max": 0.0268, "min": 0.0143}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b52_256k_c3", "arm": "tp2pp4", "summary": {"concurrency": 3, "num_requests": 3, "run_id": 9553, "corpus_window": {"start": 8400000, "end": 9186432}, "ok": 3, "failed": 0, "wall_s": 103.12, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 3, "output_throughput_tok_s": 0.03, "input_throughput_tok_s": 7626.22, "ttft_s": {"mean": 70.1194, "p50": 70.1196, "p95": 102.986, "max": 102.986, "min": 37.2526}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 70.1196, "p50": 70.1198, "p95": 102.9863, "max": 102.9863, "min": 37.2529}, "per_req_out_tok_s_e2e": {"mean": 0.0169, "p50": 0.0143, "p95": 0.0268, "max": 0.0268, "min": 0.0097}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 192, "new_tokens": 3145728, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b52_512k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9554, "corpus_window": {"start": 9300000, "end": 9824288}, "ok": 1, "failed": 0, "wall_s": 92.6, "input_len": 524288, "shared_len": 0, "unique_len": 524288, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 0.01, "input_throughput_tok_s": 5662.03, "ttft_s": {"mean": 92.5959, "p50": 92.5959, "p95": 92.5959, "max": 92.5959, "min": 92.5959}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 92.5962, "p50": 92.5962, "p95": 92.5962, "max": 92.5962, "min": 92.5962}, "per_req_out_tok_s_e2e": {"mean": 0.0108, "p50": 0.0108, "p95": 0.0108, "max": 0.0108, "min": 0.0108}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}} +{"tag": "b52_896k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9555, "corpus_window": {"start": 9900000, "end": 10817504}, "ok": 1, "failed": 0, "wall_s": 220.91, "input_len": 917504, "shared_len": 0, "unique_len": 917504, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 0.0, "input_throughput_tok_s": 4153.24, "ttft_s": {"mean": 220.9117, "p50": 220.9117, "p95": 220.9117, "max": 220.9117, "min": 220.9117}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 220.912, "p50": 220.912, "p95": 220.912, "max": 220.912, "min": 220.912}, "per_req_out_tok_s_e2e": {"mean": 0.0045, "p50": 0.0045, "p95": 0.0045, "max": 0.0045, "min": 0.0045}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 224, "new_tokens": 3670016, "cached_tokens": 0, "hit_rate": 0.0}}} diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/tp2pp4/md5_assets.txt b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/tp2pp4/md5_assets.txt new file mode 100644 index 0000000..c154f14 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/tp2pp4/md5_assets.txt @@ -0,0 +1,3 @@ +a8f3909211adb92501cf007ab71d3b29 /root/corpus_ids.json +1e34dd8d813ccb7ac8243e8d1573a49e /root/bench_corpus_v2.py +4c126d067d33b5ea27c268f37561634c /root/extract_summary.py diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/tp2pp4/server_facts.txt b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/tp2pp4/server_facts.txt new file mode 100644 index 0000000..9c4c861 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/tp2pp4/server_facts.txt @@ -0,0 +1,13 @@ +[2026-09-10 03:21:32] server_args={'model_path': '/data/hf_models/GLM-5.3-NVFP4', 'tokenizer_path': '/data/hf_models/GLM-5.3-NVFP4', 'tokenizer_mode': 'auto', 'tokenizer_backend': 'huggingface', 'tokenizer_worker_num': 1, 'detokenizer_worker_num': 1, 'skip_tokenizer_init': False, 'load_format': 'auto', 'model_loader_extra_config': '{}', 'trust_remote_code': False, 'context_length': None, 'is_embedding': False, 'enable_multimodal': None, 'revision': None, 'model_impl': 'auto', 'model_config_parser': 'auto', 'json_model_override_args': '{"index_topk_freq": 4}', 'dtype': 'auto', 'quantization': None, 'quantization_param_path': None, 'kv_cache_dtype': 'fp8_e4m3', 'enable_fp32_lm_head': False, 'modelopt_quant': None, 'modelopt_checkpoint_restore_path': None, 'modelopt_checkpoint_save_path': None, 'modelopt_export_path': None, 'quantize_and_serve': False, 'rl_quant_profile': None, 'enable_tf32_matmul': False, 'mem_fraction_static': 0.85, 'max_running_requests': 48, 'max_queued_requests': None, 'max_total_tokens': None, 'chunked_prefill_size': 16384, 'prefill_decode_interval': 0, 'enable_dynamic_chunking': False, 'max_prefill_tokens': 16384, 'prefill_max_requests': None, 'schedule_policy': 'fcfs', 'enable_priority_scheduling': False, 'disable_priority_preemption': False, 'default_priority_value': None, 'abort_on_priority_when_disabled': False, 'schedule_low_priority_values_first': False, 'priority_scheduling_preemption_threshold': 10, 'retraction_policy': 'length', 'schedule_conservativeness': 1.0, 'page_size': 64, 'c128_page_size': 16, 'swa_full_tokens_ratio': 0.8, 'disable_hybrid_swa_memory': False, 'radix_eviction_policy': 'lru', 'prefill_only_disable_kv_cache': False, 'disable_radix_cache': True, 'enable_page_major_kv_layout': False, 'enable_unified_memory': False, 'disable_chunked_prefix_cache': False, 'disable_overlap_schedule': True, 'num_continuous_decode_steps': 1, 'scheduler_recv_interval': 1, 'enable_mixed_chunk': False, 'nccl_port': None, 'dist_timeout': None, 'dist_init_addr': None, 'gated_launch_port': None, 'nnodes': 1, 'node_rank': 0, 'tp_size': 2, 'dcp_size': 1, 'pp_size': 4, 'pp_max_micro_batch_size': None, 'pp_async_batch_depth': 0, 'dp_size': 1, 'load_balance_method': 'round_robin', 'attn_cp_size': 1, 'moe_dp_size': 1, 'dwdp_size': 1, 'dcp_comm_backend': 'ag_rs', 'dcp_replicate_q_proj': None, 'enable_prefill_cp': False, 'cp_strategy': None, 'enable_dsa_cache_layer_split': False, 'enable_dsa_prefill_context_parallel': False, 'dsa_prefill_cp_mode': 'round-robin-split', 'enable_prefill_context_parallel': False, 'prefill_cp_mode': 'in-seq-split', 'enable_cp_decode_attn_tp': False, 'enable_dp_attention': False, 'enable_dp_attention_local_control_broadcast': False, 'enable_dp_lm_head': False, 'enable_tp_lm_head_all_to_all': False, 'enable_attn_tp_input_scattered': False, 'enable_shared_experts_attn_tp': False, 'enable_dense_mlp_attn_tp': False, 'disable_attn_tp_gather': False, 'enable_p2p_check': False, 'device': 'cuda', 'base_gpu_id': 0, 'gpu_id_step': 1, 'random_seed': 1037904306, 'mlx_enable_sampling': False, 'watchdog_timeout': 300, 'soft_watchdog_timeout': None, 'sleep_on_idle': False, 'use_ray': False, 'custom_sigquit_handler': None, 'numa_node': None, 'gc_threshold': None, 'host': '0.0.0.0', 'port': 30000, 'fastapi_root_path': '', 'smg_grpc_mode': False, 'grpc_mode': False, 'grpc_port': None, 'grpc_worker_threads': 4, 'sidecar': None, 'sidecar_args': None, 'skip_server_warmup': False, 'warmups': None, 'enable_http2': False, 'http2_max_concurrent_streams': 200, 'ssl_keyfile': None, 'ssl_certfile': None, 'ssl_ca_certs': None, 'ssl_keyfile_password': None, 'enable_ssl_refresh': False, 'api_key': None, 'admin_api_key': None, 'served_model_name': '/data/hf_models/GLM-5.3-NVFP4', 'weight_version': 'default', 'chat_template': None, 'hf_chat_template_name': None, 'completion_template': None, 'file_storage_path': 'sglang_storage', 'enable_cache_report': False, 'reasoning_parser': None, 'default_chat_template_kwargs': None, 'strip_thinking_cache': False, 'enable_strict_thinking': False, 'tool_call_parser': None, 'tool_server': None, 'sampling_defaults': 'model', 'asr_max_buffer_seconds': 60, 'asr_max_concurrent_sessions': 32, 'preferred_sampling_params': None, 'allow_auto_truncate': False, 'stream_interval': 1, 'batch_notify_size': 16, 'stream_response_default_include_usage': False, 'incremental_streaming_output': False, 'enable_streaming_session': False, 'enable_session_radix_cache': False, 'log_level': 'info', 'log_level_http': None, 'log_requests': False, 'log_requests_level': 2, 'log_requests_format': 'text', 'log_requests_target': None, 'uvicorn_access_log_exclude_prefixes': [], 'crash_dump_folder': None, 'show_time_cost': False, 'enable_metrics': False, 'smg_http_sidecar_port': None, 'enable_mfu_metrics': False, 'enable_metrics_for_all_schedulers': False, 'load_snapshot_publish_interval': 15, 'tokenizer_metrics_custom_labels_header': 'x-custom-labels', 'tokenizer_metrics_allowed_custom_labels': None, 'extra_metric_labels': None, 'bucket_time_to_first_token': None, 'bucket_inter_token_latency': None, 'bucket_e2e_request_latency': None, 'prompt_tokens_buckets': None, 'generation_tokens_buckets': None, 'gc_warning_threshold_secs': 0.0, 'decode_log_interval': 40, 'enable_request_time_stats_logging': False, 'kv_events_config': None, 'load_publish_endpoint': None, 'enable_forward_pass_metrics': False, 'forward_pass_metrics_worker_id': '', 'forward_pass_metrics_ipc_name': None, 'enable_trace': False, 'trace_modules': 'request', 'otlp_traces_endpoint': 'localhost:4317', 'export_metrics_to_file': False, 'export_metrics_to_file_dir': None, 'stat_loggers': None, 'constrained_json_whitespace_pattern': None, 'constrained_json_disable_any_whitespace': False, 'attention_backend': 'dsa', 'decode_attention_backend': None, 'prefill_attention_backend': None, 'sampling_backend': 'flashinfer', 'grammar_backend': 'xgrammar', 'radix_cache_backend': None, 'mm_attention_backend': None, 'fp8_gemm_runner_backend': 'auto', 'fp4_gemm_runner_backend': 'auto', 'bf16_gemm_backend': 'auto', 'dsa_prefill_backend': 'flashinfer_sparse_mla', 'dsv4_prefill_backend': 'auto', 'dsa_decode_backend': 'flashinfer_sparse_mla', 'dsa_paged_mqa_logits_backend': 'auto', 'dsa_topk_backend': 'sgl-kernel', 'disable_flashinfer_autotune': True, 'flashinfer_autotune_skip_ops': None, 'mamba_backend': 'triton', 'cuda_graph_config': {'decode': {'backend': 'full', 'max_bs': 256, 'bs': [1, 2, 4, 8, 12, 16, 24, 32, 40, 48, 56, 64, 72, 80, 88, 96, 104, 112, 120, 128, 136, 144, 152, 160, 168, 176, 184, 192, 200, 208, 216, 224, 232, 240, 248, 256], 'tc_compiler': 'eager', 'full_prefill_max_req': None, 'full_prefill_prefix_chunk_tokens': None}, 'prefill': {'backend': 'breakable', 'max_bs': 2048, 'bs': [4, 8, 12, 16, 20, 24, 28, 32, 48, 64, 80, 96, 112, 128, 144, 160, 176, 192, 208, 224, 240, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 704, 768, 832, 896, 960, 1024, 1280, 1536, 1792, 2048], 'tc_compiler': 'eager', 'full_prefill_max_req': None, 'full_prefill_prefix_chunk_tokens': None}}, 'cuda_graph_backend_decode': None, 'cuda_graph_backend_prefill': None, 'cuda_graph_max_bs_decode': None, 'cuda_graph_max_bs_prefill': None, 'cuda_graph_bs_decode': None, 'cuda_graph_bs_prefill': None, 'cuda_graph_tc_compiler': None, 'disable_prefill_cuda_graph': False, 'disable_decode_cuda_graph': False, 'disable_cuda_graph': False, 'disable_cuda_graph_padding': False, 'enable_profile_cuda_graph': False, 'enable_cudagraph_gc': False, 'debug_cuda_graph': False, 'enable_layerwise_nvtx_marker': False, 'enable_nccl_nvls': False, 'enable_symm_mem': False, 'triton_attention_reduce_in_fp32': False, 'triton_attention_num_kv_splits': 8, 'triton_attention_split_tile_size': None, 'flashinfer_mla_disable_ragged': False, 'enable_fused_qk_norm_rope': False, 'enable_precise_embedding_interpolation': False, 'enable_fused_moe_sum_all_reduce': False, 'enable_deepseek_v4_fp4_indexer': False, 'disable_custom_all_reduce': True, 'enable_mscclpp': False, 'enable_torch_symm_mem': False, 'enable_scattered_sconv': False, 'pre_warm_nccl': False, 'enable_quant_communications': False, 'enable_flashinfer_allreduce_fusion': False, 'enforce_disable_flashinfer_allreduce_fusion': False, 'flashinfer_allreduce_fusion_backend': None, 'enable_aiter_allreduce_fusion': False, 'enable_torch_compile': False, 'enable_torch_compile_debug_mode': False, 'torch_compile_max_bs': 32, 'speculative_algorithm': None, 'speculative_draft_model_path': None, 'speculative_draft_model_revision': None, 'speculative_draft_load_format': None, 'speculative_num_steps': None, 'speculative_eagle_topk': None, 'speculative_num_draft_tokens': None, 'speculative_dflash_block_size': None, 'speculative_dspark_block_size': None, 'speculative_dspark_sps_table_path': None, 'speculative_dspark_confidence_sts_path': None, 'speculative_dspark_align_verify_tokens_to_graph_tier': False, 'speculative_accept_threshold_single': 1.0, 'speculative_accept_threshold_acc': 1.0, 'speculative_use_rejection_sampling': False, 'speculative_token_map': None, 'speculative_attention_mode': 'prefill', 'speculative_draft_attention_backend': None, 'speculative_dsa_topk_backend': 'sgl-kernel', 'speculative_draft_kv_cache_dtype': None, 'speculative_draft_window_size': None, 'speculative_moe_runner_backend': 'flashinfer_cutlass', 'speculative_moe_a2a_backend': None, 'speculative_draft_model_quantization': None, '_speculative_draft_quantization_explicitly_set': False, 'speculative_skip_dp_mlp_sync': False, 'enable_multi_layer_eagle': False, 'speculative_adaptive': False, 'speculative_adaptive_config': None, 'decoupled_spec_bind_endpoint': None, 'decoupled_spec_connect_endpoints': None, 'decoupled_spec_rank': None, 'decoupled_spec_role': 'null', 'spec_trace_dir': None, 'speculative_ngram_min_bfs_breadth': 1, 'speculative_ngram_max_bfs_breadth': 10, 'speculative_ngram_match_type': 'BFS', 'speculative_ngram_max_trie_depth': 18, 'speculative_ngram_capacity': 10000000, 'speculative_ngram_external_corpus_path': None, 'speculative_ngram_external_sam_budget': 0, 'speculative_ngram_external_corpus_max_tokens': 10000000, 'ep_size': 1, 'moe_a2a_backend': 'none', 'enable_w4a4_mxfp4_megamoe': False, 'deepep_v2_mode': 'direct', 'moe_runner_backend': 'flashinfer_cutlass', 'flashinfer_mxfp4_moe_precision': 'default', 'deepep_mode': 'auto', 'fuseep_mode': 2, 'deepep_dispatcher_output_dtype': 'auto', 'ep_num_redundant_experts': 0, 'ep_dispatch_algorithm': None, 'init_expert_location': 'trivial', 'enable_eplb': False, 'eplb_algorithm': 'auto', 'eplb_rebalance_num_iterations': 1000, 'eplb_rebalance_layers_per_chunk': None, 'eplb_min_rebalancing_utilization_threshold': 1.0, 'expert_distribution_recorder_mode': None, 'expert_distribution_recorder_buffer_size': 1000, 'expert_balancedness_report_mode': 'off', 'deepep_config': None, 'moe_dense_tp_size': None, 'elastic_ep_backend': None, 'enable_elastic_expert_backup': False, 'mooncake_ib_device': None, 'enable_waterfill': False, 'ep_join_mode': None, 'ep_join_rank_offset': 0, 'elastic_ep_initial_size': None, 'max_ep_size': None, 'elastic_ep_scale_timeout': 600, 'elastic_ep_rejoin': False, 'disable_flashinfer_cutlass_moe_fp4_allgather': False, 'disable_shared_experts_fusion': True, 'enforce_shared_experts_fusion': False, 'max_mamba_cache_size': None, 'mamba_ssm_dtype': None, 'mamba_max_states_per_path': -1, 'enable_mamba_cache_stochastic_rounding': False, 'mamba_cache_philox_rounds': 0, 'mamba_full_memory_ratio': 0.9, 'mamba_radix_cache_strategy': 'auto', 'uses_mamba_radix_cache': False, 'mamba_track_interval': 256, 'enable_int8_mamba_checkpoint': False, 'int8_mamba_ckpt_size': None, 'linear_attn_backend': 'triton', 'linear_attn_decode_backend': None, 'linear_attn_prefill_backend': None, 'linear_attn_verify_backend': None, 'enable_linear_replayssm': False, 'linear_replayssm_cache_len': 16, 'enable_linear_replayssm_spec': False, 'enable_hierarchical_cache': False, 'hicache_host_memory_mode': 'cache', 'hicache_ratio': 2.0, 'hicache_size': 0, 'hicache_write_policy': 'write_through', 'hicache_io_backend': 'kernel', 'hicache_mem_layout': 'page_first', 'hicache_storage_backend': None, 'hicache_storage_prefetch_policy': 'timeout', 'hicache_storage_backend_extra_config': None, 'enable_hisparse': False, 'hisparse_config': None, 'enable_broadcast_mm_inputs_process': False, 'enable_prefix_mm_cache': False, 'mm_enable_dp_encoder': False, 'mm_process_config': {}, 'mm_processor_worker_num': 0, 'mm_io_worker_num': 0, 'allowed_media_domains': [], 'media_url_max_file_size_mb': 64, 'mm_preprocess_cache_size_mb': None, 'trust_mm_content_hashes': False, 'limit_mm_data_per_request': None, 'enable_mm_global_cache': False, 'image_processor_backend': 'auto', 'mm_global_cache_backend': 'mooncake', 'disable_fast_image_processor': False, 'mm_feature_transport': 'cpu', 'keep_mm_feature_on_device': False, 'enable_lora': None, 'enable_lora_overlap_loading': None, 'max_lora_rank': None, 'lora_target_modules': None, 'lora_paths': None, 'max_loaded_loras': None, 'max_loras_per_batch': 8, 'lora_eviction_policy': 'lru', 'lora_backend': 'csgmv', 'max_lora_chunk_size': 16, 'experts_shared_outer_loras': None, 'lora_use_virtual_experts': False, 'lora_strict_loading': False, 'lora_drain_wait_threshold': 0.0, 'enable_two_batch_overlap': False, 'enable_single_batch_overlap': False, 'tbo_token_distribution_threshold': 0.48, 'cpu_offload_gb': 0, 'offload_group_size': -1, 'offload_num_in_group': 1, 'offload_prefetch_step': 1, 'offload_mode': 'cpu', 'enable_lmcache': False, 'lmcache_config_file': None, 'enable_flexkv': False, 'flexkv_config_file': None, 'kt_weight_path': None, 'kt_method': 'AMXINT4', 'kt_cpuinfer': None, 'kt_threadpool_count': 2, 'kt_num_gpu_experts': None, 'kt_max_deferred_experts_per_token': None, 'dllm_algorithm': None, 'dllm_algorithm_config': None, 'dllm_fdfo': True, 'disaggregation_mode': 'null', 'disaggregation_transfer_backend': 'mooncake', 'disaggregation_bootstrap_port': 8998, 'disaggregation_ib_device': None, 'disaggregation_decode_enable_radix_cache': False, 'disaggregation_decode_enable_offload_kvcache': False, 'disaggregation_decode_retraction_backup': None, 'num_reserved_decode_tokens': 512, 'disaggregation_decode_extra_slots': None, 'disaggregation_decode_polling_interval': 1, 'optimistic_prefill_attempts': 0, 'encoder_only': False, 'language_only': False, 'language_model_only': False, 'encoder_transfer_backend': 'zmq_to_scheduler', 'encoder_urls': [], 'encoder_bootstrap_port': 8997, 'encoder_register_urls': [], 'enable_adaptive_dispatch_to_encoder': False, 'enable_pdmux': False, 'pdmux_config_path': None, 'sm_group_num': 8, 'startup_weight_load_mode': 'serial', 'custom_weight_loader': [], 'weight_loader_disable_mmap': False, 'weight_loader_prefetch_checkpoints': False, 'weight_loader_prefetch_num_threads': 4, 'weight_loader_drop_cache_after_load': False, 'remote_instance_weight_loader_seed_instance_ip': None, 'remote_instance_weight_loader_seed_instance_service_port': None, 'remote_instance_weight_loader_send_weights_group_ports': None, 'remote_instance_weight_loader_backend': 'nccl', 'remote_instance_weight_loader_start_seed_via_transfer_engine': False, 'engine_info_bootstrap_port': 6789, 'modelexpress_config': None, 'download_dir': None, 'model_checksum': None, 'delete_ckpt_after_loading': False, 'decrypted_config_file': None, 'decrypted_draft_config_file': None, 'checkpoint_engine_wait_weights_before_ready': False, 'enable_prefill_delayer': False, 'prefill_delayer_max_delay_passes': 30, 'prefill_delayer_token_usage_low_watermark': None, 'prefill_delayer_forward_passes_buckets': None, 'prefill_delayer_wait_seconds_buckets': None, 'prefill_delayer_queue_min_ratio': None, 'prefill_delayer_max_delay_ms': None, 'min_free_slots_delay': None, 'enable_deterministic_inference': False, 'rl_on_policy_target': None, 'kv_canary': 'none', 'kv_canary_real_data': 'none', 'kv_canary_sweep_interval': 0, 'enable_dynamic_batch_tokenizer': False, 'dynamic_batch_tokenizer_batch_size': 32, 'dynamic_batch_tokenizer_batch_timeout': 0.002, 'enable_tokenizer_batch_encode': False, 'disable_tokenizer_batch_decode': False, 'debug_tensor_dump_output_folder': None, 'debug_tensor_dump_layers': None, 'debug_tensor_dump_input_file': None, 'enable_memory_saver': False, 'enable_weights_cpu_backup': False, 'enable_draft_weights_cpu_backup': False, 'enable_custom_logit_processor': False, 'enable_return_hidden_states': False, 'return_hidden_states_mode': None, 'enable_return_routed_experts': False, 'enable_return_indexer_topk': False, 'disable_outlines_disk_cache': False, 'enable_mis': False, 'weight_cache_mode': 'off', 'weight_cache_socket': None, 'weight_cache_timeout': 1800, 'forward_hooks': None, 'msprobe_dump_config': None} +[2026-09-10 03:23:39 PP3 TP1] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 1040384, KV size: 13.35 GB +[2026-09-10 03:23:39 PP1 TP0] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 1040384, KV size: 12.59 GB +[2026-09-10 03:23:39 PP2 TP0] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 1040384, KV size: 13.35 GB +[2026-09-10 03:23:39 PP0 TP1] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 1040384, KV size: 12.97 GB +[2026-09-10 03:23:39 PP0 TP0] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 1040384, KV size: 12.97 GB +[2026-09-10 03:23:39 PP2 TP1] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 1040384, KV size: 13.35 GB +[2026-09-10 03:23:39 PP1 TP1] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 1040384, KV size: 12.59 GB +[2026-09-10 03:23:39 PP3 TP0] KV Cache is allocated. dtype: torch.float8_e4m3fn, #tokens: 1040384, KV size: 13.35 GB +[2026-09-10 03:25:26 PP0 TP0] max_total_num_tokens=1040384, chunked_prefill_size=16384, max_prefill_tokens=16384, max_running_requests=48, context_len=1048576, available_gpu_mem=20.15 GB +[2026-09-10 03:25:26 PP1 TP0] max_total_num_tokens=1040384, chunked_prefill_size=16384, max_prefill_tokens=16384, max_running_requests=48, context_len=1048576, available_gpu_mem=14.08 GB +[2026-09-10 03:25:26 PP2 TP0] max_total_num_tokens=1040384, chunked_prefill_size=16384, max_prefill_tokens=16384, max_running_requests=48, context_len=1048576, available_gpu_mem=10.52 GB +[2026-09-10 03:25:26 PP3 TP0] max_total_num_tokens=1040384, chunked_prefill_size=16384, max_prefill_tokens=16384, max_running_requests=48, context_len=1048576, available_gpu_mem=9.69 GB diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/tp2pp4/status.txt b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/tp2pp4/status.txt new file mode 100644 index 0000000..5d9c6e9 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/results/tp2pp4/status.txt @@ -0,0 +1,24 @@ +b3_16k_c1 OK +b3_16k_c8 OK +b3_16k_c16 OK +b3_16k_c32 OK +b3_16k_c64 OK +b41_1k_c1 OK +b41_1k_c8 OK +b41_1k_c32 OK +b41_1k_c64 OK +b42_1k4k_c1 OK +b42_1k4k_c8 OK +b42_1k4k_c32 OK +b42_1k4k_c64 OK +b51_64k_c1 OK +b51_64k_c4 OK +b51_64k_c8 OK +b51_128k_c1 OK +b51_128k_c2 OK +b51_128k_c4 OK +b52_256k_c1 OK +b52_256k_c2 OK +b52_256k_c3 OK +b52_512k_c1 OK +b52_896k_c1 OK diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/bench_corpus_v2.py b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/bench_corpus_v2.py new file mode 100644 index 0000000..a62ef51 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/bench_corpus_v2.py @@ -0,0 +1,319 @@ +#!/usr/bin/env python3 +"""Real-corpus (PG19) benchmark for sglang GLM-5.3-NVFP4. + +Same methodology as bench_hit90.py (input_ids direct to /generate, temp 0, +ignore_eos, stream, server-side completion_tokens counting, hit-rate verified +from scheduler logs), but prompts are token slices of REAL book text tokenized +with the served model's own tokenizer, replacing random ids. + +Corpus file: JSON {"ids": [flat token ids], "books": [[start, end), ...]} + +Fixed token-offset layout into the flat id array: + [0, 117968) shared prefix for 128k points (90% of 131072) + [0, 58976) shared prefix for 64k points (same region, shorter cut) + pool A [131072, +4*8*13104) 128k unique suffixes, run-ids 9301-9304 + pool B [550400, +4*8*6560) 64k unique suffixes, run-ids 9305-9308 + pool C [760320, 262144+524288+524288) 16k fully-unique prompts, run-ids 9311-9313 + spare [2071040, end) warmup slices / re-run margin + +Each run-id maps to one non-overlapping window (one window = one bench point); +re-running a point with fresh text = bump --pool-override past the spare base. + +v2 changes (b300-equivalent campaign, 2026-09-10): + - stats(): p95 added (nearest-rank) to match the B300 report metric contract + (TTFT P95 / TPOT P95). + - --dump-records PATH: per-request records (ttft/e2e/tpot/n_out/retractions/ + spec_accept_len) written as JSONL for post-hoc percentile checks. + - With --shared-frac 0 + --pool-override, any input length is supported + (1024 / 16384 / 65536 / 131072 / 262144 / 524288 / 917504). + +Usage: + python3 bench_corpus_v2.py --corpus /root/corpus_ids.json --input-len 16384 \ + --concurrency 64 --num-requests 128 --run-id 9505 --shared-frac 0 \ + --pool-override 2300000 --output-len 512 \ + --dump-records /root/bench_logs/xx/point_records.jsonl +""" +import argparse +import datetime +import json +import math +import re +import statistics +import subprocess +import time +from concurrent.futures import ThreadPoolExecutor + +import requests + +OUTPUT_LEN_DEFAULT = 512 +CORPUS_DEFAULT = "/root/corpus_ids.json" + +# fixed pool layout (see docstring) +S1_128K_RID0, S1_64K_RID0, S2_RID0 = 9301, 9305, 9311 +POOL_A_BASE, POOL_A_PER = 131072, 8 * 13104 # 128k suffix windows +POOL_B_BASE = POOL_A_BASE + 4 * POOL_A_PER # 550400 +POOL_B_PER = 8 * 6560 # 64k suffix windows +POOL_C_BASE = POOL_B_BASE + 4 * POOL_B_PER # 760320 +POOL_C_SIZES = [16 * 16384, 32 * 16384, 32 * 16384] # cc8 / cc16 / cc32 +SPARE_BASE = POOL_C_BASE + sum(POOL_C_SIZES) # 2071040 + +sess = requests.Session() +sess.trust_env = False # bypass any proxy env on the host + + +def split_lens(input_len, shared_frac): + # unique suffix = (1 - shared_frac) of the prompt, page-16 aligned + # (128k @0.9 -> 13104 unique; 16k @0.0 -> fully unique prompts) + unique = round(input_len * (1.0 - shared_frac) / 16) * 16 + return input_len - unique, unique + + +def pool_start_for(input_len, shared_frac, run_id, override): + if override is not None: + return override + if shared_frac > 0: + if input_len == 131072: + idx = run_id - S1_128K_RID0 + if not 0 <= idx < 4: + sys_exit_bad_runid(run_id, "128k points use run-ids 9301-9304") + return POOL_A_BASE + idx * POOL_A_PER + if input_len == 65536: + idx = run_id - S1_64K_RID0 + if not 0 <= idx < 4: + sys_exit_bad_runid(run_id, "64k points use run-ids 9305-9308") + return POOL_B_BASE + idx * POOL_B_PER + sys_exit_bad_runid(run_id, "shared-frac>0 supports 131072/65536 only") + idx = run_id - S2_RID0 + if not 0 <= idx < 3: + sys_exit_bad_runid(run_id, "16k unique points use run-ids 9311-9313") + return POOL_C_BASE + sum(POOL_C_SIZES[:idx]) + + +def sys_exit_bad_runid(run_id, msg): + raise SystemExit(f"[pool] run-id {run_id} outside expected set: {msg}") + + +def build_prompts(ids, shared_len, unique_len, num_requests, pool_start): + if shared_len: + shared = ids[0:shared_len] + else: + shared = [] + end = pool_start + num_requests * unique_len + if end > len(ids): + raise SystemExit( + f"[pool] window [{pool_start}, {end}) exceeds corpus ({len(ids)} ids); " + f"use --pool-override or a larger corpus") + prompts = [] + for i in range(num_requests): + s = pool_start + i * unique_len + prompts.append(shared + ids[s:s + unique_len]) + return shared, prompts, (pool_start, end) + + +def warmup(url, ids, shared_len): + # primes the radix cache with the shared prefix (same role as in bench_hit90); + # warm slice comes from the spare region so it never collides with a pool window + if len(ids) >= SPARE_BASE + 64: + warm_slice = ids[SPARE_BASE:SPARE_BASE + 64] + else: + warm_slice = ids[-64:] + payload = { + "input_ids": ids[0:shared_len] + warm_slice if shared_len else warm_slice, + "sampling_params": {"max_new_tokens": 8, "temperature": 0.0, "ignore_eos": True}, + } + t0 = time.perf_counter() + r = sess.post(url, json=payload, timeout=1800) + dt = time.perf_counter() - t0 + print(f"[warmup] http={r.status_code} wall={dt:.2f}s", flush=True) + + +def bench_one(url, prompt, output_len, idx, results): + payload = { + "input_ids": prompt, + "sampling_params": {"max_new_tokens": output_len, "temperature": 0.0, "ignore_eos": True}, + "stream": True, + } + rec = {"idx": idx} + t0 = time.perf_counter() + first = last = None + first_ct = None + final_meta = None + max_ct = 0 + try: + with sess.post(url, json=payload, stream=True, timeout=3600) as resp: + for raw in resp.iter_lines(): + if not raw or not raw.startswith(b"data:"): + continue + body = raw[5:].strip() + if body == b"[DONE]": + continue + now = time.perf_counter() + try: + d = json.loads(body) + except Exception: + continue + mi = d.get("meta_info") or {} + ct = mi.get("completion_tokens") or 0 + if ct: + max_ct = max(max_ct, ct) + if first is None: + first = now + first_ct = ct + last = now + if mi.get("finish_reason"): + final_meta = mi + t_end = time.perf_counter() + n_out = max_ct or ((final_meta or {}).get("completion_tokens") or 0) + decode_span = (last - first) if (first and last and last > first) else 0.0 + rec.update( + ok=n_out > 0, + ttft=(first - t0) if first else None, + e2e=t_end - t0, + n_out=n_out, + first_chunk_tokens=first_ct, + decode_span=decode_span, + tpot=(decode_span / (n_out - 1)) if n_out > 1 else None, + per_req_decode_tok_s=(n_out / decode_span) if decode_span > 0 else None, + retractions=(final_meta or {}).get("num_retractions"), + spec_accept_len=(final_meta or {}).get("spec_accept_length"), + ) + except Exception as e: + rec.update(ok=False, error=repr(e)) + results[idx] = rec + + +def verify_hit_rate(container, t_start, t_end): + def rfc3339(epoch): + return (datetime.datetime.fromtimestamp(epoch, tz=datetime.timezone.utc) + .isoformat().replace("+00:00", "Z")) + + try: + # No margin before t_start: warmup's prefill lines end strictly before it, + # and catching them would deflate the measured hit rate. + p = subprocess.run( + ["docker", "logs", container, "--since", rfc3339(t_start), "--until", rfc3339(t_end + 2)], + capture_output=True, text=True, timeout=120) + text = p.stdout + p.stderr + except Exception as e: + return {"error": repr(e)} + pat = re.compile(r"#new-token: (\d+).*?#cached-token: (\d+)") + n_batches = new_tok = cached_tok = 0 + for line in text.splitlines(): + if "TP0]" not in line or "Prefill batch" not in line: + continue + m = pat.search(line) + if m: + n_batches += 1 + new_tok += int(m.group(1)) + cached_tok += int(m.group(2)) + total = new_tok + cached_tok + return { + "prefill_batches": n_batches, + "new_tokens": new_tok, + "cached_tokens": cached_tok, + "hit_rate": round(cached_tok / total, 4) if total else None, + } + + +def stats(vals): + vals = [v for v in vals if v is not None] + if not vals: + return {"mean": None, "p50": None, "p95": None, "max": None, "min": None} + s = sorted(vals) + # nearest-rank p95: smallest value >= 95th percentile + p95_idx = max(0, math.ceil(0.95 * len(s)) - 1) + return { + "mean": round(statistics.fmean(vals), 4), + "p50": round(s[len(s) // 2], 4), + "p95": round(s[p95_idx], 4), + "max": round(s[-1], 4), + "min": round(s[0], 4), + } + + +def main(): + ap = argparse.ArgumentParser() + ap.add_argument("--concurrency", type=int, required=True) + ap.add_argument("--num-requests", type=int, required=True) + ap.add_argument("--run-id", type=int, required=True) + ap.add_argument("--input-len", type=int, required=True, help="token length of each prompt") + ap.add_argument("--output-len", type=int, default=OUTPUT_LEN_DEFAULT) + ap.add_argument("--shared-frac", type=float, default=0.9) + ap.add_argument("--corpus", default=CORPUS_DEFAULT) + ap.add_argument("--pool-override", type=int, default=None, + help="explicit corpus offset for the unique-suffix window (re-runs)") + ap.add_argument("--url", default="http://127.0.0.1:30000/generate") + ap.add_argument("--container", default="glm53-nvfp4") + ap.add_argument("--dump-records", default=None, + help="write per-request records as JSONL to this path") + args = ap.parse_args() + + with open(args.corpus) as f: + corpus = json.load(f) + ids = corpus["ids"] + + shared_len, unique_len = split_lens(args.input_len, args.shared_frac) + pool_start = pool_start_for(args.input_len, args.shared_frac, args.run_id, args.pool_override) + shared, prompts, window = build_prompts(ids, shared_len, unique_len, args.num_requests, pool_start) + print(f"[pool] window={window} shared_len={shared_len} unique_len={unique_len} " + f"corpus_total={len(ids)}", flush=True) + warmup(args.url, ids, shared_len) + + results = {} + t_start = time.time() + t0 = time.perf_counter() + with ThreadPoolExecutor(max_workers=args.concurrency) as ex: + futs = [ex.submit(bench_one, args.url, p, args.output_len, i, results) + for i, p in enumerate(prompts)] + for f in futs: + f.result() + wall = time.perf_counter() - t0 + t_end = time.time() + + if args.dump_records: + with open(args.dump_records, "w") as f: + for i in sorted(results): + f.write(json.dumps(results[i]) + "\n") + + hit = verify_hit_rate(args.container, t_start, t_end) + + ok = [r for r in results.values() if r.get("ok")] + n_out_total = sum(r["n_out"] for r in ok) + out_tps = [r["n_out"] / r["e2e"] for r in ok if r.get("e2e")] + ttft = stats([r.get("ttft") for r in ok]) + tpot = stats([r.get("tpot") for r in ok]) + e2e = stats([r.get("e2e") for r in ok]) + dec = stats([r.get("per_req_decode_tok_s") for r in ok]) + spec = [v for v in (r.get("spec_accept_len") for r in ok) if v is not None] + retr = sum(r.get("retractions") or 0 for r in ok) + + summary = { + "concurrency": args.concurrency, + "num_requests": args.num_requests, + "run_id": args.run_id, + "corpus_window": {"start": window[0], "end": window[1]}, + "ok": len(ok), + "failed": args.num_requests - len(ok), + "wall_s": round(wall, 2), + "input_len": args.input_len, + "shared_len": shared_len, + "unique_len": unique_len, + "output_len": args.output_len, + "output_tokens_total": n_out_total, + "output_throughput_tok_s": round(n_out_total / wall, 2) if wall else None, + "input_throughput_tok_s": round(args.input_len * len(ok) / wall, 2) if wall else None, + "ttft_s": ttft, + "tpot_s": tpot, + "e2e_s": e2e, + "per_req_out_tok_s_e2e": stats(out_tps), + "per_req_decode_tok_s": dec, + "spec_accept_length_mean": round(statistics.fmean(spec), 3) if spec else None, + "retractions_total": retr, + "cache_hit_from_logs": hit, + } + print("\n===== SUMMARY =====") + print(json.dumps(summary, indent=2), flush=True) + + +if __name__ == "__main__": + main() diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/extract_summary.py b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/extract_summary.py new file mode 100644 index 0000000..353f4f5 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/extract_summary.py @@ -0,0 +1,39 @@ +#!/usr/bin/env python3 +"""extract_summary.py + +Parse the trailing "===== SUMMARY =====" JSON block from a bench_corpus_v2 log, +append a tagged record to the campaign results JSONL, and print the headline +metrics. Exit codes: 0 = ok and hit_rate<=0.01 (cold-point validity), +3 = no SUMMARY block found, 4 = hit_rate above the cold-point threshold. +""" +import json +import sys + +HIT_MAX = 0.01 +MARK = "===== SUMMARY =====" + + +def main(): + logp, tag, arm, outp = sys.argv[1:5] + with open(logp, encoding="utf-8", errors="replace") as f: + log = f.read() + i = log.rfind(MARK) + if i < 0: + print("NO_SUMMARY") + sys.exit(3) + s = json.loads(log[i + len(MARK):].strip()) + with open(outp, "a", encoding="utf-8") as f: + f.write(json.dumps({"tag": tag, "arm": arm, "summary": s}) + "\n") + hit = (s.get("cache_hit_from_logs") or {}).get("hit_rate") + print(f"hit_rate={hit} out_tps={s.get('output_throughput_tok_s')} " + f"in_tps={s.get('input_throughput_tok_s')} " + f"ttft_p95={(s.get('ttft_s') or {}).get('p95')} " + f"tpot_p95={(s.get('tpot_s') or {}).get('p95')} " + f"ok={s.get('ok')}/{s.get('num_requests')} retractions={s.get('retractions_total')}") + if hit is None: + sys.exit(4) + sys.exit(0 if float(hit) <= HIT_MAX else 4) + + +if __name__ == "__main__": + main() diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/gen_report_tables.py b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/gen_report_tables.py new file mode 100644 index 0000000..268ff02 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/gen_report_tables.py @@ -0,0 +1,165 @@ +#!/usr/bin/env python3 +"""gen_report_tables.py + +Build the B300-style markdown tables for the 6000D dual-plan report from the +campaign all_results.jsonl files. Prints tables to stdout. Best value per +column within each scenario table is bolded (**...**), mirroring the B300 +report convention. +""" +import json +import sys + +ARMS = ["tp2pp4", "e7b"] +ARM_LABEL = {"tp2pp4": "TP2PP4", "e7b": "TP8+EAGLE3+AR"} + +# scenario key -> (chapter title, row label, ordered ccs, has_input_col) +SCENARIOS = [ + ("b3_16k", "主场景 16K→512", [1, 8, 16, 32, 64], True), + ("b41_1k", "4.1 短输入 1K→128", [1, 8, 32, 64], True), + ("b42_1k4k", "4.2 长输出 1K→4K", [1, 8, 32, 64], False), + ("b51_64k", "5.1 64K→512", [1, 4, 8], True), + ("b51_128k", "5.1 128K→512", [1, 2, 4], True), + ("b52_256k", "5.2 边界 256K→1", [1, 2, 3], True), + ("b52_512k", "5.2 边界 512K→1", [1], True), + ("b52_896k", "5.2 边界 896K→1", [1], True), +] + + +def load(path): + rows = {} + if not path: + return rows + with open(path, encoding="utf-8") as f: + for line in f: + rec = json.loads(line) + rows[(rec["tag"], rec["arm"])] = rec["summary"] + return rows + + +def fmt(v, kind): + if v is None: + return "—" + if kind == "tps": + return f"{v:,.0f}" if v >= 100 else f"{v:,.1f}" + if kind == "s": + return f"{v:,.2f} s" if v >= 1 else f"{v*1000:,.0f} ms" + if kind == "ms": + return f"{v*1000:,.1f} ms" + return str(v) + + +def get(s, path): + cur = s + for key in path.split("."): + if cur is None: + return None + cur = cur.get(key) + return cur + + +def table_for(prefix, title, ccs, has_input, data): + lines = [] + cols = ["并发", "方案"] + if has_input: + cols += ["Input TPS"] + cols += ["Output TPS", "TTFT P95", "TPOT P95"] + lines.append("| " + " | ".join(cols) + " |") + lines.append("|" + "---|" * len(cols)) + + # collect cells to bold best per numeric column + body = [] + for cc in ccs: + for arm in ARMS: + tag = f"{prefix}_c{cc}" + s = data.get((tag, arm)) + if s is None and cc in (2, 3): + # boundary cc>1 only exists on tp2pp4; absence = structural skip + continue + if s is None: + body.append((cc, arm, None)) + continue + body.append((cc, arm, s)) + + def val(cc, arm, path): + s = data.get((f"{prefix}_c{cc}", arm)) + return get(s, path) if s else None + + # best per column (higher tps, lower latency) + best = {} + if has_input: + ivs = [val(cc, arm, "input_throughput_tok_s") for cc, arm, _ in body if _ is not None] + ivs = [v for v in ivs if v is not None] + if ivs: + best["in"] = max(ivs) + ovs = [val(cc, arm, "output_throughput_tok_s") for cc, arm, _ in body if _ is not None] + ovs = [v for v in ovs if v is not None] + if ovs: + best["out"] = max(ovs) + tvs = [val(cc, arm, "ttft_s.p95") for cc, arm, _ in body if _ is not None] + tvs = [v for v in tvs if v is not None] + if tvs: + best["ttft"] = min(tvs) + pvs = [val(cc, arm, "tpot_s.p95") for cc, arm, _ in body if _ is not None] + pvs = [v for v in pvs if v is not None] + if pvs: + best["tpot"] = min(pvs) + + def maybe_bold(v, kind, key): + if v is None: + return "—" + cell = fmt(v, kind) + if key in best and v == best[key]: + return f"**{cell}**" + return cell + + for cc, arm, s in body: + if s is None: + lines.append(f"| {cc} | {ARM_LABEL[arm]} | " + " 结构性不可测 |" * (len(cols) - 2)) + continue + cells = [str(cc), ARM_LABEL[arm]] + if has_input: + cells.append(maybe_bold(get(s, "input_throughput_tok_s"), "tps", "in")) + cells.append(maybe_bold(get(s, "output_throughput_tok_s"), "tps", "out")) + cells.append(maybe_bold(get(s, "ttft_s.p95"), "s", "ttft")) + cells.append(maybe_bold(get(s, "tpot_s.p95"), "ms", "tpot")) + lines.append("| " + " | ".join(cells) + " |") + return "\n".join(lines) + + +def main(): + tp_data = load(sys.argv[1] if len(sys.argv) > 1 else None) + e7_data = load(sys.argv[2] if len(sys.argv) > 2 else None) + data = {**tp_data, **e7_data} + for prefix, title, ccs, has_input in SCENARIOS: + print(f"\n### {title}\n") + print(table_for(prefix, title, ccs, has_input, data)) + + # appendix: full metrics + print("\n\n## 附录:全量指标(含 mean/p50/max、回退、投机接受长度)\n") + print("| 场景点 | 方案 | ok/nreq | wall(s) | out tok/s | in tok/s | TTFT mean/p95/max(s) | TPOT mean/p95/max(ms) | retractions | accept |") + print("|---|---|---|---|---|---|---|---|---|---|") + for prefix, _, ccs, _ in SCENARIOS: + for cc in ccs: + for arm in ARMS: + s = data.get((f"{prefix}_c{cc}", arm)) + if s is None: + continue + ttft = s.get("ttft_s") or {} + tpot = s.get("tpot_s") or {} + def trio(d): + m, p, x = d.get("mean"), d.get("p95"), d.get("max") + if m is None: + return "—" + return f"{m:.2f}/{p:.2f}/{x:.2f}" + def trioms(d): + m, p, x = d.get("mean"), d.get("p95"), d.get("max") + if m is None: + return "—" + return f"{m*1000:.1f}/{p*1000:.1f}/{x*1000:.1f}" + print(f"| {prefix}_c{cc} | {ARM_LABEL[arm]} | {s.get('ok')}/{s.get('num_requests')} " + f"| {s.get('wall_s')} | {s.get('output_throughput_tok_s')} | {s.get('input_throughput_tok_s')} " + f"| {trio(ttft)} | {trioms(tpot)} | {s.get('retractions_total')} | {s.get('spec_accept_length_mean')} |") + + +if __name__ == "__main__": + main() diff --git a/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/run_b300_matrix.sh b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/run_b300_matrix.sh new file mode 100644 index 0000000..98efe49 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/scripts/run_b300_matrix.sh @@ -0,0 +1,168 @@ +#!/bin/bash +# run_b300_matrix.sh — B300-equivalent cold-cache scenario matrix on 60.8 +# +# Mirrors the B300 report scenario set (16K->512, 1K->128, 1K->4K, 64K->512, +# 128K->512, 256K/512K/896K boundary OSL=1) at reduced concurrency per team +# decision (16K/1K capped at cc64; 64K/128K at cc<=8; boundary cc<=3). +# Both arms run the same grid; E7b structurally skips 256K cc>1 / 512K / 896K +# (KV pool 276,864, ctx 270,336). +# +# Protocol: cold points (shared-frac 0), recycled corpus windows via +# --pool-override (corpus has no virgin text left), per-point idle-wait -> +# scenario-length prewarm (first point of each scenario) -> flush_cache -> +# bench -> hit-rate-from-logs verification (<=0.01 else one retry, then abort). +# run-ids 95xx are labels only (windows are explicit). +# +# Usage: nohup bash /root/run_b300_matrix.sh \ +# > /root/bench_logs/b300eq__progress.log 2>&1 & +set -u +ARM=${1:?usage: run_b300_matrix.sh tp2pp4|e7b} +case $ARM in + tp2pp4) CONTAINER=glm53-pp4 ;; + e7b) CONTAINER=glm53-nvfp4 ;; + *) echo "bad arm: $ARM"; exit 1 ;; +esac + +CORPUS=/root/corpus_ids.json +BENCH=/root/bench_corpus_v2.py +EXTRACT=/root/extract_summary.py +URL=http://127.0.0.1:30000 +STAMP=$(date +%Y%m%d_%H%M) +LOG=/root/bench_logs/b300eq_${ARM}_${STAMP} +mkdir -p "$LOG" +RESULTS="$LOG/all_results.jsonl" +STATUS="$LOG/status.txt" +: > "$STATUS" + +echo "=== [$ARM] matrix start $(date) LOGDIR=$LOG ===" + +# ---- server facts snapshot ---- +alive() { [ -n "$(docker ps --filter name=$CONTAINER --filter status=running -q)" ]; } +alive || { echo "=== [$ARM] ABORT: container $CONTAINER not running ==="; exit 1; } +docker inspect "$CONTAINER" --format '{{.Config.Cmd}}' > "$LOG/server_cmd.txt" 2>&1 +docker logs "$CONTAINER" 2>&1 | grep -E "max_total_num_tokens|KV Cache is allocated|context_len|chunked_prefill_size|max_running_request|speculative_num_steps" | head -20 > "$LOG/server_facts.txt" 2>&1 +nvidia-smi --query-gpu=index,name,memory.total,memory.used --format=csv > "$LOG/gpu_inventory_idle.csv" 2>&1 +md5sum "$CORPUS" "$BENCH" "$EXTRACT" > "$LOG/md5_assets.txt" 2>&1 +echo "--- server_cmd: $(cat "$LOG/server_cmd.txt")" +echo "--- server_facts:"; cat "$LOG/server_facts.txt" + +# ---- per-rank VRAM sampler (whole-arm timeline, 30s cadence) ---- +( while true; do + echo "# $(date +%s) $(date +%T)" + nvidia-smi --query-gpu=index,memory.used,utilization.gpu --format=csv,noheader + sleep 30 + done ) > "$LOG/vram_timeline.csv" 2>&1 & +SAMPLER=$! +trap 'kill $SAMPLER 2>/dev/null' EXIT + +idle_wait() { + for i in $(seq 1 90); do + local last + last=$(docker logs --since 90s "$CONTAINER" 2>&1 | grep 'running-req' | tail -1) + if [ -z "$last" ] || echo "$last" | grep -q 'running-req: 0'; then return 0; fi + sleep 10 + done + echo "[warn] idle_wait timeout, continuing" +} + +flush() { curl -s -m 60 -X POST $URL/flush_cache >/dev/null; sleep 3; } + +prewarm() { # $1=input_len $2=base — one uncounted request at scenario length + timeout 1800 python3 "$BENCH" --corpus "$CORPUS" --input-len "$1" --output-len 16 \ + --shared-frac 0 --concurrency 1 --num-requests 1 --run-id 9599 \ + --pool-override "$2" --url $URL/generate --container "$CONTAINER" \ + > "$LOG/prewarm_$1.log" 2>&1 || true +} + +bench_call() { # $1=log $2=tag $3=isl $4=osl $5=cc $6=nreq $7=base $8=rid + timeout 7200 python3 "$BENCH" --corpus "$CORPUS" --input-len "$3" --output-len "$4" \ + --shared-frac 0 --concurrency "$5" --num-requests "$6" --run-id "$8" \ + --pool-override "$7" --url $URL/generate --container "$CONTAINER" \ + --dump-records "$LOG/${2}_records.jsonl" > "$1" 2>&1 +} + +run_point() { # tag isl osl cc nreq base rid [prewarm=1] + local TAG=$1 ISL=$2 OSL=$3 CC=$4 NREQ=$5 BASE=$6 RID=$7 PW=${8:-0} + alive || { echo "$TAG CONTAINER_DEAD" >> "$STATUS"; echo "=== $TAG ABORT: container dead ==="; exit 1; } + echo "=== [$ARM] $TAG isl=$ISL osl=$OSL cc=$CC nreq=$NREQ base=$BASE start $(date +%T) ===" + idle_wait + [ "$PW" = "1" ] && prewarm "$ISL" "$BASE" + flush + local rc=0 + bench_call "$LOG/${TAG}.log" "$TAG" "$ISL" "$OSL" "$CC" "$NREQ" "$BASE" "$RID" || rc=$? + if [ "$rc" = "0" ]; then + python3 "$EXTRACT" "$LOG/${TAG}.log" "$TAG" "$ARM" "$RESULTS"; rc=$? + fi + if [ "$rc" = "4" ]; then + echo "=== $TAG hit_rate>0.01, retrying once ===" + idle_wait; flush + bench_call "$LOG/${TAG}_retry.log" "${TAG}_retry" "$ISL" "$OSL" "$CC" "$NREQ" "$BASE" "$RID" || rc=$? + if [ "$rc" = "0" ]; then + python3 "$EXTRACT" "$LOG/${TAG}_retry.log" "${TAG}_retry" "$ARM" "$RESULTS"; rc=$? + fi + if [ "$rc" = "4" ]; then + echo "$TAG HIT_FAIL_FINAL" >> "$STATUS" + echo "=== $TAG hit_rate still >0.01 after retry — ABORTING ARM (contamination) ===" + exit 1 + fi + fi + if [ "$rc" = "0" ]; then + echo "$TAG OK" >> "$STATUS" + else + echo "$TAG BENCH_FAIL rc=$rc" >> "$STATUS" + echo "=== $TAG bench rc=$rc (recorded, continuing) ===" + fi + echo "=== $TAG done $(date +%T) ===" +} + +# ---- recycled corpus window bases (see campaign window map; all within 21.3M) ---- +B_16K=2300000; B_1K=4500000; B_1K4=4700000; B_64K=5000000 +B_128K=6200000; B_256K=8400000; B_512K=9300000; B_896K=9900000 + +# ===== B300 §3: main scenario 16K -> 512 ===== +run_point b3_16k_c1 16384 512 1 8 $B_16K 9501 1 +run_point b3_16k_c8 16384 512 8 16 $B_16K 9502 +run_point b3_16k_c16 16384 512 16 32 $B_16K 9503 +run_point b3_16k_c32 16384 512 32 64 $B_16K 9504 +run_point b3_16k_c64 16384 512 64 128 $B_16K 9505 + +# ===== B300 §4.1: short input 1K -> 128 ===== +run_point b41_1k_c1 1024 128 1 8 $B_1K 9511 1 +run_point b41_1k_c8 1024 128 8 16 $B_1K 9512 +run_point b41_1k_c32 1024 128 32 64 $B_1K 9513 +run_point b41_1k_c64 1024 128 64 128 $B_1K 9514 + +# ===== B300 §4.2: long output 1K -> 4K ===== +run_point b42_1k4k_c1 1024 4096 1 8 $B_1K4 9521 1 +run_point b42_1k4k_c8 1024 4096 8 16 $B_1K4 9522 +run_point b42_1k4k_c32 1024 4096 32 64 $B_1K4 9523 +run_point b42_1k4k_c64 1024 4096 64 128 $B_1K4 9524 + +# ===== B300 §5.1: long context 64K -> 512 ===== +run_point b51_64k_c1 65536 512 1 8 $B_64K 9531 1 +run_point b51_64k_c4 65536 512 4 8 $B_64K 9532 +run_point b51_64k_c8 65536 512 8 16 $B_64K 9533 + +# ===== B300 §5.1: long context 128K -> 512 ===== +run_point b51_128k_c1 131072 512 1 8 $B_128K 9541 1 +run_point b51_128k_c2 131072 512 2 8 $B_128K 9542 +run_point b51_128k_c4 131072 512 4 8 $B_128K 9543 + +# ===== B300 §5.2: context boundary, OSL=1 (nreq=cc) ===== +run_point b52_256k_c1 262144 1 1 1 $B_256K 9551 1 +if [ "$ARM" = "tp2pp4" ]; then + run_point b52_256k_c2 262144 1 2 2 $B_256K 9552 + run_point b52_256k_c3 262144 1 3 3 $B_256K 9553 + run_point b52_512k_c1 524288 1 1 1 $B_512K 9554 + run_point b52_896k_c1 917504 1 1 1 $B_896K 9555 +else + # E7b structural limits: pool 276,864 (256K cc>=2 needs 524K) and ctx 270,336 (<512K) + for t in b52_256k_c2 b52_256k_c3 b52_512k_c1 b52_896k_c1; do + echo "$t SKIP_STRUCTURAL" >> "$STATUS" + done + echo "=== [e7b] 256K cc2/3, 512K, 896K skipped: structural (pool 276,864 / ctx 270,336) ===" +fi + +nvidia-smi --query-gpu=index,name,memory.total,memory.used --format=csv > "$LOG/gpu_inventory_final.csv" 2>&1 +echo "=== [$ARM] MATRIX ALL DONE $(date) — $LOG ===" +echo "--- status:"; cat "$STATUS"