Compare commits

..

2 Commits

19 changed files with 2162 additions and 2 deletions

View File

@ -1,4 +1,4 @@
# 现役部署状态页live 核验于 2026-09-09
# 现役部署状态页live 核验于 2026-09-1060.8 当日核验;其余机器 09-09 口径
> 本页记录 174.1.60.x 集群"现在跑的是什么"。改动机器前后先读这页;
> **任何变动(停/起/改配置)完成后必须更新本页**。核验命令:`docker ps` + `nvidia-smi`
@ -14,7 +14,7 @@
| 60.5 | `glm53-nvfp4`Up 2d09-09 只读核验) | **NVFP4 团队生产**deploy_glm53_605.shmd5 fcd9109b。生产机铁律不实验、不重启、不覆盖脚本。**交付升级路径已更新为 v309-09 hit90 场景优胜 TP4PP2-hicache@0.90,池 647,040/c4 并发独立文档/hit90 cc8 out +34% vs 现役):脚本 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/scripts/deploy_glm53_605_v3.sh` + 补丁束60.8:/root/glm53_r37_patch_bundle_v3.tar.gzmd5 6922e534需 scp 至 60.5+ 需 docker pull cu13 镜像;未执行、未落 60.5 磁盘60.5:/root 仅有原脚本,核验过)。此前 v2TP2PP4-hicache冷缓存口径优胜被 v3 取代,仍留仓库可作"容量优先 6 条文档/512k 单条最快"备选;回滚=原脚本** | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(方案 A 口径);待切 `glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`(容量优先备选 `..._tp2pp4_hicache.env` |
| 60.6 | 无容器,但 8 卡被外部裸金属实验占用(`/data/hzy/sparse-opd-*`09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — |
| 60.7 | 基本空4 卡仍有 `/home/user/dirA_exp` 外部小任务09-08 晚实测) | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — |
| 60.8 | `glm53-nvfp4`Up:30000restart=unless-stopped | **TP4PP2-nomtp@0.90 hit90 优胜口径在役**09-09 晚接替同日早间的 TP2PP4@0.85:同 r37 栈 cu13+9 挂载+de-GLOO仅改 tp4/pp2 + memfrac 0.90KV 池 **647,040**;实测实例原样保留切 restart 策略,未重建容器。启动 `bash /root/deploy_ppmtp_r37.sh '--tp 4 --pp-size 2 --disable-overlap-schedule --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length 1048576' nomtp 8192 0.90 1 1 0 0`。hit9090% 命中 i128k/o512out cc1/2/3/4/8/16 = **28.8/47.5/63.9/76.2/106.4/125.7**vs TP2PP4 基线 cc4 +14%/cc8 +15%);并发独立 128k 文档 **4 条**cap cc4 98.5 零排队、512k 单条 ✓ TTFT 135.8s、质量门 **7/7**同轮判决DP attention 容量负收益判死、DCP 对 DSA 静默算错禁用、MTP@128k accept 2.07 判负。让步项(知情):容量 4 条TP2PP4 为 6、512k 单条慢 33%。回滚 TP2PP4 = 同命令改 `'--tp 2 --pp-size 4'` + memfrac 0.85 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`;实验全量 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/`;前史 `experiments/pro6000/glm53_nvfp4_128k_capacity_topology/`TP2PP4 口径)+ r37 `experiments/pro6000/glm53_ppmtp_r37_verify_graph/` |
| 60.8 | `glm53-nvfp4`Up:30000restart=unless-stopped | **TP4PP2-nomtp@0.90 hit90 优胜口径在役**09-09 晚接替同日早间的 TP2PP4@0.85:同 r37 栈 cu13+9 挂载+de-GLOO仅改 tp4/pp2 + memfrac 0.90KV 池 **647,040**;实测实例原样保留切 restart 策略,未重建容器。启动 `bash /root/deploy_ppmtp_r37.sh '--tp 4 --pp-size 2 --disable-overlap-schedule --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length 1048576' nomtp 8192 0.90 1 1 0 0`。hit9090% 命中 i128k/o512out cc1/2/3/4/8/16 = **28.8/47.5/63.9/76.2/106.4/125.7**vs TP2PP4 基线 cc4 +14%/cc8 +15%);并发独立 128k 文档 **4 条**cap cc4 98.5 零排队、512k 单条 ✓ TTFT 135.8s、质量门 **7/7**同轮判决DP attention 容量负收益判死、DCP 对 DSA 静默算错禁用、MTP@128k accept 2.07 判负。让步项(知情):容量 4 条TP2PP4 为 6、512k 单条慢 33%。回滚 TP2PP4 = 同命令改 `'--tp 2 --pp-size 4'` + memfrac 0.85。**09-10 B300 对标战役**:停役(容器 rename 保全 `glm53-nvfp4-insvc`)→ 双臂 B300 场景矩阵TP2PP4-D 口径 24 点 + E7b 配方 16+1 点,判决=分界 C8/E7b 窗口≤C8/边界仅 TP2PP4 可达/与 B300 绝对差 4-5×报告飞书 wiki `A7V3wZTQeifCB4krdi6cA834nW9`)→ **原容器恢复并核验**rename 回 + starthealth 200、16K 抽测 ok、显存水位与停役前一致口径未变 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`;实验全量 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/`B300 对标 `experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/`;前史 `experiments/pro6000/glm53_nvfp4_128k_capacity_topology/`TP2PP4 口径)+ r37 `experiments/pro6000/glm53_ppmtp_r37_verify_graph/` |
## 方案 A-F 一览GLM-5.3-NVFP4 @ pro60002026-09-08 双场景报告口径)

View File

@ -0,0 +1,68 @@
# GLM-5.3-NVFP4 双方案 B300 对标场景矩阵压测 — 60.86000D
日期2026-09-10 机器174.1.60.88×RTX 6000D96GB GDDR7无 NVLink
模型GLM-5.3-NVFP4modelopt 镜像:`nightly-dev-20260828-daf63171`(两臂同)
对标基线飞书《GLM 5.3 SGLang Low-latency & High-Throughput 测试结果》B300 报告wiki UPB2w4Y5yi65qwkMxJJcZko5nUc
完整报告:本目录 `REPORT.md`= 飞书发布版 A7V3wZTQeifCB4krdi6cA834nW9
## 目标
在 6000D 上复刻 B300 报告的全部场景(主场景 16K→512、4.1 短输入、4.2 长输出、5.1 长上下文、5.2 边界),对两套在役部署方案各跑一遍完整矩阵,产出对齐 B300 8 章结构的对标报告。测后 60.8 在役服务TP4PP2@0.90原容器恢复已验证health 200 + 16K 抽测 ok + 显存水位一致)。
## 实验臂
| 臂 | 方案 | 关键配置 | 质量门 |
|---|---|---|---|
| tp2pp4 | **D 生产口径**deploy_glm53_pp4.sh | TP2PP4、mem0.85、MRR48、cps16384、radix 关、KV fp8_e4m3 池 1,040,384、无投机、index_topk_freq=4=原生默认恒等、ctx 1,048,576 | 6/7仅 tool-call无 parser历史已知 |
| e7b | **TP8+EAGLE3+AR**deploy_glm53_607_exp.sh + CAR 补丁注入) | TP8、EAGLE 4/1/5、mem0.90、MRR16、cps8192、radix 开+hicache×3、KV fp8_e4m3 GPU 池 276,480、decode 图 bs1-8、ctx 270,336、custom-AR 1stage`SGLANG_CUSTOM_ALLREDUCE_ALGO=1stage`8 rank `SSKJ_CAR_PATCH_ACTIVE` 验证) | 7/7 |
## 场景矩阵与并发档位用户裁决收敛16K 封 64、64K/128K 封 8/4
| B300 章节 | 场景 | 并发档位 | TP2PP4 活跃上限 | E7b 活跃上限 |
|---|---|---|---|---|
| §3 主场景 | 16K→512 | 1/8/16/32/64 | 48MRR | 16MRR=池贴边) |
| §4.1 | 1K→128 | 1/8/32/64 | 48 | 16 |
| §4.2 | 1K→4K | 1/8/32/64 | 48 | 16c64 用户中止 rc=143 |
| §5.1 | 64K→512 | 1/4/8 | 15 | 4 |
| §5.1 | 128K→512 | 1/2/4 | 7 | 2 |
| §5.2 | 256K→1 | 1/2/3 | 3 | 1池贴边 |
| §5.2 | 512K→1 | 1 | 1 | 结构性不可ctx |
| §5.2 | 896K→1代 B300"约1M" | 1 | 1 | 结构性不可ctx |
测量协议冷缓存shared-frac 0+ 每点 flush + 服务端命中核验 ≤0.01超限重试一次nreq=max(8, 2×cc)P95 nearest-rank 对齐 B300语料耗尽21.23M/21.30M)下用回收窗口(`--pool-override` 基址映射见 REPORT 附录 A。有效测量点 41 个tp2pp4 24 + e7b 16 + e7b 256K C=1 全新文本重测),全点 0 retraction、0 OOM。
## 判决速览(详见 REPORT.md
- **分界 C=8**E7b 全场景 C=1 占优16K out 49.6 vs 16.7=3.0×、TPOT 13.7 vs 50.3ms=3.7×1K→4K out 135 vs 20.3=6.7×C≥8 TP2PP4 全指标反超并随并发拉大16K c64 out 211 vs 84=2.5×)。
- **E7b 可用窗口 ≤C8**MRR16 + decode 掉图bs>8双击16K c16 TPOT 298ms 断崖。
- **decode 密集甜点 = E7b c8**1K→4K out 412 tok/s、TPOT 22.5ms、accept 4.0。
- **TP2PP4 甜点 c16+**16K 近线性至 c64MRR48 未饱和);全场最高输出 4.2 c64 482 tok/s但 TTFT 338s仅离线
- **边界只有 TP2PP4 可达**256K/512K/896K 全测input 7,050/5,662/4,153 tok/sE7b ctx 270,336 结构性封顶。
- **DSA 复现**C=1 TPOT 对上下文不敏感50.3/50.0/49.7ms @16/64/128K并发才是驱动c8: 70.9→161.8ms)。
- **vs B300**定性结构完全复现LL/HT 分野一致),绝对差 4-5×边界 prefill 差距收窄至 ~2×分界点本机更靠前C8 vs C64-128原因是容量上限MRR/池)而非算力。
## 关键坑位(复测必读)
1. **hicache 宿主层陷阱**256K prewarm KV 占池 94.8% 触发宿主层下放,`flush_cache` 清不掉宿主层 → 同文本测量命中 0.9998。冷缓存复测**必须换该实例从未发过的文本**(本战役 E7b 256K C=1 用窗口 9,900,000 重测达标)。
2. 语料已耗尽:回收窗口复用仅在 flush+命中核验协议下有效TP2PP4 臂 radix 本来就关,零污染。
3. 在役保全流程:`docker stop``docker rename glm53-nvfp4 glm53-nvfp4-insvc`必须先改名E7b 部署脚本会 rm -f 同名容器)→ 测毕 `rename` 回 + `start`。docker stop/rm 偶发 "zombie PID" 报错是收尾边界现象,容器终态 exited(137)、显存归零,稍等重试即可。
## 资产与 md5 台账60.8 执行件 = 本目录 = 60.7 原件 三方一致)
```
1e34dd8d813ccb7ac8243e8d1573a49e bench_corpus_v2.py (scripts/)
4c126d067d33b5ea27c268f37561634c extract_summary.py (scripts/)
5892b44610b2ce61f533f1721105625e run_b300_matrix.sh (scripts/)
def3c64c5dc19e1d507080e3366c3762 deploy_glm53_pp4.sh (见 dual_scenario_bench/scripts/md5 对照一致)
21db641f1d997e26fcf1ddc97163b87e deploy_glm53_607_exp.sh (见 dual_scenario_bench/scripts/md5 对照一致)
a8fc9a504c041acc40831a848b4985d8 patches/custom_all_reduce.py
65a4d22bd87ae343e7d68c4879c14f98 patches/custom_all_reduce_utils.py
```
gen_report_tables.py 为本地表格生成器md5 未入台账60.8 侧执行件同源)。
## 原始数据
- 本目录 `results/{tp2pp4,e7b}/`all_results.jsonl逐点 SUMMARY + 命中核验、status.txt、server_facts.txt启动参数+池分配日志摘录、gpu_inventory_idle/final.csv、vram_timeline.csv30s 采样全矩阵)
- 60.8 侧:`/root/bench_logs/b300eq_tp2pp4_20260910_1138/``/root/bench_logs/b300eq_e7b_20260910_1428/`
- E7b 256K C=1all_results.jsonl 同 tag 共 4 条,最后一条为干净重测值(生成器 dict 载入后写覆盖,天然生效)

View File

@ -0,0 +1,280 @@
# GLM-5.3-NVFP4 RTX 6000D SGLang 双方案 B300 对标场景压测报告
- 测试日期2026-09-10单日单机完成两臂
- 测试机174.1.60.86000D8 卡)
- 对标基线飞书《GLM 5.3 SGLang Low-latency & High-Throughput 测试结果》B300 报告wiki UPB2w4Y5yi65qwkMxJJcZko5nUc
- 测后状态60.8 在役服务TP4PP2@0.90 口径已原容器恢复并验证health 200 + 16K 抽测 ok + 显存水位与停役前一致)
## 1. 结论摘要
本轮在单台 8 卡 RTX 6000D 上,用 GLM-5.3-**NVFP4** 完整复刻 B300 报告的场景矩阵,测试了两套在役部署方案。两套方案代表完整部署形态,不是单参数 A/B**TP2PP4** 为 D 生产口径(吞吐/长上下文形态),**TP8+EAGLE3+AR** 为 E7b 配方(低延迟形态,含 custom allreduce 1stage 补丁)。
- **低并发优先 E7bTP8+EAGLE3+AR**:主场景 `16K→512, C=1` 输出 49.6 tok/s、TPOT 13.7 ms对 TP2PP416.7 tok/s、50.3 ms分别是 **3.0×****3.7×**;长输出 `1K→4K, C=1` 输出 135 tok/s、TPOT 8.7 ms对 TP2PP420.3、49.8 ms**6.7×****5.7×**
- **高并发优先 TP2PP4**:主场景 C=64 达到 6,748 input tok/s / 211 output tok/s对 E7b2,697 / 84.3)均为 **2.5×**;短输入 C=32/64 输出 276/341 tok/s对 E7b100/103**2.7~3.3×**
- **E7b 的可用并发窗口比 B300 Low-Latency 窄一个数量级**MRR=16 且 CUDA graph 仅覆盖 decode bs 18并发 ≥8 即掉图,主场景 C=16 TPOT P95 跳到 297.8 msC=8 为 135.1 ms同点 TP2PP4 仅 98.5 ms。E7b 的生产甜点上限 = **C≤8**
- **TP2PP4 甜点在 C=16 之后**:主场景输出吞吐从 C=8 的 101 近线性爬到 C=64 的 211 tok/sMRR48 尚未饱和4.2 长输出 C=64 达全场最高 482 tok/s但 TTFT P95 338 s需要排队预算
- **decode 密集低并发的最优解是 E7b C=8**`1K→4K, C=8` 输出 412 tok/s、TPOT 22.5 ms对 TP2PP4101 tok/s、79.4 ms为 4.1×EAGLE 实测 accept length 4.0。
- **长上下文与容量边界只有 TP2PP4 可达**128K C=1 两方案输出打平11.8 vs 11.9 tok/s但 TP2PP4 TTFT 减半18.2 s vs 37.6 s256K/512K/896K 边界 E7b 结构性不可测ctx 270,336 封顶 + KV 池 276,480 贴边TP2PP4 全部完成256K C=1/2/3、512K/896K C=1
- **DSA 特性在 6000D 复现**TP2PP4 C=1 的 TPOT 对上下文长度不敏感16K/64K/128K = 50.3/50.0/49.7 ms 恒定),并发才是 TPOT 驱动因子16K 行 C=8→C=6470.9→203.7 ms
- **与 B300 的绝对差距约 4~5×**,边界 prefill 差距收窄到约 2×256K C=1 input 7,050 vs 17,457896K 4,153 vs 约1M 行 7,897。硬件与量化口径不同B300 报告未写明量化方式),绝对值仅量级可比,两份报告的结构性结论一致(见第 9 章)。
- 全部 41 个有效测量点 **0 回退retraction、0 OOM**,冷缓存命中核验全部 ≤0.01E7b 256K C=1 首测触 hicache 宿主层陷阱,用全新文本重测达标,见 5.2 注记)。
## 2. 测试环境与配置
| 项目 | TP2PP4D 生产口径) | TP8+EAGLE3+ARE7b 配方) |
|-|-|-|
| 硬件 | 单机 8 × NVIDIA RTX 6000D96 GB GDDR7nvidia-smi 可见 85,651 MiB/卡) | 同左 |
| 模型 | GLM-5.3-NVFP4modelopt 量化,/data/hf_models/GLM-5.3-NVFP4 | 同左 |
| 镜像 | `nightly-dev-20260828-daf63171` | 同左 |
| 并行 | TP2 × PP4 | TP8 |
| 投机解码 | 无 | EAGLE3num_steps=4topk=1draft_tokens=5 |
| `mem-fraction-static` | 0.85 | 0.90 |
| 最大活跃请求MRR | 48 | 16 |
| Chunk Prefill | 16,384 | 8,192 |
| KV dtype | fp8_e4m3 | fp8_e4m3 |
| KV 池(服务端实测) | **1,040,384 tokens**12.6~13.4 GB/rank无宿主层 | **276,480 tokens GPU**15.8 GB/rank+ 分层缓存 hicache×3 宿主层write_through |
| radix cache | 关(`disable_radix_cache=True` | 开(分层缓存) |
| 上下文上限 | 1,048,576config 原生) | 270,336显存约束下的部署值 |
| CUDA graph | 常规捕获 | decode 图 bs 18bs>8 掉图) |
| custom allreduce | — | 1stage 补丁注入(`SGLANG_CUSTOM_ALLREDUCE_ALGO=1stage` 环境强制8 rank `SSKJ_CAR_PATCH_ACTIVE` 日志验证全出现) |
| `index_topk_freq` | 4override等于原生默认恒等 | 原生默认 4 |
| 质量门 | 6/7仅 tool-call 失败D 口径未配 parser历史已知其余全过 | **7/7** |
因此,下文比较回答的是"两种部署形态谁更适合该负载",不能把差异单独归因于 EAGLE、PP 流水、chunk、radix 或图覆盖中的某一项(与 B300 报告同款声明)。
**测量协议**(对齐 B300 口径):
- 冷缓存:`--shared-frac 0`,每点前 `POST /flush_cache`,服务端 Prefill 日志核算命中率,>0.01 重测一次,仍超停点排查;全矩阵命中核验最终全部达标。
- 指标Input TPS / Output TPS / TTFT P95 / TPOT P95P95 为 nearest-rankInput TPS = 输入 token / 全程墙钟,与 B300 口径一致)。
- 负载PG19 真实语料 token 切片(`corpus_ids.json`input_ids 直打 `/generate`temperature=0、ignore_eos、流式nreq = max(8, 2×并发),边界行 nreq=并发。
- 并发档位按决策收敛16K/1K 类封顶 6464K 封 8、128K 封 4128K C=2 补一档);超出活跃上限的档位是**排队观察点**(与 B300 C=256 同性质,保留为有效观察)。
- 语料已耗尽21.23M/21.30M),冷缓存口径下用**回收窗口**复用(窗口基址见附录 A逐记录 `corpus_window` 字段留档)。
**两臂活跃上限**MRR 与 KV 池决定,解释各行哪些并发是排队观察点):
| 场景 | TP2PP4 活跃上限 | E7b 活跃上限 |
|-|-|-|
| 16K / 1K | 48MRR | 16MRR=池贴边) |
| 64K | 15 | 4 |
| 128K | 7 | 2 |
| 256K | 3 | 1池 262K KV / 276K 贴边) |
| 512K / 896K | 1 | 结构性不可ctx 270,336 |
## 3. 主场景16K 输入、512 输出
B300 跑了 C=1/8/32/64/128/256本机按 MRR 上限收敛为 C=1/8/16/32/64C=16 为本机甜点档B300 无此档C=128/256 超出两臂 MRR
| 并发 | 方案 | Input TPS | Output TPS | TTFT P95 | TPOT P95 |
|-|-|-|-|-|-|
| 1 | TP2PP4 | 534 | 16.7 | 5.36 s | 50.3 ms |
| 1 | TP8+EAGLE3+AR | 1,587 | 49.6 | **3.98 s** | **13.7 ms** |
| 8 | TP2PP4 | 3,236 | **101** | **14.79 s** | **70.9 ms** |
| 8 | TP8+EAGLE3+AR | 2,605 | 81.4 | 32.52 s | 135.1 ms |
| 16 | TP2PP4 | 4,674 | **146** | **25.78 s** | **98.5 ms** |
| 16 | TP8+EAGLE3+AR | 2,217 | 69.3 | 60.86 s | 297.8 ms |
| 32 | TP2PP4 | **6,185** | **193** | **46.43 s** | **152.0 ms** |
| 32 | TP8+EAGLE3+AR | 2,527 | 79.0 | 169.68 s | 259.9 ms |
| 64 | TP2PP4 | **6,748** | **211** | **134.78 s** | **203.7 ms** |
| 64 | TP8+EAGLE3+AR | 2,697 | 84.3 | 345.69 s | 231.5 ms |
趋势:
- **分界在 C=8**C=1 E7b 全指标占优C=8 起 TP2PP4 全指标反超且输出吞吐差距随并发拉大101 vs 81 → 211 vs 84
- **E7b 在 C=8→16 输出吞吐倒退**81.4→69.3 tok/sMRR=16 开始排队 + decode 掉图bs>8 无图双击TPOT P95 从 135 ms 跳到 298 ms。EAGLE accept length 随并发从 2.14 爬到 2.99,但被掉图抵消。
- **prefill 墙的差异**E7b 的 input TPS 几乎不随并发增长C=8→642,605→2,697+3%TP2PP4 翻倍3,236→6,748+108%——chunk 8192 + TP8 无 PP 流水的 prefill 瓶颈 vs chunk 16384 + PP4 流水摊满。
- TP2PP4 到 C=64 仍在爬坡C=32→64 +9%MRR48 未饱和TTFT P95 在 C=64 达 134.8 s同 B300 一样高并发 TTFT 需要准入控制。
## 4. 短输入与长输出
### 4.1 `1K -> 128`
| 并发 | 方案 | Input TPS | Output TPS | TTFT P95 | TPOT P95 |
|-|-|-|-|-|-|
| 1 | TP2PP4 | 155 | 19.3 | 386 ms | 49.3 ms |
| 1 | TP8+EAGLE3+AR | **518** | **64.7** | **326 ms** | **14.3 ms** |
| 8 | TP2PP4 | 815 | 102 | **1.69 s** | 72.9 ms |
| 8 | TP8+EAGLE3+AR | **1,219** | **152** | 2.16 s | **53.7 ms** |
| 32 | TP2PP4 | **2,204** | **276** | **4.87 s** | **98.5 ms** |
| 32 | TP8+EAGLE3+AR | 801 | 100 | 25.60 s | 190.5 ms |
| 64 | TP2PP4 | **2,724** | **341** | **19.96 s** | **101.8 ms** |
| 64 | TP8+EAGLE3+AR | 826 | 103 | 65.81 s | 177.7 ms |
短输入下 E7b 在 C≤8 显著占优C=8 输出 152 vs 1021.5×C=32 起 TP2PP4 大幅拉开2.7~3.3×。E7b 的 input TPS 反而在 C=8 最高1,219后回落——MRR16 排队开始挤占 prefill。B300 同场景 Low-Latency 到 C=128 才被反超,本机提前到 C=8~32 之间同样是容量上限MRR16/池)而非算力所致。
### 4.2 `1K -> 4K`
| 并发 | 方案 | Output TPS | TTFT P95 | TPOT P95 |
|-|-|-|-|-|
| 1 | TP2PP4 | 20.3 | 395 ms | 49.8 ms |
| 1 | TP8+EAGLE3+AR | **135** | **320 ms** | **8.7 ms** |
| 8 | TP2PP4 | 101 | 1.87 s | 79.4 ms |
| 8 | TP8+EAGLE3+AR | **412** | **1.68 s** | **22.5 ms** |
| 32 | TP2PP4 | **367** | **4.14 s** | **90.5 ms** |
| 32 | TP8+EAGLE3+AR | 252 | 304.75 s | 74.6 ms |
| 64 | TP2PP4 | **482** | 337.92 s | 95.5 ms |
| 64 | TP8+EAGLE3+AR | 用户中止*(见注) | — | — |
\* E7b C=64 点按用户指示中止("并发 64 太高"rc=143未获得有效数据同点 TP2PP4 已完成。E7b 该点 nreq=128 远超 MRR=16属排队观察点中止不影响结论完整性。
长输出放大了两形态的差异:
- **E7b C=1/C=8 是 decode 密集负载的最优区间**C=1 输出 135 tok/s、TPOT 8.7 ms全场最低C=8 输出 412 tok/s全场第二EAGLE accept 长达 3.7~4.0——长输出让草稿模型进入"顺笔"状态accept 显著高于 4.1 短输出行2.0~2.1)。
- **E7b C=32 的 TTFT P95 304.75 s** 是纯排队nreq=64 / MRR=164 波串行),其 TPOT 74.6 ms 与掉图后水平一致。
- **TP2PP4 C=64 输出 482 tok/s 为全场最高**,但 TTFT P95 338 s 意味着该点只适合离线批处理;交互负载应压在 C=32367 tok/s、TTFT 4.1 s
## 5. 长上下文观察
### 5.1 `64K/128K -> 512`
并发档位按用户指示收敛64K 封 8、128K 封 4另补 128K C=2B300 同场景为 64K C=8/32/64、128K C=8/32。
| 场景 | 并发 | 方案 | Input TPS | Output TPS | TTFT P95 | TPOT P95 |
|-|-|-|-|-|-|-|
| 64K→512 | 1 | TP2PP4 | 1,834 | 14.3 | **10.32 s** | 50.0 ms |
| 64K→512 | 1 | TP8+EAGLE3+AR | **2,877** | **22.5** | 17.17 s | **13.5 ms** |
| 64K→512 | 4 | TP2PP4 | **4,287** | **33.5** | **28.83 s** | **99.5 ms** |
| 64K→512 | 4 | TP8+EAGLE3+AR | 3,292 | 25.7 | 68.87 s | 154.6 ms |
| 64K→512 | 8 | TP2PP4 | **5,650** | **44.1** | **53.38 s** | **161.8 ms** |
| 64K→512 | 8 | TP8+EAGLE3+AR | 3,280 | 25.6 | 149.45 s | 157.5 ms |
| 128K→512 | 1 | TP2PP4 | 3,017 | 11.8 | **18.23 s** | 49.7 ms |
| 128K→512 | 1 | TP8+EAGLE3+AR | **3,054** | **11.9** | 37.60 s | **12.3 ms** |
| 128K→512 | 2 | TP2PP4 | **4,310** | **16.8** | **32.42 s** | **83.7 ms** |
| 128K→512 | 2 | TP8+EAGLE3+AR | 3,145 | 12.3 | 75.81 s | 162.2 ms |
| 128K→512 | 4 | TP2PP4 | **5,626** | **22.0** | **60.35 s** | **146.8 ms** |
| 128K→512 | 4 | TP8+EAGLE3+AR | 3,133 | 12.2 | 159.50 s | 163.4 ms |
- 64K C=1 E7b 仍占优22.5 vs 14.3 tok/s但 128K C=1 两方案输出打平11.8 vs 11.9——prefill 逐渐成为长上下文的主导成本E7b 的 decode 优势被稀释;其 TTFT 反而慢 2×37.6 vs 18.2 s
- C≥2 起 TP2PP4 全指标占优E7b 的 output TPS 在 64K/128K 行几乎不随并发变化22.5→25.6、11.9→12.2),与主场景同一形态:容量上限 + 掉图封死并发收益。
- DSA 的 TPOT 上下文不变性TP2PP4 C=150.0 ms @64K ≈ 49.7 ms @128K ≈ 50.3 ms @16K与并发驱动性C=870.9 ms @16K → 161.8 ms @64K)在本组完整呈现。
### 5.2 上下文边界
OSL=1只验证容量与 prefill不比较 Output TPS/TPOT——与 B300 同声明。B300 完成到约 1M本机以 896K=917,504 tokens 对应 B300"约 1M"档。)
| 输入长度 | 方案 | 已完成并发 | C=1 Input TPS | 最高并发 TTFT P95 |
|-|-|-|-|-|
| 256K | TP2PP4 | 1/2/3 | **7,050** | 102.99 sC=3 |
| 256K | TP8+EAGLE3+AR | 1 | 2,867 | 91.45 sC=1 |
| 512K | TP2PP4 | 1 | **5,662** | 92.60 sC=1 |
| 512K | TP8+EAGLE3+AR | 结构性不可ctx 270,336 | — | — |
| 896K | TP2PP4 | 1 | **4,153** | 220.91 sC=1 |
| 896K | TP8+EAGLE3+AR | 结构性不可ctx 270,336 | — | — |
- 256K C=1 两臂相差 2.5×7,050 vs 2,867 tok/sTP2PP4 的 chunk 16384 + PP4 流水对超长 prefill 的摊满优势,在边界长度上比 128K 行几乎打平进一步放大E7b 的 chunk 8192 代价随长度累积。
- TP2PP4 边界 input TPS 随长度衰减平缓7,050 → 5,662 → 4,153896K 单条 220.9 s 完成、池 1,040,384 tokens 单条可容KV 917,504 + 余量)。
- **E7b 256K C=1 命中核验注记hicache 宿主层陷阱)**:首测命中率 0.9998、重试仍超——根因是该点 nreq=1 测量文本与预热完全相同256K 预热 KV262,160 tokens占池 94.8% 触发分层缓存宿主层下放,`flush_cache` 只清 GPU radix 树、清不掉宿主层。改用该服务实例从未发过的文本(窗口基址 9,900,000无预热重测命中 0.0,数据干净。此为分层缓存运维要点:**宿主层缓存不受 flush_cache 影响,冷测必须换文本**。
## 6. 显存状态
- **TP2PP4**:服务加载后空载 64.6 GiB/卡,矩阵峰值 **85.0 GiB/卡**(主场景 C=64 时逼近打满,最紧张卡余量约 0.6 GiB。mem 0.85 下 KV 池按卡容量贴满分配,属预期;继续上调 MRR 或上下文没有余量,扩容前必须先降 mem-fraction。
- **TP8+EAGLE3+AR**:空载 77.9 GiB/卡EAGLE 草稿权重 + mem 0.90 大池),矩阵峰值 **83.6 GiB/卡**(余量约 2.1 GiB
- 两臂全矩阵 **0 OOM、0 retraction**(全部 41 点 retractions_total=0——B300 未披露该指标,本机在自身容量上限内运行无回退。
- 显存时间线逐 30 s 采样留档vram_timeline.csv可复核任一时刻的卡间分布。
## 7. 建议
1. **低并发交互/agent 长思考C≤8用 E7b**:主场景 C=1 TPOT 13.7 ms、长输出 C=8 输出 412 tok/s。生产并发上限建议钉在 ≤8C=16 起 decode 掉图 + MRR16 排队使其全面劣于 TP2PP4。
2. **高并发吞吐/长上下文C≥8 或输入 ≥64K用 TP2PP4**:主场景 C=64 输出 211 tok/s、128K C=1 TTFT 18.2 s、896K 可达MRR48 内未饱和,吞吐上限即 MRR。
3. **负载形态分界线**prefill 吞吐需求 >2.2K tok/s 或并发 >8 → TP2PP4decode 为主且并发 ≤8 → E7b。两臂在 C=8 附近的输出吞吐交叉(主场景 101 vs 81、短输入 102 vs 152、长输出 101 vs 412——按输出长度分布选型不能只看并发。
4. **E7b 扩窗口的两个前置**MRR 16→更高需先扩 KV 池hicache 宿主层只救命中场景不增并发容量decode 图覆盖 bs 8→16/32 才能消掉 C=16 的 298 ms TPOT 断崖。
5. **边界与超长上下文只有 TP2PP4 口径可服务**E7b 若要对标 B300 512K/约1M 行,需要 ctx ≥524,288 与池 ≥52 万 tokens 的部署形态本版ctx 270,336 / 池 276,480结构性不可达。
6. **不要把两臂差异单归因 EAGLE**两臂同时差在并行拓扑、chunk、radix、MRR 与图覆盖;单变量消融未做(与 B300 报告建议 3 同款)。
7. **生产容量同时设吞吐和延迟 SLO**TP2PP4 主场景 C=64 输出最高但 TTFT P95 已到 135 sE7b C=32 长输出 TTFT P95 305 s。只看峰值 TPS 会掩盖排队长尾。
8. **分层缓存运维**:宿主层缓存不受 `flush_cache` 影响,任何冷缓存测量/复测必须更换输入文本(见 5.2 注记)。
## 8. 原始结果与复现
- 服务器原始结果60.8
- TP2PP4 臂:`/root/bench_logs/b300eq_tp2pp4_20260910_1138/`all_results.jsonl 24 点、status.txt、server_facts.txt、gpu_inventory、vram_timeline.csv
- E7b 臂:`/root/bench_logs/b300eq_e7b_20260910_1428/`all_results.jsonl 20 条16 点 OK + 4.2 C=64 用户中止 + 256K C=1 干净重测覆盖前 3 条污染记录status.txt 含 4 个结构性跳过与 HIT_FAIL_FINAL 首测记录)
- 资产 md5 台账:`/root/bench_logs/b300eq_md5_ledger.txt`
- 本地镜像:`D:\sskj\b300eq\{tp2pp4,e7b}\`(上述全部文件)、`D:\sskj\b300eq\report_tables.md`(表格生成器输出)
- 部署脚本:`/root/deploy_glm53_pp4.sh`md5 def3c64c…与库内 sskj main 副本一致)、`/root/deploy_glm53_607_exp.sh`CAR 补丁:`/root/patches/custom_all_reduce.py`a8fc9a50…+ `custom_all_reduce_utils.py`65a4d22b…三处60.7 原件/本地/库内md5 一致
- 测量工具:`/root/bench_corpus_v2.py`md5 1e34dd8d…p95 nearest-rank + 逐请求 dump`/root/extract_summary.py`4c126d06…`/root/run_b300_matrix.sh`5892b446…矩阵驱动alive/idle_wait/prewarm/flush/命中核验/重试/VRAM 采样)
- 复现命令(单点示例):
```bash
# 冷缓存压测E7b 256K C=1 干净版)
python3 /root/bench_corpus_v2.py --input-len 262144 --output-len 1 --shared-frac 0 \
--concurrency 1 --num-requests 1 --run-id 9551 --pool-override 9900000 \
--dump-records $L/b52_256k_c1_v2_records.jsonl
# 全矩阵nohup bash /root/run_b300_matrix.sh <arm> > <progress.log> 2>&1 &
```
## 9. 与 B300 对比观察
> **口径声明**B300 报告未写明模型量化方式(若为原始 BF16 权重,则与本机 NVFP4 非同模型形态);硬件为 8×B300288 GB HBM3evs 本机 8×RTX 6000D96 GB GDDR7镜像 v0.5.18-cu130-dev4 vs nightly-20260828-daf63171。**绝对值仅量级可比,本对比只对结构性结论负责**。
- **定性结构完全复现**低延迟配方B300 Low-Latency = TP8+EAGLE vs 本机 E7b = TP8+EAGLE3+AR在 C=1 占优、吞吐配方B300 High-Throughput = DP8+DeepEP vs 本机 TP2PP4 = D 生产口径)在高并发占优——两套硬件上"低延迟 vs 高吞吐"的分野方向一致。
- **分界点本机更靠前**B300 的交叉点在 C=64~128主场景 HT C=128 反超 33%);本机在 C=8 附近。原因不是算力而是**容量上限**:本机两臂 MRR/池上限48/16远小于 B300 配方的 256/默认,先于算力撞墙。
- **绝对差距 4~5×主场景**C=1 输出 246 vs 49.6 tok/s5.0×、input 7,882 vs 1,5875.0×);吞吐侧峰值 997 vs 2114.7×、31,889 vs 6,7484.7×)。与显存带宽硬件代差量级一致。
- **边界 prefill 差距收窄到 ~2×**256K C=1 input 7,050 vs 17,4572.5×)→ 512K 5,662 vs 12,4212.2×)→ 896K/约1M 4,153 vs 7,8971.9×)。计算密集的超长 prefill 是 6000D 相对最能打的位置PP 流水摊满 + 带宽占比下降)。
- **TPOT 差距小于吞吐差距**B300 LL C=1 4.36 ms vs E7b 13.7 ms3.1×);高并发侧 B300 HT C=128 165 ms vs TP2PP4 C=64 204 ms1.2×——NVFP4 + DSA 把 decode 单步成本压得相对不差,差距主要在吞吐面。
- **饱和形态不同**B300 LL 在 C=64 后进入 24K input tok/s 平台、HT 在 C=128 达峰后 C=256 回退 19%;本机 TP2PP4 到 C=64 仍在爬坡MRR 未饱和E7b 则被 MRR16+掉图封死在 C=8。本机没有一档出现吞吐回退——"甜点=并发上限"由 MRR 决定而非算力。
- **EAGLE 配方差异**B300 LL 为 5 steps/6 draft tokens本机 E7b 为 4 steps/topk1/5 draft tokens本机实测 accept 2.0~4.0(短输出 2.0、主场景 2.1~3.0、长输出 3.7~4.0,随 decode 深入上升。B300 未披露 accept无法直接对比投机效率。
- **容量边界差距最大**B300 两模式都完成约 1M 输入 C=1/2/4本机仅 TP2PP4 可达 896K 且 C=1 单条(池 1,040,384 刚容一条E7b 连 512K 都结构性不可测ctx 270,336。96 GB 卡上"上下文边界=显存边界"比 B300 严酷得多。
## 附录 A语料窗口映射回收窗口
语料总量 21,296,780 tokens此前场景一/二战役已消费至 21,235,008。冷缓存协议下回收复用窗口基址 `--pool-override` 显式指定,每点窗口在基址上顺序推进(逐记录 `corpus_window.start/end` 留档),每点 flush + 命中核验 ≤0.01 保证冷。
| 场景 | 窗口基址 | 备注 |
|-|-|-|
| 主场景 16K→512 | 2,300,000 | |
| 4.1 短输入 1K→128 | 4,500,000 | |
| 4.2 长输出 1K→4K | 4,700,000 | |
| 5.1 64K→512 | 5,000,000 | |
| 5.1 128K→512 | 6,200,000 | |
| 5.2 256K→1 | 8,400,000 | E7b 干净重测改用 9,900,000实例首用文本避 hicache 宿主层残留) |
| 5.2 512K→1 | 9,300,000 | 仅 TP2PP4 |
| 5.2 896K→1 | 9,900,000 | 仅 TP2PP4 |
## 附录 B全量指标含 mean/p95/max、回退、投机接受长度
| 场景点 | 方案 | ok/nreq | wall(s) | out tok/s | in tok/s | TTFT mean/p95/max(s) | TPOT mean/p95/max(ms) | retractions | accept |
|---|---|---|---|---|---|---|---|---|---|
| b3_16k_c1 | TP2PP4 | 8/8 | 245.55 | 16.68 | 533.8 | 5.35/5.36/5.36 | 49.6/50.3/50.3 | 0 | None |
| b3_16k_c1 | TP8+EAGLE3+AR | 8/8 | 82.58 | 49.6 | 1587.15 | 3.93/3.98/3.98 | 12.5/13.7/13.7 | 0 | 2.138 |
| b3_16k_c8 | TP2PP4 | 16/16 | 81.0 | 101.13 | 3236.22 | 10.05/14.79/14.79 | 59.5/70.9/70.9 | 0 | None |
| b3_16k_c8 | TP8+EAGLE3+AR | 16/16 | 100.63 | 81.41 | 2605.13 | 11.75/32.52/32.52 | 71.6/135.1/135.1 | 0 | 2.375 |
| b3_16k_c16 | TP2PP4 | 32/32 | 112.16 | 146.08 | 4674.5 | 15.53/25.78/25.84 | 79.2/98.5/101.2 | 0 | None |
| b3_16k_c16 | TP8+EAGLE3+AR | 32/32 | 236.44 | 69.29 | 2217.4 | 23.60/60.86/90.39 | 173.3/297.8/305.5 | 0 | 2.517 |
| b3_16k_c32 | TP2PP4 | 64/64 | 169.54 | 193.27 | 6184.73 | 26.52/46.43/47.89 | 113.7/152.0/157.3 | 0 | None |
| b3_16k_c32 | TP8+EAGLE3+AR | 64/64 | 414.96 | 78.97 | 2526.91 | 100.26/169.68/190.73 | 169.8/259.9/319.1 | 0 | 2.767 |
| b3_16k_c64 | TP2PP4 | 128/128 | 310.79 | 210.87 | 6747.91 | 63.05/134.78/139.99 | 139.2/203.7/214.3 | 0 | None |
| b3_16k_c64 | TP8+EAGLE3+AR | 128/128 | 777.55 | 84.29 | 2697.14 | 240.10/345.69/380.96 | 168.0/231.5/311.9 | 0 | 2.986 |
| b41_1k_c1 | TP2PP4 | 8/8 | 52.95 | 19.34 | 154.7 | 0.38/0.39/0.39 | 49.1/49.3/49.3 | 0 | None |
| b41_1k_c1 | TP8+EAGLE3+AR | 8/8 | 15.82 | 64.74 | 517.93 | 0.30/0.33/0.33 | 13.2/14.3/14.3 | 0 | 2.028 |
| b41_1k_c8 | TP2PP4 | 16/16 | 20.1 | 101.91 | 815.3 | 1.31/1.69/1.69 | 68.6/72.9/72.9 | 0 | None |
| b41_1k_c8 | TP8+EAGLE3+AR | 16/16 | 13.44 | 152.4 | 1219.24 | 1.21/2.16/2.16 | 41.4/53.7/53.7 | 0 | 2.048 |
| b41_1k_c32 | TP2PP4 | 64/64 | 29.73 | 275.55 | 2204.4 | 3.89/4.87/4.88 | 85.6/98.5/106.8 | 0 | None |
| b41_1k_c32 | TP8+EAGLE3+AR | 64/64 | 81.8 | 100.14 | 801.15 | 17.56/25.60/27.89 | 146.6/190.5/195.2 | 0 | 2.054 |
| b41_1k_c64 | TP2PP4 | 128/128 | 48.11 | 340.53 | 2724.25 | 8.14/19.96/20.01 | 93.3/101.8/125.1 | 0 | None |
| b41_1k_c64 | TP8+EAGLE3+AR | 128/128 | 158.73 | 103.22 | 825.75 | 46.74/65.81/70.71 | 144.5/177.7/212.5 | 0 | 2.087 |
| b42_1k4k_c1 | TP2PP4 | 8/8 | 1614.16 | 20.3 | 5.08 | 0.39/0.40/0.40 | 49.2/49.8/49.8 | 0 | None |
| b42_1k4k_c1 | TP8+EAGLE3+AR | 8/8 | 242.18 | 135.31 | 33.83 | 0.31/0.32/0.32 | 7.3/8.7/8.7 | 0 | 3.701 |
| b42_1k4k_c8 | TP2PP4 | 16/16 | 649.69 | 100.87 | 25.22 | 1.35/1.87/1.87 | 79.0/79.4/79.4 | 0 | None |
| b42_1k4k_c8 | TP8+EAGLE3+AR | 16/16 | 159.24 | 411.54 | 102.89 | 0.96/1.68/1.68 | 18.1/22.5/22.5 | 0 | 4.003 |
| b42_1k4k_c32 | TP2PP4 | 64/64 | 714.14 | 367.08 | 91.77 | 1.98/4.14/4.15 | 86.1/90.5/91.1 | 0 | None |
| b42_1k4k_c32 | TP8+EAGLE3+AR | 64/64 | 1040.79 | 251.87 | 62.97 | 195.07/304.75/334.69 | 60.7/74.6/92.3 | 0 | 3.852 |
| b42_1k4k_c64 | TP2PP4 | 128/128 | 1088.63 | 481.6 | 120.4 | 80.74/337.92/338.47 | 91.2/95.5/96.3 | 0 | None |
| b51_64k_c1 | TP2PP4 | 8/8 | 285.91 | 14.33 | 1833.72 | 10.28/10.32/10.32 | 49.8/50.0/50.0 | 0 | None |
| b51_64k_c1 | TP8+EAGLE3+AR | 8/8 | 182.23 | 22.48 | 2877.06 | 17.04/17.17/17.17 | 11.2/13.5/13.5 | 0 | 2.468 |
| b51_64k_c4 | TP2PP4 | 8/8 | 122.3 | 33.49 | 4286.94 | 19.55/28.83/28.83 | 81.4/99.5/99.5 | 0 | None |
| b51_64k_c4 | TP8+EAGLE3+AR | 8/8 | 159.24 | 25.72 | 3292.45 | 30.55/68.87/68.87 | 94.9/154.6/154.6 | 0 | 2.438 |
| b51_64k_c8 | TP2PP4 | 16/16 | 185.59 | 44.14 | 5650.11 | 31.83/53.38/53.38 | 119.2/161.8/161.8 | 0 | None |
| b51_64k_c8 | TP8+EAGLE3+AR | 16/16 | 319.66 | 25.63 | 3280.31 | 90.92/149.45/149.45 | 105.8/157.5/157.5 | 0 | 2.442 |
| b51_128k_c1 | TP2PP4 | 8/8 | 347.5 | 11.79 | 3017.45 | 18.21/18.23/18.23 | 49.4/49.7/49.7 | 0 | None |
| b51_128k_c1 | TP8+EAGLE3+AR | 8/8 | 343.36 | 11.93 | 3053.83 | 37.59/37.60/37.60 | 10.4/12.3/12.3 | 0 | 2.728 |
| b51_128k_c2 | TP2PP4 | 8/8 | 243.3 | 16.84 | 4309.84 | 25.27/32.42/32.42 | 69.6/83.7/83.7 | 0 | None |
| b51_128k_c2 | TP8+EAGLE3+AR | 8/8 | 333.36 | 12.29 | 3145.44 | 48.21/75.81/75.81 | 68.0/162.2/162.2 | 0 | 2.769 |
| b51_128k_c4 | TP2PP4 | 8/8 | 186.37 | 21.98 | 5626.35 | 39.29/60.35/60.35 | 105.4/146.8/146.8 | 0 | None |
| b51_128k_c4 | TP8+EAGLE3+AR | 8/8 | 334.71 | 12.24 | 3132.76 | 110.14/159.50/159.50 | 68.7/163.4/163.4 | 0 | 2.595 |
| b52_256k_c1 | TP2PP4 | 1/1 | 37.18 | 0.03 | 7050.23 | 37.18/37.18/37.18 | — | 0 | None |
| b52_256k_c1 | TP8+EAGLE3+AR | 1/1 | 91.45 | 0.01 | 2866.59 | 91.45/91.45/91.45 | — | 0 | None |
| b52_256k_c2 | TP2PP4 | 2/2 | 70.19 | 0.03 | 7469.22 | 53.69/70.12/70.12 | — | 0 | None |
| b52_256k_c3 | TP2PP4 | 3/3 | 103.12 | 0.03 | 7626.22 | 70.12/102.99/102.99 | — | 0 | None |
| b52_512k_c1 | TP2PP4 | 1/1 | 92.6 | 0.01 | 5662.03 | 92.60/92.60/92.60 | — | 0 | None |
| b52_896k_c1 | TP2PP4 | 1/1 | 220.91 | 0.0 | 4153.24 | 220.91/220.91/220.91 | — | 0 | None |
E7b 4.2 C=64 用户中止rc=143无 SUMMARY不在表内E7b 256K/512K/896K C>1 为结构性跳过E7b 256K C=1 为全新文本干净重测值(窗口 9,900,000

View File

@ -0,0 +1,424 @@
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
# Adapted from https://github.com/vllm-project/vllm/blob/v0.6.4.post1/vllm/distributed/device_communicators/custom_all_reduce.py
import ctypes
import logging
import os
from contextlib import contextmanager
from functools import partial
from typing import Any, List, Optional, Union
import torch
import torch.distributed as dist
from torch.distributed import ProcessGroup
import sglang.srt.distributed.device_communicators.custom_all_reduce_ops as ops
from sglang.srt.distributed.device_communicators.cuda_wrapper import CudaRTLibrary
from sglang.srt.distributed.device_communicators.custom_all_reduce_utils import (
can_use_custom_all_reduce_with_nvlink,
is_weak_contiguous,
)
from sglang.srt.environ import envs
from sglang.srt.model_executor.runner_backend_utils.tc_piecewise_cuda_graph import (
is_in_tc_piecewise_cuda_graph,
)
from sglang.srt.utils import (
get_bool_env_var,
is_cuda,
is_hip,
is_musa,
log_info_on_rank0,
)
_is_cuda = is_cuda()
_is_hip = is_hip()
_is_musa = is_musa()
logger = logging.getLogger(__name__)
os.environ.setdefault("SGLANG_CUSTOM_ALLREDUCE_ALGO", "1stage") # SSKJ-PATCH: C++ dispatch no-ops when full_nvlink=False; force kernel launch
class CustomAllreduce:
_SUPPORTED_WORLD_SIZES = [2, 4, 6, 8]
_MAX_CAR_SIZE = 8192 * 1024
if _is_hip:
# crossover is at 16MB buffer size for ROCm
_MAX_CAR_SIZE = 2 * 8192 * 1024
if _is_musa:
# crossover is at 128MB buffer size for MUSA
_MAX_CAR_SIZE = 16 * 8196 * 1024
# max_size: max supported allreduce size
def __init__(
self,
group: ProcessGroup,
device: Union[int, str, torch.device],
max_size=_MAX_CAR_SIZE,
) -> None:
"""
Args:
group: the process group to work on. If None, it will use the
default process group.
device: the device to bind the CustomAllreduce to. If None,
it will be bind to f"cuda:{local_rank}".
It is the caller's responsibility to make sure each communicator
is bind to a unique device, and all communicators in this group
are in the same node.
"""
self._IS_CAPTURING = False
self.disabled = True # This can be modified in-place by context manager in piecewise cuda graph runner
self.original_disabled = True # To store the original state
self.use_amd_deterministic_impl = _use_amd_deterministic_impl()
if not ops.IS_CUSTOM_AR_AVAILABLE:
# disable because of missing custom allreduce library
# e.g. in a non-cuda environment
return
rank = dist.get_rank(group=group)
world_size = dist.get_world_size(group=group)
if isinstance(device, int):
device = torch.device(f"cuda:{device}")
elif isinstance(device, str):
device = torch.device(device)
# now `device` is a `torch.device` object
assert isinstance(device, torch.device)
self.device = device
full_nvlink = can_use_custom_all_reduce_with_nvlink(
group=group,
device=device,
supported_world_size=self._SUPPORTED_WORLD_SIZES,
cls_name="CustomAllreduce",
)
if full_nvlink is None:
return # fail to get nvlink status
self.group = group
self.max_size = max_size
self.rank = rank
self.world_size = world_size
self.full_nvlink = full_nvlink
if not _is_hip:
# Buffers memory are owned by this Python class and passed to C++.
# Meta data composes of two parts: meta data for synchronization and a
# temporary buffer for storing intermediate allreduce results.
self.meta_ptrs = self.create_shared_buffer(
ops.meta_size() + max_size, group=group
)
# This is a pre-registered IPC buffer. In eager mode, input tensors
# are first copied into this buffer before allreduce is performed
self.buffer_ptrs = self.create_shared_buffer(max_size, group=group)
# This is a buffer for storing the tuples of pointers pointing to
# IPC buffers from all ranks. Each registered tuple has size of
# 8*world_size bytes where world_size is at most 8. Allocating 8MB
# is enough for 131072 such tuples. The largest model I've seen only
# needs less than 10000 of registered tuples.
self.rank_data = torch.empty(
max_size, dtype=torch.uint8, device=self.device
)
self._ptr = ops.init_custom_ar(
self.meta_ptrs, self.rank_data, rank, self.full_nvlink
)
ops.register_buffer(self._ptr, self.buffer_ptrs)
else:
# meta data buffers need to be "uncached" for signal on MI200
self.meta = ops.allocate_meta_buffer(ops.meta_size() + max_size)
self.buffer = torch.empty(max_size, dtype=torch.uint8, device=self.device)
handle = ops.get_meta_buffer_ipc_handle(self.meta)
shard_data = (
bytes(handle), # ipc handle to base ptr
0, # offset of base ptr
)
handles, offsets = self._gather_ipc_meta(shard_data)
self.rank_data = torch.empty(
max_size, dtype=torch.uint8, device=self.device
)
self._ptr = ops.init_custom_ar(
self.meta, self.rank_data, handles, offsets, rank, self.full_nvlink
)
self.register_buffer(self.buffer)
self.disabled = False
self.original_disabled = False # Ensure original_disabled == disabled
logger.warning(f"SSKJ_CAR_PATCH_ACTIVE ws={self.world_size} full_nvlink={self.full_nvlink}")
self.tms_cudagraph = envs.SGLANG_MEMORY_SAVER_CUDA_GRAPH.get()
@staticmethod
def create_shared_buffer(
size_in_bytes: int, group: Optional[ProcessGroup] = None
) -> List[int]:
"""
Creates a shared buffer and returns a list of pointers
representing the buffer on all processes in the group.
"""
lib = CudaRTLibrary()
pointer = lib.cudaMalloc(size_in_bytes)
if _is_musa:
lib.cudaMemset(pointer, 0, size_in_bytes)
handle = lib.cudaIpcGetMemHandle(pointer)
world_size = dist.get_world_size(group=group)
rank = dist.get_rank(group=group)
handles = [None] * world_size
dist.all_gather_object(handles, handle, group=group)
pointers: List[int] = []
for i, h in enumerate(handles):
if i == rank:
pointers.append(pointer.value) # type: ignore
else:
pointers.append(lib.cudaIpcOpenMemHandle(h).value) # type: ignore
return pointers
@staticmethod
def free_shared_buffer(
pointers: List[int], group: Optional[ProcessGroup] = None
) -> None:
rank = dist.get_rank(group=group)
lib = CudaRTLibrary()
lib.cudaFree(ctypes.c_void_p(pointers[rank]))
@contextmanager
def capture(self):
"""
The main responsibility of this context manager is the
`register_graph_buffers` call at the end of the context.
It records all the buffer addresses used in the CUDA graph.
"""
try:
self._IS_CAPTURING = True
yield
finally:
self._IS_CAPTURING = False
if not self.disabled:
self.register_graph_buffers()
def _get_ipc_meta(self, inp: torch.Tensor):
# _share_cuda_() doesn't accept meta buffer not allocated from
# PyTorch cache allocator, use direct HIP call to get IPC handle
handle = ops.get_meta_buffer_ipc_handle(inp)
shard_data = (
bytes(handle), # ipc handle to base ptr
0, # offset of base ptr
)
return self._gather_ipc_meta(shard_data)
def _gather_ipc_meta(self, shard_data):
# Note: don't use `[[None]] * self.world_size` here
# because it will create a list of the same reference
all_data: List[Optional[Any]] = [[None] for i in range(self.world_size)]
all_data[self.rank][0] = shard_data
ranks = dist.get_process_group_ranks(group=self.group)
ranks.sort()
for i, rank in enumerate(ranks):
dist.broadcast_object_list(
all_data[i], src=rank, group=self.group, device="cpu"
)
# we cannot directly use `dist.all_gather_object` here
# because it is incompatible with `gloo` backend under inference mode.
# see https://github.com/pytorch/pytorch/issues/126032 for details.
handles = []
offsets = []
for i in range(len(all_data)):
handles.append(all_data[i][0][0]) # type: ignore
offsets.append(all_data[i][0][1]) # type: ignore
return handles, offsets
def register_buffer(self, inp: torch.Tensor):
handles, offsets = self._get_ipc_meta(inp)
ops.register_buffer(self._ptr, inp, handles, offsets)
def register_graph_buffers(self):
if _is_hip:
handle, offset = ops.get_graph_buffer_ipc_meta(self._ptr)
handles, offsets = self._gather_ipc_meta((bytes(handle), offset))
log_info_on_rank0(logger, f"Registering {len(offset)} cuda graph addresses")
ops.register_graph_buffers(self._ptr, handles, offsets)
else:
handle, offset = ops.get_graph_buffer_ipc_meta(self._ptr)
log_info_on_rank0(logger, f"Registering {len(offset)} cuda graph addresses")
# We cannot directly use `dist.all_gather_object` here
# because it is incompatible with `gloo` backend under inference mode.
# see https://github.com/pytorch/pytorch/issues/126032 for details.
all_data = [
[None, None] for _ in range(dist.get_world_size(group=self.group))
]
all_data[self.rank] = [handle, offset]
ranks = sorted(dist.get_process_group_ranks(group=self.group))
for i, rank in enumerate(ranks):
dist.broadcast_object_list(
all_data[i], src=rank, group=self.group, device="cpu"
)
# Unpack list of tuples to tuple of lists.
handles = [d[0] for d in all_data] # type: ignore
offsets = [d[1] for d in all_data] # type: ignore
ops.register_graph_buffers(self._ptr, handles, offsets)
def should_custom_ar(self, inp: torch.Tensor):
if self.disabled:
return False
inp_size = inp.numel() * inp.element_size()
# custom allreduce requires input byte size to be multiples of 16
if inp_size % 16 != 0:
return False
if not is_weak_contiguous(inp):
return False
# for 4 or more non NVLink-capable GPUs, custom allreduce provides
# little performance improvement over NCCL.
if not _is_hip:
if True:
return inp_size <= self.max_size
return False
if _is_hip:
if self.use_amd_deterministic_impl:
return True
if self.full_nvlink:
return inp_size <= self.max_size
return False
return False
def _all_reduce_impl(self, inp: torch.Tensor, registered: bool):
out = torch.empty_like(inp)
if not _is_hip: # CUDA-like
if registered:
ops.all_reduce(self._ptr, inp, out, 0, 0)
else:
ops.all_reduce(
self._ptr, inp, out, self.buffer_ptrs[self.rank], self.max_size
)
elif self.use_amd_deterministic_impl:
inp_size = inp.numel() * inp.element_size()
if inp_size < self.max_size:
reg_buffer = self.buffer.view(inp.dtype)[: inp.numel()]
ops.deterministic_all_reduce_unreg(self._ptr, inp, reg_buffer, out)
else:
self.register_buffer(inp)
ops.deterministic_all_reduce_reg(self._ptr, inp, out)
else: # normal AMD ROCm path
if registered:
ops.all_reduce_reg(self._ptr, inp, out)
else:
ops.all_reduce_unreg(self._ptr, inp, self.buffer, out)
return out
def custom_all_reduce(self, input: torch.Tensor) -> Optional[torch.Tensor]:
"""The main allreduce API that provides support for cuda graph."""
# When custom allreduce is disabled, this will be None.
if self.disabled or not self.should_custom_ar(input):
return None
if self._IS_CAPTURING:
if torch.cuda.is_current_stream_capturing():
return self._all_reduce_impl(input, registered=not self.tms_cudagraph)
else:
# Could be warmup OR piecewise cuda graph split op execution.
# In piecewise cuda graph, split ops run eagerly outside the graph
# but _IS_CAPTURING is still True. We need to do real all-reduce.
if is_in_tc_piecewise_cuda_graph():
# Split op execution - do real all-reduce
return self._all_reduce_impl(input, registered=False)
else:
# True warmup - mimic the allocation pattern since custom
# allreduce is out-of-place.
return torch.zeros_like(input)
else:
return self._all_reduce_impl(input, registered=False)
def close(self):
if not self.disabled and self._ptr:
if ops is not None:
ops.dispose(self._ptr)
if _is_cuda:
self.free_shared_buffer(self.meta_ptrs)
self.free_shared_buffer(self.buffer_ptrs)
self._ptr = 0
def __del__(self):
self.close()
def dispatch_custom_allreduce(
group: ProcessGroup,
device: torch.device,
):
"""Return the CustomAllreduce class to use (aiter on ROCm if enabled).
On AMD with 1-stage AR enabled, use sglang's CustomAllreduce.
Otherwise use AiterCustomAllreduce if available.
On CUDA, the JIT-compiled v2 implementation is used by default.
Set SGLANG_OPT_USE_CUSTOM_ALL_REDUCE_V2=0 to fall back to the legacy CustomAllreduce.
Multi-node v2 is admitted only for a single NVLink clique (see
can_use_custom_all_reduce_v2); other cross-node groups fall back to NCCL.
"""
if _is_cuda and envs.SGLANG_OPT_USE_CUSTOM_ALL_REDUCE_V2.get():
from .custom_all_reduce_v2 import (
CustomAllReduceV2,
can_use_custom_all_reduce_v2,
)
if can_use_custom_all_reduce_v2(group=group, device=device):
logger.debug("[AR] Using CustomAllReduceV2 (JIT-compiled)")
return CustomAllReduceV2
if _is_cuda or _is_musa:
return CustomAllreduce
assert _is_hip
if envs.SGLANG_USE_1STAGE_ALLREDUCE.is_set():
if envs.SGLANG_USE_1STAGE_ALLREDUCE.get():
logger.debug(
"[AR] All-reduce: 1-stage kernel (SGLANG_USE_1STAGE_ALLREDUCE=1)"
)
else:
logger.debug("[AR] All-reduce: default (SGLANG_USE_1STAGE_ALLREDUCE=0)")
elif envs.SGLANG_ENABLE_DETERMINISTIC_INFERENCE.get():
logger.debug(
"[AR] All-reduce: 1-stage kernel (deterministic inference enabled)"
)
else:
logger.debug("[AR] All-reduce: default")
# On AMD with 1-stage AR, use sglang's CustomAllreduce
# (AiterCustomAllreduce doesn't have deterministic_all_reduce method)
if _use_amd_deterministic_impl():
return CustomAllreduce
if get_bool_env_var("SGLANG_USE_AITER_AR", default="true"):
try:
from aiter.dist.device_communicators.custom_all_reduce import (
CustomAllreduce as AiterCustomAllreduce,
)
logger.info("[AR] Using AiterCustomAllreduce (AMD default)")
tms_cudagraph = envs.SGLANG_MEMORY_SAVER_CUDA_GRAPH.get()
return partial(
AiterCustomAllreduce,
enable_register_for_capturing=not tms_cudagraph,
)
except ImportError as e:
logger.warning(
"[AR] Aiter custom all-reduce not available; "
"falling back to sglang CustomAllreduce. Details: %s",
e,
)
return CustomAllreduce
return CustomAllreduce
def _use_amd_deterministic_impl() -> bool:
if not _is_hip: # CUDA is always deterministic
return False
if envs.SGLANG_USE_1STAGE_ALLREDUCE.is_set():
return envs.SGLANG_USE_1STAGE_ALLREDUCE.get()
else:
return envs.SGLANG_ENABLE_DETERMINISTIC_INFERENCE.get()

View File

@ -0,0 +1,519 @@
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
# Adapted from https://github.com/vllm-project/vllm/blob/v0.6.4.post1/vllm/distributed/device_communicators/custom_all_reduce_utils.py
import ctypes
import json
import logging
import os
import pickle
import subprocess
import sys
import tempfile
from functools import wraps
from itertools import product
from typing import Callable, Dict, List, Optional, Sequence, TypeVar
import torch
import torch.distributed as dist
import torch.multiprocessing as mp
from typing_extensions import ParamSpec
from sglang.srt.distributed.device_communicators.cuda_wrapper import CudaRTLibrary
from sglang.srt.distributed.parallel_state import in_the_same_node_as
from sglang.srt.environ import envs as sglang_envs
from sglang.srt.utils import is_cuda, is_hip, is_musa
from sglang.srt.utils.cuda_vmm_utils import _gpu_fabric_clique
logger = logging.getLogger(__name__)
_is_cuda = is_cuda()
_is_hip = is_hip()
_is_musa = is_musa()
if _is_cuda:
try:
import pynvml
except ImportError as e:
logger.warning("Failed to import pynvml with %r", e)
if _is_musa:
try:
import pymtml as pynvml
except ImportError as e:
logger.warning("Failed to import pymtml with %r", e)
if _is_hip:
try:
from amdsmi import (
AmdSmiException,
amdsmi_get_processor_handles,
amdsmi_init,
amdsmi_shut_down,
amdsmi_topo_get_link_type,
)
except ImportError as e:
logger.warning("Failed to import amdsmi with %r", e)
_P = ParamSpec("_P")
_R = TypeVar("_R")
def update_environment_variables(envs: Dict[str, str]):
for k, v in envs.items():
if k in os.environ and os.environ[k] != v:
logger.warning(
"Overwriting environment variable %s " "from '%s' to '%s'",
k,
os.environ[k],
v,
)
os.environ[k] = v
def producer(
batch_src: Sequence[int],
producer_queue,
consumer_queue,
result_queue,
cuda_visible_devices: Optional[str] = None,
):
if cuda_visible_devices is not None:
update_environment_variables({"CUDA_VISIBLE_DEVICES": cuda_visible_devices})
lib = CudaRTLibrary()
for i in batch_src:
lib.cudaSetDevice(i)
pointer = lib.cudaMalloc(1024)
lib.cudaMemset(pointer, 1, 1024)
lib.cudaDeviceSynchronize()
handle = lib.cudaIpcGetMemHandle(pointer)
producer_queue.put(handle)
open_success = consumer_queue.get()
if open_success:
# use two queues to simulate barrier
producer_queue.put(0)
consumer_queue.get()
# check if the memory is modified
host_data = (ctypes.c_char * 1024)()
lib.cudaMemcpy(host_data, pointer, 1024) # type: ignore
for i in range(1024):
if ord(host_data[i]) != 2:
open_success = False
break
result_queue.put(open_success)
lib.cudaDeviceReset()
def consumer(
batch_tgt: Sequence[int],
producer_queue,
consumer_queue,
result_queue,
cuda_visible_devices: Optional[str] = None,
):
if cuda_visible_devices is not None:
update_environment_variables({"CUDA_VISIBLE_DEVICES": cuda_visible_devices})
lib = CudaRTLibrary()
for j in batch_tgt:
lib.cudaSetDevice(j)
handle = producer_queue.get()
open_success = False
try:
pointer = lib.cudaIpcOpenMemHandle(handle) # type: ignore
open_success = True
except RuntimeError:
# cannot error out here, because the producer process
# is still waiting for the response.
pass
consumer_queue.put(open_success)
if open_success:
# modify the memory
lib.cudaMemset(pointer, 2, 1024)
lib.cudaDeviceSynchronize()
# use two queues to simulate barrier
producer_queue.get()
consumer_queue.put(0)
# check if the memory is modified
host_data = (ctypes.c_char * 1024)()
lib.cudaMemcpy(host_data, pointer, 1024) # type: ignore
for i in range(1024):
if ord(host_data[i]) != 2:
open_success = False
break
result_queue.put(open_success)
lib.cudaDeviceReset()
def can_actually_p2p(
batch_src: Sequence[int],
batch_tgt: Sequence[int],
) -> Sequence[bool]:
"""
Usually, checking if P2P access is enabled can be done by
`torch.cuda.can_device_access_peer(src, tgt)`. However, sometimes
the driver might be broken, and `torch.cuda.can_device_access_peer(src, tgt)`
returns `True` even if P2P access is not actually possible.
See https://github.com/vllm-project/vllm/issues/2728 and
https://forums.developer.nvidia.com/t/direct-gpu-gpu-communication-does-not-seem-to-work-properly/283264/10
Therefore, we have to perform a real P2P access to check if it is actually
possible.
Note on p2p and cuda IPC:
Usually, one process uses one GPU:
GPU src --> cuda context src --> tensor src --> process src
We need to combine p2p and cuda IPC, so that:
GPU src --> cuda context src --> tensor src --> process src
|shared|
GPU tgt --> cuda context tgt --> tensor tgt --> process tgt
That is to say, process src creates a tensor in GPU src, passes IPC handle to
process tgt, and process tgt accesses the tensor in GPU tgt. Any operation on the
tensor in process tgt will be reflected in the tensor in process src, because
they are the same memory segment.
It is important to note that process tgt accesses the tensor in GPU tgt, not
GPU src. That's why we need p2p access.
The most time-consuming part is the process creation. To avoid creating
processes for every pair of GPUs, we use batched testing. We create two
processes for testing all pairs of GPUs in batch. The trick is to reset
the device after each test (which is not available in PyTorch).
""" # noqa
cuda_visible_devices = os.environ.get("CUDA_VISIBLE_DEVICES", None)
# pass the CUDA_VISIBLE_DEVICES to the child process
# to make sure they see the same set of GPUs
# make sure the processes are spawned
smp = mp.get_context("spawn")
producer_queue = smp.Queue()
consumer_queue = smp.Queue()
result_queue = smp.Queue()
p_src = smp.Process(
target=producer,
args=(
batch_src,
producer_queue,
consumer_queue,
result_queue,
cuda_visible_devices,
),
)
p_tgt = smp.Process(
target=consumer,
args=(
batch_tgt,
producer_queue,
consumer_queue,
result_queue,
cuda_visible_devices,
),
)
p_src.start()
p_tgt.start()
p_src.join()
p_tgt.join()
assert p_src.exitcode == 0 and p_tgt.exitcode == 0
result: List[bool] = []
for src, tgt in zip(batch_src, batch_tgt):
a = result_queue.get()
b = result_queue.get()
if a != b:
logger.warning(
"Two processes do not agree on the P2P access"
" status on %d -> %d, treat as disabled.",
src,
tgt,
)
result.append(False)
else:
result.append(a)
return result
# why do we need this cache?
# we are testing peer-to-peer (p2p) access between GPUs,across processes.
# if we test it every time, it will be very slow, because we need to create
# N * N * 2 processes, where N is the world size. This is very slow.
# to reduce the time, we use a cache file to store the p2p access status.
# the cache file is generated by the master process if it does not exist.
# then all the processes can read the cache file to check the p2p access status.
# Note that the cache file is suffixed by the CUDA_VISIBLE_DEVICES, so that we
# can have different cache files for different CUDA_VISIBLE_DEVICES settings,
# e.g. used by different vllm engines. The device id in the cache file is a
# **local** device id, i.e. from 0 to num_dev-1, where num_dev is the number
# of visible devices in the vllm engine.
_gpu_p2p_access_cache: Optional[Dict[str, bool]] = None
def gpu_p2p_access_check(src: int, tgt: int) -> bool:
"""Check if GPU src can access GPU tgt."""
# if the cache variable is already calculated,
# read from the cache instead of checking it again
global _gpu_p2p_access_cache
if _gpu_p2p_access_cache is not None:
return _gpu_p2p_access_cache[f"{src}->{tgt}"]
is_distributed = dist.is_initialized()
num_dev = torch.cuda.device_count()
cuda_visible_devices = os.environ.get("CUDA_VISIBLE_DEVICES", None)
if cuda_visible_devices is None:
cuda_visible_devices = ",".join(str(i) for i in range(num_dev))
# VLLM_CACHE_ROOT -> SGLANG_CACHE_ROOT
# "~/.cache/vllm" -> envs.SGLANG_CACHE_DIR
SGLANG_CACHE_ROOT = os.path.expanduser(sglang_envs.SGLANG_CACHE_DIR.get())
path = os.path.join(
SGLANG_CACHE_ROOT, f"gpu_p2p_access_cache_for_{cuda_visible_devices}.json"
)
cache_dir = os.path.dirname(path)
try:
os.makedirs(cache_dir, exist_ok=True)
except (FileExistsError, NotADirectoryError):
if not os.path.isdir(cache_dir):
# Path exists as a file (stale cache/lock). Remove and retry.
try:
os.remove(cache_dir)
except OSError:
pass
os.makedirs(cache_dir, exist_ok=True)
from sglang.srt.distributed.parallel_state import get_world_group
if (not is_distributed or get_world_group().local_rank == 0) and (
not os.path.exists(path)
):
# only the local master process (with local_rank == 0) can
# enter this block to calculate the cache
logger.info("generating GPU P2P access cache in %s", path)
cache: Dict[str, bool] = {}
ids = list(range(num_dev))
# batch of all pairs of GPUs
batch_src, batch_tgt = zip(*list(product(ids, ids)))
# NOTE: we use `subprocess` rather than `multiprocessing` here
# because the caller might not have `if __name__ == "__main__":`,
# in that case we cannot use spawn method in multiprocessing.
# However, `can_actually_p2p` requires spawn method.
# The fix is, we use `subprocess` to call the function,
# where we have `if __name__ == "__main__":` in this file.
# use a temporary file to store the result
# we don't use the output of the subprocess directly,
# because the subprocess might produce logging output
with tempfile.NamedTemporaryFile() as output_file:
input_bytes = pickle.dumps((batch_src, batch_tgt, output_file.name))
returned = subprocess.run(
[sys.executable, __file__], input=input_bytes, capture_output=True
)
# check if the subprocess is successful
try:
returned.check_returncode()
except Exception as e:
# wrap raised exception to provide more information
raise RuntimeError(
f"Error happened when batch testing "
f"peer-to-peer access from {batch_src} to {batch_tgt}:\n"
f"{returned.stderr.decode()}"
) from e
with open(output_file.name, "rb") as f:
result = pickle.load(f)
for _i, _j, r in zip(batch_src, batch_tgt, result):
cache[f"{_i}->{_j}"] = r
with open(path, "w") as f:
json.dump(cache, f, indent=4)
if is_distributed:
get_world_group().barrier()
logger.info("reading GPU P2P access cache from %s", path)
with open(path) as f:
cache = json.load(f)
_gpu_p2p_access_cache = cache
return _gpu_p2p_access_cache[f"{src}->{tgt}"]
def with_nvml_context(fn: Callable[_P, _R]) -> Callable[_P, _R]:
@wraps(fn)
def wrapper(*args: _P.args, **kwargs: _P.kwargs) -> _R:
if _is_hip:
try:
amdsmi_init()
return fn(*args, **kwargs)
finally:
amdsmi_shut_down()
else:
pynvml.nvmlInit()
try:
return fn(*args, **kwargs)
finally:
pynvml.nvmlShutdown()
return wrapper
@with_nvml_context
def is_full_nvlink(physical_device_ids: List[int], world_size: int) -> bool:
if _is_hip:
"""
query if the set of gpus are fully connected by xgmi (1 hop)
"""
handles = [amdsmi_get_processor_handles()[i] for i in physical_device_ids]
for i, handle in enumerate(handles):
for j, peer_handle in enumerate(handles):
if i < j:
try:
link_type = amdsmi_topo_get_link_type(handle, peer_handle)
# type is 2 for XGMI
if link_type["hops"] != 1 or link_type["type"] != 2:
return False
except AmdSmiException as error:
logger.error("AMD 1 hop XGMI detection failed.", exc_info=error)
return False
return True
else:
"""
query if the set of gpus are fully connected by nvlink (1 hop)
"""
handles = [pynvml.nvmlDeviceGetHandleByIndex(i) for i in physical_device_ids]
for i, handle in enumerate(handles):
for j, peer_handle in enumerate(handles):
if i < j:
try:
p2p_status = pynvml.nvmlDeviceGetP2PStatus(
handle, peer_handle, pynvml.NVML_P2P_CAPS_INDEX_NVLINK
)
if p2p_status != pynvml.NVML_P2P_STATUS_OK:
return False
except pynvml.NVMLError:
logger.exception(
"NVLink detection failed. This is normal if your"
" machine has no NVLink equipped."
)
return False
return True
@with_nvml_context
def is_one_nvlink_clique(
group: torch.distributed.ProcessGroup, device: torch.device
) -> bool:
"""True iff every rank's GPU is in the same NVLink fabric clique (one NVL72 /
MNNVL domain). Such a clique shares a single NVLink address space even across
nodes, so custom-AR v2's symm-mem storage + fabric peer VAs are valid group-wide."""
if _is_hip:
return False
try:
clique = _gpu_fabric_clique(device)
except Exception as e:
logger.warning(
"GPU fabric clique query failed (%r); custom-AR stays intra-node.", e
)
clique = None
# Always all-gather (every rank calls it once) so a failed query on any rank
# resolves to a clean False rather than a collective mismatch.
world_size = dist.get_world_size(group=group)
gathered: List[object] = [None] * world_size
dist.all_gather_object(gathered, clique, group=group)
if any(c is None for c in gathered):
return False
return len(set(gathered)) == 1
def is_weak_contiguous(inp: torch.Tensor):
return inp.is_contiguous() or (
inp.storage().nbytes() - inp.storage_offset() * inp.element_size()
== inp.numel() * inp.element_size()
)
def can_p2p(rank: int, world_size: int) -> bool:
# SGLANG_SKIP_P2P_CHECK can be set to False in sglang
SGLANG_SKIP_P2P_CHECK = os.getenv("SGLANG_SKIP_P2P_CHECK", "0") == "1"
for i in range(world_size):
if i == rank:
continue
if SGLANG_SKIP_P2P_CHECK:
logger.info("Skipping P2P check and trusting the driver's P2P report.")
return torch.cuda.can_device_access_peer(rank, i)
if not gpu_p2p_access_check(rank, i):
return False
return True
def can_use_custom_all_reduce_with_nvlink(
group: torch.distributed.ProcessGroup,
device: torch.device,
supported_world_size: List[int],
cls_name: str,
) -> Optional[bool]: # None if fail; otherwise return whether NVLink is available
assert (
dist.get_backend(group) != dist.Backend.NCCL
), f"{cls_name} should be attached to a non-NCCL group."
rank = dist.get_rank(group=group)
world_size = dist.get_world_size(group=group)
# No need to initialize custom allreduce for single GPU case.
if world_size == 1:
return
# No need to initialize custom allreduce for multi-node case.
if not all(in_the_same_node_as(group, source_rank=0)):
logger.warning(
f"{cls_name} is disabled because this process group" " spans across nodes."
)
return
# For not supported world size, we disable custom allreduce.
if world_size not in supported_world_size:
logger.warning(
f"{cls_name} is disabled due to an unsupported world"
f" size: {world_size}. Supported world sizes: {supported_world_size}. "
"To silence this warning, specify disable_custom_all_reduce=True explicitly.",
)
return
cuda_visible_devices = os.environ.get("CUDA_VISIBLE_DEVICES", None)
if cuda_visible_devices:
device_ids = list(map(int, cuda_visible_devices.split(",")))
else:
device_ids = list(range(torch.cuda.device_count()))
physical_device_id = device_ids[device.index]
tensor = torch.tensor([physical_device_id], dtype=torch.int, device="cpu")
gather_list = [
torch.tensor([0], dtype=torch.int, device="cpu") for _ in range(world_size)
]
dist.all_gather(gather_list, tensor, group=group)
physical_device_ids = [int(t) for t in gather_list]
full_nvlink = is_full_nvlink(physical_device_ids, world_size)
# test nvlink first, this will filter out most of the cases
# where custom allreduce is not supported
# this checks hardware and driver support for NVLink
if False:
logger.warning(
f"{cls_name} is disabled because it's not supported on"
" more than two PCIe-only GPUs. To silence this warning, "
"specify disable_custom_all_reduce=True explicitly."
)
return
# test P2P capability, this checks software/cudaruntime support
# this is expensive to compute at the first time
# then we cache the result
# On AMD GPU, p2p is always enabled between XGMI connected GPUs
if not _is_hip and not can_p2p(rank, world_size):
logger.warning(
f"{cls_name} is disabled because your platform lacks "
"GPU P2P capability or P2P test failed. To silence this "
"warning, specify disable_custom_all_reduce=True explicitly."
)
return
return full_nvlink
if __name__ == "__main__":
batch_src, batch_tgt, output_file = pickle.loads(sys.stdin.buffer.read())
result = can_actually_p2p(batch_src, batch_tgt)
with open(output_file, "wb") as f:
f.write(pickle.dumps(result))

View File

@ -0,0 +1,20 @@
{"tag": "b3_16k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9501, "corpus_window": {"start": 2300000, "end": 2431072}, "ok": 8, "failed": 0, "wall_s": 82.58, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 49.6, "input_throughput_tok_s": 1587.15, "ttft_s": {"mean": 3.9252, "p50": 3.9266, "p95": 3.9787, "max": 3.9787, "min": 3.8899}, "tpot_s": {"mean": 0.0125, "p50": 0.0128, "p95": 0.0137, "max": 0.0137, "min": 0.0112}, "e2e_s": {"mean": 10.3227, "p50": 10.4215, "p95": 10.9871, "max": 10.9871, "min": 9.6445}, "per_req_out_tok_s_e2e": {"mean": 49.6617, "p50": 49.8845, "p95": 53.0871, "max": 53.0871, "min": 46.6001}, "per_req_decode_tok_s": {"mean": 80.2896, "p50": 80.4694, "p95": 89.5483, "max": 89.5483, "min": 72.9492}, "spec_accept_length_mean": 2.138, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 16, "new_tokens": 131072, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b3_16k_c8", "arm": "e7b", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9502, "corpus_window": {"start": 2300000, "end": 2562144}, "ok": 16, "failed": 0, "wall_s": 100.63, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 8192, "output_throughput_tok_s": 81.41, "input_throughput_tok_s": 2605.13, "ttft_s": {"mean": 11.7518, "p50": 6.5296, "p95": 32.5233, "max": 32.5233, "min": 3.9869}, "tpot_s": {"mean": 0.0716, "p50": 0.0731, "p95": 0.1351, "max": 0.1351, "min": 0.0266}, "e2e_s": {"mean": 48.3305, "p50": 50.9027, "p95": 78.2404, "max": 78.2404, "min": 17.5929}, "per_req_out_tok_s_e2e": {"mean": 12.9171, "p50": 10.4857, "p95": 29.1027, "max": 29.1027, "min": 6.5439}, "per_req_decode_tok_s": {"mean": 16.8775, "p50": 15.7396, "p95": 37.6553, "max": 37.6553, "min": 7.4162}, "spec_accept_length_mean": 2.375, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 262144, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b3_16k_c16", "arm": "e7b", "summary": {"concurrency": 16, "num_requests": 32, "run_id": 9503, "corpus_window": {"start": 2300000, "end": 2824288}, "ok": 32, "failed": 0, "wall_s": 236.44, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 16384, "output_throughput_tok_s": 69.29, "input_throughput_tok_s": 2217.4, "ttft_s": {"mean": 23.5964, "p50": 13.173, "p95": 60.864, "max": 90.3879, "min": 4.2969}, "tpot_s": {"mean": 0.1733, "p50": 0.1784, "p95": 0.2978, "max": 0.3055, "min": 0.0668}, "e2e_s": {"mean": 112.1663, "p50": 105.5499, "p95": 170.4271, "max": 183.3467, "min": 52.7259}, "per_req_out_tok_s_e2e": {"mean": 5.0312, "p50": 4.8797, "p95": 8.0542, "max": 9.7106, "min": 2.7925}, "per_req_decode_tok_s": {"mean": 6.7269, "p50": 5.8271, "p95": 12.9785, "max": 15.0103, "min": 3.2802}, "spec_accept_length_mean": 2.517, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b3_16k_c32", "arm": "e7b", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9504, "corpus_window": {"start": 2300000, "end": 3348576}, "ok": 64, "failed": 0, "wall_s": 414.96, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 32768, "output_throughput_tok_s": 78.97, "input_throughput_tok_s": 2526.91, "ttft_s": {"mean": 100.2552, "p50": 109.5841, "p95": 169.6812, "max": 190.7286, "min": 5.2261}, "tpot_s": {"mean": 0.1698, "p50": 0.1728, "p95": 0.2599, "max": 0.3191, "min": 0.0328}, "e2e_s": {"mean": 187.0411, "p50": 190.6599, "p95": 271.5134, "max": 292.7775, "min": 86.5615}, "per_req_out_tok_s_e2e": {"mean": 2.9412, "p50": 2.6906, "p95": 4.5951, "max": 5.9149, "min": 1.7488}, "per_req_decode_tok_s": {"mean": 7.1251, "p50": 5.846, "p95": 15.414, "max": 30.5614, "min": 3.1404}, "spec_accept_length_mean": 2.767, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b3_16k_c64", "arm": "e7b", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9505, "corpus_window": {"start": 2300000, "end": 4397152}, "ok": 128, "failed": 0, "wall_s": 777.55, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 65536, "output_throughput_tok_s": 84.29, "input_throughput_tok_s": 2697.14, "ttft_s": {"mean": 240.0965, "p50": 287.3833, "p95": 345.6933, "max": 380.9581, "min": 5.2288}, "tpot_s": {"mean": 0.168, "p50": 0.1702, "p95": 0.2315, "max": 0.3119, "min": 0.0404}, "e2e_s": {"mean": 325.934, "p50": 364.8118, "p95": 432.439, "max": 474.7226, "min": 83.7008}, "per_req_out_tok_s_e2e": {"mean": 1.8253, "p50": 1.4064, "p95": 4.196, "max": 6.117, "min": 1.0785}, "per_req_decode_tok_s": {"mean": 6.6174, "p50": 5.8876, "p95": 12.5326, "max": 24.7977, "min": 3.2126}, "spec_accept_length_mean": 2.986, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b41_1k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9511, "corpus_window": {"start": 4500000, "end": 4508192}, "ok": 8, "failed": 0, "wall_s": 15.82, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 1024, "output_throughput_tok_s": 64.74, "input_throughput_tok_s": 517.93, "ttft_s": {"mean": 0.3034, "p50": 0.3015, "p95": 0.3258, "max": 0.3258, "min": 0.2959}, "tpot_s": {"mean": 0.0132, "p50": 0.0135, "p95": 0.0143, "max": 0.0143, "min": 0.0119}, "e2e_s": {"mean": 1.9769, "p50": 2.0115, "p95": 2.1177, "max": 2.1177, "min": 1.8091}, "per_req_out_tok_s_e2e": {"mean": 64.9178, "p50": 66.8893, "p95": 70.753, "max": 70.753, "min": 60.442}, "per_req_decode_tok_s": {"mean": 76.7983, "p50": 79.4095, "p95": 84.9204, "max": 84.9204, "min": 70.4319}, "spec_accept_length_mean": 2.028, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 8, "new_tokens": 8192, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b41_1k_c8", "arm": "e7b", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9512, "corpus_window": {"start": 4500000, "end": 4516384}, "ok": 16, "failed": 0, "wall_s": 13.44, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 2048, "output_throughput_tok_s": 152.4, "input_throughput_tok_s": 1219.24, "ttft_s": {"mean": 1.21, "p50": 1.7505, "p95": 2.1601, "max": 2.1601, "min": 0.387}, "tpot_s": {"mean": 0.0414, "p50": 0.042, "p95": 0.0537, "max": 0.0537, "min": 0.0319}, "e2e_s": {"mean": 6.4682, "p50": 6.5412, "p95": 8.9777, "max": 8.9777, "min": 4.4513}, "per_req_out_tok_s_e2e": {"mean": 20.4798, "p50": 20.2955, "p95": 28.7556, "max": 28.7556, "min": 14.2576}, "per_req_decode_tok_s": {"mean": 24.8432, "p50": 24.8024, "p95": 31.5592, "max": 31.5592, "min": 18.7758}, "spec_accept_length_mean": 2.048, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 10, "new_tokens": 16384, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b41_1k_c32", "arm": "e7b", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9513, "corpus_window": {"start": 4500000, "end": 4565536}, "ok": 64, "failed": 0, "wall_s": 81.8, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 8192, "output_throughput_tok_s": 100.14, "input_throughput_tok_s": 801.15, "ttft_s": {"mean": 17.5622, "p50": 21.7567, "p95": 25.6003, "max": 27.8932, "min": 1.5226}, "tpot_s": {"mean": 0.1466, "p50": 0.1501, "p95": 0.1905, "max": 0.1952, "min": 0.088}, "e2e_s": {"mean": 36.1812, "p50": 40.1766, "p95": 46.8967, "max": 47.1947, "min": 18.2249}, "per_req_out_tok_s_e2e": {"mean": 3.8358, "p50": 3.1917, "p95": 6.6145, "max": 7.0234, "min": 2.7122}, "per_req_decode_tok_s": {"mean": 7.1041, "p50": 6.7208, "p95": 10.1791, "max": 11.4499, "min": 5.164}, "spec_accept_length_mean": 2.054, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 38, "new_tokens": 65536, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b41_1k_c64", "arm": "e7b", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9514, "corpus_window": {"start": 4500000, "end": 4631072}, "ok": 128, "failed": 0, "wall_s": 158.73, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 16384, "output_throughput_tok_s": 103.22, "input_throughput_tok_s": 825.75, "ttft_s": {"mean": 46.7369, "p50": 57.65, "p95": 65.8132, "max": 70.7061, "min": 1.5163}, "tpot_s": {"mean": 0.1445, "p50": 0.144, "p95": 0.1777, "max": 0.2125, "min": 0.0759}, "e2e_s": {"mean": 65.0893, "p50": 74.2949, "p95": 84.7332, "max": 89.7485, "min": 18.3747}, "per_req_out_tok_s_e2e": {"mean": 2.4029, "p50": 1.7238, "p95": 6.272, "max": 6.9661, "min": 1.4262}, "per_req_decode_tok_s": {"mean": 7.1513, "p50": 7.0024, "p95": 9.0245, "max": 13.2811, "min": 4.7422}, "spec_accept_length_mean": 2.087, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 85, "new_tokens": 131072, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b42_1k4k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9521, "corpus_window": {"start": 4700000, "end": 4708192}, "ok": 8, "failed": 0, "wall_s": 242.18, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 32768, "output_throughput_tok_s": 135.31, "input_throughput_tok_s": 33.83, "ttft_s": {"mean": 0.3142, "p50": 0.3133, "p95": 0.3204, "max": 0.3204, "min": 0.3097}, "tpot_s": {"mean": 0.0073, "p50": 0.0078, "p95": 0.0087, "max": 0.0087, "min": 0.0059}, "e2e_s": {"mean": 30.2716, "p50": 32.1753, "p95": 35.9141, "max": 35.9141, "min": 24.509}, "per_req_out_tok_s_e2e": {"mean": 137.6058, "p50": 136.5088, "p95": 167.1221, "max": 167.1221, "min": 114.0498}, "per_req_decode_tok_s": {"mean": 139.1062, "p50": 137.9513, "p95": 169.3421, "max": 169.3421, "min": 115.0748}, "spec_accept_length_mean": 3.701, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 8, "new_tokens": 8192, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b42_1k4k_c8", "arm": "e7b", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9522, "corpus_window": {"start": 4700000, "end": 4716384}, "ok": 16, "failed": 0, "wall_s": 159.24, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 65536, "output_throughput_tok_s": 411.54, "input_throughput_tok_s": 102.89, "ttft_s": {"mean": 0.9638, "p50": 1.2709, "p95": 1.6805, "max": 1.6805, "min": 0.378}, "tpot_s": {"mean": 0.0181, "p50": 0.0175, "p95": 0.0225, "max": 0.0225, "min": 0.0153}, "e2e_s": {"mean": 75.1101, "p50": 73.2342, "p95": 93.9711, "max": 93.9711, "min": 64.3187}, "per_req_out_tok_s_e2e": {"mean": 55.3003, "p50": 58.2448, "p95": 63.6829, "max": 63.6829, "min": 43.5879}, "per_req_decode_tok_s": {"mean": 55.9977, "p50": 58.5719, "p95": 65.3811, "max": 65.3811, "min": 44.3769}, "spec_accept_length_mean": 4.003, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 10, "new_tokens": 16384, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b42_1k4k_c32", "arm": "e7b", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9523, "corpus_window": {"start": 4700000, "end": 4765536}, "ok": 64, "failed": 0, "wall_s": 1040.79, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 262144, "output_throughput_tok_s": 251.87, "input_throughput_tok_s": 62.97, "ttft_s": {"mean": 195.07, "p50": 240.7323, "p95": 304.7528, "max": 334.6933, "min": 1.3472}, "tpot_s": {"mean": 0.0607, "p50": 0.0581, "p95": 0.0746, "max": 0.0923, "min": 0.0465}, "e2e_s": {"mean": 443.6795, "p50": 478.2794, "p95": 587.5712, "max": 597.8883, "min": 219.3532}, "per_req_out_tok_s_e2e": {"mean": 10.0455, "p50": 8.6106, "p95": 17.4441, "max": 18.6731, "min": 6.8508}, "per_req_decode_tok_s": {"mean": 16.784, "p50": 17.4276, "p95": 20.0358, "max": 21.5315, "min": 10.8381}, "spec_accept_length_mean": 3.852, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 51, "new_tokens": 65536, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b51_64k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9531, "corpus_window": {"start": 5000000, "end": 5524288}, "ok": 8, "failed": 0, "wall_s": 182.23, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 22.48, "input_throughput_tok_s": 2877.06, "ttft_s": {"mean": 17.0384, "p50": 17.022, "p95": 17.1703, "max": 17.1703, "min": 17.0028}, "tpot_s": {"mean": 0.0112, "p50": 0.0117, "p95": 0.0135, "max": 0.0135, "min": 0.008}, "e2e_s": {"mean": 22.7786, "p50": 23.0017, "p95": 23.9301, "max": 23.9301, "min": 21.118}, "per_req_out_tok_s_e2e": {"mean": 22.5108, "p50": 22.5486, "p95": 24.2447, "max": 24.2447, "min": 21.3957}, "per_req_decode_tok_s": {"mean": 91.4983, "p50": 90.0463, "p95": 125.007, "max": 125.007, "min": 74.0091}, "spec_accept_length_mean": 2.468, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b51_64k_c4", "arm": "e7b", "summary": {"concurrency": 4, "num_requests": 8, "run_id": 9532, "corpus_window": {"start": 5000000, "end": 5524288}, "ok": 8, "failed": 0, "wall_s": 159.24, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 25.72, "input_throughput_tok_s": 3292.45, "ttft_s": {"mean": 30.5492, "p50": 18.5951, "p95": 68.8675, "max": 68.8675, "min": 17.0513}, "tpot_s": {"mean": 0.0949, "p50": 0.1145, "p95": 0.1546, "max": 0.1546, "min": 0.02}, "e2e_s": {"mean": 79.0616, "p50": 80.1169, "p95": 131.9282, "max": 131.9282, "min": 27.2645}, "per_req_out_tok_s_e2e": {"mean": 8.1791, "p50": 6.6417, "p95": 18.779, "max": 18.779, "min": 3.8809}, "per_req_decode_tok_s": {"mean": 16.1959, "p50": 11.2172, "p95": 50.1581, "max": 50.1581, "min": 6.4814}, "spec_accept_length_mean": 2.438, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b51_64k_c8", "arm": "e7b", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9533, "corpus_window": {"start": 5000000, "end": 6048576}, "ok": 16, "failed": 0, "wall_s": 319.66, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 8192, "output_throughput_tok_s": 25.63, "input_throughput_tok_s": 3280.31, "ttft_s": {"mean": 90.9186, "p50": 96.5085, "p95": 149.4465, "max": 149.4465, "min": 18.5948}, "tpot_s": {"mean": 0.1058, "p50": 0.1221, "p95": 0.1575, "max": 0.1575, "min": 0.0172}, "e2e_s": {"mean": 144.9847, "p50": 155.6435, "p95": 211.845, "max": 211.845, "min": 75.7166}, "per_req_out_tok_s_e2e": {"mean": 3.7942, "p50": 3.6883, "p95": 6.7621, "max": 6.7621, "min": 2.4169}, "per_req_decode_tok_s": {"mean": 13.0455, "p50": 8.4248, "p95": 58.0944, "max": 58.0944, "min": 6.3604}, "spec_accept_length_mean": 2.442, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b51_128k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9541, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 343.36, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 11.93, "input_throughput_tok_s": 3053.83, "ttft_s": {"mean": 37.5917, "p50": 37.5968, "p95": 37.6012, "max": 37.6012, "min": 37.5564}, "tpot_s": {"mean": 0.0104, "p50": 0.0104, "p95": 0.0123, "max": 0.0123, "min": 0.009}, "e2e_s": {"mean": 42.9201, "p50": 42.9071, "p95": 43.9018, "max": 43.9018, "min": 42.1336}, "per_req_out_tok_s_e2e": {"mean": 11.9312, "p50": 11.9694, "p95": 12.1518, "max": 12.1518, "min": 11.6624}, "per_req_decode_tok_s": {"mean": 97.1304, "p50": 98.8664, "p95": 111.8659, "max": 111.8659, "min": 81.2653}, "spec_accept_length_mean": 2.728, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b51_128k_c2", "arm": "e7b", "summary": {"concurrency": 2, "num_requests": 8, "run_id": 9542, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 333.36, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 12.29, "input_throughput_tok_s": 3145.44, "ttft_s": {"mean": 48.2081, "p50": 39.1403, "p95": 75.8082, "max": 75.8082, "min": 37.5834}, "tpot_s": {"mean": 0.068, "p50": 0.0858, "p95": 0.1622, "max": 0.1622, "min": 0.0099}, "e2e_s": {"mean": 82.9494, "p50": 84.1258, "p95": 122.0311, "max": 122.0311, "min": 42.6804}, "per_req_out_tok_s_e2e": {"mean": 7.0763, "p50": 6.2562, "p95": 11.9961, "max": 11.9961, "min": 4.1957}, "per_req_decode_tok_s": {"mean": 37.0402, "p50": 12.1918, "p95": 100.963, "max": 100.963, "min": 6.1768}, "spec_accept_length_mean": 2.769, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b51_128k_c4", "arm": "e7b", "summary": {"concurrency": 4, "num_requests": 8, "run_id": 9543, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 334.71, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 12.24, "input_throughput_tok_s": 3132.76, "ttft_s": {"mean": 110.1385, "p50": 120.8044, "p95": 159.4963, "max": 159.4963, "min": 39.1664}, "tpot_s": {"mean": 0.0687, "p50": 0.0152, "p95": 0.1634, "max": 0.1634, "min": 0.0097}, "e2e_s": {"mean": 145.2581, "p50": 130.4964, "p95": 204.1758, "max": 204.1758, "min": 83.563}, "per_req_out_tok_s_e2e": {"mean": 3.8065, "p50": 4.04, "p95": 6.1271, "max": 6.1271, "min": 2.5076}, "per_req_decode_tok_s": {"mean": 54.2853, "p50": 73.1655, "p95": 103.6085, "max": 103.6085, "min": 6.1313}, "spec_accept_length_mean": 2.595, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b52_256k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9551, "corpus_window": {"start": 8400000, "end": 8662144}, "ok": 1, "failed": 0, "wall_s": 0.5, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 2.01, "input_throughput_tok_s": 527883.82, "ttft_s": {"mean": 0.4955, "p50": 0.4955, "p95": 0.4955, "max": 0.4955, "min": 0.4955}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 0.4957, "p50": 0.4957, "p95": 0.4957, "max": 0.4957, "min": 0.4957}, "per_req_out_tok_s_e2e": {"mean": 2.0172, "p50": 2.0172, "p95": 2.0172, "max": 2.0172, "min": 2.0172}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 1, "new_tokens": 64, "cached_tokens": 262080, "hit_rate": 0.9998}}}
{"tag": "b52_256k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9551, "corpus_window": {"start": 9900000, "end": 10162144}, "ok": 1, "failed": 0, "wall_s": 91.45, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 0.01, "input_throughput_tok_s": 2866.59, "ttft_s": {"mean": 91.4465, "p50": 91.4465, "p95": 91.4465, "max": 91.4465, "min": 91.4465}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 91.4468, "p50": 91.4468, "p95": 91.4468, "max": 91.4468, "min": 91.4468}, "per_req_out_tok_s_e2e": {"mean": 0.0109, "p50": 0.0109, "p95": 0.0109, "max": 0.0109, "min": 0.0109}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 262144, "cached_tokens": 0, "hit_rate": 0.0}}}

View File

@ -0,0 +1,3 @@
a8f3909211adb92501cf007ab71d3b29 /root/corpus_ids.json
1e34dd8d813ccb7ac8243e8d1573a49e /root/bench_corpus_v2.py
4c126d067d33b5ea27c268f37561634c /root/extract_summary.py

File diff suppressed because one or more lines are too long

View File

@ -0,0 +1,20 @@
b3_16k_c1 OK
b3_16k_c8 OK
b3_16k_c16 OK
b3_16k_c32 OK
b3_16k_c64 OK
b41_1k_c1 OK
b41_1k_c8 OK
b41_1k_c32 OK
b41_1k_c64 OK
b42_1k4k_c1 OK
b42_1k4k_c8 OK
b42_1k4k_c32 OK
b42_1k4k_c64 BENCH_FAIL rc=143
b51_64k_c1 OK
b51_64k_c4 OK
b51_64k_c8 OK
b51_128k_c1 OK
b51_128k_c2 OK
b51_128k_c4 OK
b52_256k_c1 HIT_FAIL_FINAL

View File

@ -0,0 +1,7 @@
1e34dd8d813ccb7ac8243e8d1573a49e bench_corpus_v2.py
4c126d067d33b5ea27c268f37561634c extract_summary.py
5892b44610b2ce61f533f1721105625e run_b300_matrix.sh
def3c64c5dc19e1d507080e3366c3762 deploy_glm53_pp4.sh
21db641f1d997e26fcf1ddc97163b87e deploy_glm53_607_exp.sh
a8fc9a504c041acc40831a848b4985d8 patches/custom_all_reduce.py
65a4d22bd87ae343e7d68c4879c14f98 patches/custom_all_reduce_utils.py

View File

@ -0,0 +1,46 @@
# 数据来源与核验记录provenance
原始 csv/log 按仓库惯例不入库(.gitignore `*.csv`/`*.log`),完整文件在 60.8
`/root/bench_logs/b300eq_{tp2pp4_20260910_1138,e7b_20260910_1428}/` 与本地镜像
`D:/sskj/b300eq/`。本文件固化其中的关键事实。
## GPU 清单nvidia-smi8×RTX 6000D总 85,651 MiB/卡)
### TP2PP4 臂mem0.85,加载后空载 → 矩阵结束)
| GPU | 空载 MiB | 结束 MiB |
|---|---|---|
| 0/1 | 64,613 | 79,391 |
| 2/3 | 70,867 | 81,751 |
| 4/5 | 74,499 | 84,439 / 84,631 |
| 6/7 | 75,361 | 84,491 |
vram_timeline.csv30s 采样)全程峰值 **85,013 MiB**(主场景 C=64最紧张卡余量 ~638 MiB
### E7b 臂mem0.90 + EAGLE 草稿权重,空载更高)
| GPU | 空载 MiB | 结束 MiB |
|---|---|---|
| 0 | 77,861 | 83,477 |
| 1/2/5/6 | 77,955 | 83,551 / 83,553 |
| 3/4/7 | 77,859 | 83,477 / 83,479 |
全程峰值 **83,553 MiB**(余量 ~2.1 GiB
## 质量门判决
- TP2PP4 臂:`PASS=6 FAIL=1`(唯一失败 = tool-callD 口径无 parser历史已知GSM8K×5 + 中文推理全过)
- E7b 臂:`PASS=7 FAIL=0`(含 tool-call `get_weather{"city": "北京"}`
## 在役容器保全与恢复60.8TP4PP2-nomtp@0.90 口径)
- 停役流程:`docker stop glm53-nvfp4``docker rename glm53-nvfp4 glm53-nvfp4-insvc`(先改名,防 E7b 部署脚本 rm -f 同名容器inspect/启动命令/挂载/镜像归档于 60.8 `/root/bench_logs/b300eq_meta/`
- 镜像:`lmsysorg/sglang:nightly-dev-cu13-20260901-07c8f729`sha256:eb090e39…
- 停役前显存82,221~82,395 MiB/卡
- 恢复流程E7b 测试容器 rm显存排干 0 MiB`docker rename glm53-nvfp4-insvc glm53-nvfp4 && docker start`
- 恢复核验09-10health 200启动后 ~4 min16K/16tok 冷抽测 ok=1/1、wall 3.59 s显存 GPU4-7 与停役前持平82,2xx MiB、GPU0-3 低 ~5 GiB重启后 radix 池未回填,正常);容器口径未变
- 僵尸 PID 现象记录:`docker stop`/`rm` 偶发 "container PID xxx is zombie and can not be killed",实为收尾边界现象(容器终态 exited 137、显存归零等待 ~20s 重试即成功
## 执行资产 md560.8 = 本目录 = 60.7 原件,三方一致)
`md5_ledger.txt`。corpus 语料:`/root/corpus_ids.json`21,296,780 tokens消费至 21,235,008回收窗口协议见 REPORT.md 附录 A

View File

@ -0,0 +1,24 @@
{"tag": "b3_16k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9501, "corpus_window": {"start": 2300000, "end": 2431072}, "ok": 8, "failed": 0, "wall_s": 245.55, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 16.68, "input_throughput_tok_s": 533.8, "ttft_s": {"mean": 5.3493, "p50": 5.3474, "p95": 5.3628, "max": 5.3628, "min": 5.3438}, "tpot_s": {"mean": 0.0496, "p50": 0.0495, "p95": 0.0503, "max": 0.0503, "min": 0.0493}, "e2e_s": {"mean": 30.6931, "p50": 30.657, "p95": 31.0364, "max": 31.0364, "min": 30.5593}, "per_req_out_tok_s_e2e": {"mean": 16.6817, "p50": 16.7233, "p95": 16.7543, "max": 16.7543, "min": 16.4968}, "per_req_decode_tok_s": {"mean": 20.2031, "p50": 20.2689, "p95": 20.3081, "max": 20.3081, "min": 19.9299}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b3_16k_c8", "arm": "tp2pp4", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9502, "corpus_window": {"start": 2300000, "end": 2562144}, "ok": 16, "failed": 0, "wall_s": 81.0, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 8192, "output_throughput_tok_s": 101.13, "input_throughput_tok_s": 3236.22, "ttft_s": {"mean": 10.0539, "p50": 10.6041, "p95": 14.7948, "max": 14.7948, "min": 5.2557}, "tpot_s": {"mean": 0.0595, "p50": 0.0602, "p95": 0.0709, "max": 0.0709, "min": 0.0481}, "e2e_s": {"mean": 40.4682, "p50": 41.4728, "p95": 41.6589, "max": 41.6589, "min": 39.2731}, "per_req_out_tok_s_e2e": {"mean": 12.662, "p50": 13.0063, "p95": 13.0369, "max": 13.0369, "min": 12.2903}, "per_req_decode_tok_s": {"mean": 17.0379, "p50": 17.0019, "p95": 20.8395, "max": 20.8395, "min": 14.1226}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b3_16k_c16", "arm": "tp2pp4", "summary": {"concurrency": 16, "num_requests": 32, "run_id": 9503, "corpus_window": {"start": 2300000, "end": 2824288}, "ok": 32, "failed": 0, "wall_s": 112.16, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 16384, "output_throughput_tok_s": 146.08, "input_throughput_tok_s": 4674.5, "ttft_s": {"mean": 15.5283, "p50": 16.1378, "p95": 25.7815, "max": 25.8389, "min": 5.267}, "tpot_s": {"mean": 0.0792, "p50": 0.0797, "p95": 0.0985, "max": 0.1012, "min": 0.0571}, "e2e_s": {"mean": 55.9934, "p50": 56.8023, "p95": 57.0853, "max": 57.1347, "min": 54.9678}, "per_req_out_tok_s_e2e": {"mean": 9.1467, "p50": 9.2952, "p95": 9.3139, "max": 9.3145, "min": 8.9613}, "per_req_decode_tok_s": {"mean": 12.9884, "p50": 12.7194, "p95": 16.7725, "max": 17.5544, "min": 9.9045}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b3_16k_c32", "arm": "tp2pp4", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9504, "corpus_window": {"start": 2300000, "end": 3348576}, "ok": 64, "failed": 0, "wall_s": 169.54, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 32768, "output_throughput_tok_s": 193.27, "input_throughput_tok_s": 6184.73, "ttft_s": {"mean": 26.5164, "p50": 27.1445, "p95": 46.4256, "max": 47.8906, "min": 5.2618}, "tpot_s": {"mean": 0.1137, "p50": 0.114, "p95": 0.152, "max": 0.1573, "min": 0.07}, "e2e_s": {"mean": 84.6344, "p50": 85.322, "p95": 85.7436, "max": 85.8459, "min": 83.6335}, "per_req_out_tok_s_e2e": {"mean": 6.0502, "p50": 6.108, "p95": 6.1205, "max": 6.1219, "min": 5.9642}, "per_req_decode_tok_s": {"mean": 9.2806, "p50": 8.8387, "p95": 13.3026, "max": 14.3169, "min": 6.3708}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b3_16k_c64", "arm": "tp2pp4", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9505, "corpus_window": {"start": 2300000, "end": 4397152}, "ok": 128, "failed": 0, "wall_s": 310.79, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 65536, "output_throughput_tok_s": 210.87, "input_throughput_tok_s": 6747.91, "ttft_s": {"mean": 63.0522, "p50": 49.1362, "p95": 134.7794, "max": 139.9931, "min": 5.3652}, "tpot_s": {"mean": 0.1392, "p50": 0.1358, "p95": 0.2037, "max": 0.2143, "min": 0.0707}, "e2e_s": {"mean": 134.1654, "p50": 114.4018, "p95": 226.099, "max": 226.1544, "min": 84.0644}, "per_req_out_tok_s_e2e": {"mean": 4.1916, "p50": 4.476, "p95": 6.0862, "max": 6.0906, "min": 2.2639}, "per_req_decode_tok_s": {"mean": 7.7966, "p50": 7.3815, "p95": 11.9051, "max": 14.1753, "min": 4.6761}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 512, "new_tokens": 8388608, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b41_1k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9511, "corpus_window": {"start": 4500000, "end": 4508192}, "ok": 8, "failed": 0, "wall_s": 52.95, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 1024, "output_throughput_tok_s": 19.34, "input_throughput_tok_s": 154.7, "ttft_s": {"mean": 0.3848, "p50": 0.3852, "p95": 0.3865, "max": 0.3865, "min": 0.3825}, "tpot_s": {"mean": 0.0491, "p50": 0.0491, "p95": 0.0493, "max": 0.0493, "min": 0.0488}, "e2e_s": {"mean": 6.619, "p50": 6.62, "p95": 6.6447, "max": 6.6447, "min": 6.5863}, "per_req_out_tok_s_e2e": {"mean": 19.3385, "p50": 19.3379, "p95": 19.4343, "max": 19.4343, "min": 19.2635}, "per_req_decode_tok_s": {"mean": 20.5327, "p50": 20.5305, "p95": 20.6346, "max": 20.6346, "min": 20.4442}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 32768, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b41_1k_c8", "arm": "tp2pp4", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9512, "corpus_window": {"start": 4500000, "end": 4516384}, "ok": 16, "failed": 0, "wall_s": 20.1, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 2048, "output_throughput_tok_s": 101.91, "input_throughput_tok_s": 815.3, "ttft_s": {"mean": 1.3131, "p50": 1.2998, "p95": 1.6901, "max": 1.6901, "min": 0.6481}, "tpot_s": {"mean": 0.0686, "p50": 0.0684, "p95": 0.0729, "max": 0.0729, "min": 0.0669}, "e2e_s": {"mean": 10.024, "p50": 10.1096, "p95": 10.2887, "max": 10.2887, "min": 9.7907}, "per_req_out_tok_s_e2e": {"mean": 12.7741, "p50": 12.9718, "p95": 13.0736, "max": 13.0736, "min": 12.4409}, "per_req_decode_tok_s": {"mean": 14.7032, "p50": 14.7329, "p95": 15.0683, "max": 15.0683, "min": 13.8323}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 24, "new_tokens": 65536, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b41_1k_c32", "arm": "tp2pp4", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9513, "corpus_window": {"start": 4500000, "end": 4565536}, "ok": 64, "failed": 0, "wall_s": 29.73, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 8192, "output_throughput_tok_s": 275.55, "input_throughput_tok_s": 2204.4, "ttft_s": {"mean": 3.8882, "p50": 3.9615, "p95": 4.8674, "max": 4.8759, "min": 1.5565}, "tpot_s": {"mean": 0.0856, "p50": 0.0869, "p95": 0.0985, "max": 0.1068, "min": 0.0785}, "e2e_s": {"mean": 14.758, "p50": 14.8384, "p95": 15.0559, "max": 15.1157, "min": 14.5917}, "per_req_out_tok_s_e2e": {"mean": 8.6742, "p50": 8.7413, "p95": 8.7701, "max": 8.7721, "min": 8.468}, "per_req_decode_tok_s": {"mean": 11.8377, "p50": 11.6015, "p95": 12.822, "max": 12.8317, "min": 9.4403}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 36, "new_tokens": 262144, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b41_1k_c64", "arm": "tp2pp4", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9514, "corpus_window": {"start": 4500000, "end": 4631072}, "ok": 128, "failed": 0, "wall_s": 48.11, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 16384, "output_throughput_tok_s": 340.53, "input_throughput_tok_s": 2724.25, "ttft_s": {"mean": 8.1393, "p50": 4.8485, "p95": 19.9553, "max": 20.0096, "min": 0.3836}, "tpot_s": {"mean": 0.0933, "p50": 0.0924, "p95": 0.1018, "max": 0.1251, "min": 0.0818}, "e2e_s": {"mean": 19.9938, "p50": 16.0958, "p95": 32.0105, "max": 32.0464, "min": 15.7372}, "per_req_out_tok_s_e2e": {"mean": 6.9903, "p50": 7.9525, "p95": 8.1268, "max": 8.1336, "min": 3.9942}, "per_req_decode_tok_s": {"mean": 10.8849, "p50": 10.9086, "p95": 12.2874, "max": 12.3245, "min": 8.0552}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 52, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b42_1k4k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9521, "corpus_window": {"start": 4700000, "end": 4708192}, "ok": 8, "failed": 0, "wall_s": 1614.16, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 32768, "output_throughput_tok_s": 20.3, "input_throughput_tok_s": 5.08, "ttft_s": {"mean": 0.3926, "p50": 0.394, "p95": 0.3952, "max": 0.3952, "min": 0.3852}, "tpot_s": {"mean": 0.0492, "p50": 0.0491, "p95": 0.0498, "max": 0.0498, "min": 0.049}, "e2e_s": {"mean": 201.7695, "p50": 201.6506, "p95": 204.4013, "max": 204.4013, "min": 200.9309}, "per_req_out_tok_s_e2e": {"mean": 20.3009, "p50": 20.3439, "p95": 20.3851, "max": 20.3851, "min": 20.039}, "per_req_decode_tok_s": {"mean": 20.3407, "p50": 20.3837, "p95": 20.4254, "max": 20.4254, "min": 20.078}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 32768, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b42_1k4k_c8", "arm": "tp2pp4", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9522, "corpus_window": {"start": 4700000, "end": 4716384}, "ok": 16, "failed": 0, "wall_s": 649.69, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 65536, "output_throughput_tok_s": 100.87, "input_throughput_tok_s": 25.22, "ttft_s": {"mean": 1.352, "p50": 1.4385, "p95": 1.8726, "max": 1.8726, "min": 0.4364}, "tpot_s": {"mean": 0.079, "p50": 0.0791, "p95": 0.0794, "max": 0.0794, "min": 0.0788}, "e2e_s": {"mean": 324.8285, "p50": 325.5653, "p95": 325.6331, "max": 325.6331, "min": 324.0447}, "per_req_out_tok_s_e2e": {"mean": 12.6098, "p50": 12.6389, "p95": 12.6402, "max": 12.6402, "min": 12.5786}, "per_req_decode_tok_s": {"mean": 12.6626, "p50": 12.6734, "p95": 12.7003, "max": 12.7003, "min": 12.5955}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 20, "new_tokens": 65536, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b42_1k4k_c32", "arm": "tp2pp4", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9523, "corpus_window": {"start": 4700000, "end": 4765536}, "ok": 64, "failed": 0, "wall_s": 714.14, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 262144, "output_throughput_tok_s": 367.08, "input_throughput_tok_s": 91.77, "ttft_s": {"mean": 1.9772, "p50": 1.1755, "p95": 4.1378, "max": 4.1467, "min": 0.3867}, "tpot_s": {"mean": 0.0861, "p50": 0.0887, "p95": 0.0905, "max": 0.0911, "min": 0.0805}, "e2e_s": {"mean": 354.393, "p50": 367.2171, "p95": 373.3199, "max": 374.0319, "min": 330.2086}, "per_req_out_tok_s_e2e": {"mean": 11.5833, "p50": 11.8878, "p95": 12.2953, "max": 12.4043, "min": 10.9509}, "per_req_decode_tok_s": {"mean": 11.6457, "p50": 11.9132, "p95": 12.3308, "max": 12.4283, "min": 10.9854}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 262144, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b42_1k4k_c64", "arm": "tp2pp4", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9524, "corpus_window": {"start": 4700000, "end": 4831072}, "ok": 128, "failed": 0, "wall_s": 1088.63, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 524288, "output_throughput_tok_s": 481.6, "input_throughput_tok_s": 120.4, "ttft_s": {"mean": 80.7415, "p50": 4.7569, "p95": 337.9195, "max": 338.4726, "min": 0.3985}, "tpot_s": {"mean": 0.0912, "p50": 0.0914, "p95": 0.0955, "max": 0.0963, "min": 0.0852}, "e2e_s": {"mean": 454.4024, "p50": 385.4538, "p95": 718.7004, "max": 732.2026, "min": 351.072}, "per_req_out_tok_s_e2e": {"mean": 9.6298, "p50": 10.6531, "p95": 11.4339, "max": 11.6671, "min": 5.5941}, "per_req_decode_tok_s": {"mean": 10.9736, "p50": 10.9486, "p95": 11.4764, "max": 11.7362, "min": 10.3887}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 200, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b51_64k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9531, "corpus_window": {"start": 5000000, "end": 5524288}, "ok": 8, "failed": 0, "wall_s": 285.91, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 14.33, "input_throughput_tok_s": 1833.72, "ttft_s": {"mean": 10.2811, "p50": 10.2794, "p95": 10.3216, "max": 10.3216, "min": 10.2634}, "tpot_s": {"mean": 0.0498, "p50": 0.0499, "p95": 0.05, "max": 0.05, "min": 0.0496}, "e2e_s": {"mean": 35.739, "p50": 35.7693, "p95": 35.838, "max": 35.838, "min": 35.6073}, "per_req_out_tok_s_e2e": {"mean": 14.3261, "p50": 14.3345, "p95": 14.3791, "max": 14.3791, "min": 14.2865}, "per_req_decode_tok_s": {"mean": 20.112, "p50": 20.1156, "p95": 20.2096, "max": 20.2096, "min": 20.052}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b51_64k_c4", "arm": "tp2pp4", "summary": {"concurrency": 4, "num_requests": 8, "run_id": 9532, "corpus_window": {"start": 5000000, "end": 5524288}, "ok": 8, "failed": 0, "wall_s": 122.3, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 33.49, "input_throughput_tok_s": 4286.94, "ttft_s": {"mean": 19.5548, "p50": 22.5526, "p95": 28.8262, "max": 28.8262, "min": 10.2823}, "tpot_s": {"mean": 0.0814, "p50": 0.0873, "p95": 0.0995, "max": 0.0995, "min": 0.0632}, "e2e_s": {"mean": 61.1314, "p50": 61.2078, "p95": 61.2548, "max": 61.2548, "min": 61.002}, "per_req_out_tok_s_e2e": {"mean": 8.3754, "p50": 8.3851, "p95": 8.3932, "max": 8.3932, "min": 8.3585}, "per_req_decode_tok_s": {"mean": 12.6679, "p50": 13.2834, "p95": 15.8466, "max": 15.8466, "min": 10.0675}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b51_64k_c8", "arm": "tp2pp4", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9533, "corpus_window": {"start": 5000000, "end": 6048576}, "ok": 16, "failed": 0, "wall_s": 185.59, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 8192, "output_throughput_tok_s": 44.14, "input_throughput_tok_s": 5650.11, "ttft_s": {"mean": 31.833, "p50": 34.8143, "p95": 53.3841, "max": 53.3841, "min": 10.2802}, "tpot_s": {"mean": 0.1192, "p50": 0.1246, "p95": 0.1618, "max": 0.1618, "min": 0.0765}, "e2e_s": {"mean": 92.7517, "p50": 92.8482, "p95": 92.9902, "max": 92.9902, "min": 92.4957}, "per_req_out_tok_s_e2e": {"mean": 5.5201, "p50": 5.5226, "p95": 5.5354, "max": 5.5354, "min": 5.506}, "per_req_decode_tok_s": {"mean": 8.9039, "p50": 8.8159, "p95": 13.0909, "max": 13.0909, "min": 6.1915}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b51_128k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9541, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 347.5, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 11.79, "input_throughput_tok_s": 3017.45, "ttft_s": {"mean": 18.2074, "p50": 18.2043, "p95": 18.2252, "max": 18.2252, "min": 18.1922}, "tpot_s": {"mean": 0.0494, "p50": 0.0496, "p95": 0.0497, "max": 0.0497, "min": 0.0485}, "e2e_s": {"mean": 43.4375, "p50": 43.5633, "p95": 43.6252, "max": 43.6252, "min": 43.0007}, "per_req_out_tok_s_e2e": {"mean": 11.7874, "p50": 11.7541, "p95": 11.9068, "max": 11.9068, "min": 11.7363}, "per_req_decode_tok_s": {"mean": 20.2951, "p50": 20.1994, "p95": 20.6444, "max": 20.6444, "min": 20.1578}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b51_128k_c2", "arm": "tp2pp4", "summary": {"concurrency": 2, "num_requests": 8, "run_id": 9542, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 243.3, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 16.84, "input_throughput_tok_s": 4309.84, "ttft_s": {"mean": 25.2658, "p50": 32.1952, "p95": 32.4188, "max": 32.4188, "min": 18.173}, "tpot_s": {"mean": 0.0696, "p50": 0.0831, "p95": 0.0837, "max": 0.0837, "min": 0.0556}, "e2e_s": {"mean": 60.8183, "p50": 60.7483, "p95": 61.2374, "max": 61.2374, "min": 60.6134}, "per_req_out_tok_s_e2e": {"mean": 8.4186, "p50": 8.4352, "p95": 8.447, "max": 8.447, "min": 8.3609}, "per_req_decode_tok_s": {"mean": 14.9825, "p50": 17.8153, "p95": 18.0089, "max": 18.0089, "min": 11.965}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b51_128k_c4", "arm": "tp2pp4", "summary": {"concurrency": 4, "num_requests": 8, "run_id": 9543, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 186.37, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 21.98, "input_throughput_tok_s": 5626.35, "ttft_s": {"mean": 39.2939, "p50": 46.1927, "p95": 60.3496, "max": 60.3496, "min": 18.2774}, "tpot_s": {"mean": 0.1054, "p50": 0.1187, "p95": 0.1468, "max": 0.1468, "min": 0.0638}, "e2e_s": {"mean": 93.1534, "p50": 93.2237, "p95": 93.3139, "max": 93.3139, "min": 92.97}, "per_req_out_tok_s_e2e": {"mean": 5.4963, "p50": 5.497, "p95": 5.5072, "max": 5.5072, "min": 5.4869}, "per_req_decode_tok_s": {"mean": 10.4415, "p50": 10.8799, "p95": 15.6959, "max": 15.6959, "min": 6.8234}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b52_256k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9551, "corpus_window": {"start": 8400000, "end": 8662144}, "ok": 1, "failed": 0, "wall_s": 37.18, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 0.03, "input_throughput_tok_s": 7050.23, "ttft_s": {"mean": 37.1811, "p50": 37.1811, "p95": 37.1811, "max": 37.1811, "min": 37.1811}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 37.1814, "p50": 37.1814, "p95": 37.1814, "max": 37.1814, "min": 37.1814}, "per_req_out_tok_s_e2e": {"mean": 0.0269, "p50": 0.0269, "p95": 0.0269, "max": 0.0269, "min": 0.0269}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b52_256k_c2", "arm": "tp2pp4", "summary": {"concurrency": 2, "num_requests": 2, "run_id": 9552, "corpus_window": {"start": 8400000, "end": 8924288}, "ok": 2, "failed": 0, "wall_s": 70.19, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 2, "output_throughput_tok_s": 0.03, "input_throughput_tok_s": 7469.22, "ttft_s": {"mean": 53.6945, "p50": 70.1229, "p95": 70.1229, "max": 70.1229, "min": 37.2662}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 53.6948, "p50": 70.1232, "p95": 70.1232, "max": 70.1232, "min": 37.2664}, "per_req_out_tok_s_e2e": {"mean": 0.0205, "p50": 0.0268, "p95": 0.0268, "max": 0.0268, "min": 0.0143}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b52_256k_c3", "arm": "tp2pp4", "summary": {"concurrency": 3, "num_requests": 3, "run_id": 9553, "corpus_window": {"start": 8400000, "end": 9186432}, "ok": 3, "failed": 0, "wall_s": 103.12, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 3, "output_throughput_tok_s": 0.03, "input_throughput_tok_s": 7626.22, "ttft_s": {"mean": 70.1194, "p50": 70.1196, "p95": 102.986, "max": 102.986, "min": 37.2526}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 70.1196, "p50": 70.1198, "p95": 102.9863, "max": 102.9863, "min": 37.2529}, "per_req_out_tok_s_e2e": {"mean": 0.0169, "p50": 0.0143, "p95": 0.0268, "max": 0.0268, "min": 0.0097}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 192, "new_tokens": 3145728, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b52_512k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9554, "corpus_window": {"start": 9300000, "end": 9824288}, "ok": 1, "failed": 0, "wall_s": 92.6, "input_len": 524288, "shared_len": 0, "unique_len": 524288, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 0.01, "input_throughput_tok_s": 5662.03, "ttft_s": {"mean": 92.5959, "p50": 92.5959, "p95": 92.5959, "max": 92.5959, "min": 92.5959}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 92.5962, "p50": 92.5962, "p95": 92.5962, "max": 92.5962, "min": 92.5962}, "per_req_out_tok_s_e2e": {"mean": 0.0108, "p50": 0.0108, "p95": 0.0108, "max": 0.0108, "min": 0.0108}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}}
{"tag": "b52_896k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9555, "corpus_window": {"start": 9900000, "end": 10817504}, "ok": 1, "failed": 0, "wall_s": 220.91, "input_len": 917504, "shared_len": 0, "unique_len": 917504, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 0.0, "input_throughput_tok_s": 4153.24, "ttft_s": {"mean": 220.9117, "p50": 220.9117, "p95": 220.9117, "max": 220.9117, "min": 220.9117}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 220.912, "p50": 220.912, "p95": 220.912, "max": 220.912, "min": 220.912}, "per_req_out_tok_s_e2e": {"mean": 0.0045, "p50": 0.0045, "p95": 0.0045, "max": 0.0045, "min": 0.0045}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 224, "new_tokens": 3670016, "cached_tokens": 0, "hit_rate": 0.0}}}

View File

@ -0,0 +1,3 @@
a8f3909211adb92501cf007ab71d3b29 /root/corpus_ids.json
1e34dd8d813ccb7ac8243e8d1573a49e /root/bench_corpus_v2.py
4c126d067d33b5ea27c268f37561634c /root/extract_summary.py

File diff suppressed because one or more lines are too long

View File

@ -0,0 +1,24 @@
b3_16k_c1 OK
b3_16k_c8 OK
b3_16k_c16 OK
b3_16k_c32 OK
b3_16k_c64 OK
b41_1k_c1 OK
b41_1k_c8 OK
b41_1k_c32 OK
b41_1k_c64 OK
b42_1k4k_c1 OK
b42_1k4k_c8 OK
b42_1k4k_c32 OK
b42_1k4k_c64 OK
b51_64k_c1 OK
b51_64k_c4 OK
b51_64k_c8 OK
b51_128k_c1 OK
b51_128k_c2 OK
b51_128k_c4 OK
b52_256k_c1 OK
b52_256k_c2 OK
b52_256k_c3 OK
b52_512k_c1 OK
b52_896k_c1 OK

View File

@ -0,0 +1,319 @@
#!/usr/bin/env python3
"""Real-corpus (PG19) benchmark for sglang GLM-5.3-NVFP4.
Same methodology as bench_hit90.py (input_ids direct to /generate, temp 0,
ignore_eos, stream, server-side completion_tokens counting, hit-rate verified
from scheduler logs), but prompts are token slices of REAL book text tokenized
with the served model's own tokenizer, replacing random ids.
Corpus file: JSON {"ids": [flat token ids], "books": [[start, end), ...]}
Fixed token-offset layout into the flat id array:
[0, 117968) shared prefix for 128k points (90% of 131072)
[0, 58976) shared prefix for 64k points (same region, shorter cut)
pool A [131072, +4*8*13104) 128k unique suffixes, run-ids 9301-9304
pool B [550400, +4*8*6560) 64k unique suffixes, run-ids 9305-9308
pool C [760320, 262144+524288+524288) 16k fully-unique prompts, run-ids 9311-9313
spare [2071040, end) warmup slices / re-run margin
Each run-id maps to one non-overlapping window (one window = one bench point);
re-running a point with fresh text = bump --pool-override past the spare base.
v2 changes (b300-equivalent campaign, 2026-09-10):
- stats(): p95 added (nearest-rank) to match the B300 report metric contract
(TTFT P95 / TPOT P95).
- --dump-records PATH: per-request records (ttft/e2e/tpot/n_out/retractions/
spec_accept_len) written as JSONL for post-hoc percentile checks.
- With --shared-frac 0 + --pool-override, any input length is supported
(1024 / 16384 / 65536 / 131072 / 262144 / 524288 / 917504).
Usage:
python3 bench_corpus_v2.py --corpus /root/corpus_ids.json --input-len 16384 \
--concurrency 64 --num-requests 128 --run-id 9505 --shared-frac 0 \
--pool-override 2300000 --output-len 512 \
--dump-records /root/bench_logs/xx/point_records.jsonl
"""
import argparse
import datetime
import json
import math
import re
import statistics
import subprocess
import time
from concurrent.futures import ThreadPoolExecutor
import requests
OUTPUT_LEN_DEFAULT = 512
CORPUS_DEFAULT = "/root/corpus_ids.json"
# fixed pool layout (see docstring)
S1_128K_RID0, S1_64K_RID0, S2_RID0 = 9301, 9305, 9311
POOL_A_BASE, POOL_A_PER = 131072, 8 * 13104 # 128k suffix windows
POOL_B_BASE = POOL_A_BASE + 4 * POOL_A_PER # 550400
POOL_B_PER = 8 * 6560 # 64k suffix windows
POOL_C_BASE = POOL_B_BASE + 4 * POOL_B_PER # 760320
POOL_C_SIZES = [16 * 16384, 32 * 16384, 32 * 16384] # cc8 / cc16 / cc32
SPARE_BASE = POOL_C_BASE + sum(POOL_C_SIZES) # 2071040
sess = requests.Session()
sess.trust_env = False # bypass any proxy env on the host
def split_lens(input_len, shared_frac):
# unique suffix = (1 - shared_frac) of the prompt, page-16 aligned
# (128k @0.9 -> 13104 unique; 16k @0.0 -> fully unique prompts)
unique = round(input_len * (1.0 - shared_frac) / 16) * 16
return input_len - unique, unique
def pool_start_for(input_len, shared_frac, run_id, override):
if override is not None:
return override
if shared_frac > 0:
if input_len == 131072:
idx = run_id - S1_128K_RID0
if not 0 <= idx < 4:
sys_exit_bad_runid(run_id, "128k points use run-ids 9301-9304")
return POOL_A_BASE + idx * POOL_A_PER
if input_len == 65536:
idx = run_id - S1_64K_RID0
if not 0 <= idx < 4:
sys_exit_bad_runid(run_id, "64k points use run-ids 9305-9308")
return POOL_B_BASE + idx * POOL_B_PER
sys_exit_bad_runid(run_id, "shared-frac>0 supports 131072/65536 only")
idx = run_id - S2_RID0
if not 0 <= idx < 3:
sys_exit_bad_runid(run_id, "16k unique points use run-ids 9311-9313")
return POOL_C_BASE + sum(POOL_C_SIZES[:idx])
def sys_exit_bad_runid(run_id, msg):
raise SystemExit(f"[pool] run-id {run_id} outside expected set: {msg}")
def build_prompts(ids, shared_len, unique_len, num_requests, pool_start):
if shared_len:
shared = ids[0:shared_len]
else:
shared = []
end = pool_start + num_requests * unique_len
if end > len(ids):
raise SystemExit(
f"[pool] window [{pool_start}, {end}) exceeds corpus ({len(ids)} ids); "
f"use --pool-override or a larger corpus")
prompts = []
for i in range(num_requests):
s = pool_start + i * unique_len
prompts.append(shared + ids[s:s + unique_len])
return shared, prompts, (pool_start, end)
def warmup(url, ids, shared_len):
# primes the radix cache with the shared prefix (same role as in bench_hit90);
# warm slice comes from the spare region so it never collides with a pool window
if len(ids) >= SPARE_BASE + 64:
warm_slice = ids[SPARE_BASE:SPARE_BASE + 64]
else:
warm_slice = ids[-64:]
payload = {
"input_ids": ids[0:shared_len] + warm_slice if shared_len else warm_slice,
"sampling_params": {"max_new_tokens": 8, "temperature": 0.0, "ignore_eos": True},
}
t0 = time.perf_counter()
r = sess.post(url, json=payload, timeout=1800)
dt = time.perf_counter() - t0
print(f"[warmup] http={r.status_code} wall={dt:.2f}s", flush=True)
def bench_one(url, prompt, output_len, idx, results):
payload = {
"input_ids": prompt,
"sampling_params": {"max_new_tokens": output_len, "temperature": 0.0, "ignore_eos": True},
"stream": True,
}
rec = {"idx": idx}
t0 = time.perf_counter()
first = last = None
first_ct = None
final_meta = None
max_ct = 0
try:
with sess.post(url, json=payload, stream=True, timeout=3600) as resp:
for raw in resp.iter_lines():
if not raw or not raw.startswith(b"data:"):
continue
body = raw[5:].strip()
if body == b"[DONE]":
continue
now = time.perf_counter()
try:
d = json.loads(body)
except Exception:
continue
mi = d.get("meta_info") or {}
ct = mi.get("completion_tokens") or 0
if ct:
max_ct = max(max_ct, ct)
if first is None:
first = now
first_ct = ct
last = now
if mi.get("finish_reason"):
final_meta = mi
t_end = time.perf_counter()
n_out = max_ct or ((final_meta or {}).get("completion_tokens") or 0)
decode_span = (last - first) if (first and last and last > first) else 0.0
rec.update(
ok=n_out > 0,
ttft=(first - t0) if first else None,
e2e=t_end - t0,
n_out=n_out,
first_chunk_tokens=first_ct,
decode_span=decode_span,
tpot=(decode_span / (n_out - 1)) if n_out > 1 else None,
per_req_decode_tok_s=(n_out / decode_span) if decode_span > 0 else None,
retractions=(final_meta or {}).get("num_retractions"),
spec_accept_len=(final_meta or {}).get("spec_accept_length"),
)
except Exception as e:
rec.update(ok=False, error=repr(e))
results[idx] = rec
def verify_hit_rate(container, t_start, t_end):
def rfc3339(epoch):
return (datetime.datetime.fromtimestamp(epoch, tz=datetime.timezone.utc)
.isoformat().replace("+00:00", "Z"))
try:
# No margin before t_start: warmup's prefill lines end strictly before it,
# and catching them would deflate the measured hit rate.
p = subprocess.run(
["docker", "logs", container, "--since", rfc3339(t_start), "--until", rfc3339(t_end + 2)],
capture_output=True, text=True, timeout=120)
text = p.stdout + p.stderr
except Exception as e:
return {"error": repr(e)}
pat = re.compile(r"#new-token: (\d+).*?#cached-token: (\d+)")
n_batches = new_tok = cached_tok = 0
for line in text.splitlines():
if "TP0]" not in line or "Prefill batch" not in line:
continue
m = pat.search(line)
if m:
n_batches += 1
new_tok += int(m.group(1))
cached_tok += int(m.group(2))
total = new_tok + cached_tok
return {
"prefill_batches": n_batches,
"new_tokens": new_tok,
"cached_tokens": cached_tok,
"hit_rate": round(cached_tok / total, 4) if total else None,
}
def stats(vals):
vals = [v for v in vals if v is not None]
if not vals:
return {"mean": None, "p50": None, "p95": None, "max": None, "min": None}
s = sorted(vals)
# nearest-rank p95: smallest value >= 95th percentile
p95_idx = max(0, math.ceil(0.95 * len(s)) - 1)
return {
"mean": round(statistics.fmean(vals), 4),
"p50": round(s[len(s) // 2], 4),
"p95": round(s[p95_idx], 4),
"max": round(s[-1], 4),
"min": round(s[0], 4),
}
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--concurrency", type=int, required=True)
ap.add_argument("--num-requests", type=int, required=True)
ap.add_argument("--run-id", type=int, required=True)
ap.add_argument("--input-len", type=int, required=True, help="token length of each prompt")
ap.add_argument("--output-len", type=int, default=OUTPUT_LEN_DEFAULT)
ap.add_argument("--shared-frac", type=float, default=0.9)
ap.add_argument("--corpus", default=CORPUS_DEFAULT)
ap.add_argument("--pool-override", type=int, default=None,
help="explicit corpus offset for the unique-suffix window (re-runs)")
ap.add_argument("--url", default="http://127.0.0.1:30000/generate")
ap.add_argument("--container", default="glm53-nvfp4")
ap.add_argument("--dump-records", default=None,
help="write per-request records as JSONL to this path")
args = ap.parse_args()
with open(args.corpus) as f:
corpus = json.load(f)
ids = corpus["ids"]
shared_len, unique_len = split_lens(args.input_len, args.shared_frac)
pool_start = pool_start_for(args.input_len, args.shared_frac, args.run_id, args.pool_override)
shared, prompts, window = build_prompts(ids, shared_len, unique_len, args.num_requests, pool_start)
print(f"[pool] window={window} shared_len={shared_len} unique_len={unique_len} "
f"corpus_total={len(ids)}", flush=True)
warmup(args.url, ids, shared_len)
results = {}
t_start = time.time()
t0 = time.perf_counter()
with ThreadPoolExecutor(max_workers=args.concurrency) as ex:
futs = [ex.submit(bench_one, args.url, p, args.output_len, i, results)
for i, p in enumerate(prompts)]
for f in futs:
f.result()
wall = time.perf_counter() - t0
t_end = time.time()
if args.dump_records:
with open(args.dump_records, "w") as f:
for i in sorted(results):
f.write(json.dumps(results[i]) + "\n")
hit = verify_hit_rate(args.container, t_start, t_end)
ok = [r for r in results.values() if r.get("ok")]
n_out_total = sum(r["n_out"] for r in ok)
out_tps = [r["n_out"] / r["e2e"] for r in ok if r.get("e2e")]
ttft = stats([r.get("ttft") for r in ok])
tpot = stats([r.get("tpot") for r in ok])
e2e = stats([r.get("e2e") for r in ok])
dec = stats([r.get("per_req_decode_tok_s") for r in ok])
spec = [v for v in (r.get("spec_accept_len") for r in ok) if v is not None]
retr = sum(r.get("retractions") or 0 for r in ok)
summary = {
"concurrency": args.concurrency,
"num_requests": args.num_requests,
"run_id": args.run_id,
"corpus_window": {"start": window[0], "end": window[1]},
"ok": len(ok),
"failed": args.num_requests - len(ok),
"wall_s": round(wall, 2),
"input_len": args.input_len,
"shared_len": shared_len,
"unique_len": unique_len,
"output_len": args.output_len,
"output_tokens_total": n_out_total,
"output_throughput_tok_s": round(n_out_total / wall, 2) if wall else None,
"input_throughput_tok_s": round(args.input_len * len(ok) / wall, 2) if wall else None,
"ttft_s": ttft,
"tpot_s": tpot,
"e2e_s": e2e,
"per_req_out_tok_s_e2e": stats(out_tps),
"per_req_decode_tok_s": dec,
"spec_accept_length_mean": round(statistics.fmean(spec), 3) if spec else None,
"retractions_total": retr,
"cache_hit_from_logs": hit,
}
print("\n===== SUMMARY =====")
print(json.dumps(summary, indent=2), flush=True)
if __name__ == "__main__":
main()

View File

@ -0,0 +1,39 @@
#!/usr/bin/env python3
"""extract_summary.py <bench_log> <tag> <arm> <results_jsonl>
Parse the trailing "===== SUMMARY =====" JSON block from a bench_corpus_v2 log,
append a tagged record to the campaign results JSONL, and print the headline
metrics. Exit codes: 0 = ok and hit_rate<=0.01 (cold-point validity),
3 = no SUMMARY block found, 4 = hit_rate above the cold-point threshold.
"""
import json
import sys
HIT_MAX = 0.01
MARK = "===== SUMMARY ====="
def main():
logp, tag, arm, outp = sys.argv[1:5]
with open(logp, encoding="utf-8", errors="replace") as f:
log = f.read()
i = log.rfind(MARK)
if i < 0:
print("NO_SUMMARY")
sys.exit(3)
s = json.loads(log[i + len(MARK):].strip())
with open(outp, "a", encoding="utf-8") as f:
f.write(json.dumps({"tag": tag, "arm": arm, "summary": s}) + "\n")
hit = (s.get("cache_hit_from_logs") or {}).get("hit_rate")
print(f"hit_rate={hit} out_tps={s.get('output_throughput_tok_s')} "
f"in_tps={s.get('input_throughput_tok_s')} "
f"ttft_p95={(s.get('ttft_s') or {}).get('p95')} "
f"tpot_p95={(s.get('tpot_s') or {}).get('p95')} "
f"ok={s.get('ok')}/{s.get('num_requests')} retractions={s.get('retractions_total')}")
if hit is None:
sys.exit(4)
sys.exit(0 if float(hit) <= HIT_MAX else 4)
if __name__ == "__main__":
main()

View File

@ -0,0 +1,165 @@
#!/usr/bin/env python3
"""gen_report_tables.py <tp2pp4_all_results.jsonl> <e7b_all_results.jsonl>
Build the B300-style markdown tables for the 6000D dual-plan report from the
campaign all_results.jsonl files. Prints tables to stdout. Best value per
column within each scenario table is bolded (**...**), mirroring the B300
report convention.
"""
import json
import sys
ARMS = ["tp2pp4", "e7b"]
ARM_LABEL = {"tp2pp4": "TP2PP4", "e7b": "TP8+EAGLE3+AR"}
# scenario key -> (chapter title, row label, ordered ccs, has_input_col)
SCENARIOS = [
("b3_16k", "主场景 16K→512", [1, 8, 16, 32, 64], True),
("b41_1k", "4.1 短输入 1K→128", [1, 8, 32, 64], True),
("b42_1k4k", "4.2 长输出 1K→4K", [1, 8, 32, 64], False),
("b51_64k", "5.1 64K→512", [1, 4, 8], True),
("b51_128k", "5.1 128K→512", [1, 2, 4], True),
("b52_256k", "5.2 边界 256K→1", [1, 2, 3], True),
("b52_512k", "5.2 边界 512K→1", [1], True),
("b52_896k", "5.2 边界 896K→1", [1], True),
]
def load(path):
rows = {}
if not path:
return rows
with open(path, encoding="utf-8") as f:
for line in f:
rec = json.loads(line)
rows[(rec["tag"], rec["arm"])] = rec["summary"]
return rows
def fmt(v, kind):
if v is None:
return ""
if kind == "tps":
return f"{v:,.0f}" if v >= 100 else f"{v:,.1f}"
if kind == "s":
return f"{v:,.2f} s" if v >= 1 else f"{v*1000:,.0f} ms"
if kind == "ms":
return f"{v*1000:,.1f} ms"
return str(v)
def get(s, path):
cur = s
for key in path.split("."):
if cur is None:
return None
cur = cur.get(key)
return cur
def table_for(prefix, title, ccs, has_input, data):
lines = []
cols = ["并发", "方案"]
if has_input:
cols += ["Input TPS"]
cols += ["Output TPS", "TTFT P95", "TPOT P95"]
lines.append("| " + " | ".join(cols) + " |")
lines.append("|" + "---|" * len(cols))
# collect cells to bold best per numeric column
body = []
for cc in ccs:
for arm in ARMS:
tag = f"{prefix}_c{cc}"
s = data.get((tag, arm))
if s is None and cc in (2, 3):
# boundary cc>1 only exists on tp2pp4; absence = structural skip
continue
if s is None:
body.append((cc, arm, None))
continue
body.append((cc, arm, s))
def val(cc, arm, path):
s = data.get((f"{prefix}_c{cc}", arm))
return get(s, path) if s else None
# best per column (higher tps, lower latency)
best = {}
if has_input:
ivs = [val(cc, arm, "input_throughput_tok_s") for cc, arm, _ in body if _ is not None]
ivs = [v for v in ivs if v is not None]
if ivs:
best["in"] = max(ivs)
ovs = [val(cc, arm, "output_throughput_tok_s") for cc, arm, _ in body if _ is not None]
ovs = [v for v in ovs if v is not None]
if ovs:
best["out"] = max(ovs)
tvs = [val(cc, arm, "ttft_s.p95") for cc, arm, _ in body if _ is not None]
tvs = [v for v in tvs if v is not None]
if tvs:
best["ttft"] = min(tvs)
pvs = [val(cc, arm, "tpot_s.p95") for cc, arm, _ in body if _ is not None]
pvs = [v for v in pvs if v is not None]
if pvs:
best["tpot"] = min(pvs)
def maybe_bold(v, kind, key):
if v is None:
return ""
cell = fmt(v, kind)
if key in best and v == best[key]:
return f"**{cell}**"
return cell
for cc, arm, s in body:
if s is None:
lines.append(f"| {cc} | {ARM_LABEL[arm]} | " + " 结构性不可测 |" * (len(cols) - 2))
continue
cells = [str(cc), ARM_LABEL[arm]]
if has_input:
cells.append(maybe_bold(get(s, "input_throughput_tok_s"), "tps", "in"))
cells.append(maybe_bold(get(s, "output_throughput_tok_s"), "tps", "out"))
cells.append(maybe_bold(get(s, "ttft_s.p95"), "s", "ttft"))
cells.append(maybe_bold(get(s, "tpot_s.p95"), "ms", "tpot"))
lines.append("| " + " | ".join(cells) + " |")
return "\n".join(lines)
def main():
tp_data = load(sys.argv[1] if len(sys.argv) > 1 else None)
e7_data = load(sys.argv[2] if len(sys.argv) > 2 else None)
data = {**tp_data, **e7_data}
for prefix, title, ccs, has_input in SCENARIOS:
print(f"\n### {title}\n")
print(table_for(prefix, title, ccs, has_input, data))
# appendix: full metrics
print("\n\n## 附录:全量指标(含 mean/p50/max、回退、投机接受长度\n")
print("| 场景点 | 方案 | ok/nreq | wall(s) | out tok/s | in tok/s | TTFT mean/p95/max(s) | TPOT mean/p95/max(ms) | retractions | accept |")
print("|---|---|---|---|---|---|---|---|---|---|")
for prefix, _, ccs, _ in SCENARIOS:
for cc in ccs:
for arm in ARMS:
s = data.get((f"{prefix}_c{cc}", arm))
if s is None:
continue
ttft = s.get("ttft_s") or {}
tpot = s.get("tpot_s") or {}
def trio(d):
m, p, x = d.get("mean"), d.get("p95"), d.get("max")
if m is None:
return ""
return f"{m:.2f}/{p:.2f}/{x:.2f}"
def trioms(d):
m, p, x = d.get("mean"), d.get("p95"), d.get("max")
if m is None:
return ""
return f"{m*1000:.1f}/{p*1000:.1f}/{x*1000:.1f}"
print(f"| {prefix}_c{cc} | {ARM_LABEL[arm]} | {s.get('ok')}/{s.get('num_requests')} "
f"| {s.get('wall_s')} | {s.get('output_throughput_tok_s')} | {s.get('input_throughput_tok_s')} "
f"| {trio(ttft)} | {trioms(tpot)} | {s.get('retractions_total')} | {s.get('spec_accept_length_mean')} |")
if __name__ == "__main__":
main()

View File

@ -0,0 +1,168 @@
#!/bin/bash
# run_b300_matrix.sh <arm: tp2pp4|e7b> — B300-equivalent cold-cache scenario matrix on 60.8
#
# Mirrors the B300 report scenario set (16K->512, 1K->128, 1K->4K, 64K->512,
# 128K->512, 256K/512K/896K boundary OSL=1) at reduced concurrency per team
# decision (16K/1K capped at cc64; 64K/128K at cc<=8; boundary cc<=3).
# Both arms run the same grid; E7b structurally skips 256K cc>1 / 512K / 896K
# (KV pool 276,864, ctx 270,336).
#
# Protocol: cold points (shared-frac 0), recycled corpus windows via
# --pool-override (corpus has no virgin text left), per-point idle-wait ->
# scenario-length prewarm (first point of each scenario) -> flush_cache ->
# bench -> hit-rate-from-logs verification (<=0.01 else one retry, then abort).
# run-ids 95xx are labels only (windows are explicit).
#
# Usage: nohup bash /root/run_b300_matrix.sh <arm> \
# > /root/bench_logs/b300eq_<arm>_progress.log 2>&1 &
set -u
ARM=${1:?usage: run_b300_matrix.sh tp2pp4|e7b}
case $ARM in
tp2pp4) CONTAINER=glm53-pp4 ;;
e7b) CONTAINER=glm53-nvfp4 ;;
*) echo "bad arm: $ARM"; exit 1 ;;
esac
CORPUS=/root/corpus_ids.json
BENCH=/root/bench_corpus_v2.py
EXTRACT=/root/extract_summary.py
URL=http://127.0.0.1:30000
STAMP=$(date +%Y%m%d_%H%M)
LOG=/root/bench_logs/b300eq_${ARM}_${STAMP}
mkdir -p "$LOG"
RESULTS="$LOG/all_results.jsonl"
STATUS="$LOG/status.txt"
: > "$STATUS"
echo "=== [$ARM] matrix start $(date) LOGDIR=$LOG ==="
# ---- server facts snapshot ----
alive() { [ -n "$(docker ps --filter name=$CONTAINER --filter status=running -q)" ]; }
alive || { echo "=== [$ARM] ABORT: container $CONTAINER not running ==="; exit 1; }
docker inspect "$CONTAINER" --format '{{.Config.Cmd}}' > "$LOG/server_cmd.txt" 2>&1
docker logs "$CONTAINER" 2>&1 | grep -E "max_total_num_tokens|KV Cache is allocated|context_len|chunked_prefill_size|max_running_request|speculative_num_steps" | head -20 > "$LOG/server_facts.txt" 2>&1
nvidia-smi --query-gpu=index,name,memory.total,memory.used --format=csv > "$LOG/gpu_inventory_idle.csv" 2>&1
md5sum "$CORPUS" "$BENCH" "$EXTRACT" > "$LOG/md5_assets.txt" 2>&1
echo "--- server_cmd: $(cat "$LOG/server_cmd.txt")"
echo "--- server_facts:"; cat "$LOG/server_facts.txt"
# ---- per-rank VRAM sampler (whole-arm timeline, 30s cadence) ----
( while true; do
echo "# $(date +%s) $(date +%T)"
nvidia-smi --query-gpu=index,memory.used,utilization.gpu --format=csv,noheader
sleep 30
done ) > "$LOG/vram_timeline.csv" 2>&1 &
SAMPLER=$!
trap 'kill $SAMPLER 2>/dev/null' EXIT
idle_wait() {
for i in $(seq 1 90); do
local last
last=$(docker logs --since 90s "$CONTAINER" 2>&1 | grep 'running-req' | tail -1)
if [ -z "$last" ] || echo "$last" | grep -q 'running-req: 0'; then return 0; fi
sleep 10
done
echo "[warn] idle_wait timeout, continuing"
}
flush() { curl -s -m 60 -X POST $URL/flush_cache >/dev/null; sleep 3; }
prewarm() { # $1=input_len $2=base — one uncounted request at scenario length
timeout 1800 python3 "$BENCH" --corpus "$CORPUS" --input-len "$1" --output-len 16 \
--shared-frac 0 --concurrency 1 --num-requests 1 --run-id 9599 \
--pool-override "$2" --url $URL/generate --container "$CONTAINER" \
> "$LOG/prewarm_$1.log" 2>&1 || true
}
bench_call() { # $1=log $2=tag $3=isl $4=osl $5=cc $6=nreq $7=base $8=rid
timeout 7200 python3 "$BENCH" --corpus "$CORPUS" --input-len "$3" --output-len "$4" \
--shared-frac 0 --concurrency "$5" --num-requests "$6" --run-id "$8" \
--pool-override "$7" --url $URL/generate --container "$CONTAINER" \
--dump-records "$LOG/${2}_records.jsonl" > "$1" 2>&1
}
run_point() { # tag isl osl cc nreq base rid [prewarm=1]
local TAG=$1 ISL=$2 OSL=$3 CC=$4 NREQ=$5 BASE=$6 RID=$7 PW=${8:-0}
alive || { echo "$TAG CONTAINER_DEAD" >> "$STATUS"; echo "=== $TAG ABORT: container dead ==="; exit 1; }
echo "=== [$ARM] $TAG isl=$ISL osl=$OSL cc=$CC nreq=$NREQ base=$BASE start $(date +%T) ==="
idle_wait
[ "$PW" = "1" ] && prewarm "$ISL" "$BASE"
flush
local rc=0
bench_call "$LOG/${TAG}.log" "$TAG" "$ISL" "$OSL" "$CC" "$NREQ" "$BASE" "$RID" || rc=$?
if [ "$rc" = "0" ]; then
python3 "$EXTRACT" "$LOG/${TAG}.log" "$TAG" "$ARM" "$RESULTS"; rc=$?
fi
if [ "$rc" = "4" ]; then
echo "=== $TAG hit_rate>0.01, retrying once ==="
idle_wait; flush
bench_call "$LOG/${TAG}_retry.log" "${TAG}_retry" "$ISL" "$OSL" "$CC" "$NREQ" "$BASE" "$RID" || rc=$?
if [ "$rc" = "0" ]; then
python3 "$EXTRACT" "$LOG/${TAG}_retry.log" "${TAG}_retry" "$ARM" "$RESULTS"; rc=$?
fi
if [ "$rc" = "4" ]; then
echo "$TAG HIT_FAIL_FINAL" >> "$STATUS"
echo "=== $TAG hit_rate still >0.01 after retry — ABORTING ARM (contamination) ==="
exit 1
fi
fi
if [ "$rc" = "0" ]; then
echo "$TAG OK" >> "$STATUS"
else
echo "$TAG BENCH_FAIL rc=$rc" >> "$STATUS"
echo "=== $TAG bench rc=$rc (recorded, continuing) ==="
fi
echo "=== $TAG done $(date +%T) ==="
}
# ---- recycled corpus window bases (see campaign window map; all within 21.3M) ----
B_16K=2300000; B_1K=4500000; B_1K4=4700000; B_64K=5000000
B_128K=6200000; B_256K=8400000; B_512K=9300000; B_896K=9900000
# ===== B300 §3: main scenario 16K -> 512 =====
run_point b3_16k_c1 16384 512 1 8 $B_16K 9501 1
run_point b3_16k_c8 16384 512 8 16 $B_16K 9502
run_point b3_16k_c16 16384 512 16 32 $B_16K 9503
run_point b3_16k_c32 16384 512 32 64 $B_16K 9504
run_point b3_16k_c64 16384 512 64 128 $B_16K 9505
# ===== B300 §4.1: short input 1K -> 128 =====
run_point b41_1k_c1 1024 128 1 8 $B_1K 9511 1
run_point b41_1k_c8 1024 128 8 16 $B_1K 9512
run_point b41_1k_c32 1024 128 32 64 $B_1K 9513
run_point b41_1k_c64 1024 128 64 128 $B_1K 9514
# ===== B300 §4.2: long output 1K -> 4K =====
run_point b42_1k4k_c1 1024 4096 1 8 $B_1K4 9521 1
run_point b42_1k4k_c8 1024 4096 8 16 $B_1K4 9522
run_point b42_1k4k_c32 1024 4096 32 64 $B_1K4 9523
run_point b42_1k4k_c64 1024 4096 64 128 $B_1K4 9524
# ===== B300 §5.1: long context 64K -> 512 =====
run_point b51_64k_c1 65536 512 1 8 $B_64K 9531 1
run_point b51_64k_c4 65536 512 4 8 $B_64K 9532
run_point b51_64k_c8 65536 512 8 16 $B_64K 9533
# ===== B300 §5.1: long context 128K -> 512 =====
run_point b51_128k_c1 131072 512 1 8 $B_128K 9541 1
run_point b51_128k_c2 131072 512 2 8 $B_128K 9542
run_point b51_128k_c4 131072 512 4 8 $B_128K 9543
# ===== B300 §5.2: context boundary, OSL=1 (nreq=cc) =====
run_point b52_256k_c1 262144 1 1 1 $B_256K 9551 1
if [ "$ARM" = "tp2pp4" ]; then
run_point b52_256k_c2 262144 1 2 2 $B_256K 9552
run_point b52_256k_c3 262144 1 3 3 $B_256K 9553
run_point b52_512k_c1 524288 1 1 1 $B_512K 9554
run_point b52_896k_c1 917504 1 1 1 $B_896K 9555
else
# E7b structural limits: pool 276,864 (256K cc>=2 needs 524K) and ctx 270,336 (<512K)
for t in b52_256k_c2 b52_256k_c3 b52_512k_c1 b52_896k_c1; do
echo "$t SKIP_STRUCTURAL" >> "$STATUS"
done
echo "=== [e7b] 256K cc2/3, 512K, 896K skipped: structural (pool 276,864 / ctx 270,336) ==="
fi
nvidia-smi --query-gpu=index,name,memory.total,memory.used --format=csv > "$LOG/gpu_inventory_final.csv" 2>&1
echo "=== [$ARM] MATRIX ALL DONE $(date)$LOG ==="
echo "--- status:"; cat "$STATUS"