b300-equivalent matrix: dual-plan (TP2PP4-D vs E7b) full B300 scenario replication on 60.8 - 41 valid points, divide at C=8, E7b usable window <=C8 (MRR16+graph-drop), boundary 256K/512K/896K TP2PP4-only, 4-5x absolute gap vs B300 narrowing to ~2x at boundary prefill; hicache host-layer cold-cache pitfall documented; in-service container preserved-renamed-restored and verified (09-10)
This commit is contained in:
parent
e3476aff86
commit
21fcca5d16
@ -1,4 +1,4 @@
|
|||||||
# 现役部署状态页(live 核验于 2026-09-09)
|
# 现役部署状态页(live 核验于 2026-09-10,60.8 当日核验;其余机器 09-09 口径)
|
||||||
|
|
||||||
> 本页记录 174.1.60.x 集群"现在跑的是什么"。改动机器前后先读这页;
|
> 本页记录 174.1.60.x 集群"现在跑的是什么"。改动机器前后先读这页;
|
||||||
> **任何变动(停/起/改配置)完成后必须更新本页**。核验命令:`docker ps` + `nvidia-smi`。
|
> **任何变动(停/起/改配置)完成后必须更新本页**。核验命令:`docker ps` + `nvidia-smi`。
|
||||||
@ -14,7 +14,7 @@
|
|||||||
| 60.5 | `glm53-nvfp4`(Up 2d,09-09 只读核验) | **NVFP4 团队生产**(deploy_glm53_605.sh,md5 fcd9109b)。生产机铁律:不实验、不重启、不覆盖脚本。**交付升级路径已更新为 v3(09-09 hit90 场景优胜 TP4PP2-hicache@0.90,池 647,040/c4 并发独立文档/hit90 cc8 out +34% vs 现役):脚本 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/scripts/deploy_glm53_605_v3.sh` + 补丁束(60.8:/root/glm53_r37_patch_bundle_v3.tar.gz,md5 6922e534,需 scp 至 60.5)+ 需 docker pull cu13 镜像;未执行、未落 60.5 磁盘(60.5:/root 仅有原脚本,核验过)。此前 v2(TP2PP4-hicache,冷缓存口径优胜)被 v3 取代,仍留仓库可作"容量优先 6 条文档/512k 单条最快"备选;回滚=原脚本** | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(方案 A 口径);待切 `glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`(容量优先备选 `..._tp2pp4_hicache.env`) |
|
| 60.5 | `glm53-nvfp4`(Up 2d,09-09 只读核验) | **NVFP4 团队生产**(deploy_glm53_605.sh,md5 fcd9109b)。生产机铁律:不实验、不重启、不覆盖脚本。**交付升级路径已更新为 v3(09-09 hit90 场景优胜 TP4PP2-hicache@0.90,池 647,040/c4 并发独立文档/hit90 cc8 out +34% vs 现役):脚本 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/scripts/deploy_glm53_605_v3.sh` + 补丁束(60.8:/root/glm53_r37_patch_bundle_v3.tar.gz,md5 6922e534,需 scp 至 60.5)+ 需 docker pull cu13 镜像;未执行、未落 60.5 磁盘(60.5:/root 仅有原脚本,核验过)。此前 v2(TP2PP4-hicache,冷缓存口径优胜)被 v3 取代,仍留仓库可作"容量优先 6 条文档/512k 单条最快"备选;回滚=原脚本** | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(方案 A 口径);待切 `glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`(容量优先备选 `..._tp2pp4_hicache.env`) |
|
||||||
| 60.6 | 无容器,但 8 卡被外部裸金属实验占用(`/data/hzy/sparse-opd-*`,09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — |
|
| 60.6 | 无容器,但 8 卡被外部裸金属实验占用(`/data/hzy/sparse-opd-*`,09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — |
|
||||||
| 60.7 | 基本空(4 卡仍有 `/home/user/dirA_exp` 外部小任务,09-08 晚实测) | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — |
|
| 60.7 | 基本空(4 卡仍有 `/home/user/dirA_exp` 外部小任务,09-08 晚实测) | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — |
|
||||||
| 60.8 | `glm53-nvfp4`(Up,:30000,restart=unless-stopped) | **TP4PP2-nomtp@0.90 hit90 优胜口径在役**(09-09 晚接替同日早间的 TP2PP4@0.85:同 r37 栈 cu13+9 挂载+de-GLOO,仅改 tp4/pp2 + memfrac 0.90;KV 池 **647,040**;实测实例原样保留切 restart 策略,未重建容器。启动 `bash /root/deploy_ppmtp_r37.sh '--tp 4 --pp-size 2 --disable-overlap-schedule --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length 1048576' nomtp 8192 0.90 1 1 0 0`)。hit90(90% 命中 i128k/o512)out cc1/2/3/4/8/16 = **28.8/47.5/63.9/76.2/106.4/125.7**(vs TP2PP4 基线 cc4 +14%/cc8 +15%);并发独立 128k 文档 **4 条**(cap cc4 98.5 零排队)、512k 单条 ✓ TTFT 135.8s、质量门 **7/7**;同轮判决:DP attention 容量负收益判死、DCP 对 DSA 静默算错禁用、MTP@128k accept 2.07 判负。让步项(知情):容量 4 条(TP2PP4 为 6)、512k 单条慢 33%。回滚 TP2PP4 = 同命令改 `'--tp 2 --pp-size 4'` + memfrac 0.85 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`;实验全量 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/`;前史 `experiments/pro6000/glm53_nvfp4_128k_capacity_topology/`(TP2PP4 口径)+ r37 `experiments/pro6000/glm53_ppmtp_r37_verify_graph/` |
|
| 60.8 | `glm53-nvfp4`(Up,:30000,restart=unless-stopped) | **TP4PP2-nomtp@0.90 hit90 优胜口径在役**(09-09 晚接替同日早间的 TP2PP4@0.85:同 r37 栈 cu13+9 挂载+de-GLOO,仅改 tp4/pp2 + memfrac 0.90;KV 池 **647,040**;实测实例原样保留切 restart 策略,未重建容器。启动 `bash /root/deploy_ppmtp_r37.sh '--tp 4 --pp-size 2 --disable-overlap-schedule --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length 1048576' nomtp 8192 0.90 1 1 0 0`)。hit90(90% 命中 i128k/o512)out cc1/2/3/4/8/16 = **28.8/47.5/63.9/76.2/106.4/125.7**(vs TP2PP4 基线 cc4 +14%/cc8 +15%);并发独立 128k 文档 **4 条**(cap cc4 98.5 零排队)、512k 单条 ✓ TTFT 135.8s、质量门 **7/7**;同轮判决:DP attention 容量负收益判死、DCP 对 DSA 静默算错禁用、MTP@128k accept 2.07 判负。让步项(知情):容量 4 条(TP2PP4 为 6)、512k 单条慢 33%。回滚 TP2PP4 = 同命令改 `'--tp 2 --pp-size 4'` + memfrac 0.85。**09-10 B300 对标战役**:停役(容器 rename 保全 `glm53-nvfp4-insvc`)→ 双臂 B300 场景矩阵(TP2PP4-D 口径 24 点 + E7b 配方 16+1 点,判决=分界 C8/E7b 窗口≤C8/边界仅 TP2PP4 可达/与 B300 绝对差 4-5×,报告飞书 wiki `A7V3wZTQeifCB4krdi6cA834nW9`)→ **原容器恢复并核验**(rename 回 + start,health 200、16K 抽测 ok、显存水位与停役前一致,口径未变) | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp4pp2_hicache.env`;实验全量 `experiments/pro6000/glm53_nvfp4_hit90_dp_dcp_bench/`;B300 对标 `experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/`;前史 `experiments/pro6000/glm53_nvfp4_128k_capacity_topology/`(TP2PP4 口径)+ r37 `experiments/pro6000/glm53_ppmtp_r37_verify_graph/` |
|
||||||
|
|
||||||
## 方案 A-F 一览(GLM-5.3-NVFP4 @ pro6000,2026-09-08 双场景报告口径)
|
## 方案 A-F 一览(GLM-5.3-NVFP4 @ pro6000,2026-09-08 双场景报告口径)
|
||||||
|
|
||||||
|
|||||||
@ -0,0 +1,68 @@
|
|||||||
|
# GLM-5.3-NVFP4 双方案 B300 对标场景矩阵压测 — 60.8(6000D)
|
||||||
|
|
||||||
|
日期:2026-09-10 | 机器:174.1.60.8(8×RTX 6000D,96GB GDDR7,无 NVLink)
|
||||||
|
模型:GLM-5.3-NVFP4(modelopt)| 镜像:`nightly-dev-20260828-daf63171`(两臂同)
|
||||||
|
对标基线:飞书《GLM 5.3 | SGLang | Low-latency & High-Throughput 测试结果》(B300 报告,wiki UPB2w4Y5yi65qwkMxJJcZko5nUc)
|
||||||
|
完整报告:本目录 `REPORT.md`(= 飞书发布版 A7V3wZTQeifCB4krdi6cA834nW9)
|
||||||
|
|
||||||
|
## 目标
|
||||||
|
|
||||||
|
在 6000D 上复刻 B300 报告的全部场景(主场景 16K→512、4.1 短输入、4.2 长输出、5.1 长上下文、5.2 边界),对两套在役部署方案各跑一遍完整矩阵,产出对齐 B300 8 章结构的对标报告。测后 60.8 在役服务(TP4PP2@0.90)原容器恢复(已验证:health 200 + 16K 抽测 ok + 显存水位一致)。
|
||||||
|
|
||||||
|
## 实验臂
|
||||||
|
|
||||||
|
| 臂 | 方案 | 关键配置 | 质量门 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| tp2pp4 | **D 生产口径**(deploy_glm53_pp4.sh) | TP2PP4、mem0.85、MRR48、cps16384、radix 关、KV fp8_e4m3 池 1,040,384、无投机、index_topk_freq=4(=原生默认,恒等)、ctx 1,048,576 | 6/7(仅 tool-call:无 parser,历史已知) |
|
||||||
|
| e7b | **TP8+EAGLE3+AR**(deploy_glm53_607_exp.sh + CAR 补丁注入) | TP8、EAGLE 4/1/5、mem0.90、MRR16、cps8192、radix 开+hicache×3、KV fp8_e4m3 GPU 池 276,480、decode 图 bs1-8、ctx 270,336、custom-AR 1stage(`SGLANG_CUSTOM_ALLREDUCE_ALGO=1stage`,8 rank `SSKJ_CAR_PATCH_ACTIVE` 验证) | 7/7 |
|
||||||
|
|
||||||
|
## 场景矩阵与并发档位(用户裁决收敛:16K 封 64、64K/128K 封 8/4)
|
||||||
|
|
||||||
|
| B300 章节 | 场景 | 并发档位 | TP2PP4 活跃上限 | E7b 活跃上限 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| §3 主场景 | 16K→512 | 1/8/16/32/64 | 48(MRR) | 16(MRR=池贴边) |
|
||||||
|
| §4.1 | 1K→128 | 1/8/32/64 | 48 | 16 |
|
||||||
|
| §4.2 | 1K→4K | 1/8/32/64 | 48 | 16(c64 用户中止 rc=143) |
|
||||||
|
| §5.1 | 64K→512 | 1/4/8 | 15(池) | 4(池) |
|
||||||
|
| §5.1 | 128K→512 | 1/2/4 | 7(池) | 2(池) |
|
||||||
|
| §5.2 | 256K→1 | 1/2/3 | 3(池) | 1(池贴边) |
|
||||||
|
| §5.2 | 512K→1 | 1 | 1 | 结构性不可(ctx) |
|
||||||
|
| §5.2 | 896K→1(代 B300"约1M") | 1 | 1 | 结构性不可(ctx) |
|
||||||
|
|
||||||
|
测量协议:冷缓存(shared-frac 0)+ 每点 flush + 服务端命中核验 ≤0.01(超限重试一次);nreq=max(8, 2×cc);P95 nearest-rank 对齐 B300;语料耗尽(21.23M/21.30M)下用回收窗口(`--pool-override` 基址映射见 REPORT 附录 A)。有效测量点 41 个(tp2pp4 24 + e7b 16 + e7b 256K C=1 全新文本重测),全点 0 retraction、0 OOM。
|
||||||
|
|
||||||
|
## 判决速览(详见 REPORT.md)
|
||||||
|
|
||||||
|
- **分界 C=8**:E7b 全场景 C=1 占优(16K out 49.6 vs 16.7=3.0×、TPOT 13.7 vs 50.3ms=3.7×;1K→4K out 135 vs 20.3=6.7×);C≥8 TP2PP4 全指标反超并随并发拉大(16K c64 out 211 vs 84=2.5×)。
|
||||||
|
- **E7b 可用窗口 ≤C8**:MRR16 + decode 掉图(bs>8)双击,16K c16 TPOT 298ms 断崖。
|
||||||
|
- **decode 密集甜点 = E7b c8**:1K→4K out 412 tok/s、TPOT 22.5ms、accept 4.0。
|
||||||
|
- **TP2PP4 甜点 c16+**:16K 近线性至 c64(MRR48 未饱和);全场最高输出 4.2 c64 482 tok/s(但 TTFT 338s,仅离线)。
|
||||||
|
- **边界只有 TP2PP4 可达**:256K/512K/896K 全测(input 7,050/5,662/4,153 tok/s);E7b ctx 270,336 结构性封顶。
|
||||||
|
- **DSA 复现**:C=1 TPOT 对上下文不敏感(50.3/50.0/49.7ms @16/64/128K),并发才是驱动(c8: 70.9→161.8ms)。
|
||||||
|
- **vs B300**:定性结构完全复现(LL/HT 分野一致),绝对差 4-5×,边界 prefill 差距收窄至 ~2×;分界点本机更靠前(C8 vs C64-128),原因是容量上限(MRR/池)而非算力。
|
||||||
|
|
||||||
|
## 关键坑位(复测必读)
|
||||||
|
|
||||||
|
1. **hicache 宿主层陷阱**:256K prewarm KV 占池 94.8% 触发宿主层下放,`flush_cache` 清不掉宿主层 → 同文本测量命中 0.9998。冷缓存复测**必须换该实例从未发过的文本**(本战役 E7b 256K C=1 用窗口 9,900,000 重测达标)。
|
||||||
|
2. 语料已耗尽:回收窗口复用仅在 flush+命中核验协议下有效;TP2PP4 臂 radix 本来就关,零污染。
|
||||||
|
3. 在役保全流程:`docker stop` → `docker rename glm53-nvfp4 glm53-nvfp4-insvc`(必须先改名,E7b 部署脚本会 rm -f 同名容器)→ 测毕 `rename` 回 + `start`。docker stop/rm 偶发 "zombie PID" 报错是收尾边界现象,容器终态 exited(137)、显存归零,稍等重试即可。
|
||||||
|
|
||||||
|
## 资产与 md5 台账(60.8 执行件 = 本目录 = 60.7 原件 三方一致)
|
||||||
|
|
||||||
|
```
|
||||||
|
1e34dd8d813ccb7ac8243e8d1573a49e bench_corpus_v2.py (scripts/)
|
||||||
|
4c126d067d33b5ea27c268f37561634c extract_summary.py (scripts/)
|
||||||
|
5892b44610b2ce61f533f1721105625e run_b300_matrix.sh (scripts/)
|
||||||
|
def3c64c5dc19e1d507080e3366c3762 deploy_glm53_pp4.sh (见 dual_scenario_bench/scripts/,md5 对照一致)
|
||||||
|
21db641f1d997e26fcf1ddc97163b87e deploy_glm53_607_exp.sh (见 dual_scenario_bench/scripts/,md5 对照一致)
|
||||||
|
a8fc9a504c041acc40831a848b4985d8 patches/custom_all_reduce.py
|
||||||
|
65a4d22bd87ae343e7d68c4879c14f98 patches/custom_all_reduce_utils.py
|
||||||
|
```
|
||||||
|
|
||||||
|
gen_report_tables.py 为本地表格生成器(md5 未入台账,60.8 侧执行件同源)。
|
||||||
|
|
||||||
|
## 原始数据
|
||||||
|
|
||||||
|
- 本目录 `results/{tp2pp4,e7b}/`:all_results.jsonl(逐点 SUMMARY + 命中核验)、status.txt、server_facts.txt(启动参数+池分配日志摘录)、gpu_inventory_idle/final.csv、vram_timeline.csv(30s 采样全矩阵)
|
||||||
|
- 60.8 侧:`/root/bench_logs/b300eq_tp2pp4_20260910_1138/`、`/root/bench_logs/b300eq_e7b_20260910_1428/`
|
||||||
|
- E7b 256K C=1:all_results.jsonl 同 tag 共 4 条,最后一条为干净重测值(生成器 dict 载入后写覆盖,天然生效)
|
||||||
280
experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/REPORT.md
Normal file
280
experiments/pro6000/glm53_nvfp4_b300_equivalent_matrix/REPORT.md
Normal file
@ -0,0 +1,280 @@
|
|||||||
|
# GLM-5.3-NVFP4 | RTX 6000D | SGLang 双方案 B300 对标场景压测报告
|
||||||
|
|
||||||
|
- 测试日期:2026-09-10(单日单机完成两臂)
|
||||||
|
- 测试机:174.1.60.8(6000D,8 卡)
|
||||||
|
- 对标基线:飞书《GLM 5.3 | SGLang | Low-latency & High-Throughput 测试结果》(B300 报告,wiki UPB2w4Y5yi65qwkMxJJcZko5nUc)
|
||||||
|
- 测后状态:60.8 在役服务(TP4PP2@0.90 口径)已原容器恢复并验证(health 200 + 16K 抽测 ok + 显存水位与停役前一致)
|
||||||
|
|
||||||
|
## 1. 结论摘要
|
||||||
|
|
||||||
|
本轮在单台 8 卡 RTX 6000D 上,用 GLM-5.3-**NVFP4** 完整复刻 B300 报告的场景矩阵,测试了两套在役部署方案。两套方案代表完整部署形态,不是单参数 A/B:**TP2PP4** 为 D 生产口径(吞吐/长上下文形态),**TP8+EAGLE3+AR** 为 E7b 配方(低延迟形态,含 custom allreduce 1stage 补丁)。
|
||||||
|
|
||||||
|
- **低并发优先 E7b(TP8+EAGLE3+AR)**:主场景 `16K→512, C=1` 输出 49.6 tok/s、TPOT 13.7 ms,对 TP2PP4(16.7 tok/s、50.3 ms)分别是 **3.0×** 与 **3.7×**;长输出 `1K→4K, C=1` 输出 135 tok/s、TPOT 8.7 ms,对 TP2PP4(20.3、49.8 ms)是 **6.7×** 与 **5.7×**。
|
||||||
|
- **高并发优先 TP2PP4**:主场景 C=64 达到 6,748 input tok/s / 211 output tok/s,对 E7b(2,697 / 84.3)均为 **2.5×**;短输入 C=32/64 输出 276/341 tok/s,对 E7b(100/103)为 **2.7~3.3×**。
|
||||||
|
- **E7b 的可用并发窗口比 B300 Low-Latency 窄一个数量级**:MRR=16 且 CUDA graph 仅覆盖 decode bs 1–8,并发 ≥8 即掉图,主场景 C=16 TPOT P95 跳到 297.8 ms(C=8 为 135.1 ms;同点 TP2PP4 仅 98.5 ms)。E7b 的生产甜点上限 = **C≤8**。
|
||||||
|
- **TP2PP4 甜点在 C=16 之后**:主场景输出吞吐从 C=8 的 101 近线性爬到 C=64 的 211 tok/s(MRR48 尚未饱和),4.2 长输出 C=64 达全场最高 482 tok/s(但 TTFT P95 338 s,需要排队预算)。
|
||||||
|
- **decode 密集低并发的最优解是 E7b C=8**:`1K→4K, C=8` 输出 412 tok/s、TPOT 22.5 ms,对 TP2PP4(101 tok/s、79.4 ms)为 4.1×;EAGLE 实测 accept length 4.0。
|
||||||
|
- **长上下文与容量边界只有 TP2PP4 可达**:128K C=1 两方案输出打平(11.8 vs 11.9 tok/s)但 TP2PP4 TTFT 减半(18.2 s vs 37.6 s);256K/512K/896K 边界 E7b 结构性不可测(ctx 270,336 封顶 + KV 池 276,480 贴边),TP2PP4 全部完成(256K C=1/2/3、512K/896K C=1)。
|
||||||
|
- **DSA 特性在 6000D 复现**:TP2PP4 C=1 的 TPOT 对上下文长度不敏感(16K/64K/128K = 50.3/50.0/49.7 ms 恒定),并发才是 TPOT 驱动因子(16K 行 C=8→C=64:70.9→203.7 ms)。
|
||||||
|
- **与 B300 的绝对差距约 4~5×**,边界 prefill 差距收窄到约 2×(256K C=1 input 7,050 vs 17,457;896K 4,153 vs 约1M 行 7,897)。硬件与量化口径不同(B300 报告未写明量化方式),绝对值仅量级可比,两份报告的结构性结论一致(见第 9 章)。
|
||||||
|
- 全部 41 个有效测量点 **0 回退(retraction)、0 OOM**,冷缓存命中核验全部 ≤0.01(E7b 256K C=1 首测触 hicache 宿主层陷阱,用全新文本重测达标,见 5.2 注记)。
|
||||||
|
|
||||||
|
## 2. 测试环境与配置
|
||||||
|
|
||||||
|
| 项目 | TP2PP4(D 生产口径) | TP8+EAGLE3+AR(E7b 配方) |
|
||||||
|
|-|-|-|
|
||||||
|
| 硬件 | 单机 8 × NVIDIA RTX 6000D(96 GB GDDR7,nvidia-smi 可见 85,651 MiB/卡) | 同左 |
|
||||||
|
| 模型 | GLM-5.3-NVFP4(modelopt 量化,/data/hf_models/GLM-5.3-NVFP4) | 同左 |
|
||||||
|
| 镜像 | `nightly-dev-20260828-daf63171` | 同左 |
|
||||||
|
| 并行 | TP2 × PP4 | TP8 |
|
||||||
|
| 投机解码 | 无 | EAGLE3,num_steps=4,topk=1,draft_tokens=5 |
|
||||||
|
| `mem-fraction-static` | 0.85 | 0.90 |
|
||||||
|
| 最大活跃请求(MRR) | 48 | 16 |
|
||||||
|
| Chunk Prefill | 16,384 | 8,192 |
|
||||||
|
| KV dtype | fp8_e4m3 | fp8_e4m3 |
|
||||||
|
| KV 池(服务端实测) | **1,040,384 tokens**(12.6~13.4 GB/rank,无宿主层) | **276,480 tokens GPU**(15.8 GB/rank)+ 分层缓存 hicache×3 宿主层(write_through) |
|
||||||
|
| radix cache | 关(`disable_radix_cache=True`) | 开(分层缓存) |
|
||||||
|
| 上下文上限 | 1,048,576(config 原生) | 270,336(显存约束下的部署值) |
|
||||||
|
| CUDA graph | 常规捕获 | decode 图 bs 1–8(bs>8 掉图) |
|
||||||
|
| custom allreduce | — | 1stage 补丁注入(`SGLANG_CUSTOM_ALLREDUCE_ALGO=1stage` 环境强制;8 rank `SSKJ_CAR_PATCH_ACTIVE` 日志验证全出现) |
|
||||||
|
| `index_topk_freq` | 4(override,等于原生默认,恒等) | 原生默认 4 |
|
||||||
|
| 质量门 | 6/7(仅 tool-call 失败:D 口径未配 parser,历史已知;其余全过) | **7/7** |
|
||||||
|
|
||||||
|
因此,下文比较回答的是"两种部署形态谁更适合该负载",不能把差异单独归因于 EAGLE、PP 流水、chunk、radix 或图覆盖中的某一项(与 B300 报告同款声明)。
|
||||||
|
|
||||||
|
**测量协议**(对齐 B300 口径):
|
||||||
|
|
||||||
|
- 冷缓存:`--shared-frac 0`,每点前 `POST /flush_cache`,服务端 Prefill 日志核算命中率,>0.01 重测一次,仍超停点排查;全矩阵命中核验最终全部达标。
|
||||||
|
- 指标:Input TPS / Output TPS / TTFT P95 / TPOT P95(P95 为 nearest-rank;Input TPS = 输入 token / 全程墙钟,与 B300 口径一致)。
|
||||||
|
- 负载:PG19 真实语料 token 切片(`corpus_ids.json`),input_ids 直打 `/generate`,temperature=0、ignore_eos、流式;nreq = max(8, 2×并发),边界行 nreq=并发。
|
||||||
|
- 并发档位按决策收敛:16K/1K 类封顶 64;64K 封 8、128K 封 4(128K C=2 补一档);超出活跃上限的档位是**排队观察点**(与 B300 C=256 同性质,保留为有效观察)。
|
||||||
|
- 语料已耗尽(21.23M/21.30M),冷缓存口径下用**回收窗口**复用(窗口基址见附录 A,逐记录 `corpus_window` 字段留档)。
|
||||||
|
|
||||||
|
**两臂活跃上限**(MRR 与 KV 池决定,解释各行哪些并发是排队观察点):
|
||||||
|
|
||||||
|
| 场景 | TP2PP4 活跃上限 | E7b 活跃上限 |
|
||||||
|
|-|-|-|
|
||||||
|
| 16K / 1K | 48(MRR) | 16(MRR,=池贴边) |
|
||||||
|
| 64K | 15(池) | 4(池) |
|
||||||
|
| 128K | 7(池) | 2(池) |
|
||||||
|
| 256K | 3(池) | 1(池 262K KV / 276K 贴边) |
|
||||||
|
| 512K / 896K | 1(池) | 结构性不可(ctx 270,336) |
|
||||||
|
|
||||||
|
## 3. 主场景:16K 输入、512 输出
|
||||||
|
|
||||||
|
B300 跑了 C=1/8/32/64/128/256;本机按 MRR 上限收敛为 C=1/8/16/32/64(C=16 为本机甜点档,B300 无此档;C=128/256 超出两臂 MRR)。
|
||||||
|
|
||||||
|
| 并发 | 方案 | Input TPS | Output TPS | TTFT P95 | TPOT P95 |
|
||||||
|
|-|-|-|-|-|-|
|
||||||
|
| 1 | TP2PP4 | 534 | 16.7 | 5.36 s | 50.3 ms |
|
||||||
|
| 1 | TP8+EAGLE3+AR | 1,587 | 49.6 | **3.98 s** | **13.7 ms** |
|
||||||
|
| 8 | TP2PP4 | 3,236 | **101** | **14.79 s** | **70.9 ms** |
|
||||||
|
| 8 | TP8+EAGLE3+AR | 2,605 | 81.4 | 32.52 s | 135.1 ms |
|
||||||
|
| 16 | TP2PP4 | 4,674 | **146** | **25.78 s** | **98.5 ms** |
|
||||||
|
| 16 | TP8+EAGLE3+AR | 2,217 | 69.3 | 60.86 s | 297.8 ms |
|
||||||
|
| 32 | TP2PP4 | **6,185** | **193** | **46.43 s** | **152.0 ms** |
|
||||||
|
| 32 | TP8+EAGLE3+AR | 2,527 | 79.0 | 169.68 s | 259.9 ms |
|
||||||
|
| 64 | TP2PP4 | **6,748** | **211** | **134.78 s** | **203.7 ms** |
|
||||||
|
| 64 | TP8+EAGLE3+AR | 2,697 | 84.3 | 345.69 s | 231.5 ms |
|
||||||
|
|
||||||
|
趋势:
|
||||||
|
|
||||||
|
- **分界在 C=8**:C=1 E7b 全指标占优;C=8 起 TP2PP4 全指标反超,且输出吞吐差距随并发拉大(101 vs 81 → 211 vs 84)。
|
||||||
|
- **E7b 在 C=8→16 输出吞吐倒退**(81.4→69.3 tok/s):MRR=16 开始排队 + decode 掉图(bs>8 无图)双击;TPOT P95 从 135 ms 跳到 298 ms。EAGLE accept length 随并发从 2.14 爬到 2.99,但被掉图抵消。
|
||||||
|
- **prefill 墙的差异**:E7b 的 input TPS 几乎不随并发增长(C=8→64:2,605→2,697,+3%),TP2PP4 翻倍(3,236→6,748,+108%)——chunk 8192 + TP8 无 PP 流水的 prefill 瓶颈 vs chunk 16384 + PP4 流水摊满。
|
||||||
|
- TP2PP4 到 C=64 仍在爬坡(C=32→64 +9%),MRR48 未饱和;TTFT P95 在 C=64 达 134.8 s,同 B300 一样高并发 TTFT 需要准入控制。
|
||||||
|
|
||||||
|
## 4. 短输入与长输出
|
||||||
|
|
||||||
|
### 4.1 `1K -> 128`
|
||||||
|
|
||||||
|
| 并发 | 方案 | Input TPS | Output TPS | TTFT P95 | TPOT P95 |
|
||||||
|
|-|-|-|-|-|-|
|
||||||
|
| 1 | TP2PP4 | 155 | 19.3 | 386 ms | 49.3 ms |
|
||||||
|
| 1 | TP8+EAGLE3+AR | **518** | **64.7** | **326 ms** | **14.3 ms** |
|
||||||
|
| 8 | TP2PP4 | 815 | 102 | **1.69 s** | 72.9 ms |
|
||||||
|
| 8 | TP8+EAGLE3+AR | **1,219** | **152** | 2.16 s | **53.7 ms** |
|
||||||
|
| 32 | TP2PP4 | **2,204** | **276** | **4.87 s** | **98.5 ms** |
|
||||||
|
| 32 | TP8+EAGLE3+AR | 801 | 100 | 25.60 s | 190.5 ms |
|
||||||
|
| 64 | TP2PP4 | **2,724** | **341** | **19.96 s** | **101.8 ms** |
|
||||||
|
| 64 | TP8+EAGLE3+AR | 826 | 103 | 65.81 s | 177.7 ms |
|
||||||
|
|
||||||
|
短输入下 E7b 在 C≤8 显著占优(C=8 输出 152 vs 102,1.5×),C=32 起 TP2PP4 大幅拉开(2.7~3.3×)。E7b 的 input TPS 反而在 C=8 最高(1,219)后回落——MRR16 排队开始挤占 prefill。B300 同场景 Low-Latency 到 C=128 才被反超,本机提前到 C=8~32 之间,同样是容量上限(MRR16/池)而非算力所致。
|
||||||
|
|
||||||
|
### 4.2 `1K -> 4K`
|
||||||
|
|
||||||
|
| 并发 | 方案 | Output TPS | TTFT P95 | TPOT P95 |
|
||||||
|
|-|-|-|-|-|
|
||||||
|
| 1 | TP2PP4 | 20.3 | 395 ms | 49.8 ms |
|
||||||
|
| 1 | TP8+EAGLE3+AR | **135** | **320 ms** | **8.7 ms** |
|
||||||
|
| 8 | TP2PP4 | 101 | 1.87 s | 79.4 ms |
|
||||||
|
| 8 | TP8+EAGLE3+AR | **412** | **1.68 s** | **22.5 ms** |
|
||||||
|
| 32 | TP2PP4 | **367** | **4.14 s** | **90.5 ms** |
|
||||||
|
| 32 | TP8+EAGLE3+AR | 252 | 304.75 s | 74.6 ms |
|
||||||
|
| 64 | TP2PP4 | **482** | 337.92 s | 95.5 ms |
|
||||||
|
| 64 | TP8+EAGLE3+AR | 用户中止*(见注) | — | — |
|
||||||
|
|
||||||
|
\* E7b C=64 点按用户指示中止("并发 64 太高",rc=143),未获得有效数据;同点 TP2PP4 已完成。E7b 该点 nreq=128 远超 MRR=16,属排队观察点,中止不影响结论完整性。
|
||||||
|
|
||||||
|
长输出放大了两形态的差异:
|
||||||
|
|
||||||
|
- **E7b C=1/C=8 是 decode 密集负载的最优区间**:C=1 输出 135 tok/s、TPOT 8.7 ms(全场最低),C=8 输出 412 tok/s(全场第二),EAGLE accept 长达 3.7~4.0——长输出让草稿模型进入"顺笔"状态,accept 显著高于 4.1 短输出行(2.0~2.1)。
|
||||||
|
- **E7b C=32 的 TTFT P95 304.75 s** 是纯排队(nreq=64 / MRR=16,4 波串行),其 TPOT 74.6 ms 与掉图后水平一致。
|
||||||
|
- **TP2PP4 C=64 输出 482 tok/s 为全场最高**,但 TTFT P95 338 s 意味着该点只适合离线批处理;交互负载应压在 C=32(367 tok/s、TTFT 4.1 s)。
|
||||||
|
|
||||||
|
## 5. 长上下文观察
|
||||||
|
|
||||||
|
### 5.1 `64K/128K -> 512`
|
||||||
|
|
||||||
|
(并发档位按用户指示收敛:64K 封 8、128K 封 4,另补 128K C=2;B300 同场景为 64K C=8/32/64、128K C=8/32。)
|
||||||
|
|
||||||
|
| 场景 | 并发 | 方案 | Input TPS | Output TPS | TTFT P95 | TPOT P95 |
|
||||||
|
|-|-|-|-|-|-|-|
|
||||||
|
| 64K→512 | 1 | TP2PP4 | 1,834 | 14.3 | **10.32 s** | 50.0 ms |
|
||||||
|
| 64K→512 | 1 | TP8+EAGLE3+AR | **2,877** | **22.5** | 17.17 s | **13.5 ms** |
|
||||||
|
| 64K→512 | 4 | TP2PP4 | **4,287** | **33.5** | **28.83 s** | **99.5 ms** |
|
||||||
|
| 64K→512 | 4 | TP8+EAGLE3+AR | 3,292 | 25.7 | 68.87 s | 154.6 ms |
|
||||||
|
| 64K→512 | 8 | TP2PP4 | **5,650** | **44.1** | **53.38 s** | **161.8 ms** |
|
||||||
|
| 64K→512 | 8 | TP8+EAGLE3+AR | 3,280 | 25.6 | 149.45 s | 157.5 ms |
|
||||||
|
| 128K→512 | 1 | TP2PP4 | 3,017 | 11.8 | **18.23 s** | 49.7 ms |
|
||||||
|
| 128K→512 | 1 | TP8+EAGLE3+AR | **3,054** | **11.9** | 37.60 s | **12.3 ms** |
|
||||||
|
| 128K→512 | 2 | TP2PP4 | **4,310** | **16.8** | **32.42 s** | **83.7 ms** |
|
||||||
|
| 128K→512 | 2 | TP8+EAGLE3+AR | 3,145 | 12.3 | 75.81 s | 162.2 ms |
|
||||||
|
| 128K→512 | 4 | TP2PP4 | **5,626** | **22.0** | **60.35 s** | **146.8 ms** |
|
||||||
|
| 128K→512 | 4 | TP8+EAGLE3+AR | 3,133 | 12.2 | 159.50 s | 163.4 ms |
|
||||||
|
|
||||||
|
- 64K C=1 E7b 仍占优(22.5 vs 14.3 tok/s),但 128K C=1 两方案输出打平(11.8 vs 11.9)——prefill 逐渐成为长上下文的主导成本,E7b 的 decode 优势被稀释;其 TTFT 反而慢 2×(37.6 vs 18.2 s)。
|
||||||
|
- C≥2 起 TP2PP4 全指标占优;E7b 的 output TPS 在 64K/128K 行几乎不随并发变化(22.5→25.6、11.9→12.2),与主场景同一形态:容量上限 + 掉图封死并发收益。
|
||||||
|
- DSA 的 TPOT 上下文不变性(TP2PP4 C=1:50.0 ms @64K ≈ 49.7 ms @128K ≈ 50.3 ms @16K)与并发驱动性(C=8:70.9 ms @16K → 161.8 ms @64K)在本组完整呈现。
|
||||||
|
|
||||||
|
### 5.2 上下文边界
|
||||||
|
|
||||||
|
(OSL=1,只验证容量与 prefill,不比较 Output TPS/TPOT——与 B300 同声明。B300 完成到约 1M;本机以 896K=917,504 tokens 对应 B300"约 1M"档。)
|
||||||
|
|
||||||
|
| 输入长度 | 方案 | 已完成并发 | C=1 Input TPS | 最高并发 TTFT P95 |
|
||||||
|
|-|-|-|-|-|
|
||||||
|
| 256K | TP2PP4 | 1/2/3 | **7,050** | 102.99 s(C=3) |
|
||||||
|
| 256K | TP8+EAGLE3+AR | 1 | 2,867 | 91.45 s(C=1) |
|
||||||
|
| 512K | TP2PP4 | 1 | **5,662** | 92.60 s(C=1) |
|
||||||
|
| 512K | TP8+EAGLE3+AR | 结构性不可(ctx 270,336) | — | — |
|
||||||
|
| 896K | TP2PP4 | 1 | **4,153** | 220.91 s(C=1) |
|
||||||
|
| 896K | TP8+EAGLE3+AR | 结构性不可(ctx 270,336) | — | — |
|
||||||
|
|
||||||
|
- 256K C=1 两臂相差 2.5×(7,050 vs 2,867 tok/s):TP2PP4 的 chunk 16384 + PP4 流水对超长 prefill 的摊满优势,在边界长度上比 128K 行(几乎打平)进一步放大;E7b 的 chunk 8192 代价随长度累积。
|
||||||
|
- TP2PP4 边界 input TPS 随长度衰减平缓(7,050 → 5,662 → 4,153),896K 单条 220.9 s 完成、池 1,040,384 tokens 单条可容(KV 917,504 + 余量)。
|
||||||
|
- **E7b 256K C=1 命中核验注记(hicache 宿主层陷阱)**:首测命中率 0.9998、重试仍超——根因是该点 nreq=1 测量文本与预热完全相同,256K 预热 KV(262,160 tokens)占池 94.8% 触发分层缓存宿主层下放,`flush_cache` 只清 GPU radix 树、清不掉宿主层。改用该服务实例从未发过的文本(窗口基址 9,900,000)无预热重测,命中 0.0,数据干净。此为分层缓存运维要点:**宿主层缓存不受 flush_cache 影响,冷测必须换文本**。
|
||||||
|
|
||||||
|
## 6. 显存状态
|
||||||
|
|
||||||
|
- **TP2PP4**:服务加载后空载 64.6 GiB/卡,矩阵峰值 **85.0 GiB/卡**(主场景 C=64 时逼近打满,最紧张卡余量约 0.6 GiB)。mem 0.85 下 KV 池按卡容量贴满分配,属预期;继续上调 MRR 或上下文没有余量,扩容前必须先降 mem-fraction。
|
||||||
|
- **TP8+EAGLE3+AR**:空载 77.9 GiB/卡(EAGLE 草稿权重 + mem 0.90 大池),矩阵峰值 **83.6 GiB/卡**(余量约 2.1 GiB)。
|
||||||
|
- 两臂全矩阵 **0 OOM、0 retraction**(全部 41 点 retractions_total=0)——B300 未披露该指标,本机在自身容量上限内运行无回退。
|
||||||
|
- 显存时间线逐 30 s 采样留档(vram_timeline.csv),可复核任一时刻的卡间分布。
|
||||||
|
|
||||||
|
## 7. 建议
|
||||||
|
|
||||||
|
1. **低并发交互/agent 长思考(C≤8)用 E7b**:主场景 C=1 TPOT 13.7 ms、长输出 C=8 输出 412 tok/s。生产并发上限建议钉在 ≤8:C=16 起 decode 掉图 + MRR16 排队使其全面劣于 TP2PP4。
|
||||||
|
2. **高并发吞吐/长上下文(C≥8 或输入 ≥64K)用 TP2PP4**:主场景 C=64 输出 211 tok/s、128K C=1 TTFT 18.2 s、896K 可达;MRR48 内未饱和,吞吐上限即 MRR。
|
||||||
|
3. **负载形态分界线**:prefill 吞吐需求 >2.2K tok/s 或并发 >8 → TP2PP4;decode 为主且并发 ≤8 → E7b。两臂在 C=8 附近的输出吞吐交叉(主场景 101 vs 81、短输入 102 vs 152、长输出 101 vs 412)——按输出长度分布选型,不能只看并发。
|
||||||
|
4. **E7b 扩窗口的两个前置**:MRR 16→更高需先扩 KV 池(hicache 宿主层只救命中场景,不增并发容量);decode 图覆盖 bs 8→16/32 才能消掉 C=16 的 298 ms TPOT 断崖。
|
||||||
|
5. **边界与超长上下文只有 TP2PP4 口径可服务**:E7b 若要对标 B300 512K/约1M 行,需要 ctx ≥524,288 与池 ≥52 万 tokens 的部署形态,本版(ctx 270,336 / 池 276,480)结构性不可达。
|
||||||
|
6. **不要把两臂差异单归因 EAGLE**:两臂同时差在并行拓扑、chunk、radix、MRR 与图覆盖;单变量消融未做(与 B300 报告建议 3 同款)。
|
||||||
|
7. **生产容量同时设吞吐和延迟 SLO**:TP2PP4 主场景 C=64 输出最高但 TTFT P95 已到 135 s;E7b C=32 长输出 TTFT P95 305 s。只看峰值 TPS 会掩盖排队长尾。
|
||||||
|
8. **分层缓存运维**:宿主层缓存不受 `flush_cache` 影响,任何冷缓存测量/复测必须更换输入文本(见 5.2 注记)。
|
||||||
|
|
||||||
|
## 8. 原始结果与复现
|
||||||
|
|
||||||
|
- 服务器原始结果(60.8):
|
||||||
|
- TP2PP4 臂:`/root/bench_logs/b300eq_tp2pp4_20260910_1138/`(all_results.jsonl 24 点、status.txt、server_facts.txt、gpu_inventory、vram_timeline.csv)
|
||||||
|
- E7b 臂:`/root/bench_logs/b300eq_e7b_20260910_1428/`(all_results.jsonl 20 条:16 点 OK + 4.2 C=64 用户中止 + 256K C=1 干净重测覆盖前 3 条污染记录;status.txt 含 4 个结构性跳过与 HIT_FAIL_FINAL 首测记录)
|
||||||
|
- 资产 md5 台账:`/root/bench_logs/b300eq_md5_ledger.txt`
|
||||||
|
- 本地镜像:`D:\sskj\b300eq\{tp2pp4,e7b}\`(上述全部文件)、`D:\sskj\b300eq\report_tables.md`(表格生成器输出)
|
||||||
|
- 部署脚本:`/root/deploy_glm53_pp4.sh`(md5 def3c64c…,与库内 sskj main 副本一致)、`/root/deploy_glm53_607_exp.sh`;CAR 补丁:`/root/patches/custom_all_reduce.py`(a8fc9a50…)+ `custom_all_reduce_utils.py`(65a4d22b…),三处(60.7 原件/本地/库内)md5 一致
|
||||||
|
- 测量工具:`/root/bench_corpus_v2.py`(md5 1e34dd8d…,p95 nearest-rank + 逐请求 dump)、`/root/extract_summary.py`(4c126d06…)、`/root/run_b300_matrix.sh`(5892b446…,矩阵驱动:alive/idle_wait/prewarm/flush/命中核验/重试/VRAM 采样)
|
||||||
|
- 复现命令(单点示例):
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# 冷缓存压测(E7b 256K C=1 干净版)
|
||||||
|
python3 /root/bench_corpus_v2.py --input-len 262144 --output-len 1 --shared-frac 0 \
|
||||||
|
--concurrency 1 --num-requests 1 --run-id 9551 --pool-override 9900000 \
|
||||||
|
--dump-records $L/b52_256k_c1_v2_records.jsonl
|
||||||
|
# 全矩阵:nohup bash /root/run_b300_matrix.sh <arm> > <progress.log> 2>&1 &
|
||||||
|
```
|
||||||
|
|
||||||
|
## 9. 与 B300 对比观察
|
||||||
|
|
||||||
|
> **口径声明**:B300 报告未写明模型量化方式(若为原始 BF16 权重,则与本机 NVFP4 非同模型形态);硬件为 8×B300(288 GB HBM3e)vs 本机 8×RTX 6000D(96 GB GDDR7);镜像 v0.5.18-cu130-dev4 vs nightly-20260828-daf63171。**绝对值仅量级可比,本对比只对结构性结论负责**。
|
||||||
|
|
||||||
|
- **定性结构完全复现**:低延迟配方(B300 Low-Latency = TP8+EAGLE vs 本机 E7b = TP8+EAGLE3+AR)在 C=1 占优、吞吐配方(B300 High-Throughput = DP8+DeepEP vs 本机 TP2PP4 = D 生产口径)在高并发占优——两套硬件上"低延迟 vs 高吞吐"的分野方向一致。
|
||||||
|
- **分界点本机更靠前**:B300 的交叉点在 C=64~128(主场景 HT C=128 反超 33%);本机在 C=8 附近。原因不是算力而是**容量上限**:本机两臂 MRR/池上限(48/16)远小于 B300 配方的 256/默认,先于算力撞墙。
|
||||||
|
- **绝对差距 4~5×(主场景)**:C=1 输出 246 vs 49.6 tok/s(5.0×)、input 7,882 vs 1,587(5.0×);吞吐侧峰值 997 vs 211(4.7×)、31,889 vs 6,748(4.7×)。与显存带宽硬件代差量级一致。
|
||||||
|
- **边界 prefill 差距收窄到 ~2×**:256K C=1 input 7,050 vs 17,457(2.5×)→ 512K 5,662 vs 12,421(2.2×)→ 896K/约1M 4,153 vs 7,897(1.9×)。计算密集的超长 prefill 是 6000D 相对最能打的位置(PP 流水摊满 + 带宽占比下降)。
|
||||||
|
- **TPOT 差距小于吞吐差距**:B300 LL C=1 4.36 ms vs E7b 13.7 ms(3.1×);高并发侧 B300 HT C=128 165 ms vs TP2PP4 C=64 204 ms(1.2×)——NVFP4 + DSA 把 decode 单步成本压得相对不差,差距主要在吞吐面。
|
||||||
|
- **饱和形态不同**:B300 LL 在 C=64 后进入 24K input tok/s 平台、HT 在 C=128 达峰后 C=256 回退 19%;本机 TP2PP4 到 C=64 仍在爬坡(MRR 未饱和),E7b 则被 MRR16+掉图封死在 C=8。本机没有一档出现吞吐回退——"甜点=并发上限"由 MRR 决定而非算力。
|
||||||
|
- **EAGLE 配方差异**:B300 LL 为 5 steps/6 draft tokens,本机 E7b 为 4 steps/topk1/5 draft tokens;本机实测 accept 2.0~4.0(短输出 2.0、主场景 2.1~3.0、长输出 3.7~4.0,随 decode 深入上升)。B300 未披露 accept,无法直接对比投机效率。
|
||||||
|
- **容量边界差距最大**:B300 两模式都完成约 1M 输入 C=1/2/4;本机仅 TP2PP4 可达 896K 且 C=1 单条(池 1,040,384 刚容一条),E7b 连 512K 都结构性不可测(ctx 270,336)。96 GB 卡上"上下文边界=显存边界"比 B300 严酷得多。
|
||||||
|
|
||||||
|
## 附录 A:语料窗口映射(回收窗口)
|
||||||
|
|
||||||
|
语料总量 21,296,780 tokens,此前场景一/二战役已消费至 21,235,008。冷缓存协议下回收复用:窗口基址 `--pool-override` 显式指定,每点窗口在基址上顺序推进(逐记录 `corpus_window.start/end` 留档),每点 flush + 命中核验 ≤0.01 保证冷。
|
||||||
|
|
||||||
|
| 场景 | 窗口基址 | 备注 |
|
||||||
|
|-|-|-|
|
||||||
|
| 主场景 16K→512 | 2,300,000 | |
|
||||||
|
| 4.1 短输入 1K→128 | 4,500,000 | |
|
||||||
|
| 4.2 长输出 1K→4K | 4,700,000 | |
|
||||||
|
| 5.1 64K→512 | 5,000,000 | |
|
||||||
|
| 5.1 128K→512 | 6,200,000 | |
|
||||||
|
| 5.2 256K→1 | 8,400,000 | E7b 干净重测改用 9,900,000(实例首用文本,避 hicache 宿主层残留) |
|
||||||
|
| 5.2 512K→1 | 9,300,000 | 仅 TP2PP4 |
|
||||||
|
| 5.2 896K→1 | 9,900,000 | 仅 TP2PP4 |
|
||||||
|
|
||||||
|
## 附录 B:全量指标(含 mean/p95/max、回退、投机接受长度)
|
||||||
|
|
||||||
|
| 场景点 | 方案 | ok/nreq | wall(s) | out tok/s | in tok/s | TTFT mean/p95/max(s) | TPOT mean/p95/max(ms) | retractions | accept |
|
||||||
|
|---|---|---|---|---|---|---|---|---|---|
|
||||||
|
| b3_16k_c1 | TP2PP4 | 8/8 | 245.55 | 16.68 | 533.8 | 5.35/5.36/5.36 | 49.6/50.3/50.3 | 0 | None |
|
||||||
|
| b3_16k_c1 | TP8+EAGLE3+AR | 8/8 | 82.58 | 49.6 | 1587.15 | 3.93/3.98/3.98 | 12.5/13.7/13.7 | 0 | 2.138 |
|
||||||
|
| b3_16k_c8 | TP2PP4 | 16/16 | 81.0 | 101.13 | 3236.22 | 10.05/14.79/14.79 | 59.5/70.9/70.9 | 0 | None |
|
||||||
|
| b3_16k_c8 | TP8+EAGLE3+AR | 16/16 | 100.63 | 81.41 | 2605.13 | 11.75/32.52/32.52 | 71.6/135.1/135.1 | 0 | 2.375 |
|
||||||
|
| b3_16k_c16 | TP2PP4 | 32/32 | 112.16 | 146.08 | 4674.5 | 15.53/25.78/25.84 | 79.2/98.5/101.2 | 0 | None |
|
||||||
|
| b3_16k_c16 | TP8+EAGLE3+AR | 32/32 | 236.44 | 69.29 | 2217.4 | 23.60/60.86/90.39 | 173.3/297.8/305.5 | 0 | 2.517 |
|
||||||
|
| b3_16k_c32 | TP2PP4 | 64/64 | 169.54 | 193.27 | 6184.73 | 26.52/46.43/47.89 | 113.7/152.0/157.3 | 0 | None |
|
||||||
|
| b3_16k_c32 | TP8+EAGLE3+AR | 64/64 | 414.96 | 78.97 | 2526.91 | 100.26/169.68/190.73 | 169.8/259.9/319.1 | 0 | 2.767 |
|
||||||
|
| b3_16k_c64 | TP2PP4 | 128/128 | 310.79 | 210.87 | 6747.91 | 63.05/134.78/139.99 | 139.2/203.7/214.3 | 0 | None |
|
||||||
|
| b3_16k_c64 | TP8+EAGLE3+AR | 128/128 | 777.55 | 84.29 | 2697.14 | 240.10/345.69/380.96 | 168.0/231.5/311.9 | 0 | 2.986 |
|
||||||
|
| b41_1k_c1 | TP2PP4 | 8/8 | 52.95 | 19.34 | 154.7 | 0.38/0.39/0.39 | 49.1/49.3/49.3 | 0 | None |
|
||||||
|
| b41_1k_c1 | TP8+EAGLE3+AR | 8/8 | 15.82 | 64.74 | 517.93 | 0.30/0.33/0.33 | 13.2/14.3/14.3 | 0 | 2.028 |
|
||||||
|
| b41_1k_c8 | TP2PP4 | 16/16 | 20.1 | 101.91 | 815.3 | 1.31/1.69/1.69 | 68.6/72.9/72.9 | 0 | None |
|
||||||
|
| b41_1k_c8 | TP8+EAGLE3+AR | 16/16 | 13.44 | 152.4 | 1219.24 | 1.21/2.16/2.16 | 41.4/53.7/53.7 | 0 | 2.048 |
|
||||||
|
| b41_1k_c32 | TP2PP4 | 64/64 | 29.73 | 275.55 | 2204.4 | 3.89/4.87/4.88 | 85.6/98.5/106.8 | 0 | None |
|
||||||
|
| b41_1k_c32 | TP8+EAGLE3+AR | 64/64 | 81.8 | 100.14 | 801.15 | 17.56/25.60/27.89 | 146.6/190.5/195.2 | 0 | 2.054 |
|
||||||
|
| b41_1k_c64 | TP2PP4 | 128/128 | 48.11 | 340.53 | 2724.25 | 8.14/19.96/20.01 | 93.3/101.8/125.1 | 0 | None |
|
||||||
|
| b41_1k_c64 | TP8+EAGLE3+AR | 128/128 | 158.73 | 103.22 | 825.75 | 46.74/65.81/70.71 | 144.5/177.7/212.5 | 0 | 2.087 |
|
||||||
|
| b42_1k4k_c1 | TP2PP4 | 8/8 | 1614.16 | 20.3 | 5.08 | 0.39/0.40/0.40 | 49.2/49.8/49.8 | 0 | None |
|
||||||
|
| b42_1k4k_c1 | TP8+EAGLE3+AR | 8/8 | 242.18 | 135.31 | 33.83 | 0.31/0.32/0.32 | 7.3/8.7/8.7 | 0 | 3.701 |
|
||||||
|
| b42_1k4k_c8 | TP2PP4 | 16/16 | 649.69 | 100.87 | 25.22 | 1.35/1.87/1.87 | 79.0/79.4/79.4 | 0 | None |
|
||||||
|
| b42_1k4k_c8 | TP8+EAGLE3+AR | 16/16 | 159.24 | 411.54 | 102.89 | 0.96/1.68/1.68 | 18.1/22.5/22.5 | 0 | 4.003 |
|
||||||
|
| b42_1k4k_c32 | TP2PP4 | 64/64 | 714.14 | 367.08 | 91.77 | 1.98/4.14/4.15 | 86.1/90.5/91.1 | 0 | None |
|
||||||
|
| b42_1k4k_c32 | TP8+EAGLE3+AR | 64/64 | 1040.79 | 251.87 | 62.97 | 195.07/304.75/334.69 | 60.7/74.6/92.3 | 0 | 3.852 |
|
||||||
|
| b42_1k4k_c64 | TP2PP4 | 128/128 | 1088.63 | 481.6 | 120.4 | 80.74/337.92/338.47 | 91.2/95.5/96.3 | 0 | None |
|
||||||
|
| b51_64k_c1 | TP2PP4 | 8/8 | 285.91 | 14.33 | 1833.72 | 10.28/10.32/10.32 | 49.8/50.0/50.0 | 0 | None |
|
||||||
|
| b51_64k_c1 | TP8+EAGLE3+AR | 8/8 | 182.23 | 22.48 | 2877.06 | 17.04/17.17/17.17 | 11.2/13.5/13.5 | 0 | 2.468 |
|
||||||
|
| b51_64k_c4 | TP2PP4 | 8/8 | 122.3 | 33.49 | 4286.94 | 19.55/28.83/28.83 | 81.4/99.5/99.5 | 0 | None |
|
||||||
|
| b51_64k_c4 | TP8+EAGLE3+AR | 8/8 | 159.24 | 25.72 | 3292.45 | 30.55/68.87/68.87 | 94.9/154.6/154.6 | 0 | 2.438 |
|
||||||
|
| b51_64k_c8 | TP2PP4 | 16/16 | 185.59 | 44.14 | 5650.11 | 31.83/53.38/53.38 | 119.2/161.8/161.8 | 0 | None |
|
||||||
|
| b51_64k_c8 | TP8+EAGLE3+AR | 16/16 | 319.66 | 25.63 | 3280.31 | 90.92/149.45/149.45 | 105.8/157.5/157.5 | 0 | 2.442 |
|
||||||
|
| b51_128k_c1 | TP2PP4 | 8/8 | 347.5 | 11.79 | 3017.45 | 18.21/18.23/18.23 | 49.4/49.7/49.7 | 0 | None |
|
||||||
|
| b51_128k_c1 | TP8+EAGLE3+AR | 8/8 | 343.36 | 11.93 | 3053.83 | 37.59/37.60/37.60 | 10.4/12.3/12.3 | 0 | 2.728 |
|
||||||
|
| b51_128k_c2 | TP2PP4 | 8/8 | 243.3 | 16.84 | 4309.84 | 25.27/32.42/32.42 | 69.6/83.7/83.7 | 0 | None |
|
||||||
|
| b51_128k_c2 | TP8+EAGLE3+AR | 8/8 | 333.36 | 12.29 | 3145.44 | 48.21/75.81/75.81 | 68.0/162.2/162.2 | 0 | 2.769 |
|
||||||
|
| b51_128k_c4 | TP2PP4 | 8/8 | 186.37 | 21.98 | 5626.35 | 39.29/60.35/60.35 | 105.4/146.8/146.8 | 0 | None |
|
||||||
|
| b51_128k_c4 | TP8+EAGLE3+AR | 8/8 | 334.71 | 12.24 | 3132.76 | 110.14/159.50/159.50 | 68.7/163.4/163.4 | 0 | 2.595 |
|
||||||
|
| b52_256k_c1 | TP2PP4 | 1/1 | 37.18 | 0.03 | 7050.23 | 37.18/37.18/37.18 | — | 0 | None |
|
||||||
|
| b52_256k_c1 | TP8+EAGLE3+AR | 1/1 | 91.45 | 0.01 | 2866.59 | 91.45/91.45/91.45 | — | 0 | None |
|
||||||
|
| b52_256k_c2 | TP2PP4 | 2/2 | 70.19 | 0.03 | 7469.22 | 53.69/70.12/70.12 | — | 0 | None |
|
||||||
|
| b52_256k_c3 | TP2PP4 | 3/3 | 103.12 | 0.03 | 7626.22 | 70.12/102.99/102.99 | — | 0 | None |
|
||||||
|
| b52_512k_c1 | TP2PP4 | 1/1 | 92.6 | 0.01 | 5662.03 | 92.60/92.60/92.60 | — | 0 | None |
|
||||||
|
| b52_896k_c1 | TP2PP4 | 1/1 | 220.91 | 0.0 | 4153.24 | 220.91/220.91/220.91 | — | 0 | None |
|
||||||
|
|
||||||
|
注:E7b 4.2 C=64 用户中止(rc=143)无 SUMMARY,不在表内;E7b 256K/512K/896K C>1 为结构性跳过;E7b 256K C=1 为全新文本干净重测值(窗口 9,900,000)。
|
||||||
@ -0,0 +1,424 @@
|
|||||||
|
# SPDX-License-Identifier: Apache-2.0
|
||||||
|
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
|
||||||
|
# Adapted from https://github.com/vllm-project/vllm/blob/v0.6.4.post1/vllm/distributed/device_communicators/custom_all_reduce.py
|
||||||
|
|
||||||
|
import ctypes
|
||||||
|
import logging
|
||||||
|
import os
|
||||||
|
from contextlib import contextmanager
|
||||||
|
from functools import partial
|
||||||
|
from typing import Any, List, Optional, Union
|
||||||
|
|
||||||
|
import torch
|
||||||
|
import torch.distributed as dist
|
||||||
|
from torch.distributed import ProcessGroup
|
||||||
|
|
||||||
|
import sglang.srt.distributed.device_communicators.custom_all_reduce_ops as ops
|
||||||
|
from sglang.srt.distributed.device_communicators.cuda_wrapper import CudaRTLibrary
|
||||||
|
from sglang.srt.distributed.device_communicators.custom_all_reduce_utils import (
|
||||||
|
can_use_custom_all_reduce_with_nvlink,
|
||||||
|
is_weak_contiguous,
|
||||||
|
)
|
||||||
|
from sglang.srt.environ import envs
|
||||||
|
from sglang.srt.model_executor.runner_backend_utils.tc_piecewise_cuda_graph import (
|
||||||
|
is_in_tc_piecewise_cuda_graph,
|
||||||
|
)
|
||||||
|
from sglang.srt.utils import (
|
||||||
|
get_bool_env_var,
|
||||||
|
is_cuda,
|
||||||
|
is_hip,
|
||||||
|
is_musa,
|
||||||
|
log_info_on_rank0,
|
||||||
|
)
|
||||||
|
|
||||||
|
_is_cuda = is_cuda()
|
||||||
|
_is_hip = is_hip()
|
||||||
|
_is_musa = is_musa()
|
||||||
|
|
||||||
|
logger = logging.getLogger(__name__)
|
||||||
|
os.environ.setdefault("SGLANG_CUSTOM_ALLREDUCE_ALGO", "1stage") # SSKJ-PATCH: C++ dispatch no-ops when full_nvlink=False; force kernel launch
|
||||||
|
|
||||||
|
|
||||||
|
class CustomAllreduce:
|
||||||
|
_SUPPORTED_WORLD_SIZES = [2, 4, 6, 8]
|
||||||
|
_MAX_CAR_SIZE = 8192 * 1024
|
||||||
|
if _is_hip:
|
||||||
|
# crossover is at 16MB buffer size for ROCm
|
||||||
|
_MAX_CAR_SIZE = 2 * 8192 * 1024
|
||||||
|
if _is_musa:
|
||||||
|
# crossover is at 128MB buffer size for MUSA
|
||||||
|
_MAX_CAR_SIZE = 16 * 8196 * 1024
|
||||||
|
|
||||||
|
# max_size: max supported allreduce size
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
group: ProcessGroup,
|
||||||
|
device: Union[int, str, torch.device],
|
||||||
|
max_size=_MAX_CAR_SIZE,
|
||||||
|
) -> None:
|
||||||
|
"""
|
||||||
|
Args:
|
||||||
|
group: the process group to work on. If None, it will use the
|
||||||
|
default process group.
|
||||||
|
device: the device to bind the CustomAllreduce to. If None,
|
||||||
|
it will be bind to f"cuda:{local_rank}".
|
||||||
|
It is the caller's responsibility to make sure each communicator
|
||||||
|
is bind to a unique device, and all communicators in this group
|
||||||
|
are in the same node.
|
||||||
|
"""
|
||||||
|
self._IS_CAPTURING = False
|
||||||
|
self.disabled = True # This can be modified in-place by context manager in piecewise cuda graph runner
|
||||||
|
self.original_disabled = True # To store the original state
|
||||||
|
self.use_amd_deterministic_impl = _use_amd_deterministic_impl()
|
||||||
|
|
||||||
|
if not ops.IS_CUSTOM_AR_AVAILABLE:
|
||||||
|
# disable because of missing custom allreduce library
|
||||||
|
# e.g. in a non-cuda environment
|
||||||
|
return
|
||||||
|
|
||||||
|
rank = dist.get_rank(group=group)
|
||||||
|
world_size = dist.get_world_size(group=group)
|
||||||
|
|
||||||
|
if isinstance(device, int):
|
||||||
|
device = torch.device(f"cuda:{device}")
|
||||||
|
elif isinstance(device, str):
|
||||||
|
device = torch.device(device)
|
||||||
|
# now `device` is a `torch.device` object
|
||||||
|
assert isinstance(device, torch.device)
|
||||||
|
self.device = device
|
||||||
|
full_nvlink = can_use_custom_all_reduce_with_nvlink(
|
||||||
|
group=group,
|
||||||
|
device=device,
|
||||||
|
supported_world_size=self._SUPPORTED_WORLD_SIZES,
|
||||||
|
cls_name="CustomAllreduce",
|
||||||
|
)
|
||||||
|
if full_nvlink is None:
|
||||||
|
return # fail to get nvlink status
|
||||||
|
|
||||||
|
self.group = group
|
||||||
|
self.max_size = max_size
|
||||||
|
self.rank = rank
|
||||||
|
self.world_size = world_size
|
||||||
|
self.full_nvlink = full_nvlink
|
||||||
|
|
||||||
|
if not _is_hip:
|
||||||
|
# Buffers memory are owned by this Python class and passed to C++.
|
||||||
|
# Meta data composes of two parts: meta data for synchronization and a
|
||||||
|
# temporary buffer for storing intermediate allreduce results.
|
||||||
|
self.meta_ptrs = self.create_shared_buffer(
|
||||||
|
ops.meta_size() + max_size, group=group
|
||||||
|
)
|
||||||
|
# This is a pre-registered IPC buffer. In eager mode, input tensors
|
||||||
|
# are first copied into this buffer before allreduce is performed
|
||||||
|
self.buffer_ptrs = self.create_shared_buffer(max_size, group=group)
|
||||||
|
# This is a buffer for storing the tuples of pointers pointing to
|
||||||
|
# IPC buffers from all ranks. Each registered tuple has size of
|
||||||
|
# 8*world_size bytes where world_size is at most 8. Allocating 8MB
|
||||||
|
# is enough for 131072 such tuples. The largest model I've seen only
|
||||||
|
# needs less than 10000 of registered tuples.
|
||||||
|
self.rank_data = torch.empty(
|
||||||
|
max_size, dtype=torch.uint8, device=self.device
|
||||||
|
)
|
||||||
|
self._ptr = ops.init_custom_ar(
|
||||||
|
self.meta_ptrs, self.rank_data, rank, self.full_nvlink
|
||||||
|
)
|
||||||
|
ops.register_buffer(self._ptr, self.buffer_ptrs)
|
||||||
|
else:
|
||||||
|
# meta data buffers need to be "uncached" for signal on MI200
|
||||||
|
self.meta = ops.allocate_meta_buffer(ops.meta_size() + max_size)
|
||||||
|
self.buffer = torch.empty(max_size, dtype=torch.uint8, device=self.device)
|
||||||
|
handle = ops.get_meta_buffer_ipc_handle(self.meta)
|
||||||
|
shard_data = (
|
||||||
|
bytes(handle), # ipc handle to base ptr
|
||||||
|
0, # offset of base ptr
|
||||||
|
)
|
||||||
|
handles, offsets = self._gather_ipc_meta(shard_data)
|
||||||
|
self.rank_data = torch.empty(
|
||||||
|
max_size, dtype=torch.uint8, device=self.device
|
||||||
|
)
|
||||||
|
self._ptr = ops.init_custom_ar(
|
||||||
|
self.meta, self.rank_data, handles, offsets, rank, self.full_nvlink
|
||||||
|
)
|
||||||
|
self.register_buffer(self.buffer)
|
||||||
|
|
||||||
|
self.disabled = False
|
||||||
|
self.original_disabled = False # Ensure original_disabled == disabled
|
||||||
|
logger.warning(f"SSKJ_CAR_PATCH_ACTIVE ws={self.world_size} full_nvlink={self.full_nvlink}")
|
||||||
|
self.tms_cudagraph = envs.SGLANG_MEMORY_SAVER_CUDA_GRAPH.get()
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def create_shared_buffer(
|
||||||
|
size_in_bytes: int, group: Optional[ProcessGroup] = None
|
||||||
|
) -> List[int]:
|
||||||
|
"""
|
||||||
|
Creates a shared buffer and returns a list of pointers
|
||||||
|
representing the buffer on all processes in the group.
|
||||||
|
"""
|
||||||
|
lib = CudaRTLibrary()
|
||||||
|
pointer = lib.cudaMalloc(size_in_bytes)
|
||||||
|
if _is_musa:
|
||||||
|
lib.cudaMemset(pointer, 0, size_in_bytes)
|
||||||
|
handle = lib.cudaIpcGetMemHandle(pointer)
|
||||||
|
world_size = dist.get_world_size(group=group)
|
||||||
|
rank = dist.get_rank(group=group)
|
||||||
|
handles = [None] * world_size
|
||||||
|
dist.all_gather_object(handles, handle, group=group)
|
||||||
|
|
||||||
|
pointers: List[int] = []
|
||||||
|
for i, h in enumerate(handles):
|
||||||
|
if i == rank:
|
||||||
|
pointers.append(pointer.value) # type: ignore
|
||||||
|
else:
|
||||||
|
pointers.append(lib.cudaIpcOpenMemHandle(h).value) # type: ignore
|
||||||
|
|
||||||
|
return pointers
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def free_shared_buffer(
|
||||||
|
pointers: List[int], group: Optional[ProcessGroup] = None
|
||||||
|
) -> None:
|
||||||
|
rank = dist.get_rank(group=group)
|
||||||
|
lib = CudaRTLibrary()
|
||||||
|
lib.cudaFree(ctypes.c_void_p(pointers[rank]))
|
||||||
|
|
||||||
|
@contextmanager
|
||||||
|
def capture(self):
|
||||||
|
"""
|
||||||
|
The main responsibility of this context manager is the
|
||||||
|
`register_graph_buffers` call at the end of the context.
|
||||||
|
It records all the buffer addresses used in the CUDA graph.
|
||||||
|
"""
|
||||||
|
try:
|
||||||
|
self._IS_CAPTURING = True
|
||||||
|
yield
|
||||||
|
finally:
|
||||||
|
self._IS_CAPTURING = False
|
||||||
|
if not self.disabled:
|
||||||
|
self.register_graph_buffers()
|
||||||
|
|
||||||
|
def _get_ipc_meta(self, inp: torch.Tensor):
|
||||||
|
# _share_cuda_() doesn't accept meta buffer not allocated from
|
||||||
|
# PyTorch cache allocator, use direct HIP call to get IPC handle
|
||||||
|
handle = ops.get_meta_buffer_ipc_handle(inp)
|
||||||
|
shard_data = (
|
||||||
|
bytes(handle), # ipc handle to base ptr
|
||||||
|
0, # offset of base ptr
|
||||||
|
)
|
||||||
|
return self._gather_ipc_meta(shard_data)
|
||||||
|
|
||||||
|
def _gather_ipc_meta(self, shard_data):
|
||||||
|
# Note: don't use `[[None]] * self.world_size` here
|
||||||
|
# because it will create a list of the same reference
|
||||||
|
all_data: List[Optional[Any]] = [[None] for i in range(self.world_size)]
|
||||||
|
all_data[self.rank][0] = shard_data
|
||||||
|
|
||||||
|
ranks = dist.get_process_group_ranks(group=self.group)
|
||||||
|
ranks.sort()
|
||||||
|
for i, rank in enumerate(ranks):
|
||||||
|
dist.broadcast_object_list(
|
||||||
|
all_data[i], src=rank, group=self.group, device="cpu"
|
||||||
|
)
|
||||||
|
|
||||||
|
# we cannot directly use `dist.all_gather_object` here
|
||||||
|
# because it is incompatible with `gloo` backend under inference mode.
|
||||||
|
# see https://github.com/pytorch/pytorch/issues/126032 for details.
|
||||||
|
|
||||||
|
handles = []
|
||||||
|
offsets = []
|
||||||
|
for i in range(len(all_data)):
|
||||||
|
handles.append(all_data[i][0][0]) # type: ignore
|
||||||
|
offsets.append(all_data[i][0][1]) # type: ignore
|
||||||
|
return handles, offsets
|
||||||
|
|
||||||
|
def register_buffer(self, inp: torch.Tensor):
|
||||||
|
handles, offsets = self._get_ipc_meta(inp)
|
||||||
|
ops.register_buffer(self._ptr, inp, handles, offsets)
|
||||||
|
|
||||||
|
def register_graph_buffers(self):
|
||||||
|
if _is_hip:
|
||||||
|
handle, offset = ops.get_graph_buffer_ipc_meta(self._ptr)
|
||||||
|
handles, offsets = self._gather_ipc_meta((bytes(handle), offset))
|
||||||
|
log_info_on_rank0(logger, f"Registering {len(offset)} cuda graph addresses")
|
||||||
|
ops.register_graph_buffers(self._ptr, handles, offsets)
|
||||||
|
else:
|
||||||
|
handle, offset = ops.get_graph_buffer_ipc_meta(self._ptr)
|
||||||
|
log_info_on_rank0(logger, f"Registering {len(offset)} cuda graph addresses")
|
||||||
|
# We cannot directly use `dist.all_gather_object` here
|
||||||
|
# because it is incompatible with `gloo` backend under inference mode.
|
||||||
|
# see https://github.com/pytorch/pytorch/issues/126032 for details.
|
||||||
|
all_data = [
|
||||||
|
[None, None] for _ in range(dist.get_world_size(group=self.group))
|
||||||
|
]
|
||||||
|
all_data[self.rank] = [handle, offset]
|
||||||
|
ranks = sorted(dist.get_process_group_ranks(group=self.group))
|
||||||
|
for i, rank in enumerate(ranks):
|
||||||
|
dist.broadcast_object_list(
|
||||||
|
all_data[i], src=rank, group=self.group, device="cpu"
|
||||||
|
)
|
||||||
|
# Unpack list of tuples to tuple of lists.
|
||||||
|
handles = [d[0] for d in all_data] # type: ignore
|
||||||
|
offsets = [d[1] for d in all_data] # type: ignore
|
||||||
|
ops.register_graph_buffers(self._ptr, handles, offsets)
|
||||||
|
|
||||||
|
def should_custom_ar(self, inp: torch.Tensor):
|
||||||
|
if self.disabled:
|
||||||
|
return False
|
||||||
|
inp_size = inp.numel() * inp.element_size()
|
||||||
|
# custom allreduce requires input byte size to be multiples of 16
|
||||||
|
if inp_size % 16 != 0:
|
||||||
|
return False
|
||||||
|
if not is_weak_contiguous(inp):
|
||||||
|
return False
|
||||||
|
# for 4 or more non NVLink-capable GPUs, custom allreduce provides
|
||||||
|
# little performance improvement over NCCL.
|
||||||
|
if not _is_hip:
|
||||||
|
if True:
|
||||||
|
return inp_size <= self.max_size
|
||||||
|
return False
|
||||||
|
|
||||||
|
if _is_hip:
|
||||||
|
if self.use_amd_deterministic_impl:
|
||||||
|
return True
|
||||||
|
if self.full_nvlink:
|
||||||
|
return inp_size <= self.max_size
|
||||||
|
return False
|
||||||
|
|
||||||
|
return False
|
||||||
|
|
||||||
|
def _all_reduce_impl(self, inp: torch.Tensor, registered: bool):
|
||||||
|
out = torch.empty_like(inp)
|
||||||
|
if not _is_hip: # CUDA-like
|
||||||
|
if registered:
|
||||||
|
ops.all_reduce(self._ptr, inp, out, 0, 0)
|
||||||
|
else:
|
||||||
|
ops.all_reduce(
|
||||||
|
self._ptr, inp, out, self.buffer_ptrs[self.rank], self.max_size
|
||||||
|
)
|
||||||
|
elif self.use_amd_deterministic_impl:
|
||||||
|
inp_size = inp.numel() * inp.element_size()
|
||||||
|
if inp_size < self.max_size:
|
||||||
|
reg_buffer = self.buffer.view(inp.dtype)[: inp.numel()]
|
||||||
|
ops.deterministic_all_reduce_unreg(self._ptr, inp, reg_buffer, out)
|
||||||
|
else:
|
||||||
|
self.register_buffer(inp)
|
||||||
|
ops.deterministic_all_reduce_reg(self._ptr, inp, out)
|
||||||
|
else: # normal AMD ROCm path
|
||||||
|
if registered:
|
||||||
|
ops.all_reduce_reg(self._ptr, inp, out)
|
||||||
|
else:
|
||||||
|
ops.all_reduce_unreg(self._ptr, inp, self.buffer, out)
|
||||||
|
return out
|
||||||
|
|
||||||
|
def custom_all_reduce(self, input: torch.Tensor) -> Optional[torch.Tensor]:
|
||||||
|
"""The main allreduce API that provides support for cuda graph."""
|
||||||
|
# When custom allreduce is disabled, this will be None.
|
||||||
|
if self.disabled or not self.should_custom_ar(input):
|
||||||
|
return None
|
||||||
|
if self._IS_CAPTURING:
|
||||||
|
if torch.cuda.is_current_stream_capturing():
|
||||||
|
return self._all_reduce_impl(input, registered=not self.tms_cudagraph)
|
||||||
|
else:
|
||||||
|
# Could be warmup OR piecewise cuda graph split op execution.
|
||||||
|
# In piecewise cuda graph, split ops run eagerly outside the graph
|
||||||
|
# but _IS_CAPTURING is still True. We need to do real all-reduce.
|
||||||
|
if is_in_tc_piecewise_cuda_graph():
|
||||||
|
# Split op execution - do real all-reduce
|
||||||
|
return self._all_reduce_impl(input, registered=False)
|
||||||
|
else:
|
||||||
|
# True warmup - mimic the allocation pattern since custom
|
||||||
|
# allreduce is out-of-place.
|
||||||
|
return torch.zeros_like(input)
|
||||||
|
else:
|
||||||
|
return self._all_reduce_impl(input, registered=False)
|
||||||
|
|
||||||
|
def close(self):
|
||||||
|
if not self.disabled and self._ptr:
|
||||||
|
if ops is not None:
|
||||||
|
ops.dispose(self._ptr)
|
||||||
|
if _is_cuda:
|
||||||
|
self.free_shared_buffer(self.meta_ptrs)
|
||||||
|
self.free_shared_buffer(self.buffer_ptrs)
|
||||||
|
self._ptr = 0
|
||||||
|
|
||||||
|
def __del__(self):
|
||||||
|
self.close()
|
||||||
|
|
||||||
|
|
||||||
|
def dispatch_custom_allreduce(
|
||||||
|
group: ProcessGroup,
|
||||||
|
device: torch.device,
|
||||||
|
):
|
||||||
|
"""Return the CustomAllreduce class to use (aiter on ROCm if enabled).
|
||||||
|
|
||||||
|
On AMD with 1-stage AR enabled, use sglang's CustomAllreduce.
|
||||||
|
Otherwise use AiterCustomAllreduce if available.
|
||||||
|
|
||||||
|
On CUDA, the JIT-compiled v2 implementation is used by default.
|
||||||
|
Set SGLANG_OPT_USE_CUSTOM_ALL_REDUCE_V2=0 to fall back to the legacy CustomAllreduce.
|
||||||
|
Multi-node v2 is admitted only for a single NVLink clique (see
|
||||||
|
can_use_custom_all_reduce_v2); other cross-node groups fall back to NCCL.
|
||||||
|
"""
|
||||||
|
if _is_cuda and envs.SGLANG_OPT_USE_CUSTOM_ALL_REDUCE_V2.get():
|
||||||
|
from .custom_all_reduce_v2 import (
|
||||||
|
CustomAllReduceV2,
|
||||||
|
can_use_custom_all_reduce_v2,
|
||||||
|
)
|
||||||
|
|
||||||
|
if can_use_custom_all_reduce_v2(group=group, device=device):
|
||||||
|
logger.debug("[AR] Using CustomAllReduceV2 (JIT-compiled)")
|
||||||
|
return CustomAllReduceV2
|
||||||
|
|
||||||
|
if _is_cuda or _is_musa:
|
||||||
|
return CustomAllreduce
|
||||||
|
|
||||||
|
assert _is_hip
|
||||||
|
|
||||||
|
if envs.SGLANG_USE_1STAGE_ALLREDUCE.is_set():
|
||||||
|
if envs.SGLANG_USE_1STAGE_ALLREDUCE.get():
|
||||||
|
logger.debug(
|
||||||
|
"[AR] All-reduce: 1-stage kernel (SGLANG_USE_1STAGE_ALLREDUCE=1)"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
logger.debug("[AR] All-reduce: default (SGLANG_USE_1STAGE_ALLREDUCE=0)")
|
||||||
|
elif envs.SGLANG_ENABLE_DETERMINISTIC_INFERENCE.get():
|
||||||
|
logger.debug(
|
||||||
|
"[AR] All-reduce: 1-stage kernel (deterministic inference enabled)"
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
logger.debug("[AR] All-reduce: default")
|
||||||
|
|
||||||
|
# On AMD with 1-stage AR, use sglang's CustomAllreduce
|
||||||
|
# (AiterCustomAllreduce doesn't have deterministic_all_reduce method)
|
||||||
|
if _use_amd_deterministic_impl():
|
||||||
|
return CustomAllreduce
|
||||||
|
|
||||||
|
if get_bool_env_var("SGLANG_USE_AITER_AR", default="true"):
|
||||||
|
try:
|
||||||
|
from aiter.dist.device_communicators.custom_all_reduce import (
|
||||||
|
CustomAllreduce as AiterCustomAllreduce,
|
||||||
|
)
|
||||||
|
|
||||||
|
logger.info("[AR] Using AiterCustomAllreduce (AMD default)")
|
||||||
|
tms_cudagraph = envs.SGLANG_MEMORY_SAVER_CUDA_GRAPH.get()
|
||||||
|
return partial(
|
||||||
|
AiterCustomAllreduce,
|
||||||
|
enable_register_for_capturing=not tms_cudagraph,
|
||||||
|
)
|
||||||
|
except ImportError as e:
|
||||||
|
logger.warning(
|
||||||
|
"[AR] Aiter custom all-reduce not available; "
|
||||||
|
"falling back to sglang CustomAllreduce. Details: %s",
|
||||||
|
e,
|
||||||
|
)
|
||||||
|
return CustomAllreduce
|
||||||
|
|
||||||
|
return CustomAllreduce
|
||||||
|
|
||||||
|
|
||||||
|
def _use_amd_deterministic_impl() -> bool:
|
||||||
|
if not _is_hip: # CUDA is always deterministic
|
||||||
|
return False
|
||||||
|
if envs.SGLANG_USE_1STAGE_ALLREDUCE.is_set():
|
||||||
|
return envs.SGLANG_USE_1STAGE_ALLREDUCE.get()
|
||||||
|
else:
|
||||||
|
return envs.SGLANG_ENABLE_DETERMINISTIC_INFERENCE.get()
|
||||||
@ -0,0 +1,519 @@
|
|||||||
|
# SPDX-License-Identifier: Apache-2.0
|
||||||
|
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
|
||||||
|
# Adapted from https://github.com/vllm-project/vllm/blob/v0.6.4.post1/vllm/distributed/device_communicators/custom_all_reduce_utils.py
|
||||||
|
|
||||||
|
import ctypes
|
||||||
|
import json
|
||||||
|
import logging
|
||||||
|
import os
|
||||||
|
import pickle
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
import tempfile
|
||||||
|
from functools import wraps
|
||||||
|
from itertools import product
|
||||||
|
from typing import Callable, Dict, List, Optional, Sequence, TypeVar
|
||||||
|
|
||||||
|
import torch
|
||||||
|
import torch.distributed as dist
|
||||||
|
import torch.multiprocessing as mp
|
||||||
|
from typing_extensions import ParamSpec
|
||||||
|
|
||||||
|
from sglang.srt.distributed.device_communicators.cuda_wrapper import CudaRTLibrary
|
||||||
|
from sglang.srt.distributed.parallel_state import in_the_same_node_as
|
||||||
|
from sglang.srt.environ import envs as sglang_envs
|
||||||
|
from sglang.srt.utils import is_cuda, is_hip, is_musa
|
||||||
|
from sglang.srt.utils.cuda_vmm_utils import _gpu_fabric_clique
|
||||||
|
|
||||||
|
logger = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
_is_cuda = is_cuda()
|
||||||
|
_is_hip = is_hip()
|
||||||
|
_is_musa = is_musa()
|
||||||
|
|
||||||
|
if _is_cuda:
|
||||||
|
try:
|
||||||
|
import pynvml
|
||||||
|
except ImportError as e:
|
||||||
|
logger.warning("Failed to import pynvml with %r", e)
|
||||||
|
|
||||||
|
if _is_musa:
|
||||||
|
try:
|
||||||
|
import pymtml as pynvml
|
||||||
|
except ImportError as e:
|
||||||
|
logger.warning("Failed to import pymtml with %r", e)
|
||||||
|
|
||||||
|
if _is_hip:
|
||||||
|
try:
|
||||||
|
from amdsmi import (
|
||||||
|
AmdSmiException,
|
||||||
|
amdsmi_get_processor_handles,
|
||||||
|
amdsmi_init,
|
||||||
|
amdsmi_shut_down,
|
||||||
|
amdsmi_topo_get_link_type,
|
||||||
|
)
|
||||||
|
except ImportError as e:
|
||||||
|
logger.warning("Failed to import amdsmi with %r", e)
|
||||||
|
|
||||||
|
_P = ParamSpec("_P")
|
||||||
|
_R = TypeVar("_R")
|
||||||
|
|
||||||
|
|
||||||
|
def update_environment_variables(envs: Dict[str, str]):
|
||||||
|
for k, v in envs.items():
|
||||||
|
if k in os.environ and os.environ[k] != v:
|
||||||
|
logger.warning(
|
||||||
|
"Overwriting environment variable %s " "from '%s' to '%s'",
|
||||||
|
k,
|
||||||
|
os.environ[k],
|
||||||
|
v,
|
||||||
|
)
|
||||||
|
os.environ[k] = v
|
||||||
|
|
||||||
|
|
||||||
|
def producer(
|
||||||
|
batch_src: Sequence[int],
|
||||||
|
producer_queue,
|
||||||
|
consumer_queue,
|
||||||
|
result_queue,
|
||||||
|
cuda_visible_devices: Optional[str] = None,
|
||||||
|
):
|
||||||
|
if cuda_visible_devices is not None:
|
||||||
|
update_environment_variables({"CUDA_VISIBLE_DEVICES": cuda_visible_devices})
|
||||||
|
|
||||||
|
lib = CudaRTLibrary()
|
||||||
|
for i in batch_src:
|
||||||
|
lib.cudaSetDevice(i)
|
||||||
|
pointer = lib.cudaMalloc(1024)
|
||||||
|
lib.cudaMemset(pointer, 1, 1024)
|
||||||
|
lib.cudaDeviceSynchronize()
|
||||||
|
handle = lib.cudaIpcGetMemHandle(pointer)
|
||||||
|
producer_queue.put(handle)
|
||||||
|
open_success = consumer_queue.get()
|
||||||
|
if open_success:
|
||||||
|
# use two queues to simulate barrier
|
||||||
|
producer_queue.put(0)
|
||||||
|
consumer_queue.get()
|
||||||
|
# check if the memory is modified
|
||||||
|
host_data = (ctypes.c_char * 1024)()
|
||||||
|
lib.cudaMemcpy(host_data, pointer, 1024) # type: ignore
|
||||||
|
for i in range(1024):
|
||||||
|
if ord(host_data[i]) != 2:
|
||||||
|
open_success = False
|
||||||
|
break
|
||||||
|
result_queue.put(open_success)
|
||||||
|
lib.cudaDeviceReset()
|
||||||
|
|
||||||
|
|
||||||
|
def consumer(
|
||||||
|
batch_tgt: Sequence[int],
|
||||||
|
producer_queue,
|
||||||
|
consumer_queue,
|
||||||
|
result_queue,
|
||||||
|
cuda_visible_devices: Optional[str] = None,
|
||||||
|
):
|
||||||
|
if cuda_visible_devices is not None:
|
||||||
|
update_environment_variables({"CUDA_VISIBLE_DEVICES": cuda_visible_devices})
|
||||||
|
|
||||||
|
lib = CudaRTLibrary()
|
||||||
|
for j in batch_tgt:
|
||||||
|
lib.cudaSetDevice(j)
|
||||||
|
handle = producer_queue.get()
|
||||||
|
open_success = False
|
||||||
|
try:
|
||||||
|
pointer = lib.cudaIpcOpenMemHandle(handle) # type: ignore
|
||||||
|
open_success = True
|
||||||
|
except RuntimeError:
|
||||||
|
# cannot error out here, because the producer process
|
||||||
|
# is still waiting for the response.
|
||||||
|
pass
|
||||||
|
consumer_queue.put(open_success)
|
||||||
|
if open_success:
|
||||||
|
# modify the memory
|
||||||
|
lib.cudaMemset(pointer, 2, 1024)
|
||||||
|
lib.cudaDeviceSynchronize()
|
||||||
|
# use two queues to simulate barrier
|
||||||
|
producer_queue.get()
|
||||||
|
consumer_queue.put(0)
|
||||||
|
# check if the memory is modified
|
||||||
|
host_data = (ctypes.c_char * 1024)()
|
||||||
|
lib.cudaMemcpy(host_data, pointer, 1024) # type: ignore
|
||||||
|
for i in range(1024):
|
||||||
|
if ord(host_data[i]) != 2:
|
||||||
|
open_success = False
|
||||||
|
break
|
||||||
|
result_queue.put(open_success)
|
||||||
|
lib.cudaDeviceReset()
|
||||||
|
|
||||||
|
|
||||||
|
def can_actually_p2p(
|
||||||
|
batch_src: Sequence[int],
|
||||||
|
batch_tgt: Sequence[int],
|
||||||
|
) -> Sequence[bool]:
|
||||||
|
"""
|
||||||
|
Usually, checking if P2P access is enabled can be done by
|
||||||
|
`torch.cuda.can_device_access_peer(src, tgt)`. However, sometimes
|
||||||
|
the driver might be broken, and `torch.cuda.can_device_access_peer(src, tgt)`
|
||||||
|
returns `True` even if P2P access is not actually possible.
|
||||||
|
See https://github.com/vllm-project/vllm/issues/2728 and
|
||||||
|
https://forums.developer.nvidia.com/t/direct-gpu-gpu-communication-does-not-seem-to-work-properly/283264/10
|
||||||
|
Therefore, we have to perform a real P2P access to check if it is actually
|
||||||
|
possible.
|
||||||
|
|
||||||
|
Note on p2p and cuda IPC:
|
||||||
|
Usually, one process uses one GPU:
|
||||||
|
GPU src --> cuda context src --> tensor src --> process src
|
||||||
|
|
||||||
|
We need to combine p2p and cuda IPC, so that:
|
||||||
|
GPU src --> cuda context src --> tensor src --> process src
|
||||||
|
|shared|
|
||||||
|
GPU tgt --> cuda context tgt --> tensor tgt --> process tgt
|
||||||
|
That is to say, process src creates a tensor in GPU src, passes IPC handle to
|
||||||
|
process tgt, and process tgt accesses the tensor in GPU tgt. Any operation on the
|
||||||
|
tensor in process tgt will be reflected in the tensor in process src, because
|
||||||
|
they are the same memory segment.
|
||||||
|
It is important to note that process tgt accesses the tensor in GPU tgt, not
|
||||||
|
GPU src. That's why we need p2p access.
|
||||||
|
|
||||||
|
The most time-consuming part is the process creation. To avoid creating
|
||||||
|
processes for every pair of GPUs, we use batched testing. We create two
|
||||||
|
processes for testing all pairs of GPUs in batch. The trick is to reset
|
||||||
|
the device after each test (which is not available in PyTorch).
|
||||||
|
""" # noqa
|
||||||
|
cuda_visible_devices = os.environ.get("CUDA_VISIBLE_DEVICES", None)
|
||||||
|
# pass the CUDA_VISIBLE_DEVICES to the child process
|
||||||
|
# to make sure they see the same set of GPUs
|
||||||
|
|
||||||
|
# make sure the processes are spawned
|
||||||
|
smp = mp.get_context("spawn")
|
||||||
|
producer_queue = smp.Queue()
|
||||||
|
consumer_queue = smp.Queue()
|
||||||
|
result_queue = smp.Queue()
|
||||||
|
p_src = smp.Process(
|
||||||
|
target=producer,
|
||||||
|
args=(
|
||||||
|
batch_src,
|
||||||
|
producer_queue,
|
||||||
|
consumer_queue,
|
||||||
|
result_queue,
|
||||||
|
cuda_visible_devices,
|
||||||
|
),
|
||||||
|
)
|
||||||
|
p_tgt = smp.Process(
|
||||||
|
target=consumer,
|
||||||
|
args=(
|
||||||
|
batch_tgt,
|
||||||
|
producer_queue,
|
||||||
|
consumer_queue,
|
||||||
|
result_queue,
|
||||||
|
cuda_visible_devices,
|
||||||
|
),
|
||||||
|
)
|
||||||
|
p_src.start()
|
||||||
|
p_tgt.start()
|
||||||
|
p_src.join()
|
||||||
|
p_tgt.join()
|
||||||
|
assert p_src.exitcode == 0 and p_tgt.exitcode == 0
|
||||||
|
result: List[bool] = []
|
||||||
|
for src, tgt in zip(batch_src, batch_tgt):
|
||||||
|
a = result_queue.get()
|
||||||
|
b = result_queue.get()
|
||||||
|
if a != b:
|
||||||
|
logger.warning(
|
||||||
|
"Two processes do not agree on the P2P access"
|
||||||
|
" status on %d -> %d, treat as disabled.",
|
||||||
|
src,
|
||||||
|
tgt,
|
||||||
|
)
|
||||||
|
result.append(False)
|
||||||
|
else:
|
||||||
|
result.append(a)
|
||||||
|
return result
|
||||||
|
|
||||||
|
|
||||||
|
# why do we need this cache?
|
||||||
|
# we are testing peer-to-peer (p2p) access between GPUs,across processes.
|
||||||
|
# if we test it every time, it will be very slow, because we need to create
|
||||||
|
# N * N * 2 processes, where N is the world size. This is very slow.
|
||||||
|
# to reduce the time, we use a cache file to store the p2p access status.
|
||||||
|
# the cache file is generated by the master process if it does not exist.
|
||||||
|
# then all the processes can read the cache file to check the p2p access status.
|
||||||
|
# Note that the cache file is suffixed by the CUDA_VISIBLE_DEVICES, so that we
|
||||||
|
# can have different cache files for different CUDA_VISIBLE_DEVICES settings,
|
||||||
|
# e.g. used by different vllm engines. The device id in the cache file is a
|
||||||
|
# **local** device id, i.e. from 0 to num_dev-1, where num_dev is the number
|
||||||
|
# of visible devices in the vllm engine.
|
||||||
|
_gpu_p2p_access_cache: Optional[Dict[str, bool]] = None
|
||||||
|
|
||||||
|
|
||||||
|
def gpu_p2p_access_check(src: int, tgt: int) -> bool:
|
||||||
|
"""Check if GPU src can access GPU tgt."""
|
||||||
|
|
||||||
|
# if the cache variable is already calculated,
|
||||||
|
# read from the cache instead of checking it again
|
||||||
|
global _gpu_p2p_access_cache
|
||||||
|
if _gpu_p2p_access_cache is not None:
|
||||||
|
return _gpu_p2p_access_cache[f"{src}->{tgt}"]
|
||||||
|
|
||||||
|
is_distributed = dist.is_initialized()
|
||||||
|
|
||||||
|
num_dev = torch.cuda.device_count()
|
||||||
|
cuda_visible_devices = os.environ.get("CUDA_VISIBLE_DEVICES", None)
|
||||||
|
if cuda_visible_devices is None:
|
||||||
|
cuda_visible_devices = ",".join(str(i) for i in range(num_dev))
|
||||||
|
|
||||||
|
# VLLM_CACHE_ROOT -> SGLANG_CACHE_ROOT
|
||||||
|
# "~/.cache/vllm" -> envs.SGLANG_CACHE_DIR
|
||||||
|
SGLANG_CACHE_ROOT = os.path.expanduser(sglang_envs.SGLANG_CACHE_DIR.get())
|
||||||
|
path = os.path.join(
|
||||||
|
SGLANG_CACHE_ROOT, f"gpu_p2p_access_cache_for_{cuda_visible_devices}.json"
|
||||||
|
)
|
||||||
|
cache_dir = os.path.dirname(path)
|
||||||
|
try:
|
||||||
|
os.makedirs(cache_dir, exist_ok=True)
|
||||||
|
except (FileExistsError, NotADirectoryError):
|
||||||
|
if not os.path.isdir(cache_dir):
|
||||||
|
# Path exists as a file (stale cache/lock). Remove and retry.
|
||||||
|
try:
|
||||||
|
os.remove(cache_dir)
|
||||||
|
except OSError:
|
||||||
|
pass
|
||||||
|
os.makedirs(cache_dir, exist_ok=True)
|
||||||
|
from sglang.srt.distributed.parallel_state import get_world_group
|
||||||
|
|
||||||
|
if (not is_distributed or get_world_group().local_rank == 0) and (
|
||||||
|
not os.path.exists(path)
|
||||||
|
):
|
||||||
|
# only the local master process (with local_rank == 0) can
|
||||||
|
# enter this block to calculate the cache
|
||||||
|
logger.info("generating GPU P2P access cache in %s", path)
|
||||||
|
cache: Dict[str, bool] = {}
|
||||||
|
ids = list(range(num_dev))
|
||||||
|
# batch of all pairs of GPUs
|
||||||
|
batch_src, batch_tgt = zip(*list(product(ids, ids)))
|
||||||
|
# NOTE: we use `subprocess` rather than `multiprocessing` here
|
||||||
|
# because the caller might not have `if __name__ == "__main__":`,
|
||||||
|
# in that case we cannot use spawn method in multiprocessing.
|
||||||
|
# However, `can_actually_p2p` requires spawn method.
|
||||||
|
# The fix is, we use `subprocess` to call the function,
|
||||||
|
# where we have `if __name__ == "__main__":` in this file.
|
||||||
|
|
||||||
|
# use a temporary file to store the result
|
||||||
|
# we don't use the output of the subprocess directly,
|
||||||
|
# because the subprocess might produce logging output
|
||||||
|
with tempfile.NamedTemporaryFile() as output_file:
|
||||||
|
input_bytes = pickle.dumps((batch_src, batch_tgt, output_file.name))
|
||||||
|
returned = subprocess.run(
|
||||||
|
[sys.executable, __file__], input=input_bytes, capture_output=True
|
||||||
|
)
|
||||||
|
# check if the subprocess is successful
|
||||||
|
try:
|
||||||
|
returned.check_returncode()
|
||||||
|
except Exception as e:
|
||||||
|
# wrap raised exception to provide more information
|
||||||
|
raise RuntimeError(
|
||||||
|
f"Error happened when batch testing "
|
||||||
|
f"peer-to-peer access from {batch_src} to {batch_tgt}:\n"
|
||||||
|
f"{returned.stderr.decode()}"
|
||||||
|
) from e
|
||||||
|
with open(output_file.name, "rb") as f:
|
||||||
|
result = pickle.load(f)
|
||||||
|
for _i, _j, r in zip(batch_src, batch_tgt, result):
|
||||||
|
cache[f"{_i}->{_j}"] = r
|
||||||
|
with open(path, "w") as f:
|
||||||
|
json.dump(cache, f, indent=4)
|
||||||
|
if is_distributed:
|
||||||
|
get_world_group().barrier()
|
||||||
|
logger.info("reading GPU P2P access cache from %s", path)
|
||||||
|
with open(path) as f:
|
||||||
|
cache = json.load(f)
|
||||||
|
_gpu_p2p_access_cache = cache
|
||||||
|
return _gpu_p2p_access_cache[f"{src}->{tgt}"]
|
||||||
|
|
||||||
|
|
||||||
|
def with_nvml_context(fn: Callable[_P, _R]) -> Callable[_P, _R]:
|
||||||
|
@wraps(fn)
|
||||||
|
def wrapper(*args: _P.args, **kwargs: _P.kwargs) -> _R:
|
||||||
|
if _is_hip:
|
||||||
|
try:
|
||||||
|
amdsmi_init()
|
||||||
|
return fn(*args, **kwargs)
|
||||||
|
finally:
|
||||||
|
amdsmi_shut_down()
|
||||||
|
else:
|
||||||
|
pynvml.nvmlInit()
|
||||||
|
try:
|
||||||
|
return fn(*args, **kwargs)
|
||||||
|
finally:
|
||||||
|
pynvml.nvmlShutdown()
|
||||||
|
|
||||||
|
return wrapper
|
||||||
|
|
||||||
|
|
||||||
|
@with_nvml_context
|
||||||
|
def is_full_nvlink(physical_device_ids: List[int], world_size: int) -> bool:
|
||||||
|
if _is_hip:
|
||||||
|
"""
|
||||||
|
query if the set of gpus are fully connected by xgmi (1 hop)
|
||||||
|
"""
|
||||||
|
handles = [amdsmi_get_processor_handles()[i] for i in physical_device_ids]
|
||||||
|
for i, handle in enumerate(handles):
|
||||||
|
for j, peer_handle in enumerate(handles):
|
||||||
|
if i < j:
|
||||||
|
try:
|
||||||
|
link_type = amdsmi_topo_get_link_type(handle, peer_handle)
|
||||||
|
# type is 2 for XGMI
|
||||||
|
if link_type["hops"] != 1 or link_type["type"] != 2:
|
||||||
|
return False
|
||||||
|
except AmdSmiException as error:
|
||||||
|
logger.error("AMD 1 hop XGMI detection failed.", exc_info=error)
|
||||||
|
return False
|
||||||
|
return True
|
||||||
|
else:
|
||||||
|
"""
|
||||||
|
query if the set of gpus are fully connected by nvlink (1 hop)
|
||||||
|
"""
|
||||||
|
handles = [pynvml.nvmlDeviceGetHandleByIndex(i) for i in physical_device_ids]
|
||||||
|
for i, handle in enumerate(handles):
|
||||||
|
for j, peer_handle in enumerate(handles):
|
||||||
|
if i < j:
|
||||||
|
try:
|
||||||
|
p2p_status = pynvml.nvmlDeviceGetP2PStatus(
|
||||||
|
handle, peer_handle, pynvml.NVML_P2P_CAPS_INDEX_NVLINK
|
||||||
|
)
|
||||||
|
if p2p_status != pynvml.NVML_P2P_STATUS_OK:
|
||||||
|
return False
|
||||||
|
except pynvml.NVMLError:
|
||||||
|
logger.exception(
|
||||||
|
"NVLink detection failed. This is normal if your"
|
||||||
|
" machine has no NVLink equipped."
|
||||||
|
)
|
||||||
|
return False
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
@with_nvml_context
|
||||||
|
def is_one_nvlink_clique(
|
||||||
|
group: torch.distributed.ProcessGroup, device: torch.device
|
||||||
|
) -> bool:
|
||||||
|
"""True iff every rank's GPU is in the same NVLink fabric clique (one NVL72 /
|
||||||
|
MNNVL domain). Such a clique shares a single NVLink address space even across
|
||||||
|
nodes, so custom-AR v2's symm-mem storage + fabric peer VAs are valid group-wide."""
|
||||||
|
if _is_hip:
|
||||||
|
return False
|
||||||
|
try:
|
||||||
|
clique = _gpu_fabric_clique(device)
|
||||||
|
except Exception as e:
|
||||||
|
logger.warning(
|
||||||
|
"GPU fabric clique query failed (%r); custom-AR stays intra-node.", e
|
||||||
|
)
|
||||||
|
clique = None
|
||||||
|
# Always all-gather (every rank calls it once) so a failed query on any rank
|
||||||
|
# resolves to a clean False rather than a collective mismatch.
|
||||||
|
world_size = dist.get_world_size(group=group)
|
||||||
|
gathered: List[object] = [None] * world_size
|
||||||
|
dist.all_gather_object(gathered, clique, group=group)
|
||||||
|
if any(c is None for c in gathered):
|
||||||
|
return False
|
||||||
|
return len(set(gathered)) == 1
|
||||||
|
|
||||||
|
|
||||||
|
def is_weak_contiguous(inp: torch.Tensor):
|
||||||
|
return inp.is_contiguous() or (
|
||||||
|
inp.storage().nbytes() - inp.storage_offset() * inp.element_size()
|
||||||
|
== inp.numel() * inp.element_size()
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def can_p2p(rank: int, world_size: int) -> bool:
|
||||||
|
# SGLANG_SKIP_P2P_CHECK can be set to False in sglang
|
||||||
|
SGLANG_SKIP_P2P_CHECK = os.getenv("SGLANG_SKIP_P2P_CHECK", "0") == "1"
|
||||||
|
for i in range(world_size):
|
||||||
|
if i == rank:
|
||||||
|
continue
|
||||||
|
if SGLANG_SKIP_P2P_CHECK:
|
||||||
|
logger.info("Skipping P2P check and trusting the driver's P2P report.")
|
||||||
|
return torch.cuda.can_device_access_peer(rank, i)
|
||||||
|
if not gpu_p2p_access_check(rank, i):
|
||||||
|
return False
|
||||||
|
return True
|
||||||
|
|
||||||
|
|
||||||
|
def can_use_custom_all_reduce_with_nvlink(
|
||||||
|
group: torch.distributed.ProcessGroup,
|
||||||
|
device: torch.device,
|
||||||
|
supported_world_size: List[int],
|
||||||
|
cls_name: str,
|
||||||
|
) -> Optional[bool]: # None if fail; otherwise return whether NVLink is available
|
||||||
|
assert (
|
||||||
|
dist.get_backend(group) != dist.Backend.NCCL
|
||||||
|
), f"{cls_name} should be attached to a non-NCCL group."
|
||||||
|
|
||||||
|
rank = dist.get_rank(group=group)
|
||||||
|
world_size = dist.get_world_size(group=group)
|
||||||
|
|
||||||
|
# No need to initialize custom allreduce for single GPU case.
|
||||||
|
if world_size == 1:
|
||||||
|
return
|
||||||
|
|
||||||
|
# No need to initialize custom allreduce for multi-node case.
|
||||||
|
if not all(in_the_same_node_as(group, source_rank=0)):
|
||||||
|
logger.warning(
|
||||||
|
f"{cls_name} is disabled because this process group" " spans across nodes."
|
||||||
|
)
|
||||||
|
return
|
||||||
|
|
||||||
|
# For not supported world size, we disable custom allreduce.
|
||||||
|
if world_size not in supported_world_size:
|
||||||
|
logger.warning(
|
||||||
|
f"{cls_name} is disabled due to an unsupported world"
|
||||||
|
f" size: {world_size}. Supported world sizes: {supported_world_size}. "
|
||||||
|
"To silence this warning, specify disable_custom_all_reduce=True explicitly.",
|
||||||
|
)
|
||||||
|
return
|
||||||
|
|
||||||
|
cuda_visible_devices = os.environ.get("CUDA_VISIBLE_DEVICES", None)
|
||||||
|
if cuda_visible_devices:
|
||||||
|
device_ids = list(map(int, cuda_visible_devices.split(",")))
|
||||||
|
else:
|
||||||
|
device_ids = list(range(torch.cuda.device_count()))
|
||||||
|
physical_device_id = device_ids[device.index]
|
||||||
|
tensor = torch.tensor([physical_device_id], dtype=torch.int, device="cpu")
|
||||||
|
gather_list = [
|
||||||
|
torch.tensor([0], dtype=torch.int, device="cpu") for _ in range(world_size)
|
||||||
|
]
|
||||||
|
dist.all_gather(gather_list, tensor, group=group)
|
||||||
|
physical_device_ids = [int(t) for t in gather_list]
|
||||||
|
full_nvlink = is_full_nvlink(physical_device_ids, world_size)
|
||||||
|
|
||||||
|
# test nvlink first, this will filter out most of the cases
|
||||||
|
# where custom allreduce is not supported
|
||||||
|
# this checks hardware and driver support for NVLink
|
||||||
|
if False:
|
||||||
|
logger.warning(
|
||||||
|
f"{cls_name} is disabled because it's not supported on"
|
||||||
|
" more than two PCIe-only GPUs. To silence this warning, "
|
||||||
|
"specify disable_custom_all_reduce=True explicitly."
|
||||||
|
)
|
||||||
|
return
|
||||||
|
|
||||||
|
# test P2P capability, this checks software/cudaruntime support
|
||||||
|
# this is expensive to compute at the first time
|
||||||
|
# then we cache the result
|
||||||
|
# On AMD GPU, p2p is always enabled between XGMI connected GPUs
|
||||||
|
if not _is_hip and not can_p2p(rank, world_size):
|
||||||
|
logger.warning(
|
||||||
|
f"{cls_name} is disabled because your platform lacks "
|
||||||
|
"GPU P2P capability or P2P test failed. To silence this "
|
||||||
|
"warning, specify disable_custom_all_reduce=True explicitly."
|
||||||
|
)
|
||||||
|
return
|
||||||
|
|
||||||
|
return full_nvlink
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
batch_src, batch_tgt, output_file = pickle.loads(sys.stdin.buffer.read())
|
||||||
|
result = can_actually_p2p(batch_src, batch_tgt)
|
||||||
|
with open(output_file, "wb") as f:
|
||||||
|
f.write(pickle.dumps(result))
|
||||||
@ -0,0 +1,20 @@
|
|||||||
|
{"tag": "b3_16k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9501, "corpus_window": {"start": 2300000, "end": 2431072}, "ok": 8, "failed": 0, "wall_s": 82.58, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 49.6, "input_throughput_tok_s": 1587.15, "ttft_s": {"mean": 3.9252, "p50": 3.9266, "p95": 3.9787, "max": 3.9787, "min": 3.8899}, "tpot_s": {"mean": 0.0125, "p50": 0.0128, "p95": 0.0137, "max": 0.0137, "min": 0.0112}, "e2e_s": {"mean": 10.3227, "p50": 10.4215, "p95": 10.9871, "max": 10.9871, "min": 9.6445}, "per_req_out_tok_s_e2e": {"mean": 49.6617, "p50": 49.8845, "p95": 53.0871, "max": 53.0871, "min": 46.6001}, "per_req_decode_tok_s": {"mean": 80.2896, "p50": 80.4694, "p95": 89.5483, "max": 89.5483, "min": 72.9492}, "spec_accept_length_mean": 2.138, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 16, "new_tokens": 131072, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b3_16k_c8", "arm": "e7b", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9502, "corpus_window": {"start": 2300000, "end": 2562144}, "ok": 16, "failed": 0, "wall_s": 100.63, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 8192, "output_throughput_tok_s": 81.41, "input_throughput_tok_s": 2605.13, "ttft_s": {"mean": 11.7518, "p50": 6.5296, "p95": 32.5233, "max": 32.5233, "min": 3.9869}, "tpot_s": {"mean": 0.0716, "p50": 0.0731, "p95": 0.1351, "max": 0.1351, "min": 0.0266}, "e2e_s": {"mean": 48.3305, "p50": 50.9027, "p95": 78.2404, "max": 78.2404, "min": 17.5929}, "per_req_out_tok_s_e2e": {"mean": 12.9171, "p50": 10.4857, "p95": 29.1027, "max": 29.1027, "min": 6.5439}, "per_req_decode_tok_s": {"mean": 16.8775, "p50": 15.7396, "p95": 37.6553, "max": 37.6553, "min": 7.4162}, "spec_accept_length_mean": 2.375, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 262144, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b3_16k_c16", "arm": "e7b", "summary": {"concurrency": 16, "num_requests": 32, "run_id": 9503, "corpus_window": {"start": 2300000, "end": 2824288}, "ok": 32, "failed": 0, "wall_s": 236.44, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 16384, "output_throughput_tok_s": 69.29, "input_throughput_tok_s": 2217.4, "ttft_s": {"mean": 23.5964, "p50": 13.173, "p95": 60.864, "max": 90.3879, "min": 4.2969}, "tpot_s": {"mean": 0.1733, "p50": 0.1784, "p95": 0.2978, "max": 0.3055, "min": 0.0668}, "e2e_s": {"mean": 112.1663, "p50": 105.5499, "p95": 170.4271, "max": 183.3467, "min": 52.7259}, "per_req_out_tok_s_e2e": {"mean": 5.0312, "p50": 4.8797, "p95": 8.0542, "max": 9.7106, "min": 2.7925}, "per_req_decode_tok_s": {"mean": 6.7269, "p50": 5.8271, "p95": 12.9785, "max": 15.0103, "min": 3.2802}, "spec_accept_length_mean": 2.517, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b3_16k_c32", "arm": "e7b", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9504, "corpus_window": {"start": 2300000, "end": 3348576}, "ok": 64, "failed": 0, "wall_s": 414.96, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 32768, "output_throughput_tok_s": 78.97, "input_throughput_tok_s": 2526.91, "ttft_s": {"mean": 100.2552, "p50": 109.5841, "p95": 169.6812, "max": 190.7286, "min": 5.2261}, "tpot_s": {"mean": 0.1698, "p50": 0.1728, "p95": 0.2599, "max": 0.3191, "min": 0.0328}, "e2e_s": {"mean": 187.0411, "p50": 190.6599, "p95": 271.5134, "max": 292.7775, "min": 86.5615}, "per_req_out_tok_s_e2e": {"mean": 2.9412, "p50": 2.6906, "p95": 4.5951, "max": 5.9149, "min": 1.7488}, "per_req_decode_tok_s": {"mean": 7.1251, "p50": 5.846, "p95": 15.414, "max": 30.5614, "min": 3.1404}, "spec_accept_length_mean": 2.767, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b3_16k_c64", "arm": "e7b", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9505, "corpus_window": {"start": 2300000, "end": 4397152}, "ok": 128, "failed": 0, "wall_s": 777.55, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 65536, "output_throughput_tok_s": 84.29, "input_throughput_tok_s": 2697.14, "ttft_s": {"mean": 240.0965, "p50": 287.3833, "p95": 345.6933, "max": 380.9581, "min": 5.2288}, "tpot_s": {"mean": 0.168, "p50": 0.1702, "p95": 0.2315, "max": 0.3119, "min": 0.0404}, "e2e_s": {"mean": 325.934, "p50": 364.8118, "p95": 432.439, "max": 474.7226, "min": 83.7008}, "per_req_out_tok_s_e2e": {"mean": 1.8253, "p50": 1.4064, "p95": 4.196, "max": 6.117, "min": 1.0785}, "per_req_decode_tok_s": {"mean": 6.6174, "p50": 5.8876, "p95": 12.5326, "max": 24.7977, "min": 3.2126}, "spec_accept_length_mean": 2.986, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b41_1k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9511, "corpus_window": {"start": 4500000, "end": 4508192}, "ok": 8, "failed": 0, "wall_s": 15.82, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 1024, "output_throughput_tok_s": 64.74, "input_throughput_tok_s": 517.93, "ttft_s": {"mean": 0.3034, "p50": 0.3015, "p95": 0.3258, "max": 0.3258, "min": 0.2959}, "tpot_s": {"mean": 0.0132, "p50": 0.0135, "p95": 0.0143, "max": 0.0143, "min": 0.0119}, "e2e_s": {"mean": 1.9769, "p50": 2.0115, "p95": 2.1177, "max": 2.1177, "min": 1.8091}, "per_req_out_tok_s_e2e": {"mean": 64.9178, "p50": 66.8893, "p95": 70.753, "max": 70.753, "min": 60.442}, "per_req_decode_tok_s": {"mean": 76.7983, "p50": 79.4095, "p95": 84.9204, "max": 84.9204, "min": 70.4319}, "spec_accept_length_mean": 2.028, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 8, "new_tokens": 8192, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b41_1k_c8", "arm": "e7b", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9512, "corpus_window": {"start": 4500000, "end": 4516384}, "ok": 16, "failed": 0, "wall_s": 13.44, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 2048, "output_throughput_tok_s": 152.4, "input_throughput_tok_s": 1219.24, "ttft_s": {"mean": 1.21, "p50": 1.7505, "p95": 2.1601, "max": 2.1601, "min": 0.387}, "tpot_s": {"mean": 0.0414, "p50": 0.042, "p95": 0.0537, "max": 0.0537, "min": 0.0319}, "e2e_s": {"mean": 6.4682, "p50": 6.5412, "p95": 8.9777, "max": 8.9777, "min": 4.4513}, "per_req_out_tok_s_e2e": {"mean": 20.4798, "p50": 20.2955, "p95": 28.7556, "max": 28.7556, "min": 14.2576}, "per_req_decode_tok_s": {"mean": 24.8432, "p50": 24.8024, "p95": 31.5592, "max": 31.5592, "min": 18.7758}, "spec_accept_length_mean": 2.048, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 10, "new_tokens": 16384, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b41_1k_c32", "arm": "e7b", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9513, "corpus_window": {"start": 4500000, "end": 4565536}, "ok": 64, "failed": 0, "wall_s": 81.8, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 8192, "output_throughput_tok_s": 100.14, "input_throughput_tok_s": 801.15, "ttft_s": {"mean": 17.5622, "p50": 21.7567, "p95": 25.6003, "max": 27.8932, "min": 1.5226}, "tpot_s": {"mean": 0.1466, "p50": 0.1501, "p95": 0.1905, "max": 0.1952, "min": 0.088}, "e2e_s": {"mean": 36.1812, "p50": 40.1766, "p95": 46.8967, "max": 47.1947, "min": 18.2249}, "per_req_out_tok_s_e2e": {"mean": 3.8358, "p50": 3.1917, "p95": 6.6145, "max": 7.0234, "min": 2.7122}, "per_req_decode_tok_s": {"mean": 7.1041, "p50": 6.7208, "p95": 10.1791, "max": 11.4499, "min": 5.164}, "spec_accept_length_mean": 2.054, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 38, "new_tokens": 65536, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b41_1k_c64", "arm": "e7b", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9514, "corpus_window": {"start": 4500000, "end": 4631072}, "ok": 128, "failed": 0, "wall_s": 158.73, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 16384, "output_throughput_tok_s": 103.22, "input_throughput_tok_s": 825.75, "ttft_s": {"mean": 46.7369, "p50": 57.65, "p95": 65.8132, "max": 70.7061, "min": 1.5163}, "tpot_s": {"mean": 0.1445, "p50": 0.144, "p95": 0.1777, "max": 0.2125, "min": 0.0759}, "e2e_s": {"mean": 65.0893, "p50": 74.2949, "p95": 84.7332, "max": 89.7485, "min": 18.3747}, "per_req_out_tok_s_e2e": {"mean": 2.4029, "p50": 1.7238, "p95": 6.272, "max": 6.9661, "min": 1.4262}, "per_req_decode_tok_s": {"mean": 7.1513, "p50": 7.0024, "p95": 9.0245, "max": 13.2811, "min": 4.7422}, "spec_accept_length_mean": 2.087, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 85, "new_tokens": 131072, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b42_1k4k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9521, "corpus_window": {"start": 4700000, "end": 4708192}, "ok": 8, "failed": 0, "wall_s": 242.18, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 32768, "output_throughput_tok_s": 135.31, "input_throughput_tok_s": 33.83, "ttft_s": {"mean": 0.3142, "p50": 0.3133, "p95": 0.3204, "max": 0.3204, "min": 0.3097}, "tpot_s": {"mean": 0.0073, "p50": 0.0078, "p95": 0.0087, "max": 0.0087, "min": 0.0059}, "e2e_s": {"mean": 30.2716, "p50": 32.1753, "p95": 35.9141, "max": 35.9141, "min": 24.509}, "per_req_out_tok_s_e2e": {"mean": 137.6058, "p50": 136.5088, "p95": 167.1221, "max": 167.1221, "min": 114.0498}, "per_req_decode_tok_s": {"mean": 139.1062, "p50": 137.9513, "p95": 169.3421, "max": 169.3421, "min": 115.0748}, "spec_accept_length_mean": 3.701, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 8, "new_tokens": 8192, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b42_1k4k_c8", "arm": "e7b", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9522, "corpus_window": {"start": 4700000, "end": 4716384}, "ok": 16, "failed": 0, "wall_s": 159.24, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 65536, "output_throughput_tok_s": 411.54, "input_throughput_tok_s": 102.89, "ttft_s": {"mean": 0.9638, "p50": 1.2709, "p95": 1.6805, "max": 1.6805, "min": 0.378}, "tpot_s": {"mean": 0.0181, "p50": 0.0175, "p95": 0.0225, "max": 0.0225, "min": 0.0153}, "e2e_s": {"mean": 75.1101, "p50": 73.2342, "p95": 93.9711, "max": 93.9711, "min": 64.3187}, "per_req_out_tok_s_e2e": {"mean": 55.3003, "p50": 58.2448, "p95": 63.6829, "max": 63.6829, "min": 43.5879}, "per_req_decode_tok_s": {"mean": 55.9977, "p50": 58.5719, "p95": 65.3811, "max": 65.3811, "min": 44.3769}, "spec_accept_length_mean": 4.003, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 10, "new_tokens": 16384, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b42_1k4k_c32", "arm": "e7b", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9523, "corpus_window": {"start": 4700000, "end": 4765536}, "ok": 64, "failed": 0, "wall_s": 1040.79, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 262144, "output_throughput_tok_s": 251.87, "input_throughput_tok_s": 62.97, "ttft_s": {"mean": 195.07, "p50": 240.7323, "p95": 304.7528, "max": 334.6933, "min": 1.3472}, "tpot_s": {"mean": 0.0607, "p50": 0.0581, "p95": 0.0746, "max": 0.0923, "min": 0.0465}, "e2e_s": {"mean": 443.6795, "p50": 478.2794, "p95": 587.5712, "max": 597.8883, "min": 219.3532}, "per_req_out_tok_s_e2e": {"mean": 10.0455, "p50": 8.6106, "p95": 17.4441, "max": 18.6731, "min": 6.8508}, "per_req_decode_tok_s": {"mean": 16.784, "p50": 17.4276, "p95": 20.0358, "max": 21.5315, "min": 10.8381}, "spec_accept_length_mean": 3.852, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 51, "new_tokens": 65536, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b51_64k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9531, "corpus_window": {"start": 5000000, "end": 5524288}, "ok": 8, "failed": 0, "wall_s": 182.23, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 22.48, "input_throughput_tok_s": 2877.06, "ttft_s": {"mean": 17.0384, "p50": 17.022, "p95": 17.1703, "max": 17.1703, "min": 17.0028}, "tpot_s": {"mean": 0.0112, "p50": 0.0117, "p95": 0.0135, "max": 0.0135, "min": 0.008}, "e2e_s": {"mean": 22.7786, "p50": 23.0017, "p95": 23.9301, "max": 23.9301, "min": 21.118}, "per_req_out_tok_s_e2e": {"mean": 22.5108, "p50": 22.5486, "p95": 24.2447, "max": 24.2447, "min": 21.3957}, "per_req_decode_tok_s": {"mean": 91.4983, "p50": 90.0463, "p95": 125.007, "max": 125.007, "min": 74.0091}, "spec_accept_length_mean": 2.468, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b51_64k_c4", "arm": "e7b", "summary": {"concurrency": 4, "num_requests": 8, "run_id": 9532, "corpus_window": {"start": 5000000, "end": 5524288}, "ok": 8, "failed": 0, "wall_s": 159.24, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 25.72, "input_throughput_tok_s": 3292.45, "ttft_s": {"mean": 30.5492, "p50": 18.5951, "p95": 68.8675, "max": 68.8675, "min": 17.0513}, "tpot_s": {"mean": 0.0949, "p50": 0.1145, "p95": 0.1546, "max": 0.1546, "min": 0.02}, "e2e_s": {"mean": 79.0616, "p50": 80.1169, "p95": 131.9282, "max": 131.9282, "min": 27.2645}, "per_req_out_tok_s_e2e": {"mean": 8.1791, "p50": 6.6417, "p95": 18.779, "max": 18.779, "min": 3.8809}, "per_req_decode_tok_s": {"mean": 16.1959, "p50": 11.2172, "p95": 50.1581, "max": 50.1581, "min": 6.4814}, "spec_accept_length_mean": 2.438, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b51_64k_c8", "arm": "e7b", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9533, "corpus_window": {"start": 5000000, "end": 6048576}, "ok": 16, "failed": 0, "wall_s": 319.66, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 8192, "output_throughput_tok_s": 25.63, "input_throughput_tok_s": 3280.31, "ttft_s": {"mean": 90.9186, "p50": 96.5085, "p95": 149.4465, "max": 149.4465, "min": 18.5948}, "tpot_s": {"mean": 0.1058, "p50": 0.1221, "p95": 0.1575, "max": 0.1575, "min": 0.0172}, "e2e_s": {"mean": 144.9847, "p50": 155.6435, "p95": 211.845, "max": 211.845, "min": 75.7166}, "per_req_out_tok_s_e2e": {"mean": 3.7942, "p50": 3.6883, "p95": 6.7621, "max": 6.7621, "min": 2.4169}, "per_req_decode_tok_s": {"mean": 13.0455, "p50": 8.4248, "p95": 58.0944, "max": 58.0944, "min": 6.3604}, "spec_accept_length_mean": 2.442, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b51_128k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9541, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 343.36, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 11.93, "input_throughput_tok_s": 3053.83, "ttft_s": {"mean": 37.5917, "p50": 37.5968, "p95": 37.6012, "max": 37.6012, "min": 37.5564}, "tpot_s": {"mean": 0.0104, "p50": 0.0104, "p95": 0.0123, "max": 0.0123, "min": 0.009}, "e2e_s": {"mean": 42.9201, "p50": 42.9071, "p95": 43.9018, "max": 43.9018, "min": 42.1336}, "per_req_out_tok_s_e2e": {"mean": 11.9312, "p50": 11.9694, "p95": 12.1518, "max": 12.1518, "min": 11.6624}, "per_req_decode_tok_s": {"mean": 97.1304, "p50": 98.8664, "p95": 111.8659, "max": 111.8659, "min": 81.2653}, "spec_accept_length_mean": 2.728, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b51_128k_c2", "arm": "e7b", "summary": {"concurrency": 2, "num_requests": 8, "run_id": 9542, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 333.36, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 12.29, "input_throughput_tok_s": 3145.44, "ttft_s": {"mean": 48.2081, "p50": 39.1403, "p95": 75.8082, "max": 75.8082, "min": 37.5834}, "tpot_s": {"mean": 0.068, "p50": 0.0858, "p95": 0.1622, "max": 0.1622, "min": 0.0099}, "e2e_s": {"mean": 82.9494, "p50": 84.1258, "p95": 122.0311, "max": 122.0311, "min": 42.6804}, "per_req_out_tok_s_e2e": {"mean": 7.0763, "p50": 6.2562, "p95": 11.9961, "max": 11.9961, "min": 4.1957}, "per_req_decode_tok_s": {"mean": 37.0402, "p50": 12.1918, "p95": 100.963, "max": 100.963, "min": 6.1768}, "spec_accept_length_mean": 2.769, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b51_128k_c4", "arm": "e7b", "summary": {"concurrency": 4, "num_requests": 8, "run_id": 9543, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 334.71, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 12.24, "input_throughput_tok_s": 3132.76, "ttft_s": {"mean": 110.1385, "p50": 120.8044, "p95": 159.4963, "max": 159.4963, "min": 39.1664}, "tpot_s": {"mean": 0.0687, "p50": 0.0152, "p95": 0.1634, "max": 0.1634, "min": 0.0097}, "e2e_s": {"mean": 145.2581, "p50": 130.4964, "p95": 204.1758, "max": 204.1758, "min": 83.563}, "per_req_out_tok_s_e2e": {"mean": 3.8065, "p50": 4.04, "p95": 6.1271, "max": 6.1271, "min": 2.5076}, "per_req_decode_tok_s": {"mean": 54.2853, "p50": 73.1655, "p95": 103.6085, "max": 103.6085, "min": 6.1313}, "spec_accept_length_mean": 2.595, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b52_256k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9551, "corpus_window": {"start": 8400000, "end": 8662144}, "ok": 1, "failed": 0, "wall_s": 0.5, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 2.01, "input_throughput_tok_s": 527883.82, "ttft_s": {"mean": 0.4955, "p50": 0.4955, "p95": 0.4955, "max": 0.4955, "min": 0.4955}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 0.4957, "p50": 0.4957, "p95": 0.4957, "max": 0.4957, "min": 0.4957}, "per_req_out_tok_s_e2e": {"mean": 2.0172, "p50": 2.0172, "p95": 2.0172, "max": 2.0172, "min": 2.0172}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 1, "new_tokens": 64, "cached_tokens": 262080, "hit_rate": 0.9998}}}
|
||||||
|
{"tag": "b52_256k_c1", "arm": "e7b", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9551, "corpus_window": {"start": 9900000, "end": 10162144}, "ok": 1, "failed": 0, "wall_s": 91.45, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 0.01, "input_throughput_tok_s": 2866.59, "ttft_s": {"mean": 91.4465, "p50": 91.4465, "p95": 91.4465, "max": 91.4465, "min": 91.4465}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 91.4468, "p50": 91.4468, "p95": 91.4468, "max": 91.4468, "min": 91.4468}, "per_req_out_tok_s_e2e": {"mean": 0.0109, "p50": 0.0109, "p95": 0.0109, "max": 0.0109, "min": 0.0109}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 262144, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
@ -0,0 +1,3 @@
|
|||||||
|
a8f3909211adb92501cf007ab71d3b29 /root/corpus_ids.json
|
||||||
|
1e34dd8d813ccb7ac8243e8d1573a49e /root/bench_corpus_v2.py
|
||||||
|
4c126d067d33b5ea27c268f37561634c /root/extract_summary.py
|
||||||
File diff suppressed because one or more lines are too long
@ -0,0 +1,20 @@
|
|||||||
|
b3_16k_c1 OK
|
||||||
|
b3_16k_c8 OK
|
||||||
|
b3_16k_c16 OK
|
||||||
|
b3_16k_c32 OK
|
||||||
|
b3_16k_c64 OK
|
||||||
|
b41_1k_c1 OK
|
||||||
|
b41_1k_c8 OK
|
||||||
|
b41_1k_c32 OK
|
||||||
|
b41_1k_c64 OK
|
||||||
|
b42_1k4k_c1 OK
|
||||||
|
b42_1k4k_c8 OK
|
||||||
|
b42_1k4k_c32 OK
|
||||||
|
b42_1k4k_c64 BENCH_FAIL rc=143
|
||||||
|
b51_64k_c1 OK
|
||||||
|
b51_64k_c4 OK
|
||||||
|
b51_64k_c8 OK
|
||||||
|
b51_128k_c1 OK
|
||||||
|
b51_128k_c2 OK
|
||||||
|
b51_128k_c4 OK
|
||||||
|
b52_256k_c1 HIT_FAIL_FINAL
|
||||||
@ -0,0 +1,7 @@
|
|||||||
|
1e34dd8d813ccb7ac8243e8d1573a49e bench_corpus_v2.py
|
||||||
|
4c126d067d33b5ea27c268f37561634c extract_summary.py
|
||||||
|
5892b44610b2ce61f533f1721105625e run_b300_matrix.sh
|
||||||
|
def3c64c5dc19e1d507080e3366c3762 deploy_glm53_pp4.sh
|
||||||
|
21db641f1d997e26fcf1ddc97163b87e deploy_glm53_607_exp.sh
|
||||||
|
a8fc9a504c041acc40831a848b4985d8 patches/custom_all_reduce.py
|
||||||
|
65a4d22bd87ae343e7d68c4879c14f98 patches/custom_all_reduce_utils.py
|
||||||
@ -0,0 +1,24 @@
|
|||||||
|
{"tag": "b3_16k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9501, "corpus_window": {"start": 2300000, "end": 2431072}, "ok": 8, "failed": 0, "wall_s": 245.55, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 16.68, "input_throughput_tok_s": 533.8, "ttft_s": {"mean": 5.3493, "p50": 5.3474, "p95": 5.3628, "max": 5.3628, "min": 5.3438}, "tpot_s": {"mean": 0.0496, "p50": 0.0495, "p95": 0.0503, "max": 0.0503, "min": 0.0493}, "e2e_s": {"mean": 30.6931, "p50": 30.657, "p95": 31.0364, "max": 31.0364, "min": 30.5593}, "per_req_out_tok_s_e2e": {"mean": 16.6817, "p50": 16.7233, "p95": 16.7543, "max": 16.7543, "min": 16.4968}, "per_req_decode_tok_s": {"mean": 20.2031, "p50": 20.2689, "p95": 20.3081, "max": 20.3081, "min": 19.9299}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b3_16k_c8", "arm": "tp2pp4", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9502, "corpus_window": {"start": 2300000, "end": 2562144}, "ok": 16, "failed": 0, "wall_s": 81.0, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 8192, "output_throughput_tok_s": 101.13, "input_throughput_tok_s": 3236.22, "ttft_s": {"mean": 10.0539, "p50": 10.6041, "p95": 14.7948, "max": 14.7948, "min": 5.2557}, "tpot_s": {"mean": 0.0595, "p50": 0.0602, "p95": 0.0709, "max": 0.0709, "min": 0.0481}, "e2e_s": {"mean": 40.4682, "p50": 41.4728, "p95": 41.6589, "max": 41.6589, "min": 39.2731}, "per_req_out_tok_s_e2e": {"mean": 12.662, "p50": 13.0063, "p95": 13.0369, "max": 13.0369, "min": 12.2903}, "per_req_decode_tok_s": {"mean": 17.0379, "p50": 17.0019, "p95": 20.8395, "max": 20.8395, "min": 14.1226}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b3_16k_c16", "arm": "tp2pp4", "summary": {"concurrency": 16, "num_requests": 32, "run_id": 9503, "corpus_window": {"start": 2300000, "end": 2824288}, "ok": 32, "failed": 0, "wall_s": 112.16, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 16384, "output_throughput_tok_s": 146.08, "input_throughput_tok_s": 4674.5, "ttft_s": {"mean": 15.5283, "p50": 16.1378, "p95": 25.7815, "max": 25.8389, "min": 5.267}, "tpot_s": {"mean": 0.0792, "p50": 0.0797, "p95": 0.0985, "max": 0.1012, "min": 0.0571}, "e2e_s": {"mean": 55.9934, "p50": 56.8023, "p95": 57.0853, "max": 57.1347, "min": 54.9678}, "per_req_out_tok_s_e2e": {"mean": 9.1467, "p50": 9.2952, "p95": 9.3139, "max": 9.3145, "min": 8.9613}, "per_req_decode_tok_s": {"mean": 12.9884, "p50": 12.7194, "p95": 16.7725, "max": 17.5544, "min": 9.9045}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b3_16k_c32", "arm": "tp2pp4", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9504, "corpus_window": {"start": 2300000, "end": 3348576}, "ok": 64, "failed": 0, "wall_s": 169.54, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 32768, "output_throughput_tok_s": 193.27, "input_throughput_tok_s": 6184.73, "ttft_s": {"mean": 26.5164, "p50": 27.1445, "p95": 46.4256, "max": 47.8906, "min": 5.2618}, "tpot_s": {"mean": 0.1137, "p50": 0.114, "p95": 0.152, "max": 0.1573, "min": 0.07}, "e2e_s": {"mean": 84.6344, "p50": 85.322, "p95": 85.7436, "max": 85.8459, "min": 83.6335}, "per_req_out_tok_s_e2e": {"mean": 6.0502, "p50": 6.108, "p95": 6.1205, "max": 6.1219, "min": 5.9642}, "per_req_decode_tok_s": {"mean": 9.2806, "p50": 8.8387, "p95": 13.3026, "max": 14.3169, "min": 6.3708}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b3_16k_c64", "arm": "tp2pp4", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9505, "corpus_window": {"start": 2300000, "end": 4397152}, "ok": 128, "failed": 0, "wall_s": 310.79, "input_len": 16384, "shared_len": 0, "unique_len": 16384, "output_len": 512, "output_tokens_total": 65536, "output_throughput_tok_s": 210.87, "input_throughput_tok_s": 6747.91, "ttft_s": {"mean": 63.0522, "p50": 49.1362, "p95": 134.7794, "max": 139.9931, "min": 5.3652}, "tpot_s": {"mean": 0.1392, "p50": 0.1358, "p95": 0.2037, "max": 0.2143, "min": 0.0707}, "e2e_s": {"mean": 134.1654, "p50": 114.4018, "p95": 226.099, "max": 226.1544, "min": 84.0644}, "per_req_out_tok_s_e2e": {"mean": 4.1916, "p50": 4.476, "p95": 6.0862, "max": 6.0906, "min": 2.2639}, "per_req_decode_tok_s": {"mean": 7.7966, "p50": 7.3815, "p95": 11.9051, "max": 14.1753, "min": 4.6761}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 512, "new_tokens": 8388608, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b41_1k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9511, "corpus_window": {"start": 4500000, "end": 4508192}, "ok": 8, "failed": 0, "wall_s": 52.95, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 1024, "output_throughput_tok_s": 19.34, "input_throughput_tok_s": 154.7, "ttft_s": {"mean": 0.3848, "p50": 0.3852, "p95": 0.3865, "max": 0.3865, "min": 0.3825}, "tpot_s": {"mean": 0.0491, "p50": 0.0491, "p95": 0.0493, "max": 0.0493, "min": 0.0488}, "e2e_s": {"mean": 6.619, "p50": 6.62, "p95": 6.6447, "max": 6.6447, "min": 6.5863}, "per_req_out_tok_s_e2e": {"mean": 19.3385, "p50": 19.3379, "p95": 19.4343, "max": 19.4343, "min": 19.2635}, "per_req_decode_tok_s": {"mean": 20.5327, "p50": 20.5305, "p95": 20.6346, "max": 20.6346, "min": 20.4442}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 32768, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b41_1k_c8", "arm": "tp2pp4", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9512, "corpus_window": {"start": 4500000, "end": 4516384}, "ok": 16, "failed": 0, "wall_s": 20.1, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 2048, "output_throughput_tok_s": 101.91, "input_throughput_tok_s": 815.3, "ttft_s": {"mean": 1.3131, "p50": 1.2998, "p95": 1.6901, "max": 1.6901, "min": 0.6481}, "tpot_s": {"mean": 0.0686, "p50": 0.0684, "p95": 0.0729, "max": 0.0729, "min": 0.0669}, "e2e_s": {"mean": 10.024, "p50": 10.1096, "p95": 10.2887, "max": 10.2887, "min": 9.7907}, "per_req_out_tok_s_e2e": {"mean": 12.7741, "p50": 12.9718, "p95": 13.0736, "max": 13.0736, "min": 12.4409}, "per_req_decode_tok_s": {"mean": 14.7032, "p50": 14.7329, "p95": 15.0683, "max": 15.0683, "min": 13.8323}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 24, "new_tokens": 65536, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b41_1k_c32", "arm": "tp2pp4", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9513, "corpus_window": {"start": 4500000, "end": 4565536}, "ok": 64, "failed": 0, "wall_s": 29.73, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 8192, "output_throughput_tok_s": 275.55, "input_throughput_tok_s": 2204.4, "ttft_s": {"mean": 3.8882, "p50": 3.9615, "p95": 4.8674, "max": 4.8759, "min": 1.5565}, "tpot_s": {"mean": 0.0856, "p50": 0.0869, "p95": 0.0985, "max": 0.1068, "min": 0.0785}, "e2e_s": {"mean": 14.758, "p50": 14.8384, "p95": 15.0559, "max": 15.1157, "min": 14.5917}, "per_req_out_tok_s_e2e": {"mean": 8.6742, "p50": 8.7413, "p95": 8.7701, "max": 8.7721, "min": 8.468}, "per_req_decode_tok_s": {"mean": 11.8377, "p50": 11.6015, "p95": 12.822, "max": 12.8317, "min": 9.4403}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 36, "new_tokens": 262144, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b41_1k_c64", "arm": "tp2pp4", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9514, "corpus_window": {"start": 4500000, "end": 4631072}, "ok": 128, "failed": 0, "wall_s": 48.11, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 128, "output_tokens_total": 16384, "output_throughput_tok_s": 340.53, "input_throughput_tok_s": 2724.25, "ttft_s": {"mean": 8.1393, "p50": 4.8485, "p95": 19.9553, "max": 20.0096, "min": 0.3836}, "tpot_s": {"mean": 0.0933, "p50": 0.0924, "p95": 0.1018, "max": 0.1251, "min": 0.0818}, "e2e_s": {"mean": 19.9938, "p50": 16.0958, "p95": 32.0105, "max": 32.0464, "min": 15.7372}, "per_req_out_tok_s_e2e": {"mean": 6.9903, "p50": 7.9525, "p95": 8.1268, "max": 8.1336, "min": 3.9942}, "per_req_decode_tok_s": {"mean": 10.8849, "p50": 10.9086, "p95": 12.2874, "max": 12.3245, "min": 8.0552}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 52, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b42_1k4k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9521, "corpus_window": {"start": 4700000, "end": 4708192}, "ok": 8, "failed": 0, "wall_s": 1614.16, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 32768, "output_throughput_tok_s": 20.3, "input_throughput_tok_s": 5.08, "ttft_s": {"mean": 0.3926, "p50": 0.394, "p95": 0.3952, "max": 0.3952, "min": 0.3852}, "tpot_s": {"mean": 0.0492, "p50": 0.0491, "p95": 0.0498, "max": 0.0498, "min": 0.049}, "e2e_s": {"mean": 201.7695, "p50": 201.6506, "p95": 204.4013, "max": 204.4013, "min": 200.9309}, "per_req_out_tok_s_e2e": {"mean": 20.3009, "p50": 20.3439, "p95": 20.3851, "max": 20.3851, "min": 20.039}, "per_req_decode_tok_s": {"mean": 20.3407, "p50": 20.3837, "p95": 20.4254, "max": 20.4254, "min": 20.078}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 32, "new_tokens": 32768, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b42_1k4k_c8", "arm": "tp2pp4", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9522, "corpus_window": {"start": 4700000, "end": 4716384}, "ok": 16, "failed": 0, "wall_s": 649.69, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 65536, "output_throughput_tok_s": 100.87, "input_throughput_tok_s": 25.22, "ttft_s": {"mean": 1.352, "p50": 1.4385, "p95": 1.8726, "max": 1.8726, "min": 0.4364}, "tpot_s": {"mean": 0.079, "p50": 0.0791, "p95": 0.0794, "max": 0.0794, "min": 0.0788}, "e2e_s": {"mean": 324.8285, "p50": 325.5653, "p95": 325.6331, "max": 325.6331, "min": 324.0447}, "per_req_out_tok_s_e2e": {"mean": 12.6098, "p50": 12.6389, "p95": 12.6402, "max": 12.6402, "min": 12.5786}, "per_req_decode_tok_s": {"mean": 12.6626, "p50": 12.6734, "p95": 12.7003, "max": 12.7003, "min": 12.5955}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 20, "new_tokens": 65536, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b42_1k4k_c32", "arm": "tp2pp4", "summary": {"concurrency": 32, "num_requests": 64, "run_id": 9523, "corpus_window": {"start": 4700000, "end": 4765536}, "ok": 64, "failed": 0, "wall_s": 714.14, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 262144, "output_throughput_tok_s": 367.08, "input_throughput_tok_s": 91.77, "ttft_s": {"mean": 1.9772, "p50": 1.1755, "p95": 4.1378, "max": 4.1467, "min": 0.3867}, "tpot_s": {"mean": 0.0861, "p50": 0.0887, "p95": 0.0905, "max": 0.0911, "min": 0.0805}, "e2e_s": {"mean": 354.393, "p50": 367.2171, "p95": 373.3199, "max": 374.0319, "min": 330.2086}, "per_req_out_tok_s_e2e": {"mean": 11.5833, "p50": 11.8878, "p95": 12.2953, "max": 12.4043, "min": 10.9509}, "per_req_decode_tok_s": {"mean": 11.6457, "p50": 11.9132, "p95": 12.3308, "max": 12.4283, "min": 10.9854}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 262144, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b42_1k4k_c64", "arm": "tp2pp4", "summary": {"concurrency": 64, "num_requests": 128, "run_id": 9524, "corpus_window": {"start": 4700000, "end": 4831072}, "ok": 128, "failed": 0, "wall_s": 1088.63, "input_len": 1024, "shared_len": 0, "unique_len": 1024, "output_len": 4096, "output_tokens_total": 524288, "output_throughput_tok_s": 481.6, "input_throughput_tok_s": 120.4, "ttft_s": {"mean": 80.7415, "p50": 4.7569, "p95": 337.9195, "max": 338.4726, "min": 0.3985}, "tpot_s": {"mean": 0.0912, "p50": 0.0914, "p95": 0.0955, "max": 0.0963, "min": 0.0852}, "e2e_s": {"mean": 454.4024, "p50": 385.4538, "p95": 718.7004, "max": 732.2026, "min": 351.072}, "per_req_out_tok_s_e2e": {"mean": 9.6298, "p50": 10.6531, "p95": 11.4339, "max": 11.6671, "min": 5.5941}, "per_req_decode_tok_s": {"mean": 10.9736, "p50": 10.9486, "p95": 11.4764, "max": 11.7362, "min": 10.3887}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 200, "new_tokens": 524288, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b51_64k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9531, "corpus_window": {"start": 5000000, "end": 5524288}, "ok": 8, "failed": 0, "wall_s": 285.91, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 14.33, "input_throughput_tok_s": 1833.72, "ttft_s": {"mean": 10.2811, "p50": 10.2794, "p95": 10.3216, "max": 10.3216, "min": 10.2634}, "tpot_s": {"mean": 0.0498, "p50": 0.0499, "p95": 0.05, "max": 0.05, "min": 0.0496}, "e2e_s": {"mean": 35.739, "p50": 35.7693, "p95": 35.838, "max": 35.838, "min": 35.6073}, "per_req_out_tok_s_e2e": {"mean": 14.3261, "p50": 14.3345, "p95": 14.3791, "max": 14.3791, "min": 14.2865}, "per_req_decode_tok_s": {"mean": 20.112, "p50": 20.1156, "p95": 20.2096, "max": 20.2096, "min": 20.052}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b51_64k_c4", "arm": "tp2pp4", "summary": {"concurrency": 4, "num_requests": 8, "run_id": 9532, "corpus_window": {"start": 5000000, "end": 5524288}, "ok": 8, "failed": 0, "wall_s": 122.3, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 33.49, "input_throughput_tok_s": 4286.94, "ttft_s": {"mean": 19.5548, "p50": 22.5526, "p95": 28.8262, "max": 28.8262, "min": 10.2823}, "tpot_s": {"mean": 0.0814, "p50": 0.0873, "p95": 0.0995, "max": 0.0995, "min": 0.0632}, "e2e_s": {"mean": 61.1314, "p50": 61.2078, "p95": 61.2548, "max": 61.2548, "min": 61.002}, "per_req_out_tok_s_e2e": {"mean": 8.3754, "p50": 8.3851, "p95": 8.3932, "max": 8.3932, "min": 8.3585}, "per_req_decode_tok_s": {"mean": 12.6679, "p50": 13.2834, "p95": 15.8466, "max": 15.8466, "min": 10.0675}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b51_64k_c8", "arm": "tp2pp4", "summary": {"concurrency": 8, "num_requests": 16, "run_id": 9533, "corpus_window": {"start": 5000000, "end": 6048576}, "ok": 16, "failed": 0, "wall_s": 185.59, "input_len": 65536, "shared_len": 0, "unique_len": 65536, "output_len": 512, "output_tokens_total": 8192, "output_throughput_tok_s": 44.14, "input_throughput_tok_s": 5650.11, "ttft_s": {"mean": 31.833, "p50": 34.8143, "p95": 53.3841, "max": 53.3841, "min": 10.2802}, "tpot_s": {"mean": 0.1192, "p50": 0.1246, "p95": 0.1618, "max": 0.1618, "min": 0.0765}, "e2e_s": {"mean": 92.7517, "p50": 92.8482, "p95": 92.9902, "max": 92.9902, "min": 92.4957}, "per_req_out_tok_s_e2e": {"mean": 5.5201, "p50": 5.5226, "p95": 5.5354, "max": 5.5354, "min": 5.506}, "per_req_decode_tok_s": {"mean": 8.9039, "p50": 8.8159, "p95": 13.0909, "max": 13.0909, "min": 6.1915}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b51_128k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 8, "run_id": 9541, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 347.5, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 11.79, "input_throughput_tok_s": 3017.45, "ttft_s": {"mean": 18.2074, "p50": 18.2043, "p95": 18.2252, "max": 18.2252, "min": 18.1922}, "tpot_s": {"mean": 0.0494, "p50": 0.0496, "p95": 0.0497, "max": 0.0497, "min": 0.0485}, "e2e_s": {"mean": 43.4375, "p50": 43.5633, "p95": 43.6252, "max": 43.6252, "min": 43.0007}, "per_req_out_tok_s_e2e": {"mean": 11.7874, "p50": 11.7541, "p95": 11.9068, "max": 11.9068, "min": 11.7363}, "per_req_decode_tok_s": {"mean": 20.2951, "p50": 20.1994, "p95": 20.6444, "max": 20.6444, "min": 20.1578}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b51_128k_c2", "arm": "tp2pp4", "summary": {"concurrency": 2, "num_requests": 8, "run_id": 9542, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 243.3, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 16.84, "input_throughput_tok_s": 4309.84, "ttft_s": {"mean": 25.2658, "p50": 32.1952, "p95": 32.4188, "max": 32.4188, "min": 18.173}, "tpot_s": {"mean": 0.0696, "p50": 0.0831, "p95": 0.0837, "max": 0.0837, "min": 0.0556}, "e2e_s": {"mean": 60.8183, "p50": 60.7483, "p95": 61.2374, "max": 61.2374, "min": 60.6134}, "per_req_out_tok_s_e2e": {"mean": 8.4186, "p50": 8.4352, "p95": 8.447, "max": 8.447, "min": 8.3609}, "per_req_decode_tok_s": {"mean": 14.9825, "p50": 17.8153, "p95": 18.0089, "max": 18.0089, "min": 11.965}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b51_128k_c4", "arm": "tp2pp4", "summary": {"concurrency": 4, "num_requests": 8, "run_id": 9543, "corpus_window": {"start": 6200000, "end": 7248576}, "ok": 8, "failed": 0, "wall_s": 186.37, "input_len": 131072, "shared_len": 0, "unique_len": 131072, "output_len": 512, "output_tokens_total": 4096, "output_throughput_tok_s": 21.98, "input_throughput_tok_s": 5626.35, "ttft_s": {"mean": 39.2939, "p50": 46.1927, "p95": 60.3496, "max": 60.3496, "min": 18.2774}, "tpot_s": {"mean": 0.1054, "p50": 0.1187, "p95": 0.1468, "max": 0.1468, "min": 0.0638}, "e2e_s": {"mean": 93.1534, "p50": 93.2237, "p95": 93.3139, "max": 93.3139, "min": 92.97}, "per_req_out_tok_s_e2e": {"mean": 5.4963, "p50": 5.497, "p95": 5.5072, "max": 5.5072, "min": 5.4869}, "per_req_decode_tok_s": {"mean": 10.4415, "p50": 10.8799, "p95": 15.6959, "max": 15.6959, "min": 6.8234}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 256, "new_tokens": 4194304, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b52_256k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9551, "corpus_window": {"start": 8400000, "end": 8662144}, "ok": 1, "failed": 0, "wall_s": 37.18, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 0.03, "input_throughput_tok_s": 7050.23, "ttft_s": {"mean": 37.1811, "p50": 37.1811, "p95": 37.1811, "max": 37.1811, "min": 37.1811}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 37.1814, "p50": 37.1814, "p95": 37.1814, "max": 37.1814, "min": 37.1814}, "per_req_out_tok_s_e2e": {"mean": 0.0269, "p50": 0.0269, "p95": 0.0269, "max": 0.0269, "min": 0.0269}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 64, "new_tokens": 1048576, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b52_256k_c2", "arm": "tp2pp4", "summary": {"concurrency": 2, "num_requests": 2, "run_id": 9552, "corpus_window": {"start": 8400000, "end": 8924288}, "ok": 2, "failed": 0, "wall_s": 70.19, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 2, "output_throughput_tok_s": 0.03, "input_throughput_tok_s": 7469.22, "ttft_s": {"mean": 53.6945, "p50": 70.1229, "p95": 70.1229, "max": 70.1229, "min": 37.2662}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 53.6948, "p50": 70.1232, "p95": 70.1232, "max": 70.1232, "min": 37.2664}, "per_req_out_tok_s_e2e": {"mean": 0.0205, "p50": 0.0268, "p95": 0.0268, "max": 0.0268, "min": 0.0143}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b52_256k_c3", "arm": "tp2pp4", "summary": {"concurrency": 3, "num_requests": 3, "run_id": 9553, "corpus_window": {"start": 8400000, "end": 9186432}, "ok": 3, "failed": 0, "wall_s": 103.12, "input_len": 262144, "shared_len": 0, "unique_len": 262144, "output_len": 1, "output_tokens_total": 3, "output_throughput_tok_s": 0.03, "input_throughput_tok_s": 7626.22, "ttft_s": {"mean": 70.1194, "p50": 70.1196, "p95": 102.986, "max": 102.986, "min": 37.2526}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 70.1196, "p50": 70.1198, "p95": 102.9863, "max": 102.9863, "min": 37.2529}, "per_req_out_tok_s_e2e": {"mean": 0.0169, "p50": 0.0143, "p95": 0.0268, "max": 0.0268, "min": 0.0097}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 192, "new_tokens": 3145728, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b52_512k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9554, "corpus_window": {"start": 9300000, "end": 9824288}, "ok": 1, "failed": 0, "wall_s": 92.6, "input_len": 524288, "shared_len": 0, "unique_len": 524288, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 0.01, "input_throughput_tok_s": 5662.03, "ttft_s": {"mean": 92.5959, "p50": 92.5959, "p95": 92.5959, "max": 92.5959, "min": 92.5959}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 92.5962, "p50": 92.5962, "p95": 92.5962, "max": 92.5962, "min": 92.5962}, "per_req_out_tok_s_e2e": {"mean": 0.0108, "p50": 0.0108, "p95": 0.0108, "max": 0.0108, "min": 0.0108}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 128, "new_tokens": 2097152, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
|
{"tag": "b52_896k_c1", "arm": "tp2pp4", "summary": {"concurrency": 1, "num_requests": 1, "run_id": 9555, "corpus_window": {"start": 9900000, "end": 10817504}, "ok": 1, "failed": 0, "wall_s": 220.91, "input_len": 917504, "shared_len": 0, "unique_len": 917504, "output_len": 1, "output_tokens_total": 1, "output_throughput_tok_s": 0.0, "input_throughput_tok_s": 4153.24, "ttft_s": {"mean": 220.9117, "p50": 220.9117, "p95": 220.9117, "max": 220.9117, "min": 220.9117}, "tpot_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "e2e_s": {"mean": 220.912, "p50": 220.912, "p95": 220.912, "max": 220.912, "min": 220.912}, "per_req_out_tok_s_e2e": {"mean": 0.0045, "p50": 0.0045, "p95": 0.0045, "max": 0.0045, "min": 0.0045}, "per_req_decode_tok_s": {"mean": null, "p50": null, "p95": null, "max": null, "min": null}, "spec_accept_length_mean": null, "retractions_total": 0, "cache_hit_from_logs": {"prefill_batches": 224, "new_tokens": 3670016, "cached_tokens": 0, "hit_rate": 0.0}}}
|
||||||
@ -0,0 +1,3 @@
|
|||||||
|
a8f3909211adb92501cf007ab71d3b29 /root/corpus_ids.json
|
||||||
|
1e34dd8d813ccb7ac8243e8d1573a49e /root/bench_corpus_v2.py
|
||||||
|
4c126d067d33b5ea27c268f37561634c /root/extract_summary.py
|
||||||
File diff suppressed because one or more lines are too long
@ -0,0 +1,24 @@
|
|||||||
|
b3_16k_c1 OK
|
||||||
|
b3_16k_c8 OK
|
||||||
|
b3_16k_c16 OK
|
||||||
|
b3_16k_c32 OK
|
||||||
|
b3_16k_c64 OK
|
||||||
|
b41_1k_c1 OK
|
||||||
|
b41_1k_c8 OK
|
||||||
|
b41_1k_c32 OK
|
||||||
|
b41_1k_c64 OK
|
||||||
|
b42_1k4k_c1 OK
|
||||||
|
b42_1k4k_c8 OK
|
||||||
|
b42_1k4k_c32 OK
|
||||||
|
b42_1k4k_c64 OK
|
||||||
|
b51_64k_c1 OK
|
||||||
|
b51_64k_c4 OK
|
||||||
|
b51_64k_c8 OK
|
||||||
|
b51_128k_c1 OK
|
||||||
|
b51_128k_c2 OK
|
||||||
|
b51_128k_c4 OK
|
||||||
|
b52_256k_c1 OK
|
||||||
|
b52_256k_c2 OK
|
||||||
|
b52_256k_c3 OK
|
||||||
|
b52_512k_c1 OK
|
||||||
|
b52_896k_c1 OK
|
||||||
@ -0,0 +1,319 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""Real-corpus (PG19) benchmark for sglang GLM-5.3-NVFP4.
|
||||||
|
|
||||||
|
Same methodology as bench_hit90.py (input_ids direct to /generate, temp 0,
|
||||||
|
ignore_eos, stream, server-side completion_tokens counting, hit-rate verified
|
||||||
|
from scheduler logs), but prompts are token slices of REAL book text tokenized
|
||||||
|
with the served model's own tokenizer, replacing random ids.
|
||||||
|
|
||||||
|
Corpus file: JSON {"ids": [flat token ids], "books": [[start, end), ...]}
|
||||||
|
|
||||||
|
Fixed token-offset layout into the flat id array:
|
||||||
|
[0, 117968) shared prefix for 128k points (90% of 131072)
|
||||||
|
[0, 58976) shared prefix for 64k points (same region, shorter cut)
|
||||||
|
pool A [131072, +4*8*13104) 128k unique suffixes, run-ids 9301-9304
|
||||||
|
pool B [550400, +4*8*6560) 64k unique suffixes, run-ids 9305-9308
|
||||||
|
pool C [760320, 262144+524288+524288) 16k fully-unique prompts, run-ids 9311-9313
|
||||||
|
spare [2071040, end) warmup slices / re-run margin
|
||||||
|
|
||||||
|
Each run-id maps to one non-overlapping window (one window = one bench point);
|
||||||
|
re-running a point with fresh text = bump --pool-override past the spare base.
|
||||||
|
|
||||||
|
v2 changes (b300-equivalent campaign, 2026-09-10):
|
||||||
|
- stats(): p95 added (nearest-rank) to match the B300 report metric contract
|
||||||
|
(TTFT P95 / TPOT P95).
|
||||||
|
- --dump-records PATH: per-request records (ttft/e2e/tpot/n_out/retractions/
|
||||||
|
spec_accept_len) written as JSONL for post-hoc percentile checks.
|
||||||
|
- With --shared-frac 0 + --pool-override, any input length is supported
|
||||||
|
(1024 / 16384 / 65536 / 131072 / 262144 / 524288 / 917504).
|
||||||
|
|
||||||
|
Usage:
|
||||||
|
python3 bench_corpus_v2.py --corpus /root/corpus_ids.json --input-len 16384 \
|
||||||
|
--concurrency 64 --num-requests 128 --run-id 9505 --shared-frac 0 \
|
||||||
|
--pool-override 2300000 --output-len 512 \
|
||||||
|
--dump-records /root/bench_logs/xx/point_records.jsonl
|
||||||
|
"""
|
||||||
|
import argparse
|
||||||
|
import datetime
|
||||||
|
import json
|
||||||
|
import math
|
||||||
|
import re
|
||||||
|
import statistics
|
||||||
|
import subprocess
|
||||||
|
import time
|
||||||
|
from concurrent.futures import ThreadPoolExecutor
|
||||||
|
|
||||||
|
import requests
|
||||||
|
|
||||||
|
OUTPUT_LEN_DEFAULT = 512
|
||||||
|
CORPUS_DEFAULT = "/root/corpus_ids.json"
|
||||||
|
|
||||||
|
# fixed pool layout (see docstring)
|
||||||
|
S1_128K_RID0, S1_64K_RID0, S2_RID0 = 9301, 9305, 9311
|
||||||
|
POOL_A_BASE, POOL_A_PER = 131072, 8 * 13104 # 128k suffix windows
|
||||||
|
POOL_B_BASE = POOL_A_BASE + 4 * POOL_A_PER # 550400
|
||||||
|
POOL_B_PER = 8 * 6560 # 64k suffix windows
|
||||||
|
POOL_C_BASE = POOL_B_BASE + 4 * POOL_B_PER # 760320
|
||||||
|
POOL_C_SIZES = [16 * 16384, 32 * 16384, 32 * 16384] # cc8 / cc16 / cc32
|
||||||
|
SPARE_BASE = POOL_C_BASE + sum(POOL_C_SIZES) # 2071040
|
||||||
|
|
||||||
|
sess = requests.Session()
|
||||||
|
sess.trust_env = False # bypass any proxy env on the host
|
||||||
|
|
||||||
|
|
||||||
|
def split_lens(input_len, shared_frac):
|
||||||
|
# unique suffix = (1 - shared_frac) of the prompt, page-16 aligned
|
||||||
|
# (128k @0.9 -> 13104 unique; 16k @0.0 -> fully unique prompts)
|
||||||
|
unique = round(input_len * (1.0 - shared_frac) / 16) * 16
|
||||||
|
return input_len - unique, unique
|
||||||
|
|
||||||
|
|
||||||
|
def pool_start_for(input_len, shared_frac, run_id, override):
|
||||||
|
if override is not None:
|
||||||
|
return override
|
||||||
|
if shared_frac > 0:
|
||||||
|
if input_len == 131072:
|
||||||
|
idx = run_id - S1_128K_RID0
|
||||||
|
if not 0 <= idx < 4:
|
||||||
|
sys_exit_bad_runid(run_id, "128k points use run-ids 9301-9304")
|
||||||
|
return POOL_A_BASE + idx * POOL_A_PER
|
||||||
|
if input_len == 65536:
|
||||||
|
idx = run_id - S1_64K_RID0
|
||||||
|
if not 0 <= idx < 4:
|
||||||
|
sys_exit_bad_runid(run_id, "64k points use run-ids 9305-9308")
|
||||||
|
return POOL_B_BASE + idx * POOL_B_PER
|
||||||
|
sys_exit_bad_runid(run_id, "shared-frac>0 supports 131072/65536 only")
|
||||||
|
idx = run_id - S2_RID0
|
||||||
|
if not 0 <= idx < 3:
|
||||||
|
sys_exit_bad_runid(run_id, "16k unique points use run-ids 9311-9313")
|
||||||
|
return POOL_C_BASE + sum(POOL_C_SIZES[:idx])
|
||||||
|
|
||||||
|
|
||||||
|
def sys_exit_bad_runid(run_id, msg):
|
||||||
|
raise SystemExit(f"[pool] run-id {run_id} outside expected set: {msg}")
|
||||||
|
|
||||||
|
|
||||||
|
def build_prompts(ids, shared_len, unique_len, num_requests, pool_start):
|
||||||
|
if shared_len:
|
||||||
|
shared = ids[0:shared_len]
|
||||||
|
else:
|
||||||
|
shared = []
|
||||||
|
end = pool_start + num_requests * unique_len
|
||||||
|
if end > len(ids):
|
||||||
|
raise SystemExit(
|
||||||
|
f"[pool] window [{pool_start}, {end}) exceeds corpus ({len(ids)} ids); "
|
||||||
|
f"use --pool-override or a larger corpus")
|
||||||
|
prompts = []
|
||||||
|
for i in range(num_requests):
|
||||||
|
s = pool_start + i * unique_len
|
||||||
|
prompts.append(shared + ids[s:s + unique_len])
|
||||||
|
return shared, prompts, (pool_start, end)
|
||||||
|
|
||||||
|
|
||||||
|
def warmup(url, ids, shared_len):
|
||||||
|
# primes the radix cache with the shared prefix (same role as in bench_hit90);
|
||||||
|
# warm slice comes from the spare region so it never collides with a pool window
|
||||||
|
if len(ids) >= SPARE_BASE + 64:
|
||||||
|
warm_slice = ids[SPARE_BASE:SPARE_BASE + 64]
|
||||||
|
else:
|
||||||
|
warm_slice = ids[-64:]
|
||||||
|
payload = {
|
||||||
|
"input_ids": ids[0:shared_len] + warm_slice if shared_len else warm_slice,
|
||||||
|
"sampling_params": {"max_new_tokens": 8, "temperature": 0.0, "ignore_eos": True},
|
||||||
|
}
|
||||||
|
t0 = time.perf_counter()
|
||||||
|
r = sess.post(url, json=payload, timeout=1800)
|
||||||
|
dt = time.perf_counter() - t0
|
||||||
|
print(f"[warmup] http={r.status_code} wall={dt:.2f}s", flush=True)
|
||||||
|
|
||||||
|
|
||||||
|
def bench_one(url, prompt, output_len, idx, results):
|
||||||
|
payload = {
|
||||||
|
"input_ids": prompt,
|
||||||
|
"sampling_params": {"max_new_tokens": output_len, "temperature": 0.0, "ignore_eos": True},
|
||||||
|
"stream": True,
|
||||||
|
}
|
||||||
|
rec = {"idx": idx}
|
||||||
|
t0 = time.perf_counter()
|
||||||
|
first = last = None
|
||||||
|
first_ct = None
|
||||||
|
final_meta = None
|
||||||
|
max_ct = 0
|
||||||
|
try:
|
||||||
|
with sess.post(url, json=payload, stream=True, timeout=3600) as resp:
|
||||||
|
for raw in resp.iter_lines():
|
||||||
|
if not raw or not raw.startswith(b"data:"):
|
||||||
|
continue
|
||||||
|
body = raw[5:].strip()
|
||||||
|
if body == b"[DONE]":
|
||||||
|
continue
|
||||||
|
now = time.perf_counter()
|
||||||
|
try:
|
||||||
|
d = json.loads(body)
|
||||||
|
except Exception:
|
||||||
|
continue
|
||||||
|
mi = d.get("meta_info") or {}
|
||||||
|
ct = mi.get("completion_tokens") or 0
|
||||||
|
if ct:
|
||||||
|
max_ct = max(max_ct, ct)
|
||||||
|
if first is None:
|
||||||
|
first = now
|
||||||
|
first_ct = ct
|
||||||
|
last = now
|
||||||
|
if mi.get("finish_reason"):
|
||||||
|
final_meta = mi
|
||||||
|
t_end = time.perf_counter()
|
||||||
|
n_out = max_ct or ((final_meta or {}).get("completion_tokens") or 0)
|
||||||
|
decode_span = (last - first) if (first and last and last > first) else 0.0
|
||||||
|
rec.update(
|
||||||
|
ok=n_out > 0,
|
||||||
|
ttft=(first - t0) if first else None,
|
||||||
|
e2e=t_end - t0,
|
||||||
|
n_out=n_out,
|
||||||
|
first_chunk_tokens=first_ct,
|
||||||
|
decode_span=decode_span,
|
||||||
|
tpot=(decode_span / (n_out - 1)) if n_out > 1 else None,
|
||||||
|
per_req_decode_tok_s=(n_out / decode_span) if decode_span > 0 else None,
|
||||||
|
retractions=(final_meta or {}).get("num_retractions"),
|
||||||
|
spec_accept_len=(final_meta or {}).get("spec_accept_length"),
|
||||||
|
)
|
||||||
|
except Exception as e:
|
||||||
|
rec.update(ok=False, error=repr(e))
|
||||||
|
results[idx] = rec
|
||||||
|
|
||||||
|
|
||||||
|
def verify_hit_rate(container, t_start, t_end):
|
||||||
|
def rfc3339(epoch):
|
||||||
|
return (datetime.datetime.fromtimestamp(epoch, tz=datetime.timezone.utc)
|
||||||
|
.isoformat().replace("+00:00", "Z"))
|
||||||
|
|
||||||
|
try:
|
||||||
|
# No margin before t_start: warmup's prefill lines end strictly before it,
|
||||||
|
# and catching them would deflate the measured hit rate.
|
||||||
|
p = subprocess.run(
|
||||||
|
["docker", "logs", container, "--since", rfc3339(t_start), "--until", rfc3339(t_end + 2)],
|
||||||
|
capture_output=True, text=True, timeout=120)
|
||||||
|
text = p.stdout + p.stderr
|
||||||
|
except Exception as e:
|
||||||
|
return {"error": repr(e)}
|
||||||
|
pat = re.compile(r"#new-token: (\d+).*?#cached-token: (\d+)")
|
||||||
|
n_batches = new_tok = cached_tok = 0
|
||||||
|
for line in text.splitlines():
|
||||||
|
if "TP0]" not in line or "Prefill batch" not in line:
|
||||||
|
continue
|
||||||
|
m = pat.search(line)
|
||||||
|
if m:
|
||||||
|
n_batches += 1
|
||||||
|
new_tok += int(m.group(1))
|
||||||
|
cached_tok += int(m.group(2))
|
||||||
|
total = new_tok + cached_tok
|
||||||
|
return {
|
||||||
|
"prefill_batches": n_batches,
|
||||||
|
"new_tokens": new_tok,
|
||||||
|
"cached_tokens": cached_tok,
|
||||||
|
"hit_rate": round(cached_tok / total, 4) if total else None,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def stats(vals):
|
||||||
|
vals = [v for v in vals if v is not None]
|
||||||
|
if not vals:
|
||||||
|
return {"mean": None, "p50": None, "p95": None, "max": None, "min": None}
|
||||||
|
s = sorted(vals)
|
||||||
|
# nearest-rank p95: smallest value >= 95th percentile
|
||||||
|
p95_idx = max(0, math.ceil(0.95 * len(s)) - 1)
|
||||||
|
return {
|
||||||
|
"mean": round(statistics.fmean(vals), 4),
|
||||||
|
"p50": round(s[len(s) // 2], 4),
|
||||||
|
"p95": round(s[p95_idx], 4),
|
||||||
|
"max": round(s[-1], 4),
|
||||||
|
"min": round(s[0], 4),
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
ap = argparse.ArgumentParser()
|
||||||
|
ap.add_argument("--concurrency", type=int, required=True)
|
||||||
|
ap.add_argument("--num-requests", type=int, required=True)
|
||||||
|
ap.add_argument("--run-id", type=int, required=True)
|
||||||
|
ap.add_argument("--input-len", type=int, required=True, help="token length of each prompt")
|
||||||
|
ap.add_argument("--output-len", type=int, default=OUTPUT_LEN_DEFAULT)
|
||||||
|
ap.add_argument("--shared-frac", type=float, default=0.9)
|
||||||
|
ap.add_argument("--corpus", default=CORPUS_DEFAULT)
|
||||||
|
ap.add_argument("--pool-override", type=int, default=None,
|
||||||
|
help="explicit corpus offset for the unique-suffix window (re-runs)")
|
||||||
|
ap.add_argument("--url", default="http://127.0.0.1:30000/generate")
|
||||||
|
ap.add_argument("--container", default="glm53-nvfp4")
|
||||||
|
ap.add_argument("--dump-records", default=None,
|
||||||
|
help="write per-request records as JSONL to this path")
|
||||||
|
args = ap.parse_args()
|
||||||
|
|
||||||
|
with open(args.corpus) as f:
|
||||||
|
corpus = json.load(f)
|
||||||
|
ids = corpus["ids"]
|
||||||
|
|
||||||
|
shared_len, unique_len = split_lens(args.input_len, args.shared_frac)
|
||||||
|
pool_start = pool_start_for(args.input_len, args.shared_frac, args.run_id, args.pool_override)
|
||||||
|
shared, prompts, window = build_prompts(ids, shared_len, unique_len, args.num_requests, pool_start)
|
||||||
|
print(f"[pool] window={window} shared_len={shared_len} unique_len={unique_len} "
|
||||||
|
f"corpus_total={len(ids)}", flush=True)
|
||||||
|
warmup(args.url, ids, shared_len)
|
||||||
|
|
||||||
|
results = {}
|
||||||
|
t_start = time.time()
|
||||||
|
t0 = time.perf_counter()
|
||||||
|
with ThreadPoolExecutor(max_workers=args.concurrency) as ex:
|
||||||
|
futs = [ex.submit(bench_one, args.url, p, args.output_len, i, results)
|
||||||
|
for i, p in enumerate(prompts)]
|
||||||
|
for f in futs:
|
||||||
|
f.result()
|
||||||
|
wall = time.perf_counter() - t0
|
||||||
|
t_end = time.time()
|
||||||
|
|
||||||
|
if args.dump_records:
|
||||||
|
with open(args.dump_records, "w") as f:
|
||||||
|
for i in sorted(results):
|
||||||
|
f.write(json.dumps(results[i]) + "\n")
|
||||||
|
|
||||||
|
hit = verify_hit_rate(args.container, t_start, t_end)
|
||||||
|
|
||||||
|
ok = [r for r in results.values() if r.get("ok")]
|
||||||
|
n_out_total = sum(r["n_out"] for r in ok)
|
||||||
|
out_tps = [r["n_out"] / r["e2e"] for r in ok if r.get("e2e")]
|
||||||
|
ttft = stats([r.get("ttft") for r in ok])
|
||||||
|
tpot = stats([r.get("tpot") for r in ok])
|
||||||
|
e2e = stats([r.get("e2e") for r in ok])
|
||||||
|
dec = stats([r.get("per_req_decode_tok_s") for r in ok])
|
||||||
|
spec = [v for v in (r.get("spec_accept_len") for r in ok) if v is not None]
|
||||||
|
retr = sum(r.get("retractions") or 0 for r in ok)
|
||||||
|
|
||||||
|
summary = {
|
||||||
|
"concurrency": args.concurrency,
|
||||||
|
"num_requests": args.num_requests,
|
||||||
|
"run_id": args.run_id,
|
||||||
|
"corpus_window": {"start": window[0], "end": window[1]},
|
||||||
|
"ok": len(ok),
|
||||||
|
"failed": args.num_requests - len(ok),
|
||||||
|
"wall_s": round(wall, 2),
|
||||||
|
"input_len": args.input_len,
|
||||||
|
"shared_len": shared_len,
|
||||||
|
"unique_len": unique_len,
|
||||||
|
"output_len": args.output_len,
|
||||||
|
"output_tokens_total": n_out_total,
|
||||||
|
"output_throughput_tok_s": round(n_out_total / wall, 2) if wall else None,
|
||||||
|
"input_throughput_tok_s": round(args.input_len * len(ok) / wall, 2) if wall else None,
|
||||||
|
"ttft_s": ttft,
|
||||||
|
"tpot_s": tpot,
|
||||||
|
"e2e_s": e2e,
|
||||||
|
"per_req_out_tok_s_e2e": stats(out_tps),
|
||||||
|
"per_req_decode_tok_s": dec,
|
||||||
|
"spec_accept_length_mean": round(statistics.fmean(spec), 3) if spec else None,
|
||||||
|
"retractions_total": retr,
|
||||||
|
"cache_hit_from_logs": hit,
|
||||||
|
}
|
||||||
|
print("\n===== SUMMARY =====")
|
||||||
|
print(json.dumps(summary, indent=2), flush=True)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@ -0,0 +1,39 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""extract_summary.py <bench_log> <tag> <arm> <results_jsonl>
|
||||||
|
|
||||||
|
Parse the trailing "===== SUMMARY =====" JSON block from a bench_corpus_v2 log,
|
||||||
|
append a tagged record to the campaign results JSONL, and print the headline
|
||||||
|
metrics. Exit codes: 0 = ok and hit_rate<=0.01 (cold-point validity),
|
||||||
|
3 = no SUMMARY block found, 4 = hit_rate above the cold-point threshold.
|
||||||
|
"""
|
||||||
|
import json
|
||||||
|
import sys
|
||||||
|
|
||||||
|
HIT_MAX = 0.01
|
||||||
|
MARK = "===== SUMMARY ====="
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
logp, tag, arm, outp = sys.argv[1:5]
|
||||||
|
with open(logp, encoding="utf-8", errors="replace") as f:
|
||||||
|
log = f.read()
|
||||||
|
i = log.rfind(MARK)
|
||||||
|
if i < 0:
|
||||||
|
print("NO_SUMMARY")
|
||||||
|
sys.exit(3)
|
||||||
|
s = json.loads(log[i + len(MARK):].strip())
|
||||||
|
with open(outp, "a", encoding="utf-8") as f:
|
||||||
|
f.write(json.dumps({"tag": tag, "arm": arm, "summary": s}) + "\n")
|
||||||
|
hit = (s.get("cache_hit_from_logs") or {}).get("hit_rate")
|
||||||
|
print(f"hit_rate={hit} out_tps={s.get('output_throughput_tok_s')} "
|
||||||
|
f"in_tps={s.get('input_throughput_tok_s')} "
|
||||||
|
f"ttft_p95={(s.get('ttft_s') or {}).get('p95')} "
|
||||||
|
f"tpot_p95={(s.get('tpot_s') or {}).get('p95')} "
|
||||||
|
f"ok={s.get('ok')}/{s.get('num_requests')} retractions={s.get('retractions_total')}")
|
||||||
|
if hit is None:
|
||||||
|
sys.exit(4)
|
||||||
|
sys.exit(0 if float(hit) <= HIT_MAX else 4)
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@ -0,0 +1,165 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""gen_report_tables.py <tp2pp4_all_results.jsonl> <e7b_all_results.jsonl>
|
||||||
|
|
||||||
|
Build the B300-style markdown tables for the 6000D dual-plan report from the
|
||||||
|
campaign all_results.jsonl files. Prints tables to stdout. Best value per
|
||||||
|
column within each scenario table is bolded (**...**), mirroring the B300
|
||||||
|
report convention.
|
||||||
|
"""
|
||||||
|
import json
|
||||||
|
import sys
|
||||||
|
|
||||||
|
ARMS = ["tp2pp4", "e7b"]
|
||||||
|
ARM_LABEL = {"tp2pp4": "TP2PP4", "e7b": "TP8+EAGLE3+AR"}
|
||||||
|
|
||||||
|
# scenario key -> (chapter title, row label, ordered ccs, has_input_col)
|
||||||
|
SCENARIOS = [
|
||||||
|
("b3_16k", "主场景 16K→512", [1, 8, 16, 32, 64], True),
|
||||||
|
("b41_1k", "4.1 短输入 1K→128", [1, 8, 32, 64], True),
|
||||||
|
("b42_1k4k", "4.2 长输出 1K→4K", [1, 8, 32, 64], False),
|
||||||
|
("b51_64k", "5.1 64K→512", [1, 4, 8], True),
|
||||||
|
("b51_128k", "5.1 128K→512", [1, 2, 4], True),
|
||||||
|
("b52_256k", "5.2 边界 256K→1", [1, 2, 3], True),
|
||||||
|
("b52_512k", "5.2 边界 512K→1", [1], True),
|
||||||
|
("b52_896k", "5.2 边界 896K→1", [1], True),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def load(path):
|
||||||
|
rows = {}
|
||||||
|
if not path:
|
||||||
|
return rows
|
||||||
|
with open(path, encoding="utf-8") as f:
|
||||||
|
for line in f:
|
||||||
|
rec = json.loads(line)
|
||||||
|
rows[(rec["tag"], rec["arm"])] = rec["summary"]
|
||||||
|
return rows
|
||||||
|
|
||||||
|
|
||||||
|
def fmt(v, kind):
|
||||||
|
if v is None:
|
||||||
|
return "—"
|
||||||
|
if kind == "tps":
|
||||||
|
return f"{v:,.0f}" if v >= 100 else f"{v:,.1f}"
|
||||||
|
if kind == "s":
|
||||||
|
return f"{v:,.2f} s" if v >= 1 else f"{v*1000:,.0f} ms"
|
||||||
|
if kind == "ms":
|
||||||
|
return f"{v*1000:,.1f} ms"
|
||||||
|
return str(v)
|
||||||
|
|
||||||
|
|
||||||
|
def get(s, path):
|
||||||
|
cur = s
|
||||||
|
for key in path.split("."):
|
||||||
|
if cur is None:
|
||||||
|
return None
|
||||||
|
cur = cur.get(key)
|
||||||
|
return cur
|
||||||
|
|
||||||
|
|
||||||
|
def table_for(prefix, title, ccs, has_input, data):
|
||||||
|
lines = []
|
||||||
|
cols = ["并发", "方案"]
|
||||||
|
if has_input:
|
||||||
|
cols += ["Input TPS"]
|
||||||
|
cols += ["Output TPS", "TTFT P95", "TPOT P95"]
|
||||||
|
lines.append("| " + " | ".join(cols) + " |")
|
||||||
|
lines.append("|" + "---|" * len(cols))
|
||||||
|
|
||||||
|
# collect cells to bold best per numeric column
|
||||||
|
body = []
|
||||||
|
for cc in ccs:
|
||||||
|
for arm in ARMS:
|
||||||
|
tag = f"{prefix}_c{cc}"
|
||||||
|
s = data.get((tag, arm))
|
||||||
|
if s is None and cc in (2, 3):
|
||||||
|
# boundary cc>1 only exists on tp2pp4; absence = structural skip
|
||||||
|
continue
|
||||||
|
if s is None:
|
||||||
|
body.append((cc, arm, None))
|
||||||
|
continue
|
||||||
|
body.append((cc, arm, s))
|
||||||
|
|
||||||
|
def val(cc, arm, path):
|
||||||
|
s = data.get((f"{prefix}_c{cc}", arm))
|
||||||
|
return get(s, path) if s else None
|
||||||
|
|
||||||
|
# best per column (higher tps, lower latency)
|
||||||
|
best = {}
|
||||||
|
if has_input:
|
||||||
|
ivs = [val(cc, arm, "input_throughput_tok_s") for cc, arm, _ in body if _ is not None]
|
||||||
|
ivs = [v for v in ivs if v is not None]
|
||||||
|
if ivs:
|
||||||
|
best["in"] = max(ivs)
|
||||||
|
ovs = [val(cc, arm, "output_throughput_tok_s") for cc, arm, _ in body if _ is not None]
|
||||||
|
ovs = [v for v in ovs if v is not None]
|
||||||
|
if ovs:
|
||||||
|
best["out"] = max(ovs)
|
||||||
|
tvs = [val(cc, arm, "ttft_s.p95") for cc, arm, _ in body if _ is not None]
|
||||||
|
tvs = [v for v in tvs if v is not None]
|
||||||
|
if tvs:
|
||||||
|
best["ttft"] = min(tvs)
|
||||||
|
pvs = [val(cc, arm, "tpot_s.p95") for cc, arm, _ in body if _ is not None]
|
||||||
|
pvs = [v for v in pvs if v is not None]
|
||||||
|
if pvs:
|
||||||
|
best["tpot"] = min(pvs)
|
||||||
|
|
||||||
|
def maybe_bold(v, kind, key):
|
||||||
|
if v is None:
|
||||||
|
return "—"
|
||||||
|
cell = fmt(v, kind)
|
||||||
|
if key in best and v == best[key]:
|
||||||
|
return f"**{cell}**"
|
||||||
|
return cell
|
||||||
|
|
||||||
|
for cc, arm, s in body:
|
||||||
|
if s is None:
|
||||||
|
lines.append(f"| {cc} | {ARM_LABEL[arm]} | " + " 结构性不可测 |" * (len(cols) - 2))
|
||||||
|
continue
|
||||||
|
cells = [str(cc), ARM_LABEL[arm]]
|
||||||
|
if has_input:
|
||||||
|
cells.append(maybe_bold(get(s, "input_throughput_tok_s"), "tps", "in"))
|
||||||
|
cells.append(maybe_bold(get(s, "output_throughput_tok_s"), "tps", "out"))
|
||||||
|
cells.append(maybe_bold(get(s, "ttft_s.p95"), "s", "ttft"))
|
||||||
|
cells.append(maybe_bold(get(s, "tpot_s.p95"), "ms", "tpot"))
|
||||||
|
lines.append("| " + " | ".join(cells) + " |")
|
||||||
|
return "\n".join(lines)
|
||||||
|
|
||||||
|
|
||||||
|
def main():
|
||||||
|
tp_data = load(sys.argv[1] if len(sys.argv) > 1 else None)
|
||||||
|
e7_data = load(sys.argv[2] if len(sys.argv) > 2 else None)
|
||||||
|
data = {**tp_data, **e7_data}
|
||||||
|
for prefix, title, ccs, has_input in SCENARIOS:
|
||||||
|
print(f"\n### {title}\n")
|
||||||
|
print(table_for(prefix, title, ccs, has_input, data))
|
||||||
|
|
||||||
|
# appendix: full metrics
|
||||||
|
print("\n\n## 附录:全量指标(含 mean/p50/max、回退、投机接受长度)\n")
|
||||||
|
print("| 场景点 | 方案 | ok/nreq | wall(s) | out tok/s | in tok/s | TTFT mean/p95/max(s) | TPOT mean/p95/max(ms) | retractions | accept |")
|
||||||
|
print("|---|---|---|---|---|---|---|---|---|---|")
|
||||||
|
for prefix, _, ccs, _ in SCENARIOS:
|
||||||
|
for cc in ccs:
|
||||||
|
for arm in ARMS:
|
||||||
|
s = data.get((f"{prefix}_c{cc}", arm))
|
||||||
|
if s is None:
|
||||||
|
continue
|
||||||
|
ttft = s.get("ttft_s") or {}
|
||||||
|
tpot = s.get("tpot_s") or {}
|
||||||
|
def trio(d):
|
||||||
|
m, p, x = d.get("mean"), d.get("p95"), d.get("max")
|
||||||
|
if m is None:
|
||||||
|
return "—"
|
||||||
|
return f"{m:.2f}/{p:.2f}/{x:.2f}"
|
||||||
|
def trioms(d):
|
||||||
|
m, p, x = d.get("mean"), d.get("p95"), d.get("max")
|
||||||
|
if m is None:
|
||||||
|
return "—"
|
||||||
|
return f"{m*1000:.1f}/{p*1000:.1f}/{x*1000:.1f}"
|
||||||
|
print(f"| {prefix}_c{cc} | {ARM_LABEL[arm]} | {s.get('ok')}/{s.get('num_requests')} "
|
||||||
|
f"| {s.get('wall_s')} | {s.get('output_throughput_tok_s')} | {s.get('input_throughput_tok_s')} "
|
||||||
|
f"| {trio(ttft)} | {trioms(tpot)} | {s.get('retractions_total')} | {s.get('spec_accept_length_mean')} |")
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
main()
|
||||||
@ -0,0 +1,168 @@
|
|||||||
|
#!/bin/bash
|
||||||
|
# run_b300_matrix.sh <arm: tp2pp4|e7b> — B300-equivalent cold-cache scenario matrix on 60.8
|
||||||
|
#
|
||||||
|
# Mirrors the B300 report scenario set (16K->512, 1K->128, 1K->4K, 64K->512,
|
||||||
|
# 128K->512, 256K/512K/896K boundary OSL=1) at reduced concurrency per team
|
||||||
|
# decision (16K/1K capped at cc64; 64K/128K at cc<=8; boundary cc<=3).
|
||||||
|
# Both arms run the same grid; E7b structurally skips 256K cc>1 / 512K / 896K
|
||||||
|
# (KV pool 276,864, ctx 270,336).
|
||||||
|
#
|
||||||
|
# Protocol: cold points (shared-frac 0), recycled corpus windows via
|
||||||
|
# --pool-override (corpus has no virgin text left), per-point idle-wait ->
|
||||||
|
# scenario-length prewarm (first point of each scenario) -> flush_cache ->
|
||||||
|
# bench -> hit-rate-from-logs verification (<=0.01 else one retry, then abort).
|
||||||
|
# run-ids 95xx are labels only (windows are explicit).
|
||||||
|
#
|
||||||
|
# Usage: nohup bash /root/run_b300_matrix.sh <arm> \
|
||||||
|
# > /root/bench_logs/b300eq_<arm>_progress.log 2>&1 &
|
||||||
|
set -u
|
||||||
|
ARM=${1:?usage: run_b300_matrix.sh tp2pp4|e7b}
|
||||||
|
case $ARM in
|
||||||
|
tp2pp4) CONTAINER=glm53-pp4 ;;
|
||||||
|
e7b) CONTAINER=glm53-nvfp4 ;;
|
||||||
|
*) echo "bad arm: $ARM"; exit 1 ;;
|
||||||
|
esac
|
||||||
|
|
||||||
|
CORPUS=/root/corpus_ids.json
|
||||||
|
BENCH=/root/bench_corpus_v2.py
|
||||||
|
EXTRACT=/root/extract_summary.py
|
||||||
|
URL=http://127.0.0.1:30000
|
||||||
|
STAMP=$(date +%Y%m%d_%H%M)
|
||||||
|
LOG=/root/bench_logs/b300eq_${ARM}_${STAMP}
|
||||||
|
mkdir -p "$LOG"
|
||||||
|
RESULTS="$LOG/all_results.jsonl"
|
||||||
|
STATUS="$LOG/status.txt"
|
||||||
|
: > "$STATUS"
|
||||||
|
|
||||||
|
echo "=== [$ARM] matrix start $(date) LOGDIR=$LOG ==="
|
||||||
|
|
||||||
|
# ---- server facts snapshot ----
|
||||||
|
alive() { [ -n "$(docker ps --filter name=$CONTAINER --filter status=running -q)" ]; }
|
||||||
|
alive || { echo "=== [$ARM] ABORT: container $CONTAINER not running ==="; exit 1; }
|
||||||
|
docker inspect "$CONTAINER" --format '{{.Config.Cmd}}' > "$LOG/server_cmd.txt" 2>&1
|
||||||
|
docker logs "$CONTAINER" 2>&1 | grep -E "max_total_num_tokens|KV Cache is allocated|context_len|chunked_prefill_size|max_running_request|speculative_num_steps" | head -20 > "$LOG/server_facts.txt" 2>&1
|
||||||
|
nvidia-smi --query-gpu=index,name,memory.total,memory.used --format=csv > "$LOG/gpu_inventory_idle.csv" 2>&1
|
||||||
|
md5sum "$CORPUS" "$BENCH" "$EXTRACT" > "$LOG/md5_assets.txt" 2>&1
|
||||||
|
echo "--- server_cmd: $(cat "$LOG/server_cmd.txt")"
|
||||||
|
echo "--- server_facts:"; cat "$LOG/server_facts.txt"
|
||||||
|
|
||||||
|
# ---- per-rank VRAM sampler (whole-arm timeline, 30s cadence) ----
|
||||||
|
( while true; do
|
||||||
|
echo "# $(date +%s) $(date +%T)"
|
||||||
|
nvidia-smi --query-gpu=index,memory.used,utilization.gpu --format=csv,noheader
|
||||||
|
sleep 30
|
||||||
|
done ) > "$LOG/vram_timeline.csv" 2>&1 &
|
||||||
|
SAMPLER=$!
|
||||||
|
trap 'kill $SAMPLER 2>/dev/null' EXIT
|
||||||
|
|
||||||
|
idle_wait() {
|
||||||
|
for i in $(seq 1 90); do
|
||||||
|
local last
|
||||||
|
last=$(docker logs --since 90s "$CONTAINER" 2>&1 | grep 'running-req' | tail -1)
|
||||||
|
if [ -z "$last" ] || echo "$last" | grep -q 'running-req: 0'; then return 0; fi
|
||||||
|
sleep 10
|
||||||
|
done
|
||||||
|
echo "[warn] idle_wait timeout, continuing"
|
||||||
|
}
|
||||||
|
|
||||||
|
flush() { curl -s -m 60 -X POST $URL/flush_cache >/dev/null; sleep 3; }
|
||||||
|
|
||||||
|
prewarm() { # $1=input_len $2=base — one uncounted request at scenario length
|
||||||
|
timeout 1800 python3 "$BENCH" --corpus "$CORPUS" --input-len "$1" --output-len 16 \
|
||||||
|
--shared-frac 0 --concurrency 1 --num-requests 1 --run-id 9599 \
|
||||||
|
--pool-override "$2" --url $URL/generate --container "$CONTAINER" \
|
||||||
|
> "$LOG/prewarm_$1.log" 2>&1 || true
|
||||||
|
}
|
||||||
|
|
||||||
|
bench_call() { # $1=log $2=tag $3=isl $4=osl $5=cc $6=nreq $7=base $8=rid
|
||||||
|
timeout 7200 python3 "$BENCH" --corpus "$CORPUS" --input-len "$3" --output-len "$4" \
|
||||||
|
--shared-frac 0 --concurrency "$5" --num-requests "$6" --run-id "$8" \
|
||||||
|
--pool-override "$7" --url $URL/generate --container "$CONTAINER" \
|
||||||
|
--dump-records "$LOG/${2}_records.jsonl" > "$1" 2>&1
|
||||||
|
}
|
||||||
|
|
||||||
|
run_point() { # tag isl osl cc nreq base rid [prewarm=1]
|
||||||
|
local TAG=$1 ISL=$2 OSL=$3 CC=$4 NREQ=$5 BASE=$6 RID=$7 PW=${8:-0}
|
||||||
|
alive || { echo "$TAG CONTAINER_DEAD" >> "$STATUS"; echo "=== $TAG ABORT: container dead ==="; exit 1; }
|
||||||
|
echo "=== [$ARM] $TAG isl=$ISL osl=$OSL cc=$CC nreq=$NREQ base=$BASE start $(date +%T) ==="
|
||||||
|
idle_wait
|
||||||
|
[ "$PW" = "1" ] && prewarm "$ISL" "$BASE"
|
||||||
|
flush
|
||||||
|
local rc=0
|
||||||
|
bench_call "$LOG/${TAG}.log" "$TAG" "$ISL" "$OSL" "$CC" "$NREQ" "$BASE" "$RID" || rc=$?
|
||||||
|
if [ "$rc" = "0" ]; then
|
||||||
|
python3 "$EXTRACT" "$LOG/${TAG}.log" "$TAG" "$ARM" "$RESULTS"; rc=$?
|
||||||
|
fi
|
||||||
|
if [ "$rc" = "4" ]; then
|
||||||
|
echo "=== $TAG hit_rate>0.01, retrying once ==="
|
||||||
|
idle_wait; flush
|
||||||
|
bench_call "$LOG/${TAG}_retry.log" "${TAG}_retry" "$ISL" "$OSL" "$CC" "$NREQ" "$BASE" "$RID" || rc=$?
|
||||||
|
if [ "$rc" = "0" ]; then
|
||||||
|
python3 "$EXTRACT" "$LOG/${TAG}_retry.log" "${TAG}_retry" "$ARM" "$RESULTS"; rc=$?
|
||||||
|
fi
|
||||||
|
if [ "$rc" = "4" ]; then
|
||||||
|
echo "$TAG HIT_FAIL_FINAL" >> "$STATUS"
|
||||||
|
echo "=== $TAG hit_rate still >0.01 after retry — ABORTING ARM (contamination) ==="
|
||||||
|
exit 1
|
||||||
|
fi
|
||||||
|
fi
|
||||||
|
if [ "$rc" = "0" ]; then
|
||||||
|
echo "$TAG OK" >> "$STATUS"
|
||||||
|
else
|
||||||
|
echo "$TAG BENCH_FAIL rc=$rc" >> "$STATUS"
|
||||||
|
echo "=== $TAG bench rc=$rc (recorded, continuing) ==="
|
||||||
|
fi
|
||||||
|
echo "=== $TAG done $(date +%T) ==="
|
||||||
|
}
|
||||||
|
|
||||||
|
# ---- recycled corpus window bases (see campaign window map; all within 21.3M) ----
|
||||||
|
B_16K=2300000; B_1K=4500000; B_1K4=4700000; B_64K=5000000
|
||||||
|
B_128K=6200000; B_256K=8400000; B_512K=9300000; B_896K=9900000
|
||||||
|
|
||||||
|
# ===== B300 §3: main scenario 16K -> 512 =====
|
||||||
|
run_point b3_16k_c1 16384 512 1 8 $B_16K 9501 1
|
||||||
|
run_point b3_16k_c8 16384 512 8 16 $B_16K 9502
|
||||||
|
run_point b3_16k_c16 16384 512 16 32 $B_16K 9503
|
||||||
|
run_point b3_16k_c32 16384 512 32 64 $B_16K 9504
|
||||||
|
run_point b3_16k_c64 16384 512 64 128 $B_16K 9505
|
||||||
|
|
||||||
|
# ===== B300 §4.1: short input 1K -> 128 =====
|
||||||
|
run_point b41_1k_c1 1024 128 1 8 $B_1K 9511 1
|
||||||
|
run_point b41_1k_c8 1024 128 8 16 $B_1K 9512
|
||||||
|
run_point b41_1k_c32 1024 128 32 64 $B_1K 9513
|
||||||
|
run_point b41_1k_c64 1024 128 64 128 $B_1K 9514
|
||||||
|
|
||||||
|
# ===== B300 §4.2: long output 1K -> 4K =====
|
||||||
|
run_point b42_1k4k_c1 1024 4096 1 8 $B_1K4 9521 1
|
||||||
|
run_point b42_1k4k_c8 1024 4096 8 16 $B_1K4 9522
|
||||||
|
run_point b42_1k4k_c32 1024 4096 32 64 $B_1K4 9523
|
||||||
|
run_point b42_1k4k_c64 1024 4096 64 128 $B_1K4 9524
|
||||||
|
|
||||||
|
# ===== B300 §5.1: long context 64K -> 512 =====
|
||||||
|
run_point b51_64k_c1 65536 512 1 8 $B_64K 9531 1
|
||||||
|
run_point b51_64k_c4 65536 512 4 8 $B_64K 9532
|
||||||
|
run_point b51_64k_c8 65536 512 8 16 $B_64K 9533
|
||||||
|
|
||||||
|
# ===== B300 §5.1: long context 128K -> 512 =====
|
||||||
|
run_point b51_128k_c1 131072 512 1 8 $B_128K 9541 1
|
||||||
|
run_point b51_128k_c2 131072 512 2 8 $B_128K 9542
|
||||||
|
run_point b51_128k_c4 131072 512 4 8 $B_128K 9543
|
||||||
|
|
||||||
|
# ===== B300 §5.2: context boundary, OSL=1 (nreq=cc) =====
|
||||||
|
run_point b52_256k_c1 262144 1 1 1 $B_256K 9551 1
|
||||||
|
if [ "$ARM" = "tp2pp4" ]; then
|
||||||
|
run_point b52_256k_c2 262144 1 2 2 $B_256K 9552
|
||||||
|
run_point b52_256k_c3 262144 1 3 3 $B_256K 9553
|
||||||
|
run_point b52_512k_c1 524288 1 1 1 $B_512K 9554
|
||||||
|
run_point b52_896k_c1 917504 1 1 1 $B_896K 9555
|
||||||
|
else
|
||||||
|
# E7b structural limits: pool 276,864 (256K cc>=2 needs 524K) and ctx 270,336 (<512K)
|
||||||
|
for t in b52_256k_c2 b52_256k_c3 b52_512k_c1 b52_896k_c1; do
|
||||||
|
echo "$t SKIP_STRUCTURAL" >> "$STATUS"
|
||||||
|
done
|
||||||
|
echo "=== [e7b] 256K cc2/3, 512K, 896K skipped: structural (pool 276,864 / ctx 270,336) ==="
|
||||||
|
fi
|
||||||
|
|
||||||
|
nvidia-smi --query-gpu=index,name,memory.total,memory.used --format=csv > "$LOG/gpu_inventory_final.csv" 2>&1
|
||||||
|
echo "=== [$ARM] MATRIX ALL DONE $(date) — $LOG ==="
|
||||||
|
echo "--- status:"; cat "$STATUS"
|
||||||
Loading…
x
Reference in New Issue
Block a user