diff --git a/deploy/CURRENT.md b/deploy/CURRENT.md index 2b17433..c64e732 100644 --- a/deploy/CURRENT.md +++ b/deploy/CURRENT.md @@ -8,13 +8,13 @@ | 机器 | 在役 | 口径 / 归属 | 对应 profile | |---|---|---|---| | 60.1 (6000D-1) | `glm53-pp4`(Up,8 卡满载,:30000) | **方案 D 生产**。09-08 PD 压测窗口停机 ~2.5h 后已恢复并核验 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp2pp4.env` | -| 60.2 (6000D-2) | 空(无容器,GPU 全 0 MiB) | 方案 F PD 链已于 09-08 拆除。GPU7 曾有外部裸金属任务 main_v2.py(现已结束),动卡前仍先核实归属 | `profiles/pro6000/glm53_nvfp4_pro6000_pd_{prefill→decode 侧}.env` | -| 60.3 | 空(无容器) | — | — | +| 60.2 (6000D-2) | `glm53-nvfp4` 实验容器(09-08 晚 TP1PP8 phase,run_phase.sh 实验链进行中) | **他人实验进行中,勿动**(方案 F PD 链已拆除;GPU7 曾有外部裸金属任务)。动卡前仍先核实归属 | `profiles/pro6000/glm53_nvfp4_pro6000_pd_{prefill→decode 侧}.env` | +| 60.3 | 无容器,但 8 卡被外部裸金属训练占用(`/data/mas/larm`,09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — | | 60.4 | `dsv4_scan` 容器(Up 5min)+ 裸金属 sglang `DeepSeek-V4-Flash-0731` DSPARK DP2/TP4 :30000 | **并行会话/外部在役,勿动** | — | | 60.5 | `glm53-nvfp4`(Up 29h) | **NVFP4 团队生产**(deploy_glm53_605.sh,md5 fcd9109b)。生产机铁律:不实验、不重启、不覆盖脚本 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(方案 A 口径) | -| 60.6 | 空(无容器) | — | — | -| 60.7 | 空 | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — | -| 60.8 | 空 | 09-08 已停服清空、router 拆除,不再恢复 | — | +| 60.6 | 无容器,但 8 卡被外部裸金属实验占用(`/data/hzy/sparse-opd-*`,09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — | +| 60.7 | 基本空(4 卡仍有 `/home/user/dirA_exp` 外部小任务,09-08 晚实测) | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — | +| 60.8 | `glm53-nvfp4`(Up,:30000,restart=unless-stopped) | **A16 口径在役**(09-08 i8k/o1k/c16 压测收官保留:tp8eagle.env 场景二变体 = A + MRR32 + graph bs 4/8/12/16,池 276,864,质量门 7/7)。同场景实测最优为 B(TP4PP2,out 265 vs A16 219 tok/s),如转纯吞吐用途可切 B。压测前经授权清理了 GPU3/6 的 dirA_ext 外部 eval 进程 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(+注释中场景二高并发变体);压测记录 `experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/` | ## 方案 A-F 一览(GLM-5.3-NVFP4 @ pro6000,2026-09-08 双场景报告口径) diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/GLM53_NVFP4_i8k_o1k_c16_压测报告_2026-09-08.md b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/GLM53_NVFP4_i8k_o1k_c16_压测报告_2026-09-08.md new file mode 100644 index 0000000..30bb357 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/GLM53_NVFP4_i8k_o1k_c16_压测报告_2026-09-08.md @@ -0,0 +1,83 @@ +# GLM-5.3-NVFP4 i8k/o1k/c16 三方案部署对比压测报告 + +- 日期:2026-09-08 18:00–19:02(单机串行,174.1.60.8 / 6000D-8,8×H20) +- 场景:**input=8192 / output=1024 / concurrency=16**(全唯一输入,PG19 真实语料) +- 方案:A(生产口径 TP8+EAGLE)、B(TP4PP2 无投机)、A16(A 的 c16 调优:graph 覆盖 bs16) +- 协议:rev20 团队标准(bench_corpus.py,input_ids 直打 `/generate`,temp=0、ignore_eos、stream,nreq=32=2 波,warm nreq16 弃用,每轮全新 run-id + 全新语料窗口 `--pool-override`) +- 镜像:`lmsysorg/sglang:nightly-dev-20260828-daf63171`(digest 28e0d2607316),模型 `/data/hf_models/GLM-5.3-NVFP4` + +## 一、判决(先行) + +| 指标 | A 生产口径 | A16(graph 修复) | B(TP4PP2) | 最优 | +|---|---|---|---|---| +| **输出吞吐 tok/s** | 115.0 | 218.9 | **264.9** | B(A 的 2.30×) | +| **输入吞吐 tok/s** | 919.9 | 1,750.8 | **2,119.4** | B(A 的 2.30×) | +| **TTFT p50 s** | **3.8** | 4.0 | 10.5 | A(mean 三者打平 ~10s) | +| **TPOT p50 ms** | 123.1 | 55.0 | **50.3** | B | +| **e2e P50 s** | 139.6 | 66.2 | **61.8** | B(A 的 44%) | +| 单流 decode p50 tok/s | 8.4 | 18.9 | **20.1** | B | +| EAGLE accept | 2.52 | 2.50 | —(无投机) | — | +| 回退 / 命中 | 0 / 0.0 | 0 / 0.0 | 0 / 0.0 | 全过 | +| 质量门 | 7/7 | 7/7 | 6/7* | — | + +\* B 的 tool-call 项失败 = 脚本无 `--tool-call-parser` 的已知配置缺口(模型文本明确想调工具,非质量问题;上生产须补 parser,见双场景报告既有结论)。 + +**结论:i8k/o1k/c16 场景 B(TP4PP2)全面最优**——吞吐与 e2e 均为 A 的 2.3 倍量级,且 32 请求 e2e 几乎零离散(61.3–62.4s)。A16 把 A 与 B 的差距从 2.30× 收窄到 1.21×,是 EAGLE 家族在本场景的正确口径。**60.8 按既定安排保留最后部署的 A16 在役**(:30000,质量门 7/7、parser 完整);若该机器要转纯吞吐用途,一条命令可切 B(补 parser flags)。 + +## 二、主表(逐轮原始数据) + +列:wall | out tok/s | in tok/s | TTFT mean/p50/max s | TPOT mean/p50 ms | e2e mean/p50/max s | 单流 decode p50 tok/s | accept | 命中 | 回退 + +**方案 A:TP8 + EAGLE 4/1/5 @ mem0.90,MRR16,chunk8192,fp8KV+hicache3,graph bs≤8,KV 池 276,864**(deploy_glm53_605.sh 原样,md5 fcd9109b,与 60.5 生产同源) + +| 轮 | wall | out | in | TTFT mean/p50/max | TPOT mean/p50 | e2e mean/p50/max | 单流dec | accept | 命中 | 回退 | +|---|---|---|---|---|---|---|---|---|---|---| +| r1 | 277.4 | 118.1 | 945.0 | 10.29 / 4.37 / 31.18 | 119.6 / 118.4 | 132.6 / 136.2 / 190.1 | 8.8 | 2.589 | 0.0 | 0 | +| r2 | 293.0 | 111.8 | 894.8 | 10.02 / 3.22 / 30.94 | 126.9 / 127.8 | 139.8 / 142.9 / 193.1 | 7.9 | 2.445 | 0.0 | 0 | + +**方案 A16:A 基础上 EXTRA=`--max-running-requests 32 --cuda-graph-max-bs-decode 16 --cuda-graph-bs-decode 4 8 12 16`**(= tp8eagle.env 注释中的场景二高并发变体;KV 池不变 276,864) + +| 轮 | wall | out | in | TTFT mean/p50/max | TPOT mean/p50 | e2e mean/p50/max | 单流dec | accept | 命中 | 回退 | +|---|---|---|---|---|---|---|---|---|---|---| +| r1 | 151.3 | 216.6 | 1,732.4 | 10.06 / 4.02 / 30.63 | 60.6 / 57.7 | 72.1 / 70.0 / 108.1 | 18.6 | 2.367 | 0.0 | 0 | +| r2 | 148.2 | 221.2 | 1,769.2 | 9.88 / 3.98 / 30.67 | 57.3 / 52.2 | 68.5 / 62.3 / 109.6 | 19.2 | 2.634 | 0.0 | 0 | + +**方案 B:TP4 PP2 @ mem0.85,MRR48,chunk16384,radix-off,disable-custom-AR,index_topk_freq=4,KV 池 569,600**(deploy_glm53_optimal.sh,与 09-07/09-08 实跑版同源) + +| 轮 | wall | out | in | TTFT mean/p50/max | TPOT mean/p50 | e2e mean/p50/max | 单流dec | accept | 命中 | 回退 | +|---|---|---|---|---|---|---|---|---|---|---| +| r1 | 124.2 | 263.8 | 2,110.4 | 10.36 / 10.54 / 17.36 | 50.5 / 50.6 | 62.0 / 62.0 / 62.4 | 20.0 | — | 0.0 | 0 | +| r2 | 123.2 | 266.1 | 2,128.4 | 10.38 / 10.54 / 17.36 | 50.0 / 50.0 | 61.5 / 61.5 / 61.8 | 20.1 | — | 0.0 | 0 | + +两轮间一致性:A ±5%、A16 ±9%(TPOT,在跨窗口噪声带内)、B ±1%。 + +## 三、关键发现 + +1. **A 生产口径在 cc16 有 decode 陷阱**:decode 批 bs 9–16 落在 cuda graph 之外(graph max bs=8),EAGLE 掉图后 TPOT 123ms,比 B 的无投机裸 decode(50ms)还慢 2.4×。历史"EAGLE 大胜"结论(cc1–4)不可外推到 cc16+:掉图模式下投机收益完全反转。 +2. **A16 的 graph 修复立竿见影**:TPOT 123→55ms(−55%)、out 115→219(+90%)、e2e P50 139.6→66.2s(−53%)。这是生产配置在该场景下零成本可得的改进(EXTRA 三个参数)。 +3. **EAGLE 在 cc16 高并发下优势摊薄至零**:A16 accept 2.50,但每步墙钟 ≈ 55ms×2.5 ≈ 137ms(verify 批 16×5=80 token/步打满算力),折算单 token 55ms ≈ B 裸 decode 50ms。投机毛利被 verify 批次吃掉。 +4. **prefill 是 A 家族剩余短板**:B(chunk16384/TP4PP2)输入吞吐 2,119 tok/s,A16 1,751、A 920。B 的优势与既有"prefill 密集 cc≥16 TP4PP2 最优"结论一致;本场景 decode 占比翻倍(o1k vs 历史 o512)后 B 仍全胜。 +5. **TTFT 的 p50/mean 悖论**:TTFT mean 三方案打平(~10s),但 p50 上 A/A16(~4s)优于 B(10.5s)。机理:nreq32 两波客户端下,A/A16 的 e2e 离散大→尾部请求陆续释放槽位,第二波"到站即过";B 全体同步完成,第二波整体重排 16 条 prefill。**对 TTFT-p50 敏感的小规模交互负载,A/A16 的 chunk8192 体验更好;对稳态吞吐,B 完胜。** +6. e2e 离散度:B 最优(p50−min/max 范围仅 0.5s),A16 中等(35–110s),A 最差(79–193s)。 + +## 四、收尾状态(2026-09-08 19:02 核验) + +- **60.8 在役**:容器 `glm53-nvfp4` = A16 口径(docker Args 双值并存、生效值 MRR32/graph bs 4 8 12 16,启动日志 `max_running_requests=32`、draft graph `bs=[4,8,12,16]` 实证),:30000,health 200,82.5GB×8,池 276,864。 +- A、B 已拆除,显存归零后才部署下一方案(每步 <2000MiB 纪律)。 +- 60.1 生产 glm53-pp4、60.5 团队生产全程未动。 +- **执行前经授权清理了 60.8 GPU3/6 的外部进程**:PID 2421844/2421852(user 账号 `/home/user/dirA_exp` 的 `eval_freeze_acc.py --mode remove --step 1/4`,14:49 起,各占 7.6GB、util 35%),SIGTERM 退出,显存归零。同源任务在 60.7 尚有 4 卡占用,未动。 + +## 五、复现与资产 + +- 压测命令(60.8 本机): + `python3 /root/bench_corpus.py --corpus /root/corpus_ids.json --input-len 8192 --output-len 1024 --concurrency 16 --num-requests 32 --run-id <新> --pool-override <新窗口> --shared-frac 0 --container glm53-nvfp4` +- 轮次编排:`/root/bench_8k.sh `(warm nreq16 + r1/r2 nreq32,步进 300k) +- 部署:A=`bash /root/deploy_glm53_605.sh`;A16=`EXTRA="--max-running-requests 32 --cuda-graph-max-bs-decode 16 --cuda-graph-bs-decode 4 8 12 16" bash /root/deploy_glm53_605.sh`;B=`bash /root/deploy_glm53_optimal.sh` +- **语料窗口台账(corpus_ids.json 共 21,296,780 token)**:本轮消耗 A [17.00M,17.86M)、B [17.90M,18.76M)、A16 [18.80M,19.66M),**已消费至 ~19.66M,剩余 ~1.63M**——后续 8k 复测须从 ≥19.7M 取新窗口或重建语料。 +- 指标口径:输入吞吐=8192×ok/墙钟;输出吞吐=服务端 completion_tokens 合计/墙钟(EAGLE 下客户端计数低估 2.5–3.5×,一律服务端口径);TTFT 含排队;TPOT=(首响应→末响应)/(completion−1);P50=单请求 e2e 延迟中位数;命中由容器 TP0 Prefill 日志核算(全 0,全唯一验证通过)。 +- 日志归档:60.8 `/root/bench_logs/`(含 3×gate、9 轮 bench、runner/deploy 日志)+ 本地 tar `D:\sskj\_xfer_608\i8k_bench_20260908.tar.gz`;入库 sskj-review `experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/`。 + +## 六、与既有结论的关系 + +- 与《双场景压测报告》场景二(16k/512)一致性:B 全胜的结论在 o1k(decode 占比翻倍)下依然成立且差距未缩小——说明 B 的优势主要来自 prefill 侧与池/调度,而非 decode 占比。 +- 新增修正:**"A 生产口径"在 cc16 的历史场景二数据实际是 A16 变体口径**(tp8eagle.env 注释的场景二变体);本报告首次把 A 原口径与 A16 在同一场景下分离量化——两者差距高达 1.9×,引用历史数据时须区分口径。 diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/README.md b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/README.md new file mode 100644 index 0000000..74a7ea7 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/README.md @@ -0,0 +1,31 @@ +# GLM-5.3-NVFP4 i8k/o1k/c16 三方案部署对比压测(2026-09-08) + +- 机器:174.1.60.8(6000D-8,8×H20,单机串行三方案轮换) +- 场景:input=8192 / output=1024 / cc=16 / nreq=32(2 波),全唯一 PG19 真实语料,temp=0、ignore_eos、stream +- 完整报告:本目录 `GLM53_NVFP4_i8k_o1k_c16_压测报告_2026-09-08.md`(判决、主表、机理分析、复现命令) + +## 方案与结果(r1/r2 均值) + +| 方案 | 配置要点 | out tok/s | in tok/s | TTFT p50 | TPOT p50 | e2e P50 | 质量门 | +|---|---|---|---|---|---|---|---| +| A 生产口径 | TP8+EAGLE 4/1/5,MRR16,chunk8192,graph bs≤8,池 276,864 | 115.0 | 920 | 3.8s | 123ms | 139.6s | 7/7 | +| A16 | A + `--max-running-requests 32 --cuda-graph-max-bs-decode 16 --cuda-graph-bs-decode 4 8 12 16` | 218.9 | 1,751 | 4.0s | 55ms | 66.2s | 7/7 | +| B | TP4PP2,MRR48,chunk16384,radix-off,池 569,600 | **264.9** | **2,119** | 10.5s | **50ms** | **61.8s** | 6/7* | + +\* B 的 tool-call 失败 = 脚本无 `--tool-call-parser`(已知配置缺口,非质量问题)。 + +**判决**:B 全面最优(out 为 A 的 2.30×);A16 把差距收窄到 1.21×,是 EAGLE 家族在本场景的正确口径。 +60.8 收官保留 **A16 在役**(:30000,restart=unless-stopped)。 + +## 核心结论(引用时注意) + +1. **A 生产口径在 cc16 有 decode 陷阱**:bs 9–16 掉出 cuda graph,EAGLE 掉图后 TPOT 123ms,比 B 裸 decode(50ms)慢 2.4×——"EAGLE 大胜"仅成立于 cc≤4 或 graph 覆盖到位时。 +2. **EAGLE 在 cc16 优势摊薄至零**:A16 accept 2.50,但 verify 批 80 token/步打满算力,折算单 token 55ms ≈ B 的 50ms。 +3. **历史场景二(16k/512)的"A"行实为 A16 变体口径**,与 A 原口径相差 1.9×,引用历史数据须区分。 + +## 资产 + +- `scripts/`:bench_8k.sh(轮次编排)、bench_corpus.py(c6d92126,60.7 同源)、deploy_glm53_605.sh(fcd9109b,8 机标准)、deploy_glm53_optimal.sh(22685f56) +- `results/bench_logs/`:3×gate(A 7/7、B 6/7、A16 7/7)、9 轮 bench 日志(A/A16/B × warm/r1/r2)、runner/deploy 日志 +- 语料窗口:本轮消费 corpus_ids.json 至 **~19.66M / 21.30M**(A [17.0,17.86M)、B [17.9,18.76M)、A16 [18.8,19.66M)),剩余 ~1.63M,复测须取 ≥19.7M 新窗口 +- 压测机执行前经授权清理了 60.8 GPU3/6 的 dirA_ext 外部 eval 进程(PID 2421844/2421852,SIGTERM) diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A16_r1_summary.json b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A16_r1_summary.json new file mode 100644 index 0000000..4e75643 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A16_r1_summary.json @@ -0,0 +1,57 @@ +{ + "concurrency": 16, + "num_requests": 32, + "run_id": 9411, + "corpus_window": { + "start": 19100000, + "end": 19362144 + }, + "ok": 32, + "failed": 0, + "wall_s": 151.31, + "input_len": 8192, + "shared_len": 0, + "unique_len": 8192, + "output_len": 1024, + "output_tokens_total": 32768, + "output_throughput_tok_s": 216.55, + "input_throughput_tok_s": 1732.44, + "ttft_s": { + "mean": 10.0558, + "p50": 4.0156, + "max": 30.6315, + "min": 2.0783 + }, + "tpot_s": { + "mean": 0.0606, + "p50": 0.0577, + "max": 0.1026, + "min": 0.039 + }, + "e2e_s": { + "mean": 72.0651, + "p50": 69.9843, + "max": 108.1261, + "min": 41.9544 + }, + "per_req_out_tok_s_e2e": { + "mean": 15.3242, + "p50": 14.6458, + "max": 24.4074, + "min": 9.4704 + }, + "per_req_decode_tok_s": { + "mean": 17.6898, + "p50": 18.553, + "max": 25.6827, + "min": 9.7574 + }, + "spec_accept_length_mean": 2.367, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 32, + "new_tokens": 262144, + "cached_tokens": 0, + "hit_rate": 0.0 + } +} diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A16_r2_summary.json b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A16_r2_summary.json new file mode 100644 index 0000000..ad627f7 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A16_r2_summary.json @@ -0,0 +1,57 @@ +{ + "concurrency": 16, + "num_requests": 32, + "run_id": 9412, + "corpus_window": { + "start": 19400000, + "end": 19662144 + }, + "ok": 32, + "failed": 0, + "wall_s": 148.17, + "input_len": 8192, + "shared_len": 0, + "unique_len": 8192, + "output_len": 1024, + "output_tokens_total": 32768, + "output_throughput_tok_s": 221.15, + "input_throughput_tok_s": 1769.19, + "ttft_s": { + "mean": 9.875, + "p50": 3.9786, + "max": 30.6695, + "min": 2.0794 + }, + "tpot_s": { + "mean": 0.0573, + "p50": 0.0522, + "max": 0.102, + "min": 0.0307 + }, + "e2e_s": { + "mean": 68.4525, + "p50": 62.3499, + "max": 109.5734, + "min": 35.5019 + }, + "per_req_out_tok_s_e2e": { + "mean": 16.5606, + "p50": 16.6965, + "max": 28.8435, + "min": 9.3453 + }, + "per_req_decode_tok_s": { + "mean": 19.2407, + "p50": 19.179, + "max": 32.6448, + "min": 9.8182 + }, + "spec_accept_length_mean": 2.634, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 32, + "new_tokens": 262144, + "cached_tokens": 0, + "hit_rate": 0.0 + } +} diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A16_warm_summary.json b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A16_warm_summary.json new file mode 100644 index 0000000..e61767b --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A16_warm_summary.json @@ -0,0 +1,57 @@ +{ + "concurrency": 16, + "num_requests": 16, + "run_id": 9401, + "corpus_window": { + "start": 18800000, + "end": 18931072 + }, + "ok": 16, + "failed": 0, + "wall_s": 90.65, + "input_len": 8192, + "shared_len": 0, + "unique_len": 8192, + "output_len": 1024, + "output_tokens_total": 16384, + "output_throughput_tok_s": 180.74, + "input_throughput_tok_s": 1445.94, + "ttft_s": { + "mean": 29.6679, + "p50": 30.755, + "max": 44.1152, + "min": 15.0222 + }, + "tpot_s": { + "mean": 0.0522, + "p50": 0.0521, + "max": 0.0739, + "min": 0.0306 + }, + "e2e_s": { + "mean": 83.0523, + "p50": 87.6337, + "max": 90.6464, + "min": 69.7624 + }, + "per_req_out_tok_s_e2e": { + "mean": 12.4314, + "p50": 11.9193, + "max": 14.6784, + "min": 11.2966 + }, + "per_req_decode_tok_s": { + "mean": 20.6139, + "p50": 20.1971, + "max": 32.678, + "min": 13.5407 + }, + "spec_accept_length_mean": 2.643, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 16, + "new_tokens": 131072, + "cached_tokens": 0, + "hit_rate": 0.0 + } +} diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A_r1_summary.json b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A_r1_summary.json new file mode 100644 index 0000000..658bb65 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A_r1_summary.json @@ -0,0 +1,57 @@ +{ + "concurrency": 16, + "num_requests": 32, + "run_id": 9411, + "corpus_window": { + "start": 17300000, + "end": 17562144 + }, + "ok": 32, + "failed": 0, + "wall_s": 277.41, + "input_len": 8192, + "shared_len": 0, + "unique_len": 8192, + "output_len": 1024, + "output_tokens_total": 32768, + "output_throughput_tok_s": 118.12, + "input_throughput_tok_s": 944.97, + "ttft_s": { + "mean": 10.2921, + "p50": 4.3675, + "max": 31.1779, + "min": 2.3895 + }, + "tpot_s": { + "mean": 0.1196, + "p50": 0.1184, + "max": 0.1667, + "min": 0.075 + }, + "e2e_s": { + "mean": 132.6294, + "p50": 136.2478, + "max": 190.146, + "min": 79.1449 + }, + "per_req_out_tok_s_e2e": { + "mean": 8.1891, + "p50": 7.536, + "max": 12.9383, + "min": 5.3853 + }, + "per_req_decode_tok_s": { + "mean": 8.7968, + "p50": 8.7859, + "max": 13.3427, + "min": 6.003 + }, + "spec_accept_length_mean": 2.589, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 32, + "new_tokens": 262144, + "cached_tokens": 0, + "hit_rate": 0.0 + } +} diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A_r2_summary.json b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A_r2_summary.json new file mode 100644 index 0000000..3908b85 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A_r2_summary.json @@ -0,0 +1,57 @@ +{ + "concurrency": 16, + "num_requests": 32, + "run_id": 9412, + "corpus_window": { + "start": 17600000, + "end": 17862144 + }, + "ok": 32, + "failed": 0, + "wall_s": 292.98, + "input_len": 8192, + "shared_len": 0, + "unique_len": 8192, + "output_len": 1024, + "output_tokens_total": 32768, + "output_throughput_tok_s": 111.84, + "input_throughput_tok_s": 894.75, + "ttft_s": { + "mean": 10.0189, + "p50": 3.2165, + "max": 30.9377, + "min": 2.4035 + }, + "tpot_s": { + "mean": 0.1269, + "p50": 0.1278, + "max": 0.1762, + "min": 0.077 + }, + "e2e_s": { + "mean": 139.8173, + "p50": 142.9177, + "max": 193.13, + "min": 81.2596 + }, + "per_req_out_tok_s_e2e": { + "mean": 7.6338, + "p50": 7.6204, + "max": 12.6016, + "min": 5.3021 + }, + "per_req_decode_tok_s": { + "mean": 8.1582, + "p50": 7.9203, + "max": 13.0008, + "min": 5.6814 + }, + "spec_accept_length_mean": 2.445, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 32, + "new_tokens": 262144, + "cached_tokens": 0, + "hit_rate": 0.0 + } +} diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A_warm_summary.json b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A_warm_summary.json new file mode 100644 index 0000000..543e366 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/A_warm_summary.json @@ -0,0 +1,57 @@ +{ + "concurrency": 16, + "num_requests": 16, + "run_id": 9401, + "corpus_window": { + "start": 17000000, + "end": 17131072 + }, + "ok": 16, + "failed": 0, + "wall_s": 172.19, + "input_len": 8192, + "shared_len": 0, + "unique_len": 8192, + "output_len": 1024, + "output_tokens_total": 16384, + "output_throughput_tok_s": 95.15, + "input_throughput_tok_s": 761.21, + "ttft_s": { + "mean": 29.1555, + "p50": 30.0011, + "max": 45.0124, + "min": 14.441 + }, + "tpot_s": { + "mean": 0.1241, + "p50": 0.1326, + "max": 0.1482, + "min": 0.0785 + }, + "e2e_s": { + "mean": 156.117, + "p50": 167.158, + "max": 172.1581, + "min": 114.2218 + }, + "per_req_out_tok_s_e2e": { + "mean": 6.7017, + "p50": 6.1651, + "max": 8.965, + "min": 5.948 + }, + "per_req_decode_tok_s": { + "mean": 8.3659, + "p50": 7.9182, + "max": 12.7508, + "min": 6.7522 + }, + "spec_accept_length_mean": 2.501, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 16, + "new_tokens": 131072, + "cached_tokens": 0, + "hit_rate": 0.0 + } +} diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/B_r1_summary.json b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/B_r1_summary.json new file mode 100644 index 0000000..00976f3 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/B_r1_summary.json @@ -0,0 +1,57 @@ +{ + "concurrency": 16, + "num_requests": 32, + "run_id": 9411, + "corpus_window": { + "start": 18200000, + "end": 18462144 + }, + "ok": 32, + "failed": 0, + "wall_s": 124.22, + "input_len": 8192, + "shared_len": 0, + "unique_len": 8192, + "output_len": 1024, + "output_tokens_total": 32768, + "output_throughput_tok_s": 263.8, + "input_throughput_tok_s": 2110.38, + "ttft_s": { + "mean": 10.3634, + "p50": 10.5406, + "max": 17.3639, + "min": 1.777 + }, + "tpot_s": { + "mean": 0.0505, + "p50": 0.0506, + "max": 0.0592, + "min": 0.0434 + }, + "e2e_s": { + "mean": 62.019, + "p50": 62.0417, + "max": 62.366, + "min": 61.8 + }, + "per_req_out_tok_s_e2e": { + "mean": 16.5113, + "p50": 16.5486, + "max": 16.5696, + "min": 16.4192 + }, + "per_req_decode_tok_s": { + "mean": 19.9851, + "p50": 19.9608, + "max": 23.039, + "min": 16.9011 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 36, + "new_tokens": 524288, + "cached_tokens": 0, + "hit_rate": 0.0 + } +} diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/B_r2_summary.json b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/B_r2_summary.json new file mode 100644 index 0000000..9750e20 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/B_r2_summary.json @@ -0,0 +1,57 @@ +{ + "concurrency": 16, + "num_requests": 32, + "run_id": 9412, + "corpus_window": { + "start": 18500000, + "end": 18762144 + }, + "ok": 32, + "failed": 0, + "wall_s": 123.16, + "input_len": 8192, + "shared_len": 0, + "unique_len": 8192, + "output_len": 1024, + "output_tokens_total": 32768, + "output_throughput_tok_s": 266.05, + "input_throughput_tok_s": 2128.41, + "ttft_s": { + "mean": 10.3812, + "p50": 10.5386, + "max": 17.3626, + "min": 1.7756 + }, + "tpot_s": { + "mean": 0.05, + "p50": 0.05, + "max": 0.0587, + "min": 0.043 + }, + "e2e_s": { + "mean": 61.5036, + "p50": 61.4987, + "max": 61.8083, + "min": 61.3264 + }, + "per_req_out_tok_s_e2e": { + "mean": 16.6495, + "p50": 16.6689, + "max": 16.6975, + "min": 16.5674 + }, + "per_req_decode_tok_s": { + "mean": 20.1968, + "p50": 20.1428, + "max": 23.2803, + "min": 17.0575 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 36, + "new_tokens": 524288, + "cached_tokens": 0, + "hit_rate": 0.0 + } +} diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/B_warm_summary.json b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/B_warm_summary.json new file mode 100644 index 0000000..71b36cc --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/results/bench_logs/B_warm_summary.json @@ -0,0 +1,57 @@ +{ + "concurrency": 16, + "num_requests": 16, + "run_id": 9401, + "corpus_window": { + "start": 17900000, + "end": 18031072 + }, + "ok": 16, + "failed": 0, + "wall_s": 65.46, + "input_len": 8192, + "shared_len": 0, + "unique_len": 8192, + "output_len": 1024, + "output_tokens_total": 16384, + "output_throughput_tok_s": 250.29, + "input_throughput_tok_s": 2002.32, + "ttft_s": { + "mean": 13.2299, + "p50": 13.271, + "max": 20.2234, + "min": 4.2816 + }, + "tpot_s": { + "mean": 0.0509, + "p50": 0.0507, + "max": 0.0598, + "min": 0.0439 + }, + "e2e_s": { + "mean": 65.2664, + "p50": 65.1682, + "max": 65.4579, + "min": 65.1135 + }, + "per_req_out_tok_s_e2e": { + "mean": 15.6896, + "p50": 15.7139, + "max": 15.7264, + "min": 15.6436 + }, + "per_req_decode_tok_s": { + "mean": 19.845, + "p50": 19.7319, + "max": 22.7938, + "min": 16.7391 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 18, + "new_tokens": 262144, + "cached_tokens": 0, + "hit_rate": 0.0 + } +} diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/scripts/bench_8k.sh b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/scripts/bench_8k.sh new file mode 100644 index 0000000..d6231d1 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/scripts/bench_8k.sh @@ -0,0 +1,26 @@ +#!/bin/bash +# bench_8k.sh — i8k/o1k/c16 bench rounds for ONE scheme (run via nohup on 60.8) +# usage: bash bench_8k.sh +# Rounds: warm (nreq16, discard) + r1/r2 (nreq32, measured), fresh corpus window each. +set -u +TAG=$1 +POOL=$2 +mkdir -p /root/bench_logs + +run_round() { + local name=$1 nreq=$2 pool=$3 rid=$4 + echo "=== [$TAG] round $name nreq=$nreq pool=$pool rid=$rid start=$(date -Is) ===" + python3 /root/bench_corpus.py --corpus /root/corpus_ids.json \ + --input-len 8192 --output-len 1024 --concurrency 16 \ + --num-requests "$nreq" --run-id "$rid" --pool-override "$pool" \ + --shared-frac 0 --container glm53-nvfp4 \ + > "/root/bench_logs/${TAG}_${name}.log" 2>&1 + echo "=== [$TAG] round $name rc=$? end=$(date -Is) ===" + grep -c '"ok"' "/root/bench_logs/${TAG}_${name}.log" >/dev/null 2>&1 + tail -3 "/root/bench_logs/${TAG}_${name}.log" +} + +run_round warm 16 "$POOL" 9401 +run_round r1 32 $((POOL + 300000)) 9411 +run_round r2 32 $((POOL + 600000)) 9412 +echo "=== [$TAG] ALL DONE $(date -Is) ===" diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/scripts/bench_corpus.py b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/scripts/bench_corpus.py new file mode 100644 index 0000000..564090b --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/scripts/bench_corpus.py @@ -0,0 +1,298 @@ +#!/usr/bin/env python3 +"""Real-corpus (PG19) benchmark for sglang GLM-5.3-NVFP4. + +Same methodology as bench_hit90.py (input_ids direct to /generate, temp 0, +ignore_eos, stream, server-side completion_tokens counting, hit-rate verified +from scheduler logs), but prompts are token slices of REAL book text tokenized +with the served model's own tokenizer, replacing random ids. + +Corpus file: JSON {"ids": [flat token ids], "books": [[start, end), ...]} + +Fixed token-offset layout into the flat id array: + [0, 117968) shared prefix for 128k points (90% of 131072) + [0, 58976) shared prefix for 64k points (same region, shorter cut) + pool A [131072, +4*8*13104) 128k unique suffixes, run-ids 9301-9304 + pool B [550400, +4*8*6560) 64k unique suffixes, run-ids 9305-9308 + pool C [760320, 262144+524288+524288) 16k fully-unique prompts, run-ids 9311-9313 + spare [2071040, end) warmup slices / re-run margin + +Each run-id maps to one non-overlapping window (one window = one bench point); +re-running a point with fresh text = bump --pool-override past the spare base. + +Usage: + python3 bench_corpus.py --corpus /root/corpus_ids.json --input-len 131072 \ + --concurrency 1 --num-requests 8 --run-id 9301 --shared-frac 0.9 +""" +import argparse +import datetime +import json +import re +import statistics +import subprocess +import time +from concurrent.futures import ThreadPoolExecutor + +import requests + +OUTPUT_LEN_DEFAULT = 512 +CORPUS_DEFAULT = "/root/corpus_ids.json" + +# fixed pool layout (see docstring) +S1_128K_RID0, S1_64K_RID0, S2_RID0 = 9301, 9305, 9311 +POOL_A_BASE, POOL_A_PER = 131072, 8 * 13104 # 128k suffix windows +POOL_B_BASE = POOL_A_BASE + 4 * POOL_A_PER # 550400 +POOL_B_PER = 8 * 6560 # 64k suffix windows +POOL_C_BASE = POOL_B_BASE + 4 * POOL_B_PER # 760320 +POOL_C_SIZES = [16 * 16384, 32 * 16384, 32 * 16384] # cc8 / cc16 / cc32 +SPARE_BASE = POOL_C_BASE + sum(POOL_C_SIZES) # 2071040 + +sess = requests.Session() +sess.trust_env = False # bypass any proxy env on the host + + +def split_lens(input_len, shared_frac): + # unique suffix = (1 - shared_frac) of the prompt, page-16 aligned + # (128k @0.9 -> 13104 unique; 16k @0.0 -> fully unique prompts) + unique = round(input_len * (1.0 - shared_frac) / 16) * 16 + return input_len - unique, unique + + +def pool_start_for(input_len, shared_frac, run_id, override): + if override is not None: + return override + if shared_frac > 0: + if input_len == 131072: + idx = run_id - S1_128K_RID0 + if not 0 <= idx < 4: + sys_exit_bad_runid(run_id, "128k points use run-ids 9301-9304") + return POOL_A_BASE + idx * POOL_A_PER + if input_len == 65536: + idx = run_id - S1_64K_RID0 + if not 0 <= idx < 4: + sys_exit_bad_runid(run_id, "64k points use run-ids 9305-9308") + return POOL_B_BASE + idx * POOL_B_PER + sys_exit_bad_runid(run_id, "shared-frac>0 supports 131072/65536 only") + idx = run_id - S2_RID0 + if not 0 <= idx < 3: + sys_exit_bad_runid(run_id, "16k unique points use run-ids 9311-9313") + return POOL_C_BASE + sum(POOL_C_SIZES[:idx]) + + +def sys_exit_bad_runid(run_id, msg): + raise SystemExit(f"[pool] run-id {run_id} outside expected set: {msg}") + + +def build_prompts(ids, shared_len, unique_len, num_requests, pool_start): + if shared_len: + shared = ids[0:shared_len] + else: + shared = [] + end = pool_start + num_requests * unique_len + if end > len(ids): + raise SystemExit( + f"[pool] window [{pool_start}, {end}) exceeds corpus ({len(ids)} ids); " + f"use --pool-override or a larger corpus") + prompts = [] + for i in range(num_requests): + s = pool_start + i * unique_len + prompts.append(shared + ids[s:s + unique_len]) + return shared, prompts, (pool_start, end) + + +def warmup(url, ids, shared_len): + # primes the radix cache with the shared prefix (same role as in bench_hit90); + # warm slice comes from the spare region so it never collides with a pool window + if len(ids) >= SPARE_BASE + 64: + warm_slice = ids[SPARE_BASE:SPARE_BASE + 64] + else: + warm_slice = ids[-64:] + payload = { + "input_ids": ids[0:shared_len] + warm_slice if shared_len else warm_slice, + "sampling_params": {"max_new_tokens": 8, "temperature": 0.0, "ignore_eos": True}, + } + t0 = time.perf_counter() + r = sess.post(url, json=payload, timeout=1800) + dt = time.perf_counter() - t0 + print(f"[warmup] http={r.status_code} wall={dt:.2f}s", flush=True) + + +def bench_one(url, prompt, output_len, idx, results): + payload = { + "input_ids": prompt, + "sampling_params": {"max_new_tokens": output_len, "temperature": 0.0, "ignore_eos": True}, + "stream": True, + } + rec = {"idx": idx} + t0 = time.perf_counter() + first = last = None + first_ct = None + final_meta = None + max_ct = 0 + try: + with sess.post(url, json=payload, stream=True, timeout=3600) as resp: + for raw in resp.iter_lines(): + if not raw or not raw.startswith(b"data:"): + continue + body = raw[5:].strip() + if body == b"[DONE]": + continue + now = time.perf_counter() + try: + d = json.loads(body) + except Exception: + continue + mi = d.get("meta_info") or {} + ct = mi.get("completion_tokens") or 0 + if ct: + max_ct = max(max_ct, ct) + if first is None: + first = now + first_ct = ct + last = now + if mi.get("finish_reason"): + final_meta = mi + t_end = time.perf_counter() + n_out = max_ct or ((final_meta or {}).get("completion_tokens") or 0) + decode_span = (last - first) if (first and last and last > first) else 0.0 + rec.update( + ok=n_out > 0, + ttft=(first - t0) if first else None, + e2e=t_end - t0, + n_out=n_out, + first_chunk_tokens=first_ct, + decode_span=decode_span, + tpot=(decode_span / (n_out - 1)) if n_out > 1 else None, + per_req_decode_tok_s=(n_out / decode_span) if decode_span > 0 else None, + retractions=(final_meta or {}).get("num_retractions"), + spec_accept_len=(final_meta or {}).get("spec_accept_length"), + ) + except Exception as e: + rec.update(ok=False, error=repr(e)) + results[idx] = rec + + +def verify_hit_rate(container, t_start, t_end): + def rfc3339(epoch): + return (datetime.datetime.fromtimestamp(epoch, tz=datetime.timezone.utc) + .isoformat().replace("+00:00", "Z")) + + try: + # No margin before t_start: warmup's prefill lines end strictly before it, + # and catching them would deflate the measured hit rate. + p = subprocess.run( + ["docker", "logs", container, "--since", rfc3339(t_start), "--until", rfc3339(t_end + 2)], + capture_output=True, text=True, timeout=120) + text = p.stdout + p.stderr + except Exception as e: + return {"error": repr(e)} + pat = re.compile(r"#new-token: (\d+).*?#cached-token: (\d+)") + n_batches = new_tok = cached_tok = 0 + for line in text.splitlines(): + if "TP0]" not in line or "Prefill batch" not in line: + continue + m = pat.search(line) + if m: + n_batches += 1 + new_tok += int(m.group(1)) + cached_tok += int(m.group(2)) + total = new_tok + cached_tok + return { + "prefill_batches": n_batches, + "new_tokens": new_tok, + "cached_tokens": cached_tok, + "hit_rate": round(cached_tok / total, 4) if total else None, + } + + +def stats(vals): + vals = [v for v in vals if v is not None] + if not vals: + return {"mean": None, "p50": None, "max": None, "min": None} + s = sorted(vals) + return { + "mean": round(statistics.fmean(vals), 4), + "p50": round(s[len(s) // 2], 4), + "max": round(s[-1], 4), + "min": round(s[0], 4), + } + + +def main(): + ap = argparse.ArgumentParser() + ap.add_argument("--concurrency", type=int, required=True) + ap.add_argument("--num-requests", type=int, required=True) + ap.add_argument("--run-id", type=int, required=True) + ap.add_argument("--input-len", type=int, required=True, help="16384 / 65536 / 131072") + ap.add_argument("--output-len", type=int, default=OUTPUT_LEN_DEFAULT) + ap.add_argument("--shared-frac", type=float, default=0.9) + ap.add_argument("--corpus", default=CORPUS_DEFAULT) + ap.add_argument("--pool-override", type=int, default=None, + help="explicit corpus offset for the unique-suffix window (re-runs)") + ap.add_argument("--url", default="http://127.0.0.1:30000/generate") + ap.add_argument("--container", default="glm53-nvfp4") + args = ap.parse_args() + + with open(args.corpus) as f: + corpus = json.load(f) + ids = corpus["ids"] + + shared_len, unique_len = split_lens(args.input_len, args.shared_frac) + pool_start = pool_start_for(args.input_len, args.shared_frac, args.run_id, args.pool_override) + shared, prompts, window = build_prompts(ids, shared_len, unique_len, args.num_requests, pool_start) + print(f"[pool] window={window} shared_len={shared_len} unique_len={unique_len} " + f"corpus_total={len(ids)}", flush=True) + warmup(args.url, ids, shared_len) + + results = {} + t_start = time.time() + t0 = time.perf_counter() + with ThreadPoolExecutor(max_workers=args.concurrency) as ex: + futs = [ex.submit(bench_one, args.url, p, args.output_len, i, results) + for i, p in enumerate(prompts)] + for f in futs: + f.result() + wall = time.perf_counter() - t0 + t_end = time.time() + + hit = verify_hit_rate(args.container, t_start, t_end) + + ok = [r for r in results.values() if r.get("ok")] + n_out_total = sum(r["n_out"] for r in ok) + out_tps = [r["n_out"] / r["e2e"] for r in ok if r.get("e2e")] + ttft = stats([r.get("ttft") for r in ok]) + tpot = stats([r.get("tpot") for r in ok]) + e2e = stats([r.get("e2e") for r in ok]) + dec = stats([r.get("per_req_decode_tok_s") for r in ok]) + spec = [v for v in (r.get("spec_accept_len") for r in ok) if v is not None] + retr = sum(r.get("retractions") or 0 for r in ok) + + summary = { + "concurrency": args.concurrency, + "num_requests": args.num_requests, + "run_id": args.run_id, + "corpus_window": {"start": window[0], "end": window[1]}, + "ok": len(ok), + "failed": args.num_requests - len(ok), + "wall_s": round(wall, 2), + "input_len": args.input_len, + "shared_len": shared_len, + "unique_len": unique_len, + "output_len": args.output_len, + "output_tokens_total": n_out_total, + "output_throughput_tok_s": round(n_out_total / wall, 2) if wall else None, + "input_throughput_tok_s": round(args.input_len * len(ok) / wall, 2) if wall else None, + "ttft_s": ttft, + "tpot_s": tpot, + "e2e_s": e2e, + "per_req_out_tok_s_e2e": stats(out_tps), + "per_req_decode_tok_s": dec, + "spec_accept_length_mean": round(statistics.fmean(spec), 3) if spec else None, + "retractions_total": retr, + "cache_hit_from_logs": hit, + } + print("\n===== SUMMARY =====") + print(json.dumps(summary, indent=2), flush=True) + + +if __name__ == "__main__": + main() diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/scripts/deploy_glm53_605.sh b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/scripts/deploy_glm53_605.sh new file mode 100644 index 0000000..fe0f108 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/scripts/deploy_glm53_605.sh @@ -0,0 +1,58 @@ +#!/bin/bash +# deploy_glm53_605.sh — GLM-5.3-NVFP4 TP8 + EAGLE for 174.1.60.5 (team service) +# Target scenario: cc1-2, 64k/128k input 90% cache hit, single request up to 256k, +# output-throughput priority. Clean deploy: rm old container, wait VRAM drain, run. +# +# Env overrides: +# MEMFRAC (0.90) STEPS (4) TOPK (1) DRAFT (5) CTXLEN (270336) +# CHUNK (8192) MAXPRE (16384) RESTART (no|yes) EXTRA ("") +# Usage: +# bash deploy_glm53_605.sh # production 4/1/5 @ 0.90 +# STEPS=5 DRAFT=6 bash deploy_glm53_605.sh # tuning round +# CHUNK=16384 bash deploy_glm53_605.sh # prefill tuning round +# RESTART=yes bash deploy_glm53_605.sh # production finalize +set -e +MEMFRAC=${MEMFRAC:-0.90} +STEPS=${STEPS:-4} +TOPK=${TOPK:-1} +DRAFT=${DRAFT:-5} +CTXLEN=${CTXLEN:-270336} +CHUNK=${CHUNK:-8192} +MAXPRE=${MAXPRE:-16384} +EXTRA=${EXTRA:-} +if [ "$RESTART" = "yes" ]; then RP="--restart unless-stopped"; else RP="--restart no"; fi + +echo "[deploy] removing old container (if any)" +# rm -f times out on big GPU containers on this daemon; retry until really gone +for i in $(seq 1 45); do + CID=$(docker ps -a --filter name=glm53-nvfp4 -q) + [ -z "$CID" ] && break + docker rm -f glm53-nvfp4 >/dev/null 2>&1 || true + sleep 2 +done + +echo "[deploy] waiting for VRAM drain" +for i in $(seq 1 45); do + used=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | awk '{s+=$1} END {print s}') + [ "$used" -lt 2000 ] && break + sleep 2 +done +used=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | awk '{s+=$1} END {print s}') +echo "[deploy] VRAM now: ${used} MiB total" + +echo "[deploy] starting: TP8 EAGLE ${STEPS}/${TOPK}/${DRAFT} memfrac=${MEMFRAC} extra='${EXTRA}'" +docker run -d --name glm53-nvfp4 --gpus all --shm-size 64g --ipc=host $RP \ + -p 30000:30000 -v /data/hf_models:/data/hf_models \ + lmsysorg/sglang:nightly-dev-20260828-daf63171 \ + python3 -m sglang.launch_server \ + --model-path /data/hf_models/GLM-5.3-NVFP4 --tp 8 \ + --mem-fraction-static $MEMFRAC --max-running-requests 16 \ + --chunked-prefill-size $CHUNK --max-prefill-tokens $MAXPRE \ + --disable-shared-experts-fusion --moe-runner-backend flashinfer_cutlass --disable-flashinfer-autotune \ + --speculative-algorithm EAGLE --speculative-num-steps $STEPS --speculative-eagle-topk $TOPK --speculative-num-draft-tokens $DRAFT \ + --kv-cache-dtype fp8_e4m3 --enable-hierarchical-cache --hicache-ratio 3 \ + --cuda-graph-max-bs-decode 8 --cuda-graph-bs-decode 1 2 3 4 6 8 --cuda-graph-max-bs-prefill 8 \ + --context-length $CTXLEN --reasoning-parser glm45 --tool-call-parser glm47 \ + --host 0.0.0.0 --port 30000 $EXTRA + +echo "[deploy] container started; poll: docker logs -f glm53-nvfp4" diff --git a/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/scripts/deploy_glm53_optimal.sh b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/scripts/deploy_glm53_optimal.sh new file mode 100644 index 0000000..7faa6f8 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/scripts/deploy_glm53_optimal.sh @@ -0,0 +1,62 @@ +#!/bin/bash +# ============================================================ +# GLM-5.3 最优部署方案(6000D 8卡,TP4 PP2 + IndexCache freq=4) +# 2026-09-07 +# +# - 基线配置:TP4 PP2 + cps16k + mem0.85(131.6 tok/s 吞吐基线) +# - IndexCache(index_topk_freq=4):层轴索引复用,省 75% indexer +# 16K 场景无损失;128K 长上下文并发 1.35-1.47× 提速 +# - 禁 radix cache;禁投机解码(PP2 与投机框架不兼容,已实测) +# +# 用法:bash deploy_glm53_optimal.sh +# (60.7 试用版:增加显存排空等待 + 就绪等待加长至 20 分钟) +# ============================================================ +set -uo pipefail + +CONTAINER="glm53-nvfp4" +IMAGE="lmsysorg/sglang:nightly-dev-20260828-daf63171" +MODEL="/data/hf_models/GLM-5.3-NVFP4" +PORT=30000 +TP=4; PP=2; MEM=0.85; MRR=48; CPS=16384 + +docker rm -f ${CONTAINER} 2>/dev/null || true + +# docker rm -f 后显存释放滞后数分钟,不等会把新容器 KV 池压小 +for i in $(seq 1 60); do + used=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | awk '{s+=$1} END {print s}') + if [ "$used" -lt 1500 ]; then echo "[deploy] drained: ${used} MiB"; break; fi + echo "[deploy] drain wait ${i}: ${used} MiB" + sleep 10 +done + +docker run -d --name ${CONTAINER} --gpus all --shm-size 64g --ipc=host \ + --restart unless-stopped \ + -p ${PORT}:${PORT} \ + -v /data/hf_models:/data/hf_models \ + ${IMAGE} \ + python3 -m sglang.launch_server \ + --model-path ${MODEL} \ + --tp-size ${TP} --pp-size ${PP} \ + --mem-fraction-static ${MEM} \ + --max-running-requests ${MRR} \ + --disable-radix-cache \ + --disable-shared-experts-fusion \ + --moe-runner-backend flashinfer_cutlass \ + --disable-flashinfer-autotune \ + --disable-custom-all-reduce \ + --chunked-prefill-size ${CPS} \ + --host 0.0.0.0 --port ${PORT} \ + --json-model-override-args '{"index_topk_freq": 4}' + +echo "容器已启动,等待就绪..." +for i in $(seq 1 120); do + code=$(curl -s -o /dev/null -w '%{http_code}' http://localhost:${PORT}/health 2>/dev/null) + if [ "$code" = "200" ]; then + echo "READY after ~$((i*10))s" + docker ps --filter name=${CONTAINER} --format '{{.Names}} {{.Status}}' + echo "override args: $(docker inspect ${CONTAINER} --format '{{.Config.Cmd}}' | grep -o 'index_topk_freq[^,}]*')" + exit 0 + fi + sleep 10 +done +echo "TIMEOUT"; exit 1