feat(pro6000/GLM-5.3): i8k/o1k/c16 三方案对比压测入库(A/A16/B 轮换实测 + 60.8 转 A16 在役 + CURRENT.md 台账修正 60.2/60.3/60.6/60.8)
This commit is contained in:
parent
5c749cda03
commit
123023b6ae
@ -8,13 +8,13 @@
|
||||
| 机器 | 在役 | 口径 / 归属 | 对应 profile |
|
||||
|---|---|---|---|
|
||||
| 60.1 (6000D-1) | `glm53-pp4`(Up,8 卡满载,:30000) | **方案 D 生产**。09-08 PD 压测窗口停机 ~2.5h 后已恢复并核验 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp2pp4.env` |
|
||||
| 60.2 (6000D-2) | 空(无容器,GPU 全 0 MiB) | 方案 F PD 链已于 09-08 拆除。GPU7 曾有外部裸金属任务 main_v2.py(现已结束),动卡前仍先核实归属 | `profiles/pro6000/glm53_nvfp4_pro6000_pd_{prefill→decode 侧}.env` |
|
||||
| 60.3 | 空(无容器) | — | — |
|
||||
| 60.2 (6000D-2) | `glm53-nvfp4` 实验容器(09-08 晚 TP1PP8 phase,run_phase.sh 实验链进行中) | **他人实验进行中,勿动**(方案 F PD 链已拆除;GPU7 曾有外部裸金属任务)。动卡前仍先核实归属 | `profiles/pro6000/glm53_nvfp4_pro6000_pd_{prefill→decode 侧}.env` |
|
||||
| 60.3 | 无容器,但 8 卡被外部裸金属训练占用(`/data/mas/larm`,09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — |
|
||||
| 60.4 | `dsv4_scan` 容器(Up 5min)+ 裸金属 sglang `DeepSeek-V4-Flash-0731` DSPARK DP2/TP4 :30000 | **并行会话/外部在役,勿动** | — |
|
||||
| 60.5 | `glm53-nvfp4`(Up 29h) | **NVFP4 团队生产**(deploy_glm53_605.sh,md5 fcd9109b)。生产机铁律:不实验、不重启、不覆盖脚本 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(方案 A 口径) |
|
||||
| 60.6 | 空(无容器) | — | — |
|
||||
| 60.7 | 空 | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — |
|
||||
| 60.8 | 空 | 09-08 已停服清空、router 拆除,不再恢复 | — |
|
||||
| 60.6 | 无容器,但 8 卡被外部裸金属实验占用(`/data/hzy/sparse-opd-*`,09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — |
|
||||
| 60.7 | 基本空(4 卡仍有 `/home/user/dirA_exp` 外部小任务,09-08 晚实测) | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — |
|
||||
| 60.8 | `glm53-nvfp4`(Up,:30000,restart=unless-stopped) | **A16 口径在役**(09-08 i8k/o1k/c16 压测收官保留:tp8eagle.env 场景二变体 = A + MRR32 + graph bs 4/8/12/16,池 276,864,质量门 7/7)。同场景实测最优为 B(TP4PP2,out 265 vs A16 219 tok/s),如转纯吞吐用途可切 B。压测前经授权清理了 GPU3/6 的 dirA_ext 外部 eval 进程 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(+注释中场景二高并发变体);压测记录 `experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/` |
|
||||
|
||||
## 方案 A-F 一览(GLM-5.3-NVFP4 @ pro6000,2026-09-08 双场景报告口径)
|
||||
|
||||
|
||||
@ -0,0 +1,83 @@
|
||||
# GLM-5.3-NVFP4 i8k/o1k/c16 三方案部署对比压测报告
|
||||
|
||||
- 日期:2026-09-08 18:00–19:02(单机串行,174.1.60.8 / 6000D-8,8×H20)
|
||||
- 场景:**input=8192 / output=1024 / concurrency=16**(全唯一输入,PG19 真实语料)
|
||||
- 方案:A(生产口径 TP8+EAGLE)、B(TP4PP2 无投机)、A16(A 的 c16 调优:graph 覆盖 bs16)
|
||||
- 协议:rev20 团队标准(bench_corpus.py,input_ids 直打 `/generate`,temp=0、ignore_eos、stream,nreq=32=2 波,warm nreq16 弃用,每轮全新 run-id + 全新语料窗口 `--pool-override`)
|
||||
- 镜像:`lmsysorg/sglang:nightly-dev-20260828-daf63171`(digest 28e0d2607316),模型 `/data/hf_models/GLM-5.3-NVFP4`
|
||||
|
||||
## 一、判决(先行)
|
||||
|
||||
| 指标 | A 生产口径 | A16(graph 修复) | B(TP4PP2) | 最优 |
|
||||
|---|---|---|---|---|
|
||||
| **输出吞吐 tok/s** | 115.0 | 218.9 | **264.9** | B(A 的 2.30×) |
|
||||
| **输入吞吐 tok/s** | 919.9 | 1,750.8 | **2,119.4** | B(A 的 2.30×) |
|
||||
| **TTFT p50 s** | **3.8** | 4.0 | 10.5 | A(mean 三者打平 ~10s) |
|
||||
| **TPOT p50 ms** | 123.1 | 55.0 | **50.3** | B |
|
||||
| **e2e P50 s** | 139.6 | 66.2 | **61.8** | B(A 的 44%) |
|
||||
| 单流 decode p50 tok/s | 8.4 | 18.9 | **20.1** | B |
|
||||
| EAGLE accept | 2.52 | 2.50 | —(无投机) | — |
|
||||
| 回退 / 命中 | 0 / 0.0 | 0 / 0.0 | 0 / 0.0 | 全过 |
|
||||
| 质量门 | 7/7 | 7/7 | 6/7* | — |
|
||||
|
||||
\* B 的 tool-call 项失败 = 脚本无 `--tool-call-parser` 的已知配置缺口(模型文本明确想调工具,非质量问题;上生产须补 parser,见双场景报告既有结论)。
|
||||
|
||||
**结论:i8k/o1k/c16 场景 B(TP4PP2)全面最优**——吞吐与 e2e 均为 A 的 2.3 倍量级,且 32 请求 e2e 几乎零离散(61.3–62.4s)。A16 把 A 与 B 的差距从 2.30× 收窄到 1.21×,是 EAGLE 家族在本场景的正确口径。**60.8 按既定安排保留最后部署的 A16 在役**(:30000,质量门 7/7、parser 完整);若该机器要转纯吞吐用途,一条命令可切 B(补 parser flags)。
|
||||
|
||||
## 二、主表(逐轮原始数据)
|
||||
|
||||
列:wall | out tok/s | in tok/s | TTFT mean/p50/max s | TPOT mean/p50 ms | e2e mean/p50/max s | 单流 decode p50 tok/s | accept | 命中 | 回退
|
||||
|
||||
**方案 A:TP8 + EAGLE 4/1/5 @ mem0.90,MRR16,chunk8192,fp8KV+hicache3,graph bs≤8,KV 池 276,864**(deploy_glm53_605.sh 原样,md5 fcd9109b,与 60.5 生产同源)
|
||||
|
||||
| 轮 | wall | out | in | TTFT mean/p50/max | TPOT mean/p50 | e2e mean/p50/max | 单流dec | accept | 命中 | 回退 |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| r1 | 277.4 | 118.1 | 945.0 | 10.29 / 4.37 / 31.18 | 119.6 / 118.4 | 132.6 / 136.2 / 190.1 | 8.8 | 2.589 | 0.0 | 0 |
|
||||
| r2 | 293.0 | 111.8 | 894.8 | 10.02 / 3.22 / 30.94 | 126.9 / 127.8 | 139.8 / 142.9 / 193.1 | 7.9 | 2.445 | 0.0 | 0 |
|
||||
|
||||
**方案 A16:A 基础上 EXTRA=`--max-running-requests 32 --cuda-graph-max-bs-decode 16 --cuda-graph-bs-decode 4 8 12 16`**(= tp8eagle.env 注释中的场景二高并发变体;KV 池不变 276,864)
|
||||
|
||||
| 轮 | wall | out | in | TTFT mean/p50/max | TPOT mean/p50 | e2e mean/p50/max | 单流dec | accept | 命中 | 回退 |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| r1 | 151.3 | 216.6 | 1,732.4 | 10.06 / 4.02 / 30.63 | 60.6 / 57.7 | 72.1 / 70.0 / 108.1 | 18.6 | 2.367 | 0.0 | 0 |
|
||||
| r2 | 148.2 | 221.2 | 1,769.2 | 9.88 / 3.98 / 30.67 | 57.3 / 52.2 | 68.5 / 62.3 / 109.6 | 19.2 | 2.634 | 0.0 | 0 |
|
||||
|
||||
**方案 B:TP4 PP2 @ mem0.85,MRR48,chunk16384,radix-off,disable-custom-AR,index_topk_freq=4,KV 池 569,600**(deploy_glm53_optimal.sh,与 09-07/09-08 实跑版同源)
|
||||
|
||||
| 轮 | wall | out | in | TTFT mean/p50/max | TPOT mean/p50 | e2e mean/p50/max | 单流dec | accept | 命中 | 回退 |
|
||||
|---|---|---|---|---|---|---|---|---|---|---|
|
||||
| r1 | 124.2 | 263.8 | 2,110.4 | 10.36 / 10.54 / 17.36 | 50.5 / 50.6 | 62.0 / 62.0 / 62.4 | 20.0 | — | 0.0 | 0 |
|
||||
| r2 | 123.2 | 266.1 | 2,128.4 | 10.38 / 10.54 / 17.36 | 50.0 / 50.0 | 61.5 / 61.5 / 61.8 | 20.1 | — | 0.0 | 0 |
|
||||
|
||||
两轮间一致性:A ±5%、A16 ±9%(TPOT,在跨窗口噪声带内)、B ±1%。
|
||||
|
||||
## 三、关键发现
|
||||
|
||||
1. **A 生产口径在 cc16 有 decode 陷阱**:decode 批 bs 9–16 落在 cuda graph 之外(graph max bs=8),EAGLE 掉图后 TPOT 123ms,比 B 的无投机裸 decode(50ms)还慢 2.4×。历史"EAGLE 大胜"结论(cc1–4)不可外推到 cc16+:掉图模式下投机收益完全反转。
|
||||
2. **A16 的 graph 修复立竿见影**:TPOT 123→55ms(−55%)、out 115→219(+90%)、e2e P50 139.6→66.2s(−53%)。这是生产配置在该场景下零成本可得的改进(EXTRA 三个参数)。
|
||||
3. **EAGLE 在 cc16 高并发下优势摊薄至零**:A16 accept 2.50,但每步墙钟 ≈ 55ms×2.5 ≈ 137ms(verify 批 16×5=80 token/步打满算力),折算单 token 55ms ≈ B 裸 decode 50ms。投机毛利被 verify 批次吃掉。
|
||||
4. **prefill 是 A 家族剩余短板**:B(chunk16384/TP4PP2)输入吞吐 2,119 tok/s,A16 1,751、A 920。B 的优势与既有"prefill 密集 cc≥16 TP4PP2 最优"结论一致;本场景 decode 占比翻倍(o1k vs 历史 o512)后 B 仍全胜。
|
||||
5. **TTFT 的 p50/mean 悖论**:TTFT mean 三方案打平(~10s),但 p50 上 A/A16(~4s)优于 B(10.5s)。机理:nreq32 两波客户端下,A/A16 的 e2e 离散大→尾部请求陆续释放槽位,第二波"到站即过";B 全体同步完成,第二波整体重排 16 条 prefill。**对 TTFT-p50 敏感的小规模交互负载,A/A16 的 chunk8192 体验更好;对稳态吞吐,B 完胜。**
|
||||
6. e2e 离散度:B 最优(p50−min/max 范围仅 0.5s),A16 中等(35–110s),A 最差(79–193s)。
|
||||
|
||||
## 四、收尾状态(2026-09-08 19:02 核验)
|
||||
|
||||
- **60.8 在役**:容器 `glm53-nvfp4` = A16 口径(docker Args 双值并存、生效值 MRR32/graph bs 4 8 12 16,启动日志 `max_running_requests=32`、draft graph `bs=[4,8,12,16]` 实证),:30000,health 200,82.5GB×8,池 276,864。
|
||||
- A、B 已拆除,显存归零后才部署下一方案(每步 <2000MiB 纪律)。
|
||||
- 60.1 生产 glm53-pp4、60.5 团队生产全程未动。
|
||||
- **执行前经授权清理了 60.8 GPU3/6 的外部进程**:PID 2421844/2421852(user 账号 `/home/user/dirA_exp` 的 `eval_freeze_acc.py --mode remove --step 1/4`,14:49 起,各占 7.6GB、util 35%),SIGTERM 退出,显存归零。同源任务在 60.7 尚有 4 卡占用,未动。
|
||||
|
||||
## 五、复现与资产
|
||||
|
||||
- 压测命令(60.8 本机):
|
||||
`python3 /root/bench_corpus.py --corpus /root/corpus_ids.json --input-len 8192 --output-len 1024 --concurrency 16 --num-requests 32 --run-id <新> --pool-override <新窗口> --shared-frac 0 --container glm53-nvfp4`
|
||||
- 轮次编排:`/root/bench_8k.sh <TAG> <POOL_BASE>`(warm nreq16 + r1/r2 nreq32,步进 300k)
|
||||
- 部署:A=`bash /root/deploy_glm53_605.sh`;A16=`EXTRA="--max-running-requests 32 --cuda-graph-max-bs-decode 16 --cuda-graph-bs-decode 4 8 12 16" bash /root/deploy_glm53_605.sh`;B=`bash /root/deploy_glm53_optimal.sh`
|
||||
- **语料窗口台账(corpus_ids.json 共 21,296,780 token)**:本轮消耗 A [17.00M,17.86M)、B [17.90M,18.76M)、A16 [18.80M,19.66M),**已消费至 ~19.66M,剩余 ~1.63M**——后续 8k 复测须从 ≥19.7M 取新窗口或重建语料。
|
||||
- 指标口径:输入吞吐=8192×ok/墙钟;输出吞吐=服务端 completion_tokens 合计/墙钟(EAGLE 下客户端计数低估 2.5–3.5×,一律服务端口径);TTFT 含排队;TPOT=(首响应→末响应)/(completion−1);P50=单请求 e2e 延迟中位数;命中由容器 TP0 Prefill 日志核算(全 0,全唯一验证通过)。
|
||||
- 日志归档:60.8 `/root/bench_logs/`(含 3×gate、9 轮 bench、runner/deploy 日志)+ 本地 tar `D:\sskj\_xfer_608\i8k_bench_20260908.tar.gz`;入库 sskj-review `experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/`。
|
||||
|
||||
## 六、与既有结论的关系
|
||||
|
||||
- 与《双场景压测报告》场景二(16k/512)一致性:B 全胜的结论在 o1k(decode 占比翻倍)下依然成立且差距未缩小——说明 B 的优势主要来自 prefill 侧与池/调度,而非 decode 占比。
|
||||
- 新增修正:**"A 生产口径"在 cc16 的历史场景二数据实际是 A16 变体口径**(tp8eagle.env 注释的场景二变体);本报告首次把 A 原口径与 A16 在同一场景下分离量化——两者差距高达 1.9×,引用历史数据时须区分口径。
|
||||
31
experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/README.md
Normal file
31
experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/README.md
Normal file
@ -0,0 +1,31 @@
|
||||
# GLM-5.3-NVFP4 i8k/o1k/c16 三方案部署对比压测(2026-09-08)
|
||||
|
||||
- 机器:174.1.60.8(6000D-8,8×H20,单机串行三方案轮换)
|
||||
- 场景:input=8192 / output=1024 / cc=16 / nreq=32(2 波),全唯一 PG19 真实语料,temp=0、ignore_eos、stream
|
||||
- 完整报告:本目录 `GLM53_NVFP4_i8k_o1k_c16_压测报告_2026-09-08.md`(判决、主表、机理分析、复现命令)
|
||||
|
||||
## 方案与结果(r1/r2 均值)
|
||||
|
||||
| 方案 | 配置要点 | out tok/s | in tok/s | TTFT p50 | TPOT p50 | e2e P50 | 质量门 |
|
||||
|---|---|---|---|---|---|---|---|
|
||||
| A 生产口径 | TP8+EAGLE 4/1/5,MRR16,chunk8192,graph bs≤8,池 276,864 | 115.0 | 920 | 3.8s | 123ms | 139.6s | 7/7 |
|
||||
| A16 | A + `--max-running-requests 32 --cuda-graph-max-bs-decode 16 --cuda-graph-bs-decode 4 8 12 16` | 218.9 | 1,751 | 4.0s | 55ms | 66.2s | 7/7 |
|
||||
| B | TP4PP2,MRR48,chunk16384,radix-off,池 569,600 | **264.9** | **2,119** | 10.5s | **50ms** | **61.8s** | 6/7* |
|
||||
|
||||
\* B 的 tool-call 失败 = 脚本无 `--tool-call-parser`(已知配置缺口,非质量问题)。
|
||||
|
||||
**判决**:B 全面最优(out 为 A 的 2.30×);A16 把差距收窄到 1.21×,是 EAGLE 家族在本场景的正确口径。
|
||||
60.8 收官保留 **A16 在役**(:30000,restart=unless-stopped)。
|
||||
|
||||
## 核心结论(引用时注意)
|
||||
|
||||
1. **A 生产口径在 cc16 有 decode 陷阱**:bs 9–16 掉出 cuda graph,EAGLE 掉图后 TPOT 123ms,比 B 裸 decode(50ms)慢 2.4×——"EAGLE 大胜"仅成立于 cc≤4 或 graph 覆盖到位时。
|
||||
2. **EAGLE 在 cc16 优势摊薄至零**:A16 accept 2.50,但 verify 批 80 token/步打满算力,折算单 token 55ms ≈ B 的 50ms。
|
||||
3. **历史场景二(16k/512)的"A"行实为 A16 变体口径**,与 A 原口径相差 1.9×,引用历史数据须区分。
|
||||
|
||||
## 资产
|
||||
|
||||
- `scripts/`:bench_8k.sh(轮次编排)、bench_corpus.py(c6d92126,60.7 同源)、deploy_glm53_605.sh(fcd9109b,8 机标准)、deploy_glm53_optimal.sh(22685f56)
|
||||
- `results/bench_logs/`:3×gate(A 7/7、B 6/7、A16 7/7)、9 轮 bench 日志(A/A16/B × warm/r1/r2)、runner/deploy 日志
|
||||
- 语料窗口:本轮消费 corpus_ids.json 至 **~19.66M / 21.30M**(A [17.0,17.86M)、B [17.9,18.76M)、A16 [18.8,19.66M)),剩余 ~1.63M,复测须取 ≥19.7M 新窗口
|
||||
- 压测机执行前经授权清理了 60.8 GPU3/6 的 dirA_ext 外部 eval 进程(PID 2421844/2421852,SIGTERM)
|
||||
@ -0,0 +1,57 @@
|
||||
{
|
||||
"concurrency": 16,
|
||||
"num_requests": 32,
|
||||
"run_id": 9411,
|
||||
"corpus_window": {
|
||||
"start": 19100000,
|
||||
"end": 19362144
|
||||
},
|
||||
"ok": 32,
|
||||
"failed": 0,
|
||||
"wall_s": 151.31,
|
||||
"input_len": 8192,
|
||||
"shared_len": 0,
|
||||
"unique_len": 8192,
|
||||
"output_len": 1024,
|
||||
"output_tokens_total": 32768,
|
||||
"output_throughput_tok_s": 216.55,
|
||||
"input_throughput_tok_s": 1732.44,
|
||||
"ttft_s": {
|
||||
"mean": 10.0558,
|
||||
"p50": 4.0156,
|
||||
"max": 30.6315,
|
||||
"min": 2.0783
|
||||
},
|
||||
"tpot_s": {
|
||||
"mean": 0.0606,
|
||||
"p50": 0.0577,
|
||||
"max": 0.1026,
|
||||
"min": 0.039
|
||||
},
|
||||
"e2e_s": {
|
||||
"mean": 72.0651,
|
||||
"p50": 69.9843,
|
||||
"max": 108.1261,
|
||||
"min": 41.9544
|
||||
},
|
||||
"per_req_out_tok_s_e2e": {
|
||||
"mean": 15.3242,
|
||||
"p50": 14.6458,
|
||||
"max": 24.4074,
|
||||
"min": 9.4704
|
||||
},
|
||||
"per_req_decode_tok_s": {
|
||||
"mean": 17.6898,
|
||||
"p50": 18.553,
|
||||
"max": 25.6827,
|
||||
"min": 9.7574
|
||||
},
|
||||
"spec_accept_length_mean": 2.367,
|
||||
"retractions_total": 0,
|
||||
"cache_hit_from_logs": {
|
||||
"prefill_batches": 32,
|
||||
"new_tokens": 262144,
|
||||
"cached_tokens": 0,
|
||||
"hit_rate": 0.0
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,57 @@
|
||||
{
|
||||
"concurrency": 16,
|
||||
"num_requests": 32,
|
||||
"run_id": 9412,
|
||||
"corpus_window": {
|
||||
"start": 19400000,
|
||||
"end": 19662144
|
||||
},
|
||||
"ok": 32,
|
||||
"failed": 0,
|
||||
"wall_s": 148.17,
|
||||
"input_len": 8192,
|
||||
"shared_len": 0,
|
||||
"unique_len": 8192,
|
||||
"output_len": 1024,
|
||||
"output_tokens_total": 32768,
|
||||
"output_throughput_tok_s": 221.15,
|
||||
"input_throughput_tok_s": 1769.19,
|
||||
"ttft_s": {
|
||||
"mean": 9.875,
|
||||
"p50": 3.9786,
|
||||
"max": 30.6695,
|
||||
"min": 2.0794
|
||||
},
|
||||
"tpot_s": {
|
||||
"mean": 0.0573,
|
||||
"p50": 0.0522,
|
||||
"max": 0.102,
|
||||
"min": 0.0307
|
||||
},
|
||||
"e2e_s": {
|
||||
"mean": 68.4525,
|
||||
"p50": 62.3499,
|
||||
"max": 109.5734,
|
||||
"min": 35.5019
|
||||
},
|
||||
"per_req_out_tok_s_e2e": {
|
||||
"mean": 16.5606,
|
||||
"p50": 16.6965,
|
||||
"max": 28.8435,
|
||||
"min": 9.3453
|
||||
},
|
||||
"per_req_decode_tok_s": {
|
||||
"mean": 19.2407,
|
||||
"p50": 19.179,
|
||||
"max": 32.6448,
|
||||
"min": 9.8182
|
||||
},
|
||||
"spec_accept_length_mean": 2.634,
|
||||
"retractions_total": 0,
|
||||
"cache_hit_from_logs": {
|
||||
"prefill_batches": 32,
|
||||
"new_tokens": 262144,
|
||||
"cached_tokens": 0,
|
||||
"hit_rate": 0.0
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,57 @@
|
||||
{
|
||||
"concurrency": 16,
|
||||
"num_requests": 16,
|
||||
"run_id": 9401,
|
||||
"corpus_window": {
|
||||
"start": 18800000,
|
||||
"end": 18931072
|
||||
},
|
||||
"ok": 16,
|
||||
"failed": 0,
|
||||
"wall_s": 90.65,
|
||||
"input_len": 8192,
|
||||
"shared_len": 0,
|
||||
"unique_len": 8192,
|
||||
"output_len": 1024,
|
||||
"output_tokens_total": 16384,
|
||||
"output_throughput_tok_s": 180.74,
|
||||
"input_throughput_tok_s": 1445.94,
|
||||
"ttft_s": {
|
||||
"mean": 29.6679,
|
||||
"p50": 30.755,
|
||||
"max": 44.1152,
|
||||
"min": 15.0222
|
||||
},
|
||||
"tpot_s": {
|
||||
"mean": 0.0522,
|
||||
"p50": 0.0521,
|
||||
"max": 0.0739,
|
||||
"min": 0.0306
|
||||
},
|
||||
"e2e_s": {
|
||||
"mean": 83.0523,
|
||||
"p50": 87.6337,
|
||||
"max": 90.6464,
|
||||
"min": 69.7624
|
||||
},
|
||||
"per_req_out_tok_s_e2e": {
|
||||
"mean": 12.4314,
|
||||
"p50": 11.9193,
|
||||
"max": 14.6784,
|
||||
"min": 11.2966
|
||||
},
|
||||
"per_req_decode_tok_s": {
|
||||
"mean": 20.6139,
|
||||
"p50": 20.1971,
|
||||
"max": 32.678,
|
||||
"min": 13.5407
|
||||
},
|
||||
"spec_accept_length_mean": 2.643,
|
||||
"retractions_total": 0,
|
||||
"cache_hit_from_logs": {
|
||||
"prefill_batches": 16,
|
||||
"new_tokens": 131072,
|
||||
"cached_tokens": 0,
|
||||
"hit_rate": 0.0
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,57 @@
|
||||
{
|
||||
"concurrency": 16,
|
||||
"num_requests": 32,
|
||||
"run_id": 9411,
|
||||
"corpus_window": {
|
||||
"start": 17300000,
|
||||
"end": 17562144
|
||||
},
|
||||
"ok": 32,
|
||||
"failed": 0,
|
||||
"wall_s": 277.41,
|
||||
"input_len": 8192,
|
||||
"shared_len": 0,
|
||||
"unique_len": 8192,
|
||||
"output_len": 1024,
|
||||
"output_tokens_total": 32768,
|
||||
"output_throughput_tok_s": 118.12,
|
||||
"input_throughput_tok_s": 944.97,
|
||||
"ttft_s": {
|
||||
"mean": 10.2921,
|
||||
"p50": 4.3675,
|
||||
"max": 31.1779,
|
||||
"min": 2.3895
|
||||
},
|
||||
"tpot_s": {
|
||||
"mean": 0.1196,
|
||||
"p50": 0.1184,
|
||||
"max": 0.1667,
|
||||
"min": 0.075
|
||||
},
|
||||
"e2e_s": {
|
||||
"mean": 132.6294,
|
||||
"p50": 136.2478,
|
||||
"max": 190.146,
|
||||
"min": 79.1449
|
||||
},
|
||||
"per_req_out_tok_s_e2e": {
|
||||
"mean": 8.1891,
|
||||
"p50": 7.536,
|
||||
"max": 12.9383,
|
||||
"min": 5.3853
|
||||
},
|
||||
"per_req_decode_tok_s": {
|
||||
"mean": 8.7968,
|
||||
"p50": 8.7859,
|
||||
"max": 13.3427,
|
||||
"min": 6.003
|
||||
},
|
||||
"spec_accept_length_mean": 2.589,
|
||||
"retractions_total": 0,
|
||||
"cache_hit_from_logs": {
|
||||
"prefill_batches": 32,
|
||||
"new_tokens": 262144,
|
||||
"cached_tokens": 0,
|
||||
"hit_rate": 0.0
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,57 @@
|
||||
{
|
||||
"concurrency": 16,
|
||||
"num_requests": 32,
|
||||
"run_id": 9412,
|
||||
"corpus_window": {
|
||||
"start": 17600000,
|
||||
"end": 17862144
|
||||
},
|
||||
"ok": 32,
|
||||
"failed": 0,
|
||||
"wall_s": 292.98,
|
||||
"input_len": 8192,
|
||||
"shared_len": 0,
|
||||
"unique_len": 8192,
|
||||
"output_len": 1024,
|
||||
"output_tokens_total": 32768,
|
||||
"output_throughput_tok_s": 111.84,
|
||||
"input_throughput_tok_s": 894.75,
|
||||
"ttft_s": {
|
||||
"mean": 10.0189,
|
||||
"p50": 3.2165,
|
||||
"max": 30.9377,
|
||||
"min": 2.4035
|
||||
},
|
||||
"tpot_s": {
|
||||
"mean": 0.1269,
|
||||
"p50": 0.1278,
|
||||
"max": 0.1762,
|
||||
"min": 0.077
|
||||
},
|
||||
"e2e_s": {
|
||||
"mean": 139.8173,
|
||||
"p50": 142.9177,
|
||||
"max": 193.13,
|
||||
"min": 81.2596
|
||||
},
|
||||
"per_req_out_tok_s_e2e": {
|
||||
"mean": 7.6338,
|
||||
"p50": 7.6204,
|
||||
"max": 12.6016,
|
||||
"min": 5.3021
|
||||
},
|
||||
"per_req_decode_tok_s": {
|
||||
"mean": 8.1582,
|
||||
"p50": 7.9203,
|
||||
"max": 13.0008,
|
||||
"min": 5.6814
|
||||
},
|
||||
"spec_accept_length_mean": 2.445,
|
||||
"retractions_total": 0,
|
||||
"cache_hit_from_logs": {
|
||||
"prefill_batches": 32,
|
||||
"new_tokens": 262144,
|
||||
"cached_tokens": 0,
|
||||
"hit_rate": 0.0
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,57 @@
|
||||
{
|
||||
"concurrency": 16,
|
||||
"num_requests": 16,
|
||||
"run_id": 9401,
|
||||
"corpus_window": {
|
||||
"start": 17000000,
|
||||
"end": 17131072
|
||||
},
|
||||
"ok": 16,
|
||||
"failed": 0,
|
||||
"wall_s": 172.19,
|
||||
"input_len": 8192,
|
||||
"shared_len": 0,
|
||||
"unique_len": 8192,
|
||||
"output_len": 1024,
|
||||
"output_tokens_total": 16384,
|
||||
"output_throughput_tok_s": 95.15,
|
||||
"input_throughput_tok_s": 761.21,
|
||||
"ttft_s": {
|
||||
"mean": 29.1555,
|
||||
"p50": 30.0011,
|
||||
"max": 45.0124,
|
||||
"min": 14.441
|
||||
},
|
||||
"tpot_s": {
|
||||
"mean": 0.1241,
|
||||
"p50": 0.1326,
|
||||
"max": 0.1482,
|
||||
"min": 0.0785
|
||||
},
|
||||
"e2e_s": {
|
||||
"mean": 156.117,
|
||||
"p50": 167.158,
|
||||
"max": 172.1581,
|
||||
"min": 114.2218
|
||||
},
|
||||
"per_req_out_tok_s_e2e": {
|
||||
"mean": 6.7017,
|
||||
"p50": 6.1651,
|
||||
"max": 8.965,
|
||||
"min": 5.948
|
||||
},
|
||||
"per_req_decode_tok_s": {
|
||||
"mean": 8.3659,
|
||||
"p50": 7.9182,
|
||||
"max": 12.7508,
|
||||
"min": 6.7522
|
||||
},
|
||||
"spec_accept_length_mean": 2.501,
|
||||
"retractions_total": 0,
|
||||
"cache_hit_from_logs": {
|
||||
"prefill_batches": 16,
|
||||
"new_tokens": 131072,
|
||||
"cached_tokens": 0,
|
||||
"hit_rate": 0.0
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,57 @@
|
||||
{
|
||||
"concurrency": 16,
|
||||
"num_requests": 32,
|
||||
"run_id": 9411,
|
||||
"corpus_window": {
|
||||
"start": 18200000,
|
||||
"end": 18462144
|
||||
},
|
||||
"ok": 32,
|
||||
"failed": 0,
|
||||
"wall_s": 124.22,
|
||||
"input_len": 8192,
|
||||
"shared_len": 0,
|
||||
"unique_len": 8192,
|
||||
"output_len": 1024,
|
||||
"output_tokens_total": 32768,
|
||||
"output_throughput_tok_s": 263.8,
|
||||
"input_throughput_tok_s": 2110.38,
|
||||
"ttft_s": {
|
||||
"mean": 10.3634,
|
||||
"p50": 10.5406,
|
||||
"max": 17.3639,
|
||||
"min": 1.777
|
||||
},
|
||||
"tpot_s": {
|
||||
"mean": 0.0505,
|
||||
"p50": 0.0506,
|
||||
"max": 0.0592,
|
||||
"min": 0.0434
|
||||
},
|
||||
"e2e_s": {
|
||||
"mean": 62.019,
|
||||
"p50": 62.0417,
|
||||
"max": 62.366,
|
||||
"min": 61.8
|
||||
},
|
||||
"per_req_out_tok_s_e2e": {
|
||||
"mean": 16.5113,
|
||||
"p50": 16.5486,
|
||||
"max": 16.5696,
|
||||
"min": 16.4192
|
||||
},
|
||||
"per_req_decode_tok_s": {
|
||||
"mean": 19.9851,
|
||||
"p50": 19.9608,
|
||||
"max": 23.039,
|
||||
"min": 16.9011
|
||||
},
|
||||
"spec_accept_length_mean": null,
|
||||
"retractions_total": 0,
|
||||
"cache_hit_from_logs": {
|
||||
"prefill_batches": 36,
|
||||
"new_tokens": 524288,
|
||||
"cached_tokens": 0,
|
||||
"hit_rate": 0.0
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,57 @@
|
||||
{
|
||||
"concurrency": 16,
|
||||
"num_requests": 32,
|
||||
"run_id": 9412,
|
||||
"corpus_window": {
|
||||
"start": 18500000,
|
||||
"end": 18762144
|
||||
},
|
||||
"ok": 32,
|
||||
"failed": 0,
|
||||
"wall_s": 123.16,
|
||||
"input_len": 8192,
|
||||
"shared_len": 0,
|
||||
"unique_len": 8192,
|
||||
"output_len": 1024,
|
||||
"output_tokens_total": 32768,
|
||||
"output_throughput_tok_s": 266.05,
|
||||
"input_throughput_tok_s": 2128.41,
|
||||
"ttft_s": {
|
||||
"mean": 10.3812,
|
||||
"p50": 10.5386,
|
||||
"max": 17.3626,
|
||||
"min": 1.7756
|
||||
},
|
||||
"tpot_s": {
|
||||
"mean": 0.05,
|
||||
"p50": 0.05,
|
||||
"max": 0.0587,
|
||||
"min": 0.043
|
||||
},
|
||||
"e2e_s": {
|
||||
"mean": 61.5036,
|
||||
"p50": 61.4987,
|
||||
"max": 61.8083,
|
||||
"min": 61.3264
|
||||
},
|
||||
"per_req_out_tok_s_e2e": {
|
||||
"mean": 16.6495,
|
||||
"p50": 16.6689,
|
||||
"max": 16.6975,
|
||||
"min": 16.5674
|
||||
},
|
||||
"per_req_decode_tok_s": {
|
||||
"mean": 20.1968,
|
||||
"p50": 20.1428,
|
||||
"max": 23.2803,
|
||||
"min": 17.0575
|
||||
},
|
||||
"spec_accept_length_mean": null,
|
||||
"retractions_total": 0,
|
||||
"cache_hit_from_logs": {
|
||||
"prefill_batches": 36,
|
||||
"new_tokens": 524288,
|
||||
"cached_tokens": 0,
|
||||
"hit_rate": 0.0
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,57 @@
|
||||
{
|
||||
"concurrency": 16,
|
||||
"num_requests": 16,
|
||||
"run_id": 9401,
|
||||
"corpus_window": {
|
||||
"start": 17900000,
|
||||
"end": 18031072
|
||||
},
|
||||
"ok": 16,
|
||||
"failed": 0,
|
||||
"wall_s": 65.46,
|
||||
"input_len": 8192,
|
||||
"shared_len": 0,
|
||||
"unique_len": 8192,
|
||||
"output_len": 1024,
|
||||
"output_tokens_total": 16384,
|
||||
"output_throughput_tok_s": 250.29,
|
||||
"input_throughput_tok_s": 2002.32,
|
||||
"ttft_s": {
|
||||
"mean": 13.2299,
|
||||
"p50": 13.271,
|
||||
"max": 20.2234,
|
||||
"min": 4.2816
|
||||
},
|
||||
"tpot_s": {
|
||||
"mean": 0.0509,
|
||||
"p50": 0.0507,
|
||||
"max": 0.0598,
|
||||
"min": 0.0439
|
||||
},
|
||||
"e2e_s": {
|
||||
"mean": 65.2664,
|
||||
"p50": 65.1682,
|
||||
"max": 65.4579,
|
||||
"min": 65.1135
|
||||
},
|
||||
"per_req_out_tok_s_e2e": {
|
||||
"mean": 15.6896,
|
||||
"p50": 15.7139,
|
||||
"max": 15.7264,
|
||||
"min": 15.6436
|
||||
},
|
||||
"per_req_decode_tok_s": {
|
||||
"mean": 19.845,
|
||||
"p50": 19.7319,
|
||||
"max": 22.7938,
|
||||
"min": 16.7391
|
||||
},
|
||||
"spec_accept_length_mean": null,
|
||||
"retractions_total": 0,
|
||||
"cache_hit_from_logs": {
|
||||
"prefill_batches": 18,
|
||||
"new_tokens": 262144,
|
||||
"cached_tokens": 0,
|
||||
"hit_rate": 0.0
|
||||
}
|
||||
}
|
||||
@ -0,0 +1,26 @@
|
||||
#!/bin/bash
|
||||
# bench_8k.sh — i8k/o1k/c16 bench rounds for ONE scheme (run via nohup on 60.8)
|
||||
# usage: bash bench_8k.sh <TAG> <POOL_BASE>
|
||||
# Rounds: warm (nreq16, discard) + r1/r2 (nreq32, measured), fresh corpus window each.
|
||||
set -u
|
||||
TAG=$1
|
||||
POOL=$2
|
||||
mkdir -p /root/bench_logs
|
||||
|
||||
run_round() {
|
||||
local name=$1 nreq=$2 pool=$3 rid=$4
|
||||
echo "=== [$TAG] round $name nreq=$nreq pool=$pool rid=$rid start=$(date -Is) ==="
|
||||
python3 /root/bench_corpus.py --corpus /root/corpus_ids.json \
|
||||
--input-len 8192 --output-len 1024 --concurrency 16 \
|
||||
--num-requests "$nreq" --run-id "$rid" --pool-override "$pool" \
|
||||
--shared-frac 0 --container glm53-nvfp4 \
|
||||
> "/root/bench_logs/${TAG}_${name}.log" 2>&1
|
||||
echo "=== [$TAG] round $name rc=$? end=$(date -Is) ==="
|
||||
grep -c '"ok"' "/root/bench_logs/${TAG}_${name}.log" >/dev/null 2>&1
|
||||
tail -3 "/root/bench_logs/${TAG}_${name}.log"
|
||||
}
|
||||
|
||||
run_round warm 16 "$POOL" 9401
|
||||
run_round r1 32 $((POOL + 300000)) 9411
|
||||
run_round r2 32 $((POOL + 600000)) 9412
|
||||
echo "=== [$TAG] ALL DONE $(date -Is) ==="
|
||||
@ -0,0 +1,298 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Real-corpus (PG19) benchmark for sglang GLM-5.3-NVFP4.
|
||||
|
||||
Same methodology as bench_hit90.py (input_ids direct to /generate, temp 0,
|
||||
ignore_eos, stream, server-side completion_tokens counting, hit-rate verified
|
||||
from scheduler logs), but prompts are token slices of REAL book text tokenized
|
||||
with the served model's own tokenizer, replacing random ids.
|
||||
|
||||
Corpus file: JSON {"ids": [flat token ids], "books": [[start, end), ...]}
|
||||
|
||||
Fixed token-offset layout into the flat id array:
|
||||
[0, 117968) shared prefix for 128k points (90% of 131072)
|
||||
[0, 58976) shared prefix for 64k points (same region, shorter cut)
|
||||
pool A [131072, +4*8*13104) 128k unique suffixes, run-ids 9301-9304
|
||||
pool B [550400, +4*8*6560) 64k unique suffixes, run-ids 9305-9308
|
||||
pool C [760320, 262144+524288+524288) 16k fully-unique prompts, run-ids 9311-9313
|
||||
spare [2071040, end) warmup slices / re-run margin
|
||||
|
||||
Each run-id maps to one non-overlapping window (one window = one bench point);
|
||||
re-running a point with fresh text = bump --pool-override past the spare base.
|
||||
|
||||
Usage:
|
||||
python3 bench_corpus.py --corpus /root/corpus_ids.json --input-len 131072 \
|
||||
--concurrency 1 --num-requests 8 --run-id 9301 --shared-frac 0.9
|
||||
"""
|
||||
import argparse
|
||||
import datetime
|
||||
import json
|
||||
import re
|
||||
import statistics
|
||||
import subprocess
|
||||
import time
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
|
||||
import requests
|
||||
|
||||
OUTPUT_LEN_DEFAULT = 512
|
||||
CORPUS_DEFAULT = "/root/corpus_ids.json"
|
||||
|
||||
# fixed pool layout (see docstring)
|
||||
S1_128K_RID0, S1_64K_RID0, S2_RID0 = 9301, 9305, 9311
|
||||
POOL_A_BASE, POOL_A_PER = 131072, 8 * 13104 # 128k suffix windows
|
||||
POOL_B_BASE = POOL_A_BASE + 4 * POOL_A_PER # 550400
|
||||
POOL_B_PER = 8 * 6560 # 64k suffix windows
|
||||
POOL_C_BASE = POOL_B_BASE + 4 * POOL_B_PER # 760320
|
||||
POOL_C_SIZES = [16 * 16384, 32 * 16384, 32 * 16384] # cc8 / cc16 / cc32
|
||||
SPARE_BASE = POOL_C_BASE + sum(POOL_C_SIZES) # 2071040
|
||||
|
||||
sess = requests.Session()
|
||||
sess.trust_env = False # bypass any proxy env on the host
|
||||
|
||||
|
||||
def split_lens(input_len, shared_frac):
|
||||
# unique suffix = (1 - shared_frac) of the prompt, page-16 aligned
|
||||
# (128k @0.9 -> 13104 unique; 16k @0.0 -> fully unique prompts)
|
||||
unique = round(input_len * (1.0 - shared_frac) / 16) * 16
|
||||
return input_len - unique, unique
|
||||
|
||||
|
||||
def pool_start_for(input_len, shared_frac, run_id, override):
|
||||
if override is not None:
|
||||
return override
|
||||
if shared_frac > 0:
|
||||
if input_len == 131072:
|
||||
idx = run_id - S1_128K_RID0
|
||||
if not 0 <= idx < 4:
|
||||
sys_exit_bad_runid(run_id, "128k points use run-ids 9301-9304")
|
||||
return POOL_A_BASE + idx * POOL_A_PER
|
||||
if input_len == 65536:
|
||||
idx = run_id - S1_64K_RID0
|
||||
if not 0 <= idx < 4:
|
||||
sys_exit_bad_runid(run_id, "64k points use run-ids 9305-9308")
|
||||
return POOL_B_BASE + idx * POOL_B_PER
|
||||
sys_exit_bad_runid(run_id, "shared-frac>0 supports 131072/65536 only")
|
||||
idx = run_id - S2_RID0
|
||||
if not 0 <= idx < 3:
|
||||
sys_exit_bad_runid(run_id, "16k unique points use run-ids 9311-9313")
|
||||
return POOL_C_BASE + sum(POOL_C_SIZES[:idx])
|
||||
|
||||
|
||||
def sys_exit_bad_runid(run_id, msg):
|
||||
raise SystemExit(f"[pool] run-id {run_id} outside expected set: {msg}")
|
||||
|
||||
|
||||
def build_prompts(ids, shared_len, unique_len, num_requests, pool_start):
|
||||
if shared_len:
|
||||
shared = ids[0:shared_len]
|
||||
else:
|
||||
shared = []
|
||||
end = pool_start + num_requests * unique_len
|
||||
if end > len(ids):
|
||||
raise SystemExit(
|
||||
f"[pool] window [{pool_start}, {end}) exceeds corpus ({len(ids)} ids); "
|
||||
f"use --pool-override or a larger corpus")
|
||||
prompts = []
|
||||
for i in range(num_requests):
|
||||
s = pool_start + i * unique_len
|
||||
prompts.append(shared + ids[s:s + unique_len])
|
||||
return shared, prompts, (pool_start, end)
|
||||
|
||||
|
||||
def warmup(url, ids, shared_len):
|
||||
# primes the radix cache with the shared prefix (same role as in bench_hit90);
|
||||
# warm slice comes from the spare region so it never collides with a pool window
|
||||
if len(ids) >= SPARE_BASE + 64:
|
||||
warm_slice = ids[SPARE_BASE:SPARE_BASE + 64]
|
||||
else:
|
||||
warm_slice = ids[-64:]
|
||||
payload = {
|
||||
"input_ids": ids[0:shared_len] + warm_slice if shared_len else warm_slice,
|
||||
"sampling_params": {"max_new_tokens": 8, "temperature": 0.0, "ignore_eos": True},
|
||||
}
|
||||
t0 = time.perf_counter()
|
||||
r = sess.post(url, json=payload, timeout=1800)
|
||||
dt = time.perf_counter() - t0
|
||||
print(f"[warmup] http={r.status_code} wall={dt:.2f}s", flush=True)
|
||||
|
||||
|
||||
def bench_one(url, prompt, output_len, idx, results):
|
||||
payload = {
|
||||
"input_ids": prompt,
|
||||
"sampling_params": {"max_new_tokens": output_len, "temperature": 0.0, "ignore_eos": True},
|
||||
"stream": True,
|
||||
}
|
||||
rec = {"idx": idx}
|
||||
t0 = time.perf_counter()
|
||||
first = last = None
|
||||
first_ct = None
|
||||
final_meta = None
|
||||
max_ct = 0
|
||||
try:
|
||||
with sess.post(url, json=payload, stream=True, timeout=3600) as resp:
|
||||
for raw in resp.iter_lines():
|
||||
if not raw or not raw.startswith(b"data:"):
|
||||
continue
|
||||
body = raw[5:].strip()
|
||||
if body == b"[DONE]":
|
||||
continue
|
||||
now = time.perf_counter()
|
||||
try:
|
||||
d = json.loads(body)
|
||||
except Exception:
|
||||
continue
|
||||
mi = d.get("meta_info") or {}
|
||||
ct = mi.get("completion_tokens") or 0
|
||||
if ct:
|
||||
max_ct = max(max_ct, ct)
|
||||
if first is None:
|
||||
first = now
|
||||
first_ct = ct
|
||||
last = now
|
||||
if mi.get("finish_reason"):
|
||||
final_meta = mi
|
||||
t_end = time.perf_counter()
|
||||
n_out = max_ct or ((final_meta or {}).get("completion_tokens") or 0)
|
||||
decode_span = (last - first) if (first and last and last > first) else 0.0
|
||||
rec.update(
|
||||
ok=n_out > 0,
|
||||
ttft=(first - t0) if first else None,
|
||||
e2e=t_end - t0,
|
||||
n_out=n_out,
|
||||
first_chunk_tokens=first_ct,
|
||||
decode_span=decode_span,
|
||||
tpot=(decode_span / (n_out - 1)) if n_out > 1 else None,
|
||||
per_req_decode_tok_s=(n_out / decode_span) if decode_span > 0 else None,
|
||||
retractions=(final_meta or {}).get("num_retractions"),
|
||||
spec_accept_len=(final_meta or {}).get("spec_accept_length"),
|
||||
)
|
||||
except Exception as e:
|
||||
rec.update(ok=False, error=repr(e))
|
||||
results[idx] = rec
|
||||
|
||||
|
||||
def verify_hit_rate(container, t_start, t_end):
|
||||
def rfc3339(epoch):
|
||||
return (datetime.datetime.fromtimestamp(epoch, tz=datetime.timezone.utc)
|
||||
.isoformat().replace("+00:00", "Z"))
|
||||
|
||||
try:
|
||||
# No margin before t_start: warmup's prefill lines end strictly before it,
|
||||
# and catching them would deflate the measured hit rate.
|
||||
p = subprocess.run(
|
||||
["docker", "logs", container, "--since", rfc3339(t_start), "--until", rfc3339(t_end + 2)],
|
||||
capture_output=True, text=True, timeout=120)
|
||||
text = p.stdout + p.stderr
|
||||
except Exception as e:
|
||||
return {"error": repr(e)}
|
||||
pat = re.compile(r"#new-token: (\d+).*?#cached-token: (\d+)")
|
||||
n_batches = new_tok = cached_tok = 0
|
||||
for line in text.splitlines():
|
||||
if "TP0]" not in line or "Prefill batch" not in line:
|
||||
continue
|
||||
m = pat.search(line)
|
||||
if m:
|
||||
n_batches += 1
|
||||
new_tok += int(m.group(1))
|
||||
cached_tok += int(m.group(2))
|
||||
total = new_tok + cached_tok
|
||||
return {
|
||||
"prefill_batches": n_batches,
|
||||
"new_tokens": new_tok,
|
||||
"cached_tokens": cached_tok,
|
||||
"hit_rate": round(cached_tok / total, 4) if total else None,
|
||||
}
|
||||
|
||||
|
||||
def stats(vals):
|
||||
vals = [v for v in vals if v is not None]
|
||||
if not vals:
|
||||
return {"mean": None, "p50": None, "max": None, "min": None}
|
||||
s = sorted(vals)
|
||||
return {
|
||||
"mean": round(statistics.fmean(vals), 4),
|
||||
"p50": round(s[len(s) // 2], 4),
|
||||
"max": round(s[-1], 4),
|
||||
"min": round(s[0], 4),
|
||||
}
|
||||
|
||||
|
||||
def main():
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--concurrency", type=int, required=True)
|
||||
ap.add_argument("--num-requests", type=int, required=True)
|
||||
ap.add_argument("--run-id", type=int, required=True)
|
||||
ap.add_argument("--input-len", type=int, required=True, help="16384 / 65536 / 131072")
|
||||
ap.add_argument("--output-len", type=int, default=OUTPUT_LEN_DEFAULT)
|
||||
ap.add_argument("--shared-frac", type=float, default=0.9)
|
||||
ap.add_argument("--corpus", default=CORPUS_DEFAULT)
|
||||
ap.add_argument("--pool-override", type=int, default=None,
|
||||
help="explicit corpus offset for the unique-suffix window (re-runs)")
|
||||
ap.add_argument("--url", default="http://127.0.0.1:30000/generate")
|
||||
ap.add_argument("--container", default="glm53-nvfp4")
|
||||
args = ap.parse_args()
|
||||
|
||||
with open(args.corpus) as f:
|
||||
corpus = json.load(f)
|
||||
ids = corpus["ids"]
|
||||
|
||||
shared_len, unique_len = split_lens(args.input_len, args.shared_frac)
|
||||
pool_start = pool_start_for(args.input_len, args.shared_frac, args.run_id, args.pool_override)
|
||||
shared, prompts, window = build_prompts(ids, shared_len, unique_len, args.num_requests, pool_start)
|
||||
print(f"[pool] window={window} shared_len={shared_len} unique_len={unique_len} "
|
||||
f"corpus_total={len(ids)}", flush=True)
|
||||
warmup(args.url, ids, shared_len)
|
||||
|
||||
results = {}
|
||||
t_start = time.time()
|
||||
t0 = time.perf_counter()
|
||||
with ThreadPoolExecutor(max_workers=args.concurrency) as ex:
|
||||
futs = [ex.submit(bench_one, args.url, p, args.output_len, i, results)
|
||||
for i, p in enumerate(prompts)]
|
||||
for f in futs:
|
||||
f.result()
|
||||
wall = time.perf_counter() - t0
|
||||
t_end = time.time()
|
||||
|
||||
hit = verify_hit_rate(args.container, t_start, t_end)
|
||||
|
||||
ok = [r for r in results.values() if r.get("ok")]
|
||||
n_out_total = sum(r["n_out"] for r in ok)
|
||||
out_tps = [r["n_out"] / r["e2e"] for r in ok if r.get("e2e")]
|
||||
ttft = stats([r.get("ttft") for r in ok])
|
||||
tpot = stats([r.get("tpot") for r in ok])
|
||||
e2e = stats([r.get("e2e") for r in ok])
|
||||
dec = stats([r.get("per_req_decode_tok_s") for r in ok])
|
||||
spec = [v for v in (r.get("spec_accept_len") for r in ok) if v is not None]
|
||||
retr = sum(r.get("retractions") or 0 for r in ok)
|
||||
|
||||
summary = {
|
||||
"concurrency": args.concurrency,
|
||||
"num_requests": args.num_requests,
|
||||
"run_id": args.run_id,
|
||||
"corpus_window": {"start": window[0], "end": window[1]},
|
||||
"ok": len(ok),
|
||||
"failed": args.num_requests - len(ok),
|
||||
"wall_s": round(wall, 2),
|
||||
"input_len": args.input_len,
|
||||
"shared_len": shared_len,
|
||||
"unique_len": unique_len,
|
||||
"output_len": args.output_len,
|
||||
"output_tokens_total": n_out_total,
|
||||
"output_throughput_tok_s": round(n_out_total / wall, 2) if wall else None,
|
||||
"input_throughput_tok_s": round(args.input_len * len(ok) / wall, 2) if wall else None,
|
||||
"ttft_s": ttft,
|
||||
"tpot_s": tpot,
|
||||
"e2e_s": e2e,
|
||||
"per_req_out_tok_s_e2e": stats(out_tps),
|
||||
"per_req_decode_tok_s": dec,
|
||||
"spec_accept_length_mean": round(statistics.fmean(spec), 3) if spec else None,
|
||||
"retractions_total": retr,
|
||||
"cache_hit_from_logs": hit,
|
||||
}
|
||||
print("\n===== SUMMARY =====")
|
||||
print(json.dumps(summary, indent=2), flush=True)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@ -0,0 +1,58 @@
|
||||
#!/bin/bash
|
||||
# deploy_glm53_605.sh — GLM-5.3-NVFP4 TP8 + EAGLE for 174.1.60.5 (team service)
|
||||
# Target scenario: cc1-2, 64k/128k input 90% cache hit, single request up to 256k,
|
||||
# output-throughput priority. Clean deploy: rm old container, wait VRAM drain, run.
|
||||
#
|
||||
# Env overrides:
|
||||
# MEMFRAC (0.90) STEPS (4) TOPK (1) DRAFT (5) CTXLEN (270336)
|
||||
# CHUNK (8192) MAXPRE (16384) RESTART (no|yes) EXTRA ("")
|
||||
# Usage:
|
||||
# bash deploy_glm53_605.sh # production 4/1/5 @ 0.90
|
||||
# STEPS=5 DRAFT=6 bash deploy_glm53_605.sh # tuning round
|
||||
# CHUNK=16384 bash deploy_glm53_605.sh # prefill tuning round
|
||||
# RESTART=yes bash deploy_glm53_605.sh # production finalize
|
||||
set -e
|
||||
MEMFRAC=${MEMFRAC:-0.90}
|
||||
STEPS=${STEPS:-4}
|
||||
TOPK=${TOPK:-1}
|
||||
DRAFT=${DRAFT:-5}
|
||||
CTXLEN=${CTXLEN:-270336}
|
||||
CHUNK=${CHUNK:-8192}
|
||||
MAXPRE=${MAXPRE:-16384}
|
||||
EXTRA=${EXTRA:-}
|
||||
if [ "$RESTART" = "yes" ]; then RP="--restart unless-stopped"; else RP="--restart no"; fi
|
||||
|
||||
echo "[deploy] removing old container (if any)"
|
||||
# rm -f times out on big GPU containers on this daemon; retry until really gone
|
||||
for i in $(seq 1 45); do
|
||||
CID=$(docker ps -a --filter name=glm53-nvfp4 -q)
|
||||
[ -z "$CID" ] && break
|
||||
docker rm -f glm53-nvfp4 >/dev/null 2>&1 || true
|
||||
sleep 2
|
||||
done
|
||||
|
||||
echo "[deploy] waiting for VRAM drain"
|
||||
for i in $(seq 1 45); do
|
||||
used=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | awk '{s+=$1} END {print s}')
|
||||
[ "$used" -lt 2000 ] && break
|
||||
sleep 2
|
||||
done
|
||||
used=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | awk '{s+=$1} END {print s}')
|
||||
echo "[deploy] VRAM now: ${used} MiB total"
|
||||
|
||||
echo "[deploy] starting: TP8 EAGLE ${STEPS}/${TOPK}/${DRAFT} memfrac=${MEMFRAC} extra='${EXTRA}'"
|
||||
docker run -d --name glm53-nvfp4 --gpus all --shm-size 64g --ipc=host $RP \
|
||||
-p 30000:30000 -v /data/hf_models:/data/hf_models \
|
||||
lmsysorg/sglang:nightly-dev-20260828-daf63171 \
|
||||
python3 -m sglang.launch_server \
|
||||
--model-path /data/hf_models/GLM-5.3-NVFP4 --tp 8 \
|
||||
--mem-fraction-static $MEMFRAC --max-running-requests 16 \
|
||||
--chunked-prefill-size $CHUNK --max-prefill-tokens $MAXPRE \
|
||||
--disable-shared-experts-fusion --moe-runner-backend flashinfer_cutlass --disable-flashinfer-autotune \
|
||||
--speculative-algorithm EAGLE --speculative-num-steps $STEPS --speculative-eagle-topk $TOPK --speculative-num-draft-tokens $DRAFT \
|
||||
--kv-cache-dtype fp8_e4m3 --enable-hierarchical-cache --hicache-ratio 3 \
|
||||
--cuda-graph-max-bs-decode 8 --cuda-graph-bs-decode 1 2 3 4 6 8 --cuda-graph-max-bs-prefill 8 \
|
||||
--context-length $CTXLEN --reasoning-parser glm45 --tool-call-parser glm47 \
|
||||
--host 0.0.0.0 --port 30000 $EXTRA
|
||||
|
||||
echo "[deploy] container started; poll: docker logs -f glm53-nvfp4"
|
||||
@ -0,0 +1,62 @@
|
||||
#!/bin/bash
|
||||
# ============================================================
|
||||
# GLM-5.3 最优部署方案(6000D 8卡,TP4 PP2 + IndexCache freq=4)
|
||||
# 2026-09-07
|
||||
#
|
||||
# - 基线配置:TP4 PP2 + cps16k + mem0.85(131.6 tok/s 吞吐基线)
|
||||
# - IndexCache(index_topk_freq=4):层轴索引复用,省 75% indexer
|
||||
# 16K 场景无损失;128K 长上下文并发 1.35-1.47× 提速
|
||||
# - 禁 radix cache;禁投机解码(PP2 与投机框架不兼容,已实测)
|
||||
#
|
||||
# 用法:bash deploy_glm53_optimal.sh
|
||||
# (60.7 试用版:增加显存排空等待 + 就绪等待加长至 20 分钟)
|
||||
# ============================================================
|
||||
set -uo pipefail
|
||||
|
||||
CONTAINER="glm53-nvfp4"
|
||||
IMAGE="lmsysorg/sglang:nightly-dev-20260828-daf63171"
|
||||
MODEL="/data/hf_models/GLM-5.3-NVFP4"
|
||||
PORT=30000
|
||||
TP=4; PP=2; MEM=0.85; MRR=48; CPS=16384
|
||||
|
||||
docker rm -f ${CONTAINER} 2>/dev/null || true
|
||||
|
||||
# docker rm -f 后显存释放滞后数分钟,不等会把新容器 KV 池压小
|
||||
for i in $(seq 1 60); do
|
||||
used=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | awk '{s+=$1} END {print s}')
|
||||
if [ "$used" -lt 1500 ]; then echo "[deploy] drained: ${used} MiB"; break; fi
|
||||
echo "[deploy] drain wait ${i}: ${used} MiB"
|
||||
sleep 10
|
||||
done
|
||||
|
||||
docker run -d --name ${CONTAINER} --gpus all --shm-size 64g --ipc=host \
|
||||
--restart unless-stopped \
|
||||
-p ${PORT}:${PORT} \
|
||||
-v /data/hf_models:/data/hf_models \
|
||||
${IMAGE} \
|
||||
python3 -m sglang.launch_server \
|
||||
--model-path ${MODEL} \
|
||||
--tp-size ${TP} --pp-size ${PP} \
|
||||
--mem-fraction-static ${MEM} \
|
||||
--max-running-requests ${MRR} \
|
||||
--disable-radix-cache \
|
||||
--disable-shared-experts-fusion \
|
||||
--moe-runner-backend flashinfer_cutlass \
|
||||
--disable-flashinfer-autotune \
|
||||
--disable-custom-all-reduce \
|
||||
--chunked-prefill-size ${CPS} \
|
||||
--host 0.0.0.0 --port ${PORT} \
|
||||
--json-model-override-args '{"index_topk_freq": 4}'
|
||||
|
||||
echo "容器已启动,等待就绪..."
|
||||
for i in $(seq 1 120); do
|
||||
code=$(curl -s -o /dev/null -w '%{http_code}' http://localhost:${PORT}/health 2>/dev/null)
|
||||
if [ "$code" = "200" ]; then
|
||||
echo "READY after ~$((i*10))s"
|
||||
docker ps --filter name=${CONTAINER} --format '{{.Names}} {{.Status}}'
|
||||
echo "override args: $(docker inspect ${CONTAINER} --format '{{.Config.Cmd}}' | grep -o 'index_topk_freq[^,}]*')"
|
||||
exit 0
|
||||
fi
|
||||
sleep 10
|
||||
done
|
||||
echo "TIMEOUT"; exit 1
|
||||
Loading…
x
Reference in New Issue
Block a user