diff --git a/deploy/CURRENT.md b/deploy/CURRENT.md index 7877a00..0065fe7 100644 --- a/deploy/CURRENT.md +++ b/deploy/CURRENT.md @@ -11,10 +11,10 @@ | 60.2 (6000D-2) | `glm53-nvfp4` 实验容器(09-08 晚 TP1PP8 phase,run_phase.sh 实验链进行中) | **他人实验进行中,勿动**(方案 F PD 链已拆除;GPU7 曾有外部裸金属任务)。动卡前仍先核实归属 | `profiles/pro6000/glm53_nvfp4_pro6000_pd_{prefill→decode 侧}.env` | | 60.3 | 无容器,但 8 卡被外部裸金属训练占用(`/data/mas/larm`,09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — | | 60.4 | `glm53-nvfp4`(Up,:30000,restart=unless-stopped,09-09 部署) | **TP2PP4 GLM-5.3-NVFP4 在役**(60.1 方案 D 生产配方复刻:TP2/PP4/mem0.85/MRR48/chunk16384/radix-off/无投机/SM120 三件套/disable-custom-all-reduce/index_topk_freq=4,镜像 nightly-dev-20260828-daf63171;health 200、生成冒烟通过)。09-09 经授权清退外部 vllm 评测流水线(tmux `mas` 的 run_multiseed.sh 链,带 --resume 可续跑)后部署;启动脚本 `/root/deploy_s2_test_604.sh`(基于 exp602/deploy_s2_test.sh,唯一改动 restart=unless-stopped)。旧 `dsv4_scan` 容器仍在盘(Exited) | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp2pp4.env`(同 60.1 D 口径) | -| 60.5 | `glm53-nvfp4`(Up 29h) | **NVFP4 团队生产**(deploy_glm53_605.sh,md5 fcd9109b)。生产机铁律:不实验、不重启、不覆盖脚本 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(方案 A 口径) | +| 60.5 | `glm53-nvfp4`(Up 29h) | **NVFP4 团队生产**(deploy_glm53_605.sh,md5 fcd9109b)。生产机铁律:不实验、不重启、不覆盖脚本。**09-09 已交付容量扩容 v2 脚本(TP2PP4-hicache 优胜配置,池 909,632/c6/单条 909k,与 60.8 现役同款),未执行,择维护窗口跑 `/root/deploy_glm53_605_v2.sh`;回滚=原脚本** | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(方案 A 口径);待切 `glm53_nvfp4_pro6000_sglang_tp2pp4_hicache.env` | | 60.6 | 无容器,但 8 卡被外部裸金属实验占用(`/data/hzy/sparse-opd-*`,09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — | | 60.7 | 基本空(4 卡仍有 `/home/user/dirA_exp` 外部小任务,09-08 晚实测) | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — | -| 60.8 | `glm53-nvfp4`(Up,:30000,restart=unless-stopped) | **r37 PP+MTP verify-graph 口径在役**(09-08 深夜取代 A16:TP4PP2 + EAGLE 3/1/4 + verify CUDA graph + draft eager + de-GLOO + SYNC_MASK=127,9 挂载文件,启动 `bash /root/deploy_ppmtp_r37.sh '--tp 4 --pp-size 2 --disable-overlap-schedule --max-prefill-tokens 16384' mtp 8192 0.88 1 1 0 0`)。质量门 7/7 ×2、conc_test×3、killer×6 零崩溃;killer cc16 双种子 **63.8/64.7s vs A16 105.6/96.0s(-40%)**、TPOT 41.5/39.3ms、TTFT 33s vs 58-61s(近乎减半)。修复根因=pre-planned 早退路径缺 pp_proxy 补拷(详见 `experiments/pro6000/glm53_ppmtp_r37_verify_graph/`)。A16 切回脚本 deploy_glm53_605.sh 留盘;sglang_patch2/eagle_worker_common.py 已升级 bisect 版("1"/"0" 语义兼容) | `experiments/pro6000/glm53_ppmtp_r37_verify_graph/`(patches+scripts+results 全量);前史 `experiments/pro6000/glm53_ppmtp_r36_degloo/` | +| 60.8 | `glm53-nvfp4`(Up,:30000,restart=unless-stopped) | **TP2PP4-nomtp 容量扩容口径在役**(09-09 按用户决定取代 r37:TP2PP4 + radix/hicache 开 + ctx 1048576 + cu13 栈 9 挂载 + de-GLOO,memfrac 0.85/chunk 8192/MRR 16,KV 池 **909,632**,启动 `bash /root/deploy_ppmtp_r37.sh '--tp 2 --pp-size 4 --disable-overlap-schedule --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length 1048576' nomtp 8192 0.85`)。四臂对拍优胜(vs A-mirror/r37/B',i128k/o512 c1-4):c4 输入/输出 **5,907/23.1 tok/s**、TTFT p50 43.4s、并发上限 c6、单条 ~909k(900k 实跑)、512k 单条 TTFT 88.4s、90% 命中 c4 输入 17,625 tok/s;质量门 7/7。r37 的 MTP 优势区间是 i16k/cc≤16,切回 = deploy_ppmtp_r37.sh mtp 模式(详见 `experiments/pro6000/glm53_nvfp4_128k_capacity_topology/`) | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp2pp4_hicache.env`;实验全量 `experiments/pro6000/glm53_nvfp4_128k_capacity_topology/`;前史 r37 `experiments/pro6000/glm53_ppmtp_r37_verify_graph/` | ## 方案 A-F 一览(GLM-5.3-NVFP4 @ pro6000,2026-09-08 双场景报告口径) diff --git a/deploy/profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp2pp4_hicache.env b/deploy/profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp2pp4_hicache.env new file mode 100644 index 0000000..fad79de --- /dev/null +++ b/deploy/profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp2pp4_hicache.env @@ -0,0 +1,44 @@ +# GLM-5.3-NVFP4 SGLang TP=2 PP=4 + hicache deployment profile (single RTX 6000D node, 8 GPUs). +# 2026-09-09 128k 低并发容量扩容实验优胜配置(60.8 现役;60.5 交付 deploy_glm53_605_v2.sh 待执行)。 +# 可执行部署脚本:experiments/pro6000/glm53_nvfp4_128k_capacity_topology/scripts/deploy_glm53_605_v2.sh +# +# 与方案 D(glm53_nvfp4_pro6000_sglang_tp2pp4.env)的关键差异(勿混淆): +# - radix/hicache 保持开启(60.5 真实流量命中率 90%+,关 radix 不可接受);D 为 16k 独立输入场景关了 radix +# - cu13 镜像 + 9 补丁只读挂载(de-GLOO request_receiver / decode_cuda_graph_runner_fix 等 r37 栈遗产, +# nomtp 下 spec 相关补丁为惰性,de-GLOO 为 PP 通用修复);D 用旧镜像 20260828 无挂载 +# - --reasoning-parser glm45 --tool-call-parser glm47 齐备(质量门 7/7);D 当时 6/7 +# - --context-length 1048576(模型原生 1M;D 未设);chunk 8192(D 16384);MRR 16(D 48) +# - 实测(60.8,i128k/o512 冷缓存):KV 池 909,632 token(A 的 3.29×)、并发上限 c6、 +# 单条上限 ~909k(900k 实跑通过)、c4 输入/输出 5,907/23.1 tok/s、TTFT p50 43.4s、 +# 512k 单条 TTFT 88.4s、90% 命中 c4 输入 17,625 tok/s +# - memfrac 0.85 为验证档;PP0 stage 空闲 22GB 提示 0.88 有余量(未验证,改动须重跑质量门+容量冒烟) +PLATFORM=pro6000 +EXPERIMENT=glm53_nvfp4_128k_capacity_topology +MODEL_NAME=GLM-5.3-NVFP4 +ENGINE=sglang +RUNTIME=docker +DOCKER_IMAGE=lmsysorg/sglang:nightly-dev-cu13-20260901-07c8f729 +CONTAINER_NAME=glm53-nvfp4 +MODEL_PATH=/data/hf_models/GLM-5.3-NVFP4 +SERVED_MODEL_NAME=/data/hf_models/GLM-5.3-NVFP4 +PORT=30000 +HEALTH_PATH=/health +HEALTH_WAIT_S=1800 +CONTAINER_PYTHON=python3 + +TP=2 +PP=4 +MEM_FRACTION_STATIC=0.85 +MAX_RUNNING_REQUESTS=16 +CHUNKED_PREFILL_SIZE=8192 +CONTEXT_LENGTH=1048576 + +DEVICE_VARS="CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7" +ENGINE_ENV="PYTHONUNBUFFERED=1 SGLANG_PP_DEGLOO=1 SGLANG_PP_SPEC_FORCE_EAGER_DRAFT=1 SGLANG_PP_FORCE_EAGER_VERIFY=0 SGLANG_PP_SPEC_DEBUG=0" + +DOCKER_FLAGS="--gpus all --shm-size 64g --ipc=host --cap-add SYS_PTRACE -p ${PORT}:${PORT}" +VOLUMES="/data/hf_models:/data/hf_models + 9 patch ro-mounts (full list in scripts/deploy_glm53_605_v2.sh; files live in /root + /root/sglang_patch2 on the host)" + +BOOTSTRAP="python3 -m sglang.launch_server ${LAUNCH_ARGS}" + +LAUNCH_ARGS="--model-path ${MODEL_PATH} --tp-size ${TP} --pp-size ${PP} --mem-fraction-static ${MEM_FRACTION_STATIC} --max-running-requests ${MAX_RUNNING_REQUESTS} --chunked-prefill-size ${CHUNKED_PREFILL_SIZE} --disable-shared-experts-fusion --moe-runner-backend flashinfer_cutlass --disable-flashinfer-autotune --reasoning-parser glm45 --tool-call-parser glm47 --enable-hierarchical-cache --hicache-ratio 3 --disable-overlap-schedule --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length ${CONTEXT_LENGTH} --host 0.0.0.0 --port ${PORT}" diff --git a/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/README.md b/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/README.md new file mode 100644 index 0000000..4bc0e1f --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/README.md @@ -0,0 +1,135 @@ +# GLM-5.3-NVFP4 128k 低并发容量扩容与拓扑选型实验报告 — 2026-09-09 @ 174.1.60.8 + +## 一、结论速览 + +**优胜配置 = TP2PP4 nomtp**(radix/hicache 保持开启、`--context-length 1048576`、cu13 栈 9 挂载、memfrac 0.85、chunk 8192)。 + +| 维度 | 60.5 现役(TP8+EAGLE) | 优胜配置(TP2PP4 nomtp) | 提升 | +|---|---|---|---| +| KV 池(token) | 276,864 | **909,632** | 3.29× | +| i128k 并发上限 | c2(c3 起排队) | **c6**(c7 起排队,实测) | 3× | +| 单条请求上限 | 264k(ctx 270,336 参数所限) | **~909k**(900k 实跑通过) | 3.4× | +| c4 输入吞吐 | 3,228 tok/s | **5,907 tok/s** | 1.83× | +| c4 输出吞吐 | 12.6 tok/s | **23.1 tok/s** | 1.83× | +| c4 TTFT p50 | 119.6s | **43.4s** | −64% | +| c4 e2e p50 | 161.3s | **88.8s** | −45% | +| 90% 命中真实流量形态(c4) | — | 输入 **17,625 tok/s**、TTFT 7.3s | — | + +- 质量门 7/7 通过(GSM8K×5 + 中文推理 + tool call,PP4+radix+hicache+parser 新组合无 correctness 问题)。 +- **60.8 已按用户决定部署优胜配置留役**(restart=unless-stopped,替代 r37)。 +- **60.5 交付 `deploy_glm53_605_v2.sh`(只交未执行)**,原 `deploy_glm53_605.sh` 原样留盘即回滚路径。 + +## 二、问题诊断(60.5 生产日志实证) + +现役容器 `glm53-nvfp4`(TP8 + EAGLE 4/1/5,memfrac 0.90,fp8 KV,hicache-ratio 3,`--context-length 270336`): + +1. **"单条最高 256k" 不是模型限制**:模型原生 `max_position_embeddings = 1,048,576`(1M)。上限来自部署参数 `--context-length 270336`,且 KV 池 276,864 也刚好只装得下一条 26.4 万请求。 +2. **排队是纯容量算术**:单条 16~22 万 token 请求占池 0.58~0.90;近期一万行日志 ~13% 的 decode 步有 `#queue-req>0`,09-08 有一次池满 retract。日志实例:一条 173k(95% 命中)请求在 160k 请求 decode 期间排队无法准入。 +3. **命中不省并发容量**:radix/hicache 只加速重复前缀的 prefill;不同文档的并发请求各自全额驻留 KV,照样排队。60.5 真实流量 12万~22万 token、命中率 90%+,仍然撞容量墙。 +4. **TP8 调参无解(结构性)**:每 token KV 57.3KB(MLA 44.9 + DSA indexer ~10),fp8_e4m3 已是断言封死的硬顶(FP4 启动即 AssertionError);TP8 全系调参上限 314,944(60.6 实测 0.89+hicache4)< c3×128k 需求 394,752。 + +## 三、机理:为什么只有 PP 拓扑能解 + +GLM-5.3 是 MLA/DSA 家族,**KV 在 TP 组内每卡全量复制、只按 PP 分层切**:TP8 每卡存全部 78 层 KV(15.85GB→276,864 token);PP2 每卡存半数层(池×2);PP4 每卡存 1/4 层(池×4,实际因 hicache/图开销打折)。**KV 池容量 ≈ PP 度数倍增,与 TP 度数无关**。这就是 TP4PP2≈2.37×、TP2PP4≈3.29×(实测)的来源。另:MTP/EAGLE 的 draft KV + verify 图额外吃池(r37 同 memfrac 下 589,696→384,960,−35%)。 + +## 四、实验方法 + +- **平台**:60.8(8×RTX 6000D 85.6GB,与 60.5 同型);四配置臂各部署后跑同窗口。 +- **场景**:i131072 / o512 / cc∈{1,2,3,4},nreq=8/点,PG19 真实语料(bench_corpus.py,input_ids 直发 /generate,temp 0,ignore_eos,stream)。 +- **公平性协议**:每点独立语料窗口(cc1→base 10M,cc2→11,048,576,cc3→12,097,152,cc4→13,145,728),**所有臂用同一映射=同文本**;每点前 `POST /flush_cache`(实测连 host 层一起清,复跑命中率 0.0);每臂先做一次不计分的 131k 形状预热(吸收内核首触成本,实测不做会污染首点:16k 首跑 1653 vs 复跑 4526 tok/s)。 +- **指标口径**(沿用 rev20 双场景协议):输入吞吐 = 总输入 token/墙钟;输出吞吐 = 服务端 completion_tokens 总和/墙钟;TTFT = 发出→首个流式响应(含排队);命中率从 TP0 Prefill 日志核算。queue-req 日志用于排队取证。 +- **配置臂**: + - **arm0 r37**:60.8 在役原样(TP4PP2+EAGLE 3/1/4+verify 图,memfrac 0.88,池 384,960) + - **arm1 A-mirror**:60.5 现役逐参数复刻(旧镜像 20260828,池 276,864,ctx 270,336) + - **arm2 B'**:TP4PP2 nomtp,memfrac 0.88,ctx 524,288(池 589,696) + - **arm3 TP2PP4**:TP2PP4 nomtp,memfrac 0.85,ctx 1048576,radix/hicache 开,含 glm45/glm47 parser(池 909,632) + +## 五、主扫数据(i128k/o512,冷缓存) + +**输入吞吐(tok/s)** + +| cc | A-mirror(池276,864) | r37(池384,960) | B'(池589,696) | TP2PP4(池909,632) | +|---|---|---|---|---| +| 1 | 3,026 | **4,110** | 3,737 | 3,262 | +| 2 | 3,142 | **4,754** | 4,619 | 4,635 | +| 3 | 3,214 | 5,016 | 4,843 | **5,017** | +| 4 | 3,228 | 5,061 | 5,256 | **5,907** | + +**输出吞吐(tok/s)** + +| cc | A-mirror | r37 | B' | TP2PP4 | +|---|---|---|---|---| +| 1 | 11.8 | **16.1** | 14.6 | 12.7 | +| 2 | 12.3 | **18.6** | 18.0 | 18.1 | +| 3 | 12.6 | 19.6 | 18.9 | **19.6** | +| 4 | 12.6 | 19.8 | 20.5 | **23.1** | + +**TTFT p50 / max(秒)**(排队签名 = p50 跳升整请求时长倍数) + +| cc | A-mirror | r37 | B' | TP2PP4 | +|---|---|---|---|---| +| 1 | 37.8 | 21.1 | 21.0 | **15.5** | +| 2 | 37.8 / 76.9 | 21.1 / 41.7 | 41.3 | 29.0 | +| 3 | **77.2** / 120.3 | 43.6 / 73.8 | 41.3 / 61.6 | **29.6** / 43.4 | +| 4 | **119.6** / 156.4 | 72.2 / 93.9 | 61.6 / 81.9 | **43.4** / 57.3 | + +**e2e p50 / max(秒)** + +| cc | A-mirror | r37 | B' | TP2PP4 | +|---|---|---|---|---| +| 1 | 43.8 | **31.7** | 35.0 | 40.2 | +| 2 | 83.7 / 121.6 | **55.3** / 76.6 | 56.8 | 56.5 | +| 3 | 121.7 / 163.3 | **75.2** / 105.1 | 80.0 | 75.8 | +| 4 | 161.3 / 201.8 | 103.3 / 125.2 | 99.7 | **88.8** | + +**排队证据与机理注解**: +- A-mirror:池只容 2 条并发,c3/c4 排队(TTFT p50 77→120s);c4 输入吞吐被排队锁死在 3,228(prefill 带宽根本没用满)。 +- r37:池 384,960 同样只容 2 条 + 1 条错峰补位,c3/c4 排队(TTFT p50 43.6→72.2s;日志 518 行 queue-req>0)。**60.8 原在役配置在本场景也不合格。** +- B':c3/c4 全并发准入、零容量排队(稳态 usage 精确停在 0.22/0.45/0.67/0.89;e2e p50≈max 波内同步)。日志中的 queue-req 行是 chunk 调度/波间 radix 逐出瞬态(秒级),非容量排队——判据:TTFT 缩放=纯 prefill 带宽分摊(cc2 p50=max=41.3≈2×21.0)。 +- TP2PP4:同上零容量排队,且 TTFT 全场最优(4 级流水 prefill 重叠最深:cc1 TTFT 15.5s、串行 prefill 等效 ~8.5k tok/s);代价是 TPOT 最慢(cc1 48.4ms vs B' 27.3 vs r37 20.7 vs A 12.0),c1 短输出场景 e2e 吃亏(40.2s),c2 起被并发摊平、c3/c4 反超。 +- MTP/EAGLE 观察:accept 随并发上升(r37 2.91→3.64,A-mirror 3.31→4.26,verify batch 越大接受越高);但 MTP 的 35% 池代价在本容量场景不划算——r37 全程吞吐被 B'/TP2PP4 压制或打平。 + +## 六、优胜者(TP2PP4)附加验证 + +| 验证项 | 结果 | +|---|---| +| 质量门 | **7/7**(GSM8K 72/3/60/63/10 + 鸡兔同笼 23 + tool call get_weather北京) | +| 512k 单条(523,776+512 真实语料) | ✅ TTFT 88.4s,e2e 114.4s,TPOT 50.9ms(旧配置直接拒绝) | +| 900k 单条(900,000+512) | ✅ TTFT 210.6s,e2e 237.8s —— **单条上限 ~909k 实证**(池减 512 后的余量) | +| cc5 | ✅ in 5,747 / out 22.5,e2e p50≈max 107s,零容量排队 | +| cc6 | ✅ in 5,849 / out 22.9,e2e p50≈max 122.7s,零容量排队 —— **并发上限 = c6** | +| cc7 | ⚠️ 第 7 条排队(TTFT max 138.7s、e2e max 179.7s)—— 边界与池算术吻合(7×131,584=921,088 > 909,632) | +| 90% 命中 c4(60.5 真实流量形态,实测命中 0.8999) | in **17,625** / out 68.9 tok/s,TTFT p50 **7.3s**,e2e 29.9s —— radix 去重+hicache 对重复查询流量再放大 ~3× | + +## 七、方案对比与推荐 + +| 方案 | 池 | c4 in/out | 容量定位 | 判定 | +|---|---|---|---|---| +| A(TP8+EAGLE,60.5 现役) | 276,864 | 3,228/12.6 | c2、单条 264k | 本场景被全面支配,淘汰 | +| r37(TP4PP2+MTP,60.8 昨日在役) | 384,960 | 5,061/19.8 | c2、单条 ~380k | c3 起排队,容量不合格;仅 c1 最优 | +| B'(TP4PP2 nomtp) | 589,696 | 5,256/20.5 | c4、单条 ~588k(512k 单条会占 89% 池、堵死并发) | 强力备选:16k 短文本场景历史成绩好 | +| **TP2PP4 nomtp(优胜)** | **909,632** | **5,907/23.1** | **c6、单条 ~909k、512k 单条+2 并发共存** | 用户需求(512k 单条+长上下文为主)下的正解 | + +**推荐**:60.5 采用 TP2PP4 nomtp(`deploy_glm53_605_v2.sh`,参数与 60.8 现役完全一致)。决策依据:c3/c4/c5/c6 吞吐全场第一 + 唯一满足"512k 单条与并发共存" + 质量门 7/7 + TTFT 全档最优。若 60.5 未来 16k 短文本流量占比显著上升,再评估切 B'(其 s1 短文本历史成绩比 TP2PP4 好 27~48%)。 + +**留观旋钮**(未验证,报告只记录不推荐):memfrac 0.85→0.88 或可再抬池(PP0 空闲 22GB 最大的 stage 不均匀提示有空间);TP2PP4+MTP 可修 c1 短输出短板(池约降至 ~59 万,恰为 B' 水平);hicache-ratio 3→4 扩 host 前缀池(60.6 在 TP8 验证过)。 + +**r37 角色变化说明**:r37 的 MTP 优势区间是 i16k/cc≤16(09-09 判决 +6%@cc16),本次让位给容量优先的 TP2PP4 是按用户明确选择执行;若 60.8 未来主要服务短文本低并发,可用 deploy_ppmtp_r37.sh mtp 模式一键切回。 + +## 八、已证伪 / 排除项(本轮+引用前判) + +- TP8 任何调参:fp8 KV 断言封顶 + 上限 314,944 < c3 需求 394,752(60.6 数据)。 +- r37 现役直接顶上:池 384,960,c3 差 1 万 token 仍排队(本轮实测)。 +- MTP/EAGLE 换 KV 池:draft+verify 图 −35% 池,本场景不划算。 +- HiSparse 超池驻留:本 nightly 自旋不可用(60.6 前判)。 +- TP1PP8:池 2M 但吞吐 −12~33%(60.2 前判),仅当需要 >c6 且能接受慢时再议。 +- PD 分离/DP/EP:前判劣化或不可用,与本问题正交。 + +## 九、资产与复现 + +- **60.8**:容器 glm53-nvfp4:30000 = 优胜配置,restart=unless-stopped。启动命令 = `bash /root/deploy_ppmtp_r37.sh '--tp 2 --pp-size 4 --disable-overlap-schedule --max-prefill-tokens 16384 --disable-custom-all-reduce --context-length 1048576' nomtp 8192 0.85`。实验日志全量:`/root/bench_logs/128kexp/`(26 点 SUMMARY 已提取为 all_summaries.json)。 +- **60.5**:交付 `deploy_glm53_605_v2.sh`(内嵌完整 docker run+前置检查+显存归零等待+健康门+回滚提示);回滚 = `bash /root/deploy_glm53_605.sh`。执行前置检查会列出缺失的补丁/镜像清单(60.8 /root 均有)。 +- **压测复现**:`bash /root/arm_runner.sh `(c1-4 四点+形状预热)、`bash /root/val_runner.sh `(512k/900k/cc5-7/hit90);语料窗口映射与 flush 协议见第四节;质量门 `bash /root/quality_gate_605.sh`。 +- **仓库归档**:`experiments/pro6000/glm53_nvfp4_128k_capacity_topology/`(README=本报告、scripts、results/20260909);profile `deploy/profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp2pp4_hicache.env`;CURRENT.md 已更新 60.8 行。 + +*报告与数据:ZCode 实验 2026-09-09;压测窗口 run-id 9501-9507/9601-9604/9990-9991;全部冷缓存口径。* diff --git a/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/results/20260909/all_summaries.json b/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/results/20260909/all_summaries.json new file mode 100644 index 0000000..7d74ddc --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/results/20260909/all_summaries.json @@ -0,0 +1,1484 @@ +{ + "arm0_r37_cc1": { + "concurrency": 1, + "num_requests": 8, + "run_id": 9501, + "corpus_window": { + "start": 10000000, + "end": 11048576 + }, + "ok": 8, + "failed": 0, + "wall_s": 255.13, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 16.05, + "input_throughput_tok_s": 4109.96, + "ttft_s": { + "mean": 21.1436, + "p50": 21.1351, + "max": 21.1968, + "min": 21.1318 + }, + "tpot_s": { + "mean": 0.021, + "p50": 0.0207, + "max": 0.0255, + "min": 0.0176 + }, + "e2e_s": { + "mean": 31.8909, + "p50": 31.6992, + "max": 34.1398, + "min": 30.1727 + }, + "per_req_out_tok_s_e2e": { + "mean": 16.0805, + "p50": 16.3007, + "max": 16.969, + "min": 14.9972 + }, + "per_req_decode_tok_s": { + "mean": 48.3191, + "p50": 49.8219, + "max": 57.0435, + "min": 39.3633 + }, + "spec_accept_length_mean": 2.912, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 256, + "new_tokens": 2097152, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm0_r37_cc2": { + "concurrency": 2, + "num_requests": 8, + "run_id": 9502, + "corpus_window": { + "start": 11048576, + "end": 12097152 + }, + "ok": 8, + "failed": 0, + "wall_s": 220.58, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 18.57, + "input_throughput_tok_s": 4753.73, + "ttft_s": { + "mean": 23.7214, + "p50": 21.1489, + "max": 41.6861, + "min": 21.1264 + }, + "tpot_s": { + "mean": 0.0613, + "p50": 0.0669, + "max": 0.0689, + "min": 0.0242 + }, + "e2e_s": { + "mean": 55.0699, + "p50": 55.3137, + "max": 76.6223, + "min": 33.5372 + }, + "per_req_out_tok_s_e2e": { + "mean": 9.719, + "p50": 9.3973, + "max": 15.2666, + "min": 6.6821 + }, + "per_req_decode_tok_s": { + "mean": 18.3311, + "p50": 15.3492, + "max": 41.3304, + "min": 14.5499 + }, + "spec_accept_length_mean": 2.672, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 256, + "new_tokens": 2097152, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm0_r37_cc3": { + "concurrency": 3, + "num_requests": 8, + "run_id": 9503, + "corpus_window": { + "start": 12097152, + "end": 13145728 + }, + "ok": 8, + "failed": 0, + "wall_s": 209.05, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 19.59, + "input_throughput_tok_s": 5016.0, + "ttft_s": { + "mean": 45.9651, + "p50": 43.5894, + "max": 73.7557, + "min": 21.2775 + }, + "tpot_s": { + "mean": 0.0509, + "p50": 0.0613, + "max": 0.0655, + "min": 0.0182 + }, + "e2e_s": { + "mean": 71.9541, + "p50": 75.2035, + "max": 105.0954, + "min": 51.2337 + }, + "per_req_out_tok_s_e2e": { + "mean": 7.5561, + "p50": 6.8525, + "max": 9.9934, + "min": 4.8718 + }, + "per_req_decode_tok_s": { + "mean": 25.7245, + "p50": 16.4485, + "max": 55.2037, + "min": 15.3031 + }, + "spec_accept_length_mean": 3.493, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 256, + "new_tokens": 2097152, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm0_r37_cc4": { + "concurrency": 4, + "num_requests": 8, + "run_id": 9504, + "corpus_window": { + "start": 13145728, + "end": 14194304 + }, + "ok": 8, + "failed": 0, + "wall_s": 207.18, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 19.77, + "input_throughput_tok_s": 5061.26, + "ttft_s": { + "mean": 67.6095, + "p50": 72.2484, + "max": 93.9173, + "min": 21.2909 + }, + "tpot_s": { + "mean": 0.0498, + "p50": 0.0608, + "max": 0.1012, + "min": 0.0183 + }, + "e2e_s": { + "mean": 93.0788, + "p50": 103.3347, + "max": 125.1566, + "min": 51.1124 + }, + "per_req_out_tok_s_e2e": { + "mean": 5.8997, + "p50": 4.9549, + "max": 10.0171, + "min": 4.0909 + }, + "per_req_decode_tok_s": { + "mean": 33.4842, + "p50": 52.2793, + "max": 54.7641, + "min": 9.898 + }, + "spec_accept_length_mean": 3.635, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 256, + "new_tokens": 2097152, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm0_r37_warmup": { + "concurrency": 1, + "num_requests": 1, + "run_id": 9991, + "corpus_window": { + "start": 17000000, + "end": 17131072 + }, + "ok": 1, + "failed": 0, + "wall_s": 23.66, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 64, + "output_tokens_total": 64, + "output_throughput_tok_s": 2.7, + "input_throughput_tok_s": 5539.35, + "ttft_s": { + "mean": 21.2651, + "p50": 21.2651, + "max": 21.2651, + "min": 21.2651 + }, + "tpot_s": { + "mean": 0.038, + "p50": 0.038, + "max": 0.038, + "min": 0.038 + }, + "e2e_s": { + "mean": 23.6611, + "p50": 23.6611, + "max": 23.6611, + "min": 23.6611 + }, + "per_req_out_tok_s_e2e": { + "mean": 2.7049, + "p50": 2.7049, + "max": 2.7049, + "min": 2.7049 + }, + "per_req_decode_tok_s": { + "mean": 26.7145, + "p50": 26.7145, + "max": 26.7145, + "min": 26.7145 + }, + "spec_accept_length_mean": 2.133, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 32, + "new_tokens": 262144, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm1_Amirror_cc1": { + "concurrency": 1, + "num_requests": 8, + "run_id": 9501, + "corpus_window": { + "start": 10000000, + "end": 11048576 + }, + "ok": 8, + "failed": 0, + "wall_s": 346.56, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 11.82, + "input_throughput_tok_s": 3025.64, + "ttft_s": { + "mean": 37.7456, + "p50": 37.7503, + "max": 37.7573, + "min": 37.7043 + }, + "tpot_s": { + "mean": 0.0109, + "p50": 0.012, + "max": 0.0132, + "min": 0.0083 + }, + "e2e_s": { + "mean": 43.32, + "p50": 43.8174, + "max": 44.4887, + "min": 41.9741 + }, + "per_req_out_tok_s_e2e": { + "mean": 11.8244, + "p50": 11.7971, + "max": 12.198, + "min": 11.5085 + }, + "per_req_decode_tok_s": { + "mean": 94.6188, + "p50": 90.5656, + "max": 121.3763, + "min": 76.065 + }, + "spec_accept_length_mean": 3.314, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 128, + "new_tokens": 1048576, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm1_Amirror_cc2": { + "concurrency": 2, + "num_requests": 8, + "run_id": 9502, + "corpus_window": { + "start": 11048576, + "end": 12097152 + }, + "ok": 8, + "failed": 0, + "wall_s": 333.76, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 12.27, + "input_throughput_tok_s": 3141.72, + "ttft_s": { + "mean": 42.8869, + "p50": 37.7865, + "max": 76.9116, + "min": 37.7742 + }, + "tpot_s": { + "mean": 0.0793, + "p50": 0.0892, + "max": 0.1617, + "min": 0.0132 + }, + "e2e_s": { + "mean": 83.4182, + "p50": 83.749, + "max": 121.6486, + "min": 44.5522 + }, + "per_req_out_tok_s_e2e": { + "mean": 6.5603, + "p50": 6.1412, + "max": 11.4921, + "min": 4.2088 + }, + "per_req_decode_tok_s": { + "mean": 26.5566, + "p50": 11.3508, + "max": 75.6796, + "min": 6.1948 + }, + "spec_accept_length_mean": 2.922, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 128, + "new_tokens": 1048576, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm1_Amirror_cc3": { + "concurrency": 3, + "num_requests": 8, + "run_id": 9503, + "corpus_window": { + "start": 12097152, + "end": 13145728 + }, + "ok": 8, + "failed": 0, + "wall_s": 326.28, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 12.55, + "input_throughput_tok_s": 3213.76, + "ttft_s": { + "mean": 78.0816, + "p50": 77.2216, + "max": 120.3096, + "min": 38.9908 + }, + "tpot_s": { + "mean": 0.076, + "p50": 0.0846, + "max": 0.1619, + "min": 0.0095 + }, + "e2e_s": { + "mean": 116.8932, + "p50": 121.7292, + "max": 163.3489, + "min": 82.0943 + }, + "per_req_out_tok_s_e2e": { + "mean": 4.5817, + "p50": 4.2567, + "max": 6.2367, + "min": 3.1344 + }, + "per_req_decode_tok_s": { + "mean": 30.6847, + "p50": 11.8916, + "max": 105.0827, + "min": 6.1882 + }, + "spec_accept_length_mean": 4.019, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 128, + "new_tokens": 1048576, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm1_Amirror_cc4": { + "concurrency": 4, + "num_requests": 8, + "run_id": 9504, + "corpus_window": { + "start": 13145728, + "end": 14194304 + }, + "ok": 8, + "failed": 0, + "wall_s": 324.86, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 12.61, + "input_throughput_tok_s": 3227.8, + "ttft_s": { + "mean": 112.9849, + "p50": 119.5745, + "max": 156.4269, + "min": 38.9673 + }, + "tpot_s": { + "mean": 0.0659, + "p50": 0.0817, + "max": 0.1632, + "min": 0.0091 + }, + "e2e_s": { + "mean": 146.6423, + "p50": 161.3281, + "max": 201.7932, + "min": 80.7246 + }, + "per_req_out_tok_s_e2e": { + "mean": 3.9136, + "p50": 3.1898, + "max": 6.3426, + "min": 2.5373 + }, + "per_req_decode_tok_s": { + "mean": 43.9614, + "p50": 12.2625, + "max": 110.4456, + "min": 6.1379 + }, + "spec_accept_length_mean": 4.257, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 128, + "new_tokens": 1048576, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm1_Amirror_warmup": { + "concurrency": 1, + "num_requests": 1, + "run_id": 9991, + "corpus_window": { + "start": 17000000, + "end": 17131072 + }, + "ok": 1, + "failed": 0, + "wall_s": 52.85, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 64, + "output_tokens_total": 64, + "output_throughput_tok_s": 1.21, + "input_throughput_tok_s": 2480.16, + "ttft_s": { + "mean": 51.3163, + "p50": 51.3163, + "max": 51.3163, + "min": 51.3163 + }, + "tpot_s": { + "mean": 0.0243, + "p50": 0.0243, + "max": 0.0243, + "min": 0.0243 + }, + "e2e_s": { + "mean": 52.8471, + "p50": 52.8471, + "max": 52.8471, + "min": 52.8471 + }, + "per_req_out_tok_s_e2e": { + "mean": 1.211, + "p50": 1.211, + "max": 1.211, + "min": 1.211 + }, + "per_req_decode_tok_s": { + "mean": 41.8164, + "p50": 41.8164, + "max": 41.8164, + "min": 41.8164 + }, + "spec_accept_length_mean": 2.065, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 16, + "new_tokens": 131072, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm2_Bp_cc1": { + "concurrency": 1, + "num_requests": 8, + "run_id": 9501, + "corpus_window": { + "start": 10000000, + "end": 11048576 + }, + "ok": 8, + "failed": 0, + "wall_s": 280.57, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 14.6, + "input_throughput_tok_s": 3737.27, + "ttft_s": { + "mean": 21.0555, + "p50": 21.0457, + "max": 21.1298, + "min": 21.0325 + }, + "tpot_s": { + "mean": 0.0274, + "p50": 0.0273, + "max": 0.0282, + "min": 0.0271 + }, + "e2e_s": { + "mean": 35.0713, + "p50": 35.0062, + "max": 35.544, + "min": 34.8595 + }, + "per_req_out_tok_s_e2e": { + "mean": 14.5994, + "p50": 14.6312, + "max": 14.6875, + "min": 14.4047 + }, + "per_req_decode_tok_s": { + "mean": 36.5374, + "p50": 36.704, + "max": 37.0298, + "min": 35.5211 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 256, + "new_tokens": 2097152, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm2_Bp_cc2": { + "concurrency": 2, + "num_requests": 8, + "run_id": 9502, + "corpus_window": { + "start": 11048576, + "end": 12097152 + }, + "ok": 8, + "failed": 0, + "wall_s": 226.99, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 18.04, + "input_throughput_tok_s": 4619.49, + "ttft_s": { + "mean": 31.2138, + "p50": 41.3247, + "max": 41.3388, + "min": 21.0898 + }, + "tpot_s": { + "mean": 0.05, + "p50": 0.0695, + "max": 0.0701, + "min": 0.0298 + }, + "e2e_s": { + "mean": 56.7421, + "p50": 56.7513, + "max": 56.9058, + "min": 56.5568 + }, + "per_req_out_tok_s_e2e": { + "mean": 9.0233, + "p50": 9.0273, + "max": 9.0529, + "min": 8.9973 + }, + "per_req_decode_tok_s": { + "mean": 23.8132, + "p50": 32.8787, + "max": 33.6254, + "min": 14.2991 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 256, + "new_tokens": 2097152, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm2_Bp_cc3": { + "concurrency": 3, + "num_requests": 8, + "run_id": 9503, + "corpus_window": { + "start": 12097152, + "end": 13145728 + }, + "ok": 8, + "failed": 0, + "wall_s": 216.51, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 18.92, + "input_throughput_tok_s": 4843.07, + "ttft_s": { + "mean": 38.8062, + "p50": 41.3294, + "max": 61.5934, + "min": 21.0703 + }, + "tpot_s": { + "mean": 0.0691, + "p50": 0.0755, + "max": 0.1156, + "min": 0.0294 + }, + "e2e_s": { + "mean": 74.1346, + "p50": 79.9715, + "max": 80.2009, + "min": 56.3659 + }, + "per_req_out_tok_s_e2e": { + "mean": 7.0671, + "p50": 6.4036, + "max": 9.0835, + "min": 6.384 + }, + "per_req_decode_tok_s": { + "mean": 18.4756, + "p50": 14.4943, + "max": 34.0308, + "min": 8.6643 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 256, + "new_tokens": 2097152, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm2_Bp_cc4": { + "concurrency": 4, + "num_requests": 8, + "run_id": 9504, + "corpus_window": { + "start": 13145728, + "end": 14194304 + }, + "ok": 8, + "failed": 0, + "wall_s": 199.49, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 20.53, + "input_throughput_tok_s": 5256.26, + "ttft_s": { + "mean": 51.477, + "p50": 61.5611, + "max": 81.8537, + "min": 21.0928 + }, + "tpot_s": { + "mean": 0.0944, + "p50": 0.1142, + "max": 0.1539, + "min": 0.0349 + }, + "e2e_s": { + "mean": 99.7182, + "p50": 99.7397, + "max": 99.7733, + "min": 99.6468 + }, + "per_req_out_tok_s_e2e": { + "mean": 5.1345, + "p50": 5.1347, + "max": 5.1381, + "min": 5.1316 + }, + "per_req_decode_tok_s": { + "mean": 14.3454, + "p50": 13.4435, + "max": 28.7267, + "min": 6.5103 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 256, + "new_tokens": 2097152, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm2_Bp_warmup": { + "concurrency": 1, + "num_requests": 1, + "run_id": 9991, + "corpus_window": { + "start": 17000000, + "end": 17131072 + }, + "ok": 1, + "failed": 0, + "wall_s": 35.22, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 64, + "output_tokens_total": 64, + "output_throughput_tok_s": 1.82, + "input_throughput_tok_s": 3721.64, + "ttft_s": { + "mean": 32.9797, + "p50": 32.9797, + "max": 32.9797, + "min": 32.9797 + }, + "tpot_s": { + "mean": 0.0355, + "p50": 0.0355, + "max": 0.0355, + "min": 0.0355 + }, + "e2e_s": { + "mean": 35.2179, + "p50": 35.2179, + "max": 35.2179, + "min": 35.2179 + }, + "per_req_out_tok_s_e2e": { + "mean": 1.8173, + "p50": 1.8173, + "max": 1.8173, + "min": 1.8173 + }, + "per_req_decode_tok_s": { + "mean": 28.5963, + "p50": 28.5963, + "max": 28.5963, + "min": 28.5963 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 32, + "new_tokens": 262144, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm3_TP2PP4_cc1": { + "concurrency": 1, + "num_requests": 8, + "run_id": 9501, + "corpus_window": { + "start": 10000000, + "end": 11048576 + }, + "ok": 8, + "failed": 0, + "wall_s": 321.42, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 12.74, + "input_throughput_tok_s": 3262.35, + "ttft_s": { + "mean": 15.4876, + "p50": 15.4928, + "max": 15.5353, + "min": 15.4459 + }, + "tpot_s": { + "mean": 0.0483, + "p50": 0.0484, + "max": 0.0486, + "min": 0.0481 + }, + "e2e_s": { + "mean": 40.1768, + "p50": 40.2302, + "max": 40.3235, + "min": 40.0206 + }, + "per_req_out_tok_s_e2e": { + "mean": 12.7438, + "p50": 12.7615, + "max": 12.7934, + "min": 12.6973 + }, + "per_req_decode_tok_s": { + "mean": 20.7383, + "p50": 20.8009, + "max": 20.8346, + "min": 20.6009 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 512, + "new_tokens": 4194304, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm3_TP2PP4_cc2": { + "concurrency": 2, + "num_requests": 8, + "run_id": 9502, + "corpus_window": { + "start": 11048576, + "end": 12097152 + }, + "ok": 8, + "failed": 0, + "wall_s": 226.22, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 18.11, + "input_throughput_tok_s": 4635.23, + "ttft_s": { + "mean": 22.3167, + "p50": 28.9744, + "max": 29.1154, + "min": 15.5008 + }, + "tpot_s": { + "mean": 0.067, + "p50": 0.08, + "max": 0.0808, + "min": 0.0535 + }, + "e2e_s": { + "mean": 56.5495, + "p50": 56.4636, + "max": 56.8831, + "min": 56.4112 + }, + "per_req_out_tok_s_e2e": { + "mean": 9.0541, + "p50": 9.0686, + "max": 9.0762, + "min": 9.0009 + }, + "per_req_decode_tok_s": { + "mean": 15.566, + "p50": 18.4753, + "max": 18.7252, + "min": 12.3965 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 512, + "new_tokens": 4194304, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm3_TP2PP4_cc3": { + "concurrency": 3, + "num_requests": 8, + "run_id": 9503, + "corpus_window": { + "start": 12097152, + "end": 13145728 + }, + "ok": 8, + "failed": 0, + "wall_s": 209.0, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 19.6, + "input_throughput_tok_s": 5017.09, + "ttft_s": { + "mean": 27.8189, + "p50": 29.6062, + "max": 43.4436, + "min": 15.6499 + }, + "tpot_s": { + "mean": 0.085, + "p50": 0.09, + "max": 0.1191, + "min": 0.0533 + }, + "e2e_s": { + "mean": 71.2601, + "p50": 75.7564, + "max": 76.4978, + "min": 56.8297 + }, + "per_req_out_tok_s_e2e": { + "mean": 7.3001, + "p50": 6.7662, + "max": 9.0094, + "min": 6.693 + }, + "per_req_decode_tok_s": { + "mean": 12.6863, + "p50": 12.4942, + "max": 18.8075, + "min": 8.4145 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 512, + "new_tokens": 4194304, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm3_TP2PP4_cc4": { + "concurrency": 4, + "num_requests": 8, + "run_id": 9504, + "corpus_window": { + "start": 13145728, + "end": 14194304 + }, + "ok": 8, + "failed": 0, + "wall_s": 177.53, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 23.07, + "input_throughput_tok_s": 5906.56, + "ttft_s": { + "mean": 36.6103, + "p50": 43.4223, + "max": 57.2876, + "min": 15.9191 + }, + "tpot_s": { + "mean": 0.102, + "p50": 0.1154, + "max": 0.1425, + "min": 0.0615 + }, + "e2e_s": { + "mean": 88.7369, + "p50": 88.7669, + "max": 88.8538, + "min": 88.5975 + }, + "per_req_out_tok_s_e2e": { + "mean": 5.7699, + "p50": 5.7685, + "max": 5.7789, + "min": 5.7623 + }, + "per_req_decode_tok_s": { + "mean": 10.8261, + "p50": 11.3182, + "max": 16.279, + "min": 7.0293 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 512, + "new_tokens": 4194304, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm3_TP2PP4_cc5": { + "concurrency": 5, + "num_requests": 8, + "run_id": 9505, + "corpus_window": { + "start": 10000000, + "end": 11048576 + }, + "ok": 8, + "failed": 0, + "wall_s": 182.46, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 22.45, + "input_throughput_tok_s": 5746.99, + "ttft_s": { + "mean": 37.5353, + "p50": 42.5635, + "max": 69.4847, + "min": 15.615 + }, + "tpot_s": { + "mean": 0.1127, + "p50": 0.117, + "max": 0.1789, + "min": 0.0641 + }, + "e2e_s": { + "mean": 95.1469, + "p50": 106.9463, + "max": 107.052, + "min": 75.3993 + }, + "per_req_out_tok_s_e2e": { + "mean": 5.5367, + "p50": 4.7905, + "max": 6.7905, + "min": 4.7827 + }, + "per_req_decode_tok_s": { + "mean": 9.8884, + "p50": 10.0708, + "max": 15.6197, + "min": 5.6013 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 512, + "new_tokens": 4194304, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm3_TP2PP4_cc6": { + "concurrency": 6, + "num_requests": 8, + "run_id": 9506, + "corpus_window": { + "start": 11048576, + "end": 12097152 + }, + "ok": 8, + "failed": 0, + "wall_s": 179.26, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 22.85, + "input_throughput_tok_s": 5849.32, + "ttft_s": { + "mean": 42.7136, + "p50": 42.8146, + "max": 83.2404, + "min": 15.5862 + }, + "tpot_s": { + "mean": 0.1242, + "p50": 0.13, + "max": 0.2096, + "min": 0.0536 + }, + "e2e_s": { + "mean": 106.1613, + "p50": 122.6846, + "max": 122.8274, + "min": 56.4675 + }, + "per_req_out_tok_s_e2e": { + "mean": 5.3952, + "p50": 4.1733, + "max": 9.0672, + "min": 4.1684 + }, + "per_req_decode_tok_s": { + "mean": 9.7806, + "p50": 9.6772, + "max": 18.6901, + "min": 4.7795 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 512, + "new_tokens": 4194304, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm3_TP2PP4_cc7": { + "concurrency": 7, + "num_requests": 8, + "run_id": 9507, + "corpus_window": { + "start": 12097152, + "end": 13145728 + }, + "ok": 8, + "failed": 0, + "wall_s": 179.89, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 22.77, + "input_throughput_tok_s": 5829.09, + "ttft_s": { + "mean": 58.4235, + "p50": 56.691, + "max": 138.7352, + "min": 15.683 + }, + "tpot_s": { + "mean": 0.1241, + "p50": 0.1298, + "max": 0.2103, + "min": 0.0532 + }, + "e2e_s": { + "mean": 121.849, + "p50": 123.0753, + "max": 179.69, + "min": 56.7662 + }, + "per_req_out_tok_s_e2e": { + "mean": 4.6041, + "p50": 4.1611, + "max": 9.0195, + "min": 2.8494 + }, + "per_req_decode_tok_s": { + "mean": 9.8211, + "p50": 9.7386, + "max": 18.8168, + "min": 4.7644 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 512, + "new_tokens": 4194304, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm3_TP2PP4_hit90_cc4": { + "concurrency": 4, + "num_requests": 8, + "run_id": 9604, + "corpus_window": { + "start": 15000000, + "end": 15104832 + }, + "ok": 8, + "failed": 0, + "wall_s": 59.49, + "input_len": 131072, + "shared_len": 117968, + "unique_len": 13104, + "output_len": 512, + "output_tokens_total": 4096, + "output_throughput_tok_s": 68.85, + "input_throughput_tok_s": 17624.84, + "ttft_s": { + "mean": 6.9268, + "p50": 7.2908, + "max": 8.986, + "min": 4.5732 + }, + "tpot_s": { + "mean": 0.0446, + "p50": 0.0451, + "max": 0.0489, + "min": 0.041 + }, + "e2e_s": { + "mean": 29.7055, + "p50": 29.9081, + "max": 30.0152, + "min": 29.4371 + }, + "per_req_out_tok_s_e2e": { + "mean": 17.237, + "p50": 17.3599, + "max": 17.393, + "min": 17.058 + }, + "per_req_decode_tok_s": { + "mean": 22.5696, + "p50": 23.1193, + "max": 24.449, + "min": 20.4936 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 56, + "new_tokens": 419840, + "cached_tokens": 3774464, + "hit_rate": 0.8999 + } + }, + "arm3_TP2PP4_single512k": { + "concurrency": 1, + "num_requests": 1, + "run_id": 9601, + "corpus_window": { + "start": 10000000, + "end": 10523776 + }, + "ok": 1, + "failed": 0, + "wall_s": 114.38, + "input_len": 523776, + "shared_len": 0, + "unique_len": 523776, + "output_len": 512, + "output_tokens_total": 512, + "output_throughput_tok_s": 4.48, + "input_throughput_tok_s": 4579.23, + "ttft_s": { + "mean": 88.3831, + "p50": 88.3831, + "max": 88.3831, + "min": 88.3831 + }, + "tpot_s": { + "mean": 0.0509, + "p50": 0.0509, + "max": 0.0509, + "min": 0.0509 + }, + "e2e_s": { + "mean": 114.3797, + "p50": 114.3797, + "max": 114.3797, + "min": 114.3797 + }, + "per_req_out_tok_s_e2e": { + "mean": 4.4763, + "p50": 4.4763, + "max": 4.4763, + "min": 4.4763 + }, + "per_req_decode_tok_s": { + "mean": 19.6952, + "p50": 19.6952, + "max": 19.6952, + "min": 19.6952 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 256, + "new_tokens": 2095104, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm3_TP2PP4_single900k": { + "concurrency": 1, + "num_requests": 1, + "run_id": 9602, + "corpus_window": { + "start": 10000000, + "end": 10900000 + }, + "ok": 1, + "failed": 0, + "wall_s": 237.77, + "input_len": 900000, + "shared_len": 0, + "unique_len": 900000, + "output_len": 512, + "output_tokens_total": 512, + "output_throughput_tok_s": 2.15, + "input_throughput_tok_s": 3785.22, + "ttft_s": { + "mean": 210.587, + "p50": 210.587, + "max": 210.587, + "min": 210.587 + }, + "tpot_s": { + "mean": 0.0532, + "p50": 0.0532, + "max": 0.0532, + "min": 0.0532 + }, + "e2e_s": { + "mean": 237.7662, + "p50": 237.7662, + "max": 237.7662, + "min": 237.7662 + }, + "per_req_out_tok_s_e2e": { + "mean": 2.1534, + "p50": 2.1534, + "max": 2.1534, + "min": 2.1534 + }, + "per_req_decode_tok_s": { + "mean": 18.8381, + "p50": 18.8381, + "max": 18.8381, + "min": 18.8381 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 440, + "new_tokens": 3600128, + "cached_tokens": 0, + "hit_rate": 0.0 + } + }, + "arm3_TP2PP4_warmup": { + "concurrency": 1, + "num_requests": 1, + "run_id": 9991, + "corpus_window": { + "start": 17000000, + "end": 17131072 + }, + "ok": 1, + "failed": 0, + "wall_s": 31.74, + "input_len": 131072, + "shared_len": 0, + "unique_len": 131072, + "output_len": 64, + "output_tokens_total": 64, + "output_throughput_tok_s": 2.02, + "input_throughput_tok_s": 4129.93, + "ttft_s": { + "mean": 28.361, + "p50": 28.361, + "max": 28.361, + "min": 28.361 + }, + "tpot_s": { + "mean": 0.0536, + "p50": 0.0536, + "max": 0.0536, + "min": 0.0536 + }, + "e2e_s": { + "mean": 31.736, + "p50": 31.736, + "max": 31.736, + "min": 31.736 + }, + "per_req_out_tok_s_e2e": { + "mean": 2.0166, + "p50": 2.0166, + "max": 2.0166, + "min": 2.0166 + }, + "per_req_decode_tok_s": { + "mean": 18.9644, + "p50": 18.9644, + "max": 18.9644, + "min": 18.9644 + }, + "spec_accept_length_mean": null, + "retractions_total": 0, + "cache_hit_from_logs": { + "prefill_batches": 64, + "new_tokens": 524288, + "cached_tokens": 0, + "hit_rate": 0.0 + } + } +} diff --git a/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/results/20260909/md5s.txt b/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/results/20260909/md5s.txt new file mode 100644 index 0000000..51192f7 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/results/20260909/md5s.txt @@ -0,0 +1,3 @@ +e782a8767d4478711757669f83adb64b *scripts/arm_runner.sh +f6f5a49937ac6e8a0157ba1eb11b8ee7 *scripts/deploy_glm53_605_v2.sh +8eb75693f01f6055b15ecc3d62feff08 *scripts/val_runner.sh diff --git a/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/scripts/arm_runner.sh b/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/scripts/arm_runner.sh new file mode 100644 index 0000000..08040ed --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/scripts/arm_runner.sh @@ -0,0 +1,36 @@ +#!/bin/bash +# 128k capacity experiment - per-arm bench runner (i131072/o512, cc 1-4, nreq 8) +# usage: nohup bash arm_runner.sh >/dev/null 2>&1 & +# per-cc-point corpus windows are distinct but the mapping is identical for every arm: +# cc1 base 10000000, cc2 11048576, cc3 12097152, cc4 13145728 (all in consumed space, flushed) +ARM=${1:?arm name required} +URL=http://127.0.0.1:30000 +LOGD=/root/bench_logs/128kexp +mkdir -p $LOGD +echo "$(date +%T) [$ARM] start" >> $LOGD/${ARM}_progress.log + +# wait until server idle (no running/queued requests in recent logs) +for i in $(seq 1 60); do + R=$(docker logs --since 20s glm53-nvfp4 2>&1 | grep -cE 'running-req: [1-9]|queue-req: [1-9]') + if [ "$R" -eq 0 ]; then break; fi + sleep 10 +done + +flush() { curl -s -X POST $URL/flush_cache >/dev/null; sleep 2; } + +# shape warmup (not scored): one 128k prefill to absorb kernel first-touch cost +flush +python3 /root/bench_corpus.py --input-len 131072 --output-len 64 --shared-frac 0 \ + --concurrency 1 --num-requests 1 --run-id 9991 --pool-override 17000000 \ + > $LOGD/${ARM}_warmup.log 2>&1 +echo "$(date +%T) [$ARM] warmup done" >> $LOGD/${ARM}_progress.log + +declare -A BASE=( [1]=10000000 [2]=11048576 [3]=12097152 [4]=13145728 ) +for CC in 1 2 3 4; do + flush + python3 /root/bench_corpus.py --input-len 131072 --output-len 512 --shared-frac 0 \ + --concurrency $CC --num-requests 8 --run-id 950${CC} --pool-override ${BASE[$CC]} \ + > $LOGD/${ARM}_cc${CC}.log 2>&1 + echo "$(date +%T) [$ARM] cc$CC done rc=$?" >> $LOGD/${ARM}_progress.log +done +echo "ALL-DONE" >> $LOGD/${ARM}_progress.log diff --git a/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/scripts/deploy_glm53_605_v2.sh b/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/scripts/deploy_glm53_605_v2.sh new file mode 100644 index 0000000..e428fd5 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/scripts/deploy_glm53_605_v2.sh @@ -0,0 +1,129 @@ +#!/bin/bash +# deploy_glm53_605_v2.sh — GLM-5.3-NVFP4 60.5 生产容量扩容配置(2026-09-09 交付) +# +# 目的:解决 60.5 现役 TP8+EAGLE 配置的 KV 池容量问题(池 276,864 token): +# - 单条请求上限 264k(--context-length 270336 所致,模型原生支持 1M) +# - 并发>1 且总 token>~26 万时排队(i128k 下 c3 起必然排队) +# 本配置 = 2026-09-09 在 60.8 同型机(8×RTX 6000D)四臂对拍优胜者: +# TP2 PP4 nomtp + radix/hicache 开 + ctx 1M + cu13 栈 9 挂载 +# 实测池 909,632 token(3.29×):i128k 并发上限 c2→c6,单条上限 264k→909k, +# c4 输入/输出吞吐 5907/23.1 tok/s(现役口径 3228/12.6 的 1.83×/1.83×), +# 质量门 7/7。完整数据:experiments/pro6000/glm53_nvfp4_128k_capacity_topology/ +# +# ⚠️ 运维提示: +# 1) 在维护窗口执行:会移除现役 glm53-nvfp4 容器,停机约 10-15 分钟(权重加载+健康)。 +# 2) 回滚 = 原脚本原样留盘:bash /root/deploy_glm53_605.sh(TP8+EAGLE 原配置)。 +# 3) 与 60.8 现役(本配置已留役)一致,脚本可互相对拍 md5。 +# +# 用法: bash deploy_glm53_605_v2.sh # 部署到 30000 端口 +# ROLLBACK_ONLY=1 bash deploy_glm53_605_v2.sh # 仅回滚到原配置 + +set -u +IMAGE=lmsysorg/sglang:nightly-dev-cu13-20260901-07c8f729 +PORT=30000 + +# ---- 前置检查(缺什么给什么清单,不猜) ---- +MISSING=0 +if [ -z "$(docker images -q $IMAGE 2>/dev/null)" ]; then + echo "[MISS] 镜像不存在: $IMAGE" + echo " 拉取: docker pull $IMAGE" + echo " 或从 60.8 导: ssh 60.8 'docker save $IMAGE | gzip' | gunzip | docker load" + MISSING=1 +fi +MOUNTS=( + "/root/sglang_patch2/layer_setup.py" + "/root/sglang_patch2/validation_hook.py" + "/root/eagle_worker_v2_mask.py" + "/root/sglang_patch2/eagle_worker_common.py" + "/root/sglang_patch2/deepseek_nextn.py" + "/root/scheduler_pp_mixin_r35.py" + "/root/request_receiver_degloo.py" + "/root/sglang_patch2/deepseek_v2.py" + "/root/decode_cuda_graph_runner_fix.py" +) +for f in "${MOUNTS[@]}"; do + if [ ! -f "$f" ]; then + echo "[MISS] 补丁文件缺失: $f" + MISSING=1 + fi +done +if [ "$MISSING" = "1" ]; then + echo "[STOP] 上述文件在 60.8 /root 均有(含 sglang_patch2/ 子目录),scp 全部补齐后重跑本脚本。" + exit 1 +fi + +if [ "${ROLLBACK_ONLY:-0}" = "1" ]; then + echo "[rollback] 恢复 60.5 原生产配置(TP8+EAGLE)..." + exec bash /root/deploy_glm53_605.sh +fi + +# ---- 移除旧容器(等待显存归零,纪律:重部署前必等显存<1500MiB) ---- +docker update --restart=no glm53-nvfp4 >/dev/null 2>&1 +for i in 1 2 3 4 5 6 7 8; do + docker rm -f glm53-nvfp4 >/dev/null 2>&1 + docker ps -a --format '{{.Names}}' | grep -q '^glm53-nvfp4$' || break + sleep 5 +done +if docker ps -a --format '{{.Names}}' | grep -q '^glm53-nvfp4$'; then + echo "[ERROR] 旧容器删不掉(zombie 时重试即可)"; exit 1 +fi +echo "[wait] 等待显存释放 (<1500MiB/卡)..." +for i in $(seq 1 60); do + MAXMI=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | sort -rn | head -1) + [ "$MAXMI" -lt 1500 ] && break + sleep 5 +done +echo "[wait] 显存就绪 (${MAXMI}MiB max)" + +# ---- 部署优胜配置(与 60.8 实验臂 arm3_TP2PP4 完全一致) ---- +docker run -d --name glm53-nvfp4 --gpus all --shm-size 64g --ipc=host --cap-add SYS_PTRACE \ + -v /root/sglang_patch2/layer_setup.py:/sgl-workspace/sglang/python/sglang/srt/model_executor/model_runner_components/layer_setup.py:ro \ + -v /root/sglang_patch2/validation_hook.py:/sgl-workspace/sglang/python/sglang/srt/arg_groups/validation_hook.py:ro \ + -v /root/eagle_worker_v2_mask.py:/sgl-workspace/sglang/python/sglang/srt/speculative/eagle_worker_v2.py:ro \ + -v /root/sglang_patch2/eagle_worker_common.py:/sgl-workspace/sglang/python/sglang/srt/speculative/eagle_worker_common.py:ro \ + -v /root/sglang_patch2/deepseek_nextn.py:/sgl-workspace/sglang/python/sglang/srt/models/deepseek_nextn.py:ro \ + -v /root/scheduler_pp_mixin_r35.py:/sgl-workspace/sglang/python/sglang/srt/managers/scheduler_pp_mixin.py:ro \ + -v /root/request_receiver_degloo.py:/sgl-workspace/sglang/python/sglang/srt/managers/scheduler_components/request_receiver.py:ro \ + -v /root/sglang_patch2/deepseek_v2.py:/sgl-workspace/sglang/python/sglang/srt/models/deepseek_v2.py:ro \ + -v /root/decode_cuda_graph_runner_fix.py:/sgl-workspace/sglang/python/sglang/srt/model_executor/runner/decode_cuda_graph_runner.py:ro \ + -e SGLANG_PP_DEGLOO=1 -e SGLANG_PP_SPEC_FORCE_EAGER_DRAFT=1 -e SGLANG_PP_FORCE_EAGER_VERIFY=0 -e SGLANG_PP_SPEC_DEBUG=0 \ + --restart no -p ${PORT}:${PORT} \ + -v /data/hf_models:/data/hf_models \ + $IMAGE \ + python3 -m sglang.launch_server \ + --model-path /data/hf_models/GLM-5.3-NVFP4 \ + --tp 8 \ + --mem-fraction-static 0.85 \ + --max-running-requests 16 \ + --chunked-prefill-size 8192 \ + --disable-shared-experts-fusion \ + --moe-runner-backend flashinfer_cutlass \ + --disable-flashinfer-autotune \ + --reasoning-parser glm45 --tool-call-parser glm47 \ + --enable-hierarchical-cache --hicache-ratio 3 \ + --tp 2 --pp-size 4 --disable-overlap-schedule --max-prefill-tokens 16384 \ + --disable-custom-all-reduce --context-length 1048576 \ + --host 0.0.0.0 --port ${PORT} + +echo "[deploy] TP2PP4-nomtp 容量扩容配置启动,等待健康(最长 1800s)..." +for i in $(seq 10 10 1800); do + code=$(curl -s -o /dev/null -m3 -w '%{http_code}' http://127.0.0.1:${PORT}/health 2>/dev/null) + if [ "$code" = "200" ]; then + echo "[OK] healthy after ${i}s" + docker update --restart=unless-stopped glm53-nvfp4 >/dev/null && echo "[OK] restart=unless-stopped 已设置" + POOL=$(docker logs glm53-nvfp4 2>&1 | grep -m1 -oE 'max_total_num_tokens=[0-9]+') + echo "[INFO] ${POOL} (60.8 同型机实测 909632;若显著低于此值请停下核查)" + echo "[NEXT] 质量门: bash /root/quality_gate_605.sh 期望 PASS=7 FAIL=0" + echo "[NEXT] 容量冒烟: 300k 单条请求应能正常完成(旧配置会直接拒绝)" + exit 0 + fi + if ! docker ps --format '{{.Names}}' | grep -q '^glm53-nvfp4$'; then + echo "[DIED] 容器启动失败,最近错误:" + docker logs --tail 40 glm53-nvfp4 2>&1 | grep -iE 'error|assert|not support' | tail -8 + echo "[回滚] bash /root/deploy_glm53_605.sh" + exit 1 + fi + sleep 10 +done +echo "[TIMEOUT] 健康等待超时;回滚: bash /root/deploy_glm53_605.sh" +exit 1 diff --git a/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/scripts/val_runner.sh b/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/scripts/val_runner.sh new file mode 100644 index 0000000..67ef397 --- /dev/null +++ b/experiments/pro6000/glm53_nvfp4_128k_capacity_topology/scripts/val_runner.sh @@ -0,0 +1,48 @@ +#!/bin/bash +# winner validation runner (TP2PP4): big singles, cc boundary, 90%-hit point +# usage: nohup bash val_runner.sh >/dev/null 2>&1 & +ARM=${1:?arm name required} +URL=http://127.0.0.1:30000 +LOGD=/root/bench_logs/128kexp +echo "$(date +%T) [$ARM] val start" >> $LOGD/${ARM}_progress.log + +for i in $(seq 1 60); do + R=$(docker logs --since 20s glm53-nvfp4 2>&1 | grep -cE 'running-req: [1-9]|queue-req: [1-9]') + if [ "$R" -eq 0 ]; then break; fi + sleep 10 +done + +flush() { curl -s -X POST $URL/flush_cache >/dev/null; sleep 2; } + +# 1) 512k single (real corpus): 523776+512 = 524288 +flush +python3 /root/bench_corpus.py --input-len 523776 --output-len 512 --shared-frac 0 \ + --concurrency 1 --num-requests 1 --run-id 9601 --pool-override 10000000 \ + > $LOGD/${ARM}_single512k.log 2>&1 +echo "$(date +%T) [$ARM] single512k rc=$?" >> $LOGD/${ARM}_progress.log + +# 2) 900k single (pool ceiling probe): 900000+512 = 900512 vs pool 909632 +flush +python3 /root/bench_corpus.py --input-len 900000 --output-len 512 --shared-frac 0 \ + --concurrency 1 --num-requests 1 --run-id 9602 --pool-override 10000000 \ + > $LOGD/${ARM}_single900k.log 2>&1 +echo "$(date +%T) [$ARM] single900k rc=$?" >> $LOGD/${ARM}_progress.log + +# 3-5) cc boundary: 5 / 6 fit, 7 exceeds pool (921088 > 909632) -> expect queuing +declare -A BASEB=( [5]=10000000 [6]=11048576 [7]=12097152 ) +for CC in 5 6 7; do + flush + python3 /root/bench_corpus.py --input-len 131072 --output-len 512 --shared-frac 0 \ + --concurrency $CC --num-requests 8 --run-id 950${CC} --pool-override ${BASEB[$CC]} \ + > $LOGD/${ARM}_cc${CC}.log 2>&1 + echo "$(date +%T) [$ARM] cc$CC rc=$?" >> $LOGD/${ARM}_progress.log +done + +# 6) 90%-hit cc4 point (mirrors 60.5 real traffic: shared 117968 + unique 13104x8) +flush +python3 /root/bench_corpus.py --input-len 131072 --output-len 512 --shared-frac 0.9 \ + --concurrency 4 --num-requests 8 --run-id 9604 --pool-override 15000000 \ + > $LOGD/${ARM}_hit90_cc4.log 2>&1 +echo "$(date +%T) [$ARM] hit90_cc4 rc=$?" >> $LOGD/${ARM}_progress.log + +echo "VAL-ALL-DONE" >> $LOGD/${ARM}_progress.log