feat(pro6000/GLM-5.3): i8k/o1k/c16 三方案对比压测入库(A/A16/B 轮换实测 + 60.8 转 A16 在役 + CURRENT.md 台账修正 60.2/60.3/60.6/60.8)

This commit is contained in:
yy-fighting 2026-09-08 19:04:41 +08:00
parent 5c749cda03
commit 123023b6ae
16 changed files with 1076 additions and 5 deletions

View File

@ -8,13 +8,13 @@
| 机器 | 在役 | 口径 / 归属 | 对应 profile | | 机器 | 在役 | 口径 / 归属 | 对应 profile |
|---|---|---|---| |---|---|---|---|
| 60.1 (6000D-1) | `glm53-pp4`Up8 卡满载,:30000 | **方案 D 生产**。09-08 PD 压测窗口停机 ~2.5h 后已恢复并核验 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp2pp4.env` | | 60.1 (6000D-1) | `glm53-pp4`Up8 卡满载,:30000 | **方案 D 生产**。09-08 PD 压测窗口停机 ~2.5h 后已恢复并核验 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp2pp4.env` |
| 60.2 (6000D-2) | 无容器GPU 全 0 MiB | 方案 F PD 链已于 09-08 拆除。GPU7 曾有外部裸金属任务 main_v2.py现已结束动卡前仍先核实归属 | `profiles/pro6000/glm53_nvfp4_pro6000_pd_{prefill→decode 侧}.env` | | 60.2 (6000D-2) | `glm53-nvfp4` 实验容器09-08 晚 TP1PP8 phaserun_phase.sh 实验链进行中) | **他人实验进行中,勿动**(方案 F PD 链已拆除GPU7 曾有外部裸金属任务)。动卡前仍先核实归属 | `profiles/pro6000/glm53_nvfp4_pro6000_pd_{prefill→decode 侧}.env` |
| 60.3 | 空(无容器) | — | — | | 60.3 | 无容器,但 8 卡被外部裸金属训练占用(`/data/mas/larm`09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — |
| 60.4 | `dsv4_scan` 容器Up 5min+ 裸金属 sglang `DeepSeek-V4-Flash-0731` DSPARK DP2/TP4 :30000 | **并行会话/外部在役,勿动** | — | | 60.4 | `dsv4_scan` 容器Up 5min+ 裸金属 sglang `DeepSeek-V4-Flash-0731` DSPARK DP2/TP4 :30000 | **并行会话/外部在役,勿动** | — |
| 60.5 | `glm53-nvfp4`Up 29h | **NVFP4 团队生产**deploy_glm53_605.shmd5 fcd9109b。生产机铁律不实验、不重启、不覆盖脚本 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(方案 A 口径) | | 60.5 | `glm53-nvfp4`Up 29h | **NVFP4 团队生产**deploy_glm53_605.shmd5 fcd9109b。生产机铁律不实验、不重启、不覆盖脚本 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`(方案 A 口径) |
| 60.6 | 空(无容器) | — | — | | 60.6 | 无容器,但 8 卡被外部裸金属实验占用(`/data/hzy/sparse-opd-*`09-08 晚实测) | 外部任务,勿动(此前台账漏记) | — |
| 60.7 | 空 | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — | | 60.7 | 基本4 卡仍有 `/home/user/dirA_exp` 外部小任务09-08 晚实测) | 09-08 已拆除清空(方案 F 前身单机实验 + 场景一深优资产留盘),不再恢复 | — |
| 60.8 | 空 | 09-08 已停服清空、router 拆除,不再恢复 | — | | 60.8 | `glm53-nvfp4`Up:30000restart=unless-stopped | **A16 口径在役**09-08 i8k/o1k/c16 压测收官保留tp8eagle.env 场景二变体 = A + MRR32 + graph bs 4/8/12/16池 276,864质量门 7/7。同场景实测最优为 BTP4PP2out 265 vs A16 219 tok/s如转纯吞吐用途可切 B。压测前经授权清理了 GPU3/6 的 dirA_ext 外部 eval 进程 | `profiles/pro6000/glm53_nvfp4_pro6000_sglang_tp8eagle.env`+注释中场景二高并发变体);压测记录 `experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/` |
## 方案 A-F 一览GLM-5.3-NVFP4 @ pro60002026-09-08 双场景报告口径) ## 方案 A-F 一览GLM-5.3-NVFP4 @ pro60002026-09-08 双场景报告口径)

View File

@ -0,0 +1,83 @@
# GLM-5.3-NVFP4 i8k/o1k/c16 三方案部署对比压测报告
- 日期2026-09-08 18:0019:02单机串行174.1.60.8 / 6000D-88×H20
- 场景:**input=8192 / output=1024 / concurrency=16**全唯一输入PG19 真实语料)
- 方案A生产口径 TP8+EAGLE、BTP4PP2 无投机、A16A 的 c16 调优graph 覆盖 bs16
- 协议rev20 团队标准bench_corpus.pyinput_ids 直打 `/generate`temp=0、ignore_eos、streamnreq=32=2 波warm nreq16 弃用,每轮全新 run-id + 全新语料窗口 `--pool-override`
- 镜像:`lmsysorg/sglang:nightly-dev-20260828-daf63171`digest 28e0d2607316模型 `/data/hf_models/GLM-5.3-NVFP4`
## 一、判决(先行)
| 指标 | A 生产口径 | A16graph 修复) | BTP4PP2 | 最优 |
|---|---|---|---|---|
| **输出吞吐 tok/s** | 115.0 | 218.9 | **264.9** | BA 的 2.30× |
| **输入吞吐 tok/s** | 919.9 | 1,750.8 | **2,119.4** | BA 的 2.30× |
| **TTFT p50 s** | **3.8** | 4.0 | 10.5 | Amean 三者打平 ~10s |
| **TPOT p50 ms** | 123.1 | 55.0 | **50.3** | B |
| **e2e P50 s** | 139.6 | 66.2 | **61.8** | BA 的 44% |
| 单流 decode p50 tok/s | 8.4 | 18.9 | **20.1** | B |
| EAGLE accept | 2.52 | 2.50 | —(无投机) | — |
| 回退 / 命中 | 0 / 0.0 | 0 / 0.0 | 0 / 0.0 | 全过 |
| 质量门 | 7/7 | 7/7 | 6/7* | — |
\* B 的 tool-call 项失败 = 脚本无 `--tool-call-parser` 的已知配置缺口(模型文本明确想调工具,非质量问题;上生产须补 parser见双场景报告既有结论
**结论i8k/o1k/c16 场景 BTP4PP2全面最优**——吞吐与 e2e 均为 A 的 2.3 倍量级,且 32 请求 e2e 几乎零离散61.362.4s。A16 把 A 与 B 的差距从 2.30× 收窄到 1.21×,是 EAGLE 家族在本场景的正确口径。**60.8 按既定安排保留最后部署的 A16 在役**:30000质量门 7/7、parser 完整);若该机器要转纯吞吐用途,一条命令可切 B补 parser flags
## 二、主表(逐轮原始数据)
wall | out tok/s | in tok/s | TTFT mean/p50/max s | TPOT mean/p50 ms | e2e mean/p50/max s | 单流 decode p50 tok/s | accept | 命中 | 回退
**方案 ATP8 + EAGLE 4/1/5 @ mem0.90MRR16chunk8192fp8KV+hicache3graph bs≤8KV 池 276,864**deploy_glm53_605.sh 原样md5 fcd9109b与 60.5 生产同源)
| 轮 | wall | out | in | TTFT mean/p50/max | TPOT mean/p50 | e2e mean/p50/max | 单流dec | accept | 命中 | 回退 |
|---|---|---|---|---|---|---|---|---|---|---|
| r1 | 277.4 | 118.1 | 945.0 | 10.29 / 4.37 / 31.18 | 119.6 / 118.4 | 132.6 / 136.2 / 190.1 | 8.8 | 2.589 | 0.0 | 0 |
| r2 | 293.0 | 111.8 | 894.8 | 10.02 / 3.22 / 30.94 | 126.9 / 127.8 | 139.8 / 142.9 / 193.1 | 7.9 | 2.445 | 0.0 | 0 |
**方案 A16A 基础上 EXTRA=`--max-running-requests 32 --cuda-graph-max-bs-decode 16 --cuda-graph-bs-decode 4 8 12 16`**= tp8eagle.env 注释中的场景二高并发变体KV 池不变 276,864
| 轮 | wall | out | in | TTFT mean/p50/max | TPOT mean/p50 | e2e mean/p50/max | 单流dec | accept | 命中 | 回退 |
|---|---|---|---|---|---|---|---|---|---|---|
| r1 | 151.3 | 216.6 | 1,732.4 | 10.06 / 4.02 / 30.63 | 60.6 / 57.7 | 72.1 / 70.0 / 108.1 | 18.6 | 2.367 | 0.0 | 0 |
| r2 | 148.2 | 221.2 | 1,769.2 | 9.88 / 3.98 / 30.67 | 57.3 / 52.2 | 68.5 / 62.3 / 109.6 | 19.2 | 2.634 | 0.0 | 0 |
**方案 BTP4 PP2 @ mem0.85MRR48chunk16384radix-offdisable-custom-ARindex_topk_freq=4KV 池 569,600**deploy_glm53_optimal.sh与 09-07/09-08 实跑版同源)
| 轮 | wall | out | in | TTFT mean/p50/max | TPOT mean/p50 | e2e mean/p50/max | 单流dec | accept | 命中 | 回退 |
|---|---|---|---|---|---|---|---|---|---|---|
| r1 | 124.2 | 263.8 | 2,110.4 | 10.36 / 10.54 / 17.36 | 50.5 / 50.6 | 62.0 / 62.0 / 62.4 | 20.0 | — | 0.0 | 0 |
| r2 | 123.2 | 266.1 | 2,128.4 | 10.38 / 10.54 / 17.36 | 50.0 / 50.0 | 61.5 / 61.5 / 61.8 | 20.1 | — | 0.0 | 0 |
两轮间一致性A ±5%、A16 ±9%TPOT在跨窗口噪声带内、B ±1%。
## 三、关键发现
1. **A 生产口径在 cc16 有 decode 陷阱**decode 批 bs 916 落在 cuda graph 之外graph max bs=8EAGLE 掉图后 TPOT 123ms比 B 的无投机裸 decode50ms还慢 2.4×。历史"EAGLE 大胜"结论cc14不可外推到 cc16+:掉图模式下投机收益完全反转。
2. **A16 的 graph 修复立竿见影**TPOT 123→55ms55%、out 115→219+90%、e2e P50 139.6→66.2s53%。这是生产配置在该场景下零成本可得的改进EXTRA 三个参数)。
3. **EAGLE 在 cc16 高并发下优势摊薄至零**A16 accept 2.50,但每步墙钟 ≈ 55ms×2.5 ≈ 137msverify 批 16×5=80 token/步打满算力),折算单 token 55ms ≈ B 裸 decode 50ms。投机毛利被 verify 批次吃掉。
4. **prefill 是 A 家族剩余短板**Bchunk16384/TP4PP2输入吞吐 2,119 tok/sA16 1,751、A 920。B 的优势与既有"prefill 密集 cc≥16 TP4PP2 最优"结论一致;本场景 decode 占比翻倍o1k vs 历史 o512后 B 仍全胜。
5. **TTFT 的 p50/mean 悖论**TTFT mean 三方案打平(~10s但 p50 上 A/A16~4s优于 B10.5s。机理nreq32 两波客户端下A/A16 的 e2e 离散大→尾部请求陆续释放槽位,第二波"到站即过"B 全体同步完成,第二波整体重排 16 条 prefill。**对 TTFT-p50 敏感的小规模交互负载A/A16 的 chunk8192 体验更好对稳态吞吐B 完胜。**
6. e2e 离散度B 最优p50min/max 范围仅 0.5sA16 中等35110sA 最差79193s
## 四、收尾状态2026-09-08 19:02 核验)
- **60.8 在役**:容器 `glm53-nvfp4` = A16 口径docker Args 双值并存、生效值 MRR32/graph bs 4 8 12 16启动日志 `max_running_requests=32`、draft graph `bs=[4,8,12,16]` 实证),:30000health 20082.5GB×8池 276,864。
- A、B 已拆除,显存归零后才部署下一方案(每步 <2000MiB 纪律
- 60.1 生产 glm53-pp4、60.5 团队生产全程未动。
- **执行前经授权清理了 60.8 GPU3/6 的外部进程**PID 2421844/2421852user 账号 `/home/user/dirA_exp``eval_freeze_acc.py --mode remove --step 1/4`14:49 起,各占 7.6GB、util 35%SIGTERM 退出,显存归零。同源任务在 60.7 尚有 4 卡占用,未动。
## 五、复现与资产
- 压测命令60.8 本机):
`python3 /root/bench_corpus.py --corpus /root/corpus_ids.json --input-len 8192 --output-len 1024 --concurrency 16 --num-requests 32 --run-id <新> --pool-override <新窗口> --shared-frac 0 --container glm53-nvfp4`
- 轮次编排:`/root/bench_8k.sh <TAG> <POOL_BASE>`warm nreq16 + r1/r2 nreq32步进 300k
- 部署A=`bash /root/deploy_glm53_605.sh`A16=`EXTRA="--max-running-requests 32 --cuda-graph-max-bs-decode 16 --cuda-graph-bs-decode 4 8 12 16" bash /root/deploy_glm53_605.sh`B=`bash /root/deploy_glm53_optimal.sh`
- **语料窗口台账corpus_ids.json 共 21,296,780 token**:本轮消耗 A [17.00M,17.86M)、B [17.90M,18.76M)、A16 [18.80M,19.66M)**已消费至 ~19.66M,剩余 ~1.63M**——后续 8k 复测须从 ≥19.7M 取新窗口或重建语料。
- 指标口径:输入吞吐=8192×ok/墙钟;输出吞吐=服务端 completion_tokens 合计/墙钟EAGLE 下客户端计数低估 2.53.5×一律服务端口径TTFT 含排队TPOT=(首响应→末响应)/(completion1)P50=单请求 e2e 延迟中位数;命中由容器 TP0 Prefill 日志核算(全 0全唯一验证通过
- 日志归档60.8 `/root/bench_logs/`(含 3×gate、9 轮 bench、runner/deploy 日志)+ 本地 tar `D:\sskj\_xfer_608\i8k_bench_20260908.tar.gz`;入库 sskj-review `experiments/pro6000/glm53_nvfp4_i8k_o1k_c16_bench/`
## 六、与既有结论的关系
- 与《双场景压测报告》场景二16k/512一致性B 全胜的结论在 o1kdecode 占比翻倍)下依然成立且差距未缩小——说明 B 的优势主要来自 prefill 侧与池/调度,而非 decode 占比。
- 新增修正:**"A 生产口径"在 cc16 的历史场景二数据实际是 A16 变体口径**tp8eagle.env 注释的场景二变体);本报告首次把 A 原口径与 A16 在同一场景下分离量化——两者差距高达 1.9×,引用历史数据时须区分口径。

View File

@ -0,0 +1,31 @@
# GLM-5.3-NVFP4 i8k/o1k/c16 三方案部署对比压测2026-09-08
- 机器174.1.60.86000D-88×H20单机串行三方案轮换
- 场景input=8192 / output=1024 / cc=16 / nreq=322 波),全唯一 PG19 真实语料temp=0、ignore_eos、stream
- 完整报告:本目录 `GLM53_NVFP4_i8k_o1k_c16_压测报告_2026-09-08.md`(判决、主表、机理分析、复现命令)
## 方案与结果r1/r2 均值)
| 方案 | 配置要点 | out tok/s | in tok/s | TTFT p50 | TPOT p50 | e2e P50 | 质量门 |
|---|---|---|---|---|---|---|---|
| A 生产口径 | TP8+EAGLE 4/1/5MRR16chunk8192graph bs≤8池 276,864 | 115.0 | 920 | 3.8s | 123ms | 139.6s | 7/7 |
| A16 | A + `--max-running-requests 32 --cuda-graph-max-bs-decode 16 --cuda-graph-bs-decode 4 8 12 16` | 218.9 | 1,751 | 4.0s | 55ms | 66.2s | 7/7 |
| B | TP4PP2MRR48chunk16384radix-off池 569,600 | **264.9** | **2,119** | 10.5s | **50ms** | **61.8s** | 6/7* |
\* B 的 tool-call 失败 = 脚本无 `--tool-call-parser`(已知配置缺口,非质量问题)。
**判决**B 全面最优out 为 A 的 2.30×A16 把差距收窄到 1.21×,是 EAGLE 家族在本场景的正确口径。
60.8 收官保留 **A16 在役**:30000restart=unless-stopped
## 核心结论(引用时注意)
1. **A 生产口径在 cc16 有 decode 陷阱**bs 916 掉出 cuda graphEAGLE 掉图后 TPOT 123ms比 B 裸 decode50ms慢 2.4×——"EAGLE 大胜"仅成立于 cc≤4 或 graph 覆盖到位时。
2. **EAGLE 在 cc16 优势摊薄至零**A16 accept 2.50,但 verify 批 80 token/步打满算力,折算单 token 55ms ≈ B 的 50ms。
3. **历史场景二16k/512的"A"行实为 A16 变体口径**,与 A 原口径相差 1.9×,引用历史数据须区分。
## 资产
- `scripts/`bench_8k.sh轮次编排、bench_corpus.pyc6d9212660.7 同源、deploy_glm53_605.shfcd9109b8 机标准、deploy_glm53_optimal.sh22685f56
- `results/bench_logs/`3×gateA 7/7、B 6/7、A16 7/7、9 轮 bench 日志A/A16/B × warm/r1/r2、runner/deploy 日志
- 语料窗口:本轮消费 corpus_ids.json 至 **~19.66M / 21.30M**A [17.0,17.86M)、B [17.9,18.76M)、A16 [18.8,19.66M)),剩余 ~1.63M,复测须取 ≥19.7M 新窗口
- 压测机执行前经授权清理了 60.8 GPU3/6 的 dirA_ext 外部 eval 进程PID 2421844/2421852SIGTERM

View File

@ -0,0 +1,57 @@
{
"concurrency": 16,
"num_requests": 32,
"run_id": 9411,
"corpus_window": {
"start": 19100000,
"end": 19362144
},
"ok": 32,
"failed": 0,
"wall_s": 151.31,
"input_len": 8192,
"shared_len": 0,
"unique_len": 8192,
"output_len": 1024,
"output_tokens_total": 32768,
"output_throughput_tok_s": 216.55,
"input_throughput_tok_s": 1732.44,
"ttft_s": {
"mean": 10.0558,
"p50": 4.0156,
"max": 30.6315,
"min": 2.0783
},
"tpot_s": {
"mean": 0.0606,
"p50": 0.0577,
"max": 0.1026,
"min": 0.039
},
"e2e_s": {
"mean": 72.0651,
"p50": 69.9843,
"max": 108.1261,
"min": 41.9544
},
"per_req_out_tok_s_e2e": {
"mean": 15.3242,
"p50": 14.6458,
"max": 24.4074,
"min": 9.4704
},
"per_req_decode_tok_s": {
"mean": 17.6898,
"p50": 18.553,
"max": 25.6827,
"min": 9.7574
},
"spec_accept_length_mean": 2.367,
"retractions_total": 0,
"cache_hit_from_logs": {
"prefill_batches": 32,
"new_tokens": 262144,
"cached_tokens": 0,
"hit_rate": 0.0
}
}

View File

@ -0,0 +1,57 @@
{
"concurrency": 16,
"num_requests": 32,
"run_id": 9412,
"corpus_window": {
"start": 19400000,
"end": 19662144
},
"ok": 32,
"failed": 0,
"wall_s": 148.17,
"input_len": 8192,
"shared_len": 0,
"unique_len": 8192,
"output_len": 1024,
"output_tokens_total": 32768,
"output_throughput_tok_s": 221.15,
"input_throughput_tok_s": 1769.19,
"ttft_s": {
"mean": 9.875,
"p50": 3.9786,
"max": 30.6695,
"min": 2.0794
},
"tpot_s": {
"mean": 0.0573,
"p50": 0.0522,
"max": 0.102,
"min": 0.0307
},
"e2e_s": {
"mean": 68.4525,
"p50": 62.3499,
"max": 109.5734,
"min": 35.5019
},
"per_req_out_tok_s_e2e": {
"mean": 16.5606,
"p50": 16.6965,
"max": 28.8435,
"min": 9.3453
},
"per_req_decode_tok_s": {
"mean": 19.2407,
"p50": 19.179,
"max": 32.6448,
"min": 9.8182
},
"spec_accept_length_mean": 2.634,
"retractions_total": 0,
"cache_hit_from_logs": {
"prefill_batches": 32,
"new_tokens": 262144,
"cached_tokens": 0,
"hit_rate": 0.0
}
}

View File

@ -0,0 +1,57 @@
{
"concurrency": 16,
"num_requests": 16,
"run_id": 9401,
"corpus_window": {
"start": 18800000,
"end": 18931072
},
"ok": 16,
"failed": 0,
"wall_s": 90.65,
"input_len": 8192,
"shared_len": 0,
"unique_len": 8192,
"output_len": 1024,
"output_tokens_total": 16384,
"output_throughput_tok_s": 180.74,
"input_throughput_tok_s": 1445.94,
"ttft_s": {
"mean": 29.6679,
"p50": 30.755,
"max": 44.1152,
"min": 15.0222
},
"tpot_s": {
"mean": 0.0522,
"p50": 0.0521,
"max": 0.0739,
"min": 0.0306
},
"e2e_s": {
"mean": 83.0523,
"p50": 87.6337,
"max": 90.6464,
"min": 69.7624
},
"per_req_out_tok_s_e2e": {
"mean": 12.4314,
"p50": 11.9193,
"max": 14.6784,
"min": 11.2966
},
"per_req_decode_tok_s": {
"mean": 20.6139,
"p50": 20.1971,
"max": 32.678,
"min": 13.5407
},
"spec_accept_length_mean": 2.643,
"retractions_total": 0,
"cache_hit_from_logs": {
"prefill_batches": 16,
"new_tokens": 131072,
"cached_tokens": 0,
"hit_rate": 0.0
}
}

View File

@ -0,0 +1,57 @@
{
"concurrency": 16,
"num_requests": 32,
"run_id": 9411,
"corpus_window": {
"start": 17300000,
"end": 17562144
},
"ok": 32,
"failed": 0,
"wall_s": 277.41,
"input_len": 8192,
"shared_len": 0,
"unique_len": 8192,
"output_len": 1024,
"output_tokens_total": 32768,
"output_throughput_tok_s": 118.12,
"input_throughput_tok_s": 944.97,
"ttft_s": {
"mean": 10.2921,
"p50": 4.3675,
"max": 31.1779,
"min": 2.3895
},
"tpot_s": {
"mean": 0.1196,
"p50": 0.1184,
"max": 0.1667,
"min": 0.075
},
"e2e_s": {
"mean": 132.6294,
"p50": 136.2478,
"max": 190.146,
"min": 79.1449
},
"per_req_out_tok_s_e2e": {
"mean": 8.1891,
"p50": 7.536,
"max": 12.9383,
"min": 5.3853
},
"per_req_decode_tok_s": {
"mean": 8.7968,
"p50": 8.7859,
"max": 13.3427,
"min": 6.003
},
"spec_accept_length_mean": 2.589,
"retractions_total": 0,
"cache_hit_from_logs": {
"prefill_batches": 32,
"new_tokens": 262144,
"cached_tokens": 0,
"hit_rate": 0.0
}
}

View File

@ -0,0 +1,57 @@
{
"concurrency": 16,
"num_requests": 32,
"run_id": 9412,
"corpus_window": {
"start": 17600000,
"end": 17862144
},
"ok": 32,
"failed": 0,
"wall_s": 292.98,
"input_len": 8192,
"shared_len": 0,
"unique_len": 8192,
"output_len": 1024,
"output_tokens_total": 32768,
"output_throughput_tok_s": 111.84,
"input_throughput_tok_s": 894.75,
"ttft_s": {
"mean": 10.0189,
"p50": 3.2165,
"max": 30.9377,
"min": 2.4035
},
"tpot_s": {
"mean": 0.1269,
"p50": 0.1278,
"max": 0.1762,
"min": 0.077
},
"e2e_s": {
"mean": 139.8173,
"p50": 142.9177,
"max": 193.13,
"min": 81.2596
},
"per_req_out_tok_s_e2e": {
"mean": 7.6338,
"p50": 7.6204,
"max": 12.6016,
"min": 5.3021
},
"per_req_decode_tok_s": {
"mean": 8.1582,
"p50": 7.9203,
"max": 13.0008,
"min": 5.6814
},
"spec_accept_length_mean": 2.445,
"retractions_total": 0,
"cache_hit_from_logs": {
"prefill_batches": 32,
"new_tokens": 262144,
"cached_tokens": 0,
"hit_rate": 0.0
}
}

View File

@ -0,0 +1,57 @@
{
"concurrency": 16,
"num_requests": 16,
"run_id": 9401,
"corpus_window": {
"start": 17000000,
"end": 17131072
},
"ok": 16,
"failed": 0,
"wall_s": 172.19,
"input_len": 8192,
"shared_len": 0,
"unique_len": 8192,
"output_len": 1024,
"output_tokens_total": 16384,
"output_throughput_tok_s": 95.15,
"input_throughput_tok_s": 761.21,
"ttft_s": {
"mean": 29.1555,
"p50": 30.0011,
"max": 45.0124,
"min": 14.441
},
"tpot_s": {
"mean": 0.1241,
"p50": 0.1326,
"max": 0.1482,
"min": 0.0785
},
"e2e_s": {
"mean": 156.117,
"p50": 167.158,
"max": 172.1581,
"min": 114.2218
},
"per_req_out_tok_s_e2e": {
"mean": 6.7017,
"p50": 6.1651,
"max": 8.965,
"min": 5.948
},
"per_req_decode_tok_s": {
"mean": 8.3659,
"p50": 7.9182,
"max": 12.7508,
"min": 6.7522
},
"spec_accept_length_mean": 2.501,
"retractions_total": 0,
"cache_hit_from_logs": {
"prefill_batches": 16,
"new_tokens": 131072,
"cached_tokens": 0,
"hit_rate": 0.0
}
}

View File

@ -0,0 +1,57 @@
{
"concurrency": 16,
"num_requests": 32,
"run_id": 9411,
"corpus_window": {
"start": 18200000,
"end": 18462144
},
"ok": 32,
"failed": 0,
"wall_s": 124.22,
"input_len": 8192,
"shared_len": 0,
"unique_len": 8192,
"output_len": 1024,
"output_tokens_total": 32768,
"output_throughput_tok_s": 263.8,
"input_throughput_tok_s": 2110.38,
"ttft_s": {
"mean": 10.3634,
"p50": 10.5406,
"max": 17.3639,
"min": 1.777
},
"tpot_s": {
"mean": 0.0505,
"p50": 0.0506,
"max": 0.0592,
"min": 0.0434
},
"e2e_s": {
"mean": 62.019,
"p50": 62.0417,
"max": 62.366,
"min": 61.8
},
"per_req_out_tok_s_e2e": {
"mean": 16.5113,
"p50": 16.5486,
"max": 16.5696,
"min": 16.4192
},
"per_req_decode_tok_s": {
"mean": 19.9851,
"p50": 19.9608,
"max": 23.039,
"min": 16.9011
},
"spec_accept_length_mean": null,
"retractions_total": 0,
"cache_hit_from_logs": {
"prefill_batches": 36,
"new_tokens": 524288,
"cached_tokens": 0,
"hit_rate": 0.0
}
}

View File

@ -0,0 +1,57 @@
{
"concurrency": 16,
"num_requests": 32,
"run_id": 9412,
"corpus_window": {
"start": 18500000,
"end": 18762144
},
"ok": 32,
"failed": 0,
"wall_s": 123.16,
"input_len": 8192,
"shared_len": 0,
"unique_len": 8192,
"output_len": 1024,
"output_tokens_total": 32768,
"output_throughput_tok_s": 266.05,
"input_throughput_tok_s": 2128.41,
"ttft_s": {
"mean": 10.3812,
"p50": 10.5386,
"max": 17.3626,
"min": 1.7756
},
"tpot_s": {
"mean": 0.05,
"p50": 0.05,
"max": 0.0587,
"min": 0.043
},
"e2e_s": {
"mean": 61.5036,
"p50": 61.4987,
"max": 61.8083,
"min": 61.3264
},
"per_req_out_tok_s_e2e": {
"mean": 16.6495,
"p50": 16.6689,
"max": 16.6975,
"min": 16.5674
},
"per_req_decode_tok_s": {
"mean": 20.1968,
"p50": 20.1428,
"max": 23.2803,
"min": 17.0575
},
"spec_accept_length_mean": null,
"retractions_total": 0,
"cache_hit_from_logs": {
"prefill_batches": 36,
"new_tokens": 524288,
"cached_tokens": 0,
"hit_rate": 0.0
}
}

View File

@ -0,0 +1,57 @@
{
"concurrency": 16,
"num_requests": 16,
"run_id": 9401,
"corpus_window": {
"start": 17900000,
"end": 18031072
},
"ok": 16,
"failed": 0,
"wall_s": 65.46,
"input_len": 8192,
"shared_len": 0,
"unique_len": 8192,
"output_len": 1024,
"output_tokens_total": 16384,
"output_throughput_tok_s": 250.29,
"input_throughput_tok_s": 2002.32,
"ttft_s": {
"mean": 13.2299,
"p50": 13.271,
"max": 20.2234,
"min": 4.2816
},
"tpot_s": {
"mean": 0.0509,
"p50": 0.0507,
"max": 0.0598,
"min": 0.0439
},
"e2e_s": {
"mean": 65.2664,
"p50": 65.1682,
"max": 65.4579,
"min": 65.1135
},
"per_req_out_tok_s_e2e": {
"mean": 15.6896,
"p50": 15.7139,
"max": 15.7264,
"min": 15.6436
},
"per_req_decode_tok_s": {
"mean": 19.845,
"p50": 19.7319,
"max": 22.7938,
"min": 16.7391
},
"spec_accept_length_mean": null,
"retractions_total": 0,
"cache_hit_from_logs": {
"prefill_batches": 18,
"new_tokens": 262144,
"cached_tokens": 0,
"hit_rate": 0.0
}
}

View File

@ -0,0 +1,26 @@
#!/bin/bash
# bench_8k.sh — i8k/o1k/c16 bench rounds for ONE scheme (run via nohup on 60.8)
# usage: bash bench_8k.sh <TAG> <POOL_BASE>
# Rounds: warm (nreq16, discard) + r1/r2 (nreq32, measured), fresh corpus window each.
set -u
TAG=$1
POOL=$2
mkdir -p /root/bench_logs
run_round() {
local name=$1 nreq=$2 pool=$3 rid=$4
echo "=== [$TAG] round $name nreq=$nreq pool=$pool rid=$rid start=$(date -Is) ==="
python3 /root/bench_corpus.py --corpus /root/corpus_ids.json \
--input-len 8192 --output-len 1024 --concurrency 16 \
--num-requests "$nreq" --run-id "$rid" --pool-override "$pool" \
--shared-frac 0 --container glm53-nvfp4 \
> "/root/bench_logs/${TAG}_${name}.log" 2>&1
echo "=== [$TAG] round $name rc=$? end=$(date -Is) ==="
grep -c '"ok"' "/root/bench_logs/${TAG}_${name}.log" >/dev/null 2>&1
tail -3 "/root/bench_logs/${TAG}_${name}.log"
}
run_round warm 16 "$POOL" 9401
run_round r1 32 $((POOL + 300000)) 9411
run_round r2 32 $((POOL + 600000)) 9412
echo "=== [$TAG] ALL DONE $(date -Is) ==="

View File

@ -0,0 +1,298 @@
#!/usr/bin/env python3
"""Real-corpus (PG19) benchmark for sglang GLM-5.3-NVFP4.
Same methodology as bench_hit90.py (input_ids direct to /generate, temp 0,
ignore_eos, stream, server-side completion_tokens counting, hit-rate verified
from scheduler logs), but prompts are token slices of REAL book text tokenized
with the served model's own tokenizer, replacing random ids.
Corpus file: JSON {"ids": [flat token ids], "books": [[start, end), ...]}
Fixed token-offset layout into the flat id array:
[0, 117968) shared prefix for 128k points (90% of 131072)
[0, 58976) shared prefix for 64k points (same region, shorter cut)
pool A [131072, +4*8*13104) 128k unique suffixes, run-ids 9301-9304
pool B [550400, +4*8*6560) 64k unique suffixes, run-ids 9305-9308
pool C [760320, 262144+524288+524288) 16k fully-unique prompts, run-ids 9311-9313
spare [2071040, end) warmup slices / re-run margin
Each run-id maps to one non-overlapping window (one window = one bench point);
re-running a point with fresh text = bump --pool-override past the spare base.
Usage:
python3 bench_corpus.py --corpus /root/corpus_ids.json --input-len 131072 \
--concurrency 1 --num-requests 8 --run-id 9301 --shared-frac 0.9
"""
import argparse
import datetime
import json
import re
import statistics
import subprocess
import time
from concurrent.futures import ThreadPoolExecutor
import requests
OUTPUT_LEN_DEFAULT = 512
CORPUS_DEFAULT = "/root/corpus_ids.json"
# fixed pool layout (see docstring)
S1_128K_RID0, S1_64K_RID0, S2_RID0 = 9301, 9305, 9311
POOL_A_BASE, POOL_A_PER = 131072, 8 * 13104 # 128k suffix windows
POOL_B_BASE = POOL_A_BASE + 4 * POOL_A_PER # 550400
POOL_B_PER = 8 * 6560 # 64k suffix windows
POOL_C_BASE = POOL_B_BASE + 4 * POOL_B_PER # 760320
POOL_C_SIZES = [16 * 16384, 32 * 16384, 32 * 16384] # cc8 / cc16 / cc32
SPARE_BASE = POOL_C_BASE + sum(POOL_C_SIZES) # 2071040
sess = requests.Session()
sess.trust_env = False # bypass any proxy env on the host
def split_lens(input_len, shared_frac):
# unique suffix = (1 - shared_frac) of the prompt, page-16 aligned
# (128k @0.9 -> 13104 unique; 16k @0.0 -> fully unique prompts)
unique = round(input_len * (1.0 - shared_frac) / 16) * 16
return input_len - unique, unique
def pool_start_for(input_len, shared_frac, run_id, override):
if override is not None:
return override
if shared_frac > 0:
if input_len == 131072:
idx = run_id - S1_128K_RID0
if not 0 <= idx < 4:
sys_exit_bad_runid(run_id, "128k points use run-ids 9301-9304")
return POOL_A_BASE + idx * POOL_A_PER
if input_len == 65536:
idx = run_id - S1_64K_RID0
if not 0 <= idx < 4:
sys_exit_bad_runid(run_id, "64k points use run-ids 9305-9308")
return POOL_B_BASE + idx * POOL_B_PER
sys_exit_bad_runid(run_id, "shared-frac>0 supports 131072/65536 only")
idx = run_id - S2_RID0
if not 0 <= idx < 3:
sys_exit_bad_runid(run_id, "16k unique points use run-ids 9311-9313")
return POOL_C_BASE + sum(POOL_C_SIZES[:idx])
def sys_exit_bad_runid(run_id, msg):
raise SystemExit(f"[pool] run-id {run_id} outside expected set: {msg}")
def build_prompts(ids, shared_len, unique_len, num_requests, pool_start):
if shared_len:
shared = ids[0:shared_len]
else:
shared = []
end = pool_start + num_requests * unique_len
if end > len(ids):
raise SystemExit(
f"[pool] window [{pool_start}, {end}) exceeds corpus ({len(ids)} ids); "
f"use --pool-override or a larger corpus")
prompts = []
for i in range(num_requests):
s = pool_start + i * unique_len
prompts.append(shared + ids[s:s + unique_len])
return shared, prompts, (pool_start, end)
def warmup(url, ids, shared_len):
# primes the radix cache with the shared prefix (same role as in bench_hit90);
# warm slice comes from the spare region so it never collides with a pool window
if len(ids) >= SPARE_BASE + 64:
warm_slice = ids[SPARE_BASE:SPARE_BASE + 64]
else:
warm_slice = ids[-64:]
payload = {
"input_ids": ids[0:shared_len] + warm_slice if shared_len else warm_slice,
"sampling_params": {"max_new_tokens": 8, "temperature": 0.0, "ignore_eos": True},
}
t0 = time.perf_counter()
r = sess.post(url, json=payload, timeout=1800)
dt = time.perf_counter() - t0
print(f"[warmup] http={r.status_code} wall={dt:.2f}s", flush=True)
def bench_one(url, prompt, output_len, idx, results):
payload = {
"input_ids": prompt,
"sampling_params": {"max_new_tokens": output_len, "temperature": 0.0, "ignore_eos": True},
"stream": True,
}
rec = {"idx": idx}
t0 = time.perf_counter()
first = last = None
first_ct = None
final_meta = None
max_ct = 0
try:
with sess.post(url, json=payload, stream=True, timeout=3600) as resp:
for raw in resp.iter_lines():
if not raw or not raw.startswith(b"data:"):
continue
body = raw[5:].strip()
if body == b"[DONE]":
continue
now = time.perf_counter()
try:
d = json.loads(body)
except Exception:
continue
mi = d.get("meta_info") or {}
ct = mi.get("completion_tokens") or 0
if ct:
max_ct = max(max_ct, ct)
if first is None:
first = now
first_ct = ct
last = now
if mi.get("finish_reason"):
final_meta = mi
t_end = time.perf_counter()
n_out = max_ct or ((final_meta or {}).get("completion_tokens") or 0)
decode_span = (last - first) if (first and last and last > first) else 0.0
rec.update(
ok=n_out > 0,
ttft=(first - t0) if first else None,
e2e=t_end - t0,
n_out=n_out,
first_chunk_tokens=first_ct,
decode_span=decode_span,
tpot=(decode_span / (n_out - 1)) if n_out > 1 else None,
per_req_decode_tok_s=(n_out / decode_span) if decode_span > 0 else None,
retractions=(final_meta or {}).get("num_retractions"),
spec_accept_len=(final_meta or {}).get("spec_accept_length"),
)
except Exception as e:
rec.update(ok=False, error=repr(e))
results[idx] = rec
def verify_hit_rate(container, t_start, t_end):
def rfc3339(epoch):
return (datetime.datetime.fromtimestamp(epoch, tz=datetime.timezone.utc)
.isoformat().replace("+00:00", "Z"))
try:
# No margin before t_start: warmup's prefill lines end strictly before it,
# and catching them would deflate the measured hit rate.
p = subprocess.run(
["docker", "logs", container, "--since", rfc3339(t_start), "--until", rfc3339(t_end + 2)],
capture_output=True, text=True, timeout=120)
text = p.stdout + p.stderr
except Exception as e:
return {"error": repr(e)}
pat = re.compile(r"#new-token: (\d+).*?#cached-token: (\d+)")
n_batches = new_tok = cached_tok = 0
for line in text.splitlines():
if "TP0]" not in line or "Prefill batch" not in line:
continue
m = pat.search(line)
if m:
n_batches += 1
new_tok += int(m.group(1))
cached_tok += int(m.group(2))
total = new_tok + cached_tok
return {
"prefill_batches": n_batches,
"new_tokens": new_tok,
"cached_tokens": cached_tok,
"hit_rate": round(cached_tok / total, 4) if total else None,
}
def stats(vals):
vals = [v for v in vals if v is not None]
if not vals:
return {"mean": None, "p50": None, "max": None, "min": None}
s = sorted(vals)
return {
"mean": round(statistics.fmean(vals), 4),
"p50": round(s[len(s) // 2], 4),
"max": round(s[-1], 4),
"min": round(s[0], 4),
}
def main():
ap = argparse.ArgumentParser()
ap.add_argument("--concurrency", type=int, required=True)
ap.add_argument("--num-requests", type=int, required=True)
ap.add_argument("--run-id", type=int, required=True)
ap.add_argument("--input-len", type=int, required=True, help="16384 / 65536 / 131072")
ap.add_argument("--output-len", type=int, default=OUTPUT_LEN_DEFAULT)
ap.add_argument("--shared-frac", type=float, default=0.9)
ap.add_argument("--corpus", default=CORPUS_DEFAULT)
ap.add_argument("--pool-override", type=int, default=None,
help="explicit corpus offset for the unique-suffix window (re-runs)")
ap.add_argument("--url", default="http://127.0.0.1:30000/generate")
ap.add_argument("--container", default="glm53-nvfp4")
args = ap.parse_args()
with open(args.corpus) as f:
corpus = json.load(f)
ids = corpus["ids"]
shared_len, unique_len = split_lens(args.input_len, args.shared_frac)
pool_start = pool_start_for(args.input_len, args.shared_frac, args.run_id, args.pool_override)
shared, prompts, window = build_prompts(ids, shared_len, unique_len, args.num_requests, pool_start)
print(f"[pool] window={window} shared_len={shared_len} unique_len={unique_len} "
f"corpus_total={len(ids)}", flush=True)
warmup(args.url, ids, shared_len)
results = {}
t_start = time.time()
t0 = time.perf_counter()
with ThreadPoolExecutor(max_workers=args.concurrency) as ex:
futs = [ex.submit(bench_one, args.url, p, args.output_len, i, results)
for i, p in enumerate(prompts)]
for f in futs:
f.result()
wall = time.perf_counter() - t0
t_end = time.time()
hit = verify_hit_rate(args.container, t_start, t_end)
ok = [r for r in results.values() if r.get("ok")]
n_out_total = sum(r["n_out"] for r in ok)
out_tps = [r["n_out"] / r["e2e"] for r in ok if r.get("e2e")]
ttft = stats([r.get("ttft") for r in ok])
tpot = stats([r.get("tpot") for r in ok])
e2e = stats([r.get("e2e") for r in ok])
dec = stats([r.get("per_req_decode_tok_s") for r in ok])
spec = [v for v in (r.get("spec_accept_len") for r in ok) if v is not None]
retr = sum(r.get("retractions") or 0 for r in ok)
summary = {
"concurrency": args.concurrency,
"num_requests": args.num_requests,
"run_id": args.run_id,
"corpus_window": {"start": window[0], "end": window[1]},
"ok": len(ok),
"failed": args.num_requests - len(ok),
"wall_s": round(wall, 2),
"input_len": args.input_len,
"shared_len": shared_len,
"unique_len": unique_len,
"output_len": args.output_len,
"output_tokens_total": n_out_total,
"output_throughput_tok_s": round(n_out_total / wall, 2) if wall else None,
"input_throughput_tok_s": round(args.input_len * len(ok) / wall, 2) if wall else None,
"ttft_s": ttft,
"tpot_s": tpot,
"e2e_s": e2e,
"per_req_out_tok_s_e2e": stats(out_tps),
"per_req_decode_tok_s": dec,
"spec_accept_length_mean": round(statistics.fmean(spec), 3) if spec else None,
"retractions_total": retr,
"cache_hit_from_logs": hit,
}
print("\n===== SUMMARY =====")
print(json.dumps(summary, indent=2), flush=True)
if __name__ == "__main__":
main()

View File

@ -0,0 +1,58 @@
#!/bin/bash
# deploy_glm53_605.sh — GLM-5.3-NVFP4 TP8 + EAGLE for 174.1.60.5 (team service)
# Target scenario: cc1-2, 64k/128k input 90% cache hit, single request up to 256k,
# output-throughput priority. Clean deploy: rm old container, wait VRAM drain, run.
#
# Env overrides:
# MEMFRAC (0.90) STEPS (4) TOPK (1) DRAFT (5) CTXLEN (270336)
# CHUNK (8192) MAXPRE (16384) RESTART (no|yes) EXTRA ("")
# Usage:
# bash deploy_glm53_605.sh # production 4/1/5 @ 0.90
# STEPS=5 DRAFT=6 bash deploy_glm53_605.sh # tuning round
# CHUNK=16384 bash deploy_glm53_605.sh # prefill tuning round
# RESTART=yes bash deploy_glm53_605.sh # production finalize
set -e
MEMFRAC=${MEMFRAC:-0.90}
STEPS=${STEPS:-4}
TOPK=${TOPK:-1}
DRAFT=${DRAFT:-5}
CTXLEN=${CTXLEN:-270336}
CHUNK=${CHUNK:-8192}
MAXPRE=${MAXPRE:-16384}
EXTRA=${EXTRA:-}
if [ "$RESTART" = "yes" ]; then RP="--restart unless-stopped"; else RP="--restart no"; fi
echo "[deploy] removing old container (if any)"
# rm -f times out on big GPU containers on this daemon; retry until really gone
for i in $(seq 1 45); do
CID=$(docker ps -a --filter name=glm53-nvfp4 -q)
[ -z "$CID" ] && break
docker rm -f glm53-nvfp4 >/dev/null 2>&1 || true
sleep 2
done
echo "[deploy] waiting for VRAM drain"
for i in $(seq 1 45); do
used=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | awk '{s+=$1} END {print s}')
[ "$used" -lt 2000 ] && break
sleep 2
done
used=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | awk '{s+=$1} END {print s}')
echo "[deploy] VRAM now: ${used} MiB total"
echo "[deploy] starting: TP8 EAGLE ${STEPS}/${TOPK}/${DRAFT} memfrac=${MEMFRAC} extra='${EXTRA}'"
docker run -d --name glm53-nvfp4 --gpus all --shm-size 64g --ipc=host $RP \
-p 30000:30000 -v /data/hf_models:/data/hf_models \
lmsysorg/sglang:nightly-dev-20260828-daf63171 \
python3 -m sglang.launch_server \
--model-path /data/hf_models/GLM-5.3-NVFP4 --tp 8 \
--mem-fraction-static $MEMFRAC --max-running-requests 16 \
--chunked-prefill-size $CHUNK --max-prefill-tokens $MAXPRE \
--disable-shared-experts-fusion --moe-runner-backend flashinfer_cutlass --disable-flashinfer-autotune \
--speculative-algorithm EAGLE --speculative-num-steps $STEPS --speculative-eagle-topk $TOPK --speculative-num-draft-tokens $DRAFT \
--kv-cache-dtype fp8_e4m3 --enable-hierarchical-cache --hicache-ratio 3 \
--cuda-graph-max-bs-decode 8 --cuda-graph-bs-decode 1 2 3 4 6 8 --cuda-graph-max-bs-prefill 8 \
--context-length $CTXLEN --reasoning-parser glm45 --tool-call-parser glm47 \
--host 0.0.0.0 --port 30000 $EXTRA
echo "[deploy] container started; poll: docker logs -f glm53-nvfp4"

View File

@ -0,0 +1,62 @@
#!/bin/bash
# ============================================================
# GLM-5.3 最优部署方案6000D 8卡TP4 PP2 + IndexCache freq=4
# 2026-09-07
#
# - 基线配置TP4 PP2 + cps16k + mem0.85131.6 tok/s 吞吐基线)
# - IndexCacheindex_topk_freq=4层轴索引复用省 75% indexer
# 16K 场景无损失128K 长上下文并发 1.35-1.47× 提速
# - 禁 radix cache禁投机解码PP2 与投机框架不兼容,已实测)
#
# 用法bash deploy_glm53_optimal.sh
# 60.7 试用版:增加显存排空等待 + 就绪等待加长至 20 分钟)
# ============================================================
set -uo pipefail
CONTAINER="glm53-nvfp4"
IMAGE="lmsysorg/sglang:nightly-dev-20260828-daf63171"
MODEL="/data/hf_models/GLM-5.3-NVFP4"
PORT=30000
TP=4; PP=2; MEM=0.85; MRR=48; CPS=16384
docker rm -f ${CONTAINER} 2>/dev/null || true
# docker rm -f 后显存释放滞后数分钟,不等会把新容器 KV 池压小
for i in $(seq 1 60); do
used=$(nvidia-smi --query-gpu=memory.used --format=csv,noheader,nounits | awk '{s+=$1} END {print s}')
if [ "$used" -lt 1500 ]; then echo "[deploy] drained: ${used} MiB"; break; fi
echo "[deploy] drain wait ${i}: ${used} MiB"
sleep 10
done
docker run -d --name ${CONTAINER} --gpus all --shm-size 64g --ipc=host \
--restart unless-stopped \
-p ${PORT}:${PORT} \
-v /data/hf_models:/data/hf_models \
${IMAGE} \
python3 -m sglang.launch_server \
--model-path ${MODEL} \
--tp-size ${TP} --pp-size ${PP} \
--mem-fraction-static ${MEM} \
--max-running-requests ${MRR} \
--disable-radix-cache \
--disable-shared-experts-fusion \
--moe-runner-backend flashinfer_cutlass \
--disable-flashinfer-autotune \
--disable-custom-all-reduce \
--chunked-prefill-size ${CPS} \
--host 0.0.0.0 --port ${PORT} \
--json-model-override-args '{"index_topk_freq": 4}'
echo "容器已启动,等待就绪..."
for i in $(seq 1 120); do
code=$(curl -s -o /dev/null -w '%{http_code}' http://localhost:${PORT}/health 2>/dev/null)
if [ "$code" = "200" ]; then
echo "READY after ~$((i*10))s"
docker ps --filter name=${CONTAINER} --format '{{.Names}} {{.Status}}'
echo "override args: $(docker inspect ${CONTAINER} --format '{{.Config.Cmd}}' | grep -o 'index_topk_freq[^,}]*')"
exit 0
fi
sleep 10
done
echo "TIMEOUT"; exit 1