180 Commits

Author SHA1 Message Date
yy-fighting
97c208cfdb b300eq: TP2PP4 MRR64 retest (pass-1 + ordered v2 + MRR48 control) — c64 gains +27%/+19%, TTFT collapse, -14% 1K c32 config cost attributed
- deploy_glm53_pp4_mrr64.sh: byte-diff vs original = MRR 48->64 only; decode graph
  stack-default bs<=256 (no graph change needed, arm never fell off graph)
- pass-1 + v2 (fresh instance, retraction-prone 16K c64 ordered last): all 6 points
  aligned <=1.8% -> post-retraction contamination hypothesis disproved; anchor
  regression is reproducible MRR64 behavior
- MRR48 fresh-instance control (4.1 c32 = 266.42): decomposes -17% anchor into
  -3.3% instance freshness + -14.4% MRR64 config cost (per-req TPOT p50 87->102ms,
  three instances, mechanism unidentified)
- 16K c64: pool-capped ~59 active, 3 retractions both runs (deterministic);
  out -5% vs MRR48 queue-mode but TTFT P95 135->89s
- report/README/provenance/CURRENT.md overwritten in place per user instruction;
  feishu wiki revision 13
2026-09-10 21:29:05 +08:00
yy-fighting
2370d33c73 glm53 dsv4-migration A-baseline (TP8+EAGLE) campaign: image negative, autotune split-negative, IPC@TP8 real win +4.6~9.1% but net-neutral vs production (patch/image coupling); eager-trap profile + day-drift finding (60.1, 09-10) 2026-09-10 19:04:28 +08:00
yy-fighting
c5d91ceafe b300-equivalent matrix: E7b high-concurrency retest (MRR64 + decode-graph buckets 1-64) - graph-drop cliff fixed, report numbers overwritten in place
- deploy_glm53_e7b_hicc.sh (md5 165db732): only delta vs 607_exp = MRR 16->64 + cuda-graph-bs-decode 1..64; KV pool 276,480 unchanged, avail 6.11GB after capture
- 10 retest points (16K/4.1/4.2 at c8/16/32/64) all OK, hit=0.0 (fresh container = recycled windows virgin again), 0 retraction, QG 7/7
- verdicts: 16K output 92.1/97.6/99.6 (+18~33% vs initial, still TP2PP4-dominated, prefill wall ~100 plateau); 4.1 c32/64 229/272 (gap narrowed to 1.2x); 4.2 402/676/826.5 - E7b wins ALL cc tiers, c64 826.5 tok/s = machine-wide best output (+72% vs TP2PP4 482), TTFT 12.18s / TPOT 85.1ms; c8 anchors within +-3% prove no env drift
- REPORT.md + Feishu A7V3wZTQeifCB4krdi6cA834nW9 overwritten in place (user directive: no appended chapter); initial MRR16 run archived as baseline in results/e7b/
- provenance.md: e7b64 VRAM (idle 79.3k, peak 83,627 MiB), second in-service restore verified (fired up/health 200/16K+C4 spot/KV pool 647,040 identical)
2026-09-10 18:52:54 +08:00
yy-fighting
d063bd63c9 b300-equivalent matrix: add provenance.md (gpu inventory, vram peaks, qg verdicts, in-service preserve/restore record) 2026-09-10 17:14:38 +08:00
yy-fighting
21fcca5d16 b300-equivalent matrix: dual-plan (TP2PP4-D vs E7b) full B300 scenario replication on 60.8 - 41 valid points, divide at C=8, E7b usable window <=C8 (MRR16+graph-drop), boundary 256K/512K/896K TP2PP4-only, 4-5x absolute gap vs B300 narrowing to ~2x at boundary prefill; hicache host-layer cold-cache pitfall documented; in-service container preserved-renamed-restored and verified (09-10) 2026-09-10 17:13:50 +08:00
yy-fighting
e3476aff86 glm53 dsv4-migration bench: autotune+latest-image winner (+1.2~3.2% all 5 pts, A/B/A confirmed), PCIe-IPC pack negative (-0.4~-3.5%), page-mark kernel N/A for GLM (60.1, 09-09) 2026-09-10 03:13:02 +08:00
yy-fighting
f15003d2f5 hit90 bench: add aggregated summaries (41 scored points) + md5 manifest (70 raw logs on 60.8)
原始 .log 按仓库惯例留 60.8:/root/bench_logs/hit90exp/(md5s.txt 可核验),
仓库入聚合 SUMMARY 数据:A1/A3/A6 全套主扫+cap+决赛、A2 sanity、种子遍。
2026-09-09 17:11:37 +08:00
yy-fighting
ad1853f49b hit90 scenario bench: TP4PP2-nomtp@0.90 winner, DP-attention/DCP verdicts (60.8 serving, 60.5 v3 delivered, 09-09)
- 冠军 A6 TP4PP2-nomtp@0.90(池 647,040):hit90 out cc1-4=28.8/47.5/63.9/76.2、cc8/16=106.4/125.7(vs TP2PP4 基线 cc4+14%/cc8+15%),并发独立 128k 文档 4 条、512k 单条 ✓、质量门 7/7;60.8 在役(实测实例切 unless-stopped)
- 判决:DP attention 容量负收益判死(非 MoE 权重按 attn 组复制致 dp4 池 96,448/rank + EAGLE×DP 两层崩溃);DCP 对 DSA 静默算错禁用(dsa_backend 零引用无 guard);MTP@128k accept 2.07 判负(vs EAGLE 3.46 存活但池 276k 过不了容量门槛)
- 交付:60.5 v3 脚本(只交未执行)+ 补丁束(60.8:/root,md5 6922e534)+ profile tp4pp2_hicache.env + CURRENT.md 更新
- 归档:70 文件全量日志(含 DP 臂失败记录、A3 容器日志)+ 7 脚本 + README(双口径协议:hit90 主扫 + distinct-doc 容量探针)
2026-09-09 17:11:04 +08:00
yy-fighting
0d929948a5 60.4 redeploy: TP8+EAGLE+custom-AR 1stage (E7b recipe) replaces TP2PP4 (09-09)
CURRENT.md 60.4 行更新为 E7b 配方在役(EAGLE 4/1/5/mem0.90/MRR16/hicache3/图桶1-8,
CAR 补丁 8 rank 激活,质量门 7/7)。deploy_glm53_604_exp.sh 入库(=607 实验版逐字拷贝,
仅改 header;补丁 md5 与 car_patch/ 归档一致)。

注:本提交同时带入并行会话对 CURRENT.md 60.5(v2 容量脚本交付)与 60.8(TP2PP4-nomtp
容量口径在役)两行的未提交更新,内容为当日实测状态。
2026-09-09 15:55:43 +08:00
yy-fighting
92517e87f0 128k low-cc capacity topology: TP2PP4-nomtp winner (60.8 serving, 60.5 v2 delivered, 09-09)
诊断 60.5 排队根因:KV 池 276,864 token 纯容量算术(MLA KV 按 PP 分层切分,
池≈PP 度数倍增;TP8 调参无解,fp8 断言封顶 314,944 < c3 需求 394,752)。
四臂对拍(i128k/o512 冷缓存 cc1-4,PG19 真实语料同窗口映射):
- A-mirror(60.5 现状):c3 起容量排队,cc4 TTFT p50 119.6s,痛点复现
- r37(TP4PP2+MTP):池 384,960,c3 起排队
- B'(TP4PP2 nomtp):池 589,696,c4 零排队
- TP2PP4 nomtp 池 909,632 优胜:cc4 in 5907 / out 23.1 tok/s,
  TTFT p50 全场最优(cc1 15.5s),e2e 88.8s
优胜者验证全过:质量门 7/7;512k 单条 TTFT 88.4s;900k 单条可完成
(单条上限 ~909k 实证);并发上限 c6(cc7 第 7 条起排队);hit90 cc4
输入吞吐 17,625 tok/s、TTFT 7.3s。
60.8 末态:TP2PP4 容器在役(restart=unless-stopped)。
60.5 交付:deploy_glm53_605_v2.sh(一键部署+前置检查,原脚本=回滚路径)。
资产:arm_runner.sh / val_runner.sh / all_summaries.json(26 点全量) /
profile glm53_nvfp4_pro6000_sglang_tp2pp4_hicache.env / CURRENT.md 两行更新。
2026-09-09 12:43:11 +08:00
yy-fighting
0c89fd4fbe r37 addendum: nomtp sweep raw logs (primary evidence, force-added past *.log ignore) 2026-09-09 09:24:43 +08:00
yy-fighting
4d6dac010c r37 addendum: B'(nomtp) vs PP+MTP same-profile input/output throughput comparison (i16k/o512 random-ids cc8-64)
同栈同脚本 nomtp 对拍:cc8 打平(每token 41.0 vs 41.6ms 拐点实证)、cc16 mtp +6%+TTFT-32%、cc>=32 B' 反超 0.9-5.4%(TPOT 好2x);图修复把 r36 1.81x 惨败收敛到 ±6% 水平线;生产 r37 mtp 已恢复核验
2026-09-09 09:24:25 +08:00
yy-fighting
911a9a2fb0 r37 addendum: i16k/o512 cc 8-64 sweep (nreq=2cc) on serving container; throughput saturates 130-135 tok/s at cc16+, accept stable; rc=1 root-caused (missing bs_results dir, never existed) 2026-09-09 08:29:51 +08:00
yy-fighting
a3b8f1b2ab r37: fix PP+MTP verify CUDA graph (pre-planned path missing pp_proxy fill); beats A16 by 40% on killer, now serving on 60.8 2026-09-09 01:24:29 +08:00
yy-fighting
ffda226f5a PP+MTP r36 de-GLOO: fix + verdict (PP+MTP loses i8k 1.81x to B', graph mode correctness-broken, A16 restored) 2026-09-08 22:21:20 +08:00
yy-fighting
123023b6ae feat(pro6000/GLM-5.3): i8k/o1k/c16 三方案对比压测入库(A/A16/B 轮换实测 + 60.8 转 A16 在役 + CURRENT.md 台账修正 60.2/60.3/60.6/60.8) 2026-09-08 19:05:10 +08:00
yy-fighting
5c749cda03 feat(pro6000/GLM-5.3): 方案 D/E/F 部署资产入库(TP2PP4 生产配方 / TP8+DFlash2 / PD 分离四角色链)
- scripts: 11 个服务器原样脚本入库(md5 对照表更新至 README);D=60.1 生产原样配方、
  E=v5 DFlash 底稿、F=PD 链四角色部署+launch+双场景压测驱动
- profiles: 新增 6 个 .env(D/E 单机 + F 四角色,均带镜像 digest
  sha256:28e0d260…,对齐 kimi3 PD 多角色先例)
- deploy/PD_CHAIN.md: 方案 F 编排手册(启动顺序 mc-master→prefill→decode→router、
  基础设施依赖表、质量门口径、拆链恢复、÷2 单机等效判决)
- platforms/patches/pro6000/glm53_pd_chain/: sglang 补丁树 vs 镜像原版 11 文件
  unified diff 快照——宿主树无 .git,此为唯一版本记录(DFlash+PP+PD 解锁全集)
- deploy/manifests/: GLM-5.3-NVFP4(47分片)/GLM-5.3-DFlash2(单分片) 权重 md5 清单
- deploy/CURRENT.md: 全集群现役状态页(2026-09-08 八机实测)
- deploy/verify_profile.sh: 防漂移核验工具(digest+参数 token 比对+端口/health,
  已在 60.1 生产容器实测 PASS)
2026-09-08 16:36:15 +08:00
yy-fighting
b3165a1d3c feat(pro6000/GLM-5.3): 部署方案入库(TP8+EAGLE 生产标准 / 场景二高并发变体 / TP4PP2+IndexCache / E7b CAR 实验补丁 + deploy profiles)
- deploy_glm53_605.sh:配置 A 生产标准(60.5 在役,全 8 台 md5 fcd9109b 一致),
  支持 MEMFRAC/STEPS/TOPK/DRAFT/CHUNK/EXTRA/RESTART 调参;场景二高并发变体
  仅改 mrr32 + decode 图 bs{4,8,12,16}(KV 池 16.4 驻留上限)
- deploy_glm53_optimal(_s1).sh:配置 B TP4PP2+IndexCache(场景二最优 +41~79%;
  s1 形态唯一差异 radix-on)
- deploy_glm53_607_exp.sh + car_patch/:E7b custom-AR 1stage 补丁(cc1 decode
  每步 -14~-16%,实验性仅 cc1-2 验证;补丁文件与 60.7:/root/patches md5 一致)
- deploy/profiles/pro6000/:两个标准 profile(sskj.deploy 可消费),关键踩坑
  与场景二变体、parser 缺口均在注释中标注
2026-09-08 11:38:06 +08:00
yy-fighting
f0ab17c561 feat(pro6000/GLM-5.3): 双场景压测标准入库(bench_corpus 真实语料工具链 + run-id 窗口纪律 + 质量门禁 + 语料构建链)
- 场景一:128k/64k 输入、o512、cc1-4、90% 前缀命中(run-id 9301-9308)
- 场景二:16k 全独立输入、o512、cc8/16/32(run-id 9311-9313)
- 口径:输出吞吐取服务端 completion_tokens;命中率仅从 TP0 Prefill 日志核验;
  配置对比用 TPOT×accept;噪声带 ±8%;重跑必须 --pool-override 换新鲜窗口
- 基线(2026-09-07 真实语料)与完整纪律见实验目录 README
2026-09-08 11:37:54 +08:00
9cdbc1fd22 Revert "feat(p800): parameterize GLM5.2 single-node tuning"
This reverts commit ba20973beaed07e6f5ccf1be10e9a78e31d614e0.
2026-08-20 07:33:38 +00:00
ba20973bea feat(p800): parameterize GLM5.2 single-node tuning 2026-08-20 07:27:52 +00:00
eeec56c2cf update(P800/GLM5.2): 单节点部署测试脚本 2026-08-19 02:24:17 +00:00
shishi
987f1db4b0 feat(pro6000): Kimi-K3 DP=2 部署(方案 B:两个独立 TP32×EP32 实例 + router 负载均衡)
- 实例 A profile 加 --disable-radix-cache(bench 测量纯净)
- 新增实例 B profile(kimi3_pro6000_sglang_tp32ep32_instB,.1-.4)
- 新增 deploy_dp2.sh 编排脚本(A + B + router --worker-urls)
- 新增 README
2026-08-11 17:56:14 +08:00
shishi
c69831f258 feat(pd): PD 长上下文 adaptive concurrency bench(SLO 方案 A,并发 +16)
- matrix.json: 3 个 shape(64k/128, 16k/1k, 1k/4k)
- run_adaptive_concurrency_pd.sh: PD 专用 adaptive 脚本
  - 复用 adaptive_bench_lib.sh(并发搜索/SLO 停止/OOM 检测/完整产物)
  - server 生命周期函数 no-op(PD 服务常驻,不启停)
  - 加 --flush-cache(配合 --disable-radix-cache 测纯净 TTFT/TPOT)
  - 离线环境变量 + 本地 tokenizer(避免 HF 在线下载)
- adaptive_config.env: SEARCH_ADDEND=16 +16 递增,上限 64,回退 8/1
- config.env: 新增 get_ttft_slo_ms 分层 SLO(1k->4s, 16k->15s, 64k->30s)
2026-08-11 14:45:36 +08:00
shishi
fe375e1307 docs(pd): 补充运维一键部署完整步骤(干净环境从零到跑通)
- experiments README 新增「运维一键部署」章节:6 步完整流程
  (免密/文件分发/仓库 clone/拉镜像/启动/验证)+ 停止重启 + 文件来源
- docs/KIMI_K3_DEPLOY.md 附录 B 新增 B.2 运维准备,编号顺延 B.3-B.6
- 修复 scp 分发命令(先建父目录)
2026-08-11 11:55:39 +08:00
shishi
da0e1b4372 fix(pd): config.env 加 DOCKER_CLIENT_IMAGE(bench 用 docker client 复用 kimi-k3 镜像) 2026-08-11 11:19:42 +08:00
shishi
04dfa31583 fix(pd): PATCH_MOUNTS 补充 flashkda wheel 挂载(BOOTSTRAP 需要 /flash_kda-*.whl) 2026-08-11 10:50:43 +08:00
shishi
d72dbff689 feat(pro6000): Kimi-K3 PD 分离部署(MoonCake RDMA)- 8 节点 P/D 双 profile + deploy_pd.sh 编排 + 文档
- 新增 P 组(prefill)/D 组(decode) deploy profiles(mooncake RDMA + 计算网 mlx5_0~3)
- 新增 experiments/pro6000/kimi3_pro6000_pd_rdma/ 编排脚本(mc-master+P+D+router)
- 关键修复: 必须关闭 PYTORCH_CUDA_ALLOC_CONF=expandable_segments
  (mooncake RDMA 注册 expandable GPU 段报 Bad address,GitHub #2511)
- 容器内升级 mooncake 0.3.12.post1(含 dmabuf 修复 #2035)
- docs/KIMI_K3_DEPLOY.md 新增附录 B,README 更新实验索引
2026-08-11 10:36:09 +08:00
shishi
3bd04698bb feat(pro6000): Kimi-K3 TP32×EP32 部署 profile、sm_120 补丁与运维手册 - 4 节点 RoCE 部署 + bench 实验 2026-08-10 11:04:24 +08:00
shishi
d9a2e3b5e5 fix(910c/glm52): 镜像与 client 修正 - 部署在 910c.2 的 GLM5.2-tuned 镜像
- profile DOCKER_IMAGE 改回 local/vllm-ascend:0.23-a3-20260718-sglang
  (910c.1 的 glm5.2-a3-openeuler 缺 expert_map_manager 模块无法启动)
- config.env DOCKER_CLIENT_IMAGE 指向 910c.2 本地 tuned 镜像(带 bench_serving)
- 实测: 910c.2 TP8/DP2 smoke 40/40, TTFT 1062ms, TPOT 51ms
2026-08-03 17:22:04 +08:00
shishi
f33c5f1d3d fix(deploy): DP_FLAG 未达 BOOTSTRAP 启动命令 + bench 的 model/tokenizer 分离
- profile.py: DP_FLAG 追加到 LAUNCH_ARGS 与 BOOTSTRAP(BOOTSTRAP 的 ${LAUNCH_ARGS}
  引用在 env 解析阶段已内联,必须直接 append 到 BOOTSTRAP 末尾),并改为用
  rendered 作模板展开源;修复 910c TP4/DP4 启动退化为 TP4 单 DP 布局导致
  专家权重不分片 OOM(61.3GB/die) 的问题(顺带修复 p800/pro6000 同类隐患)
- cli.py/runner.py: bench 的 API model 名改用 SERVED_MODEL_NAME(vLLM 严格校验),
  tokenizer 独立用 MODEL_PATH 路径并 --tokenizer 透传、docker client 挂载;
  修复 910c/vLLM 场景 404 Not Found
- dsv4 910c profile 恢复标准加载参数(prefetch+multithread);config.env per-TP
  max-model-len 默认改为已验证值(32768/65536/131072)
- 实测: 910c TP4/DP4 16-worker 布局启动健康, sskj.bench smoke 40/40
  (TTFT 1667ms, TPOT 35.4ms)
2026-08-03 16:41:29 +08:00
shishi
ea8302561e fix(910c): bench client 可用化 - 本地 vllm-ascend-sglang 镜像 + torch_npu 自动加载禁用
- runner.py docker client 注入 TORCH_DEVICE_BACKEND_AUTOLOAD=0(bench 纯 HTTP
  client 不需要 NPU backend,跳过 torch_npu 加载失败)
- dsv4/glm52 config.env: DOCKER_CLIENT_IMAGE 指向本地定制镜像
  (local/vllm-ascend:0.23-a3-dsv4-sglang,内置 sglang 0.5.2 bench_serving),
  USE_DOCKER_CLIENT=1,sskj.bench/run_bench.sh 的 docker client 分支可用
- run_bench.sh docker client 分支同步注入 AUTOLOAD=0
- glm52 profile 修正 DOCKER_IMAGE 为本机存在的 glm5.2-a3-openeuler
2026-08-03 15:43:45 +08:00
shishi
6ba04325d3 feat(910c): 部署解耦 - vLLM-Ascend profile 与 deploy 层接管服务启停
- deploy/profiles/910c/ 新增 dsv4/glm52 两个 vLLM-Ascend profile
  (JSON 参数经 BOOTSTRAP base64 注入避开镜像 entrypoint 转义;per-TP 参数
  由 start 脚本导出后经 profile 模板展开)
- dsv4/glm52 start_vllm_docker.sh 改为 deploy 薄包装(保留 sg docker 重入与
  per-TP 覆盖),新增 stop_vllm_docker.sh
- run_bench.sh / run_adaptive_concurrency*.sh 的 stop/build_server_args 改走
  deploy_stop/deploy_render_args
- runtime.py 修复单节点 dry-run 未跳过健康检查的 bug
- ops/README.md 补 910c 章节;.gitignore 补 910c ops_ 输出规则
- 附带入库 910c adaptive 汇总结果
2026-08-03 15:36:11 +08:00
Zhiyi Hong
9acf9fdfdb feat(pro6000): 部署/测试解耦 - deploy 层支持多节点与 vLLM,新增 6 个 profile
- sskj.deploy runtime 支持 NODE_HOSTS 多节点编排(ssh 分发/本地 rank/LOCAL_NODE_RANK)
  与 ENGINE=vllm 启动(SERVER_CMD),容器名按 rank 自动唯一
- scripts/common/deploy_cli.sh 新增 deploy_stop/status/multinode helper 与 node-rank 透传
- src/sskj/common/env.py 修复嵌套 ${VAR:-${OTHER}/path} 展开(平衡花括号扫描)
- deploy/profiles/pro6000/ 新增 6 个 profile: tp16/tp16_eagle/glm52(多节点)、
  sglang/vllm tp_dp_matrix、qwen3(单节点)
- 6 个实验 start/stop 脚本改为 deploy 薄包装,run_bench/adaptive 的 server 启停走
  deploy_render_args/deploy_start/deploy_stop,tp16 新增 matrix.json
- 首次入库 glm52_pro6000_sglang_multinode_tp16 实验目录;ops/README.md 补 pro6000 章节
- 实测通过: 单节点 dsv4 sglang/vllm 链路 + tp16 双节点启动/bench/清理
2026-08-03 15:17:41 +08:00
yy-fighting
3761d75b00 feat(ops): add unified bench/deploy layers and P800 profile 2026-08-02 15:36:28 +08:00
shishi
e885fd0dc2 feat(adaptive): support tiered per-ISL TTFT SLO via get_ttft_slo_ms()
- adaptive_bench_lib.sh: add default get_ttft_slo_ms() fallback (flat TTFT_SLO_MS),
  use it instead of hardcoded TTFT_SLO_MS in SLO comparison and logs,
  add ttft_slo_tiers_desc to run_manifest.json
- glm52_910c config.env: define tiered SLO for GLM-5.2:
  ≤2048:5000ms, ≤8192:8000ms, ≤32768:12000ms, ≤131072:20000ms, >131072:30000ms
  (~70-80% of DSv4-Pro values, since GLM-5.2 has simpler architecture)
- glm52_910c adaptive_config.env: update TTFT_SLO_MS comment noting tiered override

Backward compatible: experiments without get_ttft_slo_ms() keep flat 4000ms behavior.
2026-07-30 10:47:56 +08:00
Zhiyi Hong
e53b2c7e4c fix: JSONL parsing in run_batch, add results/ to gitignore 2026-07-29 17:51:47 +08:00
Zhiyi Hong
c839230c0b feat: add EAGLE speculative decoding experiment (dsv4_pro6000_sglang_tp16_eagle) 2026-07-29 17:51:40 +08:00
shishi
4914ff4041 fix(dsv4): disable MTP speculative decoding for fair H20 comparison
MTP (multi-token prediction) speculative decoding was enabled, giving the
910C an unfair decode throughput advantage over the H20 baseline (which
has no MTP). Disable it so the benchmark measures pure model throughput.

- config.env: add DSV4_ENABLE_MTP=0 (default off)
- start_vllm_docker.sh: only add --speculative-config when DSV4_ENABLE_MTP=1
2026-07-29 16:50:13 +08:00
shishi
455a78161b fix(910c/glm52): 对齐官方A3教程参数(修DP die分配问题,同dsv4 99a22f0)
dsv4实验发现DP副本绑定到同一组die的问题(99a22f0),glm52存在相同问题:
缺--max-num-batched-tokens和--api-server-count导致DP worker设备分配异常
(不加api-server-count时vllm为N个DP rank启动N个API server)。

对齐docs.vllm.ai GLM5.2 A3官方教程:
- 加 --max-num-batched-tokens 8192 (官方值,影响DP调度)
- 加 --api-server-count 1 (官方值,避免多API server干扰设备分配)
- 去掉 --kv-cache-dtype fp8 (官方不指定,用默认bfloat16;且此镜像fp8本就未生效)
- 保留 --trust-remote-code / --enable-expert-parallel / enable_dsa_cp (官方有)
2026-07-29 15:28:54 +08:00
shishi
99a22f05b8 fix(dsv4): align launch params with official A3 tutorial (fixes DP die allocation)
The previous params caused all 4 DP replicas to bind to the same 4 dies
(die 0-3), leaving 12 dies idle. Root cause was a combination of missing
official params + extra non-official params that interfered with DP
worker device placement.

Verified: with aligned params, TP=4 DP=4 correctly distributes 16 workers
across all 16 dies (8 cards x 2 dies), each at ~57 GB HBM (91% util).

Changes (align to docs.vllm.ai A3 tutorial):
- Add --max-num-batched-tokens 10240 (was missing; affects DP scheduling)
- Add --api-server-count 1 (was missing; without it vllm spawns N API
  servers for N DP ranks, disturbing device assignment)
- Remove --kv-cache-dtype fp8 (official uses default bfloat16)
- Remove --trust-remote-code (official doesn't use it for DSV4)
- Remove enable_dsa_cp from additional-config (official doesn't have it)
- max-model-len: per-TP caps (32768/65536/131072) -> 1048576 for all TPs
  (official uses full 1M; the caps were over-cautious)
- max-num-seqs: per-TP (128/256/256) -> 64 for all (official value)
2026-07-29 15:09:27 +08:00
shishi
c222ed98b2 feat(910c/glm52): 并行配置改为 4 4 / 8 2 / 16 1 对标H20的 2 4 / 4 2 / 8 1
A3 910C 有16 dies(8卡x2die),H20有8卡。为公平对比,TP按卡数等效:
  A3 TP=4  DP=4 (4die/副本x4) == H20 TP=2 DP=4 (2卡/副本x4)
  A3 TP=8  DP=2 (8die/副本x2) == H20 TP=4 DP=2 (4卡/副本x2)
  A3 TP=16 DP=1 (16die/副本x1) == H20 TP=8 DP=1 (8卡/副本x1)

与dsv4实验(a65849b)保持一致的配置思路。

新增TP=4 per-TP参数覆盖(expert-parallel下专家分4份+dense复制,
KV cache极紧): gpu_mem=0.97, max_model_len=4096, max_num_seqs=32
TP=4可能OOM,若发生会自动记录并跳过。

config.env: PARALLEL_CONFIGS 8 1/16 1 -> 4 4/8 2/16 1; 加TP4_变量
start_vllm_docker.sh: case $TP 加 TP=4 分支
run_adaptive_concurrency_add16.sh: case $tp 加 TP=4 分支
2026-07-29 13:44:40 +08:00
shishi
a65849b77d fix(dsv4): use all 16 dies with TP4/DP4 + TP8/DP2 + TP16/DP1
The A3 910C has 16 dies (8 cards x 2 dies/card), and vllm-ascend's
tensor-parallel-size / data-parallel-size address DIES, not cards. The
old configs (2/4, 4/2, 8/1) all used only 8 dies = 4 cards, leaving half
the node idle. Switch to configs that use all 16 dies, and add per-TP
parameter overrides since each TP has a very different per-die memory
budget.

Why the old configs were wrong:
- A3 TP=4 (4 dies) is the per-CARD equivalent of H20 TP=4 (4 cards),
  not a fair comparison. An A3 node has 2x the compute of an 8-card H20,
  so benchmarking at 8 dies understates the A3 and never exercises the
  cross-card / full-NVLink topology.
- TP*DP must equal 16 to use all dies on the A3 node.

New PARALLEL_CONFIGS (all use 16 dies):
- TP=4  DP=4: 4 dies/replica x 4 replicas, ~39 GiB weights/die
- TP=8  DP=2: 8 dies/replica x 2 replicas, ~35 GiB weights/die
- TP=16 DP=1: 16 dies/replica x 1 replica,  ~17.5 GiB weights/die

Per-TP overrides (mirrors the glm52 6c81183 pattern):
- TP4:  max_model_len=32768  max_num_seqs=128  (tight KV cache)
- TP8:  max_model_len=65536  max_num_seqs=256  (balanced)
- TP16: max_model_len=131072 max_num_seqs=256  (max KV cache)

Changes:
- config.env: PARALLEL_CONFIGS -> "4 4"/"8 2"/"16 1"; add TP4_/TP8_/TP16_
  env vars for per-TP gpu_mem_util/max_model_len/max_num_seqs.
- start_vllm_docker.sh: case "$TP" overrides the three params after
  sourcing config.env, so the actual launch args match the per-TP budget.
- run_adaptive_concurrency_add16.sh: engine_build_server_args gets the
  same case "$TP" so the recorded server_cmd.txt stays consistent with
  the real launch.
2026-07-29 13:37:26 +08:00
shishi
63ab41b65a docs: consolidate project docs (dedup, relocate, expand 910C client guide)
Project-level documentation was scattered and duplicated across README.md,
BENCHMARK_WORKFLOW.md, and docs/EXPERIMENT_GUIDE.md (directory layout +
scripts/common component table repeated 3x). Reorganize into a clear
single-source-of-truth structure.

Changes:
- README.md: drop the 6 stale changelog entries at the top (latest was
  07-21; history lives in git log). Replace the duplicated directory-
  layout + scripts/common sections with a one-line link to
  docs/EXPERIMENT_GUIDE.md. (151 -> 99 lines)
- BENCHMARK_WORKFLOW.md -> docs/BENCHMARK_WORKFLOW.md: relocate into docs/.
  Replace its duplicated Directory Layout and Quick Start/Adding sections
  with links to EXPERIMENT_GUIDE / README / NEW_PLATFORM_GUIDE; keep the
  unique parts (Rules, Naming Conventions, Final JSON Schema, Checklist).
  (394 -> 224 lines)
- docs/EXPERIMENT_GUIDE.md: now the single authority for directory layout
  + component table + experiment conventions. Add a cross-link from the
  results.json field list to BENCHMARK_WORKFLOW's full JSON Schema and
  Naming Conventions.
- docs/H200_QUICKSTART.md: deleted (outdated, repeatedly references
  removed legacy scripts; H200 usage is covered by ADAPTIVE_CONCURRENCY_USAGE
  and experiment READMEs).
- docs/DSV4_INFERENCE_COMPARISON_REPORT.md -> experiments/h200/
  dsv4_h200_vllm_mtp_vs_default/results/20260708-160349/: this is an
  experiment report, not a project doc; relocate next to its sibling
  report.md.
- envs/ASCEND_910C_ENV_SETUP.md §8: expand the vague "pip install sglang"
  note into a full sglang client image build guide -- pin sglang 0.5.2
  (not latest; >=0.5.16 deprecates bench_serving and breaks the parser),
  --no-deps minimal install loop, docker commit to a local image, with
  the exact commands used to build local/vllm-ascend:0.23-a3-dsv4-sglang.
- experiments/h200/dsv4_h200_vllm_tp2_custom_bench/README.md: fix the
  now-broken link to BENCHMARK_WORKFLOW.md (../../ -> ../../../docs/).
- .gitignore: ignore *.bak.glm52orig scratch backups.

Also includes the add16 adaptive_results produced by the dsv4 TP=4/DP=2
runs on 910c.1.
2026-07-29 11:52:42 +08:00
shishi
70c5c57f8f fix(910c/glm52): start_vllm_docker.sh 加per-TP参数覆盖 + 修printf转义破坏JSON
问题: add16实际启动走start_vllm_docker.sh(用config.env全局MAX_MODEL_LEN=131072),
而非engine_build_server_args(有per-TP覆盖但只写server_cmd.txt不影响启动).
TP=8用131072上下文+无per-TP覆盖,且engine_build_server_args的printf %q转义
破坏了JSON(--additional-config的enable_dsa_cp等未生效),致MoE tiling失败.

修复:
1. start_vllm_docker.sh: 加per-TP case覆盖(TP=8:16384/64/0.95, TP=16:131072/256/0.92)
2. run_adaptive_concurrency_add16.sh: engine_build_server_args 用单引号包裹替代
   printf %q,避免JSON被反斜杠转义(仅影响server_cmd.txt记录,实际启动走start_vllm_docker.sh)
2026-07-29 11:00:15 +08:00
Zhiyi Hong
6d3338244b fix: cuda-graph-max-bs=64 (256 OOMs during capture) 2026-07-28 17:31:52 +08:00
Zhiyi Hong
cd362c5ca0 fix: max_running_requests=256, cuda_graph_max_bs=256 to match max concurrency 128 2026-07-28 17:20:25 +08:00
Zhiyi Hong
da1d4d3d78 add tp = 16 deepseek v4 pro sglang bench 2026-07-28 17:09:26 +08:00
shishi
d6e00d61dc fix(dsv4): sync glm52 add16 fixes (sglang 0.5.2 client + health timeout)
Sync the glm52 add16 fixes (98cdb67) into the DSV4 experiment so the
adaptive concurrency search can actually run end-to-end.

sglang client (the main blocker):
- Built local/vllm-ascend:0.23-a3-dsv4-sglang image: sglang 0.5.2 (not
  0.5.16 -- 0.5.16 deprecates bench_serving and the glm52 parser fix
  targets the 0.5.2 output format) + minimal deps (ipython/traitlets/
  stack_data/executing/asttokens/pure_eval/prompt_toolkit/wcwidth) via
  --no-deps, so the vllm env is untouched.
- Verified: python -m sglang.bench_serving --help works in the image.
- config.env DOCKER_IMAGE -> local/vllm-ascend:0.23-a3-dsv4-sglang.

config.env (sync glm52 98cdb67):
- CONTAINER_NAME: drop ${...:-} override -> fixed value (avoids the
  double-suffix bug where CONTAINER_NAME already carries _tpX_dpY).
- CONTAINER_PYTHON: /usr/local/bin/python (does not exist) ->
  /usr/local/python3.12.13/bin/python3 (matches glm52 fix).

run_adaptive_concurrency_add16.sh (sync glm52 98cdb67):
- --model $SERVED_MODEL_NAME -> --tokenizer $MODEL_PATH (bench_serving
  0.5.2 wants the tokenizer path).
- docker exec env: add TORCH_DEVICE_BACKEND_AUTOLOAD=0 so the client
  does not try to autoload torch_npu.
- export ENGINE_TP/ENGINE_DP in engine_start_server + export line;
  container_name uses ${ENGINE_TP:-${tp}} (the bench runs in a subshell
  where tp/dp are not in scope).

start_vllm_docker.sh (sync glm52 98cdb67):
- Health timeout configurable via HEALTH_MAX_RETRIES /
  HEALTH_RETRY_INTERVAL_S (default 480x5s=40min; TP=16 compiles 16
  graphs ~60min, old hardcoded 360x10s was too rigid).
- Container name: drop the double-suffix (CONTAINER_NAME no longer
  re-overridden before appending _tpX_dpY).
- Mount /mnt (bench client reads dataset from there).
2026-07-28 16:54:01 +08:00
shishi
6c81183fd7 feat(910c/glm52): 并行配置改为只测 TP=8 和 TP=16,并按TP区分服务参数
config.env: PARALLEL_CONFIGS 默认值从 "2 4"/"4 2"/"8 1" 改为 "8 1"/"16 1"
run_adaptive_concurrency_add16.sh: engine_build_server_args 按 TP 覆盖参数
- TP=8:  gpu_mem_util=0.95, max_model_len=16384,  max_num_seqs=64  (64GB/die KV cache 紧张)
- TP=16: gpu_mem_util=0.92, max_model_len=131072, max_num_seqs=256 (16 die 全用)
可通过 TP8_*/TP16_* 环境变量进一步覆盖
2026-07-28 16:51:14 +08:00