[Feat] add DSV4-Pro long-decode cases

This commit is contained in:
Zhiyi Hong 2026-07-30 23:41:08 +08:00
parent 25418ec174
commit 06b017483c
5 changed files with 38 additions and 15 deletions

View File

@ -1,5 +1,9 @@
# sskj — 多平台大模型推理性能基准测试项目 # sskj — 多平台大模型推理性能基准测试项目
> **更新2026-07-30 23:38:42 CST**
>
> 为 DeepSeek-V4-Pro 双机 Pro6000D SGLang TP16 Phase 1 增加两个最小长 Decode 补充场景:`1K → 4K, C=16` 用于观察持续长输出、KV 增长与稳态 Decode`128K → 1K, C=1` 用于观察长上下文上的 Decode Attention。继续复用唯一 `run_quick_map.sh` 入口和 `CASE_IDS` 过滤,不新增启动脚本;场景验证单测由 9 点更新为 11 点Shell、3 个 Python 单测和两点 Dry-run 已通过。
>
> **更新2026-07-30 23:06:01 CST** > **更新2026-07-30 23:06:01 CST**
> >
> 完成 DeepSeek-V4-Pro 双机 Pro6000D SGLang TP16 Phase 1 正式 quick-map。Head 与 Worker 均通过 `mlx5_0/mlx5_3` 双 Rail `NET/IB + GDRDMA` 门禁9 个固定点和 3 个混合 A/B 结果共 12/12 成功,总用时 28 分 36 秒。32K/128K 单请求 Prefill 输入吞吐为 2,652.76/2,710.16 token/s在 C=32 Decode 中注入一个 128K Prefill 后Output TPS 下降 24.08%TPOT P95 增加 66.55%。阶段 HTML 已重写为只保留成功结果,并补充正式汇总、运行清单和 Phase 2 三个诊断负载;两节点容器和 16 张 GPU 已清理。 > 完成 DeepSeek-V4-Pro 双机 Pro6000D SGLang TP16 Phase 1 正式 quick-map。Head 与 Worker 均通过 `mlx5_0/mlx5_3` 双 Rail `NET/IB + GDRDMA` 门禁9 个固定点和 3 个混合 A/B 结果共 12/12 成功,总用时 28 分 36 秒。32K/128K 单请求 Prefill 输入吞吐为 2,652.76/2,710.16 token/s在 C=32 Decode 中注入一个 128K Prefill 后Output TPS 下降 24.08%TPOT P95 增加 66.55%。阶段 HTML 已重写为只保留成功结果,并补充正式汇总、运行清单和 Phase 2 三个诊断负载;两节点容器和 16 张 GPU 已清理。

View File

@ -6,21 +6,16 @@ DeepSeek-V4-Pro on two RTX PRO 6000 Blackwell nodes. It does not modify or call
## Network precondition ## Network precondition
The 2026-07-30 audit found that the current service container does not expose The launcher exposes only the three device nodes needed by the two deployment
`/dev/infiniband`, so NCCL falls back to `NET/Socket`. The quick-map previously
selected a low-speed non-compute interface and produced invalid 1K/32K cold
Prefill baselines. The Socket default is now `eth0`, one of the two deployment
compute interfaces (`eth0/eth3`).
Declaring `NCCL_IB_HCA=mlx5_0,mlx5_3` does not by itself enable RDMA. The
launcher now exposes only the three device nodes needed by the two deployment
rails (`rdma_cm`, `uverbs0`, `uverbs3`) and fails before benchmarking unless rails (`rdma_cm`, `uverbs0`, `uverbs3`) and fails before benchmarking unless
both nodes' NCCL INFO logs prove `NET/IB` is using `mlx5_0` and `mlx5_3`. both nodes' NCCL INFO logs prove `NET/IB` is using `mlx5_0` and `mlx5_3`.
`NCCL_CROSS_NIC` has no effect while the active transport is `NET/Socket`. The 2026-07-30 Phase 1 run verified cross-node channels over
`NET/IB/*/GDRDMA`. `eth0` remains the NCCL bootstrap and Socket-fallback
interface.
## Scope ## Scope
The fixed suite covers nine points: The fixed suite covers eleven points:
| Group | ISL | OSL | Concurrency | | Group | ISL | OSL | Concurrency |
|---|---:|---:|---:| |---|---:|---:|---:|
@ -28,6 +23,8 @@ The fixed suite covers nine points:
| Prefill throughput | 32K | 1 | 16 | | Prefill throughput | 32K | 1 | 16 |
| Decode latency | 1K | 1K | 1 | | Decode latency | 1K | 1K | 1 |
| Decode throughput | 1K | 1K | 16 / 32 / 64 | | Decode throughput | 1K | 1K | 16 / 32 / 64 |
| Long-output decode | 1K | 4K | 16 |
| Long-context decode | 128K | 1K | 1 |
| Balanced pressure | 32K | 1K | 8 | | Balanced pressure | 32K | 1K | 8 |
The same main entry also runs a mixed-interference A/B: The same main entry also runs a mixed-interference A/B:
@ -47,7 +44,7 @@ attribution belongs to the later scheduler/interference phase.
|---|---| |---|---|
| `run_quick_map.sh` | The only Shell entry: service lifecycle, fixed map, mixed A/B, cleanup | | `run_quick_map.sh` | The only Shell entry: service lifecycle, fixed map, mixed A/B, cleanup |
| `config.env` | Machine, model, SGLang and benchmark configuration | | `config.env` | Machine, model, SGLang and benchmark configuration |
| `quick_map_scenarios.tsv` | Nine fixed workload points | | `quick_map_scenarios.tsv` | Eleven fixed workload points |
| `quick_map_results.py` | Validate output and generate CSV/JSONL/Markdown summaries | | `quick_map_results.py` | Validate output and generate CSV/JSONL/Markdown summaries |
| `tests/test_quick_map_results.py` | Result parser regression tests | | `tests/test_quick_map_results.py` | Result parser regression tests |
@ -77,7 +74,21 @@ CASE_IDS="short_prefill_latency_1k_c1,mid_prefill_latency_32k_c1,decode_latency_
Remove `DRY_RUN=1` only after the TP16 service has passed the `NET/IB` startup Remove `DRY_RUN=1` only after the TP16 service has passed the `NET/IB` startup
gate. An unknown `CASE_IDS` value fails before any benchmark request is sent. gate. An unknown `CASE_IDS` value fails before any benchmark request is sent.
Leave `CASE_IDS` empty to run all nine fixed points. Leave `CASE_IDS` empty to run all eleven fixed points.
To run only the two long-decode supplements:
```bash
export CASE_IDS="long_output_decode_1k_to_4k_c16,long_context_decode_128k_to_1k_c1"
export RUN_ID="dsv4pro-phase1-long-decode-$(date +%Y%m%d-%H%M%S)"
trap 'bash run_quick_map.sh stop' EXIT INT TERM
bash run_quick_map.sh start
bash run_quick_map.sh fixed
```
This uses the existing single entry for startup, the selected fixed cases, and
trap-based cleanup. It does not run the mixed A/B suite.
## GPU run ## GPU run

View File

@ -71,7 +71,7 @@ RESULT_BASE="${RESULT_BASE:-${SCRIPT_DIR}/results}"
RUNTIME_BASE="${RUNTIME_BASE:-${SCRIPT_DIR}/runtime}" RUNTIME_BASE="${RUNTIME_BASE:-${SCRIPT_DIR}/runtime}"
# Optional comma- or space-separated subset of quick_map_scenarios.tsv. # Optional comma- or space-separated subset of quick_map_scenarios.tsv.
# Leave empty for the complete nine-point fixed suite. # Leave empty for the complete eleven-point fixed suite.
CASE_IDS="${CASE_IDS:-}" CASE_IDS="${CASE_IDS:-}"
DRY_RUN="${DRY_RUN:-0}" DRY_RUN="${DRY_RUN:-0}"

View File

@ -7,4 +7,6 @@ decode_latency_1k_to_1k_c1 decode_latency 1024 1024 1 1 1 1 Single-request decod
decode_throughput_1k_to_1k_c16 decode_throughput 1024 1024 16 1 1 1 Decode throughput at C16 decode_throughput_1k_to_1k_c16 decode_throughput 1024 1024 16 1 1 1 Decode throughput at C16
decode_throughput_1k_to_1k_c32 decode_throughput 1024 1024 32 1 1 1 Decode throughput at C32 decode_throughput_1k_to_1k_c32 decode_throughput 1024 1024 32 1 1 1 Decode throughput at C32
decode_throughput_1k_to_1k_c64 decode_throughput 1024 1024 64 1 1 1 Decode throughput at C64 decode_throughput_1k_to_1k_c64 decode_throughput 1024 1024 64 1 1 1 Decode throughput at C64
long_output_decode_1k_to_4k_c16 long_output_decode 1024 4096 16 1 1 1 Sustained long-output decode and KV growth
long_context_decode_128k_to_1k_c1 long_context_decode 131072 1024 1 1 1 0 Decode attention over a 128K context
balanced_32k_to_1k_c8 balanced 32768 1024 8 1 1 0 Balanced prefill and decode pressure balanced_32k_to_1k_c8 balanced 32768 1024 8 1 1 0 Balanced prefill and decode pressure

1 # case_id stage isl osl concurrency prompt_multiplier min_prompts warmup_requests note
7 decode_throughput_1k_to_1k_c16 decode_throughput 1024 1024 16 1 1 1 Decode throughput at C16
8 decode_throughput_1k_to_1k_c32 decode_throughput 1024 1024 32 1 1 1 Decode throughput at C32
9 decode_throughput_1k_to_1k_c64 decode_throughput 1024 1024 64 1 1 1 Decode throughput at C64
10 long_output_decode_1k_to_4k_c16 long_output_decode 1024 4096 16 1 1 1 Sustained long-output decode and KV growth
11 long_context_decode_128k_to_1k_c1 long_context_decode 131072 1024 1 1 1 0 Decode attention over a 128K context
12 balanced_32k_to_1k_c8 balanced 32768 1024 8 1 1 0 Balanced prefill and decode pressure

View File

@ -16,13 +16,19 @@ import quick_map_results # noqa: E402
class QuickMapResultsTest(unittest.TestCase): class QuickMapResultsTest(unittest.TestCase):
def test_scenario_file_is_valid_and_unique(self) -> None: def test_scenario_file_is_valid_and_unique(self) -> None:
scenarios = quick_map_results.parse_scenarios(EXPERIMENT_DIR / "quick_map_scenarios.tsv") scenarios = quick_map_results.parse_scenarios(EXPERIMENT_DIR / "quick_map_scenarios.tsv")
self.assertEqual(len(scenarios), 9) self.assertEqual(len(scenarios), 11)
self.assertEqual(len({item["case_id"] for item in scenarios}), 9) self.assertEqual(len({item["case_id"] for item in scenarios}), 11)
self.assertEqual(scenarios[2]["isl"], 131072) self.assertEqual(scenarios[2]["isl"], 131072)
self.assertEqual(scenarios[2]["case_id"], "long_prefill_latency_128k_c1") self.assertEqual(scenarios[2]["case_id"], "long_prefill_latency_128k_c1")
self.assertEqual(scenarios[2]["prompt_multiplier"], 1) self.assertEqual(scenarios[2]["prompt_multiplier"], 1)
self.assertEqual(scenarios[2]["min_prompts"], 1) self.assertEqual(scenarios[2]["min_prompts"], 1)
self.assertEqual(scenarios[2]["warmup_requests"], 0) self.assertEqual(scenarios[2]["warmup_requests"], 0)
by_id = {item["case_id"]: item for item in scenarios}
self.assertEqual(by_id["long_output_decode_1k_to_4k_c16"]["osl"], 4096)
self.assertEqual(by_id["long_output_decode_1k_to_4k_c16"]["concurrency"], 16)
self.assertEqual(by_id["long_context_decode_128k_to_1k_c1"]["isl"], 131072)
self.assertEqual(by_id["long_context_decode_128k_to_1k_c1"]["osl"], 1024)
self.assertEqual(by_id["long_context_decode_128k_to_1k_c1"]["warmup_requests"], 0)
def test_latency_percentiles_fall_back_to_request_arrays(self) -> None: def test_latency_percentiles_fall_back_to_request_arrays(self) -> None:
data = { data = {