sskj/docs/BENCHMARK_WORKFLOW.md
shishi 63ab41b65a docs: consolidate project docs (dedup, relocate, expand 910C client guide)
Project-level documentation was scattered and duplicated across README.md,
BENCHMARK_WORKFLOW.md, and docs/EXPERIMENT_GUIDE.md (directory layout +
scripts/common component table repeated 3x). Reorganize into a clear
single-source-of-truth structure.

Changes:
- README.md: drop the 6 stale changelog entries at the top (latest was
  07-21; history lives in git log). Replace the duplicated directory-
  layout + scripts/common sections with a one-line link to
  docs/EXPERIMENT_GUIDE.md. (151 -> 99 lines)
- BENCHMARK_WORKFLOW.md -> docs/BENCHMARK_WORKFLOW.md: relocate into docs/.
  Replace its duplicated Directory Layout and Quick Start/Adding sections
  with links to EXPERIMENT_GUIDE / README / NEW_PLATFORM_GUIDE; keep the
  unique parts (Rules, Naming Conventions, Final JSON Schema, Checklist).
  (394 -> 224 lines)
- docs/EXPERIMENT_GUIDE.md: now the single authority for directory layout
  + component table + experiment conventions. Add a cross-link from the
  results.json field list to BENCHMARK_WORKFLOW's full JSON Schema and
  Naming Conventions.
- docs/H200_QUICKSTART.md: deleted (outdated, repeatedly references
  removed legacy scripts; H200 usage is covered by ADAPTIVE_CONCURRENCY_USAGE
  and experiment READMEs).
- docs/DSV4_INFERENCE_COMPARISON_REPORT.md -> experiments/h200/
  dsv4_h200_vllm_mtp_vs_default/results/20260708-160349/: this is an
  experiment report, not a project doc; relocate next to its sibling
  report.md.
- envs/ASCEND_910C_ENV_SETUP.md §8: expand the vague "pip install sglang"
  note into a full sglang client image build guide -- pin sglang 0.5.2
  (not latest; >=0.5.16 deprecates bench_serving and breaks the parser),
  --no-deps minimal install loop, docker commit to a local image, with
  the exact commands used to build local/vllm-ascend:0.23-a3-dsv4-sglang.
- experiments/h200/dsv4_h200_vllm_tp2_custom_bench/README.md: fix the
  now-broken link to BENCHMARK_WORKFLOW.md (../../ -> ../../../docs/).
- .gitignore: ignore *.bak.glm52orig scratch backups.

Also includes the add16 adaptive_results produced by the dsv4 TP=4/DP=2
runs on 910c.1.
2026-07-29 11:52:42 +08:00

224 lines
12 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Benchmark Workflow & Directory Conventions
## Directory Layout
仓库目录结构、`scripts/common/` 组件职责见 [`EXPERIMENT_GUIDE.md`](EXPERIMENT_GUIDE.md) §1单一权威来源避免多处维护漂移
## Rules
1. **Experiments are the primary organization unit.**
- Each experiment lives under `experiments/<platform>/<experiment_name>/` (platform ∈ {h20, h200, p800, pro6000}) and contains its scripts, configuration, and results.
- Shared orchestration code lives in `scripts/common/`; do not copy server start / health check logic into every experiment.
- 新实验必须放在 `experiments/<platform>/<name>/` 下;`scripts/` 下仅保留 `scripts/common/`(原 legacy 套件 `scripts/benchmark_dspark_0707/` 已删除)。
2. **Benchmark outputs live with their experiment.**
- For `experiments/<platform>/<name>/`, results go in `experiments/<platform>/<name>/results/<RUN_ID>/`.
- Each run directory must contain `report.md` (human-readable) and `results.json` (structured data).
- Raw outputs go in `raw_outputs/`; logs go in `logs/`.
- Legacy `bench_results/<experiment>_<timestamp>/` 目录已从仓库移除,历史结果已归档迁移至仓库外。
3. **Each experiment run directory (`results/<run>/`) must contain two final artifacts.**
- A Markdown report for human reading (e.g. `report.md`, `comparison_report.md`).
- A JSON file with the complete structured result data for programmatic analysis (e.g. `results.json`).
- The directory may also contain a `README.md` documenting provenance if the report alone does not cover it.
4. **Final JSON must contain raw/structured data, not just summary numbers.**
- Metadata: experiment name, timestamp, model, backend/inference engine, hardware/accelerator, script path, environment/commit info.
- Record the **chip/accelerator** (e.g. `NVIDIA H200`, `Kunlun XPU`) and the **inference engine** (e.g. `vllm-dspark`, `sglang`, `vllm-xpu`) explicitly. Do not infer them from directory names.
- Per-scenario/per-configuration results: all request latencies, TTFT, TPOT, ITL, token counts, throughput, accept length, success/failure counts.
- Include P50 / P90 / P95 / P99 percentiles where applicable.
- Keep the schema stable so downstream Python scripts can parse all experiments uniformly.
- See [Final JSON Schema](#final-json-schema) below for the recommended structure.
5. **Scripts should default `RESULT_ROOT` to the experiment's results directory.**
- For `experiments/<platform>/<name>/run_bench.sh`, default to `experiments/<platform>/<name>/results/${RUN_ID}/`.
- Allow override via `RESULT_ROOT` env var.
- Use `RUN_ID=$(date '+%Y%m%d-%H%M%S')` unless specified.
6. **Server start scripts write to `logs/`.**
- `logs/<service>_<timestamp>.log`
- Keep server logs separate from benchmark result logs.
7. **Scripts and outputs must record chip/accelerator and inference engine.**
- Every benchmark script should capture or accept the platform and engine it is running on (e.g. via environment variables `CHIP`, `ACCELERATOR`, `ENGINE`, `BACKEND`, or auto-detection).
- Final reports and JSON outputs must include both the accelerator/chip family and the inference engine/backend used for the run.
- Do not rely on the experiment name alone to identify the platform or engine.
8. **Use platform configuration files for chip-specific constants.**
- Put per-platform settings in `platforms/<chip>.env` (e.g. `platforms/kunlun_p800.env`, `platforms/nvidia_h200.env`).
- Scripts load the platform file via `scripts/common/platform.sh`; the active platform is selected by the `PLATFORM` env var or auto-detected.
- Keep experiment scripts free of hardcoded device IDs, image names, or model root paths.
9. **Record the exact server launch command/args for every run.**
- The exact command or full argument list used to start the server must be saved in `results.json` under `config.server_args` (or `config.phaseN_server_args` if the experiment starts the server in multiple phases).
- This is required for cross-platform reproduction: when the same experiment is run on H200 and P800, the only differences should be model paths, ports, and device IDs.
- If an experiment uses `start_server.sh`, that script should be self-contained and its command line should be reproducible from `results.json` alone.
10. **Version control: commit code and final artifacts only.**
- Always commit the experiment code (`config.env`, `run_bench.sh`, `start_*.sh`, parsers, etc.) together with the run's final artifacts.
- Final artifacts to commit: `results.json` (structured data) and `report.md` / `comparison.md` (human-readable summaries).
- Do **not** commit intermediate logs, per-request JSONL raw outputs, or GPU sampling CSVs. These are already ignored by `.gitignore` (`raw_outputs/`, `logs/`, `gpu_logs/`).
- `scenarios.tsv` and `skipped_after_oom.csv` may be committed if they are useful for reproducing the test plan, but they are optional.
## Naming Conventions
### Experiment result directories
For experiment-centric layout:
```
experiments/<platform>/<experiment>/results/<YYYYMMDD-HHMMSS>/
```
Examples:
- `experiments/p800/dsv4_p800_sglang/results/20260708-120000/`
- `experiments/h200/dsv4_h200_sglang_vs_vllm/results/20260707-132641/`
The `results.json` metadata already records `chip`/`accelerator` and `engine`, so the directory path does not need to encode them. If a single experiment must distinguish across platforms in its directory tree, use:
```
experiments/<platform>/<experiment>/results/<chip>_<engine>_<YYYYMMDD-HHMMSS>/
```
### Legacy result directories (removed)
`bench_results/<experiment>_<YYYYMMDD-HHMMSS>/``experiments/legacy_bench_results/` 均已从仓库移除,历史结果已归档迁移至仓库外;保留此节仅说明历史命名格式。
### Raw output files
Include the accelerator and inference engine in raw output filenames so files from different platforms cannot overwrite each other.
For detailed per-request JSONL outputs:
```
{chip}_{engine}_{MMDD}_{concurrency}_{input_len}_{output_len}.jsonl
```
For summary JSON outputs from `sglang.bench_serving --output-file`:
```
{chip}_{engine}_{scenario}_{params}.json
```
### Logs
```
logs/<service>_YYYYMMDD_HHMMSS.log
logs/<experiment>_orchestrator_YYYYMMDD_HHMMSS.log
```
## Final JSON Schema
The JSON file inside each experiment run directory (`experiments/<platform>/<experiment>/results/<run>/`) should follow a stable schema so that downstream Python scripts can load every experiment the same way. The file is usually named `results.json`.
### Required top-level fields
```json
{
"metadata": {
"experiment": "dsv4_h200_dspark",
"run_id": "20260707-132641",
"timestamp": "2026-07-07T13:26:41+08:00",
"model": "/data/models/DeepSeek-V4-Flash-DSpark",
"backend": "vllm-dspark",
"engine": "vllm-dspark",
"hardware": "8x NVIDIA H200 143GB",
"accelerator": "NVIDIA H200",
"chip": "NVIDIA H200",
"script": "experiments/h200/dsv4_h200_dspark/run_bench.sh",
"env": "/path/to/envs/vllm-dspark",
"git_commit": "optional git sha",
"description": "optional free-text note"
},
"config": {
"tp": 8,
"kv_cache_dtype": "fp8",
"spec_method": "dspark",
"spec_tokens": 5,
"block_size": 256,
"max_num_seqs": 256,
"extra_args": "--no-disable-hybrid-kv-cache-manager",
"server_args": "vllm serve /data/models/DeepSeek-V4-Flash-DSpark --trust-remote-code --tensor-parallel-size 8 --kv-cache-dtype fp8 --max-model-len auto --max-num-seqs 256 --spec-method dspark --spec-tokens 5 --no-disable-hybrid-kv-cache-manager --port 30004"
},
"scenarios": [
{
"name": "chat_short",
"concurrency": 64,
"input_len": 1000,
"output_len": 256,
"duration_s": 21.27,
"success": 512,
"failed": 0,
"request_throughput": 24.07,
"input_token_throughput": 6391.14,
"output_token_throughput": 3189.39,
"total_token_throughput": 9580.52,
"accept_length": 3.2,
"latencies": {
"e2e_ms": { "mean": 2547.81, "p50": 2400.0, "p90": 4800.0, "p95": 5606.07, "p99": 6543.97 },
"ttft_ms": { "mean": 268.68, "p50": 240.0, "p90": 480.0, "p95": 543.21, "p99": 588.63 },
"tpot_ms": { "mean": 18.24, "p50": 16.0, "p90": 28.0, "p95": 30.39, "p99": 39.57 },
"itl_ms": { "mean": 70.29, "p50": 60.0, "p90": 110.0, "p95": 130.0, "p99": 160.0 }
},
"slo_status": {
"ttft_p95_ok": true,
"tpot_mean_ok": true,
"overall": "✅"
},
"raw_requests": [
{
"request_id": "uuid-or-index",
"input_tokens": 1000,
"output_tokens": 256,
"e2e_ms": 2500.0,
"ttft_ms": 260.0,
"tpot_ms": 18.0,
"itl_ms": 70.0,
"accept_length": 3.0,
"success": true
}
]
}
]
}
```
### Notes
- **Do NOT embed `raw_requests` in `results.json`.** Per-request data must live only in `raw_outputs/*.jsonl` (gitignored). `results.json` must stay small (metadata + config + per-scenario summary metrics/percentiles only). Embedding raw requests caused multi-MB `results.json` bloat historically and is no longer done by `scripts/common/parse_backend.py`.
- Always include **P50 / P90 / P95 / P99** for TTFT, TPOT, E2E, and ITL. P95 is the primary SLO metric.
- Include an `slo_status` object per scenario indicating whether the scenario meets the relevant SLO (e.g. S2 tier: TTFT P95 < 3000ms, TPOT mean < 50ms). Example:
```json
"slo_status": { "ttft_p95_ok": true, "tpot_mean_ok": true, "overall": "✅" }
```
- Keep field names snake_case and consistent across experiments.
- Record hardware and engine information explicitly:
- `accelerator` / `chip`: the accelerator family, e.g. `NVIDIA H200`, `Kunlun XPU`, `AMD MI300X`.
- `engine` / `backend`: the inference engine or serving backend, e.g. `vllm-dspark`, `sglang`, `vllm-xpu`.
- `hardware`: a human-readable full hardware description, e.g. `8x NVIDIA H200 143GB`, `8x Kunlun XPU R480`.
- Keep at least one of `accelerator` or `chip`, and at least one of `engine` or `backend`, populated in every run.
- If a metric is not applicable (e.g. `accept_length` for non-speculative decoding), set it to `null` rather than omitting the key.
## Quick Start / Adding Experiments
快速复现、新增实验、新平台接入的步骤见:
- [`../README.md`](../README.md) §快速复现(命令示例)
- [`EXPERIMENT_GUIDE.md`](EXPERIMENT_GUIDE.md) §2 新增实验config.env 必备字段)
- [`NEW_PLATFORM_GUIDE.md`](NEW_PLATFORM_GUIDE.md)(新平台接入 SOP
## Checklist Before Committing / Archiving
- [ ] No `.jsonl`, `.json`, `.log`, or `.md` files left in the project root.
- [ ] For `experiments/<platform>/<name>/`, outputs live in `experiments/<platform>/<name>/results/<timestamp>/`.
- [ ] `<run>/report.md` (or equivalent human-readable `.md`) exists.
- [ ] `<run>/results.json` exists and follows the [Final JSON Schema](#final-json-schema).
- [ ] `<run>/results.json` metadata records the `chip`/`accelerator` and `engine`/`backend` used.
- [ ] `<run>/results.json` `config` records the exact server launch command/args (`server_args` or `phaseN_server_args`).
- [ ] `<run>/results.json` each scenario records `slo_status` (overall pass/fail/partial against the relevant SLO).
- [ ] Raw output filenames include the chip/accelerator and engine when cross-platform runs may collide.
- [ ] `<run>/README.md` exists and documents provenance (or the report itself covers provenance).
- [ ] Scripts either live under `experiments/<platform>/<name>/` or in `scripts/common/`.
- [ ] Script path references updated after moving.
- [ ] Commits include the experiment code plus final `results.json` and `.md` reports; logs/raw outputs are left ignored.