sskj/docs/BENCHMARK_WORKFLOW.md
shishi 63ab41b65a docs: consolidate project docs (dedup, relocate, expand 910C client guide)
Project-level documentation was scattered and duplicated across README.md,
BENCHMARK_WORKFLOW.md, and docs/EXPERIMENT_GUIDE.md (directory layout +
scripts/common component table repeated 3x). Reorganize into a clear
single-source-of-truth structure.

Changes:
- README.md: drop the 6 stale changelog entries at the top (latest was
  07-21; history lives in git log). Replace the duplicated directory-
  layout + scripts/common sections with a one-line link to
  docs/EXPERIMENT_GUIDE.md. (151 -> 99 lines)
- BENCHMARK_WORKFLOW.md -> docs/BENCHMARK_WORKFLOW.md: relocate into docs/.
  Replace its duplicated Directory Layout and Quick Start/Adding sections
  with links to EXPERIMENT_GUIDE / README / NEW_PLATFORM_GUIDE; keep the
  unique parts (Rules, Naming Conventions, Final JSON Schema, Checklist).
  (394 -> 224 lines)
- docs/EXPERIMENT_GUIDE.md: now the single authority for directory layout
  + component table + experiment conventions. Add a cross-link from the
  results.json field list to BENCHMARK_WORKFLOW's full JSON Schema and
  Naming Conventions.
- docs/H200_QUICKSTART.md: deleted (outdated, repeatedly references
  removed legacy scripts; H200 usage is covered by ADAPTIVE_CONCURRENCY_USAGE
  and experiment READMEs).
- docs/DSV4_INFERENCE_COMPARISON_REPORT.md -> experiments/h200/
  dsv4_h200_vllm_mtp_vs_default/results/20260708-160349/: this is an
  experiment report, not a project doc; relocate next to its sibling
  report.md.
- envs/ASCEND_910C_ENV_SETUP.md §8: expand the vague "pip install sglang"
  note into a full sglang client image build guide -- pin sglang 0.5.2
  (not latest; >=0.5.16 deprecates bench_serving and breaks the parser),
  --no-deps minimal install loop, docker commit to a local image, with
  the exact commands used to build local/vllm-ascend:0.23-a3-dsv4-sglang.
- experiments/h200/dsv4_h200_vllm_tp2_custom_bench/README.md: fix the
  now-broken link to BENCHMARK_WORKFLOW.md (../../ -> ../../../docs/).
- .gitignore: ignore *.bak.glm52orig scratch backups.

Also includes the add16 adaptive_results produced by the dsv4 TP=4/DP=2
runs on 910c.1.
2026-07-29 11:52:42 +08:00

12 KiB
Raw Blame History

Benchmark Workflow & Directory Conventions

Directory Layout

仓库目录结构、scripts/common/ 组件职责见 EXPERIMENT_GUIDE.md §1单一权威来源避免多处维护漂移

Rules

  1. Experiments are the primary organization unit.

    • Each experiment lives under experiments/<platform>/<experiment_name>/ (platform ∈ {h20, h200, p800, pro6000}) and contains its scripts, configuration, and results.
    • Shared orchestration code lives in scripts/common/; do not copy server start / health check logic into every experiment.
    • 新实验必须放在 experiments/<platform>/<name>/ 下;scripts/ 下仅保留 scripts/common/(原 legacy 套件 scripts/benchmark_dspark_0707/ 已删除)。
  2. Benchmark outputs live with their experiment.

    • For experiments/<platform>/<name>/, results go in experiments/<platform>/<name>/results/<RUN_ID>/.
    • Each run directory must contain report.md (human-readable) and results.json (structured data).
    • Raw outputs go in raw_outputs/; logs go in logs/.
    • Legacy bench_results/<experiment>_<timestamp>/ 目录已从仓库移除,历史结果已归档迁移至仓库外。
  3. Each experiment run directory (results/<run>/) must contain two final artifacts.

    • A Markdown report for human reading (e.g. report.md, comparison_report.md).
    • A JSON file with the complete structured result data for programmatic analysis (e.g. results.json).
    • The directory may also contain a README.md documenting provenance if the report alone does not cover it.
  4. Final JSON must contain raw/structured data, not just summary numbers.

    • Metadata: experiment name, timestamp, model, backend/inference engine, hardware/accelerator, script path, environment/commit info.
    • Record the chip/accelerator (e.g. NVIDIA H200, Kunlun XPU) and the inference engine (e.g. vllm-dspark, sglang, vllm-xpu) explicitly. Do not infer them from directory names.
    • Per-scenario/per-configuration results: all request latencies, TTFT, TPOT, ITL, token counts, throughput, accept length, success/failure counts.
    • Include P50 / P90 / P95 / P99 percentiles where applicable.
    • Keep the schema stable so downstream Python scripts can parse all experiments uniformly.
    • See Final JSON Schema below for the recommended structure.
  5. Scripts should default RESULT_ROOT to the experiment's results directory.

    • For experiments/<platform>/<name>/run_bench.sh, default to experiments/<platform>/<name>/results/${RUN_ID}/.
    • Allow override via RESULT_ROOT env var.
    • Use RUN_ID=$(date '+%Y%m%d-%H%M%S') unless specified.
  6. Server start scripts write to logs/.

    • logs/<service>_<timestamp>.log
    • Keep server logs separate from benchmark result logs.
  7. Scripts and outputs must record chip/accelerator and inference engine.

    • Every benchmark script should capture or accept the platform and engine it is running on (e.g. via environment variables CHIP, ACCELERATOR, ENGINE, BACKEND, or auto-detection).
    • Final reports and JSON outputs must include both the accelerator/chip family and the inference engine/backend used for the run.
    • Do not rely on the experiment name alone to identify the platform or engine.
  8. Use platform configuration files for chip-specific constants.

    • Put per-platform settings in platforms/<chip>.env (e.g. platforms/kunlun_p800.env, platforms/nvidia_h200.env).
    • Scripts load the platform file via scripts/common/platform.sh; the active platform is selected by the PLATFORM env var or auto-detected.
    • Keep experiment scripts free of hardcoded device IDs, image names, or model root paths.
  9. Record the exact server launch command/args for every run.

    • The exact command or full argument list used to start the server must be saved in results.json under config.server_args (or config.phaseN_server_args if the experiment starts the server in multiple phases).
    • This is required for cross-platform reproduction: when the same experiment is run on H200 and P800, the only differences should be model paths, ports, and device IDs.
    • If an experiment uses start_server.sh, that script should be self-contained and its command line should be reproducible from results.json alone.
  10. Version control: commit code and final artifacts only.

    • Always commit the experiment code (config.env, run_bench.sh, start_*.sh, parsers, etc.) together with the run's final artifacts.
    • Final artifacts to commit: results.json (structured data) and report.md / comparison.md (human-readable summaries).
    • Do not commit intermediate logs, per-request JSONL raw outputs, or GPU sampling CSVs. These are already ignored by .gitignore (raw_outputs/, logs/, gpu_logs/).
    • scenarios.tsv and skipped_after_oom.csv may be committed if they are useful for reproducing the test plan, but they are optional.

Naming Conventions

Experiment result directories

For experiment-centric layout:

experiments/<platform>/<experiment>/results/<YYYYMMDD-HHMMSS>/

Examples:

  • experiments/p800/dsv4_p800_sglang/results/20260708-120000/
  • experiments/h200/dsv4_h200_sglang_vs_vllm/results/20260707-132641/

The results.json metadata already records chip/accelerator and engine, so the directory path does not need to encode them. If a single experiment must distinguish across platforms in its directory tree, use:

experiments/<platform>/<experiment>/results/<chip>_<engine>_<YYYYMMDD-HHMMSS>/

Legacy result directories (removed)

bench_results/<experiment>_<YYYYMMDD-HHMMSS>/experiments/legacy_bench_results/ 均已从仓库移除,历史结果已归档迁移至仓库外;保留此节仅说明历史命名格式。

Raw output files

Include the accelerator and inference engine in raw output filenames so files from different platforms cannot overwrite each other.

For detailed per-request JSONL outputs:

{chip}_{engine}_{MMDD}_{concurrency}_{input_len}_{output_len}.jsonl

For summary JSON outputs from sglang.bench_serving --output-file:

{chip}_{engine}_{scenario}_{params}.json

Logs

logs/<service>_YYYYMMDD_HHMMSS.log
logs/<experiment>_orchestrator_YYYYMMDD_HHMMSS.log

Final JSON Schema

The JSON file inside each experiment run directory (experiments/<platform>/<experiment>/results/<run>/) should follow a stable schema so that downstream Python scripts can load every experiment the same way. The file is usually named results.json.

Required top-level fields

{
  "metadata": {
    "experiment": "dsv4_h200_dspark",
    "run_id": "20260707-132641",
    "timestamp": "2026-07-07T13:26:41+08:00",
    "model": "/data/models/DeepSeek-V4-Flash-DSpark",
    "backend": "vllm-dspark",
    "engine": "vllm-dspark",
    "hardware": "8x NVIDIA H200 143GB",
    "accelerator": "NVIDIA H200",
    "chip": "NVIDIA H200",
    "script": "experiments/h200/dsv4_h200_dspark/run_bench.sh",
    "env": "/path/to/envs/vllm-dspark",
    "git_commit": "optional git sha",
    "description": "optional free-text note"
  },
  "config": {
    "tp": 8,
    "kv_cache_dtype": "fp8",
    "spec_method": "dspark",
    "spec_tokens": 5,
    "block_size": 256,
    "max_num_seqs": 256,
    "extra_args": "--no-disable-hybrid-kv-cache-manager",
    "server_args": "vllm serve /data/models/DeepSeek-V4-Flash-DSpark --trust-remote-code --tensor-parallel-size 8 --kv-cache-dtype fp8 --max-model-len auto --max-num-seqs 256 --spec-method dspark --spec-tokens 5 --no-disable-hybrid-kv-cache-manager --port 30004"
  },
  "scenarios": [
    {
      "name": "chat_short",
      "concurrency": 64,
      "input_len": 1000,
      "output_len": 256,
      "duration_s": 21.27,
      "success": 512,
      "failed": 0,
      "request_throughput": 24.07,
      "input_token_throughput": 6391.14,
      "output_token_throughput": 3189.39,
      "total_token_throughput": 9580.52,
      "accept_length": 3.2,
      "latencies": {
        "e2e_ms": { "mean": 2547.81, "p50": 2400.0, "p90": 4800.0, "p95": 5606.07, "p99": 6543.97 },
        "ttft_ms": { "mean": 268.68, "p50": 240.0, "p90": 480.0, "p95": 543.21, "p99": 588.63 },
        "tpot_ms": { "mean": 18.24, "p50": 16.0, "p90": 28.0, "p95": 30.39, "p99": 39.57 },
        "itl_ms": { "mean": 70.29, "p50": 60.0, "p90": 110.0, "p95": 130.0, "p99": 160.0 }
      },
      "slo_status": {
        "ttft_p95_ok": true,
        "tpot_mean_ok": true,
        "overall": "✅"
      },
      "raw_requests": [
        {
          "request_id": "uuid-or-index",
          "input_tokens": 1000,
          "output_tokens": 256,
          "e2e_ms": 2500.0,
          "ttft_ms": 260.0,
          "tpot_ms": 18.0,
          "itl_ms": 70.0,
          "accept_length": 3.0,
          "success": true
        }
      ]
    }
  ]
}

Notes

  • Do NOT embed raw_requests in results.json. Per-request data must live only in raw_outputs/*.jsonl (gitignored). results.json must stay small (metadata + config + per-scenario summary metrics/percentiles only). Embedding raw requests caused multi-MB results.json bloat historically and is no longer done by scripts/common/parse_backend.py.
  • Always include P50 / P90 / P95 / P99 for TTFT, TPOT, E2E, and ITL. P95 is the primary SLO metric.
  • Include an slo_status object per scenario indicating whether the scenario meets the relevant SLO (e.g. S2 tier: TTFT P95 < 3000ms, TPOT mean < 50ms). Example:
    "slo_status": { "ttft_p95_ok": true, "tpot_mean_ok": true, "overall": "✅" }
    
  • Keep field names snake_case and consistent across experiments.
  • Record hardware and engine information explicitly:
    • accelerator / chip: the accelerator family, e.g. NVIDIA H200, Kunlun XPU, AMD MI300X.
    • engine / backend: the inference engine or serving backend, e.g. vllm-dspark, sglang, vllm-xpu.
    • hardware: a human-readable full hardware description, e.g. 8x NVIDIA H200 143GB, 8x Kunlun XPU R480.
    • Keep at least one of accelerator or chip, and at least one of engine or backend, populated in every run.
  • If a metric is not applicable (e.g. accept_length for non-speculative decoding), set it to null rather than omitting the key.

Quick Start / Adding Experiments

快速复现、新增实验、新平台接入的步骤见:

Checklist Before Committing / Archiving

  • No .jsonl, .json, .log, or .md files left in the project root.
  • For experiments/<platform>/<name>/, outputs live in experiments/<platform>/<name>/results/<timestamp>/.
  • <run>/report.md (or equivalent human-readable .md) exists.
  • <run>/results.json exists and follows the Final JSON Schema.
  • <run>/results.json metadata records the chip/accelerator and engine/backend used.
  • <run>/results.json config records the exact server launch command/args (server_args or phaseN_server_args).
  • <run>/results.json each scenario records slo_status (overall pass/fail/partial against the relevant SLO).
  • Raw output filenames include the chip/accelerator and engine when cross-platform runs may collide.
  • <run>/README.md exists and documents provenance (or the report itself covers provenance).
  • Scripts either live under experiments/<platform>/<name>/ or in scripts/common/.
  • Script path references updated after moving.
  • Commits include the experiment code plus final results.json and .md reports; logs/raw outputs are left ignored.