sskj/BENCHMARK_WORKFLOW.md

18 KiB
Raw Blame History

Benchmark Workflow & Directory Conventions

Directory Layout

<repo-root>/                  # 本仓库根目录(如 /data1/yy/sskj按机器调整
├── platforms/                # chip/accelerator platform configs
│   ├── kunlun_p800.env
│   ├── nvidia_h20.env
│   ├── nvidia_h200.env
│   ├── nvidia_rtx6000d.env
│   └── patches/kunlun_p800/  # runtime patches required by some images
├── scripts/
│   └── common/               # reusable components (lib.sh, platform.sh, parse_backend.py, adaptive_bench_lib.sh, ...)
├── docs/
│   └── SLO_STANDARDS.md      # 推理服务 SLO 标准
├── experiments/              # experiment-centric directories (preferred): experiments/<platform>/<experiment>/
│   ├── TEMPLATE/             # template for fixed-scenario sglang-vs-vllm experiments
│   └── <platform>/           # h20 / h200 / p800 / pro6000
│       └── dsv4_p800_sglang/
│           ├── README.md
│           ├── config.env    # experiment-level configuration
│           ├── start_server.sh
│           ├── run_bench.sh
│           ├── parse_results.py
│           └── results/
│               └── 20260708-XXXXXX/
│                   ├── report.md
│                   ├── results.json
│                   └── logs/
├── datasets/                 # benchmark datasets
└── envs/                     # Python virtual environments

legacy 的 scripts/benchmark_dspark_0707/ 已删除;顶层 bench_results/logs/ 已不在仓库中(历史结果已归档迁移至仓库外)。

Rules

  1. Experiments are the primary organization unit.

    • Each experiment lives under experiments/<platform>/<experiment_name>/ (platform ∈ {h20, h200, p800, pro6000}) and contains its scripts, configuration, and results.
    • Shared orchestration code lives in scripts/common/; do not copy server start / health check logic into every experiment.
    • 新实验必须放在 experiments/<platform>/<name>/ 下;scripts/ 下仅保留 scripts/common/(原 legacy 套件 scripts/benchmark_dspark_0707/ 已删除)。
  2. Benchmark outputs live with their experiment.

    • For experiments/<platform>/<name>/, results go in experiments/<platform>/<name>/results/<RUN_ID>/.
    • Each run directory must contain report.md (human-readable) and results.json (structured data).
    • Raw outputs go in raw_outputs/; logs go in logs/.
    • Legacy bench_results/<experiment>_<timestamp>/ 目录已从仓库移除,历史结果已归档迁移至仓库外。
  3. Each experiment run directory (results/<run>/) must contain two final artifacts.

    • A Markdown report for human reading (e.g. report.md, comparison_report.md).
    • A JSON file with the complete structured result data for programmatic analysis (e.g. results.json).
    • The directory may also contain a README.md documenting provenance if the report alone does not cover it.
  4. Final JSON must contain raw/structured data, not just summary numbers.

    • Metadata: experiment name, timestamp, model, backend/inference engine, hardware/accelerator, script path, environment/commit info.
    • Record the chip/accelerator (e.g. NVIDIA H200, Kunlun XPU) and the inference engine (e.g. vllm-dspark, sglang, vllm-xpu) explicitly. Do not infer them from directory names.
    • Per-scenario/per-configuration results: all request latencies, TTFT, TPOT, ITL, token counts, throughput, accept length, success/failure counts.
    • Include P50 / P90 / P95 / P99 percentiles where applicable.
    • Keep the schema stable so downstream Python scripts can parse all experiments uniformly.
    • See Final JSON Schema below for the recommended structure.
  5. Scripts should default RESULT_ROOT to the experiment's results directory.

    • For experiments/<platform>/<name>/run_bench.sh, default to experiments/<platform>/<name>/results/${RUN_ID}/.
    • Allow override via RESULT_ROOT env var.
    • Use RUN_ID=$(date '+%Y%m%d-%H%M%S') unless specified.
  6. Server start scripts write to logs/.

    • logs/<service>_<timestamp>.log
    • Keep server logs separate from benchmark result logs.
  7. Scripts and outputs must record chip/accelerator and inference engine.

    • Every benchmark script should capture or accept the platform and engine it is running on (e.g. via environment variables CHIP, ACCELERATOR, ENGINE, BACKEND, or auto-detection).
    • Final reports and JSON outputs must include both the accelerator/chip family and the inference engine/backend used for the run.
    • Do not rely on the experiment name alone to identify the platform or engine.
  8. Use platform configuration files for chip-specific constants.

    • Put per-platform settings in platforms/<chip>.env (e.g. platforms/kunlun_p800.env, platforms/nvidia_h200.env).
    • Scripts load the platform file via scripts/common/platform.sh; the active platform is selected by the PLATFORM env var or auto-detected.
    • Keep experiment scripts free of hardcoded device IDs, image names, or model root paths.
  9. Record the exact server launch command/args for every run.

    • The exact command or full argument list used to start the server must be saved in results.json under config.server_args (or config.phaseN_server_args if the experiment starts the server in multiple phases).
    • This is required for cross-platform reproduction: when the same experiment is run on H200 and P800, the only differences should be model paths, ports, and device IDs.
    • If an experiment uses start_server.sh, that script should be self-contained and its command line should be reproducible from results.json alone.
  10. Version control: commit code and final artifacts only.

    • Always commit the experiment code (config.env, run_bench.sh, start_*.sh, parsers, etc.) together with the run's final artifacts.
    • Final artifacts to commit: results.json (structured data) and report.md / comparison.md (human-readable summaries).
    • Do not commit intermediate logs, per-request JSONL raw outputs, or GPU sampling CSVs. These are already ignored by .gitignore (raw_outputs/, logs/, gpu_logs/).
    • scenarios.tsv and skipped_after_oom.csv may be committed if they are useful for reproducing the test plan, but they are optional.

Naming Conventions

Experiment result directories

For experiment-centric layout:

experiments/<platform>/<experiment>/results/<YYYYMMDD-HHMMSS>/

Examples:

  • experiments/p800/dsv4_p800_sglang/results/20260708-120000/
  • experiments/h200/dsv4_h200_sglang_vs_vllm/results/20260707-132641/

The results.json metadata already records chip/accelerator and engine, so the directory path does not need to encode them. If a single experiment must distinguish across platforms in its directory tree, use:

experiments/<platform>/<experiment>/results/<chip>_<engine>_<YYYYMMDD-HHMMSS>/

Legacy result directories (removed)

bench_results/<experiment>_<YYYYMMDD-HHMMSS>/experiments/legacy_bench_results/ 均已从仓库移除,历史结果已归档迁移至仓库外;保留此节仅说明历史命名格式。

Raw output files

Include the accelerator and inference engine in raw output filenames so files from different platforms cannot overwrite each other.

For detailed per-request JSONL outputs:

{chip}_{engine}_{MMDD}_{concurrency}_{input_len}_{output_len}.jsonl

For summary JSON outputs from sglang.bench_serving --output-file:

{chip}_{engine}_{scenario}_{params}.json

Logs

logs/<service>_YYYYMMDD_HHMMSS.log
logs/<experiment>_orchestrator_YYYYMMDD_HHMMSS.log

Final JSON Schema

The JSON file inside each experiment run directory (experiments/<platform>/<experiment>/results/<run>/) should follow a stable schema so that downstream Python scripts can load every experiment the same way. The file is usually named results.json.

Required top-level fields

{
  "metadata": {
    "experiment": "dsv4_h200_dspark",
    "run_id": "20260707-132641",
    "timestamp": "2026-07-07T13:26:41+08:00",
    "model": "/data/models/DeepSeek-V4-Flash-DSpark",
    "backend": "vllm-dspark",
    "engine": "vllm-dspark",
    "hardware": "8x NVIDIA H200 143GB",
    "accelerator": "NVIDIA H200",
    "chip": "NVIDIA H200",
    "script": "experiments/h200/dsv4_h200_dspark/run_bench.sh",
    "env": "/path/to/envs/vllm-dspark",
    "git_commit": "optional git sha",
    "description": "optional free-text note"
  },
  "config": {
    "tp": 8,
    "kv_cache_dtype": "fp8",
    "spec_method": "dspark",
    "spec_tokens": 5,
    "block_size": 256,
    "max_num_seqs": 256,
    "extra_args": "--no-disable-hybrid-kv-cache-manager",
    "server_args": "vllm serve /data/models/DeepSeek-V4-Flash-DSpark --trust-remote-code --tensor-parallel-size 8 --kv-cache-dtype fp8 --max-model-len auto --max-num-seqs 256 --spec-method dspark --spec-tokens 5 --no-disable-hybrid-kv-cache-manager --port 30004"
  },
  "scenarios": [
    {
      "name": "chat_short",
      "concurrency": 64,
      "input_len": 1000,
      "output_len": 256,
      "duration_s": 21.27,
      "success": 512,
      "failed": 0,
      "request_throughput": 24.07,
      "input_token_throughput": 6391.14,
      "output_token_throughput": 3189.39,
      "total_token_throughput": 9580.52,
      "accept_length": 3.2,
      "latencies": {
        "e2e_ms": { "mean": 2547.81, "p50": 2400.0, "p90": 4800.0, "p95": 5606.07, "p99": 6543.97 },
        "ttft_ms": { "mean": 268.68, "p50": 240.0, "p90": 480.0, "p95": 543.21, "p99": 588.63 },
        "tpot_ms": { "mean": 18.24, "p50": 16.0, "p90": 28.0, "p95": 30.39, "p99": 39.57 },
        "itl_ms": { "mean": 70.29, "p50": 60.0, "p90": 110.0, "p95": 130.0, "p99": 160.0 }
      },
      "slo_status": {
        "ttft_p95_ok": true,
        "tpot_mean_ok": true,
        "overall": "✅"
      },
      "raw_requests": [
        {
          "request_id": "uuid-or-index",
          "input_tokens": 1000,
          "output_tokens": 256,
          "e2e_ms": 2500.0,
          "ttft_ms": 260.0,
          "tpot_ms": 18.0,
          "itl_ms": 70.0,
          "accept_length": 3.0,
          "success": true
        }
      ]
    }
  ]
}

Notes

  • Do NOT embed raw_requests in results.json. Per-request data must live only in raw_outputs/*.jsonl (gitignored). results.json must stay small (metadata + config + per-scenario summary metrics/percentiles only). Embedding raw requests caused multi-MB results.json bloat historically and is no longer done by scripts/common/parse_backend.py.
  • Always include P50 / P90 / P95 / P99 for TTFT, TPOT, E2E, and ITL. P95 is the primary SLO metric.
  • Include an slo_status object per scenario indicating whether the scenario meets the relevant SLO (e.g. S2 tier: TTFT P95 < 3000ms, TPOT mean < 50ms). Example:
    "slo_status": { "ttft_p95_ok": true, "tpot_mean_ok": true, "overall": "✅" }
    
  • Keep field names snake_case and consistent across experiments.
  • Record hardware and engine information explicitly:
    • accelerator / chip: the accelerator family, e.g. NVIDIA H200, Kunlun XPU, AMD MI300X.
    • engine / backend: the inference engine or serving backend, e.g. vllm-dspark, sglang, vllm-xpu.
    • hardware: a human-readable full hardware description, e.g. 8x NVIDIA H200 143GB, 8x Kunlun XPU R480.
    • Keep at least one of accelerator or chip, and at least one of engine or backend, populated in every run.
  • If a metric is not applicable (e.g. accept_length for non-speculative decoding), set it to null rather than omitting the key.

Quick Start

Run P800 SGLang benchmark (Kunlun P800, Docker)

bash experiments/p800/dsv4_p800_sglang/run_bench.sh

Run H200 DSpark benchmark (NVIDIA H200, native venv)

bash experiments/h200/dsv4_h200_dspark/run_bench.sh

Note: the wrapper's default LEGACY_GRID_SCRIPT / SERVER_START_SCRIPT point at the removed legacy scripts; override them before running (see docs/H200_QUICKSTART.md).

Legacy DSpark scripts (removed)

The legacy suite scripts/benchmark_dspark_0707/ (grid benchmark, spec-tokens comparison, etc.) has been removed from this repo.

Parse results

# Experiment-centric layout
python3 experiments/p800/dsv4_p800_sglang/parse_results.py \
  experiments/p800/dsv4_p800_sglang/results/<run_id>

Adding a New Platform or Experiment

Quick Start: New Experiment on an Existing Platform (e.g. a new P800 box)

Use this when the platform already exists (platforms/<chip>.env present) and you only want to test a new model/config on a machine.

0. Machine setup (one-time, only on a brand-new machine)

git clone https://git.meta-stone.net/qqtang/sskj.git /data1/yy/sskj
cd /data1/yy/sskj
git config credential.helper store          # push once to save the token
# Verify platform auto-detection:
bash -c 'source scripts/common/platform.sh && echo "$PLATFORM | $HARDWARE | $ENGINE"'
# Ensure platform deps exist (see platforms/<chip>.env): docker image pulled, MODEL_ROOT has the model, datasets/ present.

If the platform is not auto-detected, run with PLATFORM=<chip> .... Platform-wide paths/images live in platforms/<chip>.env - edit that file (not the experiment) if a path differs on this machine.

1. Create the experiment from TEMPLATE

cd /data1/yy/sskj
cp -r experiments/TEMPLATE experiments/p800/<your_experiment_name>   # e.g. qwen3_8b_p800_sglang_tp4
cd experiments/p800/<your_experiment_name>

2. Edit config.env - the single source of experiment params (model, TP, ports, scenarios). All values use ${VAR:-default}, so they can be overridden by env vars at run time.

3. Edit start_sglang.sh / start_server.sh - the deployment script (docker run + launch flags). This is where you change deployment params (TP, mem-fraction, attention backend, cuda graph, ...). Reference config.env variables ($TP, $MEM_FRACTION_STATIC, ...) rather than hardcoding.

4. Run

bash run_bench.sh                                  # auto-detects platform
# or with overrides / explicit RUN_ID:
PLATFORM=kunlun_p800 RUN_ID=tp4_20260722 bash run_bench.sh
# quick config probe without editing files:
TP=4 MEM_FRACTION_STATIC=0.85 MAX_RUNNING=32 bash run_bench.sh

Results land in results/<RUN_ID>/ (report.md, results.json; raw_outputs/ and logs/ are gitignored). Watch logs/orchestrator.log and logs/server.log.

5. Commit code + final artifacts, push, open PR

cd /data1/yy/sskj
git checkout -b feat/p800-<your_experiment_name>
git add experiments/p800/<your_experiment_name>/config.env \
        experiments/p800/<your_experiment_name>/start_sglang.sh \
        experiments/p800/<your_experiment_name>/run_bench.sh \
        experiments/p800/<your_experiment_name>/README.md \
        experiments/p800/<your_experiment_name>/results/<RUN_ID>/report.md \
        experiments/p800/<your_experiment_name>/results/<RUN_ID>/results.json
git commit -m "feat(p800): add <your_experiment_name> experiment"
git push -u origin feat/p800-<your_experiment_name>

Do NOT git add raw_outputs/ or logs/ (already ignored). Run through the Checklist below before pushing.

Adding a brand-new chip (not just a new experiment on an existing platform)? Follow steps 1-4 below instead.

1. Add or update a platform config

Create platforms/<chip>.env with identity and platform-wide paths:

CHIP="my_chip"
ACCELERATOR="My Accelerator"
HARDWARE="8x My Accelerator"
ENGINE="vllm-myengine"
DEFAULT_PORT="30000"
MODEL_ROOT="/data/models"

For Docker-based platforms, also set DOCKER_IMAGE, CONTAINER_NAME, CONTAINER_PYTHON, and PATCH_ROOT (see platforms/kunlun_p800.env). For host-native platforms, set the relevant venv paths (see platforms/nvidia_h200.env).

2. Create an experiment directory

experiments/<platform>/<name>/
├── README.md          # Purpose and usage
├── config.env         # Model, port, scenarios, engine overrides
├── start_server.sh    # (optional) platform-specific server launch
├── run_bench.sh       # Orchestrator
└── parse_results.py   # Convert raw outputs to results.json + report.md

At minimum, run_bench.sh should:

  1. Source scripts/common/lib.sh and scripts/common/platform.sh.
  2. Read config.env.
  3. Create experiments/<platform>/<name>/results/<RUN_ID>/.
  4. Call write_metadata_json to create results.json.
  5. Record the exact server launch command/args in results.json config.server_args (or phaseN_server_args).
  6. Run the benchmark scenarios.
  7. Call parse_results.py to generate report.md.

3. Reuse shared helpers

  • scripts/common/lib.sh: logging, health checks, metadata JSON.
  • scripts/common/platform.sh: platform auto-detection and env loading.
  • scripts/common/warmup.py: server 预热。
  • scripts/common/parse_backend.py: raw jsonl -> results.json + report.md
  • scripts/common/compare.py: SGLang vs vLLM 横向对比表。

Docker/XPU 平台可以在 scripts/common/ 下新增自己的 helper但不要在实验目录复制通用逻辑。

4. Example experiments

  • Docker / XPU: experiments/p800/dsv4_p800_sglang/
  • Native / H200: experiments/h200/dsv4_h200_dspark/

See docs/H200_QUICKSTART.md for a concrete H200 porting walkthrough.

Checklist Before Committing / Archiving

  • No .jsonl, .json, .log, or .md files left in the project root.
  • For experiments/<platform>/<name>/, outputs live in experiments/<platform>/<name>/results/<timestamp>/.
  • <run>/report.md (or equivalent human-readable .md) exists.
  • <run>/results.json exists and follows the Final JSON Schema.
  • <run>/results.json metadata records the chip/accelerator and engine/backend used.
  • <run>/results.json config records the exact server launch command/args (server_args or phaseN_server_args).
  • <run>/results.json each scenario records slo_status (overall pass/fail/partial against the relevant SLO).
  • Raw output filenames include the chip/accelerator and engine when cross-platform runs may collide.
  • <run>/README.md exists and documents provenance (or the report itself covers provenance).
  • Scripts either live under experiments/<platform>/<name>/ or in scripts/common/.
  • Script path references updated after moving.
  • Commits include the experiment code plus final results.json and .md reports; logs/raw outputs are left ignored.