# Benchmark Workflow & Directory Conventions ## Directory Layout ``` /data/user1/yy/ ├── platforms/ # chip/accelerator platform configs │ ├── kunlun_p800.env │ ├── nvidia_h200.env │ └── patches/kunlun_p800/ # runtime patches required by some images ├── scripts/ # shared benchmark/orchestrator/utility scripts │ ├── common/ # reusable components (lib.sh, platform.sh, parse_backend.py, ...) │ ├── benchmark_dspark_0707/ # legacy DSpark benchmark suite (仍被 experiments/dsv4_h200_dspark/ 引用) │ └── SLO_STANDARDS.md # 推理服务 SLO 标准 ├── experiments/ # experiment-centric directories (preferred) │ └── dsv4_p800_sglang/ │ ├── README.md │ ├── config.env # experiment-level configuration │ ├── start_server.sh │ ├── run_bench.sh │ ├── parse_results.py │ └── results/ │ └── 20260708-XXXXXX/ │ ├── report.md │ ├── results.json │ └── logs/ ├── bench_results/ # legacy benchmark outputs (read-only history) │ ├── dspark_grid_20260707-132641/ │ └── ... ├── logs/ # server logs (stdout/stderr from start scripts) ├── datasets/ # benchmark datasets └── envs/ # Python virtual environments ``` ## Rules 1. **Experiments are the primary organization unit.** - Each experiment lives under `experiments//` and contains its scripts, configuration, and results. - Shared orchestration code lives in `scripts/common/`; do not copy server start / health check logic into every experiment. - 新实验必须放在 `experiments//` 下;`scripts/` 下仅保留 `scripts/common/` 和被现有实验引用的 legacy 脚本。 2. **Benchmark outputs live with their experiment.** - For `experiments//`, results go in `experiments//results//`. - Each run directory must contain `report.md` (human-readable) and `results.json` (structured data). - Raw outputs go in `raw_outputs/`; logs go in `logs/`. - Legacy `bench_results/_/` directories remain valid for archived runs. 3. **Each `bench_results//` directory must contain two final artifacts.** - A Markdown report for human reading (e.g. `report.md`, `comparison_report.md`). - A JSON file with the complete structured result data for programmatic analysis (e.g. `results.json`). - The directory may also contain a `README.md` documenting provenance if the report alone does not cover it. 4. **Final JSON must contain raw/structured data, not just summary numbers.** - Metadata: experiment name, timestamp, model, backend/inference engine, hardware/accelerator, script path, environment/commit info. - Record the **chip/accelerator** (e.g. `NVIDIA H200`, `Kunlun XPU`) and the **inference engine** (e.g. `vllm-dspark`, `sglang`, `vllm-xpu`) explicitly. Do not infer them from directory names. - Per-scenario/per-configuration results: all request latencies, TTFT, TPOT, ITL, token counts, throughput, accept length, success/failure counts. - Include P50 / P90 / P95 / P99 percentiles where applicable. - Keep the schema stable so downstream Python scripts can parse all experiments uniformly. - See [Final JSON Schema](#final-json-schema) below for the recommended structure. 5. **Scripts should default `RESULT_ROOT` to the experiment's results directory.** - For `experiments//run_bench.sh`, default to `experiments//results/${RUN_ID}/`. - For legacy scripts, default to `bench_results/_${RUN_ID}/`. - Allow override via `RESULT_ROOT` env var. - Use `RUN_ID=$(date '+%Y%m%d-%H%M%S')` unless specified. 6. **Server start scripts write to `logs/`.** - `logs/_.log` - Keep server logs separate from benchmark result logs. 7. **Scripts and outputs must record chip/accelerator and inference engine.** - Every benchmark script should capture or accept the platform and engine it is running on (e.g. via environment variables `CHIP`, `ACCELERATOR`, `ENGINE`, `BACKEND`, or auto-detection). - Final reports and JSON outputs must include both the accelerator/chip family and the inference engine/backend used for the run. - Do not rely on the experiment name alone to identify the platform or engine. 8. **Use platform configuration files for chip-specific constants.** - Put per-platform settings in `platforms/.env` (e.g. `platforms/kunlun_p800.env`, `platforms/nvidia_h200.env`). - Scripts load the platform file via `scripts/common/platform.sh`; the active platform is selected by the `PLATFORM` env var or auto-detected. - Keep experiment scripts free of hardcoded device IDs, image names, or model root paths. 9. **Record the exact server launch command/args for every run.** - The exact command or full argument list used to start the server must be saved in `results.json` under `config.server_args` (or `config.phaseN_server_args` if the experiment starts the server in multiple phases). - This is required for cross-platform reproduction: when the same experiment is run on H200 and P800, the only differences should be model paths, ports, and device IDs. - If an experiment uses `start_server.sh`, that script should be self-contained and its command line should be reproducible from `results.json` alone. ## Naming Conventions ### Experiment result directories For experiment-centric layout: ``` experiments//results// ``` Examples: - `experiments/dsv4_p800_sglang/results/20260708-120000/` - `experiments/dspark_grid/results/20260707-132641/` The `results.json` metadata already records `chip`/`accelerator` and `engine`, so the directory path does not need to encode them. If a single experiment must distinguish across platforms in its directory tree, use: ``` experiments//results/__/ ``` ### Legacy result directories ``` bench_results/_/ ``` Examples: - `experiments/legacy_bench_results/dspark_grid_20260707-132641/` (legacy result moved under `experiments/`) - `experiments/legacy_bench_results/dsv4_backend_comparison_20260707/` (legacy result moved under `experiments/`) ### Raw output files Include the accelerator and inference engine in raw output filenames so files from different platforms cannot overwrite each other. For detailed per-request JSONL outputs: ``` {chip}_{engine}_{MMDD}_{concurrency}_{input_len}_{output_len}.jsonl ``` For summary JSON outputs from `sglang.bench_serving --output-file`: ``` {chip}_{engine}_{scenario}_{params}.json ``` ### Logs ``` logs/_YYYYMMDD_HHMMSS.log logs/_orchestrator_YYYYMMDD_HHMMSS.log ``` ## Final JSON Schema The JSON file inside each `bench_results//` directory should follow a stable schema so that downstream Python scripts can load every experiment the same way. The file is usually named `results.json`. ### Required top-level fields ```json { "metadata": { "experiment": "dspark_grid", "run_id": "20260707-132641", "timestamp": "2026-07-07T13:26:41+08:00", "model": "/data/models/DeepSeek-V4-Flash-DSpark", "backend": "vllm-dspark", "engine": "vllm-dspark", "hardware": "8x NVIDIA H200 143GB", "accelerator": "NVIDIA H200", "chip": "NVIDIA H200", "script": "scripts/benchmark_dspark_0707/run_dspark_benchmark_grid.sh", "env": "/data/user1/yy/envs/vllm-dspark", "git_commit": "optional git sha", "description": "optional free-text note" }, "config": { "tp": 8, "kv_cache_dtype": "fp8", "spec_method": "dspark", "spec_tokens": 5, "block_size": 256, "max_num_seqs": 256, "extra_args": "--no-disable-hybrid-kv-cache-manager", "server_args": "vllm serve /data/models/DeepSeek-V4-Flash-DSpark --trust-remote-code --tensor-parallel-size 8 --kv-cache-dtype fp8 --max-model-len auto --max-num-seqs 256 --spec-method dspark --spec-tokens 5 --no-disable-hybrid-kv-cache-manager --port 30004" }, "scenarios": [ { "name": "chat_short", "concurrency": 64, "input_len": 1000, "output_len": 256, "duration_s": 21.27, "success": 512, "failed": 0, "request_throughput": 24.07, "input_token_throughput": 6391.14, "output_token_throughput": 3189.39, "total_token_throughput": 9580.52, "accept_length": 3.2, "latencies": { "e2e_ms": { "mean": 2547.81, "p50": 2400.0, "p90": 4800.0, "p95": 5606.07, "p99": 6543.97 }, "ttft_ms": { "mean": 268.68, "p50": 240.0, "p90": 480.0, "p95": 543.21, "p99": 588.63 }, "tpot_ms": { "mean": 18.24, "p50": 16.0, "p90": 28.0, "p95": 30.39, "p99": 39.57 }, "itl_ms": { "mean": 70.29, "p50": 60.0, "p90": 110.0, "p95": 130.0, "p99": 160.0 } }, "slo_status": { "ttft_p95_ok": true, "tpot_mean_ok": true, "overall": "✅" }, "raw_requests": [ { "request_id": "uuid-or-index", "input_tokens": 1000, "output_tokens": 256, "e2e_ms": 2500.0, "ttft_ms": 260.0, "tpot_ms": 18.0, "itl_ms": 70.0, "accept_length": 3.0, "success": true } ] } ] } ``` ### Notes - `raw_requests` is optional but recommended when the JSON size is manageable. If a single run produces millions of requests, store per-request data as `raw_outputs/*.jsonl` and keep only aggregated percentiles in `results.json`. - Always include **P50 / P90 / P95 / P99** for TTFT, TPOT, E2E, and ITL. P95 is the primary SLO metric. - Include an `slo_status` object per scenario indicating whether the scenario meets the relevant SLO (e.g. S2 tier: TTFT P95 < 3000ms, TPOT mean < 50ms). Example: ```json "slo_status": { "ttft_p95_ok": true, "tpot_mean_ok": true, "overall": "✅" } ``` - Keep field names snake_case and consistent across experiments. - Record hardware and engine information explicitly: - `accelerator` / `chip`: the accelerator family, e.g. `NVIDIA H200`, `Kunlun XPU`, `AMD MI300X`. - `engine` / `backend`: the inference engine or serving backend, e.g. `vllm-dspark`, `sglang`, `vllm-xpu`. - `hardware`: a human-readable full hardware description, e.g. `8x NVIDIA H200 143GB`, `8x Kunlun XPU R480`. - Keep at least one of `accelerator` or `chip`, and at least one of `engine` or `backend`, populated in every run. - If a metric is not applicable (e.g. `accept_length` for non-speculative decoding), set it to `null` rather than omitting the key. ## Quick Start ### Run P800 SGLang benchmark (Kunlun P800, Docker) ```bash bash experiments/dsv4_p800_sglang/run_bench.sh ``` ### Run H200 DSpark benchmark (NVIDIA H200, native venv) ```bash bash experiments/dsv4_h200_dspark/run_bench.sh ``` ### Run legacy DSpark grid benchmark ```bash bash scripts/benchmark_dspark_0707/run_dspark_benchmark_grid.sh ``` ### Run DSpark spec-tokens comparison ```bash bash scripts/benchmark_dspark_0707/run_dspark_st_comparison.sh ``` ### Parse results ```bash # Experiment-centric layout python3 experiments/dsv4_p800_sglang/parse_results.py \ experiments/dsv4_p800_sglang/results/ # Legacy layout /data/user1/yy/envs/sglang/bin/python scripts/benchmark_dspark_0707/parse_results.py \ /data/user1/yy/experiments/legacy_bench_results/dspark_grid_ ``` ## Adding a New Platform or Experiment ### 1. Add or update a platform config Create `platforms/.env` with identity and platform-wide paths: ```bash CHIP="my_chip" ACCELERATOR="My Accelerator" HARDWARE="8x My Accelerator" ENGINE="vllm-myengine" DEFAULT_PORT="30000" MODEL_ROOT="/data/models" ``` For Docker-based platforms, also set `DOCKER_IMAGE`, `CONTAINER_NAME`, `CONTAINER_PYTHON`, and `PATCH_ROOT` (see `platforms/kunlun_p800.env`). For host-native platforms, set the relevant venv paths (see `platforms/nvidia_h200.env`). ### 2. Create an experiment directory ``` experiments// ├── README.md # Purpose and usage ├── config.env # Model, port, scenarios, engine overrides ├── start_server.sh # (optional) platform-specific server launch ├── run_bench.sh # Orchestrator └── parse_results.py # Convert raw outputs to results.json + report.md ``` At minimum, `run_bench.sh` should: 1. Source `scripts/common/lib.sh` and `scripts/common/platform.sh`. 2. Read `config.env`. 3. Create `experiments//results//`. 4. Call `write_metadata_json` to create `results.json`. 5. Record the exact server launch command/args in `results.json` `config.server_args` (or `phaseN_server_args`). 6. Run the benchmark scenarios. 7. Call `parse_results.py` to generate `report.md`. ### 3. Reuse shared helpers - `scripts/common/lib.sh`: logging, health checks, metadata JSON. - `scripts/common/platform.sh`: platform auto-detection and env loading. - `scripts/common/warmup.py`: server 预热。 - `scripts/common/parse_backend.py`: raw jsonl -> `results.json` + `report.md`。 - `scripts/common/compare.py`: SGLang vs vLLM 横向对比表。 Docker/XPU 平台可以在 `scripts/common/` 下新增自己的 helper,但不要在实验目录复制通用逻辑。 ### 4. Example experiments - Docker / XPU: `experiments/dsv4_p800_sglang/` - Native / H200: `experiments/dsv4_h200_dspark/` See `docs/H200_QUICKSTART.md` for a concrete H200 porting walkthrough. ## Checklist Before Committing / Archiving - [ ] No `.jsonl`, `.json`, `.log`, or `.md` files left in the project root. - [ ] For `experiments//`, outputs live in `experiments//results//`. - [ ] For legacy runs, outputs live in `bench_results/_/`. - [ ] `/report.md` (or equivalent human-readable `.md`) exists. - [ ] `/results.json` exists and follows the [Final JSON Schema](#final-json-schema). - [ ] `/results.json` metadata records the `chip`/`accelerator` and `engine`/`backend` used. - [ ] `/results.json` `config` records the exact server launch command/args (`server_args` or `phaseN_server_args`). - [ ] `/results.json` each scenario records `slo_status` (overall pass/fail/partial against the relevant SLO). - [ ] Raw output filenames include the chip/accelerator and engine when cross-platform runs may collide. - [ ] `/README.md` exists and documents provenance (or the report itself covers provenance). - [ ] Scripts either live under `experiments//` or in `scripts/` (or `scripts//`). - [ ] Script path references updated after moving.