sskj/BENCHMARK_WORKFLOW.md

7.2 KiB

Benchmark Workflow & Directory Conventions

Directory Layout

/data/user1/yy/
├── scripts/                    # all benchmark/orchestrator/utility scripts
│   ├── benchmark_dspark_0707/  # DSpark benchmark suite
│   ├── benchmark_dsv4_backend_comparison.sh
│   ├── start_dsv4_dspark_8card.sh
│   ├── start_sglang_dsv4_8card.sh
│   └── ...
├── bench_results/              # all benchmark outputs and reports
│   ├── dsv4_backend_comparison_20260707/
│   │   ├── raw_outputs/        # JSONL raw outputs
│   │   ├── logs/               # per-run logs
│   │   └── README.md           # output manifest + provenance
│   ├── dspark_grid_20260707-132641/
│   ├── dspark_st_comparison_20260707-150649/
│   ├── eagle_grid/
│   └── ...
├── logs/                       # server logs (stdout/stderr from start scripts)
├── datasets/                   # benchmark datasets
└── envs/                       # Python virtual environments

Rules

  1. Scripts live in scripts/ only.

    • Group related scripts into subdirectories, e.g. scripts/benchmark_dspark_0707/.
    • Each script group should have its own README.md listing scripts, purpose, and outputs.
  2. Benchmark outputs live in bench_results/ only.

    • Never leave .jsonl, .json, .log, or .md reports in the project root.
    • Each benchmark run gets its own directory: bench_results/<experiment>_<timestamp>/.
    • Raw outputs go in raw_outputs/.
    • Logs go in logs/.
    • Reports (e.g. report.md, comparison_report.md) go in the run root.
  3. Each bench_results/<run>/ directory must contain two final artifacts.

    • A Markdown report for human reading (e.g. report.md, comparison_report.md).
    • A JSON file with the complete structured result data for programmatic analysis (e.g. results.json).
    • The directory may also contain a README.md documenting provenance if the report alone does not cover it.
  4. Final JSON must contain raw/structured data, not just summary numbers.

    • Metadata: experiment name, timestamp, model, backend, hardware, script path, environment/commit info.
    • Per-scenario/per-configuration results: all request latencies, TTFT, TPOT, ITL, token counts, throughput, accept length, success/failure counts.
    • Include P50 / P90 / P95 / P99 percentiles where applicable.
    • Keep the schema stable so downstream Python scripts can parse all experiments uniformly.
    • See Final JSON Schema below for the recommended structure.
  5. Scripts should default RESULT_ROOT to bench_results/<experiment>_${RUN_ID}.

    • Allow override via RESULT_ROOT env var.
    • Use RUN_ID=$(date '+%Y%m%d-%H%M%S') unless specified.
  6. Server start scripts write to logs/.

    • logs/<service>_<timestamp>.log
    • Keep server logs separate from benchmark result logs.

Naming Conventions

Result directories

bench_results/<experiment>_<YYYYMMDD-HHMMSS>/

Examples:

  • bench_results/dspark_grid_20260707-132641/
  • bench_results/dsv4_backend_comparison_20260707/

Raw output files

For detailed per-request JSONL outputs:

{backend}_{MMDD}_{concurrency}_{input_len}_{output_len}.jsonl

For summary JSON outputs from sglang.bench_serving --output-file:

{backend}_{scenario}_{params}.json

Logs

logs/<service>_YYYYMMDD_HHMMSS.log
logs/<experiment>_orchestrator_YYYYMMDD_HHMMSS.log

Final JSON Schema

The JSON file inside each bench_results/<run>/ directory should follow a stable schema so that downstream Python scripts can load every experiment the same way. The file is usually named results.json.

Required top-level fields

{
  "metadata": {
    "experiment": "dspark_grid",
    "run_id": "20260707-132641",
    "timestamp": "2026-07-07T13:26:41+08:00",
    "model": "/data/models/DeepSeek-V4-Flash-DSpark",
    "backend": "vllm-dspark",
    "script": "scripts/benchmark_dspark_0707/run_dspark_benchmark_grid.sh",
    "hardware": "8x NVIDIA H200 143GB",
    "env": "/data/user1/yy/envs/vllm-dspark",
    "git_commit": "optional git sha",
    "description": "optional free-text note"
  },
  "config": {
    "tp": 8,
    "kv_cache_dtype": "fp8",
    "spec_method": "dspark",
    "spec_tokens": 5,
    "block_size": 256,
    "max_num_seqs": 256,
    "extra_args": "--no-disable-hybrid-kv-cache-manager"
  },
  "scenarios": [
    {
      "name": "chat_short",
      "concurrency": 64,
      "input_len": 1000,
      "output_len": 256,
      "duration_s": 21.27,
      "success": 512,
      "failed": 0,
      "request_throughput": 24.07,
      "input_token_throughput": 6391.14,
      "output_token_throughput": 3189.39,
      "total_token_throughput": 9580.52,
      "accept_length": 3.2,
      "latencies": {
        "e2e_ms": { "mean": 2547.81, "p50": 2400.0, "p90": 4800.0, "p95": 5606.07, "p99": 6543.97 },
        "ttft_ms": { "mean": 268.68, "p50": 240.0, "p90": 480.0, "p95": 543.21, "p99": 588.63 },
        "tpot_ms": { "mean": 18.24, "p50": 16.0, "p90": 28.0, "p95": 30.39, "p99": 39.57 },
        "itl_ms": { "mean": 70.29, "p50": 60.0, "p90": 110.0, "p95": 130.0, "p99": 160.0 }
      },
      "raw_requests": [
        {
          "request_id": "uuid-or-index",
          "input_tokens": 1000,
          "output_tokens": 256,
          "e2e_ms": 2500.0,
          "ttft_ms": 260.0,
          "tpot_ms": 18.0,
          "itl_ms": 70.0,
          "accept_length": 3.0,
          "success": true
        }
      ]
    }
  ]
}

Notes

  • raw_requests is optional but recommended when the JSON size is manageable. If a single run produces millions of requests, store per-request data as raw_outputs/*.jsonl and keep only aggregated percentiles in results.json.
  • Always include P50 / P90 / P95 / P99 for TTFT, TPOT, E2E, and ITL. P95 is the primary SLO metric.
  • Keep field names snake_case and consistent across experiments.
  • If a metric is not applicable (e.g. accept_length for non-speculative decoding), set it to null rather than omitting the key.

Quick Start

Run DSpark grid benchmark

bash scripts/benchmark_dspark_0707/run_dspark_benchmark_grid.sh

Run DSpark spec-tokens comparison

bash scripts/benchmark_dspark_0707/run_dspark_st_comparison.sh

Run SGLang vs vLLM backend comparison

# Start SGLang on port 30000 and vLLM on port 8000, then:
bash scripts/benchmark_dsv4_backend_comparison.sh all

Parse results

/data/user1/yy/envs/sglang/bin/python scripts/benchmark_dspark_0707/parse_results.py \
  /data/user1/yy/bench_results/dspark_grid_<run_id>

Checklist Before Committing / Archiving

  • No .jsonl, .json, .log, or .md files left in /data/user1/yy/ root.
  • All outputs moved to bench_results/<experiment>_<timestamp>/.
  • bench_results/<run>/report.md (or equivalent human-readable .md) exists.
  • bench_results/<run>/results.json exists and follows the Final JSON Schema.
  • bench_results/<run>/README.md exists and documents provenance (or the report itself covers provenance).
  • Scripts moved to scripts/ (or scripts/<group>/).
  • Script path references updated after moving.