7.2 KiB
7.2 KiB
Benchmark Workflow & Directory Conventions
Directory Layout
/data/user1/yy/
├── scripts/ # all benchmark/orchestrator/utility scripts
│ ├── benchmark_dspark_0707/ # DSpark benchmark suite
│ ├── benchmark_dsv4_backend_comparison.sh
│ ├── start_dsv4_dspark_8card.sh
│ ├── start_sglang_dsv4_8card.sh
│ └── ...
├── bench_results/ # all benchmark outputs and reports
│ ├── dsv4_backend_comparison_20260707/
│ │ ├── raw_outputs/ # JSONL raw outputs
│ │ ├── logs/ # per-run logs
│ │ └── README.md # output manifest + provenance
│ ├── dspark_grid_20260707-132641/
│ ├── dspark_st_comparison_20260707-150649/
│ ├── eagle_grid/
│ └── ...
├── logs/ # server logs (stdout/stderr from start scripts)
├── datasets/ # benchmark datasets
└── envs/ # Python virtual environments
Rules
-
Scripts live in
scripts/only.- Group related scripts into subdirectories, e.g.
scripts/benchmark_dspark_0707/. - Each script group should have its own
README.mdlisting scripts, purpose, and outputs.
- Group related scripts into subdirectories, e.g.
-
Benchmark outputs live in
bench_results/only.- Never leave
.jsonl,.json,.log, or.mdreports in the project root. - Each benchmark run gets its own directory:
bench_results/<experiment>_<timestamp>/. - Raw outputs go in
raw_outputs/. - Logs go in
logs/. - Reports (e.g.
report.md,comparison_report.md) go in the run root.
- Never leave
-
Each
bench_results/<run>/directory must contain two final artifacts.- A Markdown report for human reading (e.g.
report.md,comparison_report.md). - A JSON file with the complete structured result data for programmatic analysis (e.g.
results.json). - The directory may also contain a
README.mddocumenting provenance if the report alone does not cover it.
- A Markdown report for human reading (e.g.
-
Final JSON must contain raw/structured data, not just summary numbers.
- Metadata: experiment name, timestamp, model, backend, hardware, script path, environment/commit info.
- Per-scenario/per-configuration results: all request latencies, TTFT, TPOT, ITL, token counts, throughput, accept length, success/failure counts.
- Include P50 / P90 / P95 / P99 percentiles where applicable.
- Keep the schema stable so downstream Python scripts can parse all experiments uniformly.
- See Final JSON Schema below for the recommended structure.
-
Scripts should default
RESULT_ROOTtobench_results/<experiment>_${RUN_ID}.- Allow override via
RESULT_ROOTenv var. - Use
RUN_ID=$(date '+%Y%m%d-%H%M%S')unless specified.
- Allow override via
-
Server start scripts write to
logs/.logs/<service>_<timestamp>.log- Keep server logs separate from benchmark result logs.
Naming Conventions
Result directories
bench_results/<experiment>_<YYYYMMDD-HHMMSS>/
Examples:
bench_results/dspark_grid_20260707-132641/bench_results/dsv4_backend_comparison_20260707/
Raw output files
For detailed per-request JSONL outputs:
{backend}_{MMDD}_{concurrency}_{input_len}_{output_len}.jsonl
For summary JSON outputs from sglang.bench_serving --output-file:
{backend}_{scenario}_{params}.json
Logs
logs/<service>_YYYYMMDD_HHMMSS.log
logs/<experiment>_orchestrator_YYYYMMDD_HHMMSS.log
Final JSON Schema
The JSON file inside each bench_results/<run>/ directory should follow a stable schema so that downstream Python scripts can load every experiment the same way. The file is usually named results.json.
Required top-level fields
{
"metadata": {
"experiment": "dspark_grid",
"run_id": "20260707-132641",
"timestamp": "2026-07-07T13:26:41+08:00",
"model": "/data/models/DeepSeek-V4-Flash-DSpark",
"backend": "vllm-dspark",
"script": "scripts/benchmark_dspark_0707/run_dspark_benchmark_grid.sh",
"hardware": "8x NVIDIA H200 143GB",
"env": "/data/user1/yy/envs/vllm-dspark",
"git_commit": "optional git sha",
"description": "optional free-text note"
},
"config": {
"tp": 8,
"kv_cache_dtype": "fp8",
"spec_method": "dspark",
"spec_tokens": 5,
"block_size": 256,
"max_num_seqs": 256,
"extra_args": "--no-disable-hybrid-kv-cache-manager"
},
"scenarios": [
{
"name": "chat_short",
"concurrency": 64,
"input_len": 1000,
"output_len": 256,
"duration_s": 21.27,
"success": 512,
"failed": 0,
"request_throughput": 24.07,
"input_token_throughput": 6391.14,
"output_token_throughput": 3189.39,
"total_token_throughput": 9580.52,
"accept_length": 3.2,
"latencies": {
"e2e_ms": { "mean": 2547.81, "p50": 2400.0, "p90": 4800.0, "p95": 5606.07, "p99": 6543.97 },
"ttft_ms": { "mean": 268.68, "p50": 240.0, "p90": 480.0, "p95": 543.21, "p99": 588.63 },
"tpot_ms": { "mean": 18.24, "p50": 16.0, "p90": 28.0, "p95": 30.39, "p99": 39.57 },
"itl_ms": { "mean": 70.29, "p50": 60.0, "p90": 110.0, "p95": 130.0, "p99": 160.0 }
},
"raw_requests": [
{
"request_id": "uuid-or-index",
"input_tokens": 1000,
"output_tokens": 256,
"e2e_ms": 2500.0,
"ttft_ms": 260.0,
"tpot_ms": 18.0,
"itl_ms": 70.0,
"accept_length": 3.0,
"success": true
}
]
}
]
}
Notes
raw_requestsis optional but recommended when the JSON size is manageable. If a single run produces millions of requests, store per-request data asraw_outputs/*.jsonland keep only aggregated percentiles inresults.json.- Always include P50 / P90 / P95 / P99 for TTFT, TPOT, E2E, and ITL. P95 is the primary SLO metric.
- Keep field names snake_case and consistent across experiments.
- If a metric is not applicable (e.g.
accept_lengthfor non-speculative decoding), set it tonullrather than omitting the key.
Quick Start
Run DSpark grid benchmark
bash scripts/benchmark_dspark_0707/run_dspark_benchmark_grid.sh
Run DSpark spec-tokens comparison
bash scripts/benchmark_dspark_0707/run_dspark_st_comparison.sh
Run SGLang vs vLLM backend comparison
# Start SGLang on port 30000 and vLLM on port 8000, then:
bash scripts/benchmark_dsv4_backend_comparison.sh all
Parse results
/data/user1/yy/envs/sglang/bin/python scripts/benchmark_dspark_0707/parse_results.py \
/data/user1/yy/bench_results/dspark_grid_<run_id>
Checklist Before Committing / Archiving
- No
.jsonl,.json,.log, or.mdfiles left in/data/user1/yy/root. - All outputs moved to
bench_results/<experiment>_<timestamp>/. bench_results/<run>/report.md(or equivalent human-readable.md) exists.bench_results/<run>/results.jsonexists and follows the Final JSON Schema.bench_results/<run>/README.mdexists and documents provenance (or the report itself covers provenance).- Scripts moved to
scripts/(orscripts/<group>/). - Script path references updated after moving.