diff --git a/BENCHMARK_WORKFLOW.md b/BENCHMARK_WORKFLOW.md index c7e22df..f97606e 100644 --- a/BENCHMARK_WORKFLOW.md +++ b/BENCHMARK_WORKFLOW.md @@ -37,17 +37,23 @@ - Logs go in `logs/`. - Reports (e.g. `report.md`, `comparison_report.md`) go in the run root. -3. **Each `bench_results//` directory must contain a `README.md`.** - - What was benchmarked. - - Which script(s) produced the outputs. - - File naming convention. - - Inventory of outputs (table). +3. **Each `bench_results//` directory must contain two final artifacts.** + - A Markdown report for human reading (e.g. `report.md`, `comparison_report.md`). + - A JSON file with the complete structured result data for programmatic analysis (e.g. `results.json`). + - The directory may also contain a `README.md` documenting provenance if the report alone does not cover it. -4. **Scripts should default `RESULT_ROOT` to `bench_results/_${RUN_ID}`.** +4. **Final JSON must contain raw/structured data, not just summary numbers.** + - Metadata: experiment name, timestamp, model, backend, hardware, script path, environment/commit info. + - Per-scenario/per-configuration results: all request latencies, TTFT, TPOT, ITL, token counts, throughput, accept length, success/failure counts. + - Include P50 / P90 / P95 / P99 percentiles where applicable. + - Keep the schema stable so downstream Python scripts can parse all experiments uniformly. + - See [Final JSON Schema](#final-json-schema) below for the recommended structure. + +5. **Scripts should default `RESULT_ROOT` to `bench_results/_${RUN_ID}`.** - Allow override via `RESULT_ROOT` env var. - Use `RUN_ID=$(date '+%Y%m%d-%H%M%S')` unless specified. -5. **Server start scripts write to `logs/`.** +6. **Server start scripts write to `logs/`.** - `logs/_.log` - Keep server logs separate from benchmark result logs. @@ -85,6 +91,80 @@ logs/_YYYYMMDD_HHMMSS.log logs/_orchestrator_YYYYMMDD_HHMMSS.log ``` +## Final JSON Schema + +The JSON file inside each `bench_results//` directory should follow a stable schema so that downstream Python scripts can load every experiment the same way. The file is usually named `results.json`. + +### Required top-level fields + +```json +{ + "metadata": { + "experiment": "dspark_grid", + "run_id": "20260707-132641", + "timestamp": "2026-07-07T13:26:41+08:00", + "model": "/data/models/DeepSeek-V4-Flash-DSpark", + "backend": "vllm-dspark", + "script": "scripts/benchmark_dspark_0707/run_dspark_benchmark_grid.sh", + "hardware": "8x NVIDIA H200 143GB", + "env": "/data/user1/yy/envs/vllm-dspark", + "git_commit": "optional git sha", + "description": "optional free-text note" + }, + "config": { + "tp": 8, + "kv_cache_dtype": "fp8", + "spec_method": "dspark", + "spec_tokens": 5, + "block_size": 256, + "max_num_seqs": 256, + "extra_args": "--no-disable-hybrid-kv-cache-manager" + }, + "scenarios": [ + { + "name": "chat_short", + "concurrency": 64, + "input_len": 1000, + "output_len": 256, + "duration_s": 21.27, + "success": 512, + "failed": 0, + "request_throughput": 24.07, + "input_token_throughput": 6391.14, + "output_token_throughput": 3189.39, + "total_token_throughput": 9580.52, + "accept_length": 3.2, + "latencies": { + "e2e_ms": { "mean": 2547.81, "p50": 2400.0, "p90": 4800.0, "p95": 5606.07, "p99": 6543.97 }, + "ttft_ms": { "mean": 268.68, "p50": 240.0, "p90": 480.0, "p95": 543.21, "p99": 588.63 }, + "tpot_ms": { "mean": 18.24, "p50": 16.0, "p90": 28.0, "p95": 30.39, "p99": 39.57 }, + "itl_ms": { "mean": 70.29, "p50": 60.0, "p90": 110.0, "p95": 130.0, "p99": 160.0 } + }, + "raw_requests": [ + { + "request_id": "uuid-or-index", + "input_tokens": 1000, + "output_tokens": 256, + "e2e_ms": 2500.0, + "ttft_ms": 260.0, + "tpot_ms": 18.0, + "itl_ms": 70.0, + "accept_length": 3.0, + "success": true + } + ] + } + ] +} +``` + +### Notes + +- `raw_requests` is optional but recommended when the JSON size is manageable. If a single run produces millions of requests, store per-request data as `raw_outputs/*.jsonl` and keep only aggregated percentiles in `results.json`. +- Always include **P50 / P90 / P95 / P99** for TTFT, TPOT, E2E, and ITL. P95 is the primary SLO metric. +- Keep field names snake_case and consistent across experiments. +- If a metric is not applicable (e.g. `accept_length` for non-speculative decoding), set it to `null` rather than omitting the key. + ## Quick Start ### Run DSpark grid benchmark @@ -117,6 +197,8 @@ bash scripts/benchmark_dsv4_backend_comparison.sh all - [ ] No `.jsonl`, `.json`, `.log`, or `.md` files left in `/data/user1/yy/` root. - [ ] All outputs moved to `bench_results/_/`. -- [ ] `bench_results//README.md` exists and documents provenance. +- [ ] `bench_results//report.md` (or equivalent human-readable `.md`) exists. +- [ ] `bench_results//results.json` exists and follows the [Final JSON Schema](#final-json-schema). +- [ ] `bench_results//README.md` exists and documents provenance (or the report itself covers provenance). - [ ] Scripts moved to `scripts/` (or `scripts//`). - [ ] Script path references updated after moving. diff --git a/README.md b/README.md index a4c1366..5d01ebc 100644 --- a/README.md +++ b/README.md @@ -7,7 +7,7 @@ | 目录/文件 | 说明 | |---|---| | `scripts/` | 所有 benchmark 脚本、服务启动脚本、结果解析脚本 | -| `bench_results/` | 各次实验的最终报告与 summary JSON | +| `bench_results/` | 各次实验的最终产物:人可读 `.md` 报告 + 结构化 `results.json`(含原始数据与分位统计) | | `BENCHMARK_WORKFLOW.md` | benchmark 目录与命名规范 | | `scripts/SLO_STANDARDS.md` | 推理服务 SLO 标准(TTFT/TPOT) |