docs: 规范 benchmark 结果存储格式,要求每个实验同时产出 .md 报告和结构化 results.json
This commit is contained in:
parent
eb431d3925
commit
7de88b995a
@ -37,17 +37,23 @@
|
|||||||
- Logs go in `logs/`.
|
- Logs go in `logs/`.
|
||||||
- Reports (e.g. `report.md`, `comparison_report.md`) go in the run root.
|
- Reports (e.g. `report.md`, `comparison_report.md`) go in the run root.
|
||||||
|
|
||||||
3. **Each `bench_results/<run>/` directory must contain a `README.md`.**
|
3. **Each `bench_results/<run>/` directory must contain two final artifacts.**
|
||||||
- What was benchmarked.
|
- A Markdown report for human reading (e.g. `report.md`, `comparison_report.md`).
|
||||||
- Which script(s) produced the outputs.
|
- A JSON file with the complete structured result data for programmatic analysis (e.g. `results.json`).
|
||||||
- File naming convention.
|
- The directory may also contain a `README.md` documenting provenance if the report alone does not cover it.
|
||||||
- Inventory of outputs (table).
|
|
||||||
|
|
||||||
4. **Scripts should default `RESULT_ROOT` to `bench_results/<experiment>_${RUN_ID}`.**
|
4. **Final JSON must contain raw/structured data, not just summary numbers.**
|
||||||
|
- Metadata: experiment name, timestamp, model, backend, hardware, script path, environment/commit info.
|
||||||
|
- Per-scenario/per-configuration results: all request latencies, TTFT, TPOT, ITL, token counts, throughput, accept length, success/failure counts.
|
||||||
|
- Include P50 / P90 / P95 / P99 percentiles where applicable.
|
||||||
|
- Keep the schema stable so downstream Python scripts can parse all experiments uniformly.
|
||||||
|
- See [Final JSON Schema](#final-json-schema) below for the recommended structure.
|
||||||
|
|
||||||
|
5. **Scripts should default `RESULT_ROOT` to `bench_results/<experiment>_${RUN_ID}`.**
|
||||||
- Allow override via `RESULT_ROOT` env var.
|
- Allow override via `RESULT_ROOT` env var.
|
||||||
- Use `RUN_ID=$(date '+%Y%m%d-%H%M%S')` unless specified.
|
- Use `RUN_ID=$(date '+%Y%m%d-%H%M%S')` unless specified.
|
||||||
|
|
||||||
5. **Server start scripts write to `logs/`.**
|
6. **Server start scripts write to `logs/`.**
|
||||||
- `logs/<service>_<timestamp>.log`
|
- `logs/<service>_<timestamp>.log`
|
||||||
- Keep server logs separate from benchmark result logs.
|
- Keep server logs separate from benchmark result logs.
|
||||||
|
|
||||||
@ -85,6 +91,80 @@ logs/<service>_YYYYMMDD_HHMMSS.log
|
|||||||
logs/<experiment>_orchestrator_YYYYMMDD_HHMMSS.log
|
logs/<experiment>_orchestrator_YYYYMMDD_HHMMSS.log
|
||||||
```
|
```
|
||||||
|
|
||||||
|
## Final JSON Schema
|
||||||
|
|
||||||
|
The JSON file inside each `bench_results/<run>/` directory should follow a stable schema so that downstream Python scripts can load every experiment the same way. The file is usually named `results.json`.
|
||||||
|
|
||||||
|
### Required top-level fields
|
||||||
|
|
||||||
|
```json
|
||||||
|
{
|
||||||
|
"metadata": {
|
||||||
|
"experiment": "dspark_grid",
|
||||||
|
"run_id": "20260707-132641",
|
||||||
|
"timestamp": "2026-07-07T13:26:41+08:00",
|
||||||
|
"model": "/data/models/DeepSeek-V4-Flash-DSpark",
|
||||||
|
"backend": "vllm-dspark",
|
||||||
|
"script": "scripts/benchmark_dspark_0707/run_dspark_benchmark_grid.sh",
|
||||||
|
"hardware": "8x NVIDIA H200 143GB",
|
||||||
|
"env": "/data/user1/yy/envs/vllm-dspark",
|
||||||
|
"git_commit": "optional git sha",
|
||||||
|
"description": "optional free-text note"
|
||||||
|
},
|
||||||
|
"config": {
|
||||||
|
"tp": 8,
|
||||||
|
"kv_cache_dtype": "fp8",
|
||||||
|
"spec_method": "dspark",
|
||||||
|
"spec_tokens": 5,
|
||||||
|
"block_size": 256,
|
||||||
|
"max_num_seqs": 256,
|
||||||
|
"extra_args": "--no-disable-hybrid-kv-cache-manager"
|
||||||
|
},
|
||||||
|
"scenarios": [
|
||||||
|
{
|
||||||
|
"name": "chat_short",
|
||||||
|
"concurrency": 64,
|
||||||
|
"input_len": 1000,
|
||||||
|
"output_len": 256,
|
||||||
|
"duration_s": 21.27,
|
||||||
|
"success": 512,
|
||||||
|
"failed": 0,
|
||||||
|
"request_throughput": 24.07,
|
||||||
|
"input_token_throughput": 6391.14,
|
||||||
|
"output_token_throughput": 3189.39,
|
||||||
|
"total_token_throughput": 9580.52,
|
||||||
|
"accept_length": 3.2,
|
||||||
|
"latencies": {
|
||||||
|
"e2e_ms": { "mean": 2547.81, "p50": 2400.0, "p90": 4800.0, "p95": 5606.07, "p99": 6543.97 },
|
||||||
|
"ttft_ms": { "mean": 268.68, "p50": 240.0, "p90": 480.0, "p95": 543.21, "p99": 588.63 },
|
||||||
|
"tpot_ms": { "mean": 18.24, "p50": 16.0, "p90": 28.0, "p95": 30.39, "p99": 39.57 },
|
||||||
|
"itl_ms": { "mean": 70.29, "p50": 60.0, "p90": 110.0, "p95": 130.0, "p99": 160.0 }
|
||||||
|
},
|
||||||
|
"raw_requests": [
|
||||||
|
{
|
||||||
|
"request_id": "uuid-or-index",
|
||||||
|
"input_tokens": 1000,
|
||||||
|
"output_tokens": 256,
|
||||||
|
"e2e_ms": 2500.0,
|
||||||
|
"ttft_ms": 260.0,
|
||||||
|
"tpot_ms": 18.0,
|
||||||
|
"itl_ms": 70.0,
|
||||||
|
"accept_length": 3.0,
|
||||||
|
"success": true
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
]
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
### Notes
|
||||||
|
|
||||||
|
- `raw_requests` is optional but recommended when the JSON size is manageable. If a single run produces millions of requests, store per-request data as `raw_outputs/*.jsonl` and keep only aggregated percentiles in `results.json`.
|
||||||
|
- Always include **P50 / P90 / P95 / P99** for TTFT, TPOT, E2E, and ITL. P95 is the primary SLO metric.
|
||||||
|
- Keep field names snake_case and consistent across experiments.
|
||||||
|
- If a metric is not applicable (e.g. `accept_length` for non-speculative decoding), set it to `null` rather than omitting the key.
|
||||||
|
|
||||||
## Quick Start
|
## Quick Start
|
||||||
|
|
||||||
### Run DSpark grid benchmark
|
### Run DSpark grid benchmark
|
||||||
@ -117,6 +197,8 @@ bash scripts/benchmark_dsv4_backend_comparison.sh all
|
|||||||
|
|
||||||
- [ ] No `.jsonl`, `.json`, `.log`, or `.md` files left in `/data/user1/yy/` root.
|
- [ ] No `.jsonl`, `.json`, `.log`, or `.md` files left in `/data/user1/yy/` root.
|
||||||
- [ ] All outputs moved to `bench_results/<experiment>_<timestamp>/`.
|
- [ ] All outputs moved to `bench_results/<experiment>_<timestamp>/`.
|
||||||
- [ ] `bench_results/<run>/README.md` exists and documents provenance.
|
- [ ] `bench_results/<run>/report.md` (or equivalent human-readable `.md`) exists.
|
||||||
|
- [ ] `bench_results/<run>/results.json` exists and follows the [Final JSON Schema](#final-json-schema).
|
||||||
|
- [ ] `bench_results/<run>/README.md` exists and documents provenance (or the report itself covers provenance).
|
||||||
- [ ] Scripts moved to `scripts/` (or `scripts/<group>/`).
|
- [ ] Scripts moved to `scripts/` (or `scripts/<group>/`).
|
||||||
- [ ] Script path references updated after moving.
|
- [ ] Script path references updated after moving.
|
||||||
|
|||||||
@ -7,7 +7,7 @@
|
|||||||
| 目录/文件 | 说明 |
|
| 目录/文件 | 说明 |
|
||||||
|---|---|
|
|---|---|
|
||||||
| `scripts/` | 所有 benchmark 脚本、服务启动脚本、结果解析脚本 |
|
| `scripts/` | 所有 benchmark 脚本、服务启动脚本、结果解析脚本 |
|
||||||
| `bench_results/` | 各次实验的最终报告与 summary JSON |
|
| `bench_results/` | 各次实验的最终产物:人可读 `.md` 报告 + 结构化 `results.json`(含原始数据与分位统计) |
|
||||||
| `BENCHMARK_WORKFLOW.md` | benchmark 目录与命名规范 |
|
| `BENCHMARK_WORKFLOW.md` | benchmark 目录与命名规范 |
|
||||||
| `scripts/SLO_STANDARDS.md` | 推理服务 SLO 标准(TTFT/TPOT) |
|
| `scripts/SLO_STANDARDS.md` | 推理服务 SLO 标准(TTFT/TPOT) |
|
||||||
|
|
||||||
|
|||||||
Loading…
x
Reference in New Issue
Block a user