refactor: relocate old docs and add dsv4_h200_vllm experiment
- Move dspark_deepseekv4_fix_pr_prep.md into experiments/dsv4_h200_dspark/ - Move dsv4_inference_comparison_report.md into docs/ - Delete obsolete cleanup_summary.md - Add experiments/dsv4_h200_vllm/ baseline experiment (envs/vllm + envs/sglang)
This commit is contained in:
parent
07835b9fed
commit
8e15ffe1c7
@ -1,110 +0,0 @@
|
||||
# /data/user1/yy 文件夹整理记录
|
||||
|
||||
> 整理时间:2026-07-06
|
||||
> 说明:将脚本集中到 `scripts/`,uv 虚拟环境集中到 `envs/`,并保留兼容性软链接;删除过时的 log、pid、临时文件和冗余压缩包。
|
||||
|
||||
---
|
||||
|
||||
## 目录结构
|
||||
|
||||
```
|
||||
/data/user1/yy/
|
||||
├── bench_results/
|
||||
│ ├── dsv4_comparison_20260705_152221/
|
||||
│ ├── dsv4_flash_dspark_misc/
|
||||
│ ├── pd_bench/
|
||||
│ ├── sglang_8card_max_throughput_20260705_030839/
|
||||
│ ├── sglang_8card_systematic_20260704_120819/
|
||||
│ ├── sglang_misc/
|
||||
│ ├── vllm_dspark_qwen3_20260705_121046/
|
||||
│ ├── vllm_dspark_qwen3_20260705_121218/
|
||||
│ └── vllm_dspark_qwen3_20260705_121256/
|
||||
├── datasets/
|
||||
├── envs/ # uv 虚拟环境与缓存
|
||||
│ ├── sglang/ -> 原 /data/user1/yy/sglang
|
||||
│ ├── spraseattn/ -> 原 /data/user1/yy/spraseattn
|
||||
│ ├── uv_cache/ -> 原 /data/user1/yy/uv_cache
|
||||
│ ├── vllm/ -> 原 /data/user1/yy/vllm
|
||||
│ └── vllm-dspark/ -> 原 /data/user1/yy/vllm-dspark
|
||||
├── loomeval_yy/ # 按用户要求保留
|
||||
├── scripts/ # 脚本集中存放
|
||||
│ ├── bench_dsv4_comparison.py
|
||||
│ ├── bench_vllm_dspark_qwen3.py
|
||||
│ ├── install_vllm_dspark.sh
|
||||
│ ├── run_sglang_benchmark.sh
|
||||
│ ├── run_sglang_max_throughput.sh
|
||||
│ ├── start_dsv4_dspark_8card.sh
|
||||
│ ├── start_dsv4_dspark_8card_bf16kv.sh
|
||||
│ ├── start_dsv4_dspark_8card_flashinfer.sh
|
||||
│ ├── start_pd_single_node.sh
|
||||
│ ├── start_sglang_dsv4_8card.sh
|
||||
│ ├── start_vllm_pd_single_node.sh
|
||||
│ └── test_block_sparse_attn.py
|
||||
├── tmp/
|
||||
├── vllm-main/ # vLLM 源码目录
|
||||
├── cleanup_summary.md
|
||||
├── dspark_deepseekv4_fix_pr_prep.md
|
||||
├── dsv4_inference_comparison_report.md
|
||||
├── issue_47648_comment.json
|
||||
└── issue_47648_comment.md
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## 脚本说明
|
||||
|
||||
| 脚本 | 说明 |
|
||||
|---|---|
|
||||
| `scripts/bench_dsv4_comparison.py` | vLLM DeepSeek-V4 对比测试主控脚本 |
|
||||
| `scripts/bench_vllm_dspark_qwen3.py` | Qwen3 DSpark 验证脚本 |
|
||||
| `scripts/install_vllm_dspark.sh` | vllm-dspark 安装脚本 |
|
||||
| `scripts/run_sglang_benchmark.sh` | SGLang 8卡系统测试脚本 |
|
||||
| `scripts/run_sglang_max_throughput.sh` | SGLang 极限吞吐测试脚本 |
|
||||
| `scripts/start_dsv4_dspark_8card.sh` | DSV4 DSpark 服务启动脚本 |
|
||||
| `scripts/start_dsv4_dspark_8card_bf16kv.sh` | DSV4 DSpark bf16 KV 启动脚本 |
|
||||
| `scripts/start_dsv4_dspark_8card_flashinfer.sh` | DSV4 DSpark FlashInfer 启动脚本 |
|
||||
| `scripts/start_pd_single_node.sh` | PD 分离单节点启动脚本 |
|
||||
| `scripts/start_sglang_dsv4_8card.sh` | SGLang DSV4 启动脚本 |
|
||||
| `scripts/start_vllm_pd_single_node.sh` | vLLM PD 分离单节点启动脚本 |
|
||||
| `scripts/test_block_sparse_attn.py` | block sparse attention 测试脚本 |
|
||||
|
||||
---
|
||||
|
||||
## 删除的文件/目录
|
||||
|
||||
| 路径 | 原因 |
|
||||
|---|---|
|
||||
| `__pycache__/` | Python 字节码缓存,可重新生成 |
|
||||
| `*.pid`(共 4 个) | 进程 ID 临时文件 |
|
||||
| `logs/` 目录下所有文件 | 安装、部署、下载等历史调试日志 |
|
||||
| `pd_logs/` 目录 | PD 分离测试日志 |
|
||||
| `vllm_pd_logs/` 目录 | vLLM PD 分离测试日志 |
|
||||
| `vllm-main.zip` | 与 `vllm-main/` 目录冗余,约 41MB |
|
||||
| `tmp/` 目录下内容 | 临时文件 |
|
||||
| 根目录下的 `.log` 文件 | 已归档或无需保留的日志 |
|
||||
|
||||
---
|
||||
|
||||
## 归档到 `bench_results/` 的文件
|
||||
|
||||
| 原路径 | 目标路径 | 说明 |
|
||||
|---|---|---|
|
||||
| `bench_dsv4_comparison_master.log` | `bench_results/dsv4_comparison_20260705_152221/` | 本次对比测试主日志 |
|
||||
| `sglang_8card.log` | `bench_results/sglang_misc/` | SGLang 8卡日志 |
|
||||
| `sglang_0704_1_1048576_1.jsonl` | `bench_results/sglang_misc/` | SGLang 测试原始结果 |
|
||||
| `sglang_8card_sharegpt_c*.json/log` | `bench_results/sglang_8card_systematic_20260704_120819/` | SGLang 系统测试散落结果 |
|
||||
| `pd_bench_*.json` | `bench_results/pd_bench/` | PD 分离 benchmark 结果 |
|
||||
| `dsv4_flash_dspark_c16_n50_*` | `bench_results/dsv4_flash_dspark_misc/` | DSV4 DSpark 早期单次测试 |
|
||||
|
||||
---
|
||||
|
||||
## 路径更新说明
|
||||
|
||||
脚本和文档中的环境路径已统一更新为 `envs/` 下的新路径:
|
||||
|
||||
- `/data/user1/yy/vllm-dspark/` → `/data/user1/yy/envs/vllm-dspark/`
|
||||
- `/data/user1/yy/vllm/` → `/data/user1/yy/envs/vllm/`
|
||||
- `/data/user1/yy/sglang/` → `/data/user1/yy/envs/sglang/`
|
||||
- `/data/user1/yy/spraseattn/` → `/data/user1/yy/envs/spraseattn/`
|
||||
|
||||
根目录不再保留兼容性软链接,所有环境入口统一通过 `envs/` 访问。
|
||||
43
experiments/dsv4_h200_vllm/README.md
Normal file
43
experiments/dsv4_h200_vllm/README.md
Normal file
@ -0,0 +1,43 @@
|
||||
# DSV4 H200 vLLM Baseline Benchmark
|
||||
|
||||
NVIDIA H200 + native `vllm` + `DeepSeek-V4-Flash` baseline benchmark experiment.
|
||||
|
||||
This experiment uses:
|
||||
- `/data/user1/yy/envs/vllm` as the server environment
|
||||
- `/data/user1/yy/envs/sglang` as the benchmark client environment
|
||||
|
||||
## Quick Start
|
||||
|
||||
```bash
|
||||
# Run the full experiment (start server + benchmark + parse)
|
||||
bash experiments/dsv4_h200_vllm/run_bench.sh
|
||||
|
||||
# Reuse an already-running server
|
||||
SKIP_MANAGE_SERVER=1 bash experiments/dsv4_h200_vllm/run_bench.sh
|
||||
```
|
||||
|
||||
Results land in `experiments/dsv4_h200_vllm/results/<RUN_ID>/`.
|
||||
|
||||
## Configuration
|
||||
|
||||
Edit `config.env` or override via environment variables:
|
||||
|
||||
```bash
|
||||
MODEL_PATH=/data/models/DeepSeek-V4-Flash \
|
||||
PORT=30005 \
|
||||
SCENARIOS="32 512 256 128 512 256" \
|
||||
bash experiments/dsv4_h200_vllm/run_bench.sh
|
||||
```
|
||||
|
||||
## Files
|
||||
|
||||
| File | Purpose |
|
||||
|---|---|
|
||||
| `config.env` | Experiment-level configuration (model, port, venv paths, scenarios) |
|
||||
| `start_server.sh` | Start a plain vLLM baseline server for DeepSeek-V4-Flash |
|
||||
| `run_bench.sh` | Orchestrator: metadata → server → benchmark → parse |
|
||||
| `parse_results.py` | Parse raw JSONL outputs into `results.json` + `report.md` |
|
||||
|
||||
## Platform
|
||||
|
||||
This experiment targets `platforms/nvidia_h200.env`.
|
||||
24
experiments/dsv4_h200_vllm/config.env
Normal file
24
experiments/dsv4_h200_vllm/config.env
Normal file
@ -0,0 +1,24 @@
|
||||
# Experiment-level configuration for dsv4_h200_vllm
|
||||
# All values can be overridden via environment variables.
|
||||
|
||||
EXPERIMENT="${EXPERIMENT:-dsv4_h200_vllm}"
|
||||
MODEL_NAME="${MODEL_NAME:-DeepSeek-V4-Flash}"
|
||||
MODEL_PATH="${MODEL_PATH:-/data/models/DeepSeek-V4-Flash}"
|
||||
SERVED_MODEL_NAME="${SERVED_MODEL_NAME:-deepseek-v4-flash}"
|
||||
PORT="${PORT:-30005}"
|
||||
# This experiment is a plain vLLM baseline. These values override the
|
||||
# vllm-dspark defaults coming from platforms/nvidia_h200.env.
|
||||
BACKEND="vllm"
|
||||
ENGINE="vllm"
|
||||
|
||||
# Native virtual environments on the host.
|
||||
VENV_SERVER="${VENV_SERVER:-/data/user1/yy/envs/vllm}"
|
||||
VENV_CLIENT="${VENV_CLIENT:-/data/user1/yy/envs/sglang}"
|
||||
|
||||
# Benchmark scenarios: "concurrency input_len output_len".
|
||||
# Override via SCENARIOS env var.
|
||||
SCENARIOS="${SCENARIOS:-32 512 256 128 512 256 32 4000 512}"
|
||||
NUM_PROMPTS="${NUM_PROMPTS:-128}"
|
||||
|
||||
# Server start script bundled with this experiment.
|
||||
SERVER_START_SCRIPT="${SCRIPT_DIR}/start_server.sh"
|
||||
193
experiments/dsv4_h200_vllm/parse_results.py
Executable file
193
experiments/dsv4_h200_vllm/parse_results.py
Executable file
@ -0,0 +1,193 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Parse H200 vLLM baseline benchmark JSONL outputs.
|
||||
|
||||
Reads raw JSONL files produced by `sglang.bench_serving --output-file --output-details`
|
||||
and generates:
|
||||
- results.json (appended scenarios)
|
||||
- report.md (human-readable summary)
|
||||
|
||||
Usage:
|
||||
python3 parse_results.py <result_root>
|
||||
"""
|
||||
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
|
||||
def percentile(values: list[float], p: float) -> float:
|
||||
if not values:
|
||||
return 0.0
|
||||
values = sorted(values)
|
||||
if len(values) == 1:
|
||||
return values[0]
|
||||
k = (len(values) - 1) * (p / 100.0)
|
||||
f = math.floor(k)
|
||||
c = math.ceil(k)
|
||||
if f == c:
|
||||
return values[int(k)]
|
||||
return values[f] * (c - k) + values[c] * (k - f)
|
||||
|
||||
|
||||
def parse_jsonl(path: Path) -> list[dict]:
|
||||
requests = []
|
||||
with open(path, "r", encoding="utf-8") as f:
|
||||
for line in f:
|
||||
line = line.strip()
|
||||
if not line:
|
||||
continue
|
||||
try:
|
||||
requests.append(json.loads(line))
|
||||
except json.JSONDecodeError:
|
||||
continue
|
||||
return requests
|
||||
|
||||
|
||||
def compute_metrics(requests: list[dict]) -> dict:
|
||||
success_reqs = [r for r in requests if r.get("success", True)]
|
||||
failed = len(requests) - len(success_reqs)
|
||||
|
||||
if not success_reqs:
|
||||
return {"success": 0, "failed": failed}
|
||||
|
||||
# Time boundaries.
|
||||
start_times = [r["tstamp_start"] for r in success_reqs]
|
||||
end_times = [r["tstamp_finished"] for r in success_reqs]
|
||||
duration_s = max(end_times) - min(start_times)
|
||||
|
||||
# Token counts.
|
||||
input_tokens = [r.get("prompt_tokens", 0) for r in success_reqs]
|
||||
output_tokens = [r.get("completion_tokens", 0) for r in success_reqs]
|
||||
total_input = sum(input_tokens)
|
||||
total_output = sum(output_tokens)
|
||||
|
||||
# Latencies (ms).
|
||||
e2e = [r.get("e2e_latency", 0) * 1000 for r in success_reqs]
|
||||
ttft = [r.get("ttft", 0) * 1000 for r in success_reqs]
|
||||
itls = []
|
||||
for r in success_reqs:
|
||||
itls.extend(r.get("itl", []))
|
||||
# TPOT from itl averages per request.
|
||||
tpots = []
|
||||
for r in success_reqs:
|
||||
req_itls = r.get("itl", [])
|
||||
if req_itls:
|
||||
tpots.append(sum(req_itls) / len(req_itls) * 1000)
|
||||
|
||||
def latency_stats(values: list[float]) -> dict:
|
||||
return {
|
||||
"mean": sum(values) / len(values),
|
||||
"p50": percentile(values, 50),
|
||||
"p90": percentile(values, 90),
|
||||
"p95": percentile(values, 95),
|
||||
"p99": percentile(values, 99),
|
||||
}
|
||||
|
||||
return {
|
||||
"success": len(success_reqs),
|
||||
"failed": failed,
|
||||
"duration_s": duration_s,
|
||||
"request_throughput": len(success_reqs) / duration_s if duration_s > 0 else 0.0,
|
||||
"input_token_throughput": total_input / duration_s if duration_s > 0 else 0.0,
|
||||
"output_token_throughput": total_output / duration_s if duration_s > 0 else 0.0,
|
||||
"total_token_throughput": (total_input + total_output) / duration_s if duration_s > 0 else 0.0,
|
||||
"total_input_tokens": total_input,
|
||||
"total_output_tokens": total_output,
|
||||
"e2e_ms": latency_stats(e2e),
|
||||
"ttft_ms": latency_stats(ttft),
|
||||
"tpot_ms": latency_stats(tpots),
|
||||
"itl_ms": latency_stats([v * 1000 for v in itls]),
|
||||
}
|
||||
|
||||
|
||||
def scenario_name(concurrency: int, input_len: int, output_len: int) -> str:
|
||||
return f"c{concurrency}_i{input_len}_o{output_len}"
|
||||
|
||||
|
||||
def append_scenario(results_json: Path, scenario: dict) -> None:
|
||||
with open(results_json, "r", encoding="utf-8") as f:
|
||||
data = json.load(f)
|
||||
data["scenarios"].append(scenario)
|
||||
with open(results_json, "w", encoding="utf-8") as f:
|
||||
json.dump(data, f, indent=2, ensure_ascii=False)
|
||||
|
||||
|
||||
def generate_report(result_root: Path, scenarios: list[dict]) -> None:
|
||||
report_path = result_root / "report.md"
|
||||
with open(report_path, "w", encoding="utf-8") as f:
|
||||
f.write("# H200 vLLM Baseline Benchmark Report\n\n")
|
||||
f.write(f"- Result root: `{result_root}`\n")
|
||||
f.write("- Model: `/data/models/DeepSeek-V4-Flash`\n")
|
||||
f.write("- Backend: vLLM (TP=8, FP8 KV cache, no speculative decoding)\n")
|
||||
f.write("- Benchmark client: `sglang.bench_serving --backend vllm`\n\n")
|
||||
|
||||
f.write("## Results\n\n")
|
||||
f.write("| Scenario | Concurrency | Input | Output | Duration(s) | Success | Req/s | In tok/s | Out tok/s | Total tok/s | Mean TTFT(ms) | P95 TTFT(ms) | P99 TTFT(ms) | Mean TPOT(ms) | P95 TPOT(ms) | P99 TPOT(ms) | Mean E2E(ms) | P95 E2E(ms) | P99 E2E(ms) |\n")
|
||||
f.write("|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|\n")
|
||||
|
||||
for s in scenarios:
|
||||
cfg = s["config"]
|
||||
m = s["metrics"]
|
||||
f.write(
|
||||
f"| {cfg['name']} | {cfg['concurrency']} | {cfg['input_len']} | {cfg['output_len']} | "
|
||||
f"{m['duration_s']:.2f} | {m['success']} | {m['request_throughput']:.2f} | "
|
||||
f"{m['input_token_throughput']:.2f} | {m['output_token_throughput']:.2f} | "
|
||||
f"{m['total_token_throughput']:.2f} | "
|
||||
f"{m['ttft_ms']['mean']:.2f} | {m['ttft_ms']['p95']:.2f} | {m['ttft_ms']['p99']:.2f} | "
|
||||
f"{m['tpot_ms']['mean']:.2f} | {m['tpot_ms']['p95']:.2f} | {m['tpot_ms']['p99']:.2f} | "
|
||||
f"{m['e2e_ms']['mean']:.2f} | {m['e2e_ms']['p95']:.2f} | {m['e2e_ms']['p99']:.2f} |\n"
|
||||
)
|
||||
f.write("\n")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
result_root = Path(sys.argv[1]) if len(sys.argv) > 1 else Path("results")
|
||||
raw_dir = result_root / "raw_outputs"
|
||||
results_json = result_root / "results.json"
|
||||
|
||||
if not raw_dir.exists():
|
||||
raise SystemExit(f"raw_outputs directory not found: {raw_dir}")
|
||||
|
||||
scenarios = []
|
||||
for jsonl_path in sorted(raw_dir.glob("vllm_*.jsonl")):
|
||||
# Filename: vllm_MMDD_concurrency_inputlen_outputlen.jsonl
|
||||
parts = jsonl_path.stem.split("_")
|
||||
if len(parts) < 5:
|
||||
continue
|
||||
concurrency, input_len, output_len = int(parts[2]), int(parts[3]), int(parts[4])
|
||||
|
||||
requests = parse_jsonl(jsonl_path)
|
||||
if not requests:
|
||||
continue
|
||||
|
||||
metrics = compute_metrics(requests)
|
||||
scenario = {
|
||||
"name": scenario_name(concurrency, input_len, output_len),
|
||||
"config": {
|
||||
"concurrency": concurrency,
|
||||
"input_len": input_len,
|
||||
"output_len": output_len,
|
||||
"dataset": "random",
|
||||
"num_prompts": metrics["success"] + metrics["failed"],
|
||||
},
|
||||
"metrics": metrics,
|
||||
"raw_file": str(jsonl_path),
|
||||
}
|
||||
scenarios.append(scenario)
|
||||
|
||||
if not scenarios:
|
||||
print("No benchmark outputs found to parse")
|
||||
return
|
||||
|
||||
if results_json.exists():
|
||||
for s in scenarios:
|
||||
append_scenario(results_json, s)
|
||||
|
||||
generate_report(result_root, scenarios)
|
||||
print(f"Parsed {len(scenarios)} scenarios into {result_root}/report.md")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
164
experiments/dsv4_h200_vllm/run_bench.sh
Executable file
164
experiments/dsv4_h200_vllm/run_bench.sh
Executable file
@ -0,0 +1,164 @@
|
||||
#!/usr/bin/env bash
|
||||
# H200 native vLLM baseline benchmark for DeepSeek-V4-Flash.
|
||||
set -Eeuo pipefail
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
EXPERIMENT_NAME="$(basename "$SCRIPT_DIR")"
|
||||
|
||||
# shellcheck source=/dev/null
|
||||
source "${SCRIPT_DIR}/../../scripts/common/lib.sh"
|
||||
# shellcheck source=/dev/null
|
||||
source "${SCRIPT_DIR}/../../scripts/common/platform.sh"
|
||||
# shellcheck source=/dev/null
|
||||
source "${SCRIPT_DIR}/config.env"
|
||||
|
||||
RUN_ID="${RUN_ID:-$(date '+%Y%m%d-%H%M%S')}"
|
||||
RESULT_ROOT="${RESULT_ROOT:-${SCRIPT_DIR}/results/${RUN_ID}}"
|
||||
RAW_DIR="${RESULT_ROOT}/raw_outputs"
|
||||
LOG_DIR="${RESULT_ROOT}/logs"
|
||||
|
||||
ensure_result_root "$RESULT_ROOT"
|
||||
log_init "${LOG_DIR}/orchestrator.log"
|
||||
|
||||
log "experiment=${EXPERIMENT_NAME}"
|
||||
log "run_id=${RUN_ID}"
|
||||
log "result_root=${RESULT_ROOT}"
|
||||
log "platform=${PLATFORM}"
|
||||
log "chip=${CHIP}"
|
||||
log "accelerator=${ACCELERATOR}"
|
||||
log "engine=${ENGINE}"
|
||||
log "hardware=${HARDWARE}"
|
||||
log "model=${MODEL_PATH}"
|
||||
log "server_start_script=${SERVER_START_SCRIPT}"
|
||||
|
||||
# Write initial metadata for this run.
|
||||
METADATA_JSON="${RESULT_ROOT}/results.json"
|
||||
write_metadata_json \
|
||||
"$METADATA_JSON" \
|
||||
"$EXPERIMENT_NAME" \
|
||||
"$RUN_ID" \
|
||||
"$MODEL_PATH" \
|
||||
"$BACKEND" \
|
||||
"$ENGINE" \
|
||||
"$HARDWARE" \
|
||||
"$ACCELERATOR" \
|
||||
"$CHIP" \
|
||||
"experiments/${EXPERIMENT_NAME}/run_bench.sh" \
|
||||
"$VENV_SERVER" \
|
||||
"H200 native vLLM baseline benchmark for DeepSeek-V4-Flash"
|
||||
|
||||
# Update metadata with config.
|
||||
"${VENV_CLIENT}/bin/python" - "$METADATA_JSON" <<'PY'
|
||||
import json
|
||||
import sys
|
||||
|
||||
path = sys.argv[1]
|
||||
with open(path, "r", encoding="utf-8") as f:
|
||||
data = json.load(f)
|
||||
|
||||
data["config"] = {
|
||||
"tp": 8,
|
||||
"kv_cache_dtype": "fp8",
|
||||
"block_size": 256,
|
||||
"max_num_seqs": 256,
|
||||
"port": 30005,
|
||||
"num_prompts": 128,
|
||||
"scenarios": "32 512 256 128 512 256 32 4000 512"
|
||||
}
|
||||
with open(path, "w", encoding="utf-8") as f:
|
||||
json.dump(data, f, indent=2, ensure_ascii=False)
|
||||
PY
|
||||
|
||||
is_server_healthy() {
|
||||
curl --fail --silent --show-error --max-time 5 "http://127.0.0.1:${PORT}/health" >/dev/null 2>&1
|
||||
}
|
||||
|
||||
stop_server() {
|
||||
local pid_file="/data/user1/yy/dsv4_h200_vllm.pid"
|
||||
if [[ -f "$pid_file" ]]; then
|
||||
local pid
|
||||
pid="$(cat "$pid_file")"
|
||||
if kill -0 "$pid" 2>/dev/null; then
|
||||
log "stopping existing vllm server pid=${pid}"
|
||||
kill "$pid" 2>/dev/null || true
|
||||
sleep 5
|
||||
kill -9 "$pid" 2>/dev/null || true
|
||||
fi
|
||||
rm -f "$pid_file"
|
||||
fi
|
||||
pkill -9 -f "vllm serve.*DeepSeek-V4-Flash" 2>/dev/null || true
|
||||
sleep 2
|
||||
}
|
||||
|
||||
start_server() {
|
||||
log "starting vllm server with ${SERVER_START_SCRIPT}"
|
||||
if [[ ! -x "${SERVER_START_SCRIPT}" ]]; then
|
||||
log "error: start script not found or not executable: ${SERVER_START_SCRIPT}"
|
||||
exit 1
|
||||
fi
|
||||
bash "${SERVER_START_SCRIPT}" >> "${LOG_DIR}/server.outer.log" 2>&1
|
||||
if ! is_server_healthy; then
|
||||
log "error: vllm server failed to become healthy"
|
||||
exit 1
|
||||
fi
|
||||
log "vllm server is healthy"
|
||||
}
|
||||
|
||||
on_exit() {
|
||||
local code=$?
|
||||
log "orchestrator exiting with code=${code}"
|
||||
if [[ -z "${SKIP_MANAGE_SERVER:-}" ]]; then
|
||||
stop_server
|
||||
fi
|
||||
exit "$code"
|
||||
}
|
||||
trap on_exit EXIT
|
||||
|
||||
if [[ -n "${SKIP_MANAGE_SERVER:-}" ]]; then
|
||||
log "SKIP_MANAGE_SERVER is set, assuming server is already running on port ${PORT}"
|
||||
if ! is_server_healthy; then
|
||||
log "error: no healthy server found at port ${PORT}"
|
||||
exit 1
|
||||
fi
|
||||
else
|
||||
stop_server
|
||||
start_server
|
||||
fi
|
||||
|
||||
log "===== BENCHMARK START ====="
|
||||
|
||||
for scenario in ${SCENARIOS}; do
|
||||
read -r concurrency input_len output_len <<< "$scenario"
|
||||
output_file="${RAW_DIR}/vllm_$(date '+%m%d')_${concurrency}_${input_len}_${output_len}.jsonl"
|
||||
detail_log="${LOG_DIR}/vllm_c${concurrency}_i${input_len}_o${output_len}.log"
|
||||
|
||||
log "running scenario: concurrency=${concurrency} input=${input_len} output=${output_len}"
|
||||
|
||||
"${VENV_CLIENT}/bin/python" -m sglang.bench_serving \
|
||||
--backend vllm \
|
||||
--host 127.0.0.1 \
|
||||
--port "$PORT" \
|
||||
--dataset-name random \
|
||||
--random-input-len "$input_len" \
|
||||
--random-output-len "$output_len" \
|
||||
--num-prompts "$NUM_PROMPTS" \
|
||||
--max-concurrency "$concurrency" \
|
||||
--request-rate 10000 \
|
||||
--output-file "$output_file" \
|
||||
--output-details \
|
||||
> "$detail_log" 2>&1 || {
|
||||
log "ERROR: scenario c=${concurrency} i=${input_len} o=${output_len} failed; see ${detail_log}"
|
||||
continue
|
||||
}
|
||||
|
||||
log "finished scenario: output=${output_file}"
|
||||
done
|
||||
|
||||
log "===== BENCHMARK DONE ====="
|
||||
|
||||
log "parsing results"
|
||||
"${VENV_CLIENT}/bin/python" "${SCRIPT_DIR}/parse_results.py" "$RESULT_ROOT" >> "${LOG_DIR}/parse.log" 2>&1 || {
|
||||
log "WARNING: parser failed; see ${LOG_DIR}/parse.log"
|
||||
}
|
||||
|
||||
log "all results saved to ${RESULT_ROOT}"
|
||||
65
experiments/dsv4_h200_vllm/start_server.sh
Executable file
65
experiments/dsv4_h200_vllm/start_server.sh
Executable file
@ -0,0 +1,65 @@
|
||||
#!/bin/bash
|
||||
# Start a plain vLLM baseline server for DeepSeek-V4-Flash on H200.
|
||||
set -e
|
||||
|
||||
cd /data/user1/yy
|
||||
mkdir -p logs
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
# shellcheck source=/dev/null
|
||||
source "${SCRIPT_DIR}/config.env"
|
||||
|
||||
VENV="${VENV_SERVER}"
|
||||
export PATH="$VENV/bin:$PATH"
|
||||
VLLM="$VENV/bin/vllm"
|
||||
|
||||
TP=4
|
||||
LOG="/data/user1/yy/logs/dsv4_h200_vllm_tp${TP}_$(date +%Y%m%d_%H%M%S).log"
|
||||
PID_FILE="/data/user1/yy/dsv4_h200_vllm.pid"
|
||||
|
||||
export TMPDIR=/data/user1/yy/tmp
|
||||
export CUDA_VISIBLE_DEVICES=4,5,6,7
|
||||
|
||||
echo "=== Starting DeepSeek-V4-Flash vLLM baseline (TP=$TP) ==="
|
||||
echo "Model: $MODEL_PATH"
|
||||
echo "Port: $PORT"
|
||||
echo "Log: $LOG"
|
||||
|
||||
rm -f "$PID_FILE"
|
||||
nohup "$VLLM" serve "$MODEL_PATH" \
|
||||
--trust-remote-code \
|
||||
--tensor-parallel-size "$TP" \
|
||||
--kv-cache-dtype fp8 \
|
||||
--block-size 256 \
|
||||
--max-model-len auto \
|
||||
--max-num-seqs 256 \
|
||||
--gpu-memory-utilization 0.90 \
|
||||
--tokenizer-mode deepseek_v4 \
|
||||
--reasoning-parser deepseek_v4 \
|
||||
--no-disable-hybrid-kv-cache-manager \
|
||||
--disable-uvicorn-access-log \
|
||||
--port "$PORT" \
|
||||
> "$LOG" 2>&1 &
|
||||
PID=$!
|
||||
echo $PID > "$PID_FILE"
|
||||
echo "PID: $PID"
|
||||
echo "Waiting for health..."
|
||||
|
||||
for i in $(seq 1 240); do
|
||||
if curl -s "http://127.0.0.1:$PORT/health" > /dev/null 2>&1; then
|
||||
echo "Server is ready at http://127.0.0.1:$PORT"
|
||||
echo "Log: $LOG"
|
||||
exit 0
|
||||
fi
|
||||
if ! kill -0 $PID 2>/dev/null; then
|
||||
echo "ERROR: Server exited early"
|
||||
tail -200 "$LOG"
|
||||
exit 1
|
||||
fi
|
||||
echo "Waiting... ($i/240)"
|
||||
sleep 5
|
||||
done
|
||||
|
||||
echo "ERROR: Server not healthy after 240 retries"
|
||||
tail -200 "$LOG"
|
||||
exit 1
|
||||
Loading…
x
Reference in New Issue
Block a user