update
3
.gitignore
vendored
@ -1,5 +1,6 @@
|
||||
# VBench generated videos, evaluation outputs, and logs stay on 6000D_H3.
|
||||
/vbench-base/results/
|
||||
/vbench-score/vbench-base/results/
|
||||
/vbench-score/vbench-lora/results/
|
||||
|
||||
# SQLite analysis databases stay on 6000D_H3.
|
||||
**/*.sqlite
|
||||
|
||||
51
README.md
@ -1,29 +1,48 @@
|
||||
# sskj-h3
|
||||
|
||||
MiniMax-H3 在 RTX 6000D-H3 上的部署基准、性能 Profile 与 VBench Base 评测归档。
|
||||
|
||||
本目录于 2026-08-27 从 `/data/wxy` 中的已确认源路径创建。归档采用实体副本;原脚本、结果、日志和模型目录均未移动、修改或删除。第三方 VBench 源码、模型权重、Conda 环境和无关日志不纳入归档。
|
||||
MiniMax-H3 在 RTX 6000D-H3 上的吞吐、性能 Profile、成对 SSIM 与 VBench 评测归档。
|
||||
|
||||
## 目录
|
||||
|
||||
- `sglang-base`:SGLang TP8×1、TP4×2、TP2×4,20 steps / 5s 基准。
|
||||
- `vllm-omni-base`:vLLM-Omni 1×8、2×4、4×2,20 steps / 5s 基准。
|
||||
- `sglang-profile`:SGLang TP2、768P、FL2VA/Ref2VA 输入矩阵、Torch/Nsight/NCCL 与 SDPA kernel 分析。
|
||||
- `vbench-base`:SGLang TP2×4 生成的 944 个 VBench 视频、16 维评分、独立评分环境与兼容适配记录。
|
||||
- `throughput/`:吞吐测试、Profile 与成对 SSIM。
|
||||
- `sglang-base/`:SGLang TP8×1、TP4×2、TP2×4,20 steps / 5s 基准。
|
||||
- `sglang-base-b300/`:B300 相关 SGLang 基准。
|
||||
- `sglang-lora/`:Larry LoRA 加速基准。
|
||||
- `sglang-profile/`:Torch/Nsight/NCCL 与 SDPA kernel 分析。
|
||||
- `vllm-omni-base/`:vLLM-Omni 多实例基准。
|
||||
- `common/paired_video_ssim.py`:候选运行相对 base 的逐帧成对 SSIM。
|
||||
- `vbench-score/`:VBench 视频与 16 维评分。
|
||||
- `vbench-base/`:base 生成与评分归档。
|
||||
- `vbench-lora/`:Larry LoRA 生成与评分归档。
|
||||
- `SOURCE_MAP.tsv`:源路径、归档路径、文件数和字节数。
|
||||
- `tools`:可重复执行的非破坏性归档脚本与验收脚本。
|
||||
- `tools/`:非破坏性归档与验收脚本。
|
||||
|
||||
每个实验目录的 `scripts/SHA256SUMS` 可用于校验归档脚本。结果目录保留原始层级、日志、JSONL、MP4、Torch trace 和 Nsight report。
|
||||
## SSIM 与 VBench 的分工
|
||||
|
||||
成对 SSIM 用相同 prompt、seed、任务、分辨率、时长和宽高比的 base 视频作为参考,按 `request_id` 配对。视频由 FFmpeg 解码并统一为 `yuv420p`,然后逐帧计算 Y/U/V/All SSIM;主口径是 `All`。它适合测量 Cache-DiT、TeaCache、Larry 等加速方案对 base 输出的像素/结构偏移。
|
||||
|
||||
VBench 独立衡量主体一致性、运动、审美等生成质量维度。SSIM 高不等于 VBench 高,VBench 高也不保证逐样本复现,因此两者互补。
|
||||
|
||||
吞吐 runner 只有在设置 `SSIM_REFERENCE_ROOT` 时才会在每个 phase 结束、SGLang 服务完全停止后评分,不会把解码和评分时间计入吞吐:
|
||||
|
||||
```bash
|
||||
SSIM_REFERENCE_ROOT=/data/wxy/sskj-h3/throughput/sglang-base/results/balanced-tp4-tp2-20steps-5s-20260822-175030 \
|
||||
SSIM_THRESHOLD=0.90 \
|
||||
bash /data/wxy/sskj-h3/throughput/sglang-lora/scripts/run_sglang_h3_lora_mixed_matrix_6000d.sh
|
||||
```
|
||||
|
||||
参考目录和候选目录都应包含 `tpN_replicasM/<task>/client_*/results.jsonl`。每个候选 phase 会新增:
|
||||
|
||||
- `quality/paired_ssim.json`:总体、分辨率分组和逐视频结果。
|
||||
- `quality/paired_ssim.tsv`:逐视频表。
|
||||
- `quality/paired_ssim_frames.tsv`:逐帧表。
|
||||
- `quality/paired_ssim.log`:评分日志。
|
||||
|
||||
默认阈值为 0.90,只记录是否通过,不中止完整矩阵;需要将低于阈值视作失败时设置 `SSIM_FAIL_BELOW_THRESHOLD=true`。
|
||||
|
||||
## Git 镜像边界
|
||||
|
||||
服务器归档保留全部实体文件;同步到 Git 仓库时排除 `vbench-base/results/`、全部 `*.mp4`、全部 `*.sqlite`、Nsight `*.nsys-rep` 和 Torch `*.trace.json.gz` 原始采集。代码、日志、JSON/JSONL、TSV、采集脚本、结构化汇总和 Profile 分析文档正常纳入版本库。
|
||||
|
||||
## 对应飞书报告
|
||||
|
||||
- [SGLang 多实例部署测试报告](https://gcn673xpgdxn.feishu.cn/docx/Mzh4dPPQtoFdYTxHJE6cEumRnXx)
|
||||
- [vLLM-Omni 多实例部署测试报告](https://gcn673xpgdxn.feishu.cn/docx/GYXwdRwOgoWuRsxrUKick2Ljn5s)
|
||||
- [SGLang 6000D-H3 性能 Profile 完整分析报告](https://gcn673xpgdxn.feishu.cn/docx/UmoCdnWa2okvjSxjxK0cGnIxnqe)
|
||||
服务器保留全部实体文件;同步到 Git 时排除 `vbench-score/*/results/`、全部 `*.mp4`、全部 `*.sqlite`、Nsight `*.nsys-rep` 和 Torch `*.trace.json.gz` 原始采集。代码、日志、JSON/JSONL、TSV、结构化汇总和分析文档正常纳入版本库。
|
||||
|
||||
## 验证
|
||||
|
||||
|
||||
@ -1,14 +1,14 @@
|
||||
section kind source destination files bytes
|
||||
sglang-base script /data/wxy/run_sglang_h3_mixed_matrix_6000d.sh sglang-base/scripts/run_sglang_h3_mixed_matrix_6000d.sh 1 7888
|
||||
sglang-base script /data/wxy/minimax_h3_mixed_bench.py sglang-base/scripts/minimax_h3_mixed_bench.py 1 12687
|
||||
sglang-base result /data/wxy/results/minimax_h3_mixed_matrix/mixed64-20steps-5s-20260822-100844 sglang-base/results/mixed64-20steps-5s-20260822-100844 212 342464940
|
||||
sglang-base result /data/wxy/results/minimax_h3_mixed_matrix/balanced-tp4-tp2-20steps-5s-20260822-175030 sglang-base/results/balanced-tp4-tp2-20steps-5s-20260822-175030 195 268064833
|
||||
vllm-omni-base script /data/wxy/run_vllm_omni_h3_matrix_6000d.sh vllm-omni-base/scripts/run_vllm_omni_h3_matrix_6000d.sh 1 10411
|
||||
vllm-omni-base script /data/wxy/minimax_h3_vllm_bench.py vllm-omni-base/scripts/minimax_h3_vllm_bench.py 1 14589
|
||||
vllm-omni-base result /data/wxy/results/minimax_h3_vllm_matrix/vllm-balanced64-20steps-5s-r2-20260822-225317 vllm-omni-base/results/vllm-balanced64-20steps-5s-r2-20260822-225317 271 1998342105
|
||||
sglang-profile script /data/wxy/h3_profile sglang-profile/scripts/h3_profile 15 322059
|
||||
sglang-profile result /data/wxy/profile_results/h3-quick-input-matrix-20260824-run1 sglang-profile/results/h3-quick-input-matrix-20260824-run1 71 32017378
|
||||
sglang-profile result /data/wxy/profile_results/h3-targeted-profile-20260824-run1 sglang-profile/results/h3-targeted-profile-20260824-run1 87 2124349553
|
||||
sglang-profile result /data/wxy/profile_results/h3-sdpa-kernel-matrix-20260825-run1 sglang-profile/results/h3-sdpa-kernel-matrix-20260825-run1 43 615920
|
||||
vbench-base result /data/wxy/results/h3_vbench_base/h3-vbench-base-dense-tp2x4-20260826-run1 vbench-base/results/h3-vbench-base-dense-tp2x4-20260826-run1 1983 1982500600
|
||||
vbench-base environment /data/wxy/vbench_score_setup_logs vbench-base/environment/vbench_score_setup_logs 11 49335
|
||||
sglang-base script /data/wxy/run_sglang_h3_mixed_matrix_6000d.sh throughput/sglang-base/scripts/run_sglang_h3_mixed_matrix_6000d.sh 1 7888
|
||||
sglang-base script /data/wxy/minimax_h3_mixed_bench.py throughput/sglang-base/scripts/minimax_h3_mixed_bench.py 1 12687
|
||||
sglang-base result /data/wxy/results/minimax_h3_mixed_matrix/mixed64-20steps-5s-20260822-100844 throughput/sglang-base/results/mixed64-20steps-5s-20260822-100844 212 342464940
|
||||
sglang-base result /data/wxy/results/minimax_h3_mixed_matrix/balanced-tp4-tp2-20steps-5s-20260822-175030 throughput/sglang-base/results/balanced-tp4-tp2-20steps-5s-20260822-175030 195 268064833
|
||||
vllm-omni-base script /data/wxy/run_vllm_omni_h3_matrix_6000d.sh throughput/vllm-omni-base/scripts/run_vllm_omni_h3_matrix_6000d.sh 1 10411
|
||||
vllm-omni-base script /data/wxy/minimax_h3_vllm_bench.py throughput/vllm-omni-base/scripts/minimax_h3_vllm_bench.py 1 14589
|
||||
vllm-omni-base result /data/wxy/results/minimax_h3_vllm_matrix/vllm-balanced64-20steps-5s-r2-20260822-225317 throughput/vllm-omni-base/results/vllm-balanced64-20steps-5s-r2-20260822-225317 271 1998342105
|
||||
sglang-profile script /data/wxy/h3_profile throughput/sglang-profile/scripts/h3_profile 15 322059
|
||||
sglang-profile result /data/wxy/profile_results/h3-quick-input-matrix-20260824-run1 throughput/sglang-profile/results/h3-quick-input-matrix-20260824-run1 71 32017378
|
||||
sglang-profile result /data/wxy/profile_results/h3-targeted-profile-20260824-run1 throughput/sglang-profile/results/h3-targeted-profile-20260824-run1 87 2124349553
|
||||
sglang-profile result /data/wxy/profile_results/h3-sdpa-kernel-matrix-20260825-run1 throughput/sglang-profile/results/h3-sdpa-kernel-matrix-20260825-run1 43 615920
|
||||
vbench-base result /data/wxy/results/h3_vbench_base/h3-vbench-base-dense-tp2x4-20260826-run1 vbench-score/vbench-base/results/h3-vbench-base-dense-tp2x4-20260826-run1 1983 1982500600
|
||||
vbench-base environment /data/wxy/vbench_score_setup_logs vbench-score/vbench-base/environment/vbench_score_setup_logs 11 49335
|
||||
|
||||
|
7
throughput/README.md
Normal file
@ -0,0 +1,7 @@
|
||||
# Throughput
|
||||
|
||||
MiniMax-H3 在 RTX 6000D-H3 上的 SGLang、Larry LoRA、vLLM-Omni 吞吐实验和性能 Profile。
|
||||
|
||||
吞吐与 SSIM 严格分阶段执行:客户端完成并写出吞吐结果后先停止服务,再运行 `common/paired_video_ssim.py`。因此启用 SSIM 不改变请求参数、并发方式、计时窗口或吞吐汇总。
|
||||
|
||||
SSIM 需要一个同拓扑、同任务的 base run 作为 `SSIM_REFERENCE_ROOT`。未设置时 runner 与原吞吐逻辑一致,不执行质量评分。
|
||||
450
throughput/common/paired_video_ssim.py
Executable file
@ -0,0 +1,450 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Compute paired, frame-aligned YUV420 SSIM for MiniMax-H3 runs.
|
||||
|
||||
The candidate and reference directories must each contain the throughput
|
||||
client ``results.jsonl`` files. Rows are paired by ``request_id`` and checked
|
||||
for matching prompt, seed, task, geometry, duration, and aspect ratio before
|
||||
the decoded videos are compared with FFmpeg's native ``ssim`` filter.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import math
|
||||
import shutil
|
||||
import statistics
|
||||
import subprocess
|
||||
import tempfile
|
||||
from collections import defaultdict
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
|
||||
PAIR_FIELDS = (
|
||||
"task",
|
||||
"short_edge",
|
||||
"prompt_index",
|
||||
"prompt",
|
||||
"seed",
|
||||
"duration_seconds",
|
||||
"aspect_ratio",
|
||||
)
|
||||
|
||||
|
||||
def percentile(values: list[float], q: float) -> float:
|
||||
if not values:
|
||||
return 0.0
|
||||
ordered = sorted(values)
|
||||
pos = (len(ordered) - 1) * q
|
||||
lo, hi = math.floor(pos), math.ceil(pos)
|
||||
if lo == hi:
|
||||
return ordered[lo]
|
||||
return ordered[lo] * (hi - pos) + ordered[hi] * (pos - lo)
|
||||
|
||||
|
||||
def resolve_executable(explicit: str | None, name: str) -> str:
|
||||
if explicit:
|
||||
path = Path(explicit)
|
||||
if path.is_file():
|
||||
return str(path)
|
||||
resolved = shutil.which(explicit)
|
||||
if resolved:
|
||||
return resolved
|
||||
raise SystemExit(f"{name} executable not found: {explicit}")
|
||||
|
||||
resolved = shutil.which(name)
|
||||
if resolved:
|
||||
return resolved
|
||||
candidates = [
|
||||
Path("/root/.miniconda3/envs/deploy/bin") / name,
|
||||
Path("/root/.miniconda3/envs/vllm/bin") / name,
|
||||
Path("/root/.miniconda3/envs/bbj/bin") / name,
|
||||
]
|
||||
for path in candidates:
|
||||
if path.is_file():
|
||||
return str(path)
|
||||
raise SystemExit(f"{name} is required; pass --{name} explicitly")
|
||||
|
||||
|
||||
def run_checked(command: list[str]) -> subprocess.CompletedProcess[str]:
|
||||
result = subprocess.run(command, text=True, capture_output=True, check=False)
|
||||
if result.returncode:
|
||||
rendered = " ".join(command)
|
||||
raise RuntimeError(
|
||||
f"command failed ({result.returncode}): {rendered}\n{result.stderr[-4000:]}"
|
||||
)
|
||||
return result
|
||||
|
||||
|
||||
def resolve_video_path(root: Path, raw_path: str) -> Path:
|
||||
if not raw_path:
|
||||
raise ValueError("file_path is empty")
|
||||
path = Path(raw_path)
|
||||
if path.is_file():
|
||||
return path.resolve()
|
||||
matches = [candidate for candidate in root.rglob(path.name) if candidate.is_file()]
|
||||
if len(matches) == 1:
|
||||
return matches[0].resolve()
|
||||
if not matches:
|
||||
raise ValueError(f"video does not exist: {path}")
|
||||
raise ValueError(
|
||||
f"video path {path} is stale and filename is ambiguous under {root}: "
|
||||
+ ", ".join(str(match) for match in matches[:10])
|
||||
)
|
||||
|
||||
|
||||
def load_rows(root: Path) -> dict[str, dict[str, Any]]:
|
||||
rows: dict[str, dict[str, Any]] = {}
|
||||
files = sorted(root.glob("client_*/results.jsonl"))
|
||||
if not files:
|
||||
files = sorted(root.rglob("client_*/results.jsonl"))
|
||||
if not files:
|
||||
raise ValueError(f"no client_*/results.jsonl found under {root}")
|
||||
|
||||
for path in files:
|
||||
for line_number, line in enumerate(
|
||||
path.read_text(encoding="utf-8").splitlines(), start=1
|
||||
):
|
||||
if not line.strip():
|
||||
continue
|
||||
row = json.loads(line)
|
||||
if not row.get("success"):
|
||||
continue
|
||||
request_id = str(row.get("request_id") or "")
|
||||
if not request_id:
|
||||
raise ValueError(f"missing request_id: {path}:{line_number}")
|
||||
if request_id in rows:
|
||||
raise ValueError(
|
||||
f"duplicate request_id {request_id!r} under {root}; "
|
||||
"pass one task/topology phase rather than a whole matrix"
|
||||
)
|
||||
try:
|
||||
video_path = resolve_video_path(root, str(row.get("file_path") or ""))
|
||||
except ValueError as error:
|
||||
raise ValueError(f"video for {request_id!r}: {error}") from error
|
||||
row["_resolved_file_path"] = str(video_path)
|
||||
rows[request_id] = row
|
||||
return rows
|
||||
|
||||
|
||||
def check_pair(candidate: dict[str, Any], reference: dict[str, Any]) -> None:
|
||||
mismatches = []
|
||||
for field in PAIR_FIELDS:
|
||||
if candidate.get(field) != reference.get(field):
|
||||
mismatches.append(
|
||||
f"{field}: candidate={candidate.get(field)!r} "
|
||||
f"reference={reference.get(field)!r}"
|
||||
)
|
||||
if mismatches:
|
||||
raise ValueError("pair metadata mismatch: " + "; ".join(mismatches))
|
||||
|
||||
|
||||
def probe_video(ffprobe: str, path: Path) -> dict[str, Any]:
|
||||
result = run_checked(
|
||||
[
|
||||
ffprobe,
|
||||
"-v",
|
||||
"error",
|
||||
"-select_streams",
|
||||
"v:0",
|
||||
"-count_frames",
|
||||
"-show_entries",
|
||||
"stream=width,height,pix_fmt,r_frame_rate,avg_frame_rate,nb_frames,nb_read_frames",
|
||||
"-of",
|
||||
"json",
|
||||
str(path),
|
||||
]
|
||||
)
|
||||
payload = json.loads(result.stdout)
|
||||
streams = payload.get("streams") or []
|
||||
if len(streams) != 1:
|
||||
raise ValueError(f"expected one video stream in {path}, got {len(streams)}")
|
||||
stream = streams[0]
|
||||
frame_text = stream.get("nb_read_frames") or stream.get("nb_frames")
|
||||
if frame_text in (None, "N/A"):
|
||||
raise ValueError(f"could not determine decoded frame count for {path}")
|
||||
return {
|
||||
"width": int(stream["width"]),
|
||||
"height": int(stream["height"]),
|
||||
"pix_fmt": stream.get("pix_fmt"),
|
||||
"r_frame_rate": stream.get("r_frame_rate"),
|
||||
"avg_frame_rate": stream.get("avg_frame_rate"),
|
||||
"frames": int(frame_text),
|
||||
}
|
||||
|
||||
|
||||
def check_video_contract(candidate: dict[str, Any], reference: dict[str, Any]) -> None:
|
||||
fields = ("width", "height", "r_frame_rate", "frames")
|
||||
mismatches = [
|
||||
f"{field}: candidate={candidate[field]!r} reference={reference[field]!r}"
|
||||
for field in fields
|
||||
if candidate[field] != reference[field]
|
||||
]
|
||||
if mismatches:
|
||||
raise ValueError("decoded video mismatch: " + "; ".join(mismatches))
|
||||
|
||||
|
||||
def parse_ffmpeg_stats(path: Path) -> list[dict[str, float | int]]:
|
||||
frames: list[dict[str, float | int]] = []
|
||||
for line in path.read_text(encoding="utf-8").splitlines():
|
||||
values: dict[str, str] = {}
|
||||
for token in line.split():
|
||||
if ":" in token:
|
||||
key, value = token.split(":", 1)
|
||||
values[key] = value
|
||||
if not {"n", "Y", "U", "V", "All"}.issubset(values):
|
||||
continue
|
||||
frames.append(
|
||||
{
|
||||
"frame": int(values["n"]),
|
||||
"y": float(values["Y"]),
|
||||
"u": float(values["U"]),
|
||||
"v": float(values["V"]),
|
||||
"all": float(values["All"]),
|
||||
}
|
||||
)
|
||||
if not frames:
|
||||
raise ValueError(f"FFmpeg emitted no per-frame SSIM metrics: {path}")
|
||||
return frames
|
||||
|
||||
|
||||
def compare_video_pair(
|
||||
ffmpeg: str,
|
||||
ffprobe: str,
|
||||
candidate_path: Path,
|
||||
reference_path: Path,
|
||||
stats_path: Path,
|
||||
) -> tuple[dict[str, Any], list[dict[str, float | int]]]:
|
||||
candidate_probe = probe_video(ffprobe, candidate_path)
|
||||
reference_probe = probe_video(ffprobe, reference_path)
|
||||
check_video_contract(candidate_probe, reference_probe)
|
||||
|
||||
filter_graph = (
|
||||
"[0:v]setpts=PTS-STARTPTS,format=yuv420p[candidate];"
|
||||
"[1:v]setpts=PTS-STARTPTS,format=yuv420p[reference];"
|
||||
f"[candidate][reference]ssim=stats_file={stats_path}"
|
||||
)
|
||||
run_checked(
|
||||
[
|
||||
ffmpeg,
|
||||
"-hide_banner",
|
||||
"-nostdin",
|
||||
"-loglevel",
|
||||
"error",
|
||||
"-i",
|
||||
str(candidate_path),
|
||||
"-i",
|
||||
str(reference_path),
|
||||
"-filter_complex",
|
||||
filter_graph,
|
||||
"-an",
|
||||
"-f",
|
||||
"null",
|
||||
"-",
|
||||
]
|
||||
)
|
||||
frames = parse_ffmpeg_stats(stats_path)
|
||||
if len(frames) != candidate_probe["frames"]:
|
||||
raise ValueError(
|
||||
f"SSIM frame count mismatch: metrics={len(frames)} "
|
||||
f"decoded={candidate_probe['frames']}"
|
||||
)
|
||||
return candidate_probe, frames
|
||||
|
||||
|
||||
def metric_summary(values: list[float]) -> dict[str, float]:
|
||||
return {
|
||||
"mean": statistics.fmean(values),
|
||||
"p10": percentile(values, 0.10),
|
||||
"min": min(values),
|
||||
"max": max(values),
|
||||
}
|
||||
|
||||
|
||||
def compare_command(args: argparse.Namespace) -> int:
|
||||
if not 0.0 <= args.threshold <= 1.0:
|
||||
raise SystemExit("--threshold must be between 0 and 1")
|
||||
ffmpeg = resolve_executable(args.ffmpeg, "ffmpeg")
|
||||
ffprobe = resolve_executable(args.ffprobe, "ffprobe")
|
||||
candidates = load_rows(args.candidate_dir)
|
||||
references = load_rows(args.reference_dir)
|
||||
|
||||
candidate_ids = set(candidates)
|
||||
reference_ids = set(references)
|
||||
if candidate_ids != reference_ids:
|
||||
missing_reference = sorted(candidate_ids - reference_ids)
|
||||
missing_candidate = sorted(reference_ids - candidate_ids)
|
||||
raise SystemExit(
|
||||
"request sets do not match: "
|
||||
f"missing_reference={missing_reference[:20]} "
|
||||
f"missing_candidate={missing_candidate[:20]}"
|
||||
)
|
||||
|
||||
selected_ids = sorted(candidate_ids)
|
||||
if args.limit is not None:
|
||||
if args.limit < 1:
|
||||
raise SystemExit("--limit must be at least 1")
|
||||
selected_ids = selected_ids[: args.limit]
|
||||
|
||||
args.output_dir.mkdir(parents=True, exist_ok=True)
|
||||
pair_rows: list[dict[str, Any]] = []
|
||||
all_frame_rows: list[dict[str, Any]] = []
|
||||
with tempfile.TemporaryDirectory(prefix="h3-paired-ssim-") as temporary:
|
||||
temporary_root = Path(temporary)
|
||||
for index, request_id in enumerate(selected_ids, start=1):
|
||||
candidate = candidates[request_id]
|
||||
reference = references[request_id]
|
||||
check_pair(candidate, reference)
|
||||
probe, frames = compare_video_pair(
|
||||
ffmpeg,
|
||||
ffprobe,
|
||||
Path(candidate["_resolved_file_path"]),
|
||||
Path(reference["_resolved_file_path"]),
|
||||
temporary_root / f"{index:05d}.stats",
|
||||
)
|
||||
all_values = [float(frame["all"]) for frame in frames]
|
||||
y_values = [float(frame["y"]) for frame in frames]
|
||||
u_values = [float(frame["u"]) for frame in frames]
|
||||
v_values = [float(frame["v"]) for frame in frames]
|
||||
all_summary = metric_summary(all_values)
|
||||
pair = {
|
||||
"request_id": request_id,
|
||||
"task": candidate["task"],
|
||||
"short_edge": candidate["short_edge"],
|
||||
"prompt_index": candidate["prompt_index"],
|
||||
"prompt": candidate["prompt"],
|
||||
"seed": candidate["seed"],
|
||||
"candidate_file": candidate["_resolved_file_path"],
|
||||
"reference_file": reference["_resolved_file_path"],
|
||||
"candidate_num_inference_steps": candidate.get("num_inference_steps"),
|
||||
"reference_num_inference_steps": reference.get("num_inference_steps"),
|
||||
"width": probe["width"],
|
||||
"height": probe["height"],
|
||||
"fps": probe["r_frame_rate"],
|
||||
"frames": len(frames),
|
||||
"ssim_all_mean": all_summary["mean"],
|
||||
"ssim_all_p10": all_summary["p10"],
|
||||
"ssim_all_min": all_summary["min"],
|
||||
"ssim_y_mean": statistics.fmean(y_values),
|
||||
"ssim_u_mean": statistics.fmean(u_values),
|
||||
"ssim_v_mean": statistics.fmean(v_values),
|
||||
"threshold": args.threshold,
|
||||
"passed": all_summary["mean"] >= args.threshold,
|
||||
}
|
||||
pair_rows.append(pair)
|
||||
for frame in frames:
|
||||
all_frame_rows.append({"request_id": request_id, **frame})
|
||||
print(
|
||||
f"[{index}/{len(selected_ids)}] {request_id} "
|
||||
f"mean={all_summary['mean']:.6f} p10={all_summary['p10']:.6f} "
|
||||
f"min={all_summary['min']:.6f}",
|
||||
flush=True,
|
||||
)
|
||||
|
||||
all_values = [float(row["all"]) for row in all_frame_rows]
|
||||
video_means = [float(row["ssim_all_mean"]) for row in pair_rows]
|
||||
by_short_edge: dict[str, dict[str, Any]] = {}
|
||||
grouped: dict[int, list[float]] = defaultdict(list)
|
||||
for row in pair_rows:
|
||||
grouped[int(row["short_edge"])].append(float(row["ssim_all_mean"]))
|
||||
for short_edge, values in sorted(grouped.items()):
|
||||
by_short_edge[str(short_edge)] = {
|
||||
"videos": len(values),
|
||||
"mean_video_ssim": statistics.fmean(values),
|
||||
"min_video_ssim": min(values),
|
||||
}
|
||||
|
||||
frame_summary = metric_summary(all_values)
|
||||
summary = {
|
||||
"metric": "FFmpeg decoded YUV420 SSIM All",
|
||||
"aggregation": {
|
||||
"mean_video_ssim": statistics.fmean(video_means),
|
||||
"frame_weighted_mean_ssim": frame_summary["mean"],
|
||||
"frame_p10_ssim": frame_summary["p10"],
|
||||
"min_frame_ssim": frame_summary["min"],
|
||||
},
|
||||
"candidate_dir": str(args.candidate_dir.resolve()),
|
||||
"reference_dir": str(args.reference_dir.resolve()),
|
||||
"ffmpeg": ffmpeg,
|
||||
"ffprobe": ffprobe,
|
||||
"threshold": args.threshold,
|
||||
"overall_passed": statistics.fmean(video_means) >= args.threshold,
|
||||
"videos": len(pair_rows),
|
||||
"videos_passed": sum(bool(row["passed"]) for row in pair_rows),
|
||||
"frames": len(all_frame_rows),
|
||||
"by_short_edge": by_short_edge,
|
||||
"pairs": pair_rows,
|
||||
}
|
||||
(args.output_dir / "paired_ssim.json").write_text(
|
||||
json.dumps(summary, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
|
||||
)
|
||||
|
||||
pair_columns = [
|
||||
"request_id",
|
||||
"task",
|
||||
"short_edge",
|
||||
"prompt_index",
|
||||
"seed",
|
||||
"width",
|
||||
"height",
|
||||
"fps",
|
||||
"frames",
|
||||
"ssim_all_mean",
|
||||
"ssim_all_p10",
|
||||
"ssim_all_min",
|
||||
"ssim_y_mean",
|
||||
"ssim_u_mean",
|
||||
"ssim_v_mean",
|
||||
"threshold",
|
||||
"passed",
|
||||
"candidate_file",
|
||||
"reference_file",
|
||||
]
|
||||
with (args.output_dir / "paired_ssim.tsv").open("w", encoding="utf-8") as handle:
|
||||
handle.write("\t".join(pair_columns) + "\n")
|
||||
for row in pair_rows:
|
||||
handle.write("\t".join(str(row[column]) for column in pair_columns) + "\n")
|
||||
|
||||
frame_columns = ("request_id", "frame", "y", "u", "v", "all")
|
||||
with (args.output_dir / "paired_ssim_frames.tsv").open(
|
||||
"w", encoding="utf-8"
|
||||
) as handle:
|
||||
handle.write("\t".join(frame_columns) + "\n")
|
||||
for row in all_frame_rows:
|
||||
handle.write("\t".join(str(row[column]) for column in frame_columns) + "\n")
|
||||
|
||||
print(json.dumps(summary["aggregation"], ensure_ascii=False), flush=True)
|
||||
if args.fail_below_threshold and not summary["overall_passed"]:
|
||||
return 2
|
||||
return 0
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser()
|
||||
subparsers = parser.add_subparsers(dest="command", required=True)
|
||||
compare = subparsers.add_parser("compare")
|
||||
compare.add_argument("--candidate-dir", type=Path, required=True)
|
||||
compare.add_argument("--reference-dir", type=Path, required=True)
|
||||
compare.add_argument("--output-dir", type=Path, required=True)
|
||||
compare.add_argument("--threshold", type=float, default=0.90)
|
||||
compare.add_argument(
|
||||
"--limit",
|
||||
type=int,
|
||||
help="compare only the first N matched requests (intended for smoke tests)",
|
||||
)
|
||||
compare.add_argument("--ffmpeg")
|
||||
compare.add_argument("--ffprobe")
|
||||
compare.add_argument("--fail-below-threshold", action="store_true")
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = parse_args()
|
||||
if args.command == "compare":
|
||||
return compare_command(args)
|
||||
raise AssertionError(args.command)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
167
throughput/common/summarize_ssim_comparison.py
Executable file
@ -0,0 +1,167 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Aggregate paired-video SSIM JSON files into a comparison report."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import math
|
||||
import statistics
|
||||
from collections import defaultdict
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
|
||||
def percentile(values: list[float], q: float) -> float:
|
||||
ordered = sorted(values)
|
||||
position = (len(ordered) - 1) * q
|
||||
lower, upper = math.floor(position), math.ceil(position)
|
||||
if lower == upper:
|
||||
return ordered[lower]
|
||||
return ordered[lower] * (upper - position) + ordered[upper] * (position - lower)
|
||||
|
||||
|
||||
def summarize(pairs: list[dict[str, Any]], threshold: float) -> dict[str, Any]:
|
||||
values = [float(pair["ssim_all_mean"]) for pair in pairs]
|
||||
return {
|
||||
"videos": len(values),
|
||||
"mean_video_ssim": statistics.fmean(values),
|
||||
"median_video_ssim": statistics.median(values),
|
||||
"p10_video_ssim": percentile(values, 0.10),
|
||||
"min_video_ssim": min(values),
|
||||
"videos_at_or_above_threshold": sum(value >= threshold for value in values),
|
||||
"pass_rate": sum(value >= threshold for value in values) / len(values),
|
||||
}
|
||||
|
||||
|
||||
def parse_series(text: str) -> tuple[str, str, Path]:
|
||||
try:
|
||||
scheme, task, path = text.split("=", 2)
|
||||
except ValueError as error:
|
||||
raise argparse.ArgumentTypeError("series must be SCHEME=TASK=PATH") from error
|
||||
return scheme, task, Path(path)
|
||||
|
||||
|
||||
def main() -> int:
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument("--series", action="append", type=parse_series, required=True)
|
||||
parser.add_argument("--output-dir", type=Path, required=True)
|
||||
parser.add_argument("--threshold", type=float, default=0.90)
|
||||
parser.add_argument("--reference", required=True)
|
||||
args = parser.parse_args()
|
||||
|
||||
by_scheme: dict[str, list[dict[str, Any]]] = defaultdict(list)
|
||||
sources = []
|
||||
for scheme, task, path in args.series:
|
||||
payload = json.loads(path.read_text(encoding="utf-8"))
|
||||
pairs = payload["pairs"]
|
||||
if not pairs or {pair["task"] for pair in pairs} != {task}:
|
||||
raise SystemExit(f"task mismatch for {path}: expected {task}")
|
||||
by_scheme[scheme].extend(pairs)
|
||||
sources.append({"scheme": scheme, "task": task, "path": str(path.resolve())})
|
||||
|
||||
video_counts = {scheme: len(pairs) for scheme, pairs in by_scheme.items()}
|
||||
if len(set(video_counts.values())) != 1:
|
||||
raise SystemExit(f"schemes have different video counts: {video_counts}")
|
||||
base_videos = next(iter(video_counts.values()))
|
||||
report: dict[str, Any] = {
|
||||
"metric": "FFmpeg decoded YUV420 SSIM All; base self-comparison = 1.0",
|
||||
"threshold": args.threshold,
|
||||
"reference": args.reference,
|
||||
"sources": sources,
|
||||
"schemes": {
|
||||
"base": {
|
||||
"overall": {
|
||||
"videos": base_videos,
|
||||
"mean_video_ssim": 1.0,
|
||||
"median_video_ssim": 1.0,
|
||||
"p10_video_ssim": 1.0,
|
||||
"min_video_ssim": 1.0,
|
||||
"videos_at_or_above_threshold": base_videos,
|
||||
"pass_rate": 1.0,
|
||||
}
|
||||
}
|
||||
},
|
||||
}
|
||||
table_rows = []
|
||||
for scheme, pairs in sorted(by_scheme.items()):
|
||||
task_groups: dict[str, list[dict[str, Any]]] = defaultdict(list)
|
||||
resolution_groups: dict[int, list[dict[str, Any]]] = defaultdict(list)
|
||||
for pair in pairs:
|
||||
task_groups[str(pair["task"])].append(pair)
|
||||
resolution_groups[int(pair["short_edge"])].append(pair)
|
||||
scheme_result = {
|
||||
"overall": summarize(pairs, args.threshold),
|
||||
"by_task": {
|
||||
task: summarize(group, args.threshold)
|
||||
for task, group in sorted(task_groups.items())
|
||||
},
|
||||
"by_resolution": {
|
||||
str(resolution): summarize(group, args.threshold)
|
||||
for resolution, group in sorted(resolution_groups.items())
|
||||
},
|
||||
}
|
||||
report["schemes"][scheme] = scheme_result
|
||||
for group_type, groups in (
|
||||
("overall", {"all": pairs}),
|
||||
("task", task_groups),
|
||||
("resolution", resolution_groups),
|
||||
):
|
||||
for group, members in groups.items():
|
||||
table_rows.append(
|
||||
{"scheme": scheme, "group_type": group_type, "group": group, **summarize(members, args.threshold)}
|
||||
)
|
||||
|
||||
args.output_dir.mkdir(parents=True, exist_ok=True)
|
||||
(args.output_dir / "comparison.json").write_text(
|
||||
json.dumps(report, ensure_ascii=False, indent=2) + "\n", encoding="utf-8"
|
||||
)
|
||||
columns = (
|
||||
"scheme", "group_type", "group", "videos", "mean_video_ssim",
|
||||
"median_video_ssim", "p10_video_ssim", "min_video_ssim",
|
||||
"videos_at_or_above_threshold", "pass_rate",
|
||||
)
|
||||
with (args.output_dir / "comparison.tsv").open("w", encoding="utf-8") as handle:
|
||||
handle.write("\t".join(columns) + "\n")
|
||||
for row in table_rows:
|
||||
handle.write("\t".join(str(row[column]) for column in columns) + "\n")
|
||||
|
||||
lines = [
|
||||
"# MiniMax-H3 paired SSIM: base vs Cache-DiT vs Larry LoRA",
|
||||
"",
|
||||
f"Reference: `{args.reference}`",
|
||||
"",
|
||||
"Metric: FFmpeg-decoded YUV420 `SSIM All`, paired by request_id after exact prompt/seed/task/resolution/duration/aspect-ratio checks. Base self-comparison is 1.0.",
|
||||
"",
|
||||
"| Scheme | Videos | Mean | Median | Video P10 | Worst video | >= 0.90 |",
|
||||
"|---|---:|---:|---:|---:|---:|---:|",
|
||||
f"| base | {base_videos} | 1.000000 | 1.000000 | 1.000000 | 1.000000 | {base_videos}/{base_videos} |",
|
||||
]
|
||||
for scheme in sorted(by_scheme):
|
||||
value = report["schemes"][scheme]["overall"]
|
||||
lines.append(
|
||||
f"| {scheme} | {value['videos']} | {value['mean_video_ssim']:.6f} | "
|
||||
f"{value['median_video_ssim']:.6f} | {value['p10_video_ssim']:.6f} | "
|
||||
f"{value['min_video_ssim']:.6f} | {value['videos_at_or_above_threshold']}/{value['videos']} |"
|
||||
)
|
||||
lines.extend(["", "## By task", ""])
|
||||
for scheme in sorted(by_scheme):
|
||||
for task, value in report["schemes"][scheme]["by_task"].items():
|
||||
lines.append(
|
||||
f"- {scheme} / {task}: mean={value['mean_video_ssim']:.6f}, "
|
||||
f">=0.90={value['videos_at_or_above_threshold']}/{value['videos']}"
|
||||
)
|
||||
lines.extend(["", "## By resolution", ""])
|
||||
for scheme in sorted(by_scheme):
|
||||
values = report["schemes"][scheme]["by_resolution"]
|
||||
rendered = ", ".join(
|
||||
f"{resolution}p={value['mean_video_ssim']:.6f}"
|
||||
for resolution, value in values.items()
|
||||
)
|
||||
lines.append(f"- {scheme}: {rendered}")
|
||||
(args.output_dir / "README.md").write_text("\n".join(lines) + "\n", encoding="utf-8")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
71
throughput/sglang-base-b300/README.md
Normal file
@ -0,0 +1,71 @@
|
||||
# SGLang B300 原生拓扑矩阵(部署方案一)
|
||||
|
||||
对应调研文档(本地仓库 `sskj/docs/MINIMAX_H3_B300_PLAN1_NATIVE_TOPO.md`):
|
||||
**原生 sglang、无 lossy 优化**,只调 tp / ulysses / 实例数 / 实例内批并发 / 精度档,目标节点级最高吞吐与 GPU 利用率。
|
||||
|
||||
## 口径(与 6000D 报告完全对齐,结果可直接对比)
|
||||
|
||||
- 框架/环境:SGLang(6000D 机 conda env `sglang`,sglang 0.5.17;B300 上机后按实际环境覆盖 `PYTHON`/`SGLANG_BIN`)。
|
||||
- 生成规格:20 inference steps、5 秒、16:9、`flow_shift=12.0`、`audio_flow_shift=3.0`。
|
||||
- 任务:FL2VA、Ref2VA;分辨率:480、720、768、1080;每任务每档 8 条,总量 32 条。
|
||||
- 样本按 prompt 分片均分到每个实例;实例间并行,实例内并发由 `--in-flight` 控制。
|
||||
|
||||
## 与 6000D 矩阵的差异(B300 新增轴)
|
||||
|
||||
| 轴 | 6000D(sglang-base) | B300(本目录) |
|
||||
|---|---|---|
|
||||
| 拓扑 | TP8×1 / TP4×2 / TP2×4(`--ulysses-degree 1`) | Ulysses-8×1、Ulysses-4×2、tp1×8、单卡双实例×8 |
|
||||
| 批并发 | 固定 `--batching-max-size 1` | `--batching-max-size {1,2,4}`(原生吞吐主杠杆) |
|
||||
| 精度 | BF16 | BF16 / FP8(`--quantization fp8`) |
|
||||
| 失败处理 | 严格 die | `SKIP_ON_FAIL=1` 记录后继续(探索期) |
|
||||
|
||||
## 默认矩阵(TOPO_LIST,`name|replicas|tp|ulysses|batching`)
|
||||
|
||||
| topo | 说明 | 依据 |
|
||||
|---|---|---|
|
||||
| `u8x1` | Ulysses-8 × 1 实例,batch1 | 官方 8×B300 验证拓扑(19.04s@BF16) |
|
||||
| `u8x1_b2` / `u8x1_b4` | 单实例批 2 / 批 4 | 官方吞吐档:`--encoder-parallel dp --batching-max-size N` |
|
||||
| `u4x2` / `u4x2_b2` | 2 实例 × Ulysses-4(批 1 / 2) | 折中拓扑 |
|
||||
| `tp1x8` / `tp1x8_b2` | 8 实例 × 单卡驻留(批 1 / 2) | 6000D「多实例并行」结论直译(B300 单卡 288GB 可整模型驻留) |
|
||||
| `share2x8` | 单卡双实例 × 8 卡 = 16 实例(实验项) | 用户点名方向;仅 FP8 档可行 |
|
||||
|
||||
- `GPU_MODE=partition`(默认):实例 i 用卡 `[i*K, (i+1)*K)`,K = tp×ulysses。
|
||||
- `GPU_MODE=share`(share2x8 用):每实例 1 卡,实例 i 用卡 `i % TOTAL_GPUS`(同卡多实例)。
|
||||
- `--encoder-parallel dp` 仅在 batching>1 时追加;batch=1 用默认 `auto`。
|
||||
|
||||
## 用法
|
||||
|
||||
```bash
|
||||
# dry-run:只打印矩阵计划
|
||||
DRY_RUN=1 bash scripts/run_sglang_h3_b300_matrix.sh
|
||||
|
||||
# 冒烟:单 topo、单精度、单任务、少请求
|
||||
TOPO_LIST="u8x1" QUANT_LIST="bf16" TASKS="fl2va" REQUESTS_PER_RESOLUTION=1 \
|
||||
RUN_ID=smoke bash scripts/run_sglang_h3_b300_matrix.sh
|
||||
|
||||
# 正式跑(放 tmux;默认全矩阵 = 8 topo × 2 精度 × 2 任务)
|
||||
tmux new-session -d -s b300-matrix "bash scripts/run_sglang_h3_b300_matrix.sh"
|
||||
```
|
||||
|
||||
常用覆盖变量:`TOTAL_GPUS NUM_INFERENCE_STEPS DURATION_SECONDS TOPO_LIST QUANT_LIST TASKS RESOLUTIONS REQUESTS_PER_RESOLUTION BASE_PORT MODEL REFERENCE_IMAGE PROMPT_FILE PYTHON SGLANG_BIN CLIENT_SCRIPT RUN_ID RESULT_ROOT SKIP_ON_FAIL GPU_MODE`。
|
||||
|
||||
## 结果目录
|
||||
|
||||
```
|
||||
results/<run_id>/
|
||||
├── summary.tsv # 全矩阵一行一 phase(可贴进飞书多维表格)
|
||||
├── orchestrator.log / orchestrator.pid
|
||||
└── <topo>_<quant>/
|
||||
└── <task>/
|
||||
├── server_<i>_port<p>/ # server.log / cuda_visible_devices.txt / outputs/
|
||||
├── client_<i>_port<p>/ # client.log / results.jsonl
|
||||
└── summary.json # 该 phase 汇总
|
||||
```
|
||||
|
||||
summary.tsv 列:`topo prec replicas tp ulysses batching inflight task expected recorded completed failed machine_qps latency_mean_s latency_p95_s machine_wall_s`。
|
||||
|
||||
## 备注
|
||||
|
||||
- 20 步/5s 为 6000D 对比口径;B300 官方 50 步数据见调研文档(u8x1 BF16 19.04s/请求、83.6GB/卡;FP8 18.03s、51.9GB/卡)。
|
||||
- 长片(10/15s)批容量按 token 数等比缩水,另跑专项。
|
||||
- 方案二(Turbo LoRA / SubBlock / Cache-DiT / AdaLN 缓存等优化策略)另行编排,不动本目录口径。
|
||||
191
throughput/sglang-base-b300/REPORT_TEMPLATE.md
Normal file
@ -0,0 +1,191 @@
|
||||
# MiniMax-H3 在 NVIDIA B300 上的 SGLang 部署测试报告(方案一:原生拓扑矩阵)
|
||||
|
||||
> 模板说明:结构完全对齐《MiniMax-H3 在 RTX 6000D 上的 SGLang 多实例部署测试报告》(飞书 wiki ZPtMwtunEiOb39kfV6ocGhp3nLf)。
|
||||
> 所有【待填】处由实验结果填入;数据来源:`/data/wxy/results/minimax_h3_b300_matrix/<run_id>/summary.tsv`(每 phase 一行)与各 `summary.json`。
|
||||
> 6000D 对照基线(同口径 20 步/5s/16:9/480-1080×8/seed=1101+prompt_index)已在各节标注。
|
||||
|
||||
## 1. 结论摘要
|
||||
|
||||
本轮在单台 8×NVIDIA B300 SXM6 服务器上,对 MiniMax-H3 的原生 SGLang 部署进行了等总量、等任务、等分辨率的 serving 测试,覆盖拓扑 × 批并发 × 精度三个轴:
|
||||
- 拓扑:Ulysses-8×1、Ulysses-4×2、tp1×8、单卡双实例×8(16 实例);
|
||||
- 实例内批并发:`--batching-max-size` = 1 / 2 / 4;
|
||||
- 精度:BF16 / FP8。
|
||||
|
||||
每种部署均执行 64 条正式请求(FL2VA 与 Ref2VA 各 32 条;480P/720P/768P/1080P 各 8 条;20 steps、5 秒、16:9)。
|
||||
|
||||
【待填】结论要点:
|
||||
- 吞吐优先的最优拓扑:______(预期候选:u8x1_b2/b4 或 tp1x8_b2;判定依据:machine_qps 与打包率)
|
||||
- 单请求时延最优:______(预期:u8x1 batch1)
|
||||
- FP8 相对 BF16 的吞吐/显存收益:______
|
||||
- 相对 6000D 基线(TP2×4:FL2VA 0.010769 QPS / Ref2VA 0.006234 QPS)的整机提升:______
|
||||
|
||||
## 2. 实验环境与设计
|
||||
|
||||
| 项目 | 固定配置 |
|
||||
|---|---|
|
||||
| 服务器 | Host B300,8×NVIDIA B300 SXM6,单卡 288 GB HBM3e |
|
||||
| 模型 | /data/hf_models/MiniMax-H3 |
|
||||
| 框架与环境 | SGLang;【待填】环境路径(6000D 对照为 /root/.miniconda3/envs/sglang,sglang 0.5.17) |
|
||||
| 任务 | FL2VA、Ref2VA;两个 variant 分阶段启动并顺序测试 |
|
||||
| Prompt | /root/.cache/sglang/vbench_subject_consistency.txt,取 8 条 VBench subject-consistency prompt |
|
||||
| 参考图 | /data/wxy/sskj-MiniMax-H3/assets/reference_images/landscape_mountain_lake.jpg |
|
||||
| 正式生成参数 | 20 inference steps;5 秒;16:9;flow_shift=12.0;audio_flow_shift=3.0;seed=1101+prompt_index |
|
||||
| 分辨率 | short edge 480、720、768、1080;每个任务每档 8 条 |
|
||||
| 预热 | 每实例 1 条、5 steps;预热不计入正式结果 |
|
||||
| 服务并发 | `--batching-max-size` 1/2/4(客户端 in-flight 与服务端同值);不同实例并行 |
|
||||
|
||||
### 2.1 部署方案矩阵(与 6000D 报告的对应关系)
|
||||
|
||||
6000D 报告以"每实例 GPU 数"定义方案(TP8×1 / TP4×2 / TP2×4);B300 单卡 288GB 可整模型驻留,
|
||||
以**同语义的并行档位**对应:8 卡并 = Ulysses-8,4 卡并 = Ulysses-4,单卡 = tp1(Ulysses-1)。
|
||||
|
||||
| 方案 | 实例数 | 每实例 GPU | 并行形态 | 批并发 | 服务端口 |
|
||||
|---|---|---|---|---|---|
|
||||
| u8x1 | 1 | 8 | Ulysses-8(对应 6000D TP8×1) | 1 | 30010 |
|
||||
| u8x1_b2 / u8x1_b4 | 1 | 8 | Ulysses-8 + `--encoder-parallel dp` | 2 / 4 | 30010 |
|
||||
| u4x2 / u4x2_b2 | 2 | 4 | Ulysses-4(对应 6000D TP4×2) | 1 / 2 | 30010、30020 |
|
||||
| tp1x8 / tp1x8_b2 | 8 | 1 | tp1(对应 6000D TP2×4 的"多实例"结构) | 1 / 2 | 30010–30080 |
|
||||
| share2x8 | 16 | 1(单卡双实例) | tp1 × 共享卡(实验项) | 1 | 30010–30160 |
|
||||
|
||||
每方案 × BF16 / FP8 两档;FP8 档追加 `--quantization fp8`。
|
||||
|
||||
### 2.2 样本总量与均衡分配
|
||||
|
||||
同 6000D 口径:每种部署正式总量恒定为 64 条,每任务 32 条(4 分辨率 × 8 prompt);分片按 `prompt_index % num_replicas`,
|
||||
保证每个实例拿到相同数量的 480/720/768/1080 样本。
|
||||
|
||||
### 2.3 指标口径
|
||||
|
||||
- `machine_wall_s`:同任务最早正式请求开始到最晚正式请求结束的整机墙钟时间;
|
||||
- `machine_qps`:成功请求数 ÷ machine_wall_s;
|
||||
- `latency_mean_s / latency_p95_s`:单请求端到端时延(提交→服务端生成→轮询完成),不含 server 启动与预热;
|
||||
- 分辨率级 QPS(对照 4.3/4.4 口径):并发副本数 ÷ 该分辨率平均时延;
|
||||
- 综合 QPS(FL2VA/Ref2VA 各半):`2 × 并发副本数 / (FL2VA 平均延迟 + Ref2VA 平均延迟)`。
|
||||
|
||||
## 3. 启动与测试脚本
|
||||
|
||||
编排脚本:`/data/wxy/sskj-h3/throughput/sglang-base-b300/scripts/run_sglang_h3_b300_matrix.sh`
|
||||
客户端与汇总:`/data/wxy/sskj-h3/throughput/sglang-base-b300/scripts/minimax_h3_b300_bench.py`
|
||||
|
||||
```bash
|
||||
ssh B300
|
||||
cd /data/wxy/sskj-h3/throughput/sglang-base-b300
|
||||
|
||||
DRY_RUN=1 bash scripts/run_sglang_h3_b300_matrix.sh # 打印 32 phase 计划
|
||||
|
||||
RUN_ID="b300-20steps-5s-$(date +%Y%m%d-%H%M%S)" \
|
||||
setsid bash scripts/run_sglang_h3_b300_matrix.sh \
|
||||
>/data/wxy/sglang_b300_matrix.log 2>&1 &
|
||||
|
||||
tail -f /data/wxy/sglang_b300_matrix.log
|
||||
```
|
||||
|
||||
编排脚本对每个 topo 自动计算 replicas、分配连续 GPU(`GPU_MODE=partition`;`share2x8` 需 `GPU_MODE=share`),
|
||||
独立端口,等待 `/health` 后启动同数量 client;FL2VA 完成后释放服务再切 Ref2VA。
|
||||
不可行组合(如 BF16 单卡双实例)在 `SKIP_ON_FAIL=1` 下记录后跳过。核心服务启动参数(u8x1 示例):
|
||||
|
||||
```bash
|
||||
CUDA_VISIBLE_DEVICES="0,1,2,3,4,5,6,7" \
|
||||
sglang serve \
|
||||
--model-path /data/hf_models/MiniMax-H3 \
|
||||
--model-variant "$variant" \
|
||||
--backend sglang \
|
||||
--performance-mode speed \
|
||||
--num-gpus 8 --tp-size 1 --ulysses-degree 8 \
|
||||
--use-fsdp-inference false \
|
||||
--enable-torch-compile false \
|
||||
--batching-max-size 1 --batching-delay-ms 0 \
|
||||
--warmup-resolutions 1344x768 \
|
||||
--host 0.0.0.0 --port "$port"
|
||||
```
|
||||
|
||||
客户端通过 SGLang 异步 `POST /v1/videos` 提交、`GET /v1/videos/{id}` 轮询;每条请求写入 JSONL,
|
||||
阶段结束后聚合为 summary.json 与 summary.tsv(列:topo prec replicas tp ulysses batching inflight task expected recorded completed failed machine_qps latency_mean_s latency_p95_s machine_wall_s)。
|
||||
|
||||
## 4. 测试结果
|
||||
|
||||
### 4.1 整机吞吐与端到端时延(对照 6000D 报告 4.1)
|
||||
|
||||
| 方案 | 任务 | 成功/总数 | machine QPS | 平均时延(s) | P95(s) | 墙钟(s) |
|
||||
|---|---|---|---|---|---|---|
|
||||
| u8x1 BF16 | FL2VA | 【待填】 | 【待填】 | 【待填】 | 【待填】 | 【待填】 |
|
||||
| u8x1 BF16 | Ref2VA | 【待填】 | 【待填】 | 【待填】 | 【待填】 | 【待填】 |
|
||||
| u8x1_b2 BF16 | FL2VA | 【待填】 | 【待填】 | 【待填】 | 【待填】 | 【待填】 |
|
||||
| u8x1_b4 FP8 | FL2VA | 【待填】 | 【待填】 | 【待填】 | 【待填】 | 【待填】 |
|
||||
| u4x2 BF16 | FL2VA | 【待填】 | 【待填】 | 【待填】 | 【待填】 | 【待填】 |
|
||||
| tp1x8 BF16 | FL2VA | 【待填】 | 【待填】 | 【待填】 | 【待填】 | 【待填】 |
|
||||
| tp1x8_b2 FP8 | FL2VA | 【待填】 | 【待填】 | 【待填】 | 【待填】 | 【待填】 |
|
||||
| share2x8 FP8 | FL2VA | 【待填】 | 【待填】 | 【待填】 | 【待填】 | 【待填】 |
|
||||
| ...(其余 phase 同构) | | | | | | |
|
||||
|
||||
6000D 对照(同口径):TP8×1 FL2VA 0.007646(130.79s)/ Ref2VA 0.004580;TP4×2 0.009723;TP2×4 0.010769 / 0.006234。
|
||||
|
||||
### 4.2 分辨率平均时延
|
||||
|
||||
| 任务 | 方案 | 480 | 720 | 768 | 1080 |
|
||||
|---|---|---|---|---|---|
|
||||
| FL2VA | u8x1 BF16 | 【待填】 | 【待填】 | 【待填】 | 【待填】 |
|
||||
| FL2VA | u8x1_b4 FP8 | 【待填】 | 【待填】 | 【待填】 | 【待填】 |
|
||||
| FL2VA | tp1x8 BF16 | 【待填】 | 【待填】 | 【待填】 | 【待填】 |
|
||||
| FL2VA | ... | | | | |
|
||||
| Ref2VA | ... | | | | |
|
||||
|
||||
(数据取 summary.json 的 `by_short_edge.latency_mean_s`。)
|
||||
|
||||
### 4.3 / 4.4 分辨率整机 QPS(FL2VA / Ref2VA)
|
||||
|
||||
| SGLang 配置 | 480p | 720p | 768p | 1080p |
|
||||
|---|---|---|---|---|
|
||||
| 【方案】 | 【待填】 | 【待填】 | 【待填】 | 【待填】 |
|
||||
|
||||
### 4.5 FL2VA 和 Ref2VA 各占 50% 时的综合 QPS
|
||||
|
||||
公式同 6000D 报告(2 × 并发副本数 / 时延和)。表:【待填】。
|
||||
预期观察点:1080 档是否再次成为长尾主因、批并发(batching>1)是否改变 1080 的相对惩罚。
|
||||
|
||||
## 5. 拓扑比较与建议(分析框架,结论待测)
|
||||
|
||||
| 维度 | 6000D 结论(TP8/TP4/TP2 序列) | B300 预期/待测 |
|
||||
|---|---|---|
|
||||
| 吞吐最优 | TP2×4(多实例) | 【待测】候选 u8x1_b2/b4、tp1x8_b2 |
|
||||
| 时延最优 | TP8×1 | 【待测】候选 u8x1 batch1 |
|
||||
| FP8 | 未测 | 【待测】显存 -38% → 批容量翻倍 |
|
||||
| 1080p 惩罚 | 拆分越细惩罚越明显 | 【待测】批并发是否能摊薄 |
|
||||
|
||||
建议框架(沿用 6000D 报告):
|
||||
- 若以"每台机器每天完成条数"为核心 → 选吞吐最优档;若兼顾等待时间 → 折中档(u4x2 或 u8x1_b2);
|
||||
- 1080 与低分辨率拆池,避免 head-of-line blocking;
|
||||
- 按真实流量比例做并发队列测试后再定最终档位。
|
||||
|
||||
## 6. 与 6000D / vLLM-Omni 的同口径对照
|
||||
|
||||
### 6.1 与 6000D(RTX 6000D 8 卡)横向对照
|
||||
|
||||
同任务/prompt/分辨率/steps/时长/种子、按整机口径:
|
||||
|
||||
| 部署 | 任务 | 6000D QPS | B300 QPS | B300 相对 |
|
||||
|---|---|---|---|---|
|
||||
| 8×1 卡实例(6000D TP8×1 ↔ B300 u8x1) | FL2VA | 0.007646 | 【待填】 | |
|
||||
| 8×1 卡实例 | Ref2VA | 0.004580 | 【待填】 | |
|
||||
| 4×2 卡实例(TP4×2 ↔ u4x2) | FL2VA | 0.009723 | 【待填】 | |
|
||||
| 2×1 卡实例 ×4(TP2×4 ↔ tp1x8) | FL2VA | 0.010769 | 【待填】 | |
|
||||
|
||||
### 6.2 与 vLLM-Omni 对照
|
||||
|
||||
【待填】B300 上 vLLM-Omni 同口径结果(6000D 上 vLLM-Omni 内部为 DiT TP2×USP4/2/1,多数格子快于 SGLang +25.68%/+47.59%/+24.91%/+22.09%)。
|
||||
注意两套 API/编码链路不同(SGLang 异步 job 轮询 vs vLLM-Omni 同步 MP4),仅用于整机容量判断;视频质量需另行盲评/VBench。
|
||||
|
||||
## 7. 优化策略对齐(方案二占位)
|
||||
|
||||
6000D 报告 Cache-DiT 对照(同请求同 seed,TP2×4 上叠加):
|
||||
- 配置:RDT=0.1 / MC=4 / Fn=2 / Bn=0 / W=2 / SCM=dynamic + h3_cache_patch;
|
||||
- 效果:整机 QPS +95.3%(FL2VA)/ +121.1%(Ref2VA);相对 TP8×1 基线达 2.75× / 3.01×;各分辨率一致受益(1.97×–2.39×),1080p 无长尾恶化。
|
||||
|
||||
B300 待跑:在 B300 最优拓扑上叠加同类 Cache-DiT / Turbo LoRA / SubBlock 等优化,对齐 6000D 表格结构出"表格 1~6"。
|
||||
(优化策略参考与最终档位待定,本报告先以原生拓扑矩阵收口。)
|
||||
|
||||
## 8. 结果与审计文件
|
||||
|
||||
- 本轮结果根目录:`/data/wxy/results/minimax_h3_b300_matrix/<run_id>/`
|
||||
- 汇总文件:各 run 根目录 `summary.tsv`;每个 phase 目录 `summary.json`;每个 client 目录逐请求 `results.jsonl` 与日志;每个 server 目录 `server.log`、`cuda_visible_devices.txt`、`outputs/`
|
||||
- 完整脚本:`/data/wxy/sskj-h3/throughput/sglang-base-b300/scripts/run_sglang_h3_b300_matrix.sh`、`.../minimax_h3_b300_bench.py`
|
||||
374
throughput/sglang-base-b300/scripts/minimax_h3_b300_bench.py
Executable file
@ -0,0 +1,374 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Run or summarize a stratified MiniMax-H3 FL2VA/Ref2VA serving workload.
|
||||
|
||||
B300 variant: adds per-instance in-flight concurrency (--in-flight) so the
|
||||
server-side --batching-max-size can be exercised, and extends the summary
|
||||
with topo/precision/batching columns. In-flight=1 reproduces the 6000D
|
||||
serial-per-instance behaviour exactly.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import math
|
||||
import queue
|
||||
import statistics
|
||||
import threading
|
||||
import time
|
||||
from concurrent.futures import ThreadPoolExecutor
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
import requests
|
||||
|
||||
|
||||
DEFAULT_PROMPT = "A cinematic landscape with natural motion and realistic lighting."
|
||||
|
||||
|
||||
def percentile(values: list[float], q: float) -> float:
|
||||
if not values:
|
||||
return 0.0
|
||||
values = sorted(values)
|
||||
pos = (len(values) - 1) * q
|
||||
lo, hi = math.floor(pos), math.ceil(pos)
|
||||
if lo == hi:
|
||||
return values[lo]
|
||||
return values[lo] * (hi - pos) + values[hi] * (pos - lo)
|
||||
|
||||
|
||||
def load_prompts(path: Path, count: int) -> list[str]:
|
||||
prompts: list[str] = []
|
||||
if path.is_file():
|
||||
prompts = [line.strip() for line in path.read_text(encoding="utf-8").splitlines() if line.strip()]
|
||||
if not prompts:
|
||||
prompts = [DEFAULT_PROMPT]
|
||||
repeats = (count + len(prompts) - 1) // len(prompts)
|
||||
return (prompts * repeats)[:count]
|
||||
|
||||
|
||||
def build_plan(args: argparse.Namespace) -> list[dict[str, Any]]:
|
||||
resolutions = [int(item) for item in args.resolutions.split(",") if item.strip()]
|
||||
prompts = load_prompts(args.prompt_file, args.requests_per_resolution)
|
||||
plan: list[dict[str, Any]] = []
|
||||
# Interleave resolutions so any slow drift affects every bucket similarly.
|
||||
for prompt_index, prompt in enumerate(prompts):
|
||||
for short_edge in resolutions:
|
||||
plan.append(
|
||||
{
|
||||
"request_id": f"{args.task}-r{short_edge}-p{prompt_index:02d}",
|
||||
"task": args.task,
|
||||
"short_edge": short_edge,
|
||||
"prompt_index": prompt_index,
|
||||
"prompt": prompt,
|
||||
"seed": args.seed + prompt_index,
|
||||
}
|
||||
)
|
||||
return plan
|
||||
|
||||
|
||||
def make_payload(args: argparse.Namespace, item: dict[str, Any], steps: int) -> dict[str, Any]:
|
||||
condition: dict[str, Any] = {
|
||||
"type": "image",
|
||||
"uri": str(args.reference_image),
|
||||
"role": "keyframe" if args.task == "fl2va" else "reference",
|
||||
}
|
||||
if args.task == "fl2va":
|
||||
condition["frame_index"] = 0
|
||||
return {
|
||||
"model": args.model,
|
||||
"prompt": item["prompt"],
|
||||
"num_outputs_per_prompt": 1,
|
||||
"num_inference_steps": steps,
|
||||
"flow_shift": args.flow_shift,
|
||||
"audio_flow_shift": args.audio_flow_shift,
|
||||
"seed": item["seed"],
|
||||
"task": args.task,
|
||||
"conditions": [condition],
|
||||
"target": {
|
||||
"short_edge": item["short_edge"],
|
||||
"aspect_ratio": args.aspect_ratio,
|
||||
"duration_seconds": args.duration_seconds,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def run_one(
|
||||
session: requests.Session,
|
||||
args: argparse.Namespace,
|
||||
item: dict[str, Any],
|
||||
steps: int,
|
||||
) -> dict[str, Any]:
|
||||
started_epoch = time.time()
|
||||
started = time.monotonic()
|
||||
result: dict[str, Any] = {
|
||||
**item,
|
||||
"replica_index": args.replica_index,
|
||||
"port": args.port,
|
||||
"num_inference_steps": steps,
|
||||
"duration_seconds": args.duration_seconds,
|
||||
"aspect_ratio": args.aspect_ratio,
|
||||
"started_at_epoch": started_epoch,
|
||||
"success": False,
|
||||
"error": None,
|
||||
}
|
||||
try:
|
||||
response = session.post(
|
||||
f"http://{args.host}:{args.port}/v1/videos",
|
||||
json=make_payload(args, item, steps),
|
||||
timeout=args.submit_timeout,
|
||||
)
|
||||
if response.status_code != 200:
|
||||
raise RuntimeError(f"submit HTTP {response.status_code}: {response.text[:1000]}")
|
||||
status = response.json()
|
||||
video_id = status.get("id")
|
||||
if not video_id:
|
||||
raise RuntimeError(f"submit response has no id: {status}")
|
||||
result["video_id"] = video_id
|
||||
deadline = time.monotonic() + args.request_timeout
|
||||
while status.get("status") not in {"completed", "failed"}:
|
||||
if time.monotonic() >= deadline:
|
||||
raise TimeoutError(f"video job {video_id} exceeded {args.request_timeout}s")
|
||||
time.sleep(args.poll_interval)
|
||||
poll = session.get(
|
||||
f"http://{args.host}:{args.port}/v1/videos/{video_id}",
|
||||
timeout=args.poll_timeout,
|
||||
)
|
||||
if poll.status_code != 200:
|
||||
raise RuntimeError(f"poll HTTP {poll.status_code}: {poll.text[:1000]}")
|
||||
status = poll.json()
|
||||
if status.get("status") != "completed":
|
||||
raise RuntimeError(f"job failed: {status.get('error') or status}")
|
||||
result["success"] = True
|
||||
result["inference_time_s"] = status.get("inference_time_s")
|
||||
result["peak_memory_mb"] = status.get("peak_memory_mb")
|
||||
result["file_path"] = status.get("file_path")
|
||||
except Exception as exc: # Keep the rest of the matrix running and record the cell failure.
|
||||
result["error"] = f"{type(exc).__name__}: {exc}"
|
||||
result["latency_s"] = time.monotonic() - started
|
||||
result["finished_at_epoch"] = time.time()
|
||||
return result
|
||||
|
||||
|
||||
def run_command(args: argparse.Namespace) -> int:
|
||||
if not args.reference_image.is_file():
|
||||
raise SystemExit(f"reference image not found: {args.reference_image}")
|
||||
full_plan = build_plan(args)
|
||||
# Stratify by prompt index so every replica receives the same number of
|
||||
# samples from every resolution. This avoids assigning an entire slow
|
||||
# resolution bucket (for example 1080p) to only one replica.
|
||||
shard = [
|
||||
item
|
||||
for item in full_plan
|
||||
if item["prompt_index"] % args.num_replicas == args.replica_index
|
||||
]
|
||||
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||
completed_ids: set[str] = set()
|
||||
if args.output.is_file():
|
||||
for line in args.output.read_text(encoding="utf-8").splitlines():
|
||||
try:
|
||||
completed_ids.add(json.loads(line)["request_id"])
|
||||
except (json.JSONDecodeError, KeyError):
|
||||
continue
|
||||
shard = [item for item in shard if item["request_id"] not in completed_ids]
|
||||
print(
|
||||
f"task={args.task} replica={args.replica_index}/{args.num_replicas} "
|
||||
f"requests={len(shard)} port={args.port} in_flight={args.in_flight}",
|
||||
flush=True,
|
||||
)
|
||||
failures = 0
|
||||
write_lock = threading.Lock()
|
||||
|
||||
|
||||
with requests.Session() as session:
|
||||
# Warmup stays serial so a slow first request cannot stall concurrency probes.
|
||||
for warmup_index in range(args.warmup_requests):
|
||||
warmup_item = (shard or full_plan)[warmup_index % len(shard or full_plan)].copy()
|
||||
warmup_item["request_id"] = f"warmup-{warmup_index}-{warmup_item['request_id']}"
|
||||
warmup = run_one(session, args, warmup_item, args.warmup_inference_steps)
|
||||
print(
|
||||
f"warmup {warmup_index + 1}/{args.warmup_requests}: "
|
||||
f"success={warmup['success']} latency={warmup['latency_s']:.2f}s "
|
||||
f"error={warmup['error']}",
|
||||
flush=True,
|
||||
)
|
||||
if not warmup["success"]:
|
||||
raise SystemExit("warmup failed")
|
||||
with write_lock, args.output.open("a", encoding="utf-8") as _out:
|
||||
_out.write(json.dumps(warmup, ensure_ascii=False) + "\n")
|
||||
if not shard:
|
||||
return int(failures > 0)
|
||||
if args.in_flight <= 1:
|
||||
with args.output.open("a", encoding="utf-8", buffering=1) as output:
|
||||
for index, item in enumerate(shard, start=1):
|
||||
result = run_one(session, args, item, args.num_inference_steps)
|
||||
output.write(json.dumps(result, ensure_ascii=False) + "\n")
|
||||
failures += int(not result["success"])
|
||||
print(
|
||||
f"request {index}/{len(shard)} id={item['request_id']} "
|
||||
f"success={result['success']} latency={result['latency_s']:.2f}s "
|
||||
f"error={result['error']}",
|
||||
flush=True,
|
||||
)
|
||||
return int(failures > 0)
|
||||
|
||||
# Concurrent: fixed in-flight window over a worker pool.
|
||||
task_queue: queue.Queue[dict[str, Any] | None] = queue.Queue()
|
||||
for item in shard:
|
||||
task_queue.put(item)
|
||||
for _ in range(args.in_flight):
|
||||
task_queue.put(None) # sentinel
|
||||
|
||||
completed = 0
|
||||
|
||||
def worker() -> None:
|
||||
nonlocal completed
|
||||
with args.output.open("a", encoding="utf-8", buffering=1) as output:
|
||||
while True:
|
||||
item = task_queue.get()
|
||||
if item is None:
|
||||
task_queue.task_done()
|
||||
return
|
||||
result = run_one(session, args, item, args.num_inference_steps)
|
||||
with write_lock:
|
||||
output.write(json.dumps(result, ensure_ascii=False) + "\n")
|
||||
completed += 1
|
||||
failures += int(not result["success"])
|
||||
print(
|
||||
f"request {completed}/{len(shard)} id={item['request_id']} "
|
||||
f"success={result['success']} latency={result['latency_s']:.2f}s "
|
||||
f"error={result['error']}",
|
||||
flush=True,
|
||||
)
|
||||
task_queue.task_done()
|
||||
|
||||
with ThreadPoolExecutor(max_workers=args.in_flight) as pool:
|
||||
futures = [pool.submit(worker) for _ in range(args.in_flight)]
|
||||
for future in futures:
|
||||
future.result()
|
||||
return int(failures > 0)
|
||||
|
||||
|
||||
def summarize_command(args: argparse.Namespace) -> int:
|
||||
rows: list[dict[str, Any]] = []
|
||||
for path in sorted(args.input_dir.glob("client_*/results.jsonl")):
|
||||
for line in path.read_text(encoding="utf-8").splitlines():
|
||||
if line.strip():
|
||||
rows.append(json.loads(line))
|
||||
successful = [row for row in rows if row.get("success")]
|
||||
latencies = [float(row["latency_s"]) for row in successful]
|
||||
started = [float(row["started_at_epoch"]) for row in rows]
|
||||
finished = [float(row["finished_at_epoch"]) for row in rows]
|
||||
wall_s = max(finished) - min(started) if started and finished else 0.0
|
||||
buckets: dict[str, dict[str, Any]] = {}
|
||||
for short_edge in sorted({int(row["short_edge"]) for row in rows}):
|
||||
bucket_rows = [row for row in rows if int(row["short_edge"]) == short_edge]
|
||||
bucket_success = [row for row in bucket_rows if row.get("success")]
|
||||
bucket_latencies = [float(row["latency_s"]) for row in bucket_success]
|
||||
buckets[str(short_edge)] = {
|
||||
"requests": len(bucket_rows),
|
||||
"completed": len(bucket_success),
|
||||
"failed": len(bucket_rows) - len(bucket_success),
|
||||
"latency_mean_s": statistics.fmean(bucket_latencies) if bucket_latencies else 0.0,
|
||||
"latency_p95_s": percentile(bucket_latencies, 0.95),
|
||||
}
|
||||
summary = {
|
||||
"topo": args.topo,
|
||||
"prec": args.prec,
|
||||
"replicas": args.replicas,
|
||||
"tp": args.tp,
|
||||
"ulysses": args.ulysses,
|
||||
"batching": args.batching,
|
||||
"in_flight": args.in_flight,
|
||||
"task": args.task,
|
||||
"expected_requests": args.expected_requests,
|
||||
"requests_recorded": len(rows),
|
||||
"completed": len(successful),
|
||||
"failed": len(rows) - len(successful),
|
||||
"machine_wall_s": wall_s,
|
||||
"machine_qps": len(successful) / wall_s if wall_s else 0.0,
|
||||
"latency_mean_s": statistics.fmean(latencies) if latencies else 0.0,
|
||||
"latency_p50_s": percentile(latencies, 0.50),
|
||||
"latency_p95_s": percentile(latencies, 0.95),
|
||||
"by_short_edge": buckets,
|
||||
}
|
||||
args.output.parent.mkdir(parents=True, exist_ok=True)
|
||||
args.output.write_text(json.dumps(summary, ensure_ascii=False, indent=2), encoding="utf-8")
|
||||
print(
|
||||
"\t".join(
|
||||
[
|
||||
str(args.topo),
|
||||
str(args.prec),
|
||||
str(args.replicas),
|
||||
str(args.tp),
|
||||
str(args.ulysses),
|
||||
str(args.batching),
|
||||
str(args.in_flight),
|
||||
args.task,
|
||||
str(args.expected_requests),
|
||||
str(len(rows)),
|
||||
str(len(successful)),
|
||||
str(len(rows) - len(successful)),
|
||||
f"{summary['machine_qps']:.8f}",
|
||||
f"{summary['latency_mean_s']:.6f}",
|
||||
f"{summary['latency_p95_s']:.6f}",
|
||||
f"{wall_s:.3f}",
|
||||
]
|
||||
)
|
||||
)
|
||||
return int(len(rows) != args.expected_requests or len(successful) != len(rows))
|
||||
|
||||
|
||||
def parse_args() -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser()
|
||||
subparsers = parser.add_subparsers(dest="command", required=True)
|
||||
run = subparsers.add_parser("run")
|
||||
run.add_argument("--host", default="127.0.0.1")
|
||||
run.add_argument("--port", type=int, required=True)
|
||||
run.add_argument("--model", default="/data/hf_models/MiniMax-H3")
|
||||
run.add_argument("--task", choices=["fl2va", "ref2va"], required=True)
|
||||
run.add_argument("--reference-image", type=Path, required=True)
|
||||
run.add_argument("--prompt-file", type=Path, default=Path.home() / ".cache/sglang/vbench_subject_consistency.txt")
|
||||
run.add_argument("--resolutions", default="480,720,768,1080")
|
||||
run.add_argument("--requests-per-resolution", type=int, default=8)
|
||||
run.add_argument("--replica-index", type=int, required=True)
|
||||
run.add_argument("--num-replicas", type=int, required=True)
|
||||
run.add_argument("--in-flight", type=int, default=1)
|
||||
run.add_argument("--num-inference-steps", type=int, default=20)
|
||||
run.add_argument("--warmup-requests", type=int, default=1)
|
||||
run.add_argument("--warmup-inference-steps", type=int, default=5)
|
||||
run.add_argument("--duration-seconds", type=float, default=5.0)
|
||||
run.add_argument("--aspect-ratio", default="16:9")
|
||||
run.add_argument("--flow-shift", type=float, default=12.0)
|
||||
run.add_argument("--audio-flow-shift", type=float, default=3.0)
|
||||
run.add_argument("--seed", type=int, default=1101)
|
||||
run.add_argument("--submit-timeout", type=float, default=120.0)
|
||||
run.add_argument("--poll-timeout", type=float, default=30.0)
|
||||
run.add_argument("--poll-interval", type=float, default=1.0)
|
||||
run.add_argument("--request-timeout", type=float, default=3600.0)
|
||||
run.add_argument("--output", type=Path, required=True)
|
||||
run.set_defaults(func=run_command)
|
||||
|
||||
summarize = subparsers.add_parser("summarize")
|
||||
summarize.add_argument("--input-dir", type=Path, required=True)
|
||||
summarize.add_argument("--output", type=Path, required=True)
|
||||
summarize.add_argument("--task", required=True)
|
||||
summarize.add_argument("--topo", required=True)
|
||||
summarize.add_argument("--prec", default="bf16")
|
||||
summarize.add_argument("--tp", type=int, required=True)
|
||||
summarize.add_argument("--ulysses", type=int, required=True)
|
||||
summarize.add_argument("--replicas", type=int, required=True)
|
||||
summarize.add_argument("--batching", type=int, required=True)
|
||||
summarize.add_argument("--in-flight", type=int, required=True)
|
||||
summarize.add_argument("--expected-requests", type=int, required=True)
|
||||
summarize.set_defaults(func=summarize_command)
|
||||
return parser.parse_args()
|
||||
|
||||
|
||||
def main() -> int:
|
||||
args = parse_args()
|
||||
return args.func(args)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
232
throughput/sglang-base-b300/scripts/run_sglang_h3_b300_matrix.sh
Executable file
@ -0,0 +1,232 @@
|
||||
#!/usr/bin/env bash
|
||||
# B300 native-topology matrix: topo x precision x task.
|
||||
# Topology axes: tp / ulysses / replicas / instance batching / quantization.
|
||||
# 32 requests per task: 2 tasks x 4 short-edge resolutions x 8 prompts (same gauge as 6000D).
|
||||
set -Eeuo pipefail
|
||||
|
||||
TOTAL_GPUS=${TOTAL_GPUS:-8}
|
||||
NUM_INFERENCE_STEPS=${NUM_INFERENCE_STEPS:-20}
|
||||
DURATION_SECONDS=${DURATION_SECONDS:-5}
|
||||
# name|replicas|tp|ulysses|batching
|
||||
TOPO_LIST=${TOPO_LIST:-"u8x1|1|1|8|1 u8x1_b2|1|1|8|2 u8x1_b4|1|1|8|4 u4x2|2|1|4|1 u4x2_b2|2|1|4|2 tp1x8|8|1|1|1 tp1x8_b2|8|1|1|2 share2x8|16|1|1|1"}
|
||||
QUANT_LIST=${QUANT_LIST:-"bf16 fp8"}
|
||||
TASKS=${TASKS:-"fl2va ref2va"}
|
||||
RESOLUTIONS=${RESOLUTIONS:-"480,720,768,1080"}
|
||||
REQUESTS_PER_RESOLUTION=${REQUESTS_PER_RESOLUTION:-8}
|
||||
REQUESTS_PER_TASK=$((REQUESTS_PER_RESOLUTION * 4))
|
||||
GPU_MODE=${GPU_MODE:-partition} # partition | share (share requires tp*ulysses==1)
|
||||
SKIP_ON_FAIL=${SKIP_ON_FAIL:-1} # 1: record phase failure and continue (exploration mode)
|
||||
|
||||
BASE_PORT=${BASE_PORT:-30010}
|
||||
PORT_STRIDE=${PORT_STRIDE:-10}
|
||||
MASTER_PORT_BASE=${MASTER_PORT_BASE:-31000}
|
||||
SCHEDULER_PORT_BASE=${SCHEDULER_PORT_BASE:-32000}
|
||||
HOST=${HOST:-127.0.0.1}
|
||||
MODEL=${MODEL:-/data/hf_models/MiniMax-H3}
|
||||
REFERENCE_IMAGE=${REFERENCE_IMAGE:-/data/wxy/sskj-MiniMax-H3/assets/reference_images/landscape_mountain_lake.jpg}
|
||||
PROMPT_FILE=${PROMPT_FILE:-/root/.cache/sglang/vbench_subject_consistency.txt}
|
||||
PYTHON=${PYTHON:-/root/.miniconda3/envs/sglang/bin/python}
|
||||
SGLANG_BIN=${SGLANG_BIN:-/root/.miniconda3/envs/sglang/bin/sglang}
|
||||
CLIENT_SCRIPT=${CLIENT_SCRIPT:-/data/wxy/sskj-h3/throughput/sglang-base-b300/scripts/minimax_h3_b300_bench.py}
|
||||
SERVER_START_TIMEOUT=${SERVER_START_TIMEOUT:-1800}
|
||||
RUN_ID=${RUN_ID:-b300-matrix-$(date '+%Y%m%d-%H%M%S')}
|
||||
RESULT_ROOT=${RESULT_ROOT:-/data/wxy/results/minimax_h3_b300_matrix/$RUN_ID}
|
||||
|
||||
declare -a SERVER_PIDS=()
|
||||
declare -a CLIENT_PIDS=()
|
||||
|
||||
log() { printf '[%s] %s\n' "$(date '+%F %T')" "$*"; }
|
||||
die() { log "ERROR: $*" >&2; exit 1; }
|
||||
|
||||
[[ -x "$PYTHON" ]] || die "python not executable: $PYTHON"
|
||||
[[ -x "$SGLANG_BIN" ]] || die "sglang not executable: $SGLANG_BIN"
|
||||
[[ -f "$CLIENT_SCRIPT" ]] || die "client script missing: $CLIENT_SCRIPT"
|
||||
[[ -f "$REFERENCE_IMAGE" ]] || die "reference image missing: $REFERENCE_IMAGE"
|
||||
mkdir -p "$RESULT_ROOT"
|
||||
SUMMARY_TSV="$RESULT_ROOT/summary.tsv"
|
||||
printf 'topo\tprec\treplicas\ttp\tulysses\tbatching\tinflight\ttask\texpected\trecorded\tcompleted\tfailed\tmachine_qps\tlatency_mean_s\tlatency_p95_s\tmachine_wall_s\n' > "$SUMMARY_TSV"
|
||||
|
||||
port_is_open() {
|
||||
"$PYTHON" - "$HOST" "$1" <<'PY'
|
||||
import socket, sys
|
||||
s = socket.socket(); s.settimeout(0.5)
|
||||
try: s.connect((sys.argv[1], int(sys.argv[2])))
|
||||
except OSError: raise SystemExit(1)
|
||||
else: raise SystemExit(0)
|
||||
finally: s.close()
|
||||
PY
|
||||
}
|
||||
|
||||
stop_servers() {
|
||||
local pid alive deadline
|
||||
((${#SERVER_PIDS[@]})) || return 0
|
||||
log "gracefully stopping ${#SERVER_PIDS[@]} server(s)"
|
||||
for pid in "${SERVER_PIDS[@]}"; do kill -INT "$pid" 2>/dev/null || true; done
|
||||
deadline=$((SECONDS + 120))
|
||||
while ((SECONDS < deadline)); do
|
||||
alive=0
|
||||
for pid in "${SERVER_PIDS[@]}"; do kill -0 "$pid" 2>/dev/null && alive=1; done
|
||||
((alive == 0)) && break
|
||||
sleep 2
|
||||
done
|
||||
for pid in "${SERVER_PIDS[@]}"; do
|
||||
if kill -0 "$pid" 2>/dev/null; then
|
||||
log "server pid=$pid did not exit after SIGINT; terminating process group"
|
||||
kill -TERM -- "-$pid" 2>/dev/null || kill -TERM "$pid" 2>/dev/null || true
|
||||
sleep 5
|
||||
kill -KILL -- "-$pid" 2>/dev/null || kill -KILL "$pid" 2>/dev/null || true
|
||||
fi
|
||||
wait "$pid" 2>/dev/null || true
|
||||
done
|
||||
SERVER_PIDS=()
|
||||
}
|
||||
|
||||
cleanup() {
|
||||
local rc=$? pid
|
||||
trap - EXIT INT TERM
|
||||
for pid in "${CLIENT_PIDS[@]}"; do kill -TERM "$pid" 2>/dev/null || true; done
|
||||
stop_servers
|
||||
exit "$rc"
|
||||
}
|
||||
trap cleanup EXIT INT TERM
|
||||
|
||||
wait_healthy() {
|
||||
local port=$1 pid=$2 log_file=$3 deadline=$((SECONDS + SERVER_START_TIMEOUT))
|
||||
while ((SECONDS < deadline)); do
|
||||
curl -fsS --max-time 5 "http://${HOST}:${port}/health" >/dev/null 2>&1 && return 0
|
||||
if ! kill -0 "$pid" 2>/dev/null; then tail -100 "$log_file" >&2 || true; return 1; fi
|
||||
sleep 5
|
||||
done
|
||||
tail -100 "$log_file" >&2 || true
|
||||
return 1
|
||||
}
|
||||
|
||||
gpu_csv_for() {
|
||||
# $1=replica_index $2=replicas $3=gpus_per_instance -> prints CUDA_VISIBLE_DEVICES csv
|
||||
local replica=$1 replicas=$2 k=$3 gpu_csv="" offset gpu
|
||||
if [[ "$GPU_MODE" == share ]]; then
|
||||
printf '%s' "$((replica % TOTAL_GPUS))"
|
||||
return 0
|
||||
fi
|
||||
for ((offset=0; offset<k; offset++)); do
|
||||
gpu=$((replica * k + offset)); [[ -z "$gpu_csv" ]] && gpu_csv="$gpu" || gpu_csv+=",$gpu"
|
||||
done
|
||||
printf '%s' "$gpu_csv"
|
||||
}
|
||||
|
||||
start_servers() {
|
||||
local topo=$1 prec=$2 replicas=$3 tp=$4 ulysses=$5 batching=$6 variant=$7 phase_dir=$8
|
||||
local k=$((tp * ulysses)) replica port master_port scheduler_port gpu_csv server_dir server_log candidate enc_flag
|
||||
[[ -n "$batching" && "$batching" -gt 1 ]] && enc_flag="--encoder-parallel dp" || enc_flag=""
|
||||
SERVER_PIDS=()
|
||||
for ((replica=0; replica<replicas; replica++)); do
|
||||
port=$((BASE_PORT + replica * PORT_STRIDE))
|
||||
master_port=$((MASTER_PORT_BASE + replica * PORT_STRIDE))
|
||||
scheduler_port=$((SCHEDULER_PORT_BASE + replica * PORT_STRIDE))
|
||||
for candidate in "$port" "$((port+1))" "$master_port" "$scheduler_port"; do
|
||||
port_is_open "$candidate" && die "port already in use: $candidate"
|
||||
done
|
||||
gpu_csv=$(gpu_csv_for "$replica" "$replicas" "$k")
|
||||
server_dir="$phase_dir/server_${replica}_port${port}"; mkdir -p "$server_dir/outputs"
|
||||
server_log="$server_dir/server.log"; printf '%s\n' "$gpu_csv" > "$server_dir/cuda_visible_devices.txt"
|
||||
log "starting topo=$topo prec=$prec variant=$variant replica=$replica GPUs=$gpu_csv port=$port batching=$batching$([[ -n "$enc_flag" ]] && echo " [encoder-dp]")"
|
||||
CUDA_VISIBLE_DEVICES="$gpu_csv" PYTHONUNBUFFERED=1 TOKENIZERS_PARALLELISM=false \
|
||||
SGLANG_USE_RUNAI_MODEL_STREAMER=false setsid "$SGLANG_BIN" serve \
|
||||
--model-path "$MODEL" --model-variant "$variant" --backend sglang --performance-mode speed \
|
||||
--num-gpus "$k" --tp-size "$tp" --ulysses-degree "$ulysses" --use-fsdp-inference false \
|
||||
--enable-torch-compile false --batching-max-size "$batching" --batching-delay-ms 0 \
|
||||
$([[ "$prec" == fp8 ]] && echo --quantization fp8) $enc_flag \
|
||||
--warmup-resolutions 1344x768 \
|
||||
--host 0.0.0.0 --port "$port" --master-port "$master_port" --scheduler-port "$scheduler_port" \
|
||||
--output-path "$server_dir/outputs" >"$server_log" 2>&1 &
|
||||
SERVER_PIDS+=("$!")
|
||||
done
|
||||
for ((replica=0; replica<replicas; replica++)); do
|
||||
port=$((BASE_PORT + replica * PORT_STRIDE))
|
||||
server_log="$phase_dir/server_${replica}_port${port}/server.log"
|
||||
wait_healthy "$port" "${SERVER_PIDS[$replica]}" "$server_log" || {
|
||||
log "WARN: server failed startup: topo=$topo prec=$prec variant=$variant replica=$replica"
|
||||
return 1
|
||||
}
|
||||
log "variant=$variant replica=$replica healthy port=$port"
|
||||
done
|
||||
}
|
||||
|
||||
run_clients() {
|
||||
local topo=$1 prec=$2 replicas=$3 tp=$4 ulysses=$5 batching=$6 task=$7 phase_dir=$8
|
||||
local replica port client_dir in_flight failed=0
|
||||
in_flight=$batching # in-flight jobs per instance aligns with server batching ceiling
|
||||
CLIENT_PIDS=()
|
||||
for ((replica=0; replica<replicas; replica++)); do
|
||||
port=$((BASE_PORT + replica * PORT_STRIDE)); client_dir="$phase_dir/client_${replica}_port${port}"
|
||||
mkdir -p "$client_dir"
|
||||
"$PYTHON" "$CLIENT_SCRIPT" run --host "$HOST" --port "$port" --model "$MODEL" --task "$task" \
|
||||
--reference-image "$REFERENCE_IMAGE" --prompt-file "$PROMPT_FILE" --resolutions "$RESOLUTIONS" \
|
||||
--requests-per-resolution "$REQUESTS_PER_RESOLUTION" --replica-index "$replica" --num-replicas "$replicas" \
|
||||
--in-flight "$in_flight" \
|
||||
--num-inference-steps "$NUM_INFERENCE_STEPS" --warmup-requests 1 --warmup-inference-steps 5 \
|
||||
--duration-seconds "$DURATION_SECONDS" --aspect-ratio 16:9 --output "$client_dir/results.jsonl" \
|
||||
>"$client_dir/client.log" 2>&1 &
|
||||
CLIENT_PIDS+=("$!")
|
||||
log "started task=$task client=$replica port=$port requests=$((REQUESTS_PER_TASK / replicas)) in_flight=$in_flight"
|
||||
done
|
||||
for ((replica=0; replica<replicas; replica++)); do wait "${CLIENT_PIDS[$replica]}" || failed=1; done
|
||||
CLIENT_PIDS=()
|
||||
"$PYTHON" "$CLIENT_SCRIPT" summarize --input-dir "$phase_dir" --output "$phase_dir/summary.json" \
|
||||
--task "$task" --topo "$topo" --prec "$prec" --tp "$tp" --ulysses "$ulysses" \
|
||||
--replicas "$replicas" --batching "$batching" --in-flight "$in_flight" \
|
||||
--expected-requests "$REQUESTS_PER_TASK" >> "$SUMMARY_TSV" || failed=1
|
||||
return "$failed"
|
||||
}
|
||||
|
||||
if [[ "${DRY_RUN:-0}" == 1 ]]; then
|
||||
echo "===== B300 matrix plan (DRY_RUN) ====="
|
||||
for topo_spec in $TOPO_LIST; do
|
||||
IFS='|' read -r name replicas tp ulysses batching <<< "$topo_spec"
|
||||
for prec in $QUANT_LIST; do
|
||||
for task in $TASKS; do
|
||||
echo " topo=$name prec=$prec task=$task replicas=$replicas tp=$tp ulysses=$ulysses batching=$batching gpu_mode=$GPU_MODE"
|
||||
done
|
||||
done
|
||||
done
|
||||
echo "RESULT_ROOT=$RESULT_ROOT"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
for topo_spec in $TOPO_LIST; do
|
||||
IFS='|' read -r name replicas tp ulysses batching <<< "$topo_spec"
|
||||
k=$((tp * ulysses))
|
||||
topo_skip=0
|
||||
if [[ "$GPU_MODE" == share ]]; then
|
||||
((k == 1)) || { log "WARN: skip topo $name: GPU_MODE=share requires tp*ulysses==1"; topo_skip=1; }
|
||||
((replicas % TOTAL_GPUS == 0)) || { log "WARN: skip topo $name: replicas=$replicas not multiple of TOTAL_GPUS=$TOTAL_GPUS"; topo_skip=1; }
|
||||
else
|
||||
((replicas * k == TOTAL_GPUS)) || { log "WARN: skip topo $name: replicas*K=$((replicas*k)) != TOTAL_GPUS=$TOTAL_GPUS (multi-instance-per-GPU topo needs GPU_MODE=share)"; topo_skip=1; }
|
||||
fi
|
||||
((REQUESTS_PER_TASK % replicas == 0)) || { log "WARN: skip topo $name: requests/task not divisible by replicas=$replicas"; topo_skip=1; }
|
||||
((topo_skip == 1)) && continue
|
||||
for prec in $QUANT_LIST; do
|
||||
for task in $TASKS; do
|
||||
[[ "$task" == fl2va ]] && variant=FL2VA || variant=Ref2VA
|
||||
phase_dir="$RESULT_ROOT/${name}_${prec}/${task}"; mkdir -p "$phase_dir"
|
||||
log "===== topo=$name prec=$prec task=$task replicas=$replicas tp=$tp ulysses=$ulysses batching=$batching ====="
|
||||
phase_failed=0
|
||||
if start_servers "$name" "$prec" "$replicas" "$tp" "$ulysses" "$batching" "$variant" "$phase_dir"; then
|
||||
run_clients "$name" "$prec" "$replicas" "$tp" "$ulysses" "$batching" "$task" "$phase_dir" || phase_failed=1
|
||||
else
|
||||
phase_failed=1
|
||||
fi
|
||||
stop_servers
|
||||
if ((phase_failed == 1)); then
|
||||
if [[ "$SKIP_ON_FAIL" == 1 ]]; then
|
||||
log "WARN: phase failed (topo=$name prec=$prec task=$task); skipping and continuing"
|
||||
else
|
||||
die "phase failed: topo=$name prec=$prec task=$task; inspect $phase_dir"
|
||||
fi
|
||||
fi
|
||||
done
|
||||
done
|
||||
done
|
||||
|
||||
trap - EXIT INT TERM
|
||||
log "b300 matrix complete: $SUMMARY_TSV"
|
||||
@ -19,3 +19,5 @@
|
||||
- `results/balanced-tp4-tp2-20steps-5s-20260822-175030`:均衡分片后的 TP4/TP2 最终结果。
|
||||
|
||||
报告结论:整机吞吐以 TP2×4 最优;单请求时延以 TP8×1 最优。原始源路径见根目录 `SOURCE_MAP.tsv`。
|
||||
|
||||
runner 可通过 `SSIM_REFERENCE_ROOT` 对新生成视频进行成对 SSIM。评分发生在吞吐请求结束且服务停止之后,不进入吞吐计时;具体口径和输出见 `../README.md`。
|
||||
|
After Width: | Height: | Size: 2.2 MiB |
|
After Width: | Height: | Size: 404 KiB |
|
After Width: | Height: | Size: 11 KiB |
|
After Width: | Height: | Size: 11 KiB |
|
After Width: | Height: | Size: 2.1 MiB |
|
After Width: | Height: | Size: 11 KiB |
|
After Width: | Height: | Size: 9.1 KiB |
|
After Width: | Height: | Size: 12 KiB |
|
After Width: | Height: | Size: 8.3 KiB |
|
After Width: | Height: | Size: 11 KiB |
|
After Width: | Height: | Size: 787 KiB |
|
After Width: | Height: | Size: 9.4 MiB |
|
After Width: | Height: | Size: 11 KiB |
|
After Width: | Height: | Size: 8.9 MiB |
|
After Width: | Height: | Size: 1.7 MiB |
|
After Width: | Height: | Size: 2.9 MiB |
|
After Width: | Height: | Size: 9.1 KiB |
|
After Width: | Height: | Size: 8.7 MiB |
|
After Width: | Height: | Size: 1.2 MiB |
|
After Width: | Height: | Size: 113 KiB |
|
After Width: | Height: | Size: 4.3 MiB |
|
After Width: | Height: | Size: 4.0 MiB |
|
After Width: | Height: | Size: 113 KiB |
|
After Width: | Height: | Size: 1.2 MiB |
|
After Width: | Height: | Size: 2.1 MiB |
|
After Width: | Height: | Size: 1.6 MiB |
|
After Width: | Height: | Size: 2.3 MiB |