1.1 KiB
1.1 KiB
GLM-5.2 W4A8 P800 single-node benchmark
- Validated hosts:
sskj-P800-2(gpu049) andsskj-P800-4(gpu051) - Image:
iregistry.baidu-int.com/xpu/sglang-p800-pd-disagg-0510:20260715_5746 - Default model:
/data2/glm52_deploy/models/GLM-5.2-w4a8 - Parallelism: TP=8, EP=8, DP=1
- Quantization:
w4a8_int4; KV cache:int8 - Default dataset:
/data2/ShareGPT_V3_unfiltered_cleaned_split.json - Benchmark: ShareGPT, 64 prompts, concurrency 8, output length 128, request rate unlimited
Run bash start_glm52_w4a8.sh, wait for http://127.0.0.1:30000/health, then run bash run_bench.sh.
The launcher accepts environment overrides so that tuning runs change one
variable at a time. Supported knobs include CUDA_GRAPH_MAX_BS,
CHUNKED_PREFILL_SIZE, MAX_PREFILL_TOKENS, MEM_FRACTION_STATIC,
MAX_RUNNING_REQUESTS, KV_CACHE_DTYPE, and ENABLE_SPECULATIVE.
Run the fixed four-point tuning smoke suite with:
RUN_ID=<name> bash run_tuning_smoke.sh
For EAGLE/MTP, start the service with ENABLE_SPECULATIVE=1. The default
speculative settings follow the vendor deployment document: 3 steps, top-k 1,
and 4 draft tokens.