Add experiments/dsv4_p800_max_context_length/ to find the longest input context that P800 SGLang INT8 can serve for DeepSeek-V4-Flash-INT8. - Docker-based server start with dynamic --context-length. - Single-request, single-output-token probes for each candidate length. - Records success/failure, server args, and metrics per length. - Generates results.json and report.md summarizing the max supported length. - Smoke-tested at 4096 tokens successfully.
686 B
686 B
Benchmark Report: dsv4_p800_sglang
Metadata
- Run ID: 20260708-051316
- Timestamp: 2026-07-08T05:13:16+00:00
- Chip/Accelerator: Kunlun P800 XPU / kunlun_p800
- Engine/Backend: sglang-xpu / sglang
- Hardware: 8x Kunlun P800 XPU
- Model: /data1/models/DeepSeek-V4-Flash-INT8
- Git Commit:
2c332ff
Results
| Scenario | Conc | In/Out | Success | Failed | Req/s | OutTok/s | TTFT p50 | TTFT p99 | TPOT p50 | TPOT p99 | E2E p99 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| c32_i512_o256 | 32 | 512/256 | 10 | 0 | 1.87 | 262.65 | 506.4399079710711 | 507.2712521097855 | 21.948128755653826 | 29.154062995221466 | 5314.590020017349 |