[Docs] Explain Kimi-K3 Deep PP Prefill optimization

This commit is contained in:
Zhiyi Hong 2026-08-21 14:53:44 +08:00
parent d7381abe84
commit 96ffea6d37

View File

@ -1,118 +1,294 @@
# Kimi-K3 Prefill PP baseline search # Kimi K3 推理优化 - 优化 Prefill - Deep PP
## Goal ## 1. 工作背景
Select the pipeline-parallel configuration for the next MoE A2A experiment on Kimi-K3 在 4 台 RTX 6000D、共 32 张 GPU 上部署时,最直接的方式是
four RTX 6000D nodes. The comparison fixes every workload and backend variable `PP1 / TP32 / EP4`:完整模型只有一个 Pipeline StageTensor Parallel
except PP/TP: 覆盖全部 32 张卡MoE Expert Parallel 使用 4 张卡。
- SGLang image: `local/sglang:kimi-k3-sm120-flashinfer-mxfp4-phase5` 这套配置可以正确运行,但 16K 长输入的 Prefill 延迟很高。既有 Nsight
- Model: Kimi-K3 Systems Profile 显示,在 `16K -> 1、C=8、chunk=8K` 场景中:
- GPUs: 32 GPUs on 601-604
- MoE: FlashInfer MXFP4, EP4, A2A disabled
- Workload: 16K input, one output token, C=8/16, 40 requests
- Chunked prefill: 8K
- Repeats: three per point
`run_pp_baseline_search.sh` is the only entry point. It uses the upstream - 稳定 GPU 的 Prefill 窗口中NCCL 暴露时间占比中位数为 **57.27%**
SGLang PP implementation and does not apply a runtime PP patch. - 单个 Profile 窗口约有 **557 次 AllReduce**
- AllReduce p50 约为 **7.10 ms**
- 通信与计算重叠率接近 **0%**
- 其中 MoE 相关 AllReduce 约占已分类 AllReduce 时间的 **58.73%**
## Result 这说明 PP1/TP32 并不只是“32 张卡一起算得更快”。Kimi-K3 每层中频繁出现
的 TP Collective 会跨越 4 台服务器,在没有 NVLink 的 RTX 6000D 上通过
RoCE 执行;等待通信的时间已经超过实际计算时间。
Values are the medians of three runs. PP1 is the directly comparable existing ## 2. 为什么尝试 Deep PP
run; PP2/4/8 come from the new search.
| PP / TP / EP | C | Input TPS | TTFT p50 | TTFT p95 | Input TPS vs PP1 | TTFT p50 vs PP1 | 本实验所说的 Deep PP是把一个模型划分为更多 Pipeline Stage同时相应
缩小每个 Stage 内的 TP Group
```text
总 GPU 数 = PP × TP = 32
```
在固定 32 张 GPU 的前提下,本实验考察:
| 配置 | 每个 Stage 使用的 GPU | Stage 与节点的关系 |
|---|---:|---|
| PP1 / TP32 | 32 | 一个 Stage 跨 4 个节点 |
| PP2 / TP16 | 16 | 一个 Stage 跨 2 个节点 |
| PP4 / TP8 | 8 | 一个 Stage 完整落在 1 个节点 |
| PP8 / TP4 | 4 | 每个节点放置 2 个 Stage |
PP 加深后会新增 Pipeline Activation 的 Stage-to-Stage 传输,也会产生流水线
填充和排空造成的 bubble因此 PP 并不是越大越快。它可能带来收益的原因是:
1. 每层高频 TP AllReduce 的通信域从 32 卡逐步缩小到 16、8、4 卡;
2. PP4 开始TP Group 完整落在单机内,不再让每次 TP AllReduce 都跨节点;
3. PP8 中 TP4/EP4 都位于单机,跨节点通信主要集中在少数 Pipeline 边界;
4. 只要 C=8/16 提供了足够多的请求Scheduler 就可能用不同请求交错填充
Pipeline抵消一部分 bubble。
因此这一阶段要回答的不是“PP 能不能启动”,而是:在 Kimi-K3、6000D
和固定 Prefill 压力下,缩小 TP 通信域的收益能否超过 Pipeline 开销,以及
拐点出现在 PP2、PP4 还是 PP8。
## 3. 实现方式
### 3.1 唯一实验入口
实验只保留一个入口脚本:
```text
run_pp_baseline_search.sh
```
脚本负责:
1. 检查 601-604 的模型、镜像与 `/dev/infiniband`
2. 依次启动 PP2/TP16、PP4/TP8、PP8/TP4 服务;
3. 通过 `/get_server_info` 校验实际 PP、TP、EP 和 MoE backend
4. 每个配置执行 C=8/16各重复 3 次;
5. 校验每轮是否完成 40/40 请求以及是否存在请求错误;
6. 保存完整 Docker 命令、服务日志、benchmark 日志、原始 JSONL 和 GPU 快照;
7. 生成逐轮 `results.csv` 和中位数 `summary.csv`
8. 无论成功或失败,最终清理 4 台服务器上的实验容器。
### 3.2 SGLang PP 路径
本轮使用镜像:
```text
local/sglang:kimi-k3-sm120-flashinfer-mxfp4-phase5
```
服务直接通过 SGLang 官方参数启动:
```bash
python3 -m sglang.launch_server \
--model-path /data/hf_models/Kimi-K3 \
--tp-size "${TP_SIZE}" \
--pp-size "${PP_SIZE}" \
--ep-size 4 \
--nnodes 4 \
--node-rank "${NODE_RANK}" \
--dist-init-addr 174.1.60.1:20000 \
--moe-runner-backend flashinfer_mxfp4 \
--chunked-prefill-size 8192
```
实验没有复制同事旧代码,也没有注入 `patch_pp_group.py`。准备阶段曾验证过
一版运行时 PP Group patch但它会让 Python Object 通信和 Tensor/P2P 通信
复用错误的 Gloo Group出现消息大小不匹配。回到当前 SGLang 上游实现后,
PP2、PP4 和 PP8 均能正常启动,因此该 patch 被完全移除,没有进入正式数据。
### 3.3 网络配置
4 个节点都向容器映射 `/dev/infiniband`,并使用:
```bash
NCCL_SOCKET_IFNAME=bond0
GLOO_SOCKET_IFNAME=bond0
NCCL_IB_HCA=mlx5_0,mlx5_1,mlx5_2,mlx5_3
NCCL_IB_GID_INDEX=3
```
`bond0` 承担 Bootstrap/控制连接,`mlx5_0~3` 是 NCCL 的 RoCE HCA。正式运行
中没有出现 NCCL、Gloo、RDMA、connection refused 或 communicator failure。
## 4. 实验设计
为了让结果只反映 PP/TP 变化,其他主要变量保持固定:
| 项目 | 固定值 |
|---|---|
| 模型 | Kimi-K3 |
| GPU | 601-60432 × RTX 6000D |
| MoE runner | FlashInfer MXFP4 |
| A2A backend | none |
| EP | 4 |
| 输入/输出 | 16,384 / 1 token |
| Chunked Prefill | 8,192 token |
| 并发 | C=8、C=16 |
| 每轮请求数 | 40 |
| 重复次数 | 3 |
| 数据集 | random-ids输入长度比例固定为 1.0 |
| Prefix Cache | `--disable-radix-cache` |
| 显存比例 | 0.88 |
输出长度固定为 1是为了让 TTFT 和 Input TPS 主要反映 Prefill而不是被长
Decode 混入。每个点重复 3 次,最终使用中位数,避免首轮 JIT、缓存状态或系统
瞬时抖动主导结果。
PP1 没有重复跑。它来自前一阶段已经完成的 FlashInfer MXFP4、EP4、chunk8K
实验。经逐项核对PP1 与本轮使用相同镜像、模型、MoE backend、EP、输入输出、
并发、请求数、显存比例和网络参数,区别只有 `PP1/TP32`,因此可以直接并入。
## 5. 一个关键实现差异PP Micro Batch
同事旧 PP4 配置显式设置过:
```bash
--pp-max-micro-batch-size 1
```
当前 SGLang 在该参数未设置时,会在 Scheduler 初始化后自动计算:
```python
pp_max_micro_batch_size = max(max_running_requests // pp_size, 1)
```
接纳新请求时又会执行:
```python
allocatable_requests = pp_max_micro_batch_size - running_batch_size
```
因此这里的 `1` 不只是一个普通的 Kernel batch 参数,它会把 PP Scheduler
允许同时运行的请求数压到 1。对于 C=8/16这会让请求大量排队Pipeline
难以通过请求交错保持各 Stage 忙碌bubble、P2P 和调度固定成本都会被放大。
本实验不传这个参数,使用 SGLang 当前的自动值。旧实验还混合了 Marlin、
chunk16K、EP1/EP8、PP16、旧镜像和运行时 PP patch因此旧数据不能作为
“PP 天生更慢”的单变量证据。
## 6. 实验过程与正确性
正式 Run ID
```text
kimi3-pp-upstream-20260821-120208
```
服务启动时间分别为:
| 配置 | 健康检查通过时间 |
|---|---:|
| PP2 / TP16 / EP4 | 417 s |
| PP4 / TP8 / EP4 | 405 s |
| PP8 / TP4 / EP4 | 338 s |
PP2、PP4、PP8 共形成 18 个正式样本。每个样本均完成 40/40 请求,原始结果
中的 errors 为空。服务和 benchmark 日志没有发现:
- CUDA OOM / OutOfMemory
- Traceback / Exception
- NCCL / Gloo / RDMA failure
- Engine failure
- 请求连接失败或提前终止。
PP8 在 C=8 和 C=16 下三次 Input TPS 的极差都小于 0.7%,说明结果不是某
一次运行的偶然峰值。
## 7. 性能结果
下表均为 3 次正式运行的中位数:
| PP / TP / EP | C | Input TPS | TTFT p50 | TTFT p95 | Input TPS 相对 PP1 | TTFT p50 相对 PP1 |
|---|---:|---:|---:|---:|---:|---:| |---|---:|---:|---:|---:|---:|---:|
| 1 / 32 / 4 | 8 | 3,257.96 | 39.19 s | 41.62 s | baseline | baseline | | 1 / 32 / 4 | 8 | 3,257.96 | 39.19 s | 41.62 s | 基线 | 基线 |
| 2 / 16 / 4 | 8 | 5,015.95 | 25.22 s | 26.69 s | +53.96% | -35.65% | | 2 / 16 / 4 | 8 | 5,015.95 | 25.22 s | 26.69 s | +53.96% | -35.65% |
| 4 / 8 / 4 | 8 | 7,155.20 | 17.28 s | 18.62 s | +119.62% | -55.90% | | 4 / 8 / 4 | 8 | 7,155.20 | 17.28 s | 18.62 s | +119.62% | -55.90% |
| 8 / 4 / 4 | 8 | **8,809.50** | **13.47 s** | **16.45 s** | **+170.40%** | **-65.64%** | | 8 / 4 / 4 | 8 | **8,809.50** | **13.47 s** | **16.45 s** | **+170.40%** | **-65.64%** |
| 1 / 32 / 4 | 16 | 3,260.14 | 78.39 s | 80.88 s | baseline | baseline | | 1 / 32 / 4 | 16 | 3,260.14 | 78.39 s | 80.88 s | 基线 | 基线 |
| 2 / 16 / 4 | 16 | 5,018.38 | 50.70 s | 52.14 s | +53.93% | -35.32% | | 2 / 16 / 4 | 16 | 5,018.38 | 50.70 s | 52.14 s | +53.93% | -35.32% |
| 4 / 8 / 4 | 16 | 7,152.58 | 34.84 s | 35.95 s | +119.39% | -55.56% | | 4 / 8 / 4 | 16 | 7,152.58 | 34.84 s | 35.95 s | +119.39% | -55.56% |
| 8 / 4 / 4 | 16 | **8,824.12** | **27.31 s** | **29.77 s** | **+170.67%** | **-65.16%** | | 8 / 4 / 4 | 16 | **8,824.12** | **27.31 s** | **29.77 s** | **+170.67%** | **-65.16%** |
All 18 new samples completed 40/40 requests. No OOM, traceback, NCCL failure, 从 PP1 到 PP8Input TPS 沿着 PP2、PP4、PP8 单调上升TTFT 也同步下降。
request error, or engine failure was found. The three repeats are stable: PP8 这说明在当前 C=8/16 压力下,流水线仍有足够工作可以填充;在测试范围内,
Input TPS varies by less than 0.7% at both concurrency points. 缩小 TP 通信域的收益尚未被更深 PP 的 bubble 抵消。
## Memory C=16 的 TTFT 大致是 C=8 的两倍,而同一 PP 下 Input TPS 基本不变。这表明
这批 Prefill 工作已接近该配置的稳定吞吐区:增加并发主要增加排队等待,而不是
继续提高输入吞吐。PP 改善的是服务处理这批 token 的速率,因此两个并发点都
出现了接近一致的 TPS 增益。
The table reports the highest per-GPU memory observed in the post-benchmark ## 8. 显存结果
snapshot across all four nodes.
| PP / TP / EP | Peak GPU memory | Remaining from 85,651 MiB | 下表取每个配置 benchmark 完成后4 个节点 32 张卡中的最高显存:
| PP / TP / EP | 峰值显存 | 相对 85,651 MiB 的剩余空间 |
|---|---:|---:| |---|---:|---:|
| 1 / 32 / 4 | 81,355 MiB | 4,296 MiB | | 1 / 32 / 4 | 81,355 MiB | 4,296 MiB |
| 2 / 16 / 4 | 80,335 MiB | 5,316 MiB | | 2 / 16 / 4 | 80,335 MiB | 5,316 MiB |
| 4 / 8 / 4 | 82,417 MiB | 3,234 MiB | | 4 / 8 / 4 | 82,417 MiB | 3,234 MiB |
| 8 / 4 / 4 | 79,407 MiB | 6,244 MiB | | 8 / 4 / 4 | 79,407 MiB | 6,244 MiB |
PP stages do not have identical model/state allocations, so memory is less PP 的 Stage 不会获得完全相同的层、输出头和状态,因此 PP2/4/8 的显存不再
balanced than PP1. PP8 nevertheless has the best worst-rank headroom in this 像 PP1 那样均匀。PP4 的最重 Stage 已使用约 82.4 GiBPP8 的最重 Stage
run, which matters because A2A backends allocate additional communication 约 79.4 GiB在本轮反而保留了最多的最坏 Rank 余量。
buffers.
## Why PP helps here 这一点对下一阶段 A2A 很重要,因为 DeepEP 或 FlashInfer A2A 都可能申请额外
dispatch/combine buffer。只看 TPS 而忽略最重 Rank可能出现基线可跑、打开
A2A 后启动或 Prefill OOM 的情况。
The existing PP1 Nsight profile measured a median 57.27% of the Prefill window ## 9. 通信结果如何解释
inside exposed NCCL work, with about 557 AllReduce calls and effectively no
compute/communication overlap. Its TP32 collectives span all four nodes.
Changing PP also changes the TP communication domain: 本轮 PP 搜索没有为 PP2/4/8 再录制 Nsight也没有在每个 benchmark 前后采集
独立 HCA Counter因此不能把 PP1 的 57.27% NCCL 占比直接写成 PP8 的实测
通信占比。
- PP2/TP16: each TP group spans two nodes. 能够被当前证据直接证明的是:
- PP4/TP8: each TP group fits on one node; only pipeline activation transfers
cross node boundaries.
- PP8/TP4: two pipeline stages fit per node and each TP/EP4 group is node-local.
This replaces frequent cross-node TP AllReduce with a smaller number of 1. PP1 的同口径 Nsight 确认 TP32 Prefill 以暴露的 AllReduce 为主;
pipeline activation transfers. At C=8/16 there is enough request-level work to 2. PP2/4/8 的 `/get_server_info` 证明实际 PP/TP/EP 与计划一致;
fill the deeper pipeline, so the communication saving is larger than the 3. Rank 布局使 PP4 的 TP8、PP8 的 TP4/EP4 都限制在单节点内;
pipeline bubble cost. 4. PP2/4/8 的 NCCL P2P communicator 成功建立并完成全部请求;
5. 随 TP 通信域缩小,端到端 Input TPS 按 PP1 < PP2 < PP4 < PP8 单调提升
与“跨节点 TP AllReduce 是主要瓶颈”的 Profile 结论一致。
The current run verifies the rank topology, successful NCCL P2P communicators, 因此当前结果支持通信域重构这一解释但没有伪造一个未经采样的“PP8 NCCL
and absence of communication failures. It does not contain per-PP Nsight or HCA 占比”。如需进一步拆解 PP8 内部剩余时间,应单独执行短窗口 Nsight而不是
bandwidth samples; the 57.27% NCCL attribution belongs specifically to the 把 Profile 开销混进这组性能基准。
directly comparable PP1 profile and is used to explain, not fabricate, the PP8
timeline.
## Why older PP tests could be slower ## 10. 最终判断与后续使用
The old configurations were not an apples-to-apples PP-only comparison. In 在本次固定场景中,`PP8 / TP4 / EP4` 同时得到:
particular, an old PP4 profile set:
```bash - 最高的 Input TPS
--pp-max-micro-batch-size 1 - 最低的 TTFT p50/p95
- 三次重复中稳定的结果;
- 比 PP1 和 PP4 更充足的最坏 Rank 显存余量;
- 单节点内的 TP4/EP4 通信域。
因此,下一阶段 MoE A2A 的首个兼容性和显存 Smoke 应固定使用:
```text
PP8 / TP4 / EP4
16K -> 1
C=8/16
chunked_prefill_size=8192
FlashInfer MXFP4
``` ```
Current SGLang leaves this value unset and computes: `PP4 / TP8 / EP4` 保留为回退对照:如果某个 A2A backend 对 PP8/TP4 有当前
代码未覆盖的约束,可以快速判断问题来自 backend 兼容性,还是 A2A 本身。
```python 这项结果不能外推到低并发 C=1、长 Decode、在线混合流量或更多节点。当前
pp_max_micro_batch_size = max(max_running_requests // pp_size, 1) TTFT 即使在 PP8 下仍为 13.47 sC8和 27.31 sC16距离 5 s 目标仍有
``` 明显差距。Deep PP 解决了一个主要通信瓶颈,但不是 Prefill 优化的终点。
The scheduler then limits admission with ## 11. 复现命令与证据
`pp_max_micro_batch_size - running_bs`. Setting it to one therefore permits
only one running request per PP scheduler, serializing much of a C=8/16 load
and exposing pipeline bubbles. Older tests also mixed Marlin, 16K chunks,
EP1/EP8, PP16, older images, and a runtime PP-group patch. They cannot be used
to conclude that PP itself is slower.
## Decision 在 601 执行:
Use **PP8 / TP4 / EP4** as the primary configuration for the next A2A
compatibility and memory smoke test. Keep **PP4 / TP8 / EP4** as the fallback:
it is slower and has less memory headroom here, but provides a useful check if
an A2A backend imposes TP/PP constraints not exercised by the baseline.
Do not carry forward `--pp-max-micro-batch-size 1`. Leave it unset unless a
separate controlled experiment establishes a reason to cap it.
## Reproduce and evidence
Run from 601:
```bash ```bash
cd /data/hzy/sskj/experiments/pro6000/kimi3_pro6000_sglang_pp_baseline_search cd /data/hzy/sskj/experiments/pro6000/kimi3_pro6000_sglang_pp_baseline_search
@ -121,29 +297,38 @@ RUN_ID=kimi3-pp-$(date +%Y%m%d-%H%M%S) \
bash run_pp_baseline_search.sh run bash run_pp_baseline_search.sh run
``` ```
New PP2/4/8 evidence: PP2/4/8 正式结果:
```text ```text
/data/hzy/sskj/experiments/pro6000/
kimi3_pro6000_sglang_pp_baseline_search/
results/kimi3-pp-upstream-20260821-120208/ results/kimi3-pp-upstream-20260821-120208/
summary.csv
results.csv
metadata/
bench/
raw/
gpu/
service/
``` ```
Existing PP1 evidence: 其中:
```text ```text
../kimi3_pro6000_sglang_tp32ep32_moe_backend_prefill/ summary.csv 6 个 PP×C 中位数
results.csv 18 个逐轮结果
metadata/ 实际服务配置与 Run manifest
service/ 四节点 Docker 命令和完整服务日志
bench/ benchmark 文本输出
raw/ benchmark JSONL
gpu/ 四节点启动前、健康、完成和清理后的显存快照
```
PP1 原始结果:
```text
/data/hzy/sskj/experiments/pro6000/
kimi3_pro6000_sglang_tp32ep32_moe_backend_prefill/
results/kimi3-ep4-moe-full-20260818-151349/ results/kimi3-ep4-moe-full-20260818-151349/
``` ```
Existing PP1 communication profile: PP1 Nsight 通信归因:
```text ```text
../kimi3_pro6000_sglang_prefill_communication_profile/ /data/hzy/sskj/experiments/pro6000/
kimi3_pro6000_sglang_prefill_communication_profile/
results/kimi3-prefill-comm-20260820-143749/nsys_analysis.json results/kimi3-prefill-comm-20260820-143749/nsys_analysis.json
``` ```