sskj/experiments/pro6000/kimi3_pro6000_sglang_tp32ep32_moe_backend_prefill

Kimi-K3 / Pro6000D / SGLang TP32 EP32 MoE Backend Prefill

This experiment is the real-serving acceptance gate for Kimi-K3 FlashInfer MXFP4 MoE on NVIDIA SM120. It does not modify the existing baseline experiment.

Fixed matrix

Dimension Values
Nodes 601-604, 8 GPUs per node
Parallelism TP32, EP32, DP1
MoE runner marlin, flashinfer_mxfp4
Input / output 16K / 1 token
Concurrency 8, 16
Chunked prefill 8K, 16K
Requests 40 per repeat
Repeats 3

The eight configurations are identical except for MoE runner and chunk size. The client uses SGLang random-ids, which produces the exact synthetic token length without depending on a ShareGPT file or text truncation. Radix cache is disabled. OSL=1 makes TTFT, E2E latency and input TPS the primary metrics; TPOT is not meaningful for this matrix.

The defaults above can be narrowed for a capacity-first run with EP_SIZE, BACKENDS_CSV, CHUNKED_PREFILL_SIZES_CSV, and CONCURRENCIES_CSV. For example, the TP32/EP4 maximum-pressure probe uses Marlin, 16K chunk, and C=16 before spending time on the complete backend comparison.

Unique entrypoint

Run on 601 only. The sudo password is supplied at runtime and is never written to Git or result files.

cd /data/hzy/sskj/experiments/pro6000/kimi3_pro6000_sglang_tp32ep32_moe_backend_prefill

read -rsp 'sudo password: ' SUDO_PASSWORD; echo
export SUDO_PASSWORD
export RUN_ID=kimi3-moe-prefill-$(date +%Y%m%d-%H%M%S)

bash run_moe_backend_prefill.sh build-image
bash run_moe_backend_prefill.sh run

For an unattended run, use the same entrypoint under tmux:

umask 077
printf '%s\n' "$SUDO_PASSWORD" > /tmp/kimi3-phase5-sudo
unset SUDO_PASSWORD

tmux new-session -d -s kimi3-moe-prefill \
  "cd /data/hzy/sskj/experiments/pro6000/kimi3_pro6000_sglang_tp32ep32_moe_backend_prefill && \
   SUDO_PASSWORD_FILE=/tmp/kimi3-phase5-sudo RUN_ID=\"${RUN_ID}\" \
   bash run_moe_backend_prefill.sh all 2>&1 | tee /data/hzy/${RUN_ID}.log"

Delete /tmp/kimi3-phase5-sudo after the run. The file is mode 0600 and is never copied to result directories.

Outputs

results/<RUN_ID>/
  build/            per-node image build and smoke evidence
  service/          exact docker commands and all four server logs
  raw/              one benchmark JSONL per repeat
  bench/            human-readable benchmark logs
  gpu/              per-node GPU snapshots
  results.csv       flat repeat-level metrics
  summary.json      median metrics and FlashInfer-vs-Marlin deltas
  orchestrator.log  lifecycle and progress

The run aborts on a failed service start, any failed request, a missing result, or an image capability smoke failure. Containers are removed on exit.

The candidate image preserves the Kimi base image's dependency set. It installs the patched FlashInfer wheel/AOT libraries and applies a minimal diff against the image's exact two SGLang source files; it does not replace the complete SGLang Python tree or require a newer sglang-kernel package.