sskj/platforms
yy bf03201168 feat(p800): auto-load timing module in every TP worker process
sitecustomize.py is imported automatically by Python at startup
in every process, including the SGLang scheduler/TP worker
processes spawned for tensor parallelism. Adding the p800_timing
import here ensures the monkey-patches are applied to all 8 TP
workers, not just the launch_server parent process (which never
runs the model forward and would record empty timing data).

The timing module self-disables unless SGLANG_TIMING_ENABLED=1,
so this import is a no-op when timing is not in use.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-21 02:02:06 +00:00
..

Platform Configurations

Each .env file in this directory describes one accelerator platform. They are meant to be sourced by benchmark scripts through scripts/common/platform.sh, not executed directly.

Usage

# Default platform for the current machine
bash experiments/dsv4_p800_sglang/run_bench.sh

# Explicitly select a platform
PLATFORM=kunlun_p800 bash experiments/dsv4_p800_sglang/run_bench.sh

Current platforms

File Chip/Accelerator Engine Notes
kunlun_p800.env Kunlun P800 XPU sglang-xpu Docker-based SGLang serving image
nvidia_h200.env NVIDIA H200 vllm-dspark Native host virtual environments
nvidia_h20.env NVIDIA H20 vllm / sglang Docker-basedvllm-openai / sglang 官方镜像)
nvidia_rtx6000d.env NVIDIA RTX 6000D vllm / sglang Docker-basedSM120 部署见 envs/SM120_DSV4_DEPLOYMENT_GUIDE.md

What belongs here

  • Chip/accelerator identity (CHIP, ACCELERATOR, HARDWARE, ENGINE).
  • Device selection environment variables.
  • Platform-wide paths that rarely change (model root, default port).
  • Container image / interpreter paths for Docker-based platforms.
  • Native interpreter / venv paths for host-based platforms.

What does NOT belong here

  • Specific model names or experiment scenarios — those go in experiments/<name>/config.env.
  • Engine-specific launch flags — those go in the experiment's start_server.sh or run_bench.sh.

Adding a new platform

  1. Create platforms/<chip>.env with at least CHIP, ACCELERATOR, HARDWARE, ENGINE, DEFAULT_PORT, MODEL_ROOT.
  2. If the platform runs inside Docker, set DOCKER_IMAGE, CONTAINER_NAME, CONTAINER_PYTHON, and PATCH_ROOT (see kunlun_p800.env).
  3. If the platform runs natively on the host, set the relevant venv paths (see nvidia_h200.env).
  4. Add a row to the table above and write a quick-start experiment under experiments/<name>/.