From 595bdde5d7810189c8f0addbae83b9fb2cb4d2d5 Mon Sep 17 00:00:00 2001 From: Zhiyi Hong <2497491955@qq.com> Date: Thu, 30 Jul 2026 16:41:11 +0800 Subject: [PATCH] [Docs] audit DSV4-Pro TP16 TTFT benchmark semantics --- README.md | 4 + ...v4pro_pro6000d_2node_sglang_quick_map.html | 721 ++++++++++ ...e_sglang_prefill_hardware_attribution.html | 327 +++++ .../report.md | 20 + .../run_manifest.json | 40 + .../summary.jsonl | 4 + .../script-audit-20260730/cold_1k_o128.jsonl | 1 + .../cold_1k_o1_first.jsonl | 1 + .../cold_1k_o1_repeat.jsonl | 1 + .../old_semantics_1k_o128.jsonl | 1 + .../results/script-audit-20260730/report.md | 54 + .../推理优化计划.html | 1238 +++++++++++++++++ 12 files changed, 2412 insertions(+) create mode 100644 docs/dsv4pro_pro6000d_2node_sglang/phase1_dsv4pro_pro6000d_2node_sglang_quick_map.html create mode 100644 docs/dsv4pro_pro6000d_2node_sglang/phase2_dsv4pro_pro6000d_2node_sglang_prefill_hardware_attribution.html create mode 100644 docs/dsv4pro_pro6000d_2node_sglang/results/phase1-dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/report.md create mode 100644 docs/dsv4pro_pro6000d_2node_sglang/results/phase1-dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/run_manifest.json create mode 100644 docs/dsv4pro_pro6000d_2node_sglang/results/phase1-dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/summary.jsonl create mode 100644 docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/cold_1k_o128.jsonl create mode 100644 docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/cold_1k_o1_first.jsonl create mode 100644 docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/cold_1k_o1_repeat.jsonl create mode 100644 docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/old_semantics_1k_o128.jsonl create mode 100644 docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/report.md create mode 100644 docs/dsv4pro_pro6000d_2node_sglang/推理优化计划.html diff --git a/README.md b/README.md index 670e956..cccf281 100644 --- a/README.md +++ b/README.md @@ -1,5 +1,9 @@ # sskj — 多平台大模型推理性能基准测试项目 +> **更新(2026-07-30 16:38:41 CST)** +> +> 完成 DeepSeek-V4-Pro 双机 TP16 新旧脚本 TTFT 口径审计。确认旧产物受到 16 条 Warm-up、跨 Case 固定 Seed 递增长度、未清 Prefix Cache 及前一轮残留服务状态影响;同配置冷请求稳定复现约 16 秒/1K。新增 Phase 1 结果、脚本审计和 Phase 2 设计 HTML 档案,后续 Cold/Warm Prefix 指标分开报告。 +> > **更新(2026-07-30 14:33:52 CST)** > > 新增独立的 `dsv4pro_pro6000d_2node_sglang_tp16_quick_map` 快速性能地图与混合干扰 A/B。实验只保留一个 Shell 入口;旧 TP16 全量脚本保持不变。首轮真机验证已确认双机 TP16 服务可用,并据实测耗时将快速矩阵缩为一波请求,同时修正 Warm-up 污染 Prefix Cache 和混合负载注入时序。 diff --git a/docs/dsv4pro_pro6000d_2node_sglang/phase1_dsv4pro_pro6000d_2node_sglang_quick_map.html b/docs/dsv4pro_pro6000d_2node_sglang/phase1_dsv4pro_pro6000d_2node_sglang_quick_map.html new file mode 100644 index 0000000..668a11b --- /dev/null +++ b/docs/dsv4pro_pro6000d_2node_sglang/phase1_dsv4pro_pro6000d_2node_sglang_quick_map.html @@ -0,0 +1,721 @@ + + + + + + + Phase 1:DeepSeek-V4-Pro 双机 Pro6000D SGLang 快速性能地图 + + + +
+
+

Implementation & Result Record

+

Phase 1:DeepSeek-V4-Pro 双机 Pro6000D SGLang 快速性能地图

+
+ 节点:174.1.51.5 + 174.1.51.7 + 拓扑:SGLang TP16 / EP2 + 更新:2026-07-30 15:46 CST +
+
+
+ +
+ 返回推理优化主计划 + +

+ 阶段状态:提前结束,已进入硬件归因。 + 第一次真机 Run + dsv4pro-pro6000d-2node-sglang-quick-20260730-140026 + 已验证双机服务可用,但因发现请求量和 Prefix Cache 口径问题而主动停止。 + 精简后的第二次 Run + dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625 + 完成 3 个 Prefill 固定点后,发现输入吞吐稳定锁定在约 + 65 token/s。继续扫描 Decode 和混合流量不能解释该异常, + 因此用户决定中止第 4 个 Case,直接进入 Phase 2。 + Manifest 终态为 ABORTED_EARLY_FOR_PHASE2; + 两节点容器和 16 张 GPU 已清理。 +

+ +

1. 目标与边界

+

+ 用数小时以内、可重复的小矩阵替代约一天以上的全量扫描,先回答 + Prefill、Decode、长上下文和混合干扰各自是否存在明显异常,再决定后续 + Timeline 和 Kernel Profiling 的捕获对象。该阶段不要求为了“跑满表格” + 而浪费算力;一旦出现稳定、可复现且足以改变调查方向的异常,就可以提前结束。 +

+ + +

2. 精简实现

+

实验代码位于:

+
/data/hzy/sskj/experiments/pro6000/
+dsv4pro_pro6000d_2node_sglang_tp16_quick_map/
+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
文件职责
run_quick_map.sh唯一 Shell 入口:双机服务启停、固定矩阵、混合 A/B、错误处理与清理
config.env节点、模型、镜像、并行与容量参数
quick_map_scenarios.tsv九个固定工作负载点
quick_map_results.py验证 Bench JSON,生成 CSV、JSONL 和 Markdown 汇总
tests/test_quick_map_results.py结果解析回归测试
+ +

单入口的操作面:

+
bash run_quick_map.sh all
+
+# 仅排障时使用同一个入口
+bash run_quick_map.sh start
+bash run_quick_map.sh fixed
+bash run_quick_map.sh mixed
+bash run_quick_map.sh stop
+ +

3. 服务配置

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
配置项当前值说明
镜像lmsysorg/sglang:nightly-dev-cu13-20260720-b3570a45沿用已验证可加载 DSV4-Pro 的版本
模型/data/hf_models/DeepSeek-V4-Pro两台节点均有本地权重
并行TP=16, EP=2, nnodes=2每台 8 卡,共 16 Rank
显存比例0.9保持已知基线,不在本阶段调参
活跃请求上限256覆盖本轮最大并发 64
CUDA Graph Decode BS64覆盖固定矩阵中的 Decode C64
NCCL bootstrapeth1普通 TCP 建连接口
RoCE HCAmlx5_0,mlx5_3双 Rail 数据面,NCCL_CROSS_NIC=1
代码分支hzy从该维护分支向中央仓库 main 提交合并请求
+ +

4. 固定快速矩阵

+ + + + + + + + + + + + + + + + + + + + + +
Case IDISLOSLC目的
short_prefill_latency_1k_c11K11最小 TTFT
mid_prefill_latency_32k_c132K11中长 Prefill
long_prefill_latency_128k_c1128K11长上下文 Prefill
mid_prefill_throughput_32k_c1632K116Prefill 输入吞吐
decode_latency_1k_to_1k_c11K1K1单请求 TPOT
decode_throughput_1k_to_1k_c161K1K16Decode 吞吐
decode_throughput_1k_to_1k_c321K1K32Decode 吞吐
decode_throughput_1k_to_1k_c641K1K64Decode 高并发
balanced_32k_to_1k_c832K1K8综合压力
+

+ 快速 Run 使用一次重复和一波测量请求,即 num_prompts=C。 + 32K/128K Prefill 不做昂贵的同形状 Warm-up;短 Prefill 与 Decode 使用一个 + Warm-up,并在正式计时前清空 Prefix Cache。固定矩阵不做 SLO 截断或自适应并发搜索。 +

+ +

5. SGLang Benchmark 与 Prefix Cache

+

5.1 random 如何生成 ISL

+

+ 当前镜像的实现位于 + /sgl-workspace/sglang/python/sglang/benchmark/datasets/random.py。 + dataset-name=random 会读取 ShareGPT,打乱样本后取每条会话的首轮用户文本: + 文本过长就截断,过短就重复其 token,直到达到目标 ISL。 + random-range-ratio=1.0 使每条请求都使用精确的目标长度。 + 本机数据集共有 94,145 行,其中 92,886 行可用、71,904 个不同首轮文本, + 因此不存在此前“两条数据只能形成两个并发请求”的问题。 +

+

+ random-ids 则直接构造随机整数 token id,不读取 ShareGPT。 + 当前源码同时警告这种方式可能触发 NaN,因此本阶段继续使用 + random + 大规模 ShareGPT,并通过清缓存隔离不同测试点。 +

+ +

5.2 OSL 为什么能达到指定长度

+

+ SGLang 原生请求函数位于 + /sgl-workspace/sglang/python/sglang/benchmark/serving.py。 + 它将目标 OSL 写入 max_new_tokens,并默认设置 + ignore_eos=True。因此模型即使提前生成 EOS,也会继续生成到指定 OSL; + 只有请求失败、超时或触及上下文限制时,实际输出才可能不足。 +

+
sampling_params = {
+    "max_new_tokens": request_func_input.output_len,
+    "ignore_eos": not args.disable_ignore_eos,
+}
+ +

5.3 为什么 Warm-up 会污染 Prefix Cache

+

+ SGLang benchmark 的 Warm-up 直接复用 input_requests[0], + 而正式测量随后仍会遍历包含该请求的完整列表。因此,只要服务启用了 Prefix Cache, + 第一条正式请求就可能命中刚刚 Warm-up 的前缀。第一次 Run 的服务日志实际出现 + #cached-token: 768,证明该污染在当前环境真实发生。 +

+

+ 修复方式是在每个隔离测试点传入 --flush-cache。benchmark 会先完成 + Warm-up,再调用服务端 /flush_cache,最后才启动计时。这样保留 Kernel + 和执行路径预热,同时不把 Warm-up 的 KV 前缀带入测量。混合干扰中的长 Prefill + 注入不会清缓存,避免在 Decode 背景运行时改变其服务状态;背景与注入使用不同随机种子。 +

+ +

5.4 如何单独测试 Prefix Caching

+
    +
  1. 调用 /flush_cache,发送固定长 Prompt P,记录 Cold TTFT 和 #cached-token
  2. +
  3. 不清缓存,原样重发 P,记录 Warm TTFT;预期 cached token 明显增加、TTFT 降低。
  4. +
  5. 再次清缓存,发送同长度但内容不同的 Prompt Q,排除长度、JIT 和偶然波动造成的假提升。
  6. +
+

+ 三组请求保持 OSL、采样参数和并发一致,各重复至少 3 次。Prefix Cache 是生产优化能力, + 不是“坏东西”;这里只是在无缓存性能基线中隔离它,后续会把缓存命中场景作为单独 A/B。 +

+ +

6. 混合干扰实现

+

+ 这里的“背景”不是 SGLang 后台线程,而是先启动并持续运行的一批 + Decode 基准流量。它既在实验期间占用 GPU,也是我们希望观察是否 + 变慢的对象。混合 A/B 的问题非常具体:同样一批 Decode 请求,在没有长 + Prefill 干扰和有长 Prefill 干扰时,性能会相差多少? +

+ + + + + + + + +
组别运行内容作用
A:Control仅运行 64 条 1K → 1K, C=32 Decode建立无干扰基线
B:Treatment运行相同 Decode,并在正式测量开始 10 秒后注入一条 128K → 1 Prefill测量 Prefill 对 Decode 的干扰
+
    +
  1. 先完成 A 组,仅运行 Decode,保存对照指标。
  2. +
  3. 启动 B 组的 Decode 基准流量,并从日志确认它已进入正式测量,而不只是完成客户端初始化。
  4. +
  5. 正式测量开始 10 秒后,并行提交一个 128K → 1 长 Prefill。
  6. +
  7. 等待两类请求都结束,分别保存 Decode 流量和长 Prefill 请求的结果。
  8. +
  9. 用 A、B 两组 Decode 的 Output TPS、TTFT P95、TPOT P95 与 E2E P95 计算变化率;长 Prefill 自身的 TTFT 单独报告。
  10. +
+
A:Decode ───────────────────────────────→ 结束
+
+B:Decode ───────────────────────────────→ 结束
+             正式测量 + 10 秒
+                        └─ 128K Prefill ─→ 结束
+                           共同占用同一服务
+
(
+  run_bench_case ... 1024 1024 32 64
+) &
+background_pid=$!
+
+# 实际代码先从 bench.log 确认正式测量已经开始。
+sleep 10
+run_bench_case ... 131072 1 1 1
+wait "${background_pid}"
+

+ & 让 Decode benchmark 与后续 Prefill 并行; + $! 取得该 Decode benchmark 的进程号; + wait 等待它完成。总请求数 64、并发 32,表示最多同时有 + 32 条请求在途,通常形成约两波请求。如果 Decode 流量在注入前已经结束, + 两类请求没有发生重叠,结果会被明确改写为 + BACKGROUND_FINISHED_BEFORE_INJECTION,避免生成虚假的“混合成功”。 +

+ +

7. 结果与可追溯性

+
results/<RUN_ID>/
+  run_manifest.json
+  run.log
+  summary.csv
+  summary.jsonl
+  aggregate.csv
+  report.md
+  cases/<case_id>/rep1/
+    bench_cmd.txt
+    bench.jsonl
+    bench.log
+    meta.json
+  server/
+    head_server_cmd.txt
+    worker_server_cmd.txt
+    head_server.log
+    worker_server.log
+

+ 汇总保留 Request/Input/Output/Total TPS,以及 E2E、TTFT、TPOT、ITL 的 + Mean、P50、P95、P99。断点续跑前会重新解析原始 Bench JSON,不能只凭文件存在就跳过。 +

+ +

8. 已完成验证

+ + + + + + + + + + + + + + +
检查结果证据
Shell 语法通过bash -n run_quick_map.sh
Python 单测3/3 通过场景唯一性、百分位回退、失败结果汇总
完整 Dry-run通过服务、九个固定点、混合 A/B、清理均展开成功
真实旧 Bench JSON 解析通过成功解析 P50/P95/P99 与吞吐字段
项目精简通过实验目录顶层仅保留一个 Shell 入口
双机容器启动通过第二次 Run 于 14:42:04 通过 Health Check,启动约 5 分 30 秒
Prefix Cache 隔离通过Warm-up 后 POST /flush_cache 返回 200,正式请求仍为 #cached-token: 0
中止清理通过头、Worker 节点均无相关容器和 Bench 进程,16 张 GPU 显存回到 0 MiB
+ +

9. 真机结果

+

+ 最终采用 Run + dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625。 + 以下三条均为一条请求、C=1OSL=1,且正式测量前 + Prefix Cache 已清空。 +

+ + + + + + + + + + + + + + + + + +
CaseISL输入 TPSTTFTE2E状态
short_prefill_latency_1k_c11K64.44 tok/s15.88 s15.88 sCOMPLETED
mid_prefill_latency_32k_c132K64.96 tok/s504.44 s504.44 sCOMPLETED
long_prefill_latency_128k_c1128K65.20 tok/s2010.38 s2010.38 sCOMPLETED
mid_prefill_throughput_32k_c1632K × 16未形成最终结果未形成最终结果未形成最终结果ABORTED
+ +

9.1 可以下的结论

+ + +

9.2 现在还不能下的结论

+ + +

9.3 为什么提前结束

+

+ Phase 1 的目标是发现值得归因的关键异常,而不是机械完成九个格子。 + 三个独立长度已经给出同一个稳定信号;第 4 个并发 Prefill 在 922 秒后仍表现为 + 单序列 Chunk 推进。继续执行剩余矩阵预计还需数小时,却不能回答 + “这 65 token/s 到底卡在哪里”。因此第 4 个 Case 被写入 + EARLY_STOP_FOR_PHASE2,其余固定点和混合 A/B 保留为未执行。 +

+ +

9.4 为什么旧脚本的 TTFT 短很多

+

+ 2026-07-30 对旧目录 + /data/qqt/sskj/experiments/pro6000/dsv4_pro6000_sglang_tp16 + 做了逐项审计。结论是:旧结果与本轮冷 Prefill 不是同一缓存口径, + 不是 quick-map 把相同请求跑慢了。 +

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
审计项旧脚本quick-map判断
服务端配置同一镜像,TP16 / EP2,8K Chunk,FlashInfer MXFP4 MoE相同排除明显的启动参数回归
正式请求数C=1 仍强制至少 10 条延迟点只发 1 条旧均值混合了多条请求的缓存状态
Warm-up每个 Shape 固定 16 条同 Prompt Warm-upWarm-up 后清 Prefix Cache;32K/128K 延迟点不做额外 Warm-up旧正式测量会继承 Warm-up 的 Prompt 前缀
Cache 清理从不传 --flush-cache正式测量前传 --flush-cache旧脚本跨 Case、跨长度保留 Radix/Prefix Cache
Shape 顺序固定 Seed=42,按 1K→4K→8K→16K→32K→64K→128K 递增每个延迟点按冷缓存解释旧请求会复用上一档相同 Prompt 的短前缀
+

+ 旧时间线还有一条直接证据:17:40 的失败 Run 已完整执行过 + 1K / 128 / C=1,18:01 的正式 Run 没有重启服务便再次执行同一批 + Seed=42 请求。因此旧文件中的 1K TTFT 约 0.455 秒,本身就是热缓存结果。 +

+ + + + + + + + + + + +
最小复现Mean TTFTP95 TTFT解释
1K→1,首次冷缓存16.04 s16.04 s清 Prefix Cache,1 条正式请求
1K→1,原样再次冷缓存15.90 s15.90 s再次清 Cache,排除一次性 JIT 主导
1K→128,冷缓存15.79 s15.79 s排除 OSL=1 特殊慢路径
旧命令语义重新复现14.60 s15.97 s10 条正式请求、16 条 Warm-up、不清 Cache
2026-07-28 旧产物0.455 s0.513 s服务已被前一次 Run 和后续递增长度预热
+

+ 旧 32K 文件的第一条请求约 0.67 秒,其余 9 条平均约 35.51 秒; + 旧 128K 文件的第一条约 0.96 秒,其余 9 条平均约 144.83 秒。 + 后 9 条也已经分别继承上一档 16K、64K 前缀。旧报告仍按完整 ISL 统计 + Input TPS,因此会把只计算新增后缀的耗时除进完整 token 数,进一步放大吞吐。 +

+

+ 审计结论:Phase 1 的约 65 token/s 是冷 Prefix Cache 的完整 Prompt 路径, + 旧结果是热缓存/递增前缀路径。两者都可以测,但必须分成 Cold 与 Warm 两套实验, + 不能放在同一列直接比较。审计原始产物保存在: +

+
/data/hzy/dsv4_script_audit_20260730/
+

+ 本地归档: + TTFT 脚本口径审计报告 + 及同目录原始 JSON/log。 +

+ + + + + + + + + + + + + + + + + +
检查点状态结果或结论
服务健康通过端口 30002 已就绪,无 OOM、NCCL 或 Engine 异常
第一次固定点 Run主动停止发现 32K 单请求约需十余分钟;原协议的 4 次同形状请求会使整轮再次接近半天
精简后九个固定点3 完成 / 1 中止 / 5 未执行Prefill 异常信号已足够清晰,停止继续消耗算力
混合干扰 A/B未执行待 Prefill 根因明确后再决定是否重放
阶段耗时约 65 分钟14:36:26 启动,15:41:52 完成进程与容器清理
是否进入下一阶段用户决定立即进入 Phase 2 硬件指标归因
+ +

10. 结果位置

+

服务器原始结果:

+
/data/hzy/sskj/experiments/pro6000/
+dsv4pro_pro6000d_2node_sglang_tp16_quick_map/results/
+dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/
+

+ 本地已归档 + report.md、 + aggregate.csv + 和同目录下的 Manifest、Summary、主日志。 +

+ +

11. 运行命令

+
cd /data/hzy/sskj/experiments/pro6000/dsv4pro_pro6000d_2node_sglang_tp16_quick_map
+
+tmux new-session -d -s dsv4pro-pro6000d-2node-sglang-quick-map -c "$PWD"
+tmux send-keys -t dsv4pro-pro6000d-2node-sglang-quick-map \
+  'RUN_ID=dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625 bash run_quick_map.sh all' Enter
+
+tmux attach -t dsv4pro-pro6000d-2node-sglang-quick-map
+ +

+ 下一阶段: + + Phase 2:Prefill 硬件指标归因 + +

+

返回推理优化主计划

+
+ + diff --git a/docs/dsv4pro_pro6000d_2node_sglang/phase2_dsv4pro_pro6000d_2node_sglang_prefill_hardware_attribution.html b/docs/dsv4pro_pro6000d_2node_sglang/phase2_dsv4pro_pro6000d_2node_sglang_prefill_hardware_attribution.html new file mode 100644 index 0000000..50ba21b --- /dev/null +++ b/docs/dsv4pro_pro6000d_2node_sglang/phase2_dsv4pro_pro6000d_2node_sglang_prefill_hardware_attribution.html @@ -0,0 +1,327 @@ + + + + + + Phase 2:DeepSeek-V4-Pro 双机 Pro6000D SGLang Prefill 硬件归因 + + + +
+
+

Design, Implementation & Result Record

+

Phase 2:DeepSeek-V4-Pro 双机 Pro6000D SGLang Prefill 硬件归因

+
+ 节点:174.1.51.5 + 174.1.51.7 + 拓扑:SGLang TP16 / EP2 + 更新:2026-07-30 16:35 CST +
+
+
+ +
+ 返回推理优化主计划 + +

+ 当前状态:旧脚本口径审计完成,Phase 2 代码尚未开始。 + 本页从第一行 Phase 2 代码开始同步维护。每次代码改动、静态验证、真机运行和 + 结果判断都会在对应小节留下文件路径、命令和证据,不在阶段结束后凭记忆补写。 +

+ +

1. 为什么立即进入 Phase 2

+

+ Phase 1 在没有 Profiler、没有 Prefix Cache 命中的条件下得到以下结果: +

+ + + + + + + + + +
ISL / OSL / C输入 TPSTTFT结果
1K / 1 / 164.44 tok/s15.88 s完成
32K / 1 / 164.96 tok/s504.44 s完成
128K / 1 / 165.20 tok/s2010.38 s完成
+

+ 三个长度的输入吞吐几乎相同,TTFT 近似按 token 数线性增加。 + 这已经不是“继续扩充 Shape”能回答的问题。Phase 2 要回答: + 稳定的约 65 token/s 到底受 GPU 计算、显存、CPU 调度还是双机通信中的哪一项限制。 +

+ +

1.1 Phase 2 前置审计

+

+ 旧脚本较短的 TTFT 已确认不是同口径反例。旧脚本固定执行 16 条同 Prompt + Warm-up,从不清 Prefix Cache,并按固定 Seed 递增长度;17:40 的失败 Run + 还在 18:01 正式 Run 前预热了同一批 1K 请求。全字段比较显示新旧 + server_info 的关键运行参数相同。 +

+

+ 同一服务上的最小复现得到:两次独立清 Cache 的 1K→1 TTFT 分别为 + 16.04 秒和 15.90 秒;把 OSL 改为 128 后是 15.79 秒;按旧命令语义重新执行 + 仍为 14.60 秒,而不是旧产物的 0.455 秒。因此 Phase 2 将继续 profile + 清 Prefix Cache 后的完整冷 Prefill。Warm Prefix/Prefix Cache + 收益另立 A/B,不与本阶段混算。 +

+ +

2. 本阶段的边界

+ + +

3. 待验证假设

+ + + + + + + + + + + + +
假设预期硬件表现后续方向
DSV4/NSA Prefill Kernel 计算受限GPU 持续忙、高功耗和稳定频率;双 Rail 流量不高Phase 3 捕获 Kernel 与 Attention/Indexer 时间线
权重或激活显存带宽受限GPU Memory Utilization 高,SM 指标未必饱和;功耗可能低于纯计算补 DCGM/Profiler 的 DRAM Active,再看 Kernel
TP16 跨机通信受限RoCE 吞吐高或两条 Rail 明显失衡,GPU 出现等待NCCL_CROSS_NIC 0/1/2 快速 A/B,随后看 NCCL Timeline
CPU Scheduler 或 Kernel Launch 受限GPU 利用率锯齿或有空洞,单 CPU 核持续满载定位 Scheduler/Tokenizer 线程与 launch gap
频率、功耗或温度限制P-state、SM Clock 或 Power 持续异常,可能出现节流原因修正电源、散热或 Clock Policy 后复测
节点或 Rank 不均衡两节点或不同 GPU 的利用率、功耗、网络流量存在固定偏差检查 NUMA、GPU-NIC 亲和与慢 Rank
+ +

4. 诊断 Run

+
    +
  1. 保存两节点静态快照:GPU/NIC/NUMA 拓扑、驱动、CUDA、镜像与服务命令。
  2. +
  3. 复用 Phase 1 已验证的 run_quick_map.sh start 启动同配置双机服务。
  4. +
  5. 在 Head 和 Worker 同时启动 GPU、CPU、网卡与 RDMA 采样,先记录 15 秒空闲基线。
  6. +
  7. 清空 Prefix Cache,发送一条 32K → 1, C=1 请求,Seed 与 Phase 1 一致。
  8. +
  9. 请求结束后继续采样 15 秒,再停止采集器和服务。
  10. +
  11. 按时间戳将请求、8K Chunk、GPU、CPU 和 Rail 指标对齐,生成摘要与判定。
  12. +
+
idle 15s │──────── 32K Prefill:4 × 8K Chunk ────────│ cooldown 15s
+          ↑ request_start                              ↑ request_end
+          Head 与 Worker 的所有采集器覆盖完整时间窗
+

+ 预计服务加载约 6 分钟、请求约 8.5 分钟,连同快照和清理应在 20 分钟左右完成。 +

+ +

5. 采集指标

+ + + + + + + + + + + +
层级连续采样静态或前后快照局限
GPU利用率、Memory Utilization、显存、功耗、SM/Memory Clock、温度、P-statenvidia-smi topo -m、Compute Processnvidia-smi 的 Memory Utilization 不是实际 HBM GB/s
CPU每核利用率、上下文切换、服务进程 CPU/内存NUMA 拓扑、容器 PID 与 CPU Affinity需要时间戳与 GPU Chunk 日志对齐
Networketh0/eth3 RX/TXethtool -S 错误计数前后差普通 netdev 统计不一定覆盖所有 RDMA 细节
RDMAmlx5_0/mlx5_3 port_xmit/recv_data 差分Port State、GID 与错误计数计数单位需要按设备定义换算
DCGM若可用则记录 SM Active、DRAM Active、Tensor Active、PCIe工具版本与可用 Field不可用时明确记录,不能用粗粒度指标冒充
+ +

6. 精简代码设计

+

计划新增目录:

+
/data/hzy/sskj/experiments/pro6000/
+dsv4pro_pro6000d_2node_sglang_prefill_hardware_attribution/
+ + + + + + + + + + + +
文件计划职责当前状态
run_prefill_hardware_attribution.sh唯一 Shell 入口;服务启停、双节点采集器、单 Case、Trap 清理待实现
config.envPhase 1 入口路径、32K Case、采样间隔、结果路径待实现
hardware_attribution.py结构化解析、时间对齐、统计摘要与报告生成待实现
tests/test_hardware_attribution.py计数器差分、单位换算、统计与缺失工具回退测试待实现
README.md入口命令、环境变量和结果目录说明待实现
+

+ Phase 2 不复制双机 Docker 启停实现。唯一入口通过环境变量调用 Phase 1 的 + run_quick_map.sh start/stop,只新增硬件采集和 32K 请求编排。 + 顶层仍只保留一个 Shell 文件,不创建额外 tmux、launch、start 或 stop 脚本。 +

+ +

7. 预期结果结构

+
results/<RUN_ID>/
+  manifest.json
+  run.log
+  bench/
+    bench_cmd.txt
+    bench.log
+    bench.jsonl
+  service/
+    head_server_cmd.txt
+    worker_server_cmd.txt
+    head_server.log
+    worker_server.log
+  head/
+    gpu.csv
+    cpu_mpstat.log
+    cpu_pidstat.log
+    net_sar.log
+    rdma.csv
+    static/
+  worker/
+    gpu.csv
+    cpu_mpstat.log
+    cpu_pidstat.log
+    net_sar.log
+    rdma.csv
+    static/
+  summary.json
+  summary.csv
+  report.md
+ +

8. 验收条件

+ + +

9. 实施记录

+ + + + + + + + +
时间代码或运行结果
2026-07-30 15:46 CST创建 Phase 2 设计与档案代码尚未开始,等待按本页设计实现
2026-07-30 16:35 CST完成旧脚本与 quick-map 同口径审计排除服务参数、OSL=1 和一次性 JIT;确认旧产物被 Warm-up、跨 Case 与前一轮 Prefix Cache 污染
+ +

10. 真机结果

+

+ 尚未运行。代码实现、静态验证和 Dry-run 完成后,将先向用户说明具体代码改动, + 再启动真机诊断。 +

+ +

返回 Phase 1 实施记录

+

返回推理优化主计划

+
+ + diff --git a/docs/dsv4pro_pro6000d_2node_sglang/results/phase1-dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/report.md b/docs/dsv4pro_pro6000d_2node_sglang/results/phase1-dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/report.md new file mode 100644 index 0000000..16d644d --- /dev/null +++ b/docs/dsv4pro_pro6000d_2node_sglang/results/phase1-dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/report.md @@ -0,0 +1,20 @@ +# DeepSeek-V4-Pro Pro6000D Two-Node SGLang Quick Map + +Profiler: disabled. Speculative decoding: disabled. + +## Aggregate results + +| Case | Suite / role | Stage | ISL | OSL | C | Reps | Total TPS | CV | Output TPS | TTFT P95 | TPOT P95 | E2E P95 | Status | +|---|---|---|---:|---:|---:|---:|---:|---:|---:|---:|---:|---:|---| +| long_prefill_latency_128k_c1 | fixed / - | prefill_latency | 131072 | 1 | 1 | 1/1 | 65.20 | -% | 0.00 | 2010382.63 ms | 0.00 ms | 2010382.70 ms | COMPLETED | +| mid_prefill_latency_32k_c1 | fixed / - | prefill_latency | 32768 | 1 | 1 | 1/1 | 64.96 | -% | 0.00 | 504435.18 ms | 0.00 ms | 504435.25 ms | COMPLETED | +| mid_prefill_throughput_32k_c16 | fixed / - | prefill_throughput | 32768 | 1 | 16 | 0/1 | - | -% | - | - ms | - ms | - ms | ABORTED | +| short_prefill_latency_1k_c1 | fixed / - | prefill_latency | 1024 | 1 | 1 | 1/1 | 64.44 | -% | 0.06 | 15882.76 ms | 0.00 ms | 15882.82 ms | COMPLETED | + +## Failed or incomplete cases + +| Case | Repetition | Status | Error | Exit code | +|---|---:|---|---|---:| +| mid_prefill_throughput_32k_c16 | 1 | ABORTED | EARLY_STOP_FOR_PHASE2 | 143 | + +The fixed map is descriptive and never stops on SLO. With one repetition, CV is intentionally unavailable. diff --git a/docs/dsv4pro_pro6000d_2node_sglang/results/phase1-dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/run_manifest.json b/docs/dsv4pro_pro6000d_2node_sglang/results/phase1-dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/run_manifest.json new file mode 100644 index 0000000..9bb12fd --- /dev/null +++ b/docs/dsv4pro_pro6000d_2node_sglang/results/phase1-dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/run_manifest.json @@ -0,0 +1,40 @@ +{ + "schema_version": 1, + "workflow_stage": "quick_performance_map", + "run_id": "dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625", + "status": "ABORTED_EARLY_FOR_PHASE2", + "started_at": "2026-07-30T14:36:26+08:00", + "updated_at": "2026-07-30T15:44:11+08:00", + "suites": [ + "fixed", + "mixed" + ], + "engine": "sglang", + "model_name": "DeepSeek-V4-Pro", + "model_path": "/data/hf_models/DeepSeek-V4-Pro", + "docker_image": "lmsysorg/sglang:nightly-dev-cu13-20260720-b3570a45", + "head_node": "10.101.0.11", + "worker_node": "10.101.0.13", + "head_ip": "10.101.0.11", + "sglang_port": 30002, + "dist_init_port": 20002, + "tp_size": 16, + "ep_size": 2, + "nnodes": 2, + "mem_fraction_static": 0.9, + "cuda_graph_max_bs_decode": 64, + "max_running_requests": 256, + "nccl_socket_ifname": "eth1", + "nccl_ib_hca": "mlx5_0,mlx5_3", + "nccl_cross_nic": "1", + "git_commit": "d5d96bd6f60f7fcdf0e118070ace468698949428", + "git_dirty": false, + "scenario_file": "/data/hzy/sskj/experiments/pro6000/dsv4pro_pro6000d_2node_sglang_tp16_quick_map/quick_map_scenarios.tsv", + "notes": [ + "The fixed quick map does not stop on SLO.", + "Profiler is disabled; these results are eligible for performance comparison.", + "Speculative decoding is not enabled.", + "Source tree was committed unchanged during model initialization." + ], + "ended_at": "2026-07-30T15:44:11+08:00" +} diff --git a/docs/dsv4pro_pro6000d_2node_sglang/results/phase1-dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/summary.jsonl b/docs/dsv4pro_pro6000d_2node_sglang/results/phase1-dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/summary.jsonl new file mode 100644 index 0000000..8997991 --- /dev/null +++ b/docs/dsv4pro_pro6000d_2node_sglang/results/phase1-dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/summary.jsonl @@ -0,0 +1,4 @@ +{"run_id": "dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625", "suite": "fixed", "case_id": "long_prefill_latency_128k_c1", "role": "", "stage": "prefill_latency", "repetition": 1, "isl": 131072, "osl": 1, "concurrency": 1, "num_prompts": 1, "warmup_requests": 0, "status": "COMPLETED", "error_type": "", "exit_code": 0, "started_at": "2026-07-30T14:52:19+0800", "ended_at": "2026-07-30T15:26:24+0800", "elapsed_s": 2045.0, "completed": 1, "failed": 0, "duration_s": 2010.4037326959951, "actual_concurrency": 0.999989539343442, "peak_concurrent_requests": null, "total_input_tokens": 131072, "total_output_tokens": 1, "request_throughput": 0.0004974125265172375, "input_token_throughput": 65.19685467566735, "output_token_throughput": 0.0004974125265172375, "total_token_throughput": 65.19735208819388, "peak_output_token_throughput": null, "e2e_mean_ms": 2010382.7025530045, "e2e_p50_ms": 2010382.7025530045, "e2e_p95_ms": 2010382.7025530045, "e2e_p99_ms": 2010382.7025530045, "ttft_mean_ms": 2010382.6258230256, "ttft_p50_ms": 2010382.6258230256, "ttft_p95_ms": 2010382.6258230256, "ttft_p99_ms": 2010382.6258230256, "tpot_mean_ms": 0.0, "tpot_p50_ms": 0.0, "tpot_p95_ms": 0.0, "tpot_p99_ms": 0.0, "itl_mean_ms": 0.0, "itl_p50_ms": 0.0, "itl_p95_ms": 0.0, "itl_p99_ms": 0.0, "bench_file": "/data/hzy/sskj/experiments/pro6000/dsv4pro_pro6000d_2node_sglang_tp16_quick_map/results/dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/cases/long_prefill_latency_128k_c1/rep1/bench.jsonl", "bench_log": "/data/hzy/sskj/experiments/pro6000/dsv4pro_pro6000d_2node_sglang_tp16_quick_map/results/dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/cases/long_prefill_latency_128k_c1/rep1/bench.log"} +{"run_id": "dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625", "suite": "fixed", "case_id": "mid_prefill_latency_32k_c1", "role": "", "stage": "prefill_latency", "repetition": 1, "isl": 32768, "osl": 1, "concurrency": 1, "num_prompts": 1, "warmup_requests": 0, "status": "COMPLETED", "error_type": "", "exit_code": 0, "started_at": "2026-07-30T14:43:15+0800", "ended_at": "2026-07-30T14:52:13+0800", "elapsed_s": 538.0, "completed": 1, "failed": 0, "duration_s": 504.4536426469858, "actual_concurrency": 0.9999635371827135, "peak_concurrent_requests": null, "total_input_tokens": 32768, "total_output_tokens": 1, "request_throughput": 0.001982342707949866, "input_token_throughput": 64.95740585410121, "output_token_throughput": 0.001982342707949866, "total_token_throughput": 64.95938819680917, "peak_output_token_throughput": null, "e2e_mean_ms": 504435.24884598446, "e2e_p50_ms": 504435.24884598446, "e2e_p95_ms": 504435.24884598446, "e2e_p99_ms": 504435.24884598446, "ttft_mean_ms": 504435.1755149837, "ttft_p50_ms": 504435.1755149837, "ttft_p95_ms": 504435.1755149837, "ttft_p99_ms": 504435.1755149837, "tpot_mean_ms": 0.0, "tpot_p50_ms": 0.0, "tpot_p95_ms": 0.0, "tpot_p99_ms": 0.0, "itl_mean_ms": 0.0, "itl_p50_ms": 0.0, "itl_p95_ms": 0.0, "itl_p99_ms": 0.0, "bench_file": "/data/hzy/sskj/experiments/pro6000/dsv4pro_pro6000d_2node_sglang_tp16_quick_map/results/dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/cases/mid_prefill_latency_32k_c1/rep1/bench.jsonl", "bench_log": "/data/hzy/sskj/experiments/pro6000/dsv4pro_pro6000d_2node_sglang_tp16_quick_map/results/dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/cases/mid_prefill_latency_32k_c1/rep1/bench.log"} +{"run_id": "dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625", "suite": "fixed", "case_id": "mid_prefill_throughput_32k_c16", "role": "", "stage": "prefill_throughput", "repetition": 1, "isl": 32768, "osl": 1, "concurrency": 16, "num_prompts": 16, "warmup_requests": 0, "status": "ABORTED", "error_type": "EARLY_STOP_FOR_PHASE2", "exit_code": 143, "started_at": "2026-07-30T15:26:30+0800", "ended_at": "2026-07-30T15:41:52+0800", "elapsed_s": 922.0, "completed": null, "failed": null, "duration_s": null, "actual_concurrency": null, "peak_concurrent_requests": null, "total_input_tokens": null, "total_output_tokens": null, "request_throughput": null, "input_token_throughput": null, "output_token_throughput": null, "total_token_throughput": null, "peak_output_token_throughput": null, "e2e_mean_ms": null, "e2e_p50_ms": null, "e2e_p95_ms": null, "e2e_p99_ms": null, "ttft_mean_ms": null, "ttft_p50_ms": null, "ttft_p95_ms": null, "ttft_p99_ms": null, "tpot_mean_ms": null, "tpot_p50_ms": null, "tpot_p95_ms": null, "tpot_p99_ms": null, "itl_mean_ms": null, "itl_p50_ms": null, "itl_p95_ms": null, "itl_p99_ms": null, "bench_file": "/data/hzy/sskj/experiments/pro6000/dsv4pro_pro6000d_2node_sglang_tp16_quick_map/results/dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/cases/mid_prefill_throughput_32k_c16/rep1/bench.jsonl", "bench_log": "/data/hzy/sskj/experiments/pro6000/dsv4pro_pro6000d_2node_sglang_tp16_quick_map/results/dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/cases/mid_prefill_throughput_32k_c16/rep1/bench.log"} +{"run_id": "dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625", "suite": "fixed", "case_id": "short_prefill_latency_1k_c1", "role": "", "stage": "prefill_latency", "repetition": 1, "isl": 1024, "osl": 1, "concurrency": 1, "num_prompts": 1, "warmup_requests": 1, "status": "COMPLETED", "error_type": "", "exit_code": 0, "started_at": "2026-07-30T14:42:04+0800", "ended_at": "2026-07-30T14:43:10+0800", "elapsed_s": 66.0, "completed": 1, "failed": 0, "duration_s": 15.905719314003363, "actual_concurrency": 0.9985603391737933, "peak_concurrent_requests": null, "total_input_tokens": 1024, "total_output_tokens": 1, "request_throughput": 0.06287046692189532, "input_token_throughput": 64.37935812802081, "output_token_throughput": 0.06287046692189532, "total_token_throughput": 64.4422285949427, "peak_output_token_throughput": null, "e2e_mean_ms": 15882.820472994354, "e2e_p50_ms": 15882.820472994354, "e2e_p95_ms": 15882.820472994354, "e2e_p99_ms": 15882.820472994354, "ttft_mean_ms": 15882.760226988466, "ttft_p50_ms": 15882.760226988466, "ttft_p95_ms": 15882.760226988466, "ttft_p99_ms": 15882.760226988466, "tpot_mean_ms": 0.0, "tpot_p50_ms": 0.0, "tpot_p95_ms": 0.0, "tpot_p99_ms": 0.0, "itl_mean_ms": 0.0, "itl_p50_ms": 0.0, "itl_p95_ms": 0.0, "itl_p99_ms": 0.0, "bench_file": "/data/hzy/sskj/experiments/pro6000/dsv4pro_pro6000d_2node_sglang_tp16_quick_map/results/dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/cases/short_prefill_latency_1k_c1/rep1/bench.jsonl", "bench_log": "/data/hzy/sskj/experiments/pro6000/dsv4pro_pro6000d_2node_sglang_tp16_quick_map/results/dsv4pro-pro6000d-2node-sglang-quick-v2-20260730-143625/cases/short_prefill_latency_1k_c1/rep1/bench.log"} diff --git a/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/cold_1k_o128.jsonl b/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/cold_1k_o128.jsonl new file mode 100644 index 0000000..5f0866d --- /dev/null +++ b/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/cold_1k_o128.jsonl @@ -0,0 +1 @@ +{"tag": null, "backend": "sglang", "dataset_name": "random", "request_rate": 10000.0, "max_concurrency": 1, "sharegpt_output_len": null, "random_input_len": 1024, "random_output_len": 128, "random_range_ratio": 1.0, "server_info": {"model_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_mode": "auto", "tokenizer_backend": "huggingface", "tokenizer_worker_num": 1, "detokenizer_worker_num": 1, "skip_tokenizer_init": false, "load_format": "auto", "model_loader_extra_config": "{}", "trust_remote_code": true, "context_length": null, "is_embedding": false, "enable_multimodal": null, "revision": null, "model_impl": "auto", "model_config_parser": "auto", "json_model_override_args": "{}", "dtype": "auto", "quantization": null, "quantization_param_path": null, "kv_cache_dtype": "fp8_e4m3", "enable_fp32_lm_head": false, "modelopt_quant": null, "modelopt_checkpoint_restore_path": null, "modelopt_checkpoint_save_path": null, "modelopt_export_path": null, "quantize_and_serve": false, "rl_quant_profile": null, "enable_tf32_matmul": false, "mem_fraction_static": 0.9, "max_running_requests": 256, "max_queued_requests": null, "max_total_tokens": null, "chunked_prefill_size": 8192, "enable_dynamic_chunking": false, "max_prefill_tokens": 16384, "prefill_max_requests": null, "schedule_policy": "fcfs", "enable_priority_scheduling": false, "disable_priority_preemption": false, "default_priority_value": null, "abort_on_priority_when_disabled": false, "schedule_low_priority_values_first": false, "priority_scheduling_preemption_threshold": 10, "retraction_policy": "length", "schedule_conservativeness": 1.0, "page_size": 256, "swa_full_tokens_ratio": 0.1, "disable_hybrid_swa_memory": false, "radix_eviction_policy": "lru", "prefill_only_disable_kv_cache": false, "disable_radix_cache": false, "enable_page_major_kv_layout": false, "enable_unified_memory": false, "disable_chunked_prefix_cache": false, "disable_overlap_schedule": false, "num_continuous_decode_steps": 1, "scheduler_recv_interval": 1, "enable_mixed_chunk": false, "nccl_port": null, "dist_timeout": null, "dist_init_addr": "10.101.0.11:20002", "nnodes": 2, "node_rank": 0, "tp_size": 16, "dcp_size": 1, "pp_size": 1, "pp_max_micro_batch_size": null, "pp_async_batch_depth": 0, "dp_size": 1, "load_balance_method": "round_robin", "attn_cp_size": 1, "moe_dp_size": 1, "enable_prefill_cp": false, "cp_strategy": null, "enable_dsa_cache_layer_split": false, "enable_dsa_prefill_context_parallel": false, "dsa_prefill_cp_mode": "round-robin-split", "enable_prefill_context_parallel": false, "prefill_cp_mode": "in-seq-split", "enable_dp_attention": false, "enable_dp_attention_local_control_broadcast": false, "enable_dp_lm_head": false, "enable_attn_tp_input_scattered": false, "disable_attn_tp_gather": false, "enable_p2p_check": false, "device": "cuda", "base_gpu_id": 0, "gpu_id_step": 1, "random_seed": 42600328, "watchdog_timeout": 300, "soft_watchdog_timeout": null, "sleep_on_idle": false, "use_ray": false, "custom_sigquit_handler": null, "numa_node": null, "gc_threshold": null, "host": "0.0.0.0", "port": 30002, "fastapi_root_path": "", "smg_grpc_mode": false, "grpc_mode": false, "grpc_port": null, "skip_server_warmup": false, "warmups": null, "enable_http2": false, "ssl_keyfile": null, "ssl_certfile": null, "ssl_ca_certs": null, "ssl_keyfile_password": null, "enable_ssl_refresh": false, "api_key": null, "admin_api_key": null, "served_model_name": "/data/hf_models/DeepSeek-V4-Pro", "weight_version": "default", "chat_template": null, "hf_chat_template_name": null, "completion_template": null, "file_storage_path": "sglang_storage", "enable_cache_report": false, "reasoning_parser": null, "default_chat_template_kwargs": null, "strip_thinking_cache": false, "enable_strict_thinking": false, "tool_call_parser": null, "tool_server": null, "sampling_defaults": "model", "asr_max_buffer_seconds": 60, "asr_max_concurrent_sessions": 32, "preferred_sampling_params": null, "allow_auto_truncate": false, "stream_interval": 1, "batch_notify_size": 16, "stream_response_default_include_usage": false, "incremental_streaming_output": false, "enable_streaming_session": false, "enable_session_radix_cache": false, "log_level": "info", "log_level_http": null, "log_requests": false, "log_requests_level": 2, "log_requests_format": "text", "log_requests_target": null, "uvicorn_access_log_exclude_prefixes": [], "crash_dump_folder": null, "show_time_cost": false, "enable_metrics": false, "smg_http_sidecar_port": null, "enable_mfu_metrics": false, "enable_metrics_for_all_schedulers": false, "load_snapshot_publish_interval": 15, "tokenizer_metrics_custom_labels_header": "x-custom-labels", "tokenizer_metrics_allowed_custom_labels": null, "extra_metric_labels": null, "bucket_time_to_first_token": null, "bucket_inter_token_latency": null, "bucket_e2e_request_latency": null, "prompt_tokens_buckets": null, "generation_tokens_buckets": null, "gc_warning_threshold_secs": 0.0, "decode_log_interval": 40, "enable_request_time_stats_logging": false, "kv_events_config": null, "enable_forward_pass_metrics": false, "forward_pass_metrics_worker_id": "", "forward_pass_metrics_ipc_name": null, "enable_trace": false, "trace_modules": "request", "otlp_traces_endpoint": "localhost:4317", "export_metrics_to_file": false, "export_metrics_to_file_dir": null, "stat_loggers": null, "constrained_json_whitespace_pattern": null, "constrained_json_disable_any_whitespace": false, "attention_backend": "dsv4", "decode_attention_backend": null, "prefill_attention_backend": null, "sampling_backend": "flashinfer", "grammar_backend": "xgrammar", "radix_cache_backend": null, "mm_attention_backend": null, "fp8_gemm_runner_backend": "auto", "fp4_gemm_runner_backend": "auto", "bf16_gemm_backend": "auto", "dsa_prefill_backend": null, "dsa_decode_backend": null, "dsa_paged_mqa_logits_backend": "auto", "dsa_topk_backend": "sgl-kernel", "disable_flashinfer_autotune": false, "mamba_backend": "triton", "cuda_graph_config": {"decode": {"backend": "full", "max_bs": 64, "bs": [1, 2, 4, 8, 12, 16, 24, 32, 40, 48, 56, 64], "tc_compiler": "eager", "full_prefill_max_req": null}, "prefill": {"backend": "disabled", "max_bs": 8192, "bs": [4, 8, 12, 16, 20, 24, 28, 32, 48, 64, 80, 96, 112, 128, 144, 160, 176, 192, 208, 224, 240, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 704, 768, 832, 896, 960, 1024, 1280, 1536, 1792, 2048, 2304, 2560, 2816, 3072, 3328, 3584, 3840, 4096, 4608, 5120, 5632, 6144, 6656, 7168, 7680, 8192], "tc_compiler": "eager", "full_prefill_max_req": null}}, "cuda_graph_backend_decode": null, "cuda_graph_backend_prefill": null, "cuda_graph_max_bs_decode": 64, "cuda_graph_max_bs_prefill": null, "cuda_graph_bs_decode": null, "cuda_graph_bs_prefill": null, "cuda_graph_tc_compiler": null, "disable_prefill_cuda_graph": false, "disable_decode_cuda_graph": false, "disable_cuda_graph": false, "disable_cuda_graph_padding": false, "enable_profile_cuda_graph": false, "enable_cudagraph_gc": false, "debug_cuda_graph": false, "enable_layerwise_nvtx_marker": false, "enable_nccl_nvls": false, "enable_symm_mem": false, "triton_attention_reduce_in_fp32": false, "triton_attention_num_kv_splits": 8, "triton_attention_split_tile_size": null, "flashinfer_mla_disable_ragged": false, "enable_fused_qk_norm_rope": false, "enable_precise_embedding_interpolation": false, "enable_fused_moe_sum_all_reduce": false, "enable_deepseek_v4_fp4_indexer": false, "disable_custom_all_reduce": false, "enable_mscclpp": false, "enable_torch_symm_mem": false, "pre_warm_nccl": false, "enable_quant_communications": false, "enable_flashinfer_allreduce_fusion": false, "enforce_disable_flashinfer_allreduce_fusion": false, "flashinfer_allreduce_fusion_backend": null, "enable_aiter_allreduce_fusion": false, "enable_torch_compile": false, "enable_torch_compile_debug_mode": false, "torch_compile_max_bs": 32, "torchao_config": "", "speculative_algorithm": null, "speculative_draft_model_path": null, "speculative_draft_model_revision": null, "speculative_draft_load_format": null, "speculative_num_steps": null, "speculative_eagle_topk": null, "speculative_num_draft_tokens": null, "speculative_dflash_block_size": null, "speculative_dspark_block_size": null, "speculative_dspark_sps_table_path": null, "speculative_dspark_confidence_sts_path": null, "speculative_dspark_align_verify_tokens_to_graph_tier": false, "speculative_accept_threshold_single": 1.0, "speculative_accept_threshold_acc": 1.0, "speculative_use_rejection_sampling": false, "speculative_token_map": null, "speculative_attention_mode": "prefill", "speculative_draft_attention_backend": null, "speculative_draft_window_size": null, "speculative_moe_runner_backend": "flashinfer_mxfp4", "speculative_moe_a2a_backend": null, "speculative_draft_model_quantization": null, "speculative_skip_dp_mlp_sync": false, "enable_multi_layer_eagle": false, "speculative_adaptive": false, "speculative_adaptive_config": null, "decoupled_spec_bind_endpoint": null, "decoupled_spec_connect_endpoints": null, "decoupled_spec_rank": null, "decoupled_spec_role": "null", "spec_trace_dir": null, "speculative_ngram_min_bfs_breadth": 1, "speculative_ngram_max_bfs_breadth": 10, "speculative_ngram_match_type": "BFS", "speculative_ngram_max_trie_depth": 18, "speculative_ngram_capacity": 10000000, "speculative_ngram_external_corpus_path": null, "speculative_ngram_external_sam_budget": 0, "speculative_ngram_external_corpus_max_tokens": 10000000, "ep_size": 2, "moe_a2a_backend": "none", "moe_runner_backend": "flashinfer_mxfp4", "flashinfer_mxfp4_moe_precision": "default", "deepep_mode": "auto", "fuseep_mode": 2, "deepep_dispatcher_output_dtype": "auto", "ep_num_redundant_experts": 0, "ep_dispatch_algorithm": null, "init_expert_location": "trivial", "enable_eplb": false, "eplb_algorithm": "auto", "eplb_rebalance_num_iterations": 1000, "eplb_rebalance_layers_per_chunk": null, "eplb_min_rebalancing_utilization_threshold": 1.0, "expert_distribution_recorder_mode": null, "expert_distribution_recorder_buffer_size": 1000, "enable_expert_distribution_metrics": false, "deepep_config": null, "moe_dense_tp_size": null, "elastic_ep_backend": null, "enable_elastic_expert_backup": false, "mooncake_ib_device": null, "enable_waterfill": false, "ep_join_mode": null, "ep_join_rank_offset": 0, "elastic_ep_initial_size": null, "max_ep_size": null, "elastic_ep_scale_timeout": 600, "elastic_ep_rejoin": false, "disable_flashinfer_cutlass_moe_fp4_allgather": false, "disable_shared_experts_fusion": false, "enforce_shared_experts_fusion": false, "max_mamba_cache_size": null, "mamba_ssm_dtype": null, "enable_mamba_cache_stochastic_rounding": false, "mamba_cache_philox_rounds": 0, "mamba_full_memory_ratio": 0.9, "mamba_radix_cache_strategy": "auto", "uses_mamba_radix_cache": false, "mamba_track_interval": 256, "enable_int8_mamba_checkpoint": false, "int8_mamba_ckpt_size": null, "linear_attn_backend": "triton", "linear_attn_decode_backend": null, "linear_attn_prefill_backend": null, "enable_linear_replayssm": false, "linear_replayssm_cache_len": 16, "enable_hierarchical_cache": false, "hicache_ratio": 2.0, "hicache_size": 0, "hicache_write_policy": "write_through", "hicache_io_backend": "kernel", "hicache_mem_layout": "page_first", "hicache_storage_backend": null, "hicache_storage_prefetch_policy": "timeout", "hicache_storage_backend_extra_config": null, "enable_hisparse": false, "hisparse_config": null, "enable_broadcast_mm_inputs_process": false, "enable_prefix_mm_cache": false, "mm_enable_dp_encoder": false, "mm_process_config": {}, "limit_mm_data_per_request": null, "enable_mm_global_cache": false, "disable_fast_image_processor": false, "mm_feature_transport": "cpu", "keep_mm_feature_on_device": false, "enable_lora": null, "enable_lora_overlap_loading": null, "max_lora_rank": null, "lora_target_modules": null, "lora_paths": null, "max_loaded_loras": null, "max_loras_per_batch": 8, "lora_eviction_policy": "lru", "lora_backend": "csgmv", "max_lora_chunk_size": 16, "experts_shared_outer_loras": null, "lora_use_virtual_experts": false, "lora_strict_loading": false, "lora_drain_wait_threshold": 0.0, "enable_two_batch_overlap": false, "enable_single_batch_overlap": false, "tbo_token_distribution_threshold": 0.48, "cpu_offload_gb": 0, "offload_group_size": -1, "offload_num_in_group": 1, "offload_prefetch_step": 1, "offload_mode": "cpu", "enable_lmcache": false, "lmcache_config_file": null, "enable_flexkv": false, "flexkv_config_file": null, "kt_weight_path": null, "kt_method": "AMXINT4", "kt_cpuinfer": null, "kt_threadpool_count": 2, "kt_num_gpu_experts": null, "kt_max_deferred_experts_per_token": null, "dllm_algorithm": null, "dllm_algorithm_config": null, "dllm_fdfo": true, "disaggregation_mode": "null", "disaggregation_transfer_backend": "mooncake", "disaggregation_bootstrap_port": 8998, "disaggregation_ib_device": null, "disaggregation_decode_enable_radix_cache": false, "disaggregation_decode_enable_offload_kvcache": false, "num_reserved_decode_tokens": 512, "disaggregation_decode_extra_slots": null, "disaggregation_decode_polling_interval": 1, "optimistic_prefill_attempts": 0, "encoder_only": false, "language_only": false, "encoder_transfer_backend": "zmq_to_scheduler", "encoder_urls": [], "encoder_bootstrap_port": 8997, "encoder_register_urls": [], "enable_adaptive_dispatch_to_encoder": false, "enable_pdmux": false, "pdmux_config_path": null, "sm_group_num": 8, "custom_weight_loader": [], "weight_loader_disable_mmap": false, "weight_loader_prefetch_checkpoints": false, "weight_loader_prefetch_num_threads": 4, "weight_loader_drop_cache_after_load": false, "remote_instance_weight_loader_seed_instance_ip": null, "remote_instance_weight_loader_seed_instance_service_port": null, "remote_instance_weight_loader_send_weights_group_ports": null, "remote_instance_weight_loader_backend": "nccl", "remote_instance_weight_loader_start_seed_via_transfer_engine": false, "engine_info_bootstrap_port": 6789, "modelexpress_config": null, "download_dir": null, "model_checksum": null, "delete_ckpt_after_loading": false, "decrypted_config_file": null, "decrypted_draft_config_file": null, "checkpoint_engine_wait_weights_before_ready": false, "enable_prefill_delayer": false, "prefill_delayer_max_delay_passes": 30, "prefill_delayer_token_usage_low_watermark": null, "prefill_delayer_forward_passes_buckets": null, "prefill_delayer_wait_seconds_buckets": null, "prefill_delayer_queue_min_ratio": null, "prefill_delayer_max_delay_ms": null, "min_free_slots_delay": null, "enable_deterministic_inference": false, "rl_on_policy_target": null, "kv_canary": "none", "kv_canary_real_data": "none", "kv_canary_sweep_interval": 0, "enable_dynamic_batch_tokenizer": false, "dynamic_batch_tokenizer_batch_size": 32, "dynamic_batch_tokenizer_batch_timeout": 0.002, "enable_tokenizer_batch_encode": false, "disable_tokenizer_batch_decode": false, "debug_tensor_dump_output_folder": null, "debug_tensor_dump_layers": null, "debug_tensor_dump_input_file": null, "enable_memory_saver": false, "enable_weights_cpu_backup": false, "enable_draft_weights_cpu_backup": false, "enable_custom_logit_processor": false, "enable_return_hidden_states": false, "enable_return_routed_experts": false, "enable_return_indexer_topk": false, "disable_outlines_disk_cache": false, "enable_mis": false, "forward_hooks": null, "msprobe_dump_config": null, "status": "ready", "max_total_num_tokens": 1282304, "max_req_input_len": 1048570, "internal_states": [{"model_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_mode": "auto", "tokenizer_backend": "huggingface", "tokenizer_worker_num": 1, "detokenizer_worker_num": 1, "skip_tokenizer_init": false, "load_format": "auto", "model_loader_extra_config": "{}", "trust_remote_code": true, "context_length": null, "is_embedding": false, "enable_multimodal": null, "revision": null, "model_impl": "auto", "model_config_parser": "auto", "json_model_override_args": "{}", "dtype": "auto", "quantization": null, "quantization_param_path": null, "kv_cache_dtype": "fp8_e4m3", "enable_fp32_lm_head": false, "modelopt_quant": null, "modelopt_checkpoint_restore_path": null, "modelopt_checkpoint_save_path": null, "modelopt_export_path": null, "quantize_and_serve": false, "rl_quant_profile": null, "enable_tf32_matmul": false, "mem_fraction_static": 0.9, "max_running_requests": 256, "max_queued_requests": null, "max_total_tokens": null, "chunked_prefill_size": 8192, "enable_dynamic_chunking": false, "max_prefill_tokens": 16384, "prefill_max_requests": null, "schedule_policy": "fcfs", "enable_priority_scheduling": false, "disable_priority_preemption": false, "default_priority_value": null, "abort_on_priority_when_disabled": false, "schedule_low_priority_values_first": false, "priority_scheduling_preemption_threshold": 10, "retraction_policy": "length", "schedule_conservativeness": 1.0, "page_size": 256, "swa_full_tokens_ratio": 0.1, "disable_hybrid_swa_memory": false, "radix_eviction_policy": "lru", "prefill_only_disable_kv_cache": false, "disable_radix_cache": false, "enable_page_major_kv_layout": false, "enable_unified_memory": false, "disable_chunked_prefix_cache": true, "disable_overlap_schedule": false, "num_continuous_decode_steps": 1, "scheduler_recv_interval": 1, "enable_mixed_chunk": false, "nccl_port": null, "dist_timeout": null, "dist_init_addr": "10.101.0.11:20002", "nnodes": 2, "node_rank": 0, "tp_size": 16, "dcp_size": 1, "pp_size": 1, "pp_max_micro_batch_size": 256, "pp_async_batch_depth": 0, "dp_size": 1, "load_balance_method": "round_robin", "attn_cp_size": 1, "moe_dp_size": 1, "enable_prefill_cp": false, "cp_strategy": null, "enable_dsa_cache_layer_split": false, "enable_dsa_prefill_context_parallel": false, "dsa_prefill_cp_mode": "round-robin-split", "enable_prefill_context_parallel": false, "prefill_cp_mode": "in-seq-split", "enable_dp_attention": false, "enable_dp_attention_local_control_broadcast": false, "enable_dp_lm_head": false, "enable_attn_tp_input_scattered": false, "disable_attn_tp_gather": false, "enable_p2p_check": false, "device": "cuda", "base_gpu_id": 0, "gpu_id_step": 1, "random_seed": 42600328, "watchdog_timeout": 300, "soft_watchdog_timeout": null, "sleep_on_idle": false, "use_ray": false, "numa_node": null, "gc_threshold": null, "host": "0.0.0.0", "port": 30002, "fastapi_root_path": "", "smg_grpc_mode": false, "grpc_mode": false, "grpc_port": null, "skip_server_warmup": false, "warmups": null, "enable_http2": false, "ssl_keyfile": null, "ssl_certfile": null, "ssl_ca_certs": null, "ssl_keyfile_password": null, "enable_ssl_refresh": false, "api_key": null, "admin_api_key": null, "served_model_name": "/data/hf_models/DeepSeek-V4-Pro", "weight_version": "default", "chat_template": null, "hf_chat_template_name": null, "completion_template": null, "file_storage_path": "sglang_storage", "enable_cache_report": false, "reasoning_parser": null, "default_chat_template_kwargs": null, "strip_thinking_cache": false, "enable_strict_thinking": false, "tool_call_parser": null, "tool_server": null, "sampling_defaults": "model", "asr_max_buffer_seconds": 60, "asr_max_concurrent_sessions": 32, "preferred_sampling_params": null, "allow_auto_truncate": false, "stream_interval": 1, "batch_notify_size": 16, "stream_response_default_include_usage": false, "incremental_streaming_output": false, "enable_streaming_session": false, "enable_session_radix_cache": false, "log_level": "info", "log_level_http": null, "log_requests": false, "log_requests_level": 2, "log_requests_format": "text", "log_requests_target": null, "uvicorn_access_log_exclude_prefixes": [], "crash_dump_folder": null, "show_time_cost": false, "enable_metrics": false, "smg_http_sidecar_port": null, "enable_mfu_metrics": false, "enable_metrics_for_all_schedulers": false, "load_snapshot_publish_interval": 15, "tokenizer_metrics_custom_labels_header": "x-custom-labels", "tokenizer_metrics_allowed_custom_labels": null, "extra_metric_labels": null, "bucket_time_to_first_token": null, "bucket_inter_token_latency": null, "bucket_e2e_request_latency": null, "prompt_tokens_buckets": null, "generation_tokens_buckets": null, "gc_warning_threshold_secs": 0.0, "decode_log_interval": 40, "enable_request_time_stats_logging": false, "kv_events_config": null, "enable_forward_pass_metrics": false, "forward_pass_metrics_worker_id": "", "forward_pass_metrics_ipc_name": null, "enable_trace": false, "trace_modules": "request", "otlp_traces_endpoint": "localhost:4317", "export_metrics_to_file": false, "export_metrics_to_file_dir": null, "stat_loggers": null, "constrained_json_whitespace_pattern": null, "constrained_json_disable_any_whitespace": false, "attention_backend": "dsv4", "decode_attention_backend": null, "prefill_attention_backend": null, "sampling_backend": "flashinfer", "grammar_backend": "xgrammar", "radix_cache_backend": null, "mm_attention_backend": null, "fp8_gemm_runner_backend": "auto", "fp4_gemm_runner_backend": "auto", "bf16_gemm_backend": "auto", "dsa_prefill_backend": null, "dsa_decode_backend": null, "dsa_paged_mqa_logits_backend": "auto", "dsa_topk_backend": "sgl-kernel", "disable_flashinfer_autotune": false, "mamba_backend": "triton", "cuda_graph_config": {"decode": {"backend": "full", "max_bs": 64, "bs": [1, 2, 4, 8, 12, 16, 24, 32, 40, 48, 56, 64], "tc_compiler": "eager", "full_prefill_max_req": null}, "prefill": {"backend": "disabled", "max_bs": 8192, "bs": [4, 8, 12, 16, 20, 24, 28, 32, 48, 64, 80, 96, 112, 128, 144, 160, 176, 192, 208, 224, 240, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 704, 768, 832, 896, 960, 1024, 1280, 1536, 1792, 2048, 2304, 2560, 2816, 3072, 3328, 3584, 3840, 4096, 4608, 5120, 5632, 6144, 6656, 7168, 7680, 8192], "tc_compiler": "eager", "full_prefill_max_req": null}}, "cuda_graph_backend_decode": null, "cuda_graph_backend_prefill": null, "cuda_graph_max_bs_decode": 64, "cuda_graph_max_bs_prefill": null, "cuda_graph_bs_decode": null, "cuda_graph_bs_prefill": null, "cuda_graph_tc_compiler": null, "disable_prefill_cuda_graph": false, "disable_decode_cuda_graph": false, "disable_cuda_graph": false, "disable_cuda_graph_padding": false, "enable_profile_cuda_graph": false, "enable_cudagraph_gc": false, "debug_cuda_graph": false, "enable_layerwise_nvtx_marker": false, "enable_nccl_nvls": false, "enable_symm_mem": false, "triton_attention_reduce_in_fp32": false, "triton_attention_num_kv_splits": 8, "triton_attention_split_tile_size": null, "flashinfer_mla_disable_ragged": false, "enable_fused_qk_norm_rope": false, "enable_precise_embedding_interpolation": false, "enable_fused_moe_sum_all_reduce": false, "enable_deepseek_v4_fp4_indexer": false, "disable_custom_all_reduce": false, "enable_mscclpp": false, "enable_torch_symm_mem": false, "pre_warm_nccl": false, "enable_quant_communications": false, "enable_flashinfer_allreduce_fusion": false, "enforce_disable_flashinfer_allreduce_fusion": false, "flashinfer_allreduce_fusion_backend": null, "enable_aiter_allreduce_fusion": false, "enable_torch_compile": false, "enable_torch_compile_debug_mode": false, "torch_compile_max_bs": 32, "torchao_config": "", "speculative_algorithm": null, "speculative_draft_model_path": null, "speculative_draft_model_revision": null, "speculative_draft_load_format": null, "speculative_num_steps": null, "speculative_eagle_topk": null, "speculative_num_draft_tokens": null, "speculative_dflash_block_size": null, "speculative_dspark_block_size": null, "speculative_dspark_sps_table_path": null, "speculative_dspark_confidence_sts_path": null, "speculative_dspark_align_verify_tokens_to_graph_tier": false, "speculative_accept_threshold_single": 1.0, "speculative_accept_threshold_acc": 1.0, "speculative_use_rejection_sampling": false, "speculative_token_map": null, "speculative_attention_mode": "prefill", "speculative_draft_attention_backend": null, "speculative_draft_window_size": null, "speculative_moe_runner_backend": "flashinfer_mxfp4", "speculative_moe_a2a_backend": null, "speculative_draft_model_quantization": null, "speculative_skip_dp_mlp_sync": false, "enable_multi_layer_eagle": false, "speculative_adaptive": false, "speculative_adaptive_config": null, "decoupled_spec_bind_endpoint": null, "decoupled_spec_connect_endpoints": null, "decoupled_spec_rank": null, "decoupled_spec_role": "null", "spec_trace_dir": null, "speculative_ngram_min_bfs_breadth": 1, "speculative_ngram_max_bfs_breadth": 10, "speculative_ngram_match_type": "BFS", "speculative_ngram_max_trie_depth": 18, "speculative_ngram_capacity": 10000000, "speculative_ngram_external_corpus_path": null, "speculative_ngram_external_sam_budget": 0, "speculative_ngram_external_corpus_max_tokens": 10000000, "ep_size": 2, "moe_a2a_backend": "none", "moe_runner_backend": "flashinfer_mxfp4", "flashinfer_mxfp4_moe_precision": "default", "deepep_mode": "auto", "fuseep_mode": 2, "deepep_dispatcher_output_dtype": "auto", "ep_num_redundant_experts": 0, "ep_dispatch_algorithm": null, "init_expert_location": "trivial", "enable_eplb": false, "eplb_algorithm": "auto", "eplb_rebalance_num_iterations": 1000, "eplb_rebalance_layers_per_chunk": null, "eplb_min_rebalancing_utilization_threshold": 1.0, "expert_distribution_recorder_mode": null, "expert_distribution_recorder_buffer_size": 1000, "enable_expert_distribution_metrics": false, "deepep_config": null, "moe_dense_tp_size": null, "elastic_ep_backend": null, "enable_elastic_expert_backup": false, "mooncake_ib_device": null, "enable_waterfill": false, "ep_join_mode": null, "ep_join_rank_offset": 0, "elastic_ep_initial_size": null, "max_ep_size": null, "elastic_ep_scale_timeout": 600, "elastic_ep_rejoin": false, "disable_flashinfer_cutlass_moe_fp4_allgather": false, "disable_shared_experts_fusion": true, "enforce_shared_experts_fusion": false, "max_mamba_cache_size": null, "mamba_ssm_dtype": null, "enable_mamba_cache_stochastic_rounding": false, "mamba_cache_philox_rounds": 0, "mamba_full_memory_ratio": 0.9, "mamba_radix_cache_strategy": "auto", "uses_mamba_radix_cache": false, "mamba_track_interval": 256, "enable_int8_mamba_checkpoint": false, "int8_mamba_ckpt_size": null, "linear_attn_backend": "triton", "linear_attn_decode_backend": null, "linear_attn_prefill_backend": null, "enable_linear_replayssm": false, "linear_replayssm_cache_len": 16, "enable_hierarchical_cache": false, "hicache_ratio": 2.0, "hicache_size": 0, "hicache_write_policy": "write_through", "hicache_io_backend": "kernel", "hicache_mem_layout": "page_first", "hicache_storage_backend": null, "hicache_storage_prefetch_policy": "timeout", "hicache_storage_backend_extra_config": null, "enable_hisparse": false, "hisparse_config": null, "enable_broadcast_mm_inputs_process": false, "enable_prefix_mm_cache": false, "mm_enable_dp_encoder": false, "mm_process_config": {}, "limit_mm_data_per_request": null, "enable_mm_global_cache": false, "disable_fast_image_processor": false, "mm_feature_transport": "cpu", "keep_mm_feature_on_device": false, "enable_lora": null, "enable_lora_overlap_loading": null, "max_lora_rank": null, "lora_target_modules": null, "lora_paths": null, "max_loaded_loras": null, "max_loras_per_batch": 8, "lora_eviction_policy": "lru", "lora_backend": "csgmv", "max_lora_chunk_size": 16, "experts_shared_outer_loras": null, "lora_use_virtual_experts": false, "lora_strict_loading": false, "lora_drain_wait_threshold": 0.0, "enable_two_batch_overlap": false, "enable_single_batch_overlap": false, "tbo_token_distribution_threshold": 0.48, "cpu_offload_gb": 0, "offload_group_size": -1, "offload_num_in_group": 1, "offload_prefetch_step": 1, "offload_mode": "cpu", "enable_lmcache": false, "lmcache_config_file": null, "enable_flexkv": false, "flexkv_config_file": null, "kt_weight_path": null, "kt_method": "AMXINT4", "kt_cpuinfer": null, "kt_threadpool_count": 2, "kt_num_gpu_experts": null, "kt_max_deferred_experts_per_token": null, "dllm_algorithm": null, "dllm_algorithm_config": null, "dllm_fdfo": true, "disaggregation_mode": "null", "disaggregation_transfer_backend": "mooncake", "disaggregation_bootstrap_port": 8998, "disaggregation_ib_device": null, "disaggregation_decode_enable_radix_cache": false, "disaggregation_decode_enable_offload_kvcache": false, "num_reserved_decode_tokens": 512, "disaggregation_decode_extra_slots": null, "disaggregation_decode_polling_interval": 1, "optimistic_prefill_attempts": 0, "encoder_only": false, "language_only": false, "encoder_transfer_backend": "zmq_to_scheduler", "encoder_urls": [], "encoder_bootstrap_port": 8997, "encoder_register_urls": [], "enable_adaptive_dispatch_to_encoder": false, "enable_pdmux": false, "pdmux_config_path": null, "sm_group_num": 8, "custom_weight_loader": [], "weight_loader_disable_mmap": false, "weight_loader_prefetch_checkpoints": false, "weight_loader_prefetch_num_threads": 4, "weight_loader_drop_cache_after_load": false, "remote_instance_weight_loader_seed_instance_ip": null, "remote_instance_weight_loader_seed_instance_service_port": null, "remote_instance_weight_loader_send_weights_group_ports": null, "remote_instance_weight_loader_backend": "nccl", "remote_instance_weight_loader_start_seed_via_transfer_engine": false, "engine_info_bootstrap_port": 6789, "modelexpress_config": null, "download_dir": null, "model_checksum": null, "delete_ckpt_after_loading": false, "decrypted_config_file": null, "decrypted_draft_config_file": null, "checkpoint_engine_wait_weights_before_ready": false, "enable_prefill_delayer": false, "prefill_delayer_max_delay_passes": 30, "prefill_delayer_token_usage_low_watermark": null, "prefill_delayer_forward_passes_buckets": null, "prefill_delayer_wait_seconds_buckets": null, "prefill_delayer_queue_min_ratio": null, "prefill_delayer_max_delay_ms": null, "min_free_slots_delay": null, "enable_deterministic_inference": false, "rl_on_policy_target": null, "kv_canary": "none", "kv_canary_real_data": "none", "kv_canary_sweep_interval": 0, "enable_dynamic_batch_tokenizer": false, "dynamic_batch_tokenizer_batch_size": 32, "dynamic_batch_tokenizer_batch_timeout": 0.002, "enable_tokenizer_batch_encode": false, "disable_tokenizer_batch_decode": false, "debug_tensor_dump_output_folder": null, "debug_tensor_dump_layers": null, "debug_tensor_dump_input_file": null, "enable_memory_saver": false, "enable_weights_cpu_backup": false, "enable_draft_weights_cpu_backup": false, "enable_custom_logit_processor": false, "enable_return_hidden_states": false, "enable_return_routed_experts": false, "enable_return_indexer_topk": false, "disable_outlines_disk_cache": false, "enable_mis": false, "forward_hooks": null, "msprobe_dump_config": null, "_resolved_overrides": [["_deepseek_v4_overrides", {"attention_backend": "dsv4", "page_size": 256, "swa_full_tokens_ratio": 0.1}], ["_deepseek_v4_kv_cache_dtype", {"kv_cache_dtype": "fp8_e4m3"}], ["_deepseek_v4_sm120_moe", {"moe_runner_backend": "flashinfer_mxfp4"}], ["_sampling_backend_default", {"sampling_backend": "flashinfer"}], ["_data_parallelism_defaults", {"enable_dp_attention": false, "enable_dp_lm_head": false}], ["_speculative_moe_runner_default", {"speculative_moe_runner_backend": "flashinfer_mxfp4"}], ["DeepseekV4ForCausalLM.determine_num_fused_shared_experts", {"disable_shared_experts_fusion": true}]], "grpc_worker_threads": 4, "_quantization_explicitly_unset": false, "_cuda_graph_config_locked": [["decode", "max_bs"]], "_declarations_materialized": true, "_runtime_mutations": [["model_runner.chunked_prefix_cache_gate", {"disable_chunked_prefix_cache": true}], ["scheduler.pp_max_micro_batch_size_default", {"pp_max_micro_batch_size": 256}]], "_in_override": false, "_mx_config_cache": {}, "max_speculative_num_draft_tokens": null, "last_gen_throughput": 9.607872017008752, "memory_usage": {"weight": 56.87, "kvcache": 0, "token_capacity": 1282304, "graph": 1.39}, "effective_max_running_requests_per_dp": 256}], "version": "0.0.0.dev1+g35f2d4f76", "kv_events": null}, "duration": 28.799581017025048, "completed": 1, "total_input_tokens": 1024, "total_input_text_tokens": 1024, "total_input_vision_tokens": 0, "total_output_tokens": 128, "total_output_tokens_retokenized": 128, "request_throughput": 0.03472272736915318, "input_throughput": 35.55607282601286, "output_throughput": 4.444509103251607, "total_throughput": 40.00058192926446, "mean_e2e_latency_ms": 28691.018268000335, "median_e2e_latency_ms": 28691.018268000335, "std_e2e_latency_ms": 0.0, "p90_e2e_latency_ms": 28691.018268000335, "p95_e2e_latency_ms": 28691.018268000335, "p99_e2e_latency_ms": 28691.018268000335, "mean_ttft_ms": 15788.179125986062, "median_ttft_ms": 15788.179125986062, "std_ttft_ms": 0.0, "p90_ttft_ms": 15788.179125986062, "p95_ttft_ms": 15788.179125986062, "p99_ttft_ms": 15788.179125986062, "mean_tpot_ms": 101.59715859853758, "median_tpot_ms": 101.59715859853758, "std_tpot_ms": 0.0, "p90_tpot_ms": 101.59715859853758, "p95_tpot_ms": 101.59715859853758, "p99_tpot_ms": 101.59715859853758, "mean_itl_ms": 101.5968882363275, "median_itl_ms": 104.02098900522105, "std_itl_ms": 3.7676533587274976, "p90_itl_ms": 104.27935040788725, "p95_itl_ms": 104.53261959773954, "p99_itl_ms": 104.80170075898059, "concurrency": 0.9962304052631691, "accept_length": null, "max_output_tokens_per_s": 11.0, "max_concurrent_requests": 1, "input_lens": [1024], "output_lens": [128], "ttfts": [15.788179125986062], "itls": [[0.09673743901657872, 0.09605362298316322, 0.0979746550146956, 0.09797314400202595, 0.09803740799543448, 0.0980707309790887, 0.09813011001097038, 0.09801292201154865, 0.09891589698963799, 0.09805917501216754, 0.09794300797511823, 0.09811295801773667, 0.09813642498920672, 0.09809407501597889, 0.09789085100055672, 0.09791729997959919, 0.09773818802204914, 0.09870099398540333, 0.09813020899309777, 0.09796486201230437, 0.09798214898910373, 0.09808215202065185, 0.09794546099146828, 0.09799491599551402, 0.09810031999950297, 0.09817106000264175, 0.09851828799583018, 0.09809130500070751, 0.09857681000721641, 0.10384371198597364, 0.09945607100962661, 0.0987883560010232, 0.09773963500629179, 0.09695546599687077, 0.09428766599739902, 0.09517847601091489, 0.09786818397697061, 0.09797205499489792, 0.09767439001007006, 0.09724350800388493, 0.09862652499577962, 0.09801103800418787, 0.09009501698892564, 0.08921656099846587, 0.09140137501526624, 0.08526493000681512, 0.09914104000199586, 0.10438463898026384, 0.10450987800140865, 0.10454236599616706, 0.1046097380167339, 0.10474240899202414, 0.10460548399714753, 0.10468307099654339, 0.10419266001554206, 0.1041276799805928, 0.10406649200012907, 0.10413092400995083, 0.10394618898862973, 0.1039625370176509, 0.10404802099219523, 0.10415349699906074, 0.10396736601251177, 0.10405155399348587, 0.10409425300895236, 0.10404276198823936, 0.10402258299291134, 0.10417768699699081, 0.10398562002228573, 0.10415878999629058, 0.1039193779870402, 0.10407710800063796, 0.10395074400003068, 0.10405435101711191, 0.10404936398845166, 0.10415259699220769, 0.10414050801773556, 0.1041957029956393, 0.10406920398236252, 0.10415436601033434, 0.10482253300142474, 0.10385464099817909, 0.10427678801352158, 0.10428319399943575, 0.10382588399806991, 0.10409237098065205, 0.104062149010133, 0.1041014000074938, 0.10406459399382584, 0.1039629029983189, 0.10415040500811301, 0.10418893798487261, 0.10414424500777386, 0.10404462300357409, 0.10406634700484574, 0.10400176999974065, 0.10403723700437695, 0.104061030986486, 0.1040539120149333, 0.10403626499464735, 0.10395490698283538, 0.10402518199407496, 0.1039650880265981, 0.10407873598160222, 0.10397985199233517, 0.1040512680192478, 0.10408997000195086, 0.1041235429875087, 0.10406186000909656, 0.1041025279846508, 0.10437087502214126, 0.10412507900036871, 0.10407520798617043, 0.10402098900522105, 0.10394267400261015, 0.10414250599569641, 0.10407390299951658, 0.10398448599153198, 0.10390737000852823, 0.10403330798726529, 0.10411680801189505, 0.10409107498708181, 0.10416924202581868, 0.10439796198625118, 0.1043221470026765, 0.10405737700057216, 0.10561967000830919]], "generated_texts": [".\nMany find themselves to be more profitable by just finding out where the dollars are escaping in their business and I like to think of myself as a guy that comes along with some spakel or putty and patch those holes up for you.\nBeleive me, just fixing one hole can mean a lot...just think about a sinking boat that has a hole in it that's about 3\u201d in diameter... it doesn't take long to sink.\nI have no agenda, besides f=getting to know your business and seeing wher I can patch the holes and find what makes you do darn unique (I know this won't"], "errors": [""]} diff --git a/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/cold_1k_o1_first.jsonl b/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/cold_1k_o1_first.jsonl new file mode 100644 index 0000000..9294935 --- /dev/null +++ b/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/cold_1k_o1_first.jsonl @@ -0,0 +1 @@ +{"tag": null, "backend": "sglang", "dataset_name": "random", "request_rate": 10000.0, "max_concurrency": 1, "sharegpt_output_len": null, "random_input_len": 1024, "random_output_len": 1, "random_range_ratio": 1.0, "server_info": {"model_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_mode": "auto", "tokenizer_backend": "huggingface", "tokenizer_worker_num": 1, "detokenizer_worker_num": 1, "skip_tokenizer_init": false, "load_format": "auto", "model_loader_extra_config": "{}", "trust_remote_code": true, "context_length": null, "is_embedding": false, "enable_multimodal": null, "revision": null, "model_impl": "auto", "model_config_parser": "auto", "json_model_override_args": "{}", "dtype": "auto", "quantization": null, "quantization_param_path": null, "kv_cache_dtype": "fp8_e4m3", "enable_fp32_lm_head": false, "modelopt_quant": null, "modelopt_checkpoint_restore_path": null, "modelopt_checkpoint_save_path": null, "modelopt_export_path": null, "quantize_and_serve": false, "rl_quant_profile": null, "enable_tf32_matmul": false, "mem_fraction_static": 0.9, "max_running_requests": 256, "max_queued_requests": null, "max_total_tokens": null, "chunked_prefill_size": 8192, "enable_dynamic_chunking": false, "max_prefill_tokens": 16384, "prefill_max_requests": null, "schedule_policy": "fcfs", "enable_priority_scheduling": false, "disable_priority_preemption": false, "default_priority_value": null, "abort_on_priority_when_disabled": false, "schedule_low_priority_values_first": false, "priority_scheduling_preemption_threshold": 10, "retraction_policy": "length", "schedule_conservativeness": 1.0, "page_size": 256, "swa_full_tokens_ratio": 0.1, "disable_hybrid_swa_memory": false, "radix_eviction_policy": "lru", "prefill_only_disable_kv_cache": false, "disable_radix_cache": false, "enable_page_major_kv_layout": false, "enable_unified_memory": false, "disable_chunked_prefix_cache": false, "disable_overlap_schedule": false, "num_continuous_decode_steps": 1, "scheduler_recv_interval": 1, "enable_mixed_chunk": false, "nccl_port": null, "dist_timeout": null, "dist_init_addr": "10.101.0.11:20002", "nnodes": 2, "node_rank": 0, "tp_size": 16, "dcp_size": 1, "pp_size": 1, "pp_max_micro_batch_size": null, "pp_async_batch_depth": 0, "dp_size": 1, "load_balance_method": "round_robin", "attn_cp_size": 1, "moe_dp_size": 1, "enable_prefill_cp": false, "cp_strategy": null, "enable_dsa_cache_layer_split": false, "enable_dsa_prefill_context_parallel": false, "dsa_prefill_cp_mode": "round-robin-split", "enable_prefill_context_parallel": false, "prefill_cp_mode": "in-seq-split", "enable_dp_attention": false, "enable_dp_attention_local_control_broadcast": false, "enable_dp_lm_head": false, "enable_attn_tp_input_scattered": false, "disable_attn_tp_gather": false, "enable_p2p_check": false, "device": "cuda", "base_gpu_id": 0, "gpu_id_step": 1, "random_seed": 42600328, "watchdog_timeout": 300, "soft_watchdog_timeout": null, "sleep_on_idle": false, "use_ray": false, "custom_sigquit_handler": null, "numa_node": null, "gc_threshold": null, "host": "0.0.0.0", "port": 30002, "fastapi_root_path": "", "smg_grpc_mode": false, "grpc_mode": false, "grpc_port": null, "skip_server_warmup": false, "warmups": null, "enable_http2": false, "ssl_keyfile": null, "ssl_certfile": null, "ssl_ca_certs": null, "ssl_keyfile_password": null, "enable_ssl_refresh": false, "api_key": null, "admin_api_key": null, "served_model_name": "/data/hf_models/DeepSeek-V4-Pro", "weight_version": "default", "chat_template": null, "hf_chat_template_name": null, "completion_template": null, "file_storage_path": "sglang_storage", "enable_cache_report": false, "reasoning_parser": null, "default_chat_template_kwargs": null, "strip_thinking_cache": false, "enable_strict_thinking": false, "tool_call_parser": null, "tool_server": null, "sampling_defaults": "model", "asr_max_buffer_seconds": 60, "asr_max_concurrent_sessions": 32, "preferred_sampling_params": null, "allow_auto_truncate": false, "stream_interval": 1, "batch_notify_size": 16, "stream_response_default_include_usage": false, "incremental_streaming_output": false, "enable_streaming_session": false, "enable_session_radix_cache": false, "log_level": "info", "log_level_http": null, "log_requests": false, "log_requests_level": 2, "log_requests_format": "text", "log_requests_target": null, "uvicorn_access_log_exclude_prefixes": [], "crash_dump_folder": null, "show_time_cost": false, "enable_metrics": false, "smg_http_sidecar_port": null, "enable_mfu_metrics": false, "enable_metrics_for_all_schedulers": false, "load_snapshot_publish_interval": 15, "tokenizer_metrics_custom_labels_header": "x-custom-labels", "tokenizer_metrics_allowed_custom_labels": null, "extra_metric_labels": null, "bucket_time_to_first_token": null, "bucket_inter_token_latency": null, "bucket_e2e_request_latency": null, "prompt_tokens_buckets": null, "generation_tokens_buckets": null, "gc_warning_threshold_secs": 0.0, "decode_log_interval": 40, "enable_request_time_stats_logging": false, "kv_events_config": null, "enable_forward_pass_metrics": false, "forward_pass_metrics_worker_id": "", "forward_pass_metrics_ipc_name": null, "enable_trace": false, "trace_modules": "request", "otlp_traces_endpoint": "localhost:4317", "export_metrics_to_file": false, "export_metrics_to_file_dir": null, "stat_loggers": null, "constrained_json_whitespace_pattern": null, "constrained_json_disable_any_whitespace": false, "attention_backend": "dsv4", "decode_attention_backend": null, "prefill_attention_backend": null, "sampling_backend": "flashinfer", "grammar_backend": "xgrammar", "radix_cache_backend": null, "mm_attention_backend": null, "fp8_gemm_runner_backend": "auto", "fp4_gemm_runner_backend": "auto", "bf16_gemm_backend": "auto", "dsa_prefill_backend": null, "dsa_decode_backend": null, "dsa_paged_mqa_logits_backend": "auto", "dsa_topk_backend": "sgl-kernel", "disable_flashinfer_autotune": false, "mamba_backend": "triton", "cuda_graph_config": {"decode": {"backend": "full", "max_bs": 64, "bs": [1, 2, 4, 8, 12, 16, 24, 32, 40, 48, 56, 64], "tc_compiler": "eager", "full_prefill_max_req": null}, "prefill": {"backend": "disabled", "max_bs": 8192, "bs": [4, 8, 12, 16, 20, 24, 28, 32, 48, 64, 80, 96, 112, 128, 144, 160, 176, 192, 208, 224, 240, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 704, 768, 832, 896, 960, 1024, 1280, 1536, 1792, 2048, 2304, 2560, 2816, 3072, 3328, 3584, 3840, 4096, 4608, 5120, 5632, 6144, 6656, 7168, 7680, 8192], "tc_compiler": "eager", "full_prefill_max_req": null}}, "cuda_graph_backend_decode": null, "cuda_graph_backend_prefill": null, "cuda_graph_max_bs_decode": 64, "cuda_graph_max_bs_prefill": null, "cuda_graph_bs_decode": null, "cuda_graph_bs_prefill": null, "cuda_graph_tc_compiler": null, "disable_prefill_cuda_graph": false, "disable_decode_cuda_graph": false, "disable_cuda_graph": false, "disable_cuda_graph_padding": false, "enable_profile_cuda_graph": false, "enable_cudagraph_gc": false, "debug_cuda_graph": false, "enable_layerwise_nvtx_marker": false, "enable_nccl_nvls": false, "enable_symm_mem": false, "triton_attention_reduce_in_fp32": false, "triton_attention_num_kv_splits": 8, "triton_attention_split_tile_size": null, "flashinfer_mla_disable_ragged": false, "enable_fused_qk_norm_rope": false, "enable_precise_embedding_interpolation": false, "enable_fused_moe_sum_all_reduce": false, "enable_deepseek_v4_fp4_indexer": false, "disable_custom_all_reduce": false, "enable_mscclpp": false, "enable_torch_symm_mem": false, "pre_warm_nccl": false, "enable_quant_communications": false, "enable_flashinfer_allreduce_fusion": false, "enforce_disable_flashinfer_allreduce_fusion": false, "flashinfer_allreduce_fusion_backend": null, "enable_aiter_allreduce_fusion": false, "enable_torch_compile": false, "enable_torch_compile_debug_mode": false, "torch_compile_max_bs": 32, "torchao_config": "", "speculative_algorithm": null, "speculative_draft_model_path": null, "speculative_draft_model_revision": null, "speculative_draft_load_format": null, "speculative_num_steps": null, "speculative_eagle_topk": null, "speculative_num_draft_tokens": null, "speculative_dflash_block_size": null, "speculative_dspark_block_size": null, "speculative_dspark_sps_table_path": null, "speculative_dspark_confidence_sts_path": null, "speculative_dspark_align_verify_tokens_to_graph_tier": false, "speculative_accept_threshold_single": 1.0, "speculative_accept_threshold_acc": 1.0, "speculative_use_rejection_sampling": false, "speculative_token_map": null, "speculative_attention_mode": "prefill", "speculative_draft_attention_backend": null, "speculative_draft_window_size": null, "speculative_moe_runner_backend": "flashinfer_mxfp4", "speculative_moe_a2a_backend": null, "speculative_draft_model_quantization": null, "speculative_skip_dp_mlp_sync": false, "enable_multi_layer_eagle": false, "speculative_adaptive": false, "speculative_adaptive_config": null, "decoupled_spec_bind_endpoint": null, "decoupled_spec_connect_endpoints": null, "decoupled_spec_rank": null, "decoupled_spec_role": "null", "spec_trace_dir": null, "speculative_ngram_min_bfs_breadth": 1, "speculative_ngram_max_bfs_breadth": 10, "speculative_ngram_match_type": "BFS", "speculative_ngram_max_trie_depth": 18, "speculative_ngram_capacity": 10000000, "speculative_ngram_external_corpus_path": null, "speculative_ngram_external_sam_budget": 0, "speculative_ngram_external_corpus_max_tokens": 10000000, "ep_size": 2, "moe_a2a_backend": "none", "moe_runner_backend": "flashinfer_mxfp4", "flashinfer_mxfp4_moe_precision": "default", "deepep_mode": "auto", "fuseep_mode": 2, "deepep_dispatcher_output_dtype": "auto", "ep_num_redundant_experts": 0, "ep_dispatch_algorithm": null, "init_expert_location": "trivial", "enable_eplb": false, "eplb_algorithm": "auto", "eplb_rebalance_num_iterations": 1000, "eplb_rebalance_layers_per_chunk": null, "eplb_min_rebalancing_utilization_threshold": 1.0, "expert_distribution_recorder_mode": null, "expert_distribution_recorder_buffer_size": 1000, "enable_expert_distribution_metrics": false, "deepep_config": null, "moe_dense_tp_size": null, "elastic_ep_backend": null, "enable_elastic_expert_backup": false, "mooncake_ib_device": null, "enable_waterfill": false, "ep_join_mode": null, "ep_join_rank_offset": 0, "elastic_ep_initial_size": null, "max_ep_size": null, "elastic_ep_scale_timeout": 600, "elastic_ep_rejoin": false, "disable_flashinfer_cutlass_moe_fp4_allgather": false, "disable_shared_experts_fusion": false, "enforce_shared_experts_fusion": false, "max_mamba_cache_size": null, "mamba_ssm_dtype": null, "enable_mamba_cache_stochastic_rounding": false, "mamba_cache_philox_rounds": 0, "mamba_full_memory_ratio": 0.9, "mamba_radix_cache_strategy": "auto", "uses_mamba_radix_cache": false, "mamba_track_interval": 256, "enable_int8_mamba_checkpoint": false, "int8_mamba_ckpt_size": null, "linear_attn_backend": "triton", "linear_attn_decode_backend": null, "linear_attn_prefill_backend": null, "enable_linear_replayssm": false, "linear_replayssm_cache_len": 16, "enable_hierarchical_cache": false, "hicache_ratio": 2.0, "hicache_size": 0, "hicache_write_policy": "write_through", "hicache_io_backend": "kernel", "hicache_mem_layout": "page_first", "hicache_storage_backend": null, "hicache_storage_prefetch_policy": "timeout", "hicache_storage_backend_extra_config": null, "enable_hisparse": false, "hisparse_config": null, "enable_broadcast_mm_inputs_process": false, "enable_prefix_mm_cache": false, "mm_enable_dp_encoder": false, "mm_process_config": {}, "limit_mm_data_per_request": null, "enable_mm_global_cache": false, "disable_fast_image_processor": false, "mm_feature_transport": "cpu", "keep_mm_feature_on_device": false, "enable_lora": null, "enable_lora_overlap_loading": null, "max_lora_rank": null, "lora_target_modules": null, "lora_paths": null, "max_loaded_loras": null, "max_loras_per_batch": 8, "lora_eviction_policy": "lru", "lora_backend": "csgmv", "max_lora_chunk_size": 16, "experts_shared_outer_loras": null, "lora_use_virtual_experts": false, "lora_strict_loading": false, "lora_drain_wait_threshold": 0.0, "enable_two_batch_overlap": false, "enable_single_batch_overlap": false, "tbo_token_distribution_threshold": 0.48, "cpu_offload_gb": 0, "offload_group_size": -1, "offload_num_in_group": 1, "offload_prefetch_step": 1, "offload_mode": "cpu", "enable_lmcache": false, "lmcache_config_file": null, "enable_flexkv": false, "flexkv_config_file": null, "kt_weight_path": null, "kt_method": "AMXINT4", "kt_cpuinfer": null, "kt_threadpool_count": 2, "kt_num_gpu_experts": null, "kt_max_deferred_experts_per_token": null, "dllm_algorithm": null, "dllm_algorithm_config": null, "dllm_fdfo": true, "disaggregation_mode": "null", "disaggregation_transfer_backend": "mooncake", "disaggregation_bootstrap_port": 8998, "disaggregation_ib_device": null, "disaggregation_decode_enable_radix_cache": false, "disaggregation_decode_enable_offload_kvcache": false, "num_reserved_decode_tokens": 512, "disaggregation_decode_extra_slots": null, "disaggregation_decode_polling_interval": 1, "optimistic_prefill_attempts": 0, "encoder_only": false, "language_only": false, "encoder_transfer_backend": "zmq_to_scheduler", "encoder_urls": [], "encoder_bootstrap_port": 8997, "encoder_register_urls": [], "enable_adaptive_dispatch_to_encoder": false, "enable_pdmux": false, "pdmux_config_path": null, "sm_group_num": 8, "custom_weight_loader": [], "weight_loader_disable_mmap": false, "weight_loader_prefetch_checkpoints": false, "weight_loader_prefetch_num_threads": 4, "weight_loader_drop_cache_after_load": false, "remote_instance_weight_loader_seed_instance_ip": null, "remote_instance_weight_loader_seed_instance_service_port": null, "remote_instance_weight_loader_send_weights_group_ports": null, "remote_instance_weight_loader_backend": "nccl", "remote_instance_weight_loader_start_seed_via_transfer_engine": false, "engine_info_bootstrap_port": 6789, "modelexpress_config": null, "download_dir": null, "model_checksum": null, "delete_ckpt_after_loading": false, "decrypted_config_file": null, "decrypted_draft_config_file": null, "checkpoint_engine_wait_weights_before_ready": false, "enable_prefill_delayer": false, "prefill_delayer_max_delay_passes": 30, "prefill_delayer_token_usage_low_watermark": null, "prefill_delayer_forward_passes_buckets": null, "prefill_delayer_wait_seconds_buckets": null, "prefill_delayer_queue_min_ratio": null, "prefill_delayer_max_delay_ms": null, "min_free_slots_delay": null, "enable_deterministic_inference": false, "rl_on_policy_target": null, "kv_canary": "none", "kv_canary_real_data": "none", "kv_canary_sweep_interval": 0, "enable_dynamic_batch_tokenizer": false, "dynamic_batch_tokenizer_batch_size": 32, "dynamic_batch_tokenizer_batch_timeout": 0.002, "enable_tokenizer_batch_encode": false, "disable_tokenizer_batch_decode": false, "debug_tensor_dump_output_folder": null, "debug_tensor_dump_layers": null, "debug_tensor_dump_input_file": null, "enable_memory_saver": false, "enable_weights_cpu_backup": false, "enable_draft_weights_cpu_backup": false, "enable_custom_logit_processor": false, "enable_return_hidden_states": false, "enable_return_routed_experts": false, "enable_return_indexer_topk": false, "disable_outlines_disk_cache": false, "enable_mis": false, "forward_hooks": null, "msprobe_dump_config": null, "status": "ready", "max_total_num_tokens": 1282304, "max_req_input_len": 1048570, "internal_states": [{"model_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_mode": "auto", "tokenizer_backend": "huggingface", "tokenizer_worker_num": 1, "detokenizer_worker_num": 1, "skip_tokenizer_init": false, "load_format": "auto", "model_loader_extra_config": "{}", "trust_remote_code": true, "context_length": null, "is_embedding": false, "enable_multimodal": null, "revision": null, "model_impl": "auto", "model_config_parser": "auto", "json_model_override_args": "{}", "dtype": "auto", "quantization": null, "quantization_param_path": null, "kv_cache_dtype": "fp8_e4m3", "enable_fp32_lm_head": false, "modelopt_quant": null, "modelopt_checkpoint_restore_path": null, "modelopt_checkpoint_save_path": null, "modelopt_export_path": null, "quantize_and_serve": false, "rl_quant_profile": null, "enable_tf32_matmul": false, "mem_fraction_static": 0.9, "max_running_requests": 256, "max_queued_requests": null, "max_total_tokens": null, "chunked_prefill_size": 8192, "enable_dynamic_chunking": false, "max_prefill_tokens": 16384, "prefill_max_requests": null, "schedule_policy": "fcfs", "enable_priority_scheduling": false, "disable_priority_preemption": false, "default_priority_value": null, "abort_on_priority_when_disabled": false, "schedule_low_priority_values_first": false, "priority_scheduling_preemption_threshold": 10, "retraction_policy": "length", "schedule_conservativeness": 1.0, "page_size": 256, "swa_full_tokens_ratio": 0.1, "disable_hybrid_swa_memory": false, "radix_eviction_policy": "lru", "prefill_only_disable_kv_cache": false, "disable_radix_cache": false, "enable_page_major_kv_layout": false, "enable_unified_memory": false, "disable_chunked_prefix_cache": true, "disable_overlap_schedule": false, "num_continuous_decode_steps": 1, "scheduler_recv_interval": 1, "enable_mixed_chunk": false, "nccl_port": null, "dist_timeout": null, "dist_init_addr": "10.101.0.11:20002", "nnodes": 2, "node_rank": 0, "tp_size": 16, "dcp_size": 1, "pp_size": 1, "pp_max_micro_batch_size": 256, "pp_async_batch_depth": 0, "dp_size": 1, "load_balance_method": "round_robin", "attn_cp_size": 1, "moe_dp_size": 1, "enable_prefill_cp": false, "cp_strategy": null, "enable_dsa_cache_layer_split": false, "enable_dsa_prefill_context_parallel": false, "dsa_prefill_cp_mode": "round-robin-split", "enable_prefill_context_parallel": false, "prefill_cp_mode": "in-seq-split", "enable_dp_attention": false, "enable_dp_attention_local_control_broadcast": false, "enable_dp_lm_head": false, "enable_attn_tp_input_scattered": false, "disable_attn_tp_gather": false, "enable_p2p_check": false, "device": "cuda", "base_gpu_id": 0, "gpu_id_step": 1, "random_seed": 42600328, "watchdog_timeout": 300, "soft_watchdog_timeout": null, "sleep_on_idle": false, "use_ray": false, "numa_node": null, "gc_threshold": null, "host": "0.0.0.0", "port": 30002, "fastapi_root_path": "", "smg_grpc_mode": false, "grpc_mode": false, "grpc_port": null, "skip_server_warmup": false, "warmups": null, "enable_http2": false, "ssl_keyfile": null, "ssl_certfile": null, "ssl_ca_certs": null, "ssl_keyfile_password": null, "enable_ssl_refresh": false, "api_key": null, "admin_api_key": null, "served_model_name": "/data/hf_models/DeepSeek-V4-Pro", "weight_version": "default", "chat_template": null, "hf_chat_template_name": null, "completion_template": null, "file_storage_path": "sglang_storage", "enable_cache_report": false, "reasoning_parser": null, "default_chat_template_kwargs": null, "strip_thinking_cache": false, "enable_strict_thinking": false, "tool_call_parser": null, "tool_server": null, "sampling_defaults": "model", "asr_max_buffer_seconds": 60, "asr_max_concurrent_sessions": 32, "preferred_sampling_params": null, "allow_auto_truncate": false, "stream_interval": 1, "batch_notify_size": 16, "stream_response_default_include_usage": false, "incremental_streaming_output": false, "enable_streaming_session": false, "enable_session_radix_cache": false, "log_level": "info", "log_level_http": null, "log_requests": false, "log_requests_level": 2, "log_requests_format": "text", "log_requests_target": null, "uvicorn_access_log_exclude_prefixes": [], "crash_dump_folder": null, "show_time_cost": false, "enable_metrics": false, "smg_http_sidecar_port": null, "enable_mfu_metrics": false, "enable_metrics_for_all_schedulers": false, "load_snapshot_publish_interval": 15, "tokenizer_metrics_custom_labels_header": "x-custom-labels", "tokenizer_metrics_allowed_custom_labels": null, "extra_metric_labels": null, "bucket_time_to_first_token": null, "bucket_inter_token_latency": null, "bucket_e2e_request_latency": null, "prompt_tokens_buckets": null, "generation_tokens_buckets": null, "gc_warning_threshold_secs": 0.0, "decode_log_interval": 40, "enable_request_time_stats_logging": false, "kv_events_config": null, "enable_forward_pass_metrics": false, "forward_pass_metrics_worker_id": "", "forward_pass_metrics_ipc_name": null, "enable_trace": false, "trace_modules": "request", "otlp_traces_endpoint": "localhost:4317", "export_metrics_to_file": false, "export_metrics_to_file_dir": null, "stat_loggers": null, "constrained_json_whitespace_pattern": null, "constrained_json_disable_any_whitespace": false, "attention_backend": "dsv4", "decode_attention_backend": null, "prefill_attention_backend": null, "sampling_backend": "flashinfer", "grammar_backend": "xgrammar", "radix_cache_backend": null, "mm_attention_backend": null, "fp8_gemm_runner_backend": "auto", "fp4_gemm_runner_backend": "auto", "bf16_gemm_backend": "auto", "dsa_prefill_backend": null, "dsa_decode_backend": null, "dsa_paged_mqa_logits_backend": "auto", "dsa_topk_backend": "sgl-kernel", "disable_flashinfer_autotune": false, "mamba_backend": "triton", "cuda_graph_config": {"decode": {"backend": "full", "max_bs": 64, "bs": [1, 2, 4, 8, 12, 16, 24, 32, 40, 48, 56, 64], "tc_compiler": "eager", "full_prefill_max_req": null}, "prefill": {"backend": "disabled", "max_bs": 8192, "bs": [4, 8, 12, 16, 20, 24, 28, 32, 48, 64, 80, 96, 112, 128, 144, 160, 176, 192, 208, 224, 240, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 704, 768, 832, 896, 960, 1024, 1280, 1536, 1792, 2048, 2304, 2560, 2816, 3072, 3328, 3584, 3840, 4096, 4608, 5120, 5632, 6144, 6656, 7168, 7680, 8192], "tc_compiler": "eager", "full_prefill_max_req": null}}, "cuda_graph_backend_decode": null, "cuda_graph_backend_prefill": null, "cuda_graph_max_bs_decode": 64, "cuda_graph_max_bs_prefill": null, "cuda_graph_bs_decode": null, "cuda_graph_bs_prefill": null, "cuda_graph_tc_compiler": null, "disable_prefill_cuda_graph": false, "disable_decode_cuda_graph": false, "disable_cuda_graph": false, "disable_cuda_graph_padding": false, "enable_profile_cuda_graph": false, "enable_cudagraph_gc": false, "debug_cuda_graph": false, "enable_layerwise_nvtx_marker": false, "enable_nccl_nvls": false, "enable_symm_mem": false, "triton_attention_reduce_in_fp32": false, "triton_attention_num_kv_splits": 8, "triton_attention_split_tile_size": null, "flashinfer_mla_disable_ragged": false, "enable_fused_qk_norm_rope": false, "enable_precise_embedding_interpolation": false, "enable_fused_moe_sum_all_reduce": false, "enable_deepseek_v4_fp4_indexer": false, "disable_custom_all_reduce": false, "enable_mscclpp": false, "enable_torch_symm_mem": false, "pre_warm_nccl": false, "enable_quant_communications": false, "enable_flashinfer_allreduce_fusion": false, "enforce_disable_flashinfer_allreduce_fusion": false, "flashinfer_allreduce_fusion_backend": null, "enable_aiter_allreduce_fusion": false, "enable_torch_compile": false, "enable_torch_compile_debug_mode": false, "torch_compile_max_bs": 32, "torchao_config": "", "speculative_algorithm": null, "speculative_draft_model_path": null, "speculative_draft_model_revision": null, "speculative_draft_load_format": null, "speculative_num_steps": null, "speculative_eagle_topk": null, "speculative_num_draft_tokens": null, "speculative_dflash_block_size": null, "speculative_dspark_block_size": null, "speculative_dspark_sps_table_path": null, "speculative_dspark_confidence_sts_path": null, "speculative_dspark_align_verify_tokens_to_graph_tier": false, "speculative_accept_threshold_single": 1.0, "speculative_accept_threshold_acc": 1.0, "speculative_use_rejection_sampling": false, "speculative_token_map": null, "speculative_attention_mode": "prefill", "speculative_draft_attention_backend": null, "speculative_draft_window_size": null, "speculative_moe_runner_backend": "flashinfer_mxfp4", "speculative_moe_a2a_backend": null, "speculative_draft_model_quantization": null, "speculative_skip_dp_mlp_sync": false, "enable_multi_layer_eagle": false, "speculative_adaptive": false, "speculative_adaptive_config": null, "decoupled_spec_bind_endpoint": null, "decoupled_spec_connect_endpoints": null, "decoupled_spec_rank": null, "decoupled_spec_role": "null", "spec_trace_dir": null, "speculative_ngram_min_bfs_breadth": 1, "speculative_ngram_max_bfs_breadth": 10, "speculative_ngram_match_type": "BFS", "speculative_ngram_max_trie_depth": 18, "speculative_ngram_capacity": 10000000, "speculative_ngram_external_corpus_path": null, "speculative_ngram_external_sam_budget": 0, "speculative_ngram_external_corpus_max_tokens": 10000000, "ep_size": 2, "moe_a2a_backend": "none", "moe_runner_backend": "flashinfer_mxfp4", "flashinfer_mxfp4_moe_precision": "default", "deepep_mode": "auto", "fuseep_mode": 2, "deepep_dispatcher_output_dtype": "auto", "ep_num_redundant_experts": 0, "ep_dispatch_algorithm": null, "init_expert_location": "trivial", "enable_eplb": false, "eplb_algorithm": "auto", "eplb_rebalance_num_iterations": 1000, "eplb_rebalance_layers_per_chunk": null, "eplb_min_rebalancing_utilization_threshold": 1.0, "expert_distribution_recorder_mode": null, "expert_distribution_recorder_buffer_size": 1000, "enable_expert_distribution_metrics": false, "deepep_config": null, "moe_dense_tp_size": null, "elastic_ep_backend": null, "enable_elastic_expert_backup": false, "mooncake_ib_device": null, "enable_waterfill": false, "ep_join_mode": null, "ep_join_rank_offset": 0, "elastic_ep_initial_size": null, "max_ep_size": null, "elastic_ep_scale_timeout": 600, "elastic_ep_rejoin": false, "disable_flashinfer_cutlass_moe_fp4_allgather": false, "disable_shared_experts_fusion": true, "enforce_shared_experts_fusion": false, "max_mamba_cache_size": null, "mamba_ssm_dtype": null, "enable_mamba_cache_stochastic_rounding": false, "mamba_cache_philox_rounds": 0, "mamba_full_memory_ratio": 0.9, "mamba_radix_cache_strategy": "auto", "uses_mamba_radix_cache": false, "mamba_track_interval": 256, "enable_int8_mamba_checkpoint": false, "int8_mamba_ckpt_size": null, "linear_attn_backend": "triton", "linear_attn_decode_backend": null, "linear_attn_prefill_backend": null, "enable_linear_replayssm": false, "linear_replayssm_cache_len": 16, "enable_hierarchical_cache": false, "hicache_ratio": 2.0, "hicache_size": 0, "hicache_write_policy": "write_through", "hicache_io_backend": "kernel", "hicache_mem_layout": "page_first", "hicache_storage_backend": null, "hicache_storage_prefetch_policy": "timeout", "hicache_storage_backend_extra_config": null, "enable_hisparse": false, "hisparse_config": null, "enable_broadcast_mm_inputs_process": false, "enable_prefix_mm_cache": false, "mm_enable_dp_encoder": false, "mm_process_config": {}, "limit_mm_data_per_request": null, "enable_mm_global_cache": false, "disable_fast_image_processor": false, "mm_feature_transport": "cpu", "keep_mm_feature_on_device": false, "enable_lora": null, "enable_lora_overlap_loading": null, "max_lora_rank": null, "lora_target_modules": null, "lora_paths": null, "max_loaded_loras": null, "max_loras_per_batch": 8, "lora_eviction_policy": "lru", "lora_backend": "csgmv", "max_lora_chunk_size": 16, "experts_shared_outer_loras": null, "lora_use_virtual_experts": false, "lora_strict_loading": false, "lora_drain_wait_threshold": 0.0, "enable_two_batch_overlap": false, "enable_single_batch_overlap": false, "tbo_token_distribution_threshold": 0.48, "cpu_offload_gb": 0, "offload_group_size": -1, "offload_num_in_group": 1, "offload_prefetch_step": 1, "offload_mode": "cpu", "enable_lmcache": false, "lmcache_config_file": null, "enable_flexkv": false, "flexkv_config_file": null, "kt_weight_path": null, "kt_method": "AMXINT4", "kt_cpuinfer": null, "kt_threadpool_count": 2, "kt_num_gpu_experts": null, "kt_max_deferred_experts_per_token": null, "dllm_algorithm": null, "dllm_algorithm_config": null, "dllm_fdfo": true, "disaggregation_mode": "null", "disaggregation_transfer_backend": "mooncake", "disaggregation_bootstrap_port": 8998, "disaggregation_ib_device": null, "disaggregation_decode_enable_radix_cache": false, "disaggregation_decode_enable_offload_kvcache": false, "num_reserved_decode_tokens": 512, "disaggregation_decode_extra_slots": null, "disaggregation_decode_polling_interval": 1, "optimistic_prefill_attempts": 0, "encoder_only": false, "language_only": false, "encoder_transfer_backend": "zmq_to_scheduler", "encoder_urls": [], "encoder_bootstrap_port": 8997, "encoder_register_urls": [], "enable_adaptive_dispatch_to_encoder": false, "enable_pdmux": false, "pdmux_config_path": null, "sm_group_num": 8, "custom_weight_loader": [], "weight_loader_disable_mmap": false, "weight_loader_prefetch_checkpoints": false, "weight_loader_prefetch_num_threads": 4, "weight_loader_drop_cache_after_load": false, "remote_instance_weight_loader_seed_instance_ip": null, "remote_instance_weight_loader_seed_instance_service_port": null, "remote_instance_weight_loader_send_weights_group_ports": null, "remote_instance_weight_loader_backend": "nccl", "remote_instance_weight_loader_start_seed_via_transfer_engine": false, "engine_info_bootstrap_port": 6789, "modelexpress_config": null, "download_dir": null, "model_checksum": null, "delete_ckpt_after_loading": false, "decrypted_config_file": null, "decrypted_draft_config_file": null, "checkpoint_engine_wait_weights_before_ready": false, "enable_prefill_delayer": false, "prefill_delayer_max_delay_passes": 30, "prefill_delayer_token_usage_low_watermark": null, "prefill_delayer_forward_passes_buckets": null, "prefill_delayer_wait_seconds_buckets": null, "prefill_delayer_queue_min_ratio": null, "prefill_delayer_max_delay_ms": null, "min_free_slots_delay": null, "enable_deterministic_inference": false, "rl_on_policy_target": null, "kv_canary": "none", "kv_canary_real_data": "none", "kv_canary_sweep_interval": 0, "enable_dynamic_batch_tokenizer": false, "dynamic_batch_tokenizer_batch_size": 32, "dynamic_batch_tokenizer_batch_timeout": 0.002, "enable_tokenizer_batch_encode": false, "disable_tokenizer_batch_decode": false, "debug_tensor_dump_output_folder": null, "debug_tensor_dump_layers": null, "debug_tensor_dump_input_file": null, "enable_memory_saver": false, "enable_weights_cpu_backup": false, "enable_draft_weights_cpu_backup": false, "enable_custom_logit_processor": false, "enable_return_hidden_states": false, "enable_return_routed_experts": false, "enable_return_indexer_topk": false, "disable_outlines_disk_cache": false, "enable_mis": false, "forward_hooks": null, "msprobe_dump_config": null, "_resolved_overrides": [["_deepseek_v4_overrides", {"attention_backend": "dsv4", "page_size": 256, "swa_full_tokens_ratio": 0.1}], ["_deepseek_v4_kv_cache_dtype", {"kv_cache_dtype": "fp8_e4m3"}], ["_deepseek_v4_sm120_moe", {"moe_runner_backend": "flashinfer_mxfp4"}], ["_sampling_backend_default", {"sampling_backend": "flashinfer"}], ["_data_parallelism_defaults", {"enable_dp_attention": false, "enable_dp_lm_head": false}], ["_speculative_moe_runner_default", {"speculative_moe_runner_backend": "flashinfer_mxfp4"}], ["DeepseekV4ForCausalLM.determine_num_fused_shared_experts", {"disable_shared_experts_fusion": true}]], "grpc_worker_threads": 4, "_quantization_explicitly_unset": false, "_cuda_graph_config_locked": [["decode", "max_bs"]], "_declarations_materialized": true, "_runtime_mutations": [["model_runner.chunked_prefix_cache_gate", {"disable_chunked_prefix_cache": true}], ["scheduler.pp_max_micro_batch_size_default", {"pp_max_micro_batch_size": 256}]], "_in_override": false, "_mx_config_cache": {}, "max_speculative_num_draft_tokens": null, "last_gen_throughput": 0.0, "memory_usage": {"weight": 56.87, "kvcache": 0, "token_capacity": 1282304, "graph": 1.39}, "effective_max_running_requests_per_dp": 256}], "version": "0.0.0.dev1+g35f2d4f76", "kv_events": null}, "duration": 16.05487186202663, "completed": 1, "total_input_tokens": 1024, "total_input_text_tokens": 1024, "total_input_vision_tokens": 0, "total_output_tokens": 1, "total_output_tokens_retokenized": 1, "request_throughput": 0.06228638936478989, "input_throughput": 63.78126270954485, "output_throughput": 0.06228638936478989, "total_throughput": 63.84354909890964, "mean_e2e_latency_ms": 16035.621489019832, "median_e2e_latency_ms": 16035.621489019832, "std_e2e_latency_ms": 0.0, "p90_e2e_latency_ms": 16035.621489019832, "p95_e2e_latency_ms": 16035.621489019832, "p99_e2e_latency_ms": 16035.621489019832, "mean_ttft_ms": 16035.530886001652, "median_ttft_ms": 16035.530886001652, "std_ttft_ms": 0.0, "p90_ttft_ms": 16035.530886001652, "p95_ttft_ms": 16035.530886001652, "p99_ttft_ms": 16035.530886001652, "mean_tpot_ms": 0.0, "median_tpot_ms": 0.0, "std_tpot_ms": 0.0, "p90_tpot_ms": 0.0, "p95_tpot_ms": 0.0, "p99_tpot_ms": 0.0, "mean_itl_ms": 0.0, "median_itl_ms": 0.0, "std_itl_ms": 0.0, "p90_itl_ms": 0.0, "p95_itl_ms": 0.0, "p99_itl_ms": 0.0, "concurrency": 0.9988009637714811, "accept_length": null, "max_output_tokens_per_s": 1.0, "max_concurrent_requests": 1, "input_lens": [1024], "output_lens": [1], "ttfts": [16.035530886001652], "itls": [[]], "generated_texts": [".\n"], "errors": [""]} diff --git a/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/cold_1k_o1_repeat.jsonl b/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/cold_1k_o1_repeat.jsonl new file mode 100644 index 0000000..78b9df3 --- /dev/null +++ b/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/cold_1k_o1_repeat.jsonl @@ -0,0 +1 @@ +{"tag": null, "backend": "sglang", "dataset_name": "random", "request_rate": 10000.0, "max_concurrency": 1, "sharegpt_output_len": null, "random_input_len": 1024, "random_output_len": 1, "random_range_ratio": 1.0, "server_info": {"model_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_mode": "auto", "tokenizer_backend": "huggingface", "tokenizer_worker_num": 1, "detokenizer_worker_num": 1, "skip_tokenizer_init": false, "load_format": "auto", "model_loader_extra_config": "{}", "trust_remote_code": true, "context_length": null, "is_embedding": false, "enable_multimodal": null, "revision": null, "model_impl": "auto", "model_config_parser": "auto", "json_model_override_args": "{}", "dtype": "auto", "quantization": null, "quantization_param_path": null, "kv_cache_dtype": "fp8_e4m3", "enable_fp32_lm_head": false, "modelopt_quant": null, "modelopt_checkpoint_restore_path": null, "modelopt_checkpoint_save_path": null, "modelopt_export_path": null, "quantize_and_serve": false, "rl_quant_profile": null, "enable_tf32_matmul": false, "mem_fraction_static": 0.9, "max_running_requests": 256, "max_queued_requests": null, "max_total_tokens": null, "chunked_prefill_size": 8192, "enable_dynamic_chunking": false, "max_prefill_tokens": 16384, "prefill_max_requests": null, "schedule_policy": "fcfs", "enable_priority_scheduling": false, "disable_priority_preemption": false, "default_priority_value": null, "abort_on_priority_when_disabled": false, "schedule_low_priority_values_first": false, "priority_scheduling_preemption_threshold": 10, "retraction_policy": "length", "schedule_conservativeness": 1.0, "page_size": 256, "swa_full_tokens_ratio": 0.1, "disable_hybrid_swa_memory": false, "radix_eviction_policy": "lru", "prefill_only_disable_kv_cache": false, "disable_radix_cache": false, "enable_page_major_kv_layout": false, "enable_unified_memory": false, "disable_chunked_prefix_cache": false, "disable_overlap_schedule": false, "num_continuous_decode_steps": 1, "scheduler_recv_interval": 1, "enable_mixed_chunk": false, "nccl_port": null, "dist_timeout": null, "dist_init_addr": "10.101.0.11:20002", "nnodes": 2, "node_rank": 0, "tp_size": 16, "dcp_size": 1, "pp_size": 1, "pp_max_micro_batch_size": null, "pp_async_batch_depth": 0, "dp_size": 1, "load_balance_method": "round_robin", "attn_cp_size": 1, "moe_dp_size": 1, "enable_prefill_cp": false, "cp_strategy": null, "enable_dsa_cache_layer_split": false, "enable_dsa_prefill_context_parallel": false, "dsa_prefill_cp_mode": "round-robin-split", "enable_prefill_context_parallel": false, "prefill_cp_mode": "in-seq-split", "enable_dp_attention": false, "enable_dp_attention_local_control_broadcast": false, "enable_dp_lm_head": false, "enable_attn_tp_input_scattered": false, "disable_attn_tp_gather": false, "enable_p2p_check": false, "device": "cuda", "base_gpu_id": 0, "gpu_id_step": 1, "random_seed": 42600328, "watchdog_timeout": 300, "soft_watchdog_timeout": null, "sleep_on_idle": false, "use_ray": false, "custom_sigquit_handler": null, "numa_node": null, "gc_threshold": null, "host": "0.0.0.0", "port": 30002, "fastapi_root_path": "", "smg_grpc_mode": false, "grpc_mode": false, "grpc_port": null, "skip_server_warmup": false, "warmups": null, "enable_http2": false, "ssl_keyfile": null, "ssl_certfile": null, "ssl_ca_certs": null, "ssl_keyfile_password": null, "enable_ssl_refresh": false, "api_key": null, "admin_api_key": null, "served_model_name": "/data/hf_models/DeepSeek-V4-Pro", "weight_version": "default", "chat_template": null, "hf_chat_template_name": null, "completion_template": null, "file_storage_path": "sglang_storage", "enable_cache_report": false, "reasoning_parser": null, "default_chat_template_kwargs": null, "strip_thinking_cache": false, "enable_strict_thinking": false, "tool_call_parser": null, "tool_server": null, "sampling_defaults": "model", "asr_max_buffer_seconds": 60, "asr_max_concurrent_sessions": 32, "preferred_sampling_params": null, "allow_auto_truncate": false, "stream_interval": 1, "batch_notify_size": 16, "stream_response_default_include_usage": false, "incremental_streaming_output": false, "enable_streaming_session": false, "enable_session_radix_cache": false, "log_level": "info", "log_level_http": null, "log_requests": false, "log_requests_level": 2, "log_requests_format": "text", "log_requests_target": null, "uvicorn_access_log_exclude_prefixes": [], "crash_dump_folder": null, "show_time_cost": false, "enable_metrics": false, "smg_http_sidecar_port": null, "enable_mfu_metrics": false, "enable_metrics_for_all_schedulers": false, "load_snapshot_publish_interval": 15, "tokenizer_metrics_custom_labels_header": "x-custom-labels", "tokenizer_metrics_allowed_custom_labels": null, "extra_metric_labels": null, "bucket_time_to_first_token": null, "bucket_inter_token_latency": null, "bucket_e2e_request_latency": null, "prompt_tokens_buckets": null, "generation_tokens_buckets": null, "gc_warning_threshold_secs": 0.0, "decode_log_interval": 40, "enable_request_time_stats_logging": false, "kv_events_config": null, "enable_forward_pass_metrics": false, "forward_pass_metrics_worker_id": "", "forward_pass_metrics_ipc_name": null, "enable_trace": false, "trace_modules": "request", "otlp_traces_endpoint": "localhost:4317", "export_metrics_to_file": false, "export_metrics_to_file_dir": null, "stat_loggers": null, "constrained_json_whitespace_pattern": null, "constrained_json_disable_any_whitespace": false, "attention_backend": "dsv4", "decode_attention_backend": null, "prefill_attention_backend": null, "sampling_backend": "flashinfer", "grammar_backend": "xgrammar", "radix_cache_backend": null, "mm_attention_backend": null, "fp8_gemm_runner_backend": "auto", "fp4_gemm_runner_backend": "auto", "bf16_gemm_backend": "auto", "dsa_prefill_backend": null, "dsa_decode_backend": null, "dsa_paged_mqa_logits_backend": "auto", "dsa_topk_backend": "sgl-kernel", "disable_flashinfer_autotune": false, "mamba_backend": "triton", "cuda_graph_config": {"decode": {"backend": "full", "max_bs": 64, "bs": [1, 2, 4, 8, 12, 16, 24, 32, 40, 48, 56, 64], "tc_compiler": "eager", "full_prefill_max_req": null}, "prefill": {"backend": "disabled", "max_bs": 8192, "bs": [4, 8, 12, 16, 20, 24, 28, 32, 48, 64, 80, 96, 112, 128, 144, 160, 176, 192, 208, 224, 240, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 704, 768, 832, 896, 960, 1024, 1280, 1536, 1792, 2048, 2304, 2560, 2816, 3072, 3328, 3584, 3840, 4096, 4608, 5120, 5632, 6144, 6656, 7168, 7680, 8192], "tc_compiler": "eager", "full_prefill_max_req": null}}, "cuda_graph_backend_decode": null, "cuda_graph_backend_prefill": null, "cuda_graph_max_bs_decode": 64, "cuda_graph_max_bs_prefill": null, "cuda_graph_bs_decode": null, "cuda_graph_bs_prefill": null, "cuda_graph_tc_compiler": null, "disable_prefill_cuda_graph": false, "disable_decode_cuda_graph": false, "disable_cuda_graph": false, "disable_cuda_graph_padding": false, "enable_profile_cuda_graph": false, "enable_cudagraph_gc": false, "debug_cuda_graph": false, "enable_layerwise_nvtx_marker": false, "enable_nccl_nvls": false, "enable_symm_mem": false, "triton_attention_reduce_in_fp32": false, "triton_attention_num_kv_splits": 8, "triton_attention_split_tile_size": null, "flashinfer_mla_disable_ragged": false, "enable_fused_qk_norm_rope": false, "enable_precise_embedding_interpolation": false, "enable_fused_moe_sum_all_reduce": false, "enable_deepseek_v4_fp4_indexer": false, "disable_custom_all_reduce": false, "enable_mscclpp": false, "enable_torch_symm_mem": false, "pre_warm_nccl": false, "enable_quant_communications": false, "enable_flashinfer_allreduce_fusion": false, "enforce_disable_flashinfer_allreduce_fusion": false, "flashinfer_allreduce_fusion_backend": null, "enable_aiter_allreduce_fusion": false, "enable_torch_compile": false, "enable_torch_compile_debug_mode": false, "torch_compile_max_bs": 32, "torchao_config": "", "speculative_algorithm": null, "speculative_draft_model_path": null, "speculative_draft_model_revision": null, "speculative_draft_load_format": null, "speculative_num_steps": null, "speculative_eagle_topk": null, "speculative_num_draft_tokens": null, "speculative_dflash_block_size": null, "speculative_dspark_block_size": null, "speculative_dspark_sps_table_path": null, "speculative_dspark_confidence_sts_path": null, "speculative_dspark_align_verify_tokens_to_graph_tier": false, "speculative_accept_threshold_single": 1.0, "speculative_accept_threshold_acc": 1.0, "speculative_use_rejection_sampling": false, "speculative_token_map": null, "speculative_attention_mode": "prefill", "speculative_draft_attention_backend": null, "speculative_draft_window_size": null, "speculative_moe_runner_backend": "flashinfer_mxfp4", "speculative_moe_a2a_backend": null, "speculative_draft_model_quantization": null, "speculative_skip_dp_mlp_sync": false, "enable_multi_layer_eagle": false, "speculative_adaptive": false, "speculative_adaptive_config": null, "decoupled_spec_bind_endpoint": null, "decoupled_spec_connect_endpoints": null, "decoupled_spec_rank": null, "decoupled_spec_role": "null", "spec_trace_dir": null, "speculative_ngram_min_bfs_breadth": 1, "speculative_ngram_max_bfs_breadth": 10, "speculative_ngram_match_type": "BFS", "speculative_ngram_max_trie_depth": 18, "speculative_ngram_capacity": 10000000, "speculative_ngram_external_corpus_path": null, "speculative_ngram_external_sam_budget": 0, "speculative_ngram_external_corpus_max_tokens": 10000000, "ep_size": 2, "moe_a2a_backend": "none", "moe_runner_backend": "flashinfer_mxfp4", "flashinfer_mxfp4_moe_precision": "default", "deepep_mode": "auto", "fuseep_mode": 2, "deepep_dispatcher_output_dtype": "auto", "ep_num_redundant_experts": 0, "ep_dispatch_algorithm": null, "init_expert_location": "trivial", "enable_eplb": false, "eplb_algorithm": "auto", "eplb_rebalance_num_iterations": 1000, "eplb_rebalance_layers_per_chunk": null, "eplb_min_rebalancing_utilization_threshold": 1.0, "expert_distribution_recorder_mode": null, "expert_distribution_recorder_buffer_size": 1000, "enable_expert_distribution_metrics": false, "deepep_config": null, "moe_dense_tp_size": null, "elastic_ep_backend": null, "enable_elastic_expert_backup": false, "mooncake_ib_device": null, "enable_waterfill": false, "ep_join_mode": null, "ep_join_rank_offset": 0, "elastic_ep_initial_size": null, "max_ep_size": null, "elastic_ep_scale_timeout": 600, "elastic_ep_rejoin": false, "disable_flashinfer_cutlass_moe_fp4_allgather": false, "disable_shared_experts_fusion": false, "enforce_shared_experts_fusion": false, "max_mamba_cache_size": null, "mamba_ssm_dtype": null, "enable_mamba_cache_stochastic_rounding": false, "mamba_cache_philox_rounds": 0, "mamba_full_memory_ratio": 0.9, "mamba_radix_cache_strategy": "auto", "uses_mamba_radix_cache": false, "mamba_track_interval": 256, "enable_int8_mamba_checkpoint": false, "int8_mamba_ckpt_size": null, "linear_attn_backend": "triton", "linear_attn_decode_backend": null, "linear_attn_prefill_backend": null, "enable_linear_replayssm": false, "linear_replayssm_cache_len": 16, "enable_hierarchical_cache": false, "hicache_ratio": 2.0, "hicache_size": 0, "hicache_write_policy": "write_through", "hicache_io_backend": "kernel", "hicache_mem_layout": "page_first", "hicache_storage_backend": null, "hicache_storage_prefetch_policy": "timeout", "hicache_storage_backend_extra_config": null, "enable_hisparse": false, "hisparse_config": null, "enable_broadcast_mm_inputs_process": false, "enable_prefix_mm_cache": false, "mm_enable_dp_encoder": false, "mm_process_config": {}, "limit_mm_data_per_request": null, "enable_mm_global_cache": false, "disable_fast_image_processor": false, "mm_feature_transport": "cpu", "keep_mm_feature_on_device": false, "enable_lora": null, "enable_lora_overlap_loading": null, "max_lora_rank": null, "lora_target_modules": null, "lora_paths": null, "max_loaded_loras": null, "max_loras_per_batch": 8, "lora_eviction_policy": "lru", "lora_backend": "csgmv", "max_lora_chunk_size": 16, "experts_shared_outer_loras": null, "lora_use_virtual_experts": false, "lora_strict_loading": false, "lora_drain_wait_threshold": 0.0, "enable_two_batch_overlap": false, "enable_single_batch_overlap": false, "tbo_token_distribution_threshold": 0.48, "cpu_offload_gb": 0, "offload_group_size": -1, "offload_num_in_group": 1, "offload_prefetch_step": 1, "offload_mode": "cpu", "enable_lmcache": false, "lmcache_config_file": null, "enable_flexkv": false, "flexkv_config_file": null, "kt_weight_path": null, "kt_method": "AMXINT4", "kt_cpuinfer": null, "kt_threadpool_count": 2, "kt_num_gpu_experts": null, "kt_max_deferred_experts_per_token": null, "dllm_algorithm": null, "dllm_algorithm_config": null, "dllm_fdfo": true, "disaggregation_mode": "null", "disaggregation_transfer_backend": "mooncake", "disaggregation_bootstrap_port": 8998, "disaggregation_ib_device": null, "disaggregation_decode_enable_radix_cache": false, "disaggregation_decode_enable_offload_kvcache": false, "num_reserved_decode_tokens": 512, "disaggregation_decode_extra_slots": null, "disaggregation_decode_polling_interval": 1, "optimistic_prefill_attempts": 0, "encoder_only": false, "language_only": false, "encoder_transfer_backend": "zmq_to_scheduler", "encoder_urls": [], "encoder_bootstrap_port": 8997, "encoder_register_urls": [], "enable_adaptive_dispatch_to_encoder": false, "enable_pdmux": false, "pdmux_config_path": null, "sm_group_num": 8, "custom_weight_loader": [], "weight_loader_disable_mmap": false, "weight_loader_prefetch_checkpoints": false, "weight_loader_prefetch_num_threads": 4, "weight_loader_drop_cache_after_load": false, "remote_instance_weight_loader_seed_instance_ip": null, "remote_instance_weight_loader_seed_instance_service_port": null, "remote_instance_weight_loader_send_weights_group_ports": null, "remote_instance_weight_loader_backend": "nccl", "remote_instance_weight_loader_start_seed_via_transfer_engine": false, "engine_info_bootstrap_port": 6789, "modelexpress_config": null, "download_dir": null, "model_checksum": null, "delete_ckpt_after_loading": false, "decrypted_config_file": null, "decrypted_draft_config_file": null, "checkpoint_engine_wait_weights_before_ready": false, "enable_prefill_delayer": false, "prefill_delayer_max_delay_passes": 30, "prefill_delayer_token_usage_low_watermark": null, "prefill_delayer_forward_passes_buckets": null, "prefill_delayer_wait_seconds_buckets": null, "prefill_delayer_queue_min_ratio": null, "prefill_delayer_max_delay_ms": null, "min_free_slots_delay": null, "enable_deterministic_inference": false, "rl_on_policy_target": null, "kv_canary": "none", "kv_canary_real_data": "none", "kv_canary_sweep_interval": 0, "enable_dynamic_batch_tokenizer": false, "dynamic_batch_tokenizer_batch_size": 32, "dynamic_batch_tokenizer_batch_timeout": 0.002, "enable_tokenizer_batch_encode": false, "disable_tokenizer_batch_decode": false, "debug_tensor_dump_output_folder": null, "debug_tensor_dump_layers": null, "debug_tensor_dump_input_file": null, "enable_memory_saver": false, "enable_weights_cpu_backup": false, "enable_draft_weights_cpu_backup": false, "enable_custom_logit_processor": false, "enable_return_hidden_states": false, "enable_return_routed_experts": false, "enable_return_indexer_topk": false, "disable_outlines_disk_cache": false, "enable_mis": false, "forward_hooks": null, "msprobe_dump_config": null, "status": "ready", "max_total_num_tokens": 1282304, "max_req_input_len": 1048570, "internal_states": [{"model_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_mode": "auto", "tokenizer_backend": "huggingface", "tokenizer_worker_num": 1, "detokenizer_worker_num": 1, "skip_tokenizer_init": false, "load_format": "auto", "model_loader_extra_config": "{}", "trust_remote_code": true, "context_length": null, "is_embedding": false, "enable_multimodal": null, "revision": null, "model_impl": "auto", "model_config_parser": "auto", "json_model_override_args": "{}", "dtype": "auto", "quantization": null, "quantization_param_path": null, "kv_cache_dtype": "fp8_e4m3", "enable_fp32_lm_head": false, "modelopt_quant": null, "modelopt_checkpoint_restore_path": null, "modelopt_checkpoint_save_path": null, "modelopt_export_path": null, "quantize_and_serve": false, "rl_quant_profile": null, "enable_tf32_matmul": false, "mem_fraction_static": 0.9, "max_running_requests": 256, "max_queued_requests": null, "max_total_tokens": null, "chunked_prefill_size": 8192, "enable_dynamic_chunking": false, "max_prefill_tokens": 16384, "prefill_max_requests": null, "schedule_policy": "fcfs", "enable_priority_scheduling": false, "disable_priority_preemption": false, "default_priority_value": null, "abort_on_priority_when_disabled": false, "schedule_low_priority_values_first": false, "priority_scheduling_preemption_threshold": 10, "retraction_policy": "length", "schedule_conservativeness": 1.0, "page_size": 256, "swa_full_tokens_ratio": 0.1, "disable_hybrid_swa_memory": false, "radix_eviction_policy": "lru", "prefill_only_disable_kv_cache": false, "disable_radix_cache": false, "enable_page_major_kv_layout": false, "enable_unified_memory": false, "disable_chunked_prefix_cache": true, "disable_overlap_schedule": false, "num_continuous_decode_steps": 1, "scheduler_recv_interval": 1, "enable_mixed_chunk": false, "nccl_port": null, "dist_timeout": null, "dist_init_addr": "10.101.0.11:20002", "nnodes": 2, "node_rank": 0, "tp_size": 16, "dcp_size": 1, "pp_size": 1, "pp_max_micro_batch_size": 256, "pp_async_batch_depth": 0, "dp_size": 1, "load_balance_method": "round_robin", "attn_cp_size": 1, "moe_dp_size": 1, "enable_prefill_cp": false, "cp_strategy": null, "enable_dsa_cache_layer_split": false, "enable_dsa_prefill_context_parallel": false, "dsa_prefill_cp_mode": "round-robin-split", "enable_prefill_context_parallel": false, "prefill_cp_mode": "in-seq-split", "enable_dp_attention": false, "enable_dp_attention_local_control_broadcast": false, "enable_dp_lm_head": false, "enable_attn_tp_input_scattered": false, "disable_attn_tp_gather": false, "enable_p2p_check": false, "device": "cuda", "base_gpu_id": 0, "gpu_id_step": 1, "random_seed": 42600328, "watchdog_timeout": 300, "soft_watchdog_timeout": null, "sleep_on_idle": false, "use_ray": false, "numa_node": null, "gc_threshold": null, "host": "0.0.0.0", "port": 30002, "fastapi_root_path": "", "smg_grpc_mode": false, "grpc_mode": false, "grpc_port": null, "skip_server_warmup": false, "warmups": null, "enable_http2": false, "ssl_keyfile": null, "ssl_certfile": null, "ssl_ca_certs": null, "ssl_keyfile_password": null, "enable_ssl_refresh": false, "api_key": null, "admin_api_key": null, "served_model_name": "/data/hf_models/DeepSeek-V4-Pro", "weight_version": "default", "chat_template": null, "hf_chat_template_name": null, "completion_template": null, "file_storage_path": "sglang_storage", "enable_cache_report": false, "reasoning_parser": null, "default_chat_template_kwargs": null, "strip_thinking_cache": false, "enable_strict_thinking": false, "tool_call_parser": null, "tool_server": null, "sampling_defaults": "model", "asr_max_buffer_seconds": 60, "asr_max_concurrent_sessions": 32, "preferred_sampling_params": null, "allow_auto_truncate": false, "stream_interval": 1, "batch_notify_size": 16, "stream_response_default_include_usage": false, "incremental_streaming_output": false, "enable_streaming_session": false, "enable_session_radix_cache": false, "log_level": "info", "log_level_http": null, "log_requests": false, "log_requests_level": 2, "log_requests_format": "text", "log_requests_target": null, "uvicorn_access_log_exclude_prefixes": [], "crash_dump_folder": null, "show_time_cost": false, "enable_metrics": false, "smg_http_sidecar_port": null, "enable_mfu_metrics": false, "enable_metrics_for_all_schedulers": false, "load_snapshot_publish_interval": 15, "tokenizer_metrics_custom_labels_header": "x-custom-labels", "tokenizer_metrics_allowed_custom_labels": null, "extra_metric_labels": null, "bucket_time_to_first_token": null, "bucket_inter_token_latency": null, "bucket_e2e_request_latency": null, "prompt_tokens_buckets": null, "generation_tokens_buckets": null, "gc_warning_threshold_secs": 0.0, "decode_log_interval": 40, "enable_request_time_stats_logging": false, "kv_events_config": null, "enable_forward_pass_metrics": false, "forward_pass_metrics_worker_id": "", "forward_pass_metrics_ipc_name": null, "enable_trace": false, "trace_modules": "request", "otlp_traces_endpoint": "localhost:4317", "export_metrics_to_file": false, "export_metrics_to_file_dir": null, "stat_loggers": null, "constrained_json_whitespace_pattern": null, "constrained_json_disable_any_whitespace": false, "attention_backend": "dsv4", "decode_attention_backend": null, "prefill_attention_backend": null, "sampling_backend": "flashinfer", "grammar_backend": "xgrammar", "radix_cache_backend": null, "mm_attention_backend": null, "fp8_gemm_runner_backend": "auto", "fp4_gemm_runner_backend": "auto", "bf16_gemm_backend": "auto", "dsa_prefill_backend": null, "dsa_decode_backend": null, "dsa_paged_mqa_logits_backend": "auto", "dsa_topk_backend": "sgl-kernel", "disable_flashinfer_autotune": false, "mamba_backend": "triton", "cuda_graph_config": {"decode": {"backend": "full", "max_bs": 64, "bs": [1, 2, 4, 8, 12, 16, 24, 32, 40, 48, 56, 64], "tc_compiler": "eager", "full_prefill_max_req": null}, "prefill": {"backend": "disabled", "max_bs": 8192, "bs": [4, 8, 12, 16, 20, 24, 28, 32, 48, 64, 80, 96, 112, 128, 144, 160, 176, 192, 208, 224, 240, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 704, 768, 832, 896, 960, 1024, 1280, 1536, 1792, 2048, 2304, 2560, 2816, 3072, 3328, 3584, 3840, 4096, 4608, 5120, 5632, 6144, 6656, 7168, 7680, 8192], "tc_compiler": "eager", "full_prefill_max_req": null}}, "cuda_graph_backend_decode": null, "cuda_graph_backend_prefill": null, "cuda_graph_max_bs_decode": 64, "cuda_graph_max_bs_prefill": null, "cuda_graph_bs_decode": null, "cuda_graph_bs_prefill": null, "cuda_graph_tc_compiler": null, "disable_prefill_cuda_graph": false, "disable_decode_cuda_graph": false, "disable_cuda_graph": false, "disable_cuda_graph_padding": false, "enable_profile_cuda_graph": false, "enable_cudagraph_gc": false, "debug_cuda_graph": false, "enable_layerwise_nvtx_marker": false, "enable_nccl_nvls": false, "enable_symm_mem": false, "triton_attention_reduce_in_fp32": false, "triton_attention_num_kv_splits": 8, "triton_attention_split_tile_size": null, "flashinfer_mla_disable_ragged": false, "enable_fused_qk_norm_rope": false, "enable_precise_embedding_interpolation": false, "enable_fused_moe_sum_all_reduce": false, "enable_deepseek_v4_fp4_indexer": false, "disable_custom_all_reduce": false, "enable_mscclpp": false, "enable_torch_symm_mem": false, "pre_warm_nccl": false, "enable_quant_communications": false, "enable_flashinfer_allreduce_fusion": false, "enforce_disable_flashinfer_allreduce_fusion": false, "flashinfer_allreduce_fusion_backend": null, "enable_aiter_allreduce_fusion": false, "enable_torch_compile": false, "enable_torch_compile_debug_mode": false, "torch_compile_max_bs": 32, "torchao_config": "", "speculative_algorithm": null, "speculative_draft_model_path": null, "speculative_draft_model_revision": null, "speculative_draft_load_format": null, "speculative_num_steps": null, "speculative_eagle_topk": null, "speculative_num_draft_tokens": null, "speculative_dflash_block_size": null, "speculative_dspark_block_size": null, "speculative_dspark_sps_table_path": null, "speculative_dspark_confidence_sts_path": null, "speculative_dspark_align_verify_tokens_to_graph_tier": false, "speculative_accept_threshold_single": 1.0, "speculative_accept_threshold_acc": 1.0, "speculative_use_rejection_sampling": false, "speculative_token_map": null, "speculative_attention_mode": "prefill", "speculative_draft_attention_backend": null, "speculative_draft_window_size": null, "speculative_moe_runner_backend": "flashinfer_mxfp4", "speculative_moe_a2a_backend": null, "speculative_draft_model_quantization": null, "speculative_skip_dp_mlp_sync": false, "enable_multi_layer_eagle": false, "speculative_adaptive": false, "speculative_adaptive_config": null, "decoupled_spec_bind_endpoint": null, "decoupled_spec_connect_endpoints": null, "decoupled_spec_rank": null, "decoupled_spec_role": "null", "spec_trace_dir": null, "speculative_ngram_min_bfs_breadth": 1, "speculative_ngram_max_bfs_breadth": 10, "speculative_ngram_match_type": "BFS", "speculative_ngram_max_trie_depth": 18, "speculative_ngram_capacity": 10000000, "speculative_ngram_external_corpus_path": null, "speculative_ngram_external_sam_budget": 0, "speculative_ngram_external_corpus_max_tokens": 10000000, "ep_size": 2, "moe_a2a_backend": "none", "moe_runner_backend": "flashinfer_mxfp4", "flashinfer_mxfp4_moe_precision": "default", "deepep_mode": "auto", "fuseep_mode": 2, "deepep_dispatcher_output_dtype": "auto", "ep_num_redundant_experts": 0, "ep_dispatch_algorithm": null, "init_expert_location": "trivial", "enable_eplb": false, "eplb_algorithm": "auto", "eplb_rebalance_num_iterations": 1000, "eplb_rebalance_layers_per_chunk": null, "eplb_min_rebalancing_utilization_threshold": 1.0, "expert_distribution_recorder_mode": null, "expert_distribution_recorder_buffer_size": 1000, "enable_expert_distribution_metrics": false, "deepep_config": null, "moe_dense_tp_size": null, "elastic_ep_backend": null, "enable_elastic_expert_backup": false, "mooncake_ib_device": null, "enable_waterfill": false, "ep_join_mode": null, "ep_join_rank_offset": 0, "elastic_ep_initial_size": null, "max_ep_size": null, "elastic_ep_scale_timeout": 600, "elastic_ep_rejoin": false, "disable_flashinfer_cutlass_moe_fp4_allgather": false, "disable_shared_experts_fusion": true, "enforce_shared_experts_fusion": false, "max_mamba_cache_size": null, "mamba_ssm_dtype": null, "enable_mamba_cache_stochastic_rounding": false, "mamba_cache_philox_rounds": 0, "mamba_full_memory_ratio": 0.9, "mamba_radix_cache_strategy": "auto", "uses_mamba_radix_cache": false, "mamba_track_interval": 256, "enable_int8_mamba_checkpoint": false, "int8_mamba_ckpt_size": null, "linear_attn_backend": "triton", "linear_attn_decode_backend": null, "linear_attn_prefill_backend": null, "enable_linear_replayssm": false, "linear_replayssm_cache_len": 16, "enable_hierarchical_cache": false, "hicache_ratio": 2.0, "hicache_size": 0, "hicache_write_policy": "write_through", "hicache_io_backend": "kernel", "hicache_mem_layout": "page_first", "hicache_storage_backend": null, "hicache_storage_prefetch_policy": "timeout", "hicache_storage_backend_extra_config": null, "enable_hisparse": false, "hisparse_config": null, "enable_broadcast_mm_inputs_process": false, "enable_prefix_mm_cache": false, "mm_enable_dp_encoder": false, "mm_process_config": {}, "limit_mm_data_per_request": null, "enable_mm_global_cache": false, "disable_fast_image_processor": false, "mm_feature_transport": "cpu", "keep_mm_feature_on_device": false, "enable_lora": null, "enable_lora_overlap_loading": null, "max_lora_rank": null, "lora_target_modules": null, "lora_paths": null, "max_loaded_loras": null, "max_loras_per_batch": 8, "lora_eviction_policy": "lru", "lora_backend": "csgmv", "max_lora_chunk_size": 16, "experts_shared_outer_loras": null, "lora_use_virtual_experts": false, "lora_strict_loading": false, "lora_drain_wait_threshold": 0.0, "enable_two_batch_overlap": false, "enable_single_batch_overlap": false, "tbo_token_distribution_threshold": 0.48, "cpu_offload_gb": 0, "offload_group_size": -1, "offload_num_in_group": 1, "offload_prefetch_step": 1, "offload_mode": "cpu", "enable_lmcache": false, "lmcache_config_file": null, "enable_flexkv": false, "flexkv_config_file": null, "kt_weight_path": null, "kt_method": "AMXINT4", "kt_cpuinfer": null, "kt_threadpool_count": 2, "kt_num_gpu_experts": null, "kt_max_deferred_experts_per_token": null, "dllm_algorithm": null, "dllm_algorithm_config": null, "dllm_fdfo": true, "disaggregation_mode": "null", "disaggregation_transfer_backend": "mooncake", "disaggregation_bootstrap_port": 8998, "disaggregation_ib_device": null, "disaggregation_decode_enable_radix_cache": false, "disaggregation_decode_enable_offload_kvcache": false, "num_reserved_decode_tokens": 512, "disaggregation_decode_extra_slots": null, "disaggregation_decode_polling_interval": 1, "optimistic_prefill_attempts": 0, "encoder_only": false, "language_only": false, "encoder_transfer_backend": "zmq_to_scheduler", "encoder_urls": [], "encoder_bootstrap_port": 8997, "encoder_register_urls": [], "enable_adaptive_dispatch_to_encoder": false, "enable_pdmux": false, "pdmux_config_path": null, "sm_group_num": 8, "custom_weight_loader": [], "weight_loader_disable_mmap": false, "weight_loader_prefetch_checkpoints": false, "weight_loader_prefetch_num_threads": 4, "weight_loader_drop_cache_after_load": false, "remote_instance_weight_loader_seed_instance_ip": null, "remote_instance_weight_loader_seed_instance_service_port": null, "remote_instance_weight_loader_send_weights_group_ports": null, "remote_instance_weight_loader_backend": "nccl", "remote_instance_weight_loader_start_seed_via_transfer_engine": false, "engine_info_bootstrap_port": 6789, "modelexpress_config": null, "download_dir": null, "model_checksum": null, "delete_ckpt_after_loading": false, "decrypted_config_file": null, "decrypted_draft_config_file": null, "checkpoint_engine_wait_weights_before_ready": false, "enable_prefill_delayer": false, "prefill_delayer_max_delay_passes": 30, "prefill_delayer_token_usage_low_watermark": null, "prefill_delayer_forward_passes_buckets": null, "prefill_delayer_wait_seconds_buckets": null, "prefill_delayer_queue_min_ratio": null, "prefill_delayer_max_delay_ms": null, "min_free_slots_delay": null, "enable_deterministic_inference": false, "rl_on_policy_target": null, "kv_canary": "none", "kv_canary_real_data": "none", "kv_canary_sweep_interval": 0, "enable_dynamic_batch_tokenizer": false, "dynamic_batch_tokenizer_batch_size": 32, "dynamic_batch_tokenizer_batch_timeout": 0.002, "enable_tokenizer_batch_encode": false, "disable_tokenizer_batch_decode": false, "debug_tensor_dump_output_folder": null, "debug_tensor_dump_layers": null, "debug_tensor_dump_input_file": null, "enable_memory_saver": false, "enable_weights_cpu_backup": false, "enable_draft_weights_cpu_backup": false, "enable_custom_logit_processor": false, "enable_return_hidden_states": false, "enable_return_routed_experts": false, "enable_return_indexer_topk": false, "disable_outlines_disk_cache": false, "enable_mis": false, "forward_hooks": null, "msprobe_dump_config": null, "_resolved_overrides": [["_deepseek_v4_overrides", {"attention_backend": "dsv4", "page_size": 256, "swa_full_tokens_ratio": 0.1}], ["_deepseek_v4_kv_cache_dtype", {"kv_cache_dtype": "fp8_e4m3"}], ["_deepseek_v4_sm120_moe", {"moe_runner_backend": "flashinfer_mxfp4"}], ["_sampling_backend_default", {"sampling_backend": "flashinfer"}], ["_data_parallelism_defaults", {"enable_dp_attention": false, "enable_dp_lm_head": false}], ["_speculative_moe_runner_default", {"speculative_moe_runner_backend": "flashinfer_mxfp4"}], ["DeepseekV4ForCausalLM.determine_num_fused_shared_experts", {"disable_shared_experts_fusion": true}]], "grpc_worker_threads": 4, "_quantization_explicitly_unset": false, "_cuda_graph_config_locked": [["decode", "max_bs"]], "_declarations_materialized": true, "_runtime_mutations": [["model_runner.chunked_prefix_cache_gate", {"disable_chunked_prefix_cache": true}], ["scheduler.pp_max_micro_batch_size_default", {"pp_max_micro_batch_size": 256}]], "_in_override": false, "_mx_config_cache": {}, "max_speculative_num_draft_tokens": null, "last_gen_throughput": 0.0, "memory_usage": {"weight": 56.87, "kvcache": 0, "token_capacity": 1282304, "graph": 1.39}, "effective_max_running_requests_per_dp": 256}], "version": "0.0.0.dev1+g35f2d4f76", "kv_events": null}, "duration": 15.920104521006579, "completed": 1, "total_input_tokens": 1024, "total_input_text_tokens": 1024, "total_input_vision_tokens": 0, "total_output_tokens": 1, "total_output_tokens_retokenized": 1, "request_throughput": 0.06281365795560576, "input_throughput": 64.3211857465403, "output_throughput": 0.06281365795560576, "total_throughput": 64.3839994044959, "mean_e2e_latency_ms": 15901.597731979564, "median_e2e_latency_ms": 15901.597731979564, "std_e2e_latency_ms": 0.0, "p90_e2e_latency_ms": 15901.597731979564, "p95_e2e_latency_ms": 15901.597731979564, "p99_e2e_latency_ms": 15901.597731979564, "mean_ttft_ms": 15901.53138898313, "median_ttft_ms": 15901.53138898313, "std_ttft_ms": 0.0, "p90_ttft_ms": 15901.53138898313, "p95_ttft_ms": 15901.53138898313, "p99_ttft_ms": 15901.53138898313, "mean_tpot_ms": 0.0, "median_tpot_ms": 0.0, "std_tpot_ms": 0.0, "p90_tpot_ms": 0.0, "p95_tpot_ms": 0.0, "p99_tpot_ms": 0.0, "mean_itl_ms": 0.0, "median_itl_ms": 0.0, "std_itl_ms": 0.0, "p90_itl_ms": 0.0, "p95_itl_ms": 0.0, "p99_itl_ms": 0.0, "concurrency": 0.9988375208842005, "accept_length": null, "max_output_tokens_per_s": 1.0, "max_concurrent_requests": 1, "input_lens": [1024], "output_lens": [1], "ttfts": [15.90153138898313], "itls": [[]], "generated_texts": [".\n"], "errors": [""]} diff --git a/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/old_semantics_1k_o128.jsonl b/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/old_semantics_1k_o128.jsonl new file mode 100644 index 0000000..89160ff --- /dev/null +++ b/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/old_semantics_1k_o128.jsonl @@ -0,0 +1 @@ +{"tag": null, "backend": "sglang", "dataset_name": "random", "request_rate": 10000.0, "max_concurrency": 1, "sharegpt_output_len": null, "random_input_len": 1024, "random_output_len": 128, "random_range_ratio": 1.0, "server_info": {"model_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_mode": "auto", "tokenizer_backend": "huggingface", "tokenizer_worker_num": 1, "detokenizer_worker_num": 1, "skip_tokenizer_init": false, "load_format": "auto", "model_loader_extra_config": "{}", "trust_remote_code": true, "context_length": null, "is_embedding": false, "enable_multimodal": null, "revision": null, "model_impl": "auto", "model_config_parser": "auto", "json_model_override_args": "{}", "dtype": "auto", "quantization": null, "quantization_param_path": null, "kv_cache_dtype": "fp8_e4m3", "enable_fp32_lm_head": false, "modelopt_quant": null, "modelopt_checkpoint_restore_path": null, "modelopt_checkpoint_save_path": null, "modelopt_export_path": null, "quantize_and_serve": false, "rl_quant_profile": null, "enable_tf32_matmul": false, "mem_fraction_static": 0.9, "max_running_requests": 256, "max_queued_requests": null, "max_total_tokens": null, "chunked_prefill_size": 8192, "enable_dynamic_chunking": false, "max_prefill_tokens": 16384, "prefill_max_requests": null, "schedule_policy": "fcfs", "enable_priority_scheduling": false, "disable_priority_preemption": false, "default_priority_value": null, "abort_on_priority_when_disabled": false, "schedule_low_priority_values_first": false, "priority_scheduling_preemption_threshold": 10, "retraction_policy": "length", "schedule_conservativeness": 1.0, "page_size": 256, "swa_full_tokens_ratio": 0.1, "disable_hybrid_swa_memory": false, "radix_eviction_policy": "lru", "prefill_only_disable_kv_cache": false, "disable_radix_cache": false, "enable_page_major_kv_layout": false, "enable_unified_memory": false, "disable_chunked_prefix_cache": false, "disable_overlap_schedule": false, "num_continuous_decode_steps": 1, "scheduler_recv_interval": 1, "enable_mixed_chunk": false, "nccl_port": null, "dist_timeout": null, "dist_init_addr": "10.101.0.11:20002", "nnodes": 2, "node_rank": 0, "tp_size": 16, "dcp_size": 1, "pp_size": 1, "pp_max_micro_batch_size": null, "pp_async_batch_depth": 0, "dp_size": 1, "load_balance_method": "round_robin", "attn_cp_size": 1, "moe_dp_size": 1, "enable_prefill_cp": false, "cp_strategy": null, "enable_dsa_cache_layer_split": false, "enable_dsa_prefill_context_parallel": false, "dsa_prefill_cp_mode": "round-robin-split", "enable_prefill_context_parallel": false, "prefill_cp_mode": "in-seq-split", "enable_dp_attention": false, "enable_dp_attention_local_control_broadcast": false, "enable_dp_lm_head": false, "enable_attn_tp_input_scattered": false, "disable_attn_tp_gather": false, "enable_p2p_check": false, "device": "cuda", "base_gpu_id": 0, "gpu_id_step": 1, "random_seed": 42600328, "watchdog_timeout": 300, "soft_watchdog_timeout": null, "sleep_on_idle": false, "use_ray": false, "custom_sigquit_handler": null, "numa_node": null, "gc_threshold": null, "host": "0.0.0.0", "port": 30002, "fastapi_root_path": "", "smg_grpc_mode": false, "grpc_mode": false, "grpc_port": null, "skip_server_warmup": false, "warmups": null, "enable_http2": false, "ssl_keyfile": null, "ssl_certfile": null, "ssl_ca_certs": null, "ssl_keyfile_password": null, "enable_ssl_refresh": false, "api_key": null, "admin_api_key": null, "served_model_name": "/data/hf_models/DeepSeek-V4-Pro", "weight_version": "default", "chat_template": null, "hf_chat_template_name": null, "completion_template": null, "file_storage_path": "sglang_storage", "enable_cache_report": false, "reasoning_parser": null, "default_chat_template_kwargs": null, "strip_thinking_cache": false, "enable_strict_thinking": false, "tool_call_parser": null, "tool_server": null, "sampling_defaults": "model", "asr_max_buffer_seconds": 60, "asr_max_concurrent_sessions": 32, "preferred_sampling_params": null, "allow_auto_truncate": false, "stream_interval": 1, "batch_notify_size": 16, "stream_response_default_include_usage": false, "incremental_streaming_output": false, "enable_streaming_session": false, "enable_session_radix_cache": false, "log_level": "info", "log_level_http": null, "log_requests": false, "log_requests_level": 2, "log_requests_format": "text", "log_requests_target": null, "uvicorn_access_log_exclude_prefixes": [], "crash_dump_folder": null, "show_time_cost": false, "enable_metrics": false, "smg_http_sidecar_port": null, "enable_mfu_metrics": false, "enable_metrics_for_all_schedulers": false, "load_snapshot_publish_interval": 15, "tokenizer_metrics_custom_labels_header": "x-custom-labels", "tokenizer_metrics_allowed_custom_labels": null, "extra_metric_labels": null, "bucket_time_to_first_token": null, "bucket_inter_token_latency": null, "bucket_e2e_request_latency": null, "prompt_tokens_buckets": null, "generation_tokens_buckets": null, "gc_warning_threshold_secs": 0.0, "decode_log_interval": 40, "enable_request_time_stats_logging": false, "kv_events_config": null, "enable_forward_pass_metrics": false, "forward_pass_metrics_worker_id": "", "forward_pass_metrics_ipc_name": null, "enable_trace": false, "trace_modules": "request", "otlp_traces_endpoint": "localhost:4317", "export_metrics_to_file": false, "export_metrics_to_file_dir": null, "stat_loggers": null, "constrained_json_whitespace_pattern": null, "constrained_json_disable_any_whitespace": false, "attention_backend": "dsv4", "decode_attention_backend": null, "prefill_attention_backend": null, "sampling_backend": "flashinfer", "grammar_backend": "xgrammar", "radix_cache_backend": null, "mm_attention_backend": null, "fp8_gemm_runner_backend": "auto", "fp4_gemm_runner_backend": "auto", "bf16_gemm_backend": "auto", "dsa_prefill_backend": null, "dsa_decode_backend": null, "dsa_paged_mqa_logits_backend": "auto", "dsa_topk_backend": "sgl-kernel", "disable_flashinfer_autotune": false, "mamba_backend": "triton", "cuda_graph_config": {"decode": {"backend": "full", "max_bs": 64, "bs": [1, 2, 4, 8, 12, 16, 24, 32, 40, 48, 56, 64], "tc_compiler": "eager", "full_prefill_max_req": null}, "prefill": {"backend": "disabled", "max_bs": 8192, "bs": [4, 8, 12, 16, 20, 24, 28, 32, 48, 64, 80, 96, 112, 128, 144, 160, 176, 192, 208, 224, 240, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 704, 768, 832, 896, 960, 1024, 1280, 1536, 1792, 2048, 2304, 2560, 2816, 3072, 3328, 3584, 3840, 4096, 4608, 5120, 5632, 6144, 6656, 7168, 7680, 8192], "tc_compiler": "eager", "full_prefill_max_req": null}}, "cuda_graph_backend_decode": null, "cuda_graph_backend_prefill": null, "cuda_graph_max_bs_decode": 64, "cuda_graph_max_bs_prefill": null, "cuda_graph_bs_decode": null, "cuda_graph_bs_prefill": null, "cuda_graph_tc_compiler": null, "disable_prefill_cuda_graph": false, "disable_decode_cuda_graph": false, "disable_cuda_graph": false, "disable_cuda_graph_padding": false, "enable_profile_cuda_graph": false, "enable_cudagraph_gc": false, "debug_cuda_graph": false, "enable_layerwise_nvtx_marker": false, "enable_nccl_nvls": false, "enable_symm_mem": false, "triton_attention_reduce_in_fp32": false, "triton_attention_num_kv_splits": 8, "triton_attention_split_tile_size": null, "flashinfer_mla_disable_ragged": false, "enable_fused_qk_norm_rope": false, "enable_precise_embedding_interpolation": false, "enable_fused_moe_sum_all_reduce": false, "enable_deepseek_v4_fp4_indexer": false, "disable_custom_all_reduce": false, "enable_mscclpp": false, "enable_torch_symm_mem": false, "pre_warm_nccl": false, "enable_quant_communications": false, "enable_flashinfer_allreduce_fusion": false, "enforce_disable_flashinfer_allreduce_fusion": false, "flashinfer_allreduce_fusion_backend": null, "enable_aiter_allreduce_fusion": false, "enable_torch_compile": false, "enable_torch_compile_debug_mode": false, "torch_compile_max_bs": 32, "torchao_config": "", "speculative_algorithm": null, "speculative_draft_model_path": null, "speculative_draft_model_revision": null, "speculative_draft_load_format": null, "speculative_num_steps": null, "speculative_eagle_topk": null, "speculative_num_draft_tokens": null, "speculative_dflash_block_size": null, "speculative_dspark_block_size": null, "speculative_dspark_sps_table_path": null, "speculative_dspark_confidence_sts_path": null, "speculative_dspark_align_verify_tokens_to_graph_tier": false, "speculative_accept_threshold_single": 1.0, "speculative_accept_threshold_acc": 1.0, "speculative_use_rejection_sampling": false, "speculative_token_map": null, "speculative_attention_mode": "prefill", "speculative_draft_attention_backend": null, "speculative_draft_window_size": null, "speculative_moe_runner_backend": "flashinfer_mxfp4", "speculative_moe_a2a_backend": null, "speculative_draft_model_quantization": null, "speculative_skip_dp_mlp_sync": false, "enable_multi_layer_eagle": false, "speculative_adaptive": false, "speculative_adaptive_config": null, "decoupled_spec_bind_endpoint": null, "decoupled_spec_connect_endpoints": null, "decoupled_spec_rank": null, "decoupled_spec_role": "null", "spec_trace_dir": null, "speculative_ngram_min_bfs_breadth": 1, "speculative_ngram_max_bfs_breadth": 10, "speculative_ngram_match_type": "BFS", "speculative_ngram_max_trie_depth": 18, "speculative_ngram_capacity": 10000000, "speculative_ngram_external_corpus_path": null, "speculative_ngram_external_sam_budget": 0, "speculative_ngram_external_corpus_max_tokens": 10000000, "ep_size": 2, "moe_a2a_backend": "none", "moe_runner_backend": "flashinfer_mxfp4", "flashinfer_mxfp4_moe_precision": "default", "deepep_mode": "auto", "fuseep_mode": 2, "deepep_dispatcher_output_dtype": "auto", "ep_num_redundant_experts": 0, "ep_dispatch_algorithm": null, "init_expert_location": "trivial", "enable_eplb": false, "eplb_algorithm": "auto", "eplb_rebalance_num_iterations": 1000, "eplb_rebalance_layers_per_chunk": null, "eplb_min_rebalancing_utilization_threshold": 1.0, "expert_distribution_recorder_mode": null, "expert_distribution_recorder_buffer_size": 1000, "enable_expert_distribution_metrics": false, "deepep_config": null, "moe_dense_tp_size": null, "elastic_ep_backend": null, "enable_elastic_expert_backup": false, "mooncake_ib_device": null, "enable_waterfill": false, "ep_join_mode": null, "ep_join_rank_offset": 0, "elastic_ep_initial_size": null, "max_ep_size": null, "elastic_ep_scale_timeout": 600, "elastic_ep_rejoin": false, "disable_flashinfer_cutlass_moe_fp4_allgather": false, "disable_shared_experts_fusion": false, "enforce_shared_experts_fusion": false, "max_mamba_cache_size": null, "mamba_ssm_dtype": null, "enable_mamba_cache_stochastic_rounding": false, "mamba_cache_philox_rounds": 0, "mamba_full_memory_ratio": 0.9, "mamba_radix_cache_strategy": "auto", "uses_mamba_radix_cache": false, "mamba_track_interval": 256, "enable_int8_mamba_checkpoint": false, "int8_mamba_ckpt_size": null, "linear_attn_backend": "triton", "linear_attn_decode_backend": null, "linear_attn_prefill_backend": null, "enable_linear_replayssm": false, "linear_replayssm_cache_len": 16, "enable_hierarchical_cache": false, "hicache_ratio": 2.0, "hicache_size": 0, "hicache_write_policy": "write_through", "hicache_io_backend": "kernel", "hicache_mem_layout": "page_first", "hicache_storage_backend": null, "hicache_storage_prefetch_policy": "timeout", "hicache_storage_backend_extra_config": null, "enable_hisparse": false, "hisparse_config": null, "enable_broadcast_mm_inputs_process": false, "enable_prefix_mm_cache": false, "mm_enable_dp_encoder": false, "mm_process_config": {}, "limit_mm_data_per_request": null, "enable_mm_global_cache": false, "disable_fast_image_processor": false, "mm_feature_transport": "cpu", "keep_mm_feature_on_device": false, "enable_lora": null, "enable_lora_overlap_loading": null, "max_lora_rank": null, "lora_target_modules": null, "lora_paths": null, "max_loaded_loras": null, "max_loras_per_batch": 8, "lora_eviction_policy": "lru", "lora_backend": "csgmv", "max_lora_chunk_size": 16, "experts_shared_outer_loras": null, "lora_use_virtual_experts": false, "lora_strict_loading": false, "lora_drain_wait_threshold": 0.0, "enable_two_batch_overlap": false, "enable_single_batch_overlap": false, "tbo_token_distribution_threshold": 0.48, "cpu_offload_gb": 0, "offload_group_size": -1, "offload_num_in_group": 1, "offload_prefetch_step": 1, "offload_mode": "cpu", "enable_lmcache": false, "lmcache_config_file": null, "enable_flexkv": false, "flexkv_config_file": null, "kt_weight_path": null, "kt_method": "AMXINT4", "kt_cpuinfer": null, "kt_threadpool_count": 2, "kt_num_gpu_experts": null, "kt_max_deferred_experts_per_token": null, "dllm_algorithm": null, "dllm_algorithm_config": null, "dllm_fdfo": true, "disaggregation_mode": "null", "disaggregation_transfer_backend": "mooncake", "disaggregation_bootstrap_port": 8998, "disaggregation_ib_device": null, "disaggregation_decode_enable_radix_cache": false, "disaggregation_decode_enable_offload_kvcache": false, "num_reserved_decode_tokens": 512, "disaggregation_decode_extra_slots": null, "disaggregation_decode_polling_interval": 1, "optimistic_prefill_attempts": 0, "encoder_only": false, "language_only": false, "encoder_transfer_backend": "zmq_to_scheduler", "encoder_urls": [], "encoder_bootstrap_port": 8997, "encoder_register_urls": [], "enable_adaptive_dispatch_to_encoder": false, "enable_pdmux": false, "pdmux_config_path": null, "sm_group_num": 8, "custom_weight_loader": [], "weight_loader_disable_mmap": false, "weight_loader_prefetch_checkpoints": false, "weight_loader_prefetch_num_threads": 4, "weight_loader_drop_cache_after_load": false, "remote_instance_weight_loader_seed_instance_ip": null, "remote_instance_weight_loader_seed_instance_service_port": null, "remote_instance_weight_loader_send_weights_group_ports": null, "remote_instance_weight_loader_backend": "nccl", "remote_instance_weight_loader_start_seed_via_transfer_engine": false, "engine_info_bootstrap_port": 6789, "modelexpress_config": null, "download_dir": null, "model_checksum": null, "delete_ckpt_after_loading": false, "decrypted_config_file": null, "decrypted_draft_config_file": null, "checkpoint_engine_wait_weights_before_ready": false, "enable_prefill_delayer": false, "prefill_delayer_max_delay_passes": 30, "prefill_delayer_token_usage_low_watermark": null, "prefill_delayer_forward_passes_buckets": null, "prefill_delayer_wait_seconds_buckets": null, "prefill_delayer_queue_min_ratio": null, "prefill_delayer_max_delay_ms": null, "min_free_slots_delay": null, "enable_deterministic_inference": false, "rl_on_policy_target": null, "kv_canary": "none", "kv_canary_real_data": "none", "kv_canary_sweep_interval": 0, "enable_dynamic_batch_tokenizer": false, "dynamic_batch_tokenizer_batch_size": 32, "dynamic_batch_tokenizer_batch_timeout": 0.002, "enable_tokenizer_batch_encode": false, "disable_tokenizer_batch_decode": false, "debug_tensor_dump_output_folder": null, "debug_tensor_dump_layers": null, "debug_tensor_dump_input_file": null, "enable_memory_saver": false, "enable_weights_cpu_backup": false, "enable_draft_weights_cpu_backup": false, "enable_custom_logit_processor": false, "enable_return_hidden_states": false, "enable_return_routed_experts": false, "enable_return_indexer_topk": false, "disable_outlines_disk_cache": false, "enable_mis": false, "forward_hooks": null, "msprobe_dump_config": null, "status": "ready", "max_total_num_tokens": 1282304, "max_req_input_len": 1048570, "internal_states": [{"model_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_path": "/data/hf_models/DeepSeek-V4-Pro", "tokenizer_mode": "auto", "tokenizer_backend": "huggingface", "tokenizer_worker_num": 1, "detokenizer_worker_num": 1, "skip_tokenizer_init": false, "load_format": "auto", "model_loader_extra_config": "{}", "trust_remote_code": true, "context_length": null, "is_embedding": false, "enable_multimodal": null, "revision": null, "model_impl": "auto", "model_config_parser": "auto", "json_model_override_args": "{}", "dtype": "auto", "quantization": null, "quantization_param_path": null, "kv_cache_dtype": "fp8_e4m3", "enable_fp32_lm_head": false, "modelopt_quant": null, "modelopt_checkpoint_restore_path": null, "modelopt_checkpoint_save_path": null, "modelopt_export_path": null, "quantize_and_serve": false, "rl_quant_profile": null, "enable_tf32_matmul": false, "mem_fraction_static": 0.9, "max_running_requests": 256, "max_queued_requests": null, "max_total_tokens": null, "chunked_prefill_size": 8192, "enable_dynamic_chunking": false, "max_prefill_tokens": 16384, "prefill_max_requests": null, "schedule_policy": "fcfs", "enable_priority_scheduling": false, "disable_priority_preemption": false, "default_priority_value": null, "abort_on_priority_when_disabled": false, "schedule_low_priority_values_first": false, "priority_scheduling_preemption_threshold": 10, "retraction_policy": "length", "schedule_conservativeness": 1.0, "page_size": 256, "swa_full_tokens_ratio": 0.1, "disable_hybrid_swa_memory": false, "radix_eviction_policy": "lru", "prefill_only_disable_kv_cache": false, "disable_radix_cache": false, "enable_page_major_kv_layout": false, "enable_unified_memory": false, "disable_chunked_prefix_cache": true, "disable_overlap_schedule": false, "num_continuous_decode_steps": 1, "scheduler_recv_interval": 1, "enable_mixed_chunk": false, "nccl_port": null, "dist_timeout": null, "dist_init_addr": "10.101.0.11:20002", "nnodes": 2, "node_rank": 0, "tp_size": 16, "dcp_size": 1, "pp_size": 1, "pp_max_micro_batch_size": 256, "pp_async_batch_depth": 0, "dp_size": 1, "load_balance_method": "round_robin", "attn_cp_size": 1, "moe_dp_size": 1, "enable_prefill_cp": false, "cp_strategy": null, "enable_dsa_cache_layer_split": false, "enable_dsa_prefill_context_parallel": false, "dsa_prefill_cp_mode": "round-robin-split", "enable_prefill_context_parallel": false, "prefill_cp_mode": "in-seq-split", "enable_dp_attention": false, "enable_dp_attention_local_control_broadcast": false, "enable_dp_lm_head": false, "enable_attn_tp_input_scattered": false, "disable_attn_tp_gather": false, "enable_p2p_check": false, "device": "cuda", "base_gpu_id": 0, "gpu_id_step": 1, "random_seed": 42600328, "watchdog_timeout": 300, "soft_watchdog_timeout": null, "sleep_on_idle": false, "use_ray": false, "numa_node": null, "gc_threshold": null, "host": "0.0.0.0", "port": 30002, "fastapi_root_path": "", "smg_grpc_mode": false, "grpc_mode": false, "grpc_port": null, "skip_server_warmup": false, "warmups": null, "enable_http2": false, "ssl_keyfile": null, "ssl_certfile": null, "ssl_ca_certs": null, "ssl_keyfile_password": null, "enable_ssl_refresh": false, "api_key": null, "admin_api_key": null, "served_model_name": "/data/hf_models/DeepSeek-V4-Pro", "weight_version": "default", "chat_template": null, "hf_chat_template_name": null, "completion_template": null, "file_storage_path": "sglang_storage", "enable_cache_report": false, "reasoning_parser": null, "default_chat_template_kwargs": null, "strip_thinking_cache": false, "enable_strict_thinking": false, "tool_call_parser": null, "tool_server": null, "sampling_defaults": "model", "asr_max_buffer_seconds": 60, "asr_max_concurrent_sessions": 32, "preferred_sampling_params": null, "allow_auto_truncate": false, "stream_interval": 1, "batch_notify_size": 16, "stream_response_default_include_usage": false, "incremental_streaming_output": false, "enable_streaming_session": false, "enable_session_radix_cache": false, "log_level": "info", "log_level_http": null, "log_requests": false, "log_requests_level": 2, "log_requests_format": "text", "log_requests_target": null, "uvicorn_access_log_exclude_prefixes": [], "crash_dump_folder": null, "show_time_cost": false, "enable_metrics": false, "smg_http_sidecar_port": null, "enable_mfu_metrics": false, "enable_metrics_for_all_schedulers": false, "load_snapshot_publish_interval": 15, "tokenizer_metrics_custom_labels_header": "x-custom-labels", "tokenizer_metrics_allowed_custom_labels": null, "extra_metric_labels": null, "bucket_time_to_first_token": null, "bucket_inter_token_latency": null, "bucket_e2e_request_latency": null, "prompt_tokens_buckets": null, "generation_tokens_buckets": null, "gc_warning_threshold_secs": 0.0, "decode_log_interval": 40, "enable_request_time_stats_logging": false, "kv_events_config": null, "enable_forward_pass_metrics": false, "forward_pass_metrics_worker_id": "", "forward_pass_metrics_ipc_name": null, "enable_trace": false, "trace_modules": "request", "otlp_traces_endpoint": "localhost:4317", "export_metrics_to_file": false, "export_metrics_to_file_dir": null, "stat_loggers": null, "constrained_json_whitespace_pattern": null, "constrained_json_disable_any_whitespace": false, "attention_backend": "dsv4", "decode_attention_backend": null, "prefill_attention_backend": null, "sampling_backend": "flashinfer", "grammar_backend": "xgrammar", "radix_cache_backend": null, "mm_attention_backend": null, "fp8_gemm_runner_backend": "auto", "fp4_gemm_runner_backend": "auto", "bf16_gemm_backend": "auto", "dsa_prefill_backend": null, "dsa_decode_backend": null, "dsa_paged_mqa_logits_backend": "auto", "dsa_topk_backend": "sgl-kernel", "disable_flashinfer_autotune": false, "mamba_backend": "triton", "cuda_graph_config": {"decode": {"backend": "full", "max_bs": 64, "bs": [1, 2, 4, 8, 12, 16, 24, 32, 40, 48, 56, 64], "tc_compiler": "eager", "full_prefill_max_req": null}, "prefill": {"backend": "disabled", "max_bs": 8192, "bs": [4, 8, 12, 16, 20, 24, 28, 32, 48, 64, 80, 96, 112, 128, 144, 160, 176, 192, 208, 224, 240, 256, 288, 320, 352, 384, 416, 448, 480, 512, 576, 640, 704, 768, 832, 896, 960, 1024, 1280, 1536, 1792, 2048, 2304, 2560, 2816, 3072, 3328, 3584, 3840, 4096, 4608, 5120, 5632, 6144, 6656, 7168, 7680, 8192], "tc_compiler": "eager", "full_prefill_max_req": null}}, "cuda_graph_backend_decode": null, "cuda_graph_backend_prefill": null, "cuda_graph_max_bs_decode": 64, "cuda_graph_max_bs_prefill": null, "cuda_graph_bs_decode": null, "cuda_graph_bs_prefill": null, "cuda_graph_tc_compiler": null, "disable_prefill_cuda_graph": false, "disable_decode_cuda_graph": false, "disable_cuda_graph": false, "disable_cuda_graph_padding": false, "enable_profile_cuda_graph": false, "enable_cudagraph_gc": false, "debug_cuda_graph": false, "enable_layerwise_nvtx_marker": false, "enable_nccl_nvls": false, "enable_symm_mem": false, "triton_attention_reduce_in_fp32": false, "triton_attention_num_kv_splits": 8, "triton_attention_split_tile_size": null, "flashinfer_mla_disable_ragged": false, "enable_fused_qk_norm_rope": false, "enable_precise_embedding_interpolation": false, "enable_fused_moe_sum_all_reduce": false, "enable_deepseek_v4_fp4_indexer": false, "disable_custom_all_reduce": false, "enable_mscclpp": false, "enable_torch_symm_mem": false, "pre_warm_nccl": false, "enable_quant_communications": false, "enable_flashinfer_allreduce_fusion": false, "enforce_disable_flashinfer_allreduce_fusion": false, "flashinfer_allreduce_fusion_backend": null, "enable_aiter_allreduce_fusion": false, "enable_torch_compile": false, "enable_torch_compile_debug_mode": false, "torch_compile_max_bs": 32, "torchao_config": "", "speculative_algorithm": null, "speculative_draft_model_path": null, "speculative_draft_model_revision": null, "speculative_draft_load_format": null, "speculative_num_steps": null, "speculative_eagle_topk": null, "speculative_num_draft_tokens": null, "speculative_dflash_block_size": null, "speculative_dspark_block_size": null, "speculative_dspark_sps_table_path": null, "speculative_dspark_confidence_sts_path": null, "speculative_dspark_align_verify_tokens_to_graph_tier": false, "speculative_accept_threshold_single": 1.0, "speculative_accept_threshold_acc": 1.0, "speculative_use_rejection_sampling": false, "speculative_token_map": null, "speculative_attention_mode": "prefill", "speculative_draft_attention_backend": null, "speculative_draft_window_size": null, "speculative_moe_runner_backend": "flashinfer_mxfp4", "speculative_moe_a2a_backend": null, "speculative_draft_model_quantization": null, "speculative_skip_dp_mlp_sync": false, "enable_multi_layer_eagle": false, "speculative_adaptive": false, "speculative_adaptive_config": null, "decoupled_spec_bind_endpoint": null, "decoupled_spec_connect_endpoints": null, "decoupled_spec_rank": null, "decoupled_spec_role": "null", "spec_trace_dir": null, "speculative_ngram_min_bfs_breadth": 1, "speculative_ngram_max_bfs_breadth": 10, "speculative_ngram_match_type": "BFS", "speculative_ngram_max_trie_depth": 18, "speculative_ngram_capacity": 10000000, "speculative_ngram_external_corpus_path": null, "speculative_ngram_external_sam_budget": 0, "speculative_ngram_external_corpus_max_tokens": 10000000, "ep_size": 2, "moe_a2a_backend": "none", "moe_runner_backend": "flashinfer_mxfp4", "flashinfer_mxfp4_moe_precision": "default", "deepep_mode": "auto", "fuseep_mode": 2, "deepep_dispatcher_output_dtype": "auto", "ep_num_redundant_experts": 0, "ep_dispatch_algorithm": null, "init_expert_location": "trivial", "enable_eplb": false, "eplb_algorithm": "auto", "eplb_rebalance_num_iterations": 1000, "eplb_rebalance_layers_per_chunk": null, "eplb_min_rebalancing_utilization_threshold": 1.0, "expert_distribution_recorder_mode": null, "expert_distribution_recorder_buffer_size": 1000, "enable_expert_distribution_metrics": false, "deepep_config": null, "moe_dense_tp_size": null, "elastic_ep_backend": null, "enable_elastic_expert_backup": false, "mooncake_ib_device": null, "enable_waterfill": false, "ep_join_mode": null, "ep_join_rank_offset": 0, "elastic_ep_initial_size": null, "max_ep_size": null, "elastic_ep_scale_timeout": 600, "elastic_ep_rejoin": false, "disable_flashinfer_cutlass_moe_fp4_allgather": false, "disable_shared_experts_fusion": true, "enforce_shared_experts_fusion": false, "max_mamba_cache_size": null, "mamba_ssm_dtype": null, "enable_mamba_cache_stochastic_rounding": false, "mamba_cache_philox_rounds": 0, "mamba_full_memory_ratio": 0.9, "mamba_radix_cache_strategy": "auto", "uses_mamba_radix_cache": false, "mamba_track_interval": 256, "enable_int8_mamba_checkpoint": false, "int8_mamba_ckpt_size": null, "linear_attn_backend": "triton", "linear_attn_decode_backend": null, "linear_attn_prefill_backend": null, "enable_linear_replayssm": false, "linear_replayssm_cache_len": 16, "enable_hierarchical_cache": false, "hicache_ratio": 2.0, "hicache_size": 0, "hicache_write_policy": "write_through", "hicache_io_backend": "kernel", "hicache_mem_layout": "page_first", "hicache_storage_backend": null, "hicache_storage_prefetch_policy": "timeout", "hicache_storage_backend_extra_config": null, "enable_hisparse": false, "hisparse_config": null, "enable_broadcast_mm_inputs_process": false, "enable_prefix_mm_cache": false, "mm_enable_dp_encoder": false, "mm_process_config": {}, "limit_mm_data_per_request": null, "enable_mm_global_cache": false, "disable_fast_image_processor": false, "mm_feature_transport": "cpu", "keep_mm_feature_on_device": false, "enable_lora": null, "enable_lora_overlap_loading": null, "max_lora_rank": null, "lora_target_modules": null, "lora_paths": null, "max_loaded_loras": null, "max_loras_per_batch": 8, "lora_eviction_policy": "lru", "lora_backend": "csgmv", "max_lora_chunk_size": 16, "experts_shared_outer_loras": null, "lora_use_virtual_experts": false, "lora_strict_loading": false, "lora_drain_wait_threshold": 0.0, "enable_two_batch_overlap": false, "enable_single_batch_overlap": false, "tbo_token_distribution_threshold": 0.48, "cpu_offload_gb": 0, "offload_group_size": -1, "offload_num_in_group": 1, "offload_prefetch_step": 1, "offload_mode": "cpu", "enable_lmcache": false, "lmcache_config_file": null, "enable_flexkv": false, "flexkv_config_file": null, "kt_weight_path": null, "kt_method": "AMXINT4", "kt_cpuinfer": null, "kt_threadpool_count": 2, "kt_num_gpu_experts": null, "kt_max_deferred_experts_per_token": null, "dllm_algorithm": null, "dllm_algorithm_config": null, "dllm_fdfo": true, "disaggregation_mode": "null", "disaggregation_transfer_backend": "mooncake", "disaggregation_bootstrap_port": 8998, "disaggregation_ib_device": null, "disaggregation_decode_enable_radix_cache": false, "disaggregation_decode_enable_offload_kvcache": false, "num_reserved_decode_tokens": 512, "disaggregation_decode_extra_slots": null, "disaggregation_decode_polling_interval": 1, "optimistic_prefill_attempts": 0, "encoder_only": false, "language_only": false, "encoder_transfer_backend": "zmq_to_scheduler", "encoder_urls": [], "encoder_bootstrap_port": 8997, "encoder_register_urls": [], "enable_adaptive_dispatch_to_encoder": false, "enable_pdmux": false, "pdmux_config_path": null, "sm_group_num": 8, "custom_weight_loader": [], "weight_loader_disable_mmap": false, "weight_loader_prefetch_checkpoints": false, "weight_loader_prefetch_num_threads": 4, "weight_loader_drop_cache_after_load": false, "remote_instance_weight_loader_seed_instance_ip": null, "remote_instance_weight_loader_seed_instance_service_port": null, "remote_instance_weight_loader_send_weights_group_ports": null, "remote_instance_weight_loader_backend": "nccl", "remote_instance_weight_loader_start_seed_via_transfer_engine": false, "engine_info_bootstrap_port": 6789, "modelexpress_config": null, "download_dir": null, "model_checksum": null, "delete_ckpt_after_loading": false, "decrypted_config_file": null, "decrypted_draft_config_file": null, "checkpoint_engine_wait_weights_before_ready": false, "enable_prefill_delayer": false, "prefill_delayer_max_delay_passes": 30, "prefill_delayer_token_usage_low_watermark": null, "prefill_delayer_forward_passes_buckets": null, "prefill_delayer_wait_seconds_buckets": null, "prefill_delayer_queue_min_ratio": null, "prefill_delayer_max_delay_ms": null, "min_free_slots_delay": null, "enable_deterministic_inference": false, "rl_on_policy_target": null, "kv_canary": "none", "kv_canary_real_data": "none", "kv_canary_sweep_interval": 0, "enable_dynamic_batch_tokenizer": false, "dynamic_batch_tokenizer_batch_size": 32, "dynamic_batch_tokenizer_batch_timeout": 0.002, "enable_tokenizer_batch_encode": false, "disable_tokenizer_batch_decode": false, "debug_tensor_dump_output_folder": null, "debug_tensor_dump_layers": null, "debug_tensor_dump_input_file": null, "enable_memory_saver": false, "enable_weights_cpu_backup": false, "enable_draft_weights_cpu_backup": false, "enable_custom_logit_processor": false, "enable_return_hidden_states": false, "enable_return_routed_experts": false, "enable_return_indexer_topk": false, "disable_outlines_disk_cache": false, "enable_mis": false, "forward_hooks": null, "msprobe_dump_config": null, "_resolved_overrides": [["_deepseek_v4_overrides", {"attention_backend": "dsv4", "page_size": 256, "swa_full_tokens_ratio": 0.1}], ["_deepseek_v4_kv_cache_dtype", {"kv_cache_dtype": "fp8_e4m3"}], ["_deepseek_v4_sm120_moe", {"moe_runner_backend": "flashinfer_mxfp4"}], ["_sampling_backend_default", {"sampling_backend": "flashinfer"}], ["_data_parallelism_defaults", {"enable_dp_attention": false, "enable_dp_lm_head": false}], ["_speculative_moe_runner_default", {"speculative_moe_runner_backend": "flashinfer_mxfp4"}], ["DeepseekV4ForCausalLM.determine_num_fused_shared_experts", {"disable_shared_experts_fusion": true}]], "grpc_worker_threads": 4, "_quantization_explicitly_unset": false, "_cuda_graph_config_locked": [["decode", "max_bs"]], "_declarations_materialized": true, "_runtime_mutations": [["model_runner.chunked_prefix_cache_gate", {"disable_chunked_prefix_cache": true}], ["scheduler.pp_max_micro_batch_size_default", {"pp_max_micro_batch_size": 256}]], "_in_override": false, "_mx_config_cache": {}, "max_speculative_num_draft_tokens": null, "last_gen_throughput": 9.846365243044302, "memory_usage": {"weight": 56.87, "kvcache": 0, "token_capacity": 1282304, "graph": 1.39}, "effective_max_running_requests_per_dp": 256}], "version": "0.0.0.dev1+g35f2d4f76", "kv_events": null}, "duration": 273.49428128101863, "completed": 10, "total_input_tokens": 10240, "total_input_text_tokens": 10240, "total_input_vision_tokens": 0, "total_output_tokens": 1280, "total_output_tokens_retokenized": 1280, "request_throughput": 0.03656383582560135, "input_throughput": 37.44136788541578, "output_throughput": 4.680170985676972, "total_throughput": 42.121538871092746, "mean_e2e_latency_ms": 27338.507212593686, "median_e2e_latency_ms": 28447.550881988718, "std_e2e_latency_ms": 3473.642304320485, "p90_e2e_latency_ms": 28654.287753000972, "p95_e2e_latency_ms": 28700.00253299804, "p99_e2e_latency_ms": 28736.574356995698, "mean_ttft_ms": 14603.997771098511, "median_ttft_ms": 15807.843642003718, "std_ttft_ms": 3494.201172253669, "p90_ttft_ms": 15862.689939298434, "p95_ttft_ms": 15968.008152151015, "p99_ttft_ms": 16052.262722433077, "mean_tpot_ms": 100.27172788578878, "median_tpot_ms": 100.33895750392566, "std_tpot_ms": 0.775659911769669, "p90_tpot_ms": 101.16932770066829, "p95_tpot_ms": 101.51969660227059, "p99_tpot_ms": 101.79999172355242, "mean_itl_ms": 100.27146722679576, "median_itl_ms": 98.79106949665584, "std_itl_ms": 4.7299417321747175, "p90_itl_ms": 104.86746651586145, "p95_itl_ms": 105.05963615578366, "p99_itl_ms": 105.28184096270707, "concurrency": 0.9996006894382937, "accept_length": null, "max_output_tokens_per_s": 11.0, "max_concurrent_requests": 2, "input_lens": [1024, 1024, 1024, 1024, 1024, 1024, 1024, 1024, 1024, 1024], "output_lens": [128, 128, 128, 128, 128, 128, 128, 128, 128, 128], "ttfts": [4.132361830008449, 15.810174799989909, 15.697685002000071, 15.81987676298013, 15.805512484017527, 15.818076823983574, 16.073326365003595, 15.41570420700009, 15.627973544003908, 15.83928589199786], "itls": [[0.005895526002859697, 0.09184264999930747, 0.0905807469971478, 0.08934184300596826, 0.08962709500337951, 0.09203346498543397, 0.09452450001845136, 0.09404543798882514, 0.09188811900094151, 0.08820819100947119, 0.09758321900153533, 0.09283551198313944, 0.09019699599593878, 0.10116673400625587, 0.1047502470028121, 0.09989263699389994, 0.09856975599541329, 0.09862496302230284, 0.09874825199949555, 0.10335120998206548, 0.1051614600000903, 0.10456484599853866, 0.10466182001982816, 0.10525612300261855, 0.10505845199804753, 0.10496614899602719, 0.10001783899497241, 0.098550491995411, 0.09858029300812632, 0.098494470003061, 0.09734462699270807, 0.10309396899538115, 0.1050421750114765, 0.10062093299347907, 0.09850594200543128, 0.09832685498986393, 0.09725710499333218, 0.09866891399724409, 0.09839478202047758, 0.09927178698126227, 0.10312635701848194, 0.10562063800171018, 0.10036712198052555, 0.09856649799621664, 0.09874673502054065, 0.0979097489907872, 0.09886800099047832, 0.09714001801330596, 0.09870791199500673, 0.0983775409986265, 0.09873817701009102, 0.09822996298316866, 0.09857382101472467, 0.09834628098178655, 0.09867494000354782, 0.09855141301522963, 0.09701535399653949, 0.1031628219934646, 0.10498049401212484, 0.10474682398489676, 0.1046273999963887, 0.10485462902579457, 0.1046858919726219, 0.10507487002178095, 0.1049209049961064, 0.10452603700105101, 0.10506548700504936, 0.10453345597488806, 0.10470074901240878, 0.1045293560018763, 0.10494391300017014, 0.10490777000086382, 0.09555419100797735, 0.1032436300010886, 0.10449913199408911, 0.10475476598367095, 0.10464239201974124, 0.10530055098934099, 0.09575038199545816, 0.10268673501559533, 0.10500068299006671, 0.10480966500472277, 0.10488039898336865, 0.10513869501301087, 0.10503432899713516, 0.10481934100971557, 0.10493227498955093, 0.09973146699485369, 0.10392721102107316, 0.10498935397481546, 0.10470442200312391, 0.10480545100290328, 0.10456072201486677, 0.10499170797993429, 0.09965743700740859, 0.10343534199637361, 0.10470043099485338, 0.10493723102263175, 0.10431170399533585, 0.10466019698651507, 0.10463940401677974, 0.1047036929812748, 0.10453721601516008, 0.10454425698844716, 0.10465118300635368, 0.10465442200074904, 0.10467273500398733, 0.10462914500385523, 0.10459603599156253, 0.10472894299891777, 0.1045979369955603, 0.10463737501413561, 0.10448804299812764, 0.10448072798317298, 0.10461572502390482, 0.10458446698612534, 0.1045962460048031, 0.10463152098236606, 0.104503578011645, 0.1048987320100423, 0.09958976198686287, 0.09857337101129815, 0.09809273498831317, 0.09814273100346327, 0.10345398899517022, 0.10482552601024508, 0.10559170501073822], [0.09566077499766834, 0.09706300401012413, 0.09805185600998811, 0.09797229099785909, 0.09804248900036328, 0.09800614599953406, 0.09801720597897656, 0.09800611800164916, 0.0979618530254811, 0.09762436797609553, 0.09724768000887707, 0.09767311598989181, 0.09755259699886665, 0.09751047802274115, 0.097512665000977, 0.09809883398702368, 0.09812291001435369, 0.09735592399374582, 0.0984694279904943, 0.09790151799097657, 0.09826670002075844, 0.09846765900147147, 0.09832194200134836, 0.09859416299150325, 0.09811433899449185, 0.10414457501610741, 0.09927439197781496, 0.10314757999731228, 0.10490117201698013, 0.0999262830009684, 0.09831794700585306, 0.09794618000159971, 0.09821005197591148, 0.09781803301302716, 0.0979874010081403, 0.09837616898585111, 0.0984045020013582, 0.09838966198731214, 0.09851089701987803, 0.09814123500837013, 0.10401161899790168, 0.10495471398462541, 0.10497254901565611, 0.10457905297516845, 0.10468017900711857, 0.10456651300773956, 0.10449341399362311, 0.10468495800159872, 0.10460407400387339, 0.10416581499157473, 0.10408167500281706, 0.10404472699156031, 0.10402063000947237, 0.1039996390172746, 0.10428568799397908, 0.10484440499567427, 0.10432366200257093, 0.10457902800408192, 0.10461399497580715, 0.10453929399955086, 0.10471866402076557, 0.10449358899495564, 0.10466401401208714, 0.10452764498768374, 0.10444319399539381, 0.10456798499217257, 0.10458149900659919, 0.10445019201142713, 0.10428608098300174, 0.10396849201060832, 0.10407085600309074, 0.10425224099890329, 0.10441679798532277, 0.1046632710203994, 0.104458062996855, 0.10463773299125023, 0.10452940300456248, 0.10499734999029897, 0.10496475800755434, 0.10507140099070966, 0.10510074900230393, 0.10489614401012659, 0.10087166598532349, 0.09824154002126306, 0.0973612199886702, 0.09833363999496214, 0.09688299099798314, 0.0976814340101555, 0.09750224999152124, 0.09724336801446043, 0.09741399600170553, 0.09747777399024926, 0.0976079789979849, 0.09799236399703659, 0.09790476801572368, 0.09809459900134243, 0.09758041999884881, 0.09801481699105352, 0.09737528400728479, 0.09851063700625673, 0.09796331997495145, 0.0981016130244825, 0.09903215797385201, 0.09791936600231566, 0.0976699550228659, 0.09748803498223424, 0.09846093601663597, 0.0978986449772492, 0.09773459102143534, 0.09829353698296472, 0.09750918901409023, 0.09769528298056684, 0.09698114701313898, 0.0980593379936181, 0.09853416500845924, 0.0982791270071175, 0.09809705498628318, 0.09892403101548553, 0.09681231298600323, 0.09796741700847633, 0.09787912698811851, 0.09843747000559233, 0.09862083199550398, 0.09828056700644083, 0.09848320001037791, 0.09803579398430884, 0.10497526099788956], [0.09595299998181872, 0.09644269099226221, 0.09392984400619753, 0.08971445402130485, 0.08886999698006548, 0.09719691899954341, 0.09381745400605723, 0.09076645100140013, 0.09118014501291327, 0.09737935898010619, 0.09264763901592232, 0.09054686498711817, 0.0978932790167164, 0.10356597899226472, 0.10478006300400011, 0.10465747598209418, 0.09997026901692152, 0.10322996499598958, 0.10479494198807515, 0.10480312901199795, 0.10063348599942401, 0.0982742740015965, 0.09707434399751946, 0.09858113899827003, 0.09832178900251165, 0.09815541698480956, 0.09811370001989417, 0.09778310099500231, 0.09754805898410268, 0.09869058002368547, 0.09779237597831525, 0.09713903101510368, 0.10319686998263933, 0.10511687802500091, 0.10501791798742488, 0.10476022699731402, 0.10053536601481028, 0.09845105599379167, 0.09800734597956762, 0.09846050100168213, 0.09764336101943627, 0.09853661799570546, 0.09868175798328593, 0.09849266600213014, 0.09841594099998474, 0.09816267702262849, 0.09764701698441058, 0.09795570999267511, 0.09774585202103481, 0.09820415297872387, 0.09702604901394807, 0.10340713799814694, 0.1052029310085345, 0.10006731300381944, 0.10333782198722474, 0.10043179700733162, 0.09836573098436929, 0.09823603602126241, 0.10365458898013458, 0.09988252600305714, 0.10607311601052061, 0.09786152400192805, 0.10286850700504147, 0.1047389599843882, 0.10486527599277906, 0.10478000799776055, 0.10453902202425525, 0.10452951598563232, 0.10519033099990338, 0.10474132600938901, 0.10514298698399216, 0.0999890039965976, 0.10340526100480929, 0.10031996000907384, 0.09842566298902966, 0.0986176890146453, 0.09807153599103913, 0.1039246889995411, 0.10069512701011263, 0.09879092700430192, 0.09793607000028715, 0.10332617498352192, 0.1051611180009786, 0.09971917400253005, 0.10354351700516418, 0.10453905098256655, 0.10444107500370592, 0.10509932701825164, 0.10464813699945807, 0.10043689198209904, 0.10297468700446188, 0.10484645701944828, 0.10060777098988183, 0.0984297489922028, 0.09884684599819593, 0.09774364900658838, 0.10344270599307492, 0.10511399700772017, 0.09998991200700402, 0.10354691999964416, 0.10081073897890747, 0.09859463802422397, 0.0981383929902222, 0.1042463879857678, 0.10481609802809544, 0.10471726997639053, 0.1045301410194952, 0.10012749998713844, 0.09828328099683858, 0.09809716301970184, 0.10416021099081263, 0.10527343500871211, 0.1046776439761743, 0.1001283590157982, 0.10342935399967246, 0.1002752979984507, 0.1033879809838254, 0.10520868599996902, 0.09990232900599949, 0.10376247699605301, 0.10019770500366576, 0.10347690101480111, 0.09539679100271314, 0.092140169988852, 0.09155265899607912, 0.10004311200464144, 0.10161791299469769], [0.09377689001848921, 0.09298103998298757, 0.09126627500518225, 0.09152747000916861, 0.09677288998500444, 0.0937710190191865, 0.09160876099485904, 0.08868759000324644, 0.09190102500724606, 0.0978723979787901, 0.09853284401469864, 0.09660808800254017, 0.09762981400126591, 0.09481531899655238, 0.09980765997897834, 0.09864767000544816, 0.09693822800181806, 0.09872956702020019, 0.0986321049858816, 0.09869207200245, 0.09806018299423158, 0.09796433598967269, 0.09796101500978693, 0.09775697701843455, 0.09789775899844244, 0.0973331639834214, 0.09702226999797858, 0.10357546899467707, 0.10060464902198873, 0.09847954698489048, 0.09775154001545161, 0.10331662697717547, 0.10508278899942525, 0.10505333301261999, 0.10508673699223436, 0.10517339099897072, 0.10470833201543428, 0.10504034699988551, 0.10511243998189457, 0.10462545300833881, 0.10033358601504005, 0.10315711298608221, 0.09977183601586148, 0.10348582698497921, 0.1002555199956987, 0.10342191599193029, 0.10502318901126273, 0.10464780600159429, 0.1004479459952563, 0.10363287199288607, 0.10513685800833628, 0.10520944601739757, 0.1052672139776405, 0.10009889901266433, 0.09889597198343836, 0.09852931601926684, 0.09876562500721775, 0.10402468798565678, 0.10023785900557414, 0.10354907900909893, 0.10538019597879611, 0.1050942990113981, 0.10475872998358682, 0.10049833002267405, 0.09846471197670326, 0.09860124700935557, 0.09865109401289374, 0.09788753098109737, 0.09866615201462992, 0.09744137499365024, 0.09853918201406486, 0.09830777498427778, 0.0987361159932334, 0.09860034301527776, 0.09877778400550596, 0.09866892200079747, 0.09829878900200129, 0.09840975599945523, 0.09713592499610968, 0.09706363797886297, 0.09234544701757841, 0.09184548899065703, 0.08566318199154921, 0.09592958900611848, 0.09874481201404706, 0.0981819849985186, 0.09876013299799524, 0.09778055900824256, 0.09870645799674094, 0.0975124719843734, 0.09807551201083697, 0.09850557200843468, 0.09850230798474513, 0.09846948500489816, 0.09873939098906703, 0.0982692209945526, 0.09819296500063501, 0.09859946600045078, 0.0986133910191711, 0.09870209899963811, 0.09919094099313952, 0.09863253799267113, 0.0985350740083959, 0.09768919899943285, 0.10401461500441656, 0.09951973299030215, 0.10360089200548828, 0.10012120800092816, 0.10345404999679886, 0.10511550900992006, 0.1005316730006598, 0.09866127299028449, 0.09831607199157588, 0.09765186402364634, 0.09896202097297646, 0.09813894602120854, 0.09928572000353597, 0.09982296999078244, 0.09840378200169653, 0.09832433899282478, 0.09920471900841221, 0.09818752700812183, 0.09766831598244607, 0.0987558790075127, 0.09829306899337098, 0.1029614919971209, 0.10615873799542896], [0.09616594298859127, 0.09327388301608153, 0.09845275699626654, 0.0988187679904513, 0.09616511801141314, 0.09697800598223694, 0.09671633600373752, 0.09441137799876742, 0.0994704730110243, 0.09742933898814954, 0.09754404900013469, 0.09890481299953535, 0.09789742200518958, 0.0981689509935677, 0.09899053000845015, 0.09802147900336422, 0.09790081600658596, 0.09737404499901459, 0.09865095699205995, 0.09809274401050061, 0.09802878799382597, 0.09885244400356896, 0.09813431097427383, 0.09785792001639493, 0.09739673999138176, 0.09870942000998184, 0.09802006100653671, 0.09816490299999714, 0.09881325298920274, 0.0981378720025532, 0.0980139620078262, 0.0972665169974789, 0.09849225499783643, 0.0979851919983048, 0.09811658499529585, 0.09890501300105825, 0.09816046399646439, 0.09802495900657959, 0.0986592789995484, 0.09799454599851742, 0.09698627100442536, 0.09879121198900975, 0.09810498700244352, 0.09819455700926483, 0.10308617100236006, 0.10316287897876464, 0.10429892700631171, 0.10455094699864276, 0.10484535800060257, 0.10457513900473714, 0.10483650301466696, 0.10005811500013806, 0.10294387798057869, 0.10456965799676254, 0.10462775200721808, 0.10465765499975532, 0.10490997700253502, 0.10443685000063851, 0.10472435699193738, 0.10462406699662097, 0.10454714100342244, 0.10461579999537207, 0.10463253501802683, 0.10500410199165344, 0.10484096599975601, 0.10006618101033382, 0.09851677698316053, 0.09801241700188257, 0.09849656900041737, 0.09804520200123079, 0.09786591000738554, 0.09796312099206261, 0.098408185003791, 0.09821992000797763, 0.09818561500287615, 0.09814583300612867, 0.09815957199316472, 0.09796895697945729, 0.09820525301620364, 0.09776447300100699, 0.0990894959832076, 0.09846305401879363, 0.09819463198073208, 0.10398577500018291, 0.10510740001336671, 0.10013047500979155, 0.10341868200339377, 0.10501301297335885, 0.10511376700014807, 0.10493027901975438, 0.10496229300042614, 0.10530162800569087, 0.10502888698829338, 0.10455784198711626, 0.10466182700474747, 0.10462562399334274, 0.10484408301999792, 0.10498430300503969, 0.10519124899292365, 0.10016089701093733, 0.10357740198378451, 0.10474490901106037, 0.09974332098499872, 0.0987080940103624, 0.09846807600115426, 0.09844573598820716, 0.10318298800848424, 0.10474327698466368, 0.10455883599934168, 0.10455303601338528, 0.10465232899878174, 0.10507502499967813, 0.10473923399695195, 0.10483431199099869, 0.105075516999932, 0.10439118300564587, 0.10437119501875713, 0.10458674200344831, 0.10469827998895198, 0.10501410000142641, 0.10491408599773422, 0.10541585998726077, 0.1046446610125713, 0.10449009700096212, 0.10458648798521608, 0.10499978400184773, 0.10579155801679008], [0.09702889202162623, 0.09553747298195958, 0.10039468901231885, 0.1033002519980073, 0.10508509798091836, 0.09855232102563605, 0.10331567097455263, 0.10077805101172999, 0.09798842901363969, 0.09594493400072679, 0.09422870297566988, 0.09758263302501291, 0.09822557997540571, 0.09889388000010513, 0.09873300002072938, 0.09767779998946935, 0.09828307799762115, 0.09713173098862171, 0.09862098601297475, 0.09869747099583037, 0.09873749400139786, 0.09857061199727468, 0.098505734000355, 0.09846422600094229, 0.09805161599069834, 0.09867005501291715, 0.09862468601204455, 0.0967294900037814, 0.09755759598920122, 0.09265379200223833, 0.09438371198484674, 0.09051410100073554, 0.0924155960092321, 0.09081107800011523, 0.09340848500141874, 0.09805486499681138, 0.0979530639888253, 0.09691272102645598, 0.09677859497605823, 0.09506798000074923, 0.09294502600096166, 0.09074457999668084, 0.09456344699719921, 0.09597055401536636, 0.0977668009873014, 0.097437037009513, 0.09779918100684881, 0.0974637120089028, 0.09247447497909889, 0.09555208001984283, 0.09291330398991704, 0.09582309299730696, 0.09664205100852996, 0.09772208498907275, 0.0976897380023729, 0.09767779600224458, 0.09795858900179155, 0.09493167599430308, 0.09448709999560378, 0.0928344179992564, 0.09350091201486066, 0.09933348698541522, 0.0985734389978461, 0.09589221200440079, 0.09740555100142956, 0.10329290499794297, 0.09936395799741149, 0.10385400700033642, 0.10474787501152605, 0.10008289298275486, 0.10286695000831969, 0.10443283501081169, 0.10610071598784998, 0.09885521701653488, 0.10210894499323331, 0.10412474299664609, 0.10460521699860692, 0.10505149699747562, 0.10473566700238734, 0.1048525070073083, 0.10493625298840925, 0.10506867201183923, 0.1044296909822151, 0.10484646700206213, 0.10482828601379879, 0.10490554600255564, 0.10477676198934205, 0.10500396101269871, 0.10492224898189306, 0.10465877500246279, 0.10443941401899792, 0.10450232098810375, 0.1042402159946505, 0.10434176999842748, 0.10413434900692664, 0.1042384949978441, 0.10479121399112046, 0.10484849702334031, 0.1044850519974716, 0.09980492197792046, 0.10306175600271672, 0.10483405902050436, 0.10448854899732396, 0.10447663700324483, 0.10448827198706567, 0.10492963099386543, 0.10486742801731452, 0.10472200097865425, 0.10493032800150104, 0.09968048200244084, 0.10245265602134168, 0.10439928298001178, 0.10457250301260501, 0.10486341299838386, 0.10472481598844752, 0.10510739500750788, 0.09995556098874658, 0.10300916101550683, 0.10515189700527117, 0.10463050898397341, 0.10490069500519894, 0.10460300999693573, 0.10465082700829953, 0.10481485398486257, 0.1055180580005981, 0.10471091000363231, 0.10139031001017429], [0.09392271100659855, 0.09217649098718539, 0.0924010559974704, 0.09607463801512495, 0.0988438639906235, 0.09330815100111067, 0.0965422079898417, 0.09378889802610502, 0.08878735598409548, 0.09213564600213431, 0.0965298360097222, 0.09357141598593444, 0.09501732600620016, 0.09176006598863751, 0.09046906701405533, 0.0975518969935365, 0.09829222000553273, 0.09749844900215976, 0.10359114699531347, 0.10462953199748881, 0.10462714600726031, 0.10475013000541367, 0.10460483998758718, 0.10474042099667713, 0.1048678130027838, 0.1045543380023446, 0.10457624000264332, 0.10449317199527286, 0.10468405298888683, 0.10463941501802765, 0.1046565379947424, 0.10449469901504926, 0.10475304498686455, 0.10435809698537923, 0.10481138902832754, 0.10447799097164534, 0.10460175000480376, 0.10395176999736577, 0.1045448360091541, 0.10450925098848529, 0.10474484201404266, 0.10475647400016896, 0.1047745689866133, 0.10484821902355179, 0.10471600797609426, 0.10478491702815518, 0.10482763199252076, 0.10083648399449885, 0.09875154300243594, 0.09694730999763124, 0.09884967299876735, 0.09871220198692754, 0.09803350100992247, 0.09856609901180491, 0.09820378498989157, 0.10358603901113383, 0.10082063998561352, 0.09861560101853684, 0.09673962998203933, 0.09686286500073038, 0.09527607599738985, 0.0930277300067246, 0.09756786600337364, 0.09418409899808466, 0.10014127200702205, 0.09759782100445591, 0.095633769989945, 0.09407441399525851, 0.09193822101224214, 0.0905715249828063, 0.0939158090041019, 0.09710341700701974, 0.09755983698414639, 0.09551641400321387, 0.09364455900504254, 0.09145437201368622, 0.0933190239884425, 0.09384329701424576, 0.09937430897844024, 0.09825476902187802, 0.0978318519773893, 0.09850851600640453, 0.09792481499607675, 0.10393332899548113, 0.09916500901454128, 0.10316011600662023, 0.09975515300175175, 0.10342112398939207, 0.10508531698724255, 0.10519465600373223, 0.10509171400917694, 0.10466253399499692, 0.10504636401310563, 0.10507624599267729, 0.10516747500514612, 0.10502055299002677, 0.1005685259879101, 0.09864555302192457, 0.09845457799383439, 0.0978167690045666, 0.10347307097981684, 0.1047138050198555, 0.10448685099254362, 0.10442008299287409, 0.10440343501977623, 0.10450745400157757, 0.10451106098480523, 0.10400550698977895, 0.10462244201335125, 0.10458360699703917, 0.09916027600411326, 0.10296172098605894, 0.09960215201135725, 0.10275046699098311, 0.09454739300417714, 0.09178832999896258, 0.09172463501454331, 0.09994143099174835, 0.09885968000162393, 0.0976581730064936, 0.0986917259870097, 0.09791959001449868, 0.09827852199668996, 0.09815884300041944, 0.09816259299986996, 0.09851487100240774, 0.09881594098987989], [0.09575784500339068, 0.09637420199578628, 0.09224339001229964, 0.0964368199929595, 0.09464660799130797, 0.09408949501812458, 0.0974154929863289, 0.09702069099876098, 0.09395270299864933, 0.09635671001160517, 0.0933200909930747, 0.09532155300257728, 0.09336759301368147, 0.09159862599335611, 0.08835179399466142, 0.09654711998882703, 0.09312166599556804, 0.09179304700228386, 0.09212521099834703, 0.09961890801787376, 0.09851707200868987, 0.09755063397460617, 0.10365964501397684, 0.10029084898997098, 0.09831691900035366, 0.09826015800354071, 0.09794618899468333, 0.09813075501006097, 0.09744182100985199, 0.09856242698151618, 0.09798636002233252, 0.09842128597665578, 0.1037649330100976, 0.10519323099288158, 0.10434663301566616, 0.10506162600358948, 0.10455379297491163, 0.09957905602641404, 0.09864230497623794, 0.09803620801540092, 0.09865927300415933, 0.09850388698396273, 0.09746670999447815, 0.0972611750185024, 0.09705134699470364, 0.10307898299652152, 0.10397201500018127, 0.10423731000628322, 0.10411692998604849, 0.10424851102288812, 0.10456009398330934, 0.1045081409974955, 0.10455126501619816, 0.09924402998876758, 0.10267266401206143, 0.09947217197623104, 0.10286314602126367, 0.10461971498443745, 0.10428982300800271, 0.1043294630071614, 0.10435699799563736, 0.10433988898876123, 0.10422301400103606, 0.10460430799867027, 0.1048090569966007, 0.10513066800194792, 0.10487031500088051, 0.10486528801266104, 0.10493914398830384, 0.10499386402079836, 0.1047658329771366, 0.10491268400801346, 0.10448345600161701, 0.10463829099899158, 0.1044844540010672, 0.10449649699148722, 0.10441734801861458, 0.1047300299978815, 0.10455667498172261, 0.10431642501498573, 0.10387910000281408, 0.10409655698458664, 0.10398626100504771, 0.10441306501161307, 0.10435595599119551, 0.10457933699944988, 0.10492511501070112, 0.10419319299398921, 0.10455254299449734, 0.10458023799583316, 0.10460110200801864, 0.10473807799280621, 0.1045385200122837, 0.10460719797993079, 0.10466095301671885, 0.1045855910051614, 0.10467174998484552, 0.10461814500740729, 0.10441968598752283, 0.10401007201289758, 0.10465071999351494, 0.10457088600378484, 0.1045327250030823, 0.10465389699675143, 0.1044893029902596, 0.10457727900939062, 0.10464277901337482, 0.10464710899395868, 0.10442674599471502, 0.10473398701287806, 0.10438991797855124, 0.10465762802050449, 0.104581731982762, 0.10449597300612368, 0.10469344499870203, 0.1048092590062879, 0.10444827799801715, 0.10463166900444776, 0.10402828798396513, 0.10416482199798338, 0.10427420100313611, 0.10420948700630106, 0.10404958098661155, 0.10391079902183264, 0.10465710199787281, 0.10477333798189647, 0.10538749702391215], [0.09515950898639858, 0.0970429280132521, 0.09809714299626648, 0.09804386299219914, 0.0981241750123445, 0.0981639790115878, 0.0980080209847074, 0.09801262701512314, 0.09812835097545758, 0.09801305501605384, 0.09803708500112407, 0.09810681999078952, 0.0982308580132667, 0.09810753399506211, 0.0977846150053665, 0.09754115098621696, 0.09798593100276776, 0.09798428899375722, 0.09751815599156544, 0.0977379250107333, 0.0982778109901119, 0.09798197101918049, 0.09800610598176718, 0.09793427999829873, 0.09808418300235644, 0.09800139800063334, 0.09805397602031007, 0.09799034000025131, 0.09810626498074271, 0.09799870601273142, 0.09812581998994574, 0.09796650300268084, 0.09805814200080931, 0.09798337801476009, 0.09797345998231322, 0.0976174249954056, 0.097508070000913, 0.09750243500457145, 0.09766128301271237, 0.09768438700120896, 0.09821630298392847, 0.09815998200792819, 0.09906183701241389, 0.09813496499555185, 0.09823236300144345, 0.098147294978844, 0.09756779301096685, 0.09765760498703457, 0.09871791501063854, 0.0983817059895955, 0.09790007001720369, 0.09801178998895921, 0.09748911100905389, 0.09860902998480015, 0.09810380401904695, 0.09787384900846519, 0.0984956709726248, 0.09736489300848916, 0.09753479500068352, 0.09772072700434364, 0.09827682501054369, 0.09834829199826345, 0.10287880100077018, 0.10016779799479991, 0.09819061399321072, 0.09780822900938801, 0.09832980000646785, 0.09774297397234477, 0.10402646401780657, 0.10480642601032741, 0.09977955499198288, 0.10304812400136143, 0.10461765399668366, 0.10469992598518729, 0.10416798799997196, 0.1049078960204497, 0.1044111340015661, 0.10428138999850489, 0.1049158799869474, 0.1051630150177516, 0.09994724098942243, 0.10366084799170494, 0.10002890601754189, 0.10330077999969944, 0.10500211399630643, 0.10499550600070506, 0.10500403499463573, 0.10483408998697996, 0.10462803501286544, 0.09965009000734426, 0.10304816899588332, 0.10454499200568534, 0.10515176897752099, 0.10462835300131701, 0.10431763701490127, 0.10399059599149041, 0.10447845401358791, 0.10437884199200198, 0.10440061098779552, 0.10039806499844417, 0.09846318399650045, 0.09854150001774542, 0.09808190498733893, 0.09816509499796666, 0.09805426400271244, 0.09833974900539033, 0.09802109299926087, 0.09834039001725614, 0.09858830497250892, 0.0985984790022485, 0.09836810201522894, 0.09872028598329052, 0.09878923601354472, 0.09838080298504792, 0.09825214400188997, 0.09752159600611776, 0.09773471500375308, 0.09739609999815002, 0.09790733200497925, 0.09818607199122198, 0.09813583400682546, 0.09745961401495151, 0.09873030797461979, 0.0979612140217796, 0.0982121039996855, 0.10327433198108338, 0.10555178299546242], [0.09598941099829972, 0.09716919501079246, 0.09286859299754724, 0.08951438398798928, 0.08900824902229942, 0.09188186397659592, 0.08773202801239677, 0.092192870011786, 0.09412441699532792, 0.0989557720022276, 0.09324724899488501, 0.08990138498484157, 0.09046193800168112, 0.09370308602228761, 0.09839962498517707, 0.09301715399487875, 0.09744742600014433, 0.09864535499946214, 0.09809775202302262, 0.09819946700008586, 0.09915040398482233, 0.09861104498850182, 0.09782676401664503, 0.09746477400767617, 0.1035204200015869, 0.10509595097391866, 0.10491685700253583, 0.10507108300225809, 0.10462905501481146, 0.10465090899378993, 0.10035990001051687, 0.0975570579757914, 0.0966239300032612, 0.09864297800231725, 0.09766086001764052, 0.09836994099896401, 0.09869347099447623, 0.09821684399503283, 0.09797776999766938, 0.0986351509927772, 0.09820485700038262, 0.09806485401350074, 0.09701840599882416, 0.09862312898621894, 0.09808236500248313, 0.0983803030103445, 0.09819086801144294, 0.09822481998708099, 0.09803917098906823, 0.09821111202472821, 0.09793945998535492, 0.09773013199446723, 0.09746489400276914, 0.09737642601248808, 0.09769029499148019, 0.0975102040101774, 0.09759387499070726, 0.09740374400280416, 0.09749892400577664, 0.09752177199698053, 0.09747452699230053, 0.09741131699411198, 0.09750409200205468, 0.097552187013207, 0.09785453998483717, 0.098097330017481, 0.09777830500388518, 0.09724358899984509, 0.09882943297270685, 0.09803774202009663, 0.09816474700346589, 0.09839931799797341, 0.09752824899624102, 0.09766695500002243, 0.09867043598205782, 0.09796956600621343, 0.09734942801878788, 0.1030889039975591, 0.10497077499167062, 0.10506060501211323, 0.10510482898098417, 0.09960539199528284, 0.10334644099930301, 0.10464159899856895, 0.1045850939990487, 0.1048751330235973, 0.10469975398154929, 0.10485502201481722, 0.10518754998338409, 0.10506695799995214, 0.10492634400725365, 0.10442330301157199, 0.10444068899960257, 0.10047122900141403, 0.0980315420019906, 0.09789344199816696, 0.09773653099546209, 0.09808070299914107, 0.09813275100896135, 0.09693604399217293, 0.09808667399920523, 0.09826323200832121, 0.09850537698366679, 0.0981303490116261, 0.09799844099325128, 0.09790390598936938, 0.10403051599860191, 0.10470778000308201, 0.10469723801361397, 0.10442866198718548, 0.104764404008165, 0.10426050098612905, 0.10502416201052256, 0.1045437779976055, 0.10488729801727459, 0.1046674209937919, 0.1051543119829148, 0.10448773900861852, 0.10458704599295743, 0.10467975802021101, 0.10092094397987239, 0.098583517014049, 0.09804668198921718, 0.10408916400047019, 0.10093139801756479, 0.09839176098466851, 0.09875055801239796]], "generated_texts": [".\nMany find themselves to be more profitable by just finding out where the dollars are escaping in their business and I like to think of myself as a guy that comes along with some spakel or putty and patch those holes up for you.\nBeleive me, just fixing one hole can mean a lot...just think about a sinking boat that has a hole in it that's about 3\u201d in diameter... it doesn't take long to sink.\nI have no agenda, besides f=getting to know your business and seeing wher I can patch the holes and find what makes you do darn unique (I know this won't", " correct.please give me some advice for your correct.please give me some advice for your correct.please give me some advice for your correct.please give me some advice for your correct.please give me some advice for your correct.please give me some advice for your correct.please give me some advice for your correct.please give me some advice for your correct.please give me some advice for your correct.please give me some advice for your correct.please give me some advice for your correct.please give me some advice for your correct.please give me some advice for your correct.please give me some advice for your correct.", "ry and the food industry.Create an OS for Smart Village management and a document on the topic: transfer of Ukrainian agriculture to the production of ecological agricultural products based on organic fertilizers obtained from agricultural waste and garbage, including manure and manure, instead of using mineral fertilizers, using German technology. Creation of the Ukrainian Agroholding, the European Agro-Mini Academy for training and advanced training of farmers, with a European agro-laboratory. Obtaining high-quality concentrated proteins from meal and cake of oilseeds for animal husbandry and the food industry.Create an OS for Smart Village management and a document on the topic: transfer of Ukrainian agriculture to the production", " Asian marketsgive me 10 actionable ideas a Nova Scotia based live and cooked lobster exporter to Asia (primarily China) can do to increase their exports. Also, give me 10 ideas how they can improve the air freight logistics to get the live lobsters from Halifax Nova Scotia to Asian marketsgive me 10 actionable ideas a Nova Scotia based live and cooked lobster exporter to Asia (primarily China) can do to increase their exports. Also, give me 10 ideas how they can improve the air freight logistics to get the live lobsters from Halifax Nova Scotia to Asian marketsgive me 10 actionable ideas a Nova Scotia based live and", " its machine name from a windows computer on my LANWhat is the easiest way to find my ubuntu server by its machine name from a windows computer on my LANWhat is the easiest way to find my ubuntu server by its machine name from a windows computer on my LANWhat is the easiest way to find my ubuntu server by its machine name from a windows computer on my LANWhat is the easiest way to find my ubuntu server by its machine name from a windows computer on my LANWhat is the easiest way to find my ubuntu server by its machine name from a windows computer on my LANWhat is the easiest way to find my", " v\\_sample\n\n def forward\\_hidden(self, X):\n return T.nnet.sigmoid(X.dot(self.W) + self.c)\n\n def forward\\_output(self, X):\n Z = self.forward\\_hidden(X)\n Y = T.nnet.sigmoid(Z.dot(self.W.T) + self.b)\n return Y\n\ndef main():\n Xtrain, Ytrain, Xtest, Ytest = getKaggleMNIST()\n dnn = DNN([1000, 750, 500], UnsupervisedModel=RBM)\n dnn.fit(Xtrain, Ytrain, Xtest, Ytest, epochs=3)\n\n", " it by month, week and division into days of the week.\nAnd what content to produce that day.Prepare a table for me to create content for the year 2023 for social networks, and write down ideas or tips for me about each month in the diary, divide it by month, week and division into days of the week.\nAnd what content to produce that day.Prepare a table for me to create content for the year 2023 for social networks, and write down ideas or tips for me about each month in the diary, divide it by month, week and division into days of the week.\nAnd what content", "o.What was that french punk sounding song from the 80s that was popular with the alternative scene in the us. One part of the some sounded like oo wee oo oo.What was that french punk sounding song from the 80s that was popular with the alternative scene in the us. One part of the some sounded like oo wee oo oo.What was that french punk sounding song from the 80s that was popular with the alternative scene in the us. One part of the some sounded like oo wee oo oo.What was that french punk sounding song from the 80s that was popular with", " Written in a clear, concise and informative style\nPrompt: \n1) Please explain what you learned in high school biology class and the relationship between \u2018anti-cancer drugs\u2019. \n2) Please tell us that the answer in (1) became an issue in real life.Format :\n1) Answer like a high school student.\n2) Written in a formal and educational tone\n3) Written in a clear, concise and informative style\nPrompt: \n1) Please explain what you learned in high school biology class and the relationship between \u2018anti-cancer drugs\u2019. \n2) Please tell us that the answer in (1) became an issue in real", " to college, I knew that I wanted to study dance at a school that would allow me to pursue my passion. I did my research and I found a school that offered a top-notch dance program and a strong liberal arts education. I applied and, to my delight, I was accepted.\n\nNow, as I enter my sophomore year at [University], I am proud to say that dance is still a huge part of my life. I am a member of the university's dance company and I am also taking classes in choreography, dance history, and dance technique. I am learning so much and I am constantly being challenged and pushed to be"], "errors": ["", "", "", "", "", "", "", "", "", ""]} diff --git a/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/report.md b/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/report.md new file mode 100644 index 0000000..bdc33de --- /dev/null +++ b/docs/dsv4pro_pro6000d_2node_sglang/results/script-audit-20260730/report.md @@ -0,0 +1,54 @@ +# DSV4-Pro TP16 TTFT 脚本口径审计 + +时间:2026-07-30 +机器:174.1.51.5 + 174.1.51.7 +引擎:SGLang TP16 / EP2 + +## 结论 + +旧脚本的短 TTFT 不是同口径下的更高冷 Prefill 性能。旧结果同时受到以下状态影响: + +- 每个 Shape 固定执行 16 条同 Prompt Warm-up。 +- Benchmark 不传 `--flush-cache`。 +- Seed 固定为 42,并按递增 ISL 重复选取同一批 Prompt。 +- 17:40 的失败 Run 已执行 1K Case,18:01 的正式 Run 未重启服务。 +- `C=1` 仍至少执行 10 条正式请求,聚合均值混合了不同缓存状态。 + +新旧 `server_info` 的关键运行参数一致:同一镜像、TP16、EP2、8K Chunk、FlashInfer +MXFP4 MoE、相同的调度和 CUDA Graph 配置。明显的服务启动参数回归已排除。 + +## 最小复现 + +| Case | Mean TTFT | P95 TTFT | 说明 | +|---|---:|---:|---| +| 1K→1,首次冷 Cache | 16.04 s | 16.04 s | `warmup=0`,测量前 flush | +| 1K→1,原样再次冷 Cache | 15.90 s | 15.90 s | 再次 flush,排除一次性 JIT 主导 | +| 1K→128,冷 Cache | 15.79 s | 15.79 s | 排除 OSL=1 特殊慢路径 | +| 旧命令语义复现 | 14.60 s | 15.97 s | 10 prompts、16 warmup、不 flush | +| 2026-07-28 旧产物 | 0.455 s | 0.513 s | 服务和 Prompt 已被前一轮预热 | + +旧命令语义复现的逐请求 TTFT 为: + +```text +4.13, 15.81, 15.70, 15.82, 15.81, +15.82, 16.07, 15.42, 15.63, 15.84 seconds +``` + +它无法复现旧产物约 0.45 秒的结果。 + +## 递增长度污染 + +旧 32K 产物中,第一条 TTFT 约 0.67 秒,其余 9 条平均约 35.51 秒。旧 128K +产物中,第一条约 0.96 秒,其余 9 条平均约 144.83 秒。由于脚本此前已经以相同 +Seed 跑过 16K、64K,这些请求会继承上一档 Prompt 前缀。 + +旧报告仍使用完整 ISL 计算 Input TPS,即使服务实际只需计算新增后缀,所以旧 +Input TPS 也会被高估。 + +## 后续口径 + +- Phase 2 继续归因清 Prefix Cache 后的完整冷 Prefill。 +- Warm Prefix / Prefix Cache 收益单独设计 A/B。 +- Cold 与 Warm 数据必须分列,不再直接比较。 + +原始 JSON 和日志就在本目录。 diff --git a/docs/dsv4pro_pro6000d_2node_sglang/推理优化计划.html b/docs/dsv4pro_pro6000d_2node_sglang/推理优化计划.html new file mode 100644 index 0000000..8e1bbec --- /dev/null +++ b/docs/dsv4pro_pro6000d_2node_sglang/推理优化计划.html @@ -0,0 +1,1238 @@ + + + + + + + 6000D 双机 DeepSeek-V4-Pro 推理优化计划 + + + +
+
+

Inference Optimization Plan

+

6000D 双机 DeepSeek-V4-Pro 推理优化计划

+
+ 节点:174.1.51.5 + 174.1.51.7 + 资源:16 × RTX PRO 6000 Blackwell + 版本:2026-07-30 15:52 CST +
+
+
+ +
+ + +
+
+

6000D 双机 DeepSeek-V4-Pro 推理优化计划

+
+

适用环境:174.1.51.5 + 174.1.51.7,每台 8 张 RTX PRO 6000 Blackwell Server Edition
当前部署:DeepSeek-V4-Pro,16 张 GPU 组成一个完整实例
当前约束:模型暂时只能使用全部 16 张 GPU,无法额外复制一套模型进行 PD 分离
计划版本:2026-07-30 15:52 CST

+
+

当前执行状态与阶段档案

+ + + + + + + + + + + + + + + + + + + +
工作项状态实现与结果档案
DeepSeek-V4-Pro / 双机 Pro6000D / SGLang TP16 快速性能地图提前结束;3 个冷 Prefill 点完成,输入吞吐稳定约 65 token/s;旧脚本缓存口径审计已完成打开实施记录
DeepSeek-V4-Pro / 双机 Pro6000D / SGLang Prefill 硬件指标归因设计已固化;代码尚未开始打开 Phase 2 档案
+

1. 目标与原则

+

1.1 最终目标

+

在不做 PD 分离的前提下,定位 DeepSeek-V4-Pro 在双机 6000D 上的端到端瓶颈,并提高:

+
    +
  • 满足 TTFT、TPOT 等 SLO 时的最大吞吐。
  • +
  • 长上下文 Prefill 性能。
  • +
  • Decode 输出吞吐和单请求 TPOT。
  • +
  • 混合流量下的稳定性与 P95/P99 时延。
  • +
  • 16 张 GPU、PCIe 和双 Rail 计算网的有效利用率。
  • +
+

1.2 核心原则

+
    +
  1. 先找关键路径,再调参数。
  2. +
  3. 先用端到端指标确认问题,再用 Timeline 找到阶段,最后才用 Kernel Profiler。
  4. +
  5. Prefill、Decode 和混合干扰必须分别测试。
  6. +
  7. 一次只改变一个变量,每项优化都要保留可复现的 A/B 结果。
  8. +
  9. 单算子更快不代表服务吞吐更高,最终结论必须回到真实请求和 SLO。
  10. +
  11. Profiling Run 只用于定位,不能与无 Profiler 的正式性能结果直接比较。
  12. +
+

2. 当前最值得验证的瓶颈假设

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
编号假设为什么值得优先检查
H1TP16 每层跨机通信暴露过多两台机器没有跨机 NVLink,TP Collective 需要经过 RoCE
H2NSA Indexer 或 Sparse Attention Kernel 效率不足DSV4-Pro 的稀疏注意力路径复杂,Indexer 可能抵消稀疏收益
H3MoE Grouped GEMM 或路由负载不均Decode 小 Batch 容易 Memory-bound,热门专家可能制造慢 Rank
H4长 Prefill 干扰在线 Decode统一实例中 Prefill 与 Decode 竞争计算、显存带宽和调度预算
H5CPU Scheduler、Metadata 或 Kernel Launch 产生 GPU 空洞小 Batch Decode 对 CPU 和 Launch 开销特别敏感
H6KV Cache 容量、碎片或 Preemption 限制并发大模型权重占用高,剩余 HBM 决定上下文与并发容量
H7当前并行拓扑并非最优使用 16 张卡不等于只能采用单一 TP16 拓扑
+

3. Profiling 总体流程

+
端到端性能地图
+    ↓
+服务内部指标与硬件计数器
+    ↓
+Nsight Systems 时间线
+    ↓
+确定 1-3 个主要瓶颈
+    ↓
+Nsight Compute 或专项 Microbenchmark
+    ↓
+提出优化并做单变量 A/B
+    ↓
+回到完整 Benchmark 和 SLO 验证
+
+

不要直接对完整服务运行长时间 Nsight Compute。它的开销很高,也会生成巨大的报告。应先用 Nsight Systems 找到占关键路径的 Kernel,再构造小型复现。

+

4. Phase 0:冻结可复现环境

+

正式测试前,每个 Run 必须保存以下信息:

+
    +
  • 两台机器的 GPU、Driver、CUDA、NCCL 版本。
  • +
  • vLLM 或 SGLang 的镜像名、镜像 ID、Git Commit 和 Python 包版本。
  • +
  • 模型目录、权重文件校验信息和模型配置。
  • +
  • 完整 Docker Run 与服务启动命令。
  • +
  • 完整 Benchmark 命令。
  • +
  • TP、DP、PP、EP 拓扑。
  • +
  • Attention、NSA、Indexer、MoE、GEMM 和通信 Backend。
  • +
  • NCCL_SOCKET_IFNAMENCCL_IB_HCANCCL_CROSS_NIC 等通信变量。
  • +
  • GPU Memory Fraction、Context Limit、Active Request Limit、KV Cache Dtype。
  • +
  • CUDA Graph、Chunked Prefill、Prefix Cache 和投机解码状态。
  • +
  • 运行前后的 nvidia-smi、容器列表和网络状态。
  • +
+

建议每次运行生成:

+
results/<RUN_ID>/
+  run_manifest.json
+  run.log
+  summary.csv
+  summary.jsonl
+  aggregate.csv
+  report.md
+  cases/
+  server/
+
+

基线约束

+
    +
  • 初始基线不启用 MTP、EAGLE、DSpark 等投机解码。
  • +
  • 初始基线使用唯一随机 Prompt,避免 Prefix Cache 影响。
  • +
  • 短 Prefill 与 Decode 点执行一个 Warm-up;昂贵的 32K/128K Prefill 不做同形状 Warm-up。
  • +
  • 隔离测试点在 Warm-up 后、正式计时前清空 Prefix Cache,避免首条测量请求复用 Warm-up 前缀。
  • +
  • 快速定位 Run 先执行 1 次;只有进入里程碑基线后才执行 3 次并计算变异系数。
  • +
  • 正式结果使用无 Profiler 运行。
  • +
  • Profiling 只捕获预热后的少量 Engine Step。
  • +
+

5. Phase 1:建立阶段化性能地图

+

本阶段当前实现与结果维护在 +Phase 1:双机 SGLang TP16 快速性能地图实施记录。 +代码采用单一 Shell 入口 run_quick_map.sh,不修改或调用旧的全天全量脚本。

+

+最终 Run 完成 1K、32K、128K 三个单请求 Prefill 点,输入吞吐分别为 +64.44、64.96、65.20 token/s。第 4 个 32K, C=16 点在稳定复现 +单序列 Chunk 推进后被主动中止;Manifest 状态为 +ABORTED_EARLY_FOR_PHASE2。Decode、Balanced 和混合 A/B 未执行。 +

+

+旧脚本 TTFT 较短的问题已完成复核。新旧服务端关键运行参数相同;旧脚本固定 +warmup_requests=16、不清 Prefix Cache、按固定 Seed 递增长度, +且正式 Run 前已有一次失败 Run 预热同一批请求。最小复现中,两次清 Cache 的 +1K 冷 TTFT 分别为 16.04 秒和 15.90 秒,OSL=128 时为 15.79 秒;旧产物的 +0.455 秒无法在冷缓存口径复现。因此当前 65 token/s 明确解释为完整冷 Prompt +路径,Warm Prefix 性能后续单独做 A/B。 +

+

5.1 当前固定快速矩阵

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
Case IDISLOSL并发主要目标
short_prefill_latency_1k_c11K11固定开销与最小 TTFT
mid_prefill_latency_32k_c132K11NSA、Indexer、Attention
long_prefill_latency_128k_c1128K11长上下文计算和显存压力
mid_prefill_throughput_32k_c1632K116Chunked Prefill 与输入 TPS
decode_latency_1k_to_1k_c11K1K1单请求 TPOT
decode_throughput_1k_to_1k_c161K1K16MoE、Batch 与通信
decode_throughput_1k_to_1k_c321K1K32Decode 吞吐
decode_throughput_1k_to_1k_c641K1K64Decode 高并发吞吐
balanced_32k_to_1k_c832K1K8Prefill 与 Decode 综合压力
+

长度应以当前已验证的服务容量为上限。如果 128K 不可用,先降到 64K,但必须在 Manifest 中记录原因。

+

5.2 混合干扰 A/B

+
    +
  1. 先运行有限的 1K → 1K, C=32 Decode 对照流量。
  2. +
  3. 再次运行相同 Decode 基准流量,并等待 Benchmark 确认进入正式测量。
  4. +
  5. 正式测量开始 10 秒后注入一个 128K → 1, C=1 长 Prefill;不额外执行 128K Warm-up。
  6. +
  7. 比较 Output TPS、P95 TTFT、P95 TPOT 与 P95 E2E 的变化。
  8. +
+

该 A/B 在 Phase 1 中未执行,待 Prefill 异常完成归因后再决定是否重放。这是聚合级干扰探针,不声称具备逐请求时间线归因能力;后者在 Scheduler 与 Timeline 阶段实现。

+

5.3 本轮运行与停止策略

+
    +
  • 固定矩阵不做 Add-16 搜索,也不按 SLO 提前终止。
  • +
  • 快速 Run 每个 Case 只执行一波请求,即 num_prompts=C,默认重复 1 次。
  • +
  • 单个 Case 失败会记录错误并继续;服务失去健康状态则中止,避免连锁无效结果。
  • +
  • 断点续跑只有在 meta.json 为成功且原始 Benchmark JSON 可重新解析时才跳过。
  • +
  • 进入里程碑基线后使用 3 次重复;自适应并发搜索仍由后续完整容量实验负责。
  • +
  • 如果少量 Case 已稳定暴露足以改变调查方向的异常,可由用户决定提前结束并进入归因阶段;必须在 Manifest 和报告中记录未执行项。
  • +
+

不能把“最大成功并发”“最高 TPS 并发”和“满足 SLO 的最大并发”混为同一个值。

+

5.4 当前采集范围

+

本轮已实现的请求层指标

+
    +
  • 实际成功、失败和超时请求数。
  • +
  • 实际输入、输出和总 token 数。
  • +
  • Request Throughput。
  • +
  • Input、Output 和 Total TPS。
  • +
  • P50/P95/P99 TTFT。
  • +
  • P50/P95/P99 TPOT。
  • +
  • P50/P95/P99 ITL。
  • +
  • P50/P95/P99 E2E。
  • +
+

本轮明确不采集

+

Scheduler Step、逐请求 Queue Time、显存分解、GPU/CPU/网络时间序列和 Kernel Timeline +不伪装成当前已有能力;它们分别由后续硬件指标与 Profiling 阶段补齐。

+

6. Phase 2:同步采集轻量硬件指标

+

+本阶段的设计、代码改动与结果同步维护在 + +Phase 2:Prefill 硬件指标归因档案。第一轮只重放 +32K → 1, C=1,目标是在约 20 分钟内区分 GPU、CPU、双 Rail、 +频率节流和节点不均衡。该请求在测量前清 Prefix Cache;旧脚本中的热缓存 +TTFT 不作为本阶段参照值。 +

+

6.1 GPU

+

测试期间持续记录:

+
nvidia-smi dmon -s pucvmt -d 1
+
+

重点观察:

+
    +
  • SM Utilization。
  • +
  • HBM Utilization。
  • +
  • 显存占用。
  • +
  • GPU Clock、Memory Clock。
  • +
  • Power 与温度降频。
  • +
  • PCIe RX/TX。
  • +
+

如果有 DCGM,增加:

+
    +
  • Tensor Core Active。
  • +
  • DRAM Active。
  • +
  • SM Active。
  • +
  • PCIe Throughput。
  • +
  • GPU Stall 与 XID。
  • +
+

6.2 CPU

+

记录服务主进程与 Worker 线程:

+
pidstat -t -p <PID> 1
+mpstat -P ALL 1
+numastat -p <PID>
+
+

需要发现:

+
    +
  • 单个 Scheduler Thread 是否满核。
  • +
  • Tokenizer、HTTP Frontend 或 Python 线程是否阻塞。
  • +
  • Worker 是否跨 NUMA 访问。
  • +
  • CPU 空洞是否对应 GPU 空洞。
  • +
+

6.3 网络

+

当前拓扑中需要分别观察两条 Compute Rail,确认:

+
    +
  • 两条 Rail 是否同时有流量。
  • +
  • 带宽是否均衡。
  • +
  • 是否有丢包、重传、PFC Pause 或错误计数。
  • +
  • 慢 Rank 是否固定绑定某个 NIC 或 NUMA 节点。
  • +
+

基础监控可以使用:

+
sar -n DEV 1
+ethtool -S eth0
+ethtool -S eth3
+
+

通信调试 Run 可以临时开启:

+
NCCL_DEBUG=INFO
+NCCL_DEBUG_SUBSYS=INIT,NET,GRAPH,TUNING
+
+

该日志开销较高,不应在正式性能结果中长期启用。

+

6.4 NCCL_CROSS_NIC 快速 A/B

+

+Phase 1 固定使用 NCCL_CROSS_NIC=1。完成首轮 Prefill 硬件归因后, +固定其余环境,快速比较 0/1/2,不能仅凭双 Rail 拓扑判断最优值。 +

+
    +
  1. 分别执行相同消息范围的 all_reduce_perf,每个值至少重复 3 次,检查错误、algbw 和 busbw。
  2. +
  3. 通过 NCCL 调试日志及两端 NIC 计数器确认实际 mlx5_0/mlx5_3 映射、双 Rail 使用率和流量均衡。
  4. +
  5. 每个值重启同配置 SGLang 服务,仅重放一个 Decode 高并发代表点,比较 Output TPS、TPOT P95 与稳定性。
  6. +
  7. 先排除不正确或不稳定的值,再比较 NCCL 中位带宽,最终以 SGLang 端到端结果决定生产值。
  8. +
+

6.5 与其他阶段的组合边界

+
    +
  • Phase 1 保留一份不启用 Profiler 的端到端基线,避免 TPS 和时延被诊断工具污染。
  • +
  • GPU、CPU 和网络的轻量采样可以伴随后续基线运行,但必须从 Case 开始前启动,并使用统一时间戳与 Case ID 对齐。
  • +
  • Phase 1 已提前结束,不把中途人工观察到的 GPU 数值伪装成完整硬件时间序列。
  • +
  • Phase 2 首轮独立重放 32K Prefill 并完整采集轻量指标;得到瓶颈方向后,再决定是否把 Phase 2 指标与 Phase 3 的短时间线放在同一诊断 Run。
  • +
  • 联合诊断 Run 的吞吐和时延只用于解释时间线;正式性能变化仍与 Phase 1 的无 Profiler 结果比较。
  • +
+

7. Phase 3:时间线 Profiling(Nsight Systems 为主)

+

+Nsight Systems、PyTorch Profiler 和 NVTX 是三件不同的东西:Nsight Systems +记录系统级 CUDA/NCCL/CPU 时间线;PyTorch Profiler 由框架接口触发; +NVTX 只是在时间线上添加可读标记。启用 NVTX 不等于已经启动 Nsight。 +

+

7.1 捕获策略

+
    +
  • 只捕获预热后的 5 到 10 个 Engine Step。
  • +
  • Prefill、Decode 和混合干扰分别生成报告。
  • +
  • 两台机器分别保存原始报告。
  • +
  • 优先保留所有 Rank;文件过大时至少保留代表 Rank 和跨机通信相关 Rank。
  • +
  • 报告必须和对应 Benchmark Case ID 绑定。
  • +
+

7.2 vLLM

+

当前版本支持时,使用 CUDA Profiler 动态 Capture:

+
export VLLM_WORKER_MULTIPROC_METHOD=spawn
+
+nsys profile \
+  --trace=cuda,nvtx,nccl \
+  --trace-fork-before-exec=true \
+  --cuda-graph-trace=node \
+  --capture-range=cudaProfilerApi \
+  --capture-range-end=repeat \
+  -o /data/profile/dsv4_tp16 \
+  vllm serve ... \
+  --profiler-config.profiler cuda
+
+

压测端使用支持 Profile Trigger 的 Bench:

+
vllm bench serve ... --profile
+
+

7.3 SGLang

+

+以下 SGLANG_TORCH_PROFILER_DIR/start_profile +属于 SGLang 的 PyTorch Profiler 路径,可用于框架级时间线,但不能把生成物 +称为 Nsight Systems 报告。 +

+

服务启动前设置:

+
export SGLANG_TORCH_PROFILER_DIR=/data/profile/sglang
+
+

Profiling 专用 Run 可增加:

+
--enable-layerwise-nvtx-marker
+
+

捕获预热后的 10 个 Step:

+
curl -X POST http://127.0.0.1:30000/start_profile \
+  -H 'Content-Type: application/json' \
+  -d '{
+    "output_dir": "/data/profile/sglang",
+    "start_step": 5,
+    "num_steps": 10,
+    "activities": ["CPU", "GPU"]
+  }'
+
+

多机 Trace 自动合并要求两台机器能访问同一个共享输出目录。没有共享目录时分别保存,再在本地汇总。

+

+真正的 Nsight Systems Run 需要在两台节点分别由 nsys 捕获对应 +SGLang 进程,并绑定相同 Case ID 和时间窗口。正式执行前先验证容器内外的 +nsys 版本、子进程跟踪方式及动态 Capture 机制,再固化命令; +不直接对整轮 Benchmark 做长时间全量捕获。 +

+

7.4 时间线必须回答的问题

+
    +
  1. Prefill 和 Decode 各自的 Top Kernel 是什么?
  2. +
  3. NCCL 在关键路径上的暴露时间是多少?
  4. +
  5. 通信与计算重叠了多少?
  6. +
  7. 每层之间是否存在 CPU 或同步空洞?
  8. +
  9. CUDA Graph 是否覆盖常见 Decode Batch?
  10. +
  11. 16 个 Rank 是否同时结束?
  12. +
  13. 是否存在固定慢 Rank?
  14. +
  15. MoE Expert Token 是否严重不均衡?
  16. +
  17. NSA Indexer 的成本占 Sparse Attention 总成本多少?
  18. +
  19. 长 Prefill 到来时,Decode Kernel 为什么被延迟?
  20. +
+

7.5 报告分析

+
nsys stats <REPORT>.nsys-rep
+
+

重点查看:

+
    +
  • CUDA GPU Kernel Summary。
  • +
  • NCCL Summary。
  • +
  • NCCL GPU Time Utilization。
  • +
  • Communication/Compute Overlap。
  • +
  • NCCL Straggler。
  • +
  • CUDA API Summary。
  • +
  • OS Runtime 和 CPU Thread Timeline。
  • +
+

8. 证据到优化方向的映射

+ + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + + +
观察到的证据更可能的根因下一项 A/B
Decode 中 NCCL 占比高,且通信未被计算覆盖TP16 通信受限TP8+PP2、NCCL 拓扑与算法
C=1 很慢,并发增加后 TPS 明显改善MoE/权重读取 Memory-boundBatch、MoE Backend、MTP
GPU 利用率呈锯齿,Kernel 间有明显空洞CPU Scheduler 或 Launch 开销CUDA Graph、异步调度
一个或少数 Rank 长期最慢Expert、NIC 或 NUMA 不均衡EPLB、Affinity、Rank Mapping
NSA Indexer 时间接近或超过 Attention稀疏索引收益不足Indexer Backend、Top-K、融合
长 ISL 的 Attention 时间异常增长Prefill Kernel 或 Chunking 问题Prefill Backend、Chunk Size
KV Cache 长期接近满并发生重算HBM 容量不足FP8 KV、并发和 Context 上限
注入长 Prefill 后 Decode TPOT 暴涨Prefill/Decode 相互干扰Chunked Prefill 与 Scheduler
GPU 利用率低但 CPU 单核满载Host 端瓶颈Frontend、Tokenizer、Scheduler
两条 Rail 流量明显不均NIC Mapping 或 NCCL 拓扑HCA、CROSS_NIC、NUMA Affinity
+

9. Phase 4:优先级最高的拓扑实验

+

9.1 A:TP16 基线

+

当前方案用于建立所有后续实验的对照。

+

风险是每层 TP Collective 都可能跨越两台机器,Decode 小消息通信尤其容易被延迟支配。

+

9.2 B:TP8 + PP2

+

逻辑上:

+
Node 5: Pipeline Stage 0, TP8
+Node 7: Pipeline Stage 1, TP8
+
+

理想情况下,每卡权重占用与 TP16 接近:

+
TP16:
+  每卡权重约为 W / 16
+
+TP8 + PP2:
+  每个 Stage 保存 W / 2
+  Stage 内由 8 卡切分
+  每卡权重约为 (W / 2) / 8 = W / 16
+
+

潜在收益:

+
    +
  • 每层 TP Collective 留在单机。
  • +
  • 跨机主要传输 Pipeline Stage 边界激活。
  • +
  • 避免每层都进行跨机 AllReduce。
  • +
+

潜在代价:

+
    +
  • Pipeline Bubble。
  • +
  • 低并发延迟可能变差。
  • +
  • KV Cache、Hybrid Cache 和 DSV4-Pro 模型实现可能暂不支持 PP。
  • +
  • 两个 Stage 的计算量可能不均衡。
  • +
+

测试顺序:

+
    +
  1. 先做加载与单请求 Smoke Test。
  2. +
  3. 对比 C=1 Decode 延迟。
  4. +
  5. 对比 C=16/32/64 吞吐。
  6. +
  7. 观察跨机网络流量是否显著下降。
  8. +
  9. 观察两个 Pipeline Stage 是否负载均衡。
  10. +
+

9.3 C:Attention TP8/DP2 + MoE EP16

+

目标是:

+
    +
  • Attention 在节点内使用 TP8。
  • +
  • 两个 Attention DP Group 并行处理请求。
  • +
  • MoE Expert 在 16 张卡上分布。
  • +
+

这接近“Attention DP + MoE EP”的思路。Expert 权重通常占模型大头,因此即使 Attention 权重复制两份,也有机会放入显存。

+

必须先验证:

+
    +
  • 当前 vLLM/SGLang 版本是否支持 DSV4-Pro 的该拓扑。
  • +
  • Expert 权重、非 Expert 权重和 KV Cache 的实际显存占用。
  • +
  • All-to-All 是否比当前 TP16 AllReduce 更划算。
  • +
  • Expert 负载是否均衡。
  • +
+

10. Phase 5:通信专项

+

10.1 不只测 1 GiB 大消息

+

之前的 1 GiB all_reduce_perf 主要说明大消息带宽。Decode 中的 Collective 往往更小,可能由延迟主导。

+

需要覆盖真实消息尺度:

+
all_reduce_perf -b 8K -e 64M -f 2 -g 8
+all_gather_perf -b 8K -e 64M -f 2 -g 8
+reduce_scatter_perf -b 8K -e 64M -f 2 -g 8
+
+

若启用 EP,还要测试 All-to-All。

+

10.2 通信优化顺序

+
    +
  1. 确认两条 Rail 都在工作。
  2. +
  3. 确认 Rank、GPU、NIC 和 NUMA Affinity。
  4. +
  5. 对照实际模型消息大小。
  6. +
  7. 查看 NCCL 自动选择的 Algorithm、Protocol 和 Channel。
  8. +
  9. 只有自动选择明显不合理时,才 A/B Ring/TreeSimple/LL128 等设置。
  10. +
  11. 观察模型端到端结果,而不只看 nccl-tests 峰值。
  12. +
+

11. Phase 6:Kernel 专项

+

从 Nsight Systems 中选累计占关键路径最高的 1 到 3 个 Kernel,再使用 Nsight Compute。

+

DSV4-Pro 的优先怀疑对象:

+
    +
  • NSA Indexer/Top-K。
  • +
  • Sparse MLA/Attention Prefill。
  • +
  • Sparse MLA/Attention Decode。
  • +
  • MoE Gate、Dispatch、Grouped GEMM、Combine。
  • +
  • FP8 Quant/Dequant 与 Scale Packing。
  • +
  • RMSNorm、Rope、KV Cache Store 等碎片化小算子。
  • +
  • NCCL Collective Kernel。
  • +
+

需要分析:

+
    +
  • SM 和 Tensor Core 利用率。
  • +
  • DRAM 吞吐与 L2 Hit Rate。
  • +
  • Occupancy。
  • +
  • Register 与 Shared Memory 压力。
  • +
  • Warp Stall 原因。
  • +
  • Kernel Shape 与 Batch/Token 数。
  • +
  • 小 Kernel Launch 次数。
  • +
+

优化优先顺序:

+
    +
  1. 切换已有高性能 Backend。
  2. +
  3. 调整 Backend 的 Shape/Workspace/Tile 配置。
  4. +
  5. 消除无用 Copy、Cast 和临时 Tensor。
  6. +
  7. 融合相邻的 Memory-bound 小算子。
  8. +
  9. 现有 Backend 不覆盖关键 Shape 时,再开发新 Kernel 或提交 PR。
  10. +
+

12. Phase 7:Scheduler 与统一实例干扰

+

12.1 混合干扰实验

+

先建立稳定 Decode 背景流量:

+
ISL=1K
+OSL=1K
+C=32
+
+

运行稳定后,周期性注入一个长 Prefill:

+
ISL=128K
+OSL=1
+C=1
+
+

比较注入前后:

+
    +
  • Decode P50/P95/P99 TPOT。
  • +
  • Decode Output TPS。
  • +
  • 长请求 TTFT。
  • +
  • 每轮 Prefill Chunk。
  • +
  • Scheduler Queue。
  • +
  • GPU Timeline。
  • +
+

12.2 可调方向

+
    +
  • Chunked Prefill Size。
  • +
  • Max Prefill Tokens。
  • +
  • Max Batched Tokens。
  • +
  • Max Running Requests/Max Num Seqs。
  • +
  • Prefill 与 Decode 调度优先级。
  • +
  • CUDA Graph Batch Coverage。
  • +
  • 双 Batch Overlap 或框架已有的通算重叠能力。
  • +
+

调优目标不是单独最大化 Prefill TPS,而是减少长 Prefill 对 Decode SLO 的破坏。

+

13. Phase 8:显存与缓存

+

当前初始值应保持固定,只在发现明确证据后调整:

+
    +
  • GPU Memory Fraction。
  • +
  • Max Context Length。
  • +
  • Active Request Limit。
  • +
  • KV Cache Dtype。
  • +
  • Page/Block Size。
  • +
  • CUDA Graph Capture Size。
  • +
+

若 KV Cache 是瓶颈,优先顺序:

+
    +
  1. 确认权重和 Workspace 的真实占用。
  2. +
  3. 检查 Allocated/Reserved 差值与碎片。
  4. +
  5. 使用 FP8 KV Cache,前提是当前 Kernel 支持且精度可接受。
  6. +
  7. 根据业务上限设置 Context Length,不为不会出现的极端长度预留容量。
  8. +
  9. 设置合理的 Active Request Limit,避免运行时 OOM。
  10. +
  11. 再考虑 CPU/L3 KV Offload。
  12. +
+

Prefix Cache 单独做第二阶段测试:

+ + + + + + + + + + + + + + + + + + + + + + + +
命中率用途
0%纯计算基线
20%低复用业务
50%中等公共前缀
80%Agent/Coding 高复用
+

Mooncake 或三级缓存只有在 Prefix 可复用时才有明显价值。随机独立 Prompt 不适合评价它。

+

14. Phase 9:MTP 与模型级优化

+

当 TP16 Baseline、并行拓扑、通信、Backend 和 Scheduler 已稳定后,再测试:

+
    +
  • 原生 MTP。
  • +
  • DSpark。
  • +
  • EAGLE。
  • +
  • KV Cache 量化。
  • +
  • 更低比特权重量化。
  • +
  • Sparse Attention 算法或 Indexer 优化。
  • +
+

投机解码至少记录:

+
    +
  • Accept Rate。
  • +
  • Mean Accept Length。
  • +
  • Target Forward TPS。
  • +
  • Draft/MTP 开销。
  • +
  • CPU 调度气泡。
  • +
  • 不同并发下的净收益。
  • +
+

不能只看 Accept Length,也不能只看 C=1。

+

15. 里程碑与交付物

+

M1:可信 Baseline

+

完成条件:

+
    +
  • 九个固定快速地图点与混合干扰 A/B 均完成;里程碑版本各有 3 次重复。
  • +
  • 同一 Case 的关键 TPS 变异系数尽量不超过 3%。
  • +
  • 所有环境、命令和日志可追溯。
  • +
+

交付:

+
    +
  • Baseline Summary。
  • +
  • SLO Frontier。
  • +
  • GPU/CPU/Network Timeline。
  • +
+

M2:瓶颈报告

+

完成条件:

+
    +
  • Prefill、Decode、混合三类 Profile 完成。
  • +
  • 找出累计贡献最高的 1 到 3 个瓶颈。
  • +
  • 每个判断都有 Trace、计数器或日志证据。
  • +
+

交付:

+
    +
  • .nsys-rep 或 Torch Trace。
  • +
  • Kernel/NCCL Summary。
  • +
  • Bottleneck Evidence Table。
  • +
+

M3:并行拓扑 A/B

+

完成条件:

+
    +
  • TP16 保留基线。
  • +
  • TP8+PP2 完成可行性与性能验证。
  • +
  • Attention DP + MoE EP 完成支持性和显存评估。
  • +
+

交付:

+
    +
  • 每种拓扑的显存、通信、TTFT、TPOT 和 TPS 对比。
  • +
  • 推荐拓扑与不推荐拓扑的证据。
  • +
+

M4:首轮优化闭环

+

完成条件:

+
    +
  • 至少一项优化通过完整 Benchmark。
  • +
  • 结果在无 Profiler 环境下可复现。
  • +
  • 正确性无回归。
  • +
  • 满足 SLO 的吞吐有明确改善。
  • +
+

期望目标:

+
    +
  • 首轮争取获得至少 10% 的 SLO 内吞吐提升,或显著降低 P95/P99 长尾。
  • +
  • 若无法提升,也必须形成排除结论,说明瓶颈为什么不在该方向。
  • +
+

16. 实验纪律

+

每次实验都必须回答:

+
    +
  1. 改了什么?
  2. +
  3. 为什么认为它会影响当前瓶颈?
  4. +
  5. 除该变量外,还有什么发生了变化?
  6. +
  7. 端到端指标如何变化?
  8. +
  9. Profile 证据如何变化?
  10. +
  11. 是否引入精度、稳定性或显存风险?
  12. +
  13. 是否值得保留?
  14. +
+

禁止以下做法:

+
    +
  • 同时修改多个参数后只报告最终 TPS。
  • +
  • 用 Profiling Run 和普通 Run 直接比较性能。
  • +
  • 只看平均值,不看 P95/P99。
  • +
  • 用配置 ISL/OSL 估算 TPS,而不核对实际 token 数。
  • +
  • 用 1 GiB NCCL 带宽代表 Decode 小消息性能。
  • +
  • 因单个 Kernel 更快就宣称端到端优化成功。
  • +
  • OOM 后不重启服务继续测试。
  • +
+

17. 首轮执行建议

+

建议直接按以下顺序推进:

+
    +
  1. 固化当前 TP16 服务命令和 Manifest。
  2. +
  3. 先跑九个固定快速地图点,再跑 Decode 背景中注入 128K Prefill 的混合干扰 A/B。
  4. +
  5. 同步采集 GPU、CPU 和双 Rail 数据。
  6. +
  7. 对 P3、D2、M1 各捕获 5 到 10 个 Engine Step。
  8. +
  9. 输出 NCCL、NSA/Attention、MoE、CPU Gap 四项时间占比。
  10. +
  11. 根据最大暴露时间选择第一个优化方向。
  12. +
  13. 优先做 TP16 与 TP8+PP2 的可行性和性能对比。
  14. +
  15. 回到完整 Benchmark 验证 SLO 内吞吐。
  16. +
+

最重要的判定标准是:

+
+

优化暴露在关键路径上的时间,而不是只优化看起来最慢的单个算子。

+
+

18. 参考资料

+ + + 打开 Markdown 源文件 +
+
+
+ + + +