2026-07-08 02:17:13 +00:00

8.2 KiB
Raw Blame History

vllm-dspark --spec-tokens 对比报告

  • 结果目录:/data/user1/yy/bench_results/dspark_st_comparison_20260707-150649
  • 模型:/data/models/DeepSeek-V4-Flash-DSpark
  • 后端vllm-dspark (TP=8, FP8 KV cache)
  • 对比参数:--spec-tokens 3 vs --spec-tokens 5
  • Warmup100 条
  • 压测客户端:sglang.bench_serving --backend vllm

核心指标对比

Scenario Concurrency Spec Duration(s) Req/s Out tok/s Total tok/s Mean E2E(ms) P95 E2E(ms) P99 E2E(ms) Mean TTFT(ms) P99 TTFT(ms) Mean TPOT(ms) P99 TPOT(ms)
chat_short 1 3 12.80 2.50 363.90 1002.30 398.97 757.11 948.62 43.97 68.89 2.42 3.76
chat_short 1 5 11.72 2.73 397.37 1094.51 365.29 743.89 782.98 64.42 174.36 2.06 3.40
chat_short 8 3 11.08 11.55 1525.06 4705.85 672.66 1269.82 1552.65 76.20 221.18 4.67 8.43
chat_short 8 5 11.78 10.87 1434.62 4426.78 718.40 1511.14 2008.00 94.97 220.85 4.95 9.93
chat_short 16 3 21.82 11.73 1634.03 4727.68 1333.99 2611.76 2921.26 151.65 337.53 8.85 17.66
chat_short 16 5 16.42 15.59 2171.58 6282.94 997.37 2147.86 2496.56 107.56 240.46 7.09 21.93
chat_short 32 3 24.62 20.79 2754.96 8275.55 1497.40 2944.48 3658.90 144.47 453.29 10.72 23.97
chat_short 32 5 26.32 19.45 2577.42 7742.25 1600.40 3403.00 4521.14 169.13 402.36 11.39 27.82
chat_short 64 3 21.66 23.63 3131.20 9405.73 2595.90 5230.98 6258.25 251.84 509.86 18.60 40.84
chat_short 64 5 21.26 24.09 3191.43 9586.66 2555.12 5387.59 6559.69 274.71 610.40 18.47 42.08
chat_standard 64 3 20.73 24.70 3288.58 15680.46 2483.46 4798.63 5693.88 271.41 947.15 17.93 46.05
chat_standard 64 5 20.91 24.49 3259.71 15542.80 2512.52 5394.43 6023.55 300.81 932.60 17.83 44.47
decode_heavy 1 3 68.78 0.47 446.03 564.81 2148.54 3712.03 3835.51 38.60 41.18 2.28 3.03
decode_heavy 1 5 57.70 0.55 531.76 673.37 1802.00 2957.04 3676.66 59.06 114.78 1.92 3.03
decode_heavy 32 3 96.14 5.33 5535.64 6949.57 5809.29 10830.76 12546.55 86.83 382.71 5.62 8.11
decode_heavy 32 5 92.75 5.52 5737.73 7203.28 5595.02 10412.11 12522.00 107.12 331.08 5.43 9.15
generation_standard 64 3 38.58 13.27 6705.88 13363.77 4542.22 8273.38 9355.10 190.34 813.58 9.00 15.36
generation_standard 64 5 43.39 11.80 5961.40 11880.13 5075.77 9863.55 11533.47 216.87 626.93 10.09 20.90
long_context_probe 1 3 26.10 1.23 333.92 10074.03 815.18 1382.34 1556.00 265.28 558.73 2.02 2.19
long_context_probe 1 5 22.22 1.44 392.31 11835.70 693.30 1172.84 1404.17 267.61 561.65 1.53 1.88
long_context_probe 4 3 7.34 4.36 1187.03 35811.63 888.88 1550.70 1611.29 125.51 282.82 2.77 3.23
long_context_probe 4 5 6.59 4.85 1322.32 39893.35 806.55 1390.87 1500.22 143.42 236.41 2.41 3.54
long_context_probe 8 3 38.61 3.31 838.49 27327.58 2387.05 5028.60 5667.97 406.10 1093.54 7.74 17.09
long_context_probe 8 5 35.94 3.56 900.89 29361.32 2223.29 4693.10 5345.72 425.03 1241.57 7.08 19.83
rag_medium 1 3 22.31 1.43 403.86 3317.84 696.09 1165.77 1315.00 105.51 157.60 2.10 2.81
rag_medium 1 5 18.16 1.76 496.09 4075.47 566.58 968.90 1134.15 110.34 156.35 1.64 2.34
rag_medium 8 3 20.57 6.22 1573.71 14730.01 1251.63 2472.54 2796.12 127.54 309.65 4.46 7.41
rag_medium 8 5 19.13 6.69 1692.13 15838.39 1175.26 2263.59 3106.42 154.24 352.19 4.14 8.36
rag_medium 32 3 45.02 11.37 2923.96 26176.72 2756.39 5579.00 6675.25 215.15 477.20 10.00 19.84
rag_medium 32 5 42.98 11.91 3062.92 27420.79 2626.21 5182.78 6873.72 253.09 522.80 9.17 18.55
stress_standard 64 3 21.35 23.99 3193.19 15225.66 2569.66 5006.16 5831.81 261.98 710.57 18.02 36.01
stress_standard 64 5 21.75 23.54 3133.74 14942.17 2621.52 5431.08 7161.16 305.11 734.92 18.42 41.01
stress_standard 96 3 20.36 37.72 4756.72 24069.73 2433.73 4827.07 6001.29 272.61 749.99 18.04 38.52
stress_standard 96 5 25.33 30.32 3823.91 19349.55 3032.19 6272.09 8124.32 339.22 793.18 23.06 51.46
stress_standard 128 3 21.58 47.45 5942.48 29828.36 2581.42 5162.16 6222.01 298.69 970.61 19.09 40.93
stress_standard 128 5 28.36 36.11 4522.53 22700.92 3405.74 7240.13 9176.77 382.02 1087.53 25.74 57.83

吞吐 winner 统计

Scenario Concurrency Winner (Total tok/s) st=3 Total tok/s st=5 Total tok/s 提升
chat_short 1 st=5 1002.30 1094.51 9.20%
chat_short 8 st=3 4705.85 4426.78 6.30%
chat_short 16 st=5 4727.68 6282.94 32.90%
chat_short 32 st=3 8275.55 7742.25 6.89%
chat_short 64 st=5 9405.73 9586.66 1.92%
chat_standard 64 st=3 15680.46 15542.80 0.89%
decode_heavy 1 st=5 564.81 673.37 19.22%
decode_heavy 32 st=5 6949.57 7203.28 3.65%
generation_standard 64 st=3 13363.77 11880.13 12.49%
long_context_probe 1 st=5 10074.03 11835.70 17.49%
long_context_probe 4 st=5 35811.63 39893.35 11.40%
long_context_probe 8 st=5 27327.58 29361.32 7.44%
rag_medium 1 st=5 3317.84 4075.47 22.84%
rag_medium 8 st=5 14730.01 15838.39 7.52%
rag_medium 32 st=5 26176.72 27420.79 4.75%
stress_standard 64 st=3 15225.66 14942.17 1.90%
stress_standard 96 st=3 24069.73 19349.55 24.39%
stress_standard 128 st=3 29828.36 22700.92 31.40%

关键发现

1. 并发是决定性因素

  • 低并发c=1和中低并发c=8~32st=5 在多数场景下更优,尤其是长上下文和重 decode 场景。
  • 高并发c≥96st=3 明显更优,stress_standard c=128 时 st=3 比 st=5 高 31.4%
  • 中并发c=64:两者基本持平,差异多在 2% 以内。

2. st=5 更适合长上下文和重 decode

场景 最佳 spec 原因
long_context_probe (16K) st=5 长 prefill 下 st=5 的接受长度更高
rag_medium (4K) st=5 中长输入下 st=5 延迟和吞吐都更优
decode_heavy (2K output) st=5 输出越长st=5 的投机收益越大

3. st=3 在极限并发下更稳

stress_standard 随并发增加st=3 与 st=5 的差距拉大:

并发 st=3 Total tok/s st=5 Total tok/s 差距
64 15225.66 14942.17 基本持平
96 24069.73 19349.55 st=3 高 24.4%
128 29828.36 22700.92 st=3 高 31.4%

这说明 st=5 在极限并发下验证开销和 KV cache 压力显著增加,而 st=3 的验证 batch 更小、调度更稳定。

4. chat_short c=16 的异常

chat_short c=16 时 st=5 比 st=3 高 32.9%,是一个明显的 outlier。可能原因是该并发度下 st=5 的 draft 接受率和 batch 利用率恰好达到甜点。


优化建议

推荐配置

负载特征 推荐 --spec-tokens 说明
低并发 / 在线交互c ≤ 32 5 延迟低、单请求吞吐高
中并发c ≈ 64 3 或 5 均可 差异很小
高并发 / 压测c ≥ 96 3 吞吐更高、P99 更稳
长上下文 / RAG / 重 decode 5 接受长度优势更明显

下一步

  1. 如果业务以高并发为主,将默认服务改为 --spec-tokens 3
  2. 如果业务混合,可考虑按输入长度或并发度路由到不同服务实例。
  3. 本次 warmup 已从 10 提升到 100P99 TTFT 相比首次 grid 测试有明显改善;建议保持 100 条 warmup。