2026-07-31 17:38:36 +08:00

19 KiB

Phase 2 Hardware Attribution

  • Generated: 2026-07-31T17:05:04+08:00
  • Bench rows: 8
  • Failed bench rows: 0
  • GPU summary rows: 16
  • RDMA summary rows: 4
  • Case windows: 8
  • Precise main-run windows: 8/8
  • Collector status counts: {"STARTED": 18, "STOPPED": 18}

1. Benchmark results

Run Case Status Input TPS Output TPS TTFT P95 (ms) TPOT P95 (ms)
fixed_decode_throughput_1k_to_1k_c32 decode_throughput_1k_to_1k_c32 COMPLETED 448.95107215015315 448.95107215015315 10166.788510262268 65.35758210354297
fixed_long_context_decode_128k_to_1k_c1 long_context_decode_128k_to_1k_c1 COMPLETED 1614.985848663526 12.617076942683797 48279.420554987155 32.106360114411245
fixed_long_output_decode_1k_to_4k_c16 long_output_decode_1k_to_4k_c16 COMPLETED 79.3910971002871 317.5643884011484 1723.5681610036409 49.98305734157135
fixed_long_prefill_latency_128k_c1 long_prefill_latency_128k_c1 COMPLETED 2641.3788467537956 0.02015212132838284 49610.14223104576 0.0
fixed_mid_prefill_throughput_32k_c16 mid_prefill_throughput_32k_c16 COMPLETED 3115.446435236942 0.09507587998159613 161938.13879448862 0.0
mixed_prefill_decode_interference decode_control_1k_to_1k_c32 COMPLETED 453.54762732999035 453.54762732999035 9442.68154159945 66.23843437823616
mixed_prefill_decode_interference decode_with_128k_prefill_1k_to_1k_c32 COMPLETED 344.8898309034744 344.8898309034744 9891.837346865213 110.451060616212
mixed_prefill_decode_interference long_prefill_injection_128k_to_1_c1 COMPLETED 2845.5667477625557 0.02170995138368649 45980.19455798203 0.0

2. Measurement-window validity

Case Role Duration (s) Window source
decode_throughput_1k_to_1k_c32 - 72.99 bench_main_marker_plus_duration
long_context_decode_128k_to_1k_c1 - 81.16 bench_main_marker_plus_duration
long_output_decode_1k_to_4k_c16 - 206.37 bench_main_marker_plus_duration
long_prefill_latency_128k_c1 - 49.62 bench_main_marker_plus_duration
mid_prefill_throughput_32k_c16 - 168.29 bench_main_marker_plus_duration
decode_control_1k_to_1k_c32 control 144.50 bench_main_marker_plus_duration
decode_with_128k_prefill_1k_to_1k_c32 decode_background 190.02 bench_main_marker_plus_duration
long_prefill_injection_128k_to_1_c1 prefill_injection 46.06 bench_main_marker_plus_duration

3. GPU basic state (nvidia-smi)

Data: case_gpu_node_summary.csv; raw: head|worker/gpu_samples.csv.

Case Node Samples GPU util mean/p95 (%) Memory used mean (MiB) Power mean (W) SM clock mean (MHz)
decode_throughput_1k_to_1k_c32 head 408 97.09/100.00 83300.00 230.22 2402.95
decode_throughput_1k_to_1k_c32 worker 432 97.72/100.00 83300.00 221.83 2418.55
long_context_decode_128k_to_1k_c1 head 456 97.70/100.00 83219.06 257.12 2412.63
long_context_decode_128k_to_1k_c1 worker 472 97.43/100.00 83221.32 257.90 2417.53
long_output_decode_1k_to_4k_c16 head 1184 98.83/100.00 83300.00 220.83 2402.76
long_output_decode_1k_to_4k_c16 worker 1216 99.25/100.00 83300.00 215.74 2412.57
long_prefill_latency_128k_c1 head 288 94.44/100.00 83028.22 293.40 2392.22
long_prefill_latency_128k_c1 worker 288 94.85/100.00 83057.76 297.91 2398.02
mid_prefill_throughput_32k_c16 head 952 98.21/100.00 83231.07 299.46 2417.69
mid_prefill_throughput_32k_c16 worker 968 99.07/100.00 83231.94 302.56 2419.33
decode_control_1k_to_1k_c32 head 824 98.52/100.00 83364.00 231.92 2405.39
decode_control_1k_to_1k_c32 worker 856 98.60/100.00 83364.00 227.79 2418.59
decode_with_128k_prefill_1k_to_1k_c32 head 1072 99.61/100.00 83259.71 248.31 2411.49
decode_with_128k_prefill_1k_to_1k_c32 worker 1112 98.92/100.00 83256.91 247.75 2420.77
long_prefill_injection_128k_to_1_c1 head 264 99.95/100.00 83281.49 301.20 2421.11
long_prefill_injection_128k_to_1_c1 worker 272 99.89/100.00 83282.94 307.16 2421.46

4. GPU profiling counters (DCGM)

Data: case_dcgm_summary.csv; raw: head|worker/dcgm_dmon.log.

Case Node Samples GR active SM active SM occupancy Tensor active DRAM active PCIe TX/RX mean (GB/s)
decode_throughput_1k_to_1k_c32 head 584 0.971 0.525 0.141 0.059 0.417 5.205/5.201
decode_throughput_1k_to_1k_c32 worker 576 0.979 0.523 0.141 0.058 0.415 5.236/5.232
long_context_decode_128k_to_1k_c1 head 648 0.961 0.578 0.211 0.095 0.399 6.764/6.821
long_context_decode_128k_to_1k_c1 worker 648 0.965 0.579 0.211 0.095 0.401 6.780/6.838
long_output_decode_1k_to_4k_c16 head 1648 0.989 0.490 0.133 0.045 0.414 2.621/2.634
long_output_decode_1k_to_4k_c16 worker 1640 0.990 0.489 0.133 0.045 0.414 2.619/2.633
long_prefill_latency_128k_c1 head 400 0.930 0.683 0.251 0.136 0.374 10.537/10.582
long_prefill_latency_128k_c1 worker 392 0.936 0.686 0.251 0.136 0.376 10.618/10.632
mid_prefill_throughput_32k_c16 head 1344 0.980 0.713 0.278 0.127 0.440 12.664/12.720
mid_prefill_throughput_32k_c16 worker 1344 0.984 0.715 0.278 0.127 0.441 12.699/12.755
decode_control_1k_to_1k_c32 head 1152 0.989 0.539 0.146 0.061 0.427 5.451/5.456
decode_control_1k_to_1k_c32 worker 1152 0.990 0.536 0.145 0.060 0.425 5.465/5.461
decode_with_128k_prefill_1k_to_1k_c32 head 1520 0.977 0.581 0.174 0.080 0.416 6.865/6.868
decode_with_128k_prefill_1k_to_1k_c32 worker 1520 0.978 0.578 0.174 0.080 0.416 6.865/6.872
long_prefill_injection_128k_to_1_c1 head 368 0.997 0.723 0.264 0.138 0.407 11.151/11.198
long_prefill_injection_128k_to_1_c1 worker 368 0.997 0.720 0.262 0.137 0.407 11.069/11.110

5. CPU, process, and perf

Data: case_cpu_summary.csv, case_process_summary.csv, case_perf_summary.csv; raw: mpstat.log, pidstat.log, perf_stat.log.

Case Node Samples CPU/process/perf CPU active mean/p95 (%) Hot cores max Process CPU max (%) Process wait max (%) IPC Context switches mean/interval
decode_throughput_1k_to_1k_c32 head 15/1140/84 10.13/10.42 12 1225.00 0.20 3.049 41240.93
decode_throughput_1k_to_1k_c32 worker 15/1080/84 9.84/10.09 12 1210.20 0.20 3.068 39682.00
long_context_decode_128k_to_1k_c1 head 17/1292/96 9.90/10.29 12 1222.40 0.20 2.887 44877.38
long_context_decode_128k_to_1k_c1 worker 17/1224/96 9.65/10.09 12 1215.60 0.00 2.902 43686.69
long_output_decode_1k_to_4k_c16 head 41/3116/246 10.36/10.52 12 1225.80 0.20 3.098 43749.98
long_output_decode_1k_to_4k_c16 worker 41/2952/246 9.99/10.09 12 1211.40 0.20 3.117 40973.46
long_prefill_latency_128k_c1 head 10/1129/60 10.00/10.27 12 1211.20 0.20 2.727 50995.80
long_prefill_latency_128k_c1 worker 10/720/60 9.69/10.18 12 1210.80 0.20 2.749 48210.50
mid_prefill_throughput_32k_c16 head 34/2584/204 9.99/10.27 12 1211.00 0.20 2.802 39421.29
mid_prefill_throughput_32k_c16 worker 34/2448/204 9.84/10.11 12 1210.20 0.20 2.810 39129.38
decode_control_1k_to_1k_c32 head 29/2204/174 10.26/10.44 12 1220.40 0.20 3.042 42398.83
decode_control_1k_to_1k_c32 worker 29/2088/174 9.98/10.11 12 1210.40 0.20 3.072 40484.59
decode_with_128k_prefill_1k_to_1k_c32 head 38/2888/228 10.34/11.37 13 1223.60 0.20 2.980 41223.05
decode_with_128k_prefill_1k_to_1k_c32 worker 38/2736/228 9.87/10.09 12 1211.40 0.20 3.005 39834.92
long_prefill_injection_128k_to_1_c1 head 10/749/54 10.21/10.63 13 1223.60 0.00 2.808 33399.67
long_prefill_injection_128k_to_1_c1 worker 9/648/60 10.04/10.12 12 1211.40 0.20 2.830 33911.70

6. NUMA memory placement

Data/raw: case_numa_summary.csv, head|worker/numa_samples.csv.

Case Node Samples Node 0 mean (MiB) Node 1 mean (MiB) Total mean (MiB) Imbalance mean/max (%)
decode_throughput_1k_to_1k_c32 head 13 17941.22 27565.28 45506.51 21.15/21.15
decode_throughput_1k_to_1k_c32 worker 13 19029.49 24814.32 43843.80 13.19/13.19
long_context_decode_128k_to_1k_c1 head 13 17946.50 27569.43 45515.94 21.14/21.15
long_context_decode_128k_to_1k_c1 worker 14 19037.22 24816.07 43853.29 13.18/13.18
long_output_decode_1k_to_4k_c16 head 35 17941.67 27565.20 45506.88 21.15/21.17
long_output_decode_1k_to_4k_c16 worker 37 19029.40 24814.58 43843.96 13.19/13.20
long_prefill_latency_128k_c1 head 8 17869.37 27486.32 45355.71 21.20/21.21
long_prefill_latency_128k_c1 worker 8 18970.34 24731.71 43702.05 13.18/13.21
mid_prefill_throughput_32k_c16 head 28 17940.27 27563.88 45504.17 21.15/21.15
mid_prefill_throughput_32k_c16 worker 30 19032.02 24805.59 43837.58 13.17/13.22
decode_control_1k_to_1k_c32 head 24 17946.96 27573.84 45520.81 21.15/21.21
decode_control_1k_to_1k_c32 worker 26 19039.33 24818.28 43857.62 13.18/13.18
decode_with_128k_prefill_1k_to_1k_c32 head 32 17955.73 27573.17 45528.91 21.12/21.13
decode_with_128k_prefill_1k_to_1k_c32 worker 34 19042.79 24823.40 43866.18 13.18/13.19
long_prefill_injection_128k_to_1_c1 head 8 17951.77 27569.78 45521.56 21.13/21.13
long_prefill_injection_128k_to_1_c1 worker 8 19038.43 24819.53 43857.94 13.18/13.18

7. Linux netdev and RDMA data path

Netdev data: case_netdev_summary.csv; RDMA data: case_rdma_summary.csv; raw: sar_net.log, rdma.csv.

Linux interfaces

Case Node Interface Samples RX mean/max (Gbit/s) TX mean/max (Gbit/s) Util max (%) RX/TX error max (/s)
decode_throughput_1k_to_1k_c32 head eth0 30 0.000/0.000 0.000/0.000 0.000 0.00/0.00
decode_throughput_1k_to_1k_c32 head eth3 30 0.000/0.000 0.000/0.000 0.000 0.00/0.00
decode_throughput_1k_to_1k_c32 worker eth0 30 0.000/0.000 0.000/0.000 0.000 0.00/0.00
decode_throughput_1k_to_1k_c32 worker eth3 30 0.000/0.000 0.000/0.000 0.000 0.00/0.00
long_context_decode_128k_to_1k_c1 head eth0 34 0.000/0.000 0.000/0.000 0.000 0.00/0.00
long_context_decode_128k_to_1k_c1 head eth3 34 0.000/0.000 0.000/0.000 0.000 0.00/0.00
long_context_decode_128k_to_1k_c1 worker eth0 34 0.000/0.000 0.000/0.000 0.000 0.00/0.00
long_context_decode_128k_to_1k_c1 worker eth3 34 0.000/0.000 0.000/0.000 0.000 0.00/0.00
long_output_decode_1k_to_4k_c16 head eth0 82 0.000/0.000 0.000/0.000 0.000 0.00/0.00
long_output_decode_1k_to_4k_c16 head eth3 82 0.000/0.000 0.000/0.000 0.000 0.00/0.00
long_output_decode_1k_to_4k_c16 worker eth0 82 0.000/0.000 0.000/0.000 0.000 0.00/0.00
long_output_decode_1k_to_4k_c16 worker eth3 82 0.000/0.000 0.000/0.000 0.000 0.00/0.00
long_prefill_latency_128k_c1 head eth0 20 0.000/0.001 0.000/0.000 0.000 0.00/0.00
long_prefill_latency_128k_c1 head eth3 20 0.000/0.000 0.000/0.000 0.000 0.00/0.00
long_prefill_latency_128k_c1 worker eth0 20 0.000/0.000 0.000/0.001 0.000 0.00/0.00
long_prefill_latency_128k_c1 worker eth3 20 0.000/0.000 0.000/0.000 0.000 0.00/0.00
mid_prefill_throughput_32k_c16 head eth0 68 0.000/0.000 0.000/0.000 0.000 0.00/0.00
mid_prefill_throughput_32k_c16 head eth3 68 0.000/0.000 0.000/0.000 0.000 0.00/0.00
mid_prefill_throughput_32k_c16 worker eth0 68 0.000/0.000 0.000/0.000 0.000 0.00/0.00
mid_prefill_throughput_32k_c16 worker eth3 68 0.000/0.000 0.000/0.000 0.000 0.00/0.00
decode_control_1k_to_1k_c32 head eth0 58 0.000/0.000 0.000/0.000 0.000 0.00/0.00
decode_control_1k_to_1k_c32 head eth3 58 0.000/0.000 0.000/0.000 0.000 0.00/0.00
decode_control_1k_to_1k_c32 worker eth0 58 0.000/0.000 0.000/0.000 0.000 0.00/0.00
decode_control_1k_to_1k_c32 worker eth3 58 0.000/0.000 0.000/0.000 0.000 0.00/0.00
decode_with_128k_prefill_1k_to_1k_c32 head eth0 76 0.000/0.000 0.000/0.000 0.000 0.00/0.00
decode_with_128k_prefill_1k_to_1k_c32 head eth3 76 0.000/0.000 0.000/0.000 0.000 0.00/0.00
decode_with_128k_prefill_1k_to_1k_c32 worker eth0 76 0.000/0.000 0.000/0.000 0.000 0.00/0.00
decode_with_128k_prefill_1k_to_1k_c32 worker eth3 76 0.000/0.000 0.000/0.000 0.000 0.00/0.00
long_prefill_injection_128k_to_1_c1 head eth0 20 0.000/0.000 0.000/0.000 0.000 0.00/0.00
long_prefill_injection_128k_to_1_c1 head eth3 20 0.000/0.000 0.000/0.000 0.000 0.00/0.00
long_prefill_injection_128k_to_1_c1 worker eth0 18 0.000/0.000 0.000/0.000 0.000 0.00/0.00
long_prefill_injection_128k_to_1_c1 worker eth3 18 0.000/0.000 0.000/0.000 0.000 0.00/0.00

RDMA HCAs

Case Node HCA Samples TX/RX (Gbit/s) Wait delta Discard/error delta Retry exceeded delta
decode_throughput_1k_to_1k_c32 head mlx5_0 66 36.69/36.69 0 0/0 0
decode_throughput_1k_to_1k_c32 head mlx5_3 66 36.71/36.71 0 0/0 0
decode_throughput_1k_to_1k_c32 worker mlx5_0 68 36.80/36.80 0 0/0 0
decode_throughput_1k_to_1k_c32 worker mlx5_3 68 36.79/36.79 0 0/0 0
long_context_decode_128k_to_1k_c1 head mlx5_0 75 43.83/43.83 0 0/0 0
long_context_decode_128k_to_1k_c1 head mlx5_3 75 43.95/43.95 0 0/0 0
long_context_decode_128k_to_1k_c1 worker mlx5_0 77 44.73/44.73 0 0/0 0
long_context_decode_128k_to_1k_c1 worker mlx5_3 77 44.11/44.11 0 0/0 0
long_output_decode_1k_to_4k_c16 head mlx5_0 188 18.91/18.91 0 0/0 0
long_output_decode_1k_to_4k_c16 head mlx5_3 188 18.90/18.90 0 0/0 0
long_output_decode_1k_to_4k_c16 worker mlx5_0 194 18.90/18.90 0 0/0 0
long_output_decode_1k_to_4k_c16 worker mlx5_3 194 18.90/18.90 0 0/0 0
long_prefill_latency_128k_c1 head mlx5_0 45 71.38/71.38 0 0/0 0
long_prefill_latency_128k_c1 head mlx5_3 45 71.45/71.45 0 0/0 0
long_prefill_latency_128k_c1 worker mlx5_0 47 70.54/70.54 0 0/0 0
long_prefill_latency_128k_c1 worker mlx5_3 47 70.54/70.54 0 0/0 0
mid_prefill_throughput_32k_c16 head mlx5_0 154 82.99/82.99 0 0/0 0
mid_prefill_throughput_32k_c16 head mlx5_3 154 83.01/83.01 0 0/0 0
mid_prefill_throughput_32k_c16 worker mlx5_0 158 83.44/83.44 0 0/0 0
mid_prefill_throughput_32k_c16 worker mlx5_3 158 83.45/83.45 0 0/0 0
decode_control_1k_to_1k_c32 head mlx5_0 131 38.01/38.01 0 0/0 0
decode_control_1k_to_1k_c32 head mlx5_3 131 38.00/38.00 0 0/0 0
decode_control_1k_to_1k_c32 worker mlx5_0 137 37.74/37.74 0 0/0 0
decode_control_1k_to_1k_c32 worker mlx5_3 137 37.75/37.75 0 0/0 0
decode_with_128k_prefill_1k_to_1k_c32 head mlx5_0 173 46.80/46.80 0 0/0 0
decode_with_128k_prefill_1k_to_1k_c32 head mlx5_3 173 46.81/46.81 0 0/0 0
decode_with_128k_prefill_1k_to_1k_c32 worker mlx5_0 179 46.67/46.67 0 0/0 0
decode_with_128k_prefill_1k_to_1k_c32 worker mlx5_3 179 46.67/46.67 0 0/0 0
long_prefill_injection_128k_to_1_c1 head mlx5_0 42 74.15/74.15 0 0/0 0
long_prefill_injection_128k_to_1_c1 head mlx5_3 42 74.23/74.23 0 0/0 0
long_prefill_injection_128k_to_1_c1 worker mlx5_0 43 74.55/74.54 0 0/0 0
long_prefill_injection_128k_to_1_c1 worker mlx5_3 43 74.57/74.58 0 0/0 0

8. PCIe P2P and NCCL communication baseline

Data: communication_aggregate.csv; raw: communication/*.log.

Test Scope Path/CROSS_NIC Samples/repetitions Size (MiB) Mean latency (ms) Bandwidth / busbw (GB/s) Minimum Wrong values
all_reduce head_8gpu 2 3 1 1.18 1.64 1.16 0
all_reduce head_8gpu 2 3 1024 47.26 39.76 39.67 0
all_reduce head_8gpu 2 3 64 3.11 37.79 37.25 0
all_reduce two_node_16gpu 0 3 1 1.17 1.73 1.32 0
all_reduce two_node_16gpu 0 3 1024 51.17 39.35 39.30 0
all_reduce two_node_16gpu 0 3 64 3.40 36.98 36.72 0
all_reduce two_node_16gpu 1 3 1 1.30 1.65 1.06 0
all_reduce two_node_16gpu 1 3 1024 50.73 39.68 39.59 0
all_reduce two_node_16gpu 1 3 64 3.38 37.25 36.86 0
all_reduce two_node_16gpu 2 3 1 1.14 1.76 1.41 0
all_reduce two_node_16gpu 2 3 1024 50.93 39.53 39.34 0
all_reduce two_node_16gpu 2 3 64 3.43 36.65 36.14 0
all_reduce worker_8gpu 2 3 1 0.97 1.91 1.66 0
all_reduce worker_8gpu 2 3 1024 47.27 39.75 39.56 0
all_reduce worker_8gpu 2 3 64 3.07 38.22 37.44 0
p2p_copy head cross_numa_sys 32 256 - 52.40 52.18 -
p2p_copy head same_pcie_switch 24 256 - 53.61 53.35 -
p2p_copy worker cross_numa_sys 32 256 - 52.32 52.08 -
p2p_copy worker same_pcie_switch 24 256 - 53.50 53.22 -

9. Machine-readable summaries

  • gpu_summary.csv
  • rdma_summary.csv
  • bench_summary.csv
  • case_windows.csv
  • case_gpu_summary.csv
  • case_gpu_node_summary.csv
  • case_dcgm_summary.csv
  • case_cpu_summary.csv
  • case_process_summary.csv
  • case_perf_summary.csv
  • case_numa_summary.csv
  • case_netdev_summary.csv
  • case_rdma_summary.csv
  • communication_summary.csv
  • communication_aggregate.csv
  • summary.json

Each conclusion must cite the corresponding table above and its raw file; missing samples are reported as -, never interpreted as zero.