== pm-1788716023.1756692-TP-0-PP-0.trace.json.gz ==
  span 3.522s  n_steps~12  round~293.5ms
  GPU-busy 1939.1ms (55.1%)  idle 1582.4ms (44.9%)  kernels n=20106
  host cudaStreamSynchronize: n=120 total=58.6ms (per-call avg 0.49ms)
  host cudaEventSynchronize: n=12 total=0.1ms   cudaEventRecord n=4164
  host cudaLaunchKernel: n=3072 cpu=39.9ms (avg 13us)  memcpy n=314 gpu=0.3ms
  gaps>0.3ms: n=839 total=932.7ms  top5=['239.5', '49.7', '46.1', '42.1', '39.1']ms
  gap-us attributed to cudaStreamSynchronize overlap: 39.6ms
  long host ops >1ms: n=1039 total=35837.5ms; top10:
    +    -21.0ms  3546.8ms [Trace] PyTorch Profiler (0)
    +     -8.6ms   537.4ms [user_annotation] gloo:send
    +    293.5ms   517.0ms [user_annotation] gloo:send
    +    275.2ms   513.9ms [user_annotation] gloo:send
    +    275.2ms   513.8ms [user_annotation] gloo:send
    +   2074.4ms   324.6ms [user_annotation] gloo:send
    +   2074.4ms   324.6ms [user_annotation] gloo:send
    +   2632.4ms   319.9ms [user_annotation] gloo:send
    +   2632.4ms   319.9ms [user_annotation] gloo:send
    +   3184.2ms   313.3ms [user_annotation] gloo:send
    nccl_AllReduce    1588.1ms  n=  1056  (45.1% span, 81.9% busy)
    cutlass_moe        289.5ms  n=  9612  (8.2% span, 14.9% busy)
    sparse_mla          26.9ms  n=  1008  (0.8% span, 1.4% busy)
    other               17.5ms  n=  7986  (0.5% span, 0.9% busy)
    nccl_other          14.1ms  n=    36  (0.4% span, 0.7% busy)
    mqa_logits           3.7ms  n=   204  (0.1% span, 0.2% busy)
    nccl_SendRecv        2.0ms  n=   204  (0.1% span, 0.1% busy)

== pm-1788716023.1756692-TP-0-PP-1.trace.json.gz ==
  span 11.041s  n_steps~12  round~920.1ms
  GPU-busy 947.6ms (8.6%)  idle 10093.5ms (91.4%)  kernels n=20457
  host cudaStreamSynchronize: n=120 total=36.2ms (per-call avg 0.30ms)
  host cudaEventSynchronize: n=12 total=0.1ms   cudaEventRecord n=4374
  host cudaLaunchKernel: n=3240 cpu=38.9ms (avg 12us)  memcpy n=314 gpu=0.3ms
  gaps>0.3ms: n=1432 total=8788.7ms  top5=['7772.2', '23.8', '15.3', '15.2', '14.4']ms
  gap-us attributed to cudaStreamSynchronize overlap: 7870.4ms
  long host ops >1ms: n=526 total=51095.8ms; top10:
    +     -8.3ms 11056.9ms [Trace] PyTorch Profiler (0)
    +   3250.1ms  7773.9ms [user_annotation] recv_res_dict_from_prev_stage
    +   3250.3ms  7770.3ms [user_annotation] gloo:recv
    +   2966.7ms   283.1ms [user_annotation] gloo:send
    +   2966.8ms   283.1ms [user_annotation] gloo:send
    +   1305.0ms   277.6ms [user_annotation] gloo:send
    +   1305.1ms   277.5ms [user_annotation] gloo:send
    +    754.3ms   276.8ms [user_annotation] gloo:send
    +    754.4ms   276.7ms [user_annotation] gloo:send
    +   1869.2ms   275.5ms [user_annotation] gloo:send
    nccl_AllReduce     535.6ms  n=  1044  (4.9% span, 56.5% busy)
    cutlass_moe        350.3ms  n=  9876  (3.2% span, 37.0% busy)
    sparse_mla          27.4ms  n=  1008  (0.2% span, 2.9% busy)
    other               17.5ms  n=  8094  (0.2% span, 1.8% busy)
    nccl_other          12.1ms  n=    48  (0.1% span, 1.3% busy)
    mqa_logits           3.0ms  n=   180  (0.0% span, 0.3% busy)
    nccl_SendRecv        2.7ms  n=   207  (0.0% span, 0.3% busy)

== pm-1788716023.1756692-TP-1-PP-0.trace.json.gz ==
  span 3.522s  n_steps~12  round~293.5ms
  GPU-busy 2179.6ms (61.9%)  idle 1342.0ms (38.1%)  kernels n=20106
  host cudaStreamSynchronize: n=120 total=94.4ms (per-call avg 0.79ms)
  host cudaEventSynchronize: n=12 total=0.1ms   cudaEventRecord n=4164
  host cudaLaunchKernel: n=3072 cpu=37.6ms (avg 12us)  memcpy n=314 gpu=0.3ms
  gaps>0.3ms: n=659 total=837.0ms  top5=['240.0', '49.5', '45.9', '42.3', '39.6']ms
  gap-us attributed to cudaStreamSynchronize overlap: 34.6ms
  long host ops >1ms: n=1138 total=32034.5ms; top10:
    +    -17.7ms  3543.5ms [Trace] PyTorch Profiler (0)
    +     -8.5ms   531.8ms [user_annotation] gloo:send
    +    279.8ms   507.9ms [user_annotation] gloo:send
    +    279.8ms   507.9ms [user_annotation] gloo:send
    +   2067.8ms   320.2ms [user_annotation] gloo:send
    +   2067.8ms   320.2ms [user_annotation] gloo:send
    +   2628.5ms   316.7ms [user_annotation] gloo:send
    +   2628.5ms   316.7ms [user_annotation] gloo:send
    +   3177.8ms   314.3ms [user_annotation] gloo:send
    +   3177.8ms   314.3ms [user_annotation] gloo:send
    nccl_AllReduce    1828.6ms  n=  1056  (51.9% span, 83.9% busy)
    cutlass_moe        288.3ms  n=  9612  (8.2% span, 13.2% busy)
    sparse_mla          26.7ms  n=  1008  (0.8% span, 1.2% busy)
    other               17.2ms  n=  7986  (0.5% span, 0.8% busy)
    nccl_other          16.2ms  n=    36  (0.5% span, 0.7% busy)
    mqa_logits           3.6ms  n=   204  (0.1% span, 0.2% busy)
    nccl_SendRecv        2.0ms  n=   204  (0.1% span, 0.1% busy)

== pm-1788716023.1756692-TP-1-PP-1.trace.json.gz ==
  span 11.046s  n_steps~12  round~920.5ms
  GPU-busy 2059.8ms (18.6%)  idle 8986.2ms (81.4%)  kernels n=20457
  host cudaStreamSynchronize: n=120 total=86.0ms (per-call avg 0.72ms)
  host cudaEventSynchronize: n=12 total=0.1ms   cudaEventRecord n=4374
  host cudaLaunchKernel: n=3240 cpu=37.2ms (avg 11us)  memcpy n=314 gpu=0.3ms
  gaps>0.3ms: n=827 total=8437.5ms  top5=['7771.2', '22.5', '20.4', '15.5', '15.4']ms
  gap-us attributed to cudaStreamSynchronize overlap: 7836.1ms
  long host ops >1ms: n=1091 total=52267.7ms; top10:
    +     -8.5ms 11062.3ms [Trace] PyTorch Profiler (0)
    +   3255.5ms  7772.7ms [user_annotation] recv_res_dict_from_prev_stage
    +   3255.6ms  7769.1ms [user_annotation] gloo:recv
    +   2965.9ms   289.2ms [user_annotation] gloo:send
    +   2966.0ms   289.2ms [user_annotation] gloo:send
    +   1862.4ms   287.6ms [user_annotation] gloo:send
    +   1862.4ms   287.5ms [user_annotation] gloo:send
    +   1302.8ms   285.2ms [user_annotation] gloo:send
    +   1302.8ms   285.2ms [user_annotation] gloo:send
    +    759.5ms   276.8ms [user_annotation] gloo:send
    nccl_AllReduce    1638.2ms  n=  1044  (14.8% span, 79.5% busy)
    cutlass_moe        349.2ms  n=  9876  (3.2% span, 17.0% busy)
    sparse_mla          27.4ms  n=  1008  (0.2% span, 1.3% busy)
    nccl_other          23.7ms  n=    48  (0.2% span, 1.2% busy)
    other               17.7ms  n=  8094  (0.2% span, 0.9% busy)
    mqa_logits           3.0ms  n=   180  (0.0% span, 0.1% busy)
    nccl_SendRecv        2.6ms  n=   207  (0.0% span, 0.1% busy)

== pm-1788716023.1756692-TP-2-PP-0.trace.json.gz ==
  span 3.522s  n_steps~12  round~293.5ms
  GPU-busy 1056.4ms (30.0%)  idle 2465.4ms (70.0%)  kernels n=20106
  host cudaStreamSynchronize: n=120 total=58.8ms (per-call avg 0.49ms)
  host cudaEventSynchronize: n=12 total=0.1ms   cudaEventRecord n=4164
  host cudaLaunchKernel: n=3072 cpu=41.0ms (avg 13us)  memcpy n=314 gpu=0.3ms
  gaps>0.3ms: n=1278 total=1208.5ms  top5=['240.0', '47.7', '44.1', '40.6', '37.9']ms
  gap-us attributed to cudaStreamSynchronize overlap: 36.0ms
  long host ops >1ms: n=621 total=31027.9ms; top10:
    +    -11.2ms  3537.6ms [Trace] PyTorch Profiler (0)
    +    258.3ms   538.5ms [user_annotation] gloo:send
    +    258.4ms   538.5ms [user_annotation] gloo:send
    +     -8.4ms   537.8ms [user_annotation] gloo:send
    +   2077.2ms   322.5ms [user_annotation] gloo:send
    +   2077.2ms   322.5ms [user_annotation] gloo:send
    +   3185.7ms   319.6ms [user_annotation] gloo:send
    +   3185.8ms   319.6ms [user_annotation] gloo:send
    +   2635.6ms   317.1ms [user_annotation] gloo:send
    +   2635.6ms   317.1ms [user_annotation] gloo:send
    nccl_AllReduce     707.7ms  n=  1056  (20.1% span, 67.0% busy)
    cutlass_moe        289.3ms  n=  9612  (8.2% span, 27.4% busy)
    sparse_mla          27.0ms  n=  1008  (0.8% span, 2.6% busy)
    other               17.0ms  n=  7986  (0.5% span, 1.6% busy)
    nccl_other          10.1ms  n=    36  (0.3% span, 1.0% busy)
    mqa_logits           3.7ms  n=   204  (0.1% span, 0.3% busy)
    nccl_SendRecv        2.8ms  n=   204  (0.1% span, 0.3% busy)

== pm-1788716023.1756692-TP-2-PP-1.trace.json.gz ==
  span 11.040s  n_steps~12  round~920.0ms
  GPU-busy 2123.5ms (19.2%)  idle 8916.7ms (80.8%)  kernels n=20457
  host cudaStreamSynchronize: n=120 total=134.5ms (per-call avg 1.12ms)
  host cudaEventSynchronize: n=12 total=0.1ms   cudaEventRecord n=4374
  host cudaLaunchKernel: n=3240 cpu=37.8ms (avg 12us)  memcpy n=314 gpu=0.3ms
  gaps>0.3ms: n=744 total=8389.1ms  top5=['7776.1', '24.2', '23.4', '20.7', '14.6']ms
  gap-us attributed to cudaStreamSynchronize overlap: 7893.3ms
  long host ops >1ms: n=1082 total=52177.2ms; top10:
    +     -8.1ms 11055.7ms [Trace] PyTorch Profiler (0)
    +   3249.7ms  7777.6ms [user_annotation] recv_res_dict_from_prev_stage
    +   3249.8ms  7774.0ms [user_annotation] gloo:recv
    +   1869.1ms   275.0ms [user_annotation] gloo:send
    +   1869.2ms   274.9ms [user_annotation] gloo:send
    +   2974.5ms   274.9ms [user_annotation] gloo:send
    +   2974.6ms   274.8ms [user_annotation] gloo:send
    +    758.2ms   272.3ms [user_annotation] gloo:send
    +    758.2ms   272.3ms [user_annotation] gloo:send
    +   1312.0ms   270.1ms [user_annotation] gloo:send
    nccl_AllReduce    1703.9ms  n=  1044  (15.4% span, 80.2% busy)
    cutlass_moe        348.0ms  n=  9876  (3.2% span, 16.4% busy)
    sparse_mla          27.1ms  n=  1008  (0.2% span, 1.3% busy)
    nccl_other          23.5ms  n=    48  (0.2% span, 1.1% busy)
    other               17.4ms  n=  8094  (0.2% span, 0.8% busy)
    mqa_logits           3.0ms  n=   180  (0.0% span, 0.1% busy)
    nccl_SendRecv        2.5ms  n=   207  (0.0% span, 0.1% busy)

== pm-1788716023.1756692-TP-3-PP-0.trace.json.gz ==
  span 3.522s  n_steps~12  round~293.5ms
  GPU-busy 2075.2ms (58.9%)  idle 1446.3ms (41.1%)  kernels n=20106
  host cudaStreamSynchronize: n=120 total=86.1ms (per-call avg 0.72ms)
  host cudaEventSynchronize: n=12 total=0.1ms   cudaEventRecord n=4164
  host cudaLaunchKernel: n=3072 cpu=35.9ms (avg 12us)  memcpy n=314 gpu=0.3ms
  gaps>0.3ms: n=806 total=908.2ms  top5=['239.9', '49.7', '45.9', '42.3', '39.7']ms
  gap-us attributed to cudaStreamSynchronize overlap: 34.7ms
  long host ops >1ms: n=1110 total=31972.9ms; top10:
    +    -30.9ms  3556.9ms [Trace] PyTorch Profiler (0)
    +     -8.6ms   538.8ms [user_annotation] gloo:send
    +    273.2ms   516.4ms [user_annotation] gloo:send
    +    273.3ms   516.4ms [user_annotation] gloo:send
    +   2067.2ms   325.5ms [user_annotation] gloo:send
    +   2067.2ms   325.5ms [user_annotation] gloo:send
    +   2625.7ms   323.3ms [user_annotation] gloo:send
    +   2625.7ms   323.2ms [user_annotation] gloo:send
    +   3177.4ms   316.5ms [user_annotation] gloo:send
    +   3177.4ms   316.5ms [user_annotation] gloo:send
    nccl_AllReduce    1726.3ms  n=  1056  (49.0% span, 83.2% busy)
    cutlass_moe        289.0ms  n=  9612  (8.2% span, 13.9% busy)
    sparse_mla          26.8ms  n=  1008  (0.8% span, 1.3% busy)
    other               17.4ms  n=  7986  (0.5% span, 0.8% busy)
    nccl_other          13.1ms  n=    36  (0.4% span, 0.6% busy)
    mqa_logits           3.7ms  n=   204  (0.1% span, 0.2% busy)
    nccl_SendRecv        2.0ms  n=   204  (0.1% span, 0.1% busy)

== pm-1788716023.1756692-TP-3-PP-1.trace.json.gz ==
  span 11.040s  n_steps~12  round~920.0ms
  GPU-busy 2171.9ms (19.7%)  idle 8868.2ms (80.3%)  kernels n=20457
  host cudaStreamSynchronize: n=120 total=111.6ms (per-call avg 0.93ms)
  host cudaEventSynchronize: n=12 total=0.1ms   cudaEventRecord n=4374
  host cudaLaunchKernel: n=3240 cpu=38.8ms (avg 12us)  memcpy n=314 gpu=0.3ms
  gaps>0.3ms: n=692 total=8361.0ms  top5=['7770.1', '21.1', '16.9', '15.9', '14.1']ms
  gap-us attributed to cudaStreamSynchronize overlap: 7846.8ms
  long host ops >1ms: n=1083 total=52334.1ms; top10:
    +     -7.5ms 11055.2ms [Trace] PyTorch Profiler (0)
    +   3249.5ms  7771.5ms [user_annotation] recv_res_dict_from_prev_stage
    +   3249.7ms  7768.0ms [user_annotation] gloo:recv
    +   2961.9ms   287.2ms [user_annotation] gloo:send
    +   2962.0ms   287.2ms [user_annotation] gloo:send
    +   1298.2ms   283.7ms [user_annotation] gloo:send
    +   1298.2ms   283.7ms [user_annotation] gloo:send
    +   1860.6ms   283.3ms [user_annotation] gloo:send
    +   1860.7ms   283.3ms [user_annotation] gloo:send
    +    753.5ms   276.9ms [user_annotation] gloo:send
    nccl_AllReduce    1753.4ms  n=  1044  (15.9% span, 80.7% busy)
    cutlass_moe        348.1ms  n=  9876  (3.2% span, 16.0% busy)
    sparse_mla          27.1ms  n=  1008  (0.2% span, 1.2% busy)
    nccl_other          22.3ms  n=    48  (0.2% span, 1.0% busy)
    other               17.5ms  n=  8094  (0.2% span, 0.8% busy)
    mqa_logits           3.0ms  n=   180  (0.0% span, 0.1% busy)
    nccl_SendRecv        2.4ms  n=   207  (0.0% span, 0.1% busy)

