
NOTICE: Existing SQLite export found: /report/profiles/node_5/nsys/d0.sqlite
        It is assumed file was previously exported from: /report/profiles/node_5/nsys/d0.nsys-rep
        Consider using --force-export=true if needed.

Processing [/report/profiles/node_5/nsys/d0.sqlite] with [/opt/nvidia/nsight-systems-cli/2026.4.1/target-linux-x64/reports/nvtx_kern_sum.py]... 

 ** NVTX Range Kernel Summary (nvtx_kern_sum):

 NVTX Range  Style  PID  TID  NVTX Inst  Kern Inst  Total Time (ns)  Avg (ns)  Med (ns)  Min (ns)  Max (ns)  StdDev (ns)                                              Kernel Name                                             
 ----------  -----  ---  ---  ---------  ---------  ---------------  --------  --------  --------  --------  -----------  ----------------------------------------------------------------------------------------------------
                    261  261          0         20          313,034  15,651.7  15,584.5    15,265    16,864        356.0  create_flashinfer_kv_indices_triton                                                                 
                    261  261          0         40          231,910   5,797.8   2,432.0     1,312    11,264      4,063.1  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    261  261          0         42           97,571   2,323.1   2,288.0     1,312     2,912        336.8  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    261  261          0         10           43,682   4,368.2   4,912.0     1,408     5,568      1,519.0  alloc_extend_kernel                                                                                 
                    261  261          0         36           39,618   1,100.5   1,104.5       736     1,600        211.8  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    261  261          0         23           34,784   1,512.3   1,440.0     1,408     2,176        177.2  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    261  261          0         20           30,240   1,512.0   1,504.0     1,472     1,568         34.2  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
                    261  261          0         20           23,489   1,174.5   1,184.0     1,121     1,312         37.5  _fused_replay_state_indices_kernel                                                                  
                    261  261          0         20           21,472   1,073.6   1,056.0     1,056     1,184         30.2  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
                    261  261          0         20           20,992   1,049.6   1,056.0     1,024     1,184         35.4  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
                    261  261          0         20           14,209     710.5     704.0       672       800         36.8  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
                    261  261          0          8           12,897   1,612.1   1,600.0     1,536     1,792         88.7  void at::native::elementwise_kernel<(int)128, (int)2, void at::native::gpu_kernel_impl_nocast<at::n…
                    261  261          0          8           10,018   1,252.3   1,120.0     1,089     2,240        399.4  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    261  261          0          2            2,849   1,424.5   1,424.5     1,344     1,505        113.8  assign_req_to_token_pool                                                                            
                    261  261          0          2            2,656   1,328.0   1,328.0     1,216     1,440        158.4  get_last_loc_kernel                                                                                 
                    261  261          0          1            2,112   2,112.0   2,112.0     2,112     2,112          0.0  void at::native::vectorized_elementwise_kernel<(int)4, at::native::BUnaryFunctor<int, int, int, at:…
                    261  261          0          1            1,152   1,152.0   1,152.0     1,152     1,152          0.0  void at::native::<unnamed>::CatArrayBatchedCopy_alignedK_contig<at::native::<unnamed>::OpaqueType<(…
                    262  262          0         19          287,718  15,143.1  15,136.0    14,848    15,809        231.0  create_flashinfer_kv_indices_triton                                                                 
                    262  262          0         40          156,423   3,910.6   2,432.5     1,312     7,937      2,071.1  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    262  262          0         40           93,028   2,325.7   2,368.0     1,344     3,009        314.5  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    262  262          0         10           43,266   4,326.6   4,816.5     1,377     5,728      1,511.9  alloc_extend_kernel                                                                                 
                    262  262          0         35           38,305   1,094.4   1,088.0       768     1,600        211.1  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    262  262          0         22           32,353   1,470.6   1,408.0     1,376     2,112        177.3  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    262  262          0         19           28,576   1,504.0   1,472.0     1,408     1,632         69.1  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
                    262  262          0         19           22,753   1,197.5   1,184.0     1,152     1,312         37.4  _fused_replay_state_indices_kernel                                                                  
                    262  262          0         19           19,552   1,029.1   1,024.0       992     1,152         34.2  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
                    262  262          0         19           16,096     847.2     832.0       832       864         16.4  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
                    262  262          0         19           13,089     688.9     672.0       672       768         27.0  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
                    262  262          0          8           12,416   1,552.0   1,536.0     1,472     1,696         80.2  void at::native::elementwise_kernel<(int)128, (int)2, void at::native::gpu_kernel_impl_nocast<at::n…
                    262  262          0          8           10,049   1,256.1   1,120.0     1,088     2,272        410.7  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    262  262          0          2            2,912   1,456.0   1,456.0     1,280     1,632        248.9  assign_req_to_token_pool                                                                            
                    262  262          0          2            2,624   1,312.0   1,312.0     1,184     1,440        181.0  get_last_loc_kernel                                                                                 
                    262  262          0          1            2,112   2,112.0   2,112.0     2,112     2,112          0.0  void at::native::vectorized_elementwise_kernel<(int)4, at::native::BUnaryFunctor<int, int, int, at:…
                    262  262          0          1            1,152   1,152.0   1,152.0     1,152     1,152          0.0  void at::native::<unnamed>::CatArrayBatchedCopy_alignedK_contig<at::native::<unnamed>::OpaqueType<(…
                    263  263          0         19          289,769  15,251.0  15,265.0    14,912    15,681        198.3  create_flashinfer_kv_indices_triton                                                                 
                    263  263          0         40          150,406   3,760.2   2,448.5     1,312     6,752      1,769.6  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    263  263          0         40           93,731   2,343.3   2,288.5     1,344     2,720        291.4  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    263  263          0         10           43,585   4,358.5   4,928.0     1,344     5,440      1,533.4  alloc_extend_kernel                                                                                 
                    263  263          0         35           39,330   1,123.7   1,120.0       768     1,601        209.6  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    263  263          0         22           33,410   1,518.6   1,440.0     1,344     2,176        181.2  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    263  263          0         19           28,096   1,478.7   1,472.0     1,440     1,568         29.4  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
                    263  263          0         19           22,977   1,209.3   1,184.0     1,152     1,344         46.0  _fused_replay_state_indices_kernel                                                                  
                    263  263          0         19           20,161   1,061.1   1,056.0       992     1,120         30.7  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
                    263  263          0         19           16,417     864.1     864.0       832       897         15.2  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
                    263  263          0         19           13,537     712.5     704.0       704       768         20.9  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
                    263  263          0          8           12,800   1,600.0   1,568.0     1,504     1,792        101.2  void at::native::elementwise_kernel<(int)128, (int)2, void at::native::gpu_kernel_impl_nocast<at::n…
                    263  263          0          8           10,304   1,288.0   1,152.0     1,120     2,304        410.8  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    263  263          0          2            2,944   1,472.0   1,472.0     1,344     1,600        181.0  assign_req_to_token_pool                                                                            
                    263  263          0          2            2,752   1,376.0   1,376.0     1,312     1,440         90.5  get_last_loc_kernel                                                                                 
                    263  263          0          1            2,144   2,144.0   2,144.0     2,144     2,144          0.0  void at::native::vectorized_elementwise_kernel<(int)4, at::native::BUnaryFunctor<int, int, int, at:…
                    263  263          0          1            1,120   1,120.0   1,120.0     1,120     1,120          0.0  void at::native::<unnamed>::CatArrayBatchedCopy_alignedK_contig<at::native::<unnamed>::OpaqueType<(…
                    264  264          0         19          296,970  15,630.0  15,584.0    15,105    16,256        264.1  create_flashinfer_kv_indices_triton                                                                 
                    264  264          0         40          223,749   5,593.7   2,528.0     1,344    12,000      3,720.1  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    264  264          0         40           94,594   2,364.8   2,240.0     1,408     2,912        332.4  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    264  264          0         10           43,841   4,384.1   4,928.0     1,408     5,568      1,502.1  alloc_extend_kernel                                                                                 
                    264  264          0         35           39,072   1,116.3   1,120.0       768     1,600        209.2  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    264  264          0         22           33,440   1,520.0   1,456.0     1,376     2,176        177.7  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    264  264          0         19           28,352   1,492.2   1,504.0     1,408     1,600         48.0  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
                    264  264          0         19           22,656   1,192.4   1,184.0     1,152     1,280         27.9  _fused_replay_state_indices_kernel                                                                  
                    264  264          0         19           20,225   1,064.5   1,056.0     1,024     1,152         25.8  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
                    264  264          0         19           16,640     875.8     864.0       864       896         15.9  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
                    264  264          0         19           13,664     719.2     704.0       704       832         36.0  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
                    264  264          0          8           12,737   1,592.1   1,568.0     1,504     1,761         78.2  void at::native::elementwise_kernel<(int)128, (int)2, void at::native::gpu_kernel_impl_nocast<at::n…
                    264  264          0          8           10,240   1,280.0   1,136.0     1,120     2,304        414.1  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    264  264          0          2            2,880   1,440.0   1,440.0     1,344     1,536        135.8  assign_req_to_token_pool                                                                            
                    264  264          0          2            2,688   1,344.0   1,344.0     1,280     1,408         90.5  get_last_loc_kernel                                                                                 
                    264  264          0          1            2,080   2,080.0   2,080.0     2,080     2,080          0.0  void at::native::vectorized_elementwise_kernel<(int)4, at::native::BUnaryFunctor<int, int, int, at:…
                    264  264          0          1            1,184   1,184.0   1,184.0     1,184     1,184          0.0  void at::native::<unnamed>::CatArrayBatchedCopy_alignedK_contig<at::native::<unnamed>::OpaqueType<(…
                    265  265          0         19          291,720  15,353.7  15,361.0    14,976    15,936        282.7  create_flashinfer_kv_indices_triton                                                                 
                    265  265          0         40          168,423   4,210.6   2,496.0     1,344     7,968      2,295.6  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    265  265          0         40           95,168   2,379.2   2,528.0     1,312     2,720        307.1  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    265  265          0         10           43,969   4,396.9   4,960.0     1,408     5,600      1,530.8  alloc_extend_kernel                                                                                 
                    265  265          0         35           39,362   1,124.6   1,120.0       768     1,632        212.3  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    265  265          0         22           33,056   1,502.5   1,440.0     1,408     2,144        179.3  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    265  265          0         19           28,609   1,505.7   1,504.0     1,440     1,664         43.3  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
                    265  265          0         19           22,561   1,187.4   1,184.0     1,152     1,344         48.8  _fused_replay_state_indices_kernel                                                                  
                    265  265          0         19           19,872   1,045.9   1,056.0     1,024     1,056         15.3  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
                    265  265          0         19           16,224     853.9     864.0       832       864         15.3  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
                    265  265          0         19           13,376     704.0     704.0       672       736         10.7  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
                    265  265          0          8           12,672   1,584.0   1,552.0     1,536     1,728         68.4  void at::native::elementwise_kernel<(int)128, (int)2, void at::native::gpu_kernel_impl_nocast<at::n…
                    265  265          0          8           10,400   1,300.0   1,152.0     1,152     2,272        393.0  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    265  265          0          2            2,912   1,456.0   1,456.0     1,376     1,536        113.1  assign_req_to_token_pool                                                                            
                    265  265          0          2            2,848   1,424.0   1,424.0     1,152     1,696        384.7  get_last_loc_kernel                                                                                 
                    265  265          0          1            2,112   2,112.0   2,112.0     2,112     2,112          0.0  void at::native::vectorized_elementwise_kernel<(int)4, at::native::BUnaryFunctor<int, int, int, at:…
                    265  265          0          1            1,152   1,152.0   1,152.0     1,152     1,152          0.0  void at::native::<unnamed>::CatArrayBatchedCopy_alignedK_contig<at::native::<unnamed>::OpaqueType<(…
                    266  266          0         19          288,424  15,180.2  15,168.0    14,752    15,841        278.2  create_flashinfer_kv_indices_triton                                                                 
                    266  266          0         40          162,342   4,058.6   2,400.5     1,280     7,937      2,129.1  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    266  266          0         40           92,098   2,302.4   2,176.0     1,312     2,752        324.2  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    266  266          0         10           43,712   4,371.2   4,912.0     1,344     5,664      1,514.8  alloc_extend_kernel                                                                                 
                    266  266          0         35           38,850   1,110.0   1,088.0       768     1,601        221.6  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    266  266          0         22           32,129   1,460.4   1,376.0     1,344     2,176        194.9  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    266  266          0         19           28,224   1,485.5   1,472.0     1,440     1,632         43.1  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
                    266  266          0         19           21,952   1,155.4   1,152.0     1,120     1,344         54.3  _fused_replay_state_indices_kernel                                                                  
                    266  266          0         19           19,618   1,032.5   1,024.0       992     1,344         77.5  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
                    266  266          0         19           16,001     842.2     832.0       832       865         15.4  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
                    266  266          0         19           12,930     680.5     672.0       640       800         39.7  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
                    266  266          0          8           12,578   1,572.3   1,536.5     1,504     1,728         91.1  void at::native::elementwise_kernel<(int)128, (int)2, void at::native::gpu_kernel_impl_nocast<at::n…
                    266  266          0          8           10,081   1,260.1   1,120.0     1,120     2,240        395.9  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    266  266          0          2            2,752   1,376.0   1,376.0     1,312     1,440         90.5  assign_req_to_token_pool                                                                            
                    266  266          0          2            2,657   1,328.5   1,328.5     1,152     1,505        249.6  get_last_loc_kernel                                                                                 
                    266  266          0          1            2,112   2,112.0   2,112.0     2,112     2,112          0.0  void at::native::vectorized_elementwise_kernel<(int)4, at::native::BUnaryFunctor<int, int, int, at:…
                    266  266          0          1            1,248   1,248.0   1,248.0     1,248     1,248          0.0  void at::native::<unnamed>::CatArrayBatchedCopy_alignedK_contig<at::native::<unnamed>::OpaqueType<(…
                    267  267          0         19          287,978  15,156.7  15,137.0    14,784    15,872        289.5  create_flashinfer_kv_indices_triton                                                                 
                    267  267          0         40          150,758   3,768.9   2,448.0     1,312     7,424      1,922.2  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    267  267          0         40           94,307   2,357.7   2,416.5     1,312     2,688        272.4  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    267  267          0         10           42,721   4,272.1   4,704.0     1,344     5,600      1,500.9  alloc_extend_kernel                                                                                 
                    267  267          0         35           38,630   1,103.7   1,120.0       768     1,697        218.3  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    267  267          0         22           32,896   1,495.3   1,440.0     1,376     2,144        173.9  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    267  267          0         19           28,256   1,487.2   1,472.0     1,440     1,568         39.0  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
                    267  267          0         19           22,401   1,179.0   1,184.0     1,152     1,280         35.8  _fused_replay_state_indices_kernel                                                                  
                    267  267          0         19           19,840   1,044.2   1,024.0       992     1,152         37.3  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
                    267  267          0         19           16,001     842.2     832.0       832       864         15.2  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
                    267  267          0         19           13,600     715.8     704.0       672       896         54.6  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
                    267  267          0          8           12,480   1,560.0   1,536.0     1,472     1,728         83.4  void at::native::elementwise_kernel<(int)128, (int)2, void at::native::gpu_kernel_impl_nocast<at::n…
                    267  267          0          8           10,368   1,296.0   1,152.0     1,120     2,400        446.4  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    267  267          0          2            2,912   1,456.0   1,456.0     1,344     1,568        158.4  assign_req_to_token_pool                                                                            
                    267  267          0          2            2,784   1,392.0   1,392.0     1,280     1,504        158.4  get_last_loc_kernel                                                                                 
                    267  267          0          1            2,272   2,272.0   2,272.0     2,272     2,272          0.0  void at::native::vectorized_elementwise_kernel<(int)4, at::native::BUnaryFunctor<int, int, int, at:…
                    267  267          0          1            1,216   1,216.0   1,216.0     1,216     1,216          0.0  void at::native::<unnamed>::CatArrayBatchedCopy_alignedK_contig<at::native::<unnamed>::OpaqueType<(…
                    268  268          0         19          287,625  15,138.2  15,072.0    14,816    15,681        238.3  create_flashinfer_kv_indices_triton                                                                 
                    268  268          0         40          161,350   4,033.8   2,464.0     1,312     7,840      2,053.7  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    268  268          0         40           94,273   2,356.8   2,496.0     1,312     2,784        304.2  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    268  268          0         10           43,586   4,358.6   4,800.0     1,344     5,569      1,523.1  alloc_extend_kernel                                                                                 
                    268  268          0         35           39,009   1,114.5   1,120.0       768     1,568        207.0  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    268  268          0         22           33,760   1,534.5   1,472.0     1,408     2,112        160.3  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
                    268  268          0         19           27,620   1,453.7   1,440.0     1,376     1,536         35.9  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
                    268  268          0         19           22,752   1,197.5   1,184.0     1,152     1,280         35.9  _fused_replay_state_indices_kernel                                                                  
                    268  268          0         19           20,096   1,057.7   1,056.0     1,024     1,088         19.9  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
                    268  268          0         19           16,288     857.3     864.0       832       864         13.4  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
                    268  268          0         19           13,888     730.9     704.0       704       928         54.7  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
                    268  268          0          8           12,801   1,600.1   1,584.0     1,504     1,697         66.5  void at::native::elementwise_kernel<(int)128, (int)2, void at::native::gpu_kernel_impl_nocast<at::n…
                    268  268          0          8           10,304   1,288.0   1,152.0     1,120     2,272        397.8  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
                    268  268          0          2            3,008   1,504.0   1,504.0     1,408     1,600        135.8  assign_req_to_token_pool                                                                            
                    268  268          0          2            2,721   1,360.5   1,360.5     1,152     1,569        294.9  get_last_loc_kernel                                                                                 
                    268  268          0          1            2,048   2,048.0   2,048.0     2,048     2,048          0.0  void at::native::vectorized_elementwise_kernel<(int)4, at::native::BUnaryFunctor<int, int, int, at:…
                    268  268          0          1            1,216   1,216.0   1,216.0     1,216     1,216          0.0  void at::native::<unnamed>::CatArrayBatchedCopy_alignedK_contig<at::native::<unnamed>::OpaqueType<(…
 :dflash_a…  Push…  261  261         20         20          160,771   8,038.6   8,064.0     7,936     8,160         60.2  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<float, at::native::ArgMaxOps<…
 :dflash_a…  Push…  261  261         20         20           48,066   2,403.3   2,400.0     2,336     2,593         54.0  _dflash_accept_bonus_contig_kernel                                                                  
 :dflash_a…  Push…  262  262         19         19          151,973   7,998.6   7,968.0     7,904     8,225         79.6  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<float, at::native::ArgMaxOps<…
 :dflash_a…  Push…  262  262         19         19           44,322   2,332.7   2,304.0     2,272     2,464         60.2  _dflash_accept_bonus_contig_kernel                                                                  
 :dflash_a…  Push…  263  263         19         19          152,969   8,051.0   8,033.0     7,937     8,193         70.1  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<float, at::native::ArgMaxOps<…
 :dflash_a…  Push…  263  263         19         19           46,915   2,469.2   2,432.0     2,400     2,592         65.2  _dflash_accept_bonus_contig_kernel                                                                  
 :dflash_a…  Push…  264  264         19         19          152,162   8,008.5   8,000.0     7,904     8,160         81.1  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<float, at::native::ArgMaxOps<…
 :dflash_a…  Push…  264  264         19         19           47,041   2,475.8   2,432.0     2,400     2,656         76.3  _dflash_accept_bonus_contig_kernel                                                                  
 :dflash_a…  Push…  265  265         19         19          153,381   8,072.7   8,064.0     7,968     8,256         69.9  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<float, at::native::ArgMaxOps<…
 :dflash_a…  Push…  265  265         19         19           46,624   2,453.9   2,432.0     2,400     2,592         47.8  _dflash_accept_bonus_contig_kernel                                                                  
 :dflash_a…  Push…  266  266         19         19          154,884   8,151.8   8,160.0     8,032     8,384         97.7  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<float, at::native::ArgMaxOps<…
 :dflash_a…  Push…  266  266         19         19           44,193   2,325.9   2,304.0     2,240     2,528         68.4  _dflash_accept_bonus_contig_kernel                                                                  
 :dflash_a…  Push…  267  267         19         19          153,575   8,082.9   8,065.0     7,905     8,289        103.5  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<float, at::native::ArgMaxOps<…
 :dflash_a…  Push…  267  267         19         19           45,827   2,411.9   2,400.0     2,368     2,496         38.7  _dflash_accept_bonus_contig_kernel                                                                  
 :dflash_a…  Push…  268  268         19         19          151,042   7,949.6   7,968.0     7,872     8,032         45.7  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<float, at::native::ArgMaxOps<…
 :dflash_a…  Push…  268  268         19         19           47,106   2,479.3   2,464.0     2,432     2,624         50.5  _dflash_accept_bonus_contig_kernel                                                                  
 :dflash_d…  Push…  261  261         20        260      140,794,412  541,517…  271,096…   216,455  52,394,…  3,289,463.0  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_d…  Push…  261  261         20        120       36,515,931  304,299…  305,162…   284,937   309,354      3,720.2  _fwd_kernel                                                                                         
 :dflash_d…  Push…  261  261         20         40        3,633,872  90,846.8  89,763.0    71,554   138,884     13,082.0  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_d…  Push…  261  261         20        240        2,261,312   9,422.1  10,720.0     5,888    16,768      3,072.8  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_d…  Push…  261  261         20         20        1,259,143  62,957.2  62,930.0    61,954    65,570        737.9  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_d…  Push…  261  261         20        120          854,677   7,122.3   7,024.0     6,561     9,760        492.2  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_d…  Push…  261  261         20        260          574,224   2,208.6   2,240.5     1,568     2,336        182.2  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_d…  Push…  261  261         20        240          562,578   2,344.1   2,336.0     2,240     2,560         54.8  kernel_cutlass_kernel_flashinfernormkernelsfused_add_rmsnormFusedAddRMSNormKernel_object_at__tensor…
 :dflash_d…  Push…  261  261         20        120          388,234   3,235.3   3,184.0     3,008     4,033        192.6  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_d…  Push…  261  261         20        240          377,353   1,572.3   1,536.5     1,408     1,888        123.8  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_d…  Push…  261  261         20        120          363,851   3,032.1   3,040.0     2,784     3,136         43.0  void sglang::act_and_mul_kernel<__nv_bfloat16, (sglang::ActivationKind)0, (bool)1, (bool)0, (bool)0…
 :dflash_d…  Push…  261  261         20         40          352,554   8,813.9   5,568.0     2,912    15,745      6,067.4  create_flashinfer_kv_indices_triton                                                                 
 :dflash_d…  Push…  261  261         20         20          190,054   9,502.7   9,504.0     9,409     9,632         52.3  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ma…
 :dflash_d…  Push…  261  261         20        120          148,551   1,237.9   1,216.0     1,152     1,536         76.5  _table_qk_norm_rope_kernel                                                                          
 :dflash_d…  Push…  261  261         20        120          133,187   1,109.9   1,088.0     1,024     1,888        125.1  void sglang::store_kvcache<(long)256, (long)256, (int)1, (bool)1, long>(sglang::StoreKVCacheParams) 
 :dflash_d…  Push…  261  261         20         20           76,066   3,803.3   3,744.0     3,680     4,001        124.3  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ar…
 :dflash_d…  Push…  261  261         20         60           72,642   1,210.7   1,040.5       992     1,697        262.3  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 :dflash_d…  Push…  261  261         20         60           56,805     946.8     960.0       736     1,312        153.1  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  261  261         20         20           42,402   2,120.1   2,112.0     2,080     2,208         40.0  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_d…  Push…  261  261         20         60           35,776     596.3     528.0       512       800        110.1  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 :dflash_d…  Push…  261  261         20         20           31,970   1,598.5   1,568.5     1,536     1,760         60.9  void at::native::_scatter_gather_elementwise_kernel<(int)128, (int)8, void at::native::_cuda_scatte…
 :dflash_d…  Push…  261  261         20         20           27,488   1,374.4   1,376.0     1,184     1,536        105.6  _prepare_dflash_draft_block_contig_kernel                                                           
 :dflash_d…  Push…  261  261         20         20           27,201   1,360.0   1,344.0     1,184     1,568         84.0  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  261  261         20         20           26,913   1,345.7   1,344.0     1,312     1,440         33.6  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
 :dflash_d…  Push…  261  261         20         20           24,416   1,220.8   1,344.0       896     1,536        249.9  _vocab_parallel_embedding_kernel                                                                    
 :dflash_d…  Push…  261  261         20         20           21,474   1,073.7   1,056.0     1,024     1,153         28.6  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_d…  Push…  261  261         20         20           20,865   1,043.3   1,056.0       992     1,120         26.3  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  261  261         20         20           20,640   1,032.0   1,024.0       992     1,184         41.4  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  261  261         20         20           19,745     987.3     992.0       960     1,024         15.7  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctor_add<long>, std::arra…
 :dflash_d…  Push…  261  261         20         20           13,920     696.0     704.0       672       736         20.4  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
 :dflash_d…  Push…  262  262         19        247       87,490,381  354,212…  273,575…   226,502  12,461,…    819,109.8  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_d…  Push…  262  262         19        114       34,132,683  299,409…  299,048…   295,272   305,705      1,956.1  _fwd_kernel                                                                                         
 :dflash_d…  Push…  262  262         19         38        3,422,078  90,054.7  88,434.5    73,122   142,692     13,016.6  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_d…  Push…  262  262         19        228        2,105,623   9,235.2  10,272.5     5,824    16,993      3,015.5  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_d…  Push…  262  262         19         19        1,193,856  62,834.5  62,754.0    62,338    64,002        402.7  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_d…  Push…  262  262         19        114          818,099   7,176.3   7,120.0     6,592     8,960        432.7  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_d…  Push…  262  262         19        247          538,122   2,178.6   2,208.0     1,568     2,305        179.8  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_d…  Push…  262  262         19        228          523,692   2,296.9   2,273.0     2,208     2,624         56.2  kernel_cutlass_kernel_flashinfernormkernelsfused_add_rmsnormFusedAddRMSNormKernel_object_at__tensor…
 :dflash_d…  Push…  262  262         19        114          362,118   3,176.5   3,104.0     3,008     3,809        177.5  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_d…  Push…  262  262         19        228          351,372   1,541.1   1,536.0     1,408     1,856        127.0  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_d…  Push…  262  262         19        114          340,589   2,987.6   2,992.5     2,784     3,072         44.2  void sglang::act_and_mul_kernel<__nv_bfloat16, (sglang::ActivationKind)0, (bool)1, (bool)0, (bool)0…
 :dflash_d…  Push…  262  262         19         38          337,862   8,891.1   5,584.0     2,880    15,617      6,089.8  create_flashinfer_kv_indices_triton                                                                 
 :dflash_d…  Push…  262  262         19         19          179,494   9,447.1   9,441.0     9,344     9,600         76.6  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ma…
 :dflash_d…  Push…  262  262         19        114          141,410   1,240.4   1,216.0     1,120     1,504         81.8  _table_qk_norm_rope_kernel                                                                          
 :dflash_d…  Push…  262  262         19        114          125,956   1,104.9   1,056.0     1,024     1,792        110.0  void sglang::store_kvcache<(long)256, (long)256, (int)1, (bool)1, long>(sglang::StoreKVCacheParams) 
 :dflash_d…  Push…  262  262         19         19           70,566   3,714.0   3,712.0     3,648     3,840         44.6  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ar…
 :dflash_d…  Push…  262  262         19         57           68,193   1,196.4   1,024.0       992     1,632        254.8  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 :dflash_d…  Push…  262  262         19         57           53,794     943.8     960.0       736     1,216        146.1  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  262  262         19         38           42,432   1,116.6   1,120.0     1,024     1,216         63.5  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_d…  Push…  262  262         19         19           40,193   2,115.4   2,080.0     2,016     2,304         68.2  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_d…  Push…  262  262         19         57           34,528     605.8     544.0       512       960        125.7  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 :dflash_d…  Push…  262  262         19         19           30,688   1,615.2   1,600.0     1,536     1,792         70.2  void at::native::_scatter_gather_elementwise_kernel<(int)128, (int)8, void at::native::_cuda_scatte…
 :dflash_d…  Push…  262  262         20         20           27,744   1,387.2   1,392.0     1,152     1,536         72.9  _prepare_dflash_draft_block_contig_kernel                                                           
 :dflash_d…  Push…  262  262         19         19           27,587   1,451.9   1,440.0     1,216     1,632         90.1  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  262  262         19         19           24,961   1,313.7   1,312.0     1,280     1,376         29.2  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
 :dflash_d…  Push…  262  262         19         19           20,064   1,056.0   1,024.0       992     1,184         51.2  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  262  262         19         19           19,746   1,039.3   1,024.0       992     1,216         50.4  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  262  262         20         20           19,265     963.3     928.0       896     1,344        131.7  _vocab_parallel_embedding_kernel                                                                    
 :dflash_d…  Push…  262  262         19         19           18,817     990.4     992.0       929     1,120         44.5  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctor_add<long>, std::arra…
 :dflash_d…  Push…  262  262         19         19           13,408     705.7     704.0       672       832         39.2  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
 :dflash_d…  Push…  263  263         19        247       85,594,391  346,536…  270,728…   226,983  11,108,…    738,478.8  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_d…  Push…  263  263         19        114       34,611,162  303,606…  304,121…   296,809   309,034      2,610.9  _fwd_kernel                                                                                         
 :dflash_d…  Push…  263  263         19         38        3,456,843  90,969.6  90,706.5    75,522   128,835     11,687.4  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_d…  Push…  263  263         19        228        2,112,961   9,267.4  10,464.0     5,856    17,249      3,029.8  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_d…  Push…  263  263         19         19        1,194,758  62,882.0  62,754.0    62,466    63,874        362.8  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_d…  Push…  263  263         19        114          813,916   7,139.6   7,104.0     6,592     8,992        401.9  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_d…  Push…  263  263         19        247          541,676   2,193.0   2,208.0     1,568     2,337        183.4  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_d…  Push…  263  263         19        228          529,992   2,324.5   2,304.0     2,176     2,624         64.6  kernel_cutlass_kernel_flashinfernormkernelsfused_add_rmsnormFusedAddRMSNormKernel_object_at__tensor…
 :dflash_d…  Push…  263  263         19        114          361,741   3,173.2   3,136.0     3,008     3,841        154.0  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_d…  Push…  263  263         19        228          355,566   1,559.5   1,536.5     1,408     1,952        130.2  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_d…  Push…  263  263         19        114          345,319   3,029.1   3,040.0     2,912     3,104         31.3  void sglang::act_and_mul_kernel<__nv_bfloat16, (sglang::ActivationKind)0, (bool)1, (bool)0, (bool)0…
 :dflash_d…  Push…  263  263         19         38          339,688   8,939.2   5,616.0     2,944    16,064      6,111.3  create_flashinfer_kv_indices_triton                                                                 
 :dflash_d…  Push…  263  263         19         19          181,223   9,538.1   9,536.0     9,472     9,664         44.6  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ma…
 :dflash_d…  Push…  263  263         19        114          142,565   1,250.6   1,216.0     1,152     1,728         92.2  _table_qk_norm_rope_kernel                                                                          
 :dflash_d…  Push…  263  263         19        114          127,458   1,118.1   1,088.0     1,024     1,888        132.4  void sglang::store_kvcache<(long)256, (long)256, (int)1, (bool)1, long>(sglang::StoreKVCacheParams) 
 :dflash_d…  Push…  263  263         19         19           71,681   3,772.7   3,776.0     3,712     3,936         47.6  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ar…
 :dflash_d…  Push…  263  263         19         57           68,867   1,208.2   1,056.0       960     1,664        253.3  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 :dflash_d…  Push…  263  263         19         57           54,019     947.7     960.0       768     1,344        143.4  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  263  263         19         38           43,424   1,142.7   1,152.0     1,024     1,376         80.1  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_d…  Push…  263  263         19         19           40,322   2,122.2   2,112.0     2,048     2,240         47.7  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_d…  Push…  263  263         19         57           34,370     603.0     513.0       512       928        124.6  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 :dflash_d…  Push…  263  263         19         19           30,720   1,616.8   1,600.0     1,536     1,760         50.4  void at::native::_scatter_gather_elementwise_kernel<(int)128, (int)8, void at::native::_cuda_scatte…
 :dflash_d…  Push…  263  263         20         20           28,673   1,433.7   1,424.0     1,152     1,600         95.0  _prepare_dflash_draft_block_contig_kernel                                                           
 :dflash_d…  Push…  263  263         19         19           26,752   1,408.0   1,408.0     1,216     1,568         79.1  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  263  263         19         19           24,896   1,310.3   1,312.0     1,248     1,472         51.7  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
 :dflash_d…  Push…  263  263         19         19           20,449   1,076.3   1,056.0       992     1,216         71.8  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  263  263         19         19           20,065   1,056.1   1,056.0     1,024     1,216         41.3  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  263  263         19         19           18,945     997.1     992.0       960     1,120         37.3  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctor_add<long>, std::arra…
 :dflash_d…  Push…  263  263         20         20           18,656     932.8     896.0       896     1,408        112.9  _vocab_parallel_embedding_kernel                                                                    
 :dflash_d…  Push…  263  263         19         19           13,443     707.5     704.0       672       768         33.5  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
 :dflash_d…  Push…  264  264         19        247       84,688,100  342,866…  271,015…   244,231  10,379,…    695,343.1  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_d…  Push…  264  264         19        114       34,887,481  306,030…  306,120…   302,536   309,608      1,425.5  _fwd_kernel                                                                                         
 :dflash_d…  Push…  264  264         19         38        3,210,455  84,485.7  83,506.0    74,562   101,250      6,090.0  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_d…  Push…  264  264         19        228        2,122,590   9,309.6  10,896.0     5,856    16,545      3,037.6  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_d…  Push…  264  264         19         19        1,197,184  63,009.7  62,946.0    62,242    64,514        489.4  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_d…  Push…  264  264         19        114          809,586   7,101.6   7,040.0     6,592     9,536        407.3  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_d…  Push…  264  264         19        247          542,195   2,195.1   2,240.0     1,600     2,336        174.6  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_d…  Push…  264  264         19        228          531,084   2,329.3   2,304.0     2,240     2,497         46.0  kernel_cutlass_kernel_flashinfernormkernelsfused_add_rmsnormFusedAddRMSNormKernel_object_at__tensor…
 :dflash_d…  Push…  264  264         19        114          363,658   3,190.0   3,168.0     3,008     3,584        138.1  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_d…  Push…  264  264         19        228          360,870   1,582.8   1,568.0     1,440     1,984        141.5  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_d…  Push…  264  264         19        114          344,590   3,022.7   3,040.0     2,912     3,072         32.8  void sglang::act_and_mul_kernel<__nv_bfloat16, (sglang::ActivationKind)0, (bool)1, (bool)0, (bool)0…
 :dflash_d…  Push…  264  264         19         38          342,891   9,023.4   5,856.0     2,912    15,616      6,129.0  create_flashinfer_kv_indices_triton                                                                 
 :dflash_d…  Push…  264  264         19         19          182,470   9,603.7   9,600.0     9,536     9,696         43.7  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ma…
 :dflash_d…  Push…  264  264         19        114          140,997   1,236.8   1,216.0     1,152     1,536         79.2  _table_qk_norm_rope_kernel                                                                          
 :dflash_d…  Push…  264  264         19        114          129,666   1,137.4   1,120.0     1,056     1,888        119.9  void sglang::store_kvcache<(long)256, (long)256, (int)1, (bool)1, long>(sglang::StoreKVCacheParams) 
 :dflash_d…  Push…  264  264         19         19           71,938   3,786.2   3,776.0     3,776     3,809         15.4  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ar…
 :dflash_d…  Push…  264  264         19         57           70,593   1,238.5   1,056.0     1,024     1,696        259.3  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 :dflash_d…  Push…  264  264         19         57           55,554     974.6     992.0       768     1,216        148.5  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  264  264         19         38           44,704   1,176.4   1,152.0     1,056     1,472        113.4  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_d…  Push…  264  264         19         19           40,258   2,118.8   2,081.0     2,048     2,272         67.1  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_d…  Push…  264  264         19         57           34,947     613.1     544.0       512       960        118.2  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 :dflash_d…  Push…  264  264         19         19           31,936   1,680.8   1,696.0     1,600     1,792         57.8  void at::native::_scatter_gather_elementwise_kernel<(int)128, (int)8, void at::native::_cuda_scatte…
 :dflash_d…  Push…  264  264         20         20           28,673   1,433.7   1,424.0     1,184     1,600         84.7  _prepare_dflash_draft_block_contig_kernel                                                           
 :dflash_d…  Push…  264  264         19         19           27,808   1,463.6   1,472.0     1,248     1,632         83.2  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  264  264         19         19           24,768   1,303.6   1,312.0     1,248     1,376         33.5  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
 :dflash_d…  Push…  264  264         19         19           20,449   1,076.3   1,056.0     1,024     1,312         61.5  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  264  264         19         19           19,680   1,035.8   1,024.0     1,024     1,056         15.9  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  264  264         19         19           19,648   1,034.1   1,024.0       992     1,216         48.9  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctor_add<long>, std::arra…
 :dflash_d…  Push…  264  264         20         20           18,465     923.3     896.0       896     1,312         92.4  _vocab_parallel_embedding_kernel                                                                    
 :dflash_d…  Push…  264  264         19         19           14,080     741.1     736.0       704       800         26.7  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
 :dflash_d…  Push…  265  265         19        247       84,729,271  343,033…  271,016…   262,247  10,613,…    707,602.1  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_d…  Push…  265  265         19        114       34,980,506  306,846…  306,904…   303,368   310,281      1,570.2  _fwd_kernel                                                                                         
 :dflash_d…  Push…  265  265         19         38        2,723,407  71,668.6  74,258.0    53,761    94,787     14,325.6  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_d…  Push…  265  265         19        228        2,122,915   9,311.0  10,672.0     5,888    16,928      2,951.0  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_d…  Push…  265  265         19         19        1,185,602  62,400.1  62,210.0    61,922    63,522        445.1  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_d…  Push…  265  265         19        114          814,874   7,148.0   7,088.5     6,592     8,992        375.8  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_d…  Push…  265  265         19        247          543,949   2,202.2   2,240.0     1,600     2,337        176.5  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_d…  Push…  265  265         19        228          535,185   2,347.3   2,336.0     2,240     2,656         54.8  kernel_cutlass_kernel_flashinfernormkernelsfused_add_rmsnormFusedAddRMSNormKernel_object_at__tensor…
 :dflash_d…  Push…  265  265         19        114          370,407   3,249.2   3,200.0     3,072     3,840        156.8  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_d…  Push…  265  265         19        228          359,913   1,578.6   1,552.5     1,440     1,856        121.8  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_d…  Push…  265  265         19        114          348,327   3,055.5   3,072.0     2,912     3,136         38.1  void sglang::act_and_mul_kernel<__nv_bfloat16, (sglang::ActivationKind)0, (bool)1, (bool)0, (bool)0…
 :dflash_d…  Push…  265  265         19         38          339,303   8,929.0   5,568.5     2,944    15,776      6,084.2  create_flashinfer_kv_indices_triton                                                                 
 :dflash_d…  Push…  265  265         19         19          182,949   9,628.9   9,632.0     9,536     9,792         71.3  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ma…
 :dflash_d…  Push…  265  265         19        114          141,605   1,242.1   1,216.0     1,152     1,568         78.1  _table_qk_norm_rope_kernel                                                                          
 :dflash_d…  Push…  265  265         19        114          128,293   1,125.4   1,088.0     1,024     1,856        122.6  void sglang::store_kvcache<(long)256, (long)256, (int)1, (bool)1, long>(sglang::StoreKVCacheParams) 
 :dflash_d…  Push…  265  265         19         19           72,418   3,811.5   3,712.0     3,648     4,192        170.6  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ar…
 :dflash_d…  Push…  265  265         19         57           70,018   1,228.4   1,056.0     1,024     1,824        264.5  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 :dflash_d…  Push…  265  265         19         57           54,209     951.0     960.0       768     1,216        147.6  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  265  265         19         38           43,649   1,148.7   1,168.0     1,056     1,280         82.4  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_d…  Push…  265  265         19         19           40,448   2,128.8   2,112.0     2,080     2,176         24.7  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_d…  Push…  265  265         19         57           34,145     599.0     544.0       512       800        112.4  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 :dflash_d…  Push…  265  265         19         19           30,818   1,622.0   1,600.0     1,568     1,728         40.1  void at::native::_scatter_gather_elementwise_kernel<(int)128, (int)8, void at::native::_cuda_scatte…
 :dflash_d…  Push…  265  265         20         20           29,408   1,470.4   1,472.0     1,216     1,664         83.4  _prepare_dflash_draft_block_contig_kernel                                                           
 :dflash_d…  Push…  265  265         19         19           26,369   1,387.8   1,408.0     1,248     1,440         46.7  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  265  265         19         19           25,088   1,320.4   1,312.0     1,280     1,376         29.9  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
 :dflash_d…  Push…  265  265         19         19           20,193   1,062.8   1,056.0     1,024     1,152         40.7  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  265  265         19         19           19,522   1,027.5   1,024.0       992     1,057         18.2  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  265  265         19         19           18,880     993.7     992.0       960     1,088         29.2  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctor_add<long>, std::arra…
 :dflash_d…  Push…  265  265         20         20           18,115     905.8     896.0       896       928         14.9  _vocab_parallel_embedding_kernel                                                                    
 :dflash_d…  Push…  265  265         19         19           13,697     720.9     704.0       672       800         32.6  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
 :dflash_d…  Push…  266  266         19        247       83,607,964  338,493…  272,041…   262,472  9,189,2…    626,013.9  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_d…  Push…  266  266         19        114       34,807,182  305,326…  305,418…   301,289   309,642      1,631.5  _fwd_kernel                                                                                         
 :dflash_d…  Push…  266  266         19         38        2,760,857  72,654.1  74,754.5    55,874    96,323     13,814.1  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_d…  Push…  266  266         19        228        2,114,279   9,273.2  10,592.5     5,888    16,929      3,002.3  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_d…  Push…  266  266         19         19        1,193,959  62,839.9  62,786.0    62,306    64,098        417.7  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_d…  Push…  266  266         19        114          817,878   7,174.4   7,008.5     6,560     8,640        440.8  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_d…  Push…  266  266         19        247          547,095   2,215.0   2,241.0     1,568     2,336        188.4  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_d…  Push…  266  266         19        228          532,182   2,334.1   2,336.0     2,208     2,560         58.1  kernel_cutlass_kernel_flashinfernormkernelsfused_add_rmsnormFusedAddRMSNormKernel_object_at__tensor…
 :dflash_d…  Push…  266  266         19        114          367,946   3,227.6   3,200.0     3,040     4,128        179.8  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_d…  Push…  266  266         19        228          351,340   1,541.0   1,520.5     1,408     1,952        131.7  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_d…  Push…  266  266         19        114          350,157   3,071.6   3,072.0     2,816     3,137         48.5  void sglang::act_and_mul_kernel<__nv_bfloat16, (sglang::ActivationKind)0, (bool)1, (bool)0, (bool)0…
 :dflash_d…  Push…  266  266         19         38          334,953   8,814.6   5,568.0     2,816    15,488      6,037.7  create_flashinfer_kv_indices_triton                                                                 
 :dflash_d…  Push…  266  266         19         19          180,770   9,514.2   9,504.0     9,440     9,664         59.4  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ma…
 :dflash_d…  Push…  266  266         19        114          142,949   1,253.9   1,216.0     1,152     1,472         70.6  _table_qk_norm_rope_kernel                                                                          
 :dflash_d…  Push…  266  266         19        114          125,668   1,102.4   1,056.0     1,024     1,824        127.0  void sglang::store_kvcache<(long)256, (long)256, (int)1, (bool)1, long>(sglang::StoreKVCacheParams) 
 :dflash_d…  Push…  266  266         19         19           71,269   3,751.0   3,744.0     3,712     3,872         40.6  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ar…
 :dflash_d…  Push…  266  266         19         57           67,907   1,191.4   1,024.0       992     1,664        252.4  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 :dflash_d…  Push…  266  266         19         57           53,668     941.5     960.0       736     1,248        144.5  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  266  266         19         38           42,275   1,112.5   1,120.5     1,024     1,280         60.2  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_d…  Push…  266  266         19         19           40,898   2,152.5   2,112.0     2,049     2,464        106.0  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_d…  Push…  266  266         19         57           33,955     595.7     512.0       512       800        118.0  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 :dflash_d…  Push…  266  266         19         19           30,913   1,627.0   1,632.0     1,568     1,856         66.8  void at::native::_scatter_gather_elementwise_kernel<(int)128, (int)8, void at::native::_cuda_scatte…
 :dflash_d…  Push…  266  266         20         20           27,808   1,390.4   1,376.0     1,152     1,568         75.2  _prepare_dflash_draft_block_contig_kernel                                                           
 :dflash_d…  Push…  266  266         19         19           27,138   1,428.3   1,440.0     1,216     1,600         66.8  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  266  266         19         19           25,184   1,325.5   1,312.0     1,280     1,376         26.8  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
 :dflash_d…  Push…  266  266         19         19           19,521   1,027.4   1,024.0       992     1,216         49.9  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  266  266         20         20           19,425     971.3     928.0       896     1,408        148.0  _vocab_parallel_embedding_kernel                                                                    
 :dflash_d…  Push…  266  266         19         19           19,264   1,013.9   1,024.0       992     1,056         18.6  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  266  266         19         19           18,753     987.0     992.0       960     1,056         24.5  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctor_add<long>, std::arra…
 :dflash_d…  Push…  266  266         19         19           12,865     677.1     672.0       672       704         12.0  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
 :dflash_d…  Push…  267  267         19        247       81,429,687  329,674…  275,497…   264,617  6,266,6…    464,875.6  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_d…  Push…  267  267         19        114       34,119,355  299,292…  299,338…   295,786   303,338      1,556.0  _fwd_kernel                                                                                         
 :dflash_d…  Push…  267  267         19         38        2,720,542  71,593.2  72,434.0    55,330    96,548     13,789.9  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_d…  Push…  267  267         19        228        2,114,987   9,276.3  10,448.5     5,856    17,249      3,025.9  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_d…  Push…  267  267         19         19        1,196,299  62,963.1  62,914.0    62,338    64,002        513.4  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_d…  Push…  267  267         19        114          811,192   7,115.7   7,040.0     6,593     8,865        399.0  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_d…  Push…  267  267         19        247          537,204   2,174.9   2,208.0     1,568     2,304        173.0  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_d…  Push…  267  267         19        228          521,489   2,287.2   2,272.0     2,144     2,465         57.0  kernel_cutlass_kernel_flashinfernormkernelsfused_add_rmsnormFusedAddRMSNormKernel_object_at__tensor…
 :dflash_d…  Push…  267  267         19        114          361,357   3,169.8   3,104.0     3,008     3,616        155.8  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_d…  Push…  267  267         19        228          350,639   1,537.9   1,520.5     1,408     1,824        128.7  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_d…  Push…  267  267         19        114          339,021   2,973.9   2,976.0     2,784     3,072         41.9  void sglang::act_and_mul_kernel<__nv_bfloat16, (sglang::ActivationKind)0, (bool)1, (bool)0, (bool)0…
 :dflash_d…  Push…  267  267         19         38          338,663   8,912.2   5,424.0     2,880    15,873      6,116.1  create_flashinfer_kv_indices_triton                                                                 
 :dflash_d…  Push…  267  267         19         19          179,782   9,462.2   9,472.0     9,312     9,664         66.5  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ma…
 :dflash_d…  Push…  267  267         19        114          141,283   1,239.3   1,216.0     1,152     1,504         83.1  _table_qk_norm_rope_kernel                                                                          
 :dflash_d…  Push…  267  267         19        114          126,340   1,108.2   1,057.0     1,024     1,856        134.1  void sglang::store_kvcache<(long)256, (long)256, (int)1, (bool)1, long>(sglang::StoreKVCacheParams) 
 :dflash_d…  Push…  267  267         19         19           70,498   3,710.4   3,712.0     3,680     3,808         27.1  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ar…
 :dflash_d…  Push…  267  267         19         57           67,779   1,189.1   1,056.0       960     1,633        255.8  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 :dflash_d…  Push…  267  267         19         57           53,410     937.0     960.0       736     1,217        148.4  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  267  267         19         38           43,139   1,135.2   1,104.0     1,024     1,408        102.5  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_d…  Push…  267  267         19         19           39,650   2,086.8   2,080.0     2,048     2,240         47.1  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_d…  Push…  267  267         19         57           33,761     592.3     512.0       480       768        112.0  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 :dflash_d…  Push…  267  267         19         19           30,305   1,595.0   1,568.0     1,536     1,697         41.8  void at::native::_scatter_gather_elementwise_kernel<(int)128, (int)8, void at::native::_cuda_scatte…
 :dflash_d…  Push…  267  267         20         20           27,682   1,384.1   1,376.5     1,185     1,504         56.6  _prepare_dflash_draft_block_contig_kernel                                                           
 :dflash_d…  Push…  267  267         19         19           25,986   1,367.7   1,344.0     1,216     1,569         73.1  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  267  267         19         19           25,152   1,323.8   1,312.0     1,248     1,440         46.8  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
 :dflash_d…  Push…  267  267         19         19           19,712   1,037.5   1,024.0       992     1,120         30.8  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  267  267         19         19           19,520   1,027.4   1,024.0       992     1,152         36.8  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  267  267         19         19           18,753     987.0     992.0       960     1,120         37.4  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctor_add<long>, std::arra…
 :dflash_d…  Push…  267  267         20         20           18,336     916.8     896.0       864     1,344        102.4  _vocab_parallel_embedding_kernel                                                                    
 :dflash_d…  Push…  267  267         19         19           13,249     697.3     673.0       672       768         34.7  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
 :dflash_d…  Push…  268  268         19        247       84,181,786  340,816…  273,672…   224,839  9,994,3…    672,706.8  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_d…  Push…  268  268         19        114       34,381,182  301,589…  301,497…   298,249   305,961      1,436.0  _fwd_kernel                                                                                         
 :dflash_d…  Push…  268  268         19         38        2,733,676  71,938.8  72,466.0    60,993    94,755      8,566.7  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_d…  Push…  268  268         19        228        2,117,720   9,288.2  10,592.5     5,856    16,705      3,051.6  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_d…  Push…  268  268         19         19        1,203,365  63,335.0  62,498.0    61,858    79,522      3,941.7  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_d…  Push…  268  268         19        114          814,526   7,145.0   7,104.0     6,593     9,505        414.2  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_d…  Push…  268  268         19        247          538,032   2,178.3   2,208.0     1,600     2,336        165.8  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_d…  Push…  268  268         19        228          524,687   2,301.3   2,304.0     2,208     2,496         44.8  kernel_cutlass_kernel_flashinfernormkernelsfused_add_rmsnormFusedAddRMSNormKernel_object_at__tensor…
 :dflash_d…  Push…  268  268         19        114          368,422   3,231.8   3,200.0     3,008     4,512        215.3  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_d…  Push…  268  268         19        228          356,137   1,562.0   1,552.0     1,408     1,824        131.2  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_d…  Push…  268  268         19         38          340,619   8,963.7   5,472.0     2,912    15,808      6,123.3  create_flashinfer_kv_indices_triton                                                                 
 :dflash_d…  Push…  268  268         19        114          340,425   2,986.2   2,977.0     2,816     3,040         34.1  void sglang::act_and_mul_kernel<__nv_bfloat16, (sglang::ActivationKind)0, (bool)1, (bool)0, (bool)0…
 :dflash_d…  Push…  268  268         19         19          181,923   9,574.9   9,600.0     9,344     9,696         87.1  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ma…
 :dflash_d…  Push…  268  268         19        114          139,874   1,227.0   1,184.5     1,152     1,600         82.9  _table_qk_norm_rope_kernel                                                                          
 :dflash_d…  Push…  268  268         19        114          126,370   1,108.5   1,088.0     1,024     1,856        141.8  void sglang::store_kvcache<(long)256, (long)256, (int)1, (bool)1, long>(sglang::StoreKVCacheParams) 
 :dflash_d…  Push…  268  268         19         19           70,594   3,715.5   3,712.0     3,680     3,776         25.8  void at::native::reduce_kernel<(int)512, (int)1, at::native::ReduceOp<c10::BFloat16, at::native::Ar…
 :dflash_d…  Push…  268  268         19         57           69,923   1,226.7   1,056.0       992     1,696        266.1  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 :dflash_d…  Push…  268  268         19         57           54,978     964.5     992.0       768     1,152        145.5  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  268  268         19         38           44,068   1,159.7   1,152.5     1,024     1,472        109.4  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_d…  Push…  268  268         19         19           39,235   2,065.0   2,048.0     2,016     2,208         49.4  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_d…  Push…  268  268         19         57           34,498     605.2     544.0       512       992        119.5  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 :dflash_d…  Push…  268  268         19         19           30,977   1,630.4   1,600.0     1,536     1,792         75.8  void at::native::_scatter_gather_elementwise_kernel<(int)128, (int)8, void at::native::_cuda_scatte…
 :dflash_d…  Push…  268  268         20         20           29,152   1,457.6   1,472.0     1,184     1,600        103.0  _prepare_dflash_draft_block_contig_kernel                                                           
 :dflash_d…  Push…  268  268         19         19           27,169   1,429.9   1,440.0     1,248     1,504         55.5  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_d…  Push…  268  268         19         19           24,001   1,263.2   1,280.0     1,216     1,312         28.9  void at::native::<unnamed>::multi_tensor_apply_kernel<at::native::<unnamed>::TensorListMetadata<(in…
 :dflash_d…  Push…  268  268         19         19           20,128   1,059.4   1,056.0     1,024     1,184         38.3  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  268  268         19         19           19,970   1,051.1   1,056.0       992     1,184         44.5  void at::native::vectorized_elementwise_kernel<(int)2, at::native::AUnaryFunctor<long, long, long, …
 :dflash_d…  Push…  268  268         19         19           19,712   1,037.5   1,024.0       992     1,088         22.2  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctor_add<long>, std::arra…
 :dflash_d…  Push…  268  268         20         20           17,921     896.0     896.0       864       960         27.4  _vocab_parallel_embedding_kernel                                                                    
 :dflash_d…  Push…  268  268         19         19           13,408     705.7     704.0       672       800         31.1  void <unnamed>::elementwise_kernel_with_index<int, at::native::arange_cuda_out(const c10::Scalar &,…
 :dflash_m…  Push…  261  261         20         20       10,291,457  514,572…  513,632…   512,752   531,088      3,954.8  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_128x1_tn_align8>(T1::Par…
 :dflash_m…  Push…  261  261         20         20          394,699  19,735.0  19,760.5    18,944    20,385        402.7  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_nn_align8>(T1::Para…
 :dflash_m…  Push…  261  261         20        120          122,851   1,023.8     944.0       832     1,569        218.8  set_kv_buffer_prefix_valid_tiled                                                                    
 :dflash_m…  Push…  261  261         20         40          117,765   2,944.1   2,864.0     1,824     4,416      1,111.6  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_m…  Push…  261  261         20         20          101,124   5,056.2   4,880.0     4,672     5,600        349.2  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_m…  Push…  261  261         20         20           33,635   1,681.8   1,664.0     1,601     1,793         56.3  _fused_norm_rope_kernel_stacked                                                                     
 :dflash_m…  Push…  262  262         19         19        9,778,859  514,676…  514,542…   513,582   516,558        854.4  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_128x1_tn_align8>(T1::Par…
 :dflash_m…  Push…  262  262         19         19          375,917  19,785.1  19,713.0    19,233    20,353        345.9  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_nn_align8>(T1::Para…
 :dflash_m…  Push…  262  262         19        114          117,284   1,028.8     944.0       864     1,632        233.4  set_kv_buffer_prefix_valid_tiled                                                                    
 :dflash_m…  Push…  262  262         19         38          110,788   2,915.5   2,784.5     1,792     4,320      1,096.1  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_m…  Push…  262  262         19         19           89,088   4,688.8   4,576.0     4,416     5,792        331.6  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_m…  Push…  262  262         19         19           32,256   1,697.7   1,696.0     1,600     2,048         98.6  _fused_norm_rope_kernel_stacked                                                                     
 :dflash_m…  Push…  263  263         19         19        9,782,667  514,877…  514,640…   513,328   516,656        908.2  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_128x1_tn_align8>(T1::Par…
 :dflash_m…  Push…  263  263         19         19          374,189  19,694.2  19,712.0    19,297    20,257        308.1  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_nn_align8>(T1::Para…
 :dflash_m…  Push…  263  263         19        114          118,916   1,043.1     944.0       864     1,664        231.2  set_kv_buffer_prefix_valid_tiled                                                                    
 :dflash_m…  Push…  263  263         19         38          109,987   2,894.4   2,832.0     1,824     4,384      1,063.5  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_m…  Push…  263  263         19         19           90,113   4,742.8   4,608.0     4,448     5,792        374.0  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_m…  Push…  263  263         19         19           33,313   1,753.3   1,728.0     1,632     1,953         95.9  _fused_norm_rope_kernel_stacked                                                                     
 :dflash_m…  Push…  264  264         19         19        9,769,961  514,208…  514,062…   512,653   516,462      1,055.2  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_128x1_tn_align8>(T1::Par…
 :dflash_m…  Push…  264  264         19         19          373,515  19,658.7  19,648.0    19,297    20,512        258.8  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_nn_align8>(T1::Para…
 :dflash_m…  Push…  264  264         19        114          119,586   1,049.0     960.0       864     1,632        221.4  set_kv_buffer_prefix_valid_tiled                                                                    
 :dflash_m…  Push…  264  264         19         38          111,267   2,928.1   2,848.5     1,824     4,320      1,099.9  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_m…  Push…  264  264         19         19           91,490   4,815.3   4,736.0     4,512     5,664        314.1  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_m…  Push…  264  264         19         19           32,226   1,696.1   1,696.0     1,632     1,792         44.0  _fused_norm_rope_kernel_stacked                                                                     
 :dflash_m…  Push…  265  265         19         19        9,762,651  513,823…  514,031…   512,207   516,239      1,089.1  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_128x1_tn_align8>(T1::Par…
 :dflash_m…  Push…  265  265         19         19          377,355  19,860.8  19,840.0    19,201    20,289        298.4  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_nn_align8>(T1::Para…
 :dflash_m…  Push…  265  265         19        114          118,626   1,040.6     960.0       864     1,600        223.4  set_kv_buffer_prefix_valid_tiled                                                                    
 :dflash_m…  Push…  265  265         19         38          113,412   2,984.5   2,864.5     1,824     4,353      1,138.9  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_m…  Push…  265  265         19         19           95,648   5,034.1   4,832.0     4,544     5,952        449.4  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_m…  Push…  265  265         19         19           32,481   1,709.5   1,696.0     1,664     1,856         55.8  _fused_norm_rope_kernel_stacked                                                                     
 :dflash_m…  Push…  266  266         19         19        9,785,343  515,018…  514,992…   514,033   516,401        686.9  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_128x1_tn_align8>(T1::Par…
 :dflash_m…  Push…  266  266         19         19          374,956  19,734.5  19,744.0    19,137    20,064        226.3  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_nn_align8>(T1::Para…
 :dflash_m…  Push…  266  266         19        114          118,821   1,042.3     960.5       896     1,632        214.1  set_kv_buffer_prefix_valid_tiled                                                                    
 :dflash_m…  Push…  266  266         19         38          111,233   2,927.2   2,848.0     1,856     4,352      1,077.4  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_m…  Push…  266  266         19         19           92,674   4,877.6   4,704.0     4,448     5,665        405.9  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_m…  Push…  266  266         19         19           32,417   1,706.2   1,696.0     1,600     1,888         73.2  _fused_norm_rope_kernel_stacked                                                                     
 :dflash_m…  Push…  267  267         19         19        9,771,176  514,272…  514,545…   513,553   515,058        574.0  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_128x1_tn_align8>(T1::Par…
 :dflash_m…  Push…  267  267         19         19          375,661  19,771.6  19,745.0    19,329    20,417        252.2  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_nn_align8>(T1::Para…
 :dflash_m…  Push…  267  267         19        114          117,027   1,026.6     944.0       864     1,568        225.9  set_kv_buffer_prefix_valid_tiled                                                                    
 :dflash_m…  Push…  267  267         19         38          109,571   2,883.4   2,800.0     1,824     4,448      1,067.4  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_m…  Push…  267  267         19         19           92,196   4,852.4   4,608.0     4,352     5,632        477.4  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_m…  Push…  267  267         19         19           32,256   1,697.7   1,664.0     1,632     1,792         47.1  _fused_norm_rope_kernel_stacked                                                                     
 :dflash_m…  Push…  268  268         19         19        9,778,586  514,662…  514,351…   513,422   517,007      1,007.8  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_128x1_tn_align8>(T1::Par…
 :dflash_m…  Push…  268  268         19         19          375,787  19,778.3  19,776.0    19,296    21,505        492.4  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_nn_align8>(T1::Para…
 :dflash_m…  Push…  268  268         19        114          116,804   1,024.6     928.0       832     1,600        231.3  set_kv_buffer_prefix_valid_tiled                                                                    
 :dflash_m…  Push…  268  268         19         38          111,041   2,922.1   2,864.0     1,792     4,352      1,117.3  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_m…  Push…  268  268         19         19           94,021   4,948.5   4,704.0     4,545     5,761        423.7  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_m…  Push…  268  268         19         19           33,248   1,749.9   1,696.0     1,632     1,920         88.0  _fused_norm_rope_kernel_stacked                                                                     
 :dflash_r…  Push…  261  261         20         20          280,391  14,019.5  13,856.0    13,441    17,088        753.1  _fused_mamba_state_scatter_with_mask_kernel                                                         
 :dflash_r…  Push…  261  261         20         20           80,546   4,027.3   3,936.0     3,680     5,089        331.4  _fused_conv_window_scatter_multi_kernel                                                             
 :dflash_r…  Push…  261  261         20         20           27,905   1,395.3   1,376.0     1,344     1,569         66.9  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  261  261         20         20           22,784   1,139.2   1,120.0     1,088     1,344         66.0  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  261  261         20         20           21,696   1,084.8   1,056.0     1,056     1,216         43.9  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_r…  Push…  262  262         19         19          270,248  14,223.6  14,208.0    13,536    14,784        339.4  _fused_mamba_state_scatter_with_mask_kernel                                                         
 :dflash_r…  Push…  262  262         19         19           73,189   3,852.1   3,808.0     3,585     4,256        192.9  _fused_conv_window_scatter_multi_kernel                                                             
 :dflash_r…  Push…  262  262         19         19           27,072   1,424.8   1,408.0     1,408     1,504         30.9  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  262  262         19         19           20,929   1,101.5   1,088.0     1,056     1,184         26.8  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  262  262         19         19           20,288   1,067.8   1,056.0     1,024     1,120         24.3  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_r…  Push…  263  263         19         19          266,729  14,038.4  14,112.0    13,217    14,496        356.7  _fused_mamba_state_scatter_with_mask_kernel                                                         
 :dflash_r…  Push…  263  263         19         19           73,475   3,867.1   3,808.0     3,520     4,480        263.2  _fused_conv_window_scatter_multi_kernel                                                             
 :dflash_r…  Push…  263  263         19         19           27,232   1,433.3   1,440.0     1,408     1,536         31.2  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  263  263         19         19           21,056   1,108.2   1,088.0     1,056     1,248         41.6  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  263  263         19         19           20,224   1,064.4   1,056.0     1,024     1,184         36.7  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_r…  Push…  264  264         19         19          261,702  13,773.8  13,792.0    13,473    14,240        226.9  _fused_mamba_state_scatter_with_mask_kernel                                                         
 :dflash_r…  Push…  264  264         19         19           75,620   3,980.0   4,000.0     3,616     4,352        224.6  _fused_conv_window_scatter_multi_kernel                                                             
 :dflash_r…  Push…  264  264         19         19           26,593   1,399.6   1,376.0     1,344     1,536         48.7  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  264  264         19         19           21,313   1,121.7   1,120.0     1,088     1,153         22.6  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  264  264         19         19           20,385   1,072.9   1,056.0     1,056     1,184         30.8  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_r…  Push…  265  265         19         19          273,546  14,397.2  14,337.0    13,985    14,912        226.5  _fused_mamba_state_scatter_with_mask_kernel                                                         
 :dflash_r…  Push…  265  265         19         19           74,497   3,920.9   3,840.0     3,648     4,640        262.0  _fused_conv_window_scatter_multi_kernel                                                             
 :dflash_r…  Push…  265  265         19         19           27,041   1,423.2   1,408.0     1,344     1,536         43.2  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  265  265         19         19           21,281   1,120.1   1,120.0     1,056     1,216         44.0  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  265  265         19         19           20,609   1,084.7   1,056.0     1,025     1,216         49.8  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_r…  Push…  266  266         19         19          271,561  14,292.7  14,337.0    13,697    15,201        375.3  _fused_mamba_state_scatter_with_mask_kernel                                                         
 :dflash_r…  Push…  266  266         19         19           76,739   4,038.9   4,064.0     3,712     4,385        222.2  _fused_conv_window_scatter_multi_kernel                                                             
 :dflash_r…  Push…  266  266         19         19           26,784   1,409.7   1,408.0     1,376     1,472         25.0  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  266  266         19         19           21,185   1,115.0   1,120.0     1,056     1,184         35.8  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  266  266         19         19           20,640   1,086.3   1,056.0     1,024     1,248         63.5  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_r…  Push…  267  267         19         19          265,960  13,997.9  14,144.0    13,281    14,560        398.9  _fused_mamba_state_scatter_with_mask_kernel                                                         
 :dflash_r…  Push…  267  267         19         19           75,073   3,951.2   3,936.0     3,616     4,352        227.1  _fused_conv_window_scatter_multi_kernel                                                             
 :dflash_r…  Push…  267  267         19         19           27,424   1,443.4   1,408.0     1,376     1,664         79.0  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  267  267         19         19           21,250   1,118.4   1,088.0     1,056     1,280         64.4  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  267  267         19         19           19,905   1,047.6   1,024.0     1,024     1,152         41.1  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_r…  Push…  268  268         19         19          274,248  14,434.1  14,400.0    14,016    14,817        231.0  _fused_mamba_state_scatter_with_mask_kernel                                                         
 :dflash_r…  Push…  268  268         19         19           74,306   3,910.8   3,936.0     3,616     4,320        217.6  _fused_conv_window_scatter_multi_kernel                                                             
 :dflash_r…  Push…  268  268         19         19           28,001   1,473.7   1,472.0     1,440     1,568         29.2  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  268  268         19         19           21,121   1,111.6   1,088.0     1,088     1,216         39.7  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_r…  Push…  268  268         19         19           19,681   1,035.8   1,024.0     1,024     1,056         15.8  void at::native::vectorized_elementwise_kernel<(int)2, at::native::CUDAFunctorOnSelf_add<long>, std…
 :dflash_t…  Push…  261  261         20        480    1,877,141,617  3,910,7…  3,910,0…  3,876,2…  3,942,7…      9,512.4  _fwd_kernel                                                                                         
 :dflash_t…  Push…  261  261         20      3,740    1,121,980,937  299,994…  291,977…   215,750  1,913,4…     56,045.1  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_t…  Push…  261  261         20      3,680      130,290,708  35,405.1  32,705.0    13,281    70,562     11,552.1  void cutlass::device_kernel<cutlass::gemm::kernel::GemmUniversal<cutlass::gemm::GroupProblemShape<c…
 :dflash_t…  Push…  261  261         20      1,840      110,696,805  60,161.3  60,098.0    57,634    70,243        879.6  void cutlass::Kernel2<cutlass_80_tensorop_s16816gemm_bf16_256x64_32x4_tn_align8>(T1::Params)        
 :dflash_t…  Push…  261  261         20      4,200       99,481,220  23,686.0   6,433.0     4,032    50,530     20,181.1  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_t…  Push…  261  261         20         20       67,938,338  3,396,9…  1,641,7…  1,376,3…  9,960,0…  3,012,197.1  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_t…  Push…  261  261         20      1,380       55,544,329  40,249.5  40,321.0    39,201    41,025        309.3  fused_sigmoid_gating_delta_rule_update_kernel                                                       
 :dflash_t…  Push…  261  261         20      3,240       47,146,943  14,551.5  18,369.0     5,281    23,393      7,315.5  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_t…  Push…  261  261         20      3,840       24,533,309   6,388.9   6,272.0     4,448     7,937        635.1  _score_kernel                                                                                       
 :dflash_t…  Push…  261  261         20      1,840       15,361,090   8,348.4   8,288.0     5,600    13,344      1,023.0  void tensorrt_llm::kernels::cutlass_kernels::finalizeMoeRoutingKernel<__nv_bfloat16, __nv_bfloat16,…
 :dflash_t…  Push…  261  261         20        480       13,164,976  27,427.0  27,409.0    25,889    29,377        586.2  kernel_cutlass__dsv3_fused_a_gemm_kernel_tensorptri32gmemo2112358435841_tensorptri32gmemo1635843584…
 :dflash_t…  Push…  261  261         20      2,320        9,049,727   3,900.7   4,160.0     2,432     4,832        649.9  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_t…  Push…  261  261         20      3,260        8,615,555   2,642.8   2,721.0     1,760     5,345        514.9  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_t…  Push…  261  261         20      3,840        7,985,648   2,079.6   2,112.0       928     3,136        515.4  _combine_kernel                                                                                     
 :dflash_t…  Push…  261  261         20      3,740        7,797,106   2,084.8   2,080.0     1,920     2,432         68.6  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  261  261         20      1,380        7,358,941   5,332.6   5,377.0     4,864     6,560        226.9  _causal_conv1d_update_kernel                                                                        
 :dflash_t…  Push…  261  261         20      1,840        7,062,418   3,838.3   3,839.0     3,648     4,160         72.9  void tensorrt_llm::kernels::cutlass_kernels::doActivationKernel<__nv_fp8_e4m3, __nv_bfloat16, __nv_…
 :dflash_t…  Push…  261  261         20      1,840        6,453,367   3,507.3   3,488.0     3,328     3,872         84.5  void tensorrt_llm::kernels::cutlass_kernels::expandInputRowsKernel<__nv_fp8_e4m3, __nv_fp8_e4m3, (t…
 :dflash_t…  Push…  261  261         20      1,840        4,924,853   2,676.6   2,656.0     2,368     3,233        108.2  void sglang::route_quant_fused_kernel<(bool)1, float, __nv_bfloat16>(sglang::RouteQuantFusedParams) 
 :dflash_t…  Push…  261  261         20      1,840        4,383,842   2,382.5   2,368.0     2,272     2,784         52.9  void tensorrt_llm::kernels::cutlass_kernels::blockExpertPrefixSumKernel<(int)256>(const int *, int …
 :dflash_t…  Push…  261  261         20      1,840        4,235,322   2,301.8   2,304.0     2,208     2,688         55.3  void tensorrt_llm::kernels::cutlass_kernels::computeStridesTmaWarpSpecializedKernel<__nv_fp8_e4m3, …
 :dflash_t…  Push…  261  261         20      1,840        3,825,936   2,079.3   2,080.0     1,920     2,816         66.9  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  261  261         20      1,380        3,660,171   2,652.3   2,624.0     2,368     3,136         88.2  void gemmSN_TN_kernel<float, (int)128, (int)16, (int)2, (int)4, (int)8, (int)9, (bool)0, cublasGemv…
 :dflash_t…  Push…  261  261         20      1,840        3,508,265   1,906.7   1,888.0     1,824     2,208         38.8  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  261  261         20      1,840        3,478,154   1,890.3   1,888.0     1,792     2,143         39.0  void tensorrt_llm::kernels::quantize_with_block_size<(tensorrt_llm::BlockScaleQuantizationType)2, _…
 :dflash_t…  Push…  261  261         20      1,840        3,360,673   1,826.5   1,824.0     1,760     2,048         35.8  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, float, float, float, (bool)0, floa…
 :dflash_t…  Push…  261  261         20      1,840        3,280,586   1,782.9   1,761.0     1,696     2,048         47.5  void sglang::situ_and_mul_kernel<float, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMulParams)  
 :dflash_t…  Push…  261  261         20      1,380        3,012,092   2,182.7   2,176.0     2,080     2,496         56.4  layer_norm_gated_fwd_kernel                                                                         
 :dflash_t…  Push…  261  261         20      1,560        3,001,427   1,924.0   1,888.0     1,664     5,600        416.9  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  261  261         20        480        2,942,905   6,131.1   6,113.0     5,920     6,528        102.0  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
 :dflash_t…  Push…  261  261         20      1,840        2,673,305   1,452.9   1,440.0     1,248     2,049        117.1  void sglang::add3_kernel<(bool)1, (bool)1>(sglang::Add3Params)                                      
 :dflash_t…  Push…  261  261         20      1,840        2,411,066   1,310.4   1,280.0     1,216     1,728         69.6  tensorrt_llm::kernels::cutlass_kernels::mergeExpertPrefixSumKernel(const int *, const int *, const …
 :dflash_t…  Push…  261  261         20        500        2,256,293   4,512.6   3,488.0     3,168    26,945      4,491.4  void cutlass::Kernel2<cutlass_80_wmma_tensorop_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Params)   
 :dflash_t…  Push…  261  261         20      1,840        2,130,023   1,157.6   1,152.0     1,088     1,568         41.7  void tensorrt_llm::kernels::cutlass_kernels::globalExpertPrefixSumKernel<(int)256>(const int *, int…
 :dflash_t…  Push…  261  261         20      1,720        1,673,775     973.1     960.0       800     1,184         36.8  void at::native::vectorized_elementwise_kernel<(int)4, at::native::CUDAFunctor_add<c10::BFloat16>, …
 :dflash_t…  Push…  261  261         20        960        1,588,015   1,654.2   1,648.5     1,504     1,952         98.3  void at::native::<unnamed>::CatArrayBatchedCopy<at::native::<unnamed>::OpaqueType<(unsigned int)2>,…
 :dflash_t…  Push…  261  261         20        480        1,467,601   3,057.5   3,040.0     2,400     4,160        299.9  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  261  261         20         20        1,294,634  64,731.7  64,850.0    63,714    65,890        574.4  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_t…  Push…  261  261         20        500          996,126   1,992.3   1,952.0     1,888     2,560        115.8  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, __nv_bfloat16, float, __nv_bfloat1…
 :dflash_t…  Push…  261  261         20        480          689,678   1,436.8   1,440.0     1,376     1,696         37.3  void sglang::mla_output_gate_kernel<(int)256, (bool)1>(sglang::MlaOutputGateParams)                 
 :dflash_t…  Push…  261  261         20        480          627,415   1,307.1   1,248.0     1,120     3,360        262.0  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  261  261         20         20          138,632   6,931.6   6,944.0     6,848     7,008         48.0  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_t…  Push…  261  261         20         20           70,725   3,536.3   3,521.0     3,488     3,584         35.2  void sglang::situ_and_mul_kernel<__nv_bfloat16, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMul…
 :dflash_t…  Push…  261  261         20         20           48,225   2,411.3   2,368.0     2,272     2,720        110.5  void at::native::<unnamed>::CatArrayBatchedCopy_vectorized<at::native::<unnamed>::OpaqueType<(unsig…
 :dflash_t…  Push…  261  261         20         20           27,168   1,358.4   1,360.0     1,248     1,504         62.7  _vocab_parallel_embedding_kernel                                                                    
 :dflash_t…  Push…  261  261         20         20           18,048     902.4     896.0       864     1,120         67.0  void at::native::vectorized_elementwise_kernel<(int)4, at::native::AUnaryFunctor<c10::BFloat16, c10…
 :dflash_t…  Push…  261  261         20         20           15,712     785.6     800.0       736       832         22.0  void at::native::vectorized_elementwise_kernel<(int)4, at::native::FillFunctor<float>, std::array<c…
 :dflash_t…  Push…  262  262         19        456    1,761,832,629  3,863,6…  3,862,3…  3,844,3…  3,893,2…      8,710.4  _fwd_kernel                                                                                         
 :dflash_t…  Push…  262  262         19      3,553    1,085,093,227  305,402…  292,360…   219,398  1,901,7…     59,502.1  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_t…  Push…  262  262         19      3,496      123,387,268  35,293.8  32,433.0    13,504    71,361     11,598.7  void cutlass::device_kernel<cutlass::gemm::kernel::GemmUniversal<cutlass::gemm::GroupProblemShape<c…
 :dflash_t…  Push…  262  262         19      1,748      105,055,331  60,100.3  60,033.0    57,889    70,146        989.6  void cutlass::Kernel2<cutlass_80_tensorop_s16816gemm_bf16_256x64_32x4_tn_align8>(T1::Params)        
 :dflash_t…  Push…  262  262         19      3,990       94,454,976  23,672.9   6,400.0     4,000    50,401     20,211.9  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_t…  Push…  262  262         19         19       62,325,500  3,280,2…  1,518,2…  1,375,4…  9,964,2…  3,055,437.0  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_t…  Push…  262  262         19      1,311       51,982,981  39,651.4  39,585.0    38,850    40,673        377.3  fused_sigmoid_gating_delta_rule_update_kernel                                                       
 :dflash_t…  Push…  262  262         19      3,078       44,940,929  14,600.7  18,816.5     5,184    23,873      7,352.0  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_t…  Push…  262  262         19      3,648       23,057,274   6,320.5   6,209.0     4,416     7,872        616.3  _score_kernel                                                                                       
 :dflash_t…  Push…  262  262         19      1,748       14,348,428   8,208.5   8,160.0     5,409    13,056      1,015.9  void tensorrt_llm::kernels::cutlass_kernels::finalizeMoeRoutingKernel<__nv_bfloat16, __nv_bfloat16,…
 :dflash_t…  Push…  262  262         19        456       12,528,181  27,474.1  27,489.0    26,112    28,929        536.7  kernel_cutlass__dsv3_fused_a_gemm_kernel_tensorptri32gmemo2112358435841_tensorptri32gmemo1635843584…
 :dflash_t…  Push…  262  262         19      2,204        8,513,102   3,862.6   4,128.0     2,368     4,673        653.2  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_t…  Push…  262  262         19      3,097        8,090,013   2,612.2   2,720.0     1,760     4,896        506.7  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_t…  Push…  262  262         19      3,648        7,529,017   2,063.9   2,080.0       896     3,328        516.3  _combine_kernel                                                                                     
 :dflash_t…  Push…  262  262         19      3,553        7,304,415   2,055.8   2,048.0     1,888     2,432         67.7  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  262  262         19      1,311        6,949,929   5,301.2   5,344.0     4,704     6,112        208.9  _causal_conv1d_update_kernel                                                                        
 :dflash_t…  Push…  262  262         19      1,748        6,704,508   3,835.5   3,839.0     3,616     4,224         82.1  void tensorrt_llm::kernels::cutlass_kernels::doActivationKernel<__nv_fp8_e4m3, __nv_bfloat16, __nv_…
 :dflash_t…  Push…  262  262         19      1,748        6,094,076   3,486.3   3,457.0     3,297     3,872         75.8  void tensorrt_llm::kernels::cutlass_kernels::expandInputRowsKernel<__nv_fp8_e4m3, __nv_fp8_e4m3, (t…
 :dflash_t…  Push…  262  262         19      1,748        4,609,515   2,637.0   2,624.0     2,304     3,040        107.2  void sglang::route_quant_fused_kernel<(bool)1, float, __nv_bfloat16>(sglang::RouteQuantFusedParams) 
 :dflash_t…  Push…  262  262         19      1,748        4,079,623   2,333.9   2,336.0     2,208     2,656         47.4  void tensorrt_llm::kernels::cutlass_kernels::blockExpertPrefixSumKernel<(int)256>(const int *, int …
 :dflash_t…  Push…  262  262         19      1,748        3,988,728   2,281.9   2,272.0     2,208     2,656         56.5  void tensorrt_llm::kernels::cutlass_kernels::computeStridesTmaWarpSpecializedKernel<__nv_fp8_e4m3, …
 :dflash_t…  Push…  262  262         19      1,748        3,593,559   2,055.8   2,048.0     1,920     2,400         66.1  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  262  262         19      1,311        3,441,056   2,624.8   2,592.0     2,336     3,169         89.8  void gemmSN_TN_kernel<float, (int)128, (int)16, (int)2, (int)4, (int)8, (int)9, (bool)0, cublasGemv…
 :dflash_t…  Push…  262  262         19      1,748        3,302,263   1,889.2   1,888.0     1,792     2,144         38.8  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  262  262         19      1,748        3,260,538   1,865.3   1,856.0     1,791     2,176         43.2  void tensorrt_llm::kernels::quantize_with_block_size<(tensorrt_llm::BlockScaleQuantizationType)2, _…
 :dflash_t…  Push…  262  262         19      1,748        3,165,621   1,811.0   1,792.0     1,728     2,048         34.7  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, float, float, float, (bool)0, floa…
 :dflash_t…  Push…  262  262         19      1,748        3,068,546   1,755.5   1,760.0     1,664     2,048         50.4  void sglang::situ_and_mul_kernel<float, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMulParams)  
 :dflash_t…  Push…  262  262         19      1,311        2,835,913   2,163.2   2,144.0     2,080     2,497         61.9  layer_norm_gated_fwd_kernel                                                                         
 :dflash_t…  Push…  262  262         19      1,482        2,816,436   1,900.4   1,856.0     1,632     5,536        413.2  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  262  262         19        456        2,769,108   6,072.6   6,048.0     5,824     6,433        112.4  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
 :dflash_t…  Push…  262  262         19      1,748        2,489,549   1,424.2   1,408.0     1,280     2,017        104.9  void sglang::add3_kernel<(bool)1, (bool)1>(sglang::Add3Params)                                      
 :dflash_t…  Push…  262  262         19      1,748        2,258,841   1,292.2   1,280.0     1,184     1,631         64.8  tensorrt_llm::kernels::cutlass_kernels::mergeExpertPrefixSumKernel(const int *, const int *, const …
 :dflash_t…  Push…  262  262         19        475        2,102,356   4,426.0   3,392.0     3,104    26,880      4,547.9  void cutlass::Kernel2<cutlass_80_wmma_tensorop_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Params)   
 :dflash_t…  Push…  262  262         19      1,748        1,998,709   1,143.4   1,120.0     1,088     1,408         44.1  void tensorrt_llm::kernels::cutlass_kernels::globalExpertPrefixSumKernel<(int)256>(const int *, int…
 :dflash_t…  Push…  262  262         19      1,634        1,575,559     964.2     960.0       768     1,184         35.8  void at::native::vectorized_elementwise_kernel<(int)4, at::native::CUDAFunctor_add<c10::BFloat16>, …
 :dflash_t…  Push…  262  262         19        912        1,492,717   1,636.8   1,632.0     1,472     1,984        101.5  void at::native::<unnamed>::CatArrayBatchedCopy<at::native::<unnamed>::OpaqueType<(unsigned int)2>,…
 :dflash_t…  Push…  262  262         19        456        1,351,408   2,963.6   2,944.0     2,048     4,129        294.4  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  262  262         19         19        1,233,669  64,929.9  65,058.0    64,065    66,434        615.9  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_t…  Push…  262  262         19        475          948,284   1,996.4   1,952.0     1,888     2,496        114.3  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, __nv_bfloat16, float, __nv_bfloat1…
 :dflash_t…  Push…  262  262         19        456          633,973   1,390.3   1,376.0     1,344     1,568         28.4  void sglang::mla_output_gate_kernel<(int)256, (bool)1>(sglang::MlaOutputGateParams)                 
 :dflash_t…  Push…  262  262         19        456          586,475   1,286.1   1,216.0     1,120     3,584        281.5  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  262  262         19         19          129,729   6,827.8   6,848.0     6,688     6,944         67.7  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_t…  Push…  262  262         19         19           65,987   3,473.0   3,488.0     3,360     3,521         37.5  void sglang::situ_and_mul_kernel<__nv_bfloat16, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMul…
 :dflash_t…  Push…  262  262         19         19           48,704   2,563.4   2,560.0     2,304     2,880        145.0  void at::native::<unnamed>::CatArrayBatchedCopy_vectorized<at::native::<unnamed>::OpaqueType<(unsig…
 :dflash_t…  Push…  262  262         19         19           25,152   1,323.8   1,312.0       896     1,504        124.0  _vocab_parallel_embedding_kernel                                                                    
 :dflash_t…  Push…  262  262         19         19           17,057     897.7     896.0       864       928         12.9  void at::native::vectorized_elementwise_kernel<(int)4, at::native::AUnaryFunctor<c10::BFloat16, c10…
 :dflash_t…  Push…  262  262         19         19           14,882     783.3     768.0       736       992         60.7  void at::native::vectorized_elementwise_kernel<(int)4, at::native::FillFunctor<float>, std::array<c…
 :dflash_t…  Push…  263  263         19        456    1,781,113,597  3,905,9…  3,904,8…  3,884,1…  3,934,6…     10,094.5  _fwd_kernel                                                                                         
 :dflash_t…  Push…  263  263         19      3,553    1,063,430,509  299,305…  291,785…   219,175  1,886,6…     53,338.7  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_t…  Push…  263  263         19      3,496      123,663,167  35,372.8  32,641.0    13,345    71,650     11,595.9  void cutlass::device_kernel<cutlass::gemm::kernel::GemmUniversal<cutlass::gemm::GroupProblemShape<c…
 :dflash_t…  Push…  263  263         19      1,748      105,312,942  60,247.7  60,162.0    57,698    72,610        970.3  void cutlass::Kernel2<cutlass_80_tensorop_s16816gemm_bf16_256x64_32x4_tn_align8>(T1::Params)        
 :dflash_t…  Push…  263  263         19      3,990       94,522,438  23,689.8   6,432.0     4,032    49,666     20,217.2  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_t…  Push…  263  263         19         19       62,388,528  3,283,6…  1,512,0…  1,374,7…  9,959,7…  3,057,633.5  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_t…  Push…  263  263         19      1,311       52,699,894  40,198.2  40,225.0    39,234    41,089        315.7  fused_sigmoid_gating_delta_rule_update_kernel                                                       
 :dflash_t…  Push…  263  263         19      3,078       45,024,429  14,627.8  19,024.5     5,344    23,617      7,343.8  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_t…  Push…  263  263         19      3,648       23,093,951   6,330.6   6,208.0     4,448     7,969        604.1  _score_kernel                                                                                       
 :dflash_t…  Push…  263  263         19      1,748       14,509,556   8,300.7   8,257.0     5,504    12,960      1,029.6  void tensorrt_llm::kernels::cutlass_kernels::finalizeMoeRoutingKernel<__nv_bfloat16, __nv_bfloat16,…
 :dflash_t…  Push…  263  263         19        456       12,514,082  27,443.2  27,457.0    25,857    29,121        555.4  kernel_cutlass__dsv3_fused_a_gemm_kernel_tensorptri32gmemo2112358435841_tensorptri32gmemo1635843584…
 :dflash_t…  Push…  263  263         19      2,204        8,550,931   3,879.7   4,160.0     2,400     4,736        658.8  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_t…  Push…  263  263         19      3,097        8,178,690   2,640.8   2,720.0     1,760     5,440        481.0  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_t…  Push…  263  263         19      3,648        7,666,351   2,101.5   2,112.0       928     3,233        530.1  _combine_kernel                                                                                     
 :dflash_t…  Push…  263  263         19      3,553        7,370,843   2,074.5   2,049.0     1,920     2,432         66.8  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  263  263         19      1,311        6,929,937   5,286.0   5,280.0     4,768     5,921        210.9  _causal_conv1d_update_kernel                                                                        
 :dflash_t…  Push…  263  263         19      1,748        6,720,417   3,844.6   3,839.0     3,647     4,256         81.7  void tensorrt_llm::kernels::cutlass_kernels::doActivationKernel<__nv_fp8_e4m3, __nv_bfloat16, __nv_…
 :dflash_t…  Push…  263  263         19      1,748        6,131,649   3,507.8   3,488.0     3,296     3,936         80.8  void tensorrt_llm::kernels::cutlass_kernels::expandInputRowsKernel<__nv_fp8_e4m3, __nv_fp8_e4m3, (t…
 :dflash_t…  Push…  263  263         19      1,748        4,653,544   2,662.2   2,656.0     2,336     3,136        110.0  void sglang::route_quant_fused_kernel<(bool)1, float, __nv_bfloat16>(sglang::RouteQuantFusedParams) 
 :dflash_t…  Push…  263  263         19      1,748        4,149,115   2,373.6   2,368.0     2,272     2,721         48.9  void tensorrt_llm::kernels::cutlass_kernels::blockExpertPrefixSumKernel<(int)256>(const int *, int …
 :dflash_t…  Push…  263  263         19      1,748        4,006,882   2,292.3   2,272.0     2,208     2,624         54.5  void tensorrt_llm::kernels::cutlass_kernels::computeStridesTmaWarpSpecializedKernel<__nv_fp8_e4m3, …
 :dflash_t…  Push…  263  263         19      1,748        3,594,540   2,056.4   2,048.0     1,920     2,560         64.3  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  263  263         19      1,311        3,465,504   2,643.4   2,624.0     2,336     3,424         93.9  void gemmSN_TN_kernel<float, (int)128, (int)16, (int)2, (int)4, (int)8, (int)9, (bool)0, cublasGemv…
 :dflash_t…  Push…  263  263         19      1,748        3,328,913   1,904.4   1,888.0     1,824     2,176         37.8  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  263  263         19      1,748        3,311,122   1,894.2   1,888.0     1,823     2,240         44.0  void tensorrt_llm::kernels::quantize_with_block_size<(tensorrt_llm::BlockScaleQuantizationType)2, _…
 :dflash_t…  Push…  263  263         19      1,748        3,172,346   1,814.8   1,824.0     1,760     2,049         32.1  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, float, float, float, (bool)0, floa…
 :dflash_t…  Push…  263  263         19      1,748        3,100,838   1,773.9   1,760.0     1,664     2,080         51.2  void sglang::situ_and_mul_kernel<float, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMulParams)  
 :dflash_t…  Push…  263  263         19      1,311        2,853,506   2,176.6   2,176.0     2,080     2,433         52.6  layer_norm_gated_fwd_kernel                                                                         
 :dflash_t…  Push…  263  263         19      1,482        2,824,942   1,906.2   1,856.0     1,632     5,664        418.1  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  263  263         19        456        2,791,101   6,120.8   6,112.0     5,856     6,464        103.4  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
 :dflash_t…  Push…  263  263         19      1,748        2,514,000   1,438.2   1,408.0     1,280     1,888        103.6  void sglang::add3_kernel<(bool)1, (bool)1>(sglang::Add3Params)                                      
 :dflash_t…  Push…  263  263         19      1,748        2,270,238   1,298.8   1,280.0     1,184     1,632         64.5  tensorrt_llm::kernels::cutlass_kernels::mergeExpertPrefixSumKernel(const int *, const int *, const …
 :dflash_t…  Push…  263  263         19        475        2,103,626   4,428.7   3,424.0     3,168    27,073      4,551.4  void cutlass::Kernel2<cutlass_80_wmma_tensorop_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Params)   
 :dflash_t…  Push…  263  263         19      1,748        2,017,595   1,154.2   1,152.0     1,056     1,441         42.9  void tensorrt_llm::kernels::cutlass_kernels::globalExpertPrefixSumKernel<(int)256>(const int *, int…
 :dflash_t…  Push…  263  263         19      1,634        1,592,301     974.5     960.0       800     1,376         38.4  void at::native::vectorized_elementwise_kernel<(int)4, at::native::CUDAFunctor_add<c10::BFloat16>, …
 :dflash_t…  Push…  263  263         19        912        1,511,405   1,657.2   1,648.5     1,536     1,984         97.0  void at::native::<unnamed>::CatArrayBatchedCopy<at::native::<unnamed>::OpaqueType<(unsigned int)2>,…
 :dflash_t…  Push…  263  263         19        456        1,345,866   2,951.5   2,912.5     2,304     3,904        280.4  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  263  263         19         19        1,230,470  64,761.6  64,674.0    64,066    65,922        573.5  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_t…  Push…  263  263         19        475          957,529   2,015.9   1,984.0     1,888     2,496        117.1  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, __nv_bfloat16, float, __nv_bfloat1…
 :dflash_t…  Push…  263  263         19        456          646,520   1,417.8   1,408.0     1,344     1,728         44.5  void sglang::mla_output_gate_kernel<(int)256, (bool)1>(sglang::MlaOutputGateParams)                 
 :dflash_t…  Push…  263  263         19        456          591,410   1,297.0   1,217.0     1,120     3,296        278.8  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  263  263         19         19          131,010   6,895.3   6,912.0     6,784     7,008         52.7  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_t…  Push…  263  263         19         19           66,369   3,493.1   3,488.0     3,392     3,552         34.2  void sglang::situ_and_mul_kernel<__nv_bfloat16, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMul…
 :dflash_t…  Push…  263  263         19         19           46,720   2,458.9   2,432.0     2,336     2,784        100.2  void at::native::<unnamed>::CatArrayBatchedCopy_vectorized<at::native::<unnamed>::OpaqueType<(unsig…
 :dflash_t…  Push…  263  263         19         19           22,305   1,173.9   1,249.0       832     1,632        263.7  _vocab_parallel_embedding_kernel                                                                    
 :dflash_t…  Push…  263  263         19         19           17,251     907.9     896.0       864       960         21.9  void at::native::vectorized_elementwise_kernel<(int)4, at::native::AUnaryFunctor<c10::BFloat16, c10…
 :dflash_t…  Push…  263  263         19         19           15,264     803.4     768.0       768       896         45.1  void at::native::vectorized_elementwise_kernel<(int)4, at::native::FillFunctor<float>, std::array<c…
 :dflash_t…  Push…  264  264         19        456    1,790,738,514  3,927,0…  3,927,2…  3,908,4…  3,949,6…      6,775.3  _fwd_kernel                                                                                         
 :dflash_t…  Push…  264  264         19      3,553    1,054,698,233  296,847…  290,248…   240,103  1,909,2…     50,856.6  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_t…  Push…  264  264         19      3,496      123,435,530  35,307.6  32,480.5    13,441    70,722     11,620.1  void cutlass::device_kernel<cutlass::gemm::kernel::GemmUniversal<cutlass::gemm::GroupProblemShape<c…
 :dflash_t…  Push…  264  264         19      1,748      104,960,236  60,045.9  59,985.5    58,082    73,538        912.8  void cutlass::Kernel2<cutlass_80_tensorop_s16816gemm_bf16_256x64_32x4_tn_align8>(T1::Params)        
 :dflash_t…  Push…  264  264         19      3,990       94,463,295  23,675.0   6,464.0     4,064    49,986     20,168.2  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_t…  Push…  264  264         19         19       62,610,155  3,295,2…  1,552,1…  1,381,9…  9,962,7…  3,061,572.6  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_t…  Push…  264  264         19      1,311       52,652,032  40,161.7  40,161.0    39,553    40,961        234.5  fused_sigmoid_gating_delta_rule_update_kernel                                                       
 :dflash_t…  Push…  264  264         19      3,078       44,685,152  14,517.6  18,128.5     5,280    24,097      7,301.6  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_t…  Push…  264  264         19      3,648       23,273,775   6,379.9   6,240.0     4,480     8,192        638.6  _score_kernel                                                                                       
 :dflash_t…  Push…  264  264         19      1,748       14,569,084   8,334.7   8,288.0     5,600    12,960      1,031.9  void tensorrt_llm::kernels::cutlass_kernels::finalizeMoeRoutingKernel<__nv_bfloat16, __nv_bfloat16,…
 :dflash_t…  Push…  264  264         19        456       12,495,166  27,401.7  27,456.5    25,857    29,185        587.2  kernel_cutlass__dsv3_fused_a_gemm_kernel_tensorptri32gmemo2112358435841_tensorptri32gmemo1635843584…
 :dflash_t…  Push…  264  264         19      2,204        8,575,148   3,890.7   4,160.0     2,400     4,705        644.0  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_t…  Push…  264  264         19      3,097        8,192,694   2,645.4   2,720.0     1,792     5,345        498.7  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_t…  Push…  264  264         19      3,648        7,566,148   2,074.1   2,112.0       928     3,136        517.0  _combine_kernel                                                                                     
 :dflash_t…  Push…  264  264         19      3,553        7,372,455   2,075.0   2,080.0     1,888     2,400         63.2  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  264  264         19      1,311        6,885,376   5,252.0   5,248.0     4,800     6,336        208.9  _causal_conv1d_update_kernel                                                                        
 :dflash_t…  Push…  264  264         19      1,748        6,708,139   3,837.6   3,839.0     3,616     4,224         78.9  void tensorrt_llm::kernels::cutlass_kernels::doActivationKernel<__nv_fp8_e4m3, __nv_bfloat16, __nv_…
 :dflash_t…  Push…  264  264         19      1,748        6,082,782   3,479.9   3,456.0     3,328     3,873         76.3  void tensorrt_llm::kernels::cutlass_kernels::expandInputRowsKernel<__nv_fp8_e4m3, __nv_fp8_e4m3, (t…
 :dflash_t…  Push…  264  264         19      1,748        4,649,028   2,659.6   2,656.0     2,304     3,200        106.9  void sglang::route_quant_fused_kernel<(bool)1, float, __nv_bfloat16>(sglang::RouteQuantFusedParams) 
 :dflash_t…  Push…  264  264         19      1,748        4,131,057   2,363.3   2,368.0     2,240     2,688         45.6  void tensorrt_llm::kernels::cutlass_kernels::blockExpertPrefixSumKernel<(int)256>(const int *, int …
 :dflash_t…  Push…  264  264         19      1,748        3,986,699   2,280.7   2,272.0     2,208     2,624         53.2  void tensorrt_llm::kernels::cutlass_kernels::computeStridesTmaWarpSpecializedKernel<__nv_fp8_e4m3, …
 :dflash_t…  Push…  264  264         19      1,748        3,610,785   2,065.7   2,048.0     1,952     2,432         62.1  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  264  264         19      1,311        3,506,093   2,674.4   2,656.0     2,336     3,232         96.3  void gemmSN_TN_kernel<float, (int)128, (int)16, (int)2, (int)4, (int)8, (int)9, (bool)0, cublasGemv…
 :dflash_t…  Push…  264  264         19      1,748        3,309,852   1,893.5   1,888.0     1,823     2,240         39.7  void tensorrt_llm::kernels::quantize_with_block_size<(tensorrt_llm::BlockScaleQuantizationType)2, _…
 :dflash_t…  Push…  264  264         19      1,748        3,308,673   1,892.8   1,888.0     1,824     2,144         34.0  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  264  264         19      1,748        3,167,956   1,812.3   1,793.0     1,728     2,016         32.9  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, float, float, float, (bool)0, floa…
 :dflash_t…  Push…  264  264         19      1,748        3,122,683   1,786.4   1,792.0     1,664     2,113         46.0  void sglang::situ_and_mul_kernel<float, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMulParams)  
 :dflash_t…  Push…  264  264         19      1,311        2,892,184   2,206.1   2,208.0     2,112     2,496         51.1  layer_norm_gated_fwd_kernel                                                                         
 :dflash_t…  Push…  264  264         19      1,482        2,810,787   1,896.6   1,856.0     1,632     5,600        416.9  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  264  264         19        456        2,792,584   6,124.1   6,112.0     5,920     6,464         89.9  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
 :dflash_t…  Push…  264  264         19      1,748        2,477,245   1,417.2   1,376.0     1,280     2,048         99.0  void sglang::add3_kernel<(bool)1, (bool)1>(sglang::Add3Params)                                      
 :dflash_t…  Push…  264  264         19      1,748        2,290,332   1,310.3   1,280.0     1,216     1,760         68.4  tensorrt_llm::kernels::cutlass_kernels::mergeExpertPrefixSumKernel(const int *, const int *, const …
 :dflash_t…  Push…  264  264         19        475        2,124,855   4,473.4   3,456.0     3,168    27,137      4,507.0  void cutlass::Kernel2<cutlass_80_wmma_tensorop_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Params)   
 :dflash_t…  Push…  264  264         19      1,748        2,055,778   1,176.1   1,152.0     1,120     1,504         44.6  void tensorrt_llm::kernels::cutlass_kernels::globalExpertPrefixSumKernel<(int)256>(const int *, int…
 :dflash_t…  Push…  264  264         19      1,634        1,578,389     966.0     960.0       800     1,184         37.1  void at::native::vectorized_elementwise_kernel<(int)4, at::native::CUDAFunctor_add<c10::BFloat16>, …
 :dflash_t…  Push…  264  264         19        912        1,518,251   1,664.7   1,664.0     1,536     1,952         99.4  void at::native::<unnamed>::CatArrayBatchedCopy<at::native::<unnamed>::OpaqueType<(unsigned int)2>,…
 :dflash_t…  Push…  264  264         19        456        1,377,190   3,020.2   3,008.0     2,272     3,968        305.3  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  264  264         19         19        1,229,184  64,693.9  64,834.0    63,425    65,602        572.9  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_t…  Push…  264  264         19        475          951,764   2,003.7   1,984.0     1,856     2,496        118.6  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, __nv_bfloat16, float, __nv_bfloat1…
 :dflash_t…  Push…  264  264         19        456          655,952   1,438.5   1,440.0     1,376     1,760         43.4  void sglang::mla_output_gate_kernel<(int)256, (bool)1>(sglang::MlaOutputGateParams)                 
 :dflash_t…  Push…  264  264         19        456          596,687   1,308.5   1,248.0     1,120     3,040        280.9  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  264  264         19         19          129,761   6,829.5   6,816.0     6,752     6,880         41.8  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_t…  Push…  264  264         19         19           66,401   3,494.8   3,488.0     3,360     3,648         59.9  void sglang::situ_and_mul_kernel<__nv_bfloat16, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMul…
 :dflash_t…  Push…  264  264         19         19           46,242   2,433.8   2,400.0     2,272     2,912        157.3  void at::native::<unnamed>::CatArrayBatchedCopy_vectorized<at::native::<unnamed>::OpaqueType<(unsig…
 :dflash_t…  Push…  264  264         19         19           22,048   1,160.4   1,248.0       800     1,472        232.9  _vocab_parallel_embedding_kernel                                                                    
 :dflash_t…  Push…  264  264         19         19           16,960     892.6     896.0       832       992         36.8  void at::native::vectorized_elementwise_kernel<(int)4, at::native::AUnaryFunctor<c10::BFloat16, c10…
 :dflash_t…  Push…  264  264         19         19           15,137     796.7     800.0       768       928         35.2  void at::native::vectorized_elementwise_kernel<(int)4, at::native::FillFunctor<float>, std::array<c…
 :dflash_t…  Push…  265  265         19        456    1,789,354,295  3,924,0…  3,923,5…  3,902,8…  3,948,0…      7,735.8  _fwd_kernel                                                                                         
 :dflash_t…  Push…  265  265         19      3,553    1,054,474,756  296,784…  287,368…   260,615  1,854,8…     48,107.2  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_t…  Push…  265  265         19      3,496      123,657,209  35,371.1  32,673.0    13,216    73,090     11,615.0  void cutlass::device_kernel<cutlass::gemm::kernel::GemmUniversal<cutlass::gemm::GroupProblemShape<c…
 :dflash_t…  Push…  265  265         19      1,748      105,015,992  60,077.8  60,002.0    58,306    71,042        857.0  void cutlass::Kernel2<cutlass_80_tensorop_s16816gemm_bf16_256x64_32x4_tn_align8>(T1::Params)        
 :dflash_t…  Push…  265  265         19      3,990       94,289,909  23,631.6   6,401.0     4,096    50,178     20,142.0  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_t…  Push…  265  265         19         19       62,995,528  3,315,5…  1,594,0…  1,401,0…  9,951,2…  3,055,786.8  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_t…  Push…  265  265         19      1,311       53,184,542  40,567.9  40,577.0    39,937    41,377        229.0  fused_sigmoid_gating_delta_rule_update_kernel                                                       
 :dflash_t…  Push…  265  265         19      3,078       45,042,612  14,633.7  18,736.0     5,345    23,617      7,338.9  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_t…  Push…  265  265         19      3,648       23,146,486   6,345.0   6,272.0     4,512     8,064        574.1  _score_kernel                                                                                       
 :dflash_t…  Push…  265  265         19      1,748       14,701,118   8,410.3   8,320.5     5,600    13,345      1,045.1  void tensorrt_llm::kernels::cutlass_kernels::finalizeMoeRoutingKernel<__nv_bfloat16, __nv_bfloat16,…
 :dflash_t…  Push…  265  265         19        456       12,500,081  27,412.5  27,361.0    25,953    29,889        667.4  kernel_cutlass__dsv3_fused_a_gemm_kernel_tensorptri32gmemo2112358435841_tensorptri32gmemo1635843584…
 :dflash_t…  Push…  265  265         19      2,204        8,600,169   3,902.1   4,160.0     2,400     4,768        657.0  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_t…  Push…  265  265         19      3,097        8,195,216   2,646.2   2,752.0     1,760     5,504        480.5  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_t…  Push…  265  265         19      3,648        7,659,277   2,099.6   2,112.0       928     3,201        525.3  _combine_kernel                                                                                     
 :dflash_t…  Push…  265  265         19      3,553        7,405,966   2,084.4   2,080.0     1,920     2,528         65.2  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  265  265         19      1,311        7,096,510   5,413.1   5,472.0     4,896     6,304        194.2  _causal_conv1d_update_kernel                                                                        
 :dflash_t…  Push…  265  265         19      1,748        6,681,214   3,822.2   3,808.0     3,616     4,224         80.7  void tensorrt_llm::kernels::cutlass_kernels::doActivationKernel<__nv_fp8_e4m3, __nv_bfloat16, __nv_…
 :dflash_t…  Push…  265  265         19      1,748        6,113,509   3,497.4   3,488.0     3,296     3,872         85.1  void tensorrt_llm::kernels::cutlass_kernels::expandInputRowsKernel<__nv_fp8_e4m3, __nv_fp8_e4m3, (t…
 :dflash_t…  Push…  265  265         19      1,748        4,686,121   2,680.8   2,656.0     2,336     3,104        110.1  void sglang::route_quant_fused_kernel<(bool)1, float, __nv_bfloat16>(sglang::RouteQuantFusedParams) 
 :dflash_t…  Push…  265  265         19      1,748        4,170,808   2,386.0   2,368.0     2,272     2,752         52.8  void tensorrt_llm::kernels::cutlass_kernels::blockExpertPrefixSumKernel<(int)256>(const int *, int …
 :dflash_t…  Push…  265  265         19      1,748        4,034,189   2,307.9   2,304.0     2,240     2,752         57.9  void tensorrt_llm::kernels::cutlass_kernels::computeStridesTmaWarpSpecializedKernel<__nv_fp8_e4m3, …
 :dflash_t…  Push…  265  265         19      1,748        3,639,081   2,081.9   2,080.0     1,952     4,608         87.8  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  265  265         19      1,311        3,467,718   2,645.1   2,624.0     2,400     3,200         85.9  void gemmSN_TN_kernel<float, (int)128, (int)16, (int)2, (int)4, (int)8, (int)9, (bool)0, cublasGemv…
 :dflash_t…  Push…  265  265         19      1,748        3,334,581   1,907.7   1,888.0     1,824     2,112         39.4  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  265  265         19      1,748        3,324,527   1,901.9   1,888.0     1,823     2,272         45.8  void tensorrt_llm::kernels::quantize_with_block_size<(tensorrt_llm::BlockScaleQuantizationType)2, _…
 :dflash_t…  Push…  265  265         19      1,748        3,201,396   1,831.5   1,824.0     1,760     2,048         37.7  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, float, float, float, (bool)0, floa…
 :dflash_t…  Push…  265  265         19      1,748        3,127,732   1,789.3   1,792.0     1,696     2,144         52.6  void sglang::situ_and_mul_kernel<float, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMulParams)  
 :dflash_t…  Push…  265  265         19      1,311        2,879,833   2,196.7   2,176.0     2,080     2,496         54.2  layer_norm_gated_fwd_kernel                                                                         
 :dflash_t…  Push…  265  265         19      1,482        2,834,202   1,912.4   1,857.0     1,632     5,728        423.6  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  265  265         19        456        2,827,352   6,200.3   6,176.0     5,984     6,592        116.2  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
 :dflash_t…  Push…  265  265         19      1,748        2,589,234   1,481.3   1,472.0     1,280     2,080        115.4  void sglang::add3_kernel<(bool)1, (bool)1>(sglang::Add3Params)                                      
 :dflash_t…  Push…  265  265         19      1,748        2,306,794   1,319.7   1,312.0     1,216     1,728         66.7  tensorrt_llm::kernels::cutlass_kernels::mergeExpertPrefixSumKernel(const int *, const int *, const …
 :dflash_t…  Push…  265  265         19        475        2,124,828   4,473.3   3,424.0     3,168    26,945      4,544.2  void cutlass::Kernel2<cutlass_80_wmma_tensorop_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Params)   
 :dflash_t…  Push…  265  265         19      1,748        2,039,080   1,166.5   1,152.0     1,088     1,568         47.6  void tensorrt_llm::kernels::cutlass_kernels::globalExpertPrefixSumKernel<(int)256>(const int *, int…
 :dflash_t…  Push…  265  265         19      1,634        1,592,875     974.8     960.0       800     1,184         35.3  void at::native::vectorized_elementwise_kernel<(int)4, at::native::CUDAFunctor_add<c10::BFloat16>, …
 :dflash_t…  Push…  265  265         19        912        1,518,224   1,664.7   1,664.0     1,536     1,984         95.5  void at::native::<unnamed>::CatArrayBatchedCopy<at::native::<unnamed>::OpaqueType<(unsigned int)2>,…
 :dflash_t…  Push…  265  265         19        456        1,354,314   2,970.0   2,944.0     2,400     3,936        300.3  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  265  265         19         19        1,230,594  64,768.1  64,866.0    63,554    66,466        789.3  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_t…  Push…  265  265         19        475          959,289   2,019.6   1,984.0     1,888     2,560        128.0  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, __nv_bfloat16, float, __nv_bfloat1…
 :dflash_t…  Push…  265  265         19        456          655,304   1,437.1   1,440.0     1,376     1,760         41.5  void sglang::mla_output_gate_kernel<(int)256, (bool)1>(sglang::MlaOutputGateParams)                 
 :dflash_t…  Push…  265  265         19        456          600,565   1,317.0   1,248.0     1,120     3,361        337.7  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  265  265         19         19          130,948   6,892.0   6,880.0     6,816     7,104         61.4  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_t…  Push…  265  265         19         19           66,337   3,491.4   3,488.0     3,456     3,521         26.0  void sglang::situ_and_mul_kernel<__nv_bfloat16, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMul…
 :dflash_t…  Push…  265  265         19         19           47,492   2,499.6   2,464.0     2,336     2,752        120.7  void at::native::<unnamed>::CatArrayBatchedCopy_vectorized<at::native::<unnamed>::OpaqueType<(unsig…
 :dflash_t…  Push…  265  265         19         19           21,376   1,125.1     960.0       832     2,240        359.6  _vocab_parallel_embedding_kernel                                                                    
 :dflash_t…  Push…  265  265         19         19           17,121     901.1     896.0       832     1,024         35.8  void at::native::vectorized_elementwise_kernel<(int)4, at::native::AUnaryFunctor<c10::BFloat16, c10…
 :dflash_t…  Push…  265  265         19         19           15,040     791.6     800.0       736       896         36.7  void at::native::vectorized_elementwise_kernel<(int)4, at::native::FillFunctor<float>, std::array<c…
 :dflash_t…  Push…  266  266         19        456    1,782,684,325  3,909,3…  3,909,3…  3,889,8…  3,929,3…      7,502.8  _fwd_kernel                                                                                         
 :dflash_t…  Push…  266  266         19      3,553    1,061,411,926  298,736…  286,921…   260,040  1,919,1…     49,661.0  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_t…  Push…  266  266         19      3,496      123,839,851  35,423.3  32,721.0    13,633    71,778     11,595.8  void cutlass::device_kernel<cutlass::gemm::kernel::GemmUniversal<cutlass::gemm::GroupProblemShape<c…
 :dflash_t…  Push…  266  266         19      1,748      105,112,787  60,133.2  60,066.0    57,954    72,322        950.1  void cutlass::Kernel2<cutlass_80_tensorop_s16816gemm_bf16_256x64_32x4_tn_align8>(T1::Params)        
 :dflash_t…  Push…  266  266         19      3,990       94,586,619  23,705.9   6,432.0     4,064    49,857     20,208.3  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_t…  Push…  266  266         19         19       63,010,946  3,316,3…  1,590,7…  1,406,7…  9,954,7…  3,052,598.1  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_t…  Push…  266  266         19      1,311       52,914,837  40,362.2  40,353.0    39,873    41,121        234.5  fused_sigmoid_gating_delta_rule_update_kernel                                                       
 :dflash_t…  Push…  266  266         19      3,078       45,108,898  14,655.3  18,880.0     5,184    27,009      7,360.2  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_t…  Push…  266  266         19      3,648       23,225,811   6,366.7   6,272.0     4,480     7,904        595.8  _score_kernel                                                                                       
 :dflash_t…  Push…  266  266         19      1,748       14,483,518   8,285.8   8,225.0     5,440    13,184      1,029.4  void tensorrt_llm::kernels::cutlass_kernels::finalizeMoeRoutingKernel<__nv_bfloat16, __nv_bfloat16,…
 :dflash_t…  Push…  266  266         19        456       12,516,597  27,448.7  27,489.0    25,761    28,801        565.1  kernel_cutlass__dsv3_fused_a_gemm_kernel_tensorptri32gmemo2112358435841_tensorptri32gmemo1635843584…
 :dflash_t…  Push…  266  266         19      2,204        8,594,936   3,899.7   4,160.0     2,432     4,737        657.0  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_t…  Push…  266  266         19      3,097        8,229,240   2,657.2   2,752.0     1,760     5,408        489.5  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_t…  Push…  266  266         19      3,648        7,694,929   2,109.4   2,144.0       960     3,200        533.4  _combine_kernel                                                                                     
 :dflash_t…  Push…  266  266         19      3,553        7,373,739   2,075.4   2,080.0     1,920     2,496         67.6  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  266  266         19      1,311        6,812,660   5,196.5   5,152.0     4,801     6,752        225.6  _causal_conv1d_update_kernel                                                                        
 :dflash_t…  Push…  266  266         19      1,748        6,791,266   3,885.2   3,872.0     3,648     4,288         78.6  void tensorrt_llm::kernels::cutlass_kernels::doActivationKernel<__nv_fp8_e4m3, __nv_bfloat16, __nv_…
 :dflash_t…  Push…  266  266         19      1,748        6,272,355   3,588.3   3,584.0     3,424     3,969         82.0  void tensorrt_llm::kernels::cutlass_kernels::expandInputRowsKernel<__nv_fp8_e4m3, __nv_fp8_e4m3, (t…
 :dflash_t…  Push…  266  266         19      1,748        4,680,946   2,677.9   2,656.0     2,368     3,200        107.9  void sglang::route_quant_fused_kernel<(bool)1, float, __nv_bfloat16>(sglang::RouteQuantFusedParams) 
 :dflash_t…  Push…  266  266         19      1,748        4,174,504   2,388.2   2,369.0     2,304     2,688         46.6  void tensorrt_llm::kernels::cutlass_kernels::blockExpertPrefixSumKernel<(int)256>(const int *, int …
 :dflash_t…  Push…  266  266         19      1,748        4,068,985   2,327.8   2,304.0     2,240     2,656         50.6  void tensorrt_llm::kernels::cutlass_kernels::computeStridesTmaWarpSpecializedKernel<__nv_fp8_e4m3, …
 :dflash_t…  Push…  266  266         19      1,748        3,640,311   2,082.6   2,080.0     1,920     2,432         63.6  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  266  266         19      1,311        3,497,581   2,667.9   2,656.0     2,336     3,264         89.5  void gemmSN_TN_kernel<float, (int)128, (int)16, (int)2, (int)4, (int)8, (int)9, (bool)0, cublasGemv…
 :dflash_t…  Push…  266  266         19      1,748        3,353,984   1,918.8   1,920.0     1,824     2,176         39.5  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  266  266         19      1,748        3,279,310   1,876.0   1,856.0     1,791     2,176         42.7  void tensorrt_llm::kernels::quantize_with_block_size<(tensorrt_llm::BlockScaleQuantizationType)2, _…
 :dflash_t…  Push…  266  266         19      1,748        3,213,171   1,838.2   1,824.0     1,792     2,048         32.1  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, float, float, float, (bool)0, floa…
 :dflash_t…  Push…  266  266         19      1,748        3,124,382   1,787.4   1,792.0     1,696     2,048         44.0  void sglang::situ_and_mul_kernel<float, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMulParams)  
 :dflash_t…  Push…  266  266         19      1,482        2,886,101   1,947.4   1,920.0     1,696     5,632        415.7  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  266  266         19      1,311        2,839,384   2,165.8   2,144.0     2,080     2,496         51.2  layer_norm_gated_fwd_kernel                                                                         
 :dflash_t…  Push…  266  266         19        456        2,764,597   6,062.7   6,048.0     5,856     6,464        108.2  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
 :dflash_t…  Push…  266  266         19      1,748        2,510,145   1,436.0   1,408.0     1,280     2,016        104.0  void sglang::add3_kernel<(bool)1, (bool)1>(sglang::Add3Params)                                      
 :dflash_t…  Push…  266  266         19      1,748        2,302,368   1,317.1   1,312.0     1,216     1,664         61.9  tensorrt_llm::kernels::cutlass_kernels::mergeExpertPrefixSumKernel(const int *, const int *, const …
 :dflash_t…  Push…  266  266         19        475        2,128,872   4,481.8   3,456.0     3,200    27,457      4,552.7  void cutlass::Kernel2<cutlass_80_wmma_tensorop_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Params)   
 :dflash_t…  Push…  266  266         19      1,748        2,005,764   1,147.5   1,152.0     1,088     1,408         42.8  void tensorrt_llm::kernels::cutlass_kernels::globalExpertPrefixSumKernel<(int)256>(const int *, int…
 :dflash_t…  Push…  266  266         19      1,634        1,601,594     980.2     992.0       800     1,216         36.1  void at::native::vectorized_elementwise_kernel<(int)4, at::native::CUDAFunctor_add<c10::BFloat16>, …
 :dflash_t…  Push…  266  266         19        912        1,499,824   1,644.5   1,632.0     1,504     1,952         95.1  void at::native::<unnamed>::CatArrayBatchedCopy<at::native::<unnamed>::OpaqueType<(unsigned int)2>,…
 :dflash_t…  Push…  266  266         19        456        1,354,544   2,970.5   2,944.0     2,240     4,864        302.8  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  266  266         19         19        1,239,687  65,246.7  65,186.0    64,034    66,530        697.3  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_t…  Push…  266  266         19        475          966,869   2,035.5   1,984.0     1,920     2,593        124.8  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, __nv_bfloat16, float, __nv_bfloat1…
 :dflash_t…  Push…  266  266         19        456          640,437   1,404.5   1,408.0     1,344     1,600         30.4  void sglang::mla_output_gate_kernel<(int)256, (bool)1>(sglang::MlaOutputGateParams)                 
 :dflash_t…  Push…  266  266         19        456          589,208   1,292.1   1,216.0     1,088     3,584        322.5  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  266  266         19         19          130,214   6,853.4   6,848.0     6,784     7,008         56.7  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_t…  Push…  266  266         19         19           66,945   3,523.4   3,520.0     3,392     3,584         42.5  void sglang::situ_and_mul_kernel<__nv_bfloat16, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMul…
 :dflash_t…  Push…  266  266         19         19           48,865   2,571.8   2,528.0     2,368     3,072        174.7  void at::native::<unnamed>::CatArrayBatchedCopy_vectorized<at::native::<unnamed>::OpaqueType<(unsig…
 :dflash_t…  Push…  266  266         19         19           23,137   1,217.7   1,312.0       896     1,568        235.3  _vocab_parallel_embedding_kernel                                                                    
 :dflash_t…  Push…  266  266         19         19           17,248     907.8     896.0       896       928         15.9  void at::native::vectorized_elementwise_kernel<(int)4, at::native::AUnaryFunctor<c10::BFloat16, c10…
 :dflash_t…  Push…  266  266         19         19           14,625     769.7     768.0       736       832         19.9  void at::native::vectorized_elementwise_kernel<(int)4, at::native::FillFunctor<float>, std::array<c…
 :dflash_t…  Push…  267  267         19        456    1,763,015,939  3,866,2…  3,866,7…  3,847,2…  3,892,9…      6,714.8  _fwd_kernel                                                                                         
 :dflash_t…  Push…  267  267         19      3,553    1,084,170,149  305,142…  286,922…   224,999  1,966,8…     56,108.5  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_t…  Push…  267  267         19      3,496      123,523,953  35,332.9  32,417.0    13,601    72,258     11,599.9  void cutlass::device_kernel<cutlass::gemm::kernel::GemmUniversal<cutlass::gemm::GroupProblemShape<c…
 :dflash_t…  Push…  267  267         19      1,748      105,353,209  60,270.7  60,194.0    57,794    73,187        979.6  void cutlass::Kernel2<cutlass_80_tensorop_s16816gemm_bf16_256x64_32x4_tn_align8>(T1::Params)        
 :dflash_t…  Push…  267  267         19      3,990       94,496,877  23,683.4   6,400.0     4,032    50,209     20,233.3  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_t…  Push…  267  267         19         19       63,136,141  3,322,9…  1,670,4…  1,408,8…  9,980,5…  3,053,583.7  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_t…  Push…  267  267         19      1,311       52,172,537  39,796.0  39,809.0    39,105    40,961        279.8  fused_sigmoid_gating_delta_rule_update_kernel                                                       
 :dflash_t…  Push…  267  267         19      3,078       44,969,418  14,609.9  19,040.5     5,248    23,777      7,338.9  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_t…  Push…  267  267         19      3,648       23,089,540   6,329.4   6,208.0     4,449     7,840        623.8  _score_kernel                                                                                       
 :dflash_t…  Push…  267  267         19      1,748       14,363,242   8,217.0   8,161.0     5,472    12,993      1,017.4  void tensorrt_llm::kernels::cutlass_kernels::finalizeMoeRoutingKernel<__nv_bfloat16, __nv_bfloat16,…
 :dflash_t…  Push…  267  267         19        456       12,505,545  27,424.4  27,425.0    25,953    28,993        571.6  kernel_cutlass__dsv3_fused_a_gemm_kernel_tensorptri32gmemo2112358435841_tensorptri32gmemo1635843584…
 :dflash_t…  Push…  267  267         19      2,204        8,537,923   3,873.8   4,128.0     2,400     4,800        642.9  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_t…  Push…  267  267         19      3,097        8,121,380   2,622.3   2,720.0     1,760     5,344        482.7  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_t…  Push…  267  267         19      3,648        7,614,321   2,087.3   2,112.0       928     3,200        523.8  _combine_kernel                                                                                     
 :dflash_t…  Push…  267  267         19      3,553        7,330,216   2,063.1   2,048.0     1,888     2,432         68.1  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  267  267         19      1,311        6,997,213   5,337.3   5,376.0     4,768     6,241        210.6  _causal_conv1d_update_kernel                                                                        
 :dflash_t…  Push…  267  267         19      1,748        6,734,266   3,852.6   3,840.0     3,647     4,320         77.7  void tensorrt_llm::kernels::cutlass_kernels::doActivationKernel<__nv_fp8_e4m3, __nv_bfloat16, __nv_…
 :dflash_t…  Push…  267  267         19      1,748        6,098,813   3,489.0   3,488.0     3,296     3,872         82.6  void tensorrt_llm::kernels::cutlass_kernels::expandInputRowsKernel<__nv_fp8_e4m3, __nv_fp8_e4m3, (t…
 :dflash_t…  Push…  267  267         19      1,748        4,602,366   2,632.9   2,624.0     2,336     3,136        107.8  void sglang::route_quant_fused_kernel<(bool)1, float, __nv_bfloat16>(sglang::RouteQuantFusedParams) 
 :dflash_t…  Push…  267  267         19      1,748        4,089,897   2,339.8   2,336.0     2,240     2,656         48.0  void tensorrt_llm::kernels::cutlass_kernels::blockExpertPrefixSumKernel<(int)256>(const int *, int …
 :dflash_t…  Push…  267  267         19      1,748        3,974,606   2,273.8   2,272.0     2,208     2,624         56.2  void tensorrt_llm::kernels::cutlass_kernels::computeStridesTmaWarpSpecializedKernel<__nv_fp8_e4m3, …
 :dflash_t…  Push…  267  267         19      1,748        3,574,135   2,044.7   2,017.0     1,920     3,488         75.7  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  267  267         19      1,311        3,474,831   2,650.5   2,624.0     2,304     3,456         97.7  void gemmSN_TN_kernel<float, (int)128, (int)16, (int)2, (int)4, (int)8, (int)9, (bool)0, cublasGemv…
 :dflash_t…  Push…  267  267         19      1,748        3,316,558   1,897.3   1,888.0     1,792     2,144         39.9  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  267  267         19      1,748        3,261,334   1,865.8   1,856.0     1,791     2,240         46.5  void tensorrt_llm::kernels::quantize_with_block_size<(tensorrt_llm::BlockScaleQuantizationType)2, _…
 :dflash_t…  Push…  267  267         19      1,748        3,167,790   1,812.2   1,792.0     1,728     2,048         37.3  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, float, float, float, (bool)0, floa…
 :dflash_t…  Push…  267  267         19      1,748        3,089,649   1,767.5   1,760.0     1,632     2,016         46.9  void sglang::situ_and_mul_kernel<float, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMulParams)  
 :dflash_t…  Push…  267  267         19      1,311        2,822,016   2,152.6   2,144.0     2,048     2,432         52.4  layer_norm_gated_fwd_kernel                                                                         
 :dflash_t…  Push…  267  267         19      1,482        2,820,825   1,903.4   1,856.0     1,632     5,537        414.6  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  267  267         19        456        2,770,633   6,075.9   6,049.0     5,824     6,369        105.1  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
 :dflash_t…  Push…  267  267         19      1,748        2,534,064   1,449.7   1,440.0     1,280     2,048        110.5  void sglang::add3_kernel<(bool)1, (bool)1>(sglang::Add3Params)                                      
 :dflash_t…  Push…  267  267         19      1,748        2,269,985   1,298.6   1,280.0     1,184     1,632         62.7  tensorrt_llm::kernels::cutlass_kernels::mergeExpertPrefixSumKernel(const int *, const int *, const …
 :dflash_t…  Push…  267  267         19        475        2,104,907   4,431.4   3,424.0     3,168    26,977      4,547.6  void cutlass::Kernel2<cutlass_80_wmma_tensorop_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Params)   
 :dflash_t…  Push…  267  267         19      1,748        1,988,442   1,137.6   1,120.0     1,088     1,472         42.4  void tensorrt_llm::kernels::cutlass_kernels::globalExpertPrefixSumKernel<(int)256>(const int *, int…
 :dflash_t…  Push…  267  267         19      1,634        1,571,605     961.8     960.0       800     1,152         34.2  void at::native::vectorized_elementwise_kernel<(int)4, at::native::CUDAFunctor_add<c10::BFloat16>, …
 :dflash_t…  Push…  267  267         19        912        1,511,920   1,657.8   1,648.5     1,504     1,984        100.8  void at::native::<unnamed>::CatArrayBatchedCopy<at::native::<unnamed>::OpaqueType<(unsigned int)2>,…
 :dflash_t…  Push…  267  267         19        456        1,332,520   2,922.2   2,880.0     2,336     4,032        288.6  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  267  267         19         19        1,232,139  64,849.4  64,834.0    63,971    66,083        624.7  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_t…  Push…  267  267         19        475          956,192   2,013.0   1,984.0     1,888     2,624        119.4  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, __nv_bfloat16, float, __nv_bfloat1…
 :dflash_t…  Push…  267  267         19        456          639,824   1,403.1   1,408.0     1,344     1,728         38.2  void sglang::mla_output_gate_kernel<(int)256, (bool)1>(sglang::MlaOutputGateParams)                 
 :dflash_t…  Push…  267  267         19        456          588,597   1,290.8   1,216.0     1,088     3,360        266.0  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  267  267         19         19          129,987   6,841.4   6,848.0     6,784     6,944         40.7  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_t…  Push…  267  267         19         19           66,433   3,496.5   3,488.0     3,392     3,553         39.8  void sglang::situ_and_mul_kernel<__nv_bfloat16, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMul…
 :dflash_t…  Push…  267  267         19         19           47,265   2,487.6   2,464.0     2,336     2,721        124.9  void at::native::<unnamed>::CatArrayBatchedCopy_vectorized<at::native::<unnamed>::OpaqueType<(unsig…
 :dflash_t…  Push…  267  267         19         19           19,712   1,037.5     960.0       800     2,048        319.7  _vocab_parallel_embedding_kernel                                                                    
 :dflash_t…  Push…  267  267         19         19           17,568     924.6     896.0       864     1,056         53.2  void at::native::vectorized_elementwise_kernel<(int)4, at::native::AUnaryFunctor<c10::BFloat16, c10…
 :dflash_t…  Push…  267  267         19         19           14,592     768.0     768.0       736       896         33.7  void at::native::vectorized_elementwise_kernel<(int)4, at::native::FillFunctor<float>, std::array<c…
 :dflash_t…  Push…  268  268         19        456    1,766,208,950  3,873,2…  3,872,2…  3,854,1…  3,904,1…      8,935.7  _fwd_kernel                                                                                         
 :dflash_t…  Push…  268  268         19      3,553    1,082,442,207  304,655…  300,264…   225,190  1,897,7…     56,428.2  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 :dflash_t…  Push…  268  268         19      3,496      123,488,797  35,322.9  32,513.0    12,864    70,530     11,551.5  void cutlass::device_kernel<cutlass::gemm::kernel::GemmUniversal<cutlass::gemm::GroupProblemShape<c…
 :dflash_t…  Push…  268  268         19      1,748      104,911,006  60,017.7  59,938.0    58,017    68,130        851.0  void cutlass::Kernel2<cutlass_80_tensorop_s16816gemm_bf16_256x64_32x4_tn_align8>(T1::Params)        
 :dflash_t…  Push…  268  268         19      3,990       94,301,890  23,634.6   6,369.0     4,032    50,722     20,178.5  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_16x16_128x2_tn_align8>(T1::Par…
 :dflash_t…  Push…  268  268         19         19       63,644,583  3,349,7…  1,710,4…  1,429,6…  10,032,…  3,050,186.9  ncclDevKernel_AllGather_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)                      
 :dflash_t…  Push…  268  268         19      1,311       52,134,236  39,766.8  39,745.0    38,946    40,802        308.1  fused_sigmoid_gating_delta_rule_update_kernel                                                       
 :dflash_t…  Push…  268  268         19      3,078       45,095,450  14,650.9  18,657.0     5,152    25,921      7,381.3  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x2_tn_align8>(T1::Para…
 :dflash_t…  Push…  268  268         19      3,648       22,889,158   6,274.4   6,240.0     4,481     7,744        546.9  _score_kernel                                                                                       
 :dflash_t…  Push…  268  268         19      1,748       14,350,293   8,209.5   8,192.0     5,824    12,832      1,017.1  void tensorrt_llm::kernels::cutlass_kernels::finalizeMoeRoutingKernel<__nv_bfloat16, __nv_bfloat16,…
 :dflash_t…  Push…  268  268         19        456       12,557,642  27,538.7  27,536.5    25,984    29,472        602.2  kernel_cutlass__dsv3_fused_a_gemm_kernel_tensorptri32gmemo2112358435841_tensorptri32gmemo1635843584…
 :dflash_t…  Push…  268  268         19      2,204        8,535,158   3,872.6   4,128.0     2,400     4,704        649.7  void cutlass::Kernel2<cutlass_80_wmma_tensorop_bf16_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Para…
 :dflash_t…  Push…  268  268         19      3,097        8,134,000   2,626.4   2,689.0     1,760     5,184        494.5  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, __nv_bfloat16, __nv_bfloat16, float, __nv…
 :dflash_t…  Push…  268  268         19      3,648        7,536,771   2,066.0   2,080.0       896     3,040        520.4  _combine_kernel                                                                                     
 :dflash_t…  Push…  268  268         19      3,553        7,351,510   2,069.1   2,048.0     1,888     2,464         70.6  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  268  268         19      1,311        7,008,368   5,345.8   5,408.0     4,736     6,784        199.1  _causal_conv1d_update_kernel                                                                        
 :dflash_t…  Push…  268  268         19      1,748        6,726,601   3,848.2   3,840.0     3,647     4,863         81.6  void tensorrt_llm::kernels::cutlass_kernels::doActivationKernel<__nv_fp8_e4m3, __nv_bfloat16, __nv_…
 :dflash_t…  Push…  268  268         19      1,748        6,062,874   3,468.5   3,456.0     3,296     3,936         85.2  void tensorrt_llm::kernels::cutlass_kernels::expandInputRowsKernel<__nv_fp8_e4m3, __nv_fp8_e4m3, (t…
 :dflash_t…  Push…  268  268         19      1,748        4,613,628   2,639.4   2,624.0     2,304     3,073        109.6  void sglang::route_quant_fused_kernel<(bool)1, float, __nv_bfloat16>(sglang::RouteQuantFusedParams) 
 :dflash_t…  Push…  268  268         19      1,748        4,095,255   2,342.8   2,336.0     2,240     2,656         53.0  void tensorrt_llm::kernels::cutlass_kernels::blockExpertPrefixSumKernel<(int)256>(const int *, int …
 :dflash_t…  Push…  268  268         19      1,748        3,965,994   2,268.9   2,240.0     2,176     2,624         59.8  void tensorrt_llm::kernels::cutlass_kernels::computeStridesTmaWarpSpecializedKernel<__nv_fp8_e4m3, …
 :dflash_t…  Push…  268  268         19      1,748        3,592,170   2,055.0   2,048.0     1,920     2,369         65.6  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign12…
 :dflash_t…  Push…  268  268         19      1,311        3,455,834   2,636.0   2,624.0     2,368     3,232         86.3  void gemmSN_TN_kernel<float, (int)128, (int)16, (int)2, (int)4, (int)8, (int)9, (bool)0, cublasGemv…
 :dflash_t…  Push…  268  268         19      1,748        3,294,790   1,884.9   1,888.0     1,792     2,112         38.7  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  268  268         19      1,748        3,286,118   1,879.9   1,856.0     1,823     2,240         47.1  void tensorrt_llm::kernels::quantize_with_block_size<(tensorrt_llm::BlockScaleQuantizationType)2, _…
 :dflash_t…  Push…  268  268         19      1,748        3,168,506   1,812.6   1,793.0     1,728     2,016         34.1  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, float, float, float, (bool)0, floa…
 :dflash_t…  Push…  268  268         19      1,748        3,106,742   1,777.3   1,760.0     1,664     2,112         52.3  void sglang::situ_and_mul_kernel<float, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMulParams)  
 :dflash_t…  Push…  268  268         19      1,311        2,868,096   2,187.7   2,176.0     2,080     2,496         51.7  layer_norm_gated_fwd_kernel                                                                         
 :dflash_t…  Push…  268  268         19        456        2,809,586   6,161.4   6,144.0     5,984     6,560        101.1  void at::native::index_elementwise_kernel<(int)128, (int)4, void at::native::gpu_index_kernel<void …
 :dflash_t…  Push…  268  268         19      1,482        2,783,159   1,878.0   1,824.0     1,600     5,568        418.7  void at::native::elementwise_kernel<(int)128, (int)4, void at::native::gpu_kernel_impl_nocast<at::n…
 :dflash_t…  Push…  268  268         19      1,748        2,438,092   1,394.8   1,344.0     1,280     1,856         94.7  void sglang::add3_kernel<(bool)1, (bool)1>(sglang::Add3Params)                                      
 :dflash_t…  Push…  268  268         19      1,748        2,248,739   1,286.5   1,280.0     1,184     1,664         67.8  tensorrt_llm::kernels::cutlass_kernels::mergeExpertPrefixSumKernel(const int *, const int *, const …
 :dflash_t…  Push…  268  268         19        475        2,122,197   4,467.8   3,456.0     3,200    27,360      4,552.8  void cutlass::Kernel2<cutlass_80_wmma_tensorop_s161616gemm_bf16_32x32_64x1_tn_align8>(T1::Params)   
 :dflash_t…  Push…  268  268         19      1,748        2,056,253   1,176.3   1,152.0     1,088     1,504         50.1  void tensorrt_llm::kernels::cutlass_kernels::globalExpertPrefixSumKernel<(int)256>(const int *, int…
 :dflash_t…  Push…  268  268         19      1,634        1,566,023     958.4     960.0       800     1,152         35.5  void at::native::vectorized_elementwise_kernel<(int)4, at::native::CUDAFunctor_add<c10::BFloat16>, …
 :dflash_t…  Push…  268  268         19        912        1,520,404   1,667.1   1,664.5     1,504     2,048        100.0  void at::native::<unnamed>::CatArrayBatchedCopy<at::native::<unnamed>::OpaqueType<(unsigned int)2>,…
 :dflash_t…  Push…  268  268         19        456        1,355,857   2,973.4   2,976.0     2,304     3,872        283.9  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  268  268         19         19        1,239,109  65,216.3  65,058.0    64,258    66,914        704.2  void cutlass::Kernel2<cutlass_80_tensorop_bf16_s16816gemm_relu_bf16_256x64_32x4_tn_align8>(T1::Para…
 :dflash_t…  Push…  268  268         19        475          947,547   1,994.8   1,952.0     1,888     2,497        115.4  void cublasLt::splitKreduce_kernel<(int)32, (int)16, int, float, __nv_bfloat16, float, __nv_bfloat1…
 :dflash_t…  Push…  268  268         19        456          645,334   1,415.2   1,408.0     1,344     1,696         47.4  void sglang::mla_output_gate_kernel<(int)256, (bool)1>(sglang::MlaOutputGateParams)                 
 :dflash_t…  Push…  268  268         19        456          596,588   1,308.3   1,248.0     1,088     3,488        305.0  kernel_cutlass_kernel_flashinfernormkernelsrmsnormRMSNormKernel_object_at__tensorptrbf16gmemalign16…
 :dflash_t…  Push…  268  268         19         19          129,733   6,828.1   6,816.0     6,721     6,945         54.6  void at::native::unrolled_elementwise_kernel<at::native::direct_copy_kernel_cuda(at::TensorIterator…
 :dflash_t…  Push…  268  268         19         19           66,049   3,476.3   3,488.0     3,424     3,552         34.2  void sglang::situ_and_mul_kernel<__nv_bfloat16, __nv_bfloat16, (bool)1, (bool)1>(sglang::SituAndMul…
 :dflash_t…  Push…  268  268         19         19           45,153   2,376.5   2,336.0     2,272     2,592         89.8  void at::native::<unnamed>::CatArrayBatchedCopy_vectorized<at::native::<unnamed>::OpaqueType<(unsig…
 :dflash_t…  Push…  268  268         19         19           17,184     904.4     928.0       800       960         58.3  _vocab_parallel_embedding_kernel                                                                    
 :dflash_t…  Push…  268  268         19         19           17,089     899.4     896.0       864       992         25.9  void at::native::vectorized_elementwise_kernel<(int)4, at::native::AUnaryFunctor<c10::BFloat16, c10…
 :dflash_t…  Push…  268  268         19         19           15,008     789.9     800.0       768       832         24.0  void at::native::vectorized_elementwise_kernel<(int)4, at::native::FillFunctor<float>, std::array<c…
 CCCL:cub:…  Push…  261  261          1          4           20,257   5,064.3   4,864.5     4,768     5,760        468.9  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortOnesweepKernel<at_cuda_detail::cub::de…
 CCCL:cub:…  Push…  261  261          1          1            2,272   2,272.0   2,272.0     2,272     2,272          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortHistogramKernel<at_cuda_detail::cub::d…
 CCCL:cub:…  Push…  261  261          1          1              993     993.0     993.0       993       993          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortExclusiveSumKernel<at_cuda_detail::cub…
 CCCL:cub:…  Push…  262  262          1          4           20,224   5,056.0   4,864.0     4,768     5,728        453.3  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortOnesweepKernel<at_cuda_detail::cub::de…
 CCCL:cub:…  Push…  262  262          1          1            2,272   2,272.0   2,272.0     2,272     2,272          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortHistogramKernel<at_cuda_detail::cub::d…
 CCCL:cub:…  Push…  262  262          1          1              992     992.0     992.0       992       992          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortExclusiveSumKernel<at_cuda_detail::cub…
 CCCL:cub:…  Push…  263  263          1          4           20,225   5,056.3   4,848.0     4,768     5,761        473.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortOnesweepKernel<at_cuda_detail::cub::de…
 CCCL:cub:…  Push…  263  263          1          1            2,272   2,272.0   2,272.0     2,272     2,272          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortHistogramKernel<at_cuda_detail::cub::d…
 CCCL:cub:…  Push…  263  263          1          1            1,024   1,024.0   1,024.0     1,024     1,024          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortExclusiveSumKernel<at_cuda_detail::cub…
 CCCL:cub:…  Push…  264  264          1          4           20,481   5,120.3   4,880.5     4,832     5,888        513.8  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortOnesweepKernel<at_cuda_detail::cub::de…
 CCCL:cub:…  Push…  264  264          1          1            2,368   2,368.0   2,368.0     2,368     2,368          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortHistogramKernel<at_cuda_detail::cub::d…
 CCCL:cub:…  Push…  264  264          1          1            1,024   1,024.0   1,024.0     1,024     1,024          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortExclusiveSumKernel<at_cuda_detail::cub…
 CCCL:cub:…  Push…  265  265          1          4           20,481   5,120.3   4,928.0     4,832     5,793        453.8  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortOnesweepKernel<at_cuda_detail::cub::de…
 CCCL:cub:…  Push…  265  265          1          1            2,272   2,272.0   2,272.0     2,272     2,272          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortHistogramKernel<at_cuda_detail::cub::d…
 CCCL:cub:…  Push…  265  265          1          1            1,024   1,024.0   1,024.0     1,024     1,024          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortExclusiveSumKernel<at_cuda_detail::cub…
 CCCL:cub:…  Push…  266  266          1          4           20,320   5,080.0   4,896.0     4,800     5,728        437.5  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortOnesweepKernel<at_cuda_detail::cub::de…
 CCCL:cub:…  Push…  266  266          1          1            2,272   2,272.0   2,272.0     2,272     2,272          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortHistogramKernel<at_cuda_detail::cub::d…
 CCCL:cub:…  Push…  266  266          1          1              992     992.0     992.0       992       992          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortExclusiveSumKernel<at_cuda_detail::cub…
 CCCL:cub:…  Push…  267  267          1          4           19,713   4,928.3   4,736.0     4,640     5,601        453.8  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortOnesweepKernel<at_cuda_detail::cub::de…
 CCCL:cub:…  Push…  267  267          1          1            2,240   2,240.0   2,240.0     2,240     2,240          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortHistogramKernel<at_cuda_detail::cub::d…
 CCCL:cub:…  Push…  267  267          1          1            1,120   1,120.0   1,120.0     1,120     1,120          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortExclusiveSumKernel<at_cuda_detail::cub…
 CCCL:cub:…  Push…  268  268          1          4           20,065   5,016.3   4,832.0     4,705     5,696        463.8  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortOnesweepKernel<at_cuda_detail::cub::de…
 CCCL:cub:…  Push…  268  268          1          1            2,272   2,272.0   2,272.0     2,272     2,272          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortHistogramKernel<at_cuda_detail::cub::d…
 CCCL:cub:…  Push…  268  268          1          1            1,056   1,056.0   1,056.0     1,056     1,056          0.0  void at_cuda_detail::cub::detail::radix_sort::DeviceRadixSortExclusiveSumKernel<at_cuda_detail::cub…
 CCCL:cub:…  Push…  261  261        100        100          125,349   1,253.5   1,056.0       992     1,697        266.8  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 CCCL:cub:…  Push…  261  261        100        100           61,217     612.2     544.0       512       864        114.8  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 CCCL:cub:…  Push…  262  262         95         95          117,410   1,235.9   1,056.0       992     1,632        256.7  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 CCCL:cub:…  Push…  262  262         95         95           58,785     618.8     544.0       512       992        126.6  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 CCCL:cub:…  Push…  263  263         95         95          118,469   1,247.0   1,056.0       960     1,696        259.2  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 CCCL:cub:…  Push…  263  263         95         95           58,146     612.1     544.0       512       928        119.3  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 CCCL:cub:…  Push…  264  264         95         95          120,993   1,273.6   1,056.0     1,024     1,696        266.5  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 CCCL:cub:…  Push…  264  264         95         95           59,939     630.9     544.0       512     1,024        124.2  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 CCCL:cub:…  Push…  265  265         95         95          120,324   1,266.6   1,056.0     1,024     1,824        268.7  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 CCCL:cub:…  Push…  265  265         95         95           58,849     619.5     544.0       512       896        124.8  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 CCCL:cub:…  Push…  266  266         95         95          117,379   1,235.6   1,056.0       992     1,696        258.5  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 CCCL:cub:…  Push…  266  266         95         95           58,051     611.1     512.0       512       800        120.1  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 CCCL:cub:…  Push…  267  267         95         95          116,835   1,229.8   1,056.0       960     1,696        262.7  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 CCCL:cub:…  Push…  267  267         95         95           58,114     611.7     544.0       480       896        120.5  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 CCCL:cub:…  Push…  268  268         95         95          120,228   1,265.6   1,056.0       992     1,696        273.9  void at_cuda_detail::cub::detail::scan::DeviceScanKernel<at_cuda_detail::cub::detail::scan::policy_…
 CCCL:cub:…  Push…  268  268         95         95           59,074     621.8     544.0       512       992        122.4  void at_cuda_detail::cub::detail::scan::DeviceScanInitKernel<at_cuda_detail::cub::ScanTileState<lon…
 CCCL:cub:…  Push…  261  261          1          1            4,192   4,192.0   4,192.0     4,192     4,192          0.0  void at_cuda_detail::cub::detail::select::DeviceSelectSweepKernel<at_cuda_detail::cub::detail::sele…
 CCCL:cub:…  Push…  261  261          1          1              704     704.0     704.0       704       704          0.0  void at_cuda_detail::cub::detail::scan::DeviceCompactInitKernel<at_cuda_detail::cub::ScanTileState<…
 CCCL:cub:…  Push…  262  262          1          1            3,840   3,840.0   3,840.0     3,840     3,840          0.0  void at_cuda_detail::cub::detail::select::DeviceSelectSweepKernel<at_cuda_detail::cub::detail::sele…
 CCCL:cub:…  Push…  262  262          1          1              768     768.0     768.0       768       768          0.0  void at_cuda_detail::cub::detail::scan::DeviceCompactInitKernel<at_cuda_detail::cub::ScanTileState<…
 CCCL:cub:…  Push…  263  263          1          1            3,712   3,712.0   3,712.0     3,712     3,712          0.0  void at_cuda_detail::cub::detail::select::DeviceSelectSweepKernel<at_cuda_detail::cub::detail::sele…
 CCCL:cub:…  Push…  263  263          1          1              768     768.0     768.0       768       768          0.0  void at_cuda_detail::cub::detail::scan::DeviceCompactInitKernel<at_cuda_detail::cub::ScanTileState<…
 CCCL:cub:…  Push…  264  264          1          1            4,064   4,064.0   4,064.0     4,064     4,064          0.0  void at_cuda_detail::cub::detail::select::DeviceSelectSweepKernel<at_cuda_detail::cub::detail::sele…
 CCCL:cub:…  Push…  264  264          1          1              736     736.0     736.0       736       736          0.0  void at_cuda_detail::cub::detail::scan::DeviceCompactInitKernel<at_cuda_detail::cub::ScanTileState<…
 CCCL:cub:…  Push…  265  265          1          1            3,808   3,808.0   3,808.0     3,808     3,808          0.0  void at_cuda_detail::cub::detail::select::DeviceSelectSweepKernel<at_cuda_detail::cub::detail::sele…
 CCCL:cub:…  Push…  265  265          1          1              736     736.0     736.0       736       736          0.0  void at_cuda_detail::cub::detail::scan::DeviceCompactInitKernel<at_cuda_detail::cub::ScanTileState<…
 CCCL:cub:…  Push…  266  266          1          1            3,904   3,904.0   3,904.0     3,904     3,904          0.0  void at_cuda_detail::cub::detail::select::DeviceSelectSweepKernel<at_cuda_detail::cub::detail::sele…
 CCCL:cub:…  Push…  266  266          1          1              704     704.0     704.0       704       704          0.0  void at_cuda_detail::cub::detail::scan::DeviceCompactInitKernel<at_cuda_detail::cub::ScanTileState<…
 CCCL:cub:…  Push…  267  267          1          1            3,969   3,969.0   3,969.0     3,969     3,969          0.0  void at_cuda_detail::cub::detail::select::DeviceSelectSweepKernel<at_cuda_detail::cub::detail::sele…
 CCCL:cub:…  Push…  267  267          1          1              736     736.0     736.0       736       736          0.0  void at_cuda_detail::cub::detail::scan::DeviceCompactInitKernel<at_cuda_detail::cub::ScanTileState<…
 CCCL:cub:…  Push…  268  268          1          1            4,256   4,256.0   4,256.0     4,256     4,256          0.0  void at_cuda_detail::cub::detail::select::DeviceSelectSweepKernel<at_cuda_detail::cub::detail::sele…
 CCCL:cub:…  Push…  268  268          1          1              736     736.0     736.0       736       736          0.0  void at_cuda_detail::cub::detail::scan::DeviceCompactInitKernel<at_cuda_detail::cub::ScanTileState<…
 NCCL:API …  Star…  261  261         20         20       16,619,589  830,979…  349,562…   216,455  9,565,1…  2,059,162.1  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:API …  Star…  262  262         19         19       18,889,577  994,188…  340,938…   226,502  12,461,…  2,778,620.5  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:API …  Star…  263  263         19         19       17,480,018  920,000…  343,658…   226,983  11,108,…  2,468,905.9  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:API …  Star…  264  264         19         19       16,614,565  874,450…  334,730…   244,231  10,379,…  2,303,450.2  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:API …  Star…  265  265         19         19       16,448,571  865,714…  311,433…   262,247  10,613,…  2,362,380.1  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:API …  Star…  266  266         19         19       14,995,881  789,256…  306,282…   262,472  9,189,2…  2,035,563.2  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:API …  Star…  267  267         19         19       12,042,135  633,796…  304,330…   264,617  6,266,6…  1,365,332.1  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:API …  Star…  268  268         19         19       15,079,283  793,646…  261,608…   224,839  9,994,3…  2,228,833.7  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Grou…  Star…  261  261         20         20       16,619,589  830,979…  349,562…   216,455  9,565,1…  2,059,162.1  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Grou…  Star…  262  262         19         19       18,889,577  994,188…  340,938…   226,502  12,461,…  2,778,620.5  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Grou…  Star…  263  263         19         19       17,480,018  920,000…  343,658…   226,983  11,108,…  2,468,905.9  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Grou…  Star…  264  264         19         19       16,614,565  874,450…  334,730…   244,231  10,379,…  2,303,450.2  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Grou…  Star…  265  265         19         19       16,448,571  865,714…  311,433…   262,247  10,613,…  2,362,380.1  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Grou…  Star…  266  266         19         19       14,995,881  789,256…  306,282…   262,472  9,189,2…  2,035,563.2  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Grou…  Star…  267  267         19         19       12,042,135  633,796…  304,330…   264,617  6,266,6…  1,365,332.1  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Grou…  Star…  268  268         19         19       15,079,283  793,646…  261,608…   224,839  9,994,3…  2,228,833.7  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Kern…  Star…  261  261         20         20       16,619,589  830,979…  349,562…   216,455  9,565,1…  2,059,162.1  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Kern…  Star…  262  262         19         19       18,889,577  994,188…  340,938…   226,502  12,461,…  2,778,620.5  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Kern…  Star…  263  263         19         19       17,480,018  920,000…  343,658…   226,983  11,108,…  2,468,905.9  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Kern…  Star…  264  264         19         19       16,614,565  874,450…  334,730…   244,231  10,379,…  2,303,450.2  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Kern…  Star…  265  265         19         19       16,448,571  865,714…  311,433…   262,247  10,613,…  2,362,380.1  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Kern…  Star…  266  266         19         19       14,995,881  789,256…  306,282…   262,472  9,189,2…  2,035,563.2  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Kern…  Star…  267  267         19         19       12,042,135  633,796…  304,330…   264,617  6,266,6…  1,365,332.1  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             
 NCCL:Kern…  Star…  268  268         19         19       15,079,283  793,646…  261,608…   224,839  9,994,3…  2,228,833.7  ncclDevKernel_AllReduce_Sum_bf16_RING_LL(ncclDevKernelArgsStorage<(unsigned long)4096>)             

