[Artifacts] update B300 matrix through 15:23

This commit is contained in:
Zhiyi Hong 2026-09-10 15:29:00 +08:00
parent 7984c25586
commit 9652bfdb9d
376 changed files with 246149 additions and 1119 deletions

View File

@ -13,6 +13,13 @@ text-matrix run as of the snapshot time.
- Follow-up log: `/data/b300-dsv4-glm53-dev4-20260909-084427-followup.log` - Follow-up log: `/data/b300-dsv4-glm53-dev4-20260909-084427-followup.log`
- Source archive SHA-256: `8676f28c6b9df7676ad93d5a4a71777bb504539b0ea6e3b129ccec69f5a563fd` - Source archive SHA-256: `8676f28c6b9df7676ad93d5a4a71777bb504539b0ea6e3b129ccec69f5a563fd`
Incremental synchronization:
- Delta time: `2026-09-10 07:23:10 UTC` (`2026-09-10 15:23:10 Asia/Shanghai`)
- Delta archive SHA-256: `c1bc269866152024afbdefbb15265832a773bfbf07ea2375e318dc252b25bdbb`
- At this point DeepSeek-V4-Flash had produced 53 Low-Latency and 18
Balanced point JSON files.
The DeepSeek-V4-Flash matrix was still running when this snapshot was taken. The DeepSeek-V4-Flash matrix was still running when this snapshot was taken.
Consequently, this is a complete snapshot of files present at that time, not Consequently, this is a complete snapshot of files present at that time, not
the final completed run archive. GLM-5.3 had completed its first pass. the final completed run archive. GLM-5.3 had completed its first pass.

View File

@ -1,9 +1,9 @@
index, name, memory.total [MiB], memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W] index, name, memory.total [MiB], memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 183.73 W 0, NVIDIA Graphics Device, 275040 MiB, 259411 MiB, 14703 MiB, 0 %, 242.62 W
1, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 181.64 W 1, NVIDIA Graphics Device, 275040 MiB, 259603 MiB, 14511 MiB, 0 %, 234.74 W
2, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 184.33 W 2, NVIDIA Graphics Device, 275040 MiB, 259617 MiB, 14497 MiB, 0 %, 235.85 W
3, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 186.83 W 3, NVIDIA Graphics Device, 275040 MiB, 258643 MiB, 15471 MiB, 0 %, 243.30 W
4, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 184.06 W 4, NVIDIA Graphics Device, 275040 MiB, 4 MiB, 274110 MiB, 0 %, 183.98 W
5, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 182.13 W 5, NVIDIA Graphics Device, 275040 MiB, 4 MiB, 274110 MiB, 0 %, 182.88 W
6, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 183.42 W 6, NVIDIA Graphics Device, 275040 MiB, 4 MiB, 274110 MiB, 0 %, 184.01 W
7, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 185.93 W 7, NVIDIA Graphics Device, 275040 MiB, 4 MiB, 274110 MiB, 0 %, 184.09 W

1 index name memory.total [MiB] memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 NVIDIA Graphics Device 275040 MiB 0 MiB 259411 MiB 274114 MiB 14703 MiB 0 % 183.73 W 242.62 W
3 1 NVIDIA Graphics Device 275040 MiB 0 MiB 259603 MiB 274114 MiB 14511 MiB 0 % 181.64 W 234.74 W
4 2 NVIDIA Graphics Device 275040 MiB 0 MiB 259617 MiB 274114 MiB 14497 MiB 0 % 184.33 W 235.85 W
5 3 NVIDIA Graphics Device 275040 MiB 0 MiB 258643 MiB 274114 MiB 15471 MiB 0 % 186.83 W 243.30 W
6 4 NVIDIA Graphics Device 275040 MiB 0 MiB 4 MiB 274114 MiB 274110 MiB 0 % 184.06 W 183.98 W
7 5 NVIDIA Graphics Device 275040 MiB 0 MiB 4 MiB 274114 MiB 274110 MiB 0 % 182.13 W 182.88 W
8 6 NVIDIA Graphics Device 275040 MiB 0 MiB 4 MiB 274114 MiB 274110 MiB 0 % 183.42 W 184.01 W
9 7 NVIDIA Graphics Device 275040 MiB 0 MiB 4 MiB 274114 MiB 274110 MiB 0 % 185.93 W 184.09 W

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 261419 MiB, 12695 MiB, 0 %, 241.89 W
1, 261033 MiB, 13081 MiB, 0 %, 234.70 W
2, 261481 MiB, 12633 MiB, 0 %, 235.81 W
3, 261633 MiB, 12481 MiB, 0 %, 243.99 W
4, 4 MiB, 274110 MiB, 0 %, 182.40 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.84 W
7, 4 MiB, 274110 MiB, 0 %, 183.78 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 261419 MiB 12695 MiB 0 % 241.89 W
3 1 261033 MiB 13081 MiB 0 % 234.70 W
4 2 261481 MiB 12633 MiB 0 % 235.81 W
5 3 261633 MiB 12481 MiB 0 % 243.99 W
6 4 4 MiB 274110 MiB 0 % 182.40 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.84 W
9 7 4 MiB 274110 MiB 0 % 183.78 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=1, output_file='/results/points/16k_1_c1.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=1, output_file='/results/points/16k_1_c1.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 1048576
#Output tokens: 64
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 1
Successful requests: 64
Benchmark duration (s): 62.77
Total input tokens: 1048576
Total input text tokens: 1048576
Total generated tokens: 64
Total generated tokens (retokenized): 64
Request throughput (req/s): 1.02
Input token throughput (tok/s): 16704.37
Output token throughput (tok/s): 1.02
Peak output token throughput (tok/s): 2.00
Peak concurrent requests: 3
Total token throughput (tok/s): 16705.38
Concurrency: 1.00
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 979.20
Median E2E Latency (ms): 978.74
P90 E2E Latency (ms): 998.89
P95 E2E Latency (ms): 1001.70
P99 E2E Latency (ms): 1041.31
---------------Time to First Token----------------
Mean TTFT (ms): 979.15
Median TTFT (ms): 978.69
P90 TTFT (ms): 998.85
P95 TTFT (ms): 1001.65
P99 TTFT (ms): 1041.27
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 0.00
Median TPOT (ms): 0.00
P90 TPOT (ms): 0.00
P95 TPOT (ms): 0.00
P99 TPOT (ms): 0.00
---------------Inter-Token Latency----------------
Mean ITL (ms): 0.00
Median ITL (ms): 0.00
P90 ITL (ms): 0.00
P95 ITL (ms): 0.00
P99 ITL (ms): 0.00
Max ITL (ms): 0.00
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 270303 MiB, 3811 MiB, 0 %, 245.71 W
1, 267511 MiB, 6603 MiB, 0 %, 236.66 W
2, 271933 MiB, 2181 MiB, 0 %, 238.35 W
3, 265137 MiB, 8977 MiB, 0 %, 247.74 W
4, 4 MiB, 274110 MiB, 0 %, 182.01 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.17 W
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 270303 MiB 3811 MiB 0 % 245.71 W
3 1 267511 MiB 6603 MiB 0 % 236.66 W
4 2 271933 MiB 2181 MiB 0 % 238.35 W
5 3 265137 MiB 8977 MiB 0 % 247.74 W
6 4 4 MiB 274110 MiB 0 % 182.01 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.17 W
9 7 4 MiB 274110 MiB 0 % 182.01 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=640, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=128, output_file='/results/points/16k_1_c128.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=640, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=128, output_file='/results/points/16k_1_c128.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 10485760
#Output tokens: 640
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 128
Successful requests: 640
Benchmark duration (s): 161.63
Total input tokens: 10485760
Total input text tokens: 10485760
Total generated tokens: 640
Total generated tokens (retokenized): 636
Request throughput (req/s): 3.96
Input token throughput (tok/s): 64875.87
Output token throughput (tok/s): 3.96
Peak output token throughput (tok/s): 6.00
Peak concurrent requests: 133
Total token throughput (tok/s): 64879.83
Concurrency: 115.42
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 29149.59
Median E2E Latency (ms): 31991.66
P90 E2E Latency (ms): 32464.82
P95 E2E Latency (ms): 32515.02
P99 E2E Latency (ms): 32526.70
---------------Time to First Token----------------
Mean TTFT (ms): 28973.22
Median TTFT (ms): 31988.94
P90 TTFT (ms): 32464.77
P95 TTFT (ms): 32514.97
P99 TTFT (ms): 32526.66
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 0.00
Median TPOT (ms): 0.00
P90 TPOT (ms): 0.00
P95 TPOT (ms): 0.00
P99 TPOT (ms): 0.00
---------------Inter-Token Latency----------------
Mean ITL (ms): 0.00
Median ITL (ms): 0.00
P90 ITL (ms): 0.00
P95 ITL (ms): 0.00
P99 ITL (ms): 0.00
Max ITL (ms): 0.00
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 269385 MiB, 4729 MiB, 0 %, 244.37 W
1, 267211 MiB, 6903 MiB, 0 %, 236.66 W
2, 271703 MiB, 2411 MiB, 0 %, 237.25 W
3, 265765 MiB, 8349 MiB, 0 %, 247.43 W
4, 4 MiB, 274110 MiB, 0 %, 182.01 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.25 W
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 269385 MiB 4729 MiB 0 % 244.37 W
3 1 267211 MiB 6903 MiB 0 % 236.66 W
4 2 271703 MiB 2411 MiB 0 % 237.25 W
5 3 265765 MiB 8349 MiB 0 % 247.43 W
6 4 4 MiB 274110 MiB 0 % 182.01 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.25 W
9 7 4 MiB 274110 MiB 0 % 182.01 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=1280, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=256, output_file='/results/points/16k_1_c256.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=1280, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=256, output_file='/results/points/16k_1_c256.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 20971520
#Output tokens: 1280
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 256
Successful requests: 1280
Benchmark duration (s): 326.43
Total input tokens: 20971520
Total input text tokens: 20971520
Total generated tokens: 1280
Total generated tokens (retokenized): 1274
Request throughput (req/s): 3.92
Input token throughput (tok/s): 64245.92
Output token throughput (tok/s): 3.92
Peak output token throughput (tok/s): 5.00
Peak concurrent requests: 260
Total token throughput (tok/s): 64249.84
Concurrency: 230.13
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 58687.97
Median E2E Latency (ms): 64913.32
P90 E2E Latency (ms): 65018.57
P95 E2E Latency (ms): 65033.21
P99 E2E Latency (ms): 65149.17
---------------Time to First Token----------------
Mean TTFT (ms): 58398.08
Median TTFT (ms): 64913.17
P90 TTFT (ms): 65018.54
P95 TTFT (ms): 65033.17
P99 TTFT (ms): 65149.12
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 0.00
Median TPOT (ms): 0.00
P90 TPOT (ms): 0.00
P95 TPOT (ms): 0.00
P99 TPOT (ms): 0.00
---------------Inter-Token Latency----------------
Mean ITL (ms): 0.00
Median ITL (ms): 0.00
P90 ITL (ms): 0.00
P95 ITL (ms): 0.00
P99 ITL (ms): 0.00
Max ITL (ms): 0.00
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 269243 MiB, 4871 MiB, 0 %, 243.84 W
1, 264807 MiB, 9307 MiB, 0 %, 236.54 W
2, 268279 MiB, 5835 MiB, 0 %, 236.92 W
3, 266383 MiB, 7731 MiB, 0 %, 247.24 W
4, 4 MiB, 274110 MiB, 0 %, 182.13 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.13 W
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 269243 MiB 4871 MiB 0 % 243.84 W
3 1 264807 MiB 9307 MiB 0 % 236.54 W
4 2 268279 MiB 5835 MiB 0 % 236.92 W
5 3 266383 MiB 7731 MiB 0 % 247.24 W
6 4 4 MiB 274110 MiB 0 % 182.13 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.13 W
9 7 4 MiB 274110 MiB 0 % 182.01 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=160, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=32, output_file='/results/points/16k_1_c32.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=160, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=32, output_file='/results/points/16k_1_c32.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 2621440
#Output tokens: 160
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 32
Successful requests: 160
Benchmark duration (s): 40.27
Total input tokens: 2621440
Total input text tokens: 2621440
Total generated tokens: 160
Total generated tokens (retokenized): 158
Request throughput (req/s): 3.97
Input token throughput (tok/s): 65093.70
Output token throughput (tok/s): 3.97
Peak output token throughput (tok/s): 5.00
Peak concurrent requests: 36
Total token throughput (tok/s): 65097.67
Concurrency: 29.23
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 7356.74
Median E2E Latency (ms): 7919.05
P90 E2E Latency (ms): 7940.64
P95 E2E Latency (ms): 7943.24
P99 E2E Latency (ms): 8597.74
---------------Time to First Token----------------
Mean TTFT (ms): 7257.63
Median TTFT (ms): 7918.42
P90 TTFT (ms): 7940.60
P95 TTFT (ms): 7943.21
P99 TTFT (ms): 8597.70
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 0.00
Median TPOT (ms): 0.00
P90 TPOT (ms): 0.00
P95 TPOT (ms): 0.00
P99 TPOT (ms): 0.00
---------------Inter-Token Latency----------------
Mean ITL (ms): 0.00
Median ITL (ms): 0.00
P90 ITL (ms): 0.00
P95 ITL (ms): 0.00
P99 ITL (ms): 0.00
Max ITL (ms): 0.00
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 270311 MiB, 3803 MiB, 0 %, 245.60 W
1, 266159 MiB, 7955 MiB, 0 %, 236.66 W
2, 267307 MiB, 6807 MiB, 0 %, 238.36 W
3, 264547 MiB, 9567 MiB, 0 %, 247.74 W
4, 4 MiB, 274110 MiB, 0 %, 182.08 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.37 W
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 270311 MiB 3803 MiB 0 % 245.60 W
3 1 266159 MiB 7955 MiB 0 % 236.66 W
4 2 267307 MiB 6807 MiB 0 % 238.36 W
5 3 264547 MiB 9567 MiB 0 % 247.74 W
6 4 4 MiB 274110 MiB 0 % 182.08 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.37 W
9 7 4 MiB 274110 MiB 0 % 182.01 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=320, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=64, output_file='/results/points/16k_1_c64.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=320, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=64, output_file='/results/points/16k_1_c64.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 5242880
#Output tokens: 320
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 64
Successful requests: 320
Benchmark duration (s): 80.80
Total input tokens: 5242880
Total input text tokens: 5242880
Total generated tokens: 320
Total generated tokens (retokenized): 318
Request throughput (req/s): 3.96
Input token throughput (tok/s): 64887.20
Output token throughput (tok/s): 3.96
Peak output token throughput (tok/s): 6.00
Peak concurrent requests: 68
Total token throughput (tok/s): 64891.16
Concurrency: 58.04
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 14654.14
Median E2E Latency (ms): 15916.11
P90 E2E Latency (ms): 16382.05
P95 E2E Latency (ms): 16398.86
P99 E2E Latency (ms): 16407.45
---------------Time to First Token----------------
Mean TTFT (ms): 14554.54
Median TTFT (ms): 15915.15
P90 TTFT (ms): 16382.00
P95 TTFT (ms): 16398.83
P99 TTFT (ms): 16407.41
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 0.00
Median TPOT (ms): 0.00
P90 TPOT (ms): 0.00
P95 TPOT (ms): 0.00
P99 TPOT (ms): 0.00
---------------Inter-Token Latency----------------
Mean ITL (ms): 0.00
Median ITL (ms): 0.00
P90 ITL (ms): 0.00
P95 ITL (ms): 0.00
P99 ITL (ms): 0.00
Max ITL (ms): 0.00
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 269233 MiB, 4881 MiB, 0 %, 243.84 W
1, 263439 MiB, 10675 MiB, 0 %, 236.01 W
2, 268945 MiB, 5169 MiB, 0 %, 237.17 W
3, 265381 MiB, 8733 MiB, 0 %, 245.79 W
4, 4 MiB, 274110 MiB, 0 %, 182.52 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.99 W
7, 4 MiB, 274110 MiB, 0 %, 183.85 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 269233 MiB 4881 MiB 0 % 243.84 W
3 1 263439 MiB 10675 MiB 0 % 236.01 W
4 2 268945 MiB 5169 MiB 0 % 237.17 W
5 3 265381 MiB 8733 MiB 0 % 245.79 W
6 4 4 MiB 274110 MiB 0 % 182.52 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.99 W
9 7 4 MiB 274110 MiB 0 % 183.85 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=8, output_file='/results/points/16k_1_c8.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=8, output_file='/results/points/16k_1_c8.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 1048576
#Output tokens: 64
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 8
Successful requests: 64
Benchmark duration (s): 16.41
Total input tokens: 1048576
Total input text tokens: 1048576
Total generated tokens: 64
Total generated tokens (retokenized): 64
Request throughput (req/s): 3.90
Input token throughput (tok/s): 63912.26
Output token throughput (tok/s): 3.90
Peak output token throughput (tok/s): 4.00
Peak concurrent requests: 12
Total token throughput (tok/s): 63916.16
Concurrency: 7.76
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 1988.23
Median E2E Latency (ms): 1986.87
P90 E2E Latency (ms): 1993.74
P95 E2E Latency (ms): 2334.31
P99 E2E Latency (ms): 2638.62
---------------Time to First Token----------------
Mean TTFT (ms): 1988.19
Median TTFT (ms): 1986.83
P90 TTFT (ms): 1993.70
P95 TTFT (ms): 2334.27
P99 TTFT (ms): 2638.58
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 0.00
Median TPOT (ms): 0.00
P90 TPOT (ms): 0.00
P95 TPOT (ms): 0.00
P99 TPOT (ms): 0.00
---------------Inter-Token Latency----------------
Mean ITL (ms): 0.00
Median ITL (ms): 0.00
P90 ITL (ms): 0.00
P95 ITL (ms): 0.00
P99 ITL (ms): 0.00
Max ITL (ms): 0.00
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 261315 MiB, 12799 MiB, 0 %, 240.90 W
1, 260915 MiB, 13199 MiB, 0 %, 234.32 W
2, 261347 MiB, 12767 MiB, 0 %, 234.96 W
3, 261503 MiB, 12611 MiB, 0 %, 243.99 W
4, 4 MiB, 274110 MiB, 0 %, 183.47 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.22 W
7, 4 MiB, 274110 MiB, 0 %, 183.97 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 261315 MiB 12799 MiB 0 % 240.90 W
3 1 260915 MiB 13199 MiB 0 % 234.32 W
4 2 261347 MiB 12767 MiB 0 % 234.96 W
5 3 261503 MiB 12611 MiB 0 % 243.99 W
6 4 4 MiB 274110 MiB 0 % 183.47 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.22 W
9 7 4 MiB 274110 MiB 0 % 183.97 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=1, output_file='/results/points/16k_512_c1.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=1, output_file='/results/points/16k_512_c1.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 1048576
#Output tokens: 32768
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 1
Successful requests: 64
Benchmark duration (s): 416.51
Total input tokens: 1048576
Total input text tokens: 1048576
Total generated tokens: 32768
Total generated tokens (retokenized): 32447
Request throughput (req/s): 0.15
Input token throughput (tok/s): 2517.53
Output token throughput (tok/s): 78.67
Peak output token throughput (tok/s): 108.00
Peak concurrent requests: 2
Total token throughput (tok/s): 2596.20
Concurrency: 1.00
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 6506.17
Median E2E Latency (ms): 6469.01
P90 E2E Latency (ms): 6740.62
P95 E2E Latency (ms): 6816.94
P99 E2E Latency (ms): 6875.99
---------------Time to First Token----------------
Mean TTFT (ms): 1044.49
Median TTFT (ms): 1043.08
P90 TTFT (ms): 1065.97
P95 TTFT (ms): 1067.18
P99 TTFT (ms): 1078.33
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 10.69
Median TPOT (ms): 10.60
P90 TPOT (ms): 11.18
P95 TPOT (ms): 11.26
P99 TPOT (ms): 11.38
---------------Inter-Token Latency----------------
Mean ITL (ms): 10.69
Median ITL (ms): 10.82
P90 ITL (ms): 11.59
P95 ITL (ms): 11.82
P99 ITL (ms): 12.29
Max ITL (ms): 16.37
==================================================

View File

@ -0,0 +1,817 @@
[2026-09-10 05:25:31 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.78
[2026-09-10 05:25:31 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17796.50
[2026-09-10 05:25:31 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17815.89
[2026-09-10 05:25:31 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 376057.58
[2026-09-10 05:25:31 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
[2026-09-10 05:25:32 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.03, #queue-req: 0
[2026-09-10 05:25:32 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.86, #queue-req: 0
[2026-09-10 05:25:33 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.73, #queue-req: 0
[2026-09-10 05:25:33 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.06, #queue-req: 0
[2026-09-10 05:25:34 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.13, #queue-req: 0
[2026-09-10 05:25:34 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.50, #queue-req: 0
[2026-09-10 05:25:35 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.06, #queue-req: 0
[2026-09-10 05:25:35 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.82, #queue-req: 0
[2026-09-10 05:25:35 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.64, #queue-req: 0
[2026-09-10 05:25:36 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.21, #queue-req: 0
[2026-09-10 05:25:36 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.39, #queue-req: 0
[2026-09-10 05:25:37 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.66, #queue-req: 0
[2026-09-10 05:25:37] INFO: 127.0.0.1:44620 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:25:37 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.63
[2026-09-10 05:25:38 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17725.79
[2026-09-10 05:25:38 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17725.44
[2026-09-10 05:25:38 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 394303.47
[2026-09-10 05:25:38 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
[2026-09-10 05:25:38 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.84, #queue-req: 0
[2026-09-10 05:25:39 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 103.15, #queue-req: 0
[2026-09-10 05:25:39 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.24, #queue-req: 0
[2026-09-10 05:25:40 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.48, #queue-req: 0
[2026-09-10 05:25:40 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.04, #queue-req: 0
[2026-09-10 05:25:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.90, #queue-req: 0
[2026-09-10 05:25:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.07, #queue-req: 0
[2026-09-10 05:25:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.71, #queue-req: 0
[2026-09-10 05:25:42 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.81, #queue-req: 0
[2026-09-10 05:25:42 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.67, #queue-req: 0
[2026-09-10 05:25:43 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.06, #queue-req: 0
[2026-09-10 05:25:43 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.96, #queue-req: 0
[2026-09-10 05:25:43] INFO: 127.0.0.1:51576 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:25:44 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.43
[2026-09-10 05:25:44 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 16995.45
[2026-09-10 05:25:44 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 16964.93
[2026-09-10 05:25:44 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 379714.50
[2026-09-10 05:25:44 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
[2026-09-10 05:25:45 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 102.63, #queue-req: 0
[2026-09-10 05:25:45 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.93, #queue-req: 0
[2026-09-10 05:25:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.32, #queue-req: 0
[2026-09-10 05:25:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.12, #queue-req: 0
[2026-09-10 05:25:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.60, #queue-req: 0
[2026-09-10 05:25:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.87, #queue-req: 0
[2026-09-10 05:25:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.95, #queue-req: 0
[2026-09-10 05:25:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.35, #queue-req: 0
[2026-09-10 05:25:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.68, #queue-req: 0
[2026-09-10 05:25:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.43, #queue-req: 0
[2026-09-10 05:25:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.30, #queue-req: 0
[2026-09-10 05:25:50] INFO: 127.0.0.1:59534 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:25:50 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.33
[2026-09-10 05:25:50 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17384.11
[2026-09-10 05:25:51 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17358.92
[2026-09-10 05:25:51 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 393361.18
[2026-09-10 05:25:51 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.92, #queue-req: 0
[2026-09-10 05:25:51 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 104.68, #queue-req: 0
[2026-09-10 05:25:52 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.05, #queue-req: 0
[2026-09-10 05:25:52 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.47, #queue-req: 0
[2026-09-10 05:25:53 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.94, #queue-req: 0
[2026-09-10 05:25:53 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.10, #queue-req: 0
[2026-09-10 05:25:53 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.51, #queue-req: 0
[2026-09-10 05:25:54 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.33, #queue-req: 0
[2026-09-10 05:25:54 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.94, #queue-req: 0
[2026-09-10 05:25:55 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.56, #queue-req: 0
[2026-09-10 05:25:55 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.49, #queue-req: 0
[2026-09-10 05:25:56 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.70, #queue-req: 0
[2026-09-10 05:25:56] INFO: 127.0.0.1:59544 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:25:56 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.89
[2026-09-10 05:25:57 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17506.33
[2026-09-10 05:25:57 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17698.03
[2026-09-10 05:25:57 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 392426.55
[2026-09-10 05:25:57 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.95, #queue-req: 0
[2026-09-10 05:25:58 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.34, #queue-req: 0
[2026-09-10 05:25:58 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.02, #queue-req: 0
[2026-09-10 05:25:59 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.99, #queue-req: 0
[2026-09-10 05:25:59 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.44, #queue-req: 0
[2026-09-10 05:26:00 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.45, #queue-req: 0
[2026-09-10 05:26:00 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.92, #queue-req: 0
[2026-09-10 05:26:00 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.80, #queue-req: 0
[2026-09-10 05:26:01 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.58, #queue-req: 0
[2026-09-10 05:26:01 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.71, #queue-req: 0
[2026-09-10 05:26:02 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.90, #queue-req: 0
[2026-09-10 05:26:02 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.25, #queue-req: 0
[2026-09-10 05:26:03] INFO: 127.0.0.1:36546 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:26:03 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.48
[2026-09-10 05:26:03 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17221.91
[2026-09-10 05:26:04 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17618.48
[2026-09-10 05:26:04 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 410228.80
[2026-09-10 05:26:04 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
[2026-09-10 05:26:04 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 108.85, #queue-req: 0
[2026-09-10 05:26:05 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.93, #queue-req: 0
[2026-09-10 05:26:05 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.49, #queue-req: 0
[2026-09-10 05:26:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.88, #queue-req: 0
[2026-09-10 05:26:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.41, #queue-req: 0
[2026-09-10 05:26:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.71, #queue-req: 0
[2026-09-10 05:26:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.55, #queue-req: 0
[2026-09-10 05:26:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.97, #queue-req: 0
[2026-09-10 05:26:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.38, #queue-req: 0
[2026-09-10 05:26:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.10, #queue-req: 0
[2026-09-10 05:26:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.10, #queue-req: 0
[2026-09-10 05:26:09] INFO: 127.0.0.1:47830 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:26:10 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.00
[2026-09-10 05:26:10 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 16721.27
[2026-09-10 05:26:10 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 16889.70
[2026-09-10 05:26:10 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 368892.56
[2026-09-10 05:26:10 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
[2026-09-10 05:26:10 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.43, #queue-req: 0
[2026-09-10 05:26:11 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.04, #queue-req: 0
[2026-09-10 05:26:11 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.33, #queue-req: 0
[2026-09-10 05:26:12 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.55, #queue-req: 0
[2026-09-10 05:26:12 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.56, #queue-req: 0
[2026-09-10 05:26:13 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.67, #queue-req: 0
[2026-09-10 05:26:13 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.06, #queue-req: 0
[2026-09-10 05:26:14 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.12, #queue-req: 0
[2026-09-10 05:26:14 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.92, #queue-req: 0
[2026-09-10 05:26:14 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.26, #queue-req: 0
[2026-09-10 05:26:15 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.30, #queue-req: 0
[2026-09-10 05:26:15 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.94, #queue-req: 0
[2026-09-10 05:26:16] INFO: 127.0.0.1:47846 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:26:16 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.81
[2026-09-10 05:26:16 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17033.75
[2026-09-10 05:26:17 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17310.55
[2026-09-10 05:26:17 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 384089.58
[2026-09-10 05:26:17 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
[2026-09-10 05:26:17 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 112.77, #queue-req: 0
[2026-09-10 05:26:17 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 100.14, #queue-req: 0
[2026-09-10 05:26:18 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.95, #queue-req: 0
[2026-09-10 05:26:18 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.95, #queue-req: 0
[2026-09-10 05:26:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.64, #queue-req: 0
[2026-09-10 05:26:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.10, #queue-req: 0
[2026-09-10 05:26:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.87, #queue-req: 0
[2026-09-10 05:26:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.80, #queue-req: 0
[2026-09-10 05:26:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.44, #queue-req: 0
[2026-09-10 05:26:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.71, #queue-req: 0
[2026-09-10 05:26:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.03, #queue-req: 0
[2026-09-10 05:26:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.32, #queue-req: 0
[2026-09-10 05:26:22] INFO: 127.0.0.1:40652 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:26:22 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.46
[2026-09-10 05:26:23 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17864.62
[2026-09-10 05:26:23 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17840.13
[2026-09-10 05:26:23 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 417235.83
[2026-09-10 05:26:23 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
[2026-09-10 05:26:23 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 110.66, #queue-req: 0
[2026-09-10 05:26:24 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.17, #queue-req: 0
[2026-09-10 05:26:24 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.25, #queue-req: 0
[2026-09-10 05:26:25 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.47, #queue-req: 0
[2026-09-10 05:26:25 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.30, #queue-req: 0
[2026-09-10 05:26:26 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.61, #queue-req: 0
[2026-09-10 05:26:26 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.22, #queue-req: 0
[2026-09-10 05:26:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.35, #queue-req: 0
[2026-09-10 05:26:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.51, #queue-req: 0
[2026-09-10 05:26:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.82, #queue-req: 0
[2026-09-10 05:26:28 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.67, #queue-req: 0
[2026-09-10 05:26:28 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.28, #queue-req: 0
[2026-09-10 05:26:29] INFO: 127.0.0.1:44996 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:26:29 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.77
[2026-09-10 05:26:29 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17803.49
[2026-09-10 05:26:30 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17766.36
[2026-09-10 05:26:30 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 412728.85
[2026-09-10 05:26:30 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
[2026-09-10 05:26:30 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.84, #queue-req: 0
[2026-09-10 05:26:30 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 109.76, #queue-req: 0
[2026-09-10 05:26:31 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 100.04, #queue-req: 0
[2026-09-10 05:26:31 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.17, #queue-req: 0
[2026-09-10 05:26:32 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.54, #queue-req: 0
[2026-09-10 05:26:32 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.49, #queue-req: 0
[2026-09-10 05:26:32 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.87, #queue-req: 0
[2026-09-10 05:26:33 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.58, #queue-req: 0
[2026-09-10 05:26:33 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.82, #queue-req: 0
[2026-09-10 05:26:34 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.60, #queue-req: 0
[2026-09-10 05:26:34 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.12, #queue-req: 0
[2026-09-10 05:26:35 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.73, #queue-req: 0
[2026-09-10 05:26:35] INFO: 127.0.0.1:45008 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:26:35 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.63
[2026-09-10 05:26:36 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17430.73
[2026-09-10 05:26:36 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17464.24
[2026-09-10 05:26:36 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 397323.61
[2026-09-10 05:26:36 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
[2026-09-10 05:26:36 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 114.44, #queue-req: 0
[2026-09-10 05:26:37 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.90, #queue-req: 0
[2026-09-10 05:26:37 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 100.25, #queue-req: 0
[2026-09-10 05:26:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.05, #queue-req: 0
[2026-09-10 05:26:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.03, #queue-req: 0
[2026-09-10 05:26:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.31, #queue-req: 0
[2026-09-10 05:26:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.02, #queue-req: 0
[2026-09-10 05:26:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.51, #queue-req: 0
[2026-09-10 05:26:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.85, #queue-req: 0
[2026-09-10 05:26:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.22, #queue-req: 0
[2026-09-10 05:26:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.41, #queue-req: 0
[2026-09-10 05:26:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.27, #queue-req: 0
[2026-09-10 05:26:41] INFO: 127.0.0.1:41330 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:26:42 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 162.52
[2026-09-10 05:26:42 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17395.21
[2026-09-10 05:26:42 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17358.20
[2026-09-10 05:26:42 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 392883.77
[2026-09-10 05:26:42 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
[2026-09-10 05:26:43 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 112.70, #queue-req: 0
[2026-09-10 05:26:43 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.53, #queue-req: 0
[2026-09-10 05:26:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.32, #queue-req: 0
[2026-09-10 05:26:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.61, #queue-req: 0
[2026-09-10 05:26:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.51, #queue-req: 0
[2026-09-10 05:26:45 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.15, #queue-req: 0
[2026-09-10 05:26:45 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.52, #queue-req: 0
[2026-09-10 05:26:46 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.59, #queue-req: 0
[2026-09-10 05:26:46 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.05, #queue-req: 0
[2026-09-10 05:26:47 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.99, #queue-req: 0
[2026-09-10 05:26:47 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.42, #queue-req: 0
[2026-09-10 05:26:47 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.30, #queue-req: 0
[2026-09-10 05:26:48] INFO: 127.0.0.1:41346 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:26:48 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 162.45
[2026-09-10 05:26:48 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 16997.55
[2026-09-10 05:26:49 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17005.57
[2026-09-10 05:26:49 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 412761.43
[2026-09-10 05:26:49 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.96, #queue-req: 0
[2026-09-10 05:26:49 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 106.99, #queue-req: 0
[2026-09-10 05:26:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.87, #queue-req: 0
[2026-09-10 05:26:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.10, #queue-req: 0
[2026-09-10 05:26:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.71, #queue-req: 0
[2026-09-10 05:26:51 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.29, #queue-req: 0
[2026-09-10 05:26:51 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.48, #queue-req: 0
[2026-09-10 05:26:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.68, #queue-req: 0
[2026-09-10 05:26:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.94, #queue-req: 0
[2026-09-10 05:26:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.10, #queue-req: 0
[2026-09-10 05:26:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.83, #queue-req: 0
[2026-09-10 05:26:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.42, #queue-req: 0
[2026-09-10 05:26:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.96, #queue-req: 0
[2026-09-10 05:26:54] INFO: 127.0.0.1:36030 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:26:55 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 162.39
[2026-09-10 05:26:55 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17713.01
[2026-09-10 05:26:55 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17792.55
[2026-09-10 05:26:55 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 421200.80
[2026-09-10 05:26:55 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.92, #queue-req: 0
[2026-09-10 05:26:56 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.94, #queue-req: 0
[2026-09-10 05:26:56 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 110.85, #queue-req: 0
[2026-09-10 05:26:57 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 101.09, #queue-req: 0
[2026-09-10 05:26:57 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.24, #queue-req: 0
[2026-09-10 05:26:57 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.28, #queue-req: 0
[2026-09-10 05:26:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.14, #queue-req: 0
[2026-09-10 05:26:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.69, #queue-req: 0
[2026-09-10 05:26:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.70, #queue-req: 0
[2026-09-10 05:26:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.50, #queue-req: 0
[2026-09-10 05:27:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.11, #queue-req: 0
[2026-09-10 05:27:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.22, #queue-req: 0
[2026-09-10 05:27:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.99, #queue-req: 0
[2026-09-10 05:27:01] INFO: 127.0.0.1:48624 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:27:01 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 162.38
[2026-09-10 05:27:01 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17377.25
[2026-09-10 05:27:02 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17445.25
[2026-09-10 05:27:02 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 369621.82
[2026-09-10 05:27:02 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
[2026-09-10 05:27:02 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 103.91, #queue-req: 0
[2026-09-10 05:27:03 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.59, #queue-req: 0
[2026-09-10 05:27:03 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.65, #queue-req: 0
[2026-09-10 05:27:04 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.30, #queue-req: 0
[2026-09-10 05:27:04 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.02, #queue-req: 0
[2026-09-10 05:27:04 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.05, #queue-req: 0
[2026-09-10 05:27:05 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.68, #queue-req: 0
[2026-09-10 05:27:05 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.26, #queue-req: 0
[2026-09-10 05:27:06 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.65, #queue-req: 0
[2026-09-10 05:27:06 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.81, #queue-req: 0
[2026-09-10 05:27:07 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.06, #queue-req: 0
[2026-09-10 05:27:07 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.70, #queue-req: 0
[2026-09-10 05:27:07] INFO: 127.0.0.1:48628 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:27:08 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.17
[2026-09-10 05:27:08 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17333.87
[2026-09-10 05:27:08 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17041.99
[2026-09-10 05:27:08 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 379084.60
[2026-09-10 05:27:08 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
[2026-09-10 05:27:09 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 107.04, #queue-req: 0
[2026-09-10 05:27:09 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.07, #queue-req: 0
[2026-09-10 05:27:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.71, #queue-req: 0
[2026-09-10 05:27:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.69, #queue-req: 0
[2026-09-10 05:27:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.16, #queue-req: 0
[2026-09-10 05:27:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.55, #queue-req: 0
[2026-09-10 05:27:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.45, #queue-req: 0
[2026-09-10 05:27:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.56, #queue-req: 0
[2026-09-10 05:27:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.48, #queue-req: 0
[2026-09-10 05:27:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.22, #queue-req: 0
[2026-09-10 05:27:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.72, #queue-req: 0
[2026-09-10 05:27:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.86, #queue-req: 0
[2026-09-10 05:27:14] INFO: 127.0.0.1:37054 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:27:14 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.58
[2026-09-10 05:27:14 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17402.19
[2026-09-10 05:27:15 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17381.88
[2026-09-10 05:27:15 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 409661.08
[2026-09-10 05:27:15 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
[2026-09-10 05:27:15 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 101.85, #queue-req: 0
[2026-09-10 05:27:16 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.27, #queue-req: 0
[2026-09-10 05:27:16 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.18, #queue-req: 0
[2026-09-10 05:27:17 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.41, #queue-req: 0
[2026-09-10 05:27:17 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.62, #queue-req: 0
[2026-09-10 05:27:17 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.17, #queue-req: 0
[2026-09-10 05:27:18 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.24, #queue-req: 0
[2026-09-10 05:27:18 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.01, #queue-req: 0
[2026-09-10 05:27:19 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.40, #queue-req: 0
[2026-09-10 05:27:19 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.47, #queue-req: 0
[2026-09-10 05:27:20 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.46, #queue-req: 0
[2026-09-10 05:27:20 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.96, #queue-req: 0
[2026-09-10 05:27:20] INFO: 127.0.0.1:35934 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:27:21 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.40
[2026-09-10 05:27:21 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17616.27
[2026-09-10 05:27:21 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17663.97
[2026-09-10 05:27:21 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 418037.71
[2026-09-10 05:27:22 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
[2026-09-10 05:27:22 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 112.73, #queue-req: 0
[2026-09-10 05:27:22 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 101.37, #queue-req: 0
[2026-09-10 05:27:23 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.16, #queue-req: 0
[2026-09-10 05:27:23 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.66, #queue-req: 0
[2026-09-10 05:27:24 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.14, #queue-req: 0
[2026-09-10 05:27:24 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.76, #queue-req: 0
[2026-09-10 05:27:24 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.94, #queue-req: 0
[2026-09-10 05:27:25 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.04, #queue-req: 0
[2026-09-10 05:27:25 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.23, #queue-req: 0
[2026-09-10 05:27:26 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.62, #queue-req: 0
[2026-09-10 05:27:26 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.45, #queue-req: 0
[2026-09-10 05:27:27 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.76, #queue-req: 0
[2026-09-10 05:27:27] INFO: 127.0.0.1:35946 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:27:27 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.78
[2026-09-10 05:27:27 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17369.58
[2026-09-10 05:27:28 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17362.70
[2026-09-10 05:27:28 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 380751.03
[2026-09-10 05:27:28 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
[2026-09-10 05:27:28 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 105.20, #queue-req: 0
[2026-09-10 05:27:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.34, #queue-req: 0
[2026-09-10 05:27:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.37, #queue-req: 0
[2026-09-10 05:27:30 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.55, #queue-req: 0
[2026-09-10 05:27:30 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.73, #queue-req: 0
[2026-09-10 05:27:31 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.76, #queue-req: 0
[2026-09-10 05:27:31 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.96, #queue-req: 0
[2026-09-10 05:27:31 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.05, #queue-req: 0
[2026-09-10 05:27:32 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.21, #queue-req: 0
[2026-09-10 05:27:32 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.10, #queue-req: 0
[2026-09-10 05:27:33 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.13, #queue-req: 0
[2026-09-10 05:27:33 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.48, #queue-req: 0
[2026-09-10 05:27:33] INFO: 127.0.0.1:39016 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:27:34 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.09
[2026-09-10 05:27:34 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17363.73
[2026-09-10 05:27:34 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17411.52
[2026-09-10 05:27:34 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 389176.04
[2026-09-10 05:27:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
[2026-09-10 05:27:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 109.81, #queue-req: 0
[2026-09-10 05:27:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.05, #queue-req: 0
[2026-09-10 05:27:36 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.08, #queue-req: 0
[2026-09-10 05:27:36 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.08, #queue-req: 0
[2026-09-10 05:27:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.84, #queue-req: 0
[2026-09-10 05:27:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.60, #queue-req: 0
[2026-09-10 05:27:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.85, #queue-req: 0
[2026-09-10 05:27:38 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.69, #queue-req: 0
[2026-09-10 05:27:38 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.54, #queue-req: 0
[2026-09-10 05:27:39 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.92, #queue-req: 0
[2026-09-10 05:27:39 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.63, #queue-req: 0
[2026-09-10 05:27:40 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.11, #queue-req: 0
[2026-09-10 05:27:40] INFO: 127.0.0.1:46422 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:27:40 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.69
[2026-09-10 05:27:40 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17021.02
[2026-09-10 05:27:41 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 16924.05
[2026-09-10 05:27:41 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 394127.88
[2026-09-10 05:27:41 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
[2026-09-10 05:27:41 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 98.08, #queue-req: 0
[2026-09-10 05:27:42 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.83, #queue-req: 0
[2026-09-10 05:27:42 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.91, #queue-req: 0
[2026-09-10 05:27:43 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.48, #queue-req: 0
[2026-09-10 05:27:43 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.25, #queue-req: 0
[2026-09-10 05:27:44 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.29, #queue-req: 0
[2026-09-10 05:27:44 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.75, #queue-req: 0
[2026-09-10 05:27:44 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.24, #queue-req: 0
[2026-09-10 05:27:45 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.23, #queue-req: 0
[2026-09-10 05:27:45 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.46, #queue-req: 0
[2026-09-10 05:27:46 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.35, #queue-req: 0
[2026-09-10 05:27:46 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.69, #queue-req: 0
[2026-09-10 05:27:46] INFO: 127.0.0.1:46432 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:27:47 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.41
[2026-09-10 05:27:47 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17578.22
[2026-09-10 05:27:47 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17677.17
[2026-09-10 05:27:47 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 396939.69
[2026-09-10 05:27:48 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
[2026-09-10 05:27:48 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 112.61, #queue-req: 0
[2026-09-10 05:27:48 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 99.63, #queue-req: 0
[2026-09-10 05:27:49 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.89, #queue-req: 0
[2026-09-10 05:27:49 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.06, #queue-req: 0
[2026-09-10 05:27:50 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.47, #queue-req: 0
[2026-09-10 05:27:50 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.00, #queue-req: 0
[2026-09-10 05:27:50 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.13, #queue-req: 0
[2026-09-10 05:27:51 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.45, #queue-req: 0
[2026-09-10 05:27:51 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.03, #queue-req: 0
[2026-09-10 05:27:52 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.70, #queue-req: 0
[2026-09-10 05:27:52 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.27, #queue-req: 0
[2026-09-10 05:27:53 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.01, #queue-req: 0
[2026-09-10 05:27:53] INFO: 127.0.0.1:50828 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:27:53 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.07
[2026-09-10 05:27:53 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17228.48
[2026-09-10 05:27:54 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17207.25
[2026-09-10 05:27:54 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 410922.47
[2026-09-10 05:27:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.92, #queue-req: 0
[2026-09-10 05:27:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 99.44, #queue-req: 0
[2026-09-10 05:27:55 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.80, #queue-req: 0
[2026-09-10 05:27:55 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.61, #queue-req: 0
[2026-09-10 05:27:56 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.42, #queue-req: 0
[2026-09-10 05:27:56 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.16, #queue-req: 0
[2026-09-10 05:27:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.03, #queue-req: 0
[2026-09-10 05:27:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.03, #queue-req: 0
[2026-09-10 05:27:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.96, #queue-req: 0
[2026-09-10 05:27:58 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.27, #queue-req: 0
[2026-09-10 05:27:58 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.32, #queue-req: 0
[2026-09-10 05:27:59 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.85, #queue-req: 0
[2026-09-10 05:27:59] INFO: 127.0.0.1:41758 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:28:00 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.31
[2026-09-10 05:28:00 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17338.35
[2026-09-10 05:28:00 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17315.81
[2026-09-10 05:28:00 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 391036.78
[2026-09-10 05:28:00 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
[2026-09-10 05:28:01 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 105.44, #queue-req: 0
[2026-09-10 05:28:01 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.51, #queue-req: 0
[2026-09-10 05:28:02 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.12, #queue-req: 0
[2026-09-10 05:28:02 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.85, #queue-req: 0
[2026-09-10 05:28:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.30, #queue-req: 0
[2026-09-10 05:28:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.13, #queue-req: 0
[2026-09-10 05:28:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.22, #queue-req: 0
[2026-09-10 05:28:04 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.75, #queue-req: 0
[2026-09-10 05:28:04 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.40, #queue-req: 0
[2026-09-10 05:28:05 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.83, #queue-req: 0
[2026-09-10 05:28:05 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.66, #queue-req: 0
[2026-09-10 05:28:05] INFO: 127.0.0.1:41770 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:28:06 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.24
[2026-09-10 05:28:06 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17057.57
[2026-09-10 05:28:06 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 16873.65
[2026-09-10 05:28:06 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 410991.66
[2026-09-10 05:28:07 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.95, #queue-req: 0
[2026-09-10 05:28:07 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.17, #queue-req: 0
[2026-09-10 05:28:08 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.73, #queue-req: 0
[2026-09-10 05:28:08 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.49, #queue-req: 0
[2026-09-10 05:28:09 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.48, #queue-req: 0
[2026-09-10 05:28:09 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.23, #queue-req: 0
[2026-09-10 05:28:10 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.79, #queue-req: 0
[2026-09-10 05:28:10 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.48, #queue-req: 0
[2026-09-10 05:28:10 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.89, #queue-req: 0
[2026-09-10 05:28:11 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.87, #queue-req: 0
[2026-09-10 05:28:11 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.32, #queue-req: 0
[2026-09-10 05:28:12 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.00, #queue-req: 0
[2026-09-10 05:28:12] INFO: 127.0.0.1:55260 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:28:13 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.39
[2026-09-10 05:28:13 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17695.04
[2026-09-10 05:28:13 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17662.37
[2026-09-10 05:28:13 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 437870.59
[2026-09-10 05:28:14 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
[2026-09-10 05:28:14 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.11, #queue-req: 0
[2026-09-10 05:28:14 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.93, #queue-req: 0
[2026-09-10 05:28:15 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.88, #queue-req: 0
[2026-09-10 05:28:15 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.11, #queue-req: 0
[2026-09-10 05:28:16 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.93, #queue-req: 0
[2026-09-10 05:28:16 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.21, #queue-req: 0
[2026-09-10 05:28:17 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.39, #queue-req: 0
[2026-09-10 05:28:17 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.06, #queue-req: 0
[2026-09-10 05:28:17 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.00, #queue-req: 0
[2026-09-10 05:28:18 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.62, #queue-req: 0
[2026-09-10 05:28:18 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.05, #queue-req: 0
[2026-09-10 05:28:19] INFO: 127.0.0.1:59888 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:28:19 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.25
[2026-09-10 05:28:19 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17140.02
[2026-09-10 05:28:20 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17220.37
[2026-09-10 05:28:20 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 387351.35
[2026-09-10 05:28:20 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
[2026-09-10 05:28:20 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.76, #queue-req: 0
[2026-09-10 05:28:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.82, #queue-req: 0
[2026-09-10 05:28:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.32, #queue-req: 0
[2026-09-10 05:28:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.85, #queue-req: 0
[2026-09-10 05:28:22 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.69, #queue-req: 0
[2026-09-10 05:28:22 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.88, #queue-req: 0
[2026-09-10 05:28:23 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.80, #queue-req: 0
[2026-09-10 05:28:23 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.91, #queue-req: 0
[2026-09-10 05:28:24 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.49, #queue-req: 0
[2026-09-10 05:28:24 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.99, #queue-req: 0
[2026-09-10 05:28:24 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.13, #queue-req: 0
[2026-09-10 05:28:25 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.81, #queue-req: 0
[2026-09-10 05:28:25] INFO: 127.0.0.1:59902 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:28:26 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.08
[2026-09-10 05:28:26 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17350.40
[2026-09-10 05:28:26 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17335.52
[2026-09-10 05:28:26 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 396368.12
[2026-09-10 05:28:26 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
[2026-09-10 05:28:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.17, #queue-req: 0
[2026-09-10 05:28:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 103.86, #queue-req: 0
[2026-09-10 05:28:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.88, #queue-req: 0
[2026-09-10 05:28:28 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.39, #queue-req: 0
[2026-09-10 05:28:28 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.59, #queue-req: 0
[2026-09-10 05:28:29 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.83, #queue-req: 0
[2026-09-10 05:28:29 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.65, #queue-req: 0
[2026-09-10 05:28:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.39, #queue-req: 0
[2026-09-10 05:28:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.83, #queue-req: 0
[2026-09-10 05:28:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.08, #queue-req: 0
[2026-09-10 05:28:31 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.75, #queue-req: 0
[2026-09-10 05:28:31 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.12, #queue-req: 0
[2026-09-10 05:28:32] INFO: 127.0.0.1:60442 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:28:32 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.94
[2026-09-10 05:28:32 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17017.54
[2026-09-10 05:28:33 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17011.39
[2026-09-10 05:28:33 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 398319.36
[2026-09-10 05:28:33 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
[2026-09-10 05:28:33 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 109.96, #queue-req: 0
[2026-09-10 05:28:33 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.71, #queue-req: 0
[2026-09-10 05:28:34 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.57, #queue-req: 0
[2026-09-10 05:28:34 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.96, #queue-req: 0
[2026-09-10 05:28:35 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.13, #queue-req: 0
[2026-09-10 05:28:35 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.85, #queue-req: 0
[2026-09-10 05:28:36 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.15, #queue-req: 0
[2026-09-10 05:28:36 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.79, #queue-req: 0
[2026-09-10 05:28:37 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.43, #queue-req: 0
[2026-09-10 05:28:37 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.43, #queue-req: 0
[2026-09-10 05:28:38 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.16, #queue-req: 0
[2026-09-10 05:28:38 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.87, #queue-req: 0
[2026-09-10 05:28:38] INFO: 127.0.0.1:51206 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:28:39 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 158.80
[2026-09-10 05:28:39 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17702.47
[2026-09-10 05:28:39 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17638.20
[2026-09-10 05:28:39 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 392293.20
[2026-09-10 05:28:39 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
[2026-09-10 05:28:40 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.63, #queue-req: 0
[2026-09-10 05:28:40 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.20, #queue-req: 0
[2026-09-10 05:28:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.62, #queue-req: 0
[2026-09-10 05:28:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.31, #queue-req: 0
[2026-09-10 05:28:42 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.75, #queue-req: 0
[2026-09-10 05:28:42 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.51, #queue-req: 0
[2026-09-10 05:28:42 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.04, #queue-req: 0
[2026-09-10 05:28:43 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.46, #queue-req: 0
[2026-09-10 05:28:43 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.30, #queue-req: 0
[2026-09-10 05:28:44 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.93, #queue-req: 0
[2026-09-10 05:28:44 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.25, #queue-req: 0
[2026-09-10 05:28:45 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.30, #queue-req: 0
[2026-09-10 05:28:45] INFO: 127.0.0.1:51216 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:28:45 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 158.91
[2026-09-10 05:28:46 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17041.01
[2026-09-10 05:28:46 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17065.29
[2026-09-10 05:28:46 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 379012.97
[2026-09-10 05:28:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
[2026-09-10 05:28:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 110.72, #queue-req: 0
[2026-09-10 05:28:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.25, #queue-req: 0
[2026-09-10 05:28:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.53, #queue-req: 0
[2026-09-10 05:28:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.02, #queue-req: 0
[2026-09-10 05:28:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.82, #queue-req: 0
[2026-09-10 05:28:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.68, #queue-req: 0
[2026-09-10 05:28:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.21, #queue-req: 0
[2026-09-10 05:28:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.74, #queue-req: 0
[2026-09-10 05:28:50 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.48, #queue-req: 0
[2026-09-10 05:28:50 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.47, #queue-req: 0
[2026-09-10 05:28:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.88, #queue-req: 0
[2026-09-10 05:28:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.74, #queue-req: 0
[2026-09-10 05:28:51] INFO: 127.0.0.1:48420 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:28:52 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.07
[2026-09-10 05:28:52 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17348.12
[2026-09-10 05:28:52 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17338.37
[2026-09-10 05:28:52 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 393722.24
[2026-09-10 05:28:53 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.88, #queue-req: 0
[2026-09-10 05:28:53 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 110.81, #queue-req: 0
[2026-09-10 05:28:53 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.58, #queue-req: 0
[2026-09-10 05:28:54 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.89, #queue-req: 0
[2026-09-10 05:28:54 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.16, #queue-req: 0
[2026-09-10 05:28:55 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.96, #queue-req: 0
[2026-09-10 05:28:55 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.17, #queue-req: 0
[2026-09-10 05:28:55 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.13, #queue-req: 0
[2026-09-10 05:28:56 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.94, #queue-req: 0
[2026-09-10 05:28:56 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.61, #queue-req: 0
[2026-09-10 05:28:57 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.43, #queue-req: 0
[2026-09-10 05:28:57 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.14, #queue-req: 0
[2026-09-10 05:28:58 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.35, #queue-req: 0
[2026-09-10 05:28:58] INFO: 127.0.0.1:48932 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:28:58 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 158.73
[2026-09-10 05:28:59 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17549.97
[2026-09-10 05:28:59 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17455.86
[2026-09-10 05:28:59 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 391661.25
[2026-09-10 05:28:59 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
[2026-09-10 05:28:59 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 107.06, #queue-req: 0
[2026-09-10 05:29:00 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.85, #queue-req: 0
[2026-09-10 05:29:00 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.87, #queue-req: 0
[2026-09-10 05:29:01 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.70, #queue-req: 0
[2026-09-10 05:29:01 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.79, #queue-req: 0
[2026-09-10 05:29:02 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.64, #queue-req: 0
[2026-09-10 05:29:02 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.60, #queue-req: 0
[2026-09-10 05:29:03 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.28, #queue-req: 0
[2026-09-10 05:29:03 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.57, #queue-req: 0
[2026-09-10 05:29:04 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.36, #queue-req: 0
[2026-09-10 05:29:04 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.60, #queue-req: 0
[2026-09-10 05:29:04 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.26, #queue-req: 0
[2026-09-10 05:29:05] INFO: 127.0.0.1:48936 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:29:05 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.06
[2026-09-10 05:29:05 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17735.57
[2026-09-10 05:29:06 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17689.19
[2026-09-10 05:29:06 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 418243.76
[2026-09-10 05:29:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
[2026-09-10 05:29:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.43, #queue-req: 0
[2026-09-10 05:29:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 100.83, #queue-req: 0
[2026-09-10 05:29:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.53, #queue-req: 0
[2026-09-10 05:29:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.45, #queue-req: 0
[2026-09-10 05:29:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.68, #queue-req: 0
[2026-09-10 05:29:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.88, #queue-req: 0
[2026-09-10 05:29:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.53, #queue-req: 0
[2026-09-10 05:29:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.67, #queue-req: 0
[2026-09-10 05:29:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.28, #queue-req: 0
[2026-09-10 05:29:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.07, #queue-req: 0
[2026-09-10 05:29:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.23, #queue-req: 0
[2026-09-10 05:29:11 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.37, #queue-req: 0
[2026-09-10 05:29:11] INFO: 127.0.0.1:58922 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:29:12 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.64
[2026-09-10 05:29:12 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17012.62
[2026-09-10 05:29:12 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17002.83
[2026-09-10 05:29:12 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 380234.28
[2026-09-10 05:29:12 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
[2026-09-10 05:29:13 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 110.93, #queue-req: 0
[2026-09-10 05:29:13 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.34, #queue-req: 0
[2026-09-10 05:29:14 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.01, #queue-req: 0
[2026-09-10 05:29:14 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.84, #queue-req: 0
[2026-09-10 05:29:14 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.67, #queue-req: 0
[2026-09-10 05:29:15 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.14, #queue-req: 0
[2026-09-10 05:29:15 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.86, #queue-req: 0
[2026-09-10 05:29:16 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.64, #queue-req: 0
[2026-09-10 05:29:16 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.55, #queue-req: 0
[2026-09-10 05:29:17 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.19, #queue-req: 0
[2026-09-10 05:29:17 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.60, #queue-req: 0
[2026-09-10 05:29:17 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.50, #queue-req: 0
[2026-09-10 05:29:18] INFO: 127.0.0.1:58924 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:29:18 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.69
[2026-09-10 05:29:18 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17380.48
[2026-09-10 05:29:19 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17372.91
[2026-09-10 05:29:19 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 376968.20
[2026-09-10 05:29:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
[2026-09-10 05:29:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 112.62, #queue-req: 0
[2026-09-10 05:29:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 100.19, #queue-req: 0
[2026-09-10 05:29:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.50, #queue-req: 0
[2026-09-10 05:29:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.97, #queue-req: 0
[2026-09-10 05:29:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.28, #queue-req: 0
[2026-09-10 05:29:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.74, #queue-req: 0
[2026-09-10 05:29:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.04, #queue-req: 0
[2026-09-10 05:29:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.05, #queue-req: 0
[2026-09-10 05:29:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.24, #queue-req: 0
[2026-09-10 05:29:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.07, #queue-req: 0
[2026-09-10 05:29:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.46, #queue-req: 0
[2026-09-10 05:29:24 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.14, #queue-req: 0
[2026-09-10 05:29:24] INFO: 127.0.0.1:54110 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:29:24 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.94
[2026-09-10 05:29:25 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17537.93
[2026-09-10 05:29:25 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17564.52
[2026-09-10 05:29:25 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 385204.31
[2026-09-10 05:29:25 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
[2026-09-10 05:29:26 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 100.03, #queue-req: 0
[2026-09-10 05:29:26 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.20, #queue-req: 0
[2026-09-10 05:29:26 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.30, #queue-req: 0
[2026-09-10 05:29:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.22, #queue-req: 0
[2026-09-10 05:29:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.82, #queue-req: 0
[2026-09-10 05:29:28 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.27, #queue-req: 0
[2026-09-10 05:29:28 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.17, #queue-req: 0
[2026-09-10 05:29:29 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.10, #queue-req: 0
[2026-09-10 05:29:29 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.80, #queue-req: 0
[2026-09-10 05:29:30 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.01, #queue-req: 0
[2026-09-10 05:29:30 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.53, #queue-req: 0
[2026-09-10 05:29:31 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.39, #queue-req: 0
[2026-09-10 05:29:31] INFO: 127.0.0.1:45216 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:29:31 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.88
[2026-09-10 05:29:32 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17620.10
[2026-09-10 05:29:32 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17730.20
[2026-09-10 05:29:32 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 413367.81
[2026-09-10 05:29:32 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
[2026-09-10 05:29:32 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 111.07, #queue-req: 0
[2026-09-10 05:29:33 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.31, #queue-req: 0
[2026-09-10 05:29:33 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.34, #queue-req: 0
[2026-09-10 05:29:34 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.48, #queue-req: 0
[2026-09-10 05:29:34 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.02, #queue-req: 0
[2026-09-10 05:29:34 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.32, #queue-req: 0
[2026-09-10 05:29:35 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.19, #queue-req: 0
[2026-09-10 05:29:35 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.15, #queue-req: 0
[2026-09-10 05:29:36 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.87, #queue-req: 0
[2026-09-10 05:29:36 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.78, #queue-req: 0
[2026-09-10 05:29:37 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.04, #queue-req: 0
[2026-09-10 05:29:37 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.12, #queue-req: 0
[2026-09-10 05:29:37] INFO: 127.0.0.1:45224 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:29:38 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.90
[2026-09-10 05:29:38 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17313.96
[2026-09-10 05:29:38 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17374.01
[2026-09-10 05:29:38 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 364144.24
[2026-09-10 05:29:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
[2026-09-10 05:29:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 105.17, #queue-req: 0
[2026-09-10 05:29:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.47, #queue-req: 0
[2026-09-10 05:29:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.79, #queue-req: 0
[2026-09-10 05:29:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.80, #queue-req: 0
[2026-09-10 05:29:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.67, #queue-req: 0
[2026-09-10 05:29:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.32, #queue-req: 0
[2026-09-10 05:29:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.02, #queue-req: 0
[2026-09-10 05:29:42 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.43, #queue-req: 0
[2026-09-10 05:29:42 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.73, #queue-req: 0
[2026-09-10 05:29:43 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.67, #queue-req: 0
[2026-09-10 05:29:43 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.06, #queue-req: 0
[2026-09-10 05:29:44 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.73, #queue-req: 0
[2026-09-10 05:29:44] INFO: 127.0.0.1:45950 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:29:44 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.61
[2026-09-10 05:29:44 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17307.80
[2026-09-10 05:29:45 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17436.66
[2026-09-10 05:29:45 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 398465.33
[2026-09-10 05:29:45 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
[2026-09-10 05:29:45 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 108.94, #queue-req: 0
[2026-09-10 05:29:46 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 98.77, #queue-req: 0
[2026-09-10 05:29:46 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.00, #queue-req: 0
[2026-09-10 05:29:47 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.99, #queue-req: 0
[2026-09-10 05:29:47 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.07, #queue-req: 0
[2026-09-10 05:29:47 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.58, #queue-req: 0
[2026-09-10 05:29:48 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.86, #queue-req: 0
[2026-09-10 05:29:48 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.20, #queue-req: 0
[2026-09-10 05:29:49 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.27, #queue-req: 0
[2026-09-10 05:29:49 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.66, #queue-req: 0
[2026-09-10 05:29:50 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.47, #queue-req: 0
[2026-09-10 05:29:50 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.08, #queue-req: 0
[2026-09-10 05:29:50] INFO: 127.0.0.1:43446 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:29:51 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.99
[2026-09-10 05:29:51 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17050.57
[2026-09-10 05:29:51 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17024.07
[2026-09-10 05:29:51 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 377001.13
[2026-09-10 05:29:51 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
[2026-09-10 05:29:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.93, #queue-req: 0
[2026-09-10 05:29:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.53, #queue-req: 0
[2026-09-10 05:29:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.25, #queue-req: 0
[2026-09-10 05:29:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.44, #queue-req: 0
[2026-09-10 05:29:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.55, #queue-req: 0
[2026-09-10 05:29:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.26, #queue-req: 0
[2026-09-10 05:29:55 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.63, #queue-req: 0
[2026-09-10 05:29:55 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.24, #queue-req: 0
[2026-09-10 05:29:55 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.19, #queue-req: 0
[2026-09-10 05:29:56 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.55, #queue-req: 0
[2026-09-10 05:29:56 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.83, #queue-req: 0
[2026-09-10 05:29:57 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.44, #queue-req: 0
[2026-09-10 05:29:57] INFO: 127.0.0.1:43456 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:29:57 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.74
[2026-09-10 05:29:58 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17673.58
[2026-09-10 05:29:58 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17444.02
[2026-09-10 05:29:58 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 442709.39
[2026-09-10 05:29:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
[2026-09-10 05:29:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 108.41, #queue-req: 0
[2026-09-10 05:29:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.25, #queue-req: 0
[2026-09-10 05:29:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.60, #queue-req: 0
[2026-09-10 05:30:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.15, #queue-req: 0
[2026-09-10 05:30:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.47, #queue-req: 0
[2026-09-10 05:30:01 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.45, #queue-req: 0
[2026-09-10 05:30:01 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.72, #queue-req: 0
[2026-09-10 05:30:02 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.06, #queue-req: 0
[2026-09-10 05:30:02 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.50, #queue-req: 0
[2026-09-10 05:30:02 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.57, #queue-req: 0
[2026-09-10 05:30:03 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.53, #queue-req: 0
[2026-09-10 05:30:03 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.21, #queue-req: 0
[2026-09-10 05:30:03] INFO: 127.0.0.1:49560 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:30:04 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.87
[2026-09-10 05:30:04 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17024.45
[2026-09-10 05:30:04 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 16899.80
[2026-09-10 05:30:04 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 376913.94
[2026-09-10 05:30:05 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
[2026-09-10 05:30:05 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 101.37, #queue-req: 0
[2026-09-10 05:30:05 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.39, #queue-req: 0
[2026-09-10 05:30:06 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.55, #queue-req: 0
[2026-09-10 05:30:06 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.16, #queue-req: 0
[2026-09-10 05:30:07 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.31, #queue-req: 0
[2026-09-10 05:30:07 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.23, #queue-req: 0
[2026-09-10 05:30:08 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.37, #queue-req: 0
[2026-09-10 05:30:08 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.51, #queue-req: 0
[2026-09-10 05:30:09 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.18, #queue-req: 0
[2026-09-10 05:30:09 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.07, #queue-req: 0
[2026-09-10 05:30:09 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.34, #queue-req: 0
[2026-09-10 05:30:10] INFO: 127.0.0.1:35116 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:30:10 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.85
[2026-09-10 05:30:11 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17331.73
[2026-09-10 05:30:11 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17380.05
[2026-09-10 05:30:11 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 376651.23
[2026-09-10 05:30:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
[2026-09-10 05:30:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 104.75, #queue-req: 0
[2026-09-10 05:30:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.22, #queue-req: 0
[2026-09-10 05:30:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.80, #queue-req: 0
[2026-09-10 05:30:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.49, #queue-req: 0
[2026-09-10 05:30:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.23, #queue-req: 0
[2026-09-10 05:30:14 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.53, #queue-req: 0
[2026-09-10 05:30:14 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.36, #queue-req: 0
[2026-09-10 05:30:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.65, #queue-req: 0
[2026-09-10 05:30:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.52, #queue-req: 0
[2026-09-10 05:30:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.56, #queue-req: 0
[2026-09-10 05:30:16 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.12, #queue-req: 0
[2026-09-10 05:30:16] INFO: 127.0.0.1:35126 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:30:17 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.92
[2026-09-10 05:30:17 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17012.10
[2026-09-10 05:30:17 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17006.24
[2026-09-10 05:30:17 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 385755.05
[2026-09-10 05:30:18 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
[2026-09-10 05:30:18 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.34, #queue-req: 0
[2026-09-10 05:30:18 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.97, #queue-req: 0
[2026-09-10 05:30:19 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.89, #queue-req: 0
[2026-09-10 05:30:19 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.97, #queue-req: 0
[2026-09-10 05:30:20 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.72, #queue-req: 0
[2026-09-10 05:30:20 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.53, #queue-req: 0
[2026-09-10 05:30:21 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.45, #queue-req: 0
[2026-09-10 05:30:21 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.08, #queue-req: 0
[2026-09-10 05:30:21 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.85, #queue-req: 0
[2026-09-10 05:30:22 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.87, #queue-req: 0
[2026-09-10 05:30:22 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.07, #queue-req: 0
[2026-09-10 05:30:23] INFO: 127.0.0.1:45272 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 05:30:23 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.10
[2026-09-10 05:30:24 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17663.48
[2026-09-10 05:30:24 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17724.99
[2026-09-10 05:30:24 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 402718.07
[2026-09-10 05:30:24 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.92, #queue-req: 0
[2026-09-10 05:30:24 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 108.94, #queue-req: 0
[2026-09-10 05:30:25 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.26, #queue-req: 0
[2026-09-10 05:30:25 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.09, #queue-req: 0
[2026-09-10 05:30:26 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.33, #queue-req: 0
[2026-09-10 05:30:26 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.07, #queue-req: 0
[2026-09-10 05:30:27 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.00, #queue-req: 0
[2026-09-10 05:30:27 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.90, #queue-req: 0
[2026-09-10 05:30:27 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.09, #queue-req: 0
[2026-09-10 05:30:28 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.07, #queue-req: 0
[2026-09-10 05:30:28 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.68, #queue-req: 0
[2026-09-10 05:30:29 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.71, #queue-req: 0
[2026-09-10 05:30:29] INFO: 127.0.0.1:37684 - "GET /server_info HTTP/1.1" 200 OK
[2026-09-10 05:30:29] INFO: 127.0.0.1:37692 - "GET /server_info HTTP/1.1" 200 OK

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 270267 MiB, 3847 MiB, 0 %, 245.21 W
1, 266529 MiB, 7585 MiB, 0 %, 236.04 W
2, 267221 MiB, 6893 MiB, 0 %, 237.03 W
3, 267217 MiB, 6897 MiB, 0 %, 247.66 W
4, 4 MiB, 274110 MiB, 0 %, 182.01 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.13 W
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 270267 MiB 3847 MiB 0 % 245.21 W
3 1 266529 MiB 7585 MiB 0 % 236.04 W
4 2 267221 MiB 6893 MiB 0 % 237.03 W
5 3 267217 MiB 6897 MiB 0 % 247.66 W
6 4 4 MiB 274110 MiB 0 % 182.01 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.13 W
9 7 4 MiB 274110 MiB 0 % 182.01 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=640, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=128, output_file='/results/points/16k_512_c128.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=640, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=128, output_file='/results/points/16k_512_c128.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 10485760
#Output tokens: 327680
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 128
Successful requests: 640
Benchmark duration (s): 201.20
Total input tokens: 10485760
Total input text tokens: 10485760
Total generated tokens: 327680
Total generated tokens (retokenized): 324267
Request throughput (req/s): 3.18
Input token throughput (tok/s): 52115.96
Output token throughput (tok/s): 1628.62
Peak output token throughput (tok/s): 8444.00
Peak concurrent requests: 256
Total token throughput (tok/s): 53744.58
Concurrency: 127.47
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 40073.80
Median E2E Latency (ms): 39933.15
P90 E2E Latency (ms): 40867.69
P95 E2E Latency (ms): 40965.30
P99 E2E Latency (ms): 41054.20
---------------Time to First Token----------------
Mean TTFT (ms): 17331.84
Median TTFT (ms): 17343.29
P90 TTFT (ms): 30366.98
P95 TTFT (ms): 31955.08
P99 TTFT (ms): 32531.95
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 44.50
Median TPOT (ms): 44.47
P90 TPOT (ms): 69.86
P95 TPOT (ms): 73.50
P99 TPOT (ms): 75.57
---------------Inter-Token Latency----------------
Mean ITL (ms): 44.51
Median ITL (ms): 15.27
P90 ITL (ms): 19.67
P95 ITL (ms): 21.54
P99 ITL (ms): 22.49
Max ITL (ms): 31081.58
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 269243 MiB, 4871 MiB, 0 %, 244.80 W
1, 268011 MiB, 6103 MiB, 0 %, 236.31 W
2, 268177 MiB, 5937 MiB, 0 %, 236.89 W
3, 266263 MiB, 7851 MiB, 0 %, 246.29 W
4, 4 MiB, 274110 MiB, 0 %, 181.89 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.13 W
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 269243 MiB 4871 MiB 0 % 244.80 W
3 1 268011 MiB 6103 MiB 0 % 236.31 W
4 2 268177 MiB 5937 MiB 0 % 236.89 W
5 3 266263 MiB 7851 MiB 0 % 246.29 W
6 4 4 MiB 274110 MiB 0 % 181.89 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.13 W
9 7 4 MiB 274110 MiB 0 % 182.01 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=1280, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=256, output_file='/results/points/16k_512_c256.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=1280, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=256, output_file='/results/points/16k_512_c256.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 20971520
#Output tokens: 655360
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 256
Successful requests: 1280
Benchmark duration (s): 365.13
Total input tokens: 20971520
Total input text tokens: 20971520
Total generated tokens: 655360
Total generated tokens (retokenized): 647182
Request throughput (req/s): 3.51
Input token throughput (tok/s): 57436.40
Output token throughput (tok/s): 1794.89
Peak output token throughput (tok/s): 16320.00
Peak concurrent requests: 512
Total token throughput (tok/s): 59231.29
Concurrency: 254.97
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 72730.55
Median E2E Latency (ms): 72592.18
P90 E2E Latency (ms): 74451.94
P95 E2E Latency (ms): 74517.86
P99 E2E Latency (ms): 74606.27
---------------Time to First Token----------------
Mean TTFT (ms): 33211.70
Median TTFT (ms): 33119.78
P90 TTFT (ms): 59093.67
P95 TTFT (ms): 62849.59
P99 TTFT (ms): 65119.95
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 77.34
Median TPOT (ms): 77.14
P90 TPOT (ms): 128.18
P95 TPOT (ms): 134.28
P99 TPOT (ms): 139.41
---------------Inter-Token Latency----------------
Mean ITL (ms): 77.34
Median ITL (ms): 16.59
P90 ITL (ms): 22.04
P95 ITL (ms): 24.22
P99 ITL (ms): 25.93
Max ITL (ms): 64829.80
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 267151 MiB, 6963 MiB, 0 %, 243.92 W
1, 264711 MiB, 9403 MiB, 0 %, 236.50 W
2, 267957 MiB, 6157 MiB, 0 %, 236.94 W
3, 266287 MiB, 7827 MiB, 0 %, 246.99 W
4, 4 MiB, 274110 MiB, 0 %, 182.20 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.13 W
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 267151 MiB 6963 MiB 0 % 243.92 W
3 1 264711 MiB 9403 MiB 0 % 236.50 W
4 2 267957 MiB 6157 MiB 0 % 236.94 W
5 3 266287 MiB 7827 MiB 0 % 246.99 W
6 4 4 MiB 274110 MiB 0 % 182.20 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.13 W
9 7 4 MiB 274110 MiB 0 % 182.01 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=160, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=32, output_file='/results/points/16k_512_c32.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=160, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=32, output_file='/results/points/16k_512_c32.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 2621440
#Output tokens: 81920
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 32
Successful requests: 160
Benchmark duration (s): 72.62
Total input tokens: 2621440
Total input text tokens: 2621440
Total generated tokens: 81920
Total generated tokens (retokenized): 81120
Request throughput (req/s): 2.20
Input token throughput (tok/s): 36100.44
Output token throughput (tok/s): 1128.14
Peak output token throughput (tok/s): 2668.00
Peak concurrent requests: 64
Total token throughput (tok/s): 37228.58
Concurrency: 31.86
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 14458.39
Median E2E Latency (ms): 14360.77
P90 E2E Latency (ms): 14889.61
P95 E2E Latency (ms): 14940.54
P99 E2E Latency (ms): 14964.61
---------------Time to First Token----------------
Mean TTFT (ms): 4936.55
Median TTFT (ms): 5135.45
P90 TTFT (ms): 8067.95
P95 TTFT (ms): 8087.71
P99 TTFT (ms): 8603.75
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 18.63
Median TPOT (ms): 18.13
P90 TPOT (ms): 25.17
P95 TPOT (ms): 25.54
P99 TPOT (ms): 25.85
---------------Inter-Token Latency----------------
Mean ITL (ms): 18.63
Median ITL (ms): 12.17
P90 ITL (ms): 13.64
P95 ITL (ms): 14.52
P99 ITL (ms): 16.03
Max ITL (ms): 7200.39
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 270175 MiB, 3939 MiB, 0 %, 245.79 W
1, 265185 MiB, 8929 MiB, 0 %, 236.66 W
2, 267183 MiB, 6931 MiB, 0 %, 238.58 W
3, 265317 MiB, 8797 MiB, 0 %, 247.74 W
4, 4 MiB, 274110 MiB, 0 %, 182.05 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.21 W
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 270175 MiB 3939 MiB 0 % 245.79 W
3 1 265185 MiB 8929 MiB 0 % 236.66 W
4 2 267183 MiB 6931 MiB 0 % 238.58 W
5 3 265317 MiB 8797 MiB 0 % 247.74 W
6 4 4 MiB 274110 MiB 0 % 182.05 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.21 W
9 7 4 MiB 274110 MiB 0 % 182.01 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=320, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=64, output_file='/results/points/16k_512_c64.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=320, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=64, output_file='/results/points/16k_512_c64.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 5242880
#Output tokens: 163840
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 64
Successful requests: 320
Benchmark duration (s): 115.64
Total input tokens: 5242880
Total input text tokens: 5242880
Total generated tokens: 163840
Total generated tokens (retokenized): 161636
Request throughput (req/s): 2.77
Input token throughput (tok/s): 45336.60
Output token throughput (tok/s): 1416.77
Peak output token throughput (tok/s): 4928.00
Peak concurrent requests: 128
Total token throughput (tok/s): 46753.36
Concurrency: 63.66
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 23007.55
Median E2E Latency (ms): 22960.73
P90 E2E Latency (ms): 23535.50
P95 E2E Latency (ms): 23633.89
P99 E2E Latency (ms): 23711.10
---------------Time to First Token----------------
Mean TTFT (ms): 8990.44
Median TTFT (ms): 9165.76
P90 TTFT (ms): 15556.71
P95 TTFT (ms): 16042.56
P99 TTFT (ms): 16476.52
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 27.43
Median TPOT (ms): 27.23
P90 TPOT (ms): 40.16
P95 TPOT (ms): 41.92
P99 TPOT (ms): 42.70
---------------Inter-Token Latency----------------
Mean ITL (ms): 27.43
Median ITL (ms): 13.36
P90 ITL (ms): 17.01
P95 ITL (ms): 18.35
P99 ITL (ms): 19.25
Max ITL (ms): 15156.70
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 268149 MiB, 5965 MiB, 0 %, 241.92 W
1, 263337 MiB, 10777 MiB, 0 %, 234.70 W
2, 267143 MiB, 6971 MiB, 0 %, 236.84 W
3, 263745 MiB, 10369 MiB, 0 %, 245.26 W
4, 4 MiB, 274110 MiB, 0 %, 182.51 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.29 W
7, 4 MiB, 274110 MiB, 0 %, 182.40 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 268149 MiB 5965 MiB 0 % 241.92 W
3 1 263337 MiB 10777 MiB 0 % 234.70 W
4 2 267143 MiB 6971 MiB 0 % 236.84 W
5 3 263745 MiB 10369 MiB 0 % 245.26 W
6 4 4 MiB 274110 MiB 0 % 182.51 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.29 W
9 7 4 MiB 274110 MiB 0 % 182.40 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=8, output_file='/results/points/16k_512_c8.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=8, output_file='/results/points/16k_512_c8.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 1048576
#Output tokens: 32768
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 8
Successful requests: 64
Benchmark duration (s): 61.85
Total input tokens: 1048576
Total input text tokens: 1048576
Total generated tokens: 32768
Total generated tokens (retokenized): 32406
Request throughput (req/s): 1.03
Input token throughput (tok/s): 16954.74
Output token throughput (tok/s): 529.84
Peak output token throughput (tok/s): 792.00
Peak concurrent requests: 16
Total token throughput (tok/s): 17484.58
Concurrency: 7.98
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 7714.34
Median E2E Latency (ms): 7658.51
P90 E2E Latency (ms): 8116.15
P95 E2E Latency (ms): 8130.24
P99 E2E Latency (ms): 8147.10
---------------Time to First Token----------------
Mean TTFT (ms): 1875.84
Median TTFT (ms): 1937.05
P90 TTFT (ms): 2255.93
P95 TTFT (ms): 2578.91
P99 TTFT (ms): 2648.89
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 11.43
Median TPOT (ms): 11.34
P90 TPOT (ms): 12.56
P95 TPOT (ms): 12.69
P99 TPOT (ms): 12.87
---------------Inter-Token Latency----------------
Mean ITL (ms): 11.43
Median ITL (ms): 10.52
P90 ITL (ms): 11.94
P95 ITL (ms): 12.30
P99 ITL (ms): 12.93
Max ITL (ms): 1265.14
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 260225 MiB, 13889 MiB, 0 %, 241.91 W
1, 260417 MiB, 13697 MiB, 0 %, 234.35 W
2, 260529 MiB, 13585 MiB, 0 %, 234.97 W
3, 259501 MiB, 14613 MiB, 0 %, 243.95 W
4, 4 MiB, 274110 MiB, 0 %, 183.34 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.17 W
7, 4 MiB, 274110 MiB, 0 %, 183.89 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 260225 MiB 13889 MiB 0 % 241.91 W
3 1 260417 MiB 13697 MiB 0 % 234.35 W
4 2 260529 MiB 13585 MiB 0 % 234.97 W
5 3 259501 MiB 14613 MiB 0 % 243.95 W
6 4 4 MiB 274110 MiB 0 % 183.34 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.17 W
9 7 4 MiB 274110 MiB 0 % 183.89 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=1, output_file='/results/points/1k_128_c1.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=1, output_file='/results/points/1k_128_c1.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 65536
#Output tokens: 8192
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 1
Successful requests: 64
Benchmark duration (s): 101.91
Total input tokens: 65536
Total input text tokens: 65536
Total generated tokens: 8192
Total generated tokens (retokenized): 8079
Request throughput (req/s): 0.63
Input token throughput (tok/s): 643.09
Output token throughput (tok/s): 80.39
Peak output token throughput (tok/s): 107.00
Peak concurrent requests: 2
Total token throughput (tok/s): 723.47
Concurrency: 1.00
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 1590.91
Median E2E Latency (ms): 1591.12
P90 E2E Latency (ms): 1655.93
P95 E2E Latency (ms): 1682.16
P99 E2E Latency (ms): 1732.89
---------------Time to First Token----------------
Mean TTFT (ms): 331.86
Median TTFT (ms): 331.96
P90 TTFT (ms): 338.18
P95 TTFT (ms): 342.42
P99 TTFT (ms): 346.96
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 9.91
Median TPOT (ms): 9.91
P90 TPOT (ms): 10.43
P95 TPOT (ms): 10.62
P99 TPOT (ms): 11.06
---------------Inter-Token Latency----------------
Mean ITL (ms): 9.92
Median ITL (ms): 9.93
P90 ITL (ms): 11.44
P95 ITL (ms): 11.75
P99 ITL (ms): 12.22
Max ITL (ms): 15.98
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 267981 MiB, 6133 MiB, 0 %, 243.80 W
1, 268261 MiB, 5853 MiB, 0 %, 235.13 W
2, 267361 MiB, 6753 MiB, 0 %, 236.93 W
3, 264909 MiB, 9205 MiB, 0 %, 245.82 W
4, 4 MiB, 274110 MiB, 0 %, 182.47 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 183.17 W
7, 4 MiB, 274110 MiB, 0 %, 182.83 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 267981 MiB 6133 MiB 0 % 243.80 W
3 1 268261 MiB 5853 MiB 0 % 235.13 W
4 2 267361 MiB 6753 MiB 0 % 236.93 W
5 3 264909 MiB 9205 MiB 0 % 245.82 W
6 4 4 MiB 274110 MiB 0 % 182.47 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 183.17 W
9 7 4 MiB 274110 MiB 0 % 182.83 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=640, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=128, output_file='/results/points/1k_128_c128.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=640, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=128, output_file='/results/points/1k_128_c128.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 655360
#Output tokens: 81920
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 128
Successful requests: 640
Benchmark duration (s): 22.24
Total input tokens: 655360
Total input text tokens: 655360
Total generated tokens: 81920
Total generated tokens (retokenized): 80729
Request throughput (req/s): 28.78
Input token throughput (tok/s): 29472.67
Output token throughput (tok/s): 3684.08
Peak output token throughput (tok/s): 8508.00
Peak concurrent requests: 256
Total token throughput (tok/s): 33156.76
Concurrency: 126.32
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 4388.74
Median E2E Latency (ms): 4274.89
P90 E2E Latency (ms): 4888.74
P95 E2E Latency (ms): 4916.76
P99 E2E Latency (ms): 4964.03
---------------Time to First Token----------------
Mean TTFT (ms): 1748.02
Median TTFT (ms): 1769.41
P90 TTFT (ms): 2354.73
P95 TTFT (ms): 2841.20
P99 TTFT (ms): 2851.83
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 20.79
Median TPOT (ms): 19.82
P90 TPOT (ms): 27.51
P95 TPOT (ms): 27.68
P99 TPOT (ms): 28.92
---------------Inter-Token Latency----------------
Mean ITL (ms): 20.80
Median ITL (ms): 15.31
P90 ITL (ms): 19.24
P95 ITL (ms): 20.39
P99 ITL (ms): 52.25
Max ITL (ms): 1851.24
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 266755 MiB, 7359 MiB, 0 %, 243.84 W
1, 266199 MiB, 7915 MiB, 0 %, 235.27 W
2, 269157 MiB, 4957 MiB, 0 %, 236.85 W
3, 266425 MiB, 7689 MiB, 0 %, 245.84 W
4, 4 MiB, 274110 MiB, 0 %, 182.24 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.17 W
7, 4 MiB, 274110 MiB, 0 %, 182.05 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 266755 MiB 7359 MiB 0 % 243.84 W
3 1 266199 MiB 7915 MiB 0 % 235.27 W
4 2 269157 MiB 4957 MiB 0 % 236.85 W
5 3 266425 MiB 7689 MiB 0 % 245.84 W
6 4 4 MiB 274110 MiB 0 % 182.24 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.17 W
9 7 4 MiB 274110 MiB 0 % 182.05 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=1280, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=256, output_file='/results/points/1k_128_c256.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=1280, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=256, output_file='/results/points/1k_128_c256.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 1310720
#Output tokens: 163840
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 256
Successful requests: 1280
Benchmark duration (s): 36.14
Total input tokens: 1310720
Total input text tokens: 1310720
Total generated tokens: 163840
Total generated tokens (retokenized): 161449
Request throughput (req/s): 35.42
Input token throughput (tok/s): 36265.26
Output token throughput (tok/s): 4533.16
Peak output token throughput (tok/s): 16224.00
Peak concurrent requests: 512
Total token throughput (tok/s): 40798.42
Concurrency: 251.46
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 7100.38
Median E2E Latency (ms): 6570.99
P90 E2E Latency (ms): 9276.25
P95 E2E Latency (ms): 9295.96
P99 E2E Latency (ms): 9342.45
---------------Time to First Token----------------
Mean TTFT (ms): 2807.50
Median TTFT (ms): 2728.95
P90 TTFT (ms): 4249.06
P95 TTFT (ms): 4505.51
P99 TTFT (ms): 7117.16
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 33.80
Median TPOT (ms): 32.71
P90 TPOT (ms): 51.77
P95 TPOT (ms): 59.65
P99 TPOT (ms): 67.15
---------------Inter-Token Latency----------------
Mean ITL (ms): 33.81
Median ITL (ms): 16.52
P90 ITL (ms): 21.80
P95 ITL (ms): 24.21
P99 ITL (ms): 62.30
Max ITL (ms): 6732.10
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 265741 MiB, 8373 MiB, 0 %, 241.89 W
1, 263601 MiB, 10513 MiB, 0 %, 234.70 W
2, 265273 MiB, 8841 MiB, 0 %, 236.67 W
3, 264477 MiB, 9637 MiB, 0 %, 245.12 W
4, 4 MiB, 274110 MiB, 0 %, 183.15 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.52 W
7, 4 MiB, 274110 MiB, 0 %, 183.93 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 265741 MiB 8373 MiB 0 % 241.89 W
3 1 263601 MiB 10513 MiB 0 % 234.70 W
4 2 265273 MiB 8841 MiB 0 % 236.67 W
5 3 264477 MiB 9637 MiB 0 % 245.12 W
6 4 4 MiB 274110 MiB 0 % 183.15 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.52 W
9 7 4 MiB 274110 MiB 0 % 183.93 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=160, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=32, output_file='/results/points/1k_128_c32.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=160, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=32, output_file='/results/points/1k_128_c32.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 163840
#Output tokens: 20480
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 32
Successful requests: 160
Benchmark duration (s): 11.91
Total input tokens: 163840
Total input text tokens: 163840
Total generated tokens: 20480
Total generated tokens (retokenized): 20159
Request throughput (req/s): 13.43
Input token throughput (tok/s): 13757.02
Output token throughput (tok/s): 1719.63
Peak output token throughput (tok/s): 2656.00
Peak concurrent requests: 64
Total token throughput (tok/s): 15476.65
Concurrency: 31.68
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 2358.10
Median E2E Latency (ms): 2329.83
P90 E2E Latency (ms): 2490.41
P95 E2E Latency (ms): 2504.46
P99 E2E Latency (ms): 2511.69
---------------Time to First Token----------------
Mean TTFT (ms): 790.50
Median TTFT (ms): 764.38
P90 TTFT (ms): 935.92
P95 TTFT (ms): 943.32
P99 TTFT (ms): 947.21
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 12.34
Median TPOT (ms): 12.28
P90 TPOT (ms): 12.45
P95 TPOT (ms): 12.48
P99 TPOT (ms): 14.38
---------------Inter-Token Latency----------------
Mean ITL (ms): 12.35
Median ITL (ms): 12.20
P90 ITL (ms): 14.33
P95 ITL (ms): 14.88
P99 ITL (ms): 19.57
Max ITL (ms): 289.57
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 265741 MiB, 8373 MiB, 0 %, 243.69 W
1, 265289 MiB, 8825 MiB, 0 %, 235.99 W
2, 265307 MiB, 8807 MiB, 0 %, 237.03 W
3, 264623 MiB, 9491 MiB, 0 %, 245.81 W
4, 4 MiB, 274110 MiB, 0 %, 183.38 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.74 W
7, 4 MiB, 274110 MiB, 0 %, 183.81 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 265741 MiB 8373 MiB 0 % 243.69 W
3 1 265289 MiB 8825 MiB 0 % 235.99 W
4 2 265307 MiB 8807 MiB 0 % 237.03 W
5 3 264623 MiB 9491 MiB 0 % 245.81 W
6 4 4 MiB 274110 MiB 0 % 183.38 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.74 W
9 7 4 MiB 274110 MiB 0 % 183.81 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=320, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=64, output_file='/results/points/1k_128_c64.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=320, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=64, output_file='/results/points/1k_128_c64.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 327680
#Output tokens: 40960
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 64
Successful requests: 320
Benchmark duration (s): 15.44
Total input tokens: 327680
Total input text tokens: 327680
Total generated tokens: 40960
Total generated tokens (retokenized): 40414
Request throughput (req/s): 20.72
Input token throughput (tok/s): 21219.84
Output token throughput (tok/s): 2652.48
Peak output token throughput (tok/s): 4771.00
Peak concurrent requests: 128
Total token throughput (tok/s): 23872.32
Concurrency: 62.85
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 3032.75
Median E2E Latency (ms): 3019.70
P90 E2E Latency (ms): 3197.66
P95 E2E Latency (ms): 3218.95
P99 E2E Latency (ms): 3245.89
---------------Time to First Token----------------
Mean TTFT (ms): 1104.28
Median TTFT (ms): 1027.77
P90 TTFT (ms): 1355.06
P95 TTFT (ms): 1451.23
P99 TTFT (ms): 1457.66
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 15.18
Median TPOT (ms): 13.88
P90 TPOT (ms): 17.80
P95 TPOT (ms): 17.88
P99 TPOT (ms): 19.35
---------------Inter-Token Latency----------------
Mean ITL (ms): 15.19
Median ITL (ms): 13.44
P90 ITL (ms): 17.03
P95 ITL (ms): 18.35
P99 ITL (ms): 26.81
Max ITL (ms): 808.18
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 262295 MiB, 11819 MiB, 0 %, 241.92 W
1, 260983 MiB, 13131 MiB, 0 %, 234.70 W
2, 262043 MiB, 12071 MiB, 0 %, 234.99 W
3, 261853 MiB, 12261 MiB, 0 %, 243.96 W
4, 4 MiB, 274110 MiB, 0 %, 183.07 W
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
6, 4 MiB, 274110 MiB, 0 %, 182.49 W
7, 4 MiB, 274110 MiB, 0 %, 183.93 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 262295 MiB 11819 MiB 0 % 241.92 W
3 1 260983 MiB 13131 MiB 0 % 234.70 W
4 2 262043 MiB 12071 MiB 0 % 234.99 W
5 3 261853 MiB 12261 MiB 0 % 243.96 W
6 4 4 MiB 274110 MiB 0 % 183.07 W
7 5 4 MiB 274110 MiB 0 % 182.13 W
8 6 4 MiB 274110 MiB 0 % 182.49 W
9 7 4 MiB 274110 MiB 0 % 183.93 W

View File

@ -0,0 +1,59 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=8, output_file='/results/points/1k_128_c8.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
Server ready in 0.0s.
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=8, output_file='/results/points/1k_128_c8.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
#Input tokens: 65536
#Output tokens: 8192
Starting warmup with 1 sequences...
Warmup completed with 1 sequences. Starting main benchmark run...
============ Serving Benchmark Result ============
Backend: sglang
Traffic request rate: inf
Max request concurrency: 8
Successful requests: 64
Benchmark duration (s): 14.92
Total input tokens: 65536
Total input text tokens: 65536
Total generated tokens: 8192
Total generated tokens (retokenized): 8085
Request throughput (req/s): 4.29
Input token throughput (tok/s): 4391.78
Output token throughput (tok/s): 548.97
Peak output token throughput (tok/s): 792.00
Peak concurrent requests: 16
Total token throughput (tok/s): 4940.76
Concurrency: 7.96
----------------End-to-End Latency----------------
Mean E2E Latency (ms): 1856.61
Median E2E Latency (ms): 1838.64
P90 E2E Latency (ms): 1935.91
P95 E2E Latency (ms): 1939.07
P99 E2E Latency (ms): 1944.16
---------------Time to First Token----------------
Mean TTFT (ms): 558.89
Median TTFT (ms): 548.42
P90 TTFT (ms): 646.48
P95 TTFT (ms): 649.01
P99 TTFT (ms): 651.38
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms): 10.22
Median TPOT (ms): 10.13
P90 TPOT (ms): 10.70
P95 TPOT (ms): 10.73
P99 TPOT (ms): 10.76
---------------Inter-Token Latency----------------
Mean ITL (ms): 10.22
Median ITL (ms): 10.10
P90 ITL (ms): 11.38
P95 ITL (ms): 11.71
P99 ITL (ms): 12.88
Max ITL (ms): 22.24
==================================================

View File

@ -0,0 +1,9 @@
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
0, 260225 MiB, 13889 MiB, 79 %, 367.26 W
1, 260419 MiB, 13695 MiB, 80 %, 357.83 W
2, 260531 MiB, 13583 MiB, 78 %, 356.71 W
3, 259501 MiB, 14613 MiB, 87 %, 366.72 W
4, 4 MiB, 274110 MiB, 0 %, 183.82 W
5, 4 MiB, 274110 MiB, 0 %, 182.17 W
6, 4 MiB, 274110 MiB, 0 %, 183.67 W
7, 4 MiB, 274110 MiB, 0 %, 183.97 W
1 index memory.used [MiB] memory.free [MiB] utilization.gpu [%] power.draw [W]
2 0 260225 MiB 13889 MiB 79 % 367.26 W
3 1 260419 MiB 13695 MiB 80 % 357.83 W
4 2 260531 MiB 13583 MiB 78 % 356.71 W
5 3 259501 MiB 14613 MiB 87 % 366.72 W
6 4 4 MiB 274110 MiB 0 % 183.82 W
7 5 4 MiB 274110 MiB 0 % 182.17 W
8 6 4 MiB 274110 MiB 0 % 183.67 W
9 7 4 MiB 274110 MiB 0 % 183.97 W

View File

@ -0,0 +1,2 @@
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
warnings.warn(

View File

@ -0,0 +1,704 @@
[2026-09-10 06:30:56 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.05, #queue-req: 0
[2026-09-10 06:30:56 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.69, #queue-req: 0
[2026-09-10 06:30:57 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.03, #queue-req: 0
[2026-09-10 06:30:57 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.40, #queue-req: 0
[2026-09-10 06:30:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.21, #queue-req: 0
[2026-09-10 06:30:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.15, #queue-req: 0
[2026-09-10 06:30:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.72, #queue-req: 0
[2026-09-10 06:30:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.89, #queue-req: 0
[2026-09-10 06:30:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.64, #queue-req: 0
[2026-09-10 06:31:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.45, #queue-req: 0
[2026-09-10 06:31:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.63, #queue-req: 0
[2026-09-10 06:31:01 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.53, #queue-req: 0
[2026-09-10 06:31:01 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.26, #queue-req: 0
[2026-09-10 06:31:01 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.30, #queue-req: 0
[2026-09-10 06:31:02 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.74, #queue-req: 0
[2026-09-10 06:31:02 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.82, #queue-req: 0
[2026-09-10 06:31:03 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.98, #queue-req: 0
[2026-09-10 06:31:03 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.45, #queue-req: 0
[2026-09-10 06:31:04 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.18, #queue-req: 0
[2026-09-10 06:31:04 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.02, #queue-req: 0
[2026-09-10 06:31:04 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.39, #queue-req: 0
[2026-09-10 06:31:05 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.93, #queue-req: 0
[2026-09-10 06:31:05 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.72, #queue-req: 0
[2026-09-10 06:31:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.81, #queue-req: 0
[2026-09-10 06:31:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.94, #queue-req: 0
[2026-09-10 06:31:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.02, #queue-req: 0
[2026-09-10 06:31:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.01, #queue-req: 0
[2026-09-10 06:31:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.60, #queue-req: 0
[2026-09-10 06:31:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.32, #queue-req: 0
[2026-09-10 06:31:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.13, #queue-req: 0
[2026-09-10 06:31:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.24, #queue-req: 0
[2026-09-10 06:31:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.07, #queue-req: 0
[2026-09-10 06:31:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.77, #queue-req: 0
[2026-09-10 06:31:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.88, #queue-req: 0
[2026-09-10 06:31:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.47, #queue-req: 0
[2026-09-10 06:31:11 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.60, #queue-req: 0
[2026-09-10 06:31:11 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.11, #queue-req: 0
[2026-09-10 06:31:12 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.02, #queue-req: 0
[2026-09-10 06:31:12 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.48, #queue-req: 0
[2026-09-10 06:31:13 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.16, #queue-req: 0
[2026-09-10 06:31:13 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.70, #queue-req: 0
[2026-09-10 06:31:13 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.69, #queue-req: 0
[2026-09-10 06:31:14 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.38, #queue-req: 0
[2026-09-10 06:31:14 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.65, #queue-req: 0
[2026-09-10 06:31:15 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.73, #queue-req: 0
[2026-09-10 06:31:15 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.15, #queue-req: 0
[2026-09-10 06:31:16 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.26, #queue-req: 0
[2026-09-10 06:31:16 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.52, #queue-req: 0
[2026-09-10 06:31:16 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.43, #queue-req: 0
[2026-09-10 06:31:17] INFO: 127.0.0.1:49772 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 06:31:17 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 1024, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 5.71
[2026-09-10 06:31:17 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 0.29, #queue-req: 0
[2026-09-10 06:31:17 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 112.36, #queue-req: 0
[2026-09-10 06:31:18 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.50, #queue-req: 0
[2026-09-10 06:31:18 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.76, #queue-req: 0
[2026-09-10 06:31:19 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.48, #queue-req: 0
[2026-09-10 06:31:19 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.39, #queue-req: 0
[2026-09-10 06:31:20 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.42, #queue-req: 0
[2026-09-10 06:31:20 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.01, #queue-req: 0
[2026-09-10 06:31:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.41, #queue-req: 0
[2026-09-10 06:31:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.01, #queue-req: 0
[2026-09-10 06:31:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.55, #queue-req: 0
[2026-09-10 06:31:22 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.21, #queue-req: 0
[2026-09-10 06:31:22 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.61, #queue-req: 0
[2026-09-10 06:31:23 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.20, #queue-req: 0
[2026-09-10 06:31:23 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.26, #queue-req: 0
[2026-09-10 06:31:24 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.26, #queue-req: 0
[2026-09-10 06:31:24 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.11, #queue-req: 0
[2026-09-10 06:31:25 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.82, #queue-req: 0
[2026-09-10 06:31:25 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.41, #queue-req: 0
[2026-09-10 06:31:25 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.18, #queue-req: 0
[2026-09-10 06:31:26 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.00, #queue-req: 0
[2026-09-10 06:31:26 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.42, #queue-req: 0
[2026-09-10 06:31:27 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.93, #queue-req: 0
[2026-09-10 06:31:27 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.20, #queue-req: 0
[2026-09-10 06:31:28 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.07, #queue-req: 0
[2026-09-10 06:31:28 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.46, #queue-req: 0
[2026-09-10 06:31:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.21, #queue-req: 0
[2026-09-10 06:31:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.66, #queue-req: 0
[2026-09-10 06:31:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.66, #queue-req: 0
[2026-09-10 06:31:30 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.16, #queue-req: 0
[2026-09-10 06:31:30 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.24, #queue-req: 0
[2026-09-10 06:31:31 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.15, #queue-req: 0
[2026-09-10 06:31:31 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.74, #queue-req: 0
[2026-09-10 06:31:32 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.71, #queue-req: 0
[2026-09-10 06:31:32 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.75, #queue-req: 0
[2026-09-10 06:31:33 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.88, #queue-req: 0
[2026-09-10 06:31:33 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.02, #queue-req: 0
[2026-09-10 06:31:33 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.77, #queue-req: 0
[2026-09-10 06:31:34 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.51, #queue-req: 0
[2026-09-10 06:31:34 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.54, #queue-req: 0
[2026-09-10 06:31:35 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.98, #queue-req: 0
[2026-09-10 06:31:35 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.14, #queue-req: 0
[2026-09-10 06:31:36 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.59, #queue-req: 0
[2026-09-10 06:31:36 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.02, #queue-req: 0
[2026-09-10 06:31:36 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.00, #queue-req: 0
[2026-09-10 06:31:37 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.67, #queue-req: 0
[2026-09-10 06:31:37 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.41, #queue-req: 0
[2026-09-10 06:31:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.77, #queue-req: 0
[2026-09-10 06:31:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.84, #queue-req: 0
[2026-09-10 06:31:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.46, #queue-req: 0
[2026-09-10 06:31:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.39, #queue-req: 0
[2026-09-10 06:31:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.14, #queue-req: 0
[2026-09-10 06:31:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.00, #queue-req: 0
[2026-09-10 06:31:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.31, #queue-req: 0
[2026-09-10 06:31:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.85, #queue-req: 0
[2026-09-10 06:31:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.19, #queue-req: 0
[2026-09-10 06:31:42 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.44, #queue-req: 0
[2026-09-10 06:31:42 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.75, #queue-req: 0
[2026-09-10 06:31:43 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.96, #queue-req: 0
[2026-09-10 06:31:43 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.37, #queue-req: 0
[2026-09-10 06:31:44 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.59, #queue-req: 0
[2026-09-10 06:31:44 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.70, #queue-req: 0
[2026-09-10 06:31:44 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.44, #queue-req: 0
[2026-09-10 06:31:45 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.37, #queue-req: 0
[2026-09-10 06:31:45 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.00, #queue-req: 0
[2026-09-10 06:31:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.82, #queue-req: 0
[2026-09-10 06:31:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.76, #queue-req: 0
[2026-09-10 06:31:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.64, #queue-req: 0
[2026-09-10 06:31:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.25, #queue-req: 0
[2026-09-10 06:31:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.44, #queue-req: 0
[2026-09-10 06:31:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.67, #queue-req: 0
[2026-09-10 06:31:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.53, #queue-req: 0
[2026-09-10 06:31:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.30, #queue-req: 0
[2026-09-10 06:31:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.72, #queue-req: 0
[2026-09-10 06:31:50 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.49, #queue-req: 0
[2026-09-10 06:31:50 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.74, #queue-req: 0
[2026-09-10 06:31:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.23, #queue-req: 0
[2026-09-10 06:31:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.64, #queue-req: 0
[2026-09-10 06:31:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.05, #queue-req: 0
[2026-09-10 06:31:52 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.81, #queue-req: 0
[2026-09-10 06:31:52 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.05, #queue-req: 0
[2026-09-10 06:31:53 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.60, #queue-req: 0
[2026-09-10 06:31:53 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.62, #queue-req: 0
[2026-09-10 06:31:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.41, #queue-req: 0
[2026-09-10 06:31:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.43, #queue-req: 0
[2026-09-10 06:31:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.65, #queue-req: 0
[2026-09-10 06:31:55 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.38, #queue-req: 0
[2026-09-10 06:31:55 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.41, #queue-req: 0
[2026-09-10 06:31:56 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.56, #queue-req: 0
[2026-09-10 06:31:56 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.42, #queue-req: 0
[2026-09-10 06:31:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.32, #queue-req: 0
[2026-09-10 06:31:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.28, #queue-req: 0
[2026-09-10 06:31:58 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.95, #queue-req: 0
[2026-09-10 06:31:58 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.38, #queue-req: 0
[2026-09-10 06:31:58 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.68, #queue-req: 0
[2026-09-10 06:31:59 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.69, #queue-req: 0
[2026-09-10 06:31:59 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.57, #queue-req: 0
[2026-09-10 06:32:00 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.44, #queue-req: 0
[2026-09-10 06:32:00 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.63, #queue-req: 0
[2026-09-10 06:32:01 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.79, #queue-req: 0
[2026-09-10 06:32:01 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.58, #queue-req: 0
[2026-09-10 06:32:02 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.28, #queue-req: 0
[2026-09-10 06:32:02 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.37, #queue-req: 0
[2026-09-10 06:32:02] INFO: 127.0.0.1:38660 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 06:32:02 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 1024, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 5.68
[2026-09-10 06:32:02 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 0.29, #queue-req: 0
[2026-09-10 06:32:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.83, #queue-req: 0
[2026-09-10 06:32:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 100.54, #queue-req: 0
[2026-09-10 06:32:04 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.85, #queue-req: 0
[2026-09-10 06:32:04 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.67, #queue-req: 0
[2026-09-10 06:32:04 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.52, #queue-req: 0
[2026-09-10 06:32:05 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.89, #queue-req: 0
[2026-09-10 06:32:05 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.27, #queue-req: 0
[2026-09-10 06:32:06 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.17, #queue-req: 0
[2026-09-10 06:32:06 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.92, #queue-req: 0
[2026-09-10 06:32:07 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.44, #queue-req: 0
[2026-09-10 06:32:07 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.90, #queue-req: 0
[2026-09-10 06:32:07 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.74, #queue-req: 0
[2026-09-10 06:32:08 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.41, #queue-req: 0
[2026-09-10 06:32:08 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.03, #queue-req: 0
[2026-09-10 06:32:09 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.85, #queue-req: 0
[2026-09-10 06:32:09 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.27, #queue-req: 0
[2026-09-10 06:32:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.74, #queue-req: 0
[2026-09-10 06:32:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.30, #queue-req: 0
[2026-09-10 06:32:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.39, #queue-req: 0
[2026-09-10 06:32:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.37, #queue-req: 0
[2026-09-10 06:32:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.44, #queue-req: 0
[2026-09-10 06:32:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.98, #queue-req: 0
[2026-09-10 06:32:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.34, #queue-req: 0
[2026-09-10 06:32:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.58, #queue-req: 0
[2026-09-10 06:32:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.13, #queue-req: 0
[2026-09-10 06:32:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.12, #queue-req: 0
[2026-09-10 06:32:14 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.40, #queue-req: 0
[2026-09-10 06:32:14 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.03, #queue-req: 0
[2026-09-10 06:32:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.31, #queue-req: 0
[2026-09-10 06:32:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.07, #queue-req: 0
[2026-09-10 06:32:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.63, #queue-req: 0
[2026-09-10 06:32:16 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.60, #queue-req: 0
[2026-09-10 06:32:16 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.37, #queue-req: 0
[2026-09-10 06:32:17 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.07, #queue-req: 0
[2026-09-10 06:32:17 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.26, #queue-req: 0
[2026-09-10 06:32:18 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.85, #queue-req: 0
[2026-09-10 06:32:18 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.74, #queue-req: 0
[2026-09-10 06:32:18 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.82, #queue-req: 0
[2026-09-10 06:32:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.86, #queue-req: 0
[2026-09-10 06:32:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.31, #queue-req: 0
[2026-09-10 06:32:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.95, #queue-req: 0
[2026-09-10 06:32:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.25, #queue-req: 0
[2026-09-10 06:32:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.22, #queue-req: 0
[2026-09-10 06:32:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.04, #queue-req: 0
[2026-09-10 06:32:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.75, #queue-req: 0
[2026-09-10 06:32:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.93, #queue-req: 0
[2026-09-10 06:32:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.29, #queue-req: 0
[2026-09-10 06:32:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.54, #queue-req: 0
[2026-09-10 06:32:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.25, #queue-req: 0
[2026-09-10 06:32:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.14, #queue-req: 0
[2026-09-10 06:32:24 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.42, #queue-req: 0
[2026-09-10 06:32:24 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.54, #queue-req: 0
[2026-09-10 06:32:25 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.87, #queue-req: 0
[2026-09-10 06:32:25 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.13, #queue-req: 0
[2026-09-10 06:32:26 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.81, #queue-req: 0
[2026-09-10 06:32:26 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.34, #queue-req: 0
[2026-09-10 06:32:26 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.97, #queue-req: 0
[2026-09-10 06:32:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.19, #queue-req: 0
[2026-09-10 06:32:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.89, #queue-req: 0
[2026-09-10 06:32:28 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.23, #queue-req: 0
[2026-09-10 06:32:28 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.99, #queue-req: 0
[2026-09-10 06:32:29 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.38, #queue-req: 0
[2026-09-10 06:32:29 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.36, #queue-req: 0
[2026-09-10 06:32:29 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.74, #queue-req: 0
[2026-09-10 06:32:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.43, #queue-req: 0
[2026-09-10 06:32:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.30, #queue-req: 0
[2026-09-10 06:32:31 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.31, #queue-req: 0
[2026-09-10 06:32:31 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.29, #queue-req: 0
[2026-09-10 06:32:31 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.17, #queue-req: 0
[2026-09-10 06:32:32 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.37, #queue-req: 0
[2026-09-10 06:32:32 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.64, #queue-req: 0
[2026-09-10 06:32:33 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.00, #queue-req: 0
[2026-09-10 06:32:33 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.17, #queue-req: 0
[2026-09-10 06:32:34 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.00, #queue-req: 0
[2026-09-10 06:32:34 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.59, #queue-req: 0
[2026-09-10 06:32:34 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.13, #queue-req: 0
[2026-09-10 06:32:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.29, #queue-req: 0
[2026-09-10 06:32:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.49, #queue-req: 0
[2026-09-10 06:32:36 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.53, #queue-req: 0
[2026-09-10 06:32:36 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.90, #queue-req: 0
[2026-09-10 06:32:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.64, #queue-req: 0
[2026-09-10 06:32:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.35, #queue-req: 0
[2026-09-10 06:32:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.59, #queue-req: 0
[2026-09-10 06:32:38 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.47, #queue-req: 0
[2026-09-10 06:32:38 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.75, #queue-req: 0
[2026-09-10 06:32:39 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.68, #queue-req: 0
[2026-09-10 06:32:39 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.40, #queue-req: 0
[2026-09-10 06:32:39 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.45, #queue-req: 0
[2026-09-10 06:32:40 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.91, #queue-req: 0
[2026-09-10 06:32:40 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.37, #queue-req: 0
[2026-09-10 06:32:41 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.01, #queue-req: 0
[2026-09-10 06:32:41 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.94, #queue-req: 0
[2026-09-10 06:32:42 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.54, #queue-req: 0
[2026-09-10 06:32:42 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.03, #queue-req: 0
[2026-09-10 06:32:42 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.33, #queue-req: 0
[2026-09-10 06:32:43 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.90, #queue-req: 0
[2026-09-10 06:32:43 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.72, #queue-req: 0
[2026-09-10 06:32:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.66, #queue-req: 0
[2026-09-10 06:32:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.54, #queue-req: 0
[2026-09-10 06:32:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.84, #queue-req: 0
[2026-09-10 06:32:45 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.88, #queue-req: 0
[2026-09-10 06:32:45 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.48, #queue-req: 0
[2026-09-10 06:32:45] INFO: 127.0.0.1:48278 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 06:32:46 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 1024, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 5.69
[2026-09-10 06:32:46 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 0.30, #queue-req: 0
[2026-09-10 06:32:46 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 108.73, #queue-req: 0
[2026-09-10 06:32:47 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.12, #queue-req: 0
[2026-09-10 06:32:47 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.06, #queue-req: 0
[2026-09-10 06:32:48 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.76, #queue-req: 0
[2026-09-10 06:32:48 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.26, #queue-req: 0
[2026-09-10 06:32:48 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.92, #queue-req: 0
[2026-09-10 06:32:49 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.72, #queue-req: 0
[2026-09-10 06:32:49 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.92, #queue-req: 0
[2026-09-10 06:32:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.25, #queue-req: 0
[2026-09-10 06:32:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.80, #queue-req: 0
[2026-09-10 06:32:51 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.13, #queue-req: 0
[2026-09-10 06:32:51 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.39, #queue-req: 0
[2026-09-10 06:32:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.15, #queue-req: 0
[2026-09-10 06:32:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.85, #queue-req: 0
[2026-09-10 06:32:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.47, #queue-req: 0
[2026-09-10 06:32:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.16, #queue-req: 0
[2026-09-10 06:32:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.35, #queue-req: 0
[2026-09-10 06:32:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.05, #queue-req: 0
[2026-09-10 06:32:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.81, #queue-req: 0
[2026-09-10 06:32:55 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.03, #queue-req: 0
[2026-09-10 06:32:55 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.34, #queue-req: 0
[2026-09-10 06:32:56 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.52, #queue-req: 0
[2026-09-10 06:32:56 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.22, #queue-req: 0
[2026-09-10 06:32:57 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.28, #queue-req: 0
[2026-09-10 06:32:57 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.95, #queue-req: 0
[2026-09-10 06:32:58 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.90, #queue-req: 0
[2026-09-10 06:32:58 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.89, #queue-req: 0
[2026-09-10 06:32:59 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.44, #queue-req: 0
[2026-09-10 06:32:59 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.51, #queue-req: 0
[2026-09-10 06:32:59 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.06, #queue-req: 0
[2026-09-10 06:33:00 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.37, #queue-req: 0
[2026-09-10 06:33:00 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.14, #queue-req: 0
[2026-09-10 06:33:01 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.15, #queue-req: 0
[2026-09-10 06:33:01 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.82, #queue-req: 0
[2026-09-10 06:33:02 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.21, #queue-req: 0
[2026-09-10 06:33:02 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.44, #queue-req: 0
[2026-09-10 06:33:03 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.18, #queue-req: 0
[2026-09-10 06:33:03 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.38, #queue-req: 0
[2026-09-10 06:33:04 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.46, #queue-req: 0
[2026-09-10 06:33:04 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.63, #queue-req: 0
[2026-09-10 06:33:05 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.06, #queue-req: 0
[2026-09-10 06:33:05 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.29, #queue-req: 0
[2026-09-10 06:33:05 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.75, #queue-req: 0
[2026-09-10 06:33:06 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.86, #queue-req: 0
[2026-09-10 06:33:06 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.09, #queue-req: 0
[2026-09-10 06:33:07 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.30, #queue-req: 0
[2026-09-10 06:33:07 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.51, #queue-req: 0
[2026-09-10 06:33:08 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.14, #queue-req: 0
[2026-09-10 06:33:08 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.59, #queue-req: 0
[2026-09-10 06:33:09 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.27, #queue-req: 0
[2026-09-10 06:33:09 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.04, #queue-req: 0
[2026-09-10 06:33:10 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.42, #queue-req: 0
[2026-09-10 06:33:10 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.67, #queue-req: 0
[2026-09-10 06:33:11 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.26, #queue-req: 0
[2026-09-10 06:33:11 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.97, #queue-req: 0
[2026-09-10 06:33:11 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.60, #queue-req: 0
[2026-09-10 06:33:12 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.61, #queue-req: 0
[2026-09-10 06:33:12 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.92, #queue-req: 0
[2026-09-10 06:33:13 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.42, #queue-req: 0
[2026-09-10 06:33:13 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.88, #queue-req: 0
[2026-09-10 06:33:14 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.70, #queue-req: 0
[2026-09-10 06:33:14 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.90, #queue-req: 0
[2026-09-10 06:33:15 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.45, #queue-req: 0
[2026-09-10 06:33:15 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.09, #queue-req: 0
[2026-09-10 06:33:16 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.89, #queue-req: 0
[2026-09-10 06:33:16 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.88, #queue-req: 0
[2026-09-10 06:33:16 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.64, #queue-req: 0
[2026-09-10 06:33:17 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.52, #queue-req: 0
[2026-09-10 06:33:17 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.45, #queue-req: 0
[2026-09-10 06:33:18 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.57, #queue-req: 0
[2026-09-10 06:33:18 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.65, #queue-req: 0
[2026-09-10 06:33:19 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.70, #queue-req: 0
[2026-09-10 06:33:19 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.02, #queue-req: 0
[2026-09-10 06:33:20 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.18, #queue-req: 0
[2026-09-10 06:33:20 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.64, #queue-req: 0
[2026-09-10 06:33:21 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.25, #queue-req: 0
[2026-09-10 06:33:21 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.42, #queue-req: 0
[2026-09-10 06:33:22 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.23, #queue-req: 0
[2026-09-10 06:33:22 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.68, #queue-req: 0
[2026-09-10 06:33:22 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.45, #queue-req: 0
[2026-09-10 06:33:23 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.95, #queue-req: 0
[2026-09-10 06:33:23 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.85, #queue-req: 0
[2026-09-10 06:33:24 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.21, #queue-req: 0
[2026-09-10 06:33:24 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.10, #queue-req: 0
[2026-09-10 06:33:25 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.56, #queue-req: 0
[2026-09-10 06:33:25 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.82, #queue-req: 0
[2026-09-10 06:33:26 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.35, #queue-req: 0
[2026-09-10 06:33:26 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.11, #queue-req: 0
[2026-09-10 06:33:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.59, #queue-req: 0
[2026-09-10 06:33:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.78, #queue-req: 0
[2026-09-10 06:33:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.82, #queue-req: 0
[2026-09-10 06:33:28 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.32, #queue-req: 0
[2026-09-10 06:33:28 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.77, #queue-req: 0
[2026-09-10 06:33:29 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.32, #queue-req: 0
[2026-09-10 06:33:29 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.10, #queue-req: 0
[2026-09-10 06:33:30 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.03, #queue-req: 0
[2026-09-10 06:33:30 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.62, #queue-req: 0
[2026-09-10 06:33:31 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.21, #queue-req: 0
[2026-09-10 06:33:31 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.59, #queue-req: 0
[2026-09-10 06:33:32 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.37, #queue-req: 0
[2026-09-10 06:33:32 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.42, #queue-req: 0
[2026-09-10 06:33:33 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.81, #queue-req: 0
[2026-09-10 06:33:33] INFO: 127.0.0.1:51152 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 06:33:33 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 1024, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 5.69
[2026-09-10 06:33:33 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 0.29, #queue-req: 0
[2026-09-10 06:33:33 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 115.10, #queue-req: 0
[2026-09-10 06:33:34 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 103.33, #queue-req: 0
[2026-09-10 06:33:34 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.77, #queue-req: 0
[2026-09-10 06:33:35 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.59, #queue-req: 0
[2026-09-10 06:33:35 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.14, #queue-req: 0
[2026-09-10 06:33:35 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.95, #queue-req: 0
[2026-09-10 06:33:36 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.67, #queue-req: 0
[2026-09-10 06:33:36 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.39, #queue-req: 0
[2026-09-10 06:33:37 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.24, #queue-req: 0
[2026-09-10 06:33:37 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.22, #queue-req: 0
[2026-09-10 06:33:38 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.35, #queue-req: 0
[2026-09-10 06:33:38 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.87, #queue-req: 0
[2026-09-10 06:33:38 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.94, #queue-req: 0
[2026-09-10 06:33:39 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.40, #queue-req: 0
[2026-09-10 06:33:39 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.57, #queue-req: 0
[2026-09-10 06:33:40 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.60, #queue-req: 0
[2026-09-10 06:33:40 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.60, #queue-req: 0
[2026-09-10 06:33:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.72, #queue-req: 0
[2026-09-10 06:33:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.04, #queue-req: 0
[2026-09-10 06:33:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.39, #queue-req: 0
[2026-09-10 06:33:42 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.57, #queue-req: 0
[2026-09-10 06:33:42 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.61, #queue-req: 0
[2026-09-10 06:33:43 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.72, #queue-req: 0
[2026-09-10 06:33:43 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.33, #queue-req: 0
[2026-09-10 06:33:44 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.53, #queue-req: 0
[2026-09-10 06:33:44 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.72, #queue-req: 0
[2026-09-10 06:33:44 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.95, #queue-req: 0
[2026-09-10 06:33:45 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.11, #queue-req: 0
[2026-09-10 06:33:45 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.38, #queue-req: 0
[2026-09-10 06:33:46 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.42, #queue-req: 0
[2026-09-10 06:33:46 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.90, #queue-req: 0
[2026-09-10 06:33:47 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.46, #queue-req: 0
[2026-09-10 06:33:47 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.18, #queue-req: 0
[2026-09-10 06:33:47 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.49, #queue-req: 0
[2026-09-10 06:33:48 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.37, #queue-req: 0
[2026-09-10 06:33:48 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.79, #queue-req: 0
[2026-09-10 06:33:49 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.72, #queue-req: 0
[2026-09-10 06:33:49 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.94, #queue-req: 0
[2026-09-10 06:33:50 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.32, #queue-req: 0
[2026-09-10 06:33:50 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.05, #queue-req: 0
[2026-09-10 06:33:50 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.49, #queue-req: 0
[2026-09-10 06:33:51 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.43, #queue-req: 0
[2026-09-10 06:33:51 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.50, #queue-req: 0
[2026-09-10 06:33:52 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.27, #queue-req: 0
[2026-09-10 06:33:52 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.46, #queue-req: 0
[2026-09-10 06:33:53 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.68, #queue-req: 0
[2026-09-10 06:33:53 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.32, #queue-req: 0
[2026-09-10 06:33:53 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.72, #queue-req: 0
[2026-09-10 06:33:54 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.74, #queue-req: 0
[2026-09-10 06:33:54 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.42, #queue-req: 0
[2026-09-10 06:33:55 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.57, #queue-req: 0
[2026-09-10 06:33:55 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.38, #queue-req: 0
[2026-09-10 06:33:56 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.27, #queue-req: 0
[2026-09-10 06:33:56 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.74, #queue-req: 0
[2026-09-10 06:33:56 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.03, #queue-req: 0
[2026-09-10 06:33:57 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.65, #queue-req: 0
[2026-09-10 06:33:57 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.29, #queue-req: 0
[2026-09-10 06:33:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.18, #queue-req: 0
[2026-09-10 06:33:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.70, #queue-req: 0
[2026-09-10 06:33:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.08, #queue-req: 0
[2026-09-10 06:33:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.33, #queue-req: 0
[2026-09-10 06:34:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.83, #queue-req: 0
[2026-09-10 06:34:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.83, #queue-req: 0
[2026-09-10 06:34:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.22, #queue-req: 0
[2026-09-10 06:34:01 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.23, #queue-req: 0
[2026-09-10 06:34:01 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.73, #queue-req: 0
[2026-09-10 06:34:02 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.19, #queue-req: 0
[2026-09-10 06:34:02 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.22, #queue-req: 0
[2026-09-10 06:34:03 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.58, #queue-req: 0
[2026-09-10 06:34:03 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.51, #queue-req: 0
[2026-09-10 06:34:03 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.98, #queue-req: 0
[2026-09-10 06:34:04 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.07, #queue-req: 0
[2026-09-10 06:34:04 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.02, #queue-req: 0
[2026-09-10 06:34:05 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.21, #queue-req: 0
[2026-09-10 06:34:05 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.61, #queue-req: 0
[2026-09-10 06:34:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.61, #queue-req: 0
[2026-09-10 06:34:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.52, #queue-req: 0
[2026-09-10 06:34:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.94, #queue-req: 0
[2026-09-10 06:34:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.06, #queue-req: 0
[2026-09-10 06:34:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.76, #queue-req: 0
[2026-09-10 06:34:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.22, #queue-req: 0
[2026-09-10 06:34:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.38, #queue-req: 0
[2026-09-10 06:34:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.49, #queue-req: 0
[2026-09-10 06:34:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.06, #queue-req: 0
[2026-09-10 06:34:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.16, #queue-req: 0
[2026-09-10 06:34:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.26, #queue-req: 0
[2026-09-10 06:34:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.77, #queue-req: 0
[2026-09-10 06:34:11 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.94, #queue-req: 0
[2026-09-10 06:34:11 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.93, #queue-req: 0
[2026-09-10 06:34:12 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.05, #queue-req: 0
[2026-09-10 06:34:12 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.90, #queue-req: 0
[2026-09-10 06:34:12 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.46, #queue-req: 0
[2026-09-10 06:34:13 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.32, #queue-req: 0
[2026-09-10 06:34:13 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.37, #queue-req: 0
[2026-09-10 06:34:14 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.52, #queue-req: 0
[2026-09-10 06:34:14 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.31, #queue-req: 0
[2026-09-10 06:34:15 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.06, #queue-req: 0
[2026-09-10 06:34:15 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.78, #queue-req: 0
[2026-09-10 06:34:15 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.75, #queue-req: 0
[2026-09-10 06:34:16 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.05, #queue-req: 0
[2026-09-10 06:34:16 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.38, #queue-req: 0
[2026-09-10 06:34:17 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.75, #queue-req: 0
[2026-09-10 06:34:17] INFO: 127.0.0.1:39600 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 06:34:17 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 1024, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 5.69
[2026-09-10 06:34:17 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 0.30, #queue-req: 0
[2026-09-10 06:34:18 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 106.96, #queue-req: 0
[2026-09-10 06:34:18 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.12, #queue-req: 0
[2026-09-10 06:34:19 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.34, #queue-req: 0
[2026-09-10 06:34:19 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.06, #queue-req: 0
[2026-09-10 06:34:19 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.04, #queue-req: 0
[2026-09-10 06:34:20 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.36, #queue-req: 0
[2026-09-10 06:34:20 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.85, #queue-req: 0
[2026-09-10 06:34:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.01, #queue-req: 0
[2026-09-10 06:34:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.91, #queue-req: 0
[2026-09-10 06:34:22 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.79, #queue-req: 0
[2026-09-10 06:34:22 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.36, #queue-req: 0
[2026-09-10 06:34:23 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.04, #queue-req: 0
[2026-09-10 06:34:23 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.78, #queue-req: 0
[2026-09-10 06:34:23 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.07, #queue-req: 0
[2026-09-10 06:34:24 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.20, #queue-req: 0
[2026-09-10 06:34:24 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.39, #queue-req: 0
[2026-09-10 06:34:25 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.82, #queue-req: 0
[2026-09-10 06:34:25 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.34, #queue-req: 0
[2026-09-10 06:34:26 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.30, #queue-req: 0
[2026-09-10 06:34:26 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.99, #queue-req: 0
[2026-09-10 06:34:26 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.59, #queue-req: 0
[2026-09-10 06:34:27 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.94, #queue-req: 0
[2026-09-10 06:34:27 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.88, #queue-req: 0
[2026-09-10 06:34:28 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.37, #queue-req: 0
[2026-09-10 06:34:28 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.92, #queue-req: 0
[2026-09-10 06:34:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.74, #queue-req: 0
[2026-09-10 06:34:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.37, #queue-req: 0
[2026-09-10 06:34:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.90, #queue-req: 0
[2026-09-10 06:34:30 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.51, #queue-req: 0
[2026-09-10 06:34:30 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.64, #queue-req: 0
[2026-09-10 06:34:31 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.05, #queue-req: 0
[2026-09-10 06:34:31 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.10, #queue-req: 0
[2026-09-10 06:34:32 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.74, #queue-req: 0
[2026-09-10 06:34:32 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.59, #queue-req: 0
[2026-09-10 06:34:32 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.67, #queue-req: 0
[2026-09-10 06:34:33 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.73, #queue-req: 0
[2026-09-10 06:34:33 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.54, #queue-req: 0
[2026-09-10 06:34:34 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.40, #queue-req: 0
[2026-09-10 06:34:34 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.97, #queue-req: 0
[2026-09-10 06:34:35 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.04, #queue-req: 0
[2026-09-10 06:34:35 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.46, #queue-req: 0
[2026-09-10 06:34:35 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.07, #queue-req: 0
[2026-09-10 06:34:36 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.92, #queue-req: 0
[2026-09-10 06:34:36 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.49, #queue-req: 0
[2026-09-10 06:34:37 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.24, #queue-req: 0
[2026-09-10 06:34:37 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.70, #queue-req: 0
[2026-09-10 06:34:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.73, #queue-req: 0
[2026-09-10 06:34:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.29, #queue-req: 0
[2026-09-10 06:34:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.61, #queue-req: 0
[2026-09-10 06:34:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.87, #queue-req: 0
[2026-09-10 06:34:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.14, #queue-req: 0
[2026-09-10 06:34:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.98, #queue-req: 0
[2026-09-10 06:34:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.51, #queue-req: 0
[2026-09-10 06:34:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.62, #queue-req: 0
[2026-09-10 06:34:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.08, #queue-req: 0
[2026-09-10 06:34:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.55, #queue-req: 0
[2026-09-10 06:34:42 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.53, #queue-req: 0
[2026-09-10 06:34:42 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.56, #queue-req: 0
[2026-09-10 06:34:43 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.93, #queue-req: 0
[2026-09-10 06:34:43 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.28, #queue-req: 0
[2026-09-10 06:34:44 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.53, #queue-req: 0
[2026-09-10 06:34:44 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.62, #queue-req: 0
[2026-09-10 06:34:45 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.41, #queue-req: 0
[2026-09-10 06:34:45 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.00, #queue-req: 0
[2026-09-10 06:34:45 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.07, #queue-req: 0
[2026-09-10 06:34:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.77, #queue-req: 0
[2026-09-10 06:34:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.57, #queue-req: 0
[2026-09-10 06:34:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.25, #queue-req: 0
[2026-09-10 06:34:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.84, #queue-req: 0
[2026-09-10 06:34:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.50, #queue-req: 0
[2026-09-10 06:34:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.39, #queue-req: 0
[2026-09-10 06:34:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.37, #queue-req: 0
[2026-09-10 06:34:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.57, #queue-req: 0
[2026-09-10 06:34:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.39, #queue-req: 0
[2026-09-10 06:34:50 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.12, #queue-req: 0
[2026-09-10 06:34:50 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.00, #queue-req: 0
[2026-09-10 06:34:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.26, #queue-req: 0
[2026-09-10 06:34:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.50, #queue-req: 0
[2026-09-10 06:34:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.04, #queue-req: 0
[2026-09-10 06:34:52 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.96, #queue-req: 0
[2026-09-10 06:34:52 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.30, #queue-req: 0
[2026-09-10 06:34:53 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.03, #queue-req: 0
[2026-09-10 06:34:53 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.75, #queue-req: 0
[2026-09-10 06:34:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.30, #queue-req: 0
[2026-09-10 06:34:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.19, #queue-req: 0
[2026-09-10 06:34:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.72, #queue-req: 0
[2026-09-10 06:34:55 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.26, #queue-req: 0
[2026-09-10 06:34:55 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.37, #queue-req: 0
[2026-09-10 06:34:56 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.94, #queue-req: 0
[2026-09-10 06:34:56 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.95, #queue-req: 0
[2026-09-10 06:34:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.30, #queue-req: 0
[2026-09-10 06:34:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.59, #queue-req: 0
[2026-09-10 06:34:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.80, #queue-req: 0
[2026-09-10 06:34:58 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.80, #queue-req: 0
[2026-09-10 06:34:58 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.14, #queue-req: 0
[2026-09-10 06:34:59 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.51, #queue-req: 0
[2026-09-10 06:34:59 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.34, #queue-req: 0
[2026-09-10 06:35:00 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.27, #queue-req: 0
[2026-09-10 06:35:00 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.77, #queue-req: 0
[2026-09-10 06:35:01 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.17, #queue-req: 0
[2026-09-10 06:35:01 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.18, #queue-req: 0
[2026-09-10 06:35:01] INFO: 127.0.0.1:52588 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 06:35:01 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 1024, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 5.72
[2026-09-10 06:35:02 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 0.29, #queue-req: 0
[2026-09-10 06:35:02 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 109.85, #queue-req: 0
[2026-09-10 06:35:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 98.85, #queue-req: 0
[2026-09-10 06:35:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.30, #queue-req: 0
[2026-09-10 06:35:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.74, #queue-req: 0
[2026-09-10 06:35:04 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.12, #queue-req: 0
[2026-09-10 06:35:04 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.79, #queue-req: 0
[2026-09-10 06:35:05 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.26, #queue-req: 0
[2026-09-10 06:35:05 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.94, #queue-req: 0
[2026-09-10 06:35:06 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.01, #queue-req: 0
[2026-09-10 06:35:06 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.90, #queue-req: 0
[2026-09-10 06:35:06 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.88, #queue-req: 0
[2026-09-10 06:35:07 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.23, #queue-req: 0
[2026-09-10 06:35:07 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.51, #queue-req: 0
[2026-09-10 06:35:08 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.30, #queue-req: 0
[2026-09-10 06:35:08 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.36, #queue-req: 0
[2026-09-10 06:35:08 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.24, #queue-req: 0
[2026-09-10 06:35:09 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.20, #queue-req: 0
[2026-09-10 06:35:09 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.66, #queue-req: 0
[2026-09-10 06:35:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.50, #queue-req: 0
[2026-09-10 06:35:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.96, #queue-req: 0
[2026-09-10 06:35:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.51, #queue-req: 0
[2026-09-10 06:35:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.20, #queue-req: 0
[2026-09-10 06:35:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.11, #queue-req: 0
[2026-09-10 06:35:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.99, #queue-req: 0
[2026-09-10 06:35:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.19, #queue-req: 0
[2026-09-10 06:35:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.15, #queue-req: 0
[2026-09-10 06:35:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.04, #queue-req: 0
[2026-09-10 06:35:14 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.83, #queue-req: 0
[2026-09-10 06:35:14 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.67, #queue-req: 0
[2026-09-10 06:35:14 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.77, #queue-req: 0
[2026-09-10 06:35:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.68, #queue-req: 0
[2026-09-10 06:35:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.44, #queue-req: 0
[2026-09-10 06:35:16 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.29, #queue-req: 0
[2026-09-10 06:35:16 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.64, #queue-req: 0
[2026-09-10 06:35:16 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.98, #queue-req: 0
[2026-09-10 06:35:17 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.42, #queue-req: 0
[2026-09-10 06:35:17 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.50, #queue-req: 0
[2026-09-10 06:35:18 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.41, #queue-req: 0
[2026-09-10 06:35:18 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.59, #queue-req: 0
[2026-09-10 06:35:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.34, #queue-req: 0
[2026-09-10 06:35:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.24, #queue-req: 0
[2026-09-10 06:35:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.86, #queue-req: 0
[2026-09-10 06:35:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.89, #queue-req: 0
[2026-09-10 06:35:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.38, #queue-req: 0
[2026-09-10 06:35:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.60, #queue-req: 0
[2026-09-10 06:35:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.38, #queue-req: 0
[2026-09-10 06:35:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.16, #queue-req: 0
[2026-09-10 06:35:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.52, #queue-req: 0
[2026-09-10 06:35:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.44, #queue-req: 0
[2026-09-10 06:35:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.95, #queue-req: 0
[2026-09-10 06:35:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.45, #queue-req: 0
[2026-09-10 06:35:24 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.96, #queue-req: 0
[2026-09-10 06:35:24 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.19, #queue-req: 0
[2026-09-10 06:35:24 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.76, #queue-req: 0
[2026-09-10 06:35:25 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.02, #queue-req: 0
[2026-09-10 06:35:25 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.91, #queue-req: 0
[2026-09-10 06:35:26 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.48, #queue-req: 0
[2026-09-10 06:35:26 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.15, #queue-req: 0
[2026-09-10 06:35:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.59, #queue-req: 0
[2026-09-10 06:35:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.88, #queue-req: 0
[2026-09-10 06:35:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.53, #queue-req: 0
[2026-09-10 06:35:28 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.74, #queue-req: 0
[2026-09-10 06:35:28 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.50, #queue-req: 0
[2026-09-10 06:35:29 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.43, #queue-req: 0
[2026-09-10 06:35:29 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.77, #queue-req: 0
[2026-09-10 06:35:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.26, #queue-req: 0
[2026-09-10 06:35:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.63, #queue-req: 0
[2026-09-10 06:35:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.08, #queue-req: 0
[2026-09-10 06:35:31 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.20, #queue-req: 0
[2026-09-10 06:35:31 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.97, #queue-req: 0
[2026-09-10 06:35:32 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.81, #queue-req: 0
[2026-09-10 06:35:32 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.17, #queue-req: 0
[2026-09-10 06:35:32 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.46, #queue-req: 0
[2026-09-10 06:35:33 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.11, #queue-req: 0
[2026-09-10 06:35:33 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.96, #queue-req: 0
[2026-09-10 06:35:34 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.02, #queue-req: 0
[2026-09-10 06:35:34 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.95, #queue-req: 0
[2026-09-10 06:35:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.12, #queue-req: 0
[2026-09-10 06:35:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.18, #queue-req: 0
[2026-09-10 06:35:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.97, #queue-req: 0
[2026-09-10 06:35:36 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.36, #queue-req: 0
[2026-09-10 06:35:36 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.52, #queue-req: 0
[2026-09-10 06:35:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.87, #queue-req: 0
[2026-09-10 06:35:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.90, #queue-req: 0
[2026-09-10 06:35:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.15, #queue-req: 0
[2026-09-10 06:35:38 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.33, #queue-req: 0
[2026-09-10 06:35:38 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.95, #queue-req: 0
[2026-09-10 06:35:39 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.19, #queue-req: 0
[2026-09-10 06:35:39 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.97, #queue-req: 0
[2026-09-10 06:35:40 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.36, #queue-req: 0
[2026-09-10 06:35:40 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.08, #queue-req: 0
[2026-09-10 06:35:40 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.82, #queue-req: 0
[2026-09-10 06:35:41 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.67, #queue-req: 0
[2026-09-10 06:35:41 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.92, #queue-req: 0
[2026-09-10 06:35:42 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.79, #queue-req: 0
[2026-09-10 06:35:42 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.30, #queue-req: 0
[2026-09-10 06:35:43 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.60, #queue-req: 0
[2026-09-10 06:35:43 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.50, #queue-req: 0
[2026-09-10 06:35:43 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.95, #queue-req: 0
[2026-09-10 06:35:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.25, #queue-req: 0
[2026-09-10 06:35:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.69, #queue-req: 0
[2026-09-10 06:35:44] INFO: 127.0.0.1:46194 - "POST /generate HTTP/1.1" 200 OK
[2026-09-10 06:35:45 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 1024, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 5.72
[2026-09-10 06:35:45 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 0.30, #queue-req: 0
[2026-09-10 06:35:45 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.27, #queue-req: 0
[2026-09-10 06:35:46 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.65, #queue-req: 0
[2026-09-10 06:35:46 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.88, #queue-req: 0
[2026-09-10 06:35:47 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.15, #queue-req: 0
[2026-09-10 06:35:47 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.87, #queue-req: 0
[2026-09-10 06:35:48 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.11, #queue-req: 0
[2026-09-10 06:35:48 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.26, #queue-req: 0
[2026-09-10 06:35:49 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.00, #queue-req: 0
[2026-09-10 06:35:49 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.81, #queue-req: 0
[2026-09-10 06:35:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.91, #queue-req: 0
[2026-09-10 06:35:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.60, #queue-req: 0
[2026-09-10 06:35:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.63, #queue-req: 0
[2026-09-10 06:35:51 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.64, #queue-req: 0
[2026-09-10 06:35:51 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.12, #queue-req: 0
[2026-09-10 06:35:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.53, #queue-req: 0
[2026-09-10 06:35:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.78, #queue-req: 0
[2026-09-10 06:35:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.21, #queue-req: 0
[2026-09-10 06:35:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.99, #queue-req: 0
[2026-09-10 06:35:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.84, #queue-req: 0
[2026-09-10 06:35:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.96, #queue-req: 0
[2026-09-10 06:35:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.41, #queue-req: 0
[2026-09-10 06:35:55 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.42, #queue-req: 0
[2026-09-10 06:35:55 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.24, #queue-req: 0
[2026-09-10 06:35:56 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.83, #queue-req: 0

Some files were not shown because too many files have changed in this diff Show More