[Artifacts] update B300 matrix through 15:23
This commit is contained in:
parent
7984c25586
commit
9652bfdb9d
@ -13,6 +13,13 @@ text-matrix run as of the snapshot time.
|
||||
- Follow-up log: `/data/b300-dsv4-glm53-dev4-20260909-084427-followup.log`
|
||||
- Source archive SHA-256: `8676f28c6b9df7676ad93d5a4a71777bb504539b0ea6e3b129ccec69f5a563fd`
|
||||
|
||||
Incremental synchronization:
|
||||
|
||||
- Delta time: `2026-09-10 07:23:10 UTC` (`2026-09-10 15:23:10 Asia/Shanghai`)
|
||||
- Delta archive SHA-256: `c1bc269866152024afbdefbb15265832a773bfbf07ea2375e318dc252b25bdbb`
|
||||
- At this point DeepSeek-V4-Flash had produced 53 Low-Latency and 18
|
||||
Balanced point JSON files.
|
||||
|
||||
The DeepSeek-V4-Flash matrix was still running when this snapshot was taken.
|
||||
Consequently, this is a complete snapshot of files present at that time, not
|
||||
the final completed run archive. GLM-5.3 had completed its first pass.
|
||||
|
||||
File diff suppressed because one or more lines are too long
@ -1,9 +1,9 @@
|
||||
index, name, memory.total [MiB], memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 183.73 W
|
||||
1, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 181.64 W
|
||||
2, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 184.33 W
|
||||
3, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 186.83 W
|
||||
4, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 184.06 W
|
||||
5, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 182.13 W
|
||||
6, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 183.42 W
|
||||
7, NVIDIA Graphics Device, 275040 MiB, 0 MiB, 274114 MiB, 0 %, 185.93 W
|
||||
0, NVIDIA Graphics Device, 275040 MiB, 259411 MiB, 14703 MiB, 0 %, 242.62 W
|
||||
1, NVIDIA Graphics Device, 275040 MiB, 259603 MiB, 14511 MiB, 0 %, 234.74 W
|
||||
2, NVIDIA Graphics Device, 275040 MiB, 259617 MiB, 14497 MiB, 0 %, 235.85 W
|
||||
3, NVIDIA Graphics Device, 275040 MiB, 258643 MiB, 15471 MiB, 0 %, 243.30 W
|
||||
4, NVIDIA Graphics Device, 275040 MiB, 4 MiB, 274110 MiB, 0 %, 183.98 W
|
||||
5, NVIDIA Graphics Device, 275040 MiB, 4 MiB, 274110 MiB, 0 %, 182.88 W
|
||||
6, NVIDIA Graphics Device, 275040 MiB, 4 MiB, 274110 MiB, 0 %, 184.01 W
|
||||
7, NVIDIA Graphics Device, 275040 MiB, 4 MiB, 274110 MiB, 0 %, 184.09 W
|
||||
|
||||
|
@ -0,0 +1 @@
|
||||
2026-09-10T05:47:02+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 261419 MiB, 12695 MiB, 0 %, 241.89 W
|
||||
1, 261033 MiB, 13081 MiB, 0 %, 234.70 W
|
||||
2, 261481 MiB, 12633 MiB, 0 %, 235.81 W
|
||||
3, 261633 MiB, 12481 MiB, 0 %, 243.99 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 182.40 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.84 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 183.78 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=1, output_file='/results/points/16k_1_c1.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=1, output_file='/results/points/16k_1_c1.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 1048576
|
||||
#Output tokens: 64
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 1
|
||||
Successful requests: 64
|
||||
Benchmark duration (s): 62.77
|
||||
Total input tokens: 1048576
|
||||
Total input text tokens: 1048576
|
||||
Total generated tokens: 64
|
||||
Total generated tokens (retokenized): 64
|
||||
Request throughput (req/s): 1.02
|
||||
Input token throughput (tok/s): 16704.37
|
||||
Output token throughput (tok/s): 1.02
|
||||
Peak output token throughput (tok/s): 2.00
|
||||
Peak concurrent requests: 3
|
||||
Total token throughput (tok/s): 16705.38
|
||||
Concurrency: 1.00
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 979.20
|
||||
Median E2E Latency (ms): 978.74
|
||||
P90 E2E Latency (ms): 998.89
|
||||
P95 E2E Latency (ms): 1001.70
|
||||
P99 E2E Latency (ms): 1041.31
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 979.15
|
||||
Median TTFT (ms): 978.69
|
||||
P90 TTFT (ms): 998.85
|
||||
P95 TTFT (ms): 1001.65
|
||||
P99 TTFT (ms): 1041.27
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 0.00
|
||||
Median TPOT (ms): 0.00
|
||||
P90 TPOT (ms): 0.00
|
||||
P95 TPOT (ms): 0.00
|
||||
P99 TPOT (ms): 0.00
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 0.00
|
||||
Median ITL (ms): 0.00
|
||||
P90 ITL (ms): 0.00
|
||||
P95 ITL (ms): 0.00
|
||||
P99 ITL (ms): 0.00
|
||||
Max ITL (ms): 0.00
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T05:53:09+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 270303 MiB, 3811 MiB, 0 %, 245.71 W
|
||||
1, 267511 MiB, 6603 MiB, 0 %, 236.66 W
|
||||
2, 271933 MiB, 2181 MiB, 0 %, 238.35 W
|
||||
3, 265137 MiB, 8977 MiB, 0 %, 247.74 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 182.01 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.17 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=640, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=128, output_file='/results/points/16k_1_c128.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=640, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=128, output_file='/results/points/16k_1_c128.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 10485760
|
||||
#Output tokens: 640
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 128
|
||||
Successful requests: 640
|
||||
Benchmark duration (s): 161.63
|
||||
Total input tokens: 10485760
|
||||
Total input text tokens: 10485760
|
||||
Total generated tokens: 640
|
||||
Total generated tokens (retokenized): 636
|
||||
Request throughput (req/s): 3.96
|
||||
Input token throughput (tok/s): 64875.87
|
||||
Output token throughput (tok/s): 3.96
|
||||
Peak output token throughput (tok/s): 6.00
|
||||
Peak concurrent requests: 133
|
||||
Total token throughput (tok/s): 64879.83
|
||||
Concurrency: 115.42
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 29149.59
|
||||
Median E2E Latency (ms): 31991.66
|
||||
P90 E2E Latency (ms): 32464.82
|
||||
P95 E2E Latency (ms): 32515.02
|
||||
P99 E2E Latency (ms): 32526.70
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 28973.22
|
||||
Median TTFT (ms): 31988.94
|
||||
P90 TTFT (ms): 32464.77
|
||||
P95 TTFT (ms): 32514.97
|
||||
P99 TTFT (ms): 32526.66
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 0.00
|
||||
Median TPOT (ms): 0.00
|
||||
P90 TPOT (ms): 0.00
|
||||
P95 TPOT (ms): 0.00
|
||||
P99 TPOT (ms): 0.00
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 0.00
|
||||
Median ITL (ms): 0.00
|
||||
P90 ITL (ms): 0.00
|
||||
P95 ITL (ms): 0.00
|
||||
P99 ITL (ms): 0.00
|
||||
Max ITL (ms): 0.00
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T05:58:59+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 269385 MiB, 4729 MiB, 0 %, 244.37 W
|
||||
1, 267211 MiB, 6903 MiB, 0 %, 236.66 W
|
||||
2, 271703 MiB, 2411 MiB, 0 %, 237.25 W
|
||||
3, 265765 MiB, 8349 MiB, 0 %, 247.43 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 182.01 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.25 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=1280, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=256, output_file='/results/points/16k_1_c256.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=1280, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=256, output_file='/results/points/16k_1_c256.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 20971520
|
||||
#Output tokens: 1280
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 256
|
||||
Successful requests: 1280
|
||||
Benchmark duration (s): 326.43
|
||||
Total input tokens: 20971520
|
||||
Total input text tokens: 20971520
|
||||
Total generated tokens: 1280
|
||||
Total generated tokens (retokenized): 1274
|
||||
Request throughput (req/s): 3.92
|
||||
Input token throughput (tok/s): 64245.92
|
||||
Output token throughput (tok/s): 3.92
|
||||
Peak output token throughput (tok/s): 5.00
|
||||
Peak concurrent requests: 260
|
||||
Total token throughput (tok/s): 64249.84
|
||||
Concurrency: 230.13
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 58687.97
|
||||
Median E2E Latency (ms): 64913.32
|
||||
P90 E2E Latency (ms): 65018.57
|
||||
P95 E2E Latency (ms): 65033.21
|
||||
P99 E2E Latency (ms): 65149.17
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 58398.08
|
||||
Median TTFT (ms): 64913.17
|
||||
P90 TTFT (ms): 65018.54
|
||||
P95 TTFT (ms): 65033.17
|
||||
P99 TTFT (ms): 65149.12
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 0.00
|
||||
Median TPOT (ms): 0.00
|
||||
P90 TPOT (ms): 0.00
|
||||
P95 TPOT (ms): 0.00
|
||||
P99 TPOT (ms): 0.00
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 0.00
|
||||
Median ITL (ms): 0.00
|
||||
P90 ITL (ms): 0.00
|
||||
P95 ITL (ms): 0.00
|
||||
P99 ITL (ms): 0.00
|
||||
Max ITL (ms): 0.00
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T05:48:31+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 269243 MiB, 4871 MiB, 0 %, 243.84 W
|
||||
1, 264807 MiB, 9307 MiB, 0 %, 236.54 W
|
||||
2, 268279 MiB, 5835 MiB, 0 %, 236.92 W
|
||||
3, 266383 MiB, 7731 MiB, 0 %, 247.24 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=160, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=32, output_file='/results/points/16k_1_c32.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=160, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=32, output_file='/results/points/16k_1_c32.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 2621440
|
||||
#Output tokens: 160
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 32
|
||||
Successful requests: 160
|
||||
Benchmark duration (s): 40.27
|
||||
Total input tokens: 2621440
|
||||
Total input text tokens: 2621440
|
||||
Total generated tokens: 160
|
||||
Total generated tokens (retokenized): 158
|
||||
Request throughput (req/s): 3.97
|
||||
Input token throughput (tok/s): 65093.70
|
||||
Output token throughput (tok/s): 3.97
|
||||
Peak output token throughput (tok/s): 5.00
|
||||
Peak concurrent requests: 36
|
||||
Total token throughput (tok/s): 65097.67
|
||||
Concurrency: 29.23
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 7356.74
|
||||
Median E2E Latency (ms): 7919.05
|
||||
P90 E2E Latency (ms): 7940.64
|
||||
P95 E2E Latency (ms): 7943.24
|
||||
P99 E2E Latency (ms): 8597.74
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 7257.63
|
||||
Median TTFT (ms): 7918.42
|
||||
P90 TTFT (ms): 7940.60
|
||||
P95 TTFT (ms): 7943.21
|
||||
P99 TTFT (ms): 8597.70
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 0.00
|
||||
Median TPOT (ms): 0.00
|
||||
P90 TPOT (ms): 0.00
|
||||
P95 TPOT (ms): 0.00
|
||||
P99 TPOT (ms): 0.00
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 0.00
|
||||
Median ITL (ms): 0.00
|
||||
P90 ITL (ms): 0.00
|
||||
P95 ITL (ms): 0.00
|
||||
P99 ITL (ms): 0.00
|
||||
Max ITL (ms): 0.00
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T05:50:09+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 270311 MiB, 3803 MiB, 0 %, 245.60 W
|
||||
1, 266159 MiB, 7955 MiB, 0 %, 236.66 W
|
||||
2, 267307 MiB, 6807 MiB, 0 %, 238.36 W
|
||||
3, 264547 MiB, 9567 MiB, 0 %, 247.74 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 182.08 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.37 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=320, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=64, output_file='/results/points/16k_1_c64.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=320, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=64, output_file='/results/points/16k_1_c64.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 5242880
|
||||
#Output tokens: 320
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 64
|
||||
Successful requests: 320
|
||||
Benchmark duration (s): 80.80
|
||||
Total input tokens: 5242880
|
||||
Total input text tokens: 5242880
|
||||
Total generated tokens: 320
|
||||
Total generated tokens (retokenized): 318
|
||||
Request throughput (req/s): 3.96
|
||||
Input token throughput (tok/s): 64887.20
|
||||
Output token throughput (tok/s): 3.96
|
||||
Peak output token throughput (tok/s): 6.00
|
||||
Peak concurrent requests: 68
|
||||
Total token throughput (tok/s): 64891.16
|
||||
Concurrency: 58.04
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 14654.14
|
||||
Median E2E Latency (ms): 15916.11
|
||||
P90 E2E Latency (ms): 16382.05
|
||||
P95 E2E Latency (ms): 16398.86
|
||||
P99 E2E Latency (ms): 16407.45
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 14554.54
|
||||
Median TTFT (ms): 15915.15
|
||||
P90 TTFT (ms): 16382.00
|
||||
P95 TTFT (ms): 16398.83
|
||||
P99 TTFT (ms): 16407.41
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 0.00
|
||||
Median TPOT (ms): 0.00
|
||||
P90 TPOT (ms): 0.00
|
||||
P95 TPOT (ms): 0.00
|
||||
P99 TPOT (ms): 0.00
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 0.00
|
||||
Median ITL (ms): 0.00
|
||||
P90 ITL (ms): 0.00
|
||||
P95 ITL (ms): 0.00
|
||||
P99 ITL (ms): 0.00
|
||||
Max ITL (ms): 0.00
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T05:47:34+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 269233 MiB, 4881 MiB, 0 %, 243.84 W
|
||||
1, 263439 MiB, 10675 MiB, 0 %, 236.01 W
|
||||
2, 268945 MiB, 5169 MiB, 0 %, 237.17 W
|
||||
3, 265381 MiB, 8733 MiB, 0 %, 245.79 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 182.52 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.99 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 183.85 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=8, output_file='/results/points/16k_1_c8.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=1, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=8, output_file='/results/points/16k_1_c8.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 1048576
|
||||
#Output tokens: 64
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 8
|
||||
Successful requests: 64
|
||||
Benchmark duration (s): 16.41
|
||||
Total input tokens: 1048576
|
||||
Total input text tokens: 1048576
|
||||
Total generated tokens: 64
|
||||
Total generated tokens (retokenized): 64
|
||||
Request throughput (req/s): 3.90
|
||||
Input token throughput (tok/s): 63912.26
|
||||
Output token throughput (tok/s): 3.90
|
||||
Peak output token throughput (tok/s): 4.00
|
||||
Peak concurrent requests: 12
|
||||
Total token throughput (tok/s): 63916.16
|
||||
Concurrency: 7.76
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 1988.23
|
||||
Median E2E Latency (ms): 1986.87
|
||||
P90 E2E Latency (ms): 1993.74
|
||||
P95 E2E Latency (ms): 2334.31
|
||||
P99 E2E Latency (ms): 2638.62
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 1988.19
|
||||
Median TTFT (ms): 1986.83
|
||||
P90 TTFT (ms): 1993.70
|
||||
P95 TTFT (ms): 2334.27
|
||||
P99 TTFT (ms): 2638.58
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 0.00
|
||||
Median TPOT (ms): 0.00
|
||||
P90 TPOT (ms): 0.00
|
||||
P95 TPOT (ms): 0.00
|
||||
P99 TPOT (ms): 0.00
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 0.00
|
||||
Median ITL (ms): 0.00
|
||||
P90 ITL (ms): 0.00
|
||||
P95 ITL (ms): 0.00
|
||||
P99 ITL (ms): 0.00
|
||||
Max ITL (ms): 0.00
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T05:30:32+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 261315 MiB, 12799 MiB, 0 %, 240.90 W
|
||||
1, 260915 MiB, 13199 MiB, 0 %, 234.32 W
|
||||
2, 261347 MiB, 12767 MiB, 0 %, 234.96 W
|
||||
3, 261503 MiB, 12611 MiB, 0 %, 243.99 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 183.47 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.22 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 183.97 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=1, output_file='/results/points/16k_512_c1.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=1, output_file='/results/points/16k_512_c1.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 1048576
|
||||
#Output tokens: 32768
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 1
|
||||
Successful requests: 64
|
||||
Benchmark duration (s): 416.51
|
||||
Total input tokens: 1048576
|
||||
Total input text tokens: 1048576
|
||||
Total generated tokens: 32768
|
||||
Total generated tokens (retokenized): 32447
|
||||
Request throughput (req/s): 0.15
|
||||
Input token throughput (tok/s): 2517.53
|
||||
Output token throughput (tok/s): 78.67
|
||||
Peak output token throughput (tok/s): 108.00
|
||||
Peak concurrent requests: 2
|
||||
Total token throughput (tok/s): 2596.20
|
||||
Concurrency: 1.00
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 6506.17
|
||||
Median E2E Latency (ms): 6469.01
|
||||
P90 E2E Latency (ms): 6740.62
|
||||
P95 E2E Latency (ms): 6816.94
|
||||
P99 E2E Latency (ms): 6875.99
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 1044.49
|
||||
Median TTFT (ms): 1043.08
|
||||
P90 TTFT (ms): 1065.97
|
||||
P95 TTFT (ms): 1067.18
|
||||
P99 TTFT (ms): 1078.33
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 10.69
|
||||
Median TPOT (ms): 10.60
|
||||
P90 TPOT (ms): 11.18
|
||||
P95 TPOT (ms): 11.26
|
||||
P99 TPOT (ms): 11.38
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 10.69
|
||||
Median ITL (ms): 10.82
|
||||
P90 ITL (ms): 11.59
|
||||
P95 ITL (ms): 11.82
|
||||
P99 ITL (ms): 12.29
|
||||
Max ITL (ms): 16.37
|
||||
==================================================
|
||||
@ -0,0 +1,817 @@
|
||||
[2026-09-10 05:25:31 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.78
|
||||
[2026-09-10 05:25:31 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17796.50
|
||||
[2026-09-10 05:25:31 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17815.89
|
||||
[2026-09-10 05:25:31 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 376057.58
|
||||
[2026-09-10 05:25:31 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
|
||||
[2026-09-10 05:25:32 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.03, #queue-req: 0
|
||||
[2026-09-10 05:25:32 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.86, #queue-req: 0
|
||||
[2026-09-10 05:25:33 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.73, #queue-req: 0
|
||||
[2026-09-10 05:25:33 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.06, #queue-req: 0
|
||||
[2026-09-10 05:25:34 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.13, #queue-req: 0
|
||||
[2026-09-10 05:25:34 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.50, #queue-req: 0
|
||||
[2026-09-10 05:25:35 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.06, #queue-req: 0
|
||||
[2026-09-10 05:25:35 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.82, #queue-req: 0
|
||||
[2026-09-10 05:25:35 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.64, #queue-req: 0
|
||||
[2026-09-10 05:25:36 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.21, #queue-req: 0
|
||||
[2026-09-10 05:25:36 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.39, #queue-req: 0
|
||||
[2026-09-10 05:25:37 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.66, #queue-req: 0
|
||||
[2026-09-10 05:25:37] INFO: 127.0.0.1:44620 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:25:37 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.63
|
||||
[2026-09-10 05:25:38 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17725.79
|
||||
[2026-09-10 05:25:38 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17725.44
|
||||
[2026-09-10 05:25:38 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 394303.47
|
||||
[2026-09-10 05:25:38 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
|
||||
[2026-09-10 05:25:38 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.84, #queue-req: 0
|
||||
[2026-09-10 05:25:39 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 103.15, #queue-req: 0
|
||||
[2026-09-10 05:25:39 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.24, #queue-req: 0
|
||||
[2026-09-10 05:25:40 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.48, #queue-req: 0
|
||||
[2026-09-10 05:25:40 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.04, #queue-req: 0
|
||||
[2026-09-10 05:25:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.90, #queue-req: 0
|
||||
[2026-09-10 05:25:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.07, #queue-req: 0
|
||||
[2026-09-10 05:25:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.71, #queue-req: 0
|
||||
[2026-09-10 05:25:42 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.81, #queue-req: 0
|
||||
[2026-09-10 05:25:42 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.67, #queue-req: 0
|
||||
[2026-09-10 05:25:43 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.06, #queue-req: 0
|
||||
[2026-09-10 05:25:43 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.96, #queue-req: 0
|
||||
[2026-09-10 05:25:43] INFO: 127.0.0.1:51576 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:25:44 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.43
|
||||
[2026-09-10 05:25:44 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 16995.45
|
||||
[2026-09-10 05:25:44 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 16964.93
|
||||
[2026-09-10 05:25:44 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 379714.50
|
||||
[2026-09-10 05:25:44 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
|
||||
[2026-09-10 05:25:45 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 102.63, #queue-req: 0
|
||||
[2026-09-10 05:25:45 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.93, #queue-req: 0
|
||||
[2026-09-10 05:25:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.32, #queue-req: 0
|
||||
[2026-09-10 05:25:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.12, #queue-req: 0
|
||||
[2026-09-10 05:25:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.60, #queue-req: 0
|
||||
[2026-09-10 05:25:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.87, #queue-req: 0
|
||||
[2026-09-10 05:25:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.95, #queue-req: 0
|
||||
[2026-09-10 05:25:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.35, #queue-req: 0
|
||||
[2026-09-10 05:25:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.68, #queue-req: 0
|
||||
[2026-09-10 05:25:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.43, #queue-req: 0
|
||||
[2026-09-10 05:25:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.30, #queue-req: 0
|
||||
[2026-09-10 05:25:50] INFO: 127.0.0.1:59534 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:25:50 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.33
|
||||
[2026-09-10 05:25:50 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17384.11
|
||||
[2026-09-10 05:25:51 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17358.92
|
||||
[2026-09-10 05:25:51 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 393361.18
|
||||
[2026-09-10 05:25:51 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.92, #queue-req: 0
|
||||
[2026-09-10 05:25:51 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 104.68, #queue-req: 0
|
||||
[2026-09-10 05:25:52 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.05, #queue-req: 0
|
||||
[2026-09-10 05:25:52 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.47, #queue-req: 0
|
||||
[2026-09-10 05:25:53 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.94, #queue-req: 0
|
||||
[2026-09-10 05:25:53 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.10, #queue-req: 0
|
||||
[2026-09-10 05:25:53 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.51, #queue-req: 0
|
||||
[2026-09-10 05:25:54 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.33, #queue-req: 0
|
||||
[2026-09-10 05:25:54 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.94, #queue-req: 0
|
||||
[2026-09-10 05:25:55 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.56, #queue-req: 0
|
||||
[2026-09-10 05:25:55 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.49, #queue-req: 0
|
||||
[2026-09-10 05:25:56 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.70, #queue-req: 0
|
||||
[2026-09-10 05:25:56] INFO: 127.0.0.1:59544 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:25:56 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.89
|
||||
[2026-09-10 05:25:57 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17506.33
|
||||
[2026-09-10 05:25:57 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17698.03
|
||||
[2026-09-10 05:25:57 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 392426.55
|
||||
[2026-09-10 05:25:57 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.95, #queue-req: 0
|
||||
[2026-09-10 05:25:58 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.34, #queue-req: 0
|
||||
[2026-09-10 05:25:58 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.02, #queue-req: 0
|
||||
[2026-09-10 05:25:59 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.99, #queue-req: 0
|
||||
[2026-09-10 05:25:59 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.44, #queue-req: 0
|
||||
[2026-09-10 05:26:00 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.45, #queue-req: 0
|
||||
[2026-09-10 05:26:00 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.92, #queue-req: 0
|
||||
[2026-09-10 05:26:00 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.80, #queue-req: 0
|
||||
[2026-09-10 05:26:01 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.58, #queue-req: 0
|
||||
[2026-09-10 05:26:01 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.71, #queue-req: 0
|
||||
[2026-09-10 05:26:02 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.90, #queue-req: 0
|
||||
[2026-09-10 05:26:02 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.25, #queue-req: 0
|
||||
[2026-09-10 05:26:03] INFO: 127.0.0.1:36546 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:26:03 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.48
|
||||
[2026-09-10 05:26:03 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17221.91
|
||||
[2026-09-10 05:26:04 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17618.48
|
||||
[2026-09-10 05:26:04 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 410228.80
|
||||
[2026-09-10 05:26:04 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
|
||||
[2026-09-10 05:26:04 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 108.85, #queue-req: 0
|
||||
[2026-09-10 05:26:05 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.93, #queue-req: 0
|
||||
[2026-09-10 05:26:05 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.49, #queue-req: 0
|
||||
[2026-09-10 05:26:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.88, #queue-req: 0
|
||||
[2026-09-10 05:26:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.41, #queue-req: 0
|
||||
[2026-09-10 05:26:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.71, #queue-req: 0
|
||||
[2026-09-10 05:26:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.55, #queue-req: 0
|
||||
[2026-09-10 05:26:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.97, #queue-req: 0
|
||||
[2026-09-10 05:26:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.38, #queue-req: 0
|
||||
[2026-09-10 05:26:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.10, #queue-req: 0
|
||||
[2026-09-10 05:26:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.10, #queue-req: 0
|
||||
[2026-09-10 05:26:09] INFO: 127.0.0.1:47830 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:26:10 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.00
|
||||
[2026-09-10 05:26:10 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 16721.27
|
||||
[2026-09-10 05:26:10 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 16889.70
|
||||
[2026-09-10 05:26:10 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 368892.56
|
||||
[2026-09-10 05:26:10 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
|
||||
[2026-09-10 05:26:10 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.43, #queue-req: 0
|
||||
[2026-09-10 05:26:11 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.04, #queue-req: 0
|
||||
[2026-09-10 05:26:11 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.33, #queue-req: 0
|
||||
[2026-09-10 05:26:12 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.55, #queue-req: 0
|
||||
[2026-09-10 05:26:12 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.56, #queue-req: 0
|
||||
[2026-09-10 05:26:13 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.67, #queue-req: 0
|
||||
[2026-09-10 05:26:13 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.06, #queue-req: 0
|
||||
[2026-09-10 05:26:14 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.12, #queue-req: 0
|
||||
[2026-09-10 05:26:14 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.92, #queue-req: 0
|
||||
[2026-09-10 05:26:14 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.26, #queue-req: 0
|
||||
[2026-09-10 05:26:15 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.30, #queue-req: 0
|
||||
[2026-09-10 05:26:15 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.94, #queue-req: 0
|
||||
[2026-09-10 05:26:16] INFO: 127.0.0.1:47846 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:26:16 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.81
|
||||
[2026-09-10 05:26:16 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17033.75
|
||||
[2026-09-10 05:26:17 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17310.55
|
||||
[2026-09-10 05:26:17 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 384089.58
|
||||
[2026-09-10 05:26:17 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
|
||||
[2026-09-10 05:26:17 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 112.77, #queue-req: 0
|
||||
[2026-09-10 05:26:17 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 100.14, #queue-req: 0
|
||||
[2026-09-10 05:26:18 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.95, #queue-req: 0
|
||||
[2026-09-10 05:26:18 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.95, #queue-req: 0
|
||||
[2026-09-10 05:26:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.64, #queue-req: 0
|
||||
[2026-09-10 05:26:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.10, #queue-req: 0
|
||||
[2026-09-10 05:26:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.87, #queue-req: 0
|
||||
[2026-09-10 05:26:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.80, #queue-req: 0
|
||||
[2026-09-10 05:26:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.44, #queue-req: 0
|
||||
[2026-09-10 05:26:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.71, #queue-req: 0
|
||||
[2026-09-10 05:26:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.03, #queue-req: 0
|
||||
[2026-09-10 05:26:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.32, #queue-req: 0
|
||||
[2026-09-10 05:26:22] INFO: 127.0.0.1:40652 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:26:22 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.46
|
||||
[2026-09-10 05:26:23 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17864.62
|
||||
[2026-09-10 05:26:23 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17840.13
|
||||
[2026-09-10 05:26:23 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 417235.83
|
||||
[2026-09-10 05:26:23 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
|
||||
[2026-09-10 05:26:23 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 110.66, #queue-req: 0
|
||||
[2026-09-10 05:26:24 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.17, #queue-req: 0
|
||||
[2026-09-10 05:26:24 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.25, #queue-req: 0
|
||||
[2026-09-10 05:26:25 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.47, #queue-req: 0
|
||||
[2026-09-10 05:26:25 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.30, #queue-req: 0
|
||||
[2026-09-10 05:26:26 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.61, #queue-req: 0
|
||||
[2026-09-10 05:26:26 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.22, #queue-req: 0
|
||||
[2026-09-10 05:26:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.35, #queue-req: 0
|
||||
[2026-09-10 05:26:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.51, #queue-req: 0
|
||||
[2026-09-10 05:26:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.82, #queue-req: 0
|
||||
[2026-09-10 05:26:28 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.67, #queue-req: 0
|
||||
[2026-09-10 05:26:28 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.28, #queue-req: 0
|
||||
[2026-09-10 05:26:29] INFO: 127.0.0.1:44996 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:26:29 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.77
|
||||
[2026-09-10 05:26:29 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17803.49
|
||||
[2026-09-10 05:26:30 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17766.36
|
||||
[2026-09-10 05:26:30 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 412728.85
|
||||
[2026-09-10 05:26:30 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
|
||||
[2026-09-10 05:26:30 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.84, #queue-req: 0
|
||||
[2026-09-10 05:26:30 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 109.76, #queue-req: 0
|
||||
[2026-09-10 05:26:31 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 100.04, #queue-req: 0
|
||||
[2026-09-10 05:26:31 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.17, #queue-req: 0
|
||||
[2026-09-10 05:26:32 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.54, #queue-req: 0
|
||||
[2026-09-10 05:26:32 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.49, #queue-req: 0
|
||||
[2026-09-10 05:26:32 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.87, #queue-req: 0
|
||||
[2026-09-10 05:26:33 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.58, #queue-req: 0
|
||||
[2026-09-10 05:26:33 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.82, #queue-req: 0
|
||||
[2026-09-10 05:26:34 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.60, #queue-req: 0
|
||||
[2026-09-10 05:26:34 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.12, #queue-req: 0
|
||||
[2026-09-10 05:26:35 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.73, #queue-req: 0
|
||||
[2026-09-10 05:26:35] INFO: 127.0.0.1:45008 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:26:35 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.63
|
||||
[2026-09-10 05:26:36 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17430.73
|
||||
[2026-09-10 05:26:36 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17464.24
|
||||
[2026-09-10 05:26:36 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 397323.61
|
||||
[2026-09-10 05:26:36 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
|
||||
[2026-09-10 05:26:36 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 114.44, #queue-req: 0
|
||||
[2026-09-10 05:26:37 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.90, #queue-req: 0
|
||||
[2026-09-10 05:26:37 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 100.25, #queue-req: 0
|
||||
[2026-09-10 05:26:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.05, #queue-req: 0
|
||||
[2026-09-10 05:26:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.03, #queue-req: 0
|
||||
[2026-09-10 05:26:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.31, #queue-req: 0
|
||||
[2026-09-10 05:26:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.02, #queue-req: 0
|
||||
[2026-09-10 05:26:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.51, #queue-req: 0
|
||||
[2026-09-10 05:26:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.85, #queue-req: 0
|
||||
[2026-09-10 05:26:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.22, #queue-req: 0
|
||||
[2026-09-10 05:26:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.41, #queue-req: 0
|
||||
[2026-09-10 05:26:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.27, #queue-req: 0
|
||||
[2026-09-10 05:26:41] INFO: 127.0.0.1:41330 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:26:42 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 162.52
|
||||
[2026-09-10 05:26:42 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17395.21
|
||||
[2026-09-10 05:26:42 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17358.20
|
||||
[2026-09-10 05:26:42 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 392883.77
|
||||
[2026-09-10 05:26:42 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
|
||||
[2026-09-10 05:26:43 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 112.70, #queue-req: 0
|
||||
[2026-09-10 05:26:43 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.53, #queue-req: 0
|
||||
[2026-09-10 05:26:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.32, #queue-req: 0
|
||||
[2026-09-10 05:26:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.61, #queue-req: 0
|
||||
[2026-09-10 05:26:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.51, #queue-req: 0
|
||||
[2026-09-10 05:26:45 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.15, #queue-req: 0
|
||||
[2026-09-10 05:26:45 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.52, #queue-req: 0
|
||||
[2026-09-10 05:26:46 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.59, #queue-req: 0
|
||||
[2026-09-10 05:26:46 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.05, #queue-req: 0
|
||||
[2026-09-10 05:26:47 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.99, #queue-req: 0
|
||||
[2026-09-10 05:26:47 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.42, #queue-req: 0
|
||||
[2026-09-10 05:26:47 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.30, #queue-req: 0
|
||||
[2026-09-10 05:26:48] INFO: 127.0.0.1:41346 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:26:48 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 162.45
|
||||
[2026-09-10 05:26:48 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 16997.55
|
||||
[2026-09-10 05:26:49 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17005.57
|
||||
[2026-09-10 05:26:49 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 412761.43
|
||||
[2026-09-10 05:26:49 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.96, #queue-req: 0
|
||||
[2026-09-10 05:26:49 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 106.99, #queue-req: 0
|
||||
[2026-09-10 05:26:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.87, #queue-req: 0
|
||||
[2026-09-10 05:26:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.10, #queue-req: 0
|
||||
[2026-09-10 05:26:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.71, #queue-req: 0
|
||||
[2026-09-10 05:26:51 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.29, #queue-req: 0
|
||||
[2026-09-10 05:26:51 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.48, #queue-req: 0
|
||||
[2026-09-10 05:26:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.68, #queue-req: 0
|
||||
[2026-09-10 05:26:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.94, #queue-req: 0
|
||||
[2026-09-10 05:26:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.10, #queue-req: 0
|
||||
[2026-09-10 05:26:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.83, #queue-req: 0
|
||||
[2026-09-10 05:26:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.42, #queue-req: 0
|
||||
[2026-09-10 05:26:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.96, #queue-req: 0
|
||||
[2026-09-10 05:26:54] INFO: 127.0.0.1:36030 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:26:55 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 162.39
|
||||
[2026-09-10 05:26:55 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17713.01
|
||||
[2026-09-10 05:26:55 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17792.55
|
||||
[2026-09-10 05:26:55 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 421200.80
|
||||
[2026-09-10 05:26:55 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.92, #queue-req: 0
|
||||
[2026-09-10 05:26:56 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.94, #queue-req: 0
|
||||
[2026-09-10 05:26:56 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 110.85, #queue-req: 0
|
||||
[2026-09-10 05:26:57 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 101.09, #queue-req: 0
|
||||
[2026-09-10 05:26:57 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.24, #queue-req: 0
|
||||
[2026-09-10 05:26:57 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.28, #queue-req: 0
|
||||
[2026-09-10 05:26:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.14, #queue-req: 0
|
||||
[2026-09-10 05:26:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.69, #queue-req: 0
|
||||
[2026-09-10 05:26:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.70, #queue-req: 0
|
||||
[2026-09-10 05:26:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.50, #queue-req: 0
|
||||
[2026-09-10 05:27:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.11, #queue-req: 0
|
||||
[2026-09-10 05:27:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.22, #queue-req: 0
|
||||
[2026-09-10 05:27:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.99, #queue-req: 0
|
||||
[2026-09-10 05:27:01] INFO: 127.0.0.1:48624 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:27:01 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 162.38
|
||||
[2026-09-10 05:27:01 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17377.25
|
||||
[2026-09-10 05:27:02 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17445.25
|
||||
[2026-09-10 05:27:02 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 369621.82
|
||||
[2026-09-10 05:27:02 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
|
||||
[2026-09-10 05:27:02 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 103.91, #queue-req: 0
|
||||
[2026-09-10 05:27:03 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.59, #queue-req: 0
|
||||
[2026-09-10 05:27:03 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.65, #queue-req: 0
|
||||
[2026-09-10 05:27:04 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.30, #queue-req: 0
|
||||
[2026-09-10 05:27:04 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.02, #queue-req: 0
|
||||
[2026-09-10 05:27:04 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.05, #queue-req: 0
|
||||
[2026-09-10 05:27:05 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.68, #queue-req: 0
|
||||
[2026-09-10 05:27:05 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.26, #queue-req: 0
|
||||
[2026-09-10 05:27:06 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.65, #queue-req: 0
|
||||
[2026-09-10 05:27:06 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.81, #queue-req: 0
|
||||
[2026-09-10 05:27:07 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.06, #queue-req: 0
|
||||
[2026-09-10 05:27:07 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.70, #queue-req: 0
|
||||
[2026-09-10 05:27:07] INFO: 127.0.0.1:48628 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:27:08 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.17
|
||||
[2026-09-10 05:27:08 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17333.87
|
||||
[2026-09-10 05:27:08 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17041.99
|
||||
[2026-09-10 05:27:08 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 379084.60
|
||||
[2026-09-10 05:27:08 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
|
||||
[2026-09-10 05:27:09 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 107.04, #queue-req: 0
|
||||
[2026-09-10 05:27:09 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.07, #queue-req: 0
|
||||
[2026-09-10 05:27:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.71, #queue-req: 0
|
||||
[2026-09-10 05:27:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.69, #queue-req: 0
|
||||
[2026-09-10 05:27:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.16, #queue-req: 0
|
||||
[2026-09-10 05:27:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.55, #queue-req: 0
|
||||
[2026-09-10 05:27:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.45, #queue-req: 0
|
||||
[2026-09-10 05:27:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.56, #queue-req: 0
|
||||
[2026-09-10 05:27:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.48, #queue-req: 0
|
||||
[2026-09-10 05:27:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.22, #queue-req: 0
|
||||
[2026-09-10 05:27:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.72, #queue-req: 0
|
||||
[2026-09-10 05:27:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.86, #queue-req: 0
|
||||
[2026-09-10 05:27:14] INFO: 127.0.0.1:37054 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:27:14 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.58
|
||||
[2026-09-10 05:27:14 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17402.19
|
||||
[2026-09-10 05:27:15 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17381.88
|
||||
[2026-09-10 05:27:15 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 409661.08
|
||||
[2026-09-10 05:27:15 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
|
||||
[2026-09-10 05:27:15 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 101.85, #queue-req: 0
|
||||
[2026-09-10 05:27:16 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.27, #queue-req: 0
|
||||
[2026-09-10 05:27:16 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.18, #queue-req: 0
|
||||
[2026-09-10 05:27:17 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.41, #queue-req: 0
|
||||
[2026-09-10 05:27:17 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.62, #queue-req: 0
|
||||
[2026-09-10 05:27:17 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.17, #queue-req: 0
|
||||
[2026-09-10 05:27:18 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.24, #queue-req: 0
|
||||
[2026-09-10 05:27:18 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.01, #queue-req: 0
|
||||
[2026-09-10 05:27:19 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.40, #queue-req: 0
|
||||
[2026-09-10 05:27:19 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.47, #queue-req: 0
|
||||
[2026-09-10 05:27:20 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.46, #queue-req: 0
|
||||
[2026-09-10 05:27:20 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.96, #queue-req: 0
|
||||
[2026-09-10 05:27:20] INFO: 127.0.0.1:35934 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:27:21 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.40
|
||||
[2026-09-10 05:27:21 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17616.27
|
||||
[2026-09-10 05:27:21 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17663.97
|
||||
[2026-09-10 05:27:21 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 418037.71
|
||||
[2026-09-10 05:27:22 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
|
||||
[2026-09-10 05:27:22 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 112.73, #queue-req: 0
|
||||
[2026-09-10 05:27:22 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 101.37, #queue-req: 0
|
||||
[2026-09-10 05:27:23 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.16, #queue-req: 0
|
||||
[2026-09-10 05:27:23 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.66, #queue-req: 0
|
||||
[2026-09-10 05:27:24 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.14, #queue-req: 0
|
||||
[2026-09-10 05:27:24 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.76, #queue-req: 0
|
||||
[2026-09-10 05:27:24 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.94, #queue-req: 0
|
||||
[2026-09-10 05:27:25 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.04, #queue-req: 0
|
||||
[2026-09-10 05:27:25 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.23, #queue-req: 0
|
||||
[2026-09-10 05:27:26 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.62, #queue-req: 0
|
||||
[2026-09-10 05:27:26 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.45, #queue-req: 0
|
||||
[2026-09-10 05:27:27 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.76, #queue-req: 0
|
||||
[2026-09-10 05:27:27] INFO: 127.0.0.1:35946 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:27:27 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.78
|
||||
[2026-09-10 05:27:27 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17369.58
|
||||
[2026-09-10 05:27:28 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17362.70
|
||||
[2026-09-10 05:27:28 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 380751.03
|
||||
[2026-09-10 05:27:28 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
|
||||
[2026-09-10 05:27:28 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 105.20, #queue-req: 0
|
||||
[2026-09-10 05:27:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.34, #queue-req: 0
|
||||
[2026-09-10 05:27:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.37, #queue-req: 0
|
||||
[2026-09-10 05:27:30 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.55, #queue-req: 0
|
||||
[2026-09-10 05:27:30 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.73, #queue-req: 0
|
||||
[2026-09-10 05:27:31 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.76, #queue-req: 0
|
||||
[2026-09-10 05:27:31 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.96, #queue-req: 0
|
||||
[2026-09-10 05:27:31 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.05, #queue-req: 0
|
||||
[2026-09-10 05:27:32 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.21, #queue-req: 0
|
||||
[2026-09-10 05:27:32 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.10, #queue-req: 0
|
||||
[2026-09-10 05:27:33 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.13, #queue-req: 0
|
||||
[2026-09-10 05:27:33 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.48, #queue-req: 0
|
||||
[2026-09-10 05:27:33] INFO: 127.0.0.1:39016 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:27:34 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.09
|
||||
[2026-09-10 05:27:34 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17363.73
|
||||
[2026-09-10 05:27:34 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17411.52
|
||||
[2026-09-10 05:27:34 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 389176.04
|
||||
[2026-09-10 05:27:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
|
||||
[2026-09-10 05:27:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 109.81, #queue-req: 0
|
||||
[2026-09-10 05:27:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.05, #queue-req: 0
|
||||
[2026-09-10 05:27:36 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.08, #queue-req: 0
|
||||
[2026-09-10 05:27:36 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.08, #queue-req: 0
|
||||
[2026-09-10 05:27:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.84, #queue-req: 0
|
||||
[2026-09-10 05:27:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.60, #queue-req: 0
|
||||
[2026-09-10 05:27:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.85, #queue-req: 0
|
||||
[2026-09-10 05:27:38 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.69, #queue-req: 0
|
||||
[2026-09-10 05:27:38 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.54, #queue-req: 0
|
||||
[2026-09-10 05:27:39 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.92, #queue-req: 0
|
||||
[2026-09-10 05:27:39 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.63, #queue-req: 0
|
||||
[2026-09-10 05:27:40 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.11, #queue-req: 0
|
||||
[2026-09-10 05:27:40] INFO: 127.0.0.1:46422 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:27:40 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.69
|
||||
[2026-09-10 05:27:40 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17021.02
|
||||
[2026-09-10 05:27:41 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 16924.05
|
||||
[2026-09-10 05:27:41 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 394127.88
|
||||
[2026-09-10 05:27:41 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
|
||||
[2026-09-10 05:27:41 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 98.08, #queue-req: 0
|
||||
[2026-09-10 05:27:42 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.83, #queue-req: 0
|
||||
[2026-09-10 05:27:42 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.91, #queue-req: 0
|
||||
[2026-09-10 05:27:43 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.48, #queue-req: 0
|
||||
[2026-09-10 05:27:43 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.25, #queue-req: 0
|
||||
[2026-09-10 05:27:44 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.29, #queue-req: 0
|
||||
[2026-09-10 05:27:44 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.75, #queue-req: 0
|
||||
[2026-09-10 05:27:44 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.24, #queue-req: 0
|
||||
[2026-09-10 05:27:45 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.23, #queue-req: 0
|
||||
[2026-09-10 05:27:45 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.46, #queue-req: 0
|
||||
[2026-09-10 05:27:46 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.35, #queue-req: 0
|
||||
[2026-09-10 05:27:46 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.69, #queue-req: 0
|
||||
[2026-09-10 05:27:46] INFO: 127.0.0.1:46432 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:27:47 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.41
|
||||
[2026-09-10 05:27:47 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17578.22
|
||||
[2026-09-10 05:27:47 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17677.17
|
||||
[2026-09-10 05:27:47 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 396939.69
|
||||
[2026-09-10 05:27:48 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
|
||||
[2026-09-10 05:27:48 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 112.61, #queue-req: 0
|
||||
[2026-09-10 05:27:48 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 99.63, #queue-req: 0
|
||||
[2026-09-10 05:27:49 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.89, #queue-req: 0
|
||||
[2026-09-10 05:27:49 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.06, #queue-req: 0
|
||||
[2026-09-10 05:27:50 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.47, #queue-req: 0
|
||||
[2026-09-10 05:27:50 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.00, #queue-req: 0
|
||||
[2026-09-10 05:27:50 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.13, #queue-req: 0
|
||||
[2026-09-10 05:27:51 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.45, #queue-req: 0
|
||||
[2026-09-10 05:27:51 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.03, #queue-req: 0
|
||||
[2026-09-10 05:27:52 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.70, #queue-req: 0
|
||||
[2026-09-10 05:27:52 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.27, #queue-req: 0
|
||||
[2026-09-10 05:27:53 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.01, #queue-req: 0
|
||||
[2026-09-10 05:27:53] INFO: 127.0.0.1:50828 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:27:53 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.07
|
||||
[2026-09-10 05:27:53 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17228.48
|
||||
[2026-09-10 05:27:54 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17207.25
|
||||
[2026-09-10 05:27:54 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 410922.47
|
||||
[2026-09-10 05:27:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.92, #queue-req: 0
|
||||
[2026-09-10 05:27:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 99.44, #queue-req: 0
|
||||
[2026-09-10 05:27:55 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.80, #queue-req: 0
|
||||
[2026-09-10 05:27:55 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.61, #queue-req: 0
|
||||
[2026-09-10 05:27:56 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.42, #queue-req: 0
|
||||
[2026-09-10 05:27:56 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.16, #queue-req: 0
|
||||
[2026-09-10 05:27:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.03, #queue-req: 0
|
||||
[2026-09-10 05:27:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.03, #queue-req: 0
|
||||
[2026-09-10 05:27:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.96, #queue-req: 0
|
||||
[2026-09-10 05:27:58 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.27, #queue-req: 0
|
||||
[2026-09-10 05:27:58 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.32, #queue-req: 0
|
||||
[2026-09-10 05:27:59 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.85, #queue-req: 0
|
||||
[2026-09-10 05:27:59] INFO: 127.0.0.1:41758 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:28:00 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.31
|
||||
[2026-09-10 05:28:00 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17338.35
|
||||
[2026-09-10 05:28:00 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17315.81
|
||||
[2026-09-10 05:28:00 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 391036.78
|
||||
[2026-09-10 05:28:00 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
|
||||
[2026-09-10 05:28:01 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 105.44, #queue-req: 0
|
||||
[2026-09-10 05:28:01 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.51, #queue-req: 0
|
||||
[2026-09-10 05:28:02 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.12, #queue-req: 0
|
||||
[2026-09-10 05:28:02 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.85, #queue-req: 0
|
||||
[2026-09-10 05:28:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.30, #queue-req: 0
|
||||
[2026-09-10 05:28:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.13, #queue-req: 0
|
||||
[2026-09-10 05:28:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.22, #queue-req: 0
|
||||
[2026-09-10 05:28:04 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.75, #queue-req: 0
|
||||
[2026-09-10 05:28:04 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.40, #queue-req: 0
|
||||
[2026-09-10 05:28:05 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.83, #queue-req: 0
|
||||
[2026-09-10 05:28:05 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.66, #queue-req: 0
|
||||
[2026-09-10 05:28:05] INFO: 127.0.0.1:41770 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:28:06 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.24
|
||||
[2026-09-10 05:28:06 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17057.57
|
||||
[2026-09-10 05:28:06 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 16873.65
|
||||
[2026-09-10 05:28:06 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 410991.66
|
||||
[2026-09-10 05:28:07 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.95, #queue-req: 0
|
||||
[2026-09-10 05:28:07 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.17, #queue-req: 0
|
||||
[2026-09-10 05:28:08 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.73, #queue-req: 0
|
||||
[2026-09-10 05:28:08 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.49, #queue-req: 0
|
||||
[2026-09-10 05:28:09 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.48, #queue-req: 0
|
||||
[2026-09-10 05:28:09 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.23, #queue-req: 0
|
||||
[2026-09-10 05:28:10 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.79, #queue-req: 0
|
||||
[2026-09-10 05:28:10 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.48, #queue-req: 0
|
||||
[2026-09-10 05:28:10 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.89, #queue-req: 0
|
||||
[2026-09-10 05:28:11 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.87, #queue-req: 0
|
||||
[2026-09-10 05:28:11 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.32, #queue-req: 0
|
||||
[2026-09-10 05:28:12 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.00, #queue-req: 0
|
||||
[2026-09-10 05:28:12] INFO: 127.0.0.1:55260 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:28:13 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.39
|
||||
[2026-09-10 05:28:13 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17695.04
|
||||
[2026-09-10 05:28:13 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17662.37
|
||||
[2026-09-10 05:28:13 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 437870.59
|
||||
[2026-09-10 05:28:14 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
|
||||
[2026-09-10 05:28:14 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.11, #queue-req: 0
|
||||
[2026-09-10 05:28:14 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.93, #queue-req: 0
|
||||
[2026-09-10 05:28:15 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.88, #queue-req: 0
|
||||
[2026-09-10 05:28:15 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.11, #queue-req: 0
|
||||
[2026-09-10 05:28:16 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.93, #queue-req: 0
|
||||
[2026-09-10 05:28:16 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.21, #queue-req: 0
|
||||
[2026-09-10 05:28:17 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.39, #queue-req: 0
|
||||
[2026-09-10 05:28:17 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.06, #queue-req: 0
|
||||
[2026-09-10 05:28:17 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.00, #queue-req: 0
|
||||
[2026-09-10 05:28:18 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.62, #queue-req: 0
|
||||
[2026-09-10 05:28:18 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.05, #queue-req: 0
|
||||
[2026-09-10 05:28:19] INFO: 127.0.0.1:59888 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:28:19 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.25
|
||||
[2026-09-10 05:28:19 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17140.02
|
||||
[2026-09-10 05:28:20 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17220.37
|
||||
[2026-09-10 05:28:20 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 387351.35
|
||||
[2026-09-10 05:28:20 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
|
||||
[2026-09-10 05:28:20 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.76, #queue-req: 0
|
||||
[2026-09-10 05:28:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.82, #queue-req: 0
|
||||
[2026-09-10 05:28:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.32, #queue-req: 0
|
||||
[2026-09-10 05:28:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.85, #queue-req: 0
|
||||
[2026-09-10 05:28:22 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.69, #queue-req: 0
|
||||
[2026-09-10 05:28:22 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.88, #queue-req: 0
|
||||
[2026-09-10 05:28:23 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.80, #queue-req: 0
|
||||
[2026-09-10 05:28:23 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.91, #queue-req: 0
|
||||
[2026-09-10 05:28:24 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.49, #queue-req: 0
|
||||
[2026-09-10 05:28:24 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.99, #queue-req: 0
|
||||
[2026-09-10 05:28:24 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.13, #queue-req: 0
|
||||
[2026-09-10 05:28:25 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.81, #queue-req: 0
|
||||
[2026-09-10 05:28:25] INFO: 127.0.0.1:59902 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:28:26 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 160.08
|
||||
[2026-09-10 05:28:26 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17350.40
|
||||
[2026-09-10 05:28:26 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17335.52
|
||||
[2026-09-10 05:28:26 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 396368.12
|
||||
[2026-09-10 05:28:26 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
|
||||
[2026-09-10 05:28:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.17, #queue-req: 0
|
||||
[2026-09-10 05:28:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 103.86, #queue-req: 0
|
||||
[2026-09-10 05:28:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.88, #queue-req: 0
|
||||
[2026-09-10 05:28:28 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.39, #queue-req: 0
|
||||
[2026-09-10 05:28:28 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.59, #queue-req: 0
|
||||
[2026-09-10 05:28:29 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.83, #queue-req: 0
|
||||
[2026-09-10 05:28:29 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.65, #queue-req: 0
|
||||
[2026-09-10 05:28:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.39, #queue-req: 0
|
||||
[2026-09-10 05:28:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.83, #queue-req: 0
|
||||
[2026-09-10 05:28:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.08, #queue-req: 0
|
||||
[2026-09-10 05:28:31 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.75, #queue-req: 0
|
||||
[2026-09-10 05:28:31 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.12, #queue-req: 0
|
||||
[2026-09-10 05:28:32] INFO: 127.0.0.1:60442 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:28:32 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.94
|
||||
[2026-09-10 05:28:32 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17017.54
|
||||
[2026-09-10 05:28:33 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17011.39
|
||||
[2026-09-10 05:28:33 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 398319.36
|
||||
[2026-09-10 05:28:33 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
|
||||
[2026-09-10 05:28:33 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 109.96, #queue-req: 0
|
||||
[2026-09-10 05:28:33 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.71, #queue-req: 0
|
||||
[2026-09-10 05:28:34 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.57, #queue-req: 0
|
||||
[2026-09-10 05:28:34 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.96, #queue-req: 0
|
||||
[2026-09-10 05:28:35 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.13, #queue-req: 0
|
||||
[2026-09-10 05:28:35 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.85, #queue-req: 0
|
||||
[2026-09-10 05:28:36 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.15, #queue-req: 0
|
||||
[2026-09-10 05:28:36 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.79, #queue-req: 0
|
||||
[2026-09-10 05:28:37 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.43, #queue-req: 0
|
||||
[2026-09-10 05:28:37 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.43, #queue-req: 0
|
||||
[2026-09-10 05:28:38 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.16, #queue-req: 0
|
||||
[2026-09-10 05:28:38 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.87, #queue-req: 0
|
||||
[2026-09-10 05:28:38] INFO: 127.0.0.1:51206 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:28:39 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 158.80
|
||||
[2026-09-10 05:28:39 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17702.47
|
||||
[2026-09-10 05:28:39 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17638.20
|
||||
[2026-09-10 05:28:39 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 392293.20
|
||||
[2026-09-10 05:28:39 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
|
||||
[2026-09-10 05:28:40 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.63, #queue-req: 0
|
||||
[2026-09-10 05:28:40 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.20, #queue-req: 0
|
||||
[2026-09-10 05:28:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.62, #queue-req: 0
|
||||
[2026-09-10 05:28:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.31, #queue-req: 0
|
||||
[2026-09-10 05:28:42 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.75, #queue-req: 0
|
||||
[2026-09-10 05:28:42 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.51, #queue-req: 0
|
||||
[2026-09-10 05:28:42 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.04, #queue-req: 0
|
||||
[2026-09-10 05:28:43 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.46, #queue-req: 0
|
||||
[2026-09-10 05:28:43 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.30, #queue-req: 0
|
||||
[2026-09-10 05:28:44 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.93, #queue-req: 0
|
||||
[2026-09-10 05:28:44 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.25, #queue-req: 0
|
||||
[2026-09-10 05:28:45 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.30, #queue-req: 0
|
||||
[2026-09-10 05:28:45] INFO: 127.0.0.1:51216 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:28:45 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 158.91
|
||||
[2026-09-10 05:28:46 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17041.01
|
||||
[2026-09-10 05:28:46 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17065.29
|
||||
[2026-09-10 05:28:46 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 379012.97
|
||||
[2026-09-10 05:28:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
|
||||
[2026-09-10 05:28:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 110.72, #queue-req: 0
|
||||
[2026-09-10 05:28:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.25, #queue-req: 0
|
||||
[2026-09-10 05:28:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.53, #queue-req: 0
|
||||
[2026-09-10 05:28:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.02, #queue-req: 0
|
||||
[2026-09-10 05:28:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.82, #queue-req: 0
|
||||
[2026-09-10 05:28:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.68, #queue-req: 0
|
||||
[2026-09-10 05:28:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.21, #queue-req: 0
|
||||
[2026-09-10 05:28:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.74, #queue-req: 0
|
||||
[2026-09-10 05:28:50 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.48, #queue-req: 0
|
||||
[2026-09-10 05:28:50 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.47, #queue-req: 0
|
||||
[2026-09-10 05:28:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.88, #queue-req: 0
|
||||
[2026-09-10 05:28:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.74, #queue-req: 0
|
||||
[2026-09-10 05:28:51] INFO: 127.0.0.1:48420 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:28:52 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.07
|
||||
[2026-09-10 05:28:52 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17348.12
|
||||
[2026-09-10 05:28:52 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17338.37
|
||||
[2026-09-10 05:28:52 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 393722.24
|
||||
[2026-09-10 05:28:53 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.88, #queue-req: 0
|
||||
[2026-09-10 05:28:53 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 110.81, #queue-req: 0
|
||||
[2026-09-10 05:28:53 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.58, #queue-req: 0
|
||||
[2026-09-10 05:28:54 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.89, #queue-req: 0
|
||||
[2026-09-10 05:28:54 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.16, #queue-req: 0
|
||||
[2026-09-10 05:28:55 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.96, #queue-req: 0
|
||||
[2026-09-10 05:28:55 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.17, #queue-req: 0
|
||||
[2026-09-10 05:28:55 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.13, #queue-req: 0
|
||||
[2026-09-10 05:28:56 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.94, #queue-req: 0
|
||||
[2026-09-10 05:28:56 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.61, #queue-req: 0
|
||||
[2026-09-10 05:28:57 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.43, #queue-req: 0
|
||||
[2026-09-10 05:28:57 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.14, #queue-req: 0
|
||||
[2026-09-10 05:28:58 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.35, #queue-req: 0
|
||||
[2026-09-10 05:28:58] INFO: 127.0.0.1:48932 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:28:58 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 158.73
|
||||
[2026-09-10 05:28:59 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17549.97
|
||||
[2026-09-10 05:28:59 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17455.86
|
||||
[2026-09-10 05:28:59 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 391661.25
|
||||
[2026-09-10 05:28:59 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.91, #queue-req: 0
|
||||
[2026-09-10 05:28:59 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 107.06, #queue-req: 0
|
||||
[2026-09-10 05:29:00 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.85, #queue-req: 0
|
||||
[2026-09-10 05:29:00 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.87, #queue-req: 0
|
||||
[2026-09-10 05:29:01 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.70, #queue-req: 0
|
||||
[2026-09-10 05:29:01 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.79, #queue-req: 0
|
||||
[2026-09-10 05:29:02 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.64, #queue-req: 0
|
||||
[2026-09-10 05:29:02 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.60, #queue-req: 0
|
||||
[2026-09-10 05:29:03 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.28, #queue-req: 0
|
||||
[2026-09-10 05:29:03 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.57, #queue-req: 0
|
||||
[2026-09-10 05:29:04 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.36, #queue-req: 0
|
||||
[2026-09-10 05:29:04 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.60, #queue-req: 0
|
||||
[2026-09-10 05:29:04 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.26, #queue-req: 0
|
||||
[2026-09-10 05:29:05] INFO: 127.0.0.1:48936 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:29:05 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.06
|
||||
[2026-09-10 05:29:05 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17735.57
|
||||
[2026-09-10 05:29:06 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17689.19
|
||||
[2026-09-10 05:29:06 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 418243.76
|
||||
[2026-09-10 05:29:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
|
||||
[2026-09-10 05:29:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.43, #queue-req: 0
|
||||
[2026-09-10 05:29:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 100.83, #queue-req: 0
|
||||
[2026-09-10 05:29:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.53, #queue-req: 0
|
||||
[2026-09-10 05:29:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.45, #queue-req: 0
|
||||
[2026-09-10 05:29:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.68, #queue-req: 0
|
||||
[2026-09-10 05:29:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.88, #queue-req: 0
|
||||
[2026-09-10 05:29:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.53, #queue-req: 0
|
||||
[2026-09-10 05:29:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.67, #queue-req: 0
|
||||
[2026-09-10 05:29:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.28, #queue-req: 0
|
||||
[2026-09-10 05:29:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.07, #queue-req: 0
|
||||
[2026-09-10 05:29:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.23, #queue-req: 0
|
||||
[2026-09-10 05:29:11 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.37, #queue-req: 0
|
||||
[2026-09-10 05:29:11] INFO: 127.0.0.1:58922 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:29:12 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.64
|
||||
[2026-09-10 05:29:12 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17012.62
|
||||
[2026-09-10 05:29:12 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17002.83
|
||||
[2026-09-10 05:29:12 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 380234.28
|
||||
[2026-09-10 05:29:12 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
|
||||
[2026-09-10 05:29:13 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 110.93, #queue-req: 0
|
||||
[2026-09-10 05:29:13 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.34, #queue-req: 0
|
||||
[2026-09-10 05:29:14 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.01, #queue-req: 0
|
||||
[2026-09-10 05:29:14 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.84, #queue-req: 0
|
||||
[2026-09-10 05:29:14 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.67, #queue-req: 0
|
||||
[2026-09-10 05:29:15 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.14, #queue-req: 0
|
||||
[2026-09-10 05:29:15 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.86, #queue-req: 0
|
||||
[2026-09-10 05:29:16 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.64, #queue-req: 0
|
||||
[2026-09-10 05:29:16 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.55, #queue-req: 0
|
||||
[2026-09-10 05:29:17 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.19, #queue-req: 0
|
||||
[2026-09-10 05:29:17 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.60, #queue-req: 0
|
||||
[2026-09-10 05:29:17 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.50, #queue-req: 0
|
||||
[2026-09-10 05:29:18] INFO: 127.0.0.1:58924 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:29:18 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.69
|
||||
[2026-09-10 05:29:18 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17380.48
|
||||
[2026-09-10 05:29:19 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17372.91
|
||||
[2026-09-10 05:29:19 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 376968.20
|
||||
[2026-09-10 05:29:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
|
||||
[2026-09-10 05:29:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 112.62, #queue-req: 0
|
||||
[2026-09-10 05:29:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 100.19, #queue-req: 0
|
||||
[2026-09-10 05:29:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.50, #queue-req: 0
|
||||
[2026-09-10 05:29:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.97, #queue-req: 0
|
||||
[2026-09-10 05:29:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.28, #queue-req: 0
|
||||
[2026-09-10 05:29:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.74, #queue-req: 0
|
||||
[2026-09-10 05:29:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.04, #queue-req: 0
|
||||
[2026-09-10 05:29:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.05, #queue-req: 0
|
||||
[2026-09-10 05:29:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.24, #queue-req: 0
|
||||
[2026-09-10 05:29:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.07, #queue-req: 0
|
||||
[2026-09-10 05:29:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.46, #queue-req: 0
|
||||
[2026-09-10 05:29:24 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.14, #queue-req: 0
|
||||
[2026-09-10 05:29:24] INFO: 127.0.0.1:54110 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:29:24 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.94
|
||||
[2026-09-10 05:29:25 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17537.93
|
||||
[2026-09-10 05:29:25 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17564.52
|
||||
[2026-09-10 05:29:25 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 385204.31
|
||||
[2026-09-10 05:29:25 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
|
||||
[2026-09-10 05:29:26 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 100.03, #queue-req: 0
|
||||
[2026-09-10 05:29:26 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.20, #queue-req: 0
|
||||
[2026-09-10 05:29:26 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.30, #queue-req: 0
|
||||
[2026-09-10 05:29:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.22, #queue-req: 0
|
||||
[2026-09-10 05:29:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.82, #queue-req: 0
|
||||
[2026-09-10 05:29:28 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.27, #queue-req: 0
|
||||
[2026-09-10 05:29:28 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.17, #queue-req: 0
|
||||
[2026-09-10 05:29:29 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.10, #queue-req: 0
|
||||
[2026-09-10 05:29:29 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.80, #queue-req: 0
|
||||
[2026-09-10 05:29:30 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.01, #queue-req: 0
|
||||
[2026-09-10 05:29:30 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.53, #queue-req: 0
|
||||
[2026-09-10 05:29:31 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.39, #queue-req: 0
|
||||
[2026-09-10 05:29:31] INFO: 127.0.0.1:45216 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:29:31 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.88
|
||||
[2026-09-10 05:29:32 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17620.10
|
||||
[2026-09-10 05:29:32 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17730.20
|
||||
[2026-09-10 05:29:32 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 413367.81
|
||||
[2026-09-10 05:29:32 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
|
||||
[2026-09-10 05:29:32 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 111.07, #queue-req: 0
|
||||
[2026-09-10 05:29:33 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.31, #queue-req: 0
|
||||
[2026-09-10 05:29:33 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.34, #queue-req: 0
|
||||
[2026-09-10 05:29:34 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.48, #queue-req: 0
|
||||
[2026-09-10 05:29:34 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.02, #queue-req: 0
|
||||
[2026-09-10 05:29:34 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.32, #queue-req: 0
|
||||
[2026-09-10 05:29:35 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.19, #queue-req: 0
|
||||
[2026-09-10 05:29:35 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.15, #queue-req: 0
|
||||
[2026-09-10 05:29:36 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.87, #queue-req: 0
|
||||
[2026-09-10 05:29:36 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.78, #queue-req: 0
|
||||
[2026-09-10 05:29:37 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.04, #queue-req: 0
|
||||
[2026-09-10 05:29:37 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.12, #queue-req: 0
|
||||
[2026-09-10 05:29:37] INFO: 127.0.0.1:45224 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:29:38 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.90
|
||||
[2026-09-10 05:29:38 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17313.96
|
||||
[2026-09-10 05:29:38 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17374.01
|
||||
[2026-09-10 05:29:38 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 364144.24
|
||||
[2026-09-10 05:29:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
|
||||
[2026-09-10 05:29:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 105.17, #queue-req: 0
|
||||
[2026-09-10 05:29:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.47, #queue-req: 0
|
||||
[2026-09-10 05:29:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.79, #queue-req: 0
|
||||
[2026-09-10 05:29:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.80, #queue-req: 0
|
||||
[2026-09-10 05:29:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.67, #queue-req: 0
|
||||
[2026-09-10 05:29:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.32, #queue-req: 0
|
||||
[2026-09-10 05:29:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.02, #queue-req: 0
|
||||
[2026-09-10 05:29:42 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.43, #queue-req: 0
|
||||
[2026-09-10 05:29:42 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.73, #queue-req: 0
|
||||
[2026-09-10 05:29:43 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.67, #queue-req: 0
|
||||
[2026-09-10 05:29:43 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.06, #queue-req: 0
|
||||
[2026-09-10 05:29:44 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.73, #queue-req: 0
|
||||
[2026-09-10 05:29:44] INFO: 127.0.0.1:45950 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:29:44 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.61
|
||||
[2026-09-10 05:29:44 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17307.80
|
||||
[2026-09-10 05:29:45 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17436.66
|
||||
[2026-09-10 05:29:45 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 398465.33
|
||||
[2026-09-10 05:29:45 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
|
||||
[2026-09-10 05:29:45 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 108.94, #queue-req: 0
|
||||
[2026-09-10 05:29:46 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 98.77, #queue-req: 0
|
||||
[2026-09-10 05:29:46 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.00, #queue-req: 0
|
||||
[2026-09-10 05:29:47 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.99, #queue-req: 0
|
||||
[2026-09-10 05:29:47 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.07, #queue-req: 0
|
||||
[2026-09-10 05:29:47 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.58, #queue-req: 0
|
||||
[2026-09-10 05:29:48 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.86, #queue-req: 0
|
||||
[2026-09-10 05:29:48 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.20, #queue-req: 0
|
||||
[2026-09-10 05:29:49 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.27, #queue-req: 0
|
||||
[2026-09-10 05:29:49 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.66, #queue-req: 0
|
||||
[2026-09-10 05:29:50 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.47, #queue-req: 0
|
||||
[2026-09-10 05:29:50 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.08, #queue-req: 0
|
||||
[2026-09-10 05:29:50] INFO: 127.0.0.1:43446 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:29:51 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.99
|
||||
[2026-09-10 05:29:51 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17050.57
|
||||
[2026-09-10 05:29:51 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17024.07
|
||||
[2026-09-10 05:29:51 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 377001.13
|
||||
[2026-09-10 05:29:51 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
|
||||
[2026-09-10 05:29:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.93, #queue-req: 0
|
||||
[2026-09-10 05:29:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.53, #queue-req: 0
|
||||
[2026-09-10 05:29:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.25, #queue-req: 0
|
||||
[2026-09-10 05:29:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.44, #queue-req: 0
|
||||
[2026-09-10 05:29:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.55, #queue-req: 0
|
||||
[2026-09-10 05:29:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.26, #queue-req: 0
|
||||
[2026-09-10 05:29:55 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.63, #queue-req: 0
|
||||
[2026-09-10 05:29:55 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.24, #queue-req: 0
|
||||
[2026-09-10 05:29:55 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.19, #queue-req: 0
|
||||
[2026-09-10 05:29:56 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.55, #queue-req: 0
|
||||
[2026-09-10 05:29:56 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.83, #queue-req: 0
|
||||
[2026-09-10 05:29:57 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.44, #queue-req: 0
|
||||
[2026-09-10 05:29:57] INFO: 127.0.0.1:43456 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:29:57 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.74
|
||||
[2026-09-10 05:29:58 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17673.58
|
||||
[2026-09-10 05:29:58 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17444.02
|
||||
[2026-09-10 05:29:58 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 442709.39
|
||||
[2026-09-10 05:29:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
|
||||
[2026-09-10 05:29:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 108.41, #queue-req: 0
|
||||
[2026-09-10 05:29:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.25, #queue-req: 0
|
||||
[2026-09-10 05:29:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.60, #queue-req: 0
|
||||
[2026-09-10 05:30:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.15, #queue-req: 0
|
||||
[2026-09-10 05:30:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.47, #queue-req: 0
|
||||
[2026-09-10 05:30:01 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.45, #queue-req: 0
|
||||
[2026-09-10 05:30:01 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.72, #queue-req: 0
|
||||
[2026-09-10 05:30:02 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.06, #queue-req: 0
|
||||
[2026-09-10 05:30:02 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.50, #queue-req: 0
|
||||
[2026-09-10 05:30:02 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.57, #queue-req: 0
|
||||
[2026-09-10 05:30:03 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.53, #queue-req: 0
|
||||
[2026-09-10 05:30:03 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.21, #queue-req: 0
|
||||
[2026-09-10 05:30:03] INFO: 127.0.0.1:49560 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:30:04 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.87
|
||||
[2026-09-10 05:30:04 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17024.45
|
||||
[2026-09-10 05:30:04 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 16899.80
|
||||
[2026-09-10 05:30:04 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 376913.94
|
||||
[2026-09-10 05:30:05 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.90, #queue-req: 0
|
||||
[2026-09-10 05:30:05 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 101.37, #queue-req: 0
|
||||
[2026-09-10 05:30:05 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.39, #queue-req: 0
|
||||
[2026-09-10 05:30:06 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.55, #queue-req: 0
|
||||
[2026-09-10 05:30:06 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.16, #queue-req: 0
|
||||
[2026-09-10 05:30:07 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.31, #queue-req: 0
|
||||
[2026-09-10 05:30:07 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.23, #queue-req: 0
|
||||
[2026-09-10 05:30:08 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.37, #queue-req: 0
|
||||
[2026-09-10 05:30:08 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.51, #queue-req: 0
|
||||
[2026-09-10 05:30:09 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.18, #queue-req: 0
|
||||
[2026-09-10 05:30:09 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.07, #queue-req: 0
|
||||
[2026-09-10 05:30:09 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.34, #queue-req: 0
|
||||
[2026-09-10 05:30:10] INFO: 127.0.0.1:35116 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:30:10 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.85
|
||||
[2026-09-10 05:30:11 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17331.73
|
||||
[2026-09-10 05:30:11 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17380.05
|
||||
[2026-09-10 05:30:11 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 376651.23
|
||||
[2026-09-10 05:30:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.89, #queue-req: 0
|
||||
[2026-09-10 05:30:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 104.75, #queue-req: 0
|
||||
[2026-09-10 05:30:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.22, #queue-req: 0
|
||||
[2026-09-10 05:30:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.80, #queue-req: 0
|
||||
[2026-09-10 05:30:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.49, #queue-req: 0
|
||||
[2026-09-10 05:30:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.23, #queue-req: 0
|
||||
[2026-09-10 05:30:14 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.53, #queue-req: 0
|
||||
[2026-09-10 05:30:14 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.36, #queue-req: 0
|
||||
[2026-09-10 05:30:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.65, #queue-req: 0
|
||||
[2026-09-10 05:30:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.52, #queue-req: 0
|
||||
[2026-09-10 05:30:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.56, #queue-req: 0
|
||||
[2026-09-10 05:30:16 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.12, #queue-req: 0
|
||||
[2026-09-10 05:30:16] INFO: 127.0.0.1:35126 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:30:17 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 159.92
|
||||
[2026-09-10 05:30:17 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17012.10
|
||||
[2026-09-10 05:30:17 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17006.24
|
||||
[2026-09-10 05:30:17 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 385755.05
|
||||
[2026-09-10 05:30:18 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.93, #queue-req: 0
|
||||
[2026-09-10 05:30:18 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.34, #queue-req: 0
|
||||
[2026-09-10 05:30:18 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.97, #queue-req: 0
|
||||
[2026-09-10 05:30:19 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.89, #queue-req: 0
|
||||
[2026-09-10 05:30:19 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.97, #queue-req: 0
|
||||
[2026-09-10 05:30:20 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.72, #queue-req: 0
|
||||
[2026-09-10 05:30:20 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.53, #queue-req: 0
|
||||
[2026-09-10 05:30:21 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.45, #queue-req: 0
|
||||
[2026-09-10 05:30:21 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.08, #queue-req: 0
|
||||
[2026-09-10 05:30:21 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.85, #queue-req: 0
|
||||
[2026-09-10 05:30:22 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.87, #queue-req: 0
|
||||
[2026-09-10 05:30:22 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.07, #queue-req: 0
|
||||
[2026-09-10 05:30:23] INFO: 127.0.0.1:45272 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:30:23 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 12288, cuda graph: False, input throughput (token/s): 161.10
|
||||
[2026-09-10 05:30:24 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 8192, cuda graph: False, input throughput (token/s): 17663.48
|
||||
[2026-09-10 05:30:24 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 4096, cuda graph: False, input throughput (token/s): 17724.99
|
||||
[2026-09-10 05:30:24 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 4096, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 402718.07
|
||||
[2026-09-10 05:30:24 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 1.92, #queue-req: 0
|
||||
[2026-09-10 05:30:24 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 108.94, #queue-req: 0
|
||||
[2026-09-10 05:30:25 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.26, #queue-req: 0
|
||||
[2026-09-10 05:30:25 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.09, #queue-req: 0
|
||||
[2026-09-10 05:30:26 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.33, #queue-req: 0
|
||||
[2026-09-10 05:30:26 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16640, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.07, #queue-req: 0
|
||||
[2026-09-10 05:30:27 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.00, #queue-req: 0
|
||||
[2026-09-10 05:30:27 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.90, #queue-req: 0
|
||||
[2026-09-10 05:30:27 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.09, #queue-req: 0
|
||||
[2026-09-10 05:30:28 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.07, #queue-req: 0
|
||||
[2026-09-10 05:30:28 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.68, #queue-req: 0
|
||||
[2026-09-10 05:30:29 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 16896, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.71, #queue-req: 0
|
||||
[2026-09-10 05:30:29] INFO: 127.0.0.1:37684 - "GET /server_info HTTP/1.1" 200 OK
|
||||
[2026-09-10 05:30:29] INFO: 127.0.0.1:37692 - "GET /server_info HTTP/1.1" 200 OK
|
||||
@ -0,0 +1 @@
|
||||
2026-09-10T05:39:14+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 270267 MiB, 3847 MiB, 0 %, 245.21 W
|
||||
1, 266529 MiB, 7585 MiB, 0 %, 236.04 W
|
||||
2, 267221 MiB, 6893 MiB, 0 %, 237.03 W
|
||||
3, 267217 MiB, 6897 MiB, 0 %, 247.66 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 182.01 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
|
||||
|
Binary file not shown.
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=640, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=128, output_file='/results/points/16k_512_c128.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=640, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=128, output_file='/results/points/16k_512_c128.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 10485760
|
||||
#Output tokens: 327680
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 128
|
||||
Successful requests: 640
|
||||
Benchmark duration (s): 201.20
|
||||
Total input tokens: 10485760
|
||||
Total input text tokens: 10485760
|
||||
Total generated tokens: 327680
|
||||
Total generated tokens (retokenized): 324267
|
||||
Request throughput (req/s): 3.18
|
||||
Input token throughput (tok/s): 52115.96
|
||||
Output token throughput (tok/s): 1628.62
|
||||
Peak output token throughput (tok/s): 8444.00
|
||||
Peak concurrent requests: 256
|
||||
Total token throughput (tok/s): 53744.58
|
||||
Concurrency: 127.47
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 40073.80
|
||||
Median E2E Latency (ms): 39933.15
|
||||
P90 E2E Latency (ms): 40867.69
|
||||
P95 E2E Latency (ms): 40965.30
|
||||
P99 E2E Latency (ms): 41054.20
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 17331.84
|
||||
Median TTFT (ms): 17343.29
|
||||
P90 TTFT (ms): 30366.98
|
||||
P95 TTFT (ms): 31955.08
|
||||
P99 TTFT (ms): 32531.95
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 44.50
|
||||
Median TPOT (ms): 44.47
|
||||
P90 TPOT (ms): 69.86
|
||||
P95 TPOT (ms): 73.50
|
||||
P99 TPOT (ms): 75.57
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 44.51
|
||||
Median ITL (ms): 15.27
|
||||
P90 ITL (ms): 19.67
|
||||
P95 ITL (ms): 21.54
|
||||
P99 ITL (ms): 22.49
|
||||
Max ITL (ms): 31081.58
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T05:45:44+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 269243 MiB, 4871 MiB, 0 %, 244.80 W
|
||||
1, 268011 MiB, 6103 MiB, 0 %, 236.31 W
|
||||
2, 268177 MiB, 5937 MiB, 0 %, 236.89 W
|
||||
3, 266263 MiB, 7851 MiB, 0 %, 246.29 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 181.89 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
|
||||
|
Binary file not shown.
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=1280, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=256, output_file='/results/points/16k_512_c256.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=1280, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=256, output_file='/results/points/16k_512_c256.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 20971520
|
||||
#Output tokens: 655360
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 256
|
||||
Successful requests: 1280
|
||||
Benchmark duration (s): 365.13
|
||||
Total input tokens: 20971520
|
||||
Total input text tokens: 20971520
|
||||
Total generated tokens: 655360
|
||||
Total generated tokens (retokenized): 647182
|
||||
Request throughput (req/s): 3.51
|
||||
Input token throughput (tok/s): 57436.40
|
||||
Output token throughput (tok/s): 1794.89
|
||||
Peak output token throughput (tok/s): 16320.00
|
||||
Peak concurrent requests: 512
|
||||
Total token throughput (tok/s): 59231.29
|
||||
Concurrency: 254.97
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 72730.55
|
||||
Median E2E Latency (ms): 72592.18
|
||||
P90 E2E Latency (ms): 74451.94
|
||||
P95 E2E Latency (ms): 74517.86
|
||||
P99 E2E Latency (ms): 74606.27
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 33211.70
|
||||
Median TTFT (ms): 33119.78
|
||||
P90 TTFT (ms): 59093.67
|
||||
P95 TTFT (ms): 62849.59
|
||||
P99 TTFT (ms): 65119.95
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 77.34
|
||||
Median TPOT (ms): 77.14
|
||||
P90 TPOT (ms): 128.18
|
||||
P95 TPOT (ms): 134.28
|
||||
P99 TPOT (ms): 139.41
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 77.34
|
||||
Median ITL (ms): 16.59
|
||||
P90 ITL (ms): 22.04
|
||||
P95 ITL (ms): 24.22
|
||||
P99 ITL (ms): 25.93
|
||||
Max ITL (ms): 64829.80
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T05:33:19+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 267151 MiB, 6963 MiB, 0 %, 243.92 W
|
||||
1, 264711 MiB, 9403 MiB, 0 %, 236.50 W
|
||||
2, 267957 MiB, 6157 MiB, 0 %, 236.94 W
|
||||
3, 266287 MiB, 7827 MiB, 0 %, 246.99 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 182.20 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=160, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=32, output_file='/results/points/16k_512_c32.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=160, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=32, output_file='/results/points/16k_512_c32.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 2621440
|
||||
#Output tokens: 81920
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 32
|
||||
Successful requests: 160
|
||||
Benchmark duration (s): 72.62
|
||||
Total input tokens: 2621440
|
||||
Total input text tokens: 2621440
|
||||
Total generated tokens: 81920
|
||||
Total generated tokens (retokenized): 81120
|
||||
Request throughput (req/s): 2.20
|
||||
Input token throughput (tok/s): 36100.44
|
||||
Output token throughput (tok/s): 1128.14
|
||||
Peak output token throughput (tok/s): 2668.00
|
||||
Peak concurrent requests: 64
|
||||
Total token throughput (tok/s): 37228.58
|
||||
Concurrency: 31.86
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 14458.39
|
||||
Median E2E Latency (ms): 14360.77
|
||||
P90 E2E Latency (ms): 14889.61
|
||||
P95 E2E Latency (ms): 14940.54
|
||||
P99 E2E Latency (ms): 14964.61
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 4936.55
|
||||
Median TTFT (ms): 5135.45
|
||||
P90 TTFT (ms): 8067.95
|
||||
P95 TTFT (ms): 8087.71
|
||||
P99 TTFT (ms): 8603.75
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 18.63
|
||||
Median TPOT (ms): 18.13
|
||||
P90 TPOT (ms): 25.17
|
||||
P95 TPOT (ms): 25.54
|
||||
P99 TPOT (ms): 25.85
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 18.63
|
||||
Median ITL (ms): 12.17
|
||||
P90 ITL (ms): 13.64
|
||||
P95 ITL (ms): 14.52
|
||||
P99 ITL (ms): 16.03
|
||||
Max ITL (ms): 7200.39
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T05:35:32+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 270175 MiB, 3939 MiB, 0 %, 245.79 W
|
||||
1, 265185 MiB, 8929 MiB, 0 %, 236.66 W
|
||||
2, 267183 MiB, 6931 MiB, 0 %, 238.58 W
|
||||
3, 265317 MiB, 8797 MiB, 0 %, 247.74 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 182.05 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.21 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 182.01 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=320, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=64, output_file='/results/points/16k_512_c64.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=320, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=64, output_file='/results/points/16k_512_c64.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 5242880
|
||||
#Output tokens: 163840
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 64
|
||||
Successful requests: 320
|
||||
Benchmark duration (s): 115.64
|
||||
Total input tokens: 5242880
|
||||
Total input text tokens: 5242880
|
||||
Total generated tokens: 163840
|
||||
Total generated tokens (retokenized): 161636
|
||||
Request throughput (req/s): 2.77
|
||||
Input token throughput (tok/s): 45336.60
|
||||
Output token throughput (tok/s): 1416.77
|
||||
Peak output token throughput (tok/s): 4928.00
|
||||
Peak concurrent requests: 128
|
||||
Total token throughput (tok/s): 46753.36
|
||||
Concurrency: 63.66
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 23007.55
|
||||
Median E2E Latency (ms): 22960.73
|
||||
P90 E2E Latency (ms): 23535.50
|
||||
P95 E2E Latency (ms): 23633.89
|
||||
P99 E2E Latency (ms): 23711.10
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 8990.44
|
||||
Median TTFT (ms): 9165.76
|
||||
P90 TTFT (ms): 15556.71
|
||||
P95 TTFT (ms): 16042.56
|
||||
P99 TTFT (ms): 16476.52
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 27.43
|
||||
Median TPOT (ms): 27.23
|
||||
P90 TPOT (ms): 40.16
|
||||
P95 TPOT (ms): 41.92
|
||||
P99 TPOT (ms): 42.70
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 27.43
|
||||
Median ITL (ms): 13.36
|
||||
P90 ITL (ms): 17.01
|
||||
P95 ITL (ms): 18.35
|
||||
P99 ITL (ms): 19.25
|
||||
Max ITL (ms): 15156.70
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T05:31:50+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 268149 MiB, 5965 MiB, 0 %, 241.92 W
|
||||
1, 263337 MiB, 10777 MiB, 0 %, 234.70 W
|
||||
2, 267143 MiB, 6971 MiB, 0 %, 236.84 W
|
||||
3, 263745 MiB, 10369 MiB, 0 %, 245.26 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 182.51 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.29 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 182.40 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=8, output_file='/results/points/16k_512_c8.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=16384, random_output_len=512, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=8, output_file='/results/points/16k_512_c8.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 1048576
|
||||
#Output tokens: 32768
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 8
|
||||
Successful requests: 64
|
||||
Benchmark duration (s): 61.85
|
||||
Total input tokens: 1048576
|
||||
Total input text tokens: 1048576
|
||||
Total generated tokens: 32768
|
||||
Total generated tokens (retokenized): 32406
|
||||
Request throughput (req/s): 1.03
|
||||
Input token throughput (tok/s): 16954.74
|
||||
Output token throughput (tok/s): 529.84
|
||||
Peak output token throughput (tok/s): 792.00
|
||||
Peak concurrent requests: 16
|
||||
Total token throughput (tok/s): 17484.58
|
||||
Concurrency: 7.98
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 7714.34
|
||||
Median E2E Latency (ms): 7658.51
|
||||
P90 E2E Latency (ms): 8116.15
|
||||
P95 E2E Latency (ms): 8130.24
|
||||
P99 E2E Latency (ms): 8147.10
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 1875.84
|
||||
Median TTFT (ms): 1937.05
|
||||
P90 TTFT (ms): 2255.93
|
||||
P95 TTFT (ms): 2578.91
|
||||
P99 TTFT (ms): 2648.89
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 11.43
|
||||
Median TPOT (ms): 11.34
|
||||
P90 TPOT (ms): 12.56
|
||||
P95 TPOT (ms): 12.69
|
||||
P99 TPOT (ms): 12.87
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 11.43
|
||||
Median ITL (ms): 10.52
|
||||
P90 ITL (ms): 11.94
|
||||
P95 ITL (ms): 12.30
|
||||
P99 ITL (ms): 12.93
|
||||
Max ITL (ms): 1265.14
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T06:00:58+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 260225 MiB, 13889 MiB, 0 %, 241.91 W
|
||||
1, 260417 MiB, 13697 MiB, 0 %, 234.35 W
|
||||
2, 260529 MiB, 13585 MiB, 0 %, 234.97 W
|
||||
3, 259501 MiB, 14613 MiB, 0 %, 243.95 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 183.34 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.17 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 183.89 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=1, output_file='/results/points/1k_128_c1.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=1, output_file='/results/points/1k_128_c1.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 65536
|
||||
#Output tokens: 8192
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 1
|
||||
Successful requests: 64
|
||||
Benchmark duration (s): 101.91
|
||||
Total input tokens: 65536
|
||||
Total input text tokens: 65536
|
||||
Total generated tokens: 8192
|
||||
Total generated tokens (retokenized): 8079
|
||||
Request throughput (req/s): 0.63
|
||||
Input token throughput (tok/s): 643.09
|
||||
Output token throughput (tok/s): 80.39
|
||||
Peak output token throughput (tok/s): 107.00
|
||||
Peak concurrent requests: 2
|
||||
Total token throughput (tok/s): 723.47
|
||||
Concurrency: 1.00
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 1590.91
|
||||
Median E2E Latency (ms): 1591.12
|
||||
P90 E2E Latency (ms): 1655.93
|
||||
P95 E2E Latency (ms): 1682.16
|
||||
P99 E2E Latency (ms): 1732.89
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 331.86
|
||||
Median TTFT (ms): 331.96
|
||||
P90 TTFT (ms): 338.18
|
||||
P95 TTFT (ms): 342.42
|
||||
P99 TTFT (ms): 346.96
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 9.91
|
||||
Median TPOT (ms): 9.91
|
||||
P90 TPOT (ms): 10.43
|
||||
P95 TPOT (ms): 10.62
|
||||
P99 TPOT (ms): 11.06
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 9.92
|
||||
Median ITL (ms): 9.93
|
||||
P90 ITL (ms): 11.44
|
||||
P95 ITL (ms): 11.75
|
||||
P99 ITL (ms): 12.22
|
||||
Max ITL (ms): 15.98
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T06:03:03+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 267981 MiB, 6133 MiB, 0 %, 243.80 W
|
||||
1, 268261 MiB, 5853 MiB, 0 %, 235.13 W
|
||||
2, 267361 MiB, 6753 MiB, 0 %, 236.93 W
|
||||
3, 264909 MiB, 9205 MiB, 0 %, 245.82 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 182.47 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 183.17 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 182.83 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=640, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=128, output_file='/results/points/1k_128_c128.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=640, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=128, output_file='/results/points/1k_128_c128.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 655360
|
||||
#Output tokens: 81920
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 128
|
||||
Successful requests: 640
|
||||
Benchmark duration (s): 22.24
|
||||
Total input tokens: 655360
|
||||
Total input text tokens: 655360
|
||||
Total generated tokens: 81920
|
||||
Total generated tokens (retokenized): 80729
|
||||
Request throughput (req/s): 28.78
|
||||
Input token throughput (tok/s): 29472.67
|
||||
Output token throughput (tok/s): 3684.08
|
||||
Peak output token throughput (tok/s): 8508.00
|
||||
Peak concurrent requests: 256
|
||||
Total token throughput (tok/s): 33156.76
|
||||
Concurrency: 126.32
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 4388.74
|
||||
Median E2E Latency (ms): 4274.89
|
||||
P90 E2E Latency (ms): 4888.74
|
||||
P95 E2E Latency (ms): 4916.76
|
||||
P99 E2E Latency (ms): 4964.03
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 1748.02
|
||||
Median TTFT (ms): 1769.41
|
||||
P90 TTFT (ms): 2354.73
|
||||
P95 TTFT (ms): 2841.20
|
||||
P99 TTFT (ms): 2851.83
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 20.79
|
||||
Median TPOT (ms): 19.82
|
||||
P90 TPOT (ms): 27.51
|
||||
P95 TPOT (ms): 27.68
|
||||
P99 TPOT (ms): 28.92
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 20.80
|
||||
Median ITL (ms): 15.31
|
||||
P90 ITL (ms): 19.24
|
||||
P95 ITL (ms): 20.39
|
||||
P99 ITL (ms): 52.25
|
||||
Max ITL (ms): 1851.24
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T06:03:55+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 266755 MiB, 7359 MiB, 0 %, 243.84 W
|
||||
1, 266199 MiB, 7915 MiB, 0 %, 235.27 W
|
||||
2, 269157 MiB, 4957 MiB, 0 %, 236.85 W
|
||||
3, 266425 MiB, 7689 MiB, 0 %, 245.84 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 182.24 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.17 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 182.05 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=1280, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=256, output_file='/results/points/1k_128_c256.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=1280, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=256, output_file='/results/points/1k_128_c256.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 1310720
|
||||
#Output tokens: 163840
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 256
|
||||
Successful requests: 1280
|
||||
Benchmark duration (s): 36.14
|
||||
Total input tokens: 1310720
|
||||
Total input text tokens: 1310720
|
||||
Total generated tokens: 163840
|
||||
Total generated tokens (retokenized): 161449
|
||||
Request throughput (req/s): 35.42
|
||||
Input token throughput (tok/s): 36265.26
|
||||
Output token throughput (tok/s): 4533.16
|
||||
Peak output token throughput (tok/s): 16224.00
|
||||
Peak concurrent requests: 512
|
||||
Total token throughput (tok/s): 40798.42
|
||||
Concurrency: 251.46
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 7100.38
|
||||
Median E2E Latency (ms): 6570.99
|
||||
P90 E2E Latency (ms): 9276.25
|
||||
P95 E2E Latency (ms): 9295.96
|
||||
P99 E2E Latency (ms): 9342.45
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 2807.50
|
||||
Median TTFT (ms): 2728.95
|
||||
P90 TTFT (ms): 4249.06
|
||||
P95 TTFT (ms): 4505.51
|
||||
P99 TTFT (ms): 7117.16
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 33.80
|
||||
Median TPOT (ms): 32.71
|
||||
P90 TPOT (ms): 51.77
|
||||
P95 TPOT (ms): 59.65
|
||||
P99 TPOT (ms): 67.15
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 33.81
|
||||
Median ITL (ms): 16.52
|
||||
P90 ITL (ms): 21.80
|
||||
P95 ITL (ms): 24.21
|
||||
P99 ITL (ms): 62.30
|
||||
Max ITL (ms): 6732.10
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T06:01:55+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 265741 MiB, 8373 MiB, 0 %, 241.89 W
|
||||
1, 263601 MiB, 10513 MiB, 0 %, 234.70 W
|
||||
2, 265273 MiB, 8841 MiB, 0 %, 236.67 W
|
||||
3, 264477 MiB, 9637 MiB, 0 %, 245.12 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 183.15 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.52 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 183.93 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=160, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=32, output_file='/results/points/1k_128_c32.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=160, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=32, output_file='/results/points/1k_128_c32.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 163840
|
||||
#Output tokens: 20480
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 32
|
||||
Successful requests: 160
|
||||
Benchmark duration (s): 11.91
|
||||
Total input tokens: 163840
|
||||
Total input text tokens: 163840
|
||||
Total generated tokens: 20480
|
||||
Total generated tokens (retokenized): 20159
|
||||
Request throughput (req/s): 13.43
|
||||
Input token throughput (tok/s): 13757.02
|
||||
Output token throughput (tok/s): 1719.63
|
||||
Peak output token throughput (tok/s): 2656.00
|
||||
Peak concurrent requests: 64
|
||||
Total token throughput (tok/s): 15476.65
|
||||
Concurrency: 31.68
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 2358.10
|
||||
Median E2E Latency (ms): 2329.83
|
||||
P90 E2E Latency (ms): 2490.41
|
||||
P95 E2E Latency (ms): 2504.46
|
||||
P99 E2E Latency (ms): 2511.69
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 790.50
|
||||
Median TTFT (ms): 764.38
|
||||
P90 TTFT (ms): 935.92
|
||||
P95 TTFT (ms): 943.32
|
||||
P99 TTFT (ms): 947.21
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 12.34
|
||||
Median TPOT (ms): 12.28
|
||||
P90 TPOT (ms): 12.45
|
||||
P95 TPOT (ms): 12.48
|
||||
P99 TPOT (ms): 14.38
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 12.35
|
||||
Median ITL (ms): 12.20
|
||||
P90 ITL (ms): 14.33
|
||||
P95 ITL (ms): 14.88
|
||||
P99 ITL (ms): 19.57
|
||||
Max ITL (ms): 289.57
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T06:02:25+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 265741 MiB, 8373 MiB, 0 %, 243.69 W
|
||||
1, 265289 MiB, 8825 MiB, 0 %, 235.99 W
|
||||
2, 265307 MiB, 8807 MiB, 0 %, 237.03 W
|
||||
3, 264623 MiB, 9491 MiB, 0 %, 245.81 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 183.38 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.74 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 183.81 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=320, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=64, output_file='/results/points/1k_128_c64.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=320, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=64, output_file='/results/points/1k_128_c64.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 327680
|
||||
#Output tokens: 40960
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 64
|
||||
Successful requests: 320
|
||||
Benchmark duration (s): 15.44
|
||||
Total input tokens: 327680
|
||||
Total input text tokens: 327680
|
||||
Total generated tokens: 40960
|
||||
Total generated tokens (retokenized): 40414
|
||||
Request throughput (req/s): 20.72
|
||||
Input token throughput (tok/s): 21219.84
|
||||
Output token throughput (tok/s): 2652.48
|
||||
Peak output token throughput (tok/s): 4771.00
|
||||
Peak concurrent requests: 128
|
||||
Total token throughput (tok/s): 23872.32
|
||||
Concurrency: 62.85
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 3032.75
|
||||
Median E2E Latency (ms): 3019.70
|
||||
P90 E2E Latency (ms): 3197.66
|
||||
P95 E2E Latency (ms): 3218.95
|
||||
P99 E2E Latency (ms): 3245.89
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 1104.28
|
||||
Median TTFT (ms): 1027.77
|
||||
P90 TTFT (ms): 1355.06
|
||||
P95 TTFT (ms): 1451.23
|
||||
P99 TTFT (ms): 1457.66
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 15.18
|
||||
Median TPOT (ms): 13.88
|
||||
P90 TPOT (ms): 17.80
|
||||
P95 TPOT (ms): 17.88
|
||||
P99 TPOT (ms): 19.35
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 15.19
|
||||
Median ITL (ms): 13.44
|
||||
P90 ITL (ms): 17.03
|
||||
P95 ITL (ms): 18.35
|
||||
P99 ITL (ms): 26.81
|
||||
Max ITL (ms): 808.18
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
2026-09-10T06:01:28+00:00
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 262295 MiB, 11819 MiB, 0 %, 241.92 W
|
||||
1, 260983 MiB, 13131 MiB, 0 %, 234.70 W
|
||||
2, 262043 MiB, 12071 MiB, 0 %, 234.99 W
|
||||
3, 261853 MiB, 12261 MiB, 0 %, 243.96 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 183.07 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.13 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 182.49 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 183.93 W
|
||||
|
File diff suppressed because one or more lines are too long
@ -0,0 +1,59 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
benchmark_args=Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=8, output_file='/results/points/1k_128_c8.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
Waiting up to 60s for http://127.0.0.1:30000/v1/models to become ready...
|
||||
Server ready in 0.0s.
|
||||
|
||||
WARNING It is recommended to use the `Chat` or `Instruct` model for benchmarking.
|
||||
Because when the tokenizer counts the output tokens, if there is gibberish, it might count incorrectly.
|
||||
|
||||
Namespace(backend='sglang', base_url=None, host='127.0.0.1', port=30000, ready_check_timeout_sec=60, dataset_name='random-ids', dataset_path='', dataset_offset=0, agentic_max_turns=None, speed_bench_category=None, speed_bench_output_len=512, model='/models', served_model_name=None, tokenizer='/models', num_prompts=64, sharegpt_output_len=None, sharegpt_context_len=None, random_input_len=1024, random_output_len=128, random_range_ratio=1.0, image_count=1, image_resolution='1080p', random_image_count=False, image_format='jpeg', image_content='random', request_rate=inf, use_trace_timestamps=False, max_concurrency=8, output_file='/results/points/1k_128_c8.json', output_details=True, print_requests=False, disable_tqdm=True, disable_stream=False, return_logprob=False, top_logprobs_num=0, token_ids_logprob=None, logprob_start_len=-1, return_routed_experts=False, cache_report=False, seed=20260904, disable_ignore_eos=False, temperature=0.0, top_p=1.0, extra_request_body=None, apply_chat_template=False, profile=False, plot_throughput=False, profile_activities=['CPU', 'GPU'], profile_start_step=None, profile_steps=None, profile_num_steps=None, profile_by_stage=False, profile_stages=None, profile_output_dir=None, profile_prefix=None, lora_name=None, lora_request_distribution='uniform', lora_zipf_alpha=1.5, prompt_suffix='', pd_separated=False, profile_prefill_url=None, profile_decode_url=None, flush_cache=True, flush_cache_timeout=60.0, warmup_requests=1, tokenize_prompt=True, gsp_num_groups=64, gsp_prompts_per_group=16, gsp_system_prompt_len=2048, gsp_question_len=128, gsp_output_len=256, gsp_range_ratio=1.0, gsp_fast_prepare=False, gsp_send_routing_key=False, gsp_num_turns=1, gsp_ordered=False, gsp_group_distribution='uniform', gsp_zipf_alpha=None, mooncake_slowdown_factor=1.0, mooncake_num_rounds=1, mooncake_workload='conversation', fake_prefill=False, tag=None, header=None)
|
||||
|
||||
#Input tokens: 65536
|
||||
#Output tokens: 8192
|
||||
Starting warmup with 1 sequences...
|
||||
Warmup completed with 1 sequences. Starting main benchmark run...
|
||||
|
||||
============ Serving Benchmark Result ============
|
||||
Backend: sglang
|
||||
Traffic request rate: inf
|
||||
Max request concurrency: 8
|
||||
Successful requests: 64
|
||||
Benchmark duration (s): 14.92
|
||||
Total input tokens: 65536
|
||||
Total input text tokens: 65536
|
||||
Total generated tokens: 8192
|
||||
Total generated tokens (retokenized): 8085
|
||||
Request throughput (req/s): 4.29
|
||||
Input token throughput (tok/s): 4391.78
|
||||
Output token throughput (tok/s): 548.97
|
||||
Peak output token throughput (tok/s): 792.00
|
||||
Peak concurrent requests: 16
|
||||
Total token throughput (tok/s): 4940.76
|
||||
Concurrency: 7.96
|
||||
----------------End-to-End Latency----------------
|
||||
Mean E2E Latency (ms): 1856.61
|
||||
Median E2E Latency (ms): 1838.64
|
||||
P90 E2E Latency (ms): 1935.91
|
||||
P95 E2E Latency (ms): 1939.07
|
||||
P99 E2E Latency (ms): 1944.16
|
||||
---------------Time to First Token----------------
|
||||
Mean TTFT (ms): 558.89
|
||||
Median TTFT (ms): 548.42
|
||||
P90 TTFT (ms): 646.48
|
||||
P95 TTFT (ms): 649.01
|
||||
P99 TTFT (ms): 651.38
|
||||
-----Time per Output Token (excl. 1st token)------
|
||||
Mean TPOT (ms): 10.22
|
||||
Median TPOT (ms): 10.13
|
||||
P90 TPOT (ms): 10.70
|
||||
P95 TPOT (ms): 10.73
|
||||
P99 TPOT (ms): 10.76
|
||||
---------------Inter-Token Latency----------------
|
||||
Mean ITL (ms): 10.22
|
||||
Median ITL (ms): 10.10
|
||||
P90 ITL (ms): 11.38
|
||||
P95 ITL (ms): 11.71
|
||||
P99 ITL (ms): 12.88
|
||||
Max ITL (ms): 22.24
|
||||
==================================================
|
||||
File diff suppressed because it is too large
Load Diff
@ -0,0 +1 @@
|
||||
TIMEOUT_PASS
|
||||
@ -0,0 +1,9 @@
|
||||
index, memory.used [MiB], memory.free [MiB], utilization.gpu [%], power.draw [W]
|
||||
0, 260225 MiB, 13889 MiB, 79 %, 367.26 W
|
||||
1, 260419 MiB, 13695 MiB, 80 %, 357.83 W
|
||||
2, 260531 MiB, 13583 MiB, 78 %, 356.71 W
|
||||
3, 259501 MiB, 14613 MiB, 87 %, 366.72 W
|
||||
4, 4 MiB, 274110 MiB, 0 %, 183.82 W
|
||||
5, 4 MiB, 274110 MiB, 0 %, 182.17 W
|
||||
6, 4 MiB, 274110 MiB, 0 %, 183.67 W
|
||||
7, 4 MiB, 274110 MiB, 0 %, 183.97 W
|
||||
|
Binary file not shown.
@ -0,0 +1,2 @@
|
||||
/sgl-workspace/sglang/python/sglang/bench_serving.py:13: FutureWarning: `sglang.bench_serving` is deprecated and will be removed in a future release; use `sglang.benchmark.serving` instead (e.g. `python -m sglang.benchmark.serving`).
|
||||
warnings.warn(
|
||||
@ -0,0 +1,704 @@
|
||||
[2026-09-10 06:30:56 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.05, #queue-req: 0
|
||||
[2026-09-10 06:30:56 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.69, #queue-req: 0
|
||||
[2026-09-10 06:30:57 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.03, #queue-req: 0
|
||||
[2026-09-10 06:30:57 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.40, #queue-req: 0
|
||||
[2026-09-10 06:30:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.21, #queue-req: 0
|
||||
[2026-09-10 06:30:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.15, #queue-req: 0
|
||||
[2026-09-10 06:30:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.72, #queue-req: 0
|
||||
[2026-09-10 06:30:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.89, #queue-req: 0
|
||||
[2026-09-10 06:30:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.64, #queue-req: 0
|
||||
[2026-09-10 06:31:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.45, #queue-req: 0
|
||||
[2026-09-10 06:31:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.63, #queue-req: 0
|
||||
[2026-09-10 06:31:01 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.53, #queue-req: 0
|
||||
[2026-09-10 06:31:01 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.26, #queue-req: 0
|
||||
[2026-09-10 06:31:01 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.30, #queue-req: 0
|
||||
[2026-09-10 06:31:02 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.74, #queue-req: 0
|
||||
[2026-09-10 06:31:02 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.82, #queue-req: 0
|
||||
[2026-09-10 06:31:03 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.98, #queue-req: 0
|
||||
[2026-09-10 06:31:03 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.45, #queue-req: 0
|
||||
[2026-09-10 06:31:04 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.18, #queue-req: 0
|
||||
[2026-09-10 06:31:04 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.02, #queue-req: 0
|
||||
[2026-09-10 06:31:04 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.39, #queue-req: 0
|
||||
[2026-09-10 06:31:05 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.93, #queue-req: 0
|
||||
[2026-09-10 06:31:05 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.72, #queue-req: 0
|
||||
[2026-09-10 06:31:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.81, #queue-req: 0
|
||||
[2026-09-10 06:31:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.94, #queue-req: 0
|
||||
[2026-09-10 06:31:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.02, #queue-req: 0
|
||||
[2026-09-10 06:31:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.01, #queue-req: 0
|
||||
[2026-09-10 06:31:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.60, #queue-req: 0
|
||||
[2026-09-10 06:31:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.32, #queue-req: 0
|
||||
[2026-09-10 06:31:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.13, #queue-req: 0
|
||||
[2026-09-10 06:31:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.24, #queue-req: 0
|
||||
[2026-09-10 06:31:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.07, #queue-req: 0
|
||||
[2026-09-10 06:31:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.77, #queue-req: 0
|
||||
[2026-09-10 06:31:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.88, #queue-req: 0
|
||||
[2026-09-10 06:31:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.47, #queue-req: 0
|
||||
[2026-09-10 06:31:11 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.60, #queue-req: 0
|
||||
[2026-09-10 06:31:11 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.11, #queue-req: 0
|
||||
[2026-09-10 06:31:12 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.02, #queue-req: 0
|
||||
[2026-09-10 06:31:12 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.48, #queue-req: 0
|
||||
[2026-09-10 06:31:13 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.16, #queue-req: 0
|
||||
[2026-09-10 06:31:13 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.70, #queue-req: 0
|
||||
[2026-09-10 06:31:13 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.69, #queue-req: 0
|
||||
[2026-09-10 06:31:14 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.38, #queue-req: 0
|
||||
[2026-09-10 06:31:14 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.65, #queue-req: 0
|
||||
[2026-09-10 06:31:15 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.73, #queue-req: 0
|
||||
[2026-09-10 06:31:15 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.15, #queue-req: 0
|
||||
[2026-09-10 06:31:16 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.26, #queue-req: 0
|
||||
[2026-09-10 06:31:16 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.52, #queue-req: 0
|
||||
[2026-09-10 06:31:16 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.43, #queue-req: 0
|
||||
[2026-09-10 06:31:17] INFO: 127.0.0.1:49772 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 06:31:17 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 1024, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 5.71
|
||||
[2026-09-10 06:31:17 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 0.29, #queue-req: 0
|
||||
[2026-09-10 06:31:17 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 112.36, #queue-req: 0
|
||||
[2026-09-10 06:31:18 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.50, #queue-req: 0
|
||||
[2026-09-10 06:31:18 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.76, #queue-req: 0
|
||||
[2026-09-10 06:31:19 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.48, #queue-req: 0
|
||||
[2026-09-10 06:31:19 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.39, #queue-req: 0
|
||||
[2026-09-10 06:31:20 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.42, #queue-req: 0
|
||||
[2026-09-10 06:31:20 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.01, #queue-req: 0
|
||||
[2026-09-10 06:31:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.41, #queue-req: 0
|
||||
[2026-09-10 06:31:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.01, #queue-req: 0
|
||||
[2026-09-10 06:31:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.55, #queue-req: 0
|
||||
[2026-09-10 06:31:22 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.21, #queue-req: 0
|
||||
[2026-09-10 06:31:22 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.61, #queue-req: 0
|
||||
[2026-09-10 06:31:23 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.20, #queue-req: 0
|
||||
[2026-09-10 06:31:23 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.26, #queue-req: 0
|
||||
[2026-09-10 06:31:24 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.26, #queue-req: 0
|
||||
[2026-09-10 06:31:24 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.11, #queue-req: 0
|
||||
[2026-09-10 06:31:25 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.82, #queue-req: 0
|
||||
[2026-09-10 06:31:25 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.41, #queue-req: 0
|
||||
[2026-09-10 06:31:25 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.18, #queue-req: 0
|
||||
[2026-09-10 06:31:26 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.00, #queue-req: 0
|
||||
[2026-09-10 06:31:26 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.42, #queue-req: 0
|
||||
[2026-09-10 06:31:27 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.93, #queue-req: 0
|
||||
[2026-09-10 06:31:27 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.20, #queue-req: 0
|
||||
[2026-09-10 06:31:28 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.07, #queue-req: 0
|
||||
[2026-09-10 06:31:28 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.46, #queue-req: 0
|
||||
[2026-09-10 06:31:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.21, #queue-req: 0
|
||||
[2026-09-10 06:31:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.66, #queue-req: 0
|
||||
[2026-09-10 06:31:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.66, #queue-req: 0
|
||||
[2026-09-10 06:31:30 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.16, #queue-req: 0
|
||||
[2026-09-10 06:31:30 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.24, #queue-req: 0
|
||||
[2026-09-10 06:31:31 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.15, #queue-req: 0
|
||||
[2026-09-10 06:31:31 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.74, #queue-req: 0
|
||||
[2026-09-10 06:31:32 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.71, #queue-req: 0
|
||||
[2026-09-10 06:31:32 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.75, #queue-req: 0
|
||||
[2026-09-10 06:31:33 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.88, #queue-req: 0
|
||||
[2026-09-10 06:31:33 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.02, #queue-req: 0
|
||||
[2026-09-10 06:31:33 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.77, #queue-req: 0
|
||||
[2026-09-10 06:31:34 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.51, #queue-req: 0
|
||||
[2026-09-10 06:31:34 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.54, #queue-req: 0
|
||||
[2026-09-10 06:31:35 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.98, #queue-req: 0
|
||||
[2026-09-10 06:31:35 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.14, #queue-req: 0
|
||||
[2026-09-10 06:31:36 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.59, #queue-req: 0
|
||||
[2026-09-10 06:31:36 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.02, #queue-req: 0
|
||||
[2026-09-10 06:31:36 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.00, #queue-req: 0
|
||||
[2026-09-10 06:31:37 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.67, #queue-req: 0
|
||||
[2026-09-10 06:31:37 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.41, #queue-req: 0
|
||||
[2026-09-10 06:31:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.77, #queue-req: 0
|
||||
[2026-09-10 06:31:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.84, #queue-req: 0
|
||||
[2026-09-10 06:31:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.46, #queue-req: 0
|
||||
[2026-09-10 06:31:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.39, #queue-req: 0
|
||||
[2026-09-10 06:31:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.14, #queue-req: 0
|
||||
[2026-09-10 06:31:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.00, #queue-req: 0
|
||||
[2026-09-10 06:31:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.31, #queue-req: 0
|
||||
[2026-09-10 06:31:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.85, #queue-req: 0
|
||||
[2026-09-10 06:31:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.19, #queue-req: 0
|
||||
[2026-09-10 06:31:42 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.44, #queue-req: 0
|
||||
[2026-09-10 06:31:42 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.75, #queue-req: 0
|
||||
[2026-09-10 06:31:43 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.96, #queue-req: 0
|
||||
[2026-09-10 06:31:43 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.37, #queue-req: 0
|
||||
[2026-09-10 06:31:44 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.59, #queue-req: 0
|
||||
[2026-09-10 06:31:44 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.70, #queue-req: 0
|
||||
[2026-09-10 06:31:44 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.44, #queue-req: 0
|
||||
[2026-09-10 06:31:45 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.37, #queue-req: 0
|
||||
[2026-09-10 06:31:45 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.00, #queue-req: 0
|
||||
[2026-09-10 06:31:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.82, #queue-req: 0
|
||||
[2026-09-10 06:31:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.76, #queue-req: 0
|
||||
[2026-09-10 06:31:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.64, #queue-req: 0
|
||||
[2026-09-10 06:31:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.25, #queue-req: 0
|
||||
[2026-09-10 06:31:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.44, #queue-req: 0
|
||||
[2026-09-10 06:31:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.67, #queue-req: 0
|
||||
[2026-09-10 06:31:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.53, #queue-req: 0
|
||||
[2026-09-10 06:31:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.30, #queue-req: 0
|
||||
[2026-09-10 06:31:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.72, #queue-req: 0
|
||||
[2026-09-10 06:31:50 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.49, #queue-req: 0
|
||||
[2026-09-10 06:31:50 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.74, #queue-req: 0
|
||||
[2026-09-10 06:31:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.23, #queue-req: 0
|
||||
[2026-09-10 06:31:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.64, #queue-req: 0
|
||||
[2026-09-10 06:31:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.05, #queue-req: 0
|
||||
[2026-09-10 06:31:52 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.81, #queue-req: 0
|
||||
[2026-09-10 06:31:52 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.05, #queue-req: 0
|
||||
[2026-09-10 06:31:53 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.60, #queue-req: 0
|
||||
[2026-09-10 06:31:53 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.62, #queue-req: 0
|
||||
[2026-09-10 06:31:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.41, #queue-req: 0
|
||||
[2026-09-10 06:31:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.43, #queue-req: 0
|
||||
[2026-09-10 06:31:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.65, #queue-req: 0
|
||||
[2026-09-10 06:31:55 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.38, #queue-req: 0
|
||||
[2026-09-10 06:31:55 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.41, #queue-req: 0
|
||||
[2026-09-10 06:31:56 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.56, #queue-req: 0
|
||||
[2026-09-10 06:31:56 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.42, #queue-req: 0
|
||||
[2026-09-10 06:31:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.32, #queue-req: 0
|
||||
[2026-09-10 06:31:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.28, #queue-req: 0
|
||||
[2026-09-10 06:31:58 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.95, #queue-req: 0
|
||||
[2026-09-10 06:31:58 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.38, #queue-req: 0
|
||||
[2026-09-10 06:31:58 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.68, #queue-req: 0
|
||||
[2026-09-10 06:31:59 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.69, #queue-req: 0
|
||||
[2026-09-10 06:31:59 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.57, #queue-req: 0
|
||||
[2026-09-10 06:32:00 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.44, #queue-req: 0
|
||||
[2026-09-10 06:32:00 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.63, #queue-req: 0
|
||||
[2026-09-10 06:32:01 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.79, #queue-req: 0
|
||||
[2026-09-10 06:32:01 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.58, #queue-req: 0
|
||||
[2026-09-10 06:32:02 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.28, #queue-req: 0
|
||||
[2026-09-10 06:32:02 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.37, #queue-req: 0
|
||||
[2026-09-10 06:32:02] INFO: 127.0.0.1:38660 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 06:32:02 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 1024, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 5.68
|
||||
[2026-09-10 06:32:02 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 0.29, #queue-req: 0
|
||||
[2026-09-10 06:32:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 113.83, #queue-req: 0
|
||||
[2026-09-10 06:32:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 100.54, #queue-req: 0
|
||||
[2026-09-10 06:32:04 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.85, #queue-req: 0
|
||||
[2026-09-10 06:32:04 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.67, #queue-req: 0
|
||||
[2026-09-10 06:32:04 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.52, #queue-req: 0
|
||||
[2026-09-10 06:32:05 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.89, #queue-req: 0
|
||||
[2026-09-10 06:32:05 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.27, #queue-req: 0
|
||||
[2026-09-10 06:32:06 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.17, #queue-req: 0
|
||||
[2026-09-10 06:32:06 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.92, #queue-req: 0
|
||||
[2026-09-10 06:32:07 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.44, #queue-req: 0
|
||||
[2026-09-10 06:32:07 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.90, #queue-req: 0
|
||||
[2026-09-10 06:32:07 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.74, #queue-req: 0
|
||||
[2026-09-10 06:32:08 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.41, #queue-req: 0
|
||||
[2026-09-10 06:32:08 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.03, #queue-req: 0
|
||||
[2026-09-10 06:32:09 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.85, #queue-req: 0
|
||||
[2026-09-10 06:32:09 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.27, #queue-req: 0
|
||||
[2026-09-10 06:32:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.74, #queue-req: 0
|
||||
[2026-09-10 06:32:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.30, #queue-req: 0
|
||||
[2026-09-10 06:32:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.39, #queue-req: 0
|
||||
[2026-09-10 06:32:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.37, #queue-req: 0
|
||||
[2026-09-10 06:32:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.44, #queue-req: 0
|
||||
[2026-09-10 06:32:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.98, #queue-req: 0
|
||||
[2026-09-10 06:32:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.34, #queue-req: 0
|
||||
[2026-09-10 06:32:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.58, #queue-req: 0
|
||||
[2026-09-10 06:32:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.13, #queue-req: 0
|
||||
[2026-09-10 06:32:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.12, #queue-req: 0
|
||||
[2026-09-10 06:32:14 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.40, #queue-req: 0
|
||||
[2026-09-10 06:32:14 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.03, #queue-req: 0
|
||||
[2026-09-10 06:32:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.31, #queue-req: 0
|
||||
[2026-09-10 06:32:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.07, #queue-req: 0
|
||||
[2026-09-10 06:32:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.63, #queue-req: 0
|
||||
[2026-09-10 06:32:16 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.60, #queue-req: 0
|
||||
[2026-09-10 06:32:16 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.37, #queue-req: 0
|
||||
[2026-09-10 06:32:17 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.07, #queue-req: 0
|
||||
[2026-09-10 06:32:17 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.26, #queue-req: 0
|
||||
[2026-09-10 06:32:18 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.85, #queue-req: 0
|
||||
[2026-09-10 06:32:18 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.74, #queue-req: 0
|
||||
[2026-09-10 06:32:18 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.82, #queue-req: 0
|
||||
[2026-09-10 06:32:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.86, #queue-req: 0
|
||||
[2026-09-10 06:32:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.31, #queue-req: 0
|
||||
[2026-09-10 06:32:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.95, #queue-req: 0
|
||||
[2026-09-10 06:32:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.25, #queue-req: 0
|
||||
[2026-09-10 06:32:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.22, #queue-req: 0
|
||||
[2026-09-10 06:32:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.04, #queue-req: 0
|
||||
[2026-09-10 06:32:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.75, #queue-req: 0
|
||||
[2026-09-10 06:32:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.93, #queue-req: 0
|
||||
[2026-09-10 06:32:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.29, #queue-req: 0
|
||||
[2026-09-10 06:32:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.54, #queue-req: 0
|
||||
[2026-09-10 06:32:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.25, #queue-req: 0
|
||||
[2026-09-10 06:32:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.14, #queue-req: 0
|
||||
[2026-09-10 06:32:24 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.42, #queue-req: 0
|
||||
[2026-09-10 06:32:24 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.54, #queue-req: 0
|
||||
[2026-09-10 06:32:25 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.87, #queue-req: 0
|
||||
[2026-09-10 06:32:25 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.13, #queue-req: 0
|
||||
[2026-09-10 06:32:26 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.81, #queue-req: 0
|
||||
[2026-09-10 06:32:26 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.34, #queue-req: 0
|
||||
[2026-09-10 06:32:26 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.97, #queue-req: 0
|
||||
[2026-09-10 06:32:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.19, #queue-req: 0
|
||||
[2026-09-10 06:32:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.89, #queue-req: 0
|
||||
[2026-09-10 06:32:28 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.23, #queue-req: 0
|
||||
[2026-09-10 06:32:28 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.99, #queue-req: 0
|
||||
[2026-09-10 06:32:29 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.38, #queue-req: 0
|
||||
[2026-09-10 06:32:29 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.36, #queue-req: 0
|
||||
[2026-09-10 06:32:29 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.74, #queue-req: 0
|
||||
[2026-09-10 06:32:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.43, #queue-req: 0
|
||||
[2026-09-10 06:32:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.30, #queue-req: 0
|
||||
[2026-09-10 06:32:31 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.31, #queue-req: 0
|
||||
[2026-09-10 06:32:31 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.29, #queue-req: 0
|
||||
[2026-09-10 06:32:31 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.17, #queue-req: 0
|
||||
[2026-09-10 06:32:32 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.37, #queue-req: 0
|
||||
[2026-09-10 06:32:32 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.64, #queue-req: 0
|
||||
[2026-09-10 06:32:33 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.00, #queue-req: 0
|
||||
[2026-09-10 06:32:33 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.17, #queue-req: 0
|
||||
[2026-09-10 06:32:34 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.00, #queue-req: 0
|
||||
[2026-09-10 06:32:34 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.59, #queue-req: 0
|
||||
[2026-09-10 06:32:34 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.13, #queue-req: 0
|
||||
[2026-09-10 06:32:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.29, #queue-req: 0
|
||||
[2026-09-10 06:32:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.49, #queue-req: 0
|
||||
[2026-09-10 06:32:36 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.53, #queue-req: 0
|
||||
[2026-09-10 06:32:36 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.90, #queue-req: 0
|
||||
[2026-09-10 06:32:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.64, #queue-req: 0
|
||||
[2026-09-10 06:32:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.35, #queue-req: 0
|
||||
[2026-09-10 06:32:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.59, #queue-req: 0
|
||||
[2026-09-10 06:32:38 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.47, #queue-req: 0
|
||||
[2026-09-10 06:32:38 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.75, #queue-req: 0
|
||||
[2026-09-10 06:32:39 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.68, #queue-req: 0
|
||||
[2026-09-10 06:32:39 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.40, #queue-req: 0
|
||||
[2026-09-10 06:32:39 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.45, #queue-req: 0
|
||||
[2026-09-10 06:32:40 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.91, #queue-req: 0
|
||||
[2026-09-10 06:32:40 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.37, #queue-req: 0
|
||||
[2026-09-10 06:32:41 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.01, #queue-req: 0
|
||||
[2026-09-10 06:32:41 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.94, #queue-req: 0
|
||||
[2026-09-10 06:32:42 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.54, #queue-req: 0
|
||||
[2026-09-10 06:32:42 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.03, #queue-req: 0
|
||||
[2026-09-10 06:32:42 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.33, #queue-req: 0
|
||||
[2026-09-10 06:32:43 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.90, #queue-req: 0
|
||||
[2026-09-10 06:32:43 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.72, #queue-req: 0
|
||||
[2026-09-10 06:32:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.66, #queue-req: 0
|
||||
[2026-09-10 06:32:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 97.54, #queue-req: 0
|
||||
[2026-09-10 06:32:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.84, #queue-req: 0
|
||||
[2026-09-10 06:32:45 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.88, #queue-req: 0
|
||||
[2026-09-10 06:32:45 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.48, #queue-req: 0
|
||||
[2026-09-10 06:32:45] INFO: 127.0.0.1:48278 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 06:32:46 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 1024, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 5.69
|
||||
[2026-09-10 06:32:46 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 0.30, #queue-req: 0
|
||||
[2026-09-10 06:32:46 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 108.73, #queue-req: 0
|
||||
[2026-09-10 06:32:47 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.12, #queue-req: 0
|
||||
[2026-09-10 06:32:47 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.06, #queue-req: 0
|
||||
[2026-09-10 06:32:48 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.76, #queue-req: 0
|
||||
[2026-09-10 06:32:48 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.26, #queue-req: 0
|
||||
[2026-09-10 06:32:48 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.92, #queue-req: 0
|
||||
[2026-09-10 06:32:49 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.72, #queue-req: 0
|
||||
[2026-09-10 06:32:49 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.92, #queue-req: 0
|
||||
[2026-09-10 06:32:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.25, #queue-req: 0
|
||||
[2026-09-10 06:32:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.80, #queue-req: 0
|
||||
[2026-09-10 06:32:51 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.13, #queue-req: 0
|
||||
[2026-09-10 06:32:51 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.39, #queue-req: 0
|
||||
[2026-09-10 06:32:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.15, #queue-req: 0
|
||||
[2026-09-10 06:32:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.85, #queue-req: 0
|
||||
[2026-09-10 06:32:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.47, #queue-req: 0
|
||||
[2026-09-10 06:32:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.16, #queue-req: 0
|
||||
[2026-09-10 06:32:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.35, #queue-req: 0
|
||||
[2026-09-10 06:32:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.05, #queue-req: 0
|
||||
[2026-09-10 06:32:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.81, #queue-req: 0
|
||||
[2026-09-10 06:32:55 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.03, #queue-req: 0
|
||||
[2026-09-10 06:32:55 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.34, #queue-req: 0
|
||||
[2026-09-10 06:32:56 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.52, #queue-req: 0
|
||||
[2026-09-10 06:32:56 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.22, #queue-req: 0
|
||||
[2026-09-10 06:32:57 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.28, #queue-req: 0
|
||||
[2026-09-10 06:32:57 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.95, #queue-req: 0
|
||||
[2026-09-10 06:32:58 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.90, #queue-req: 0
|
||||
[2026-09-10 06:32:58 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.89, #queue-req: 0
|
||||
[2026-09-10 06:32:59 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.44, #queue-req: 0
|
||||
[2026-09-10 06:32:59 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.51, #queue-req: 0
|
||||
[2026-09-10 06:32:59 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.06, #queue-req: 0
|
||||
[2026-09-10 06:33:00 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.37, #queue-req: 0
|
||||
[2026-09-10 06:33:00 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.14, #queue-req: 0
|
||||
[2026-09-10 06:33:01 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.15, #queue-req: 0
|
||||
[2026-09-10 06:33:01 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.82, #queue-req: 0
|
||||
[2026-09-10 06:33:02 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.21, #queue-req: 0
|
||||
[2026-09-10 06:33:02 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.44, #queue-req: 0
|
||||
[2026-09-10 06:33:03 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.18, #queue-req: 0
|
||||
[2026-09-10 06:33:03 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.38, #queue-req: 0
|
||||
[2026-09-10 06:33:04 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.46, #queue-req: 0
|
||||
[2026-09-10 06:33:04 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.63, #queue-req: 0
|
||||
[2026-09-10 06:33:05 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.06, #queue-req: 0
|
||||
[2026-09-10 06:33:05 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.29, #queue-req: 0
|
||||
[2026-09-10 06:33:05 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.75, #queue-req: 0
|
||||
[2026-09-10 06:33:06 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.86, #queue-req: 0
|
||||
[2026-09-10 06:33:06 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.09, #queue-req: 0
|
||||
[2026-09-10 06:33:07 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.30, #queue-req: 0
|
||||
[2026-09-10 06:33:07 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.51, #queue-req: 0
|
||||
[2026-09-10 06:33:08 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.14, #queue-req: 0
|
||||
[2026-09-10 06:33:08 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.59, #queue-req: 0
|
||||
[2026-09-10 06:33:09 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.27, #queue-req: 0
|
||||
[2026-09-10 06:33:09 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.04, #queue-req: 0
|
||||
[2026-09-10 06:33:10 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.42, #queue-req: 0
|
||||
[2026-09-10 06:33:10 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.67, #queue-req: 0
|
||||
[2026-09-10 06:33:11 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.26, #queue-req: 0
|
||||
[2026-09-10 06:33:11 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.97, #queue-req: 0
|
||||
[2026-09-10 06:33:11 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.60, #queue-req: 0
|
||||
[2026-09-10 06:33:12 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 85.61, #queue-req: 0
|
||||
[2026-09-10 06:33:12 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.92, #queue-req: 0
|
||||
[2026-09-10 06:33:13 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.42, #queue-req: 0
|
||||
[2026-09-10 06:33:13 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.88, #queue-req: 0
|
||||
[2026-09-10 06:33:14 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.70, #queue-req: 0
|
||||
[2026-09-10 06:33:14 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.90, #queue-req: 0
|
||||
[2026-09-10 06:33:15 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.45, #queue-req: 0
|
||||
[2026-09-10 06:33:15 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.09, #queue-req: 0
|
||||
[2026-09-10 06:33:16 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.89, #queue-req: 0
|
||||
[2026-09-10 06:33:16 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.88, #queue-req: 0
|
||||
[2026-09-10 06:33:16 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.64, #queue-req: 0
|
||||
[2026-09-10 06:33:17 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.52, #queue-req: 0
|
||||
[2026-09-10 06:33:17 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.45, #queue-req: 0
|
||||
[2026-09-10 06:33:18 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.57, #queue-req: 0
|
||||
[2026-09-10 06:33:18 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.65, #queue-req: 0
|
||||
[2026-09-10 06:33:19 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.70, #queue-req: 0
|
||||
[2026-09-10 06:33:19 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.02, #queue-req: 0
|
||||
[2026-09-10 06:33:20 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.18, #queue-req: 0
|
||||
[2026-09-10 06:33:20 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.64, #queue-req: 0
|
||||
[2026-09-10 06:33:21 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.25, #queue-req: 0
|
||||
[2026-09-10 06:33:21 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.42, #queue-req: 0
|
||||
[2026-09-10 06:33:22 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.23, #queue-req: 0
|
||||
[2026-09-10 06:33:22 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.68, #queue-req: 0
|
||||
[2026-09-10 06:33:22 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.45, #queue-req: 0
|
||||
[2026-09-10 06:33:23 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.95, #queue-req: 0
|
||||
[2026-09-10 06:33:23 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.85, #queue-req: 0
|
||||
[2026-09-10 06:33:24 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.21, #queue-req: 0
|
||||
[2026-09-10 06:33:24 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.10, #queue-req: 0
|
||||
[2026-09-10 06:33:25 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.56, #queue-req: 0
|
||||
[2026-09-10 06:33:25 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.82, #queue-req: 0
|
||||
[2026-09-10 06:33:26 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.35, #queue-req: 0
|
||||
[2026-09-10 06:33:26 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.11, #queue-req: 0
|
||||
[2026-09-10 06:33:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.59, #queue-req: 0
|
||||
[2026-09-10 06:33:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.78, #queue-req: 0
|
||||
[2026-09-10 06:33:27 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.82, #queue-req: 0
|
||||
[2026-09-10 06:33:28 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.32, #queue-req: 0
|
||||
[2026-09-10 06:33:28 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.77, #queue-req: 0
|
||||
[2026-09-10 06:33:29 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.32, #queue-req: 0
|
||||
[2026-09-10 06:33:29 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.10, #queue-req: 0
|
||||
[2026-09-10 06:33:30 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.03, #queue-req: 0
|
||||
[2026-09-10 06:33:30 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.62, #queue-req: 0
|
||||
[2026-09-10 06:33:31 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.21, #queue-req: 0
|
||||
[2026-09-10 06:33:31 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.59, #queue-req: 0
|
||||
[2026-09-10 06:33:32 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 86.37, #queue-req: 0
|
||||
[2026-09-10 06:33:32 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.42, #queue-req: 0
|
||||
[2026-09-10 06:33:33 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 87.81, #queue-req: 0
|
||||
[2026-09-10 06:33:33] INFO: 127.0.0.1:51152 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 06:33:33 DP0 TP0 EP0] Prefill batch, #new-seq: 1, #new-token: 1024, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 5.69
|
||||
[2026-09-10 06:33:33 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 0.29, #queue-req: 0
|
||||
[2026-09-10 06:33:33 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 115.10, #queue-req: 0
|
||||
[2026-09-10 06:33:34 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 103.33, #queue-req: 0
|
||||
[2026-09-10 06:33:34 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.77, #queue-req: 0
|
||||
[2026-09-10 06:33:35 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.59, #queue-req: 0
|
||||
[2026-09-10 06:33:35 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.14, #queue-req: 0
|
||||
[2026-09-10 06:33:35 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.95, #queue-req: 0
|
||||
[2026-09-10 06:33:36 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.67, #queue-req: 0
|
||||
[2026-09-10 06:33:36 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.39, #queue-req: 0
|
||||
[2026-09-10 06:33:37 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.24, #queue-req: 0
|
||||
[2026-09-10 06:33:37 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.22, #queue-req: 0
|
||||
[2026-09-10 06:33:38 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.35, #queue-req: 0
|
||||
[2026-09-10 06:33:38 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.87, #queue-req: 0
|
||||
[2026-09-10 06:33:38 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.94, #queue-req: 0
|
||||
[2026-09-10 06:33:39 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.40, #queue-req: 0
|
||||
[2026-09-10 06:33:39 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.57, #queue-req: 0
|
||||
[2026-09-10 06:33:40 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.60, #queue-req: 0
|
||||
[2026-09-10 06:33:40 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.60, #queue-req: 0
|
||||
[2026-09-10 06:33:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.72, #queue-req: 0
|
||||
[2026-09-10 06:33:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.04, #queue-req: 0
|
||||
[2026-09-10 06:33:41 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.39, #queue-req: 0
|
||||
[2026-09-10 06:33:42 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.57, #queue-req: 0
|
||||
[2026-09-10 06:33:42 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.61, #queue-req: 0
|
||||
[2026-09-10 06:33:43 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.72, #queue-req: 0
|
||||
[2026-09-10 06:33:43 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.33, #queue-req: 0
|
||||
[2026-09-10 06:33:44 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.53, #queue-req: 0
|
||||
[2026-09-10 06:33:44 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.72, #queue-req: 0
|
||||
[2026-09-10 06:33:44 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.95, #queue-req: 0
|
||||
[2026-09-10 06:33:45 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.11, #queue-req: 0
|
||||
[2026-09-10 06:33:45 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.38, #queue-req: 0
|
||||
[2026-09-10 06:33:46 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.42, #queue-req: 0
|
||||
[2026-09-10 06:33:46 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.90, #queue-req: 0
|
||||
[2026-09-10 06:33:47 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.46, #queue-req: 0
|
||||
[2026-09-10 06:33:47 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.18, #queue-req: 0
|
||||
[2026-09-10 06:33:47 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.49, #queue-req: 0
|
||||
[2026-09-10 06:33:48 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.37, #queue-req: 0
|
||||
[2026-09-10 06:33:48 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.79, #queue-req: 0
|
||||
[2026-09-10 06:33:49 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.72, #queue-req: 0
|
||||
[2026-09-10 06:33:49 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.94, #queue-req: 0
|
||||
[2026-09-10 06:33:50 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.32, #queue-req: 0
|
||||
[2026-09-10 06:33:50 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.05, #queue-req: 0
|
||||
[2026-09-10 06:33:50 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.49, #queue-req: 0
|
||||
[2026-09-10 06:33:51 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.43, #queue-req: 0
|
||||
[2026-09-10 06:33:51 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.50, #queue-req: 0
|
||||
[2026-09-10 06:33:52 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.27, #queue-req: 0
|
||||
[2026-09-10 06:33:52 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.46, #queue-req: 0
|
||||
[2026-09-10 06:33:53 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.68, #queue-req: 0
|
||||
[2026-09-10 06:33:53 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.32, #queue-req: 0
|
||||
[2026-09-10 06:33:53 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.72, #queue-req: 0
|
||||
[2026-09-10 06:33:54 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.74, #queue-req: 0
|
||||
[2026-09-10 06:33:54 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.42, #queue-req: 0
|
||||
[2026-09-10 06:33:55 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.57, #queue-req: 0
|
||||
[2026-09-10 06:33:55 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.38, #queue-req: 0
|
||||
[2026-09-10 06:33:56 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.27, #queue-req: 0
|
||||
[2026-09-10 06:33:56 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.74, #queue-req: 0
|
||||
[2026-09-10 06:33:56 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.03, #queue-req: 0
|
||||
[2026-09-10 06:33:57 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.65, #queue-req: 0
|
||||
[2026-09-10 06:33:57 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.29, #queue-req: 0
|
||||
[2026-09-10 06:33:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.18, #queue-req: 0
|
||||
[2026-09-10 06:33:58 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.70, #queue-req: 0
|
||||
[2026-09-10 06:33:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.08, #queue-req: 0
|
||||
[2026-09-10 06:33:59 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.33, #queue-req: 0
|
||||
[2026-09-10 06:34:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.83, #queue-req: 0
|
||||
[2026-09-10 06:34:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.83, #queue-req: 0
|
||||
[2026-09-10 06:34:00 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.22, #queue-req: 0
|
||||
[2026-09-10 06:34:01 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.23, #queue-req: 0
|
||||
[2026-09-10 06:34:01 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.73, #queue-req: 0
|
||||
[2026-09-10 06:34:02 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.19, #queue-req: 0
|
||||
[2026-09-10 06:34:02 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.22, #queue-req: 0
|
||||
[2026-09-10 06:34:03 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.58, #queue-req: 0
|
||||
[2026-09-10 06:34:03 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.51, #queue-req: 0
|
||||
[2026-09-10 06:34:03 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.98, #queue-req: 0
|
||||
[2026-09-10 06:34:04 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.07, #queue-req: 0
|
||||
[2026-09-10 06:34:04 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.02, #queue-req: 0
|
||||
[2026-09-10 06:34:05 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.21, #queue-req: 0
|
||||
[2026-09-10 06:34:05 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.61, #queue-req: 0
|
||||
[2026-09-10 06:34:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.61, #queue-req: 0
|
||||
[2026-09-10 06:34:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.52, #queue-req: 0
|
||||
[2026-09-10 06:34:06 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.94, #queue-req: 0
|
||||
[2026-09-10 06:34:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.06, #queue-req: 0
|
||||
[2026-09-10 06:34:07 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.76, #queue-req: 0
|
||||
[2026-09-10 06:34:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.22, #queue-req: 0
|
||||
[2026-09-10 06:34:08 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.38, #queue-req: 0
|
||||
[2026-09-10 06:34:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.49, #queue-req: 0
|
||||
[2026-09-10 06:34:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.06, #queue-req: 0
|
||||
[2026-09-10 06:34:09 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.16, #queue-req: 0
|
||||
[2026-09-10 06:34:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.26, #queue-req: 0
|
||||
[2026-09-10 06:34:10 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.77, #queue-req: 0
|
||||
[2026-09-10 06:34:11 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.94, #queue-req: 0
|
||||
[2026-09-10 06:34:11 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.93, #queue-req: 0
|
||||
[2026-09-10 06:34:12 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.05, #queue-req: 0
|
||||
[2026-09-10 06:34:12 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.90, #queue-req: 0
|
||||
[2026-09-10 06:34:12 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.46, #queue-req: 0
|
||||
[2026-09-10 06:34:13 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.32, #queue-req: 0
|
||||
[2026-09-10 06:34:13 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.37, #queue-req: 0
|
||||
[2026-09-10 06:34:14 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.52, #queue-req: 0
|
||||
[2026-09-10 06:34:14 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.31, #queue-req: 0
|
||||
[2026-09-10 06:34:15 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.06, #queue-req: 0
|
||||
[2026-09-10 06:34:15 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.78, #queue-req: 0
|
||||
[2026-09-10 06:34:15 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.75, #queue-req: 0
|
||||
[2026-09-10 06:34:16 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.05, #queue-req: 0
|
||||
[2026-09-10 06:34:16 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.38, #queue-req: 0
|
||||
[2026-09-10 06:34:17 DP0 TP0 EP0] Decode batch, #running-req: 1, #full token: 0, full token usage: 0.00, #swa token: 0, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.75, #queue-req: 0
|
||||
[2026-09-10 06:34:17] INFO: 127.0.0.1:39600 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 06:34:17 DP1 TP1 EP1] Prefill batch, #new-seq: 1, #new-token: 1024, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 5.69
|
||||
[2026-09-10 06:34:17 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 0.30, #queue-req: 0
|
||||
[2026-09-10 06:34:18 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 106.96, #queue-req: 0
|
||||
[2026-09-10 06:34:18 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.12, #queue-req: 0
|
||||
[2026-09-10 06:34:19 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.34, #queue-req: 0
|
||||
[2026-09-10 06:34:19 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.06, #queue-req: 0
|
||||
[2026-09-10 06:34:19 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.04, #queue-req: 0
|
||||
[2026-09-10 06:34:20 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.36, #queue-req: 0
|
||||
[2026-09-10 06:34:20 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.85, #queue-req: 0
|
||||
[2026-09-10 06:34:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.01, #queue-req: 0
|
||||
[2026-09-10 06:34:21 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.91, #queue-req: 0
|
||||
[2026-09-10 06:34:22 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.79, #queue-req: 0
|
||||
[2026-09-10 06:34:22 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.36, #queue-req: 0
|
||||
[2026-09-10 06:34:23 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.04, #queue-req: 0
|
||||
[2026-09-10 06:34:23 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.78, #queue-req: 0
|
||||
[2026-09-10 06:34:23 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.07, #queue-req: 0
|
||||
[2026-09-10 06:34:24 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.20, #queue-req: 0
|
||||
[2026-09-10 06:34:24 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.39, #queue-req: 0
|
||||
[2026-09-10 06:34:25 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.82, #queue-req: 0
|
||||
[2026-09-10 06:34:25 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.34, #queue-req: 0
|
||||
[2026-09-10 06:34:26 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.30, #queue-req: 0
|
||||
[2026-09-10 06:34:26 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.99, #queue-req: 0
|
||||
[2026-09-10 06:34:26 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.59, #queue-req: 0
|
||||
[2026-09-10 06:34:27 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.94, #queue-req: 0
|
||||
[2026-09-10 06:34:27 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.88, #queue-req: 0
|
||||
[2026-09-10 06:34:28 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.37, #queue-req: 0
|
||||
[2026-09-10 06:34:28 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.92, #queue-req: 0
|
||||
[2026-09-10 06:34:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.74, #queue-req: 0
|
||||
[2026-09-10 06:34:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.37, #queue-req: 0
|
||||
[2026-09-10 06:34:29 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.90, #queue-req: 0
|
||||
[2026-09-10 06:34:30 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.51, #queue-req: 0
|
||||
[2026-09-10 06:34:30 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.64, #queue-req: 0
|
||||
[2026-09-10 06:34:31 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.05, #queue-req: 0
|
||||
[2026-09-10 06:34:31 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.10, #queue-req: 0
|
||||
[2026-09-10 06:34:32 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.74, #queue-req: 0
|
||||
[2026-09-10 06:34:32 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.59, #queue-req: 0
|
||||
[2026-09-10 06:34:32 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.67, #queue-req: 0
|
||||
[2026-09-10 06:34:33 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.73, #queue-req: 0
|
||||
[2026-09-10 06:34:33 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.54, #queue-req: 0
|
||||
[2026-09-10 06:34:34 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.40, #queue-req: 0
|
||||
[2026-09-10 06:34:34 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.97, #queue-req: 0
|
||||
[2026-09-10 06:34:35 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.04, #queue-req: 0
|
||||
[2026-09-10 06:34:35 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.46, #queue-req: 0
|
||||
[2026-09-10 06:34:35 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.07, #queue-req: 0
|
||||
[2026-09-10 06:34:36 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.92, #queue-req: 0
|
||||
[2026-09-10 06:34:36 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.49, #queue-req: 0
|
||||
[2026-09-10 06:34:37 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.24, #queue-req: 0
|
||||
[2026-09-10 06:34:37 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.70, #queue-req: 0
|
||||
[2026-09-10 06:34:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.73, #queue-req: 0
|
||||
[2026-09-10 06:34:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.29, #queue-req: 0
|
||||
[2026-09-10 06:34:38 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.61, #queue-req: 0
|
||||
[2026-09-10 06:34:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.87, #queue-req: 0
|
||||
[2026-09-10 06:34:39 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.14, #queue-req: 0
|
||||
[2026-09-10 06:34:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.98, #queue-req: 0
|
||||
[2026-09-10 06:34:40 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.51, #queue-req: 0
|
||||
[2026-09-10 06:34:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.62, #queue-req: 0
|
||||
[2026-09-10 06:34:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.08, #queue-req: 0
|
||||
[2026-09-10 06:34:41 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.55, #queue-req: 0
|
||||
[2026-09-10 06:34:42 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.53, #queue-req: 0
|
||||
[2026-09-10 06:34:42 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.56, #queue-req: 0
|
||||
[2026-09-10 06:34:43 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.93, #queue-req: 0
|
||||
[2026-09-10 06:34:43 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.28, #queue-req: 0
|
||||
[2026-09-10 06:34:44 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.53, #queue-req: 0
|
||||
[2026-09-10 06:34:44 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.62, #queue-req: 0
|
||||
[2026-09-10 06:34:45 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.41, #queue-req: 0
|
||||
[2026-09-10 06:34:45 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.00, #queue-req: 0
|
||||
[2026-09-10 06:34:45 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.07, #queue-req: 0
|
||||
[2026-09-10 06:34:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.77, #queue-req: 0
|
||||
[2026-09-10 06:34:46 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.57, #queue-req: 0
|
||||
[2026-09-10 06:34:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.25, #queue-req: 0
|
||||
[2026-09-10 06:34:47 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.84, #queue-req: 0
|
||||
[2026-09-10 06:34:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.50, #queue-req: 0
|
||||
[2026-09-10 06:34:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.39, #queue-req: 0
|
||||
[2026-09-10 06:34:48 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.37, #queue-req: 0
|
||||
[2026-09-10 06:34:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.57, #queue-req: 0
|
||||
[2026-09-10 06:34:49 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.39, #queue-req: 0
|
||||
[2026-09-10 06:34:50 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.12, #queue-req: 0
|
||||
[2026-09-10 06:34:50 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.00, #queue-req: 0
|
||||
[2026-09-10 06:34:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.26, #queue-req: 0
|
||||
[2026-09-10 06:34:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.50, #queue-req: 0
|
||||
[2026-09-10 06:34:51 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.04, #queue-req: 0
|
||||
[2026-09-10 06:34:52 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.96, #queue-req: 0
|
||||
[2026-09-10 06:34:52 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.30, #queue-req: 0
|
||||
[2026-09-10 06:34:53 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.03, #queue-req: 0
|
||||
[2026-09-10 06:34:53 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.75, #queue-req: 0
|
||||
[2026-09-10 06:34:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.30, #queue-req: 0
|
||||
[2026-09-10 06:34:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.19, #queue-req: 0
|
||||
[2026-09-10 06:34:54 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 91.72, #queue-req: 0
|
||||
[2026-09-10 06:34:55 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.26, #queue-req: 0
|
||||
[2026-09-10 06:34:55 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.37, #queue-req: 0
|
||||
[2026-09-10 06:34:56 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.94, #queue-req: 0
|
||||
[2026-09-10 06:34:56 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.95, #queue-req: 0
|
||||
[2026-09-10 06:34:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.30, #queue-req: 0
|
||||
[2026-09-10 06:34:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.59, #queue-req: 0
|
||||
[2026-09-10 06:34:57 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.80, #queue-req: 0
|
||||
[2026-09-10 06:34:58 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.80, #queue-req: 0
|
||||
[2026-09-10 06:34:58 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.14, #queue-req: 0
|
||||
[2026-09-10 06:34:59 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.51, #queue-req: 0
|
||||
[2026-09-10 06:34:59 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.34, #queue-req: 0
|
||||
[2026-09-10 06:35:00 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.27, #queue-req: 0
|
||||
[2026-09-10 06:35:00 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.77, #queue-req: 0
|
||||
[2026-09-10 06:35:01 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.17, #queue-req: 0
|
||||
[2026-09-10 06:35:01 DP1 TP1 EP1] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 92.18, #queue-req: 0
|
||||
[2026-09-10 06:35:01] INFO: 127.0.0.1:52588 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 06:35:01 DP2 TP2 EP2] Prefill batch, #new-seq: 1, #new-token: 1024, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 5.72
|
||||
[2026-09-10 06:35:02 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 0.29, #queue-req: 0
|
||||
[2026-09-10 06:35:02 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 109.85, #queue-req: 0
|
||||
[2026-09-10 06:35:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 98.85, #queue-req: 0
|
||||
[2026-09-10 06:35:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.30, #queue-req: 0
|
||||
[2026-09-10 06:35:03 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.74, #queue-req: 0
|
||||
[2026-09-10 06:35:04 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.12, #queue-req: 0
|
||||
[2026-09-10 06:35:04 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.79, #queue-req: 0
|
||||
[2026-09-10 06:35:05 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.26, #queue-req: 0
|
||||
[2026-09-10 06:35:05 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.94, #queue-req: 0
|
||||
[2026-09-10 06:35:06 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.01, #queue-req: 0
|
||||
[2026-09-10 06:35:06 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.90, #queue-req: 0
|
||||
[2026-09-10 06:35:06 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.88, #queue-req: 0
|
||||
[2026-09-10 06:35:07 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.23, #queue-req: 0
|
||||
[2026-09-10 06:35:07 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.51, #queue-req: 0
|
||||
[2026-09-10 06:35:08 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.30, #queue-req: 0
|
||||
[2026-09-10 06:35:08 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.36, #queue-req: 0
|
||||
[2026-09-10 06:35:08 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.24, #queue-req: 0
|
||||
[2026-09-10 06:35:09 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.20, #queue-req: 0
|
||||
[2026-09-10 06:35:09 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.66, #queue-req: 0
|
||||
[2026-09-10 06:35:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.50, #queue-req: 0
|
||||
[2026-09-10 06:35:10 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.96, #queue-req: 0
|
||||
[2026-09-10 06:35:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.51, #queue-req: 0
|
||||
[2026-09-10 06:35:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.20, #queue-req: 0
|
||||
[2026-09-10 06:35:11 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.11, #queue-req: 0
|
||||
[2026-09-10 06:35:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.99, #queue-req: 0
|
||||
[2026-09-10 06:35:12 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.19, #queue-req: 0
|
||||
[2026-09-10 06:35:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.15, #queue-req: 0
|
||||
[2026-09-10 06:35:13 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.04, #queue-req: 0
|
||||
[2026-09-10 06:35:14 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.83, #queue-req: 0
|
||||
[2026-09-10 06:35:14 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.67, #queue-req: 0
|
||||
[2026-09-10 06:35:14 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2304, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.77, #queue-req: 0
|
||||
[2026-09-10 06:35:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.68, #queue-req: 0
|
||||
[2026-09-10 06:35:15 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.44, #queue-req: 0
|
||||
[2026-09-10 06:35:16 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.29, #queue-req: 0
|
||||
[2026-09-10 06:35:16 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.64, #queue-req: 0
|
||||
[2026-09-10 06:35:16 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.98, #queue-req: 0
|
||||
[2026-09-10 06:35:17 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.42, #queue-req: 0
|
||||
[2026-09-10 06:35:17 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2560, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.50, #queue-req: 0
|
||||
[2026-09-10 06:35:18 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.41, #queue-req: 0
|
||||
[2026-09-10 06:35:18 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.59, #queue-req: 0
|
||||
[2026-09-10 06:35:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.34, #queue-req: 0
|
||||
[2026-09-10 06:35:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.24, #queue-req: 0
|
||||
[2026-09-10 06:35:19 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.86, #queue-req: 0
|
||||
[2026-09-10 06:35:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 2816, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.89, #queue-req: 0
|
||||
[2026-09-10 06:35:20 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.38, #queue-req: 0
|
||||
[2026-09-10 06:35:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.60, #queue-req: 0
|
||||
[2026-09-10 06:35:21 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.38, #queue-req: 0
|
||||
[2026-09-10 06:35:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.16, #queue-req: 0
|
||||
[2026-09-10 06:35:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.52, #queue-req: 0
|
||||
[2026-09-10 06:35:22 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.44, #queue-req: 0
|
||||
[2026-09-10 06:35:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3072, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.95, #queue-req: 0
|
||||
[2026-09-10 06:35:23 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.45, #queue-req: 0
|
||||
[2026-09-10 06:35:24 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.96, #queue-req: 0
|
||||
[2026-09-10 06:35:24 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.19, #queue-req: 0
|
||||
[2026-09-10 06:35:24 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.76, #queue-req: 0
|
||||
[2026-09-10 06:35:25 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.02, #queue-req: 0
|
||||
[2026-09-10 06:35:25 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3328, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.91, #queue-req: 0
|
||||
[2026-09-10 06:35:26 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.48, #queue-req: 0
|
||||
[2026-09-10 06:35:26 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.15, #queue-req: 0
|
||||
[2026-09-10 06:35:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.59, #queue-req: 0
|
||||
[2026-09-10 06:35:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.88, #queue-req: 0
|
||||
[2026-09-10 06:35:27 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.53, #queue-req: 0
|
||||
[2026-09-10 06:35:28 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3584, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.74, #queue-req: 0
|
||||
[2026-09-10 06:35:28 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.50, #queue-req: 0
|
||||
[2026-09-10 06:35:29 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.43, #queue-req: 0
|
||||
[2026-09-10 06:35:29 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.77, #queue-req: 0
|
||||
[2026-09-10 06:35:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.26, #queue-req: 0
|
||||
[2026-09-10 06:35:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.63, #queue-req: 0
|
||||
[2026-09-10 06:35:30 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.08, #queue-req: 0
|
||||
[2026-09-10 06:35:31 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 3840, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.20, #queue-req: 0
|
||||
[2026-09-10 06:35:31 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.97, #queue-req: 0
|
||||
[2026-09-10 06:35:32 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.81, #queue-req: 0
|
||||
[2026-09-10 06:35:32 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.17, #queue-req: 0
|
||||
[2026-09-10 06:35:32 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.46, #queue-req: 0
|
||||
[2026-09-10 06:35:33 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.11, #queue-req: 0
|
||||
[2026-09-10 06:35:33 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4096, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.96, #queue-req: 0
|
||||
[2026-09-10 06:35:34 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.02, #queue-req: 0
|
||||
[2026-09-10 06:35:34 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.95, #queue-req: 0
|
||||
[2026-09-10 06:35:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.12, #queue-req: 0
|
||||
[2026-09-10 06:35:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.18, #queue-req: 0
|
||||
[2026-09-10 06:35:35 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.97, #queue-req: 0
|
||||
[2026-09-10 06:35:36 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.36, #queue-req: 0
|
||||
[2026-09-10 06:35:36 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4352, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.52, #queue-req: 0
|
||||
[2026-09-10 06:35:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.87, #queue-req: 0
|
||||
[2026-09-10 06:35:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.90, #queue-req: 0
|
||||
[2026-09-10 06:35:37 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.15, #queue-req: 0
|
||||
[2026-09-10 06:35:38 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.33, #queue-req: 0
|
||||
[2026-09-10 06:35:38 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.95, #queue-req: 0
|
||||
[2026-09-10 06:35:39 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4608, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 96.19, #queue-req: 0
|
||||
[2026-09-10 06:35:39 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.97, #queue-req: 0
|
||||
[2026-09-10 06:35:40 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.36, #queue-req: 0
|
||||
[2026-09-10 06:35:40 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.08, #queue-req: 0
|
||||
[2026-09-10 06:35:40 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.82, #queue-req: 0
|
||||
[2026-09-10 06:35:41 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.67, #queue-req: 0
|
||||
[2026-09-10 06:35:41 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 4864, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 93.92, #queue-req: 0
|
||||
[2026-09-10 06:35:42 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 1024, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.79, #queue-req: 0
|
||||
[2026-09-10 06:35:42 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.30, #queue-req: 0
|
||||
[2026-09-10 06:35:43 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.60, #queue-req: 0
|
||||
[2026-09-10 06:35:43 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.50, #queue-req: 0
|
||||
[2026-09-10 06:35:43 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.95, #queue-req: 0
|
||||
[2026-09-10 06:35:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.25, #queue-req: 0
|
||||
[2026-09-10 06:35:44 DP2 TP2 EP2] Decode batch, #running-req: 1, #full token: 5120, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 95.69, #queue-req: 0
|
||||
[2026-09-10 06:35:44] INFO: 127.0.0.1:46194 - "POST /generate HTTP/1.1" 200 OK
|
||||
[2026-09-10 06:35:45 DP3 TP3 EP3] Prefill batch, #new-seq: 1, #new-token: 1024, #cached-token: 0, full token usage: 0.00, swa token usage: 0.00, #running-req: 0, #queue-req: 0, #pending-token: 0, cuda graph: False, input throughput (token/s): 5.72
|
||||
[2026-09-10 06:35:45 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 0.30, #queue-req: 0
|
||||
[2026-09-10 06:35:45 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 94.27, #queue-req: 0
|
||||
[2026-09-10 06:35:46 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.65, #queue-req: 0
|
||||
[2026-09-10 06:35:46 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.88, #queue-req: 0
|
||||
[2026-09-10 06:35:47 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.15, #queue-req: 0
|
||||
[2026-09-10 06:35:47 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1280, full token usage: 0.00, #swa token: 512, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.87, #queue-req: 0
|
||||
[2026-09-10 06:35:48 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.11, #queue-req: 0
|
||||
[2026-09-10 06:35:48 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.26, #queue-req: 0
|
||||
[2026-09-10 06:35:49 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.00, #queue-req: 0
|
||||
[2026-09-10 06:35:49 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.81, #queue-req: 0
|
||||
[2026-09-10 06:35:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.91, #queue-req: 0
|
||||
[2026-09-10 06:35:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1536, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.60, #queue-req: 0
|
||||
[2026-09-10 06:35:50 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.63, #queue-req: 0
|
||||
[2026-09-10 06:35:51 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.64, #queue-req: 0
|
||||
[2026-09-10 06:35:51 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 90.12, #queue-req: 0
|
||||
[2026-09-10 06:35:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.53, #queue-req: 0
|
||||
[2026-09-10 06:35:52 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.78, #queue-req: 0
|
||||
[2026-09-10 06:35:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.21, #queue-req: 0
|
||||
[2026-09-10 06:35:53 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 1792, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.99, #queue-req: 0
|
||||
[2026-09-10 06:35:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.84, #queue-req: 0
|
||||
[2026-09-10 06:35:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.96, #queue-req: 0
|
||||
[2026-09-10 06:35:54 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.41, #queue-req: 0
|
||||
[2026-09-10 06:35:55 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.42, #queue-req: 0
|
||||
[2026-09-10 06:35:55 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 89.24, #queue-req: 0
|
||||
[2026-09-10 06:35:56 DP3 TP3 EP3] Decode batch, #running-req: 1, #full token: 2048, full token usage: 0.00, #swa token: 768, swa token usage: 0.00, cuda graph: True, gen throughput (token/s): 88.83, #queue-req: 0
|
||||
@ -0,0 +1,3 @@
|
||||
timestamp=2026-09-10T06:35:57+00:00
|
||||
rc=137
|
||||
timeout_seconds=1800
|
||||
@ -0,0 +1 @@
|
||||
2026-09-10T06:59:43+00:00
|
||||
Some files were not shown because too many files have changed in this diff Show More
Loading…
x
Reference in New Issue
Block a user