Failed to get device capability: SM 12.x requires CUDA >= 12.9. Failed to get device capability: SM 12.x requires CUDA >= 12.9. [08-31 15:10:18] Applying performance_mode=speed [08-31 15:10:18] server_args: {"model_path": "/data/hf_models/MiniMax-H3", "model_subfolder": null, "model_variant": "Ref2VA", "model_id": null, "backend": "sglang", "attention_backend": null, "attention_backend_config": {}, "component_attention_backends": {}, "cache_dit_config": null, "nccl_port": null, "trust_remote_code": false, "revision": null, "num_gpus": 2, "performance_mode": "speed", "base_gpu_id": 0, "gpu_ids": null, "tp_size": 2, "sp_degree": 1, "ulysses_degree": 1, "ring_degree": 1, "dp_size": 1, "dp_degree": 1, "enable_cfg_parallel": false, "cfg_parallel_degree": 1, "encoder_parallel": "auto", "hsdp_replicate_dim": 1, "hsdp_shard_dim": 2, "dist_timeout": 3600, "pipeline_class_name": null, "lora_path": null, "lora_nickname": "default", "lora_scale": 1.0, "lora_merge_mode": "auto", "lora_weight_name": null, "component_paths": {}, "transformer_weights_path": null, "component_transformer_weights_paths": {}, "quantization": null, "quantization_ignored_layers": null, "lora_target_modules": null, "dit_cpu_offload": false, "dit_layerwise_offload": false, "layerwise_offload_components": null, "dit_offload_prefetch_size": 0.0, "dit_layerwise_resident_layers": 0.0, "offload_during_compile": true, "text_encoder_cpu_offload": false, "image_encoder_cpu_offload": false, "vae_cpu_offload": false, "use_fsdp_inference": false, "pin_cpu_memory": true, "ltx2_two_stage_device_mode": null, "comfyui_mode": false, "enable_torch_compile": false, "regional_compile": false, "enable_breakable_cuda_graph": false, "bcg_text_buckets": null, "enable_layerwise_nvtx_marker": false, "warmup_mode": "server", "warmup": true, "server_warmup": true, "warmup_resolutions": null, "warmup_steps": 1, "disable_autocast": false, "master_port": 35010, "host": "0.0.0.0", "port": 34020, "webui": false, "webui_port": 12312, "scheduler_port": 36010, "batching_mode": "dynamic", "batching_max_size": 1, "batching_delay_ms": 0.0, "batching_config": null, "enable_batching_metrics": false, "strict_ports": false, "output_path": "/data/wxy/sskj-h3/throughput/sglang-base/results/ref2va-feishu-base-tp2x4-768p-15s-20steps-20260831-150959/server_1_port34020/outputs", "input_save_path": "inputs/uploads", "prompt_file_path": null, "model_paths": {}, "model_loaded": {"transformer": true, "vae": true, "video_vae": true, "audio_vae": true, "video_dit": true, "audio_dit": true, "dual_tower_bridge": true}, "boundary_ratio": null, "disagg_role": "monolithic", "disagg_timeout": 3600, "disagg_downstream_wait_timeout": 1800, "disagg_dispatch_policy": "round_robin", "disagg_mode": false, "disagg_instance_id": 0, "disagg_max_slots_per_instance": 8, "disagg_transfer_redundancy": 1.25, "disagg_role_device": "auto", "disagg_transfer_backend": "auto", "disagg_transfer_pool_size": 268435456, "disagg_transfer_pin_memory": "auto", "disagg_p2p_hostname": "127.0.0.1", "disagg_ib_device": null, "disagg_server_addr": null, "encoder_urls": null, "denoiser_urls": null, "decoder_urls": null, "encoder_tp": null, "denoiser_tp": null, "denoiser_sp": null, "denoiser_ulysses": null, "denoiser_ring": null, "decoder_sp": null, "decoder_tp": null, "pool_work_endpoint": null, "pool_result_endpoint": null, "pool_control_endpoint": null, "pool_control_advertised_endpoint": null, "log_level": "info", "log_requests": false, "log_requests_level": 2, "log_requests_format": "text", "log_requests_target": null, "uvicorn_access_log_exclude_prefixes": [], "enable_trace": false, "otlp_traces_endpoint": "localhost:4317", "srt_encoder_url": null, "srt_encoder_connect_timeout": 3.05, "srt_encoder_timeout": 100, "pe_server_url": null} [08-31 15:10:18] Starting server... Failed to get device capability: SM 12.x requires CUDA >= 12.9. Failed to get device capability: SM 12.x requires CUDA >= 12.9. Failed to get device capability: SM 12.x requires CUDA >= 12.9. Failed to get device capability: SM 12.x requires CUDA >= 12.9. [08-31 15:10:35] Scheduler bind at endpoint: tcp://0.0.0.0:36010 [08-31 15:10:36] torch.compile cache: TORCHINDUCTOR_CACHE_DIR=/root/.cache/sgl_diffusion/torch_compile_cache/inductor TRITON_CACHE_DIR=/root/.cache/sgl_diffusion/torch_compile_cache/triton [08-31 15:10:36] Initializing distributed environment with world_size=2, device=cuda:0, timeout=3600 [08-31 15:10:36] Setting distributed timeout to 3600 seconds [08-31 15:10:36] Found nccl from library libnccl.so.2 [08-31 15:10:36] sglang-diffusion is using nccl==2.28.9 [08-31 15:10:39] reading GPU P2P access cache from /root/.cache/sglang/gpu_p2p_access_cache_for_2,3.json [08-31 15:10:39] reading GPU P2P access cache from /root/.cache/sglang/gpu_p2p_access_cache_for_2,3.json [08-31 15:10:39] Found nccl from library libnccl.so.2 [08-31 15:10:39] sglang-diffusion is using nccl==2.28.9 [08-31 15:10:39] No pipeline_class_name specified, using model_index.json [08-31 15:10:39] Using pipeline from model_index.json: MiniMaxH3Pipeline [08-31 15:10:39] Loading pipeline modules... [08-31 15:10:39] Model path: /data/hf_models/MiniMax-H3/Ref2VA [08-31 15:10:39] Diffusers version: 0.32.2 [08-31 15:10:39] Loading pipeline modules from config: {'_class_name': 'MiniMaxH3Pipeline', '_diffusers_version': '0.32.2', 'text_encoder': ['transformers', 'MiniMaxH3Qwen3VLHFEncoder'], 'tokenizer': ['transformers', 'Qwen2TokenizerFast'], 'video_vae': ['diffusers', 'MiniMaxH3VideoVAE'], 'audio_vae': ['diffusers', 'MiniMaxH3AudioVAE'], 'scheduler': None, 'transformer': ['diffusers', 'MiniMaxH3DiTModel'], 'processor': ['transformers', 'Qwen3VLProcessor'], '_minimax_h3': {'schema_version': 1, 'partition': 'ref2va', 'tasks': ['ref2va'], 'task_aliases': {}, 'sigma_shift_scales': {'video': 12.0, 'audio': 3.0}}} [08-31 15:10:39] Loading required components: ['processor', 'text_encoder', 'tokenizer', 'video_vae', 'audio_vae', 'transformer'] [08-31 15:10:39] Memory-aware component load order: ['text_encoder', 'transformer', 'audio_vae', 'video_vae', 'processor', 'tokenizer'] Loading required modules: 0%| | 0/6 [00:00 neg_prompt: None seed: 1102 infer_steps: 5 num_outputs_per_prompt: 1 guidance_scale: 1.0 embedded_guidance_scale: 6.0 n_tokens: None flow_shift: 12.0 image_path: None save_output: True output_file_path: /data/wxy/sskj-h3/throughput/sglang-base/results/ref2va-feishu-base-tp2x4-768p-15s-20steps-20260831-150959/server_1_port34020/outputs/c703762c-465c-4991-8ccd-7c4623250706.mp4 [08-31 15:11:45] Running pipeline stages: ['InputValidationStage', 'MiniMaxH3PartitionAdmissionStage', 'MiniMaxH3TextEncodingStage', 'MiniMaxH3VisualEncodingStage', 'MiniMaxH3AudioEncodingStage', 'MiniMaxH3LatentPreparationStage', 'MiniMaxH3TimestepPreparationStage', 'MiniMaxH3DenoisingStage', 'MiniMaxH3DecodingStage'] [08-31 15:11:45] [InputValidationStage] started... [08-31 15:11:45] [InputValidationStage] finished in 0.0001 seconds [08-31 15:11:45] [MiniMaxH3PartitionAdmissionStage] started... [08-31 15:11:45] [MiniMaxH3PartitionAdmissionStage] finished in 0.0001 seconds [08-31 15:11:45] [MiniMaxH3TextEncodingStage] started... [2026-08-31 15:11:46] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:11:47] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:11:48] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:11:49] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:11:50] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:11:51] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:11:52] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:11:53] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:11:54] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:11:55] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [08-31 15:11:55] [MiniMaxH3TextEncodingStage] finished in 10.1368 seconds [2026-08-31 15:11:56] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:11:57] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [08-31 15:11:58] [MiniMaxH3VisualEncodingStage] started... [2026-08-31 15:11:58] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:11:59] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:12:00] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:12:01] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:12:02] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [08-31 15:12:03] [MiniMaxH3VisualEncodingStage] finished in 4.8130 seconds [08-31 15:12:03] [MiniMaxH3AudioEncodingStage] started... [2026-08-31 15:12:03] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [08-31 15:12:04] [MiniMaxH3AudioEncodingStage] finished in 0.4585 seconds [08-31 15:12:04] [MiniMaxH3LatentPreparationStage] started... [08-31 15:12:04] [MiniMaxH3LatentPreparationStage] finished in 0.0984 seconds [08-31 15:12:04] [MiniMaxH3TimestepPreparationStage] started... [08-31 15:12:04] [MiniMaxH3TimestepPreparationStage] finished in 0.0005 seconds [08-31 15:12:04] [MiniMaxH3DenoisingStage] started... [2026-08-31 15:12:04] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:12:05] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK minimax_h3 denoise: 0%| | 0/4 [00:00 save_output_paths=lambda output_batch: self._save_output_paths( ^^^^^^^^^^^^^^^^^^^^^^^^ File "/root/.miniconda3/envs/sglang/lib/python3.12/site-packages/sglang/multimodal_gen/runtime/managers/gpu_worker.py", line 651, in _save_output_paths output_batch.output_file_paths = save_outputs( ^^^^^^^^^^^^^ File "/root/.miniconda3/envs/sglang/lib/python3.12/site-packages/sglang/multimodal_gen/runtime/entrypoints/utils.py", line 1057, in save_outputs frames = post_process_sample( ^^^^^^^^^^^^^^^^^^^^ File "/root/.miniconda3/envs/sglang/lib/python3.12/site-packages/sglang/multimodal_gen/runtime/entrypoints/utils.py", line 1107, in post_process_sample materialized = materialize_output_sample( ^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/root/.miniconda3/envs/sglang/lib/python3.12/site-packages/sglang/multimodal_gen/runtime/entrypoints/utils.py", line 888, in materialize_output_sample frames = _sample_to_uint8_frames(sample_without_audio) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/root/.miniconda3/envs/sglang/lib/python3.12/site-packages/sglang/multimodal_gen/runtime/entrypoints/utils.py", line 827, in _sample_to_uint8_frames sample = (sample * 255).clamp(0, 255).to(torch.uint8) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^ torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 4.18 GiB. GPU 0 has a total capacity of 83.05 GiB of which 3.99 GiB is free. Including non-PyTorch memory, this process has 79.04 GiB memory in use. Of the allocated memory 70.81 GiB is allocated by PyTorch, and 4.02 GiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://docs.pytorch.org/docs/stable/notes/cuda.html#optimizing-memory-usage-with-pytorch-cuda-alloc-conf) [08-31 15:23:01] OOM detected. Possible solutions: - If the OOM occurs during loading: 1. Check available memory on every selected GPU, not only total capacity. In multi-GPU runs, the least-free selected GPU is the bottleneck. 2. For single-GPU deployment, use `--performance-mode memory`, component CPU offload, or `--dit-layerwise-offload` for supported Wan/MOVA DiTs. 3. For multi-GPU deployment, keep the default `--performance-mode auto` or set `--use-fsdp-inference true` to shard DiT weights with FSDP. FSDP is not a single-GPU substitute for CPU offload. - If the OOM occurs during runtime: 1. Reduce resolution, `--num-frames`, or batch size. 2. Use `--performance-mode memory` for lower memory usage. 3. Enable SP/Ulysses/Ring for sequence-heavy workloads in multi-GPU setups. 4. Use FSDP, with CFG parallelism when supported, for validated multi-GPU workloads. 5. Use a lower-memory attention backend or quantization when available. Or, open an issue on GitHub https://github.com/sgl-project/sglang/issues/new/choose  [2026-08-31 15:23:02] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:23:03] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:23:04] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:23:05] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:23:06] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:23:07] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:23:08] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:23:09] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:23:10] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:23:11] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:23:12] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:23:13] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [08-31 15:23:25] Failed to generate output for prompt: CUDA out of memory. Tried to allocate 4.18 GiB. GPU 0 has a total capacity of 83.05 GiB of which 3.67 GiB is free. Including non-PyTorch memory, this process has 4.65 GiB memory in use. Process 2258197 has 74.71 GiB memory in use. Of the allocated memory 4.18 GiB is allocated by PyTorch, and 18.19 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://docs.pytorch.org/docs/stable/notes/cuda.html#optimizing-memory-usage-with-pytorch-cuda-alloc-conf) Traceback (most recent call last): File "/root/.miniconda3/envs/sglang/lib/python3.12/site-packages/sglang/multimodal_gen/runtime/utils/logging_utils.py", line 628, in log_generation_timer yield timer File "/root/.miniconda3/envs/sglang/lib/python3.12/site-packages/sglang/multimodal_gen/runtime/entrypoints/openai/utils.py", line 368, in process_generation_batch save_file_path_list = save_outputs( ^^^^^^^^^^^^^ File "/root/.miniconda3/envs/sglang/lib/python3.12/site-packages/sglang/multimodal_gen/runtime/entrypoints/utils.py", line 1057, in save_outputs frames = post_process_sample( ^^^^^^^^^^^^^^^^^^^^ File "/root/.miniconda3/envs/sglang/lib/python3.12/site-packages/sglang/multimodal_gen/runtime/entrypoints/utils.py", line 1107, in post_process_sample materialized = materialize_output_sample( ^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/root/.miniconda3/envs/sglang/lib/python3.12/site-packages/sglang/multimodal_gen/runtime/entrypoints/utils.py", line 888, in materialize_output_sample frames = _sample_to_uint8_frames(sample_without_audio) ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ File "/root/.miniconda3/envs/sglang/lib/python3.12/site-packages/sglang/multimodal_gen/runtime/entrypoints/utils.py", line 827, in _sample_to_uint8_frames sample = (sample * 255).clamp(0, 255).to(torch.uint8) ~~~~~~~^~~~~ torch.OutOfMemoryError: CUDA out of memory. Tried to allocate 4.18 GiB. GPU 0 has a total capacity of 83.05 GiB of which 3.67 GiB is free. Including non-PyTorch memory, this process has 4.65 GiB memory in use. Process 2258197 has 74.71 GiB memory in use. Of the allocated memory 4.18 GiB is allocated by PyTorch, and 18.19 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://docs.pytorch.org/docs/stable/notes/cuda.html#optimizing-memory-usage-with-pytorch-cuda-alloc-conf) [08-31 15:23:25] CUDA out of memory. Tried to allocate 4.18 GiB. GPU 0 has a total capacity of 83.05 GiB of which 3.67 GiB is free. Including non-PyTorch memory, this process has 4.65 GiB memory in use. Process 2258197 has 74.71 GiB memory in use. Of the allocated memory 4.18 GiB is allocated by PyTorch, and 18.19 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://docs.pytorch.org/docs/stable/notes/cuda.html#optimizing-memory-usage-with-pytorch-cuda-alloc-conf) [2026-08-31 15:23:25] INFO: 127.0.0.1:38202 - "GET /v1/videos/c703762c-465c-4991-8ccd-7c4623250706 HTTP/1.1" 200 OK [2026-08-31 15:28:10] INFO: Shutting down [2026-08-31 15:28:10] INFO: Waiting for application shutdown. [08-31 15:28:10] FastAPI app is shutting down... [2026-08-31 15:28:10] INFO: Application shutdown complete. [2026-08-31 15:28:10] INFO: Finished server process [2257455] [08-31 15:28:14] Worker 0: Shutdown complete. [08-31 15:28:17] kill_process_tree called: parent_pid=2257455, include_parent=False, pid=2257455