104 lines
13 KiB
Plaintext
104 lines
13 KiB
Plaintext
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
|
||
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
|
||
[08-31 15:03:40] Applying performance_mode=speed
|
||
[08-31 15:03:40] server_args: {"model_path": "/data/hf_models/MiniMax-H3", "model_subfolder": null, "model_variant": "Ref2VA", "model_id": null, "backend": "sglang", "attention_backend": null, "attention_backend_config": {}, "component_attention_backends": {}, "cache_dit_config": null, "nccl_port": null, "trust_remote_code": false, "revision": null, "num_gpus": 2, "performance_mode": "speed", "base_gpu_id": 0, "gpu_ids": null, "tp_size": 2, "sp_degree": 1, "ulysses_degree": 1, "ring_degree": 1, "dp_size": 1, "dp_degree": 1, "enable_cfg_parallel": false, "cfg_parallel_degree": 1, "encoder_parallel": "auto", "hsdp_replicate_dim": 1, "hsdp_shard_dim": 2, "dist_timeout": 3600, "pipeline_class_name": null, "lora_path": null, "lora_nickname": "default", "lora_scale": 1.0, "lora_merge_mode": "auto", "lora_weight_name": null, "component_paths": {}, "transformer_weights_path": null, "component_transformer_weights_paths": {}, "quantization": null, "quantization_ignored_layers": null, "lora_target_modules": null, "dit_cpu_offload": false, "dit_layerwise_offload": false, "layerwise_offload_components": null, "dit_offload_prefetch_size": 0.0, "dit_layerwise_resident_layers": 0.0, "offload_during_compile": true, "text_encoder_cpu_offload": false, "image_encoder_cpu_offload": false, "vae_cpu_offload": false, "use_fsdp_inference": false, "pin_cpu_memory": true, "ltx2_two_stage_device_mode": null, "comfyui_mode": false, "enable_torch_compile": false, "regional_compile": false, "enable_breakable_cuda_graph": false, "bcg_text_buckets": null, "enable_layerwise_nvtx_marker": false, "warmup_mode": "server", "warmup": true, "server_warmup": true, "warmup_resolutions": null, "warmup_steps": 1, "disable_autocast": false, "master_port": 35030, "host": "0.0.0.0", "port": 34040, "webui": false, "webui_port": 12312, "scheduler_port": 36030, "batching_mode": "dynamic", "batching_max_size": 1, "batching_delay_ms": 0.0, "batching_config": null, "enable_batching_metrics": false, "strict_ports": false, "output_path": "/data/wxy/sskj-h3/throughput/sglang-base/results/ref2va-feishu-base-tp2x4-768p-15s-20steps-20260831-150318/server_3_port34040/outputs", "input_save_path": "inputs/uploads", "prompt_file_path": null, "model_paths": {}, "model_loaded": {"transformer": true, "vae": true, "video_vae": true, "audio_vae": true, "video_dit": true, "audio_dit": true, "dual_tower_bridge": true}, "boundary_ratio": null, "disagg_role": "monolithic", "disagg_timeout": 3600, "disagg_downstream_wait_timeout": 1800, "disagg_dispatch_policy": "round_robin", "disagg_mode": false, "disagg_instance_id": 0, "disagg_max_slots_per_instance": 8, "disagg_transfer_redundancy": 1.25, "disagg_role_device": "auto", "disagg_transfer_backend": "auto", "disagg_transfer_pool_size": 268435456, "disagg_transfer_pin_memory": "auto", "disagg_p2p_hostname": "127.0.0.1", "disagg_ib_device": null, "disagg_server_addr": null, "encoder_urls": null, "denoiser_urls": null, "decoder_urls": null, "encoder_tp": null, "denoiser_tp": null, "denoiser_sp": null, "denoiser_ulysses": null, "denoiser_ring": null, "decoder_sp": null, "decoder_tp": null, "pool_work_endpoint": null, "pool_result_endpoint": null, "pool_control_endpoint": null, "pool_control_advertised_endpoint": null, "log_level": "info", "log_requests": false, "log_requests_level": 2, "log_requests_format": "text", "log_requests_target": null, "uvicorn_access_log_exclude_prefixes": [], "enable_trace": false, "otlp_traces_endpoint": "localhost:4317", "srt_encoder_url": null, "srt_encoder_connect_timeout": 3.05, "srt_encoder_timeout": 100, "pe_server_url": null}
|
||
[08-31 15:03:40] Starting server...
|
||
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
|
||
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
|
||
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
|
||
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
|
||
[08-31 15:03:59] Scheduler bind at endpoint: tcp://0.0.0.0:36030
|
||
[08-31 15:03:59] torch.compile cache: TORCHINDUCTOR_CACHE_DIR=/root/.cache/sgl_diffusion/torch_compile_cache/inductor TRITON_CACHE_DIR=/root/.cache/sgl_diffusion/torch_compile_cache/triton
|
||
[08-31 15:03:59] Initializing distributed environment with world_size=2, device=cuda:0, timeout=3600
|
||
[08-31 15:03:59] Setting distributed timeout to 3600 seconds
|
||
[08-31 15:04:00] Found nccl from library libnccl.so.2
|
||
[08-31 15:04:00] sglang-diffusion is using nccl==2.28.9
|
||
[08-31 15:04:03] reading GPU P2P access cache from /root/.cache/sglang/gpu_p2p_access_cache_for_6,7.json
|
||
[08-31 15:04:03] reading GPU P2P access cache from /root/.cache/sglang/gpu_p2p_access_cache_for_6,7.json
|
||
[08-31 15:04:03] Found nccl from library libnccl.so.2
|
||
[08-31 15:04:03] sglang-diffusion is using nccl==2.28.9
|
||
[08-31 15:04:03] No pipeline_class_name specified, using model_index.json
|
||
|
||
Loading required modules: 0%| | 0/6 [00:00<?, ?it/s][08-31 15:04:04] Using pipeline from model_index.json: MiniMaxH3Pipeline
|
||
[08-31 15:04:04] Loading pipeline modules...
|
||
[08-31 15:04:04] Model path: /data/hf_models/MiniMax-H3/Ref2VA
|
||
[08-31 15:04:04] Diffusers version: 0.32.2
|
||
[08-31 15:04:04] Loading pipeline modules from config: {'_class_name': 'MiniMaxH3Pipeline', '_diffusers_version': '0.32.2', 'text_encoder': ['transformers', 'MiniMaxH3Qwen3VLHFEncoder'], 'tokenizer': ['transformers', 'Qwen2TokenizerFast'], 'video_vae': ['diffusers', 'MiniMaxH3VideoVAE'], 'audio_vae': ['diffusers', 'MiniMaxH3AudioVAE'], 'scheduler': None, 'transformer': ['diffusers', 'MiniMaxH3DiTModel'], 'processor': ['transformers', 'Qwen3VLProcessor'], '_minimax_h3': {'schema_version': 1, 'partition': 'ref2va', 'tasks': ['ref2va'], 'task_aliases': {}, 'sigma_shift_scales': {'video': 12.0, 'audio': 3.0}}}
|
||
[08-31 15:04:04] Loading required components: ['processor', 'text_encoder', 'tokenizer', 'video_vae', 'audio_vae', 'transformer']
|
||
[08-31 15:04:04] Memory-aware component load order: ['text_encoder', 'transformer', 'audio_vae', 'video_vae', 'processor', 'tokenizer']
|
||
|
||
Loading required modules: 0%| | 0/6 [00:00<?, ?it/s][08-31 15:04:04] Loading text_encoder from /data/hf_models/MiniMax-H3/Ref2VA/text_encoder. avail mem: 78.97 GB
|
||
[08-31 15:04:04] Defaulting to Torch SDPA backend on SM12.x
|
||
|
||
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/12 [00:00<?, ?it/s]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 8% Completed | 1/12 [00:02<00:23, 2.16s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 17% Completed | 2/12 [00:03<00:15, 1.58s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 25% Completed | 3/12 [00:04<00:12, 1.36s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 33% Completed | 4/12 [00:05<00:10, 1.27s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 42% Completed | 5/12 [00:06<00:08, 1.19s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 50% Completed | 6/12 [00:07<00:06, 1.11s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 58% Completed | 7/12 [00:08<00:05, 1.06s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 67% Completed | 8/12 [00:09<00:04, 1.04s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 75% Completed | 9/12 [00:10<00:03, 1.01s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 83% Completed | 10/12 [00:11<00:01, 1.00it/s]
|
||
[A
|
||
Loading required modules: 17%|█▋ | 1/6 [00:12<01:02, 12.53s/it]
|
||
|
||
Loading safetensors checkpoint shards: 92% Completed | 11/12 [00:11<00:00, 1.28it/s]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 100% Completed | 12/12 [00:12<00:00, 1.58it/s]
|
||
[A
|
||
Loading safetensors checkpoint shards: 100% Completed | 12/12 [00:12<00:00, 1.00s/it]
|
||
|
||
[08-31 15:04:17] Loaded text_encoder: MiniMaxH3Qwen3VLEncoder (sgl-diffusion version). model size: 25.26 GB, consumed GPU mem: 25.45 GB, avail GPU mem: 53.52 GB
|
||
[08-31 15:04:17] Attention backends for text_encoder: torch_sdpa
|
||
|
||
Loading required modules: 17%|█▋ | 1/6 [00:13<01:05, 13.05s/it][08-31 15:04:17] Loading transformer from /data/hf_models/MiniMax-H3/Ref2VA/transformer. avail mem: 53.52 GB
|
||
[08-31 15:04:17] Loading MiniMaxH3DiTModel from 13 safetensors file(s) , param_dtype: torch.bfloat16
|
||
|
||
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/13 [00:00<?, ?it/s]
|
||
[A
|
||
Loading safetensors checkpoint shards: 100% Completed | 13/13 [00:00<00:00, 1075.89it/s]
|
||
|
||
[08-31 15:04:35] Loaded transformer: MiniMaxH3DiTModel (sgl-diffusion version). model size: 30.86 GB, consumed GPU mem: 33.87 GB, avail GPU mem: 19.65 GB
|
||
|
||
Loading required modules: 33%|███▎ | 2/6 [00:31<01:04, 16.24s/it][08-31 15:04:35] Loading audio_vae from /data/hf_models/MiniMax-H3/Ref2VA/audio_vae. avail mem: 19.65 GB
|
||
|
||
Loading required modules: 33%|███▎ | 2/6 [00:31<01:05, 16.36s/it][08-31 15:04:36] Loaded audio_vae: MiniMaxH3AudioVAE (sgl-diffusion version). model size: 0.56 GB, consumed GPU mem: 0.10 GB, avail GPU mem: 19.55 GB
|
||
|
||
Loading required modules: 50%|█████ | 3/6 [00:31<00:27, 9.03s/it][08-31 15:04:36] Loading video_vae from /data/hf_models/MiniMax-H3/Ref2VA/video_vae. avail mem: 19.55 GB
|
||
|
||
Loading required modules: 50%|█████ | 3/6 [00:32<00:27, 9.11s/it][08-31 15:04:41] Loaded video_vae: MiniMaxH3VideoVAE (sgl-diffusion version). model size: 9.7 GB, consumed GPU mem: 7.60 GB, avail GPU mem: 11.95 GB
|
||
|
||
Loading required modules: 67%|██████▋ | 4/6 [00:36<00:14, 7.39s/it][08-31 15:04:41] Loading processor from /data/hf_models/MiniMax-H3/Ref2VA/processor. avail mem: 11.95 GB
|
||
[08-31 15:04:41] Loaded processor: Qwen3VLProcessor (sgl-diffusion version). model size: NA GB, consumed GPU mem: 0.00 GB, avail GPU mem: 11.95 GB
|
||
|
||
Loading required modules: 83%|████████▎ | 5/6 [00:37<00:04, 4.85s/it][08-31 15:04:41] Loading tokenizer from /data/hf_models/MiniMax-H3/Ref2VA/tokenizer. avail mem: 11.95 GB
|
||
[08-31 15:04:41] Loaded tokenizer: Qwen2Tokenizer (sgl-diffusion version). model size: NA GB, consumed GPU mem: 0.00 GB, avail GPU mem: 11.95 GB
|
||
|
||
Loading required modules: 100%|██████████| 6/6 [00:37<00:00, 3.30s/it]
|
||
Loading required modules: 100%|██████████| 6/6 [00:37<00:00, 6.25s/it]
|
||
[08-31 15:04:41] Creating pipeline stages...
|
||
[08-31 15:04:41] Defaulting to Torch SDPA backend on SM12.x
|
||
[08-31 15:04:41] Using torch_sdpa attention backend
|
||
[08-31 15:04:41] Pipeline instantiated
|
||
[08-31 15:04:41] Worker 0: Initialized device, model, and distributed environment.
|
||
[08-31 15:04:41] Worker 0: Scheduler loop started.
|
||
|
||
Loading required modules: 67%|██████▋ | 4/6 [00:37<00:15, 7.73s/it]
|
||
Loading required modules: 83%|████████▎ | 5/6 [00:38<00:05, 5.09s/it]
|
||
Loading required modules: 100%|██████████| 6/6 [00:38<00:00, 3.48s/it]
|
||
Loading required modules: 100%|██████████| 6/6 [00:38<00:00, 6.40s/it]
|
||
[08-31 15:04:42] Starting FastAPI server.
|
||
[2026-08-31 15:04:42] [32mINFO[0m: Started server process [[36m2252070[0m]
|
||
[2026-08-31 15:04:42] [32mINFO[0m: Waiting for application startup.
|
||
[08-31 15:04:42] ZMQ Broker is listening for offline jobs on tcp://127.0.0.1:34041
|
||
[2026-08-31 15:04:42] [32mINFO[0m: Application startup complete.
|
||
[2026-08-31 15:04:42] [32mINFO[0m: Uvicorn running on [1mhttp://0.0.0.0:34040[0m (Press CTRL+C to quit)
|
||
[2026-08-31 15:04:43] [32mINFO[0m: 127.0.0.1:51114 - "[1mGET /health HTTP/1.1[0m" [32m200 OK[0m
|
||
|
||
Warmup requests: 0%| | 0/1 [00:00<?, ?req/s][2026-08-31 15:04:47] [32mINFO[0m: 127.0.0.1:51124 - "[1mGET /health HTTP/1.1[0m" [32m200 OK[0m
|
||
[08-31 15:04:47] Defaulting to Torch SDPA backend on SM12.x
|
||
|
||
Warmup requests: 100%|██████████| 1/1 [00:31<00:00, 31.04s/req, server warmup req (1344x768x124f, 2/50 steps), last=31.04s]
|
||
Warmup requests: 100%|██████████| 1/1 [00:31<00:00, 31.04s/req, server warmup req (1344x768x124f, 2/50 steps), last=31.04s]
|
||
[08-31 15:05:15] The server is fired up and ready to roll!
|
||
[08-31 15:05:15] Adjusting number of frames from 1 to 1 based on model
|
||
[08-31 15:05:15] Adjusting number of frames from 1 to 5 based on number of GPUs (2)
|
||
[2026-08-31 15:05:15] [32mINFO[0m: 127.0.0.1:51132 - "[1mPOST /v1/videos HTTP/1.1[0m" [31m400 Bad Request[0m
|
||
[2026-08-31 15:05:16] [32mINFO[0m: Shutting down
|
||
[2026-08-31 15:05:16] [32mINFO[0m: Waiting for application shutdown.
|
||
[08-31 15:05:16] FastAPI app is shutting down...
|
||
[2026-08-31 15:05:16] [32mINFO[0m: Application shutdown complete.
|
||
[2026-08-31 15:05:16] [32mINFO[0m: Finished server process [[36m2252070[0m]
|
||
[08-31 15:05:20] Worker 0: Shutdown complete.
|
||
[08-31 15:05:23] kill_process_tree called: parent_pid=2252070, include_parent=False, pid=2252070
|