104 lines
13 KiB
Plaintext
104 lines
13 KiB
Plaintext
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
|
||
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
|
||
[08-31 15:03:40] Applying performance_mode=speed
|
||
[08-31 15:03:40] server_args: {"model_path": "/data/hf_models/MiniMax-H3", "model_subfolder": null, "model_variant": "Ref2VA", "model_id": null, "backend": "sglang", "attention_backend": null, "attention_backend_config": {}, "component_attention_backends": {}, "cache_dit_config": null, "nccl_port": null, "trust_remote_code": false, "revision": null, "num_gpus": 2, "performance_mode": "speed", "base_gpu_id": 0, "gpu_ids": null, "tp_size": 2, "sp_degree": 1, "ulysses_degree": 1, "ring_degree": 1, "dp_size": 1, "dp_degree": 1, "enable_cfg_parallel": false, "cfg_parallel_degree": 1, "encoder_parallel": "auto", "hsdp_replicate_dim": 1, "hsdp_shard_dim": 2, "dist_timeout": 3600, "pipeline_class_name": null, "lora_path": null, "lora_nickname": "default", "lora_scale": 1.0, "lora_merge_mode": "auto", "lora_weight_name": null, "component_paths": {}, "transformer_weights_path": null, "component_transformer_weights_paths": {}, "quantization": null, "quantization_ignored_layers": null, "lora_target_modules": null, "dit_cpu_offload": false, "dit_layerwise_offload": false, "layerwise_offload_components": null, "dit_offload_prefetch_size": 0.0, "dit_layerwise_resident_layers": 0.0, "offload_during_compile": true, "text_encoder_cpu_offload": false, "image_encoder_cpu_offload": false, "vae_cpu_offload": false, "use_fsdp_inference": false, "pin_cpu_memory": true, "ltx2_two_stage_device_mode": null, "comfyui_mode": false, "enable_torch_compile": false, "regional_compile": false, "enable_breakable_cuda_graph": false, "bcg_text_buckets": null, "enable_layerwise_nvtx_marker": false, "warmup_mode": "server", "warmup": true, "server_warmup": true, "warmup_resolutions": null, "warmup_steps": 1, "disable_autocast": false, "master_port": 35010, "host": "0.0.0.0", "port": 34020, "webui": false, "webui_port": 12312, "scheduler_port": 36010, "batching_mode": "dynamic", "batching_max_size": 1, "batching_delay_ms": 0.0, "batching_config": null, "enable_batching_metrics": false, "strict_ports": false, "output_path": "/data/wxy/sskj-h3/throughput/sglang-base/results/ref2va-feishu-base-tp2x4-768p-15s-20steps-20260831-150318/server_1_port34020/outputs", "input_save_path": "inputs/uploads", "prompt_file_path": null, "model_paths": {}, "model_loaded": {"transformer": true, "vae": true, "video_vae": true, "audio_vae": true, "video_dit": true, "audio_dit": true, "dual_tower_bridge": true}, "boundary_ratio": null, "disagg_role": "monolithic", "disagg_timeout": 3600, "disagg_downstream_wait_timeout": 1800, "disagg_dispatch_policy": "round_robin", "disagg_mode": false, "disagg_instance_id": 0, "disagg_max_slots_per_instance": 8, "disagg_transfer_redundancy": 1.25, "disagg_role_device": "auto", "disagg_transfer_backend": "auto", "disagg_transfer_pool_size": 268435456, "disagg_transfer_pin_memory": "auto", "disagg_p2p_hostname": "127.0.0.1", "disagg_ib_device": null, "disagg_server_addr": null, "encoder_urls": null, "denoiser_urls": null, "decoder_urls": null, "encoder_tp": null, "denoiser_tp": null, "denoiser_sp": null, "denoiser_ulysses": null, "denoiser_ring": null, "decoder_sp": null, "decoder_tp": null, "pool_work_endpoint": null, "pool_result_endpoint": null, "pool_control_endpoint": null, "pool_control_advertised_endpoint": null, "log_level": "info", "log_requests": false, "log_requests_level": 2, "log_requests_format": "text", "log_requests_target": null, "uvicorn_access_log_exclude_prefixes": [], "enable_trace": false, "otlp_traces_endpoint": "localhost:4317", "srt_encoder_url": null, "srt_encoder_connect_timeout": 3.05, "srt_encoder_timeout": 100, "pe_server_url": null}
|
||
[08-31 15:03:40] Starting server...
|
||
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
|
||
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
|
||
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
|
||
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
|
||
[08-31 15:03:59] Scheduler bind at endpoint: tcp://0.0.0.0:36010
|
||
[08-31 15:04:00] torch.compile cache: TORCHINDUCTOR_CACHE_DIR=/root/.cache/sgl_diffusion/torch_compile_cache/inductor TRITON_CACHE_DIR=/root/.cache/sgl_diffusion/torch_compile_cache/triton
|
||
[08-31 15:04:00] Initializing distributed environment with world_size=2, device=cuda:0, timeout=3600
|
||
[08-31 15:04:00] Setting distributed timeout to 3600 seconds
|
||
[08-31 15:04:01] Found nccl from library libnccl.so.2
|
||
[08-31 15:04:01] sglang-diffusion is using nccl==2.28.9
|
||
[08-31 15:04:04] reading GPU P2P access cache from /root/.cache/sglang/gpu_p2p_access_cache_for_2,3.json
|
||
[08-31 15:04:04] reading GPU P2P access cache from /root/.cache/sglang/gpu_p2p_access_cache_for_2,3.json
|
||
[08-31 15:04:04] Found nccl from library libnccl.so.2
|
||
[08-31 15:04:04] sglang-diffusion is using nccl==2.28.9
|
||
[08-31 15:04:04] No pipeline_class_name specified, using model_index.json
|
||
|
||
Loading required modules: 0%| | 0/6 [00:00<?, ?it/s][08-31 15:04:04] Using pipeline from model_index.json: MiniMaxH3Pipeline
|
||
[08-31 15:04:04] Loading pipeline modules...
|
||
[08-31 15:04:04] Model path: /data/hf_models/MiniMax-H3/Ref2VA
|
||
[08-31 15:04:04] Diffusers version: 0.32.2
|
||
[08-31 15:04:04] Loading pipeline modules from config: {'_class_name': 'MiniMaxH3Pipeline', '_diffusers_version': '0.32.2', 'text_encoder': ['transformers', 'MiniMaxH3Qwen3VLHFEncoder'], 'tokenizer': ['transformers', 'Qwen2TokenizerFast'], 'video_vae': ['diffusers', 'MiniMaxH3VideoVAE'], 'audio_vae': ['diffusers', 'MiniMaxH3AudioVAE'], 'scheduler': None, 'transformer': ['diffusers', 'MiniMaxH3DiTModel'], 'processor': ['transformers', 'Qwen3VLProcessor'], '_minimax_h3': {'schema_version': 1, 'partition': 'ref2va', 'tasks': ['ref2va'], 'task_aliases': {}, 'sigma_shift_scales': {'video': 12.0, 'audio': 3.0}}}
|
||
[08-31 15:04:04] Loading required components: ['processor', 'text_encoder', 'tokenizer', 'video_vae', 'audio_vae', 'transformer']
|
||
[08-31 15:04:04] Memory-aware component load order: ['text_encoder', 'transformer', 'audio_vae', 'video_vae', 'processor', 'tokenizer']
|
||
|
||
Loading required modules: 0%| | 0/6 [00:00<?, ?it/s][08-31 15:04:04] Loading text_encoder from /data/hf_models/MiniMax-H3/Ref2VA/text_encoder. avail mem: 78.97 GB
|
||
[08-31 15:04:04] Defaulting to Torch SDPA backend on SM12.x
|
||
|
||
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/12 [00:00<?, ?it/s]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 8% Completed | 1/12 [00:02<00:24, 2.23s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 17% Completed | 2/12 [00:03<00:15, 1.58s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 25% Completed | 3/12 [00:04<00:12, 1.37s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 33% Completed | 4/12 [00:05<00:10, 1.31s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 42% Completed | 5/12 [00:06<00:08, 1.26s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 50% Completed | 6/12 [00:07<00:06, 1.14s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 58% Completed | 7/12 [00:08<00:05, 1.06s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 67% Completed | 8/12 [00:09<00:04, 1.01s/it]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 75% Completed | 9/12 [00:10<00:02, 1.03it/s]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 83% Completed | 10/12 [00:11<00:01, 1.05it/s]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 92% Completed | 11/12 [00:11<00:00, 1.35it/s]
|
||
[A
|
||
|
||
Loading safetensors checkpoint shards: 100% Completed | 12/12 [00:11<00:00, 1.66it/s]
|
||
[A
|
||
Loading safetensors checkpoint shards: 100% Completed | 12/12 [00:11<00:00, 1.01it/s]
|
||
|
||
[08-31 15:04:16] Loaded text_encoder: MiniMaxH3Qwen3VLEncoder (sgl-diffusion version). model size: 25.26 GB, consumed GPU mem: 25.45 GB, avail GPU mem: 53.52 GB
|
||
[08-31 15:04:16] Attention backends for text_encoder: torch_sdpa
|
||
|
||
Loading required modules: 17%|█▋ | 1/6 [00:12<01:01, 12.22s/it][08-31 15:04:16] Loading transformer from /data/hf_models/MiniMax-H3/Ref2VA/transformer. avail mem: 53.52 GB
|
||
[08-31 15:04:16] Loading MiniMaxH3DiTModel from 13 safetensors file(s) , param_dtype: torch.bfloat16
|
||
|
||
|
||
Loading safetensors checkpoint shards: 0% Completed | 0/13 [00:00<?, ?it/s]
|
||
[A
|
||
Loading safetensors checkpoint shards: 100% Completed | 13/13 [00:00<00:00, 1035.47it/s]
|
||
|
||
|
||
Loading required modules: 17%|█▋ | 1/6 [00:12<01:02, 12.54s/it]
|
||
Loading required modules: 33%|███▎ | 2/6 [00:31<01:04, 16.18s/it][08-31 15:04:35] Loaded transformer: MiniMaxH3DiTModel (sgl-diffusion version). model size: 30.86 GB, consumed GPU mem: 33.87 GB, avail GPU mem: 19.65 GB
|
||
|
||
Loading required modules: 33%|███▎ | 2/6 [00:31<01:04, 16.20s/it][08-31 15:04:35] Loading audio_vae from /data/hf_models/MiniMax-H3/Ref2VA/audio_vae. avail mem: 19.65 GB
|
||
|
||
Loading required modules: 50%|█████ | 3/6 [00:31<00:27, 9.02s/it][08-31 15:04:36] Loaded audio_vae: MiniMaxH3AudioVAE (sgl-diffusion version). model size: 0.56 GB, consumed GPU mem: 0.10 GB, avail GPU mem: 19.55 GB
|
||
|
||
Loading required modules: 50%|█████ | 3/6 [00:31<00:27, 9.04s/it][08-31 15:04:36] Loading video_vae from /data/hf_models/MiniMax-H3/Ref2VA/video_vae. avail mem: 19.55 GB
|
||
|
||
Loading required modules: 67%|██████▋ | 4/6 [00:37<00:15, 7.53s/it][08-31 15:04:41] Loaded video_vae: MiniMaxH3VideoVAE (sgl-diffusion version). model size: 9.7 GB, consumed GPU mem: 7.60 GB, avail GPU mem: 11.95 GB
|
||
|
||
Loading required modules: 67%|██████▋ | 4/6 [00:37<00:15, 7.62s/it][08-31 15:04:41] Loading processor from /data/hf_models/MiniMax-H3/Ref2VA/processor. avail mem: 11.95 GB
|
||
|
||
Loading required modules: 83%|████████▎ | 5/6 [00:37<00:04, 4.97s/it][08-31 15:04:42] Loaded processor: Qwen3VLProcessor (sgl-diffusion version). model size: NA GB, consumed GPU mem: 0.00 GB, avail GPU mem: 11.95 GB
|
||
|
||
Loading required modules: 83%|████████▎ | 5/6 [00:37<00:05, 5.02s/it][08-31 15:04:42] Loading tokenizer from /data/hf_models/MiniMax-H3/Ref2VA/tokenizer. avail mem: 11.95 GB
|
||
|
||
Loading required modules: 100%|██████████| 6/6 [00:37<00:00, 3.40s/it]
|
||
Loading required modules: 100%|██████████| 6/6 [00:37<00:00, 6.30s/it]
|
||
[08-31 15:04:42] Loaded tokenizer: Qwen2Tokenizer (sgl-diffusion version). model size: NA GB, consumed GPU mem: 0.00 GB, avail GPU mem: 11.95 GB
|
||
|
||
Loading required modules: 100%|██████████| 6/6 [00:37<00:00, 3.43s/it]
|
||
Loading required modules: 100%|██████████| 6/6 [00:37<00:00, 6.32s/it]
|
||
[08-31 15:04:42] Creating pipeline stages...
|
||
[08-31 15:04:42] Defaulting to Torch SDPA backend on SM12.x
|
||
[08-31 15:04:42] Using torch_sdpa attention backend
|
||
[08-31 15:04:42] Pipeline instantiated
|
||
[08-31 15:04:42] Worker 0: Initialized device, model, and distributed environment.
|
||
[08-31 15:04:42] Worker 0: Scheduler loop started.
|
||
[08-31 15:04:42] Starting FastAPI server.
|
||
[2026-08-31 15:04:42] [32mINFO[0m: Started server process [[36m2251930[0m]
|
||
[2026-08-31 15:04:42] [32mINFO[0m: Waiting for application startup.
|
||
[08-31 15:04:42] ZMQ Broker is listening for offline jobs on tcp://127.0.0.1:34021
|
||
[2026-08-31 15:04:42] [32mINFO[0m: Application startup complete.
|
||
[2026-08-31 15:04:42] [32mINFO[0m: Uvicorn running on [1mhttp://0.0.0.0:34020[0m (Press CTRL+C to quit)
|
||
[2026-08-31 15:04:43] [32mINFO[0m: 127.0.0.1:42776 - "[1mGET /health HTTP/1.1[0m" [32m200 OK[0m
|
||
|
||
Warmup requests: 0%| | 0/1 [00:00<?, ?req/s][2026-08-31 15:04:47] [32mINFO[0m: 127.0.0.1:42792 - "[1mGET /health HTTP/1.1[0m" [32m200 OK[0m
|
||
[08-31 15:04:47] Defaulting to Torch SDPA backend on SM12.x
|
||
|
||
Warmup requests: 100%|██████████| 1/1 [00:31<00:00, 31.20s/req, server warmup req (1344x768x124f, 2/50 steps), last=31.20s]
|
||
Warmup requests: 100%|██████████| 1/1 [00:31<00:00, 31.20s/req, server warmup req (1344x768x124f, 2/50 steps), last=31.20s]
|
||
[08-31 15:05:15] The server is fired up and ready to roll!
|
||
[08-31 15:05:15] Adjusting number of frames from 1 to 1 based on model
|
||
[08-31 15:05:15] Adjusting number of frames from 1 to 5 based on number of GPUs (2)
|
||
[2026-08-31 15:05:15] [32mINFO[0m: 127.0.0.1:42798 - "[1mPOST /v1/videos HTTP/1.1[0m" [31m400 Bad Request[0m
|
||
[2026-08-31 15:05:16] [32mINFO[0m: Shutting down
|
||
[2026-08-31 15:05:16] [32mINFO[0m: Waiting for application shutdown.
|
||
[08-31 15:05:16] FastAPI app is shutting down...
|
||
[2026-08-31 15:05:16] [32mINFO[0m: Application shutdown complete.
|
||
[2026-08-31 15:05:16] [32mINFO[0m: Finished server process [[36m2251930[0m]
|
||
[08-31 15:05:21] Worker 0: Shutdown complete.
|
||
[08-31 15:05:24] kill_process_tree called: parent_pid=2251930, include_parent=False, pid=2251930
|