2026-08-31 15:57:13 +08:00

104 lines
13 KiB
Plaintext
Raw Blame History

This file contains invisible Unicode characters

This file contains invisible Unicode characters that are indistinguishable to humans but may be processed differently by a computer. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

Failed to get device capability: SM 12.x requires CUDA >= 12.9.
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
[08-31 15:03:40] Applying performance_mode=speed
[08-31 15:03:40] server_args: {"model_path": "/data/hf_models/MiniMax-H3", "model_subfolder": null, "model_variant": "Ref2VA", "model_id": null, "backend": "sglang", "attention_backend": null, "attention_backend_config": {}, "component_attention_backends": {}, "cache_dit_config": null, "nccl_port": null, "trust_remote_code": false, "revision": null, "num_gpus": 2, "performance_mode": "speed", "base_gpu_id": 0, "gpu_ids": null, "tp_size": 2, "sp_degree": 1, "ulysses_degree": 1, "ring_degree": 1, "dp_size": 1, "dp_degree": 1, "enable_cfg_parallel": false, "cfg_parallel_degree": 1, "encoder_parallel": "auto", "hsdp_replicate_dim": 1, "hsdp_shard_dim": 2, "dist_timeout": 3600, "pipeline_class_name": null, "lora_path": null, "lora_nickname": "default", "lora_scale": 1.0, "lora_merge_mode": "auto", "lora_weight_name": null, "component_paths": {}, "transformer_weights_path": null, "component_transformer_weights_paths": {}, "quantization": null, "quantization_ignored_layers": null, "lora_target_modules": null, "dit_cpu_offload": false, "dit_layerwise_offload": false, "layerwise_offload_components": null, "dit_offload_prefetch_size": 0.0, "dit_layerwise_resident_layers": 0.0, "offload_during_compile": true, "text_encoder_cpu_offload": false, "image_encoder_cpu_offload": false, "vae_cpu_offload": false, "use_fsdp_inference": false, "pin_cpu_memory": true, "ltx2_two_stage_device_mode": null, "comfyui_mode": false, "enable_torch_compile": false, "regional_compile": false, "enable_breakable_cuda_graph": false, "bcg_text_buckets": null, "enable_layerwise_nvtx_marker": false, "warmup_mode": "server", "warmup": true, "server_warmup": true, "warmup_resolutions": null, "warmup_steps": 1, "disable_autocast": false, "master_port": 35010, "host": "0.0.0.0", "port": 34020, "webui": false, "webui_port": 12312, "scheduler_port": 36010, "batching_mode": "dynamic", "batching_max_size": 1, "batching_delay_ms": 0.0, "batching_config": null, "enable_batching_metrics": false, "strict_ports": false, "output_path": "/data/wxy/sskj-h3/throughput/sglang-base/results/ref2va-feishu-base-tp2x4-768p-15s-20steps-20260831-150318/server_1_port34020/outputs", "input_save_path": "inputs/uploads", "prompt_file_path": null, "model_paths": {}, "model_loaded": {"transformer": true, "vae": true, "video_vae": true, "audio_vae": true, "video_dit": true, "audio_dit": true, "dual_tower_bridge": true}, "boundary_ratio": null, "disagg_role": "monolithic", "disagg_timeout": 3600, "disagg_downstream_wait_timeout": 1800, "disagg_dispatch_policy": "round_robin", "disagg_mode": false, "disagg_instance_id": 0, "disagg_max_slots_per_instance": 8, "disagg_transfer_redundancy": 1.25, "disagg_role_device": "auto", "disagg_transfer_backend": "auto", "disagg_transfer_pool_size": 268435456, "disagg_transfer_pin_memory": "auto", "disagg_p2p_hostname": "127.0.0.1", "disagg_ib_device": null, "disagg_server_addr": null, "encoder_urls": null, "denoiser_urls": null, "decoder_urls": null, "encoder_tp": null, "denoiser_tp": null, "denoiser_sp": null, "denoiser_ulysses": null, "denoiser_ring": null, "decoder_sp": null, "decoder_tp": null, "pool_work_endpoint": null, "pool_result_endpoint": null, "pool_control_endpoint": null, "pool_control_advertised_endpoint": null, "log_level": "info", "log_requests": false, "log_requests_level": 2, "log_requests_format": "text", "log_requests_target": null, "uvicorn_access_log_exclude_prefixes": [], "enable_trace": false, "otlp_traces_endpoint": "localhost:4317", "srt_encoder_url": null, "srt_encoder_connect_timeout": 3.05, "srt_encoder_timeout": 100, "pe_server_url": null}
[08-31 15:03:40] Starting server...
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
Failed to get device capability: SM 12.x requires CUDA >= 12.9.
[08-31 15:03:59] Scheduler bind at endpoint: tcp://0.0.0.0:36010
[08-31 15:04:00] torch.compile cache: TORCHINDUCTOR_CACHE_DIR=/root/.cache/sgl_diffusion/torch_compile_cache/inductor TRITON_CACHE_DIR=/root/.cache/sgl_diffusion/torch_compile_cache/triton
[08-31 15:04:00] Initializing distributed environment with world_size=2, device=cuda:0, timeout=3600
[08-31 15:04:00] Setting distributed timeout to 3600 seconds
[08-31 15:04:01] Found nccl from library libnccl.so.2
[08-31 15:04:01] sglang-diffusion is using nccl==2.28.9
[08-31 15:04:04] reading GPU P2P access cache from /root/.cache/sglang/gpu_p2p_access_cache_for_2,3.json
[08-31 15:04:04] reading GPU P2P access cache from /root/.cache/sglang/gpu_p2p_access_cache_for_2,3.json
[08-31 15:04:04] Found nccl from library libnccl.so.2
[08-31 15:04:04] sglang-diffusion is using nccl==2.28.9
[08-31 15:04:04] No pipeline_class_name specified, using model_index.json
Loading required modules: 0%| | 0/6 [00:00<?, ?it/s][08-31 15:04:04] Using pipeline from model_index.json: MiniMaxH3Pipeline
[08-31 15:04:04] Loading pipeline modules...
[08-31 15:04:04] Model path: /data/hf_models/MiniMax-H3/Ref2VA
[08-31 15:04:04] Diffusers version: 0.32.2
[08-31 15:04:04] Loading pipeline modules from config: {'_class_name': 'MiniMaxH3Pipeline', '_diffusers_version': '0.32.2', 'text_encoder': ['transformers', 'MiniMaxH3Qwen3VLHFEncoder'], 'tokenizer': ['transformers', 'Qwen2TokenizerFast'], 'video_vae': ['diffusers', 'MiniMaxH3VideoVAE'], 'audio_vae': ['diffusers', 'MiniMaxH3AudioVAE'], 'scheduler': None, 'transformer': ['diffusers', 'MiniMaxH3DiTModel'], 'processor': ['transformers', 'Qwen3VLProcessor'], '_minimax_h3': {'schema_version': 1, 'partition': 'ref2va', 'tasks': ['ref2va'], 'task_aliases': {}, 'sigma_shift_scales': {'video': 12.0, 'audio': 3.0}}}
[08-31 15:04:04] Loading required components: ['processor', 'text_encoder', 'tokenizer', 'video_vae', 'audio_vae', 'transformer']
[08-31 15:04:04] Memory-aware component load order: ['text_encoder', 'transformer', 'audio_vae', 'video_vae', 'processor', 'tokenizer']
Loading required modules: 0%| | 0/6 [00:00<?, ?it/s][08-31 15:04:04] Loading text_encoder from /data/hf_models/MiniMax-H3/Ref2VA/text_encoder. avail mem: 78.97 GB
[08-31 15:04:04] Defaulting to Torch SDPA backend on SM12.x
Loading safetensors checkpoint shards: 0% Completed | 0/12 [00:00<?, ?it/s]

Loading safetensors checkpoint shards: 8% Completed | 1/12 [00:02<00:24, 2.23s/it]

Loading safetensors checkpoint shards: 17% Completed | 2/12 [00:03<00:15, 1.58s/it]

Loading safetensors checkpoint shards: 25% Completed | 3/12 [00:04<00:12, 1.37s/it]

Loading safetensors checkpoint shards: 33% Completed | 4/12 [00:05<00:10, 1.31s/it]

Loading safetensors checkpoint shards: 42% Completed | 5/12 [00:06<00:08, 1.26s/it]

Loading safetensors checkpoint shards: 50% Completed | 6/12 [00:07<00:06, 1.14s/it]

Loading safetensors checkpoint shards: 58% Completed | 7/12 [00:08<00:05, 1.06s/it]

Loading safetensors checkpoint shards: 67% Completed | 8/12 [00:09<00:04, 1.01s/it]

Loading safetensors checkpoint shards: 75% Completed | 9/12 [00:10<00:02, 1.03it/s]

Loading safetensors checkpoint shards: 83% Completed | 10/12 [00:11<00:01, 1.05it/s]

Loading safetensors checkpoint shards: 92% Completed | 11/12 [00:11<00:00, 1.35it/s]

Loading safetensors checkpoint shards: 100% Completed | 12/12 [00:11<00:00, 1.66it/s]

Loading safetensors checkpoint shards: 100% Completed | 12/12 [00:11<00:00, 1.01it/s]
[08-31 15:04:16] Loaded text_encoder: MiniMaxH3Qwen3VLEncoder (sgl-diffusion version). model size: 25.26 GB, consumed GPU mem: 25.45 GB, avail GPU mem: 53.52 GB
[08-31 15:04:16] Attention backends for text_encoder: torch_sdpa
Loading required modules: 17%|█▋ | 1/6 [00:12<01:01, 12.22s/it][08-31 15:04:16] Loading transformer from /data/hf_models/MiniMax-H3/Ref2VA/transformer. avail mem: 53.52 GB
[08-31 15:04:16] Loading MiniMaxH3DiTModel from 13 safetensors file(s) , param_dtype: torch.bfloat16
Loading safetensors checkpoint shards: 0% Completed | 0/13 [00:00<?, ?it/s]

Loading safetensors checkpoint shards: 100% Completed | 13/13 [00:00<00:00, 1035.47it/s]
Loading required modules: 17%|█▋ | 1/6 [00:12<01:02, 12.54s/it]
Loading required modules: 33%|███▎ | 2/6 [00:31<01:04, 16.18s/it][08-31 15:04:35] Loaded transformer: MiniMaxH3DiTModel (sgl-diffusion version). model size: 30.86 GB, consumed GPU mem: 33.87 GB, avail GPU mem: 19.65 GB
Loading required modules: 33%|███▎ | 2/6 [00:31<01:04, 16.20s/it][08-31 15:04:35] Loading audio_vae from /data/hf_models/MiniMax-H3/Ref2VA/audio_vae. avail mem: 19.65 GB
Loading required modules: 50%|█████ | 3/6 [00:31<00:27, 9.02s/it][08-31 15:04:36] Loaded audio_vae: MiniMaxH3AudioVAE (sgl-diffusion version). model size: 0.56 GB, consumed GPU mem: 0.10 GB, avail GPU mem: 19.55 GB
Loading required modules: 50%|█████ | 3/6 [00:31<00:27, 9.04s/it][08-31 15:04:36] Loading video_vae from /data/hf_models/MiniMax-H3/Ref2VA/video_vae. avail mem: 19.55 GB
Loading required modules: 67%|██████▋ | 4/6 [00:37<00:15, 7.53s/it][08-31 15:04:41] Loaded video_vae: MiniMaxH3VideoVAE (sgl-diffusion version). model size: 9.7 GB, consumed GPU mem: 7.60 GB, avail GPU mem: 11.95 GB
Loading required modules: 67%|██████▋ | 4/6 [00:37<00:15, 7.62s/it][08-31 15:04:41] Loading processor from /data/hf_models/MiniMax-H3/Ref2VA/processor. avail mem: 11.95 GB
Loading required modules: 83%|████████▎ | 5/6 [00:37<00:04, 4.97s/it][08-31 15:04:42] Loaded processor: Qwen3VLProcessor (sgl-diffusion version). model size: NA GB, consumed GPU mem: 0.00 GB, avail GPU mem: 11.95 GB
Loading required modules: 83%|████████▎ | 5/6 [00:37<00:05, 5.02s/it][08-31 15:04:42] Loading tokenizer from /data/hf_models/MiniMax-H3/Ref2VA/tokenizer. avail mem: 11.95 GB
Loading required modules: 100%|██████████| 6/6 [00:37<00:00, 3.40s/it]
Loading required modules: 100%|██████████| 6/6 [00:37<00:00, 6.30s/it]
[08-31 15:04:42] Loaded tokenizer: Qwen2Tokenizer (sgl-diffusion version). model size: NA GB, consumed GPU mem: 0.00 GB, avail GPU mem: 11.95 GB
Loading required modules: 100%|██████████| 6/6 [00:37<00:00, 3.43s/it]
Loading required modules: 100%|██████████| 6/6 [00:37<00:00, 6.32s/it]
[08-31 15:04:42] Creating pipeline stages...
[08-31 15:04:42] Defaulting to Torch SDPA backend on SM12.x
[08-31 15:04:42] Using torch_sdpa attention backend
[08-31 15:04:42] Pipeline instantiated
[08-31 15:04:42] Worker 0: Initialized device, model, and distributed environment.
[08-31 15:04:42] Worker 0: Scheduler loop started.
[08-31 15:04:42] Starting FastAPI server.
[2026-08-31 15:04:42] INFO: Started server process [2251930]
[2026-08-31 15:04:42] INFO: Waiting for application startup.
[08-31 15:04:42] ZMQ Broker is listening for offline jobs on tcp://127.0.0.1:34021
[2026-08-31 15:04:42] INFO: Application startup complete.
[2026-08-31 15:04:42] INFO: Uvicorn running on http://0.0.0.0:34020 (Press CTRL+C to quit)
[2026-08-31 15:04:43] INFO: 127.0.0.1:42776 - "GET /health HTTP/1.1" 200 OK
Warmup requests: 0%| | 0/1 [00:00<?, ?req/s][2026-08-31 15:04:47] INFO: 127.0.0.1:42792 - "GET /health HTTP/1.1" 200 OK
[08-31 15:04:47] Defaulting to Torch SDPA backend on SM12.x
Warmup requests: 100%|██████████| 1/1 [00:31<00:00, 31.20s/req, server warmup req (1344x768x124f, 2/50 steps), last=31.20s]
Warmup requests: 100%|██████████| 1/1 [00:31<00:00, 31.20s/req, server warmup req (1344x768x124f, 2/50 steps), last=31.20s]
[08-31 15:05:15] The server is fired up and ready to roll!
[08-31 15:05:15] Adjusting number of frames from 1 to 1 based on model
[08-31 15:05:15] Adjusting number of frames from 1 to 5 based on number of GPUs (2)
[2026-08-31 15:05:15] INFO: 127.0.0.1:42798 - "POST /v1/videos HTTP/1.1" 400 Bad Request
[2026-08-31 15:05:16] INFO: Shutting down
[2026-08-31 15:05:16] INFO: Waiting for application shutdown.
[08-31 15:05:16] FastAPI app is shutting down...
[2026-08-31 15:05:16] INFO: Application shutdown complete.
[2026-08-31 15:05:16] INFO: Finished server process [2251930]
[08-31 15:05:21] Worker 0: Shutdown complete.
[08-31 15:05:24] kill_process_tree called: parent_pid=2251930, include_parent=False, pid=2251930