# Multimodal Large Models This framework supports two custom multimodal evaluation methods: - **General-VQA Format**: Suitable for Q&A-based multimodal evaluation tasks. Supports two input styles: **OpenAI Messages Data**, **MMMU-style Data with Media Placeholders**. - **General-VMCQ Format**: Suitable for multiple-choice multimodal evaluation tasks. Uses the [media placeholders][mp-feature] to embed images, videos, and audio in questions and options, similar to MMMU format. ## General-VQA Format General-VQA supports **two input styles**: 1. **OpenAI Messages Data** — full structured content with explicit media parts (images, audio, video) in the OpenAI message schema. Supports multi-turn conversations, system prompts, and fine-grained control over each content part. 2. **MMMU-style Data with Media Placeholders** — a simpler approach where the user message is a plain-text string containing ``, `