# General-VQA ## Overview General-VQA is a customizable visual question answering benchmark for evaluating multimodal models. It supports OpenAI-compatible message format with flexible image/video/audio input (local paths, URLs, or base64). ## Task Description - **Task Type**: Visual Question Answering - **Input**: Images/videos/audio + questions in OpenAI chat format - **Output**: Free-form text answer - **Flexibility**: Supports custom datasets via TSV/JSONL files ## Key Features - OpenAI-compatible message format - Supports multiple image/video/audio input methods (path, URL, base64) - **Media placeholders**: Use ```` / ``