sora 13274243a0 Bump vendored EvalScope and add K3-ready DPV4 configs.
Keep K3 suite selection and report-schema scoring in bash, merge K3/vision dataset_args into dpv4 yamls, and pin EvalScope at 735d920ee911 with local patches.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-02 07:30:48 +00:00

120 lines
3.6 KiB
Markdown

# MVBench
## Overview
MVBench is a public multimodal video understanding benchmark covering temporal perception,
attribute/state reasoning, symbolic ordering, and high-level cognition. This native adapter uses
the ModelScope `PKU-Alignment/MVBench` mirror by default, which provides JSON annotations plus
optimized video archives.
## Task Description
- **Task Type**: Video multiple-choice question answering
- **Input**: Video + question + answer choices
- **Output**: Single correct answer letter
- **Subsets**: 20 MVBench tasks; the default smoke-test subset is `action_antonym`
## Evaluation Notes
- Default configuration uses **0-shot** evaluation
- Primary metric: **Accuracy**
- The default `action_antonym` subset downloads a small public MP4 archive for quick validation
- Full benchmark evaluation can be requested by setting `subset_list` to additional MVBench subsets
- Time-bounded records keep start/end metadata and add a short segment instruction to the prompt
## Properties
| Property | Value |
|----------|-------|
| **Benchmark Name** | `mvbench` |
| **Dataset ID** | [PKU-Alignment/MVBench](https://modelscope.cn/datasets/PKU-Alignment/MVBench/summary) |
| **Paper** | [Paper](https://arxiv.org/abs/2311.17005) |
| **Tags** | `MCQ`, `MultiModal`, `Video` |
| **Metrics** | `accuracy` |
| **Default Shots** | 0-shot |
| **Evaluation Split** | `train` |
## Data Statistics
| Metric | Value |
|--------|-------|
| Total Samples | 4,000 |
**Per-Subset Statistics:**
| Subset | Samples | Prompt Mean | Prompt Min | Prompt Max |
|--------|---------|-------------|------------|------------|
| `action_antonym` | 200 | N/A | N/A | N/A |
| `action_count` | 200 | N/A | N/A | N/A |
| `action_localization` | 200 | N/A | N/A | N/A |
| `action_prediction` | 200 | N/A | N/A | N/A |
| `action_sequence` | 200 | N/A | N/A | N/A |
| `character_order` | 200 | N/A | N/A | N/A |
| `counterfactual_inference` | 200 | N/A | N/A | N/A |
| `egocentric_navigation` | 200 | N/A | N/A | N/A |
| `episodic_reasoning` | 200 | N/A | N/A | N/A |
| `fine_grained_action` | 200 | N/A | N/A | N/A |
| `fine_grained_pose` | 200 | N/A | N/A | N/A |
| `moving_attribute` | 200 | N/A | N/A | N/A |
| `moving_count` | 200 | N/A | N/A | N/A |
| `moving_direction` | 200 | N/A | N/A | N/A |
| `object_existence` | 200 | N/A | N/A | N/A |
| `object_interaction` | 200 | N/A | N/A | N/A |
| `object_shuffle` | 200 | N/A | N/A | N/A |
| `scene_transition` | 200 | N/A | N/A | N/A |
| `state_change` | 200 | N/A | N/A | N/A |
| `unexpected_action` | 200 | N/A | N/A | N/A |
## Sample Example
*Sample example not available.*
## Prompt Template
**Prompt Template:**
```text
Answer the following multiple choice question. The last line of your response should be of the following format: 'ANSWER: [LETTER]' (without quotes) where [LETTER] is one of {letters}. Think step by step before answering.
{question}
{choices}
```
## Usage
### Using CLI
```bash
evalscope eval \
--model YOUR_MODEL \
--api-url OPENAI_API_COMPAT_URL \
--api-key EMPTY_TOKEN \
--datasets mvbench \
--limit 10 # Remove this line for formal evaluation
```
### Using Python
```python
from evalscope import run_task
from evalscope.config import TaskConfig
task_cfg = TaskConfig(
model='YOUR_MODEL',
api_url='OPENAI_API_COMPAT_URL',
api_key='EMPTY_TOKEN',
datasets=['mvbench'],
dataset_args={
'mvbench': {
# subset_list: ['action_antonym', 'action_count', 'action_localization'] # optional, evaluate specific subsets
}
},
limit=10, # Remove this line for formal evaluation
)
run_task(task_cfg=task_cfg)
```