2026-07-08 08:57:50 +00:00

149 lines
5.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# MathVerse
## 概述
MathVerse 是一个全面的视觉数学基准测试旨在对多模态大语言模型MLLMs进行公平且深入的评估。它包含 2,612 道高质量、多学科的带图数学题,并根据不同信息模态转换为 15K 个测试样本。
## 任务描述
- **任务类型**:视觉数学推理
- **输入**:带图的数学问题 + 问题(多选题或自由作答)
- **输出**:答案(多选题为选项字母,自由作答为数值或表达式)
- **领域**:带视觉图表的多学科数学
## 核心特性
- 2,612 道题目,每道题生成 6 个版本(共 15K 个样本)
- 测试 MLLM 是否真正理解用于数学推理的视觉图表
- 题目版本根据对视觉信息的依赖程度划分:
- **Text Dominant文本主导**:大部分信息在文本中
- **Text Lite文本轻量**:文本与视觉信息均衡
- **Vision Intensive视觉密集**:更依赖视觉信息
- **Vision Dominant视觉主导**:主要依赖视觉信息
- **Vision Only纯视觉**:所有信息均在图表中
- 支持多选题和自由作答两种形式
## 评估说明
- 默认评估使用 **testmini** 划分
- 主要指标:**准确率Accuracy**,采用数值比较
- 自由作答答案使用 \boxed{} 格式
- 使用 LLM 作为裁判进行答案验证
- 按题目版本分别报告结果,便于深入分析
## 属性
| 属性 | 值 |
|----------|-------|
| **基准测试名称** | `math_verse` |
| **数据集ID** | [evalscope/MathVerse](https://modelscope.cn/datasets/evalscope/MathVerse/summary) |
| **论文** | N/A |
| **标签** | `MCQ`, `Math`, `MultiModal`, `Reasoning` |
| **指标** | `acc` |
| **默认提示数量** | 0-shot |
| **评估划分** | `testmini` |
## 数据统计
| 指标 | 值 |
|--------|-------|
| 总样本数 | 3,940 |
| 提示词长度(平均) | 274.2 字符 |
| 提示词长度(最小/最大) | 70 / 1535 字符 |
**各子集统计:**
| 子集 | 样本数 | 提示平均长度 | 提示最小长度 | 提示最大长度 |
|--------|---------|-------------|------------|------------|
| `Text Dominant` | 788 | 369.63 | 122 | 1535 |
| `Text Lite` | 788 | 294.77 | 78 | 1397 |
| `Vision Intensive` | 788 | 280.39 | 78 | 1350 |
| `Vision Dominant` | 788 | 272.11 | 78 | 1356 |
| `Vision Only` | 788 | 154.1 | 70 | 222 |
**图像统计:**
| 指标 | 值 |
|--------|-------|
| 总图像数 | 3,940 |
| 每样本图像数 | 最小: 1, 最大: 1, 平均: 1 |
| 分辨率范围 | 63x70 - 6840x3549 |
| 格式 | jpeg, png |
## 样例示例
**子集**: `Text Dominant`
```json
{
"input": [
{
"id": "a3189330",
"content": [
{
"text": "Answer the following multiple choice question. The last line of your response should be of the following format: 'ANSWER: [LETTER]' (without quotes) where [LETTER] is one of A, B, C, D. Think step by step before answering.\n\nAs shown in the figure, in triangle ABC, it is known that angle A = 80.0, angle B = 60.0, point D is on AB and point E is on AC, DE parallel BC, then the size of angle CED is ()\nChoices:\nA:40°\nB:60°\nC:120°\nD:140°"
},
{
"image": "[BASE64_IMAGE: png, ~1.6KB]"
}
]
}
],
"target": "D",
"id": 0,
"group_id": 0,
"subset_key": "Text Dominant",
"metadata": {
"sample_index": "1",
"problem_index": "1",
"problem_version": "Text Dominant",
"question_type": "multi-choice",
"query_wo": "Please directly answer the question and provide the correct option letter, e.g., A, B, C, D.\nQuestion: As shown in the figure, in triangle ABC, it is known that angle A = 80.0, angle B = 60.0, point D is on AB and point E is on AC, DE parallel BC, then the size of angle CED is ()\nChoices:\nA:40°\nB:60°\nC:120°\nD:140°",
"query_cot": "Please first conduct reasoning, and then answer the question and provide the correct option letter, e.g., A, B, C, D, at the end.\nQuestion: As shown in the figure, in triangle ABC, it is known that angle A = 80.0, angle B = 60.0, point D is on AB and point E is on AC, DE parallel BC, then the size of angle CED is ()\nChoices:\nA:40°\nB:60°\nC:120°\nD:140°",
"question_for_eval": "As shown in the figure, in triangle ABC, it is known that angle A = 80.0, angle B = 60.0, point D is on AB and point E is on AC, DE parallel BC, then the size of angle CED is ()\nChoices:\nA:40°\nB:60°\nC:120°\nD:140°"
}
}
```
## 提示模板
**提示模板:**
```text
{question}
Please reason step by step, and put your final answer within \boxed{{}}.
```
## 使用方法
### 使用 CLI
```bash
evalscope eval \
--model YOUR_MODEL \
--api-url OPENAI_API_COMPAT_URL \
--api-key EMPTY_TOKEN \
--datasets math_verse \
--limit 10 # 正式评估时请删除此行
```
### 使用 Python
```python
from evalscope import run_task
from evalscope.config import TaskConfig
task_cfg = TaskConfig(
model='YOUR_MODEL',
api_url='OPENAI_API_COMPAT_URL',
api_key='EMPTY_TOKEN',
datasets=['math_verse'],
dataset_args={
'math_verse': {
# subset_list: ['Text Dominant', 'Text Lite', 'Vision Intensive'] # 可选,用于评估特定子集
}
},
limit=10, # 正式评估时请删除此行
)
run_task(task_cfg=task_cfg)
```