Keep K3 suite selection and report-schema scoring in bash, merge K3/vision dataset_args into dpv4 yamls, and pin EvalScope at 735d920ee911 with local patches. Co-authored-by: Cursor <cursoragent@cursor.com>
5.1 KiB
5.1 KiB
PerceptionBench
Overview
PerceptionBench is a benchmark from Moonshot AI that evaluates the atomic visual perception capabilities of multimodal large language models. It is built bottom-up: the earliest failure points of frontier MLLMs on 42 existing benchmarks were diagnosed to derive an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Each question isolates a single capability, so difficulty stems from perception rather than reasoning or knowledge.
Task Description
- Task Type: Visual Perception (open-ended question answering)
- Input: One or more images interleaved with a question
- Output: Free-form short answer with a uniquely determined reference
- Domain: Atomic visual perception across ten capabilities
Key Features
- 3,000 verified questions covering ten atomic perceptual capabilities
- 1,800 questions (60%) are atomic sub-questions decomposed from attributed failures on source benchmarks; 1,200 (40%) are newly authored on supplemented images
- Subsets follow the ten
error_categorylabels: visual relation, counting, attribute, depth & 3D perception, localization, comparison, fine-grained recognition, contextual integration, OCR, and perception-related hallucination - Multi-image questions are supported: images are interleaved into the question via
<|image_N|>placeholders - Samples carrying a
hint(coordinate convention or image dimensions) pass it as a system message, matching the official message builder
Evaluation Notes
- Default evaluation uses the train split (3,000 samples, single split dataset)
- Primary metric: Accuracy, reported overall and per capability
- Scoring follows the official protocol: an LLM judge grades the free-form answer against the
reference with the teacher-grading prompt and returns a strict 0/1 verdict per item
(
[reason]/[judge] True|False); the paper uses GPT-oss-120B, whose agreement with human judgment is 99.7% on a 300-sample audit - Empty or failed generations are scored 0 without invoking the judge
- Requires an LLM judge configured through
judge.models - The dataset embeds images as base64 data URIs (~1.6 GB download on first use)
Properties
| Property | Value |
|---|---|
| Benchmark Name | perception_bench |
| Dataset ID | moonshotai/PerceptionBench |
| Paper | Paper |
| Tags | MultiModal, QA |
| Metrics | accuracy |
| Default Shots | 0-shot |
| Evaluation Split | train |
Data Statistics
| Metric | Value |
|---|---|
| Total Samples | 3,000 |
| Prompt Length (Mean) | 233.87 chars |
| Prompt Length (Min/Max) | 29 / 1076 chars |
Per-Subset Statistics:
| Subset | Samples | Prompt Mean | Prompt Min | Prompt Max |
|---|---|---|---|---|
visual_relation_error |
330 | 275.62 | 43 | 876 |
visual_counting_error |
330 | 161.11 | 37 | 831 |
visual_attribute_error |
330 | 225.58 | 34 | 1006 |
depth_3d_perception_error |
330 | 278.5 | 60 | 976 |
visual_localization_error |
330 | 284.79 | 62 | 1076 |
visual_comparison_error |
279 | 270.14 | 39 | 801 |
fine_grained_recognition_error |
290 | 225.91 | 44 | 917 |
context_integration_error |
255 | 277.04 | 58 | 845 |
ocr_error |
255 | 175.39 | 29 | 934 |
hallucination |
271 | 150.94 | 42 | 515 |
Image Statistics:
| Metric | Value |
|---|---|
| Total Images | 3,567 |
| Images per Sample | min: 1, max: 8, mean: 1.19 |
| Resolution Range | 101x64 - 5712x4953 |
| Formats | jpeg, png, webp |
Sample Example
Subset: visual_relation_error
{
"input": [
{
"id": "28bf28ec",
"content": [
{
"image": "[BASE64_IMAGE: png, ~97.9KB]"
},
{
"text": "How many arrows does the dashed box intersect with? Just answer with the number."
}
]
}
],
"target": "4",
"id": 0,
"group_id": 0,
"subset_key": "visual_relation_error",
"metadata": {
"index": 5,
"problem": "<|image_1|>How many arrows does the dashed box intersect with? Just answer with the number.",
"error_category": "visual_relation_error",
"source_bmk": "NA",
"source_idx": null
}
}
Prompt Template
No prompt template defined.
Usage
Using CLI
evalscope eval \
--model YOUR_MODEL \
--api-url OPENAI_API_COMPAT_URL \
--api-key EMPTY_TOKEN \
--datasets perception_bench \
--limit 10 # Remove this line for formal evaluation
Using Python
from evalscope import run_task
from evalscope.config import TaskConfig
task_cfg = TaskConfig(
model='YOUR_MODEL',
api_url='OPENAI_API_COMPAT_URL',
api_key='EMPTY_TOKEN',
datasets=['perception_bench'],
dataset_args={
'perception_bench': {
# subset_list: ['visual_relation_error', 'visual_counting_error', 'visual_attribute_error'] # optional, evaluate specific subsets
}
},
limit=10, # Remove this line for formal evaluation
)
run_task(task_cfg=task_cfg)