Keep K3 suite selection and report-schema scoring in bash, merge K3/vision dataset_args into dpv4 yamls, and pin EvalScope at 735d920ee911 with local patches. Co-authored-by: Cursor <cursoragent@cursor.com>
3.9 KiB
3.9 KiB
CountQA
概述
CountQA 用于测试物体计数能力,这是一种基础的感知技能,而多模态模型在此方面基本未经过充分评估。该基准的数据集图像均在日常环境中手工拍摄,并刻意包含高密度物体、杂乱背景和遮挡情况,使得仅靠检测少量分离良好的物体无法完成计数任务。
任务描述
- 任务类型:自由形式视觉问答(物体计数)
- 输入:一张真实世界照片 + 一个计数问题(例如:“有多少件夹克?”)
- 输出:一个整数
- 领域:日常场景——包括食品杂货、厨房用具、工具、衣物、办公用品及户外物品
主要特点
- 包含 1,528 个问答对,覆盖 1,001 张图像;每张图像可能对应多个问题
- 真实计数值在拍摄过程中现场标注(而非事后标注),范围从 0 到 400
- 问题中包含组合型问题,需对多种物体类型分别计数后求和
- 约一半图像属于杂乱场景而非聚焦于单一主体(每个样本元数据中标记为
is_focused),场景类别记录在categories字段中
评估说明
-
默认评估使用 test 划分作为单一子集
-
主要指标:Accuracy (
accuracy) —— 与真实整数值完全匹配 -
次要指标:relaxed_acc —— 论文中定义的宽松准确率,当预测值在真实值的 ±5% 范围内即视为正确
-
使用论文中的系统提示词原样不变;该提示强制模型仅输出一个纯整数
-
答案解析规则:若回复本身是整数则直接采用;否则提取其中第一个整数——此规则与论文中重写器 LLM 所用规则一致。若回复中不含任何数字,则得分为 0,因此
max_tokens必须为模型留出足够空间输出答案;若模型以叙述方式计数(如“第一行有 3 个……”),则以其提到的第一个数字评分,而非其最终陈述的总数 -
评分基于确定性算术运算,无需 LLM 评判器:请保持
judge.strategy为rule或auto,因为若设为llm将会用通用评判分数替代上述两个指标。若需从忽略输出格式的模型回复中提取不同数字,可通过dataset_args添加运行时过滤器(例如filters={'regex': {'regex_pattern': '(\d+)', 'group_select': -1}}提取最后一个数字),而非修改适配器
属性
| 属性 | 值 |
|---|---|
| 基准测试名称 | count_qa |
| 数据集ID | evalscope/CountQA |
| 论文 | Paper |
| 标签 | MultiModal, QA, Reasoning |
| 指标 | accuracy, relaxed_acc |
| 默认示例数量 | 0-shot |
| 评估划分 | test |
数据统计
统计数据暂不可用。
样例示例
样例示例暂不可用。
提示模板
系统提示词:
You are a helpful assistant that counts the number of items in an image. The user will provide an image and ask a question about the number of a certain type of item in the image. If the user question is referring to multiple objects, it means that you need to provide a sum of the number of items. You will count the number of items and return the number as an integer. Your output should STRICTLY be a single integer and nothing else.
未定义提示模板。
使用方法
使用 CLI
evalscope eval \
--model YOUR_MODEL \
--api-url OPENAI_API_COMPAT_URL \
--api-key EMPTY_TOKEN \
--datasets count_qa \
--limit 10 # 正式评估时请删除此行
使用 Python
from evalscope import run_task
from evalscope.config import TaskConfig
task_cfg = TaskConfig(
model='YOUR_MODEL',
api_url='OPENAI_API_COMPAT_URL',
api_key='EMPTY_TOKEN',
datasets=['count_qa'],
limit=10, # 正式评估时请删除此行
)
run_task(task_cfg=task_cfg)