sora 13274243a0 Bump vendored EvalScope and add K3-ready DPV4 configs.
Keep K3 suite selection and report-schema scoring in bash, merge K3/vision dataset_args into dpv4 yamls, and pin EvalScope at 735d920ee911 with local patches.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-09-02 07:30:48 +00:00

3.9 KiB
Raw Permalink Blame History

CountQA

概述

CountQA 用于测试物体计数能力,这是一种基础的感知技能,而多模态模型在此方面基本未经过充分评估。该基准的数据集图像均在日常环境中手工拍摄,并刻意包含高密度物体、杂乱背景和遮挡情况,使得仅靠检测少量分离良好的物体无法完成计数任务。

任务描述

  • 任务类型:自由形式视觉问答(物体计数)
  • 输入:一张真实世界照片 + 一个计数问题(例如:“有多少件夹克?”)
  • 输出:一个整数
  • 领域:日常场景——包括食品杂货、厨房用具、工具、衣物、办公用品及户外物品

主要特点

  • 包含 1,528 个问答对,覆盖 1,001 张图像;每张图像可能对应多个问题
  • 真实计数值在拍摄过程中现场标注(而非事后标注),范围从 0 到 400
  • 问题中包含组合型问题,需对多种物体类型分别计数后求和
  • 约一半图像属于杂乱场景而非聚焦于单一主体(每个样本元数据中标记为 is_focused),场景类别记录在 categories 字段中

评估说明

  • 默认评估使用 test 划分作为单一子集

  • 主要指标:Accuracy (accuracy) —— 与真实整数值完全匹配

  • 次要指标:relaxed_acc —— 论文中定义的宽松准确率,当预测值在真实值的 ±5% 范围内即视为正确

  • 使用论文中的系统提示词原样不变;该提示强制模型仅输出一个纯整数

  • 答案解析规则:若回复本身是整数则直接采用;否则提取其中第一个整数——此规则与论文中重写器 LLM 所用规则一致。若回复中不含任何数字,则得分为 0因此 max_tokens 必须为模型留出足够空间输出答案;若模型以叙述方式计数(如“第一行有 3 个……”),则以其提到的第一个数字评分,而非其最终陈述的总数

  • 评分基于确定性算术运算,无需 LLM 评判器:请保持 judge.strategyruleauto,因为若设为 llm 将会用通用评判分数替代上述两个指标。若需从忽略输出格式的模型回复中提取不同数字,可通过 dataset_args 添加运行时过滤器(例如 filters={'regex': {'regex_pattern': '(\d+)', 'group_select': -1}} 提取最后一个数字),而非修改适配器

  • 论文

属性

属性
基准测试名称 count_qa
数据集ID evalscope/CountQA
论文 Paper
标签 MultiModal, QA, Reasoning
指标 accuracy, relaxed_acc
默认示例数量 0-shot
评估划分 test

数据统计

统计数据暂不可用。

样例示例

样例示例暂不可用。

提示模板

系统提示词:

You are a helpful assistant that counts the number of items in an image. The user will provide an image and ask a question about the number of a certain type of item in the image. If the user question is referring to multiple objects, it means that you need to provide a sum of the number of items. You will count the number of items and return the number as an integer. Your output should STRICTLY be a single integer and nothing else.

未定义提示模板。

使用方法

使用 CLI

evalscope eval \
    --model YOUR_MODEL \
    --api-url OPENAI_API_COMPAT_URL \
    --api-key EMPTY_TOKEN \
    --datasets count_qa \
    --limit 10  # 正式评估时请删除此行

使用 Python

from evalscope import run_task
from evalscope.config import TaskConfig

task_cfg = TaskConfig(
    model='YOUR_MODEL',
    api_url='OPENAI_API_COMPAT_URL',
    api_key='EMPTY_TOKEN',
    datasets=['count_qa'],
    limit=10,  # 正式评估时请删除此行
)

run_task(task_cfg=task_cfg)