- accurate output structure (summary.xlsx/csv + per-bench xlsx; drop stale detail.md/summary.md mentions) - dedicated YAML-config section documenting the keys that actually work (generation params + repeats + max_input_tokens) and precedence - perf-stats table: what is always collected vs --perf streaming-only - FAQ: endpoint probe, context-overflow shrink, thinking-mode notes (non-stream chat_template_kwargs vs cloud-API param), auto-stream - config/README.md rewritten to match the flat-key reality (old file documented a directory scheme + judge/env keys that are not consumed) Co-Authored-By: Claude <noreply@anthropic.com>
配置目录
逐 bench 生成参数的 YAML,eval run 时自动应用。
config/
├── README.md ← 本文件
└── default.yaml ← 当前唯一配置(自动加载)
YAML 结构(单文件,扁平键)
default: # 协议级默认,所有 bench 继承
temperature: 0.0
top_p: 1.0
max_tokens: 32768
aime25: # 逐 bench 覆盖(键与 default 同级合并)
temperature: 1.0 # temp=1 方差测量
max_tokens: 8192
repeats: 12 # 跑 12 遍报均值;断点按轮隔离(:rep2 :rep3 ...)
longbench_v2:
max_tokens: 8192
max_input_tokens: 128000 # 超长输入的 tokenizer 中段截断预算
支持的键
| 键 | 作用 |
|---|---|
temperature top_p max_tokens |
生成参数,直传模型 |
max_input_tokens |
输入截断预算(tokenized 中段截断,头部尾部保留) |
repeats |
该 bench 重复轮数;summary 报均值,时间/token 报总和 |
优先级
DatasetSpec.gen_config < default 段 < bench 段 < 命令行显式参数
使用
# 目录里只有一个 yaml 时自动加载;--config 显式指定(省略 .yaml 后缀)
evalharness eval run aime25 --config default --api-url ... --model ...
# 多 bench:各自读自己的段
evalharness eval run aime24 aime25 aime26 hmmt26 --config default ...
新增一套配置
cp default.yaml glm53-nothink.yaml # 编辑后
evalharness eval run aime25 --config glm53-nothink ...
注意:同时存在多个 yaml 时不自动加载,必须 --config 指明。