The adapter dropped reasoning_effort/thinking from the payload, so the
middle rung of the ladder (es reference: full 98.3 / effort_low 94 /
no-think 82.3 on humaneval) was unreachable. Both keys now pass
through; --reasoning-effort {minimal,low,medium,high,max} overrides
the YAML, and config/effort_low.yaml mirrors default.yaml with
reasoning_effort: low for one-command low-thinking runs.
Probe on the endpoint: same question, default = 319 chars reasoning /
262 tok, low = 47 chars / 117 tok -- the server honors it.
Co-Authored-By: Claude <noreply@anthropic.com>
配置目录
逐 bench 生成参数的 YAML,eval run 时自动应用。
config/
├── README.md ← 本文件
└── default.yaml ← 当前唯一配置(自动加载)
YAML 结构(单文件,扁平键)
default: # 协议级默认,所有 bench 继承
temperature: 0.0
top_p: 1.0
max_tokens: 32768
aime25: # 逐 bench 覆盖(键与 default 同级合并)
temperature: 1.0 # temp=1 方差测量
max_tokens: 8192
repeats: 12 # 跑 12 遍报均值;断点按轮隔离(:rep2 :rep3 ...)
longbench_v2:
max_tokens: 8192
max_input_tokens: 128000 # 超长输入的 tokenizer 中段截断预算
支持的键
| 键 | 作用 |
|---|---|
temperature top_p max_tokens |
生成参数,直传模型 |
max_input_tokens |
输入截断预算(tokenized 中段截断,头部尾部保留) |
repeats |
该 bench 重复轮数;summary 报均值,时间/token 报总和 |
优先级
DatasetSpec.gen_config < default 段 < bench 段 < 命令行显式参数
使用
# 目录里只有一个 yaml 时自动加载;--config 显式指定(省略 .yaml 后缀)
evalharness eval run aime25 --config default --api-url ... --model ...
# 多 bench:各自读自己的段
evalharness eval run aime24 aime25 aime26 hmmt26 --config default ...
新增一套配置
cp default.yaml glm53-nothink.yaml # 编辑后
evalharness eval run aime25 --config glm53-nothink ...
注意:同时存在多个 yaml 时不自动加载,必须 --config 指明。