sora 7921688149 Run plan: concrete sample counts + 'auto' concurrency display/alias
- sample-counts manifest (config/sample_counts.yaml, harvested from
  real runs): uncached benches still show exact numbers in the plan
  instead of 'counts when datasets load' -- 'cache+est.' marks the mix
- '--concurrency auto' is now an alias for --auto-concurrency
- Concurrency row shows 'auto (start 8, gate decides)' when the gate
  drives, instead of a bare misleading 8

Also verified end-to-end: thinking-mode humaneval rep1/rep2 both
pass 98.8%, matching the es reference runs (98.17/98.78/98.78) on the
same model -- framework alignment holds on the thinking path too.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-14 09:11:07 +00:00
..

配置目录

逐 bench 生成参数的 YAMLeval run 时自动应用。

config/
├── README.md        ← 本文件
└── default.yaml     ← 当前唯一配置(自动加载)

YAML 结构(单文件,扁平键)

default:                # 协议级默认,所有 bench 继承
  temperature: 0.0
  top_p: 1.0
  max_tokens: 32768

aime25:                 # 逐 bench 覆盖(键与 default 同级合并)
  temperature: 1.0      # temp=1 方差测量
  max_tokens: 8192
  repeats: 12           # 跑 12 遍报均值;断点按轮隔离(:rep2 :rep3 ...

longbench_v2:
  max_tokens: 8192
  max_input_tokens: 128000   # 超长输入的 tokenizer 中段截断预算

支持的键

作用
temperature top_p max_tokens 生成参数,直传模型
max_input_tokens 输入截断预算tokenized 中段截断,头部尾部保留)
repeats 该 bench 重复轮数summary 报均值,时间/token 报总和

优先级

DatasetSpec.gen_config  <  default 段  <  bench 段  <  命令行显式参数

使用

# 目录里只有一个 yaml 时自动加载;--config 显式指定(省略 .yaml 后缀)
evalharness eval run aime25 --config default --api-url ... --model ...

# 多 bench各自读自己的段
evalharness eval run aime24 aime25 aime26 hmmt26 --config default ...

新增一套配置

cp default.yaml glm53-nothink.yaml   # 编辑后
evalharness eval run aime25 --config glm53-nothink ...

注意:同时存在多个 yaml 时不自动加载,必须 --config 指明。