humaneval/gpqa_diamond: repeats 3 (align with the es reference runs)

syy's es runs: humaneval x3, gpqa x2, aime25/26 x12 (already aligned),
mmlu_pro/longbench_v2 x1 (temp=0 deterministic -- repeats are noise,
and 3x 12k samples is pure cost). temp=1 benches get 3.

Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
sora 2026-09-14 07:39:55 +00:00
parent 77d2c4569a
commit acb94e3e20

View File

@ -23,6 +23,7 @@ imo_answerbench:
temperature: 1.0
gpqa_diamond:
temperature: 1.0
repeats: 3
max_tokens: 8192
mmlu:
max_tokens: 8192
@ -42,6 +43,7 @@ trivia_qa:
max_tokens: 8192
humaneval:
temperature: 1.0
repeats: 3
live_code_bench:
temperature: 1.0
longbench_v2: