humaneval/gpqa_diamond: repeats 3 (align with the es reference runs)
syy's es runs: humaneval x3, gpqa x2, aime25/26 x12 (already aligned), mmlu_pro/longbench_v2 x1 (temp=0 deterministic -- repeats are noise, and 3x 12k samples is pure cost). temp=1 benches get 3. Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
parent
77d2c4569a
commit
acb94e3e20
@ -23,6 +23,7 @@ imo_answerbench:
|
||||
temperature: 1.0
|
||||
gpqa_diamond:
|
||||
temperature: 1.0
|
||||
repeats: 3
|
||||
max_tokens: 8192
|
||||
mmlu:
|
||||
max_tokens: 8192
|
||||
@ -42,6 +43,7 @@ trivia_qa:
|
||||
max_tokens: 8192
|
||||
humaneval:
|
||||
temperature: 1.0
|
||||
repeats: 3
|
||||
live_code_bench:
|
||||
temperature: 1.0
|
||||
longbench_v2:
|
||||
|
||||
Loading…
x
Reference in New Issue
Block a user