humaneval/gpqa_diamond: repeats 3 (align with the es reference runs)
syy's es runs: humaneval x3, gpqa x2, aime25/26 x12 (already aligned), mmlu_pro/longbench_v2 x1 (temp=0 deterministic -- repeats are noise, and 3x 12k samples is pure cost). temp=1 benches get 3. Co-Authored-By: Claude <noreply@anthropic.com>
This commit is contained in:
parent
77d2c4569a
commit
acb94e3e20
@ -23,6 +23,7 @@ imo_answerbench:
|
|||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
gpqa_diamond:
|
gpqa_diamond:
|
||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
|
repeats: 3
|
||||||
max_tokens: 8192
|
max_tokens: 8192
|
||||||
mmlu:
|
mmlu:
|
||||||
max_tokens: 8192
|
max_tokens: 8192
|
||||||
@ -42,6 +43,7 @@ trivia_qa:
|
|||||||
max_tokens: 8192
|
max_tokens: 8192
|
||||||
humaneval:
|
humaneval:
|
||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
|
repeats: 3
|
||||||
live_code_bench:
|
live_code_bench:
|
||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
longbench_v2:
|
longbench_v2:
|
||||||
|
|||||||
Loading…
x
Reference in New Issue
Block a user