time_s/time_h summed per-prediction latency_s, which includes RESTORED
predictions' original generation time -- days old and from a slower
setup, it once reported 15.9h for a one-hour aime25 run. All rows now
report the bench's actual wall clock; token totals stay as the true
cost of the predictions used.
n for repeats>1 is num_samples x repeats (12 runs over 30 problems is
360 generations, not 30).
Co-Authored-By: Claude <noreply@anthropic.com>