- sample-counts manifest (config/sample_counts.yaml, harvested from
real runs): uncached benches still show exact numbers in the plan
instead of 'counts when datasets load' -- 'cache+est.' marks the mix
- '--concurrency auto' is now an alias for --auto-concurrency
- Concurrency row shows 'auto (start 8, gate decides)' when the gate
drives, instead of a bare misleading 8
Also verified end-to-end: thinking-mode humaneval rep1/rep2 both
pass 98.8%, matching the es reference runs (98.17/98.78/98.78) on the
same model -- framework alignment holds on the thinking path too.
Co-Authored-By: Claude <noreply@anthropic.com>