- Add --thinking-max-tokens-scale to multiply max_tokens when --thinking is enabled. - Add 'official' suite covering all benchmarks in the public comparison tables (Kimi/GLM/DS etc). - Print scale factor in run summary when thinking is on.
- Add --thinking-max-tokens-scale to multiply max_tokens when --thinking is enabled. - Add 'official' suite covering all benchmarks in the public comparison tables (Kimi/GLM/DS etc). - Print scale factor in run summary when thinking is on.