- perf_stats aggregator lives in eval/, not model/: the import failed silently and EVERY perf column was empty (not just ttft). Now warns on stderr instead of swallowing. - repeats > 1 get their own checkpoint key (:rep2, :rep3, ...): repeat 2 previously restored repeat 1's predictions and finished instantly with identical scores. rep1 keeps the legacy key (existing checkpoints still resume). - repeats summary: report the MEAN score and aggregate time/tokens over ALL runs (was: last run only). - README: six-benchmark command as the primary example. Co-Authored-By: Claude <noreply@anthropic.com>
9 lines
387 B
Python
9 lines
387 B
Python
"""Vendored BFCL official AST checker (from bfcl-eval, Apache-2.0).
|
|
|
|
Source: bfcl_eval/eval_checker/ast_eval/{ast_checker.py,type_convertor/}
|
|
+ bfcl_eval/constants/type_mappings.py
|
|
Only change: imports rerouted locally and the MODEL_CONFIG_MAPPING lookup
|
|
replaced by an explicit ``underscore_to_dot`` argument. Upstream license
|
|
and notice apply to the files in this directory.
|
|
"""
|