- cli: rich Run Summary table for multi-benchmark runs (green/red rows,
fallback to aligned plain text); unified _fmt_score (fractions render
as percentages everywhere -- was 1.0 in summary vs 100.0% in detail);
fix the stray "summary csv -> None/viz/..." print without --out-dir;
summary.md upgraded to a proper table with model/timestamp/ok-count
header -- one table for a whole N-benchmark run
- text renderer: single-bench headline deduped (dataset==recipe) and
compacted to one facts line; adaptive metric-name column (long names
no longer break alignment)
- md_compare: auto-switches to one-row-per-benchmark when comparing
different benchmarks with different metrics; same-bench model
comparison gains baseline delta markers (+/- percentage points)
- README: conda create/activate in the install block
Co-Authored-By: Claude <noreply@anthropic.com>