sora
5af8e2a7fd
Pluginize the run shell: @register_progress (rich/plain), @register_theme, @register_hook (on_benchmark_failed/done), @register_prober; CLI gains --progress-plugin/--theme
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-11 03:26:19 +00:00
sora
98007cccd7
Color scheme: green for facts/phrases/scores, blue for paths
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-11 03:18:52 +00:00
sora
0ca8aedefe
Scoring line: colon restored, full sentence (score over N samples), score keeps solid green with yellow phrase tail; path coloring covers 'to /abs' endings
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-11 03:16:08 +00:00
sora
d7f06395a0
Highlight whole noun phrases (1319 samples from cache / 4/4 predictions / 0 samples left to run), not bare numbers; score rule runs first so scores keep solid green
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-11 03:09:15 +00:00
sora
a8c04311c0
Semantic numbers in narration: bold yellow (bold alone was hard to see)
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-11 03:02:03 +00:00
sora
a1fa75bf89
All narration lines as full sentences (Checkpoint: 4/4 predictions already generated, 0 samples left to run / Generation skipped: ... / Scoring complete: acc 100.0% / Writing results to ...)
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-11 02:59:22 +00:00
sora
470496b7ef
Probe line carries model/endpoint/instances/probe-latency (ANSI-colored on tty); semantic numbers (counts, checkpoint fractions) bold in narration lines
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-11 02:43:10 +00:00
sora
b8ff031be7
Scores in narration lines get bold green (the key fact, previously lost in all-white lines); paths stay cyan
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-11 02:35:49 +00:00
sora
9eb1f7e58b
Fix KeyError 'cur' regression (new-task branch missed the field); live bars hard-disable unless the real stream is a tty (FORCE_COLOR env made rich claim terminal-ness on pipes -> refresh thread stalled the whole run)
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 12:02:27 +00:00
sora
e7792a559c
Narration lines get stage icons (⬇ download / 📦 ready / ✳ few-shot / ◷ checkpoint / 🤖 generating / ⏭ skipped / ★ scoring / 📝 writing / 🔗 endpoint) and cyan-highlighted paths
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 11:29:07 +00:00
sora
3911019315
writing-results line shows the destination directory
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 11:22:04 +00:00
sora
5b595f3a10
Narration lines all plain white (color scheme toggle kept in _phase_color for quick restore)
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 11:19:58 +00:00
sora
5734f7fbe6
Uniform [i/N] prefix on every narration line (status_callback restores the tag; checkpoint-restored detail routes through status_callback; probe line remains pre-callback)
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 11:15:04 +00:00
sora
392cc8daca
Persistent benchmark counter [i/N] on the single progress bar; phase lines all route through the live channel for strict ordering
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 11:13:34 +00:00
sora
9818515247
Narration colors by stage semantics: start/in-progress events yellow, completion events green (dataset ready / generation complete / scoring complete / checkpoint restored)
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 11:09:38 +00:00
sora
89f950d8a0
Narration lines uniformly uncolored (phase prints drop cyan, rich number auto-highlight off); color reserved for results/errors only
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 11:06:41 +00:00
sora
8632b2735b
Drop the overall benches bar from the CLI (per-benchmark result lines already track progress; reporter keeps the capability for API users)
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 11:04:00 +00:00
sora
b7be2a8df3
Progress bars self-describing: overall bar carries the current benchmark (benches · 1/6 humaneval), sample bar carries the stage tag (humaneval · generating/scoring/writing)
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 11:02:41 +00:00
sora
8ea83fea4f
Pause the rich live display while datasets materialize so hub tqdm progress (download/generating splits) is visible instead of being erased; reporter gains pause()/resume()
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 10:54:04 +00:00
sora
afe2fb8d28
Progress coverage for minute-scale work: overall bench bar (which benchmark of N, covers loading/scoring), byte-level download bars for dataset fetches (ModelScope/HF raw, tty-only), one shared reporter across benchmarks (caller-owned lifecycle)
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 10:47:21 +00:00
sora
c54207a180
Error panel centered; comma-joined benchmark names get an immediate fix hint
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 10:43:41 +00:00
sora
a3269252fd
Closing line simplified to one line (status + result path); tree removed
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 10:42:04 +00:00
sora
7f3c42d39c
Drop summary.md from outputs; reports switch to report.jsonl (header line + one sample per line, round-trip verified)
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 10:40:00 +00:00
sora
f9cb34fc17
Output dir restructure (evalscope/inspect-style): top-level summary.{xlsx,md,csv} + one directory per benchmark (report.json + detail.md); drops the opaque viz/ and reports/ layers
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 10:31:38 +00:00
sora
bd35b5b724
Auto-export consolidated excel workbook (results.xlsx: Summary/Perf/Categories/Samples) on every --out-dir run; per-bench artifact .md instead of .txt; xlsxwriter joins core deps
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 10:29:13 +00:00
sora
8cd45e4d16
Single-bench runs use the unified Run Summary table (panel behind --verbose); short model names in panel titles; plain-text paths in closing block (OSC8 links unreliable across terminals)
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 10:25:56 +00:00
sora
1ef230f4c1
Run Plan Samples row: show the actual sampling config instead of the confusing placeholder
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 10:23:49 +00:00
sora
e9c0b5be77
Fail-fast endpoint probe (English, full curl hint, red panel on failure); slim run narration (drop duplicate generating/materializing lines); probe success line
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 10:21:27 +00:00
sora
3a1bfb9fde
Simplify closing block to one clean status+directory tree; drop lat/trunc columns from summary
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 10:08:21 +00:00
sora
ed36fda365
Auto-save results by default (evalharness-results/<stamp>-<model>/); final notice: run finished + saved location with clickable links
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 10:05:29 +00:00
sora
c1896e0a63
Center run plan, result panel, and Run Summary table; add lat p50/p90 + truncation columns
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 10:03:19 +00:00
sora
b4d39c560e
Run Summary: tok in/out columns with per-direction throughput (in/s, out/s), totals; clickable file:// links in artifacts notice; drop note column
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 09:59:02 +00:00
sora
82d9d03b88
Run Summary: tokens/throughput/note columns + totals row; final artifacts notice (where results were saved, or how to save)
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 09:49:36 +00:00
sora
075e8728cb
Score bars: rich-style half-cell glyphs (━━━╸╌╌) replacing block chars
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 09:30:25 +00:00
sora
5c3d622f41
Rich result panel: metrics+bar with score colors, run stats (wall/model-time/throughput/tokens in-out/tok-s), top+bottom group highlights, perf row
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 09:20:44 +00:00
sora
b6c473ac26
Output polish: 1-based phase indices everywhere, drop [1/1] prefix on single-bench runs, slim progress bar (drop Waiting/Last columns)
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 09:16:33 +00:00
sora
9991c15816
README rewritten in evalscope style (numbered flow, code-first, 289->190 lines); config entry points: --hf-endpoint flag, ckpt follows cache root, EVALHARNESS_DOCKER_MIRRORS override
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 07:27:19 +00:00
sora
5bb4b75b07
Rename --judge to --judge-model (alias kept) for symmetry with --model
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 06:52:31 +00:00
sora
c78b0d6f0f
Add --api-key/--judge-api-key: explicit keys override env resolution, never serialized
...
Explicit key flows only into request headers (adapter attribute / pool
member), so it cannot leak into the spec string, EvalReport, or logs --
verified by scanning a report produced with a sentinel key. Two-key
setups run twice with different --api-key, or use per-host env vars.
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 06:50:18 +00:00
sora
ad1cd18f04
Add --judge-provider flag; document key resolution rules (env per host, judge included)
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 06:45:26 +00:00
sora
08605d9ab2
CLI flag symmetry: --judge-api-url split; mock-boxed hyphen names
...
- --judge now accepts a bare model name paired with --judge-api-url,
mirroring --model/--api-url; full legacy specs keep working
(_compose_judge_spec, verified: name+url -> spec, spec passthrough,
empty -> None; end-to-end on simple_qa with a live judge endpoint)
- mock adapter spellings: mock-boxed / mock-oracle / mock-fc preferred,
colon forms still accepted; bare 'mock' stays echo
- fix mock adapter singleton mode pollution: resolve_adapter memoizes
one instance, so mock-boxed then mock in one process leaked the
boxed mode into the echo run -- each mock spec now builds a fresh
instance
- README: mock-boxed in examples, judge flags row updated
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 06:43:27 +00:00
sora
ea93602dfa
Unified, nicer result tables + conda env setup in README
...
- cli: rich Run Summary table for multi-benchmark runs (green/red rows,
fallback to aligned plain text); unified _fmt_score (fractions render
as percentages everywhere -- was 1.0 in summary vs 100.0% in detail);
fix the stray "summary csv -> None/viz/..." print without --out-dir;
summary.md upgraded to a proper table with model/timestamp/ok-count
header -- one table for a whole N-benchmark run
- text renderer: single-bench headline deduped (dataset==recipe) and
compacted to one facts line; adaptive metric-name column (long names
no longer break alignment)
- md_compare: auto-switches to one-row-per-benchmark when comparing
different benchmarks with different metrics; same-bench model
comparison gains baseline delta markers (+/- percentage points)
- README: conda create/activate in the install block
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 06:30:22 +00:00
sora
27cf8b3c7e
Usability round: progress plugin, CLI provider flags, top-level run(), vendored BFCL checker
...
- progress/: Rich per-sample terminal progress plugin (Run Plan panel,
in-flight/rate/ETA bar); shared console + log-through-live to avoid
interleaved writes, rollback() pairs begin_sample on the retry path,
begin moved inside the semaphore (in-flight = actually generating),
graceful degradation when rich is absent
- cli.py: --provider/--api-url/--model composition (openai-chat |
openai-pool), --disable-thinking/--perf/--textools as first-class
flags, per-bench phase lines and done/failed result lines
- __init__: top-level run()/arun() entries (event-loop safe for notebooks)
- third_party/bfcl: vendored official BFCL ast_checker + type mappings
(Apache-2.0, provenance in __init__.py); imports rerouted locally,
underscore_to_dot parameterized; verified bit-identical with the
bfcl-eval package on 100 real rows -- removes the heavy extra
(pinned numpy + cloud SDK wall) from the install path
- runner: progress/status hooks through generate+evaluate, checkpoint
key scheme fix (empty-store falsy bug), tiered retry backoff,
multi-segment pool {range} expansion fix, adapter-instance passthrough
- pyproject: tree_sitter family joins core deps; [bfcl] extra retired
- README: rewritten (zh) -- install/quickstart/flags reference/bench
table/reliability/extension/architecture/validation
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-10 05:46:45 +00:00
sora
46bef7d3dd
dp4-flash 28-bench alignment: renderer plugin layer, exec_workers, gen_profiles, AIMD pool, BCB/LCB/bfcl/gfc judge fixes
...
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-08 05:55:04 +00:00
3c63338ac1
Excel renderer (xlsxwriter, 4 sheets: Summary dashboard with 21 columns, Perf detail, Categories, Samples drill-down); !perf spec flag; viz --style excel
2026-08-27 02:56:07 +00:00
e77ea4ce5c
!perf spec flag; perf_stats full dashboard columns (success_rate, latency/ttft/tpot mean+P90+P99, tokens mean/total, output TPS, request QPS); summary.csv perf columns aligned
2026-08-27 02:49:55 +00:00
78459c974e
tau2 via OFFICIAL engine: self-running env plugin (run_task hook + needs_adapter generic dispatch, no name hardcoding), deep generate() patch, official ToolCall shape, data plugin keeps verbatim Task json, official reward scoring; checkpoint plugin (per-sample resume) + summary csv with time/categories
2026-08-25 11:06:08 +00:00
1da665fec4
Add --limit-per-task (per-subset cap, evalscope --limit semantics; composable with --limit as intersection); execution-bench parity verified: official human-eval core 5/5 == our docker sandbox on same completions
2026-08-25 05:18:23 +00:00
3d16ab9103
Add SWE-bench single-turn harness (patch apply + FAIL_TO_PASS in official sweb image, sh entries in docker sandbox), bigcodebench official all-libs image, docker default for humaneval/LCB, input truncation (max_input_chars) + context assembly (passage/context), API retry with backoff, judge thread-safe bridge, parquet blob integrity check, summary.csv in out-dir, sandbox prefetch CLI
2026-08-24 08:43:49 +00:00
b2e7133b20
Add sandbox layer (docker exec hard-isolation + serve envs, refcounted acquire/release, atexit teardown, bind-mount sharing) and agent evaluation driver (message pump drive(), Trajectory, bfcl_mock env with official call-sequence scoring, mock:fc oracle); Deployer delegates to sandbox.serve_env; code_any extractor fixes indentation-stripping; CLI --env
2026-08-24 07:11:03 +00:00