8 Commits

Author SHA1 Message Date
sora
841b7b1ffb Bound the sandbox rm -f cleanup calls (60s)
An unbounded docker rm against a bloated daemon hangs for minutes and
silently eats the worker pool: 7 of 8 scoring workers were observed
stuck in cleanup while only 1 execution ran.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-14 03:34:13 +00:00
sora
92b36c5e6c Sandbox: reap timed-out containers, retry daemon-side failures; scoring milestones
- docker exec: named containers; a timed-out/killed 'docker run' only
  kills the CLI client while the container lives on (--rm fires on
  EXIT) -- rm -f the name on timeout/interrupt so runs stop leaking
- exit 125 = daemon-side failure, not model failure: retry up to 2x
  (a bloated daemon was turning healthy samples into pass=0)
- scoring milestones: first completion logs immediately, then every
  ~5% (10% was too sparse when docker is slow: minutes of silence
  right after the 'scoring' phase starts, looks hung)

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-14 03:26:35 +00:00
sora
46bef7d3dd dp4-flash 28-bench alignment: renderer plugin layer, exec_workers, gen_profiles, AIMD pool, BCB/LCB/bfcl/gfc judge fixes
Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-08 05:55:04 +00:00
f7c7536400 Fix images_for_samples export (swe prefetch import error) 2026-08-27 15:27:30 +00:00
7114564301 Multi-endpoint model pool (openai-pool/{lo..hi} round-robin with failover), nothink/textools spec flags (chat_template_kwargs enable_thinking=false; text-protocol tool calls for backends without --enable-auto-tool-choice), global cross-bench queue in ladder runs 2026-08-26 06:26:52 +00:00
a32d902e94 Comparison-driven fixes: MCQ choices in prompt + letter contract, Answer: suffix for QA, markdown answer cleaning, trivia_qa answer_phrase priority, DROP gold-as-alternatives (OR) official semantics, retry on 5xx, numpy/scipy compat 2026-08-24 15:48:42 +00:00
3d16ab9103 Add SWE-bench single-turn harness (patch apply + FAIL_TO_PASS in official sweb image, sh entries in docker sandbox), bigcodebench official all-libs image, docker default for humaneval/LCB, input truncation (max_input_chars) + context assembly (passage/context), API retry with backoff, judge thread-safe bridge, parquet blob integrity check, summary.csv in out-dir, sandbox prefetch CLI 2026-08-24 08:43:49 +00:00
b2e7133b20 Add sandbox layer (docker exec hard-isolation + serve envs, refcounted acquire/release, atexit teardown, bind-mount sharing) and agent evaluation driver (message pump drive(), Trajectory, bfcl_mock env with official call-sequence scoring, mock:fc oracle); Deployer delegates to sandbox.serve_env; code_any extractor fixes indentation-stripping; CLI --env 2026-08-24 07:11:03 +00:00