One ImageService per process. When tasks arrive (runner) or a sample
starts (env), its images are REGISTERED; a background worker pool
delivers them -- local tar shipments first (es's swebench_v 500-image
batch set, disk-cached index), network mirror chain second. Sample
execution waits on a readiness barrier instead of the old failing
timings (docker-run implicit pull killed at 120s; score-phase batch
pull ran after generation had already failed).
- runner registers every pending sample's image up front (pull-ahead
overlaps generation)
- env blocks on wait_ready(1800s) before docker run -- a slow pull
delays that sample, never fails it
- tar index cached at /tmp/evalharness_tar_index.json (full scan costs
minutes; only the first process pays)
- images the service loaded are released at exit (atexit; opt out with
EVALHARNESS_KEEP_SWE_IMAGES); pre-existing local images never touched
- EVALHARNESS_IMAGE_WORKERS (default 2) tunes the pool
Verified E2E: register -> background load from swebench_batch_001.tar.gz
-> image present locally (matplotlib-14623); wait barrier semantics
confirmed (blocks until load completes).
Co-Authored-By: Claude <noreply@anthropic.com>