Logo
Explore Help
Register Sign In
Meta-Eval/EvalHarness
3
0
Fork 0
You've already forked EvalHarness
Code Issues Pull Requests Actions Packages Projects Releases Wiki Activity
EvalHarness/evalharness/agent
History
sora 7085624da0 swe_agentic: env_state carries version/hints/env_commit (make_test_spec need)
MAP_REPO_VERSION_TO_SPECS[repo][version] KeyError'd on empty version --
env_state didn't store the version field, the scorer's fallback to
sample.metadata hit the positional misalign again. The four official
metadata fields now stored alongside the rest.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-18 16:03:21 +00:00
..
envs
swe_agentic: env_state carries version/hints/env_commit (make_test_spec need)
2026-09-18 16:03:21 +00:00
__init__.py
Comparison-driven fixes: MCQ choices in prompt + letter contract, Answer: suffix for QA, markdown answer cleaning, trivia_qa answer_phrase priority, DROP gold-as-alternatives (OR) official semantics, retry on 5xx, numpy/scipy compat
2026-08-24 15:48:42 +00:00
loop.py
tau2 via OFFICIAL engine: self-running env plugin (run_task hook + needs_adapter generic dispatch, no name hardcoding), deep generate() patch, official ToolCall shape, data plugin keeps verbatim Task json, official reward scoring; checkpoint plugin (per-sample resume) + summary csv with time/categories
2026-08-25 11:06:08 +00:00
Powered by Gitea Version: 1.23.5 Page: 64ms Template: 5ms
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API