Logo
Explore Help
Register Sign In
Meta-Eval/EvalHarness
3
0
Fork 0
You've already forked EvalHarness
Code Issues Pull Requests Actions Packages Projects Releases Wiki Activity
EvalHarness/evalharness/agent
History
sora a93996094d swe_agentic: sentinel patch must end with a newline
The captured diff was stripped, and git apply rejects diffs whose last
line lacks a trailing newline ('corrupt patch at line N') -- the first
agentic score was resolved=0 with patch_apply_failed despite a correct
fix. Verified: same payload + newline applies clean in the official
container.

Co-Authored-By: Claude <noreply@anthropic.com>
2026-09-18 08:22:59 +00:00
..
envs
swe_agentic: sentinel patch must end with a newline
2026-09-18 08:22:59 +00:00
__init__.py
Comparison-driven fixes: MCQ choices in prompt + letter contract, Answer: suffix for QA, markdown answer cleaning, trivia_qa answer_phrase priority, DROP gold-as-alternatives (OR) official semantics, retry on 5xx, numpy/scipy compat
2026-08-24 15:48:42 +00:00
loop.py
tau2 via OFFICIAL engine: self-running env plugin (run_task hook + needs_adapter generic dispatch, no name hardcoding), deep generate() patch, official ToolCall shape, data plugin keeps verbatim Task json, official reward scoring; checkpoint plugin (per-sample resume) + summary csv with time/categories
2026-08-25 11:06:08 +00:00
Powered by Gitea Version: 1.23.5 Page: 60ms Template: 5ms
English
Bahasa Indonesia Deutsch English Español Français Gaeilge Italiano Latviešu Magyar nyelv Nederlands Polski Português de Portugal Português do Brasil Suomi Svenska Türkçe Čeština Ελληνικά Български Русский Українська فارسی മലയാളം 日本語 简体中文 繁體中文(台灣) 繁體中文(香港) 한국어
Licenses API