Keep K3 suite selection and report-schema scoring in bash, merge K3/vision dataset_args into dpv4 yamls, and pin EvalScope at 735d920ee911 with local patches. Co-authored-by: Cursor <cursoragent@cursor.com>
28 lines
731 B
Markdown
28 lines
731 B
Markdown
# Extended Benchmarks
|
|
|
|
This section introduces evaluation benchmarks that require additional dependency packages or special configurations. These benchmarks are typically designed for specific domains or tasks, providing more specialized evaluation capabilities.
|
|
|
|
Before using these benchmarks, please follow the instructions in each benchmark's documentation to install the corresponding dependency packages and complete the necessary environment configuration.
|
|
|
|
:::{toctree}
|
|
:maxdepth: 1
|
|
|
|
terminal_bench.md
|
|
miniwob.md
|
|
skillsbench.md
|
|
toolathlon.md
|
|
gaia.md
|
|
wide_search.md
|
|
deepsearchqa.md
|
|
swe_bench.md
|
|
swe_bench_pro.md
|
|
tau_bench.md
|
|
tau2_bench.md
|
|
tau3_bench.md
|
|
bfcl_v3.md
|
|
bfcl_v4.md
|
|
needle_haystack.md
|
|
toolbench.md
|
|
longwriter.md
|
|
:::
|