codewhale/scripts/benchmarks/harbor/codewhale_agent.py at b4edb4e1ef91e5523df4c96629e27c38233ddc41 - codewhale - Gitea: Git with a cup of tea

dgf1988/codewhale

Files

T

Hunter B b329a532f5 feat(benchmarks): add SWE-bench, Terminal-Bench, and PinchBench integration

Benchmark harness for evaluating CodeWhale against three external
benchmarks:

- SWE-bench: batch driver wrapping existing codewhale swebench commands
- Terminal-Bench: Harbor adapter (BaseInstalledAgent) for container eval
- PinchBench: runner with auto-install for real-world agent tasks

Includes docs/BENCHMARKS.md umbrella doc with setup, usage, and
reproducibility checklist. Scripts record version/commit/timestamp
metadata for each run.

Branch: codex/v0.8.53-benchmarks (based on v0.8.53)

2026-06-04 19:22:06 -07:00

5 lines

145 B

Python

Raw Blame History

 """Harbor adapter entry point for CodeWhale."""
 from scripts.benchmarks.harbor import CodeWhaleAgent  # noqa: F401
 __all__ = ["CodeWhaleAgent"]