dgf1988/codewhale

Files

T

Hunter B a5f27aae3a feat(benchmarks): default PinchBench to MiMo v2.5 Pro, add direct-mimo routing

PinchBench runner now defaults to openrouter/xiaomi/mimo-v2.5-pro instead
of deepseek/deepseek-chat. Adds --direct-mimo flag for routing through
Xiaomi's API directly (bypasses OpenRouter), with tp-/sk- key type
detection and endpoint mismatch warnings.

Harbor adapter gains --provider CLI flag for MiMo provider routing.

Known issues documented in docs/MIMO_BENCHMARK_ISSUES.md:
- PinchBench model validation requires OpenRouter prefix
- OPENROUTER_API_KEY needed even for some direct-provider paths
- Token Plan vs pay-as-you-go key/endpoint mismatch
- PinchBench runs through OpenClaw, not CodeWhale

2026-06-04 19:33:43 -07:00

harbor

feat(benchmarks): default PinchBench to MiMo v2.5 Pro, add direct-mimo routing

2026-06-04 19:33:43 -07:00

README.md

feat(benchmarks): add SWE-bench, Terminal-Bench, and PinchBench integration

2026-06-04 19:22:06 -07:00

run-pinchbench.sh

feat(benchmarks): default PinchBench to MiMo v2.5 Pro, add direct-mimo routing

2026-06-04 19:33:43 -07:00

run-swebench.sh

feat(benchmarks): add SWE-bench, Terminal-Bench, and PinchBench integration

2026-06-04 19:22:06 -07:00

run-terminal-bench.sh

feat(benchmarks): add SWE-bench, Terminal-Bench, and PinchBench integration

2026-06-04 19:22:06 -07:00

README.md

Benchmark Scripts

Convenience runners for evaluating CodeWhale against external benchmarks.

Quick Start

# Set your API key
export DEEPSEEK_API_KEY="sk-..."

# SWE-bench (single instance)
./scripts/benchmarks/run-swebench.sh \
  --instance-id django__django-12345 \
  --issue-file ./issue.md

# Terminal-Bench (via Harbor)
./scripts/benchmarks/run-terminal-bench.sh \
  --model deepseek/deepseek-chat

# PinchBench (auto-install + run)
./scripts/benchmarks/run-pinchbench.sh \
  --install \
  --model deepseek/deepseek-chat

Files

run-swebench.sh — SWE-bench batch driver and evaluator
run-terminal-bench.sh — Terminal-Bench runner via Harbor
run-pinchbench.sh — PinchBench runner with auto-install
harbor/__init__.py — Harbor adapter for CodeWhale (Python)
harbor/codewhale_agent.py — Adapter entry point

Documentation

See docs/BENCHMARKS.md for full setup instructions, reproducibility checklists, and references.