dgf1988/codewhale

Files

T

Claude 8284f395e6 test(integration): add signature field to ContentBlock::Thinking initializers and pattern in tests/ — the bins-only local run missed integration-test targets when the field landed

Co-Authored-By: Claude <noreply@anthropic.com>
https://claude.ai/code/session_018zaP8vUfTAsrE38L6h6fw5

2026-06-11 04:38:39 +00:00

features

Add Gherkin acceptance E2E harness example

2026-06-07 16:12:12 +02:00

fixtures

feat(tools): add image_ocr tool — extract text from images via tesseract

2026-05-12 00:58:48 -05:00

support

fix(release): check-versions validates the generated TUI changelog slice, not byte equality

2026-06-09 23:32:40 -07:00

cache_guard.rs

fix: resolve clippy warnings in harvested PRs (needless-borrow, is_multiple_of, dead unwrap)

2026-06-01 21:24:38 -07:00

directory_listing_acceptance.rs

Add Gherkin acceptance E2E harness example

2026-06-07 16:12:12 +02:00

eval_harness.rs

Add Gherkin acceptance E2E harness example

2026-06-07 16:12:12 +02:00

integration_mock_llm.rs

test(integration): add signature field to ContentBlock::Thinking initializers and pattern in tests/ — the bins-only local run missed integration-test targets when the field landed

2026-06-11 04:38:39 +00:00

palette_audit.rs

refactor(palette): remove unused backward-compat aliases and add module docs (#2445 )

2026-05-31 10:47:32 -07:00

protocol_recovery.rs

refactor(strings): rebrand user-facing strings to codewhale

2026-05-23 11:48:43 -05:00

qa_pty.rs

chore(tui): harden exec harness signals

2026-06-06 22:55:23 -07:00

README.md

fix(release): check-versions validates the generated TUI changelog slice, not byte equality

2026-06-09 23:32:40 -07:00

reasoning_content_replayed_after_tool_call.rs

test(integration): add signature field to ContentBlock::Thinking initializers and pattern in tests/ — the bins-only local run missed integration-test targets when the field landed

2026-06-11 04:38:39 +00:00

skill_install.rs

fix(skills): accept workflow pack archive layouts (#1164 )

2026-05-08 02:37:21 -05:00

tool_lifecycle_acceptance.rs

Address acceptance harness review feedback

2026-06-07 16:29:40 +02:00

README.md

`crates/tui/tests/`

Integration tests for the TUI binary. Per CONTRIBUTING.md, each crate's integration tests live in its own tests/ directory; the repository-root tests/ directory is unused.

Mock LLM client (`integration_mock_llm.rs`)

crates/tui/src/llm_client/mock.rs provides a MockLlmClient that implements the LlmClient trait by replaying queue-driven canned responses and capturing every outgoing MessageRequest. Tests mock at the trait boundary — never at the reqwest HTTP layer — because the trait is the durable abstraction the runtime is meant to depend on.

Coverage today exercises the trait surface end-to-end:

streaming turn loop
reasoning-content replay across tool-call rounds (V4 §5.1.1, the bug that broke v0.4.9-v0.5.1)
tool-call round-trip with chunked input JSON
multi-tool-call ordering inside a single turn
compaction-style non-streaming create_message
sub-agent style independent parent/child mocks
capacity-gate observation of a captured request before stream drain

Four full-engine tests (engine_full_*) are #[ignore]-marked. They unblock when core::engine::Engine is refactored to take Arc<dyn LlmClient> instead of a concrete Option<DeepSeekClient>. See the comment block at the bottom of integration_mock_llm.rs for the exact refactor surface.

`--record` mode for `deepseek eval`

The offline deepseek eval harness now accepts --record <DIR>. When set, each tool step appends one JSON Lines record to <DIR>/<scenario>.jsonl (default scenario: offline-tool-loop.jsonl). Each line is a self-contained JSON object with the schema:

{ "request":  { "step": "list_dir", "kind": "List" },
  "response_events": [ { "type": "ok", "output": "…" } ] }

The mock LLM client (crate::llm_client::mock) replays these fixtures by mapping each response_events array onto a canned Vec<StreamEvent>. Drop generated fixtures into crates/tui/tests/fixtures/ so they ride the repo and feed the mock in CI.

Quick example:

cargo run --bin codewhale -- eval --record crates/tui/tests/fixtures
cat crates/tui/tests/fixtures/offline-tool-loop.jsonl | jq .

The scenario name is sanitized to [A-Za-z0-9_-] before forming the filename, so unusual scenario strings stay portable across platforms.

README.md

crates/tui/tests/

Mock LLM client (integration_mock_llm.rs)

--record mode for deepseek eval

`crates/tui/tests/`

Mock LLM client (`integration_mock_llm.rs`)

`--record` mode for `deepseek eval`