Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Drive a change through a red-green-refactor loop - failing test first, minimal code to pass, then clean up. Use when implementing a feature or fixing a bug where correctness matters and a test can pin the behavior. Says "TDD", "test first", "red green refactor", "write the test first".
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 175% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -7% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 20% | 0% |
| case-14 | ✗→✓ | ▲ Improved | -17% | 0% |
The agent codes better with a tight feedback loop than with a long specification. A failing test is the tightest loop there is: it states the target, and the target either goes green or it does not.
Run it. Confirm it fails for the right reason - a test that passes before you write the code is testing nothing. If it errors instead of failing on the assertion, fix the test setup first.
solution, not the abstraction - the minimal move. Run the test. Green.
name things, simplify. Re-run after every edit. Behavior is frozen; only the shape changes.
speed limit - never take on a step too big to hold a single test.
not on which private method got called. A test that breaks on every refactor is a liability.
should know what broke without reading the body.
filesystem), not internal collaborators. A test that mocks the thing under test asserts nothing.
expect(x).toBe(x)after setting x) verifies nothing. Assert the value the behavior should produce.
Reach for TDD when the behavior is specifiable and a test can pin it: business logic, parsers, state machines, bug fixes (write the failing case first, then fix). Skip it for pure exploration, throwaway spikes, and layout-only UI where a test asserts nothing a human would not eyeball.
Show each red-green transition, not just the final green. If a test is hard to write, say what the difficulty reveals about the design - untestable code is usually badly-seamed code, and that is a finding, not an obstacle.
Other measured skills in the registry, with their headline benchmark lift.