Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when the user explicitly requests strict or test-first TDD, or when the current conversation already contains an explicit `TDD Route: strict` decision from another Aegis workflow.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 150% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 2488% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 473% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 174% | 0% |
| case-18 | ✗→✓ | ▲ Improved | 171% | 0% |
→ False-positive entry on a native-direct-skill host? → Exit immediately unless the user explicitly asked for TDD or the conversation already contains TDD Route: strict. In off mode, do not start RED / GREEN / REFACTOR from generic bugfix, contract, shared-module, or risky-code wording alone. Hand control back to using-aegis, systematic-debugging, writing-plans, or the fast path with verification. → Implementing a feature or bugfix under TDD Route strict? → No production code without a failing test first. Gate: medium/high complexity? → route to brainstorming or writing-plans first. Mode: default off disables automatic TDD, not completion verification; auto chooses strict/light/skipped by risk. Change Necessity: before strict RED/GREEN enters production edits, confirm the slice really needs a code change. Cycle: RED (write test → watch it fail) → GREEN (minimal code → watch it pass) → REFACTOR (clean up → keep green) Regression: shared module → related tests. contract change → producer + consumer. core logic → old + new tests. Ripple signal hit → cover producer+consumer or real user path before claiming green. GREEN proves the currently expressed behavior slice only. GREEN does not by itself prove parent-task acceptance, business-value completion, or final completion. → Done when: chosen TDD Route is recorded, strict-route tests pass, TDD preflight gate passed when applicable, pre-edit complexity risk was checked for non-trivial source edits, and verification-before-completion has fresh evidence.
Under TDD Route: strict, write the test first. Watch it fail. Write minimal code to pass.
If you didn't watch the test fail, you don't know if it tests the right thing.
TDD Mode has two values: off and auto. The default off mode disables automatic TDD routing but never disables verification-before-completion. auto lets Aegis choose a TDD Route by task risk.
On native-direct-skill hosts, automatic entry must stay anchored to literal conversation markers such as TDD Route: strict, strict TDD, test-first, or RED / GREEN / REFACTOR, not generic risky-implementation wording.
Only enter this skill after one of these explicit entry signals exists:
TDD Route: strict from another Aegis workflowTypical strict-route shapes once entry is already justified: new features, bug fixes, refactoring, behavior or logic changes, interface/data contract changes, cross-module or shared-module changes, and core logic refactors.
Exceptions (ask your human partner): throwaway prototypes, generated code, config files, pure docs cleanup, read-only diagnosis, comment-only changes.
Before source edits, decide:
textAegis Visibility: - Why this TDD route is strict, light, or skipped: - What RED/GREEN proves: - What still needs verification: TDD Route: - Mode: auto | off - Decision: strict | light | skipped - Strict authority: explicit user/project request | recorded auto decision | not applicable - Test posture: diagnostic reproduction | post-change regression | strict RED test - Reason: - Verification:
In auto, use strict for behavior, bugfix, contract, shared/core, producer / consumer, persistence, permission, migration, or meaningful regression risk. Use light for tiny low-risk edits with an obvious readback or command check. Use skipped for read-only, docs-only, generated, throwaway, comment-only, or environment-bound work where TDD does not fit.
In off, do not automatically require TDD, create a strict route, or infer one from risk alone. Explicit user/project TDD requests still apply; risky work may still need regression coverage and verification-before-completion before any completion claim. For plan or execution review, Mode: off / Decision: skipped is the normal record unless an explicit user/project strict request overrides it. That record does not load this skill or turn a diagnostic reproduction into RED. An approved plan does not supply strict authority by itself. If this skill was loaded anyway without an explicit TDD request or a visible TDD Route: strict marker, exit instead of improvising an automatic strict route from risk words alone.
Keep Aegis Visibility task-specific: explain the route decision and the regression boundary, not a generic claim that TDD was used.
TDD is the implementation discipline for an approved behavior or atomic task. It is not a substitute for task routing, product clarification, or planning.
Before writing tests or production code, stop and route to brainstorming or writing-plans if the current request has any medium- or high-complexity signal:
recovery paths
boundaries, migrations, permissions, or persistence
For these tasks, require a baseline read-set, plan, and atomic tasks before TDD. High-complexity or ambiguous tasks also need a spec/design review before planning. Only proceed directly with TDD for low-complexity work whose intent, owner, compatibility boundary, verification path, and slice goal / success evidence are already clear.
Before strict RED/GREEN enters production code edits, make the code-change decision visible. Any new source-code path needs this check before RED/GREEN normalizes it as work to implement. This is the "should code change at all?" check; it is not a new artifact and does not belong in the using-aegis hot path.
This is behavior-triggered, not prompt-triggered. If strict TDD is about to add any new source-code path or enter production source edits, expose a natural readback even when the user did not ask for it. A tiny helper, small guard, new branch, fallback, adapter, or owner is not exempt. Example: "Code necessity check: a non-code path is insufficient because <reason>; the minimum change boundary is <owner/files>, so the decision is code-change."
textChange Necessity: - User-visible need: - No-change / non-code option: - Why code change is necessary: - Minimum change boundary: - Decision: no-change | docs/config-only | code-change | needs-clarification
If the decision is no-change, do not write tests or production code for a non-change. If the decision is docs/config-only, route to that narrower surface and verify it. If the decision is needs-clarification, pause before RED/GREEN. If the decision is code-change, carry the minimum boundary into TDD Route, RED, and regression scope.
Before strict TDD on non-trivial work, record the planned complexity budget so RED/GREEN does not silently normalize a wrong or overloaded owner.
textComplexity Budget: - Artifact class: - Current pressure: - Projected post-change pressure: - Planned governance:
Use using-aegis/references/complexity-governance.md for shared artifact classes, pressure signals, and the meaning of planned governance.
Before production code edits, check whether the intended source edit would add logic to an overloaded or wrong owner. Tiny edits can keep this to one line.
Use using-aegis/references/complexity-governance.md for shared pressure signals and the meaning of over-budget.
textPre-Edit Complexity Check: - Target edit file: - Existing pressure signal: - Owner fit: - Safer edit boundary: - Decision: edit-in-place | extract helper | add owner file | split task | pause for plan update Pre-Edit Owner-Fit Decision: - Edit intent: wiring-only | move-out / extract-first | local-fix-without-new-responsibility | new-responsibility | emergency / compatibility patch - Owner fit: - Safer edit boundary: - Decision: edit-in-place | extract helper | add owner file | split task | pause for plan update
If the decision is pause for plan update, stop TDD and return to writing-plans or brainstorming with the evidence.
If the predicted result is that this slice would push a maintained artifact over budget and the slice does not also govern that overrun, do not continue with RED/GREEN as if the task were safely scoped. Pause and update the plan.
When the target edit file is over-budget or mixed-purpose, classify edit intent before production source edits. new-responsibility must not be added in place by default. wiring-only, move-out / extract-first, and local-fix-without-new-responsibility may proceed only when they do not add a new responsibility and the verification boundary is clear. emergency / compatibility patch requires residual risk and a retirement trigger.
When a medium- or high-complexity task needs project records, use configured Aegis workspace support lazily. Prefer the installed Aegis workspace helper (python <aegis-workspace-helper> init --root <target-project-root>) when it is available. If the task needs a process trail under work/, prefer python <aegis-workspace-helper> new-work --root <target-project-root> ... so the intent, checkpoint, drift, and evidence paths are indexed and structurally checkable:
textdocs/aegis/ README.md INDEX.md BASELINE-GOVERNANCE.md adr/ baseline/ specs/ plans/ work/YYYY-MM-DD-<task-slug>/ 10-intent.md 20-checkpoint.md 90-evidence.md 99-reflection.md
Do not promote reusable project facts, decisions, specs, or plans into those directories unless the workflow needs them and no existing project authority already owns them.
State: input | output | boundary | acceptance criteria. Check existing test coverage first. Write one minimal test showing what should happen. A minimal test anchors the next behavior slice; it does not by itself define whole-task completeness unless the parent acceptance is already fully pinned.
<Good>
typescripttest('retries failed operations 3 times', async () => { let attempts = 0; const operation = () => { attempts++; if (attempts < 3) throw new Error('fail'); return 'success'; }; const result = await retryOperation(operation); expect(result).toBe('success'); expect(attempts).toBe(3); });
Clear name, tests real behavior, one thing </Good>
<Bad>
typescripttest('retry works', async () => { const mock = jest.fn() .mockRejectedValueOnce(new Error()) .mockRejectedValueOnce(new Error()) .mockResolvedValueOnce('success'); await retryOperation(mock); expect(mock).toHaveBeenCalledTimes(3); });
Vague name, tests mock not code </Bad>
Requirements:
end-to-end or integration test for the main path before narrower unit tests
or operational floor before treating unit tests as sufficient
MANDATORY. Never skip.
bashnpm test path/to/test.test.ts
Confirm:
Test passes? You're testing existing behavior. Fix test.
Test errors? Fix error, re-run until it fails correctly.
Write simplest code to pass the test.
<Good>
typescriptasync function retryOperation<T>(fn: () => Promise<T>): Promise<T> { for (let i = 0; i < 3; i++) { try { return await fn(); } catch (e) { if (i === 2) throw e; } } throw new Error('unreachable'); }
Just enough to pass </Good>
<Bad>
typescriptasync function retryOperation<T>( fn: () => Promise<T>, options?: { maxRetries?: number; backoff?: 'linear' | 'exponential'; onRetry?: (attempt: number) => void; } ): Promise<T> { // YAGNI }
Over-engineered </Bad>
Don't add features, refactor other code, or "improve" beyond the test.
Fix the real owner of the behavior. Do not add a new fallback, adapter, or branch unless the debugging or design workflow identifies why it is necessary and what old path retires.
MANDATORY.
bashnpm test path/to/test.test.ts
Confirm:
Test fails? Fix code, not test.
Other tests fail? Fix now.
After green only:
Keep tests green. Don't add behavior.
Next failing test for next feature.
At minimum, run the target test you just changed or added. Broaden regression based on impact:
the downstream effect remains bounded
If the current environment cannot run automated tests, state the blocker and provide reproducible manual verification steps.
| Quality | Good | Bad | |---------|------|-----| | Minimal | One thing. "and" in name? Split it. | test('validates email and domain and whitespace') | | Clear | Name describes behavior | test('test1') | | Shows intent | Demonstrates desired API | Obscures what code should do |
These red flags apply only after this skill has validly entered under TDD Route: strict. Do not project them onto debugging or regression work whose route is light or skipped.
All of these mean: Delete code. Start over with TDD.
Bug: Empty email accepted
RED
typescripttest('rejects empty email', async () => { const result = await submitForm({ email: '' }); expect(result.error).toBe('Email required'); });
Verify RED
bash$ npm test FAIL: expected 'Email required', got undefined
GREEN
typescriptfunction submitForm(data: FormData) { if (!data.email?.trim()) { return { error: 'Email required' }; } // ... }
Verify GREEN
bash$ npm test PASS
REFACTOR Extract validation for multiple fields if needed.
TaskIntentDraft, parent plan/spec, or Slice Card exists, covered and uncovered scope are explicit before any done claimCan't check all boxes? Start over.
Exploratory spikes are allowed only as throwaway learning. When the spike ends, convert confirmed behavior into tests before formal implementation.
Emergency hotfixes may prioritize the smallest safe repair when delay is more dangerous than incomplete TDD. Record the reason, keep the change narrow, and add the missing regression test in the same slice or the next nearest slice.
Don't know how to test → write wished-for API first. Test too complicated → simplify design. Must mock everything → reduce coupling.
Bug found? Start with systematic-debugging: reproduce, trace the owner, and choose the smallest proof that supports the diagnosis. Under recorded TDD Route: strict, the reproduction becomes the required failing test before production edits. With TDD Mode: off and no strict route, use diagnostic reproduction and targeted post-change regression as fit the repair; do not start RED / GREEN by inference.
Other measured skills in the registry, with their headline benchmark lift.