Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Assesses existing module and interface design for complexity symptoms: information leakage, shallow interfaces, pass-through layers, and unknown unknowns. Produces a structured assessment — not transformations (use aposd-simplifying-complexity to edit) and not new-design generation (use aposd-designing-deep-modules).
.claude/skills/ryanthedev-aposd-reviewing-module-design/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 64% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 60% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 64% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 21% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 47% | 0% |
Run the checklist. The checklist exists because intuition misses structural problems.
Unknown unknowns are highest severity. If it's unclear what code/info is needed for changes, flag immediately.
Use this systematic checklist when reviewing code:
| Symptom | Question | If Yes | |---------|----------|--------| | Change Amplification | Does a simple change require modifications in many places? | Flag dependency problem | | Cognitive Load | Must developer know too much to work here? | Flag obscurity or leaky abstraction | | Unknown Unknowns | Is it unclear what code/info is needed for changes? | Highest severity - flag immediately |
| Check | Deep (Good) | Shallow (Bad) | |-------|-------------|---------------| | Interface vs implementation | Interface much simpler | Interface rivals implementation | | Method count | Few, powerful methods | Many, limited methods | | Hidden information | High | Low | | Common case | Simple to use | Complex to use |
Red flag: If understanding the interface isn't much simpler than understanding the implementation, the module is shallow.
| Red Flag | Detection | Severity | |----------|-----------|----------| | Information Leakage | Same knowledge in multiple modules | High | | Temporal Decomposition | Structure mirrors execution order rather than knowledge | Medium | | Back-Door Leakage | Shared knowledge not visible in interfaces but both depend on it | High | | Overexposure | Common use forces learning rare features | Medium | | Silent Failure | Module swallows errors, returns defaults, or hides failure states from callers | High |
| Red Flag | Detection | Severity | |----------|-----------|----------| | Pass-Through Method | Method only passes arguments to another with same API | High | | Adjacent Similar Abstractions | Following operation through layers, abstractions don't change | High | | Shallow Decorator | Large boilerplate, small functionality gain | Medium |
Test: Follow a single operation through layers. Does the abstraction change with each method call? If not, there's a layer problem.
| Red Flag | Detection | Severity | |----------|-----------|----------| | Conjoined Methods | Can't understand one method without another's implementation | High | | Special-General Mixture | General mechanism contains use-case specific code | High | | Code Repetition | Same code appears in multiple places | Medium | | Shallow Split | Method split resulted in interface ≈ implementation | Medium |
When evaluating whether code should be combined or separated:
1. Do pieces share information?
YES → Should probably be together
2. Would combining simplify the interface?
YES → Should probably be together
3. Is there repeated code?
YES → Extract shared method (if long snippet, simple signature)
4. Does module mix general-purpose with special-purpose?
YES → Should be separatedKey principle: Depth > Length. Never sacrifice depth for length.
| Situation | Correct Action | |-----------|---------------| | Long method with clean abstraction | Keep together | | Short method requiring another's impl to understand | Combine them | | Method split creating conjoined pair | Undo the split | | Long method with extractable subtask | Extract subtask only |
Test for valid split: Can the pieces be understood independently AND reused separately?
When reporting findings, use:
## Design Review: [Component Name]
### Critical Issues (Must Address)
- [Red flag]: [Specific location] - [Why it's a problem]
### Moderate Issues (Should Address)
- [Red flag]: [Specific location] - [Why it's a problem]
### Observations (Consider)
- [Pattern noticed] - [Potential concern]
### Positive Patterns
- [What's working well]Before reporting any red flag, validate:
Ask before concluding:
| Question | If Yes | |----------|--------| | Is there an existing pattern for this type of problem? | Compare approaches | | Does this introduce a second way to do the same thing? | Flag unless justified | | Would a maintainer be surprised by the difference? | Requires explicit documentation |
Balance: Evaluate patterns on merit, but don't create gratuitous inconsistency. The goal is maintainability, not conformance.
| Conflict | Resolution | |----------|------------| | Depth vs Cohesion | Prefer cohesion. A focused shallow module beats a bloated deep one. | | Information Hiding vs Testability | Testing seams (injectable dependencies) are acceptable "leakage" | | Simple Interface vs Configurability | Real systems need configuration; penalize only unnecessary complexity |
Detailed per-dimension checklists: Read(${CLAUDE_SKILL_DIR}/checklists.md)
| After | Next | |-------|------| | Issues found, transformation needed | Skill(code-foundations:aposd-simplifying-complexity) (transformation vs assessment) | | Issues found, plan needed | Flag for /code-foundations:plan | | No issues | Done |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 20,330 | 3,652 | -82% | 1 | 1 | 0% | 3,264 | 1,991 | -39% | 0 | 0 | — |
case-02 | fail→fail | 20,413 | 3,778 | -81% | 1 | 1 | 0% | 3,411 | 2,040 | -40% | 0 | 0 | — |
case-03 | pass→pass | 7,804 | 9,763 | +25% | 1 | 1 | 0% | 1,204 | 2,319 | +93% | 0 | 0 | — |
case-04 | pass→pass | 9,612 | 7,015 | -27% | 1 | 1 | 0% | 1,615 | 2,524 | +56% | 0 | 0 | — |
case-05 | pass→pass | 10,011 | 4,752 | -53% | 1 | 1 | 0% | 1,405 | 2,083 | +48% | 0 | 0 | — |
case-06 | fail→pass | 9,785 | 8,041 | -18% | 1 | 1 | 0% | 1,561 | 2,564 | +64% | 0 | 0 | — |
case-07 | pass→pass | 13,027 | 7,253 | -44% | 1 | 1 | 0% | 2,037 | 2,540 | +25% | 0 | 0 | — |
case-08 | fail→pass | 12,719 | 11,814 | -7% | 1 | 1 | 0% | 2,017 | 3,220 | +60% | 0 | 0 | — |
case-09 | pass→pass | 9,534 | 3,076 | -68% | 1 | 1 | 0% | 1,422 | 1,910 | +34% | 0 | 0 | — |
case-10 | pass→pass | 14,379 | 8,544 | -41% | 1 | 1 | 0% | 2,221 | 2,606 | +17% | 0 | 0 | — |
case-11 | fail→pass | 8,119 | 3,725 | -54% | 1 | 1 | 0% | 1,264 | 2,076 | +64% | 0 | 0 | — |
case-12 | pass→pass | 8,996 | 6,944 | -23% | 1 | 1 | 0% | 1,312 | 2,451 | +87% | 0 | 0 | — |
case-13 | pass→pass | 11,259 | 4,079 | -64% | 1 | 1 | 0% | 1,701 | 1,924 | +13% | 0 | 0 | — |
case-14 | fail→pass | 15,046 | 8,052 | -46% | 1 | 1 | 0% | 2,151 | 2,595 | +21% | 0 | 0 | — |
case-15 | fail→pass | 9,549 | 4,719 | -51% | 1 | 1 | 0% | 1,381 | 2,029 | +47% | 0 | 0 | — |
case-16 | pass→pass | 12,447 | 9,666 | -22% | 1 | 1 | 0% | 1,894 | 2,821 | +49% | 0 | 0 | — |
case-17 | pass→pass | 11,141 | 6,074 | -45% | 1 | 1 | 0% | 1,561 | 2,315 | +48% | 0 | 0 | — |
case-18 | pass→pass | 8,364 | 5,116 | -39% | 1 | 1 | 0% | 1,275 | 2,154 | +69% | 0 | 0 | — |
case-19 | fail→pass | 7,815 | 3,153 | -60% | 1 | 1 | 0% | 1,059 | 1,833 | +73% | 0 | 0 | — |
case-20 | fail→fail | 2,882 | 3,685 | +28% | 1 | 1 | 0% | 431 | 1,927 | +347% | 0 | 0 | — |
case-21 | pass→pass | 20,663 | 22,439 | +9% | 1 | 1 | 0% | 3,517 | 4,936 | +40% | 0 | 0 | — |
case-22 | pass→pass | 22,796 | 20,683 | -9% | 1 | 1 | 0% | 4,617 | 5,585 | +21% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.