Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Verifies implementation completeness across functional correctness, error handling, concurrency, and security dimensions using APOSD's post-implementation checklist. Run after a coding task is nominally complete, not during active bug investigation (use cc-debugging).
.claude/skills/ryanthedev-aposd-verifying-correctness/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 104% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-19 | ✓→✗ | ▼ Worse | 88% | 0% |
| case-21 | ✓→✗ | ▼ Worse | 79% | 0% |
Design quality ≠ correctness. Well-designed code can still have bugs, missing requirements, or safety issues.
Run every dimension check before claiming done — "I think I covered everything" without explicit mapping is exactly the gap this skill exists to close.
For each dimension: detect if it applies, then verify.
Detect: Were requirements stated? (explicit list, user request, spec)
If YES, verify:
Red flag: "I think I covered everything" without explicit mapping
Detect: Any of these present?
If YES, verify:
Red flag: "It's probably fine" or "Python GIL handles it"
Detect: Can any operation fail?
If YES, verify:
except: or except Exception: passRed flag: "Errors are rare" or "caller handles it" without checking caller
Detect: Does code acquire resources?
If YES, verify:
Red flag: "It cleans up eventually" or daemon threads without shutdown
Detect: Does code handle variable-size input?
If YES, verify:
[], "", None, 0?Red flag: "Nobody would pass that" or "that's an edge case"
Detect: Does code handle untrusted input?
If YES, verify:
../ exploitation)Red flag: "It's internal only" (internals get exposed)
Detailed per-dimension checklists: Read(${CLAUDE_SKILL_DIR}/checklists.md)
When verifying, output:
## Correctness Verification
### Requirements: [PASS/FAIL/N/A]
- Requirement 1 → implemented in X
- Requirement 2 → implemented in Y
### Concurrency: [PASS/FAIL/N/A]
- Shared state: [list]
- Protection: [how]
### Errors: [PASS/FAIL/N/A]
- Failure points: [list]
- Handling: [approach]
### Resources: [PASS/FAIL/N/A]
- Acquired: [list]
- Released: [how]
### Boundaries: [PASS/FAIL/N/A]
- Edge cases: [list]
- Handling: [approach]
### Security: [PASS/FAIL/N/A]
- Untrusted input: [list]
- Validation: [approach]
**Verdict:** [DONE / NOT DONE - list blockers]| Skill | Focus | When | |-------|-------|------| | aposd-designing-deep-modules | Design quality | FIRST—during design | | aposd-verifying-correctness | Actual correctness | BEFORE "done" | | cc-quality-practices | Testing/debugging | Throughout |
Order: Design → Implement → Verify (this skill) → Done
| After | Next | |-------|------| | All dimensions pass | Done (pre-commit gate) |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→pass | 14,327 | 6,159 | -57% | 1 | 1 | 0% | 1,396 | 2,215 | +59% | 0 | 0 | — |
case-05 | fail→pass | 7,849 | 8,074 | +3% | 1 | 1 | 0% | 1,298 | 2,653 | +104% | 0 | 0 | — |
case-01 | fail→fail | 22,233 | 8,228 | -63% | 1 | 1 | 0% | 3,393 | 2,621 | -23% | 0 | 0 | — |
case-02 | pass→pass | 11,159 | 10,418 | -7% | 1 | 1 | 0% | 1,807 | 3,025 | +67% | 0 | 0 | — |
case-03 | pass→pass | 12,977 | 10,043 | -23% | 1 | 1 | 0% | 2,031 | 2,980 | +47% | 0 | 0 | — |
case-06 | fail→pass | 12,344 | 9,774 | -21% | 1 | 1 | 0% | 1,910 | 2,861 | +50% | 0 | 0 | — |
case-07 | pass→pass | 11,849 | 7,141 | -40% | 1 | 1 | 0% | 1,731 | 2,329 | +35% | 0 | 0 | — |
case-08 | pass→pass | 10,207 | 10,659 | +4% | 1 | 1 | 0% | 1,826 | 3,156 | +73% | 0 | 0 | — |
case-09 | pass→pass | 12,079 | 8,606 | -29% | 1 | 1 | 0% | 1,976 | 2,715 | +37% | 0 | 0 | — |
case-10 | fail→fail | 8,967 | 6,910 | -23% | 1 | 1 | 0% | 1,440 | 2,462 | +71% | 0 | 0 | — |
case-11 | pass→pass | 16,014 | 21,145 | +32% | 1 | 1 | 0% | 2,754 | 4,580 | +66% | 0 | 0 | — |
case-12 | pass→pass | 12,427 | 10,380 | -16% | 1 | 1 | 0% | 2,177 | 3,088 | +42% | 0 | 0 | — |
case-13 | fail→pass | 8,687 | 6,824 | -21% | 1 | 1 | 0% | 1,429 | 2,187 | +53% | 0 | 0 | — |
case-14 | pass→pass | 9,250 | 5,869 | -37% | 1 | 1 | 0% | 1,647 | 2,366 | +44% | 0 | 0 | — |
case-15 | pass→pass | 12,524 | 8,731 | -30% | 1 | 1 | 0% | 2,174 | 2,691 | +24% | 0 | 0 | — |
case-16 | pass→pass | 10,970 | 10,237 | -7% | 1 | 1 | 0% | 1,802 | 2,880 | +60% | 0 | 0 | — |
case-17 | pass→pass | 12,126 | 75,432 | +522% | 1 | 1 | 0% | 1,987 | 3,482 | +75% | 0 | 0 | — |
case-18 | pass→pass | 14,063 | 11,671 | -17% | 1 | 1 | 0% | 2,180 | 2,850 | +31% | 0 | 0 | — |
case-19 | pass→fail | 42,110 | 18,963 | -55% | 1 | 1 | 0% | 2,376 | 4,475 | +88% | 0 | 0 | — |
case-20 | fail→fail | 13,424 | 12,350 | -8% | 1 | 1 | 0% | 2,221 | 3,389 | +53% | 0 | 0 | — |
case-21 | pass→fail | 18,070 | 20,152 | +12% | 1 | 1 | 0% | 2,912 | 5,198 | +79% | 0 | 0 | — |
case-22 | fail→fail | 14,709 | 11,323 | -23% | 1 | 1 | 0% | 2,392 | 3,062 | +28% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +5 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.