Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Review an entire codebase for architecture, engineering health, and exploitable risk; generate a prioritized remediation plan, an evidence-anchored system knowledge document, or both.
.claude/skills/hoangnguyen0403-codebase-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-19 | ✗→✓ | ▲ Improved | -4% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 98% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 1% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -14% | 0% |
> !IMPORTANT] > Review an entire codebase for architecture, engineering health, and exploitable risk; generate a prioritized remediation plan, an evidence-anchored system knowledge document, or both.
Optional args: slug=<feature>, ticket=<id/url>, mode=interactive|autonomous|channel, channel=<id>, auto_continue=true|false, profile=business|hybrid|technical.
When the user asks to perform this workflow, execute the following steps:
> Goal: Map a codebase from evidence, expose systemic risk, and produce the review and/or knowledge artifact requested.
analysis=fast|deep and deliverable=review|knowledge|both; default to fast + review to preserve the existing audit behavior. Use deep for knowledge or both unless the user explicitly requests otherwise.package.json, go.mod, pubspec.yaml, pom.xml) and locate source, tests, docs, IaC, runtime config, entry points, data stores, and generated paths.common-architecture-audit, common-security-audit, common-owasp, and common-llm-security.trusted, semi-trusted, or untrusted; record missing or inaccessible evidence.trigger -> validation/auth -> state mutation -> side effect -> consumer. Map cross-cutting logging, caching, error handling, authentication, and authorization.fast: inspect largest non-generated files, changed hotspots, auth surfaces, execution/config chokepoints, and the highest-centrality modules.deep: also inspect service-to-service flows, state lifecycle, persistence/migrations, jobs/events, feature boundaries, architecture drift, compliance-sensitive paths, and LLM/agent runtime risks.reviewContext for the pass: analysisMode, promptInjectionRisk, delegationMode, assignedRoles, and false-positive controls used by the human or agent team.confirmed.confirmed, needs validation, and not enough evidence separate.design-solution with explicit security constraints and follow-up questions.review or both, write artifacts/codebase-review.md with engineering health, architecture, delivery risk, severity-ranked findings, evidence gaps, and phased remediation. Score from 100: Critical -15, High -8, Medium -3, Low -1; cap at 40 for any P0.knowledge or both, write docs/architecture/codebase-knowledge.md with system purpose, evidence/assumptions, component map, critical flows and state ownership, integrations/trust boundaries, interaction matrix, change cautions, risks, glossary, and coverage/next-read queue. Use Mermaid only when it clarifies a real relationship.artifacts/security-review.md with scope, trust boundaries, review context, runtime contract, findings, evidence gaps, source provenance, confidence, exploit path, control mapping, and handoff notes.partial, preserve the ordered next-read queue, and do not present the knowledge document as complete.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-14 | pass→pass | 13,742 | 4,486 | -67% | 1 | 1 | 0% | 2,270 | 1,998 | -12% | 0 | 0 | — |
case-19 | fail→pass | 11,889 | 3,779 | -68% | 1 | 1 | 0% | 1,879 | 1,803 | -4% | 0 | 0 | — |
case-01 | fail→fail | 9,306 | 18,828 | +102% | 1 | 1 | 0% | 1,338 | 4,231 | +216% | 0 | 0 | — |
case-02 | fail→fail | 4,115 | 3,871 | -6% | 1 | 1 | 0% | 249 | 1,535 | +516% | 0 | 0 | — |
case-03 | fail→fail | 13,209 | 16,814 | +27% | 1 | 1 | 0% | 1,532 | 3,587 | +134% | 0 | 0 | — |
case-04 | fail→pass | 13,339 | 14,037 | +5% | 1 | 1 | 0% | 1,457 | 2,878 | +98% | 0 | 0 | — |
case-05 | pass→pass | 11,329 | 5,890 | -48% | 1 | 1 | 0% | 1,787 | 2,197 | +23% | 0 | 0 | — |
case-06 | fail→pass | 9,796 | 4,017 | -59% | 1 | 1 | 0% | 1,603 | 1,716 | +7% | 0 | 0 | — |
case-07 | fail→pass | 10,927 | 3,315 | -70% | 1 | 1 | 0% | 1,865 | 1,891 | +1% | 0 | 0 | — |
case-13 | fail→pass | 13,698 | 4,727 | -65% | 1 | 1 | 0% | 2,510 | 2,154 | -14% | 0 | 0 | — |
case-08 | fail→pass | 10,584 | 2,854 | -73% | 1 | 1 | 0% | 1,985 | 1,814 | -9% | 0 | 0 | — |
case-09 | pass→pass | 12,443 | 8,458 | -32% | 1 | 1 | 0% | 2,201 | 2,786 | +27% | 0 | 0 | — |
case-10 | fail→pass | 7,583 | 4,247 | -44% | 1 | 1 | 0% | 1,259 | 2,005 | +59% | 0 | 0 | — |
case-11 | pass→pass | 13,566 | 7,504 | -45% | 1 | 1 | 0% | 2,104 | 2,480 | +18% | 0 | 0 | — |
case-12 | pass→pass | 9,776 | 3,260 | -67% | 1 | 1 | 0% | 1,650 | 1,831 | +11% | 0 | 0 | — |
case-15 | fail→fail | 6,765 | 4,395 | -35% | 1 | 1 | 0% | 1,165 | 1,440 | +24% | 0 | 0 | — |
case-16 | pass→pass | 11,533 | 7,211 | -37% | 1 | 1 | 0% | 2,128 | 2,602 | +22% | 0 | 0 | — |
case-17 | pass→pass | 14,644 | 22,116 | +51% | 1 | 1 | 0% | 2,577 | 4,940 | +92% | 0 | 0 | — |
case-18 | pass→pass | 13,110 | 5,421 | -59% | 1 | 1 | 0% | 1,970 | 2,153 | +9% | 0 | 0 | — |
case-20 | pass→fail | 7,543 | 12,746 | +69% | 1 | 1 | 0% | 1,469 | 3,203 | +118% | 0 | 0 | — |
case-21 | pass→pass | 14,181 | 15,367 | +8% | 1 | 1 | 0% | 3,045 | 4,665 | +53% | 0 | 0 | — |
case-22 | fail→fail | 1,789 | 4,782 | +167% | 1 | 1 | 0% | 287 | 1,550 | +440% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 21 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.