Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Performs a comprehensive code review of a spec implementation and generates a review round directory with issue files compatible with cy-fix-reviews. Use when reviewing implemented spec tasks, creating a manual review round without an external provider, or performing a quality audit of code changes. Do not use for fetching reviews from external providers, fixing existing review issues, executing spec tasks, or editing source code.
.claude/skills/compozy-cy-review-round/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 48% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 329% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 121% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 59% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 94% | 0% |
Perform a structured code review of a spec implementation and produce a review round directory that the cy-fix-reviews workflow can process.
.compozy/tasks/<name>/ directory..compozy/tasks/<name>/.reviews-NNN/ subdirectories to determine the next round number. If none exist, use round 1..compozy/tasks/<name>/reviews-NNN/ with the round number zero-padded to 3 digits. Do NOT create it yet — wait until step 4 confirms there are issues to write. This avoids leaving empty directories when the review finds no issues._spec.md and _tasks.md from the spec directory to understand what was implemented and why, plus the contract catalogs _user_stories.md and _tests.md when present..compozy/tasks/<name>/adrs/ for architectural decision context._spec.md is missing, warn that the review will lack requirements context but proceed with a code-quality-only review.references/review-criteria.md for severity definitions and evaluation areas._tasks.md slices) and review each partition against its task's contract, core implementation files first. When the scope cannot be partitioned and still exceeds what a complete read can honestly cover, say so in the summary and recommend splitting the delivery — a review that silently degrades to sampling at scale certifies nothing._spec.md was available in step 2, cross-check the implementation against every stated requirement, acceptance criterion, and architectural decision — including every acceptance criterion and edge case in _user_stories.md when it exists. Flag any requirement that is missing, partially implemented, or implemented differently than specified. These are correctness issues — assign severity based on the gap's impact (critical if a core feature is missing, high if behavior deviates from spec, medium if an edge case from the spec is unhandled)._spec.md states a Motivating Problem, verify the delivered work solves it and name the slice that does — checking the implementation against the spec alone cannot catch a spec that drifted from its own mission. An ADR that narrowed or deferred the Motivating Problem without the user's recorded sign-off is itself a finding, at the severity of the gap it created._tests.md exists, verify that every test ID assigned in completed tasks' ## Tests sections is implemented in the suite and asserts the behavior the contract specifies. A missing case, or a hollow one that exists without asserting the contracted behavior, is an issue — assign severity based on the impact of the behavior left unverified.// nolint: intentionally ignoring close error on read-only file), do not create an issue. Only flag patterns that are genuinely problematic, not merely unconventional.references/issue-template.md for the canonical format.issue_NNN.md file in the review round directory.001 and increments sequentially. --- provider: manual pr: round: <N> round_created_at: <UTC timestamp in RFC3339 format> status: pending file: path/to/file.go line: 42 severity: high author: claude-code provider_ref: ---
# Issue NNN: <title>
## Review Comment
<detailed review body>
## Triage
UNREVIEWED<author> field must be claude-code.provider_ref field must be empty.provider field must be manual.pr field is empty for manual reviews. If the user provides a PR number, include it.round field must match the directory number as an integer (not zero-padded).round_created_at field must use the same current UTC RFC3339 timestamp in every issue in this round.severity field must be exactly one of: critical, high, medium, low.compozy loop run --workspace <ref> --name review-and-fix --input task_name=<name> to process the review round.cy-final-verify before claiming the review round is complete.provider, pr, round, and round_created_at values.reviews-NNN naming convention.cy-fix-reviews workflow handles remediation.prompt.ParseReviewContext()._meta.md; round metadata lives in each issue file frontmatter.gh mutations._spec.md is missing, warn about the lack of requirements context but proceed with code-quality-only review.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 7,768 | 5,630 | -28% | 1 | 1 | 0% | 119 | 2,549 | +2042% | 0 | 0 | — |
case-02 | fail→fail | 6,519 | 39,696 | +509% | 1 | 1 | 0% | 243 | 2,503 | +930% | 0 | 0 | — |
case-03 | fail→fail | 5,561 | 11,232 | +102% | 1 | 1 | 0% | 276 | 2,462 | +792% | 0 | 0 | — |
case-04 | fail→fail | 12,015 | 8,883 | -26% | 1 | 1 | 0% | 400 | 2,907 | +627% | 0 | 0 | — |
case-05 | fail→pass | 60,728 | 21,830 | -64% | 1 | 1 | 0% | 3,193 | 4,722 | +48% | 0 | 0 | — |
case-06 | fail→pass | 5,895 | 21,497 | +265% | 1 | 1 | 0% | 899 | 3,855 | +329% | 0 | 0 | — |
case-07 | fail→pass | 8,590 | 6,001 | -30% | 1 | 1 | 0% | 1,338 | 2,961 | +121% | 0 | 0 | — |
case-08 | fail→pass | 65,746 | 3,944 | -94% | 1 | 1 | 0% | 1,781 | 2,835 | +59% | 0 | 0 | — |
case-09 | fail→pass | 11,988 | 4,997 | -58% | 1 | 1 | 0% | 1,573 | 3,048 | +94% | 0 | 0 | — |
case-10 | pass→pass | 12,650 | 5,356 | -58% | 1 | 1 | 0% | 1,739 | 2,881 | +66% | 0 | 0 | — |
case-11 | pass→pass | 25,679 | 5,587 | -78% | 1 | 1 | 0% | 1,888 | 3,116 | +65% | 0 | 0 | — |
case-12 | pass→pass | 11,177 | 5,568 | -50% | 1 | 1 | 0% | 1,541 | 2,872 | +86% | 0 | 0 | — |
case-13 | fail→pass | 19,625 | 4,000 | -80% | 1 | 1 | 0% | 1,420 | 2,854 | +101% | 0 | 0 | — |
case-14 | pass→pass | 10,604 | 3,857 | -64% | 1 | 1 | 0% | 1,295 | 2,771 | +114% | 0 | 0 | — |
case-15 | fail→pass | 11,584 | 8,030 | -31% | 1 | 1 | 0% | 1,872 | 3,102 | +66% | 0 | 0 | — |
case-16 | fail→pass | 47,143 | 6,116 | -87% | 1 | 1 | 0% | 2,663 | 3,215 | +21% | 0 | 0 | — |
case-17 | fail→pass | 12,058 | 3,009 | -75% | 1 | 1 | 0% | 1,440 | 2,650 | +84% | 0 | 0 | — |
case-18 | pass→pass | 24,512 | 5,698 | -77% | 1 | 1 | 0% | 1,453 | 2,980 | +105% | 0 | 0 | — |
case-19 | pass→pass | 17,522 | 7,707 | -56% | 1 | 1 | 0% | 2,096 | 3,341 | +59% | 0 | 0 | — |
case-20 | pass→pass | 8,261 | 3,969 | -52% | 1 | 1 | 0% | 1,340 | 2,925 | +118% | 0 | 0 | — |
case-21 | pass→pass | 12,015 | 5,248 | -56% | 1 | 1 | 0% | 1,588 | 3,056 | +92% | 0 | 0 | — |
case-22 | fail→pass | 9,562 | 5,633 | -41% | 1 | 1 | 0% | 1,349 | 2,916 | +116% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +45 percentage points is the difference between those two pass rates over the 18 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.