Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Review changed Elixir/Phoenix code read-only. Check requirements, cite evidence, deduplicate findings, and return a severity-based verdict.
.claude/skills/oliver-kriska-phx-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 111% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 5% | 0% |
| case-19 | ✗→✓ | ▲ Improved | -15% | 0% |
| case-01 | ✓→✗ | ▼ Worse | 200% | 0% |
| case-05 | ✓→✗ | ▼ Worse | 72% | 0% |
Perform an evidence-based, read-only review of changed code. Find and explain issues; do not edit files, create tasks, or fix findings.
text/skill:phx-review /skill:phx-review test /skill:phx-review security /skill:phx-review .claude/plans/auth/plan.md /skill:phx-review --no-requirements
Treat the text after the skill name as a focus area, issue identifier, or path to a plan/specification.
describe the concrete failure mode.
highest justified severity.
optional runtime capabilities only when present.
Determine the merge base or user-specified base, then inspect:
bashgit status --short git diff --name-only <base>...HEAD git diff --stat <base>...HEAD git diff <base>...HEAD -- <changed-files>
Do not assume HEAD~5 is the correct base. Include uncommitted changes when the user asks to review the current worktree. Record the chosen scope in the result.
Unless --no-requirements is set, look for an explicit plan/spec path, current conversation requirements, a branch or commit issue identifier, or the latest relevant plan. Use available integrations or gh issue view when configured; otherwise mark requirements NOT AVAILABLE and continue.
Read references/requirements-detection.md for detection order. Never let a missing Linear, GitHub, hook, or MCP integration block code review.
Select only concerns relevant to the diff:
Native Pi subagents may run independent read-only concern tracks in parallel. Use generic subagents with the complete diff scope and return findings to this session; do not depend on separately installed named agents. If subagents are unavailable or unnecessary, run every selected concern sequentially here. A sequential review is fully valid.
For each candidate:
PRE-EXISTING.Run targeted read-only verification when it materially changes confidence. Do not alter files or suppress failures. If a check cannot run, report that clearly.
Return one verdict:
PASSPASS WITH WARNINGSREQUIRES CHANGESBLOCKEDList findings in descending severity as BLOCKER, WARNING, or SUGGESTION. Each finding must include path:line, evidence, impact, and the smallest appropriate correction. Add requirements coverage before findings; any UNMET requirement requires REQUIRES CHANGES.
If there are no findings, say so explicitly and list residual risks or checks not run. Stop after presenting the review. Suggest /skill:phx-triage, /skill:phx-plan, or /skill:phx-compound as optional next steps without invoking them automatically.
references/requirements-detection.md — requirements source and coverage rulesreferences/agent-spawning.md — Pi concern selection and optional parallelism| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→fail | 3,068 | 4,116 | +34% | 1 | 1 | 0% | 415 | 1,244 | +200% | 0 | 0 | — |
case-02 | fail→fail | 7,409 | 11,470 | +55% | 1 | 1 | 0% | 900 | 2,511 | +179% | 0 | 0 | — |
case-03 | fail→fail | 4,871 | 30,982 | +536% | 1 | 1 | 0% | 427 | 1,233 | +189% | 0 | 0 | — |
case-04 | fail→pass | 6,639 | 16,112 | +143% | 1 | 1 | 0% | 1,214 | 2,564 | +111% | 0 | 0 | — |
case-05 | pass→fail | 4,649 | 7,202 | +55% | 1 | 1 | 0% | 875 | 1,502 | +72% | 0 | 0 | — |
case-06 | pass→fail | 28,000 | 6,125 | -78% | 1 | 1 | 0% | 5,362 | 1,285 | -76% | 0 | 0 | — |
case-07 | fail→fail | 9,298 | 4,053 | -56% | 1 | 1 | 0% | 1,402 | 1,568 | +12% | 0 | 0 | — |
case-08 | fail→pass | 11,993 | 5,698 | -52% | 1 | 1 | 0% | 1,787 | 1,872 | +5% | 0 | 0 | — |
case-09 | pass→pass | 12,368 | 4,808 | -61% | 1 | 1 | 0% | 2,060 | 1,752 | -15% | 0 | 0 | — |
case-10 | pass→pass | 10,143 | 3,105 | -69% | 1 | 1 | 0% | 1,505 | 1,406 | -7% | 0 | 0 | — |
case-11 | pass→pass | 18,501 | 22,385 | +21% | 1 | 1 | 0% | 2,859 | 4,509 | +58% | 0 | 0 | — |
case-12 | pass→pass | 16,242 | 12,252 | -25% | 1 | 1 | 0% | 2,575 | 2,834 | +10% | 0 | 0 | — |
case-13 | pass→pass | 15,835 | 11,834 | -25% | 1 | 1 | 0% | 2,377 | 2,746 | +16% | 0 | 0 | — |
case-14 | pass→pass | 10,014 | 5,488 | -45% | 1 | 1 | 0% | 1,559 | 1,847 | +18% | 0 | 0 | — |
case-15 | pass→pass | 9,576 | 2,862 | -70% | 1 | 1 | 0% | 1,406 | 1,338 | -5% | 0 | 0 | — |
case-16 | pass→pass | 12,252 | 3,436 | -72% | 1 | 1 | 0% | 1,813 | 1,433 | -21% | 0 | 0 | — |
case-17 | pass→pass | 8,695 | 6,913 | -20% | 1 | 1 | 0% | 1,421 | 2,072 | +46% | 0 | 0 | — |
case-18 | pass→pass | 12,029 | 2,371 | -80% | 1 | 1 | 0% | 1,637 | 1,308 | -20% | 0 | 0 | — |
case-19 | fail→pass | 12,572 | 5,121 | -59% | 1 | 1 | 0% | 2,011 | 1,719 | -15% | 0 | 0 | — |
case-20 | pass→pass | 8,076 | 2,324 | -71% | 1 | 1 | 0% | 1,257 | 1,341 | +7% | 0 | 0 | — |
case-21 | pass→pass | 12,422 | 2,528 | -80% | 1 | 1 | 0% | 1,838 | 1,276 | -31% | 0 | 0 | — |
case-22 | pass→pass | 19,238 | 18,559 | -4% | 1 | 1 | 0% | 2,748 | 4,047 | +47% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 18 counted toward the lift figure. The other 4 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of 0 percentage points is the difference between those two pass rates over the 18 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.