Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Challenge mode reviews - rigorous questioning before approving changes. Use when you want thorough scrutiny of Ecto changes, LiveView events, OTP designs, or PR readiness.
.claude/skills/oliver-kriska-challenge/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 97% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 62% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 19% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 45% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 39% | 0% |
Rigorous, critical review patterns inspired by Boris Cherny's "Grill me" approach. Push beyond first solutions to ensure quality.
/phx:challenge ecto)Grill the developer on database changes:
Migration Safety
Query Performance
Schema Integrity
Backward Compatibility
/phx:challenge liveview)Prove the LiveView handles all cases:
Event Coverage
handle_event clause and expected socket statePubSub Handling
handle_info clause and when it's triggeredState Transitions
Memory & Performance
/phx:challenge pr)Senior engineer review checklist:
Must Pass
Performance
OTP
Security
CRITICAL: Prevents re-discovering identical issues across consecutive runs.
.claude/plans/*/reviews/ and .claude/reviews/ for prior findingsmarkdown## Challenge: Ecto — Orders Migration ### FINDING 1: Table lock risk (HIGH) AddColumn on `orders` (2.1M rows) will lock table during deploy. **Proof needed**: Run `SELECT count(*) FROM orders` — if >1M, use `ALTER TABLE ... ADD COLUMN ... DEFAULT NULL` (no lock). ### FINDING 2: Missing index (MEDIUM) New `WHERE status = ?` query on line 45 has no index. **Action**: Add `create index(:orders, [:status])` to migration. ### Status: BLOCKED — 2 unresolved findings
Run /phx:challenge [mode] to initiate a rigorous review. The reviewer will not approve until all concerns are addressed with evidence.
Example workflow:
/phx:challenge ecto after migration changes| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 21,952 | 5,435 | -75% | 1 | 1 | 0% | 3,591 | 1,538 | -57% | 0 | 0 | — |
case-02 | fail→fail | 12,384 | 8,146 | -34% | 1 | 1 | 0% | 1,828 | 1,784 | -2% | 0 | 0 | — |
case-03 | fail→fail | 9,529 | 6,531 | -31% | 1 | 1 | 0% | 1,381 | 1,755 | +27% | 0 | 0 | — |
case-04 | pass→pass | 14,116 | 14,283 | +1% | 1 | 1 | 0% | 2,348 | 3,404 | +45% | 0 | 0 | — |
case-05 | pass→pass | 11,548 | 10,119 | -12% | 1 | 1 | 0% | 1,992 | 2,764 | +39% | 0 | 0 | — |
case-06 | fail→pass | 8,792 | 11,181 | +27% | 1 | 1 | 0% | 1,418 | 2,790 | +97% | 0 | 0 | — |
case-07 | pass→pass | 16,061 | 9,919 | -38% | 1 | 1 | 0% | 2,563 | 2,592 | +1% | 0 | 0 | — |
case-08 | pass→pass | 16,942 | 13,004 | -23% | 1 | 1 | 0% | 2,728 | 3,139 | +15% | 0 | 0 | — |
case-09 | pass→pass | 16,298 | 18,070 | +11% | 1 | 1 | 0% | 2,688 | 4,122 | +53% | 0 | 0 | — |
case-10 | pass→pass | 13,150 | 10,953 | -17% | 1 | 1 | 0% | 1,951 | 2,792 | +43% | 0 | 0 | — |
case-11 | pass→pass | 13,100 | 11,931 | -9% | 1 | 1 | 0% | 2,237 | 2,987 | +34% | 0 | 0 | — |
case-12 | pass→pass | 11,796 | 12,382 | +5% | 1 | 1 | 0% | 1,824 | 3,032 | +66% | 0 | 0 | — |
case-13 | pass→pass | 9,488 | 6,362 | -33% | 1 | 1 | 0% | 1,652 | 2,145 | +30% | 0 | 0 | — |
case-14 | fail→pass | 12,712 | 13,230 | +4% | 1 | 1 | 0% | 1,919 | 3,114 | +62% | 0 | 0 | — |
case-15 | fail→fail | 14,270 | 16,463 | +15% | 1 | 1 | 0% | 2,291 | 3,639 | +59% | 0 | 0 | — |
case-16 | pass→pass | 12,868 | 14,710 | +14% | 1 | 1 | 0% | 1,202 | 2,724 | +127% | 0 | 0 | — |
case-17 | pass→pass | 3,748 | 2,232 | -40% | 1 | 1 | 0% | 594 | 1,476 | +148% | 0 | 0 | — |
case-23 | pass→pass | 10,076 | 9,499 | -6% | 1 | 1 | 0% | 1,880 | 2,982 | +59% | 0 | 0 | — |
case-18 | pass→pass | 5,843 | 4,149 | -29% | 1 | 1 | 0% | 959 | 1,759 | +83% | 0 | 0 | — |
case-19 | fail→pass | 21,288 | 16,618 | -22% | 1 | 1 | 0% | 3,022 | 3,587 | +19% | 0 | 0 | — |
case-20 | pass→pass | 15,481 | 13,567 | -12% | 1 | 1 | 0% | 2,316 | 3,264 | +41% | 0 | 0 | — |
case-21 | pass→pass | 4,564 | 6,002 | +32% | 1 | 1 | 0% | 988 | 2,216 | +124% | 0 | 0 | — |
case-22 | pass→pass | 4,076 | 4,354 | +7% | 1 | 1 | 0% | 752 | 1,989 | +164% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +13 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.