Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Build your MVP following the AGENTS.md plan. Use when the user wants to start building, implement features, or says "build my MVP", "start coding", or "implement the project".
.claude/skills/khazp-vibe-build/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | -67% | 0% |
| case-20 | ✓→✗ | ▼ Worse | -70% | 0% |
| case-07 | ✓→✗ | ▼ Worse | -77% | 0% |
| case-10 | ✓→✗ | ▼ Worse | -71% | 0% |
| case-11 | ✓→✗ | ▼ Worse | -87% | 0% |
Read AGENTS.md, MEMORY.md, the manifest's product documents, and relevant agent_docs. Establish acceptance criteria for one usable slice and implement within existing authorization. Preserve uncommitted work and record an actual recovery checkpoint before risky changes; never fabricate a commit or backup.
Start with a functioning screen or equivalent observable output. Add authentication, a database, infrastructure, and AI only when requirements justify them. Run the project's applicable checks after reviewing commands; doctor is setup validation only. Use ../vibe-verify/SKILL.md for the actual journey. For AI features also check failure behavior, data boundaries, and permission denial where relevant.
Update current progress and next steps in MEMORY.md; stable rules remain in AGENTS.md. Report Changed, Checked with commands/results, Not checked, Next decision, and Recovery. Do not treat a passing build as proof of working behavior.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 14,356 | 13,793 | -4% | 1 | 1 | 0% | 186 | 342 | +84% | 0 | 0 | — |
case-20 | pass→fail | 13,921 | 15,788 | +13% | 1 | 1 | 0% | 1,555 | 459 | -70% | 0 | 0 | — |
case-02 | fail→fail | 51,960 | 14,533 | -72% | 1 | 1 | 0% | 8,352 | 364 | -96% | 0 | 0 | — |
case-03 | fail→fail | 14,818 | 15,326 | +3% | 1 | 1 | 0% | 170 | 421 | +148% | 0 | 0 | — |
case-04 | fail→fail | 9,956 | 14,630 | +47% | 1 | 1 | 0% | 884 | 434 | -51% | 0 | 0 | — |
case-05 | fail→fail | 7,491 | 15,508 | +107% | 1 | 1 | 0% | 252 | 473 | +88% | 0 | 0 | — |
case-06 | fail→fail | 13,498 | 14,937 | +11% | 1 | 1 | 0% | 1,532 | 387 | -75% | 0 | 0 | — |
case-07 | pass→fail | 17,179 | 15,821 | -8% | 1 | 1 | 0% | 2,148 | 484 | -77% | 0 | 0 | — |
case-08 | pass→pass | 16,362 | 18,177 | +11% | 1 | 1 | 0% | 1,744 | 1,159 | -34% | 0 | 0 | — |
case-09 | fail→fail | 20,805 | 13,289 | -36% | 1 | 1 | 0% | 2,568 | 369 | -86% | 0 | 0 | — |
case-10 | pass→fail | 15,062 | 15,908 | +6% | 1 | 1 | 0% | 1,656 | 487 | -71% | 0 | 0 | — |
case-11 | pass→fail | 23,981 | 15,646 | -35% | 1 | 1 | 0% | 3,212 | 413 | -87% | 0 | 0 | — |
case-12 | fail→fail | 15,876 | 14,150 | -11% | 1 | 1 | 0% | 1,816 | 495 | -73% | 0 | 0 | — |
case-13 | pass→pass | 9,394 | 6,870 | -27% | 1 | 1 | 0% | 818 | 518 | -37% | 0 | 0 | — |
case-14 | fail→fail | 19,318 | 15,002 | -22% | 1 | 1 | 0% | 2,001 | 397 | -80% | 0 | 0 | — |
case-15 | fail→pass | 13,506 | 7,226 | -46% | 1 | 1 | 0% | 1,394 | 454 | -67% | 0 | 0 | — |
case-16 | pass→pass | 13,887 | 18,214 | +31% | 1 | 1 | 0% | 1,455 | 825 | -43% | 0 | 0 | — |
case-17 | pass→fail | 8,360 | 16,149 | +93% | 1 | 1 | 0% | 1,334 | 518 | -61% | 0 | 0 | — |
case-18 | pass→pass | 15,418 | 17,676 | +15% | 1 | 1 | 0% | 1,610 | 911 | -43% | 0 | 0 | — |
case-19 | pass→fail | 23,998 | 16,006 | -33% | 1 | 1 | 0% | 2,892 | 338 | -88% | 0 | 0 | — |
case-21 | fail→fail | 13,109 | 14,068 | +7% | 1 | 1 | 0% | 1,270 | 506 | -60% | 0 | 0 | — |
case-22 | fail→fail | 16,475 | 14,723 | -11% | 1 | 1 | 0% | 1,810 | 426 | -76% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 5 counted toward the lift figure. The other 17 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -23 percentage points is the difference between those two pass rates over the 5 comparable cases. 10 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/26/2026 | +5% |
| gemini-3.6-flash | verified | 8/12/2026 | +14% |
Other measured skills in the registry, with their headline benchmark lift.