Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Try to bootstrap and start a repository like a cold agent, then report where the path breaks down
.claude/skills/kunanonj-cursor-plugin-agent-compat-agent-startup-review/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -27% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 2% | 0% |
| case-03 | ✗→✓ | ▲ Improved | -26% | 0% |
| case-04 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -30% | 0% |
Tries the cold-start path and reports how much work it takes to get the repo running.
Use when the user wants to know whether a repo is actually easy to start, not just whether it claims to be.
README, scripts, toolchain files, env examples, and workflow docs.93/100 if the main startup path works inside the time budget, even if it needs ordinary local prerequisites such as Docker or a database.84/100 if the repo starts, but only after some digging, a recovery step, or heavier setup than the docs suggest.68/100 if a startup path probably exists but stays too manual, too ambiguous, or too expensive for normal agent use.27/100 if you cannot get a credible startup path working from the repo and docs you have.12/100 if the path is blocked on secrets, accounts, or infrastructure you cannot reasonably access.82, 85, or 91 over a multiple of ten when that is the more honest read.Reply in plain text only (no markdown fences, no # headings, no emphasis syntax). Use this layout:
First line: Startup Compatibility Score: <score>/100
Then a short summary paragraph.
Then the line Problems followed by one bullet per line using - .
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 15,246 | 4,587 | -70% | 1 | 1 | 0% | 2,314 | 1,686 | -27% | 0 | 0 | — |
case-02 | fail→pass | 7,300 | 3,294 | -55% | 1 | 1 | 0% | 1,308 | 1,331 | +2% | 0 | 0 | — |
case-03 | fail→pass | 14,115 | 4,652 | -67% | 1 | 1 | 0% | 2,347 | 1,738 | -26% | 0 | 0 | — |
case-04 | fail→pass | 14,563 | 4,383 | -70% | 1 | 1 | 0% | 2,587 | 1,655 | -36% | 0 | 0 | — |
case-05 | fail→pass | 13,921 | 5,213 | -63% | 1 | 1 | 0% | 2,497 | 1,759 | -30% | 0 | 0 | — |
case-06 | fail→pass | 24,694 | 15,642 | -37% | 1 | 1 | 0% | 1,121 | 1,507 | +34% | 0 | 0 | — |
case-07 | fail→pass | 6,785 | 4,009 | -41% | 1 | 1 | 0% | 1,164 | 1,486 | +28% | 0 | 0 | — |
case-08 | fail→pass | 9,325 | 4,268 | -54% | 1 | 1 | 0% | 1,884 | 1,565 | -17% | 0 | 0 | — |
case-09 | fail→pass | 6,690 | 4,339 | -35% | 1 | 1 | 0% | 1,234 | 1,596 | +29% | 0 | 0 | — |
case-10 | fail→pass | 5,289 | 4,513 | -15% | 1 | 1 | 0% | 1,137 | 1,601 | +41% | 0 | 0 | — |
case-11 | fail→pass | 10,633 | 5,155 | -52% | 1 | 1 | 0% | 1,933 | 1,675 | -13% | 0 | 0 | — |
case-12 | fail→pass | 15,983 | 4,112 | -74% | 1 | 1 | 0% | 2,499 | 1,473 | -41% | 0 | 0 | — |
case-13 | fail→pass | 7,884 | 3,169 | -60% | 1 | 1 | 0% | 1,538 | 1,327 | -14% | 0 | 0 | — |
case-14 | fail→pass | 9,229 | 4,174 | -55% | 1 | 1 | 0% | 1,751 | 1,519 | -13% | 0 | 0 | — |
case-15 | fail→pass | 11,257 | 3,604 | -68% | 1 | 1 | 0% | 2,369 | 1,396 | -41% | 0 | 0 | — |
case-16 | fail→pass | 14,331 | 4,617 | -68% | 1 | 1 | 0% | 2,568 | 1,644 | -36% | 0 | 0 | — |
case-17 | fail→pass | 12,927 | 5,921 | -54% | 1 | 1 | 0% | 2,263 | 1,850 | -18% | 0 | 0 | — |
case-18 | fail→fail | 12,343 | 4,792 | -61% | 1 | 1 | 0% | 2,276 | 1,668 | -27% | 0 | 0 | — |
case-19 | fail→fail | 8,502 | 2,932 | -66% | 1 | 1 | 0% | 1,597 | 1,298 | -19% | 0 | 0 | — |
case-20 | pass→fail | 11,892 | 11,708 | -2% | 1 | 1 | 0% | 1,861 | 3,219 | +73% | 0 | 0 | — |
case-21 | pass→pass | 6,052 | 5,838 | -4% | 1 | 1 | 0% | 1,443 | 2,064 | +43% | 0 | 0 | — |
case-22 | pass→fail | 9,215 | 5,872 | -36% | 1 | 1 | 0% | 2,146 | 1,947 | -9% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +68 percentage points is the difference between those two pass rates over the 21 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.