Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Turn a one-line objective into a step-by-step construction plan any coding agent can execute cold. Each step has a self-contained context brief — a fresh agent in a new session can pick up any step without reading prior steps.
.claude/skills/lingxling-blueprint/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -22% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 135% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 82% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 62% | 0% |
Turn a one-line objective into a step-by-step plan any coding agent can execute cold.
Blueprint is for multi-session, multi-agent engineering projects where each step must be independently executable by a fresh agent that has never seen the conversation history. Install it once, invoke it with /blueprint <project> <objective>.
/blueprint myapp "migrate database to PostgreSQL"/blueprint antbot "extract providers into plugins"bashmkdir -p ~/.claude/skills git clone https://github.com/antbotlab/blueprint.git ~/.claude/skills/blueprint
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-08 | fail→fail | 13,667 | 12,164 | -11% | 1 | 1 | 0% | 2,171 | 2,630 | +21% | 0 | 0 | — |
case-01 | fail→fail | 51,194 | 36,677 | -28% | 1 | 1 | 0% | 8,232 | 5,720 | -31% | 0 | 0 | — |
case-02 | fail→pass | 54,803 | 29,311 | -47% | 1 | 1 | 0% | 7,061 | 5,532 | -22% | 0 | 0 | — |
case-03 | fail→fail | 30,777 | 27,889 | -9% | 1 | 1 | 0% | 4,229 | 4,890 | +16% | 0 | 0 | — |
case-04 | pass→fail | 4,208 | 3,156 | -25% | 1 | 1 | 0% | 534 | 1,102 | +106% | 0 | 0 | — |
case-05 | pass→fail | 6,848 | 3,528 | -48% | 1 | 1 | 0% | 832 | 1,167 | +40% | 0 | 0 | — |
case-06 | pass→pass | 24,399 | 6,780 | -72% | 1 | 1 | 0% | 1,076 | 1,974 | +83% | 0 | 0 | — |
case-07 | fail→fail | 5,413 | 7,584 | +40% | 1 | 1 | 0% | 928 | 1,799 | +94% | 0 | 0 | — |
case-09 | fail→fail | 26,337 | 24,428 | -7% | 1 | 1 | 0% | 3,702 | 4,033 | +9% | 0 | 0 | — |
case-10 | fail→pass | 16,224 | 14,127 | -13% | 1 | 1 | 0% | 2,195 | 3,152 | +44% | 0 | 0 | — |
case-11 | fail→fail | 25,836 | 22,028 | -15% | 1 | 1 | 0% | 2,932 | 4,054 | +38% | 0 | 0 | — |
case-12 | fail→fail | 35,028 | 38,813 | +11% | 1 | 1 | 0% | 3,071 | 6,730 | +119% | 0 | 0 | — |
case-13 | fail→fail | 31,637 | 24,258 | -23% | 1 | 1 | 0% | 3,297 | 4,760 | +44% | 0 | 0 | — |
case-14 | fail→pass | 25,748 | 49,995 | +94% | 1 | 1 | 0% | 3,717 | 8,717 | +135% | 0 | 0 | — |
case-15 | fail→fail | 17,919 | 25,298 | +41% | 1 | 1 | 0% | 2,935 | 4,062 | +38% | 0 | 0 | — |
case-16 | fail→pass | 23,862 | 28,301 | +19% | 1 | 1 | 0% | 3,015 | 5,486 | +82% | 0 | 0 | — |
case-17 | fail→pass | 25,889 | 39,454 | +52% | 1 | 1 | 0% | 3,256 | 5,271 | +62% | 0 | 0 | — |
case-18 | fail→pass | 24,832 | 32,991 | +33% | 1 | 1 | 0% | 3,393 | 5,332 | +57% | 0 | 0 | — |
case-19 | fail→fail | 22,816 | 104,364 | +357% | 1 | 1 | 0% | 3,693 | 4,906 | +33% | 0 | 0 | — |
case-20 | fail→fail | 19,179 | 28,001 | +46% | 1 | 1 | 0% | 3,014 | 5,529 | +83% | 0 | 0 | — |
case-21 | fail→pass | 36,834 | 32,687 | -11% | 1 | 1 | 0% | 4,507 | 4,629 | +3% | 0 | 0 | — |
case-22 | fail→pass | 19,710 | 25,257 | +28% | 1 | 1 | 0% | 3,122 | 5,050 | +62% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +27 percentage points is the difference between those two pass rates over the 22 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.