Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Figure out why the loaf came out dense, flat, or gummy — and what to change next bake. Use when asked why is my sourdough [dense/flat/gummy/not rising], my starter isn't bubbling, help fix my bread, or troubleshoot my sourdough. Produces a likely-cause diagnosis from your symptoms and process, the specific fix for the next bake, a starter-health check, and a simple timing/temperature adjustment — no dogma, just the variable that's actually off.
.claude/skills/mohitagw15856-sourdough-troubleshooter/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -8% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -3% | 0% |
| case-21 | ✗→✓ | ▲ Improved | 4% | 0% |
| case-10 | ✓→✓ | = Same ✓ | -7% | 0% |
Sourdough fails for a short list of reasons — an underpowered starter, under- or over-proofing, weak shaping, or temperature. This reads your symptoms and your process, names the single variable most likely at fault, and tells you what to change next time, instead of drowning you in conflicting internet advice.
Ask for these if not provided:
Most likely cause: diagnosis] — because symptom + process detail]. Change this next bake: one specific fix]. Keep the same: so the test is clean].
Starter check: strong / needs building — do X]. Proofing tweak: more/less time for your kitchen temp; how to judge doneness].
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-10 | pass→pass | 19,660 | 14,609 | -26% | 1 | 1 | 0% | 2,389 | 2,222 | -7% | 0 | 0 | — |
case-01 | fail→fail | 19,699 | 17,831 | -9% | 1 | 1 | 0% | 2,581 | 3,053 | +18% | 0 | 0 | — |
case-02 | pass→pass | 16,578 | 8,199 | -51% | 1 | 1 | 0% | 2,076 | 2,365 | +14% | 0 | 0 | — |
case-03 | fail→pass | 15,252 | 15,690 | +3% | 1 | 1 | 0% | 1,910 | 2,949 | +54% | 0 | 0 | — |
case-04 | pass→pass | 15,180 | 22,791 | +50% | 1 | 1 | 0% | 2,994 | 4,131 | +38% | 0 | 0 | — |
case-05 | pass→pass | 15,332 | 13,034 | -15% | 1 | 1 | 0% | 2,645 | 2,912 | +10% | 0 | 0 | — |
case-06 | pass→pass | 17,757 | 15,343 | -14% | 1 | 1 | 0% | 2,065 | 2,389 | +16% | 0 | 0 | — |
case-07 | fail→fail | 19,700 | 18,514 | -6% | 1 | 1 | 0% | 2,345 | 2,944 | +26% | 0 | 0 | — |
case-08 | pass→pass | 18,344 | 8,455 | -54% | 1 | 1 | 0% | 2,145 | 2,173 | +1% | 0 | 0 | — |
case-09 | pass→pass | 10,991 | 9,800 | -11% | 1 | 1 | 0% | 1,912 | 2,481 | +30% | 0 | 0 | — |
case-11 | pass→pass | 19,633 | 12,019 | -39% | 1 | 1 | 0% | 2,431 | 2,011 | -17% | 0 | 0 | — |
case-12 | fail→pass | 15,901 | 9,402 | -41% | 1 | 1 | 0% | 2,541 | 2,344 | -8% | 0 | 0 | — |
case-13 | pass→pass | 12,716 | 9,054 | -29% | 1 | 1 | 0% | 2,011 | 2,218 | +10% | 0 | 0 | — |
case-14 | pass→pass | 16,055 | 14,010 | -13% | 1 | 1 | 0% | 2,638 | 2,442 | -7% | 0 | 0 | — |
case-15 | pass→pass | 19,319 | 7,508 | -61% | 1 | 1 | 0% | 2,282 | 2,176 | -5% | 0 | 0 | — |
case-16 | fail→pass | 18,999 | 8,410 | -56% | 1 | 1 | 0% | 2,256 | 2,195 | -3% | 0 | 0 | — |
case-17 | pass→pass | 11,677 | 14,392 | +23% | 1 | 1 | 0% | 2,061 | 2,444 | +19% | 0 | 0 | — |
case-18 | pass→pass | 19,285 | 10,438 | -46% | 1 | 1 | 0% | 2,338 | 2,353 | +1% | 0 | 0 | — |
case-19 | pass→pass | 17,772 | 12,795 | -28% | 1 | 1 | 0% | 1,983 | 2,150 | +8% | 0 | 0 | — |
case-20 | pass→pass | 21,417 | 14,294 | -33% | 1 | 1 | 0% | 2,641 | 2,279 | -14% | 0 | 0 | — |
case-21 | fail→pass | 15,265 | 14,359 | -6% | 1 | 1 | 0% | 2,242 | 2,323 | +4% | 0 | 0 | — |
case-22 | pass→pass | 20,593 | 14,074 | -32% | 1 | 1 | 0% | 2,670 | 2,390 | -10% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.