Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Find what a change could break somewhere else before it ships, beyond the diff, and prove the one fact it's safe because of by running real code instead of writing it up. Use for 'blast radius of X', 'what could this break', or reviewing a small diff you don't trust.
.claude/skills/kunanonj-cursor-plugin-pstack-blast-radius/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -30% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 43% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-20 | ✗→✓ | ▲ Improved | 31% | 0% |
Find what a change breaks somewhere else, before it ships. Use for "blast radius of X", "what could this break", or reviewing a small diff you don't trust yet.
Companion to how and why. how tells you what the code does. why tells you why it's shaped that way. Blast radius tells you what it breaks somewhere else.
Listing the callers is not the job. The agent can grep those in a second. The job is the breakage grep won't show you.
A blast-radius writeup that sounds right is worthless. It reads as convincing whether or not it's true, and that is the trap you are walking into. So don't hand back the writeup. Find the one or two facts the whole thing depends on and prove them by running code. Words are where you start, not what you ship.
For each fact the change's safety depends on, get it as far down this list as is cheap, and say where it stopped.
file:line, or the library's own source.Any safety fact you can't get to step 4, say so out loud. Don't write it up as settled. Step 4 is usually one small script that imports the same library the app ships and calls the exact function you're worried about.
why step 2 to pull the PR and commits.why. Cite a real file:line, a search that finds nothing is still an answer, and never make up a caller or an API.arena. Ask several models the same question and merge the answers. Different models catch different real bugs.file:line, how likely and how bad, and how to check. Paste the proof for the ones that matter.Write it through unslop, cite real code, and strip anything private before it goes anywhere public.
Reply: the writeup above, with the one safety fact either proven or marked unproven.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→pass | 13,039 | 16,730 | +28% | 1 | 1 | 0% | 2,469 | 4,106 | +66% | 0 | 0 | — |
case-01 | fail→fail | 26,479 | 23,277 | -12% | 1 | 1 | 0% | 2,741 | 3,739 | +36% | 0 | 0 | — |
case-02 | fail→pass | 24,800 | 13,257 | -47% | 1 | 1 | 0% | 4,759 | 3,309 | -30% | 0 | 0 | — |
case-03 | fail→fail | 14,588 | 15,612 | +7% | 1 | 1 | 0% | 2,803 | 3,447 | +23% | 0 | 0 | — |
case-05 | pass→pass | 14,287 | 16,288 | +14% | 1 | 1 | 0% | 2,335 | 3,642 | +56% | 0 | 0 | — |
case-06 | pass→fail | 6,305 | 10,907 | +73% | 1 | 1 | 0% | 1,254 | 2,713 | +116% | 0 | 0 | — |
case-07 | fail→fail | 16,935 | 22,384 | +32% | 1 | 1 | 0% | 2,945 | 4,765 | +62% | 0 | 0 | — |
case-08 | fail→fail | 17,291 | 16,282 | -6% | 1 | 1 | 0% | 2,916 | 3,932 | +35% | 0 | 0 | — |
case-09 | fail→pass | 13,250 | 14,349 | +8% | 1 | 1 | 0% | 2,363 | 3,380 | +43% | 0 | 0 | — |
case-10 | fail→pass | 16,231 | 19,206 | +18% | 1 | 1 | 0% | 2,799 | 4,202 | +50% | 0 | 0 | — |
case-11 | pass→pass | 15,783 | 28,353 | +80% | 1 | 1 | 0% | 2,897 | 5,585 | +93% | 0 | 0 | — |
case-12 | fail→fail | 27,567 | 3,773 | -86% | 1 | 1 | 0% | 2,220 | 1,174 | -47% | 0 | 0 | — |
case-13 | fail→fail | 19,686 | 19,385 | -2% | 1 | 1 | 0% | 3,247 | 4,135 | +27% | 0 | 0 | — |
case-14 | fail→fail | 15,776 | 18,348 | +16% | 1 | 1 | 0% | 2,693 | 4,294 | +59% | 0 | 0 | — |
case-15 | fail→fail | 13,695 | 20,821 | +52% | 1 | 1 | 0% | 2,490 | 4,171 | +68% | 0 | 0 | — |
case-16 | fail→pass | 14,359 | 15,962 | +11% | 1 | 1 | 0% | 2,424 | 3,744 | +54% | 0 | 0 | — |
case-17 | pass→pass | 17,258 | 15,148 | -12% | 1 | 1 | 0% | 2,982 | 3,596 | +21% | 0 | 0 | — |
case-18 | fail→fail | 14,381 | 5,468 | -62% | 1 | 1 | 0% | 2,437 | 1,217 | -50% | 0 | 0 | — |
case-19 | fail→fail | 17,285 | 4,976 | -71% | 1 | 1 | 0% | 3,036 | 1,149 | -62% | 0 | 0 | — |
case-20 | fail→pass | 18,281 | 20,639 | +13% | 1 | 1 | 0% | 3,301 | 4,315 | +31% | 0 | 0 | — |
case-21 | fail→pass | 16,863 | 31,379 | +86% | 1 | 1 | 0% | 2,846 | 6,344 | +123% | 0 | 0 | — |
case-22 | fail→pass | 16,129 | 10,942 | -32% | 1 | 1 | 0% | 2,845 | 2,905 | +2% | 0 | 0 | — |
case-23 | fail→pass | 18,187 | 24,494 | +35% | 1 | 1 | 0% | 3,157 | 4,841 | +53% | 0 | 0 | — |
case-24 | fail→pass | 15,490 | 18,256 | +18% | 1 | 1 | 0% | 2,677 | 3,669 | +37% | 0 | 0 | — |
case-25 | fail→fail | 17,576 | 24,262 | +38% | 1 | 1 | 0% | 2,909 | 5,012 | +72% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 25 cases were attempted, and 22 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +32 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.