Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Apply an advisory edit scope in this session. Use for read-only or directory-scoped work; no enforcement hook is installed.
.claude/skills/oliver-kriska-phx-freeze/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 10% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 37% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 49% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -60% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -38% | 0% |
Apply a current-session instruction that limits which files this agent may edit. This generated runtime does not install an enforcement hook, so the scope is advisory rather than a technical lock. Never claim that edits are blocked by the runtime.
text/skill:phx-freeze /skill:phx-freeze lib/app_web priv/repo /skill:phx-freeze status /skill:phx-freeze off
Treat the text after the skill invocation as follows:
| Invocation | Current-session behavior | |---|---| | No arguments | Do not edit files; investigation and reporting remain read-only. | | Path prefixes | Edit only files under the listed project-relative prefixes. | | status | Report the advisory scope currently established in this conversation. | | off | Clear the advisory scope for subsequent work. |
Do not create .claude/.freeze. That sentinel belongs to the canonical Claude Code plugin and could affect a later Claude Code session even though this runtime cannot enforce or clear it reliably.
the current agent, not a runtime or security boundary.
.claude/.freeze in a generated runtime — nomatching enforcement component is installed here.
ask before editing outside listed prefixes.
lib/foo includes that directory and itsdescendants, not a sibling such as lib/foobar.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 3,902 | 1,667 | -57% | 1 | 1 | 0% | 590 | 649 | +10% | 0 | 0 | — |
case-02 | pass→pass | 9,793 | 2,035 | -79% | 1 | 1 | 0% | 1,543 | 673 | -56% | 0 | 0 | — |
case-03 | fail→pass | 3,928 | 2,927 | -25% | 1 | 1 | 0% | 581 | 798 | +37% | 0 | 0 | — |
case-04 | fail→pass | 2,959 | 1,646 | -44% | 1 | 1 | 0% | 410 | 611 | +49% | 0 | 0 | — |
case-05 | pass→pass | 1,307 | 1,868 | +43% | 1 | 1 | 0% | 182 | 636 | +249% | 0 | 0 | — |
case-06 | fail→pass | 13,058 | 3,035 | -77% | 1 | 1 | 0% | 2,141 | 858 | -60% | 0 | 0 | — |
case-07 | fail→pass | 8,952 | 3,098 | -65% | 1 | 1 | 0% | 1,423 | 884 | -38% | 0 | 0 | — |
case-08 | pass→pass | 3,844 | 2,900 | -25% | 1 | 1 | 0% | 577 | 797 | +38% | 0 | 0 | — |
case-09 | pass→pass | 7,642 | 2,664 | -65% | 1 | 1 | 0% | 1,229 | 717 | -42% | 0 | 0 | — |
case-10 | fail→pass | 3,105 | 3,265 | +5% | 1 | 1 | 0% | 510 | 899 | +76% | 0 | 0 | — |
case-11 | fail→pass | 11,876 | 3,030 | -74% | 1 | 1 | 0% | 1,946 | 774 | -60% | 0 | 0 | — |
case-12 | pass→pass | 5,721 | 1,935 | -66% | 1 | 1 | 0% | 964 | 672 | -30% | 0 | 0 | — |
case-13 | pass→pass | 8,208 | 1,952 | -76% | 1 | 1 | 0% | 1,157 | 667 | -42% | 0 | 0 | — |
case-14 | fail→pass | 9,537 | 2,216 | -77% | 1 | 1 | 0% | 1,518 | 645 | -58% | 0 | 0 | — |
case-15 | fail→pass | 10,002 | 3,368 | -66% | 1 | 1 | 0% | 1,488 | 820 | -45% | 0 | 0 | — |
case-16 | fail→pass | 10,905 | 2,164 | -80% | 1 | 1 | 0% | 1,651 | 682 | -59% | 0 | 0 | — |
case-17 | pass→pass | 5,464 | 2,050 | -62% | 1 | 1 | 0% | 864 | 645 | -25% | 0 | 0 | — |
case-18 | fail→pass | 8,775 | 3,875 | -56% | 1 | 1 | 0% | 1,424 | 906 | -36% | 0 | 0 | — |
case-19 | pass→pass | 8,721 | 2,025 | -77% | 1 | 1 | 0% | 1,706 | 619 | -64% | 0 | 0 | — |
case-20 | pass→pass | 5,364 | 4,015 | -25% | 1 | 1 | 0% | 968 | 1,072 | +11% | 0 | 0 | — |
case-21 | pass→pass | 5,027 | 2,736 | -46% | 1 | 1 | 0% | 838 | 768 | -8% | 0 | 0 | — |
case-22 | pass→pass | 5,195 | 4,449 | -14% | 1 | 1 | 0% | 860 | 1,097 | +28% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +50 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.