Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Configure which models pstack uses per role. Detects your available models and writes an always-applied rule that overrides the skill defaults. Use for /setup-pstack, \"configure pstack models\", or changing pstack's model choices.
.claude/skills/kunanonj-cursor-plugin-pstack-setup-pstack/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-10 | ✗→✓ | ▲ Improved | 7% | 0% |
| case-14 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-19 | ✗→✓ | ▲ Improved | -35% | 0% |
| case-04 | ✓→✗ | ▼ Worse | -6% | 0% |
| case-11 | ✓→✗ | ▼ Worse | 166% | 0% |
Write ~/.cursor/rules/pstack-models.mdc, an always-applied rule that sets pstack's model per role. The skills read it and fall back to their inline defaults when a line is absent, so this is an override layer, not a requirement.
Enumerate the model slugs you can pass to a Task subagent in this session; that is the dependable source. If Cursor also exposes a models API or CLI that lists the user's entitled models, prefer it for completeness. If you cannot detect any, ask the user to paste the slugs they have access to. Never write a slug you have not confirmed is available.
The default role-to-model mapping is the rule shape shown in step 5 below. If ~/.cursor/rules/pstack-models.mdc already exists, read it and treat its values as the current choices. Otherwise start from those defaults.
Show every role with its current model, marking any whose model is not in the detected set as needing a choice. Ask whether to accept as-is or change specific roles, offering the detected models as the options. Prefer AskQuestion over free text. For panel roles (how critics, arena runners, architect runners, interrogate reviewers) the value is a list, and one subagent runs per model, so the list length sets the count.
Every slug written must be in the detected set. If a chosen slug is not available, stop and ask again. A rule pointing at a model the user cannot use breaks every delegation that reads it.
Write ~/.cursor/rules/pstack-models.mdc with alwaysApply: true and one line per role, using the same labels poteto-mode uses. Overwrite the whole file so re-runs stay idempotent. Shape:
---
description: pstack per-role model choices (overrides skill defaults)
alwaysApply: true
---
# pstack model configuration. One line per role. Delete a line to fall back to the skill default.
feature, refactoring: composer-2.5-fast
bug-fix: gpt-5.5-high-fast
perf-issue: gpt-5.5-high-fast
hillclimb: gpt-5.5-high-fast
judgment and prose: claude-opus-4-8-thinking-xhigh
how explorer: composer-2.5-fast
how explainer: claude-opus-4-8-thinking-xhigh
how critics: claude-opus-4-8-thinking-xhigh, gpt-5.5-high-fast, composer-2.5-fast
why investigators: composer-2.5-fast
why synthesizer: claude-opus-4-8-thinking-xhigh
reflect tooling: composer-2.5-fast
reflect judgment, divergent, synthesizer: claude-opus-4-8-thinking-xhigh
arena runners: claude-opus-4-8-thinking-xhigh, gpt-5.5-high-fast, composer-2.5-fast
architect runners: claude-opus-4-8-thinking-xhigh, gpt-5.5-high-fast, composer-2.5-fast
interrogate reviewers: claude-opus-4-8-thinking-xhigh, gpt-5.5-high-fast, composer-2.5-fastTell the user the rule was written and that it applies to new sessions. Re-running this skill updates it.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→fail | 5,389 | 4,997 | -7% | 1 | 1 | 0% | 380 | 1,106 | +191% | 0 | 0 | — |
case-01 | fail→fail | 3,526 | 5,473 | +55% | 1 | 1 | 0% | 434 | 1,096 | +153% | 0 | 0 | — |
case-02 | fail→fail | 12,423 | 5,428 | -56% | 1 | 1 | 0% | 2,550 | 1,173 | -54% | 0 | 0 | — |
case-04 | pass→fail | 8,093 | 6,454 | -20% | 1 | 1 | 0% | 1,341 | 1,260 | -6% | 0 | 0 | — |
case-05 | pass→pass | 7,953 | 17,422 | +119% | 1 | 1 | 0% | 1,647 | 4,045 | +146% | 0 | 0 | — |
case-06 | pass→pass | 5,213 | 25,162 | +383% | 1 | 1 | 0% | 837 | 3,983 | +376% | 0 | 0 | — |
case-07 | fail→fail | 7,071 | 10,439 | +48% | 1 | 1 | 0% | 1,309 | 1,504 | +15% | 0 | 0 | — |
case-08 | fail→fail | 4,332 | 3,262 | -25% | 1 | 1 | 0% | 768 | 1,410 | +84% | 0 | 0 | — |
case-09 | pass→pass | 7,129 | 10,661 | +50% | 1 | 1 | 0% | 1,066 | 2,963 | +178% | 0 | 0 | — |
case-10 | fail→pass | 8,064 | 3,983 | -51% | 1 | 1 | 0% | 1,562 | 1,676 | +7% | 0 | 0 | — |
case-11 | pass→fail | 5,672 | 7,733 | +36% | 1 | 1 | 0% | 941 | 2,505 | +166% | 0 | 0 | — |
case-12 | pass→pass | 3,267 | 1,412 | -57% | 1 | 1 | 0% | 560 | 1,085 | +94% | 0 | 0 | — |
case-13 | pass→pass | 7,601 | 4,856 | -36% | 1 | 1 | 0% | 1,370 | 1,781 | +30% | 0 | 0 | — |
case-14 | fail→pass | 9,442 | 8,444 | -11% | 1 | 1 | 0% | 1,728 | 2,325 | +35% | 0 | 0 | — |
case-15 | pass→pass | 5,755 | 3,163 | -45% | 1 | 1 | 0% | 1,079 | 1,570 | +46% | 0 | 0 | — |
case-16 | pass→pass | 5,116 | 2,107 | -59% | 1 | 1 | 0% | 819 | 1,158 | +41% | 0 | 0 | — |
case-17 | pass→pass | 4,713 | 1,849 | -61% | 1 | 1 | 0% | 938 | 1,149 | +22% | 0 | 0 | — |
case-18 | pass→pass | 5,636 | 2,522 | -55% | 1 | 1 | 0% | 941 | 1,252 | +33% | 0 | 0 | — |
case-19 | fail→pass | 9,477 | 1,351 | -86% | 1 | 1 | 0% | 1,577 | 1,019 | -35% | 0 | 0 | — |
case-20 | fail→fail | 7,263 | 5,074 | -30% | 1 | 1 | 0% | 1,242 | 1,581 | +27% | 0 | 0 | — |
case-21 | pass→pass | 4,470 | 2,993 | -33% | 1 | 1 | 0% | 827 | 1,405 | +70% | 0 | 0 | — |
case-22 | pass→pass | 8,272 | 5,153 | -38% | 1 | 1 | 0% | 1,428 | 1,702 | +19% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 17 counted toward the lift figure. The other 5 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +5 percentage points is the difference between those two pass rates over the 17 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.