Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Recommends the skill or flow that fits the user's situation. A router over the skills in this repo. Use when the user asks which skill to use, is unsure what fits, or needs a pointer to the right workflow.
.claude/skills/fradser-ask-matt/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 50% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 91% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 80% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 60% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 79% | 0% |
You don't remember every skill, so ask.
A flow is a path through the skills. Most paths run along one main flow, and two on-ramps merge onto it. Everything else is standalone, or a vocabulary layer that runs underneath.
The route most work travels. You have an idea and want it built.
/mattpocock:grill-with-docs sharpens the idea by interview. Start here when you have a codebase: it's stateful, retaining what it learns in CONTEXT.md and ADRs. (No codebase? Use /mattpocock:grill-me instead, covered under Standalone. Both run the same /mattpocock:grilling primitive; grill-with-docs is the one that leaves a paper trail.)/mattpocock:handoff in both directions (see Crossing sessions):/mattpocock:handoff out, then open a fresh session against that file,/mattpocock:prototype to answer the question with throwaway code,/mattpocock:handoff back what you learned, and reference it from the original idea thread./mattpocock:to-spec (turn the thread into a spec), then /mattpocock:to-tickets to split it into tracer-bullet tickets, each declaring its blocking edges. On a local tracker that's one file per ticket under .scratch/<feature>/issues/, worked blockers-first by hand; on a real tracker the edges become native blocking links, so any ticket whose blockers are done can be grabbed: kick off /mattpocock:implement per ticket, clearing context between each one./mattpocock:implement right here, in the same context window.Either way, /mattpocock:implement builds each issue by driving /mattpocock:bdd internally (Discovery → Formulation, then Automation delegated to /mattpocock:tdd) (one vertical slice at a time), then closes out by running /mattpocock:code-review, a two-axis review (Standards + Spec) of the diff, before committing. Reach for /mattpocock:bdd on its own when you want to define Gherkin scenarios for a feature, and /mattpocock:tdd on its own when you want to write tests within the BDD Automation phase. Reach for /mattpocock:code-review on its own whenever you want to review a branch or PR against a fixed point.
Keep steps 1–3 in one unbroken context window (don't compact or clear until after /mattpocock:to-tickets) so the grilling, spec, and tickets all build on the same thinking. Each /mattpocock:implement then starts fresh, working from the ticket.
The limit on this is the smart zone: the window (~120k tokens on state-of-the-art models) within which the model still reasons sharply. If a session approaches it before /mattpocock:to-tickets, don't push on degraded; /mattpocock:handoff and continue in a fresh thread.
Wayfinder is only for work too big for one session — never a well-scoped feature. Triage is only for issues the user didn't create — never for the tickets /mattpocock:to-tickets produced. Keep steps 1–3 in one unbroken context window: no compact or clear until after /mattpocock:to-tickets.
A starting situation that generates work, then merges onto the main flow.
/mattpocock:triage. It moves issues through triage roles and produces agent-ready issues, which /mattpocock:implement later picks up.Triage is only for issues you didn't create: bug reports, incoming feature requests, anything that arrives raw. Tickets that /mattpocock:to-tickets produced are already agent-ready, so don't triage them.
/mattpocock:diagnosing-bugs. For the hard ones: the bug that resists a first glance, the intermittent flake, the regression that crept in between two known-good states. It refuses to theorise until it has a tight feedback loop (one command that already goes red on this bug), then fixes with a regression test. Its post-mortem hands off to /mattpocock:improve-codebase-architecture when the real finding is that there's no good seam to lock the bug down./mattpocock:wayfinder, the most cognitively demanding flow here. When the way from here to the destination isn't visible yet, it charts a shared map of decision tickets on the issue tracker and resolves them one at a time, producing decisions, not deliverables, until the fog is pushed back and the way is clear. Where /mattpocock:grill-with-docs sharpens an idea you can hold in one session, wayfinder is for the idea you can't, and it's slower and denser, so save it for exactly that, never a well-scoped feature.When the map clears, it hands off, it doesn't build: merge onto the main flow at /mattpocock:to-spec, which collapses the map's linked decisions into a buildable plan, then /mattpocock:to-tickets and /mattpocock:implement as usual. Looping the map straight into /mattpocock:implement skips that collapse and throws the linked detail away, so go straight to /mattpocock:implement only when the effort turned out genuinely small.
Not feature work, just upkeep.
/mattpocock:improve-codebase-architecture runs whenever you have a spare moment to keep the codebase good for agents to operate in. It surfaces deepening opportunities; picking one _generates an idea_ you can take into the main flow at /mattpocock:grill-with-docs. It's the survey that finds the candidates; /mattpocock:codebase-design (below) is the bench you design the chosen one on.Model-invoked references that run beneath the other skills, each the single source of truth for its vocabulary. Reach for them directly when the words, not the process, are the problem; or let the skills above pull them in.
/mattpocock:codebase-design: the deep-module vocabulary (module, interface, depth, seam, adapter, leverage, locality) for designing a module's shape: a lot of behaviour behind a small interface at a clean seam. /mattpocock:bdd, /mattpocock:tdd (BDD-driven), and /mattpocock:improve-codebase-architecture all speak it./mattpocock:domain-modeling: sharpen the project's domain language: challenge a fuzzy term, resolve an overloaded word ("account" doing three jobs), record a hard-to-reverse decision as an ADR. It's the active discipline /mattpocock:grill-with-docs drives to keep CONTEXT.md a clean glossary.A phase is a chunk of work inside a session: the grilling, the implementation, the QA. At the boundary between two of them you have five options, and picking between them is the fuzziest decision in this whole map:
/clear: empty the window, when nothing here matters to what's next./mattpocock:handoff writes a portable markdown file. Narrow: only for a new harness, a new directory, a colleague, or forking a side task mid-phase. What it buys is portability. It's the bridge between context windows, in either direction./compact (built-in) compresses this context and seeds a fresh session with it. The default, at the bottom of the tree rather than the first reach.Read PHASE-BOUNDARIES.md for the ordered tree: the five questions, the reasoning behind each branch, and why the primary-source cost makes Continue the one to rule out first. Make the decision at a boundary; mid-phase, continue or split the rest into subagents. /mattpocock:handoff forks; /compact continues.
Off the main flow entirely.
/mattpocock:grill-me: the same relentless interview as /mattpocock:grill-with-docs, but for when you have no codebase. Stateless: it saves nothing locally, builds no CONTEXT.md. Reach for it to sharpen any plan or design that doesn't live in a repo./mattpocock:grilling is the interview primitive itself: one question at a time via the AskUserQuestion tool, facts are the agent's job and decisions are yours. /mattpocock:grill-me and /mattpocock:grill-with-docs are the two named ways in, and /mattpocock:triage, /mattpocock:wayfinder and /mattpocock:improve-codebase-architecture all run it internally. Reach for it directly only when you want the interview with no wrapper around it./mattpocock:resolving-merge-conflicts works an in-progress merge or rebase conflict hunk by hunk, resolving by intent traced to each side's primary source rather than by picking lines, then finishes the operation. It never runs --abort. Standalone and off every flow: reach for it when you are already mid-conflict./mattpocock:prototype is a small, throwaway program that answers one design question: does this state model feel right, or what should this UI look like. It's the detour in step 2 of the main flow, but reach for it any time a design question is hard to settle on paper./mattpocock:research: delegate reading legwork to a background agent: it investigates a question against primary sources, then leaves a cited Markdown file in the repo. Keep working while it reads. The file it produces is something to take into the main flow at /mattpocock:grill-with-docs, since research feeds the thinking rather than replacing it./mattpocock:to-questionnaire comes in when the thing blocking you isn't in your head or the codebase but in someone else's, and it writes them a questionnaire to fill in. It's the inverse of /mattpocock:grill-me: instead of interviewing you about the subject, it interviews you about the send (who it's going to, what you need back) and aims the questions at the gap. What comes back is material for /mattpocock:grill-with-docs or /mattpocock:to-spec./mattpocock:wizard is for the steps only a human can take: provisioning infrastructure, setting up credentials or CI secrets, clicking through an unfamiliar third-party dashboard, running a one-off migration or cutover. It generates an interactive bash script that opens each URL, captures each value, and writes it into .env and GitHub secrets, so the procedure stops being something you re-explain to an agent every time. Model-invoked, so the agent reaches for it the moment it hits a wall only you can pass. If the agent could just do it itself, it should; this is for where a human is genuinely in the loop./mattpocock:wait-what is the corrective for a message that didn't land. Use it mid-conversation, inside any other skill, and the agent re-pitches what it just said with the context you were missing, in plain English, using the CONTEXT.md vocabulary. It works after the fact; /mattpocock:grill-with-docs is the upfront cure, because a shared language agreed early is what stops the jargon arriving at all./mattpocock:teach: learn a concept over multiple sessions, using the current directory as a stateful workspace./mattpocock:writing-for-agents is the reference for writing documents agents consume: skills, AGENTS.md/CLAUDE.md, pointed-at docs./mattpocock:writing-great-skills — reference for writing and editing skills well./mattpocock:bdd — the BDD lifecycle: Discovery → Formulation → Automation (delegated to /mattpocock:tdd). Reach for it directly when you want to define feature behavior via Gherkin scenarios./mattpocock:tdd — the BDD-driven Automation phase: test quality, seams, mocking, anti-patterns, and the red-green loop rules. Reach for it directly when you want to write or improve tests during BDD implementation./mattpocock:setup-matt-pocock-skills: run before your first engineering flow to configure the issue tracker, triage labels, and doc layout the other skills assume. Custom issue trackers also work.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 21,896 | 11,276 | -49% | 1 | 1 | 0% | 3,403 | 5,091 | +50% | 0 | 0 | — |
case-02 | fail→pass | 14,551 | 9,013 | -38% | 1 | 1 | 0% | 2,416 | 4,603 | +91% | 0 | 0 | — |
case-03 | fail→pass | 37,778 | 20,240 | -46% | 1 | 1 | 0% | 2,583 | 4,637 | +80% | 0 | 0 | — |
case-04 | fail→pass | 17,522 | 5,508 | -69% | 1 | 1 | 0% | 2,464 | 3,952 | +60% | 0 | 0 | — |
case-05 | fail→pass | 15,386 | 5,672 | -63% | 1 | 1 | 0% | 2,249 | 4,018 | +79% | 0 | 0 | — |
case-06 | fail→pass | 26,884 | 4,927 | -82% | 1 | 1 | 0% | 4,041 | 3,816 | -6% | 0 | 0 | — |
case-07 | fail→pass | 19,707 | 7,297 | -63% | 1 | 1 | 0% | 2,724 | 3,835 | +41% | 0 | 0 | — |
case-08 | fail→pass | 21,155 | 3,413 | -84% | 1 | 1 | 0% | 3,148 | 3,629 | +15% | 0 | 0 | — |
case-09 | fail→pass | 8,159 | 4,241 | -48% | 1 | 1 | 0% | 1,107 | 3,652 | +230% | 0 | 0 | — |
case-10 | fail→pass | 16,055 | 6,625 | -59% | 1 | 1 | 0% | 2,259 | 3,982 | +76% | 0 | 0 | — |
case-11 | fail→pass | 17,872 | 5,625 | -69% | 1 | 1 | 0% | 1,721 | 3,802 | +121% | 0 | 0 | — |
case-12 | fail→pass | 10,083 | 3,714 | -63% | 1 | 1 | 0% | 1,023 | 3,476 | +240% | 0 | 0 | — |
case-13 | fail→pass | 9,180 | 5,041 | -45% | 1 | 1 | 0% | 1,301 | 3,411 | +162% | 0 | 0 | — |
case-14 | fail→pass | 22,107 | 4,845 | -78% | 1 | 1 | 0% | 1,082 | 3,542 | +227% | 0 | 0 | — |
case-15 | fail→pass | 7,332 | 2,616 | -64% | 1 | 1 | 0% | 1,129 | 3,475 | +208% | 0 | 0 | — |
case-16 | fail→pass | 13,856 | 8,070 | -42% | 1 | 1 | 0% | 2,160 | 4,356 | +102% | 0 | 0 | — |
case-17 | fail→pass | 12,252 | 9,619 | -21% | 1 | 1 | 0% | 1,885 | 4,150 | +120% | 0 | 0 | — |
case-18 | fail→pass | 9,079 | 9,517 | +5% | 1 | 1 | 0% | 1,567 | 3,732 | +138% | 0 | 0 | — |
case-19 | fail→pass | 10,578 | 3,413 | -68% | 1 | 1 | 0% | 1,557 | 3,605 | +132% | 0 | 0 | — |
case-20 | pass→pass | 4,437 | 7,844 | +77% | 1 | 1 | 0% | 677 | 3,593 | +431% | 0 | 0 | — |
case-21 | pass→pass | 13,416 | 9,313 | -31% | 1 | 1 | 0% | 2,125 | 4,669 | +120% | 0 | 0 | — |
case-22 | pass→pass | 11,057 | 7,869 | -29% | 1 | 1 | 0% | 2,018 | 4,421 | +119% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +86 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/29/2026 | +73% |
Other measured skills in the registry, with their headline benchmark lift.