Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when drafting a QIP, after qv-qip-triage recommends creating one, shaping a fuzzy technical proposal with alternatives and consequences, or invoking /qv-qip-create.
.claude/skills/tetherto-qv-qip-create/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 51% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 204% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 135% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 12% | 0% |
Help an author draft a QIP before posting to Slack Canvas.
Core principle: a QIP draft is earned by evidence, decision context, and explicit trade-offs. Do not treat this as a template-filling or content-generation task.
Use when:
qv-qip-triage recommended a QIP and the user confirmed/qv-qip-createDo NOT use for:
qv-qip-triage)qv-qip-review)Read before drafting:
docs/architecture/PRINCIPLES.md for lightweight principle checksFor every new QIP request, ask at least one clarifying question before drafting unless the user explicitly says:
draft with assumptionsmake a first passno questionsuse my brief as final contextA terse slash-command prompt with only a title, technology list, package name, or desired outcome is fuzzy idea mode. Examples:
/qv-qip-create Kotlin SDK Android Coroutines JNI Maven Central/qv-qip-create native mobile SDK/qv-qip-create improve registry replicationIf the prompt names several possible motivations, ask the user to choose the primary driver instead of guessing.
Use only when the user already provided:
Do not classify a request as clear proposal mode just because it contains many solution details. A rich technology list is still fuzzy idea mode if the problem, timing, affected surface, and trade-offs are not explicit.
Ask only for missing essentials.
Use when the user has a problem or direction but no settled solution.
Ask short questions one at a time until enough context exists:
Do not dump all questions at once unless the user asks for a batch.
Before saving or presenting a QIP draft, confirm these are known:
If two or more are unknown, do not draft. Ask the next most important question. If exactly one is unknown, either ask or label it clearly as an assumption.
Do not invent:
Every substantive claim in the QIP must come from one of:
Prefer questions over assumptions for architectural proposals. If the draft includes assumptions, keep them explicit and easy for the author to confirm or delete.
If the proposal affects SDK API, native bindings, runtime, mobile support, storage, transport, model registry, release flow, or security boundaries, inspect relevant repo docs or code before drafting.
Minimum verification:
Do not produce a full QIP from general knowledge alone.
Before the final draft, produce a short People to consult before posting note.
Use this advice rule: consult everyone meaningfully affected and people with relevant expertise.
Include:
Advice is direction plus reasoning, not a vote.
arch/qips/<short-slug>.md is only an example local path, not a required repo path.idle timeout, whole-stream deadline, idempotent, terminal failureAlternatives considered separate from Solution. Include obvious options reviewers would expect to see, but keep each to 1-2 sentences or link to research for detailed analysis.Consequences as positive impact plus trade-offs for proposal review. Avoid a probable-production-bugs list; state what reviewers must accept, then add mitigation only where it affects whether the proposal should be accepted, changed, or split.Before finalizing, check whether the proposal needs any of these:
If the rationale or impact is unclear, investigate the existing code/docs enough to form a grounded draft or ask the user for the missing decision context.
Stop drafting and ask a question when you notice any of these:
For file-based drafts, reply with the saved path and a brief summary of what changed. Do not paste the whole QIP unless the user asks.
For chat-only drafts, first show People to consult before posting, then the draft:
markdownPeople to consult before posting - <role or team>: <why> QIP Template :clipboard: Approvers ...
End with this author checklist:
markdownAuthor checklist - [ ] Problem is clear and timely - [ ] Solution is concrete enough to review - [ ] Chosen solution is justified against obvious alternatives - [ ] Trust boundaries and security properties are explicit when affected - [ ] Compatibility, migration, and release impact are explicit when affected - [ ] Alternatives considered is brief or links to detailed research - [ ] Consequences state positive impact and trade-offs reviewers must accept - [ ] Out of scope is explicit - [ ] Approvers table preserved - [ ] Consultation note reflects affected teams and expertise
If the user asks whether the draft is ready to post, suggest qv-qip-review.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-17 | fail→fail | 14,128 | 18,743 | +33% | 1 | 1 | 0% | 2,392 | 4,270 | +79% | 0 | 0 | — |
case-01 | pass→pass | 17,143 | 15,353 | -10% | 1 | 1 | 0% | 2,751 | 4,432 | +61% | 0 | 0 | — |
case-02 | fail→pass | 20,114 | 6,794 | -66% | 1 | 1 | 0% | 3,190 | 3,118 | -2% | 0 | 0 | — |
case-03 | fail→fail | 10,660 | 4,514 | -58% | 1 | 1 | 0% | 1,629 | 2,711 | +66% | 0 | 0 | — |
case-04 | fail→pass | 13,072 | 6,848 | -48% | 1 | 1 | 0% | 2,006 | 3,039 | +51% | 0 | 0 | — |
case-05 | fail→pass | 5,642 | 4,777 | -15% | 1 | 1 | 0% | 915 | 2,781 | +204% | 0 | 0 | — |
case-06 | pass→pass | 7,955 | 4,546 | -43% | 1 | 1 | 0% | 646 | 2,764 | +328% | 0 | 0 | — |
case-07 | pass→pass | 16,894 | 13,604 | -19% | 1 | 1 | 0% | 2,702 | 4,231 | +57% | 0 | 0 | — |
case-08 | fail→fail | 12,573 | 6,379 | -49% | 1 | 1 | 0% | 1,919 | 2,246 | +17% | 0 | 0 | — |
case-09 | pass→pass | 17,878 | 21,435 | +20% | 1 | 1 | 0% | 2,727 | 4,812 | +76% | 0 | 0 | — |
case-10 | fail→pass | 15,129 | 21,237 | +40% | 1 | 1 | 0% | 2,375 | 5,574 | +135% | 0 | 0 | — |
case-11 | fail→pass | 13,545 | 3,736 | -72% | 1 | 1 | 0% | 2,241 | 2,512 | +12% | 0 | 0 | — |
case-12 | fail→pass | 18,647 | 13,410 | -28% | 1 | 1 | 0% | 2,937 | 4,041 | +38% | 0 | 0 | — |
case-13 | pass→pass | 13,448 | 17,683 | +31% | 1 | 1 | 0% | 2,242 | 4,826 | +115% | 0 | 0 | — |
case-14 | pass→pass | 16,313 | 15,461 | -5% | 1 | 1 | 0% | 2,575 | 3,980 | +55% | 0 | 0 | — |
case-15 | fail→pass | 17,117 | 11,636 | -32% | 1 | 1 | 0% | 2,627 | 3,826 | +46% | 0 | 0 | — |
case-16 | fail→pass | 16,452 | 20,326 | +24% | 1 | 1 | 0% | 2,836 | 4,686 | +65% | 0 | 0 | — |
case-18 | fail→pass | 12,036 | 22,931 | +91% | 1 | 1 | 0% | 1,756 | 4,842 | +176% | 0 | 0 | — |
case-19 | pass→pass | 10,172 | 4,226 | -58% | 1 | 1 | 0% | 1,584 | 2,703 | +71% | 0 | 0 | — |
case-20 | pass→pass | 18,896 | 16,858 | -11% | 1 | 1 | 0% | 3,028 | 4,660 | +54% | 0 | 0 | — |
case-21 | pass→pass | 17,326 | 14,675 | -15% | 1 | 1 | 0% | 2,759 | 4,336 | +57% | 0 | 0 | — |
case-22 | fail→pass | 13,084 | 13,171 | +1% | 1 | 1 | 0% | 2,107 | 3,588 | +70% | 0 | 0 | — |
case-23 | pass→fail | 16,668 | 4,614 | -72% | 1 | 1 | 0% | 2,788 | 2,326 | -17% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 21 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +39 percentage points is the difference between those two pass rates over the 21 comparable cases. 2 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.