Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Answer a bounded question with current cited
.claude/skills/boshu2-research/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-17 | ✗→✓ | ▲ Improved | 98% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 2% | 0% |
| case-04 | ✓→✗ | ▼ Worse | 179% | 0% |
| case-11 | ✓→✗ | ▼ Worse | -46% | 0% |
| case-15 | ✓→✗ | ▼ Worse | -24% | 0% |
Answer one bounded question with current evidence. Research informs a caller; it does not select work, approve a plan, mutate lifecycle state, or decide what happens next.
required for a useful answer.
use current primary sources.
Use the current agent inline by default. Parallel readers or alternate runtimes are optional execution choices only when the caller authorizes them. Prior research, CASS, MS, codebase recon, and pattern mining are advisory context sources, not required phases. Hydrate only the sources the current decision needs and return cited evidence with source identity and freshness; never build or maintain a merged context store.
A claim about what code does cites the commit it was observed at, plus file:line — code moves, and a citation without a revision decays silently into a claim about a repository that no longer exists. For the working tree, record the current HEAD and whether the cited file carries uncommitted changes. The named failure mode is the floating citation: a path and line that resolved when written, drifted after a refactor, and now lends false authority to a stale answer. A reader must be able to run git show <commit>:<path> and see the cited lines; a code claim that cannot survive that replay is reported as unverified, not asserted.
Research is done when its capability flags are answerable, not when effort feels sufficient. At the start, derive from the bounded question a short list of capability statements — "can name the module that owns X, with citation", "can state whether Y is reachable from Z, or that this is unknown". The stop condition: every flag is either satisfied with evidence or explicitly reported unknown with what was searched. Hours spent and files read are not flags. The named failure mode is effort-shaped doneness — stopping because the search was long, and shipping an answer whose load-bearing claim was never actually established. If a flag stays unsatisfiable inside scope, say so and stop; widening the question mid-search is a new question, and the caller owns it.
When the caller supplies several reports for one bounded question, synthesize them as evidence inside this same Research invocation:
supplied identifier, title, author/runtime when known, and revision or date when supplied. Assign a short local label without replacing that identity.
reference. Normalize wording only for comparison; never merge citations or make agreement erase provenance.
Agreement means independent reports support the same claim. Contradiction preserves the conflicting claims and evidence. Unknown means the reports do not establish the fact or the underlying source was not checked. Reports that repeat one upstream source are agreement in wording, not independent corroboration; preserve that shared provenance.
question requires it. A report's conclusion is advisory, not authority.
where they disagree, and what remains unknown. Report checked and unchecked sources, then stop.
Do not recursively launch another Research pass, invent a synthesis umbrella, or start a new runtime merely because multiple reports exist. Additional readers remain caller-authorized execution choices, not part of this procedure.
For a quick question, return the cited answer directly. When the caller asks for a durable artifact, write one report containing:
For a durable synthesis of multiple reports, also include source_ledger and comparison (agreements, contradictions, and unknowns) as defined by the output schema. Single-report outputs may omit those optional fields.
Do not emit approval, confidence gates, retry instructions, owner, next action, or delivery state.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→fail | 7,032 | 18,680 | +166% | 1 | 1 | 0% | 1,157 | 3,231 | +179% | 0 | 0 | — |
case-01 | fail→fail | 4,885 | 5,204 | +7% | 1 | 1 | 0% | 211 | 1,200 | +469% | 0 | 0 | — |
case-02 | fail→fail | 24,316 | 3,020 | -88% | 1 | 1 | 0% | 3,870 | 1,171 | -70% | 0 | 0 | — |
case-03 | fail→fail | 15,611 | 5,392 | -65% | 1 | 1 | 0% | 2,249 | 1,337 | -41% | 0 | 0 | — |
case-05 | fail→fail | 5,234 | 9,589 | +83% | 1 | 1 | 0% | 798 | 2,590 | +225% | 0 | 0 | — |
case-06 | fail→fail | 5,171 | 7,100 | +37% | 1 | 1 | 0% | 862 | 2,135 | +148% | 0 | 0 | — |
case-07 | fail→fail | 5,378 | 5,704 | +6% | 1 | 1 | 0% | 295 | 1,324 | +349% | 0 | 0 | — |
case-08 | fail→fail | 9,028 | 5,725 | -37% | 1 | 1 | 0% | 1,307 | 1,249 | -4% | 0 | 0 | — |
case-09 | pass→pass | 7,654 | 10,486 | +37% | 1 | 1 | 0% | 1,330 | 2,250 | +69% | 0 | 0 | — |
case-10 | fail→fail | 6,914 | 5,461 | -21% | 1 | 1 | 0% | 1,158 | 1,246 | +8% | 0 | 0 | — |
case-11 | pass→fail | 13,345 | 5,529 | -59% | 1 | 1 | 0% | 2,203 | 1,199 | -46% | 0 | 0 | — |
case-12 | fail→fail | 4,323 | 4,817 | +11% | 1 | 1 | 0% | 627 | 1,282 | +104% | 0 | 0 | — |
case-13 | fail→fail | 16,860 | 5,928 | -65% | 1 | 1 | 0% | 2,584 | 1,249 | -52% | 0 | 0 | — |
case-14 | fail→fail | 14,167 | 7,277 | -49% | 1 | 1 | 0% | 2,209 | 1,364 | -38% | 0 | 0 | — |
case-15 | pass→fail | 10,507 | 5,692 | -46% | 1 | 1 | 0% | 1,653 | 1,249 | -24% | 0 | 0 | — |
case-16 | pass→fail | 4,901 | 5,709 | +16% | 1 | 1 | 0% | 742 | 1,271 | +71% | 0 | 0 | — |
case-17 | fail→pass | 12,210 | 16,960 | +39% | 1 | 1 | 0% | 1,926 | 3,809 | +98% | 0 | 0 | — |
case-18 | pass→fail | 16,052 | 5,492 | -66% | 1 | 1 | 0% | 2,628 | 1,199 | -54% | 0 | 0 | — |
case-19 | fail→fail | 3,429 | 4,121 | +20% | 1 | 1 | 0% | 323 | 1,161 | +259% | 0 | 0 | — |
case-20 | pass→fail | 10,300 | 3,430 | -67% | 1 | 1 | 0% | 1,772 | 1,205 | -32% | 0 | 0 | — |
case-21 | fail→fail | 4,860 | 4,899 | +1% | 1 | 1 | 0% | 265 | 1,368 | +416% | 0 | 0 | — |
case-22 | fail→pass | 11,250 | 4,898 | -56% | 1 | 1 | 0% | 1,759 | 1,799 | +2% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 6 counted toward the lift figure. The other 16 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of -18 percentage points is the difference between those two pass rates over the 6 comparable cases. 8 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.