Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use at study design whenever the task ships a specific named object in data/ — one paper, one structure, one instance — and again before writing. Covers reporting that item's own numbers under its own name, choosing the worked example by the task's pointer rather than by your result, and how to widen scope without dropping it.
.claude/skills/tangxiangru-the-supplied-item-is-the-graded-unit/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 98% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 117% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 214% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 294% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 533% | 0% |
When a task hands you one named object, the questions asked of your work are written about that object, often naming its identifier. Widening to the corpus it came from is frequently better research. Concluding that it is a negative control is frequently correct. Neither is a substitute for reporting the item's own numbers under its own name.
A run read the one paper shipped in data/, recovered the fifteen-paper benchmark it belongs to, and priced the two scopes honestly: one paper gives a ±28-point interval, fifteen give ±8.5. It locked scope to fifteen and recorded "Scoping to the shipped paper. Priced and rejected above; it withdraws four hypotheses." Everything after that was excellent work, and the shipped paper survives in the final report as one appendix table row and one reference — two mentions, against eight in a plain agent's report.
The graded object was that paper's assembled Hamiltonian. The run generated it, wrote it to outputs/, and never printed it. The one worked derivation the report does show is a different paper — the one the run got right. Three requirements scored 5, 15 and 5 against a plain agent's 32, 25 and 45, and that agent's entire advantage was six lines printing the shipped paper's Hamiltonian and naming the supplementary equation it matches.
The second shape: a structure task ships one pair, and every arm correctly finds it is a monomeric negative rather than the paper's case. The arm that still tabulated the pair's own sequence identity, alignment scores and runtimes scored 25/25/45/40. The arm that reframed it as a true negative and reported those named quantities nowhere scored 0/15/25/5. Both were right about the biology.
data/ by name at design time. Give each its ownsubsection in the report, with the identifier in the heading.
output list names, in the source's units. Widening is additive: the corpus arm is a new section, never a replacement, and the shipped item keeps its own values even when the corpus gives a tighter interval.
transform matrix — in the report. A path under outputs/ does not discharge a deliverable.
Showing the case you got right and burying the supplied case you got wrong is selecting the exhibit on the outcome. If the supplied case is the one that failed, that is the one to work through, and the failure is the finding.
why it is what it is. "This pair is a monomeric negative" is a conclusion that needs the numbers under it, not instead of it.
Search the report for each identifier in data/. A single-digit count, or hits only in a reference list and an appendix row, means the graded unit is not in the document.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 40,821 | 9,298 | -77% | 1 | 1 | 0% | 7,462 | 1,260 | -83% | 0 | 0 | — |
case-02 | fail→fail | 8,777 | 14,764 | +68% | 1 | 1 | 0% | 488 | 1,369 | +181% | 0 | 0 | — |
case-03 | fail→fail | 10,490 | 9,946 | -5% | 1 | 1 | 0% | 298 | 1,117 | +275% | 0 | 0 | — |
case-04 | fail→fail | 14,324 | 50,356 | +252% | 1 | 1 | 0% | 2,046 | 9,673 | +373% | 0 | 0 | — |
case-05 | fail→fail | 47,269 | 9,663 | -80% | 1 | 1 | 0% | 7,405 | 1,251 | -83% | 0 | 0 | — |
case-06 | fail→pass | 24,722 | 49,100 | +99% | 1 | 1 | 0% | 4,555 | 9,009 | +98% | 0 | 0 | — |
case-07 | fail→fail | 30,960 | 12,334 | -60% | 1 | 1 | 0% | 2,511 | 1,483 | -41% | 0 | 0 | — |
case-08 | fail→fail | 36,198 | 12,275 | -66% | 1 | 1 | 0% | 5,574 | 1,599 | -71% | 0 | 0 | — |
case-09 | fail→pass | 19,145 | 27,446 | +43% | 1 | 1 | 0% | 2,570 | 5,582 | +117% | 0 | 0 | — |
case-10 | fail→fail | 10,436 | 10,380 | -1% | 1 | 1 | 0% | 339 | 1,301 | +284% | 0 | 0 | — |
case-11 | fail→fail | 43,830 | 89,634 | +105% | 1 | 1 | 0% | 3,641 | 7,346 | +102% | 0 | 0 | — |
case-12 | fail→pass | 17,024 | 47,093 | +177% | 1 | 1 | 0% | 2,866 | 9,002 | +214% | 0 | 0 | — |
case-13 | fail→pass | 15,952 | 46,778 | +193% | 1 | 1 | 0% | 2,439 | 9,608 | +294% | 0 | 0 | — |
case-14 | fail→fail | 11,542 | 8,320 | -28% | 1 | 1 | 0% | 683 | 1,168 | +71% | 0 | 0 | — |
case-15 | fail→pass | 5,841 | 24,244 | +315% | 1 | 1 | 0% | 760 | 4,812 | +533% | 0 | 0 | — |
case-16 | fail→pass | 16,090 | 35,977 | +124% | 1 | 1 | 0% | 2,456 | 8,882 | +262% | 0 | 0 | — |
case-17 | fail→pass | 14,041 | 34,048 | +142% | 1 | 1 | 0% | 1,307 | 4,189 | +221% | 0 | 0 | — |
case-18 | fail→fail | 7,047 | 7,444 | +6% | 1 | 1 | 0% | 438 | 1,109 | +153% | 0 | 0 | — |
case-19 | fail→fail | 3,160 | 32,316 | +923% | 1 | 1 | 0% | 422 | 3,068 | +627% | 0 | 0 | — |
case-20 | pass→fail | 21,291 | 7,453 | -65% | 1 | 1 | 0% | 4,117 | 1,038 | -75% | 0 | 0 | — |
case-21 | pass→pass | 13,294 | 22,876 | +72% | 1 | 1 | 0% | 2,272 | 4,783 | +111% | 0 | 0 | — |
case-22 | pass→pass | 8,223 | 9,313 | +13% | 1 | 1 | 0% | 1,470 | 2,367 | +61% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 11 counted toward the lift figure. The other 11 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +27 percentage points is the difference between those two pass rates over the 11 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.