Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Triage unexpected TiDB test diffs that seem unrelated to the current PR. Use when plan/result/testdata changes appear after merge/rebase or only in specific local runs, especially to quickly rule in/out failpoint enablement issues.
.claude/skills/pingcap-tidb-test-diff-triage/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 34% | 0% |
| case-02 | ✗→✓ | ▲ Improved | -9% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -58% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -49% | 0% |
| case-13 | ✗→✓ | ▲ Improved | -52% | 0% |
Use this workflow when:
In TiDB, many test behaviors rely on failpoint instrumentation. -tags=intest,deadlock does not enable failpoints.
docs/agents/testing-flow.md -> Failpoint decision for unit tests against the affected package.Failpoint-enabled run and add -count=1 to the go test command for reproducibility.If failpoint is not the cause:
-run <TestName> -count=1).Useful commands:
bashgit bisect start git bisect bad <bad_commit> git bisect good <good_commit>
Only sync expected plan/result when one of these is true:
Do not record/update testdata before root cause is identified.
Symptom: what diff changed.Scope: single test vs full suite.Failpoint check: commands and result.First bad commit: hash and title (if bisected).Conclusion: setup issue / upstream behavior change / local regression.Action: rerun with failpoint, sync expected, or continue fixing code.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 47,718 | 10,665 | -78% | 1 | 1 | 0% | 1,502 | 2,011 | +34% | 0 | 0 | — |
case-02 | fail→pass | 44,355 | 30,471 | -31% | 1 | 1 | 0% | 3,115 | 2,843 | -9% | 0 | 0 | — |
case-03 | fail→fail | 7,236 | 10,157 | +40% | 1 | 1 | 0% | 525 | 2,075 | +295% | 0 | 0 | — |
case-04 | pass→pass | 17,328 | 31,496 | +82% | 1 | 1 | 0% | 3,002 | 2,396 | -20% | 0 | 0 | — |
case-05 | pass→pass | 11,714 | 9,599 | -18% | 1 | 1 | 0% | 2,090 | 2,086 | -0% | 0 | 0 | — |
case-06 | pass→pass | 17,940 | 17,049 | -5% | 1 | 1 | 0% | 2,936 | 3,397 | +16% | 0 | 0 | — |
case-07 | fail→pass | 52,175 | 4,808 | -91% | 1 | 1 | 0% | 2,648 | 1,121 | -58% | 0 | 0 | — |
case-08 | pass→pass | 4,671 | 2,772 | -41% | 1 | 1 | 0% | 685 | 700 | +2% | 0 | 0 | — |
case-09 | fail→pass | 14,685 | 5,255 | -64% | 1 | 1 | 0% | 2,025 | 1,035 | -49% | 0 | 0 | — |
case-10 | fail→fail | 13,813 | 8,069 | -42% | 1 | 1 | 0% | 1,828 | 1,684 | -8% | 0 | 0 | — |
case-11 | pass→pass | 14,048 | 5,246 | -63% | 1 | 1 | 0% | 1,984 | 1,227 | -38% | 0 | 0 | — |
case-12 | pass→pass | 10,034 | 9,449 | -6% | 1 | 1 | 0% | 1,453 | 1,513 | +4% | 0 | 0 | — |
case-13 | fail→pass | 13,090 | 3,880 | -70% | 1 | 1 | 0% | 1,884 | 912 | -52% | 0 | 0 | — |
case-14 | fail→pass | 15,178 | 4,287 | -72% | 1 | 1 | 0% | 2,155 | 1,060 | -51% | 0 | 0 | — |
case-15 | fail→pass | 15,325 | 8,268 | -46% | 1 | 1 | 0% | 2,178 | 1,628 | -25% | 0 | 0 | — |
case-16 | fail→pass | 9,770 | 4,603 | -53% | 1 | 1 | 0% | 1,261 | 1,083 | -14% | 0 | 0 | — |
case-17 | fail→pass | 19,375 | 7,719 | -60% | 1 | 1 | 0% | 1,673 | 1,384 | -17% | 0 | 0 | — |
case-18 | pass→pass | 12,003 | 5,235 | -56% | 1 | 1 | 0% | 1,826 | 1,111 | -39% | 0 | 0 | — |
case-19 | fail→pass | 12,337 | 3,144 | -75% | 1 | 1 | 0% | 2,176 | 828 | -62% | 0 | 0 | — |
case-20 | fail→pass | 11,847 | 2,683 | -77% | 1 | 1 | 0% | 1,794 | 685 | -62% | 0 | 0 | — |
case-21 | fail→pass | 13,587 | 3,889 | -71% | 1 | 1 | 0% | 1,738 | 1,032 | -41% | 0 | 0 | — |
case-22 | fail→pass | 10,342 | 4,827 | -53% | 1 | 1 | 0% | 1,503 | 1,169 | -22% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +59 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.