Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when choosing local validation scope in TiDB work, especially to separate fast coding-loop checks from completion checks and avoid unnecessary slow commands.
.claude/skills/pingcap-tidb-verify-profile/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | -36% | 0% |
| case-05 | ✗→✓ | ▲ Improved | -33% | 0% |
| case-06 | ✗→✓ | ▲ Improved | -73% | 0% |
| case-07 | ✗→✓ | ▲ Improved | -31% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -60% | 0% |
Use this skill to select validation for repository changes, including code, formatting, documentation, testdata, and build configuration. Read-only analysis does not require build/test checks. Policy requirements still come from AGENTS.md; this skill is the execution guide.
WIP (coding loop)Use while iterating on a change.
go test -run <TestName> -tags=intest,deadlock).make lint, package-wide runs, realtikvtest).Ready (completion gate)Use when delivering changes or preparing a PR, as defined in AGENTS.md -> Quick Decision Matrix. Select checks from the actual change type; status wording neither adds nor waives checks.
AGENTS.md -> Task -> Validation Matrix and the applicable special cases in Quick Decision Matrix.make lint.AGENTS.md -> Agent Output Contract for final reporting.Reuse completed checks that still cover the delivered changes. Rerun affected checks when relevant changes or new failures invalidate the results, not merely because another status update is due.
Heavy (explicitly required)Use only when scope or user request requires expensive checks.
make bazel_lint_changed unless the user explicitly requests it.| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 15,684 | 6,040 | -61% | 1 | 1 | 0% | 2,157 | 1,378 | -36% | 0 | 0 | — |
case-02 | fail→fail | 21,350 | 5,371 | -75% | 1 | 1 | 0% | 2,033 | 585 | -71% | 0 | 0 | — |
case-03 | fail→fail | 20,476 | 15,912 | -22% | 1 | 1 | 0% | 3,143 | 1,824 | -42% | 0 | 0 | — |
case-04 | fail→fail | 15,001 | 2,810 | -81% | 1 | 1 | 0% | 2,011 | 786 | -61% | 0 | 0 | — |
case-05 | fail→pass | 9,463 | 5,132 | -46% | 1 | 1 | 0% | 1,344 | 897 | -33% | 0 | 0 | — |
case-06 | fail→pass | 30,301 | 3,405 | -89% | 1 | 1 | 0% | 2,955 | 808 | -73% | 0 | 0 | — |
case-07 | fail→pass | 15,604 | 10,484 | -33% | 1 | 1 | 0% | 2,085 | 1,445 | -31% | 0 | 0 | — |
case-08 | fail→pass | 30,355 | 4,232 | -86% | 1 | 1 | 0% | 2,530 | 1,008 | -60% | 0 | 0 | — |
case-09 | fail→pass | 29,132 | 6,988 | -76% | 1 | 1 | 0% | 2,160 | 1,581 | -27% | 0 | 0 | — |
case-10 | fail→pass | 32,345 | 3,752 | -88% | 1 | 1 | 0% | 2,258 | 820 | -64% | 0 | 0 | — |
case-11 | fail→pass | 12,390 | 4,497 | -64% | 1 | 1 | 0% | 1,775 | 955 | -46% | 0 | 0 | — |
case-12 | pass→pass | 17,003 | 4,106 | -76% | 1 | 1 | 0% | 1,748 | 931 | -47% | 0 | 0 | — |
case-13 | fail→pass | 31,245 | 7,947 | -75% | 1 | 1 | 0% | 1,191 | 1,566 | +31% | 0 | 0 | — |
case-14 | fail→pass | 12,759 | 3,829 | -70% | 1 | 1 | 0% | 1,865 | 906 | -51% | 0 | 0 | — |
case-15 | fail→pass | 8,988 | 2,844 | -68% | 1 | 1 | 0% | 1,389 | 738 | -47% | 0 | 0 | — |
case-16 | fail→fail | 15,681 | 8,636 | -45% | 1 | 1 | 0% | 2,530 | 1,662 | -34% | 0 | 0 | — |
case-17 | pass→pass | 10,544 | 3,876 | -63% | 1 | 1 | 0% | 1,604 | 922 | -43% | 0 | 0 | — |
case-18 | fail→fail | 12,014 | 9,458 | -21% | 1 | 1 | 0% | 2,180 | 1,729 | -21% | 0 | 0 | — |
case-19 | pass→pass | 12,732 | 4,767 | -63% | 1 | 1 | 0% | 1,823 | 769 | -58% | 0 | 0 | — |
case-20 | fail→fail | 13,422 | 9,333 | -30% | 1 | 1 | 0% | 2,021 | 1,735 | -14% | 0 | 0 | — |
case-21 | pass→pass | 31,065 | 10,287 | -67% | 1 | 1 | 0% | 2,779 | 2,177 | -22% | 0 | 0 | — |
case-22 | pass→pass | 7,478 | 6,636 | -11% | 1 | 1 | 0% | 1,060 | 1,185 | +12% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 20 counted toward the lift figure. The other 2 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +50 percentage points is the difference between those two pass rates over the 20 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.