Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when a research task needs reproducible Kaggle discovery, metadata inspection, bounded public-data downloads, competition or kernel discovery, model discovery, or an explicitly approved Kaggle write/delete operation through the official CLI.
.claude/skills/brycewang-stanford-kaggle-research/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-03 | ✗→✓ | ▲ Improved | 654% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 35% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -48% | 0% |
| case-10 | ✗→✓ | ▲ Improved | -2% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 86% | 0% |
Use the official Kaggle CLI through the policy-enforcing wrapper in this skill. The wrapper records bounded, redacted audit data; confines downloads to an approved directory; and blocks remote mutation unless the user explicitly authorizes it.
download, write, or delete. Do not broaden user authority.
kaggle>=2.2,<3.read, print, echo, log, or commit credential values.
bash python scripts/kaggle_research.py doctor --json
under an explicitly chosen output root.
generated artifacts. Distinguish verified observations from assumptions.
Pass Kaggle arguments after -- so their order is preserved:
bashpython scripts/kaggle_research.py run --audit artifacts/audit.json -- datasets list -s iris -v python scripts/kaggle_research.py run --output-root artifacts/kaggle -- datasets download -d owner/dataset
Preview any potentially mutating command first:
bashpython scripts/kaggle_research.py run --dry-run --allow-write -- datasets create -p dataset-package
An actual remote write additionally requires explicit user authorization and --allow-write. A delete additionally requires --allow-delete and --confirm-resource matching the exact resource classified by the wrapper. The runtime never retries writes or deletes.
The live smoke workflow calls Kaggle's real service, inspects all supported resource groups, downloads a small public dataset, and verifies its hash:
bashpython scripts/kaggle_research.py smoke-readonly --output-root artifacts/kaggle-smoke --report artifacts/kaggle-smoke-report.json
The corresponding integration test is opt-in so normal unit tests do not depend on network access:
bashAERS_KAGGLE_LIVE=1 python -m unittest discover -s tests -p "test_live_readonly.py" -v
Only run the live lane when credentials are already available in the process environment. It must remain read/download-only.
references/authentication.md
references/datasets.md
references/competitions.md
references/kernels.md
references/models.md
references/testing-and-safety.md
Read only the reference page required for the active task.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 6,164 | 4,610 | -25% | 1 | 1 | 0% | 248 | 1,072 | +332% | 0 | 0 | — |
case-02 | fail→fail | 13,914 | 4,909 | -65% | 1 | 1 | 0% | 2,335 | 1,176 | -50% | 0 | 0 | — |
case-03 | fail→pass | 9,925 | 26,286 | +165% | 1 | 1 | 0% | 747 | 5,632 | +654% | 0 | 0 | — |
case-04 | pass→pass | 14,021 | 1,345 | -90% | 1 | 1 | 0% | 1,221 | 919 | -25% | 0 | 0 | — |
case-05 | pass→pass | 5,343 | 2,750 | -49% | 1 | 1 | 0% | 1,021 | 1,244 | +22% | 0 | 0 | — |
case-06 | pass→pass | 7,712 | 4,151 | -46% | 1 | 1 | 0% | 1,437 | 1,523 | +6% | 0 | 0 | — |
case-07 | fail→pass | 7,094 | 5,477 | -23% | 1 | 1 | 0% | 1,241 | 1,671 | +35% | 0 | 0 | — |
case-08 | pass→pass | 7,950 | 3,907 | -51% | 1 | 1 | 0% | 1,429 | 1,458 | +2% | 0 | 0 | — |
case-09 | fail→pass | 13,345 | 2,554 | -81% | 1 | 1 | 0% | 2,346 | 1,222 | -48% | 0 | 0 | — |
case-10 | fail→pass | 8,311 | 2,762 | -67% | 1 | 1 | 0% | 1,387 | 1,360 | -2% | 0 | 0 | — |
case-11 | fail→pass | 3,695 | 3,775 | +2% | 1 | 1 | 0% | 737 | 1,371 | +86% | 0 | 0 | — |
case-12 | fail→pass | 9,343 | 2,220 | -76% | 1 | 1 | 0% | 1,483 | 1,095 | -26% | 0 | 0 | — |
case-13 | fail→pass | 9,386 | 2,978 | -68% | 1 | 1 | 0% | 1,780 | 1,224 | -31% | 0 | 0 | — |
case-14 | pass→pass | 10,454 | 4,766 | -54% | 1 | 1 | 0% | 1,822 | 1,579 | -13% | 0 | 0 | — |
case-15 | pass→pass | 12,266 | 5,282 | -57% | 1 | 1 | 0% | 2,033 | 1,570 | -23% | 0 | 0 | — |
case-16 | fail→pass | 11,786 | 2,391 | -80% | 1 | 1 | 0% | 1,991 | 1,099 | -45% | 0 | 0 | — |
case-17 | fail→pass | 7,277 | 3,485 | -52% | 1 | 1 | 0% | 1,175 | 1,416 | +21% | 0 | 0 | — |
case-18 | fail→pass | 10,608 | 3,827 | -64% | 1 | 1 | 0% | 1,860 | 1,332 | -28% | 0 | 0 | — |
case-19 | fail→pass | 11,070 | 4,871 | -56% | 1 | 1 | 0% | 2,078 | 1,607 | -23% | 0 | 0 | — |
case-20 | pass→pass | 8,654 | 5,589 | -35% | 1 | 1 | 0% | 1,723 | 1,874 | +9% | 0 | 0 | — |
case-21 | pass→pass | 8,515 | 5,564 | -35% | 1 | 1 | 0% | 1,651 | 1,729 | +5% | 0 | 0 | — |
case-22 | pass→pass | 12,967 | 7,450 | -43% | 1 | 1 | 0% | 2,811 | 2,137 | -24% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +50 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.