Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Download datasets, manage competitions and notebooks via Kaggle API
.claude/skills/brycewang-stanford-kaggle-api-guide/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 133% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 157% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 243% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 173% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 213% | 0% |
Kaggle is the world's largest data science and machine learning community, hosting thousands of datasets, competitions, and computational notebooks. The Kaggle API provides programmatic access to these resources, enabling researchers to download datasets, submit competition entries, manage kernels (notebooks), and explore the Kaggle ecosystem from the command line or scripts.
For academic researchers, Kaggle is a valuable resource for accessing curated, well-documented datasets across diverse domains including healthcare, natural language processing, computer vision, economics, and social sciences. Many published research papers use Kaggle datasets as benchmarks, and the platform's competition infrastructure provides standardized evaluation frameworks for comparing methods.
The Kaggle API is available as a Python CLI tool and library. It requires a free Kaggle account and API token for authentication. The API supports dataset search and download, competition data retrieval, kernel management, and model access.
A free Kaggle API token is required. Generate one from your Kaggle account settings at https://www.kaggle.com/settings.
Download the kaggle.json credentials file and place it in the standard location:
bash# The kaggle.json file should be at ~/.kaggle/kaggle.json # It contains your username and key from your Kaggle account settings mkdir -p ~/.kaggle # Move your downloaded kaggle.json to ~/.kaggle/kaggle.json chmod 600 ~/.kaggle/kaggle.json
Alternatively, use environment variables:
bashexport KAGGLE_USERNAME=$KAGGLE_USERNAME export KAGGLE_KEY=$KAGGLE_KEY
Install the CLI tool:
bashpip install kaggle
Find datasets by keyword, file type, or license.
bash# Search for datasets kaggle datasets list -s "climate change" --sort-by votes # Search with specific criteria kaggle datasets list -s "medical imaging" --file-type csv --max-size 1000000
bash# Download and unzip a dataset kaggle datasets download -d "heptapod/titanic" --unzip -p ./data/titanic/ # Download a specific file from a dataset kaggle datasets download -d "yelp-dataset/yelp-dataset" -f "yelp_academic_dataset_review.json" -p ./data/
bash# List active competitions kaggle competitions list # Download competition data (must accept rules on kaggle.com first) kaggle competitions download -c "house-prices-advanced-regression-techniques" -p ./data/house-prices/
bash# Submit predictions kaggle competitions submit -c "house-prices-advanced-regression-techniques" \ -f ./submission.csv -m "Random forest baseline v1" # Check submission status kaggle competitions submissions -c "house-prices-advanced-regression-techniques"
bash# Search for notebooks kaggle kernels list -s "transformer nlp" --sort-by voteCount # Pull a notebook to local kaggle kernels pull "username/notebook-name" -p ./notebooks/ # Push a notebook to Kaggle kaggle kernels push -p ./my-notebook/
pythonimport subprocess import json import os def search_kaggle_datasets(query, sort_by="votes", max_results=10): """Search Kaggle datasets and return structured results.""" cmd = [ "kaggle", "datasets", "list", "-s", query, "--sort-by", sort_by, "--max-size", "50000000", "--csv" ] result = subprocess.run(cmd, capture_output=True, text=True) lines = result.stdout.strip().split("\n") if len(lines) < 2: return [] headers = lines[0].split(",") datasets = [] for line in lines[1:max_results + 1]: values = line.split(",") dataset = dict(zip(headers, values)) datasets.append(dataset) return datasets def download_dataset(dataset_ref, output_dir="./data"): """Download a Kaggle dataset by reference.""" os.makedirs(output_dir, exist_ok=True) cmd = [ "kaggle", "datasets", "download", "-d", dataset_ref, "--unzip", "-p", output_dir ] result = subprocess.run(cmd, capture_output=True, text=True) if result.returncode == 0: print(f"Downloaded {dataset_ref} to {output_dir}") else: print(f"Error: {result.stderr}") # Search for NLP benchmark datasets datasets = search_kaggle_datasets("nlp text classification benchmark") for ds in datasets[:5]: print(f" {ds.get('ref', 'N/A')}") print(f" Size: {ds.get('totalBytes', 'N/A')} bytes") print(f" Votes: {ds.get('voteCount', 'N/A')}") print()
pythonfrom kaggle.api.kaggle_api_extended import KaggleApi api = KaggleApi() api.authenticate() # Search datasets datasets = api.dataset_list(search="genomics", sort_by="updated") for ds in datasets[:5]: print(f"{ds.ref}: {ds.title} ({ds.size})") # Get dataset metadata metadata = api.dataset_view("nih-chest-xrays/data") print(f"Title: {metadata.title}") print(f"Size: {metadata.totalBytes}") print(f"Description: {metadata.description[:200]}") # Download dataset files api.dataset_download_files( "nih-chest-xrays/sample", path="./data/chest-xrays/", unzip=True )
Benchmark Dataset Access: Download well-established datasets used in published research for reproducibility studies. Kaggle hosts canonical versions of many benchmark datasets referenced in ML papers.
Competition as Evaluation Framework: Use Kaggle competitions as standardized evaluation environments with leaderboards and held-out test sets. Submit predictions from novel methods to compare against state-of-the-art approaches.
Data Exploration Notebooks: Search for and pull community notebooks that explore datasets relevant to your research. These often contain valuable preprocessing code, exploratory analysis, and baseline models.
Collaborative Research Datasets: Upload processed research datasets to Kaggle for sharing with collaborators and the broader community, enabling others to reproduce and extend your work.
Cross-Domain Transfer: Search across Kaggle's diverse dataset collection to find datasets from adjacent domains that could be useful for transfer learning or cross-domain validation studies.
kernel-metadata.json file specifying the kernel type, language, and datasetskaggle.json to version control; use environment variables in CI/CD pipelines| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 3,407 | 26,307 | +672% | 1 | 1 | 0% | 714 | 2,447 | +243% | 0 | 0 | — |
case-02 | fail→pass | 35,510 | 2,710 | -92% | 1 | 1 | 0% | 1,007 | 2,350 | +133% | 0 | 0 | — |
case-03 | pass→pass | 5,666 | 5,368 | -5% | 1 | 1 | 0% | 1,056 | 2,879 | +173% | 0 | 0 | — |
case-04 | pass→pass | 4,267 | 4,067 | -5% | 1 | 1 | 0% | 765 | 2,395 | +213% | 0 | 0 | — |
case-05 | pass→pass | 2,761 | 2,048 | -26% | 1 | 1 | 0% | 437 | 2,210 | +406% | 0 | 0 | — |
case-18 | pass→pass | 23,601 | 12,370 | -48% | 1 | 1 | 0% | 2,466 | 4,004 | +62% | 0 | 0 | — |
case-06 | pass→pass | 6,027 | 2,750 | -54% | 1 | 1 | 0% | 920 | 2,417 | +163% | 0 | 0 | — |
case-07 | pass→pass | 6,077 | 3,935 | -35% | 1 | 1 | 0% | 878 | 2,609 | +197% | 0 | 0 | — |
case-08 | pass→pass | 3,679 | 2,886 | -22% | 1 | 1 | 0% | 661 | 2,362 | +257% | 0 | 0 | — |
case-09 | pass→pass | 3,888 | 2,910 | -25% | 1 | 1 | 0% | 843 | 2,369 | +181% | 0 | 0 | — |
case-10 | pass→pass | 7,068 | 3,328 | -53% | 1 | 1 | 0% | 1,279 | 2,513 | +96% | 0 | 0 | — |
case-11 | pass→pass | 4,526 | 2,624 | -42% | 1 | 1 | 0% | 842 | 2,278 | +171% | 0 | 0 | — |
case-12 | pass→pass | 3,911 | 4,221 | +8% | 1 | 1 | 0% | 686 | 2,479 | +261% | 0 | 0 | — |
case-13 | pass→pass | 7,042 | 4,512 | -36% | 1 | 1 | 0% | 1,386 | 2,572 | +86% | 0 | 0 | — |
case-19 | fail→pass | 17,217 | 4,241 | -75% | 1 | 1 | 0% | 945 | 2,431 | +157% | 0 | 0 | — |
case-14 | pass→pass | 7,086 | 2,936 | -59% | 1 | 1 | 0% | 1,142 | 2,379 | +108% | 0 | 0 | — |
case-15 | pass→pass | 3,881 | 3,146 | -19% | 1 | 1 | 0% | 612 | 2,353 | +284% | 0 | 0 | — |
case-16 | pass→pass | 5,343 | 4,103 | -23% | 1 | 1 | 0% | 893 | 2,558 | +186% | 0 | 0 | — |
case-17 | pass→pass | 14,078 | 8,430 | -40% | 1 | 1 | 0% | 2,306 | 3,210 | +39% | 0 | 0 | — |
case-20 | pass→pass | 6,499 | 3,873 | -40% | 1 | 1 | 0% | 1,180 | 2,474 | +110% | 0 | 0 | — |
case-21 | pass→pass | 7,852 | 5,130 | -35% | 1 | 1 | 0% | 1,194 | 2,703 | +126% | 0 | 0 | — |
case-22 | pass→pass | 11,552 | 10,318 | -11% | 1 | 1 | 0% | 1,905 | 3,941 | +107% | 0 | 0 | — |
case-23 | pass→pass | 8,368 | 7,415 | -11% | 1 | 1 | 0% | 1,345 | 3,094 | +130% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.