Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Access harmonized census and survey microdata via the IPUMS API
.claude/skills/brycewang-stanford-ipums-microdata-api/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 61% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 53% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 107% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 25% | 0% |
IPUMS (Integrated Public Use Microdata Series) provides the world's largest collection of harmonized census and survey microdata. Hosted by the University of Minnesota, it covers demographic, health, labor, and geographic data across 100+ countries and 100+ years. The API enables programmatic extract creation, metadata queries, and data retrieval. Free registration required.
| Collection | Coverage | Records | |-----------|----------|---------| | IPUMS USA | U.S. Census & ACS (1850-present) | 16B+ person-records | | IPUMS CPS | Current Population Survey (1962-present) | Labor force data | | IPUMS International | Census data from 100+ countries | 2B+ person-records | | IPUMS NHGIS | U.S. geographic/aggregate data | County-level stats | | IPUMS DHS | Demographic and Health Surveys | 300+ surveys, 90 countries | | IPUMS Time Use | American Time Use Survey | Time diary data | | IPUMS Health | NHIS health surveys | Health/disability data | | IPUMS Higher Ed | NSCG/SDR science workforce | S&E workforce data |
https://api.ipums.org/extracts/bash# Register at https://www.ipums.org/ # API key from your account settings export IPUMS_KEY="..."
bash# Request a data extract (IPUMS USA example) curl -X POST "https://api.ipums.org/extracts/?collection=usa&version=2" \ -H "Authorization: $IPUMS_KEY" \ -H "Content-Type: application/json" \ -d '{ "description": "Income by education, 2020 ACS", "data_structure": {"rectangular": {"on": "P"}}, "data_format": "csv", "samples": {"us2020a": {}}, "variables": { "AGE": {}, "SEX": {}, "RACE": {}, "EDUC": {}, "INCTOT": {}, "EMPSTAT": {} } }'
bashcurl "https://api.ipums.org/extracts/42?collection=usa&version=2" \ -H "Authorization: $IPUMS_KEY"
bash# When status is "completed" curl -O "https://api.ipums.org/extracts/42/download?collection=usa&version=2" \ -H "Authorization: $IPUMS_KEY"
bash# List available variables curl "https://api.ipums.org/metadata/usa/variables?version=2" \ -H "Authorization: $IPUMS_KEY" # Get variable details curl "https://api.ipums.org/metadata/usa/variables/EDUC?version=2" \ -H "Authorization: $IPUMS_KEY" # List available samples curl "https://api.ipums.org/metadata/usa/samples?version=2" \ -H "Authorization: $IPUMS_KEY"
pythonimport os import time import requests BASE_URL = "https://api.ipums.org" HEADERS = {"Authorization": os.environ.get("IPUMS_KEY", "")} def create_extract(collection: str, samples: dict, variables: list, description: str = "", data_format: str = "csv") -> int: """Create an IPUMS data extract request.""" var_dict = {v: {} for v in variables} body = { "description": description, "data_format": data_format, "data_structure": {"rectangular": {"on": "P"}}, "samples": {s: {} for s in samples} if isinstance(samples, list) else samples, "variables": var_dict, } resp = requests.post( f"{BASE_URL}/extracts/?collection={collection}&version=2", headers={**HEADERS, "Content-Type": "application/json"}, json=body, ) resp.raise_for_status() return resp.json()["number"] def wait_for_extract(extract_id: int, collection: str, poll_interval: int = 30) -> str: """Poll until extract is ready, return download URL.""" while True: resp = requests.get( f"{BASE_URL}/extracts/{extract_id}" f"?collection={collection}&version=2", headers=HEADERS, ) resp.raise_for_status() data = resp.json() status = data.get("status") if status == "completed": return data["download_links"]["data"]["url"] elif status == "failed": raise RuntimeError(f"Extract failed: {data}") print(f"Status: {status}, waiting {poll_interval}s...") time.sleep(poll_interval) def get_variable_info(collection: str, variable: str) -> dict: """Get metadata about a variable.""" resp = requests.get( f"{BASE_URL}/metadata/{collection}/variables/{variable}" f"?version=2", headers=HEADERS, ) resp.raise_for_status() return resp.json() # Example: request 2020 ACS income data extract_id = create_extract( collection="usa", samples=["us2020a"], variables=["AGE", "SEX", "RACE", "EDUC", "INCTOT", "EMPSTAT"], description="Education-income analysis 2020", ) print(f"Extract #{extract_id} submitted. Waiting...") download_url = wait_for_extract(extract_id, "usa") print(f"Ready: {download_url}")
| Variable | Description | |----------|-------------| | AGE | Age | | SEX | Sex | | RACE | Race | | EDUC | Education level | | INCTOT | Total income | | EMPSTAT | Employment status | | OCC | Occupation | | IND | Industry | | POVERTY | Poverty status | | MIGRATE1 | Migration status | | MARST | Marital status | | NCHILD | Number of children |
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 12,251 | 12,891 | +5% | 1 | 1 | 0% | 2,212 | 3,571 | +61% | 0 | 0 | — |
case-02 | pass→pass | 16,957 | 16,348 | -4% | 1 | 1 | 0% | 3,349 | 5,321 | +59% | 0 | 0 | — |
case-03 | fail→fail | 10,003 | 7,234 | -28% | 1 | 1 | 0% | 1,772 | 2,923 | +65% | 0 | 0 | — |
case-04 | fail→pass | 12,995 | 9,764 | -25% | 1 | 1 | 0% | 2,342 | 3,586 | +53% | 0 | 0 | — |
case-05 | fail→pass | 6,787 | 3,028 | -55% | 1 | 1 | 0% | 1,137 | 2,353 | +107% | 0 | 0 | — |
case-06 | fail→pass | 15,277 | 5,650 | -63% | 1 | 1 | 0% | 2,534 | 2,678 | +6% | 0 | 0 | — |
case-07 | pass→pass | 12,641 | 7,987 | -37% | 1 | 1 | 0% | 2,303 | 3,255 | +41% | 0 | 0 | — |
case-08 | fail→fail | 15,336 | 15,381 | +0% | 1 | 1 | 0% | 2,400 | 4,617 | +92% | 0 | 0 | — |
case-09 | fail→pass | 22,186 | 7,963 | -64% | 1 | 1 | 0% | 2,644 | 3,318 | +25% | 0 | 0 | — |
case-10 | pass→pass | 3,630 | 5,027 | +38% | 1 | 1 | 0% | 587 | 2,670 | +355% | 0 | 0 | — |
case-11 | pass→pass | 6,063 | 3,912 | -35% | 1 | 1 | 0% | 1,055 | 2,394 | +127% | 0 | 0 | — |
case-12 | pass→pass | 5,960 | 3,804 | -36% | 1 | 1 | 0% | 928 | 2,333 | +151% | 0 | 0 | — |
case-13 | pass→pass | 6,500 | 3,566 | -45% | 1 | 1 | 0% | 867 | 2,316 | +167% | 0 | 0 | — |
case-14 | pass→pass | 4,904 | 4,369 | -11% | 1 | 1 | 0% | 741 | 2,365 | +219% | 0 | 0 | — |
case-15 | pass→pass | 3,546 | 2,355 | -34% | 1 | 1 | 0% | 569 | 2,091 | +267% | 0 | 0 | — |
case-16 | pass→pass | 3,116 | 2,983 | -4% | 1 | 1 | 0% | 483 | 2,213 | +358% | 0 | 0 | — |
case-17 | pass→pass | 8,775 | 4,876 | -44% | 1 | 1 | 0% | 1,193 | 2,591 | +117% | 0 | 0 | — |
case-18 | pass→pass | 3,272 | 2,752 | -16% | 1 | 1 | 0% | 501 | 2,162 | +332% | 0 | 0 | — |
case-19 | fail→pass | 2,877 | 1,966 | -32% | 1 | 1 | 0% | 376 | 2,005 | +433% | 0 | 0 | — |
case-20 | pass→pass | 5,304 | 2,516 | -53% | 1 | 1 | 0% | 868 | 2,052 | +136% | 0 | 0 | — |
case-21 | pass→pass | 9,493 | 11,088 | +17% | 1 | 1 | 0% | 1,418 | 3,583 | +153% | 0 | 0 | — |
case-22 | pass→pass | 7,090 | 6,328 | -11% | 1 | 1 | 0% | 1,377 | 2,965 | +115% | 0 | 0 | — |
case-23 | pass→pass | 9,547 | 10,235 | +7% | 1 | 1 | 0% | 1,752 | 3,677 | +110% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of +26 percentage points is the difference between those two pass rates over the 23 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.