Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Fetch weather/climate data via Earth2Studio data sources for specific variables and times. Do NOT use for inference pipelines, model discovery, or installation.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 90% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 60% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 55% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 149% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 59% | 0% |
Guide a user through downloading weather/climate data via Earth2Studio data source APIs. Identifies compatible sources by checking the lexicon, verifies variable support, and produces a working fetch script outputting an xarray DataArray.
uv pip install earth2studio or equivalent)~/.cdsapirc)You are helping a user download specific weather/climate data using Earth2Studio's data source APIs. Your job is to identify which data source(s) can provide the requested variables, verify compatibility via the lexicon system, and produce a working fetch script.
Data source APIs, available variables, and the lexicon evolve between releases. Before recommending a data source or writing a fetch script:
and constructor arguments.
that data source.
Live doc references (fetch only what the user's request requires):
<https://nvidia.github.io/earth2studio/modules/datasources_analysis.html>
<https://nvidia.github.io/earth2studio/modules/datasources_forecast.html>
<https://nvidia.github.io/earth2studio/modules/datasources_dataframe.html>
<https://github.com/NVIDIA/earth2studio/blob/main/earth2studio/lexicon/base.py>
<https://github.com/NVIDIA/earth2studio/tree/main/earth2studio/lexicon>
Extract from what the user has said (ask follow-ups if needed, cap at 3 questions):
(e.g. t2m, u500, z850, tp, msl). If the user uses plain language ("500 hPa geopotential height"), map it to the E2Studio name by checking the live base.py E2STUDIO_VOCAB.
discrete times?
Based on the request type, narrow candidates:
Analysis/reanalysis (historical state at a specific time):
IFS/IFS_ENS (ECMWF), ARCO/CDS/WB2ERA5/NCAR_ERA5 (ERA5 reanalysis), GOES/MRMS/JPSS (observational)
Forecast (predictions from an initialization time with lead times):
AIFS_FX, CFS_FX
Key differentiators to surface:
history; reanalysis (ERA5 via ARCO/CDS/WB2) goes back decades
WB2ERA5_32x64 is 5.625° global
This is critical. Each data source has a lexicon file that defines which E2Studio variables it can provide.
To verify:
https://github.com/NVIDIA/earth2studio/blob/main/earth2studio/lexicon/<source>.py (e.g. gfs.py, hrrr.py, cds.py, arco.py, wb2.py)
source's VOCAB dict
it — try another
The lexicon VOCAB maps Earth2Studio variable names → source-specific identifiers. If a variable key exists in the VOCAB, the source supports it.
Present the results clearly: "GFS supports `t2m`, `u500`, `z850`. HRRR also supports these but is limited to North America. ARCO (ERA5) supports all three and has data back to 1959."
Present the viable options with tradeoffs:
| Source | Variables | Coverage | Resolution | Time Range | |--------|-----------|----------|------------|------------| | ... | ... | ... | ... | ... |
Let the user pick. If there's one obvious choice, recommend it and ask for confirmation.
Write a Python script that uses the selected data source to fetch the requested data. The script structure depends on whether it's an analysis or forecast source.
Analysis source pattern:
pythonimport datetime from earth2studio.data import <SourceClass> # Initialize data source ds = <SourceClass>() # Fetch data # Analysis sources use: ds(time, variable) -> xr.DataArray time = [datetime.datetime(YYYY, M, D, H)] # or array of times variable = ["var1", "var2"] # E2Studio variable names data = ds(time, variable)
Forecast source pattern:
pythonimport datetime from earth2studio.data import <SourceClass> # Initialize data source ds = <SourceClass>() # Forecast sources use: ds(time, lead_time, variable) -> xr.DataArray time = [datetime.datetime(YYYY, M, D, H)] # initialization time lead_time = [datetime.timedelta(hours=H)] # or array of lead times variable = ["var1", "var2"] data = ds(time, lead_time, variable)
Always fetch the specific data source's API doc page to confirm the exact constructor arguments and call signature before writing the script — they can vary (some need auth tokens, cache paths, specific parameters).
Include in the script:
print(data), data.shape, data.coords)After delivering the script, mention:
discover skill
EARTH2STUDIO_CACHE)
Owns: identifying data sources for a user's variable/time request, verifying variable support via lexicon, generating data fetch scripts, explaining analysis vs. forecast source differences.
Does not own: installation (earth2studio-install), model selection (earth2studio-discover), inference pipelines, custom data source creation (point to extend examples), data source authentication setup beyond what the docs describe.
Typical invocation:
> "I need 500 hPa geopotential height and 2m temperature from ERA5 > for January 1, 2020 at 00Z."
The skill would:
z500, t2m(GCS, S3, CDS API)
DataArrayFile/DataSetFile directly
sources in a single call
variables; always verify via lexicon
are generally faster
| Error | Cause | Solution | |-------|-------|----------| | KeyError: '<var>' | Not in lexicon | Check lexicon; try another source | | FileNotFoundError / 404 | Time not available | Verify temporal coverage | | CDS API timeout | Queue congestion | Retry or use ARCO for ERA5 | | ModuleNotFoundError | Not installed | uv pip install earth2studio | | Empty DataArray | Time/var mismatch | Check datetime and variable name |
Other measured skills in the registry, with their headline benchmark lift.