Create and validate Earth2Studio data source wrappers (DataSource, ForecastSource, DataFrameSource, ForecastFrameSource) from remote stores. May add Python dependencies to pyproject.toml as part of development. Do NOT use for fetching data with existing sources, model inference, or Earth2Studio installation/setup tasks.
Instrucciones de origen · Vista previa de solo lectura
name
earth2studio-create-datasource
metadata
{"version":"0.17.1","author":"NVIDIA Earth-2 Team <agent-skills@nvidia.com>","tags":["earth2studio","earth2","python","data-source","forecast-source","integration"]}
license
Apache-2.0
permissions
["network","filesystem","shell"]
description
Create and validate Earth2Studio data source wrappers (DataSource, ForecastSource, DataFrameSource, ForecastFrameSource) from remote stores. May add Python dependencies to pyproject.toml as part of development. Do NOT use for fetching data with existing sources, model inference, or Earth2Studio installation/setup tasks.
argument-hint
URL or description of remote data store (optional)
Create and Validate Data Source
Purpose
End-to-end workflow for implementing a new Earth2Studio data source wrapper
that connects a remote data store (S3, GCS, Azure, HTTP, HuggingFace) to
Earth2Studio's async data fetching infrastructure — from analysis through
implementation, testing, validation, and PR submission.
Prerequisites
Earth2Studio dev environment with ( must work)
uv
uv run python
Git configured with fork (origin) and upstream (upstream) remotes
Access to the target remote data store (credentials if private)
Python 3.10+
⚠️ Credential Safety: If the remote store requires authentication,
credentials must be passed via environment variables only. Never hard-code
secrets in source files. Before any git commit, verify no credentials are
staged (check git diff --cached for API keys, tokens, or passwords).
Workspace
Use the directory containing pyproject.toml. For Harbor evals, write to
/workspace/output/ preserving paths. Never read evals/targets/.
Instructions
Python Environment: Always use uv run python or the local .venv.
Never use the system Python directly.
Follow every step in order.
[CONFIRM] gates: Step 1 (Source Type), Step 12 (Sanity-Check Plots),
and Step 13 (Ready to Submit) require explicit user approval before
proceeding. All other [CONFIRM] markers are advisory — present
decisions inline and proceed without blocking.
Deliverables first: Write the source file and test file (Steps 6–7)
before extended exploration, documentation, registration, CHANGELOG, or PR
work. Skip Steps 8–14 when the user asks for implementation only.
Before you finish: Run verification commands in the repo root so results
appear in the session log:
uv run pytest test/data/test_<source>.py -x
make format && make lint
Be concise: Avoid long architecture reports; summarize decisions in a
few sentences and move on to file writes.
Hangs or User Feedback If agent becomes stuck or user provides a
correction during this skills use, conservatively review relevant part of
the skill and improve. Be concise.
One source type per invocation. Invoke again for companion types.
Reference Files
Load these on demand during the relevant steps:
File
Content
Load at
references/reference-implementation.py
Obstore-backed source skeleton with FILL comments
Steps 6–10
references/reference-lexicon.py
Lexicon skeleton with FILL comments
Step 4
references/testing-guide.py
Test skeleton with FILL comments
Step 11
references/validation-guide.md
Plot templates, PR body template, Greptile handling
If $ARGUMENTS is provided, use it (URL → WebFetch; file path → read).
If empty, ask:
Please provide a URL, API documentation link, or description of the
remote data store. This will be used to understand storage format,
access pattern, variable inventory, temporal/spatial resolution.
Step 1 — Determine Source Type
Protocol
Returns
Has lead_time?
Use
DataSource
xr.DataArray
No
Gridded analysis/reanalysis
ForecastSource
xr.DataArray
Yes
Gridded forecast
DataFrameSource
pd.DataFrame
No
Sparse/station obs
ForecastFrameSource
pd.DataFrame
Yes
Sparse forecast obs
Key factors: gridded vs sparse → DataArray vs DataFrame; analysis vs forecast → Source vs ForecastSource.
[CONFIRM — Source Type]
Present recommended type with justification. Ask for confirmation.
Step 2 — Examine Remote Store & Propose Dependencies
Add the class alphabetically to the appropriate autosummary block in
docs/modules/datasources_analysis.md, datasources_forecast.md, or
datasources_dataframe.md
Keep the existing region, dataclass, and product badges accurate. Add
provider:<name> when the service provider is known, and add
dataset:<family> only when the source belongs to a named dataset family;
do not infer a family from the provider. Use dataset:gfs for GEFS sources.
Reuse lowercase badge keys already defined in mkdocs.yml. If a new
provider or dataset family is required, add its badge definition there and
use hide_in: [autosummary, filter]; the catalog reads these badges even
though data-source API tables do not display or filter on them.
Add entry under the current unreleased version. See
references/reference-implementation.py REGISTRATION CHECKLIST for the format.
One line per source. Do NOT add separate lexicon entries.
Step 11 — Verify Style & Expand Tests
Run make format && make lint && make license. Load references/testing-guide.py
for test skeletons. Required tests: test_<source>_fetch (slow), _cache (slow),
_call_mock, _exceptions, _available. Target 90%+ coverage with --slow.
[CONFIRM — Tests]
Present test file, functions, coverage.
Step 12 — Validate Variables & Sanity-Check
Validate all lexicon vars against real data (run script, do NOT commit)
Remove variables with < 10% valid data
Create sanity-check plot (gridded or sparse template)
Tell user the plot path and ask for visual confirmation
[CONFIRM — Sanity-Check Plots]
User MUST visually inspect plots. Do not proceed without confirmation.
Step 13 — Branch, Commit & Open PR
⚠️ WARNING: This step performs irreversible git operations (push, PR
creation). Do NOT proceed without explicit user confirmation. Present all
planned git commands to the user and wait for approval.
Create branch feat/data-source-<name>
Commit (do NOT add sanity-check script/images)
Present full command list to user for review
Push to fork (only after user approves)
gh pr create --repo NVIDIA/earth2studio (only after user approves)
Immediately post sanity-check validation as PR comment with: