Skip to main content

data-engineering

Stars8
Forks1
UpdatedJune 6, 2026 at 14:54

Use this when: build a data pipeline, my pipeline is not idempotent, clean messy data, convert CSV to Parquet, my data has duplicates, validate schema at ingestion, pipeline fails on re-run, process files larger than memory, schedule a recurring job, upstream schema changed and broke my pipeline, migrate data between systems, query Parquet without loading it, deduplicate records, batch vs stream processing, DuckDB for analytics, choose an orchestrator, slow pandas pipeline

Installation

Install with Codex or Claude Copy this prompt, paste it into Codex, Claude, or another assistant, and let it review the skill page and install it for you.

SKILL.md
readonly