| name | data-catalog-and-discovery |
| description | Guides agents through data catalog, discovery, and metadata quality workflows. Use when publishing datasets, improving discoverability, curating lineage metadata, or making data products easier for other teams to find and trust. |
| metadata | {"category":"data","source":{"repository":"https://github.com/vaquarkhan/data-engineering-agent-skills","path":"skills/data-catalog-and-discovery","license_path":"LICENSE","commit":"421ef57e8d42c464b29339193c18dd5bd2946bc2"}} |
Data Catalog And Discovery
Overview
Use this skill when the challenge is not only building data, but making it understandable and discoverable. It helps agents treat metadata, ownership, lineage, and usage context as delivery artifacts instead of afterthoughts.
When to Use
- publishing a new shared dataset
- improving catalog metadata quality
- curating lineage, tags, or ownership information
- reducing duplicate datasets created because teams cannot find trusted ones
Do not stop at filling in a title and description. Discovery quality requires operational context too.
Workflow
-
Define the discovery contract.
Include:
- owner
- business description
- technical description
- grain
- freshness expectation
- intended consumers
-
Link the asset to its lineage.
Show upstream sources, transformation layers, and major downstream uses where possible.
-
Add trust signals.
Typical signals:
- quality status
- SLA or freshness status
- certification or review state