Skip to main content

data-pipeline-quality

Automated data quality checks for pipelines. Testing pyramids, dbt test patterns, data contracts, circuit breakers, and monitoring. Use when implementing data quality checks, writing dbt tests, defining data contracts, setting up pipeline validation, building automated quality monitoring, or when someone asks "how do I test my data pipeline?" For quality scoring methodology (the 5-dimension rubric), see data-quality-assessment.

跳到安装

来源信息

仓库
hollandkevint/data-product-operator
最近来源活动
2026年9月11日 01:52
检测到的 SKILL.md 语言
英语
星标
3
分支
1

安装方式

默认使用会先检查来源的 Prompt;你也可以切换为直接命令,或下载本地副本。

检查来源文件

决定是否安装前,请先阅读 SKILL.md,以及 SkillsMP 当前展示的配套文件。

正在显示 SKILL.md

SKILL.md
来源说明 · 只读预览
name
data-pipeline-quality
version
0.1.1
description
Automated data quality checks for pipelines. Testing pyramids, dbt test patterns, data contracts, circuit breakers, and monitoring. Use when implementing data quality checks, writing dbt tests, defining data contracts, setting up pipeline validation, building automated quality monitoring, or when someone asks "how do I test my data pipeline?" For quality scoring methodology (the 5-dimension rubric), see data-quality-assessment.
user-invocable
false
## Testing Pyramid for Data Run tests in this order. Cheapest and fastest first: | Layer | What It Catches | Examples | |-------|----------------|---------| | Schema tests (run first) | Structural failures | Column types, not-null, uniqueness, accepted values | | Business rule tests | Logic errors | Cross-field validation, referential integrity, range checks | | Integration tests (run last) | System-level drift | Cross-system reconciliation, end-to-end row counts | Schema tests are cheap. Run them on every pipeline execution. Business rule tests are mid-tier — run them on staging and production. Integration tests are expensive — run them on a schedule (daily or pre-release). ## dbt Test Patterns **Generic tests** for reusable checks. Apply across models: ```yaml models: - name: fct_encounters columns: - name: encounter_id tests: [not_null, unique] - name: encounter_type tests: - accepted_values: values: ['inpatient', 'outpatient', 'emergency', 'observation'] - name: patient_id tests: - relationships: to: ref('dim_patient') field: patient_id ``` **Custom generic test** for row count tolerance: ```sql {% test row_count_within_tolerance(model, min_count, max_count) %} select count(*) as row_count from {{ model }} having count(*) < {{ min_count }} or count(*) > {{ max_count }} {% endtest %} ``` **Singular tests** for business logic specific to one model. Use singular tests when the logic doesn't generalize. ## Data Contracts A data contract is a product spec for your data. It defines what consumers can depend on. ```yaml contract: name: fct_encounters version: 2 owner: data-platform-team sla: freshness: "< 4 hours from source update" completeness: ">= 99.5% of expected rows" accuracy: ">= 99.9% match to source of record" schema: encounter_id: {type: bigint, nullable: false, unique: true} patient_id: {type: bigint, nullable: false} encounter_date: {type: date, nullable: false} ``` **Producer responsibilities**: Meet the SLA, notify consumers before breaking changes, version the schema. **Consumer expectations**: Query only contracted fields, respect the grain, report quality issues. ## Circuit Breakers Wire quality gates into pipeline stages. When a check fails, block the pipeline and alert. - **Schema violation**: Block immediately. Bad schema corrupts everything downstream. - **Row count outside tolerance**: Block and alert. Investigate before proceeding. - **Freshness SLA breach**: Alert the on-call. Don't block unless downstream consumers can't tolerate stale data. - **Business rule failure**: Block if the failure rate exceeds threshold (e.g., >1% of rows). Alert if below. Cross-reference `data-quality-assessment` for the 4-stage quality model (detect → assess → respond → prevent). CRITICAL: Never auto-heal data quality issues in production. Alert, block, investigate. Auto-fixes mask root causes and erode trust faster than stale data. ## Monitoring Track quality over time, not just point-in-time pass/fail: - **Row count trends**: Compare expected and observed counts within the same scope. Abrupt or gradual changes suggest checks; their shape does not establish source or pipeline root cause. Test competing explanations. - **Freshness**: Separate source-event time, arrival time and processing time against the consumer's requirement. A successful run can process stale or incomplete inputs. Track the relevant lag distribution and coverage. - **Anomaly detection**: Choose statistical or fixed thresholds for the risk, seasonality and available history. Confirm expected ranges with the owner; a statistical alert does not replace an acceptance rule. ## Cross-References For quality scoring methodology (the 5-dimension rubric and maturity model), see `data-quality-assessment`. This skill covers how to automate those checks in your pipeline.
在 GitHub 查看