Skip to main content

synthetic-data-workflow

Privacy-preserving workflow for building synthetic datasets and data source skills when raw data must never enter the DAAF container. The user profiles their sensitive data locally with a disclosure-controlled script; only a summary profile report crosses the boundary; DAAF builds a synthetic dataset and skill from the report alone. Use whenever data is sensitive, proprietary, PII-bearing, HIPAA/FERPA-governed, held in a secure enclave, or the user says the data cannot leave their environment, they cannot upload it, or asks to profile it locally. Covers a four-tier disclosure ladder (T1 schema, T2 marginals, T3 relationships, T4 local high-fidelity synthesis), profile-only generation with simstudy (R) / NumPy-SciPy copulas (Python), and three-part QA (disclosure-safety, report consistency, synthetic-vs-profile validation). Not for synthetic control, the causal-inference method; for that see data-scientist. Synthetic data here is a code-development scaffold, not an analytic substitute.

インストールへ移動

ソース情報

リポジトリ
DAAF-Contribution-Community/daaf
ソースの最終更新活動
2026年8月3日 18:21
検出された SKILL.md の言語
英語
スター
232
フォーク
34

インストール方法

デフォルトでは、最初にソースを確認する Prompt が選択されています。直接コマンドに切り替えるか、ローカルコピーをダウンロードすることもできます。

ソースファイルを確認

インストールを決める前に、SKILL.md と SkillsMP に表示されている付属ファイルをお読みください。