Skip to main content

migrate-catalog-flow

Guided catalog migration flow. Extracts Databricks Unity Catalog / HMS metadata, previews the 18-rule DDL rewrite, then batched replay on AIDP. Asks before the destructive replay step. Use when the user wants to migrate the catalog layer (schemas + tables) with checkpoints, rather than running aidp-migrate-catalog directly.

소스 정보

저장소
oracle-samples/oracle-aidp-samples
최근 소스 활동
2026년 6월 24일 07:21
감지된 SKILL.md 언어
영어
스타
47
포크
32

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
migrate-catalog-flow
description
Guided catalog migration flow. Extracts Databricks Unity Catalog / HMS metadata, previews the 18-rule DDL rewrite, then batched replay on AIDP. Asks before the destructive replay step. Use when the user wants to migrate the catalog layer (schemas + tables) with checkpoints, rather than running aidp-migrate-catalog directly.
# `migrate-catalog-flow` — guided catalog migration Walk the user through extracting Unity Catalog / HMS metadata from Databricks, previewing the rewritten DDL, and replaying it on AIDP. ## When to use - User says "migrate the catalog", "port the schemas", "migrate Unity Catalog DDL". - User wants the guided extract → dry-run → confirm → replay flow with stop points. - BEFORE [`migrate-job-flow`](../migrate-job-flow/SKILL.md) — schemas + tables must exist on AIDP first so notebook reads have targets. For a one-shot non-interactive run, prefer [`aidp-migrate-catalog`](../aidp-migrate-catalog/SKILL.md) directly. ## Workflow 1. **Confirm prereqs** — invoke [`aidp-migrator-bootstrap`](../aidp-migrator-bootstrap/SKILL.md). Especially check `DATABRICKS_HOST` + `DATABRICKS_TOKEN` are set (catalog extract needs them). 2. **Confirm scope** — ask which catalogs / schemas the user wants to migrate. Default to "everything in this catalog" but accept a filter list. 3. **Confirm bucket mapping** — if any external tables have `s3://` locations, the bucket-map config must exist. Route to [`aidp-bucket-mapping`](../aidp-bucket-mapping/SKILL.md) if missing. 4. **Stage 1: extract** — `extract_catalog_databricks.py` → `reports/catalog_pack.json`. Show the table count. 5. **Stage 2 dry-run: rewrite preview** — `migrate_catalog.py --dry-run`. Surface: - Total CREATE SCHEMA statements - Total CREATE TABLE statements - Any rejections (materialized views, streaming, unsupported) - Any bucket-map misses Ask the user "ready to replay on the cluster?" 6. **Stage 2 replay** — `migrate_catalog.py` (no `--dry-run`). Surface per-chunk status. 7. **Verify** — for each migrated schema, run `SHOW TABLES IN default.<schema>` and surface the count. Compare to extract. ## Args If the user supplied `<catalog>` or `<catalog>:<schema>` filters in their prompt, use them; else ask in step 2. ## Checkpoints ``` [Phase 1/4] Extracted N catalogs, M schemas, K tables → reports/catalog_pack.json [Phase 2/4] Dry-run: would create 23 schemas + 412 tables. Rejected: 3 MVs, 1 streaming table. About to run live DDL replay — proceed? (y/N) [Phase 3/4] Replayed in 4 chunks of 25 statements each. All chunks committed. [Phase 4/4] Verify: SHOW TABLES across each schema. Discrepancies: 0 ``` ## When to stop - User aborts at the dry-run checkpoint. - Bucket-map is missing buckets — fix first via [`aidp-bucket-mapping`](../aidp-bucket-mapping/SKILL.md). - Dry-run shows >10% rejected (MVs / streaming) — likely a structural mismatch; review with the user before proceeding. ## After this - Verify with [`aidp-check-data`](../aidp-check-data/SKILL.md) — schemas + tables should now resolve. - Proceed to [`migrate-job-flow`](../migrate-job-flow/SKILL.md) for the notebook layer.
GitHub에서 보기