Skip to main content

databricks-autonomous-operations

End-to-end autonomous deployment and operations skill. Deploys Databricks Asset Bundles, runs jobs and pipelines, polls for completion, diagnoses failures, applies fixes, redeploys, and verifies — all without human intervention. Also serves as SDK/CLI/REST API reference. Use when deploying bundles, running jobs, monitoring pipelines, troubleshooting ANY failure, or operating in a self-healing deploy-fix-redeploy cycle. Triggers on "deploy", "bundle deploy", "run job", "run pipeline", "make it work", "job failed", "troubleshoot", "fix and redeploy".

설치로 이동

소스 정보

저장소
databricks-solutions/vibe-coding-workshop-template
최근 소스 활동
2026년 8월 31일 04:03
감지된 SKILL.md 언어
영어
스타
6
포크
8

설치 방법

기본적으로 소스를 먼저 확인하는 Prompt가 선택됩니다. 직접 명령으로 전환하거나 로컬 사본을 다운로드할 수도 있습니다.

소스 파일 검토

설치 여부를 결정하기 전에 SKILL.md와 SkillsMP에 표시된 보조 파일을 읽어 보세요.

파일 탐색기
15 개 파일

SKILL.md 표시 중

SKILL.md
소스 지침 · 읽기 전용 미리보기
name
databricks-autonomous-operations
description
End-to-end autonomous deployment and operations skill. Deploys Databricks Asset Bundles, runs jobs and pipelines, polls for completion, diagnoses failures, applies fixes, redeploys, and verifies — all without human intervention. Also serves as SDK/CLI/REST API reference. Use when deploying bundles, running jobs, monitoring pipelines, troubleshooting ANY failure, or operating in a self-healing deploy-fix-redeploy cycle. Triggers on "deploy", "bundle deploy", "run job", "run pipeline", "make it work", "job failed", "troubleshoot", "fix and redeploy".
metadata
{"author":"prashanth subrahmanyam","version":"3.2","domain":"operations","role":"shared","used_by_stages":[1,2,3,4,5,6,7,8,9],"called_by":["bronze/00-bronze-layer-setup","silver/00-silver-layer-setup","gold/01-gold-layer-setup","semantic-layer/00-semantic-layer-setup","monitoring/00-observability-setup","ml/00-ml-pipeline-setup","genai-agents/00-course-orchestrator"],"dependencies":["admin/self-improvement"],"triggers":["deploy","bundle deploy","bundle run","run job","run pipeline","make it work","deploy and run","deploy and fix","job failed","monitor job","troubleshoot","get-run","run output","task failed","redeploy","job status","pipeline failed","DLT error","cluster issue","monitor failed","alert failed","deploy failed","databricks-sdk","databricks-connect","CLI","INTERNAL_ERROR","UNRESOLVED_COLUMN","ModuleNotFoundError","TABLE_OR_VIEW_NOT_FOUND","ResourceAlreadyExists","PERMISSION_DENIED","self-heal","[Truncated]"],"last_verified":"2026-06-02","volatility":"medium","clients":["ide_cli","genie_code"],"deploy_verb":"bundle deploy --target dev","deploy_note":"operations/CLI reference — on Genie Code every databricks command routes via runDatabricksCli FROM THE BUNDLE EDITOR (dp_bundle_root); a blocked bundle deploy/run is a page-context signal, NEVER substitute SDK/REST creation (RULE_10); see genie-code-environment","coverage":"all_stages","upstream_sources":[{"name":"databricks-agent-skills","repo":"databricks/databricks-agent-skills","paths":"[Truncated]","relationship":"extended","last_synced":"2026-08-30","sync_commit":"ca92a6c"}]}
# Databricks Autonomous Operations ## 1. Overview > **Context Loading Rule:** Load this skill at **deployment time**, not during planning. > During earlier phases, retain only this 3-step summary in working memory: > `Deploy (bundle deploy) → Poll (get-run / pipelines get) → Verify (get-run-output)`. > Full skill activation should happen when the first `bundle deploy` or `bundle run` is invoked. This skill is both an **SDK/CLI/Connect reference** and an **autonomous operations playbook**. It teaches the AI agent to operate as an SRE — independently deploying, monitoring, diagnosing failures, applying fixes, redeploying, and verifying results across all Databricks resource types. **Core Loop:** Deploy → Poll → Diagnose → Fix → Redeploy → Verify (max 3 iterations before escalation to user). ### When to Activate This Skill - Deploying or running Databricks Asset Bundles (`bundle deploy`, `bundle run`) - Monitoring job or pipeline runs for completion - Troubleshooting ANY failure: jobs, DLT pipelines, monitors, alerts, clusters, Genie Spaces - Using the Databricks Python SDK, CLI, Connect, or REST API - Encountering error messages from Databricks services - Operating in a self-healing deploy-fix-redeploy cycle - Checking job/task/pipeline status or retrieving run output **SDK Docs:** https://databricks-sdk-py.readthedocs.io/en/latest/ **GitHub:** https://github.com/databricks/databricks-sdk-py **Runnable Examples:** See `examples/` directory for complete, copy-pasteable scripts: - `examples/1-authentication.py` — All auth patterns (env vars, profiles, Azure SP, AccountClient, notebook context) - `examples/2-clusters-and-jobs.py` — Cluster auto-selection, autoscaling, job CRUD, submit_and_wait, run_now_and_wait - `examples/3-sql-and-warehouses.py` — Parameterized queries, chunked results, query_to_dataframe helper - `examples/4-unity-catalog.py` — Tables, schemas, catalogs, volumes, file operations, pattern matching - `examples/5-serving-and-vector-search.py` — Endpoint creation with TrafficConfig, chat/embedding queries, vector search with filters - `examples/6-autonomous-operations.py` — Job monitoring, multi-task output retrieval, failure diagnosis, self-healing loop --- ## 2. Environment & Authentication > **Client routing (RULE_1/2/4 — read once, applies to every command in this skill).** This skill is the > CLI/SDK operations surface; the same operations run on both clients via different channels — **deploy > mechanics are owned by `databricks-asset-bundles` (the spine); reference it rather than re-deriving them.** > - **IDE (Cursor):** the local `databricks` CLI; auth via `databricks auth login` / `~/.databrickscfg`; > `databricks-connect` is available for local Spark. > - **Genie Code has *three* execution paths — try them in order; "blocked on one path ≠ impossible":** > 1. **`runDatabricksCli`** — the allow-listed, pre-authenticated CLI path (no `auth login`). The path > for `bundle validate` / `summary` / `deploy --target dev` and read verbs. `--version`/`help`/ > `auth token`/`aitools`/`apps validate`/`apps manifest` are hard-blocked; `apps deploy` is > **unreliable here** (page-dependent + CWD-defeated). Use a `bundle validate` behavior probe instead > of a numeric `--version` compare. > 2. **Python SDK** (`WorkspaceClient` via `executeCode`) — the **most capable** path: it **bypasses the > CLI allow-list** and is the reliable way to `w.apps.deploy(...)`, retrieve `w.config.token`, and poll > deployment/run state. **Caveat:** the SDK has **no bundle-deploy equivalent** (`bundle deploy` is a > composite client-side operation) — keep `bundle deploy` on `runDatabricksCli`. > 3. **Native tools** (`createAsset`/`readTable`/…) for governed asset operations. > > No local Spark on Genie Code (use workspace **serverless**; `databricks-connect` is IDE-only). Full > allow-list / CWD / FUSE / escape-hatch detail is in the **`genie-code-environment`** skill — load it on > demand. The CLI command examples below run via path 1 on Genie Code (local shell on the IDE); where a > verb is CLI-blocked, reach for the SDK (path 2). > > **⛔ The carve-out the "fall to the SDK" rule does NOT cover — `bundle deploy`/`run` and resource creation.** > "Blocked ≠ impossible, try the SDK" is for **read-only / ad-hoc** ops (polling, inspecting schemas, lineage, > token retrieval, `apps deploy`). It is **NOT** a license to substitute the SDK/REST for the bundle. When > `bundle deploy`/`run` is blocked, that is a **page-context signal, not a dead end**: the verb is gated to the > **bundle-folder page**, which on Genie Code you reach by opening the **"Open in bundle editor"** affordance on > the `dp_bundle_root` folder (the bundle editor's CWD *is* the bundle root, where `validate`/`deploy`/`run` are > pre-approved). **FIELD-CONFIRMED:** the same `bundle deploy` that returned "blocked by safety guardrails" / > "`databricks.yml` not found" from a file/notebook page returned "Deployment complete!" and `bundle run … > SUCCESS` from the bundle editor. So the fix is **navigate to the bundle editor**, never `w.jobs.create()` / > `POST /api/2.1/jobs/create` / `POST /api/2.0/pipelines` / `CREATE TABLE` via `executeCode`. Creating > jobs/pipelines/tables directly is the **RULE_10 authoring-discipline violation** — it produces live, > un-versioned state that diverges from the bundle and is the exact regression this spine prevents. If > `bundle deploy`/`run` **still** fails *from the bundle editor*, STOP and report the blocker — the SDK/REST > creation route is an **escape hatch only on explicit operator authorization.** Surface a clickable > bundle-editor link: `file_id = w.workspace.get_status("<dp_bundle_root>/databricks.yml").object_id`, > `folder_id = w.workspace.get_status("<dp_bundle_root>").object_id`, link = > `{w.config.host}/editor/files/{file_id}?o={w.get_workspace_id()}&contextId=folder%3A{folder_id}`. Detail in > `genie-code-environment` §3/§8. ### Setup - SDK: `uv pip install databricks-sdk`; **Connect (IDE-only):** `uv pip install databricks-connect` — not used on Genie Code (serverless; no local Spark) - CLI version: **>= 0.278.0** on the IDE (`databricks --version`); on Genie Code the version is not introspectable (use a `bundle validate` behavior probe) - Config (IDE): `~/.databrickscfg` or env vars `DATABRICKS_HOST`, `DATABRICKS_TOKEN`; Genie Code is pre-authenticated in-session ### Quick Auth ```python from databricks.sdk import WorkspaceClient w = WorkspaceClient() # Auto-detect (env, config file, or notebook) w = WorkspaceClient(profile="MY_PROFILE") # Named profile ``` - **Token expired? (IDE path)** `databricks auth login --host <url> --profile <name>` — N/A on Genie Code (pre-authenticated) - **Profile-based CLI:** `DATABRICKS_CONFIG_PROFILE=<name> databricks <command>` - **Full auth patterns** (Azure SP, AccountClient, etc.): see `references/sdk-api-reference.md` --- ## 3. SDK API Quick Reference **Full code examples and patterns:** `references/sdk-api-reference.md` **Runnable scripts:** `examples/` directory (see file listing in Section 1) | API | Key Operations | Critical Notes | |-----|---------------|----------------| | **Clusters** | `list`, `get`, `create_and_wait`, `ensure_cluster_is_running`, `start`, `stop` | Use `select_spark_version()` and `select_node_type()` | | **Jobs** | `list`, `run_now_and_wait`, `submit_and_wait`, `get_run`, `get_run_output`, `cancel_run` | `get_run_output` needs **task** `run_id`, not parent job `run_id` | | **SQL Execution** | `execute_statement` with `wait_timeout` | Use `StatementParameterListItem` for parameterized queries | | **Warehouses** | `list`, `start`, `stop`, `create_and_wait` | Use `auto_stop_mins` to control costs | | **Unity Catalog** | `tables.list`, `tables.get`, `tables.exists`, `catalogs.list`, `schemas.list` | `list_summaries` for fast pattern matching | | **Volumes/Files** | `files.upload`, `files.download`, `files.list_directory_contents` | Path format: `/Volumes/catalog/schema/volume/file` | | **Serving** | `create_and_wait`, `query` (custom/chat/embeddings), `get_open_ai_client` | `scale_to_zero_enabled` for cost control | | **Vector Search** | `query_index`, `upsert_data_vector_index`, `sync_index` | Use `filters_json` for filtered queries | | **Pipelines** | `list_pipelines`, `get`, `start_update`, `stop_and_wait`, `list_pipeline_events` | Events contain error details for troubleshooting | | **Monitors** | `quality_monitors.get`, `delete`, `update`, `run_refresh` | Update MUST include ALL `custom_metrics` | | **Alerts V2** | `alerts_v2.create_alert`, `update_alert`, `delete_alert` | Update requires full `update_mask` | | **Secrets** | `create_scope`, `put_secret`, `get_secret` | Use for API keys, tokens | ### Critical Patterns - **Async apps (FastAPI):** SDK is synchronous — wrap with `asyncio.to_thread()` - **Long-running ops:** Use `_and_wait()` variants with `timeout=timedelta(...)` - **Error handling:** `from databricks.sdk.errors import NotFound, PermissionDenied, ResourceAlreadyExists` - **Direct REST:** `w.api_client.do(method="GET", path="/api/2.0/...")` --- ## 4. CLI Operations Reference ### Command Matrix | Category | Command | Purpose | |----------|---------|---------| | **Bundle** | `databricks bundle validate` | Pre-deploy validation (catches ~80% of errors) | | **Bundle** | `databricks bundle deploy -t <target>` | Deploy all resources | | **Bundle** | `databricks bundle run -t <target> <job>` | Start a job run | | **Bundle** | `databricks bundle destroy -t <target>` | Cleanup all resources | | **Jobs** | `databricks jobs get-run <RUN_ID> --output json` | Get run status | | **Jobs** | `databricks jobs get-run-output <TASK_RUN_ID> --output json` | Get task notebook output | | **Jobs** | `databricks jobs list-runs --job-id <JOB_ID> --output json` | List recent runs | | **Jobs** | `databricks jobs cancel-run <RUN_ID>` | Cancel running job | | **Pipelines** | `databricks pipelines get <PIPELINE_ID> --output json` | Get pipeline status | | **Pipelines** | `databricks pipelines get-update <PID> <UID> --output json` | Get update details | | **Pipelines** | `databricks pipelines list-pipeline-events <PID> --output json` | Get events/errors | | **Pipelines** | `databricks pipelines start-update <PIPELINE_ID>` | Start pipeline update | | **Clusters** | `databricks clusters get <CLUSTER_ID> --output json` | Get cluster status | | **Clusters** | `databricks clusters events <CLUSTER_ID> --output json` | Get cluster events | | **Warehouses** | `databricks warehouses get <WH_ID> --output json` | Get warehouse status | | **Auth** (IDE only) | `databricks auth login --host <url> --profile <name>` | Re-authenticate (Genie Code is pre-authenticated) | | **Workspace** | `databricks workspace export <path> --format SOURCE` | Export notebook | | **Apps** | `databricks apps logs <app-name>` | App deployment/runtime logs | ### Key jq Patterns See `references/cli-jq-patterns.md` for the complete catalog. Most critical: ```bash # Job state databricks jobs get-run <RUN_ID> --output json | jq '.state' # Task summary for multi-task jobs databricks jobs get-run <RUN_ID> --output json | jq '.tasks[] | {task: .task_key, run_id: .run_id, result: .state.result_state}' # Failed tasks only databricks jobs get-run <RUN_ID> --output json | jq '.tasks[] | select(.state.result_state == "FAILED") | {task: .task_key, error: .state.state_message, url: .run_page_url}' # Task output (MUST use task run_id, not parent job run_id) databricks jobs get-run-output <TASK_RUN_ID> --output json | jq -r '.notebook_output.result // "No output"' ``` ---
GitHub에서 보기
이 SKILL.md는 매우 커서 SkillsMP가 여기에는 첫 섹션만 미리 보여줍니다. GitHub에서 보기