- name
- databricks-autonomous-operations
- description
- End-to-end autonomous deployment and operations skill. Deploys Databricks Asset Bundles, runs jobs and pipelines, polls for completion, diagnoses failures, applies fixes, redeploys, and verifies — all without human intervention. Also serves as SDK/CLI/REST API reference. Use when deploying bundles, running jobs, monitoring pipelines, troubleshooting ANY failure, or operating in a self-healing deploy-fix-redeploy cycle. Triggers on "deploy", "bundle deploy", "run job", "run pipeline", "make it work", "job failed", "troubleshoot", "fix and redeploy".
- metadata
- {"author":"prashanth subrahmanyam","version":"3.2","domain":"operations","role":"shared","used_by_stages":[1,2,3,4,5,6,7,8,9],"called_by":["bronze/00-bronze-layer-setup","silver/00-silver-layer-setup","gold/01-gold-layer-setup","semantic-layer/00-semantic-layer-setup","monitoring/00-observability-setup","ml/00-ml-pipeline-setup","genai-agents/00-course-orchestrator"],"dependencies":["admin/self-improvement"],"triggers":["deploy","bundle deploy","bundle run","run job","run pipeline","make it work","deploy and run","deploy and fix","job failed","monitor job","troubleshoot","get-run","run output","task failed","redeploy","job status","pipeline failed","DLT error","cluster issue","monitor failed","alert failed","deploy failed","databricks-sdk","databricks-connect","CLI","INTERNAL_ERROR","UNRESOLVED_COLUMN","ModuleNotFoundError","TABLE_OR_VIEW_NOT_FOUND","ResourceAlreadyExists","PERMISSION_DENIED","self-heal","[Truncated]"],"last_verified":"2026-06-02","volatility":"medium","clients":["ide_cli","genie_code"],"deploy_verb":"bundle deploy --target dev","deploy_note":"operations/CLI reference — on Genie Code every databricks command routes via runDatabricksCli FROM THE BUNDLE EDITOR (dp_bundle_root); a blocked bundle deploy/run is a page-context signal, NEVER substitute SDK/REST creation (RULE_10); see genie-code-environment","coverage":"all_stages","upstream_sources":[{"name":"databricks-agent-skills","repo":"databricks/databricks-agent-skills","paths":"[Truncated]","relationship":"extended","last_synced":"2026-08-30","sync_commit":"ca92a6c"}]}
# Databricks Autonomous Operations
## 1. Overview
> **Context Loading Rule:** Load this skill at **deployment time**, not during planning.
> During earlier phases, retain only this 3-step summary in working memory:
> `Deploy (bundle deploy) → Poll (get-run / pipelines get) → Verify (get-run-output)`.
> Full skill activation should happen when the first `bundle deploy` or `bundle run` is invoked.
This skill is both an **SDK/CLI/Connect reference** and an **autonomous operations playbook**. It teaches the AI agent to operate as an SRE — independently deploying, monitoring, diagnosing failures, applying fixes, redeploying, and verifying results across all Databricks resource types.
**Core Loop:** Deploy → Poll → Diagnose → Fix → Redeploy → Verify (max 3 iterations before escalation to user).
### When to Activate This Skill
- Deploying or running Databricks Asset Bundles (`bundle deploy`, `bundle run`)
- Monitoring job or pipeline runs for completion
- Troubleshooting ANY failure: jobs, DLT pipelines, monitors, alerts, clusters, Genie Spaces
- Using the Databricks Python SDK, CLI, Connect, or REST API
- Encountering error messages from Databricks services
- Operating in a self-healing deploy-fix-redeploy cycle
- Checking job/task/pipeline status or retrieving run output
**SDK Docs:** https://databricks-sdk-py.readthedocs.io/en/latest/
**GitHub:** https://github.com/databricks/databricks-sdk-py
**Runnable Examples:** See `examples/` directory for complete, copy-pasteable scripts:
- `examples/1-authentication.py` — All auth patterns (env vars, profiles, Azure SP, AccountClient, notebook context)
- `examples/2-clusters-and-jobs.py` — Cluster auto-selection, autoscaling, job CRUD, submit_and_wait, run_now_and_wait
- `examples/3-sql-and-warehouses.py` — Parameterized queries, chunked results, query_to_dataframe helper
- `examples/4-unity-catalog.py` — Tables, schemas, catalogs, volumes, file operations, pattern matching
- `examples/5-serving-and-vector-search.py` — Endpoint creation with TrafficConfig, chat/embedding queries, vector search with filters
- `examples/6-autonomous-operations.py` — Job monitoring, multi-task output retrieval, failure diagnosis, self-healing loop
---
## 2. Environment & Authentication
> **Client routing (RULE_1/2/4 — read once, applies to every command in this skill).** This skill is the
> CLI/SDK operations surface; the same operations run on both clients via different channels — **deploy
> mechanics are owned by `databricks-asset-bundles` (the spine); reference it rather than re-deriving them.**
> - **IDE (Cursor):** the local `databricks` CLI; auth via `databricks auth login` / `~/.databrickscfg`;
> `databricks-connect` is available for local Spark.
> - **Genie Code has *three* execution paths — try them in order; "blocked on one path ≠ impossible":**
> 1. **`runDatabricksCli`** — the allow-listed, pre-authenticated CLI path (no `auth login`). The path
> for `bundle validate` / `summary` / `deploy --target dev` and read verbs. `--version`/`help`/
> `auth token`/`aitools`/`apps validate`/`apps manifest` are hard-blocked; `apps deploy` is
> **unreliable here** (page-dependent + CWD-defeated). Use a `bundle validate` behavior probe instead
> of a numeric `--version` compare.
> 2. **Python SDK** (`WorkspaceClient` via `executeCode`) — the **most capable** path: it **bypasses the
> CLI allow-list** and is the reliable way to `w.apps.deploy(...)`, retrieve `w.config.token`, and poll
> deployment/run state. **Caveat:** the SDK has **no bundle-deploy equivalent** (`bundle deploy` is a
> composite client-side operation) — keep `bundle deploy` on `runDatabricksCli`.
> 3. **Native tools** (`createAsset`/`readTable`/…) for governed asset operations.
>
> No local Spark on Genie Code (use workspace **serverless**; `databricks-connect` is IDE-only). Full
> allow-list / CWD / FUSE / escape-hatch detail is in the **`genie-code-environment`** skill — load it on
> demand. The CLI command examples below run via path 1 on Genie Code (local shell on the IDE); where a
> verb is CLI-blocked, reach for the SDK (path 2).
>
> **⛔ The carve-out the "fall to the SDK" rule does NOT cover — `bundle deploy`/`run` and resource creation.**
> "Blocked ≠ impossible, try the SDK" is for **read-only / ad-hoc** ops (polling, inspecting schemas, lineage,
> token retrieval, `apps deploy`). It is **NOT** a license to substitute the SDK/REST for the bundle. When
> `bundle deploy`/`run` is blocked, that is a **page-context signal, not a dead end**: the verb is gated to the
> **bundle-folder page**, which on Genie Code you reach by opening the **"Open in bundle editor"** affordance on
> the `dp_bundle_root` folder (the bundle editor's CWD *is* the bundle root, where `validate`/`deploy`/`run` are
> pre-approved). **FIELD-CONFIRMED:** the same `bundle deploy` that returned "blocked by safety guardrails" /
> "`databricks.yml` not found" from a file/notebook page returned "Deployment complete!" and `bundle run …
> SUCCESS` from the bundle editor. So the fix is **navigate to the bundle editor**, never `w.jobs.create()` /
> `POST /api/2.1/jobs/create` / `POST /api/2.0/pipelines` / `CREATE TABLE` via `executeCode`. Creating
> jobs/pipelines/tables directly is the **RULE_10 authoring-discipline violation** — it produces live,
> un-versioned state that diverges from the bundle and is the exact regression this spine prevents. If
> `bundle deploy`/`run` **still** fails *from the bundle editor*, STOP and report the blocker — the SDK/REST
> creation route is an **escape hatch only on explicit operator authorization.** Surface a clickable
> bundle-editor link: `file_id = w.workspace.get_status("<dp_bundle_root>/databricks.yml").object_id`,
> `folder_id = w.workspace.get_status("<dp_bundle_root>").object_id`, link =
> `{w.config.host}/editor/files/{file_id}?o={w.get_workspace_id()}&contextId=folder%3A{folder_id}`. Detail in
> `genie-code-environment` §3/§8.
### Setup
- SDK: `uv pip install databricks-sdk`; **Connect (IDE-only):** `uv pip install databricks-connect` —
not used on Genie Code (serverless; no local Spark)
- CLI version: **>= 0.278.0** on the IDE (`databricks --version`); on Genie Code the version is not
introspectable (use a `bundle validate` behavior probe)
- Config (IDE): `~/.databrickscfg` or env vars `DATABRICKS_HOST`, `DATABRICKS_TOKEN`; Genie Code is
pre-authenticated in-session
### Quick Auth
```python
from databricks.sdk import WorkspaceClient
w = WorkspaceClient() # Auto-detect (env, config file, or notebook)
w = WorkspaceClient(profile="MY_PROFILE") # Named profile
```
- **Token expired? (IDE path)** `databricks auth login --host <url> --profile <name>` — N/A on Genie Code
(pre-authenticated)
- **Profile-based CLI:** `DATABRICKS_CONFIG_PROFILE=<name> databricks <command>`
- **Full auth patterns** (Azure SP, AccountClient, etc.): see `references/sdk-api-reference.md`
---
## 3. SDK API Quick Reference
**Full code examples and patterns:** `references/sdk-api-reference.md`
**Runnable scripts:** `examples/` directory (see file listing in Section 1)
| API | Key Operations | Critical Notes |
|-----|---------------|----------------|
| **Clusters** | `list`, `get`, `create_and_wait`, `ensure_cluster_is_running`, `start`, `stop` | Use `select_spark_version()` and `select_node_type()` |
| **Jobs** | `list`, `run_now_and_wait`, `submit_and_wait`, `get_run`, `get_run_output`, `cancel_run` | `get_run_output` needs **task** `run_id`, not parent job `run_id` |
| **SQL Execution** | `execute_statement` with `wait_timeout` | Use `StatementParameterListItem` for parameterized queries |
| **Warehouses** | `list`, `start`, `stop`, `create_and_wait` | Use `auto_stop_mins` to control costs |
| **Unity Catalog** | `tables.list`, `tables.get`, `tables.exists`, `catalogs.list`, `schemas.list` | `list_summaries` for fast pattern matching |
| **Volumes/Files** | `files.upload`, `files.download`, `files.list_directory_contents` | Path format: `/Volumes/catalog/schema/volume/file` |
| **Serving** | `create_and_wait`, `query` (custom/chat/embeddings), `get_open_ai_client` | `scale_to_zero_enabled` for cost control |
| **Vector Search** | `query_index`, `upsert_data_vector_index`, `sync_index` | Use `filters_json` for filtered queries |
| **Pipelines** | `list_pipelines`, `get`, `start_update`, `stop_and_wait`, `list_pipeline_events` | Events contain error details for troubleshooting |
| **Monitors** | `quality_monitors.get`, `delete`, `update`, `run_refresh` | Update MUST include ALL `custom_metrics` |
| **Alerts V2** | `alerts_v2.create_alert`, `update_alert`, `delete_alert` | Update requires full `update_mask` |
| **Secrets** | `create_scope`, `put_secret`, `get_secret` | Use for API keys, tokens |
### Critical Patterns
- **Async apps (FastAPI):** SDK is synchronous — wrap with `asyncio.to_thread()`
- **Long-running ops:** Use `_and_wait()` variants with `timeout=timedelta(...)`
- **Error handling:** `from databricks.sdk.errors import NotFound, PermissionDenied, ResourceAlreadyExists`
- **Direct REST:** `w.api_client.do(method="GET", path="/api/2.0/...")`
---
## 4. CLI Operations Reference
### Command Matrix
| Category | Command | Purpose |
|----------|---------|---------|
| **Bundle** | `databricks bundle validate` | Pre-deploy validation (catches ~80% of errors) |
| **Bundle** | `databricks bundle deploy -t <target>` | Deploy all resources |
| **Bundle** | `databricks bundle run -t <target> <job>` | Start a job run |
| **Bundle** | `databricks bundle destroy -t <target>` | Cleanup all resources |
| **Jobs** | `databricks jobs get-run <RUN_ID> --output json` | Get run status |
| **Jobs** | `databricks jobs get-run-output <TASK_RUN_ID> --output json` | Get task notebook output |
| **Jobs** | `databricks jobs list-runs --job-id <JOB_ID> --output json` | List recent runs |
| **Jobs** | `databricks jobs cancel-run <RUN_ID>` | Cancel running job |
| **Pipelines** | `databricks pipelines get <PIPELINE_ID> --output json` | Get pipeline status |
| **Pipelines** | `databricks pipelines get-update <PID> <UID> --output json` | Get update details |
| **Pipelines** | `databricks pipelines list-pipeline-events <PID> --output json` | Get events/errors |
| **Pipelines** | `databricks pipelines start-update <PIPELINE_ID>` | Start pipeline update |
| **Clusters** | `databricks clusters get <CLUSTER_ID> --output json` | Get cluster status |
| **Clusters** | `databricks clusters events <CLUSTER_ID> --output json` | Get cluster events |
| **Warehouses** | `databricks warehouses get <WH_ID> --output json` | Get warehouse status |
| **Auth** (IDE only) | `databricks auth login --host <url> --profile <name>` | Re-authenticate (Genie Code is pre-authenticated) |
| **Workspace** | `databricks workspace export <path> --format SOURCE` | Export notebook |
| **Apps** | `databricks apps logs <app-name>` | App deployment/runtime logs |
### Key jq Patterns
See `references/cli-jq-patterns.md` for the complete catalog. Most critical:
```bash
# Job state
databricks jobs get-run <RUN_ID> --output json | jq '.state'
# Task summary for multi-task jobs
databricks jobs get-run <RUN_ID> --output json | jq '.tasks[] | {task: .task_key, run_id: .run_id, result: .state.result_state}'
# Failed tasks only
databricks jobs get-run <RUN_ID> --output json | jq '.tasks[] | select(.state.result_state == "FAILED") | {task: .task_key, error: .state.state_message, url: .run_page_url}'
# Task output (MUST use task run_id, not parent job run_id)
databricks jobs get-run-output <TASK_RUN_ID> --output json | jq -r '.notebook_output.result // "No output"'
```
---
GitHub에서 보기