Skip to main content

migrate-from-model-serving

Migrate an MLflow ResponsesAgent from Databricks Model Serving to Databricks Apps. Use when: (1) User wants to migrate from Model Serving to Apps, (2) User has a ResponsesAgent with predict()/predict_stream() methods, (3) User wants to convert to @invoke/@stream decorators.

Jump to install

Source facts

Repository
databricks-solutions/lakebase-for-ai-developers
Last source activity
July 30, 2026 at 18:26
Detected SKILL.md language
English
Stars
0
Forks
2

Install options

The review-first prompt is selected by default. You can switch to a direct command or download a local copy.

Review the source files

Read SKILL.md and any companion files shown by SkillsMP before deciding whether to install.

Showing SKILL.md

SKILL.md
Source instructions · Read-only preview
name
migrate-from-model-serving
description
Migrate an MLflow ResponsesAgent from Databricks Model Serving to Databricks Apps. Use when: (1) User wants to migrate from Model Serving to Apps, (2) User has a ResponsesAgent with predict()/predict_stream() methods, (3) User wants to convert to @invoke/@stream decorators.
# Model Serving to Databricks Apps Migration Guide This guide instructs LLM coding agents how to migrate an MLflow ResponsesAgent from Databricks Model Serving to Databricks Apps. --- ## Overview **Goal:** Migrate an agent deployed on Databricks Model Serving (using `ResponsesAgent` with `predict()`/`predict_stream()`) to Databricks Apps (using MLflow GenAI Server with `@invoke`/`@stream` decorators). **Key Transformation:** - Model Serving: Synchronous `predict()` and `predict_stream()` methods on a class - Apps: Functions with `@invoke` and `@stream` decorators (sync or async, based on user preference) **Deliverables:** After migration is complete, you will have: ``` <working-directory>/ ├── original_mlflow_model/ # Downloaded artifacts from Model Serving │ ├── MLmodel │ ├── code/ │ │ └── agent.py │ ├── input_example.json │ └── requirements.txt │ └── <app-name>/ # New Databricks App (ready to deploy) ├── agent_server/ │ ├── agent.py # Migrated agent code │ └── ... ├── databricks.yml # Bundle config with resources ├── pyproject.toml ├── uv.lock └── ... ``` > **`<app-name>`** is the name the user provides at the start of the migration. It is used as both the directory name and the Databricks App name at deploy time. --- ## Before You Begin: Gather User Inputs **Before doing anything else, ask the user three questions.** Use structured user input when available to collect all answers at once; otherwise ask directly, then execute the rest of the migration autonomously. **Questions to ask:** 1. **Databricks profile:** Which Databricks CLI profile should be used for the workspace where the Model Serving endpoint lives? (Run `databricks auth profiles` first to list available profiles and their workspaces, then present the options to the user.) 2. **App name:** What should the new Databricks App be named? (Must be lowercase, can contain letters, numbers, and hyphens, and must be unique within the workspace.) 3. **Async migration:** Would you like to migrate your agent code to be fully async? - **Yes (Recommended):** Converts all I/O operations to async (`await`/`async for`), enabling higher concurrency on smaller compute — no more threads sitting idle while waiting for LLM responses or long-running tool calls. - **No:** Keeps your existing synchronous code with minimal changes — just extracts the logic from the `ResponsesAgent` class and wraps it with `@invoke`/`@stream` decorators. Simpler migration, but each request blocks a thread while waiting for I/O. Store the answers as: - `<profile>` — used for ALL `databricks` CLI commands throughout the migration (via `--profile <profile>`) - `<app-name>` — used as both the directory name for the migrated app AND the app name when deploying with `databricks bundle deploy` - `<async>` — `yes` or `no`, determines whether to convert the agent code to async or keep it synchronous ### Validate Authentication After receiving the user's answers, validate the selected profile: ```bash databricks current-user me --profile <profile> ``` If this fails with an authentication error, prompt the user to re-authenticate: ```bash databricks auth login --profile <profile> ``` > **Important:** Remember to include `--profile <profile>` on every `databricks` CLI command throughout the migration. ### Create the App Directory Copy all scaffold files from the current working directory into a new directory named `<app-name>/`. Exclude instruction files (`AGENTS.md`, `CLAUDE.md`), hidden directories (`.agents/`, `.claude/`, `.git/`), and any migration artifacts (e.g., `original_mlflow_model/`, `.migration-venv/`). Do NOT search for or copy scaffold files from other directories or templates — everything you need is right here. All subsequent migration steps operate inside the `<app-name>/` directory. > **Note:** The `agent_server/agent.py` scaffold is intentionally framework-agnostic — it contains the `@invoke`/`@stream` decorator pattern with TODO placeholders. Step 3 (Migrate the Agent Code) will replace these placeholders with the actual agent logic from the original Model Serving endpoint. ### Create Task List **Create a task list to track progress.** This helps the user follow along and see what's completed, in progress, and pending. > **User tip:** Press `Ctrl+T` to toggle the task list view in your terminal. The display shows up to 10 tasks at a time with status indicators. Create the following tasks using the `TaskCreate` tool: | Task | Description | |------|-------------| | **Authenticate to Databricks** | Verify Databricks CLI authentication and validate the selected profile | | **Download original agent artifacts** | Download the MLflow model artifacts from Model Serving endpoint | | **Analyze and understand agent code** | Examine the original agent code, identify tools, resources, and dependencies | | **Migrate agent code to Apps format** | Transform ResponsesAgent class to @invoke/@stream decorated functions | | **Set up and configure the app** | Install dependencies, run quickstart, configure environment | | **Test agent locally** | Start local server and verify the agent works correctly | | **Deploy to Databricks Apps** | Configure databricks.yml resources and deploy with Databricks Asset Bundles | | **Test deployed app** | Verify the deployed app responds correctly | Update task status as you progress: - Mark tasks as `in_progress` when starting each step - Mark tasks as `completed` when finished - This gives the user visibility into migration progress --- ## Step 1: Download the Original Agent Code > **Task:** Mark "Authenticate to Databricks" as `completed`. Mark "Download original agent artifacts" as `in_progress`. > > **Note:** The `<profile>` and `<app-name>` values were collected from the user in the "Before You Begin" section. Use them throughout. Download the original agent code from the Model Serving endpoint. This requires setting up a virtual environment with MLflow to access the model artifacts. ### 1.1 Get Model Info from Endpoint If you have a serving endpoint name, extract the model details: ```bash # Get endpoint info (remember to include --profile if using non-default) databricks serving-endpoints get <endpoint-name> --profile <profile> --output json ``` Look for `served_entities[0].entity_name` (model name) and `entity_version` in the response. Find the entity with 100% traffic in `traffic_config.routes`. ### 1.2 Download Model Artifacts Use `uv run --with` to download artifacts without creating a separate virtual environment. The `mlflow[databricks]` extra includes `boto3` for Unity Catalog artifact access: ```bash DATABRICKS_CONFIG_PROFILE=<profile> uv run --no-project \ --with "mlflow[databricks]>=2.15.0" \ --with "databricks-sdk>=0.30.0" \ python3 << 'EOF' import mlflow mlflow.set_tracking_uri("databricks") # Replace with actual values from step 1.1 MODEL_NAME = "<model-name>" VERSION = "<version>" print(f"Downloading model: models:/{MODEL_NAME}/{VERSION}") mlflow.artifacts.download_artifacts( artifact_uri=f"models:/{MODEL_NAME}/{VERSION}", dst_path="./original_mlflow_model" ) print("Download complete! Artifacts saved to ./original_mlflow_model") EOF ``` ### 1.3 Verify Downloaded Artifacts Check that the key files exist and understand the full structure: ```bash # List all downloaded files recursively find ./original_mlflow_model -type f | head -50 # Check for MLmodel file (contains resource requirements) cat ./original_mlflow_model/MLmodel # Check for input example (useful for testing) cat ./original_mlflow_model/input_example.json 2>/dev/null ``` **Examine the `/code` folder** - contains all code dependencies logged via `code_paths=["..."]`: ```bash # List all code files ls -la ./original_mlflow_model/code/ # The main agent is typically agent.py, but there may be additional modules find ./original_mlflow_model/code -name "*.py" -type f ``` **Examine the `/artifacts` folder** (if present) - contains artifacts logged via `artifacts={...}`: ```bash # Check for artifacts folder ls -la ./original_mlflow_model/artifacts/ 2>/dev/null # List all artifacts find ./original_mlflow_model/artifacts -type f 2>/dev/null ``` > **Important:** Take note of ALL files in `/code` and `/artifacts`. You will need to copy these to the migrated app and ensure imports still work correctly. ### Expected Output Structure After successful download, you should have: ``` ./original_mlflow_model/ ├── MLmodel # Model metadata and resource requirements ├── code/ # Code logged via code_paths=["..."] │ ├── agent.py # Main agent implementation │ ├── utils.py # (optional) Helper modules │ ├── tools.py # (optional) Custom tool definitions │ └── ... # Any other code dependencies ├── artifacts/ # (optional) Artifacts logged via artifacts={...} │ ├── config.yaml # (optional) Configuration files │ ├── prompts/ # (optional) Prompt templates │ └── ... # Any other artifacts (data files, etc.) ├── input_example.json # Sample request for testing ├── requirements.txt # Original dependencies └── ... ``` ### Key Files to Examine 1. **`code/agent.py`** - Contains the `ResponsesAgent` class with `predict()` and `predict_stream()` methods 2. **`code/*.py`** - Any additional Python modules the agent imports 3. **`MLmodel`** - Contains the `resources` section listing required Databricks resources 4. **`artifacts/`** - Any configuration files, prompts, or data files the agent uses 5. **`input_example.json`** - Use this to test the migrated agent ### Troubleshooting Model Download **"Unable to import necessary dependencies to access model version files in Unity Catalog"** This means `boto3` is missing. Ensure you're using `mlflow[databricks]` (not just `mlflow`) in the `--with` flag — the `[databricks]` extra includes `boto3`. **"INVALID_PARAMETER_VALUE" or authentication errors** Re-authenticate with Databricks (include profile if non-default): ```bash databricks auth login --profile <profile> ``` **Wrong workspace / Model not found** Make sure you're using the correct profile that corresponds to the workspace where the model is deployed: ```bash # List profiles to see which workspace each points to databricks auth profiles # Verify you can access the workspace databricks current-user me --profile <profile> # List models in that workspace databricks registered-models list --profile <profile> databricks model-versions list --name "<model-name>" --profile <profile> ``` --- ## Step 2: Understand the Key Transformations > **Task:** Mark "Download original agent artifacts" as `completed`. Mark "Analyze and understand agent code" as `in_progress`. ### Entry Point Transformation In both cases, the `ResponsesAgent` class is replaced with decorated functions. The difference is whether those functions are async or sync. **Model Serving (OLD):** ```python from mlflow.pyfunc import ResponsesAgent, ResponsesAgentRequest, ResponsesAgentResponse class MyAgent(ResponsesAgent): def predict(self, request: ResponsesAgentRequest, params=None) -> ResponsesAgentResponse: # Synchronous implementation ... return ResponsesAgentResponse(output=outputs) def predict_stream(self, request: ResponsesAgentRequest, params=None): # Synchronous generator for chunk in ...: yield ResponsesAgentStreamEvent(...) ``` **Apps — Async (if `<async>` = yes):** ```python from mlflow.genai.agent_server import invoke, stream from mlflow.types.responses import ( ResponsesAgentRequest, ResponsesAgentResponse, ResponsesAgentStreamEvent, ) @invoke() async def non_streaming(request: ResponsesAgentRequest) -> ResponsesAgentResponse: # Async implementation - typically calls streaming() and collects results outputs = [ event.item async for event in streaming(request) if event.type == "response.output_item.done" ] return ResponsesAgentResponse(output=outputs) @stream() async def streaming(request: ResponsesAgentRequest) -> AsyncGenerator[ResponsesAgentStreamEvent, None]: # Async generator async for event in ...: yield event ``` **Apps — Sync (if `<async>` = no):** ```python from mlflow.genai.agent_server import invoke, stream from mlflow.types.responses import ( ResponsesAgentRequest, ResponsesAgentResponse, ResponsesAgentStreamEvent, ) @invoke() def non_streaming(request: ResponsesAgentRequest) -> ResponsesAgentResponse: # Same sync logic from original predict(), extracted from the class ... return ResponsesAgentResponse(output=outputs) @stream() def streaming(request: ResponsesAgentRequest): # Same sync generator from original predict_stream(), extracted from the class for chunk in ...: yield ResponsesAgentStreamEvent(...) ``` ### Key Differences | Aspect | Model Serving | Apps (async) | Apps (sync) | |--------|--------------|------|------| | Structure | `class MyAgent(ResponsesAgent)` | Decorated functions | Decorated functions | | Functions | `def predict()` / `def predict_stream()` | `async def` with `await` | `def` (same as original) | | Streaming | Sync generator (`yield`) | Async generator (`async for` / `yield`) | Sync generator (`yield`) | | Server | MLflow Model Server | MLflow GenAI Server (FastAPI) | MLflow GenAI Server (FastAPI) | | Deployment | `databricks_agents.deploy()` | `databricks bundle deploy` + `bundle run` | `databricks bundle deploy` + `bundle run` | ### Async Patterns (only if `<async>` = yes) > **Skip this section if the user chose synchronous migration.** The sync path keeps all original I/O calls as-is. All I/O operations must be converted to async: ```python # OLD (sync) response = client.chat(messages) # NEW (async) response = await client.achat(messages) # OLD (sync iteration) for chunk in stream: yield chunk # NEW (async iteration) async for chunk in stream: yield chunk ``` --- ## Step 3: Migrate the Agent Code > **Task:** Mark "Analyze and understand agent code" as `completed`. Mark "Migrate agent code to Apps format" as `in_progress`. ### 3.1 Copy Code Dependencies and Artifacts The original MLflow model may contain multiple code files and artifacts that need to be migrated. **Copy all code files from `/code` to `agent_server/`:** ```bash # Copy all Python files from original code folder cp ./original_mlflow_model/code/*.py ./<app-name>/agent_server/ # If there are subdirectories with code, copy those too # cp -r ./original_mlflow_model/code/submodule ./<app-name>/agent_server/ ``` **Copy artifacts (if present):** ```bash # Create an artifacts directory in the migrated app if needed mkdir -p ./<app-name>/agent_server/artifacts # Copy all artifacts
View on GitHub
This SKILL.md is very large, so SkillsMP previews the first section here. View on GitHub