What you will be able to do
- Choose between AI Playground, Supervisor Agent, custom code agents and the Agent Bricks CLI for building an agent on Databricks
- Describe the parts of the Databricks Apps agent template: chat UI, MLflow AgentServer, ResponsesAgent and an agent framework
- Explain why wrapping an agent of any framework in MLflow ResponsesAgent makes it work with Databricks evaluation, tracing and monitoring
- Add tools to an agent through MCP servers or local function tools, and know which of them need resource grants
- Set up streaming output, error propagation and custom inputs and outputs in a ResponsesAgent
Key concept
MLflow ResponsesAgent — ResponsesAgent is the MLflow interface you wrap your agent in, whatever framework you wrote it with. Once wrapped, the agent works with Databricks logging, tracing, evaluation, deployment and monitoring without further changes.
1.Where agent development happens on Databricks
Databricks gives you several ways to build an agent, from no-code to fully custom. AI Playground is the no-code option. You pick an LLM, add tools and chat with the agent to test it, then export the result to code. Knowledge Assistant lets you build and optimize domain-specific AI chatbots through an interface. The Supervisor Agent is a managed way to orchestrate other agents and tools, Genie Agents among them. For full control you build a custom agent in Python, and Databricks says this route supports agents written with any authoring library and is integrated with MLflow Tracing. The Agent Bricks CLI (Beta) is a code-first alternative to building custom agents on Databricks Apps. It scaffolds, runs and deploys custom code agents from the command line.
| Option | What it gives you |
|---|---|
| AI Playground | No-code prototyping of agent behaviour and tool integrations before you generate code |
| Knowledge Assistant | Build and optimize domain-specific AI chatbots using an intuitive interface |
| Supervisor Agent | Orchestrates Genie Agents, agent endpoints, Unity Catalog functions, MCP servers and custom agents |
| Build custom agents | Python authoring with LangGraph, LangChain, OpenAI or LlamaIndex, integrated with MLflow Tracing and iterated on with Databricks Apps |
| Agent Bricks CLI (Beta) | Scaffold, run and deploy custom code agents from the command line, with managed memory, sessions, tools and tracing wired up |
Checkpoint 1 of 6· Match them up
Match each Databricks option to its description
Tap a term, then the definition that fits it.
The Databricks agent overview table describes each option this way. Only the custom-agent route names specific authoring libraries.
“Build a supervisor agent that orchestrates Genie Agents, agent endpoints, Unity Catalog functions, MCP servers, and custom agents.”Source: docs.databricks.com
Sources1
2.The Databricks Apps agent template
If you build a custom agent and deploy it on Databricks Apps, you control the agent code, the server configuration and the deployment workflow. Start from a template rather than wiring the pieces together yourself. The agent-openai-agents-sdk template contains three things: an agent built with the OpenAI Agents SDK, starter code for a conversational REST API with a chat UI, and code to evaluate the agent using MLflow. Its server is the MLflow AgentServer, an async FastAPI server with built-in tracing and observability. The AgentServer exposes a /responses endpoint and takes care of request routing, logging and error handling.
git clone https://github.com/databricks/app-templates.git
cd app-templates/agent-openai-agents-sdkCheckpoint 2 of 6· Check yourself
In the Databricks Apps agent template, which component provides the /responses endpoint and manages request routing, logging and error handling?
The chat UI is the frontend and the OpenAI Agents SDK is the agent framework. The MLflow AgentServer is the async FastAPI server that serves /responses.
“The AgentServer provides the /responses endpoint for querying your agent”Source: docs.databricks.com
Sources2
3.ResponsesAgent: one interface for any framework
This is what connects the agent framework to MLflow. You can author an agent with any framework and then wrap it in ResponsesAgent, and the wrapped agent works out of the box with AI Playground, Agent Evaluation and Agent Monitoring. You also get typed Python authoring classes, automatic tracing that combines streamed responses into a single trace, and compatibility with the OpenAI Responses schema. ResponsesAgent also supports multi-agent systems, streaming output, tool-calling confirmation and long-running tools. It can return several messages per turn, intermediate tool-calling messages included, which the documentation says improves quality and conversation management.
In code, the class-based form is a subclass of ResponsesAgent. You implement predict, which takes a ResponsesAgentRequest and returns a ResponsesAgentResponse, and predict_stream, which yields ResponsesAgentStreamEvent objects. On Model Serving agents use this class-based structure. On Databricks Apps the MLflow AgentServer instead serves module-level functions decorated with @invoke() and @stream(), which use the same request, response and stream-event types.
from mlflow.pyfunc import ResponsesAgent, ResponsesAgentRequest, ResponsesAgentResponse
class MyAgent(ResponsesAgent):
def predict(self, request: ResponsesAgentRequest, params=None) -> ResponsesAgentResponse:
# Synchronous implementation
...
return ResponsesAgentResponse(output=outputs)
def predict_stream(self, request: ResponsesAgentRequest, params=None):
# Synchronous generator
for chunk in ...:
yield ResponsesAgentStreamEvent(...)Checkpoint 3 of 6· Exam question
A team is building a tool-calling LangChain agent on Databricks that must invoke a Python function already registered and governed in Unity Catalog. They want the function's permissions to remain enforced by Unity Catalog while exposing it as a callable tool to the agent. Which approach should they use?
Correct answer: A — Wrap the Unity Catalog function with `UCFunctionToolkit(function_names=[...])` and pass the toolkit's `tools` property into `create_tool_calling_agent()`.
- A. This is correct because `UCFunctionToolkit` wraps a governed Unity Catalog function as a LangChain-compatible tool while Unity Catalog continues to enforce access control on the underlying function. Extracting tools via the toolkit's `tools` property is the documented way to pass governed functions into a tool-calling agent.
- B. This is incorrect because duplicating the function's logic outside Unity Catalog removes it from catalog-enforced permissions and auditing, defeating the goal of keeping governance intact.
- C. This is incorrect because loading a PyFunc model is meant for serving previously logged MLflow models, not for exposing a Unity Catalog Python function as a callable agent tool.
- D. This is incorrect because Genie Spaces and the conversation API are designed for natural-language querying of structured data, not for wrapping arbitrary governed Python functions as agent tools.
4.Giving the agent tools
An agent becomes useful when it can query databases, search documents or call APIs. You give it those abilities by connecting MCP servers. The template already includes one MCP server connection. To add more, you configure additional MCP servers in your agent code and grant the permissions they need in databricks.yml. Some operations don't need external data, such as transformations, calculations and utilities. For those you define local function tools that run in the same process as the agent.
Databricks also provides managed MCP servers. For example, Unity Catalog functions are reachable through a managed server for their catalog and schema: the built-in function system.ai.python_exec, which lets an agent run Python code, is available through the managed MCP server for the system.ai schema. The databricks-mcp package provides DatabricksMCPClient to list a server's tools and to work out the Unity Catalog resources the agent needs. The same managed-server approach covers Genie and AI Search, where the agent needs access to the resources behind each server, such as CAN_RUN on a Genie space or SELECT on an AI Search index.
mcp_client = DatabricksMCPClient(
server_url=f"{host}/api/2.0/mcp/functions/system/ai/python_exec",
workspace_client=workspace_client,
)
tools = mcp_client.list_tools()from langchain_core.tools import tool
from langgraph.prebuilt import create_react_agent
from databricks_langchain import ChatDatabricks
@tool
def get_current_time() -> str:
"""Get the current date and time."""
from datetime import datetime
return datetime.now().isoformat()
agent = create_react_agent(
ChatDatabricks(endpoint="databricks-claude-sonnet-4-5"),
tools=[get_current_time],
)Checkpoint 4 of 6· Fill the gap
This is the same tool written for the OpenAI Agents SDK. Which decorator goes in the blank?
@ ?
def get_current_time() -> str:
"""Get the current date and time."""
from datetime import datetime
return datetime.now().isoformat()
agent = Agent(
name="My agent",
instructions="You are a helpful assistant.",
model="databricks-claude-sonnet-4-5",
tools=[get_current_time],
)The OpenAI Agents SDK uses @function_tool. @tool is LangChain's decorator, and @mlflow.trace adds tracing but does not register a tool.
Source: docs.databricks.com5.Streaming, errors and custom inputs
With streaming, the agent sends its answer in chunks as it generates it instead of waiting for the full response. In ResponsesAgent you send a series of output_text.delta events that all share one item_id, then finish with a response.output_item.done event that has the same item_id and the complete text. The done event tells Databricks to record the output with MLflow tracing, aggregate the stream in Unity Gateway inference tables, and show the full answer in AI Playground. If an error happens during streaming, it arrives with the last token under databricks_output.error, and the client is responsible for handling it.
Checkpoint 5 of 6· Put it in order
Put the streaming sequence in order
- 1.Send a response.output_item.done event with the same item_id and the complete text
- 2.Databricks traces the output and shows the complete response in AI Playground
- 3.Emit multiple output_text.delta events sharing one item_id
The delta events stream the chunks. Only the final done event carries the complete text and triggers tracing and aggregation.
“Send a final response.output_item.done event with the same item_id as the delta events containing the complete final output text.”Source: docs.databricks.com
Sometimes an agent needs extra inputs, such as client_type or session_id, or extra outputs, such as links to retrieval sources, that should stay out of the chat history. ResponsesAgent has two fields for this: custom_inputs and custom_outputs. In your agent code you read the extra inputs from request.custom_inputs. One documented limit: the Agent Evaluation review app does not support rendering traces for agents with additional input fields.
Checkpoint 6 of 6· Check yourself
Your agent needs a session_id from the caller, and that value should not appear in future chat history. Where does it go?
custom_inputs carries extra fields such as session_id without putting them in the chat history.
“MLflow ResponsesAgent natively supports the fields custom_inputs and custom_outputs.”Source: docs.databricks.com
To deploy to Model Serving, you log the agent with MLflow, register it in Unity Catalog and deploy it. When you log, pass the resources the agent depends on, such as its serving endpoint, Unity Catalog functions and vector search indexes, so Model Serving can grant access at deployment. For managed MCP servers, DatabricksMCPClient().get_databricks_resources() returns the resources the server needs. The example below sets the registry URI to Unity Catalog beforehand with mlflow.set_registry_uri.
with mlflow.start_run():
logged_model_info = mlflow.pyfunc.log_model(
artifact_path="mcp_agent",
python_model=agent_script,
resources=resources,
)
UC_MODEL_NAME = "main.default.databricks_docs_mcp_agent"
registered_model = mlflow.register_model(logged_model_info.model_uri, UC_MODEL_NAME)
agents.deploy(
model_name=UC_MODEL_NAME,
model_version=registered_model.version,
)Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Every tool an agent uses must have a resource grant in databricks.yml.Why is that wrong?
Only external resources such as additional MCP servers need grants. Local function tools run inside the agent process and need none.
Covered in Giving the agent tools
2.Adding custom_inputs to an agent has no effect on any of the evaluation tooling.Why is that wrong?
The documentation names one limit: the Agent Evaluation review app cannot render traces for agents that have additional input fields.
Covered in Streaming, errors and custom inputs
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Supports agents written with any authoring library, including LangGraph, LangChain, OpenAI, and LlamaIndex.”
↩︎ Where agent development happens on Databricks“Prototype and test agents in a no-code environment.”
↩︎ Where agent development happens on Databricks“Build and optimize domain-specific AI chatbots using an intuitive interface.”
↩︎ Where agent development happens on Databricks“Build a supervisor agent that orchestrates Genie Agents, agent endpoints, Unity Catalog functions, MCP servers, and custom agents.”
↩︎ Checkpoint - 2.
“Databricks Apps gives you full control over the agent code, server configuration, and deployment workflow.”
↩︎ The Databricks Apps agent template“An async FastAPI server that handles agent requests with built-in tracing and observability.”
↩︎ The Databricks Apps agent template“The AgentServer provides the /responses endpoint for querying your agent”
↩︎ The Databricks Apps agent template“Wrap any existing agent using the ResponsesAgent interface to get out-of-the-box compatibility with AI Playground, Agent Evaluation, and Agent Monitoring.”
↩︎ ResponsesAgent: one interface for any framework“Return multiple messages, including intermediate tool-calling messages, for improved quality and conversation management.”
↩︎ ResponsesAgent: one interface for any framework“To add more tools, configure additional MCP servers in your agent code and grant the required permissions in databricks.yml.”
↩︎ Giving the agent tools“pass the Unity Gateway endpoint name as the model argument and set use_ai_gateway=True on the Databricks LLM client.”
↩︎ Giving the agent tools“Databricks propagates any errors encountered while streaming with the last token under databricks_output.error.”
↩︎ Streaming, errors and custom inputs“Databricks recommends MLflow ResponsesAgent to build agents.”
↩︎ Key concept“Local function tools don't require resource grants in databricks.yml because they run within the agent process.”
↩︎ Exam trap 1“The Agent Evaluation review app does not support rendering traces for agents with additional input fields.”
↩︎ Exam trap 2“The key is wrapping your agent with MLflow ResponsesAgent interface.”
↩︎ Prediction“Local function tools don't require resource grants in databricks.yml because they run within the agent process.”
↩︎ Prediction“Send a final response.output_item.done event with the same item_id as the delta events containing the complete final output text.”
↩︎ Checkpoint“MLflow ResponsesAgent natively supports the fields custom_inputs and custom_outputs.”
↩︎ Checkpoint - 3.
“On Model Serving, agents use a class-based ResponsesAgent with predict() and predict_stream() methods.”
↩︎ ResponsesAgent: one interface for any framework - 4.
“connect to the managed MCP server for the system.ai Unity Catalog schema.”
↩︎ Giving the agent tools - 5.
“Log the agent with all the resources it needs at logging time, then deploy.”
↩︎ Streaming, errors and custom inputs