What you will be able to do
- Choose between a local function tool, a Unity Catalog function tool, a retriever tool and an MCP server for a given capability
- Wrap a Unity Catalog function with UCFunctionToolkit and attach it to a LangChain agent
- Explain why Databricks recommends MCP for adding Unity Catalog functions to agents
- Route structured-data questions to Genie agents and unstructured ones to an AI Search retriever
1.Local function tools: the lightest option
In an agent built with LangChain or a similar library, every capability the LLM can call is a tool. The simplest kind lives in your own code. "For operations that don't require external data sources or APIs, define tools directly in your agent code." Typical uses are data transformations, calculations and utility functions. Each library has its own decorator: LangChain uses @tool and the OpenAI Agents SDK uses @function_tool. The function's docstring tells the model what the tool does.
Checkpoint 1 of 6· Fill the gap
Which decorator turns this function into a LangChain tool?
from langchain_core.tools import tool
from langgraph.prebuilt import create_react_agent
from databricks_langchain import ChatDatabricks
@ ?
def get_current_time() -> str:
"""Get the current date and time."""
from datetime import datetime
return datetime.now().isoformat()
agent = create_react_agent(
ChatDatabricks(endpoint="databricks-claude-sonnet-4-5"),
tools=[get_current_time],
)LangChain's decorator is @tool, imported from langchain_core.tools. @function_tool is the OpenAI Agents SDK equivalent.
Source: docs.databricks.comLocal tools also cost the least to deploy. "Local function tools don't require resource grants in databricks.yml because they run within the agent process." They become the wrong choice once a tool needs governed data, sharing across teams, or reuse from more than one library. That is when you move to Unity Catalog.
Sources1
2.Unity Catalog function tools: governed and portable
A Unity Catalog function is a SQL or Python function registered under catalog.schema. It is governed like any other securable, and many libraries can call it. Before wiring a function into an agent, test it directly.
result = client.execute_function(
function_name=f"{CATALOG}.{SCHEMA}.add_numbers",
parameters={"number_1": 36939.0, "number_2": 8922.4}
)
result.value # OUTPUT: '45861.4'There are two ways to expose the function to a LangChain agent. The first is UCFunctionToolkit from databricks_langchain, which wraps one or more functions as LangChain tools. "The toolkit ensures consistency across different AI libraries and adds helpful features like auto-tracing for retrievers."
from databricks_langchain import UCFunctionToolkit
# Create a toolkit with the Unity Catalog function
func_name = f"{CATALOG}.{SCHEMA}.add_numbers"
toolkit = UCFunctionToolkit(function_names=[func_name])
tools = toolkit.toolsThe documentation then builds an agent with LangChain's create_tool_calling_agent. The tools are exposed to the model through a prompt template that has an agent_scratchpad placeholder. That sample uses AgentExecutor only "for simplicity". For production workloads, Databricks points you to the agent authoring workflow on Databricks Apps.
# Define the prompt
prompt = ChatPromptTemplate.from_messages(
[
(
"system",
"You are a helpful assistant. Make sure to use tools for additional functionality.",
),
("placeholder", "{chat_history}"),
("human", "{input}"),
("placeholder", "{agent_scratchpad}"),
]
)
# Enable automatic tracing
mlflow.langchain.autolog()
# Define the agent, specifying the tools from the toolkit above
agent = create_tool_calling_agent(llm, tools, prompt)Checkpoint 2 of 6· Put it in order
Put the steps for giving a LangChain agent a Unity Catalog tool in order
- 1.Create the Unity Catalog function under a catalog and schema
- 2.Pass toolkit.tools to create_tool_calling_agent with the LLM and prompt
- 3.Wrap it with UCFunctionToolkit(function_names=[...])
- 4.Test it with client.execute_function using its fully qualified name
You create and test the function first, then expose it through the toolkit. The agent can only use the tools after they are passed in.
“Once you have created and tested your Unity Catalog function, choose one of the following approaches to add it to your agent.”Source: docs.databricks.com
The second way is the one Databricks prefers: "Databricks recommends using MCP servers to add Unity Catalog functions to your agent." The managed MCP URL /api/2.0/mcp/functions/{catalog}/{schema} gives automatic tool discovery and built-in authentication. In LangGraph you connect to it with DatabricksMultiServerMCPClient. Unlike local tools, an app needs an EXECUTE grant on the function in databricks.yml.
mcp_client = DatabricksMultiServerMCPClient([
DatabricksMCPServer(
name="uc-functions",
url=f"{host}/api/2.0/mcp/functions/<catalog>/<schema>",
workspace_client=workspace_client,
),
])| Library | Helper for UC function tools | Agent construct shown |
|---|---|---|
| LangChain / LangGraph | databricks_langchain UCFunctionToolkit, or DatabricksMultiServerMCPClient (MCP) | create_tool_calling_agent / create_react_agent |
| OpenAI Agents SDK | McpServer.from_uc_function, or unitycatalog.ai.openai.toolkit UCFunctionToolkit | Agent with mcp_servers |
| LlamaIndex | unitycatalog.ai.llama_index.toolkit UCFunctionToolkit | ReActAgent.from_tools |
Checkpoint 3 of 6· Exam question
A team has already registered a Python function in Unity Catalog that looks up current inventory levels by SKU, and access to it is controlled through Unity Catalog grants. They want a LangChain agent to call this exact governed function as a tool, without re-implementing the lookup logic or bypassing UC access control. Which approach should they use?
Correct answer: A — Wrap the function with `UCFunctionToolkit(function_names=["catalog.schema.inventory_lookup"])` so LangChain calls the governed lookup through its existing tool list.
- A. This is correct: `UCFunctionToolkit` wraps a registered Unity Catalog function directly, and the resulting tools list can be passed straight into the agent so the original governed function is invoked with its access controls intact.
- B. Reimplementing the lookup as a local `@tool` duplicates the logic outside Unity Catalog and means the agent's calls are no longer subject to UC grants or lineage tracking, defeating the governance goal.
- C. Routing through a separate REST endpoint on Model Serving adds an extra service to build and maintain and does not use the existing Unity Catalog function client, so it does not meet the requirement to call the function as-is.
- D. Exporting results to a scheduled Delta table introduces staleness and an unnecessary data pipeline instead of directly calling the already-registered, access-controlled function.
Sources2
3.Custom retriever tools for unstructured data
For document Q&A, the agent needs a tool that searches an index. The Databricks AI Bridge packages (databricks-langchain, databricks-openai) include "helper functions like from_vector_search and from_uc_function to create retrievers from existing Databricks resources." In LangChain, VectorSearchRetrieverTool wraps an AI Search index as a tool you can bind to any chat model.
from databricks_langchain import VectorSearchRetrieverTool, ChatDatabricks
# Initialize the retriever tool.
vs_tool = VectorSearchRetrieverTool(
index_name="catalog.schema.my_databricks_docs_index",
tool_name="databricks_docs_retriever",
tool_description="Retrieves information about Databricks products from official Databricks documentation."
)
# Run a query against the vector search index locally for testing
vs_tool.invoke("Databricks Agent Framework?")
# Bind the retriever tool to your Langchain LLM of choice
llm = ChatDatabricks(endpoint="databricks-claude-sonnet-4-5")
llm_with_tools = llm.bind_tools([vs_tool])Checkpoint 4 of 6· Check yourself
An agent with a VectorSearchRetrieverTool rarely calls it, even for questions the index can answer. Which change does the documentation point to?
The agent decides whether to call a tool from its description, so a vague description leads to missed calls.
“Provide a descriptive tool_description to help the agent understand the tool and determine when to invoke it.”Source: docs.databricks.com
Sources3
4.Structured data, Genie agents and orchestration
The type of data decides which tool you need. For unstructured text, "AI Search automatically indexes your knowledge base at scale for semantic or hybrid search," which is what a retriever tool queries. For tables, the options are serverless SQL, online feature stores, and Genie: "Genie agents can be used in multi-agent systems to answer natural language queries about your structured data." (These sources do not describe the Genie conversation API itself; this lesson covers only where Genie fits among the tool choices.)
When one application needs several of these, you can use a Supervisor Agent, which "orchestrates Genie Agents, agent endpoints, Unity Catalog functions, MCP servers, and custom agents." A code-first alternative is a DSPy multi-agent system over Genie Agents and Model Serving agents. Both MCP servers and Unity Catalog functions "provide Unity Catalog-based governance", so access control stays the same whichever library calls them.
Checkpoint 5 of 6· Exam question
An engineering team already has a Mosaic AI Vector Search (AI Search) index over chunked support documents and wants to give a LangChain agent read-only access to that index as a callable retrieval tool, with minimal custom code and while keeping MLflow tracing on retrieval calls. Which option best satisfies this?
Correct answer: A — Instantiate `VectorSearchRetrieverTool` from the `databricks-langchain` package pointed at the existing AI Search index and bind it to the agent's tools.
- A. This is correct: the `VectorSearchRetrieverTool` helper in the LangChain integration package is built specifically to point at an existing AI Search index and expose it as an agent tool with built-in MLflow tracing, requiring minimal custom code.
- B. Hand-writing a UC function that makes raw REST calls to the vector search API duplicates functionality the retriever tool already provides and adds unnecessary custom code and maintenance burden.
- C. Exporting the index into a Delta table for keyword filtering abandons the vector similarity search the index was built for and replaces it with a much weaker keyword-matching approach.
- D. Reimplementing similarity search locally ignores the already-built AI Search index entirely and requires substantial custom code, which conflicts with the minimal-code requirement.
Checkpoint 6 of 6· Check yourself
A multi-agent app must answer "What were Q3 sales by region?" from Delta tables and "What does our refund policy say?" from PDFs. Which pairing of tools fits?
Genie agents answer natural-language questions over structured data. AI Search indexes unstructured documents for semantic or hybrid search.
“Genie agents can be used in multi-agent systems to answer natural language queries about your structured data.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.UCFunctionToolkit is the recommended way to add Unity Catalog functions to an agent.Why is that wrong?
UCFunctionToolkit works, but Databricks recommends MCP servers, which provide automatic tool discovery and built-in authentication.
Covered in Unity Catalog function tools: governed and portable
2.Every agent tool, including a local @tool function, needs a resource grant in databricks.yml.Why is that wrong?
Local function tools run inside the agent process and need no grants. Unity Catalog functions reached through MCP do need an EXECUTE grant.
Covered in Local function tools: the lightest option
3.The LangChain AgentExecutor sample in the docs is the production pattern for Databricks agents.Why is that wrong?
The AgentExecutor example is there for simplicity. Production workloads should follow the agent authoring workflow on Databricks Apps.
Covered in Unity Catalog function tools: governed and portable
Practise it for real
Give a LangChain tool-calling agent a governed Unity Catalog function and confirm it calls the function.
1.Run client.execute_function on your {CATALOG}.{SCHEMA}.add_numbers function with number_1=36939.0 and number_2=8922.4.
Why: Testing the function on its own first separates function bugs from agent bugs.
You should see: result.value returns '45861.4'.
2.Create UCFunctionToolkit(function_names=[func_name]) and read toolkit.tools.
Why: The toolkit turns the UC function into LangChain-compatible tools.
You should see: A list of tools containing the add_numbers function.
3.Call mlflow.langchain.autolog(), then build the agent with create_tool_calling_agent(llm, tools, prompt) using a ChatDatabricks LLM and the tool-calling prompt template.
Why: Autologging records each tool call as a trace, which matters most on serverless compute, where it is not on by default.
You should see: An agent object, with no errors.
4.Wrap it in AgentExecutor(agent=agent, tools=tools, verbose=True) and invoke it with {"input": "What is 36939.0 + 8922.4?"}.
Why: This checks that the model chooses the tool instead of doing the arithmetic itself.
You should see: Verbose output shows the add_numbers tool being called, and a trace appears in the active MLflow experiment.
Stuck? Get a nudge
If the agent answers without calling the tool, check that tools were passed to both create_tool_calling_agent and AgentExecutor.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“For operations that don't require external data sources or APIs, define tools directly in your agent code.”
↩︎ Local function tools: the lightest option“Local function tools don't require resource grants in databricks.yml because they run within the agent process.”
↩︎ Local function tools: the lightest option“Local function tools don't require resource grants in databricks.yml because they run within the agent process.”
↩︎ Exam trap 2 - 2.
“The toolkit ensures consistency across different AI libraries and adds helpful features like auto-tracing for retrievers.”
↩︎ Unity Catalog function tools: governed and portable“Databricks recommends using MCP servers to add Unity Catalog functions to your agent.”
↩︎ Unity Catalog function tools: governed and portable“The MCP approach provides a simpler integration with automatic tool discovery and built-in authentication support.”
↩︎ Unity Catalog function tools: governed and portable“Databricks recommends using MCP servers to add Unity Catalog functions to your agent.”
↩︎ Exam trap 1“This example authors a simple agent using LangChain AgentExecutor API for simplicity. For production workloads, use the agent authoring workflow”
↩︎ Exam trap 3“Once you have created and tested your Unity Catalog function, choose one of the following approaches to add it to your agent.”
↩︎ Checkpoint - 3.
“helper functions like from_vector_search and from_uc_function to create retrievers from existing Databricks resources”
↩︎ Custom retriever tools for unstructured data“Provide a descriptive tool_description to help the agent understand the tool and determine when to invoke it.”
↩︎ Checkpoint - 4.
“AI Search automatically indexes your knowledge base at scale for semantic or hybrid search.”
↩︎ Structured data, Genie agents and orchestration“Tool support includes MCP servers and Unity Catalog Functions, both of which provide Unity Catalog-based governance.”
↩︎ Structured data, Genie agents and orchestration“Genie agents can be used in multi-agent systems to answer natural language queries about your structured data.”
↩︎ Checkpoint - 5.
“Build a supervisor agent that orchestrates Genie Agents, agent endpoints, Unity Catalog functions, MCP servers, and custom agents.”
↩︎ Structured data, Genie agents and orchestration