What you will be able to do
- Match a data source to the right Databricks agent tool: Genie, a Unity Catalog function, AI Search or the SQL MCP server
- Write tool definitions (type hints, docstrings, descriptions) that let an LLM decide when and in what order to call tools
- Explain how one tool's output becomes the next tool's input in a bounded tool-calling loop
1.Matching each knowledge source to a tool
Ordering tools well depends on defining them well, and the first choice is which kind of tool fits each data source. Databricks gives a recommendation for each case.
For questions about data in Unity Catalog tables, Databricks recommends Genie Agents. A Genie Agent holds up to 25 Unity Catalog tables that Genie can keep in context and query in natural language, and your agent reaches it through a managed MCP URL. When the agent should run AI-generated SQL through a SQL warehouse, use the managed Databricks SQL MCP server rather than building a custom tool. When the query is known ahead of time and only the parameters change, use a Unity Catalog SQL function. For unstructured documents, the managed AI Search MCP server gives the agent access to an AI Search index (formerly Vector Search).
Unity Catalog functions can also take actions, running custom logic that goes beyond language generation. Outside the structured-retrieval case, though, Databricks recommends MCP servers or logic defined directly in agent code, for faster execution and per-user authentication.
| Need | Recommended tool | Managed MCP path |
|---|---|---|
| Natural-language questions over Unity Catalog tables | Genie Agent | /api/2.0/mcp/genie/{genie_space_id} |
| Run AI-generated SQL through a SQL warehouse | Databricks SQL MCP server | /api/2.0/mcp/sql |
| Known query; the agent supplies parameters | Unity Catalog SQL function | /api/2.0/mcp/functions/{catalog}/{schema} |
| Search unstructured documents | AI Search index | /api/2.0/mcp/ai-search/{catalog}/{schema}/{index_name} |
Checkpoint 1 of 3· Match them up
Match each agent requirement to the tool Databricks recommends
Tap a term, then the definition that fits it.
Genie handles natural-language table queries and UC functions handle fixed, parameterised queries. AI Search covers unstructured text, and the SQL MCP server runs AI-generated SQL.
“Create a structured retrieval tool using Unity Catalog SQL functions when the query is known ahead of time and the agent provides the parameters.”Source: docs.databricks.com
2.Descriptions are how the LLM decides order
A tool's description is the LLM's only guide to when to call it. A Python callable registered with create_python_function must meet several rules. It needs type hints on every argument and on the return value. It cannot use *args or **kwargs. Its docstring must follow Google syntax, because the toolkit parses the docstring to tell the LLM how and when to use the function. Any libraries must be imported inside the function body.
In a multi-stage agent, the description can also say how a tool's output feeds the next step. In the SQL function below, the COMMENT tells the LLM that the returned customer ID can be used in other queries. That gives the LLM a reason to call this lookup first and pass its result into later tools.
CREATE OR REPLACE FUNCTION main.default.lookup_customer_info(
customer_name_input STRING COMMENT 'Name of the customer whose info to look up'
)
RETURNS STRING
COMMENT 'Returns metadata about a particular customer, given the customer''s name, including the customer''s email and ID. The
customer ID can be used for other queries.'
RETURN SELECT CONCAT(
'Customer ID: ', customer_id, ', ',
'Customer Email: ', customer_email
)
FROM main.default.customer_data
WHERE customer_name = customer_name_input
LIMIT 1;Checkpoint 2 of 3· Fill the gap
Which parameter gives the LLM the text it uses to decide when to call this retriever tool?
vs_tool = VectorSearchRetrieverTool(
index_name="catalog.schema.my_databricks_docs_index",
tool_name="databricks_docs_retriever",
? ="Retrieves information about Databricks products from official Databricks documentation."
)tool_description carries the text the LLM reads when it decides whether to call this retriever. tool_name identifies the tool, and index_name points to the index.
Checkpoint 3 of 3· Check yourself
A developer registers a Python tool with create_python_function, but it fails when the agent runs it. Which rule is the developer most likely to have broken?
Unity Catalog does not resolve imports placed outside the function, so the tool fails at run time. The other three options describe correct practice.
“Libraries must be imported within the function's body. Imports outside the function will not be resolved when running the tool.”Source: docs.databricks.com
Sources3
3.The tool-calling loop that sequences the calls
Once tools are defined, they are bound to the LLM and run inside a loop. In the Databricks volume example, the user asks the agent to list the files in a volume, then read one and summarise it. The agent has two tools, list_volume_files and read_volume_file. Neither the code nor the developer fixes their order. On each pass the LLM returns tool calls, the loop runs them and adds the results to the message history as tool messages, and the LLM is called again. The loop ends when the LLM responds without requesting a tool, or when the iteration cap is reached.
for _ in range(5): # max iterations
response = llm_with_tools.invoke(messages)
messages.append(response)
if not response.tool_calls:
break
for tc in response.tool_calls:
result = tool_map[tc["name"]].invoke(tc["args"])
messages.append(ToolMessage(content=result, tool_call_id=tc["id"]))Action tools carry a risk that retrieval tools do not. Databricks warns that executing arbitrary code in an agent tool can expose sensitive or private information the agent has access to. You are responsible for running only trusted code and for setting guardrails and permissions. When an agent runs in Databricks Apps, it gets access to a Unity Catalog function through an explicit EXECUTE grant on that function.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Unity Catalog functions are the recommended way to build every agent tool on Databricks.Why is that wrong?
Databricks recommends UC functions specifically for structured retrieval with a known query. For most other cases it recommends MCP servers or logic written in agent code.
Covered in Matching each knowledge source to a tool
2.Calling Genie through its managed MCP server automatically passes the earlier conversation to Genie.Why is that wrong?
The Genie MCP server invokes Genie as a tool, so no history is passed. To pass conversation context deterministically, call Genie as an agent in a multi-agent system.
Covered in Matching each knowledge source to a tool
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“If your agent needs to query data in Unity Catalog tables, Databricks recommends using Genie Agents.”
↩︎ Matching each knowledge source to a tool“The managed MCP server for Genie invokes Genie as an MCP tool, which means history isn't passed when invoking Genie APIs.”
↩︎ Exam trap 2“Create a structured retrieval tool using Unity Catalog SQL functions when the query is known ahead of time and the agent provides the parameters.”
↩︎ Checkpoint - 2.
“Use the Databricks-managed AI Search MCP server to give your agent access to a Databricks AI Search index.”
↩︎ Matching each knowledge source to a tool“Bind the tools to an LLM and run a tool-calling loop:”
↩︎ The tool-calling loop that sequences the calls“Provide a descriptive tool_description to help the agent understand the tool and determine when to invoke it.”
↩︎ Prediction - 3.
“perform specific tasks that extend the capabilities of LLMs beyond language generation”
↩︎ Matching each knowledge source to a tool“Write clear descriptions for your function and its arguments to help the LLM understand how and when to use the function.”
↩︎ Descriptions are how the LLM decides order“Variable arguments such as *args and **kwargs are not supported.”
↩︎ Descriptions are how the LLM decides order“Executing arbitrary code in an agent tool can expose sensitive or private information that the agent has access to.”
↩︎ The tool-calling loop that sequences the calls“In most other use cases, Databricks recommends MCP servers or defining the logic directly in agent code”
↩︎ Exam trap 1“Libraries must be imported within the function's body. Imports outside the function will not be resolved when running the tool.”
↩︎ Checkpoint