What you will be able to do
- Pick the simplest design pattern that fits a GenAI use case: LLM + prompt, deterministic chain, single agent or multi-agent
- Explain why the choice of authoring framework (LangChain, LangGraph, OpenAI Agents SDK, LlamaIndex, DSPy) stays open on Databricks, and what ResponsesAgent adds
- Recognise when DSPy fits better than hand-written LangChain prompts
- Turn on MLflow tracing for LangChain and know where autologging is not automatic
Key concept
Framework-agnostic agent authoring — On Databricks, you choose an orchestration library (LangChain, LangGraph, OpenAI Agents SDK, LlamaIndex, DSPy) for how well it fits the task, not for platform lock-in. You wrap the agent in MLflow's ResponsesAgent interface so that tracing, evaluation, deployment and monitoring work the same whichever library you picked.
1.Pick the pattern before the library
Asking "should I use LangChain?" too early skips a step. First decide how much the LLM should control. Databricks describes a range of options, from a single LLM call to a multi-agent system, and its advice is to start at the simple end: "When building any AI-powered application, start simple." Add agentic behaviour only when the application actually needs the extra flexibility or model-driven decisions.
| Pattern | When to use | Pros | Cons |
|---|---|---|---|
| LLM + prompt | Generic Q&A; quick prototype for short-term use | Very simple; easy to create | Minimal customization |
| Deterministic chain | Well-defined tasks; static pipelines such as basic RAG | Simple; easy to audit | Inflexible; requires code changes to adapt |
| Single-agent system | Moderate to complex queries in the same domain | Flexible; a good "default" | Less predictable; must guard against repeated or incorrect tool calls |
| Multi-agent system | Large or cross-functional domains | Highly modular; scales to big domains | Complex to orchestrate; harder to trace and debug |
What separates a chain from an agent is who decides which tools run. In a deterministic chain, the developer sets which tools or models are called, in what order and with which parameters. "The LLM does not make decisions about which tools to call or in what order." Agentic designs hand those decisions to the model. That buys flexibility, but more agentic approaches "come with the cost of extra complexity and potential latency." So the pattern you pick largely determines which framework features you need: a fixed pipeline (prompt | llm | parser) for a chain, or a tool-calling agent loop for an agent.
Checkpoint 1 of 5· Check yourself
Which statement about a deterministic chain is correct?
A chain follows a predefined workflow for every request. Tool selection and ordering are developer decisions, which makes chains predictable and easy to audit.
“The LLM does not make decisions about which tools to call or in what order.”Source: docs.databricks.com
Sources1
2.LangChain or something else? The platform leaves it open
Once you know the pattern, you choose a library to write it in. Custom agents on Databricks are not tied to one library. The "Build custom agents" capability "Supports agents written with any authoring library, including LangGraph, LangChain, OpenAI, and LlamaIndex." The Databricks Apps agent template, for example, uses the OpenAI Agents SDK for conversation management and tool orchestration, and the same pages show LangGraph equivalents side by side.
The library is interchangeable because of a wrapper. Databricks recommends MLflow ResponsesAgent: "You can author agents using any framework. The key is wrapping your agent with MLflow ResponsesAgent interface." A wrapped agent gets out-of-the-box compatibility with AI Playground, Agent Evaluation and Agent Monitoring. It also gets streaming output, multi-agent support and tool-calling message history. In practice, you should choose LangChain or a similar library for what it does well, such as its tool abstractions, prebuilt agent loops or retriever integrations. The rest of the Databricks lifecycle works either way.
In AI Playground, the no-code option. You select an LLM, add tools and chat with the agent to test its responses. Then you export it to code for deployment.
Checkpoint 2 of 5· Exam question
Which Python package provides purpose-built LangChain tool wrappers for Unity Catalog functions, letting a governed UC function be surfaced directly as a LangChain-compatible tool for an agent?
Correct answer: A — databricks-langchain
- A. This package provides the LangChain integrations, including a toolkit that wraps Unity Catalog functions so they can be passed directly into a LangChain agent's tool list.
- B. This package is a client for running Spark workloads against a remote Databricks cluster from a local environment; it has nothing to do with wrapping UC functions as LangChain tools.
- C. This package is a general-purpose REST API client for managing Databricks workspace resources such as clusters and jobs, not a LangChain tool integration library.
- D. This package is used for managing feature tables and feature lookups in the Feature Store, which is unrelated to exposing UC functions as LangChain tools.
Checkpoint 3 of 5· Check yourself
A team has a working LlamaIndex agent and wants Agent Evaluation and Agent Monitoring on Databricks. What is the recommended approach?
ResponsesAgent is the compatibility layer. Any existing agent wrapped in it gains Playground, Evaluation and Monitoring support.
“Wrap any existing agent using the ResponsesAgent interface to get out-of-the-box compatibility with AI Playground, Agent Evaluation, and Agent Monitoring.”Source: docs.databricks.com
3.When DSPy beats hand-written prompts
LangChain-style libraries mostly have you write prompts and wire them into chains. DSPy takes a different approach: "DSPy is a framework for programmatically defining and optimizing AI agents. DSPy can automate prompt engineering and orchestrate LLM fine-tuning to improve performance." Instead of hand-tuning a prompt string, you declare what each step takes in and produces, then let an optimizer tune the prompts against a metric.
| Component | Role |
|---|---|
| Module | Handles a specific text transformation (answering, summarizing); replaces hand-written prompts and can learn from examples |
| Signature | Natural-language description of a module's inputs and outputs, e.g. "question -> answer" |
| Compiler | Optimizer that adjusts modules to meet a performance metric, by better prompts or fine-tuning |
| Program | A set of modules connected into a pipeline for complex tasks |
DSPy is not an all-or-nothing switch. Databricks publishes notebooks that "show how to migrate LangChain model code to DSPy and optimize it for better performance". Other DSPy examples build RAG over an AI Search index, and a multi-agent system that orchestrates Genie Agents, Model Serving agents and UC function-calling agents. DSPy is a good choice when prompt quality is the bottleneck and you can measure it. Outside DSPy, Databricks also offers prompt work interactively in AI Playground or through MLflow Prompt Optimization.
Checkpoint 4 of 5· Match them up
Match each DSPy component to what it does
Tap a term, then the definition that fits it.
Signatures describe behaviour, modules carry it out, programs chain modules together, and the compiler optimizes the whole pipeline toward a metric.
“It improves LM pipelines by adjusting modules to meet a performance metric, either by generating better prompts or fine-tuning models.”Source: docs.databricks.com
4.Pairing LangChain with MLflow Tracing
Observability is one practical reason to pick LangChain on Databricks. MLflow provides automatic tracing for it. With a single call, "nested traces are automatically logged to the active MLflow Experiment upon invocation of chains."
import mlflow
mlflow.langchain.autolog()Here is the catch: "On serverless compute clusters, autologging is not automatically enabled." You must call mlflow.langchain.autolog() explicitly. Auto tracing covers invoke, batch, stream, their async variants, __call__ for Chains and AgentExecutors, and get_relevant_documents for retrievers. Packaging differs by stage. Use the full mlflow[databricks] package for development. For production, install mlflow-tracing, which "is optimized for production use." If you need custom span attributes, you can subclass MlflowLangchainTracer.
Checkpoint 5 of 5· Check yourself
Which package does the documentation recommend for tracing a LangChain app in production?
mlflow[databricks] is the full development package. mlflow-tracing is the lighter package intended for production deployments.
“The mlflow-tracing package is optimized for production use.”Source: docs.databricks.com
Sources6
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The most agentic design (multi-agent, or at least a tool-calling agent) is always the best choice for a GenAI app.Why is that wrong?
Databricks advises starting simple. Deterministic chains are preferred for well-defined tasks where auditing and latency matter, and agentic designs add complexity and latency.
Covered in Pick the pattern before the library
2.LangChain traces are always captured automatically on Databricks, so there is no need to call mlflow.langchain.autolog().Why is that wrong?
On serverless compute, autologging is off by default and you must enable it explicitly.
Covered in Pairing LangChain with MLflow Tracing
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“When building any AI-powered application, start simple.”
↩︎ Pick the pattern before the library“The LLM does not make decisions about which tools to call or in what order.”
↩︎ Pick the pattern before the library“More agentic approaches offer greater flexibility and potential, but they come with the cost of extra complexity and potential latency.”
↩︎ Exam trap 1“When you want to minimize latency by avoiding multiple LLM calls for orchestration decisions.”
↩︎ Prediction - 2.
“Supports agents written with any authoring library, including LangGraph, LangChain, OpenAI, and LlamaIndex.”
↩︎ LangChain or something else? The platform leaves it open - 3.
“You can author agents using any framework. The key is wrapping your agent with MLflow ResponsesAgent interface.”
↩︎ LangChain or something else? The platform leaves it open“ResponsesAgent lets you build agents with any third-party framework, then integrate it with Databricks AI features”
↩︎ Key concept“Wrap any existing agent using the ResponsesAgent interface to get out-of-the-box compatibility with AI Playground, Agent Evaluation, and Agent Monitoring.”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/agents/dspyOfficial docs
“DSPy is a framework for programmatically defining and optimizing AI agents. DSPy can automate prompt engineering and orchestrate LLM fine-tuning to improve performance.”
↩︎ When DSPy beats hand-written prompts“These notebooks show how to migrate LangChain model code to DSPy and optimize it for better performance.”
↩︎ When DSPy beats hand-written prompts“It improves LM pipelines by adjusting modules to meet a performance metric, either by generating better prompts or fine-tuning models.”
↩︎ Checkpoint - 5.
“Prompt engineering can be done interactively using AI Playground, or through data-driven optimization using MLflow Prompt Optimization.”
↩︎ When DSPy beats hand-written prompts - 6.
“nested traces are automatically logged to the active MLflow Experiment upon invocation of chains.”
↩︎ Pairing LangChain with MLflow Tracing“get_relevant_documents (for retrievers)”
↩︎ Pairing LangChain with MLflow Tracing“On serverless compute clusters, autologging is not automatically enabled.”
↩︎ Exam trap 2“The mlflow-tracing package is optimized for production use.”
↩︎ Checkpoint