What you will be able to do
- Recognise the business problems Knowledge Assistant is built for and the knowledge sources it accepts
- Recognise when Information Extraction is the right brick, and how its schema, evaluation and deployment work
- Choose between a cited chat answer and structured per-document fields for a given requirement
Key concept
The output shape picks the brick — Knowledge Assistant answers people's questions about a document collection, with citations. Information Extraction turns each document into the fields a schema defines. Decide which shape the business needs before you pick a brick.
1.Knowledge Assistant: cited answers over your documents
Agent Bricks gives you ready-made agents for common problems, so you don't have to hand-assemble a retriever, a prompt and an LLM. Knowledge Assistant covers the most common one: people asking questions about a body of documents. It builds a chatbot that answers with citations. It uses an Instructed Retriever approach, which addresses limitations of traditional RAG. The documentation's own examples are answering from product documentation, answering employee HR-policy questions, and answering customers from a support knowledge base. The result is an agent endpoint. You can chat with it in AI Playground or call it from your own applications. Subject matter experts improve it by adding questions and natural-language guidelines in the build experience, so you don't retrain anything by hand.
You can add up to 10 knowledge sources. Each one is one of three types, and each type comes with its own constraints. For every source you also write a "Describe the content" note. This is not just a label: the agent reads it to decide which source to search for a given question.
| Source type | What it requires |
|---|---|
| Files in a Volume | Supported types: txt, pdf, md, ppt/pptx, doc/docx. Files over 100 MB are skipped, and so are PDF/DOC/DOCX/PPT/PPTX files over 500 pages (for PowerPoint, each slide counts as a page) |
| Files in a Table | A streaming table or a table with Change data feed enabled, a content column of BINARY or STRING, and a metadata or _metadata STRUCT column (file_path, file_name, file_size, file_modification_time) |
| AI Search Index | Must use databricks-gte-large-en, databricks-bge-large-en, or databricks-qwen3-embedding-0-6b. The Doc URI Column supplies citations and the Text Column supplies the retrieved text |
Checkpoint 1 of 5· Check yourself
A team already has an AI Search index over its support articles, built with an embedding model that is not one of the three Databricks-hosted models listed for Knowledge Assistant. What is true?
AI Search index sources are supported only when the index uses databricks-gte-large-en, databricks-bge-large-en, or databricks-qwen3-embedding-0-6b.
“AI Search indexes are only supported if the index uses one of the following embedding models”Source: docs.databricks.com
Sources1
2.Information Extraction: from documents to schema fields
Some requirements aren't about answering questions. They're about turning a pile of documents into rows. Information Extraction does this. You define a schema, and each document or text value becomes a set of structured fields that you can use for analysis, reporting, or downstream agents. The documented examples are legal parties and terms from contracts, line items and payment terms from invoices, and key details from medical records. You can use it from the Agent Bricks UI, in SQL, and through the REST API. Your input data must be in a Unity Catalog volume or table, with at least one file or row.
There are three ways to define the schema. You can describe what you want in natural language and click Generate Schema, which produces a JSON schema with field names and definitions. You can add fields by hand with a name, type and description. Or you can edit the JSON directly. Three optional settings change the output: Precision mode (better accuracy for nuanced fields and long documents), Citations, and Confidence scores. To refine the results, give natural-language feedback on individual inputs, which auto-tunes the field descriptions, or edit the descriptions yourself. Versions let you compare configurations and restore an earlier one.
Run an evaluation against a labeled Unity Catalog table. The table needs an input column (STRING text or the VARIANT output of ai_parse_document) and a ground-truth column holding the expected result as a JSON string that conforms to your schema. The run reports overall field accuracy, documents fully correct, and per-field accuracy, precision, recall and F1, and it compares the results with the previous run.
When the agent performs well enough, click Use Agent. You can either open a SQL query that applies ai_extract with your schema to the whole volume or table, or create a Lakeflow pipeline that runs on a schedule and writes new extractions to a streaming table. Applications call the REST endpoint instead. Its default limit is 120 requests per minute per workspace.
Checkpoint 2 of 5· Fill the gap
Which SQL function completes this extraction query?
SELECT ? (
'Invoice #12345 from Acme Corp for $1,250.00 dated 2024-01-15',
'{
"invoice_id": {"type": "string", "description": "Unique invoice identifier"},
"vendor_name": {"type": "string", "description": "Legal business name"},
"total_amount": {"type": "number", "description": "Total invoice amount"},
"invoice_date": {"type": "string", "description": "Date in YYYY-MM-DD format"}
}',
options => map('version', '2.1')
);Information Extraction runs at scale through the ai_extract SQL function, which takes the content and a JSON schema of the fields to return.
Source: docs.databricks.comCheckpoint 3 of 5· Put it in order
Put the Information Extraction workflow in order
- 1.Click Save and run extraction and refine with feedback
- 2.Run an evaluation against a ground-truth table
- 3.Click Use Agent to run in SQL or create a Lakeflow pipeline
- 4.Define the extraction schema (generated, manual, or JSON)
- 5.Select a Unity Catalog volume, a table column, or uploaded files and click Create Agent
You pick the data before you can create the agent, the schema drives each extraction run, evaluation measures the current version, and you deploy only once you're satisfied with performance.
“After you're happy with the agent's performance, use the agent to extract information.”Source: docs.databricks.com
Sources2
3.Choosing between the two
Both bricks read unstructured documents, which is why they get confused. The deciding question is who consumes the output. When a person asks a free-form question and wants a grounded answer with a source, use Knowledge Assistant. When a table, report or downstream system needs the same fields from every document, use Information Extraction. One gives a conversation over the whole collection. The other gives a schema applied to each document.
Checkpoint 4 of 5· Match them up
Match each requirement to the brick that fits it
Tap a term, then the definition that fits it.
Free-form questions answered with citations go to Knowledge Assistant. Fixed fields pulled from every document go to Information Extraction. Coordinating several agents and tools on one task is the job of Supervisor Agent.
“Extracting line items and payment terms from invoices.”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
A retail company's HR team receives repetitive questions about vacation policy, benefits enrollment, and reimbursement procedures. The answers exist across dozens of PDF and Word policy documents stored in a Unity Catalog volume, and HR wants a chatbot that cites the specific document and section it used to answer, with minimal custom engineering. Which Agent Bricks component should the team use to build this solution?
Correct answer: A — Knowledge Assistant, which ingests the PDF and Word files from Unity Catalog volumes and produces citation-backed answers via a guided Instructed Retriever approach
- A. Knowledge Assistant is designed exactly for this scenario: it ingests unstructured documents from Unity Catalog volumes and generates chatbot answers with citations back to the source documents, using a guided, low-code workflow rather than hand-built retrieval code.
- B. Multi-Agent Supervisor coordinates multiple already-built subagents to handle different task domains, but this scenario has a single knowledge domain (HR policy documents), so there is nothing to route between and no need for a supervisor layer.
- C. Information Extraction is built to pull specific structured fields out of documents into a schema, not to power a conversational, citation-backed question-answering chatbot, so it does not match the stated requirement.
- D. A hand-built LangChain retrieval chain could technically answer the questions, but it requires engineers to author, tune, and maintain the retrieval and prompt logic themselves, which conflicts with the team's goal of minimal custom engineering.
Sources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A Knowledge Assistant chatbot is the right tool for pulling invoice totals and dates into a reporting table.Why is that wrong?
Knowledge Assistant returns cited conversational answers. Information Extraction applies a defined schema to each document and produces structured fields, which you can run at scale with ai_extract or a scheduled Lakeflow pipeline.
Covered in Information Extraction: from documents to schema fields
2.Any Unity Catalog table of documents can be attached to Knowledge Assistant as a 'Files in a Table' source.Why is that wrong?
The table must be a streaming table or have Change data feed enabled, and it needs a content column and a metadata STRUCT column. Tables from sources like Jira or Confluence may need to be transformed first.
Covered in Knowledge Assistant: cited answers over your documents
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Use Knowledge Assistant to create a chatbot that can answer questions about your documents and provide high-quality responses with citations.”
↩︎ Knowledge Assistant: cited answers over your documents“describe what content the knowledge source contains to help the agent understand when to use this data source.”
↩︎ Knowledge Assistant: cited answers over your documents“You can provide up to 10 knowledge sources.”
↩︎ Knowledge Assistant: cited answers over your documents“Answer employee questions related to HR policies.”
↩︎ Choosing between the two“The table is either a streaming table or a table with Change data feed enabled.”
↩︎ Exam trap 2“Files larger than 100 MB are automatically skipped during ingestion and are not included in the knowledge base.”
↩︎ Prediction“AI Search indexes are only supported if the index uses one of the following embedding models”
↩︎ Checkpoint - 2.
“Information extraction is available through the Agent Bricks UI, in SQL, and with the REST API.”
↩︎ Information Extraction: from documents to schema fields“A ground truth column with the expected extraction result as a JSON string.”
↩︎ Information Extraction: from documents to schema fields“The default rate limit for ai_extract REST API requests is 120 requests per minute per workspace.”
↩︎ Information Extraction: from documents to schema fields“Information Extraction transforms unstructured documents and text into key, structured insights using a defined schema.”
↩︎ Key concept“Create a Lakeflow pipeline that runs on scheduled intervals to invoke your agent on new data.”
↩︎ Exam trap 1“After you're happy with the agent's performance, use the agent to extract information.”
↩︎ Checkpoint“Extracting line items and payment terms from invoices.”
↩︎ Checkpoint