What you will be able to do
- Explain why the data corpus limits what a RAG application can answer, no matter how well retrieval and prompting are tuned
- Work out which document types a RAG use case needs, starting from the questions users will ask
- Explain why domain experts and an evaluation set are the way to check corpus coverage
- Choose between candidate sources on relevance, freshness, ownership, structure, duplication and noise
Key concept
Corpus composition — Corpus composition means deciding which sources and content go into the knowledge base that a RAG application retrieves from. Retrieval can only return what the corpus contains, so this choice limits the quality of every answer the application gives.
1.The corpus sets the limit
A RAG application has two parts: an offline data pipeline that prepares and indexes documents, and an online chain that retrieves from that index and generates an answer. People usually tune the online part first, adjusting the prompt, the number of chunks, or the search type. All of those settings work on whatever the pipeline put into the index. If the documents that answer a question were never collected, the retriever cannot return them, and the LLM either says it does not know or, worse, fills the gap with a plausible guess.
That is why Databricks puts corpus selection first in the data pipeline, before parsing, chunking or embedding. The first component listed is "Corpus composition and ingestion: Select the right data sources and content based on the specific use case." Note the phrase "based on the specific use case". There is no corpus that is correct for every application. The right one follows from what the application is meant to answer.
The Databricks quality guidance treats this as a retrieval-quality problem with two failure modes. Context can be missing important information, or it can contain superfluous information. Leaving documents out causes the first. Adding documents carelessly causes the second. Choosing sources well means avoiding both.
Checkpoint 1 of 4· Check yourself
According to Databricks' RAG quality guidance, which two context problems make it hard to produce high-quality RAG output?
Retrieval quality fails when needed information is absent or when irrelevant material crowds the context. Choosing source documents affects both.
“missing important information or contains superfluous information”Source: docs.databricks.com
2.Start from the questions users will ask
Because "The correct data depends entirely on your application's specific requirements and goals", work backwards. List the kinds of questions the application must answer, then ask which documents contain those answers. The Databricks cookbook uses a customer support bot as its example and names four document types to consider. Each covers a different kind of user need. The table below matches each type to the kind of question it is likely to answer. The document types come from the source; the question column is an illustration.
| Document type | Typical user question it covers |
|---|---|
| Knowledge base documents | How do I do X? What is the policy on Y? |
| Frequently Asked Questions (FAQs) | Short, common questions with short, standard answers |
| Product manuals and specifications | What are the limits, settings or features of a specific product? |
| Troubleshooting guides | I see this error or symptom. How do I fix it? |
Use the same reasoning on any scenario the exam gives you. An HR benefits assistant needs current policy handbooks and plan documents, not last year's versions or unrelated engineering wikis. A contract-review assistant needs the contracts themselves and the clause playbooks. The data also does not have to be unstructured. Databricks notes that RAG can use unstructured data such as PDFs, Office documents and wikis, or structured data such as customer records and transaction tables. Which you need depends on whether the answers live in prose or in rows. Exam questions about choosing sources usually reward the option that matches the source to the task, not the one that includes the most data.
Checkpoint 2 of 4· Check yourself
A team is building a RAG assistant that answers questions about order status and recent purchases for logged-in customers. Which source belongs at the centre of its corpus?
The answers to order-status questions live in structured transaction data, and RAG can retrieve from structured sources. Blog posts and archives don't contain those answers.
“The data you use with RAG depends on your use case.”Source: docs.databricks.com
3.Domain experts decide what coverage means
Engineers building a RAG app are often not experts in its content, so they can't reliably tell which documents matter. Databricks' advice is direct: "Engage domain experts and stakeholders from the beginning of any project to help identify and curate relevant content that could improve the quality and coverage of your data corpus." Experts add two things. They know the queries users are likely to submit, and they can rank which information is most critical to include.
Coverage is something you test, not something you assume. The evaluation set links the corpus decision to measurement. Databricks describes an evaluation set as representative queries with ground-truth answers and, optionally, "the correct supporting documents that should be retrieved". When writing those labels, a question whose supporting document is missing from every candidate source shows a gap in the corpus before users find it. The guidance also says the evaluation set should be continually updated to reflect the changing nature of the indexed data. So corpus selection is reviewed again as the application and its data change, not done once.
Checkpoint 3 of 4· Check yourself
Early in a RAG project, what does Databricks say domain experts and stakeholders contribute to the corpus?
The guidance involves experts specifically in identifying and prioritising content and in predicting user queries. Technical settings are not their role here.
“They can provide insights into the types of queries that users are likely to submit”Source: docs.databricks.com
4.Judging source quality: relevant, current, unique and clean
Once you know which kinds of documents you need, compare candidate sources on quality. Four properties come up most often.
Ownership and freshness. RAG is most useful for knowledge the base model lacks: proprietary documents, and information that changes. Databricks notes that "A RAG application can supply the LLM with information from an updated knowledge base." Choose the authoritative, current version of a source over a stale copy. Prefer sources that can be ingested incrementally so the index stays current.
Uniqueness. Copies of the same document, such as repeated policy versions or mirrored wiki pages, use up retrieval slots with repeated text. The pipeline guidance includes a deduplication step to find and remove duplicates and near-duplicates.
Relevance. Off-topic documents add the superfluous context discussed in the first section. The pipeline also includes a filtering step to "Eliminate irrelevant or unwanted documents from the collection."
Cleanliness. Even a relevant source can be noisy. When parsing problems keep occurring, the guidance says they often point to upstream issues with the quality of the source data, and that can be a reason to choose a better source rather than keep patching the parser.
It treats coverage as volume. Indexing everything brings in duplicates, outdated versions and irrelevant material, which is the superfluous context that lowers retrieval quality. The better plan is to start from the use case, include the sources that answer its questions, then deduplicate and filter the rest.
Checkpoint 4 of 4· Match them up
Match each data-pipeline step to what it does to the document collection
Tap a term, then the definition that fits it.
These three steps decide which documents end up in the knowledge base: selecting by use case, removing repeats, and removing irrelevant material.
“Deduplication: Analyze the documents to identify and eliminate duplicates or near-duplicate documents.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.If a RAG app answers a topic wrongly, the fix is always to tune retrieval settings or the prompt.Why is that wrong?
Retrieval can only return what was ingested. If the source documents for that topic are missing, no chain setting can recover them, so check corpus coverage first.
Covered in The corpus sets the limit
2.Adding more documents to the corpus always improves answer quality.Why is that wrong?
Irrelevant and duplicate documents add superfluous context, which lowers retrieval quality. Filtering and deduplication exist to remove them.
Covered in Judging source quality: relevant, current, unique and clean
3.The engineering team can choose the right source documents without input from the business.Why is that wrong?
Developers are often not experts in the content. Databricks recommends involving domain experts and stakeholders from the start to identify and prioritise content.
Covered in Domain experts decide what coverage means
4.RAG source documents must be unstructured text such as PDFs.Why is that wrong?
RAG can use structured data, such as transaction tables or customer records, as well as unstructured documents. The choice depends on where the answers are.
Covered in Start from the questions users will ask
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/quality-data-pipeline-ragOfficial docs
“Corpus composition and ingestion: Select the right data sources and content based on the specific use case.”
↩︎ The corpus sets the limit“The correct data depends entirely on your application's specific requirements and goals”
↩︎ Start from the questions users will ask“help prioritize the most critical information to include”
↩︎ Domain experts decide what coverage means“Doing so often points to upstream issues with the quality of the source data.”
↩︎ Judging source quality: relevant, current, unique and clean“Your RAG application can't retrieve the information required to answer a user query without the right data corpus”
↩︎ Key concept“Your RAG application can't retrieve the information required to answer a user query without the right data corpus”
↩︎ Exam trap 1“Filtering: Eliminate irrelevant or unwanted documents from the collection.”
↩︎ Exam trap 2“Engage domain experts and stakeholders from the beginning of any project to help identify and curate relevant content”
↩︎ Exam trap 3“They can provide insights into the types of queries that users are likely to submit”
↩︎ Checkpoint“Deduplication: Analyze the documents to identify and eliminate duplicates or near-duplicate documents.”
↩︎ Checkpoint - 2.
“It's difficult to generate high-quality RAG output if the context provided to the LLM is missing important information or contains superfluous information.”
↩︎ The corpus sets the limit“missing important information or contains superfluous information”
↩︎ Checkpoint - 3.
“Transaction data from a SQL database”
↩︎ Start from the questions users will ask“A RAG application can supply the LLM with information from an updated knowledge base.”
↩︎ Judging source quality: relevant, current, unique and clean“RAG architecture can work with either unstructured or structured supporting data.”
↩︎ Exam trap 4“The data you use with RAG depends on your use case.”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/evaluate-define-qualityOfficial docs
“the correct supporting documents that should be retrieved”
↩︎ Domain experts decide what coverage means
Also cited
“Missing key context leads to incorrect answers or hallucinations”
↩︎ Prediction