CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 2 · Lesson 11/56

    Choosing Source Documents for a RAG Application

    Identify needed source documents that provide necessary knowledge and quality for a given RAG application

    11 min read
    1.79% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain why the data corpus limits what a RAG application can answer, no matter how well retrieval and prompting are tuned
    • Work out which document types a RAG use case needs, starting from the questions users will ask
    • Explain why domain experts and an evaluation set are the way to check corpus coverage
    • Choose between candidate sources on relevance, freshness, ownership, structure, duplication and noise

    Key concept

    Corpus composition — Corpus composition means deciding which sources and content go into the knowledge base that a RAG application retrieves from. Retrieval can only return what the corpus contains, so this choice limits the quality of every answer the application gives.

    1.The corpus sets the limit

    A RAG application has two parts: an offline data pipeline that prepares and indexes documents, and an online chain that retrieves from that index and generates an answer. People usually tune the online part first, adjusting the prompt, the number of chunks, or the search type. All of those settings work on whatever the pipeline put into the index. If the documents that answer a question were never collected, the retriever cannot return them, and the LLM either says it does not know or, worse, fills the gap with a plausible guess.

    That is why Databricks puts corpus selection first in the data pipeline, before parsing, chunking or embedding. The first component listed is "Corpus composition and ingestion: Select the right data sources and content based on the specific use case." Note the phrase "based on the specific use case". There is no corpus that is correct for every application. The right one follows from what the application is meant to answer.

    The Databricks quality guidance treats this as a retrieval-quality problem with two failure modes. Context can be missing important information, or it can contain superfluous information. Leaving documents out causes the first. Adding documents carelessly causes the second. Choosing sources well means avoiding both.

    Checkpoint 1 of 4· Check yourself

    According to Databricks' RAG quality guidance, which two context problems make it hard to produce high-quality RAG output?

    Sources12

    2.Start from the questions users will ask

    Because "The correct data depends entirely on your application's specific requirements and goals", work backwards. List the kinds of questions the application must answer, then ask which documents contain those answers. The Databricks cookbook uses a customer support bot as its example and names four document types to consider. Each covers a different kind of user need. The table below matches each type to the kind of question it is likely to answer. The document types come from the source; the question column is an illustration.

    Customer support bot: candidate document types (from the Databricks example) and the kinds of questions each would cover
    Document typeTypical user question it covers
    Knowledge base documentsHow do I do X? What is the policy on Y?
    Frequently Asked Questions (FAQs)Short, common questions with short, standard answers
    Product manuals and specificationsWhat are the limits, settings or features of a specific product?
    Troubleshooting guidesI see this error or symptom. How do I fix it?

    Use the same reasoning on any scenario the exam gives you. An HR benefits assistant needs current policy handbooks and plan documents, not last year's versions or unrelated engineering wikis. A contract-review assistant needs the contracts themselves and the clause playbooks. The data also does not have to be unstructured. Databricks notes that RAG can use unstructured data such as PDFs, Office documents and wikis, or structured data such as customer records and transaction tables. Which you need depends on whether the answers live in prose or in rows. Exam questions about choosing sources usually reward the option that matches the source to the task, not the one that includes the most data.

    Checkpoint 2 of 4· Check yourself

    A team is building a RAG assistant that answers questions about order status and recent purchases for logged-in customers. Which source belongs at the centre of its corpus?

    Sources13

    3.Domain experts decide what coverage means

    Engineers building a RAG app are often not experts in its content, so they can't reliably tell which documents matter. Databricks' advice is direct: "Engage domain experts and stakeholders from the beginning of any project to help identify and curate relevant content that could improve the quality and coverage of your data corpus." Experts add two things. They know the queries users are likely to submit, and they can rank which information is most critical to include.

    Coverage is something you test, not something you assume. The evaluation set links the corpus decision to measurement. Databricks describes an evaluation set as representative queries with ground-truth answers and, optionally, "the correct supporting documents that should be retrieved". When writing those labels, a question whose supporting document is missing from every candidate source shows a gap in the corpus before users find it. The guidance also says the evaluation set should be continually updated to reflect the changing nature of the indexed data. So corpus selection is reviewed again as the application and its data change, not done once.

    Checkpoint 3 of 4· Check yourself

    Early in a RAG project, what does Databricks say domain experts and stakeholders contribute to the corpus?

    Sources14

    4.Judging source quality: relevant, current, unique and clean

    Once you know which kinds of documents you need, compare candidate sources on quality. Four properties come up most often.

    Ownership and freshness. RAG is most useful for knowledge the base model lacks: proprietary documents, and information that changes. Databricks notes that "A RAG application can supply the LLM with information from an updated knowledge base." Choose the authoritative, current version of a source over a stale copy. Prefer sources that can be ingested incrementally so the index stays current.

    Uniqueness. Copies of the same document, such as repeated policy versions or mirrored wiki pages, use up retrieval slots with repeated text. The pipeline guidance includes a deduplication step to find and remove duplicates and near-duplicates.

    Relevance. Off-topic documents add the superfluous context discussed in the first section. The pipeline also includes a filtering step to "Eliminate irrelevant or unwanted documents from the collection."

    Cleanliness. Even a relevant source can be noisy. When parsing problems keep occurring, the guidance says they often point to upstream issues with the quality of the source data, and that can be a reason to choose a better source rather than keep patching the parser.

    Checkpoint 4 of 4· Match them up

    Match each data-pipeline step to what it does to the document collection

    Tap a term, then the definition that fits it.

    Sources31

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.If a RAG app answers a topic wrongly, the fix is always to tune retrieval settings or the prompt.Why is that wrong?

      Retrieval can only return what was ingested. If the source documents for that topic are missing, no chain setting can recover them, so check corpus coverage first.

      Covered in The corpus sets the limit

    2. 2.Adding more documents to the corpus always improves answer quality.Why is that wrong?

      Irrelevant and duplicate documents add superfluous context, which lowers retrieval quality. Filtering and deduplication exist to remove them.

      Covered in Judging source quality: relevant, current, unique and clean

    3. 3.The engineering team can choose the right source documents without input from the business.Why is that wrong?

      Developers are often not experts in the content. Databricks recommends involving domain experts and stakeholders from the start to identify and prioritise content.

      Covered in Domain experts decide what coverage means

    4. 4.RAG source documents must be unstructured text such as PDFs.Why is that wrong?

      RAG can use structured data, such as transaction tables or customer records, as well as unstructured documents. The choice depends on where the answers are.

      Covered in Start from the questions users will ask

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Corpus composition and ingestion: Select the right data sources and content based on the specific use case.”
      ↩︎ The corpus sets the limit
      “The correct data depends entirely on your application's specific requirements and goals”
      ↩︎ Start from the questions users will ask
      “help prioritize the most critical information to include”
      ↩︎ Domain experts decide what coverage means
      “Doing so often points to upstream issues with the quality of the source data.”
      ↩︎ Judging source quality: relevant, current, unique and clean
      “Your RAG application can't retrieve the information required to answer a user query without the right data corpus”
      ↩︎ Key concept
      “Your RAG application can't retrieve the information required to answer a user query without the right data corpus”
      ↩︎ Exam trap 1
      “Filtering: Eliminate irrelevant or unwanted documents from the collection.”
      ↩︎ Exam trap 2
      “Engage domain experts and stakeholders from the beginning of any project to help identify and curate relevant content”
      ↩︎ Exam trap 3
      “They can provide insights into the types of queries that users are likely to submit”
      ↩︎ Checkpoint
      “Deduplication: Analyze the documents to identify and eliminate duplicates or near-duplicate documents.”
      ↩︎ Checkpoint
    2. 2.
      “It's difficult to generate high-quality RAG output if the context provided to the LLM is missing important information or contains superfluous information.”
      ↩︎ The corpus sets the limit
      “missing important information or contains superfluous information”
      ↩︎ Checkpoint
    3. 3.
      “Transaction data from a SQL database”
      ↩︎ Start from the questions users will ask
      “A RAG application can supply the LLM with information from an updated knowledge base.”
      ↩︎ Judging source quality: relevant, current, unique and clean
      “RAG architecture can work with either unstructured or structured supporting data.”
      ↩︎ Exam trap 4
      “The data you use with RAG depends on your use case.”
      ↩︎ Checkpoint

    Also cited

    Spotted a mistake, or was something unclear? Tell us.