CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 2 · Lesson 12/56

    Retrieval Metrics: Precision, Recall, DCG@10 and NDCG

    Use tools and metrics to evaluate retrieval performance

    10 min read
    1.79% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Calculate precision and recall for a retrieval result and say which one needs ground truth
    • Choose recall@k or precision@k based on the use case
    • Explain the 4-point graded relevance scale and the score that counts as relevant
    • Explain why Databricks recommends DCG@10 over NDCG, and when to use MRR or MAP instead

    Key concept

    Graded relevance — Each query and retrieved-document pair gets a relevance score from 0 to 3, not just relevant or not relevant. Metrics such as DCG then reward results by how useful they are and how high they rank.

    1.Precision and recall: the two basic retrieval questions

    A RAG application can fail in two places: the retriever and the generator. Databricks says you need to measure both separately, because a good answer can come from bad retrieval, and a bad answer can come from good retrieval. If you only measure the final answer, you can't tell which component to fix. This lesson covers the retriever. Its two basic metrics answer two different questions.

    Precision asks: of the chunks I retrieved, what share is relevant to the query? Recall asks: of all the documents I know are relevant, what share did I retrieve a chunk from? Precision measures how clean the results are. Recall measures how complete they are.

    The two metrics have different data requirements. To compute recall, your ground truth must list every relevant item. That is why Agent Evaluation computes document_recall deterministically from labelled documents. Precision only looks at what was retrieved, so an LLM judge can score each chunk with no labels at all.

    Here is the cookbook's worked example. Three results are retrieved and two are relevant, so precision is 2/3. Four relevant documents exist and two were retrieved, so recall is 2/4 = 0.5.

    Checkpoint 1 of 5· Check yourself

    A retriever returns 3 chunks. 2 are relevant. The labelled ground truth lists 4 relevant documents, and the 2 relevant chunks come from 2 of them. What are precision and recall?

    Sources1

    2.Choosing recall@k or precision@k for the use case

    Both metrics are usually reported at a cutoff k, meaning they only look at the top k results. Which one you track depends on what a miss costs. In RAG, a missing chunk means the LLM never sees that context, so the answer can be incomplete or made up. Recall@k is the default for RAG agents. When a user or downstream system needs only the exact match near the top, precision@k matters more.

    Metric to track by use case, from the AI Search retrieval quality guide
    PriorityExample use casesMetric to track
    Recall matters mostRAG agents, pharma clinical trial matching, financial compliance search, manufacturing root cause analysisRecall@k (e.g. recall@10, recall@50)
    Precision matters mostEntity resolution/fuzzy matching, financial services deduplication, supply chain part matching, tech support knowledge basePrecision@k (e.g. precision@3, precision@10)
    BalancedM&A due diligence, patent prior art search, Customer 360 matchingBoth recall and precision

    Checkpoint 2 of 5· Check yourself

    A team builds a RAG agent that answers compliance questions. Their main worry is the agent hallucinating because a key passage never reached the prompt. Which metric should they track first?

    Checkpoint 3 of 5· Exam question

    According to Databricks AI Search retrieval quality guidance, which metric is recommended as the primary metric for evaluating overall retrieval quality?

    Sources2

    3.Graded relevance: why 0–3 beats yes/no

    Binary precision treats a document that answers the question the same as one that only mentions the topic. The Databricks AI Search evaluation instead has an LLM judge score every query and document pair on a 4-point scale. Several metrics then turn that scale back into a yes/no answer at a fixed threshold, and the threshold is easy to get wrong.

    The LLM-judge relevance scale used by AI Search retrieval evaluation
    ScoreLabelMeaning
    3Highly RelevantDirectly answers the query or gives exactly the information sought
    2RelevantRelated and useful, but may not fully answer the query
    1Partially RelevantMentions the topic but gives no useful information for the query
    0Not RelevantUnrelated, or written in a different language from the query

    For precision@k and MRR, a result counts as relevant only if it scores 2 or higher. The relevance distribution metric follows the same rule: "Relevant+ %" counts scores of 2 and 3, and "Not Relevant %" counts scores of 0 and 1. A score-1 document is on topic but still counts against you. Language counts too: a correct answer in French to an English query scores 0.

    Checkpoint 4 of 5· Check yourself

    In the AI Search relevance distribution, which bucket does a result scored 1 (Partially Relevant) fall into?

    Sources3

    4.DCG@10, NDCG, MRR and MAP

    Databricks recommends DCG@10 as the primary metric for overall retrieval quality. It combines two things: how relevant each result is and where it ranks. Each of the top 10 results adds a gain of 2^relevance − 1. A score of 3 adds 7, and a score of 1 adds 1. That gain is then divided by a logarithmic discount, so results lower in the list count for less. If every result scored 3, DCG@10 would reach its maximum of 31.80. If all 10 results score 2, DCG@10 is 13.63. At that level, a 1-point gain is about a 7% relative improvement.

    NDCG divides DCG by the ideal DCG, which is what DCG would be if the same results were sorted by relevance. Its range is 0 to 1. NDCG tells you whether the system ranks well, but not how much useful information the user actually got. Two more ranking metrics fill specific needs. MRR looks only at the first relevant result. MAP averages precision at the position of every relevant result. The docs warn that no single metric tells the whole story.

    Which ranking metric answers which question
    MetricWhat it measuresWhen to use
    DCG@10Total utility of the top 10, weighted by positionPrimary metric for overall retrieval quality
    NDCGOrdering relative to the ideal ordering (0 to 1)Checking that ranking is correct, independent of how many relevant docs exist
    MRRAverage of 1/rank of the first result scoring 2 or moreWhen the top result matters most, e.g. question answering
    MAPPrecision at each relevant result's position, averagedA single number for ranking quality across all relevant documents
    Recall@kFraction of known relevant documents in the top kWhen completeness matters, e.g. RAG

    Checkpoint 5 of 5· Match them up

    Match each metric to what it captures

    Tap a term, then the definition that fits it.

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.An NDCG of 1.0 means retrieval quality is as good as it can be.Why is that wrong?

      NDCG only says the results are in ideal order. A set with 2 good results among 3 irrelevant ones can score 1.0, so use DCG@10 to see how much useful content came back.

      Covered in DCG@10, NDCG, MRR and MAP

    2. 2.Recall can be computed by an LLM judge looking only at the retrieved chunks.Why is that wrong?

      Recall's denominator is the full set of known relevant items, so it needs ground truth. Precision is the metric a judge can score without labels.

      Covered in Precision and recall: the two basic retrieval questions

    3. 3.Any result the judge scores above 0 counts as relevant for precision@k.Why is that wrong?

      Precision@k counts a result as relevant only at a score of 2 or more. A score-1 (Partially Relevant) result does not count.

      Covered in Graded relevance: why 0–3 beats yes/no

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.”
      ↩︎ Precision and recall: the two basic retrieval questions
      “two out of the three retrieved results were relevant to the user's query, so the precision was 0.66 (2/3).”
      ↩︎ Precision and recall: the two basic retrieval questions
      “Computing recall requires your ground-truth to contain all relevant items.”
      ↩︎ Exam trap 2
      “Computing precision does not require knowing all relevant items.”
      ↩︎ Prediction
      “The retrieved docs included two out of a total of four relevant docs, so the recall was 0.5 (2/4).”
      ↩︎ Checkpoint
    2. 2.
      “Metric to track: Recall@k (for example, recall@10, recall@50).”
      ↩︎ Choosing recall@k or precision@k for the use case
      “Metric to track: Precision@k (for example, precision@3, precision@10).”
      ↩︎ Choosing recall@k or precision@k for the use case
      “RAG agents: Missing key context leads to incorrect answers or hallucinations.”
      ↩︎ Checkpoint
    3. 3.
      “An LLM judge evaluates every query and retrieved document pair on a 4-point relevance scale.”
      ↩︎ Graded relevance: why 0–3 beats yes/no
      “Databricks recommends using DCG@10 as the primary metric for evaluating overall retrieval quality.”
      ↩︎ DCG@10, NDCG, MRR and MAP
      “a score-3 result contributes 7, while a score-1 result contributes 1.”
      ↩︎ DCG@10, NDCG, MRR and MAP
      “NDCG normalizes DCG by dividing it by the ideal DCG (the DCG if results were sorted in descending order of relevance).”
      ↩︎ DCG@10, NDCG, MRR and MAP
      “This granularity flows through to the metrics, particularly DCG, which weights higher-quality results more heavily.”
      ↩︎ Key concept
      “Both result sets achieve a perfect NDCG of 1.00 because each has results in ideal descending order.”
      ↩︎ Exam trap 1
      “The fraction of top-k results that are relevant (relevance score >= 2).”
      ↩︎ Exam trap 3
      “Not Relevant %: Results scoring 0 or 1 (not useful).”
      ↩︎ Checkpoint
      “Both result sets achieve a perfect NDCG of 1.00 because each has results in ideal descending order.”
      ↩︎ Prediction
      “MRR is the average of 1/rank across queries, where rank is the position of the first relevant result (score >= 2).”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Evaluating Retrieval with AI Search Evaluation and Agent Evaluation Judges

    Spotted a mistake, or was something unclear? Tell us.