What you will be able to do
- Calculate precision and recall for a retrieval result and say which one needs ground truth
- Choose recall@k or precision@k based on the use case
- Explain the 4-point graded relevance scale and the score that counts as relevant
- Explain why Databricks recommends DCG@10 over NDCG, and when to use MRR or MAP instead
Key concept
Graded relevance — Each query and retrieved-document pair gets a relevance score from 0 to 3, not just relevant or not relevant. Metrics such as DCG then reward results by how useful they are and how high they rank.
1.Precision and recall: the two basic retrieval questions
A RAG application can fail in two places: the retriever and the generator. Databricks says you need to measure both separately, because a good answer can come from bad retrieval, and a bad answer can come from good retrieval. If you only measure the final answer, you can't tell which component to fix. This lesson covers the retriever. Its two basic metrics answer two different questions.
Precision asks: of the chunks I retrieved, what share is relevant to the query? Recall asks: of all the documents I know are relevant, what share did I retrieve a chunk from? Precision measures how clean the results are. Recall measures how complete they are.
The two metrics have different data requirements. To compute recall, your ground truth must list every relevant item. That is why Agent Evaluation computes document_recall deterministically from labelled documents. Precision only looks at what was retrieved, so an LLM judge can score each chunk with no labels at all.
Here is the cookbook's worked example. Three results are retrieved and two are relevant, so precision is 2/3. Four relevant documents exist and two were retrieved, so recall is 2/4 = 0.5.
Checkpoint 1 of 5· Check yourself
A retriever returns 3 chunks. 2 are relevant. The labelled ground truth lists 4 relevant documents, and the 2 relevant chunks come from 2 of them. What are precision and recall?
Precision divides by the number retrieved (2/3). Recall divides by the number of known relevant documents (2/4).
“The retrieved docs included two out of a total of four relevant docs, so the recall was 0.5 (2/4).”Source: docs.databricks.com
Sources1
2.Choosing recall@k or precision@k for the use case
Both metrics are usually reported at a cutoff k, meaning they only look at the top k results. Which one you track depends on what a miss costs. In RAG, a missing chunk means the LLM never sees that context, so the answer can be incomplete or made up. Recall@k is the default for RAG agents. When a user or downstream system needs only the exact match near the top, precision@k matters more.
| Priority | Example use cases | Metric to track |
|---|---|---|
| Recall matters most | RAG agents, pharma clinical trial matching, financial compliance search, manufacturing root cause analysis | Recall@k (e.g. recall@10, recall@50) |
| Precision matters most | Entity resolution/fuzzy matching, financial services deduplication, supply chain part matching, tech support knowledge base | Precision@k (e.g. precision@3, precision@10) |
| Balanced | M&A due diligence, patent prior art search, Customer 360 matching | Both recall and precision |
Checkpoint 2 of 5· Check yourself
A team builds a RAG agent that answers compliance questions. Their main worry is the agent hallucinating because a key passage never reached the prompt. Which metric should they track first?
Missing context is a completeness failure, and recall measures completeness. The guide lists RAG agents under 'recall matters most'.
“RAG agents: Missing key context leads to incorrect answers or hallucinations.”Source: docs.databricks.com
Checkpoint 3 of 5· Exam question
According to Databricks AI Search retrieval quality guidance, which metric is recommended as the primary metric for evaluating overall retrieval quality?
Correct answer: A — DCG@10
- A. Discounted Cumulative Gain at rank 10 is the metric Databricks recommends as the primary measure of overall retrieval quality, since it rewards relevant documents appearing near the top of a fixed-size result window using graded relevance.
- B. Normalized DCG rescales scores to a 0-1 range to compare across queries with different numbers of relevant documents, which is useful as a supplementary metric but is not the recommended primary metric.
- C. Recall at k measures the fraction of all relevant documents retrieved and is best suited to coverage-sensitive use cases, not the general primary quality metric Databricks recommends.
- D. Mean Reciprocal Rank captures how quickly the first relevant result appears and is a supplementary metric for scenarios where only one good result matters, not the recommended primary metric.
Sources2
3.Graded relevance: why 0–3 beats yes/no
Binary precision treats a document that answers the question the same as one that only mentions the topic. The Databricks AI Search evaluation instead has an LLM judge score every query and document pair on a 4-point scale. Several metrics then turn that scale back into a yes/no answer at a fixed threshold, and the threshold is easy to get wrong.
| Score | Label | Meaning |
|---|---|---|
| 3 | Highly Relevant | Directly answers the query or gives exactly the information sought |
| 2 | Relevant | Related and useful, but may not fully answer the query |
| 1 | Partially Relevant | Mentions the topic but gives no useful information for the query |
| 0 | Not Relevant | Unrelated, or written in a different language from the query |
For precision@k and MRR, a result counts as relevant only if it scores 2 or higher. The relevance distribution metric follows the same rule: "Relevant+ %" counts scores of 2 and 3, and "Not Relevant %" counts scores of 0 and 1. A score-1 document is on topic but still counts against you. Language counts too: a correct answer in French to an English query scores 0.
Checkpoint 4 of 5· Check yourself
In the AI Search relevance distribution, which bucket does a result scored 1 (Partially Relevant) fall into?
The distribution groups scores 0 and 1 together as not useful. Only scores of 2 or higher count as Relevant+.
“Not Relevant %: Results scoring 0 or 1 (not useful).”Source: docs.databricks.com
Sources3
4.DCG@10, NDCG, MRR and MAP
Databricks recommends DCG@10 as the primary metric for overall retrieval quality. It combines two things: how relevant each result is and where it ranks. Each of the top 10 results adds a gain of 2^relevance − 1. A score of 3 adds 7, and a score of 1 adds 1. That gain is then divided by a logarithmic discount, so results lower in the list count for less. If every result scored 3, DCG@10 would reach its maximum of 31.80. If all 10 results score 2, DCG@10 is 13.63. At that level, a 1-point gain is about a 7% relative improvement.
NDCG divides DCG by the ideal DCG, which is what DCG would be if the same results were sorted by relevance. Its range is 0 to 1. NDCG tells you whether the system ranks well, but not how much useful information the user actually got. Two more ranking metrics fill specific needs. MRR looks only at the first relevant result. MAP averages precision at the position of every relevant result. The docs warn that no single metric tells the whole story.
| Metric | What it measures | When to use |
|---|---|---|
| DCG@10 | Total utility of the top 10, weighted by position | Primary metric for overall retrieval quality |
| NDCG | Ordering relative to the ideal ordering (0 to 1) | Checking that ranking is correct, independent of how many relevant docs exist |
| MRR | Average of 1/rank of the first result scoring 2 or more | When the top result matters most, e.g. question answering |
| MAP | Precision at each relevant result's position, averaged | A single number for ranking quality across all relevant documents |
| Recall@k | Fraction of known relevant documents in the top k | When completeness matters, e.g. RAG |
Checkpoint 5 of 5· Match them up
Match each metric to what it captures
Tap a term, then the definition that fits it.
DCG measures absolute utility, NDCG measures ranking relative to the ideal, MRR looks only at the first hit, and MAP covers every relevant hit.
“MRR is the average of 1/rank across queries, where rank is the position of the first relevant result (score >= 2).”Source: docs.databricks.com
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.An NDCG of 1.0 means retrieval quality is as good as it can be.Why is that wrong?
NDCG only says the results are in ideal order. A set with 2 good results among 3 irrelevant ones can score 1.0, so use DCG@10 to see how much useful content came back.
Covered in DCG@10, NDCG, MRR and MAP
2.Recall can be computed by an LLM judge looking only at the retrieved chunks.Why is that wrong?
Recall's denominator is the full set of known relevant items, so it needs ground truth. Precision is the metric a judge can score without labels.
Covered in Precision and recall: the two basic retrieval questions
3.Any result the judge scores above 0 counts as relevant for precision@k.Why is that wrong?
Precision@k counts a result as relevant only at a score of 2 or more. A score-1 (Partially Relevant) result does not count.
Covered in Graded relevance: why 0–3 beats yes/no
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/evaluate-assess-performanceOfficial docs
“A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.”
↩︎ Precision and recall: the two basic retrieval questions“two out of the three retrieved results were relevant to the user's query, so the precision was 0.66 (2/3).”
↩︎ Precision and recall: the two basic retrieval questions“Computing recall requires your ground-truth to contain all relevant items.”
↩︎ Exam trap 2“Computing precision does not require knowing all relevant items.”
↩︎ Prediction“The retrieved docs included two out of a total of four relevant docs, so the recall was 0.5 (2/4).”
↩︎ Checkpoint - 2.
“Metric to track: Recall@k (for example, recall@10, recall@50).”
↩︎ Choosing recall@k or precision@k for the use case“Metric to track: Precision@k (for example, precision@3, precision@10).”
↩︎ Choosing recall@k or precision@k for the use case“RAG agents: Missing key context leads to incorrect answers or hallucinations.”
↩︎ Checkpoint - 3.
“An LLM judge evaluates every query and retrieved document pair on a 4-point relevance scale.”
↩︎ Graded relevance: why 0–3 beats yes/no“Databricks recommends using DCG@10 as the primary metric for evaluating overall retrieval quality.”
↩︎ DCG@10, NDCG, MRR and MAP“a score-3 result contributes 7, while a score-1 result contributes 1.”
↩︎ DCG@10, NDCG, MRR and MAP“NDCG normalizes DCG by dividing it by the ideal DCG (the DCG if results were sorted in descending order of relevance).”
↩︎ DCG@10, NDCG, MRR and MAP“This granularity flows through to the metrics, particularly DCG, which weights higher-quality results more heavily.”
↩︎ Key concept“Both result sets achieve a perfect NDCG of 1.00 because each has results in ideal descending order.”
↩︎ Exam trap 1“The fraction of top-k results that are relevant (relevance score >= 2).”
↩︎ Exam trap 3“Not Relevant %: Results scoring 0 or 1 (not useful).”
↩︎ Checkpoint“Both result sets achieve a perfect NDCG of 1.00 because each has results in ideal descending order.”
↩︎ Prediction“MRR is the average of 1/rank across queries, where rank is the position of the first relevant result (score >= 2).”
↩︎ Checkpoint