What you will be able to do
- Explain where re-ranking sits in a retrieval pipeline and what it changes about the results
- Tell a cross-encoder reranker apart from rank fusion (RRF) in hybrid search
- Decide when the built-in reranker fits a workload and when a faster reranker is needed
Key concept
Re-ranking (second-pass retrieval) — First, a fast retrieval step (ANN, full-text or hybrid) collects a set of candidate chunks. Then a reranker scores each candidate again against the query and puts them in a new order, so the most relevant chunks end up at the top of the context passed to the LLM.
1.Where re-ranking sits in the retrieval pipeline
How well a RAG application answers depends on how well it retrieves. The Databricks quality guidance lists the retrieval approach as one of the things you can tune, next to chunking, embeddings and prompt formatting, and re-ranking is one of the options it names. The guidance describes that approach as "keyword vs. hybrid vs. semantic search, rewriting the user's query, transforming a user's query into filters, or re-ranking."
Re-ranking is a separate step from retrieval. The first stage (ANN vector search, full-text keyword search, or a hybrid of the two) collects a set of candidates. The reranker then works only on that set. It doesn't search the index again. It scores each candidate a second time and changes the order. Databricks describes the goal as evaluating the retrieved documents "to identify the ones that are semantically most relevant."
Order matters because the LLM sees a fixed number of chunks. If the best passage comes back in position eight and only the top five go into the prompt, the model never sees it. Re-ranking moves passages like that to the top. The AI Search guide says the payoff is precision: the reranker "Boosts precision by achieving high recall with fewer candidates." It also adds that the reranker "Works best when combined with techniques like hybrid search and filtering". Those techniques produce a better candidate set, and the reranker produces a better order within it.
Checkpoint 1 of 5· Check yourself
In a Databricks AI Search query that uses the reranker, what does the reranker work on?
Re-ranking runs after initial retrieval. It scores the retrieved candidates again in the context of the query and returns a more accurate order. Query rewriting and re-embedding are different techniques.
“After initial retrieval, a cross-encoder model re-evaluates each result in the context of the query, producing more accurate relevance ordering.”Source: docs.databricks.com
Checkpoint 2 of 5· Exam question
In a two-stage retrieval pipeline for a RAG application, what is the primary role of a re-ranking step?
Correct answer: A — Re-scoring the initial candidate set returned by vector search so the most relevant passages are ranked highest before being passed to the LLM
- A. Re-ranking is a post-retrieval refinement step: it takes the candidates already surfaced by the vector index and re-scores them, typically with a model that jointly evaluates query-passage relevance, so the strongest matches float to the top before generation. This improves precision on the small candidate set that ultimately reaches the LLM.
- B. Re-ranking does not swap out or retrain the embedding model behind the vector index; the embedding model continues to drive the initial approximate-nearest-neighbor search unchanged. Re-ranking operates as a separate downstream scoring pass on the results that search already returned.
- C. De-duplicating chunks is a data preparation and indexing concern handled before or during ingestion, not something a reranker performs at query time. A reranker only reorders the results of a specific query's retrieval, it does not modify the stored index contents.
- D. Re-ranking narrows the candidate set rather than expanding it; the pattern is to retrieve a broader set first and then reduce it down to the highest-quality subset. Expanding the candidate pool is the job of the initial retrieval step, not the reranker.
2.Two meanings of "re-ranking": rank fusion vs a reranker model
Exam questions often hinge on the fact that "re-ranking" is used for more than one thing. The RAG cookbook uses it in three ways.
1. Combining result lists in hybrid search. Hybrid search runs semantic and keyword search and then merges the two lists. The cookbook says the merge can use "reciprocal rank fusion or a re-ranking model." In AI Search, the hybrid query type merges with Reciprocal Rank Fusion (RRF), which combines rank positions from the two lists. RRF doesn't read the text again. 2. Applying extra ranking rules. You can re-order retrieved chunks by a rule that doesn't involve a model. The cookbook's example is "sort by time". 3. Running a reranker model. A cross-encoder reads the query and each candidate together and gives the pair a relevance score. The cookbook names mxbai-rerank and ColBERTv2 as examples of models that "can yield an uplift in retrieval performance."
The Databricks built-in reranker is the third kind. In the AI Search strategy table it appears as its own row, separate from hybrid.
| Strategy | How it works | Best for |
|---|---|---|
| Hybrid | Combines ANN and full-text results using Reciprocal Rank Fusion (RRF). | General-purpose retrieval. The recommended starting point for most use cases. |
| Hybrid + reranker | Runs hybrid search, then re-scores results with a cross-encoder reranker model. | Higher precision when latency allows (typically under 1 second additional per query). |
Checkpoint 3 of 5· Match them up
Match each technique to what it does
Tap a term, then the definition that fits it.
RRF is how hybrid search merges its two lists. A cross-encoder reranker is a model that runs afterwards as a second pass. A rule-based re-order such as sorting by time involves no model.
“After retrieving an initial set of chunks, apply additional ranking criteria (for example, sort by time) or a reranker model to re-order the results.”Source: docs.databricks.com
3.The quality–latency trade-off
Scoring every query–candidate pair with a cross-encoder costs time. Databricks states it plainly: "Reranking incurs a small latency delay but can significantly improve retrieval quality and agent performance." The size of the gain varies between Databricks pages. The query guide says "approximately 10%", and the retrieval-quality guide describes a "~15% quality improvement" from enabling it. What both agree on is the direction: a clear quality gain in exchange for added latency.
Whether that exchange is worth it depends on the workload. In a RAG agent, the LLM takes most of the response time, so a reranking step of under a second barely shows. That is why Databricks recommends "trying out reranking for any RAG agent use case." Interactive search behaves differently, and the table below gives the guidance for the built-in reranker. Its performance figure is that it "Reranks 50 results in under 1 second in typical workloads," and it can be as fast as about 250 ms for shorter chunks.
| Good fit | Not a good fit |
|---|---|
| RAG agents (latency is dominated by LLM generation). | High QPS applications (>5 QPS without additional scaling). |
| Quality-first applications. | Real-time search bars requiring <100 msec latency. |
| Low-to-moderate QPS (~5 QPS out of the box). | Applications where sub-second reranking time is unacceptable. |
Checkpoint 4 of 5· Check yourself
Why is the built-in reranker a good fit for a RAG chatbot even though it adds up to about a second per query?
Databricks lists RAG agents as a good fit because LLM generation takes most of the response time. Above 5 QPS without extra scaling is listed as a poor fit, not an optimal load.
“RAG agents (latency is dominated by LLM generation).”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
A team builds a RAG pipeline where the AI Search index performs ANN retrieval and a reranker narrows the results before they reach the LLM. To get the benefit of the reranker, how should the initial ANN retrieval step be configured relative to the final number of chunks sent to the LLM?
Correct answer: A — Retrieve a larger candidate set than what's ultimately sent to the LLM, then let the reranker score and cut it down (e.g., retrieve 50, keep the top 5)
- A. Re-ranking is only useful when it has a meaningfully larger pool to choose from than what's finally used, since its value comes from achieving high recall with a wide initial candidate set and then narrowing to a smaller high-precision set. Retrieving on the order of tens of candidates and reranking down to the few actually sent to the LLM is the standard two-stage pattern.
- B. If ANN retrieval already returns exactly the number of chunks needed, the reranker has no extra candidates to filter out, so it can only reorder a set that's already fixed in size. This defeats the precision benefit of reranking, since low-relevance chunks that should have been dropped stay in the final set.
- C. A reranker scores existing candidates for relevance; it does not generate or expand new passages from a single retrieved chunk. Retrieving only one chunk upfront also removes any pool for the reranker to meaningfully evaluate.
- D. Candidate set size for a two-stage retrieval pipeline is a retrieval-quality decision, not something tied to the embedding model's context window length. Sizing batches to the embedding context window doesn't relate to how many chunks the reranker needs to choose among.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Hybrid search already re-ranks with a model, so adding the reranker is redundant.Why is that wrong?
Hybrid search merges its ANN and full-text lists with Reciprocal Rank Fusion, which is based on rank positions. The reranker is a separate cross-encoder that runs after hybrid search and scores the results again.
Covered in Two meanings of "re-ranking": rank fusion vs a reranker model
2.Re-ranking is never usable for low-latency search bars or high-QPS applications.Why is that wrong?
Only the built-in reranker is a poor fit for those workloads. A lightweight cross-encoder served on Model Serving can rerank in under 100 ms and still improve quality.
Covered in The quality–latency trade-off
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Reranking is a technique that improves retrieval quality by evaluating the retrieved documents to identify the ones that are semantically most relevant.”
↩︎ Where re-ranking sits in the retrieval pipeline“Combines ANN and full-text results using Reciprocal Rank Fusion (RRF).”
↩︎ Two meanings of "re-ranking": rank fusion vs a reranker model“Reranking incurs a small latency delay but can significantly improve retrieval quality and agent performance.”
↩︎ The quality–latency trade-off“The reranker typically improves quality by approximately 10% but adds latency.”
↩︎ The quality–latency trade-off“Databricks recommends trying out reranking for any RAG agent use case.”
↩︎ The quality–latency trade-off“The reranker is an optional second pass applied on top of any strategy.”
↩︎ Key concept“Runs hybrid search, then re-scores results with a cross-encoder reranker model.”
↩︎ Exam trap 1“After initial retrieval, a cross-encoder model re-evaluates each result in the context of the query, producing more accurate relevance ordering.”
↩︎ Checkpoint - 2.
“The retrieval approach (for example, keyword vs. hybrid vs. semantic search, rewriting the user's query, transforming a user's query into filters, or re-ranking).”
↩︎ Where re-ranking sits in the retrieval pipeline - 3.
“Boosts precision by achieving high recall with fewer candidates.”
↩︎ Where re-ranking sits in the retrieval pipeline“Works best when combined with techniques like hybrid search and filtering.”
↩︎ Where re-ranking sits in the retrieval pipeline“One-line change for ~15% quality improvement.”
↩︎ The quality–latency trade-off“Reranks 50 results in under 1 second in typical workloads.”
↩︎ The quality–latency trade-off“Reranking can still provide significant quality improvements for search bars and high-QPS applications - you just need a faster reranker.”
↩︎ Exam trap 2“Consider deploying a lightweight reranking model (for example, cross-encoder/ms-marco-TinyBERT-L-2-v2) as a custom model on Databricks Model Serving for sub-100 msec reranking.”
↩︎ Prediction“RAG agents (latency is dominated by LLM generation).”
↩︎ Checkpoint - 4.
“Use a re-ranking approach to combine the results, such as reciprocal rank fusion or a re-ranking model.”
↩︎ Two meanings of "re-ranking": rank fusion vs a reranker model“Reranking with cross-encoder models such as mxbai-rerank and ColBERTv2 can yield an uplift in retrieval performance.”
↩︎ Two meanings of "re-ranking": rank fusion vs a reranker model“After retrieving an initial set of chunks, apply additional ranking criteria (for example, sort by time) or a reranker model to re-order the results.”
↩︎ Checkpoint