What you will be able to do
- Configure columns_to_rerank and know how num_results, column order and the 2000-character limit affect re-ranking
- Use semantic metadata columns to give the reranker more context
- Confirm the reranker ran, detect a fallback, and measure its impact with and without the reranker
- Describe what reranker finetuning does and its column rule
1.Turning on the reranker in a query
In AI Search, you turn on re-ranking by adding a reranker argument to similarity_search. You don't need to rebuild the index. Two parameters do two different jobs. columns sets which fields are returned to the caller. columns_to_rerank sets which fields the reranker reads when it scores each candidate. The two lists can differ.
results = index.similarity_search(
query_text="How to create an AI Search index",
num_results=10,
columns=["id", "text", "parent_doc_summary"],
reranker={
"model": "databricks_reranker",
"parameters": {
"columns_to_rerank": ["text", "parent_doc_summary"]
}
}
)Two details about how this behaves come up in exam questions. First, num_results controls only how many results are returned at the end. Databricks says it "does not affect the number of results used for reranking." Second, the order of columns_to_rerank matters. The reranker reads the columns in the order you list them and stops after the first 2000 characters. If a long text column comes first, it can use up the whole budget before the reranker reaches a summary column listed after it.
The SDK also has a typed helper, DatabricksReranker, which you can combine with any query type, including hybrid:
from databricks.ai_search.reranker import DatabricksReranker
results = index.similarity_search(
query_text = "How to create an AI Search index",
columns = ["id", "text", "parent_doc_summary", "date"],
num_results = 10,
query_type = "hybrid",
reranker=DatabricksReranker(columns_to_rerank=["text", "parent_doc_summary", "other_column"])
)Checkpoint 1 of 5· Fill the gap
Which key tells the reranker which columns to read when it re-scores each result?
results = index.similarity_search(
query_text="How to create an AI Search index",
num_results=10,
columns=["id", "text", "parent_doc_summary"],
reranker={
"model": "databricks_reranker",
"parameters": {
" ? ": ["text", "parent_doc_summary"]
}
}
)columns_to_rerank sets the content the reranker scores. columns only sets which fields are returned, and num_results only sets how many final results come back.
Checkpoint 2 of 5· Exam question
An engineering team is building a real-time product search bar that must return results in under 100 milliseconds and serves more than 5 queries per second at peak. Should they enable a cross-encoder-style reranker on top of ANN retrieval?
Correct answer: A — No, because the added scoring latency and query volume make the reranker a poor fit for this strict sub-100-millisecond, high-QPS search experience
- A. Reranking adds a scoring pass on top of the initial retrieval, which introduces measurable latency and doesn't scale well at high query-per-second volumes. For a strict sub-100-millisecond, high-QPS search bar, that added cost typically exceeds the available latency budget, making reranking unsuitable here.
- B. Reranking adds an extra scoring step rather than reducing latency; it can improve result quality but at the cost of additional processing time per query, not less. Treating it as a universal latency reducer misstates how the two-stage pipeline actually behaves.
- C. ANN retrieval is fully capable of returning ranked results on its own; a reranker is an optional refinement layer, not a requirement for the initial search to function. Search applications commonly run ANN-only pipelines when low latency matters more than the precision gain from reranking.
- D. A cross-encoder-style reranker scores query-passage pairs directly rather than reusing precomputed embedding vectors, which is exactly why it adds meaningful per-query latency compared to embedding-only similarity search. It does not run for free on top of existing vector computations.
Sources1
2.Giving the reranker more context with metadata
A chunk read on its own often lacks context. It may not say which manual it came from or which procedure it belongs to. Databricks recommends adding semantic context during data preparation, and one stated reason is that it "Gives rerankers more context for scoring." You can do this in two ways: prepend the summary and section to the chunk text, or store them as separate metadata columns.
Separate columns don't help unless the query uses them. The guidance says to use reranking with columns_to_rerank for semantic metadata such as summaries. For keyword-only metadata, it says to use hybrid (full-text) search. If a column isn't listed in columns_to_rerank, the reranker doesn't read it.
Checkpoint 3 of 5· Check yourself
An index has a doc_summary column describing each chunk's parent document. How do you make the built-in reranker use it when scoring?
For semantic metadata, Databricks says to list the column in columns_to_rerank. Returning a column with columns doesn't send it to the reranker. Hybrid full-text search is the advice for keyword-only metadata.
“For semantic metadata: Use reranking with columns_to_rerank parameter to consider these columns.”Source: docs.databricks.com
3.Confirming the reranker ran and measuring its effect
Set debug_level to at least 1 and the response includes timing for each stage. The example below shows the trade-off from the previous section in real numbers: ANN retrieval took 29 ms and reranking took 619 ms.
'debug_info': {'response_time': 693.0, 'ann_time': 29.0, 'reranker_time': 619.0}If the reranker call fails, the query still returns results. They are the first-stage results in their original order, and the debug info includes a RERANKER_TEMPORARILY_UNAVAILABLE warning saying "Results returned have not been processed by the reranker." Since the query doesn't raise an error, you have to check the debug info to know whether re-ranking actually happened.
To decide whether re-ranking pays off on your own data, use AI Search's built-in retrieval-quality evaluation (Beta). It generates queries from your documents, runs ANN, hybrid and full-text retrieval, and evaluates each strategy with and without the reranker on the same set of queries. An LLM judge then scores how relevant the results are. The results dashboard shows a bar chart comparing DCG@10 for each query type with and without the reranker, plus a table comparing base and reranker performance.
Checkpoint 4 of 5· Check yourself
A query that requests the reranker returns results normally, but debug_info contains a RERANKER_TEMPORARILY_UNAVAILABLE warning. What do the results represent?
When the reranker is unavailable, the query still returns results. They come straight from initial retrieval, and the warning says so.
“Results returned have not been processed by the reranker.”Source: docs.databricks.com
4.Finetuning a reranker on your own corpus
The built-in reranker is a general-purpose model. Reranker finetuning (Beta) "trains a custom reranker on your own index data," and Databricks handles training data generation, training, registration and deployment. It supports managed Delta Sync indexes only. Training data is built by sampling queries, from payload logs, a table you provide, or synthesized from your data. Databricks then retrieves candidates for each query and has an LLM judge their relevance. A base reranker model is trained on that data, registered in Unity Catalog, and deployed to Model Serving.
Querying with the finetuned reranker has a stricter column rule than the built-in one. columns_to_rerank must be exactly the index's embedding source column, given as a single-element list. In the REST API you also have to pass the Model Serving endpoint name with model_type set to MODEL_TYPE_FINETUNED, because REST doesn't work out the endpoint for you.
Checkpoint 5 of 5· Put it in order
Put the reranker finetuning pipeline in order
- 1.Train the reranker, starting from a base reranker model
- 2.Register the model in Unity Catalog and deploy it to a Model Serving endpoint
- 3.Query the index with the finetuned reranker through the Python SDK or REST API
- 4.Generate training data: sample queries, retrieve candidates, and have an LLM judge relevance
Databricks documents four stages in this order: generate training data, train, register and deploy, then query.
“It then retrieves candidates and uses an LLM to judge relevance and build training data.”Source: docs.databricks.com
Sources4
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Raising num_results makes the reranker score more candidates.Why is that wrong?
num_results only sets how many final results are returned. It doesn't change how many results the reranker scores.
Covered in Turning on the reranker in a query
2.The reranker reads every column in columns_to_rerank in full, so order doesn't matter.Why is that wrong?
The reranker reads the columns in the order listed and stops after the first 2000 characters, so a long column listed first can push out the ones after it.
Covered in Turning on the reranker in a query
3.A finetuned reranker accepts the same multi-column columns_to_rerank list as the built-in reranker, including metadata columns.Why is that wrong?
For a finetuned reranker, columns_to_rerank must be exactly the embedding source column, given as a single-element list.
Covered in Finetuning a reranker on your own corpus
Practise it for real
Measure what the built-in reranker changes, in both order and latency, on one of your AI Search indexes
1.Run index.similarity_search with a query_text, num_results=10, the columns you want returned, and debug_level=1, without a reranker argument.
Why: This gives you the first-stage order and timing as a baseline.
You should see: Ten results, and debug_info showing response_time and ann_time but no reranker_time.
2.Run the same query again with reranker={"model": "databricks_reranker", "parameters": {"columns_to_rerank": ["text"]}}.
Why: Changing only the reranker means any difference in order comes from re-ranking.
You should see: The same number of results, possibly in a different order, and debug_info now including reranker_time.
3.Add a summary or section metadata column after text in columns_to_rerank and run the query again.
Why: Metadata columns give the reranker more context, but only if they are listed and fit within its 2000-character budget.
You should see: The order may change again. If the text column is long, the extra column may have little effect.
4.Inspect debug_info for a warnings entry on each reranked run.
Why: If the reranker is unavailable, the query returns un-reranked results without raising an error.
You should see: No warnings if re-ranking succeeded. A RERANKER_TEMPORARILY_UNAVAILABLE warning if it fell back.
Stuck? Get a nudge
Compare reranker_time with ann_time. On a RAG agent, compare it with the LLM's generation time too, before you decide whether the added latency matters.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“num_results is the final number of results to return. This does not affect the number of results used for reranking.”
↩︎ Turning on the reranker in a query“The reranking calculation takes the columns in the order they are listed, and considers only the first 2000 characters it finds.”
↩︎ Turning on the reranker in a query“You can also specify columns containing metadata that you want the reranker to use for additional context as it assesses each document's relevance.”
↩︎ Giving the reranker more context with metadata“To ensure that you get latency information, set debug_level to at least 1.”
↩︎ Confirming the reranker ran and measuring its effect“num_results is the final number of results to return. This does not affect the number of results used for reranking.”
↩︎ Exam trap 1“The reranking calculation takes the columns in the order they are listed, and considers only the first 2000 characters it finds.”
↩︎ Exam trap 2“Results returned have not been processed by the reranker.”
↩︎ Checkpoint - 2.
“Gives rerankers more context for scoring.”
↩︎ Giving the reranker more context with metadata“For semantic metadata: Use reranking with columns_to_rerank parameter to consider these columns.”
↩︎ Checkpoint - 3.
“Each strategy is also evaluated with and without the reranker.”
↩︎ Confirming the reranker ran and measuring its effect“the dashboard shows a bar chart that compares DCG@10 scores for each query type, with and without using the reranker.”
↩︎ Confirming the reranker ran and measuring its effect - 4.
“Reranker finetuning trains a custom reranker on your own index data, then serves it so a model tuned to your corpus reranks your queries.”
↩︎ Finetuning a reranker on your own corpus“Databricks trains a reranker on that data, starting from a base reranker model.”
↩︎ Finetuning a reranker on your own corpus“This feature supports managed indexes only.”
↩︎ Finetuning a reranker on your own corpus“columns_to_rerank: Must be exactly the index's embedding source column, as a single-element list.”
↩︎ Exam trap 3“It then retrieves candidates and uses an LLM to judge relevance and build training data.”
↩︎ Checkpoint