CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 3 · Lesson 18/56

    Augmenting Prompts with User Intents, Key Fields and Terms

    Augment a prompt with additional context from a user's input based on key fields, terms, and intents

    15 min read
    1.79% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Place query understanding and prompt augmentation correctly within a RAG chain
    • Classify user intent into predefined categories and use it to drive later steps
    • Extract key fields from a user query as retrieval filters, and name the pipeline changes this requires
    • Rewrite and expand the terms in a query so they match the vocabulary of the documents
    • Assemble an augmented prompt from a template whose variables come from the user's input

    Key concept

    Query understanding — A step at the start of the chain. It reads the user's raw request and breaks it into structured signals: what the user wants (intent), concrete constraints (key fields such as products, dates or regions) and better search wording (terms). Retrieval and prompt augmentation then use those signals instead of the raw text.

    1.Where user-input context enters the chain

    When a user sends a request to a RAG application, a fixed series of steps runs. Databricks calls this series the RAG chain. In outline, the chain has to understand the user's question, retrieve supporting data, augment the prompt with that data, and generate a response. Augmentation means combining the supporting data with the user's request, usually inside a template that also carries formatting rules and instructions for the LLM.

    This lesson is about the first step and how it shapes the third. The raw text a user types is rarely the best retrieval query, and it is rarely the best way to fill a prompt template either. The Databricks RAG inference guide describes an optional user query preprocessing step. It can format the query in a template, use another model to rewrite it, or pull out keywords. Its output is a *retrieval query*, which the retrieval step uses. Prompt augmentation then puts together the final prompt. That prompt joins the user's query with the retrieved context in a template that tells the model how to use each part.

    The RAG chain steps that use the user's input
    StepWhat happens to the user's input
    (Optional) User query preprocessingFormatted in a template, rewritten by another model, or mined for keywords. The output is a retrieval query
    RetrievalThe retrieval query is embedded with the same embedding model used for the document chunks, and the most similar chunks are returned
    Prompt augmentationThe user's query is combined with the retrieved context in a template that tells the model how to use each part
    LLM generationThe model answers from the augmented prompt, grounded on the added context

    These choices affect quality. The cookbook warns that the way you combine retrieved chunks with the user's query can significantly affect the model's ability to produce good responses. So building the augmented prompt is a real design decision. The rest of this lesson covers the three kinds of signal you can extract from user input to improve it: intents, key fields and terms.

    Checkpoint 1 of 5· Check yourself

    In the Databricks RAG inference chain, what does the optional user query preprocessing step produce?

    Sources12

    2.Intents: deciding what the user is asking for

    You could pass a user query straight to retrieval, and for some queries that works. Databricks' guidance, though, is that reformulating the query before retrieval is generally beneficial. The first thing to get from the query is its intent: which kind of request this is. In the cookbook's customer-support example, an LLM sorts each query into predefined categories such as "product information", "troubleshooting" or "account management".

    Intent is useful because it shapes every later step. A troubleshooting question needs different entities pulled out of it, and different supporting context, than an account-management question.

    This point is easy to get backwards. A single carefully written prompt is efficient. But the cookbook offers a general rule of thumb: when you try to pack several complex logic steps into one prompt, splitting them up often works better. For query understanding, that means one LLM call to classify intent, another to extract entities, and a third to rewrite the query. Each call depends on the one before. Entity extraction runs *based on the identified intent*, and rewriting *uses the extracted intent and entities*. The trade-off is clear: the extra calls add latency, and in return you get finer control and possibly better retrieved documents.

    Checkpoint 2 of 5· Put it in order

    Put the steps of the cookbook's multi-step query understanding component for a customer support bot in order.

    1. 1.Query rewriting: the extracted intent and entities are used to rewrite the query into a more specific, targeted form
    2. 2.Entity extraction: a second LLM call pulls out product names, reported errors or account numbers, based on the intent
    3. 3.Intent classification: an LLM sorts the query into predefined categories

    Sources3

    3.Key fields: turning parts of the query into retrieval filters

    Many queries contain hard constraints mixed into everyday language, for example "reports from 2023", "laptops" or a city name. The cookbook calls the technique for handling these filter extraction. You identify the constraints in the query and pass them to the retrieval step as extra parameters, so search covers only the relevant subset of data. The cookbook gives three kinds of example:

    - Time periods, such as articles from the last 6 months or reports from a given year - Products, services or categories, such as a named service or a product type - Geographic entities, such as city names or country codes

    In the support-bot example, the entities are product names, reported errors and account numbers. The AI Search retrieval quality guide describes an agent that does this automatically. It extracts relevant entities from the query, generates SQL-like filter strings, and runs a search that combines semantic matching with precise filtering. The guide calls filtering the biggest lever for retrieval quality and says it can reduce the search space by 90% or more. For customer support it suggests filtering by product line, region and issue category.

    The reflect question points to the catch. An extracted filter only helps if there is something to filter on. The cookbook says filter extraction has to be built together with two other parts. The data pipeline's metadata extraction step must make the relevant fields available on every document or chunk. The retrieval step must also be built to accept and apply the extracted filters. If you extract region = EMEA from the query but no chunk carries a region field, the filter cannot work.

    Checkpoint 3 of 5· Check yourself

    Your team adds an LLM step that extracts product names from user queries and passes them as filters, but retrieval quality does not improve. According to the Databricks cookbook, what else must change for filter extraction to work?

    Sources34

    4.Terms: rewriting and expanding the user's wording

    Intents and key fields are structured. Terms are the actual words of the query, and they often don't match the documents' vocabulary. Query rewriting turns a user query into one or more queries that better express the original intent. It helps most with complex or ambiguous queries whose wording differs from the wording in the documents. The cookbook gives three examples:

    - Paraphrasing conversation history in a multi-turn chat, so a follow-up such as "what about the second one?" becomes a standalone query - Correcting spelling mistakes in the user's query - Replacing words or phrases with synonyms to capture a broader range of relevant documents

    The AI Search retrieval quality guide puts the synonym idea into practice as query expansion: generate several variations of the query to improve recall. An LLM writes the variations, and the system searches with the original query plus each variation.

    Query expansion: an LLM generates synonym variations of the user's querypython
    # Use LLM to expand query with synonyms and related terms
    def expand_query(user_query):
        prompt = f"""Generate 3 variations of this search query including synonyms:
        Query: {user_query}
        Return only the variations, one per line."""

    Checkpoint 4 of 5· Fill the gap

    In the guide's expand_query function, which value completes the loop so that the original query and its variations are all searched?

    for query in  ?  + variations:
            results = index.similarity_search(query_text=query, num_results=10)
            all_results.extend(results)

    In the guide's example, "car maintenance" also searches "automobile repair", "vehicle servicing" and "auto maintenance". The stated benefit is better recall, because the search finds documents that use different wording. As with filters, rewriting is not a standalone fix. The cookbook says query rewriting must be done together with changes to the retrieval component. A rewritten query only helps if the retriever is set up to use it well. For example, it may need to search with several variations and merge the results, as expand_query does.

    Sources34

    5.Assembling the augmented prompt from extracted context

    After query understanding has produced an intent, key fields and better search terms, the final step is to put them into the prompt the LLM will see. The usual approach is a template with named variables. The MLflow Prompt Registry example registers a customer-support template whose variables are the company name, the customer's topic and the customer's question. The application fills in each variable at request time.

    A registered prompt template whose variables are filled from the user's requestpython
    initial_template = """\
    You are a helpful customer support assistant for {{company_name}}.
    
    Please help the customer with their inquiry about: {{topic}}
    
    Customer Question: {{question}}
    
    Provide a friendly, professional response that addresses their concern.
    """

    The template keeps two things apart: what the customer is asking about ({{topic}}) and what they actually wrote ({{question}}). That is exactly the split query understanding gives you. A classified intent is one value you could put in a slot like {{topic}}, while the original question stays as written. Retrieved context, chosen using the extracted filters and rewritten terms, is added alongside them. The template's instructions then tell the model how to use each part, which matches the Databricks description of prompt augmentation. The registry example's second version adds explicit response guidelines (acknowledge the concern, give clear next steps, keep a professional tone) on top of the same variables. Because the template is versioned, changes like these can be tracked against the application version.

    Signals extracted from user input and where each one goes
    Signal in the user's inputTechniqueWhere its output goes
    Intent (what kind of request)Intent classification into predefined categoriesGuides entity extraction and query rewriting, and can fill a topic-style template variable
    Key fields (time periods, products, geography, account numbers)Filter extraction / entity extractionPassed to the retrieval step as additional filter parameters
    Terms (the user's wording)Query rewriting and query expansionA retrieval query (or several variations) that better matches the documents' terminology

    Checkpoint 5 of 5· Match them up

    Match each query-understanding technique to what it does

    Tap a term, then the definition that fits it.

    Sources35

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Adding an LLM step that extracts filters from the user's query is enough to make filtered retrieval work.Why is that wrong?

      Filters only work if each chunk carries the matching metadata fields and the retriever accepts and applies them. Both the data pipeline and the retriever have to change.

      Covered in Key fields: turning parts of the query into retrieval filters

    2. 2.Query rewriting is a self-contained prompt change that does not affect the rest of the chain.Why is that wrong?

      Databricks states that rewriting must be paired with changes to the retrieval component, for example searching with several variations and deduplicating the results.

      Covered in Terms: rewriting and expanding the user's wording

    3. 3.A single LLM call that classifies intent, extracts entities and rewrites the query together is always the better design, because fewer calls means better quality.Why is that wrong?

      A single call is efficient, but splitting query understanding into several calls can give finer control and better retrieved documents, at the cost of some latency.

      Covered in Intents: deciding what the user is asking for

    4. 4.The user's raw query is the best retrieval query because it states exactly what the user wants.Why is that wrong?

      Using the raw query works for some cases, but Databricks advises reformulating it before retrieval so it better expresses intent.

      Covered in Intents: deciding what the user is asking for

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Understand the user's question. Retrieve supporting data. Augment the prompt with supporting data. Generate a response from an LLM using the augmented prompt.”
      ↩︎ Where user-input context enters the chain
      “supporting data is combined with the user's request, often using a template with additional formatting and instructions to the LLM”
      ↩︎ Where user-input context enters the chain
    2. 2.
      “formed by augmenting the user's query with the retrieved context, in a template that instructs the model how to use each component”
      ↩︎ Where user-input context enters the chain
      “how to combine them with the user's query in step 3 can significantly impact the model's ability to generate quality responses”
      ↩︎ Where user-input context enters the chain
      “This can involve formatting the query within a template, using another model to rewrite the request, or extracting keywords to aid retrieval.”
      ↩︎ Checkpoint
    3. 3.
      “Use an LLM to classify the user's query into predefined categories”
      ↩︎ Intents: deciding what the user is asking for
      “this approach may add some latency to the overall process”
      ↩︎ Intents: deciding what the user is asking for
      “Extracting geographic entities from the query, such as city names or country codes.”
      ↩︎ Key fields: turning parts of the query into retrieval filters
      “use another LLM call to extract relevant entities from the query, such as product names, reported errors, or account numbers”
      ↩︎ Key fields: turning parts of the query into retrieval filters
      “This can be particularly useful when dealing with complex or ambiguous queries that might not directly match the terminology used in the retrieval documents.”
      ↩︎ Terms: rewriting and expanding the user's wording
      “Replacing words or phrases in the user query with synonyms to capture a broader range of relevant documents”
      ↩︎ Terms: rewriting and expanding the user's wording
      “Combining a user query with retrieved information and instructions to guide the LLM towards generating high-quality responses.”
      ↩︎ Assembling the augmented prompt from extracted context
      “Use the extracted intent and entities to rewrite the original query into a more specific and targeted format”
      ↩︎ Assembling the augmented prompt from extracted context
      “Analyzing and transforming user queries to better represent intent and extract relevant information, such as filters or keywords, to improve the retrieval process.”
      ↩︎ Key concept
      “the retrieval step should be implemented to accept and apply extracted filters”
      ↩︎ Exam trap 1
      “Query rewriting must be done in conjunction with changes to the retrieval component”
      ↩︎ Exam trap 2
      “it can allow for more fine-grained control and potentially improve the quality of the retrieved documents”
      ↩︎ Exam trap 3
      “it is generally beneficial to reformulate the query before the retrieval step”
      ↩︎ Exam trap 4
      “breaking down the query understanding process into multiple LLM calls can lead to better results”
      ↩︎ Prediction
      “you might use one LLM call to classify the query intent, another to extract relevant entities, and a third to rewrite the query”
      ↩︎ Checkpoint
      “Filter extraction must be done in conjunction with changes to both metadata extraction data pipeline and retriever chain components.”
      ↩︎ Checkpoint
      “Filter extraction involves identifying and extracting these filters from the query and passing them to the retrieval step as additional parameters.”
      ↩︎ Checkpoint
    4. 4.
      “Executes the search with both semantic understanding and precise filtering.”
      ↩︎ Key fields: turning parts of the query into retrieval filters
      “Filtering dramatically reduces search space and improves both precision and recall.”
      ↩︎ Key fields: turning parts of the query into retrieval filters
      “Impact: Improves recall by finding documents with different terminology.”
      ↩︎ Terms: rewriting and expanding the user's wording

    Spotted a mistake, or was something unclear? Tell us.