CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 3 · Lesson 19/56

    Prompt Engineering: Moving an LLM from Baseline to Desired Output

    Create a prompt that adjusts an LLM's response from a baseline to a desired output

    11 min read
    1.79% of exam
    6 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Tell an accurate baseline response apart from a response that matches the task
    • Rewrite a bare prompt using a role, explicit output constraints, numbered guidelines and a priority instruction
    • Choose between few-shot examples, template parameters and chain-of-thought for a given output gap
    • Explain why a prompt tuned for one model may need reworking on another model

    Key concept

    Baseline vs. desired output — A baseline is what the model returns for a bare prompt. It can be correct and still miss the format, length, tone or label set the application needs. Prompt engineering adds instructions, constraints and examples to close that specific gap.

    1.What a baseline response is, and why it falls short

    Every prompt-engineering task starts with a baseline: the response the model gives to the simplest version of the prompt. The Databricks examples start from one-line templates such as Summarize this text: {{content}}, Answer this question: {{question}} and classify this: {{query}}. These prompts give the model the task and nothing more.

    This is the main idea of the subdomain. A baseline response is often not wrong. It just isn't what you asked for. In the classification tutorial, the application needs exactly one label from a fixed set of five: CONCLUSIONS, RESULTS, METHODS, OBJECTIVE or BACKGROUND. The tutorial says the model "should only output that word without any further explanation." A correct answer wrapped in an explanation still fails that requirement. The Markdown example has the same kind of problem: asked What is the capital of France? with a bare prompt, the model gets the facts right, but "the model doesn't output in a Markdown format."

    Before editing anything, write the gap down precisely. "Make it better" gives you nothing to work from. "Return exactly one of five labels, with no explanation" does. Each kind of gap points to a particular fix: format, length, label set, tone, ignored instructions or missing facts.

    Checkpoint 1 of 3· Check yourself

    A bare classification prompt returns the correct label, followed by a paragraph explaining the choice. The application needs only the label. How should you describe the gap?

    Sources12

    2.Anatomy of an adjusted prompt

    The Databricks prompt-evaluation guide registers two versions of a summarization prompt, which shows what the adjustments look like in practice. Version 1 is the baseline, Summarize this text: {{content}}. Version 2 adds a role, a hard length constraint (stated twice), a bulleted list of guidelines, and a trailing Summary: cue that marks where the answer should start.

    Version 2 of the summarization prompt: role, an explicit sentence count, guidelines and an output cuepython
    prompt_v2 = mlflow.genai.register_prompt(
        name=PROMPT_NAME,
        template="""You are an expert summarizer. Create a summary of the following content in *exactly* 2 sentences (no more, no less - be very careful about the number of sentences).
    
    Guidelines:
    - Include ALL core facts and key findings
    - Use clear, concise language
    - Maintain factual accuracy
    - Cover all main points mentioned
    - Write for a general audience
    - Use exactly 2 sentences
    
    Content: {{content}}
    
    Summary:""",
        commit_message="v2: Added comprehensive fact coverage with 2-sentence requirement"
    )
    How v1 and v2 of the summarization prompt differ, element by element
    Elementv1 (baseline)v2 (adjusted)
    RoleNone"You are an expert summarizer"
    LengthUnspecified"*exactly* 2 sentences", repeated in the guidelines
    Content coverageUnspecified"Include ALL core facts and key findings"
    Audience and styleUnspecified"Write for a general audience", "clear, concise language"
    Output cueNoneTemplate ends with "Summary:"

    Each addition targets one property of the output. The sentence count fixes length. The coverage bullets protect factual completeness. The audience bullet sets the register. The Create-and-edit-prompts guide does the same thing with its own template: it asks for a "neutral, objective tone" and asks the model to "Maintain the same level of formality as the original text." Because each rule is specific, an evaluator can later check whether the output follows it.

    Checkpoint 2 of 3· Check yourself

    A summarizer's outputs are factually fine, but their length varies from one to six sentences. The product needs exactly two. Which change to the prompt targets this gap most directly?

    Sources34

    3.Making the model follow the instruction that matters

    Sometimes the baseline problem is not format but priorities: the model has the right inputs and weighs them wrongly. In the Databricks evaluate-and-improve tutorial, a sales-email generator received a rep's instructions plus retrieved customer data. Evaluation showed the emails didn't follow the instructions closely enough. The v2 prompt reorders and labels its inputs so the model knows what comes first.

    The v2 sales-email prompt puts user instructions above the retrieved contextpython
        prompt = f"""You are a sales representative writing an email.
    
    MOST IMPORTANT: Follow these specific user instructions exactly:
    {user_instructions}
    
    Customer context (only use what's relevant to the instructions):
    {context}
    
    Guidelines:
    1. PRIORITIZE the user instructions above all else
    2. Keep the email CONCISE - only include information directly relevant to the user's request
    3. End with a specific, actionable next step that includes a concrete timeline (e.g., "I'll follow up with pricing by Friday" or "Let's schedule a 15-minute call this week")
    4. Only reference customer information if it's directly relevant to the user's instructions
    
    Write a brief, focused email that satisfies the user's exact request."""

    The prompt uses several techniques together:

    - Ordering. The user instructions come before the context. - Labelling. The instructions are marked "MOST IMPORTANT". - Scoping. The context is fenced with "only use what's relevant". - Concrete closing. A concrete example shows what a good closing step looks like.

    The call also sends a system message, "You are a helpful sales assistant who writes concise, instruction-focused emails." This repeats the same priorities at the system level.

    The tutorial's guidance on improving a prompt matches what this example does: "Refine system prompts to address specific failure patterns, add explicit guidelines for edge cases, include examples demonstrating correct handling." Note the word *specific*. Each edit should answer a failure that evaluation actually observed. A general rewrite doesn't do that.

    Sources5

    4.Few-shot examples, parameters and chain-of-thought

    Instructions describe the output you want. Sometimes it works better to show it, or to change how the model reaches its answer. The Databricks RAG-quality guide calls iterating on the prompt template "(AKA prompt engineering)" and lists several techniques beyond plain instructions.

    Prompt-template techniques from the Databricks RAG-quality guide
    TechniqueWhat you add to the promptOutput gap it targets
    Few-shot examplesWell-formed queries paired with ideal responsesFormat, style and content the model can't infer from instructions
    Parameterized templateVariables such as current date, user context or other metadataGeneric answers that should be personalized or context-aware
    Chain-of-ThoughtA request to work through the problem step by stepShallow answers to multi-step or open-ended queries

    The guide also explains how to choose few-shot examples: find the query types your chain struggles with, then "Create gold-standard responses for those queries and include them as examples in the prompt." The examples should be representative of the queries you expect at inference time and should cover a varied range of them. Chain-of-thought "breaks down complicated questions into simpler, sequential steps, guiding the LLM through a logical reasoning process."

    Checkpoint 3 of 3· Match them up

    Match each observed problem to the prompt technique that addresses it

    Tap a term, then the definition that fits it.

    One caveat applies to every technique here. The guide warns that a prompt tuned for one model may be less effective on another. It says to experiment with different prompt formats and lengths, and to be prepared to adapt your prompts when you switch models. If you change models, your desired-output prompt has to be checked again.

    Sources6

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.If the baseline answer is factually correct, the prompt doesn't need adjusting.Why is that wrong?

      A baseline can be accurate and still fail the application's format or label requirements. The gap you are closing is alignment with the use case, not only correctness.

      Covered in What a baseline response is, and why it falls short

    2. 2.A prompt that produces the desired output on one model will produce it on any model you swap in.Why is that wrong?

      Prompts often don't carry over between models. After switching, expect to rework formats, lengths and wording.

      Covered in Few-shot examples, parameters and chain-of-thought

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “While accurate, it's not aligned to any task or use case you're looking for.”
      ↩︎ What a baseline response is, and why it falls short
      “While accurate, it's not aligned to any task or use case you're looking for.”
      ↩︎ Key concept
      “It should only output that word without any further explanation.”
      ↩︎ Exam trap 1
      “It should only output that word without any further explanation.”
      ↩︎ Checkpoint
    2. 3.
      “Register different prompt versions representing different approaches to your task”
      ↩︎ Anatomy of an adjusted prompt
      “Create a summary of the following content in *exactly* 2 sentences (no more, no less - be very careful about the number of sentences).”
      ↩︎ Checkpoint
    3. 5.
      “Refine system prompts to address specific failure patterns, add explicit guidelines for edge cases, include examples demonstrating correct handling”
      ↩︎ Making the model follow the instruction that matters
    4. 6.
      “Include examples of well-formed queries and their corresponding ideal responses within the prompt template itself (few-shot learning).”
      ↩︎ Few-shot examples, parameters and chain-of-thought
      “This prompt engineering strategy breaks down complicated questions into simpler, sequential steps, guiding the LLM through a logical reasoning process.”
      ↩︎ Few-shot examples, parameters and chain-of-thought
      “Recognize that prompts often do not transfer seamlessly across different language models.”
      ↩︎ Exam trap 2
      “This helps the model understand the desired format, style, and content of the responses.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Evaluating and Optimizing Prompt Versions with MLflow

    Spotted a mistake, or was something unclear? Tell us.