What you will be able to do
- Tell an accurate baseline response apart from a response that matches the task
- Rewrite a bare prompt using a role, explicit output constraints, numbered guidelines and a priority instruction
- Choose between few-shot examples, template parameters and chain-of-thought for a given output gap
- Explain why a prompt tuned for one model may need reworking on another model
Key concept
Baseline vs. desired output — A baseline is what the model returns for a bare prompt. It can be correct and still miss the format, length, tone or label set the application needs. Prompt engineering adds instructions, constraints and examples to close that specific gap.
1.What a baseline response is, and why it falls short
Every prompt-engineering task starts with a baseline: the response the model gives to the simplest version of the prompt. The Databricks examples start from one-line templates such as Summarize this text: {{content}}, Answer this question: {{question}} and classify this: {{query}}. These prompts give the model the task and nothing more.
This is the main idea of the subdomain. A baseline response is often not wrong. It just isn't what you asked for. In the classification tutorial, the application needs exactly one label from a fixed set of five: CONCLUSIONS, RESULTS, METHODS, OBJECTIVE or BACKGROUND. The tutorial says the model "should only output that word without any further explanation." A correct answer wrapped in an explanation still fails that requirement. The Markdown example has the same kind of problem: asked What is the capital of France? with a bare prompt, the model gets the facts right, but "the model doesn't output in a Markdown format."
Before editing anything, write the gap down precisely. "Make it better" gives you nothing to work from. "Return exactly one of five labels, with no explanation" does. Each kind of gap points to a particular fix: format, length, label set, tone, ignored instructions or missing facts.
Checkpoint 1 of 3· Check yourself
A bare classification prompt returns the correct label, followed by a paragraph explaining the choice. The application needs only the label. How should you describe the gap?
The tutorial wants one word and no further explanation, so a correct label wrapped in explanation still misses the requirement. The gap is format, not correctness.
“It should only output that word without any further explanation.”Source: docs.databricks.com
2.Anatomy of an adjusted prompt
The Databricks prompt-evaluation guide registers two versions of a summarization prompt, which shows what the adjustments look like in practice. Version 1 is the baseline, Summarize this text: {{content}}. Version 2 adds a role, a hard length constraint (stated twice), a bulleted list of guidelines, and a trailing Summary: cue that marks where the answer should start.
prompt_v2 = mlflow.genai.register_prompt(
name=PROMPT_NAME,
template="""You are an expert summarizer. Create a summary of the following content in *exactly* 2 sentences (no more, no less - be very careful about the number of sentences).
Guidelines:
- Include ALL core facts and key findings
- Use clear, concise language
- Maintain factual accuracy
- Cover all main points mentioned
- Write for a general audience
- Use exactly 2 sentences
Content: {{content}}
Summary:""",
commit_message="v2: Added comprehensive fact coverage with 2-sentence requirement"
)| Element | v1 (baseline) | v2 (adjusted) |
|---|---|---|
| Role | None | "You are an expert summarizer" |
| Length | Unspecified | "*exactly* 2 sentences", repeated in the guidelines |
| Content coverage | Unspecified | "Include ALL core facts and key findings" |
| Audience and style | Unspecified | "Write for a general audience", "clear, concise language" |
| Output cue | None | Template ends with "Summary:" |
Each addition targets one property of the output. The sentence count fixes length. The coverage bullets protect factual completeness. The audience bullet sets the register. The Create-and-edit-prompts guide does the same thing with its own template: it asks for a "neutral, objective tone" and asks the model to "Maintain the same level of formality as the original text." Because each rule is specific, an evaluator can later check whether the output follows it.
Checkpoint 2 of 3· Check yourself
A summarizer's outputs are factually fine, but their length varies from one to six sentences. The product needs exactly two. Which change to the prompt targets this gap most directly?
Length is the gap, so the fix is an explicit, repeated length constraint, as in v2. A generic role line and extra context don't constrain length.
“Create a summary of the following content in *exactly* 2 sentences (no more, no less - be very careful about the number of sentences).”Source: docs.databricks.com
3.Making the model follow the instruction that matters
Sometimes the baseline problem is not format but priorities: the model has the right inputs and weighs them wrongly. In the Databricks evaluate-and-improve tutorial, a sales-email generator received a rep's instructions plus retrieved customer data. Evaluation showed the emails didn't follow the instructions closely enough. The v2 prompt reorders and labels its inputs so the model knows what comes first.
prompt = f"""You are a sales representative writing an email.
MOST IMPORTANT: Follow these specific user instructions exactly:
{user_instructions}
Customer context (only use what's relevant to the instructions):
{context}
Guidelines:
1. PRIORITIZE the user instructions above all else
2. Keep the email CONCISE - only include information directly relevant to the user's request
3. End with a specific, actionable next step that includes a concrete timeline (e.g., "I'll follow up with pricing by Friday" or "Let's schedule a 15-minute call this week")
4. Only reference customer information if it's directly relevant to the user's instructions
Write a brief, focused email that satisfies the user's exact request."""The prompt uses several techniques together:
- Ordering. The user instructions come before the context. - Labelling. The instructions are marked "MOST IMPORTANT". - Scoping. The context is fenced with "only use what's relevant". - Concrete closing. A concrete example shows what a good closing step looks like.
The call also sends a system message, "You are a helpful sales assistant who writes concise, instruction-focused emails." This repeats the same priorities at the system level.
The tutorial's guidance on improving a prompt matches what this example does: "Refine system prompts to address specific failure patterns, add explicit guidelines for edge cases, include examples demonstrating correct handling." Note the word *specific*. Each edit should answer a failure that evaluation actually observed. A general rewrite doesn't do that.
An abstract instruction like "end with a next step" still leaves the model to guess how concrete that step should be. A worked example shows the level of detail you expect: a specific action plus a specific timeline. That shrinks the gap between what the model produces and what you want.
Sources5
4.Few-shot examples, parameters and chain-of-thought
Instructions describe the output you want. Sometimes it works better to show it, or to change how the model reaches its answer. The Databricks RAG-quality guide calls iterating on the prompt template "(AKA prompt engineering)" and lists several techniques beyond plain instructions.
| Technique | What you add to the prompt | Output gap it targets |
|---|---|---|
| Few-shot examples | Well-formed queries paired with ideal responses | Format, style and content the model can't infer from instructions |
| Parameterized template | Variables such as current date, user context or other metadata | Generic answers that should be personalized or context-aware |
| Chain-of-Thought | A request to work through the problem step by step | Shallow answers to multi-step or open-ended queries |
The guide also explains how to choose few-shot examples: find the query types your chain struggles with, then "Create gold-standard responses for those queries and include them as examples in the prompt." The examples should be representative of the queries you expect at inference time and should cover a varied range of them. Chain-of-thought "breaks down complicated questions into simpler, sequential steps, guiding the LLM through a logical reasoning process."
Checkpoint 3 of 3· Match them up
Match each observed problem to the prompt technique that addresses it
Tap a term, then the definition that fits it.
Parameters inject context, few-shot examples show format and style, and chain-of-thought guides multi-step reasoning.
“This helps the model understand the desired format, style, and content of the responses.”Source: docs.databricks.com
One caveat applies to every technique here. The guide warns that a prompt tuned for one model may be less effective on another. It says to experiment with different prompt formats and lengths, and to be prepared to adapt your prompts when you switch models. If you change models, your desired-output prompt has to be checked again.
Sources6
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.If the baseline answer is factually correct, the prompt doesn't need adjusting.Why is that wrong?
A baseline can be accurate and still fail the application's format or label requirements. The gap you are closing is alignment with the use case, not only correctness.
Covered in What a baseline response is, and why it falls short
2.A prompt that produces the desired output on one model will produce it on any model you swap in.Why is that wrong?
Prompts often don't carry over between models. After switching, expect to rework formats, lengths and wording.
Covered in Few-shot examples, parameters and chain-of-thought
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/mlflow3/genai/tutorials/examples/prompt-optimization-quickstartOfficial docs
“While accurate, it's not aligned to any task or use case you're looking for.”
↩︎ What a baseline response is, and why it falls short“While accurate, it's not aligned to any task or use case you're looking for.”
↩︎ Key concept“It should only output that word without any further explanation.”
↩︎ Exam trap 1“It should only output that word without any further explanation.”
↩︎ Checkpoint - 2.
“In the following example, the model doesn't output in a Markdown format.”
↩︎ What a baseline response is, and why it falls short - 3.https://docs.databricks.com/aws/en/mlflow3/genai/prompt-version-mgmt/prompt-registry/evaluate-promptsOfficial docs
“Register different prompt versions representing different approaches to your task”
↩︎ Anatomy of an adjusted prompt“Create a summary of the following content in *exactly* 2 sentences (no more, no less - be very careful about the number of sentences).”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/mlflow3/genai/prompt-version-mgmt/prompt-registry/create-and-edit-promptsOfficial docs
“- Maintain the same level of formality as the original text”
↩︎ Anatomy of an adjusted prompt - 5.
“Refine system prompts to address specific failure patterns, add explicit guidelines for edge cases, include examples demonstrating correct handling”
↩︎ Making the model follow the instruction that matters - 6.
“Include examples of well-formed queries and their corresponding ideal responses within the prompt template itself (few-shot learning).”
↩︎ Few-shot examples, parameters and chain-of-thought“This prompt engineering strategy breaks down complicated questions into simpler, sequential steps, guiding the LLM through a logical reasoning process.”
↩︎ Few-shot examples, parameters and chain-of-thought“Recognize that prompts often do not transfer seamlessly across different language models.”
↩︎ Exam trap 2“This helps the model understand the desired format, style, and content of the responses.”
↩︎ Checkpoint