CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 3 · Lesson 19/56

    Evaluating and Optimizing Prompt Versions with MLflow

    Create a prompt that adjusts an LLM's response from a baseline to a desired output

    9 min read
    1.79% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Register baseline and adjusted prompts as versions in the MLflow Prompt Registry
    • Write a pass/fail judge for the specific output property a prompt change targets
    • Compare prompt versions on the same evaluation dataset and judges
    • Describe how mlflow.genai.optimize_prompts and GEPA improve a prompt automatically

    1.Registering baseline and adjusted prompts as versions

    To show that a prompt change moved the output toward the target, you need both prompts saved somewhere you can reload them. In Databricks, that place is the MLflow Prompt Registry, backed by Unity Catalog. You create each prompt with mlflow.genai.register_prompt(), and "Prompts use double-brace syntax ({{variable}}) for template variables." If you register again under the same name, MLflow creates a new version instead of overwriting the old one. Your baseline is version 1 and your adjusted prompt is version 2.

    Checkpoint 1 of 3· Fill the gap

    Which call creates a new version of the baseline summarization prompt?

    # Version 1: Basic prompt
    prompt_v1 = mlflow.genai. ? (
        name=PROMPT_NAME,
        template="Summarize this text: {{content}}",
        commit_message="v1: Basic summarization prompt"
    )

    The commit_message records why you made the change, for example "v2: Added comprehensive fact coverage with 2-sentence requirement." Databricks' best-practice list says the same: "Document changes: Use meaningful commit messages to track why changes were made." Your application then loads a specific version by URI, so switching between baseline and adjusted prompts needs no code change:

    Loading a specific prompt version by URI or by name plus versionpython
    # Load a specific version using URI syntax
    prompt = mlflow.genai.load_prompt(name_or_uri=f"prompts:/{uc_schema}.{prompt_name}/2")
    
    # Or load from specific version
    prompt = mlflow.genai.load_prompt(name_or_uri=f"{uc_schema}.{prompt_name}", version="2")

    Sources12

    2.Measuring whether the output moved

    The evaluation guide follows a fixed sequence: "create prompt versions, build evaluation datasets with expected facts, and use MLflow's evaluation framework to compare performance." The scorers should measure the exact properties your prompt change targeted. For the two-sentence summarizer, that means two scorers. The built-in Correctness() checks expected facts. A custom judge built with make_judge checks the sentence count.

    A custom pass/fail judge for the one property v2 was designed to fixpython
    from mlflow.genai import make_judge
    
    # Create a custom judge using make_judge
    sentence_count_judge = make_judge(
        name="sentence_count_compliance",
        instructions="""Evaluate if this summary follows the 2-sentence requirement.
    
    Summary: {{ outputs }}
    
    Count the sentences carefully. Return true if the summary has exactly 2 sentences, and false otherwise.""",
        feedback_value_type=bool,
    )

    Databricks recommends feedback_value_type=bool for pass/fail judges because "MLflow aggregates boolean feedback into an average, so you get a pass rate to compare versions on." Each version is evaluated with mlflow.genai.evaluate() inside its own named run, and the result is metrics such as correctness/mean and sentence_count_compliance/mean for each version. In the sales-email tutorial, the improved v2 is run "using the same judges and dataset to see if you've successfully addressed the issues." You then select both runs in the Evaluation runs view and choose Compare. Databricks calls this view "the primary way to verify that a new prompt version outperforms the previous one."

    Checkpoint 2 of 3· Put it in order

    Put the steps for proving a prompt adjustment worked in order

    1. 1.Register the baseline and adjusted prompts as versions
    2. 2.Compare the evaluation runs
    3. 3.Build an evaluation dataset with expected facts
    4. 4.Evaluate each version with the same scorers

    Sources23

    3.Letting MLflow optimize the prompt for you

    Editing prompts by hand works, but MLflow can also run the adjustment loop itself. mlflow.genai.optimize_prompts() (Beta, requires MLflow >= 3.5.0) improves prompts using evaluation metrics and training data. It supports the GEPA algorithm through GepaPromptOptimizer, which "iteratively refines prompts using LLM-driven reflection and automated feedback." The optimizer still needs a definition of the desired output. You supply it as data: inputs, expected outputs and expected facts. In the classification tutorial, every row says the label must be one of the five allowed values.

    Optimizing the bare classification prompt with GEPA against labelled examplespython
    # Optimize the prompt
    result = mlflow.genai.optimize_prompts(
        predict_fn=predict_fn,
        train_data=dataset,
        prompt_uris=[prompt.uri],
        optimizer=GepaPromptOptimizer(reflection_model="databricks:/databricks-claude-sonnet-4-5"),
        scorers=[Correctness(model="databricks:/databricks-gpt-5")],
    )
    
    # Use the optimized prompt
    optimized_prompt = result.optimized_prompts[0]
    print(f"Optimized template: {optimized_prompt.template}")

    The result looks like a careful hand-written adjustment. Starting from Answer this question: {{question}}, the optimizer produced a prompt with a precision instruction, numbered rules and format examples:

    Opening of the GEPA-optimized prompt (the baseline was the first line alone)text
    Answer this question: {{question}}.
    Focus on providing precise,
    factual information without additional commentary or explanations.
    
    1. **Identify the Subject**: Clearly determine the specific subject
    of the question (e.g., geography, history)
    and provide a concise answer.
    
    2. **Clarity and Precision**: Your response should be a single,
    clear statement that directly addresses the question.
    Do not add extra details, context, or alternatives.

    The optimized prompt is saved as a new version in the Prompt Registry, so you can evaluate and compare it the same way as a hand-written v2.

    Checkpoint 3 of 3· Check yourself

    What must you supply for mlflow.genai.optimize_prompts() to steer a prompt toward the desired output?

    Sources4

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.You can judge a new prompt version on a new or harder dataset and compare that score with the old version's earlier score.Why is that wrong?

      If the data changes, the comparison no longer isolates the prompt change. Evaluate every version on the same dataset with the same judges.

      Covered in Measuring whether the output moved

    2. 2.Automatic prompt optimization removes the need to define the desired output.Why is that wrong?

      optimize_prompts learns the target from training data and scorers that you supply. Without expected outputs or facts, it has nothing to optimize toward.

      Covered in Letting MLflow optimize the prompt for you

    Practise it for real

    Show with metrics that an adjusted summarization prompt hits a two-sentence target more often than the baseline

    1. 1.Register Summarize this text: {{content}} with mlflow.genai.register_prompt under a Unity Catalog name, with commit message "v1: Basic summarization prompt".

      Why: The baseline has to be a stored version so you can evaluate it later.

      You should see: Version 1 of the prompt is created.

    2. 2.Register the expert-summarizer template, with the exactly-2-sentences rule and guidelines, under the same name.

      Why: Registering under the same name creates a new version instead of a new prompt.

      You should see: Version 2 of the prompt is created.

    3. 3.Create sentence_count_judge with make_judge and feedback_value_type=bool.

      Why: A boolean judge aggregates into a pass rate for the exact property v2 targets.

      You should see: A judge you can pass in the scorers list next to Correctness().

    4. 4.For versions 1 and 2, run mlflow.genai.evaluate inside a named run, using the same eval_dataset and scorers.

      Why: With data and judges held constant, any difference in scores comes from the prompt.

      You should see: Two runs, each reporting correctness/mean and sentence_count_compliance/mean.

    5. 5.Select both runs under Evaluation runs and choose Compare from the Actions menu.

      Why: Comparing runs is how you confirm the new version outperforms the old one.

      You should see: Results shown per trace for both runs, with higher sentence compliance expected for v2.

    Stuck? Get a nudge

    If v2 doesn't improve compliance, tighten the length rule and register a v3. Don't change the dataset.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 2.
      “Document changes: Use meaningful commit messages to track why changes were made.”
      ↩︎ Registering baseline and adjusted prompts as versions
      “MLflow aggregates boolean feedback into an average, so you get a pass rate to compare versions on”
      ↩︎ Measuring whether the output moved
      “Use consistent datasets: Evaluate all versions against the same data for fair comparison.”
      ↩︎ Exam trap 1
      “Use consistent datasets: Evaluate all versions against the same data for fair comparison.”
      ↩︎ Prediction
      “create prompt versions, build evaluation datasets with expected facts, and use MLflow's evaluation framework to compare performance”
      ↩︎ Checkpoint
    2. 3.
      “This comparison view is the primary way to verify that a new prompt version outperforms the previous one.”
      ↩︎ Measuring whether the output moved
    3. 4.
      “GEPA iteratively refines prompts using LLM-driven reflection and automated feedback, leading to systematic and data-driven improvements.”
      ↩︎ Letting MLflow optimize the prompt for you
      “Version Control: Automatically registers optimized prompts in MLflow Prompt Registry.”
      ↩︎ Letting MLflow optimize the prompt for you
      “Data-Driven Optimization: Uses your training data and custom scorers to guide optimization.”
      ↩︎ Checkpoint

    Also cited

    Spotted a mistake, or was something unclear? Tell us.