CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 3 · Lesson 21/56

    Choosing an LLM for Production: Serving Mode, Throughput and Metrics

    Select the best LLM based on the attributes of the application to be developed

    8 min read
    1.79% of exam
    6 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Pick pay-per-token, priority mode, provisioned throughput or AI Functions based on an application's production needs
    • Use rate limits and regional availability to confirm a candidate model can serve the expected load
    • Confirm a model choice by comparing candidates on quality, cost and latency metrics

    1.Match the serving mode to the workload

    Choosing an LLM involves more than the model's capabilities. It also depends on how the model will be served. Databricks Foundation Model APIs offer several modes, and each one suits a different kind of application. When you pick a model, you also need to pick a mode that model supports.

    Foundation Model APIs modes and the application attributes they fit
    ModeFits applications that…Notes from the docs
    Pay-per-tokenare just starting out or are proofs of conceptPreconfigured endpoints in the workspace; recommended for getting started
    Priority pay-per-token (priority mode)are latency-sensitive and real-timeOpt-in per request, no commitment, higher per-token rate
    Provisioned throughputneed high throughput, performance guarantees, fine-tuned models or extra securityOn-demand (no term) or reserved (fixed 1- or 3-month term); HIPAA available
    AI Functions optimized modelsrun batch inferenceRecommended for batch inference workloads

    Provisioned throughput also matters when the best model is one you customised yourself. It supports "Fine-tuned variants of base models, such as models that are fine-tuned on proprietary data," and base models registered from Hugging Face or another external source, an approach that "works with any fine-tuned variant of the supported models." If your application needs a model that isn't hosted by Databricks at all, external models let you reach third-party providers such as OpenAI, Anthropic, Cohere, Amazon Bedrock and Google Cloud Vertex AI through one endpoint interface.

    Checkpoint 1 of 5· Match them up

    Match each application to the Foundation Model APIs mode recommended for it

    Tap a term, then the definition that fits it.

    Checkpoint 2 of 5· Exam question

    A legal team wants to summarize full contracts, some running 50 to 100 pages, in a single prompt without splitting the document into chunks. Which model attribute should be prioritized when selecting an LLM for this application?

    Sources12

    2.Check that the model can carry the load

    A model that fits the task can still fail the application if it can't handle the traffic. Pay-per-token endpoints have per-model limits on input tokens per minute (ITPM), output tokens per minute (OTPM) and queries per hour (QPH), and those limits differ a lot between models. The figures below are for Enterprise tier workspaces.

    Selected pay-per-token rate limits (Enterprise tier)
    ModelITPM limitOTPM limitQPH limit
    GPT OSS 20B1,000,000100,000360,000
    GPT OSS 120B1,000,000100,000360,000
    Claude Opus 51,000,000100,000360,000
    Claude Opus 4.8200,00020,000360,000
    GLM 5.32,000,00040,0007,200
    DeepSeek V4 Pro (0813)200,0004,0007,200

    Read the table against your application's traffic. GLM 5.3 accepts a very high input-token rate but allows only 7,200 queries per hour, so an app that sends many short requests will hit the QPH limit long before the token limit. The limits page also links model size to workload: "Smaller models for high-volume tasks: Use models like GPT OSS 20B for tasks that require higher throughput," and reserve GPT OSS 120B for tasks that require maximum capability. Some models also cap request size; for Gemini 2.5 Pro and Gemini 2.5 Flash, "Requests must be smaller than 200K input tokens or 400KB."

    If a model is otherwise the right fit but its limits are too low, you don't have to switch models. The page says: "For models that support provisioned throughput, create a provisioned throughput endpoint if the pay-per-token limits don't meet your use case requirements."

    Finally, confirm that the model is actually available to you. Model assets in system.ai are listed globally, and "Seeing a model there does not mean it is available in your workspace." Availability depends on region, cross-geo settings and the model itself, so check the Unity Gateway UI and your permissions before you commit.

    Checkpoint 3 of 5· Check yourself

    Evaluation shows Claude Opus 4.8 is the best-quality fit, but projected traffic exceeds its pay-per-token limits. The model supports provisioned throughput. What do the docs recommend?

    Sources34

    3.Confirm the choice by comparing candidates

    Model descriptions and limit tables give you a shortlist, not a final answer. Every entry on the supported-models page carries the same warning: "Databricks recommends customers evaluate their models to ensure they're operating as intended." The Foundation Model APIs make it cheap to do that. One listed use is to "Efficiently compare LLMs to see the best candidate for your use case, or swap a production model with a better performing one." Another is to replace proprietary models with open alternatives "to optimize for cost and performance."

    A comparison needs the same metrics for every candidate. Quality is one of them. System performance is another: "Overall latency and token consumption are examples of chain performance metrics," and both can be measured deterministically from the application's outputs. Token counts matter twice: they drive cost, and they show whether a candidate stays within its ITPM and OTPM limits. The limits page shows how to read them from each response.

    Checkpoint 4 of 5· Fill the gap

    Complete the snippet so it records the number of output tokens a candidate model produced

    # Example: Track token usage
    response = model.generate(prompt)
    input_tokens = response.usage.prompt_tokens
    output_tokens = response.usage. ? 
    total_tokens = response.usage.total_tokens

    Checkpoint 5 of 5· Exam question

    A team has fine-tuned a foundation model and is preparing to move a customer-facing application to production, where consistent response latency and guaranteed throughput under load are required. How should the fine-tuned model be served?

    Sources156

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Priority pay-per-token is the right mode for any production workload, including ones with compliance requirements such as HIPAA.Why is that wrong?

      Priority mode is for latency-sensitive traffic. Provisioned throughput is the mode recommended for production and the one listed with compliance certifications like HIPAA.

      Covered in Match the serving mode to the workload

    2. 2.If a model appears in system.ai, it can be used from my workspace.Why is that wrong?

      system.ai is a global listing. Whether you can query a model depends on workspace region, cross-geo settings and model availability.

      Covered in Check that the model can carry the load

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “This mode is recommended for all production workloads, especially those that require high throughput, performance guarantees, fine-tuned models, or have additional security requirements.”
      ↩︎ Match the serving mode to the workload
      “Priority mode is opt-in per request, requires no commitment, and is billed at a higher per-token rate.”
      ↩︎ Match the serving mode to the workload
      “AI Functions optimized models: This mode is recommended for batch inference workloads.”
      ↩︎ Match the serving mode to the workload
      “Efficiently compare LLMs to see the best candidate for your use case, or swap a production model with a better performing one.”
      ↩︎ Confirm the choice by comparing candidates
      “Provisioned throughput endpoints are available with compliance certifications like HIPAA.”
      ↩︎ Exam trap 1
      “Provisioned throughput endpoints are available with compliance certifications like HIPAA.”
      ↩︎ Prediction
    2. 2.
      “External models are third-party models hosted outside of Databricks.”
      ↩︎ Match the serving mode to the workload
    3. 3.
      “Smaller models for high-volume tasks: Use models like GPT OSS 20B for tasks that require higher throughput.”
      ↩︎ Check that the model can carry the load
      “Requests must be smaller than 200K input tokens or 400KB.”
      ↩︎ Check that the model can carry the load
      “create a provisioned throughput endpoint if the pay-per-token limits don't meet your use case requirements.”
      ↩︎ Checkpoint
    4. 4.
      “Availability depends on your workspace region, cross-geo settings, and model availability.”
      ↩︎ Check that the model can carry the load
      “Seeing a model there does not mean it is available in your workspace.”
      ↩︎ Exam trap 2
    5. 5.
      “Databricks recommends customers evaluate their models to ensure they're operating as intended.”
      ↩︎ Confirm the choice by comparing candidates
    6. 6.
      “Overall latency and token consumption are examples of chain performance metrics.”
      ↩︎ Confirm the choice by comparing candidates

    Spotted a mistake, or was something unclear? Tell us.