CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 4 · Lesson 34/56

    Foundation Model APIs: Serving Modes and Provisioned Throughput Endpoints

    Identify how to serve an LLM application that leverages Foundation Model APIs

    10 min read
    1.79% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain what Foundation Model APIs give an LLM application compared with deploying and maintaining your own model
    • Choose between pay-per-token, priority pay-per-token, provisioned throughput (on-demand or reserved), and AI Functions for a given workload
    • List the requirements for using Foundation Model APIs
    • Create a provisioned throughput endpoint from system.ai, through the UI or the REST API, and size it using the model optimization information API

    Key concept

    Foundation Model APIs — A Model Serving capability where Databricks hosts the foundation models and exposes them as serving endpoints. Your LLM application calls the endpoint and never deploys or operates the model. What you do choose is the serving mode, and that choice sets the cost, latency consistency, and capacity guarantees.

    1.What Foundation Model APIs give an LLM application

    Every LLM application needs a model it can call. On Databricks the shortest path is Foundation Model APIs, a feature of Model Serving that lets you "access and query state-of-the-art open models from a serving endpoint." Databricks hosts the models, so your work is the application code and not the model infrastructure.

    The documentation lists what this is for: checking whether a project is viable before you invest more, building a quick proof of concept, building a RAG chatbot by pairing a foundation model with a vector index, replacing proprietary models with open ones, comparing LLMs or swapping in a better one, and running production apps on "a scalable, SLA-backed LLM serving solution." The models are also a Databricks Designated Service, which means Databricks Geos manage data residency when customer content is processed.

    Checkpoint 1 of 5· Check yourself

    Compared with deploying your own model, what do Foundation Model APIs let a team skip?

    Sources1

    2.Picking a serving mode

    The model is the same in every mode. What changes is how you pay for capacity and what performance you get. Pay-per-token is the starting point: Databricks preconfigures endpoints for the pay-per-token models, and you find them by clicking AI Gateway in the sidebar, where they are listed as system-provided model services.

    For latency-sensitive production traffic on pay-per-token there is priority mode. You set the service_tier request parameter to "priority", and Databricks admits the request ahead of standard best-effort pay-per-token traffic. You opt in per request, there is no commitment, and you pay a higher per-token rate.

    Provisioned throughput gives you dedicated capacity. It is the recommended mode for production, and it is required when you serve fine-tuned or fully custom weights. It comes in two options: on-demand, with no term, and reserved, with a fixed 1- or 3-month term at a lower per-unit rate. Provisioned throughput endpoints are also available with compliance certifications such as HIPAA. The last mode, AI Functions optimized models, is meant for batch inference.

    Foundation Model APIs serving modes compared
    ModeCommitment / billingRecommended for
    Pay-per-tokenNo commitment; billed per tokenGetting started; preconfigured endpoints are already in the workspace
    Priority pay-per-token (service_tier set to "priority")Opt-in per request, no commitment, higher per-token rateLatency-sensitive production workloads that need consistent performance under load
    Provisioned throughput: on-demandDedicated capacity, no term; change or remove it at any timeProduction workloads that need high throughput, performance guarantees, fine-tuned models, or additional security
    Provisioned throughput: reservedDedicated capacity reserved for a fixed 1- or 3-month termThe same production needs, with capacity committed for a fixed term
    AI Functions optimized modelsRun through AI FunctionsBatch inference workloads

    Checkpoint 2 of 5· Check yourself

    A real-time support assistant runs on pay-per-token and has inconsistent latency at peak. The team does not want dedicated capacity or a term commitment. What should they do?

    Sources1

    3.Requirements and model task types

    Before an application can call Foundation Model APIs, three things must be in place. You need a Databricks API token to authenticate requests. You need serverless compute if you use provisioned throughput models. Your workspace must be in a supported region, and the pay-per-token and provisioned throughput modes each have their own list of regions.

    The models fall into task types, and the task type tells you which model a feature needs. General purpose (chat) models handle multi-turn conversation, for example virtual assistants and support bots. Embedding models turn data into vectors for semantic search and RAG; the Databricks-hosted ones are databricks-qwen3-embedding-0-6b, databricks-gte-large-en and databricks-bge-large-en. Vision models analyse images and documents, and reasoning models cover code generation and agent orchestration. A RAG application usually needs two of these at once: an embedding model for retrieval and a chat model for generation.

    Checkpoint 3 of 5· Match them up

    Match each task type to the use case the documentation recommends it for

    Tap a term, then the definition that fits it.

    Sources1

    4.Creating a provisioned throughput endpoint

    Databricks recommends serving the foundation models that come pre-installed in Unity Catalog, in the system catalog under the ai schema (system.ai). In the UI you open system.ai in Catalog Explorer, click the model, then click Serve this model. That opens the Create serving endpoint page. In the Up to dropdown you set the maximum tokens per second. The endpoint scales automatically, and Modify shows the minimum it can scale down to. One restriction: for Meta Llama models you must choose an Instruct version, because base versions cannot be deployed from Unity Catalog.

    Checkpoint 4 of 5· Put it in order

    Put the UI steps for serving a foundation model from Unity Catalog in order

    1. 1.Navigate to system.ai in Catalog Explorer
    2. 2.Configure the endpoint on the Create serving endpoint page
    3. 3.Click on the name of the model to deploy
    4. 4.On the model page, click the Serve this model button

    With the REST API, the request body depends on how the model measures capacity: in throughput bands or in model units. In both cases you first call the model optimization information API, GET api/2.0/serving-endpoints/get-model-optimization-info/{registered_model_name}/{version}. Its response says whether the model is optimizable and gives the chunk size, which is the increment you provision in.

    Example optimization info response for a throughput-band modeljson
    {
      "optimizable": true,
      "model_type": "llama",
      "throughput_chunk_size": 980
    }
    Throughput bands vs model units in the REST request
    Capacity measureChunk size fieldRequest fieldsPOST path
    Throughput bandsthroughput_chunk_sizemin_provisioned_throughput and max_provisioned_throughput/api/2.0/serving-endpoints
    Model unitsmodel_unit_chunk_sizeprovisioned_model_units/api/2.0/serving-endpoints/pt
    Creating a model-unit endpoint (here for system.ai.gpt-oss-120b): note the /pt pathpython
    # Send the POST request to create the serving endpoint
    data = {
      "name": endpoint_name,
      "config": {
        "served_entities": [
          {
            "entity_name": model_name,
            "entity_version": model_version,
            "provisioned_model_units": desired_model_units,
          }
        ]
      },
    }
    
    response = requests.post(
      url=f"{API_ROOT}/api/2.0/serving-endpoints/pt", json=data, headers=headers
    )

    Checkpoint 5 of 5· Fill the gap

    For a model that measures capacity in model units, which key of the optimization info response gives the increment?

    chunk_size = optimizable_info[' ? ']
    
    # Desired provisioned throughput
    desired_model_units = 2 * chunk_size

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Provisioned throughput is the easiest way to start using Foundation Model APIs.Why is that wrong?

      Pay-per-token is the recommended starting point, and its endpoints are already preconfigured in the workspace. Provisioned throughput is the recommendation for production workloads that need guarantees.

      Covered in Picking a serving mode

    2. 2.Every provisioned throughput endpoint is created by POSTing min/max throughput to /api/2.0/serving-endpoints.Why is that wrong?

      Only throughput-band models use min_provisioned_throughput and max_provisioned_throughput. Model-unit models use provisioned_model_units and POST to the /pt path.

      Covered in Creating a provisioned throughput endpoint

    3. 3.Any Meta Llama variant in system.ai can be served on provisioned throughput.Why is that wrong?

      Only Instruct versions can be deployed from Unity Catalog. Base versions are not supported.

      Covered in Creating a provisioned throughput endpoint

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “These models are hosted by Databricks and you can quickly and easily build applications that use them without maintaining your own model deployment.”
      ↩︎ What Foundation Model APIs give an LLM application
      “This mode is recommended for all production workloads, especially those that require high throughput, performance guarantees, fine-tuned models, or have additional security requirements.”
      ↩︎ Picking a serving mode
      “AI Functions optimized models: This mode is recommended for batch inference workloads.”
      ↩︎ Picking a serving mode
      “Reserved provisioned throughput: reserve dedicated capacity for a fixed 1- or 3-month term.”
      ↩︎ Picking a serving mode
      “Serverless compute (for provisioned throughput models).”
      ↩︎ Requirements and model task types
      “Databricks API token to authenticate endpoint requests.”
      ↩︎ Requirements and model task types
      “These models are hosted by Databricks and you can quickly and easily build applications that use them without maintaining your own model deployment.”
      ↩︎ Key concept
      “This is the easiest way to start accessing foundation models on Databricks and is recommended for beginning your journey with Foundation Model APIs.”
      ↩︎ Exam trap 1
      “Use a foundation model, along with a vector index, to build a chatbot using retrieval augmented generation (RAG).”
      ↩︎ Prediction
      “Priority mode is opt-in per request, requires no commitment, and is billed at a higher per-token rate.”
      ↩︎ Checkpoint
    2. 2.
      “Databricks recommends using the foundation models that are pre-installed in Unity Catalog.”
      ↩︎ Creating a provisioned throughput endpoint
      “To identify the suitable range for your needs, Databricks recommends using the model optimization information API within the platform.”
      ↩︎ Creating a provisioned throughput endpoint
      “Provisioned throughput endpoints automatically scale, so you can select Modify to view the minimum tokens per second your endpoint can scale down to.”
      ↩︎ Creating a provisioned throughput endpoint
      “For these models, you specify the provisioned_model_units field in your request and send the POST request to the /api/2.0/serving-endpoints/pt endpoint.”
      ↩︎ Exam trap 2
      “Base versions of the Meta Llama models are not supported for deployment from Unity Catalog.”
      ↩︎ Exam trap 3
      “On the model page, click the Serve this model button.”
      ↩︎ Checkpoint

    Also cited

    Continue to page 2 of 2

    Querying Foundation Model APIs from an LLM Application

    Spotted a mistake, or was something unclear? Tell us.