CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 4 · Lesson 34/56

    Querying Foundation Model APIs from an LLM Application

    Identify how to serve an LLM application that leverages Foundation Model APIs

    11 min read
    1.79% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Choose a query method (OpenAI client, AI Functions, Serving UI, REST API, MLflow Deployments SDK, Databricks Python SDK) for an application
    • Apply the recommended authentication for development and for production
    • Write a chat request with the OpenAI-compatible client and read the response
    • Run batch inference over a foundation model from SQL with ai_query
    • Route an agent's LLM calls through a Unity Gateway endpoint

    1.Ways an application can call a foundation model

    After you pick a serving mode, the application needs a way to send requests. Model Serving uses one OpenAI-compatible API for Databricks-hosted Foundation Model APIs and for external models. That lets the same calling code work whichever provider is behind the endpoint. All of these requests go through Unity Gateway, where you can apply rate limits, budgets and guardrails.

    If you just want to call an LLM, the docs suggest model services in Unity Gateway. These give you governed access to Databricks-hosted models without configuring a serving endpoint. If you serve the model on an endpoint you host yourself, the calling code is the same except for one value: you replace the model service name (for example system.ai.claude-sonnet-4-5) with your endpoint name.

    Query methods for foundation model endpoints
    MethodHow it worksSetup
    OpenAI clientPass the endpoint (or model service) name as the model input; supports chat, embeddings and completionspip install -U databricks-openai
    AI FunctionsCall model inference from SQL with ai_queryNone beyond SQL access
    Serving UIQuery endpoint on the Serving endpoint page; Show Example loads a logged input exampleNone
    REST APIPOST /serving-endpoints/{name}/invocationsDatabricks API token
    MLflow Deployments SDKCall predict() on a deployments clientpip install mlflow
    Databricks Python SDKA layer over the REST API that handles authenticationPreinstalled on Databricks Runtime 13.3 LTS or above

    Checkpoint 1 of 6· Check yourself

    A developer is building a long-running chat application on Foundation Model APIs. Which interface do the docs recommend for extended interactions?

    Sources1

    2.Authenticating the application

    The OpenAI client, the REST API and the MLflow Deployments SDK all need a Databricks API token to send scoring requests. Which token depends on the stage. For testing and development, use a personal access token that belongs to a service principal rather than to a workspace user. In production, switch to machine-to-machine OAuth tokens.

    The databricks-openai package means you don't have to wire up credentials by hand: it "provides an OpenAI client with authorization automatically configured." In a notebook you install it and then run dbutils.library.restartPython(). You only use the plain openai client, with your token as api_key and the workspace's /ai-gateway/mlflow/v1 URL as base_url, when you query foundation models from outside the workspace.

    Checkpoint 2 of 6· Check yourself

    A data scientist is testing an endpoint from a dev notebook. Which token do the docs recommend?

    Checkpoint 3 of 6· Exam question

    A developer wants to minimize code changes while switching their application between a Databricks-hosted Foundation Model APIs endpoint and a third-party hosted LLM, and also wants to reuse existing tooling built around a widely adopted client library. Which client interface does Databricks recommend for querying Foundation Model APIs endpoints in this scenario?

    Sources1

    3.Sending a chat request and reading the response

    Inside the workspace, a chat call is a standard OpenAI chat.completions.create call. Only the model value is Databricks-specific.

    Querying a pay-per-token chat model with DatabricksOpenAIpython
    from databricks_openai import DatabricksOpenAI
    
    client = DatabricksOpenAI()
    
    response = client.chat.completions.create(
        model="system.ai.claude-sonnet-4-5",
        messages=[
          {
            "role": "system",
            "content": "You are a helpful assistant."
          },
          {
            "role": "user",
            "content": "What is a mixture of experts model?",
          }
        ],
        max_tokens=256
    )

    A REST call carries the same fields in a JSON body: messages, max_tokens and temperature. The response follows the OpenAI chat completion shape, including a usage object that counts prompt and completion tokens. On pay-per-token, that token count is what you are billed for.

    Expected REST response format for a chat modeljson
    {
      "model": "databricks-claude-sonnet-4-5",
      "choices": [
        {
          "message": {},
          "index": 0,
          "finish_reason": null
        }
      ],
      "usage": {
        "prompt_tokens": 7,
        "completion_tokens": 74,
        "total_tokens": 81
      },
      "object": "chat.completion",
      "id": null,
      "created": 1698824353
    }

    Checkpoint 4 of 6· Check yourself

    Your app queries the model service system.ai.claude-sonnet-4-5. You move it to a model serving endpoint that you host. What changes in the client call?

    Sources2

    4.Batch inference from SQL with ai_query

    Not every LLM application serves one user at a time. To transform rows in a table, you can call the same foundation models from SQL with the built-in ai_query function. This is the AI Functions path, and the serving mode recommended for batch inference is AI Functions optimized models. Note that the function is in Public Preview and its definition might change.

    Checkpoint 5 of 6· Fill the gap

    Which SQL function sends this prompt to a foundation model?

    SELECT  ? (
        "system.ai.claude-sonnet-4-5",
        "Can you explain AI in ten words?"
      )

    Sources21

    5.Serving an agent app that calls Foundation Model APIs

    A whole LLM application, such as an agent with a chat UI, can be deployed on Databricks Apps. Apps "gives you full control over the agent code, server configuration, and deployment workflow." The app still gets its model from a foundation model endpoint. The recommended way to wire this up is to send every LLM call through Unity Gateway: in the agent code, pass the gateway endpoint name as model and set use_ai_gateway=True on the Databricks LLM client (AsyncDatabricksOpenAI or ChatDatabricks in LangChain). The client then handles authentication. Because the gateway sits in the request path, you can swap models, attribute cost per app, and inspect or replay traffic without changing agent code or rotating provider credentials.

    Checkpoint 6 of 6· Check yourself

    An agent on Databricks Apps uses ChatDatabricks. What makes its LLM calls go through Unity Gateway?

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A developer's personal access token is the right credential for an app querying Foundation Model APIs in production.Why is that wrong?

      Production should use machine-to-machine OAuth tokens. Even for development, the recommendation is service-principal PATs, not tokens tied to workspace users.

      Covered in Authenticating the application

    2. 2.DatabricksOpenAI works the same way from outside the workspace.Why is that wrong?

      To query from outside the workspace you use the plain OpenAI client, configured with your token and the workspace URL.

      Covered in Authenticating the application

    Practise it for real

    Query a pay-per-token foundation model from a notebook, first with the OpenAI-compatible client and then from SQL

    1. 1.In a notebook, run pip install -U databricks-openai, then dbutils.library.restartPython()

      Why: The package provides an OpenAI client with authorization configured for Databricks

      You should see: The package installs and the Python process restarts

    2. 2.Run the DatabricksOpenAI chat.completions.create example with model="system.ai.claude-sonnet-4-5"

      Why: This is the recommended OpenAI-compatible path for calling a model service

      You should see: A chat completion object whose usage shows prompt, completion and total tokens

    3. 3.Change the model argument to a different general purpose model from the supported list and run it again

      Why: Shows that the shared API lets you compare or swap models without changing any other code

      You should see: A response in the same format from the other model

    4. 4.In a SQL cell, run SELECT ai_query("system.ai.claude-sonnet-4-5", "Can you explain AI in ten words?")

      Why: ai_query is the AI Functions path for model inference from SQL, used for batch workloads

      You should see: A single text response returned as the query result

    Stuck? Get a nudge

    If the model name is rejected, check whether you are calling a model service (system.ai.*) or a serving endpoint, and use the matching name.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “routed through Unity Gateway, which lets you apply rate limits, budgets, and guardrails to control cost and access.”
      ↩︎ Ways an application can call a foundation model
      “Model services give you governed access to foundation models served natively by Databricks through a single interface, without configuring a serving endpoint.”
      ↩︎ Ways an application can call a foundation model
      “This package provides an OpenAI client with authorization automatically configured to query AI models.”
      ↩︎ Authenticating the application
      “Invoke model inference directly from SQL using the ai_query SQL function.”
      ↩︎ Batch inference from SQL with ai_query
      “As a security best practice for production scenarios, Databricks recommends that you use machine-to-machine OAuth tokens for authentication during production.”
      ↩︎ Exam trap 1
      “As a security best practice for production scenarios, Databricks recommends that you use machine-to-machine OAuth tokens for authentication during production.”
      ↩︎ Prediction
      “For testing and development, Databricks recommends using a personal access token belonging to service principals instead of workspace users.”
      ↩︎ Checkpoint
    2. 2.
      “If you use model serving endpoints instead of model services, replace the model service name with an endpoint name.”
      ↩︎ Sending a chat request and reading the response
      “This function is in Public Preview and the definition might change.”
      ↩︎ Batch inference from SQL with ai_query
      “To query foundation models outside of your workspace, you must use the OpenAI client directly.”
      ↩︎ Exam trap 2
    3. 3.
      “you can centralize permissions, attribute cost per app, swap models, and inspect or replay traffic without modifying agent code or rotating provider credentials.”
      ↩︎ Serving an agent app that calls Foundation Model APIs
      “Databricks Apps gives you full control over the agent code, server configuration, and deployment workflow.”
      ↩︎ Serving an agent app that calls Foundation Model APIs
      “In your agent code, pass the Unity Gateway endpoint name as the model argument and set use_ai_gateway=True on the Databricks LLM client.”
      ↩︎ Checkpoint

    Also cited

    Spotted a mistake, or was something unclear? Tell us.