What you will be able to do
- Explain which identity a model serving endpoint uses to reach Unity Catalog resources, and why that identity can't be changed after creation
- List the Unity Catalog grants the endpoint creator needs, and say which are checked at creation time and which only fail at query time
- Grant CAN QUERY on an endpoint and name the permission level needed to change endpoint permissions
- Restrict which foundation models can be invoked, and describe the network controls that apply to serving endpoints
Key concept
Endpoint's recorded creator — Databricks records the identity that created a serving endpoint and uses it to reach Unity Catalog resources for that endpoint. It can't be changed, so its grants and workspace membership are re-checked on every update.
1.Whose permissions does the endpoint run with?
A model serving endpoint doesn't borrow the permissions of whoever happens to query it. When you create an endpoint, Databricks records the calling identity as the endpoint's creator. The documentation describes this as typically a service principal, but it is whichever identity made the call, including a person's user account as in the scenario above. It is the identity used to reach Unity Catalog resources for the endpoint, and you can't reassign it later.
Two checks apply to both the caller and the recorded creator before an endpoint can be created or updated. Each must be a member of the workspace, and each must hold the workspace-access entitlement. Configuration and served-entity updates check the creator's membership and grants again. If the recorded creator loses grants or is removed from the workspace, there is no way to transfer the endpoint. You must delete the endpoint and recreate it under a service principal that has the required permissions and is still a workspace member. This is why production endpoints are usually created by a service principal rather than by a person.
The recorded creator needs these grants on each served Unity Catalog model: USE CATALOG on the catalog, USE SCHEMA on the schema and EXECUTE on the model. They are validated when the endpoint is created or updated, and a missing grant fails the request with PERMISSION_DENIED. If the model declares transitive function dependencies, the creator also needs EXECUTE on those upstream functions.
The timing matters for troubleshooting. Grants that are only needed at query time are not checked in advance, so the endpoint deploys without complaint and then fails with runtime errors once it serves traffic.
| Resource | Required grant | When validated |
|---|---|---|
| Catalog containing the model | USE CATALOG | Endpoint creation or update |
| Schema containing the model | USE SCHEMA | Endpoint creation or update |
| Unity Catalog model | EXECUTE | Endpoint creation or update |
Checkpoint 1 of 5· Check yourself
An endpoint deployed without errors, but it returns errors as soon as it receives traffic. The served model's grants were validated at deployment. What is the most likely cause?
Creation-time grants would have failed the deployment. Grants needed only at query time are checked when traffic arrives, so they show up as runtime errors.
“Grants required at query time are not validated upfront — missing grants cause runtime errors when the endpoint serves traffic.”Source: docs.databricks.com
Sources1
2.Who may query or manage the endpoint
The creator's grants control what the endpoint can reach. A separate set of permissions on the endpoint itself controls who can use it. To grant a user permission to call the endpoint, give them CAN QUERY. To change the endpoint's permissions at all, you need at least CAN MANAGE on it. You can view and change permissions with the Permissions button in the UI, the Databricks CLI or the Permissions API. In the CLI, databricks permissions get serving-endpoints <endpoint-id> lists the current permissions. Granting a permission looks like this:
databricks permissions update serving-endpoints <endpoint-id> --json '{
"access_control_list": [
{
"user_name": "jsmith@example.com",
"permission_level": "CAN_QUERY"
}
]
}'Checkpoint 2 of 5· Check yourself
A team lead wants to give an analyst query access to an existing serving endpoint. What is the minimum permission the team lead needs on that endpoint?
Changing endpoint permissions requires at least CAN MANAGE. CAN QUERY is the permission being handed out, and the Unity Catalog grants apply to the model, not to the endpoint's access list.
“You must have at least the CAN MANAGE permission on a serving endpoint to modify permissions.”Source: docs.databricks.com
Agents that call other workspace resources, such as AI Search indexes or LLM endpoints, rely on automatic authentication passthrough. For this to work, the agent must be logged with the resources it needs. If a deployed agent hits authentication errors reaching those resources, inspect the logged resources by loading the model with mlflow.models.Model.load(model_uri) and printing its resources. To re-add missing or incorrect resources, log the agent again and redeploy it.
If you use manual authentication instead, check that the environment variables are set correctly. Manual settings override any automatic authentication configuration.
Checkpoint 3 of 5· Exam question
A GenAI engineer is deploying an agent to a Model Serving endpoint. The agent must query a Unity Catalog Vector Search index and call a Unity Catalog function, and the team wants Databricks to manage and rotate credentials automatically instead of storing any secrets by hand. How should the engineer configure this?
Correct answer: A — Declare the index and function as resource dependencies via mlflow.models.resources, prompting Databricks to issue a system-generated service principal with short-lived, rotated credentials.
- A. Declaring the Vector Search index and function as resource dependencies at logging time triggers automatic authentication passthrough: Databricks provisions a system-generated service principal and issues short-lived, auto-rotated M2M OAuth tokens scoped to those resources, with no manual secret management required.
- B. Manually storing a personal access token as a secret and referencing it works for external systems, but it requires the engineer to manage and rotate the token themselves, which is exactly the manual overhead the team wants to avoid for Databricks-native resources.
- C. On-behalf-of-user authentication is designed to have each end user's own identity govern access at request time; caching a single token in `__init__` defeats that purpose and is also the wrong lifecycle point, since user identity is only known inside the prediction call.
- D. Granting a broad admin role to the endpoint's identity violates least-privilege access and is not how Databricks scopes credentials for declared resources; automatic passthrough intentionally limits access to only the resources that were declared.
3.Restricting which foundation models can be called
For Databricks-hosted foundation models, access is controlled through Unity Catalog on the system.ai schema. This is a feature your account must already be enrolled in. You also need Unity Catalog enabled and account admin or metastore admin privileges. A metastore admin can grant and revoke system.ai schema permissions without an account admin. To build an allow-list:
1. Revoke EXECUTE on system.ai from All Users. This clears all default access to models, so no user can invoke any model until permissions are explicitly re-granted.
2. In Catalog Explorer, under system > ai > models, grant EXECUTE on each approved model to All Users or to specific groups.
3. Delete every provisioned throughput endpoint that serves a model you don't allow.
Step 3 is needed because enforcement isn't the same across the three workload types. Pay-per-token endpoints and batch inference through AI Functions enforce the new permissions automatically, and calls to them stop immediately after the revoke. Provisioned throughput endpoints keep serving until you delete them manually.
Checkpoint 4 of 5· Put it in order
Put the steps for restricting an organization to approved foundation models in order
- 1.Revoke EXECUTE on the system.ai schema from All Users
- 2.Delete provisioned throughput endpoints that serve disallowed models
- 3.Grant EXECUTE on each approved model under system.ai models
Revoking at the schema level removes all default access first. Per-model grants then build the allow-list. Provisioned throughput endpoints don't enforce the change, so you delete them by hand.
“Removing EXECUTE from the system.ai schema clears all default access to models.”Source: docs.databricks.com
Sources4
4.Network controls around the endpoint
Network rules sit outside identity and permissions. Model Serving endpoints are protected by access control, and they also respect the networking rules configured on the workspace. For inbound traffic, that means ingress rules such as IP allowlists and PrivateLink. For outbound traffic, you can limit where an endpoint may connect by configuring network policies for serverless egress control. One gap to know: by default, Model Serving does not support PrivateLink to external endpoints. Support is evaluated and rolled out region by region through your Databricks account team.
None of them do. The IP allowlist is an ingress rule for inbound traffic, CAN QUERY decides who can call the endpoint, and Unity Catalog grants decide which governed objects the creator identity can use. Only network policies for serverless egress control restrict outbound connections from the endpoint.
Checkpoint 5 of 5· Check yourself
Which statement about Model Serving networking is supported by the documentation?
Endpoints follow workspace ingress rules. Outbound traffic can be restricted with network policies. PrivateLink to external endpoints is not supported by default.
“Model Serving endpoints are protected by access control and respect networking-related ingress rules configured on the workspace, like IP allowlists and PrivateLink.”Source: docs.databricks.com
Sources5
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.If the person who created an endpoint leaves, an admin can reassign the endpoint to a different owner.Why is that wrong?
The recorded creator can't be changed. If that identity loses grants or workspace membership, you must delete the endpoint and recreate it under a service principal that has the required permissions.
2.A successful deployment proves the endpoint has every permission it needs.Why is that wrong?
Only some grants are checked at creation or update. Grants needed at query time aren't validated in advance and fail as runtime errors once traffic arrives.
3.Revoking EXECUTE on a foundation model immediately stops every endpoint that serves it.Why is that wrong?
Pay-per-token endpoints and AI Functions batch inference enforce the change automatically. Provisioned throughput endpoints keep serving until you delete them.
Covered in Restricting which foundation models can be called
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/model-serving/create-manage-serving-endpointsOfficial docs
“When you create an endpoint, Databricks records the calling identity as the endpoint's creator.”
↩︎ Whose permissions does the endpoint run with?“USE CATALOG on the catalog, USE SCHEMA on the schema, EXECUTE on the model”
↩︎ Whose permissions does the endpoint run with?“the recorded creator also needs EXECUTE on those upstream functions”
↩︎ Whose permissions does the endpoint run with?“is used to access Unity Catalog resources on behalf of the endpoint and cannot be changed after creation”
↩︎ Key concept“you must delete the endpoint and recreate it under a service principal that has the required permissions”
↩︎ Exam trap 1“Grants required at query time are not validated upfront — missing grants cause runtime errors when the endpoint serves traffic.”
↩︎ Exam trap 2“Updates fail with PERMISSION_DENIED if the recorded creator is no longer a workspace member, even when the caller has valid permissions.”
↩︎ Prediction“Grants required at query time are not validated upfront — missing grants cause runtime errors when the endpoint serves traffic.”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/machine-learning/model-serving/manage-serving-endpointsOfficial docs
“Grant user jsmith@example.com the CAN QUERY permission on the serving endpoint.”
↩︎ Who may query or manage the endpoint“You must have at least the CAN MANAGE permission on a serving endpoint to modify permissions.”
↩︎ Checkpoint - 3.
“verify that it was logged with the necessary resources for automatic authentication passthrough”
↩︎ Who may query or manage the endpoint“To re-add missing or incorrect resources, log the agent and deploy it again.”
↩︎ Who may query or manage the endpoint - 4.https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/model-uc-permissionsOfficial docs
“Pay-per-token and batch inference (AI Functions) endpoints automatically enforce the new permissions.”
↩︎ Restricting which foundation models can be called“Repeat for each approved model. This creates an allow-list of permitted models.”
↩︎ Restricting which foundation models can be called“A metastore admin can grant and revoke system.ai schema permissions without an account admin.”
↩︎ Restricting which foundation models can be called“Your account must already be enrolled in foundation model Unity Catalog permissions.”
↩︎ Restricting which foundation models can be called“Provisioned throughput endpoints continue serving until manually deleted.”
↩︎ Exam trap 3“Removing EXECUTE from the system.ai schema clears all default access to models.”
↩︎ Checkpoint - 5.
“You can restrict outbound network access from Model Serving endpoints by configuring network policies.”
↩︎ Network controls around the endpoint“By default, Model Serving does not support PrivateLink to external endpoints.”
↩︎ Network controls around the endpoint“Model Serving endpoints are protected by access control and respect networking-related ingress rules configured on the workspace, like IP allowlists and PrivateLink.”
↩︎ Checkpoint