CertSafari
    Databricks Certified Data Analyst Associate· Lessons

    Domain 3 · Lesson 7/39

    Delta Sharing, Databricks Marketplace and API-Driven Data Intake

    Explain the approaches for bringing data into Databricks, covering ingestion from S3, data sharing with external systems via Delta Sharing, API-driven data intake, the Auto Loader feature, and Marketplace.

    9 min read
    2.56% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain shares, providers and recipients, and how sharing data with external organizations works
    • Distinguish Databricks-to-Databricks sharing from Databricks-to-Open sharing
    • Describe how consumers find, request and access Databricks Marketplace data products
    • Explain why to check for a managed connector before writing custom API ingestion

    1.Delta Sharing: shares, providers and recipients

    Some data comes into Databricks from another organization rather than from a bucket. The exam guide calls this capability Delta Sharing. The current documentation it points to calls the platform OpenSharing, so expect both names. OpenSharing is an open protocol for secure data sharing that lets you share data and AI assets with users outside your organization, whether or not they use Databricks. It also underpins Databricks Marketplace and Clean Rooms.

    Three objects do the work. A share is a read-only collection of tables and table partitions. If the recipient uses Unity Catalog, a share can also include views, volumes, models and notebook files. A provider shares the data. A recipient receives it. In Unity Catalog, a recipient is a securable object tied to a credential or a sharing identifier. Access is revoked at either end. Remove a share and every recipient of it loses access. Delete a recipient and it loses access to every share. To share from several metastores, you define the recipient separately in each one.

    Checkpoint 1 of 7· Match them up

    Match each sharing concept to its role

    Tap a term, then the definition that fits it.

    A provider with account admin or metastore admin privileges sets up sharing in four steps: enable OpenSharing on the metastore that manages the data, create a share and add data to it, create a recipient, and grant the recipient access. You only need the enable step when sharing with other accounts or with non-Databricks clients. Sharing between metastores in the same account is always on.

    Checkpoint 2 of 7· Put it in order

    Put the provider's setup steps for sharing with an external organization in order

    1. 1.Create a share and add data to it
    2. 2.Enable OpenSharing on the Unity Catalog metastore that manages the data
    3. 3.Grant the recipient access to the share
    4. 4.Create a recipient

    Sources1

    2.Databricks-to-Databricks vs Databricks-to-Open sharing

    From a Unity Catalog workspace, which protocol you use depends on who receives the data. Databricks-to-Databricks sharing reaches users on a different Unity Catalog metastore, in any account and on AWS, Azure or GCP. Databricks-to-Open sharing reaches anyone, including people who don't use Databricks at all. They read the data with tools such as Apache Spark, pandas or Power BI.

    The two sharing protocols available from a Unity Catalog-enabled workspace
    AspectDatabricks-to-DatabricksDatabricks-to-Open
    RecipientUsers on a Unity Catalog-enabled workspace with a different metastoreAny user, on any computing platform
    AuthenticationNo recipient token; handled by the Databricks platformLong-lived bearer token, or OIDC federation with short-lived OAuth tokens
    Shareable assetsTables plus notebooks, volumes and modelsTabular data

    Checkpoint 3 of 7· Check yourself

    A partner analyses data in pandas and has no Databricks workspace. How do they get access to your shared table?

    Checkpoint 4 of 7· Exam question

    A data analyst needs to give a partner organization ongoing, read-only access to a set of Unity Catalog tables so the partner can query the latest data directly from their own Power BI environment, without Databricks ever copying the underlying data files to the partner's infrastructure. The partner does not use Databricks at all. Which approach meets this requirement?

    Sources1

    3.Databricks Marketplace

    Databricks Marketplace is built on the same sharing protocol. It's an open exchange where providers publish listings and Databricks customers discover them. Listings include datasets, AI models, notebooks, apps and MCP servers. Datasets usually arrive as catalogs of tabular data, but volumes of non-tabular data are supported too. Some listings are public. Others belong to a private exchange that only member consumers can see.

    To consume listings in a Unity Catalog workspace, you need a Premium plan or above and the USE MARKETPLACE ASSETS privilege on the metastore. That privilege is enabled for all users by default. Free, instantly available listings need only Get instant access and acceptance of the terms. You can rename the suggested catalog, and the data then appears as a read-only catalog in Catalog Explorer. Listings marked By request need provider approval, usually because a commercial transaction is involved. Once data is shared, you don't need a workspace to work with it. External platforms such as Power BI, pandas or Apache Spark can read it through Databricks-to-Open connectors.

    Checkpoint 5 of 7· Check yourself

    After you click Get instant access on a free listing, where does the data appear?

    Checkpoint 6 of 7· Exam question

    An analyst wants to bring a third-party demographic dataset into their workspace to enrich internal sales tables, and would prefer to browse a catalog of vetted external providers, request access from within the workspace UI, and have the resulting tables appear directly in Unity Catalog without writing any ingestion code. Which capability is designed for this?

    Sources23

    4.API-driven data intake

    The last way in is pulling data from a web service: fetching it over HTTP, usually as paginated JSON. Files and message buses have built-in sources, but there's no generic API source. If you write your own API ingestion, you handle authentication, pagination and rate limits yourself. Lakeflow pipelines support three patterns for ingesting from an arbitrary API, and which one fits depends on data volume and how often you need to refresh. The sources available for this lesson don't describe the three patterns, so they aren't covered here.

    Before writing any of that code, check whether a managed connector already exists. Lakeflow Connect has built-in connectors for common SaaS APIs such as Salesforce, Workday, ServiceNow and Google Analytics, plus a growing set of partner connectors. A connector handles authentication, pagination and incremental extraction for you. Custom code is for when no connector fits. Either way, store credentials as Databricks secrets, never in pipeline code, and make sure the pipeline compute can reach the endpoint over the network.

    Checkpoint 7 of 7· Check yourself

    A team needs Salesforce data in Databricks and plans to write a Python loop that calls the REST API. What should they do first?

    Sources4

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Giving another workspace on the same metastore access to your tables requires a share.Why is that wrong?

      Sharing is for crossing metastores or organizations. Within one metastore, Unity Catalog grants already control access across workspaces.

      Covered in Databricks-to-Databricks vs Databricks-to-Open sharing

    2. 2.You need a Databricks workspace to use Marketplace data at all.Why is that wrong?

      A workspace is needed to request a data product. Once the data is shared, you can work with it from other platforms.

      Covered in Databricks Marketplace

    3. 3.Pipelines have a built-in generic REST API source, just like they have for files.Why is that wrong?

      There's no generic API source. Use a managed connector, or write custom code that handles authentication, pagination and rate limits yourself.

      Covered in API-driven data intake

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “lets you share data and AI assets with users outside your organization, regardless of whether they use Databricks”
      ↩︎ Delta Sharing: shares, providers and recipients
      “If you remove a share from your Unity Catalog metastore, all recipients of that share lose the ability to access it.”
      ↩︎ Delta Sharing: shares, providers and recipients
      “you must define the recipient separately for each metastore”
      ↩︎ Delta Sharing: shares, providers and recipients
      “Databricks-to-Databricks sharing between Unity Catalog metastores in the same account is always enabled.”
      ↩︎ Delta Sharing: shares, providers and recipients
      “the share recipient doesn't need a token to access the share, and the provider doesn't need to manage recipient tokens”
      ↩︎ Databricks-to-Databricks vs Databricks-to-Open sharing
      “Databricks-to-Databricks also supports notebook, volume, and model sharing, which is not available in Databricks-to-Open sharing.”
      ↩︎ Databricks-to-Databricks vs Databricks-to-Open sharing
      “there is no need to use OpenSharing to share data between workspaces attached to the same Unity Catalog metastore”
      ↩︎ Exam trap 1
      “a share is a read-only collection of tables and table partitions that a provider wants to share with one or more recipients”
      ↩︎ Checkpoint
      “Enable OpenSharing on the Unity Catalog metastore that manages the data.”
      ↩︎ Checkpoint
      “there is no need to use OpenSharing to share data between workspaces attached to the same Unity Catalog metastore”
      ↩︎ Prediction
      “You generate a long-lived bearer token and share it securely with the recipient.”
      ↩︎ Checkpoint
    2. 2.
      “Listings include datasets, AI models, notebooks, apps, and Model Context Protocol (MCP) servers.”
      ↩︎ Databricks Marketplace
      “The Open Marketplace, which does not require access to a Databricks workspace.”
      ↩︎ Databricks Marketplace
      “You do not need a Databricks workspace to access and work with data once it is shared”
      ↩︎ Exam trap 2
      “To request access to data products in the Marketplace, you must use the Marketplace on a Databricks workspace.”
      ↩︎ Prediction
    3. 3.
      “This privilege is enabled for all users on all Unity Catalog metastores by default.”
      ↩︎ Databricks Marketplace
      “Some data products require provider approval, typically because a commercial transaction is involved”
      ↩︎ Databricks Marketplace
      “Click the Open button to view the data product, which appears as a read-only catalog in Catalog Explorer.”
      ↩︎ Checkpoint
    4. 4.
      “pulling data over HTTP from a web service, usually as paginated JSON”
      ↩︎ API-driven data intake
      “Lakeflow Connect ships built-in connectors for many common software as a service (SaaS) APIs, such as Salesforce, Workday, ServiceNow, and Google Analytics”
      ↩︎ API-driven data intake
      “Never hardcode credentials in pipeline source code.”
      ↩︎ API-driven data intake
      “there's no built-in generic API source, so you handle authentication, pagination, and rate limits yourself”
      ↩︎ Exam trap 3
      “Before you write any custom API-ingestion code, check whether a managed connector already exists for your source.”
      ↩︎ Checkpoint

    Spotted a mistake, or was something unclear? Tell us.