CertSafari

    Free Microsoft Certified: Azure Databricks Data Engineer Associate (DP-750) Sample Questions

    35 free sample questions from our bank of 348+, covering every exam domain, with answers and detailed explanations. Updated September 2026.

    Domain 1: Set up and configure an Azure Databricks environment

    Subdomain 1.2: Create and organize objects in Unity Catalog

    1.A data analyst is preparing a Genie space before sharing it with business users so that natural-language answers are accurate. Which configuration elements should the analyst add to improve response quality? (Select all that apply.)(Select 3)

    1. A.General instructions written in plain language that define ambiguous business terms and clarify how metrics should be calculated.
    2. B.Example SQL queries, reviewed and accepted from suggested workspace queries, that Genie can reuse as trusted patterns for similar questions.
    3. C.A requirement that every business user first learn SQL, since natural-language questions are only a fallback once SQL access is unavailable.
    4. D.Assigning `MANAGE` privileges on the Genie space to every business user so they can edit instructions if an answer looks wrong.
    5. E.Reviewing suggested queries surfaced from popular workspace queries associated with the added tables and accepting the relevant ones.
    Show answer & explanation

    Correct answers: A, B, E — General instructions written in plain language that define ambiguous business terms and clarify how metrics should be calculated.; Example SQL queries, reviewed and accepted from suggested workspace queries, that Genie can reuse as trusted patterns for similar questions.; Reviewing suggested queries surfaced from popular workspace queries associated with the added tables and accepting the relevant ones.

    • A. General instructions let the analyst spell out business terminology and calculation rules so Genie consistently interprets ambiguous natural-language phrases the same way.
    • B. Accepted example SQL queries become trusted patterns that Genie can reuse for structurally similar questions, improving the accuracy of generated SQL.
    • C. Requiring SQL literacy defeats the purpose of a natural-language interface and is not a documented step for improving a Genie space's answer quality.
    • D. Broadly granting `MANAGE` to every business user is a permissions escalation unrelated to answer accuracy, and it risks unintended edits to shared instructions.
    • E. Reviewing and accepting Genie's suggested queries, drawn from relevant existing workspace query history, is a documented step for building out trustworthy example queries.

    Subdomain 1.2: Create and organize objects in Unity Catalog

    2.To register a database in an external analytics platform as a Unity Catalog object using Lakehouse Federation, run `CREATE ___ CATALOG ... USING CONNECTION ...`.

    1. A.FOREIGN
    2. B.SHARED
    3. C.MANAGED
    Show answer & explanation

    Correct answer: A — FOREIGN

    • A. `CREATE FOREIGN CATALOG ... USING CONNECTION` is the documented statement for mirroring an external database through Lakehouse Federation.
    • B. A shared catalog is created with `USING SHARE` to consume an OpenSharing share, not to mirror a database through a connection.
    • C. `MANAGED` is not a catalog type keyword in the `CREATE CATALOG` statement; managed storage is instead configured with the `MANAGED LOCATION` clause on a standard catalog.

    Subdomain 1.1: Select and configure compute in a workspace

    3.A cost-conscious team configures an instance pool with several idle instances on standby so future clusters start faster. No cluster is currently attached to the pool. What is true about the cost of those idle pooled instances?

    1. A.Azure Databricks does not charge DBUs for idle pooled instances, though the underlying cloud instance provider billing still applies to them.
    2. B.Azure Databricks charges full DBUs for idle pooled instances at the same rate as instances actively running Spark workloads on a cluster.
    3. C.Idle pooled instances accrue DBU charges only for the driver node equivalent, while worker-equivalent idle instances remain completely unbilled.
    4. D.Idle pooled instances are entirely free of charge, including cloud provider VM billing, until a cluster attaches and starts using them.
    Show answer & explanation

    Correct answer: A — Azure Databricks does not charge DBUs for idle pooled instances, though the underlying cloud instance provider billing still applies to them.

    • A. Databricks doesn't charge DBUs while instances sit idle in a pool, but the cloud provider still bills for the underlying VM capacity.
    • B. Idle pooled instances specifically avoid DBU charges; only once a cluster attaches and runs workloads do DBU charges for that usage begin.
    • C. Pools don't distinguish idle instances by driver/worker role for DBU purposes; DBU charges are waived across all idle pooled instances alike.
    • D. Cloud provider VM billing still applies to idle instances in a pool; only the Databricks DBU portion is waived while they're idle.

    Subdomain 1.1: Select and configure compute in a workspace

    4.An engineer configures autoscaling on a multi-node classic compute resource and expects it can be scaled down to zero worker nodes when idle, the same way single-node compute has zero worker nodes. Evaluate this statement: this expectation matches how autoscaling behaves on multi-node compute.

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: B — False

    • A. This statement is incorrect, so True does not match how autoscaling limits apply to multi-node compute resources.
    • B. This statement is false — a multi-node compute resource cannot be scaled down to zero workers; single-node compute exists specifically for that zero-worker case.

    Subdomain 1.1: Select and configure compute in a workspace

    5.A workspace admin is provisioning compute that many analysts will share for concurrent notebook work, and needs data isolated per user: compute using ___ access mode allows many users to attach to and concurrently execute workloads on the same resource, with per-user isolation from Lakeguard.

    1. A.Dedicated
    2. B.No isolation shared
    3. C.Standard
    Show answer & explanation

    Correct answer: C — Standard

    • A. Dedicated access mode restricts the compute resource to a single assigned user or group, which doesn't match a many-user concurrent scenario.
    • B. No isolation shared access mode is a legacy configuration that lacks the per-user workload isolation Lakeguard provides under standard access mode.
    • C. Standard access mode is the mode that supports many concurrent users attaching to one resource, with Lakeguard isolating each user's workload.

    Subdomain 1.2: Create and organize objects in Unity Catalog

    6.Within a single `analytics` catalog shared by multiple teams, a platform admin wants clear separation between development and production data without creating additional catalogs, while still allowing workspace-catalog binding at the catalog level later if needed. What is the most appropriate approach?

    1. A.Prefix schema names consistently, such as `dev_orders` and `prod_orders`, so environment is visible in the schema name while both stay in the same catalog.
    2. B.Store development and production tables together in one shared schema, using the table comments on each object to note which rows belong to which environment.
    3. C.Grant every user `MANAGE` privilege on the entire catalog so each team can informally track which schemas are currently safe to treat as production.
    4. D.Create duplicate table names across schemas without any naming distinction, relying on team memory to know which schema is currently used for production.
    Show answer & explanation

    Correct answer: A — Prefix schema names consistently, such as `dev_orders` and `prod_orders`, so environment is visible in the schema name while both stay in the same catalog.

    • A. A consistent schema-name prefix like `dev_` or `prod_` gives immediate visual separation of environments within one catalog, while leaving room to promote to catalog-level isolation later.
    • B. Table comments describe data for discovery, but they don't create the structural naming boundary needed to tell environments apart at a glance or in access policies.
    • C. Broadly granting `MANAGE` increases the blast radius of accidental changes and does nothing to establish a naming convention that distinguishes environments.
    • D. Identical, undistinguished names across schemas invite accidental cross-environment queries and provide no naming signal at all.

    Domain 2: Secure and govern Unity Catalog objects

    Subdomain 2.1: Secure Unity Catalog objects

    7.An architect is designing how an Azure Databricks workspace will authenticate to an Azure Data Lake Storage Gen2 account for Unity Catalog managed storage, and wants to avoid maintaining or rotating any credential secrets for this connection. Which authentication approach best satisfies this requirement?

    1. A.Create an access connector for Azure Databricks, which deploys with a managed identity, and grant that identity a role on the storage account.
    2. B.Register an Entra ID application as a service principal, generate a client secret, and store the secret in the storage credential definition.
    3. C.Create a shared access signature token scoped to the storage account and embed the token string directly in the storage credential.
    4. D.Store the storage account's access key as a Databricks secret and reference that secret from every job that reads the storage account.
    Show answer & explanation

    Correct answer: A — Create an access connector for Azure Databricks, which deploys with a managed identity, and grant that identity a role on the storage account.

    • A. An access connector for Azure Databricks deploys with a system-assigned managed identity by default, and granting that identity a role on the storage account lets Unity Catalog authenticate without any credential to store or rotate. This is the credential-free pattern the architect wants.
    • B. A service principal's client secret is a credential that must be stored securely and rotated periodically, which is exactly the maintenance burden the architect wants to avoid. This approach still relies on a secret.
    • C. A shared access signature token is itself a time-limited credential that must eventually be regenerated and re-embedded, so it does not eliminate credential rotation. It also is not the storage credential mechanism Unity Catalog uses for managed identities.
    • D. A storage account access key is a long-lived shared secret that requires careful rotation and distribution to every consuming job. Storing it as a Databricks secret still means the workspace depends on a credential rather than a managed identity.

    Subdomain 2.1: Secure Unity Catalog objects

    8.A column mask UDF bound to a table's `salary` column can accept additional columns from the same row as extra input parameters, letting the mask vary its output based on more than just the `salary` value itself.

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: A — True

    • A. This statement is correct: column masks can take other columns as inputs to vary the masking behavior based on multiple attributes of the same row, such as masking `salary` differently depending on a `department` column.
    • B. This statement is factually accurate, so labeling it false would be incorrect. Column masks are documented as supporting multiple input columns, not just the masked column alone.

    Subdomain 2.1: Secure Unity Catalog objects

    9.An Azure Key Vault-backed secret scope in Azure Databricks is a read-only interface into the vault, meaning secrets referenced through the scope must be created, updated, and deleted directly in Azure rather than through the Databricks CLI.

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: A — True

    • A. This statement is correct: Azure Key Vault-backed secret scopes are documented as a read-only interface to the vault, so secret values themselves must be managed in Azure, not through Databricks secret-writing commands.
    • B. This statement matches the documented behavior of Key Vault-backed scopes, so labeling it false would be incorrect. Only Databricks-backed scopes support creating and deleting secrets directly through the Databricks CLI.

    Subdomain 2.2: Govern Unity Catalog objects

    10.A data platform team manages a Unity Catalog metastore where the number of tables grows by dozens each week across multiple schemas. Compliance requires that any table tagged as containing PII automatically receives a mandatory column mask on identifier columns, without any table owner intervention. Which approach satisfies this requirement?

    1. A.Attach an ABAC policy at the catalog level that matches columns via a governed tag, so newly tagged tables are covered automatically.
    2. B.Ask each table owner to run `ALTER TABLE ... ALTER COLUMN ... SET MASK` on every PII column as soon as a new table is created in their schema.
    3. C.Build a scheduled job that scans the catalog nightly and issues `ALTER TABLE` commands to add column masks to any table it finds containing PII columns.
    4. D.Create a dynamic view over each PII table that hides the identifier columns, and require downstream consumers to query the view instead of the base table.
    Show answer & explanation

    Correct answer: A — Attach an ABAC policy at the catalog level that matches columns via a governed tag, so newly tagged tables are covered automatically.

    • A. Attaching a policy at the catalog level lets Unity Catalog evaluate governed tag conditions dynamically, so any table or column later tagged as PII is protected without further action. This is the ABAC approach and it scales because policy authors and data stewards can work independently.
    • B. Manually running the mask statement on each new table depends on every owner remembering to do it, which does not meet a hands-off, automatic requirement and does not scale as table volume grows.
    • C. A nightly scan-and-alter job still leaves PII columns unmasked for up to a day after creation, and it duplicates work that a tag-driven policy already automates without a custom scheduler.
    • D. A dynamic view can hide identifier columns, but it only protects consumers who are redirected to query the view; anyone who still queries the base table directly sees the unmasked PII, so it does not enforce the mask automatically at the table level.

    Subdomain 2.2: Govern Unity Catalog objects

    11.A schema has ten stable tables that rarely change, and each table needs its own custom row-level logic tailored to that specific table's business rules. The team wants each table owner to manage and adjust their own filter directly, without introducing a governed tag taxonomy. Which approach fits best?

    1. A.Apply table-level row filters with `ALTER TABLE ... SET ROW FILTER` so each owner manages logic directly.
    2. B.Create an ABAC row filter policy at the schema level so a single tag-driven rule automatically covers all ten tables.
    3. C.Build a governed tag taxonomy and assign tags to each table so future ABAC policies can match them dynamically.
    4. D.Wrap each table in a dynamic view with `is_account_group_member()` checks so the row logic can span multiple base tables.
    Show answer & explanation

    Correct answer: A — Apply table-level row filters with `ALTER TABLE ... SET ROW FILTER` so each owner manages logic directly.

    • A. Table-level row filters are bound to a single table with `ALTER TABLE ... SET ROW FILTER` and are managed directly by the table owner, which fits a small, stable set of tables where each one needs its own tailored logic.
    • B. A schema-level ABAC policy is designed for consistent, tag-driven enforcement across many tables and takes decision-making away from individual owners, which conflicts with the requirement that each owner manage their own filter.
    • C. Building a tag taxonomy is the ABAC path, which trades per-table owner control for centralized, automatic coverage — the opposite of what a small set of independently managed tables needs here.
    • D. A dynamic view can encode custom logic, but it introduces a new object that consumers must be redirected to query, and it does not let the owner manage protection on the original table the way a bound row filter does.

    Subdomain 2.2: Govern Unity Catalog objects

    12.A workspace admin configured Azure diagnostic settings to send Databricks logs to a Log Analytics workspace, but a security review later finds no record of group membership changes anywhere in Log Analytics. What explains the gap?

    1. A.Group events are captured only in the audit log system table, not through Azure diagnostic settings, so the admin must query `system.access.audit` for that category.
    2. B.Group membership changes are not audited by Azure Databricks at all, in either the system table or diagnostic settings, regardless of configuration.
    3. C.Diagnostic settings only forward logs generated in the last 24 hours, so any older group-membership events were already dropped before the review took place.
    4. D.Log Analytics requires a separate Azure AD Premium license to display Unity Catalog-related events, which is unrelated to the Databricks diagnostic settings configuration.
    Show answer & explanation

    Correct answer: A — Group events are captured only in the audit log system table, not through Azure diagnostic settings, so the admin must query `system.access.audit` for that category.

    • A. Some categories, including group events, cluster policy events, and certain account-level activity, are recorded in the audit log system table but are not among the subset of categories exported through Azure diagnostic settings, so the system table must be queried directly.
    • B. Azure Databricks does audit group creation and membership changes; the gap is about which delivery channel carries that category, not whether the platform tracks it at all.
    • C. Diagnostic settings forward events continuously as they occur rather than applying a rolling 24-hour cutoff, so an age-based drop does not explain a category being absent entirely.
    • D. The audit log system table is a Databricks Premium-plan feature, not a feature gated by a separate Azure AD Premium license, so licensing at that layer is not the cause of the missing events.

    Domain 3: Prepare and process data

    Subdomain 3.1: Design and implement data modeling in Unity Catalog

    13.A finance team requests a fact table for expense approvals where every report must show one row per individual expense line item, including partial approvals within a single expense report. Which grain should the table be designed at?

    1. A.One row per expense line item, since that is the lowest level of detail the reports require.
    2. B.One row per expense report, since aggregating line items keeps the table smaller and easier to maintain.
    3. C.One row per employee per month, since expense reports can be summarized at a monthly reporting cadence.
    4. D.One row per approval batch, since approvals are typically processed together in scheduled batch runs.
    Show answer & explanation

    Correct answer: A — One row per expense line item, since that is the lowest level of detail the reports require.

    • A. Because reports must show partial approvals within a single expense report, the table needs to preserve the line-item level of detail rather than summarizing it away.
    • B. One row per report hides the individual line items and their separate approval states, which the requirement explicitly needs to see.
    • C. Summarizing to employee-per-month discards both the report and line-item detail needed to show partial approvals accurately.
    • D. Grouping by approval batch mixes line items from potentially different reports into one row and still loses the required line-item granularity.

    Subdomain 3.1: Design and implement data modeling in Unity Catalog

    14.Which statement accurately compares managed Apache Iceberg tables and Delta Lake tables in Unity Catalog?

    1. A.Both formats can use liquid clustering, though Iceberg support requires a newer Databricks Runtime and remains in public preview.
    2. B.Iceberg tables cannot be managed by Unity Catalog at all and must always be registered as external tables.
    3. C.Delta Lake tables do not support ACID transactions, which is the main advantage Iceberg offers over Delta.
    4. D.Liquid clustering is exclusive to Delta Lake tables and has no equivalent capability planned for Iceberg tables.
    Show answer & explanation

    Correct answer: A — Both formats can use liquid clustering, though Iceberg support requires a newer Databricks Runtime and remains in public preview.

    • A. Liquid clustering is generally available for Delta Lake tables and available in public preview for managed Apache Iceberg tables on a newer Databricks Runtime, so both formats support it today under different maturity levels.
    • B. Unity Catalog supports managed Apache Iceberg tables directly, so Iceberg data does not have to be registered only as an external table.
    • C. Delta Lake tables do support ACID transactions through their transaction log, so this is not a real advantage Iceberg holds over Delta.
    • D. Liquid clustering has already been extended to managed Iceberg tables in preview, so it is not an exclusively Delta capability.

    Subdomain 3.2: Ingest data into Unity Catalog

    15.A source database exposes nightly full-table exports to cloud storage but has no change-data-feed or CDC log enabled. A pipeline needs to detect inserts, updates, and deletes between each night's export and turn them into SCD Type 2 history. Which approach fits?

    1. A.`create_auto_cdc_from_snapshot_flow`, comparing each new export against the previous one to derive the changes.
    2. B.`create_auto_cdc_flow` reading a Delta change data feed generated automatically from the nightly exports.
    3. C.A SQL `AUTO CDC INTO` statement reading the raw export files with `SEQUENCE BY` on the file's modification time.
    4. D.`COPY INTO` loading each night's export as new rows appended to a single history table.
    Show answer & explanation

    Correct answer: A — `create_auto_cdc_from_snapshot_flow`, comparing each new export against the previous one to derive the changes.

    • A. `create_auto_cdc_from_snapshot_flow` is designed for exactly this case: it compares successive in-order snapshots to derive inserts, updates, and deletes when no CDC feed exists, and can store the result as SCD Type 2.
    • B. `create_auto_cdc_flow` processes an existing change feed; a nightly full-table export is a snapshot, not a change feed, so there is nothing for this API to consume without first deriving changes.
    • C. `AUTO CDC INTO` expects a stream of already-identified change records with keys and a sequence column; raw full-table exports do not carry per-row change type or sequencing information on their own.
    • D. Appending each night's full export to one table produces a stack of duplicate snapshots rather than a computed history of what changed between them.

    Subdomain 3.2: Ingest data into Unity Catalog

    16.True or False: The `AUTO CDC` APIs replace the older `APPLY CHANGES` APIs and use the same syntax, so pipelines written against `APPLY CHANGES` continue to work even though Databricks recommends migrating to `AUTO CDC`.

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: A — True

    • A. Databricks documentation states that `AUTO CDC` replaces `APPLY CHANGES` with identical syntax, and that `APPLY CHANGES` remains available even as `AUTO CDC` becomes the recommended API.
    • B. This statement matches documented behavior, so treating it as false would misrepresent how Databricks handles the API transition and backward compatibility.

    Subdomain 3.2: Ingest data into Unity Catalog

    17.Which three benefits does Auto Loader provide over reading files directly with `spark.readStream.format(fileFormat).load(directory)`? (Select 3)(Select 3)

    1. A.It eliminates the need to maintain a Delta table target, since Auto Loader can write query results directly to arbitrary JDBC destinations.
    2. B.It automatically detects schema drift and can rescue data that would otherwise be dropped or lost.
    3. C.It guarantees files are always processed in the exact order they were written to storage.
    4. D.It uses native cloud APIs, and file notification mode can avoid directory listing entirely to reduce cost.
    5. E.It removes the need for a checkpoint location, since Auto Loader stores all state in the destination Delta table itself.
    6. F.It can scale to ingesting billions of files, including asynchronous backfills, without wasting compute resources.
    Show answer & explanation

    Correct answers: B, D, F — It automatically detects schema drift and can rescue data that would otherwise be dropped or lost.; It uses native cloud APIs, and file notification mode can avoid directory listing entirely to reduce cost.; It can scale to ingesting billions of files, including asynchronous backfills, without wasting compute resources.

    • A. Auto Loader is a Structured Streaming source (`cloudFiles`) that writes into Delta tables like any other streaming source; it does not remove the need for a Delta target or add native JDBC sink support.
    • B. Auto Loader can detect schema drift as new columns appear and rescue data that doesn't match the expected schema instead of silently dropping or losing it.
    • C. Auto Loader explicitly does not guarantee the order in which files are discovered or processed, whether using directory listing or file notification mode.
    • D. Auto Loader uses native cloud APIs to list files, and file notification mode can avoid directory listing altogether, which further reduces cloud API costs.
    • E. Auto Loader still requires a checkpoint location to persist ingestion progress in a scalable key-value store; it does not eliminate this requirement.
    • F. Auto Loader is built to scale to billions of files and supports asynchronous backfills so that large-scale historical loads don't waste ongoing compute resources.

    Subdomain 3.4: Implement and manage data quality constraints in Unity Catalog

    18.Which statement correctly describes the purpose of the overwriteSchema write option on a Delta table?

    1. A.It allows a write to fully replace the table's schema, including changing a column's data type or removing columns, rather than only adding new ones.
    2. B.It allows a write to add new columns from the source to the end of the target schema while leaving all existing columns and their types unchanged.
    3. C.It allows a streaming query to resume from its last checkpoint after the source topology changes, without requiring a new checkpoint location.
    4. D.It allows a table to accept writes from multiple concurrent streams by automatically partitioning incoming data across the writers.
    Show answer & explanation

    Correct answer: A — It allows a write to fully replace the table's schema, including changing a column's data type or removing columns, rather than only adding new ones.

    • A. Correct — overwriteSchema replaces the table's schema entirely with the schema of the incoming write, which is how column type changes, renames, or removals are applied by rewriting the table.
    • B. Incorrect — additive column growth without touching existing columns describes mergeSchema behavior, not overwriteSchema.
    • C. Incorrect — checkpoint recovery and resuming a stream is a Structured Streaming concept unrelated to a table write option that governs schema replacement.
    • D. Incorrect — overwriteSchema controls how a write's schema is applied to the table, not how concurrent writers partition or coordinate incoming data.

    Subdomain 3.4: Implement and manage data quality constraints in Unity Catalog

    19.A streaming table declares CONSTRAINT valid_count EXPECT (count > 0) ON VIOLATION FAIL UPDATE, and during a triggered pipeline run a batch contains one row where count is zero. What happens to the table update?

    1. A.The entire update is rolled back atomically, so none of the batch's rows, valid or invalid, are committed to the target table.
    2. B.Only the single invalid row is excluded from the commit, while every other valid row in the same batch is still written to the target table.
    3. C.The invalid row is written with count forced to 1 so the constraint passes, and the pipeline continues processing the rest of the batch.
    4. D.The update completes normally and an alert email is sent to the pipeline owner listing the row that violated the constraint.
    Show answer & explanation

    Correct answer: A — The entire update is rolled back atomically, so none of the batch's rows, valid or invalid, are committed to the target table.

    • A. Correct — the FAIL UPDATE action stops execution immediately on a violation and atomically rolls back the transaction, so nothing from that batch is committed.
    • B. Incorrect — partially committing valid rows while excluding only the invalid one describes the drop action, not fail, which rolls back the whole update instead.
    • C. Incorrect — a fail expectation never rewrites a record's values to force it to pass; it stops the update instead of silently correcting data.
    • D. Incorrect — Lakeflow does not send email alerts as part of expectation handling; the failure is surfaced through the pipeline's error message and event log, not email.

    Subdomain 3.3: Cleanse, transform, and load data into Unity Catalog

    20.A `customers` table has a `phone_number` column that is null for about 5% of rows because the field was optional at signup, and the marketing team wants to keep every customer record while flagging which ones lack contact information. What should the engineer do?

    1. A.Keep the null values as-is and add a derived boolean column flagging whether `phone_number` is null.
    2. B.Run `dropna` on the `phone_number` column so that every remaining row has a complete contact record.
    3. C.Replace every null `phone_number` value with an empty string so the column always contains printable text.
    4. D.Delete the `phone_number` column entirely, since a partially populated field adds no analytical value.
    Show answer & explanation

    Correct answer: A — Keep the null values as-is and add a derived boolean column flagging whether `phone_number` is null.

    • A. Adding a derived flag preserves every customer row while giving marketing a clean way to filter on missing contact information, which matches the stated requirement exactly.
    • B. `dropna` on `phone_number` would discard roughly 5% of legitimate customer records, violating the requirement to keep every customer row.
    • C. Replacing nulls with empty strings hides the missing-data signal instead of flagging it, and downstream filters would need extra logic to distinguish empty strings from real values.
    • D. Deleting the column removes information the marketing team explicitly wants to use for flagging incomplete records.

    Subdomain 3.3: Cleanse, transform, and load data into Unity Catalog

    21.A team needs every customer in the `customers` table who has never placed an order in the `orders` table, without returning any columns from the `orders` table itself. Which join type produces exactly this result?

    1. A.A `LEFT ANTI JOIN` from `customers` to `orders` on the customer key, returning only unmatched customers.
    2. B.A `LEFT OUTER JOIN` from `customers` to `orders` on the customer key, returning unmatched customers with no further filter step.
    3. C.An `INNER JOIN` between `customers` and `orders` restricted to customers with a matching order key value.
    4. D.A `FULL OUTER JOIN` between `customers` and `orders`, then discarding rows that matched on either side.
    Show answer & explanation

    Correct answer: A — A `LEFT ANTI JOIN` from `customers` to `orders` on the customer key, returning only unmatched customers.

    • A. `LEFT ANTI JOIN` is purpose-built to return only rows from the left table that have no match in the right table, giving exactly the unmatched customers with no order columns at all.
    • B. A plain `LEFT OUTER JOIN` alone returns both matched and unmatched customer rows with null order columns for the unmatched ones; without a further filter it does not isolate only the customers with no orders.
    • C. An `INNER JOIN` returns only customers that do have a matching order, which is the opposite of what is needed here.
    • D. A `FULL OUTER JOIN` followed by discarding matched rows would still require extra filtering logic and would also include unmatched order rows that have no customer, which is out of scope for this request.

    Subdomain 3.3: Cleanse, transform, and load data into Unity Catalog

    22.In a `MERGE INTO` statement, the clause that updates or deletes target rows which have no matching row in the source table is called `WHEN _____`.

    1. A.NOT MATCHED BY SOURCE
    2. B.NOT MATCHED BY TARGET
    3. C.MATCHED BY DEFAULT
    Show answer & explanation

    Correct answer: A — NOT MATCHED BY SOURCE

    • A. `WHEN NOT MATCHED BY SOURCE` fires for target rows that have no matching source row, and only supports `UPDATE` or `DELETE` actions on those target rows.
    • B. `WHEN NOT MATCHED BY TARGET` (or plain `WHEN NOT MATCHED`) fires for unmatched source rows and inserts them, which is the opposite side of the merge from what this clause describes.
    • C. `MATCHED BY DEFAULT` is not a real `MERGE INTO` clause; the valid clause keywords are `MATCHED`, `NOT MATCHED [BY TARGET]`, and `NOT MATCHED BY SOURCE`.

    Subdomain 3.1: Design and implement data modeling in Unity Catalog

    23.An analytics team ingests a nightly export from an on-premises ERP system where downstream reports only need the current day's rows, and the source system cannot provide a change log. Which extraction and file type design best fits this constraint?

    1. A.Perform a full extract of the source table on each run and land it as Parquet files, then overwrite the target table from the latest extract.
    2. B.Perform an incremental extract using a change data feed from the ERP system and land the deltas as Avro files, then merge into the target table.
    3. C.Perform a full extract of the source table on each run and land it as compressed CSV files, then append every run to the target table.
    4. D.Perform an incremental extract based on a `modified_date` column and land the deltas as JSON files, then merge into the target table.
    Show answer & explanation

    Correct answer: A — Perform a full extract of the source table on each run and land it as Parquet files, then overwrite the target table from the latest extract.

    • A. Since the source offers no reliable change-tracking mechanism, extracting the full table each run and overwriting the target with the latest snapshot avoids relying on data that doesn't exist, and Parquet's columnar format loads efficiently.
    • B. The scenario states the source cannot provide a change log, so designing extraction around a change data feed will not work against this ERP system.
    • C. Appending a full extract every run without overwriting duplicates every previous day's rows in the target table instead of reflecting only the current day's snapshot.
    • D. The scenario gives no indication that a trustworthy `modified_date` column exists on the source, so an incremental extract built around it risks missing or double-counting rows.

    Subdomain 3.4: Implement and manage data quality constraints in Unity Catalog

    24.A streaming job writes to the Delta table `iot.readings`, whose `temperature_celsius` column is typed DOUBLE. A malformed upstream record arrives with the string "N/A" in that field, and the write is attempted without any schema evolution options set. What happens?

    1. A.The write is rejected because Delta's schema enforcement checks that incoming types match the column's declared type, and "N/A" cannot be cast to DOUBLE.
    2. B.The write succeeds because Delta silently widens the column's type to permissive text, storing "N/A" as a literal string so the streaming job keeps ingesting mixed values going forward.
    3. C.The write succeeds because Azure Databricks silently coerces the invalid "N/A" string to NULL during the streaming merge, so ingestion continues without pausing the job or raising an error.
    4. D.The write succeeds because Delta's schema evolution automatically routes the malformed value into a `_rescued_data` column while storing NULL in temperature_celsius for that record.
    Show answer & explanation

    Correct answer: A — The write is rejected because Delta's schema enforcement checks that incoming types match the column's declared type, and "N/A" cannot be cast to DOUBLE.

    • A. Correct — Delta's schema enforcement validates that every incoming value matches the target column's declared type before it commits a write, so a non-numeric string cannot be written into a DOUBLE column and the transaction fails with a schema mismatch error.
    • B. Incorrect — Delta does not silently widen a column's type to accept mixed values; enforcement blocks the type mismatch rather than accommodating it.
    • C. Incorrect — Delta does not convert unparseable strings to NULL on write; the transaction fails so the engineer can address the malformed value.
    • D. Incorrect — the _rescued_data column is an Auto Loader ingestion feature that captures unparsed data during schema inference from files, not a behavior of a direct write against an already-typed Delta table.

    Domain 4: Deploy and maintain data pipelines and workloads

    Subdomain 4.1: Design and implement data pipelines

    25.A job task partially wrote output rows to a Delta table using an `INSERT` before it failed. An engineer clicks Repair run to re-run only the failed task. What should the engineer verify before trusting the repaired result?

    1. A.That the task's write logic is idempotent, for instance using `MERGE` or overwrite, since a repair re-runs the task from the very beginning.
    2. B.That the job's maximum concurrent runs setting has been increased, since repair runs otherwise queue behind the original failed run indefinitely.
    3. C.That the notebook's cluster has been switched from a job cluster to an all-purpose cluster, since repair runs cannot execute on job clusters.
    4. D.That the task's retry count has been set to zero, since a nonzero retry count silently disables the Repair run feature for that task.
    Show answer & explanation

    Correct answer: A — That the task's write logic is idempotent, for instance using `MERGE` or overwrite, since a repair re-runs the task from the very beginning.

    • A. Databricks documentation warns that a repair re-runs each unsuccessful task from the beginning, and because Lakeflow Jobs doesn't make tasks idempotent, an append-style write that already wrote partial output before failing can duplicate that data on repair.
    • B. Repair runs are not blocked by the maximum concurrent runs setting queuing behind the original run; that setting limits how many separate runs of the same job can be active, which is unrelated to the safety of re-running a single failed task.
    • C. Repair runs are supported on job clusters; in fact a repair on a shared job cluster creates a new versioned cluster rather than requiring a switch to all-purpose compute.
    • D. The retry count and the Repair run feature are independent controls; a nonzero retry count governs automatic re-attempts during the original run and does not disable the ability to manually repair the run afterward.

    Subdomain 4.1: Design and implement data pipelines

    26.A pipeline task type within a Lakeflow Job can run an existing Lakeflow Spark Declarative Pipeline as one step in that job's larger task graph, alongside notebook and other task types.

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: A — True

    • A. This statement is correct: a pipeline task is one of the supported task types in Lakeflow Jobs, letting a job orchestrate a materialized view or streaming table pipeline as a dependency alongside notebook tasks and other task types in the same DAG.
    • B. This statement is not the false case here; pipeline tasks are a documented, supported way to include a Lakeflow Spark Declarative Pipeline as a node in a job's task graph.

    Subdomain 4.3: Implement development lifecycle processes in Azure Databricks

    27.Before a pull request that changes `resources/jobs/*.yml` is allowed to merge, a CI job must confirm the bundle configuration is syntactically valid and report the resolved bundle identity, without deploying anything or touching the remote workspace state. Which command satisfies this check?

    1. A.`databricks bundle validate`, which checks the configuration files and prints a summary of the bundle's identity without deploying resources.
    2. B.`databricks bundle plan`, which builds the bundle and lists every create, update, and delete action that a subsequent deploy would perform.
    3. C.`databricks bundle summary`, which reads the deployment state already recorded in the target workspace and prints resource links.
    4. D.`databricks bundle deploy --dry-run`, which the reviewer assumes walks through the full deployment process while skipping only the final write.
    Show answer & explanation

    Correct answer: A — `databricks bundle validate`, which checks the configuration files and prints a summary of the bundle's identity without deploying resources.

    • A. Correct. `bundle validate` checks that configuration files are syntactically correct and outputs the resolved bundle identity, and it never writes anything to the remote workspace.
    • B. `bundle plan` does build the bundle and preview planned actions, but it requires the direct deployment engine and reports pending changes against remote state rather than acting purely as a syntax check.
    • C. `bundle summary` depends on resources that have already been deployed; it reads back existing state and deep links, so it cannot validate configuration for a change that has not been deployed yet.
    • D. The `bundle deploy` command has no `--dry-run` flag; deploy either performs the deployment (subject to `--force`/`--auto-approve`) or fails outright, so this option does not exist.

    Subdomain 4.3: Implement development lifecycle processes in Azure Databricks

    28.Which statement best describes the scope of a unit test in a Databricks Asset Bundle project's testing strategy?

    1. A.It exercises a single function or transformation in isolation, using controlled inputs and asserting a specific expected output.
    2. B.It exercises the entire deployed bundle across every configured target to confirm each workspace produces identical results.
    3. C.It exercises the interaction between two or more deployed job tasks to confirm data passed between them stays correctly formatted.
    4. D.It exercises the finished pipeline against production-scale data so business stakeholders can accept the final report as delivered.
    Show answer & explanation

    Correct answer: A — It exercises a single function or transformation in isolation, using controlled inputs and asserting a specific expected output.

    • A. Correct. A unit test targets one function or transformation in isolation with controlled inputs, which keeps it fast and lets failures be attributed to a single piece of logic.
    • B. Running the same checks across every target describes a cross-environment validation approach, not the defining characteristic of a unit test, which is isolation of one component.
    • C. Checking how two or more tasks interact and hand off data describes integration testing, which sits above the unit level rather than testing a single function alone.
    • D. Running against production-scale data for stakeholder sign-off describes user acceptance testing, not a unit test, which stays small and isolated by design.

    Subdomain 4.3: Implement development lifecycle processes in Azure Databricks

    29.A release engineer wants a deployment to skip every interactive confirmation prompt so it can run unattended in a CI pipeline. They should add the ___ flag to the `databricks bundle deploy` command.

    1. A.`--auto-approve`
    2. B.`--fail-on-active-runs`
    3. C.`--force-lock`
    Show answer & explanation

    Correct answer: A — `--auto-approve`

    • A. Correct. This flag skips interactive approvals that might otherwise be required for deployment, which is exactly what an unattended pipeline needs.
    • B. This flag causes the deployment to fail if there are active runs, a safety check unrelated to suppressing interactive approval prompts.
    • C. This flag forces acquisition of the deployment lock left over from a crashed prior deployment; it does not suppress interactive approval prompts.

    Subdomain 4.4: Monitor, troubleshoot, and optimize workloads in Azure Databricks

    30.An engineer is about to repair a failed multi-task Lakeflow Jobs run and wants to understand what to expect. Select the statements that correctly describe repair run behavior.(Select 3)

    1. A.If the original run used a shared job cluster, the repair run provisions a new job cluster with a version suffix rather than reusing the exact original cluster instance.
    2. B.Repair is only available for jobs that orchestrate two or more tasks; a failed single-task job must instead be triggered again with Run now.
    3. C.Because Lakeflow Jobs does not make tasks idempotent by default, re-running a task that partially wrote output before failing can duplicate that data.
    4. D.Repair run always resets every task's parameters back to the values used in the very first attempt, ignoring any settings changed since then.
    5. E.Repair run permanently deletes the run history for the original failed attempt so only the repaired run's outcome remains visible afterward.
    Show answer & explanation

    Correct answers: A, B, C — If the original run used a shared job cluster, the repair run provisions a new job cluster with a version suffix rather than reusing the exact original cluster instance.; Repair is only available for jobs that orchestrate two or more tasks; a failed single-task job must instead be triggered again with Run now.; Because Lakeflow Jobs does not make tasks idempotent by default, re-running a task that partially wrote output before failing can duplicate that data.

    • A. A new job cluster with a version suffix is correct because when tasks share a job cluster, the repair run provisions a fresh cluster instance rather than reattaching to the original one, so the initial and repaired runs stay distinguishable.
    • B. Repair being limited to jobs with two or more tasks is correct because a single-task job has no dependent tasks to selectively rerun, so recovering it simply means triggering it again with Run now.
    • C. The idempotency caution is correct because Lakeflow Jobs reruns each unsuccessful task from the beginning without guaranteeing exactly-once writes, so a task that already wrote partial output before failing can produce duplicate data on repair.
    • D. Repair run does not force parameters back to the first attempt's values; it re-runs unsuccessful tasks with the current job and task settings, and any parameters entered in the repair dialog override those current values.
    • E. Repair run does not delete run history; the matrix view keeps the original failed attempt visible alongside the new repair column so both are still available to review.

    Subdomain 4.4: Monitor, troubleshoot, and optimize workloads in Azure Databricks

    31.While reviewing a long-running stage's Summary Metrics, an engineer should suspect skew when the Max task duration is at least ___ % higher than the 75th percentile duration for that stage.

    1. A.50
    2. B.5
    3. C.500
    Show answer & explanation

    Correct answer: A — 50

    • A. 50 is correct because the Spark UI guide specifically calls out that a Max duration at least 50% higher than the 75th percentile is the threshold at which skew becomes a reasonable suspicion.
    • B. 5 is incorrect because a gap that small is well within normal task-to-task variance and does not indicate the kind of imbalance the guide associates with skew.
    • C. 500 is incorrect because that far overstates the guide's threshold; requiring a five-times gap would miss skew that is already clearly affecting the stage's runtime at more moderate gaps.

    Subdomain 4.4: Monitor, troubleshoot, and optimize workloads in Azure Databricks

    32.A team runs OPTIMIZE followed by VACUUM on a large Delta table as part of a weekly maintenance job. Select the benefits this combination provides.(Select 3)

    1. A.OPTIMIZE reduces the number of small files a query must open, which lowers per-file read overhead and speeds up subsequent queries.
    2. B.VACUUM reclaims storage cost by physically deleting old, unreferenced data files once they fall outside the configured retention window.
    3. C.If predictive optimization is enabled for the table, Azure Databricks can trigger this same compaction and cleanup automatically without the weekly job.
    4. D.Running OPTIMIZE guarantees that every future query against the table will complete in constant time regardless of how much new data is later appended.
    5. E.Running VACUUM immediately deletes every version older than the current one, permanently disabling time travel to any past version of the table.
    Show answer & explanation

    Correct answers: A, B, C — OPTIMIZE reduces the number of small files a query must open, which lowers per-file read overhead and speeds up subsequent queries.; VACUUM reclaims storage cost by physically deleting old, unreferenced data files once they fall outside the configured retention window.; If predictive optimization is enabled for the table, Azure Databricks can trigger this same compaction and cleanup automatically without the weekly job.

    • A. Reduced per-file read overhead from OPTIMIZE is correct because compacting small files into larger ones means queries open fewer files, directly speeding up subsequent reads.
    • B. Reclaimed storage cost from VACUUM is correct because it physically deletes data files that are no longer referenced by the transaction log once they age past the retention window, freeing up storage that was otherwise billed.
    • C. Predictive optimization automating this maintenance is correct because when it is enabled for a table, Azure Databricks can trigger OPTIMIZE and VACUUM operations itself, reducing the need to schedule this weekly job manually.
    • D. Constant-time future queries is incorrect because OPTIMIZE compacts the files that exist at the time it runs; newly appended data after that point can reintroduce small files, so query time is not permanently fixed.
    • E. Immediately deleting every version older than the current one is incorrect because VACUUM respects the configured retention window, defaulting to seven days, rather than deleting all history the instant it runs.

    Subdomain 4.2: Implement Lakeflow Jobs

    33.A task is configured with both a ten-minute timeout and a retry policy allowing two additional attempts. The task's first attempt runs for fifteen minutes without finishing. Based on documented behavior, what happens next?

    1. A.The task is marked Timed Out at ten minutes, and each retry attempt gets its own ten-minute timeout
    2. B.The ten-minute timeout is ignored entirely because a retry policy has also been configured for the task
    3. C.The task keeps running past fifteen minutes because the timeout only applies to the very first attempt
    4. D.Retries are skipped completely once a timeout occurs, and the run is marked failed after one attempt
    Show answer & explanation

    Correct answer: A — The task is marked Timed Out at ten minutes, and each retry attempt gets its own ten-minute timeout

    • A. When both a timeout and retries are configured, the timeout is documented to apply to each retry individually, so the first attempt stops at ten minutes and every later retry gets its own fresh ten-minute window.
    • B. Configuring retries does not disable or override the timeout; the two settings are documented to work together rather than one canceling out the other.
    • C. Letting the task run indefinitely past its configured timeout would defeat the purpose of setting one, and the timeout is documented to apply to the first attempt just as it does to retries.
    • D. A timeout that ends an attempt counts as a failure like any other, and a configured retry policy still applies retries after a timeout-caused failure rather than skipping them outright.

    Subdomain 4.2: Implement Lakeflow Jobs

    34.An engineer is managing an existing trigger on a job that occasionally needs to be paused for scheduled maintenance. Based on documented behavior for managing an existing trigger, which of the following statements are accurate? Select all that apply.(Select 3)

    1. A.Clicking Pause stops the trigger from starting new runs, while any run already active keeps running to completion
    2. B.Clicking Resume causes the trigger to continue on the same schedule or event configuration it had before pausing
    3. C.If a run is still active when a continuous trigger is resumed, the scheduler waits for it to finish first
    4. D.Clicking Delete on a trigger immediately cancels and erases every past run that trigger ever started
    5. E.Editing a trigger's configuration is only possible while the trigger is sitting in the Paused state
    6. F.A single job can have a Scheduled trigger and a Continuous trigger both active on it at once
    Show answer & explanation

    Correct answers: A, B, C — Clicking Pause stops the trigger from starting new runs, while any run already active keeps running to completion; Clicking Resume causes the trigger to continue on the same schedule or event configuration it had before pausing; If a run is still active when a continuous trigger is resumed, the scheduler waits for it to finish first

    • A. Pausing an active trigger is documented to stop it from starting new runs going forward, while any run that is already active is allowed to continue running to completion.
    • B. Resuming a trigger is documented to continue the previously configured schedule or event behavior rather than resetting to some different default configuration.
    • C. If a run is still active when a continuous trigger is resumed, the job scheduler is documented to wait until that run completes before triggering a new run.
    • D. Deleting a trigger removes it from starting future runs; it does not retroactively cancel or erase runs that already occurred in the past under that trigger.
    • E. A trigger's configuration can be edited using Edit trigger regardless of whether the trigger is currently active or paused, so pausing first is not a prerequisite for editing it.
    • F. A job has a single trigger configuration at any given time, chosen from types such as Scheduled, Table update, File arrival, Model update, or Continuous, not several active simultaneously.

    Subdomain 4.2: Implement Lakeflow Jobs

    35.A platform team owns a nightly batch ETL job that must finish before six in the morning so downstream dashboards stay fresh. They want to know the moment something goes wrong, be warned if a run risks missing the deadline, and avoid alert floods when the job is intentionally skipped during a maintenance freeze. Which configurations should they set up? Select all that apply.(Select 3)

    1. A.A Failure notification so they are alerted whenever a run ends in an unsuccessful state
    2. B.A Duration warning tied to a configured expected completion time before the six-AM deadline
    3. C.Mute notifications for skipped runs so intentional maintenance-freeze skips stay quiet
    4. D.A Streaming backlog notification tied to a Kafka source backlog metric threshold
    5. E.A Model update trigger that starts a new run when a registered model version is ready
    6. F.Mute notifications until the last retry so no retry attempts are ever recorded in run history
    Show answer & explanation

    Correct answers: A, B, C — A Failure notification so they are alerted whenever a run ends in an unsuccessful state; A Duration warning tied to a configured expected completion time before the six-AM deadline; Mute notifications for skipped runs so intentional maintenance-freeze skips stay quiet

    • A. A failure notification is the direct mechanism for being alerted the moment a run ends unsuccessfully, matching the team's first stated requirement.
    • B. Configuring a duration warning against an expected completion time is how a team gets alerted when a run risks missing a hard deadline like six AM, without waiting for outright failure.
    • C. Muting notifications for skipped runs specifically suppresses alerts for runs skipped due to reasons like a maintenance freeze, directly addressing the team's concern about alert flooding.
    • D. A streaming backlog notification concerns backlog metrics on streaming sources; this nightly batch ETL scenario has no streaming backlog dimension described, so it does not address the stated requirements.
    • E. A model update trigger starts new runs based on registered model changes and has nothing to do with alerting on failures, durations, or skipped runs for an existing scheduled job.
    • F. Muting until the last retry only suppresses notifications for intermediate retry attempts; it does not touch run history visibility and does not address the team's stated failure, duration, or skip-noise goals.

    Want the full experience?

    These are just samples. Practice the full Microsoft Certified: Azure Databricks Data Engineer Associate (DP-750) question bank in quiz mode — free, no signup, with domain practice and exam simulation.