CertSafari

    Free Microsoft Certified: Azure Data Fundamentals (DP-900) Sample Questions

    35 free sample questions from our bank of 361+, covering every exam domain, with answers and detailed explanations. Updated September 2026.

    Domain 1: Describe core data concepts

    Subdomain 1.1: Describe ways to represent data

    1.Which file format stores a JSON-based header that describes the record schema, followed by the actual data as binary blocks of records?

    1. A.Avro
    2. B.Parquet
    3. C.Delta Lake
    4. D.Comma-separated values (CSV)
    Show answer & explanation

    Correct answer: A — Avro

    • A. Correct: this row-based format stores a JSON header describing the schema, then stores the records themselves as binary data in one or more blocks.
    • B. Incorrect: this format stores data by column within row groups and uses its own metadata structure, not a JSON header followed by row-based binary blocks.
    • C. Incorrect: this format adds a transaction log on top of a columnar format to enable ACID transactions, rather than storing a JSON schema header with row blocks.
    • D. Incorrect: this format stores plain text fields separated by commas and has no binary record blocks or JSON schema header.

    Subdomain 1.1: Describe ways to represent data

    2.A support team stores thousands of free-text customer chat transcripts with no predefined fields or schema. Which type of data are these transcripts?

    1. A.Unstructured data
    2. B.Structured data
    3. C.Semi-structured data
    4. D.Relational data
    Show answer & explanation

    Correct answer: A — Unstructured data

    • A. Correct: free-text transcripts with no predefined fields or schema are a classic example of unstructured data.
    • B. Incorrect: this term requires a fixed schema with consistent fields, which does not apply to free-form chat transcripts.
    • C. Incorrect: this term applies to data with some structure and field variation, such as JSON, not to unstructured free text.
    • D. Incorrect: this term describes data organized into related tables with key values, which does not apply to free-text transcripts.

    Subdomain 1.1: Describe ways to represent data

    3.A data platform team is choosing a file format for a new analytics data lake that will store billions of rows and needs fast column-level queries with compression. Which of the following statements about Parquet support this choice? (Select all that apply)(Select 3)

    1. A.Parquet is a columnar format, so a query that reads only a few columns can retrieve just the chunks it needs instead of scanning entire rows.
    2. B.Parquet files include metadata describing which rows are found in each chunk, letting applications locate and read the correct data quickly.
    3. C.Parquet specializes in efficiently storing and processing nested data types and supports strong compression and encoding schemes.
    4. D.Parquet is a plain-text, comma-delimited format, so every value must be parsed character by character before it can be compressed.
    5. E.Parquet was designed primarily for humans to read directly in a text editor rather than for applications to process at scale.
    Show answer & explanation

    Correct answers: A, B, C — Parquet is a columnar format, so a query that reads only a few columns can retrieve just the chunks it needs instead of scanning entire rows.; Parquet files include metadata describing which rows are found in each chunk, letting applications locate and read the correct data quickly.; Parquet specializes in efficiently storing and processing nested data types and supports strong compression and encoding schemes.

    • A. Correct: Parquet's columnar layout lets a query retrieve only the needed column chunks instead of scanning full rows, which speeds up analytical queries.
    • B. Correct: Parquet's chunk-level metadata lets applications locate and retrieve the correct data quickly without scanning the whole file.
    • C. Correct: Parquet's support for nested data types and strong compression and encoding schemes makes it well suited to large-scale analytics.
    • D. Incorrect: Parquet is a binary, columnar format rather than a plain-text, comma-delimited format like CSV.
    • E. Incorrect: Parquet is optimized for application processing at scale, not for direct human reading in a text editor.

    Subdomain 1.2: Identify options for data storage

    4.A gaming company needs a data store that can retrieve a player's session state in under a millisecond, looked up only by the player's unique session ID with no other query pattern required. Which type of database best fits this requirement?

    1. A.Key-value store
    2. B.Graph database
    3. C.Relational database
    4. D.Columnar database
    Show answer & explanation

    Correct answer: A — Key-value store

    • A. This is correct because storing a value against a single unique key, with no secondary queries needed, is exactly the access pattern this database type is optimized for.
    • B. A store optimized for traversing relationships between connected entities adds overhead this simple, single-key lookup does not need.
    • C. Enforcing a fixed schema and relationships across tables adds overhead that is unnecessary when the only access pattern is a single-key lookup.
    • D. A store built for scanning large volumes of column-oriented data is optimized for a different access pattern than a single sub-millisecond key lookup.

    Subdomain 1.2: Identify options for data storage

    5.An application's primary access pattern is traversing multi-hop relationships between users, such as finding friends of friends who share a common interest. Would a key-value store be the best database type for this access pattern?

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: B — False

    • A. This is incorrect — a store optimized for looking up a single value by one key has no built-in way to traverse chains of relationships between records efficiently.
    • B. This is correct — traversing multi-hop relationships is exactly what a graph database is purpose-built for, since it models entities as nodes and their connections as edges, unlike a key-value store.

    Subdomain 1.2: Identify options for data storage

    6.A columnar file format such as Parquet organizes data by column instead of by row, so an analytical query that only needs a few columns out of many can avoid reading the unused columns from disk. Is this statement correct?

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: A — True

    • A. This is correct — grouping values by column on disk is exactly what lets a query skip columns it does not reference, reducing the data scanned.
    • B. This is incorrect — this accurately describes how columnar formats reduce scanned data compared with row-based formats, which read every column of every row regardless of what the query needs.

    Subdomain 1.4: Identify roles and responsibilities for data workloads

    7.A company is hiring a data engineer to support its new analytics platform. Which of the following tasks would fall within this role's responsibilities? (Select all that apply)(Select 3)

    1. A.Building and scheduling pipelines that ingest data from source systems into a data lake.
    2. B.Cleansing and transforming raw data so it is ready for downstream analytical processing.
    3. C.Provisioning and configuring storage structures that support large-scale analytical workloads.
    4. D.Granting and revoking database user permissions to control access to sensitive tables.
    5. E.Publishing interactive Power BI dashboards that summarize key performance indicators.
    6. F.Applying database engine patches and monitoring server uptime for production systems.
    Show answer & explanation

    Correct answers: A, B, C — Building and scheduling pipelines that ingest data from source systems into a data lake.; Cleansing and transforming raw data so it is ready for downstream analytical processing.; Provisioning and configuring storage structures that support large-scale analytical workloads.

    • A. Building and scheduling ingestion pipelines that move data from source systems into a data lake is a core, hands-on part of this role's work.
    • B. Cleansing and transforming raw data so downstream systems and analysts can use it is a defining part of preparing data for analytics.
    • C. Provisioning and configuring storage structures sized for large analytical workloads is squarely part of building the analytics platform's foundation.
    • D. Granting and revoking user permissions on sensitive tables is database administration work focused on security, not pipeline or storage design.
    • E. Publishing dashboards that summarize key performance indicators is analyst work that happens after the data has already been prepared.
    • F. Applying engine patches and monitoring server uptime is operational database administration, separate from building ingestion and transformation pipelines.

    Subdomain 1.4: Identify roles and responsibilities for data workloads

    8.A new hire joins the database operations team supporting the company's production systems. Which of the following tasks are commonly performed by a database administrator? (Select all that apply)(Select 3)

    1. A.Applying security patches and monitoring uptime for production database servers.
    2. B.Configuring backup schedules and testing disaster recovery restoration procedures.
    3. C.Tuning query performance by adjusting indexes and reviewing query execution plans.
    4. D.Building Power BI dashboards that visualize monthly sales performance.
    5. E.Designing ETL pipelines that transform raw data for analytical workloads.
    Show answer & explanation

    Correct answers: A, B, C — Applying security patches and monitoring uptime for production database servers.; Configuring backup schedules and testing disaster recovery restoration procedures.; Tuning query performance by adjusting indexes and reviewing query execution plans.

    • A. Applying security patches and watching server uptime is routine operational work for keeping production databases healthy and secure.
    • B. Configuring backup schedules and confirming restores actually succeed is a core disaster-recovery duty owned by this role.
    • C. Adjusting indexes and reviewing execution plans to speed up slow queries is standard performance-tuning work this role performs.
    • D. Building Power BI dashboards to visualize sales performance is analyst work focused on business reporting, not database operations.
    • E. Designing ETL pipelines that transform raw data is data engineering work focused on data preparation, separate from database administration.

    Subdomain 1.4: Identify roles and responsibilities for data workloads

    9.A company assigns its Power BI report developer the task of designing and scheduling the nightly data ingestion pipeline that pulls raw files into the data lake. Is this a typical data analyst responsibility?

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: B — False

    • A. Designing and scheduling an ingestion pipeline is not typical analyst work; it is a data engineering task focused on moving raw data into storage before it is analysis-ready.
    • B. This task belongs to a data engineer, who builds and schedules the pipelines that ingest raw files, while a data analyst works with the data once it has already been prepared.

    Subdomain 1.3: Describe common data workloads

    10.A bank's mobile app initiates a transfer of $40 from Account A to Account B. The system must debit Account A and credit Account B, and if either step cannot complete, neither change should be saved. Which ACID property is this transaction relying on?

    1. A.Consistency
    2. B.Atomicity
    3. C.Isolation
    4. D.Durability
    Show answer & explanation

    Correct answer: B — Atomicity

    • A. Consistency describes a transaction moving the database from one valid state to another, not the all-or-nothing guarantee this scenario needs.
    • B. This guarantee treats the debit and credit as a single unit of work that either completes fully or has no effect at all, which matches the requirement exactly.
    • C. Isolation is about concurrent transactions not interfering with each other, which is a different concern from whether one transaction's steps all succeed or all fail.
    • D. Durability guarantees a committed transaction survives a later system failure, not that its individual steps succeed or fail as a unit.

    Subdomain 1.3: Describe common data workloads

    11.Scenario: Your finance team needs to run heavy trend-aggregation queries over three years of historical sales data every night. To meet this need, you point their reporting queries directly at the live order-processing database used for point-of-sale transactions, with no separate analytical store. Does this design meet the requirement without risking contention with live transactions?

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: B — False

    • A. This is not correct: running heavy aggregation queries directly against a live transactional database risks contending with ongoing purchase transactions, and that database is not optimized for this kind of historical aggregation.
    • B. Running the trend-aggregation queries directly against the live transactional database risks resource contention with ongoing sales transactions, and that database is optimized for small record-level operations rather than large historical aggregations, so a separate analytical store is the better fit.

    Subdomain 1.3: Describe common data workloads

    12.Which of the following best describes an OLTP system?

    1. A.A system optimized for read-mostly queries over large historical datasets aggregated for reporting.
    2. B.A system optimized for high volumes of read and write transactions that satisfy ACID guarantees.
    3. C.A system that stores raw, unstructured files for later exploration by data scientists.
    4. D.A system that only supports scheduled batch loads of data with no live transaction support.
    Show answer & explanation

    Correct answer: B — A system optimized for high volumes of read and write transactions that satisfy ACID guarantees.

    • A. Read-mostly querying over large historical datasets describes an analytical workload, not the create-retrieve-update-delete pattern an OLTP system is built for.
    • B. This is the defining trait of online transaction processing: high-volume, small, discrete transactions that must satisfy atomicity, consistency, isolation, and durability.
    • C. Storing raw unstructured files for exploratory work describes an early data-lake stage of an analytics pipeline, not a transactional system.
    • D. OLTP systems are built for live, immediate transaction processing, not restricted to scheduled batch loads with no live activity.

    Domain 2: Identify considerations for relational data on Azure

    Subdomain 2.1: Describe relational concepts

    13.A developer wants a reusable piece of SQL logic that accepts an order's subtotal and tax rate as parameters and returns the calculated total, so it can be called directly within a SELECT statement's column list. Which database object fits this purpose?

    1. A.A trigger, which runs automatically in response to insert, update, or delete events
    2. B.A stored procedure, which is executed as a standalone call and cannot be embedded in a SELECT
    3. C.An index, which organizes column values to speed up filtering and sorting operations
    4. D.A function, which accepts parameters and returns a computed value usable inside a query
    Show answer & explanation

    Correct answer: D — A function, which accepts parameters and returns a computed value usable inside a query

    • A. A trigger fires automatically on data change events, it is not called directly within a SELECT statement's column list to return a value.
    • B. A stored procedure is typically executed with a separate call and, unlike a function, generally cannot be embedded inside a SELECT statement's expression.
    • C. An index only affects how quickly rows are located or sorted, it does not accept parameters or return a computed value.
    • D. This is correct - a function takes input parameters, computes a value, and returns it, which allows it to be embedded directly inside a SELECT statement's column list.

    Subdomain 2.1: Describe relational concepts

    14.Which of the following SQL statements belong to the Data Manipulation Language (DML) category? (Select all that apply.)(Select 3)

    1. A.SELECT, which retrieves rows and columns that match specified conditions
    2. B.INSERT, which adds one or more new rows of data into a table
    3. C.UPDATE, which changes the values of existing rows that match a condition
    4. D.GRANT, which gives a user or role permission to access a database object
    5. E.CREATE TABLE, which defines the structure of a new table
    Show answer & explanation

    Correct answers: A, B, C — SELECT, which retrieves rows and columns that match specified conditions; INSERT, which adds one or more new rows of data into a table; UPDATE, which changes the values of existing rows that match a condition

    • A. Correct - SELECT retrieves and filters existing data, which is a core Data Manipulation Language operation.
    • B. Correct - INSERT adds new rows of data into an existing table, working with the data rather than the structure.
    • C. Correct - UPDATE changes the values already stored in existing rows, another Data Manipulation Language operation.
    • D. Incorrect - GRANT assigns access permissions to a user or role, which places it in the Data Control Language category rather than DML.
    • E. Incorrect - CREATE TABLE defines a new table's structure, which is a Data Definition Language operation, not data manipulation.

    Subdomain 2.1: Describe relational concepts

    15.Which of the following are benefits of normalizing a relational database schema? (Select all that apply.)(Select 3)

    1. A.Reduces redundant storage of the same data across multiple rows
    2. B.Lowers the risk of update anomalies where a change must be repeated in many places
    3. C.Improves overall data integrity by organizing related data into separate tables
    4. D.Eliminates the need to ever define primary or foreign keys in the schema
    5. E.Guarantees that every query will run faster without needing any indexes
    Show answer & explanation

    Correct answers: A, B, C — Reduces redundant storage of the same data across multiple rows; Lowers the risk of update anomalies where a change must be repeated in many places; Improves overall data integrity by organizing related data into separate tables

    • A. Correct - normalization organizes data so the same fact is not stored redundantly across many rows, reducing wasted storage and inconsistency.
    • B. Correct - by removing redundant copies, normalization lowers the risk that an update to one fact is missed in some of its duplicate copies.
    • C. Correct - splitting data into related tables based on their dependencies improves overall data integrity across the schema.
    • D. Incorrect - normalization actually relies on primary and foreign keys to define the relationships between the newly separated tables.
    • E. Incorrect - normalization does not guarantee faster queries by itself, and highly normalized schemas can require more joins that may need indexes to perform well.

    Subdomain 2.2: Describe relational Azure data services

    16.A healthcare company must keep its SQL Server workload fully isolated inside its own virtual network with only a private IP address, while still avoiding management of patching and backups. Which Azure data service satisfies both requirements?

    1. A.Azure SQL Managed Instance
    2. B.Azure SQL Database
    3. C.SQL Server on Azure Virtual Machines
    4. D.Azure Database for MySQL flexible server
    Show answer & explanation

    Correct answer: A — Azure SQL Managed Instance

    • A. Managed Instance has a native virtual network implementation with a private IP endpoint by default and single-tenant isolated infrastructure, all while remaining a fully managed PaaS service.
    • B. Azure SQL Database can be restricted with private endpoints, but it does not provide the same instance-scoped, single-tenant VNet isolation model that Managed Instance is built around.
    • C. This IaaS option can also run inside a private VNet, but the customer must patch the OS and manage backups themselves, which the company wants to avoid.
    • D. This service runs the MySQL engine, so it cannot host a SQL Server workload at all regardless of its networking configuration.

    Subdomain 2.2: Describe relational Azure data services

    17.A company needs to run an application that requires linked servers, CLR integration, and global temporal tables against a fully managed instance, without administering the underlying operating system. Which Azure data service supports all of these features together?

    1. A.Azure SQL Managed Instance
    2. B.Azure SQL Database
    3. C.Azure Database for MariaDB
    4. D.Azure Database for PostgreSQL flexible server
    Show answer & explanation

    Correct answer: A — Azure SQL Managed Instance

    • A. Managed Instance's near-100% Database Engine compatibility list includes linked servers, CLR modules, and global temporal tables, all delivered as a fully managed PaaS service.
    • B. Azure SQL Database does not support linked servers or CLR integration in the same way Managed Instance does, since it targets a more modern, simplified single-database model.
    • C. This service runs the MariaDB engine, which has no equivalent to SQL Server linked servers, CLR integration, or temporal tables.
    • D. This service runs PostgreSQL, which does not provide SQL Server-specific features such as linked servers, CLR modules, or temporal tables.

    Subdomain 2.2: Describe relational Azure data services

    18.Which of the following are open-source relational database engines offered as fully managed Azure Database services? (Select all that apply.)(Select 3)

    1. A.PostgreSQL
    2. B.MySQL
    3. C.Oracle Database
    4. D.IBM Db2
    5. E.MariaDB
    Show answer & explanation

    Correct answers: A, B, E — PostgreSQL; MySQL; MariaDB

    • A. Azure Database for PostgreSQL flexible server is a fully managed offering built on the open-source PostgreSQL community engine.
    • B. Azure Database for MySQL flexible server is a fully managed offering built on the open-source MySQL community engine.
    • C. Oracle Database is a commercial, closed-source engine and is not offered as an Azure Database open-source managed service; it is typically run on a virtual machine instead.
    • D. IBM Db2 is a commercial engine and is not one of the Azure Database open-source managed offerings.
    • E. Azure Database for MariaDB is also an open-source engine offering in the Azure Database family, built on the community MariaDB fork of MySQL.

    Subdomain 2.2: Describe relational Azure data services

    19.Which of the following are valid Azure SQL Database purchasing or compute options? (Select all that apply.)(Select 4)

    1. A.vCore-based purchasing model
    2. B.DTU-based purchasing model
    3. C.Serverless compute tier
    4. D.Cassandra API
    5. E.Provisioned compute tier
    Show answer & explanation

    Correct answers: A, B, C, E — vCore-based purchasing model; DTU-based purchasing model; Serverless compute tier; Provisioned compute tier

    • A. The vCore-based purchasing model lets you choose vCores, memory, and storage independently and supports Azure Hybrid Benefit.
    • B. The DTU-based purchasing model blends compute, memory, and I/O into fixed tiers for lighter configuration effort.
    • C. The serverless compute tier automatically scales compute based on workload activity and bills per second of use.
    • D. The Cassandra API is a Cosmos DB API for wide-column workloads and is unrelated to Azure SQL Database purchasing options.
    • E. The provisioned compute tier allocates a fixed amount of compute continuously and bills at a fixed hourly price.

    Domain 3: Describe considerations for working with non-relational data on Azure

    Subdomain 3.2: Describe the capabilities and features of Azure Cosmos DB

    20.Which Azure Cosmos DB API is the native, original API that most closely matches the underlying database engine's own data and query model?

    1. A.The API for NoSQL
    2. B.The API for MongoDB
    3. C.The API for Cassandra
    4. D.The API for Gremlin
    Show answer & explanation

    Correct answer: A — The API for NoSQL

    • A. The API for NoSQL is Cosmos DB's native interface, built directly on its document engine rather than emulating another database.
    • B. The MongoDB API is a compatibility layer that translates MongoDB wire protocol calls onto the underlying engine.
    • C. The Cassandra API is a compatibility layer for CQL, not the engine's native interface.
    • D. The Gremlin API is a compatibility layer for graph traversal, not the engine's native interface.

    Subdomain 3.2: Describe the capabilities and features of Azure Cosmos DB

    21.A new internal tool receives only a handful of requests per hour during a pilot phase, and the team wants to avoid paying for reserved throughput capacity it won't use. Which Cosmos DB throughput mode fits this pilot best?

    1. A.Serverless throughput, which charges per request consumed rather than for pre-provisioned capacity.
    2. B.Manually provisioned throughput fixed at the maximum RU/s the tool might ever need in production.
    3. C.Autoscale throughput configured at its highest scale tier to guarantee headroom during the pilot.
    4. D.Reserved capacity purchased for a one-year term to lock in a discounted throughput rate.
    Show answer & explanation

    Correct answer: A — Serverless throughput, which charges per request consumed rather than for pre-provisioned capacity.

    • A. Serverless billing charges only for the request units consumed, matching a low, sporadic pilot workload without reserving capacity.
    • B. Provisioning for a hypothetical future production maximum would pay for far more capacity than a low-traffic pilot actually uses.
    • C. Autoscale still bills for a minimum reserved RU/s band, so setting it to the highest tier wastes money on an idle pilot.
    • D. A one-year reserved capacity commitment does not fit a short pilot with uncertain, low traffic.

    Subdomain 3.2: Describe the capabilities and features of Azure Cosmos DB

    22.Azure Cosmos DB can scale throughput and storage independently and elastically across Azure regions. Does this meet the goal?

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: A — True

    • A. Storage and throughput in Cosmos DB scale independently per region, expanding or contracting elastically as demand changes.
    • B. Elastic, independent scaling of storage and throughput is a core, published capability of Cosmos DB, so this statement is accurate.

    Subdomain 3.2: Describe the capabilities and features of Azure Cosmos DB

    23.The API for NoSQL in Azure Cosmos DB is purpose-built for graph traversal queries over vertices and edges. Does this meet the goal?

    1. A.True
    2. B.False
    Show answer & explanation

    Correct answer: B — False

    • A. The API for NoSQL uses a document model queried with SQL-like syntax; graph traversal over vertices and edges is the Gremlin API's role, not this API's.
    • B. Graph traversal over vertices and edges is handled by the Gremlin API, not the API for NoSQL, which serves document data instead, so this statement is not accurate.

    Subdomain 3.1: Describe the capabilities of Azure storage

    24.A company keeps its primary file share in an Azure Files account so branch-office file servers can stop growing local disks, but each branch still needs the files to open with local-disk speed and to keep working briefly if the WAN link drops. Which approach meets both requirements?

    1. A.Deploy Azure File Sync so each branch server caches the Azure file share locally and continues serving cached files during short WAN outages.
    2. B.Move the shared files into Azure Blob storage and have each branch server download the full dataset nightly through a scheduled AzCopy job.
    3. C.Replicate the file share into an Azure Table storage account so branch offices query file contents as key/attribute entities instead of files.
    4. D.Configure geo-redundant storage on the file share alone, without any branch-side component, and have branch users connect directly over the WAN each time.
    Show answer & explanation

    Correct answer: A — Deploy Azure File Sync so each branch server caches the Azure file share locally and continues serving cached files during short WAN outages.

    • A. Azure File Sync installs an agent on the branch server that caches the Azure file share locally, giving local-disk read speed and letting cached files stay available briefly if the WAN connection drops.
    • B. A nightly full download to Blob storage does not keep files continuously current, does not give a native file-share interface for local apps, and does not survive a WAN outage that happens between downloads.
    • C. Table storage stores discrete key/attribute entities rather than file content, so moving the share there would break normal file access instead of speeding it up during a WAN outage.
    • D. Geo-redundant storage protects the data against a datacenter outage, but it does nothing for branch-side read latency or continued access when the branch itself loses its WAN connection.

    Subdomain 3.1: Describe the capabilities of Azure storage

    25.A disaster-recovery plan for a heavily used Azure file share requires the data to survive an entire datacenter outage while also protecting against a single zone failure within the primary region. Which two Azure Files redundancy options should the team evaluate? (Select 2)(Select 2)

    1. A.Locally redundant storage (LRS)
    2. B.Zone-redundant storage (ZRS)
    3. C.Geo-redundant storage (GRS)
    4. D.Archive redundancy
    5. E.Partition redundancy
    Show answer & explanation

    Correct answers: B, C — Zone-redundant storage (ZRS); Geo-redundant storage (GRS)

    • A. Locally redundant storage only replicates copies within a single datacenter, so it does not protect against a zone failure or a full datacenter outage.
    • B. Zone-redundant storage replicates the file share synchronously across availability zones in the primary region, protecting against a single zone failure.
    • C. Geo-redundant storage replicates data to a secondary, geographically distant region, protecting the file share even if the entire primary datacenter is lost.
    • D. Archive redundancy is not a real Azure Files redundancy option; the archive access tier is unrelated to how many physical copies of data are kept.
    • E. Partition redundancy is not an Azure Storage term; partition key is a Table storage concept for grouping entities, not a redundancy setting.

    Subdomain 3.1: Describe the capabilities of Azure storage

    26.An operations team is choosing Azure Files instead of Blob storage for a migration project. Which three of the following are accurate reasons to prefer Azure Files for this project? (Select 3)(Select 3)

    1. A.The application is a lift-and-shift of an on-premises file server that still expects a mapped drive letter
    2. B.Client machines already mount the share using the SMB or NFS protocol and need that to keep working
    3. C.Multiple VMs must read shared application configuration files through normal file-system I/O calls
    4. D.The workload stores petabytes of unstructured images that are served directly to a public website
    5. E.The application performs fast key/attribute lookups that are filtered by a partition key value
    6. F.The workload feeds a Spark cluster that needs a hierarchical namespace for big-data analytics
    Show answer & explanation

    Correct answers: A, B, C — The application is a lift-and-shift of an on-premises file server that still expects a mapped drive letter; Client machines already mount the share using the SMB or NFS protocol and need that to keep working; Multiple VMs must read shared application configuration files through normal file-system I/O calls

    • A. A lift-and-shift file server that still expects a mapped drive letter is a primary Azure Files scenario, since applications can keep working without code changes.
    • B. Client machines that already mount shares over SMB or NFS can keep doing so against Azure Files, since it supports both protocols natively.
    • C. Letting multiple VMs read shared configuration files through normal file-system I/O calls is exactly what a centralized Azure file share provides.
    • D. Serving petabytes of unstructured images to a public website is a Blob storage scenario, since Blob storage's object model is optimized for that access pattern, not file shares.
    • E. Fast key/attribute lookups filtered by a partition key value describe Azure Table storage's data model, not the file-share model Azure Files provides.
    • F. A hierarchical namespace for big-data analytics describes Azure Data Lake Storage built on top of Blob storage, not Azure Files.

    Domain 4: Describe an analytics workload

    Subdomain 4.1: Describe common elements of large-scale analytics

    27.A logistics company wants to detect a delivery truck's temperature sensor going out of range within seconds of the reading arriving, so a driver can be alerted immediately. Which processing approach fits this requirement?

    1. A.Streaming analytics
    2. B.Nightly batch analytics
    3. C.Weekly aggregation jobs
    4. D.Manual spreadsheet review
    Show answer & explanation

    Correct answer: A — Streaming analytics

    • A. Correct: streaming analytics continuously ingests and evaluates events as they occur, giving the near-instant detection the alert scenario needs.
    • B. Incorrect: a nightly batch job would only surface the out-of-range reading hours later, missing the requirement for a within-seconds alert.
    • C. Incorrect: weekly aggregation summarizes data over a long window and is far too infrequent for a real-time driver alert.
    • D. Incorrect: manual spreadsheet review depends on a person opening a file and cannot react to sensor readings within seconds.

    Subdomain 4.1: Describe common elements of large-scale analytics

    28.A team loads raw sales data directly into a cloud data warehouse and then runs SQL transformation scripts inside the warehouse to clean and reshape it for reporting. Which pipeline pattern are they using?

    1. A.ELT (extract, load, transform)
    2. B.ETL (extract, transform, load)
    3. C.OLTP replication
    4. D.Change data capture only
    Show answer & explanation

    Correct answer: A — ELT (extract, load, transform)

    • A. Correct: loading raw data first and transforming it afterward using the target warehouse's own compute is the defining trait of ELT.
    • B. Incorrect: ETL transforms the data on separate compute before it is loaded, which is the opposite order from running SQL inside the warehouse after loading.
    • C. Incorrect: OLTP replication copies transactional database state for availability or reporting, not a warehouse transformation workflow.
    • D. Incorrect: change data capture only tracks and streams row-level changes from a source, it does not describe the load-then-transform pattern here.

    Subdomain 4.1: Describe common elements of large-scale analytics

    29.Which of the following are workloads included within Microsoft Fabric's unified SaaS platform? Select all that apply.(Select 3)

    1. A.Fabric Data Factory
    2. B.Fabric Data Engineering
    3. C.Power BI
    4. D.On-premises AD domain controller
    5. E.Self-managed SQL Server VM
    Show answer & explanation

    Correct answers: A, B, C — Fabric Data Factory; Fabric Data Engineering; Power BI

    • A. Correct: Data Factory in Fabric provides a data integration experience with connectors for ingesting and preparing data from many sources.
    • B. Correct: Fabric Data Engineering provides Apache Spark for processing and transforming large datasets with notebooks and scheduled jobs.
    • C. Correct: Power BI is included as Fabric's workload for connecting to data, building reports, and sharing dashboards.
    • D. Incorrect: an on-premises Active Directory domain controller is infrastructure a customer runs themselves, it is not a Fabric workload.
    • E. Incorrect: a self-managed SQL Server virtual machine is customer-run IaaS infrastructure, it is not part of the Fabric SaaS platform.

    Subdomain 4.2: Describe considerations for real-time data analytics

    30.A telecommunications provider ingests call detail records continuously through Event Hubs and uses Stream Analytics to compute rolling network usage statistics, then sends the results to Power BI for a live operations dashboard. Which two components of this pipeline are correctly matched to their role? (Select 2)(Select 2)

    1. A.Event Hubs receives the continuous stream of call detail records generated by network elements in real time
    2. B.Stream Analytics applies a windowed SQL-like query to compute rolling usage statistics from ingested events
    3. C.Power BI ingests the raw call detail records directly from network elements before any processing occurs
    4. D.Event Hubs computes the rolling network usage statistics before forwarding them to Stream Analytics
    5. E.Stream Analytics permanently stores every raw call detail record as its primary function
    Show answer & explanation

    Correct answers: A, B — Event Hubs receives the continuous stream of call detail records generated by network elements in real time; Stream Analytics applies a windowed SQL-like query to compute rolling usage statistics from ingested events

    • A. Correct: Event Hubs is the ingestion layer designed to absorb a continuous, high-volume stream of events like call detail records from many network elements at once.
    • B. Correct: Stream Analytics is the processing layer that runs the windowed SQL-like query against the ingested stream, producing the rolling usage statistics described in the scenario.
    • C. Incorrect: in this pipeline Power BI receives already-processed output from Stream Analytics for visualization; it is not the component that ingests raw records directly from network elements.
    • D. Incorrect: Event Hubs is an ingestion and buffering service, not a query engine; it does not compute rolling aggregates itself, that responsibility belongs to Stream Analytics.
    • E. Incorrect: Stream Analytics processes data in-memory as it streams through and does not store the incoming data; long-term storage of raw records would go to a service such as Blob Storage or Data Lake Storage.

    Subdomain 4.2: Describe considerations for real-time data analytics

    31.Which of the following best describes Microsoft Fabric's role in real-time analytics, as distinct from a service that only supports batch pipelines?

    1. A.Fabric includes a Real-Time Intelligence experience for ingesting and visualizing continuously arriving event data
    2. B.Fabric can only load data on a fixed nightly schedule and has no capability to process data as it arrives
    3. C.Fabric replaces the need for any data ingestion service entirely because it generates its own event data internally
    4. D.Fabric is limited to relational data warehousing only and cannot connect to any streaming data source at all
    Show answer & explanation

    Correct answer: A — Fabric includes a Real-Time Intelligence experience for ingesting and visualizing continuously arriving event data

    • A. Correct: Microsoft Fabric offers a dedicated Real-Time Intelligence experience that handles ingesting, analyzing, and visualizing event data as it streams in, giving it a genuine real-time analytics capability.
    • B. Incorrect: Fabric supports both batch pipelines and continuous streaming ingestion through Real-Time Intelligence, so describing it as limited to a fixed nightly schedule misrepresents its capabilities.
    • C. Incorrect: Fabric consumes event data from external producers such as applications, sensors, or Event Hubs; it does not generate its own data internally in place of an ingestion source.
    • D. Incorrect: Fabric's unified platform spans a data warehouse, a lakehouse, and Real-Time Intelligence, so it is not restricted to relational warehousing and does connect to streaming sources.

    Subdomain 4.2: Describe considerations for real-time data analytics

    32.A financial services firm wants to monitor stock trades in real time to detect suspicious trading patterns and also wants to keep a permanent, queryable archive of every trade for a compliance audit that runs once per quarter. Which architecture best satisfies both needs?

    1. A.Ingest trades with Event Hubs, detect patterns with Stream Analytics, and use Capture to copy records to Data Lake Storage for the quarterly audit
    2. B.Ingest trades with Event Hubs and analyze them only with Stream Analytics, discarding all raw data immediately after each query result gets produced
    3. C.Store all trades directly within Power BI Desktop and run the full compliance audit by manually scrolling through the report each quarter
    4. D.Wait until the end of each quarter to load all trades into Stream Analytics in one large batch for both detection and archival purposes
    Show answer & explanation

    Correct answer: A — Ingest trades with Event Hubs, detect patterns with Stream Analytics, and use Capture to copy records to Data Lake Storage for the quarterly audit

    • A. Correct: this architecture uses Event Hubs and Stream Analytics for the immediate pattern detection requirement, while Capture simultaneously persists a durable copy to Data Lake Storage that the quarterly batch audit can query later, satisfying both latency needs with one pipeline.
    • B. Incorrect: Stream Analytics does not persist incoming data since processing happens in-memory, so discarding everything after each query would leave nothing for the quarterly compliance audit to review.
    • C. Incorrect: Power BI Desktop is a report authoring tool, not a durable storage or query system for archiving every individual trade, and manually scrolling through a report is not a workable audit process.
    • D. Incorrect: Stream Analytics is designed for continuous low-latency processing of data in motion, not for loading a full quarter of accumulated trades in one large batch, and this approach would also miss the real-time detection requirement entirely.

    Subdomain 4.3: Describe data visualization in Microsoft Power BI

    33.A sales manager wants to check the latest regional revenue dashboard while walking through a warehouse, using only a smartphone and no laptop nearby. Which Power BI component fits this need?

    1. A.Power BI Mobile
    2. B.Power BI Desktop
    3. C.Power BI Report Builder
    4. D.Power Query Editor
    Show answer & explanation

    Correct answer: A — Power BI Mobile

    • A. Correct: this app is optimized for phones and tablets, letting users view and interact with already-published dashboards while away from a computer, matching the warehouse scenario.
    • B. Incorrect: this Windows desktop application is designed for authoring data models and reports on a PC, not for lightweight on-the-go viewing on a phone.
    • C. Incorrect: this tool creates pixel-perfect paginated reports such as invoices and is not the app used for casual dashboard viewing on a mobile device.
    • D. Incorrect: this component transforms and shapes source data inside Power BI Desktop; it has no dashboard-viewing interface for mobile users.

    Subdomain 4.3: Describe data visualization in Microsoft Power BI

    34.A finance department needs to generate a strictly formatted, printable invoice from Power BI, where every field lands in an exact position on the page for compliance reasons. Which Power BI feature is designed for this?

    1. A.Paginated reports built in Power BI Report Builder for print
    2. B.A standard interactive report built in Power BI Desktop pages
    3. C.A Power BI dashboard assembled from tiles pinned by a user
    4. D.A Q&A visual that auto-generates charts from typed questions
    Show answer & explanation

    Correct answer: A — Paginated reports built in Power BI Report Builder for print

    • A. Correct: paginated reports are built for pixel-perfect, print-ready layouts such as invoices, where every field must appear in an exact fixed position across pages.
    • B. Incorrect: interactive reports are designed for on-screen exploration with resizable, dynamic visuals, not for guaranteeing an exact printed layout.
    • C. Incorrect: a dashboard is a single-page collection of pinned tiles summarizing multiple reports; it is not built for exact, page-by-page printable formatting.
    • D. Incorrect: this feature lets users type questions and get auto-generated visuals; it does not produce a fixed, compliance-ready printed layout.

    Subdomain 4.3: Describe data visualization in Microsoft Power BI

    35.Which of the following are core features available in Power BI Desktop when building a report? (Select all that apply.)(Select 3)

    1. A.Connecting to more than 100 different data sources
    2. B.Creating data models with relationships and DAX calculations
    3. C.Access to more than 30 built-in visualization types
    4. D.Setting up Azure Active Directory password policies org-wide
    5. E.Provisioning and managing new Azure virtual machine instances
    Show answer & explanation

    Correct answers: A, B, C — Connecting to more than 100 different data sources; Creating data models with relationships and DAX calculations; Access to more than 30 built-in visualization types

    • A. Correct: Power BI Desktop can connect to well over 100 data sources spanning databases, files, cloud services, and web sources.
    • B. Correct: Desktop lets authors define relationships between tables and write DAX calculated columns and measures as part of the data model.
    • C. Correct: the Visualizations pane in Desktop ships with more than 30 built-in visual types, with more available from AppSource.
    • D. Incorrect: tenant-wide identity password policy configuration is an identity administration task handled outside Power BI, not a Desktop report-authoring feature.
    • E. Incorrect: provisioning virtual machine infrastructure is an Azure compute administration task unrelated to Power BI Desktop's report and data modeling capabilities.

    Want the full experience?

    These are just samples. Practice the full Microsoft Certified: Azure Data Fundamentals (DP-900) question bank in quiz mode — free, no signup, with domain practice and exam simulation.