CertSafari

    Free CompTIA Data+ Sample Questions

    35 free sample questions from our bank of 340+, covering every exam domain, with answers and detailed explanations. Updated September 2026.

    Domain 1: Data Concepts and Environments

    Subdomain 1.4: Identify common data analysis tools.

    1.A marketing analyst wants to drag and drop fields onto a canvas to build ad hoc exploratory visuals, using a tool widely known for its large library of connectors and its polished chart gallery for quick visual exploration. Which BI tool matches this description?

    1. A.Power BI
    2. B.Looker
    3. C.Anaconda
    4. D.Tableau
    Show answer & explanation

    Correct answer: D — Tableau

    • A. Power BI is a strong drag-and-drop BI tool, but it is best known for deep Microsoft 365/Azure integration rather than the connector breadth and chart-gallery reputation described here.
    • B. Looker centers on a governed semantic modeling layer (LookML) for consistent metrics rather than freeform drag-and-drop chart building, so it does not match this description.
    • C. Anaconda is a Python/R distribution and package manager, not a business intelligence visualization tool.
    • D. Tableau is widely recognized for its broad connector library and drag-and-drop chart gallery aimed at fast, ad hoc visual exploration, matching the scenario.

    Subdomain 1.4: Identify common data analysis tools.

    2.Which of the following are primarily GUI tools for managing and querying database management systems (DBMS), rather than programming languages or BI platforms? (Select 3)(Select 3)

    1. A.DBeaver
    2. B.Toad
    3. C.MySQL Workbench
    4. D.tidyverse
    5. E.Looker
    6. F.Scala
    Show answer & explanation

    Correct answers: A, B, C — DBeaver; Toad; MySQL Workbench

    • A. DBeaver is a generic GUI client for connecting to, browsing, and querying a wide range of relational and non-relational databases, making it a DBMS management tool.
    • B. Toad is a commercial GUI for writing queries, tuning performance, and administering database schemas, making it a DBMS management tool.
    • C. MySQL Workbench is Oracle's official GUI for designing, querying, and administering MySQL databases, making it a DBMS management tool.
    • D. tidyverse is a collection of R packages for data wrangling and visualization, not a database management GUI.
    • E. Looker is a business intelligence platform for building governed dashboards, not a database administration GUI.
    • F. Scala is a general-purpose programming language, not a GUI tool for managing databases.

    Subdomain 1.5: Identify artificial intelligence (AI) concepts.

    3.A support team feeds incoming customer emails into a tool that writes a full draft reply for a human agent to review and send. Which AI concept best describes what the tool is doing?

    1. A.Generative AI
    2. B.Robotic process automation
    3. C.Deep learning image classification
    4. D.Traditional rule-based scripting
    Show answer & explanation

    Correct answer: A — Generative AI

    • A. Generative AI is correct because the tool is producing new, original text tailored to each email rather than following a fixed template. Drafting a full reply from scratch is exactly the kind of content-creation task this category of AI performs.
    • B. This is incorrect because robotic process automation automates repetitive clicks and data transfers between systems; it does not compose original written language in response to varied input.
    • C. This is incorrect because image classification categorizes visual inputs like photos, and the scenario involves generating written text, not analyzing images.
    • D. This is incorrect because a fixed set of hand-coded rules cannot produce a unique, context-appropriate reply for every different email; that requires a model that generates language dynamically.

    Subdomain 1.5: Identify artificial intelligence (AI) concepts.

    4.A marketing team is evaluating a generative AI tool to help draft blog post outlines. Which of the following are accurate considerations about generative AI behavior that the team should weigh before adopting it? (Select all that apply.)(Select 3)

    1. A.The tool can produce fluent, human-like text without a person writing each sentence from scratch
    2. B.The tool can occasionally generate fabricated facts or citations that sound plausible but are not accurate
    3. C.The tool requires a labeled training example for every possible topic before it can generate any text on that topic
    4. D.The tool only works by clicking through existing software screens and cannot generate any new written content
    5. E.The tool's output quality depends heavily on how clearly and specifically the prompt describes the desired outline
    6. F.The tool guarantees perfectly factual output because it was trained exclusively on peer-reviewed academic sources
    Show answer & explanation

    Correct answers: A, B, E — The tool can produce fluent, human-like text without a person writing each sentence from scratch; The tool can occasionally generate fabricated facts or citations that sound plausible but are not accurate; The tool's output quality depends heavily on how clearly and specifically the prompt describes the desired outline

    • A. This is correct because generating fluent, original text without a person drafting every sentence is a core, well-documented capability of generative AI tools like the one being evaluated.
    • B. This is correct because generative AI models are known to sometimes hallucinate, producing confident but fabricated facts or citations, which the marketing team should account for when reviewing drafts.
    • C. This is incorrect because generative AI models generalize from broad pretraining and do not need a dedicated labeled example for every specific topic before producing text on it.
    • D. This is incorrect because clicking through existing software screens without generating new content describes robotic process automation, not the generative AI tool being evaluated.
    • E. This is correct because the specificity and clarity of the prompt strongly influences how relevant and useful the generated outline turns out to be, which is a practical consideration for the team.
    • F. This is incorrect because no generative AI tool guarantees perfectly factual output, and claiming exclusive training on peer-reviewed sources does not eliminate the risk of fabricated or inaccurate content.

    Subdomain 1.3: Identify infrastructure concepts.

    5.An accounting department wants a shared network location where employees can browse nested folders organized by client and year, and open documents directly from their office applications over a standard network file-sharing protocol. Which storage type should be provisioned?

    1. A.File storage
    2. B.Block storage
    3. C.Object storage
    4. D.Local storage
    Show answer & explanation

    Correct answer: A — File storage

    • A. Organizing data in nested folders and serving it to multiple users over a network file-sharing protocol like SMB or NFS is exactly what this storage type provides.
    • B. A raw volume formatted for one server's file system does not natively support simultaneous folder browsing by many employees over a network protocol.
    • C. Files addressed by a unique identifier over an API do not present the nested folder browsing experience the department wants from their office applications.
    • D. Storage confined to a single device cannot be shared across the accounting department's employees or accessed over the network.

    Subdomain 1.3: Identify infrastructure concepts.

    6.A development team wants to package an application together with its exact runtime, libraries, and configuration so it runs identically on a developer's laptop, in the test environment, and in production, without the overhead of running a full separate guest operating system for each instance. Which infrastructure concept addresses this need?

    1. A.Virtual machines
    2. B.Hybrid cloud
    3. C.Containerization
    4. D.Block storage
    Show answer & explanation

    Correct answer: C — Containerization

    • A. Each virtual machine runs its own full guest operating system, which is the overhead the team specifically wants to avoid.
    • B. Combining on-premises and provider infrastructure addresses where workloads run, not how an application and its dependencies are packaged for consistent execution.
    • C. Packaging an application with its runtime, libraries, and configuration into a portable unit that shares the host operating system kernel is exactly what this concept provides.
    • D. This concept describes a raw storage volume format, not a way to package an application and its dependencies for consistent execution.

    Subdomain 1.1: Explain data concepts.

    7.An analyst runs an aggregate function to calculate the average order value across a table of transactions, but several rows have no value recorded in the discount column. How do most SQL aggregate functions such as AVG handle these missing entries by default?

    1. A.They exclude rows containing a null value from the calculation, rather than treating the missing value as zero.
    2. B.They automatically substitute a zero for every null value before performing the aggregate calculation across the column.
    3. C.They raise an error and halt execution whenever a null value appears anywhere in the aggregated column.
    4. D.They convert every null value in the column to the average of the remaining non-null values first.
    Show answer & explanation

    Correct answer: A — They exclude rows containing a null value from the calculation, rather than treating the missing value as zero.

    • A. This is correct because aggregate functions such as AVG typically ignore rows where the aggregated column is null, rather than treating the missing value as a zero.
    • B. This is incorrect because substituting zero for a null would understate the average, and standard aggregate functions do not perform this substitution automatically.
    • C. This is incorrect because encountering a null value in an aggregated column does not typically cause the query to fail or halt execution.
    • D. This is incorrect because aggregate functions do not pre-fill null values with a computed average before running; they simply exclude the null rows from the calculation.

    Subdomain 1.1: Explain data concepts.

    8.A finance team is designing a column to store monetary transaction amounts that must always retain exact precision to the cent, with no rounding errors introduced during storage or calculation. Which numeric data type should the column use?

    1. A.Decimal, because it stores an exact fixed-precision value, avoiding the rounding errors of floating-point storage.
    2. B.Float, because it offers the widest possible range of representable values at the cost of some storage space.
    3. C.Integer, because it stores whole numbers only, requiring the application to track cents in a separate column.
    4. D.Boolean, because it uses the least storage space of any numeric type, simplifying downstream calculations considerably.
    Show answer & explanation

    Correct answer: A — Decimal, because it stores an exact fixed-precision value, avoiding the rounding errors of floating-point storage.

    • A. Decimal is correct because it stores numbers with exact fixed precision, which avoids the rounding imprecision that floating-point types can introduce in monetary calculations.
    • B. Float is incorrect because it approximates values using binary floating-point representation, which can introduce small rounding errors unsuitable for exact monetary amounts.
    • C. Integer is incorrect because it cannot represent fractional cent values on its own, forcing an awkward workaround that decimal handles natively.
    • D. Boolean is incorrect because it can only represent two states, such as true or false, and cannot store a numeric monetary amount at all.

    Subdomain 1.2: Identify types of data sources.

    9.A developer wants an application to pull current currency exchange rates automatically several times per day from a third-party financial data provider, without manually downloading files or scraping any web pages. Which data source type fits this requirement?

    1. A.An application programming interface that returns structured exchange-rate data on each programmatic request
    2. B.A nightly CSV export that the finance team manually emails to the development team each day
    3. C.A screen-scraped copy of the provider's public exchange-rate webpage saved as an HTML file
    4. D.A read replica of the provider's internal relational database that the developer cannot access directly
    Show answer & explanation

    Correct answer: A — An application programming interface that returns structured exchange-rate data on each programmatic request

    • A. An API lets the application request current exchange-rate data programmatically as often as needed, matching the automated, several-times-per-day requirement without manual steps.
    • B. A manually emailed export depends on a person sending a file, which breaks the requirement for automated, on-demand refreshes multiple times a day.
    • C. Scraping a webpage into a saved HTML file introduces manual or brittle parsing steps and was explicitly ruled out by the scenario.
    • D. A database replica the developer cannot access directly is not a usable data source for this application, regardless of how current its data is.

    Subdomain 1.2: Identify types of data sources.

    10.A manufacturing company wants to ingest and store raw sensor readings, images, and unstructured maintenance notes from factory equipment in their original formats before deciding how to analyze them later. Which repository type best fits this need?

    1. A.A data lake, storing raw structured, semi-structured, and unstructured data without a fixed schema upfront
    2. B.A data mart, which stores only cleaned, department-specific data already summarized for reporting
    3. C.A data warehouse, which requires all incoming data to conform to a predefined relational schema
    4. D.A data silo, which isolates a single team's data from the rest of the organization's systems
    Show answer & explanation

    Correct answer: A — A data lake, storing raw structured, semi-structured, and unstructured data without a fixed schema upfront

    • A. A data lake accepts raw sensor readings, images, and unstructured notes in their native format without demanding a schema before ingestion, which matches this company's need to store first and decide analysis later.
    • B. A data mart holds already-cleaned, department-scoped data prepared for reporting, which contradicts the need to keep raw, undecided-format data.
    • C. A data warehouse requires incoming data to conform to a predefined relational schema, which raw images and unstructured notes cannot satisfy without prior transformation.
    • D. A data silo describes an isolated store that other teams cannot access; it is an organizational problem rather than a repository designed for raw multi-format ingestion.

    Domain 2: Data Acquisition and Preparation

    Subdomain 2.2: Given a scenario, perform data exploration to identify possible inconsistencies with a data set.

    11.An analyst is asked to define what makes a data set "complete" before it is loaded into a reporting dashboard. Which definition best matches how completeness is understood in data exploration?

    1. A.The data set contains all the records and field values that are expected to be present for its intended purpose
    2. B.The data set contains no two rows that are exact duplicates of each other across every column
    3. C.The data set contains only values that fall within three standard deviations of the column mean
    4. D.The data set contains fields whose values have been converted to a single consistent unit of measure
    Show answer & explanation

    Correct answer: A — The data set contains all the records and field values that are expected to be present for its intended purpose

    • A. This definition is correct because completeness measures whether the expected records and field values are actually present, which is the standard meaning of the term in data exploration.
    • B. This definition is incorrect because the absence of duplicate rows describes deduplication, not completeness.
    • C. This definition is incorrect because staying within a statistical range describes the absence of outliers, not whether expected data is present.
    • D. This definition is incorrect because consistent units of measure describe a standardization or transformation step, not whether the data set has all its expected values.

    Subdomain 2.2: Given a scenario, perform data exploration to identify possible inconsistencies with a data set.

    12.A payroll analyst is validating a new employee data feed and defines a rule that the `salary` field must be a positive number and the `employment_status` field must be one of `Active`, `Leave`, or `Terminated`. After running the check, 12 rows fail the `employment_status` rule with a value of `Retired`. What is the most appropriate next step?

    1. A.Confirm with the source system owner whether `Retired` is a legitimate status that should be added to the accepted value list
    2. B.Automatically convert every `Retired` value to `Terminated` without further review, since both indicate the employee no longer works there
    3. C.Delete the 12 rows from the feed immediately, since any value outside the defined list cannot be loaded under any circumstance
    4. D.Ignore the failed rows and proceed with the load, since validation rules are only advisory and do not affect downstream reporting
    Show answer & explanation

    Correct answer: A — Confirm with the source system owner whether `Retired` is a legitimate status that should be added to the accepted value list

    • A. Confirming with the source system owner is correct because a validation rule that rejects a value not previously anticipated should be reviewed against real business meaning before deciding whether to expand the rule or correct the data.
    • B. Automatically converting to `Terminated` without review is incorrect because it assumes equivalence between two distinct statuses without confirming that assumption is actually accurate for payroll purposes.
    • C. Deleting the rows immediately is incorrect because it removes legitimate employee records before determining whether `Retired` should simply be added as a valid status.
    • D. Ignoring the failed rows is incorrect because validation rules exist specifically to catch values that do not conform to expected business logic, and skipping the review defeats the purpose of running the check.

    Subdomain 2.1: Given a scenario, use data acquisition methods.

    13.A team maintains identical monthly sales export files, `sales_jan.csv` through `sales_dec.csv`, each with the same eleven columns in the same order. The analyst wants a single result set containing every row from all twelve files stacked into one table for the year. Which querying technique accomplishes this?

    1. A.Concatenate the twelve files by appending their rows into a single result set, since the column structure is identical across all of them.
    2. B.Join the twelve files on a shared date column, matching each row in one file to a corresponding row in the other eleven files.
    3. C.Apply a nested query that selects the January file as an inner query and filters the other eleven files against its results.
    4. D.Group the twelve files by month and aggregate each column into a single summarized total row per file, discarding all individual sales records.
    Show answer & explanation

    Correct answer: A — Concatenate the twelve files by appending their rows into a single result set, since the column structure is identical across all of them.

    • A. Concatenating, also called a union, appends the rows of each file to the others when the column layout matches, producing one continuous table of all twelve months of sales rows, which is what the scenario requires.
    • B. Joining on a shared column combines columns from two sources side by side for matching keys, which would widen the table with duplicate or mismatched columns rather than stack the same eleven columns across all twelve files.
    • C. A nested query embeds one query inside another to filter or compute a value, but it does not stack independent files with identical structure into one combined row set.
    • D. Grouping and aggregating collapses rows into summary totals per group, which would lose the individual sales records rather than preserve every row from all twelve files.

    Subdomain 2.1: Given a scenario, use data acquisition methods.

    14.A retailer wants to understand how satisfied customers are with a new checkout process. Rather than analyzing existing transaction logs, the analyst designs a short questionnaire with rating-scale and open-text questions and sends it to customers who recently checked out. Which data collection method is being used?

    1. A.Surveying, since the analyst is directly asking customers structured questions to gather their opinions and feedback.
    2. B.Sampling, since the analyst is selecting a representative subset of the transaction log to review in detail.
    3. C.Query optimization, since the analyst is designing indexed questions to speed up future data retrieval.
    4. D.Data integration, since the analyst is combining the new questionnaire responses with the existing transaction log table.
    Show answer & explanation

    Correct answer: A — Surveying, since the analyst is directly asking customers structured questions to gather their opinions and feedback.

    • A. Surveying is the practice of collecting information directly from people through structured questions, which matches the analyst designing and sending a questionnaire with rating-scale and open-text items to gather customer opinions.
    • B. Sampling refers to selecting a representative subset of an existing dataset for analysis, not to designing and distributing new questions to gather fresh opinions from customers.
    • C. Query optimization is a set of techniques for improving database query performance, such as indexing and parameterization, and has no relationship to designing a customer questionnaire.
    • D. Data integration means combining data from multiple existing sources into a unified view, which is a separate later activity from the act of collecting new opinion data through a questionnaire.

    Subdomain 2.3: Given a scenario, perform appropriate data transformation and cleansing techniques.

    15.A statistician is comparing exam scores across two different tests that use different scales and different score distributions. To make the scores comparable, the statistician subtracts each test's mean score from every value and divides the result by that test's standard deviation. Which technique is being used?

    1. A.Standardization, which rescales values based on the mean and standard deviation of the distribution so they can be compared across differing scales.
    2. B.Binning, which groups the exam scores into a fixed number of discrete performance categories, such as low, medium, and high tiers.
    3. C.Imputation, which estimates and fills in exam scores that are entirely missing from the dataset using values calculated from the remaining recorded scores.
    4. D.Augmentation, which enriches the exam score dataset with additional attributes obtained from an external student records system.
    Show answer & explanation

    Correct answer: A — Standardization, which rescales values based on the mean and standard deviation of the distribution so they can be compared across differing scales.

    • A. Subtracting the mean and dividing by the standard deviation is the defining calculation of standardization, which puts differently distributed variables onto a common, comparable scale.
    • B. Binning groups values into discrete performance categories and does not involve the mean-and-standard-deviation calculation used to make two distributions comparable.
    • C. Imputation addresses gaps caused by missing scores and has no bearing on scores that are already present and simply need to be rescaled.
    • D. Augmentation brings in attributes from an outside source, which is unrelated to a purely mathematical rescaling of an existing numeric column.

    Subdomain 2.3: Given a scenario, perform appropriate data transformation and cleansing techniques.

    16.What is the key distinction between merging and appending two datasets?

    1. A.Merging joins datasets side by side using a shared key column, while appending stacks datasets with matching structure on top of each other.
    2. B.Merging always requires deleting any existing duplicate rows first, while appending never removes rows from either of the two source datasets.
    3. C.Merging can only be performed on numeric columns, while appending can only be performed on text or categorical columns.
    4. D.Merging changes a column's data type during the join, while appending changes a column's data type during the stack.
    Show answer & explanation

    Correct answer: A — Merging joins datasets side by side using a shared key column, while appending stacks datasets with matching structure on top of each other.

    • A. Merging aligns two datasets horizontally by matching a shared key, while appending stacks compatible datasets vertically, adding more rows rather than more columns.
    • B. Neither operation inherently requires deleting duplicates or guarantees no rows are removed; duplicate handling is a separate decision made independently of whether the operation is a merge or an append.
    • C. Merging can join on keys of any data type, including text identifiers, so the operation is not restricted to numeric columns alone.
    • D. Neither merging nor appending is defined by changing a column's data type; type conversion is a separate transformation applied independently of combining datasets.

    Domain 3: Data Analysis

    Subdomain 3.2: Given a scenario, select the appropriate statistical method or function.

    17.In a spreadsheet, an analyst wants a column that displays `Yes` when a customer's order total exceeds $500 and `No` otherwise. Which function accomplishes this?

    1. A.`IF`, a logical function that tests a condition and returns one value when true and another when false.
    2. B.`TRIM`, a string function that removes leading and trailing spaces from a text value.
    3. C.`YEAR`, a date function that extracts the calendar year portion from a date value.
    4. D.`AVERAGE`, a mathematical function that calculates the mean of a range of numeric values.
    Show answer & explanation

    Correct answer: A — `IF`, a logical function that tests a condition and returns one value when true and another when false.

    • A. Testing whether the order total exceeds $500 and returning a different result for true versus false matches exactly what this function is designed to do.
    • B. Removing extra spaces from text has no ability to test a numeric threshold or return a conditional Yes or No result.
    • C. Extracting the calendar year from a date value has no relationship to testing whether an order total exceeds a threshold.
    • D. Calculating a mean across a range of values produces a single number, not a conditional Yes or No label for each order.

    Subdomain 3.2: Given a scenario, select the appropriate statistical method or function.

    18.A dataset imported from a legacy system contains customer names with extra leading and trailing spaces, causing exact-match lookups to fail. Which type of function should the analyst apply to fix this?

    1. A.A string function such as `TRIM`, which removes leading and trailing spaces from a text value.
    2. B.A logical function such as `AND`, which returns true only when every one of its specified conditions is also true.
    3. C.A mathematical function such as `ROUND`, which rounds a numeric value to a specified number of decimal places.
    4. D.A date function such as `YEAR`, which extracts only the calendar year component from a date value.
    Show answer & explanation

    Correct answer: A — A string function such as `TRIM`, which removes leading and trailing spaces from a text value.

    • A. Removing the extra leading and trailing spaces directly fixes the exact-match lookup problem described for these imported customer names.
    • B. Requiring multiple conditions to be true evaluates logical conditions and does not remove extra spaces from imported text values.
    • C. Adjusting decimal precision applies to numeric values and has no effect on extra spaces within a text field.
    • D. Extracting a calendar year applies to date values and has no relationship to fixing extra spaces in a text name field.

    Subdomain 3.1: Given a set of requirements, determine the appropriate communication approach for data analysis.

    19.A report author wants to define a KPI for a customer support team's weekly performance review. Which of the following is the best-formed KPI statement?

    1. A.First-contact resolution rate, targeted at 85% or higher, tracked weekly against the prior period.
    2. B.Number of coffee breaks taken by support agents during each eight-hour shift on the schedule.
    3. C.The support team's preferred meeting room location for the weekly performance review session.
    4. D.A general statement that the support team should try to do better than it did last week.
    Show answer & explanation

    Correct answer: A — First-contact resolution rate, targeted at 85% or higher, tracked weekly against the prior period.

    • A. This is correct because it names a specific, measurable indicator (first-contact resolution rate), a target threshold, and a tracking cadence, all core traits of a well-formed KPI.
    • B. Counting coffee breaks is measurable but has no established connection to support team performance or service goals, so it does not function as a meaningful KPI.
    • C. A meeting room preference is a logistical detail, not a measurable performance indicator, so it cannot serve as a KPI for the team's weekly review.
    • D. A vague statement to 'do better' lacks any specific, measurable indicator or target, which fails the basic definition of a KPI.

    Subdomain 3.3: Given a scenario, troubleshoot basic issues using the appropriate tool or method.

    20.A query that aggregates sales by region fails with an error indicating that a selected column is not part of an aggregate function or the GROUP BY clause. What is the most direct fix?

    1. A.Add the non-aggregated column to the GROUP BY clause so every selected column is either grouped or aggregated.
    2. B.Restart the database connection, since ambiguous aggregate errors are typically caused by dropped network sessions.
    3. C.Validate that the source table has not been corrupted by comparing row counts against a recent backup file.
    4. D.Search the vendor's community forum for a firmware update that resolves aggregate function limitations.
    Show answer & explanation

    Correct answer: A — Add the non-aggregated column to the GROUP BY clause so every selected column is either grouped or aggregated.

    • A. This specific error means a selected column must be grouped or aggregated, so adding it to the GROUP BY clause directly satisfies that requirement.
    • B. This is a query-parsing error unrelated to the network session, so restarting the connection would not change the outcome.
    • C. The error occurs regardless of the data's integrity, since it is about how the query is structured, not about corrupted values.
    • D. GROUP BY behavior is standard SQL syntax, not a firmware limitation, so no vendor patch is needed to resolve this error.

    Subdomain 3.3: Given a scenario, troubleshoot basic issues using the appropriate tool or method.

    21.A report that ran successfully for months suddenly fails after an upstream team migrates their system to a new platform, though no changes were made to the reporting query. What should the analyst validate first?

    1. A.Whether the source table's schema, column names, or data types changed as part of the upstream migration.
    2. B.Whether the analyst's own laptop has enough local disk space to render the report's visualizations.
    3. C.Whether the report's color scheme still meets the organization's branding and accessibility guidelines.
    4. D.Whether the vendor's community forum has published a new certification exam covering the platform.
    Show answer & explanation

    Correct answer: A — Whether the source table's schema, column names, or data types changed as part of the upstream migration.

    • A. A migration timed exactly with the failure, and no query changes on this end, strongly points to a schema or structure change made by the upstream system.
    • B. Local disk space affects how a report renders on one machine, not why a server-side query would suddenly fail after an unrelated migration.
    • C. Branding and accessibility of the report's visuals have no bearing on whether the query executing against the migrated source succeeds.
    • D. A certification exam listing is unrelated to diagnosing a technical schema change caused by an infrastructure migration.

    Subdomain 3.1: Given a set of requirements, determine the appropriate communication approach for data analysis.

    22.A company is preparing a public-facing annual report that will include several data visualizations. A visually impaired employee on the review team relies on a screen reader to consume the report. Which action best supports auditory accessibility for this reviewer?

    1. A.Add descriptive alt text and a written data summary for each chart so a screen reader can convey the chart's meaning aloud.
    2. B.Increase the font size of chart titles, axis labels, and data callouts so the visualizations remain legible when reviewed on screen.
    3. C.Switch every chart's color palette to a colorblind-safe set with textured fills so hue and shape differences remain distinguishable.
    4. D.Export the report as a high-resolution, vector-based PDF with embedded fonts so the charts remain sharp when the document is zoomed in on screen.
    Show answer & explanation

    Correct answer: A — Add descriptive alt text and a written data summary for each chart so a screen reader can convey the chart's meaning aloud.

    • A. Alt text and written summaries are correct because screen readers convert text to speech, so a chart without a text equivalent conveys nothing to a user relying on audio output.
    • B. Increasing font size addresses low vision, a visual accessibility need, not the auditory access provided by screen-reader-compatible text descriptions.
    • C. A colorblind-safe palette addresses a visual color-perception need and does not help a screen reader convey chart content through audio.
    • D. A sharper, zoomable PDF improves visual clarity at higher resolution but does not add any text a screen reader could announce, so it does not address auditory access.

    Domain 4: Visualization and Reporting

    Subdomain 4.2: Given a scenario, use the appropriate delivery or consumption method.

    23.A finance director requests a report comparing this year's audit findings to last year's, but only wants it produced once for an upcoming board meeting rather than every month going forward. Which report frequency should the analytics team apply?

    1. A.An ad hoc report, since it is generated for this single, unplanned occasion rather than on a repeating schedule
    2. B.A recurring report, since audit comparisons should always be scheduled to run automatically every month
    3. C.A dynamic dashboard, since only continuously refreshing visuals can compare two years of audit data
    4. D.A self-service portal, since the director should build the comparison independently rather than request it
    Show answer & explanation

    Correct answer: A — An ad hoc report, since it is generated for this single, unplanned occasion rather than on a repeating schedule

    • A. A one-time request tied to a specific upcoming meeting, rather than an ongoing schedule, is the defining trait of an ad hoc report.
    • B. The director explicitly does not want this produced every month, so scheduling it as a recurring report would contradict the stated requirement.
    • C. Comparing two fixed prior periods does not require continuous refreshing, so a dynamic dashboard is not necessary to satisfy this one-time comparison.
    • D. The director asked the analytics team to produce the report rather than build it themselves, so a self-service portal does not match the stated request.

    Subdomain 4.2: Given a scenario, use the appropriate delivery or consumption method.

    24.A product team wants a weekly automated report of feature adoption metrics sent to stakeholders every Monday, using data collected as of the previous Sunday night, without needing anyone to manually request it. Which combination of frequency and versioning best fits this need?

    1. A.A recurring report paired with snapshot versioning captured at a consistent weekly point in time
    2. B.An ad hoc report paired with real-time versioning that updates continuously throughout the week
    3. C.A recurring report paired with real-time versioning that changes every time a user views it
    4. D.An ad hoc report paired with snapshot versioning captured only when a stakeholder asks for it
    Show answer & explanation

    Correct answer: A — A recurring report paired with snapshot versioning captured at a consistent weekly point in time

    • A. Automatically sending the report every Monday without manual requests matches recurring frequency, and using data as of a fixed prior point (Sunday night) matches snapshot versioning.
    • B. The scenario explicitly states the report goes out automatically every Monday rather than being requested, which rules out describing it as ad hoc.
    • C. The report uses a fixed Sunday-night capture rather than data that changes every time it is viewed, so real-time versioning does not match the described behavior.
    • D. The report is sent automatically on a schedule rather than only when requested, so labeling it ad hoc contradicts the stated weekly automation.

    Subdomain 4.1: Given a scenario, use the appropriate visual elements.

    25.An analyst finishes a quarterly report using default software colors and a generic font, but the company's style guide specifies a particular logo placement, color palette, and typeface for all external documents. What should the analyst do before distribution?

    1. A.Apply the company's approved logo, color palette, and typeface to the report so it matches other official external materials
    2. B.Leave the report in its default styling, since visual branding has no bearing on how external stakeholders interpret the data
    3. C.Replace all charts with plain text summaries, since removing visuals avoids any need to apply a consistent brand style
    4. D.Add multiple unrelated logos from different departments to the report, since more branding elements signal more legitimacy
    Show answer & explanation

    Correct answer: A — Apply the company's approved logo, color palette, and typeface to the report so it matches other official external materials

    • A. Applying the approved logo, palette, and typeface is correct because consistent branding across external materials builds recognition and signals that the report is official company output.
    • B. Branding affects how credible and recognizable a report appears to external stakeholders, so leaving default styling in place contradicts the style guide's purpose.
    • C. Replacing charts with plain text avoids the branding question but sacrifices the clarity that visualizations provide, which is not what the style guide requires.
    • D. Adding multiple unrelated logos creates visual confusion about which department or brand the report represents, rather than the single consistent identity the guide specifies.

    Subdomain 4.3: Given a scenario, troubleshoot issues using report validation techniques.

    26.A finance dashboard refreshes noticeably slower than it did last month, even though the underlying warehouse table size has stayed roughly the same. Which technique would most directly help isolate the cause of a degrading refresh rate over time?

    1. A.Review the calculation logic and query steps that run during the refresh to find a newly introduced inefficient join or nested calculation
    2. B.Have a peer inspect the dashboard's color palette to confirm it still matches the corporate branding guide
    3. C.Reduce the number of visible legends and labels so the rendering engine has fewer elements to draw
    4. D.Change the report's delivery frequency from daily to weekly so refreshes happen less often
    Show answer & explanation

    Correct answer: A — Review the calculation logic and query steps that run during the refresh to find a newly introduced inefficient join or nested calculation

    • A. A calc/code review of the refresh logic can surface a recently added inefficient join, a recalculated measure, or an unoptimized query step that is now taking longer even though data volume is stable, which directly targets the refresh-rate regression.
    • B. Checking the color palette against branding guidelines is a visual design concern and has no bearing on how long the refresh query takes to execute.
    • C. Reducing visible legends changes rendering load slightly but does not address a refresh that is slow because of the underlying query or calculation logic, which is the described symptom.
    • D. Lowering delivery frequency hides the slow refresh from users less often but does not diagnose or fix why the refresh itself has gotten slower.

    Subdomain 4.3: Given a scenario, troubleshoot issues using report validation techniques.

    27.A retail chain's dashboard queries live point-of-sale data across 500 stores at the individual transaction level for every page load, and users complain about both excessive load time and an occasional filter that seems to ignore the selected store. The team can only fix one thing before the next release. Which fix addresses the higher-impact, more fundamental issue?

    1. A.Restructure the source data with pre-aggregated, filterable summary tables so the load time and the filter's underlying field references are corrected at once.
    2. B.Add a data versioning toggle so users can switch between a cached snapshot and the live real-time view, but each mode still queries unfiltered transaction-level records across all 500 stores.
    3. C.Draft a more detailed executive summary that explains to leadership why the dashboard has been slow, citing store-level query volume and transaction counts as the root cause of the delay.
    4. D.Change the dashboard's color scheme and add loading-state indicators so slow-loading sections are visually highlighted while store filters and transaction fields render unchanged.
    Show answer & explanation

    Correct answer: A — Restructure the source data with pre-aggregated, filterable summary tables so the load time and the filter's underlying field references are corrected at once.

    • A. Restructuring the source into pre-aggregated summary tables with correct filterable fields reduces the row-level volume driving the load-time problem and fixes the store field the filter depends on, addressing both symptoms from their shared root cause.
    • B. Adding a versioning toggle changes how current the displayed data is but still queries the same unfiltered, transaction-level volume, so it would not resolve the load time or the filter mismatch.
    • C. Explaining the slowness to leadership communicates the problem but does not reduce data volume or fix the broken filter reference.
    • D. Highlighting slow sections with color is a cosmetic indicator of the problem and does not reduce query volume or correct the filter's underlying issue.

    Subdomain 4.1: Given a scenario, use the appropriate visual elements.

    28.An analyst plots website traffic and revenue on the same chart using two different vertical axes, one scaled in thousands of visits and the other in dollars. A reviewer says the chart makes it look like traffic and revenue move together more closely than the raw numbers support. What is the most likely cause?

    1. A.The two axes were scaled independently, so aligning their ranges visually can overstate or understate how closely the two lines actually move together.
    2. B.The chart was built from a pivot table instead of a native chart object, so it cannot render two data series on independently scaled axes without merging them onto one axis.
    3. C.The chart uses a categorical color palette instead of a sequential scale, and that mismatched color choice is what visually aligns the traffic and revenue lines on the page.
    4. D.The chart is missing a legend identifying which axis belongs to which series, and adding one would automatically correct the mismatch between the two independently scaled axes.
    Show answer & explanation

    Correct answer: A — The two axes were scaled independently, so aligning their ranges visually can overstate or understate how closely the two lines actually move together.

    • A. Independently scaled dual axes are correct because choosing each axis range separately can visually align two unrelated lines, creating an impression of correlation that the underlying numbers may not support.
    • B. The scenario describes an existing chart with two visible lines and axes, not a pivot table, so this option does not match what is actually being described.
    • C. Color scheme choice affects how categories are distinguished visually but has no bearing on how the two numeric axis ranges were scaled relative to each other.
    • D. A missing legend would make it harder to identify which line is which, but it does not change the independent axis scaling that is causing the misleading visual alignment.

    Domain 5: Data Governance

    Subdomain 5.1: Explain data management concepts.

    29.A retail company's product catalog organizes items into categories, subcategories, and individual SKUs, where each SKU belongs to exactly one subcategory and each subcategory belongs to exactly one category. Which documentation concept describes this parent-child organization of data?

    1. A.Hierarchy structure
    2. B.Data lineage
    3. C.Source of truth
    4. D.Data dictionary
    Show answer & explanation

    Correct answer: A — Hierarchy structure

    • A. Modeling categories, subcategories, and SKUs as nested parent-child levels is exactly what this concept documents, capturing how data elements relate to one another structurally.
    • B. This concept traces where a data element originated and how it was transformed, not how categories and subcategories nest within one another.
    • C. This concept identifies the single authoritative dataset for a given fact, which is unrelated to describing parent-child category relationships.
    • D. This artifact defines field names, types, and valid values for a dataset, not the nested organization of categories and subcategories.

    Subdomain 5.1: Explain data management concepts.

    30.In the context of data versioning, what does a "refresh interval" define?

    1. A.How frequently a dataset or report is updated with new data from its source, such as hourly, daily, or weekly.
    2. B.The exact point in time at which a full copy of a dataset was captured and archived for historical comparison.
    3. C.The set of permissions that determines which users are allowed to view or modify a shared dataset.
    4. D.The specific database engine setting that controls how long a query result stays cached before eviction.
    Show answer & explanation

    Correct answer: A — How frequently a dataset or report is updated with new data from its source, such as hourly, daily, or weekly.

    • A. A refresh interval specifies the cadence at which a dataset or report pulls new data from its source, whether that cadence is hourly, daily, or weekly.
    • B. This describes a snapshot, which captures a fixed historical copy, rather than the recurring cadence at which live data is refreshed.
    • C. This describes access control, which governs who can view or modify data, and is unrelated to how often a dataset's contents are updated.
    • D. This describes a query cache setting, which controls result reuse at the engine level, not the business cadence at which a dataset is refreshed with new source data.

    Subdomain 5.3: Compare and contrast data privacy and protection practices.

    31.A data engineering team is building a pipeline that pushes customer order records from an on-premises database to a cloud data warehouse every night over the public internet. Which control most directly protects the records while they travel between the two systems?

    1. A.Encrypt the connection between the source database and the cloud warehouse using TLS so the transferred records are unreadable if intercepted in transit.
    2. B.Compress the nightly export files before transfer so the transmission completes faster and consumes less network bandwidth overall.
    3. C.Store the customer order records in an encrypted database volume at the source system so files sitting on disk cannot be read by anyone lacking the decryption key.
    4. D.Restrict which employee accounts can query the cloud warehouse after the nightly load finishes so only authorized analysts see the data.
    Show answer & explanation

    Correct answer: A — Encrypt the connection between the source database and the cloud warehouse using TLS so the transferred records are unreadable if intercepted in transit.

    • A. TLS encryption on the transfer channel is exactly what encryption in transit means: it protects data while it moves across a network, which is the exposure this nightly transfer creates. Anyone intercepting the traffic would only see ciphertext.
    • B. Compression changes file size and transfer speed but does nothing to prevent an interceptor from reading the contents once decompressed. It is a performance technique, not a protection against exposure during transmission.
    • C. Encrypting the source database volume protects data at rest on that server's disks, not data while it is moving across the network to the warehouse. The two controls address different stages of the data's lifecycle.
    • D. Restricting who can query the warehouse after the load is an access control measure that governs use once the data has already arrived. It does not protect the records during the nightly transfer itself.

    Subdomain 5.3: Compare and contrast data privacy and protection practices.

    32.A government statistics agency wants to publish a dataset of survey responses to a public research repository, where the identity of any respondent must never be recoverable, even by combining the release with other public data sources. Which combination of practices best supports this goal? (Select all that apply.)(Select 2)

    1. A.Anonymize the dataset by removing or generalizing fields, such as exact birth dates, that could be combined with outside data to re-identify a respondent.
    2. B.Suppress or aggregate small population subgroups so that no combination of remaining fields narrows the data down to a single identifiable respondent.
    3. C.Mask the respondent names with realistic fictitious substitutes while keeping the original values stored separately in a secure archive for future internal reference.
    4. D.Grant read-only role-based access to the published file so that only the statistics agency's own analysts can open the public repository listing.
    5. E.Encrypt the published file at rest using a key retained by the agency, and share that key with any researcher who requests the dataset.
    Show answer & explanation

    Correct answers: A, B — Anonymize the dataset by removing or generalizing fields, such as exact birth dates, that could be combined with outside data to re-identify a respondent.; Suppress or aggregate small population subgroups so that no combination of remaining fields narrows the data down to a single identifiable respondent.

    • A. Removing or generalizing quasi-identifying fields like exact birth dates is a core anonymization step that prevents re-identification through combination with outside datasets, which directly matches the agency's stated goal.
    • B. Suppressing or aggregating small subgroups prevents a respondent from being isolated through a rare combination of attributes, a standard anonymization technique for public statistical releases where no individual should be re-identifiable.
    • C. Retaining the original identifying values elsewhere means the release is reversible in principle, which fails the goal of an anonymized public dataset where identity must never be recoverable by anyone, including the publisher.
    • D. A public research repository is meant to be openly accessible; restricting it to the agency's own analysts contradicts the purpose of a public release and does not itself address re-identification risk within the published content.
    • E. Distributing the decryption key to any requesting researcher effectively makes the file's contents as accessible as if it were unencrypted, and encryption does nothing to remove identifying detail from the dataset itself.

    Subdomain 5.4: Compare and contrast data quality assurance practices.

    33.After a vendor delivers a newly built customer segmentation report, the business intelligence lead schedules a session where the marketing team logs into the tool, runs their typical filters, and confirms the segments and totals match what they need for the upcoming campaign before sign-off. Which practice does this describe?

    1. A.Stress testing, subjecting the report's query engine to a simulated surge of concurrent marketing users
    2. B.User acceptance testing, having the actual business users validate that the delivered report meets their working requirements
    3. C.Source control review, checking the report's version history to confirm the vendor's changes were properly tracked
    4. D.Data profiling, scanning the underlying tables to measure completeness and uniqueness before the report was built
    Show answer & explanation

    Correct answer: B — User acceptance testing, having the actual business users validate that the delivered report meets their working requirements

    • A. Subjecting the query engine to a simulated user surge measures performance under load, not whether the report's content satisfies the marketing team's working needs.
    • B. Having the actual business users run their typical filters and confirm the segments and totals meet their needs before sign-off is the defining activity of user acceptance testing.
    • C. Checking version history confirms that changes were tracked and approved, which is unrelated to marketing staff validating the report's content against their needs.
    • D. Scanning the underlying tables for completeness and uniqueness happens earlier, before the report is built, and does not involve business users validating the finished deliverable.

    Subdomain 5.4: Compare and contrast data quality assurance practices.

    34.A data quality manager wants to align the organization's data quality management practices with an internationally recognized standards framework rather than relying solely on internally developed rules. Which organization publishes standards, including ones specifically covering data quality, that the manager should reference?

    1. A.International Organization for Standardization (ISO), which publishes internationally recognized data quality management standards
    2. B.Project Management Institute (PMI), which publishes standards and certifications focused on project management methodologies and processes
    3. C.Institute of Electrical and Electronics Engineers (IEEE), which publishes technical standards primarily for electrical and electronics engineering
    4. D.National Institute of Standards and Technology (NIST), which publishes cybersecurity and privacy framework guidance for U.S. federal agencies
    Show answer & explanation

    Correct answer: A — International Organization for Standardization (ISO), which publishes internationally recognized data quality management standards

    • A. The International Organization for Standardization publishes internationally recognized standards that specifically address data quality management, making it the reference the manager is looking for.
    • B. This organization's standards and certifications focus on project management methodologies, which is a different discipline than data quality management practices.
    • C. This organization's technical standards are primarily aimed at electrical and electronics engineering rather than data quality management practices.
    • D. This organization's frameworks are best known for cybersecurity and privacy guidance for U.S. federal agencies, not a general-purpose international data quality management standard.

    Subdomain 5.2: Summarize concepts related to data compliance.

    35.An analyst is designing a data retention schedule for a hospital's patient records system that must satisfy both regulatory requirements and storage cost constraints. Which approach best reflects sound retention practice?

    1. A.Define retention periods per record type based on applicable regulations, then automatically archive or delete records once each period expires
    2. B.Keep every patient record indefinitely in the primary production database so nothing is ever at risk of being unavailable
    3. C.Delete records as soon as a patient's treatment episode formally ends, since older records add no further analytical value to the ongoing care team
    4. D.Let each department decide informally how long to keep records, since retention needs vary too much to standardize across a hospital
    Show answer & explanation

    Correct answer: A — Define retention periods per record type based on applicable regulations, then automatically archive or delete records once each period expires

    • A. Sound retention practice ties specific retention periods to record type and the regulations governing that type, then enforces disposal or archival automatically once the period lapses.
    • B. Keeping everything indefinitely in production ignores regulatory minimums and maximums, inflates storage cost, and increases exposure if a breach occurs, since more data sits at risk longer.
    • C. Deleting records the moment treatment ends can violate legally mandated minimum retention periods that many healthcare regulations impose for patient records.
    • D. Leaving retention to informal departmental judgment creates inconsistent compliance postures and makes it difficult to demonstrate adherence to a regulator during an audit.

    Want the full experience?

    These are just samples. Practice the full CompTIA Data+ question bank in quiz mode — free, no signup, with domain practice and exam simulation.