CertSafari

    Free AWS Certified Data Engineer - Associate (DEA-C01) Sample Questions

    35 free sample questions from our bank of 349+, covering every exam domain, with answers and detailed explanations. Updated September 2026.

    Domain 1: Data Ingestion and Transformation

    Subdomain 1.4: Apply programming concepts

    1.A team wants to detect when a scheduled Lambda-based ingestion job silently stops producing new records in its destination table, rather than only being alerted when the function throws an error. Which combination of practices addresses this need?

    1. A.Emit a custom CloudWatch metric for records processed per run and alarm when the metric drops below an expected threshold
    2. B.Increase the function's memory allocation so it has more headroom before approaching its configured execution timeout
    3. C.Wrap the business logic in a broad `try`/`except` block that logs any exception message before returning normally
    4. D.Reduce the function's reserved concurrency to one so only a single invocation can run across all scheduled triggers
    Show answer & explanation

    Correct answer: A — Emit a custom CloudWatch metric for records processed per run and alarm when the metric drops below an expected threshold

    • A. A custom metric tracking records processed, with an alarm on it dropping below expectation, catches a job that succeeds while producing unexpectedly few records.
    • B. Increasing memory can help memory-bound performance, but it does nothing to detect a job that completes successfully while silently under-producing records.
    • C. Catching and logging exceptions surfaces real errors, but a job that runs successfully without throwing raises no exception for this pattern to catch.
    • D. Limiting concurrency to one controls how many instances run simultaneously, with no bearing on detecting whether a completed run produced expected volume.

    Subdomain 1.4: Apply programming concepts

    2.A data engineering team is comparing infrastructure-as-code tools for a new serverless pipeline that must be redeployed consistently across dev, test, and production accounts. Which two statements about infrastructure as code are accurate for this scenario? (Select TWO.)(Select 2)

    1. A.Storing templates in version control lets every deployed change be traced back to a specific commit, author, and timestamp
    2. B.Declaring resources in a template lets the same definition be deployed repeatably across multiple accounts without manual steps
    3. C.Infrastructure as code requires every resource to be created exclusively through the AWS Management Console before automation applies
    4. D.A template-based deployment guarantees zero downtime for every resource type regardless of how the update changes the resource
    5. E.Infrastructure as code eliminates the need for any testing of the deployed resources once a template passes syntax validation
    Show answer & explanation

    Correct answers: A, B — Storing templates in version control lets every deployed change be traced back to a specific commit, author, and timestamp; Declaring resources in a template lets the same definition be deployed repeatably across multiple accounts without manual steps

    • A. Version-controlled templates tie every change to a commit, author, and timestamp, giving the team traceability across accounts and over time for incidents.
    • B. A declarative template can be applied the same way in each account, producing consistent resources without engineers repeating manual console steps.
    • C. Infrastructure as code is defined in templates and applied through automation instead of the console; it does not require console-first resource creation.
    • D. Some resource updates require replacement or cause brief interruption depending on the change type, so zero downtime is not guaranteed by templating alone.
    • E. Passing syntax validation only confirms the template parses correctly; it does not verify the deployed resources behave correctly, so testing still matters.

    Subdomain 1.2: Transform and process data

    3.A team needs an AWS Glue ETL job to connect to a legacy database engine for which AWS Glue does not provide a built-in JDBC connection type. The team has obtained a JDBC driver JAR file from the database vendor. What is the correct way to make this driver available to the Glue job?

    1. A.Upload the JDBC driver JAR to an Amazon S3 location and reference that S3 path when creating the Glue connection, so the job loads the custom driver at runtime.
    2. B.Install the driver JAR directly onto the AWS Glue service's underlying infrastructure by opening a support case, since customers cannot add drivers to a fully managed service themselves.
    3. C.Convert the JDBC driver into an AWS Lambda layer and attach that layer to the Glue job, because Glue jobs share the same runtime environment as Lambda functions.
    4. D.Rewrite the connection using Amazon Data Firehose instead of JDBC, since Firehose can substitute for any custom JDBC driver when connecting to unsupported database engines.
    Show answer & explanation

    Correct answer: A — Upload the JDBC driver JAR to an Amazon S3 location and reference that S3 path when creating the Glue connection, so the job loads the custom driver at runtime.

    • A. AWS Glue lets you upload a custom JDBC driver JAR to Amazon S3 and reference its path when defining the connection, so the job loads that driver at runtime for engines Glue does not support natively.
    • B. Glue is fully managed, but customers add custom JDBC drivers themselves through the S3-referenced connection mechanism rather than through a support case that modifies the underlying service infrastructure.
    • C. Glue jobs run in a Spark environment, not the AWS Lambda runtime, so Lambda layers are not a mechanism for supplying dependencies to a Glue ETL job.
    • D. Amazon Data Firehose delivers streaming data to a fixed set of destinations and does not connect to arbitrary relational databases over JDBC, so it cannot substitute for a custom driver.

    Subdomain 1.2: Transform and process data

    4.A data engineering team wants to enrich its AWS Glue Data Catalog with generated descriptions of table contents and also classify free-text customer support tickets by topic as part of an existing Lambda-based transformation pipeline. Which two approaches correctly use a large language model through Amazon Bedrock for these tasks?(Select 2)

    1. A.Invoke a foundation model through Bedrock from a Lambda function triggered after a Glue crawler run, using the model to generate a description of each updated table's schema.
    2. B.Call a Bedrock foundation model from within the existing Lambda transformation step to classify each incoming support ticket's free text into a topic category before writing the record downstream.
    3. C.Replace AWS Glue crawlers entirely with Amazon Bedrock, since Bedrock foundation models can directly update AWS Glue Data Catalog table definitions without any crawler or Lambda invocation.
    4. D.Configure Amazon Bedrock to run as a Spark worker type inside an AWS Glue job, since Bedrock foundation models execute natively as DPU-based Glue Spark executors.
    5. E.Use Amazon Bedrock exclusively for numeric aggregation of ticket volume by day, since foundation models are the only AWS mechanism capable of counting rows grouped by date value.
    Show answer & explanation

    Correct answers: A, B — Invoke a foundation model through Bedrock from a Lambda function triggered after a Glue crawler run, using the model to generate a description of each updated table's schema.; Call a Bedrock foundation model from within the existing Lambda transformation step to classify each incoming support ticket's free text into a topic category before writing the record downstream.

    • A. Invoking a Bedrock foundation model from a Lambda function triggered after a crawler run lets the model read the updated schema and generate a descriptive summary, which is a documented pattern for enriching the Glue Data Catalog with generative AI metadata.
    • B. Calling a Bedrock foundation model from inside an existing Lambda transformation step to classify free-text ticket content fits naturally into a pipeline that already processes each record with Lambda, adding topic classification without a separate service.
    • C. Bedrock foundation models generate text and other content; they do not directly write to or replace the Glue Data Catalog's crawler-based schema discovery mechanism.
    • D. Bedrock is a managed API for invoking foundation models over HTTPS; it is not a Glue Spark worker type and cannot run as a DPU-based executor inside a Glue job.
    • E. Simple numeric aggregation such as counting rows by date is a standard SQL or Spark operation; a foundation model is not required and is not the only AWS mechanism capable of that calculation.

    Subdomain 1.1: Perform data ingestion

    5.A data engineer is deciding how to trigger AWS Glue crawlers and jobs for two different pipelines: pipeline A must catalog new S3 partitions exactly once every hour regardless of activity, and pipeline B must start processing within seconds whenever a new file lands in S3, at unpredictable times throughout the day. Which combination of triggers correctly matches each pipeline? (Select two.)(Select 2)

    1. A.For pipeline A, configure the crawler with a time-based schedule, such as an hourly cron expression, so it runs once every hour regardless of file activity.
    2. B.For pipeline B, configure an S3 Event Notification (or an EventBridge rule matching S3 object-created events) that invokes the job as soon as a new file arrives.
    3. C.For pipeline A, configure an S3 Event Notification so the crawler only runs when a new object happens to be created in the bucket.
    4. D.For pipeline B, configure the job with a fixed hourly schedule so it only checks for newly arrived files once every hour instead of reacting to them immediately.
    5. E.For both pipelines, configure only a manual on-demand trigger so an operator must always start every run by hand from the console.
    Show answer & explanation

    Correct answers: A, B — For pipeline A, configure the crawler with a time-based schedule, such as an hourly cron expression, so it runs once every hour regardless of file activity.; For pipeline B, configure an S3 Event Notification (or an EventBridge rule matching S3 object-created events) that invokes the job as soon as a new file arrives.

    • A. A fixed hourly schedule directly matches pipeline A's requirement to catalog partitions once every hour regardless of whether or when files actually arrive.
    • B. An S3 Event Notification or an EventBridge rule on object-created events fires immediately when a new file lands, which matches pipeline B's need to start processing within seconds of unpredictable arrivals.
    • C. An event-driven trigger would run the crawler only when an object happens to be created, which would not guarantee the fixed once-per-hour cadence pipeline A requires, especially during periods with no new files.
    • D. A fixed hourly check would introduce up to an hour of delay before new files are noticed, which fails pipeline B's requirement to react within seconds of an unpredictable file arrival.
    • E. Manual on-demand triggers require a person to start every run, which meets neither pipeline A's unattended hourly cadence nor pipeline B's need for immediate, automatic reaction to new files.

    Subdomain 1.1: Perform data ingestion

    6.A data engineer connects a Lambda function to a Kinesis data stream and leaves the event source mapping at its default settings. Which statement correctly describes the default behavior of this Lambda-Kinesis integration?

    1. A.Lambda polls each shard as a standard iterator consumer, sharing that shard's read throughput with any other consumer applications reading the same shard.
    2. B.Lambda automatically registers itself as an enhanced fan-out consumer with a dedicated read throughput connection to every shard by default.
    3. C.Lambda invokes the function once for every individual record as soon as that single record arrives, never grouping records into batches.
    4. D.Lambda requires the existing data stream to be fully deleted and recreated with a special Lambda-only stream type before any event source mapping can be created.
    Show answer & explanation

    Correct answer: A — Lambda polls each shard as a standard iterator consumer, sharing that shard's read throughput with any other consumer applications reading the same shard.

    • A. Without configuring enhanced fan-out, the default Lambda-Kinesis event source mapping uses the standard iterator, polling each shard for records and sharing that shard's read throughput with any other consumer applications on the stream.
    • B. Enhanced fan-out with a dedicated per-shard connection must be explicitly registered as a stream consumer and referenced in the event source mapping; it is not the default behavior.
    • C. By default Lambda reads records in batches from each shard and invokes the function per batch rather than per individual record, which is more efficient than one invocation per record.
    • D. Lambda's Kinesis event source mapping works with any standard Kinesis data stream; there is no special Lambda-only stream type or requirement to recreate the stream.

    Subdomain 1.3: Orchestrate data pipelines

    7.A company wants its ETL orchestration layer to require zero server patching or capacity planning from the data engineering team while still supporting complex branching and long-running, multi-hour workflows coordinating Glue jobs, Lambda functions, and manual approval steps. Which statement correctly identifies a fully serverless orchestration option that meets this requirement?

    1. A.Glue triggers alone are the correct choice here because they natively support multi-hour executions, complex conditional branching, and pausing for human approvals from staff
    2. B.Step Functions Standard Workflows are fully managed with no servers to provision, support year-long executions, and integrate natively with Glue, Lambda, and approval tokens
    3. C.EMR on EC2 with a self-managed Airflow install on the master node is serverless because clusters can be set to auto-terminate once each job run completes
    4. D.An EC2 Auto Scaling group running a custom Python scheduler is serverless as long as the minimum instance count is set to zero when no jobs are running
    Show answer & explanation

    Correct answer: B — Step Functions Standard Workflows are fully managed with no servers to provision, support year-long executions, and integrate natively with Glue, Lambda, and approval tokens

    • A. Glue triggers can start jobs and crawlers on a schedule, on demand, or on an event, but they provide no Choice-state branching logic and no mechanism to pause and wait for human approval.
    • B. Step Functions Standard Workflows require no infrastructure provisioning from the team, support executions lasting up to a year, and have native integrations for Glue, Lambda, and human approval through waitForTaskToken.
    • C. EMR on EC2 provisions actual EC2 instances the team is responsible for patching and sizing, and installing Airflow on the master node still means managing that node's software lifecycle, which is not serverless.
    • D. An Auto Scaling group still consists of managed EC2 instances the team must patch and configure when scaled up; a zero minimum reduces idle cost but does not make the underlying compute serverless.

    Subdomain 1.3: Orchestrate data pipelines

    8.A pipeline architect must design orchestration for a data platform that spans three AWS accounts belonging to different business units, where each unit's DAG occasionally needs to trigger a downstream DAG owned by another unit once its own upstream data assets are ready, without the units tightly coupling their schedules together. Which Amazon MWAA capability directly supports this cross-account, asset-driven triggering pattern?

    1. A.Configuring every DAG's schedule_interval to the same cron expression, since identical schedules guarantee that upstream data is always ready beforehand
    2. B.Asset Watchers, which let a producer DAG publish asset events that trigger dependent workflows in downstream MWAA environments, including across separate accounts
    3. C.Increasing the MWAA environment class to the largest available size, since a bigger environment class alone enables DAGs in one account to trigger DAGs in another account
    4. D.Disabling DAG catchup across all three environments, since disabling catchup is the mechanism MWAA uses to coordinate cross-account asset dependencies
    Show answer & explanation

    Correct answer: B — Asset Watchers, which let a producer DAG publish asset events that trigger dependent workflows in downstream MWAA environments, including across separate accounts

    • A. Matching cron schedules across units only guarantees DAGs start at the same time, not that upstream data is actually ready, and creates no real dependency between the units' workflows.
    • B. Asset Watchers let a producer DAG publish events about data asset readiness that trigger dependent workflows in downstream MWAA environments, including in other accounts, giving decoupled, event-driven cross-account triggering.
    • C. The MWAA environment class controls compute capacity for DAGs within a single environment; it has no bearing on whether one account's environment can trigger a DAG in a separate account's environment.
    • D. The catchup setting controls whether missed past DAG runs are backfilled when a DAG is unpaused; it has no role in coordinating asset-driven triggering between DAGs in different accounts.

    Subdomain 1.3: Orchestrate data pipelines

    9.A pipeline's Step Functions state machine invokes a Lambda function using the .sync-equivalent callback pattern so a long-running external ETL system can report completion asynchronously. If the external system never calls SendTaskSuccess or SendTaskFailure because of an unexpected crash, which built-in Step Functions mechanism prevents the state machine execution from waiting indefinitely?

    1. A.The external system's IAM role expiring is what ends the wait, since IAM credential expiration is the only mechanism able to terminate a waitForTaskToken state
    2. B.Configuring a TimeoutSeconds or HeartbeatSeconds value on the task, so the machine fails the task with a States.Timeout error if no response arrives in the window
    3. C.Step Functions automatically cancels any waitForTaskToken execution after exactly sixty seconds with no configuration required, regardless of the nature of the external system
    4. D.Step Functions has no way to bound a waitForTaskToken wait, so the team must build a separate external polling function to detect the stuck execution
    Show answer & explanation

    Correct answer: B — Configuring a TimeoutSeconds or HeartbeatSeconds value on the task, so the machine fails the task with a States.Timeout error if no response arrives in the window

    • A. IAM credential or role expiration is unrelated to how Step Functions bounds a waitForTaskToken wait; the timeout behavior is governed by the task's own Timeout and Heartbeat configuration.
    • B. Task states support a TimeoutSeconds field bounding total wait time and a HeartbeatSeconds field requiring periodic heartbeats; if neither arrives in the configured window, Step Functions fails the task with States.Timeout.
    • C. There is no automatic, unconfigurable sixty-second cutoff on waitForTaskToken tasks; the wait duration is governed by the explicitly configured TimeoutSeconds and HeartbeatSeconds fields.
    • D. Step Functions does provide native TimeoutSeconds and HeartbeatSeconds fields specifically to bound these waits, so building a separate external polling mechanism is unnecessary.

    Domain 2: Data Store Management

    Subdomain 2.1: Choose a data store

    10.A healthcare partner requires that inbound claims files be digitally signed, encrypted in transit, and delivered with built-in support for message integrity checks and receipt acknowledgments as part of a compliance-driven B2B exchange. Which AWS Transfer Family protocol should the data engineer configure?

    1. A.Configure an AS2 endpoint, which is built for compliance-oriented B2B exchanges with signing, encryption, and message-disposition receipts.
    2. B.Configure a plain FTP endpoint, since FTP is the simplest legacy protocol available for exchanging unencrypted files with external partners.
    3. C.Configure an SFTP endpoint and rely solely on the partner's SSH key pair to satisfy the compliance requirement without AS2.
    4. D.Configure browser-based transfers so business users can manually upload claims files through a web console each business day.
    Show answer & explanation

    Correct answer: A — Configure an AS2 endpoint, which is built for compliance-oriented B2B exchanges with signing, encryption, and message-disposition receipts.

    • Configure an AS2 endpoint, which is built for compliance-oriented B2B exchanges with signing, encryption, and message-disposition receipts.. AS2 is specifically designed for regulated B2B data exchange, natively supporting digital signatures, encryption, and message-disposition acknowledgments that confirm receipt and integrity.
    • Configure a plain FTP endpoint, since FTP is the simplest legacy protocol available for exchanging unencrypted files with external partners.. Plain FTP transmits data unencrypted and has no built-in signing or receipt mechanism, so it cannot satisfy the compliance and integrity requirements described.
    • Configure an SFTP endpoint and rely solely on the partner's SSH key pair to satisfy the compliance requirement without AS2.. SFTP secures the transport channel with SSH but does not provide message-level signing or delivery receipts the way AS2 does for B2B compliance workflows.
    • Configure browser-based transfers so business users can manually upload claims files through a web console each business day.. Browser-based transfers are meant for ad hoc human uploads to S3 and do not provide the automated signing, encryption, and receipt features required for this B2B exchange.

    Subdomain 2.1: Choose a data store

    11.A dashboard in Amazon QuickSight repeatedly runs the same expensive multi-join aggregate query against Redshift fact tables, and stakeholders need the dashboard to load quickly even though the underlying data only changes a few times per day. Which Redshift feature should the data engineer configure?

    1. A.Create a materialized view over the aggregate query and enable automatic refresh so it stays current as base tables change.
    2. B.Create a regular view over the aggregate query so Redshift always recomputes the join and aggregation on each dashboard load.
    3. C.Increase the cluster's concurrency scaling limit so more dashboard queries can run in parallel against the base tables.
    4. D.Add a sort key to each fact table so the same multi-join aggregate scans fewer blocks on every dashboard load.
    Show answer & explanation

    Correct answer: A — Create a materialized view over the aggregate query and enable automatic refresh so it stays current as base tables change.

    • Create a materialized view over the aggregate query and enable automatic refresh so it stays current as base tables change.. A materialized view stores the precomputed join and aggregate result, and automatic refresh keeps it in sync with the infrequent base-table updates, making dashboard loads fast.
    • Create a regular view over the aggregate query so Redshift always recomputes the join and aggregation on each dashboard load.. A regular view is just stored SQL; it still runs the full expensive join and aggregation every time the dashboard queries it, so load time does not improve.
    • Increase the cluster's concurrency scaling limit so more dashboard queries can run in parallel against the base tables.. Concurrency scaling adds capacity for handling more simultaneous queries but does not reduce the cost of each individual expensive aggregate query.
    • Add a sort key to each fact table so the same multi-join aggregate scans fewer blocks on every dashboard load.. A sort key can reduce scanned blocks somewhat, but the query still recomputes the full multi-table join and aggregation on every load rather than reusing a precomputed result.

    Subdomain 2.2: Understand data cataloging systems

    12.A Glue ETL job reads source data with from_catalog, transforms it, and writes the result to S3 partitioned by region and load_date using write_dynamic_frame with partitionKeys set. Currently, analysts cannot query the newly written partitions in Athena until a separate crawler runs afterward, which delays availability by hours. The team wants the ETL job itself to register new partitions in the Data Catalog as part of the write, without adding a second crawler step. Which change to the job accomplishes this?

    1. A.Set the job's catalog update options, such as enableUpdateCatalog and the target database and table, on the sink so the job registers new partitions as it writes.
    2. B.Increase the number of workers assigned to the Glue job so it finishes writing to S3 faster, which causes the Data Catalog to detect the new partitions sooner on its own.
    3. C.Call from_options instead of from_catalog when reading the source data, since from_options automatically propagates any newly written partitions back to the catalog.
    4. D.Add an Athena MSCK REPAIR TABLE statement to the S3 bucket's event notification configuration so partitions register automatically whenever new objects land.
    Show answer & explanation

    Correct answer: A — Set the job's catalog update options, such as enableUpdateCatalog and the target database and table, on the sink so the job registers new partitions as it writes.

    • A. Enabling catalog updates on the sink, along with specifying the target database and table, makes the Glue job itself create or update partition metadata in the Data Catalog as it writes output, removing the need for a separate crawler pass.
    • B. Worker count controls parallelism for the transform and write stages of the job, not whether or when the Data Catalog is notified about new partitions, so scaling workers does not solve the metadata registration gap.
    • C. from_options reads data without consulting the catalog at all and has no bearing on write-side catalog registration; it is unrelated to whether new output partitions get registered after a write.
    • D. S3 event notifications can trigger Lambda functions or other targets, but they cannot directly embed an Athena DDL statement, and this is not a supported mechanism for catalog partition synchronization.

    Subdomain 2.2: Understand data cataloging systems

    13.A team wants an AWS Glue crawler to catalog an existing Amazon DynamoDB table so that the table's item structure becomes queryable through Athena federated queries. When configuring the crawler's data source, which detail is specific to crawling a DynamoDB table rather than an S3-based data store?

    1. A.The crawler is pointed directly at the DynamoDB table name rather than an S3 path or JDBC connection, since DynamoDB is added as its own native crawler data source type.
    2. B.The crawler requires a custom classifier written specifically for DynamoDB's binary item format, because DynamoDB tables cannot be classified using any of the built-in classifiers.
    3. C.The crawler must first export the DynamoDB table to S3 using AWS Data Pipeline, since Glue crawlers can only read from S3 prefixes and JDBC-compatible connections.
    4. D.The crawler needs a VPC-based JDBC connection configured with the DynamoDB table's endpoint URL and port number, identical to how it would connect to an RDS database.
    Show answer & explanation

    Correct answer: A — The crawler is pointed directly at the DynamoDB table name rather than an S3 path or JDBC connection, since DynamoDB is added as its own native crawler data source type.

    • A. DynamoDB is one of the native data source types a Glue crawler can target directly by table name, without needing a JDBC connection string or an S3 path, which is what distinguishes it from file-based or relational sources.
    • B. Glue infers a schema from a sample of DynamoDB items using its built-in handling for that data source type; it does not require a hand-written custom classifier just to read a DynamoDB table's structure.
    • C. Exporting to S3 first would be a workaround, not a requirement; Glue crawlers support DynamoDB as a first-class data source and can read from the table directly.
    • D. DynamoDB is accessed through the AWS API rather than a JDBC endpoint, so configuring a JDBC connection with a host and port, as would be done for RDS, does not apply to DynamoDB crawling.

    Subdomain 2.3: Manage the lifecycle of data

    14.A company applies an S3 Lifecycle rule that transitions objects tagged `project=archive` to S3 Glacier Flexible Retrieval 60 days after creation, then expires them at 365 days. A batch job runs monthly to remove the `project=archive` tag from objects that turn out to still be needed, and the team wants a reliable way to confirm that removing a tag actually prevented the scheduled transition before they discard their own record of which objects were tagged. Which two statements correctly describe how to handle this? (Select 2)(Select 2)

    1. A.S3 re-evaluates the object's current tags when each lifecycle action executes, so a tag removed before the action runs means it no longer matches
    2. B.To reliably confirm an expiration or transition already ran before removing a tag, the team should wait for an S3 Lifecycle Event Notification first
    3. C.S3 caches the tag value from when the object was first created, so removing the tag later has no effect on the scheduled action
    4. D.Removing the tag immediately halts any lifecycle action already in progress, including a transition that started earlier the same day
    5. E.Tag-based rules are evaluated only once, at rule creation time, so objects uploaded after tags are removed remain permanently exempt
    Show answer & explanation

    Correct answers: A, B — S3 re-evaluates the object's current tags when each lifecycle action executes, so a tag removed before the action runs means it no longer matches; To reliably confirm an expiration or transition already ran before removing a tag, the team should wait for an S3 Lifecycle Event Notification first

    • A. S3 evaluates tag-based filters against an object's current tags daily and again at the moment a queued action executes, so if the triggering tag was removed before the transition or expiration actually runs, the object no longer matches and the action does not proceed.
    • B. AWS documentation recommends waiting for a Lifecycle Event Notification to reliably confirm that an expiration or transition already executed before removing the triggering tag, since tag changes and the daily evaluation cycle make timing assumptions unreliable otherwise.
    • C. S3 does not cache tag values at creation time; the tag filter is a live condition re-checked before each action executes, so a tag change made any time before that point is honored.
    • D. Once a lifecycle action has been queued for asynchronous processing and begins executing, removing a tag does not interrupt an action that is already underway; the re-evaluation happens before an action is applied, not mid-execution.
    • E. Lifecycle rules apply continuously to both existing and newly created objects that match the filter at evaluation time, not just once when the rule was authored, so the rule keeps applying to new uploads going forward.

    Subdomain 2.3: Manage the lifecycle of data

    15.An engineer configures an S3 Lifecycle rule intending to first transition objects to S3 Standard-IA at 30 days and then to S3 Glacier Flexible Retrieval at 45 days. After reviewing S3's documented rules for chaining storage class transitions, which statement is accurate?

    1. A.This configuration is invalid because S3 requires a minimum gap between certain storage class transitions, and 15 days between Standard-IA and Glacier does not satisfy it
    2. B.This configuration is valid because S3 allows any number of chained transitions between storage classes with no minimum gap requirement whatsoever
    3. C.This configuration is invalid because an object can never transition through more than one storage class within a single lifecycle rule at all
    4. D.This configuration is valid, but S3 will silently ignore the Standard-IA transition and send the objects directly to Glacier at 45 days instead
    Show answer & explanation

    Correct answer: A — This configuration is invalid because S3 requires a minimum gap between certain storage class transitions, and 15 days between Standard-IA and Glacier does not satisfy it

    • A. S3 Lifecycle enforces minimum transition timing between certain storage class pairs, and Amazon's documented guidance requires at least 30 days between a Standard-IA (or One Zone-IA) transition and a subsequent Glacier Flexible Retrieval transition, so a 15-day gap violates that constraint and the configuration would be rejected or corrected.
    • B. S3 does not allow arbitrary transition timing between every storage class pair; documented minimum-gap constraints exist specifically for transitions into and out of the IA classes, so 'no minimum gap requirement' misstates the documented behavior.
    • C. A single S3 Lifecycle rule can define multiple ordered transition actions that chain an object through several storage classes over its lifetime; splitting each transition into a separate rule is not a requirement.
    • D. S3 does not silently skip a configured transition action; an invalid chained-transition configuration is rejected or must be corrected rather than partially applied by ignoring one of the two transitions.

    Subdomain 2.4: Design data models and schema evolution

    16.The `region_lookup` dimension table in a Redshift cluster has only 40 rows and is joined against several large fact tables, including `sales_fact` and `inventory_fact`, which use KEY distribution on different columns. Which distribution style keeps the join local on every node without redistributing `region_lookup` at query time?

    1. A.DISTSTYLE ALL, which stores a full copy of the small table on every compute node so no join ever needs to redistribute rows
    2. B.DISTSTYLE KEY on the region_id column, matching the distribution key used by exactly one of the two fact tables
    3. C.DISTSTYLE EVEN, which spreads the forty rows evenly across node slices using a round-robin placement algorithm
    4. D.DISTSTYLE AUTO with a nightly VACUUM REEVALUATE command that manually redistributes rows across every node slice each night
    Show answer & explanation

    Correct answer: A — DISTSTYLE ALL, which stores a full copy of the small table on every compute node so no join ever needs to redistribute rows

    • A. A full copy on every node means whichever fact table's distribution key is used, the matching lookup row is always present locally, so no redistribution is needed for either join.
    • B. Matching the key of one fact table only solves collocation for that table; the join against the other fact table, which distributes on a different column, would still require redistribution.
    • C. Even distribution spreads rows without regard to any join column, so both fact table joins would still need to redistribute the lookup rows across the network.
    • D. There is no VACUUM REEVALUATE command in Redshift, and manual reshuffling under AUTO distribution does not guarantee collocation with two differently distributed fact tables.

    Subdomain 2.4: Design data models and schema evolution

    17.A model risk review requires the data science team to determine exactly which training dataset, preprocessing job, and feature transformations produced a specific model currently deployed to a SageMaker endpoint. Which AWS capability directly provides this traceability?

    1. A.Amazon SageMaker ML Lineage Tracking, which records associations between datasets, processing jobs, training jobs, and models
    2. B.Amazon CloudWatch Logs Insights, which stores the raw console output text produced during each SageMaker training job run
    3. C.AWS CloudTrail, which records the API calls made against the SageMaker service but not the specific artifacts those calls operated on
    4. D.Amazon S3 Object Versioning on the training data bucket, which retains prior versions of files after they are overwritten
    Show answer & explanation

    Correct answer: A — Amazon SageMaker ML Lineage Tracking, which records associations between datasets, processing jobs, training jobs, and models

    • A. ML Lineage Tracking is purpose-built to capture the graph of artifacts and their associations across an ML workflow, letting a reviewer query from a deployed model back through training jobs to the exact dataset and transformations used.
    • B. CloudWatch Logs Insights can search log text emitted during a job, but it does not model the structured relationships between datasets, jobs, and models that a lineage query needs to trace directly.
    • C. CloudTrail records who called which API and when, which is useful for auditing actions, but it does not track the artifact-level relationships between a dataset, a training job, and a resulting model.
    • D. S3 Versioning preserves prior file contents in a bucket, but on its own it does not connect a specific dataset version to the training job and model that consumed it.

    Domain 3: Data Operations and Support

    Subdomain 3.1: Automate data processing by using AWS services

    18.A data platform team needs a nightly Glue job to start at 2 AM in a specific customer time zone, with the exact invocation payload varying by customer, across hundreds of customer-specific schedules that change occasionally as customers are added or removed. Which approach is best suited to managing this at scale?

    1. A.Create per-customer schedules using Amazon EventBridge Scheduler, which manages individually recurring schedules with timezone awareness and per-schedule target payloads.
    2. B.Create a single EventBridge rule with one fixed cron expression in UTC, and have the Glue job itself compute each customer's local time before deciding whether to run at all.
    3. C.Create one Lambda function containing a loop for every customer's schedule, and invoke that Lambda from an MWAA DAG that runs once a day at midnight UTC.
    4. D.Create a single Step Functions Standard workflow with one Wait state per customer, running continuously for a year to reach every customer's execution time.
    Show answer & explanation

    Correct answer: A — Create per-customer schedules using Amazon EventBridge Scheduler, which manages individually recurring schedules with timezone awareness and per-schedule target payloads.

    • A. EventBridge Scheduler is designed to manage large numbers of individual, timezone-aware schedules with per-schedule target payloads, which fits hundreds of customer-specific nightly invocations that change over time.
    • B. A single UTC-only rule pushes all the timezone and per-customer payload logic into the Glue job itself, which is exactly the per-schedule management work EventBridge Scheduler already handles natively.
    • C. Hardcoding hundreds of customer schedules inside one Lambda's loop, invoked once daily by an Airflow DAG, creates a single point of maintenance and does not honor individually varying customer times.
    • D. A long-running Standard workflow parked in Wait states for up to a year per customer is an inefficient, hard-to-manage substitute for a purpose-built scheduling service.

    Subdomain 3.1: Automate data processing by using AWS services

    19.A team wants a Lambda function to run whenever any AWS Glue job in the account transitions to a `FAILED` state, so an on-call engineer is paged automatically, without the Lambda function polling the Glue API. Which approach should they configure? (Select TWO.)(Select 2)

    1. A.Create an EventBridge rule with an event pattern matching Glue Job State Change events where the state is `FAILED`, and set the Lambda function as its target.
    2. B.Configure the rule's target with an on-failure destination or dead-letter queue so the alert is retained if the Lambda invocation itself later fails.
    3. C.Have the Lambda function call `get_job_runs` for every Glue job in the account once a minute and compare each run's state to the previous check.
    4. D.Subscribe the Lambda function directly to AWS CloudTrail management events, since CloudTrail natively filters and delivers only Glue job failure events.
    5. E.Grant the Lambda function's execution role s3:GetObject on every bucket in the account, since that permission is what allows EventBridge to deliver Glue events.
    Show answer & explanation

    Correct answers: A, B — Create an EventBridge rule with an event pattern matching Glue Job State Change events where the state is `FAILED`, and set the Lambda function as its target.; Configure the rule's target with an on-failure destination or dead-letter queue so the alert is retained if the Lambda invocation itself later fails.

    • A. Glue emits job state change events to EventBridge, so a rule with an event pattern matching `FAILED` states and the Lambda function as its target reacts immediately without any polling.
    • B. Adding an on-failure destination or DLQ to the rule's target protects against losing the failure alert if the Lambda invocation itself errors, keeping the page reliable end to end.
    • C. Polling `get_job_runs` for every job once a minute is exactly the polling approach the team wants to avoid, and it adds latency and unnecessary API calls compared to an event-driven rule.
    • D. CloudTrail records API activity and does not itself filter and deliver targeted service state-change events to a Lambda function the way an EventBridge rule with an event pattern does.
    • E. S3 object permissions have no bearing on whether EventBridge delivers Glue job state change events; that delivery is governed by the rule's event pattern and target permissions, not S3 access.

    Subdomain 3.2: Analyze data by using AWS services

    20.A data engineer needs to run a one-time ad hoc query against Parquet files in S3 to count distinct customers per region for a report due within the hour, and does not want to provision or manage any compute infrastructure. Which AWS service best fits this need?

    1. A.Amazon Athena, running a serverless SQL query directly against the Parquet files through the Glue Data Catalog.
    2. B.Amazon EMR, launching a transient cluster sized for the query and terminating it once the report is generated.
    3. C.Amazon Redshift provisioned clusters, resizing the cluster to a larger node type to handle the ad hoc workload.
    4. D.AWS Glue ETL jobs, writing a Spark script that aggregates the Parquet files and writes the result back to S3.
    Show answer & explanation

    Correct answer: A — Amazon Athena, running a serverless SQL query directly against the Parquet files through the Glue Data Catalog.

    • Amazon Athena, running a serverless SQL query directly against the Parquet files through the Glue Data Catalog.. Athena is a serverless interactive query service that runs standard SQL directly against S3 data registered in the Glue Data Catalog, with no clusters to provision, matching a one-off report needed within the hour.
    • Amazon EMR, launching a transient cluster sized for the query and terminating it once the report is generated.. Even a transient EMR cluster requires choosing an instance type and cluster size and waiting for it to bootstrap, adding provisioning overhead that is unnecessary for a single ad hoc aggregate query.
    • Amazon Redshift provisioned clusters, resizing the cluster to a larger node type to handle the ad hoc workload.. A provisioned Redshift cluster requires the data to already be loaded or spectrum-configured and involves managing fixed compute nodes, which is heavier infrastructure than a single ad hoc query justifies.
    • AWS Glue ETL jobs, writing a Spark script that aggregates the Parquet files and writes the result back to S3.. Writing and deploying a Glue Spark job adds development and job-run overhead that is disproportionate to a simple distinct-count query needed quickly for a report.

    Subdomain 3.2: Analyze data by using AWS services

    21.A data engineer is choosing between running exploratory Apache Spark analysis in an Athena Spark notebook versus launching a long-running Amazon EMR cluster for the same exploration. The workload is bursty: heavy for a few hours a week and idle the rest of the time. Which factor most strongly favors the Athena Spark notebook for this workload?

    1. A.Athena Spark bills only for the compute consumed during active sessions, avoiding the cost of an idle cluster during the many hours with no exploratory workload.
    2. B.Athena Spark supports installing arbitrary custom bootstrap actions and low-level YARN tuning that a managed EMR cluster does not expose to users.
    3. C.Athena Spark stores notebook files directly on cluster-attached HDFS volumes, giving faster local disk access than EMR's S3-backed storage.
    4. D.Athena Spark allows full root SSH access to the underlying worker nodes, which a managed EMR cluster restricts for administrators by default.
    Show answer & explanation

    Correct answer: A — Athena Spark bills only for the compute consumed during active sessions, avoiding the cost of an idle cluster during the many hours with no exploratory workload.

    • Athena Spark bills only for the compute consumed during active sessions, avoiding the cost of an idle cluster during the many hours with no exploratory workload.. Athena Spark is serverless and charges for the compute used during active notebook sessions, so a bursty workload that is idle most of the week avoids paying for cluster capacity that sits unused, unlike a long-running EMR cluster.
    • Athena Spark supports installing arbitrary custom bootstrap actions and low-level YARN tuning that a managed EMR cluster does not expose to users.. It is EMR, not Athena Spark, that exposes bootstrap actions and low-level YARN configuration; Athena Spark trades that fine-grained control for a simpler, serverless session model.
    • Athena Spark stores notebook files directly on cluster-attached HDFS volumes, giving faster local disk access than EMR's S3-backed storage.. Athena Spark is serverless and does not maintain a persistent HDFS cluster for notebook storage, so this describes neither how Athena Spark nor typical EMR-on-S3 storage patterns work.
    • Athena Spark allows full root SSH access to the underlying worker nodes, which a managed EMR cluster restricts for administrators by default.. Athena Spark does not expose SSH access to underlying compute at all, since AWS manages the infrastructure entirely; EMR is the service that can allow SSH access to cluster nodes.

    Subdomain 3.3: Maintain and monitor data pipelines

    22.A team ingests terabytes of daily application logs from dozens of microservices supporting a real-time analytics pipeline. Analysts need full-text search across free-form log messages, custom dashboards with drill-down visualizations, and the ability to retain and search years of historical log data cost-effectively. Which approach best fits these requirements?

    1. A.Stream the logs into an Amazon OpenSearch Service domain sized for the retention and query load, and build OpenSearch Dashboards for full-text search and drill-down visualization.
    2. B.Keep all logs in CloudWatch Logs and rely exclusively on CloudWatch Logs Insights queries, since Logs Insights already provides unlimited retention and long-term dashboarding out of the box.
    3. C.Move the logs into Amazon DynamoDB with a single partition key per microservice, then use DynamoDB Streams to build custom full-text search logic in a downstream Lambda function.
    4. D.Load the logs into an Amazon Redshift provisioned cluster and build materialized views that analysts query with SQL to perform free-text search across log messages.
    Show answer & explanation

    Correct answer: A — Stream the logs into an Amazon OpenSearch Service domain sized for the retention and query load, and build OpenSearch Dashboards for full-text search and drill-down visualization.

    • A. OpenSearch Service is purpose-built for full-text search at scale and pairs with Dashboards for interactive drill-down visualization, and its storage tiers can retain years of data more cost-effectively than keeping everything hot in CloudWatch Logs.
    • B. CloudWatch Logs Insights is well suited to ad hoc queries, but retention is governed by the log group's configured setting rather than unlimited, and it lacks the rich dashboarding this scenario requires.
    • C. DynamoDB is a key-value and document store without native full-text search, so free-form log message search would require substantial custom indexing logic built entirely outside DynamoDB itself.
    • D. Redshift SQL can filter structured text with pattern matching, but it is not designed for full-text search relevance ranking or the interactive drill-down visualization experience analysts need here.

    Subdomain 3.3: Maintain and monitor data pipelines

    23.An engineer needs to write a CloudWatch Logs Insights query against a log group receiving JSON-formatted logs from several microservices, to count errors per `service` field over 5-minute buckets, sorted by count in descending order, and limited to the top 20 results. Which two query pipeline components should the query include to accomplish this? (Select TWO)(Select 3)

    1. A.A `filter` command early in the pipeline that restricts processing to log entries where the `level` field equals `ERROR`, before any aggregation runs.
    2. B.A `stats count(*) by service, bin(5m)` command that groups the filtered entries into 5-minute buckets per service and counts the errors in each bucket.
    3. C.A `parse` command that extracts the `service` field using a regular expression, since CloudWatch Logs Insights cannot read fields directly from JSON-formatted log entries.
    4. D.A `dedup service` command placed before the `stats` command, since Logs Insights requires duplicate service names to be removed before any aggregation can be performed.
    5. E.A `display` command that renders the query results as a line chart directly within the query pipeline, since `sort` and `limit` only work on visualized results.
    6. F.A `sort count desc` and `limit 20` command at the end of the pipeline to order the aggregated results and cap the output to the top 20 buckets.
    Show answer & explanation

    Correct answers: A, B, F — A `filter` command early in the pipeline that restricts processing to log entries where the `level` field equals `ERROR`, before any aggregation runs.; A `stats count(*) by service, bin(5m)` command that groups the filtered entries into 5-minute buckets per service and counts the errors in each bucket.; A `sort count desc` and `limit 20` command at the end of the pipeline to order the aggregated results and cap the output to the top 20 buckets.

    • A. Filtering to ERROR-level entries as early as possible in the pipeline reduces the data volume before the more expensive aggregation step runs, and it directly implements the requirement to count only errors.
    • B. A stats command with a by clause that includes bin(5m) is exactly how Logs Insights groups counts into fixed time buckets per service, matching the 5-minute bucket and per-service breakdown requested.
    • C. Logs Insights automatically discovers fields in JSON-formatted log events without a parse command, so an explicit regular expression extraction is unnecessary for this JSON-based log group.
    • D. Dedup removes duplicate log records matching specified fields; it is not a prerequisite for aggregation and would incorrectly collapse repeated service names that should each contribute to the count.
    • E. Sort and limit are pipeline commands that operate directly on the query's tabular results; visualization is a separate rendering option and is not required for sort or limit to function.
    • F. Sort and limit are the standard commands for ordering aggregated results by count and capping the returned rows to the requested top 20 buckets.

    Subdomain 3.4: Ensure data quality

    24.A DataBrew dataset has five numeric columns — `rate`, `pay`, `bonus`, `increase`, and `deduction` — that must each stay at or below 100. Instead of authoring five nearly identical rules, an analyst wants to define the "value at most 100" check once and apply it across all five columns in a single rule. Which ruleset configuration achieves this?

    1. A.Set the data quality check scope to apply the check to a selected group of columns rather than to an individual column.
    2. B.Duplicate the ruleset five times, once for each column, and run all five rulesets together inside a single profile job.
    3. C.Add a `CustomSql` rule that concatenates the five column names into one SQL expression evaluated as text.
    4. D.Configure five separate profile jobs, each scheduled to run against one of the five numeric columns.
    Show answer & explanation

    Correct answer: A — Set the data quality check scope to apply the check to a selected group of columns rather than to an individual column.

    • A. The check scope setting lets a single check be applied to a group of selected columns at once, which is exactly the mechanism for enforcing the same "at most 100" condition across all five numeric columns without repeating the rule.
    • B. Duplicating the whole ruleset five times still creates five separate rule definitions to maintain, which is the repetitive approach the analyst is trying to avoid.
    • C. Concatenating column names into a single SQL string does not express a per-column numeric comparison and would not correctly validate each column's values.
    • D. Running five separate profile jobs multiplies operational overhead and still evaluates each column independently rather than through one shared rule.

    Subdomain 3.4: Ensure data quality

    25.A team running Spark ETL jobs on AWS Glue 4.0 notices that shuffle stages sometimes produce a few oversized partitions that slow down the whole job, while at other times the same job runs evenly. Rather than manually tuning partition counts or salting keys for every job, the team wants Spark to detect and correct skewed shuffle partitions automatically at runtime based on actual data statistics. Which two statements about the capability that provides this are correct? (Select TWO.)(Select 2)

    1. A.Adaptive Query Execution (AQE) uses runtime statistics gathered during execution to dynamically resize shuffle partitions and split skewed ones.
    2. B.AQE ships enabled by default starting with AWS Glue version 3.0 and continuing in 4.0, with no separate feature flag needed to turn it on.
    3. C.AQE only takes effect during the planning phase before a job starts, relying solely on static statistics gathered ahead of time from the Data Catalog.
    4. D.AQE must be manually configured with a salting key for every join before it can detect or correct any kind of partition-level data skew.
    5. E.AQE replaces the need for a Data Catalog entirely, since it infers table schema directly from statistics gathered at runtime.
    6. F.AQE is exclusive to AWS Glue Data Quality evaluation runs and has no effect at all on standard Spark ETL transform jobs.
    Show answer & explanation

    Correct answers: A, B — Adaptive Query Execution (AQE) uses runtime statistics gathered during execution to dynamically resize shuffle partitions and split skewed ones.; AQE ships enabled by default starting with AWS Glue version 3.0 and continuing in 4.0, with no separate feature flag needed to turn it on.

    • A. AQE reoptimizes the query plan at runtime using actual data statistics gathered during execution, including detecting and splitting skewed shuffle partitions, which is the automatic runtime correction the team wants.
    • B. AQE ships enabled by default from Glue 3.0 onward and continues in 4.0, so teams get this runtime correction without having to opt in through a separate flag.
    • C. AQE's defining feature is that it adapts the plan during execution using runtime statistics, not just static statistics collected before the job starts.
    • D. AQE does not require manual salting configuration; it detects and corrects skew automatically based on runtime task statistics without a developer defining a salt key.
    • E. AQE optimizes query execution plans and does not replace or infer the Data Catalog's schema metadata role.
    • F. AQE is a general Spark SQL execution optimization that applies to standard ETL transform jobs; it is not limited to data quality evaluation runs.

    Domain 4: Data Security and Governance

    Subdomain 4.1: Apply authentication mechanisms

    26.A data engineer configures an Amazon S3 access point so that requests are only accepted when they originate from inside a specific VPC. Which statements correctly describe this configuration? (Select TWO)(Select 2)

    1. A.The access point policy works alongside the underlying bucket policy, and a request must be permitted by both for the operation to succeed.
    2. B.Restricting the access point to a VPC means requests routed through the public internet to that access point's alias are automatically rejected.
    3. C.A VPC-restricted access point replaces the need for any IAM permissions policy on the calling principal, since network origin alone grants access.
    4. D.Once an access point is created, the underlying bucket policy is automatically deleted because the access point becomes the sole entry point for the bucket.
    5. E.Access points only support object-level operations, so they cannot delete the underlying bucket or modify its Amazon S3 Replication configuration.
    Show answer & explanation

    Correct answers: A, B — The access point policy works alongside the underlying bucket policy, and a request must be permitted by both for the operation to succeed.; Restricting the access point to a VPC means requests routed through the public internet to that access point's alias are automatically rejected.

    • A. An access point attached to a bucket enforces its own policy in conjunction with the bucket policy, so both must allow the request; the access point narrows access, it does not replace the bucket-level policy.
    • B. A VPC-only access point is configured to accept requests only through the VPC endpoint tied to that VPC, so connection attempts arriving over the public internet to that access point are rejected regardless of the caller's identity.
    • C. The calling principal still needs an IAM identity policy that allows the S3 action on the access point; network origin restriction is an additional layer, not a substitute for identity-based authorization.
    • D. Creating an access point does not delete or disable the bucket policy; the bucket policy continues to apply to direct bucket requests and to requests made through any access point.
    • E. Access points only support object-level operations such as `GetObject` and `PutObject`; bucket-level actions like deleting the bucket or changing its replication configuration are outside what an access point can be used for.

    Subdomain 4.1: Apply authentication mechanisms

    27.Which AWS service is best described as an unmanaged option where the customer is responsible for patching the operating system and the database engine, in contrast to a managed alternative that handles those tasks automatically?

    1. A.Self-managed database on EC2
    2. B.Amazon RDS for PostgreSQL
    3. C.Amazon Redshift Serverless
    4. D.AWS Glue
    Show answer & explanation

    Correct answer: A — Self-managed database on EC2

    • A. A database engine installed directly on an EC2 instance is unmanaged: the customer owns operating system patching, engine upgrades, and backup configuration, since AWS only manages the underlying virtualization layer.
    • B. Amazon RDS is a managed relational database service where AWS handles engine patching, automated backups, and failover, which is the opposite of the unmanaged responsibility split described in the stem.
    • C. Redshift Serverless is a fully managed, serverless data warehouse where AWS provisions and scales compute automatically, so the customer has no operating system or engine patching responsibility.
    • D. AWS Glue is a fully managed ETL service where AWS provisions and manages the underlying Spark or Python execution environment, leaving the customer with no infrastructure or engine patching to perform.

    Subdomain 4.2: Apply authorization mechanisms

    28.A company runs a nightly ETL job on Amazon EC2 that connects to an Amazon RDS for PostgreSQL database using a hard-coded username and password stored in the application's configuration file. Security flags this as a risk and asks the data engineer to remove the hard-coded credentials and enable periodic credential rotation without requiring application redeployment when the password changes. Which two actions should the data engineer take? (Choose 2.)(Select 2)

    1. A.Store the database credentials in AWS Secrets Manager and configure automatic rotation using the built-in RDS rotation function so the password is updated on a schedule.
    2. B.Move the credentials into an SSM Parameter Store SecureString parameter and have the job re-read the parameter on every scheduled run without configuring any rotation.
    3. C.Update the ETL job's code to call the Secrets Manager get-secret-value API at runtime instead of reading static credentials from the configuration file.
    4. D.Embed the credentials in an EC2 instance user-data script so they are injected into the environment at boot instead of stored in the configuration file.
    5. E.Grant the EC2 instance's IAM role the RDS master password directly through an inline policy so the application can authenticate without relying on any external secret store.
    Show answer & explanation

    Correct answers: A, C — Store the database credentials in AWS Secrets Manager and configure automatic rotation using the built-in RDS rotation function so the password is updated on a schedule.; Update the ETL job's code to call the Secrets Manager get-secret-value API at runtime instead of reading static credentials from the configuration file.

    • A. Secrets Manager natively integrates with RDS to rotate database passwords on a schedule using a managed Lambda rotation function, which directly satisfies the rotation requirement.
    • B. This removes the hard-coded value from the config file, but Parameter Store alone does not rotate the password, so the static password is simply relocated rather than rotated.
    • C. Retrieving the secret dynamically at runtime means the job always picks up the latest rotated password, removing the hard-coded value without needing an application redeploy.
    • D. User-data scripts are visible to anyone with EC2 describe permissions on the instance and provide no rotation capability, leaving a static, exposed credential in place.
    • E. IAM policies do not carry database passwords as a permission grant, and this still leaves a static credential outside a managed secret store with no rotation configured for it.

    Subdomain 4.2: Apply authorization mechanisms

    29.A company organizes its AWS resources with a project tag whose value matches one of several active data engineering initiatives. Leadership wants engineers to automatically gain permissions on any S3 bucket and Glue database tagged with the same project they are assigned to, without editing IAM policies whenever a new project starts or an engineer changes projects. Which two elements should the data engineer combine to implement this attribute-based access control (ABAC) pattern? (Choose 2.)(Select 2)

    1. A.Tag each engineer's IAM role with a project principal tag matching their current assignment, updating only that tag value when the engineer switches projects.
    2. B.Write an IAM policy condition comparing the principal's project tag against the resource's project tag so access is allowed only when the two values match.
    3. C.Create a customer managed policy naming the explicit ARNs of every bucket and database per project, attaching the matching set whenever an engineer changes projects.
    4. D.Configure organization-wide service control policies that deny all Glue and S3 actions account-wide, then permit exceptions per engineer through individual resource-based policies.
    5. E.Grant every engineer the AWS managed PowerUserAccess policy and depend on config rules to flag access outside an assigned project only after it has occurred.
    Show answer & explanation

    Correct answers: A, B — Tag each engineer's IAM role with a project principal tag matching their current assignment, updating only that tag value when the engineer switches projects.; Write an IAM policy condition comparing the principal's project tag against the resource's project tag so access is allowed only when the two values match.

    • A. A principal tag on the role is the attribute that ABAC conditions compare against a resource tag, and changing the value on a project switch never touches the IAM policy.
    • B. This condition is the mechanism that makes ABAC work: a single policy grants access dynamically whenever the principal's tag matches the resource's tag, with no per-project policy edit.
    • C. Naming explicit ARNs per project is the role-based, named-resource pattern ABAC is meant to replace, and it still requires a policy change on every reassignment.
    • D. Service control policies set permission boundaries across accounts but do not themselves perform tag-matching authorization, and per-engineer resource policies still need manual updates as assignments change.
    • E. PowerUserAccess grants broad permissions across nearly all services regardless of project, and config rules only report violations after access has already occurred, not at request time.

    Subdomain 4.3: Ensure data encryption and masking

    30.A data engineering team runs nightly batch jobs that connect to an Amazon Redshift cluster over the public internet from an on-premises ETL server. A security audit requires that all connections to the cluster be encrypted in transit, and that unencrypted connection attempts be rejected outright rather than merely discouraged. What should the team configure?

    1. A.Set the cluster's parameter group `require_ssl` parameter to true, which causes Redshift to reject any client connection that does not negotiate SSL/TLS before it can run queries.
    2. B.Instruct the ETL server's database driver to prefer SSL when available, so connections use encryption whenever both sides support it but silently fall back to a plaintext connection otherwise.
    3. C.Move the on-premises ETL server into the same VPC as the Redshift cluster using a Site-to-Site VPN, since traffic within a VPC is automatically encrypted by AWS at the network layer.
    4. D.Enable encryption at rest for the Redshift cluster's underlying storage volumes, since data encrypted at rest is also protected while it travels between the client and the cluster.
    Show answer & explanation

    Correct answer: A — Set the cluster's parameter group `require_ssl` parameter to true, which causes Redshift to reject any client connection that does not negotiate SSL/TLS before it can run queries.

    • A. Setting `require_ssl` to true in the cluster's parameter group makes Redshift enforce SSL/TLS on every connection and refuse any client that attempts to connect without it, matching the requirement to reject unencrypted attempts.
    • B. A driver-side preference for SSL negotiates encryption opportunistically but still permits a plaintext connection when the server does not enforce TLS, which does not satisfy a hard requirement to reject unencrypted attempts.
    • C. VPC network traffic is not automatically encrypted at the packet level by default, and a VPN secures the tunnel between networks, not the application-layer connection between the ETL client and the cluster's SSL listener.
    • D. Encryption at rest protects data stored on disk and is a separate control from transport encryption; it has no effect on whether a network connection between the client and cluster is encrypted.

    Subdomain 4.3: Ensure data encryption and masking

    31.A data platform team runs a short-lived AWS Glue job that must decrypt objects protected by a customer managed KMS key for exactly the duration of one ETL run, without permanently modifying the key policy or the Glue job role's IAM policy. Which KMS mechanism, and which of its properties, correctly support this temporary, narrowly scoped access pattern? (Select 2)(Select 2)

    1. A.Use the KMS `CreateGrant` operation to issue a grant for the specific principal and operations needed, since a grant can be programmatically created and later retired or revoked once the job finishes.
    2. B.A grant can be scoped with constraints, such as an encryption context constraint, so the permission it confers applies only to requests matching that specific context, not to every use of the key.
    3. C.Grants permanently modify the key policy document itself, so retiring a grant after the job completes always requires editing and republishing the key policy to remove the corresponding grant statement.
    4. D.A grant is the only way to allow any principal to use a KMS key at all, since key policies and IAM policies cannot independently grant permission to use a customer managed key.
    5. E.Grants can only be created by the AWS account root user through the AWS Management Console, which makes them unsuitable for permissions issued programmatically inside an automated ETL pipeline.
    Show answer & explanation

    Correct answers: A, B — Use the KMS `CreateGrant` operation to issue a grant for the specific principal and operations needed, since a grant can be programmatically created and later retired or revoked once the job finishes.; A grant can be scoped with constraints, such as an encryption context constraint, so the permission it confers applies only to requests matching that specific context, not to every use of the key.

    • A. `CreateGrant` lets a permitted principal programmatically delegate specific key operations to another principal for a bounded purpose, and the grant can be retired or revoked afterward, making it well suited to temporary access during a single job run.
    • B. Grants support constraints such as an encryption context constraint, which narrows the grant so it only authorizes requests carrying that specific context, giving fine-grained scoping beyond a blanket key permission.
    • C. A grant is a separate access control object from the key policy document; creating or retiring a grant does not edit or republish the key policy, which is exactly why grants are convenient for temporary, programmatic access.
    • D. Key policies and IAM policies can and routinely do grant permission to use a KMS key on their own; grants are an additional, more dynamic mechanism layered on top, not the only way to authorize key usage.
    • E. Grants are designed specifically for programmatic delegation and can be created by any principal that itself has grant-creation permission on the key, not solely the root user through the console, which is what makes them practical for automated pipelines.

    Subdomain 4.4: Prepare logs for audit

    32.A data engineering team already has an organization CloudTrail trail delivering log files to a centralized Amazon S3 bucket, partitioned by account, Region, and date. Auditors need to run one-off SQL queries this quarter to find every `DeleteBucket` call, without provisioning a new managed logging service or ingesting the data elsewhere. Which approach satisfies this with the least new infrastructure?

    1. A.Define an Amazon Athena table over the existing S3 prefix using partition projection so queries can filter on the account, Region, and date partitions.
    2. B.Provision an Amazon OpenSearch Service domain and write a custom ingestion pipeline to reindex every historical log file from the S3 bucket before searching.
    3. C.Launch a persistent Amazon EMR cluster configured with Hive to reprocess the full history of log files nightly into a new internal table format for querying.
    4. D.Copy every log file out of the centralized bucket into a new CloudTrail Lake event data store so the existing partition structure can be queried with SQL.
    Show answer & explanation

    Correct answer: A — Define an Amazon Athena table over the existing S3 prefix using partition projection so queries can filter on the account, Region, and date partitions.

    • A. Athena can query CloudTrail JSON log files in place using `CREATE EXTERNAL TABLE` with partition projection over the account, Region, and date structure already in the S3 path, requiring no data movement or new managed service.
    • B. Standing up an OpenSearch domain and a custom reindexing pipeline is significant new infrastructure for a one-off quarterly query need, and it duplicates data that Athena could already query directly from S3.
    • C. A persistent EMR cluster with nightly Hive reprocessing is far more operational overhead than a one-off SQL query requires, and it leaves a cluster running that must be managed and paid for continuously.
    • D. CloudTrail Lake event data stores ingest events directly from CloudTrail or via copy-trail operations, not by manually copying arbitrary S3 objects, and building this pipeline is unnecessary new infrastructure for a single query.

    Subdomain 4.4: Prepare logs for audit

    33.A team is deciding whether to adopt AWS CloudTrail Lake instead of building and maintaining their own Amazon Athena and AWS Glue Data Catalog pipeline over CloudTrail logs in Amazon S3. Which statements accurately describe CloudTrail Lake's properties? (Select THREE.)(Select 3)

    1. A.It lets you run SQL queries across multiple event data stores spanning several accounts and Regions in the CloudTrail console, without configuring Athena tables.
    2. B.It converts ingested events into columnar Apache ORC storage, so ad hoc SQL queries typically scan less data than an equivalent query over row-based JSON files.
    3. C.It supports configurable retention of up to seven years on an event data store, meeting long compliance retention windows without a separate S3 lifecycle policy.
    4. D.It stores events exclusively inside an S3 bucket that you create and manage yourself, giving you direct control over the bucket's encryption keys and lifecycle rules.
    5. E.Running queries in CloudTrail Lake is always free of charge, since the service only bills for the volume of events ingested into an event data store.
    6. F.It automatically forwards any event flagged as high risk to Amazon Macie for sensitive-data classification before the event becomes queryable in the console.
    Show answer & explanation

    Correct answers: A, B, C — It lets you run SQL queries across multiple event data stores spanning several accounts and Regions in the CloudTrail console, without configuring Athena tables.; It converts ingested events into columnar Apache ORC storage, so ad hoc SQL queries typically scan less data than an equivalent query over row-based JSON files.; It supports configurable retention of up to seven years on an event data store, meeting long compliance retention windows without a separate S3 lifecycle policy.

    • A. Event data stores can span multiple accounts and Regions, and the CloudTrail console itself provides the SQL query editor, so no separate Athena table or Glue Data Catalog setup is required.
    • B. CloudTrail Lake converts events to columnar ORC storage internally, which is why SQL queries against it are typically more efficient than scanning equivalent row-based JSON files directly.
    • C. Event data stores support pricing options with retention up to roughly seven years, letting compliance teams meet long retention requirements without managing a separate S3 lifecycle configuration.
    • D. CloudTrail Lake manages its own internal storage for event data stores rather than writing events into a customer-owned S3 bucket, so this description of the storage model is incorrect.
    • E. CloudTrail Lake bills separately for data ingested into an event data store and for the amount of data scanned when running queries, so query execution is not free.
    • F. CloudTrail Lake has no built-in integration that automatically routes flagged events to Amazon Macie; Macie is a separate service focused on discovering sensitive data in Amazon S3.

    Subdomain 4.5: Understand data privacy and governance

    34.An auditor asks a data engineer to determine exactly when the encryption setting on a specific Amazon S3 bucket was changed last quarter, what the setting's value was immediately before and after the change, and which other resources referenced that bucket at the time. Which AWS service is purpose-built to answer this question?

    1. A.AWS Config, using its configuration history and item timeline for the bucket to show setting values and related resources over time.
    2. B.AWS CloudTrail, using its event history to show which principal made API calls against the bucket during the quarter in question.
    3. C.Amazon CloudWatch Logs, using a log group that captures S3 server access logs to help reconstruct the bucket's past setting history.
    4. D.AWS Trusted Advisor, using its security checks to display the bucket's current encryption configuration and flag any risks.
    Show answer & explanation

    Correct answer: A — AWS Config, using its configuration history and item timeline for the bucket to show setting values and related resources over time.

    • A. AWS Config records configuration items and configuration history for a resource over time, letting the engineer view the exact value of a setting before and after a change and see the relationships to other resources at that point in time.
    • B. CloudTrail records who called which API and when, which is useful for identifying the actor, but it does not by itself present the before-and-after configuration state or the resource's relationship timeline the way Config's configuration history does.
    • C. S3 server access logs capture data-plane requests like object reads and writes; they do not track changes to bucket-level settings such as encryption configuration.
    • D. Trusted Advisor evaluates the current state of a resource against best-practice checks; it does not retain a historical timeline of past configuration values for a specific bucket.

    Subdomain 4.5: Understand data privacy and governance

    35.A data platform team uses Amazon SageMaker Unified Studio and wants a marketing analyst to be able to browse and query a governed customer dataset for a specific initiative, without giving that analyst standing access to every dataset registered in the organization's catalog. Which approach matches how SageMaker Catalog projects are designed to manage this?

    1. A.Add the analyst as a project member for the initiative, and have them request the specific catalog asset so its owner can approve access scoped to that project.
    2. B.Grant the analyst an IAM policy with sagemaker:* permissions on the account so they can reach every dataset registered anywhere in the SageMaker Catalog directly.
    3. C.Export the entire catalog's metadata to a shared spreadsheet and email it to the analyst so they can identify the tables relevant to their initiative themselves.
    4. D.Create a new Redshift datashare containing every table in the catalog and grant the analyst's IAM user direct SELECT access to the datashare.
    Show answer & explanation

    Correct answer: A — Add the analyst as a project member for the initiative, and have them request the specific catalog asset so its owner can approve access scoped to that project.

    • A. SageMaker Catalog projects scope membership and asset subscriptions to a specific initiative, so adding the analyst to the project and having them request the specific asset lets the dataset owner grant access limited to that project rather than the whole catalog.
    • B. A broad sagemaker:* IAM policy would expose every dataset in the catalog to the analyst, which is the opposite of the least-privilege, project-scoped access the team wants to provide.
    • C. Exporting metadata to a spreadsheet abandons the catalog's governed access and subscription model entirely and provides no enforcement of what the analyst is actually permitted to query.
    • D. Building a Redshift datashare over the whole catalog and granting direct SELECT access bypasses the project-based subscription and approval workflow and exposes far more data than the single initiative requires.

    Want the full experience?

    These are just samples. Practice the full AWS Certified Data Engineer - Associate (DEA-C01) question bank in quiz mode — free, no signup, with domain practice and exam simulation.