CertSafari

    Free AWS Certified Machine Learning Engineer - Associate (MLA-C01) Sample Questions

    35 free sample questions from our bank of 352+, covering every exam domain, with answers and detailed explanations. Updated September 2026.

    Domain 1: Data Preparation for Machine Learning (ML)

    Subdomain 1.1: Ingest and store data

    1.A team needs to merge nightly sales data extracted from an on-premises Oracle database over JDBC with clickstream data already stored as Parquet in S3, apply a multi-step Spark transformation to deduplicate and join the datasets, and load the merged result back into S3, all on a fully managed, serverless schedule without provisioning a persistent cluster. Which service fits this requirement?

    1. A.A persistent Amazon EMR cluster running Apache Spark, with both the Oracle JDBC driver and the S3 connector installed, kept running continuously to handle the nightly merge.
    2. B.A SageMaker Processing job running a single-instance pandas script that opens a JDBC connection to Oracle and reads the S3 Parquet files into memory for the join.
    3. C.An AWS Glue ETL job configured with both a JDBC connection to the Oracle database and an S3 data source, running the Spark-based merge logic on Glue's serverless infrastructure.
    4. D.An Amazon Redshift Spectrum query that joins an external table pointing at the S3 Parquet files with a federated query against the on-premises Oracle database.
    Show answer & explanation

    Correct answer: C — An AWS Glue ETL job configured with both a JDBC connection to the Oracle database and an S3 data source, running the Spark-based merge logic on Glue's serverless infrastructure.

    • A. A persistent EMR cluster can run the same Spark logic but requires the team to provision, patch, and pay for cluster infrastructure continuously, which directly conflicts with the fully managed, serverless requirement stated.
    • B. A single-instance pandas script would struggle to scale to nightly sales and clickstream data volumes and would require the team to manage retries, scheduling, and JDBC driver packaging themselves rather than using a managed ETL service.
    • C. AWS Glue provides serverless, managed Spark execution along with built-in JDBC connections for relational sources and native S3 support, so it can run the described multi-step merge on a schedule without the team provisioning or managing any cluster.
    • D. Redshift Spectrum federated queries can join across sources, but this pattern is built around SQL-based analytical queries rather than orchestrating a scheduled, multi-step Spark-based deduplication and merge ETL pipeline.

    Subdomain 1.1: Ingest and store data

    2.A team's Kinesis Data Streams producer application begins throwing `ProvisionedThroughputExceededException` errors during a traffic spike, even though the total data volume being written is well within the stream's aggregate provisioned throughput across all shards. Investigation shows most records share the same partition key. Which change would resolve this specific throttling pattern?

    1. A.Use a more varied or higher-cardinality partition key strategy so records spread evenly across all shards instead of concentrating writes on a single shard's throughput limit.
    2. B.Increase the stream's data retention period from the default twenty-four hours to a longer window so that throttled records can be retried over more time.
    3. C.Switch the producer to use the Kinesis Producer Library's synchronous PutRecord calls instead of batched PutRecords calls to reduce per-call throughput consumption.
    4. D.Increase the total number of shards on the stream without changing the partition key strategy, since more shards always increase the throughput available to every individual key.
    Show answer & explanation

    Correct answer: A — Use a more varied or higher-cardinality partition key strategy so records spread evenly across all shards instead of concentrating writes on a single shard's throughput limit.

    • A. A low-cardinality partition key routes nearly all records to one shard, and that single shard's fixed 1 MB/s and 1,000 records/s write limits get exceeded even while the stream's aggregate capacity is underused; distributing the key spreads writes across shards.
    • B. Retention period controls how long records remain available for consumers to read, not how much write throughput a single shard can accept, so extending it does not address a write-side throughput exception.
    • C. Switching from batched to individual synchronous PutRecord calls does not change which shard a given partition key is routed to, so the same single shard would still receive the concentrated write load and remain throttled.
    • D. Adding shards increases the stream's total aggregate throughput, but a single hot partition key still hashes to only one shard, so the per-shard limit that is actually being exceeded remains unaffected by adding unrelated shards.

    Subdomain 1.1: Ingest and store data

    3.A team is choosing where to persist a nightly-refreshed batch of engineered features that will be read primarily by SageMaker training jobs and occasional ad hoc analysis, with no requirement for millisecond-latency single-record lookups, but with a strong preference for the lowest possible per-gigabyte storage cost and easy integration with existing S3-based training pipelines. Which storage choice best matches these priorities?

    1. A.Store the engineered features on a provisioned io2 Block Express EBS volume attached to a dedicated always-on EC2 instance, since block storage delivers the highest IOPS for any downstream reader.
    2. B.Store the engineered features as partitioned Parquet files in S3 Standard, since it offers low per-gigabyte cost, direct integration with SageMaker training input channels, and easy ad hoc querying through Athena.
    3. C.Store the engineered features in Amazon DynamoDB, since its single-digit-millisecond key-value access pattern is well suited to nightly-batch feature data consumed by training jobs and analysts.
    4. D.Store the engineered features in the SageMaker Feature Store online store, since its low-millisecond read latency ensures the fastest possible access for any downstream training or analysis workload.
    Show answer & explanation

    Correct answer: B — Store the engineered features as partitioned Parquet files in S3 Standard, since it offers low per-gigabyte cost, direct integration with SageMaker training input channels, and easy ad hoc querying through Athena.

    • A. A dedicated always-on EC2 instance with a provisioned IOPS volume incurs continuous compute and storage cost for a workload accessed only nightly and occasionally, which works against the goal of the lowest possible per-gigabyte cost.
    • B. Partitioned Parquet in S3 Standard combines low per-gigabyte storage cost, native integration as a SageMaker training data source, and straightforward ad hoc querying through Athena, matching every stated priority without adding unnecessary latency-oriented infrastructure.
    • C. DynamoDB is priced and optimized around low-latency key-value access patterns, which is more expensive per gigabyte than S3 and unnecessary for a training and ad hoc analysis pattern that has no millisecond lookup requirement.
    • D. The online store is optimized for millisecond single-record lookups at a higher cost per stored feature value, which is unnecessary and more expensive infrastructure for a workload that explicitly has no low-latency single-record requirement.

    Subdomain 1.2: Transform data and perform feature engineering

    4.A fraud detection system ingests transaction events through Amazon Kinesis Data Streams and needs each event's fields normalized and enriched with a computed risk score within a few seconds of arrival, before the transformed event is written downstream for real-time scoring. Which lightweight, event-driven service is best suited to perform this per-record transformation?

    1. A.Attach an AWS Lambda function as a stream consumer to transform and enrich each Kinesis record within seconds of it arriving.
    2. B.Schedule a nightly AWS Glue batch ETL job to read the full day's accumulated Kinesis records and transform them all at once.
    3. C.Configure a SageMaker Ground Truth labeling job to manually review and annotate each incoming Kinesis transaction event in real time.
    4. D.Run a weekly Amazon EMR Spark batch job against an S3 export of the stream to normalize and enrich the accumulated transaction events.
    Show answer & explanation

    Correct answer: A — Attach an AWS Lambda function as a stream consumer to transform and enrich each Kinesis record within seconds of it arriving.

    • A. AWS Lambda can be attached directly as a Kinesis Data Streams consumer, invoking a function per batch of records within seconds, which fits the low-latency per-record enrichment this fraud scenario requires.
    • B. A nightly batch Glue job introduces up to a day of latency before records are transformed, which fails the requirement to enrich each transaction within a few seconds of arrival for real-time scoring.
    • C. Ground Truth is a human labeling service for producing training datasets; it cannot perform automated, low-latency enrichment of streaming transaction events in real time.
    • D. A weekly EMR Spark batch job operates on accumulated historical exports and introduces far too much latency to meet a near-real-time, few-seconds transformation requirement for fraud scoring.

    Subdomain 1.2: Transform data and perform feature engineering

    5.After building and validating a feature engineering flow in SageMaker Data Wrangler, an ML engineer wants the resulting features to be automatically written into a managed store that both a real-time inference endpoint and a batch training job can later read from, without hand-writing a separate export script. What should the engineer do from within the Data Wrangler flow?

    1. A.Use the Data Wrangler export option to send the transformed features directly to a SageMaker Feature Store feature group as the destination.
    2. B.Manually copy each transformed column's values into a spreadsheet and re-upload that spreadsheet separately to an Amazon RDS database instance.
    3. C.Configure an AWS Glue crawler to scan the Data Wrangler flow file itself so it infers a schema that a training job can somehow later read.
    4. D.Set up a SageMaker Ground Truth labeling job to re-review the transformed features so they become suitable for real-time inference and training.
    Show answer & explanation

    Correct answer: A — Use the Data Wrangler export option to send the transformed features directly to a SageMaker Feature Store feature group as the destination.

    • A. Data Wrangler has a built-in export destination for Feature Store, letting a flow's transformed output be ingested directly into a feature group without a hand-written export script.
    • B. Manually copying data into a spreadsheet and re-uploading it to RDS reintroduces the manual export process the engineer explicitly wants to avoid, and RDS is not a Feature Store online/offline destination.
    • C. A Glue crawler infers schema from data files; a Data Wrangler flow file describes transformation steps rather than tabular data, so crawling it would not produce a usable feature store or training dataset.
    • D. Ground Truth manages human labeling workflows for creating training labels; it has no role in exporting already-transformed numeric features into a store usable by inference and training systems.

    Subdomain 1.2: Transform data and perform feature engineering

    6.An ML engineer is writing a custom AWS Glue ETL job in a Glue Studio notebook to combine order records with customer records stored in two separate Glue Data Catalog tables, matching rows on a shared `customer_id` column, before the combined DynamicFrame is written out for feature engineering. Which Glue capability should the engineer use to perform this combination?

    1. A.Use the Join transform on the two DynamicFrames within the Glue ETL job, matching rows on the shared `customer_id` column.
    2. B.Use an AWS Glue crawler to scan both tables again, since re-running the crawler automatically merges two separate tables into one.
    3. C.Use SageMaker Ground Truth to have human workers manually copy matching customer rows from one table into the other by hand.
    4. D.Use Amazon Mechanical Turk workers to visually compare the order and customer tables and report which rows share a customer ID.
    Show answer & explanation

    Correct answer: A — Use the Join transform on the two DynamicFrames within the Glue ETL job, matching rows on the shared `customer_id` column.

    • A. Glue's Join transform combines two DynamicFrames on a specified key, such as `customer_id`, which is the standard, code-native way to merge order and customer tables inside a Glue ETL job.
    • B. Re-running a crawler updates catalog metadata like schema and partitions for existing tables; it does not merge the row-level contents of two separate tables into a combined dataset.
    • C. Manually copying rows between tables through a human labeling workflow does not scale to a production ETL job and is not what Ground Truth is designed to perform.
    • D. Using a crowdsourced workforce to visually compare table rows for matching IDs is impractical at scale and is not a capability Mechanical Turk or Glue ETL jobs provide for structured table joins.

    Subdomain 1.3: Ensure data integrity and prepare data for modeling

    7.A lending company is preparing a training dataset for a loan-approval model and wants to check whether applicants over age 60 receive positive loan outcomes at a different rate than younger applicants. Which SageMaker Clarify pre-training bias metric directly measures this difference in the proportion of positive labels between the two age facets?

    1. A.Difference in Proportions of Labels (DPL), which compares the rate of positive observed outcomes between the favored and disfavored age facets
    2. B.Class Imbalance (CI), which compares the number of training samples collected for each age facet regardless of their loan outcomes
    3. C.Kolmogorov-Smirnov (KS) metric, which measures the maximum divergence across every outcome category between the two facets
    4. D.Conditional Demographic Disparity (CDD), which measures rejection versus acceptance disparity separately within each subgroup
    Show answer & explanation

    Correct answer: A — Difference in Proportions of Labels (DPL), which compares the rate of positive observed outcomes between the favored and disfavored age facets

    • A. Difference in Proportions of Labels measures exactly this: the gap in the rate of positive observed outcomes between a favored and a disfavored facet, which fits an age-based approval-rate comparison.
    • B. Class Imbalance only compares how many training samples exist for each facet and ignores the outcome labels entirely, so it cannot reveal an approval-rate gap.
    • C. The Kolmogorov-Smirnov metric measures the maximum divergence across the full outcome distribution rather than a single proportion-of-positive-outcomes comparison between two facets.
    • D. Conditional Demographic Disparity examines rejection-versus-acceptance patterns within subgroups conditioned on additional attributes, which is a finer-grained analysis than a direct two-facet proportion comparison.

    Subdomain 1.3: Ensure data integrity and prepare data for modeling

    8.A financial services company runs a nightly job that uploads transaction datasets from an on-premises data center to an S3 bucket for model training. An internal audit finds that some uploads occur over unencrypted HTTP connections. Which bucket-level configuration should the team apply to ensure that data is always encrypted while in transit to S3?

    1. A.Attach a bucket policy that denies any S3 request where the `aws:SecureTransport` condition key evaluates to false, forcing all uploads to use HTTPS
    2. B.Enable default server-side encryption on the bucket so that every object is automatically encrypted with SSE-KMS once it arrives in S3
    3. C.Turn on S3 Transfer Acceleration for the bucket so uploads route through CloudFront edge locations closer to the on-premises data center
    4. D.Enable S3 Object Lock in compliance mode so uploaded transaction objects cannot be overwritten or deleted for a fixed retention period
    Show answer & explanation

    Correct answer: A — Attach a bucket policy that denies any S3 request where the `aws:SecureTransport` condition key evaluates to false, forcing all uploads to use HTTPS

    • A. A bucket policy that denies requests where `aws:SecureTransport` is false rejects any non-HTTPS request outright, which directly enforces encryption in transit for every upload.
    • B. Default server-side encryption protects data once it is stored in S3, but it does not control or enforce the protocol used to transmit the data to S3 in the first place.
    • C. Transfer Acceleration improves upload speed by routing traffic through edge locations but does not by itself enforce that connections use HTTPS instead of HTTP.
    • D. Object Lock prevents deletion or modification of stored objects for compliance retention purposes and has no effect on whether the upload connection itself is encrypted.

    Subdomain 1.3: Ensure data integrity and prepare data for modeling

    9.A team training a facial attribute classifier notices the model performs noticeably worse for one skin-tone group that is underrepresented in the training images. Beyond collecting more real photos, which data preparation technique can most directly reduce this prediction bias by expanding the effective diversity of examples for the underrepresented group?

    1. A.Apply targeted data augmentation, such as varying lighting, angle, and background, specifically to images from the underrepresented skin-tone group
    2. B.Increase the image resolution used for every photo in the dataset so that finer facial details are preserved during training
    3. C.Apply the same augmentation transformations uniformly and equally across every skin-tone group regardless of how many images each group currently has
    4. D.Reduce the number of training epochs so the model has less opportunity to overfit to the more heavily represented skin-tone group
    Show answer & explanation

    Correct answer: A — Apply targeted data augmentation, such as varying lighting, angle, and background, specifically to images from the underrepresented skin-tone group

    • A. Targeted augmentation applied specifically to the underrepresented group's existing images creates additional diverse, realistic training variations for that group, directly narrowing the representation gap driving the prediction bias.
    • B. Increasing resolution uniformly changes image detail for every group but does not change the relative number of training examples available for the underrepresented group.
    • C. Applying the same amount of augmentation to every group proportionally increases each group by the same multiple, so the underrepresented group's relative share of the dataset does not improve.
    • D. Reducing training epochs affects how much the model fits the data overall but does not change the underlying imbalance in how many examples exist for each skin-tone group.

    Domain 2: ML Model Development

    Subdomain 2.1: Choose a modeling approach

    10.A payments company wants to flag a transaction as likely fraud by comparing it to the outcomes of the transactions that most closely resemble it in amount, merchant category, and device fingerprint, without assuming any particular functional relationship between those features and fraud. Data scientists want a non-parametric method that can be updated by simply adding new labeled transactions to the reference set. Which SageMaker AI built-in algorithm fits this requirement?

    1. A.Use the Latent Dirichlet Allocation algorithm, since discovering latent topics across the transaction reference set will surface which topic corresponds to fraud.
    2. B.Use the IP Insights algorithm, since learning historical associations between IP addresses and user accounts will classify the specific transaction as fraudulent or legitimate.
    3. C.Use the k-Nearest Neighbors algorithm, since it classifies a new transaction from the labels of its closest neighbors in the reference set without assuming a functional form.
    4. D.Use the Factorization Machines algorithm, since modeling pairwise interactions between sparse categorical features will directly assign a fraud label to the transaction.
    Show answer & explanation

    Correct answer: C — Use the k-Nearest Neighbors algorithm, since it classifies a new transaction from the labels of its closest neighbors in the reference set without assuming a functional form.

    • A. LDA is an unsupervised topic-modeling algorithm for discovering themes in a collection of documents; it has no mechanism for producing a fraud classification for a single transaction based on nearest labeled neighbors.
    • B. IP Insights learns usage patterns between IP addresses and entities to flag unusual pairings, which is a narrower, address-focused anomaly signal rather than the general similarity-based classification across amount, merchant, and device features described here.
    • C. k-Nearest Neighbors is a non-parametric algorithm that predicts a label by looking at the k most similar labeled points in the reference set, which matches the requirement to compare a new transaction to similar past outcomes and update simply by adding new labeled data.
    • D. Factorization Machines are well suited to sparse, high-dimensional categorical interactions such as click-through prediction, but the scenario specifically calls for a similarity-based, non-parametric comparison to a reference set, which is not how factorization machines make predictions.

    Subdomain 2.1: Choose a modeling approach

    11.A helpdesk team wants to automatically find, for every new incoming ticket, the most similar previously resolved ticket so agents can reuse the earlier resolution. The approach must learn low-dimensional embeddings from labeled pairs of tickets that were previously identified as related, so those embeddings can then feed a downstream similarity search. Which SageMaker AI built-in algorithm is designed for this embedding task?

    1. A.Train the Image Classification algorithm, since categorizing a screenshot attached to each ticket will surface the most similar prior support resolution found.
    2. B.Train the Latent Dirichlet Allocation algorithm, since assigning each ticket a distribution over latent topics will directly identify the single most similar past resolved ticket.
    3. C.Train the Linear Learner algorithm in regression mode, since predicting a similarity score per pair of tickets from raw text features is the standard, well-known approach.
    4. D.Train the Object2Vec algorithm, since it learns low-dimensional embeddings of paired objects from the labeled ticket-relationship data, improving downstream similarity search.
    Show answer & explanation

    Correct answer: D — Train the Object2Vec algorithm, since it learns low-dimensional embeddings of paired objects from the labeled ticket-relationship data, improving downstream similarity search.

    • A. Image Classification assigns a category label to an image and is not applicable to comparing the textual similarity of two support tickets, and most tickets in this scenario are text, not images.
    • B. LDA discovers latent topics across a corpus of documents in an unsupervised way, but it does not use labeled pairwise relationships between tickets and does not directly output a similarity ranking between two specific tickets.
    • C. A generic Linear Learner regression on raw text features would require substantial manual feature engineering and is not designed to learn embeddings from labeled object pairs the way Object2Vec is purpose-built to do.
    • D. Object2Vec is a highly customizable algorithm for learning low-dimensional dense embeddings of paired objects from labeled relationship data, which is exactly the mechanism needed to embed tickets so a downstream nearest-neighbor search can find similar past resolutions.

    Subdomain 2.1: Choose a modeling approach

    12.A media company wants to launch a text classifier that tags incoming articles by topic within two weeks, has only a few thousand labeled examples of its own, and wants to start from a model that has already learned general language patterns rather than training a transformer from random initialization. Which SageMaker AI capability should the team use to meet this timeline?

    1. A.Train the SageMaker AI Semantic Segmentation built-in algorithm from randomly initialized weights on these labeled articles, since that is the fastest way to learn accurate topic tags.
    2. B.Build a custom transformer architecture from scratch and train it end to end on the few thousand labeled articles before the tight two-week deadline arrives.
    3. C.Train a Random Cut Forest model directly on the article text, since it will assign each article the correct topic tag after seeing enough labeled training examples.
    4. D.Use SageMaker JumpStart to fine-tune a pretrained foundation model on the company's own labeled articles, adapting existing language knowledge to the topic-tagging task quickly.
    Show answer & explanation

    Correct answer: D — Use SageMaker JumpStart to fine-tune a pretrained foundation model on the company's own labeled articles, adapting existing language knowledge to the topic-tagging task quickly.

    • A. Semantic Segmentation is an image algorithm for pixel-level labeling and has no applicability to text topic classification, and training any model from randomly initialized weights forgoes the benefit of transfer learning the team needs given limited data.
    • B. Building and training a custom transformer from scratch typically requires far more labeled data and compute time than a few thousand examples and two weeks allow, making it impractical compared to fine-tuning an existing pretrained model.
    • C. Random Cut Forest is an unsupervised anomaly-detection algorithm that outputs an anomaly score, not a topic classification, so it cannot be trained to assign topic tags to articles.
    • D. SageMaker JumpStart provides pretrained, open-source foundation models that can be fine-tuned on a smaller labeled dataset, letting the team adapt existing language understanding to topic tagging quickly instead of learning language from scratch, which fits the tight two-week timeline and limited labeled data.

    Subdomain 2.2: Train and refine models

    13.A team is training a large computer vision model on a dataset that fits in memory on a single GPU, but training on one instance takes far too long to iterate on experiments. Which distributed training approach is most appropriate for reducing wall-clock training time in this situation?

    1. A.Use model parallel training to split the layers of the model itself across multiple instances, since the whole model is too large to fit into the memory of any single GPU.
    2. B.Use SageMaker AI distributed data parallel training to split each batch across multiple instances, with each instance computing gradients on its data shard before the gradients are synchronized.
    3. C.Increase the batch size on a single instance until GPU memory is fully saturated, then rely on gradient accumulation to simulate the effect of training across multiple instances.
    4. D.Run several independent single-instance training jobs with different random seeds in parallel across the training fleet, then average the final model weights once every job finishes a full training run.
    Show answer & explanation

    Correct answer: B — Use SageMaker AI distributed data parallel training to split each batch across multiple instances, with each instance computing gradients on its data shard before the gradients are synchronized.

    • A. Model parallelism is meant for models too large to fit in a single GPU's memory, but the scenario states the model already fits on one GPU, so it does not address the actual bottleneck.
    • B. Data parallel training shards each batch across multiple instances so gradients are computed and synchronized in parallel, which directly reduces wall-clock time when a single instance is the bottleneck.
    • C. Gradient accumulation on one instance still performs the same total computation on the same hardware, so it does not reduce wall-clock time the way adding more instances does.
    • D. Naively averaging weights from independently trained models does not reproduce synchronized data-parallel training and typically produces a worse model than coordinated gradient updates.

    Subdomain 2.2: Train and refine models

    14.A linear model is trained on a dataset with thousands of candidate features, many of which are believed to be irrelevant to the target. The team wants a regularization technique that will drive the coefficients of irrelevant features to exactly zero, effectively performing feature selection during training. Which regularization technique should they apply?

    1. A.L2 regularization, which adds a penalty proportional to the squared magnitude of the coefficients and shrinks all coefficients smoothly toward zero without eliminating any of them.
    2. B.L1 regularization, which adds a penalty proportional to the absolute value of the coefficients and tends to push many coefficients to exactly zero, removing irrelevant features.
    3. C.Dropout, which randomly deactivates a fraction of the input features during each training pass so the model cannot rely on any single feature too heavily.
    4. D.Early stopping, which halts training once the validation metric stops improving, which indirectly limits how large any individual coefficient can grow during training.
    Show answer & explanation

    Correct answer: B — L1 regularization, which adds a penalty proportional to the absolute value of the coefficients and tends to push many coefficients to exactly zero, removing irrelevant features.

    • A. L2 regularization shrinks coefficients smoothly toward zero but rarely drives them to exactly zero, so it does not perform the same sparse feature selection as L1.
    • B. L1 regularization's absolute-value penalty produces sparse solutions where many coefficients are driven to exactly zero, which is the feature-selection effect described in the scenario.
    • C. Dropout is a neural network technique for deactivating neurons during training and does not apply to shrinking or zeroing out linear model coefficients.
    • D. Early stopping controls training duration based on validation performance and does not directly zero out any model coefficients.

    Subdomain 2.2: Train and refine models

    15.A team wants to combine multiple weak, shallow decision trees into a strong learner by training each new tree specifically to correct the residual errors made by the ensemble of previously trained trees, rather than training all trees independently in parallel. Which ensembling technique are they describing?

    1. A.Bagging, which trains multiple learners independently and in parallel on different bootstrapped samples of the training data, then averages or votes on all of their individual predictions.
    2. B.Stacking, which trains several different types of base models independently and then trains a separate meta-model to combine their individual predictions.
    3. C.Boosting, which trains an ensemble of weak learners sequentially, where each new learner focuses specifically on the errors made by the combined ensemble of previous learners so far.
    4. D.Grid search, which exhaustively evaluates every combination of a predefined set of discrete hyperparameter values to find the best performing configuration.
    Show answer & explanation

    Correct answer: C — Boosting, which trains an ensemble of weak learners sequentially, where each new learner focuses specifically on the errors made by the combined ensemble of previous learners so far.

    • A. Bagging trains learners independently and in parallel on bootstrapped samples rather than having each new learner correct the errors of previous ones.
    • B. Stacking combines different model types through a separately trained meta-model rather than having each new learner sequentially correct the ensemble's residual errors.
    • C. Boosting builds an ensemble sequentially, with each new weak learner specifically targeting the residual errors of the combined previous learners, matching the described training process exactly.
    • D. Grid search is a hyperparameter search strategy that evaluates discrete combinations and has no role in how an ensemble of learners is trained together.

    Subdomain 2.3: Analyze model performance

    16.A team trained a multiclass classifier that routes support tickets into five categories and visualizes results as a confusion matrix heat map. The heat map shows a dense off-diagonal cell where actual `billing` tickets are frequently predicted as `account_access`. What should the team conclude from this pattern?

    1. A.The model separates categories well overall, and a single dense off-diagonal cell is a normal rendering artifact that appears in every heat map regardless of the model's actual accuracy.
    2. B.The model confuses `billing` and `account_access` specifically, which points to overlapping feature signals between those two categories that the current model cannot reliably distinguish.
    3. C.The `billing` category must contain far too few labeled examples anywhere in the dataset, since only class imbalance can ever produce an off-diagonal cluster in a confusion matrix.
    4. D.The heat map proves the `account_access` labels in the training data were assigned incorrectly, because a well performing classifier never misclassifies tickets into a neighboring category.
    Show answer & explanation

    Correct answer: B — The model confuses `billing` and `account_access` specifically, which points to overlapping feature signals between those two categories that the current model cannot reliably distinguish.

    • A. A dense off-diagonal cell is a genuine signal of confusion between two specific classes, not a generic rendering artifact that would appear regardless of model quality.
    • B. A concentrated off-diagonal cell between two specific classes shows the model struggles to separate those categories, which typically reflects overlapping or ambiguous features between them rather than a random error pattern.
    • C. Class imbalance can contribute to poor performance on a category, but it is not the only cause of an off-diagonal cluster, and the pattern described points to feature overlap between two named classes instead.
    • D. A confusion matrix reflects model prediction errors, not necessarily a labeling defect in the underlying dataset, so jumping to a labeling error conclusion is not supported by this pattern alone.

    Subdomain 2.3: Analyze model performance

    17.A trained classifier for a dataset where 97% of examples belong to the negative class reports 97% accuracy. Before celebrating, the team compares this result against a majority-class baseline that always predicts the negative class. What does this comparison reveal?

    1. A.The majority-class baseline also achieves roughly 97% accuracy, so the trained model has not demonstrated any meaningful improvement over always predicting the dominant class.
    2. B.The majority-class baseline would score close to 0% accuracy on this dataset, so the trained model's 97% result clearly demonstrates strong, meaningful predictive skill.
    3. C.Comparing against a majority-class baseline is not a valid technique on imbalanced data, since baselines are only meaningful when the two classes are roughly equal in size.
    4. D.The 97% accuracy figure is invalid on its own regardless of any baseline, because accuracy can never be computed correctly on a dataset with any class imbalance present at all.
    Show answer & explanation

    Correct answer: A — The majority-class baseline also achieves roughly 97% accuracy, so the trained model has not demonstrated any meaningful improvement over always predicting the dominant class.

    • A. A baseline that always predicts the dominant class matches the class proportion, so it would also score near 97% accuracy here, showing the trained model has not clearly outperformed a trivial strategy on this metric alone.
    • B. A majority-class baseline on a 97% dominant negative class would itself score close to 97%, not near zero, so this option misstates how the baseline would actually perform.
    • C. Majority-class baselines are specifically useful for exposing when accuracy is misleading on imbalanced data, so dismissing the comparison as invalid ignores exactly the situation it is meant to catch.
    • D. Accuracy can still be computed on imbalanced data; the issue is that the number becomes less informative on its own, not that the calculation itself becomes invalid.

    Subdomain 2.3: Analyze model performance

    18.A team has a model that scores slightly higher on validation accuracy than a simpler alternative, but it takes four times longer to train and requires a more expensive GPU instance type for both training and inference. The business use case does not have a strict accuracy requirement beyond the simpler model's current score. How should the team weigh this tradeoff?

    1. A.Always choose the model with the highest validation accuracy regardless of training time or inference cost, since accuracy is the only factor that should ever influence a model selection decision.
    2. B.Favor the simpler, cheaper model when its accuracy already meets the business requirement, since the slight accuracy gain does not justify the added training time and inference cost here.
    3. C.Choose the more expensive model because a larger inference instance type always produces lower total cost at any request volume once the system reaches production scale.
    4. D.Ignore both training time and cost entirely, and select whichever model was easier to implement in code regardless of what either candidate model actually achieves on accuracy.
    Show answer & explanation

    Correct answer: B — Favor the simpler, cheaper model when its accuracy already meets the business requirement, since the slight accuracy gain does not justify the added training time and inference cost here.

    • A. Optimizing for accuracy alone ignores the training time and inference cost differences described, which the exam expects a candidate to weigh against the marginal accuracy benefit.
    • B. When the simpler model already meets the business accuracy requirement, the marginal accuracy gain from the costlier model does not offset its added training time and higher inference cost, favoring the cheaper option.
    • C. A larger, more expensive instance type does not automatically produce lower total cost at scale; cost depends on the actual tradeoff between instance price and throughput, which is not established here.
    • D. Implementation ease is a separate concern from the performance, time, and cost tradeoff the scenario asks about, and ignoring accuracy and cost entirely skips the actual decision being evaluated.

    Domain 3: Deployment and Orchestration of ML Workflows

    Subdomain 3.3: Use automated orchestration tools to set up continuous integration and continuous delivery (CI/CD) pipelines

    19.A SageMaker AI pipeline needs to pause while a separate, non-AWS labeling vendor manually reviews a sample of predictions and sends back a quality signal before the pipeline continues to the next training step. The vendor's system communicates asynchronously and cannot be invoked as a direct API call from within the pipeline. Which step type is designed for this kind of external, asynchronous wait?

    1. A.A ProcessingStep, which runs a managed processing job inside the pipeline's own AWS account and has no mechanism to wait on an external vendor system
    2. B.A CallbackStep, which sends a token to an SQS queue and pauses the pipeline until an external process signals completion
    3. C.A ConditionStep, which only evaluates existing pipeline properties or parameters and cannot pause execution to wait for input from an outside party
    4. D.A LambdaStep, which synchronously invokes an AWS Lambda function and must receive a response within the function's own execution timeout to proceed
    Show answer & explanation

    Correct answer: B — A CallbackStep, which sends a token to an SQS queue and pauses the pipeline until an external process signals completion

    • A. A ProcessingStep only runs a job within the pipeline's own AWS environment and has no built-in way to pause for a signal coming from a system outside AWS.
    • B. A CallbackStep publishes a token to an Amazon SQS queue and halts the pipeline until an external system, such as a third-party vendor's process, sends that token back with a completion signal, which fits an asynchronous external wait.
    • C. A ConditionStep only branches based on values already available to the pipeline and cannot introduce a wait for an external party to supply new input.
    • D. A LambdaStep is a synchronous invocation tied to the function's own timeout, so it is not suited to waiting on an external vendor process that may take an unpredictable amount of time to respond.

    Subdomain 3.3: Use automated orchestration tools to set up continuous integration and continuous delivery (CI/CD) pipelines

    20.New labeled training data lands in an Amazon S3 bucket at unpredictable times throughout the day, and the ML team wants a retraining pipeline to start automatically as soon as a new object is uploaded, rather than on a fixed schedule. Which approach lets an object upload event directly trigger the retraining pipeline?

    1. A.An Amazon EventBridge rule that matches S3 object-created events for the bucket and targets the SageMaker AI pipeline to start a new execution
    2. B.A CodeDeploy deployment group with an in-place deployment configuration that redeploys the pipeline whenever an S3 lifecycle rule transitions an object
    3. C.A CodePipeline source stage configured to poll the S3 bucket once per day and start the pipeline only if new objects were found during that check
    4. D.A CodeBuild project with a scheduled cron trigger set to run every five minutes, checking the S3 bucket contents for any newly uploaded objects
    Show answer & explanation

    Correct answer: A — An Amazon EventBridge rule that matches S3 object-created events for the bucket and targets the SageMaker AI pipeline to start a new execution

    • A. Amazon S3 emits object-created events that Amazon EventBridge can match with a rule, allowing the rule to target and start a SageMaker AI pipeline execution immediately whenever a new labeled data object is uploaded.
    • B. CodeDeploy deployment groups manage application traffic shifting and have no mechanism tied to S3 lifecycle transitions or object-created events for starting an ML pipeline.
    • C. A daily polling source stage would delay the retraining start by up to a full day after the new data arrives, which does not meet the requirement to react as soon as the object is uploaded.
    • D. A cron-based CodeBuild trigger checks on a fixed interval rather than reacting immediately to the upload event itself, introducing unnecessary polling delay instead of event-driven triggering.

    Subdomain 3.3: Use automated orchestration tools to set up continuous integration and continuous delivery (CI/CD) pipelines

    21.A CodeBuild project must run integration tests that connect to a private Amazon RDS instance and an internal feature store endpoint, neither of which is reachable from the public internet. What should the team configure on the CodeBuild project so its build containers can reach these private resources during the test phase?

    1. A.A public build environment compute type upgrade, since increasing the vCPU and memory allocation is what grants a build container access to private subnets
    2. B.VPC configuration on the CodeBuild project, specifying the private subnets and security groups so build containers can reach the internal resources
    3. C.An IAM policy attached to the CodeBuild service role that grants `rds:Connect` and `es:ESHttpGet` actions, which alone establishes network connectivity to private endpoints
    4. D.A CodePipeline artifact store change to a customer-managed KMS key, since encrypting artifacts is what allows the build container to route into a private VPC
    Show answer & explanation

    Correct answer: B — VPC configuration on the CodeBuild project, specifying the private subnets and security groups so build containers can reach the internal resources

    • A. Changing the compute type only affects the build container's CPU and memory allocation; it has no effect on which network the container is launched into and does not grant access to private subnets.
    • B. Configuring a CodeBuild project with VPC settings, including the private subnets and security groups, launches the build container inside that VPC with network-level access to resources like a private RDS instance or an internal feature store endpoint.
    • C. IAM permissions control which AWS API calls the build role is authorized to make, but they do not establish network-layer connectivity into a VPC; without VPC configuration, the container still cannot route to private subnets.
    • D. Changing the KMS key used to encrypt pipeline artifacts affects encryption of stored artifacts and has no bearing on whether the build container can reach resources inside a private VPC.

    Subdomain 3.1: Select deployment infrastructure based on existing architecture and requirements

    22.A SaaS company wants to host hundreds of small XGBoost models, one per customer, behind a single endpoint to minimize hosting cost. Traffic to any individual customer's model is infrequent, but the company cannot predict which models will be called next. Which SageMaker AI deployment pattern best fits this scenario?

    1. A.A multi-model endpoint that dynamically loads each customer's model from Amazon S3 into a shared serving container on demand
    2. B.One dedicated real-time endpoint per customer, each running on its own small instance sized for that customer's traffic
    3. C.A multi-container endpoint where each of the hundreds of models runs in its own container within one endpoint configuration
    4. D.A single-model real-time endpoint that reloads a different customer's model weights on every incoming invocation request
    Show answer & explanation

    Correct answer: A — A multi-model endpoint that dynamically loads each customer's model from Amazon S3 into a shared serving container on demand

    • A. Multi-model endpoints host many similarly sized models on a shared fleet, dynamically downloading and caching each model from S3 on first use, which is exactly the cost-efficient pattern for many infrequently used, similar models sharing one endpoint.
    • B. Provisioning a dedicated endpoint per customer multiplies fixed hosting costs across hundreds of low-traffic models, which directly contradicts the goal of minimizing cost for infrequent per-customer usage.
    • C. Multi-container endpoints are intended for a small number of distinct containers serving different frameworks or steps, not for scaling to hundreds of per-customer models, which would require an unmanageable number of container definitions.
    • D. Manually reloading model weights on every request is not a supported endpoint pattern and would add prohibitive per-invocation latency compared to the caching behavior multi-model endpoints provide natively.

    Subdomain 3.1: Select deployment infrastructure based on existing architecture and requirements

    23.A data scientist trained a PyTorch model and needs a custom Python package with a compiled C extension installed at inference time that is not available in the standard pre-built PyTorch container and cannot be installed through a `requirements.txt` file at runtime. Which approach is appropriate?

    1. A.Extend the pre-built SageMaker AI PyTorch container with a custom Dockerfile that installs the compiled dependency during the image build
    2. B.Add the package name to a `requirements.txt` file so SageMaker AI installs the compiled extension automatically at container startup
    3. C.Deploy using the base pre-built PyTorch container unmodified, since all system-level dependencies are already included by default
    4. D.Switch to a serverless endpoint, since serverless containers automatically resolve any missing compiled dependency at invocation time
    Show answer & explanation

    Correct answer: A — Extend the pre-built SageMaker AI PyTorch container with a custom Dockerfile that installs the compiled dependency during the image build

    • A. When a dependency needs a build step, such as compiling a C extension, that a runtime `requirements.txt` install cannot perform, extending the pre-built container with a Dockerfile that compiles and installs it at image-build time is the documented approach.
    • B. A `requirements.txt` file only triggers a pip install at runtime and cannot perform the compilation step a C extension needs, so it does not work for dependencies the scenario says cannot be installed this way.
    • C. The unmodified pre-built container only ships with the packages SageMaker AI includes by default, so a custom compiled dependency that is not part of that image will still be missing at inference time.
    • D. Serverless endpoints run the same container image logic as real-time endpoints and do not automatically resolve missing compiled dependencies; the image itself still needs to contain the package.

    Subdomain 3.1: Select deployment infrastructure based on existing architecture and requirements

    24.Which statement about SageMaker Neo model optimization is accurate?

    1. A.Neo converts a trained model into a framework-agnostic intermediate representation, then compiles optimized binary code for the chosen target platform
    2. B.Neo retrains the model from raw training data using a smaller architecture to reduce the model's memory footprint before deployment
    3. C.Neo only supports compiling models for AWS Inferentia chips and cannot target any general-purpose edge device processors
    4. D.Neo requires the original training container to remain running alongside the compiled model at inference time for compatibility
    Show answer & explanation

    Correct answer: A — Neo converts a trained model into a framework-agnostic intermediate representation, then compiles optimized binary code for the chosen target platform

    • A. SageMaker Neo's compiler reads models from supported frameworks, converts them into a framework-agnostic intermediate representation, applies optimizations, and generates target-specific binary code plus a matching runtime for that platform.
    • B. Neo compiles and optimizes an already-trained model's computation graph; it does not retrain the model from raw data or redesign its architecture to shrink it.
    • C. Neo supports compilation for a broad set of targets, including cloud instances such as Inferentia as well as edge device processors from vendors like ARM, Intel, and NXP, not Inferentia exclusively.
    • D. A Neo-compiled model runs using Neo's own lightweight runtime on the target platform and does not depend on the original training container remaining active during inference.

    Subdomain 3.2: Create and script infrastructure based on existing architecture and requirements

    25.A team is packaging a scikit-learn model with a custom preprocessing step that has no equivalent in any SageMaker AI pre-built framework container. They must build a bring-your-own-container (BYOC) image so SageMaker AI can serve real-time inference from it. Which requirement must the custom image satisfy for SageMaker AI to invoke it correctly at inference time?

    1. A.The image must implement an HTTP server that responds to `/invocations` for inference requests and `/ping` for health checks, and it must be pushed to Amazon ECR so SageMaker AI can pull it when creating the endpoint
    2. B.The image only needs to include the trained model file and the scikit-learn library; SageMaker AI automatically injects a matching HTTP server into any container at endpoint creation time
    3. C.The image must be uploaded directly to the SageMaker AI console as a ZIP archive, since Amazon ECR only accepts images built from AWS-managed base layers
    4. D.The image must define a `lambda_handler` function as its entry point, since SageMaker AI real-time endpoints invoke containers using the same interface as AWS Lambda
    Show answer & explanation

    Correct answer: A — The image must implement an HTTP server that responds to `/invocations` for inference requests and `/ping` for health checks, and it must be pushed to Amazon ECR so SageMaker AI can pull it when creating the endpoint

    • A. A BYOC image for real-time inference must run a web server exposing the /invocations endpoint for scoring requests and /ping for health checks, and the built image needs to be stored in Amazon ECR because that is where SageMaker AI pulls container images from when it launches endpoint instances.
    • B. SageMaker AI does not inject any serving code into a custom container; a bring-your-own image is entirely responsible for implementing its own inference server, so omitting one leaves the container unable to respond to invocation requests.
    • C. Container images for SageMaker AI are stored and pulled from Amazon ECR as standard Docker images, not uploaded as ZIP archives through the console, and Amazon ECR accepts any properly built Docker image regardless of base layer source.
    • D. SageMaker AI real-time endpoints use an HTTP-based container contract, not the Lambda runtime API's handler-function interface, so a lambda_handler entry point has no effect on how a SageMaker AI container serves requests.

    Subdomain 3.2: Create and script infrastructure based on existing architecture and requirements

    26.A team hosts a latency-critical fraud-scoring endpoint where user-facing requests must stay under 100 ms even as traffic grows. They currently scale on `CPUUtilization`, but during load tests, latency degrades well before CPU utilization reaches the configured target, since large batch payloads saturate the model's inference compute before the instance's overall CPU appears busy. Which change to the scaling policy would better protect the latency requirement?

    1. A.Switch the target tracking policy to track `ModelLatency`, so new instances are added as soon as per-request inference time rises, directly reflecting the metric the team cares about
    2. B.Lower the `CPUUtilization` target value from 70 percent to 60 percent, which reduces the average load per instance without changing which metric drives scaling decisions
    3. C.Switch to a scheduled scaling policy that raises minimum capacity at fixed times of day based on typical traffic history for this endpoint
    4. D.Disable auto scaling entirely and permanently run the maximum instance count observed during the worst historical load test
    Show answer & explanation

    Correct answer: A — Switch the target tracking policy to track `ModelLatency`, so new instances are added as soon as per-request inference time rises, directly reflecting the metric the team cares about

    • A. Scaling directly on ModelLatency ties capacity decisions to the actual per-request inference time that is degrading, so instances are added in response to the latency symptom itself rather than waiting for an indirect proxy metric like overall CPU to catch up.
    • B. Lowering the CPU target still scales on the same metric that lags behind the real bottleneck in this scenario, so it may add capacity somewhat earlier but does not fix the underlying mismatch between CPU utilization and actual inference latency.
    • C. Scheduled scaling only changes capacity at fixed times based on historical patterns and does not respond to real-time latency degradation caused by payload characteristics that vary independently of time of day.
    • D. Running permanently at worst-case capacity avoids the metric mismatch but pays for peak-load capacity at all times, which is a costly workaround rather than a fix to the scaling policy's choice of metric.

    Subdomain 3.2: Create and script infrastructure based on existing architecture and requirements

    27.An ML engineer is provisioning the IAM execution role that a SageMaker AI endpoint will assume to load model artifacts and publish CloudWatch metrics. Following least-privilege practice while scripting this role in CloudFormation, which approach is most appropriate?

    1. A.Grant `s3:GetObject` scoped to the specific model artifact bucket and prefix, plus the managed CloudWatch metrics permissions the endpoint needs, rather than attaching a broad administrator policy
    2. B.Attach the `AdministratorAccess` managed policy to the role so the endpoint is guaranteed to have every permission it could ever need without further changes
    3. C.Skip attaching any IAM policy to the role, since SageMaker AI endpoints inherit the permissions of the AWS account root user automatically at runtime
    4. D.Grant `s3:*` on all buckets in the account so that any future model artifact location the team chooses will already be accessible without updating the role
    Show answer & explanation

    Correct answer: A — Grant `s3:GetObject` scoped to the specific model artifact bucket and prefix, plus the managed CloudWatch metrics permissions the endpoint needs, rather than attaching a broad administrator policy

    • A. Scoping s3:GetObject to only the specific bucket and prefix holding the model artifacts, alongside the narrow CloudWatch metrics permissions actually required, follows least privilege by granting exactly the access the endpoint needs and nothing more.
    • B. Attaching AdministratorAccess grants the endpoint's role far more permission than reading model artifacts and publishing metrics requires, which violates least privilege and expands the damage possible if the role's credentials are ever misused.
    • C. SageMaker AI endpoints have no implicit inheritance of root account permissions; without an attached policy granting explicit access, the endpoint's role would be unable to read model artifacts or publish metrics at all.
    • D. Granting s3:* across every bucket in the account is far broader than the endpoint's actual need to read from one artifact location, and it violates least privilege by exposing unrelated data the endpoint has no reason to touch.

    Domain 4: ML Solution Monitoring, Maintenance, and Security

    Subdomain 4.1: Monitor model inference

    28.A SageMaker Model Monitor data quality job publishes custom CloudWatch metrics after each scheduled execution. The ML engineer wants the retraining pipeline to start automatically whenever the violation count for a specific feature exceeds a threshold across two consecutive executions, without a human checking the console. Which combination accomplishes this?

    1. A.Create a CloudWatch alarm on the monitoring job's custom metric with the required evaluation periods, and have the alarm invoke an EventBridge rule that starts retraining
    2. B.Enable AWS CloudTrail data events for the monitoring schedule and configure a CloudTrail Lake query that runs daily to check whether violations occurred in the prior executions
    3. C.Set the monitoring schedule's `MonitoringExecutionSummary` status to `Failed` whenever violations occur, and manually restart the pipeline the next time someone reviews the schedule history
    4. D.Increase the monitoring job's instance type so that violation detection completes faster, which shortens the delay before anyone notices the issue by checking the console manually
    Show answer & explanation

    Correct answer: A — Create a CloudWatch alarm on the monitoring job's custom metric with the required evaluation periods, and have the alarm invoke an EventBridge rule that starts retraining

    • A. A CloudWatch alarm evaluated over the required number of periods against the monitoring job's published metric can transition to an alarm state that triggers an EventBridge rule, giving a fully automated path from detected violations to a retraining pipeline execution.
    • B. CloudTrail records API activity for auditing purposes, not the statistical violation counts produced by a Model Monitor execution, so it cannot be used to evaluate whether a metric threshold was crossed.
    • C. Relying on someone to notice a failed status and manually restart the pipeline defeats the requirement for an automatic response and introduces a delay tied to human review rather than immediate action.
    • D. A larger instance type can shorten how long the monitoring job takes to run, but it does nothing to automatically trigger a downstream pipeline once violations are detected.

    Subdomain 4.1: Monitor model inference

    29.A logistics company scores shipment-delay risk using a SageMaker AI batch transform job that runs every night against the day's shipment records, rather than a real-time endpoint. They want to monitor these batch predictions for data quality drift the same way they would for a real-time endpoint. What is required to set this up?

    1. A.Enable data capture for the batch transform job's inputs and outputs and create an on-schedule Model Monitor job analyzing the captured data against the baseline
    2. B.Convert the batch transform job into a real-time endpoint, since Model Monitor only supports analyzing traffic sent to a persistent, continuously running real-time endpoint
    3. C.Enable SageMaker Clarify bias drift monitoring on the batch transform job, since bias drift monitoring is the only Model Monitor type that supports asynchronous batch workloads
    4. D.Write the batch transform output to Amazon S3 and manually run a Jupyter notebook comparing summary statistics each morning, since Model Monitor has no batch transform support
    Show answer & explanation

    Correct answer: A — Enable data capture for the batch transform job's inputs and outputs and create an on-schedule Model Monitor job analyzing the captured data against the baseline

    • A. Model Monitor supports on-schedule monitoring for batch transform jobs by capturing the job's inputs and outputs and running an analysis against the baseline, mirroring the workflow used for real-time endpoints without requiring an endpoint.
    • B. Converting the workload to a real-time endpoint is unnecessary and changes the deployment architecture; Model Monitor explicitly supports on-schedule monitoring of batch transform jobs as an alternative to real-time endpoint monitoring.
    • C. Bias drift monitoring is one specific monitoring type focused on fairness metrics, not the only type that works with batch transform; data quality and model quality monitoring also support batch transform jobs.
    • D. A manual notebook-based comparison ignores the built-in on-schedule batch transform monitoring capability that Model Monitor already provides, requiring ongoing manual effort instead of an automated schedule.

    Subdomain 4.1: Monitor model inference

    30.A financial services company's compliance team requires a record of every API call made to invoke a SageMaker endpoint hosting a model used in credit decisions, including who or what made the call, the source IP, and the timestamp, for a regulatory audit. Which AWS service should be enabled to satisfy this requirement?

    1. A.AWS CloudTrail, configured to log data events for the SageMaker endpoint so each invocation is recorded with the caller identity, source IP address, and timestamp
    2. B.SageMaker Model Monitor, configured with a data quality baseline, since its violation reports include the identity of every caller that sent a request to the endpoint
    3. C.AWS X-Ray, configured with active tracing enabled on the endpoint, since trace segments include the IAM identity that initiated each traced inference request
    4. D.Amazon CloudWatch Logs Insights, configured to run a saved query hourly against the endpoint's application logs to reconstruct the identity of each caller
    Show answer & explanation

    Correct answer: A — AWS CloudTrail, configured to log data events for the SageMaker endpoint so each invocation is recorded with the caller identity, source IP address, and timestamp

    • A. CloudTrail is the AWS service purpose-built for auditing API activity, and enabling data events for a SageMaker endpoint captures each invocation with the caller's identity, source IP address, and timestamp, which is exactly what a compliance audit requires.
    • B. Model Monitor's violation reports describe statistical deviations in feature values or predictions; they do not record caller identity or source IP information needed for an audit trail of who invoked the endpoint.
    • C. X-Ray trace segments capture timing and service call information for performance troubleshooting, not the caller identity and source IP metadata that a compliance audit trail specifically requires.
    • D. Application logs typically do not include structured caller identity and source IP fields for every invocation unless custom logging is added, making this an unreliable and incomplete substitute for CloudTrail's built-in audit records.

    Subdomain 4.3: Secure AWS resources

    31.A shared AWS account hosts ML resources for two internal product teams, Fraud and Search. The platform team wants engineers on each team to only be able to access SageMaker AI resources tagged for their own project, without maintaining a separate IAM policy per resource as new resources are created. Which access control pattern fits this requirement?

    1. A.Attribute-based access control using an aws:ResourceTag condition that only allows an action when the resource's Project tag matches the caller's own tag.
    2. B.Role-based access control using a single shared IAM role attached to every engineer on both teams, since a shared role removes the need to track tags on resources.
    3. C.Network-based access control using a security group rule that only allows traffic between resources that were both created in the same calendar month and Region.
    4. D.Discretionary access control using S3 bucket ACLs on every object individually, updated by hand each time a new SageMaker AI resource is created for either team.
    Show answer & explanation

    Correct answer: A — Attribute-based access control using an aws:ResourceTag condition that only allows an action when the resource's Project tag matches the caller's own tag.

    • A. An aws:ResourceTag condition compares a resource's Project tag against the caller's own tag at request time, which scales to new resources automatically without hand-writing a new policy statement for each one.
    • B. A single shared role attached to both teams grants the same access to everyone regardless of which project they belong to, which is the opposite of separating Fraud and Search access.
    • C. Security groups control network traffic between resources based on IP and port rules; they have no concept of a resource's creation date or of separating access by project ownership.
    • D. Per-object ACLs updated by hand for every new resource is exactly the manual, non-scaling maintenance burden the team is trying to avoid by adopting tag-based access control.

    Subdomain 4.3: Secure AWS resources

    32.A SageMaker AI Studio domain is shared by several user profiles belonging to different data scientists. The platform team wants each user's notebook kernel to be able to write only to that specific user's own output prefix in S3, so one data scientist cannot read or overwrite another's files, even though everyone shares the same Studio domain. What should the team configure?

    1. A.Assign each user profile its own dedicated execution role scoped to that user's S3 output prefix, rather than having every profile share one domain-wide execution role.
    2. B.Create a single execution role for the entire Studio domain with s3:* permissions on the whole bucket, and instruct each data scientist to only use their own prefix voluntarily.
    3. C.Enable network isolation on the Studio domain so that data scientists cannot make any outbound network call to Amazon S3 or any other AWS service.
    4. D.Give each data scientist an individual S3 bucket ACL entry directly on the shared execution role's IAM user, since ACLs apply to IAM users rather than to S3 objects.
    Show answer & explanation

    Correct answer: A — Assign each user profile its own dedicated execution role scoped to that user's S3 output prefix, rather than having every profile share one domain-wide execution role.

    • A. Per-user-profile execution roles scoped to each user's own prefix mean the permissions themselves stop cross-user access, rather than relying on data scientists to self-restrict which prefix they touch.
    • B. A single shared role with s3:* on the whole bucket technically allows every user to read and overwrite every other user's files, so voluntary compliance is the only thing preventing exactly the access the team wants to block.
    • C. Network isolation would prevent Studio notebooks from reaching S3 at all, including the user's own output prefix, so it blocks legitimate access rather than separating access between users.
    • D. S3 ACLs are attached to buckets and objects, not to IAM users, so there is no such thing as an ACL entry placed directly on an IAM user or execution role.

    Subdomain 4.2: Monitor and optimize infrastructure and costs

    33.An ML engineer manages an Amazon SageMaker AI real-time endpoint serving a fraud-detection model. Users report intermittent slow responses during peak hours, and the team wants an automatic notification within minutes whenever per-request processing time exceeds an acceptable threshold, without manually checking logs. Which approach meets this requirement?

    1. A.Enable CloudWatch Lambda Insights on the endpoint container and review the collected cold-start diagnostics only after an incident is reported.
    2. B.Configure CloudTrail to record `InvokeEndpoint` calls and run a scheduled Athena query that emails a summary report of call volume once per day.
    3. C.Set up an AWS Trusted Advisor check that scans EC2 instance utilization weekly and flags endpoints running below 10 percent CPU usage.
    4. D.Create a CloudWatch alarm on the endpoint's `ModelLatency` metric that notifies an SNS topic once the average exceeds a defined threshold.
    Show answer & explanation

    Correct answer: D — Create a CloudWatch alarm on the endpoint's `ModelLatency` metric that notifies an SNS topic once the average exceeds a defined threshold.

    • A. Incorrect. CloudWatch Lambda Insights applies to AWS Lambda functions, not a SageMaker AI hosted endpoint container, and reviewing diagnostics only after the fact is reactive rather than proactive alerting.
    • B. Incorrect. CloudTrail records control-plane API activity, not inference request latency, and a once-daily scheduled report does not provide the near-real-time notification the team needs.
    • C. Incorrect. Trusted Advisor's low-utilization check flags idle EC2 instances for cost savings on a weekly cadence; it does not monitor per-request inference latency or alert within minutes.
    • D. Correct. A CloudWatch alarm on `ModelLatency` evaluates the metric continuously and can trigger an SNS notification within minutes once the threshold is breached, exactly matching the near-real-time alerting requirement.

    Subdomain 4.2: Monitor and optimize infrastructure and costs

    34.An ML platform team supports five product teams that each run their own SageMaker AI training jobs and endpoints in a shared AWS account. Finance wants a monthly breakdown of exactly how much each product team spent, without splitting the account into separate accounts. What should the platform team implement first?

    1. A.Create a CloudWatch Logs Insights query across all log groups that filters invocation counts by the IAM role each team assumes.
    2. B.Enable AWS X-Ray tracing on every training job and endpoint so that request-level trace spans can be summed into a per-team cost figure.
    3. C.Configure AWS Trusted Advisor to run its fault-tolerance checks weekly and export the results as a proxy for team-level spend.
    4. D.Apply a consistent `team` cost allocation tag to every resource, then activate that tag in AWS Billing and Cost Management.
    Show answer & explanation

    Correct answer: D — Apply a consistent `team` cost allocation tag to every resource, then activate that tag in AWS Billing and Cost Management.

    • A. Incorrect. Logs Insights queries application log data, not billing records, so filtering invocation counts by role would only approximate usage, not actual dollar cost per team.
    • B. Incorrect. X-Ray traces request latency and call chains for troubleshooting performance; it has no concept of billing cost and cannot produce a per-team dollar breakdown.
    • C. Incorrect. Trusted Advisor's fault-tolerance checks assess resiliency configuration, not spend, and are unrelated to producing a team-level cost breakdown.
    • D. Correct. Tagging every resource with a consistent team identifier and activating that tag as a cost allocation tag in Billing and Cost Management lets costs be grouped and reported per team within a single shared account.

    Subdomain 4.2: Monitor and optimize infrastructure and costs

    35.A Lambda-based inference API in front of a SageMaker AI model serves a mobile app with sharp, predictable traffic spikes at 9 AM each weekday. Cold starts during these spikes push p99 latency above the app's SLA. The team wants execution environments pre-initialized ahead of the known spike. Which configuration addresses this directly?

    1. A.Configure provisioned concurrency for the function's alias and use Application Auto Scaling scheduled scaling to raise it before 9 AM each day.
    2. B.Configure reserved concurrency for the function at its current peak value so no other function in the shared account can consume that capacity.
    3. C.Enable AWS X-Ray active tracing on the function so that cold-start segments are visible in the trace map after the spike has passed.
    4. D.Enable CloudWatch Lambda Insights on the function so that cold starts are logged in detail immediately as they occur each morning.
    Show answer & explanation

    Correct answer: A — Configure provisioned concurrency for the function's alias and use Application Auto Scaling scheduled scaling to raise it before 9 AM each day.

    • A. Correct. Provisioned concurrency pre-initializes execution environments so invocations respond without a cold start, and Application Auto Scaling scheduled scaling can raise that provisioned amount ahead of the known 9 AM spike.
    • B. Incorrect. Reserved concurrency only sets a maximum and reserves capacity from the shared pool; it does not pre-initialize execution environments, so cold starts still occur on first invocation.
    • C. Incorrect. X-Ray tracing helps visualize where cold-start time is spent after the fact; it does not pre-warm execution environments to prevent the cold start from occurring.
    • D. Incorrect. Lambda Insights reports on cold starts after they happen for diagnostic purposes; it does not prevent or reduce cold starts during a traffic spike.

    Want the full experience?

    These are just samples. Practice the full AWS Certified Machine Learning Engineer - Associate (MLA-C01) question bank in quiz mode — free, no signup, with domain practice and exam simulation.