CertSafari

    Free AWS Certified Machine Learning Engineer - Associate (MLA-C02) Sample Questions

    35 free sample questions from our bank of 354+, covering every exam domain, with answers and detailed explanations. Updated October 2026.

    Domain 1: Data Preparation for ML and AI

    Subdomain 1.3: Validate data quality and manage bias

    1.A medical imaging dataset has several scans per patient. After a random image-level split, the model scores 99 percent in validation but only 78 percent on new patients. What is the most likely cause and fix?

    1. A.The classes are balanced too evenly across partitions, so remove stratification and let the image order in the manifest decide the split.
    2. B.Scans of one patient sit in both train and validation sets, so split by patient identifier to keep each patient in one partition.
    3. C.The learning rate is too high for the optimizer, so reduce it and keep the same split because the gap comes from unstable gradient updates.
    4. D.The validation set is too small to be reliable, so apply more aggressive image augmentation to the validation images until scores decrease.
    Show answer & explanation

    Correct answer: B — Scans of one patient sit in both train and validation sets, so split by patient identifier to keep each patient in one partition.

    • A. Removing stratification does not address duplicated patients across partitions and may worsen class proportions.
    • B. Group leakage lets the model memorize patient-specific traits; grouping on patient ID keeps validation data truly unseen.
    • C. A high learning rate would hurt training behavior in general; it does not create a gap that appears on new patients only.
    • D. Augmenting validation data does not remove the leakage and distorts the evaluation that should reflect real images.

    Subdomain 1.3: Validate data quality and manage bias

    2.A churn dataset in SageMaker Data Wrangler has 4 percent positive examples across numeric features. The team wants to generate synthetic minority records rather than duplicate rows. Which transform should they use?

    1. A.The Balance data transform with random undersampling, which draws additional minority records by sampling randomly from the minority class.
    2. B.The Featurize text transform with a character-level vectorizer, which expands minority rows into several variants from their string columns.
    3. C.The Balance data transform with the SMOTE option, which interpolates between a minority record and its nearest minority neighbors.
    4. D.The Handle outliers transform with a quantile range, which adds new rows whenever a value falls within the extreme tail of the distribution.
    Show answer & explanation

    Correct answer: C — The Balance data transform with the SMOTE option, which interpolates between a minority record and its nearest minority neighbors.

    • A. Undersampling discards majority rows; it does not create new minority records at all.
    • B. Text featurization produces numeric vectors from strings and does not rebalance numeric records.
    • C. SMOTE builds new synthetic minority rows along lines between neighbors, avoiding exact duplication of existing records.
    • D. Outlier handling edits or removes extreme values and has no function for creating minority class rows.

    Subdomain 1.3: Validate data quality and manage bias

    3.A team prepares a tabular training set in SageMaker Data Wrangler and wants an automated report that flags target leakage, duplicate rows, and anomalous samples before modeling. What should they generate?

    1. A.A SageMaker Model Monitor baseline job, which compares live endpoint traffic with captured statistics to find data problems.
    2. B.The Data Quality and Insights Report, which analyzes the dataset for target leakage, duplicates, anomalous samples, and feature statistics.
    3. C.A SageMaker Clarify explainability report, which computes SHAP values for a deployed endpoint and highlights leakage in served predictions.
    4. D.A Data Wrangler Bias Report on the label column, which reports duplicate rows and leakage alongside the class imbalance of the facet.
    Show answer & explanation

    Correct answer: B — The Data Quality and Insights Report, which analyzes the dataset for target leakage, duplicates, anomalous samples, and feature statistics.

    • A. Model Monitor watches deployed endpoints and is not a pre-training dataset analysis.
    • B. This report summarizes dataset problems including target leakage and duplicate rows and can run a quick model for early signals.
    • C. Explainability reports need a trained model and endpoint, and they do not scan raw data for duplicates.
    • D. The bias report computes metrics such as CI and DPL for a facet and does not cover duplicates or leakage.

    Subdomain 1.1: Collect and store data

    4.Which statement describes the main purpose of the SageMaker Feature Store online store?

    1. A.It detects drift between training and serving feature distributions and raises CloudWatch alarms for any violation found
    2. B.It serves the latest feature values per record identifier with single-digit millisecond latency for inference lookups
    3. C.It orchestrates feature transformation code that runs on a schedule inside SageMaker Pipelines processing steps each day
    4. D.It keeps the full history of feature values in S3 as Parquet so training jobs can rebuild features at past timestamps
    Show answer & explanation

    Correct answer: B — It serves the latest feature values per record identifier with single-digit millisecond latency for inference lookups

    • A. Drift detection is a Model Monitor capability, not a function of the online store.
    • B. The online store keeps the most recent record for each identifier and is built for low-latency reads by a deployed endpoint.
    • C. Transformation orchestration belongs to Pipelines and processing jobs, not the storage layer of a feature group.
    • D. Full history in S3 describes the offline store, which supports point-in-time training queries.

    Subdomain 1.1: Collect and store data

    5.An application team already runs Amazon Aurora PostgreSQL and wants to store 2 million document embeddings next to relational metadata and run similarity queries with low recall loss and fast approximate search. What is the right configuration?

    1. A.Enable the `pgvector` extension with `CREATE EXTENSION vector`, add a `vector(n)` column, and build an HNSW index on it
    2. B.Store the embeddings as JSONB arrays and compute cosine similarity in a PL/pgSQL function over a full table scan
    3. C.Enable Aurora Parallel Query so the cluster computes distance scores for all rows quickly in the storage layer
    4. D.Export the embeddings to DynamoDB and use a global secondary index on the vector attribute for nearest neighbor queries
    Show answer & explanation

    Correct answer: A — Enable the `pgvector` extension with `CREATE EXTENSION vector`, add a `vector(n)` column, and build an HNSW index on it

    • A. pgvector adds the vector type and approximate indexes, and HNSW offers strong recall-speed tradeoffs without a training step.
    • B. Scanning every row in a function cannot meet low-latency requirements at two million vectors.
    • C. Parallel Query accelerates analytic scans, not approximate nearest neighbor search, and does not define a vector type.
    • D. DynamoDB indexes support key lookups and cannot rank items by vector distance.

    Subdomain 1.1: Collect and store data

    6.An ingestion job writes 5 million tiny JSON files, each about 2 KB, to S3 per day. Glue ETL and Athena jobs over the data run slowly due to per-file overhead. Which TWO actions improve performance?(Select 2)

    1. A.Read with `groupFiles` and `groupSize` options in the Glue job so many small files merge into larger input partitions
    2. B.Compact the data into larger Parquet files, in the range of hundreds of MB, when writing the curated layer
    3. C.Switch the bucket to S3 One Zone-IA so object retrieval for each small file is served from one single zone
    4. D.Increase the Athena workgroup data-scan limit so each query is allowed to open more files at the same time
    5. E.Enable S3 Replication to a second bucket so queries can run against the replica's object listing instead
    Show answer & explanation

    Correct answers: A, B — Read with `groupFiles` and `groupSize` options in the Glue job so many small files merge into larger input partitions; Compact the data into larger Parquet files, in the range of hundreds of MB, when writing the curated layer

    • A. Grouping small files cuts task-scheduling and listing overhead during reads.
    • B. Fewer, larger columnar files reduce S3 request counts and let engines read efficiently.
    • C. A storage class change does not remove per-file overhead and may add minimum-size charges.
    • D. Scan limits cap spend per query; they do not reduce small-file overhead.
    • E. Replication duplicates the same small files and only adds storage cost.

    Subdomain 1.2: Perform data transformation, feature engineering, and pre-processing

    7.Business analysts with no coding experience must standardize date formats, remove duplicate rows and profile the quality of CSV files that land in S3, then repeat the same steps on every new file. Which approach fits best?

    1. A.Write a PySpark script in an AWS Glue interactive session that applies the cleanup logic, then schedule it with a Glue trigger
    2. B.Build a recipe in AWS Glue DataBrew's visual interface, profile the dataset, and run the recipe as a scheduled job
    3. C.Run pandas code in an Amazon EMR notebook and ask the analysts to rerun the notebook cells manually whenever a new file arrives
    4. D.Attach an Amazon S3 event notification to a Lambda function that uses regular expressions to rewrite dates and drop duplicates
    Show answer & explanation

    Correct answer: B — Build a recipe in AWS Glue DataBrew's visual interface, profile the dataset, and run the recipe as a scheduled job

    • A. This works technically but demands PySpark skills the analysts do not have. The requirement is a no-code experience.
    • B. DataBrew offers a no-code visual recipe editor with built-in profiling. Recipes can be saved and run as scheduled jobs against new data.
    • C. A notebook needs coding knowledge and manual reruns, so it is neither no-code nor repeatable on a schedule.
    • D. A Lambda function needs custom code and cannot profile data quality. It would also struggle with large files and deduplication across rows.

    Subdomain 1.2: Perform data transformation, feature engineering, and pre-processing

    8.A demand forecasting dataset has a single timestamp column. Which TWO feature engineering steps best let the model learn daily and weekly seasonality?(Select 2)

    1. A.Encode hour of day as a plain integer from zero to twenty-three, which keeps midnight and eleven at night adjacent in distance
    2. B.Split the timestamp into separate hour of day, day of week and month features so seasonal effects become explicit inputs
    3. C.Keep only raw epoch seconds, min-max scaled, and drop the calendar parts so the model infers weekly cycles from magnitude alone
    4. D.One-hot encode the full timestamp so each distinct second becomes its own column, giving the model an exact time indicator
    5. E.Encode hour of day as a sine and cosine pair so that 23:00 and 00:00 end up close together in feature space
    Show answer & explanation

    Correct answers: B, E — Split the timestamp into separate hour of day, day of week and month features so seasonal effects become explicit inputs; Encode hour of day as a sine and cosine pair so that 23:00 and 00:00 end up close together in feature space

    • A. As a plain integer, hour zero and hour twenty-three are the farthest apart even though they are neighbors in time. This misrepresents the cycle.
    • B. Splitting a timestamp exposes calendar components that drive seasonality. The model can then learn patterns by hour, weekday and month.
    • C. A monotonically growing epoch value carries no repeating pattern. A model cannot easily derive weekly cycles from it.
    • D. Nearly every timestamp is unique, so this creates millions of sparse columns that generalize to nothing. It carries no seasonal structure.
    • E. Cyclical encoding preserves the circular nature of time. Plain integers would place the end and start of the day far apart.

    Subdomain 1.2: Perform data transformation, feature engineering, and pre-processing

    9.A data engineering team must find and hash columns containing emails and phone numbers while running an AWS Glue ETL job over tables in the Data Catalog. Which Glue capability should they use?

    1. A.Run a Glue crawler with custom classifiers, which anonymize detected sensitive columns while they populate the Data Catalog
    2. B.Create a Glue Data Quality ruleset in DQDL whose rules mask any column that contains an email address or phone number
    3. C.Apply the FindMatches machine learning transform, which hashes columns holding personal identifiers during the deduplication pass
    4. D.Add the Detect Sensitive Data transform in Glue Studio, select entity types, and choose an action that hashes or replaces values
    Show answer & explanation

    Correct answer: D — Add the Detect Sensitive Data transform in Glue Studio, select entity types, and choose an action that hashes or replaces values

    • A. Classifiers help infer schemas and formats. They never change the underlying data.
    • B. Data Quality rules evaluate and report on data, but do not alter or mask values. They would only flag the issues.
    • C. FindMatches links duplicate records and does not mask fields. Its purpose is entity resolution.
    • D. The Detect Sensitive Data transform identifies sensitive data in columns and can replace values with a fixed string or a hash. It runs inside the ETL job.

    Domain 2: ML Model and Foundation Model (FM) Development

    Subdomain 2.2: Train, fine-tune, and customize models for ML and AI solutions

    10.A regression model has a high RMSE on both the training set and the validation set, and the two errors are nearly identical. The model is a shallow linear model with strong L2 regularization on a dataset with complex nonlinear relationships. What should the engineer do?

    1. A.Increase the L2 penalty further and collect more training rows so that the model variance decreases on the validation set.
    2. B.Apply dropout layers to the input features so that the model becomes robust to noisy columns in the training dataset.
    3. C.Stop training earlier using validation loss monitoring so that the model stops before memorizing noise in the training data.
    4. D.Add polynomial features or move to a higher-capacity model, and weaken the regularization so the model can fit the patterns.
    Show answer & explanation

    Correct answer: D — Add polynomial features or move to a higher-capacity model, and weaken the regularization so the model can fit the patterns.

    • A. Variance is not the problem here; stronger regularization would restrict an already underfitting model further.
    • B. Dropout is a regularization technique that further reduces effective capacity, which worsens underfitting rather than correcting it.
    • C. Early stopping combats overfitting, but training error is already high, so stopping earlier would only leave the model less fitted.
    • D. Similar high error on both sets signals underfitting; raising capacity and relaxing regularization lets the model capture nonlinear structure.

    Subdomain 2.2: Train, fine-tune, and customize models for ML and AI solutions

    11.An application uses a large teacher model on Amazon Bedrock for customer-intent classification at high accuracy, but inference costs and latency are too high for 2 million daily requests. The team wants a cheaper model with comparable quality on this task. Which technique should they use?

    1. A.Use Bedrock continued pre-training on the smaller model with the intent prompts so that it matches the larger model on classification quality.
    2. B.Cache each prompt in Amazon DynamoDB and serve cached answers to all requests, retraining the teacher only when the cache misses occur.
    3. C.Use Bedrock model distillation, where the teacher generates responses to the team's prompts and a smaller student model is fine-tuned on them.
    4. D.Raise the `temperature` inference parameter of the teacher model so the generated tokens finish faster and cost drops per request.
    Show answer & explanation

    Correct answer: C — Use Bedrock model distillation, where the teacher generates responses to the team's prompts and a smaller student model is fine-tuned on them.

    • A. Continued pre-training uses unlabeled text for domain knowledge and gives no supervised signal on the intent classification outputs.
    • B. Caching helps only for repeated identical prompts; most customer messages are unique, and the underlying teacher cost remains for each miss.
    • C. Distillation transfers task behavior from a large teacher to a smaller student, cutting per-request cost and latency while preserving much of the accuracy.
    • D. Temperature changes the randomness of sampling, not the model size or compute per token, so cost and latency do not decrease.

    Subdomain 2.2: Train, fine-tune, and customize models for ML and AI solutions

    12.A RAG application stores 200 million chunks in Amazon OpenSearch Service using 1,024-dimension vectors from Amazon Titan Text Embeddings V2. Storage and search latency costs are high, while retrieval quality is already well above target. What should the team do first?

    1. A.Re-embed the corpus with the 256-dimension output setting of the same embedding model, then compare retrieval metrics to the current index.
    2. B.Disable approximate nearest neighbor search and use exact k-NN brute force so the engine can safely compress the stored vectors in memory.
    3. C.Fine-tune a new embedding model with 4,096 dimensions so that each chunk is represented with more detail than the current vectors allow.
    4. D.Query the index with the 1,024-dimension model but index with 256 dimensions by truncating stored vectors, leaving the query vectors untouched.
    Show answer & explanation

    Correct answer: A — Re-embed the corpus with the 256-dimension output setting of the same embedding model, then compare retrieval metrics to the current index.

    • A. Titan Text Embeddings V2 supports smaller output dimensions, which shrink index size and speed up search with limited accuracy loss.
    • B. Exact k-NN increases compute per query and does not compress vectors, so latency and cost would rise instead of fall.
    • C. Higher dimensions would increase storage and latency, working against the stated goal of reducing cost.
    • D. Query and document vectors must share the same dimension and embedding space; mismatched sizes cannot be compared by the vector search engine.

    Subdomain 2.1: Choose appropriate modeling approaches for ML and AI solutions

    13.A parts-catalog assistant uses a Bedrock knowledge base. Searches for exact part numbers like "XJ-4471-B" keep returning semantically similar but wrong parts. How should retrieval be adjusted?

    1. A.Fine-tune the generation model on the catalog so it memorizes every part number inside its weights
    2. B.Raise the model temperature so the generator explores more candidate part numbers before writing its answer
    3. C.Reduce embedding dimensionality so near-identical identifiers collapse into the same cluster in the vector space
    4. D.Switch to hybrid search so keyword matching on exact part numbers combines with semantic vector similarity
    Show answer & explanation

    Correct answer: D — Switch to hybrid search so keyword matching on exact part numbers combines with semantic vector similarity

    • A. Memorizing identifiers in weights is unreliable and does not fix the retrieval step that supplies the wrong passages.
    • B. Temperature only changes sampling randomness at generation time and cannot correct which passages were retrieved.
    • C. Collapsing vectors makes similar identifiers harder to tell apart, worsening the exact-match problem.
    • D. Hybrid search merges lexical matching with vector search, so identifiers that embeddings treat as similar are still matched exactly.

    Subdomain 2.1: Choose appropriate modeling approaches for ML and AI solutions

    14.A nightly 20-hour GPU training job can tolerate interruptions, and the team wants to cut training cost without changing the algorithm. What should they do?

    1. A.Enable managed spot training with an S3 checkpoint location so interrupted jobs resume from the last checkpoint
    2. B.Reserve On-Demand Capacity for the training instances, which SageMaker bills at a discount of up to 90 percent
    3. C.Turn off checkpointing so each run avoids S3 write overhead and finishes before any Spot interruption can occur
    4. D.Replace the GPU instances with CPU instances of equal vCPU count so the unchanged code runs at a lower hourly price
    Show answer & explanation

    Correct answer: A — Enable managed spot training with an S3 checkpoint location so interrupted jobs resume from the last checkpoint

    • A. Managed Spot uses spare capacity at a steep discount, and checkpoints let an interrupted job resume instead of restarting.
    • B. Capacity reservations guarantee availability but are billed at On-Demand rates; the up-to-90-percent discount applies to Spot.
    • C. Without checkpoints, an interruption restarts the 20-hour job from scratch, wasting the savings.
    • D. CPU instances would train a GPU-optimized job far more slowly, so total cost would likely rise.

    Subdomain 2.1: Choose appropriate modeling approaches for ML and AI solutions

    15.A factory streams unlabeled vibration readings from thousands of sensors and has never recorded labeled failures. It wants to flag unusual readings as they arrive. Which approach fits?

    1. A.Train a Linear Learner binary classifier using failure labels for events the plant has never recorded
    2. B.Train an XGBoost classifier with the sensor's own reading as the target so it flags readings that differ
    3. C.Train a BlazingText model in Skip-gram mode on the readings to learn embeddings of failures
    4. D.Train a SageMaker Random Cut Forest model on the readings and flag points with high anomaly scores
    Show answer & explanation

    Correct answer: D — Train a SageMaker Random Cut Forest model on the readings and flag points with high anomaly scores

    • A. Supervised classification needs labeled failures, which do not exist in this scenario.
    • B. Predicting the input as its own target teaches nothing about anomalies and has no failure labels to learn from.
    • C. BlazingText produces word embeddings from text and does not suit numeric sensor streams.
    • D. Random Cut Forest is an unsupervised anomaly detection algorithm that needs no labeled failures.

    Subdomain 2.3: Analyze and evaluate the performance of ML and AI systems

    16.An engineer sets up SageMaker Model Monitor data quality monitoring for a real-time endpoint and runs a baselining job on the training dataset. Which output does the baselining job produce for later comparison against captured traffic?

    1. A.A `model.tar.gz` archive containing the trained weights and the inference script, which each scheduled monitoring run extracts for scoring.
    2. B.A Feature Store feature group that holds every training row, which scheduled monitoring runs join against to compute feature distance.
    3. C.A `statistics.json` file with per-feature statistics and a `constraints.json` file with suggested constraints, written to S3.
    4. D.A CloudWatch dashboard JSON that lists thresholds for CPU and latency, which the monitoring schedule evaluates against endpoint invocation metrics.
    Show answer & explanation

    Correct answer: C — A `statistics.json` file with per-feature statistics and a `constraints.json` file with suggested constraints, written to S3.

    • A. That archive is the trained model package, not a monitoring baseline; baselining analyzes the dataset rather than packaging the model.
    • B. Feature Store is not the baseline output; baselining writes statistics and constraints files to S3 instead of creating a feature group.
    • C. The baseline job computes dataset statistics and suggests constraints, and later monitoring executions compare captured data against these two files to flag violations.
    • D. Dashboards track infrastructure metrics; data quality baselines describe feature distributions and constraints, not CPU or latency thresholds.

    Subdomain 2.3: Analyze and evaluate the performance of ML and AI systems

    17.A media company generates short summaries of news articles with a foundation model and has human-written reference summaries. They want an automated metric that measures how much of the reference content appears in each generated summary. Which should they use?

    1. A.ROUGE, which measures n-gram and longest-common-subsequence overlap, with an emphasis on recall of reference content.
    2. B.Word error rate, which counts insertions, deletions and substitutions needed to turn the generated summary into the reference text.
    3. C.Perplexity of the generated summary under the model, where a lower value shows the summary matches the reference content closely.
    4. D.Mean reciprocal rank of the reference summary among the top results returned for the article in a vector search.
    Show answer & explanation

    Correct answer: A — ROUGE, which measures n-gram and longest-common-subsequence overlap, with an emphasis on recall of reference content.

    • A. ROUGE was designed for summarization and reports recall-oriented overlap between generated and reference text.
    • B. Word error rate is the standard for speech recognition transcripts and penalizes legitimate rewording in summaries.
    • C. Perplexity measures how well a model predicts text, not overlap with a reference summary.
    • D. Mean reciprocal rank evaluates ranked retrieval and says nothing about the content coverage of a generated summary.

    Subdomain 2.3: Analyze and evaluate the performance of ML and AI systems

    18.A production RAG application on Bedrock Knowledge Bases must be monitored so that drops in retrieval quality are noticed before users complain. Which TWO actions provide this monitoring?(Select 2)

    1. A.Enable knowledge base logging to CloudWatch Logs or S3 so queries, retrieved chunks and relevance scores can be analyzed over time.
    2. B.Rely on endpoint 5xx error alarms alone, since an HTTP success status confirms that retrieved passages were relevant to each question.
    3. C.Run scheduled retrieval evaluation jobs on a golden question set or sampled logs, and publish the scores to CloudWatch alarms.
    4. D.Attach a Model Monitor bias drift schedule to the vector store so changes in embedding dimensions raise a retrieval accuracy alert.
    5. E.Disable logging of retrieved passages to lower storage costs and track only the monthly invoice for embedding model usage.
    Show answer & explanation

    Correct answers: A, C — Enable knowledge base logging to CloudWatch Logs or S3 so queries, retrieved chunks and relevance scores can be analyzed over time.; Run scheduled retrieval evaluation jobs on a golden question set or sampled logs, and publish the scores to CloudWatch alarms.

    • A. Logged queries and retrieved chunks with scores provide the data needed to spot rising empty-result rates or falling relevance.
    • B. A successful HTTP response says nothing about relevance; retrieval can return poor chunks with no errors raised.
    • C. Recurring evaluations against known questions produce relevance and coverage trends that alarms can act on.
    • D. Bias drift monitoring covers deployed model endpoints and fairness metrics, not retrieval accuracy of a knowledge base.
    • E. Cost data does not show retrieval quality, and removing the retrieved passages from logs removes the evidence needed to diagnose it.

    Domain 3: Deployment and Orchestration of ML and AI Workflows

    Subdomain 3.1: Manage deployment infrastructure for ML and AI model types

    19.A SaaS provider trains a separate XGBoost model for each of 6,000 customers. All models use the same container and framework, each model is called rarely, and total artifacts exceed what fits in memory on one instance. Which design is MOST cost-effective?

    1. A.Create 6,000 separate real-time endpoints with one ml.t3.medium instance each so every customer's model is isolated and can be scaled independently.
    2. B.Host all models on one SageMaker multi-model endpoint and pass `TargetModel` in each `InvokeEndpoint` call so artifacts load from S3 on demand.
    3. C.Package all 6,000 artifacts into one model.tar.gz file on a single real-time endpoint and have the inference script load every model into memory when the container starts.
    4. D.Build one SageMaker inference pipeline that chains all 6,000 models as serial containers so each request flows through every model and the relevant output is picked at the end.
    Show answer & explanation

    Correct answer: B — Host all models on one SageMaker multi-model endpoint and pass `TargetModel` in each `InvokeEndpoint` call so artifacts load from S3 on demand.

    • A. Thousands of dedicated endpoints each pay for at least one running instance, which is far more expensive than sharing hosting resources for rarely used models.
    • B. A multi-model endpoint shares one container and fleet across many same-framework models, loading artifacts lazily from S3 and evicting idle ones from memory. `TargetModel` selects the model per request.
    • C. The combined artifacts exceed instance memory, so loading everything at startup fails, and a single archive of that size makes startup and updates impractical.
    • D. Inference pipelines chain a small number of containers serially on one endpoint and cannot hold thousands. Running every model for every request would also be wasteful.

    Subdomain 3.1: Manage deployment infrastructure for ML and AI model types

    20.A bank builds a Bedrock multi-agent system with a billing agent, a fraud agent and a card-services agent, each owned and versioned by a different team. Simple single-topic questions must be answered with the lowest latency. Which TWO configuration choices are appropriate?(Select 2)

    1. A.Create a supervisor agent with multi-agent collaboration enabled and attach each team's agent as a collaborator through its alias.
    2. B.Merge all three agents' instructions into one monolithic agent with every action group attached, since collaborators cannot be versioned by separate teams.
    3. C.Set the supervisor's mode to supervisor with routing, so clear single-topic requests go straight to one collaborator, skipping full orchestration.
    4. D.Attach each collaborator to the supervisor as a knowledge base, so the supervisor retrieves answers from the other agents' stored conversation transcripts.
    5. E.Chain the three agents with a hard-coded sequence in a Lambda function that calls `InvokeAgent` on each in turn for every user message.
    Show answer & explanation

    Correct answers: A, C — Create a supervisor agent with multi-agent collaboration enabled and attach each team's agent as a collaborator through its alias.; Set the supervisor's mode to supervisor with routing, so clear single-topic requests go straight to one collaborator, skipping full orchestration.

    • A. A supervisor coordinates collaborators attached by alias, so each team can version and update its own agent independently.
    • B. Collaborators do support aliases and separate ownership, and a monolithic agent loses independent versioning and enlarges the prompt.
    • C. Routing mode forwards simple requests directly to the matching collaborator, cutting latency, and falls back to full orchestration for complex multi-topic questions.
    • D. Knowledge bases index documents. Linking an agent as one does not make it a collaborator that can act on requests.
    • E. A fixed sequence invokes every agent each time, which raises latency for simple requests and removes dynamic delegation.

    Subdomain 3.1: Manage deployment infrastructure for ML and AI model types

    21.Users ask compound questions such as "Compare the warranty terms and shipping policies for the Pro and Lite plans." A Bedrock knowledge base returns passages about only one topic. Which TWO changes improve retrieval for these questions?(Select 2)

    1. A.Replace the vector store with an Amazon DynamoDB table keyed by plan name so questions are answered by exact key lookup.
    2. B.Tag documents with metadata such as plan and document type, then apply metadata filters so each sub-query is scoped to the right source.
    3. C.Set the generation temperature to 0 so the model reconstructs the missing topic passages from its training data instead of retrieved text.
    4. D.Merge every document into a single large file before ingestion so that all plans appear in the same chunk for each query.
    5. E.Enable query decomposition in the orchestration configuration so the question is split into sub-queries retrieved separately.
    Show answer & explanation

    Correct answers: B, E — Tag documents with metadata such as plan and document type, then apply metadata filters so each sub-query is scoped to the right source.; Enable query decomposition in the orchestration configuration so the question is split into sub-queries retrieved separately.

    • A. DynamoDB key lookups cannot interpret free-form compound questions or semantically match relevant policy passages.
    • B. Metadata filtering narrows candidates to the right plan and document category, improving precision for each sub-query.
    • C. Temperature does not influence retrieval, and recalling policy text from training data risks wrong facts.
    • D. Merging documents blurs distinctions between plans and produces oversized chunks, which hurts precision.
    • E. Query decomposition breaks a compound question into smaller queries, each retrieving its own relevant passages for the final answer.

    Subdomain 3.2: Provision and configure resources for ML and AI workloads based on existing architecture and requirements

    22.A team packages its own inference server into a custom Docker image for a SageMaker real-time endpoint. Health checks fail and the endpoint stays in Creating until it times out. Which container behavior does SageMaker require for hosting?

    1. A.Listen on port 80 and answer GET /health with HTTP 200, while accepting inference requests as POST /predict on that same port.
    2. B.Listen on port 8080 and answer GET /healthz with HTTP 204, while accepting inference requests as POST /predict through the container ENTRYPOINT.
    3. C.Listen on port 8080 and answer GET /ping with HTTP 200, while accepting inference requests as POST /invocations on the same port.
    4. D.Listen on port 8501 and answer GET /v1/models with HTTP 200, while accepting inference requests as POST /v1/models:predict on that port.
    Show answer & explanation

    Correct answer: C — Listen on port 8080 and answer GET /ping with HTTP 200, while accepting inference requests as POST /invocations on the same port.

    • A. SageMaker does not probe /health or route to /predict. Using these paths and port 80 leaves the health check unanswered.
    • B. The port is correct but the health path and route are not recognized. SageMaker requires /ping returning 200 and /invocations for inference.
    • C. The SageMaker hosting contract expects the container to serve /ping and /invocations on port 8080. A healthy /ping response is what moves the endpoint to InService.
    • D. Those are TensorFlow Serving defaults. SageMaker still calls /ping and /invocations on 8080 unless a wrapper translates them.

    Subdomain 3.2: Provision and configure resources for ML and AI workloads based on existing architecture and requirements

    23.An application in a VPC with no internet access must call a SageMaker real-time endpoint. Which configuration lets InvokeEndpoint calls stay on the AWS network?

    1. A.Place a public Application Load Balancer in front of the endpoint and restrict its HTTPS listener to the application subnet CIDR range.
    2. B.Create a gateway VPC endpoint for the sagemaker.runtime service and add it to the route tables of the application subnets used by the callers.
    3. C.Create an interface VPC endpoint for com.amazonaws.eu-central-1.sagemaker.runtime and enable private DNS for the application.
    4. D.Create an interface VPC endpoint for com.amazonaws.eu-central-1.sagemaker.api and enable private DNS for the application's invocations.
    Show answer & explanation

    Correct answer: C — Create an interface VPC endpoint for com.amazonaws.eu-central-1.sagemaker.runtime and enable private DNS for the application.

    • A. SageMaker endpoints cannot be registered as ALB targets, and a public load balancer would not provide private connectivity.
    • B. Gateway endpoints exist only for Amazon S3 and Amazon DynamoDB. SageMaker Runtime is reached through an interface endpoint.
    • C. InvokeEndpoint calls go to the SageMaker Runtime service. An interface endpoint with private DNS resolves the regular hostname to private addresses in the VPC.
    • D. The sagemaker.api service covers control-plane calls such as CreateEndpoint. Invocation uses the separate runtime service.

    Subdomain 3.2: Provision and configure resources for ML and AI workloads based on existing architecture and requirements

    24.A multi-tenant support assistant uses a single Bedrock knowledge base holding documents for 40 product lines. Answers must only draw on the product line selected by the user, with no separate knowledge base per line. What should the team do?

    1. A.Tag each S3 object with a productLine object tag and let the knowledge base honor those tags when retrieving chunks for each query.
    2. B.Prefix every user question with the product line name so the embedding model naturally favors documents from that line.
    3. C.Add a `<file>.metadata.json` per document with a productLine attribute, then apply an equals filter on it at retrieval.
    4. D.Rename every source file with the product line as a prefix and set the number of results to one to avoid mixing lines.
    Show answer & explanation

    Correct answer: C — Add a `<file>.metadata.json` per document with a productLine attribute, then apply an equals filter on it at retrieval.

    • A. Knowledge bases do not use S3 object tags for filtering. Filtering relies on metadata files in the data source.
    • B. Prompt wording only shifts similarity scores and does not enforce a boundary, so other product lines can still be retrieved.
    • C. Metadata files uploaded next to the source documents become filterable attributes. A retrieval filter restricts results to matching documents.
    • D. File names are not a filter, and a single result limits recall without guaranteeing the right product line.

    Subdomain 3.3: Implement automated orchestration and continuous integration and continuous delivery (CI/CD) pipelines for MLOps and AI workloads

    25.Auditors ask the team to show, for any production endpoint, which training data, source commit and container image produced the deployed model. Which approach gives repeatable, auditable evidence?

    1. A.Store each model.tar.gz in S3 under a date-stamped key and keep a spreadsheet that maps each key to the dataset and commit used in training.
    2. B.Register each trained model as a version in a model package group with image digest, data and commit metadata, and enable ML Lineage Tracking.
    3. C.Rely on CloudTrail records of CreateModel and CreateEndpoint calls, which list the S3 artifact location used for each deployment over time.
    4. D.Tag each endpoint with a semantic version string and keep the training code in a Git tag carrying the same version number for traceability.
    Show answer & explanation

    Correct answer: B — Register each trained model as a version in a model package group with image digest, data and commit metadata, and enable ML Lineage Tracking.

    • A. A spreadsheet is a manual record that can drift from reality. It gives no enforced version history or linkage to jobs.
    • B. The registry stores immutable versions with their inference specification. Lineage tracking links each version to its datasets, jobs and artifacts for audits.
    • C. CloudTrail shows who deployed and which S3 path was used. It does not record training data, commit or image provenance.
    • D. Tags help humans but record nothing about the data or image used, and tags can be edited after the fact.

    Subdomain 3.3: Implement automated orchestration and continuous integration and continuous delivery (CI/CD) pipelines for MLOps and AI workloads

    26.A workflow runs a Glue job, a SageMaker training job, an evaluation, and then a deployment. Wait times for training vary from minutes to hours. What is the simplest way to avoid polling code?

    1. A.Insert a Wait state set to the average training duration and then run the evaluation state directly after it finishes waiting.
    2. B.Use an Express workflow with a Lambda function that loops, sleeping and calling DescribeTrainingJob until the training status is Completed.
    3. C.Start the training job from an EventBridge rule and have the evaluation Lambda function subscribe to that same rule as another target.
    4. D.Use a Step Functions Standard workflow and call the SageMaker CreateTrainingJob integration with the .sync service integration pattern.
    Show answer & explanation

    Correct answer: D — Use a Step Functions Standard workflow and call the SageMaker CreateTrainingJob integration with the .sync service integration pattern.

    • A. A fixed wait either proceeds before training ends or wastes time, because duration varies by run.
    • B. Express workflows are limited to five minutes and a sleeping Lambda loop is the polling code to avoid.
    • C. Both targets would run at the same time, so evaluation starts before the training job produces a model.
    • D. The .sync pattern makes Step Functions wait until the training job finishes, with Retry and Catch for errors, and no polling code.

    Subdomain 3.3: Implement automated orchestration and continuous integration and continuous delivery (CI/CD) pipelines for MLOps and AI workloads

    27.A CI stage for a traditional ML project must validate each change before promotion. Which THREE tests belong in the pipeline?(Select 3)

    1. A.Retrain the full model from scratch inside every commit's unit test job, so each commit is verified end to end.
    2. B.Use accuracy measured on the training dataset as the promotion metric, as it is available immediately after fitting.
    3. C.Skip pre-production checks and send 100% of production traffic to the candidate, relying on rollback if errors are reported.
    4. D.Evaluate the candidate on a held-out dataset and block promotion if the metric falls below the current production baseline.
    5. E.Deploy to a staging endpoint and run integration tests that send schema-valid and malformed payloads and check response codes.
    6. F.Run pytest unit tests on the feature engineering and preprocessing functions in a CodeBuild stage on every commit.
    Show answer & explanation

    Correct answers: D, E, F — Evaluate the candidate on a held-out dataset and block promotion if the metric falls below the current production baseline.; Deploy to a staging endpoint and run integration tests that send schema-valid and malformed payloads and check response codes.; Run pytest unit tests on the feature engineering and preprocessing functions in a CodeBuild stage on every commit.

    • A. Full training per commit is slow and expensive for a unit test and belongs in the training pipeline.
    • B. Training-set accuracy rewards overfitting and cannot identify generalization problems.
    • C. Exposing every user to an unvalidated model is the opposite of automated testing before release.
    • D. A baseline gate stops regressions that pass code tests but reduce prediction quality.
    • E. Staging tests validate the container, serialization and endpoint behavior in a production-like setup.
    • F. Unit tests find logic errors in data transformation code cheaply and quickly, before any compute is spent on training.

    Domain 4: Operating, Monitoring, and Securing ML and AI Solutions

    Subdomain 4.1: Monitor ML and AI model inference and performance

    28.A support-chat team wants to score thousands of responses for helpfulness, coherence and faithfulness each week without paying for human reviewers. Programmatic metrics like BERTScore do not capture these qualities. Which Amazon Bedrock capability applies?

    1. A.An automatic evaluation job that selects the accuracy metric and compares each response with the reference using embedding similarity.
    2. B.An invocation logging configuration that stores each response in S3 so that an analyst can skim a random sample once a month.
    3. C.A model evaluation job that uses an LLM-as-a-judge, where an evaluator model scores each response against built-in or custom quality metrics.
    4. D.A Bedrock Guardrails policy with a denied-topic filter that blocks any response judged unhelpful before it is returned to the user.
    Show answer & explanation

    Correct answer: C — A model evaluation job that uses an LLM-as-a-judge, where an evaluator model scores each response against built-in or custom quality metrics.

    • A. Incorrect. That measures closeness to a reference answer, which is not a judgment of helpfulness, coherence or faithfulness.
    • B. Incorrect. Logging gives raw data but no scoring, and monthly manual sampling neither scales nor avoids reviewer cost.
    • C. Correct. Judge-model evaluations rate qualitative dimensions at scale and return scores with explanations for each prompt.
    • D. Incorrect. Guardrails enforce safety and policy boundaries on content; they do not produce quality scores across weekly batches.

    Subdomain 4.1: Monitor ML and AI model inference and performance

    29.A retailer's churn model has no labels at inference time, and customer behavior is changing. Which TWO approaches detect covariate shift in the incoming features? (Select TWO.)(Select 2)

    1. A.Run a Model Monitor data quality schedule against the baseline from the training data and review the violations report for drifted features.
    2. B.Run Clarify pre-training bias analysis on the live requests to confirm the distribution of each feature matches the training set.
    3. C.Alarm on the endpoint's Invocations metric, because a change in request counts implies the feature values changed as well.
    4. D.Add a SageMaker Processing step that computes the population stability index or a KS statistic between captured features and training data.
    5. E.Create a model quality schedule and alert on accuracy, since accuracy drops are the earliest sign that the feature distribution moved.
    Show answer & explanation

    Correct answers: A, D — Run a Model Monitor data quality schedule against the baseline from the training data and review the violations report for drifted features.; Add a SageMaker Processing step that computes the population stability index or a KS statistic between captured features and training data.

    • A. Correct. Data quality monitoring compares live feature statistics with training baselines and needs no labels.
    • B. Incorrect. Pre-training bias analysis measures imbalance in labeled training data, not changes in live feature distributions.
    • C. Incorrect. Request counts can stay constant while the feature values change, and the reverse is equally common.
    • D. Correct. A custom statistical test provides a drift score per feature and can be scheduled in a pipeline.
    • E. Incorrect. Model quality needs labels the team does not have, and it detects performance changes later than distribution shifts.

    Subdomain 4.3: Secure ML and AI workloads and model endpoints

    30.An administrator wants data scientists to start SageMaker AI training jobs using only a specific execution role named ml-train-role. Several other roles in the account trust sagemaker.amazonaws.com and have broader permissions. Which IAM statement prevents privilege escalation through the training API?

    1. A.Allow sts:AssumeRole on every role that trusts sagemaker.amazonaws.com so that the training job inherits the narrowest role it can find.
    2. B.Allow iam:GetRole on all roles in the account so that SageMaker can verify role permissions before it accepts a training job request.
    3. C.Allow iam:PassRole only on the ml-train-role ARN, with the condition key iam:PassedToService set to sagemaker.amazonaws.com.
    4. D.Allow sagemaker:CreateTrainingJob on all resources and rely on the SageMaker service to reject execution roles that exceed the caller's permissions.
    Show answer & explanation

    Correct answer: C — Allow iam:PassRole only on the ml-train-role ARN, with the condition key iam:PassedToService set to sagemaker.amazonaws.com.

    • A. Granting sts:AssumeRole on every SageMaker-trusted role widens the attack surface. Training jobs use the passed execution role, not roles assumed by the user.
    • B. iam:GetRole only reads role metadata. It does not control which roles a caller can attach to a job, so escalation remains possible.
    • C. A user needs iam:PassRole to hand an execution role to SageMaker. Restricting that permission to one role ARN and the SageMaker service principal stops users from passing more privileged roles.
    • D. SageMaker does not compare the execution role to the caller's permissions. Without a PassRole restriction, a user allowed to pass any role could borrow broader privileges.

    Subdomain 4.3: Secure ML and AI workloads and model endpoints

    31.A team plans to deploy a third-party model container from AWS Marketplace to a SageMaker AI real-time endpoint. Security requires that the container cannot make any outbound network calls, even if it contains malicious code. Which setting satisfies this?

    1. A.Place the endpoint in a public subnet and attach a security group with no inbound rules, which blocks every outbound connection from the container.
    2. B.Enable inter-container traffic encryption on the model so that any outbound connection initiated by the container is rejected by the platform.
    3. C.Attach an IAM policy to the execution role that denies all s3:* and sts:* actions, which prevents the container from opening network sockets.
    4. D.Set EnableNetworkIsolation to true on the model so the container runs without outbound network access, including calls to AWS services.
    Show answer & explanation

    Correct answer: D — Set EnableNetworkIsolation to true on the model so the container runs without outbound network access, including calls to AWS services.

    • A. Security groups are stateful, so an empty inbound rule set does not block outbound connections. A public subnet also adds exposure.
    • B. Inter-container encryption protects traffic between training containers. It does not block outbound connections from a single container.
    • C. IAM denies restrict authorization for AWS API calls. They do not stop a container from opening arbitrary network connections to other hosts.
    • D. Network isolation prevents the container from making outbound calls, which contains a malicious or compromised third-party image. SageMaker still handles artifact download itself.

    Subdomain 4.3: Secure ML and AI workloads and model endpoints

    32.A distributed SageMaker AI training job processes regulated data. The team must encrypt the attached training volumes and the output artifacts with a customer managed KMS key, and also encrypt traffic between the training instances. Which THREE settings should be applied?(Select 3)

    1. A.Specify a customer managed key as the VolumeKmsKeyId for the training instances so attached storage is encrypted.
    2. B.Turn on network isolation for the training job, which is the setting that encrypts traffic between nodes of a distributed job.
    3. C.Set the KmsKeyId in the output data configuration so that model artifacts written to S3 are encrypted with the same key.
    4. D.Enable inter-container traffic encryption on the training job so communication between distributed nodes is encrypted.
    5. E.Store the KMS key material in the training container image so the job can decrypt its data without calling AWS KMS.
    6. F.Enable S3 Transfer Acceleration on the output bucket so that artifacts are encrypted in transit between instances.
    Show answer & explanation

    Correct answers: A, C, D — Specify a customer managed key as the VolumeKmsKeyId for the training instances so attached storage is encrypted.; Set the KmsKeyId in the output data configuration so that model artifacts written to S3 are encrypted with the same key.; Enable inter-container traffic encryption on the training job so communication between distributed nodes is encrypted.

    • A. The volume KMS key parameter encrypts the ML storage volume attached to the training instances.
    • B. Network isolation blocks outbound network access and is not an encryption setting.
    • C. The output configuration key encrypts what the job writes to S3, giving a consistent customer-managed key for artifacts.
    • D. Inter-container traffic encryption protects data exchanged between nodes in distributed training, which is off by default.
    • E. Embedding key material in an image exposes it to anyone who can pull the image, and KMS keys are not meant to be exported.
    • F. Transfer Acceleration speeds up transfers over long distances and does not encrypt traffic between training instances.

    Subdomain 4.2: Optimize and manage ML and AI infrastructure costs and performance

    33.A retail company serves a gradient-boosted tree model built with XGBoost from a real-time SageMaker endpoint. Inference is CPU-bound, latency targets are comfortably met, and the team wants to cut the instance bill without adding accelerators. What should they do?

    1. A.Move the endpoint to ml.p3.2xlarge GPU instances, because GPUs always give the lowest cost per prediction for tree-based models.
    2. B.Recompile the model with the AWS Neuron SDK and host it on ml.inf1 instances, which accelerate tree ensembles on Inferentia.
    3. C.Redeploy on ml.g5.xlarge instances so the CUDA cores can be shared across concurrent prediction requests from many clients.
    4. D.Deploy on Graviton-based ml.c7g instances with an ARM64-compatible XGBoost container for better CPU price-performance.
    Show answer & explanation

    Correct answer: D — Deploy on Graviton-based ml.c7g instances with an ARM64-compatible XGBoost container for better CPU price-performance.

    • A. GPU instances carry a higher hourly rate, and tree-ensemble inference on CPU rarely benefits enough to offset it.
    • B. Neuron targets neural network workloads; gradient-boosted trees are not a supported Inferentia use case.
    • C. A g5 instance adds GPU cost that a CPU-bound XGBoost model will not use, so the bill grows without a latency benefit.
    • D. Graviton-based ml.c7g instances run CPU inference at a lower price than comparable x86 instances, provided the container image is built for ARM64.

    Subdomain 4.2: Optimize and manage ML and AI infrastructure costs and performance

    34.A finance team wants to be warned about unexpected Amazon Bedrock spend spikes and to block further spend automatically once a hard monthly limit is reached. Which TWO should they configure?(Select 2)

    1. A.Create an AWS Cost Anomaly Detection monitor on Amazon Bedrock with an alert subscription through SNS or email.
    2. B.Create an AWS Budgets cost budget filtered to Amazon Bedrock with a budget action applying an IAM deny policy at 100%.
    3. C.Request a Service Quotas decrease on every Bedrock model and rely on throttling to make costs reach exactly the budget.
    4. D.Enable Compute Optimizer for Bedrock to recommend cheaper models and automatically stop requests that exceed the limit.
    5. E.Create a CloudWatch alarm on EstimatedCharges in each linked account to remove Bedrock permissions the moment it breaches.
    Show answer & explanation

    Correct answers: A, B — Create an AWS Cost Anomaly Detection monitor on Amazon Bedrock with an alert subscription through SNS or email.; Create an AWS Budgets cost budget filtered to Amazon Bedrock with a budget action applying an IAM deny policy at 100%.

    • A. Cost Anomaly Detection learns spend patterns and notifies on unusual increases without a fixed threshold.
    • B. Budget actions can attach a restrictive IAM or SCP policy when the threshold is crossed, which enforces the limit beyond mere notification.
    • C. Throttling quotas control rate, not dollar spend, so they cannot produce an exact budget stop.
    • D. Compute Optimizer does not offer Bedrock model recommendations or automatic request blocking.
    • E. An EstimatedCharges alarm only notifies and does not change permissions by itself.

    Subdomain 4.2: Optimize and manage ML and AI infrastructure costs and performance

    35.A chatbot averages 3,000 output tokens per reply but customers need answers of about 150 words. Which TWO actions reduce token spend while keeping answer quality?(Select 2)

    1. A.Set `maxTokens` to a limit that fits the required answer length and instruct the model to respond concisely in the system prompt.
    2. B.Track InputTokenCount and OutputTokenCount in the AWS/Bedrock CloudWatch namespace to confirm the change lowers token consumption.
    3. C.Raise the temperature to 1.0 so the model produces shorter and more varied completions for each of the customer questions asked.
    4. D.Append the full chat history to every request so the model remembers context and answers each follow-up question faster than before.
    5. E.Switch to a model with a larger context window, since larger windows are known to reduce the output tokens produced per reply.
    Show answer & explanation

    Correct answers: A, B — Set `maxTokens` to a limit that fits the required answer length and instruct the model to respond concisely in the system prompt.; Track InputTokenCount and OutputTokenCount in the AWS/Bedrock CloudWatch namespace to confirm the change lowers token consumption.

    • A. Bounding output length and asking for concision directly reduce billable output tokens.
    • B. Token-count metrics provide the feedback loop to verify savings and spot regressions.
    • C. Temperature adjusts sampling randomness and has no reliable effect on reply length.
    • D. Sending full history on each call increases input tokens, so the cost rises.
    • E. A larger context window allows bigger inputs but does not shorten outputs, and may carry a higher price.

    Want the full experience?

    These are just samples. Practice the full AWS Certified Machine Learning Engineer - Associate (MLA-C02) question bank in quiz mode — free, no signup, with domain practice and exam simulation.