CertSafari

    Free AWS Certified Generative AI Developer - Professional (AIP-C01) Sample Questions

    35 free sample questions from our bank of 361+, covering every exam domain, with answers and detailed explanations. Updated September 2026.

    Domain 1: Foundation Model Integration, Data Management, and Compliance

    Subdomain 1.2: Select and configure FMs.

    1.A media company invokes Amazon Bedrock foundation models directly from its application backend using hardcoded model IDs. Leadership now wants the ability to swap between Bedrock model providers, or roll out a new model version to a subset of traffic, without redeploying application code. Which architecture best achieves this?

    1. A.Put an AWS Lambda function behind Amazon API Gateway to abstract the invocation, and source the active model ID from AWS AppConfig so it can change without a code deploy.
    2. B.Keep the current hardcoded model IDs in the application, and rely on manual code reviews to catch any place that needs to change when a model swap is requested.
    3. C.Move the model IDs into environment variables set at container build time, and redeploy the application container whenever a different model needs to be used.
    4. D.Store the model ID inside a configuration file bundled into the deployment package, and cut and ship a brand-new application release each time that file needs any change.
    Show answer & explanation

    Correct answer: A — Put an AWS Lambda function behind Amazon API Gateway to abstract the invocation, and source the active model ID from AWS AppConfig so it can change without a code deploy.

    • A. Fronting the invocation logic with Lambda and API Gateway, and externalizing the model choice to AppConfig, lets the routing and model ID change through a configuration deployment rather than an application code deploy.
    • B. Relying on manual review to catch every hardcoded reference does not remove the underlying coupling and leaves the process error-prone and still tied to a code change.
    • C. Environment variables set at build time still require rebuilding and redeploying the container image for every model change, which is the exact friction the team wants removed.
    • D. Bundling the model ID inside the deployment package still requires cutting and shipping a new release for every change, so it does not decouple configuration from code deploys.

    Subdomain 1.4: Design and implement vector store solutions.

    2.A conglomerate is building a single Amazon Bedrock knowledge base that must serve five business units, each with deeply nested product manuals containing parent sections and child sub-sections. Retrieval must return a focused child chunk for precise answers while still letting the model pull the surrounding parent section for context when the child alone is insufficient. Which approach best satisfies this requirement?

    1. A.Configure the knowledge base with hierarchical chunking so parent and child chunk relationships are preserved, letting retrieval return a child chunk and expand to its parent when more context is needed.
    2. B.Configure the knowledge base with fixed-size chunking at a large chunk size so each chunk already contains an entire section, removing any need to track parent-child relationships during retrieval flows.
    3. C.Split the documentation into five separate knowledge bases, one per business unit, and rely on the FM to merge results from whichever knowledge base returns the highest similarity score.
    4. D.Store the entire manual as a single embedding per document and let the FM use its context window to locate the relevant subsection after retrieval returns the whole document text.
    Show answer & explanation

    Correct answer: A — Configure the knowledge base with hierarchical chunking so parent and child chunk relationships are preserved, letting retrieval return a child chunk and expand to its parent when more context is needed.

    • A. Hierarchical chunking is designed exactly for this case: it indexes smaller child chunks for precise semantic matching while retaining a link to the larger parent chunk, so retrieval can return the focused answer and still surface surrounding context on demand.
    • B. Large fixed-size chunks trade away retrieval precision because a single chunk mixes multiple topics; the model gets more text per hit but loses the ability to zero in on the specific passage that actually answers the question.
    • C. Splitting by business unit addresses organizational separation but does nothing for the parent-child depth problem within a single unit's manuals, and juggling five knowledge bases adds retrieval-orchestration complexity the scenario does not call for.
    • D. Embedding whole documents produces very coarse similarity matches and forces the FM to consume large amounts of irrelevant text from the context window, which is inefficient and does not target the sub-section the user actually needs.

    Subdomain 1.4: Design and implement vector store solutions.

    3.A team evaluating vector store options for a Bedrock knowledge base needs metadata filtering, such as filtering by product category, to work correctly the moment the knowledge base goes live, without any additional manual configuration inside the vector store's own index settings. Which statement about metadata filtering behavior across vector store choices is accurate?

    1. A.MongoDB Atlas requires manually configuring filters directly in its own vector index settings before metadata filtering will work, since it is not enabled automatically by default.
    2. B.Amazon OpenSearch Serverless requires no metadata field configuration whatsoever, since every field in a vector index is automatically filterable no matter how that field was originally mapped.
    3. C.Aurora PostgreSQL with pgvector filters on metadata automatically without any indexing, because sequential scans on a jsonb column are always fast enough at any table size.
    4. D.Amazon S3 Vectors treats every attached metadata key as filterable by default, with no option anywhere to mark any key as non-filterable at index creation time.
    Show answer & explanation

    Correct answer: A — MongoDB Atlas requires manually configuring filters directly in its own vector index settings before metadata filtering will work, since it is not enabled automatically by default.

    • A. MongoDB Atlas is documented as requiring the customer to manually configure filters within their own Atlas vector index settings, since metadata filtering does not work out of the box the way it does for some other supported stores.
    • B. OpenSearch requires custom metadata fields intended for filtering to be defined with a keyword type or a text field with a keyword subfield; fields mapped without that structure fail filtering queries with a rewrite error rather than working automatically.
    • C. Aurora PostgreSQL benefits significantly from a GIN index on the jsonb metadata column for filter performance; relying on sequential scans without an index does not scale and is not how the documented setup recommends configuring filtering.
    • D. Amazon S3 Vectors lets metadata be configured as either filterable or non-filterable at index creation time, so treating every key as filterable by default with no option otherwise misstates how S3 Vectors metadata is actually configured.

    Subdomain 1.1: Analyze requirements and design GenAI solutions.

    4.A team has just been asked to build a GenAI feature and is meeting with business stakeholders for the first time to define what success looks like, before any model has been selected or any prompt has been written. According to the Well-Architected Generative AI Lens lifecycle, which phase is the team currently in?

    1. A.Scoping the problem to determine whether generative AI is an appropriate way to address the stated business need at all.
    2. B.Selecting a specific foundation model that sufficiently addresses the technical requirements already agreed upon for the task.
    3. C.Customizing a chosen model's behavior using prompts, retrieved data, or updated weights to improve its output quality.
    4. D.Iterating on a released generative AI capability based on production usage data and feedback collected after launch.
    Show answer & explanation

    Correct answer: A — Scoping the problem to determine whether generative AI is an appropriate way to address the stated business need at all.

    • A. Defining what success looks like with stakeholders before any model has been chosen or any prompt written is the scoping phase, where the team determines whether generative AI is even the right approach to the business need.
    • B. Model selection happens after scoping has established the requirements a model needs to satisfy, but this team has not yet reached that point since no requirements have been agreed upon yet.
    • C. Customization work like prompting or fine-tuning happens after a model has already been selected, which has not occurred yet in a meeting held before any model choice or prompt has been discussed.
    • D. Iterating on a released capability requires the feature to already be in production collecting usage data, which is far downstream of an initial stakeholder meeting that precedes model selection entirely.

    Subdomain 1.3: Implement data validation and processing pipelines for FM consumption.

    5.A retailer stores product photos in S3 alongside free-text seller descriptions and wants a single pipeline step that reads both a product image and its accompanying description together to generate a standardized, SEO-optimized product title and category tag before the item is indexed for search. Which approach processes both modalities together in a single inference call?

    1. A.Send the image bytes and the description text together in one request to a multimodal foundation model on Amazon Bedrock that accepts combined image and text content in a single prompt.
    2. B.Run the image through an Amazon Rekognition label detection API to produce tags, then send only those tags along with the description text to a text-only foundation model on Amazon Bedrock.
    3. C.Convert the product image to a base64 string and store it as metadata in Amazon DynamoDB, then send only the seller description text to a foundation model for title generation.
    4. D.Use Amazon Textract to extract any text visible inside the product photo, then merge that extracted text with the seller description before sending it to a text-only foundation model.
    Show answer & explanation

    Correct answer: A — Send the image bytes and the description text together in one request to a multimodal foundation model on Amazon Bedrock that accepts combined image and text content in a single prompt.

    • A. A multimodal Bedrock model can accept image and text content blocks in the same request, letting it reason jointly over the visual and textual signal in a single inference call.
    • B. Rekognition labels only capture generic object tags, discarding the visual detail a multimodal model would otherwise see directly, and the text-only model never receives the image itself.
    • C. Storing the image as base64 metadata does not feed it into the model at all, so the title-generation call only ever sees the seller description text.
    • D. Textract extracts printed or handwritten text from an image but does not capture non-text visual attributes like color or style, and it still requires a separate text-only inference call.

    Subdomain 1.3: Implement data validation and processing pipelines for FM consumption.

    6.A team is building a multi-turn customer support chatbot on top of Amazon Bedrock and wants to preserve conversational context across turns while remaining compatible with any Bedrock model that supports the Converse API, without hand-building a model-specific prompt template for each provider. How should each turn be formatted?

    1. A.Append each turn as a message object with a `role` of `user` or `assistant` and a `content` array holding the turn's text or media blocks, passed to the Converse API.
    2. B.Concatenate the entire conversation history into a single long text string labeled with "User:" and "Assistant:" prefixes, then pass that string as the sole prompt in an InvokeModel request.
    3. C.Store each turn as a separate item in Amazon DynamoDB and send only the item's primary key to the foundation model, expecting it to fetch and join prior turns itself.
    4. D.Send only the most recent user message to the model on every turn, relying on the model's general training data to somehow infer everything said earlier in this specific conversation.
    Show answer & explanation

    Correct answer: A — Append each turn as a message object with a `role` of `user` or `assistant` and a `content` array holding the turn's text or media blocks, passed to the Converse API.

    • A. The Converse API standardizes conversation state as a list of messages with alternating user/assistant roles and content blocks, working consistently across Bedrock models without a model-specific template.
    • B. Manually concatenating turns into a labeled text blob abandons the Converse API's structured message format and reintroduces the model-specific prompt engineering the Converse API is meant to avoid.
    • C. Foundation models cannot query external databases mid-inference on their own; the application must retrieve and include the actual turn content in the request.
    • D. Foundation models have no memory of a specific ongoing conversation from training data; omitting prior turns means the model loses all context from earlier in the chat.

    Subdomain 1.6: Implement prompt engineering strategies and governance for FM interactions.

    7.A prompt engineering team is redesigning several existing prompts and Bedrock Flows for a customer-facing GenAI application. Which of the following changes each correctly apply an Amazon Bedrock Guardrail to protect model input or output within the described component? (Select TWO.)(Select 2)

    1. A.Adding a `guardrailConfiguration` field with the guardrail identifier and version to a flow's prompt node so both the input prompt and the generated response are evaluated against the guardrail.
    2. B.Calling the `ApplyGuardrail` API directly against user input text before it is ever sent to a foundation model, to evaluate the input independently of any model invocation.
    3. C.Attaching a guardrail identifier to an S3 storage node's configuration, expecting the guardrail to scan files after they are written to the bucket for compliance purposes.
    4. D.Configuring a guardrail inside a Lambda function node's `lambdaFunction` block, expecting Bedrock to automatically apply content filtering to the function's return value.
    5. E.Setting a guardrail identifier on a Lex node, expecting the guardrail to filter the predicted intent output before it reaches any downstream node in the flow.
    Show answer & explanation

    Correct answers: A, B — Adding a `guardrailConfiguration` field with the guardrail identifier and version to a flow's prompt node so both the input prompt and the generated response are evaluated against the guardrail.; Calling the `ApplyGuardrail` API directly against user input text before it is ever sent to a foundation model, to evaluate the input independently of any model invocation.

    • A. A prompt node's guardrail configuration is a documented, supported way to have Bedrock evaluate both the prompt input and the model's completion against the specified guardrail identifier and version as part of flow execution.
    • B. The `ApplyGuardrail` API is specifically designed to evaluate arbitrary input or output text against a guardrail independently of invoking a foundation model, making this a valid way to screen user input before any model call.
    • C. An S3 storage node's configuration only specifies the target bucket for storing flow data and has no guardrail parameter or built-in mechanism to scan stored files after they are written.
    • D. A Lambda function node's configuration only specifies the function ARN to invoke and has no guardrail field, so content filtering on a Lambda's return value is not something Bedrock applies automatically through that node.
    • E. A Lex node's configuration specifies the bot alias and locale for intent recognition and has no guardrail parameter, so there is no built-in mechanism for a guardrail to filter a predicted intent through that node type.

    Subdomain 1.6: Implement prompt engineering strategies and governance for FM interactions.

    8.A pharmaceutical company's internal research assistant retrieves excerpts from clinical trial reports and generates a summary answer for a researcher's question. The compliance team wants the assistant to automatically flag any generated answer that introduces a claim not actually present in the retrieved trial excerpts, so a reviewer can catch potential fabrication before the researcher relies on it. Which two design choices together best satisfy this requirement? (Select TWO.)(Select 2)

    1. A.Apply a Bedrock Guardrail with contextual grounding checks configured on the prompt or knowledge base node so ungrounded claims in the generated answer are automatically flagged.
    2. B.Add an explicit instruction in the prompt directing the model to answer strictly from the supplied trial excerpts and to state clearly when a claim cannot be supported by them.
    3. C.Increase the maximum output token limit so the model has more room to generate a longer, more thorough summary answer from the excerpts retrieved for the researcher's question.
    4. D.Switch the retrieval step to return excerpts from a broader, unfiltered set of external medical literature instead of the specific trial reports being reviewed.
    5. E.Disable the knowledge base node's model-generation step so only raw retrieved excerpts are returned to the researcher, removing summarization from the workflow entirely.
    Show answer & explanation

    Correct answers: A, B — Apply a Bedrock Guardrail with contextual grounding checks configured on the prompt or knowledge base node so ungrounded claims in the generated answer are automatically flagged.; Add an explicit instruction in the prompt directing the model to answer strictly from the supplied trial excerpts and to state clearly when a claim cannot be supported by them.

    • A. Contextual grounding checks are the Guardrails feature specifically designed to detect when a generated response introduces information not supported by the retrieved source, which is exactly the automated flagging mechanism compliance is asking for.
    • B. An explicit instruction constraining the model to the supplied excerpts and requiring it to acknowledge unsupported claims directly reduces the model's tendency to fabricate content and works alongside the automated grounding check as a complementary safeguard.
    • C. A larger output token limit only changes how much text the model is allowed to generate and has no mechanism for detecting or flagging whether a claim in that text is actually grounded in the retrieved excerpts.
    • D. Broadening retrieval to unfiltered external literature moves further away from the specific trial reports under review and increases the risk of introducing ungrounded claims rather than helping the assistant stay grounded in the intended source.
    • E. Removing summarization entirely would return only raw excerpts without a generated answer, which eliminates the fabrication risk by eliminating the feature itself rather than providing the requested flagging mechanism for generated summaries.

    Subdomain 1.5: Design retrieval mechanisms for FM augmentation.

    9.A knowledge team is configuring semantic chunking for a data source of long-form research articles so that chunk boundaries fall at genuine topic shifts rather than arbitrary token counts. During testing, the resulting chunks are inconsistent: some chunks merge unrelated paragraphs together, while others split a single argument across multiple chunks. Which pair of adjustments should the team make to the semantic chunking configuration to produce more coherent, topic-aligned chunks?

    1. A.Increase the buffer size so more surrounding sentences are combined before embedding, and raise the breakpoint percentile threshold so only clearly dissimilar sentences trigger a new chunk.
    2. B.Decrease the buffer size all the way to embed each sentence in isolation, and also lower the breakpoint percentile threshold so nearly every sentence pair starts a brand-new chunk boundary.
    3. C.Switch the maximum tokens parameter to a much smaller value while leaving the buffer size and the breakpoint percentile threshold at their current settings unchanged.
    4. D.Disable semantic chunking's dependency on a foundation model entirely and instead apply a fixed overlap percentage to the existing chunk boundaries from ingestion.
    5. E.Set the breakpoint percentile threshold to its lowest possible value while also reducing the maximum tokens parameter so every chunk stays extremely short and fragmented.
    Show answer & explanation

    Correct answer: A — Increase the buffer size so more surrounding sentences are combined before embedding, and raise the breakpoint percentile threshold so only clearly dissimilar sentences trigger a new chunk.

    • A. A larger buffer size gives each sentence more surrounding context before it is embedded, reducing false splits from momentary wording shifts, while a higher breakpoint percentile threshold requires sentences to be more clearly dissimilar before a boundary is drawn, which together curb both merged and over-split chunks.
    • B. Shrinking the buffer size removes the surrounding context that helps distinguish a genuine topic shift from ordinary sentence variation, and a lower breakpoint threshold makes boundaries trigger too easily, worsening the over-splitting problem rather than fixing it.
    • C. The maximum tokens parameter only caps chunk size after boundaries are already chosen from sentence dissimilarity, so changing it alone does not address why unrelated paragraphs are merging or a single argument is being split.
    • D. Semantic chunking depends on a foundation model to compute sentence dissimilarity by design, and replacing that mechanism with fixed-percentage overlap reverts to a token-based approach that cannot align boundaries with topic shifts.
    • E. Pushing the breakpoint threshold to its lowest setting combined with a very small token cap produces excessively fragmented chunks, which trades the merging problem for aggressive over-splitting instead of balancing the two.

    Subdomain 1.5: Design retrieval mechanisms for FM augmentation.

    10.An architecture reviewer is evaluating a proposal that describes the Model Context Protocol (MCP) as a way to store and index vector embeddings for retrieval, replacing the knowledge base's existing vector store. Which statement correctly describes what MCP actually is?

    1. A.MCP is a standardized protocol for how agents discover and invoke tools, such as a vector search tool, not a vector database or embedding storage mechanism itself.
    2. B.MCP is a vector indexing engine that fully replaces OpenSearch, Aurora pgvector, and other vector stores as the actual place where embeddings get physically stored.
    3. C.MCP is an embeddings model family, similar to Titan Text Embeddings, that converts text chunks into vector representations for semantic search.
    4. D.MCP is a chunking strategy, similar to fixed-size or hierarchical chunking, that determines how documents are split before embeddings are generated.
    Show answer & explanation

    Correct answer: A — MCP is a standardized protocol for how agents discover and invoke tools, such as a vector search tool, not a vector database or embedding storage mechanism itself.

    • A. MCP is documented as a standardized way for agents to discover and invoke tools, such as a vector search tool exposed behind a gateway, and it does not itself store or index vectors, so it sits alongside a vector store rather than replacing one.
    • B. MCP does not physically store or index vector embeddings; that responsibility remains with a vector store such as OpenSearch or Aurora pgvector, so describing MCP as a vector indexing engine misstates its role.
    • C. MCP is not an embeddings model and does not convert text into vectors; that conversion is performed by a separate embeddings model such as Titan Text Embeddings, which MCP has no role in replacing.
    • D. MCP has nothing to do with how documents are split into chunks during ingestion; chunking strategies operate entirely upstream of, and independently from, the tool-invocation protocol MCP defines.

    Domain 2: Implementation and Integration

    Subdomain 2.3: Design and implement enterprise integration architectures.

    11.For an enterprise integrating thousands of workforce users with Bedrock-backed applications, which statement correctly compares identity federation against creating individual long-lived IAM users?

    1. A.Identity federation issues short-term credentials tied to an existing corporate identity provider, avoiding the burden of provisioning, rotating, and deprovisioning an IAM user per employee.
    2. B.Individual IAM users are the recommended approach at workforce scale because each user's access key can be rotated automatically by AWS without any administrator involvement whatsoever.
    3. C.Identity federation and individual IAM users provide identical security properties, so the choice between them has no meaningful effect on credential lifecycle or exposure at enterprise scale.
    4. D.Individual IAM users are required for Bedrock access specifically, since Bedrock does not accept temporary security credentials obtained through an identity provider or assumed role.
    Show answer & explanation

    Correct answer: A — Identity federation issues short-term credentials tied to an existing corporate identity provider, avoiding the burden of provisioning, rotating, and deprovisioning an IAM user per employee.

    • A. Federated sign-in through an existing identity provider issues temporary credentials scoped to a session, which removes the need to manually create, rotate, and delete a long-lived IAM user for every one of thousands of employees.
    • B. AWS access keys for IAM users are not rotated automatically by AWS; rotation is an operational task the account owner must schedule and perform, so this claim misstates how IAM user credentials are managed.
    • C. Long-lived IAM user credentials and short-lived federated session credentials differ meaningfully in exposure window and rotation burden, so the two approaches are not equivalent from a security standpoint.
    • D. Bedrock accepts standard AWS IAM authentication, including temporary credentials obtained via federation or assumed roles, so there is no restriction forcing the use of individual long-lived IAM users.

    Subdomain 2.3: Design and implement enterprise integration architectures.

    12.In a centralized GenAI gateway architecture that abstracts foundation model access for many internal applications, which capability is most directly an example of an observability and control mechanism the gateway provides?

    1. A.The gateway records per-request metrics such as latency, token usage, and error rate for every calling application, and enforces per-application rate limits from that recorded usage.
    2. B.The gateway automatically rewrites every application's prompt text to a shorter version before forwarding it, regardless of whether the application requested any prompt modification.
    3. C.The gateway stores a permanent, unencrypted copy of every request and response payload in a publicly readable location so any employee can browse call history without authentication.
    4. D.The gateway removes the need for any application team to specify which foundation model it wants to call, since the gateway always selects a model automatically without any application input.
    Show answer & explanation

    Correct answer: A — The gateway records per-request metrics such as latency, token usage, and error rate for every calling application, and enforces per-application rate limits from that recorded usage.

    • A. Recording per-request metrics like latency, token usage, and error rate, and using that data to enforce rate limits, is a direct observability and control mechanism, giving the platform team visibility and enforcement power across every consuming application.
    • B. Automatically rewriting prompt text without being asked is a content-manipulation behavior, not an observability or control mechanism, and it risks changing application behavior in ways the calling team never requested.
    • C. Storing payloads in a publicly readable, unauthenticated location is a data-exposure risk rather than a control mechanism, and it contradicts the access-control goals a gateway is meant to help enforce.
    • D. Removing an application's ability to specify a model takes away a legitimate integration choice rather than adding observability or control, and most gateway designs still let the caller request a specific approved model.

    Subdomain 2.2: Implement model deployment strategies.

    13.A machine learning team fine-tuned a foundation model in Amazon Bedrock using their own labeled dataset to improve its performance on domain-specific terminology. Before they can run inference against this customized model, which requirement must they satisfy?

    1. A.They must purchase Amazon Bedrock Provisioned Throughput for the customized model, because Bedrock does not support on-demand invocation for any customized model.
    2. B.They must deploy the customized model to a SageMaker AI real-time endpoint instead, because Bedrock only serves the original base-model weights and cannot host any customized model artifacts at all.
    3. C.They must request a service quota increase for on-demand throughput, because customized models consume a noticeably larger per-request token allowance than the equivalent base foundation model.
    4. D.They must convert the customized model into a Bedrock Agent action group, because agent invocation is the only supported runtime path for models that carry additional fine-tuned weights.
    Show answer & explanation

    Correct answer: A — They must purchase Amazon Bedrock Provisioned Throughput for the customized model, because Bedrock does not support on-demand invocation for any customized model.

    • A. Bedrock requires a Provisioned Throughput purchase before a customized model can be invoked at all — on-demand, pay-per-token invocation is only available for unmodified base foundation models.
    • B. Bedrock does host customized models directly; moving to a separate SageMaker AI endpoint is unnecessary and ignores that Bedrock's own customization and Provisioned Throughput features exist for this purpose.
    • C. The blocking requirement is Provisioned Throughput, not a token-allowance quota — a quota increase for on-demand throughput would not unlock invocation of a customized model since on-demand access simply is not offered for it.
    • D. Agent action groups are for connecting a model to external tools and APIs; they are unrelated to whether a customized model can be invoked, and they do not substitute for the Provisioned Throughput requirement.

    Subdomain 2.2: Implement model deployment strategies.

    14.A platform serves three distinct foundation-model workloads: a live chat assistant that needs sub-second responses at unpredictable volume, a nightly batch job that scores hundreds of thousands of stored documents with no user waiting on the result, and an image-captioning feature that receives infrequent requests but must not incur cost while idle. Which pairing of workload to inference option correctly matches all three?

    1. A.Serve the chat assistant with a low-latency real-time endpoint kept warm for interactive traffic, score the batch job through batch or asynchronous inference, and serve image captioning through on-demand invocation billed only per request.
    2. B.Serve the chat assistant through nightly batch inference so responses accumulate overnight, score the batch job with a real-time endpoint kept warm around the clock, and serve captioning through a full-month Provisioned Throughput commitment.
    3. C.Serve the chat assistant through on-demand invocation with no warm capacity reserved in advance, score the batch job on a real-time endpoint kept running and sized for its peak scoring rate, and delay captioning until the next batch run.
    4. D.Serve all three workloads through one Provisioned Throughput commitment sized for the chat assistant's peak traffic, since a single reserved-capacity purchase can absorb scoring and captioning without further configuration.
    Show answer & explanation

    Correct answer: A — Serve the chat assistant with a low-latency real-time endpoint kept warm for interactive traffic, score the batch job through batch or asynchronous inference, and serve image captioning through on-demand invocation billed only per request.

    • A. A warm real-time endpoint gives the chat assistant the low, consistent latency it needs; batch or asynchronous inference fits document scoring that no user is waiting on; and per-request on-demand invocation avoids any charge while the captioning feature sits idle.
    • B. Routing the chat assistant through overnight batch inference would make interactive users wait until the next batch run, and reserving a real-time endpoint continuously for document scoring pays for always-on capacity a background job does not need.
    • C. Unpredictable on-demand invocation with no warm capacity risks added latency for a sub-second chat requirement, and sizing a continuously running real-time endpoint for the batch job's peak scoring rate pays for idle capacity between runs.
    • D. A single reserved-capacity commitment sized for interactive chat traffic bills continuously regardless of the batch job's and captioning feature's very different, mostly idle usage patterns, which wastes the discount Provisioned Throughput offers steady workloads.

    Subdomain 2.1: Implement agentic AI solutions and tool integrations.

    15.A developer is defining a new tool for a Bedrock agent that looks up a customer's order status. The tool must accept an order ID and an optional region code, validate that the order ID matches an expected format before any downstream call is made, and return a response the agent can parse reliably regardless of which foundation model is orchestrating the agent. Which practice best achieves reliable tool operation here?

    1. A.Define the tool with a standardized function schema specifying required and optional parameter types, and implement input validation and a consistent, structured response format in the backing Lambda function.
    2. B.Describe the order-lookup tool only in free-form prose within the agent's instructions, without a structured parameter schema, and let the model infer the order ID's expected format from context each time.
    3. C.Skip parameter validation in the Lambda function entirely and let the downstream order-status service reject malformed order IDs, then surface that service's raw error message back to the agent unchanged and unparsed.
    4. D.Return whatever raw response shape the downstream order-status service happens to produce, so its structure can change between calls, and require each orchestrating foundation model to adapt its own parsing.
    Show answer & explanation

    Correct answer: A — Define the tool with a standardized function schema specifying required and optional parameter types, and implement input validation and a consistent, structured response format in the backing Lambda function.

    • A. A standardized function schema with typed required and optional parameters, paired with input validation and a consistent structured response, gives any orchestrating model a reliable contract for calling and interpreting the tool.
    • B. Relying on free-form prose without a structured schema leaves the expected order ID format and optional region code ambiguous, which increases the chance of malformed calls compared to an explicit schema.
    • C. Skipping validation in the Lambda function pushes error handling onto the downstream service and exposes its raw error format to the agent, which is a less controlled and less predictable failure path than validating first.
    • D. Passing through an unstable, changing response shape forces every orchestrating model to adapt its parsing individually, which undermines the consistent, predictable tool behavior reliable operation requires.

    Subdomain 2.1: Implement agentic AI solutions and tool integrations.

    16.A platform architect is comparing Amazon Bedrock Agents (Classic) with Amazon Bedrock AgentCore for a new agent build that requires fine-grained observability, identity management, and a gateway for tool access. Which statement accurately reflects the current relationship between these two offerings?

    1. A.Amazon Bedrock Agents Classic is closed to new customers, and Amazon Bedrock AgentCore is the offering AWS points new builds toward for similar capabilities plus added observability and gateway components.
    2. B.Amazon Bedrock Agents Classic remains the only agent-building service AWS offers, and Amazon Bedrock AgentCore is an unrelated, deprecated preview feature that new agent builds should never consider at all.
    3. C.Amazon Bedrock AgentCore is strictly a rebranding of Amazon Bedrock Agents Classic with an identical feature set, so the choice between the two has no bearing on which capabilities a new agent build can access.
    4. D.Amazon Bedrock Agents Classic and Amazon Bedrock AgentCore are both open to new customers today with fully interchangeable feature sets, so choosing between them is purely a matter of team naming preference.
    Show answer & explanation

    Correct answer: A — Amazon Bedrock Agents Classic is closed to new customers, and Amazon Bedrock AgentCore is the offering AWS points new builds toward for similar capabilities plus added observability and gateway components.

    • A. Bedrock Agents Classic is closed to new customers, and AWS documentation directs new builds toward AgentCore for similar orchestration capabilities along with added observability, identity, and gateway functionality.
    • B. This reverses the actual status: Agents Classic is the one closed to new customers, and AgentCore is the actively developed offering AWS points new builds toward, not an unrelated deprecated preview.
    • C. AgentCore adds observability, identity, and gateway components beyond what Classic provided, so describing it as an identical rebrand misrepresents the expanded feature set new builds gain by choosing it.
    • D. Agents Classic is no longer open to new customers, so the two offerings are not both currently available to a new build in the same way, and the choice is not merely a naming preference.

    Subdomain 2.5: Implement application integration patterns and development tools.

    17.A platform team is exposing an internal foundation model service to five other engineering teams across the company. Each consuming team builds its own client in a different language, and the platform team wants a contract-first process where the request and response shapes, required parameters, and error codes are agreed upon and validated before any client or server code is written. Which approach best supports this API-first development style?

    1. A.Author an OpenAPI specification that defines the request and response schemas, parameters, and error codes up front, and generate client SDKs and server stubs from that shared contract.
    2. B.Have each consuming team read the platform team's Lambda function source code directly and infer the expected request and response shapes from the implementation as it evolves.
    3. C.Publish a written prose document describing the API in a wiki page, and ask each consuming team to write their own client based on their individual interpretation of the wording.
    4. D.Skip formal API design and let the first team that integrates define the request shape by trial and error, then have later teams copy whatever worked for that first client.
    Show answer & explanation

    Correct answer: A — Author an OpenAPI specification that defines the request and response schemas, parameters, and error codes up front, and generate client SDKs and server stubs from that shared contract.

    • A. An OpenAPI specification is a machine-readable contract that defines schemas, parameters, and errors before implementation begins, which is exactly what API-first development uses to keep multiple independent clients aligned and to generate consistent SDKs and stubs.
    • B. Reading implementation code as the source of truth couples every consuming team to internal details that can change at any time, and it provides no agreed contract to validate requests and responses against.
    • C. A prose wiki description is not machine-validated, so it is easy for each team's interpretation to drift from the others, producing inconsistent clients that a formal schema would have caught.
    • D. Letting the first integration define the contract by trial and error skips the up-front agreement step entirely and risks baking an accidental, undocumented shape into every later client that copies it.

    Subdomain 2.5: Implement application integration patterns and development tools.

    18.A team is designing a multi-agent system where a coordinating agent delegates subtasks to specialized worker agents, and they want to use an AWS-native, open-source framework specifically built for orchestrating this kind of agent collaboration rather than hand-rolling their own agent-to-agent messaging. Which pairing of requirement to AWS-native capability is correct?

    1. A.Multi-agent coordination and delegation maps to AWS Agent Squad, an AWS-native framework purpose-built for routing tasks between specialized agents.
    2. B.Multi-agent coordination and delegation maps to Amazon CloudWatch Synthetics, a service for running scripted canaries against application endpoints on a schedule.
    3. C.Multi-agent coordination and delegation maps to AWS Config, a service for recording and evaluating the configuration state of AWS resources over time.
    4. D.Multi-agent coordination and delegation maps to Amazon EventBridge Scheduler, a service for invoking targets on a fixed or recurring time-based schedule.
    Show answer & explanation

    Correct answer: A — Multi-agent coordination and delegation maps to AWS Agent Squad, an AWS-native framework purpose-built for routing tasks between specialized agents.

    • A. AWS Agent Squad is an AWS-native orchestration framework built specifically to route tasks and manage collaboration between multiple specialized agents, which matches the coordinator/worker delegation pattern described.
    • B. CloudWatch Synthetics runs scripted health-check canaries against endpoints on a schedule; it has no role in coordinating task delegation between cooperating AI agents.
    • C. AWS Config tracks and evaluates the configuration state of AWS resources for compliance purposes, which is unrelated to orchestrating collaboration between AI agents.
    • D. EventBridge Scheduler invokes a target on a time-based schedule; it does not provide the task-routing or delegation logic that a multi-agent coordination framework needs.

    Subdomain 2.4: Implement FM API integrations.

    19.A team wants to deliver a long Bedrock-generated report to a client incrementally over a plain HTTP connection through API Gateway's REST API, using chunked transfer encoding rather than adopting a WebSocket API, because their client library already handles chunked HTTP well. Which two statements correctly describe constraints they must plan around? (Select 2)(Select 2)

    1. A.API Gateway's HTTP APIs (the newer, lighter-weight API type) do not support chunked transfer encoding, so relying on chunked delivery through that API type will fail even if the backend produces a chunked stream.
    2. B.A Lambda proxy integration behind API Gateway must still fully assemble its complete response object before returning it to API Gateway, which limits how much true incremental delivery reaches the client through this path.
    3. C.Chunked transfer encoding automatically converts any REST API integration into a full duplex WebSocket connection, so the team gets bidirectional streaming for free without further design changes.
    4. D.Every foundation model hosted in Bedrock produces output using the chunked transfer encoding format natively, so no additional response formatting is required on the backend at all.
    5. E.API Gateway strips the `Transfer-Encoding` header on affected API types regardless of what the backend integration sends, so the client cannot rely on that header surviving the round trip.
    6. F.Chunked transfer encoding is only usable when the client and backend both run inside the same VPC, since API Gateway cannot forward chunked responses across a public internet connection.
    Show answer & explanation

    Correct answers: A, E — API Gateway's HTTP APIs (the newer, lighter-weight API type) do not support chunked transfer encoding, so relying on chunked delivery through that API type will fail even if the backend produces a chunked stream.; API Gateway strips the `Transfer-Encoding` header on affected API types regardless of what the backend integration sends, so the client cannot rely on that header surviving the round trip.

    • A. API Gateway's newer HTTP API type does not support chunked transfer encoding, so a design that depends on chunked delivery must avoid that API type and use a different integration path instead.
    • B. A standard Lambda proxy integration returns one complete response object to API Gateway, so the function itself must finish assembling output before anything is returned, constraining how incremental the delivery can actually be through that integration style.
    • C. Chunked transfer encoding is a one-way HTTP response mechanism and has no relationship to WebSocket connections; it does not convert a REST integration into a bidirectional channel.
    • D. Foundation models return their output through the Bedrock runtime's own response format, such as the streaming event structure; nothing about that format is inherently chunked-transfer-encoded HTTP, so backend formatting work is still required.
    • E. API Gateway is known to drop or strip the `Transfer-Encoding` header on the API types where chunked encoding is unsupported, so a design cannot depend on that header reaching the client intact.
    • F. Chunked transfer encoding is a standard HTTP mechanism that works over ordinary internet connections; it has no dependency on client and backend sharing a VPC.

    Domain 3: AI Safety, Security, and Governance

    Subdomain 3.1: Implement input and output safety controls.

    20.A legal research assistant built on a Bedrock Knowledge Base occasionally returns a fluent answer that states a case outcome not actually present in the retrieved source documents. The compliance team wants an automated way to detect when a generated answer deviates from or adds information beyond what the retrieved passages actually support, and to block that specific answer, without waiting for a human reviewer to read every response. Which capability should the team configure?

    1. A.A contextual grounding check in a Bedrock guardrail, configured with a threshold that compares the generated answer against the retrieved source passages and blocks responses that are not sufficiently grounded.
    2. B.A higher top-k retrieval setting configured on the knowledge base itself, so that many more source passages get retrieved and passed into the prompt context for every single user question that gets submitted overall.
    3. C.A word filter in the guardrail listing specific legal terms that should never appear in a generated response, regardless of what the retrieved source content actually contains.
    4. D.A scheduled Amazon CloudWatch alarm on the knowledge base's document ingestion job failure rate, meant to catch problems with the underlying source corpus feeding the assistant.
    Show answer & explanation

    Correct answer: A — A contextual grounding check in a Bedrock guardrail, configured with a threshold that compares the generated answer against the retrieved source passages and blocks responses that are not sufficiently grounded.

    • A. A contextual grounding check specifically scores whether a generated answer is factually supported by and relevant to the retrieved source passages, and can block responses that fall below a configured grounding threshold, which directly targets ungrounded hallucinated claims.
    • B. Retrieving more passages can improve the material available to the model but does not itself check whether the eventual generated answer stayed faithful to that material, so ungrounded claims could still pass through unchecked.
    • C. A word filter blocks or allows specific literal terms and phrases; it cannot evaluate whether a factual claim in an answer is actually supported by the retrieved documents, so it does not address hallucinated content that uses none of the listed words.
    • D. An ingestion failure alarm detects when documents fail to load into the knowledge base, which is a data-pipeline health signal and does not evaluate whether an individual generated answer is grounded in retrieved content.

    Subdomain 3.1: Implement input and output safety controls.

    21.A SaaS company's public-facing coding assistant, built on a Bedrock model with a Standard-tier guardrail, is meant to refuse requests to generate malicious scripts. An attacker submits a request asking the model to write a Python function whose docstring and variable names describe steps for exfiltrating credentials, hoping that framing the harmful instructions as code comments rather than plain prose will slip past the content filter. Which guardrail behavior most directly addresses this specific evasion attempt?

    1. A.With the Standard tier, content filter and denied topics detection extends into code elements such as comments, variable names, and string literals, so intent hidden in identifiers is still evaluated.
    2. B.The guardrail's sensitive information policy scans code output for entities such as email addresses and phone numbers, which happens to also catch a request describing credential exfiltration steps in detail.
    3. C.A word filter listing the literal phrase 'exfiltrate credentials' blocks the request outright, since any request containing that exact phrase anywhere in the code gets denied regardless of context.
    4. D.The Classic tier's content filter already treats all text within a request identically regardless of whether it appears in code comments or plain prose, so tier choice is not relevant here.
    Show answer & explanation

    Correct answer: A — With the Standard tier, content filter and denied topics detection extends into code elements such as comments, variable names, and string literals, so intent hidden in identifiers is still evaluated.

    • A. The Standard safeguard tier explicitly extends content filter and denied topics detection into code elements including comments, variable names, and string literals, which is the capability that specifically prevents an attacker from hiding harmful intent inside code identifiers to evade detection.
    • B. The sensitive information policy is built to detect structured PII entities like emails and phone numbers, not to recognize harmful intent expressed through descriptive docstrings or variable naming, so it would not reliably catch this framing.
    • C. An exact-match word filter only catches the literal phrase it lists, so an attacker who rephrases the same intent with different wording in the docstring or identifiers would evade it entirely, unlike a category-based detection extended into code.
    • D. The scenario specifically distinguishes tier behavior, and the Classic tier does not extend content and topic detection into code elements the way the Standard tier does, so treating them as equivalent misses the actual difference relevant here.

    Subdomain 3.4: Implement responsible AI principles.

    22.A team fine-tuned a foundation model for internal use and, before releasing it to production, wants a single governed artifact that documents the model's intended use cases, known limitations, and evaluation results so that risk and compliance stakeholders can review it without digging through training notebooks. Which capability is designed for this purpose?

    1. A.An Amazon SageMaker Model Card documenting intended use, known limitations, and evaluation results in one record.
    2. B.An Amazon Bedrock guardrail configured with content filters covering the categories most relevant to the model's use case.
    3. C.A CloudWatch dashboard graphing token usage and latency metrics collected from the model's production invocation logs.
    4. D.An IAM policy restricting which roles are permitted to invoke the fine-tuned model through the Bedrock runtime endpoint.
    Show answer & explanation

    Correct answer: A — An Amazon SageMaker Model Card documenting intended use, known limitations, and evaluation results in one record.

    • A. A model card is purpose-built to centralize intended use, risk rating, limitations, and evaluation results in one governed record, which is exactly the single documentation artifact risk and compliance stakeholders can review.
    • B. A guardrail enforces safety and policy rules at inference time; it does not document intended use, known limitations, or evaluation results as a governance artifact for stakeholder review.
    • C. An operational dashboard reports runtime metrics like latency and token usage, which is useful for monitoring but does not capture the model's intended use, limitations, or evaluation findings.
    • D. An IAM policy controls access permissions to the model endpoint; it has no role in documenting the model's use cases, limitations, or evaluation results.

    Subdomain 3.4: Implement responsible AI principles.

    23.A pharmaceutical customer support assistant retrieves passages from an approved clinical reference and generates a response summarizing drug interaction warnings. In one exchange, the generated response states an interaction risk that does not appear anywhere in the retrieved passages, effectively inventing a claim not supported by the source. Which Amazon Bedrock Guardrails capability is specifically designed to detect this kind of unsupported claim?

    1. A.A contextual grounding check that compares the generated response against the retrieved source and flags unsupported claims.
    2. B.A denied topics filter configured to block any response that mentions a drug interaction, regardless of whether it is accurate or invented.
    3. C.A word filter listing specific drug names so that any response naming a blocked drug is stopped before it reaches the customer.
    4. D.A content filter set to the highest strength for the misconduct category so that any medical claim is treated as high risk.
    Show answer & explanation

    Correct answer: A — A contextual grounding check that compares the generated response against the retrieved source and flags unsupported claims.

    • A. A contextual grounding check exists precisely to compare a generated response against the retrieved source and flag or block content that is not grounded in it, which is what catches an invented interaction claim.
    • B. A denied topics filter blocking all drug-interaction mentions would also block accurate, correctly grounded warnings, which is not targeted at distinguishing an invented claim from a supported one.
    • C. A word filter matches specific listed terms; it has no mechanism for comparing a generated statement against retrieved source material to determine whether the claim is actually supported.
    • D. The misconduct content category targets a specific class of harmful content unrelated to factual grounding, so raising its strength does nothing to detect whether a medical claim matches the retrieved source.

    Subdomain 3.2: Implement data security and privacy controls.

    24.A retail company centralizes structured sales data in an AWS Lake Formation data catalog and wants a Bedrock Agent to answer analyst questions by querying only the tables relevant to each analyst's department, without duplicating the data per department. The security team also wants column-level restriction so agents never expose customer payment fields even to authorized departments. Which of the following AWS Lake Formation capabilities directly satisfy these requirements?(Select 3)

    1. A.Table-level and column-level permissions that let administrators grant a department read access to specific tables while excluding sensitive columns such as payment fields.
    2. B.Data filters that provide row and cell-level security so the same underlying table returns different rows or masked cells depending on the requesting principal's grant.
    3. C.LF-Tags that let administrators tag tables and columns once and grant access by tag expression, so department access scales without per-table permission edits.
    4. D.Cross-account resource links that copy the sales tables into a separate AWS account per department, duplicating the full dataset to guarantee physical isolation between departments.
    5. E.Query result caching that stores each analyst's most recent query output in a shared cache layer to reduce repeated full-table scans of the underlying sales dataset.
    6. F.Blueprint workflows that automatically ingest new sales data from on-premises relational databases into the catalog on a recurring, administrator-defined schedule.
    Show answer & explanation

    Correct answers: A, B, C — Table-level and column-level permissions that let administrators grant a department read access to specific tables while excluding sensitive columns such as payment fields.; Data filters that provide row and cell-level security so the same underlying table returns different rows or masked cells depending on the requesting principal's grant.; LF-Tags that let administrators tag tables and columns once and grant access by tag expression, so department access scales without per-table permission edits.

    • A. Column-level permissions let an administrator grant a department access to a table while excluding specific columns such as payment fields, which is exactly the fine-grained exclusion the requirement describes. This works on the single centralized copy of the data, so no duplication is needed.
    • B. Data filters apply row and cell-level rules so different departments querying the same physical table see only their department's rows or masked versions of restricted cells. This satisfies both the department scoping and payment-field protection without copying data.
    • C. LF-Tags let administrators tag resources once and grant access through tag expressions, which scales department-based access as new tables are added without re-granting permissions table by table. This directly supports centralized, non-duplicated access management.
    • D. Copying tables into per-department accounts creates physical duplicates of the same data, which directly contradicts the requirement to avoid duplicating data per department. It also multiplies the operational burden of keeping copies synchronized.
    • E. Query result caching only affects performance by reducing repeated scans; it has no access control mechanism and does nothing to restrict which tables, rows, or columns a department can see.
    • F. Blueprint workflows automate data ingestion from source systems into the catalog, which addresses how data arrives rather than how access to that data is restricted once it is there. This does not provide any department- or column-level access control.

    Subdomain 3.2: Implement data security and privacy controls.

    25.A fintech company wants its Bedrock-powered assistant to help support agents troubleshoot account issues using real customer data, but does not want the foundation model to ever see raw account numbers or full names during the conversation. The masked values must map back to the same customer consistently across multiple messages in a session so the agent's follow-up questions still make sense. Which masking approach satisfies this?

    1. A.Apply deterministic tokenization at the application layer before sending text to the model, replacing each account number and name with a consistent token per customer for the session.
    2. B.Apply random redaction at the application layer, replacing each account number and name with a freshly generated random string on every single message sent to the model.
    3. C.Rely entirely on the foundation model's context window to automatically forget account numbers and names after each individual message, with no masking of any kind applied before submission.
    4. D.Apply Bedrock Guardrails contextual grounding checks to the conversation so that any response referencing an account number is flagged as a potential hallucination.
    Show answer & explanation

    Correct answer: A — Apply deterministic tokenization at the application layer before sending text to the model, replacing each account number and name with a consistent token per customer for the session.

    • A. Deterministic tokenization maps the same account number or name to the same token every time within a session, so the model never sees the raw value but can still track that later messages refer to the same customer, which follow-up questions depend on.
    • B. Generating a fresh random string for every message breaks the mapping between a customer and their masked identifier from one message to the next, so the model cannot recognize that later messages refer to the same account.
    • C. The model's context window retains whatever raw text is sent to it during a session; it does not automatically strip or forget sensitive values, so raw account numbers and names would still reach the model with no masking applied at all.
    • D. Contextual grounding checks evaluate whether a response is factually supported by retrieved source content; they are designed to catch hallucinations, not to mask or replace sensitive values before the model ever sees them.

    Subdomain 3.3: Implement AI governance and compliance mechanisms.

    26.A multinational bank is standing up an internal governance program covering every GenAI application it builds, across multiple business units, each currently choosing its own foundation models and safety controls independently. The chief risk officer wants a system that keeps oversight consistent across units while still letting each team ship. Which elements should the governance program include? (Select THREE.)(Select 3)

    1. A.A requirement that every deployed model have a completed model card documenting its intended use, evaluation results, and risk rating on file.
    2. B.A cross-functional review board that approves new GenAI use cases against the policy framework and regulatory requirements before production deployment begins.
    3. C.A one-time policy document published at program kickoff, with no defined cadence for revisiting or updating it as regulations change.
    4. D.A rule that only the central risk team may write any prompt or configure any guardrail, with every business unit submitting requests by email each single time.
    5. E.A central policy framework defining minimum required controls, like guardrail categories and logging destinations, every application must satisfy at launch.
    Show answer & explanation

    Correct answers: A, B, E — A requirement that every deployed model have a completed model card documenting its intended use, evaluation results, and risk rating on file.; A cross-functional review board that approves new GenAI use cases against the policy framework and regulatory requirements before production deployment begins.; A central policy framework defining minimum required controls, like guardrail categories and logging destinations, every application must satisfy at launch.

    • A. Requiring a completed model card on file for every deployed model gives the program a consistent, auditable artifact per model, aligning organizational oversight with the same documentation standard across every business unit.
    • B. A cross-functional review board that checks each new use case against the policy framework before production is the structural gate that prevents inconsistent, ungoverned launches while still letting teams ship once they pass review.
    • C. A policy with no review cadence goes stale as regulations and business needs change, so it fails to provide the consistent, current oversight a governance program is meant to maintain over time.
    • D. Centralizing every prompt and guardrail configuration change through email requests to one team creates a bottleneck that directly contradicts the risk officer's requirement to keep teams able to ship, rather than establishing scalable oversight.
    • E. A central policy framework setting minimum required controls gives every business unit a shared, consistent baseline to build against, which is the mechanism that keeps oversight consistent across units that otherwise choose their own models and controls independently.

    Subdomain 3.3: Implement AI governance and compliance mechanisms.

    27.An auditor asks whether AWS CloudTrail alone is sufficient to reconstruct the exact text a foundation model returned to a specific user last month. What is the correct answer, and why?

    1. A.Yes, because Bedrock automatically forwards every model response body into the same CloudTrail event that records the InvokeModel call.
    2. B.No, because CloudTrail logs that InvokeModel ran and by whom, not the response text; reconstructing it needs model invocation logging.
    3. C.No, because CloudTrail only records console sign-in events and cannot see any programmatic API calls made through the SDK.
    4. D.Yes, because the CloudTrail Event history console shows the model's full response text inside the InvokeModel event's response fields.
    Show answer & explanation

    Correct answer: B — No, because CloudTrail logs that InvokeModel ran and by whom, not the response text; reconstructing it needs model invocation logging.

    • A. Bedrock does not forward model response bodies into CloudTrail events; response content capture is the purpose of the separate model invocation logging feature, which writes to CloudWatch Logs or S3, not to CloudTrail.
    • B. CloudTrail is an API activity log: it records that a principal called InvokeModel, on which model, and when, but the event's responseElements field does not carry the model's response body, so reconstructing the exact returned text requires Bedrock model invocation logging instead, delivered to CloudWatch Logs or S3.
    • C. CloudTrail records both console actions and programmatic SDK or CLI API calls; it is not limited to console sign-in events, so this understates what CloudTrail actually covers even though it correctly concludes CloudTrail cannot reconstruct the response text.
    • D. CloudTrail's own published example log entry for the InvokeModel action shows the responseElements field as null: the event records that the call happened and its request parameters, not the model's output text, so there is no response text to view in Event history even though InvokeModel is logged by default as a management event.

    Domain 4: Operational Efficiency and Optimization for GenAI Applications

    Subdomain 4.1: Implement cost optimization and resource efficiency strategies.

    28.A team is preparing to purchase Bedrock Provisioned Throughput for a custom fine-tuned model that will serve a mobile app's real-time chat feature. They need to determine the right number of Model Units to avoid both throttling during peak usage and paying for capacity that sits unused most of the time. Which two steps should be part of this capacity-planning process?(Select 2)

    1. A.Measure the peak input and output token rate the workload actually needs per minute during the busiest observed traffic period, and size the Model Unit count to cover that peak rather than the daily average.
    2. B.Review historical or projected traffic patterns to choose a commitment term, no commitment, one month, or six months, matching how confident the team is in sustained demand versus flexibility.
    3. C.Provision the maximum number of Model Units the account quota allows, regardless of measured traffic, since Provisioned Throughput cannot be modified or deleted once purchased.
    4. D.Skip any traffic measurement and provision one Model Unit as a default starting point, since a single Model Unit provides unlimited throughput for any foundation model regardless of request volume.
    5. E.Assume on-demand and Provisioned Throughput always cost the same per token processed, so the number of Model Units purchased has no effect on the workload's total monthly bill.
    Show answer & explanation

    Correct answers: A, B — Measure the peak input and output token rate the workload actually needs per minute during the busiest observed traffic period, and size the Model Unit count to cover that peak rather than the daily average.; Review historical or projected traffic patterns to choose a commitment term, no commitment, one month, or six months, matching how confident the team is in sustained demand versus flexibility.

    • A. Provisioned Throughput capacity is defined by how many input and output tokens per minute a Model Unit can sustain, so sizing to the actual peak rate is what prevents throttling during the busiest periods without over-provisioning for the rest of the day.
    • B. The commitment term directly trades flexibility for a lower hourly rate, so matching the term to how confident the team is in sustained demand avoids either paying a premium for unnecessary flexibility or locking into capacity longer than justified.
    • C. Provisioned Throughput can be modified or deleted, subject to the chosen commitment term, and provisioning the account maximum regardless of measured traffic wastes money on capacity the workload does not use.
    • D. A Model Unit delivers a specific, bounded throughput level, not unlimited throughput, so a single unit sized without reference to actual measured traffic will throttle a workload whose real demand exceeds that unit's capacity.
    • E. On-demand and Provisioned Throughput use different pricing models entirely, hourly commitment versus per-token, so the number of Model Units purchased directly determines fixed hourly cost independent of how many tokens are actually processed.

    Subdomain 4.3: Implement monitoring systems for GenAI applications.

    29.A knowledge-base-backed RAG application indexes product manuals into an Amazon OpenSearch Serverless vector collection. Over several months of continuous ingestion, query latency has gradually increased even though query volume and document count have stayed roughly constant. Which vector store operational management practice would have surfaced this degradation before users noticed slower responses?

    1. A.Switch the knowledge base's vector store from OpenSearch Serverless to a relational database with a JSON column, since relational engines do not exhibit segment fragmentation over time.
    2. B.Increase the chunk size used during document ingestion so fewer, larger chunks are stored, since larger chunks always reduce the total number of vectors and therefore reduce latency.
    3. C.Periodically re-embed every document in the knowledge base using a newer embedding model version, since embedding model upgrades are the primary cause of vector search latency growth.
    4. D.Continuously monitor OpenSearch Serverless CloudWatch latency and request-rate metrics, and schedule automated index optimization routines such as force-merge to consolidate segments.
    Show answer & explanation

    Correct answer: D — Continuously monitor OpenSearch Serverless CloudWatch latency and request-rate metrics, and schedule automated index optimization routines such as force-merge to consolidate segments.

    • A. A relational database with a JSON column has no native approximate nearest neighbor vector index at all, so it is not a comparable vector store and does not address the operational monitoring gap the scenario describes.
    • B. Larger chunks reduce vector count but change retrieval granularity and answer quality, and chunk size has no direct relationship to segment fragmentation accumulated from continuous document ingestion over months.
    • C. Re-embedding with a newer model changes vector representations and could even require reindexing, but it does not address segment fragmentation from ongoing ingestion, which is the mechanism behind steadily growing search latency with stable volume.
    • D. Tracking OpenSearch Serverless's own latency and request-rate metrics over time would show the gradual trend directly, and scheduled index optimization routines address the underlying cause of segment fragmentation from continuous ingestion that typically drives this kind of slow latency creep.

    Subdomain 4.2: Optimize application performance.

    30.A product team shipped a new prompt template for their support chatbot and wants to confirm, with evidence rather than opinion, whether it actually resolves more customer questions on the first reply than the template it is replacing before rolling it out to all traffic. Which approach best fits this goal?

    1. A.Run an A/B test that splits live traffic between the current and new templates and compares first-reply resolution rate between the two groups over the same period.
    2. B.Adjust the temperature parameter upward on the new template so its responses appear more varied and confident to the support team reviewing transcripts by hand.
    3. C.Enable prompt caching on the new template so its responses return faster, then assume the lower latency reflects a quality improvement over the old template.
    4. D.Deploy the new template to 100% of traffic immediately and compare its resolution rate against last quarter's historical average from the old template's performance.
    Show answer & explanation

    Correct answer: A — Run an A/B test that splits live traffic between the current and new templates and compares first-reply resolution rate between the two groups over the same period.

    • A. An A/B test that runs both templates concurrently against comparable live traffic isolates the effect of the prompt change itself, giving a direct, evidence-based comparison of first-reply resolution rate between the two versions.
    • B. Raising temperature changes how varied the wording sounds but has no bearing on whether the underlying prompt template actually resolves more customer questions, and it does not produce a comparison against the old template at all.
    • C. Prompt caching can reduce response latency, but faster responses are not evidence of better resolution quality, and this approach never compares the two templates against each other.
    • D. Comparing the new template's live performance against a prior quarter's historical average confounds seasonal and traffic differences with the prompt change, so it does not isolate whether the new template caused the difference.

    Subdomain 4.2: Optimize application performance.

    31.An observability team wants to right-size a GenAI application's auto-scaling policy but currently only tracks CPU and memory utilization on the application servers that call Bedrock, not anything about the model traffic itself. Requests are timing out during load spikes even though CPU and memory stay comfortably low. Which two monitoring additions would most directly explain and help correct this scaling gap? (Select TWO.)(Select 2)

    1. A.Track prompt and completion token volume and invocation counts against throughput quotas, since GenAI cost is driven by token throughput, not server CPU.
    2. B.Track queue depth and the number of concurrent in-flight Bedrock requests against the account's configured concurrency limits during traffic spikes.
    3. C.Add disk I/O monitoring to the application servers, since GenAI workloads are typically constrained by local storage read and write speed.
    4. D.Increase the CPU and memory utilization thresholds that trigger auto-scaling, so the policy waits longer before adding capacity during a spike.
    5. E.Replace CPU-based auto-scaling with a fixed schedule that adds server capacity at the same hours daily regardless of traffic patterns, ignoring token volume.
    Show answer & explanation

    Correct answers: A, B — Track prompt and completion token volume and invocation counts against throughput quotas, since GenAI cost is driven by token throughput, not server CPU.; Track queue depth and the number of concurrent in-flight Bedrock requests against the account's configured concurrency limits during traffic spikes.

    • A. Because the application servers are just thin callers to Bedrock, the real constraint during a spike is token throughput against the account's per-minute quotas, so surfacing prompt/completion token volume and invocation-rate metrics reveals the actual bottleneck that CPU and memory cannot show.
    • B. Queue depth and concurrent in-flight request counts reveal exactly when the fleet is approaching its effective concurrency ceiling, giving a second leading indicator that CPU and memory metrics cannot surface.
    • C. Disk I/O on the calling servers is unrelated to Bedrock's model-side throughput limits; the application servers making lightweight API calls are unlikely to be storage-bound, so this metric would not explain the timeouts.
    • D. Raising the CPU and memory thresholds would delay scaling further and does not address that CPU and memory were never the constraint, since they were already low when the timeouts occurred.
    • E. A fixed daily schedule adds capacity independent of actual traffic or token-throughput conditions, so it would not target the load spikes that are the actual cause of the timeouts.

    Domain 5: Testing, Validation, and Troubleshooting

    Subdomain 5.1: Implement evaluation systems for GenAI.

    32.A subscription-box company's internal chatbot receives a steady stream of thumbs-up and thumbs-down ratings from employees after each answer. Product wants this signal to actually change which prompt template and model configuration ships next, rather than sitting unused in a database. Which mechanism turns this raw feedback into a repeatable improvement loop?

    1. A.Route the rated responses into an annotation workflow that groups low-rated responses by failure pattern, then feed the insights into targeted changes that are re-evaluated before rollout.
    2. B.Delete all thumbs-down ratings from the database on a weekly schedule, so the visible average rating trends steadily upward regardless of the underlying response quality employees actually saw.
    3. C.Display only the raw thumbs-up count on an internal dashboard as the sole success metric, without ever inspecting the content of the responses employees separately rated poorly.
    4. D.Require every employee to submit a written justification alongside each thumbs-up rating, but leave every thumbs-down rating completely unstructured and effectively impossible to categorize.
    Show answer & explanation

    Correct answer: A — Route the rated responses into an annotation workflow that groups low-rated responses by failure pattern, then feed the insights into targeted changes that are re-evaluated before rollout.

    • A. An annotation workflow that clusters low-rated responses by the specific way they failed converts raw ratings into actionable patterns, letting the team make targeted changes and re-evaluate them, which closes the feedback loop the scenario asks for.
    • B. Deleting negative ratings hides the underlying quality problem rather than fixing it, and it removes exactly the signal that would drive a genuine improvement in the prompt or configuration.
    • C. A raw thumbs-up count with no analysis of the negative responses provides a vanity metric but no insight into what is going wrong, so it cannot drive targeted prompt or configuration changes.
    • D. Structuring only the positive feedback while leaving negative feedback unstructured discards the responses most useful for identifying failure patterns, since thumbs-down ratings are the ones that need categorization to act on.

    Subdomain 5.1: Implement evaluation systems for GenAI.

    33.A team is deciding which evaluation approach to apply first when validating a new prompt template for a customer-facing GenAI assistant before it reaches production. Which sequence of evaluation stages reflects a sound, systematic quality assurance process, matching automated checks before the update reaches broader validation and rollout?

    1. A.Run automated quality gates and regression tests against known cases first, then proceed to broader human or LLM-as-a-judge review and a staged canary rollout only after the automated checks pass.
    2. B.Route every prompt change directly into a full production rollout first, and only run automated regression tests afterward if customers report a problem with the new template.
    3. C.Have a single engineer manually approve the change based on personal impression alone, and skip both automated regression testing and any staged rollout entirely.
    4. D.Run the staged canary rollout to 100% of production traffic first, then run automated regression tests afterward purely for historical record-keeping with no ability to halt the rollout once begun.
    Show answer & explanation

    Correct answer: A — Run automated quality gates and regression tests against known cases first, then proceed to broader human or LLM-as-a-judge review and a staged canary rollout only after the automated checks pass.

    • A. Running automated quality gates and regression tests first catches known-failure patterns cheaply and quickly, and only promoting a change to broader review and staged rollout after passing those checks forms the systematic order a mature QA process follows.
    • B. Rolling out to full production before any automated regression testing removes the early, low-cost detection step and turns customer complaints into the first line of defense, which is the reverse of a systematic quality assurance process.
    • C. A single engineer's personal impression is not a repeatable or systematic check, and skipping both automated testing and staged rollout removes the structured gates a quality assurance process depends on.
    • D. Rolling out to all traffic before running any regression tests defeats the purpose of a canary rollout, which is meant to limit exposure while the change is still being validated, not to be run at full scale first.

    Subdomain 5.2: Troubleshoot GenAI applications.

    34.A developer integrating Amazon Bedrock's InvokeModel API receives a `ValidationException` on a subset of requests, while functionally similar requests from other application components succeed. What is the most likely general cause of a `ValidationException` from a Bedrock model API, and what should the developer check first?

    1. A.The request body does not conform to the target model's expected schema or parameter constraints, so the developer should compare the request's JSON body against the model's documented schema.
    2. B.The AWS account has been temporarily suspended for non-payment, so the developer should check the AWS Billing console for an overdue invoice before touching the request payload.
    3. C.The Bedrock service in that Region is undergoing scheduled maintenance, so the developer should check the AWS Health Dashboard and simply wait for the maintenance window to end.
    4. D.The IAM role lacks the `bedrock:InvokeModel` permission, so the developer should check the role's attached IAM policy for a missing `Allow` statement covering that specific Bedrock invocation action.
    Show answer & explanation

    Correct answer: A — The request body does not conform to the target model's expected schema or parameter constraints, so the developer should compare the request's JSON body against the model's documented schema.

    • A. A ValidationException signals a malformed request such as wrong parameter names, types, or out-of-range values for that specific model, so comparing the payload against the documented schema is the right first check.
    • B. A billing suspension produces an account-level access failure affecting all requests, not a per-request ValidationException tied to a specific malformed body, so it would not explain selective failures.
    • C. Regional service maintenance typically surfaces as throttling or service-unavailable errors across most requests, not a validation error scoped to a particular request's content.
    • D. A missing invoke permission produces an AccessDeniedException rather than a ValidationException, so a policy gap does not explain this specific error code.

    Subdomain 5.2: Troubleshoot GenAI applications.

    35.A team runs an A/B comparison between two versions of a customer-support prompt template using Bedrock Model Evaluation. The results show the new version scores higher on conciseness but lower on factual accuracy and consistency across repeated runs of the same input. The team wants a refinement workflow that addresses the regression without discarding the conciseness gain. Select all statements that describe a valid step in a systematic refinement workflow for this situation.(Select 3)

    1. A.Isolate which specific instruction changed between the two template versions and test that change independently against the same evaluation set to confirm it is the source of the accuracy regression.
    2. B.Re-run the lower-scoring template version several times against the same inputs to check whether the inconsistency is inherent to the prompt wording or attributable to sampling variance at the current temperature setting.
    3. C.Iterate on the instruction wording that caused the regression, for example by adding explicit grounding or verification steps, then re-score the revised version against the same evaluation set before promoting it.
    4. D.Discard the evaluation results entirely and promote the new template to production immediately, since a conciseness improvement outweighs a factual accuracy regression in a support context.
    5. E.Permanently pin the model's temperature to zero for every future prompt template built in the application, since a single template's consistency regression means temperature should never be allowed to vary again anywhere.
    Show answer & explanation

    Correct answers: A, B, C — Isolate which specific instruction changed between the two template versions and test that change independently against the same evaluation set to confirm it is the source of the accuracy regression.; Re-run the lower-scoring template version several times against the same inputs to check whether the inconsistency is inherent to the prompt wording or attributable to sampling variance at the current temperature setting.; Iterate on the instruction wording that caused the regression, for example by adding explicit grounding or verification steps, then re-score the revised version against the same evaluation set before promoting it.

    • A. Isolating the specific instruction change and testing it independently is the systematic way to confirm which wording caused the accuracy regression before touching anything else.
    • B. Checking whether inconsistency persists across repeated runs distinguishes an inherent prompt problem from ordinary sampling variance, which changes what fix is actually needed.
    • C. Iteratively refining the identified problem wording and re-scoring against the same evaluation set is the core loop of a systematic, repeatable refinement workflow.
    • D. Promoting a template with a confirmed factual accuracy regression in a customer-support context without addressing the cause ignores the evaluation results rather than acting on them.
    • E. A single template's consistency issue does not justify a blanket, permanent temperature change across every other template in the application; that generalizes one finding far beyond its evidence.

    Want the full experience?

    These are just samples. Practice the full AWS Certified Generative AI Developer - Professional (AIP-C01) question bank in quiz mode — free, no signup, with domain practice and exam simulation.