CertSafari

    Free AWS Certified Solutions Architect - Associate (SAA-C03) Sample Questions

    35 free sample questions from our bank of 349+, covering every exam domain, with answers and detailed explanations. Updated September 2026.

    Domain 1: Design Secure Architectures

    Subdomain 1.2: Design secure workloads and applications

    1.A company is building a mobile application that must let customers sign up and sign in with an email address and password, and after authentication the app must obtain temporary AWS credentials so it can upload photos directly to an Amazon S3 bucket. Which combination of Amazon Cognito components should the solutions architect use?

    1. A.Use a Cognito user pool to handle customer sign-up and sign-in, then configure a Cognito identity pool that accepts the user pool's tokens and exchanges them for temporary IAM credentials scoped to the S3 upload permissions.
    2. B.Use a Cognito identity pool to handle customer sign-up and sign-in directly, then configure a Cognito user pool to exchange the identity pool's session tokens for temporary IAM credentials scoped to S3.
    3. C.Use a single Cognito user pool for both authentication and AWS credential exchange, since user pools issue IAM-scoped temporary credentials directly to authenticated app users without needing a separate identity pool at all.
    4. D.Use AWS IAM Identity Center to authenticate the mobile application's customers against a corporate directory, then assign each customer a permission set that grants direct access to the S3 bucket.
    Show answer & explanation

    Correct answer: A — Use a Cognito user pool to handle customer sign-up and sign-in, then configure a Cognito identity pool that accepts the user pool's tokens and exchanges them for temporary IAM credentials scoped to the S3 upload permissions.

    • A. A user pool is a user directory that handles sign-up, sign-in, and issues authentication tokens after a successful login. An identity pool takes those tokens as a trusted identity provider and exchanges them for temporary IAM credentials, which is exactly the pairing needed to authenticate customers and then let the app call S3 directly.
    • B. This reverses the roles of the two components: identity pools do not provide sign-up or sign-in functionality, and user pools do not exchange session tokens for AWS credentials. Sign-up and sign-in belong to the user pool, and credential exchange belongs to the identity pool.
    • C. A user pool issues JSON web tokens after authentication, but it does not issue temporary AWS IAM credentials on its own; that exchange is the specific job of an identity pool, which is required to grant the app access to call S3.
    • D. IAM Identity Center is designed for workforce access to AWS accounts and business applications through a corporate or external directory, not for authenticating a public mobile application's customers, so it is not the appropriate service for this consumer-facing sign-up flow.

    Subdomain 1.2: Design secure workloads and applications

    2.Application servers running in a private subnet need to read and write objects in Amazon S3 as part of their normal processing. The company wants this traffic to stay on the AWS network instead of traversing the NAT gateway and the public internet, and wants to avoid the per-hour and per-gigabyte charges the NAT gateway incurs for this traffic. Which change should the solutions architect make?

    1. A.Create a gateway VPC endpoint for Amazon S3 and add a route to it in the private subnet's route table, so S3 traffic reaches the service over the AWS network without passing through the NAT gateway.
    2. B.Create an interface VPC endpoint for Amazon S3 in the public subnet, and update the private subnet's security group to allow outbound traffic to the endpoint's public IP address.
    3. C.Increase the NAT gateway's bandwidth allocation and add a second NAT gateway in the same subnet, so S3 traffic can load balance across both gateways and reduce the per-gateway data processing charges.
    4. D.Configure AWS Direct Connect between the VPC and the nearest AWS Region, and route all S3 traffic over the dedicated connection instead of through the NAT gateway.
    Show answer & explanation

    Correct answer: A — Create a gateway VPC endpoint for Amazon S3 and add a route to it in the private subnet's route table, so S3 traffic reaches the service over the AWS network without passing through the NAT gateway.

    • A. Amazon S3 supports a gateway-type VPC endpoint, which is added as a route table target rather than a network interface; adding that route to the private subnet's route table sends S3 traffic over the AWS network and bypasses the NAT gateway entirely, along with its data processing charges.
    • B. S3 uses a gateway endpoint, not an interface endpoint, and VPC endpoints use private IP addresses within the VPC rather than a public IP address, so this configuration describes the wrong endpoint type and an incorrect addressing model.
    • C. Adding NAT gateway capacity or a second NAT gateway increases resilience and throughput, but the traffic still passes through a NAT gateway and continues to incur its hourly and per-gigabyte charges, so it does not meet the goal of avoiding those costs.
    • D. Direct Connect provides a dedicated network connection between an on-premises location and AWS; it is not used to route traffic between a VPC's private subnet and S3 within the same Region, and setting it up for this purpose would add unnecessary cost and complexity.

    Subdomain 1.1: Design secure access to AWS resources

    3.A company's security team must grant a third-party auditing firm temporary, read-only access to CloudTrail logs and IAM configuration data in the production AWS account. The auditors already manage their own separate AWS account and the access must not require long-term credentials in the production account. Which approach meets this requirement with the least ongoing operational overhead?

    1. A.Create an IAM role in the production account with a trust policy naming the auditor's account, then let auditors call `sts:AssumeRole` to obtain temporary, read-only scoped credentials.
    2. B.Create an individual IAM user in the production account for each auditor, attach a read-only managed policy, and require the team to rotate every user's access keys every 90 days by hand.
    3. C.Share the production account's root user credentials with the auditing firm for the engagement, then rotate the root password once the audit work has been completed.
    4. D.Enable IAM Identity Center in the production account and build a permission set that the auditing firm's external identity provider signs users into directly.
    Show answer & explanation

    Correct answer: A — Create an IAM role in the production account with a trust policy naming the auditor's account, then let auditors call `sts:AssumeRole` to obtain temporary, read-only scoped credentials.

    • A. A cross-account IAM role with a trust policy scoped to the auditor's account lets the auditors assume the role via AWS STS and receive short-lived credentials, with no standing user or keys to manage or leak. Revoking access is as simple as removing the trust relationship.
    • B. Creating per-auditor IAM users introduces long-term credentials and a key-rotation burden that the requirement explicitly rules out. It also requires manual user lifecycle management every time an auditor joins or leaves the engagement.
    • C. Sharing root credentials violates the principle of least privilege and AWS security best practice, since the root user has unrestricted access to the entire account far beyond what an audit requires. Root credentials should never be distributed to external parties.
    • D. IAM Identity Center centralizes workforce access to accounts a business already manages under AWS Organizations; wiring an external firm's separate identity provider into the audited account's Identity Center instance adds federation setup overhead the scenario does not call for.

    Subdomain 1.1: Design secure access to AWS resources

    4.A platform team lets application teams create their own IAM roles but wants to cap the maximum permissions any self-service role can ever have, without restricting the platform team's own administrative roles in the same account. Which mechanism fits this single-account requirement?

    1. A.Attach a permissions boundary to the self-service roles that sets the maximum permissions those roles can have, leaving the platform team's roles unaffected.
    2. B.Apply a service control policy in AWS Organizations to the account, since SCPs can target individual roles within a single account without exceptions.
    3. C.Add an IAM group for application teams and rely on group membership alone to cap permissions, since groups enforce a hard permission ceiling by default.
    4. D.Configure AWS Config to automatically detach any policy exceeding a defined permission threshold from roles created by application teams after the fact.
    Show answer & explanation

    Correct answer: A — Attach a permissions boundary to the self-service roles that sets the maximum permissions those roles can have, leaving the platform team's roles unaffected.

    • A. This is correct because a permissions boundary sets the maximum permissions a role can ever have regardless of the policies later attached to it, and it can be applied selectively to self-service roles only.
    • B. Service control policies apply to entire accounts or organizational units, not to individual roles within a single account, so they cannot selectively cap only self-service roles.
    • C. IAM group membership controls which policies are attached to a user but does not impose any ceiling on what permissions a role created by that user can have.
    • D. AWS Config can detect policies exceeding a threshold and alert on them, but it does not proactively cap what permissions a role can be granted at creation time.

    Subdomain 1.1: Design secure access to AWS resources

    5.An engineer with an IAM user in a development account needs to periodically review resources in a separate production account through the AWS Management Console, without maintaining separate login credentials for the production account. What should be configured?

    1. A.Create a cross-account IAM role in the production account that trusts the development account, so the engineer can switch roles from the console.
    2. B.Share the production account's root user credentials with the engineer over an encrypted channel, restricting the password rotation frequency to weekly.
    3. C.Create an identical IAM user with the same username and password in the production account so credentials stay consistent across both accounts.
    4. D.Grant the engineer's development-account user an inline policy naming resources in the production account, since IAM policies apply across account boundaries by default.
    Show answer & explanation

    Correct answer: A — Create a cross-account IAM role in the production account that trusts the development account, so the engineer can switch roles from the console.

    • A. This is correct because a cross-account role with a trust policy naming the development account lets the engineer switch roles from the console using their existing development-account credentials.
    • B. Sharing root user credentials, even over an encrypted channel, creates unrestricted long-term access shared between two people, which is the opposite of scoped, auditable access.
    • C. Creating a duplicate IAM user with matching credentials in the production account requires maintaining a second identity and does not use temporary, auditable cross-account access.
    • D. IAM policies are scoped to the account that owns them and do not grant access to resources in a different account without an explicit cross-account trust relationship.

    Subdomain 1.3: Determine appropriate data security controls

    6.A company runs a production Amazon RDS for MySQL database and must be able to restore the database to any point in time within the past 7 days after an accidental data deletion, without relying on operators remembering to take manual snapshots. Which configuration meets this requirement?

    1. A.Configure a maintenance window during which an administrator manually creates a database snapshot every night and deletes any snapshot older than 7 days.
    2. B.Enable automated backups on the RDS instance with a 7-day backup retention period, which captures transaction logs continuously to support point-in-time restore.
    3. C.Enable RDS read replicas in two additional Availability Zones, since replica databases automatically retain 7 days of restorable transaction history for point-in-time recovery.
    4. D.Rely on the default database engine configuration, since Amazon RDS enables continuous backups and point-in-time restore automatically with no setup required.
    Show answer & explanation

    Correct answer: B — Enable automated backups on the RDS instance with a 7-day backup retention period, which captures transaction logs continuously to support point-in-time restore.

    • A. Manually triggered nightly snapshots only allow restoring to the moment each snapshot was taken, not to an arbitrary point in time between snapshots, and they depend on an operator remembering to run and prune them. This does not provide true point-in-time restore.
    • B. Automated backups take a daily snapshot and continuously capture transaction logs, letting the instance be restored to any second within the configured retention window, up to 7 days here, without any manual intervention. This is the built-in feature designed for exactly this recovery requirement.
    • C. Read replicas exist to offload read traffic and improve availability, but they are not a backup retention mechanism and do not provide a 7-day restorable transaction history for point-in-time recovery. Data deleted on the primary also propagates to replicas.
    • D. Automated backups and point-in-time restore are available on RDS, but they are not enabled by default and require an administrator to configure a non-zero backup retention period on the instance. Assuming default behavior would leave the database without the needed recovery window.

    Subdomain 1.2: Design secure workloads and applications

    7.A media company's live-streaming platform behind CloudFront and an Application Load Balancer has suffered repeated large-scale volumetric DDoS attacks that degrade performance for legitimate viewers. Leadership wants proactive mitigation, real-time attack visibility, and direct access to AWS's DDoS response team. Which option should the architect recommend?

    1. A.Subscribe to AWS Shield Advanced on the CloudFront distribution and ALB, which adds application-layer DDoS mitigation, attack visibility, and Shield Response Team support.
    2. B.Rely on AWS Shield Standard, which already includes application-layer mitigation and direct engagement with the AWS Shield Response Team at no added cost at all.
    3. C.Deploy AWS Firewall Manager alone across the accounts, since it independently detects and mitigates volumetric DDoS attacks without needing any other supporting service at all.
    4. D.Configure Amazon GuardDuty with S3 protection enabled, which extends its findings to cover network and transport layer DDoS traffic patterns.
    Show answer & explanation

    Correct answer: A — Subscribe to AWS Shield Advanced on the CloudFront distribution and ALB, which adds application-layer DDoS mitigation, attack visibility, and Shield Response Team support.

    • A. Shield Advanced adds expanded DDoS protection including automatic application-layer mitigation, detailed attack diagnostics, and access to the Shield Response Team, which matches the proactive support the company is asking for.
    • B. Shield Standard is included automatically and covers common network and transport layer attacks, but it lacks the advanced application-layer mitigation, attack visibility dashboards, and Response Team access that Shield Advanced provides.
    • C. Firewall Manager centrally manages policies such as WAF rules and Shield Advanced protections across accounts, but it is a management layer and does not itself perform DDoS detection or mitigation.
    • D. GuardDuty's S3 protection detects suspicious access patterns against S3 buckets and is unrelated to mitigating network or application layer DDoS traffic against a streaming platform.

    Subdomain 1.2: Design secure workloads and applications

    8.A security team notices unusual API activity in CloudTrail logs, including calls originating from an EC2 instance's credentials being used from an IP address associated with a known malicious host. They want continuous, managed detection of this kind of anomaly across their accounts without deploying and maintaining their own detection infrastructure. Which service should they enable?

    1. A.Amazon GuardDuty, which continuously analyzes CloudTrail, VPC Flow Logs, and DNS logs using threat intelligence to flag anomalous account and instance behavior.
    2. B.Amazon Macie, which continuously scans CloudTrail management events to identify compromised IAM credentials being reused from unfamiliar geographic regions worldwide.
    3. C.AWS WAF, which inspects API Gateway request logs for credential reuse patterns and automatically revokes any IAM role session found to be compromised.
    4. D.AWS Secrets Manager, which monitors credential usage patterns across services and rotates any access key flagged as originating from a suspicious IP.
    Show answer & explanation

    Correct answer: A — Amazon GuardDuty, which continuously analyzes CloudTrail, VPC Flow Logs, and DNS logs using threat intelligence to flag anomalous account and instance behavior.

    • A. GuardDuty is a managed threat-detection service that continuously analyzes CloudTrail management and data events, VPC Flow Logs, and DNS query logs against threat intelligence feeds to surface exactly this kind of anomalous credential use.
    • B. Macie is built to discover and classify sensitive data such as PII in S3 buckets using machine learning; it does not analyze CloudTrail events for anomalous IAM credential usage across accounts.
    • C. AWS WAF filters and inspects web requests to protected resources like CloudFront or API Gateway; it has no capability to analyze CloudTrail logs or revoke IAM role sessions based on credential misuse.
    • D. Secrets Manager stores and rotates secrets on a schedule the administrator defines, but it does not monitor CloudTrail for suspicious usage patterns or automatically react to detected anomalies.

    Subdomain 1.3: Determine appropriate data security controls

    9.An engineering team is designing envelope encryption for a custom application that must encrypt large files before uploading them to S3, without sending the full file through AWS KMS for every encryption operation. Which sequence correctly describes how AWS KMS envelope encryption works for this use case?

    1. A.The application calls KMS to encrypt the entire file directly, then stores the returned ciphertext blob in S3 next to the key ID used.
    2. B.The application requests a data key from KMS, encrypts the file locally with the plaintext copy, discards it, then stores the encrypted data key.
    3. C.The application generates its own AES key entirely without ever contacting KMS at all, encrypts the file locally, then later imports that key into KMS.
    4. D.The application asks KMS for a key pair, encrypts the file with the public key directly, then stores only the private key with the file.
    Show answer & explanation

    Correct answer: B — The application requests a data key from KMS, encrypts the file locally with the plaintext copy, discards it, then stores the encrypted data key.

    • A. AWS KMS enforces a request size limit that makes direct encryption of large files through the `Encrypt` API impractical; envelope encryption exists specifically to avoid sending full file contents to KMS.
    • B. This is the correct envelope encryption pattern: `GenerateDataKey` returns both a plaintext and an encrypted copy of a data key, the plaintext key encrypts the file locally and is then discarded, and only the encrypted data key needs to be stored alongside the ciphertext.
    • C. Generating a key entirely outside KMS defeats the purpose of centralized key management and auditing, and importing a key that has already been used for encryption is not how KMS's data key workflow is intended to operate.
    • D. Envelope encryption for bulk data typically uses a symmetric data key for performance, and storing the private key alongside the encrypted file would expose the very key needed to decrypt it, defeating the purpose of encryption.

    Subdomain 1.3: Determine appropriate data security controls

    10.A financial services company needs to enforce that data uploaded to a shared S3 bucket is always encrypted using a specific customer managed KMS key, rejecting any upload that uses SSE-S3, a different KMS key, or no encryption at all. Which control achieves this at upload time?

    1. A.A bucket policy `Deny` statement using the KMS key ID condition key to reject any PUT that omits or mismatches the required key or method.
    2. B.S3 default bucket encryption set to the required key, which on its own rejects any PUT that specifies a mismatched header value.
    3. C.An S3 Lifecycle rule that deletes any object found encrypted with the wrong key or method within 24 hours of the upload.
    4. D.IAM permissions boundaries placed on every uploading role, restricting `s3:PutObject` so it only succeeds while default encryption stays on.
    Show answer & explanation

    Correct answer: A — A bucket policy `Deny` statement using the KMS key ID condition key to reject any PUT that omits or mismatches the required key or method.

    • A. A bucket policy Deny statement conditioned on the KMS key ID header actively rejects the PUT request itself when the wrong key, wrong method, or no encryption header is present, enforcing the requirement at the moment of upload.
    • B. Default bucket encryption only supplies an encryption setting when a request omits one; it does not reject or block requests that explicitly specify a different encryption method or key, so noncompliant uploads with an explicit header would still succeed.
    • C. A lifecycle rule acts after the fact on a schedule and cannot inspect or evaluate the encryption method used on an object, nor can it prevent the noncompliant object from being stored and accessed in the interim.
    • D. Permissions boundaries constrain the maximum permissions of an IAM identity but cannot evaluate a request's encryption headers, so they cannot conditionally block a PUT request based on which encryption key was specified.

    Domain 2: Design Resilient Architectures

    Subdomain 2.1: Design scalable and loosely coupled architectures

    11.An e-commerce platform stores product images, application binaries deployed to EC2 instances via a shared mount used by multiple instances concurrently, and transaction logs for a relational database engine. Which storage type pairing correctly matches product images and the shared application mount to their appropriate AWS storage service?

    1. A.Product images in Amazon S3 as object storage, and the shared application mount on Amazon EFS as file storage accessible from multiple instances at once.
    2. B.Product images in Amazon EBS as block storage, and the shared application mount on Amazon S3 as object storage mounted directly by each instance's operating system.
    3. C.Product images in Amazon EFS as file storage, and the shared application mount on Amazon EBS as block storage attached simultaneously to every instance.
    4. D.Product images in Amazon S3 as object storage, and the shared application mount on Amazon EBS as block storage shared concurrently across all instances.
    Show answer & explanation

    Correct answer: A — Product images in Amazon S3 as object storage, and the shared application mount on Amazon EFS as file storage accessible from multiple instances at once.

    • A. S3 is object storage well suited to storing discrete files like product images accessed over HTTP, and EFS is a file system built for concurrent, shared read/write access from many EC2 instances at once, which matches both requirements correctly.
    • B. EBS volumes attach to a single instance at a time in nearly all cases and are not object storage, and S3 does not natively mount as a POSIX file system for concurrent instance access, so this pairing inverts both storage types incorrectly.
    • C. EFS is file storage rather than the best fit for storing large volumes of individual product image objects served over the web, and a standard EBS volume cannot be attached to multiple instances simultaneously for shared read/write access.
    • D. S3 correctly fits the images, but a standard EBS volume is designed to attach to a single EC2 instance and cannot be shared concurrently across an entire fleet the way the shared mount requires.

    Subdomain 2.1: Design scalable and loosely coupled architectures

    12.A subscription service processes customer sign-up requests as Lambda functions triggered from an SQS queue. Occasionally, a malformed message causes the function to fail repeatedly, and the team wants to isolate these problematic messages after a set number of failed processing attempts so they do not block other messages from being processed. Which SQS feature should they configure?

    1. A.A dead-letter queue with a maxReceiveCount redrive policy, so messages that fail processing that many times are moved out of the main queue automatically.
    2. B.A FIFO queue with content-based deduplication, so any message that fails is automatically deduplicated and discarded after the first failure.
    3. C.A shorter message retention period on the main queue, so failing messages expire and disappear from the queue faster than they currently do.
    4. D.A higher maximum message size limit on the queue, since larger size limits automatically route failing messages to a separate holding area.
    Show answer & explanation

    Correct answer: A — A dead-letter queue with a maxReceiveCount redrive policy, so messages that fail processing that many times are moved out of the main queue automatically.

    • A. Configuring a dead-letter queue with a redrive policy and a maxReceiveCount threshold automatically moves a message out of the main queue once it has failed processing that many times, isolating problem messages without blocking the rest of the queue.
    • B. Content-based deduplication in a FIFO queue prevents duplicate message bodies from being enqueued within a deduplication interval; it has no relationship to isolating messages that repeatedly fail processing.
    • C. Shortening retention only changes how long an untouched message survives before automatic deletion; it does not target specifically the messages that are failing processing versus healthy ones.
    • D. The maximum message size setting controls how large a message payload can be and has no effect on routing failing messages anywhere; this is not a real behavior of that setting.

    Subdomain 2.1: Design scalable and loosely coupled architectures

    13.An order-processing application writes each order directly to a fleet of EC2 worker instances over HTTP. During flash sales, the worker fleet is overwhelmed and orders are dropped before workers can catch up. A solutions architect must decouple order submission from order processing so that request spikes no longer cause data loss, and each order is processed by exactly one worker. What should the architect do?

    1. A.Place an Amazon SQS standard queue between the order producers and the worker fleet, and have workers poll and delete messages after processing.
    2. B.Configure an Application Load Balancer in front of the worker fleet to spread incoming order requests evenly across all currently healthy instances.
    3. C.Increase the EC2 instance type of each worker to a larger size so that more orders can be processed per second during traffic spikes.
    4. D.Write each incoming order to an Amazon S3 bucket and configure workers to list the bucket on a fixed schedule to discover new orders.
    Show answer & explanation

    Correct answer: A — Place an Amazon SQS standard queue between the order producers and the worker fleet, and have workers poll and delete messages after processing.

    • A. Amazon SQS buffers incoming orders durably and lets producers publish independently of how fast workers can consume, so a spike in submissions no longer risks dropped requests. Workers poll at their own pace and delete each message only after successful processing, giving reliable at-least-once handling.
    • B. A load balancer distributes live requests across instances but still requires a worker to be immediately available to accept the connection; it does not buffer or persist orders, so a spike that outpaces the fleet still causes failed or dropped requests.
    • C. Vertical scaling raises the ceiling for a single instance but does not remove the tight coupling between producers and workers, and it does not protect against sudden spikes that exceed even the larger instance's capacity.
    • D. Polling S3 on a fixed schedule introduces processing latency and does not guarantee that only one worker picks up a given order, so it neither decouples cleanly nor prevents duplicate processing the way a managed queue does.

    Subdomain 2.1: Design scalable and loosely coupled architectures

    14.A team is containerizing an existing application and wants to run the containers on AWS without provisioning, patching, or scaling any underlying EC2 instances themselves, while still using Amazon ECS as the container orchestrator. Which combination of services meets this requirement?

    1. A.Amazon ECS with the AWS Fargate launch type, which runs each task on serverless compute capacity that AWS provisions and manages.
    2. B.Amazon ECS with the EC2 launch type, using an Auto Scaling group with a launch template that the team configures and patches.
    3. C.Amazon EKS with self-managed worker nodes, where the team installs the Kubernetes control plane components on EC2 instances directly.
    4. D.AWS Elastic Beanstalk with a Docker platform, where the team manually selects and resizes the underlying EC2 instance fleet.
    Show answer & explanation

    Correct answer: A — Amazon ECS with the AWS Fargate launch type, which runs each task on serverless compute capacity that AWS provisions and manages.

    • A. The Fargate launch type lets Amazon ECS run containers on serverless compute that AWS provisions, patches, and scales behind the scenes, so the team defines task definitions and never manages an EC2 fleet, which is exactly what the no-server-management requirement calls for.
    • B. The EC2 launch type still requires the team to size, patch, and scale the underlying Auto Scaling group of container instances themselves, which is the operational burden the requirement explicitly wants to avoid.
    • C. Self-managed EKS worker nodes put the team in charge of provisioning and patching the EC2 instances that back the Kubernetes nodes, and it also does not use Amazon ECS as the orchestrator, which the requirement specifies.
    • D. Elastic Beanstalk with manually sized EC2 instances still leaves the team responsible for the underlying instance fleet and does not orchestrate containers through Amazon ECS, so it does not satisfy either stated requirement.

    Subdomain 2.2: Design highly available and/or fault-tolerant architectures

    15.A company runs a warm standby environment in a secondary Region sized at 20 percent of production capacity. Its DR runbook calls for scaling the standby Auto Scaling group up to full production capacity within minutes of a declared disaster. During a recent DR test, the scale-out stalled because the secondary Region's default EC2 vCPU service quota was too low to launch the additional instances. Which action prevents this from recurring?

    1. A.Submit a Service Quotas increase request for the EC2 vCPU quota in the secondary Region ahead of time, sized to cover the instance count and type needed at full capacity.
    2. B.Configure the Auto Scaling group's health check grace period to a longer value so newly launched replacement instances have more time to pass checks before being marked unhealthy.
    3. C.Switch the standby Auto Scaling group to launch Spot Instances instead of On-Demand Instances, since Spot Instances are assumed to draw from a separate, effectively unlimited quota.
    4. D.Reduce the desired capacity target used during the DR test so the scale-out event requests fewer instances than the secondary Region's current default vCPU quota currently allows.
    Show answer & explanation

    Correct answer: A — Submit a Service Quotas increase request for the EC2 vCPU quota in the secondary Region ahead of time, sized to cover the instance count and type needed at full capacity.

    • A. Requesting the vCPU quota increase in advance, sized for the full failover fleet, ensures the account can actually launch the required number and type of instances the moment a disaster is declared, which is exactly what caused the test to stall.
    • B. A longer health check grace period only affects how long Auto Scaling waits before marking a launched instance unhealthy; it does not change whether the account is permitted to launch enough instances in the first place, so the underlying quota limit remains.
    • C. Spot Instances draw from their own separate EC2 vCPU-based service quotas rather than being exempt from quotas entirely, so switching purchasing options does not remove the risk of hitting a limit during a large scale-out.
    • D. Lowering the desired capacity used during testing would let the test pass, but it does not resolve the actual problem: the DR runbook still requires scaling to full production capacity, which would hit the same quota ceiling during a real disaster.

    Subdomain 2.2: Design highly available and/or fault-tolerant architectures

    16.A company maintains a warm standby environment in a secondary Region sized at 10% of production capacity. During a failover drill, scaling the standby environment up to full production capacity fails because the account's EC2 vCPU service quota in that Region is too low. What should the architect do to prevent this during a real disaster?

    1. A.Switch the standby environment to use Spot Instances exclusively, since Spot Instance capacity is not subject to the same account-level service quotas as On-Demand Instance capacity in that Region.
    2. B.Proactively request a service quota increase for EC2 vCPUs and any other constrained resources in the secondary Region so full failover capacity is available before it is needed.
    3. C.Configure AWS Auto Scaling to automatically request quota increases from AWS Support the moment a scaling activity is blocked by an insufficient service quota.
    4. D.Reduce the production Region's service quota to match the secondary Region's current quota, so both Regions stay within identical limits during normal operation.
    Show answer & explanation

    Correct answer: B — Proactively request a service quota increase for EC2 vCPUs and any other constrained resources in the secondary Region so full failover capacity is available before it is needed.

    • A. Spot Instances are still governed by their own vCPU-based service quotas in a Region, so switching purchase options does not remove the risk of hitting a quota limit during a large scale-out.
    • B. Service quota increases must be requested and approved ahead of time; requesting the needed vCPU and related quotas for the secondary Region before a disaster ensures the standby environment can actually scale to full capacity when failover is triggered.
    • C. There is no built-in mechanism for Auto Scaling to automatically request and receive a quota increase from AWS Support in real time during a scaling event, so this cannot be relied on during an active disaster.
    • D. Lowering the production Region's quota does not raise the secondary Region's quota and would instead risk constraining the primary environment, which does not solve the standby capacity problem at all.

    Subdomain 2.2: Design highly available and/or fault-tolerant architectures

    17.An Auto Scaling group uses only the default EC2 status checks to determine instance health. An application-level bug periodically hangs the web server process while the underlying EC2 instance continues to report healthy status checks, so the Auto Scaling group never replaces the hung instances. What should the architect change?

    1. A.Lower the EC2 status check grace period so the Auto Scaling group evaluates the underlying hardware and instance reachability status more frequently and reacts faster to those specific issues.
    2. B.Attach the Auto Scaling group to an Elastic Load Balancer and enable ELB health checks, so instance health is based on the load balancer's application-level checks instead of just EC2 status.
    3. C.Increase the Auto Scaling group's desired capacity so extra instances compensate for any individual instance that hangs without needing to detect or replace it.
    4. D.Replace EC2 status checks with CloudWatch billing alarms that notify the operations team by email whenever an instance's estimated hourly charge changes unexpectedly during a scaling event.
    Show answer & explanation

    Correct answer: B — Attach the Auto Scaling group to an Elastic Load Balancer and enable ELB health checks, so instance health is based on the load balancer's application-level checks instead of just EC2 status.

    • A. Shortening the grace period only changes how soon status checks begin being evaluated; it does not make EC2 status checks aware of an application-level hang, since they only detect underlying hardware and instance reachability issues.
    • B. Enabling ELB health checks makes the Auto Scaling group evaluate instance health based on the load balancer's application-level probe instead of just EC2 status, so a hung web server process that fails that probe is correctly detected and the instance is replaced.
    • C. Adding more instances masks the symptom by diluting the impact of a hung instance but does not detect or replace the hung instance itself, leaving it running indefinitely and continuing to receive some traffic.
    • D. CloudWatch billing alarms track cost metrics and are unrelated to instance health monitoring, so they would not help the Auto Scaling group detect or react to an application-level hang.

    Domain 3: Design High-Performing Architectures

    Subdomain 3.2: Design high-performing and elastic compute solutions

    18.A gaming company runs a self-managed NoSQL database cluster on EC2 that requires very high random-access IOPS from local disk and low storage latency, with data replicated across nodes so instance-local durability is not a concern. Which instance family best matches this requirement?

    1. A.Storage optimized instances, because they provide directly attached NVMe SSD storage tuned for high random I/O throughput and low latency, which suits a replicated NoSQL cluster.
    2. B.General purpose instances, because they balance compute, memory, and network resources evenly but do not offer the high local-disk random IOPS this database workload depends on.
    3. C.Compute optimized instances, because they maximize vCPU throughput for CPU-bound processing but provide comparatively limited local storage performance for a disk-heavy database.
    4. D.Memory optimized instances, because they provide large RAM pools for caching working sets in memory but do not specialize in high-throughput local disk I/O for a replicated cluster.
    Show answer & explanation

    Correct answer: A — Storage optimized instances, because they provide directly attached NVMe SSD storage tuned for high random I/O throughput and low latency, which suits a replicated NoSQL cluster.

    • A. Storage optimized instances include locally attached NVMe SSDs designed for very high random I/O operations per second and low latency, which is the defining requirement of this disk-heavy replicated NoSQL cluster.
    • B. General purpose instances offer a balanced mix of resources for varied workloads, but their local storage throughput does not match the high random IOPS a storage optimized family delivers for this database.
    • C. Compute optimized instances prioritize sustained CPU throughput over local storage performance, so they would leave the database's disk I/O requirement as the limiting factor.
    • D. Memory optimized instances add RAM capacity for caching but the requirement here is fast, high-volume local disk access, which memory optimized instances are not specifically built to provide.

    Subdomain 3.1: Determine high-performing and/or scalable storage solutions

    19.A media analytics company ingests several petabytes of clickstream and video metadata files daily from multiple regional data centers. Analysts run ad hoc big data queries across the entire historical dataset using Amazon Athena and Amazon EMR, and the object count is expected to grow past a billion objects within two years. Which storage solution meets these scalability and access requirements with the least operational overhead?

    1. A.Amazon S3 automatically scales storage capacity and object count into the billions while supporting many concurrent Athena and EMR readers without pre-provisioning throughput.
    2. B.Mounting an Amazon EFS General Purpose file system to EMR worker nodes lets Athena query files over NFS directly, removing the need for any object storage layer.
    3. C.Attaching Provisioned IOPS SSD (io2) EBS volumes to an EC2 fleet gives Athena a block device to read from, maximizing per-volume query throughput for this workload.
    4. D.Attaching Throughput Optimized HDD (st1) EBS volumes to the EMR master node is designed for exactly this kind of large-scale streaming analytics access pattern for big data workloads.
    Show answer & explanation

    Correct answer: A — Amazon S3 automatically scales storage capacity and object count into the billions while supporting many concurrent Athena and EMR readers without pre-provisioning throughput.

    • A. Amazon S3 is purpose-built object storage that scales capacity and object count virtually without limit, and both Athena and EMR read from S3 natively as their primary data source, so no throughput needs to be pre-provisioned.
    • B. Athena is a serverless SQL engine that queries data registered in the AWS Glue Data Catalog against Amazon S3; it does not read files over an NFS mount, so pointing EMR at an EFS file system does not give Athena access to that data.
    • C. EBS volumes attach to a single EC2 instance (aside from Multi-Attach), so they cannot serve as a shared, queryable data store for Athena, and this approach adds instance management overhead the scenario is trying to avoid.
    • D. st1 volumes are block storage tied to one instance and Availability Zone; they are not a shared analytics data store, and EMR clusters read their source data from Amazon S3 through EMRFS rather than from an attached HDD volume.

    Subdomain 3.3: Determine high-performing database solutions

    20.A mobile app backend uses Amazon DynamoDB, and traffic is highly unpredictable — daily active users can spike 20 times over during a viral event with no advance warning. The company wants the table to handle sudden bursts without throttling and without a team manually forecasting or adjusting capacity. Which DynamoDB configuration should be used?

    1. A.Configure the table with on-demand capacity mode, which instantly accommodates request-rate increases and bills per request without requiring any capacity planning.
    2. B.Configure the table with provisioned capacity mode and enable auto scaling, which raises capacity units gradually as CloudWatch alarms detect sustained utilization increases.
    3. C.Configure the table with provisioned capacity mode set to the highest historical peak, so capacity always exceeds any traffic the application has previously observed.
    4. D.Migrate the table's data to Amazon RDS with Provisioned IOPS storage, which allocates a fixed, guaranteed number of I/O operations per second for the workload.
    Show answer & explanation

    Correct answer: A — Configure the table with on-demand capacity mode, which instantly accommodates request-rate increases and bills per request without requiring any capacity planning.

    • A. On-demand capacity mode scales DynamoDB throughput instantly to match incoming request rates and charges per request, with no capacity planning needed. This directly matches an unpredictable workload with sudden, unforecastable spikes.
    • B. Provisioned capacity with auto scaling reacts to CloudWatch alarms that require sustained utilization over several minutes before adding capacity, so a sudden 20-times spike can throttle requests before capacity catches up. It is built for gradually changing load, not instantaneous bursts.
    • C. Sizing provisioned capacity to the highest historical peak still assumes the future spike will not exceed past observations, which contradicts a viral event with no advance warning, and it pays for unused capacity the rest of the time. It does not eliminate the need for forecasting.
    • D. Migrating to a relational engine with fixed Provisioned IOPS storage replaces one capacity-planning problem with another, since the storage IOPS value must still be chosen and does not scale itself. It also requires re-architecting the application's data model away from DynamoDB.

    Subdomain 3.4: Determine high-performing and/or scalable network architectures

    21.A SaaS company hosts separate customer-facing web applications at `tenant-a.example.com` and `tenant-b.example.com`, both served by identical fleets of EC2 instances behind the same load balancer. Traffic for each domain must be routed to that tenant's own target group without deploying a separate load balancer per customer. What should the architect configure?

    1. A.An Application Load Balancer listener with host-based routing rules that match the Host header and forward it to the matching tenant's target group
    2. B.A Network Load Balancer listener bound to a single port that forwards every incoming connection to one shared target group holding both tenants' instances
    3. C.A separate Auto Scaling group per tenant with no load balancer, relying on Route 53 weighted records to split traffic evenly between the two fleets
    4. D.An Application Load Balancer with one target group and a Lambda function invoked per request to manually redirect traffic based on the domain name
    Show answer & explanation

    Correct answer: A — An Application Load Balancer listener with host-based routing rules that match the Host header and forward it to the matching tenant's target group

    • A. This is correct because an Application Load Balancer can inspect the HTTP Host header and apply host-based listener rules, forwarding requests for each domain to that tenant's own dedicated target group from a single load balancer.
    • B. A Network Load Balancer routes at the connection level and cannot read the HTTP Host header, so it cannot separate tenant-a and tenant-b traffic into different target groups on its own.
    • C. Removing the load balancer and relying on weighted Route 53 records splits traffic by a fixed weight, not by which tenant domain was requested, so it does not guarantee each tenant only reaches its own fleet.
    • D. Adding a Lambda function to manually redirect every request duplicates work the Application Load Balancer's built-in host-based routing already performs natively, adding latency and operational overhead for no benefit.

    Subdomain 3.2: Design high-performing and elastic compute solutions

    22.A web application runs behind an Application Load Balancer on an Amazon EC2 Auto Scaling group. Traffic is unpredictable throughout the day, and the operations team wants the group to add or remove instances automatically so that average CPU utilization across the group stays close to 50%, without the team having to define scaling thresholds or step adjustments manually. Which Auto Scaling configuration meets this requirement?

    1. A.Configure a scheduled scaling policy that increases the desired capacity of the Auto Scaling group at fixed times each day based on historical traffic
    2. B.Configure a step scaling policy with multiple CPU utilization thresholds, each mapped to a specific number of instances to add or remove
    3. C.Configure a simple scaling policy that adds one instance whenever a CloudWatch alarm for CPU utilization above 50% enters the alarm state
    4. D.Configure a target tracking scaling policy on the Auto Scaling group using the average CPU utilization metric with a target value of 50%
    Show answer & explanation

    Correct answer: D — Configure a target tracking scaling policy on the Auto Scaling group using the average CPU utilization metric with a target value of 50%

    • A. Scheduled scaling adjusts desired capacity at predetermined times derived from historical patterns, which does not react to real-time unpredictable CPU load and still requires the team to define the schedule and capacity values manually.
    • B. Step scaling requires the team to manually define multiple thresholds and corresponding capacity adjustments, which is exactly the manual threshold definition the operations team wants to avoid, unlike a target tracking policy.
    • C. Simple scaling reacts to a single alarm with a fixed adjustment and then waits out a cooldown period before evaluating again, requiring manual threshold and increment configuration rather than continuously tracking a target metric value.
    • D. A target tracking policy is correct because the operator only specifies a target value for a metric such as average CPU utilization, and Amazon EC2 Auto Scaling automatically creates and manages the CloudWatch alarms and capacity adjustments needed to keep the metric near that target.

    Subdomain 3.3: Determine high-performing database solutions

    23.A company is choosing a database engine for a new internal application and wants to understand a key operational difference between Amazon Aurora and standard Amazon RDS for MySQL before deciding. Which statement correctly describes a distinguishing characteristic of Amazon Aurora compared with standard Amazon RDS for MySQL?

    1. A.Aurora storage automatically scales up in increments as data grows, up to a much larger maximum cluster volume size than a standard RDS for MySQL instance.
    2. B.Aurora requires manual provisioning of fixed storage volumes up front, while standard RDS for MySQL storage scales automatically without limit.
    3. C.Aurora only supports NoSQL key-value access patterns, while standard RDS for MySQL supports full relational SQL querying and joins.
    4. D.Aurora cannot be deployed across multiple Availability Zones, while standard RDS for MySQL Multi-AZ deployments span up to six Availability Zones.
    Show answer & explanation

    Correct answer: A — Aurora storage automatically scales up in increments as data grows, up to a much larger maximum cluster volume size than a standard RDS for MySQL instance.

    • A. Aurora's storage layer is a distributed, auto-scaling volume that grows automatically in increments as data is written, up to a maximum cluster storage size far larger than what a single standard RDS for MySQL instance's provisioned storage supports, which is a well-documented distinguishing characteristic between the two.
    • B. This reverses the actual behavior: Aurora is the engine that scales storage automatically, while standard RDS for MySQL requires you to provision a storage amount up front and does not scale without limit, so the roles described here are backwards.
    • C. Aurora is a relational database engine that is compatible with MySQL and PostgreSQL and supports the same full SQL querying and joins as those engines; it is not a NoSQL key-value store, so this statement misdescribes Aurora's data model entirely.
    • D. Aurora clusters replicate storage across multiple Availability Zones by design, and it is standard RDS for MySQL Multi-AZ that provides one standby in a second AZ, not six, so both halves of this statement are incorrect.

    Subdomain 3.4: Determine high-performing and/or scalable network architectures

    24.An application running in private subnets of a VPC needs to call the Amazon S3 API to read configuration objects. Security policy prohibits the instances from having public IP addresses or routing any traffic through an internet gateway or NAT device, but the calls must still reach the S3 API over AWS's private network. Which solution meets this requirement?

    1. A.Create a gateway VPC endpoint for Amazon S3 and add a route to it in the private subnet's route table so S3 API calls stay on the AWS network.
    2. B.Attach a NAT gateway in a public subnet and update the private route table to send 0.0.0.0/0 traffic through it so instances can reach the public S3 endpoint.
    3. C.Assign each instance an Elastic IP address and attach an internet gateway to the VPC so the instances can reach the public S3 API endpoint directly.
    4. D.Create a Site-to-Site VPN connection between the VPC and the on-premises data center and route S3 API traffic out through the customer gateway device.
    Show answer & explanation

    Correct answer: A — Create a gateway VPC endpoint for Amazon S3 and add a route to it in the private subnet's route table so S3 API calls stay on the AWS network.

    • A. A gateway VPC endpoint for S3 is correct because it adds a prefix-list route that keeps S3 API traffic entirely on the AWS private network, reaching the service without an internet gateway, NAT device, or public IP address on the instances.
    • B. A NAT gateway still sends traffic out toward the internet-routable S3 endpoint through an internet gateway, which violates the requirement to avoid internet gateways and NAT devices even though the instances themselves stay private.
    • C. Assigning Elastic IP addresses and attaching an internet gateway directly contradicts the stated policy against public IP addresses and internet gateway routing for these instances.
    • D. A Site-to-Site VPN connects the VPC to an on-premises network over IPsec tunnels and has no role in reaching the AWS-hosted S3 API, so it does not provide a path to the service at all.

    Subdomain 3.5: Determine high-performing data ingestion and transformation solutions

    25.A manufacturing company needs to perform a one-time migration of 80 TB of historical sensor log files from an on-premises NFS file server into Amazon S3, followed by an incremental sync of only newly modified files every night. Bandwidth is limited, so the transfer should minimize redundant data movement and verify data integrity automatically. Which service best fits this requirement?

    1. A.AWS DataSync configured with a task that transfers the NFS share to S3 and is scheduled to run nightly, moving only files that have changed since the previous run
    2. B.AWS Storage Gateway deployed as a File Gateway that presents an NFS mount to the on-premises server, continuously caching every accessed file locally for low-latency reads
    3. C.Amazon Kinesis Data Firehose configured with a custom HTTP endpoint on the file server that streams each modified file as an individual record into an S3 bucket
    4. D.AWS Glue configured with a JDBC connection to the NFS file server, running a nightly crawler job that copies any newly detected files into the target S3 bucket
    Show answer & explanation

    Correct answer: A — AWS DataSync configured with a task that transfers the NFS share to S3 and is scheduled to run nightly, moving only files that have changed since the previous run

    • A. DataSync is purpose-built for online, automated bulk transfer between on-premises NFS/SMB storage and AWS, with built-in incremental scanning that copies only changed files and automatic data integrity verification, matching both the one-time migration and nightly sync needs.
    • B. File Gateway is designed for ongoing, continuous local file access backed by S3 with local caching, not for a scheduled bulk migration followed by incremental nightly transfers of changed files, so it does not fit this transfer pattern.
    • C. Firehose delivers streaming records to destinations like S3 and does not have a mechanism to connect to an on-premises NFS server or transfer whole files as part of a file-migration workflow.
    • D. Glue connects to structured data sources such as JDBC databases for cataloging and ETL; it does not provide file-level transfer from an NFS server, and crawlers do not copy files into S3 on a schedule.

    Subdomain 3.1: Determine high-performing and/or scalable storage solutions

    26.A team is deploying a two-node clustered database on Amazon EC2 that requires both nodes to have simultaneous block-level read and write access to the same underlying volume using Amazon EBS Multi-Attach. Select TWO statements that correctly describe how Multi-Attach must be used in this design.(Select 2)

    1. A.Multi-Attach is available only on Provisioned IOPS SSD volumes — io1 or io2 — and is never supported on gp3, st1, or sc1 volumes.
    2. B.Enabling Multi-Attach installs a lock manager on the volume so the two operating systems never write to the same block simultaneously.
    3. C.The cluster software must coordinate writes with a cluster-aware file system, since Amazon EBS itself never arbitrates conflicting I/O.
    4. D.A Multi-Attach volume can be shared by EC2 instances placed in different Availability Zones as long as both instances sit in the same AWS Region.
    5. E.Turning on Multi-Attach doubles the volume's provisioned IOPS so each attached instance independently receives the full baseline performance.
    Show answer & explanation

    Correct answers: A, C — Multi-Attach is available only on Provisioned IOPS SSD volumes — io1 or io2 — and is never supported on gp3, st1, or sc1 volumes.; The cluster software must coordinate writes with a cluster-aware file system, since Amazon EBS itself never arbitrates conflicting I/O.

    • A. Multi-Attach is a capability of the Provisioned IOPS SSD family — io1 and io2 — and cannot be enabled on gp3 or on either HDD-backed volume type, which limits the volume choice for this cluster design.
    • B. Amazon EBS does not install or manage any locking mechanism on a Multi-Attach volume; it simply allows multiple instances to send I/O to the same volume, leaving conflict avoidance entirely to the attached operating systems or application.
    • C. Because EBS performs no write arbitration between attached instances, the cluster must run a cluster-aware file system or equivalent coordination layer so both nodes never corrupt each other's writes.
    • D. A Multi-Attach volume can only be attached to instances within the same Availability Zone as the volume, so this cluster's two nodes must be placed in that single zone rather than spread across zones.
    • E. Provisioned IOPS on a Multi-Attach volume is a fixed property of the volume itself and is shared across all attached instances; attaching more instances does not add or duplicate the volume's IOPS allocation.

    Subdomain 3.5: Determine high-performing data ingestion and transformation solutions

    27.Which statement accurately describes Amazon EMR's role in processing large-scale datasets as part of a data ingestion and transformation architecture?

    1. A.Amazon EMR is a fully managed NoSQL database service that automatically partitions large tables across nodes and exposes a key-value API for high-throughput reads and writes at single-digit millisecond latencies.
    2. B.Amazon EMR is a serverless SQL query engine that scans files stored directly in Amazon S3 and charges strictly per gigabyte of data scanned, with no cluster nodes for the team to configure or manage.
    3. C.Amazon EMR is a hybrid storage gateway appliance that caches on-premises file shares locally and asynchronously replicates the cached content into Amazon S3 for durable long-term retention and backup.
    4. D.Amazon EMR provisions managed clusters running open-source frameworks such as Apache Spark and Hadoop, so teams reuse existing big-data code while EMR handles provisioning, scaling, and Spot Instance integration.
    Show answer & explanation

    Correct answer: D — Amazon EMR provisions managed clusters running open-source frameworks such as Apache Spark and Hadoop, so teams reuse existing big-data code while EMR handles provisioning, scaling, and Spot Instance integration.

    • A. Automatic partitioning across nodes with a key-value API at single-digit millisecond latency describes a managed NoSQL database, not a big-data processing cluster service.
    • B. A serverless engine that scans S3 files with no cluster to manage and bills per gigabyte scanned describes an interactive query service, not a managed cluster platform.
    • C. Caching on-premises file shares locally and replicating them into object storage describes a hybrid storage gateway appliance, not a big-data cluster processing service.
    • D. This is correct: this service provisions and manages clusters that run familiar open-source frameworks, so teams reuse existing code while the service handles scaling and can mix in Spot Instances to lower cost.

    Domain 4: Design Cost-Optimized Architectures

    Subdomain 4.4: Design cost-optimized network architectures

    28.A manufacturing company needs a steady, predictable connection between its data center and AWS to move several terabytes of sensor data every day at consistent throughput, and it can tolerate a multi-week lead time to provision the link. Cost per gigabyte transferred matters more than setup speed. Which connectivity option best fits this requirement?

    1. A.Provision two Site-to-Site VPN connections in an active-active configuration to double available bandwidth to match a dedicated line's throughput.
    2. B.Provision an AWS Direct Connect connection, which offers lower data transfer rates and more consistent throughput for sustained high-volume traffic than a VPN.
    3. C.Provision a Site-to-Site VPN connection, which can be established within hours and offers the lowest per-gigabyte data transfer rate of any connectivity option.
    4. D.Route all sensor data over the public internet directly to a VPC endpoint, since internet paths offer the same throughput consistency as a dedicated line.
    Show answer & explanation

    Correct answer: B — Provision an AWS Direct Connect connection, which offers lower data transfer rates and more consistent throughput for sustained high-volume traffic than a VPN.

    • A. Running two VPN tunnels adds redundancy and some aggregate throughput, but each tunnel is still capped well below dedicated-line speeds and both remain subject to public internet variability.
    • B. Direct Connect provides a dedicated physical link with consistent throughput and lower data transfer rates than internet-based paths, which fits a steady, high-volume workload where setup lead time is acceptable.
    • C. A Site-to-Site VPN can be provisioned quickly, but it runs over the public internet, so its throughput is variable and its per-gigabyte data transfer pricing is not lower than a dedicated Direct Connect connection.
    • D. Routing sensor data over the public internet is subject to variable latency and throughput because it competes with other internet traffic, which does not match a requirement for consistent, predictable performance.

    Subdomain 4.1: Design cost-optimized storage solutions

    29.An enterprise runs a legacy on-premises file server that departments access over SMB every day, and it is running out of local disk space. The company wants to keep the existing SMB file share experience for end users unchanged while moving the bulk of the data to lower-cost cloud storage, caching only the most recently used files locally. Which service should the company deploy?

    1. A.AWS Storage Gateway configured as a File Gateway, because it presents an SMB share backed by S3 and caches frequently accessed files locally while the rest lives in the cloud.
    2. B.AWS DataSync deployed as a scheduled agent task, because DataSync exposes an SMB endpoint on-premises and transparently redirects file reads and writes to an S3 bucket.
    3. C.AWS Transfer Family configured with an SMB endpoint, because Transfer Family lets on-premises Windows clients continue using SMB paths while files are actually stored in S3.
    4. D.Amazon FSx for Lustre deployed on-premises, because Lustre file systems natively speak SMB and were designed as a drop-in replacement for aging Windows file servers.
    Show answer & explanation

    Correct answer: A — AWS Storage Gateway configured as a File Gateway, because it presents an SMB share backed by S3 and caches frequently accessed files locally while the rest lives in the cloud.

    • A. A Storage Gateway File Gateway presents a local SMB or NFS share while storing the actual data as objects in S3, caching the most recently used data on local storage, which matches the exact use case of extending an on-premises file share into S3 without changing the end-user experience.
    • B. AWS DataSync is a data transfer service for moving files between on-premises storage and AWS storage on a schedule; it does not present a live SMB share endpoint that end users can browse day to day.
    • C. AWS Transfer Family provides managed SFTP, FTPS, and FTP endpoints for partner file exchange; it does not expose an SMB protocol endpoint and is not designed to extend an internal Windows file share.
    • D. Amazon FSx for Lustre is a high-performance file system built for compute-intensive workloads such as machine learning and HPC, is not deployed on-premises, and does not support the SMB protocol.

    Subdomain 4.2: Design cost-optimized compute solutions

    30.A finance department wants to see AWS spend broken out by engineering team and by project, across all accounts in the organization, so they can charge each team's budget accurately. What should be configured to make this breakdown possible?

    1. A.Apply consistent cost allocation tags, such as team and project keys, to resources and activate those tags in AWS Cost Explorer so costs can be filtered and grouped by tag value.
    2. B.Enable AWS Trusted Advisor cost checks across every account, since its recommendations automatically group all historical spend by team and project without any tagging work required.
    3. C.Move every team into its own AWS Region, since AWS bills usage separately per Region and that separation is the mechanism used to attribute cost to a specific team or project.
    4. D.Turn on detailed CloudWatch monitoring for every resource, since per-minute metrics let Cost Explorer infer which team or project generated a given charge after the fact.
    Show answer & explanation

    Correct answer: A — Apply consistent cost allocation tags, such as team and project keys, to resources and activate those tags in AWS Cost Explorer so costs can be filtered and grouped by tag value.

    • A. Cost allocation tags let teams attach key-value metadata such as team or project to resources, and once activated in the billing console those tags become dimensions that Cost Explorer can filter and group cost data by.
    • B. Trusted Advisor surfaces cost optimization checks like idle resources and underutilized instances; it does not attribute historical spend to an organizational team or project the way tag-based cost allocation does.
    • C. Splitting teams across Regions changes where resources run and adds unrelated latency and data-residency considerations, but Region alone is not a supported way to attribute cost to a team or project in the billing tools.
    • D. Detailed monitoring increases the frequency of CloudWatch metric data points for troubleshooting performance, but metrics carry no team or project attribution and cannot substitute for cost allocation tags in Cost Explorer.

    Subdomain 4.3: Design cost-optimized database solutions

    31.An e-commerce site's product catalog page issues the same handful of read queries against Amazon RDS for MySQL millions of times per day, since most shoppers browse the same popular items. To reduce load-driven database costs without changing the underlying relational schema, which approach should the team add in front of the database?

    1. A.Add an Amazon ElastiCache (Redis) layer that caches the frequent catalog query results so most reads never reach the database.
    2. B.Add more RDS read replicas so the repeated catalog queries can be spread out across several database instances instead of just one.
    3. C.Enable RDS Performance Insights so the team can see which catalog queries run most often and manually optimize their SQL.
    4. D.Switch the catalog table's storage engine settings to increase the buffer pool size available for caching data in memory.
    Show answer & explanation

    Correct answer: A — Add an Amazon ElastiCache (Redis) layer that caches the frequent catalog query results so most reads never reach the database.

    • A. An in-memory cache like ElastiCache Redis stores the results of the repeated catalog queries so subsequent identical requests are served from cache instead of hitting the database, cutting both load and the read capacity the team must pay for.
    • B. Adding more read replicas still sends every repeated query to a database instance and multiplies the cost of provisioned compute, rather than eliminating the redundant reads entirely.
    • C. Performance Insights is a monitoring tool that helps identify slow or frequent queries, but it does not itself reduce the number of times those queries hit the database.
    • D. Increasing the buffer pool size can speed up reads that still reach the instance, but it does not stop the same query from being re-executed against the database engine each time a shopper loads the page.

    Subdomain 4.1: Design cost-optimized storage solutions

    32.A media platform ingests thousands of new video thumbnail images every day. Some thumbnails go viral and are requested constantly for weeks, while others are never viewed again after the first few days, and the platform cannot predict in advance which pattern any given thumbnail will follow. The platform wants storage costs to adjust automatically to actual access patterns without engineers writing or maintaining lifecycle rules. Which storage approach best fits this requirement?

    1. A.Store thumbnails in S3 Intelligent-Tiering, which monitors access at the object level and moves each one between tiers without retrieval fees
    2. B.Store thumbnails in S3 Standard-IA, which lowers the per-GB rate for any thumbnail that is not requested within the first thirty days of upload
    3. C.Store thumbnails in S3 One Zone-IA, which reduces storage cost by keeping a single copy in one Availability Zone for infrequently viewed files
    4. D.Write a scheduled Lambda function that inspects CloudWatch access metrics nightly and moves each thumbnail to a cheaper class based on the results
    Show answer & explanation

    Correct answer: A — Store thumbnails in S3 Intelligent-Tiering, which monitors access at the object level and moves each one between tiers without retrieval fees

    • A. S3 Intelligent-Tiering automatically monitors each object's access pattern and moves it between frequent- and infrequent-access tiers with no retrieval fees, which directly matches an unpredictable, per-object access pattern with no engineering upkeep.
    • B. S3 Standard-IA applies one fixed class to every thumbnail regardless of whether it later goes viral, so a thumbnail that becomes popular after the 30-day mark still incurs per-GB retrieval fees on every frequent request, which does not adapt automatically.
    • C. S3 One Zone-IA is a single fixed class with single-AZ resilience trade-offs, and like Standard-IA it does not adjust itself when a thumbnail's popularity changes over time, so it fails the automatic-adaptation requirement.
    • D. A custom scheduled function that inspects metrics and moves objects would require engineers to build and maintain the exact lifecycle logic the platform explicitly wants to avoid, and it reintroduces the operational overhead the built-in tiering service is meant to remove.

    Subdomain 4.2: Design cost-optimized compute solutions

    33.A data science team runs a single large EC2 instance that loads a multi-gigabyte dataset into memory and builds an in-memory index that takes about 40 minutes to rebuild from scratch. The team only needs the instance during business hours and wants to avoid paying for compute overnight, but rebuilding the in-memory index every morning is too slow for their workflow. Which action lets the team stop paying for compute overnight while preserving the in-memory state for a fast restart the next morning?

    1. A.Stop the instance at the end of each business day, which halts billing for compute while the operating system automatically reloads the in-memory index from an EBS snapshot on the next boot.
    2. B.Hibernate the instance at the end of each business day, which persists the contents of RAM to the root EBS volume and restores that memory state when the instance resumes.
    3. C.Reboot the instance at the end of each business day, which restarts the operating system but keeps the underlying EC2 host and hourly compute billing running overnight.
    4. D.Terminate the instance at the end of each business day and launch a fresh instance each morning from an AMI that has the in-memory index baked into its root volume.
    Show answer & explanation

    Correct answer: B — Hibernate the instance at the end of each business day, which persists the contents of RAM to the root EBS volume and restores that memory state when the instance resumes.

    • A. Stopping the instance does halt compute billing, but there is no mechanism that reloads RAM contents from an EBS snapshot on boot; the operating system starts fresh and the 40-minute index rebuild would still be required each morning.
    • B. Hibernating the instance is correct because AWS copies the in-memory RAM contents to the root EBS volume at hibernation and restores that exact memory state on resume, so billing stops overnight while the rebuilt index survives.
    • C. Rebooting restarts the operating system but leaves the underlying EC2 host allocated and billed for compute the entire time, so it does not achieve the goal of avoiding overnight charges.
    • D. An AMI captures a snapshot of the root volume at the moment it was created, not the live in-memory state at the end of each business day, so relaunching from it would not reflect that day's rebuilt index.

    Subdomain 4.3: Design cost-optimized database solutions

    34.A mobile gaming company launches a new DynamoDB table to store player session state. Traffic is highly unpredictable, with sudden 50x spikes during viral moments and near-zero traffic overnight, and the team wants to avoid manual capacity planning while paying only for the requests the application actually makes. Which capacity mode should they configure for the table?

    1. A.Configure provisioned capacity with a fixed high RCU and WCU ceiling sized for the largest expected spike so throughput is always guaranteed.
    2. B.Configure on-demand capacity mode so the table is billed per read and write request and automatically absorbs sudden traffic spikes.
    3. C.Configure provisioned capacity with application auto scaling enabled so throughput adjusts gradually between a minimum and maximum you define.
    4. D.Add DynamoDB Accelerator as an in-memory cache layer in front of a fixed-capacity table to absorb the overnight and spike traffic.
    Show answer & explanation

    Correct answer: B — Configure on-demand capacity mode so the table is billed per read and write request and automatically absorbs sudden traffic spikes.

    • A. Sizing a fixed provisioned ceiling for the worst-case 50x spike means paying for that peak capacity around the clock, including the near-zero overnight hours, which is the opposite of the pay-per-use goal.
    • B. On-demand capacity mode bills per read and write request and scales automatically to handle sudden surges without any capacity planning, matching an unpredictable, spiky traffic pattern at the lowest ongoing cost.
    • C. Auto scaling adjusts provisioned throughput gradually based on utilization targets over minutes, so it lags behind a sudden 50x spike and still requires the team to plan and tune minimum and maximum thresholds.
    • D. An accelerator cache reduces read latency and load on the table but does not change how write capacity is billed, and it adds an extra cluster to manage and pay for on top of the underlying capacity mode problem.

    Subdomain 4.4: Design cost-optimized network architectures

    35.A logistics company connects a branch office to its AWS environment using a single AWS Site-to-Site VPN connection. Traffic has grown to the point where the single connection's per-tunnel throughput ceiling is regularly saturated during business hours, and the company wants to increase available bandwidth quickly without the multi-week procurement and installation process that a new dedicated line would require. Which approach increases available bandwidth while keeping the solution VPN-based?

    1. A.Attach the VPN connection to a transit gateway and establish multiple Site-to-Site VPN connections from the branch office, using equal-cost multi-path routing across the added tunnels.
    2. B.Request an increase to the branch office's public internet circuit speed only, since a Site-to-Site VPN connection's throughput is limited exclusively by the customer gateway's uplink.
    3. C.Replace the Site-to-Site VPN with a VPC peering connection between the branch office network and the VPC to remove the per-tunnel throughput ceiling entirely.
    4. D.Migrate the branch office connection to AWS PrivateLink so branch office traffic reaches the VPC through an interface endpoint instead of a VPN tunnel.
    Show answer & explanation

    Correct answer: A — Attach the VPN connection to a transit gateway and establish multiple Site-to-Site VPN connections from the branch office, using equal-cost multi-path routing across the added tunnels.

    • A. Attaching multiple VPN connections to a transit gateway and spreading traffic across the resulting tunnels with equal-cost multi-path routing adds bandwidth beyond a single tunnel's ceiling in days rather than the weeks a new dedicated circuit typically needs.
    • B. Each Site-to-Site VPN tunnel has its own throughput ceiling that is independent of the customer gateway's raw internet uplink speed, so increasing the branch office's internet circuit alone does not raise the per-tunnel limit that is being saturated.
    • C. VPC peering connects two VPCs directly and has no mechanism for terminating an on-premises branch office network connection, so it cannot take the place of a Site-to-Site VPN in this scenario.
    • D. PrivateLink exposes a specific service through an interface endpoint inside a VPC; it is not a way to terminate a branch office network connection or carry general branch-to-VPC traffic.

    Want the full experience?

    These are just samples. Practice the full AWS Certified Solutions Architect - Associate (SAA-C03) question bank in quiz mode — free, no signup, with domain practice and exam simulation.