CertSafari

    Free Google Professional Cloud Architect Sample Questions

    35 free sample questions from our bank of 350+, covering every exam domain, with answers and detailed explanations. Updated September 2026.

    Domain 1: Designing and planning a cloud solution architecture

    Subdomain 1.1: Designing a cloud solution infrastructure that meets business requirements.

    1.An enterprise architecture team is evaluating whether to build a custom internal expense-reporting system on Google Cloud or buy a SaaS expense-management product, as part of a workload disposition review. Which two factors should most strongly influence a decision to buy rather than build?(Select 2)

    1. A.The capability is not part of the company's core competitive differentiation, so engineering effort is better spent on revenue-generating features.
    2. B.A mature SaaS vendor already meets the required compliance certifications and integrates with the existing HR and finance systems out of the box.
    3. C.The internal team has already written several proprietary services in the same programming language used by the candidate SaaS vendor's API client library.
    4. D.Google Cloud offers a managed database service that could host a custom application's schema with minimal operational overhead.
    5. E.The company's engineers have attended training on Compute Engine and are comfortable provisioning virtual machines for internal tools.
    Show answer & explanation

    Correct answers: A, B — The capability is not part of the company's core competitive differentiation, so engineering effort is better spent on revenue-generating features.; A mature SaaS vendor already meets the required compliance certifications and integrates with the existing HR and finance systems out of the box.

    • A. Expense reporting is a back-office capability that does not differentiate the company competitively, which is a classic signal to buy so engineering time goes toward features that move the business forward. This is one of the two factors that should tip the disposition decision toward buying.
    • B. A vendor that already carries the needed compliance certifications and integrates with existing HR and finance systems removes both compliance risk and integration effort that a custom build would otherwise require. This directly supports buying over building and is the other deciding factor.
    • C. Language familiarity with an API client library is a minor implementation convenience and does not address whether the capability is strategic or how much compliance and integration work a custom build would require. It is not a strong driver of the build-versus-buy decision.
    • D. The availability of a managed database service lowers the operational cost of building, which would argue for building rather than buying. It does not, by itself, make the case that buying is the better disposition for this capability.
    • E. Engineer comfort with Compute Engine reflects existing build capability rather than a reason to prefer a purchased solution. Team skill availability supports either disposition and is not decisive for choosing to buy.

    Subdomain 1.1: Designing a cloud solution infrastructure that meets business requirements.

    2.Case study: TerramEarth. The telemetry-ingestion pipeline that receives equipment data from hundreds of thousands of vehicles is central to the company's predictive-maintenance product, and the architecture team is defining business continuity requirements so a regional outage does not silently drop incoming telemetry. Which two design decisions correctly support this continuity requirement?(Select 2)

    1. A.Buffer incoming telemetry in a durable, replicated messaging service so data persists and can be reprocessed if a downstream consumer becomes unavailable.
    2. B.Deploy the ingestion pipeline across multiple regions with automated failover, so vehicle telemetry keeps being accepted if one region has an outage.
    3. C.Accept telemetry only in a single region and instruct vehicles to discard any readings that fail to deliver during an outage, keeping the pipeline simple.
    4. D.Store incoming telemetry only in each ingestion server's local memory, since durability can be added later once the predictive-maintenance product matures.
    5. E.Have field technicians manually collect vehicle telemetry from onboard storage during any outage and upload it once the pipeline recovers.
    Show answer & explanation

    Correct answers: A, B — Buffer incoming telemetry in a durable, replicated messaging service so data persists and can be reprocessed if a downstream consumer becomes unavailable.; Deploy the ingestion pipeline across multiple regions with automated failover, so vehicle telemetry keeps being accepted if one region has an outage.

    • A. A durable, replicated messaging layer holds incoming telemetry safely even if a downstream processing component is temporarily unavailable, ensuring the data can still be consumed once processing resumes rather than being lost.
    • B. Deploying ingestion across multiple regions with automated failover means the pipeline keeps accepting vehicle telemetry even if one region experiences an outage, which is exactly what prevents a regional failure from silently dropping data.
    • C. A single-region pipeline that discards undelivered readings during an outage is the exact failure mode the continuity requirement is meant to prevent, since telemetry would be silently and permanently lost during any regional disruption.
    • D. Keeping telemetry only in local server memory means any server restart or failure loses that data immediately, which provides no durability at all and directly conflicts with a business continuity requirement for a central data pipeline.
    • E. Relying on manual field collection during an outage does not scale to hundreds of thousands of vehicles and introduces long delays before data becomes available again, which does not meet the pipeline's continuity needs.

    Subdomain 1.3: Designing network, storage, and compute resources.

    3.Case study – TerramEarth: TerramEarth manufactures heavy agricultural and mining equipment and collects telemetry from 20 million connected vehicles; it runs a nightly batch job that reprocesses a full day of historical sensor data into aggregated reports, and the job can restart from its last checkpoint if interrupted. The data engineering team wants to minimize compute cost for this fault-tolerant, non-customer-facing job, which can tolerate individual worker restarts. Which Compute Engine VM provisioning option should they use for the worker pool?

    1. A.Spot VMs, accepting that Compute Engine can reclaim capacity with a short shutdown notice in exchange for a steep discount off the on-demand price.
    2. B.Sole-tenant nodes, dedicating physical hardware to the workload so no other customer's virtual machines share the same server.
    3. C.Reserved, committed-use standard VMs, locking in a one- or three-year term to guarantee capacity and a discounted hourly rate.
    4. D.On-demand standard VMs with automatic restart enabled, paying full price so the workload is covered by the Compute Engine SLA.
    Show answer & explanation

    Correct answer: A — Spot VMs, accepting that Compute Engine can reclaim capacity with a short shutdown notice in exchange for a steep discount off the on-demand price.

    • A. Spot VMs offer the deepest discount off on-demand pricing in exchange for the possibility of preemption, and because this checkpointed batch job tolerates worker restarts, it can absorb that tradeoff to minimize cost.
    • B. Sole-tenant nodes address hardware isolation and compliance requirements, not cost minimization, and they carry a premium rather than a discount, so they do not fit a job optimizing purely for lowest compute cost.
    • C. Committed-use discounts require a one- or three-year commitment to steady capacity, which suits predictable always-on workloads rather than a nightly job that could instead exploit spare, interruptible capacity for a deeper discount.
    • D. Paying full on-demand price for SLA coverage is unnecessary here because the job is fault-tolerant and not customer-facing, so the extra cost of guaranteed uptime does not buy anything the workload needs.

    Subdomain 1.5: Envisioning future solution improvements.

    4.A company exposes a public REST API that several partner integrations depend on. The architecture team expects to introduce breaking changes to the API's response format as the product evolves over the next few years and wants existing partner integrations to keep working when that happens. Which practice best supports this future evolution?

    1. A.Introduce a new versioned API path for breaking changes, while still supporting the existing version until partners migrate.
    2. B.Change the existing API's response format in place whenever a new feature requires it, without ever adding a version indicator.
    3. C.Ask every partner to rebuild their entire integration from scratch each time the API's response format changes in any way.
    4. D.Remove the previous API version immediately whenever a new version is released, regardless of each partner's migration status.
    Show answer & explanation

    Correct answer: A — Introduce a new versioned API path for breaking changes, while still supporting the existing version until partners migrate.

    • A. Introducing a new versioned path for breaking changes while keeping the existing version available lets partner integrations keep working against the version they built for, until they choose to migrate on their own schedule.
    • B. Changing the existing response format in place without a version indicator breaks every partner integration built against the prior format the moment the change ships, with no way for them to keep working unchanged.
    • C. Requiring partners to rebuild their integration for any format change, however small, creates unnecessary partner churn and does not reflect a practice for evolving an API in a compatible way.
    • D. Removing the previous version immediately on release gives partners no migration window, which breaks their integrations rather than letting them keep working while the API evolves.

    Subdomain 1.4: Creating a migration plan (i.e., documents and architectural diagrams).

    5.During the discovery phase of a migration plan, an architect deploys the Migration Center Discovery Client into an on-premises VMware environment instead of manually uploading spreadsheets. What is the primary purpose of this component?

    1. A.It automatically collects configuration and utilization data from on-premises servers and databases and uploads it to Migration Center for assessment.
    2. B.It provisions a landing zone in Google Cloud and creates the folder and project hierarchy needed before any application workloads can be migrated there.
    3. C.It executes the actual data transfer of virtual machine disks from the on-premises hypervisor into Compute Engine persistent disks, separate from inventory collection.
    4. D.It generates the final architectural diagrams and migration wave sequencing document for stakeholder sign-off after assessment completes.
    Show answer & explanation

    Correct answer: A — It automatically collects configuration and utilization data from on-premises servers and databases and uploads it to Migration Center for assessment.

    • A. The Discovery Client is the on-premises component that automatically gathers configuration and utilization data from servers and databases and sends it to Migration Center, which is exactly the automated collection step needed before an accurate assessment can be produced. This removes the manual effort of building inventory spreadsheets by hand.
    • B. Provisioning a landing zone with a folder and project hierarchy is a resource-hierarchy design activity, not something the Discovery Client performs. That work happens separately, typically before or in parallel with discovery, using organization and folder design guidance.
    • C. Migration Center's discovery tooling collects assessment data; it does not perform the live data transfer of virtual machine disks into Compute Engine. Disk migration is handled by dedicated migration tooling once the assessment and planning phases are complete.
    • D. Architectural diagrams and wave sequencing documents are outputs an architect produces using the assessment data, not something the discovery component generates automatically. The Discovery Client's role ends once it uploads the collected inventory data for analysis.

    Subdomain 1.4: Creating a migration plan (i.e., documents and architectural diagrams).

    6.A company is planning hybrid connectivity to support a multi-year, phased migration in which some applications remain on-premises indefinitely while others move to Google Cloud in successive waves. The migration plan must recommend connectivity options that provide both the bandwidth needed for large data transfers and ongoing low-latency access for hybrid applications. Which options should the plan include? (Select all that apply)(Select 3)

    1. A.Dedicated Interconnect for high-bandwidth, low-latency private connectivity directly into the Google Cloud network, suitable for large-scale data transfer and long-term hybrid application traffic.
    2. B.Cloud VPN as a lower-bandwidth, encrypted fallback path over the public internet, useful for redundancy or for workloads that do not need Interconnect-level throughput.
    3. C.Public internet connectivity with no encryption at all for all traffic between the on-premises data center and Google Cloud, since encryption is assumed unnecessary for purely internal business applications.
    4. D.A single unencrypted static route through a home broadband connection, since migration data volumes are always small enough for consumer-grade internet links to handle reliably.
    5. E.Network Connectivity Center to manage and centralize multiple hybrid connections, such as Interconnect and VPN, into a hub-and-spoke topology as connected sites grow in number.
    Show answer & explanation

    Correct answers: A, B, E — Dedicated Interconnect for high-bandwidth, low-latency private connectivity directly into the Google Cloud network, suitable for large-scale data transfer and long-term hybrid application traffic.; Cloud VPN as a lower-bandwidth, encrypted fallback path over the public internet, useful for redundancy or for workloads that do not need Interconnect-level throughput.; Network Connectivity Center to manage and centralize multiple hybrid connections, such as Interconnect and VPN, into a hub-and-spoke topology as connected sites grow in number.

    • A. Dedicated Interconnect offers a private, dedicated, high-bandwidth path suitable for both the large data transfers involved in migration and the ongoing low-latency needs of applications that stay hybrid for the long term.
    • B. Cloud VPN provides an encrypted path over the public internet that works well as a redundant path or for workloads that do not need Interconnect-level throughput, complementing a primary Interconnect connection.
    • C. Sending all traffic over the public internet with no encryption at all ignores the security risk of exposing application data in transit, regardless of whether the traffic is described internally as business application traffic.
    • D. A single unencrypted consumer-grade broadband link has neither the reliability nor the bandwidth guarantees needed for enterprise migration data volumes or ongoing hybrid application traffic, especially at multi-year, multi-wave scale.
    • E. Network Connectivity Center centralizes multiple hybrid connections such as Interconnect and VPN into a manageable hub-and-spoke topology, which becomes increasingly valuable as the number of connected sites grows over a multi-year migration.

    Subdomain 1.2: Designing a cloud solution infrastructure that meets technical requirements.

    7.Case study — Mountkirk Games: Mountkirk Games is a mobile game studio that migrated its on-premises multiplayer backend to Google Cloud and is launching a new real-time multiplayer title expected to draw millions of concurrent players across North America, Europe, and Asia. The business requires the game to stay available even if an entire Google Cloud region becomes unreachable, and technical leadership wants gameplay telemetry captured for near-real-time analytics. Which design best meets the requirement to keep the game available if an entire Google Cloud region fails, while keeping latency low for players on three continents?

    1. A.Run autoscaled instance groups or GKE clusters behind a global external Application Load Balancer in two regions per continent, backed by a multi-region Spanner instance for consistent writes.
    2. B.Place the backend in the region closest to the largest player base and rely on Cloud CDN in front of the load balancer, caching responses so players on other continents get acceptable latency during normal play.
    3. C.Run regional external load balancers in one region per continent and replicate session state to Cloud Storage on a nightly schedule, so a failed region can be rebuilt from the most recent nightly backup file.
    4. D.Spread compute across multiple zones within one region near most players and pair it with a Cloud SQL high-availability instance whose standby replica sits in the same region to guard against zone failures only.
    Show answer & explanation

    Correct answer: A — Run autoscaled instance groups or GKE clusters behind a global external Application Load Balancer in two regions per continent, backed by a multi-region Spanner instance for consistent writes.

    • A. A global load balancer fanning out to instance groups or clusters in multiple regions removes any single region as a point of failure, and a multi-region Spanner instance keeps session data consistent even when one region drops out.
    • B. Serving from a single region with CDN caching in front reduces latency for cacheable content but does nothing for a full regional outage, since the origin backend and its state still live in one place.
    • C. Nightly replication to Cloud Storage means up to a day of session data could be lost, and manual rebuild from backups takes far longer than the near-continuous availability the studio needs during an active outage.
    • D. Multi-zone redundancy inside a single region protects against a zone failure but not a regional failure, and a same-region Cloud SQL standby replica goes down along with the primary if the whole region fails.

    Subdomain 1.2: Designing a cloud solution infrastructure that meets technical requirements.

    8.An architect is scoping a new three-tier application and wants to iterate on the design conversationally, then have the tool generate the corresponding Terraform to provision it, before writing any infrastructure code by hand. Which capability should the architect use?

    1. A.Gemini Cloud Assist's natural-language design experience in Application Design Center, which lets the architect refine the architecture in conversation, then generate Terraform.
    2. B.Cloud Monitoring's dashboard editor, which lets the architect assemble charts for the planned services only once the tier has been manually provisioned and is already producing metrics.
    3. C.The Cloud Billing budget console, which lets the architect set spending alerts for the project the new application will eventually run in, once resources exist there.
    4. D.Cloud Deploy's delivery pipeline configuration, which promotes an already-built container image through defined targets after the application has already been designed and coded.
    Show answer & explanation

    Correct answer: A — Gemini Cloud Assist's natural-language design experience in Application Design Center, which lets the architect refine the architecture in conversation, then generate Terraform.

    • A. Gemini Cloud Assist supports designing, configuring, and provisioning architectures through natural-language interaction in Application Design Center, including generating Terraform, kubectl manifests, or gcloud commands from the resulting design.
    • B. A monitoring dashboard editor visualizes metrics from resources that already exist; it has no role in conversationally designing an architecture or generating its infrastructure code.
    • C. A billing budget console tracks spend against a threshold and cannot design an architecture or produce Terraform for resources that do not exist yet.
    • D. A delivery pipeline promotes a built artifact through deployment targets; it assumes the infrastructure and application already exist rather than helping design them from scratch.

    Domain 2: Managing and provisioning a cloud solution infrastructure

    Subdomain 2.2: Configuring individual storage systems.

    9.An architect is provisioning boot and data disks for a fleet of Compute Engine VMs running a latency-sensitive OLTP database. The workload needs IOPS and throughput that can be tuned independently of the provisioned capacity, and the team wants to avoid resizing the volume every time performance requirements change. Which disk choice best satisfies this requirement?

    1. A.Provision Hyperdisk volumes and configure IOPS and throughput as separate parameters from the disk capacity
    2. B.Provision pd-balanced Persistent Disk volumes and increase the disk size whenever more IOPS or throughput is needed
    3. C.Provision Local SSD volumes attached to the VM to guarantee the highest possible IOPS for the database files
    4. D.Provision pd-standard Persistent Disk volumes and rely on automatic snapshot schedules to preserve performance headroom
    Show answer & explanation

    Correct answer: A — Provision Hyperdisk volumes and configure IOPS and throughput as separate parameters from the disk capacity

    • A. Hyperdisk decouples performance provisioning from capacity, letting an architect raise or lower IOPS and throughput independently of the volume size and without resizing the disk, which is exactly the tunable-performance requirement described.
    • B. Persistent Disk performance, including pd-balanced, scales with the provisioned capacity, so the only way to raise IOPS or throughput is to increase the disk size, which contradicts the requirement to tune performance without resizing.
    • C. Local SSD data does not survive a VM stop, suspend, or crash, so it cannot serve as durable storage for an OLTP database that must retain committed transactions, even though its raw IOPS is high.
    • D. Snapshot schedules protect against data loss but do not change a persistent disk's IOPS or throughput ceiling, and pd-standard still ties its performance to the size of the provisioned volume.

    Subdomain 2.3: Configuring compute systems.

    10.A team building a new containerized service wants a fully managed node experience with no node provisioning or patching to perform themselves, and wants billing based only on the CPU and memory their pods actually request rather than the size of statically provisioned node pools. Which two characteristics correctly describe why GKE Autopilot mode fits this team better than GKE Standard mode? (Select two.)(Select 2)

    1. A.Google manages node provisioning, upgrades, and repairs automatically, removing the need for the team to size or patch the underlying node pools
    2. B.Billing is based on the CPU, memory, and storage that pods actually request rather than the capacity of statically sized node pools the team provisions
    3. C.The team retains full SSH access to every node and can install custom kernel modules and daemonsets that need direct host-level privileges
    4. D.Autopilot lets the team choose and configure every node's machine type individually, giving the same granular control as manually managed node pools
    5. E.Autopilot requires the team to manually renew node TLS certificates and rotate kubelet credentials on a recurring schedule to keep the cluster compliant
    Show answer & explanation

    Correct answers: A, B — Google manages node provisioning, upgrades, and repairs automatically, removing the need for the team to size or patch the underlying node pools; Billing is based on the CPU, memory, and storage that pods actually request rather than the capacity of statically sized node pools the team provisions

    • A. Autopilot's defining characteristic is that Google operates node provisioning, upgrades, and repairs on the team's behalf, which is exactly the hands-off node management this team is asking for.
    • B. Autopilot bills by the resources pods actually request rather than by the size of a fixed node pool, matching the team's requirement to pay only for what their workloads consume.
    • C. Autopilot does not give users SSH access to nodes or allow privileged host-level customization such as custom kernel modules, since Google manages the node layer entirely, so this describes Standard mode's flexibility, not Autopilot's.
    • D. Autopilot does not expose per-node machine type selection to the team; node shape is managed automatically based on workload requests, so this describes the manual control available in Standard mode, not Autopilot.
    • E. Node-level credential rotation and certificate renewal are part of the node management Google performs automatically in Autopilot, so requiring the team to do this manually is not how Autopilot works.

    Subdomain 2.3: Configuring compute systems.

    11.A team is building a new event-driven microservice that receives bursts of order-confirmation events from Pub/Sub during flash sales. The service must run an existing multi-container image, keep a long-lived in-memory cache warmed at startup across requests on the same instance, and scale to zero between sales events to avoid idle cost. Which compute option best fits these requirements?

    1. A.Deploy the microservice on Cloud Run, since it runs the existing container image, can hold the in-memory cache across requests on a warm instance, and scales to zero when idle
    2. B.Deploy the microservice on Cloud Functions, since it accepts arbitrary multi-container images and preserves in-memory cache state across cold starts by design
    3. C.Deploy the microservice on App Engine standard environment, since it natively runs arbitrary multi-container images without any repackaging of the existing service
    4. D.Deploy the microservice on a single always-on Compute Engine VM, since that is the only Google Cloud compute option that supports scale-to-zero behavior for this bursty flash-sale traffic
    Show answer & explanation

    Correct answer: A — Deploy the microservice on Cloud Run, since it runs the existing container image, can hold the in-memory cache across requests on a warm instance, and scales to zero when idle

    • A. Cloud Run runs an existing container image directly, keeps an instance warm to serve subsequent requests so an in-memory cache persists between them, and scales instance count down to zero when there is no traffic, matching all three requirements.
    • B. Cloud Functions is built around single-purpose functions rather than an existing multi-container image, and a cold start replaces the execution environment, so in-memory state is not reliably preserved the way this workload needs.
    • C. App Engine standard environment runs applications in Google-managed language runtimes rather than arbitrary multi-container images, so the existing container would need to be repackaged rather than deployed as-is.
    • D. A single always-on Compute Engine VM keeps running and billing continuously; it does not scale to zero, and Cloud Run is not the only Google Cloud compute option with scale-to-zero behavior, so this description is inaccurate on both points.

    Subdomain 2.1: Configuring network topologies.

    12.An architect is designing firewall rules for a VPC where instances are frequently recreated by autoscaling groups and where engineers occasionally need to relabel workloads without security review. The design must ensure that a compromised engineer credential cannot silently widen network access by relabeling an instance. Which firewall targeting approach best meets this goal?

    1. A.Target firewall rules using the service account attached to each instance, since changing an instance's service account requires an IAM permission that can be tightly restricted.
    2. B.Target firewall rules using network tags applied to each instance, since tags can be added quickly during incidents without waiting for an IAM role change to propagate.
    3. C.Target firewall rules using the subnet each instance is created in, since every workload in a given subnet automatically receives the same intended access level.
    4. D.Target firewall rules using the external IP address range assigned to each instance, since IP-based rules do not depend on any instance metadata that an engineer could quickly edit or relabel.
    Show answer & explanation

    Correct answer: A — Target firewall rules using the service account attached to each instance, since changing an instance's service account requires an IAM permission that can be tightly restricted.

    • A. Service accounts are access-controlled through IAM, so changing which service account an instance uses requires a permission that security can restrict, preventing an engineer from silently widening firewall scope by editing instance metadata.
    • B. Network tags are plain instance metadata that most engineers can edit directly, so an attacker or careless engineer can add a permissive tag to an instance without any IAM approval, widening its effective access.
    • C. Subnet-based targeting forces every workload in that subnet to share one access level, which does not scale for autoscaled groups with mixed roles and still lets an engineer move an instance into a broader subnet.
    • D. Autoscaled instances are recreated with new external IPs and this design likely uses Cloud NAT or no external IPs at all, so IP-based rules would be unstable and would not solve the metadata-editing risk described.

    Subdomain 2.1: Configuring network topologies.

    13.A security team must add intrusion detection for east-west traffic between subnets in a production VPC, prevent sensitive BigQuery datasets from being exfiltrated even by a principal with valid IAM credentials operating from an unauthorized network, and retain a record of every firewall connection for audit review. Which combination of controls should the architect deploy? (Select 3)(Select 3)

    1. A.Deploy Cloud IDS with a mirroring policy on the subnets carrying east-west traffic so a managed engine can inspect it for known attack signatures.
    2. B.Wrap the BigQuery datasets in a VPC Service Controls perimeter so data cannot be exfiltrated to unauthorized networks even by a caller holding valid IAM permissions.
    3. C.Enable VPC Flow Logs and firewall rule logging on the relevant networks so every accepted and denied connection is recorded for later audit review.
    4. D.Grant every analyst project-level Editor permissions on the BigQuery datasets so they no longer need to individually request access through IAM before running each query.
    5. E.Delete all deny firewall rules from the VPC so that Cloud IDS has an unobstructed view of all traffic without any packets being blocked beforehand.
    6. F.Replace IAM-based authentication for BigQuery with a single shared service account credential distributed to every analyst who needs dataset access.
    Show answer & explanation

    Correct answers: A, B, C — Deploy Cloud IDS with a mirroring policy on the subnets carrying east-west traffic so a managed engine can inspect it for known attack signatures.; Wrap the BigQuery datasets in a VPC Service Controls perimeter so data cannot be exfiltrated to unauthorized networks even by a caller holding valid IAM permissions.; Enable VPC Flow Logs and firewall rule logging on the relevant networks so every accepted and denied connection is recorded for later audit review.

    • A. Cloud IDS mirrors traffic on the targeted subnets to a managed detection engine that inspects it against known attack signatures, directly satisfying the requirement to add intrusion detection for east-west traffic between subnets.
    • B. A VPC Service Controls perimeter around the BigQuery datasets blocks data movement to unauthorized networks even when the caller presents valid IAM credentials, which is exactly the exfiltration protection the scenario requires.
    • C. Turning on VPC Flow Logs together with firewall rule logging records every accepted and denied connection through the network, giving the audit trail the security team needs to review later.
    • D. Granting broad project-level Editor permissions removes the least-privilege access control the datasets need and increases exfiltration risk rather than preventing it, working against the stated security goal.
    • E. Removing deny firewall rules eliminates network-level access control entirely and would let unwanted traffic reach the subnets unblocked, which weakens security instead of improving intrusion detection visibility.
    • F. Sharing one service account credential among every analyst removes individual accountability and makes IAM-based access review impossible, undermining both the audit and least-privilege goals of the scenario.

    Subdomain 2.5: Configuring prebuilt solutions or APIs with Agent Platform.

    14.Case study: TerramEarth manufactures heavy equipment and maintains a global network of independent dealers. Dealers currently call a support line to check parts availability and compatibility against TerramEarth's internal parts catalog and service manuals. TerramEarth wants a self-service chat experience that answers dealer questions grounded in that internal content and cites the source manual page for each answer, without the team hand-building retrieval and ranking logic. Which approach best satisfies this?

    1. A.Build a grounded search-and-chat agent with Vertex AI Search and Conversation, pointed at a data store built from the catalog and manuals.
    2. B.Deploy the Natural Language API's syntax analysis feature against the service manuals to produce a parts glossary dealers can browse by category.
    3. C.Call the Speech-to-Text API on recordings of past support calls to build a searchable archive of previously answered dealer questions.
    4. D.Run the Vision API's document text detection feature over scanned manual pages to extract raw text into a shared internal spreadsheet.
    Show answer & explanation

    Correct answer: A — Build a grounded search-and-chat agent with Vertex AI Search and Conversation, pointed at a data store built from the catalog and manuals.

    • A. Vertex AI Search and Conversation is built for exactly this pattern: it ingests a data store such as the parts catalog and manuals, then serves a grounded chat interface that answers questions and cites the source passages.
    • B. Syntax analysis parses grammatical structure such as parts of speech, and produces no ranked retrieval or chat interface, so it cannot power a question-answering experience on its own.
    • C. Transcribing past call recordings only creates more raw audio-derived text; it does not deliver a grounded chat interface that answers new dealer questions against the catalog and manuals.
    • D. Extracting raw text from scanned pages into a spreadsheet produces unstructured text a dealer would still have to search manually, with no chat interface, ranking, or citation of source pages.

    Subdomain 2.5: Configuring prebuilt solutions or APIs with Agent Platform.

    15.A team has fine-tuned a model from Model Garden and is preparing to deploy it to serve production traffic for an internal application. Which practices should they follow for this deployment? (Select 3.)(Select 3)

    1. A.Deploy the model behind a private endpoint reachable only from the application's VPC, rather than exposing a public endpoint to the internet.
    2. B.Configure autoscaling with a minimum and maximum replica count sized to the application's expected traffic pattern and latency targets.
    3. C.Enable model monitoring to detect prediction drift over time so the team is alerted before accuracy degrades unnoticed in production.
    4. D.Deploy the endpoint publicly with no authentication configured, since internal applications are assumed to be trusted by default.
    5. E.Provision a single fixed replica with no autoscaling enabled, since inference workloads are assumed to have constant, predictable demand.
    6. F.Skip endpoint access logging entirely, since production monitoring is expected to rely only on the application's own client-side metrics.
    Show answer & explanation

    Correct answers: A, B, C — Deploy the model behind a private endpoint reachable only from the application's VPC, rather than exposing a public endpoint to the internet.; Configure autoscaling with a minimum and maximum replica count sized to the application's expected traffic pattern and latency targets.; Enable model monitoring to detect prediction drift over time so the team is alerted before accuracy degrades unnoticed in production.

    • A. A private endpoint scoped to the application's VPC keeps the deployed model reachable only from trusted network paths, which is standard practice for a production endpoint serving an internal application.
    • B. Sizing autoscaling minimum and maximum replicas to expected traffic and latency targets lets the endpoint absorb load spikes while avoiding the cost of over-provisioned idle capacity, a core production deployment practice.
    • C. Model monitoring for prediction drift catches gradual accuracy degradation as real-world input distributions shift away from the fine-tuning data, which a team would otherwise only discover from downstream complaints.
    • D. Assuming an internal application is trusted by default and skipping authentication leaves the endpoint exposed to anything on the network path, which a production deployment should never assume is safe.
    • E. A single fixed replica with no autoscaling leaves no headroom for traffic spikes and no resilience if that replica fails, which is a fragile choice for a production-facing endpoint.
    • F. Skipping endpoint access logging removes the server-side record needed to diagnose failed requests or investigate anomalous traffic, which client-side application metrics alone cannot reconstruct.

    Subdomain 2.4: Leveraging Gemini Enterprise Agent Platform for end-to-end ML workflows.

    16.A retail analytics team is preparing to onboard several structured and unstructured data sources into Agent Platform Pipelines for a churn-prediction workflow. Which three actions correctly prepare the data integration layer before pipeline runs begin? (Select three.)(Select 3)

    1. A.Load structured customer transaction data into BigQuery and expose curated views so pipeline components can query it directly through the BigQuery connector.
    2. B.Register the engineered churn features in Vertex AI Feature Store so both training pipeline steps and online serving can read consistent feature values.
    3. C.Stage raw unstructured documents and images in Cloud Storage buckets organized by ingestion date so pipeline components can reference them as artifacts.
    4. D.Convert every data source into a single relational schema stored only in Cloud SQL, since Agent Platform Pipelines can only read from a relational database connector.
    5. E.Skip access control configuration on the new data sources until after the first pipeline run succeeds, then apply IAM policies retroactively once the schema is validated.
    6. F.Require all incoming data to be fully labeled by a human reviewer before it can be staged anywhere, even for exploratory pipeline runs used for schema validation.
    Show answer & explanation

    Correct answers: A, B, C — Load structured customer transaction data into BigQuery and expose curated views so pipeline components can query it directly through the BigQuery connector.; Register the engineered churn features in Vertex AI Feature Store so both training pipeline steps and online serving can read consistent feature values.; Stage raw unstructured documents and images in Cloud Storage buckets organized by ingestion date so pipeline components can reference them as artifacts.

    • A. Curated BigQuery views give pipeline components a governed, queryable interface to structured transaction data without each component needing to know the raw table layout.
    • B. Registering features in Feature Store ensures training and online serving read the same computed values, which avoids the training-serving skew that comes from recomputing features in two separate places.
    • C. Cloud Storage staged by ingestion date gives unstructured artifacts a predictable, referenceable location that pipeline components can read as inputs without depending on a database connector.
    • D. This is a false constraint: Agent Platform Pipelines can read from BigQuery, Cloud Storage, Feature Store, and other connectors, so forcing every source into one relational schema throws away structure best suited to each data type.
    • E. Deferring access control until after a pipeline run succeeds exposes sensitive data during the run itself; IAM and access boundaries should be configured before any pipeline reads the data, not retroactively.
    • F. Requiring full human labeling before any staging, even for exploratory schema validation, blocks the team from simply inspecting a data source's structure and unnecessarily delays early integration work.

    Subdomain 2.4: Leveraging Gemini Enterprise Agent Platform for end-to-end ML workflows.

    17.A team needs to run two different workloads: a lightweight function that resizes and validates uploaded images before they enter a pipeline, invoked sporadically throughout the day, and a multi-day, GPU-intensive training run for a large computer-vision model. Which pairing of services correctly matches each workload?

    1. A.Use Cloud Run functions for the sporadic image-resizing step, and use AI Hypercomputer GPU resources orchestrated through Agent Platform Pipelines for the multi-day training run.
    2. B.Use AI Hypercomputer GPU resources for the sporadic image-resizing step, and use Cloud Run functions to execute the entire multi-day computer-vision training run.
    3. C.Use Cloud Run functions for both workloads, since Cloud Run functions can run continuously for multiple days at a time without any maximum execution time limit.
    4. D.Use AI Hypercomputer GPU resources for both workloads, since provisioning always-on GPU capacity is by far the most cost-effective choice for a sporadic, short-lived function overall.
    Show answer & explanation

    Correct answer: A — Use Cloud Run functions for the sporadic image-resizing step, and use AI Hypercomputer GPU resources orchestrated through Agent Platform Pipelines for the multi-day training run.

    • A. Cloud Run functions fit a sporadic, short-lived task like image resizing because it only runs when invoked, while a multi-day GPU training run needs the sustained accelerator capacity that AI Hypercomputer, orchestrated by Agent Platform Pipelines, is built to provide.
    • B. This reverses the fit: provisioning GPU capacity for a sporadic function wastes idle accelerator time, and Cloud Run functions are not designed to sustain a multi-day, resource-intensive training run.
    • C. Cloud Run functions have a maximum execution time per invocation, so they cannot host a multi-day training run regardless of how the workload is structured.
    • D. Keeping GPU capacity always on for a sporadic, short-lived function pays for idle accelerator time between invocations, which is the opposite of cost-effective for that workload.

    Domain 3: Designing for security and compliance

    Subdomain 3.2: Designing for compliance.

    18.EHR Healthcare provides software services to hospital and clinic customers and stores patient electronic health records that fall under HIPAA. The architecture team is designing the data platform for a new patient portal that will run on Google Cloud. Which control is the most important prerequisite before any protected health information can be processed on the platform?

    1. A.Sign a Business Associate Agreement with Google Cloud and configure the workloads to use only HIPAA-covered products under that agreement.
    2. B.Enable Cloud CDN in front of the patient portal so that page load times stay low for clinicians accessing records during appointments.
    3. C.Move all patient record storage into a single multi-region Cloud Storage bucket to maximize object durability across the organization.
    4. D.Grant the clinical support team project-level Owner roles so they can troubleshoot patient record issues without escalation delays.
    Show answer & explanation

    Correct answer: A — Sign a Business Associate Agreement with Google Cloud and configure the workloads to use only HIPAA-covered products under that agreement.

    • A. HIPAA requires a signed Business Associate Agreement between the covered entity and any vendor that processes protected health information, and Google Cloud only extends HIPAA eligibility to the specific products listed under that agreement, so this must be in place before any PHI touches the platform.
    • B. A content delivery network improves latency for static assets but has no bearing on whether the organization has the legal and contractual coverage HIPAA requires for handling protected health information.
    • C. Storage durability and region configuration are architectural choices that matter after HIPAA eligibility and contractual coverage are established; durability alone does not satisfy the regulatory prerequisite for processing PHI.
    • D. Broad Owner-level access directly violates the least-privilege principle HIPAA compliance programs require and increases the risk of unauthorized disclosure rather than addressing the prerequisite for handling PHI at all.

    Subdomain 3.2: Designing for compliance.

    19.A retail company must demonstrate to a PCI DSS Qualified Security Assessor that only a small, defined set of engineers can approve changes to firewall rules protecting the cardholder-data environment, and that every change request and approval is recorded. Which two practices best support passing this control review? (Choose 2)(Select 2)

    1. A.Require firewall rule changes to pass through a change ticketing system that records the requester, approver, and justification.
    2. B.Restrict the IAM roles that can modify VPC firewall rules to a small, named group of network security engineers only.
    3. C.Allow any engineer with project Editor access to modify firewall rules directly, since Editor already implies senior trust.
    4. D.Disable Cloud Logging on the network project to reduce noise in the security operations team's alerting dashboard.
    5. E.Rotate the shared root service account key used for firewall automation every twelve months instead of every six months.
    Show answer & explanation

    Correct answers: A, B — Require firewall rule changes to pass through a change ticketing system that records the requester, approver, and justification.; Restrict the IAM roles that can modify VPC firewall rules to a small, named group of network security engineers only.

    • A. A ticketing-based change management process that records who requested a change, who approved it, and why creates the auditable trail a PCI DSS assessor expects to see for controls protecting the cardholder-data environment.
    • B. Limiting the IAM roles capable of modifying firewall rules to a small, named group directly enforces the requirement that only a defined set of engineers can approve changes to controls guarding cardholder data.
    • C. The broad Editor role grants far more permissions than firewall management alone and is not scoped to a small, defined group, which works against the principle of least privilege the assessor is testing.
    • D. Disabling logging removes the evidence trail needed to prove who made changes and when, which is the opposite of what a PCI DSS control review requires for the cardholder-data environment.
    • E. A shared service account undermines individual accountability regardless of key rotation cadence, since a review cannot tie a specific change back to a specific engineer when credentials are shared.

    Subdomain 3.1: Designing for security.

    20.A company wants to extend zero-trust access controls to SaaS applications that employees reach through their browser, not just applications hosted on Google Cloud, including real-time data-loss-prevention scanning of files being downloaded from those SaaS apps. Which Google Cloud offering is designed for this?

    1. A.Chrome Enterprise Premium, which applies context-aware access and data-loss-prevention scanning to browser sessions reaching SaaS apps.
    2. B.Identity-Aware Proxy, since it can be configured in front of any third-party SaaS application regardless of where that application is hosted.
    3. C.VPC Service Controls, since perimeters can be extended to include external SaaS vendors as long as they share a Google Cloud organization.
    4. D.Access Context Manager alone, since access levels are automatically enforced by every browser without any additional client software.
    Show answer & explanation

    Correct answer: A — Chrome Enterprise Premium, which applies context-aware access and data-loss-prevention scanning to browser sessions reaching SaaS apps.

    • A. Chrome Enterprise Premium applies BeyondCorp-style context-aware access controls and real-time data-loss-prevention scanning directly in the browser, covering SaaS destinations that the organization does not host, not just Google Cloud applications.
    • B. IAP is a reverse proxy that sits in front of applications the organization deploys behind a Google Cloud load balancer; it cannot be inserted in front of a third-party SaaS vendor's own infrastructure.
    • C. VPC Service Controls protects Google Cloud API calls within a defined perimeter of Google Cloud projects; it has no mechanism to extend that perimeter to an external SaaS vendor's own services.
    • D. Access levels are policy objects evaluated by services like IAP and VPC Service Controls; a browser does not automatically enforce them without an integration point such as Chrome Enterprise Premium.

    Subdomain 3.1: Designing for security.

    21.An architect wants build artifacts to carry cryptographically verifiable proof of exactly which source commit and build steps produced them, so downstream consumers can confirm an image was not tampered with after the build. Which Google Cloud capability provides this?

    1. A.Cloud Storage object versioning on the build output bucket, which keeps a history of every artifact ever produced by the pipeline.
    2. B.Artifact Registry's vulnerability scanning feature, which reports known CVEs in an image's operating system packages after each push.
    3. C.Cloud Logging's build logs, which record the console output of each build step for later review by the security team.
    4. D.Cloud Build's SLSA-compliant build provenance, which signs an attestation describing the source, steps, and inputs used.
    Show answer & explanation

    Correct answer: D — Cloud Build's SLSA-compliant build provenance, which signs an attestation describing the source, steps, and inputs used.

    • A. Object versioning preserves a history of files in a bucket but carries no cryptographic claim about which source commit or build steps produced any particular version.
    • B. Vulnerability scanning identifies known CVEs in installed packages, which is useful for patch management but says nothing about whether the artifact was built by the expected pipeline from the expected source.
    • C. Build logs record console output for debugging and review, but plain-text logs are not a signed, tamper-evident attestation that downstream consumers can cryptographically verify.
    • D. SLSA-compliant build provenance from Cloud Build produces a signed, non-falsifiable record of the exact source commit, build steps, and inputs that produced the artifact, which is precisely the cryptographic proof of origin the architect needs.

    Domain 4: Analyzing and optimizing technical and business processes

    Subdomain 4.2: Analyzing and defining business processes.

    22.A finance team is comparing the total cost of continuing to purchase on-premises servers every three years against moving the workload to Google Cloud with committed use discounts. Which framing correctly distinguishes the CapEx and OpEx implications of this decision?

    1. A.On-premises purchases are capital expenditures depreciated over time, while committed use discounts convert cloud spend into predictable, recurring operating expenses.
    2. B.Both on-premises servers and committed use discounts are capital expenditures, so the only meaningful difference between them is the total dollar amount spent.
    3. C.Committed use discounts are capital expenditures because they require an upfront one-to-three-year financial commitment similar to buying physical hardware outright.
    4. D.On-premises servers are operating expenses because they are paid for through the IT department's annual budget rather than a separate capital budget line.
    Show answer & explanation

    Correct answer: A — On-premises purchases are capital expenditures depreciated over time, while committed use discounts convert cloud spend into predictable, recurring operating expenses.

    • A. On-premises hardware purchases are capitalized and depreciated over their useful life, while committed use discounts are still consumption-based operating spend, just with a predictable, discounted recurring cost.
    • B. Treating both options as capital expenditure ignores the fundamental accounting distinction between owned depreciable assets and ongoing operating spend, which materially affects budgeting and tax treatment.
    • C. A multi-year committed use discount is still billed as recurring operating expense under standard cloud accounting, not capitalized like a physical asset purchase.
    • D. How a purchase is funded internally does not determine its accounting classification; depreciable hardware purchases are capital expenditures regardless of which budget line pays for them.

    Subdomain 4.2: Analyzing and defining business processes.

    23.After a failed production deployment caused a multi-hour outage, the organization wants to strengthen its change management process without slowing down every future release. Which action best balances these goals?

    1. A.Introduce automated pre-deployment checks and staged rollouts for changes above a defined risk threshold, leaving low-risk changes on the existing fast path.
    2. B.Require every future change, regardless of size or risk, to go through the same full manual review process used for the failed deployment.
    3. C.Freeze all production deployments indefinitely until leadership decides on a new process, with no interim guidance for the engineering teams.
    4. D.Remove change review entirely and trust individual engineers to self-certify that their changes are safe before deploying to production.
    Show answer & explanation

    Correct answer: A — Introduce automated pre-deployment checks and staged rollouts for changes above a defined risk threshold, leaving low-risk changes on the existing fast path.

    • A. Adding automated checks and staged rollouts specifically for higher-risk changes strengthens oversight where it is needed while keeping the existing fast path for changes that carry little risk.
    • B. Applying the same heavy manual review to every change regardless of risk slows down low-risk releases without necessarily preventing the kind of failure that caused the outage.
    • C. An indefinite freeze with no interim guidance stalls all delivery and leaves teams without a clear path forward, which is a disproportionate response to one failed deployment.
    • D. Removing review entirely and relying on self-certification removes exactly the safeguard whose absence likely contributed to the outage in the first place.

    Subdomain 4.1: Analyzing and defining technical processes.

    24.A team is designing the disaster recovery approach for a Cloud SQL for PostgreSQL instance that backs a customer-facing order system. Regional outages must not put the order history at risk of loss. Which combination of building blocks should the team include? (Choose 3.)(Select 3)

    1. A.Schedule automated backups and export them to a multi-region Cloud Storage bucket so backups survive the loss of the primary region entirely.
    2. B.Write and rehearse a documented runbook for promoting the replica to primary and repointing the application's database connection string.
    3. C.Rely on the application's in-memory cache to reconstruct recent order history after a regional failure takes the primary database completely offline.
    4. D.Provision a cross-region read replica so a promotable, continuously updated copy of the order data exists outside the primary region.
    5. E.Disable point-in-time recovery to reduce storage costs, on the assumption that the cross-region replica alone covers every recovery need.
    6. F.Store the only backup copy on a persistent disk attached to the primary instance, since that gives the fastest possible restore time.
    Show answer & explanation

    Correct answers: A, B, D — Schedule automated backups and export them to a multi-region Cloud Storage bucket so backups survive the loss of the primary region entirely.; Write and rehearse a documented runbook for promoting the replica to primary and repointing the application's database connection string.; Provision a cross-region read replica so a promotable, continuously updated copy of the order data exists outside the primary region.

    • A. Automated backups exported to a multi-region bucket give the team a recovery path that is independent of the primary region's availability, protecting against a scenario where the region itself is lost.
    • B. A rehearsed promotion runbook turns disaster recovery from a theoretical capability into a tested procedure, so the team can actually execute the failover within its recovery time objective when it matters.
    • C. An application cache holds only a transient, incomplete view of recent activity and is not a durable or complete record of order history, so it cannot substitute for a database recovery strategy.
    • D. A cross-region read replica keeps continuously replicated data outside the primary region, so it can be promoted to become the new primary if that region becomes unavailable.
    • E. Point-in-time recovery protects against logical errors such as accidental deletes within the primary region, and disabling it removes a recovery capability the cross-region replica does not provide.
    • F. A backup on a disk attached to the same instance is lost along with that instance during a regional outage, so it does not protect the order history against the failure the team is planning for.

    Subdomain 4.1: Analyzing and defining technical processes.

    25.Cymbal Retail expects a large surge in e-commerce traffic during an upcoming holiday shopping event and wants confidence that its checkout service can handle the projected peak load before the event begins. Which testing practice should be built into the release process ahead of the event?

    1. A.Run the existing unit test suite one additional time before the event, since unit tests already cover the checkout service's business logic.
    2. B.Run automated load tests against a staging environment that mirrors production, driving traffic beyond the projected peak volume before release.
    3. C.Ask the customer support team to estimate whether the checkout service seems fast enough based on their day-to-day interactions with customers.
    4. D.Skip load testing this cycle and instead monitor production closely during the event so the team can react quickly if the checkout service is overwhelmed.
    Show answer & explanation

    Correct answer: B — Run automated load tests against a staging environment that mirrors production, driving traffic beyond the projected peak volume before release.

    • A. Unit tests validate business logic in isolation and do not exercise the service under realistic concurrent load, so they cannot reveal whether the checkout path will hold up at peak volume.
    • B. Driving simulated load up to and beyond the expected peak against a production-like staging environment surfaces capacity limits, such as database connection exhaustion or autoscaling delays, while there is still time to fix them before real customers hit the same limits.
    • C. Support team impressions from routine interactions are not a substitute for measured load against realistic peak traffic volumes, since day-to-day usage does not resemble a holiday traffic surge.
    • D. Reacting during the event itself means any capacity problem is discovered while real customers are trying to check out, which is exactly the outcome Cymbal Retail wants to avoid by testing ahead of time.

    Domain 5: Managing implementation

    Subdomain 5.1: Advising development and operation teams to ensure the successful deployment of the solution.

    26.A team deploys a new version of a customer-facing service to Google Kubernetes Engine and wants to reduce the blast radius of a bad release. They want new traffic to be gradually shifted to the new revision in small percentage increments, with automated rollback triggered if error rates or latency exceed defined thresholds during the rollout, and without building and maintaining their own custom rollout tooling. Which approach should the team adopt?

    1. A.Configure a Cloud Deploy delivery pipeline with a canary deployment strategy and Cloud Monitoring-based verification so traffic shifts gradually and an unhealthy release rolls back automatically.
    2. B.Replace the entire Kubernetes Deployment object in a single `kubectl apply` command each release, relying on the default rolling update to gradually swap out old pods.
    3. C.Deploy the new version to a completely separate GKE cluster and manually update the DNS record to point all traffic to the new cluster once testing is complete.
    4. D.Tag the new container image and push it to Artifact Registry, since pushing an image to the registry is sufficient to gradually shift live production traffic to it.
    Show answer & explanation

    Correct answer: A — Configure a Cloud Deploy delivery pipeline with a canary deployment strategy and Cloud Monitoring-based verification so traffic shifts gradually and an unhealthy release rolls back automatically.

    • A. Cloud Deploy's canary strategy shifts a configurable percentage of traffic to the new revision in stages and can be paired with automated verification against Cloud Monitoring metrics, triggering a rollback if error rate or latency thresholds are breached, without custom rollout logic.
    • B. A default rolling update replaces pods based on readiness rather than shifting a controlled percentage of live traffic, and it has no built-in automated rollback tied to error-rate or latency thresholds, so a bad release can still reach most users.
    • C. Standing up a parallel cluster and cutting DNS over manually is an all-or-nothing switch rather than a gradual, percentage-based traffic shift, and DNS changes propagate on their own timeline, which does not give the fine-grained control the team wants.
    • D. Pushing a tagged image to Artifact Registry only makes the new version available to be deployed; it has no effect on live traffic routing and does nothing to gradually shift users or trigger a rollback.

    Subdomain 5.1: Advising development and operation teams to ensure the successful deployment of the solution.

    27.A logistics company built its shipment-tracking backend as a set of internal microservices with no external-facing access layer. The company now wants to expose a subset of that functionality to external carrier partners as a stable, versioned API, while retaining the ability to change the internal microservices' implementation without breaking partner integrations, and while applying consistent authentication and rate limiting across every partner-facing endpoint. What should the architecture team introduce?

    1. A.An Apigee API management layer in front of the internal microservices, exposing a stable proxy interface to partners while decoupling it from internal implementation changes.
    2. B.Direct external network access from partner systems straight to each internal microservice's private IP address, since a management layer would only add unnecessary latency.
    3. C.A shared internal service account credential distributed to every external partner, since a single shared credential is sufficient for both authentication and rate limiting.
    4. D.A one-time data export of the tracking database delivered to each partner nightly, since exporting data avoids the need to expose any API surface at all.
    Show answer & explanation

    Correct answer: A — An Apigee API management layer in front of the internal microservices, exposing a stable proxy interface to partners while decoupling it from internal implementation changes.

    • A. An Apigee layer in front of the internal microservices gives partners a stable, versioned proxy interface while the underlying implementation can change freely behind it, and it provides a single place to apply consistent authentication and rate-limiting policy across every partner-facing endpoint.
    • B. Giving partner systems direct network access to each internal microservice's private IP ties every partner integration to the current internal implementation and topology, so any internal refactor would directly break external partners.
    • C. A single shared credential for every partner cannot distinguish which partner made a given call, which makes it impossible to apply per-partner rate limits or to revoke one partner's access without affecting all the others.
    • D. A nightly data export replaces real-time API access with a stale batch file, which does not meet the goal of exposing live functionality as a versioned API and does nothing to enforce authentication or rate limiting per partner.

    Subdomain 5.2: Interacting with Google Cloud programmatically.

    28.Your team currently applies Terraform configurations by hand from each engineer's laptop and wants managed state storage, deployment previews, and audit logging without switching to a different CI system. Which Google Cloud capability provides this directly?

    1. A.Infrastructure Manager, which runs your Terraform blueprint as a managed service, handling state storage, execution, previews, and audit logging.
    2. B.Infrastructure Manager, which replaces Terraform's configuration language with a proprietary syntax that must be rewritten before any blueprint can be deployed.
    3. C.Infrastructure Manager, which only supports resources originally created through the Cloud Console and cannot manage anything provisioned by `terraform apply`.
    4. D.Infrastructure Manager, which only estimates the monthly cost of a Terraform plan and never actually executes any part of the underlying deployment itself.
    Show answer & explanation

    Correct answer: A — Infrastructure Manager, which runs your Terraform blueprint as a managed service, handling state storage, execution, previews, and audit logging.

    • A. Infrastructure Manager takes an existing Terraform root module as a blueprint and runs it as a managed service, storing state, generating previews before apply, and logging every deployment action for audit purposes.
    • B. Infrastructure Manager consumes standard Terraform configuration written in HashiCorp Configuration Language directly; it does not require rewriting configurations into a different syntax.
    • C. Infrastructure Manager deploys whatever resources a Terraform blueprint defines, regardless of whether they were originally created through the console or an earlier `terraform apply`, as long as state is imported correctly.
    • D. Infrastructure Manager actually executes the deployment and manages its state; cost estimation is not its core function, and it does far more than produce a price estimate.

    Subdomain 5.2: Interacting with Google Cloud programmatically.

    29.Which statement correctly describes the relationship between the Google Cloud SDK and the `gcloud`, `gsutil`, and `bq` command-line tools?

    1. A.The Cloud SDK is the installable package that bundles `gcloud`, `gsutil`, and `bq` together, along with the libraries and dependencies they share.
    2. B.`gcloud`, `gsutil`, and `bq` are three unrelated third-party projects that happen to share a name but ship from separate, independent installers.
    3. C.Only `gcloud` is officially part of the Cloud SDK; `gsutil` and `bq` must always be installed separately from entirely unrelated package repositories.
    4. D.The Cloud SDK was fully discontinued and replaced entirely by `gcloud storage`, which now also handles every BigQuery and Compute Engine task.
    Show answer & explanation

    Correct answer: A — The Cloud SDK is the installable package that bundles `gcloud`, `gsutil`, and `bq` together, along with the libraries and dependencies they share.

    • A. The Cloud SDK installer bundles `gcloud`, `gsutil`, and `bq` along with their shared client libraries, giving a single installation path for all three command-line tools.
    • B. All three tools are first-party Google-maintained components distributed together as part of the same SDK, not independent third-party projects with separate installers.
    • C. `gsutil` and `bq` are included in the same Cloud SDK installation as `gcloud`, rather than requiring separate installs from unrelated repositories.
    • D. The Cloud SDK remains the active distribution mechanism for these tools; `gcloud storage` is a newer command group for Cloud Storage operations within `gcloud`, not a replacement for the entire SDK or for `bq`.

    Domain 6: Ensuring solution and operations excellence

    Subdomain 6.1: Understanding the principles and recommendations of the operational excellence pillar of the Google Cloud Well-Architected Framework

    30.A platform team at a growing SaaS company experienced a production outage. After restoring service, the team wants a practice that documents the incident timeline, root cause, and contributing factors without assigning blame to any individual, so preventive fixes get prioritized as follow-up work. Which practice should the team adopt?

    1. A.Conduct a blameless postmortem that captures the incident timeline, root cause, and contributing factors, and track the resulting action items to completion.
    2. B.Ask the on-call engineer who was paged during the outage to write a private summary for their manager, and close the incident once that summary is submitted.
    3. C.Schedule a retrospective meeting only when an outage breaches a four-hour service level agreement, and skip the review process for every shorter incident.
    4. D.Wait until the next quarterly business review to discuss the outage alongside other operational metrics, instead of analyzing it immediately after resolution.
    Show answer & explanation

    Correct answer: A — Conduct a blameless postmortem that captures the incident timeline, root cause, and contributing factors, and track the resulting action items to completion.

    • A. A blameless postmortem that records the timeline, root cause, and contributing factors, with tracked action items, is the operational excellence practice for turning an incident into concrete preventive work without singling out individuals.
    • B. A private summary written by only the paged engineer and sent to a manager assigns the review to one individual and keeps the findings out of shared team knowledge, which undermines a blameless, organization-wide review.
    • C. Reviewing only outages that cross a fixed duration threshold means shorter but still preventable incidents never get analyzed, so recurring risks in those shorter incidents go unaddressed.
    • D. Delaying the review to a quarterly cadence separates the analysis from the incident by months, by which point details of the timeline and contributing factors are harder to reconstruct accurately.

    Subdomain 6.2: Familiarity with Google Cloud Observability solutions.

    31.Case study: TerramEarth manufactures heavy equipment and is migrating telemetry ingestion from its vehicles to Google Cloud. Each vehicle streams sensor readings through Pub/Sub into a Dataflow pipeline that writes to BigQuery, and the operations team wants to detect ingestion pipeline problems before the business-critical predictive-maintenance dashboards go stale, without generating alerts for every transient Pub/Sub backlog blip that resolves on its own. Which two monitoring and alerting practices should the team implement?(Select 2)

    1. A.Create an alerting policy on the Pub/Sub subscription's oldest unacknowledged message age that only fires once the age sustains above a threshold for several consecutive minutes, filtering out transient backlog spikes that clear on their own.
    2. B.Define an SLO for the predictive-maintenance dashboards based on end-to-end data freshness, and configure a burn-rate alert that pages engineers only when the freshness SLI is trending toward breaching that SLO within the error budget window.
    3. C.Configure Cloud Monitoring to page an engineer immediately whenever the Pub/Sub subscription's backlog exceeds zero messages, since any queued message represents an ingestion failure that must be investigated in real time.
    4. D.Set a single alerting policy on Dataflow worker CPU utilization and rely on it exclusively to represent the health of the entire ingestion pipeline, from the vehicle's initial Pub/Sub publish through the final BigQuery write.
    5. E.Disable all automated alerting on the ingestion pipeline and instead have an engineer manually query BigQuery once per day to confirm that new telemetry rows are present for every vehicle fleet in the program.
    6. F.Route every Dataflow worker log entry at INFO severity into a paging notification channel for the on-call engineer, so that normal, expected pipeline operating activity itself repeatedly triggers an unnecessary page.
    Show answer & explanation

    Correct answers: A, B — Create an alerting policy on the Pub/Sub subscription's oldest unacknowledged message age that only fires once the age sustains above a threshold for several consecutive minutes, filtering out transient backlog spikes that clear on their own.; Define an SLO for the predictive-maintenance dashboards based on end-to-end data freshness, and configure a burn-rate alert that pages engineers only when the freshness SLI is trending toward breaching that SLO within the error budget window.

    • A. Requiring the oldest unacknowledged message age to sustain above a threshold for several minutes before firing filters out normal, self-resolving backlog fluctuations while still catching a genuinely stuck subscription.
    • B. An SLO on end-to-end data freshness with a burn-rate alert ties paging directly to whether the business-critical dashboards are actually at risk of going stale, which is the symptom that matters rather than an intermediate pipeline metric.
    • C. Paging on any nonzero backlog treats every routine, momentary queuing event as an incident, which will page the team constantly for conditions that resolve on their own and is exactly the alert fatigue the team wants to avoid.
    • D. Worker CPU utilization is only one stage of a multi-stage pipeline and does not by itself capture failures at the Pub/Sub ingestion point or the BigQuery write step, so it is an incomplete proxy for overall pipeline health.
    • E. A once-daily manual check leaves the team blind to problems for up to a full day, which is too slow to catch an ingestion failure before the predictive-maintenance dashboards have already gone stale for customers.
    • F. INFO-level logs represent normal operation, so paging on every such entry generates a continuous stream of unnecessary pages for expected activity rather than surfacing genuine anomalies.

    Subdomain 6.3: Deployment and release management

    32.Case study: TerramEarth. Firmware-triggered backend updates must reach a production Cloud Run service that processes telemetry from manufacturing equipment, and the operations team requires that a named engineer explicitly sign off before any release reaches the production target, while staging deployments proceed without intervention. Which combination of Cloud Deploy capabilities addresses this requirement? (Choose 2)(Select 2)

    1. A.Configure the production target in the delivery pipeline to require approval, so the rollout pauses until an authorized user approves it before deploying.
    2. B.Leave the staging target's promotion configuration without a required-approval setting so its rollout proceeds automatically once the prior stage's rollout succeeds.
    3. C.Grant the reviewing engineer an IAM role that includes the rollout-approval permission, since approval authority is controlled through IAM policy rather than target configuration alone.
    4. D.Set the production target's `strategy` field to `canary` with a 100% initial phase, expecting this alone to pause every rollout until it is manually resumed by a reviewer.
    5. E.Enable Cloud Audit Logs on the delivery pipeline resource, expecting the logging configuration by itself to insert a human approval gate before any rollout can proceed.
    Show answer & explanation

    Correct answers: A, C — Configure the production target in the delivery pipeline to require approval, so the rollout pauses until an authorized user approves it before deploying.; Grant the reviewing engineer an IAM role that includes the rollout-approval permission, since approval authority is controlled through IAM policy rather than target configuration alone.

    • A. Marking the production target as requiring approval causes its rollout to pause in an awaiting-approval state until an authorized user approves or rejects it, which is exactly the sign-off gate the operations team wants. Cloud Deploy generates approval-related Pub/Sub notifications so the reviewer can be alerted when action is needed.
    • B. A target without a required-approval setting advances automatically once the previous rollout succeeds, so leaving staging unconfigured this way lets it proceed without manual intervention. This matches the requirement that only production needs sign-off.
    • C. Requiring approval on a target is not sufficient by itself; the specific engineer must also hold an IAM role or permission that grants rollout-approval rights, since Cloud Deploy authorizes the approve action through IAM. Configuring the target alone would leave no one able to actually approve the pending rollout.
    • D. A canary strategy phases traffic percentages and can include verification jobs, but a 100% initial phase does not by itself create a manual, named-approver sign-off gate independent of IAM permissions. Pausing for verification is a different mechanism than the explicit approval workflow requested here.
    • E. Cloud Audit Logs record who took which action against a resource for compliance and forensic review, but enabling logging does not insert an approval requirement into a rollout. A target must be explicitly configured to require approval for that gate to exist.

    Subdomain 6.4: Assisting with the support of deployed solutions

    33.In Google Cloud's operational excellence practices, what is the primary purpose of conducting a blameless postmortem after a production incident?

    1. A.To identify the contributing factors and systemic gaps behind an incident and produce concrete follow-up actions, without assigning individual fault to the engineers involved.
    2. B.To determine which specific engineer made the mistake that caused the incident, so that corrective performance feedback can be delivered to that person.
    3. C.To calculate the exact financial cost of the outage for the finance department, kept separate from any technical analysis of what actually caused the incident to occur.
    4. D.To satisfy a one-time reporting obligation, after which the incident's documentation can be discarded entirely since the underlying issue has already been fixed in production.
    Show answer & explanation

    Correct answer: A — To identify the contributing factors and systemic gaps behind an incident and produce concrete follow-up actions, without assigning individual fault to the engineers involved.

    • A. A blameless postmortem focuses on the systemic and process factors that allowed an incident to happen, producing action items that reduce the chance of recurrence, rather than pointing at an individual.
    • B. Assigning individual fault is the opposite of the blameless approach; naming a specific engineer discourages honest reporting in future incidents and is not the postmortem's purpose.
    • C. Financial cost accounting may sometimes accompany incident follow-up, but it is not the primary purpose of a postmortem, which is centered on understanding and preventing the technical and process failure.
    • D. Postmortems are meant to be retained as institutional knowledge that informs future reliability work, not discarded once the immediate fix is in place.

    Subdomain 6.6: Ensuring the reliability of solutions in production (e.g., chaos engineering, penetration testing, and load testing)

    34.Case study: Mountkirk Games is preparing to launch a new multiplayer mobile game expected to reach 500,000 concurrent players within its first month. The game backend runs on GKE and uses Cloud Spanner for player state and Memorystore for session caching. The SRE team must validate that the backend can sustain the expected peak concurrency and surface bottlenecks before the public launch, while keeping the testing infrastructure temporary and cost-controlled. Which approach best validates production reliability before launch?

    1. A.Deploy a distributed load-testing cluster on GKE running an open-source tool such as Locust across many worker pods to generate concurrent simulated player traffic against staging, then tear the cluster down after the test.
    2. B.Configure Cloud Monitoring uptime checks against the production API endpoints, then rely on the resulting latency graphs alone to infer whether the backend supports the expected 500,000 concurrent players.
    3. C.Increase the Cloud Spanner instance's processing units to the maximum supported size ahead of launch, and treat that capacity increase alone as proof that peak concurrency has been validated.
    4. D.Enable Cloud CDN in front of the game's API gateway, then treat the resulting drop in origin request counts as confirmation the backend can handle the expected player concurrency.
    Show answer & explanation

    Correct answer: A — Deploy a distributed load-testing cluster on GKE running an open-source tool such as Locust across many worker pods to generate concurrent simulated player traffic against staging, then tear the cluster down after the test.

    • A. A distributed load-testing cluster on GKE can generate realistic concurrent traffic that mirrors expected player load, letting the SRE team observe Cloud Spanner and Memorystore behavior under stress and surface bottlenecks before launch. Running the worker pods temporarily on GKE and deleting them afterward keeps the exercise cost-controlled, which is exactly what the scenario asks for.
    • B. Uptime checks confirm an endpoint is reachable and record baseline latency under whatever traffic is already occurring, but they do not generate the concurrent load needed to reveal how the system behaves at 500,000 concurrent players. Relying on them alone leaves capacity and bottleneck questions completely unanswered before launch.
    • C. Raising Cloud Spanner processing units increases available compute capacity, but capacity alone does not prove the rest of the stack, including GKE autoscaling, Memorystore, and application logic, behaves correctly under realistic concurrent load. Sizing decisions without an actual load test amount to an untested assumption about launch readiness.
    • D. Enabling Cloud CDN can reduce origin traffic for cacheable content, but a multiplayer game's real-time session and state APIs are largely non-cacheable, so a drop in origin requests says little about backend concurrency handling. This does not substitute for actually generating and measuring peak simulated load.

    Subdomain 6.5: Evaluating quality control measures

    35.Case study: Mountkirk Games. The studio wants to test a new matchmaking algorithm with real players before rolling it out to everyone, while guaranteeing that if the new algorithm causes matchmaking failures, only a small fraction of players are affected and the change can be reversed within minutes. Which release strategy should the team implement?

    1. A.Configure a Cloud Deploy pipeline with a canary strategy that routes a small, configurable share of matchmaking traffic to the new version, monitors error and latency metrics, and promotes or rolls back automatically.
    2. B.Deploy the new algorithm to a fully separate GKE cluster, manually redirect all player traffic to that cluster at once, and redirect traffic back to the original cluster once players start reporting problems.
    3. C.Merge the new algorithm's feature branch directly into main and deploy it through the existing pipeline with no staged rollout at all, so every player receives it in the same release window.
    4. D.Ship the new algorithm behind a client-side configuration flag that individual players toggle manually, and wait for player-submitted bug reports before deciding whether the release should stay live.
    Show answer & explanation

    Correct answer: A — Configure a Cloud Deploy pipeline with a canary strategy that routes a small, configurable share of matchmaking traffic to the new version, monitors error and latency metrics, and promotes or rolls back automatically.

    • A. A canary strategy exposes only a small, controllable slice of players to the new matchmaking version while metric-based automation watches error rate and latency, which limits the blast radius exactly as required. Because the pipeline can halt or reverse the rollout automatically once thresholds are crossed, recovery happens in minutes rather than waiting on a person to notice.
    • B. Redirecting all traffic to the new cluster at once exposes every player to the new algorithm simultaneously, which is the opposite of limiting impact to a small fraction of players. Relying on manual redirection back to the original cluster also depends on someone noticing and acting quickly, which is slower than an automated canary rollback.
    • C. Deploying directly to every player with no staged rollout means a flawed algorithm affects the entire player base immediately, with no way to limit exposure beforehand. There is also no automated mechanism described here to detect a problem or reverse it quickly.
    • D. A player-controlled toggle depends on players opting in and reporting bugs themselves, which is slow, unpredictable, and does not guarantee only a small fraction of players are exposed at any given time. Waiting for bug reports is also far slower than automated metric-based detection and rollback.

    Want the full experience?

    These are just samples. Practice the full Google Professional Cloud Architect question bank in quiz mode — free, no signup, with domain practice and exam simulation.