Subdomain 1.1: Ingest and store data
1.A team needs to merge nightly sales data extracted from an on-premises Oracle database over JDBC with clickstream data already stored as Parquet in S3, apply a multi-step Spark transformation to deduplicate and join the datasets, and load the merged result back into S3, all on a fully managed, serverless schedule without provisioning a persistent cluster. Which service fits this requirement?
- A.A persistent Amazon EMR cluster running Apache Spark, with both the Oracle JDBC driver and the S3 connector installed, kept running continuously to handle the nightly merge.
- B.A SageMaker Processing job running a single-instance pandas script that opens a JDBC connection to Oracle and reads the S3 Parquet files into memory for the join.
- C.An AWS Glue ETL job configured with both a JDBC connection to the Oracle database and an S3 data source, running the Spark-based merge logic on Glue's serverless infrastructure.
- D.An Amazon Redshift Spectrum query that joins an external table pointing at the S3 Parquet files with a federated query against the on-premises Oracle database.
Show answer & explanation
Correct answer: C — An AWS Glue ETL job configured with both a JDBC connection to the Oracle database and an S3 data source, running the Spark-based merge logic on Glue's serverless infrastructure.
- A. A persistent EMR cluster can run the same Spark logic but requires the team to provision, patch, and pay for cluster infrastructure continuously, which directly conflicts with the fully managed, serverless requirement stated.
- B. A single-instance pandas script would struggle to scale to nightly sales and clickstream data volumes and would require the team to manage retries, scheduling, and JDBC driver packaging themselves rather than using a managed ETL service.
- C. AWS Glue provides serverless, managed Spark execution along with built-in JDBC connections for relational sources and native S3 support, so it can run the described multi-step merge on a schedule without the team provisioning or managing any cluster.
- D. Redshift Spectrum federated queries can join across sources, but this pattern is built around SQL-based analytical queries rather than orchestrating a scheduled, multi-step Spark-based deduplication and merge ETL pipeline.