Subdomain 1.1: Construct distributed feature engineering pipelines.
1.An engineer loads an 800-million-row feature table with `session.table("FEATURES")`, applies several `filter` and `with_column` calls, then runs `df.to_pandas()` to feed a local scaler. The notebook kernel runs out of memory. What is the best redesign?
- A.Increase the notebook kernel memory and call `to_pandas()` in a loop over `limit` and `offset` slices to process rows in chunks
- B.Keep the work as Snowpark DataFrame operations and persist results with `write.save_as_table` so it runs on the warehouse
- C.Replace the Snowpark calls with a Python loop that issues a `session.sql` SELECT per customer id and appends each result locally
- D.Export the table to CSV files in an internal stage with `COPY INTO`, then download the files and concatenate them using pandas
Show answer & explanation
Correct answer: B — Keep the work as Snowpark DataFrame operations and persist results with `write.save_as_table` so it runs on the warehouse
- A. Incorrect: slicing with limit and offset still pulls every row to the client, repeats the scan per slice, and gives unstable ordering without a sort.
- B. Correct: Snowpark DataFrames are lazy and compile into SQL, so transformations and the final write run in the warehouse and nothing large reaches the client.
- C. Incorrect: issuing one query per customer id creates huge numbers of round trips and still concentrates all rows in client memory.
- D. Incorrect: downloading unloaded CSV files moves the entire dataset out of Snowflake and recreates the same memory limitation locally.