Subdomain 1.2: Scaling and Tuning
1.An engineer configures `TorchDistributor.init(num_processes=4, local_mode=True, use_gpu=True)` and calls `.run(train_fn)` on a Databricks cluster with a driver that has 4 GPUs and several worker nodes attached. Where does the training actually execute?
- A.All 4 processes run on the single driver machine, since `local_mode=True` confines TorchDistributor to launching multiple processes on one node.
- B.The 4 processes are spread one per worker node across the cluster, since `local_mode=True` always maps each process to a distinct physical node.
- C.TorchDistributor ignores `local_mode` when `use_gpu=True` and always distributes processes across every available GPU in the entire cluster.
- D.The call raises an error because `num_processes` must equal the total worker node count whenever `local_mode` is explicitly set to `True`.
Show answer & explanation
Correct answer: A — All 4 processes run on the single driver machine, since `local_mode=True` confines TorchDistributor to launching multiple processes on one node.
- A. Correct: `local_mode=True` tells TorchDistributor to launch the specified number of processes on a single machine, so all 4 GPU processes run on the driver rather than being spread across workers.
- B. Spreading one process per worker node is the behavior of multi-node mode (`local_mode=False`), not local mode, which intentionally keeps every process on one machine.
- C. `use_gpu` only controls whether processes use GPU devices; it does not override `local_mode`, which independently governs whether execution stays on one node or spans the cluster.
- D. `num_processes` in local mode specifies how many processes run on the single local machine and has no required relationship to the number of worker nodes in the cluster.