Subdomain 1.4: Identify the use cases and strengths of Google’s foundation models.
1.A film studio experimenting with previsualization wants to generate rough scene footage from a written script excerpt before committing budget to actual filming. Which model is designed for this?
- A.Veo, because it generates video sequences directly from descriptive text, useful for rough scene previsualization
- B.Imagen, because its multi-frame mode stitches sequential still images into a low-frame-rate scene preview
- C.Gemini, because its multimodal output can render a fully animated video summary of the script
- D.Gemma, because its open architecture lets studios train a custom scene-generation model overnight
Show answer & explanation
Correct answer: A — Veo, because it generates video sequences directly from descriptive text, useful for rough scene previsualization
- A. This model's core purpose is generating video sequences from text prompts, which is exactly what turning a script excerpt into rough scene footage requires. That makes it the model built for text-to-video previsualization work.
- B. This model generates still images rather than video, and it does not offer a mode that stitches stills into motion footage. Previsualization needs actual generated video, which is outside this model's core capability.
- C. This general reasoning and text model is not built to render animated video output from a script; its strength is multimodal understanding and conversation, not video generation. The studio needs a dedicated video-generation model for this task.
- D. This open model is meant for self-hosted fine-tuning of general or text models, not for training a scene-generation capability overnight, and it is not the model built for producing video. It would not deliver rough scene footage from a script.