Domain 1 of 4 · Chapter 4 of 4

Google Foundation Models

Unlock the complete study guide + 1,040 practice questions across 16 full exams.

Bundled into the existing Generative AI Leader premium course — no separate purchase.

14-day money-back guarantee — no questions asked.

Included in this chapter:

  • The four families and what each is for
  • Gemini: the multimodal workhorse
  • Gemma, Imagen, and Veo up close
  • Choosing a model for a use case

Google's four generative foundation model families at a glance

DimensionGeminiGemmaImagenVeo
Primary outputText plus multimodal reasoningText (open-weight LLM)ImagesVideo with audio
OpennessProprietary, managedOpen-weight, self-hostableProprietary, managedProprietary, managed
Where it runsGoogle-managed APIYour infra, on-device, or Model GardenGoogle-managed APIGoogle-managed API
Signature strengthMultimodal input, long context, agentic reasoningLightweight, private, customizablePhotorealism and text-in-imageCinematic clips with native audio
Reach for it whenChatbots, analysis, agents, summarizationData must stay in your control, offline, or fine-tunedMarketing, product, and packaging imageryVideo ads, storyboards, previsualization

Decision tree

What must the modelproduce?ImageVideoText or chatImagenimagesVeovideo + audioData must stay inyour control?YesNoGemmaopen-weight, self-hostedHardest reasoning orlong inputs?YesNoGemini Prodeepest reasoningGemini Flashfast, low cost

Cheat sheet

  • Pick the model family by what it must produce
  • Gemini is Google's multimodal large language model
  • Choose the Gemini tier by trading capability against cost and latency
  • Gemini Nano is the compact tier for on-device use
  • Gemma is the open-weight choice for control and privacy
  • Imagen generates images from text
  • Imagen can render legible text inside a generated image
  • Veo generates video with synchronized native audio
  • Proprietary models are managed APIs; Gemma is the one you self-host
  • Vertex AI Model Garden is the single catalog for all four generative families
  • Understanding an image is Gemini; creating one is Imagen
  • Chirp is Google's speech recognition model spanning many languages, including rare ones

Unlock with Premium — includes all practice exams and the complete study guide.

Also tested in

References

  1. Imagen — Google's text-to-image model
  2. Veo — Google's video generation model
  3. Gemma — Google's open models
  4. Gemini — Google's multimodal model family
  5. Vertex AI foundation models and Model Garden