2.7 Emerging multimodal trends

NCA-GENM · Core Machine Learning and AI Knowledge (20% of the exam) · Official objective: “Track emerging multimodal trends and technologies.”

Vision language models and world foundation models such as NVIDIA Cosmos.

Key points

  1. A VLM combines a large language model with a vision encoder so the LLM can 'see'. Unlike fixed-class vision models, it can follow natural-language instructions.

    What NVIDIA says (1)

    “Vision language models (VLMs) are multimodal, generative AI models capable of understanding and processing video, image, and text.”

    — NVIDIA Glossary: Vision Language Models

  2. Zero-shot means doing a task with no task-specific training examples. Traditional CV models must be retrained when a new class is added. VLMs means vision language models.

    What NVIDIA says (1)

    “Out of the box, VLMs have strong zero-shot performance on a variety of vision tasks”

    — NVIDIA Glossary: Vision Language Models

  3. Physical AI is AI that perceives and acts in the physical world, such as robots and vehicles. Cosmos provides world foundation models that generate and predict video of the world.

    What NVIDIA says (1)

    “NVIDIA Cosmos is a developer-first platform for designing Physical AI systems.”

    — NVIDIA Cosmos: Introduction

Key terms

Sample question

What is a vision language model (VLM)?

Show the answer

Answer: A multimodal generative AI model that understands video, image and text

A VLM combines a large language model with a vision encoder so the LLM can 'see'. Unlike fixed-class vision models, it can follow natural-language instructions.

What NVIDIA says (1)

“Vision language models (VLMs) are multimodal, generative AI models capable of understanding and processing video, image, and text.”

— NVIDIA Glossary: Vision Language Models

Practice 2.7 (3 questions) Full Core Machine Learning and AI Knowledge guide

← 2.6 Multimodal transfer learning · 2.8 Energy-efficient, trustworthy models →