2.7 Emerging multimodal trends
Vision language models and world foundation models such as NVIDIA Cosmos.
Key points
A VLM combines a large language model with a vision encoder so the LLM can 'see'. Unlike fixed-class vision models, it can follow natural-language instructions.
What NVIDIA says (1)
“Vision language models (VLMs) are multimodal, generative AI models capable of understanding and processing video, image, and text.”
Zero-shot means doing a task with no task-specific training examples. Traditional CV models must be retrained when a new class is added. VLMs means vision language models.
What NVIDIA says (1)
“Out of the box, VLMs have strong zero-shot performance on a variety of vision tasks”
Physical AI is AI that perceives and acts in the physical world, such as robots and vehicles. Cosmos provides world foundation models that generate and predict video of the world.
What NVIDIA says (1)
“NVIDIA Cosmos is a developer-first platform for designing Physical AI systems.”
Key terms
- Multimodal model: A model that works with more than one type of data, such as text, images, audio or video.
- Vision language model: A generative model that combines a large language model with a vision encoder so it can answer questions about images and video.
- NVIDIA Cosmos: An NVIDIA platform of world foundation models for building Physical AI systems.
- Zero-shot: Doing a task with no examples given in the prompt or for training.
Sample question
What is a vision language model (VLM)?
Show the answer
Answer: A multimodal generative AI model that understands video, image and text
A VLM combines a large language model with a vision encoder so the LLM can 'see'. Unlike fixed-class vision models, it can follow natural-language instructions.
What NVIDIA says (1)
“Vision language models (VLMs) are multimodal, generative AI models capable of understanding and processing video, image, and text.”
Practice 2.7 (3 questions) Full Core Machine Learning and AI Knowledge guide
← 2.6 Multimodal transfer learning · 2.8 Energy-efficient, trustworthy models →