AI Glossary
Multimodal AI
AI that processes different data types – text, image, audio, video – simultaneously.
Definition
Multimodal AI processes and generates different types of data simultaneously – text, images, audio, and video. While early models handled only one modality (e.g., text only), multimodal models can describe an image, convert a sketch into code, or link spoken language with visual information. GPT-4o and Gemini are examples of multimodal models.
Related Terms
Sources: a16z AI Glossary · CNET AI Terms Glossary