Start for Free Book a Demo Contact
Platform Login
AI Glossary

Multimodal AI

AI that processes different data types – text, image, audio, video – simultaneously.

Definition

Multimodal AI processes and generates different types of data simultaneously – text, images, audio, and video. While early models handled only one modality (e.g., text only), multimodal models can describe an image, convert a sketch into code, or link spoken language with visual information. GPT-4o and Gemini are examples of multimodal models.

Sources: a16z AI Glossary · CNET AI Terms Glossary

← Back to AI Glossary