Multimodal AI
Artificial intelligence system that can process different types of data such as text, visual, audio.
Technically multi-mode models often use a separate encoder for every fashionity (e.g. a visual Transformer for images, a language model encoder for text) and works by combining this encoder output in a common space. Systems such as GPT-4o, Gemini and Claude’s visual comprehension capabilities, CLIP are examples of this category, such as visual-language alignment models and text generating visual/video from Sora, Midjourney. Multi-mode artificial intelligence plays an increasingly central role in the solution of real world problems, from diagnostic systems that evaluate patient notes with medical image, visual disabilities users to voice-described assistants.
