M Harfi 👁 23 views

Multimodal AI

Artificial intelligence system that can process different types of data such as text, visual, audio.

Multi-mode artificial intelligence (multimodal AI) refers to the artificial intelligence systems that can understand, associate and produce together in the same model of text, visual, audio, video. Traditional models were often customised to a single fashionity: a model only processed the text, another model. Multi-mode models can produce text that describe a visual for example, bringing together different fashionities in a common representation (Pinkdding) space, for example, a text can create visual or answer a voicemail with a visual context. This is a more close approach to human balance that integrates information from different senses.

Technically multi-mode models often use a separate encoder for every fashionity (e.g. a visual Transformer for images, a language model encoder for text) and works by combining this encoder output in a common space. Systems such as GPT-4o, Gemini and Claude’s visual comprehension capabilities, CLIP are examples of this category, such as visual-language alignment models and text generating visual/video from Sora, Midjourney. Multi-mode artificial intelligence plays an increasingly central role in the solution of real world problems, from diagnostic systems that evaluate patient notes with medical image, visual disabilities users to voice-described assistants.