AI Dictionary › AI Fundamentals

Multi-modal AI

Multi-modal AI is the capability of an AI model to process and generate different types of media, not just text, but also images, audio, video, and code, in an integrated way. Multi-modal models can receive an image as input and respond in text, or conversely generate images from text descriptions.

Definition

The major modern AI models are already multi-modal: GPT-4V can "see" and analyze images; Gemini was natively designed to be multi-modal; Claude can analyze PDF documents and images.

Related terms

More in AI Fundamentals

Put it into practice

From our network

AGORÀ Intelligence: Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.