Multimodal

AI 101: The 2026 Foundation — lesson 11 of 20 (DEFINITION)

Models that handle more than text — images, audio, video, documents — in and out. Practical upshot: you can photograph a contract, a spreadsheet, or a broken part and ask AI about it.

Source: Google DeepMind — Gemini

Previous lesson · Next lesson · Course overview · All courses