# Multimodal

Models that handle more than text — images, audio, video, documents — in and out. Practical upshot: you can photograph a contract, a spreadsheet, or a broken part and ask AI about it.

Canonical: https://robauto.ai/learn/ai-101/11

_AI 101: The 2026 Foundation — lesson 11 of 20 (DEFINITION)_

Models that handle more than text — images, audio, video, documents — in and out. Practical upshot: you can photograph a contract, a spreadsheet, or a broken part and ask AI about it.

Source: [Google DeepMind — Gemini](https://deepmind.google/technologies/gemini/?utm_source=robauto)

[Previous lesson](/learn/ai-101/10) · [Next lesson](/learn/ai-101/12) · [Course overview](/learn/ai-101) · [All courses](/learn)

---

(c) 2026 Robauto, Inc. — support@robauto.ai
Machine surfaces: https://robauto.ai/llms.txt · https://robauto.ai/llms-full.txt · https://robauto.ai/.well-known/api-catalog
