VLA (Vision-Language-Action)

IoT & Robotics: AI in the Physical World — lesson 7 of 20 (DEFINITION)

Foundation models that take camera input + a plain-English instruction (‘put the red parts in bin 3’) and output robot movements. VLA adoption tripled in a year — now in ~40% of new robot deployments. Programming is becoming prompting.

Source: State of Robotics 2026

Previous lesson · Next lesson · Course overview · All courses