Robotics and Embodied AI511 · Module I · Lesson 01 of 3
Article · 12 min

Vision-language-action models

The RT-2 lineage.

Summary

The RT-2 lineage. This lesson sits inside Module I — The field — of Robotics and Embodied AI, the course that anchors the Advanced AI Research program. It is not a survey; it is the specific, working understanding of "Vision-language-action models" that the rest of the course assumes you carry forward.

Objectives
  • 01Define Vision-language-action models in the precise sense used across Robotics and Embodied AI.
  • 02Recognize when Vision-language-action models is the correct lens for the situation in front of you, and when it is not.
  • 03Apply Vision-language-action models to a concrete case drawn from The field, and defend the result in plain language.
  • 04Connect Vision-language-action models to the adjacent lessons in this module without collapsing the distinctions between them.
The Lesson

The idea, stated plainly

The RT-2 lineage. That single sentence is the whole lesson in compressed form. The rest of the reading unfolds it — what it means when the terms are taken seriously, where it comes from, and what work it does inside Robotics and Embodied AI. Read the sentence, then read it again after the sections below; it should carry more weight the second time.

Why it belongs in The field

Module I exists because the current state. "Vision-language-action models" is one of the pillars of that module: without it, the later lessons either become memorization or lose their bite. Notice which earlier lessons this one leans on, and which later lessons will lean on it — the shape of the module is easier to see once you place this piece.

How the School of Artificial Intelligence faculty use it

In practice, working school of artificial intelligence professionals reach for this idea before they reach for a formula or a tool. It is a way of framing the problem so that the right question comes first. The mark of understanding is not that you can recite Vision-language-action models; it is that you catch yourself using it, unprompted, when the situation calls for it.

Common misreadings

The most frequent error is to treat Vision-language-action models as a slogan and skip the mechanics. The second most frequent is the opposite — treating the mechanics as the point, when the mechanics are only there to make the idea usable. Both errors collapse the same distinction, and both are correctable by returning to the one-line summary and asking what it actually claims.

Key Ideas
  • Vision-language-action models is a working tool, not a slogan.
  • Its meaning is set by the module it lives in: The field.
  • Understanding is demonstrated by unprompted use in the correct situation.
  • The adjacent lessons in this module are its natural context; read them together.
References
  • 511 — Robotics and Embodied AI, Module I: The fieldThe parent module for this lesson. Re-read the module blurb after finishing the lesson.
  • The Anabasis Academy — School of Artificial Intelligence, Advanced AI ResearchThe wider program this lesson serves; the Certificate in Advanced AI Research (Expert tier). Admission requires the AI Foundations and AI Engineering certificates or the equivalent in industry. credential ultimately certifies mastery of ideas like this one.