Answers

What is a VLA (vision-language-action) model?

The change it brings is economic as much as technical. With classical robotics, every new task cost months of expert programming; with a VLA, the new task is taught with demonstrations and data, and the same model generalizes to objects it never saw. That is why the competitive edge has shifted from the body to the brain and the training data: whoever has hours of real recorded work (factories, warehouses) trains better models than whoever only has a lab. Both sides of the coin showed in 2026: heads, robots sorting packages for nearly 40 hours live; tails, a model that learns from data also degrades when the data changes, and on a physical robot that is managed with even more care.

The best-documented example in this house is Figure's Helix, and the map of who uses which brain is in what AI humanoid robots use. To tell when a VLA robot is truly autonomous and when there is a person behind it, our guide to teleoperated vs. autonomous.

The full context, in the full story.

Sources