Industry
What is a vision-language-action (VLA) model, and who is building them?
A vision-language-action (VLA) model is one neural network that reads camera images and a plain-language instruction and outputs a robot's next motor commands, built by training a vision-language model on robot demonstrations. Google DeepMind, Physical Intelligence, NVIDIA, Figure AI, Skild AI and AgiBot lead, and the commercial race now turns on the cost of action data rather than model design.