Tackling Robotics with (V)LM Agents | Nishanth J. Kumar

-

-

-

-

-

Recently, there have been some interesting new results showing that some of the latest large-scale VLMs (Vision-Language Models) can directly control robots to solve a variety of real-world tasks. This has led to some claims and excitement that progress in robotics might happen as an emergent effect of scaling current multi-modal models instead of training different robotics-specific models. As someone who’s been doing research in the field for a few years now, I decided to dive into trying to understand these results and sort through the various potential implications and future directions.

What are the new results?

One widely-circulated result that also has a clear description of the testing protocol is from Robocurve . Researchers asked a few recent frontier VLMs (Claude Fable 5, Claude Fable 5.1, GPT-6 Astra) to perform some simple tasks (placing a block into a bowl, placing a puzzle piece into a groove) by controlling some YAM arms . The model takes the task description and camera images (and a history of previous images and actions if available) and outputs end-effector positions and orientation for each end-effector 1 . Another related result is from the RoboDojo team 2 , who ran GPT-6 Astra and GPT-5.5 through much the same kind of interface on their official simulation benchmark: 42 tasks, 50 episodes each, ranked against the 43 policies on their public leaderboard. Here, Astra placed first, ahead of every learned policy on the board. In both cases there is no other learned policy anywhere between the model and the robot.