{"article":{"slug":"tackling-robotics-with-v-lm-agents","title":"Tackling Robotics with (V)LM Agents","subtitle":null,"summary":"Nishanth J. Kumar surveys recent demos and ideas around GPT-6 and other vision-language models solving robotics tasks—summarizing approaches and offering thoughts on what works and what still breaks.","content_type":"blog_post","language":"en","canonical_url":"https://nishanthjkumar.com/blog/2026/Tackling-Robotics-with-(V)LM-Agents/","author":{"name":"Nishanth J. Kumar","url":"https://nishanthjkumar.com","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"nishanthjkumar.com","url":"https://nishanthjkumar.com","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Robotics","slug":"robotics","url":"https://listedarticles.com/topics/robotics"},{"name":"LLMs","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":2304,"reading_minutes":10,"published_at":"2026-09-22T15:00:00.000Z","added_at":"2026-09-24T03:17:12.485Z","updated_at":"2026-09-24T03:17:12.485Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/tackling-robotics-with-v-lm-agents","markdown_url":"https://listedarticles.com/articles/tackling-robotics-with-v-lm-agents.md","example":false,"citation":"Nishanth J. Kumar, nishanthjkumar.com. \"Tackling Robotics with (V)LM Agents.\" 22 Sept 2026. https://nishanthjkumar.com/blog/2026/Tackling-Robotics-with-(V)LM-Agents/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://nishanthjkumar.com/blog/2026/Tackling-Robotics-with-(V)LM-Agents/"},"body_markdown":"Tackling Robotics with (V)LM Agents | Nishanth J. Kumar\n\n-\n\n-\n\n-\n\n-\n\n-\n\nRecently, there have been some interesting new results showing that some of the latest large-scale VLMs (Vision-Language Models) can directly control robots to solve a variety of real-world tasks. This has led to some claims and excitement that progress in robotics might happen as an emergent effect of scaling current multi-modal models instead of training different robotics-specific models. As someone who’s been doing research in the field for a few years now, I decided to dive into trying to understand these results and sort through the various potential implications and future directions.\n\n# What are the new results?\n\nOne widely-circulated result that also has a clear description of the testing protocol is from Robocurve . Researchers asked a few recent frontier VLMs (Claude Fable 5, Claude Fable 5.1, GPT-6 Astra) to perform some simple tasks (placing a block into a bowl, placing a puzzle piece into a groove) by controlling some YAM arms . The model takes the task description and camera images (and a history of previous images and actions if available) and outputs end-effector positions and orientation for each end-effector 1 . Another related result is from the RoboDojo team 2 , who ran GPT-6 Astra and GPT-5.5 through much the same kind of interface on their official simulation benchmark: 42 tasks, 50 episodes each, ranked against the 43 policies on their public leaderboard. Here, Astra placed first, ahead of every learned policy on the board. In both cases there is no other learned policy anywhere between the model and the robot.\n\nThe full reported results are worth looking at because they yield at least two useful trends:\n\nRobocurve (real arms, 20 trials/task) RoboDojo (simulation, 42 tasks × 50 episodes)\nOverall — Astra 28.97, 1st of 43; best learned policy DM0.5 24.90, π0.5 3 11.41\nGeometry and semantics 4 95% Astra, 40% Fable 5.1, 5% Fable 5 Generalization 33.36 and Open 34.36, both 1st (next best 23.54 and 6.50)\nPrecision and contact 5 10% Astra, 10% Fable 5.1, 0% Fable 5 Precision 12.65 (4.0% SR) against 28.25 for the best VLA\nLong-horizon not evaluated 21.45 (8.25% SR) against 44.12 for the best VLA\n\nThe first thing is the trend line across model generations on the easy task(s): (e.g. 5% → 40% → 95% from Robocurve’s results), over models released within about a year of each other, none of which was built with these specific robot arms in mind. RoboDojo’s results show that GPT-5.5 and Astra scored 1.13 and 28.97 respectively on the same set of tasks with the same harness, which suggests the improvement is directly due to capability improvements in the underlying model. The second thing is that this trend does not seem to manifest on the harder task, where the newest and by far the strongest model does no better than one two generations older. In Robocurve’s testing of a contact-rich puzzle-piece insertion task, both models get the puzzle piece to the groove and then stall at the insertion — the part that requires reacting to contact rather than reasoning about geometry. RoboDojo’s per-axis scores show the same split, and show it more starkly: Astra places first overall only by dominating the generalization and open-instruction axes, and is beaten on precision (12.65 against 28.25) and long-horizon execution (21.45 against 44.12) by policies it otherwise outranks. Its success rate on the precision tasks is 4%. 6\n\nThere have been a handful of more anecdotal results. One X user found that GPT-6 Astra is able to solve a range of tasks much more quickly (in terms of number of turns) and successfully than other recent frontier VLMs. However, they allowed the model to write functions and use helper tools (e.g. SAM for perception) instead of directly having it output end-effector commands. A number of other users demonstrated Astra performing a variety of other impressive real-world tasks (e.g. painting the golden-gate bridge , setting up a MuJoCo simulation and drawing a dove inside it , training a dexterous pen-spinning behavior , maneuvering complex interlocking puzzle pieces to be unstuck , or turning the knob on a real-world washing machine ). However, these were largely demonstrations lacking thorough empirical results, comparisons to baselines, or clarity on the robot’s exact I/O specification or prompting setup.\n\nIt is important for context to note that this is not the first time that VLMs or LLMs have been used to directly control robots. There are a number of research works dating back to at least 2022 ( SayCan 7 , Code as Policies 8 , ProgPrompt 9 , Inner Monologue 10 ) that explore this idea, and it remains an extremely active area of investigation (e.g. this 11 , or this 12 ). Anthropic has even studied their latest models’ capabilities on robotics tasks for the past year or so ( Project Fetch 13 , and more recently an evaluation across quadrupeds, humanoids, and arms 14 ). What is noteworthy about these new results though is that the VLM is controlling the robot at a fairly low-level (end-effector targets) instead of via an intricate harness or a number of purpose-built tools, and that performance seems to have improved somewhat dramatically with recent improved model releases.\n\n# Why are these results exciting?\n\n15 Recent prevailing wisdom in robotics has been that robotics fundamentally requires capabilities that VLMs and LLMs cannot possess because of their structure , training data 16 , and/or training objectives. The vast majority of recent state-of-the-art robotics demos (e.g. this , or this 17 , or this 18 ) have come from training models on robotics data that are specialized to robotics. However, these new results and demos referenced above illustrate that VLMs might possess sparks of embodied intelligence and robotic control despite not being explicitly trained for these things. Moreover, it seems that newer generations of models are increasingly capable at these tasks. Given this, perhaps scaling up existing models and training techniques will lead to models that can do ever more complex robotics tasks. Thus, perhaps progress in robotics could happen as a by-product of large-scale foundation model training instead of requiring purpose-built robotics-specific models.\n\nThis is quite significant if true, and could have several important implications for current directions in the field.\n\n- Perhaps very little robotics data will be required . Many recent results have been built atop extremely large-scale, often proprietary, data-collection efforts, and there is significant interest and effort being expended to obtain the largest useful robotics dataset ( RealOmni-Open 19 , Index 20 ). However, if VLMs can learn to do well on robotics tasks from being trained (largely) on internet text, image, and coding data, then perhaps we don’t need much robotics data and these data collection efforts are unnecessary.\n\n- Transferring across embodiments might not be hard . A frequent challenge for robotics models has been having the same model control a variety of different robot hardware form-factors. There is significant effort (e.g. described here 21 , or here 22 ) required to train models to exhibit cross-embodiment generalization. If VLMs are able to solve robotics problems by writing programs or outputting controls specific to the problem and environment, then perhaps cross-embodiment generalization emerges without any significant or explicit attempts to train for it.\n\n- Harnesses and tools might be invaluable . A significant enabler for VLMs to be useful at coding has been good tooling and harnesses for developers (e.g. Claude Code ). Indeed, research in leveraging such models for robotics (which has been ongoing for a number of years) has had similar findings ( Code as Policies , CaP-X 12 , Waddle 23 ): building the right tools and abstractions around the model might determine success or failure for end-to-end robot behavior.\n\n## A related thought experiment\n\nIn some sense, building a general-purpose robot by training a model to be extremely good at programming and physical reasoning is not an altogether surprising strategy. We’ve known for a long time that it’s possible to program any specific robot to do any specific task. Assuming the environment is roughly static, and given enough time to think through the specific motions and measure out the relevant distances, a skilled robot programmer or automation integrator can write out a program that performs that specific task in that specific environment. This is, in fact, most of what industrial robotics is : the integration, programming and commissioning work wrapped around an arm routinely costs as much as the arm itself or more , and it has to be redone for every new robot and environment. The trouble has been generalizing the program to produce the correct motion in new tasks and environments.\n\nImagine for a moment a future where a general-purpose robot is widely and easily available. From any customer’s point of view, they can simply buy a robot, bring it to their environment of choice (e.g. a home, cafe, factory, etc.) and just ask it to start doing useful tasks. However, unbeknownst to the customer, each robot has a little (and invisible) elf that operates it. Whenever a human asks the robot to do something, the elf quickly runs around with a measuring tape and then writes out the precise program that generates the motions that accomplish that specific task in that environment.\n\nThis is an admittedly contrived setup, but I believe it is a useful lens for thinking about a particular approach to building such a general-purpose robot. Instead of trying to mimic/automate human brains (i.e., training a robotics-specific model) by having a model go directly from pixels to torques, we have a model that automates the programmer of robots, which is able to achieve tasks in a fully general way by writing hyper-specific programs for each task and environment it encounters.\n\n# What’s missing?\n\nWhile current results and what they promise are certainly exciting, there are several important and substantial hurdles to be cleared before it is clearly practical to build and deploy general-purpose robots by scaling VLMs.\n\n### Thorough experimentation and (strong) evidence of generalization\n\nThe bulk of current evidence is either a demonstration with no quantitative results, or a set of preliminary quantitative results that lacks scale and statistical rigor (i.e., hundreds or more trials with results over several random seeds). Moreover, many of them do not directly compare against established baselines from the research literature: Code as Policies or more recent improvements 12 run on the same frontier models, recent VLAs like π0.7 24 , MolmoAct2 25 or GR00T N1.5 26 , and recent WAMs like DreamZero 27 . These are currently far from the type of thorough result that could be published in a robotics research paper that would be accepted at a top conference or journal.\n\nAdditionally, a significant part of the excitement around these results comes from the idea that the underlying models were trained on a very small amount (if any) of robotics data. If in fact they were trained on large quantities of relevant robotics data (perhaps even on the tested embodiments) 28 , then these results do not really demonstrate that physical commonsense is coming as an emergent phenomenon of large-scale training on non-robotics data. Even if they were not trained on robotics data, it is not guaranteed that performance on robotics tasks will continue to scale with training in any meaningful way (i.e., the scaling laws 29 of these models on robotics tasks are entirely unclear). Thorough experimentation and some insight into the training data for these latest frontier models would help determine the extent to which there is substantial evidence of physical commonsense 30 emerging from scaling these models.\n\n### Task complexity\n\nCurrent results demonstrate behavior on relatively simple, often short-horizon tasks with limited contact and physical interaction. However, a significant challenge in robotics is solving long-horizon, contact-rich tasks that require significant dexterity — folding laundry out of a dryer 31 , assembling and packing deformable goods , making a bed or a coffee end-to-end , or cracking + beating eggs and using them to make an omelette . There is, as of yet, limited evidence that VLMs will be able to solve such tasks on robots directly 32 . Indeed, VLMs - even augmented with tools and harnesses - might be simply incapable of solving certain tasks. In the above-discussed thought experiment of the elf, it is possible that even an extremely competent elf would be unable to write the correct program without precisely measuring distances involved, or without a sense of touch, which could be impossible for a VLM operating just from robot cameras without any additional sensors. However, several research works have demonstrated that such models are able to leverage tools or ML (e.g. Eureka 33 , DrEureka 34 ) to perform dexterous behavior, so perhaps this could be a feasible path forward.\n\n### Reliability and safety on real hardware\n\nThere is much evidence that current models are not particularly safe or reliable at executing useful behavior on real robots. Robodojo results show that Astra solves their tasks with an average success rate of 22.48%, which is far from reliable completion. Moreover, the RoboDojo team had to halt their real-robot campaign for safety after Astra repeatedly issued physically unreasonable or unsafe actions, including incidents that damaged hardware. Robocurve found that several recent VLMs will execute harmful and dangerous tasks on real hardware without refusal. Model alignment remains a challenging problem even for disembodied VLMs : ensuring models will safely execute actions on hardware could be even more challenging.\n\n### Speed and cost in deployment\n\nOne issue noted as part of all the recent results is speed. Frontier VLMs take on the order of seconds (at best) to produce a turn of output. However, some tasks - such as walking even a quadruped robot - require commands to be output at a much higher frequency (e.g. 50 Hz 35 ). This control frequency is not feasible for modern large frontier VLMs 36 , and it is unclear that it will ever be. A related issue is the current paradigm for querying large models: it may not be feasible for every robot to assume constant connection to a large centralized model server simply because wifi might have too much latency or too little bandwidth to be reliable.","body_html":"<p>Tackling Robotics with (V)LM Agents | Nishanth J. Kumar</p>\n<p>-</p>\n<p>-</p>\n<p>-</p>\n<p>-</p>\n<p>-</p>\n<p>Recently, there have been some interesting new results showing that some of the latest large-scale VLMs (Vision-Language Models) can directly control robots to solve a variety of real-world tasks. This has led to some claims and excitement that progress in robotics might happen as an emergent effect of scaling current multi-modal models instead of training different robotics-specific models. As someone who’s been doing research in the field for a few years now, I decided to dive into trying to understand these results and sort through the various potential implications and future directions.</p>\n<h1 id=\"what-are-the-new-results\">What are the new results?</h1>\n<p>One widely-circulated result that also has a clear description of the testing protocol is from Robocurve . Researchers asked a few recent frontier VLMs (Claude Fable 5, Claude Fable 5.1, GPT-6 Astra) to perform some simple tasks (placing a block into a bowl, placing a puzzle piece into a groove) by controlling some YAM arms . The model takes the task description and camera images (and a history of previous images and actions if available) and outputs end-effector positions and orientation for each end-effector 1 . Another related result is from the RoboDojo team 2 , who ran GPT-6 Astra and GPT-5.5 through much the same kind of interface on their official simulation benchmark: 42 tasks, 50 episodes each, ranked against the 43 policies on their public leaderboard. Here, Astra placed first, ahead of every learned policy on the board. In both cases there is no other learned policy anywhere between the model and the robot.</p>\n<p>The full reported results are worth looking at because they yield at least two useful trends:</p>\n<p>Robocurve (real arms, 20 trials/task) RoboDojo (simulation, 42 tasks × 50 episodes)\nOverall — Astra 28.97, 1st of 43; best learned policy DM0.5 24.90, π0.5 3 11.41\nGeometry and semantics 4 95% Astra, 40% Fable 5.1, 5% Fable 5 Generalization 33.36 and Open 34.36, both 1st (next best 23.54 and 6.50)\nPrecision and contact 5 10% Astra, 10% Fable 5.1, 0% Fable 5 Precision 12.65 (4.0% SR) against 28.25 for the best VLA\nLong-horizon not evaluated 21.45 (8.25% SR) against 44.12 for the best VLA</p>\n<p>The first thing is the trend line across model generations on the easy task(s): (e.g. 5% → 40% → 95% from Robocurve’s results), over models released within about a year of each other, none of which was built with these specific robot arms in mind. RoboDojo’s results show that GPT-5.5 and Astra scored 1.13 and 28.97 respectively on the same set of tasks with the same harness, which suggests the improvement is directly due to capability improvements in the underlying model. The second thing is that this trend does not seem to manifest on the harder task, where the newest and by far the strongest model does no better than one two generations older. In Robocurve’s testing of a contact-rich puzzle-piece insertion task, both models get the puzzle piece to the groove and then stall at the insertion — the part that requires reacting to contact rather than reasoning about geometry. RoboDojo’s per-axis scores show the same split, and show it more starkly: Astra places first overall only by dominating the generalization and open-instruction axes, and is beaten on precision (12.65 against 28.25) and long-horizon execution (21.45 against 44.12) by policies it otherwise outranks. Its success rate on the precision tasks is 4%. 6</p>\n<p>There have been a handful of more anecdotal results. One X user found that GPT-6 Astra is able to solve a range of tasks much more quickly (in terms of number of turns) and successfully than other recent frontier VLMs. However, they allowed the model to write functions and use helper tools (e.g. SAM for perception) instead of directly having it output end-effector commands. A number of other users demonstrated Astra performing a variety of other impressive real-world tasks (e.g. painting the golden-gate bridge , setting up a MuJoCo simulation and drawing a dove inside it , training a dexterous pen-spinning behavior , maneuvering complex interlocking puzzle pieces to be unstuck , or turning the knob on a real-world washing machine ). However, these were largely demonstrations lacking thorough empirical results, comparisons to baselines, or clarity on the robot’s exact I/O specification or prompting setup.</p>\n<p>It is important for context to note that this is not the first time that VLMs or LLMs have been used to directly control robots. There are a number of research works dating back to at least 2022 ( SayCan 7 , Code as Policies 8 , ProgPrompt 9 , Inner Monologue 10 ) that explore this idea, and it remains an extremely active area of investigation (e.g. this 11 , or this 12 ). Anthropic has even studied their latest models’ capabilities on robotics tasks for the past year or so ( Project Fetch 13 , and more recently an evaluation across quadrupeds, humanoids, and arms 14 ). What is noteworthy about these new results though is that the VLM is controlling the robot at a fairly low-level (end-effector targets) instead of via an intricate harness or a number of purpose-built tools, and that performance seems to have improved somewhat dramatically with recent improved model releases.</p>\n<h1 id=\"why-are-these-results-exciting\">Why are these results exciting?</h1>\n<p>15 Recent prevailing wisdom in robotics has been that robotics fundamentally requires capabilities that VLMs and LLMs cannot possess because of their structure , training data 16 , and/or training objectives. The vast majority of recent state-of-the-art robotics demos (e.g. this , or this 17 , or this 18 ) have come from training models on robotics data that are specialized to robotics. However, these new results and demos referenced above illustrate that VLMs might possess sparks of embodied intelligence and robotic control despite not being explicitly trained for these things. Moreover, it seems that newer generations of models are increasingly capable at these tasks. Given this, perhaps scaling up existing models and training techniques will lead to models that can do ever more complex robotics tasks. Thus, perhaps progress in robotics could happen as a by-product of large-scale foundation model training instead of requiring purpose-built robotics-specific models.</p>\n<p>This is quite significant if true, and could have several important implications for current directions in the field.</p>\n<ul><li>Perhaps very little robotics data will be required . Many recent results have been built atop extremely large-scale, often proprietary, data-collection efforts, and there is significant interest and effort being expended to obtain the largest useful robotics dataset ( RealOmni-Open 19 , Index 20 ). However, if VLMs can learn to do well on robotics tasks from being trained (largely) on internet text, image, and coding data, then perhaps we don’t need much robotics data and these data collection efforts are unnecessary.</li><li>Transferring across embodiments might not be hard . A frequent challenge for robotics models has been having the same model control a variety of different robot hardware form-factors. There is significant effort (e.g. described here 21 , or here 22 ) required to train models to exhibit cross-embodiment generalization. If VLMs are able to solve robotics problems by writing programs or outputting controls specific to the problem and environment, then perhaps cross-embodiment generalization emerges without any significant or explicit attempts to train for it.</li><li>Harnesses and tools might be invaluable . A significant enabler for VLMs to be useful at coding has been good tooling and harnesses for developers (e.g. Claude Code ). Indeed, research in leveraging such models for robotics (which has been ongoing for a number of years) has had similar findings ( Code as Policies , CaP-X 12 , Waddle 23 ): building the right tools and abstractions around the model might determine success or failure for end-to-end robot behavior.</li></ul>\n<h2 id=\"a-related-thought-experiment\">A related thought experiment</h2>\n<p>In some sense, building a general-purpose robot by training a model to be extremely good at programming and physical reasoning is not an altogether surprising strategy. We’ve known for a long time that it’s possible to program any specific robot to do any specific task. Assuming the environment is roughly static, and given enough time to think through the specific motions and measure out the relevant distances, a skilled robot programmer or automation integrator can write out a program that performs that specific task in that specific environment. This is, in fact, most of what industrial robotics is : the integration, programming and commissioning work wrapped around an arm routinely costs as much as the arm itself or more , and it has to be redone for every new robot and environment. The trouble has been generalizing the program to produce the correct motion in new tasks and environments.</p>\n<p>Imagine for a moment a future where a general-purpose robot is widely and easily available. From any customer’s point of view, they can simply buy a robot, bring it to their environment of choice (e.g. a home, cafe, factory, etc.) and just ask it to start doing useful tasks. However, unbeknownst to the customer, each robot has a little (and invisible) elf that operates it. Whenever a human asks the robot to do something, the elf quickly runs around with a measuring tape and then writes out the precise program that generates the motions that accomplish that specific task in that environment.</p>\n<p>This is an admittedly contrived setup, but I believe it is a useful lens for thinking about a particular approach to building such a general-purpose robot. Instead of trying to mimic/automate human brains (i.e., training a robotics-specific model) by having a model go directly from pixels to torques, we have a model that automates the programmer of robots, which is able to achieve tasks in a fully general way by writing hyper-specific programs for each task and environment it encounters.</p>\n<h1 id=\"what-s-missing\">What’s missing?</h1>\n<p>While current results and what they promise are certainly exciting, there are several important and substantial hurdles to be cleared before it is clearly practical to build and deploy general-purpose robots by scaling VLMs.</p>\n<h3 id=\"thorough-experimentation-and-strong-evidence-of-generalization\">Thorough experimentation and (strong) evidence of generalization</h3>\n<p>The bulk of current evidence is either a demonstration with no quantitative results, or a set of preliminary quantitative results that lacks scale and statistical rigor (i.e., hundreds or more trials with results over several random seeds). Moreover, many of them do not directly compare against established baselines from the research literature: Code as Policies or more recent improvements 12 run on the same frontier models, recent VLAs like π0.7 24 , MolmoAct2 25 or GR00T N1.5 26 , and recent WAMs like DreamZero 27 . These are currently far from the type of thorough result that could be published in a robotics research paper that would be accepted at a top conference or journal.</p>\n<p>Additionally, a significant part of the excitement around these results comes from the idea that the underlying models were trained on a very small amount (if any) of robotics data. If in fact they were trained on large quantities of relevant robotics data (perhaps even on the tested embodiments) 28 , then these results do not really demonstrate that physical commonsense is coming as an emergent phenomenon of large-scale training on non-robotics data. Even if they were not trained on robotics data, it is not guaranteed that performance on robotics tasks will continue to scale with training in any meaningful way (i.e., the scaling laws 29 of these models on robotics tasks are entirely unclear). Thorough experimentation and some insight into the training data for these latest frontier models would help determine the extent to which there is substantial evidence of physical commonsense 30 emerging from scaling these models.</p>\n<h3 id=\"task-complexity\">Task complexity</h3>\n<p>Current results demonstrate behavior on relatively simple, often short-horizon tasks with limited contact and physical interaction. However, a significant challenge in robotics is solving long-horizon, contact-rich tasks that require significant dexterity — folding laundry out of a dryer 31 , assembling and packing deformable goods , making a bed or a coffee end-to-end , or cracking + beating eggs and using them to make an omelette . There is, as of yet, limited evidence that VLMs will be able to solve such tasks on robots directly 32 . Indeed, VLMs - even augmented with tools and harnesses - might be simply incapable of solving certain tasks. In the above-discussed thought experiment of the elf, it is possible that even an extremely competent elf would be unable to write the correct program without precisely measuring distances involved, or without a sense of touch, which could be impossible for a VLM operating just from robot cameras without any additional sensors. However, several research works have demonstrated that such models are able to leverage tools or ML (e.g. Eureka 33 , DrEureka 34 ) to perform dexterous behavior, so perhaps this could be a feasible path forward.</p>\n<h3 id=\"reliability-and-safety-on-real-hardware\">Reliability and safety on real hardware</h3>\n<p>There is much evidence that current models are not particularly safe or reliable at executing useful behavior on real robots. Robodojo results show that Astra solves their tasks with an average success rate of 22.48%, which is far from reliable completion. Moreover, the RoboDojo team had to halt their real-robot campaign for safety after Astra repeatedly issued physically unreasonable or unsafe actions, including incidents that damaged hardware. Robocurve found that several recent VLMs will execute harmful and dangerous tasks on real hardware without refusal. Model alignment remains a challenging problem even for disembodied VLMs : ensuring models will safely execute actions on hardware could be even more challenging.</p>\n<h3 id=\"speed-and-cost-in-deployment\">Speed and cost in deployment</h3>\n<p>One issue noted as part of all the recent results is speed. Frontier VLMs take on the order of seconds (at best) to produce a turn of output. However, some tasks - such as walking even a quadruped robot - require commands to be output at a much higher frequency (e.g. 50 Hz 35 ). This control frequency is not feasible for modern large frontier VLMs 36 , and it is unclear that it will ever be. A related issue is the current paradigm for querying large models: it may not be feasible for every robot to assume constant connection to a large centralized model server simply because wifi might have too much latency or too little bandwidth to be reliable.</p>","headings":[{"level":1,"text":"What are the new results?","id":"what-are-the-new-results"},{"level":1,"text":"Why are these results exciting?","id":"why-are-these-results-exciting"},{"level":2,"text":"A related thought experiment","id":"a-related-thought-experiment"},{"level":1,"text":"What’s missing?","id":"what-s-missing"},{"level":3,"text":"Thorough experimentation and (strong) evidence of generalization","id":"thorough-experimentation-and-strong-evidence-of-generalization"},{"level":3,"text":"Task complexity","id":"task-complexity"},{"level":3,"text":"Reliability and safety on real hardware","id":"reliability-and-safety-on-real-hardware"},{"level":3,"text":"Speed and cost in deployment","id":"speed-and-cost-in-deployment"}]}}