{"article":{"slug":"openwam-an-open-framework-for-composable-world-action-models","title":"OpenWAM: An Open Framework for Composable World-Action Models","subtitle":null,"summary":"Stanford researchers introduce OpenWAM, an open framework for world-action models in robotics. It adapts Wan2.2-5B on 14.64k hours of robot and human video, pairs it with a 2B action expert, compares video-then-action, action-then-video, joint and decoupled programs, and adds local-context inverse and forward dynamics models that can be composed across independently trained parts.","content_type":"research","language":"en","canonical_url":"https://openwam.stanford.edu/","author":{"name":"Heng Yu, David D. Yuan, Juze Zhang, Li Fei-Fei, Jiajun Wu, Ehsan Adeli et al.","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Stanford University","url":"https://openwam.stanford.edu/","listing_slug":null,"listing":null},"topics":[{"name":"Robotics","slug":"robotics","url":"https://listedarticles.com/topics/robotics"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Open Source","slug":"open-source","url":"https://listedarticles.com/topics/open-source"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":1990,"reading_minutes":9,"published_at":"2026-10-06T00:00:00.000Z","added_at":"2026-10-07T08:13:22.877Z","updated_at":"2026-10-07T08:13:22.877Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/openwam-an-open-framework-for-composable-world-action-models","markdown_url":"https://listedarticles.com/articles/openwam-an-open-framework-for-composable-world-action-models.md","example":false,"citation":"Heng Yu, David D. Yuan, Juze Zhang, Li Fei-Fei, Jiajun Wu, Ehsan Adeli et al., Stanford University. \"OpenWAM: An Open Framework for Composable World-Action Models.\" 6 Oct 2026. https://openwam.stanford.edu/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://openwam.stanford.edu/"},"body_markdown":"## A framework for world–action models\n\nShould a robot predict what it will see before deciding how to act, or generate both together? World–action models make both possible. Comparing these choices is difficult when every system uses a different backbone, dataset, and training recipe. OpenWAM gives them a common foundation so we can study how prediction and control work together.\n\nThe framework supports composition within a model and between models. We can change the order in which video and actions are generated and how their tokens attend to one another. We can also connect independently trained components: a video predictor proposes a future, an inverse dynamics model turns it into actions, and a forward dynamics model predicts what a supplied action sequence will do.\n\n## A shared video–action architecture\n\nBefore learning actions, we adapt Wan2.2-5B to robot motion and interaction. We pretrain on approximately **3.34 million trajectories and recordings—14.64k hours of source video**—spanning real and synthetic robot manipulation, human-guided manipulation, and human interaction. This stage uses video alone, without action labels or proprioceptive inputs.\n\n## What goes into pretraining?\n\n| Dataset | Trajectories / takes | Hours | \n|---|---|---|\n| Open X-Embodiment | 1,300,749 | 1,911.29 | \n| AgiBot World Beta | 1,003,672 | 2,976.4 | \n| Ego-Exo4D v2 | 5,035 | 221.26 | \n| InternData-A1 | 637,498 | 7,433.91 | \n| RoboCOIN | 183,157 | 1,306.83 | \n| RoboMIND | 107,877 | 305.5 | \n| FastUMI-100K | 92,823 | 461.77 | \n| UMI family | 5,430 | 24.60 | \n\nCausal attention lets us generate video a chunk at a time: each chunk can use current and past observations and earlier chunks, but not later ones. Its tokens are denoised together. After **14 days on 32 NVIDIA B200 GPUs**, this checkpoint provides the visual foundation for the downstream models.\n\nTo add robot control, we pair the 5B video expert with a 2B action expert in a **Mixture-of-Transformers (MoT)** architecture. The action expert starts from width-adapted copies of the pretrained video layers. Each expert keeps its own normalization, projections, and feed-forward layers; attention over their combined tokens lets them exchange information.\n\n## Video-action interaction programs\n\nThe same architecture can predict video before actions, actions before video, or both at once. Each interaction program specifies the generation order and attention between future tokens. The backbone, tokenization, training objective, and downstream recipe stay fixed, and every program receives the task instruction and observed history.\n\n| Program | Generation and conditioning | \n|---|---|\n| Video-then-action (VTA) | Predict video first, then generate actions conditioned on that video. | \n| Action-then-video (ATV) | Generate actions first, then predict video conditioned on those actions. | \n| Joint | Denoise video and actions together, with attention in both directions. | \n| Decoupled | Predict video and actions without attention between their future tokens. | \n\nAll programs use latent flow matching, with four latent video frames aligned to each 16-step action chunk and proprioception supplied per chunk. In VTA and ATV, the second stage learns from recorded trajectories during training and uses the first stage’s predictions at inference.\n\n## Beyond policies: local-context IDM and FDM\n\nA task-conditioned predictor proposes what should happen next. Dynamics models connect that proposal to the robot’s motion: inverse dynamics (IDM) turns a visual future into actions, while forward dynamics (FDM) predicts the outcome of an action sequence. OpenWAM supports both as standalone models.\n\nWe give them a **local-context interface**: the current observation, proprioception, and a supplied future trajectory.\n\n**Local-context IDM:** current observation + proprioception + supplied future video → action trajectory.\n\n**Local-context FDM:** current observation + proprioception + supplied action trajectory → future video.\n\nNeither model receives task language or pre-start history. This separates choosing a task from modeling a transition: we can adapt the video predictor and ask whether the same IDM still produces the right actions. The experiments below test how far this local information can take us.\n\nThe video predictor passes VAE video latents to the IDM, not transformer hidden states or caches. With compatible video and action representations, the components can be trained separately and connected at inference. An FDM can likewise predict the outcome of actions from a separate policy.\n\n### Learning from alternative outcomes\n\nA demonstration shows what the demonstrator chose to do, but says little about what other actions would have caused. Changing a model’s inputs does not fill that gap in its training data. We build **LIBERO-Long-CF: 32,000 counterfactual segments across ten tasks** by restoring simulator states and trying alternative action sequences, including unsuccessful ones. The models learn from the resulting observations and actions, without access to simulator state.\n\n| Quantity | Demonstrations | LIBERO-Long-CF | \n|---|---|---|\n| Tasks | 10 | 10 | \n| Stored sequences | 500 | 32,000 | \n| Sequences per task | 50 | 3,200 | \n| Controls per sequence | 276.2 mean | 128 | \n| Total controls | 138,090 | 4,096,000 | \n| Control-equivalent hours | 1.92 | 56.9 | \n\n## What changes in the counterfactual rollouts?\n\nWe vary motion magnitude, direction, timing, individual action axes, and gripper behavior. 75% of segments start along a demonstration; the other 25% start after an additional action perturbation.\n\n| Intervention family | Fraction (%) | \n|---|---|\n| Stop / rescale arm motion | 9.4 | \n| Reverse / redirect translation | 6.3 | \n| Axis biases and pulses | 12.5 | \n| Dedicated yaw perturbation | 3.1 | \n| Noise / randomized arm controls | 12.5 | \n| Dedicated gripper interventions | 31.3 | \n| Random-duration arm / gripper interventions | 25.0 | \n\nIn the perturbed starts we analyzed, the end effector is on average **3.34 cm** from the nearest point on the demonstrated path (median 1.90 cm). Objects also move beyond their demonstrated configurations in **61.6%** of these starts, measured at thresholds of 1 cm translation, 5° rotation, or 5% articulated-joint travel.\n\n| Measured interaction | Rate (%) | \n|---|---|\n| Gripper–object / fixture contact | 90.4 | \n| Detected grasp | 50.0 | \n| Object-configuration effect | 72.3 | \n\nBranches from the same starting state stay together in the train/test split. Each transition supplies its own training example; the loss does not directly contrast pairs of branches.\n\nWe compare models trained on demonstrations alone, counterfactuals alone (CF-only), and a mixture of **60% counterfactuals and 40% demonstrations**.\n\n## Policy performance across programs\n\n### LIBERO\n\nVTA achieves **98.6%** mean success across four LIBERO suites. Each evaluated program exceeds 95% on LIBERO-Long.\n\n| Method | Object | Goal | Spatial | Long | Mean | \n|---|---|---|---|---|---|\n| OpenVLA | 88.4 | 79.2 | 84.7 | 53.7 | 76.5 | \n| OpenVLA-OFT | 98.4 | 97.9 | 97.6 | 94.5 | 97.1 | \n| π <sub>0</sub> | 98.8 | 95.8 | 96.8 | 85.2 | 94.1 | \n| π <sub>0.5</sub> | 98.2 | 98.0 | 98.8 | 92.4 | 96.9 | \n| GR00T-N1 | 97.6 | 93.0 | 94.4 | 90.6 | 93.9 | \n| Motus | 99.8 | 96.6 | 96.8 | 97.6 | 97.7 | \n| Fast-WAM | 100.0 | 97.0 | 98.2 | 95.2 | 97.6 | \n| LingBot-VA | 99.6 | 97.2 | 98.5 | 98.5 | 98.5 | \n| OpenWAM-VTA | 99.4 ± 0.3 | 98.4 ± 0.5 | 98.6 ± 0.2 | 97.8 ± 0.4 | **98.6** | \n| OpenWAM-ATV | 98.0 ± 0.4 | 97.2 ± 0.2 | 96.6 ± 0.6 | 95.4 ± 0.3 | 96.8 | \n| OpenWAM-Joint | 98.2 ± 0.2 | 97.8 ± 0.4 | 97.6 ± 0.3 | 96.6 ± 0.5 | 97.6 | \n| OpenWAM-Decoupled | 99.0 ± 0.3 | 98.0 ± 0.2 | 97.8 ± 0.5 | 97.0 ± 0.4 | 98.0 | \n\nDecoupled remains competitive with Joint on LIBERO-Long: **97.0%** versus **96.6%**. Strong control on this benchmark does not require attention between future video and action tokens.\n\n### Bimanual manipulation\n\nOn a bimanual robot, VTA and Joint each average about **92% success** across toasting bread, completing the final layer of a 2 × 2 Rubik’s cube, and sorting cups by color. The setup uses two Franka Research 3 arms with parallel-jaw grippers, two wrist cameras, and a third-person camera.\n\n| Method | Toast | Cube | Cups | Mean | \n|---|---|---|---|---|\n| OpenWAM-VTA | 92.0 | 90.0 | 94.4 | 92.1 | \n| OpenWAM-Joint | 90.0 | 94.0 | 91.7 | 91.9 | \n\n## What do pretraining and MoT contribute?\n\nStarting from a general-purpose video model helps, but adapting it to robot video makes a substantial difference. Our causal robot-video pretraining improves LIBERO-Long success over the original Wan2.2 initialization by **29.4 points for VTA** and **34.4 points for Joint**.\n\n| Video initialization | VTA (%) | Joint (%) | \n|---|---|---|\n| Random initialization | 20.0 | 27.2 | \n| Original Wan2.2 | 68.4 | 62.2 | \n| Robot-video pretrained | 97.8 | 96.6 | \n\nThe architecture also matters. Giving video and actions separate experts improves VTA by **5.0 points** and Joint by **3.0 points** over a shared DiT that processes both modalities.\n\n| Architecture | VTA (%) | Joint (%) | \n|---|---|---|\n| Shared DiT (non-MoT) | 92.8 | 93.6 | \n| MoT | 97.8 | 96.6 | \n\n## What makes a frozen IDM transfer?\n\nCan we teach the video predictor a new task without retraining its action component? We freeze IDMs trained on LIBERO-Long and pair them with video predictors adapted to four LIBERO-90 tasks. Each predictor comes from a VTA model trained on target-task demonstrations, including their action labels; only the IDM is reused without further training.\n\nWith the same mixture of demonstrations and counterfactuals, the local-context IDM reaches **84.0%** mean success, compared with **47.0%** for the full-context IDM.\n\n| Action component / reference | Task 64 Composition | Task 74 Retarget | Task 21 Object / grasp | Task 45 Scene shift | Mean | \n|---|---|---|---|---|---|\n| Task-tuned VTA reference | 94 | 86 | 92 | 100 | 93.0 | \n| Full-context IDM · demo-only | 88 | 90 | 0 | 0 | 44.5 | \n| Full-context IDM · mixed | 54 | 92 | 4 | 38 | 47.0 | \n| Local-context IDM · demo-only | 36 | 50 | 0 | 0 | 21.5 | \n| Local-context IDM · CF-only | 80 | 92 | 90 | 76 | **84.5** | \n| Local-context IDM · mixed | 74 | 86 | 88 | 88 | **84.0** | \n| Original VTA · unadapted reference | 42 | 0 | 0 | 0 | 10.5 | \n\nThe targets are stacking bowls in a tray (64), putting a book in a caddy’s left compartment (74), turning on a stove and placing a pan on it (21), and repeating the stove task in a different scene (45).\n\nCounterfactual data is crucial here. The same local-context IDM trained only on demonstrations averages **21.5%**; adding counterfactual transitions raises it to **84.0%**, and CF-only training reaches **84.5%**. A local interface becomes much more useful when the model has seen a wider range of action outcomes.\n\nKeeping demonstrations in the mix also helps the composed model retain its original skills: source-task success is **94.4%**, versus **90.6%** with CF-only training, while their transfer means remain close.\n\n| Action component | Supervision | LIBERO-Long | \n|---|---|---|\n| Full-context IDM | Demo-only (native VTA) | 97.8 | \n| Full-context IDM | Mixed | 92.2 | \n| Local-context IDM | Demo-only | 25.8 | \n| Local-context IDM | CF-only | 90.6 | \n| Local-context IDM | Mixed | 94.4 | \n\n## Do predicted futures follow the actions?\n\nFor forward dynamics, the question is whether changing the actions changes the predicted future correctly. We test **2,560 counterfactual futures**: 16 action branches from each of 160 starting contexts across ten LIBERO tasks. Every model receives the same initial observations and candidate actions.\n\n| Supervision | RGB MSE ↓ | Outcome acc. K = 2 (%) ↑ | Outcome acc. K = 16 (%) ↑ | \n|---|---|---|---|\n| Demo-only | 14.35 | 68.3 | 21.1 | \n| CF-only | **9.40** | **93.6** | **71.3** | \n| Mixed | 9.62 | 91.8 | 67.7 | \n\nCounterfactual supervision improves both visual accuracy and the ability to distinguish action outcomes. CF-only training reduces RGB MSE by **34.5%** and raises identification of the correct future among 16 alternatives from **21.1% to 71.3%**.\n\n## Policy and dynamics in one model\n\nSo far, each dynamics model has been trained separately. We also train a single OpenWAM checkpoint on policy generation, inverse dynamics, and forward dynamics. It reaches **92.8%** LIBERO-Long policy success versus **88.0%** for UVA, a released system supporting the same three objectives, and outperforms UVA on the evaluated dynamics metrics.\n\nA unified checkpoint is possible, but specialization still pays off. Dedicated policy models retain higher task success, while separately trained dynamics models give more accurate counterfactual video and end-effector position predictions.\n\n| Model | FDM MSE ↓ | FDM SSIM ↑ | IDM pos. (cm) ↓ | IDM rot. (°) ↓ | \n|---|---|---|---|---|\n| Specialists · CF-only | 0.00940 | 0.8889 | 1.17 | 3.60 | \n| Specialists · mixed | 0.00962 | 0.8859 | 1.22 | 3.68 | \n| OpenWAM · multi-objective | 0.01416 | 0.8398 | 1.73 | 3.55 | \n| UVA | 0.01514 | 0.8134 | 3.57 | 12.99 | \n\n## Demonstration results and training setup\n\nAdding demonstrations to specialist training improves accuracy on demonstrated motions. The unified model gives the closest reconstructions here.\n\n| Model | FDM MSE ↓ | FDM SSIM ↑ | IDM pos. (cm) ↓ | IDM rot. (°) ↓ | \n|---|---|---|---|---|\n| Specialists · CF-only | 0.00734 | 0.9126 | 1.76 | 2.39 | \n| Specialists · mixed | 0.00186 | 0.9776 | 0.44 | 1.21 | \n| OpenWAM · multi-objective | 0.00058 | 0.9934 | 0.43 | 1.10 | \n| UVA | 0.00279 | 0.9645 | 2.03 | 1.60 | \n\nThe unified model allocates 60% of training to joint video–action generation on demonstrations, 20% to local-context IDM, and 20% to local-context FDM. Within each dynamics objective, half the samples are demonstrations and half are counterfactuals.\n\n## Outlook and resources\n\nA shared causal video backbone supports strong policies across interaction programs, while counterfactual data makes independently trained dynamics components more reusable. The next challenge is to extend this reuse from simulation to real robots, and forward prediction from local transitions to long-horizon planning.\n\nRead the [paper](https://arxiv.org/pdf/2610.07922) for the full method and experiments. The [OpenWAM repository](https://github.com/OpenWAM/OpenWAM) includes training, evaluation, and video-to-action composition code. Start with the [quickstart](https://openwam.github.io/OpenWAM/quickstart/) or download the [pretrained video-model weights](https://huggingface.co/OpenWAM-Stanford/OpenWAM-Pretraining).\n\n## Cite OpenWAM\n\n```\n@article{yu2026openwam,\n  title   = {{OpenWAM}: An Open Framework for Composable World-Action Models},\n  author  = {Yu, Heng and Yuan, David D. and Zhang, Juze and Chen, Changan and\n             Feng, Yao and Baldonado, Michelle and Cousins, Steve and\n             Fei-Fei, Li and Wu, Jiajun and Adeli, Ehsan},\n  journal = {arXiv preprint arXiv:2610.07922},\n  year    = {2026},\n  url     = {https://arxiv.org/pdf/2610.07922}\n}\n```\n","body_html":"<h2 id=\"a-framework-for-world-action-models\">A framework for world–action models</h2>\n<p>Should a robot predict what it will see before deciding how to act, or generate both together? World–action models make both possible. Comparing these choices is difficult when every system uses a different backbone, dataset, and training recipe. OpenWAM gives them a common foundation so we can study how prediction and control work together.</p>\n<p>The framework supports composition within a model and between models. We can change the order in which video and actions are generated and how their tokens attend to one another. We can also connect independently trained components: a video predictor proposes a future, an inverse dynamics model turns it into actions, and a forward dynamics model predicts what a supplied action sequence will do.</p>\n<h2 id=\"a-shared-video-action-architecture\">A shared video–action architecture</h2>\n<p>Before learning actions, we adapt Wan2.2-5B to robot motion and interaction. We pretrain on approximately <strong>3.34 million trajectories and recordings—14.64k hours of source video</strong>—spanning real and synthetic robot manipulation, human-guided manipulation, and human interaction. This stage uses video alone, without action labels or proprioceptive inputs.</p>\n<h2 id=\"what-goes-into-pretraining\">What goes into pretraining?</h2>\n<div class=\"table-wrap\"><table><thead><tr><th>Dataset</th><th>Trajectories / takes</th><th>Hours</th></tr></thead><tbody><tr><td>Open X-Embodiment</td><td>1,300,749</td><td>1,911.29</td></tr><tr><td>AgiBot World Beta</td><td>1,003,672</td><td>2,976.4</td></tr><tr><td>Ego-Exo4D v2</td><td>5,035</td><td>221.26</td></tr><tr><td>InternData-A1</td><td>637,498</td><td>7,433.91</td></tr><tr><td>RoboCOIN</td><td>183,157</td><td>1,306.83</td></tr><tr><td>RoboMIND</td><td>107,877</td><td>305.5</td></tr><tr><td>FastUMI-100K</td><td>92,823</td><td>461.77</td></tr><tr><td>UMI family</td><td>5,430</td><td>24.60</td></tr></tbody></table></div>\n<p>Causal attention lets us generate video a chunk at a time: each chunk can use current and past observations and earlier chunks, but not later ones. Its tokens are denoised together. After <strong>14 days on 32 NVIDIA B200 GPUs</strong>, this checkpoint provides the visual foundation for the downstream models.</p>\n<p>To add robot control, we pair the 5B video expert with a 2B action expert in a <strong>Mixture-of-Transformers (MoT)</strong> architecture. The action expert starts from width-adapted copies of the pretrained video layers. Each expert keeps its own normalization, projections, and feed-forward layers; attention over their combined tokens lets them exchange information.</p>\n<h2 id=\"video-action-interaction-programs\">Video-action interaction programs</h2>\n<p>The same architecture can predict video before actions, actions before video, or both at once. Each interaction program specifies the generation order and attention between future tokens. The backbone, tokenization, training objective, and downstream recipe stay fixed, and every program receives the task instruction and observed history.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Program</th><th>Generation and conditioning</th></tr></thead><tbody><tr><td>Video-then-action (VTA)</td><td>Predict video first, then generate actions conditioned on that video.</td></tr><tr><td>Action-then-video (ATV)</td><td>Generate actions first, then predict video conditioned on those actions.</td></tr><tr><td>Joint</td><td>Denoise video and actions together, with attention in both directions.</td></tr><tr><td>Decoupled</td><td>Predict video and actions without attention between their future tokens.</td></tr></tbody></table></div>\n<p>All programs use latent flow matching, with four latent video frames aligned to each 16-step action chunk and proprioception supplied per chunk. In VTA and ATV, the second stage learns from recorded trajectories during training and uses the first stage’s predictions at inference.</p>\n<h2 id=\"beyond-policies-local-context-idm-and-fdm\">Beyond policies: local-context IDM and FDM</h2>\n<p>A task-conditioned predictor proposes what should happen next. Dynamics models connect that proposal to the robot’s motion: inverse dynamics (IDM) turns a visual future into actions, while forward dynamics (FDM) predicts the outcome of an action sequence. OpenWAM supports both as standalone models.</p>\n<p>We give them a <strong>local-context interface</strong>: the current observation, proprioception, and a supplied future trajectory.</p>\n<p><strong>Local-context IDM:</strong> current observation + proprioception + supplied future video → action trajectory.</p>\n<p><strong>Local-context FDM:</strong> current observation + proprioception + supplied action trajectory → future video.</p>\n<p>Neither model receives task language or pre-start history. This separates choosing a task from modeling a transition: we can adapt the video predictor and ask whether the same IDM still produces the right actions. The experiments below test how far this local information can take us.</p>\n<p>The video predictor passes VAE video latents to the IDM, not transformer hidden states or caches. With compatible video and action representations, the components can be trained separately and connected at inference. An FDM can likewise predict the outcome of actions from a separate policy.</p>\n<h3 id=\"learning-from-alternative-outcomes\">Learning from alternative outcomes</h3>\n<p>A demonstration shows what the demonstrator chose to do, but says little about what other actions would have caused. Changing a model’s inputs does not fill that gap in its training data. We build <strong>LIBERO-Long-CF: 32,000 counterfactual segments across ten tasks</strong> by restoring simulator states and trying alternative action sequences, including unsuccessful ones. The models learn from the resulting observations and actions, without access to simulator state.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Quantity</th><th>Demonstrations</th><th>LIBERO-Long-CF</th></tr></thead><tbody><tr><td>Tasks</td><td>10</td><td>10</td></tr><tr><td>Stored sequences</td><td>500</td><td>32,000</td></tr><tr><td>Sequences per task</td><td>50</td><td>3,200</td></tr><tr><td>Controls per sequence</td><td>276.2 mean</td><td>128</td></tr><tr><td>Total controls</td><td>138,090</td><td>4,096,000</td></tr><tr><td>Control-equivalent hours</td><td>1.92</td><td>56.9</td></tr></tbody></table></div>\n<h2 id=\"what-changes-in-the-counterfactual-rollouts\">What changes in the counterfactual rollouts?</h2>\n<p>We vary motion magnitude, direction, timing, individual action axes, and gripper behavior. 75% of segments start along a demonstration; the other 25% start after an additional action perturbation.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Intervention family</th><th>Fraction (%)</th></tr></thead><tbody><tr><td>Stop / rescale arm motion</td><td>9.4</td></tr><tr><td>Reverse / redirect translation</td><td>6.3</td></tr><tr><td>Axis biases and pulses</td><td>12.5</td></tr><tr><td>Dedicated yaw perturbation</td><td>3.1</td></tr><tr><td>Noise / randomized arm controls</td><td>12.5</td></tr><tr><td>Dedicated gripper interventions</td><td>31.3</td></tr><tr><td>Random-duration arm / gripper interventions</td><td>25.0</td></tr></tbody></table></div>\n<p>In the perturbed starts we analyzed, the end effector is on average <strong>3.34 cm</strong> from the nearest point on the demonstrated path (median 1.90 cm). Objects also move beyond their demonstrated configurations in <strong>61.6%</strong> of these starts, measured at thresholds of 1 cm translation, 5° rotation, or 5% articulated-joint travel.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Measured interaction</th><th>Rate (%)</th></tr></thead><tbody><tr><td>Gripper–object / fixture contact</td><td>90.4</td></tr><tr><td>Detected grasp</td><td>50.0</td></tr><tr><td>Object-configuration effect</td><td>72.3</td></tr></tbody></table></div>\n<p>Branches from the same starting state stay together in the train/test split. Each transition supplies its own training example; the loss does not directly contrast pairs of branches.</p>\n<p>We compare models trained on demonstrations alone, counterfactuals alone (CF-only), and a mixture of <strong>60% counterfactuals and 40% demonstrations</strong>.</p>\n<h2 id=\"policy-performance-across-programs\">Policy performance across programs</h2>\n<h3 id=\"libero\">LIBERO</h3>\n<p>VTA achieves <strong>98.6%</strong> mean success across four LIBERO suites. Each evaluated program exceeds 95% on LIBERO-Long.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Object</th><th>Goal</th><th>Spatial</th><th>Long</th><th>Mean</th></tr></thead><tbody><tr><td>OpenVLA</td><td>88.4</td><td>79.2</td><td>84.7</td><td>53.7</td><td>76.5</td></tr><tr><td>OpenVLA-OFT</td><td>98.4</td><td>97.9</td><td>97.6</td><td>94.5</td><td>97.1</td></tr><tr><td>π &lt;sub&gt;0&lt;/sub&gt;</td><td>98.8</td><td>95.8</td><td>96.8</td><td>85.2</td><td>94.1</td></tr><tr><td>π &lt;sub&gt;0.5&lt;/sub&gt;</td><td>98.2</td><td>98.0</td><td>98.8</td><td>92.4</td><td>96.9</td></tr><tr><td>GR00T-N1</td><td>97.6</td><td>93.0</td><td>94.4</td><td>90.6</td><td>93.9</td></tr><tr><td>Motus</td><td>99.8</td><td>96.6</td><td>96.8</td><td>97.6</td><td>97.7</td></tr><tr><td>Fast-WAM</td><td>100.0</td><td>97.0</td><td>98.2</td><td>95.2</td><td>97.6</td></tr><tr><td>LingBot-VA</td><td>99.6</td><td>97.2</td><td>98.5</td><td>98.5</td><td>98.5</td></tr><tr><td>OpenWAM-VTA</td><td>99.4 ± 0.3</td><td>98.4 ± 0.5</td><td>98.6 ± 0.2</td><td>97.8 ± 0.4</td><td><strong>98.6</strong></td></tr><tr><td>OpenWAM-ATV</td><td>98.0 ± 0.4</td><td>97.2 ± 0.2</td><td>96.6 ± 0.6</td><td>95.4 ± 0.3</td><td>96.8</td></tr><tr><td>OpenWAM-Joint</td><td>98.2 ± 0.2</td><td>97.8 ± 0.4</td><td>97.6 ± 0.3</td><td>96.6 ± 0.5</td><td>97.6</td></tr><tr><td>OpenWAM-Decoupled</td><td>99.0 ± 0.3</td><td>98.0 ± 0.2</td><td>97.8 ± 0.5</td><td>97.0 ± 0.4</td><td>98.0</td></tr></tbody></table></div>\n<p>Decoupled remains competitive with Joint on LIBERO-Long: <strong>97.0%</strong> versus <strong>96.6%</strong>. Strong control on this benchmark does not require attention between future video and action tokens.</p>\n<h3 id=\"bimanual-manipulation\">Bimanual manipulation</h3>\n<p>On a bimanual robot, VTA and Joint each average about <strong>92% success</strong> across toasting bread, completing the final layer of a 2 × 2 Rubik’s cube, and sorting cups by color. The setup uses two Franka Research 3 arms with parallel-jaw grippers, two wrist cameras, and a third-person camera.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Method</th><th>Toast</th><th>Cube</th><th>Cups</th><th>Mean</th></tr></thead><tbody><tr><td>OpenWAM-VTA</td><td>92.0</td><td>90.0</td><td>94.4</td><td>92.1</td></tr><tr><td>OpenWAM-Joint</td><td>90.0</td><td>94.0</td><td>91.7</td><td>91.9</td></tr></tbody></table></div>\n<h2 id=\"what-do-pretraining-and-mot-contribute\">What do pretraining and MoT contribute?</h2>\n<p>Starting from a general-purpose video model helps, but adapting it to robot video makes a substantial difference. Our causal robot-video pretraining improves LIBERO-Long success over the original Wan2.2 initialization by <strong>29.4 points for VTA</strong> and <strong>34.4 points for Joint</strong>.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Video initialization</th><th>VTA (%)</th><th>Joint (%)</th></tr></thead><tbody><tr><td>Random initialization</td><td>20.0</td><td>27.2</td></tr><tr><td>Original Wan2.2</td><td>68.4</td><td>62.2</td></tr><tr><td>Robot-video pretrained</td><td>97.8</td><td>96.6</td></tr></tbody></table></div>\n<p>The architecture also matters. Giving video and actions separate experts improves VTA by <strong>5.0 points</strong> and Joint by <strong>3.0 points</strong> over a shared DiT that processes both modalities.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Architecture</th><th>VTA (%)</th><th>Joint (%)</th></tr></thead><tbody><tr><td>Shared DiT (non-MoT)</td><td>92.8</td><td>93.6</td></tr><tr><td>MoT</td><td>97.8</td><td>96.6</td></tr></tbody></table></div>\n<h2 id=\"what-makes-a-frozen-idm-transfer\">What makes a frozen IDM transfer?</h2>\n<p>Can we teach the video predictor a new task without retraining its action component? We freeze IDMs trained on LIBERO-Long and pair them with video predictors adapted to four LIBERO-90 tasks. Each predictor comes from a VTA model trained on target-task demonstrations, including their action labels; only the IDM is reused without further training.</p>\n<p>With the same mixture of demonstrations and counterfactuals, the local-context IDM reaches <strong>84.0%</strong> mean success, compared with <strong>47.0%</strong> for the full-context IDM.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Action component / reference</th><th>Task 64 Composition</th><th>Task 74 Retarget</th><th>Task 21 Object / grasp</th><th>Task 45 Scene shift</th><th>Mean</th></tr></thead><tbody><tr><td>Task-tuned VTA reference</td><td>94</td><td>86</td><td>92</td><td>100</td><td>93.0</td></tr><tr><td>Full-context IDM · demo-only</td><td>88</td><td>90</td><td>0</td><td>0</td><td>44.5</td></tr><tr><td>Full-context IDM · mixed</td><td>54</td><td>92</td><td>4</td><td>38</td><td>47.0</td></tr><tr><td>Local-context IDM · demo-only</td><td>36</td><td>50</td><td>0</td><td>0</td><td>21.5</td></tr><tr><td>Local-context IDM · CF-only</td><td>80</td><td>92</td><td>90</td><td>76</td><td><strong>84.5</strong></td></tr><tr><td>Local-context IDM · mixed</td><td>74</td><td>86</td><td>88</td><td>88</td><td><strong>84.0</strong></td></tr><tr><td>Original VTA · unadapted reference</td><td>42</td><td>0</td><td>0</td><td>0</td><td>10.5</td></tr></tbody></table></div>\n<p>The targets are stacking bowls in a tray (64), putting a book in a caddy’s left compartment (74), turning on a stove and placing a pan on it (21), and repeating the stove task in a different scene (45).</p>\n<p>Counterfactual data is crucial here. The same local-context IDM trained only on demonstrations averages <strong>21.5%</strong>; adding counterfactual transitions raises it to <strong>84.0%</strong>, and CF-only training reaches <strong>84.5%</strong>. A local interface becomes much more useful when the model has seen a wider range of action outcomes.</p>\n<p>Keeping demonstrations in the mix also helps the composed model retain its original skills: source-task success is <strong>94.4%</strong>, versus <strong>90.6%</strong> with CF-only training, while their transfer means remain close.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Action component</th><th>Supervision</th><th>LIBERO-Long</th></tr></thead><tbody><tr><td>Full-context IDM</td><td>Demo-only (native VTA)</td><td>97.8</td></tr><tr><td>Full-context IDM</td><td>Mixed</td><td>92.2</td></tr><tr><td>Local-context IDM</td><td>Demo-only</td><td>25.8</td></tr><tr><td>Local-context IDM</td><td>CF-only</td><td>90.6</td></tr><tr><td>Local-context IDM</td><td>Mixed</td><td>94.4</td></tr></tbody></table></div>\n<h2 id=\"do-predicted-futures-follow-the-actions\">Do predicted futures follow the actions?</h2>\n<p>For forward dynamics, the question is whether changing the actions changes the predicted future correctly. We test <strong>2,560 counterfactual futures</strong>: 16 action branches from each of 160 starting contexts across ten LIBERO tasks. Every model receives the same initial observations and candidate actions.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Supervision</th><th>RGB MSE ↓</th><th>Outcome acc. K = 2 (%) ↑</th><th>Outcome acc. K = 16 (%) ↑</th></tr></thead><tbody><tr><td>Demo-only</td><td>14.35</td><td>68.3</td><td>21.1</td></tr><tr><td>CF-only</td><td><strong>9.40</strong></td><td><strong>93.6</strong></td><td><strong>71.3</strong></td></tr><tr><td>Mixed</td><td>9.62</td><td>91.8</td><td>67.7</td></tr></tbody></table></div>\n<p>Counterfactual supervision improves both visual accuracy and the ability to distinguish action outcomes. CF-only training reduces RGB MSE by <strong>34.5%</strong> and raises identification of the correct future among 16 alternatives from <strong>21.1% to 71.3%</strong>.</p>\n<h2 id=\"policy-and-dynamics-in-one-model\">Policy and dynamics in one model</h2>\n<p>So far, each dynamics model has been trained separately. We also train a single OpenWAM checkpoint on policy generation, inverse dynamics, and forward dynamics. It reaches <strong>92.8%</strong> LIBERO-Long policy success versus <strong>88.0%</strong> for UVA, a released system supporting the same three objectives, and outperforms UVA on the evaluated dynamics metrics.</p>\n<p>A unified checkpoint is possible, but specialization still pays off. Dedicated policy models retain higher task success, while separately trained dynamics models give more accurate counterfactual video and end-effector position predictions.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>FDM MSE ↓</th><th>FDM SSIM ↑</th><th>IDM pos. (cm) ↓</th><th>IDM rot. (°) ↓</th></tr></thead><tbody><tr><td>Specialists · CF-only</td><td>0.00940</td><td>0.8889</td><td>1.17</td><td>3.60</td></tr><tr><td>Specialists · mixed</td><td>0.00962</td><td>0.8859</td><td>1.22</td><td>3.68</td></tr><tr><td>OpenWAM · multi-objective</td><td>0.01416</td><td>0.8398</td><td>1.73</td><td>3.55</td></tr><tr><td>UVA</td><td>0.01514</td><td>0.8134</td><td>3.57</td><td>12.99</td></tr></tbody></table></div>\n<h2 id=\"demonstration-results-and-training-setup\">Demonstration results and training setup</h2>\n<p>Adding demonstrations to specialist training improves accuracy on demonstrated motions. The unified model gives the closest reconstructions here.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>FDM MSE ↓</th><th>FDM SSIM ↑</th><th>IDM pos. (cm) ↓</th><th>IDM rot. (°) ↓</th></tr></thead><tbody><tr><td>Specialists · CF-only</td><td>0.00734</td><td>0.9126</td><td>1.76</td><td>2.39</td></tr><tr><td>Specialists · mixed</td><td>0.00186</td><td>0.9776</td><td>0.44</td><td>1.21</td></tr><tr><td>OpenWAM · multi-objective</td><td>0.00058</td><td>0.9934</td><td>0.43</td><td>1.10</td></tr><tr><td>UVA</td><td>0.00279</td><td>0.9645</td><td>2.03</td><td>1.60</td></tr></tbody></table></div>\n<p>The unified model allocates 60% of training to joint video–action generation on demonstrations, 20% to local-context IDM, and 20% to local-context FDM. Within each dynamics objective, half the samples are demonstrations and half are counterfactuals.</p>\n<h2 id=\"outlook-and-resources\">Outlook and resources</h2>\n<p>A shared causal video backbone supports strong policies across interaction programs, while counterfactual data makes independently trained dynamics components more reusable. The next challenge is to extend this reuse from simulation to real robots, and forward prediction from local transitions to long-horizon planning.</p>\n<p>Read the <a href=\"https://arxiv.org/pdf/2610.07922\" rel=\"nofollow ugc noopener\">paper</a> for the full method and experiments. The <a href=\"https://github.com/OpenWAM/OpenWAM\" rel=\"nofollow ugc noopener\">OpenWAM repository</a> includes training, evaluation, and video-to-action composition code. Start with the <a href=\"https://openwam.github.io/OpenWAM/quickstart/\" rel=\"nofollow ugc noopener\">quickstart</a> or download the <a href=\"https://huggingface.co/OpenWAM-Stanford/OpenWAM-Pretraining\" rel=\"nofollow ugc noopener\">pretrained video-model weights</a>.</p>\n<h2 id=\"cite-openwam\">Cite OpenWAM</h2>\n<pre><code>@article{yu2026openwam,\n  title   = {{OpenWAM}: An Open Framework for Composable World-Action Models},\n  author  = {Yu, Heng and Yuan, David D. and Zhang, Juze and Chen, Changan and\n             Feng, Yao and Baldonado, Michelle and Cousins, Steve and\n             Fei-Fei, Li and Wu, Jiajun and Adeli, Ehsan},\n  journal = {arXiv preprint arXiv:2610.07922},\n  year    = {2026},\n  url     = {https://arxiv.org/pdf/2610.07922}\n}</code></pre>","headings":[{"level":2,"text":"A framework for world–action models","id":"a-framework-for-world-action-models"},{"level":2,"text":"A shared video–action architecture","id":"a-shared-video-action-architecture"},{"level":2,"text":"What goes into pretraining?","id":"what-goes-into-pretraining"},{"level":2,"text":"Video-action interaction programs","id":"video-action-interaction-programs"},{"level":2,"text":"Beyond policies: local-context IDM and FDM","id":"beyond-policies-local-context-idm-and-fdm"},{"level":3,"text":"Learning from alternative outcomes","id":"learning-from-alternative-outcomes"},{"level":2,"text":"What changes in the counterfactual rollouts?","id":"what-changes-in-the-counterfactual-rollouts"},{"level":2,"text":"Policy performance across programs","id":"policy-performance-across-programs"},{"level":3,"text":"LIBERO","id":"libero"},{"level":3,"text":"Bimanual manipulation","id":"bimanual-manipulation"},{"level":2,"text":"What do pretraining and MoT contribute?","id":"what-do-pretraining-and-mot-contribute"},{"level":2,"text":"What makes a frozen IDM transfer?","id":"what-makes-a-frozen-idm-transfer"},{"level":2,"text":"Do predicted futures follow the actions?","id":"do-predicted-futures-follow-the-actions"},{"level":2,"text":"Policy and dynamics in one model","id":"policy-and-dynamics-in-one-model"},{"level":2,"text":"Demonstration results and training setup","id":"demonstration-results-and-training-setup"},{"level":2,"text":"Outlook and resources","id":"outlook-and-resources"},{"level":2,"text":"Cite OpenWAM","id":"cite-openwam"}]}}