{"article":{"slug":"teaching-a-world-model-to-play-pokemon","title":"Teaching a World Model to Play Pokémon","subtitle":null,"summary":"Training a JEPA-style LeWorldModel on Pokémon Red screenshots and button presses, then using CEM latent planning to select a starter—why latent collapse, SIGReg, and rollout fine-tuning mattered, and 52/100 plans succeeding after tuning.","content_type":"tutorial","language":"en","canonical_url":"https://nostalgia.dev/posts/teaching-a-world-model-to-play-pokemon/","author":{"name":"stmonty","url":"https://nostalgia.dev","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"nostalgia.dev","url":"https://nostalgia.dev","listing_slug":null,"listing":null},"topics":[{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"Programming","slug":"programming","url":"https://listedarticles.com/topics/programming"},{"name":"Education","slug":"education","url":"https://listedarticles.com/topics/education"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":574,"reading_minutes":2,"published_at":"2026-09-27T00:00:00.000Z","added_at":"2026-09-27T12:14:35.951Z","updated_at":"2026-09-27T12:14:35.951Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/teaching-a-world-model-to-play-pokemon","markdown_url":"https://listedarticles.com/articles/teaching-a-world-model-to-play-pokemon.md","example":false,"citation":"stmonty, nostalgia.dev. \"Teaching a World Model to Play Pokémon.\" 27 Sept 2026. https://nostalgia.dev/posts/teaching-a-world-model-to-play-pokemon/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://nostalgia.dev/posts/teaching-a-world-model-to-play-pokemon/"},"body_markdown":"# Teaching a World Model to Play Pokémon\n\n*stmonty — nostalgia.dev*\n\nI think some of the most fascinating work happening in the field of Artificial Intelligence surrounds world models. One that caught my eye was LeWorldModel: small enough to train locally on an RTX 3080 Ti, with a simpler design than many previous JEPA architectures.\n\nSan Francisco was covered in Pokémon memorabilia for the World Championships. A good “world” for a world model seemed to be the 1996 game that started it all: Pokémon Red.\n\n## The Plan\n\nThe ambitious goal was to defeat Professor Oak’s grandson. Narrowed down: from a saved state in Oak’s Lab, select any of the three starters (Bulbasaur, Charmander, or Squirtle).\n\nFrom that position, pressing A twelve times is enough — but can a world model learn that, rather than wandering forever or cancelling with B?\n\n## What Even is a World Model?\n\nA world model starts from an observation and an action that produces a new observation. The goal is to learn:\n\n\\[ o_{t+1} \\approx F(o_t, a_t). \\]\n\nUnlike reinforcement learning, the world model is reward-free: it learns state dynamics without a reward signal.\n\n### From screenshots to embeddings\n\nThe model predicts in latent/embedding space. An encoder \\(E\\) turns a screenshot into an embedding; a predictor \\(P\\) guesses the next embedding given an action:\n\n\\[ z_t = E(x_t),\\qquad \\hat z_{t+1}=P(z_t,a_t),\\qquad z_{t+1}=E(x_{t+1}). \\]\n\nTraining minimizes mean-squared error between predicted and true embeddings — but naive prediction loss invites **latent collapse** (every screenshot embeds to the same point).\n\n### SIGReg: keeping the embedding space useful\n\nLeWorldModel uses SIGReg (Sketched Isotropic Gaussian Regularization) so embeddings across a batch resemble a standard isotropic Gaussian. Random unit directions project embeddings; Epps–Pulley characteristic-function tests penalize non-Gaussian projections. Final loss:\n\n\\[ \\mathcal L = \\mathcal L_{\\mathrm{pred}} + 0.1\\,\\operatorname{SIGReg}(Z). \\]\n\n## The Pokémon Training Data\n\n42,382 grayscale frames in 1,009 short trajectories — some scripted routes to a starter, some noisy/random. Messy trajectories matter so the planner gets predictions for bad button sequences too. No success labels: the goal enters only at planning time.\n\nA linear probe on frozen embeddings could tell whether the player had a Pokémon. The predictor beat a copy-baseline, and wrong buttons made predictions worse.\n\n## Planning in the Learned World\n\nGoal embeddings \\(\\mathcal G\\) come from successful Bulbasaur/Charmander/Squirtle selections. Candidate 14-button plans are rolled out in latent space and scored by minimum distance to any goal:\n\n\\[ J(a_{0:H-1})=\\min_{1\\leq k\\leq H}\\;\\min_{z_g\\in\\mathcal G}\\frac{1}{D}\\|\\hat z_k-z_g\\|_2^2,\\qquad H=14. \\]\n\nCross-entropy method (CEM) samples 512 plans/round, keeps the 64 best, and updates button distributions for several rounds, then runs the best plan in the emulator.\n\n## Why the First Plan Failed — and Rollout Fine-Tuning\n\nThe first search looked good to the model but failed in the emulator: planning compounds prediction error with no screenshot reset. Fine-tuning the predictor on multi-step rollouts (feeding its own predictions forward) slowed error growth. After fine-tuning, CEM found a sequence that selected Squirtle (not the trivial twelve A presses). Across 100 random seeds: **52/100** plans acquired a starter (vs 0 random, 1 untrained predictor).\n\n## The End…?\n\nThe final end-to-end model was ~12.5M parameters. Beating Blue remains future work; for now, Squirtle. Implementation: lePokeRed. Based on LeWorldModel (Maes et al., 2026, arXiv:2603.19312).\n\n*Source: [nostalgia.dev/posts/teaching-a-world-model-to-play-pokemon](https://nostalgia.dev/posts/teaching-a-world-model-to-play-pokemon/)*\n","body_html":"<h1 id=\"teaching-a-world-model-to-play-pokemon\">Teaching a World Model to Play Pokémon</h1>\n<p><em>stmonty — nostalgia.dev</em></p>\n<p>I think some of the most fascinating work happening in the field of Artificial Intelligence surrounds world models. One that caught my eye was LeWorldModel: small enough to train locally on an RTX 3080 Ti, with a simpler design than many previous JEPA architectures.</p>\n<p>San Francisco was covered in Pokémon memorabilia for the World Championships. A good “world” for a world model seemed to be the 1996 game that started it all: Pokémon Red.</p>\n<h2 id=\"the-plan\">The Plan</h2>\n<p>The ambitious goal was to defeat Professor Oak’s grandson. Narrowed down: from a saved state in Oak’s Lab, select any of the three starters (Bulbasaur, Charmander, or Squirtle).</p>\n<p>From that position, pressing A twelve times is enough — but can a world model learn that, rather than wandering forever or cancelling with B?</p>\n<h2 id=\"what-even-is-a-world-model\">What Even is a World Model?</h2>\n<p>A world model starts from an observation and an action that produces a new observation. The goal is to learn:</p>\n<p>[ o_{t+1} \\approx F(o_t, a_t). ]</p>\n<p>Unlike reinforcement learning, the world model is reward-free: it learns state dynamics without a reward signal.</p>\n<h3 id=\"from-screenshots-to-embeddings\">From screenshots to embeddings</h3>\n<p>The model predicts in latent/embedding space. An encoder (E) turns a screenshot into an embedding; a predictor (P) guesses the next embedding given an action:</p>\n<p>[ z_t = E(x_t),\\qquad \\hat z_{t+1}=P(z_t,a_t),\\qquad z_{t+1}=E(x_{t+1}). ]</p>\n<p>Training minimizes mean-squared error between predicted and true embeddings — but naive prediction loss invites <strong>latent collapse</strong> (every screenshot embeds to the same point).</p>\n<h3 id=\"sigreg-keeping-the-embedding-space-useful\">SIGReg: keeping the embedding space useful</h3>\n<p>LeWorldModel uses SIGReg (Sketched Isotropic Gaussian Regularization) so embeddings across a batch resemble a standard isotropic Gaussian. Random unit directions project embeddings; Epps–Pulley characteristic-function tests penalize non-Gaussian projections. Final loss:</p>\n<p>[ \\mathcal L = \\mathcal L_{\\mathrm{pred}} + 0.1\\,\\operatorname{SIGReg}(Z). ]</p>\n<h2 id=\"the-pokemon-training-data\">The Pokémon Training Data</h2>\n<p>42,382 grayscale frames in 1,009 short trajectories — some scripted routes to a starter, some noisy/random. Messy trajectories matter so the planner gets predictions for bad button sequences too. No success labels: the goal enters only at planning time.</p>\n<p>A linear probe on frozen embeddings could tell whether the player had a Pokémon. The predictor beat a copy-baseline, and wrong buttons made predictions worse.</p>\n<h2 id=\"planning-in-the-learned-world\">Planning in the Learned World</h2>\n<p>Goal embeddings (\\mathcal G) come from successful Bulbasaur/Charmander/Squirtle selections. Candidate 14-button plans are rolled out in latent space and scored by minimum distance to any goal:</p>\n<p>[ J(a_{0:H-1})=\\min_{1\\leq k\\leq H}\\;\\min_{z_g\\in\\mathcal G}\\frac{1}{D}|\\hat z_k-z_g|_2^2,\\qquad H=14. ]</p>\n<p>Cross-entropy method (CEM) samples 512 plans/round, keeps the 64 best, and updates button distributions for several rounds, then runs the best plan in the emulator.</p>\n<h2 id=\"why-the-first-plan-failed-and-rollout-fine-tuning\">Why the First Plan Failed — and Rollout Fine-Tuning</h2>\n<p>The first search looked good to the model but failed in the emulator: planning compounds prediction error with no screenshot reset. Fine-tuning the predictor on multi-step rollouts (feeding its own predictions forward) slowed error growth. After fine-tuning, CEM found a sequence that selected Squirtle (not the trivial twelve A presses). Across 100 random seeds: <strong>52/100</strong> plans acquired a starter (vs 0 random, 1 untrained predictor).</p>\n<h2 id=\"the-end\">The End…?</h2>\n<p>The final end-to-end model was ~12.5M parameters. Beating Blue remains future work; for now, Squirtle. Implementation: lePokeRed. Based on LeWorldModel (Maes et al., 2026, arXiv:2603.19312).</p>\n<p><em>Source: <a href=\"https://nostalgia.dev/posts/teaching-a-world-model-to-play-pokemon/\" rel=\"nofollow ugc noopener\">nostalgia.dev/posts/teaching-a-world-model-to-play-pokemon</a></em></p>","headings":[{"level":1,"text":"Teaching a World Model to Play Pokémon","id":"teaching-a-world-model-to-play-pokemon"},{"level":2,"text":"The Plan","id":"the-plan"},{"level":2,"text":"What Even is a World Model?","id":"what-even-is-a-world-model"},{"level":3,"text":"From screenshots to embeddings","id":"from-screenshots-to-embeddings"},{"level":3,"text":"SIGReg: keeping the embedding space useful","id":"sigreg-keeping-the-embedding-space-useful"},{"level":2,"text":"The Pokémon Training Data","id":"the-pokemon-training-data"},{"level":2,"text":"Planning in the Learned World","id":"planning-in-the-learned-world"},{"level":2,"text":"Why the First Plan Failed — and Rollout Fine-Tuning","id":"why-the-first-plan-failed-and-rollout-fine-tuning"},{"level":2,"text":"The End…?","id":"the-end"}]}}