---
title: "Teaching a World Model to Play Pokémon"
slug: teaching-a-world-model-to-play-pokemon
url: https://listedarticles.com/articles/teaching-a-world-model-to-play-pokemon
canonical_url: https://nostalgia.dev/posts/teaching-a-world-model-to-play-pokemon/
content_type: tutorial
language: en
published_at: 2026-09-27T00:00:00.000Z
updated_at: 2026-09-27T12:14:35.951Z
author: "stmonty"
author_url: https://nostalgia.dev
authored_by: human
publisher: "nostalgia.dev"
publisher_url: https://nostalgia.dev
topics: ["AI", "Machine Learning", "Research", "Programming", "Education"]
license: all-rights-reserved
word_count: 574
reading_minutes: 2
citation: "stmonty, nostalgia.dev. \"Teaching a World Model to Play Pokémon.\" 27 Sept 2026. https://nostalgia.dev/posts/teaching-a-world-model-to-play-pokemon/ (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Teaching a World Model to Play Pokémon

> Training a JEPA-style LeWorldModel on Pokémon Red screenshots and button presses, then using CEM latent planning to select a starter—why latent collapse, SIGReg, and rollout fine-tuning mattered, and 52/100 plans succeeding after tuning.

# Teaching a World Model to Play Pokémon

*stmonty — nostalgia.dev*

I think some of the most fascinating work happening in the field of Artificial Intelligence surrounds world models. One that caught my eye was LeWorldModel: small enough to train locally on an RTX 3080 Ti, with a simpler design than many previous JEPA architectures.

San Francisco was covered in Pokémon memorabilia for the World Championships. A good “world” for a world model seemed to be the 1996 game that started it all: Pokémon Red.

## The Plan

The ambitious goal was to defeat Professor Oak’s grandson. Narrowed down: from a saved state in Oak’s Lab, select any of the three starters (Bulbasaur, Charmander, or Squirtle).

From that position, pressing A twelve times is enough — but can a world model learn that, rather than wandering forever or cancelling with B?

## What Even is a World Model?

A world model starts from an observation and an action that produces a new observation. The goal is to learn:

\[ o_{t+1} \approx F(o_t, a_t). \]

Unlike reinforcement learning, the world model is reward-free: it learns state dynamics without a reward signal.

### From screenshots to embeddings

The model predicts in latent/embedding space. An encoder \(E\) turns a screenshot into an embedding; a predictor \(P\) guesses the next embedding given an action:

\[ z_t = E(x_t),\qquad \hat z_{t+1}=P(z_t,a_t),\qquad z_{t+1}=E(x_{t+1}). \]

Training minimizes mean-squared error between predicted and true embeddings — but naive prediction loss invites **latent collapse** (every screenshot embeds to the same point).

### SIGReg: keeping the embedding space useful

LeWorldModel uses SIGReg (Sketched Isotropic Gaussian Regularization) so embeddings across a batch resemble a standard isotropic Gaussian. Random unit directions project embeddings; Epps–Pulley characteristic-function tests penalize non-Gaussian projections. Final loss:

\[ \mathcal L = \mathcal L_{\mathrm{pred}} + 0.1\,\operatorname{SIGReg}(Z). \]

## The Pokémon Training Data

42,382 grayscale frames in 1,009 short trajectories — some scripted routes to a starter, some noisy/random. Messy trajectories matter so the planner gets predictions for bad button sequences too. No success labels: the goal enters only at planning time.

A linear probe on frozen embeddings could tell whether the player had a Pokémon. The predictor beat a copy-baseline, and wrong buttons made predictions worse.

## Planning in the Learned World

Goal embeddings \(\mathcal G\) come from successful Bulbasaur/Charmander/Squirtle selections. Candidate 14-button plans are rolled out in latent space and scored by minimum distance to any goal:

\[ J(a_{0:H-1})=\min_{1\leq k\leq H}\;\min_{z_g\in\mathcal G}\frac{1}{D}\|\hat z_k-z_g\|_2^2,\qquad H=14. \]

Cross-entropy method (CEM) samples 512 plans/round, keeps the 64 best, and updates button distributions for several rounds, then runs the best plan in the emulator.

## Why the First Plan Failed — and Rollout Fine-Tuning

The first search looked good to the model but failed in the emulator: planning compounds prediction error with no screenshot reset. Fine-tuning the predictor on multi-step rollouts (feeding its own predictions forward) slowed error growth. After fine-tuning, CEM found a sequence that selected Squirtle (not the trivial twelve A presses). Across 100 random seeds: **52/100** plans acquired a starter (vs 0 random, 1 untrained predictor).

## The End…?

The final end-to-end model was ~12.5M parameters. Beating Blue remains future work; for now, Squirtle. Implementation: lePokeRed. Based on LeWorldModel (Maes et al., 2026, arXiv:2603.19312).

*Source: [nostalgia.dev/posts/teaching-a-world-model-to-play-pokemon](https://nostalgia.dev/posts/teaching-a-world-model-to-play-pokemon/)*
