I think some of the most fascinating work happening in the field of Artificial Intelligence surrounds world models. One that caught my eye was LeWorldModel: small enough to train locally on an RTX 3080 Ti, with a simpler design than many previous JEPA architectures.
San Francisco was covered in Pokémon memorabilia for the World Championships. A good “world” for a world model seemed to be the 1996 game that started it all: Pokémon Red.
The ambitious goal was to defeat Professor Oak’s grandson. Narrowed down: from a saved state in Oak’s Lab, select any of the three starters (Bulbasaur, Charmander, or Squirtle).
From that position, pressing A twelve times is enough — but can a world model learn that, rather than wandering forever or cancelling with B?
What Even is a World Model?
A world model starts from an observation and an action that produces a new observation. The goal is to learn:
[ o_{t+1} \approx F(o_t, a_t). ]
Unlike reinforcement learning, the world model is reward-free: it learns state dynamics without a reward signal.
From screenshots to embeddings
The model predicts in latent/embedding space. An encoder (E) turns a screenshot into an embedding; a predictor (P) guesses the next embedding given an action:
[ z_t = E(x_t),\qquad \hat z_{t+1}=P(z_t,a_t),\qquad z_{t+1}=E(x_{t+1}). ]
Training minimizes mean-squared error between predicted and true embeddings — but naive prediction loss invites latent collapse (every screenshot embeds to the same point).
SIGReg: keeping the embedding space useful
LeWorldModel uses SIGReg (Sketched Isotropic Gaussian Regularization) so embeddings across a batch resemble a standard isotropic Gaussian. Random unit directions project embeddings; Epps–Pulley characteristic-function tests penalize non-Gaussian projections. Final loss:
[ \mathcal L = \mathcal L_{\mathrm{pred}} + 0.1\,\operatorname{SIGReg}(Z). ]
The Pokémon Training Data
42,382 grayscale frames in 1,009 short trajectories — some scripted routes to a starter, some noisy/random. Messy trajectories matter so the planner gets predictions for bad button sequences too. No success labels: the goal enters only at planning time.
A linear probe on frozen embeddings could tell whether the player had a Pokémon. The predictor beat a copy-baseline, and wrong buttons made predictions worse.
Planning in the Learned World
Goal embeddings (\mathcal G) come from successful Bulbasaur/Charmander/Squirtle selections. Candidate 14-button plans are rolled out in latent space and scored by minimum distance to any goal:
[ J(a_{0:H-1})=\min_{1\leq k\leq H}\;\min_{z_g\in\mathcal G}\frac{1}{D}|\hat z_k-z_g|_2^2,\qquad H=14. ]
Cross-entropy method (CEM) samples 512 plans/round, keeps the 64 best, and updates button distributions for several rounds, then runs the best plan in the emulator.
Why the First Plan Failed — and Rollout Fine-Tuning
The first search looked good to the model but failed in the emulator: planning compounds prediction error with no screenshot reset. Fine-tuning the predictor on multi-step rollouts (feeding its own predictions forward) slowed error growth. After fine-tuning, CEM found a sequence that selected Squirtle (not the trivial twelve A presses). Across 100 random seeds: 52/100 plans acquired a starter (vs 0 random, 1 untrained predictor).
The End…?
The final end-to-end model was ~12.5M parameters. Beating Blue remains future work; for now, Squirtle. Implementation: lePokeRed. Based on LeWorldModel (Maes et al., 2026, arXiv:2603.19312).
Source: nostalgia.dev/posts/teaching-a-world-model-to-play-pokemon