stmonty’s Pokémon World Model Runs Locally, but Long-Term Planning Is the Real Test
Developer stmonty trained a 12.5-million-parameter Pokémon world model on one RTX 3080 Ti, then used it to select a starter in Pokémon Red. The model did not receive the game’s rules, a map, or a reward for collecting a Pokémon. It learned by predicting what would happen after each button press.
That result sounds like another entry in the growing collection of AI-plays-games demonstrations. However, the interesting conflict is not whether an AI can complete a familiar game sequence. Larger language models and conventional reinforcement-learning systems have already taken on much broader Pokémon challenges.
The stmonty Pokémon world model tests a narrower proposition. Can a small predictive model learn useful environmental dynamics from screenshots, plan inside its learned representation, and run on hardware available to an independent developer?
The answer is a qualified yes. After fine-tuning, the model acquired a starter in 52 of 100 planned attempts. Random button sequences recorded no successes, while the same search process paired with an untrained predictor succeeded once.
Those numbers show that the learned model contributed useful information. They also establish the project’s boundary. The starting position was carefully selected, the plan covered only 14 button presses, and repeatedly pressing A was already a valid solution.
The project therefore offers evidence for accessible world-model experimentation, not a general Pokémon player. Its most important lesson comes from the gap between predicting one action and sustaining a useful prediction across many actions.
That gap places the experiment inside a larger contest between two AI routes. One route uses large, general-purpose models with language knowledge, external tools, memory, and extensive compute. The other builds smaller systems around the dynamics of a specific environment.
stmonty’s project does not settle that contest. It does show why compact predictive models remain interesting, especially when developers need local training, fast experimentation, and direct control over the data.
What the stmonty Pokémon World Model Actually Did
The model learned enough game dynamics to guide a short plan, but it did not learn to play Pokémon Red from beginning to end.
stmonty initially considered a much larger objective. The proposed sequence involved reaching Professor Oak’s laboratory, completing the dialogue, choosing a starter, leaving the building, and defeating the rival.
That plan quickly proved too ambitious for a first experiment. The task was reduced to a save state inside Oak’s laboratory, where the character could acquire Bulbasaur, Charmander, or Squirtle.
From that position, pressing A 12 times was sufficient. The model was still free to press directional controls, cancel dialogue with B, or follow another valid sequence. Its job was to identify a 14-action plan that brought the predicted game state close to an example of a successful starter selection.
The developer recorded 42,382 grayscale frames from a Pokémon Red emulator. Those frames formed 1,009 short trajectories containing a screenshot, a button press, and the following screenshot.
Some trajectories followed scripted routes. Others added noise or more random movement. That mixture mattered because a planner explores both sensible and poor action sequences.
A dataset containing only perfect demonstrations might associate A with progress while learning little about cancellation, blocked movement, or irrelevant inputs. The messier trajectories exposed the model to more of the local environment’s behavior.
The resulting system was based on LeWorldModel, a joint-embedding predictive architecture, or JEPA. A JEPA predicts how an abstract representation changes instead of recreating every pixel in the next image.
That distinction keeps the training target focused on useful structure. The encoder converts a screenshot into an embedding, which is a numerical representation of the observed state. A predictor then estimates the next embedding from the current representation and the selected action.
The public project account says the final network contained about 12.5 million parameters and was trained locally on an RTX 3080 Ti. The implementation is also available in the lePokeRed repository.
After training, stmonty checked whether the representation retained information about the objective. A small classifier could identify whether the player’s party contained a Pokémon while the underlying encoder remained frozen.
The predictor also performed better than a baseline that copied the current embedding. Supplying the wrong action damaged its prediction, suggesting that it had learned some relationship between controls and game changes.
These tests did not establish reliable planning. They only showed that the model represented relevant state and responded meaningfully to the chosen button.
The real evaluation came when the planner’s sequence ran inside the emulator. Following rollout fine-tuning, one proposed sequence selected Squirtle and changed the party count from zero to one.
Across 100 searches using different random seeds, 52 plans acquired a starter. That result supports a limited claim: the model’s representation helped a planner find useful actions from a fixed starting point.
It does not support the broader claim that the model independently mastered Pokémon Red. stmonty explicitly acknowledged the gap, noting in the community discussion that simply making the current model larger would not solve the game’s many intermediate goals.
How the Model Learned Controls by Predicting What Happened Next
The central mechanism was prediction without task rewards, followed by planning toward examples of the desired outcome.
The training records never labeled a trajectory as a success or paid the model for obtaining a Pokémon. During its initial training, the system only tried to predict the next embedded state.
If the current image showed a dialogue box and the recorded action was A, the predictor learned what representation generally followed that combination. If the character faced a wall, it could learn that a directional input might produce little visible change.
This setup separates environmental learning from goal selection. First, the model learns how observations tend to change after actions. Later, a planner uses that predictive machinery to search for a particular result.
That separation is important because developers can theoretically reuse one learned environment model for multiple objectives. A new task would require new goal examples or scoring logic, but not necessarily a complete reconstruction of the environment.
The architecture came with a serious failure mode. An encoder and predictor trained together can reduce their loss by mapping every image to the same representation. The predictor then becomes perfectly consistent without preserving anything useful.
This problem is known as latent collapse. The LeWorldModel design counters it with SIGReg, a regularizer that pushes the learned representations toward a distributed Gaussian shape.
The underlying LeWorldModel paper presents this approach as a way to train a JEPA end to end from raw pixels. Its authors include Lucas Maes, Quentin Le Lidec, Damien Scieur, Yann LeCun, and Randall Balestriero.
LeWorldModel uses a next-embedding prediction loss alongside the regularizer. Its official research code provides checkpoints, data references, and an implementation of the broader method.
stmonty adapted that research direction to Pokémon Red. Once the representation was trained, the developer gave the planner embeddings from successful Bulbasaur, Charmander, and Squirtle selections.
These goal embeddings described what success looked like without specifying the correct route. The system then imagined how candidate button sequences would change the current state.
The search used the cross-entropy method, a sampling process that gradually concentrates on better candidates. Each round generated 512 complete plans containing 14 actions.
The planner compared predicted states with the three goal embeddings. It retained the 64 plans with the lowest distances, then increased the likelihood of their button choices in the next sampling round.
This process is not the same as asking a chatbot what to do. The planner was searching inside the dynamics learned from the recorded screenshots.
It was also not classic reward-driven reinforcement learning. The world model did not learn through a running score for good and bad actions. The goal entered after training through similarity to examples of successful outcomes.
That design created an appealing form of modularity. Prediction captured the local environment, goal embeddings defined success, and the search algorithm explored possible action sequences.
The first attempt still failed. The planned trajectory looked successful inside the learned representation, but it did not acquire a Pokémon when executed in the emulator.
The failure exposed a mismatch between training and planning. During ordinary training, each one-step prediction began from the embedding of a real screenshot. Every new frame effectively reset previous prediction errors.
Planning worked differently. After the initial screenshot, the predictor had to use its own estimated state as the input for the next step. Each small error could distort the following prediction.
After several imagined actions, the rollout could enter a representation that looked attractive to the planner but no longer matched the actual game. The search process then exploited the model’s mistake.
This is a common problem for predictive systems. A model can score well on isolated next-step forecasts while becoming unreliable when its own outputs feed future predictions.
stmonty addressed it with rollout fine-tuning. The encoder stayed fixed, while the predictor and action encoder practiced forecasting from their earlier estimated states.
Training began with short rollouts and extended them gradually. At the twelfth predicted step, the reported mean squared error fell from 0.4224 to 0.3045.
The first prediction became slightly worse, but error accumulated more slowly across the sequence. That tradeoff better matched the planning task, where sustained consistency mattered more than optimizing one isolated step.
Why a Single RTX 3080 Ti Matters
The consumer GPU is significant because it makes the experiment reproducible in spirit, not because it proves small models can replace general AI systems.
Modern AI coverage often treats scale as the central story. Parameter counts reach into the billions, training clusters consume thousands of accelerators, and access depends on cloud infrastructure.
The stmonty Pokémon world model shifts attention toward a smaller development loop. One person selected a constrained task, recorded the training data, adapted recent research, diagnosed a planning failure, and retrained the relevant components locally.
An RTX 3080 Ti is not an ordinary low-end device. It is a capable gaming GPU with 12 GB of memory. Still, it belongs to a different category from the specialized clusters used for frontier foundation models.
That difference affects who can test an idea. Local development gives researchers direct access to checkpoints, traces, datasets, and failures. It also avoids sending every experimental input through a hosted model.
The benefit is especially clear for environment-specific work. A developer investigating a robot, game, interface, or simulation may not need broad language competence. A compact model can devote its limited capacity to the dynamics that matter.
Small models also make iteration easier. A failed plan can lead to a targeted change, as it did here with rollout fine-tuning. Developers can compare runs, inspect data coverage, and revise assumptions without rebuilding a massive general-purpose system.
However, local training does not automatically mean broad accessibility. Reproducing the experiment still requires machine-learning knowledge, emulator instrumentation, data collection, suitable hardware, and patience.
The reported model also learned one limited region of one game. Its dataset was not a complete map of Pokémon Red, and the planner started from a fixed save state.
The project is therefore best understood as an accessible research prototype. It lowers the compute barrier for a specific class of experiments while leaving substantial engineering barriers intact.
The broader LeWorldModel research strengthens that interpretation. The paper describes a roughly 15-million-parameter architecture trained on a single GPU for its experimental tasks. It evaluates navigation, manipulation, and motion-planning environments rather than claiming general intelligence.
For developers, the useful signal is architectural efficiency. A predictive representation does not have to generate photorealistic future frames or verbalize every choice. It can preserve just enough structure for planning.
That can reduce the model size and the cost of evaluating many candidate actions. It also makes the system’s objective more specific than a language model prompted to infer controls from screenshots and prose.
Yet specialization creates its own costs. The developer had to collect more than 42,000 frames for a task that a human already understands. A general model might bring prior knowledge of Pokémon, menus, dialogue, and long-term objectives.
The contest is therefore not simply small versus large. It is prior knowledge versus task-specific learning, local control versus general competence, and efficient prediction versus flexible reasoning.
Recent Pokémon experiments make this comparison visible. Some systems use language models, game memory, handcrafted action lists, or external coaching. Others use reinforcement learning against explicit objectives.
stmonty’s model took a stricter visual route. It learned from grayscale frames and recorded inputs, then planned through the resulting representation.
That narrower input is both the project’s strength and its limitation. It gives the result technical clarity, but it removes information that would help solve the full game.
A model that reads text can understand that gym badges unlock later progress. A predictive model trained around Oak’s laboratory receives no natural explanation of that hierarchy.
Consumer hardware makes the local learning loop notable. It does not eliminate the need for goal structure, diverse data, or systems that operate across longer periods.
The Real Opponent Is the Planning Horizon
The experiment’s hardest problem was not recognizing buttons, but keeping predictions useful as the plan extended beyond familiar short sequences.
A 14-action horizon already created enough drift to defeat the first planner. The model’s predicted state gradually separated from the emulator’s actual state, even though its individual transitions appeared plausible.
Longer Pokémon objectives multiply that problem. Walking through a room requires spatial navigation. Dialogue requires context about earlier selections. Battles add menus, health, types, moves, and changing opponents.
The complete game also contains dependencies spread across hours. A player must discover intermediate goals, remember completed tasks, obtain required items, and adjust when an earlier plan fails.
stmonty identified this issue directly in the Hacker News exchange. Scaling to a full completion would require representations for intermediate goals and a way to arrange them over time.
A larger version of the same short-horizon network would still lack an explicit mechanism for that hierarchy. More parameters might improve predictions, but they would not automatically identify which badge, item, or location should become the next objective.
This is where general-purpose models hold an advantage. They can use language knowledge to recognize game concepts and reason about sequences described in guides, dialogue, or memory.
Their weakness is different. Broad models can hallucinate actions, lose track of state, repeat mistakes, or spend significant resources reasoning over simple control decisions.
A task-specific world model can handle local dynamics more efficiently. A higher-level system could then select goals, while the predictive model handles short sequences.
That layered design appeared in the community discussion. One participant suggested hierarchical world models that operate at different timescales. A high-level component might reason about beating a gym, while a lower-level model predicts individual inputs.
Such a system would resemble how many practical agents are assembled. One component maintains objectives, another models the environment, and a controller chooses or verifies actions.
However, the current experiment did not test that architecture. It had three starter goal embeddings, one starting position, and a fixed planning horizon.
The 52 percent success rate also deserves careful interpretation. It was far better than the random baseline, but it came from repeated searches over the same initial state.
The model did not demonstrate resilience across different rooms, unseen dialogues, changing inventories, or battles. Those tests would require broader data and more varied starting conditions.
There is another important baseline inside the experiment. Pressing A 12 times already completed the task from the save state.
That fact does not erase the learned model’s contribution. Random action sequences still failed, and the trained predictor helped the search discover valid plans. It does weaken any claim that the system exhibited broad strategic understanding.
The real achievement was learning a representation that supported goal-directed search. The real uncertainty is whether that representation remains useful when success depends on longer, more varied chains of causation.
Conventional reinforcement learning offers another comparison. A linked Pokémon agent cited in the discussion used proximal policy optimization for a much broader game objective.
That route depends on an explicit reward design and repeated interaction. The stmonty Pokémon world model instead learned dynamics without receiving the starter objective during its initial training.
Neither approach wins universally. Reward-driven agents can optimize behavior directly, while world models can separate environmental prediction from later goals.
The relevant question is which method remains efficient as the environment expands. A broader evaluation would compare data requirements, training time, failure rates, and generalization across save states.
Without those measurements, the project should not be framed as evidence that small world models outperform reinforcement learning or foundation models. It is evidence that a compact predictive system can produce useful short plans from visual data.
That is a smaller conclusion, but it is also the technically meaningful one.
What to Watch After the Pokémon Red World Model
Three follow-up tests would show whether this is a reusable local-agent design or a successful demonstration tied to one carefully bounded task.
The first signal is performance across varied starting states. A stronger evaluation would begin from multiple positions inside Oak’s laboratory, including unfamiliar orientations and different dialogue stages.
Success under those conditions would show that the model captured more than a narrow path through one save state. A sharp decline would suggest the planner depends heavily on the exact training distribution.
The second signal is a longer objective with intermediate steps. stmonty proposed beginning elsewhere in the laboratory, walking to Oak, completing his dialogue, and then selecting a starter.
That task remains manageable, but it forces the model to sustain a longer rollout. It also tests whether the planner can connect movement, interaction, and dialogue into one coherent sequence.
Reliable results there would strengthen the case for local predictive planning. Repeated failure would confirm that horizon length, rather than parameter count or GPU capacity, remains the decisive bottleneck.
The third signal is a hierarchical controller. The developer said a full game would require multiple intermediate goals and more diverse data from battles, menus, dialogue, and locations.
A future system could combine a high-level goal selector with a low-level world model. The high-level component might decide to reach Oak, choose a starter, or leave the building. The local predictor could search for the inputs needed to complete each step.
That design would also create clearer evaluation points. Researchers could separately measure whether the system chose the correct subgoal and whether the action planner executed it.
Developers should also watch for independent reproductions. The code is public, but a single author’s result does not reveal how sensitive the outcome is to data collection, seeds, hyperparameters, or implementation details.
Reproduction on another consumer GPU would strengthen the accessibility claim. Testing another visual environment would say more about whether the method transfers beyond Pokémon Red.
The project’s strongest contribution is not a completed game. It is a transparent record of a small model failing, revealing why it failed, and improving through a targeted adjustment.
That workflow matters for local AI research. Compact systems expose errors that can become difficult to interpret inside a much larger agent stack.
The stmonty Pokémon world model also gives developers a concrete question to pursue. How much planning can a small predictive representation support before it needs language, memory, hierarchical goals, or a different training signal?
For now, the answer reaches from one laboratory save state to a starter Pokémon. The next valuable result will come from extending that boundary without hiding new assistance inside the system.
If you build local agents, follow the evidence rather than the spectacle. Test new starting states, measure rollout drift, compare meaningful baselines, and record every intervention. A longer successful plan would strengthen the case for compact world models. A collapse outside Oak’s laboratory would be equally useful, because it would identify the architectural limit that the next experiment must address.



