Bobcat Plays Pokémon Gold
Bobcat never writes a word: it reads a state and returns one probability per option. We put the Korean edition of Pokémon Gold behind a harness that only states facts and lists the actions the game allows, gave Bobcat a strategy guide, and let it make every choice from a new game. Twenty new games at one commit, with nothing changed while they ran: 17 won all 16 badges, a median of 19.3 game hours each. A second batch at another commit ended the same, 17 of 20. Random answers on the same harness won no badge. One more game was recorded in full; the six that stalled are listed with why.
On this page
Bobcat is a model that never writes a word. Given a state and a list of options, it returns one probability per option, and that is all it can do. We put the Korean edition of Pokémon Gold behind a harness that only states facts and lists the actions the game allows at that moment, gave Bobcat a strategy guide to read, and let it make every choice from a new game: where to go, what to do there, every move in every battle, what to buy, whom to catch, and when a step of the guide is done.
Then we started 20 new games at once, all at one commit, and changed nothing while they ran. 17 of the 20 won all 16 badges, Johto's eight and Kanto's eight, beating the Elite Four and the Champion on the way. A game that finished took a median of 19.3 hours on the game's own clock and about 5,100 decisions. We did it twice, at two commits a few hours apart, and both batches ended at 17 of 20. The same harness answering at random won no badge.
Our own bar was nine games in ten, and we are not there: the six games that stalled are listed below with the reason for each. After the batches we started three more games at the second batch's commit and recorded them frame by frame; all three won the sixteen badges, beat Red on Mt. Silver and caught Ho-Oh. The videos here are of the one that finished first: 4,887 decisions and 19.9 game hours.

The recorded game in highlights
8 moments of the recorded game, each the dashboard as Bobcat saw it, sped up (one recorded frame is one second of the game): on the left the game, beside it the state Bobcat was sent, then the options it rated highest with their buttons and probabilities, and the log of its choices. First all 25 of them, every gym leader included, in four minutes; then 8 of them one by one.
How often does it finish?
A batch is twenty games started together from the same new game, in the bedroom after the opening, at one commit of the harness and the guide. Nothing is edited while they run, and no game is continued from a save: a game that stalls is a game that did not finish. Each stops when it wins its sixteenth badge. A game that sat on one step of the guide for thousands of decisions without progress was stopped and counted as stalled.
| Batch | Commit | All 16 badges | Game hours to finish (fastest, median, slowest) |
|---|---|---|---|
| A | 13803fe | 17 of 20 | 17.1, 20.3, 26.0 |
| B | 6c233f1 | 17 of 20 | 17.0, 19.3, 24.2 |
Together that is 34 of 40, 85%. Forty games do not pin a rate down closely: the 95% interval runs from 71% to 93%. The games that did not finish:
- Batch A, game 7, at 15 badges: it lost to Blue, the last gym leader, again and again, and wandered off between fights instead of training. Stopped after 4,977 decisions on that step, 56.6 game hours in.
- Batch A, game 11, at 7 badges: it kept losing to Clair with one strong Pokémon and the rest of the team far behind it, and wandered away from her gym. Stopped after 3,514 decisions on that step, 22.1 game hours in.
- Batch A, game 15, at 7 badges: it had the Ice Path's boulders set aside as signs already read, though Strength must be used again on every visit (fixed before batch B). Stopped after 3,215 decisions on that step, 19.9 game hours in.
- Batch B, game 2, at 8 badges: it could not teach Waterfall to its Gyarados, which already knew four moves: two of the game's questions went unread, and the teach never happened. Stopped after 5,677 decisions on that step, 27.6 game hours in.
- Batch B, game 14, at 7 badges: it lost to Clair the way one of batch A's games did, with a team far behind its starter. Stopped after 6,450 decisions on that step, 25.6 game hours in.
- Batch B, game 20, at 11 badges: it took the cycling road to Fuchsia without a bicycle, and the gate's guard turned it back each time. Stopped after 4,681 decisions on that step, 24.6 game hours in.
Three of the six lost a battle Bobcat had to win: a strong gym leader against a team built around one Pokémon. The other three were gaps in the harness: boulders it set aside, fixed between the batches; questions it did not read; a road it did not know was closed. The last two are the next fixes; the battles are Bobcat's to win.
What does the harness do, and what does it not do?
A speedrunner would call the harness a status checker. It reads the game and says what is true: the dialog on screen, read from pixels against the cartridge's own font; menus and where the cursor is; the place and tile, money and badges; the party, with each Pokémon's level, HP and moves with the PP left; the bag and the Pokégear's cards; in battle, both Pokémon and our own types; and what was tried and did not work (a move with no effect, an exit that led nowhere, a battle lost, with the party's levels). It lists what the game allows now: the exits the player can walk to, the people and objects in reach, signs, water and trees for a field move the party knows, grass, the Pokégear radio where a sleeping Snorlax blocks the road, and in battle the moves with PP left, switching, healing items, balls and running. Each option shows the buttons that carry it out (“→5 ↑3 A”), as a player would press them.
It does not choose, and it withholds what the game does not show. The opponent's types are never given, and neither is whether a move is effective; Bobcat has to know that, or read it from what the game says after a move. When Bobcat picks an action, the harness carries it out mechanically: “go to the Radio Tower’s 2F” is a shortest path to the stairs and the button presses to climb them.
What does Bobcat decide?
In the recorded game:
- Where to go and what to do, in one request with two Choices, each time the situation changes (2,245 overworld decisions). The place it chose is kept for its next few decisions, and each exit says whether it leads there.
- Every battle turn: which move, whether to switch, heal, throw a ball or run (2,234 decisions).
- Every prompt and menu: yes or no, what to buy, which move to forget, who learns an HM (323 decisions).
- Whether to catch a wild Pokémon when the guide says to catch one, and with which ball (13 decisions).
- Who leads, asked at a Pokémon Center when one member is far ahead of or behind the rest (30 times).
- Whether a step of the guide is done, where no fact the game records settles it (41 looks).
Why a strategy guide, and what is in it?
A person playing a long game for the first time reads a walkthrough, so Bobcat gets one: a Korean text with one chapter per badge count, published exactly as the games used it. It gives the story order, the place names, the team to catch (Mareep on Route 32, a Water Pokémon that can learn Surf, the Red Gyarados), which types do well against each leader, and for the game's few puzzles with a single answer (the Ice Path's boulders, the underground switches, Saffron's warp pads) that answer, checked tile by tile against the game's disassembly. Bobcat decides what to do with it. The harness turns the page only on facts: an item the step names is in the bag, a place it ends at was entered, a trainer it names was beaten, an event the game records (the Power Plant's part returned, Snorlax fought, Blue spoken to on Cinnabar) is set.
Which choices were not Bobcat's?
The harness has fixed rules, each in the published code with its threshold and each counted in the game's log. On the overworld they set Bobcat's top action aside when it repeats itself: an action picked four times in fifty on a map, an exit that failed twice, someone who has said the same thing three times; the action taken is then Bobcat's most probable remaining one, or, when everything is set aside, the one done least lately. In the recorded game 79% of the 2,245 overworld actions were Bobcat's own top choice. Battles, prompts and catches are Bobcat's answers. A few things are done without asking: the intro, declining nicknames, confirming a purchase or a teach Bobcat chose, and backing out of a menu loop the harness does not understand.

How do we know the harness is not playing?
With every Choice drawn at random on the same harness, at the same commit and from the same new game, it won no badge in 9,517 decisions. The harness is not neutral: it walks to what is chosen and marks the exits that lead towards the guide's places, which is why random answers get anywhere at all. The difference between that and sixteen badges is Bobcat's. The harness also records, for every decision, the state Bobcat was sent, every option with its probability and which rule if any set its choice aside, so the claim can be checked line by line.
How did it get here?
The harness was written against this game with Bobcat in the loop, and most of the work was finding what it read wrongly deep into a game: a PC menu left open, a party screen after a faint, a pocket of the bag read half drawn, an empty items pocket taken for the ball pocket, a shop list that ignored short button presses, a ship that arrived while the guide still said to find the captain. Each showed up only after hundreds or thousands of decisions. When a game stalled on one, we fixed the harness and continued that game from its last save, which is how the first games reached sixteen badges, Red and Ho-Oh. Then came the batches, twenty new games at a time: each batch's stalls were the next commit's fixes, and changes to the guide's hints were tried on and off from the same saves before they stayed. The two batches above are the last two. Between them we fixed one thing batch A showed, the Ice Path's boulders, and renamed the stairs between Goldenrod and its underground, which games had taken for the switch room.
Where is Bobcat weak?
- Battles against a strong leader with a lopsided team, as above. In the recorded game it lost 30 battles and won 229 against trainers.
- It leans on its starter (the recorded game ended with Typhlosion at level 93, Togetic at level 30, Mareep at level 8, Slowpoke at level 10, Gyarados at level 34 and Ho-Oh at level 40); asked who should lead, it keeps the order most of the time.
- Loop rules are needed: 21.0% of the recorded game's overworld actions were set by one.
What did it cost?
The recorded game used 9,015,052 input tokens over 4,887 decisions. It ran for 0.9 hours on one rented RTX PRO 6000 (96 GB) that it shared with the other recorded games; at $2.76 an hour, that time costs $2.37. A finished game of the batches used a median of 9.4 million input tokens; twenty share one GPU. Everything behind the two batches, the batches before them and every fix's test included, took 39.8 hours of one GPU at $4.12 an hour, $164.13. For scale, the recorded game's tokens at Jev's published list price of $0.042 per million input tokens would be $0.38; that is a reference point, not a measurement of either service on this workload.
Limitations
- 85% of 40 new games finished, below our own bar of nine in ten, and forty games are a small sample.
- The twenty games of a batch start from the same new game and differ only through the small numeric differences of a busy model server. They are not twenty independent players: games often played the first chapters alike, and stalls came in kinds rather than one by one.
- Both commits were tuned on this game: every fix and hint came from earlier games of it. There is no held-out game, and another game would need its own harness.
- We wrote the guide, and it contains the answers to the game's fixed puzzles and the team to catch.
- The recorded games were played after the batches, at batch B's commit, because the batches were not recorded; the videos show the first of them to finish.
The code, the guide and the decision logs are in github.com/foxl-ai/bobcat under demos/pokemon. You need your own cartridge dump; no ROM or game asset is published. Pokémon is a trademark of Nintendo, Creatures and GAME FREAK. This is an independent research demo, not affiliated with or endorsed by them, or by TypeSafe AI. The model is described in Bobcat: Typed Decisions from One Forward Pass.
This post is also in Korean.
References and further reading
- Code, guide and decision logs (GitHub)Reference
- Bobcat weights on Hugging FaceReference
- Bobcat Flash weightsReference
- Continual Harness (arXiv)Reference
- PyBoy emulatorReference
- pret/pokegold disassemblyReference