Game Reinforcement Learning isn't Playing Around
Or: why superintelligence is just Runescape with a new coat of paint. From the outside, a lot of reinforcement learning research looks like we're wasting our time and playing too many games. A world class professor told me so verbatim as a new PhD admit. Another told me not to join their university. I dedicate this article in their honor: here's why you didn't play enough games!
Learning through Interaction is Fundamental
Imagine trying to learn about the world only by reading books and watching videos. It wouldn't work in virtually any job, right? Except for some very recent reinforcement learning work, this is exactly how large AI models are trained. It's technically possible that AI might be able to learn this way even though humans can't, just because of the sheer volume of training data. But it's a poor bet. According to all the bigwigs selling you AI, we now have "graduate level" models in every field. At least, if you don't mind them making stuff up a good chunk of the time and failing on simple iterated tasks. So where is the disconnect? How can a "general" model get an IMO gold medal and fail to play simple games? Because book smarts lack verification. Interaction is a grounded mechanism where you get feedback by seeing how your actions alter the world. It's not that AI keeps going off the rails: it's that it doesn't have any in the first place!
Why not Robots?
That's a much more natural formulation of interaction after all, isn't it? It depends what you care about! Picking up and moving around objects is not what I would describe as intelligence. It's control. The majority of robotics tasks used in research are locomotion, manipulation, and very narrowly scoped household chores. How would you simulate a construction crew building a house? You lose on the physics before even getting to training.
Control is important! It is a hard requirement for real-world intelligence unless we want a bunch of smart computers ordering us around. But is it sufficient as the sole paradigm for studying interactive intelligence? No! It's missing all the high-level reasoning bits that separate us from animals. Wouldn't it be great if we had a broad, readily available suite of tasks that require various forms of high-level reasoning? Ideally one with tons of human data, good interpretability, and possibly even some quick-to-simulate aspects of control? Wait...
Game On!
If all of this has been incredibly obvious, chances are you have played some games. Not all AI researchers have, and some are very smug about it. It is impossible to comprehend the monumental accomplishments of ~2019 era OpenAI and DeepMind solving DoTA and SC2 if you think of all games like Atari. Multiagent coordination, precise control coupled with long-term strategy, progressive acquisition of domain knowledge through repetition? Nope, it's just another Atari! I tried to explain to one of the surly professors that both of these games are professionally played, and that top players practice for thousands of hours and continue to innovate on new strategies for years ... "that's their problem."
The topic of my PhD dissertation was Neural MMO, an RL environment I wrote inspired by classic Massively Multiplayer Online games like Runescape and World of Warcraft. The key realization I had was that MMOs are the longest form genre of games, with progression over hundreds and thousands of hours, rich many-agent interactions, emergent strategies, and dynamic economies. Because of this, they are also among the most computationally efficient by design. To me, this was a clear path to emergent intelligence in simulation. Just borrow mechanics from the games industry that we know result in an interesting strategy space, write fast versions for training, and develop increasingly competent agents across all forms of reasoning. Compared to virtually any other benchmark, games are also more interpretable because they have to be. You can play a game to understand how it works, and the games industry has spent decades making complex tasks intuitive and accessible. This is not true of some of the more abstract benchmarks used in academic research. I tried every conceivable way to explain all of this, but there was always a fraction of my audience that couldn't get past the word "game."
How do we do it?
I started PufferAI after graduating to make RL fast and sane, with an engineering-heavy emphasis on scaling to more complex tasks. After a little over a year of work, our library has become a widely used RL toolkit with 25+ ultra-fast environments for research and testing. My favorite is still Neural MMO 3, which I wrote as a 1000x faster follow-up to my PhD that doesn't sacrifice any of the complexity of the original (it's actually a fair bit harder).
It doesn't make practical sense for me to exclusively work on Neural MMO anymore, but scaling to more complex games is still a key piece of our approach. I missed a few important details during my PhD. First, you also have to demonstrate results on non-game simulations - not just for game skeptics, but for people outside of the field. We do this now, both in-house and for clients. Second, many ideas can be investigated at smaller scale with specialized probe environments. I was always skeptical of this during my PhD, because these results usually didn't transfer. The advancements you can read about in the rest of my articles have mostly fixed that. And finally, I underestimated the sheer amount of engineering required to make RL fast. Neural MMO 3 was trained on over a petabyte of compressed data (12,000 years of games) on a single server.
PufferLib is free and open-source with an active community of contributors. Many are building new games for research. Some are building industry sims based on their experience in other fields. If you'd like to join us, join discord.gg/puffer. And if you'd like to support my work for free, star pufferlib on github.