PufferLib is a fast and sane reinforcement learning library that can train tiny, super-human models in seconds. The included learning algorithm, hyperparameter tuning, and simulation methods are the product of our own research. All our tools are free and open source. Need a high performance environment for your application? We build them professionally and offer training + extended support. Contact jsuarez🐡puffer🐡ai.
Blog
- A High Throughput Recurrent Network for PufferLib 4.0
- Visualizing 20k RL Experiments with Constellation
- The Engineering behind PufferLib 4.0
- PufferLib 4.0
- Offline RL is not RL
- Why RL Failed to Bootstrap
- Game Reinforcement Learning isn't Playing Around
- The Tragedy of Reinforcement Learning
- An Ultra Opinionated Guide to Reinforcement Learning
- Reinforcement Learning on a Petabyte of Data at Home
- PufferLib 3.0: Better Reinforcement Learning at 4M sps
- Stronger Hyperparameters with Protein
- Puffing Up PPO
- Neural MMO 3.0
- PufferLib 2.0: Reinforcement Learning at 1M sps
- My thoughts on the current state of AI
- Reinforcement Learning Quickstart Guide
- Sticky Actions Considered Harmful
- CARBS Dominates Hyperparameter Tuning
- The Fireball Definition of AGI
- The Puffer Stack
- Reinforcement Learning Infra is Trash and it's Our Fault
- PufferLib 1.0: Now Stable
- PufferLib 0.7: Puffing Up Performance with Shared Memory
- PufferLib 0.6: An Ocean of Environments for Learning Pufferfish
- PufferLib 0.5: A Bigger EnvPool for Growing Puffers
- PufferLib 0.4: Ready to Take on Bigger Fish
- PufferLib 0.2: Ready to Take on the Big Fish
BibTeX
Repo / RLC Best Paper 2025 / Original Whitepaper@misc{pufferlib,
title = {{PufferLib}: Fast and Sane Simplifying Reinforcement Learning for Complex Game Environments},
author = {Joseph Suarez},
howpublished = {\url{https://github.com/PufferAI/PufferLib}},
year = {2024},
note = {GitHub repository}
}
@article{suarez2025pufferlib,
title = {{PufferLib} 2.0: {R}einforcement Learning at 1M steps/s},
author = {Suarez, Joseph},
journal = {Reinforcement Learning Journal},
volume = {6},
pages = {1378--1388},
year = {2025}
}
@misc{suarez2024pufferlibmakingreinforcementlearning,
title={PufferLib: Making Reinforcement Learning
Libraries and Environments Play Nice},
author={Joseph Suarez},
year={2024},
eprint={2406.12905},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2406.12905},
}
Contributors
Joseph Suarez
Founder & Head Puffer. Writes a lot of code.
Spencer Cheng
Research Scientist & Engineer at PufferAI. Drive, Terraform, Tower Climb, Go, Connect4,
TripleTriad, RWare
Kyoung Whan Choe (최경환)
Mujoco bindings, Testing and bug fixes.
Perumaal Shanmugam
4.0 Environment vectorization, kernel dev
Jonah
kernel dev
David Rubinstein
Several performance improvements w/ torch compilation, lead pokerl contributor.
Victor M. Yeom-Song
Porting Protein to CUDA C
Andrew LeFevre
Impulse Wars
Alok Singh
Docking
Nathan Lichtlé
Tactics
Daniel Addis
Enduro, testing, bug fixes, outreach, recruitment; major pokerl contributor.
Hadrien Crassous
Tetris, freeway
Finlay Sanders
Drone
Sam Turner
Drone
Kinvert
Whisker race
Keelan Donovan
Major pokerl contributor.
Gabe
Pacman
Joao Abrantes
Slimevolley
Yannik
2048
Xander
Trash Pickup
Noah Farr
Breakout
Jake Forsey
Connect4 rewrite with fast minmax AI opponent
David (dmoore101)
Improved breakout physics
haterade
Website design for demo page
arb8020
Website improvements for environment demos
David Bloomin
CARBS integration improvements, 0.4 policy pool/store/selector
Black Ink South
Character art for MOBA
Nick Jenkins
Layout for the system architecture diagram. Adversary.design
Andranik Tigranyan
Streamline and animate the pufferfish. Hire him on UpWork if you like what you see here.
Sara Earle
Original pufferfish model. Hire her on UpWork if you like what you see here.