Reinforcement Learning from Scratch
Six RL algorithms built by hand, all learning to balance a pole.
Across three assignments for the Reinforcement Learning course at Leiden University, my group of three built the core deep RL algorithms from scratch in PyTorch and ran them all on CartPole. No external libraries, just the concepts turned into code. We covered DQN, REINFORCE, Actor Critic, A2C, PPO and discrete SAC.
Links 路 馃捇 DQN 路 馃捇 REINFORCE & AC 路 馃捇 PPO & SAC 路 馃搫 DQN 路 馃搫 REINFORCE & AC 路 馃搫 PPO & SAC
What鈥檚 in it
The first assignment is a Deep Q Network written from the ground up, paired with a hyperparameter ablation. I toggled experience replay and the target network on and off, swept a range of settings, and found that on a problem this simple the target network tended to hurt more than it helped.
The later assignments move to policy gradient and actor critic methods. My PPO uses a clipped objective with generalised advantage estimation, minibatched updates and an entropy bonus that decays over training. My discrete SAC has twin Q networks, a replay buffer, soft target updates and automatic temperature tuning. Both essentially solve CartPole, climbing to the maximum return, while the simpler REINFORCE and Actor Critic only get part of the way.
Writing these by hand rather than calling a library, as well as performing our own ablations, is the main reason the ideas actually stuck.