Read: Multi-Armed Bandits: Exploration vs Exp…
Start with the chapter spine
Open practice dock
Drills and weak-topic replay
Study library
Search all imported chapters
Reinforcement Learning
Continue from the last thing you actually touched, instead of hunting through the registry again.
ContinueNo assignments matching your criteria.
Try a broader filter or jump back to the course hub to browse the full assignment set.
Multi-Armed Bandits: Exploration vs Exploitation
Deep dive into patterns and axioms.
Markov Decision Processes & Bellman Equations
Deep dive into patterns and axioms.
Dynamic Programming: Policy Evaluation, Value Iteration, and Policy Iteration
Deep dive into patterns and axioms.
Monte Carlo Methods: First-Visit, Every-Visit, and Monte Carlo Control
Deep dive into patterns and axioms.
Temporal Difference Learning: TD(0), TD(λ), and Eligibility Traces
Deep dive into patterns and axioms.
Q-Learning and SARSA: Off-Policy and On-Policy TD Control
Deep dive into patterns and axioms.
Deep Q-Networks: Experience Replay, Target Networks, Rainbow
Deep dive into patterns and axioms.
Policy Gradients: REINFORCE and Actor-Critic
Deep dive into patterns and axioms.
Actor-Critic Methods: A2C, A3C, and Advantage Estimation
Deep dive into patterns and axioms.
Proximal Policy Optimization: Clipped Surrogate and Trust Regions
Deep dive into patterns and axioms.