BSDA5007 · workspace
Reinforcement Learning
Syllabus, study tools, lectures, and curriculum map.
← Back to hubMore courses
Browse course shelfWeekly outline
Syllabus
Week topics from the course map
00W00
Topic
Incomplete
01W01
Review of ML fundamentals – Classification, Regression. Review of probability theory and optimization concepts.
Incomplete
02W02
RL Framework; Supervised learning vs. RL; Explore-Exploit Dilemma; Examples.
Incomplete
03W03
MAB: Definition, Uses, Algorithms, Contextual Bandits, Transition to full RL, Intro to full RL problem
Incomplete
04W04
Intro to MDPs: Definitions , Returns, Value function, Q-function.
Incomplete
05W05
Bellman Equation, DP, Value Iteration, Policy Iteration, Generalized Policy Iteration.
Incomplete
06W06
Evaluation and Control: TD learning, SARSA, Q-learning, Monte Carlo, TD Lambda, Eligibility Traces.
Incomplete
07W07
Maximization-Bias & Representations: Double Q learning, Tabular learning vs. Parameterized, Q-learning with NNs
Incomplete
08W08
Function approximation: Semi-gradient methods, SGD, DQNs, Replay Buffer.
Incomplete
09W09
Policy Gradients: Introduction, Motivation, REINFORCE, PG theorem, Introduction to AC methods
Incomplete
010W10
Actor-Critic Methods, Baselines, Advantage AC, A3C
Incomplete
011W11
Advanced PG/A-C methods: Deterministic PG and DDPG, Soft Actor-Critic (SAC)
Incomplete
012W12
Model-Based RL: Introduction, Motivation, Connections to Planning, Types of MBRL, Benefits, RL with a Learnt Model, Dyna-style models, Latent variable
Incomplete
Study tools
Memorize first
Compact recall strip
5 min
Deep Q-Networks: Experience Replay, Target Networks, Rainbow
Open this first for the core logic and traps.
Dive in
5 min
Markov Decision Processes & Bellman Equations
Open this first for the core logic and traps.
Dive in
5 min
Multi-Armed Bandits: Exploration vs Exploitation
Open this first for the core logic and traps.
Dive in
Practice Dock
One drill at a time, with Theo hints, teacher lanes, and a replay queue.
Use this when the theory is already clear and the goal is fast, clean retrieval under pressure. Mark weak questions, copy a revision pack, and move on.
TheoTeacherTextbookSpeedrun
Hint ladder
5 steps
Weak replay
Local
Revision pack
Copy
Mode lanes
4
Ask less. Recall more.
Open Practice Dock
SM-2 Active
Spaced Repetition
Flashcards
Spaced repetition engine for long-term retention of core concepts.
Start session
Not Available
Interactive Simulations
Visual Labs
Exam Mode
Adaptive Practice
Mock Exam
Set paper
Code Protocol
OPPE Simulator
Secure Comm-Link Terminal
Syllabus Matrix Registry
Global Course Index