Quiz 2

BSDA5007 · workspace

Reinforcement Learning

Syllabus, study tools, lectures, and curriculum map.

← Back to hub
Weekly outline

Syllabus

Week topics from the course map

00W00

Topic

Incomplete
01W01

Review of ML fundamentals – Classification, Regression. Review of probability theory and optimization concepts.

Incomplete
02W02

RL Framework; Supervised learning vs. RL; Explore-Exploit Dilemma; Examples.

Incomplete
03W03

MAB: Definition, Uses, Algorithms, Contextual Bandits, Transition to full RL, Intro to full RL problem

Incomplete
04W04

Intro to MDPs: Definitions , Returns, Value function, Q-function.

Incomplete
05W05

Bellman Equation, DP, Value Iteration, Policy Iteration, Generalized Policy Iteration.

Incomplete
06W06

Evaluation and Control: TD learning, SARSA, Q-learning, Monte Carlo, TD Lambda, Eligibility Traces.

Incomplete
07W07

Maximization-Bias & Representations: Double Q learning, Tabular learning vs. Parameterized, Q-learning with NNs

Incomplete
08W08

Function approximation: Semi-gradient methods, SGD, DQNs, Replay Buffer.

Incomplete
09W09

Policy Gradients: Introduction, Motivation, REINFORCE, PG theorem, Introduction to AC methods

Incomplete
010W10

Actor-Critic Methods, Baselines, Advantage AC, A3C

Incomplete
011W11

Advanced PG/A-C methods: Deterministic PG and DDPG, Soft Actor-Critic (SAC)

Incomplete
012W12

Model-Based RL: Introduction, Motivation, Connections to Planning, Types of MBRL, Benefits, RL with a Learnt Model, Dyna-style models, Latent variable

Incomplete

Secure Comm-Link Terminal

Secure Comm-Link // Playlist Connected
Uplink 12ms
Syllabus Synchronization: Active

Syllabus Matrix Registry

Global Course Index

Open Full Frame

Master Registry

v6.4 Directory

Foundational

Diploma

BSc Degree

BS Degree

PG / MTech

BSDA5007
PG / MTech
4 Credits

Special topics in Machine Learning (Reinforcement Learning)

To enable the student to understand the reinforcement learning paradigm, to be able to identify when an RL formulation is appropriate, to understan...

Execution Protocol

Module 0

Topic

Module 1

Review of ML fundamentals – Classification, Regression. Review of probability theory and optimization concepts.

Module 2

RL Framework; Supervised learning vs. RL; Explore-Exploit Dilemma; Examples.

Module 3

MAB: Definition, Uses, Algorithms, Contextual Bandits, Transition to full RL, Intro to full RL problem

Module 4

Intro to MDPs: Definitions , Returns, Value function, Q-function.

Module 5

Bellman Equation, DP, Value Iteration, Policy Iteration, Generalized Policy Iteration.

Module 6

Evaluation and Control: TD learning, SARSA, Q-learning, Monte Carlo, TD Lambda, Eligibility Traces.

Module 7

Maximization-Bias & Representations: Double Q learning, Tabular learning vs. Parameterized, Q-learning with NNs

Module 8

Function approximation: Semi-gradient methods, SGD, DQNs, Replay Buffer.

Module 9

Policy Gradients: Introduction, Motivation, REINFORCE, PG theorem, Introduction to AC methods

Module 10

Actor-Critic Methods, Baselines, Advantage AC, A3C

Module 11

Advanced PG/A-C methods: Deterministic PG and DDPG, Soft Actor-Critic (SAC)

Module 12

Model-Based RL: Introduction, Motivation, Connections to Planning, Types of MBRL, Benefits, RL with a Learnt Model, Dyna-style models, Latent variable

Document outline

Keep your place and jump directly to a heading.

Table of Contents
System Normal // Awaiting Context

Intelligence Hub

Navigate the knowledge graph to generate context. The Hub adapts dynamically to surface backlinks, related notes, and metadata insights.