I gave a talk at our internal machine learning seminar, together with Max Golightly, on the road that leads from optimal control to reinforcement learning.

We started from the deterministic picture — variational calculus, Hamilton–Jacobi, the Pontryagin maximum principle — and then moved to the probabilistic one, going through Markov decision processes, dynamic programming and policy gradient methods.

I’m sharing the slides here in case you’re interested in these topics Link.