Skip to main content

RL 007

· One min read

Temporal-Difference Learning

Updates a guess towards a guess.

  • Model-free: No MDP Dynamics or Rewards
  • Learns directly from episodes of experience.
  • Bootstrapping: Estimate value function using estimated value from incomplete episodes.
  • Updates the estimate of VV immediately after each step.
  • Combines Monte Carlo and Dynamic Programming (Bootstrapping idea).