RL 007
· 약 1분
Temporal-Difference Learning
Updates a guess towards a guess.
- Model-free: No MDP Dynamics or Rewards
- Learns directly from episodes of experience.
- Bootstrapping: Estimate value function using estimated value from incomplete episodes.
- Updates the estimate of immediately after each step.
- Combines Monte Carlo and Dynamic Programming (Bootstrapping idea).