Function Approximation
v^(s,w)≈vπ(s)
q^(s,a,w)≈qπ(s,a)
- Incremental methods for prediction
- Model-free VFA: MC and TD targets
GPI, Generalized Policy Iteration
- Iterate between policy evaluation and policy improvement
- Policy Evaluation: Approximate policy evaluation. q^(⋅,⋅,w)≈qπ
- Policy Improvement: ϵ-greedy policy improvement.
Types of Action-Value FA
s→FA(w)→q^(s,w)
- State Value Function Approximation
(s,a)→FA(w)→q^(s,a,w)
- Action Value Function Approximation
- q^(s,Left,w)=3.2
- q^(s,Right,w)=8.7
s→FA(w)→q^(s,a1,w)⋯q^(s,an,w)
- The network output is a vector of Q-values for each action.
- More convenient to use discrete action spaces.
| Action | Q-value |
|---|
| Left | 3.2 |
| Right | 8.7 |
| Jump | 5.3 |
Action VFA
q^(s,a,w)≈qπ(s,a)
- Approximate the Action-value function.
- Doing action a at state s, Approximate the Q-value of that action.
J=Eπ[(qπ(s,a)−q^(s,a,w))2]
- Minimize mean-squared error (MSE) between the approximate Q-value q^(s,a,w) and the true Q-value qπ(s,a).
Δw=α(qπ(s,a)−q^(s,a,w))∇wq^(s,a,w)
- Use SGD to find the local minimum.
- qπ−q^ tells us how much the prediction is off from the target.
- ∇wq^(s,a,w) tells us how each weight affects the predicted Q-value.
Linear Action VFA
x(s,a)=x1(s,a)x2(s,a)⋮xn(s,a)
- Feature vector is used to represent the state and action pair.
q^(s,a,w)≈x(s,a)Tw=∑jxj(s,a)wj
- Predicted Q-value is a linear combination of the feature values. (sum of each feature value xj multiplied by the weight wj)
q^(s,a,w)=x(s,a)Twq^=x1w1+x2w2+⋯+xnwn
- x1,x2,⋯: features of the state and action pair.
- w1,w2,⋯: weights that need to be learned.
- derivative q^ with respect to w.
- ∂w1∂q^=x1
- ∂w2∂q^=x2
∇wq^=x1x2⋮xn=x(s,a)
- e.g. q^=2w1+5w2, then ∇wq^=[25]
Δw=α(qπ−q^)∇wq^=α(qπ−q^)x(s,a)=α(qπ(s,a)−q^(s,a,w))x(s,a)