RL 010
· 약 5분
Policy Search Methods
Instead of and functions, we directly estimate the optimal policy.
- Policy was generated directly from the value function (e.g. -greedy)
- Parameterize the policy directly.
- Instead of or , we have
- Given a particular state , what actions should we take? (What is probability of that particular action?)
- is a learnable parameter which could be a neural network.
Value-based and Policy-based Methods
- Value-based: Learnt value function, Implicit policy (e.g. -greedy)
- Policy-based: No value function, Learnt policy.
- Good convergence properties
- Effective in high-dimensional or continuous action spaces
- if it is really continuous and large, it is not efficient as well.
- Can learn stochastic policies.
- Sometimes policies are relatively simple, values and models are complex.
- Typically reaches local optima.
- Obtained knowledge is specific and does not generalize well.
- Actor-Critic: Learnt value function, Learnt policy.
Policy Gradient Techniques
- Trust-Region Based: Optimize the policy in a trust zone (closer circle).
Stochastic policies
Rock-Paper-Scissors
- A deterministic policy is easily exploited
- Deterministic policy: If we follow a set of action or sequence of actions, we will definitely reach the goal and will always reach teh same location.
- A uniform random policy is optimal policy.
Grid World

- Agent cannot differentiate between the grey states.
- features describing state and action .
- is the North and the South are walls and the agent will move to the East.
- If , then when the agent is in the left grey state, it will move west, get stuck in a loop, and never reach the goal.

- Therefore, an optimal stochastic policy moves randomly or in grey states.
- The agent will reach the goal with high probability.
- Policy-based methods can learn Stochastic policies. 4563
Objective Function in Policy-based Methods
Given a policy with parameters , the goal is to find the best .
- In an episodic task (with an end state), is for that particular start state and measures how much expected return we can get.
- is the start state.
- is the measure of the quality of the policy or the objective function to optimize.
- In continuing environments (doesn't have an end state), consider the states and take the average value weighted by how often the agent visits each state.
- is how often the agent is visiting that particular state .
- Use the average reward per time step.
- Small batches can be used to estimate this average reward during training.
Policy-based Methods
- Policy-based RL is an optimization problem, find that maximize .
- is the policy objective function.
- Gradient-free methods: Hill Climbing, Genetic Algorithms, ...
- Gradient-based
- Policy gradient method will search for a local minimum in by ascending the gradient of the Policy.
- why ascending? because we want to maximize the objective function's value/rewards.
- is the step size.
- Assume the policy is differentiable.
- Compute an estimate of the policy gradient.
- Compute using Monte Carlo samples since we need to explore the environment, gather information, and then learn the policy.
Score Function
- Assume the policy is differentiable.
- is the gradient.
- Multiplying and dividing by is called the Likelihood Ratio Trick.
- Since , this does not change the original value.
- is Gradient of the policy
- is the policy.
- measures how sensitive the policy is to changes in , relative to the policy value itself.
- From the derivative rule of the logarithm, .
- Since , this does not change the original value.
- is called the Score Function.