Shlakhter, Oleksandr; Lee, Chi-Guhn - In: Computational Statistics 78 (2013) 1, pp. 61-76
We propose a new approach to accelerate the convergence of the modified policy iteration method for Markov decision processes with the total expected discounted reward. In the new policy iteration an additional operator is applied to the iterate generated by Markov operator, resulting in a...