Linear Fitted-Q Iteration with Multiple Reward Functions

Daniel J Lizotte; Michael Bowling; Susan A Murphy

Linear Fitted-Q Iteration with Multiple Reward Functions

J Mach Learn Res. 2012 Nov;13(Nov):3253-3295.

Authors

Daniel J Lizotte¹, Michael Bowling, Susan A Murphy

Affiliation

¹ David R. Cheriton School of Computer Science, University of Waterloo, Waterloo, ON N2L 3G1, Canada, DLIZOTTE@UWATERLOO.CA.

PMID: 23741197
PMCID: PMC3670261

Abstract

We present a general and detailed development of an algorithm for finite-horizon fitted-Q iteration with an arbitrary number of reward signals and linear value function approximation using an arbitrary number of state features. This includes a detailed treatment of the 3-reward function case using triangulation primitives from computational geometry and a method for identifying globally dominated actions. We also present an example of how our methods can be used to construct a real-world decision aid by considering symptom reduction, weight gain, and quality of life in sequential treatments for schizophrenia. Finally, we discuss future directions in which to take this work that will further enable our methods to make a positive impact on the field of evidence-based clinical decision support.

Keywords: decision making; dynamic programming; linear regression; preference elicitation; reinforcement learning.

Abstract

Grants and funding