상세 보기
초록
In this paper, we propose a method to reduce the learning time of Q-learning by combining the method of updating even to Q-values of unexecuted actions with the method of adding a terminal reward to unvisited Q-values. To verify the method, its performance was compared to that of conventional Q-learning. The proposed approach showed the same performance as conventional Q-learning, with only 27 % of the learning episodes required for conventional Q-learning. Accordingly, we verified that the proposed method reduced learning time by updating more Q-values in the early stage of learning and distributing a terminal reward to more Q-values. © 2013 Springer Science+Business Media.
키워드
Propagation; Q-learning; Q-value; Terminal reward
- 제목
- Enhanced reinforcement learning by recursive updating of Q-values for reward propagation
- 저자
- Sung, Y.; Ahn, E.; Cho, K.
- 발행일
- 2013
- 유형
- Conference Paper
- 권
- 215 LNEE
- 페이지
- 1003 ~ 1008