Self-organizing state aggregation for architecture design of Q-learning

Authors:
Kao-Shing Hwang;Hsin-Yi Lin;Yuan-Pao Hsu;Hung-Hsiu Yu
Affiliations:
Department of Electrical Engineering, National Chung-Cheng University, Chiayi 711, Taiwan and Department of Electrical Engineering, National Sun Yat-sen University, Kaohisung 800, Taiwan;Robotics Control Department, Intelligent Robotics Technology Division, Mechanical and Systems Research Laboratories (MSL), Industrial Technology Research Institute (ITRI), Hsinchu 310, Taiwan;Department of Computer Science and Information Engineering, National Formosa University, Yunlin 632, Taiwan;Robotics Control Department, Intelligent Robotics Technology Division, Mechanical and Systems Research Laboratories (MSL), Industrial Technology Research Institute (ITRI), Hsinchu 310, Taiwan
Venue:
Information Sciences: an International Journal
Year:
2011

Citing 12
Cited 6

The ART of Adaptive Pattern Recognition by a Self-Organizing Neural Network

Computer
Technical Note: \cal Q-Learning

Machine Learning
Incremental multi-step Q-learning

Machine Learning - Special issue on reinforcement learning
Introduction to Reinforcement Learning

Introduction to Reinforcement Learning
The FAST Architecture: A Neural Network with Flexible Adaptable-Size Topology

MICRONEURO '96 Proceedings of the 5th International Conference on Microelectronics for Neural Networks and Fuzzy Systems
Reinforcement learning based on local state feature learning and policy adjustment

Information Sciences—Informatics and Computer Science: An International Journal - Special issue: Introduction to multimedia and mobile agents
Anti-swing and positioning control of overhead traveling crane

Information Sciences: an International Journal
A generic architecture for adaptive agents based on reinforcement learning

Information Sciences—Informatics and Computer Science: An International Journal - Special issue: Bio-inspired systems (BIS)
Self-organizing learning array and its application to economic and financial problems

Information Sciences: an International Journal
A fuzzy Actor-Critic reinforcement learning network

Information Sciences: an International Journal
A novel approach for multi-agent-based Intelligent Manufacturing System

Information Sciences: an International Journal
Cooperative strategy based on adaptive Q-learning for robot soccer systems

IEEE Transactions on Fuzzy Systems

Induced states in a decision tree constructed by Q-learning

Information Sciences: an International Journal
An iterative adaptive dynamic programming algorithm for optimal control of unknown discrete-time nonlinear systems with constrained inputs

Information Sciences: an International Journal
Simultaneous policy update algorithms for learning the solution of linear continuous-time H∞ state feedback control

Information Sciences: an International Journal
Undesired state-action prediction in multi-agent reinforcement learning for linked multi-component robotic system control

Information Sciences: an International Journal
Policy sharing between multiple mobile robots using decision trees

Information Sciences: an International Journal
Intelligent controllers for bi-objective dynamic scheduling on a single machine with sequence-dependent setups

Applied Soft Computing

Quantified Score

Hi-index	0.07

Visualization

Abstract

This work describes a novel algorithm that integrates an adaptive resonance method (ARM), i.e. an ART-based algorithm with a self-organized design, and a Q-learning algorithm. By dynamically adjusting the size of sensitivity regions of each neuron and adaptively eliminating one of the redundant neurons, ARM can preserve resources, i.e. available neurons, to accommodate additional categories. As a dynamic programming-based reinforcement learning method, Q-learning involves use of the learned action-value function, Q, which directly approximates Q^*, i.e. the optimal action-value function, which is independent of the policy followed. In the proposed algorithm, ARM functions as a cluster to classify input vectors from the outside world. Clustered results are then sent to the Q-learning design in order to learn how to implement the optimum actions to the outside world. Simulation results of the well-known control algorithm of balancing an inverted pendulum on a cart demonstrates the effectiveness of the proposed algorithm.