VLDB 2026 Research / reviewers in the wild / expert
Kao-Shing Hwang
dblp:02/1218
· DBLP profile ↗
79ranked-venue papers
37as first author
17since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 32 · 21 first-author · 4 since 2021Artificial intelligence and machine learning · 25 · 7 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 22 · 20 first-author · 1 since 2021Databases, data management, data science and information retrieval · 11 · 2 first-author · 3 since 2021Systems, architecture and hardware · 3Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Adaptive path planning for wafer second probing via an attention-based hierarchical reinforcement learning framework with shared memory
Haobin Shi, Ziming He, Kao-Shing Hwang |
Inf. Sci. | 3 |
| 2025 | Domain adaptation with temporal ensembling to local attention region search for object detection
Haobin Shi, Ziming He, Kao-Shing Hwang |
Knowl. Based Syst. | 3 |
| 2025 | A segmented motion synthesis method for robotic task-oriented locomotion imitation system
Haobin Shi, Ziming He, Jianning Zhan, Kao-Shing Hwang |
Knowl. Based Syst. | 4 |
| 2025 | Robotic Locomotion Skill Learning Using Unsupervised Reinforcement Learning With Controllable Latent Space PartitionabstractEffective skill learning in an unsupervised manner is one of the capabilities an intelligent agent or robot should have. The discovered task-agnostic skills can be fine-tuned to downstream long-horizon tasks to improve execution efficiency. Unfortunately, the self-learning of locomotion skills, which occurs naturally in infancy, has been slow to develop in robotics. The instability exhibited by existing skill-learning methods makes it difficult to directly apply to complex control tasks, such as humanoid robots. To acquire reliable robotic locomotion skills, this article proposes a controllable latent space partition framework to assist reinforcement learning in accomplishing practicability-oriented unsupervised skill discovery (PoSD). Specifically, we use the distance similarity measure of the trajectory feature space to introduce the indicative information of the expert demonstrations into the partitioning and mapping process of the latent space. In addition, the intrinsic subrewards based on contrastive learning and particle entropy are designed to promote skill diversity and encourage exploration. Finally, reinforcement learning completes the generation of skill-conditioned policy driven by composite intrinsic rewards. The performance investigation of our method is conducted on five robots with more than 15 skills. The results indicate that PoSD achieves noticeable improvements in adaptation efficiency and practicability compared with other SOTA unsupervised skill discovery methods. Ziming He, Haobin Shi, Jingchen Li 0003, Kao-Shing Hwang |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Semi-Supervised Detection Model Based on Adaptive Ensemble Learning for Medical ImagesabstractIntroducing deep learning technologies into the medical image processing field requires accuracy guarantee, especially for high-resolution images relayed through endoscopes. Moreover, works relying on supervised learning are powerless in the case of inadequate labeled samples. Therefore, for end-to-end medical image detection with overcritical efficiency and accuracy in endoscope detection, an ensemble-learning-based model with a semi-supervised mechanism is developed in this work. To gain a more accurate result through multiple detection models, we propose a new ensemble mechanism, termed alternative adaptive boosting method (Al-Adaboost), combining the decision-making of two hierarchical models. Specifically, the proposal consists of two modules. One is a local region proposal model with attentive temporal-spatial pathways for bounding box regression and classification, and the other one is a recurrent attention model (RAM) to provide more precise inferences for further classification according to the regression result. The proposal Al-Adaboost will adjust the weights of labeled samples and the two classifiers adaptively, and the nonlabel samples are assigned pseudolabels by our model. We investigate the performance of Al-Adaboost on both the colonoscopy and laryngoscopy data coming from CVC-ClinicDB and the affiliated hospital of Kaohsiung Medical University. The experimental results prove the feasibility and superiority of our model. Jingchen Li 0003, Haobin Shi, Naijun Liu, Kao-Shing Hwang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Eliminating Primacy Bias in Online Reinforcement Learning by Self-DistillationabstractExcessive invalid explorations at the beginning of training lead deep reinforcement learning process to fall into the risk of overfitting, further resulting in spurious decisions, which obstruct agents in the following states and explorations. This phenomenon is termed primacy bias in online reinforcement learning. This work systematically investigates the primacy bias in online reinforcement learning, discussing the reason for primacy bias, while the characteristic of primacy bias is also analyzed. Besides, to learn a policy generalized to the following states and explorations, we develop an online reinforcement learning framework, termed self-distillation reinforcement learning (SDRL), based on knowledge distillation, allowing the agent to transfer the learned knowledge into a randomly initialized policy at regular intervals, and the new policy network is used to replace the original one in the following training. The core idea for this work is distilling knowledge from the trained policy to another policy can filter biases out, generating a more generalized policy in the learning process. Moreover, to avoid the overfitting of the new policy due to excessive distillations, we add an additional loss in the knowledge distillation process, using L2 regularization to improve the generalization, and the self-imitation mechanism is introduced to accelerate the learning on the current experiences. The results of several experiments in DMC and Atari 100k suggest the proposal has the ability to eliminate primacy bias for reinforcement learning methods, and the policy after knowledge distillation can urge agents to get higher scores more quickly. Jingchen Li 0003, Haobin Shi, Hua-Rui Wu 0001, Chunjiang Zhao 0001, Kao-Shing Hwang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Using Goal-Conditioned Reinforcement Learning With Deep Imitation to Control Robot Arm in Flexible Flat Cable Assembly TaskabstractLeveraging reinforcement learning on high-precision decision-making in Robot Arm assembly scenes is a desired goal in the industrial community. However, tasks like Flexible Flat Cable (FFC) assembly, which require highly trained workers, pose significant challenges due to sparse rewards and limited learning conditions. In this work, we propose a goal-conditioned self-imitation reinforcement learning method for FFC assembly without relying on a specific end-effector, where both perception and behavior plannings are learned through reinforcement learning. We analyze the challenges faced by Robot Arm in high-precision assembly scenarios and balance the breadth and depth of exploration during training. Our end-to-end model consists of hindsight and self-imitation modules, allowing the Robot Arm to leverage futile exploration and optimize successful trajectories. Our method does not require rule-based or manual rewards, and it enables the Robot Arm to quickly find feasible solutions through experience relabeling, while unnecessary explorations are avoided. We train the FFC assembly policy in a simulation environment and transfer it to the real scenario by using domain adaptation. We explore various combinations of hindsight and self-imitation learning, and discuss the results comprehensively. Experimental findings demonstrate that our model achieves fast and advanced flexible flat cable assembly, surpassing other reinforcement learning-based methods.Note to Practitioners—The motivation of this article stems from the need to develop an efficient and accurate FFC assembly policy for 3C (Computer, Communication, and Consumer Electronic) industry, promoting the development of intelligent manufacturing. Traditional control methods are incompetent to complete such a high-precision task with Robot Arm due to the difficult-to-model connectors, and existing reinforcement learning methods cannot converge with restricted epochs because of the difficult goals or trajectories. To quickly learn a high-quality assembly for Robot Arm and accelerate the convergence speed, we combine the goal-conditioned reinforcement learning and self-imitation mechanism, balancing the depth and breadth of exploration. The proposal takes visual information and six-dimensions force as state, obtaining satisfactory assembly policies. We build a simulation scene by the Pybullet platform and pre-train the Robot Arm on it, and then the pre-trained policies can be reused in real scenarios with finetuning. Jingchen Li 0003, Haobin Shi, Kao-Shing Hwang |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2023 | Lateral Transfer Learning for Multiagent Reinforcement LearningabstractSome researchers have introduced transfer learning mechanisms to multiagent reinforcement learning (MARL). However, the existing works devoted to cross-task transfer for multiagent systems were designed just for homogeneous agents or similar domains. This work proposes an all-purpose cross-transfer method, called multiagent lateral transfer (MALT), assisting MARL with alleviating the training burden. We discuss several challenges in developing an all-purpose multiagent cross-task transfer learning method and provide a feasible way of reusing knowledge for MARL. In the developed method, we take features as the transfer object rather than policies or experiences, inspired by the progressive network. To achieve more efficient transfer, we assign pretrained policy networks for agents based on clustering, while an attention module is introduced to enhance the transfer framework. The proposed method has no strict requirements for the source task and target task. Compared with the existing works, our method can transfer knowledge among heterogeneous agents and also avoid negative transfer in the case of fully different tasks. As far as we know, this article is the first work denoted to all-purpose cross-task transfer for MARL. Several experiments in various scenarios have been conducted to compare the performance of the proposed method with baselines. The results demonstrate that the method is sufficiently flexible for most settings, including cooperative, competitive, homogeneous, and heterogeneous configurations. Haobin Shi, Jingchen Li 0003, Jiahui Mao, Kao-Shing Hwang |
IEEE Trans. Cybern. | 4 |
| 2023 | Path Planning of Randomly Scattering Waypoints for Wafer Probing Based on Deep Attention MechanismabstractWafer probing is a critical process employed to measure the yield of wafer fabrication. The primary object of wafer probing is to find the defect grain on the wafer. After a full coverage check, there are always some suspected grains existing for further inspection. However, this second probing result could be affected by the shape of the probe card and the setting actions (path planning) of operators for grains randomly scattering on the wafer. Good grains can be damaged by reprobe actions, which decrease production performance and customer trust. In general, it also requires manpower to perform reprobing, which dramatically deteriorates the throughput of production. This article has studied this problem, and an adaptive coverage path planning (CPP) method for randomly scattering grains using an attention interface is proposed. The proposed randomly scattering waypoints method uses deep reinforcement learning (DRL) for automatic real-time path planning of the second detection. A soft attention interface accelerates the process with a less overlapped check. The experimental results demonstrate the efficiency of the proposed method in terms of less overlapping and steps, and this method learns a better CPP strategy for wafer probing than programmed paths and other RL-based methods. Haobin Shi, Jingchen Li 0003, Meng Liang, Maxwell Hwang, Kao-Shing Hwang, Yun-Yu Hsu |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2022 | Multi-agent reinforcement learning by the actor-critic model with an attention interface
Lixiang Zhang, Jingchen Li 0003, Haobin Shi, Kao-Shing Hwang |
Neurocomputing | 5 |
| 2022 | A collaboration of multi-agent model using an interactive interface
Jingchen Li 0003, Haobin Shi, Kao-Shing Hwang |
Inf. Sci. | 4 |
| 2022 | A behavior fusion method based on inverse reinforcement learning
Haobin Shi, Jingchen Li 0003, Shicong Chen, Kao-Shing Hwang |
Inf. Sci. | 4 |
| 2022 | An adaptive multi-sensor visual attention model
Jingchen Li 0003, Haobin Shi, Kao-Shing Hwang |
Neural Comput. Appl. | 4 |
| 2022 | Using Fuzzy Logic to Learn Abstract Policies in Large-Scale Multiagent Reinforcement LearningabstractLarge-scale multiagent reinforcement learning requires huge computation and space costs, and the too-long execution process makes it hard to train policies for agents. This work proposes a concept of fuzzy agent, which is a new paradigm for training homogeneous agents. Aiming at a lightweight and affordable reinforcement learning mechanism for large-scale homogeneous multiagent systems, we break the one-to-one correspondence between agent and policy, designing abstract agents as the substitute for the multiagent to interact with the environment. The Markov decision process models for these abstract agents are conducted by fuzzy logic, which also acts on the behavior mapping from abstract agent to entity. Specifically, just the abstract agents execute their policy at a time step, and the concrete behaviors are generated by simple matrix operations. The proposal has lower space and computation complexities because the number of abstract agents is far less than that of entities, and the coupling among agents is retained implicitly. Compared with other approximation and simplification methods, the proposed fuzzy agent not only greatly reduces required computing resources but also ensures the effectiveness of the learned policies. Several experiments are conducted to validate our method. The results show that the proposal outperforms the baseline methods, while it has satisfactory zero-shot and few-shot transfer abilities. Jingchen Li 0003, Haobin Shi, Kao-Shing Hwang |
IEEE Trans. Fuzzy Syst. | 3 |
| 2021 | A deep learning method for counting white blood cells in bone marrow imagesabstractBACKGROUND: Differentiating and counting various types of white blood cells (WBC) in bone marrow smears allows the detection of infection, anemia, and leukemia or analysis of a process of treatment. However, manually locating, identifying, and counting the different classes of WBC is time-consuming and fatiguing. Classification and counting accuracy depends on the capability and experience of operators. RESULTS: This paper uses a deep learning method to count cells in color bone marrow microscopic images automatically. The proposed method uses a Faster RCNN and a Feature Pyramid Network to construct a system that deals with various illumination levels and accounts for color components' stability. The dataset of The Second Affiliated Hospital of Zhejiang University is used to train and test. CONCLUSIONS: The experiments test the effectiveness of the proposed white blood cell classification system using a total of 609 white blood cell images with a resolution of 2560 × 1920. The highest overall correct recognition rate could reach 98.8% accuracy. The experimental results show that the proposed system is comparable to some state-of-art systems. A user interface allows pathologists to operate the system easily. Maxwell Hwang, Wei-Cheng Jiang, Kefeng Ding, Hsiao Chien Chang, Kao-Shing Hwang |
BMC Bioinform. | 6 |
| 2021 | Application of artificial intelligence ensemble learning model in early prediction of atrial fibrillationabstractBACKGROUND: Atrial fibrillation is a paroxysmal heart disease without any obvious symptoms for most people during the onset. The electrocardiogram (ECG) at the time other than the onset of this disease is not significantly different from that of normal people, which makes it difficult to detect and diagnose. However, if atrial fibrillation is not detected and treated early, it tends to worsen the condition and increase the possibility of stroke. In this paper, P-wave morphology parameters and heart rate variability feature parameters were simultaneously extracted from the ECG. A total of 31 parameters were used as input variables to perform the modeling of artificial intelligence ensemble learning model. RESULTS: This paper applied three artificial intelligence ensemble learning methods, namely Bagging ensemble learning method, AdaBoost ensemble learning method, and Stacking ensemble learning method. The prediction results of these three artificial intelligence ensemble learning methods were compared. As a result of the comparison, the Stacking ensemble learning method combined with various models finally obtained the best prediction effect with the accuracy of 92%, sensitivity of 88%, specificity of 96%, positive predictive value of 95.7%, negative predictive value of 88.9%, F1 score of 0.9231 and area under receiver operating characteristic curve value of 0.911. CONCLUSION: In feature extraction, this paper combined P-wave morphology parameters and heart rate variability parameters as input parameters for model training, and validated the value of the proposed parameters combination for the improvement of the model's predicting effect. In the calculation of the P-wave morphology parameters, the hybrid Taguchi-genetic algorithm was used to obtain more accurate Gaussian function fitting parameters. The prediction model was trained using the Stacking ensemble learning method, so that the model accuracy had better results, which can further improve the early prediction of atrial fibrillation. Cai Wu, Maxwell Hwang, Tian-Hsiang Huang, Yenming J. Chen, Yiu-Jen Chang, Tsung-Han Ho, Kao-Shing Hwang, Wen-Hsien Ho |
BMC Bioinform. | 8 |
| 2021 | An explainable ensemble feedforward method with Gaussian convolutional filter
Jingchen Li 0003, Haobin Shi, Kao-Shing Hwang |
Knowl. Based Syst. | 3 |
| 2020 | An ensemble method for inverse reinforcement learning
Jin-Ling Lin, Kao-Shing Hwang, Haobin Shi, Wei Pan 0007 |
Inf. Sci. | 2 |
| 2020 | An end-to-end inverse reinforcement learning by a boosting approach with relative entropy
Tao Zhang 0089, Maxwell Hwang, Kao-Shing Hwang |
Inf. Sci. | 4 |
| 2020 | Adaptive Image-Based Visual Servoing Using Reinforcement Learning With Fuzzy State CodingabstractImage-based visual servoing (IBVS) allows precise control of positioning and motion for relatively stationary targets using visual feedback. For IBVS, a mixture parameter β allows better approximation of the image Jacobian matrix, which has a significant effect on the performance of IBVS. However, the setting for the mixture parameter depends on the camera's realtime posture; there is no clear way to define the change rules for most IBVS applications. Using simple model-free reinforcement learning, Q-learning, this article proposes a method to adaptively adjust the image Jacobian matrix for IBVS. If the state-space is discretized, traditional Q-learning encounters problems with the resolution that can cause sudden changes in the action, so the visual servoing system performs poorly. Besides, a robot in a real-world environment also cannot learn on as large a scale as virtual agents, so the efficiency with which agents learn must be increased. This article proposes a method that uses fuzzy state coding to accelerate learning during the training phase and to produce a smooth output in the application phase of the learning experience. A method that compensates for delay also allows more accurate extraction of features in a real environment. The results for simulation and experiment demonstrate that the proposed method performs better than other methods, in terms of learning speed, movement trajectory, and convergence time. Haobin Shi, Haibo Wu 0007, Jinhui Zhu, Maxwell Hwang, Kao-Shing Hwang |
IEEE Trans. Fuzzy Syst. | 6 |
| 2020 | A Fuzzy Adaptive Approach to Decoupled Visual Servoing for a Wheeled Mobile RobotabstractTo address the performance bottleneck for image-based visual servoing (IBVS), it is necessary to have appropriate servoing control laws, increased accuracy for image feature detection, and minimal approximation errors. This article proposes a fuzzy adaptive method for decoupled IBVS that allows the efficient control of a wheeled mobile robot (WMR). To address the under-actuated dynamics of the WMR, a decoupled controller is used and translation and rotation are decoupled by using two independent servoing gains, instead of the single servoing gain that is used for traditional IBVS. To reduce the effect of image noise, this article develops an improved bagging method for the decoupled controller that calculates the inverse kinematics and does not use the Moore-Penrose pseudoinverse method. To improve convergence, improved Q-learning is used to adaptively adjust the mixture parameter for the image Jacobian matrix (IQ-IBVS). This allows the mixture parameter can be adjusted while the robot moves under the influence of servo control. A fuzzy method is used to tune the learning rate for the IQ-IBVS method, which ensures effective learning. The results of simulation and experiments show that the proposed method performs better than other methods, in terms of convergence. Haobin Shi, Meng Xu 0005, Kao-Shing Hwang |
IEEE Trans. Fuzzy Syst. | 3 |
| 2020 | End-to-End Navigation Strategy With Deep Reinforcement Learning for Mobile RobotsabstractIn this article, we develop a navigation strategy based on deep reinforcement learning (DRL) for mobile robots. Because of the large difference between simulation and reality, most of the trained DRL models cannot be directly migrated into real robots. Moreover, how to explore in a sparsely rewarded environment is also a long-standing problem of DRL. This article proposes an end-to-end navigation planner that translates sparse laser ranging results into movement actions. Using this highly abstract data as input, agents trained by simulation can be extended to the real scene for practical application. For map-less navigation across obstacles and traps, it is difficult to reach the target via random exploration. Curiosity is used to encourage agents to explore the state of an environment that has not been visited and as an additional reward for exploring behavior. The agent relies on the self-supervised model to predict the next state, based on the current state and the executed action. The prediction error is used as a measure of curiosity. The experimental results demonstrate that without any manual design features and previous demonstrations, the proposed method accomplishes map-less navigation in complex environments. Through a reward signal that is enhanced by intrinsic motivation, the agent explores more efficiently, and the learned strategy is more reliable. Haobin Shi, Meng Xu 0005, Kao-Shing Hwang |
IEEE Trans. Ind. Informatics | 4 |
| 2019 | A learning approach to image-based visual servoing with a bagging method of velocity calculations
Haobin Shi, Kao-Shing Hwang, Xuesi Li |
Inf. Sci. | 2 |
| 2019 | Tracking and proximity detection for robotic operations by multiple depth cameras
Haobin Shi, Kao-Shing Hwang, Juting Lai |
Inf. Sci. | 2 |
| 2019 | Adaptive Image-Based Visual Servoing With Temporary Loss of the Visual SignalabstractImage-based visual servoing (IBVS) can reach a desired position for a relatively stationary target using continuous visual feedback. Proper feature extraction and appropriate servoing control laws are essential to performance for IBVS. IBVS control can be interrupted or interfered abruptly if no features are extracted when the observed object is occluded. To address the problem of missing feature points in current images during a visual navigation task, a homography method that uses a priori visual information is proposed to predict all of the missing feature points and to ensure the execution of IBVS. The mixture parameter for the image Jacobian matrix can also affect the control of IBVS. The settings for the mixture parameter are heuristic so there is no a systematic approach for most IBVS applications. An adaptive control approach is proposed to determine the mixture parameter. The proposed method uses a reinforcement learning (RL) method to adaptively adjust the mixture parameter during the robot movement, which allows more efficient control than a constant parameter. A logarithmic interval state-space partition for RL is used to ensure efficient learning. The integrated visual servoing control system is validated by several experiments that involve wheeled mobile robots reaching a target with a desired configuration. The results for simulation and experiment demonstrate that the proposed method has a faster convergence rate than other methods. Haobin Shi, Gang Sun 0003, Yuanpeng Wang, Kao-Shing Hwang |
IEEE Trans. Ind. Informatics | 4 |
| 2018 | A Visual Servoing with Collision Avoidance Mechanism for Redundant ManipulatorsabstractThis paper presents a visual servoing method for a redundant manipulator to control quickly and accurately by image error in a complicated environment. The proposed approach combines two parts, one is position-based visual servoing, and another is avoiding obstacles by using virtual repulsive torque. The image features are found and mapped by Oriented FAST and Rotated BRIEF (ORB) and the image errors are calculated. Furthermore, the object pose is estimated in the world coordinate by the image errors and the current pose of the manipulator and the target EOAT pose compute will be obtained. After confirming the EOAT pose, this method introduces a virtual repulsive torque for solving generalized inverse kinematics problems by analytical and numerical methods. The proposed approach uses the property of the redundant manipulator to maintain the distance between links and obstacles. For the applicability, we transform the programs to the nodes and use Robot Operating System as the core to build the communication protocol for exchanging information. The validity and efficiency of the proposed method are measured in simulations, by which it verifies that the method is not only feasible to constrained motions but also has good performances in collision avoidance. Kao-Shing Hwang, Yi-Yun Cho, Wei-Cheng Jiang |
SMC | 1 |
| 2018 | An adaptive decision-making method with fuzzy Bayesian reinforcement learning for robot soccer
Haobin Shi, Shuge Zhang, Xuesi Li, Kao-Shing Hwang |
Inf. Sci. | 5 |
| 2018 | Decoupled Visual Servoing With Fuzzy Q-LearningabstractThe objective of visual servoing aims to control an object's motion with visual feedbacks and becomes popular recently. Problems of complex modeling and instability always exist in visual servoing methods. Moreover, there are few research works on selection of the servoing gain in image-based visual servoing (IBVS) methods. This paper proposes an IBVS method with Q-Learning, where the learning rate is adjusted by a fuzzy system. Meanwhile, a synthetic preprocess is introduced to perform feature extraction. The extraction method is actually a combination of a color-based recognition algorithm and an improved contour-based recognition algorithm. For dealing with underactuated dynamics of the unmanned aerial vehicles (UAVs), a decoupled controller is designed, where the velocity and attitude are decoupled through attenuating the effects of underactuation in roll and pitch and two independent servoing gains, for linear and angular motion servoing, respectively, are designed in place of single servoing gain in traditional methods. For further improvement in convergence and stability, a reinforcement learning method, Q-Learning, is taken for adaptive servoing gain adjustment. The Q-Learning is composed of two independent learning agents for adjusting two serving gains, respectively. In order to improve the performance of the Q-Learning, a fuzzy-based method is proposed for tuning the learning rate. The results of simulations and experiments on control of UAVs demonstrate that the proposed method has better properties in stability and convergence than the competing methods. Haobin Shi, Xuesi Li, Kao-Shing Hwang, Wei Pan 0007, Genjiu Xu |
IEEE Trans. Ind. Informatics | 3 |
| 2018 | Model Learning for Multistep Backward Prediction in Dyna-Q LearningabstractA model-based reinforcement learning (RL) method which interplays direct and indirect learning to update${Q}$functions is proposed. The environment is approximated by a virtual model that can predict the transition to the next state and the reward of the domain. This virtual model is used to train${Q}$functions to accelerate policy learning. Lookup table methods are usually used to establish such environmental models, but these methods need to collect tremendous amounts of experiences to enumerate responses of the environment. In this paper, a stochastic model learning method based on tree structures is presented. To model the transition probability, an online clustering method is applied to equip the model learning method with the abilities to evaluate the transition probability. By the virtual model, the RL method produces simulated experience in the stage of indirect learning. Since simulated transitions and backups are more usefully focused by working backward from the state-action, the pair estimated${Q}$value of which changes significantly, the useful one-step backups are actions that lead directly into the one state whose value has already obviously been changed. This, however, may induce a false positive; that is, a backup state may be an invalid state, such as an absorbing or terminal state, especially in cases where the changes of${Q}$values at the planning stage are still needed to put back for ranking even though they are based on a simulated experience and are possibly erroneous. It is obvious that when the agent is attracted to generate simulated experience around the area of these absorbing states, the learning efficiency is deteriorated. This paper proposes three detecting methods to solve this problem. Moreover, the policy learning can speed up. The effectiveness and generality of our method is further demonstrated in three numerical simulations. The simulation results demonstrate that the training rate of our method is obviously improved. Kao-Shing Hwang, Wei-Cheng Jiang, Iris Hwang |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2017 | A Reinforcement Learning Method with Implicit Critics from a Bystander
Kao-Shing Hwang, Chi-Wei Hsieh, Wei-Cheng Jiang, Jin-Ling Lin |
ISNN (1) | 1 |
| 2017 | A shaped-q learning for multi-agents systemsabstractThis paper proposes an architecture where each agent maintains a cooperative tendency table (CTT). In the process of learning, agents need not communicate with each other but observe partners' actions while taking actions. If one of the agents meets a bad situation, such as bumping onto obstacles after taking an action. In such a case, agents will receive a bad reward from the environment. Similarly, if one agent reaches a goal after taking an action, agents obtain a good reward instead. Rewards are used to update the policy and to adjust cooperative tendency values which are recorded in the individual CTT. When an agent perceives a state, the corresponding cooperative tendency value, and the Q-value are merged to a Shaped-Q value. The action with maximal Shaped-Q value in this state will be selected. After agents take actions and receive a reward, agents update their own CTTs. Therefore, agents could use this method to reach a consensus more quickly to enhance learning efficiency and reduce the occurrence of stagnation. The simulation results demonstrate that the proposed method can speed up the learning process and solve the problem of huge memory space consumption to some degrees. As well, it can make agents complete the task together more efficiently. Kao-Shing Hwang, Wei-Cheng Jiang |
SMC | 1 |
| 2017 | Pheromone-Based Planning Strategies in Dyna-Q LearningabstractA Dyna-Q algorithm is known as model-based reinforcement learning, so the learning agent not only interacts with the environment to learn an optimal policy, but also builds an environmental model simultaneously. To deal with the shortage of online samples, the environmental model is introduced to achieve the goal. To enhance the efficiency of the model, this paper proposes a model shaping method to compensate for bleak states scarcely visited during neighbor information. After acquiring an accurate model, many virtual experiences are sampled from this shaping model and indirect learning is thereby performed. However, how to use the model to speed up learning is an important issue. To increase the learning speed of the Dyna-Q algorithm based on the prioritized sweeping that can actually be regarded as a breadth-first search method, this paper introduces a depth-first search method that applies the techniques of ant colony algorithms to an exploration factor for selecting candidates in indirect learning. The strategy evolves to a hybrid planning approach by proportionally interleaving executions of depth-first planning and breadth-first planning. To verify the validity and applicability of the proposed method, simulations with a mountain car and maze problem are conducted. The simulation results show that the proposed method can achieve the objectives of sample efficiency and learning acceleration for the Dyna-Q learning algorithm. Kao-Shing Hwang, Wei-Cheng Jiang |
IEEE Trans. Ind. Informatics | 1 |
| 2017 | Motion Segmentation and Balancing for a Biped Robot's Imitation LearningabstractTechniques for transferring human behaviors to robots through learning by imitation/demonstration have been the subject of much study. However, direct transfer of human motion trajectories to humanoid robots does not result in dynamically stable robot movements because of the differences in human and humanoid robot kinematics and dynamics. An imitating algorithm called posture-based imitation with balance learning (Post-BL) is proposed in this paper. This Post-BL algorithm consists of three parts: a key posture identification method is used to capture key postures as knots to reconstruct the motion imitated; a clustering method classifies key postures with high similarity; and a learning method enhances the static stability of balance during imitation. In motion reproduction, the proposed system smoothly transits between key poses and the robot learns to maintain balance by slightly adjusting the leg joints. The developed balance controller uses a reinforcement learning mechanism, which is sufficient to stabilize the robot during online imitation. The experimental results for simulation and a real humanoid robot show that the Post-BL algorithm allows demonstrated motions to be imitated balance to be preserved. Kao-Shing Hwang, Wei-Cheng Jiang, Haobin Shi |
IEEE Trans. Ind. Informatics | 1 |
| 2016 | Adaboost-like method for inverse reinforcement learningabstractReinforcement learning allows agents to use trial and error method to learn intelligent behaviors which like human beings. However, when the learning tasks become difficult, how to define the reward function is an imperative issue. So, inverse reinforcement learning is proposed to form the reward function that imitates the process of interaction between the expert and the environment. In this paper, an Adaboost-like inverse reinforcement learning methods is proposed. This method uses Adaboost classifier and upper confidence bounds to generate the reward function for a complex task. In the imitating process, the agent continuously compares the difference between itself and the expert, and then the difference decides a specific weight for each state through Adaboost classifier. The weight combines with state confidence by upper confidence bounds to form an approximate reward function. Finally, a simulation, maze environment is used to demonstrate that the proposed method can decrease the computation time. Kao-Shing Hwang, Hsuan-yi Chiang, Wei-Cheng Jiang |
FUZZ-IEEE | 1 |
| 2015 | Image Based Visual Servoing Using Proportional Controller with CompensatorabstractThe main objective is to design a proportional controller of a robot manipulator using the fuzzy cerebellar model articulation controller based on Takagi -- Sugeno (T -- S) framework with a compensator. The controller and compensator apply in visual servoing, including system identification of image and kinematic Jacobians. The proposed approach is basically as a function of the visual error and extent from the error with respect to desire visual feature. This approach leads to enormous reduction on computational expense compared to the image-based approaches of model inverse kinematics. The design of the controller architecture will make it possible to implement in general case. Proportional control variable are learned offline with the help of FCMAC-T-S model, and online compensator scheme has been proposed for adapting possible uncertainties in the unknown system and environment. Stimulation results have shown that visual servoing for tracking static target can be achieved using the proposed controller with compensator. Kao-Shing Hwang, Ming-Han Chung, Wei-Cheng Jiang |
SMC | 1 |
| 2015 | Model Learning and Knowledge Sharing for a Multiagent System With Dyna-Q LearningabstractIn a multiagent system, if agents' experiences could be accessible and assessed between peers for environmental modeling, they can alleviate the burden of exploration for unvisited states or unseen situations so as to accelerate the learning process. Since how to build up an effective and accurate model within a limited time is an important issue, especially for complex environments, this paper introduces a model-based reinforcement learning method based on a tree structure to achieve efficient modeling and less memory consumption. The proposed algorithm tailored a Dyna-Q architecture to multiagent systems by means of a tree structure for modeling. The tree-model built from real experiences is used to generate virtual experiences such that the elapsed time in learning could be reduced. As well, this model is suitable for knowledge sharing. This paper is inspired by the concept of knowledge sharing methods in multiagent systems where an agent could construct a global model from scattered local models held by individual agents. Consequently, it can increase modeling accuracy so as to provide valid simulated experiences for indirect learning at the early stage of learning. To simplify the sharing process, the proposed method applies resampling techniques to grafting partial branches of trees containing required and useful experiences disseminated from experienced peers, instead of merging the whole trees. The simulation results demonstrate that the proposed sharing method can achieve the objectives of sample efficiency and learning acceleration in multiagent cooperation applications. Kao-Shing Hwang, Wei-Cheng Jiang |
IEEE Trans. Cybern. | 1 |
| 2015 | Learning to Adjust and Refine Gait Patterns for a Biped RobotabstractIn this paper, a reinforced learning method for biped walking is proposed, where the robot learns to appropriately modulate an observed walking pattern. The biped robot was equipped with two Q -learning mechanisms. First, the robot learns a policy to adjust a defective walking pattern, gait-by-gait, into a more stable one. To avoid the complexity of adjusting too many joints of a humanoid robot and to speed up the learning process, the dimensionality of the action space was reduced. In turn, the other learning mechanism trained the robot to walk in a refined pattern, allowing it to walk faster without the loss of other required criteria, such as walking straight. This approach was implemented with both a simulated robot model and an actual biped robot. The results from the simulations and experiments show that successful walking policies were obtained. The learning system works quickly enough so that the robot was able to continually adapt to the terrain as it walked. Kao-Shing Hwang, Jin-Ling Lin, Keng-Hao Yeh |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2014 | Reward shaping for reinforcement learning by emotion expressionsabstractIn this paper, a non-expert learning system was proposed to guide the robots learn their behaviors by humans' emotional expressions. The proposed system used interval fuzzy type-2 algorithm to recognize the human's facial expressions, which were captured by a web camera. Furthermore, emotion value (E-value), generated based on non-expert human's facial expressions, was applied to the reinforcement learning to train robots. Two kinds of problems were experimented. One was the human being know the exact solution to train robots and could clearly observe good or bad choice robots had been made. The other one was human being did not know the exact solution but robots could still learn from human's experience. The experiment results show that no matter the learning environment could be clearly observed by human being or not, robots could learn from human's facial expressions by the proposed learning system. Kao-Shing Hwang, J. L. Ling, Yu-Ying Chen, Wei-Han Wang |
SMC | 1 |
| 2014 | A Simple Scheme for Formation Control Based on Weighted Behavior LearningabstractSeveral correlated issues of autonomy and simplicity regarding formation control for robots with a self-awareness mechanism in unstructured environments are considered. To achieve autonomy and simplicity, a hybrid scheme is derived for robot maneuvering based on a multibehavioral system. The system holds some self-awareness capabilities ensuring precision and robustness in the presence of internal and external disturbances within the limited capacity of interrobot communication. This is to ensure that the robots can march and simultaneously maintain their assigned formation and avoid hazardous collisions on the way to their destination. These self-awareness capabilities are achieved through a layered reinforcement learning algorithm. At the bottom level, robots are equipped with a set of primitive behaviors learned prior to team formation. The high-level combined behavior is generated by multiplying the outputs of each primitive behavior by its weight, and then summing and normalizing the results. The weights keep adaptively adjusting to the proposed reinforcement learning method. Once a robot receives a command providing the formation shape and the location of the destination, the robot approaches the destination autonomously and keeps an appropriate distance from its neighbors to maintain the assigned pattern. Leadership was given to a robot occupying the lead position. The volunteer leader takes the responsibility for keeping the formation, reporting its existence to the other robots, and resetting the position assignment after passing obstacles. Simulations show the practicality and performance of the proposed approach in both static and dynamic obstacle environments. Jin-Ling Lin, Kao-Shing Hwang, Ya-Ling Wang 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2013 | Model-Based Indirect Learning Method Based on Dyna-Q ArchitectureabstractIn this paper, a model learning method based on tree structures is present to achieve the sample efficiency in stochastic environment. The proposed method is composed of Q-Learning algorithm to form a Dyna agent that can used to speed up learning. The Q-Learning is used to learn the policy, and the proposed method is for model learning. The model builds the environment model and simulates the virtual experience. The virtual experience can decrease the interaction between the agent and the environment and make the agent perform value iterations quickly. Thus, the proposed agent has additional experience for updating the policy. The simulation task, a mobile robot in a maze, is introduced to compare the methods, Q-Learning, Dyna-Q and the proposed method. The result of simulation confirms the proposed method that can achieve the goal of sample efficiency. Kao-Shing Hwang, Wei-Cheng Jiang, Wei-Han Wang |
SMC | 1 |
| 2013 | Adaptive Model Learning Based on Dyna-Q LearningabstractDyna-Q, a well-known model-based reinforcement learning (RL) method, interplays offline simulations and action executions to update Q functions. It creates a world model that predicts the feature values in the next state and the reward function of the domain directly from the data and uses the model to train Q functions to accelerate policy learning. In general, tabular methods are always used in Dyna-Q to establish the model, but a tabular model needs many more samples of experience to approximate the environment concisely. In this article, an adaptive model learning method based on tree structures is presented to enhance sampling efficiency in modeling the world model. The proposed method is to produce simulated experiences for indirect learning. Thus, the proposed agent has additional experience for updating the policy. The agent works backwards from collections of state transition and associated rewards, utilizing coarse coding to learn their definitions for the region of state space that tracks back to the precedent states. The proposed method estimates the reward and transition probabilities between states from past experience. Because the resultant tree is always concise and small, the agent can use value iteration to quickly estimate the Q-values of each action in the induced states and determine a policy. The effectiveness and generality of our method is further demonstrated in two numerical simulations. Two simulations, a mountain car and a mobile robot in a maze, are used to verify the proposed methods. The simulation result demonstrates that the training rate of our method can improve obviously. Kao-Shing Hwang, Wei-Cheng Jiang |
Cybern. Syst. | 1 |
| 2013 | Policy sharing between multiple mobile robots using decision trees
Kao-Shing Hwang, Wei-Cheng Jiang |
Inf. Sci. | 2 |
| 2013 | Classification-based video super-resolution using artificial neural networks
Ming-Hui Cheng, Kao-Shing Hwang, Jyh-Horng Jeng, Nai-Wei Lin |
Signal Process. | 2 |
| 2013 | Policy Improvement by a Model-Free Dyna ArchitectureabstractThe objective of this paper is to accelerate the process of policy improvement in reinforcement learning. The proposed Dyna-style system combines two learning schemes, one of which utilizes a temporal difference method for direct learning; the other uses relative values for indirect learning in planning between two successive direct learning cycles. Instead of establishing a complicated world model, the approach introduces a simple predictor of average rewards to actor-critic architecture in the simulation (planning) mode. The relative value of a state, defined as the accumulated differences between immediate reward and average reward, is used to steer the improvement process in the right direction. The proposed learning scheme is applied to control a pendulum system for tracking a desired trajectory to demonstrate its adaptability and robustness. Through reinforcement signals from the environment, the system takes the appropriate action to drive an unknown dynamic to track desired outputs in few learning cycles. Comparisons are made between the proposed model-free method, a connectionist adaptive heuristic critic, and an advanced method of Dyna-Q learning in the experiments of labyrinth exploration. The proposed method outperforms its counterparts in terms of elapsed time and convergence rate. Kao-Shing Hwang, Chia-Yue Lo |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2012 | Adaptive state aggregation for reinforcement learningabstractState partition is an important issue in reinforcement learning, because it has a significant effect on the performance. In this paper, an adaptive state partition method is presented for discretizing the state space adaptively and makes use of decision trees effectively. The proposed method splits the state space according to the temporal difference generated by the reinforcement learning. Consequently, the reinforcement learning uses the state space partitioned by the decision tree to learn the policy simultaneously. For avoiding a trivial partition, sibling nodes are pruned according to the Activity and the Reliability. A Monte-Carlo Tree Search (MCTS) is also proposed to explore the policy. A simulation for approaching goal has been conducted to demonstrate that the proposed method can achieve the design goal. Kao-Shing Hwang, Wei-Cheng Jiang |
SMC | 1 |
| 2012 | Continuous Q-Learning for Multi-Agent CooperationabstractThe structure of the proposed Q-learning with a continuous actions set, as well as the states, consists of three parts: a fuzzy quantizer, whose output is used to learn the action policy; an action evaluation module, which models and produces the expected evaluation signal; and a stochastic action selection unit that generates an action with the expectation of better performance using a probability distribution function to estimate an optimal action selection policy. Further, the algorithm is applied to cooperative tasks where two robots must consider their partner's action before taking their own actions. Conventional Q-learning requires a predefined and discrete state space but fails to identify the variances in different situations in the same state. The proposed Q-learning with the stochastic recording real-valued unit can differentiate the actions corresponding to different state inputs but categorized to the same state. Therefore, this unit can be regarded as an action evaluation module, which models and produces the expected evaluation signal, and an action selection unit that generates an action with the expectation of better performance using a probability distribution function to estimate an optimal action selection policy. The results from the simulations demonstrate better performance and applicability of the proposed learning model. Kao-Shing Hwang, Wei-Cheng Jiang, Yu-Hong Lin, Li-Hsin Lai |
Cybern. Syst. | 1 |
| 2012 | Variable patrol Planning of Multi-Robot Systems by a Cooperative Auction SystemabstractA cooperative auction system (CAS) is proposed to solve the large-scale multi-robot patrol planning problem. Each robot picks its own patrol points via the cooperative auction system and the system continuously re-auctions, based on the team work performance. The proposed method not only works in static environments but also considers variable path planning when the number of mobile robots increases or decreases during patrol. From the results of the simulation, the proposed approach demonstrates decreased time complexity, a lower routing path cost, improved balance of workload among robots, and the potential to scale to a large number of robots and is adaptive to environmental perturbations when the number of robots changes during patrol. Jin-Ling Lin, Kao-Shing Hwang, Hui-Ling Huang |
Cybern. Syst. | 2 |
| 2012 | Induced states in a decision tree constructed by Q-learning
Kao-Shing Hwang, Wei-Cheng Jiang, Tsung-Wen Yang |
Inf. Sci. | 1 |
| 2012 | Fusion of Multiple Behaviors Using Layered Reinforcement LearningabstractThis study introduces a method to enable a robot to learn how to perform new tasks through human demonstration and independent practice. The proposed process consists of two interconnected phases; in the first phase, state-action data are obtained from human demonstrations, and an aggregated state space is learned in terms of a decision tree that groups similar states together through reinforcement learning. Without the postprocess of trimming, in tree induction, the tree encodes a control policy that can be used to control the robot by means of repeatedly improving itself. Once a variety of behaviors is learned, more elaborate behaviors can be generated by selectively organizing several behaviors using another Q-learning algorithm. The composed outputs of the organized basic behaviors on the motor level are weighted using the policy learned through Q-learning. This approach uses three diverse Q-learning algorithms to learn complex behaviors. The experimental results show that the learned complicated behaviors, organized according to individual basic behaviors by the three Q-learning algorithms on different levels, can function more adaptively in a dynamic environment. Kao-Shing Hwang, Chun-Ju Wu |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2011 | Intelligent adaptive trajectory tracking control using fuzzy basis function networks for an autonomous small-scale helicopterabstractThis paper presents an intelligent adaptive trajectory tracking controller using fuzzy basis function networks (FBFN) for an autonomous small-scale helicopter. With the on-line FBFN approximation to the vehicle mass and the coupling effect between the force and the moments, the intelligent adaptive controller is systematically synthesized using backstepping technique. This controller is shown to achieve the semi-global ultimate boundedness of the closed-loop helicopter dynamics and accommodate agile flight maneuvers in trajectory tracking. The effectiveness and merit of the proposed method are exemplified by performing one nonlinear simulation and by performance comparison with a well-known controller. Ching-Chih Tsai 0001, Chi-Tai Lee, Kao-Shing Hwang |
SMC | 3 |
| 2011 | Self-organizing state aggregation for architecture design of Q-learning
Kao-Shing Hwang, Hsin-Yi Lin, Yuan-Pao Hsu, Hung-Hsiu Yu |
Inf. Sci. | 1 |
| 2010 | An adaptive state aggregation approach to Q-learning with real-valued action functionabstractThe fundamental approach of Q-learning is based on finite discrete state spaces, and incrementally estimating Q-values based on the reward received from the environment and the agent's previous Q-value estimates. Unfortunately, robots always learn and behave in a continuous perceptual space where the observed perceptions are transformed into or coarsely regarded as states. Nowadays, there is no elegant way to combine discrete actions with continuous observations or states. Therefore, accommodating continuous states with a finite discrete set of actions has become an important and intriguing issue in this research area. We proposed an algorithm to define an action policy from a discrete space to a real valued domain; that is, the method selects a real-valued action from a discrete set, the magnitude of which is immediately imposed a slight bias before this determined action is taken. From the viewpoint of exploration and exploitation, the method searches for a better action based on a paradigm action in the solution space with a variation within the biased region. Further, the proposed method uses the renown epsilon-greedy to explore a better trace but with a narrowized Tabu search. Kao-Shing Hwang |
SMC | 1 |
| 2010 | Dyna-like reinforcement learning based on accumulative and average rewardsabstractAn approach to accelerating the learning process of the actor-critic learning algorithm for reinforcement learning is presented. The algorithm was derived from principles based on the prediction of average rewards and temporal difference (TD) learning with averaged and discounted rewards. The derived algorithm was applied to neural networks, demonstrating their effective operation in nonlinear control problems. The motivation of the proposed algorithm was to elaborate how a learning scheme, implemented by artificial neural networks (ANNs), can speed up learning processes based on an arrangement akin to the Dyna-Q learning, where a simulative model of the controlled plant is established for virtual learning between two control cycles. Instead of modeling the complicated plant, the approach just introduced a simple predictor of rewards for virtual learning in simulation mode. Two TD learning methods based discounted and averaged rewards respectively, are used alternatively in the control and simulation mode to facilitate the derived algorithm. The proposed Alternative Learning Critic (ALC) algorithm consists of two sub-systems: one is Evaluation Predictor (EP), which performs an approximation of a long-term evaluation function, and the other is an immediate action selector, which is composed of two ANNs: Action Controller (AC) and Reinforcement Predictor (RP). The proposed learning scheme is then applied to control a pendulum system for tracking a desired trajectory to demonstrate its applausive performance and robustness. Through reinforcement signals from the environment, the system takes an appropriate action to a plant with unknown dynamics so the actual output of the plant can track the desired one concisely within a short learning cycles. Further, ALC is used as the compensator of a PI controller, which is actually only working well on a linear system, to control that pendulum system. The results show the affined system, trained ALC and the PI controller can manipulate together on a nonlinear system with unknown dynamics. Kao-Shing Hwang, Chia-Yue Lo |
SMC | 1 |
| 2009 | Behavioral-Fusion Control Based on Reinforcement LearningabstractTo design appropriate actions of mobile robots, the designers usually observe the sensory signals on the robots and decide the actions from the viewpoint of some desired purposes. This approach needs deliberative consideration and abundant knowledge on robotics for a variety of situations. To improve the actions of robots, it is hard to sense the error by human eyes and takes time in trial-and-error. In this article, we propose a novel learning algorithm, fused behavior Q-learning algorithm (FBQL) to deal with such situations. The proposed algorithm has the merit of simplicity in designing individual behavior by means of a decision tree approach to state aggregation which is eventually recoding the domain knowledge. Furthermore, these learned behaviors are fused into a more complicated behavior by a set of appropriate weighting parameters through a Q-learning mechanism such that the robots can behave adaptively and optimally in a dynamic environment. Kao-Shing Hwang, Chun-Ju Wu, Cheng-Shong Wu |
SMC | 1 |
| 2009 | Real-Valued Q-learning in Multi-agent CooperationabstractIn this paper, we propose a Q-learning with continuous action policy and extend this algorithm to a multi-agent system. We examine this algorithm in a task that there are two robots taking action independently but connected with a straight bar. The robots must cooperate to move to the goal and avoid the obstacles in the environment. Conventional Q-learning needs a pre-defined and discrete state space but fails to identify the variances of the different situation in the same state. We introduce a stochastic recording real-valued unit to Q-learning to differentiate the actions corresponding to different state inputs but categorized to the same state. This unit can be regarded as an action evaluation module, which models and produces the expected evaluation signal and an action selection unit that generates an action with the expectation of better performance using a probability distribution function that estimates an optimal action selection policy. The results from both the simulation and experiment demonstrate better performance and applicability of the proposed learning model. Kao-Shing Hwang, Chia-Yue Lo, Kim-Joan Chen |
SMC | 1 |
| 2009 | Cooperative Learning by Policy-Sharing in Multiple AgentsabstractReinforcement learning is one of the more prominent machine-learning technologies due to its unsupervised learning structure and ability to continually learn, even in a dynamic operating environment. Applying this learning to cooperative multi-agent systems not only allows each individual agent to learn from its own experience, but also offers the opportunity for the individual agents to learn from the other agents in the system so the speed of learning can be accelerated. In the proposed learning algorithm, an agent adapts to comply with its peers by learning carefully when it obtains a positive reinforcement feedback signal, but should learn more aggressively if a negative reward follows the action just taken. These two properties are applied to develop the proposed cooperative learning method. This research presents the novel use of the fastest policy hill-climbing methods of Win or Lose Fast (WoLF) with policy-sharing. Results from the multi-agent cooperative domain illustrate that the proposed algorithms perform better than Q-learning alone in a piano mover environment. It also demonstrates that agents can learn to accomplish a task together efficiently through repetitive trials. Kao-Shing Hwang, Chia-Ju Lin, Chia-Yue Lo |
Cybern. Syst. | 1 |
| 2007 | An ARM-Based Q-Learning Algorithm
Yuan-Pao Hsu, Kao-Shing Hwang, Hsin-Yi Lin |
ICIC (3) | 2 |
| 2007 | Cooperation Between Multiple Agents Based on Partially Sharing Policy
Kao-Shing Hwang, Chia-Ju Lin, Chun-Ju Wu, Chia-Yue Lo |
ICIC (1) | 1 |
| 2007 | Fast-handoff schemes for inter-subnet handoffin IEEE 802.11 WLANs for SIP/RTP applicationsabstractBecause of the popularity of portable devices and wireless LANs based on IEEE 802.11 standard, the mobility in wireless network environment become more important. There are two important issues to be addressed in order to support real-time service across wireless networks. One is high transmission data rate; the other is low handoff delay latency. In this thesis, we propose a cross-layer fast-handoff scheme in order to provide seamless mobility support to the mobile hosts in IEEE 802.11 wireless networks. We modify enhanced Inter-Access Point Protocol (IAPP) to fit real-time traffic and extend the function of an IEEE 802.11 Access Point to co-operate for cross-layerconsiderations. The link-layer handoff latency, network-layer handoff latency and application-layer handoff latency have been improved by our proposed scheme. Cheng-Shong Wu, Ming-Ta Yang, Kao-Shing Hwang |
IWCMC | 3 |
| 2007 | Cooperation in multiple agents based on sharing policyabstractIn human society, learning is essential to intelligent behavior. However, people do not need to learn everything from scratch by their own discovery. Instead, they exchange information and knowledge with one another and learn from their peers and teachers. When a task is too complex for an individual to handle, one may cooperate with its partners in order to accomplish it. Like human society, cooperation exists in the other species, such as ants that are known to communicate about the locations of food and move it cooperatively. Using the experience and knowledge of other agents, a learning agent may learn faster, make fewer mistakes, and create rules for unstructured situations. In the proposed learning algorithm, an agent adapts to comply with its peers by learning carefully when it obtains a positive reinforcement feedback signal, but should learn more aggressively if a negative reward follows the action just taken. These two properties are applied to develop the proposed cooperative learning method conceptually. The algorithm is implemented in some cooperative tasks and demonstrates that agents can learn to accomplish a task together efficiently through a repetitive trials. Kao-Shing Hwang, Chia-Ju Lin, Chun-Ju Wu, Chia-Yue Lo |
SMC | 1 |
| 2007 | Reinforcement Learning in Strategy Selection for a Coordinated Multirobot SystemabstractThis correspondence presents a multistrategy decision making system for robot soccer games. Through reinforcement processes, the coordination between robots is learned in the course of game. Meanwhile, a better action can be granted after an iterative learning process. The experimental scenario is a five-versus-five soccer game, where the proposed system dynamically assigns each player to a position in a primitive role, such as attacker, goalkeeper, etc. The responsibility of each player varies along with the change of the role in state transitions. Therefore, the system uses several strategies, such as offensive strategy, defensive strategy, and so on, for a variety of scenarios. Thus, the decision-making mechanism can choose a better strategy according to the circumstances encountered. In each strategy, a robot should behave in coordination with its teammates and resolve conflicts aggressively. The major task assignment to robots in each strategy is simply to catch good positions. Therefore, the problem of dispatching robots to good positions in a reasonable manner should be effectively handled with. This kind of problem is similar to assignment problems in linear programming research. Utilizing the Hungarian method, each robot can be assigned to its assigned spot with minimal cost. Consequently, robots based on the proposed decision-making system can accomplish each situational task in coordination. Kao-Shing Hwang, Ching-Huang Lee |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2006 | Q-Learning with FCMAC in Multi-agent Cooperation
Kao-Shing Hwang, Tzung-Feng Lin |
ISNN (1) | 1 |
| 2006 | Self Organizing Decision Tree Based on Reinforcement Learning and its Application on State Space PartitionabstractMost of tree induction algorithms are typically based on a top-down greedy strategy that sometimes makes local optimal decision at each node. Meanwhile, this strategy may induce a larger tree than needed such that requires more redundant computation. To tackle the greedy problem, a reinforcement learning method is applied to grow the decision tree. The splitting criterion is based on long-term evaluations of payoff instead of immediate evaluations. In this work, a tree induction problem is regarded as a reinforcement learning problem and solved by the technique in that problem domain. The proposed method consists of two cycles: split estimation and tree growing. In split estimation cycle, an inducer estimates long-term evaluations of splits at visited nodes. In the second cycle, the inducer grows the tree by the learned long-term evaluations. A comparison with CART on several datasets is reported. The proposed method is then applied to tree-based reinforcement learning. The state spare partition in a critic actor model, adaptive heuristic critic (AHC), is replaced by a regression tree, which is constructed by the proposed method. The experimental results are also demonstrated to show the feasibility and high performance of the proposed system. Kao-Shing Hwang, Tsung-Wen Yang, Chia-Ju Lin |
SMC | 1 |
| 2005 | Multi-agent Congestion Control for High-Speed Networks Using Reinforcement Co-learning
Kao-Shing Hwang, Mingchang Hsiao, Cheng-Shong Wu, Shunwen Tan |
ISNN (3) | 1 |
| 2005 | A Reinforcement Learning Approach To Congestion Control Of High-Speed Multimedia NetworksabstractA reinforcement learning scheme on congestion control in a high-speed network is presented. Traditional methods for congestion control always monitor the queue length, on which the source rate depends. However, the determination of the congested threshold and sending rate is difficult to couple with each other in these methods. We proposed a simple and robust reinforcement learning congestion controller (RLCC) to solve the problem. The scheme consists of two subsystems: the expectation-return predictor is a long-term policy evaluator and the other is a short-term rate selector, which is composed of action-value evaluator and stochastic action selector elements. RLCC receives reinforcement signals generated by an immediate reward evaluator and takes the best action to control source flow in consideration of high throughput and low cell loss rate. Through on-line learning processes, RLCC can adaptively take more and more correct actions under time-varying environments. Simulation results have shown that the proposed approach can increase system utilization and decrease packet losses simultaneously in comparison with the popular best-effort scheme. Ming-Chang Shaio, Shun-Wen Tan, Kao-Shing Hwang, Cheng-Shong Wu |
Cybern. Syst. | 3 |
| 2005 | Cooperative multiagent congestion control for high-speed networksabstractAn adaptive multiagent reinforcement learning method for solving congestion control problems on dynamic high-speed networks is presented. Traditional reactive congestion control selects a source rate in terms of the queue length restricted to a predefined threshold. However, the determination of congestion threshold and sending rate is difficult and inaccurate due to the propagation delay and the dynamic nature of the networks. A simple and robust cooperative multiagent congestion controller (CMCC), which consists of two subsystems: a long-term policy evaluator, expectation-return predictor and a short-term rate selector composed of action-value evaluator and stochastic action selector elements has been proposed to solve the problem. After receiving cooperative reinforcement signals generated by a cooperative fuzzy reward evaluator using game theory, CMCC takes the best action to regulate source flow with the features of high throughput and low packet loss rate. By means of learning procedures, CMCC can learn to take correct actions adaptively under time-varying environments. Simulation results showed that the proposed approach can promote the system utilization and decrease packet losses simultaneously. Kao-Shing Hwang, Shun-Wen Tan, Ming-Chang Hsiao, Cheng-Shong Wu |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2004 | Cooperative strategy based on adaptive Q-learning for robot soccer systemsabstractThe objective of this paper is to develop a self-learning cooperative strategy for robot soccer systems. The strategy enables robots to cooperate and coordinate with each other to achieve the objectives of offense and defense. Through the mechanism of learning, the robots can learn from experiences in either successes or failures, and utilize these experiences to improve the performance gradually. The cooperative strategy is built using a hierarchical architecture. The first layer of the structure is responsible for assigning each role, that is, how many defenders and sidekicks should be played according to the positional states. The second layer is for the role assignment related to the decision from the previous layer. We develop two algorithms for assignment of the roles, the attacker, the defenders, and the sidekicks. The last layer is the behavior layer in which robots execute their behavior commands and tasks based on their roles. The attacker is responsible for chasing the ball and attacking. The sidekicks are responsible for finding good positions, and the defenders are responsible for defending competitor scoring. The robots' roles are not fixed. They can dynamically exchange their roles with each other. In the aspect of learning, we develop an adaptive Q-learning method which is modified form the traditional Q-learning. A simple ant experiment shows that Q-learning is more effective than the traditional techniques, and it is also successfully applied to the learning of the cooperative strategy. Kao-Shing Hwang, Shun-Wen Tan, Chien-Cheng Chen |
IEEE Trans. Fuzzy Syst. | 1 |
| 2003 | Reinforcement learning congestion controller for multimedia surveillance systemabstractThe use of reinforcement learning scheme for congestion control in factory surveillance network is presented in this paper. Traditional methods perform congestion control by means of monitoring the queue length. When the queue length is greater than a predefined threshold, the source rate is decreased at a fixed rate. However, the determination of the congested threshold and sending rate is difficult for these methods. We adopted a simple reinforcement learning method, called Adaptive Heuristic Critic (AHC), to solve the problem. The AHC controller maintains an expectation of reward and takes the best policy to control source flow. By way of learning and then taking right actions, simulation results have shown that the approach can promote the system utilization and decrease packet loss. Ming-Chang Hsiao, Kao-Shing Hwang, Shun-Wen Tan, Cheng-Shong Wu |
ICRA | 2 |
| 2003 | Reinforcement learning to adaptive control of nonlinear systemsabstractBased on the feedback linearization theory, this paper presents how a reinforcement learning scheme that is adopted to construct artificial neural networks (ANNs) can linearize a nonlinear system effectively. The proposed reinforcement linearization learning system (RLLS) consists of two sub-systems: The evaluation predictor (EP) is a long-term policy selector, and the other is a short-term action selector composed of linearizing control (LC) and reinforce predictor (RP) elements. In addition, a reference model plays the role of the environment, which provides the reinforcement signal to the linearizing process. The RLLS thus receives reinforcement signals to accomplish the linearizing behavior to control a nonlinear system such that it can behave similarly to the reference model. Eventually, the RLLS performs identification and linearization concurrently. Simulation results demonstrate that the proposed learning scheme, which is applied to linearizing a pendulum system, provides better control reliability and robustness than conventional ANN schemes. Furthermore, a PI controller is used to control the linearized plant where the affine system behaves like a linear system. Kao-Shing Hwang, Shun-Wen Tan, Min-Cheng Tsai |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2002 | Hybrid differential evolution with multiplier updating method for nonlinear constrained optimization problemsabstractIn this paper, we introduce hybrid differential evolution with a multiplier updating method to solve constrained optimization problems. An adaptive scheme for penalty parameters is involved in the proposed algorithm so that smaller penalty parameters can be used and does not affect the final search results. Computational examples reveal that nearly identical minimum solutions can be obtained using the proposed algorithm even under wide variation of the initial penalty parameters. Yung-Chien Lin, Kao-Shing Hwang, Feng-Sheng Wang |
IEEE Congress on Evolutionary Computation | 2 |
| 2002 | A Self-Improving Fuzzy Cerebellar Model Articulation Controller with Stochastic Action GenerationabstractA modified fuzzy cerebellar model articulation controller (FCMAC) with reinforcement learning capability is introduced in this article. This model utilizes the likelihood scheme to predict the evaluation of successive actions. Based on an approximating evaluation model, the proper output (action) is always selected. The structure of the proposed FCMAC consists of three parts: a fuzzy quantizer, which is used to represent the associative mapping function from the receptive field to the actual memory; an action evaluation module, which models and produces the expected evaluation signal and an action selection unit that generates an action with the expectation of better performance using a probability distribution function that estimates an optimal action selection policy. To demonstrate its excellent performance, the proposed self-improving model is implemented as a neural network controller for the swing control of a pendulum system. The results from both the simulation and experiment demonstrates better performance and applicability of the proposed learning model. Kao-Shing Hwang, Yuan-Pao Hsu |
Cybern. Syst. | 1 |
| 2000 | A new CMAC neural network architecture and its ASIC realizationabstract-, -" " " ( " $# , " .#$ , " , ( .* , / ( , ( " " 0 (( , #$ ,( ( ( " / ! ( )*+ ( " 1 ( "-)2+ .. .* , 3 (( 4 ( $ ( ( )5+ " ( 0 ,( , ( #$ , " "3 , 670! , #$ ( , (" " (( %$#& $# , ( Yuan-Bao Hsu, Kao-Shing Hwang, Chien-Yuan Pao, Jinn-Shyan Wang |
ASP-DAC | 2 |
| 2000 | Plant scheduling and planning using mixed-integer hybrid differential evolution with multiplier updatingabstractPlant scheduling and planning are two of the most important decision-making problems in manufacturing industry. In general, these two decision-making problems are complex, due to the features of combinatorial nature for production-strategy selection and coupling properties for constrained requirements. In this paper, we have developed two general mixed-integer nonlinear programming models to formulate the scheduling and planning problems. In order to obtain a global solution, mixed-integer hybrid differential evolution with a multiplier updating method is introduced to solve both constrained problems. The proposed method can use parameters to obtain a feasible solution as compared with the penalty function approach. Yung-Chien Lin, Kao-Shing Hwang, Feng-Sheng Wang |
CEC | 2 |
| 2000 | Reinforcement Linearization Control SystemabstractThe objective of the article is to provide an effective linearization control approach for a nonlinear system. Three reinforcement back propagation learning algorithms (RBPs), based on different step-ahead predictions, are proposed to build the affine linear model of a nonlinear system by means of a composed neural network structure. The approach is used to cancel the effect of nonlinearity of a plant. Reinforcement back propagations can compensate the nonlinearity of the system dynamics between the outputs of the reference model and the system responses. In other words, the role of the composed neural plant is to perform model matching for a linearized system. Based on the derivation of RBPs, a synthetic model, a reinforcement nonlinear control system (RNCS) is developed. This scheme excels the conventional approaches and RBPs. The proposed learning schemes are implemented to linearize a pendulum system. The simulation has been done to illustrate the performance of the proposed learning schemes. Kao-Shing Hwang, Horng-Jen Chao |
Cybern. Syst. | 1 |
| 1999 | A hybrid method of evolutionary algorithms for mixed-integer nonlinear optimization problemsabstractA hybrid method of evolutionary algorithms, called mixed-integer hybrid differential evolution (MIHDE), is proposed in this study. In the hybrid method, a mixed coding is used to represent the continuous and discrete variables. A rounding operation in the mutation is introduced to handle the integer variables so that the method is not only used to solve mixed-integer nonlinear optimization problems, but also used to solve the real or integer nonlinear optimization problems. The accelerated phase and migrating phase are implemented in MIHDE. These two phases acted as a balancing operator are used to explore the search space and to exploit the best solution. Both examples of mechanical design are tested by the MIHDE. The computation results demonstrate that the MIHDE is superior to other methods in terms of solution quality and robustness property. Yung-Chien Lin, Feng-Sheng Wang, Kao-Shing Hwang |
CEC | 3 |
| 1998 | A multi-agent architecture for mobile robot navigation controlabstractFor mobile robot navigation, controlling the components such as sensors and motors is not enough. Besides controlling such components, a navigation control system needs to provide decisions in choosing a path to go and carrying out requests from the user. There are many tasks which need to be performed concurrently. Intelligent agents are suitable for being responsible for carrying out each task and performing cooperation between the agents. The paper describes an architecture which uses agents in a cooperative environment. Jia-Houng Shyu, Alan Liu, Kao-Shing Hwang |
ICTAI | 3 |
| 1998 | Reinforcement nonlinear control systemabstractIn the paper a method called the reinforcement nonlinear control system (RNCS) is proposed for intelligent control of nonlinear systems. The RNCS integrates two artificial neural networks (ANNs) into a learning system. One of the ANNs plays a role of identification predictor and the other as a linearization controller. The RNCS is trained to linearize a plant by a desired linear dynamics imposed. The plant-network system is treated as a linear system and the wealth of linear control theories can be applied after linearizing the control affine plant. Kao-Shing Hwang, Hung-Jen Chao |
SMC | 1 |
| 1998 | Smooth trajectory tracking of three-link robot: a self-organizing CMAC approachabstractA neuro fuzzy system which is embedded in the conventional control theory is proposed to tackle physical learning control problems. The control scheme is composed of two elements. The first element, the fuzzy sliding mode controller (FSMC), is used to drive the state variables to a specific switching hyperplane or a desired trajectory. The second one is developed based on the concept of the self organizing fuzzy cerebellar model articulation controller (FCMAC) and adaptive heuristic critic (AHC). Both compose a forward compensator to reduce the chattering effect or cancel the influence of system uncertainties. A geometrical explanation on how the FCMAC algorithm works is provided and some refined procedures of the AHC are presented as well. Simulations on smooth motion of a three-link robot is given to illustrate the performance and applicability of the proposed control scheme. Kao-Shing Hwang, Ching-Shun Lin |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 1991 | Analysis of voluntary movements for robotic controlabstractThis study focuses on the characterization and control of voluntary movements by analyzing motor command patterns, ensuing analysis from involuntary limb movements. As a result, a neuromuscular-like model with two muscle-reflex models is proposed, and motor command patterns are approximated identified from the experimental data. Based on the identified motor command patterns, a control strategy, muscle modulation control, is proposed for possible applications in robotic control using neuromuscular-like models.> Chi-Haur Wu, Kuu-Young Young, Kao-Shing Hwang |
ICRA | 3 |