Eiji Uchibe

dblp:37/6513 · DBLP profile ↗
← Back
49ranked-venue papers
18as first author
10since 2021 · last 2025
0000-0001-7908-0258ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 48 · 17 first-author · 10 since 2021Systems, architecture and hardware · 9 · 6 first-authorDatabases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Theoretical Guarantees for Minimum Bayes Risk Decoding
abstract
Minimum Bayes Risk (MBR) decoding optimizes output selection by maximizing the expected utility value of an underlying human distribution.While prior work has shown the effectiveness of MBR decoding through empirical evaluation, few studies have analytically investigated why the method is effective.As a result of our analysis, we show that, given the size n of the reference hypothesis set used in computation, MBR decoding approaches the optimal solution with high probability at a rate of O n -1 2 , under certain assumptions, even though the language space Y is significantly larger |Y| ≫ n.This result helps to theoretically explain the strong performance observed in several prior empirical studies on MBR decoding.In addition, we provide the performance gap for maximum-a-posteriori (MAP) decoding and compare it to MBR decoding.The result of this paper indicates that MBR decoding tends to converge to the optimal solution faster than MAP decoding in several cases.
Yuki Ichihara, Yuu Jinnai, Kaito Ariu, Tetsuro Morimura, Eiji Uchibe
ACL (1)5
2025 Human-in-the-Loop Generative Policy Learning from Demonstrations and Preferences
Eiji Uchibe
ICONIP (5)1
2024 Reward-Punishment Reinforcement Learning with Maximum Entropy
abstract
We introduce the "soft Deep MaxPain" (softDMP) algorithm, which integrates the optimization of long-term policy entropy into reward-punishment reinforcement learning objectives. Our motivation is to facilitate a smoother variation of operators utilized in the updating of action values beyond traditional "max" and "min" operators, where the goal is enhancing sample efficiency and robustness. We also address two unresolved issues from the previous Deep MaxPain method. Firstly, we investigate how the negated ("flipped") pain-seeking sub-policy, derived from the punishment action value, collaborates with the "min" operator to effectively learn the punishment module and how softDMP’s smooth learning operator provides insights into the "flipping" trick. Secondly, we tackle the challenge of data collection for learning the punishment module to mitigate inconsistencies arising from the involvement of the "flipped" sub-policy (pain-avoidance sub-policy) in the unified behavior policy. We empirically explore the first issue in two discrete Markov Decision Process (MDP) environments, elucidating the crucial advancements of the DMP approach and the necessity for soft treatments on the hard operators. For the second issue, we propose a probabilistic classifier based on the ratio of the pain-seeking sub-policy to the sum of the pain-seeking and goalreaching sub-policies. This classifier assigns roll-outs to separate replay buffers for updating reward and punishment action-value functions respectively. Our framework demonstrates superior performance in Turtlebot 3’s maze navigation tasks under the ROS Gazebo simulation.
Jiexin Wang 0001, Eiji Uchibe
IJCNN2
2024 Estimating cost function of expert players in differential games: A model-based method and its data-driven extension
abstract
In this paper, we introduce two algorithms for estimating the cost function of expert players engaged in optimal performance within linear continuous-time differential games. Initially, we propose a model-based algorithm, followed by its data-driven model-free extension. Both methods rely on optimal policy gains, obtained by observing Nash equilibrium trajectories of the expert players. The model-free method also utilizes the trajectories of the learner system. This method addresses the limitations found in existing model-free approaches, which may suffer from either high computational costs, limited applicability to specific systems, or both. The effectiveness of the proposed methods is demonstrated through numerical simulations.
Hamed Jabbari, Eiji Uchibe
Expert Syst. Appl.2
2024 Online estimation of objective function for continuous-time deterministic systems
abstract
We developed two online data-driven methods for estimating an objective function in continuous-time linear and nonlinear deterministic systems. The primary focus addressed the challenge posed by unknown input dynamics (control mapping function) in the expert system, a critical element for an online solution of the problem. Our methods leverage both the learner's and expert's data for effective problem-solving. The first approach, which is model-free, estimates the expert's policy and integrates it into the learner agent to approximate the objective function associated with the optimal policy. The second approach estimates the input dynamics from the learner's data and combines it with the expert's input-state observations to tackle the objective function estimation problem. Compared to other methods for deterministic systems that rely on both the learner's and expert's data, our approaches offer reduced complexity by eliminating the need to estimate an optimal policy after each objective function update. We conduct a convergence analysis of the estimation techniques using Lyapunov-based methods. Numerical experiments validate the effectiveness of our developed methods.
Hamed Jabbari, Eiji Uchibe
Neural Networks2
2023 Online Reinforcement Learning Control of Nonlinear Dynamic Systems: A State-action Value Function Based Solution
Hamed Jabbari, Eiji Uchibe
Neurocomputing2
2022 Deep learning, reinforcement learning, and world models
abstract
Deep learning (DL) and reinforcement learning (RL) methods seem to be a part of indispensable factors to achieve human-level or super-human AI systems. On the other hand, both DL and RL have strong connections with our brain functions and with neuroscientific findings. In this review, we summarize talks and discussions in the "Deep Learning and Reinforcement Learning" session of the symposium, International Symposium on Artificial Intelligence and Brain Science. In this session, we discussed whether we can achieve comprehensive understanding of human intelligence based on the recent advances of deep learning and reinforcement learning algorithms. Speakers contributed to provide talks about their recent studies that can be key technologies to achieve human-level intelligence.
Yutaka Matsuo, Yann LeCun, Maneesh Sahani, Doina Precup, David Silver 0001, Masashi Sugiyama, Eiji Uchibe, Jun Morimoto
Neural Networks7
2021 Parallel and hierarchical neural mechanisms for adaptive and predictive behavioral control
abstract
Our brain can be recognized as a network of largely hierarchically organized neural circuits that operate to control specific functions, but when acting in parallel, enable the performance of complex and simultaneous behaviors. Indeed, many of our daily actions require concurrent information processing in sensorimotor, associative, and limbic circuits that are dynamically and hierarchically modulated by sensory information and previous learning. This organization of information processing in biological organisms has served as a major inspiration for artificial intelligence and has helped to create in silico systems capable of matching or even outperforming humans in several specific tasks, including visual recognition and strategy-based games. However, the development of human-like robots that are able to move as quickly as humans and respond flexibly in various situations remains a major challenge and indicates an area where further use of parallel and hierarchical architectures may hold promise. In this article we review several important neural and behavioral mechanisms organizing hierarchical and predictive processing for the acquisition and realization of flexible behavioral control. Then, inspired by the organizational features of brain circuits, we introduce a multi-timescale parallel and hierarchical learning framework for the realization of versatile and agile movement in humanoid robots.
Tom Macpherson, Masayuki Matsumoto, Hiroaki Gomi, Jun Morimoto, Eiji Uchibe, Takatoshi Hikida
Neural Networks5
2021 Forward and inverse reinforcement learning sharing network weights and hyperparameters
abstract
This paper proposes model-free imitation learning named Entropy-Regularized Imitation Learning (ERIL) that minimizes the reverse Kullback-Leibler (KL) divergence. ERIL combines forward and inverse reinforcement learning (RL) under the framework of an entropy-regularized Markov decision process. An inverse RL step computes the log-ratio between two distributions by evaluating two binary discriminators. The first discriminator distinguishes the state generated by the forward RL step from the expert's state. The second discriminator, which is structured by the theory of entropy regularization, distinguishes the state-action-next-state tuples generated by the learner from the expert ones. One notable feature is that the second discriminator shares hyperparameters with the forward RL, which can be used to control the discriminator's ability. A forward RL step minimizes the reverse KL estimated by the inverse RL step. We show that minimizing the reverse KL divergence is equivalent to finding an optimal policy. Our experimental results on MuJoCo-simulated environments and vision-based reaching tasks with a robotic arm show that ERIL is more sample-efficient than the baseline methods. We apply the method to human behaviors that perform a pole-balancing task and describe how the estimated reward functions show how every subject achieves her goal.
Eiji Uchibe, Kenji Doya
Neural Networks1
2021 Modular deep reinforcement learning from reward and punishment for robot navigation
abstract
Modular Reinforcement Learning decomposes a monolithic task into several tasks with sub-goals and learns each one in parallel to solve the original problem. Such learning patterns can be traced in the brains of animals. Recent evidence in neuroscience shows that animals utilize separate systems for processing rewards and punishments, illuminating a different perspective for modularizing Reinforcement Learning tasks. MaxPain and its deep variant, Deep MaxPain, showed the advances of such dichotomy-based decomposing architecture over conventional Q-learning in terms of safety and learning efficiency. These two methods differ in policy derivation. MaxPain linearly unified the reward and punishment value functions and generated a joint policy based on unified values; Deep MaxPain tackled scaling problems in high-dimensional cases by linearly forming a joint policy from two sub-policies obtained from their value functions. However, the mixing weights in both methods were determined manually, causing inadequate use of the learned modules. In this work, we discuss the signal scaling of reward and punishment related to discounting factor γ, and propose a weak constraint for signaling design. To further exploit the learning models, we propose a state-value dependent weighting scheme that automatically tunes the mixing weights: hard-max and softmax based on a case analysis of Boltzmann distribution. We focus on maze-solving navigation tasks and investigate how two metrics (pain-avoiding and goal-reaching) influence each other's behaviors during learning. We propose a sensor fusion network structure that utilizes lidar and images captured by a monocular camera instead of lidar-only and image-only sensing. Our results, both in the simulation of three types of mazes with different complexities and a real robot experiment of an L-maze on Turtlebot3 Waffle Pi, showed the improvements of our methods.
Jiexin Wang 0001, Stefan Elfwing, Eiji Uchibe
Neural Networks3
2019 Theoretical Analysis of Efficiency and Robustness of Softmax and Gap-Increasing Operators in Reinforcement Learning
abstract
In this paper, we propose and analyze conservative value iteration, which unifies value iteration, soft value iteration, advantage learning, and dynamic policy programming. Our analysis shows that algorithms using a combination of gap-increasing and max operators are resilient to stochastic errors, but not to non-stochastic errors. In contrast, algorithms using a softmax operator without a gap-increasing operator are less susceptible to all types of errors, but may display poor asymptotic performance. Algorithms using a combination of gap-increasing and softmax operators are much more effective and may asymptotically outperform algorithms with the max operator. Not only do these theoretical results provide a deep understanding of various reinforcement learning algorithms, but they also highlight the effectiveness of gap-increasing operators, as well as the limitations of traditional greedy value updates by the max operator.
Tadashi Kozuno, Eiji Uchibe, Kenji Doya
AISTATS2
2018 Online meta-learning by parallel algorithm competition
abstract
The efficiency of reinforcement learning algorithms depends critically on a few meta-parameters that modulate the learning updates and the trade-off between exploration and exploitation. The adaptation of the meta-parameters is an open question, which arguably has become a more important issue recently with the success of deep reinforcement learning. The long learning times in domains such as Atari 2600 video games makes it not feasible to perform comprehensive searches of appropriate meta-parameter values. In this study, we propose the Online Meta-learning by Parallel Algorithm Competition (OMPAC) method, which is a novel Lamarckian evolutionary approach to online meta-parameter adaptation. The population consists of several instances of a reinforcement learning algorithm which are run in parallel with small differences in initial meta-parameter values. After a fixed number of learning episodes, the instances are selected based on their performance on the task at hand, i.e., the fitness. Before continuing the learning, Gaussian noise is added to the meta-parameters with a predefined probability. We validate the OMPAC method by improving the state-of-the-art results in stochastic SZ-Tetris and in 10x10 Tetris by 31% and 84%, respectively, and by improving the learning speed and performance for deep Sarsa(λ) agents in the Atari 2600 domain.
Stefan Elfwing, Eiji Uchibe, Kenji Doya
GECCO2
2018 Efficient sample reuse in policy search by multiple importance sampling
abstract
Policy search such as reinforcement learning and evolutionary computation is a framework for finding an optimal policy of control problems, but it usually requires a huge number of samples. Importance sampling is a common tool to use samples drawn from a proposal distribution different from the targeted one, and it is widely used by the policy search methods to update the policy from a set of datasets that are collected by previous sampling distributions. However, the proposal distribution is created by a mixture of previous distributions with fixed mixing weights in most of previous studies, and it is often numerically unstable. To overcome this problem, we propose the method of adaptive multiple importance sampling that optimizes the mixing coefficients to minimize the variance of the importance sampling estimator while utilizing as many samples as possible. We apply the proposed method to the five policy search methods such as PGPE, PoWER, CMA-ES, REPS, and NES, and their algorithms are evaluated by some benchmark control tasks. Experimental results show that all the five methods improve sample efficiency. In addition, we show that optimizing the mixing weights achieves stable learning.
Eiji Uchibe
GECCO1
2018 Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
abstract
In recent years, neural networks have enjoyed a renaissance as function approximators in reinforcement learning. Two decades after Tesauro's TD-Gammon achieved near top-level human performance in backgammon, the deep reinforcement learning algorithm DQN achieved human-level performance in many Atari 2600 games. The purpose of this study is twofold. First, we propose two activation functions for neural network function approximation in reinforcement learning: the sigmoid-weighted linear unit (SiLU) and its derivative function (dSiLU). The activation of the SiLU is computed by the sigmoid function multiplied by its input. Second, we suggest that the more traditional approach of using on-policy learning with eligibility traces, instead of experience replay, and softmax action selection can be competitive with DQN, without the need for a separate target network. We validate our proposed approach by, first, achieving new state-of-the-art results in both stochastic SZ-Tetris and Tetris with a small 10 × 10 board, using TD(λ) learning and shallow dSiLU network agents, and, then, by outperforming DQN in the Atari 2600 domain by using a deep Sarsa(λ) agent with SiLU and dSiLU hidden units.
Stefan Elfwing, Eiji Uchibe, Kenji Doya
Neural Networks2
2018 Model-Free Deep Inverse Reinforcement Learning by Logistic Regression
abstract
This paper proposes model-free deep inverse reinforcement learning to find nonlinear reward function structures. We formulate inverse reinforcement learning as a problem of density ratio estimation, and show that the log of the ratio between an optimal state transition and a baseline one is given by a part of reward and the difference of the value functions under the framework of linearly solvable Markov decision processes. The logarithm of density ratio is efficiently calculated by binomial logistic regression, of which the classifier is constructed by the reward and state value function. The classifier tries to discriminate between samples drawn from the optimal state transition probability and those from the baseline one. Then, the estimated state value function is used to initialize the part of the deep neural networks for forward reinforcement learning. The proposed deep forward and inverse reinforcement learning is applied into two benchmark games: Atari 2600 and Reversi. Simulation results show that our method reaches the best performance substantially faster than the standard combination of forward and inverse reinforcement learning as well as behavior cloning.
Eiji Uchibe
Neural Process. Lett.1
2017 Average Reward Optimization with Multiple Discounting Reinforcement Learners
Chris Reinke, Eiji Uchibe, Kenji Doya
ICONIP (1)2
2017 Deep dynamic policy programming for robot control with raw images
abstract
Deep reinforcement learning has drawn much attention in robot control since it enables agents to learn control policies from very high dimensional states such as raw images. On the other hand, its dependency upon the availability of a significant quantity of training samples and its fragility in learning makes it difficult to apply for real world robot tasks. To alleviate these issues we propose Deep Dynamic Policy Programming (DDPP), which combines the sample efficiency and smooth policy updates of dynamic policy programming with the contemporary deep reinforcement learning framework. The effectiveness of the proposed method is first demonstrated in a simulation of the robot arm control problem, with comparison to Deep Q-Networks. As validation on a real robot system, DDPP also successfully learned the flipping of a handkerchief with a NEXTAGE humanoid robot using a reduced number of learning samples, whereas Deep Q-Networks failed to learn the task.
Yoshihisa Tsurumine, Yunduan Cui, Eiji Uchibe, Takamitsu Matsubara
IROS3
2016 Deep Inverse Reinforcement Learning by Logistic Regression
Eiji Uchibe
ICONIP (1)1
2016 From free energy to expected energy: Improving energy-based value function approximation in reinforcement learning
abstract
Free-energy based reinforcement learning (FERL) was proposed for learning in high-dimensional state and action spaces. However, the FERL method does only really work well with binary, or close to binary, state input, where the number of active states is fewer than the number of non-active states. In the FERL method, the value function is approximated by the negative free energy of a restricted Boltzmann machine (RBM). In our earlier study, we demonstrated that the performance and the robustness of the FERL method can be improved by scaling the free energy by a constant that is related to the size of network. In this study, we propose that RBM function approximation can be further improved by approximating the value function by the negative expected energy (EERL), instead of the negative free energy, as well as being able to handle continuous state input. We validate our proposed method by demonstrating that EERL: (1) outperforms FERL, as well as standard neural network and linear function approximation, for three versions of a gridworld task with high-dimensional image state input; (2) achieves new state-of-the-art results in stochastic SZ-Tetris in both model-free and model-based learning settings; and (3) significantly outperforms FERL and standard neural network function approximation for a robot navigation task with raw and noisy RGB images as state input and a large number of actions.
Stefan Elfwing, Eiji Uchibe, Kenji Doya
Neural Networks2
2015 Expected energy-based restricted Boltzmann machine for classification
abstract
In classification tasks, restricted Boltzmann machines (RBMs) have predominantly been used in the first stage, either as feature extractors or to provide initialization of neural networks. In this study, we propose a discriminative learning approach to provide a self-contained RBM method for classification, inspired by free-energy based function approximation (FE-RBM), originally proposed for reinforcement learning. For classification, the FE-RBM method computes the output for an input vector and a class vector by the negative free energy of an RBM. Learning is achieved by stochastic gradient-descent using a mean-squared error training objective. In an earlier study, we demonstrated that the performance and the robustness of FE-RBM function approximation can be improved by scaling the free energy by a constant that is related to the size of network. In this study, we propose that the learning performance of RBM function approximation can be further improved by computing the output by the negative expected energy (EE-RBM), instead of the negative free energy. To create a deep learning architecture, we stack several RBMs on top of each other. We also connect the class nodes to all hidden layers to try to improve the performance even further. We validate the classification performance of EE-RBM using the MNIST data set and the NORB data set, achieving competitive performance compared with other classifiers such as standard neural networks, deep belief networks, classification RBMs, and support vector machines. The purpose of using the NORB data set is to demonstrate that EE-RBM with binary input nodes can achieve high performance in the continuous input domain.
Stefan Elfwing, Eiji Uchibe, Kenji Doya
Neural Networks2
2014 Combining learned controllers to achieve new goals based on linearly solvable MDPs
abstract
Learning complicated behaviors usually involves intensive manual tuning and expensive computational optimization because we have to solve a nonlinear Hamilton-Jacobi-Bellman (HJB) equation. Recently, Todorov proposed a class of the so-called Linearly solvable Markov Decision Process (LMDP) which converts a nonlinear HJB equation to a linear differential equation. Linearity of the simplified HJB equation allows us to apply superposition to derive a new composite controller from a set of learned primitive controllers. However, his method was a model-based approach and it was not evaluated in a real domain. This study proposes a model-free method which is similar to the Least Squares Temporal Difference (LSTD) learning. In this method, the exponentially transformed cost function can be regarded as the discount factor in LSTD. Our proposed method is applied to learning walking behaviors with the quadruped robot to evaluate in real robot experiments. The goal of each primitive task is to go to the specific target position in the environment and that of the composite task is to approach arbitrary region represented by the primitives' target positions. Experimental results show that the composite policy can be used as a good initial policy for the new task.
Eiji Uchibe, Kenji Doya
ICRA1
2010 Free-Energy Based Reinforcement Learning for Vision-Based Navigation with High-Dimensional Sensory Inputs
Stefan Elfwing, Makoto Otsuka, Eiji Uchibe, Kenji Doya
ICONIP (1)3
2010 Derivatives of Logarithmic Stationary Distributions for Policy Gradient Reinforcement Learning
abstract
Most conventional policy gradient reinforcement learning (PGRL) algorithms neglect (or do not explicitly make use of) a term in the average reward gradient with respect to the policy parameter. That term involves the derivative of the stationary state distribution that corresponds to the sensitivity of its distribution to changes in the policy parameter. Although the bias introduced by this omission can be reduced by setting the forgetting rate gamma for the value functions close to 1, these algorithms do not permit gamma to be set exactly at gamma = 1. In this article, we propose a method for estimating the log stationary state distribution derivative (LSD) as a useful form of the derivative of the stationary state distribution through backward Markov chain formulation and a temporal difference learning framework. A new policy gradient (PG) framework with an LSD is also proposed, in which the average reward gradient can be estimated by setting gamma = 0, so it becomes unnecessary to learn the value functions. We also test the performance of the proposed algorithms using simple benchmark tasks and show that these can improve the performances of existing PG methods.
Tetsuro Morimura, Eiji Uchibe, Junichiro Yoshimoto, Jan Peters 0001, Kenji Doya
Neural Comput.2
2009 Emergence of Different Mating Strategies in Artificial Embodied Evolution
Stefan Elfwing, Eiji Uchibe, Kenji Doya
ICONIP (2)2
2009 A Generalized Natural Actor-Critic Algorithm
abstract
Policy gradient Reinforcement Learning (RL) algorithms have received much attention in seeking stochastic policies that maximize the average rewards. In addition, extensions based on the concept of the Natural Gradient (NG) show promising learning efficiency because these regard metrics for the task. Though there are two candidate metrics, Kakades Fisher Information Matrix (FIM) and Morimuras FIM, all RL algorithms with NG have followed the Kakades approach. In this paper, we describe a generalized Natural Gradient (gNG) by linearly interpolating the two FIMs and propose an efficient implementation for the gNG learning based on a theory of the estimating function, generalized Natural Actor-Critic (gNAC). The gNAC algorithm involves a near optimal auxiliary function to reduce the variance of the gNG estimates. Interestingly, the gNAC can be regarded as a natural extension of the current state-of-the-art NAC algorithm, as long as the interpolating parameter is appropriately selected. Numerical experiments showed that the proposed gNAC algorithm can estimate gNG efficiently and outperformed the NAC algorithm.
Tetsuro Morimura, Eiji Uchibe, Junichiro Yoshimoto, Kenji Doya
NIPS2
2008 NeuroEvolution Based on Reusable and Hierarchical Modular Representation
Takumi Kamioka, Eiji Uchibe, Kenji Doya
ICONIP (1)2
2008 A New Natural Policy Gradient by Stationary Distribution Metric
Tetsuro Morimura, Eiji Uchibe, Junichiro Yoshimoto, Kenji Doya
ECML/PKDD (2)2
2008 Finding intrinsic rewards by embodied evolution and constrained reinforcement learning
Eiji Uchibe, Kenji Doya
Neural Networks1
2007 Finding Exploratory Rewards by Embodied Evolution and Constrained Reinforcement Learning in the Cyber Rodents
Eiji Uchibe, Kenji Doya
ICONIP (2)1
2007 Evolutionary Development of Hierarchical Learning Structures
abstract
Hierarchical reinforcement learning (RL) algorithms can learn a policy faster than standard RL algorithms. However, the applicability of hierarchical RL algorithms is limited by the fact that the task decomposition has to be performed in advance by the human designer. We propose a Lamarckian evolutionary approach for automatic development of the learning structure in hierarchical RL. The proposed method combines the MAXQ hierarchical RL method and genetic programming (GP). In the MAXQ framework, a subtask can optimize the policy independently of its parent task's policy, which makes it possible to reuse learned policies of the subtasks. In the proposed method, the MAXQ method learns the policy based on the task hierarchies obtained by GP, while the GP explores the appropriate hierarchies using the result of the MAXQ method. To show the validity of the proposed method, we have performed simulation experiments for a foraging task in three different environmental settings. The results show strong interconnection between the obtained learning structures and the given task environments. The main conclusion of the experiments is that the GP can find a minimal strategy, i.e., a hierarchy that minimizes the number of primitive subtasks that can be executed for each type of situation. The experimental results for the most challenging environment also show that the policies of the subtasks can continue to improve, even after the structure of the hierarchy has been evolutionary stabilized, as an effect of Lamarckian mechanisms
Stefan Elfwing, Eiji Uchibe, Kenji Doya, Henrik I. Christensen
IEEE Trans. Evol. Comput.2
2006 Incremental Coevolution With Competitive and Cooperative Tasks in a Multirobot Environment
abstract
Coevolution has been receiving increased attention as a method for simultaneously developing the control structures of multiple agents. Our ultimate goal is the mutual development of skills through coevolution. The coevolutionary process is, however, often prone to settle into suboptimal strategies. The key to successful coevolution has thus far been unclear. This paper discusses how several robots can emerge cooperative and competitive behavior through coevolutionary processes. In order to realize successful coevolution, we propose two ideas: multiple schedules for incremental evolution and fitness sharing based on the method of importance sampling. To examine this issue, we conducted a series of computer simulations. We have chosen a simplified soccer game consisting of two or three robots as a testbed for analyzing a problem in which both competitive and cooperative tasks are involved. We show that the proposed fitness evaluation allows robots to evolve robust behaviors in cooperative and competitive situations
Eiji Uchibe, Minoru Asada
Proc. IEEE1
2005 Biologically inspired embodied evolution of survival
abstract
Embodied evolution is a methodology for evolutionary robotics that mimics the distributed, asynchronous and autonomous properties of biological evolution. The evaluation, selection and reproduction are carried out by and between the robots, without any need for human intervention. In this paper, we propose a biologically inspired embodied evolution framework, which fully integrates self-preservation, recharging from external batteries in the environment, and self-reproduction, pair-wise exchange of genetic material, into a survival system. The individuals are explicitly evaluated for the performance of the battery capturing task, but also implicitly for the mating task by the fact that an individual that mates frequently has larger probability to spread its gene in the population. We have evaluated our method in simulation experiments and the simulation results show that the solutions obtained by our embodied evolution method were able to optimize the two survival tasks, battery capturing and mating, simultaneously. We have also performed preliminary experiments in hardware, with promising results.
Stefan Elfwing, Eiji Uchibe, Kenji Doya, Henrik I. Christensen
Congress on Evolutionary Computation2
2004 Multi-agent reinforcement learning: using macro actions to learn a mating task
abstract
Standard reinforcement learning methods are inefficient and often inadequate for learning cooperative multi-agent tasks. For these kinds of tasks the behavior of one agent strongly depends on dynamic interaction with other agents, not only with the interaction with a static environment as in standard reinforcement learning. The success of the learning is therefore coupled to the agents' ability to predict the other agents behaviors. In this study we try to overcome this problem by adding a few simple macro actions, actions that are extended in time for more than one time step. The macro actions improve the learning by making search of the state space more effective and thereby making the behavior more predictable for the other agent. In this study we have considered a cooperative mating task, which is the first step towards our aim to perform embodied evolution, where the evolutionary selection process is an integrated part of the task. We show, in simulation and hardware, that in the case of learning without macro actions, the agents fail to learn a meaningful behavior. In contrast, for the learning with macro action the agents learn a good mating behavior in reasonable time, in both simulation and hardware.
Stefan Elfwing, Eiji Uchibe, Kenji Doya, Henrik I. Christensen
IROS2
2003 An Evolutionary Approach to Automatic Construction of the Structure in Hierarchical Reinforcement Learning
Stefan Elfwing, Eiji Uchibe, Kenji Doya
GECCO2
2001 Dynamic Task Assignment in a Multiagent/Multitask Environment based on Module Conflict Resolution
abstract
It is necessary to coordinate multiple tasks in order to cope with larger-scaled and more complicated tasks. However, it seems very hard to accomplish the multiple tasks at the same time. The paper proposes a method to resolve a conflict between task modules through the processes of their executions. Based on the proposed method, the robot can select an appropriate module according to the priority. In addition, we apply the module conflict resolution to a multiagent environment. Consequently, multiple tasks are automatically allocated to the multiple robots. As a task example, a soccer game is selected to show the validity of the proposed method. Real experiments are shown, and a discussion is given.
Eiji Uchibe, Tatsunori Kato, Koh Hosoda, Minoru Asada
ICRA1
2001 Evolutionary Behavior Selection with Activation/Termination Constraints
Eiji Uchibe, Masakazu Yanase, Minoru Asada
RoboCup1
2000 Osaka University "Trackies 2000"
Yasutake Takahashi, Eiji Uchibe, Takahashi Tamura, Masakazu Yanase, Shoichi Ikenoue, Shujiro Inui, Minoru Asada
RoboCup2
1999 The Team Description of Osaka University "Trackies-99"
Sho'ji Suzuki, Tatsunori Kato, Hiroshi Ishizuka, Hiroyoshi Kawanishi, Takashi Tamura, Masakazu Yanase, Yasutake Takahashi, Eiji Uchibe, Minoru Asada
RoboCup8
1999 Multiple Reward Criterion for Cooperative Behavior Acquisition in a Muliagent Environment
Eiji Uchibe, Minoru Asada
RoboCup1
1999 Cooperative Behavior Acquisition for Mobile Robots in Dynamically Changing Real Worlds Via Vision-Based Reinforcement Learning and Development
Minoru Asada, Eiji Uchibe, Koh Hosoda
Artif. Intell.2
1998 State Space Construction for Behavior Acquisition in Multi Agent Environments with Vision and Action
abstract
This paper proposes a method which estimates the relationships between learner's behaviors and other agents' ones in the environment through interactions (observation and action) using the method of system identification. In order to identify the model of each agent, Akaike's Information Criterion is applied to the results of Canonical Variate Analysis for the relationship between the observed data in terms of action and future observation. Next, reinforcement learning based on the estimated state vectors is performed to obtain the optimal behavior. The proposed method is applied to a soccer playing situation, where a rolling ball and other moving agents are well modeled and the learner's behaviors are successfully acquired by the method. Computer simulations and real experiments are shown and a discussion is given.
Eiji Uchibe, Minoru Asada, Koh Hosoda
ICCV1
1998 Cooperative Behavior Acquisition in Multi Mobile Robots Environment by Reinforcement Learning Based on State Vector Estimation
abstract
This paper proposes a method that acquires robots' behaviors based on the estimation of the state vectors. In order to acquire the cooperative behaviors in multi-robot environments, each learning robot estimates the local predictive model between the learner and the other objects separately. Based on the local predictive models, the robots learn the desired behaviors using reinforcement learning. The proposed method is applied to a soccer playing situation, where a rolling ball and other moving robots are well modeled and the learner's behaviors are successfully acquired by the method. Computer simulations and real experiments are shown and a discussion is given.
Eiji Uchibe, Minoru Asada, Koh Hosoda
ICRA1
1998 Environmental Complexity Control for Vision-Based Learning Mobile Robot
abstract
Discusses how a robot can develop its state vector according to the complexity of the interactions with its environment. A method for controlling the complexity is proposed for a vision-based mobile robot whose task is to shoot a ball into a goal avoiding collisions with a goalkeeper. First, we provide the most difficult situation (the maximum speed of the goalkeeper with chasing-a-ball behavior), and the robot estimates the full set of state vectors with the order of the major vector components by a method of system identification. The environmental complexity is defined in terms of the speed of the goalkeeper while the complexity of the state vector is the number of the dimensions of the state vector. According to the increase of the speed of the goalkeeper, the dimension of the state vector is increased by taking a trade-off between the size of the state space (the dimension) and the learning time. Simulations are shown, and other issues for the complexity control are discussed.
Eiji Uchibe, Minoru Asada, Koh Hosoda
ICRA1
1998 Co-evolution for cooperative behavior acquisition in a multiple mobile robot environment
abstract
Co-evolution has been receiving increased attention as a method for multi agent simultaneous learning. This paper discusses how multiple robots can emerge cooperative behaviors through co-evolutionary processes. As an example task, a simplified soccer game with three learning robots is selected and a genetic programming method is applied to individual population corresponding to each robot so as to obtain cooperative and competitive behaviors. The complexity of the problem can be explained twofold: co-evolution for cooperative behaviors needs exact synchronization of mutual evolutions, and three robot co-evolution requires well-complicated environment setups that may gradually change from, simpler to more complicated situations. Simulation results are shown, and a discussion is given.
Eiji Uchibe, Masateru Nakamura, Minoru Asada
IROS1
1998 An Application of Vision-Based Learning in RoboCup for a Real Robot with an Omnidirectional Vision System and the Team Description of Osaka University "Trackies"
Sho'ji Suzuki, Tatsunori Kato, Hiroshi Ishizuka, Yasutake Takahashi, Eiji Uchibe, Minoru Asada
RoboCup5
1998 Cooperative Behavior Acquisition in a Multiple Mobile Robot Environment by Co-evolution
Eiji Uchibe, Masateru Nakamura, Minoru Asada
RoboCup1
1997 Vision-Based Robot Learning Towards RoboCup: Osaka University "Trackies"
Sho'ji Suzuki, Yasutake Takahashi, Eiji Uchibe, Masateru Nakamura, Chizuko Mishima, Hiroshi Ishizuka, Tatsunori Kato, Minoru Asada
RoboCup3
1996 Behavior coordination for a mobile robot using modular reinforcement learning
abstract
Coordination of multiple behaviors independently obtained by a reinforcement learning method is one of the issues in order for the method to be scaled to larger and more complex robot learning tasks. Direct combination of all the state spaces for individual modules (subtasks) needs enormous learning time, and it causes hidden states. This paper presents a method of modular learning which coordinates multiple behaviors taking account of a trade-off between learning time and performance. First, in order to reduce the learning time the whole state space is classified into two categories based on the action values separately obtained by Q learning: the area where one of the learned behaviors is directly applicable (no more learning area), and the area where learning is necessary due to competition of multiple behaviors (re-learning area). Second, hidden states are detected by model fitting to the learned action values based on the information criterion. Finally, the initial action valves in the re-learning area are adjusted so that they can be consistent with the values in the no more learning area. The method is applied to one to one soccer playing robots. Computer simulation and real robot experiments are given, to show the validity of the proposed method.
Eiji Uchibe, Minoru Asada, Koh Hosoda
IROS1
1994 Coordination of multiple behaviors acquired by a vision-based reinforcement learning
abstract
A method is proposed which accomplishes a whole task consisting of plural subtasks by coordinating multiple behaviors acquired by a vision-based reinforcement learning. First, individual behaviors which achieve the corresponding subtasks are independently acquired by Q-learning, a widely used reinforcement learning method. Each learned behavior can be represented by an action-value function in terms of state of the environment and robot action. Next, three kinds of coordinations of multiple behaviors are considered; simple summation of different action-value functions, switching action-value functions according to situations, and learning with previously obtained action-value functions as initial values of a new action-value function. A task of shooting a ball into the goal avoiding collisions with an enemy is examined. The task can be decomposed into a ball shooting subtask and a collision avoiding subtask. These subtasks should be accomplished simultaneously, but they are not independent of each other.>
Minoru Asada, Eiji Uchibe, Shoichi Noda, Sukoya Tawaratsumida, Koh Hosoda
IROS2