VLDB 2026 Research / reviewers in the wild / expert
Eiji Uchibe
dblp:37/6513
· DBLP profile ↗
49ranked-venue papers
18as first author
10since 2021 · last 2025
0000-0001-7908-0258ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 48 · 17 first-author · 10 since 2021Systems, architecture and hardware · 9 · 6 first-authorDatabases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Theoretical Guarantees for Minimum Bayes Risk DecodingabstractMinimum Bayes Risk (MBR) decoding optimizes output selection by maximizing the expected utility value of an underlying human distribution.While prior work has shown the effectiveness of MBR decoding through empirical evaluation, few studies have analytically investigated why the method is effective.As a result of our analysis, we show that, given the size n of the reference hypothesis set used in computation, MBR decoding approaches the optimal solution with high probability at a rate of O n -1 2 , under certain assumptions, even though the language space Y is significantly larger |Y| ≫ n.This result helps to theoretically explain the strong performance observed in several prior empirical studies on MBR decoding.In addition, we provide the performance gap for maximum-a-posteriori (MAP) decoding and compare it to MBR decoding.The result of this paper indicates that MBR decoding tends to converge to the optimal solution faster than MAP decoding in several cases. Yuki Ichihara, Yuu Jinnai, Kaito Ariu, Tetsuro Morimura, Eiji Uchibe |
ACL (1) | 5 |
| 2025 | Human-in-the-Loop Generative Policy Learning from Demonstrations and Preferences
Eiji Uchibe |
ICONIP (5) | 1 |
| 2024 | Reward-Punishment Reinforcement Learning with Maximum EntropyabstractWe introduce the "soft Deep MaxPain" (softDMP) algorithm, which integrates the optimization of long-term policy entropy into reward-punishment reinforcement learning objectives. Our motivation is to facilitate a smoother variation of operators utilized in the updating of action values beyond traditional "max" and "min" operators, where the goal is enhancing sample efficiency and robustness. We also address two unresolved issues from the previous Deep MaxPain method. Firstly, we investigate how the negated ("flipped") pain-seeking sub-policy, derived from the punishment action value, collaborates with the "min" operator to effectively learn the punishment module and how softDMP’s smooth learning operator provides insights into the "flipping" trick. Secondly, we tackle the challenge of data collection for learning the punishment module to mitigate inconsistencies arising from the involvement of the "flipped" sub-policy (pain-avoidance sub-policy) in the unified behavior policy. We empirically explore the first issue in two discrete Markov Decision Process (MDP) environments, elucidating the crucial advancements of the DMP approach and the necessity for soft treatments on the hard operators. For the second issue, we propose a probabilistic classifier based on the ratio of the pain-seeking sub-policy to the sum of the pain-seeking and goalreaching sub-policies. This classifier assigns roll-outs to separate replay buffers for updating reward and punishment action-value functions respectively. Our framework demonstrates superior performance in Turtlebot 3’s maze navigation tasks under the ROS Gazebo simulation. Jiexin Wang 0001, Eiji Uchibe |
IJCNN | 2 |
| 2024 | Estimating cost function of expert players in differential games: A model-based method and its data-driven extensionabstractIn this paper, we introduce two algorithms for estimating the cost function of expert players engaged in optimal performance within linear continuous-time differential games. Initially, we propose a model-based algorithm, followed by its data-driven model-free extension. Both methods rely on optimal policy gains, obtained by observing Nash equilibrium trajectories of the expert players. The model-free method also utilizes the trajectories of the learner system. This method addresses the limitations found in existing model-free approaches, which may suffer from either high computational costs, limited applicability to specific systems, or both. The effectiveness of the proposed methods is demonstrated through numerical simulations. Hamed Jabbari, Eiji Uchibe |
Expert Syst. Appl. | 2 |
| 2024 | Online estimation of objective function for continuous-time deterministic systemsabstractWe developed two online data-driven methods for estimating an objective function in continuous-time linear and nonlinear deterministic systems. The primary focus addressed the challenge posed by unknown input dynamics (control mapping function) in the expert system, a critical element for an online solution of the problem. Our methods leverage both the learner's and expert's data for effective problem-solving. The first approach, which is model-free, estimates the expert's policy and integrates it into the learner agent to approximate the objective function associated with the optimal policy. The second approach estimates the input dynamics from the learner's data and combines it with the expert's input-state observations to tackle the objective function estimation problem. Compared to other methods for deterministic systems that rely on both the learner's and expert's data, our approaches offer reduced complexity by eliminating the need to estimate an optimal policy after each objective function update. We conduct a convergence analysis of the estimation techniques using Lyapunov-based methods. Numerical experiments validate the effectiveness of our developed methods. Hamed Jabbari, Eiji Uchibe |
Neural Networks | 2 |
| 2023 | Online Reinforcement Learning Control of Nonlinear Dynamic Systems: A State-action Value Function Based Solution
Hamed Jabbari, Eiji Uchibe |
Neurocomputing | 2 |
| 2022 | Deep learning, reinforcement learning, and world modelsabstractDeep learning (DL) and reinforcement learning (RL) methods seem to be a part of indispensable factors to achieve human-level or super-human AI systems. On the other hand, both DL and RL have strong connections with our brain functions and with neuroscientific findings. In this review, we summarize talks and discussions in the "Deep Learning and Reinforcement Learning" session of the symposium, International Symposium on Artificial Intelligence and Brain Science. In this session, we discussed whether we can achieve comprehensive understanding of human intelligence based on the recent advances of deep learning and reinforcement learning algorithms. Speakers contributed to provide talks about their recent studies that can be key technologies to achieve human-level intelligence. Yutaka Matsuo, Yann LeCun, Maneesh Sahani, Doina Precup, David Silver 0001, Masashi Sugiyama, Eiji Uchibe, Jun Morimoto |
Neural Networks | 7 |
| 2021 | Parallel and hierarchical neural mechanisms for adaptive and predictive behavioral controlabstractOur brain can be recognized as a network of largely hierarchically organized neural circuits that operate to control specific functions, but when acting in parallel, enable the performance of complex and simultaneous behaviors. Indeed, many of our daily actions require concurrent information processing in sensorimotor, associative, and limbic circuits that are dynamically and hierarchically modulated by sensory information and previous learning. This organization of information processing in biological organisms has served as a major inspiration for artificial intelligence and has helped to create in silico systems capable of matching or even outperforming humans in several specific tasks, including visual recognition and strategy-based games. However, the development of human-like robots that are able to move as quickly as humans and respond flexibly in various situations remains a major challenge and indicates an area where further use of parallel and hierarchical architectures may hold promise. In this article we review several important neural and behavioral mechanisms organizing hierarchical and predictive processing for the acquisition and realization of flexible behavioral control. Then, inspired by the organizational features of brain circuits, we introduce a multi-timescale parallel and hierarchical learning framework for the realization of versatile and agile movement in humanoid robots. Tom Macpherson, Masayuki Matsumoto, Hiroaki Gomi, Jun Morimoto, Eiji Uchibe, Takatoshi Hikida |
Neural Networks | 5 |
| 2021 | Forward and inverse reinforcement learning sharing network weights and hyperparametersabstractThis paper proposes model-free imitation learning named Entropy-Regularized Imitation Learning (ERIL) that minimizes the reverse Kullback-Leibler (KL) divergence. ERIL combines forward and inverse reinforcement learning (RL) under the framework of an entropy-regularized Markov decision process. An inverse RL step computes the log-ratio between two distributions by evaluating two binary discriminators. The first discriminator distinguishes the state generated by the forward RL step from the expert's state. The second discriminator, which is structured by the theory of entropy regularization, distinguishes the state-action-next-state tuples generated by the learner from the expert ones. One notable feature is that the second discriminator shares hyperparameters with the forward RL, which can be used to control the discriminator's ability. A forward RL step minimizes the reverse KL estimated by the inverse RL step. We show that minimizing the reverse KL divergence is equivalent to finding an optimal policy. Our experimental results on MuJoCo-simulated environments and vision-based reaching tasks with a robotic arm show that ERIL is more sample-efficient than the baseline methods. We apply the method to human behaviors that perform a pole-balancing task and describe how the estimated reward functions show how every subject achieves her goal. Eiji Uchibe, Kenji Doya |
Neural Networks | 1 |
| 2021 | Modular deep reinforcement learning from reward and punishment for robot navigationabstractModular Reinforcement Learning decomposes a monolithic task into several tasks with sub-goals and learns each one in parallel to solve the original problem. Such learning patterns can be traced in the brains of animals. Recent evidence in neuroscience shows that animals utilize separate systems for processing rewards and punishments, illuminating a different perspective for modularizing Reinforcement Learning tasks. MaxPain and its deep variant, Deep MaxPain, showed the advances of such dichotomy-based decomposing architecture over conventional Q-learning in terms of safety and learning efficiency. These two methods differ in policy derivation. MaxPain linearly unified the reward and punishment value functions and generated a joint policy based on unified values; Deep MaxPain tackled scaling problems in high-dimensional cases by linearly forming a joint policy from two sub-policies obtained from their value functions. However, the mixing weights in both methods were determined manually, causing inadequate use of the learned modules. In this work, we discuss the signal scaling of reward and punishment related to discounting factor γ, and propose a weak constraint for signaling design. To further exploit the learning models, we propose a state-value dependent weighting scheme that automatically tunes the mixing weights: hard-max and softmax based on a case analysis of Boltzmann distribution. We focus on maze-solving navigation tasks and investigate how two metrics (pain-avoiding and goal-reaching) influence each other's behaviors during learning. We propose a sensor fusion network structure that utilizes lidar and images captured by a monocular camera instead of lidar-only and image-only sensing. Our results, both in the simulation of three types of mazes with different complexities and a real robot experiment of an L-maze on Turtlebot3 Waffle Pi, showed the improvements of our methods. Jiexin Wang 0001, Stefan Elfwing, Eiji Uchibe |
Neural Networks | 3 |
| 2019 | Theoretical Analysis of Efficiency and Robustness of Softmax and Gap-Increasing Operators in Reinforcement LearningabstractIn this paper, we propose and analyze conservative value iteration, which unifies value iteration, soft value iteration, advantage learning, and dynamic policy programming. Our analysis shows that algorithms using a combination of gap-increasing and max operators are resilient to stochastic errors, but not to non-stochastic errors. In contrast, algorithms using a softmax operator without a gap-increasing operator are less susceptible to all types of errors, but may display poor asymptotic performance. Algorithms using a combination of gap-increasing and softmax operators are much more effective and may asymptotically outperform algorithms with the max operator. Not only do these theoretical results provide a deep understanding of various reinforcement learning algorithms, but they also highlight the effectiveness of gap-increasing operators, as well as the limitations of traditional greedy value updates by the max operator. Tadashi Kozuno, Eiji Uchibe, Kenji Doya |
AISTATS | 2 |
| 2018 | Online meta-learning by parallel algorithm competitionabstractThe efficiency of reinforcement learning algorithms depends critically on a few meta-parameters that modulate the learning updates and the trade-off between exploration and exploitation. The adaptation of the meta-parameters is an open question, which arguably has become a more important issue recently with the success of deep reinforcement learning. The long learning times in domains such as Atari 2600 video games makes it not feasible to perform comprehensive searches of appropriate meta-parameter values. In this study, we propose the Online Meta-learning by Parallel Algorithm Competition (OMPAC) method, which is a novel Lamarckian evolutionary approach to online meta-parameter adaptation. The population consists of several instances of a reinforcement learning algorithm which are run in parallel with small differences in initial meta-parameter values. After a fixed number of learning episodes, the instances are selected based on their performance on the task at hand, i.e., the fitness. Before continuing the learning, Gaussian noise is added to the meta-parameters with a predefined probability. We validate the OMPAC method by improving the state-of-the-art results in stochastic SZ-Tetris and in 10x10 Tetris by 31% and 84%, respectively, and by improving the learning speed and performance for deep Sarsa(λ) agents in the Atari 2600 domain. Stefan Elfwing, Eiji Uchibe, Kenji Doya |
GECCO | 2 |
| 2018 | Efficient sample reuse in policy search by multiple importance samplingabstractPolicy search such as reinforcement learning and evolutionary computation is a framework for finding an optimal policy of control problems, but it usually requires a huge number of samples. Importance sampling is a common tool to use samples drawn from a proposal distribution different from the targeted one, and it is widely used by the policy search methods to update the policy from a set of datasets that are collected by previous sampling distributions. However, the proposal distribution is created by a mixture of previous distributions with fixed mixing weights in most of previous studies, and it is often numerically unstable. To overcome this problem, we propose the method of adaptive multiple importance sampling that optimizes the mixing coefficients to minimize the variance of the importance sampling estimator while utilizing as many samples as possible. We apply the proposed method to the five policy search methods such as PGPE, PoWER, CMA-ES, REPS, and NES, and their algorithms are evaluated by some benchmark control tasks. Experimental results show that all the five methods improve sample efficiency. In addition, we show that optimizing the mixing weights achieves stable learning. Eiji Uchibe |
GECCO | 1 |
| 2018 | Sigmoid-weighted linear units for neural network function approximation in reinforcement learningabstractIn recent years, neural networks have enjoyed a renaissance as function approximators in reinforcement learning. Two decades after Tesauro's TD-Gammon achieved near top-level human performance in backgammon, the deep reinforcement learning algorithm DQN achieved human-level performance in many Atari 2600 games. The purpose of this study is twofold. First, we propose two activation functions for neural network function approximation in reinforcement learning: the sigmoid-weighted linear unit (SiLU) and its derivative function (dSiLU). The activation of the SiLU is computed by the sigmoid function multiplied by its input. Second, we suggest that the more traditional approach of using on-policy learning with eligibility traces, instead of experience replay, and softmax action selection can be competitive with DQN, without the need for a separate target network. We validate our proposed approach by, first, achieving new state-of-the-art results in both stochastic SZ-Tetris and Tetris with a small 10 × 10 board, using TD(λ) learning and shallow dSiLU network agents, and, then, by outperforming DQN in the Atari 2600 domain by using a deep Sarsa(λ) agent with SiLU and dSiLU hidden units. Stefan Elfwing, Eiji Uchibe, Kenji Doya |
Neural Networks | 2 |
| 2018 | Model-Free Deep Inverse Reinforcement Learning by Logistic RegressionabstractThis paper proposes model-free deep inverse reinforcement learning to find nonlinear reward function structures. We formulate inverse reinforcement learning as a problem of density ratio estimation, and show that the log of the ratio between an optimal state transition and a baseline one is given by a part of reward and the difference of the value functions under the framework of linearly solvable Markov decision processes. The logarithm of density ratio is efficiently calculated by binomial logistic regression, of which the classifier is constructed by the reward and state value function. The classifier tries to discriminate between samples drawn from the optimal state transition probability and those from the baseline one. Then, the estimated state value function is used to initialize the part of the deep neural networks for forward reinforcement learning. The proposed deep forward and inverse reinforcement learning is applied into two benchmark games: Atari 2600 and Reversi. Simulation results show that our method reaches the best performance substantially faster than the standard combination of forward and inverse reinforcement learning as well as behavior cloning. Eiji Uchibe |
Neural Process. Lett. | 1 |
| 2017 | Average Reward Optimization with Multiple Discounting Reinforcement Learners
Chris Reinke, Eiji Uchibe, Kenji Doya |
ICONIP (1) | 2 |
| 2017 | Deep dynamic policy programming for robot control with raw imagesabstractDeep reinforcement learning has drawn much attention in robot control since it enables agents to learn control policies from very high dimensional states such as raw images. On the other hand, its dependency upon the availability of a significant quantity of training samples and its fragility in learning makes it difficult to apply for real world robot tasks. To alleviate these issues we propose Deep Dynamic Policy Programming (DDPP), which combines the sample efficiency and smooth policy updates of dynamic policy programming with the contemporary deep reinforcement learning framework. The effectiveness of the proposed method is first demonstrated in a simulation of the robot arm control problem, with comparison to Deep Q-Networks. As validation on a real robot system, DDPP also successfully learned the flipping of a handkerchief with a NEXTAGE humanoid robot using a reduced number of learning samples, whereas Deep Q-Networks failed to learn the task. Yoshihisa Tsurumine, Yunduan Cui, Eiji Uchibe, Takamitsu Matsubara |
IROS | 3 |
| 2016 | Deep Inverse Reinforcement Learning by Logistic Regression
Eiji Uchibe |
ICONIP (1) | 1 |
| 2016 | From free energy to expected energy: Improving energy-based value function approximation in reinforcement learningabstractFree-energy based reinforcement learning (FERL) was proposed for learning in high-dimensional state and action spaces. However, the FERL method does only really work well with binary, or close to binary, state input, where the number of active states is fewer than the number of non-active states. In the FERL method, the value function is approximated by the negative free energy of a restricted Boltzmann machine (RBM). In our earlier study, we demonstrated that the performance and the robustness of the FERL method can be improved by scaling the free energy by a constant that is related to the size of network. In this study, we propose that RBM function approximation can be further improved by approximating the value function by the negative expected energy (EERL), instead of the negative free energy, as well as being able to handle continuous state input. We validate our proposed method by demonstrating that EERL: (1) outperforms FERL, as well as standard neural network and linear function approximation, for three versions of a gridworld task with high-dimensional image state input; (2) achieves new state-of-the-art results in stochastic SZ-Tetris in both model-free and model-based learning settings; and (3) significantly outperforms FERL and standard neural network function approximation for a robot navigation task with raw and noisy RGB images as state input and a large number of actions. Stefan Elfwing, Eiji Uchibe, Kenji Doya |
Neural Networks | 2 |
| 2015 | Expected energy-based restricted Boltzmann machine for classificationabstractIn classification tasks, restricted Boltzmann machines (RBMs) have predominantly been used in the first stage, either as feature extractors or to provide initialization of neural networks. In this study, we propose a discriminative learning approach to provide a self-contained RBM method for classification, inspired by free-energy based function approximation (FE-RBM), originally proposed for reinforcement learning. For classification, the FE-RBM method computes the output for an input vector and a class vector by the negative free energy of an RBM. Learning is achieved by stochastic gradient-descent using a mean-squared error training objective. In an earlier study, we demonstrated that the performance and the robustness of FE-RBM function approximation can be improved by scaling the free energy by a constant that is related to the size of network. In this study, we propose that the learning performance of RBM function approximation can be further improved by computing the output by the negative expected energy (EE-RBM), instead of the negative free energy. To create a deep learning architecture, we stack several RBMs on top of each other. We also connect the class nodes to all hidden layers to try to improve the performance even further. We validate the classification performance of EE-RBM using the MNIST data set and the NORB data set, achieving competitive performance compared with other classifiers such as standard neural networks, deep belief networks, classification RBMs, and support vector machines. The purpose of using the NORB data set is to demonstrate that EE-RBM with binary input nodes can achieve high performance in the continuous input domain. Stefan Elfwing, Eiji Uchibe, Kenji Doya |
Neural Networks | 2 |
| 2014 | Combining learned controllers to achieve new goals based on linearly solvable MDPsabstractLearning complicated behaviors usually involves intensive manual tuning and expensive computational optimization because we have to solve a nonlinear Hamilton-Jacobi-Bellman (HJB) equation. Recently, Todorov proposed a class of the so-called Linearly solvable Markov Decision Process (LMDP) which converts a nonlinear HJB equation to a linear differential equation. Linearity of the simplified HJB equation allows us to apply superposition to derive a new composite controller from a set of learned primitive controllers. However, his method was a model-based approach and it was not evaluated in a real domain. This study proposes a model-free method which is similar to the Least Squares Temporal Difference (LSTD) learning. In this method, the exponentially transformed cost function can be regarded as the discount factor in LSTD. Our proposed method is applied to learning walking behaviors with the quadruped robot to evaluate in real robot experiments. The goal of each primitive task is to go to the specific target position in the environment and that of the composite task is to approach arbitrary region represented by the primitives' target positions. Experimental results show that the composite policy can be used as a good initial policy for the new task. Eiji Uchibe, Kenji Doya |
ICRA | 1 |
| 2010 | Free-Energy Based Reinforcement Learning for Vision-Based Navigation with High-Dimensional Sensory Inputs
Stefan Elfwing, Makoto Otsuka, Eiji Uchibe, Kenji Doya |
ICONIP (1) | 3 |
| 2010 | Derivatives of Logarithmic Stationary Distributions for Policy Gradient Reinforcement LearningabstractMost conventional policy gradient reinforcement learning (PGRL) algorithms neglect (or do not explicitly make use of) a term in the average reward gradient with respect to the policy parameter. That term involves the derivative of the stationary state distribution that corresponds to the sensitivity of its distribution to changes in the policy parameter. Although the bias introduced by this omission can be reduced by setting the forgetting rate gamma for the value functions close to 1, these algorithms do not permit gamma to be set exactly at gamma = 1. In this article, we propose a method for estimating the log stationary state distribution derivative (LSD) as a useful form of the derivative of the stationary state distribution through backward Markov chain formulation and a temporal difference learning framework. A new policy gradient (PG) framework with an LSD is also proposed, in which the average reward gradient can be estimated by setting gamma = 0, so it becomes unnecessary to learn the value functions. We also test the performance of the proposed algorithms using simple benchmark tasks and show that these can improve the performances of existing PG methods. Tetsuro Morimura, Eiji Uchibe, Junichiro Yoshimoto, Jan Peters 0001, Kenji Doya |
Neural Comput. | 2 |
| 2009 | Emergence of Different Mating Strategies in Artificial Embodied Evolution
Stefan Elfwing, Eiji Uchibe, Kenji Doya |
ICONIP (2) | 2 |
| 2009 | A Generalized Natural Actor-Critic AlgorithmabstractPolicy gradient Reinforcement Learning (RL) algorithms have received much attention in seeking stochastic policies that maximize the average rewards. In addition, extensions based on the concept of the Natural Gradient (NG) show promising learning efficiency because these regard metrics for the task. Though there are two candidate metrics, Kakades Fisher Information Matrix (FIM) and Morimuras FIM, all RL algorithms with NG have followed the Kakades approach. In this paper, we describe a generalized Natural Gradient (gNG) by linearly interpolating the two FIMs and propose an efficient implementation for the gNG learning based on a theory of the estimating function, generalized Natural Actor-Critic (gNAC). The gNAC algorithm involves a near optimal auxiliary function to reduce the variance of the gNG estimates. Interestingly, the gNAC can be regarded as a natural extension of the current state-of-the-art NAC algorithm, as long as the interpolating parameter is appropriately selected. Numerical experiments showed that the proposed gNAC algorithm can estimate gNG efficiently and outperformed the NAC algorithm. Tetsuro Morimura, Eiji Uchibe, Junichiro Yoshimoto, Kenji Doya |
NIPS | 2 |
| 2008 | NeuroEvolution Based on Reusable and Hierarchical Modular Representation
Takumi Kamioka, Eiji Uchibe, Kenji Doya |
ICONIP (1) | 2 |
| 2008 | A New Natural Policy Gradient by Stationary Distribution Metric
Tetsuro Morimura, Eiji Uchibe, Junichiro Yoshimoto, Kenji Doya |
ECML/PKDD (2) | 2 |
| 2008 | Finding intrinsic rewards by embodied evolution and constrained reinforcement learning
Eiji Uchibe, Kenji Doya |
Neural Networks | 1 |
| 2007 | Finding Exploratory Rewards by Embodied Evolution and Constrained Reinforcement Learning in the Cyber Rodents
Eiji Uchibe, Kenji Doya |
ICONIP (2) | 1 |
| 2007 | Evolutionary Development of Hierarchical Learning StructuresabstractHierarchical reinforcement learning (RL) algorithms can learn a policy faster than standard RL algorithms. However, the applicability of hierarchical RL algorithms is limited by the fact that the task decomposition has to be performed in advance by the human designer. We propose a Lamarckian evolutionary approach for automatic development of the learning structure in hierarchical RL. The proposed method combines the MAXQ hierarchical RL method and genetic programming (GP). In the MAXQ framework, a subtask can optimize the policy independently of its parent task's policy, which makes it possible to reuse learned policies of the subtasks. In the proposed method, the MAXQ method learns the policy based on the task hierarchies obtained by GP, while the GP explores the appropriate hierarchies using the result of the MAXQ method. To show the validity of the proposed method, we have performed simulation experiments for a foraging task in three different environmental settings. The results show strong interconnection between the obtained learning structures and the given task environments. The main conclusion of the experiments is that the GP can find a minimal strategy, i.e., a hierarchy that minimizes the number of primitive subtasks that can be executed for each type of situation. The experimental results for the most challenging environment also show that the policies of the subtasks can continue to improve, even after the structure of the hierarchy has been evolutionary stabilized, as an effect of Lamarckian mechanisms Stefan Elfwing, Eiji Uchibe, Kenji Doya, Henrik I. Christensen |
IEEE Trans. Evol. Comput. | 2 |
| 2006 | Incremental Coevolution With Competitive and Cooperative Tasks in a Multirobot EnvironmentabstractCoevolution has been receiving increased attention as a method for simultaneously developing the control structures of multiple agents. Our ultimate goal is the mutual development of skills through coevolution. The coevolutionary process is, however, often prone to settle into suboptimal strategies. The key to successful coevolution has thus far been unclear. This paper discusses how several robots can emerge cooperative and competitive behavior through coevolutionary processes. In order to realize successful coevolution, we propose two ideas: multiple schedules for incremental evolution and fitness sharing based on the method of importance sampling. To examine this issue, we conducted a series of computer simulations. We have chosen a simplified soccer game consisting of two or three robots as a testbed for analyzing a problem in which both competitive and cooperative tasks are involved. We show that the proposed fitness evaluation allows robots to evolve robust behaviors in cooperative and competitive situations Eiji Uchibe, Minoru Asada |
Proc. IEEE | 1 |
| 2005 | Biologically inspired embodied evolution of survivalabstractEmbodied evolution is a methodology for evolutionary robotics that mimics the distributed, asynchronous and autonomous properties of biological evolution. The evaluation, selection and reproduction are carried out by and between the robots, without any need for human intervention. In this paper, we propose a biologically inspired embodied evolution framework, which fully integrates self-preservation, recharging from external batteries in the environment, and self-reproduction, pair-wise exchange of genetic material, into a survival system. The individuals are explicitly evaluated for the performance of the battery capturing task, but also implicitly for the mating task by the fact that an individual that mates frequently has larger probability to spread its gene in the population. We have evaluated our method in simulation experiments and the simulation results show that the solutions obtained by our embodied evolution method were able to optimize the two survival tasks, battery capturing and mating, simultaneously. We have also performed preliminary experiments in hardware, with promising results. Stefan Elfwing, Eiji Uchibe, Kenji Doya, Henrik I. Christensen |
Congress on Evolutionary Computation | 2 |
| 2004 | Multi-agent reinforcement learning: using macro actions to learn a mating taskabstractStandard reinforcement learning methods are inefficient and often inadequate for learning cooperative multi-agent tasks. For these kinds of tasks the behavior of one agent strongly depends on dynamic interaction with other agents, not only with the interaction with a static environment as in standard reinforcement learning. The success of the learning is therefore coupled to the agents' ability to predict the other agents behaviors. In this study we try to overcome this problem by adding a few simple macro actions, actions that are extended in time for more than one time step. The macro actions improve the learning by making search of the state space more effective and thereby making the behavior more predictable for the other agent. In this study we have considered a cooperative mating task, which is the first step towards our aim to perform embodied evolution, where the evolutionary selection process is an integrated part of the task. We show, in simulation and hardware, that in the case of learning without macro actions, the agents fail to learn a meaningful behavior. In contrast, for the learning with macro action the agents learn a good mating behavior in reasonable time, in both simulation and hardware. Stefan Elfwing, Eiji Uchibe, Kenji Doya, Henrik I. Christensen |
IROS | 2 |
| 2003 | An Evolutionary Approach to Automatic Construction of the Structure in Hierarchical Reinforcement Learning
Stefan Elfwing, Eiji Uchibe, Kenji Doya |
GECCO | 2 |
| 2001 | Dynamic Task Assignment in a Multiagent/Multitask Environment based on Module Conflict ResolutionabstractIt is necessary to coordinate multiple tasks in order to cope with larger-scaled and more complicated tasks. However, it seems very hard to accomplish the multiple tasks at the same time. The paper proposes a method to resolve a conflict between task modules through the processes of their executions. Based on the proposed method, the robot can select an appropriate module according to the priority. In addition, we apply the module conflict resolution to a multiagent environment. Consequently, multiple tasks are automatically allocated to the multiple robots. As a task example, a soccer game is selected to show the validity of the proposed method. Real experiments are shown, and a discussion is given. Eiji Uchibe, Tatsunori Kato, Koh Hosoda, Minoru Asada |
ICRA | 1 |
| 2001 | Evolutionary Behavior Selection with Activation/Termination Constraints
Eiji Uchibe, Masakazu Yanase, Minoru Asada |
RoboCup | 1 |
| 2000 | Osaka University "Trackies 2000"
Yasutake Takahashi, Eiji Uchibe, Takahashi Tamura, Masakazu Yanase, Shoichi Ikenoue, Shujiro Inui, Minoru Asada |
RoboCup | 2 |
| 1999 | The Team Description of Osaka University "Trackies-99"
Sho'ji Suzuki, Tatsunori Kato, Hiroshi Ishizuka, Hiroyoshi Kawanishi, Takashi Tamura, Masakazu Yanase, Yasutake Takahashi, Eiji Uchibe, Minoru Asada |
RoboCup | 8 |
| 1999 | Multiple Reward Criterion for Cooperative Behavior Acquisition in a Muliagent Environment
Eiji Uchibe, Minoru Asada |
RoboCup | 1 |
| 1999 | Cooperative Behavior Acquisition for Mobile Robots in Dynamically Changing Real Worlds Via Vision-Based Reinforcement Learning and Development
Minoru Asada, Eiji Uchibe, Koh Hosoda |
Artif. Intell. | 2 |
| 1998 | State Space Construction for Behavior Acquisition in Multi Agent Environments with Vision and ActionabstractThis paper proposes a method which estimates the relationships between learner's behaviors and other agents' ones in the environment through interactions (observation and action) using the method of system identification. In order to identify the model of each agent, Akaike's Information Criterion is applied to the results of Canonical Variate Analysis for the relationship between the observed data in terms of action and future observation. Next, reinforcement learning based on the estimated state vectors is performed to obtain the optimal behavior. The proposed method is applied to a soccer playing situation, where a rolling ball and other moving agents are well modeled and the learner's behaviors are successfully acquired by the method. Computer simulations and real experiments are shown and a discussion is given. Eiji Uchibe, Minoru Asada, Koh Hosoda |
ICCV | 1 |
| 1998 | Cooperative Behavior Acquisition in Multi Mobile Robots Environment by Reinforcement Learning Based on State Vector EstimationabstractThis paper proposes a method that acquires robots' behaviors based on the estimation of the state vectors. In order to acquire the cooperative behaviors in multi-robot environments, each learning robot estimates the local predictive model between the learner and the other objects separately. Based on the local predictive models, the robots learn the desired behaviors using reinforcement learning. The proposed method is applied to a soccer playing situation, where a rolling ball and other moving robots are well modeled and the learner's behaviors are successfully acquired by the method. Computer simulations and real experiments are shown and a discussion is given. Eiji Uchibe, Minoru Asada, Koh Hosoda |
ICRA | 1 |
| 1998 | Environmental Complexity Control for Vision-Based Learning Mobile RobotabstractDiscusses how a robot can develop its state vector according to the complexity of the interactions with its environment. A method for controlling the complexity is proposed for a vision-based mobile robot whose task is to shoot a ball into a goal avoiding collisions with a goalkeeper. First, we provide the most difficult situation (the maximum speed of the goalkeeper with chasing-a-ball behavior), and the robot estimates the full set of state vectors with the order of the major vector components by a method of system identification. The environmental complexity is defined in terms of the speed of the goalkeeper while the complexity of the state vector is the number of the dimensions of the state vector. According to the increase of the speed of the goalkeeper, the dimension of the state vector is increased by taking a trade-off between the size of the state space (the dimension) and the learning time. Simulations are shown, and other issues for the complexity control are discussed. Eiji Uchibe, Minoru Asada, Koh Hosoda |
ICRA | 1 |
| 1998 | Co-evolution for cooperative behavior acquisition in a multiple mobile robot environmentabstractCo-evolution has been receiving increased attention as a method for multi agent simultaneous learning. This paper discusses how multiple robots can emerge cooperative behaviors through co-evolutionary processes. As an example task, a simplified soccer game with three learning robots is selected and a genetic programming method is applied to individual population corresponding to each robot so as to obtain cooperative and competitive behaviors. The complexity of the problem can be explained twofold: co-evolution for cooperative behaviors needs exact synchronization of mutual evolutions, and three robot co-evolution requires well-complicated environment setups that may gradually change from, simpler to more complicated situations. Simulation results are shown, and a discussion is given. Eiji Uchibe, Masateru Nakamura, Minoru Asada |
IROS | 1 |
| 1998 | An Application of Vision-Based Learning in RoboCup for a Real Robot with an Omnidirectional Vision System and the Team Description of Osaka University "Trackies"
Sho'ji Suzuki, Tatsunori Kato, Hiroshi Ishizuka, Yasutake Takahashi, Eiji Uchibe, Minoru Asada |
RoboCup | 5 |
| 1998 | Cooperative Behavior Acquisition in a Multiple Mobile Robot Environment by Co-evolution
Eiji Uchibe, Masateru Nakamura, Minoru Asada |
RoboCup | 1 |
| 1997 | Vision-Based Robot Learning Towards RoboCup: Osaka University "Trackies"
Sho'ji Suzuki, Yasutake Takahashi, Eiji Uchibe, Masateru Nakamura, Chizuko Mishima, Hiroshi Ishizuka, Tatsunori Kato, Minoru Asada |
RoboCup | 3 |
| 1996 | Behavior coordination for a mobile robot using modular reinforcement learningabstractCoordination of multiple behaviors independently obtained by a reinforcement learning method is one of the issues in order for the method to be scaled to larger and more complex robot learning tasks. Direct combination of all the state spaces for individual modules (subtasks) needs enormous learning time, and it causes hidden states. This paper presents a method of modular learning which coordinates multiple behaviors taking account of a trade-off between learning time and performance. First, in order to reduce the learning time the whole state space is classified into two categories based on the action values separately obtained by Q learning: the area where one of the learned behaviors is directly applicable (no more learning area), and the area where learning is necessary due to competition of multiple behaviors (re-learning area). Second, hidden states are detected by model fitting to the learned action values based on the information criterion. Finally, the initial action valves in the re-learning area are adjusted so that they can be consistent with the values in the no more learning area. The method is applied to one to one soccer playing robots. Computer simulation and real robot experiments are given, to show the validity of the proposed method. Eiji Uchibe, Minoru Asada, Koh Hosoda |
IROS | 1 |
| 1994 | Coordination of multiple behaviors acquired by a vision-based reinforcement learningabstractA method is proposed which accomplishes a whole task consisting of plural subtasks by coordinating multiple behaviors acquired by a vision-based reinforcement learning. First, individual behaviors which achieve the corresponding subtasks are independently acquired by Q-learning, a widely used reinforcement learning method. Each learned behavior can be represented by an action-value function in terms of state of the environment and robot action. Next, three kinds of coordinations of multiple behaviors are considered; simple summation of different action-value functions, switching action-value functions according to situations, and learning with previously obtained action-value functions as initial values of a new action-value function. A task of shooting a ball into the goal avoiding collisions with an enemy is examined. The task can be decomposed into a ball shooting subtask and a collision avoiding subtask. These subtasks should be accomplished simultaneously, but they are not independent of each other.> Minoru Asada, Eiji Uchibe, Shoichi Noda, Sukoya Tawaratsumida, Koh Hosoda |
IROS | 2 |