Kenji Doya

dblp:00/100 · DBLP profile ↗
← Back
105ranked-venue papers
31as first author
14since 2021 · last 2026
0000-0002-2446-6820ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 95 · 31 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 1 first-author · 3 since 2021Systems, architecture and hardware · 5Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CEBRA-Enabled Latent Embeddings of Wearable Biosignals for Personalized Biorhythm Modeling
Sutashu Tomonaga, Jo Fujimori, Yi-Shan Cheng, Haruo Mizutani, Kenji Doya
PERSUASIVE5
2025 Training Recurrent Neural Networks with Inherent Missing Data for Wearable Device Applications (Student Abstract)
abstract
Wearable devices are transforming healthcare by providing continuous, real-time physiological data for monitoring and analysis. However, data often suffer from noise and significant missing values due to operational constraints and user compliance. Traditional approaches address these issues through data imputation during pre-processing, introducing biases and inaccuracies. We propose a novel method enabling Recurrent Neural Networks (RNNs) to inherently handle missing data without imputation. By implementing teacher-forcing during Backpropagation Through Time (BPTT) when data are available and switching to autonomous mode otherwise, our approach leverages RNNs' dynamics to model physiological signals accurately. We demonstrate our method's effectiveness using the Lorenz 63 system as a surrogate dataset, achieving robust reconstructions with 80% missing data.
Sutashu Tomonaga, Haruo Mizutani, Kenji Doya
AAAI3
2025 Optical Neuroimage Studio (OptiNiSt): Intuitive, scalable, extendable framework for optical neuroimage data analysis
abstract
Advancements in calcium indicators and optical techniques have made optical neural recording common in neuroscience. As data volumes grow, streamlining the analysis pipelines for image preprocessing, signal extraction, and subsequent neural activity analyses becomes essential. Challenges in analysis includes 1) ensuring data quality of original and processed data at each step, 2) selecting optimal algorithms and their parameters from numerous options, each with its own pros and cons, by implementing or installing them manually, 3) systematically recording each analysis step for reproducibility, and 4) adopting standard data formats for data sharing and meta-analyses. To address these challenges, we developed Optical Neuroimage Studio (OptiNiSt), a scalable, extendable, and reproducible framework for creating calcium data analysis pipelines. OptiNiSt includes the following features. 1) Researchers can easily create analysis pipelines by selecting multiple processing modules, tuning their parameters, and visualizing the results at each step through a graphic user interface in a web browser. 2) In addition to pre-installed tools, new analysis algorithms can be easily added. 3) Once a processing pipeline is designed, the entire workflow with its modules and parameters are stored in a YAML file, which makes the pipeline reproducible and deployable on high-performance computing clusters. 4) OptiNiSt can read image data in a variety of file formats and store the analysis results in NWB (Neurodata Without Borders), a standard data format for data sharing. We expect that this framework will be helpful in standardizing optical neural data analysis protocols.
Yukako Yamane, Keita Matsumoto, Ryota Kanai, Miles Desforges, Carlos Enrique Gutierrez, Kenji Doya
PLoS Comput. Biol.7
2024 Intrinsic Rewards for Exploration Without Harm From Observational Noise: A Simulation Study Based on the Free Energy Principle
abstract
In reinforcement learning (RL), artificial agents are trained to maximize numerical rewards by performing tasks. Exploration is essential in RL because agents must discover information before exploiting it. Two rewards encouraging efficient exploration are the entropy of action policy and curiosity for information gain. Entropy is well established in the literature, promoting randomized action selection. Curiosity is defined in a broad variety of ways in literature, promoting discovery of novel experiences. One example, prediction error curiosity, rewards agents for discovering observations they cannot accurately predict. However, such agents may be distracted by unpredictable observational noises known as curiosity traps. Based on the free energy principle (FEP), this letter proposes hidden state curiosity, which rewards agents by the KL divergence between the predictive prior and posterior probabilities of latent variables. We trained six types of agents to navigate mazes: baseline agents without rewards for entropy or curiosity and agents rewarded for entropy and/or either prediction error curiosity or hidden state curiosity. We find that entropy and curiosity result in efficient exploration, especially both employed together. Notably, agents with hidden state curiosity demonstrate resilience against curiosity traps, which hinder agents with prediction error curiosity. This suggests implementing the FEP that may enhance the robustness and generalization of RL models, potentially aligning the learning processes of artificial and biological agents.
Theodore Jerome Tinker, Kenji Doya, Jun Tani
Neural Comput.2
2023 Enhancing reinforcement learning models by including direct and indirect pathways improves performance on striatal dependent tasks
abstract
A major advance in understanding learning behavior stems from experiments showing that reward learning requires dopamine inputs to striatal neurons and arises from synaptic plasticity of cortico-striatal synapses. Numerous reinforcement learning models mimic this dopamine-dependent synaptic plasticity by using the reward prediction error, which resembles dopamine neuron firing, to learn the best action in response to a set of cues. Though these models can explain many facets of behavior, reproducing some types of goal-directed behavior, such as renewal and reversal, require additional model components. Here we present a reinforcement learning model, TD2Q, which better corresponds to the basal ganglia with two Q matrices, one representing direct pathway neurons (G) and another representing indirect pathway neurons (N). Unlike previous two-Q architectures, a novel and critical aspect of TD2Q is to update the G and N matrices utilizing the temporal difference reward prediction error. A best action is selected for N and G using a softmax with a reward-dependent adaptive exploration parameter, and then differences are resolved using a second selection step applied to the two action probabilities. The model is tested on a range of multi-step tasks including extinction, renewal, discrimination; switching reward probability learning; and sequence learning. Simulations show that TD2Q produces behaviors similar to rodents in choice and sequence learning tasks, and that use of the temporal difference reward prediction error is required to learn multi-step tasks. Blocking the update rule on the N matrix blocks discrimination learning, as observed experimentally. Performance in the sequence learning task is dramatically improved with two matrices. These results suggest that including additional aspects of basal ganglia physiology can improve the performance of reinforcement learning models, better reproduce animal behaviors, and provide insight as to the role of direct- and indirect-pathway striatal neurons.
Kim T. Blackwell, Kenji Doya
PLoS Comput. Biol.2
2022 Variational oracle guiding for reinforcement learning
Tadashi Kozuno, Xufang Luo, Zhao-Yun Chen, Kenji Doya, Yuqing Yang 0001, Dongsheng Li 0002
ICLR5
2022 Numerical Data Imputation: Choose kNN over Deep Learning
Florian Lalande, Kenji Doya
SISAP2
2022 Social impact and governance of AI and neurotechnologies
abstract
Advances in artificial intelligence (AI) and brain science are going to have a huge impact on society. While technologies based on those advances can provide enormous social benefits, adoption of new technologies poses various risks. This article first reviews the co-evolution of AI and brain science and the benefits of brain-inspired AI in sustainability, healthcare, and scientific discoveries. We then consider possible risks from those technologies, including intentional abuse, autonomous weapons, cognitive enhancement by brain-computer interfaces, insidious effects of social media, inequity, and enfeeblement. We also discuss practical ways to bring ethical principles into practice. One proposal is to stop giving explicit goals to AI agents and to enable them to keep learning human preferences. Another is to learn from democratic mechanisms that evolved in human society to avoid over-consolidation of power. Finally, we emphasize the importance of open discussions not only by experts, but also including a diverse array of lay opinions.
Kenji Doya, Arisa Ema, Hiroaki Kitano, Masamichi Sakagami, Stuart Russell 0001
Neural Networks1
2022 Neural Networks special issue on Artificial Intelligence and Brain Science
Kenji Doya, Karl J. Friston, Masashi Sugiyama, Josh Tenenbaum
Neural Networks1
2022 Continual growth and a transition
Kenji Doya, Taro Toyoizumi, DeLiang Wang
Neural Networks1
2022 Announcement of the Neural Networks Best Paper Award
Kenji Doya, DeLiang Wang
Neural Networks1
2022 A whole brain probabilistic generative model: Toward realizing cognitive architectures for developmental robots
abstract
Building a human-like integrative artificial cognitive system, that is, an artificial general intelligence (AGI), is the holy grail of the artificial intelligence (AI) field. Furthermore, a computational model that enables an artificial system to achieve cognitive development will be an excellent reference for brain and cognitive science. This paper describes an approach to develop a cognitive architecture by integrating elemental cognitive modules to enable the training of the modules as a whole. This approach is based on two ideas: (1) brain-inspired AI, learning human brain architecture to build human-level intelligence, and (2) a probabilistic generative model (PGM)-based cognitive architecture to develop a cognitive system for developmental robots by integrating PGMs. The proposed development framework is called a whole brain PGM (WB-PGM), which differs fundamentally from existing cognitive architectures in that it can learn continuously through a system based on sensory-motor information. In this paper, we describe the rationale for WB-PGM, the current status of PGM-based elemental cognitive modules, their relationship with the human brain, the approach to the integration of the cognitive modules, and future challenges. Our findings can serve as a reference for brain studies. As PGMs describe explicit informational relationships between variables, WB-PGM provides interpretable guidance from computational sciences to brain science. By providing such information, researchers in neuroscience can provide feedback to researchers in AI and robotics on what the current models lack with reference to the brain. Further, it can facilitate collaboration among researchers in neuro-cognitive sciences as well as AI and robotics.
Tadahiro Taniguchi, Hiroshi Yamakawa, Takayuki Nagai, Kenji Doya, Masamichi Sakagami, Tomoaki Nakamura, Akira Taniguchi
Neural Networks4
2021 Maintaining the Publication Infrastructure in a Worldwide Pandemic
Kenji Doya, DeLiang Wang
Neural Networks1
2021 Forward and inverse reinforcement learning sharing network weights and hyperparameters
abstract
This paper proposes model-free imitation learning named Entropy-Regularized Imitation Learning (ERIL) that minimizes the reverse Kullback-Leibler (KL) divergence. ERIL combines forward and inverse reinforcement learning (RL) under the framework of an entropy-regularized Markov decision process. An inverse RL step computes the log-ratio between two distributions by evaluating two binary discriminators. The first discriminator distinguishes the state generated by the forward RL step from the expert's state. The second discriminator, which is structured by the theory of entropy regularization, distinguishes the state-action-next-state tuples generated by the learner from the expert ones. One notable feature is that the second discriminator shares hyperparameters with the forward RL, which can be used to control the discriminator's ability. A forward RL step minimizes the reverse KL estimated by the inverse RL step. We show that minimizing the reverse KL divergence is equivalent to finding an optimal policy. Our experimental results on MuJoCo-simulated environments and vision-based reaching tasks with a robotic arm show that ERIL is more sample-efficient than the baseline methods. We apply the method to human behaviors that perform a pole-balancing task and describe how the estimated reward functions show how every subject achieves her goal.
Eiji Uchibe, Kenji Doya
Neural Networks2
2020 Variational Recurrent Models for Solving Partially Observable Control Tasks
Kenji Doya, Jun Tani
ICLR2
2020 Announcement of the Neural Networks Best Paper Award
Kenji Doya, DeLiang Wang
Neural Networks1
2020 Self-organization of action hierarchy and compositionality by reinforcement learning with recurrent neural networks
abstract
Recurrent neural networks (RNNs) for reinforcement learning (RL) have shown distinct advantages, e.g., solving memory-dependent tasks and meta-learning. However, little effort has been spent on improving RNN architectures and on understanding the underlying neural mechanisms for performance gain. In this paper, we propose a novel, multiple-timescale, stochastic RNN for RL. Empirical results show that the network can autonomously learn to abstract sub-goals and can self-develop an action hierarchy using internal dynamics in a challenging continuous control task. Furthermore, we show that the self-developed compositionality of the network enhances faster re-learning when adapting to a new task that is a re-composition of previously learned sub-goals, than when starting from scratch. We also found that improved performance can be achieved when neural activities are subject to stochastic rather than deterministic dynamics.
Kenji Doya, Jun Tani
Neural Networks2
2019 Theoretical Analysis of Efficiency and Robustness of Softmax and Gap-Increasing Operators in Reinforcement Learning
abstract
In this paper, we propose and analyze conservative value iteration, which unifies value iteration, soft value iteration, advantage learning, and dynamic policy programming. Our analysis shows that algorithms using a combination of gap-increasing and max operators are resilient to stochastic errors, but not to non-stochastic errors. In contrast, algorithms using a softmax operator without a gap-increasing operator are less susceptible to all types of errors, but may display poor asymptotic performance. Algorithms using a combination of gap-increasing and softmax operators are much more effective and may asymptotically outperform algorithms with the max operator. Not only do these theoretical results provide a deep understanding of various reinforcement learning algorithms, but they also highlight the effectiveness of gap-increasing operators, as well as the limitations of traditional greedy value updates by the max operator.
Tadashi Kozuno, Eiji Uchibe, Kenji Doya
AISTATS3
2018 Online meta-learning by parallel algorithm competition
abstract
The efficiency of reinforcement learning algorithms depends critically on a few meta-parameters that modulate the learning updates and the trade-off between exploration and exploitation. The adaptation of the meta-parameters is an open question, which arguably has become a more important issue recently with the success of deep reinforcement learning. The long learning times in domains such as Atari 2600 video games makes it not feasible to perform comprehensive searches of appropriate meta-parameter values. In this study, we propose the Online Meta-learning by Parallel Algorithm Competition (OMPAC) method, which is a novel Lamarckian evolutionary approach to online meta-parameter adaptation. The population consists of several instances of a reinforcement learning algorithm which are run in parallel with small differences in initial meta-parameter values. After a fixed number of learning episodes, the instances are selected based on their performance on the task at hand, i.e., the fitness. Before continuing the learning, Gaussian noise is added to the meta-parameters with a predefined probability. We validate the OMPAC method by improving the state-of-the-art results in stochastic SZ-Tetris and in 10x10 Tetris by 31% and 84%, respectively, and by improving the learning speed and performance for deep Sarsa(λ) agents in the Atari 2600 domain.
Stefan Elfwing, Eiji Uchibe, Kenji Doya
GECCO3
2018 PIPPS: Flexible Model-Based Policy Search Robust to the Curse of Chaos
abstract
Previously, the exploding gradient problem has been explained to be central in deep learning and model-based reinforcement learning, because it causes numerical issues and instability in optimization. Our experiments in model-based reinforcement learning imply that the problem is not just a numerical issue, but it may be caused by a fundamental chaos-like nature of long chains of nonlinear computations. Not only do the magnitudes of the gradients become large, the direction of the gradients becomes essentially random. We show that reparameterization gradients suffer from the problem, while likelihood ratio gradients are robust. Using our insights, we develop a model-based policy search framework, Probabilistic Inference for Particle-Based Policy Search (PIPPS), which is easily extensible, and allows for almost arbitrary models and policies, while simultaneously matching the performance of previous data-efficient learning algorithms. Finally, we invent the total propagation algorithm, which efficiently computes a union over all pathwise derivative depths during a single backwards pass, automatically giving greater weight to estimators with lower variance, sometimes improving over reparameterization gradients by $10^6$ times.
Paavo Parmas, Carl E. Rasmussen, Jan Peters 0001, Kenji Doya
ICML4
2018 Connectivity inference from neural recording data: Challenges, mathematical bases and research directions
abstract
This article presents a review of computational methods for connectivity inference from neural activity data derived from multi-electrode recordings or fluorescence imaging. We first identify biophysical and technical challenges in connectivity inference along the data processing pipeline. We then review connectivity inference methods based on two major mathematical foundations, namely, descriptive model-free approaches and generative model-based approaches. We investigate representative studies in both categories and clarify which challenges have been addressed by which method. We further identify critical open issues and possible research directions.
Ildefons Magrans de Abril, Junichiro Yoshimoto, Kenji Doya
Neural Networks3
2018 Fostering deep learning and beyond
Kenji Doya, DeLiang Wang
Neural Networks1
2018 Sigmoid-weighted linear units for neural network function approximation in reinforcement learning
abstract
In recent years, neural networks have enjoyed a renaissance as function approximators in reinforcement learning. Two decades after Tesauro's TD-Gammon achieved near top-level human performance in backgammon, the deep reinforcement learning algorithm DQN achieved human-level performance in many Atari 2600 games. The purpose of this study is twofold. First, we propose two activation functions for neural network function approximation in reinforcement learning: the sigmoid-weighted linear unit (SiLU) and its derivative function (dSiLU). The activation of the SiLU is computed by the sigmoid function multiplied by its input. Second, we suggest that the more traditional approach of using on-policy learning with eligibility traces, instead of experience replay, and softmax action selection can be competitive with DQN, without the need for a separate target network. We validate our proposed approach by, first, achieving new state-of-the-art results in both stochastic SZ-Tetris and Tetris with a small 10 × 10 board, using TD(λ) learning and shallow dSiLU network agents, and, then, by outperforming DQN in the Atari 2600 domain by using a deep Sarsa(λ) agent with SiLU and dSiLU hidden units.
Stefan Elfwing, Eiji Uchibe, Kenji Doya
Neural Networks3
2017 Average Reward Optimization with Multiple Discounting Reinforcement Learners
Chris Reinke, Eiji Uchibe, Kenji Doya
ICONIP (1)3
2017 Sparse kernel canonical correlation analysis for discovery of nonlinear interactions in high-dimensional data
abstract
BACKGROUND: Advance in high-throughput technologies in genomics, transcriptomics, and metabolomics has created demand for bioinformatics tools to integrate high-dimensional data from different sources. Canonical correlation analysis (CCA) is a statistical tool for finding linear associations between different types of information. Previous extensions of CCA used to capture nonlinear associations, such as kernel CCA, did not allow feature selection or capturing of multiple canonical components. Here we propose a novel method, two-stage kernel CCA (TSKCCA) to select appropriate kernels in the framework of multiple kernel learning. RESULTS: TSKCCA first selects relevant kernels based on the HSIC criterion in the multiple kernel learning framework. Weights are then derived by non-negative matrix decomposition with L1 regularization. Using artificial datasets and nutrigenomic datasets, we show that TSKCCA can extract multiple, nonlinear associations among high-dimensional data and multiplicative interactions among variables. CONCLUSIONS: TSKCCA can identify nonlinear associations among high-dimensional data more reliably than previous nonlinear CCA methods.
Kosuke Yoshida, Junichiro Yoshimoto, Kenji Doya
BMC Bioinform.3
2017 Promoting Further Developments of Neural Networks
Kenji Doya, DeLiang Wang
Neural Networks1
2017 Announcement of the Neural Networks Best Paper Award
Kenji Doya, DeLiang Wang
Neural Networks1
2016 State of Neural Networks Is Strong
Kenji Doya, DeLiang Wang
Neural Networks1
2016 From free energy to expected energy: Improving energy-based value function approximation in reinforcement learning
abstract
Free-energy based reinforcement learning (FERL) was proposed for learning in high-dimensional state and action spaces. However, the FERL method does only really work well with binary, or close to binary, state input, where the number of active states is fewer than the number of non-active states. In the FERL method, the value function is approximated by the negative free energy of a restricted Boltzmann machine (RBM). In our earlier study, we demonstrated that the performance and the robustness of the FERL method can be improved by scaling the free energy by a constant that is related to the size of network. In this study, we propose that RBM function approximation can be further improved by approximating the value function by the negative expected energy (EERL), instead of the negative free energy, as well as being able to handle continuous state input. We validate our proposed method by demonstrating that EERL: (1) outperforms FERL, as well as standard neural network and linear function approximation, for three versions of a gridworld task with high-dimensional image state input; (2) achieves new state-of-the-art results in stochastic SZ-Tetris in both model-free and model-based learning settings; and (3) significantly outperforms FERL and standard neural network function approximation for a robot navigation task with raw and noisy RGB images as state input and a large number of actions.
Stefan Elfwing, Eiji Uchibe, Kenji Doya
Neural Networks3
2015 Resting state functional connectivity explains individual scores of multiple clinical measures for major depression
abstract
Recent studies have revealed that resting state functional connectivity is associated with major depressive disorder (MDD). However, the relationship between functional connectivity and clinical measures for the detailed assessment of depression remains unclear. The objective of our study is thus to associate functional connectivity of depressed patients and healthy controls with their individual clinical measures, using a statistical method called partial least squares analysis (PLS). We demonstrated that this method could predict certain clinical measures based on a limited number of functional connections and provided benefits to the prediction performance through incorporation of the subject's age and the estimation of multiple measures simultaneously. Generalizability of the prediction model was assured through leave one out cross validation. The results showed that for BDI-II and SHAPS the most contributing connections concerned cuneus, precuneus and middle frontal cortex and areas of the cerebellum. While the relationship was similar for PANAS(n), it showed its strongest relation with functional connection between calcarine and insula.
Kosuke Yoshida, Yu Shimizu, Junichiro Yoshimoto, Shigeru Toki, Go Okada, Masahiro Takamura, Yasumasa Okamoto, Shigeto Yamawaki, Kenji Doya
BIBM9
2015 Computational Complexity Reduction for Functional Connectivity Estimation in Large Scale Neural Network
Jeonghun Baek, Shigeyuki Oba, Junichiro Yoshimoto, Kenji Doya, Shin Ishii
ICONIP (3)4
2015 Expected energy-based restricted Boltzmann machine for classification
abstract
In classification tasks, restricted Boltzmann machines (RBMs) have predominantly been used in the first stage, either as feature extractors or to provide initialization of neural networks. In this study, we propose a discriminative learning approach to provide a self-contained RBM method for classification, inspired by free-energy based function approximation (FE-RBM), originally proposed for reinforcement learning. For classification, the FE-RBM method computes the output for an input vector and a class vector by the negative free energy of an RBM. Learning is achieved by stochastic gradient-descent using a mean-squared error training objective. In an earlier study, we demonstrated that the performance and the robustness of FE-RBM function approximation can be improved by scaling the free energy by a constant that is related to the size of network. In this study, we propose that the learning performance of RBM function approximation can be further improved by computing the output by the negative expected energy (EE-RBM), instead of the negative free energy. To create a deep learning architecture, we stack several RBMs on top of each other. We also connect the class nodes to all hidden layers to try to improve the performance even further. We validate the classification performance of EE-RBM using the MNIST data set and the NORB data set, achieving competitive performance compared with other classifiers such as standard neural networks, deep belief networks, classification RBMs, and support vector machines. The purpose of using the NORB data set is to demonstrate that EE-RBM with binary input nodes can achieve high performance in the continuous input domain.
Stefan Elfwing, Eiji Uchibe, Kenji Doya
Neural Networks3
2015 Parallel Representation of Value-Based and Finite State-Based Strategies in the Ventral and Dorsal Striatum
abstract
Previous theoretical studies of animal and human behavioral learning have focused on the dichotomy of the value-based strategy using action value functions to predict rewards and the model-based strategy using internal models to predict environmental states. However, animals and humans often take simple procedural behaviors, such as the "win-stay, lose-switch" strategy without explicit prediction of rewards or states. Here we consider another strategy, the finite state-based strategy, in which a subject selects an action depending on its discrete internal state and updates the state depending on the action chosen and the reward outcome. By analyzing choice behavior of rats in a free-choice task, we found that the finite state-based strategy fitted their behavioral choices more accurately than value-based and model-based strategies did. When fitted models were run autonomously with the same task, only the finite state-based strategy could reproduce the key feature of choice sequences. Analyses of neural activity recorded from the dorsolateral striatum (DLS), the dorsomedial striatum (DMS), and the ventral striatum (VS) identified significant fractions of neurons in all three subareas for which activities were correlated with individual states of the finite state-based strategy. The signal of internal states at the time of choice was found in DMS, and for clusters of states was found in VS. In addition, action values and state values of the value-based strategy were encoded in DMS and VS, respectively. These results suggest that both the value-based strategy and the finite state-based strategy are implemented in the striatum.
Makoto Ito, Kenji Doya
PLoS Comput. Biol.2
2014 Inter Subject Correlation of Brain Activity during Visuo-Motor Sequence Learning
Krishna P. Miyapuram, Ujjval Pamnani, Kenji Doya, Raju S. Bapi
ICONIP (1)3
2014 Combining learned controllers to achieve new goals based on linearly solvable MDPs
abstract
Learning complicated behaviors usually involves intensive manual tuning and expensive computational optimization because we have to solve a nonlinear Hamilton-Jacobi-Bellman (HJB) equation. Recently, Todorov proposed a class of the so-called Linearly solvable Markov Decision Process (LMDP) which converts a nonlinear HJB equation to a linear differential equation. Linearity of the simplified HJB equation allows us to apply superposition to derive a new composite controller from a set of learned primitive controllers. However, his method was a model-based approach and it was not evaluated in a real domain. This study proposes a model-free method which is similar to the Least Squares Temporal Difference (LSTD) learning. In this method, the exponentially transformed cost function can be regarded as the discount factor in LSTD. Our proposed method is applied to learning walking behaviors with the quadruped robot to evaluate in real robot experiments. The goal of each primitive task is to go to the specific target position in the environment and that of the composite task is to approach arbitrary region represented by the primitives' target positions. Experimental results show that the composite policy can be used as a good initial policy for the new task.
Eiji Uchibe, Kenji Doya
ICRA2
2012 Neural Computations Supporting Cognition: Rumelhart Prize Symposium in Honor of Peter Dayan
Kenji Doya, John P. O'Doherty, Alexandre Pouget, Peter Bossaerts, Nathaniel D. Daw, Yael Niv
CogSci1
2012 MOSAIC for Multiple-Reward Environments
abstract
Reinforcement learning (RL) can provide a basic framework for autonomous robots to learn to control and maximize future cumulative rewards in complex environments. To achieve high performance, RL controllers must consider the complex external dynamics for movements and task (reward function) and optimize control commands. For example, a robot playing tennis and squash needs to cope with the different dynamics of a tennis or squash racket and such dynamic environmental factors as the wind. In addition, this robot has to tailor its tactics simultaneously under the rules of either game. This double complexity of the external dynamics and reward function sometimes becomes more complex when both the multiple dynamics and multiple reward functions switch implicitly, as in the situation of a real (multi-agent) game of tennis where one player cannot observe the intention of her opponents or her partner. The robot must consider its opponent's and its partner's unobservable behavioral goals (reward function). In this article, we address how an RL agent should be designed to handle such double complexity of dynamics and reward. We have previously proposed modular selection and identification for control (MOSAIC) to cope with nonstationary dynamics where appropriate controllers are selected and learned among many candidates based on the error of its paired dynamics predictor: the forward model. Here we extend this framework for RL and propose MOSAIC-MR architecture. It resembles MOSAIC in spirit and selects and learns an appropriate RL controller based on the RL controller's TD error using the errors of the dynamics (the forward model) and the reward predictors. Furthermore, unlike other MOSAIC variants for RL, RL controllers are not a priori paired with the fixed predictors of dynamics and rewards. The simulation results demonstrate that MOSAIC-MR outperforms other counterparts because of this flexible association ability among RL controllers, forward models, and reward predictors.
Norikazu Sugimoto, Masahiko Haruno, Kenji Doya, Mitsuo Kawato
Neural Comput.3
2012 Expedited review process
Kenji Doya, John G. Taylor, DeLiang Wang
Neural Networks1
2012 Loss of a Co-Editor-in-Chief and friend
Kenji Doya, DeLiang Wang
Neural Networks1
2011 Neurocomputational models of brain disorders
Vassilis Cutsuridis, Tjitske Heida, Wlodzislaw Duch, Kenji Doya
Neural Networks4
2011 An excellent year and a transition
Kenji Doya, Stephen Grossberg, John G. Taylor, DeLiang Wang
Neural Networks1
2011 Multi-scale, multi-modal neural modeling and simulation
Shin Ishii, Markus Diesmann, Kenji Doya
Neural Networks3
2010 Free-energy-based reinforcement learning in a partially observable environment
Makoto Otsuka, Junichiro Yoshimoto, Kenji Doya
ESANN3
2010 Free-Energy Based Reinforcement Learning for Vision-Based Navigation with High-Dimensional Sensory Inputs
Stefan Elfwing, Makoto Otsuka, Eiji Uchibe, Kenji Doya
ICONIP (1)4
2010 Derivatives of Logarithmic Stationary Distributions for Policy Gradient Reinforcement Learning
abstract
Most conventional policy gradient reinforcement learning (PGRL) algorithms neglect (or do not explicitly make use of) a term in the average reward gradient with respect to the policy parameter. That term involves the derivative of the stationary state distribution that corresponds to the sensitivity of its distribution to changes in the policy parameter. Although the bias introduced by this omission can be reduced by setting the forgetting rate gamma for the value functions close to 1, these algorithms do not permit gamma to be set exactly at gamma = 1. In this article, we propose a method for estimating the log stationary state distribution derivative (LSD) as a useful form of the derivative of the stationary state distribution through backward Markov chain formulation and a temporal difference learning framework. A new policy gradient (PG) framework with an LSD is also proposed, in which the average reward gradient can be estimated by setting gamma = 0, so it becomes unnecessary to learn the value functions. We also test the performance of the proposed algorithms using simple benchmark tasks and show that these can improve the performances of existing PG methods.
Tetsuro Morimura, Eiji Uchibe, Junichiro Yoshimoto, Jan Peters 0001, Kenji Doya
Neural Comput.5
2010 Editorial for 2010
Kenji Doya, Stephen Grossberg, John G. Taylor
Neural Networks1
2010 A computational neural model of goal-directed utterance selection
Michael Klein, Hans Kamp, Günther Palm, Kenji Doya
Neural Networks4
2010 A Kinetic Model of Dopamine- and Calcium-Dependent Striatal Synaptic Plasticity
abstract
Corticostriatal synapse plasticity of medium spiny neurons is regulated by glutamate input from the cortex and dopamine input from the substantia nigra. While cortical stimulation alone results in long-term depression (LTD), the combination with dopamine switches LTD to long-term potentiation (LTP), which is known as dopamine-dependent plasticity. LTP is also induced by cortical stimulation in magnesium-free solution, which leads to massive calcium influx through NMDA-type receptors and is regarded as calcium-dependent plasticity. Signaling cascades in the corticostriatal spines are currently under investigation. However, because of the existence of multiple excitatory and inhibitory pathways with loops, the mechanisms regulating the two types of plasticity remain poorly understood. A signaling pathway model of spines that express D1-type dopamine receptors was constructed to analyze the dynamic mechanisms of dopamine- and calcium-dependent plasticity. The model incorporated all major signaling molecules, including dopamine- and cyclic AMP-regulated phosphoprotein with a molecular weight of 32 kDa (DARPP32), as well as AMPA receptor trafficking in the post-synaptic membrane. Simulations with dopamine and calcium inputs reproduced dopamine- and calcium-dependent plasticity. Further in silico experiments revealed that the positive feedback loop consisted of protein kinase A (PKA), protein phosphatase 2A (PP2A), and the phosphorylation site at threonine 75 of DARPP-32 (Thr75) served as the major switch for inducing LTD and LTP. Calcium input modulated this loop through the PP2B (phosphatase 2B)-CK1 (casein kinase 1)-Cdk5 (cyclin-dependent kinase 5)-Thr75 pathway and PP2A, whereas calcium and dopamine input activated the loop via PKA activation by cyclic AMP (cAMP). The positive feedback loop displayed robust bi-stable responses following changes in the reaction parameters. Increased basal dopamine levels disrupted this dopamine-dependent plasticity. The present model elucidated the mechanisms involved in bidirectional regulation of corticostriatal synapses and will allow for further exploration into causes and therapies for dysfunctions such as drug addiction.
Takashi Nakano, Tomokazu Doi, Junichiro Yoshimoto, Kenji Doya
PLoS Comput. Biol.4
2009 Calcium Responses Model in Striatum Dependent on Timed Input Sources
Takashi Nakano, Junichiro Yoshimoto, Jeffery R. Wickens, Kenji Doya
ICANN (1)4
2009 Emergence of Different Mating Strategies in Artificial Embodied Evolution
Stefan Elfwing, Eiji Uchibe, Kenji Doya
ICONIP (2)3
2009 A Generalized Natural Actor-Critic Algorithm
abstract
Policy gradient Reinforcement Learning (RL) algorithms have received much attention in seeking stochastic policies that maximize the average rewards. In addition, extensions based on the concept of the Natural Gradient (NG) show promising learning efficiency because these regard metrics for the task. Though there are two candidate metrics, Kakades Fisher Information Matrix (FIM) and Morimuras FIM, all RL algorithms with NG have followed the Kakades approach. In this paper, we describe a generalized Natural Gradient (gNG) by linearly interpolating the two FIMs and propose an efficient implementation for the gNG learning based on a theory of the estimating function, generalized Natural Actor-Critic (gNAC). The gNAC algorithm involves a near optimal auxiliary function to reduce the variance of the gNG estimates. Interestingly, the gNAC can be regarded as a natural extension of the current state-of-the-art NAC algorithm, as long as the interpolating parameter is appropriately selected. Numerical experiments showed that the proposed gNAC algorithm can estimate gNG efficiently and outperformed the NAC algorithm.
Tetsuro Morimura, Eiji Uchibe, Junichiro Yoshimoto, Kenji Doya
NIPS4
2009 New action editors join the journal! Five exciting special issues in the works!
Kenji Doya, Stephen Grossberg, John G. Taylor
Neural Networks1
2008 Robust Population Coding in Free-Energy-Based Reinforcement Learning
Makoto Otsuka, Junichiro Yoshimoto, Kenji Doya
ICANN (1)3
2008 NeuroEvolution Based on Reusable and Hierarchical Modular Representation
Takumi Kamioka, Eiji Uchibe, Kenji Doya
ICONIP (1)3
2008 A New Natural Policy Gradient by Stationary Distribution Metric
Tetsuro Morimura, Eiji Uchibe, Junichiro Yoshimoto, Kenji Doya
ECML/PKDD (2)4
2008 Neural Networks goes electronic at twenty!
Kenji Doya, Stephen Grossberg, John G. Taylor
Neural Networks1
2008 Mini-special issue: ICONIP 2007
Masumi Ishikawa, Kenji Doya
Neural Networks2
2008 Finding intrinsic rewards by embodied evolution and constrained reinforcement learning
Eiji Uchibe, Kenji Doya
Neural Networks2
2007 Estimating Internal Variables of a Decision Maker's Brain: A Model-Based Approach for Neuroscience
Kazuyuki Samejima, Kenji Doya
ICONIP (1)2
2007 Finding Exploratory Rewards by Embodied Evolution and Constrained Reinforcement Learning in the Cyber Rodents
Eiji Uchibe, Kenji Doya
ICONIP (2)2
2007 Bayesian System Identification of Molecular Cascades
Junichiro Yoshimoto, Kenji Doya
ICONIP (1)2
2007 Reinforcement Learning State Estimator
abstract
In this study, we propose a novel use of reinforcement learning for estimating hidden variables and parameters of nonlinear dynamical systems. A critical issue in hidden-state estimation is that we cannot directly observe estimation errors. However, by defining errors of observable variables as a delayed penalty, we can apply a reinforcement learning frame-work to state estimation problems. Specifically, we derive a method to construct a nonlinear state estimator by finding an appropriate feedback input gain using the policy gradient method. We tested the proposed method on single pendulum dynamics and show that the joint angle variable could be successfully estimated by observing only the angular velocity, and vice versa. In addition, we show that we could acquire a state estimator for the pendulum swing-up task in which a swing-up controller is also acquired by reinforcement learning simultaneously. Furthermore, we demonstrate that it is possible to estimate the dynamics of the pendulum itself while the hidden variables are estimated in the pendulum swing-up task. Application of the proposed method to a two-linked biped model is also presented.
Jun Morimoto, Kenji Doya
Neural Comput.2
2007 Multiple model-based reinforcement learning explains dopamine neuronal activity
Mathieu Bertin, Nicolas Schweighofer, Kenji Doya
Neural Networks3
2007 Erratum to "Brain mechanism of reward prediction under predictable and unpredictable environmental dynamics" [Neural Networks 19 (8) (2006) 1233-1241]
Saori C. Tanaka, Kazuyuki Samejima, Go Okada, Kazutaka Ueda, Yasumasa Okamoto, Shigeto Yamawaki, Kenji Doya
Neural Networks7
2007 Nitric Oxide Regulates Input Specificity of Long-Term Depression and Context Dependence of Cerebellar Learning
abstract
Recent studies have shown that multiple internal models are acquired in the cerebellum and that these can be switched under a given context of behavior. It has been proposed that long-term depression (LTD) of parallel fiber (PF)-Purkinje cell (PC) synapses forms the cellular basis of cerebellar learning, and that the presynaptically synthesized messenger nitric oxide (NO) is a crucial "gatekeeper" for LTD. Because NO diffuses freely to neighboring synapses, this volume learning is not input-specific and brings into question the biological significance of LTD as the basic mechanism for efficient supervised learning. To better characterize the role of NO in cerebellar learning, we simulated the sequence of electrophysiological and biochemical events in PF-PC LTD by combining established simulation models of the electrophysiology, calcium dynamics, and signaling pathways of the PC. The results demonstrate that the local NO concentration is critical for induction of LTD and for its input specificity. Pre- and postsynaptic coincident firing is not sufficient for a PF-PC synapse to undergo LTD, and LTD is induced only when a sufficient amount of NO is provided by activation of the surrounding PFs. On the other hand, above-adequate levels of activity in nearby PFs cause accumulation of NO, which also allows LTD in neighboring synapses that were not directly stimulated, ruining input specificity. These findings lead us to propose the hypothesis that NO represents the relevance of a given context and enables context-dependent selection of internal models to be updated. We also predict sparse PF activity in vivo because, otherwise, input specificity would be lost.
Hideaki Ogasawara, Tomokazu Doi, Kenji Doya, Mitsuo Kawato
PLoS Comput. Biol.3
2007 Evolutionary Development of Hierarchical Learning Structures
abstract
Hierarchical reinforcement learning (RL) algorithms can learn a policy faster than standard RL algorithms. However, the applicability of hierarchical RL algorithms is limited by the fact that the task decomposition has to be performed in advance by the human designer. We propose a Lamarckian evolutionary approach for automatic development of the learning structure in hierarchical RL. The proposed method combines the MAXQ hierarchical RL method and genetic programming (GP). In the MAXQ framework, a subtask can optimize the policy independently of its parent task's policy, which makes it possible to reuse learned policies of the subtasks. In the proposed method, the MAXQ method learns the policy based on the task hierarchies obtained by GP, while the GP explores the appropriate hierarchies using the result of the MAXQ method. To show the validity of the proposed method, we have performed simulation experiments for a foraging task in three different environmental settings. The results show strong interconnection between the obtained learning structures and the given task environments. The main conclusion of the experiments is that the GP can find a minimal strategy, i.e., a hierarchy that minimizes the number of primitive subtasks that can be executed for each type of situation. The experimental results for the most challenging environment also show that the policies of the subtasks can continue to improve, even after the structure of the hierarchy has been evolutionary stabilized, as an effect of Lamarckian mechanisms
Stefan Elfwing, Eiji Uchibe, Kenji Doya, Henrik I. Christensen
IEEE Trans. Evol. Comput.3
2006 Hierarchical Chunking during Learning of Visuomotor Sequences
abstract
It is well known that learning a sequential skill involves chaining a number of primitive actions together into chunks. We describe three different experiments using an explicit visuomotor sequence learning paradigm called the m times n task. The m times n task enables hierarchical learning of sequences by presenting m elements of the sequence at a time (called the set). The entire sequence to be learned is composed of n such sets and is called a hyperset. In the first experiment, we showed the chunking phenomenon while learning a sequence as opposed to following randomly generated visual cues. We further explored the nature of chunking across sets using complex sequences in the second experiment. Finally, we investigated effector dependence of the chunking patterns in the third experiment. Our results point out the facilitating factors for chunk formation in visuomotor sequence learning.
Krishna P. Miyapuram, Raju S. Bapi, Chandrasekhar V. S. Pammi, Ahmed, Kenji Doya
IJCNN5
2006 Brain mechanism of reward prediction under predictable and unpredictable environmental dynamics
Saori C. Tanaka, Kazuyuki Samejima, Go Okada, Kazutaka Ueda, Yasumasa Okamoto, Shigeto Yamawaki, Kenji Doya
Neural Networks7
2006 Humans Can Adopt Optimal Discounting Strategy under Real-Time Constraints
abstract
Critical to our many daily choices between larger delayed rewards, and smaller more immediate rewards, are the shape and the steepness of the function that discounts rewards with time. Although research in artificial intelligence favors exponential discounting in uncertain environments, studies with humans and animals have consistently shown hyperbolic discounting. We investigated how humans perform in a reward decision task with temporal constraints, in which each choice affects the time remaining for later trials, and in which the delays vary at each trial. We demonstrated that most of our subjects adopted exponential discounting in this experiment. Further, we confirmed analytically that exponential discounting, with a decay rate comparable to that used by our subjects, maximized the total reward gain in our task. Our results suggest that the particular shape and steepness of temporal discounting is determined by the task that the subject is facing, and question the notion of hyperbolic reward discounting as a universal principle.
Nicolas Schweighofer, K. Shishida, Cheol E. Han, Yasumasa Okamoto, Saori C. Tanaka, Shigeto Yamawaki, Kenji Doya
PLoS Comput. Biol.7
2005 Biologically inspired embodied evolution of survival
abstract
Embodied evolution is a methodology for evolutionary robotics that mimics the distributed, asynchronous and autonomous properties of biological evolution. The evaluation, selection and reproduction are carried out by and between the robots, without any need for human intervention. In this paper, we propose a biologically inspired embodied evolution framework, which fully integrates self-preservation, recharging from external batteries in the environment, and self-reproduction, pair-wise exchange of genetic material, into a survival system. The individuals are explicitly evaluated for the performance of the battery capturing task, but also implicitly for the mating task by the fact that an individual that mates frequently has larger probability to spread its gene in the population. We have evaluated our method in simulation experiments and the simulation results show that the solutions obtained by our embodied evolution method were able to optimize the two survival tasks, battery capturing and mating, simultaneously. We have also performed preliminary experiments in hardware, with promising results.
Stefan Elfwing, Eiji Uchibe, Kenji Doya, Henrik I. Christensen
Congress on Evolutionary Computation3
2005 Learning Sensory Feedback to CPG with Policy Gradient for Biped Locomotion
abstract
This paper proposes a learning framework for a CPG-based biped locomotion controller using a policy gradient method. Our goal in this study is to develop an efficient learning algorithm by reducing the dimensionality of the state space used for learning. We demonstrate that an appropriate feedback controller in the CPG-based controller can be acquired using the proposed method within a few thousand trials by numerical simulations. Furthermore, we implement the learned controller on the physical biped robot to experimentally show that the learned controller successfully works in the real environment.
Takamitsu Matsubara, Jun Morimoto, Jun Nakanishi, Masa-aki Sato, Kenji Doya
ICRA5
2005 Robust Reinforcement Learning
abstract
This letter proposes a new reinforcement learning (RL) paradigm that explicitly takes into account input disturbance as well as modeling errors. The use of environmental models in RL is quite popular for both offline learning using simulations and for online action planning. However, the difference between the model and the real environment can lead to unpredictable, and often unwanted, results. Based on the theory of H(infinity) control, we consider a differential game in which a "disturbing" agent tries to make the worst possible disturbance while a "control" agent tries to make the best control input. The problem is formulated as finding a min-max solution of a value function that takes into account the amount of the reward and the norm of the disturbance. We derive online learning algorithms for estimating the value function and for calculating the worst disturbance and the best control in reference to the value function. We tested the paradigm, which we call robust reinforcement learning (RRL), on the control task of an inverted pendulum. In the linear domain, the policy and the value function learned by online algorithms coincided with those derived analytically by the linear H(infinity) control theory. For a fully nonlinear swing-up task, RRL achieved robust performance with changes in the pendulum weight and friction, while a standard reinforcement learning algorithm could not deal with these changes. We also applied RRL to the cart-pole swing-up task, and a robust swing-up policy was acquired.
Jun Morimoto, Kenji Doya
Neural Comput.2
2004 Chunking Phenomenon in Complex Sequential Skill Learning in Humans
Chandrasekhar V. S. Pammi, Krishna P. Miyapuram, Raju S. Bapi, Kenji Doya
ICONIP4
2004 Multi-agent reinforcement learning: using macro actions to learn a mating task
abstract
Standard reinforcement learning methods are inefficient and often inadequate for learning cooperative multi-agent tasks. For these kinds of tasks the behavior of one agent strongly depends on dynamic interaction with other agents, not only with the interaction with a static environment as in standard reinforcement learning. The success of the learning is therefore coupled to the agents' ability to predict the other agents behaviors. In this study we try to overcome this problem by adding a few simple macro actions, actions that are extended in time for more than one time step. The macro actions improve the learning by making search of the state space more effective and thereby making the behavior more predictable for the other agent. In this study we have considered a cooperative mating task, which is the first step towards our aim to perform embodied evolution, where the evolutionary selection process is an integrated part of the task. We show, in simulation and hardware, that in the case of learning without macro actions, the agents fail to learn a meaningful behavior. In contrast, for the learning with macro action the agents learn a good mating behavior in reasonable time, in both simulation and hardware.
Stefan Elfwing, Eiji Uchibe, Kenji Doya, Henrik I. Christensen
IROS3
2004 Responding to Modalities with Different Latencies
abstract
Motor control depends on sensory feedback in multiple modalities with different latencies. In this paper we consider within the framework of re- inforcement learning how different sensory modalities can be combined and selected for real-time, optimal movement control. We propose an actor-critic architecture with multiple modules, whose output are com- bined using a softmax function. We tested our architecture in a simu- lation of a sequential reaching task. Reaching was initially guided by visual feedback with a long latency. Our learning scheme allowed the agent to utilize the somatosensory feedback with shorter latency when the hand is near the experienced trajectory. In simulations with different latencies for visual and somatosensory feedback, we found that the agent depended more on feedback with shorter latency.
Fredrik Bissmarck, Hiroyuki Nakahara, Kenji Doya, Okihide Hikosaka
NIPS3
2004 Reinforcement learning with via-point representation
Hiroyuki Miyamoto, Jun Morimoto, Kenji Doya, Mitsuo Kawato
Neural Networks3
2003 Evolving recurrent neural controllers for sequential tasks: a parallel implementation
abstract
Evolution of complex behaviours requires a careful selection of genetic algorithm parameters and a large number of computations. In this paper, we considered evolution of recurrent neural controllers for nonMarkovian sequential tasks using a regional model genetic algorithm. The subpopulations apply different strategies and compete with each other. Simulation and experimental results using cyber rodent robot indicate that regional model outperformed single population genetic algorithm by distributing the genetic resources effectively as different strategies successful during the course of evolution.
Genci Capi, Kenji Doya
IEEE Congress on Evolutionary Computation2
2003 An Evolutionary Approach to Automatic Construction of the Structure in Hierarchical Reinforcement Learning
Stefan Elfwing, Eiji Uchibe, Kenji Doya
GECCO3
2003 Evolution of meta-parameters in reinforcement learning algorithm
abstract
A crucial issue in reinforcement learning applications is how to set meta-parameters, such as the learning rate and "temperature" for exploration, to match the demands of the task and the environment. In this paper, we propose a method to adjust meta-parameters of reinforcement learning by real-number genetic algorithm. It was shown in simulations of foraging tasks that appropriate settings of meta-parameters, which are strongly dependent on each other, can be found by evolution. Furthermore, we verified in hardware experiments using cyber rodent (CR) robots that the meta-parameters evolved in simulation are helpful for learning in real hardware.
Anders Eriksson, Genci Capi, Kenji Doya
IROS3
2003 Estimating Internal Variables and Paramters of a Learning Agent by a Particle Filter
abstract
When we model a higher order functions, such as learning and memory, we face a difficulty of comparing neural activities with hidden variables that depend on the history of sensory and motor signals and the dynam- ics of the network. Here, we propose novel method for estimating hidden variables of a learning agent, such as connection weights from sequences of observable variables. Bayesian estimation is a method to estimate the posterior probability of hidden variables from observable data sequence using a dynamic model of hidden and observable variables. In this pa- per, we apply particle filter for estimating internal parameters and meta- parameters of a reinforcement learning model. We verified the effective- ness of the method using both artificial data and real animal behavioral data.
Kazuyuki Samejima, Kenji Doya, Yasumasa Ueda, Minoru Kimura
NIPS2
2003 Different Cortico-Basal Ganglia Loops Specialize in Reward Prediction at Different Time Scales
abstract
To understand the brain mechanisms involved in reward prediction on different time scales, we developed a Markov decision task that requires prediction of both immediate and future rewards, and ana- lyzed subjects’ brain activities using functional MRI. We estimated the time course of reward prediction and reward prediction error on different time scales from subjects' performance data, and used them as the explanatory variables for SPM analysis. We found topog- raphic maps of different time scales in medial frontal cortex and striatum. The result suggests that different cortico-basal ganglia loops are specialized for reward prediction on different time scales.
Saori C. Tanaka, Kenji Doya, Go Okada, Kazutaka Ueda, Yasumasa Okamoto, Shigeto Yamawaki
NIPS2
2003 Inter-module credit assignment in modular reinforcement learning
Kazuyuki Samejima, Kenji Doya, Mitsuo Kawato
Neural Networks2
2003 Meta-learning in Reinforcement Learning
Nicolas Schweighofer, Kenji Doya
Neural Networks2
2002 Multiple Model-Based Reinforcement Learning
abstract
We propose a modular reinforcement learning architecture for nonlinear, nonstationary control tasks, which we call multiple model-based reinforcement learning (MMRL). The basic idea is to decompose a complex task into multiple domains in space and time based on the predictability of the environmental dynamics. The system is composed of multiple modules, each of which consists of a state prediction model and a reinforcement learning controller. The "responsibility signal," which is given by the softmax function of the prediction errors, is used to weight the outputs of multiple modules, as well as to gate the learning of the prediction models and the reinforcement learning controllers. We formulate MMRL for both discrete-time, finite-state case and continuous-time, continuous-state case. The performance of MMRL was demonstrated for discrete case in a nonstationary hunting task in a grid world and for continuous case in a nonlinear, nonstationary control task of swinging up a pendulum with variable physical parameters.
Kenji Doya, Kazuyuki Samejima, Ken-ichi Katagiri, Mitsuo Kawato
Neural Comput.1
2002 Metalearning and neuromodulation
Kenji Doya
Neural Networks1
2002 Introduction for 2002 Special Issue: Computational Models of Neuromodulation
Kenji Doya, Peter Dayan, Michael E. Hasselmo
Neural Networks1
2000 Acquisition of Stand-up Behavior by a Real Robot using Hierarchical Reinforcement Learning
Jun Morimoto, Kenji Doya
ICML2
2000 Robust Reinforcement Learning
abstract
This paper proposes a new reinforcement learning (RL) paradigm that explicitly takes into account input disturbance as well as mod(cid:173) eling errors. The use of environmental models in RL is quite pop(cid:173) ular for both off-line learning by simulations and for on-line ac(cid:173) tion planning. However, the difference between the model and the real environment can lead to unpredictable, often unwanted results. Based on the theory of H oocontrol, we consider a differential game in which a 'disturbing' agent (disturber) tries to make the worst possible disturbance while a 'control' agent (actor) tries to make the best control input. The problem is formulated as finding a min(cid:173) max solution of a value function that takes into account the norm of the output deviation and the norm of the disturbance. We derive on-line learning algorithms for estimating the value function and for calculating the worst disturbance and the best control in refer(cid:173) ence to the value function. We tested the paradigm, which we call "Robust Reinforcement Learning (RRL)," in the task of inverted pendulum. In the linear domain, the policy and the value func(cid:173) tion learned by the on-line algorithms coincided with those derived analytically by the linear H ootheory. For a fully nonlinear swing(cid:173) up task, the control by RRL achieved robust performance against changes in the pendulum weight and friction while a standard RL control could not deal with such environmental changes.
Jun Morimoto, Kenji Doya
NIPS2
2000 Reinforcement Learning in Continuous Time and Space
abstract
This article presents a reinforcement learning framework for continuous-time dynamical systems without a priori discretization of time, state, and action. Based on the Hamilton-Jacobi-Bellman (HJB) equation for infinite-horizon, discounted reward problems, we derive algorithms for estimating value functions and improving policies with the use of function approximators. The process of value function estimation is formulated as the minimization of a continuous-time form of the temporal difference (TD) error. Update methods based on backward Euler approximation and exponential eligibility traces are derived, and their correspondences with the conventional residual gradient, TD(0), and TD(lambda) algorithms are shown. For policy improvement, two methods-a continuous actor-critic method and a value-gradient-based greedy policy-are formulated. As a special case of the latter, a nonlinear feedback control law using the value gradient and the model of the input gain is derived. The advantage updating, a model-free algorithm derived previously, is also formulated in the HJB-based framework. The performance of the proposed algorithms is first tested in a nonlinear control task of swinging a pendulum up with limited torque. It is shown in the simulations that (1) the task is accomplished by the continuous actor-critic method in a number of trials several times fewer than by the conventional discrete actor-critic method; (2) among the continuous policy update methods, the value-gradient-based policy with a known or learned dynamic model performs several times better than the actor-critic method; and (3) a value function update using exponential eligibility traces is more efficient and stable than that based on Euler approximation. The algorithms are then tested in a higher-dimensional task: cart-pole swing-up. This task is accomplished in several hundred trials using the value-gradient-based policy with a learned dynamic model.
Kenji Doya
Neural Comput.1
1999 What are the computations of the cerebellum, the basal ganglia and the cerebral cortex?
Kenji Doya
Neural Networks1
1998 A Sequence Learning Architecture Based on Cortico-Basal Ganglionic Loops and Reinforcement Learning
Raju S. Bapi, Kenji Doya
ICONIP2
1998 Hierarchical Reinforcement Learning of Low-Dimensional Subgoals and High-Dimensional Trajectories
Jun Morimoto, Kenji Doya
ICONIP2
1998 A Model of the Electrophysiological Properties of the Inferior Olive Neurons
Nicolas Schweighofer, Kenji Doya, Mitsuo Kawato
ICONIP2
1998 Reinforcement learning of dynamic motor sequence: learning to stand up
abstract
We propose a learning method for implementing human-like sequential movements in robots. As an example of dynamic sequential movement, we consider the "stand-up" task for a two-joint, three-link robot. In contrast to the case of steady walking or standing, the desired trajectory for such a transient behavior is very difficult to derive. The goal of the task is to find a path that links a lying state to an upright state under the constraints of the system dynamics. The geometry of the robot is such that there is no static solution; the robot has to stand up dynamically utilizing the momentum of its body. We use reinforcement learning, in particular, a continuous time and state temporal difference (TD) learning method. For successful results, we use 1) an efficient method of value function approximation in a high-dimensional state space, and 2) a hierarchical architecture which divides a large state space into a few smaller pieces.
Jun Morimoto, Kenji Doya
IROS2
1998 Near Saddle-Node Bifurcation Behavior as Dynamics in Working Memory for Goal-Directed Behavior
abstract
In consideration of working memory as a means for goal-directed behavior in nonstationary environments, we argue that the dynamics of working memory should satisfy two opposing demands: long-term maintenance and quick transition. These two characteristics are contradictory within the linear domain. We propose the near-saddle-node bifurcation behavior of a sigmoidal unit with a self-connection as a candidate of the dynamical mechanism that satisfies both of these demands. It is shown in evolutionary programming experiments that the near-saddle-node bifurcation behavior can be found in recurrent networks optimized for a task that requires efficient use of working memory. The results suggests that the near-saddle-node bifurcation behavior may be a functional necessity for survival in nonstationary environments.
Hiroyuki Nakahara, Kenji Doya
Neural Comput.2
1996 Efficient Nonlinear Control with Actor-Tutor Architecture
Kenji Doya
NIPS1
1995 Temporal Difference Learning in Continuous Time and Space
Kenji Doya
NIPS1
1995 Dynamics of Attention as Near Saddle-Node Bifurcation Behavior
Hiroyuki Nakahara, Kenji Doya
NIPS2
1994 A Novel Reinforcement Model of Birdsong Vocalization Learning
abstract
Songbirds learn to imitate a tutor song through auditory and motor learn(cid:173) ing. We have developed a theoretical framework for song learning that accounts for response properties of neurons that have been observed in many of the nuclei that are involved in song learning. Specifically, we suggest that the anteriorforebrain pathway, which is not needed for song production in the adult but is essential for song acquisition, provides synaptic perturbations and adaptive evaluations for syllable vocalization learning. A computer model based on reinforcement learning was con(cid:173) structed that could replicate a real zebra finch song with 90% accuracy based on a spectrographic measure. The second generation of the bird(cid:173) song model replicated the tutor song with 96% accuracy.
Kenji Doya, Terrence J. Sejnowski
NIPS1
1994 Dimension Reduction of Biological Neuron Models by Artificial Neural Networks
abstract
An artificial neural network approach to dimension reduction of dynamical systems is proposed and applied to conductance-based neuron models. Networks with bottleneck layers of continuous-time dynamical units could make a two-dimensional model from the trajectories of the Hodgkin-Huxley model and a three-dimensional model from the trajectories of a six-dimensional bursting neuron model. Nullcline analysis of these reduced models revealed the bifurcations of the dynamical system underlying firing and bursting behaviors.
Kenji Doya, Allen I. Selverston
Neural Comput.1
1993 A Hodgkin-Huxley Type Neuron Model That Learns Slow Non-Spike Oscillations
Kenji Doya, Allen I. Selverston, Peter F. Rowat
NIPS1
1992 Maaping Between Neural and Physical Activities of the Lobster Gastric Mill
Kenji Doya, Mary E. T. Boyle, Allen I. Selverston
NIPS1
1991 Adaptive Synchronization of Neural and Physical Oscillators
Kenji Doya, Shuji Yoshizawa
NIPS1
1990 Memorizing hierarchical temporal patterns in analog neuron networks
abstract
A neural network model of a hierarchical temporal pattern generator is proposed. The learning algorithms for simple temporal pattern-generator networks are discussed. Learning schemes for constructing hierarchical networks of the temporal pattern generators are investigated. A simple simulation of the hierarchical network is shown
Kenji Doya, Shuji Yoshizawa
IJCNN1
1989 Adaptive neural oscillator using continuous-time back-propagation learning
Kenji Doya, Shuji Yoshizawa
Neural Networks1