EDBT 2026 Demo / reviewers in the wild / expert
Taisuke Kobayashi
dblp:123/6873
· DBLP profile ↗
26ranked-venue papers
18as first author
12since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 16 first-author · 11 since 2021Systems, architecture and hardware · 10 · 7 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Consolidated Adaptive T-soft Update for Deep Reinforcement LearningabstractDemand for deep reinforcement learning (DRL) is gradually increased to enable robots to perform complex tasks, while DRL is known to be unstable. As a technique to stabilize its learning, a target network that slowly and asymptotically matches a main network is widely employed to generate stable pseudo-supervised signals. Recently, T-soft update has been proposed as a noise-robust update rule for the target network and has contributed to improving the DRL performance. However, the noise robustness of T-soft update is specified by a hyperparameter, which should be tuned for each task, and suppression of updates would cause deviation of the two networks. This study develops a consolidated and adaptive T-soft (CAT-soft) update based on approximate maximum likelihood estimation of student-t distribution and an additional consolidation. Since the noise robustness is represented by a model parameter of the student-t distribution, this method makes the noise robustness adaptive. In addition, the parameters of the main network, those that deviate from the target network, are consolidated to the target network. The proposed method outperformed the conventional methods in numerical simulations. Taisuke Kobayashi |
IJCNN | 1 |
| 2024 | Revisiting experience replayable conditions
Taisuke Kobayashi |
Appl. Intell. | 1 |
| 2023 | Domains as Objectives: Multi-Domain Reinforcement Learning with Convex-Coverage Set Learning for Domain Uncertainty AwarenessabstractDomain randomization (DR) is a powerful framework that has allowed the transfer of policies from randomized domain (a.k.a. simulation) to real robots with little to no retraining requirement. However, because the policy has to perform well for many different domain conditions, DR tends to produce sub-optimal policies that can be too conservative on the target real system. This problem is further exacerbated the larger the randomized domain is. To tackle this issue, recent works have proposed to learn universal policies (UP) with domain knowledge such that they can adapt their behavior to each domain when paired with an online system identifier (OSI). However, in most applications, perfect identifications of the target domain can be impossible. In this paper, by drawing similarities between DR as a multi-domain reinforcement learning and multi-objective reinforcement learning (MORL), we propose to learn a UP over the convex coverage set borrowed from the MORL theory. Thanks to this, our method learns a UP that effectively captures different sub-domains of the uncertainty set and can therefore adapt its behavior based on an OSI uncertainty, unlocking the power of stochastic system identification with no retraining requirement. This pseudo-MORL framework also contains previous works in DR and robust reinforcement learning. We conduct simulations on Mujoco tasks and experiments on a real D'Claw robot, revealing the effectiveness of our domain-uncertainty-aware UP for sim-to-real transfer. Wendyam Eric Lionel Ilboudo, Taisuke Kobayashi, Takamitsu Matsubara |
IROS | 2 |
| 2023 | AdaTerm: Adaptive T-distribution estimated robust moments for Noise-Robust stochastic gradient optimization
Wendyam Eric Lionel Ilboudo, Taisuke Kobayashi, Takamitsu Matsubara |
Neurocomputing | 2 |
| 2022 | L2C2: Locally Lipschitz Continuous Constraint towards Stable and Smooth Reinforcement LearningabstractThis paper proposes a new regularization technique for reinforcement learning (RL) towards making policy and value functions smooth and stable. RL is known for the instability of the learning process and the sensitivity of the acquired policy to noise. Several methods have been proposed to resolve these problems, and in summary, the smoothness of policy and value functions learned mainly in RL contributes to these problems. However, if these functions are extremely smooth, their expressiveness would be lost, resulting in not obtaining the global optimal solution. This paper therefore considers RL under local Lipschitz continuity constraint, so-called L2C2. By designing the spatio-temporal locally compact space for L2C2 from the state transition at each time step, the moderate smoothness can be achieved without loss of expressiveness. Numerical noisy simulations verified that the proposed L2C2 outperforms the task performance while smoothing out the robot action generated from the learned policy. Taisuke Kobayashi |
IROS | 1 |
| 2022 | Optimistic reinforcement learning by forward Kullback-Leibler divergence optimization
Taisuke Kobayashi |
Neural Networks | 1 |
| 2022 | Latent Representation in Human-Robot Interaction With Explicit Consideration of Periodic DynamicsabstractThis article presents a new data-driven framework for analyzing periodic physical human–robot interaction (pHRI) in latent state space. The model representing pHRI is critical for elaborating human understanding and/or robot control during pHRI. Recent advancements in deep learning technology would allow us to train such a model on a dataset collected from the actual pHRI. Our framework is based on a variational recurrent neural network (VRNN), which can process time-series data generated by a pHRI. This study modifies VRNN to explicitly integrate the latent dynamics from robot to human and to distinguish it from a human state estimate module. Furthermore, to analyze periodic motions, such as walking, we integrate VRNN with a new recurrent network based on reservoir computing (RC), which has random and fixed connections between numerous neurons. By boosting RC into a complex domain, periodic behavior can be represented as phase rotation in the complex domain without decaying the amplitude. A rope rotation/swinging experiment was used to validate the proposed framework. The proposed framework, trained on the collected experiment dataset, achieved the latent state space in which variation in periodic motions can be distinguished. The best prediction accuracy of the human observations and robot actions was obtained in such a well-distinguished space. Taisuke Kobayashi, Shingo Murata, Tetsunari Inamura |
IEEE Trans. Hum. Mach. Syst. | 1 |
| 2022 | Robust Stochastic Gradient Descent With Student-t Distribution Based First-Order MomentumabstractRemarkable achievements by deep neural networks stand on the development of excellent stochastic gradient descent methods. Deep-learning-based machine learning algorithms, however, have to find patterns between observations and supervised signals, even though they may include some noise that hides the true relationship between them, more or less especially in the robotics domain. To perform well even with such noise, we expect them to be able to detect outliers and discard them when needed. We, therefore, propose a new stochastic gradient optimization method, whose robustness is directly built in the algorithm, using the robust student-t distribution as its core idea. We integrate our method to some of the latest stochastic gradient algorithms, and in particular, Adam, the popular optimizer, is modified through our method. The resultant algorithm, called t-Adam, along with the other stochastic gradient methods integrated with our core idea is shown to effectively outperform Adam and their original versions in terms of robustness against noise on diverse tasks, ranging from regression and classification to reinforcement learning problems. Wendyam Eric Lionel Ilboudo, Taisuke Kobayashi, Kenji Sugimoto |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | Proximal Policy Optimization with Relative Pearson DivergenceabstractThe recent remarkable progress of deep reinforcement learning (DRL) stands on regularization of policy for stable and efficient learning. A popular method, named proximal policy optimization (PPO), has been introduced for this purpose. PPO clips density ratio of the latest and baseline policies with a threshold, while its minimization target is unclear. As another problem of PPO, the symmetric threshold is given numerically while the density ratio itself is in asymmetric domain, thereby causing unbalanced regularization of the policy. This paper therefore proposes a new variant of PPO by considering a regularization problem of relative Pearson (RPE) divergence, so-called PPO-RPE. This regularization yields the clear minimization target, which constrains the latest policy to the baseline one. Through its analysis, the intuitive threshold-based design consistent with the asymmetry of the threshold and the domain of density ratio can be derived. Through four benchmark tasks, PPO-RPE performed as well as or better than the conventional methods in terms of the task performance by the learned policy. Taisuke Kobayashi |
ICRA | 1 |
| 2021 | Adaptive t-Momentum-based Optimization for Unknown Ratio of Outliers in Amateur Data in Imitation LearningabstractBehavioral cloning (BC) bears a high potential for safe and direct transfer of human skills to robots. However, demonstrations performed by human operators often contain noise or imperfect behaviors that can affect the efficiency of the imitator if left unchecked. In order to allow the imitators to effectively learn from imperfect demonstrations, we propose to employ the robust t-momentum optimization algorithm. This algorithm builds on the Student's t-distribution in order to deal with heavy-tailed data and reduce the effect of outlying observations. We extend the t-momentum algorithm to allow for an adaptive and automatic robustness and show empirically how the algorithm can be used to produce robust BC imitators against datasets with unknown heaviness. Indeed, the imitators trained with the t-momentum-based Adam optimizers displayed robustness to imperfect demonstrations on two different manipulation tasks with different robots and revealed the capability to take advantage of the additional data while reducing the adverse effect of non-optimal behaviors. Wendyam Eric Lionel Ilboudo, Taisuke Kobayashi, Kenji Sugimoto |
IROS | 2 |
| 2021 | Bottom-up multi-agent reinforcement learning by reward shaping for cooperative-competitive tasks
Takumi Aotani, Taisuke Kobayashi, Kenji Sugimoto |
Appl. Intell. | 2 |
| 2021 | t-soft update of target network for deep reinforcement learning
Taisuke Kobayashi, Wendyam Eric Lionel Ilboudo |
Neural Networks | 1 |
| 2020 | Reinforcement learning for quadrupedal locomotion with design of continual-hierarchical curriculum
Taisuke Kobayashi, Toshiki Sugino |
Eng. Appl. Artif. Intell. | 1 |
| 2019 | Variational Deep Embedding with Regularized Student-t Mixture Model
Taisuke Kobayashi |
ICANN (3) | 1 |
| 2019 | Student-t policy in reinforcement learning to acquire global optimum of robot control
Taisuke Kobayashi |
Appl. Intell. | 1 |
| 2018 | Check Regularization: Combining Modularity and Elasticity for Memory Consolidation
Taisuke Kobayashi |
ICANN (2) | 1 |
| 2018 | Practical Fractional-Order Neuron Dynamics for Reservoir Computing
Taisuke Kobayashi |
ICANN (3) | 1 |
| 2018 | Bottom-up Multi-agent Reinforcement Learning for Selective CooperationabstractApplications of multi-agent system like cooperative transport are found in various domains of real world. Due to the complexity inherent in multi-agent system, however, handling with preprogramming is difficult. Multi-agent reinforcement learning (MARL), which is a framework to make multiple agents in the same environment learn their policies simultaneously using reinforcement learning, is receiving attention. In the conventional MARL, although decentralization is essential for feasible learning, rewards for the agents have been allocated from a centralized system in the environment. Instead of such "top-down" MARL, to achieve the completely distributed autonomous systems, we tackle a new paradigm named "bottom-up" MARL, where the agents get their own rewards. The bottom-up MARL requires to share the respective rewards for emerging orderly group behaviors, which cannot be acquired merely by maximizing the mean of them. We therefore propose the architecture that has three components: estimating rewards of other agents; selecting rewards to reinforce from the correlation, and; promoting the exploration to find unknown correlation. The proposed architecture is verified that every element is essential by numerical simulation performed in stages. A similar task is also accomplished in dynamical simulation under the same conditions as the actual robots. Takumi Aotani, Taisuke Kobayashi, Kenji Sugimoto |
SMC | 2 |
| 2016 | Unified bipedal gait for walking and running by dynamics-based virtual holonomic constraint in PDACabstractConventional humanoids have achieved walking and running by independent controllers, even though a transition between these independent motions should be added to connect them while ensuring stability. In contrast, human selects walking/running at low/high speed in terms of energy efficiency, and transits between them naturally. This fact implies that human gait shares the inherent controller among walking and running despite their quite different appearances. Hence, we propose a “unified bipedal gait,” which includes walking, running, and the transition. The unified bipedal gait has the inherent controller: passive dynamic autonomous control (PDAC) with a damping and spring loaded inverted pendulum (D-SLIP) model. The PDAC constrains the humanoid natural dynamics by a virtual holonomic constraint (VHC) that degenerates the natural manifold of the states for stabilization. Compliance in the D-SLIP is capable to yield the required characteristics of walking/running: low/high compliant legs for walking/running. Thus, a novel VHC is designed to extract the required characteristics of walking/running from the D-SLIP dynamics and form the proper manifold. As a result, we achieved the unified bipedal gait that bifurcates to walking and running via the natural transition. The high energy efficiency was confirmed in this unified bipedal gait at any gait speed. Taisuke Kobayashi, Yasuhisa Hasegawa, Kousuke Sekiyama, Tadayoshi Aoyama, Toshio Fukuda |
ICRA | 1 |
| 2016 | Quasi-passive dynamic autonomous control to enhance horizontal and turning gait speed controlabstractThis paper proposes a quasi-passive dynamic autonomous control (Q-PDAC) for a three-dimensional (3-D) bipedal gait of humanoid robots from start points to goal points. The major approach for 3-D traveling is currently footstep planning by a constantly stable gait with an emphasis on its accurate and secure traveling. However, energy would potentially be wasted when the robot accurately travels according to the planned footsteps. In contrast, a limit-cycle-based gait possesses the good efficiency by a gait speed control, although the accurate and secure traveling is difficult for it. Its gait speed control is unfortunately not enough to freely travel on 3-D spaces: shortages of a turning speed control, trackability of horizontal speed, and stability of the bipedal gait. Hence, the Q-PDAC supplies three proper angular momenta by hip and ankle joints to achieve the turning motion and enhance the trackability of horizontal motion. Three angular momenta are simply designed consistent in the PDAC dynamics, and achieved the sufficient gait speed control for 3-D traveling. As a result, the robot can efficiently travel from the start point to the goal point while following a leader point not to collide with walls. Taisuke Kobayashi, Kousuke Sekiyama, Yasuhisa Hasegawa, Tadayoshi Aoyama, Toshio Fukuda |
IROS | 1 |
| 2015 | Optimal use of arm-swing for bipedal walking controlabstractWalking capability composed of stability and efficiency is one of the most important issues in the field of humanoid robots. An effective swing of the arms is expected to enhance the walking capability under the constraints from the limited body. We propose an arm-swing method to enhance the stability and efficiency by selecting optimal arm-swing strategy depending on the walking conditions. In this research, we select the optimal strategy between the support of the center of gravity (COG) tracking for stability and the walk without arm-swing for efficiency. To support the COG tracking, we employ a predictive control. States are defined as an inverted pendulum model and inputs are given as an inertial force of arm-swing. Input and output weights in the predictive control are adjustable by a support weight introduced in this paper. Selection of the optimal support weight by a selection algorithm for locomotion (Su-SAL) switches the two strategies by adjusting the ratio of input and output (I/O) weights. Su-SAL maximizes the efficiency while keeping the stability in comparison with the case of the constant support weight. Taisuke Kobayashi, Kousuke Sekiyama, Tadayoshi Aoyama, Yasuhisa Hasegawa, Toshio Fukuda |
ICRA | 1 |
| 2015 | Selection Algorithm for Locomotion Based on the Evaluation of Falling RiskabstractAn environmentally specific type of locomotion (e.g., bipedal or quadrupedal walking) is effective only under the specified environments. However, other conditions could cause physical body constraints and decrease mobility. Despite these constraints, legged robots are desired with high overall mobility such that they can walk under various conditions. Thus, a combination of types of locomotion is needed to maximize overall mobility. We have developed a gorilla-type robot, which can switch between bipedal and quadrupedal walking. A selection technique to optimize locomotion choice would be beneficial to the robot, which will experience challenging situations when walking through complex terrains, receiving disturbances, or malfunctioning. We present a selection algorithm for locomotion (SAL) that improves overall mobility by autonomously selecting the optimal locomotion. The falling risk of each locomotion mode is evaluated with a Bayesian network to represent the robot's situation. The evaluation function for the SAL determines the optimal locomotion choice based on falling risk and moving speed. In this paper, the SAL is used for two state variables of locomotion: gait (Ga-SAL) and speed (Sp-SAL). Both the simulations and experiments validated that the robot traveled efficiently in complex environments. Taisuke Kobayashi, Tadayoshi Aoyama, Kousuke Sekiyama, Toshio Fukuda |
IEEE Trans. Robotics | 1 |
| 2013 | Locomotion selection strategy for multi-locomotion robot based on stability and efficiencyabstractThis paper shows improvement of stability and efficiency for mobility using locomotion selection strategy. First strategy is the selection of a gait relying on locomotion rewards. The locomotion reward has been proposed as an indicator for selection algorithm based on Falling Risk and the moving speed. This strategy has achieved a capability of large changes of uncertainties, such as a steep slope. Second strategy is adjustment of moving speed by the extended locomotion reward that explicitly shows the relationship between the moving speed and Falling Risk. The robot aims at the maximum moving speed without a falling, and removes small changes of uncertainties as a result. We performed an experiment in order to confirm effects of two strategies in an environment that includes a rough terrain as a small uncertainty and two steps as a large uncertainty. The robot improved the moving speed about 37.5% from the case of only using the gait selection strategy. Taisuke Kobayashi, Tadayoshi Aoyama, Masafumi Sobajima, Kousuke Sekiyama, Toshio Fukuda |
IROS | 1 |
| 2012 | Locomotion selection of Multi-Locomotion Robot based on Falling Risk and moving efficiencyabstractThis paper deals with a method of locomotion selection based on Falling Risk and moving efficiency. The robot estimates information from sensors by solving state equation. The robot evaluates the Falling Risk as an indicator of uncertainty. Falling Risk is derived from measured information by using Bayesian Network. Locomotion selection during walking is modeled as a Semi-Markov Decision Process and the most appropriate locomotion is selected by using the greedy algorithm. As a result, the robot can move in the environment that is difficult to travel by single locomotion mode, maintaining the maximum moving efficiency. Taisuke Kobayashi, Tadayoshi Aoyama, Kousuke Sekiyama, Zhiguo Lu, Yasuhisa Hasegawa, Toshio Fukuda |
IROS | 1 |
| 2012 | Optimal control of energetically efficient ladder decent motion with internal stress adjustment using key joint methodabstractFor multi-contact robot motion, a closed chain is formed by robot links and the environment. This paper proposes a new methodology named “key joint method” for reducing the energy cost by adjusting an internal stress inside a closed chain. Firstly, we analyze the internal stress theoretically taking the degrees of freedom (DOF) and the number of position actuated joints into consideration, then a practical key joint method is proposed by changing a suitable redundant position controlled joint to be force control. After that, a parametric family is introduced for representing various of possible motions subjected to the robot dynamics and other constraints. Finally, a general optimization method is proposed for planning an energetically efficient multi-contact robot motion taking the motion trajectories and internal stress into consideration. As an example, the pace gait ladder decent motion is taken to explain the principle and realization of the proposed method. As experimental evaluation shows, the key joint method is effective for reducing the energy cost in the multi-contact motion. Zhiguo Lu, Kousuke Sekiyama, Tadayoshi Aoyama, Yasuhisa Hasegawa, Taisuke Kobayashi, Toshio Fukuda |
IROS | 5 |
| 1994 | Computer Graphics System for Reproducing Three-dimensional Shape from Idea SketchabstractAbstract This paper describes the technical features and method of implementation of our designer support system which enables reproducing a three‐dimensional shape from an idea sketch speedily and modifying the reproduced shape easily, thereby facilitating the consistency of the shape to be checked from a design viewpoint. The designer support system was developed as a tool to cut the time and labor designers have to spend in the design process and help designers to fully display their creativity–the essential attribute of designers. This system uses cross‐section lines of an idea sketch of an object drawn by an industrial designer as input data to reproduce automatically a three‐dimensional wireframe model of the object by a newly developed three‐dimensional graphic algorithm including graphic constraints. In addition, it automatically creates planes by scanning the area enclosed by reproduced curved lines and subjects them to shading. In this process, the dimensions of the object need not be input to the system. The designer can operate the system simply by using the mouse to input appropriate data. Thus, the system permits the designer to easily express his idea in the form of a three‐dimensional shape and study idea variations on the display screen without creating a mockup. Makoto Akeo, Hiroshi Hashimoto, Taisuke Kobayashi, Tetsuo Shibusawa |
Comput. Graph. Forum | 3 |