Motoki Omura

dblp:371/5840 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 81% Optimization for machine learning · 19%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
value function estimation
1.022025
Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning · AAAI 2024
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning · ICML 2025
Machine learning › Reinforcement learning
actor-critic methods
0.912025
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning · ICML 2025
Machine learning › Optimization for machine learning › stochastic search
annealing
0.912025
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning · ICML 2025
Machine learning › Reinforcement learning › dynamic programming
bellman operator
0.912025
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning · ICML 2025
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning
0.812024
Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning · AAAI 2024
Machine learning › Reinforcement learning › value function estimation
overestimation bias
0.312025
Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning · ICML 2025

Methods — techniques the papers use, named apart from their topics

annealing · 0.9TD3 · 0.9SAC · 0.9synthetic noise · 0.8symmetric q-learning · 0.8
YearPublicationVenuePosition
2025 Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning
abstract
For continuous action spaces, actor-critic methods are widely used in online reinforcement learning (RL). However, unlike RL algorithms for discrete actions, which generally model the optimal value function using the Bellman optimality operator, RL algorithms for continuous actions typically model Q-values for the current policy using the Bellman operator. These algorithms for continuous actions rely exclusively on policy updates for improvement, which often results in low sample efficiency. This study examines the effectiveness of incorporating the Bellman optimality operator into actor-critic frameworks. Experiments in a simple environment show that modeling optimal values accelerates learning but leads to overestimation bias. To address this, we propose an annealing approach that gradually transitions from the Bellman optimality operator to the Bellman operator, thereby accelerating learning while mitigating bias. Our method, combined with TD3 and SAC, significantly outperforms existing approaches across various locomotion and manipulation tasks, demonstrating improved performance and robustness to hyperparameters related to optimality. The code for this study is available at https://github.com/motokiomura/annealed-q-learning.
Motoki Omura, Kazuki Ota, Takayuki Osa, Yusuke Mukuta, Tatsuya Harada
ICML1
2024 Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning
abstract
In deep reinforcement learning, estimating the value function to evaluate the quality of states and actions is essential. The value function is often trained using the least squares method, which implicitly assumes a Gaussian error distribution. However, a recent study suggested that the error distribution for training the value function is often skewed because of the properties of the Bellman operator, and violates the implicit assumption of normal error distribution in the least squares method. To address this, we proposed a method called Symmetric Q-learning, in which the synthetic noise generated from a zero-mean distribution is added to the target values to generate a Gaussian error distribution. We evaluated the proposed method on continuous control benchmark tasks in MuJoCo. It improved the sample efficiency of a state-of-the-art reinforcement learning method by reducing the skewness of the error distribution.
Motoki Omura, Takayuki Osa, Yusuke Mukuta, Tatsuya Harada
AAAI1