VLDB 2026 Research / reviewers in the wild / expert
Motoki Omura
dblp:371/5840
· DBLP profile ↗
2ranked-venue papers
2as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Reinforcement learning · 81% Optimization for machine learning · 19% |
Topics — the 6 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning
value function estimation |
1.0 | 2 | 2025 | Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning · AAAI 2024 Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning · ICML 2025 |
Machine learning › Reinforcement learning
actor-critic methods |
0.9 | 1 | 2025 | Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning · ICML 2025 |
Machine learning › Optimization for machine learning › stochastic search
annealing |
0.9 | 1 | 2025 | Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning · ICML 2025 |
Machine learning › Reinforcement learning › dynamic programming
bellman operator |
0.9 | 1 | 2025 | Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning · ICML 2025 |
Machine learning › Reinforcement learning › value-based reinforcement learning
q-learning |
0.8 | 1 | 2024 | Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement Learning · AAAI 2024 |
Machine learning › Reinforcement learning › value function estimation
overestimation bias |
0.3 | 1 | 2025 | Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement Learning · ICML 2025 |
Methods — techniques the papers use, named apart from their topics
annealing · 0.9TD3 · 0.9SAC · 0.9synthetic noise · 0.8symmetric q-learning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gradual Transition from Bellman Optimality Operator to Bellman Operator in Online Reinforcement LearningabstractFor continuous action spaces, actor-critic methods are widely used in online reinforcement learning (RL). However, unlike RL algorithms for discrete actions, which generally model the optimal value function using the Bellman optimality operator, RL algorithms for continuous actions typically model Q-values for the current policy using the Bellman operator. These algorithms for continuous actions rely exclusively on policy updates for improvement, which often results in low sample efficiency. This study examines the effectiveness of incorporating the Bellman optimality operator into actor-critic frameworks. Experiments in a simple environment show that modeling optimal values accelerates learning but leads to overestimation bias. To address this, we propose an annealing approach that gradually transitions from the Bellman optimality operator to the Bellman operator, thereby accelerating learning while mitigating bias. Our method, combined with TD3 and SAC, significantly outperforms existing approaches across various locomotion and manipulation tasks, demonstrating improved performance and robustness to hyperparameters related to optimality. The code for this study is available at https://github.com/motokiomura/annealed-q-learning. Motoki Omura, Kazuki Ota, Takayuki Osa, Yusuke Mukuta, Tatsuya Harada |
ICML | 1 |
| 2024 | Symmetric Q-learning: Reducing Skewness of Bellman Error in Online Reinforcement LearningabstractIn deep reinforcement learning, estimating the value function to evaluate the quality of states and actions is essential. The value function is often trained using the least squares method, which implicitly assumes a Gaussian error distribution. However, a recent study suggested that the error distribution for training the value function is often skewed because of the properties of the Bellman operator, and violates the implicit assumption of normal error distribution in the least squares method. To address this, we proposed a method called Symmetric Q-learning, in which the synthetic noise generated from a zero-mean distribution is added to the target values to generate a Gaussian error distribution. We evaluated the proposed method on continuous control benchmark tasks in MuJoCo. It improved the sample efficiency of a state-of-the-art reinforcement learning method by reducing the skewness of the error distribution. Motoki Omura, Takayuki Osa, Yusuke Mukuta, Tatsuya Harada |
AAAI | 1 |