Toshinori Kitamura

dblp:55/9917 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Reinforcement learning · 68% Learning theory · 15% Robot manipulation · 9%
Human-computer interaction and pervasive computing
1 paper
Human-robot interaction · 100%

Topics — the 12 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
reinforcement learning theory
1.522025
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation · NeurIPS 2025
Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice · ICML 2023
Machine learning › Reinforcement learning › markov decision process
constrained markov decision process
0.912025
Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form · ICLR 2025
Machine learning › Reinforcement learning
constrained reinforcement learning
0.912025
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation · NeurIPS 2025
Machine learning › Reinforcement learning
policy optimization
0.912025
Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form · ICLR 2025
Machine learning › Learning theory › online learning
regret bounds
0.912025
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation · NeurIPS 2025
Machine learning › Reinforcement learning
robust reinforcement learning
0.912025
Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form · ICLR 2025
Machine learning › Reinforcement learning › markov decision process › low-rank MDP
linear MDP
0.712023
Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice · ICML 2023
Machine learning › Learning theory
sample complexity
0.712023
Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice · ICML 2023
Machine learning › Reinforcement learning
value-based reinforcement learning
0.712023
Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice · ICML 2023
Machine learning › Reinforcement learning › function approximation
linear function approximation
0.312025
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation · NeurIPS 2025
Machine learning › Reinforcement learning
reinforcement learning with function approximation
0.312025
Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation · NeurIPS 2025
Human-robot interaction
safe human-robot interaction
0.312025
A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics · IJCAI 2025

Methods — techniques the papers use, named apart from their topics

policy gradient · 0.9optimism-based exploration · 0.9linear MDP · 0.9foundation models · 0.9foundation model · 0.9epigraph form · 0.9bisection search · 0.9variance weighting · 0.7mirror descent · 0.7least squares regression · 0.7
YearPublicationVenuePosition
2025 Near-Optimal Policy Identification in Robust Constrained Markov Decision Processes via Epigraph Form
abstract
Designing a safe policy for uncertain environments is crucial in real-world control systems. However, this challenge remains inadequately addressed within the Markov decision process (MDP) framework. This paper presents the first algorithm guaranteed to identify a near-optimal policy in a robust constrained MDP (RCMDP), where an optimal policy minimizes cumulative cost while satisfying constraints in the worst-case scenario across a set of environments. We first prove that the conventional policy gradient approach to the Lagrangian max-min formulation can become trapped in suboptimal solutions. This occurs when its inner minimization encounters a sum of conflicting gradients from the objective and constraint functions. To address this, we leverage the epigraph form of the RCMDP problem, which resolves the conflict by selecting a single gradient from either the objective or the constraints. Building on the epigraph form, we propose a bisection search algorithm with a policy gradient subroutine and prove that it identifies an $\varepsilon$-optimal policy in an RCMDP with $\widetilde{\mathcal{O}}(\varepsilon^{-4})$ robust policy evaluations.
Toshinori Kitamura, Tadashi Kozuno, Wataru Kumagai, Kenta Hoshino, Yohei Hosoe, Kazumi Kasaura, Masashi Hamaya, Paavo Parmas, Yutaka Matsuo
ICLR1
2025 A Comprehensive Survey on Physical Risk Control in the Era of Foundation Model-enabled Robotics
abstract
Recent Foundation Model-enabled robotics (FMRs) display greatly improved general-purpose skills, enabling more adaptable automation than conventional robotics. Their ability to handle diverse tasks thus creates new opportunities to replace human labor. However, unlike general foundation models, FMRs interact with the physical world, where their actions directly affect the safety of humans and surrounding objects, requiring careful deployment and control. Based on this proposition, our survey comprehensively summarizes robot control approaches to mitigate physical risks by covering all the lifespan of FMRs ranging from pre-deployment to post-accident stage. Specifically, we broadly divide the timeline into the following three phases: (1) pre-deployment phase, (2) pre-incident phase, and (3) post-incident phase. Throughout this survey, we find that there is much room to study (i) pre-incident risk mitigation strategies, (ii) research that assumes physical interaction with humans, and (iii) essential issues of foundation models themselves. We hope that this survey will be a milestone in providing a high-resolution analysis of the physical risks of FMRs and their control, contributing to the realization of a good human-robot relationship.
Takeshi Kojima, Yaonan Zhu, Yusuke Iwasawa, Toshinori Kitamura, Gang Yan 0003, Shu Morikuni, Ryosuke Takanami, Alfredo Solano, Tatsuya Matsushima, Akiko Murakami, Yutaka Matsuo
IJCAI4
2025 Provably Efficient RL under Episode-Wise Safety in Constrained MDPs with Linear Function Approximation
abstract
We study the reinforcement learning (RL) problem in a constrained Markov decision process (CMDP), where an agent explores the environment to maximize the expected cumulative reward while satisfying a single constraint on the expected total utility value in every episode. While this problem is well understood in the tabular setting, theoretical results for function approximation remain scarce. This paper closes the gap by proposing an RL algorithm for linear CMDPs that achieves $\widetilde{\mathcal{O}}(\sqrt{K})$ regret with an episode-wise zero-violation guarantee. Furthermore, our method is computationally efficient, scaling polynomially with problem-dependent parameters while remaining independent of the state space size. Our results significantly improve upon recent linear CMDP algorithms, which either violate the constraint or incur exponential computational costs.
Toshinori Kitamura, Arnob Ghosh, Tadashi Kozuno, Wataru Kumagai, Kazumi Kasaura, Kenta Hoshino, Yohei Hosoe, Yutaka Matsuo
NeurIPS1
2023 Regularization and Variance-Weighted Regression Achieves Minimax Optimality in Linear MDPs: Theory and Practice
abstract
Mirror descent value iteration (MDVI), an abstraction of Kullback-Leibler (KL) and entropy-regularized reinforcement learning (RL), has served as the basis for recent high-performing practical RL algorithms. However, despite the use of function approximation in practice, the theoretical understanding of MDVI has been limited to tabular Markov decision processes (MDPs). We study MDVI with linear function approximation through its sample complexity required to identify an $\varepsilon$-optimal policy with probability $1-\delta$ under the settings of an infinite-horizon linear MDP, generative model, and G-optimal design. We demonstrate that least-squares regression weighted by the variance of an estimated optimal value function of the next state is crucial to achieving minimax optimality. Based on this observation, we present Variance-Weighted Least-Squares MDVI (VWLS-MDVI), the first theoretical algorithm that achieves nearly minimax optimal sample complexity for infinite-horizon linear MDPs. Furthermore, we propose a practical VWLS algorithm for value-based deep RL, Deep Variance Weighting (DVW). Our experiments demonstrate that DVW improves the performance of popular value-based deep RL algorithms on a set of MinAtar benchmarks.
Toshinori Kitamura, Tadashi Kozuno, Yunhao Tang, Nino Vieillard, Michal Valko, Jincheng Mei, Pierre Ménard, Mohammad Gheshlaghi Azar, Rémi Munos, Olivier Pietquin, Matthieu Geist, Csaba Szepesvári, Wataru Kumagai, Yutaka Matsuo
ICML1
2021 Geometric Value Iteration: Dynamic Error-Aware KL Regularization for Reinforcement Learning
abstract
The recent boom in the literature on entropy-regularized reinforcement learning (RL) approaches reveals that Kullback-Leibler (KL) regularization brings advantages to RL algorithms by canceling out errors under mild assumptions. However, existing analyses focus on fixed regularization with a constant weighting coefficient and do not consider cases where the coefficient is allowed to change dynamically. In this paper, we study the dynamic coefficient scheme and present the first asymptotic error bound. Based on the dynamic coefficient error bound, we propose an effective scheme to tune the coefficient according to the magnitude of error in favor of more robust learning. Complementing this development, we propose a novel algorithm, Geometric Value Iteration (GVI), that features a dynamic error-aware KL coefficient design with the aim of mitigating the impact of errors on performance. Our experiments demonstrate that GVI can effectively exploit the trade-off between learning speed and robustness over uniform averaging of a constant KL coefficient. The combination of GVI and deep networks shows stable learning behavior even in the absence of a target network, where algorithms with a constant KL coefficient would greatly oscillate or even fail to converge.
Toshinori Kitamura, Lingwei Zhu, Takamitsu Matsubara
ACML1
2021 Cautious Actor-Critic
abstract
The oscillating performance of off-policy learning and persisting errors in the actor-critic(AC) setting call for algorithms that can conservatively learn to suit the stability-critical applications better. In this paper, we propose a novel off-policy AC algorithm cautious actor-critic (CAC). The name cautious comes from the doubly conservative nature that we exploit the classic policy interpolation from conservative policy iteration for the actor and the entropy-regularization of conservative value iteration for the critic. Our key observation is the entropy-regularized critic facilitates and simplifies the unwieldy interpolated actor update while still ensuring robust policy improvement. We compare CAC to state-of-the-art AC methods on a set of challenging continuous control problems and demonstrate thatCAC achieves comparable performance while significantly stabilizes learning.
Lingwei Zhu, Toshinori Kitamura, Takamitsu Matsubara
ACML2