Kazumune Hashimoto

dblp:166/3737 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-9376-5760ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Reinforcement learning · 79% Trustworthy machine learning · 10% Motion planning and robot control · 10%

Topics — the 9 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning
safe reinforcement learning
1.522024
Flipping-based Policy for Chance-Constrained Markov Decision Processes · NeurIPS 2024
Long-Term Safe Reinforcement Learning with Binary Feedback · AAAI 2024
Machine learning › Reinforcement learning › markov decision process
constrained markov decision process
0.812024
Long-Term Safe Reinforcement Learning with Binary Feedback · AAAI 2024
Machine learning › Reinforcement learning › safe reinforcement learning
constrained policy optimization
0.812024
Flipping-based Policy for Chance-Constrained Markov Decision Processes · NeurIPS 2024
Machine learning › Reinforcement learning
constrained reinforcement learning
0.712023
Safe Exploration in Reinforcement Learning: A Generalized Formulation and Algorithms · NeurIPS 2023
Machine learning › Reinforcement learning › safe reinforcement learning
safe exploration
0.712023
Safe Exploration in Reinforcement Learning: A Generalized Formulation and Algorithms · NeurIPS 2023
Robotics › Motion planning and robot control
safety guarantees
0.712023
Safe Exploration in Reinforcement Learning: A Generalized Formulation and Algorithms · NeurIPS 2023
Machine learning › Trustworthy machine learning
uncertainty estimation
0.712023
Safe Exploration in Reinforcement Learning: A Generalized Formulation and Algorithms · NeurIPS 2023
Machine learning › Reinforcement learning
bellman equation
0.212024
Flipping-based Policy for Chance-Constrained Markov Decision Processes · NeurIPS 2024
Machine learning › Reinforcement learning
markov decision process
0.212024
Flipping-based Policy for Chance-Constrained Markov Decision Processes · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

generalized linear model · 1.4flipping-based policy · 0.8constrained policy optimization · 0.8LoBiSaRL · 0.8gaussian process · 0.7deep reinforcement learning · 0.7
YearPublicationVenuePosition
2026 Signal temporal logic-based neural network for driving task generation in Advanced Driver Assistance Systems
Kazumune Hashimoto, Yusuke Yokokawa, Norika Arai, Xun Shen, Xingguo Zhang, Pongsathorn Raksincharoensak
Eng. Appl. Artif. Intell.1
2025 Learning-Based Event-Triggered MPC With Gaussian Processes Under Terminal Constraints
abstract
The event-triggered control strategy is capable of significantly reducing the number of control task executions while achieving desired control objectives, such as stability. In this article, we introduce a novel learning-based method for event-triggered model predictive control with initially unknown dynamics. The formulation of optimal control problems (OCPs) is based on predictive states derived from Gaussian process (GP) regression under terminal constraints. The event-triggered condition proposed in this article is derived from the recursive feasibility, so that the OCPs are solved only when an error between the predictive and the actual states exceeds a certain threshold. This article analyzes the convergence of the closed-loop system under the event-triggered condition, demonstrating that the system's state will enter the terminal set within a finite time, assuming small-enough uncertainty in the GP model. We validate this approach through a tracking control problem, illustrating its practical effectiveness.
Kazumune Hashimoto, Yuga Onoue, Akifumi Wachi, Xun Shen
IEEE Trans. Cybern.1
2025 Sample-Based Continuous Approximate Method for Constructing Interval Neural Network
abstract
In safety-critical engineering applications, such as robust prediction against adversarial noise, it is necessary to quantify neural networks' uncertainty. Interval neural networks (INNs) are effective models for uncertainty quantification, giving an interval of predictions instead of a single value for a corresponding input. This article formulates the problem of training an INN as a chance-constrained optimization problem. The optimal solution of the formulated chance-constrained optimization naturally forms an INN that gives the tightest interval of predictions with a required confidence level. Since the chance-constrained optimization problem is intractable, a sample-based continuous approximate method is used to obtain approximate solutions to the chance-constrained optimization problem. We prove the uniform convergence of the approximation, showing that it gives the optimal INN consistently with the original ones. Additionally, we investigate the reliability of the approximation with finite samples, giving the probability bound for violation with finite samples. Through a numerical example and an application case study of anomaly detection in wind power data, we evaluate the effectiveness of the proposed INN against existing approaches, including Bayesian neural networks, highlighting its capability to significantly improve the performance of applying INNs for regression and unsupervised anomaly detection.
Xun Shen, Tinghui Ouyang, Kazumune Hashimoto, Yuhu Wu
IEEE Trans. Neural Networks Learn. Syst.3
2024 Long-Term Safe Reinforcement Learning with Binary Feedback
abstract
Safety is an indispensable requirement for applying reinforcement learning (RL) to real problems. Although there has been a surge of safe RL algorithms proposed in recent years, most existing work typically 1) relies on receiving numeric safety feedback; 2) does not guarantee safety during the learning process; 3) limits the problem to a priori known, deterministic transition dynamics; and/or 4) assume the existence of a known safe policy for any states. Addressing the issues mentioned above, we thus propose Long-term Binary-feedback Safe RL (LoBiSaRL), a safe RL algorithm for constrained Markov decision processes (CMDPs) with binary safety feedback and an unknown, stochastic state transition function. LoBiSaRL optimizes a policy to maximize rewards while guaranteeing long-term safety that an agent executes only safe state-action pairs throughout each episode with high probability. Specifically, LoBiSaRL models the binary safety function via a generalized linear model (GLM) and conservatively takes only a safe action at every time step while inferring its effect on future safety under proper assumptions. Our theoretical results show that LoBiSaRL guarantees the long-term safety constraint, with high probability. Finally, our empirical results demonstrate that our algorithm is safer than existing methods without significantly compromising performance in terms of reward.
Akifumi Wachi, Wataru Hashimoto 0001, Kazumune Hashimoto
AAAI3
2024 Flipping-based Policy for Chance-Constrained Markov Decision Processes
abstract
Safe reinforcement learning (RL) is a promising approach for many real-world decision-making problems where ensuring safety is a critical necessity. In safe RL research, while expected cumulative safety constraints (ECSCs) are typically the first choices, chance constraints are often more pragmatic for incorporating safety under uncertainties. This paper proposes a \textit{flipping-based policy} for Chance-Constrained Markov Decision Processes (CCMDPs). The flipping-based policy selects the next action by tossing a potentially distorted coin between two action candidates. The probability of the flip and the two action candidates vary depending on the state. We establish a Bellman equation for CCMDPs and further prove the existence of a flipping-based policy within the optimal solution sets. Since solving the problem with joint chance constraints is challenging in practice, we then prove that joint chance constraints can be approximated into Expected Cumulative Safety Constraints (ECSCs) and that there exists a flipping-based policy in the optimal solution sets for constrained MDPs with ECSCs. As a specific instance of practical implementations, we present a framework for adapting constrained policy optimization to train a flipping-based policy. This framework can be applied to other safe RL algorithms. We demonstrate that the flipping-based policy can improve the performance of the existing safe RL algorithms under the same limits of safety constraints on Safety Gym benchmarks.
Xun Shen, Akifumi Wachi, Kazumune Hashimoto, Sebastien Gros
NeurIPS4
2024 A Robust Traffic Flow Control Using Connected Vehicle Technology: Signal Spatio-Temporal Logic-Based Approach
abstract
This study examines traffic signal optimization at an intersection by means of formal methods for mobility and safety of vehicular and pedestrian movements. In view of this, the long-established pre-timed and actuated control methods for traffic signals are inadequate to make extensive use of the comprehensive traffic states of conventional and connected entities provided by installed sensors and vehicular ad hoc network. To compensate for this shortcoming, signal spatio-temporal logic (SSTL) specifications are proposed as functions of traffic states to express the mobility and safety objectives to be achieved via model predictive control (MPC). MPC allows for a robust optimization of traffic signal timing against all realizations of additive bounded disturbances, meaning random incoming traffic, such that an SSTL specification is satisfied, if feasible; otherwise, it is violated as little as possible.
Sagar V. Patil, Kazumune Hashimoto, Masako Kishida
IEEE Trans. Intell. Transp. Syst.2
2023 Safe Exploration in Reinforcement Learning: A Generalized Formulation and Algorithms
abstract
Safe exploration is essential for the practical use of reinforcement learning (RL) in many real-world scenarios. In this paper, we present a generalized safe exploration (GSE) problem as a unified formulation of common safe exploration problems. We then propose a solution of the GSE problem in the form of a meta-algorithm for safe exploration, MASE, which combines an unconstrained RL algorithm with an uncertainty quantifier to guarantee safety in the current episode while properly penalizing unsafe explorations before actual safety violation to discourage them in future episodes. The advantage of MASE is that we can optimize a policy while guaranteeing with a high probability that no safety constraint will be violated under proper assumptions. Specifically, we present two variants of MASE with different constructions of the uncertainty quantifier: one based on generalized linear models with theoretical guarantees of safety and near-optimality, and another that combines a Gaussian process to ensure safety with a deep RL algorithm to maximize the reward. Finally, we demonstrate that our proposed algorithm achieves better performance than state-of-the-art algorithms on grid-world and Safety Gym benchmarks without violating any safety constraints, even during training.
Akifumi Wachi, Wataru Hashimoto 0001, Xun Shen, Kazumune Hashimoto
NeurIPS4
2022 Collaborative Rover-copter Path Planning and Exploration with Temporal Logic Specifications Based on Bayesian Update Under Uncertain Environments
abstract
This article investigates a collaborative rover-copter path planning and exploration with temporal logic specifications under uncertain environments. The objective of the rover is to complete a mission expressed by a syntactically co-safe linear temporal logic (scLTL) formula, while the objective of the copter is to actively explore the environment and reduce its uncertainties, aiming at assisting the rover and enhancing the efficiency of the mission completion. To formalize our approach, we first capture the environmental uncertainties by environmental beliefs of the atomic propositions, under an assumption that it is unknown which properties (or, atomic propositions) are satisfied in each area of the environment. The environmental beliefs of the atomic propositions are updated according to the Bayes rule based on the Bernoulli-type sensor measurements provided by both the rover and the copter. Then, the optimal policy for the rover is synthesized by maximizing a belief of the satisfaction of the scLTL formula through an implementation of an automata-based model checking. An exploration policy for the copter is then synthesized by employing the notion of an entropy that is evaluated based on the environmental beliefs of the atomic propositions, and a path that the rover intends to follow according to the optimal policy. As such, the copter can actively explore regions whose uncertainties are high and that are relevant to the mission completion. Finally, some numerical examples illustrate the effectiveness of the proposed approach.
Kazumune Hashimoto, Natsuko Tsumagari, Toshimitsu Ushio
ACM Trans. Cyber Phys. Syst.1
2021 Learning Self-Triggered Controllers With Gaussian Processes
abstract
This article investigates the design of self-triggered controllers for networked control systems (NCSs), where the dynamics of the plant are unknown a priori. To deal with the unknown transition dynamics, we employ the Gaussian process (GP) regression in order to learn the dynamics of the plant. To design the self-triggered controller, we formulate an optimal control problem, such that the optimal control and communication policies can be jointly designed based on the GP model of the plant. Moreover, we provide an overall implementation algorithm that jointly learns the dynamics of the plant and the self-triggered controller based on a reinforcement learning framework. Finally, a numerical simulation illustrates the effectiveness of the proposed approach.
Kazumune Hashimoto, Yuichi Yoshimura, Toshimitsu Ushio
IEEE Trans. Cybern.1