VLDB 2026 Research / reviewers in the wild / expert
Masahiro Kato
dblp:83/5680
· DBLP profile ↗
17ranked-venue papers
9as first author
11since 2021 · last 2025
0000-0003-4050-023XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 9 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LC-Tsallis-INF: Generalized Best-of-Both-Worlds Linear Contextual BanditsabstractWe investigate the \emph{linear contextual bandit problem} with independent and identically distributed (i.i.d.) contexts. In this problem, we aim to develop a \emph{Best-of-Both-Worlds} (BoBW) algorithm with regret upper bounds in both stochastic and adversarial regimes. We develop an algorithm based on \emph{Follow-The-Regularized-Leader} (FTRL) with Tsallis entropy, referred to as the $\alpha$-\emph{Linear-Contextual (LC)-Tsallis-INF}. We show that its regret is at most $O(\log(T))$ in the stochastic regime under the assumption that the suboptimality gap is uniformly bounded from below, and at most $O(\sqrt{T})$ in the adversarial regime. Furthermore, our regret analysis is extended to more general regimes characterized by the \emph{margin condition} with a parameter $\beta \in (1, \infty]$, which imposes a milder assumption on the suboptimality gap than in previous studies. We show that the proposed algorithm achieves $O\left(\log(T)^{\frac{1+\beta}{2+\beta}}T^{\frac{1}{2+\beta}}\right)$ regret under the margin condition. Masahiro Kato, Shinji Ito |
AISTATS | 1 |
| 2025 | PUATE: Efficient ATE Estimation from Treated (Positive) and Unlabeled UnitsabstractThe estimation of average treatment effects (ATEs), defined as the difference in expected outcomes between treatment and control groups, is a central topic in causal inference. This study develops semiparametric efficient estimators for ATE in a setting where only a treatment group and an unlabeled group—consisting of units whose treatment status is unknown—are observed. This scenario constitutes a variant of learning from positive and unlabeled data (PU learning) and can be viewed as a special case of ATE estimation with missing data. For this setting, we derive the semiparametric efficiency bounds, which characterize the lowest achievable asymptotic variance for regular estimators. We then construct semiparametric efficient ATE estimators that attain these bounds. Our results contribute to the literature on causal inference with missing data and weakly supervised learning. Masahiro Kato, Fumiaki Kozai, Ryo Inokuchi |
NeurIPS | 1 |
| 2024 | Active Adaptive Experimental Design for Treatment Effect Estimation with Covariate ChoiceabstractThis study designs an adaptive experiment for efficiently estimating *average treatment effects* (ATEs). In each round of our adaptive experiment, an experimenter sequentially samples an experimental unit, assigns a treatment, and observes the corresponding outcome immediately. At the end of the experiment, the experimenter estimates an ATE using the gathered samples. The objective is to estimate the ATE with a smaller asymptotic variance. Existing studies have designed experiments that adaptively optimize the propensity score (treatment-assignment probability). As a generalization of such an approach, we propose optimizing the covariate density as well as the propensity score. First, we derive the efficient covariate density and propensity score that minimize the semiparametric efficiency bound and find that optimizing both covariate density and propensity score minimizes the semiparametric efficiency bound more effectively than optimizing only the propensity score. Next, we design an adaptive experiment using the efficient covariate density and propensity score sequentially estimated during the experiment. Lastly, we propose an ATE estimator whose asymptotic variance aligns with the minimized semiparametric efficiency bound. Masahiro Kato, Akihiro Oga, Wataru Komatsubara, Ryo Inokuchi |
ICML | 1 |
| 2024 | Reconsidering Stochastic Policy Gradient Methods for Traffic Signal Control
Masahiro Kato, Ryosuke Kojima |
IEA/AIE | 1 |
| 2023 | Unified Perspective on Probability Divergence via the Density-Ratio Likelihood: Bridging KL-Divergence and Integral Probability MetricsabstractThis paper provides a unified perspective for the Kullback-Leibler (KL)-divergence and the integral probability metrics (IPMs) from the perspective of maximum likelihood density-ratio estimation (DRE). Both the KL-divergence and the IPMs are widely used in various fields in applications such as generative modeling. However, a unified understanding of these concepts has still been unexplored. In this paper, we show that the KL-divergence and the IPMs can be represented as maximal likelihoods differing only by sampling schemes, and use this result to derive a unified form of the IPMs and a relaxed estimation method. To develop the estimation problem, we construct an unconstrained maximum likelihood estimator to perform DRE with a stratified sampling scheme. We further propose a novel class of probability divergences, called the Density Ratio Metrics (DRMs), that interpolates the KL-divergence and the IPMs. In addition to these findings, we also introduce some applications of the DRMs, such as DRE and generative adversarial networks. In experiments, we validate the effectiveness of our proposed methods. Masahiro Kato, Masaaki Imaizumi, Kentaro Minami |
AISTATS | 1 |
| 2022 | Learning Causal Models from Conditional Moment Restrictions by Importance Weighting
Masahiro Kato, Masaaki Imaizumi, Kenichiro McAlinn, Shota Yasui, Haruo Kakehi |
ICLR | 1 |
| 2022 | Learning Classifiers under Delayed Feedback with a Time Window AssumptionabstractWe consider training a binary classifier under delayed feedback (DF learning). For example, in the conversion prediction in online ads, we initially receive negative samples that clicked the ads but did not buy an item; subsequently, some samples among them buy an item then change to positive. In the setting of DF learning, we observe samples over time, then learn a classifier at some point. We initially receive negative samples; subsequently, some samples among them change to positive. This problem is conceivable in various real-world applications such as online advertisements, where the user action takes place long after the first click. Owing to the delayed feedback, naive classification of the positive and negative samples returns a biased classifier. One solution is to use samples that have been observed for more than a certain time window assuming these samples are correctly labeled. However, existing studies reported that simply using a subset of all samples based on the time window assumption does not perform well, and that using all samples along with the time window assumption improves empirical performance. We extend these existing studies and propose a method with the unbiased and convex empirical risk that is constructed from all samples under the time window assumption. To demonstrate the soundness of the proposed method, we provide experimental results on a synthetic and open dataset that is the real traffic log datasets in online advertising. Shota Yasui, Masahiro Kato |
KDD | 2 |
| 2021 | Non-Negative Bregman Divergence Minimization for Deep Direct Density Ratio EstimationabstractDensity ratio estimation (DRE) is at the core of various machine learning tasks such as anomaly detection and domain adaptation. In the DRE literature, existing studies have extensively studied methods based on Bregman divergence (BD) minimization. However, when we apply the BD minimization with highly flexible models, such as deep neural networks, it tends to suffer from what we call train-loss hacking, which is a source of over-fitting caused by a typical characteristic of empirical BD estimators. In this paper, to mitigate train-loss hacking, we propose non-negative correction for empirical BD estimators. Theoretically, we confirm the soundness of the proposed method through a generalization error bound. In our experiments, the proposed methods show favorable performances in inlier-based outlier detection. Masahiro Kato, Takeshi Teshima |
ICML | 1 |
| 2021 | The Adaptive Doubly Robust Estimator and a Paradox Concerning Logging PolicyabstractThe doubly robust (DR) estimator, which consists of two nuisance parameters, the conditional mean outcome and the logging policy (the probability of choosing an action), is crucial in causal inference. This paper proposes a DR estimator for dependent samples obtained from adaptive experiments. To obtain an asymptotically normal semiparametric estimator from dependent samples without non-Donsker nuisance estimators, we propose adaptive-fitting as a variant of sample-splitting. We also report an empirical paradox that our proposed DR estimator tends to show better performances compared to other estimators utilizing the true logging policy. While a similar phenomenon is known for estimators with i.i.d. samples, traditional explanations based on asymptotic efficiency cannot elucidate our case with dependent samples. We confirm this hypothesis through simulation studies. Masahiro Kato, Kenichiro McAlinn, Shota Yasui |
NeurIPS | 1 |
| 2021 | Scalable Personalised Item Ranking through Parametric Density EstimationabstractLearning from implicit feedback is challenging because of the difficult nature of the one-class problem: we can observe only positive examples. Most conventional methods use a pairwise ranking approach and negative samplers to cope with the one-class problem. However, such methods have two main drawbacks particularly in large-scale applications; (1) the pairwise approach is severely inefficient due to the quadratic computational cost; and (2) even recent model-based samplers (e.g. IRGAN) cannot achieve practical efficiency due to the training of an extra model. Riku Togashi, Masahiro Kato, Mayu Otani, Tetsuya Sakai, Shin'ichi Satoh 0001 |
SIGIR | 2 |
| 2021 | Density-Ratio Based Personalised Ranking from Implicit FeedbackabstractLearning from implicit user feedback is challenging as we can only observe positive samples but never access negative ones. Most conventional methods cope with this issue by adopting a pairwise ranking approach with negative sampling. However, the pairwise ranking approach has a severe disadvantage in the convergence time owing to the quadratically increasing computational cost with respect to the sample size; it is problematic, particularly for large-scale datasets and complex models such as neural networks. By contrast, a pointwise approach does not directly solve a ranking problem, and is therefore inferior to a pairwise counterpart in top-K ranking tasks; however, it is generally advantageous in regards to the convergence time. This study aims to establish an approach to learn personalised ranking from implicit feedback, which reconciles the training efficiency of the pointwise approach and ranking effectiveness of the pairwise counterpart. The key idea is to estimate the ranking of items in a pointwise manner; we first reformulate the conventional pointwise approach based on density ratio estimation and then incorporate the essence of ranking-oriented approaches (e.g. the pairwise approach) into our formulation. Through experiments on three real-world datasets, we demonstrate that our approach dramatically reduces the convergence time (one to two orders of magnitude faster) and significantly improves the ranking performance. Riku Togashi, Masahiro Kato, Mayu Otani, Shin'ichi Satoh 0001 |
WWW | 2 |
| 2020 | Off-Policy Evaluation and Learning for External Validity under a Covariate ShiftabstractWe consider the evaluation and training of a new policy for the evaluation data by using the historical data obtained from a different policy. The goal of off-policy evaluation (OPE) is to estimate the expected reward of a new policy over the evaluation data, and that of off-policy learning (OPL) is to find a new policy that maximizes the expected reward over the evaluation data. Although the standard OPE and OPL assume the same distribution of covariate between the historical and evaluation data, there often exists a problem of a covariate shift,i.e., the distribution of the covariate of the historical data is different from that of the evaluation data. In this paper, we derive the efficiency bound of OPE under a covariate shift. Then, we propose doubly robust and efficient estimators for OPE and OPL under a covariate shift by using an estimator of the density ratio between the distributions of the historical and evaluation data. We also discuss other possible estimators and compare their theoretical properties. Finally, we confirm the effectiveness of the proposed estimators through experiments. Masatoshi Uehara, Masahiro Kato, Shota Yasui |
NeurIPS | 2 |
| 2019 | Learning from Positive and Unlabeled Data with a Selection Bias
Masahiro Kato, Takeshi Teshima, Junya Honda |
ICLR (Poster) | 1 |
| 2018 | Teach-and-Replay of Mobile Robot with Particle Filter on EpisodeabstractA novel method for replaying behavior of a mobile robot from its memory of past experiences is presented in this paper. The method is a version of a particle filter on episode (PFoE), which applies a particle filter on the memory so as to efficiently find some similar situations with the current one. Though the original PFoE was proposed as a reinforcement learning method, we once removed the reward system from the original one so as to apply it to task teaching. In the experiment, we gave several kinds of motion to a micromouse type robot with the proposed method through a gamepad. The robot replayed the behaviors robustly with sensor feedback after several number of repetitive teaching. Ryuichi Ueda, Masahiro Kato, Atsushi Saito, Ryo Okazaki |
ICRA | 2 |
| 2008 | Local-spectrum-based distinction between handwritten and machine-printed charactersabstractIn this paper, we propose a method to distinguish between handwritten and machine-printed characters with no need to locate character or text-line positions. We transform a local region in a document image into frequency domain to extract feature values including fluctuations caused by handwriting. We feed the feature values to an optimized multilayer perceptron (MLP) to get likelihood of handwriting. We call this method the spectrum-domain local fluctuation detection (SDLFD) method. Experimental results show that our method distinguishes handwritten characters from machine-printed ones with no need of text-line position information. We also found that the scheme is robust against the change in scanning resolution. Jumpei Koyama, Akira Hirose 0001, Masahiro Kato |
ICIP | 3 |
| 2008 | Distinction between handwritten and machine-printed characters with no need to locate character or text line positionabstractIn this paper, we propose a method for distinction between handwritten and machine-printed characters with no need to locate positions of characters or text lines. We call the proposed method psilaspectrum-based local fluctuation detection method. The method transforms local regions in document images into power spectrum to extract feature values which represent fluctuations caused by handwriting. We employ a multilayer perceptron for the distinction. We feed the obtained feature values to a preliminarily optimized multilayer perceptron (MLP), and the MLP yields likelihood of handwriting. We prepare a document image which has randomly aligned characters for an experiment. The experimental result shows that our method can distinguish handwritten and machine-printed characters with no need to locate positions of characters or text lines. Jumpei Koyama, Masahiro Kato, Akira Hirose 0001 |
IJCNN | 2 |
| 2007 | Handwritten Character Distinction Method Inspired by Human Vision Mechanism
Jumpei Koyama, Masahiro Kato, Akira Hirose 0001 |
ICONIP (1) | 2 |