VLDB 2026 Research / reviewers in the wild / expert
Masanari Kimura
dblp:215/4359
· DBLP profile ↗
11ranked-venue papers
10as first author
10since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 9 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Density Ratio Estimation via Sampling along Generalized Geodesics on Statistical ManifoldsabstractThe density ratio of two probability distributions is one of the fundamental tools in mathematical and computational statistics and machine learning, and it has a variety of known applications. Therefore, density ratio estimation from finite samples is a very important task, but it is known to be unstable when the distributions are distant from each other. One approach to address this problem is density ratio estimation using incremental mixtures of the two distributions. We geometrically reinterpret existing methods for density ratio estimation based on incremental mixtures. We show that these methods can be regarded as iterating on the Riemannian manifold along a particular curve between the two probability distributions. Making use of the geometry of the manifold, we propose to consider incremental density ratio estimation along generalized geodesics on this manifold. To achieve such a method requires Monte Carlo sampling along geodesics via transformations of the two distributions. We show how to implement an iterative algorithm to sample along these geodesics and show how changing the distances along the geodesic affect the variance and accuracy of the estimation of the density ratio. Masanari Kimura, Howard D. Bondell |
AISTATS | 1 |
| 2025 | Difference-of-submodular Bregman DivergenceabstractThe Bregman divergence, which is generated from a convex function, is commonly used as a pseudo-distance for comparing vectors or functions in continuous spaces. In contrast, defining an analog of the Bregman divergence for discrete spaces is nontrivial. Iyer & Bilmes (2012b) considered Bregman divergences on discrete domains using submodular functions as generating functions, the discrete analogs of convex functions. In this paper, we further generalize this framework to cases where the generating function is neither submodular nor supermodular, thus increasing the flexibility and representational capacity of the resulting divergence, which we term the difference-of-submodular Bregman divergence. Additionally, we introduce a learnable form of this divergence using permutation-invariant neural networks (NNs) and demonstrate through experiments that it effectively captures key structural properties in discrete data. As a result, the proposed method significantly improves the performance of existing methods on tasks such as clustering and set retrieval problems. This work addresses the challenge of defining meaningful divergences in discrete settings and provides a new tool for tasks requiring structure-preserving distance measures. Masanari Kimura, Takahiro Kawashima, Tasuku Soma, Hideitsu Hino |
ICLR | 1 |
| 2025 | Explaining Black-box Model Predictions via Two-level Nested Feature Attributions with Consistency PropertyabstractTechniques that explain the predictions of black-box machine learning models are crucial to make the models transparent, thereby increasing trust in AI systems. The input features to the models often have a nested structure that consists of high- and low-level features, and each high-level feature is decomposed into multiple low-level features. For such inputs, both high-level feature attributions (HiFAs) and low-level feature attributions (LoFAs) are important for better understanding the model's decision. In this paper, we propose a model-agnostic local explanation method that effectively exploits the nested structure of the input to estimate the two-level feature attributions simultaneously. A key idea of the proposed method is to introduce the consistency property that should exist between the HiFAs and LoFAs, thereby bridging the separate optimization problems for estimating them. Thanks to this consistency property, the proposed method can produce HiFAs and LoFAs that are both faithful to the black-box models and consistent with each other, using a smaller number of queries to the models. In experiments on image classification in multiple instance learning and text classification using language models, we demonstrate that the HiFAs and LoFAs estimated by the proposed method are accurate, faithful to the behaviors of the black-box models, and provide consistent explanations. Yuya Yoshikawa, Masanari Kimura, Ryotaro Shimizu, Yuki Saito 0002 |
IJCAI | 2 |
| 2025 | Graph-Smoothed Bayesian Black-Box Shift Estimator and Its Information GeometryabstractLabel shift adaptation aims to recover target class priors when the labelled source distribution $P$ and the unlabelled target distribution $Q$ share $P(X \mid Y) = Q(X \mid Y)$ but $P(Y) \neq Q(Y)$. Classical black‑box shift estimators invert an empirical confusion matrix of a frozen classifier, producing a brittle point estimate that ignores sampling noise and similarity among classes. We present Graph‑Smoothed Bayesian BBSE (GS‑B$^3$SE), a fully probabilistic alternative that places Laplacian–Gaussian priors on both target log‑priors and confusion‑matrix columns, tying them together on a label‑similarity graph. The resulting posterior is tractable with HMC or a fast block Newton–CG scheme. We prove identifiability, $N^{-1/2}$ contraction, variance bounds that shrink with the graph’s algebraic connectivity, and robustness to Laplacian misspecification. We also reinterpret GS‑B$^3$SE through information geometry, showing that it generalizes existing shift estimators. Masanari Kimura |
NeurIPS | 1 |
| 2025 | Geometric insights into focal loss: Reducing curvature for enhanced model calibrationabstractThe key factor in implementing machine learning algorithms in decision-making situations is not only the accuracy of the model but also its confidence level. The confidence level of a model in a classification problem is often given by the output vector of a softmax function for convenience. However, these values are known to deviate significantly from the actual expected model confidence. This problem is called model calibration and has been studied extensively. One of the simplest techniques to tackle this task is focal loss, a generalization of cross-entropy by introducing one positive parameter. Although many related studies exist because of the simplicity of the idea and its formalization, the theoretical analysis of its behavior is still insufficient. In this study, our objective is to understand the behavior of focal loss by reinterpreting this function geometrically. Our analysis suggests that focal loss reduces the curvature of the loss surface in training the model. This indicates that curvature may be one of the essential factors in achieving model calibration. We design numerical experiments to support this conjecture to reveal the behavior of focal loss and the relationship between calibration performance and curvature. • We reinterpret focal loss and show that it behaves as the curvature reduction. • We provide the conjecture of the connection between curvature and model calibration. • We design the numerical experiments to demonstrate our theoretical hypothesis. Masanari Kimura, Hiroki Naganuma |
Pattern Recognit. Lett. | 1 |
| 2022 | Generalization Bounds for Set-to-Set Matching with Negative Sampling
Masanari Kimura |
ICONIP (4) | 1 |
| 2022 | Information Geometrically Generalized Covariate Shift AdaptationabstractMany machine learning methods assume that the training and test data follow the same distribution. However, in the real world, this assumption is often violated. In particular, the marginal distribution of the data changes, called covariate shift, is one of the most important research topics in machine learning. We show that the well-known family of covariate shift adaptation methods is unified in the framework of information geometry. Furthermore, we show that parameter search for a geometrically generalized covariate shift adaptation method can be achieved efficiently. Numerical experiments show that our generalization can achieve better performance than the existing methods it encompasses. Masanari Kimura, Hideitsu Hino |
Neural Comput. | 1 |
| 2021 | Why Mixup Improves the Model Performance
Masanari Kimura |
ICANN (2) | 1 |
| 2021 | Understanding Test-Time Augmentation
Masanari Kimura |
ICONIP (1) | 1 |
| 2021 | Density-Fixing: Simple yet Effective Regularization Method based on the Class PriorsabstractMachine learning models suffer from overfitting, which is caused by a lack of labeled data. We proposed a framework of regularization methods, called density-fixing, that can be used commonly for supervised and semi-supervised learning to tackle this problem. Our proposed regularization method improves the generalization performance by forcing the model to approximate the class prior distribution or occurrence frequency. This regularization term is naturally derived from the formula of maximum likelihood estimation and is theoretically justified. We further investigated the asymptotic behavior of the proposed method and how the regularization term behaves when assuming a prior distribution of several classes in practice. We provide several theoretical analyses of the proposed method including asymptotic behavior. Our experimental results on multiple benchmark datasets are sufficient to support our argument, and we suggest that this simple and effective regularization method is useful in real-world machine learning problems. Masanari Kimura, Ryohei Izawa |
IJCNN | 1 |
| 2020 | Batch Prioritization of Data Labeling Tasks for Training ClassifiersabstractIn a data labeling process for building machine learning, the choice of labeling data instances is known to have a significant impact on the performance of classifiers. So far, the study of active learning has addressed the issue of how to choose the subset by prioritizing the data instances based on the state of the current classifier. However, the active learning approach has two drawbacks that (i) require a training loop to update the priorities of labeling tasks and (ii) require us to choose a specific active learner while we do not know the optimal classification model. In this paper, we propose a new framework of priority-aware labeling system that allows a parallel task assignment to crowd workers without assuming a particular classifier, which is based on novel methods called “batch prioritization” and “label expansion”. We conducted experiments with multiple datasets to examine the effectiveness of the approach and found that the proposed method improves the performance of the final classifiers more quickly than the active learning approach despite that the labeling tasks can be processed in a fully parallel manner. Masanari Kimura, Kei Wakabayashi, Atsuyuki Morishima |
HCOMP | 1 |