VLDB 2026 Research / reviewers in the wild / expert
Tomoya Sakai 0001
dblp:75/4734-1
· DBLP profile ↗
16ranked-venue papers
8as first author
6since 2021 · last 2022
0000-0003-3510-0979ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 7 first-author · 4 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | A Generalized Backward Compatibility MetricabstractRetraining a classifier with new data is inseparable from ML/AI applications, but most of the existing ML methods do not take into account the backward compatibility of predictions. That is, although the overall performance of a new classifier is improved, users will be confused by the wrong predictions of the new classifier, especially when the predictions of the old classifier are correct for the same samples. To this end, several metrics and learning methods for backward compatibility have been actively studied recently. Despite significant interest in backward compatibility, the metrics and methods are not well known from a theoretical perspective. In this paper, we first analyze the existing backward compatibility metrics and reveal that these metrics essentially assess the same quantity between old and new models. In addition, to obtain a unified view of backward compatibility metrics, we propose a generalized backward compatibility (GBC) metric that can represent the existing backward compatibility metrics. We formulate a learning objective based on the GBC metric and derive the estimation error bound, and the result is applied to one of the existing methods. Through further analysis, we reveal that the existing backward compatibility metrics are not suitable for imbalanced classification. We then design a backward compatibility metric for imbalanced classification on the basis of the GBC metric and empirically demonstrate the practicality of the proposed metric. Tomoya Sakai 0001 |
KDD | 1 |
| 2021 | Regret Minimization for Causal Inference on Large Treatment SpaceabstractPredicting which action (treatment) will lead to a better outcome is a central task in decision support systems. To build a prediction model in real situations, learning from observational data with a sampling bias is a critical issue due to the lack of randomized controlled trial (RCT) data. To handle such biased observational data, recent efforts in causal inference and counterfactual machine learning have focused on debiased estimation of the potential outcomes on a binary action space and the difference between them, namely, the individual treatment effect. When it comes to a large action space (e.g., selecting an appropriate combination of medicines for a patient), however, the regression accuracy of the potential outcomes is no longer sufficient in practical terms to achieve a good decision-making performance. This is because a high mean accuracy on the large action space does not guarantee the nonexistence of a single potential outcome misestimation that misleads the whole decision. Our proposed loss minimizes the classification error of whether or not the action is relatively good for the individual target among all feasible actions, which further improves the decision-making performance, as we demonstrate. We also propose a network architecture and a regularizer that extracts a debiased representation not only from the individual feature but also from the biased action for better generalization in large action spaces. Extensive experiments on synthetic and semi-synthetic datasets demonstrate the superiority of our method for large combinatorial action spaces. Akira Tanimoto, Tomoya Sakai 0001, Takashi Takenouchi, Hisashi Kashima |
AISTATS | 2 |
| 2021 | Causal Combinatorial Factorization Machines for Set-Wise Recommendation
Akira Tanimoto, Tomoya Sakai 0001, Takashi Takenouchi, Hisashi Kashima |
PAKDD (2) | 2 |
| 2021 | Source Hypothesis Transfer for Zero-Shot Domain Adaptation
Tomoya Sakai 0001 |
ECML/PKDD (1) | 1 |
| 2021 | Predictive Optimization with Zero-Shot Domain AdaptationabstractPrediction in a new domain without any training sample, called zero-shot domain adaptation (ZSDA), is an important task in domain adaptation.While prediction in a new domain has gained much attention in recent years, in this paper, we investigate another potential of ZSDA.Specifically, instead of predicting responses in a new domain, we find a description of a new domain given a prediction.The task is regarded as predictive optimization, but existing predictive optimization methods have not been extended to handling multiple domains.We propose a simple framework for predictive optimization with ZSDA and analyze the condition in which the optimization problem becomes convex optimization.We also discuss how to handle the interaction of characteristics of a domain in predictive optimization.Through numerical experiments, we demonstrate the potential usefulness of our proposed framework. Tomoya Sakai 0001, Naoto Ohsaka |
SDM | 1 |
| 2021 | Information-Theoretic Representation Learning for Positive-Unlabeled ClassificationabstractRecent advances in weakly supervised classification allow us to train a classifier from only positive and unlabeled (PU) data. However, existing PU classification methods typically require an accurate estimate of the class-prior probability, a critical bottleneck particularly for high-dimensional data. This problem has been commonly addressed by applying principal component analysis in advance, but such unsupervised dimension reduction can collapse the underlying class structure. In this letter, we propose a novel representation learning method from PU data based on the information-maximization principle. Our method does not require class-prior estimation and thus can be used as a preprocessing method for PU classification. Through experiments, we demonstrate that our method, combined with deep neural networks, highly improves the accuracy of PU class-prior estimation, leading to state-of-the-art PU classification performance. Tomoya Sakai 0001, Gang Niu 0001, Masashi Sugiyama |
Neural Comput. | 1 |
| 2020 | Do We Need Zero Training Loss After Achieving Zero Training Error?abstractOverparameterized deep networks have the capacity to memorize training data with zero \emph{training error}. Even after memorization, the \emph{training loss} continues to approach zero, making the model overconfident and the test performance degraded. Since existing regularizers do not directly aim to avoid zero training loss, it is hard to tune their hyperparameters in order to maintain a fixed/preset level of training loss. We propose a direct solution called \emph{flooding} that intentionally prevents further reduction of the training loss when it reaches a reasonably small value, which we call the \emph{flood level}. Our approach makes the loss float around the flood level by doing mini-batched gradient descent as usual but gradient ascent if the training loss is below the flood level. This can be implemented with one line of code and is compatible with any stochastic optimizer and other regularizers. With flooding, the model will continue to “random walk” with the same non-zero training loss, and we expect it to drift into an area with a flat loss landscape that leads to better generalization. We experimentally show that flooding improves performance and, as a byproduct, induces a double descent curve of the test loss. Takashi Ishida 0001, Ikko Yamane, Tomoya Sakai 0001, Gang Niu 0001, Masashi Sugiyama |
ICML | 3 |
| 2020 | A Predictive Optimization Framework for Hierarchical Demand MatchingabstractPredictive optimization is a framework for designing an entire data-analysis pipeline that comprises both prediction and optimization, to be able to maximize overall throughput performance. In practical demand analysis, a knowledge of hierarchies, which might be geographical or categorical, is recognized as useful, though such additional knowledge has not been taken into account in existing predictive optimization. In this paper, we propose a novel hierarchical predictive optimization pipeline that is able to deal with a wide range of applications including inventory management. Based on an existing hierarchical demand prediction model, we present a stochastic matching framework that can manage prediction-uncertainty in decision making. We further provide a greedy approximation algorithm for solving demand matching on hierarchical structures. In experimental evaluations on both artificial and real-world data, we demonstrate the effectiveness of our proposed hierarchical-predictive-optimization pipeline. Naoto Ohsaka, Tomoya Sakai 0001, Akihiro Yabe |
SDM | 2 |
| 2020 | Robust modal regression with direct gradient approximation of modal regression riskabstractModal regression is aimed at estimating the global mode (i.e., global maximum) of the conditional density function of the output variable given input variables, and has led to regression methods robust against a wide-range of noises. A typical approach for modal regression takes a two-step approach of firstly approximating the modal regression risk (MRR) and of secondly maximizing the approximated MRR with some gradient method. However, this two-step approach can be suboptimal in gradient-based maximization methods because a good MRR approximator does not necessarily give a good gradient approximator of MRR. In this paper, we take a novel approach of \emph{directly} approximating the gradient of MRR in modal regression. Based on the direct approach, we first propose a modal regression method with reproducing kernels where a new update rule to estimate the conditional mode is derived based on a fixed-point method. Then, the derived update rule is theoretically investigated. Furthermore, since our direct approach is compatible with recent sophisticated stochastic gradient methods (e.g., Adam), another modal regression method is also proposed based on neural networks. Finally, the superior performance of the proposed methods is demonstrated on various artificial and benchmark datasets. Hiroaki Sasaki, Tomoya Sakai 0001, Takafumi Kanamori |
UAI | 2 |
| 2019 | Covariate Shift Adaptation on Learning from Positive and Unlabeled DataabstractThe goal of binary classification is to identify whether an input sample belongs to positive or negative classes. Usually, supervised learning is applied to obtain a classification rule, but in real-world applications, it is conceivable that only positive and unlabeled data are accessible for learning, which is called learning from positive and unlabeled data (PU learning). Furthermore, in practice, data distributions are likely to differ between training and testing due to, for example, time variation and domain shift. The covariate shift is a dataset shift situation, where distributions of covariates (inputs) differ between training and testing, but the input-output relation is the same. In this paper, we address the PU learning problem under the covariate shift. We propose an importanceweighted PU learning method and reveal in which situations the importance-weighting is necessary. Moreover, we derive the convergence rate of the proposed method under mild conditions and experimentally demonstrate its effectiveness. Tomoya Sakai 0001, Nobuyuki Shimizu |
AAAI | 1 |
| 2018 | Semi-supervised AUC optimization based on positive-unlabeled learning
Tomoya Sakai 0001, Gang Niu 0001, Masashi Sugiyama |
Mach. Learn. | 1 |
| 2018 | Correction to: Semi-supervised AUC optimization based on positive-unlabeled learning
Tomoya Sakai 0001, Gang Niu 0001, Masashi Sugiyama |
Mach. Learn. | 1 |
| 2018 | Convex formulation of multiple instance learning from positive and unlabeled bags
Han Bao 0002, Tomoya Sakai 0001, Issei Sato, Masashi Sugiyama |
Neural Networks | 2 |
| 2017 | Least-Squares Log-Density Gradient Clustering for Riemannian ManifoldsabstractMean shift is a mode-seeking clustering algorithm that has been successfully used in a wide range of applications such as image segmentation and object tracking. To further improve the clustering performance, mean shift has been extended to various directions, including generalization to handle data on Riemannian manifolds and extension to directly estimating the density gradient without density estimation. In this paper, we combine these ideas and propose a novel mode-seeking algorithm for Riemannian manifolds with direct density-gradient estimation. Although the idea of combining the two extensions is rather straightforward, directly estimating the density gradient on Riemannian manifolds is mathematically challenging. We will provide a mathematically sound algorithm and demonstrate its usefulness through experiments. Mina Ashizawa, Hiroaki Sasaki, Tomoya Sakai 0001, Masashi Sugiyama |
AISTATS | 3 |
| 2017 | Semi-Supervised Classification Based on Classification from Positive and Unlabeled DataabstractMost of the semi-supervised classification methods developed so far use unlabeled data for regularization purposes under particular distributional assumptions such as the cluster assumption. In contrast, recently developed methods of classification from positive and unlabeled data (PU classification) use unlabeled data for risk evaluation, i.e., label information is directly extracted from unlabeled data. In this paper, we extend PU classification to also incorporate negative data and propose a novel semi-supervised learning approach. We establish generalization error bounds for our novel methods and show that the bounds decrease with respect to the number of unlabeled data without the distributional assumptions that are required in existing semi-supervised learning methods. Through experiments, we demonstrate the usefulness of the proposed methods. Tomoya Sakai 0001, Marthinus Christoffel du Plessis, Gang Niu 0001, Masashi Sugiyama |
ICML | 1 |
| 2016 | Theoretical Comparisons of Positive-Unlabeled Learning against Positive-Negative LearningabstractIn PU learning, a binary classifier is trained from positive (P) and unlabeled (U) data without negative (N) data. Although N data is missing, it sometimes outperforms PN learning (i.e., ordinary supervised learning). Hitherto, neither theoretical nor experimental analysis has been given to explain this phenomenon. In this paper, we theoretically compare PU (and NU) learning against PN learning based on the upper bounds on estimation errors. We find simple conditions when PU and NU learning are likely to outperform PN learning, and we prove that, in terms of the upper bounds, either PU or NU learning (depending on the class-prior probability and the sizes of P and N data) given infinite U data will improve on PN learning. Our theoretical findings well agree with the experimental results on artificial and benchmark data even when the experimental setup does not match the theoretical assumptions exactly. Gang Niu 0001, Marthinus Christoffel du Plessis, Tomoya Sakai 0001, Masashi Sugiyama |
NIPS | 3 |