Takafumi Kanamori

dblp:76/6882 · DBLP profile ↗
← Back
58ranked-venue papers
15as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 52 · 14 first-author · 11 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021Theory of computation · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Noiseless Diffusion-GAN: Scaling-based data augmentation for generative models
abstract
This paper explores stable learning methods for generative models designed to facilitate high-quality data generation. Noise injection is a commonly employed technique to enhance learning stability; however, selecting an appropriate noise distribution remains a significant challenge. Diffusion-GAN, a recently proposed approach, addresses this issue by leveraging the diffusion process alongside a timestep-dependent discriminator. In this study, we analyze Diffusion-GAN and identify data scaling as a critical factor for achieving stable learning and high-quality data generation. Based on these insights, we introduce a learning algorithm, termed Scale-GAN, which incorporates data scaling and variance-based regularization. Moreover, we provide a theoretical proof demonstrating that data scaling effectively manages the bias-variance trade-off within the estimation error bound. Experimental evaluations on standard benchmark datasets highlight the proposed method's efficacy in enhancing both stability and accuracy.
Yoshitaka Koike, Takumi Nakagawa, Hiroki Waida, Takafumi Kanamori
Neural Networks4
2025 TULiP: Test-Time Uncertainty Estimation via Linearization and Weight Perturbation
Dongshen Wu, Yuichiro Wada, Takafumi Kanamori
ICONIP (1)4
2025 Tensor dictionary-based heterogeneous transfer learning to study emotion-related gender differences in brain
Lan Yang 0010, Chen Qiao, Takafumi Kanamori, Vince D. Calhoun, Julia M. Stephen, Tony W. Wilson, Yu-Ping Wang 0002
Neural Networks3
2024 Robust VAEs via Generating Process of Noise Augmented Data
abstract
Advancing defensive mechanisms against adversar-ial attacks in generative models is a critical research topic in machine learning. Our study focuses on a specific type of generative models - Variational Auto-Encoders (VAEs). Contrary to common beliefs and existing literature which suggest that noise injection towards training data can make models more robust, our preliminary experiments revealed that naive usage of noise augmentation technique did not substantially improve VAE robustness. In fact, it even degraded the quality of learned representations, making VAEs more susceptible to adversarial perturbations. This paper introduces a novel framework that enhances robustness by regularizing the latent space divergence between original and noise-augmented data. Through incorpo-rating a paired probabilistic prior into the standard variational lower bound, our method significantly boosts defense against adversarial attacks. Our empirical evaluations demonstrate that this approach, termed Robust Augmented Variational Auto-EN coder (RAVEN), yields superior performance in resisting adversarial inputs on widely-recognized benchmark datasets.
Hiroo Irobe, Wataru Aoki, Kimihiro Yamazaki, Takumi Nakagawa, Hiroki Waida, Yuichiro Wada, Takafumi Kanamori
ISIT8
2024 Denoising cosine similarity: A theory-driven approach for efficient representation learning
Takumi Nakagawa, Yutaro Sanada, Hiroki Waida, Yuichiro Wada, Kosaku Takanashi, Tomonori Yamada, Takafumi Kanamori
Neural Networks8
2023 Towards Understanding the Mechanism of Contrastive Learning via Similarity Structure: A Theoretical Analysis
Hiroki Waida, Yuichiro Wada, Léo Andéol, Takumi Nakagawa, Takafumi Kanamori
ECML/PKDD (4)6
2023 Deep Clustering With a Constraint for Topological Invariance Based on Symmetric InfoNCE
abstract
We consider the scenario of deep clustering, in which the available prior knowledge is limited. In this scenario, few existing state-of-the-art deep clustering methods can perform well for both noncomplex topology and complex topology data sets. To address the problem, we propose a constraint utilizing symmetric InfoNCE, which helps an objective of the deep clustering method in the scenario of training the model so as to be efficient for not only noncomplex topology but also complex topology data sets. Additionally, we provide several theoretical explanations of the reason that the constraint can enhances the performance of deep clustering methods. To confirm the effectiveness of the proposed constraint, we introduce a deep clustering method named MIST, which is a combination of an existing deep clustering method and our constraint. Our numerical experiments via MIST demonstrate that the constraint is effective. In addition, MIST outperforms other state-of-the-art deep clustering methods for most of the commonly used 10 benchmark data sets.
Yuichiro Wada, Hiroki Waida, Kaito Goto, Yusaku Hino, Takafumi Kanamori
Neural Comput.6
2023 Learning domain invariant representations by joint Wasserstein distance minimization
abstract
Domain shifts in the training data are common in practical applications of machine learning; they occur for instance when the data is coming from different sources. Ideally, a ML model should work well independently of these shifts, for example, by learning a domain-invariant representation. However, common ML losses do not give strong guarantees on how consistently the ML model performs for different domains, in particular, whether the model performs well on a domain at the expense of its performance on another domain. In this paper, we build new theoretical foundations for this problem, by contributing a set of mathematical relations between classical losses for supervised ML and the Wasserstein distance in joint space (i.e. representation and output space). We show that classification or regression losses, when combined with a GAN-type discriminator between domains, form an upper-bound to the true Wasserstein distance between domains. This implies a more invariant representation and also more stable prediction performance across domains. Theoretical results are corroborated empirically on several image datasets. Our proposed approach systematically produces the highest minimum classification accuracy across domains, and the most invariant representation.
Léo Andéol, Yusei Kawakami, Yuichiro Wada, Takafumi Kanamori, Klaus-Robert Müller, Grégoire Montavon
Neural Networks4
2022 Mode estimation on matrix manifolds: Convergence and robustness
abstract
Data on matrix manifolds are ubiquitous on a wide range of research fields. The key issue is estimation of the modes (i.e., maxima) of the probability density function underlying the data. For instance, local modes (i.e., local maxima) can be used for clustering, while the global mode (i.e., the global maximum) is a robust alternative to the Frechet mean. Previously, to estimate the modes, an iterative method has been proposed based on a Riemannian gradient estimator and empirically showed the superior performance in clustering (Ashizawa et al., 2017). However, it has not been theoretically investigated if the iterative method is able to capture the modes based on the gradient estimator. In this paper, we propose simple iterative methods for mode estimation on matrix manifolds based on the Euclidean metric. A key contribution is to perform theoretical analysis and establish sufficient conditions for the monotonic ascending and convergence of the proposed iterative methods. In addition, for the previous method, we prove the monotonic ascending property towards a mode. Thus, our work can be also regarded as compensating for the lack of theoretical analysis in the previous method. Furthermore, the robustness of the iterative methods is theoretically investigated in terms of the breakdown point. Finally, the proposed methods are experimentally demonstrated to work well in clustering and robust mode estimation on matrix manifolds.
Hiroaki Sasaki, Takafumi Kanamori
AISTATS3
2022 Deep Self-Supervised Learning of Speech Denoising from Noisy Speeches
Yutaro Sanada, Takumi Nakagawa, Yuichiro Wada, Kosaku Takanashi, Kiichi Tokuyama, Takafumi Kanamori, Tomonori Yamada
INTERSPEECH7
2022 Estimating Density Models with Truncation Boundaries using Score Matching
abstract
Truncated densities are probability density functions defined on truncated domains. They share the same parametric form with their non-truncated counterparts up to a normalizing constant. Since the computation of their normalizing constants is usually infeasible, Maximum Likelihood Estimation cannot be easily applied to estimate truncated density models. Score Matching (SM) is a powerful tool for fitting parameters using only unnormalized models. However, it cannot be directly applied here as boundary conditions that derive a tractable SM objective are not satisfied by truncated densities. This paper studies parameter estimation for truncated probability densities using SM. The estimator minimizes a weighted Fisher divergence. The weight function is simply the shortest distance from a data point to the domain's boundary. We show this choice of weight function naturally arises from minimizing the Stein discrepancy and upper bounding the finite-sample estimation error. We demonstrate the usefulness of our method via numerical experiments and a study on the Chicago crime data set. We also show that the proposed density estimation can correct the outlier-trimming bias caused by aggressive outlier detection methods.
Song Liu 0002, Takafumi Kanamori, Daniel J. Williams
J. Mach. Learn. Res.2
2021 Uncertainty propagation for dropout-based Bayesian neural networks
abstract
Uncertainty evaluation is a core technique when deep neural networks (DNNs) are used in real-world problems. In practical applications, we often encounter unexpected samples that have not seen in the training process. Not only achieving the high-prediction accuracy but also detecting uncertain data is significant for safety-critical systems. In statistics and machine learning, Bayesian inference has been exploited for uncertainty evaluation. The Bayesian neural networks (BNNs) have recently attracted considerable attention in this context, as the DNN trained using dropout is interpreted as a Bayesian method. Based on this interpretation, several methods to calculate the Bayes predictive distribution for DNNs have been developed. Though the Monte-Carlo method called MC dropout is a popular method for uncertainty evaluation, it requires a number of repeated feed-forward calculations of DNNs with randomly sampled weight parameters. To overcome the computational issue, we propose a sampling-free method to evaluate uncertainty. Our method converts a neural network trained using dropout to the corresponding Bayesian neural network with variance propagation. Our method is available not only to feed-forward NNs but also to recurrent NNs such as LSTM. We report the computational efficiency and statistical reliability of our method in numerical experiments of language modeling using RNNs, and the out-of-distribution detection with DNNs.
Yuki Mae, Wataru Kumagai, Takafumi Kanamori
Neural Networks3
2020 A Unified Statistically Efficient Estimation Framework for Unnormalized Models
abstract
The parameter estimation of unnormalized models is a challenging problem. The maximum likelihood estimation (MLE) is computationally infeasible for these models since normalizing constants are not explicitly calculated. Although some consistent estimators have been proposed earlier, the problem of statistical efficiency remains. In this study, we propose a unified, statistically efficient estimation framework for unnormalized models and several efficient estimators, whose asymptotic variance is the same as the MLE. The computational cost of these estimators is also reasonable and they can be employed whether the sample space is discrete or continuous. The loss functions of the proposed estimators are derived by combining the following two methods: (1) density-ratio matching using Bregman divergence, and (2) plugging-in nonparametric estimators. We also analyze the properties of the proposed estimators when the unnormalized models are misspecified. The experimental results demonstrate the advantages of our method over existing approaches.
Masatoshi Uehara, Takafumi Kanamori, Takashi Takenouchi, Takeru Matsuda
AISTATS2
2020 Robust modal regression with direct gradient approximation of modal regression risk
abstract
Modal regression is aimed at estimating the global mode (i.e., global maximum) of the conditional density function of the output variable given input variables, and has led to regression methods robust against a wide-range of noises. A typical approach for modal regression takes a two-step approach of firstly approximating the modal regression risk (MRR) and of secondly maximizing the approximated MRR with some gradient method. However, this two-step approach can be suboptimal in gradient-based maximization methods because a good MRR approximator does not necessarily give a good gradient approximator of MRR. In this paper, we take a novel approach of \emph{directly} approximating the gradient of MRR in modal regression. Based on the direct approach, we first propose a modal regression method with reproducing kernels where a new update rule to estimate the conditional mode is derived based on a fixed-point method. Then, the derived update rule is theoretically investigated. Furthermore, since our direct approach is compatible with recent sophisticated stochastic gradient methods (e.g., Adam), another modal regression method is also proposed based on neural networks. Finally, the superior performance of the proposed methods is demonstrated on various artificial and benchmark datasets.
Hiroaki Sasaki, Tomoya Sakai 0001, Takafumi Kanamori
UAI3
2019 Fisher Efficient Inference of Intractable Models
abstract
Maximum Likelihood Estimators (MLE) has many good properties. For example, the asymptotic variance of MLE solution attains equality of the asymptotic Cram{\'e}r-Rao lower bound (efficiency bound), which is the minimum possible variance for an unbiased estimator. However, obtaining such MLE solution requires calculating the likelihood function which may not be tractable due to the normalization term of the density model. In this paper, we derive a Discriminative Likelihood Estimator (DLE) from the Kullback-Leibler divergence minimization criterion implemented via density ratio estimation and a Stein operator. We study the problem of model inference using DLE. We prove its consistency and show that the asymptotic variance of its solution can attain the equality of the efficiency bound under mild regularity conditions. We also propose a dual formulation of DLE which can be easily optimized. Numerical studies validate our asymptotic theorems and we give an example where DLE successfully estimates an intractable model constructed using a pre-trained deep neural network.
Song Liu 0002, Takafumi Kanamori, Wittawat Jitkrittum
NeurIPS2
2019 Risk bound of transfer learning using parametric feature mapping and its application to sparse coding
abstract
In this study, we consider a transfer-learning problem using the parameter transfer approach, in which a suitable parameter of feature mapping is learned through one task and applied to another objective task. We introduce the notion of local stability and parameter transfer learnability of parametric feature mapping, and derive an excess risk bound for parameter transfer algorithms. As an application of parameter transfer learning, we discuss the performance of sparse coding in self-taught learning. Although self-taught learning algorithms with a large volume of unlabeled data often show excellent empirical performance, their theoretical analysis has not yet been studied. In this paper, we also provide a theoretical excess risk bound for self-taught learning. In addition, we show that the results of numerical experiments agree with our theoretical analysis.
Wataru Kumagai, Takafumi Kanamori
Mach. Learn.2
2019 Variable Selection for Nonparametric Learning with Power Series Kernels
abstract
In this letter, we propose a variable selection method for general nonparametric kernel-based estimation. The proposed method consists of two-stage estimation: (1) construct a consistent estimator of the target function, and (2) approximate the estimator using a few variables by [Formula: see text]-type penalized estimation. We see that the proposed method can be applied to various kernel nonparametric estimation such as kernel ridge regression, kernel-based density, and density-ratio estimation. We prove that the proposed method has the property of variable selection consistency when the power series kernel is used. Here, the power series kernel is a certain class of kernels containing polynomial and exponential kernels. This result is regarded as an extension of the variable selection consistency for the nonnegative garrote (NNG), a special case of the adaptive Lasso, to the kernel-based estimators. Several experiments, including simulation studies and real data applications, show the effectiveness of the proposed method.
Kota Matsui, Wataru Kumagai, Kenta Kanamori, Mitsuaki Nishikimi, Takafumi Kanamori
Neural Comput.5
2017 Estimating Density Ridges by Direct Estimation of Density-Derivative-Ratios
abstract
Estimation of \emphdensity ridges has been gathering a great deal of attention since it enables us to reveal lower-dimensional structures hidden in data. Recently, \emphsubspace constrained mean shift (SCMS) was proposed as a practical algorithm for density ridge estimation. A key technical ingredient in SCMS is to accurately estimate the ratios of the density derivatives to the density. SCMS takes a three-step approach for this purpose — first estimating the data density, then computing its derivatives, and finally taking their ratios. However, this three-step approach can be unreliable because a good density estimator does not necessarily mean a good density derivative estimator and division by an estimated density could significantly magnify the estimation error. To overcome these problems, we propose a novel method that directly estimates the ratios without going through density estimation and division. Our proposed estimator has an analytic-form solution and it can be computed efficiently. We further establish a non-parametric convergence bound for the proposed ratio estimator. Finally, based on this direct ratio estimator, we develop a practical algorithm for density ridge estimation and experimentally demonstrate its usefulness on a variety of datasets.
Hiroaki Sasaki, Takafumi Kanamori, Masashi Sugiyama
AISTATS2
2017 Parallel distributed block coordinate descent methods based on pairwise comparison oracle
Kota Matsui, Wataru Kumagai, Takafumi Kanamori
J. Glob. Optim.3
2017 Mode-Seeking Clustering and Density Ridge Estimation via Direct Estimation of Density-Derivative-Ratios
Hiroaki Sasaki, Takafumi Kanamori, Aapo Hyvärinen, Gang Niu 0001, Masashi Sugiyama
J. Mach. Learn. Res.2
2017 Statistical Inference with Unnormalized Discrete Models and Localized Homogeneous Divergences
abstract
In this paper, we focus on parameters estimation of probabilistic models in discrete space. A naive calculation of the normalization constant of the probabilistic model on discrete space is often infeasible and statistical inference based on such probabilistic models has difficulty. In this paper, we propose a novel estimator for probabilistic models on discrete space, which is derived from an empirically localized homogeneous divergence. The idea of the empirical localization makes it possible to ignore an unobserved domain on sample space, and the homogeneous divergence is a discrepancy measure between two positive measures and has a weak coincidence axiom. The proposed estimator can be constructed without calculating the normalization constant and is asymptotically consistent and Fisher efficient. We investigate statistical properties of the proposed estimator and reveal a relationship between the empirically localized homogeneous divergence and a mixture of the $\alpha$-divergence. The $\alpha$-divergence is a non- homogeneous discrepancy measure that is frequently discussed in the context of information geometry. Using the relationship, we also propose an asymptotically consistent estimator of the normalization constant. Experiments showed that the proposed estimator comparably performs to the maximum likelihood estimator but with drastically lower computational cost.
Takashi Takenouchi, Takafumi Kanamori
J. Mach. Learn. Res.2
2017 DC Algorithm for Extended Robust Support Vector Machine
abstract
Nonconvex variants of support vector machines (SVMs) have been developed for various purposes. For example, robust SVMs attain robustness to outliers by using a nonconvex loss function, while extended [Formula: see text]-SVM (E[Formula: see text]-SVM) extends the range of the hyperparameter by introducing a nonconvex constraint. Here, we consider an extended robust support vector machine (ER-SVM), a robust variant of E[Formula: see text]-SVM. ER-SVM combines two types of nonconvexity from robust SVMs and E[Formula: see text]-SVM. Because of the two nonconvexities, the existing algorithm we proposed needs to be divided into two parts depending on whether the hyperparameter value is in the extended range or not. The algorithm also heuristically solves the nonconvex problem in the extended range. In this letter, we propose a new, efficient algorithm for ER-SVM. The algorithm deals with two types of nonconvexity while never entailing more computations than either E[Formula: see text]-SVM or robust SVM, and it finds a critical point of ER-SVM. Furthermore, we show that ER-SVM includes the existing robust SVMs as special cases. Numerical experiments confirm the effectiveness of integrating the two nonconvexities.
Shuhei Fujiwara, Akiko Takeda, Takafumi Kanamori
Neural Comput.3
2017 Robustness of learning algorithms using hinge loss with outlier indicators
Takafumi Kanamori, Shuhei Fujiwara, Akiko Takeda
Neural Networks1
2017 Graph-based composite local Bregman divergences on discrete sample spaces
Takafumi Kanamori, Takashi Takenouchi
Neural Networks1
2015 Empirical Localization of Homogeneous Divergences on Discrete Sample Spaces
abstract
In this paper, we propose a novel parameter estimator for probabilistic models on discrete space. The proposed estimator is derived from minimization of homogeneous divergence and can be constructed without calculation of the normalization constant, which is frequently infeasible for models in the discrete space. We investigate statistical properties of the proposed estimator such as consistency and asymptotic normality, and reveal a relationship with the alpha-divergence. Small experiments show that the proposed estimator attains comparable performance to the MLE with drastically lower computational cost.
Takashi Takenouchi, Takafumi Kanamori
NIPS2
2014 Extended Robust Support Vector Machine Based on Financial Risk Minimization
abstract
Financial risk measures have been used recently in machine learning. For example, ν-support vector machine ν-SVM) minimizes the conditional value at risk (CVaR) of margin distribution. The measure is popular in finance because of the subadditivity property, but it is very sensitive to a few outliers in the tail of the distribution. We propose a new classification method, extended robust SVM (ER-SVM), which minimizes an intermediate risk measure between the CVaR and value at risk (VaR) by expecting that the resulting model becomes less sensitive than ν-SVM to outliers. We can regard ER-SVM as an extension of robust SVM, which uses a truncated hinge loss. Numerical experiments imply the ER-SVM's possibility of achieving a better prediction performance with proper parameter setting.
Akiko Takeda, Shuhei Fujiwara, Takafumi Kanamori
Neural Comput.3
2014 Using financial risk measures for analyzing generalization performance of machine learning models
Akiko Takeda, Takafumi Kanamori
Neural Networks2
2013 Conjugate relation between loss functions and uncertainty sets in classification problems
Takafumi Kanamori, Akiko Takeda, Taiji Suzuki
J. Mach. Learn. Res.1
2013 Computational complexity of kernel-based density-ratio estimation: a condition number analysis
abstract
In this study, the computational properties of a kernel-based least-squares density-ratio estimator are investigated from the viewpoint of condition numbers . The condition number of the Hessian matrix of the loss function is closely related to the convergence rate of optimization and the numerical stability. We use smoothed analysis techniques and theoretically demonstrate that the kernel least-squares method has a smaller condition number than other M-estimators. This implies that the kernel least-squares method has desirable computational properties. In addition, an alternate formulation of the kernel least-squares estimator that possesses an even smaller condition number is presented. The validity of the theoretical analysis is verified through numerical experiments.
Takafumi Kanamori, Taiji Suzuki, Masashi Sugiyama
Mach. Learn.1
2013 Semi-supervised learning with density-ratio estimation
Masanori Kawakita, Takafumi Kanamori
Mach. Learn.2
2013 Density-Difference Estimation
abstract
We address the problem of estimating the difference between two probability densities. A naive approach is a two-step procedure of first estimating two densities separately and then computing their difference. However, this procedure does not necessarily work well because the first step is performed without regard to the second step, and thus a small estimation error incurred in the first stage can cause a big error in the second stage. In this letter, we propose a single-shot procedure for directly estimating the density difference without separately estimating two densities. We derive a nonparametric finite-sample error bound for the proposed single-shot density-difference estimator and show that it achieves the optimal convergence rate. We then show how the proposed density-difference estimator can be used in L²-distance approximation. Finally, we experimentally demonstrate the usefulness of the proposed method in robust distribution comparison such as class-prior estimation and change-point detection.
Masashi Sugiyama, Takafumi Kanamori, Taiji Suzuki, Marthinus Christoffel du Plessis, Song Liu 0002, Ichiro Takeuchi
Neural Comput.2
2013 A Unified Classification Model Based on Robust Optimization
abstract
A wide variety of machine learning algorithms such as the support vector machine (SVM), minimax probability machine (MPM), and Fisher discriminant analysis (FDA) exist for binary classification. The purpose of this letter is to provide a unified classification model that includes these models through a robust optimization approach. This unified model has several benefits. One is that the extensions and improvements intended for SVMs become applicable to MPM and FDA, and vice versa. For example, we can obtain nonconvex variants of MPM and FDA by mimicking Perez-Cruz, Weston, Hermann, and Schölkopf's (2003) extension from convex ν-SVM to nonconvex Eν-SVM. Another benefit is to provide theoretical results concerning these learning methods at once by dealing with the unified model. We give a statistical interpretation of the unified classification model and prove that the model is a good approximation for the worst-case minimization of an expected loss with respect to the uncertain probability distribution. We also propose a nonconvex optimization algorithm that can be applied to nonconvex variants of existing learning methods and show promising numerical results.
Akiko Takeda, Hiroyuki Mitsugi, Takafumi Kanamori
Neural Comput.3
2013 Relative Density-Ratio Estimation for Robust Distribution Comparison
abstract
Divergence estimators based on direct approximation of density ratios without going through separate approximation of numerator and denominator densities have been successfully applied to machine learning tasks that involve distribution comparison such as outlier detection, transfer learning, and two-sample homogeneity test. However, since density-ratio functions often possess high fluctuation, divergence estimation is a challenging task in practice. In this letter, we use relative divergences for distribution comparison, which involves approximation of relative density ratios. Since relative density ratios are always smoother than corresponding ordinary density ratios, our proposed method is favorable in terms of nonparametric convergence speed. Furthermore, we show that the proposed divergence estimator has asymptotic variance independent of the model complexity under a parametric setup, implying that the proposed estimator hardly overfits even with complex models. Through experiments, we demonstrate the usefulness of the proposed approach.
Makoto Yamada, Taiji Suzuki, Takafumi Kanamori, Hirotaka Hachiya, Masashi Sugiyama
Neural Comput.3
2012 A Unified Robust Classification Model
Akiko Takeda, Hiroyuki Mitsugi, Takafumi Kanamori
ICML3
2012 Non-convex Optimization on Stiefel Manifold and Applications to Machine Learning
Takafumi Kanamori, Akiko Takeda
ICONIP (1)1
2012 Density-Difference Estimation
abstract
We address the problem of estimating the difference between two probability densities. A naive approach is a two-step procedure of first estimating two densities separately and then computing their difference. However, such a two-step procedure does not necessarily work well because the first step is performed without regard to the second step and thus a small estimation error incurred in the first stage can cause a big error in the second stage. In this paper, we propose a single-shot procedure for directly estimating the density difference without separately estimating two densities. We derive a non-parametric finite-sample error bound for the proposed single-shot density-difference estimator and show that it achieves the optimal convergence rate. We then show how the proposed density-difference estimator can be utilized in L2-distance approximation. Finally, we experimentally demonstrate the usefulness of the proposed method in robust distribution comparison such as class-prior estimation and change-point detection.
Masashi Sugiyama, Takafumi Kanamori, Taiji Suzuki, Marthinus Christoffel du Plessis, Song Liu 0002, Ichiro Takeuchi
NIPS2
2012 Statistical analysis of kernel-based least-squares density-ratio estimation
Takafumi Kanamori, Taiji Suzuki, Masashi Sugiyama
Mach. Learn.1
2012 f-Divergence Estimation and Two-Sample Homogeneity Test Under Semiparametric Density-Ratio Models
abstract
A density ratio is defined by the ratio of two probability densities. We study the inference problem of density ratios and apply a semiparametric density-ratio estimator to the two-sample homogeneity test. In the proposed test procedure, the$f$-divergence between two probability densities is estimated using a density-ratio estimator. The$f$-divergence estimator is then exploited for the two-sample homogeneity test. We derive an optimal estimator of$f$-divergence in the sense of the asymptotic variance in a semiparametric setting, and provide a statistic for two-sample homogeneity test based on the optimal estimator. We prove that the proposed test dominates the existing empirical likelihood score test. Through numerical studies, we illustrate the adequacy of the asymptotic theory for finite-sample inference.
Takafumi Kanamori, Taiji Suzuki, Masashi Sugiyama
IEEE Trans. Inf. Theory1
2011 Relative Density-Ratio Estimation for Robust Distribution Comparison
abstract
Divergence estimators based on direct approximation of density-ratios without going through separate approximation of numerator and denominator densities have been successfully applied to machine learning tasks that involve distribution comparison such as outlier detection, transfer learning, and two-sample homogeneity test. However, since density-ratio functions often possess high fluctuation, divergence estimation is still a challenging task in practice. In this paper, we propose to use relative divergences for distribution comparison, which involves approximation of relative density-ratios. Since relative density-ratios are always smoother than corresponding ordinary density-ratios, our proposed method is favorable in terms of the non-parametric convergence speed. Furthermore, we show that the proposed divergence estimator has asymptotic variance independent of the model complexity under a parametric setup, implying that the proposed estimator hardly overfits even with complex models. Through experiments, we demonstrate the usefulness of the proposed approach.
Makoto Yamada, Taiji Suzuki, Takafumi Kanamori, Hirotaka Hachiya, Masashi Sugiyama
NIPS3
2011 Statistical outlier detection using direct density ratio estimation
Shohei Hido, Yuta Tsuboi, Hisashi Kashima, Masashi Sugiyama, Takafumi Kanamori
Knowl. Inf. Syst.5
2011 Least-squares two-sample test
Masashi Sugiyama, Taiji Suzuki, Yuta Itoh 0001, Takafumi Kanamori, Manabu Kimura
Neural Networks4
2011 Direct density-ratio estimation with dimensionality reduction via least-squares hetero-distributional subspace search
Masashi Sugiyama, Makoto Yamada, Paul von Bünau, Taiji Suzuki, Takafumi Kanamori, Motoaki Kawanabe
Neural Networks5
2010 Direct Density Ratio Estimation with Dimensionality Reduction
abstract
Methods for directly estimating the ratio of two probability density functions without going through density estimation have been actively explored recently since they can be used for various data processing tasks such as non-stationarity adaptation, outlier detection, conditional density estimation, feature selection, and independent component analysis. However, even the state-of-the-art density ratio estimation methods still perform rather poorly in high-dimensional problems. In this paper, we propose a new density ratio estimation method which incorporates dimensionality reduction into a density ratio estimation procedure. Our key idea is to identify a low-dimensional subspace in which the two densities corresponding to the denominator and the numerator in the density ratio are significantly different. Then the density ratio is estimated only within this low-dimensional subspace. Through numerical examples, we illustrate the effectiveness of the proposed method.
Masashi Sugiyama, Satoshi Hara 0001, Paul von Bünau, Taiji Suzuki, Takafumi Kanamori, Motoaki Kawanabe
SDM5
2010 Deformation of log-likelihood loss function for multiclass boosting
Takafumi Kanamori
Neural Networks1
2009 Mutual information estimation reveals global associations between stimuli and biological processes
abstract
BACKGROUND: Although microarray gene expression analysis has become popular, it remains difficult to interpret the biological changes caused by stimuli or variation of conditions. Clustering of genes and associating each group with biological functions are often used methods. However, such methods only detect partial changes within cell processes. Herein, we propose a method for discovering global changes within a cell by associating observed conditions of gene expression with gene functions. RESULTS: To elucidate the association, we introduce a novel feature selection method called Least-Squares Mutual Information (LSMI), which computes mutual information without density estimaion, and therefore LSMI can detect nonlinear associations within a cell. We demonstrate the effectiveness of LSMI through comparison with existing methods. The results of the application to yeast microarray datasets reveal that non-natural stimuli affect various biological processes, whereas others are no significant relation to specific cell processes. Furthermore, we discover that biological processes can be categorized into four types according to the responses of various stimuli: DNA/RNA metabolism, gene expression, protein metabolism, and protein localization. CONCLUSION: We proposed a novel feature selection method called LSMI, and applied LSMI to mining the association between conditions of yeast and biological processes through microarray datasets. In fact, LSMI allows us to elucidate the global organization of cellular process control.
Taiji Suzuki, Masashi Sugiyama, Takafumi Kanamori, Jun Sese
BMC Bioinform.3
2009 A Least-squares Approach to Direct Importance Estimation
Takafumi Kanamori, Shohei Hido, Masashi Sugiyama
J. Mach. Learn. Res.1
2009 Nonparametric Conditional Density Estimation Using Piecewise-Linear Solution Path of Kernel Quantile Regression
abstract
The goal of regression analysis is to describe the stochastic relationship between an input vector x and a scalar output y. This can be achieved by estimating the entire conditional density p(y / x). In this letter, we present a new approach for nonparametric conditional density estimation. We develop a piecewise-linear path-following method for kernel-based quantile regression. It enables us to estimate the cumulative distribution function of p(y / x) in piecewise-linear form for all x in the input domain. Theoretical analyses and experimental results are presented to show the effectiveness of the approach.
Ichiro Takeuchi, Kaname Nomura, Takafumi Kanamori
Neural Comput.3
2008 Inlier-Based Outlier Detection via Direct Density Ratio Estimation
abstract
We propose a new statistical approach to the problem of inlier-based outlier detection, i.e.,finding outliers in the test set based on the training set consisting only of inliers. Our key idea is to use the ratio of training and test data densities as an outlier score; we estimate the ratio directly in a semi-parametric fashion without going through density estimation. Thus our approach is expected to have better performance in high-dimensional problems. Furthermore, the applied algorithm for density ratio estimation is equipped with a natural cross-validation procedure, allowing us to objectively optimize the value of tuning parameters such as the regularization parameter and the kernel width. The algorithm offers a closed-form solution as well as a closed-form formula for the leave-one-out error. Thanks to this, the proposed outlier detection method is computationally very efficient and is scalable to massive datasets. Simulations with benchmark and real-world datasets illustrate the usefulness of the proposed approach.
Shohei Hido, Yuta Tsuboi, Hisashi Kashima, Masashi Sugiyama, Takafumi Kanamori
ICDM5
2008 Efficient Direct Density Ratio Estimation for Non-stationarity Adaptation and Outlier Detection
abstract
We address the problem of estimating the ratio of two probability density functions (a.k.a.~the importance). The importance values can be used for various succeeding tasks such as non-stationarity adaptation or outlier detection. In this paper, we propose a new importance estimation method that has a closed-form solution; the leave-one-out cross-validation score can also be computed analytically. Therefore, the proposed method is computationally very efficient and numerically stable. We also elucidate theoretical properties of the proposed method such as the convergence rate and approximation error bound. Numerical experiments show that the proposed method is comparable to the best existing method in accuracy, while it is computationally more efficient than competing approaches.
Takafumi Kanamori, Shohei Hido, Masashi Sugiyama
NIPS1
2008 Robust Boosting Algorithm Against Mislabeling in Multiclass Problems
abstract
We discuss robustness against mislabeling in multiclass labels for classification problems and propose two algorithms of boosting, the normalized Eta-Boost.M and Eta-Boost.M, based on the Eta-divergence. Those two boosting algorithms are closely related to models of mislabeling in which the label is erroneously exchanged for others. For the two boosting algorithms, theoretical aspects supporting the robustness for mislabeling are explored. We apply the proposed two boosting methods for synthetic and real data sets to investigate the performance of these methods, focusing on robustness, and confirm the validity of the proposed methods.
Takashi Takenouchi, Shinto Eguchi, Noboru Murata, Takafumi Kanamori
Neural Comput.4
2007 Multiclass Boosting Algorithms for Shrinkage Estimators of Class Probability
Takafumi Kanamori
ALT1
2007 Pool-based active learning with optimal sampling distribution and its information geometrical interpretation
Takafumi Kanamori
Neurocomputing1
2007 Robust Loss Functions for Boosting
abstract
Boosting is known as a gradient descent algorithm over loss functions. It is often pointed out that the typical boosting algorithm, Adaboost, is highly affected by outliers. In this letter, loss functions for robust boosting are studied. Based on the concept of robust statistics, we propose a transformation of loss functions that makes boosting algorithms robust against extreme outliers. Next, the truncation of loss functions is applied to contamination models that describe the occurrence of mislabels near decision boundaries. Numerical experiments illustrate that the proposed loss functions derived from the contamination models are useful for handling highly noisy data in comparison with other loss functions.
Takafumi Kanamori, Takashi Takenouchi, Shinto Eguchi, Noboru Murata
Neural Comput.1
2006 The Entire Solution Path of Kernel-based Nonparametric Conditional Quantile Estimator
abstract
The goal of regression analysis is to describe the relationship between an output y and a vector of inputs x. Least squares regression provides how the mean of y changes with x, i.e. it estimates the conditional mean function. Estimating a set of conditional quantile functions provides a more complete view of the relationship between y and x. Quantile regression [1] is one of the promising approaches to estimate conditional quantile functions. Several types of quantile regression estimator have been studied in the literature. In this paper, we are particularly concerned with kernel-based nonparametric quantile regression formulated as a quadratic programing problem similar to those in support vector machine literature [2]. A group of conditional quantile functions, say, at the orders q = 0.1, 0.2, . . .. 0.9, can provide a nonparametric description of the conditional probability density p (y|x ). This requires us to solve many quadratic programming problems and it could be computationally demanding for large-scale problems. In this paper, inspired by the recently developed path following strategy [3][4], we derive an algorithm to solve a sequence of quadratic programming problems for the entire range of quantile orders q ∈ (0, 1). As well as the computational efficiency, the derived algorithm provides the full nonparametric description of the conditional distribution p (y|x). A few examples are given to illustrate the algorithm.
Ichiro Takeuchi, Kaname Nomura, Takafumi Kanamori
IJCNN3
2004 The Most Robust Loss Function for Boosting
Takafumi Kanamori, Takashi Takenouchi, Shinto Eguchi, Noboru Murata
ICONIP1
2004 Information Geometry of U-Boost and Bregman Divergence
abstract
We aim at an extension of AdaBoost to U-Boost, in the paradigm to build a stronger classification machine from a set of weak learning machines. A geometric understanding of the Bregman divergence defined by a generic convex function U leads to the U-Boost method in the framework of information geometry extended to the space of the finite measures over a label set. We propose two versions of U-Boost learning algorithms by taking account of whether the domain is restricted to the space of probability functions. In the sequential step, we observe that the two adjacent and the initial classifiers are associated with a right triangle in the scale via the Bregman divergence, called the Pythagorean relation. This leads to a mild convergence property of the U-Boost algorithm as seen in the expectation-maximization algorithm. Statistical discussions for consistency and robustness elucidate the properties of the U-Boost methods based on a stochastic assumption for training data.
Noboru Murata, Takashi Takenouchi, Takafumi Kanamori, Shinto Eguchi
Neural Comput.3
2002 A New Sequential Algorithm for Regression Problems by Using Mixture Distribution
Takafumi Kanamori
ICANN1
2002 Robust Regression with Asymmetric Heavy-Tail Noise Distributions
abstract
In the presence of a heavy-tail noise distribution, regression becomes much more difficult. Traditional robust regression methods assume that the noise distribution is symmetric, and they downweight the influence of so-called outliers. When the noise distribution is asymmetric, these methods yield biased regression estimators. Motivated by data-mining problems for the insurance industry, we propose a new approach to robust regression tailored to deal with asymmetric noise distribution. The main idea is to learn most of the parameters of the model using conditional quantile estimators (which are biased but robust estimators of the regression) and to learn a few remaining parameters to combine and correct these estimators, to minimize the average squared error in an unbiased way. Theoretical analysis and experiments show the clear advantages of the approach. Results are on artificial data as well as insurance data, using both linear and neural network predictors.
Ichiro Takeuchi, Yoshua Bengio, Takafumi Kanamori
Neural Comput.3