Kazuho Watanabe

dblp:12/3438 · DBLP profile ↗
← Back
46ranked-venue papers
25as first author
6since 2021 · last 2025
0000-0001-6357-5141ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 25 · 15 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 5 first-author · 4 since 2021Theory of computation · 7 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 2Security and privacy · 1
YearPublicationVenuePosition
2025 Empirical Bayes Estimation for Lasso-Type Regularizers: Analysis of Automatic Relevance Determination
abstract
This paper focuses on linear regression models with non-conjugate sparsity-inducing regularizers such as lasso and group lasso. Although the empirical Bayes approach enables us to estimate the regularization parameter, little is known on the properties of the estimators. In particular, many aspects regarding the specific conditions under which the mechanism of automatic relevance determination (ARD) occurs remain unexplained. In this paper, we derive the empirical Bayes estimators for the group lasso regularized linear regression models with limited parameters. It is shown that the estimators diverge under a specific condition, giving rise to the ARD mechanism. We also prove that empirical Bayes methods can produce the ARD mechanism in general regularized linear regression models and clarify the conditions under which models such as ridge, lasso, and group lasso can do so. The full version of this paper, including the Appendix, is accessible at https://arxiv.org/abs/2501.11280.
Tsukasa Yoshida, Kazuho Watanabe
ISIT2
2024 Unbiased Estimating Equation on Inverse Divergence and its Conditions
abstract
This paper focuses on the Bregman divergence defined by the reciprocal function, called the inverse divergence. For the loss function defined by the monotonically increasing function$f$and inverse divergence, the conditions for the statistical model and function$f$under which the estimating equation is unbiased are clarified. Specifically, we characterize two types of statistical models, an inverse Gaussian type and a mixture of generalized inverse Gaussian type distributions, to show that the conditions for the function$f$are different for each model. We also define Bregman divergence as a linear sum over the dimensions of the inverse divergence and extend the results to the multi-dimensional case.
Masahiro Kobayashi, Kazuho Watanabe
ISIT2
2024 Unbiased Estimating Equation and Latent Bias Under f-Separable Bregman Distortion Measures
abstract
We discuss unbiased estimating equations in a class of objective functions using a monotonically increasing function f and Bregman divergence. The choice of the function f gives desirable properties, such as robustness against outliers. To obtain unbiased estimating equations, analytically intractable integrals are generally required as bias correction terms. In this study, we clarify the combination of Bregman divergence, statistical model, and function f in which the bias correction term vanishes. Focusing on Mahalanobis and Itakura-Saito distances, we generalize fundamental existing results and characterize a class of distributions of positive reals with a scale parameter, including the gamma distribution as a special case. We also generalized these results to general model classes characterized by one-dimensional Bregman divergence. Furthermore, we discuss the possibility of latent bias minimization when the proportion of outliers is large, which is induced by the extinction of the bias correction term. We conducted numerical experiments to show that the latent bias can approach zero under heavy contamination of outliers or very small inliers.
Masahiro Kobayashi, Kazuho Watanabe
IEEE Trans. Inf. Theory2
2022 Approximate Empirical Bayes Estimation of the Regularization Parameter in ℓ1 Trend Filtering
abstract
Trend filtering is often used in economics and other fields. ℓ1trend filtering was proposed as a derivative of Hodrick-Prescott filtering based on the sparsity in the changes of trends. Although it has a regularization parameter, which needs to be set in advance, the non-conjugacy arising from the ℓ1regularization term prohibits a tractable Bayesian treatment including the sequence-dependent estimation of the regularization parameter. In this study, we consider the empirical Bayes estimation of the regularization parameter by approximating the non-conjugate prior distribution in ℓ1trend filtering by variational approximation.
Akiharu Omae, Kazuho Watanabe
ISIT2
2021 Statistical Learning of the Insensitive Parameter in Support Vector Models
abstract
We consider the estimation of the insensitive parameter$\varepsilon$in statistical models with$\varepsilon$-insensitive loss functions. The properties of the maximum likelihood estimators are studied for the$\varepsilon$-insensitive hyperbolic secant model. Focusing on the$\varepsilon$-insensitive Laplace and Gauss models, we analyze the average generalization errors of maximum likelihood and Bayesian learning. It is shown that$\varepsilon$-insensitive models behave as regular statistical models if the true generating distribution is in the interior of the parameter space, whereas non-regularity arises at the endpoint of the parameter space.
Kazuho Watanabe
ISIT1
2021 Generalized Dirichlet-process-means for f-separable distortion measures
abstract
DP-means clustering was obtained as an extension of K-means clustering. While it is implemented with a simple and efficient algorithm, it can estimate the number of clusters simultaneously. However, DP-means is specifically designed for the average distortion measure. Therefore, it is vulnerable to outliers in data, and can cause large maximum distortion in clusters. In this work, we extend the objective function of the DP-means to f-separable distortion measures and propose a unified learning algorithm to overcome the above problems by selecting the function f. Further, the influence function of the estimated cluster center is analyzed to evaluate the robustness against outliers. We demonstrate the performance of the generalized method by numerical experiments using real datasets.
Masahiro Kobayashi, Kazuho Watanabe
Neurocomputing2
2020 Multi-Decoder RNN Autoencoder Based on Variational Bayes Method
abstract
Clustering algorithms have wide applications and play an important role in data analysis fields including time series data analysis. However, in time series analysis, most of the algorithms used signal shape features or the initial value of hidden variable of a neural network. Little has been discussed on the methods based on the generative model of the time series. In this paper, we propose a new clustering algorithm focusing on the generative process of the signal with a recurrent neural network and the variational Bayes method. Our experiments show that the proposed algorithm not only has a robustness against for phase shift, amplitude and signal length variations but also provide a flexible clustering based on the property of the variational Bayes method.
Daisuke Kaji, Kazuho Watanabe, Masahiro Kobayashi
IJCNN2
2020 Discrete Optimal Reconstruction Distributions for Itakura-Saito Distortion Measure
abstract
The optimal reconstruction distribution achieving the rate-distortion function is elusive except for limited examples of sources and distortion measures if the rate-distortion function is strictly greater than the Shannon lower bound. In this paper, focusing on the Itakura-Saito distortion measure, we prove that if the Shannon lower bound is not tight, the optimal reconstruction distribution is purely discrete. Combined with the fact that the Shannon lower bound is tight for the gamma source, this result shows that it is the only source that has continuous optimal reconstruction distributions for the range of entire positive rate.
Kazuho Watanabe
ISIT1
2020 Unbiased Estimation Equation under f-Separable Bregman Distortion Measures
abstract
We discuss unbiased estimation equations in a class of objective function using a monotonically increasing function f and Bregman divergence. The choice of the function f gives desirable properties such as robustness against outliers. In order to obtain unbiased estimation equations, analytically intractable integrals are generally required as bias correction terms. In this study, we clarify the combination of Bregman divergence, statistical model, and function f in which the bias correction term vanishes. Focusing on Mahalanobis and Itakura-Saito distances, we provide a generalization of fundamental existing results and characterize a class of distributions of positive reals with a scale parameter, which includes the gamma distribution as a special case. We discuss the possibility of latent bias minimization when the proportion of outliers is large, which is induced by the extinction of the bias correction term.
Masahiro Kobayashi, Kazuho Watanabe
ITW2
2020 Context-aware placement of items with gaze-based interaction
abstract
Appropriate product placement significantly influences how viewers easily find their favorites, especially when they try to select products from digital signage displays. This increases the demand for dynamic categorization and optimal placement of items according to the context in which viewers explore their preferred choices. In this paper, we present an approach for optimizing the placement of items by respecting the underlying context in the search for favorites. Our approach starts with formulating the static placement of items as a constrained optimization problem, in which we incorporate design rules that highlight the underlying categorization of the items. We then extend this idea to accommodate dynamic placement according to the context in which users explore their preferred choices. This is accomplished by adaptively adjusting the priority of each item based on the distribution of visual attention obtained by an eye-tracking device. In particular, we construct a context map for understanding the relationship between the items by taking advantage of topic-based text mining techniques. We provide several examples of gaze-based interaction to demonstrate the capability of the proposed approach, which is followed by a discussion on possible directions for future research.
Shigeo Takahashi, Akane Uchita, Kazuho Watanabe, Masatoshi Arikawa
VINCI3
2019 Minimax Online Prediction of Varying Bernoulli Process under Variational Approximation
abstract
We consider the online prediction of a varying Bernoulli process (sequence of varying Bernoulli probabilities) from a single binary sequence. A real-valued online prediction method has been proposed as a prior work that incorporates the smoothness of the prediction sequence into the concept of the regret. Also, a Bayesian prediction method for the varying Bernoulli processes has been developed based on the variational inference. However, the former is not applicable to loss functions other than the squared error function, and the latter has no guarantee on the regret as an online prediction method. We propose a new online prediction method of a varying Bernoulli process from a single binary sequence with a guarantee to minimize the maximum regret under variational approximation. Through numerical experiments, we compare the Bayesian prediction method with the proposed method by using the regret with/without approximation and the KL divergence from the true underlying process. We discuss the prediction accuracy and influences of the approximation of the proposed method.
Kenta Konagayoshi, Kazuho Watanabe
ACML2
2018 Automatic DNN Node Pruning Using Mixture Distribution-based Group Regularization
Tsukasa Yoshida, Takafumi Moriya, Kazuho Watanabe, Yusuke Shinohara, Yoshikazu Yamaguchi, Yushi Aono
INTERSPEECH3
2018 Sparse Regression Code with Sparse Dictionary for Absolute Error Criterion
abstract
The Sparse Regression Code (SPARC) was proposed as an efficient lossy compression method for continuous sources, and was proved to achieve the rate-distortion curve for the i.i.d. Gaussian source. However, the original SPARC is specialized to the squared distortion criterion, and it is unknown how to adapt the SPARC to other distortion criteria. In this study, focusing on the absolute distortion criterion, we improve the original SPARC. This is achieved by designing the dictionary matrix sparsely as suggested by the rate-distortion theory of absolute distortion, and by deriving a proper sequence of regression coefficients. It is demonstrated that our algorithm yields smaller distortion than the original SPARC at all rates under the absolute distortion criterion.
Ryota Konabe, Kazuho Watanabe
ISIT2
2018 Generalized Dirichlet-Process-Means for Robust and Maximum Distortion Criteria
abstract
DP-means clustering was obtained as an extension of K-means clustering. While it is implemented with a simple and efficient algorithm, it can estimate the number of clusters simultaneously. However, DP-means is specifically designed for the average distortion criterion. Therefore, it is vulnerable to outliers in data, and can cause large maximum distortion in clusters. This study introduces a new parameter to the objective function of DP-means to provide an extension of DP-means, which bridges robust estimation of cluster centers and minimization of the maximum distortion criterion.
Masahiro Kobayashi, Kazuho Watanabe
ISITA2
2017 Making many-to-many parallel coordinate plots scalable by asymmetric biclustering
abstract
Datasets obtained through recently advanced measurement techniques tend to possess a large number of dimensions. This leads to explosively increasing computation costs for analyzing such datasets, thus making formulation and verification of scientific hypotheses very difficult. Therefore, an efficient approach to identifying feature subspaces of target datasets, that is, the subspaces of dimension variables or subsets of the data samples, is required to describe the essence hidden in the original dataset. This paper proposes a visual data mining framework for supporting semiautomatic data analysis that builds upon asymmetric biclustering to explore highly correlated feature subspaces. For this purpose, a variant of parallel coordinate plots, many-to-many parallel coordinate plots, is extended to visually assist appropriate selections of feature subspaces as well as to avoid intrinsic visual clutter. In this framework, biclustering is applied to dimension variables and data samples of the dataset simultaneously and asymmetrically. A set of variable axes are projected to a single composite axis while data samples between two consecutive variable axes are bundled using polygonal strips. This makes the visualization method scalable and enables it to play a key role in the framework. The effectiveness of the proposed framework has been empirically proven, and it is remarkably useful for many-to-many parallel coordinate plots.
Hsiang-Yun Wu, Yusuke Niibe, Kazuho Watanabe, Shigeo Takahashi, Makoto Uemura, Issei Fujishiro
PacificVis3
2017 Rate-distortion tradeoffs under Kernel-based distortion measures
abstract
Kernel methods have been used for turning linear learning algorithms into nonlinear ones. These nonlinear algorithms measure distances between data points by the distance in the kernel-induced feature space. In lossy data compression, the optimal tradeoff between the number of quantized points and the incurred distortion is characterized by the rate-distortion function. However, the rate-distortion functions associated with distortion measures involving kernel feature mapping have yet to be analyzed. We consider two reconstruction schemes, reconstruction in input space and reconstruction in feature space, and provide bounds to the rate-distortion functions for these schemes. Comparison of the derived bounds to the quantizer performance obtained by the kernel K-means method suggests that the rate-distortion bounds for input space and feature space reconstructions are informative at low and high distortion levels, respectively.
Kazuho Watanabe
ISIT1
2016 Constant-width rate-distortion bounds for power distortion measures
abstract
The explicit form of the rate-distortion function has rarely been obtained except for few cases where the Shannon lower bound coincides with the rate-distortion function for the entire range of the positive rate. In this paper, we consider the β-th power distortion measure, and prove that β-generalized Gaussian distribution is the only source that can make the Shannon lower bound tight for all distortion levels. We demonstrate that the tightness of the Shannon lower bound for β = 1 (Laplacian source) and β = 2 (Gaussian source) yields upper bounds to the rate-distortion function of power distortion measures with a different exponent, which have a constant gap from the Shannon lower bound for all distortion levels. Applying similar arguments to ε-insensitive distortion measures, we consider the tightness of the Shannon lower bound and derive an upper bound to the distortion-rate function, which is accurate at low rates.
Kazuho Watanabe
ITW1
2016 Rate-Distortion Functions for Gamma-Type Sources Under Absolute-Log Distortion Measure
abstract
When the information source is a continuous distribution and the rate-distortion function is strictly larger than the Shannon lower bound, the explicit evaluation of the rate-distortion function is not straightforward. We evaluate the rate-distortion function for an independent identically distributed gamma source with respect to the absolute-log distortion measure. The logarithmic transformation reduces this rate-distortion problem to that under the absolute distortion measure. Extending the explicit evaluation of the rate-distortion function for the Gaussian sources, we obtain the parametric form of the rate-distortion function. We show that the optimal distribution of reconstruction consists of a continuous component enclosed by left and right discrete components, and the left discrete component vanishes when the acceptable distortion is small. We further extend the result for a wider class of source distributions.
Kazuho Watanabe, Shiro Ikeda
IEEE Trans. Inf. Theory1
2015 Biclustering multivariate data for correlated subspace mining
abstract
Exploring feature subspaces is one of promising approaches to analyzing and understanding the important patterns in multivariate data. If relying too much on effective enhancements in manual interventions, the associated results depend heavily on the knowledge and skills of users performing the data analysis. This paper presents a novel approach to extracting feature subspaces from multivariate data by incorporating biclustering techniques. The approach has been maximally automated in the sense that highly-correlated dimensions are automatically grouped to form subspaces, which effectively supports further exploration of them. A key idea behind our approach lies in a new mathematical formulation of asymmetric biclustering, by combining spherical k-means clustering for grouping highly-correlated dimensions, together with ordinary k-means clustering for identifying subsets of data samples. Lower-dimensional representations of data in feature subspaces are successfully visualized by parallel coordinate plot, where we project the data samples of correlated dimensions to one composite axis through dimensionality reduction schemes. Several experimental results of our data analysis together with discussions will be provided to assess the capability of our approach.
Kazuho Watanabe, Hsiang-Yun Wu, Yusuke Niibe, Shigeo Takahashi, Issei Fujishiro
PacificVis1
2015 Vector quantization based on ε-insensitive mixture models
abstract
Laplacian mixture models have been used to deal with heavy-tailed distributions in data modeling problems. We consider an extension of Laplacian mixture models, which consists of ε-insensitive component distributions. An EM-type learning algorithm is derived for the maximum likelihood estimation of the proposed mixture model. The E-step is formulated in the usual way, while the M-step is formulated as the dual optimization problem instead of the primal optimization problem. Additionally, the convergence proof for ε=0 is accomplished. As an analogy to the k-means algorithm, we obtain what we call the ei-means algorithm in a certain limit of the learning algorithm. The derived algorithm is applied to approximate computation of rate-distortion functions associated with the ε-insensitive loss function. Then, it is demonstrated by synthetic data and real-world Spambase data that with appropriate selection of the ε value, the model is able to tolerate small percentage of noisy data.
Kazuho Watanabe
Neurocomputing1
2015 Achievability of asymptotic minimax regret by horizon-dependent and horizon-independent strategies
Kazuho Watanabe, Teemu Roos
J. Mach. Learn. Res.1
2015 Entropic risk minimization for nonparametric estimation of mixing distributions
abstract
We discuss a nonparametric estimation method for the mixing distributions in mixture models. The problem is formalized as a minimization of a one-parameter objective functional, which becomes the maximum likelihood estimation or the kernel vector quantization in special cases. Generalizing the theorem for the nonparametric maximum likelihood estimation, we prove the existence and discreteness of the optimal mixing distribution and provide an algorithm to calculate it. It is demonstrated that with an appropriate choice of the parameter, the proposed method is less prone to overfitting than the maximum likelihood method. We further discuss the connection between the unifying estimation framework and the rate-distortion problem.
Kazuho Watanabe, Shiro Ikeda
Mach. Learn.1
2015 Variational Bayesian Inference Algorithms for Infinite Relational Model of Network Data
abstract
Network data show the relationship among one kind of objects, such as social networks and hyperlinks on the Web. Many statistical models have been proposed for analyzing these data. For modeling cluster structures of networks, the infinite relational model (IRM) was proposed as a Bayesian nonparametric extension of the stochastic block model. In this brief, we derive the inference algorithms for the IRM of network data based on the variational Bayesian (VB) inference methods. After showing the standard VB inference, we derive the collapsed VB (CVB) inference and its variant called the zeroth-order CVB inference. We compared the performances of the inference algorithms using six real network datasets. The CVB inference outperformed the VB inference in most of the datasets, and the differences were especially larger in dense networks.
Takuya Konishi, Takatomi Kubo, Kazuho Watanabe, Kazushi Ikeda
IEEE Trans. Neural Networks Learn. Syst.3
2015 Variational Inference With ARD Prior for NIRS Diffuse Optical Tomography
abstract
Diffuse optical tomography (DOT) reconstructs 3-D tomographic images of brain activities from observations by near-infrared spectroscopy (NIRS) that is formulated as an ill-posed inverse problem. This brief presents a method for NIRS DOT based on a hierarchical Bayesian approach introducing the automatic relevance determination prior and the variational Bayes technique. Although the sparseness of the estimation strongly depends on the hyperparameters, in general, our method has less dependency on the hyperparameters. We confirm through numerical experiments that a schematic phase diagram of sparseness with respect to the hyperparameters has two regions: in one region hyperparameters give sparse solutions and in the other they give dense ones. The experimental results are supported by our theoretical analyses in simple cases.
Atsushi Miyamoto, Kazuho Watanabe, Kazushi Ikeda, Masa-aki Sato
IEEE Trans. Neural Networks Learn. Syst.2
2014 Bayesian properties of normalized maximum likelihood and its fast computation
abstract
The normalized maximized likelihood (NML) provides the minimax regret solution in universal data compression, gambling, and prediction, and it plays an essential role in the minimum description length (MDL) method of statistical modeling and estimation. Here we show that when the sample space is finite, a generic condition on the linear independence of the component models implies that the normalized maximum likelihood has an exact Bayes-like representation as a mixture of the component models, even in finite samples, though the weights of linear combination may be both positive and negative. This addresses in part the relationship between MDL and Bayes modeling. The representation also has the practical advantage of speeding the calculation of marginals and conditionals required for coding and prediction applications.
Andrew R. Barron, Teemu Roos, Kazuho Watanabe
ISIT3
2014 Spectral-Based Contractible Parallel Coordinates
abstract
Parallel coordinates is well-known as a popular tool for visualizing the underlying relationships among variables in high-dimension datasets. However, this representation still suffers from visual clutter arising from intersections among poly line plots especially when the number of data samples and their associated dimension become high. This paper presents a method of alleviating such visual clutter by contracting multiple axes through the analysis of correlation between every pair of variables. In this method, we first construct a graph by connecting axis nodes with an edge weighted by data correlation between the corresponding pair of dimensions, and then reorder the multiple axes by projecting the nodes onto the primary axis obtained through the spectral graph analysis. This allows us to compose a dendrogram tree by recursively merging a pair of the closest axes one by one. Our visualization platform helps the visual interpretation of such axis contraction by plotting the principal component of each data sample along the composite axis. Smooth animation of the associated axis contraction and expansion has also been implemented to enhance the visual readability of behavior inherent in the given high-dimensional datasets.
Koto Nohno, Hsiang-Yun Wu, Kazuho Watanabe, Shigeo Takahashi, Issei Fujishiro
IV3
2014 Analysis of Variational Bayesian Latent Dirichlet Allocation: Weaker Sparsity Than MAP
Shinichi Nakajima, Issei Sato, Masashi Sugiyama, Kazuho Watanabe, Hiroko Kobayashi
NIPS4
2013 Achievability of Asymptotic Minimax Regret in Online and Batch Prediction
abstract
The normalized maximum likelihood model achieves the minimax coding (log-loss) regret for data of fixed sample size n. However, it is a batch strategy, i.e., it requires that n be known in advance. Furthermore, it is computationally infeasible for most statistical models, and several computationally feasible alternative strategies have been devised. We characterize the achievability of asymptotic minimaxity by batch strategies (i.e., strategies that depend on n) as well as online strategies (i.e., strategies independent of n). On one hand, we conjecture that for a large class of models, no online strategy can be asymptotically minimax. We prove that this holds under a slightly stronger definition of asymptotic minimaxity. Our numerical experiments support the conjecture about non-achievability by so called last-step minimax algorithms, which are independent of n. On the other hand, we show that in the multinomial model, a Bayes mixture defined by the conjugate Dirichlet prior with a simple dependency on n achieves asymptotic minimaxity for all sequences, thus providing a simpler asymptotic minimax strategy compared to earlier work by Xie and Barron. The numerical results also demonstrate superior finite-sample behavior by a number of novel batch and online algorithms.
Kazuho Watanabe, Teemu Roos, Petri Myllymäki
ACML1
2013 Vector Quantization Using Mixture of Epsilon-Insensitive Components
Kazuho Watanabe
ICONIP (3)1
2013 Rate-distortion function for gamma sources under absolute-log distortion measure
abstract
We evaluate the rate-distortion function for the i.i.d. gamma sources with respect to the absolute-log distortion measure. The logarithmic transformation reduces this rate-distortion problem to that under the absolute distortion measure. Extending the explicit evaluation of the rate-distortion function for the Gaussian sources, we obtain the parametric form of the rate-distortion function. We show that the optimal distribution of reconstruction consists of a continuous component enclosed by left and right discrete components and the left discrete component vanishes when the allowed distortion is small.
Kazuho Watanabe, Shiro Ikeda
ISIT1
2013 Rate-distortion bounds for an ε-insensitive distortion measure
abstract
Direct evaluation of the rate-distortion function has rarely been achieved when it is strictly greater than its Shannon lower bound. In this paper, we consider the ratedistortion function for the distortion measure defined by an ε-insensitive loss function. We first present the Shannon lower bound applicable to any source distribution with finite differential entropy. Then, focusing on the Laplacian and Gaussian sources, we prove that the rate-distortion functions of these sources are strictly greater than their Shannon lower bounds and obtain analytic upper bounds for the rate-distortion functions. Small distortion limit and numerical evaluation of the bounds suggest that the Shannon lower bound provides a good approximation to the rate-distortion function for the ε-insensitive distortion measure.
Kazuho Watanabe
ITW1
2012 An alternative view of variational Bayes and asymptotic approximations of free energy
Kazuho Watanabe
Mach. Learn.1
2011 Phase diagrams of a variational Bayesian approach with ARD prior in NIRS-DOT
abstract
Diffuse optical tomography is a method used to reconstruct tomographic images from brain activities observed by near-infrared spectroscopy. This is useful for brain-machine interface and is formulated as an ill-posed inverse problem. We apply a hierarchical Bayesian approach, automatic relevance determination (ARD) prior and the variational Bayes method, that can introduce localization into the estimation of the problem. Although ARD enables sparse estimation, it is still open how hyperparameters affect the sparseness and accuracy of the estimation. Through numerical experiments, we present a schematic phase diagram of sparseness with respect to the hyperparameters in the method, which indicates the region of the hyperparameters where sparse estimation is achievable.
Atsushi Miyamoto, Kazuho Watanabe, Kazushi Ikeda, Masa-aki Sato
IJCNN2
2011 Divergence measures and a general framework for local variational approximation
Kazuho Watanabe, Masato Okada, Kazushi Ikeda
Neural Networks1
2009 Upper bound for variational free energy of Bayesian networks
Kazuho Watanabe, Motoki Shiga, Sumio Watanabe
Mach. Learn.1
2009 Variational Bayesian Mixture Model on a Subspace of Exponential Family Distributions
abstract
Exponential principal component analysis (e-PCA) has been proposed to reduce the dimension of the parameters of probability distributions using Kullback information as a distance between two distributions. It also provides a framework for dealing with various data types such as binary and integer for which the Gaussian assumption on the data distribution is inappropriate. In this paper, we introduce a latent variable model for the e-PCA. Assuming the discrete distribution on the latent variable leads to mixture models with constraint on their parameters. This provides a framework for clustering on the lower dimensional subspace of exponential family distributions. We derive a learning algorithm for those mixture models based on the variational Bayes (VB) method. Although intractable integration is required to implement the algorithm for a subspace, an approximation technique using Laplace's method allows us to carry out clustering on an arbitrary subspace. Combined with the estimation of the subspace, the resulting algorithm performs simultaneous dimensionality reduction and clustering. Numerical experiments on synthetic and real data demonstrate its effectiveness for extracting the structures of data as a visualization technique and its high generalization ability as a density estimation model.
Kazuho Watanabe, Shotaro Akaho, Shinichiro Omachi, Masato Okada
IEEE Trans. Neural Networks1
2008 Firing Rate Estimation Using an Approximate Bayesian Method
Kazuho Watanabe, Masato Okada
ICONIP (1)1
2007 Stochastic complexities of general mixture models in variational Bayesian learning
Kazuho Watanabe, Sumio Watanabe
Neural Networks1
2007 Stochastic complexity for mixture of exponential families in generalized variational Bayes
Kazuho Watanabe, Sumio Watanabe
Theor. Comput. Sci.1
2006 Free Energy of Stochastic Context Free Grammar on Variational Bayes
Tikara Hosino, Kazuho Watanabe, Sumio Watanabe
ICONIP (1)2
2006 Upper Bounds for Variational Stochastic Complexities of Bayesian Networks
Kazuho Watanabe, Motoki Shiga, Sumio Watanabe
IDEAL1
2006 Stochastic Complexities of Gaussian Mixtures in Variational Bayesian Approximation
abstract
Bayesian learning has been widely used and proved to be effective in many data modeling problems. However, computations involved in it require huge costs and generally cannot be performed exactly. The variational Bayesian approach, proposed as an approximation of Bayesian learning, has provided computational tractability and good generalization performance in many applications. The properties and capabilities of variational Bayesian learning itself have not been clarified yet. It is still unknown how good approximation the variational Bayesian approach can achieve. In this paper, we discuss variational Bayesian learning of Gaussian mixture models and derive upper and lower bounds of variational stochastic complexities. The variational stochastic complexity, which corresponds to the minimum variational free energy and a lower bound of the Bayesian evidence, not only becomes important in addressing the model selection problem, but also enables us to discuss the accuracy of the variational Bayesian approach as an approximation of true Bayesian learning.
Kazuho Watanabe, Sumio Watanabe
J. Mach. Learn. Res.1
2005 Stochastic Complexity for Mixture of Exponential Families in Variational Bayes
Kazuho Watanabe, Sumio Watanabe
ALT1
2005 Stochastic complexity of variational Bayesian hidden Markov models
abstract
Variational Bayesian learning was proposed as the approximation method of Bayesian learning. Inspite of efficiency and experimental good performance, their mathematical property has not yet been clarified. In this paper we analyze variational Bayesian hidden Markov models which include the true one thus the models are non-identifiable. We derive their asymptotic stochastic complexity. It is shown that, in some prior condition, the stochastic complexity is much smaller than those of identifiable models.
Tikara Hosino, Kazuho Watanabe, Sumio Watanabe
IJCNN2
2005 Variational Bayesian Stochastic Complexity of Mixture Models
abstract
The Variational Bayesian framework has been widely used to approximate the Bayesian learning. In various applications, it has provided computational tractability and good generalization performance. In this paper, we discuss the Variational Bayesian learning of the mixture of exponential families and provide some additional theoretical support by deriving the asymptotic form of the stochastic complexity. The stochastic complexity, which corresponds to the minimum free energy and a lower bound of the marginal likelihood, is a key quantity for model selection. It also enables us to discuss the effect of hyperparameters and the accuracy of the Variational Bayesian approach as an approximation of the true Bayesian learning.
Kazuho Watanabe, Sumio Watanabe
NIPS1
2004 Estimation of the Data Region Using Extreme-Value Distributions
Kazuho Watanabe, Sumio Watanabe
ALT1