VLDB 2026 Research / reviewers in the wild / expert
Takeru Matsuda
dblp:194/4550
· DBLP profile ↗
12ranked-venue papers
6as first author
5since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 3 since 2021Theory of computation · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Exploring Intra and Inter-language Consistency in Embeddings with ICAabstractWord embeddings represent words as multidimensional real vectors, facilitating data analysis and processing, but are often challenging to interpret.Independent Component Analysis (ICA) creates clearer semantic axes by identifying independent key features.Previous research has shown ICA's potential to reveal universal semantic axes across languages.However, it lacked verification of the consistency of independent components within and across languages.We investigated the consistency of semantic axes in two ways: both within a single language and across multiple languages.We first probed into intra-language consistency, focusing on the reproducibility of axes by running the ICA algorithm multiple times and clustering the outcomes.Then, we statistically examined inter-language consistency by verifying those axes' correspondences using statistical tests.We newly applied statistical methods to establish a robust framework that ensures the reliability and universality of semantic axes. Rongzhi Li, Takeru Matsuda, Hitomi Yanaka |
EMNLP | 2 |
| 2024 | Adapting to General Quadratic Loss via Singular Value ShrinkageabstractThe Gaussian sequence model is a canonical model in nonparametric estimation. In this study, we introduce a multivariate version of the Gaussian sequence model and investigate adaptive estimation over the multivariate Sobolev ellipsoids, where adaptation is not only to unknown smoothness but also to arbitrary quadratic loss. First, we derive an oracle inequality for the singular value shrinkage estimator by Efron and Morris, which is a matrix generalization of the James–Stein estimator. Next, we develop an asymptotically minimax estimator on the multivariate Sobolev ellipsoid for each quadratic loss, which can be viewed as a generalization of Pinsker’s theorem. Then, we show that the blockwise Efron–Morris estimator is exactly adaptive minimax over the multivariate Sobolev ellipsoids under the corresponding quadratic loss. It attains sharp adaptive estimation of any linear combination of the mean sequences simultaneously. Takeru Matsuda |
IEEE Trans. Inf. Theory | 1 |
| 2022 | Oscillator decomposition of infant fNIRS dataabstractThe functional near-infrared spectroscopy (fNIRS) can detect hemodynamic responses in the brain and the data consist of bivariate time series of oxygenated hemoglobin (oxy-Hb) and deoxygenated hemoglobin (deoxy-Hb) on each channel. In this study, we investigate oscillatory changes in infant fNIRS signals by using the oscillator decompisition method (OSC-DECOMP), which is a statistical method for extracting oscillators from time series data based on Gaussian linear state space models. OSC-DECOMP provides a natural decomposition of fNIRS data into oscillation components in a data-driven manner and does not require the arbitrary selection of band-pass filters. We analyzed 18-ch fNIRS data (3 minutes) acquired from 21 sleeping 3-month-old infants. Five to seven oscillators were extracted on most channels, and their frequency distribution had three peaks in the vicinity of 0.01-0.1 Hz, 1.6-2.4 Hz and 3.6-4.4 Hz. The first peak was considered to reflect hemodynamic changes in response to the brain activity, and the phase difference between oxy-Hb and deoxy-Hb for the associated oscillators was at approximately 230 degrees. The second peak was attributed to cardiac pulse waves and mirroring noise. Although these oscillators have close frequencies, OSC-DECOMP can separate them through estimating their different projection patterns on oxy-Hb and deoxy-Hb. The third peak was regarded as the harmonic of the second peak. By comparing the Akaike Information Criterion (AIC) of two state space models, we determined that the time series of oxy-Hb and deoxy-Hb on each channel originate from common oscillatory activity. We also utilized the result of OSC-DECOMP to investigate the frequency-specific functional connectivity. Whereas the brain oscillator exhibited functional connectivity, the pulse waves and mirroring noise oscillators showed spatially homogeneous and independent changes. OSC-DECOMP is a promising tool for data-driven extraction of oscillation components from biological time series data. Takeru Matsuda, Fumitaka Homae, Hama Watanabe, Gentaro Taga, Fumiyasu Komaki |
PLoS Comput. Biol. | 1 |
| 2021 | Interpretable Stein Goodness-of-fit Tests on Riemannian ManifoldabstractIn many applications, we encounter data on Riemannian manifolds such as torus and rotation groups. Standard statistical procedures for multivariate data are not applicable to such data. In this study, we develop goodness-of-fit testing and interpretable model criticism methods for general distributions on Riemannian manifolds, including those with an intractable normalization constant. The proposed methods are based on extensions of kernel Stein discrepancy, which are derived from Stein operators on Riemannian manifolds. We discuss the connections between the proposed tests with existing ones and provide a theoretical analysis of their asymptotic Bahadur efficiency. Simulation results and real data applications show the validity and usefulness of the proposed methods. Takeru Matsuda |
ICML | 2 |
| 2021 | Information criteria for non-normalized modelsabstractMany statistical models are given in the form of non-normalized densities with an intractable normalization constant. Since maximum likelihood estimation is computationally intensive for these models, several estimation methods have been developed which do not require explicit computation of the normalization constant, such as noise contrastive estimation (NCE) and score matching. However, model selection methods for general nonnormalized models have not been proposed so far. In this study, we develop information criteria for non-normalized models estimated by NCE or score matching. They are approximately unbiased estimators of discrepancy measures for non-normalized models. Simulation results and applications to real data demonstrate that the proposed criteria enable selection of the appropriate non-normalized model in a data-driven manner. Takeru Matsuda, Masatoshi Uehara, Aapo Hyvärinen |
J. Mach. Learn. Res. | 1 |
| 2020 | A Unified Statistically Efficient Estimation Framework for Unnormalized ModelsabstractThe parameter estimation of unnormalized models is a challenging problem. The maximum likelihood estimation (MLE) is computationally infeasible for these models since normalizing constants are not explicitly calculated. Although some consistent estimators have been proposed earlier, the problem of statistical efficiency remains. In this study, we propose a unified, statistically efficient estimation framework for unnormalized models and several efficient estimators, whose asymptotic variance is the same as the MLE. The computational cost of these estimators is also reasonable and they can be employed whether the sample space is discrete or continuous. The loss functions of the proposed estimators are derived by combining the following two methods: (1) density-ratio matching using Bregman divergence, and (2) plugging-in nonparametric estimators. We also analyze the properties of the proposed estimators when the unnormalized models are misspecified. The experimental results demonstrate the advantages of our method over existing approaches. Masatoshi Uehara, Takafumi Kanamori, Takashi Takenouchi, Takeru Matsuda |
AISTATS | 4 |
| 2020 | Imputation estimators for unnormalized models with missing dataabstractSeveral statistical models are given in the form of unnormalized densities and calculation of the normalization constant is intractable. We propose estimation methods for such unnormalized models with missing data. The key concept is to combine imputation techniques with estimators for unnormalized models including noise contrastive estimation and score matching. Further, we derive asymptotic distributions of the proposed estimators and construct confidence intervals. Simulation results with truncated Gaussian graphical models and the application to real data of wind direction demonstrate that the proposed methods enable statistical inference from missing data properly. Masatoshi Uehara, Takeru Matsuda, Jae Kwang Kim |
AISTATS | 2 |
| 2020 | A Stein Goodness-of-fit Test for Directional DistributionsabstractIn many fields, data appears in the form of direction (unit vector) and usual statistical procedures are not applicable to such directional data. In this study, we propose nonparametric goodness-of-fit testing procedures for general directional distributions based on kernel Stein discrepancy. Our method is based on Stein’s operator on spheres, which is derived by using Stokes’ theorem. Notably, the proposed method is applicable to distributions with an intractable normalization constant, which commonly appear in directional statistics. Experimental results demonstrate that the proposed methods control type-I error well and have larger power than existing tests, including the test based on the maximum mean discrepancy. Takeru Matsuda |
AISTATS | 2 |
| 2019 | Estimation of Non-Normalized Mixture ModelsabstractWe develop a general method for estimating a finite mixture of non-normalized models. A non-normalized model is defined to be a parametric distribution with an intractable normalization constant. Existing methods for estimating non-normalized models without computing the normalization constant are not applicable to mixture models because they contain more than one intractable normalization constant. The proposed method is derived by extending noise contrastive estimation (NCE), which estimates non-normalized models by discriminating between the observed data and some artificially generated noise. In particular, the proposed method provides a probabilistically principled clustering method that is able to utilize a deep representation. Applications to clustering of natural images and neuroimaging data give promising results. Takeru Matsuda, Aapo Hyvärinen |
AISTATS | 1 |
| 2019 | Harmonic Bayesian Prediction Under $\alpha$ -DivergenceabstractWe investigate Bayesian shrinkage methods for constructing predictive distributions. We consider the multivariate normal model with a known covariance matrix and show that the Bayesian predictive density with respect to Stein's harmonic prior dominates the best invariant Bayesian predictive density when the dimension is greater than or equal to 3. Alpha divergence from the true distribution to a predictive distribution is adopted as a loss function. Yuzo Maruyama, Takeru Matsuda, Toshio Ohnishi |
IEEE Trans. Inf. Theory | 2 |
| 2017 | Time Series Decomposition into Oscillation Components and Phase EstimationabstractMany time series are naturally considered as a superposition of several oscillation components. For example, electroencephalogram (EEG) time series include oscillation components such as alpha, beta, and gamma. We propose a method for decomposing time series into such oscillation components using state-space models. Based on the concept of random frequency modulation, gaussian linear state-space models for oscillation components are developed. In this model, the frequency of an oscillator fluctuates by noise. Time series decomposition is accomplished by this model like the Bayesian seasonal adjustment method. Since the model parameters are estimated from data by the empirical Bayes' method, the amplitudes and the frequencies of oscillation components are determined in a data-driven manner. Also, the appropriate number of oscillation components is determined with the Akaike information criterion (AIC). In this way, the proposed method provides a natural decomposition of the given time series into oscillation components. In neuroscience, the phase of neural time series plays an important role in neural information processing. The proposed method can be used to estimate the phase of each oscillation component and has several advantages over a conventional method based on the Hilbert transform. Thus, the proposed method enables an investigation of the phase dynamics of time series. Numerical results show that the proposed method succeeds in extracting intermittent oscillations like ripples and detecting the phase reset phenomena. We apply the proposed method to real data from various fields such as astronomy, ecology, tidology, and neuroscience. Takeru Matsuda, Fumiyasu Komaki |
Neural Comput. | 1 |
| 2017 | Multivariate Time Series Decomposition into Oscillation ComponentsabstractMany time series are considered to be a superposition of several oscillation components. We have proposed a method for decomposing univariate time series into oscillation components and estimating their phases (Matsuda & Komaki, 2017 ). In this study, we extend that method to multivariate time series. We assume that several oscillators underlie the given multivariate time series and that each variable corresponds to a superposition of the projections of the oscillators. Thus, the oscillators superpose on each variable with amplitude and phase modulation. Based on this idea, we develop gaussian linear state-space models and use them to decompose the given multivariate time series. The model parameters are estimated from data using the empirical Bayes method, and the number of oscillators is determined using the Akaike information criterion. Therefore, the proposed method extracts underlying oscillators in a data-driven manner and enables investigation of phase dynamics in a given multivariate time series. Numerical results show the effectiveness of the proposed method. From monthly mean north-south sunspot number data, the proposed method reveals an interesting phase relationship. Takeru Matsuda, Fumiyasu Komaki |
Neural Comput. | 1 |