VLDB 2026 Research / reviewers in the wild / expert
Norihiro Takamune
dblp:148/9599
· DBLP profile ↗
21ranked-venue papers
0as first author
8since 2021 · last 2026
0000-0002-8102-3110ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 since 2021Artificial intelligence and machine learning · 8 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Stride conversion algorithms for convolutional layers and its application to sampling-frequency-independent deep neural networksabstractWe propose interpolation-based algorithms that enable convolutional and transposed convolutional layers to operate with arbitrary (including non-integer) strides. A primary motivation for the proposed algorithms is to maintain a consistent temporal resolution when adapting deep neural networks (DNNs) to different sampling frequencies (SFs). To handle untrained SFs, we previously introduced SF-independent (SFI) convolutional layers, which adjust kernel weights in accordance with the target SF. However, achieving full consistency across SFs also requires the proportional adjustment of the stride, which results in non-integer values in many practical cases. Conventional algorithms for convolutional layers cannot handle such strides directly, and commonly used approaches (e.g., stride rounding or signal resampling) lead to performance degradation. To solve this problem, we propose a feature-domain interpolation framework that constructs continuous-time representations of intermediate features. This enables sampling at arbitrary stride intervals without modifying the network architecture. Through music source separation experiments, we show that the proposed algorithms maintain a strong performance across a range of SFs, including those where the stride becomes non-integer. Our analysis reveals that the proposed algorithms are robust to the choice of interpolation method and are especially effective for sources containing pitched sounds. Kanami Imamura, Tomohiko Nakamura, Norihiro Takamune, Kohei Yatabe, Hiroshi Saruwatari |
Signal Process. | 3 |
| 2024 | Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
Kentaro Seki, Shinnosuke Takamichi, Norihiro Takamune, Yuki Saito 0001, Kanami Imamura, Hiroshi Saruwatari |
INTERSPEECH | 3 |
| 2023 | HumanDiffusion: diffusion model using perceptual gradients
Yota Ueda, Shinnosuke Takamichi, Yuki Saito 0001, Norihiro Takamune, Hiroshi Saruwatari |
INTERSPEECH | 4 |
| 2023 | PoP-IDLMA: Product-of-Prior Independent Deeply Learned Matrix Analysis for Multichannel Music Source SeparationabstractIndependent deeply learned matrix analysis (IDLMA) is a state-of-the-art determined audio source separation method based on pretrained deep neural networks (DNNs). Owing to the excellent expression power of DNNs, IDLMA can handle a wider range of sources than conventional source models such as nonegative matrix factorization (NMF). However, owing to its supervised nature, the separation performance of IDLMA often degrades in the presence of timbral mismatches between the training data and the to-be-separated data. In this paper, we propose two source models that encompass the NMF- and DNN-based source models by constructing a prior distribution of the source power spectrogram (product of priors: PoP) on the basis of the product-of-expert concept. Since the NMF-based source model works well for a fully blind situation, the proposed models can handle the timbral mismatch without losing the expression power of DNNs. By introducing the PoP-based source models into IDLMA, we propose IDLMA extensions (PoP-IDLMAs) and derive their efficient parameter estimation algorithms on the basis of the majorization–minimization algorithm. Experimental results demonstrated the effectiveness of the proposed PoP-IDLMAs and that the proposed models greatly improve the source power estimation in frequency bands above 500 Hz. Takuya Hasumi, Tomohiko Nakamura, Norihiro Takamune, Hiroshi Saruwatari, Daichi Kitamura, Yu Takahashi, Kazunobu Kondo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Deficient Basis Estimation of Noise Spatial Covariance Matrix for Rank-Constrained Spatial Covariance Matrix Estimation Method in Blind Speech ExtractionabstractRank-constrained spatial covariance matrix estimation (RCSCME) is a state-of-the-art blind speech extraction method applied to cases where one directional target speech and diffuse noise are mixed. In this paper, we proposed a new algorithmic extension of RCSCME. RCSCME complements a deficient one rank of the diffuse noise spatial covariance matrix, which cannot be estimated via preprocessing such as independent low-rank matrix analysis, and estimates the source model parameters simultaneously. In the conventional RC- SCME, a direction of the deficient basis is fixed in advance and only the scale is estimated; however, the candidate of this deficient basis is not unique in general. In the proposed RCSCM model, the deficient basis itself can be accurately estimated as a vector variable by solving a vector optimization problem. Also, we derive new update rules based on the EM algorithm. We confirm that the proposed method outperforms conventional methods under several noise conditions. Yuto Kondo, Yuki Kubo, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari |
ICASSP | 3 |
| 2021 | Joint-diagonalizability-constrained multichannel nonnegative matrix factorization based on time-variant multivariate complex sub-Gaussian distributionabstractMultichannel nonnegative matrix factorization (MNMF) is a common blind source separation technique that employs full-rank spatial covariance matrices (SCMs). The full-rank SCMs can simulate reverberant mixing systems where the sources are spatially spread. In conventional MNMF, spectrograms of observed signals are modeled by some types of distribution, e.g., the Gaussian distribution and Student’s t distribution. However, MNMF based on the sub-Gaussian distribution has not been proposed because its cost function is difficult to minimize. In this paper, we address the statistical model extension of MNMF to the sub-Gaussian distribution to improve the source separation accuracy. In the proposed method, the generalized Gaussian distribution is utilized as the sub-Gaussian model. Moreover, to design an auxiliary function for the proposed cost function, we introduce the joint-diagonalizability constraint to SCMs similarly to FastMNMF. Two types of update rule for the proposed MNMF are derived on the basis of the majorization-minimization (MM) and majorization-equalization (ME) algorithms. Since the optimization speed of each parameter affects the source separation performance, we experimentally analyze the best combination of MM- and ME-algorithm-based update rules in the proposed method. Experiments of blind source separation reveal that the proposed MNMF based on the sub-Gaussian model can outperform conventional methods. Keigo Kamo, Yoshiki Mitsui, Yuki Kubo, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo |
Signal Process. | 4 |
| 2021 | Independent deeply learned matrix analysis with automatic selection of stable microphone-wise update and fast sourcewise update of demixing matrixabstractIndependent deeply learned matrix analysis (IDLMA) is a fast and high-performance method for multichannel audio source separation. IDLMA utilizes the deep neural network inference of source models and the blind estimation of demixing filters based on source independence. In conventional IDLMA, iterative projection (IP) is exploited to estimate the demixing filters. Although IP is a fast algorithm, it sometimes fails to estimate an appropriate solution. This is because IP updates the demixing filters in a sourcewise manner, where only one source model is used for each update, and the update sometimes becomes unstable owing to the specific low-quality source models. In this paper, we first derive a new numerically stable microphone-wise update algorithm that exploits all source model information simultaneously. The microphone-wise update problem cannot be solved by IP; instead, a new type of vectorwise coordinate descent algorithm is introduced. Next, comparison analysis of the proposed microphone-wise update and IP reveals the tradeoff w.r.t. convergence speed and numerical stability. To resolve this tradeoff problem, we propose the automatic selection of update rules on the basis of the likelihood function of observed signals. Finally, experimental results show the efficacy of the proposed IDLMA with the automatic selection of update rules. Naoki Makishima, Yoshiki Mitsui, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo |
Signal Process. | 3 |
| 2021 | Multichannel Blind Source Separation Based on Evanescent-Region-Aware Non-Negative Tensor Factorization in Spherical Harmonic DomainabstractThere is growing interest in new audio formats in the context of virtual reality (VR), and higher-order ambisonics (HOA) is preferred for VR systems to transmit recorded scenes owing to its transmission efficiency and its flexibility to work with different loudspeaker setups. However, the conversion between another well-known format, i.e., object format, and the HOA format is not fully addressed in the literature. To address this issue, blind source separation in a spherical harmonic (SH) domain can be considered as the best way to extract objects in terms of efficiency, i.e., decoding HOA signals for separation can be omitted. A few authors attempted to extract objects from encoded HOA signals directly by using multichannel non-negative matrix factorization (MNMF), but these approaches either assume only far-field sources or do not take array characteristics into account, which make these methods difficult to use for VR in practical situations where singers or speakers often perform close to microphones. Furthermore, MNMF generally requires a huge computational cost, although dimensional reduction to the SH domain is performed. In this work, we also model near-field sources by estimating the model parameters of non-negative tensor factorization (NTF) in the SH domain assuming that microphone signals can be obtained with a rigid spherical array. We propose a masking scheme to exclude noisy evanescent regions in the SH domain from the NTF cost function. Evaluations show that our method outperforms existing methods devised for the HOA format and that our masking approach is effective in improving the separation quality. Yuki Mitsufuji, Norihiro Takamune, Shoichi Koyama, Hiroshi Saruwatari |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Regularized Fast Multichannel Nonnegative Matrix Factorization with ILRMA-Based Prior Distribution of Joint-Diagonalization ProcessabstractIn this paper, we address a convolutive blind source separation (BSS) problem and propose a new extended framework of FastMNMF by introducing prior information for joint diagonalization of the spatial covariance matrix model. Recently, FastMNMF has been proposed as a fast version of multichannel nonnegative matrix factorization under the assumption that the spatial covariance matrices of multiple sources can be jointly diagonalized. However, its source-separation performance was not improved and the physical meaning of the joint-diagonalization process was unclear. To resolve these problems, we first reveal a close relationship between the joint-diagonalization process and the demixing system used in independent low-rank matrix analysis (ILRMA). Next, motivated by this fact, we propose a new regularized FastMNMF supported by ILRMA and derive convergence-guaranteed parameter update rules. From BSS experiments, we show that the proposed method outperforms the conventional FastMNMF in source-separation accuracy with almost the same computation time. Keigo Kamo, Yuki Kubo, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo |
ICASSP | 3 |
| 2020 | Convergence-Guaranteed Independent Positive Semidefinite Tensor Analysis Based on Student's T DistributionabstractIn this paper, we address a blind source separation (BSS) problem and propose a new extended framework of independent positive semidefinite tensor analysis (IPSDTA). IPSDTA is a state-of-the-art BSS method that enables us to take interfrequency correlations into account, but the generative model is limited within the multivariate Gaussian distribution and its parameter optimization algorithm does not guarantee stable convergence. To resolve these problems, first, we propose to extend the generative model to a parametric multivariate Student’s t distribution that can deal with various types of signal. Secondly, we derive a new parameter optimization algorithm that guarantees the monotonic nonincrease in the cost function, providing stable convergence. Experimental results reveal that the cost function in the conventional IPSDTA does not display monotonically nonincreasing properties. On the other hand, the proposed method guarantees the monotonic nonincrease in the cost function and outperforms the conventional ILRMA and IPSDTA in the source-separation performance. Tatsuki Kondo, Kanta Fukushige, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Rintaro Ikeshita, Tomohiro Nakatani |
ICASSP | 3 |
| 2020 | Phase reconstruction from amplitude spectrograms based on directional-statistics deep neural networksabstractThis paper presents a deep neural network (DNN)-based phase reconstruction method from amplitude spectrograms. In speech processing, an amplitude spectrogram is often used for processing, and the corresponding phases are reconstructed from the amplitude spectrogram by using the Griffin-Lim method. However, the Griffin-Lim method causes unnatural artifacts in synthetic speech. To solve this problem, we propose the directional-statistics DNNs for predicting phases from the amplitude spectrograms. We first propose the von Mises distribution DNN, which is a generative model having the von Mises distribution and models histograms of a periodic variable. We extend it for modeling group delay that has a stronger connection to the amplitude spectrograms. Furthermore, we generalize the group-delay modeling and propose another DNN called the sine-skewed generalized cardioid distribution DNN for modeling asymmetric histograms such as a group delay. Results from objective and subjective evaluations indicate that (1) our von Mises distribution DNN can predict group delay more accurately than predicting phases, (2) our DNN works as better initialization of the Griffin-Lim method, (3) the phase reconstruction methods based on our von Mises distribution DNN achieve better speech quality than the conventional Griffin-Lim method, and (4) our sine-skewed generalized cardioid distribution DNN models the group delay more accurately than our von Mises distribution DNN. Shinnosuke Takamichi, Yuki Saito 0001, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari |
Signal Process. | 3 |
| 2020 | Acoustic model-based subword tokenization and prosodic-context extraction without language knowledge for text-to-speech synthesisabstractThis paper presents text tokenization and context extraction without using language knowledge for text-to-speech (TTS) synthesis. To generate prosody, statistical parametric TTS synthesis typically requires the professional knowledge of the target language. Therefore, languages suitable for TTS synthesis are limited to only rich-resource languages. To achieve TTS synthesis without using language knowledge, we propose acoustic model-based subword tokenization and unsupervised extraction of prosodic contexts. The subword tokenization can determine language units suitable for prosody generation. The context extraction can retrieve contexts from pairs of subwords and prosody. The proposed methods function without language knowledge and can improve F0 prediction accuracy. Experimental evaluation demonstrates that 1) the training of proposed subword tokenization, which uses the expectation-maximization algorithm and deep neural networks, is empirically stable, 2) the proposed subword tokenization tokenizes text into subwords that are close to language-specific units, and 3) the proposed methods outperform the conventional methods using language model-based tokenization in terms of synthetic speech quality. Masashi Aso, Shinnosuke Takamichi, Norihiro Takamune, Hiroshi Saruwatari |
Speech Commun. | 3 |
| 2020 | Blind Speech Extraction Based on Rank-Constrained Spatial Covariance Matrix Estimation With Multivariate Generalized Gaussian DistributionabstractIn this article, we propose a new blind speech extraction (BSE) method that robustly extracts a directional speech from background diffuse noise by combining independent low-rank matrix analysis (ILRMA) and efficient rank-constrained spatial covariance matrix (SCM) estimation. To achieve more accurate BSE than ILRMA, which assumes each source to be a point source (rank-1 spatial model), the proposed method restores the lost spatial basis for the full-rank SCM of diffuse noise. We adopt the multivariate complex generalized Gaussian distribution (GGD) as the statistical generative model to express various types of observed signal. To estimate the model parameters for an arbitrary shape parameter of the multivariate GGD, we derive a new inequality for rank-constrained SCMs. Also, we propose new acceleration methods to accomplish much faster extraction than conventional blind source separation methods. In BSE experiments using simulated and real recorded data, we confirm that the proposed method achieves more accurate and faster speech extraction than conventional methods. Yuki Kubo, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Multichannel Non-Negative Matrix Factorization Using Banded Spatial Covariance Matrices in Wavenumber DomainabstractBlind source separation exploiting multichannel information has long been a popular topic, and recently proposed methods based on the local Gaussian model have shown promising results despite its high computational cost for the case of many microphone signals. The low updating speed for such a model is mainly due to the inversion of a spatial covariance matrix, for which the complexity increases with the number of microphones, M, and is generally of order O(M3). Several projection-based approaches that attempt to concentrate energy on the diagonal part of the spatial covariance matrix have been introduced to circumvent the matrix inversion, which can reduce the complexity to O(M). In this article, we focus on the fast Fourier transform as a projection method because the energy concentration on the diagonal can be efficiently achieved compared with other projection-based methods. For the case where the diagonalization is imperfect, for example, owing to discontinuities at the edge of a linear array, we also developed a more robust algorithm approximating the tri-diagonal part of the spatial covariance matrix, which requires a complexity of O(M2) for the inversion by applying the Thomas algorithm. To remove the ad-hoc integration of post clustering after the decomposition, we also examine a self-clustering algorithm. Our evaluation shows better results than other previously proposed methods in terms of the separation quality under reverberant conditions as well as higher efficiency than multichannel non-negative matrix factorization. Yuki Mitsufuji, Stefan Uhlich, Norihiro Takamune, Daichi Kitamura, Shoichi Koyama, Hiroshi Saruwatari |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Independent Low-Rank Matrix Analysis Based on Time-Variant Sub-Gaussian Source Model for Determined Blind Source SeparationabstractIndependent low-rank matrix analysis (ILRMA) is a fast and stable method of blind audio source separation. Conventional ILRMAs assume time-variant (super-)Gaussian source models, which can only represent signals that follow a super-Gaussian distribution. In this article, we focus on ILRMA based on a generalized Gaussian distribution (GGD-ILRMA) and propose a new type of GGD-ILRMA that adopts a time-variant sub-Gaussian distribution for the source model. We propose a new update scheme called generalized iterative projection for homogeneous source models (GIP-HSM) and obtain a convergence-guaranteed update rule for demixing spatial parameters by combining the GIP-HSM scheme and the majorization-minimization (MM) algorithm. Furthermore, a new extension of the MM algorithm is proposed for the convergence acceleration by applying the majorization-equalization algorithm to a multivariate case. In the experimental evaluation, we show the versatility of the proposed method, i.e., the proposed time-variant sub-Gaussian source model can be applied to various types of source signal. Shinichi Mogami, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo, Nobutaka Ono |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Bilevel Optimization Using Stationary Point of Lower-Level Objective Function for Discriminative Basis Learning in Nonnegative Matrix FactorizationabstractIn this letter, we address an audio signal separation problem and propose a new effective algorithm for solving a bilevel optimization in discriminative nonnegative matrix factorization (NMF). Recently, discriminative training of NMF bases has been developed for better signal separation in supervised NMF (SNMF), which exploits a priori training of given sample signals. The optimization in this method consists of a simultaneous minimization of two objective functions, resulting in a bilevel optimization problem with SNMF (BiSNMF), where conventional methods approximately solve this optimization. To strictly solve BiSNMF, we introduce a new algorithm with the following two features: (a) conversion of the optimization constraint into a penalty term and (b) optimization of the reformulated problem on the basis of a multiplicative steepest descent, ensuring the nonnegativity of variables. Experiments on music signal separation show the efficacy of the proposed algorithm. Hiroaki Nakajima, Daichi Kitamura, Norihiro Takamune, Hiroshi Saruwatari, Nobutaka Ono |
IEEE Signal Process. Lett. | 3 |
| 2019 | Independent Deeply Learned Matrix Analysis for Determined Audio Source SeparationabstractIn this paper, we propose a new framework called independent deeply learned matrix analysis (IDLMA), which unifies a deep neural network (DNN) and independence-based multichannel audio source separation. IDLMA utilizes both pretrained DNN source models and statistical independence between sources for the separation, where the time-frequency structures of each source are iteratively optimized by a DNN while enhancing the estimation accuracy of the spatial demixing filters. As the source generative model, we introduce a complex heavy-tailed distribution to improve the separation performance. In addition, we address a semi-supervised situation; namely, a solo-recorded audio dataset can be prepared for only one source in the mixture signal. To solve the limited-data problem, we propose an appropriate data augmentation method to adapt the DNN source models to the observed signal, which enables IDLMA to work even in the semi-supervised situation. Experiments are conducted using music signals with a training dataset in both supervised and semi-supervised situations. The results show the validity of the proposed method in terms of the separation accuracy. Naoki Makishima, Shinichi Mogami, Norihiro Takamune, Daichi Kitamura, Hayato Sumino, Shinnosuke Takamichi, Hiroshi Saruwatari, Nobutaka Ono |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Vectorwise Coordinate Descent Algorithm for Spatially Regularized Independent Low-Rank Matrix AnalysisabstractAudio source separation is an important problem for many audio applications. Independent low-rank matrix analysis (ILRMA) is a recently proposed algorithm that employs the statistical independence between sources and the low-rankness of the time-frequency structure in each source. As reported in this paper, we have developed a new framework that enables us to introduce a spatial regularization of the demixing matrix in ILRMA. Since the conventional optimization cannot be applied to this regularized ILRMA, we derive a novel approach based on vectorwise coordinate descent, which does not require a step-size parameter and guarantees convergence. In experiments, ILRMA with beamforming-based regularization is evaluated as an application of the proposed framework. Yoshiki Mitsui, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Yu Takahashi, Kazunobu Kondo |
ICASSP | 2 |
| 2017 | Spatio-temporal sparse sound field decomposition considering acoustic source signal characteristicsabstractWe propose a sound field decomposition method that takes into consideration spatio-temporal sparsity. It has been proved that sparse representation of a sound field is effective in reducing errors originating from spatial aliasing artifacts compared with conventional plane wave decomposition. In most current methods of sparse sound field decomposition, the spatial sparsity of the sound source distribution is only assumed. However, it is known that the temporal structure of the source signal to be decomposed can also be sparse in the time-frequency domain. We formulate an objective function for sparse sound field decomposition by using the ℓp,q-norm to simultaneously induce sparsity in the space and time domains. An optimization algorithm on the auxiliary function method is derived to solve it. Numerical simulations of acoustic holography indicate that the reconstruction accuracy can be improved by controlling the parameter of temporal sparsity. We also demonstrate that a statistical measure of the source signals can be used as an indicator to determine a nearly optimal parameter. Naoki Murata, Shoichi Koyama, Norihiro Takamune, Hiroshi Saruwatari |
ICASSP | 3 |
| 2016 | Sparse sound field decomposition with multichannel extension of complex NMFabstractA sparse sound field decomposition method using prior information on source signals in the time-frequency domain is proposed. Sparse sound field decomposition has been proved to be effective for various acoustic signal processing applications. Current methods for sparse decomposition are based only on the spatial sparsity of the source distribution. However, it can be assumed that possible source signals to be decomposed are approximately known in advance. To exploit this prior information, we incorporated the complex nonnegative factorization model into sparse sound field decomposition. Since the magnitude spectrum of the possible source signals can be trained in advance, accuracy of the sparse decomposition can be improved even when the source signals are highly correlated and the sources are in a highly noisy environment. In addition, the proposed decomposition algorithm is derived using the auxiliary function method. Numerical experiments indicated that the sparse decomposition performance was significantly improved using the proposed method. Naoki Murata, Shoichi Koyama, Hirokazu Kameoka, Norihiro Takamune, Hiroshi Saruwatari |
ICASSP | 4 |
| 2014 | Underdetermined blind separation and tracking of moving sources based ONDOA-HMMabstractThis paper deals with the problem of the underdetermined blind separation and tracking of moving sources. In practical situations, sound sources such as human speakers can move freely and so blind separation algorithms must be designed to track the temporal changes of the impulse responses. We propose solving this problem through the posterior inference of the parameters in a generative model of an observed multichannel signal, formulated under the assumption of the sparsity of time-frequency components of speech and the continuity of speakers' movements. Specifically, we describe a generative model of mixture signals by incorporating a generative model of a time-varying frequency array response for each source, described using a path-restricted hidden Markov model (HMM). Each hidden state of the present HMM represents the direction of arrival (DOA) of each source, and so we call it a “DOA-HMM.” Through the posterior inference of the overall generative model, we can simultaneously track the DOAs of sources, separate source signals and perform permutation alignment. The experiment showed that the proposed algorithm provided a 6.20 dB improvement compared with the conventional method in terms of the signal-to-interference ratio. Takuya Higuchi, Norihiro Takamune, Tomohiko Nakamura, Hirokazu Kameoka |
ICASSP | 2 |