VLDB 2026 Research / reviewers in the wild / expert
Rintaro Ikeshita
dblp:175/8983
· DBLP profile ↗
27ranked-venue papers
5as first author
18since 2021 · last 2026
0000-0003-2608-1999ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 10 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Microphone array geometry-independent multi-talker distant ASR: NTT system for DASR task of the CHiME-8 challenge
Naoyuki Kamo, Naohiro Tawara, Atsushi Ando, Takatomo Kano, Hiroshi Sato 0002, Rintaro Ikeshita, Takafumi Moriya, Shota Horiguchi, Kohei Matsuura, Atsunori Ogawa, Alexis Plaquet, Takanori Ashihara, Tsubasa Ochiai, Masato Mimura, Marc Delcroix, Tomohiro Nakatani, Taichi Asami, Shoko Araki |
Comput. Speech Lang. | 6 |
| 2024 | How Does End-To-End Speech Recognition Training Impact Speech Enhancement Artifacts?abstractJointly training a speech enhancement (SE) front-end and an automatic speech recognition (ASR) back-end has been investigated as a way to mitigate the influence of processing distortion generated by single-channel SE on ASR. In this paper, we investigate the effect of such joint training on the signal-level characteristics of the enhanced signals from the viewpoint of the decomposed noise and artifact errors. The experimental analyses provide two novel findings: 1) ASR-level training of the SE front-end reduces the artifact errors while increasing the noise errors, and 2) simply interpolating the enhanced and observed signals, which achieves a similar effect of reducing artifacts and increasing noise, improves ASR performance without jointly modifying the SE and ASR modules, even for a strong ASR back-end using a WavLM feature extractor. Our findings provide a better understanding of the effect of joint training and a novel insight for designing an ASR agnostic SE front-end. Kazuma Iwamoto, Tsubasa Ochiai, Marc Delcroix, Rintaro Ikeshita, Hiroshi Sato 0002, Shoko Araki, Shigeru Katagiri |
ICASSP | 4 |
| 2024 | Neural Network-Based Virtual Microphone Estimation with Virtual Microphone and Beamformer-Level Multi-Task LossabstractArray processing performance depends on the number of microphones available. Virtual microphone estimation (VME) has been proposed to increase the number of microphone signals artificially. Neural network-based VME (NN-VME) trains an NN with a VM-level loss to predict a signal at a microphone location that is available during training but not at inference. However, this training objective may not be optimal for a specific array processing back-end, such as beamforming. An alternative approach is to use a training objective considering the array-processing back-end, such as a loss on the beamformer output. This approach may generate signals optimal for beamforming but not physically grounded. To combine the advantages of both approaches, this paper proposes a multi-task loss for NN-VME that combines both VM-level and beamformer-level losses. We evaluate the proposed multi-task NN-VME on multi-talker underdetermined conditions and show that it achieves a 33.1 % relative WER improvement compared to using only real microphones and 10.8 % compared to using a prior NN-VME approach. Hanako Segawa, Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Rintaro Ikeshita, Shoko Araki, Takeshi Yamada, Shoji Makino |
ICASSP | 5 |
| 2024 | Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition PerformanceabstractIt is challenging to improve automatic speech recognition (ASR) performance in noisy conditions with a single-channel speech enhancement (SE) front-end. This is generally attributed to the processing distortions caused by the nonlinear processing of single-channel SE front-ends. However, the causes of such degraded ASR performance have not been fully investigated. How to design single-channel SE front-ends in a way that significantly improves ASR performance remains an open research question. In this study, we investigate a signal-level numerical metric that can explain the cause of degradation in ASR performance. To this end, we propose a novel analysis scheme based on the orthogonal projection-based decomposition of SE errors. This scheme manually modifies the ratio of the decomposed interference, noise, and artifact errors, and it enables us to directly evaluate the impact of each error type on ASR performance. Our analysis reveals the particularly detrimental effect of artifact errors on ASR performance compared to the other types of errors. This provides us with a more principled definition of processing distortions that cause the ASR performance degradation. Then, we study two practical approaches for reducing the impact of artifact errors. First, we prove that the simple observation adding (OA) post-processing (i.e., interpolating the enhanced and observed signals) can improve the signal-to-artifact ratio. Second, we propose a novel training objective, called artifact-boosted signal-to-distortion ratio (AB-SDR), which forces the model to estimate the enhanced signals with fewer artifact errors. Through experiments, we confirm that both the OA and AB-SDR approaches are effective in decreasing artifact errors caused by single-channel SE front-ends, allowing them to significantly improve ASR performance. Tsubasa Ochiai, Kazuma Iwamoto, Marc Delcroix, Rintaro Ikeshita, Hiroshi Sato 0002, Shoko Araki, Shigeru Katagiri |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Blind and Spatially-Regularized Online Joint Optimization of Source Separation, Dereverberation, and Noise ReductionabstractThis paper proposes a computationally efficient joint optimization algorithm that performs online source separation, dereverberation, and noise reduction based on blind and spatially-regularized processing. When applying such online Blind Source Separation (BSS) as online Independent Vector Extraction (IVE) to a speech application, we must focus on the trade-off between the algorithmic delay and separation accuracy, both of which depend on the analysis frame length. In addition, to separate the sources with specified source permutation, researchers introduced spatial regularization based on the Directions-of-Arrival (DOAs) of the sources into IVE. However, the scale ambiguity of IVE often makes the spatial regularization work inappropriately. To solve these problems, we first propose a blind online joint optimization algorithm of IVE and weighted prediction error dereverberation (WPE). This online algorithm can achieve accurate separation even using short analysis frames because reverberation can be reduced using WPE. We then extend the online joint optimization with robust spatial regularization. We reveal that regularizing the scale of the separated signals is very effective in making the DOA-based spatial regularization work reliably. Our experiments confirm that our blind online joint optimization algorithm can significantly improve the separation accuracy with an algorithmic delay of 8 ms. In addition, we confirm that the proposed spatially-regularized online joint optimization algorithm reduces the rate of the source permutation error to zero percent. Tetsuya Ueda, Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Shoko Araki, Shoji Makino |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Fast Online Source Steering Algorithm for Tracking Single Moving Source Using Online Independent Vector AnalysisabstractWe address the problem of separating moving sources using online independent vector analysis (IVA). To solve this problem, researchers have extended the iterative projection (IP) and iterative source steering (ISS) algorithms developed for batch auxiliary-function-based IVA (AuxIVA) to online scenarios and showed their effectiveness. However, the conventional online IP and ISS are slow because they update K × K covariance matrices for all sources, where K is the number of microphones. Here, we show that, in a target-source tracking scenario in which only one source moves, there exists an inexpensive formula for online ISS that avoids updating the full covariance matrices without changing the behavior of the algorithm. The time complexity of the proposed algorithm, which we call online source steering (OSS), is K times smaller than that of the conventional online IP and ISS for the target-source tracking task. A numerical experiment on separating a moving source demonstrates that the proposed OSS is significantly faster than the conventional online IP and ISS. Taishi Nakashima, Rintaro Ikeshita, Nobutaka Ono, Shoko Araki, Tomohiro Nakatani |
ICASSP | 2 |
| 2023 | Multi-Frame Full-Rank Spatial Covariance Analysis for Underdetermined Blind Source Separation and DereverberationabstractFull-rank spatial covariance analysis (FCA) is a technique for blind source separation (BSS), and can be applied to underdetermined situations where the sources outnumber the microphones. This paper proposes multi-frame FCA as an extension of FCA to improve the BSS performance when the room reverberations are not so short that multiple time frames are needed to cover the dominant parts of the reverberations. There has already been proposed an FCA model that considers delayed source components. However, the existing FCA model does not take the correlation between different time frames into account. In contrast, our new extension models multiple time frames with multivariate Gaussian distributions of larger dimensionality than the existing FCA models, aiming to better model the source components spanning multiple time frames. We derive an expectation-maximization (EM) algorithm to optimize the model parameters. Experimental results show that the proposed multi-frame FCA performed clearly better than the existing FCA techniques in BSS tasks and also joint BSS and blind dereverberation tasks. Hiroshi Sawada, Rintaro Ikeshita, Keisuke Kinoshita, Tomohiro Nakatani |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Importance of Switch Optimization Criterion in Switching WPE DereverberationabstractWeighted prediction error (WPE) is a fundamental dereverberation method to predict the late reverberation component of an observed signal based on linear prediction (LP). Recently, WPE was extended to Switching WPE (SwWPE), which optimizes (i) multiple LP filters and (ii) switching parameters to determine the best LP filter used for each time-frequency bin. Conventionally, these parameters are optimized based on the maximum likelihood (ML) criterion, but this is not optimal in terms of signal quality, such as signal-to-distortion ratio (SDR) and word error rate (WER) of automatic speech recognition. We thus propose a new SwWPE processing flow that enables us to optimize switching parameters based on an arbitrary optimization criterion. Using oracle clean signals, we demonstrate the potential performance of our new approach with an SDR maximization criterion, revealing that it can significantly improve the SDR and WER obtained by the conventional ML-based SwWPE. This motivates us to propose new SwWPE processing in which the switching parameters are externally estimated using a deep neural network (DNN) that is trained with an end-to-end SDR maximization criterion. The experimental result clearly demonstrates the improved SDR performance of the new approach compared to the conventional WPE and SwWPE. Naoyuki Kamo, Rintaro Ikeshita, Keisuke Kinoshita, Tomohiro Nakatani |
ICASSP | 2 |
| 2022 | Multi-Frame Full-Rank Spatial Covariance Analysis for Underdetermined BSS in Reverberant EnvironmentsabstractFull-rank spatial covariance analysis (FCA) is a blind source separation (BSS) method, and can be applied to underdetermined cases where the sources outnumber the microphones. This paper proposes a new extension of FCA, aiming to improve BSS performance for mixtures in which the length of reverberation exceeds the analysis frame. There has already been proposed a model that considers delayed source components as the exceeded parts. In contrast, our new extension models multiple time frames with multivariate Gaussian distributions of larger dimensionality than the existing FCA models. We derive an expectation-maximization algorithm to optimize the model parameters. Experiments to separate four speech sources with three microphones show that the proposed extension outperforms the existing models, and more specifically, the original FCA by around 2 dB measured with signal-to-distortion ratio. Hiroshi Sawada, Rintaro Ikeshita, Keisuke Kinoshita, Tomohiro Nakatani |
ICASSP | 2 |
| 2022 | How bad are artifacts?: Analyzing the impact of speech enhancement errors on ASRabstractIt is challenging to improve automatic speech recognition (ASR) performance in noisy conditions with single-channel speech enhancement (SE).In this paper, we investigate the causes of ASR performance degradation by decomposing the SE errors using orthogonal projection-based decomposition (OPD).OPD decomposes the SE errors into noise and artifact components.The artifact component is defined as the SE error signal that cannot be represented as a linear combination of speech and noise sources.We propose manually scaling the error components to analyze their impact on ASR.We experimentally identify the artifact component as the main cause of performance degradation, and we find that mitigating the artifact can greatly improve ASR performance.Furthermore, we demonstrate that the simple observation adding (OA) technique (i.e., adding a scaled version of the observed signal to the enhanced speech) can monotonically increase the signal-to-artifact ratio under a mild condition.Accordingly, we experimentally confirm that OA improves ASR performance for both simulated and real recordings.The findings of this paper provide a better understanding of the influence of SE errors on ASR and open the door to future research on novel approaches for designing effective single-channel SE front-ends for ASR. Kazuma Iwamoto, Tsubasa Ochiai, Marc Delcroix, Rintaro Ikeshita, Hiroshi Sato 0002, Shoko Araki, Shigeru Katagiri |
INTERSPEECH | 4 |
| 2022 | Switching Independent Vector Analysis and its Extension to Blind and Spatially Guided Convolutional Beamforming AlgorithmsabstractThis paper develops a framework that can accurately perform denoising, dereverberation, and source separation using a relatively small number of microphones. It has been empirically confirmed that Independent Vector Analysis (IVA) can blindly separate$N$sources from their sound mixture even with diffuse noise when a sufficiently large number ($=M$) of microphones are available (i.e.,$M\gg N)$. However, the estimation accuracy is seriously degraded when the number of microphones, or more specifically$M-N$$(\geq 0)$, decreases. To overcome this IVA limitation, we propose switching IVA (swIVA) in this paper. With swIVA, the time frames of an observed signal with time-varying characteristics are clustered into several groups, each of which can be well handled by IVA with a small number of microphones, and thus accurate estimation can be achieved by individually applying IVA to each group. Conventionally, a switching mechanism was introduced into a Minimum-Variance Distortionless Response (MVDR) beamformer, and this paper extends the mechanism to work with a blind source separation algorithm. To incorporate dereverberation capability, we further extend swIVA to a blind Convolutional beamforming algorithm (swCIVA) that integrates swIVA and switching Weighted Prediction Error-based dereverberation (swWPE) in a jointly optimal way. With swCIVA, two different time-varying characteristics of an observed signal are captured for dereverberation and source separation to achieve effective estimation. We show that both swIVA and swCIVA can be optimized effectively based on blind signal processing, and their performance can be further improved using a spatial guide for initialization. Experiments demonstrate that both the proposed methods largely outperformed conventional IVA and its convolutional beamforming extension (CIVA) in terms of objective signal quality and automatic speech recognition scores when using relatively few microphones. Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Hiroshi Sawada, Naoyuki Kamo, Shoko Araki |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Blind and Neural Network-Guided Convolutional Beamformer for Joint Denoising, Dereverberation, and Source SeparationabstractThis paper proposes an approach for optimizing a Convolutional BeamFormer (CBF) that can jointly perform denoising (DN), dereverberation (DR), and source separation (SS). First, we develop a blind CBF optimization algorithm that requires no prior information on the sources or the room acoustics, by extending a conventional joint DR and SS method. For making the optimization computationally tractable, we incorporate two techniques into the approach: the Source-Wise Factorization (SW-Fact) of a CBF and the Independent Vector Extraction (IVE). To further improve the performance, we develop a method that integrates a neural network (NN) based source power spectra estimation with CBF optimization by an inverse-Gamma prior. Experiments using noisy reverberant mixtures reveal that our proposed method with both blind and NN-guided scenarios greatly outperforms the conventional state-of-the-art NN-supported mask-based CBF in terms of the improvement in automatic speech recognition and signal distortion reduction performance. Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Hiroshi Sawada, Shoko Araki |
ICASSP | 2 |
| 2021 | Neural Network-Based Virtual Microphone EstimatorabstractDeveloping microphone array technologies for a small number of microphones is important due to the constraints of many devices. One direction to address this situation consists of virtually augmenting the number of microphone signals, e.g., based on several physical model assumptions. However, such assumptions are not necessarily met in realistic conditions. In this paper, as an alternative approach, we propose a neural network-based virtual microphone estimator (NN-VME). The NN-VME estimates virtual microphone signals directly in the time domain, by utilizing the precise estimation capability of the recent time-domain neural networks. We adopt a fully supervised learning framework that uses actual observations at the locations of the virtual microphones at training time. Consequently, the NN-VME can be trained using only multi-channel observations and thus directly on real recordings, avoiding the need for unrealistic physical model-based assumptions. Experiments on the CHiME-4 corpus show that the proposed NN-VME achieves high virtual microphone estimation performance even for real recordings and that a beamformer augmented with the NN-VME improves both the speech enhancement and recognition performance. Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Shoko Araki |
ICASSP | 4 |
| 2021 | Low Latency Online Blind Source Separation Based on Joint Optimization with Blind DereverberationabstractThis paper presents a new low-latency online blind source separation (BSS) algorithm. Although algorithmic delay of a frequency domain online BSS can be reduced simply by shortening the short-time Fourier transform (STFT) frame length, it degrades the source separation performance in the presence of reverberation. This paper proposes a method to solve this problem by integrating BSS with Weighted Prediction Error (WPE) based dereverberation. Although a simple cascade of online BSS after online WPE upgrades the separation performance, the overall optimality is not guaranteed. Instead, this paper extends a recently proposed batch processing algorithm that can jointly optimize dereverberation and separation so that it can perform online processing with low computational cost and little processing delay (< 12 ms). The results of a source separation experiment in a noisy car environment suggest that the proposed online method has better separation performance than the simple cascaded methods. Tetsuya Ueda, Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Shoko Araki, Shoji Makino |
ICASSP | 3 |
| 2021 | Online Speech Dereverberation Using Mixture of Multichannel Linear Prediction ModelsabstractWe extend the state-of-the-art online dereverberation method, online weighted prediction error (WPE), which predicts late reverberation components using a multichannel linear prediction (MCLP) filter. The multi-input/output inverse theorem states that in general such an MCLP filter for WPE exists only if the number of sources is less than that of microphones M and there is no additive noise; otherwise, the WPE model has error, degrading its performance especially when M is small. To mitigate this WPE drawback, we recently developed an offline dereverberation method called switching WPE (SwWPE), which has multiple MCLP filters and switches them for each time-frequency bin. We here propose a recursive least squares (RLS) algorithm for the online optimization of SwWPE. Our proposed RLS has forgetting weights for each MCLP filter, and their decay rates are controlled based on a switching mechanism of SwWPE to stably improve the dereverberation performance of the online WPE. Experimental results show that our online SwWPE clearly outperforms online WPE in speech dereverberation tasks. Rintaro Ikeshita, Keisuke Kinoshita, Naoyuki Kamo, Tomohiro Nakatani |
IEEE Signal Process. Lett. | 1 |
| 2021 | Blind Signal Dereverberation Based on Mixture of Weighted Prediction Error ModelsabstractWe extend the linear prediction-based dereverberation method called weighted prediction error (WPE). WPE optimizes a causal finite impulse response (FIR) filter that predicts the late reverberation components of an observed signal. However, by the multi-input/output inverse (MINT) theorem, in general, such FIR filters exist only when the number of sources is fewer than that of the microphones and no ambient noise exists. To mitigate the model error of WPE in adverse environments, we propose a mixture model of multiple WPEs in which the time frames are divided into clusters in each frequency bin and the WPE's causal FIR filter is optimized in each cluster. Experimental results show that our proposed method significantly improves the dereverberation performance of WPE when the noise level is high or the number of microphones is small. Rintaro Ikeshita, Naoyuki Kamo, Tomohiro Nakatani |
IEEE Signal Process. Lett. | 1 |
| 2021 | Independent Vector Extraction for Fast Joint Blind Source Separation and DereverberationabstractWe address a blind source separation (BSS) problem in a noisy reverberant environment in which the number of microphones$M$is greater than the number of sources of interest, and the other noise components can be approximated as stationary and Gaussian distributed. Conventional BSS algorithms for the optimization of a multi-input multi-output convolutional beamformer have suffered from a huge computational cost when$M$is large. We here propose a computationally efficient method that integrates a weighted prediction error (WPE) dereverberation method and a fast BSS method called independent vector extraction (IVE), which has been developed for less reverberant environments. We show that, given the power spectrum for each source, the optimization problem of the new method can be reduced to that of IVE by exploiting the stationary condition, which makes the optimization easy to handle and computationally efficient. An experiment of speech signal separation shows that, compared to a conventional method that integrates WPE and independent vector analysis, our proposed method achieves much faster convergence while maintaining its separation performance. Rintaro Ikeshita, Tomohiro Nakatani |
IEEE Signal Process. Lett. | 1 |
| 2021 | A Joint Diagonalization Based Efficient Approach to Underdetermined Blind Audio Source Separation Using the Multichannel Wiener FilterabstractBlind source separation (BSS) of audio signals aims to separate original source signals from their mixtures recorded by microphones. The applications include automatic speech recognition in a noisy/multi-speaker environment, hearing aids, and music analysis. Independent component analysis (ICA) can perform BSS efficiently, but it is basically inapplicable to the underdetermined case-the number of sources > the number of microphones. In contrast, a BSS approach using the multichannel Wiener filter (MWF) is applicable even to the underdetermined case, but conventional methods based on this approach-including full-rank spatial covariance analysis (FCA)-are highly inefficient. This is because these methods require massive numbers of matrix inversions to design the MWF. To obtain the best of both worlds, we take a joint diagonalization approach: We restrict spatial covariance matrices of all sources to the class of jointly diagonalizable matrices. This enables the above matrix inversions to be replaced by mere scalar inversions of the diagonal elements of diagonal matrices. Based on this, we present FastFCA and FastMNMF-efficient methods for underdetermined BSS. In an experiment, FastFCA was several orders of magnitude faster than FCA without sacrificing separation performance. We also present a unified framework for underdetermined and determined BSS, which highlights theoretical connections between various methods including ours. The efficiency of our BSS methods makes them suitable for large data (e.g., data augmentation for machine learning) or limited computational resources encountered in, e.g., hearing aids, distributed microphone arrays, and online BSS. Nobutaka Ito, Rintaro Ikeshita, Hiroshi Sawada, Tomohiro Nakatani |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Overdetermined Independent Vector AnalysisabstractWe address the convolutive blind source separation problem for the (over-)determined case where (i) the number of nonstationary target-sources K is less than that of microphones M, and (ii) there are up to M - K stationary Gaussian noises that need not to be extracted. Independent vector analysis (IVA) can solve the problem by separating into M sources and selecting the top K highly nonstationary signals among them, but this approach suffers from a waste of computation especially when K ≪ M. Channel reductions in preprocessing of IVA by, e.g., principle component analysis have the risk of removing the target signals. We here extend IVA to resolve these issues. One such extension has been attained by assuming the orthogonality constraint (OC) that the sample correlation between the target and noise signals is to be zero. The proposed IVA, on the other hand, does not rely on OC and exploits only the independence between sources and the stationarity of the noises. This enables us to develop several efficient algorithms based on block coordinate descent methods with a problem specific acceleration. We clarify that one such algorithm exactly coincides with the conventional IVA with OC, and also explain that the other newly developed algorithms are faster than it. Experimental results show the improved computational load of the new algorithms compared to the conventional methods. In particular, a new algorithm specialized for K = 1 outperforms the others. Rintaro Ikeshita, Tomohiro Nakatani, Shoko Araki |
ICASSP | 1 |
| 2020 | Convergence-Guaranteed Independent Positive Semidefinite Tensor Analysis Based on Student's T DistributionabstractIn this paper, we address a blind source separation (BSS) problem and propose a new extended framework of independent positive semidefinite tensor analysis (IPSDTA). IPSDTA is a state-of-the-art BSS method that enables us to take interfrequency correlations into account, but the generative model is limited within the multivariate Gaussian distribution and its parameter optimization algorithm does not guarantee stable convergence. To resolve these problems, first, we propose to extend the generative model to a parametric multivariate Student’s t distribution that can deal with various types of signal. Secondly, we derive a new parameter optimization algorithm that guarantees the monotonic nonincrease in the cost function, providing stable convergence. Experimental results reveal that the cost function in the conventional IPSDTA does not display monotonically nonincreasing properties. On the other hand, the proposed method guarantees the monotonic nonincrease in the cost function and outperforms the conventional ILRMA and IPSDTA in the source-separation performance. Tatsuki Kondo, Kanta Fukushige, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Rintaro Ikeshita, Tomohiro Nakatani |
ICASSP | 6 |
| 2020 | DNN-supported Mask-based Convolutional Beamforming for Simultaneous Denoising, Dereverberation, and Source SeparationabstractIn this article, we investigate an integrated mask-based convolutional beamforming method for performing simultaneous denoising, dereverberation, and source separation. Conventionally, it is difficult for neural network-supported mask-based source separation to perform denoising and dereverberation at the same time and for spatial clustering-based source separation to reliably solve the permutation problem in the presence of noise and reverberation. This greatly limits the application of mask-based source separation. To address this issue, we propose a method to integrate state-of-the-art techniques for mask-based beamforming into a single optimization framework. These techniques include frequency-domain Convolutional Neural Network based utterance-level Permutation Invariant Training with a large receptive field (CNN-uPIT), noisy Complex Gaussian Mixture Model based spatial clustering (noisyCGMM), and Weighted Power minimization Distortionless response (WPD) convolutional beamforming. Our experiments show that all these components are essential for accurately estimating desired speech signals in noisy reverberant multisource environments. Tomohiro Nakatani, Riki Takahashi, Tsubasa Ochiai, Keisuke Kinoshita, Rintaro Ikeshita, Marc Delcroix, Shoko Araki |
ICASSP | 5 |
| 2020 | Beam-TasNet: Time-domain Audio Separation Network Meets Frequency-domain BeamformerabstractRecent studies have shown that acoustic beamforming using a microphone array plays an important role in the construction of high-performance automatic speech recognition (ASR) systems, especially for noisy and overlapping speech conditions. In parallel with the success of multichannel beamforming for ASR, in the speech separation field, the time-domain audio separation network (TasNet), which accepts a time-domain mixture as input and directly estimates the time-domain waveforms for each source, achieves remarkable speech separation performance. In light of these two recent trends, the question of whether TasNet can benefit from beamforming to achieve high ASR performance in overlapping speech conditions naturally arises. Motivated by this question, this paper proposes a novel speech separation scheme, i.e., Beam-TasNet, which combines TasNet with the frequency-domain beamformer, i.e., a minimum variance distortionless response (MVDR) beamformer, through spatial covariance computation to achieve better ASR performance. Experiments on the spatialized WSJ0-2mix corpus show that our proposed Beam-TasNet significantly outperforms the conventional TasNet without beamforming and, moreover, successfully achieves a word error rate comparable to an oracle mask-based MVDR beamformer. Tsubasa Ochiai, Marc Delcroix, Rintaro Ikeshita, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki |
ICASSP | 3 |
| 2020 | Computationally Efficient and Versatile Framework for Joint Optimization of Blind Speech Separation and Dereverberation
Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Hiroshi Sawada, Shoko Araki |
INTERSPEECH | 2 |
| 2020 | Jointly Optimal Denoising, Dereverberation, and Source SeparationabstractThis article proposes methods that can optimize a Convolutional BeamFormer (CBF) for jointly performing denoising, dereverberation, and source separation (DN+DR+SS) in a computationally efficient way. Conventionally, a cascade configuration, composed of a Weighted Prediction Error minimization (WPE) dereverberation filter followed by a Minimum Variance Distortionless Response (MVDR) beamformer, has been used as the state-of-the-art frontend of far-field speech recognition, even though this approach's overall optimality is not guaranteed. In the blind signal processing area, an approach for jointly optimizing dereverberation and source separation (DR+SS) has been proposed; however, it requires huge computing cost, and has not been extended for applications to DN+DR+SS. To overcome the above limitations, this paper develops new approaches for jointly optimizing DN+DR+SS in a computationally much more efficient way. To this end, we first present an objective function to optimize a CBF for performing DN+DR+SS based on maximum likelihood estimation on an assumption that the steering vectors of the target signals are given or can be estimated, e.g., using a neural network. This paper refers to a CBF optimized by this objective function as a weighted Minimum-Power Distortionless Response (wMPDR) CBF. Then, we derive two algorithms for optimizing a wMPDR CBF based on two different ways of factorizing a CBF into WPE filters and beamformers: one based on an extension of the conventional joint optimization approach proposed for DR+SS and another based on a novel technique. Experiments using noisy reverberant sound mixtures show that the proposed optimization approaches greatly improve the performance of the speech enhancement in comparison with the conventional cascade configuration in terms of signal distortion measures and ASR performance. The proposed approaches also greatly reduce the computing cost with improved estimation accuracy in comparison with the conventional joint optimization approach. Tomohiro Nakatani, Christoph Böddeker, Keisuke Kinoshita, Rintaro Ikeshita, Marc Delcroix, Reinhold Häb-Umbach |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2019 | Acoustic Modeling for Distant Multi-talker Speech Recognition with Single- and Multi-channel BranchesabstractThis paper presents a novel heterogeneous-input multi-channel acoustic model (AM) that has both single-channel and multi-channel input branches. In our proposed training pipeline, a single-channel AM is trained first, then a multi-channel AM is trained starting from the single-channel AM with a randomly initialized multi-channel input branch. Our model uniquely uses the power of a complemen-tal speech enhancement (SE) module while exploiting the power of jointly trained AM and SE architecture. Our method was the foundation for the Hitachi/JHU CHiME-5 system that achieved the second-best result in the CHiME-5 competition, and this paper details various investigation results that we were not able to present during the competition period. We also evaluated and reconfirmed our method's effectiveness with the AMI Meeting Corpus. Our AM achieved a 30.12% word error rate (WER) for the development set and a 32.33% WER for the evaluation set for the AMI Corpus, both of which are the best results ever reported to the best of our knowledge. Naoyuki Kanda, Yusuke Fujita, Shota Horiguchi, Rintaro Ikeshita, Kenji Nagamatsu, Shinji Watanabe 0001 |
ICASSP | 4 |
| 2018 | Independent Low-Rank Matrix Analysis Based on Multivariate Complex Exponential Power DistributionabstractIndependent low-rank matrix analysis (ILRMA), a unified method of independent vector analysis (IVA) and nonnegative matrix factorization (NMF), is a state-of-the-art blind source separation method for convolutive mixtures. Although ILRMA provides high separation performance for music signals whose spectra can be well modeled by NMF, speech spectra do not have low-rank properties, and modeling them by NMF is not appropriate. In this paper, to stably improve the separation performance of ILRMA for speech mixtures, a source spectrum model in ILRMA is generalized to explicitly model the strong higher-order correlations between neighboring frequency bins of speech signals. In addition, multivariate complex exponential power distributions, which are recognized to have high performance with IVA, are introduced as source distributions assumed in ILRMA. Experimental results show the effectiveness of the proposed method over the original ILRMA when separating speech mixtures. Rintaro Ikeshita, Yohei Kawaguchi |
ICASSP | 1 |
| 2015 | Unified ASR system using LGM-based source separation, noise-robust feature extraction, and word hypothesis selectionabstractIn this paper, we propose a unified system that incorporates speech source separation and automatic speech recognition for various noise environments. There are three features in the proposed system. The first feature of the proposed method is the LGM (local Gaussian modeling) based source separation with the efficient permutation alignment method that integrates a power spectrum correlation based method and a direction-of-arrival (DOA) based method. Evaluation results show that using the separated speech with the baseline acoustic modeling method reduces the word error rate (WER) significantly. The second feature of the proposed method is multi-condition training with per-utterance normalized features and noise-aware features in the acoustic modeling step. In this paper, we show that the proposed training method is effective even when an input signal has been distorted through the source separation step. The third feature is the word hypothesis selection method for integrating multiple recognition results. The proposed selection method estimates correct words based on a recognizer's confidence and co-occurrence characteristics. The evaluation results show that the proposed selection method outperforms the conventional recognizer output voting error reduction (ROVER) method. The proposed system is evaluated using the third CHiME challenge dataset. Evaluation results show that the proposed system resulted in an improvement of 66.1% over the baseline system. Yusuke Fujita, Ryoichi Takashima, Takeshi Homma, Rintaro Ikeshita, Yohei Kawaguchi, Takashi Sumiyoshi, Takashi Endo, Masahito Togami |
ASRU | 4 |