VLDB 2026 Research / reviewers in the wild / expert
Tomohiro Nakatani
dblp:38/5179
· DBLP profile ↗
227ranked-venue papers
37as first author
34since 2021 · last 2026
0000-0002-7487-7150ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 191 · 28 first-author · 26 since 2021Artificial intelligence and machine learning · 101 · 18 first-author · 15 since 2021Systems, architecture and hardware · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Microphone array geometry-independent multi-talker distant ASR: NTT system for DASR task of the CHiME-8 challenge
Naoyuki Kamo, Naohiro Tawara, Atsushi Ando, Takatomo Kano, Hiroshi Sato 0002, Rintaro Ikeshita, Takafumi Moriya, Shota Horiguchi, Kohei Matsuura, Atsunori Ogawa, Alexis Plaquet, Takanori Ashihara, Tsubasa Ochiai, Masato Mimura, Marc Delcroix, Tomohiro Nakatani, Taichi Asami, Shoko Araki |
Comput. Speech Lang. | 16 |
| 2025 | SoundBeam meets M2D: Target Sound Extraction with Audio Foundation ModelabstractTarget sound extraction (TSE) consists of isolating a desired sound from a mixture of arbitrary sounds using clues to identify it. A TSE system requires solving two problems at once, identifying the target source and extracting the target signal from the mixture. For increased practicability, the same system should work with various types of sound. The duality of the problem and the wide variety of sounds make it challenging to train a powerful TSE system from scratch. In this paper, to tackle this problem, we explore using a pre-trained audio foundation model that can provide rich feature representations of sounds within a TSE system. We chose the masked-modeling duo (M2D) foundation model, which appears especially suited for the TSE task, as it is trained using a dual objective consisting of sound-label predictions and improved masked prediction. These objectives are related to sound identification and the signal extraction problems of TSE. We propose a new TSE system that integrates the feature representation from M2D into SoundBeam, which is a strong TSE system that can exploit both target sound class labels and pre-recorded enrollments (or audio queries) as clues. We show experimentally that using M2D can increase extraction performance, especially when employing enrollment clues. Carlos Hernandez-Olivan, Marc Delcroix, Tsubasa Ochiai, Daisuke Niizumi, Naohiro Tawara, Tomohiro Nakatani, Shoko Araki |
ICASSP | 6 |
| 2025 | A Hybrid Probabilistic-Deterministic Model Recursively Enhancing SpeechabstractThis paper introduces Probabilistic-Deterministic Recursive Enhancement (PDRE), an innovative iterative Speech Enhancement (SE) approach that integrates probabilistic and deterministic methodologies. Recent advancements in diffusion models have demonstrated the exceptional effectiveness of probabilistic model-based iterative estimation in SE, especially when combined with deterministic Neural Network (NN)-based methods. However, these models often require extensive iterations, leading to significant computational costs. To tackle this issue, we propose PDRE as a more efficient alternative. PDRE progressively refines the clean speech density estimates by recursively applying an Enhancement Network (EN), which is trained using a maximum likelihood objective. A single application of the EN can substantially improve the estimation, enabling PDRE to achieve high SE accuracy with significantly fewer iterations. Additionally, PDRE synergizes recursive enhancement with deterministic signal estimation, resulting in even greater accuracy. Our experiments demonstrate that PDRE significantly reduces iteration counts and computational costs compared to diffusion model-based SEs while maintaining or improving the remarkably high estimation accuracy. Tomohiro Nakatani, Naoyuki Kamo, Marc Delcroix, Shoko Araki |
ICASSP | 1 |
| 2025 | MOVER: Combining Multiple Meeting Recognition Systems
Naoyuki Kamo, Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2024 | Neural Network-Based Virtual Microphone Estimation with Virtual Microphone and Beamformer-Level Multi-Task LossabstractArray processing performance depends on the number of microphones available. Virtual microphone estimation (VME) has been proposed to increase the number of microphone signals artificially. Neural network-based VME (NN-VME) trains an NN with a VM-level loss to predict a signal at a microphone location that is available during training but not at inference. However, this training objective may not be optimal for a specific array processing back-end, such as beamforming. An alternative approach is to use a training objective considering the array-processing back-end, such as a loss on the beamformer output. This approach may generate signals optimal for beamforming but not physically grounded. To combine the advantages of both approaches, this paper proposes a multi-task loss for NN-VME that combines both VM-level and beamformer-level losses. We evaluate the proposed multi-task NN-VME on multi-talker underdetermined conditions and show that it achieves a 33.1 % relative WER improvement compared to using only real microphones and 10.8 % compared to using a prior NN-VME approach. Hanako Segawa, Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Rintaro Ikeshita, Shoko Araki, Takeshi Yamada, Shoji Makino |
ICASSP | 4 |
| 2024 | Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers
Marvin Tammen, Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Shoko Araki, Simon Doclo |
INTERSPEECH | 4 |
| 2024 | Blind and Spatially-Regularized Online Joint Optimization of Source Separation, Dereverberation, and Noise ReductionabstractThis paper proposes a computationally efficient joint optimization algorithm that performs online source separation, dereverberation, and noise reduction based on blind and spatially-regularized processing. When applying such online Blind Source Separation (BSS) as online Independent Vector Extraction (IVE) to a speech application, we must focus on the trade-off between the algorithmic delay and separation accuracy, both of which depend on the analysis frame length. In addition, to separate the sources with specified source permutation, researchers introduced spatial regularization based on the Directions-of-Arrival (DOAs) of the sources into IVE. However, the scale ambiguity of IVE often makes the spatial regularization work inappropriately. To solve these problems, we first propose a blind online joint optimization algorithm of IVE and weighted prediction error dereverberation (WPE). This online algorithm can achieve accurate separation even using short analysis frames because reverberation can be reduced using WPE. We then extend the online joint optimization with robust spatial regularization. We reveal that regularizing the scale of the separated signals is very effective in making the DOA-based spatial regularization work reliably. Our experiments confirm that our blind online joint optimization algorithm can significantly improve the separation accuracy with an algorithmic delay of 8 ms. In addition, we confirm that the proposed spatially-regularized online joint optimization algorithm reduces the rate of the source permutation error to zero percent. Tetsuya Ueda, Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Shoko Araki, Shoji Makino |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Fast Online Source Steering Algorithm for Tracking Single Moving Source Using Online Independent Vector AnalysisabstractWe address the problem of separating moving sources using online independent vector analysis (IVA). To solve this problem, researchers have extended the iterative projection (IP) and iterative source steering (ISS) algorithms developed for batch auxiliary-function-based IVA (AuxIVA) to online scenarios and showed their effectiveness. However, the conventional online IP and ISS are slow because they update K × K covariance matrices for all sources, where K is the number of microphones. Here, we show that, in a target-source tracking scenario in which only one source moves, there exists an inexpensive formula for online ISS that avoids updating the full covariance matrices without changing the behavior of the algorithm. The time complexity of the proposed algorithm, which we call online source steering (OSS), is K times smaller than that of the conventional online IP and ISS for the target-source tracking task. A numerical experiment on separating a moving source demonstrates that the proposed OSS is significantly faster than the conventional online IP and ISS. Taishi Nakashima, Rintaro Ikeshita, Nobutaka Ono, Shoko Araki, Tomohiro Nakatani |
ICASSP | 5 |
| 2023 | Impact of Residual Noise and Artifacts in Speech Enhancement Errors on Intelligibility of Human and Machine
Shoko Araki, Ayako Yamamoto, Tsubasa Ochiai, Kenichi Arai, Atsunori Ogawa, Tomohiro Nakatani, Toshio Irino |
INTERSPEECH | 6 |
| 2023 | Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization
Marc Delcroix, Naohiro Tawara, Mireia Díez, Federico Landini, Anna Silnova, Atsunori Ogawa, Tomohiro Nakatani, Lukás Burget, Shoko Araki |
INTERSPEECH | 7 |
| 2023 | Target Speech Extraction with Conditional Diffusion Model
Naoyuki Kamo, Marc Delcroix, Tomohiro Nakatani |
INTERSPEECH | 3 |
| 2023 | Mask-Based Neural Beamforming for Moving Speakers With Self-Attention-Based TrackingabstractBeamforming is a powerful tool designed to enhance speech signals from the direction of a target source. Computing the beamforming filter requires estimating spatial covariance matrices (SCMs) of the source and noise signals. Time-frequency masks are often used to compute these SCMs. Most studies of mask-based beamforming have assumed that the sources do not move. However, sources often move in practice, which causes performance degradation. In this paper, we address the problem of mask-based beamforming for moving sources. We first review classical approaches to tracking a moving source, which perform online or blockwise computation of the SCMs. We show that these approaches can be interpreted as computing a sum of instantaneous SCMs weighted by attention weights. These weights indicate which time frames of the signal to consider in the SCM computation. Online or blockwise computation assumes a heuristic and deterministic way of computing these attention weights that, although simple, may not result in optimal performance. We thus introduce a learning-based framework that computes optimal attention weights for beamforming. We achieve this using a neural network implemented with self-attention layers. We show experimentally that our proposed framework can greatly improve beamforming performance in moving source situations while maintaining high performance in non-moving situations, thus enabling the development of mask-based beamformers robust to source movements. Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Shoko Araki |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Multi-Frame Full-Rank Spatial Covariance Analysis for Underdetermined Blind Source Separation and DereverberationabstractFull-rank spatial covariance analysis (FCA) is a technique for blind source separation (BSS), and can be applied to underdetermined situations where the sources outnumber the microphones. This paper proposes multi-frame FCA as an extension of FCA to improve the BSS performance when the room reverberations are not so short that multiple time frames are needed to cover the dominant parts of the reverberations. There has already been proposed an FCA model that considers delayed source components. However, the existing FCA model does not take the correlation between different time frames into account. In contrast, our new extension models multiple time frames with multivariate Gaussian distributions of larger dimensionality than the existing FCA models, aiming to better model the source components spanning multiple time frames. We derive an expectation-maximization (EM) algorithm to optimize the model parameters. Experimental results show that the proposed multi-frame FCA performed clearly better than the existing FCA techniques in BSS tasks and also joint BSS and blind dereverberation tasks. Hiroshi Sawada, Rintaro Ikeshita, Keisuke Kinoshita, Tomohiro Nakatani |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Importance of Switch Optimization Criterion in Switching WPE DereverberationabstractWeighted prediction error (WPE) is a fundamental dereverberation method to predict the late reverberation component of an observed signal based on linear prediction (LP). Recently, WPE was extended to Switching WPE (SwWPE), which optimizes (i) multiple LP filters and (ii) switching parameters to determine the best LP filter used for each time-frequency bin. Conventionally, these parameters are optimized based on the maximum likelihood (ML) criterion, but this is not optimal in terms of signal quality, such as signal-to-distortion ratio (SDR) and word error rate (WER) of automatic speech recognition. We thus propose a new SwWPE processing flow that enables us to optimize switching parameters based on an arbitrary optimization criterion. Using oracle clean signals, we demonstrate the potential performance of our new approach with an SDR maximization criterion, revealing that it can significantly improve the SDR and WER obtained by the conventional ML-based SwWPE. This motivates us to propose new SwWPE processing in which the switching parameters are externally estimated using a deep neural network (DNN) that is trained with an end-to-end SDR maximization criterion. The experimental result clearly demonstrates the improved SDR performance of the new approach compared to the conventional WPE and SwWPE. Naoyuki Kamo, Rintaro Ikeshita, Keisuke Kinoshita, Tomohiro Nakatani |
ICASSP | 4 |
| 2022 | Multi-Frame Full-Rank Spatial Covariance Analysis for Underdetermined BSS in Reverberant EnvironmentsabstractFull-rank spatial covariance analysis (FCA) is a blind source separation (BSS) method, and can be applied to underdetermined cases where the sources outnumber the microphones. This paper proposes a new extension of FCA, aiming to improve BSS performance for mixtures in which the length of reverberation exceeds the analysis frame. There has already been proposed a model that considers delayed source components as the exceeded parts. In contrast, our new extension models multiple time frames with multivariate Gaussian distributions of larger dimensionality than the existing FCA models. We derive an expectation-maximization algorithm to optimize the model parameters. Experiments to separate four speech sources with three microphones show that the proposed extension outperforms the existing models, and more specifically, the original FCA by around 2 dB measured with signal-to-distortion ratio. Hiroshi Sawada, Rintaro Ikeshita, Keisuke Kinoshita, Tomohiro Nakatani |
ICASSP | 4 |
| 2022 | Listen only to me! How well can target speech extraction handle false alarms?
Marc Delcroix, Keisuke Kinoshita, Tsubasa Ochiai, Katerina Zmolíková, Hiroshi Sato 0002, Tomohiro Nakatani |
INTERSPEECH | 6 |
| 2022 | Switching Independent Vector Analysis and its Extension to Blind and Spatially Guided Convolutional Beamforming AlgorithmsabstractThis paper develops a framework that can accurately perform denoising, dereverberation, and source separation using a relatively small number of microphones. It has been empirically confirmed that Independent Vector Analysis (IVA) can blindly separate$N$sources from their sound mixture even with diffuse noise when a sufficiently large number ($=M$) of microphones are available (i.e.,$M\gg N)$. However, the estimation accuracy is seriously degraded when the number of microphones, or more specifically$M-N$$(\geq 0)$, decreases. To overcome this IVA limitation, we propose switching IVA (swIVA) in this paper. With swIVA, the time frames of an observed signal with time-varying characteristics are clustered into several groups, each of which can be well handled by IVA with a small number of microphones, and thus accurate estimation can be achieved by individually applying IVA to each group. Conventionally, a switching mechanism was introduced into a Minimum-Variance Distortionless Response (MVDR) beamformer, and this paper extends the mechanism to work with a blind source separation algorithm. To incorporate dereverberation capability, we further extend swIVA to a blind Convolutional beamforming algorithm (swCIVA) that integrates swIVA and switching Weighted Prediction Error-based dereverberation (swWPE) in a jointly optimal way. With swCIVA, two different time-varying characteristics of an observed signal are captured for dereverberation and source separation to achieve effective estimation. We show that both swIVA and swCIVA can be optimized effectively based on blind signal processing, and their performance can be further improved using a spatial guide for initialization. Experiments demonstrate that both the proposed methods largely outperformed conventional IVA and its convolutional beamforming extension (CIVA) in terms of objective signal quality and automatic speech recognition scores when using relatively few microphones. Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Hiroshi Sawada, Naoyuki Kamo, Shoko Araki |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | End-to-End Dereverberation, Beamforming, and Speech Recognition in a Cocktail PartyabstractFar-field multi-speaker automatic speech recognition (ASR) has drawn increasing attention in recent years. Most existing methods feature a signal processing frontend and an ASR backend. In realistic scenarios, these modules are usually trained separately or progressively, which suffers from either inter-module mismatch or a complicated training process. In this paper, we propose an end-to-end multi-channel model that jointly optimizes the speech enhancement (including speech dereverberation, denoising, and separation) frontend and the ASR backend as a single system. To the best of our knowledge, this is the first work that proposes to optimize dereverberation, beamforming, and multi-speaker ASR in a fully end-to-end manner. The frontend module consists of a weighted prediction error (WPE) based submodule for dereverberation and a neural beamformer for denoising and speech separation. For the backend, we adopt a widely used end-to-end (E2E) ASR architecture. It is worth noting that the entire model is differentiable and can be optimized in a fully end-to-end manner using only the ASR criterion, without the need of parallel signal-level labels. We evaluate the proposed model on several multi-speaker benchmark datasets, and experimental results show that the fully E2E ASR model can achieve competitive performance on both noisy and reverberant conditions, with over 30% relative word error rate (WER) reduction over the single-channel baseline systems. Wangyou Zhang, Xuankai Chang, Christoph Böddeker, Tomohiro Nakatani, Shinji Watanabe 0001, Yanmin Qian |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Convolutive Transfer Function Invariant SDR Training Criteria for Multi-Channel Reverberant Speech SeparationabstractTime-domain training criteria have proven to be very effective for the separation of single-channel non-reverberant speech mixtures. Likewise, mask-based beamforming has shown impressive performance in multi-channel reverberant speech enhancement and source separation. Here, we propose to combine neural network supported multi-channel source separation with a time-domain training objective function. For the objective we propose to use a convolutive transfer function invariant Signal-to-Distortion Ratio (CI-SDR) based loss. While this is a well-known evaluation metric (BSS Eval), it has not been used as a training objective before. To show the effectiveness, we demonstrate the performance on LibriSpeech based reverberant mixtures. On this task, the proposed system approaches the error rate obtained on single-source non-reverberant input, i.e., LibriSpeech test clean, with a difference of only 1.2 percentage points, thus outperforming a conventional permutation invariant training based system and alternative objectives like Scale Invariant Signal-to-Distortion Ratio by a large margin. Christoph Böddeker, Wangyou Zhang, Tomohiro Nakatani, Keisuke Kinoshita, Tsubasa Ochiai, Marc Delcroix, Naoyuki Kamo, Yanmin Qian, Reinhold Häb-Umbach |
ICASSP | 3 |
| 2021 | Speaker Activity Driven Neural Speech ExtractionabstractTarget speech extraction, which extracts the speech of a target speaker in a mixture given auxiliary speaker clues, has recently received increased interest. Various clues have been investigated such as pre-recorded enrollment utterances, direction information, or video of the target speaker. In this paper, we explore the use of speaker activity information as an auxiliary clue for single-channel neural network-based speech extraction. We propose a speaker activity driven speech extraction neural network (ADEnet) and show that it can achieve performance levels competitive with enrollment-based approaches, without the need for pre-recordings. We further demonstrate the potential of the proposed approach for processing meeting-like recordings, where the speaker activity is obtained from a diarization system. We show that this simple yet practical approach can successfully extract speakers after diarization, which results in improved ASR performance, especially in high overlapping conditions, with a relative word error rate reduction of up to 25%. Marc Delcroix, Katerina Zmolíková, Tsubasa Ochiai, Keisuke Kinoshita, Tomohiro Nakatani |
ICASSP | 5 |
| 2021 | Blind and Neural Network-Guided Convolutional Beamformer for Joint Denoising, Dereverberation, and Source SeparationabstractThis paper proposes an approach for optimizing a Convolutional BeamFormer (CBF) that can jointly perform denoising (DN), dereverberation (DR), and source separation (SS). First, we develop a blind CBF optimization algorithm that requires no prior information on the sources or the room acoustics, by extending a conventional joint DR and SS method. For making the optimization computationally tractable, we incorporate two techniques into the approach: the Source-Wise Factorization (SW-Fact) of a CBF and the Independent Vector Extraction (IVE). To further improve the performance, we develop a method that integrates a neural network (NN) based source power spectra estimation with CBF optimization by an inverse-Gamma prior. Experiments using noisy reverberant mixtures reveal that our proposed method with both blind and NN-guided scenarios greatly outperforms the conventional state-of-the-art NN-supported mask-based CBF in terms of the improvement in automatic speech recognition and signal distortion reduction performance. Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Hiroshi Sawada, Shoko Araki |
ICASSP | 1 |
| 2021 | Neural Network-Based Virtual Microphone EstimatorabstractDeveloping microphone array technologies for a small number of microphones is important due to the constraints of many devices. One direction to address this situation consists of virtually augmenting the number of microphone signals, e.g., based on several physical model assumptions. However, such assumptions are not necessarily met in realistic conditions. In this paper, as an alternative approach, we propose a neural network-based virtual microphone estimator (NN-VME). The NN-VME estimates virtual microphone signals directly in the time domain, by utilizing the precise estimation capability of the recent time-domain neural networks. We adopt a fully supervised learning framework that uses actual observations at the locations of the virtual microphones at training time. Consequently, the NN-VME can be trained using only multi-channel observations and thus directly on real recordings, avoiding the need for unrealistic physical model-based assumptions. Experiments on the CHiME-4 corpus show that the proposed NN-VME achieves high virtual microphone estimation performance even for real recordings and that a beamformer augmented with the NN-VME improves both the speech enhancement and recognition performance. Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Shoko Araki |
ICASSP | 3 |
| 2021 | Low Latency Online Blind Source Separation Based on Joint Optimization with Blind DereverberationabstractThis paper presents a new low-latency online blind source separation (BSS) algorithm. Although algorithmic delay of a frequency domain online BSS can be reduced simply by shortening the short-time Fourier transform (STFT) frame length, it degrades the source separation performance in the presence of reverberation. This paper proposes a method to solve this problem by integrating BSS with Weighted Prediction Error (WPE) based dereverberation. Although a simple cascade of online BSS after online WPE upgrades the separation performance, the overall optimality is not guaranteed. Instead, this paper extends a recently proposed batch processing algorithm that can jointly optimize dereverberation and separation so that it can perform online processing with low computational cost and little processing delay (< 12 ms). The results of a source separation experiment in a noisy car environment suggest that the proposed online method has better separation performance than the simple cascaded methods. Tetsuya Ueda, Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Shoko Araki, Shoji Makino |
ICASSP | 2 |
| 2021 | Data Fusion for Audiovisual Speaker Localization: Extending Dynamic Stream Weights to the Spatial DomainabstractEstimating the positions of multiple speakers can be helpful for tasks like automatic speech recognition or speaker diarization. Both applications benefit from a known speaker position when, for instance, applying beamforming or assigning unique speaker identities. Recently, several approaches utilizing acoustic signals augmented with visual data have been proposed for this task. However, both the acoustic and the visual modality may be corrupted in specific spatial regions, for instance due to poor lighting conditions or to the presence of background noise. This paper proposes a novel audiovisual data fusion framework for speaker localization by assigning individual dynamic stream weights to specific regions in the localization space. This fusion is achieved via a neural network, which combines the predictions of individual audio and video trackers based on their time- and location-dependent reliability. A performance evaluation using audiovisual recordings yields promising results, with the proposed fusion approach outperforming all baseline models. Julio Wissing, Benedikt T. Boenninghoff, Dorothea Kolossa, Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki, Christopher Schymura |
ICASSP | 7 |
| 2021 | End-to-End Dereverberation, Beamforming, and Speech Recognition with Improved Numerical Stability and Advanced FrontendabstractRecently, the end-to-end approach has been successfully applied to multi-speaker speech separation and recognition in both single-channel and multichannel conditions. However, severe performance degradation is still observed in the reverberant and noisy scenarios, and there is still a large performance gap between anechoic and reverberant conditions. In this work, we focus on the multichannel multi-speaker reverberant condition, and propose to extend our previous framework for end-to-end dereverberation, beamforming, and speech recognition with improved numerical stability and advanced frontend subnetworks including voice activity detection like masks. The techniques significantly stabilize the end-to-end training process. The experiments on the spatialized wsj1-2mix corpus show that the proposed system achieves about 35% WER relative reduction compared to our conventional multi-channel E2E ASR system, and also obtains decent speech dereverberation and separation performance (SDR=12.5 dB) in the reverberant multi-speaker condition while trained only with the ASR criterion. Wangyou Zhang, Christoph Böddeker, Shinji Watanabe 0001, Tomohiro Nakatani, Marc Delcroix, Keisuke Kinoshita, Tsubasa Ochiai, Naoyuki Kamo, Reinhold Häb-Umbach, Yanmin Qian |
ICASSP | 4 |
| 2021 | PILOT: Introducing Transformers for Probabilistic Sound Event LocalizationabstractSound event localization aims at estimating the positions of sound sources in the environment with respect to an acoustic receiver (e.g. a microphone array). Recent advances in this domain most prominently focused on utilizing deep recurrent neural networks. Inspired by the success of transformer architectures as a suitable alternative to classical recurrent neural networks, this paper introduces a novel transformer-based sound event localization framework, where temporal dependencies in the received multi-channel audio signals are captured via self-attention mechanisms. Additionally, the estimated sound event positions are represented as multivariate Gaussian variables, yielding an additional notion of uncertainty, which many previously proposed deep learning-based systems designed for this application do not provide. The framework is evaluated on three publicly available multi-source sound event localization datasets and compared against state-of-the-art methods in terms of localization error and event detection accuracy. It outperforms all competing systems on all datasets with statistical significant differences in performance. Christopher Schymura, Benedikt T. Boenninghoff, Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki, Dorothea Kolossa |
Interspeech | 6 |
| 2021 | Comparison of Remote Experiments Using Crowdsourcing and Laboratory Experiments on Speech IntelligibilityabstractMany subjective experiments have been performed to develop objective speech intelligibility measures, but the novel coronavirus outbreak has made it very difficult to conduct experiments in a laboratory. One solution is to perform remote testing using crowdsourcing; however, because we cannot control the listening conditions, it is unclear whether the results are entirely reliable. In this study, we compared speech intelligibility scores obtained in remote and laboratory experiments. The results showed that the mean and standard deviation (SD) of the remote experiments' speech reception threshold (SRT) were higher than those of the laboratory experiments. However, the variance in the SRTs across the speech-enhancement conditions revealed similarities, implying that remote testing results may be as useful as laboratory experiments to develop an objective measure. We also show that the practice session scores correlate with the SRT values. This is a priori information before performing the main tests and would be useful for data screening to reduce the variability of the SRT distribution. Ayako Yamamoto, Toshio Irino, Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani |
Interspeech | 7 |
| 2021 | Multimodal Attention Fusion for Target Speaker ExtractionabstractTarget speaker extraction, which aims at extracting a target speaker's voice from a mixture of voices using audio, visual or locational clues, has received much interest. Recently an audio-visual target speaker extraction has been proposed that extracts target speech by using complementary audio and visual clues. Although audio-visual target speaker extraction offers a more stable performance than single modality methods for simulated data, its adaptation towards realistic situations has not been fully explored as well as evaluations on real recorded mixtures. One of the major issues to handle realistic situations is how to make the system robust to clue corruption because in real recordings both clues may not be equally reliable, e.g. visual clues may be affected by occlusions. In this work, we propose a novel attention mechanism for multi-modal fusion and its training methods that enable to effectively capture the reliability of the clues and weight the more reliable ones. Our proposals improve signal to distortion ratio (SDR) by 1.0 dB over conventional fusion mechanisms on simulated data. Moreover, we also record an audio-visual dataset of simultaneous speech with realistic visual clue corruption and show that audio-visual target speaker extraction with our proposals successfully work on real data. Hiroshi Sato 0002, Tsubasa Ochiai, Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani, Shoko Araki |
SLT | 5 |
| 2021 | Integration of Variational Autoencoder and Spatial Clustering for Adaptive Multi-Channel Neural Speech SeparationabstractIn this paper, we propose a method combining variational autoencoder model of speech with a spatial clustering approach for multi-channel speech separation. The advantage of integrating spatial clustering with a spectral model was shown in several works. As the spectral model, previous works used either factorial generative models of the mixed speech or discriminative neural networks. In our work, we combine the strengths of both approaches, by building a factorial model based on a generative neural network, a variational autoencoder. By doing so, we can exploit the modeling power of neural networks, but at the same time, keep a structured model. Such a model can be advantageous when adapting to new noise conditions as only the noise part of the model needs to be modified. We show experimentally, that our model significantly outperforms previous factorial model based on Gaussian mixture model (DOLPHIN), performs comparably to integration of permutation invariant training with spatial clustering, and enables us to easily adapt to new noise conditions. Katerina Zmolíková, Marc Delcroix, Lukás Burget, Tomohiro Nakatani, Jan Cernocký |
SLT | 4 |
| 2021 | Far-Field Automatic Speech RecognitionabstractThe machine recognition of speech spoken at a distance from the microphones, known as far-field automatic speech recognition (ASR), has received a significant increase in attention in science and industry, which caused or was caused by an equally significant improvement in recognition accuracy. Meanwhile, it has entered the consumer market with digital home assistants with a spoken language interface being its most prominent application. Speech recorded at a distance is affected by various acoustic distortions, and consequently, quite different processing pipelines have emerged compared with ASR for close-talk speech. A signal enhancement front end for dereverberation, source separation, and acoustic beamforming is employed to clean up the speech, and the back-end ASR engine is robustified by multicondition training and adaptation. We will also describe the so-called end-to-end approach to ASR, which is a new promising architecture that has recently been extended to the far-field scenario. This tutorial article gives an account of the algorithms used to enable accurate speech recognition from a distance, and it will be seen that, although deep learning has a significant share in the technological breakthroughs, a clever combination with traditional signal processing can lead to surprisingly effective solutions. Reinhold Häb-Umbach, Jahn Heymann, Lukas Drude, Shinji Watanabe 0001, Marc Delcroix, Tomohiro Nakatani |
Proc. IEEE | 6 |
| 2021 | Online Speech Dereverberation Using Mixture of Multichannel Linear Prediction ModelsabstractWe extend the state-of-the-art online dereverberation method, online weighted prediction error (WPE), which predicts late reverberation components using a multichannel linear prediction (MCLP) filter. The multi-input/output inverse theorem states that in general such an MCLP filter for WPE exists only if the number of sources is less than that of microphones M and there is no additive noise; otherwise, the WPE model has error, degrading its performance especially when M is small. To mitigate this WPE drawback, we recently developed an offline dereverberation method called switching WPE (SwWPE), which has multiple MCLP filters and switches them for each time-frequency bin. We here propose a recursive least squares (RLS) algorithm for the online optimization of SwWPE. Our proposed RLS has forgetting weights for each MCLP filter, and their decay rates are controlled based on a switching mechanism of SwWPE to stably improve the dereverberation performance of the online WPE. Experimental results show that our online SwWPE clearly outperforms online WPE in speech dereverberation tasks. Rintaro Ikeshita, Keisuke Kinoshita, Naoyuki Kamo, Tomohiro Nakatani |
IEEE Signal Process. Lett. | 4 |
| 2021 | Blind Signal Dereverberation Based on Mixture of Weighted Prediction Error ModelsabstractWe extend the linear prediction-based dereverberation method called weighted prediction error (WPE). WPE optimizes a causal finite impulse response (FIR) filter that predicts the late reverberation components of an observed signal. However, by the multi-input/output inverse (MINT) theorem, in general, such FIR filters exist only when the number of sources is fewer than that of the microphones and no ambient noise exists. To mitigate the model error of WPE in adverse environments, we propose a mixture model of multiple WPEs in which the time frames are divided into clusters in each frequency bin and the WPE's causal FIR filter is optimized in each cluster. Experimental results show that our proposed method significantly improves the dereverberation performance of WPE when the noise level is high or the number of microphones is small. Rintaro Ikeshita, Naoyuki Kamo, Tomohiro Nakatani |
IEEE Signal Process. Lett. | 3 |
| 2021 | Independent Vector Extraction for Fast Joint Blind Source Separation and DereverberationabstractWe address a blind source separation (BSS) problem in a noisy reverberant environment in which the number of microphones$M$is greater than the number of sources of interest, and the other noise components can be approximated as stationary and Gaussian distributed. Conventional BSS algorithms for the optimization of a multi-input multi-output convolutional beamformer have suffered from a huge computational cost when$M$is large. We here propose a computationally efficient method that integrates a weighted prediction error (WPE) dereverberation method and a fast BSS method called independent vector extraction (IVE), which has been developed for less reverberant environments. We show that, given the power spectrum for each source, the optimization problem of the new method can be reduced to that of IVE by exploiting the stationary condition, which makes the optimization easy to handle and computationally efficient. An experiment of speech signal separation shows that, compared to a conventional method that integrates WPE and independent vector analysis, our proposed method achieves much faster convergence while maintaining its separation performance. Rintaro Ikeshita, Tomohiro Nakatani |
IEEE Signal Process. Lett. | 2 |
| 2021 | A Joint Diagonalization Based Efficient Approach to Underdetermined Blind Audio Source Separation Using the Multichannel Wiener FilterabstractBlind source separation (BSS) of audio signals aims to separate original source signals from their mixtures recorded by microphones. The applications include automatic speech recognition in a noisy/multi-speaker environment, hearing aids, and music analysis. Independent component analysis (ICA) can perform BSS efficiently, but it is basically inapplicable to the underdetermined case-the number of sources > the number of microphones. In contrast, a BSS approach using the multichannel Wiener filter (MWF) is applicable even to the underdetermined case, but conventional methods based on this approach-including full-rank spatial covariance analysis (FCA)-are highly inefficient. This is because these methods require massive numbers of matrix inversions to design the MWF. To obtain the best of both worlds, we take a joint diagonalization approach: We restrict spatial covariance matrices of all sources to the class of jointly diagonalizable matrices. This enables the above matrix inversions to be replaced by mere scalar inversions of the diagonal elements of diagonal matrices. Based on this, we present FastFCA and FastMNMF-efficient methods for underdetermined BSS. In an experiment, FastFCA was several orders of magnitude faster than FCA without sacrificing separation performance. We also present a unified framework for underdetermined and determined BSS, which highlights theoretical connections between various methods including ours. The efficiency of our BSS methods makes them suitable for large data (e.g., data augmentation for machine learning) or limited computational resources encountered in, e.g., hearing aids, distributed microphone arrays, and online BSS. Nobutaka Ito, Rintaro Ikeshita, Hiroshi Sawada, Tomohiro Nakatani |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Jointly Optimal Dereverberation and BeamformingabstractWe previously proposed an optimal (in the maximum likelihood sense) convolutional beamformer that can perform simultaneous denoising and dereverberation, and showed its superiority over the widely used cascade of a Weighted Prediction Error (WPE) dereverberation filter and a conventional Minimum-Power Distortionless Response (MPDR) beamformer. However, it has not been fully investigated which components in the convolutional beamformer yield such superiority. To this end, this paper presents a new derivation of the convolutional beamformer that allows us to factorize it into a WPE dereverberation filter, and a special type of a (non-convolutional) beamformer, referred to as a weighted MPDR (wM-PDR) beamformer, without loss of optimality. With experiments, we show that the superiority of the convolutional beamformer in fact comes from its wMPDR part. Christoph Böddeker, Tomohiro Nakatani, Keisuke Kinoshita, Reinhold Häb-Umbach |
ICASSP | 2 |
| 2020 | Improving Speaker Discrimination of Target Speech Extraction With Time-Domain SpeakerbeamabstractTarget speech extraction, which extracts a single target source in a mixture given clues about the target speaker, has attracted increasing attention. We have recently proposed SpeakerBeam, which exploits an adaptation utterance of the target speaker to extract his/her voice characteristics that are then used to guide a neural network towards extracting speech of that speaker. SpeakerBeam presents a practical alternative to speech separation as it enables tracking speech of a target speaker across utterances, and achieves promising speech extraction performance. However, it sometimes fails when speakers have similar voice characteristics, such as in same-gender mixtures, because it is difficult to discriminate the target speaker from the interfering speakers. In this paper, we investigate strategies for improving the speaker discrimination capability of SpeakerBeam. First, we propose a time-domain implementation of SpeakerBeam similar to that proposed for a time-domain audio separation network (TasNet), which has achieved state-of-the-art performance for speech separation. Besides, we investigate (1) the use of spatial features to better discriminate speakers when microphone array recordings are available, (2) adding an auxiliary speaker identification loss for helping to learn more discriminative voice characteristics. We show experimentally that these strategies greatly improve speech extraction performance, especially for same-gender mixtures, and outperform TasNet in terms of target speech extraction. Marc Delcroix, Tsubasa Ochiai, Katerina Zmolíková, Keisuke Kinoshita, Naohiro Tawara, Tomohiro Nakatani, Shoko Araki |
ICASSP | 6 |
| 2020 | Overdetermined Independent Vector AnalysisabstractWe address the convolutive blind source separation problem for the (over-)determined case where (i) the number of nonstationary target-sources K is less than that of microphones M, and (ii) there are up to M - K stationary Gaussian noises that need not to be extracted. Independent vector analysis (IVA) can solve the problem by separating into M sources and selecting the top K highly nonstationary signals among them, but this approach suffers from a waste of computation especially when K ≪ M. Channel reductions in preprocessing of IVA by, e.g., principle component analysis have the risk of removing the target signals. We here extend IVA to resolve these issues. One such extension has been attained by assuming the orthogonality constraint (OC) that the sample correlation between the target and noise signals is to be zero. The proposed IVA, on the other hand, does not rely on OC and exploits only the independence between sources and the stationarity of the noises. This enables us to develop several efficient algorithms based on block coordinate descent methods with a problem specific acceleration. We clarify that one such algorithm exactly coincides with the conventional IVA with OC, and also explain that the other newly developed algorithms are faster than it. Experimental results show the improved computational load of the new algorithms compared to the conventional methods. In particular, a new algorithm specialized for K = 1 outperforms the others. Rintaro Ikeshita, Tomohiro Nakatani, Shoko Araki |
ICASSP | 2 |
| 2020 | Tackling Real Noisy Reverberant Meetings with All-Neural Source Separation, Counting, and Diarization SystemabstractAutomatic meeting analysis is an essential fundamental technology required to let, e.g. smart devices follow and respond to our conversations. To achieve an optimal automatic meeting analysis, we previously proposed an all-neural approach that jointly solves source separation, speaker diarization and source counting problems in an optimal way (in a sense that all the 3 tasks can be jointly optimized through error back-propagation). It was shown that the method could well handle simulated clean (noiseless and anechoic) dialog-like data, and achieved very good performance in comparison with several conventional methods. However, it was not clear whether such all-neural approach would be successfully generalized to more complicated real meeting data containing more spontaneously-speaking speakers, severe noise and reverberation, and how it performs in comparison with the state-of-the-art systems in such scenarios. In this paper, we first consider practical issues required for improving the robustness of the all-neural approach, and then experimentally show that, even in real meeting scenarios, the all-neural approach can perform effective speech enhancement, and simultaneously outperform state-of-the-art systems. Keisuke Kinoshita, Marc Delcroix, Shoko Araki, Tomohiro Nakatani |
ICASSP | 4 |
| 2020 | Improving Noise Robust Automatic Speech Recognition with Single-Channel Time-Domain Enhancement NetworkabstractWith the advent of deep learning, research on noise-robust automatic speech recognition (ASR) has progressed rapidly. However, ASR performance in noisy conditions of single-channel systems remains unsatisfactory. Indeed, most single-channel speech enhancement (SE) methods (denoising) have brought only limited performance gains over state-of-the-art ASR back-end trained on multi-condition training data. Recently, there has been much research on neural network-based SE methods working in the time-domain showing levels of performance never attained before. However, it has not been established whether the high enhancement performance achieved by such time-domain approaches could be translated into ASR. In this paper, we show that a single-channel time-domain denoising approach can significantly improve ASR performance, providing more than 30 % relative word error reduction over a strong ASR back-end on the real evaluation data of the single-channel track of the CHiME-4 dataset. These positive results demonstrate that single-channel noise reduction can still improve ASR performance, which should open the door to more research in that direction. Keisuke Kinoshita, Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani |
ICASSP | 4 |
| 2020 | Convergence-Guaranteed Independent Positive Semidefinite Tensor Analysis Based on Student's T DistributionabstractIn this paper, we address a blind source separation (BSS) problem and propose a new extended framework of independent positive semidefinite tensor analysis (IPSDTA). IPSDTA is a state-of-the-art BSS method that enables us to take interfrequency correlations into account, but the generative model is limited within the multivariate Gaussian distribution and its parameter optimization algorithm does not guarantee stable convergence. To resolve these problems, first, we propose to extend the generative model to a parametric multivariate Student’s t distribution that can deal with various types of signal. Secondly, we derive a new parameter optimization algorithm that guarantees the monotonic nonincrease in the cost function, providing stable convergence. Experimental results reveal that the cost function in the conventional IPSDTA does not display monotonically nonincreasing properties. On the other hand, the proposed method guarantees the monotonic nonincrease in the cost function and outperforms the conventional ILRMA and IPSDTA in the source-separation performance. Tatsuki Kondo, Kanta Fukushige, Norihiro Takamune, Daichi Kitamura, Hiroshi Saruwatari, Rintaro Ikeshita, Tomohiro Nakatani |
ICASSP | 7 |
| 2020 | DNN-supported Mask-based Convolutional Beamforming for Simultaneous Denoising, Dereverberation, and Source SeparationabstractIn this article, we investigate an integrated mask-based convolutional beamforming method for performing simultaneous denoising, dereverberation, and source separation. Conventionally, it is difficult for neural network-supported mask-based source separation to perform denoising and dereverberation at the same time and for spatial clustering-based source separation to reliably solve the permutation problem in the presence of noise and reverberation. This greatly limits the application of mask-based source separation. To address this issue, we propose a method to integrate state-of-the-art techniques for mask-based beamforming into a single optimization framework. These techniques include frequency-domain Convolutional Neural Network based utterance-level Permutation Invariant Training with a large receptive field (CNN-uPIT), noisy Complex Gaussian Mixture Model based spatial clustering (noisyCGMM), and Weighted Power minimization Distortionless response (WPD) convolutional beamforming. Our experiments show that all these components are essential for accurately estimating desired speech signals in noisy reverberant multisource environments. Tomohiro Nakatani, Riki Takahashi, Tsubasa Ochiai, Keisuke Kinoshita, Rintaro Ikeshita, Marc Delcroix, Shoko Araki |
ICASSP | 1 |
| 2020 | End-to-End Training of Time Domain Audio Separation and RecognitionabstractThe rising interest in single-channel multi-speaker speech separation sparked development of End-to-End (E2E) approaches to multi-speaker speech recognition. However, up until now, state-of-the-art neural network-based time domain source separation has not yet been combined with E2E speech recognition. We here demonstrate how to combine a separation module based on a Convolutional Time domain Audio Separation Network (Conv-TasNet) with an E2E speech recognizer and how to train such a model jointly by distributing it over multiple GPUs or by approximating truncated back-propagation for the convolutional front-end. To put this work into perspective and illustrate the complexity of the design space, we provide a compact overview of single-channel multi-speaker recognition systems. Our experiments show a word error rate of 11.0% on WSJ0-2mix and indicate that our joint time domain model can yield substantial improvements over cascade DNN-HMM and monolithic E2E frequency domain systems proposed so far. Thilo von Neumann, Keisuke Kinoshita, Lukas Drude, Christoph Böddeker, Marc Delcroix, Tomohiro Nakatani, Reinhold Häb-Umbach |
ICASSP | 6 |
| 2020 | Beam-TasNet: Time-domain Audio Separation Network Meets Frequency-domain BeamformerabstractRecent studies have shown that acoustic beamforming using a microphone array plays an important role in the construction of high-performance automatic speech recognition (ASR) systems, especially for noisy and overlapping speech conditions. In parallel with the success of multichannel beamforming for ASR, in the speech separation field, the time-domain audio separation network (TasNet), which accepts a time-domain mixture as input and directly estimates the time-domain waveforms for each source, achieves remarkable speech separation performance. In light of these two recent trends, the question of whether TasNet can benefit from beamforming to achieve high ASR performance in overlapping speech conditions naturally arises. Motivated by this question, this paper proposes a novel speech separation scheme, i.e., Beam-TasNet, which combines TasNet with the frequency-domain beamformer, i.e., a minimum variance distortionless response (MVDR) beamformer, through spatial covariance computation to achieve better ASR performance. Experiments on the spatialized WSJ0-2mix corpus show that our proposed Beam-TasNet significantly outperforms the conventional TasNet without beamforming and, moreover, successfully achieves a word error rate comparable to an oracle mask-based MVDR beamformer. Tsubasa Ochiai, Marc Delcroix, Rintaro Ikeshita, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki |
ICASSP | 5 |
| 2020 | A Dynamic Stream Weight Backprop Kalman Filter for Audiovisual Speaker TrackingabstractAudiovisual speaker tracking is an application that has been tackled by a wide range of classical approaches based on Gaussian filters, most notably the well-known Kalman filter. Recently, a specific Kalman filter implementation was proposed for this task, which incorporated dynamic stream weights to explicitly control the influence of acoustic and visual observations during estimation. Inspired by recent progress in the context of integrating uncertainty estimates into modern deep learning frameworks, this paper proposes a deep neural-network-based implementation of the Kalman filter with dynamic stream weights, whose parameters can be learned via standard backpropagation. This allows for jointly optimizing the parameters of the model and the dynamic stream weight estimator in a unified framework. An experimental study on audiovisual speaker tracking shows that the proposed model shows comparable performance to state-of-the-art recurrent neural networks with the additional advantage of requiring a smaller number of parameters and providing explicit uncertainty information. Christopher Schymura, Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki, Dorothea Kolossa |
ICASSP | 5 |
| 2020 | Predicting Intelligibility of Enhanced Speech Using Posteriors Derived from DNN-Based ASR System
Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Toshio Irino |
INTERSPEECH | 5 |
| 2020 | Multi-Path RNN for Hierarchical Modeling of Long Sequential Data and its Application to Speaker Stream SeparationabstractRecently, the source separation performance was greatly improved by time-domain audio source separation based on dualpath recurrent neural network (DPRNN).DPRNN is a simple but effective model for a long sequential data.While DPRNN is quite efficient in modeling a sequential data of the length of an utterance, i.e., about 5 to 10 second data, it is harder to apply it to longer sequences such as whole conversations consisting of multiple utterances.It is simply because, in such a case, the number of time steps consumed by its internal module called inter-chunk RNN becomes extremely large.To mitigate this problem, this paper proposes a multi-path RNN (MPRNN), a generalized version of DPRNN, that models the input data in a hierarchical manner.In the MPRNN framework, the input data is represented at several (≥ 3) time-resolutions, each of which is modeled by a specific RNN sub-module.For example, the RNN sub-module that deals with the finest resolution may model temporal relationship only within a phoneme, while the RNN sub-module handling the most coarse resolution may capture only the relationship between utterances such as speaker information.We perform experiments using simulated dialogue-like mixtures and show that MPRNN has greater model capacity, and it outperforms the current state-of-the-art DPRNN framework especially in online processing scenarios. Keisuke Kinoshita, Thilo von Neumann, Marc Delcroix, Tomohiro Nakatani, Reinhold Häb-Umbach |
INTERSPEECH | 4 |
| 2020 | Computationally Efficient and Versatile Framework for Joint Optimization of Blind Speech Separation and Dereverberation
Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Hiroshi Sawada, Shoko Araki |
INTERSPEECH | 1 |
| 2020 | Multi-Talker ASR for an Unknown Number of Sources: Joint Training of Source Counting, Separation and ASRabstractMost approaches to multi-talker overlapped speech separation and recognition assume that the number of simultaneously active speakers is given, but in realistic situations, it is typically unknown. To cope with this, we extend an iterative speech extraction system with mechanisms to count the number of sources and combine it with a single-talker speech recognizer to form the first end-to-end multi-talker automatic speech recognition system for an unknown number of active speakers. Our experiments show very promising performance in counting accuracy, source separation and speech recognition on simulated clean mixtures from WSJ0-2mix and WSJ0-3mix. Among others, we set a new state-of-the-art word error rate on the WSJ0-2mix database. Furthermore, our system generalizes well to a larger number of speakers than it ever saw during training, as shown in experiments with the WSJ0-4mix database. Thilo von Neumann, Christoph Böddeker, Lukas Drude, Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani, Reinhold Häb-Umbach |
INTERSPEECH | 6 |
| 2020 | GEDI: Gammachirp envelope distortion index for predicting intelligibility of enhanced speech
Katsuhiko Yamamoto, Toshio Irino, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani |
Speech Commun. | 5 |
| 2020 | Jointly Optimal Denoising, Dereverberation, and Source SeparationabstractThis article proposes methods that can optimize a Convolutional BeamFormer (CBF) for jointly performing denoising, dereverberation, and source separation (DN+DR+SS) in a computationally efficient way. Conventionally, a cascade configuration, composed of a Weighted Prediction Error minimization (WPE) dereverberation filter followed by a Minimum Variance Distortionless Response (MVDR) beamformer, has been used as the state-of-the-art frontend of far-field speech recognition, even though this approach's overall optimality is not guaranteed. In the blind signal processing area, an approach for jointly optimizing dereverberation and source separation (DR+SS) has been proposed; however, it requires huge computing cost, and has not been extended for applications to DN+DR+SS. To overcome the above limitations, this paper develops new approaches for jointly optimizing DN+DR+SS in a computationally much more efficient way. To this end, we first present an objective function to optimize a CBF for performing DN+DR+SS based on maximum likelihood estimation on an assumption that the steering vectors of the target signals are given or can be estimated, e.g., using a neural network. This paper refers to a CBF optimized by this objective function as a weighted Minimum-Power Distortionless Response (wMPDR) CBF. Then, we derive two algorithms for optimizing a wMPDR CBF based on two different ways of factorizing a CBF into WPE filters and beamformers: one based on an extension of the conventional joint optimization approach proposed for DR+SS and another based on a novel technique. Experiments using noisy reverberant sound mixtures show that the proposed optimization approaches greatly improve the performance of the speech enhancement in comparison with the conventional cascade configuration in terms of signal distortion measures and ASR performance. The proposed approaches also greatly reduce the computing cost with improved estimation accuracy in comparison with the conventional joint optimization approach. Tomohiro Nakatani, Christoph Böddeker, Keisuke Kinoshita, Rintaro Ikeshita, Marc Delcroix, Reinhold Häb-Umbach |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2019 | Compact Network for Speakerbeam Target Speaker ExtractionabstractSpeech separation that separates a mixture of speech signals into each of its sources has been an active research topic for a long time and has seen recent progress with the advent of deep learning. A related problem is target speaker extraction, i.e. extraction of only speech of a target speaker out of a mixture, given characteristics of his/her voice. We have recently proposed SpeakerBeam, which is a neural network-based target speaker extraction method. Speaker-Beam uses a speech extraction network that is adapted to the target speaker using auxiliary features derived from an adaptation utterance of that speaker. Initially, we implemented SpeakerBeam with a factorized adaptation layer, which consists of several parallel linear transformations weighted by weights derived from the auxiliary features. The factorized layer is effective for target speech extraction, but it requires a large number of parameters. In this paper, we propose to simply scale the activations of a hidden layer of the speech extraction network with weights derived from the auxiliary features. This simpler approach greatly reduces the number of model parameters by up to 60%, making it much more practical, while maintaining a similar level of performance. We tested our approach on simulated and real noisy and reverberant mixtures, showing the potential of SpeakerBeam for real-life applications. Moreover, we showed that speech extraction performance of SpeakerBeam compares favorably with that of a state-of-the-art speech separation method with a similar network configuration. Marc Delcroix, Katerina Zmolíková, Tsubasa Ochiai, Keisuke Kinoshita, Shoko Araki, Tomohiro Nakatani |
ICASSP | 6 |
| 2019 | A Unified Framework for Feature-based Domain Adaptation of Neural Network Language ModelsabstractAn important task for language models is the adaptation of general-domain models to specific target domains. For neural network-based language models, feature-based domain adaptation has been a popular method in previous research. Conventional methods use an adaptation feature providing context information that is calculated from a topic model. However, such a topic model needs to be trained separately from the language model. To unify the language and context model training, we present an approach that combines an extractor network and a domain adaptation layer. The extractor network learns a context representation from a fixed-size window of past words and provides the context information for the adaptation layer. The benefit of our method is that the extractor network can be trained jointly with the language model in a single training step. Our proposed method showed superior performance over conventional domain adaptation with topic features on a dataset of TED talks with respect to perplexity and word error rate after 100-best rescoring. Michael Hentschel, Marc Delcroix, Atsunori Ogawa, Tomoharu Iwata, Tomohiro Nakatani |
ICASSP | 5 |
| 2019 | Joint Optimization of Neural Network-based WPE Dereverberation and Acoustic Model for Robust Online ASRabstractSignal dereverberation using the Weighted Prediction Error (WPE) method has been proven to be an effective means to raise the accuracy of far-field speech recognition. First proposed as an iterative algorithm, follow-up works have reformulated it as a recursive least squares algorithm and therefore enabled its use in online applications. For this algorithm, the estimation of the power spectral density (PSD) of the anechoic signal plays an important role and strongly influences its performance. Recently, we showed that using a neural network PSD estimator leads to improved performance for online automatic speech recognition. This, however, comes at a price. To train the network, we require parallel data, i.e., utterances simultaneously available in clean and reverberated form. Here we propose to overcome this limitation by training the network jointly with the acoustic model of the speech recognizer. To be specific, the gradients computed from the cross-entropy loss between the target senone sequence and the acoustic model network output is backpropagated through the complex-valued dereverberation filter estimation to the neural network for PSD estimation. Evaluation on two databases demonstrates improved performance for on-line processing scenarios while imposing fewer requirements on the available training data and thus widening the range of applications. Jahn Heymann, Lukas Drude, Reinhold Häb-Umbach, Keisuke Kinoshita, Tomohiro Nakatani |
ICASSP | 5 |
| 2019 | FastMNMF: Joint Diagonalization Based Accelerated Algorithms for Multichannel Nonnegative Matrix FactorizationabstractA multichannel extension of nonnegative matrix factorization (NMF) for audio/music data, called multichannel NMF (MNMF), has been proposed by Sawada et al ["Multichannel extensions of non-negative matrix factorization with complex-valued data IEEE Trans. ASLP, vol. 21, no. 5, pp. 971-982, May 2013]. However, conventional MNMF algorithms have a major drawback of a heavy computational load due to numerous matrix operations, such as matrix inversions and matrix multiplications. Here we propose FastMNMF, accelerated algorithms for the MNMF based on joint diagonalization of matrices. It is well known that, for diagonal matrices, matrix operations reduce to mere scalar operations on diagonal entries. Because of this property, the joint diagonalization results in a significantly reduced computational load compared to conventional MNMF algorithms. This makes the proposed FastMNMF even applicable to a situation with alarge database or restricted computational resources. Nobutaka Ito, Tomohiro Nakatani |
ICASSP | 2 |
| 2019 | Semi-supervised End-to-end Speech Recognition Using Text-to-speech and AutoencodersabstractWe introduce speech and text autoencoders that share encoders and decoders with an automatic speech recognition (ASR) model to improve ASR performance with large speech only and text only training datasets. To build the speech and text autoencoders, we leverage state-of-the-art ASR and text-to-speech (TTS) encoder decoder architectures. These autoencoders learn features from speech only and text only datasets by switching the encoders and decoders used in the ASR and TTS models. Simultaneously, they aim to encode features to be compatible with ASR and TTS models by a multi-task loss. Additionally, we anticipate that TTS joint training can also improve the ASR performance because both ASR and TTS models learn transformations between speech and text. The experimental result we obtained with our semi-supervised end-to-end ASR/TTS training revealed reductions from a model initially trained with a small paired subset of the LibriSpeech corpus in the character error rate from 10.4% to 8.4% and word error rate from 20.6% to 18.0% by retraining the model with a large unpaired subset of the corpus. Shigeki Karita, Shinji Watanabe 0001, Tomoharu Iwata, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani |
ICASSP | 6 |
| 2019 | Mask-based MVDR Beamformer for Noisy Multisource Environments: Introduction of Time-varying Spatial Covariance ModelabstractThis paper proposes a method for designing a time-varying minimum variance distortionless response (MVDR) beamformer using time-frequency masks, with the aim of improving speech enhancement in noisy multi-speaker environments. A key to successful beamforming is to estimate accurately a time-varying spatial covariance matrix (SCM) for noise composed of both stationary diffuse noise and highly time-varying speech. For this purpose, we introduce a stochastic model that can represent the time-varying characteristics of the noise SCM, and derive a method for estimating a time-varying noise SCM based on the model. Experiments show that the proposed method can substantially improve the performance of the beamformer in terms of automatic speech recognition (ASR) accuracy and source-to-distortion ratio compared with a conventional time-invariant MVDR beamformer. Yuki Kubo, Tomohiro Nakatani, Marc Delcroix, Keisuke Kinoshita, Shoko Araki |
ICASSP | 2 |
| 2019 | All-neural Online Source Separation, Counting, and Diarization for Meeting AnalysisabstractAutomatic meeting analysis comprises the tasks of speaker counting, speaker diarization, and the separation of overlapped speech, followed by automatic speech recognition. This all has to be carried out on arbitrarily long sessions and, ideally, in an online or block-online manner. While significant progress has been made on individual tasks, this paper presents for the first time an all-neural approach to simultaneous speaker counting, diarization and source separation. The NN-based estimator operates in a block-online fashion and tracks speakers even if they remain silent for a number of time blocks, thus learning a stable output order for the separated sources. The neural network is recurrent over time as well as over the number of sources. The simulation experiments show that state of the art separation performance is achieved, while at the same time delivering good diarization and source counting results. It even generalizes well to an unseen large number of blocks. Thilo von Neumann, Keisuke Kinoshita, Marc Delcroix, Shoko Araki, Tomohiro Nakatani, Reinhold Häb-Umbach |
ICASSP | 5 |
| 2019 | A Unified Framework for Neural Speech Separation and ExtractionabstractThe development of deep learning techniques has triggered the active investigation of neural network-based speech enhancement approaches. In particular, single-channel blind (uninformed) speech separation and speaker-aware (informed) speech extraction have received increased interest. Blind speech separation separates a speech mixture into all source signals without requiring any auxiliary information about the speakers. In contrast, speaker-aware speech extraction focuses on extracting speech from a target speaker using prior knowledge, such as an utterance spoken by the target speaker. Speaker extraction is therefore not fully blind, but it can mitigate the source permutation problem faced by blind source separation, and potentially achieve better speech quality by exploiting the auxiliary information. In this paper, to take advantage of both approaches, we propose a unified framework for both speech separation and speech extraction using a single model. This is realized by incorporating a speaker attention mechanism within a generalized permutation invariant training (PIT)-based blind speech separation model, and introducing a multitask separation/extraction objective for training the model. Experiments on the WSJ0-2mix dataset show that our proposed framework realizes both uninformed separation and informed extraction, and achieves better separation/extraction performance than a baseline PIT-based model. Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Tomohiro Nakatani |
ICASSP | 5 |
| 2019 | ILP-based Compressive Speech Summarization with Content Word Coverage Maximization and Its Oracle Performance AnalysisabstractWe propose an integer linear programming (ILP)-based compressive speech summarization method that maximizes the coverage of content words in a resultant summary. It is an unsupervised method and, under the designed constraints, it performs a single-step globally optimal summarization of a given long speech recording, which is decoded as a confusion network form of an automatic speech recognition (ASR) hypothesis sequence. It selects as many different content words as possible from the speech input that inevitably includes a high level of redundancy (e.g. the repetition of the same word) under a given length constraint. In experiments using a lecture speech corpus, we obtained higher summarization performance in terms of ROUGE scores than with a baseline extractive summarization method. We further conduct experimental analyses to obtain the oracle (upper bound) performance of the summarization methods. The analysis results show that the oracle performance is very high even though the ASR hypotheses include recognition errors. It is significantly higher than the system performance and, in addition, the oracle performance of the compressive method is significantly higher than that of the extractive method. These results confirm that our method is a promising approach. Atsunori Ogawa, Tsutomu Hirao, Tomohiro Nakatani, Masaaki Nagata |
ICASSP | 3 |
| 2019 | Predicting Speech Intelligibility of Enhanced Speech Using Phone Accuracy of DNN-Based ASR System
Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Katsuhiko Yamamoto, Toshio Irino |
INTERSPEECH | 5 |
| 2019 | End-to-End SpeakerBeam for Single Channel Target Speech Recognition
Marc Delcroix, Shinji Watanabe 0001, Tsubasa Ochiai, Keisuke Kinoshita, Shigeki Karita, Atsunori Ogawa, Tomohiro Nakatani |
INTERSPEECH | 7 |
| 2019 | Improving Transformer-Based End-to-End Speech Recognition with Connectionist Temporal Classification and Language Model Integration
Shigeki Karita, Nelson Enrique Yalta Soplin, Shinji Watanabe 0001, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani |
INTERSPEECH | 6 |
| 2019 | Simultaneous Denoising and Dereverberation for Low-Latency Applications Using Frame-by-Frame Online Unified Convolutional Beamformer
Tomohiro Nakatani, Keisuke Kinoshita |
INTERSPEECH | 1 |
| 2019 | Multimodal SpeakerBeam: Single Channel Target Speech Extraction with Audio-Visual Speaker Clues
Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Tomohiro Nakatani |
INTERSPEECH | 5 |
| 2019 | Improved Deep Duel Model for Rescoring N-Best Speech Recognition List Using Backward LSTMLM and Ensemble Encoders
Atsunori Ogawa, Marc Delcroix, Shigeki Karita, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2019 | A Unified Convolutional Beamformer for Simultaneous Denoising and DereverberationabstractThis letter proposes a method for estimating a convolutional beamformer that can perform denoising and dereverberation simultaneously in an optimal way. The application of dereverberation based on a weighted prediction error (WPE) method followed by denoising based on a minimum variance distortionless response (MVDR) beamformer has conventionally been considered a promising approach, however, the optimality of this approach cannot be guaranteed. To realize the optimal integration of denoising and dereverberation, we present a method that unifies the WPE dereverberation method and a variant of the MVDR beamformer, namely a minimum power distortionless response beamformer, into a single convolutional beamformer, and we optimize it based on a single unified optimization criterion. The proposed beamformer is referred to as a weighted power minimization distortionless response beamformer. Experiments show that the proposed method substantially improves the speech enhancement performance in terms of both objective speech enhancement measures and automatic speech recognition performance. Tomohiro Nakatani, Keisuke Kinoshita |
IEEE Signal Process. Lett. | 1 |
| 2018 | Permutation-Free Cgmm: Complex Gaussian Mixture Model with Inverse Wishart Mixture Model Based Spatial Prior for Permutation-Free Source Separation and Source CountingabstractHere we propose a permutation-free cGMM (PF-cGMM), a new probabilistic model of observed mixtures, which can resolve permutation ambiguity between frequency bins, and is applicable even when the number of sources is unknown. A recently proposed complex Gaussian mixture model (cGMM) is highly effective for frequency bin-wise clustering when the number of sources is known. However, it cannot resolve the permutation ambiguity, and is inapplicable when the number of sources is unknown. The proposed PF-cGMM is an extension of the cGMM, which resolves these issues. The resolution of the permutation ambiguity can be realized by a spatial prior called a complex inverse Wishart mixture model (cIWMM). The absence of the permutation ambiguity facilitates source counting, which is performed by hierarchical clustering in this paper. Experiments showed that the PF-cGMM was able to (1) resolve the permutation ambiguity and (2) realize source separation even when the number of sources was unknown with little performance degradation compared to when it was known. Juan Azcarreta, Nobutaka Ito, Shoko Araki, Tomohiro Nakatani |
ICASSP | 4 |
| 2018 | Single Channel Target Speaker Extraction and Recognition with Speaker BeamabstractThis paper addresses the problem of single channel speech recognition of a target speaker in a mixture of speech signals. We propose to exploit auxiliary speaker information provided by an adaptation utterance from the target speaker to extract and recognize only that speaker. Using such auxiliary information, we can build a speaker extraction neural network (NN) that is independent of the number of sources in the mixture, and that can track speakers across different utterances, which are two challenging issues occurring with conventional approaches for speech recognition of mixtures. We call such an informed speaker extraction scheme “SpeakerBeam”. SpeakerBeam exploits a recently developed context adaptive deep NN (CADNN) that allows tracking speech from a target speaker using a speaker adaptation layer, whose parameters are adjusted depending on auxiliary features representing the target speaker characteristics. SpeakerBeam was previously investigated for speaker extraction using a microphone array. In this paper, we demonstrate that it is also efficient for single channel speaker extraction. The speaker adaptation layer can be employed either to build a speaker adaptive acoustic model that recognizes only the target speaker or a mask-based speaker extraction network that extracts the target speech from the speech mixture signal prior to recognition. We also show that the latter speaker extraction network can be optimized jointly with an acoustic model to further improve ASR performance. Marc Delcroix, Katerina Zmolíková, Keisuke Kinoshita, Atsunori Ogawa, Tomohiro Nakatani |
ICASSP | 5 |
| 2018 | Dual Frequency- and Block-Permutation Alignment for Deep Learning Based Block-Online Blind Source SeparationabstractDeep attractor networks (DANs) are a recently introduced method to blindly separate sources from spectral features of a monaural recording using bidirectional long short-term memory networks (BLSTMs). Due to the nature of BLSTMs, this is inherently not online-ready and resorting to operating on blocks yields a block permutation problem in that the index of each speaker may change between blocks. We here propose the joint modeling of spatial and spectral features to solve the block permutation problem and generalize DANs to multi-channel meeting recordings: The DAN acts as a spectral feature extractor for a subsequent model-based clustering approach. We first analyze different joint models in batch-processing scenarios and finally propose a block-online blind source separation algorithm. The efficacy of the proposed models is demonstrated on reverberant mixtures corrupted by real recordings of multi-channel background noise. We demonstrate that both the proposed batch-processing and the proposed block-online system outperform (a) a spatial-only model with a state-of-the-art frequency permutation solver and (b) a spectral-only model with an oracle block permutation solver in terms of signal to distortion ratio (SDR) gains. Lukas Drude, Takuya Higuchi, Keisuke Kinoshita, Tomohiro Nakatani, Reinhold Häb-Umbach |
ICASSP | 4 |
| 2018 | Frame-by-Frame Closed-Form Update for Mask-Based Adaptive MVDR BeamformingabstractBeamforming approaches using time-frequency masks have recently been investigated and have shown promising results for noise robust automatic speech recognition (ASR) in many tasks. The time-frequency masks are estimated to compute the spatial statistics of target speech and noise signals, and then the statistics are used to derive a beamformer. Although its effectiveness has been clearly shown in batch and blockwise processing, it has not been well extended to frame-by-frame processing, which is a very important procedure for many actual applications. In this paper, we derive a frame-by-frame update rule for a mask-based minimum variance distortion-less response (MVDR) beamformer, which enables us to obtain enhanced signals without a long delay by combining it with uni-directional recurrent neural network-based mask estimation. Based on the Woodbury matrix identity, our algorithm achieves a closed-form solution of the mask-based MVDR beamformer at every time frame without any matrix inversion. Experimental results show that our frame-by-frame beamformer outperforms baseline block-wise beamforming on the CHiME-3 simulation dataset even with a shorter time delay. Takuya Higuchi, Keisuke Kinoshita, Nobutaka Ito, Shigeki Karita, Tomohiro Nakatani |
ICASSP | 5 |
| 2018 | Maximum-Likelihood Online Speaker Diarization in Noisy Meetings Based on Categorical Mixture Model and Probabilistic Spatial DictionaryabstractIn this paper, we propose a maximum-likelihood online diarization method based on a probabilistic spatial dictionary. This dictionary consists of the given probability distribution of spatial features for each possible direction of arrival (DOA) of source signals. Recently, we have developed an online, noise-robust diarization method by utilizing this dictionary as spatial prior information. In this method, DOA estimation is first performed frame-wise based on the dictionary, and subsequently diarization is performed. Although the DOA estimation is performed optimally in the maximum-likelihood sense, the diarization is performed suboptimally based on some heuristics. In contrast, the proposed method performs DOA estimation and diarization jointly and optimally in the maximum-likelihood sense. This is realized by introducing a categorical mixture model (CMM), which has source-wise DOA information and diarization information as unknown parameters. We conducted an experiment on a real-world meeting dataset, and confirmed that the proposed method reduced a diarization error rate by absolute 2.7% compared to the above conventional method. Nobutaka Ito, Takashi Makino, Shoko Araki, Tomohiro Nakatani |
ICASSP | 4 |
| 2018 | Sequence Training of Encoder-Decoder Model Using Policy Gradient for End-to-End Speech RecognitionabstractThe standard evaluation metric of automatic speech recognition (ASR) is the word error rate (WER), which measures the dissimilarity between recognized word sequences and their ground truth. Many training algorithms designed to reduce sequence-level errors such as WER have been proposed for hidden Markov model (HMM)-based ASR, e.g., state-level minimum Bayes risk (sMBR). However, these approaches cannot be used directly for encoder-decoder model based end-to-end ASR, because the encoder-decoder model employs very different mechanisms from HMM-based approaches. In this paper, we propose a new method for optimizing the encoder-decoder model based on a sequence-level evaluation metric. Since the WER is not directly differentiable, we adopt a policy gradient objective function to train the encoder-decoder model, which enables us to minimize the expected WER of the model predictions. This training method employs the scoring of multiple hypotheses as in the decoding stage while usual cross entropy training uses only the ground truth. Therefore, we can expect it to improve the decoding results of the encoder-decoder model. We perform experiments using the Tedlium corpus to demonstrate the potential of our proposed method for improving the recognition performance of the encoder-decoder model. Shigeki Karita, Atsunori Ogawa, Marc Delcroix, Tomohiro Nakatani |
ICASSP | 4 |
| 2018 | Listening to Each Speaker One by One with Recurrent Selective Hearing NetworksabstractDeep learning-based single-channel source separation algorithms are currently being actively investigated. Among them, Deep Clustering (DC) and Deep Attractor Networks (DANs) have made it possible to separate an arbitrary number of speakers. In particular, they cleverly combine a neural network and a K-means clustering algorithm to obtain source separation masks with the assumption that the correct number of speakers at the test time is known in advance. Unlike DC and DAN, Permutation Invariant Training (PIT) was proposed as a purely neural network-based mask estimator. Essentially, however, PIT can deal with only a fixed number of speakers, given the strong relationship between the dimensions of the output nodes and the assumed number of sources. Considering these limitations and merits of such conventional methods, this paper proposes a purely neural-network based mask estimator that can handle an arbitrary number of sources, and simultaneously estimate the number of sources in the test signal. To accomplish this, while the conventional methods deal with the source separation problem as a one-pass problem, we cast the problem as a recursive multi-pass source extraction problem based on a recurrent neural network (RNN) that can learn and determine how many computational steps/iterations have to be performed depending on the input signals. In this paper, we describe our proposed method in detail, and experimentally show its efficacy in terms of source separation and source counting performance. Keisuke Kinoshita, Lukas Drude, Marc Delcroix, Tomohiro Nakatani |
ICASSP | 4 |
| 2018 | Rescoring N-Best Speech Recognition List Based on One-on-One Hypothesis Comparison Using Encoder-Classifier ModelabstractThis paper proposes a new model for accurately rescoring (reranking) N-best speech recognition hypothesis lists. The model is based on state-of-the-art neural networks (NNs) and provides the minimum necessary functionality to perform N-best rescoring, i.e. one-on-one hypothesis comparison on a given N-best list in terms of word error rates (WERs). The model is composed of a long short-term memory (LSTM)-based encoder network followed by a fully-connected feedforward NN-based binary-class classifier network. Given the feature vector sequences of two hypotheses to be compared, this encoder-classifier (EC) model encodes these features and outputs binary-class probabilities that indicate which hypothesis has the lower WER. Then, depending on the output, the ranks of these hypotheses can be swapped. By repeating this one-on-one hypothesis comparison, a reranked N-best list can be obtained. In N-best rescoring experiments using a large scale speech corpus, the proposed EC model steadily outperforms an LSTM-based language model (LSTMLM), which is a strong and widely-used competitor. In addition, by incorporating the LSTMLM scores as an additional feature vector dimension, the N-best rescoring performance of the EC model is further improved. The improved EC model achieves a 10% relative WER reduction from the LSTMLM baseline. Atsunori Ogawa, Marc Delcroix, Shigeki Karita, Tomohiro Nakatani |
ICASSP | 4 |
| 2018 | Optimization of Speaker-Aware Multichannel Speech Extraction with ASR CriterionabstractThis paper addresses the problem of recognizing speech corrupted by overlapping speakers in a multichannel setting. To extract a target speaker from the mixture, we use a neural network based beamformer which uses masks estimated by a neural network to compute statistically optimal spatial filters. Following our previous work, we inform the neural network about the target speaker using information extracted from an adaptation utterance’ enabling the network to track the target speaker. While in the previous work, this method was used to separately extract the speaker and then pass such preprocessed speech to a speech recognition system, here we explore training both systems jointly with a common speech recognition criterion. We show that integrating the two systems and training for the final objective improves the performance. In addition, the integration enables further sharing of information between the acoustic model and the speaker extraction system, by making use of the predicted HMM-state posteriors to refine the masks used for beamforming. Katerina Zmolíková, Marc Delcroix, Keisuke Kinoshita, Takuya Higuchi, Tomohiro Nakatani, Jan Cernocký |
ICASSP | 5 |
| 2018 | Auxiliary Feature Based Adaptation of End-to-end ASR Systems
Marc Delcroix, Shinji Watanabe 0001, Atsunori Ogawa, Shigeki Karita, Tomohiro Nakatani |
INTERSPEECH | 5 |
| 2018 | Integrating Neural Network Based Beamforming and Weighted Prediction Error DereverberationabstractThe weighted prediction error (WPE) algorithm has proven to be a very successful dereverberation method for the REVERB challenge. Likewise, neural network based mask estimation for beamforming demonstrated very good noise suppression in the CHiME 3 and CHiME 4 challenges. Recently, it has been shown that this estimator can also be trained to perform dereverberation and denoising jointly. However, up to now a comparison of a neural beamformer and WPE is still missing, so is an investigation into a combination of the two. Therefore, we here provide an extensive evaluation of both and consequently propose variants to integrate deep neural network based beamforming with WPE. For these integrated variants we identify a consistent word error rate (WER) reduction on two distinct databases. In particular, our study shows that deep learning based beamforming benefits from a model-based dereverberation technique (i.e. WPE) and vice versa. Our key findings are: (a) Neural beamforming yields the lower WERs in comparison to WPE the more channels and noise are present. (b) Integration of WPE and a neural beamformer consistently outperforms all stand-alone systems. Lukas Drude, Christoph Böddeker, Jahn Heymann, Reinhold Häb-Umbach, Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani |
INTERSPEECH | 7 |
| 2018 | Multi-resolution Gammachirp Envelope Distortion Index for Intelligibility Prediction of Noisy Speech
Katsuhiko Yamamoto, Toshio Irino, Narumi Ohashi, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani |
INTERSPEECH | 6 |
| 2018 | Distortionless Beamforming Optimized With ℓ1-Norm MinimizationabstractWe propose beamforming method that minimizes the ℓ1norm of a beamformer output vector under the same distortionless constraint as that of the conventional minimum power distortionless response (MPDR) beamformer. Using the ℓ1norm makes the beamformer output sparse. This leads to reducing the residual elements of the interference signal. In addition, the sensitivity of the proposed beamformer can be controlled by adding a norm constraint as in the MPDR beamformer. The proposed method improved the signal-to-interference-noise ratio by 7 dB from that of the MPDR beamformer for reverberation time T60= 300 ms in a simulation. Satoru Emura, Shoko Araki, Tomohiro Nakatani, Noboru Harada |
IEEE Signal Process. Lett. | 3 |
| 2018 | Context Adaptive Neural Network Based Acoustic Models for Rapid AdaptationabstractThe adaptation of automatic speech recognition systems to a speaker or an environment is important if we are to achieve high speech recognition performance ubiquitously. Recently, deep neural network (DNN) based acoustic models have been made adaptive to speakers or environments by the addition of an auxiliary feature representing the acoustic context information such as speaker or noise characteristics to the network input. The addition of such auxiliary features to the input realizes only the adaptation of the bias term of the input layer. In this paper, we introduce “context adaptive neural networks,” which are an alternative approach for exploiting auxiliary features that can achieve adaptation of all the parameters of a layer including the linear transformation matrices and the bias terms. A context adaptive neural network is a neural network with one of its layers factorized into sublayers, each associated with an acoustic context class representing a class of speakers or noise conditions. The output of the factorized layer is obtained as a weighted sum of the contributions of all of the sublayers. The weighting coefficients, or context class weights, are derived from the auxiliary features, by transforming them through an auxiliary network. The auxiliary network and the main network can be trained jointly, which enables the context classes that optimize the training criterion to be learned automatically. We perform experiments on three tasks, i.e., two speaker adaptation experiments using DNN models with medium-sized (Wall Street Journal) and large (Continuous Spontaneous Japanese) training datasets, and one environmental adaptation of a convolutional neural network based acoustic model with CHiME3 data. These experiments confirm the potential of the proposed approach in various settings. Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Christian Huemmer 0001, Tomohiro Nakatani |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | Adversarial training for data-driven speech enhancement without parallel corpusabstractThis paper describes a way of performing data-driven speech enhancement for noise robust automatic speech recognition (ASR), where we train a model for speech enhancement without a parallel corpus. Data-driven speech enhancement with deep models has recently been investigated and proven to be a promising approach for ASR. However, for model training, we need a parallel corpus consisting of noisy speech signals and corresponding clean speech signals for supervision. Therefore a deep model can be trained only with a simulated dataset, and we cannot take advantage of a large number of noisy recordings that do not have corresponding clean speech signals. As a first step towards model training without supervision, this paper proposes a novel approach introducing adversarial training for a time-frequency mask estimator. Our cost function for model training is defined by discriminators instead of by using the distance between the model outputs and the supervision. The discriminators distinguish between true signals and enhanced signals obtained with time-frequency masks estimated with a mask estimator. The mask estimator is trained to cheat the discriminators, which enables the mask estimator to estimate the appropriate time-frequency masks without a parallel corpus. The enhanced signal is finally obtained with masking-based beamforming. Experimental results show that, even without exploiting parallel data, our speech enhancement approach achieves improved ASR performance compared with results obtained with unprocessed signals and achieves comparable ASR performance to that obtained with a model trained with a parallel corpus based on a minimum mean squared error (MMSE) criterion. Takuya Higuchi, Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani |
ASRU | 4 |
| 2017 | Learning speaker representation for neural network based multichannel speaker extractionabstractRecently, schemes employing deep neural networks (DNNs) for extracting speech from noisy observation have demonstrated great potential for noise robust automatic speech recognition. However, these schemes are not well suited when the interfering noise is another speaker. To enable extracting a target speaker from a mixture of speakers, we have recently proposed to inform the neural network using speaker information extracted from an adaptation utterance from the same speaker. In our previous work, we explored ways how to inform the network about the speaker and found a speaker adaptive layer approach to be suitable for this task. In our experiments, we used speaker features designed for speaker recognition tasks as the additional speaker information, which may not be optimal for the speaker extraction task. In this paper, we propose a usage of a sequence summarizing scheme enabling to learn the speaker representation jointly with the network. Furthermore, we extend the previous experiments to demonstrate the potential of our proposed method as a front-end for speech recognition and explore the effect of additional noise on the performance of the method. Katerina Zmolíková, Marc Delcroix, Keisuke Kinoshita, Takuya Higuchi, Atsunori Ogawa, Tomohiro Nakatani |
ASRU | 6 |
| 2017 | Unsupervised utterance-wise beamformer estimation with speech recognition-level criterionabstractIn this paper, we perform beamforming with a speech recognition-level criterion. A beamformer is usually designed by optimizing signal-level criteria, e.g., by minimizing the beamformer output covariance or by maximizing the signal-to-noise ratio (SNR). Such signal-level criteria do not always guarantee that the optimized beamformer is the best for noise robust automatic speech recognition. Recently, a few approaches have been proposed for performing beamforming with a speech recognition-level criterion. These approaches train beamformers along with an acoustic model by using multichannel training data and a parallel corpus of noisy and clean data. This paper proposes a novel approach for estimating the beamformer for every test utterance with a speech recognition-level criterion. We use an unsupervised acoustic model adaptation scheme to optimize our beamformer. Specifically, we first obtain decoding results with an initialized beamformer, and then we optimize our beamformer using back propagation to minimize the cross entropy between the first-pass decoding results and actual network outputs. With this approach, our beamformer can be trained to discriminate hidden Markov model states more clearly for every test utterance. Experimental results show that our beamformer outperforms a beamformer designed with a signal-level criterion. Takuya Higuchi, Takuya Yoshioka, Keisuke Kinoshita, Tomohiro Nakatani |
ICASSP | 4 |
| 2017 | Online environmental adaptation of CNN-based acoustic models using spatial diffuseness featuresabstractWe propose a new concept for adapting CNN-based acoustic models using spatial diffuseness features as auxiliary information about the acoustic environment: the spatial diffuseness features are simultaneously employed as acoustic-model input features and to estimate environmental cues for context adaptation, where one convolutional layer is factorized into several sub-layers to represent different acoustic conditions. This context-adaptive CNN-based acoustic model facilitates an online environmental adaptation and is experimentally verified for the real-world recordings provided by the CHiME-3 task. The best performing setup reduces the average word error rate scores achieved by the baseline system (without using spatial diffuseness features) from 19.4% to 15.9% and 12.2% to 10.7% considering two experimental setups with and without front-end signal enhancement, respectively. Christian Huemmer 0001, Marc Delcroix, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Walter Kellermann |
ICASSP | 5 |
| 2017 | Probabilistic spatial dictionary based online adaptive beamforming for meeting recognition in noisy and reverberant environmentsabstractHere we propose online adaptive beamforming for automatic speech recognition (ASR) in meetings in noisy, reverberant environments. The proposed method is based on recently developed mask-based beamforming, in which accurate mask estimation and diarization are paramount. Real-world experiments have shown that mask-based beamforming enables accurate ASR in meetings in small noise and reverberation with a signal-to-noise ratio (SNR) of 15–25 dB and a reverberation time (RT) of 120–350 ms. In this paper, we deal with a more adverse condition: meetings in large noise and reverberation with an SNR of 3–15 dB and an RT of 500 ms. To this end, we exploit a probabilistic spatial dictionary, a dictionary that consists of a pre-trained probability distribution of source location features for each potential speaker location. This dictionary enables us to perform mask estimation and diarization for beamforming accurately, even in the above adverse condition. The proposed method reduced the word error rate (WER) on real meeting data by 54.8% relative to our previous beamforming method. Nobutaka Ito, Shoko Araki, Marc Delcroix, Tomohiro Nakatani |
ICASSP | 4 |
| 2017 | Deep mixture density network for statistical model-based feature enhancementabstractWe propose a novel framework designed to extend conventional deep neural network (DNN)-based feature enhancement approaches. In general, the conventional DNN-based feature enhancement framework aims to map input noisy observation to clean speech or a binary/ soft mask in a deterministic way, assuming that there is one-to-one mapping between the input and the output without any uncertainty. However, when we consider that the general feature enhancement problem to be an ill-posed inverse problem where the mapping cannot be uniquely determined given an input signal, the assumption in the conventional approaches is not theoretically correct and potentially limits the performance of DNN-based feature enhancement. To overcome this problem, this paper proposes utilizing a mixture density network (MDN), which is a neural network that maps an input feature to a set of Gaussian mixture model (GMM) parameters representing the distribution of a target variable. By estimating the distribution of clean speech feature based on MDN, we are now able to explicitly consider the uncertainty in the parameter estimation. Then, we further utilizes the estimated GMM to obtain a refined clean speech estimate in the framework of statistical model-based feature enhancement. In this paper, after detailing the proposed framework and the MDN, we show mathematically and experimentally how MDN appropriately models the uncertainty information. We also show that the proposed method can outperform a conventional DNN-based feature enhancement method. Keisuke Kinoshita, Marc Delcroix, Atsunori Ogawa, Takuya Higuchi, Tomohiro Nakatani |
ICASSP | 5 |
| 2017 | Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamformingabstractRecently, time-frequency mask-based beamforming has been extensively studied as the frontend of deep neural network (DNN) based automatic speech recognition (ASR) in noisy environments. Two mask estimation approaches have been separately developed for this beamforming method, namely the the DNN-based approach, which exploits the time-frequency features of the signal, and the spatial clustering-based approach, which exploits the spatial features of the signal. This paper proposes a new method that integrates the two approaches in a probabilistic way to further improve mask estimation by exploiting the advantages of both approaches. Experiments using the real data of the CHiME-3 multichannel noisy speech corpus show that the proposed method almost always outperforms the conventional approaches in terms of word error rate (WER) improvement. Tomohiro Nakatani, Nobutaka Ito, Takuya Higuchi, Shoko Araki, Keisuke Kinoshita |
ICASSP | 1 |
| 2017 | Cumulative moving averaged bottleneck speaker vectors for online speaker adaptation of CNN-based acoustic modelsabstractAdapting acoustic models to speakers have shown to greatly improve performance for many tasks. Among the adaptation approaches, exploiting auxiliary features characterizing speakers or environments has received great attention because they allow rapid adaptation, i.e. adaptation with limited amount of speech data such as a single utterance. However, the auxiliary features are usually computed in batch mode, which causes some inevitable latency. In this paper we explore an extension of the auxiliary feature-based adaptation to online processing. We employ auxiliary features obtained from bottleneck speaker vectors and extend their computation to online processing using cumulative moving averaging. We test our proposed approach for deep CNN-based acoustic models, using context adaptive networks to exploit the auxiliary features. Experimental results on the CHiME-3 task demonstrate that the proposed approach can realize online speaker adaptation. Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Taichi Asami, Shigeru Katagiri, Tomohiro Nakatani |
ICASSP | 7 |
| 2017 | Feedback connection for deep neural network-based acoustic modelingabstractThe use of auxiliary features is an effective way to improve the performance of deep neural network (DNN)-based acoustic models. Most approaches use auxiliary features that represent the speaker or the environment. These auxiliary features are usually computed independently of the acoustic model. This paper investigates a types of auxiliary features obtained from the output of a hidden layer that feeds back to the input layer of the network. Since the auxiliary features are extracted from the hidden layer of the network no external information is required such as the speaker or the environment. Experimentally, by forcing the extraction of the auxiliary features from the same networks, we can further improve the performance of the overall network and reduce the total number of parameters used. We tested this approach with different deep neural network architectures including: deep neural networks, convolutional neural networks and unfolded recurrent convolutional networks. We confirmed the effectiveness of this approach on the CHiME3 dataset. Dung T. Tran, Marc Delcroix, Atsunori Ogawa, Christian Huemmer 0001, Tomohiro Nakatani |
ICASSP | 5 |
| 2017 | Deep Clustering-Based Beamforming for Separation with Unknown Number of Sources
Takuya Higuchi, Keisuke Kinoshita, Marc Delcroix, Katerina Zmolíková, Tomohiro Nakatani |
INTERSPEECH | 5 |
| 2017 | Forward-Backward Convolutional LSTM for Acoustic Modeling
Shigeki Karita, Atsunori Ogawa, Marc Delcroix, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2017 | Neural Network-Based Spectrum Estimation for Online WPE Dereverberation
Keisuke Kinoshita, Marc Delcroix, Haeyong Kwon, Takuma Mori, Tomohiro Nakatani |
INTERSPEECH | 5 |
| 2017 | Improved Example-Based Speech Enhancement by Using Deep Neural Network Acoustic Model for Noise Robust Example Search
Atsunori Ogawa, Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2017 | Unfolded Deep Recurrent Convolutional Neural Network with Jump Ahead Connections for Acoustic Modeling
Dung T. Tran, Marc Delcroix, Shigeki Karita, Michael Hentschel, Atsunori Ogawa, Tomohiro Nakatani |
INTERSPEECH | 6 |
| 2017 | Uncertainty Decoding with Adaptive Sampling for Noise Robust DNN-Based Acoustic Modeling
Dung T. Tran, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2017 | Predicting Speech Intelligibility Using a Gammachirp Envelope Distortion Index Based on the Signal-to-Distortion Ratio
Katsuhiko Yamamoto, Toshio Irino, Toshie Matsui, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani |
INTERSPEECH | 6 |
| 2017 | Speaker-Aware Neural Network Based Beamformer for Speaker Extraction in Speech Mixtures
Katerina Zmolíková, Marc Delcroix, Keisuke Kinoshita, Takuya Higuchi, Atsunori Ogawa, Tomohiro Nakatani |
INTERSPEECH | 6 |
| 2017 | Online MVDR Beamformer Based on Complex Gaussian Mixture Model With Spatial Prior for Noise Robust ASRabstractThis paper considers acoustic beamforming for noise robust automatic speech recognition. A beamformer attenuates background noise by enhancing sound components coming from a direction specified by a steering vector. Hence, accurate steering vector estimation is paramount for successful noise reduction. Recently, time-frequency masking has been proposed to estimate the steering vectors that are used for a beamformer. In particular, we have developed a new form of this approach, which uses a speech spectral model based on a complex Gaussian mixture model (CGMM) to estimate the time-frequency masks needed for steering vector estimation, and extended the CGMM-based beamformer to an online speech enhancement scenario. Our previous experiments showed that the proposed CGMM-based approach outperforms a recently proposed mask estimator based on a Watson mixture model and the baseline speech enhancement system of the CHiME-3 challenge. This paper provides additional experimental results for our online processing, which achieves performance comparable to that of batch processing with a suitable block-batch size. This online version reduces the CHiME-3 word error rate (WER) on the evaluation set from 8.37% to 8.06%. Moreover, in this paper, we introduce a probabilistic prior distribution for a spatial correlation matrix (a CGMM parameter), which enables more stable steering vector estimation in the presence of interfering speakers. In practice, the performance of the proposed online beamformer degrades with observations that contain only noise or/and interference because of the failure of the CGMM parameter estimation. The introduced spatial prior enables the target speaker's parameter to avoid overfitting to noise or/and interference. Experimental results show that the spatial prior reduces the WER from 38.4% to 29.2% in a conversation recognition task compared with the CGMM-based approach without the prior, and outperforms a conventional online speech enhancement approach. Takuya Higuchi, Nobutaka Ito, Shoko Araki, Takuya Yoshioka, Marc Delcroix, Tomohiro Nakatani |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2016 | Spatial correlation model based observation vector clustering and MVDR beamforming for meeting recognitionabstractThis paper addresses a minimum variance distortionless response (MVDR) beamforming based speech enhancement approach for meeting speech recognition. In a meeting situation, speaker overlaps and noise signals are not negligible. To handle these issues, we employ MVDR beamforming, where accurate estimation of the steering vector is paramount. We recently found that steering vector estimation by clustering the time-frequency components of microphone observation vectors performs well as regards real-world noise reduction. The clustering is performed by taking a cue from the spatial correlation matrix of each speaker, which is realized by modeling the time-frequency components of the observation vectors with a complex Gaussian mixture model (CGMM). Experimental results with real recordings show that the proposed MVDR scheme outperforms conventional null-beamformer based speech enhancement in a meeting situation. Shoko Araki, Masahiro Okada, Takuya Higuchi, Atsunori Ogawa, Tomohiro Nakatani |
ICASSP | 5 |
| 2016 | Context adaptive deep neural networks for fast acoustic model adaptation in noisy conditionsabstractDeep neural network (DNN) based acoustic models have greatly improved the performance of automatic speech recognition (ASR) for various tasks. Further performance improvements have been reported when making DNNs aware of the acoustic context (e.g. speaker or environment) for example by adding auxiliary features to the input, such as noise estimates or speaker i-vectors. We have recently proposed a context adaptive DNN (CA-DNN), which is another approach to exploit the acoustic context information within a DNN. A CA-DNN is a DNN that has one or several factorized layers, i.e. layers that use a different set of parameters to process each acoustic context class. The output of a factorized layer is obtained by the weighted sum over the contribution of the different context classes, given weights over the context classes. In our previous work, the class weights were computed independently of the recognizer. In this paper, we extend our previous work by introducing the joint training of the CA-DNN parameters and the class weights computation. Consequently, the class weights and the associated class definitions can be optimized for ASR. We report experimental results on the AURORA4 noisy speech recognition task showing the potential of our approach for fast unsupervised adaptation. Marc Delcroix, Keisuke Kinoshita, Chengzhu Yu, Atsunori Ogawa, Takuya Yoshioka, Tomohiro Nakatani |
ICASSP | 6 |
| 2016 | Multi-pass feature enhancement based on generative-discriminative hybrid approach for noise robust speech recognitionabstractThis paper presents multi-pass feature enhancement technique that consists of three processing passes. In the proposed method, the first pass was described in our previous work, and consists of model-based feature enhancement realized by employing a generative-discriminative hybrid approach with Gaussian mixture models and deep neural networks (DNNs). As an extension of the previous work, the second pass of the proposed method utilizes DNNs retrained with iterative realignment and auxiliary features obtained from intermediate parameters of the first processing pass. In the third pass, we apply unsupervised DNN adaptation and system combination to the results of the second pass. Therefore, the proposed multi-pass technique realizes stepwise improvements in feature enhancement. For CHiME3 task evaluations, the proposed method provided noticeable improvements in noisy speech recognition accuracy compared with results obtained using the previous one-pass feature enhancement technique. Masakiyo Fujimoto, Tomohiro Nakatani |
ICASSP | 2 |
| 2016 | Robust MVDR beamforming using time-frequency masks for online/offline ASR in noiseabstractThis paper considers acoustic beamforming for noise robust automatic speech recognition (ASR). A beamformer attenuates background noise by enhancing sound components coming from a direction specified by a steering vector. Hence, accurate steering vector estimation is paramount for successful noise reduction. Recently, a beamforming approach was proposed that employs time-frequency masks. In the speech recognition system we submitted to the CHiME-3 Challenge, we employed a new form of this approach that uses a speech spectral model based on a complex Gaussian mixture model (CGMM) to estimate the time-frequency masks and the steering vector without providing technical details. This paper elaborates on this technique and examines its effectiveness for ASR. Experimental results show that the CGMM-based approach outperforms a recently proposed mask estimator based on a Watson mixture model. In addition, the CGMM-based approach is extended to an online speech enhancement scenario, which allows this technique to be used in an online recognition setup. This online version reduces the CHiME-3 evaluation error rate from 15.60% to 8.47%, which is a comparable improvement to that obtained by batch processing. Takuya Higuchi, Nobutaka Ito, Takuya Yoshioka, Tomohiro Nakatani |
ICASSP | 4 |
| 2016 | Modeling audio directional statistics using a complex bingham mixture model for blind source extraction from diffuse noiseabstractMask estimation is a central task in blind signal processing including source separation, denoising, and multi-source localization. In this paper, we define a complex Bingham mixture model (cBMM), and propose it as a model of directional statistics for mask estimation. The complex Bingham distribution can represent not only rotationally symmetric but also rotationally asymmetric distributions. Therefore, it can precisely model stochastic variation of the directional statistics due to reverberation, noise, source movement, etc., which is not necessarily rotationally symmetric. In an experimental evaluation, the proposed cBMM outperformed a conventional complex Watson mixture model (cWMM) in terms of blind source extraction from diffuse noise, reducing the word error rate by 0.91% absolute on CHiME-3 challenge data. Nobutaka Ito, Shoko Araki, Tomohiro Nakatani |
ICASSP | 3 |
| 2016 | Real-time integration of statistical model-based speech enhancement with unsupervised noise PSD estimation using microphone arrayabstractWe propose a technique of multi-channel speech enhancement based on integration of beamforming and statistical model-based speech enhancement to clearly extract the target speech, even in very noisy environments. Conventional microphone array-based techniques estimate speech and noise power spectral densities (PSDs) from the spatial cues of the sound sources; however, their estimation errors dramatically increase when there are many noise sources. We integrated clean speech models trained in advance and the noise PSDs estimated in beamspace to compose observation models and designed a precise Wiener filter. Experiments under adverse noise conditions showed that the proposed technique significantly improved the signal-to-noise ratios (SNRs) compared with the conventional microphone array processing technique. Tomoko Kawase, Kenta Niwa, Masakiyo Fujimoto, Noriyoshi Kamado, Kazunori Kobayashi, Shoko Araki, Tomohiro Nakatani |
ICASSP | 7 |
| 2016 | A generative-discriminative hybrid approach to multi-channel noise reduction for robust automatic speech recognitionabstractIn the recent years, discriminative models have become a very attractive utility and gained a lot of attention in the speech research community, encompassing both front and back-end methods, thanks to their prominent discriminative power and the availability of improved training strategies. When it comes to the recognition of speech that is distorted by highly non-stationary environmental noise, robust front and backend methods are required in order to achieve a satisfactorily high speech recognition performance. Furthermore, when dealing with severe noise conditions, multi-channel front-end methods can be advantageous for suppressing environmental background noise, as compared to single-channel methods. In this work, we improve an existing multi-channel noise reduction approach, referred to as DOminance-based Loca-tional and Power-spectral cHaracteristics INtegration (DOLPHIN), by using a generative-discriminative hybrid model, that makes use of spatial and spectral features. We show that the proposed method outperforms the existing DOLPHIN approach, which is solely based on generative models, in terms of the word error rate reduction achieved on the CHiME-3 challenge data. Hendrik Meutzner, Shoko Araki, Masakiyo Fujimoto, Tomohiro Nakatani |
ICASSP | 4 |
| 2016 | Noise robust speech recognition using recent developments in neural networks for computer visionabstractConvolutional Neural Networks (CNNs) are superior to fully connected neural networks in various speech recognition tasks and the advantage is pronounced in noisy environments. In recent years, many techniques have been proposed in the computer vision community to improve CNN's classification performance. This paper considers two approaches recently developed for image classification and examines their impacts on noisy speech recognition performance. The first approach is to increase the depth of convolution layers. Different approaches to deepening the CNNs are compared. In particular, the usefulness of learning dynamic features with small convolution layers that perform convolution in time is shown along with a modulation frequency analysis of the learned convolution filters. The second approach is to use trainable activation functions. Specifically, the use of a Parametric Rectified Linear Unit (PReLU) is investigated. Experimental results show that both approaches yield significant improvements in performance. Combining the two approaches further reduces recognition errors, producing a word error rate of 11.1% in the Aurora4 task, the best published result for this corpus, with a standard one-pass bi-gram decoding set-up. Takuya Yoshioka, Katsunori Ohnishi, Fuming Fang, Tomohiro Nakatani |
ICASSP | 4 |
| 2016 | Context Adaptive Neural Network for Rapid Adaptation of Deep CNN Based Acoustic Models
Marc Delcroix, Keisuke Kinoshita, Atsunori Ogawa, Takuya Yoshioka, Dung T. Tran, Tomohiro Nakatani |
INTERSPEECH | 6 |
| 2016 | Optimization of Speech Enhancement Front-End with Speech Recognition-Level Criterion
Takuya Higuchi, Takuya Yoshioka, Tomohiro Nakatani |
INTERSPEECH | 3 |
| 2016 | Robust Example Search Using Bottleneck Features for Example-Based Speech Enhancement
Atsunori Ogawa, Shogo Seki, Keisuke Kinoshita, Marc Delcroix, Takuya Yoshioka, Tomohiro Nakatani, Kazuya Takeda |
INTERSPEECH | 6 |
| 2016 | Factorized Linear Input Network for Acoustic Model Adaptation in Noisy Conditions
Dung T. Tran, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2016 | Speech Intelligibility Prediction Based on the Envelope Power Spectrum Model with the Dynamic Compressive Gammachirp Auditory Filterbank
Katsuhiko Yamamoto, Toshio Irino, Toshie Matsui, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani |
INTERSPEECH | 6 |
| 2016 | Differenced maximum mutual information criterion for robust unsupervised acoustic model adaptation
Marc Delcroix, Atsunori Ogawa, Seong-Jun Hahm, Tomohiro Nakatani, Atsushi Nakamura |
Comput. Speech Lang. | 4 |
| 2015 | The NTT CHiME-3 system: Advances in speech enhancement and recognition for mobile multi-microphone devicesabstractCHiME-3 is a research community challenge organised in 2015 to evaluate speech recognition systems for mobile multi-microphone devices used in noisy daily environments. This paper describes NTT's CHiME-3 system, which integrates advanced speech enhancement and recognition techniques. Newly developed techniques include the use of spectral masks for acoustic beam-steering vector estimation and acoustic modelling with deep convolutional neural networks based on the "network in network" concept. In addition to these improvements, our system has several key differences from the official baseline system. The differences include multi-microphone training, dereverberation, and cross adaptation of neural networks with different architectures. The impacts that these techniques have on recognition performance are investigated. By combining these advanced techniques, our system achieves a 3.45% development error rate and a 5.83% evaluation error rate. Three simpler systems are also developed to perform evaluations with constrained set-ups. Takuya Yoshioka, Nobutaka Ito, Marc Delcroix, Atsunori Ogawa, Keisuke Kinoshita, Masakiyo Fujimoto, Chengzhu Yu, Wojciech J. Fabian, Miquel Espi, Takuya Higuchi, Shoko Araki, Tomohiro Nakatani |
ASRU | 12 |
| 2015 | Exploring multi-channel features for denoising-autoencoder-based speech enhancementabstractThis paper investigates a multi-channel denoising autoencoder (DAE)-based speech enhancement approach. In recent years, deep neural network (DNN)-based monaural speech enhancement and robust automatic speech recognition (ASR) approaches have attracted much attention due to their high performance. Although multi-channel speech enhancement usually outperforms single channel approaches, there has been little research on the use of multi-channel processing in the context of DAE. In this paper, we explore the use of several multi-channel features as DAE input to confirm whether multi-channel information can improve performance. Experimental results show that certain multi-channel features outperform both a monaural DAE and a conventional time-frequency-mask-based speech enhancement method. Shoko Araki, Tomoki Hayashi, Marc Delcroix, Masakiyo Fujimoto, Kazuya Takeda, Tomohiro Nakatani |
ICASSP | 6 |
| 2015 | Context adaptive deep neural networks for fast acoustic model adaptationabstractDeep neural networks (DNNs) are widely used for acoustic modeling in automatic speech recognition (ASR), since they greatly outperform legacy Gaussian mixture model-based systems. However, the levels of performance achieved by current DNN-based systems remain far too low in many tasks, e.g. when the training and testing acoustic contexts differ due to ambient noise, reverberation or speaker variability. Consequently, research on DNN adaptation has recently attracted much interest. In this paper, we present a novel approach for the fast adaptation of a DNN-based acoustic model to the acoustic context. We introduce a context adaptive DNN with one or several layers depending on external factors that represent the acoustic conditions. This is realized by introducing a factorized layer that uses a different set of parameters to process each class of factors. The output of the factorized layer is then obtained by weighted averaging over the contribution of the different factor classes, given posteriors over the factor classes. This paper introduces the concept of context adaptive DNN and describes preliminary experiments with the TIMIT phoneme recognition task showing consistent improvement with the proposed approach. Marc Delcroix, Keisuke Kinoshita, Takaaki Hori, Tomohiro Nakatani |
ICASSP | 4 |
| 2015 | Feature enhancement based on generative-discriminative hybrid approach with gmms and DNNS for noise robust speech recognitionabstractThis paper presents a technique that combines generative and discriminative approaches with Gaussian mixture models (GMMs) and deep neural networks (DNNs) for model-based feature enhancement. Typical model-based feature enhancement employs a generative model approach. The enhanced features are obtained by using the weighted sum of linear transformations given by each Gaussian component contained in GMMs and corresponding posterior probabilities. The computation of posterior probabilities is a crucial factor for this kind of feature enhancement, and can also be formulated as the class discrimination problem of observed noisy features. The prominent discriminability of DNNs is a well-known solution to this discrimination problem. Therefore, we propose the use of DNNs for computing posterior probabilities. The proposed method incorporates the benefit of the discriminative approach into the generative approach. For AURORA2 task evaluations, the proposed method provided noticeable improvements compared with results obtained using the conventional generative model approach. Masakiyo Fujimoto, Tomohiro Nakatani |
ICASSP | 2 |
| 2015 | Modeling inter-node acoustic dependencies with Restricted Boltzmann Machine for distributed microphone array based BSSabstractAn accurate estimation of a source activity information is essential for many speech enhancement algorithms including blind source separation (BSS). In this paper, we propose a novel BSS method that accurately models and estimates the source activity in distributed microphone array (DMA) scenarios. In DMA scenarios, microphones (or in more general term, microphone-nodes) are often spatially distributed to a great degree. If there are multiple source signals in such an environment, the level of each source signal at each microphone-node varies significantly, thus the source activities observable at one microphone-node should be significantly different from those of other nodes. Therefore, it is essential to assume node-specific source activities in DMA scenarios. In the proposed method, the estimation of the node-specific source activities are done by integrating node-wise clustering-based BSS processings based on inter-node acoustic dependencies, i.e., a co-occurrence of the source activities among nodes. To model the co-occurrence relationship, we employ Restricted Boltzmann Machine (RBM) in a similar manner as it is used for collaborative filtering. This paper introduces a probabilistic formulation of the proposed method, and experimentally demonstrates how essential it is to estimate the node-specific source activities for distributed microphone array based BSS. Keisuke Kinoshita, Tomohiro Nakatani |
ICASSP | 2 |
| 2015 | Far-field speech recognition using CNN-DNN-HMM with convolution in timeabstractRecent studies in speech recognition have shown that the performance of convolutional neural networks (CNNs) is superior to that of fully connected deep neural networks (DNNs). In this paper, we explore the use of CNNs in far-field speech recognition for dealing with reverberation, which blurs spectral energies along the time axis. Unlike most previous CNN applications to speech recognition, we consider convolution in time to examine whether it provides an improved reverberation modelling capability. Experimental results show that a CNN coupled with a fully connected DNN can model short time correlations in feature vectors with fewer parameters than a DNN and thus generalise better to unseen test environments. Combining this approach with signal-space dereverberation, which copes with long-term correlations, is shown to result in further improvement, where the gains from both approaches are almost additive. An initial investigation of the use of restricted convolution forms is also undertaken. Takuya Yoshioka, Shigeki Karita, Tomohiro Nakatani |
ICASSP | 3 |
| 2015 | Feature extraction strategies in deep learning based acoustic event detection
Miquel Espi, Masakiyo Fujimoto, Keisuke Kinoshita, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2015 | Text-informed speech enhancement with deep neural networksabstractA speech signal captured by a distant microphone is generally contaminated by background noise, which severely degrades the audible quality and intelligibility of the observed signal. To resolve this issue, speech enhancement has been intensively studied. In this paper, we consider a text-informed speech enhancement, where the enhancement process is guided by the corresponding text information, i.e., a correct transcription of the target utterance. The proposed deep neural network (DNN)based framework is motivated by the recent success in the textto-speech (TTS) research in employing DNN as well as high audible-quality output signal of the corpus-based speech enhancement which borrows knowledge from the TTS research field. Taking advantage of the nature of DNN that allows us to utilize disparate features in an inference stage, the proposed method infers the clean speech features by jointly using the observed signal and widely-used TTS features derived from the corresponding text. In this paper, we first introduce the background and the details of the proposed method. Then, we show how the text information can be naturally integrated into speech enhancement by utilizing DNN and improve the enhancement performance. Index Terms: speech enhancement, text-to-speech, deep neural network Keisuke Kinoshita, Marc Delcroix, Atsunori Ogawa, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2015 | Robust i-vector extraction for neural network adaptation in noisy environment
Chengzhu Yu, Atsunori Ogawa, Marc Delcroix, Takuya Yoshioka, Tomohiro Nakatani, John H. L. Hansen |
INTERSPEECH | 5 |
| 2014 | Unsupervised non-parametric Bayesian modeling of non-stationary noise for model-based noise suppressionabstractThe accurate modeling of non-stationary noise plays an important role in model-based noise suppression for noise robust speech recognition. We have already proposed methods for unsupervised noise modeling with a Gaussian mixture model or a hidden Markov model by using a minimum mean squared error estimate of the noise. However, our previous work fixed the structure of the noise model empirically without any consideration of noise characteristics; thus, optimization of the noise model structure is required if we are to obtain further improvements. Although the Bayesian information criterion (BIC) has been widely used as a conventional approach to model structure estimation, it is not always the optimal criterion. Therefore, this paper presents a way of modeling non-stationary noise with a non-parametric Bayesian approach that estimates the model structure depending on the characteristics of given observations. The proposed method provided improved results for the evaluations of two different speech recognition tasks compared with results obtained using the conventional BIC-based approach. Masakiyo Fujimoto, Yotaro Kubo, Tomohiro Nakatani |
ICASSP | 3 |
| 2014 | Probabilistic integration of diffuse noise suppression and dereverberationabstractThis paper deals with joint suppression of diffuse noise and reverberation, to enhance perceived speech quality and speech recognition performance. Although diffuse noise and reverberation are both omnipresent in the real world, conventional methods have modeled only one while neglecting the other. In contrast, we propose a novel joint suppression method that employs a unified probabilistic model of observed signals affected by both diffuse noise and reverberation. Through likelihood maximization, this unified model enables proper parameter estimation that takes into account both diffuse noise and reverberation. As a byproduct, we also propose a novel method for diffuse noise suppression. Experimental results demonstrate the effectiveness of the proposed joint suppression method in terms of dereverberation and denoising. Nobutaka Ito, Shoko Araki, Tomohiro Nakatani |
ICASSP | 3 |
| 2014 | Fast segment search for corpus-based speech enhancement based on speech recognition technologyabstractCorpus-based speech enhancement has received increasing attention recently since it shows high enhancement performance in highly non-stationary noisy environments by precisely modeling the long-term temporal dynamics of speech. However, it has a disadvantage in that the cost is very high for searching the longest matching clean speech segments from a multi-condition parallel speech corpus. This paper proposes a fast segment search method for corpus-based speech enhancement. It is mainly based on two techniques derived from speech recognition technology. The first is an A* search like segment evaluation function for accurately finding the longest matching segments. The second is a tree and linear connected search space for efficiently sharing the segment likelihood calculations. In the experiments for non-stationary noisy observations using the 26 multi-condition TIMIT parallel speech corpus, the proposed search method found the segments almost in real-time without degrading the quality of the enhanced speech. Our method was about 7 to 13 times faster than the conventional segment search method. Atsunori Ogawa, Keisuke Kinoshita, Takaaki Hori, Tomohiro Nakatani, Atsushi Nakamura |
ICASSP | 4 |
| 2014 | Location Feature Integration for Clustering-Based Speech Separation in Distributed Microphone ArraysabstractIn distributed microphone arrays (DMAs) the source location information can be defined at the intra and inter-node levels. Indeed, while the first type of information results from the diversity of acoustic channels recorded by microphones embedded in the same node, the second is attributed to the differences between the acoustic channels observed by spatially distributed nodes. Both cues are very useful in DMA processing, and the aim of this paper is to utilize both of them to cluster and separate multiple competing speech signals. To capture the intra-node information, we employ the normalized recording vector, while at the inter-node level, we consider different features including the energy level differences with and without the phase differences between nodes. We model the intra-node information using the Watson mixture model (WMM), and propose using the Gamma mixture model (GaMM), Dirichlet mixture model (DMM), and WMM to model different inter-node location features. Furthermore, we propose several integrations of the intra-node and inter-node feature contributions to cluster speech recordings using the expectation maximization algorithm. Finally, simulation results are provided to demonstrate the performance of all ensuing methods. Mehrez Souden, Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2013 | Unsupervised discriminative adaptation using differenced maximum mutual information based linear regressionabstractThis paper proposes a new approach for unsupervised model adaptation using a discriminative criterion. Discriminative criteria for acoustic model training have been widely used and have provided significantly improved performance compared with models trained using maximum likelihood. However, discriminative criteria are sensitive to errors in reference transcriptions, which limits their applicability to unsupervised adaptation. In this paper, we apply the recently proposed differenced maximum mutual information (dMMI) criteria to unsupervised linear regression based adaptation because dMMI has an intrinsic mechanism that mitigates the influence of transcription errors. We report unsupervised adaptation results for a large vocabulary continuous speech recognition task showing a significant improvement over maximum likelihood based linear regression. Marc Delcroix, Atsunori Ogawa, Seong-Jun Hahm, Tomohiro Nakatani, Atsushi Nakamura |
ICASSP | 4 |
| 2013 | Permutation-free convolutive blind source separation via full-band clustering based on frequency-independent source presence priorsabstractWe propose permutation-free frequency-domain blind source separation (BSS) via full-band clustering of the time-frequency (T-F) components based on time-varying signal presence priors. Frequency-domain methods of BSS usually process each frequency bin separately, and therefore necessitate the subsequent alignment of the permutation ambiguity that arises between frequency bins. In contrast, the proposed method simultaneously processes all frequency bins by using a mixture model with time-varying, frequency-independent mixture weights. We propose to assume non-sparse priors on the mixture weights to prevent the degradation of source separation performance by the time-varying mixture weights. We propose a customized expectation-maximization (EM) algorithm for the maximum a posteriori (MAP) estimation of the model parameters, to which we introduce a novel technique to avoid convergence to local maxima. For audio source separation, we use the normalized observation vector as the feature vector, and theWatson mixture model (WMM) as the mixture model. Evaluations confirm that the proposed permutation-free BSS results in source separation performance comparable to the state-of-the-art clustering-based BSS composed of bin-wise clustering and permutation alignment. Nobutaka Ito, Shoko Araki, Tomohiro Nakatani |
ICASSP | 3 |
| 2013 | Coupling beamforming with spatial and spectral feature based spectral enhancement and its application to meeting recognitionabstractThis paper discusses microphone array based interference reduction approaches for robust automatic speech recognition. A model based multichannel spectral enhancement approach has recently been proposed for effectively reducing interference by exploiting both the spatial and spectral features of the signals. With the goal of further improving the effectiveness of this approach, we propose a new framework that combines this approach with a microphone-array based beamforming approach. Because the two approaches can work in a complementary manner in the proposed framework, they can greatly improve the interference reduction performance. We apply the proposed framework to the recognition of actual meetings, and show that it is superior to the use of beamforming or spectral enhancement alone in terms of the word error rates. Tomohiro Nakatani, Mehrez Souden, Shoko Araki, Takuya Yoshioka, Takaaki Hori, Atsunori Ogawa |
ICASSP | 1 |
| 2013 | An integration of source location cues for speech clustering in distributed microphone arraysabstractWe propose a new approach for clustering competing speech sources using distributed microphone arrays. In this approach, we first define two feature vectors where the first captures the intra-node location information while the second captures the level difference of speech energy recorded at different nodes. Then, we introduce Watson and Dirichlet mixture models to model the first and second features, respectively. We integrate both types of information in an expectation maximization algorithm to cluster the simultaneous speech sources. The performance of the proposed approach is superior to best node selection and comparable to centralized processing in terms of conventional blind source separation metrics. Mehrez Souden, Keisuke Kinoshita, Tomohiro Nakatani |
ICASSP | 3 |
| 2013 | Noise model transfer using affine transformation with application to large vocabulary reverberant speech recognitionabstractThis paper considers using the feature enhancement approach for automatic recognition of speech corrupted by severely nonstationary noise, caused for example by interfering talkers and inter-frame distortion induced by reverberation. In particular, we focus on the issue of feature-domain noise model estimation and investigate a recently proposed approach, called noise model transfer (NMT), for estimating the rapidly changing noise model parameter values. Based on the fact that noise spectral changes can be detected more easily in the power spectrum domain than in the feature domain, NMT estimates the noise model parameter values for each time frame by using both observed feature vectors and noise power spectral estimates, on the assumption that a separate noise power spectrum estimator is available. This is achieved by finding the best transformation that maps the power spectra onto the noise model parameter space in the maximum likelihood sense. Whereas the transformation was previously modeled using a bias vector, this paper employs a more flexible affine transformation model. The results of 20,000-word reverberant speech recognition experiments show the advantage of the affine transformation model. Takuya Yoshioka, Tomohiro Nakatani |
ICASSP | 2 |
| 2013 | Is speech enhancement pre-processing still relevant when using deep neural networks for acoustic modeling?
Marc Delcroix, Yotaro Kubo, Tomohiro Nakatani, Atsushi Nakamura |
INTERSPEECH | 3 |
| 2013 | Model-based noise suppression using unsupervised estimation of hidden Markov model for non-stationary noiseabstractAlthough typical model-based noise suppression including the vector Taylor series-based approach employs a single Gaussian distribution for the noise model, it is insufficient for nonstationary noises which have a complex structured distribution. As a solution to this problem, we have already proposed a method for estimating a Gaussian mixture model (GMM)-based noise model by using a minimum mean squared error (MMSE) estimate of the noise. However, the state transition process of the non-stationary noise is not modeled in the noise GMM. In this paper, we propose a way of modeling the noise with a hidden Markov model (HMM) as an extension of our previous method. The proposed method proves that the HMM-based noise model outperforms a GMM-based noise model composed of the same number of Gaussian components. In addition, we discuss the appropriate topology for the noise HMM, i.e., a leftto-right HMM and an ergodic HMM. Index Terms: noise suppression, noise modeling, model topology, MMSE estimation Masakiyo Fujimoto, Tomohiro Nakatani |
INTERSPEECH | 2 |
| 2013 | Blind source separation using spatially distributed microphones based on microphone-location dependent source activities
Keisuke Kinoshita, Mehrez Souden, Tomohiro Nakatani |
INTERSPEECH | 3 |
| 2013 | Conditional emission densities for combining speech enhancement and recognition systems
Armin Sehr, Takuya Yoshioka, Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Roland Maas, Walter Kellermann |
INTERSPEECH | 5 |
| 2013 | On the robustness of distributed EM based BSS in asynchronous distributed microphone array scenarios
Yasufumi Uezu, Keisuke Kinoshita, Mehrez Souden, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2013 | Speech recognition in living rooms: Integrated speech enhancement and recognition system based on spatial, spectral and temporal modeling of sounds
Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki, Atsunori Ogawa, Takaaki Hori, Shinji Watanabe 0001, Masakiyo Fujimoto, Takuya Yoshioka, Takanobu Oba, Yotaro Kubo, Mehrez Souden, Seong-Jun Hahm, Atsushi Nakamura |
Comput. Speech Lang. | 3 |
| 2013 | Cluster-based dynamic variance adaptation for interconnecting speech enhancement pre-processor and speech recognizer
Marc Delcroix, Shinji Watanabe 0001, Tomohiro Nakatani, Atsushi Nakamura |
Comput. Speech Lang. | 3 |
| 2013 | Dominance Based Integration of Spatial and Spectral Features for Speech EnhancementabstractThis paper proposes a versatile technique for integrating two conventional speech enhancement approaches, a spatial clustering approach (SCA) and a factorial model approach (FMA), which are based on two different features of signals, namely spatial and spectral features, respectively. When used separately the conventional approaches simply identify time frequency (TF) bins that are dominated by interference for speech enhancement. Integration of the two approaches makes identification more reliable, and allows us to estimate speech spectra more accurately even in highly nonstationary interference environments. This paper also proposes extensions of the FMA for further elaboration of the proposed technique, including one that uses spectral models based on mel-frequency cepstral coefficients and another to cope with mismatches, such as channel mismatches, between captured signals and the spectral models. Experiments using simulated and real recordings show that the proposed technique can effectively improve audible speech quality and the automatic speech recognition score. Tomohiro Nakatani, Shoko Araki, Takuya Yoshioka, Marc Delcroix, Masakiyo Fujimoto |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2013 | A Multichannel MMSE-Based Framework for Speech Source Separation and Noise ReductionabstractWe propose a new framework for joint multichannel speech source separation and acoustic noise reduction. In this framework, we start by formulating the minimum-mean-square error (MMSE)-based solution in the context of multiple simultaneous speakers and background noise, and outline the importance of the estimation of the activities of the speakers. The latter is accurately achieved by introducing a latent variable that takes N+1 possible discrete states for a mixture of N speech signals plus additive noise. Each state characterizes the dominance of one of the N+1 signals. We determine the posterior probability of this latent variable, and show how it plays a twofold role in the MMSE-based speech enhancement. First, it allows the extraction of the second order statistics of the noise and each of the speech signals from the noisy data. These statistics are needed to formulate the multichannel Wiener-based filters (including the minimum variance distortionless response). Second, it weighs the outputs of these linear filters to shape the spectral contents of the signals' estimates following the associated target speakers' activities. We use the spatial and spectral cues contained in the multichannel recordings of the sound mixtures to compute the posterior probability of this latent variable. The spatial cue is acquired by using the normalized observation vector whose distribution is well approximated by a Gaussian-mixture-like model, while the spectral cue can be captured by using a pre-trained Gaussian mixture model for the log-spectra of speech. The parameters of the investigated models and the speakers' activities (posterior probabilities of the different states of the latent variable) are estimated via expectation maximization. Experimental results including comparisons with the well-known independent component analysis and masking are provided to demonstrate the efficiency of the proposed framework. Mehrez Souden, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani, Hiroshi Sawada |
IEEE Trans. Speech Audio Process. | 4 |
| 2013 | Noise Model Transfer: Novel Approach to Robustness Against Nonstationary NoiseabstractThis paper proposes an approach, called noise model transfer (NMT), for estimating the rapidly changing parameter values of a feature-domain noise model, which can be used to enhance feature vectors corrupted by highly nonstationary noise. Unlike conventional methods, the proposed approach can exploit both observed feature vectors, representing spectral envelopes, and other signal properties that are usually discarded during feature extraction but that are useful for separating nonstationary noise from speech. Specifically, we assume the availability of a noise power spectrum estimator that can capture rapid changes in noise characteristics by leveraging such signal properties. NMT determines the optimal transformation from the estimated noise power spectra into the feature-domain noise model parameter values in the sense of maximum likelihood. NMT is successfully applied to meeting speech recognition, where the main noise sources are competing talkers; and reverberant speech recognition, where the late reverberation is regarded as highly nonstationary additive noise. Takuya Yoshioka, Tomohiro Nakatani |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Sparse vector factorization for underdetermined BSS using wrapped-phase GMM and source log-spectral priorabstractWe propose a sparse vector factorization (SVF) approach for blind source separation, which inherently avoids the permutation problem. The SVF assumes the sparseness of sources, and defines a sparse vector (SV) that consists of the locational and spectral features of each source at all the frequencies. Then, by assuming that the locational and spectral SVs are generated by frequency-independent parameters, the method executes the SVF. Our locational feature is the phase difference (PD) between two microphone observations, and we model it with a frequency-independent time-difference of arrival (TDOA) parameter. Moreover, we employ the wrapped-phase GMM in order to take the spatial aliasing problem into account. On the other hand, the spectral feature is the log spectrum, and we provide a prior for a spectral parameter. The SVF is formulated with a maximum a posteriori (MAP) estimation framework, where the locational and spectral parameters are inferred by the EM algorithm. Experimental results show that our proposed method can separate signals successfully even for an underdetermined case. Shoko Araki, Tomohiro Nakatani |
ICASSP | 2 |
| 2012 | Discriminative feature transforms using differenced maximum mutual informationabstractRecently feature compensation techniques that train feature transforms using a discriminative criterion have attracted much interest in the speech recognition community. Typically, the acoustic feature space is modeled by a Gaussian mixture model (GMM), and a feature transform is assigned to each Gaussian of the GMM. Feature compensation is then performed by transforming features using the transformation associated with each Gaussian, then summing up the transformed features weighted by the posterior probability of each Gaussian. Several discriminative criteria have been investigated for estimating the feature transformation parameters including maximum mutual information (MMI) and minimum phone error (MPE). Recently, the differenced MMI (dMMI) criterion that generalizes MMI andMPE, has been shown to provide competitive performance for acoustic model training. In this paper, we investigate the use of the dMMI criterion for discriminative feature transforms and demonstrate in a noisy speech recognition experiment that dMMI achieves recognition performance superior to that of MMI or MPE. Marc Delcroix, Atsunori Ogawa, Shinji Watanabe 0001, Tomohiro Nakatani, Atsushi Nakamura |
ICASSP | 4 |
| 2012 | Noise suppression with unsupervised joint speaker adaptation and noise mixture model estimationabstractThe estimation of an accurate noise model is a crucial problem for model-based noise suppression including a vector Taylor series (VTS)-based approach. The variation of the speaker characteristics is also a crucial factor as regards the model-based noise suppression. As a result, a speaker adaptation technique plays an important role in the model-based noise suppression. To deal with former problem, we have already proposed an unsupervised estimation method for a noise mixture model. Therefore, this paper proposes a joint processing method that simultaneously achieves speaker adaptation and noise mixture model estimation. This joint processing is realized by using minimum mean squared error (MMSE) estimates of clean speech and noise. Although VTS-based approach involves nonlinear transformation, the MMSE estimates make it possible to flexibly estimate accurate parameters for the joint processing without the influences of non-linear VTS transformation. In the evaluation, the proposed method provided an improvement compared with results obtained using only noise mixture model estimation. Masakiyo Fujimoto, Shinji Watanabe 0001, Tomohiro Nakatani |
ICASSP | 3 |
| 2012 | Introduction of speech log-spectral priors into dereverberation based on Itakura-Saito distance minimizationabstractIt has recently been shown that a multi-channel linear prediction can effectively achieve blind speech dereverberation based on maximum-likelihood (ML) estimation. This approach can estimate and cancel unknown reverberation processes from only a few seconds of observation. However, one problem with this approach is that speech distortion may increase if we iterate the dereverberation more than once based on Itakura-Saito (IS) distance minimization to further reduce the reverberation. To overcome this problem, we introduce speech log-spectral priors into this approach, and reformulate it based on maximum a posteriori (MAP) estimation. Two types of priors are introduced, a Gaussian mixture model (GMM) of speech log spectra, and a GMM of speech mel-frequency cepstral coefficients. In the formulation, we also propose a new versatile technique to integrate such log-spectral priors with the IS distance minimization in a computationally efficient manner. Preliminary experiments show the effectiveness of the proposed approach. Yasuaki Iwata, Tomohiro Nakatani |
ICASSP | 2 |
| 2012 | New analytical update rule for TDOA inference for underdetermined BSS in noisy environmentsabstractIn this paper, we propose a new technique for sparseness-based underdetermined BSS that is based on the clustering of the frequency-dependent time difference of arrival (TDOA) information and that can cope with diffused noise environments. Such a method with an EM algorithm has already been proposed, however, it required a time-consuming exhaust search for TDOA inference. To remove the need for such an exhaust search, we propose a new technique by focusing on a stereo case. We derive an update rule for analytical TDOA estimation. This update rule eliminates the need for the exhaustive TDOA search, and therefore reduces the computational load. We show experimental results for separation performance and calculation time in comparison with those obtained with the conventional approach. Our reported results validate our proposed method, that is, our proposed method achieves high performance without a high computational cost. Takuro Maruyama, Shoko Araki, Tomohiro Nakatani, Shigeki Miyabe, Takeshi Yamada, Shoji Makino, Atsushi Nakamura |
ICASSP | 3 |
| 2012 | LogMax observation model with MFCC-based spectral prior for reduction of highly nonstationary ambient noiseabstractThis paper proposes a new single/multi-channel speech enhancement approach based on a LogMax observation model integrated with Gaussian mixture models of speech and noise mel-frequency cepstral coefficients (MFCC-GMM). It has been reported that the LogMax observation model has high potential for reducing highly nonstationary noise, for example, when it is combined with factorial hidden Markov models. In addition, it has recently been shown that a source location based speech enhancement approach can be easily incorporated into this model for more efficient and reliable estimation. However, the unique structure of the LogMax model has prevented us from using it with MFCC-GMMs, which is a fundamental limitation of this approach. Our proposal in this paper is aimed at overcoming this limitation. Experiments using the PASCAL CHiME separation and recognition challenge task show the superiority of the proposed approach as regards both speech quality and automatic speech recognition performance. Tomohiro Nakatani, Takuya Yoshioka, Shoko Araki, Marc Delcroix, Masakiyo Fujimoto |
ICASSP | 1 |
| 2012 | A multichannel MMSE-based framework for joint blind source separation and noise reductionabstractIn this paper, we propose a new framework to separate multiple speech signals and reduce the additive acoustic noise using multiple microphones. In this framework, we start by formulating the minimum-mean-square error (MMSE) criterion to retrieve each of the desired speech signals from the observed mixtures of sounds and outline the importance of multi-speaker activity detection. The latter is modeled by introducing a latent variable whose posterior probability is computed via expectation maximization (EM) combining both the spatial and spectral cues of the multichannel speech observations. We experimentally demonstrate that the resulting joint blind source separation (BSS) and noise reduction solution performs remarkably well in reverberant and noisy environments. Mehrez Souden, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani, Hiroshi Sawada |
ICASSP | 4 |
| 2012 | Time-varying residual noise feature model estimation for multi-microphone speech recognitionabstractThis paper proposes a method for compensating for the effect of noise remaining in a signal generated by a multi-microphone signal enhancer in the feature domain as a post-processing. The proposed method assumes that the multi-microphone signal enhancer generates estimates of both the target and original environmental noise signals. To obtain a time-varying residual noise feature model that responds to noise changes quickly and is consistent with a clean feature model, the proposed method leverages both the multiple signal estimates provided by the signal enhancer and the clean feature model. Specifically, the proposed method first roughly estimates residual noise features on a frame-by-frame basis by comparing the target and noise signal estimates. Then, these rough estimates are refined by using the clean feature model to yield a time-varying residual noise feature model. Experimental results show the effectiveness of the proposed method and its wide applicability. Takuya Yoshioka, Emmanuel Ternon, Tomohiro Nakatani |
ICASSP | 3 |
| 2012 | Example-based speech enhancement with joint utilization of spatial, spectral & temporal cues of speech and noise
Keisuke Kinoshita, Marc Delcroix, Mehrez Souden, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2012 | Frame-wise model re-estimation method based on Gaussian pruning with weight normalization for noise robust voice activity detection
Masakiyo Fujimoto, Shinji Watanabe 0001, Tomohiro Nakatani |
Speech Commun. | 3 |
| 2012 | Noise Power Spectral Density Tracking: A Maximum Likelihood PerspectiveabstractWe propose a new approach for online noise power spectral density (psd) tracking. In this approach, the prior and posterior probabilities of speech absence and also noise statistics are analytically retrieved from a maximum-likelihood-based criterion at every time-frequency slot. The recursive update rules of these three terms are performed in a unified manner and without relying on the conventional tracking of speech psd minima. A single parameter (a forgetting factor) is needed in this process. Comparisons with state of the art methods demonstrate the effectiveness of our proposal. Mehrez Souden, Marc Delcroix, Keisuke Kinoshita, Takuya Yoshioka, Tomohiro Nakatani |
IEEE Signal Process. Lett. | 5 |
| 2012 | Low-Latency Real-Time Meeting Recognition and Understanding Using Distant Microphones and Omni-Directional CameraabstractThis paper presents our real-time meeting analyzer for monitoring conversations in an ongoing group meeting. The goal of the system is to recognize automatically “who is speaking what” in an online manner for meeting assistance. Our system continuously captures the utterances and face poses of each speaker using a microphone array and an omni-directional camera positioned at the center of the meeting table. Through a series of advanced audio processing operations, an overlapping speech signal is enhanced and the components are separated into individual speaker's channels. Then the utterances are sequentially transcribed by our speech recognizer with low latency. In parallel with speech recognition, the activity of each participant (e.g., speaking, laughing, watching someone) and the circumstances of the meeting (e.g., topic, activeness, casualness) are detected and displayed on a browser together with the transcripts. In this paper, we describe our techniques and our attempt to achieve the low-latency monitoring of meetings, and we show our experimental results for real-time meeting transcription. Takaaki Hori, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto, Shinji Watanabe 0001, Takanobu Oba, Atsunori Ogawa, Kazuhiro Otsuka, Dan Mikami, Keisuke Kinoshita, Tomohiro Nakatani, Atsushi Nakamura, Junji Yamato |
IEEE Trans. Speech Audio Process. | 11 |
| 2012 | Probabilistic Speaker Diarization With Bag-of-Words Representations of Speaker Angle InformationabstractSpeaker diarization determines “who spoke when” from the recorded conversations of an unknown number of people. In general, we have no a priori information about the number, the locations, or even the characteristics of the speakers. Additionally, speakers' speech utterances vary dynamically because of turn-taking during the conversations. These conditions make the speaker-clustering task extremely difficult. The problem becomes even harder if online (incremental) processing is required. In this paper, we formulate the speaker-clustering problem as the clustering of the sequential audio features generated by an unknown number of latent mixture components (speakers). We employ a probabilistic model that assumes time-sensitive speaker mixtures at every time frame, which, surprisingly, suits the diarization scenario. We combine the time-varying probabilistic model with direction of arrival (DOA) information calculated from a microphone array in a bag-of-words (BoW)-style feature representation. The proposed system effectively estimates the number and locations of the speakers in an online manner based on the standard Bayes inference scheme. Experiments confirm that the proposed model can successfully infer the number and features of speakers and yield better or comparable speaker diarization results compared with conventional methods in several datasets. Katsuhiko Ishiguro, Takeshi Yamada, Shoko Araki, Tomohiro Nakatani, Hiroshi Sawada |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | Generalization of Multi-Channel Linear Prediction Methods for Blind MIMO Impulse Response ShorteningabstractThe performance of many microphone array processing techniques deteriorates in the presence of reverberation. To provide a widely applicable solution to this longstanding problem, this paper generalizes existing dereverberation methods using subband-domain multi-channel linear prediction filters so that the resultant generalized algorithm can blindly shorten a multiple-input multiple-output (MIMO) room impulse response between a set of unknown number of sources and a microphone array. Unlike existing dereverberation methods, the presented algorithm is developed without assuming specific acoustic conditions, and provides a firm theoretical underpinning for the applicability of the subband-domain multi-channel linear prediction methods. The generalization is achieved by using a new cost function for estimating the prediction filter and an efficient optimization algorithm. The proposed generalized algorithm makes it easier to understand the common background underlying different dereverberation methods and future technical development. Indeed, this paper also derives two alternative dereverberation methods from the proposed algorithm, which are advantageous in terms of computational complexity. Experimental results are reported, showing that the proposed generalized algorithm effectively achieves blind MIMO impulse response shortening especially in a mid-to-high frequency range. Takuya Yoshioka, Tomohiro Nakatani |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Hybrid approach for multichannel source separation combining time-frequency mask with multi-channel Wiener filterabstractThis paper discusses a hybrid approach for the multi-channel source separation, where both a time-frequency (t-f) mask and a multi-channel Wiener filter (WF) are utilized. T-f mask based approaches have been widely studied, because they can separate signals with a low calculation cost. However, the separated signals with a t-f mask usually contain a non-linear distortion. On the other hand, a new multi-channel WF framework employing a spatial covariance matrix model has recently been proposed. With the WF method, we can obtain separated signals of better quality than with a t-f mask, however, the method is computationally expensive because it requires many iterations for the optimization. In this paper, in order to take advantages of both approaches, first we explain the hybrid algorithm by introducing the t-f mask concept to the WF approach. Then we show that the hybrid approach achieves high performance without the iterative calculation for the WF. We also present the way for applying the hybrid method to the case where the number of sources is unavailable. Shoko Araki, Tomohiro Nakatani |
ICASSP | 2 |
| 2011 | Non-stationary noise estimation method based on bias-residual component decomposition for robust speech recognitionabstractThis paper addresses a noise suppression problem, namely the estimation of non-stationary noise sequences. In this problem, we assume that non-stationary noise can be decomposed into stationary and non-stationary components. These components are described respectively as the bias factor and the residual signal between the bias component and noise at each frame. This decomposition clarifies the role of each component, thus enabling us to apply a suitable parameter estimation technique to each component. In this paper, tile bias component is estimated by the EM algorithm with the entire observed signal sequence. On the other hand, the residual component is sequentially estimated by multiplying the extended Kalman filter with the EM algorithm. In the evaluation results, we confirmed that the proposed method improved speech recognition accuracy compared with the noise estimation methods without component decomposition. Masakiyo Fujimoto, Shinji Watanabe 0001, Tomohiro Nakatani |
ICASSP | 3 |
| 2011 | Joint unsupervised learning of hidden Markov source models and source location models for multichannel source separationabstractThis paper discusses a multichannel source separation approach that exploits the statistical characteristics of source location cues characterized by steering vector models (SM) and those of source log spectra characterized by hidden Markov models (spectral HMM). Recently, it was shown that the use of speaker independent spectral HMMs trained in advance substantially improves the quality of speech signals separated based on source location cues in a computationally efficient manner. However, with this approach, mismatches between the spectral HMMs and the observation may substantially degrade the separation quality, which limits the applicability of this approach. To overcome this problem, this paper proposes a method for learning the parameters of the spectral HMMs jointly with those of the SMs from the observed sound mixtures. Experimental results show that the proposed method works effectively for separation of convolutive sound mixtures. Tomohiro Nakatani, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto |
ICASSP | 1 |
| 2011 | Speech enhancement based on log spectral envelope model and harmonicity-derived spectral mask, and its coupling with feature compensationabstractThe use of a speech spectral envelope model defined in the log spectrum-type domain is a common approach to feature enhancement for noise robust speech recognition. However, from the noise reduction viewpoint, this approach ignores non-peak components of a spectrum and thus suffers from the poor SNR improvement during voiced periods. This paper proposes a speech enhancement method that exploits a log spectral envelope model and a harmonic structure. The key to the method is its use of a harmonic structure to define the prior distribution of a spectral mask, which is used for both accurate noise estimation and attenuation. In addition, we combine log mel-frequency feature enhancement with the above method to take advantage of low dimensionality. The whole proposed method outperforms a state-of-the-art speech enhancement method in four different noise environments. Takuya Yoshioka, Tomohiro Nakatani |
ICASSP | 2 |
| 2011 | A Robust Estimation Method of Noise Mixture Model for Noise Suppression
Masakiyo Fujimoto, Shinji Watanabe 0001, Tomohiro Nakatani |
INTERSPEECH | 3 |
| 2011 | Single Channel Dereverberation Using Example-Based Speech Enhancement with Uncertainty Decoding Technique
Keisuke Kinoshita, Mehrez Souden, Marc Delcroix, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2011 | Reduction of Highly Nonstationary Ambient Noise by Integrating Spectral and Locational Characteristics of Speech and Noise for Robust ASR
Tomohiro Nakatani, Shoko Araki, Marc Delcroix, Takuya Yoshioka, Masakiyo Fujimoto |
INTERSPEECH | 1 |
| 2011 | A Multichannel Feature-Based Processing for Robust Speech Recognition
Mehrez Souden, Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2011 | Blind Separation and Dereverberation of Speech Mixtures by Joint OptimizationabstractThis paper proposes a method for performing blind source separation (BSS) and blind dereverberation (BD) at the same time for speech mixtures. In most previous studies, BSS and BD have been investigated separately. The separation performance of conventional BSS methods deteriorates as the reverberation time increases while many existing BD methods rely on the assumption that there is only one sound source in a room. Therefore, it has been difficult to perform both BSS and BD when the reverberation time is long. The proposed method uses a network, in which dereverberation and separation networks are connected in tandem, to estimate source signals. The parameters for the dereverberation network (prediction matrices) and those for the separation network (separation matrices) are jointly optimized. This enables a BD process to take a BSS process into account. The prediction and separation matrices are alternately optimized with each depending on the other; hence, we call the proposed method the conditional separation and dereverberation (CSD) method. Comprehensive evaluation results are reported, where all the speech materials contained in the complete test set of the TIMIT corpus are used. The CSD method improves the signal-to-interference ratio by an average of about 4 dB over the conventional frequency-domain BSS approach for reverberation times of 0.3 and 0.5 s. The direct-to-reverberation ratio is also improved by about 10 dB. Takuya Yoshioka, Tomohiro Nakatani, Masato Miyoshi, Hiroshi G. Okuno |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Simultaneous clustering of mixing and spectral model parameters for blind sparse source separationabstractThis paper proposes a sparse source separation method which clusters the phase difference between the microphone observations and the amplitude modulation (AM) of the source spectrum simultaneously. The phase difference clustering separates the signals in each frequency bin, and the AM clustering corresponds to permutation alignment. Because the proposed method has an inherent ability to align the permutation of frequency components, the proposed method can be applied even when the spatial aliasing problem occurs. Moreover, because the common AM property collects the synchronized frequency components, we can model the microphone observations with a small number of sources. This property enables us to count the number of sources. That is, the proposed method can be applied even if the number of sources is unknown. The experimental results confirm the effectiveness of our proposed method. Shoko Araki, Tomohiro Nakatani, Hiroshi Sawada |
ICASSP | 2 |
| 2010 | Blind upmix of stereo music signals using multi-step linear prediction based reverberation extractionabstractWe propose a blind upmixing method for stereo music signals that utilizes multi-step linear prediction and decomposes the input signals into reverberation and the dereverberated signals. The proposed method is directly motivated by our previously proposed dereverberation algorithm that was shown to dereverberate speech signals well. In this paper, we first analyze the behavior of the multi-step linear prediction and investigate the reverberation reduction/extraction strategy using a stereo music signal model. Based on the analysis, we show that the proposed method can perform a dereverberation and reverberation extraction based on the stereo music signal, and achieve an efficient blind upmix of stereo music by assigning its dereverberated signal to the front channels and extracted reverberation to the rear channels. In the experiment, we apply the proposed upmixing method to real stereo music signals, and confirm its effectiveness with an objective evaluation and a preference test. Keisuke Kinoshita, Tomohiro Nakatani, Masato Miyoshi |
ICASSP | 2 |
| 2010 | Single channel source separation based on sparse source observation model with harmonic constraintabstractThis paper proposes a general single channel source separation approach that exploits statistical characteristics of the source including sparseness. A new observation model for a mixture of sparse sources is introduced for this purpose. With this approach, source separation is achieved by iterating two simple sub-procedures, namely the clustering of the time-frequency (TF) bins into individual sources and the separate updating of the model parameters of each source. An advantage of this approach is that we can update the model parameters of each source assuming each cluster to contain a single source, and thus we can utilize the various model parameter estimation algorithms used for single source analysis, which can be simple and accurate, in an efficient and unified manner. We implement a harmonicity based source separation method with this approach using a robust fundamental frequency (F0) estimation algorithm. The experimental results confirm the effectiveness of the proposed method. Tomohiro Nakatani, Shoko Araki |
ICASSP | 1 |
| 2010 | Music dereverberation using harmonic structure source model and Wiener filterabstractThis paper proposes a dereverberation method for musical audio signals. Existing dereverberation methods are designed for speech signals and are not necessarily effective for suppressing long and dense reverberation in musical audio signals because: 1) an all-pole model and a non-parametric model, which are used to represent source spectra, do not match musical tones, and 2) the conventional inverse-filter-based dereverberation is not effective for suppressing long and dense reverberation. To overcome the two problems, an appropriate dereverberation approach for musical audio signals is established. The first problem is resolved by using a harmonic Gaussian mixture model (GMM) to accurately model the harmonic structure of a source spectrum. The second problem is resolved by performing dereverberation with a Wiener filter based on both an estimated inverse filter and an estimated source spectrum model. Experimental results reveal the effectiveness of the proposed dereverberation method using these two solutions. Naoki Yasuraoka, Takuya Yoshioka, Tomohiro Nakatani, Atsushi Nakamura, Hiroshi G. Okuno |
ICASSP | 3 |
| 2010 | Noisy speech enhancement based on prior knowledge about spectral envelope and harmonic structureabstractThis paper considers the enhancement of noisy speech. Earlier studies have revealed that an approach that enhances spectral envelopes by using prior knowledge about the all-pole (AP) model parameters of clean speech learnt from speech corpora is advantageous in terms of the amount of musical noise and speech distortion. This paper proposes a new speech enhancement method, in which harmonic structure enhancement is incorporated in learning-based spectral envelope enhancement to further improve performance. The harmonic structure is represented by using a harmonic Gaussian mixture model (GMM), which is parameterized by a voicing indicator and a fundamental frequency. The parameters of the AP model and the harmonic GMM are jointly estimated by maximum a posteriori estimation, thus enabling the enhancement of spectral envelopes and harmonic structures in a unified framework. The proposed method outperforms the spectral envelope enhancement approach by 0.85 dB in cepstral distance. Takuya Yoshioka, Tomohiro Nakatani, Hiroshi G. Okuno |
ICASSP | 2 |
| 2010 | Voice activity detection using frame-wise model re-estimation method based on Gaussian pruning with weight normalizationabstractThis paper proposes a frame-wise model re-estimation method based on Gaussian pruning with weight normalization for noise robust voice activity detection (VAD). Our previous work, switching Kalman filter-based VAD, sequentially estimates a non-stationary noise Gaussian mixture model (GMM) and constructs GMMs of observed noisy speech signals by composing pre-trained silence and clean GMMs and sequentially estimated noise GMMs. However, the composed models are not optimal, because they do not fully reflect the characteristics of the observed signal. Thus, to ensure the optimality of the composed models, we investigate a method for re-estimating the composed model. Since our VAD method works under the frame-wise sequential processing, there are insufficient re-training data for re-estimation of whole model parameters. To solve this problem, we propose a model re-estimation method that involves the extraction of reliable information using Gaussian pruning with weight normalization. Namely, the proposed method reestimates the model by pruning non-dominant Gaussian distributions in expressing the local characteristics of each frame and by normalizing Gaussian weights of remaining distributions. Index Terms: voice activity detection, switching Kalman filter, Gaussian pruning, Gaussian weight normalization Masakiyo Fujimoto, Shinji Watanabe 0001, Tomohiro Nakatani |
INTERSPEECH | 3 |
| 2010 | Multichannel source separation based on source location cue with log-spectral shaping by hidden Markov source model
Tomohiro Nakatani, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto |
INTERSPEECH | 1 |
| 2010 | Cepstral smoothing of separated signals for underdetermined speech separationabstractMusical noise is a typical problem with blind source separation using a time-frequency mask. Recently, the cepstral smoothing of spectral masks (CSM) was proposed. Based on the idea of smoothing in the cepstral domain, this paper proposes the cepstral smoothing of separated signals (CSS) on the assumption that a cepstral representation better reflects the characteristics of speech signals than those of masks (or filter gains). We also report a comparative evaluation study of CSM and CSS with other musical noise reduction methods. Our experimental results show that CSM is effective for musical noise reduction, but the target speech was relatively distorted. On the other hand, our proposed CSS produced less distorted target signals with the same musical noise reduction as CSM. Yumi Ansa, Shoko Araki, Shoji Makino, Tomohiro Nakatani, Takeshi Yamada, Atsushi Nakamura, Nobuhiko Kitawaki |
ISCAS | 4 |
| 2010 | Real-time meeting recognition and understanding using distant microphones and omni-directional cameraabstractThis paper presents our newly developed real-time meeting analyzer for monitoring conversations in an ongoing group meeting. The goal of the system is to automatically recognize “who is speaking what” in an online manner for meeting assistance. Our system continuously captures the utterances and the face pose of each speaker using a distant microphone array and an omni-directional camera at the center of the meeting table. Through a series of advanced audio processing operations, an overlapping speech signal is enhanced and the components are separated into individual speaker's channels. Then the utterances are sequentially transcribed by our speech recognizer with low latency. In parallel with speech recognition, the activity of each participant (e.g. speaking, laughing, watching someone) and the situation of the meeting (e.g. topic, activeness, casualness) are detected and displayed on a browser together with the transcripts. In this paper, we describe our techniques and our attempt to achieve the low-latency monitoring of meetings, and we show our experimental results for real-time meeting transcription. Takaaki Hori, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto, Shinji Watanabe 0001, Takanobu Oba, Atsunori Ogawa, Kazuhiro Otsuka, Dan Mikami, Keisuke Kinoshita, Tomohiro Nakatani, Atsushi Nakamura, Junji Yamato |
SLT | 11 |
| 2010 | Noise robust voice activity detection based on periodic to aperiodic component ratio
Kentaro Ishizuka, Tomohiro Nakatani, Masakiyo Fujimoto, Noboru Miyazaki |
Speech Commun. | 2 |
| 2010 | Introduction to the Special Issue on Processing Reverberant Speech: Methodologies and ApplicationsabstractThe 17 papers in this special issue focus on the methodologies and applications of processing reverberant speech. The issue highlights some major aspects of the recent progress in the field. Tomohiro Nakatani, Walter Kellermann, Patrick A. Naylor, Masato Miyoshi, Biing-Hwang Juang |
IEEE Trans. Speech Audio Process. | 1 |
| 2010 | Speech Dereverberation Based on Variance-Normalized Delayed Linear PredictionabstractThis paper proposes a statistical model-based speech dereverberation approach that can cancel the late reverberation of a reverberant speech signal captured by distant microphones without prior knowledge of the room impulse responses. With this approach, the generative model of the captured signal is composed of a source process, which is assumed to be a Gaussian process with a time-varying variance, and an observation process modeled by a delayed linear prediction (DLP). The optimization objective for the dereverberation problem is derived to be the sum of the squared prediction errors normalized by the source variances; hence, this approach is referred to as variance-normalized delayed linear prediction (NDLP). Inheriting the characteristic of DLP, NDLP can robustly estimate an inverse system for late reverberation in the presence of noise without greatly distorting a direct speech signal. In addition, owing to the use of variance normalization, NDLP allows us to improve the dereverberation result especially with relatively short (of the order of a few seconds) observations. Furthermore, NDLP can be implemented in a computationally efficient manner in the time-frequency domain. Experimental results demonstrate the effectiveness and efficiency of the proposed approach in comparison with two existing approaches. Tomohiro Nakatani, Takuya Yoshioka, Keisuke Kinoshita, Masato Miyoshi, Biing-Hwang Juang |
IEEE Trans. Speech Audio Process. | 1 |
| 2009 | Blind sparse source separation for unknown number of sources using Gaussian mixture model fitting with Dirichlet priorabstractIn this paper, we propose a novel sparse source separation method that can be applied even if the number of sources is unknown. Recently, many sparse source separation approaches with time-frequency masks have been proposed. However, most of these approaches require information on the number of sources in advance. In our proposed method, we model the histogram of the estimated direction of arrival (DOA) with a Gaussian mixture model (GMM) with a Dirichlet prior. Then we estimate the model parameters by using the maximum a posteriori estimation based on the EM algorithm. In order to avoid one cluster being modeled by two or more Gaussians, we utilize a sparse distribution modeled by the Dirichlet distributions as the prior of the GMM mixture weight. By using this prior, without any specific model selection process, our proposed method can estimate the number of sources and time-frequency masks simultaneously. Experimental results show the performance of our proposed method. Shoko Araki, Tomohiro Nakatani, Hiroshi Sawada, Shoji Makino |
ICASSP | 2 |
| 2009 | Robust speech dereverberation based on non-negativity and sparse nature of speech spectrogramsabstractThis paper presents a blind dereverberation method designed to recover the subband envelope of an original speech signal from its reverberant version. The problem is formulated as a blind deconvolution problem with non-negative constraints, regularized by the sparse nature of speech spectrograms. We derive an iterative algorithm for its optimization, which can be seen as a special case of the non-negative matrix factor deconvolution. We confirmed through experiments that the algorithm is fast and robust to speaker movement. Hirokazu Kameoka, Tomohiro Nakatani, Takuya Yoshioka |
ICASSP | 2 |
| 2009 | Real-time speech enhancement in noisy reverberant multi-talker environments based on a location-independent room acoustics modelabstractThis paper describes a new real-time speech enhancement method that reduces signal distortion caused by stationary noise and late reflections of reverberation in speech signals captured by a single distant microphone under multi-talker conditions. A major problem here is how to estimate the energy of the late reflections in real time when the room impulse responses from individual talkers to the microphone are not given or fixed in advance. To solve this problem, we introduce a probabilistic room acoustics model, and provide a method for estimating the energy of late reflections based on this model. In this method, parameters of the model for a room can be fixed in advance only from a few seconds of observation. By incorporating the proposed approach into a conventional frequency domain noise reduction scheme, we realize an integrated real-time speech enhancement framework. The effectiveness of the proposed method is confirmed experimentally for a case where there are two talkers in a room. Tomohiro Nakatani, Takuya Yoshioka, Keisuke Kinoshita, Masato Miyoshi, Biing-Hwang Juang |
ICASSP | 1 |
| 2009 | Adaptive dereverberation of speech signals with speaker-position change detectionabstractThis paper proposes a method for adaptive speech dereverberation and speaker-position change detection, which have not previously been addressed. Signal transmission channels in rooms are modeled as auto-regressive systems in individual frequency bands. The proposed method adaptively estimates the regression coefficients of this model, which are called room regression coefficients (RRCs). The proposed method has two distinguishing features: (1) The method is based on the weighted recursive least squares algorithm, which enables an efficient RRC-estimate update as well as a fast convergence rate; (2) The method detects changes in speaker position and so can quickly catch up with the sudden channel changes that such position changes cause. Detection is realized by finding time frames where the power of dereverberated speech is anomalously amplified. Experimental results showed that the proposed method attained convergence in 5 seconds and successfully detected changes in speaker position. Takuya Yoshioka, Hideyuki Tachibana, Tomohiro Nakatani, Masato Miyoshi |
ICASSP | 3 |
| 2009 | A speaker diarization method based on the probabilistic fusion of audio-visual location informationabstractThis paper proposes a speaker diarization method for determining ""who spoke when"" in multi-party conversations, based on the probabilistic fusion of audio and visual location information. The audio and visual information is obtained from a compact system designed to analyze round table multi-party conversations. The system consists of two cameras and a triangular microphone array with three microphones, and can cover a spherical region. Speaker locations are estimated from audio and visual observations in terms of azimuths from this recording system. Unlike conventional speech diarization methods, our proposed method estimates the probability of the presence of multiple simultaneous speakers in a physical space with a small microphone setup instead of using a cascade consisting of speech activity detection, direction of arrival estimation, acoustic feature extraction, and information criteria based speaker segmentation. To estimate the speaker presence more correctly, the speech presence probabilities in a physical space are integrated with the probabilities estimated from participants' face locations obtained with a robust particle filtering based face tracker with two cameras equipped with fisheye lenses. The locations in a physical space with highly integrated probabilities are then classified into a certain number of speaker classes by using on-line classification to realize speaker diarization. The probability calculations and speaker classifications are conducted on-line, making it unnecessary to observe all the conversation data. An experiment using real casual conversations, which include more overlaps and short speech segments than formal meetings, showed the advantages of the proposed method. Kentaro Ishizuka, Shoko Araki, Kazuhiro Otsuka, Tomohiro Nakatani, Masakiyo Fujimoto |
ICMI | 4 |
| 2009 | A study of mutual front-end processing method based on statistical model for noise robust speech recognitionabstractThis paper addresses robust front-end processing for automatic speech recognition (ASR) in noise. Accurate recognition of corrupted speech requires noise robust front-end processing, e.g., voice activity detection (VAD) and noise suppression (NS). Typically, VAD and NS are combined as one-way processing, and are developed independently. However, VAD and NS should not be assumed to be independent techniques, because sharing each others’ information is important for the improvement of front-end processing. Thus, we investigate the mutual front-end processing by integrating VAD and NS, which can beneficially share each others’ information. In an evaluation of a concatenated speech corpus, CENSREC-1-C database, the proposed method improves the performance of both VAD and ASR compared with the conventional method. Index Terms: voice activity detection, noise suppression, mutual front-end processing, speech recognition Masakiyo Fujimoto, Kentaro Ishizuka, Tomohiro Nakatani |
INTERSPEECH | 3 |
| 2009 | Development of Japanese infant speech database from longitudinal recordings
Shigeaki Amano, Tadahisa Kondo, Kazumi Kato, Tomohiro Nakatani |
Speech Commun. | 4 |
| 2009 | Static and Dynamic Variance Compensation for Recognition of Reverberant Speech With Dereverberation PreprocessingabstractThe performance of automatic speech recognition is severely degraded in the presence of noise or reverberation. Much research has been undertaken on noise robustness. In contrast, the problem of the recognition of reverberant speech has received far less attention and remains very challenging. In this paper, we use a dereverberation method to reduce reverberation prior to recognition. Such a preprocessor may remove most reverberation effects. However, it often introduces distortion, causing a dynamic mismatch between speech features and the acoustic model used for recognition. Model adaptation could be used to reduce this mismatch. However, conventional model adaptation techniques assume a static mismatch and may therefore not cope well with a dynamic mismatch arising from dereverberation. This paper proposes a novel adaptation scheme that is capable of managing both static and dynamic mismatches. We introduce a parametric model for variance adaptation that includes static and dynamic components in order to realize an appropriate interconnection between dereverberation and a speech recognizer. The model parameters are optimized using adaptive training implemented with the expectation maximization algorithm. An experiment using the proposed method with reverberant speech for a reverberation time of 0.5 s revealed that it was possible to achieve an 80% reduction in the relative error rate compared with the recognition of dereverberated speech (word error rate of 31%), and the final error rate was 5.4%, which was obtained by combining the proposed variance compensation and MLLR adaptation. Marc Delcroix, Tomohiro Nakatani, Shinji Watanabe 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Suppression of Late Reverberation Effect on Speech Signal Using Long-Term Multiple-step Linear PredictionabstractA speech signal captured by a distant microphone is generally smeared by reverberation, which severely degrades automatic speech recognition (ASR) performance. One way to solve this problem is to dereverberate the observed signal prior to ASR. In this paper, a room impulse response is assumed to consist of three parts: a direct-path response, early reflections and late reverberations. Since late reverberations are known to be a major cause of ASR performance degradation, this paper focuses on dealing with the effect of late reverberations. The proposed method first estimates the late reverberations using long-term multi-step linear prediction, and then reduces the late reverberation effect by employing spectral subtraction. The algorithm provided good dereverberation with training data corresponding to the duration of one speech utterance, in our case, less than 6 s. This paper describes the proposed framework for both single-channel and multichannel scenarios. Experimental results showed substantial improvements in ASR performance with real recordings under severe reverberant conditions. Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani, Masato Miyoshi |
IEEE Trans. Speech Audio Process. | 3 |
| 2009 | Integrated Speech Enhancement Method Using Noise Suppression and DereverberationabstractThis paper proposes a method for enhancing speech signals contaminated by room reverberation and additive stationary noise. The following conditions are assumed. 1) Short-time spectral components of speech and noise are statistically independent Gaussian random variables. 2) A room's convolutive system is modeled as an autoregressive system in each frequency band. 3) A short-time power spectral density of speech is modeled as an all-pole spectrum, while that of noise is assumed to be time-invariant and known in advance. Under these conditions, the proposed method estimates the parameters of the convolutive system and those of the all-pole speech model based on the maximum likelihood estimation method. The estimated parameters are then used to calculate the minimum mean square error estimates of the speech spectral components. The proposed method has two significant features. 1) The parameter estimation part performs noise suppression and dereverberation alternately. (2) Noise-free reverberant speech spectrum estimates, which are transferred by the noise suppression process to the dereverberation process, are represented in the form of a probability distribution. This paper reports the experimental results of 1500 trials conducted using 500 different utterances. The reverberation time RT60was 0.6 s, and the reverberant signal to noise ratio was 20, 15, or 10 dB. The experimental results show the superiority of the proposed method over the sequential performance of the noise suppression and dereverberation processes. Takuya Yoshioka, Tomohiro Nakatani, Masato Miyoshi |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Combined static and dynamic variance adaptation for efficient interconnection of speech enhancement pre-processor with speech recognizerabstractIt is well known that automatic speech recognition performs poorly in presence of noise or reverberation. Much research has been undertaken on model adaptation and speech enhancement to increase the robustness of speech recognizers. Model adaptation is effective to remove static mismatch between speech features and acoustic model parameters, but may not cope well with dynamic mismatch. Speech enhancement approaches can reduce dynamic perturbations, but often do not interconnect well with speech recognizer. There seems to be a lack of optimal way to combine these two approaches. In this paper we propose introducing the dynamic capabilities of speech enhancement into a static adaptation scheme. We focus on variance adaptation, and propose a novel parametric variance model that includes static and dynamic components. The dynamic component is derived from a speech enhancement pre-process, and the parameters of the model are optimized using an adaptive training scheme. An evaluation of the method with a speech dereverberation for preprocessing revealed that a 80 % relative error rate reduction was possible compared with the recognition of dereverberated speech, and the final error rate was 5.4 % which is close to that of clean speech (1.2%). Marc Delcroix, Tomohiro Nakatani, Shinji Watanabe 0001 |
ICASSP | 2 |
| 2008 | A voice activity detection based on the adaptive integration of multiple speech features and a signal decision schemeabstractThis paper addresses the problem of voice activity detection (VAD) in noisy environments. The VAD method proposed in this paper integrates multiple speech features and a signal decision scheme, namely the speech periodic to aperiodic component ratio and a switching Kalman filter. The integration is carried out by using the weighted sum of likelihoods outputted from each VAD (stream). The stream weight is decided adaptively each short time frame. The evaluation is carried out by using a VAD evaluation framework, CENSREC1-C. The evaluation results revealed that the proposed method significantly outperforms the baseline results of CENSREC-1-C as regards VAD accuracy in real environments. In addition, we carried out speech recognition evaluations by using detected speech signals, and confirmed that the proposed method contributes to an improvement in speech recognition accuracy. Masakiyo Fujimoto, Kentaro Ishizuka, Tomohiro Nakatani |
ICASSP | 3 |
| 2008 | Blind speech dereverberation with multi-channel linear prediction based on short time fourier transform representationabstractIt has recently been shown that the use of the time-varying nature of speech signals allows us to achieve high quality speech dereverberation based on multi-channel linear prediction (MCLP). However, this approach requires a huge computing cost for calculating large covariance matrices in the time domain. In addition, we face the important problem of how to combine the speech dereverberation efficiently with many other useful speech enhancement techniques in the short time Fourier transform (STFT) domain. As the first step to overcoming these problems, this paper presents methods for implementing MCLP based speech dereverberation that allow it to work in the STFT domain with much less computing cost. The effectiveness of the present methods is confirmed by experiments in terms of the recovered signal quality and the computing time. Tomohiro Nakatani, Takuya Yoshioka, Keisuke Kinoshita, Masato Miyoshi, Biing-Hwang Juang |
ICASSP | 1 |
| 2008 | Maximum likelihood approach to speech enhancement for noisy reverberant signalsabstractThis paper proposes a speech enhancement method for signals contaminated by room reverberation and additive background noise. The following conditions are assumed: (1) The spectral components of speech and noise are statistically independent Gaussian random variables. (2) The convolutive distortion channel is modeled as an auto-regressive system in each frequency bin. (3) The power spectral density of speech is modeled as an all-pole spectrum, while that of noise is assumed to be stationary and given in advance. Under these conditions, the proposed method estimates the parameters of the channel and those of the all-pole speech model based on the maximum likelihood estimation method. Experimental results showed that the proposed method successfully suppressed the reverberation and additive noise from three-second noisy reverberant signals when the reverberation time was 0.5 seconds and the reverberant signal to noise ratio was 10 dB. Takuya Yoshioka, Tomohiro Nakatani, Takafumi Hikichi, Masato Miyoshi |
ICASSP | 2 |
| 2008 | Study of integration of statistical model-based voice activity detection and noise suppression
Masakiyo Fujimoto, Kentaro Ishizuka, Tomohiro Nakatani |
INTERSPEECH | 3 |
| 2008 | Missing feature speech recognition in a meeting situation with maximum SNR beamformingabstractEspecially for tasks like automatic meeting transcription, it would be useful to automatically recognize speech also while multiple speakers are talking simultaneously. For this purpose, speech separation can be performed, for example by using maximum SNR beamforming. However, even when good interferer suppression is attained, the interfering speech will still be recognizable during those intervals, where the target speaker is silent. In order to avoid the consequential insertion errors, a new soft masking scheme is proposed, which works in the time domain by inducing a large damping on those temporal periods, where the observed direction of arrival does not correspond to that of the target speaker. Even though the masking scheme is aggressive, by means of missing feature recognition the recognition accuracy can be improved significantly, with relative error reductions in the order of 60% compared to maximum SNR beamforming alone, and it is successful also for three simultaneously active speakers. Results are reported based on the SOLON speech recognizer, NTT’s large vocabulary system [1], which is applied here for the recognition of artificially mixed data using real-room impulse responses and the entire clean test set of the Aurora 2 database. Dorothea Kolossa, Shoko Araki, Marc Delcroix, Tomohiro Nakatani, Reinhold Orglmeister, Shoji Makino |
ISCAS | 4 |
| 2008 | A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments
Tomohiro Nakatani, Shigeaki Amano, Toshio Irino, Kentaro Ishizuka, Tadahisa Kondo |
Speech Commun. | 1 |
| 2008 | Speech Dereverberation Based on Maximum-Likelihood Estimation With Time-Varying Gaussian Source ModelabstractDistant acquisition of acoustic signals in an enclosed space often produces reverberant components due to acoustic reflections in the room. Speech dereverberation is in general desirable when the signal is acquired through distant microphones in such applications as hands-free speech recognition, teleconferencing, and meeting recording. This paper proposes a new speech dereverberation approach based on a statistical speech model. A time-varying Gaussian source model (TVGSM) is introduced as a model that represents the dynamic short time characteristics of nonreverberant speech segments, including the time and frequency structures of the speech spectrum. With this model, dereverberation of the speech signal is formulated as a maximum-likelihood (ML) problem based on multichannel linear prediction, in which the speech signal is recovered by transforming the observed signal into one that is probabilistically more like nonreverberant speech. We first present a general ML solution based on TVGSM, and derive several dereverberation algorithms based on various source models. Specifically, we present a source model consisting of a finite number of states, each of which is manifested by a short time speech spectrum, defined by a corresponding autocorrelation (AC) vector. The dereverberation algorithm based on this model involves a finite collection of spectral patterns that form a codebook. We confirm experimentally that both the time and frequency characteristics represented in the source models are very important for speech dereverberation, and that the prior knowledge represented by the codebook allows us to further improve the dereverberated speech quality. We also confirm that the quality of reverberant speech signals can be greatly improved in terms of the spectral shape and energy time-pattern distortions from simply a short speech signal using a speaker-independent codebook. Tomohiro Nakatani, Biing-Hwang Juang, Takuya Yoshioka, Keisuke Kinoshita, Marc Delcroix, Masato Miyoshi |
IEEE Trans. Speech Audio Process. | 1 |
| 2007 | Study on Speech Dereverberation with Autocorrelation CodebookabstractThis paper proposes a new speech dereverberation approach based on a statistical speech model. An autocorrelation codebook is introduced as a model that can represent time-varying short-time speech characteristics corresponding to the cepstrum and harmonics. The speech dereverberation is formulated as a likelihood maximization problem, in which the quality of a speech signal is recovered by turning the signal into one that is probabilistically more like a clean speech. Two dereverberation algorithms are derived based on different scenarios, regularized inversion and inverse filter estimation. Experimental results show that the proposed approach allows us to reduce both reverberation and noise with the regularized inversion, and to estimate inverse filters that can dereverberate signals effectively from just a small number of observed signals. Tomohiro Nakatani, Biing-Hwang Juang, Takafumi Hikichi, Takuya Yoshioka, Keisuke Kinoshita, Marc Delcroix, Masato Miyoshi |
ICASSP (1) | 1 |
| 2007 | Two-Microphone Voice Activity Detection Based on the Homogeneity of the Direction of Arrival EstimatesabstractVoice activity detection (VAD) systems have been the object of continuous research during the last three decades. While single microphone systems cannot take advantage of certain spatial properties of speech signals, microphone array systems consisting of many elements based on beamforming techniques can be difficult to implement in reality due to cost and complexity issues. The aim of the work described in this paper was to achieve both practical feasibility and spatial discrimination ability. A new approach is developed for two-microphone VAD capable of profiting from the concentration of speech energy in time, frequency and space. The algorithm is implemented and compared with several standard VAD algorithms, such as AFE, AMR and G.729B, and other recently proposed systems, revealing promising results under real-world noise conditions. The main advantage of the proposed approach is its capacity to outperform the above methods without the need for any spatial or spectral constraints, which makes it both versatile and capable of further improvement. Juan E. Rubio, Kentaro Ishizuka, Hiroshi Sawada, Shoko Araki, Tomohiro Nakatani, Masakiyo Fujimoto |
ICASSP (4) | 5 |
| 2007 | Noise robust front-end processing with voice activity detection based on periodic to aperiodic component ratio
Kentaro Ishizuka, Tomohiro Nakatani, Masakiyo Fujimoto, Noboru Miyazaki |
INTERSPEECH | 2 |
| 2007 | Multi-step linear prediction based speech dereverberation in noisy reverberant environment
Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani, Masato Miyoshi |
INTERSPEECH | 3 |
| 2007 | Joint Source-Channel Modeling and Estimation for Speech DereverberationabstractSpeech dereverberation is an important challenge in acoustic signal processing because of the detrimental effects upon the signal quality brought by the reverberant components due to acoustic reflections from the walls of the room enclosure. Most of the methods proposed so far such as microphone array beamforming and room impulse response inversion for deconvolution make very little use of the knowledge about the characteristics of the source signal. This paper presents a formulation of speech dereverberation as a probabilistic modeling problem in which joint estimation of the source and the channel (or its inverse) can be accomplished. We discuss reasonable representations of the source prior for inclusion in such a formulation and propose the corresponding solutions. Biing-Hwang Juang, Tomohiro Nakatani |
ISCAS | 2 |
| 2007 | Robust blind dereverberation of speech signals based on characteristics of short-time speech segmentsabstractThis paper addresses blind dereverberation techniques based on the inherent characteristics of speech signals. Two challenging issues for speech dereverberation involve decomposing reverberant observed signals into colored sources and room transfer functions (RTFs), and making the inverse filtering robust as regards acoustic and system noise. We show that short-time speech characteristics are very important for this task, and that multi-channel linear prediction (MCLP) is a useful tool for achieving robust inverse filtering. As examples, we detail three recently proposed robust dereverberation methods. By assuming the source to be a small order autoregressive process, we can present an efficient source estimation method that reduces late reverberation reflections of the reverberation using multi-step linear prediction. By exploiting the time-varying nature of the speech signals, we can also develop a method that can estimate both the source and the inverse filters of the RTFs. Furthermore, we can achieve high quality speech dereverberation by formulating the problem as a likelihood maximization problem using a statistical speech model that represents the spectral characteristics of short-time speech segments including harmonicity. Tomohiro Nakatani, Takafumi Hikichi, Keisuke Kinoshita, Takuya Yoshioka, Marc Delcroix, Masato Miyoshi, Biing-Hwang Juang |
ISCAS | 1 |
| 2007 | Harmonicity-Based Blind Dereverberation for Single-Channel Speech SignalsabstractThe distant acquisition of acoustic signals in an enclosed space often produces reverberant artifacts due to the room impulse response. Speech dereverberation is desirable in situations where the distant acquisition of acoustic signals is involved. These situations include hands-free speech recognition, teleconferencing, and meeting recording, to name a few. This paper proposes a processing method, named Harmonicity-based dEReverBeration (HERB), to reduce the amount of reverberation in the signal picked up by a single microphone. The method makes extensive use of harmonicity, a unique characteristic of speech, in the design of a dereverberation filter. In particular, harmonicity enhancement is proposed and demonstrated as an effective way of estimating a filter that approximates an inverse filter corresponding to the room impulse response. Two specific harmonicity enhancement techniques are presented and compared; one based on an average transfer function and the other on the minimization of a mean squared error function. Prototype HERB systems are implemented by introducing several techniques to improve the accuracy of dereverberation filter estimation, including time warping analysis. Experimental results show that the proposed methods can achieve high-quality speech dereverberation, when the reverberation time is between 0.1 and 1.0 s, in terms of reverberation energy decay curves and automatic speech recognition accuracy Tomohiro Nakatani, Keisuke Kinoshita, Masato Miyoshi |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Spectral Subtraction Steered by Multi-Step Forward Linear Prediction For Single Channel Speech DereverberationabstractA speech signal captured by a distant microphone is generally smeared by reverberation, which severely degrades automatic speech recognition (ASR) performance. In this paper, we propose a novel dereverberation method utilizing multi-step forward linear prediction. It precisely estimates and suppresses the late reflections, which constitute a major cause of ASR performance degradation. Our experimental results showed that the proposed method can improve ASR performance significantly even without using special adaptation methods such as multi-condition acoustic model training Keisuke Kinoshita, Tomohiro Nakatani, Masato Miyoshi |
ICASSP (1) | 2 |
| 2006 | Speech Dereverberation Based on Probabilistic Models of Source and Room AcousticsabstractThis paper proposes a new single channel speech dereverberation method, in which the features of source signals and room acoustics are represented by probabilistic density functions (pdf) and the source signals are estimated by maximizing a likelihood function defined based on the pdfs. Two types of pdfs are introduced for the source signals, based on two essential speech signal features, harmonicity and sparseness, while the pdf for the room acoustics is defined based on an inverse filtering operation. The EM algorithm is used to solve this maximum likelihood problem efficiently. The resultant algorithm elaborates the initial source signal estimate given solely based on its source signal features by integrating them with the room acoustics feature through the EM iteration. The effectiveness of the present method is shown in terms of the energy decay curves of the dereverberated impulse responses Tomohiro Nakatani, Biing-Hwang Juang, Keisuke Kinoshita, Masato Miyoshi |
ICASSP (1) | 1 |
| 2006 | A feature extraction method using subband based periodicity and aperiodicity decomposition with noise robust frontend processing for automatic speech recognition
Kentaro Ishizuka, Tomohiro Nakatani |
Speech Commun. | 2 |
| 2005 | Speech Signal Analysis with Exponential Autoregressive ModelabstractThe paper proposes a speech signal analysis approach that uses an exponential autoregressive (ExpAR) model. In real speech signals, the amplitude and frequency fluctuate randomly. These fluctuations are non-Gaussian and have nonlinear dynamics. This means that they cannot be modeled adequately with linear AR models or compositions of sine/cosine waves, as these analysis methods are known to be affected by such fluctuations. Our proposed approach, using the ExpAR model, can deal with such fluctuations, and it is autoregressive in form with amplitude dependent exponential coefficients. Studies to fit the ExpAR model to real speech data have shown that AIC (Akaike's information criteria) values achieved by the ExpAR model are better (lower) than those obtained with a linear AR model, and that the ExpAR model provides a good model of speech fluctuations as movements of the position of its poles. The coefficients change with time depending on the amplitude of the speech signals, and so this model is also capable of realizing a fine instantaneous spectral estimation. The modeling of such speech fluctuations has the potential to be used for improving automatic speech recognition performance in clean or noisy environments, and the naturalness of synthesized speech. Kentaro Ishizuka, Hiroko Kato Solvang, Tomohiro Nakatani |
ICASSP (1) | 3 |
| 2005 | Fast Estimation of a Precise Dereverberation Filter based on Speech HarmonicityabstractA speech signal captured by a distant microphone is generally smeared by reverberation. This severely degrades both the speech intelligibility and automatic speech recognition (ASR) performance. In this paper, we propose a new dereverberation scheme based on harmonicity based dereverberation (HERB), aiming primarily at reducing the amount of training data needed to estimate an inverse filter. We show experimentally that our new dereverberation scheme successfully achieves high quality dereverberation with much smaller amounts of training data, and is very effective at improving both audible quality and ASR performance, even in unknown severely reverberant environments. Keisuke Kinoshita, Tomohiro Nakatani, Masato Miyoshi |
ICASSP (1) | 2 |
| 2005 | Efficient blind dereverberation framework for automatic speech recognition
Keisuke Kinoshita, Tomohiro Nakatani, Masato Miyoshi |
INTERSPEECH | 2 |
| 2004 | Developmental changes in voiced-segment ratio for Japanese infants and parentsabstractUtterances of five Japanese infants and their parents were recorded longitudinally and used to develop an infant speech database. This database was used to analyze the voicedsegment ratio to investigate the developmental changes in utterances produced by infants and parents. The voicedsegment ratio is the ratio of the summed duration of a voiced segment to the total utterance duration. The analyses showed that the ratio tends to increase in an infant's utterances before the infant starts to produce 2-word sentences. The analyses also showed that, at the same stage, the ratio is higher in parents' infant-directed speech than in parents' adult-directed speech. These results suggest that the voiced-segment ratio reflects the development of an infant's utterance ability, and that a higher voiced-segment ratio is one of the characteristics of infant-directed speech. Shigeaki Amano, Tomohiro Nakatani, Tadahisa Kondo |
INTERSPEECH | 2 |
| 2004 | Disambiguation in determining phonemes of sound-imitation words for environmental sound recognitionabstractOnomatopoeia, or sound-imitation words (SIWs) are important in informing sound events in human-computer communication. One problem is listener-dependency in recognizing environmental sounds by means of SIWs, that is, different listener hears the same environmental sound as a different SIW even under the same condition. Therefore, the use of usual Japanese phonemes is not adequate to express SIWs. To cope with this ambiguity problem of phoneme determination, we designed a set of new phonemes, referred to as the basic phoneme-groups, to represent environmental sounds. The basic phonemegroup consists of one or more Japanese phonemes, and thus the ambiguity problem is resolved based on it by generating one or more SIWs for a sound event. An HMM-based scheme is adopted to recognize SIWs using the phoneme-groups. Listening experiments with seven subjects showed that automatic SIW recognition based on the basic phoneme-groups outperformed ones based on the other types of phonemes. The recall and precision rate were 56.4% and 72.2%, respectively. Kazushi Ishihara, Yuya Hattori, Tomohiro Nakatani, Kazunori Komatani, Tetsuya Ogata, Hiroshi G. Okuno |
INTERSPEECH | 3 |
| 2004 | Improvement in robustness of speech feature extraction method using sub-band based periodicity and aperiodicity decompositionabstractThis paper shows improvements in robustness of a speech feature extraction method using Sub-band based Periodicity and Aperiodicity DEcomposition or SPADE. With SPADE, the speech signal is divided into sub-band signals through bandpass filter banks, after which the sub-band signal is decomposed into its periodic and aperiodic features by the comb filter. The comb filters are designed individually based on estimated periodicities of each sub-band signal. Both the periodic and aperiodic features are used as speech feature parameters. The evaluation experiment conducted with AURORA-2J (Japanese AURORA2) shows that SPADE certainly reduces the averaged word error rate (WER) under clean-speech training. SPADE also improves the performance under multicondition-speech training when the noise condition of the test data is open. However, SPADE degrades the performance when the channel condition of the test data is open. To cope with this problem, in this paper we apply the cepstral mean normalization (CMN) to SPADE. The result shows that CMN greatly improves the performance not only for test data under the open-channel condition but also for data under the closed-channel condition. SPADE with CMN achieves an averaged word accuracy of 89.96 %, and an averaged WER reduction of 28.61 %. This word accuracy is better than that achieved by using MFCC with CMN. Kentaro Ishizuka, Noboru Miyazaki, Tomohiro Nakatani, Yasuhiro Minami |
INTERSPEECH | 3 |
| 2004 | Improving automatic speech recognition performance and speech inteligibility with harmonicity based dereverberationabstractA speech signal captured by a distant microphone is generally smeared by reverberation, that severely degrades both the speech intelligibility and Automatic Speech Recognition (ASR) performance. Previously, we proposed a novel dereverberation method, named “Harmonicity based dEReverBeration (HERB)”, which estimates the inverse filter of an unknown impulse response by utilizing the inherent speech property, harmonics. In this paper, we carry out a formal evaluation of speech intelligibility for dereverberated speech, and further investigate HERB’s possibilities to improve ASR performance. Experimental results show that HERB is able to improve speech intelligibility to the level of clean speech. HERB is also found to be very effective at improving ASR performance, even under unknown severe reverberant environments by being used with MLLR and a multicondition acoustic model. Keisuke Kinoshita, Tomohiro Nakatani, Masato Miyoshi |
INTERSPEECH | 2 |
| 2004 | Harmonicity based monaural speech dereverberation with time warping and F0 adaptive windowabstractAlthough a number of dereverberation methods have been reported, dereverberation is still a challenging problem especially when a single microphone is used. To overcome this problem, we proposed a harmonicity based dereverberation method (HERB). HERB can blindly estimate the inverse filter of a room transfer function based on the harmonicity of speech signals and dereverberate the signals. However, HERB uses an imprecise assumption that hinders the dereverberation performance, that is, the fundamental frequency ( ) of a speech signal is assumed to be constant within a short time frame when extracting the features of harmonic components. In this paper, we combine HERB with time warping analysis and an adaptive window to remove this bottleneck. This extension makes it possible to estimate harmonic components precisely even when their frequencies change rapidly. Experiments show that time warping analysis with an adaptive window can effectively improve the dereverberation effect of HERB. Tomohiro Nakatani, Keisuke Kinoshita, Masato Miyoshi, Parham Zolfaghari |
INTERSPEECH | 1 |
| 2004 | Automatic Sound-Imitation Word Recognition from Environmental Sounds Focusing on Ambiguity Problem in Determining Phonemes
Kazushi Ishihara, Tomohiro Nakatani, Tetsuya Ogata, Hiroshi G. Okuno |
PRICAI | 2 |
| 2003 | Blind dereverberation of single channel speech signal based on harmonic structureabstractThe paper presents a new method for dereverberation of speech signals with a single microphone. For applications such as speech recognition, reverberant speech causes serious problems when a distant microphone is used in recording. This is especially severe when the reverberation time exceeds 0.5 s. We propose a method which uses the fundamental frequency (F/sub 0/) of the target speech as the primary feature for dereverberation. This method initially estimates F/sub 0/ and the harmonic structure of the speech signal and then obtains a dereverberation operator. This operator transforms the reverberant signal to its direct signal based on an inverse filtering operation. Dereverberation is achieved without prior knowledge of either the room acoustics or the target speech. Experimental results show that the dereverberation operator estimated from 5240 Japanese word utterances could effectively reduce the reverberation when the reverberation time is longer than 0.1 s. Tomohiro Nakatani, Masato Miyoshi |
ICASSP (1) | 1 |
| 2003 | Dominance spectrum based v/UV classification and f_0 estimation
Tomohiro Nakatani, Toshio Irino, Parham Zolfaghari |
INTERSPEECH | 1 |
| 2003 | Glottal closure instant synchronous sinusoidal model for high quality speech analysis/synthesis
Parham Zolfaghari, Tomohiro Nakatani, Toshio Irino, Hideki Kawahara, Fumitada Itakura |
INTERSPEECH | 2 |
| 2003 | One Microphone Blind Dereverberation Based on Quasi-periodicity of Speech SignalsabstractSpeech dereverberation is desirable with a view to achieving, for exam- ple, robust speech recognition in the real world. However, it is still a chal- lenging problem, especially when using a single microphone. Although blind equalization techniques have been exploited, they cannot deal with speech signals appropriately because their assumptions are not satisfied by speech signals. We propose a new dereverberation principle based on an inherent property of speech signals, namely quasi-periodicity. The present methods learn the dereverberation filter from a lot of speech data with no prior knowledge of the data, and can achieve high quality speech dereverberation especially when the reverberation time is long. Tomohiro Nakatani, Masato Miyoshi, Keisuke Kinoshita |
NIPS | 1 |
| 2002 | Evaluation of a speech recognition / generation method based on HMM and straightabstractWe propose a method for integrating speech recognition and generation within a unified framework. The method consists of STRAIGHT, warped-frequency DCT, and an HMM engine. The warped-frequency DCT i s used to derive a kind of mel-cepstral coefficient from the smoothed spectrum of STRAIGHT, which is known as a high-quality vocoder. This analysis/synthesis method has potential to improve the performance beyond a conventional method using the MFCC derived from the STFT. We evaluated the method by using speakerdependent speech recognition as well as by the perceptual evaluation of sounds generated by HMM text-to-speech. The recognition rate using the coefficients from the warped-DCT of the STRAIGHT spectrum was almost the same as that obtained using conventional MFCCs. The sound quality was sufficiently good for a fundamental system. 1. Toshio Irino, Yasuhiro Minami, Tomohiro Nakatani, Minoru Tsuzaki, H. Tagawa |
INTERSPEECH | 3 |
| 2002 | Robust fundamental frequency estimation against background noise and spectral distortionabstractThis paper presents a new method for robust fundamental frequency (F0) estimation in the presence of background noise and spectral distortion. We define degree of dominance and a dominance spectrum based on instantaneous frequencies. The degree of dominance allows us to evaluate the magnitude of individual harmonic components of speech signals relative to background noise while eliminating the influence of spectral distortion. The fundamental frequency is robustly estimated from reliable harmonic components easily selected from the dominance spectra. Experiments are performed using white and multi-talker background noise with and without spectral distortion produced by a SRAEN filter. Results show that the present method is better than the commonlyused methods in terms of correct F0 rates. Tomohiro Nakatani, Toshio Irino |
INTERSPEECH | 1 |
| 1999 | Harmonic sound stream segregation using localization and its application to speech stream segregation
Tomohiro Nakatani, Hiroshi G. Okuno |
Speech Commun. | 1 |
| 1999 | Listening to two simultaneous speeches
Hiroshi G. Okuno, Tomohiro Nakatani, Takeshi Kawabata |
Speech Commun. | 2 |
| 1997 | Understanding Three Simultaneous Speeches
Hiroshi G. Okuno, Tomohiro Nakatani, Takeshi Kawabata |
IJCAI (1) | 2 |
| 1996 | Localization by harmonic structure and its application to harmonic sound stream segregationabstractSound stream segregation is essential for understanding auditory events in the real-world. In this paper, we present a new method for sound stream segregation using harmonic structure and localization, or direction, in the horizontal plane. The direction of the sound source is determined by using the harmonic structure extracted from binaural inputs. The fundamental frequency of each sound is then refined by using the direction of its source. This paper discusses how the effectiveness of the harmonic-based stream segregation system (HBSS) is improved by incorporating the new method and presents the binaural HBSS (Bi-HBSS). In particular, experimental results show that the Bi-HBSS reduces the spectrum distortions and the fundamental frequency errors, compared with the HBSS and with a direction-based stream segregation system. Tomohiro Nakatani, Masataka Goto, Hiroshi G. Okuno |
ICASSP | 1 |
| 1996 | A new speech enhancement: speech stream segregationabstractSpeech stream segregation is presented as a new speech enhancement for automatic speech recognition.Two issues are addressed: speech stream segregation from a mixture of sounds, and interfacing speech stream segregation with automatic speech recognition.Speech stream segregation is modeled as a process of extracting harmonic fragments, grouping these extracted harmonic fragments, and substituting non-harmonic residue for non-harmonic parts of groups.The main problem in interfacing speech stream segregation with HMM-based speech recognition is how to improve the degradation of recognition performance due to spectral distortion of segregated sounds, which is caused mainly by transfer function of a binaural input.Our solution is to re-train the parameters of HMM with training data binauralized for four directions.Experiments with 500 mixtures of two women' s utterances of a word showed that the cumulative accuracy of word recognition up to the 10th candidate of each woman' s utterance is, on average, 75%. Hiroshi G. Okuno, Tomohiro Nakatani, Takeshi Kawabata |
ICSLP | 2 |
| 1995 | A computational model of sound stream segregation with multi-agent paradigmabstractThis paper presents a new computation model for sound stream segregation based on a multi-agent paradigm. Sound streams are thought to play a key role in auditory scene analysis, which provides a general framework for auditory research including voiced speech and music. Each agent is dynamically allocated to a sound stream, and it segregates the stream by focusing on consistent attributes. Agents interact with each other to resolve stream interference. We design agents to segregate harmonic streams and a noise stream. The presented system can segregate all the streams from a mixture of a male and a female voiced speech and a background non-harmonic noise. Tomohiro Nakatani, Takeshi Kawabata, Hiroshi G. Okuno |
ICASSP | 1 |
| 1995 | Residue-Driven Architecture for Computational Auditory Scene Analysis
Tomohiro Nakatani, Hiroshi G. Okuno, Takeshi Kawabata |
IJCAI | 1 |
| 1994 | Auditory Stream Segregation in Auditory Scene Analysis with a Multi-Agent System
Tomohiro Nakatani, Hiroshi G. Okuno, Takeshi Kawabata |
AAAI | 1 |
| 1994 | Unified architecture for auditory scene analysis and spoken language processing
Tomohiro Nakatani, Takeshi Kawabata, Hiroshi G. Okuno |
ICSLP | 1 |