VLDB 2026 Research / reviewers in the wild / expert
Shoko Araki
dblp:62/1595
· DBLP profile ↗
134ranked-venue papers
24as first author
34since 2021 · last 2026
0000-0003-4363-4305ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 104 · 22 first-author · 27 since 2021Artificial intelligence and machine learning · 47 · 5 first-author · 17 since 2021Systems, architecture and hardware · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Construction and Analysis of Japanese Parent-Child Dialogic Reading Corpus for Conversational Agents
Yuko Nakagi, Yuya Chiba, Sanae Fujita, Shoko Araki |
LREC | 4 |
| 2026 | Microphone array geometry-independent multi-talker distant ASR: NTT system for DASR task of the CHiME-8 challenge
Naoyuki Kamo, Naohiro Tawara, Atsushi Ando, Takatomo Kano, Hiroshi Sato 0002, Rintaro Ikeshita, Takafumi Moriya, Shota Horiguchi, Kohei Matsuura, Atsunori Ogawa, Alexis Plaquet, Takanori Ashihara, Tsubasa Ochiai, Masato Mimura, Marc Delcroix, Tomohiro Nakatani, Taichi Asami, Shoko Araki |
Comput. Speech Lang. | 18 |
| 2025 | 30+ Years of Source Separation Research: Achievements and Future ChallengesabstractSource separation (SS) of acoustic signals is a research field that emerged in the mid-1990s and has flourished ever since. On the occasion of ICASSP’s 50thanniversary, we review the major contributions and advancements in the past three decades in the speech, audio, and music SS research field. We will cover both single- and multi-channel SS approaches. We will also look back on key efforts to foster a culture of scientific evaluation in the research field, including challenges, performance metrics, and datasets. We will conclude by discussing current trends and future research directions. Shoko Araki, Nobutaka Ito, Reinhold Häb-Umbach, Gordon Wichern, Yuki Mitsufuji |
ICASSP | 1 |
| 2025 | SoundBeam meets M2D: Target Sound Extraction with Audio Foundation ModelabstractTarget sound extraction (TSE) consists of isolating a desired sound from a mixture of arbitrary sounds using clues to identify it. A TSE system requires solving two problems at once, identifying the target source and extracting the target signal from the mixture. For increased practicability, the same system should work with various types of sound. The duality of the problem and the wide variety of sounds make it challenging to train a powerful TSE system from scratch. In this paper, to tackle this problem, we explore using a pre-trained audio foundation model that can provide rich feature representations of sounds within a TSE system. We chose the masked-modeling duo (M2D) foundation model, which appears especially suited for the TSE task, as it is trained using a dual objective consisting of sound-label predictions and improved masked prediction. These objectives are related to sound identification and the signal extraction problems of TSE. We propose a new TSE system that integrates the feature representation from M2D into SoundBeam, which is a strong TSE system that can exploit both target sound class labels and pre-recorded enrollments (or audio queries) as clues. We show experimentally that using M2D can increase extraction performance, especially when employing enrollment clues. Carlos Hernandez-Olivan, Marc Delcroix, Tsubasa Ochiai, Daisuke Niizumi, Naohiro Tawara, Tomohiro Nakatani, Shoko Araki |
ICASSP | 7 |
| 2025 | A Hybrid Probabilistic-Deterministic Model Recursively Enhancing SpeechabstractThis paper introduces Probabilistic-Deterministic Recursive Enhancement (PDRE), an innovative iterative Speech Enhancement (SE) approach that integrates probabilistic and deterministic methodologies. Recent advancements in diffusion models have demonstrated the exceptional effectiveness of probabilistic model-based iterative estimation in SE, especially when combined with deterministic Neural Network (NN)-based methods. However, these models often require extensive iterations, leading to significant computational costs. To tackle this issue, we propose PDRE as a more efficient alternative. PDRE progressively refines the clean speech density estimates by recursively applying an Enhancement Network (EN), which is trained using a maximum likelihood objective. A single application of the EN can substantially improve the estimation, enabling PDRE to achieve high SE accuracy with significantly fewer iterations. Additionally, PDRE synergizes recursive enhancement with deterministic signal estimation, resulting in even greater accuracy. Our experiments demonstrate that PDRE significantly reduces iteration counts and computational costs compared to diffusion model-based SEs while maintaining or improving the remarkably high estimation accuracy. Tomohiro Nakatani, Naoyuki Kamo, Marc Delcroix, Shoko Araki |
ICASSP | 4 |
| 2025 | TS-SUPERB: A Target Speech Processing Benchmark for Speech Self-Supervised Learning ModelsabstractSelf-supervised learning (SSL) models have significantly advanced speech processing tasks, and several benchmarks have been proposed to validate their effectiveness. However, previous benchmarks have primarily focused on single-speaker scenarios, with less exploration of target-speaker tasks in noisy, multi-talker conditions—a more challenging yet practical case. In this paper, we introduce the Target-Speaker Speech Processing Universal Performance Benchmark (TS-SUPERB), which includes four widely recognized target-speaker processing tasks that require identifying the target speaker and extracting information from the speech mixture. In our benchmark, the speaker embedding extracted from enrollment speech is used as a clue to condition downstream models. The benchmark result reveals the importance of evaluating SSL models in target speaker scenarios, demonstrating that performance cannot be easily inferred from related single-speaker tasks. Moreover, by using a unified SSL-based target speech encoder, consisting of a speaker encoder and an extractor module, we also investigate joint optimization across TS tasks to leverage mutual information and demonstrate its effectiveness.1 Junyi Peng, Takanori Ashihara, Marc Delcroix, Tsubasa Ochiai, Oldrich Plchot, Shoko Araki, Jan Cernocký |
ICASSP | 6 |
| 2025 | Mamba-based Segmentation Model for Speaker DiarizationabstractMamba is a newly proposed architecture that behaves like a recurrent neural network (RNN) with attention-like capabilities. These properties are promising for speaker diarization, as attention-based models have unsuitable memory requirements for long-form audio, and traditional RNN capabilities are too limited. In this paper, we propose to assess the potential of Mamba for diarization by comparing the state-of-the-art neural segmentation of the pyannote pipeline with our proposed Mamba-based variant. Mamba’s stronger processing capabilities allow usage of longer local windows, which significantly improve diarization quality by making the speaker embedding extraction more reliable. We find Mamba to be a superior alternative to both traditional RNN and the tested attention-based model. Our proposed Mamba-based system achieves state-of-the-art performance on three widely used diarization datasets. Alexis Plaquet, Naohiro Tawara, Marc Delcroix, Shota Horiguchi, Atsushi Ando, Shoko Araki |
ICASSP | 6 |
| 2025 | Why is children's ASR so difficult? Analyzing children's phonological error patterns using SSL-based phoneme recognizers
Koharu Horii, Naohiro Tawara, Atsunori Ogawa, Shoko Araki |
INTERSPEECH | 4 |
| 2025 | Real-time TSE demonstration via SoundBeam with KD
Keigo Wakayama, Tomoko Kawase, Takafumi Moriya, Marc Delcroix, Hiroshi Sato 0002, Tsubasa Ochiai, Masahiro Yasuda, Shoko Araki |
INTERSPEECH | 8 |
| 2024 | How Does End-To-End Speech Recognition Training Impact Speech Enhancement Artifacts?abstractJointly training a speech enhancement (SE) front-end and an automatic speech recognition (ASR) back-end has been investigated as a way to mitigate the influence of processing distortion generated by single-channel SE on ASR. In this paper, we investigate the effect of such joint training on the signal-level characteristics of the enhanced signals from the viewpoint of the decomposed noise and artifact errors. The experimental analyses provide two novel findings: 1) ASR-level training of the SE front-end reduces the artifact errors while increasing the noise errors, and 2) simply interpolating the enhanced and observed signals, which achieves a similar effect of reducing artifacts and increasing noise, improves ASR performance without jointly modifying the SE and ASR modules, even for a strong ASR back-end using a WavLM feature extractor. Our findings provide a better understanding of the effect of joint training and a novel insight for designing an ASR agnostic SE front-end. Kazuma Iwamoto, Tsubasa Ochiai, Marc Delcroix, Rintaro Ikeshita, Hiroshi Sato 0002, Shoko Araki, Shigeru Katagiri |
ICASSP | 6 |
| 2024 | Target Speech Extraction with Pre-Trained Self-Supervised Learning ModelsabstractPre-trained self-supervised learning (SSL) models have achieved remarkable success in various speech tasks. However, their potential in target speech extraction (TSE) has not been fully exploited. TSE aims to extract the speech of a target speaker in a mixture guided by enrollment utterances. We exploit pre-trained SSL models for two purposes within a TSE framework, i.e., to process the input mixture and to derive speaker embeddings from the enrollment. In this paper, we focus on how to effectively use SSL models for TSE. We first introduce a novel TSE downstream task following the SUPERB principles. This simple experiment shows the potential of SSL models for TSE, but extraction performance remains far behind the state-of-the-art. We then extend a powerful TSE architecture by incorporating two SSL-based modules: an Adaptive Input Enhancer (AIE) and a speaker encoder. Specifically, the proposed AIE utilizes intermediate representations from the CNN encoder by adjusting the time resolution of CNN encoder and transformer blocks through progressive upsampling, capturing both fine-grained and hierarchical features. Our method outperforms current TSE systems achieving a SI-SDR improvement of 14.0 dB on LibriMix. Moreover, we can further improve performance by 0.7 dB by fine-tuning the whole model including the SSL model parameters. Junyi Peng, Marc Delcroix, Tsubasa Ochiai, Oldrich Plchot, Shoko Araki, Jan Cernocký |
ICASSP | 5 |
| 2024 | Neural Network-Based Virtual Microphone Estimation with Virtual Microphone and Beamformer-Level Multi-Task LossabstractArray processing performance depends on the number of microphones available. Virtual microphone estimation (VME) has been proposed to increase the number of microphone signals artificially. Neural network-based VME (NN-VME) trains an NN with a VM-level loss to predict a signal at a microphone location that is available during training but not at inference. However, this training objective may not be optimal for a specific array processing back-end, such as beamforming. An alternative approach is to use a training objective considering the array-processing back-end, such as a loss on the beamformer output. This approach may generate signals optimal for beamforming but not physically grounded. To combine the advantages of both approaches, this paper proposes a multi-task loss for NN-VME that combines both VM-level and beamformer-level losses. We evaluate the proposed multi-task NN-VME on multi-talker underdetermined conditions and show that it achieves a 33.1 % relative WER improvement compared to using only real microphones and 10.8 % compared to using a prior NN-VME approach. Hanako Segawa, Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Rintaro Ikeshita, Shoko Araki, Takeshi Yamada, Shoji Makino |
ICASSP | 6 |
| 2024 | Online Target Sound Extraction with Knowledge Distillation from Partially Non-Causal TeacherabstractTarget Sound Extraction (TSE) is a technique for extracting sound events belonging to a target sound class in a mixture using a Deep Neural Network (DNN). Offline TSE that uses non-causal models has achieved high extraction performance. However, many applications require online processing. Simply converting the non-causal TSE model architecture to a causal one leads to significant performance degradation. To mitigate this problem, we propose using Knowledge Distillation (KD) from a non-causal teacher to a causal student for TSE. In particular, we investigate different options for the non-causal teacher. We identify that a causal network with a non-causal layer normalization provides a strong teacher from which it is easier to transfer knowledge to the student. We conduct experiments with simulated sound mixtures and show that training a causal TSE with the proposed KD scheme can improve the signal-to-distortion ratio (SDR) by 0.9 dB compared to a baseline causal system. Keigo Wakayama, Tsubasa Ochiai, Marc Delcroix, Masahiro Yasuda, Shoichiro Saito, Shoko Araki, Akira Nakayama |
ICASSP | 6 |
| 2024 | Frontier of Frontend for Conversational Speech Processing
Shoko Araki |
INTERSPEECH | 1 |
| 2024 | Array Geometry-Robust Attention-Based Neural Beamformer for Moving Speakers
Marvin Tammen, Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Shoko Araki, Simon Doclo |
INTERSPEECH | 5 |
| 2024 | Rethinking Processing Distortions: Disentangling the Impact of Speech Enhancement Errors on Speech Recognition PerformanceabstractIt is challenging to improve automatic speech recognition (ASR) performance in noisy conditions with a single-channel speech enhancement (SE) front-end. This is generally attributed to the processing distortions caused by the nonlinear processing of single-channel SE front-ends. However, the causes of such degraded ASR performance have not been fully investigated. How to design single-channel SE front-ends in a way that significantly improves ASR performance remains an open research question. In this study, we investigate a signal-level numerical metric that can explain the cause of degradation in ASR performance. To this end, we propose a novel analysis scheme based on the orthogonal projection-based decomposition of SE errors. This scheme manually modifies the ratio of the decomposed interference, noise, and artifact errors, and it enables us to directly evaluate the impact of each error type on ASR performance. Our analysis reveals the particularly detrimental effect of artifact errors on ASR performance compared to the other types of errors. This provides us with a more principled definition of processing distortions that cause the ASR performance degradation. Then, we study two practical approaches for reducing the impact of artifact errors. First, we prove that the simple observation adding (OA) post-processing (i.e., interpolating the enhanced and observed signals) can improve the signal-to-artifact ratio. Second, we propose a novel training objective, called artifact-boosted signal-to-distortion ratio (AB-SDR), which forces the model to estimate the enhanced signals with fewer artifact errors. Through experiments, we confirm that both the OA and AB-SDR approaches are effective in decreasing artifact errors caused by single-channel SE front-ends, allowing them to significantly improve ASR performance. Tsubasa Ochiai, Kazuma Iwamoto, Marc Delcroix, Rintaro Ikeshita, Hiroshi Sato 0002, Shoko Araki, Shigeru Katagiri |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2024 | Blind and Spatially-Regularized Online Joint Optimization of Source Separation, Dereverberation, and Noise ReductionabstractThis paper proposes a computationally efficient joint optimization algorithm that performs online source separation, dereverberation, and noise reduction based on blind and spatially-regularized processing. When applying such online Blind Source Separation (BSS) as online Independent Vector Extraction (IVE) to a speech application, we must focus on the trade-off between the algorithmic delay and separation accuracy, both of which depend on the analysis frame length. In addition, to separate the sources with specified source permutation, researchers introduced spatial regularization based on the Directions-of-Arrival (DOAs) of the sources into IVE. However, the scale ambiguity of IVE often makes the spatial regularization work inappropriately. To solve these problems, we first propose a blind online joint optimization algorithm of IVE and weighted prediction error dereverberation (WPE). This online algorithm can achieve accurate separation even using short analysis frames because reverberation can be reduced using WPE. We then extend the online joint optimization with robust spatial regularization. We reveal that regularizing the scale of the separated signals is very effective in making the DOA-based spatial regularization work reliably. Our experiments confirm that our blind online joint optimization algorithm can significantly improve the separation accuracy with an algorithmic delay of 8 ms. In addition, we confirm that the proposed spatially-regularized online joint optimization algorithm reduces the rate of the source permutation error to zero percent. Tetsuya Ueda, Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Shoko Araki, Shoji Makino |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Fast Online Source Steering Algorithm for Tracking Single Moving Source Using Online Independent Vector AnalysisabstractWe address the problem of separating moving sources using online independent vector analysis (IVA). To solve this problem, researchers have extended the iterative projection (IP) and iterative source steering (ISS) algorithms developed for batch auxiliary-function-based IVA (AuxIVA) to online scenarios and showed their effectiveness. However, the conventional online IP and ISS are slow because they update K × K covariance matrices for all sources, where K is the number of microphones. Here, we show that, in a target-source tracking scenario in which only one source moves, there exists an inexpensive formula for online ISS that avoids updating the full covariance matrices without changing the behavior of the algorithm. The time complexity of the proposed algorithm, which we call online source steering (OSS), is K times smaller than that of the conventional online IP and ISS for the target-source tracking task. A numerical experiment on separating a moving source demonstrates that the proposed OSS is significantly faster than the conventional online IP and ISS. Taishi Nakashima, Rintaro Ikeshita, Nobutaka Ono, Shoko Araki, Tomohiro Nakatani |
ICASSP | 4 |
| 2023 | Impact of Residual Noise and Artifacts in Speech Enhancement Errors on Intelligibility of Human and Machine
Shoko Araki, Ayako Yamamoto, Tsubasa Ochiai, Kenichi Arai, Atsunori Ogawa, Tomohiro Nakatani, Toshio Irino |
INTERSPEECH | 1 |
| 2023 | Multi-Stream Extension of Variational Bayesian HMM Clustering (MS-VBx) for Combined End-to-End and Vector Clustering-based Diarization
Marc Delcroix, Naohiro Tawara, Mireia Díez, Federico Landini, Anna Silnova, Atsunori Ogawa, Tomohiro Nakatani, Lukás Burget, Shoko Araki |
INTERSPEECH | 9 |
| 2023 | SoundBeam: Target Sound Extraction Conditioned on Sound-Class Labels and Enrollment Clues for Increased Performance and Continuous LearningabstractIn many situations, we would like to hear desired sound events (SEs) while being able to ignore interference. Target sound extraction (TSE) tackles this problem by estimating the audio signal of the sounds of target SE classes in a mixture of sounds while suppressing all other sounds. We can achieve this with a neural network that extracts the target SEs by conditioning it on clues representing the target SE classes. Two types of clues have been proposed, i.e., targetSE class labelsandenrollment audio samples(or audio queries), which are pre-recorded audio samples of sounds from the target SE classes. Systems based on SE class labels can directly optimize embedding vectors representing the SE classes, resulting in high extraction performance. However, extending these systems to extract new SE classes not encountered during training is not easy. Enrollment-based approaches extract SEs by finding sounds in the mixtures that share similar characteristics to the enrollment audio samples. These approaches do not explicitly rely on SE class definitions and can thus handle new SE classes. In this paper, we introduce a TSE framework, SoundBeam, that combines the advantages of both approaches. We also perform an extensive evaluation of the different TSE schemes using synthesized and real mixtures, which shows the potential of SoundBeam. Marc Delcroix, Jorge Bennasar Vázquez, Tsubasa Ochiai, Keisuke Kinoshita, Yasunori Ohishi, Shoko Araki |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Mask-Based Neural Beamforming for Moving Speakers With Self-Attention-Based TrackingabstractBeamforming is a powerful tool designed to enhance speech signals from the direction of a target source. Computing the beamforming filter requires estimating spatial covariance matrices (SCMs) of the source and noise signals. Time-frequency masks are often used to compute these SCMs. Most studies of mask-based beamforming have assumed that the sources do not move. However, sources often move in practice, which causes performance degradation. In this paper, we address the problem of mask-based beamforming for moving sources. We first review classical approaches to tracking a moving source, which perform online or blockwise computation of the SCMs. We show that these approaches can be interpreted as computing a sum of instantaneous SCMs weighted by attention weights. These weights indicate which time frames of the signal to consider in the SCM computation. Online or blockwise computation assumes a heuristic and deterministic way of computing these attention weights that, although simple, may not result in optimal performance. We thus introduce a learning-based framework that computes optimal attention weights for beamforming. We achieve this using a neural network implemented with self-attention layers. We show experimentally that our proposed framework can greatly improve beamforming performance in moving source situations while maintaining high performance in non-moving situations, thus enabling the development of mask-based beamformers robust to source movements. Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Shoko Araki |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | Lattice Rescoring Based on Large Ensemble of Complementary Neural Language ModelsabstractWe investigate the effectiveness of using a large ensemble of advanced neural language models (NLMs) for lattice rescoring on automatic speech recognition (ASR) hypotheses. Previous studies have reported the effectiveness of combining a small number of NLMs. In contrast, in this study, we combine up to eight NLMs, i.e., forward/backward long short-term memory/Transformer-LMs that are trained with two different random initialization seeds. We combine these NLMs through iterative lattice generation. Since these NLMs work complementarily with each other, by combining them one by one at each rescoring iteration, language scores attached to given lattice arcs can be gradually refined. Consequently, errors of the ASR hypotheses can be gradually reduced. We also investigate the effectiveness of carrying over contextual information (previous rescoring results) across a lattice sequence of a long speech such as a lecture speech. In experiments using a lecture speech corpus, by combining the eight NLMs and using context carry-over, we obtained a 24.4% relative word error rate reduction from the ASR 1-best baseline. For further comparison, we performed simultaneous (i.e., non-iterative) NLM combination and 100-best rescoring using the large ensemble of NLMs, which confirmed the advantage of lattice rescoring with iterative NLM combination. Atsunori Ogawa, Naohiro Tawara, Marc Delcroix, Shoko Araki |
ICASSP | 4 |
| 2022 | How bad are artifacts?: Analyzing the impact of speech enhancement errors on ASRabstractIt is challenging to improve automatic speech recognition (ASR) performance in noisy conditions with single-channel speech enhancement (SE).In this paper, we investigate the causes of ASR performance degradation by decomposing the SE errors using orthogonal projection-based decomposition (OPD).OPD decomposes the SE errors into noise and artifact components.The artifact component is defined as the SE error signal that cannot be represented as a linear combination of speech and noise sources.We propose manually scaling the error components to analyze their impact on ASR.We experimentally identify the artifact component as the main cause of performance degradation, and we find that mitigating the artifact can greatly improve ASR performance.Furthermore, we demonstrate that the simple observation adding (OA) technique (i.e., adding a scaled version of the observed signal to the enhanced speech) can monotonically increase the signal-to-artifact ratio under a mild condition.Accordingly, we experimentally confirm that OA improves ASR performance for both simulated and real recordings.The findings of this paper provide a better understanding of the influence of SE errors on ASR and open the door to future research on novel approaches for designing effective single-channel SE front-ends for ASR. Kazuma Iwamoto, Tsubasa Ochiai, Marc Delcroix, Rintaro Ikeshita, Hiroshi Sato 0002, Shoko Araki, Shigeru Katagiri |
INTERSPEECH | 6 |
| 2022 | ConceptBeam: Concept Driven Target Speech ExtractionabstractWe propose a novel framework for target speech extraction based on semantic information, called ConceptBeam. Target speech extraction means extracting the speech of a target speaker in a mixture. Typical approaches have been exploiting properties of audio signals, such as harmonic structure and direction of arrival. In contrast, ConceptBeam tackles the problem with semantic clues. Specifically, we extract the speech of speakers speaking about a concept, i.e., a topic of interest, using a concept specifier such as an image or speech. Solving this novel problem would open the door to innovative applications such as listening systems that focus on a particular topic discussed in a conversation. Unlike keywords, concepts are abstract notions, making it challenging to directly represent a target concept. In our scheme, a concept is encoded as a semantic embedding by mapping the concept specifier to a shared embedding space. This modality-independent space can be built by means of deep metric learning using paired data consisting of images and their spoken captions. We use it to bridge modality-dependent information, i.e., the speech segments in the mixture, and the specified, modality-independent concept. As a proof of our scheme, we performed experiments using a set of images associated with spoken captions. That is, we generated speech mixtures from these spoken captions and used the images or speech signals as the concept specifiers. We then extracted the target speech using the acoustic characteristics of the identified segments. We compare ConceptBeam with two methods: one based on keywords obtained from recognition systems and another based on sound source separation. We show that ConceptBeam clearly outperforms the baseline methods and effectively extracts speech based on the semantic representation. Yasunori Ohishi, Marc Delcroix, Tsubasa Ochiai, Shoko Araki, Daiki Takeuchi, Daisuke Niizumi, Akisato Kimura, Noboru Harada, Kunio Kashino |
ACM Multimedia | 4 |
| 2022 | Switching Independent Vector Analysis and its Extension to Blind and Spatially Guided Convolutional Beamforming AlgorithmsabstractThis paper develops a framework that can accurately perform denoising, dereverberation, and source separation using a relatively small number of microphones. It has been empirically confirmed that Independent Vector Analysis (IVA) can blindly separate$N$sources from their sound mixture even with diffuse noise when a sufficiently large number ($=M$) of microphones are available (i.e.,$M\gg N)$. However, the estimation accuracy is seriously degraded when the number of microphones, or more specifically$M-N$$(\geq 0)$, decreases. To overcome this IVA limitation, we propose switching IVA (swIVA) in this paper. With swIVA, the time frames of an observed signal with time-varying characteristics are clustered into several groups, each of which can be well handled by IVA with a small number of microphones, and thus accurate estimation can be achieved by individually applying IVA to each group. Conventionally, a switching mechanism was introduced into a Minimum-Variance Distortionless Response (MVDR) beamformer, and this paper extends the mechanism to work with a blind source separation algorithm. To incorporate dereverberation capability, we further extend swIVA to a blind Convolutional beamforming algorithm (swCIVA) that integrates swIVA and switching Weighted Prediction Error-based dereverberation (swWPE) in a jointly optimal way. With swCIVA, two different time-varying characteristics of an observed signal are captured for dereverberation and source separation to achieve effective estimation. We show that both swIVA and swCIVA can be optimized effectively based on blind signal processing, and their performance can be further improved using a spatial guide for initialization. Experiments demonstrate that both the proposed methods largely outperformed conventional IVA and its convolutional beamforming extension (CIVA) in terms of objective signal quality and automatic speech recognition scores when using relatively few microphones. Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Hiroshi Sawada, Naoyuki Kamo, Shoko Araki |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2021 | Blind and Neural Network-Guided Convolutional Beamformer for Joint Denoising, Dereverberation, and Source SeparationabstractThis paper proposes an approach for optimizing a Convolutional BeamFormer (CBF) that can jointly perform denoising (DN), dereverberation (DR), and source separation (SS). First, we develop a blind CBF optimization algorithm that requires no prior information on the sources or the room acoustics, by extending a conventional joint DR and SS method. For making the optimization computationally tractable, we incorporate two techniques into the approach: the Source-Wise Factorization (SW-Fact) of a CBF and the Independent Vector Extraction (IVE). To further improve the performance, we develop a method that integrates a neural network (NN) based source power spectra estimation with CBF optimization by an inverse-Gamma prior. Experiments using noisy reverberant mixtures reveal that our proposed method with both blind and NN-guided scenarios greatly outperforms the conventional state-of-the-art NN-supported mask-based CBF in terms of the improvement in automatic speech recognition and signal distortion reduction performance. Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Hiroshi Sawada, Shoko Araki |
ICASSP | 5 |
| 2021 | Neural Network-Based Virtual Microphone EstimatorabstractDeveloping microphone array technologies for a small number of microphones is important due to the constraints of many devices. One direction to address this situation consists of virtually augmenting the number of microphone signals, e.g., based on several physical model assumptions. However, such assumptions are not necessarily met in realistic conditions. In this paper, as an alternative approach, we propose a neural network-based virtual microphone estimator (NN-VME). The NN-VME estimates virtual microphone signals directly in the time domain, by utilizing the precise estimation capability of the recent time-domain neural networks. We adopt a fully supervised learning framework that uses actual observations at the locations of the virtual microphones at training time. Consequently, the NN-VME can be trained using only multi-channel observations and thus directly on real recordings, avoiding the need for unrealistic physical model-based assumptions. Experiments on the CHiME-4 corpus show that the proposed NN-VME achieves high virtual microphone estimation performance even for real recordings and that a beamformer augmented with the NN-VME improves both the speech enhancement and recognition performance. Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Shoko Araki |
ICASSP | 6 |
| 2021 | Low Latency Online Blind Source Separation Based on Joint Optimization with Blind DereverberationabstractThis paper presents a new low-latency online blind source separation (BSS) algorithm. Although algorithmic delay of a frequency domain online BSS can be reduced simply by shortening the short-time Fourier transform (STFT) frame length, it degrades the source separation performance in the presence of reverberation. This paper proposes a method to solve this problem by integrating BSS with Weighted Prediction Error (WPE) based dereverberation. Although a simple cascade of online BSS after online WPE upgrades the separation performance, the overall optimality is not guaranteed. Instead, this paper extends a recently proposed batch processing algorithm that can jointly optimize dereverberation and separation so that it can perform online processing with low computational cost and little processing delay (< 12 ms). The results of a source separation experiment in a noisy car environment suggest that the proposed online method has better separation performance than the simple cascaded methods. Tetsuya Ueda, Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Shoko Araki, Shoji Makino |
ICASSP | 5 |
| 2021 | Data Fusion for Audiovisual Speaker Localization: Extending Dynamic Stream Weights to the Spatial DomainabstractEstimating the positions of multiple speakers can be helpful for tasks like automatic speech recognition or speaker diarization. Both applications benefit from a known speaker position when, for instance, applying beamforming or assigning unique speaker identities. Recently, several approaches utilizing acoustic signals augmented with visual data have been proposed for this task. However, both the acoustic and the visual modality may be corrupted in specific spatial regions, for instance due to poor lighting conditions or to the presence of background noise. This paper proposes a novel audiovisual data fusion framework for speaker localization by assigning individual dynamic stream weights to specific regions in the localization space. This fusion is achieved via a neural network, which combines the predictions of individual audio and video trackers based on their time- and location-dependent reliability. A performance evaluation using audiovisual recordings yields promising results, with the proposed fusion approach outperforming all baseline models. Julio Wissing, Benedikt T. Boenninghoff, Dorothea Kolossa, Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki, Christopher Schymura |
ICASSP | 8 |
| 2021 | Few-Shot Learning of New Sound Classes for Target Sound ExtractionabstractTarget sound extraction consists of extracting the sound of a target acoustic event (AE) class from a mixture of AE sounds. It can be realized using a neural network that extracts the target sound conditioned on a 1-hot vector that represents the desired AE class. With this approach, embedding vectors associated with the AE classes are directly optimized for the extraction of sound classes seen during training. However, it is not easy to extend this framework to new AE classes, i.e. unseen during training. Recently, speech, music, or AE sound extraction based on enrollment audio of the desired sound offers the potential of extracting any target sound in a mixture given only a short audio signal of a similar sound. In this work, we propose combining 1-hot- and enrollment-based target sound extraction, allowing optimal performance for seen AE classes and simple extension to new classes. In experiments with synthesized sound mixtures generated with the Freesound Dataset (FSD) datasets, we demonstrate the benefit of the combined framework for both seen and new AE classes. Besides, we also propose adapting the embedding vectors obtained from a few enrollment audio samples (few-shot) to further improve performance on new classes. Marc Delcroix, Jorge Bennasar Vázquez, Tsubasa Ochiai, Keisuke Kinoshita, Shoko Araki |
Interspeech | 5 |
| 2021 | PILOT: Introducing Transformers for Probabilistic Sound Event LocalizationabstractSound event localization aims at estimating the positions of sound sources in the environment with respect to an acoustic receiver (e.g. a microphone array). Recent advances in this domain most prominently focused on utilizing deep recurrent neural networks. Inspired by the success of transformer architectures as a suitable alternative to classical recurrent neural networks, this paper introduces a novel transformer-based sound event localization framework, where temporal dependencies in the received multi-channel audio signals are captured via self-attention mechanisms. Additionally, the estimated sound event positions are represented as multivariate Gaussian variables, yielding an additional notion of uncertainty, which many previously proposed deep learning-based systems designed for this application do not provide. The framework is evaluated on three publicly available multi-source sound event localization datasets and compared against state-of-the-art methods in terms of localization error and event detection accuracy. It outperforms all competing systems on all datasets with statistical significant differences in performance. Christopher Schymura, Benedikt T. Boenninghoff, Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki, Dorothea Kolossa |
Interspeech | 7 |
| 2021 | Comparison of Remote Experiments Using Crowdsourcing and Laboratory Experiments on Speech IntelligibilityabstractMany subjective experiments have been performed to develop objective speech intelligibility measures, but the novel coronavirus outbreak has made it very difficult to conduct experiments in a laboratory. One solution is to perform remote testing using crowdsourcing; however, because we cannot control the listening conditions, it is unclear whether the results are entirely reliable. In this study, we compared speech intelligibility scores obtained in remote and laboratory experiments. The results showed that the mean and standard deviation (SD) of the remote experiments' speech reception threshold (SRT) were higher than those of the laboratory experiments. However, the variance in the SRTs across the speech-enhancement conditions revealed similarities, implying that remote testing results may be as useful as laboratory experiments to develop an objective measure. We also show that the practice session scores correlate with the SRT values. This is a priori information before performing the main tests and would be useful for data screening to reduce the variability of the SRT distribution. Ayako Yamamoto, Toshio Irino, Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani |
Interspeech | 4 |
| 2021 | Multimodal Attention Fusion for Target Speaker ExtractionabstractTarget speaker extraction, which aims at extracting a target speaker's voice from a mixture of voices using audio, visual or locational clues, has received much interest. Recently an audio-visual target speaker extraction has been proposed that extracts target speech by using complementary audio and visual clues. Although audio-visual target speaker extraction offers a more stable performance than single modality methods for simulated data, its adaptation towards realistic situations has not been fully explored as well as evaluations on real recorded mixtures. One of the major issues to handle realistic situations is how to make the system robust to clue corruption because in real recordings both clues may not be equally reliable, e.g. visual clues may be affected by occlusions. In this work, we propose a novel attention mechanism for multi-modal fusion and its training methods that enable to effectively capture the reliability of the clues and weight the more reliable ones. Our proposals improve signal to distortion ratio (SDR) by 1.0 dB over conventional fusion mechanisms on simulated data. Moreover, we also record an audio-visual dataset of simultaneous speech with realistic visual clue corruption and show that audio-visual target speaker extraction with our proposals successfully work on real data. Hiroshi Sato 0002, Tsubasa Ochiai, Keisuke Kinoshita, Marc Delcroix, Tomohiro Nakatani, Shoko Araki |
SLT | 6 |
| 2020 | Improving Speaker Discrimination of Target Speech Extraction With Time-Domain SpeakerbeamabstractTarget speech extraction, which extracts a single target source in a mixture given clues about the target speaker, has attracted increasing attention. We have recently proposed SpeakerBeam, which exploits an adaptation utterance of the target speaker to extract his/her voice characteristics that are then used to guide a neural network towards extracting speech of that speaker. SpeakerBeam presents a practical alternative to speech separation as it enables tracking speech of a target speaker across utterances, and achieves promising speech extraction performance. However, it sometimes fails when speakers have similar voice characteristics, such as in same-gender mixtures, because it is difficult to discriminate the target speaker from the interfering speakers. In this paper, we investigate strategies for improving the speaker discrimination capability of SpeakerBeam. First, we propose a time-domain implementation of SpeakerBeam similar to that proposed for a time-domain audio separation network (TasNet), which has achieved state-of-the-art performance for speech separation. Besides, we investigate (1) the use of spatial features to better discriminate speakers when microphone array recordings are available, (2) adding an auxiliary speaker identification loss for helping to learn more discriminative voice characteristics. We show experimentally that these strategies greatly improve speech extraction performance, especially for same-gender mixtures, and outperform TasNet in terms of target speech extraction. Marc Delcroix, Tsubasa Ochiai, Katerina Zmolíková, Keisuke Kinoshita, Naohiro Tawara, Tomohiro Nakatani, Shoko Araki |
ICASSP | 7 |
| 2020 | A Frequency-Domain BSS Method Based on ℓ1 Norm, Unitary Constraint, and Cayley TransformabstractWe propose a frequency-domain blind source separation method that uses (a) the ℓ1norm of orthonormal vectors of estimated source signals as a sparsity measure and (b) Cayley transform for optimizing the objective function under the unitary constraint in the Riemannian geometry approach. The orthonormal vectors of estimated source signals, obtained by the sphering of observed mixed signals and the unitary constraint on the separation filters, enables us to use the ℓ1norm properly as a sparsity measure. The Cayley transform enables us to handle the geometrical aspects of the unitary constraint efficiently. According to the simulation of a two-channel case, the proposed method achieved a 20-dB improvement in the source-to-interference ratio in a room with a reverberation time of T60= 300ms. Satoru Emura, Hiroshi Sawada, Shoko Araki, Noboru Harada |
ICASSP | 3 |
| 2020 | Overdetermined Independent Vector AnalysisabstractWe address the convolutive blind source separation problem for the (over-)determined case where (i) the number of nonstationary target-sources K is less than that of microphones M, and (ii) there are up to M - K stationary Gaussian noises that need not to be extracted. Independent vector analysis (IVA) can solve the problem by separating into M sources and selecting the top K highly nonstationary signals among them, but this approach suffers from a waste of computation especially when K ≪ M. Channel reductions in preprocessing of IVA by, e.g., principle component analysis have the risk of removing the target signals. We here extend IVA to resolve these issues. One such extension has been attained by assuming the orthogonality constraint (OC) that the sample correlation between the target and noise signals is to be zero. The proposed IVA, on the other hand, does not rely on OC and exploits only the independence between sources and the stationarity of the noises. This enables us to develop several efficient algorithms based on block coordinate descent methods with a problem specific acceleration. We clarify that one such algorithm exactly coincides with the conventional IVA with OC, and also explain that the other newly developed algorithms are faster than it. Experimental results show the improved computational load of the new algorithms compared to the conventional methods. In particular, a new algorithm specialized for K = 1 outperforms the others. Rintaro Ikeshita, Tomohiro Nakatani, Shoko Araki |
ICASSP | 3 |
| 2020 | Tackling Real Noisy Reverberant Meetings with All-Neural Source Separation, Counting, and Diarization SystemabstractAutomatic meeting analysis is an essential fundamental technology required to let, e.g. smart devices follow and respond to our conversations. To achieve an optimal automatic meeting analysis, we previously proposed an all-neural approach that jointly solves source separation, speaker diarization and source counting problems in an optimal way (in a sense that all the 3 tasks can be jointly optimized through error back-propagation). It was shown that the method could well handle simulated clean (noiseless and anechoic) dialog-like data, and achieved very good performance in comparison with several conventional methods. However, it was not clear whether such all-neural approach would be successfully generalized to more complicated real meeting data containing more spontaneously-speaking speakers, severe noise and reverberation, and how it performs in comparison with the state-of-the-art systems in such scenarios. In this paper, we first consider practical issues required for improving the robustness of the all-neural approach, and then experimentally show that, even in real meeting scenarios, the all-neural approach can perform effective speech enhancement, and simultaneously outperform state-of-the-art systems. Keisuke Kinoshita, Marc Delcroix, Shoko Araki, Tomohiro Nakatani |
ICASSP | 3 |
| 2020 | DNN-supported Mask-based Convolutional Beamforming for Simultaneous Denoising, Dereverberation, and Source SeparationabstractIn this article, we investigate an integrated mask-based convolutional beamforming method for performing simultaneous denoising, dereverberation, and source separation. Conventionally, it is difficult for neural network-supported mask-based source separation to perform denoising and dereverberation at the same time and for spatial clustering-based source separation to reliably solve the permutation problem in the presence of noise and reverberation. This greatly limits the application of mask-based source separation. To address this issue, we propose a method to integrate state-of-the-art techniques for mask-based beamforming into a single optimization framework. These techniques include frequency-domain Convolutional Neural Network based utterance-level Permutation Invariant Training with a large receptive field (CNN-uPIT), noisy Complex Gaussian Mixture Model based spatial clustering (noisyCGMM), and Weighted Power minimization Distortionless response (WPD) convolutional beamforming. Our experiments show that all these components are essential for accurately estimating desired speech signals in noisy reverberant multisource environments. Tomohiro Nakatani, Riki Takahashi, Tsubasa Ochiai, Keisuke Kinoshita, Rintaro Ikeshita, Marc Delcroix, Shoko Araki |
ICASSP | 7 |
| 2020 | Beam-TasNet: Time-domain Audio Separation Network Meets Frequency-domain BeamformerabstractRecent studies have shown that acoustic beamforming using a microphone array plays an important role in the construction of high-performance automatic speech recognition (ASR) systems, especially for noisy and overlapping speech conditions. In parallel with the success of multichannel beamforming for ASR, in the speech separation field, the time-domain audio separation network (TasNet), which accepts a time-domain mixture as input and directly estimates the time-domain waveforms for each source, achieves remarkable speech separation performance. In light of these two recent trends, the question of whether TasNet can benefit from beamforming to achieve high ASR performance in overlapping speech conditions naturally arises. Motivated by this question, this paper proposes a novel speech separation scheme, i.e., Beam-TasNet, which combines TasNet with the frequency-domain beamformer, i.e., a minimum variance distortionless response (MVDR) beamformer, through spatial covariance computation to achieve better ASR performance. Experiments on the spatialized WSJ0-2mix corpus show that our proposed Beam-TasNet significantly outperforms the conventional TasNet without beamforming and, moreover, successfully achieves a word error rate comparable to an oracle mask-based MVDR beamformer. Tsubasa Ochiai, Marc Delcroix, Rintaro Ikeshita, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki |
ICASSP | 6 |
| 2020 | A Dynamic Stream Weight Backprop Kalman Filter for Audiovisual Speaker TrackingabstractAudiovisual speaker tracking is an application that has been tackled by a wide range of classical approaches based on Gaussian filters, most notably the well-known Kalman filter. Recently, a specific Kalman filter implementation was proposed for this task, which incorporated dynamic stream weights to explicitly control the influence of acoustic and visual observations during estimation. Inspired by recent progress in the context of integrating uncertainty estimates into modern deep learning frameworks, this paper proposes a deep neural-network-based implementation of the Kalman filter with dynamic stream weights, whose parameters can be learned via standard backpropagation. This allows for jointly optimizing the parameters of the model and the dynamic stream weight estimator in a unified framework. An experimental study on audiovisual speaker tracking shows that the proposed model shows comparable performance to state-of-the-art recurrent neural networks with the additional advantage of requiring a smaller number of parameters and providing explicit uncertainty information. Christopher Schymura, Tsubasa Ochiai, Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki, Dorothea Kolossa |
ICASSP | 6 |
| 2020 | Predicting Intelligibility of Enhanced Speech Using Posteriors Derived from DNN-Based ASR System
Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Toshio Irino |
INTERSPEECH | 2 |
| 2020 | Computationally Efficient and Versatile Framework for Joint Optimization of Blind Speech Separation and Dereverberation
Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Hiroshi Sawada, Shoko Araki |
INTERSPEECH | 5 |
| 2020 | Listen to What You Want: Neural Network-Based Universal Sound SelectorabstractBeing able to control the acoustic events (AEs) to which we want to listen would allow the development of more controllable hearable devices.This paper addresses the AE sound selection (or removal) problems, that we define as the extraction (or suppression) of all the sounds that belong to one or multiple desired AE classes.Although this problem could be addressed with a combination of source separation followed by AE classification, this is a sub-optimal way of solving the problem.Moreover, source separation usually requires knowing the maximum number of sources, which may not be practical when dealing with AEs.In this paper, we propose instead a universal sound selection neural network that enables to directly select AE sounds from a mixture given user-specified target AE classes.The proposed framework can be explicitly optimized to simultaneously select sounds from multiple desired AE classes, independently of the number of sources in the mixture.We experimentally show that the proposed method achieves promising AE sound selection performance and could be generalized to mixtures with a number of sources that are unseen during training. Tsubasa Ochiai, Marc Delcroix, Yuma Koizumi, Hiroaki Ito, Keisuke Kinoshita, Shoko Araki |
INTERSPEECH | 6 |
| 2020 | GEDI: Gammachirp envelope distortion index for predicting intelligibility of enhanced speech
Katsuhiko Yamamoto, Toshio Irino, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani |
Speech Commun. | 3 |
| 2020 | Multi-Delay Sparse Approach to Residual Crosstalk Reduction for Blind Source SeparationabstractFor reducing residual crosstalk in the output of blind source separation, we propose a frequency-domain post-filtering method that uses a multi-delay model of complex-valued residual crosstalk and sparsifies the estimates of the source signals. We formulate the reduction of residual crosstalk as an optimization problem using ℓ1norm and solve it using the alternating direction method of multiplier. The proposed method improved the source-to-interference ratio from 17.8 to 20.5 dB and the source-to-distortion ratio from 10.2 to 11.3 dB when it was combined with a brute-force solver of FastICA and reverberation time T60was 300 ms. Satoru Emura, Hiroshi Sawada, Shoko Araki, Noboru Harada |
IEEE Signal Process. Lett. | 3 |
| 2019 | Estimation of Sampling Frequency Mismatch between Distributed Asynchronous Microphones under Existence of Source Movements with Stationary Time Periods DetectionabstractIn this paper, we propose a method of estimating the sampling frequency mismatch among asynchronous recording devices, even when the sources sometimes move. For a spatially stationary source, there is a method of estimating the sampling frequency mismatch, which appears in the drift of the time difference among the observed digitized signals. When the source moves, however, the change of its location also affects the drift, and the method fails to estimate the mismatch. In the meantime, looking at the practical recording situations, sources sometimes move but sometimes do not move. That is, there should be a set of time frames in which we can assume the spatial stationarity of sources, and to which we are still able to apply the sampling frequency mismatch estimation method. Based on this idea, our proposed method first detects a set of time frames where we can assume the spatial stationary by clustering the time frames using the covariance matrix of each recording device, and then estimates the mismatch by using the detected stationary time frames. Using real recordings with several IC recorders, we show that the proposed method can estimate the sampling frequency mismatch accurately even when the sources sometimes move. Shoko Araki, Nobutaka Ono, Keisuke Kinoshita, Marc Delcroix |
ICASSP | 1 |
| 2019 | Compact Network for Speakerbeam Target Speaker ExtractionabstractSpeech separation that separates a mixture of speech signals into each of its sources has been an active research topic for a long time and has seen recent progress with the advent of deep learning. A related problem is target speaker extraction, i.e. extraction of only speech of a target speaker out of a mixture, given characteristics of his/her voice. We have recently proposed SpeakerBeam, which is a neural network-based target speaker extraction method. Speaker-Beam uses a speech extraction network that is adapted to the target speaker using auxiliary features derived from an adaptation utterance of that speaker. Initially, we implemented SpeakerBeam with a factorized adaptation layer, which consists of several parallel linear transformations weighted by weights derived from the auxiliary features. The factorized layer is effective for target speech extraction, but it requires a large number of parameters. In this paper, we propose to simply scale the activations of a hidden layer of the speech extraction network with weights derived from the auxiliary features. This simpler approach greatly reduces the number of model parameters by up to 60%, making it much more practical, while maintaining a similar level of performance. We tested our approach on simulated and real noisy and reverberant mixtures, showing the potential of SpeakerBeam for real-life applications. Moreover, we showed that speech extraction performance of SpeakerBeam compares favorably with that of a state-of-the-art speech separation method with a similar network configuration. Marc Delcroix, Katerina Zmolíková, Tsubasa Ochiai, Keisuke Kinoshita, Shoko Araki, Tomohiro Nakatani |
ICASSP | 5 |
| 2019 | Mask-based MVDR Beamformer for Noisy Multisource Environments: Introduction of Time-varying Spatial Covariance ModelabstractThis paper proposes a method for designing a time-varying minimum variance distortionless response (MVDR) beamformer using time-frequency masks, with the aim of improving speech enhancement in noisy multi-speaker environments. A key to successful beamforming is to estimate accurately a time-varying spatial covariance matrix (SCM) for noise composed of both stationary diffuse noise and highly time-varying speech. For this purpose, we introduce a stochastic model that can represent the time-varying characteristics of the noise SCM, and derive a method for estimating a time-varying noise SCM based on the model. Experiments show that the proposed method can substantially improve the performance of the beamformer in terms of automatic speech recognition (ASR) accuracy and source-to-distortion ratio compared with a conventional time-invariant MVDR beamformer. Yuki Kubo, Tomohiro Nakatani, Marc Delcroix, Keisuke Kinoshita, Shoko Araki |
ICASSP | 5 |
| 2019 | All-neural Online Source Separation, Counting, and Diarization for Meeting AnalysisabstractAutomatic meeting analysis comprises the tasks of speaker counting, speaker diarization, and the separation of overlapped speech, followed by automatic speech recognition. This all has to be carried out on arbitrarily long sessions and, ideally, in an online or block-online manner. While significant progress has been made on individual tasks, this paper presents for the first time an all-neural approach to simultaneous speaker counting, diarization and source separation. The NN-based estimator operates in a block-online fashion and tracks speakers even if they remain silent for a number of time blocks, thus learning a stable output order for the separated sources. The neural network is recurrent over time as well as over the number of sources. The simulation experiments show that state of the art separation performance is achieved, while at the same time delivering good diarization and source counting results. It even generalizes well to an unseen large number of blocks. Thilo von Neumann, Keisuke Kinoshita, Marc Delcroix, Shoko Araki, Tomohiro Nakatani, Reinhold Häb-Umbach |
ICASSP | 4 |
| 2019 | Predicting Speech Intelligibility of Enhanced Speech Using Phone Accuracy of DNN-Based ASR System
Kenichi Arai, Shoko Araki, Atsunori Ogawa, Keisuke Kinoshita, Tomohiro Nakatani, Katsuhiko Yamamoto, Toshio Irino |
INTERSPEECH | 2 |
| 2018 | Meeting Recognition with Asynchronous Distributed Microphone Array Using Block-Wise Refinement of Mask-Based MVDR BeamformerabstractThis paper addresses a front-end system for speech recognition of spontaneous conversational speech signals that are recorded with asynchronous distributed microphones such as smartphones. In our previous work, we proposed combining blind synchronization and a state-of-the-art microphone array speech enhancement technique, e.g., a time-frequency mask based minimum variance distortionless response (MVDR) beamformer. This approach has provided reasonably high recognition performance even if we use asynchronous microphones. However, because the previous speech enhancement method was applied in a full-batch mode, it has been difficult to track speaker position movement in a real meeting conversation. To make it possible to handle the speaker movement, this paper describes our attempt to refine the mask-based MVDR beamformer in a blockwise manner, and reports that such a refinement reduces the word error rate from 31.4% to 28.8% for real meeting recordings. Shoko Araki, Nobutaka Ono, Keisuke Kinoshita, Marc Delcroix |
ICASSP | 1 |
| 2018 | Permutation-Free Cgmm: Complex Gaussian Mixture Model with Inverse Wishart Mixture Model Based Spatial Prior for Permutation-Free Source Separation and Source CountingabstractHere we propose a permutation-free cGMM (PF-cGMM), a new probabilistic model of observed mixtures, which can resolve permutation ambiguity between frequency bins, and is applicable even when the number of sources is unknown. A recently proposed complex Gaussian mixture model (cGMM) is highly effective for frequency bin-wise clustering when the number of sources is known. However, it cannot resolve the permutation ambiguity, and is inapplicable when the number of sources is unknown. The proposed PF-cGMM is an extension of the cGMM, which resolves these issues. The resolution of the permutation ambiguity can be realized by a spatial prior called a complex inverse Wishart mixture model (cIWMM). The absence of the permutation ambiguity facilitates source counting, which is performed by hierarchical clustering in this paper. Experiments showed that the PF-cGMM was able to (1) resolve the permutation ambiguity and (2) realize source separation even when the number of sources was unknown with little performance degradation compared to when it was known. Juan Azcarreta, Nobutaka Ito, Shoko Araki, Tomohiro Nakatani |
ICASSP | 3 |
| 2018 | Maximum-Likelihood Online Speaker Diarization in Noisy Meetings Based on Categorical Mixture Model and Probabilistic Spatial DictionaryabstractIn this paper, we propose a maximum-likelihood online diarization method based on a probabilistic spatial dictionary. This dictionary consists of the given probability distribution of spatial features for each possible direction of arrival (DOA) of source signals. Recently, we have developed an online, noise-robust diarization method by utilizing this dictionary as spatial prior information. In this method, DOA estimation is first performed frame-wise based on the dictionary, and subsequently diarization is performed. Although the DOA estimation is performed optimally in the maximum-likelihood sense, the diarization is performed suboptimally based on some heuristics. In contrast, the proposed method performs DOA estimation and diarization jointly and optimally in the maximum-likelihood sense. This is realized by introducing a categorical mixture model (CMM), which has source-wise DOA information and diarization information as unknown parameters. We conducted an experiment on a real-world meeting dataset, and confirmed that the proposed method reduced a diarization error rate by absolute 2.7% compared to the above conventional method. Nobutaka Ito, Takashi Makino, Shoko Araki, Tomohiro Nakatani |
ICASSP | 3 |
| 2018 | Multi-resolution Gammachirp Envelope Distortion Index for Intelligibility Prediction of Noisy Speech
Katsuhiko Yamamoto, Toshio Irino, Narumi Ohashi, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2018 | Distortionless Beamforming Optimized With ℓ1-Norm MinimizationabstractWe propose beamforming method that minimizes the ℓ1norm of a beamformer output vector under the same distortionless constraint as that of the conventional minimum power distortionless response (MPDR) beamformer. Using the ℓ1norm makes the beamformer output sparse. This leads to reducing the residual elements of the interference signal. In addition, the sensitivity of the proposed beamformer can be controlled by adding a norm constraint as in the MPDR beamformer. The proposed method improved the signal-to-interference-noise ratio by 7 dB from that of the MPDR beamformer for reverberation time T60= 300 ms in a simulation. Satoru Emura, Shoko Araki, Tomohiro Nakatani, Noboru Harada |
IEEE Signal Process. Lett. | 2 |
| 2017 | Meeting recognition with asynchronous distributed microphone arrayabstractRecently, recognition of conversational speech such as meetings has widely been studied. However, most existing approaches rely on using a single close talking microphone or a distant microphone array where all the microphones are synchronous. In contrast, this paper tackles a recognition task of conversational speech recorded with asynchronous distributed microphones, to which conventional array processing is not directly applicable. We demonstrate that we can significantly improve recognition performance even when microphones are asynchronous by combining blind synchronization and state-of-the-art microphone array speech enhancement techniques such as independent vector analysis (IVA) and a time-frequency mask based minimum variance distortionless response (MVDR) beamformer. Using such a front-end, we could reduce the word error rate from 42.2 % to 29.9 % for real meeting recordings. Shoko Araki, Nobutaka Ono, Keisuke Kinoshita, Marc Delcroix |
ASRU | 1 |
| 2017 | Probabilistic spatial dictionary based online adaptive beamforming for meeting recognition in noisy and reverberant environmentsabstractHere we propose online adaptive beamforming for automatic speech recognition (ASR) in meetings in noisy, reverberant environments. The proposed method is based on recently developed mask-based beamforming, in which accurate mask estimation and diarization are paramount. Real-world experiments have shown that mask-based beamforming enables accurate ASR in meetings in small noise and reverberation with a signal-to-noise ratio (SNR) of 15–25 dB and a reverberation time (RT) of 120–350 ms. In this paper, we deal with a more adverse condition: meetings in large noise and reverberation with an SNR of 3–15 dB and an RT of 500 ms. To this end, we exploit a probabilistic spatial dictionary, a dictionary that consists of a pre-trained probability distribution of source location features for each potential speaker location. This dictionary enables us to perform mask estimation and diarization for beamforming accurately, even in the above adverse condition. The proposed method reduced the word error rate (WER) on real meeting data by 54.8% relative to our previous beamforming method. Nobutaka Ito, Shoko Araki, Marc Delcroix, Tomohiro Nakatani |
ICASSP | 2 |
| 2017 | Integrating DNN-based and spatial clustering-based mask estimation for robust MVDR beamformingabstractRecently, time-frequency mask-based beamforming has been extensively studied as the frontend of deep neural network (DNN) based automatic speech recognition (ASR) in noisy environments. Two mask estimation approaches have been separately developed for this beamforming method, namely the the DNN-based approach, which exploits the time-frequency features of the signal, and the spatial clustering-based approach, which exploits the spatial features of the signal. This paper proposes a new method that integrates the two approaches in a probabilistic way to further improve mask estimation by exploiting the advantages of both approaches. Experiments using the real data of the CHiME-3 multichannel noisy speech corpus show that the proposed method almost always outperforms the conventional approaches in terms of word error rate (WER) improvement. Tomohiro Nakatani, Nobutaka Ito, Takuya Higuchi, Shoko Araki, Keisuke Kinoshita |
ICASSP | 4 |
| 2017 | Predicting Speech Intelligibility Using a Gammachirp Envelope Distortion Index Based on the Signal-to-Distortion Ratio
Katsuhiko Yamamoto, Toshio Irino, Toshie Matsui, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2017 | Online MVDR Beamformer Based on Complex Gaussian Mixture Model With Spatial Prior for Noise Robust ASRabstractThis paper considers acoustic beamforming for noise robust automatic speech recognition. A beamformer attenuates background noise by enhancing sound components coming from a direction specified by a steering vector. Hence, accurate steering vector estimation is paramount for successful noise reduction. Recently, time-frequency masking has been proposed to estimate the steering vectors that are used for a beamformer. In particular, we have developed a new form of this approach, which uses a speech spectral model based on a complex Gaussian mixture model (CGMM) to estimate the time-frequency masks needed for steering vector estimation, and extended the CGMM-based beamformer to an online speech enhancement scenario. Our previous experiments showed that the proposed CGMM-based approach outperforms a recently proposed mask estimator based on a Watson mixture model and the baseline speech enhancement system of the CHiME-3 challenge. This paper provides additional experimental results for our online processing, which achieves performance comparable to that of batch processing with a suitable block-batch size. This online version reduces the CHiME-3 word error rate (WER) on the evaluation set from 8.37% to 8.06%. Moreover, in this paper, we introduce a probabilistic prior distribution for a spatial correlation matrix (a CGMM parameter), which enables more stable steering vector estimation in the presence of interfering speakers. In practice, the performance of the proposed online beamformer degrades with observations that contain only noise or/and interference because of the failure of the CGMM parameter estimation. The introduced spatial prior enables the target speaker's parameter to avoid overfitting to noise or/and interference. Experimental results show that the spatial prior reduces the WER from 38.4% to 29.2% in a conversation recognition task compared with the CGMM-based approach without the prior, and outperforms a conventional online speech enhancement approach. Takuya Higuchi, Nobutaka Ito, Shoko Araki, Takuya Yoshioka, Marc Delcroix, Tomohiro Nakatani |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Spatial correlation model based observation vector clustering and MVDR beamforming for meeting recognitionabstractThis paper addresses a minimum variance distortionless response (MVDR) beamforming based speech enhancement approach for meeting speech recognition. In a meeting situation, speaker overlaps and noise signals are not negligible. To handle these issues, we employ MVDR beamforming, where accurate estimation of the steering vector is paramount. We recently found that steering vector estimation by clustering the time-frequency components of microphone observation vectors performs well as regards real-world noise reduction. The clustering is performed by taking a cue from the spatial correlation matrix of each speaker, which is realized by modeling the time-frequency components of the observation vectors with a complex Gaussian mixture model (CGMM). Experimental results with real recordings show that the proposed MVDR scheme outperforms conventional null-beamformer based speech enhancement in a meeting situation. Shoko Araki, Masahiro Okada, Takuya Higuchi, Atsunori Ogawa, Tomohiro Nakatani |
ICASSP | 1 |
| 2016 | Modeling audio directional statistics using a complex bingham mixture model for blind source extraction from diffuse noiseabstractMask estimation is a central task in blind signal processing including source separation, denoising, and multi-source localization. In this paper, we define a complex Bingham mixture model (cBMM), and propose it as a model of directional statistics for mask estimation. The complex Bingham distribution can represent not only rotationally symmetric but also rotationally asymmetric distributions. Therefore, it can precisely model stochastic variation of the directional statistics due to reverberation, noise, source movement, etc., which is not necessarily rotationally symmetric. In an experimental evaluation, the proposed cBMM outperformed a conventional complex Watson mixture model (cWMM) in terms of blind source extraction from diffuse noise, reducing the word error rate by 0.91% absolute on CHiME-3 challenge data. Nobutaka Ito, Shoko Araki, Tomohiro Nakatani |
ICASSP | 2 |
| 2016 | Real-time integration of statistical model-based speech enhancement with unsupervised noise PSD estimation using microphone arrayabstractWe propose a technique of multi-channel speech enhancement based on integration of beamforming and statistical model-based speech enhancement to clearly extract the target speech, even in very noisy environments. Conventional microphone array-based techniques estimate speech and noise power spectral densities (PSDs) from the spatial cues of the sound sources; however, their estimation errors dramatically increase when there are many noise sources. We integrated clean speech models trained in advance and the noise PSDs estimated in beamspace to compose observation models and designed a precise Wiener filter. Experiments under adverse noise conditions showed that the proposed technique significantly improved the signal-to-noise ratios (SNRs) compared with the conventional microphone array processing technique. Tomoko Kawase, Kenta Niwa, Masakiyo Fujimoto, Noriyoshi Kamado, Kazunori Kobayashi, Shoko Araki, Tomohiro Nakatani |
ICASSP | 6 |
| 2016 | A generative-discriminative hybrid approach to multi-channel noise reduction for robust automatic speech recognitionabstractIn the recent years, discriminative models have become a very attractive utility and gained a lot of attention in the speech research community, encompassing both front and back-end methods, thanks to their prominent discriminative power and the availability of improved training strategies. When it comes to the recognition of speech that is distorted by highly non-stationary environmental noise, robust front and backend methods are required in order to achieve a satisfactorily high speech recognition performance. Furthermore, when dealing with severe noise conditions, multi-channel front-end methods can be advantageous for suppressing environmental background noise, as compared to single-channel methods. In this work, we improve an existing multi-channel noise reduction approach, referred to as DOminance-based Loca-tional and Power-spectral cHaracteristics INtegration (DOLPHIN), by using a generative-discriminative hybrid model, that makes use of spatial and spectral features. We show that the proposed method outperforms the existing DOLPHIN approach, which is solely based on generative models, in terms of the word error rate reduction achieved on the CHiME-3 challenge data. Hendrik Meutzner, Shoko Araki, Masakiyo Fujimoto, Tomohiro Nakatani |
ICASSP | 2 |
| 2016 | Speech Intelligibility Prediction Based on the Envelope Power Spectrum Model with the Dynamic Compressive Gammachirp Auditory Filterbank
Katsuhiko Yamamoto, Toshio Irino, Toshie Matsui, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani |
INTERSPEECH | 4 |
| 2015 | The NTT CHiME-3 system: Advances in speech enhancement and recognition for mobile multi-microphone devicesabstractCHiME-3 is a research community challenge organised in 2015 to evaluate speech recognition systems for mobile multi-microphone devices used in noisy daily environments. This paper describes NTT's CHiME-3 system, which integrates advanced speech enhancement and recognition techniques. Newly developed techniques include the use of spectral masks for acoustic beam-steering vector estimation and acoustic modelling with deep convolutional neural networks based on the "network in network" concept. In addition to these improvements, our system has several key differences from the official baseline system. The differences include multi-microphone training, dereverberation, and cross adaptation of neural networks with different architectures. The impacts that these techniques have on recognition performance are investigated. By combining these advanced techniques, our system achieves a 3.45% development error rate and a 5.83% evaluation error rate. Three simpler systems are also developed to perform evaluations with constrained set-ups. Takuya Yoshioka, Nobutaka Ito, Marc Delcroix, Atsunori Ogawa, Keisuke Kinoshita, Masakiyo Fujimoto, Chengzhu Yu, Wojciech J. Fabian, Miquel Espi, Takuya Higuchi, Shoko Araki, Tomohiro Nakatani |
ASRU | 11 |
| 2015 | Exploring multi-channel features for denoising-autoencoder-based speech enhancementabstractThis paper investigates a multi-channel denoising autoencoder (DAE)-based speech enhancement approach. In recent years, deep neural network (DNN)-based monaural speech enhancement and robust automatic speech recognition (ASR) approaches have attracted much attention due to their high performance. Although multi-channel speech enhancement usually outperforms single channel approaches, there has been little research on the use of multi-channel processing in the context of DAE. In this paper, we explore the use of several multi-channel features as DAE input to confirm whether multi-channel information can improve performance. Experimental results show that certain multi-channel features outperform both a monaural DAE and a conventional time-frequency-mask-based speech enhancement method. Shoko Araki, Tomoki Hayashi, Marc Delcroix, Masakiyo Fujimoto, Kazuya Takeda, Tomohiro Nakatani |
ICASSP | 1 |
| 2014 | Probabilistic integration of diffuse noise suppression and dereverberationabstractThis paper deals with joint suppression of diffuse noise and reverberation, to enhance perceived speech quality and speech recognition performance. Although diffuse noise and reverberation are both omnipresent in the real world, conventional methods have modeled only one while neglecting the other. In contrast, we propose a novel joint suppression method that employs a unified probabilistic model of observed signals affected by both diffuse noise and reverberation. Through likelihood maximization, this unified model enables proper parameter estimation that takes into account both diffuse noise and reverberation. As a byproduct, we also propose a novel method for diffuse noise suppression. Experimental results demonstrate the effectiveness of the proposed joint suppression method in terms of dereverberation and denoising. Nobutaka Ito, Shoko Araki, Tomohiro Nakatani |
ICASSP | 2 |
| 2013 | Permutation-free convolutive blind source separation via full-band clustering based on frequency-independent source presence priorsabstractWe propose permutation-free frequency-domain blind source separation (BSS) via full-band clustering of the time-frequency (T-F) components based on time-varying signal presence priors. Frequency-domain methods of BSS usually process each frequency bin separately, and therefore necessitate the subsequent alignment of the permutation ambiguity that arises between frequency bins. In contrast, the proposed method simultaneously processes all frequency bins by using a mixture model with time-varying, frequency-independent mixture weights. We propose to assume non-sparse priors on the mixture weights to prevent the degradation of source separation performance by the time-varying mixture weights. We propose a customized expectation-maximization (EM) algorithm for the maximum a posteriori (MAP) estimation of the model parameters, to which we introduce a novel technique to avoid convergence to local maxima. For audio source separation, we use the normalized observation vector as the feature vector, and theWatson mixture model (WMM) as the mixture model. Evaluations confirm that the proposed permutation-free BSS results in source separation performance comparable to the state-of-the-art clustering-based BSS composed of bin-wise clustering and permutation alignment. Nobutaka Ito, Shoko Araki, Tomohiro Nakatani |
ICASSP | 2 |
| 2013 | Coupling beamforming with spatial and spectral feature based spectral enhancement and its application to meeting recognitionabstractThis paper discusses microphone array based interference reduction approaches for robust automatic speech recognition. A model based multichannel spectral enhancement approach has recently been proposed for effectively reducing interference by exploiting both the spatial and spectral features of the signals. With the goal of further improving the effectiveness of this approach, we propose a new framework that combines this approach with a microphone-array based beamforming approach. Because the two approaches can work in a complementary manner in the proposed framework, they can greatly improve the interference reduction performance. We apply the proposed framework to the recognition of actual meetings, and show that it is superior to the use of beamforming or spectral enhancement alone in terms of the word error rates. Tomohiro Nakatani, Mehrez Souden, Shoko Araki, Takuya Yoshioka, Takaaki Hori, Atsunori Ogawa |
ICASSP | 3 |
| 2013 | Speech recognition in living rooms: Integrated speech enhancement and recognition system based on spatial, spectral and temporal modeling of sounds
Marc Delcroix, Keisuke Kinoshita, Tomohiro Nakatani, Shoko Araki, Atsunori Ogawa, Takaaki Hori, Shinji Watanabe 0001, Masakiyo Fujimoto, Takuya Yoshioka, Takanobu Oba, Yotaro Kubo, Mehrez Souden, Seong-Jun Hahm, Atsushi Nakamura |
Comput. Speech Lang. | 4 |
| 2013 | Dominance Based Integration of Spatial and Spectral Features for Speech EnhancementabstractThis paper proposes a versatile technique for integrating two conventional speech enhancement approaches, a spatial clustering approach (SCA) and a factorial model approach (FMA), which are based on two different features of signals, namely spatial and spectral features, respectively. When used separately the conventional approaches simply identify time frequency (TF) bins that are dominated by interference for speech enhancement. Integration of the two approaches makes identification more reliable, and allows us to estimate speech spectra more accurately even in highly nonstationary interference environments. This paper also proposes extensions of the FMA for further elaboration of the proposed technique, including one that uses spectral models based on mel-frequency cepstral coefficients and another to cope with mismatches, such as channel mismatches, between captured signals and the spectral models. Experiments using simulated and real recordings show that the proposed technique can effectively improve audible speech quality and the automatic speech recognition score. Tomohiro Nakatani, Shoko Araki, Takuya Yoshioka, Marc Delcroix, Masakiyo Fujimoto |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Multichannel Extensions of Non-Negative Matrix Factorization With Complex-Valued DataabstractThis paper presents new formulations and algorithms for multichannel extensions of non-negative matrix factorization (NMF). The formulations employ Hermitian positive semidefinite matrices to represent a multichannel version of non-negative elements. Multichannel Euclidean distance and multichannel Itakura-Saito (IS) divergence are defined based on appropriate statistical models utilizing multivariate complex Gaussian distributions. To minimize this distance/divergence, efficient optimization algorithms in the form of multiplicative updates are derived by using properly designed auxiliary functions. Two methods are proposed for clustering NMF bases according to the estimated spatial property. Convolutive blind source separation (BSS) is performed by the multichannel extensions of NMF with the clustering mechanism. Experimental results show that 1) the derived multiplicative update rules exhibited good convergence behavior, and 2) BSS tasks for several music sources with two microphones and three instrumental parts were evaluated successfully. Hiroshi Sawada, Hirokazu Kameoka, Shoko Araki, Naonori Ueda |
IEEE Trans. Speech Audio Process. | 3 |
| 2013 | A Multichannel MMSE-Based Framework for Speech Source Separation and Noise ReductionabstractWe propose a new framework for joint multichannel speech source separation and acoustic noise reduction. In this framework, we start by formulating the minimum-mean-square error (MMSE)-based solution in the context of multiple simultaneous speakers and background noise, and outline the importance of the estimation of the activities of the speakers. The latter is accurately achieved by introducing a latent variable that takes N+1 possible discrete states for a mixture of N speech signals plus additive noise. Each state characterizes the dominance of one of the N+1 signals. We determine the posterior probability of this latent variable, and show how it plays a twofold role in the MMSE-based speech enhancement. First, it allows the extraction of the second order statistics of the noise and each of the speech signals from the noisy data. These statistics are needed to formulate the multichannel Wiener-based filters (including the minimum variance distortionless response). Second, it weighs the outputs of these linear filters to shape the spectral contents of the signals' estimates following the associated target speakers' activities. We use the spatial and spectral cues contained in the multichannel recordings of the sound mixtures to compute the posterior probability of this latent variable. The spatial cue is acquired by using the normalized observation vector whose distribution is well approximated by a Gaussian-mixture-like model, while the spectral cue can be captured by using a pre-trained Gaussian mixture model for the log-spectra of speech. The parameters of the investigated models and the speakers' activities (posterior probabilities of the different states of the latent variable) are estimated via expectation maximization. Experimental results including comparisons with the well-known independent component analysis and masking are provided to demonstrate the efficiency of the proposed framework. Mehrez Souden, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani, Hiroshi Sawada |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Sparse vector factorization for underdetermined BSS using wrapped-phase GMM and source log-spectral priorabstractWe propose a sparse vector factorization (SVF) approach for blind source separation, which inherently avoids the permutation problem. The SVF assumes the sparseness of sources, and defines a sparse vector (SV) that consists of the locational and spectral features of each source at all the frequencies. Then, by assuming that the locational and spectral SVs are generated by frequency-independent parameters, the method executes the SVF. Our locational feature is the phase difference (PD) between two microphone observations, and we model it with a frequency-independent time-difference of arrival (TDOA) parameter. Moreover, we employ the wrapped-phase GMM in order to take the spatial aliasing problem into account. On the other hand, the spectral feature is the log spectrum, and we provide a prior for a spectral parameter. The SVF is formulated with a maximum a posteriori (MAP) estimation framework, where the locational and spectral parameters are inferred by the EM algorithm. Experimental results show that our proposed method can separate signals successfully even for an underdetermined case. Shoko Araki, Tomohiro Nakatani |
ICASSP | 1 |
| 2012 | New analytical update rule for TDOA inference for underdetermined BSS in noisy environmentsabstractIn this paper, we propose a new technique for sparseness-based underdetermined BSS that is based on the clustering of the frequency-dependent time difference of arrival (TDOA) information and that can cope with diffused noise environments. Such a method with an EM algorithm has already been proposed, however, it required a time-consuming exhaust search for TDOA inference. To remove the need for such an exhaust search, we propose a new technique by focusing on a stereo case. We derive an update rule for analytical TDOA estimation. This update rule eliminates the need for the exhaustive TDOA search, and therefore reduces the computational load. We show experimental results for separation performance and calculation time in comparison with those obtained with the conventional approach. Our reported results validate our proposed method, that is, our proposed method achieves high performance without a high computational cost. Takuro Maruyama, Shoko Araki, Tomohiro Nakatani, Shigeki Miyabe, Takeshi Yamada, Shoji Makino, Atsushi Nakamura |
ICASSP | 2 |
| 2012 | LogMax observation model with MFCC-based spectral prior for reduction of highly nonstationary ambient noiseabstractThis paper proposes a new single/multi-channel speech enhancement approach based on a LogMax observation model integrated with Gaussian mixture models of speech and noise mel-frequency cepstral coefficients (MFCC-GMM). It has been reported that the LogMax observation model has high potential for reducing highly nonstationary noise, for example, when it is combined with factorial hidden Markov models. In addition, it has recently been shown that a source location based speech enhancement approach can be easily incorporated into this model for more efficient and reliable estimation. However, the unique structure of the LogMax model has prevented us from using it with MFCC-GMMs, which is a fundamental limitation of this approach. Our proposal in this paper is aimed at overcoming this limitation. Experiments using the PASCAL CHiME separation and recognition challenge task show the superiority of the proposed approach as regards both speech quality and automatic speech recognition performance. Tomohiro Nakatani, Takuya Yoshioka, Shoko Araki, Marc Delcroix, Masakiyo Fujimoto |
ICASSP | 3 |
| 2012 | Efficient algorithms for multichannel extensions of Itakura-Saito nonnegative matrix factorizationabstractThis paper proposes new algorithms for multichannel extensions of nonnegative matrix factorization (NMF) with the Itakura-Saito (IS) divergence. We employ Hermitian positive definite matrices for modeling the covariance matrix of a multivariate complex Gaussian distribution. Such matrices are basically estimated for NMF bases, but a source separation task can be performed by introducing variables that relate NMF bases and sources. The new algorithms are derived by using a majorization scheme with properly designed auxiliary functions. The algorithms are in the form of multiplicative updates, and exhibit good convergence behavior. We have succeeded in separating a professionally produced music recording into its vocal and guitar components. Hiroshi Sawada, Hirokazu Kameoka, Shoko Araki, Naonori Ueda |
ICASSP | 3 |
| 2012 | A multichannel MMSE-based framework for joint blind source separation and noise reductionabstractIn this paper, we propose a new framework to separate multiple speech signals and reduce the additive acoustic noise using multiple microphones. In this framework, we start by formulating the minimum-mean-square error (MMSE) criterion to retrieve each of the desired speech signals from the observed mixtures of sounds and outline the importance of multi-speaker activity detection. The latter is modeled by introducing a latent variable whose posterior probability is computed via expectation maximization (EM) combining both the spatial and spectral cues of the multichannel speech observations. We experimentally demonstrate that the resulting joint blind source separation (BSS) and noise reduction solution performs remarkably well in reverberant and noisy environments. Mehrez Souden, Shoko Araki, Keisuke Kinoshita, Tomohiro Nakatani, Hiroshi Sawada |
ICASSP | 2 |
| 2012 | The signal separation evaluation campaign (2007-2010): Achievements and remaining challenges
Emmanuel Vincent 0001, Shoko Araki, Fabian J. Theis, Guido Nolte, Pau Bofill, Hiroshi Sawada, Alexey Ozerov, Vikrham Gowreesunker, Dominik Lutter, Ngoc Q. K. Duong |
Signal Process. | 2 |
| 2012 | Low-Latency Real-Time Meeting Recognition and Understanding Using Distant Microphones and Omni-Directional CameraabstractThis paper presents our real-time meeting analyzer for monitoring conversations in an ongoing group meeting. The goal of the system is to recognize automatically “who is speaking what” in an online manner for meeting assistance. Our system continuously captures the utterances and face poses of each speaker using a microphone array and an omni-directional camera positioned at the center of the meeting table. Through a series of advanced audio processing operations, an overlapping speech signal is enhanced and the components are separated into individual speaker's channels. Then the utterances are sequentially transcribed by our speech recognizer with low latency. In parallel with speech recognition, the activity of each participant (e.g., speaking, laughing, watching someone) and the circumstances of the meeting (e.g., topic, activeness, casualness) are detected and displayed on a browser together with the transcripts. In this paper, we describe our techniques and our attempt to achieve the low-latency monitoring of meetings, and we show our experimental results for real-time meeting transcription. Takaaki Hori, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto, Shinji Watanabe 0001, Takanobu Oba, Atsunori Ogawa, Kazuhiro Otsuka, Dan Mikami, Keisuke Kinoshita, Tomohiro Nakatani, Atsushi Nakamura, Junji Yamato |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Probabilistic Speaker Diarization With Bag-of-Words Representations of Speaker Angle InformationabstractSpeaker diarization determines “who spoke when” from the recorded conversations of an unknown number of people. In general, we have no a priori information about the number, the locations, or even the characteristics of the speakers. Additionally, speakers' speech utterances vary dynamically because of turn-taking during the conversations. These conditions make the speaker-clustering task extremely difficult. The problem becomes even harder if online (incremental) processing is required. In this paper, we formulate the speaker-clustering problem as the clustering of the sequential audio features generated by an unknown number of latent mixture components (speakers). We employ a probabilistic model that assumes time-sensitive speaker mixtures at every time frame, which, surprisingly, suits the diarization scenario. We combine the time-varying probabilistic model with direction of arrival (DOA) information calculated from a microphone array in a bag-of-words (BoW)-style feature representation. The proposed system effectively estimates the number and locations of the speakers in an online manner based on the standard Bayes inference scheme. Experiments confirm that the proposed model can successfully infer the number and features of speakers and yield better or comparable speaker diarization results compared with conventional methods in several datasets. Katsuhiko Ishiguro, Takeshi Yamada, Shoko Araki, Tomohiro Nakatani, Hiroshi Sawada |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Hybrid approach for multichannel source separation combining time-frequency mask with multi-channel Wiener filterabstractThis paper discusses a hybrid approach for the multi-channel source separation, where both a time-frequency (t-f) mask and a multi-channel Wiener filter (WF) are utilized. T-f mask based approaches have been widely studied, because they can separate signals with a low calculation cost. However, the separated signals with a t-f mask usually contain a non-linear distortion. On the other hand, a new multi-channel WF framework employing a spatial covariance matrix model has recently been proposed. With the WF method, we can obtain separated signals of better quality than with a t-f mask, however, the method is computationally expensive because it requires many iterations for the optimization. In this paper, in order to take advantages of both approaches, first we explain the hybrid algorithm by introducing the t-f mask concept to the WF approach. Then we show that the hybrid approach achieves high performance without the iterative calculation for the WF. We also present the way for applying the hybrid method to the case where the number of sources is unavailable. Shoko Araki, Tomohiro Nakatani |
ICASSP | 1 |
| 2011 | Joint unsupervised learning of hidden Markov source models and source location models for multichannel source separationabstractThis paper discusses a multichannel source separation approach that exploits the statistical characteristics of source location cues characterized by steering vector models (SM) and those of source log spectra characterized by hidden Markov models (spectral HMM). Recently, it was shown that the use of speaker independent spectral HMMs trained in advance substantially improves the quality of speech signals separated based on source location cues in a computationally efficient manner. However, with this approach, mismatches between the spectral HMMs and the observation may substantially degrade the separation quality, which limits the applicability of this approach. To overcome this problem, this paper proposes a method for learning the parameters of the spectral HMMs jointly with those of the SMs from the observed sound mixtures. Experimental results show that the proposed method works effectively for separation of convolutive sound mixtures. Tomohiro Nakatani, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto |
ICASSP | 2 |
| 2011 | Formulations and algorithms for multichannel complex NMFabstractThis paper studies some formulations and algorithms for the multichannel extension of nonnegative matrix factorization (NMF). We model the inter-channel characteristics of each NMF basis, including both the amplitude ratios and the phase differences on a channel pair. The learned inter-channel characteristics provide useful information for binding each NMF basis to each source component in such a situation that multiple sources are mixed in a convolutive manner and observed at multiple microphones. Effective optimization algorithms based on majorization are derived by using properly designed auxiliary functions. Experimental results show that the algorithms converged favorably regardless of the initialization. Hiroshi Sawada, Hirokazu Kameoka, Shoko Araki, Naonori Ueda |
ICASSP | 3 |
| 2011 | Reduction of Highly Nonstationary Ambient Noise by Integrating Spectral and Locational Characteristics of Speech and Noise for Robust ASR
Tomohiro Nakatani, Shoko Araki, Marc Delcroix, Takuya Yoshioka, Masakiyo Fujimoto |
INTERSPEECH | 2 |
| 2011 | Underdetermined Convolutive Blind Source Separation via Frequency Bin-Wise Clustering and Permutation AlignmentabstractThis paper presents a blind source separation method for convolutive mixtures of speech/audio sources. The method can even be applied to an underdetermined case where there are fewer microphones than sources. The separation operation is performed in the frequency domain and consists of two stages. In the first stage, frequency-domain mixture samples are clustered into each source by an expectation-maximization (EM) algorithm. Since the clustering is performed in a frequency bin-wise manner, the permutation ambiguities of the bin-wise clustered samples should be aligned. This is solved in the second stage by using the probability on how likely each sample belongs to the assigned class. This two-stage structure makes it possible to attain a good separation even under reverberant conditions. Experimental results for separating four speech signals with three microphones under reverberant conditions show the superiority of the new method over existing methods. We also report separation results for a benchmark data set and live recordings of speech mixtures. Hiroshi Sawada, Shoko Araki, Shoji Makino |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Simultaneous clustering of mixing and spectral model parameters for blind sparse source separationabstractThis paper proposes a sparse source separation method which clusters the phase difference between the microphone observations and the amplitude modulation (AM) of the source spectrum simultaneously. The phase difference clustering separates the signals in each frequency bin, and the AM clustering corresponds to permutation alignment. Because the proposed method has an inherent ability to align the permutation of frequency components, the proposed method can be applied even when the spatial aliasing problem occurs. Moreover, because the common AM property collects the synchronized frequency components, we can model the microphone observations with a small number of sources. This property enables us to count the number of sources. That is, the proposed method can be applied even if the number of sources is unknown. The experimental results confirm the effectiveness of our proposed method. Shoko Araki, Tomohiro Nakatani, Hiroshi Sawada |
ICASSP | 1 |
| 2010 | Single channel source separation based on sparse source observation model with harmonic constraintabstractThis paper proposes a general single channel source separation approach that exploits statistical characteristics of the source including sparseness. A new observation model for a mixture of sparse sources is introduced for this purpose. With this approach, source separation is achieved by iterating two simple sub-procedures, namely the clustering of the time-frequency (TF) bins into individual sources and the separate updating of the model parameters of each source. An advantage of this approach is that we can update the model parameters of each source assuming each cluster to contain a single source, and thus we can utilize the various model parameter estimation algorithms used for single source analysis, which can be simple and accurate, in an efficient and unified manner. We implement a harmonicity based source separation method with this approach using a robust fundamental frequency (F0) estimation algorithm. The experimental results confirm the effectiveness of the proposed method. Tomohiro Nakatani, Shoko Araki |
ICASSP | 2 |
| 2010 | Multichannel source separation based on source location cue with log-spectral shaping by hidden Markov source model
Tomohiro Nakatani, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto |
INTERSPEECH | 2 |
| 2010 | Cepstral smoothing of separated signals for underdetermined speech separationabstractMusical noise is a typical problem with blind source separation using a time-frequency mask. Recently, the cepstral smoothing of spectral masks (CSM) was proposed. Based on the idea of smoothing in the cepstral domain, this paper proposes the cepstral smoothing of separated signals (CSS) on the assumption that a cepstral representation better reflects the characteristics of speech signals than those of masks (or filter gains). We also report a comparative evaluation study of CSM and CSS with other musical noise reduction methods. Our experimental results show that CSM is effective for musical noise reduction, but the target speech was relatively distorted. On the other hand, our proposed CSS produced less distorted target signals with the same musical noise reduction as CSM. Yumi Ansa, Shoko Araki, Shoji Makino, Tomohiro Nakatani, Takeshi Yamada, Atsushi Nakamura, Nobuhiko Kitawaki |
ISCAS | 2 |
| 2010 | Real-time meeting recognition and understanding using distant microphones and omni-directional cameraabstractThis paper presents our newly developed real-time meeting analyzer for monitoring conversations in an ongoing group meeting. The goal of the system is to automatically recognize “who is speaking what” in an online manner for meeting assistance. Our system continuously captures the utterances and the face pose of each speaker using a distant microphone array and an omni-directional camera at the center of the meeting table. Through a series of advanced audio processing operations, an overlapping speech signal is enhanced and the components are separated into individual speaker's channels. Then the utterances are sequentially transcribed by our speech recognizer with low latency. In parallel with speech recognition, the activity of each participant (e.g. speaking, laughing, watching someone) and the situation of the meeting (e.g. topic, activeness, casualness) are detected and displayed on a browser together with the transcripts. In this paper, we describe our techniques and our attempt to achieve the low-latency monitoring of meetings, and we show our experimental results for real-time meeting transcription. Takaaki Hori, Shoko Araki, Takuya Yoshioka, Masakiyo Fujimoto, Shinji Watanabe 0001, Takanobu Oba, Atsunori Ogawa, Kazuhiro Otsuka, Dan Mikami, Keisuke Kinoshita, Tomohiro Nakatani, Atsushi Nakamura, Junji Yamato |
SLT | 2 |
| 2010 | Speech Activity Detection for Multi-Party Conversation Analyses Based on Likelihood Ratio Test on Spatial MagnitudeabstractThis paper proposes a microphone array-based speech activity detection (SAD) method for analyzing multi-party conversations recorded in the presence of noise. In particular, the proposed method considers conversations where the number of speakers and speaker locations cannot be restricted, such as when standing and talking, and at poster sessions. When we observe such conversations, there are directional noise sources and diffuse noise that affect the direction of arrival estimations of the target speech signals. To detect speech activity without a priori knowledge about the speakers and noise environments, a likelihood ratio test (LRT)-based SAD method is applied to spatial magnitude, which are estimated by using the time-frequency masking of the observed spectra. The proposed method can exploit the enhanced signals obtained from time-frequency masking, and works even in the presence of environmental noise. Experiments with recorded simulated poster sessions confirmed that the proposed method could outperform conventional methods based on the LRT for a single channel, magnitude coherence, or crosspower spectrum phase. Kentaro Ishizuka, Shoko Araki, Tatsuya Kawahara |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | An Optical Access Network System without a Power Supply Using Blind Speech Separation and a Loopback TechniqueabstractThis paper proposes an optical access system that enables voice communication services to continue when a power failure occurs. In particular, we propose applying a digital signal processing technique namely blind speech separation (BSS) to an optical multiple access system for demultiplexing randomly mixture signals received from an upstream user. We present a brief overview and the target of proposed system and describe the key technique, namely the BSS procedure. We then report system feasibility studies for the first step of a numerical simulation (assuming two wavelengths and three users) and experimental tests (assuming two wavelengths and two users) using actual sample voice signals. We confirmed that separation was realized by using a subjective assessment that consisted of listening to obtained signals. Takayoshi Tashiro, Shoko Araki, Yasuhiko Nakanishi, Hideaki Kimura 0002, Kiyomi Kumozaki, Masato Miyoshi |
GLOBECOM | 2 |
| 2009 | Blind sparse source separation for unknown number of sources using Gaussian mixture model fitting with Dirichlet priorabstractIn this paper, we propose a novel sparse source separation method that can be applied even if the number of sources is unknown. Recently, many sparse source separation approaches with time-frequency masks have been proposed. However, most of these approaches require information on the number of sources in advance. In our proposed method, we model the histogram of the estimated direction of arrival (DOA) with a Gaussian mixture model (GMM) with a Dirichlet prior. Then we estimate the model parameters by using the maximum a posteriori estimation based on the EM algorithm. In order to avoid one cluster being modeled by two or more Gaussians, we utilize a sparse distribution modeled by the Dirichlet distributions as the prior of the GMM mixture weight. By using this prior, without any specific model selection process, our proposed method can estimate the number of sources and time-frequency masks simultaneously. Experimental results show the performance of our proposed method. Shoko Araki, Tomohiro Nakatani, Hiroshi Sawada, Shoji Makino |
ICASSP | 1 |
| 2009 | A speaker diarization method based on the probabilistic fusion of audio-visual location informationabstractThis paper proposes a speaker diarization method for determining ""who spoke when"" in multi-party conversations, based on the probabilistic fusion of audio and visual location information. The audio and visual information is obtained from a compact system designed to analyze round table multi-party conversations. The system consists of two cameras and a triangular microphone array with three microphones, and can cover a spherical region. Speaker locations are estimated from audio and visual observations in terms of azimuths from this recording system. Unlike conventional speech diarization methods, our proposed method estimates the probability of the presence of multiple simultaneous speakers in a physical space with a small microphone setup instead of using a cascade consisting of speech activity detection, direction of arrival estimation, acoustic feature extraction, and information criteria based speaker segmentation. To estimate the speaker presence more correctly, the speech presence probabilities in a physical space are integrated with the probabilities estimated from participants' face locations obtained with a robust particle filtering based face tracker with two cameras equipped with fisheye lenses. The locations in a physical space with highly integrated probabilities are then classified into a certain number of speaker classes by using on-line classification to realize speaker diarization. The probability calculations and speaker classifications are conducted on-line, making it unnecessary to observe all the conversation data. An experiment using real casual conversations, which include more overlaps and short speech segments than formal meetings, showed the advantages of the proposed method. Kentaro Ishizuka, Shoko Araki, Kazuhiro Otsuka, Tomohiro Nakatani, Masakiyo Fujimoto |
ICMI | 2 |
| 2009 | Realtime meeting analysis and 3D meeting viewer based on omnidirectional multimodal sensorsabstractThis demo presents a realtime system for analyzing group meetings. Targeting round-table meetings, this system employs an omnidirectional camera-microphone system. The goal of this system is to automatically discover "who is talking to whom and when". To that purpose, the face pose/position of meeting participants are tracked on panorama images acquired from fisheye-based omnidirectional cameras. From audio signals obtained with microphone array, speaker diarization, i.e. the estimation of "who is speaking and when", is carried out. The visual focus of attention, i.e. "who is looking at whom", is esimated from the result of face tracking. The results are displayed based on a 3D visualization scheme. The advantage of our system is its realtimeness. We will demonstrate the portable version of the system consisting of two laptop PCs. In addition, we will showcase our meeting playback viewer with man-machine interfaces that allow users to freely control space and time of meeting scenes. With this viewer, users can also experince 3D positional sound effect linked with 3D viewpoint, using enhanced audio tracks for each participant. Kazuhiro Otsuka, Shoko Araki, Dan Mikami, Kentaro Ishizuka, Masakiyo Fujimoto, Junji Yamato |
ICMI | 2 |
| 2009 | Frequency-Domain Pearson Distribution Approach for Independent Component Analysis (FD-Pearson-ICA) in Blind Source SeparationabstractIn frequency-domain blind source separation (BSS) for speech with independent component analysis (ICA), a practical parametric Pearson distribution system is used to model the distribution of frequency-domain source signals. ICA adaptation rules have a score function determined by an approximated signal distribution. Approximation based on the data may produce better separation performance than we can obtain with ICA. Previously, conventional hyperbolic tangent$(tanh)$or generalized Gaussian distribution (GGD) was uniformly applied to the score function for all frequency bins, even though a wideband speech signal has different distributions at different frequencies. To deal with this, we propose modeling the signal distribution at each frequency by adopting a parametric Pearson distribution and employing it to optimize the separation matrix in the ICA learning process. The score function is estimated by the appropriate Pearson distribution parameters for each frequency bin. We devised three methods for Pearson distribution parameter estimation and conducted separation experiments with real speech signals convolved with actual room impulse responses$(T_{60}=130\ {\hbox {ms}})$. Our experimental results show that the proposed frequency-domain Pearson-ICA (FD-Pearson-ICA) adapted well to the characteristics of frequency-domain source signals. By applying the FD-Pearson-ICA performance, the signal-to-interference ratio significantly improved by around 2–3 dB compared with conventional nonlinear functions. Even if the signal-to-interference ratio (SIR) values of FD-Pearson-ICA were poor, the performance based on a disparity measure between the true score function and estimated parametric score function clearly showed the advantage of FD-Pearson-ICA. Furthermore, we confirmed the optimum of the proposed approach for/optimized the proposed approach as regards separation performance. By combining individual distribution parameters directly estimated at low frequency with the appropriate parameters optimized at high frequency, it was possible to both reasonably improve the FD-Pearson-ICA performance without any significant increase in the computational burden by comparison with conventional nonlinear functions. Hiroko Kato Solvang, Yuichi Nagahara, Shoko Araki, Hiroshi Sawada, Shoji Makino |
IEEE Trans. Speech Audio Process. | 3 |
| 2008 | Speaker indexing and speech enhancement in real meetings / conversationsabstractThis paper presents a speaker indexing method that uses a small number of microphones to estimate who spoke when. Our proposed speaker indexing is realized by using a noise robust voice activity detector (VAD), a QCC-PHAT based direction of arrival (DOA) estimator, and a DOA classifier. Using the estimated speaker indexing information, we can also enhance the utterances of each speaker with a maximum signal-to-noise-ratio (MaxSNR) beamformer. This paper applies our system to real recorded meetings / conversations recorded in a room with a reverberation time of 350 ms, and evaluates the performance by a standard measure: the diarization error rate (DER). Even for the real conversations, which have many speaker turn-takings and overlaps, the speaker error time was very small with our proposed system. We are planning to demonstrate a real-time speaker indexing system at ICASSP2008. Shoko Araki, Masakiyo Fujimoto, Kentaro Ishizuka, Hiroshi Sawada, Shoji Makino |
ICASSP | 1 |
| 2008 | A realtime multimodal system for analyzing group meetings by combining face pose tracking and speaker diarizationabstractThis paper presents a realtime system for analyzing group meetings that uses a novel omnidirectional camera-microphone system. The goal is to automatically discover the visual focus of attention (VFOA), i.e. "who is looking at whom", in addition to speaker diarization, i.e. "who is speaking and when". First, a novel tabletop sensing device for round-table meetings is presented; it consists of two cameras with two fisheye lenses and a triangular microphone array. Second, from high-resolution omnidirectional images captured with the cameras, the position and pose of people's faces are estimated by STCTracker (Sparse Template Condensation Tracker); it realizes realtime robust tracking of multiple faces by utilizing GPUs (Graphics Processing Units). The face position/pose data output by the face tracker is used to estimate the focus of attention in the group. Using the microphone array, robust speaker diarization is carried out by a VAD (Voice Activity Detection) and a DOA (Direction of Arrival) estimation followed by sound source clustering. This paper also presents new 3-D visualization schemes for meeting scenes and the results of an analysis. Using two PCs, one for vision and one for audio processing, the system runs at about 20 frames per second for 5-person meetings. Kazuhiro Otsuka, Shoko Araki, Kentaro Ishizuka, Masakiyo Fujimoto, Martin Heinrich, Junji Yamato |
ICMI | 2 |
| 2008 | Statistical speech activity detection based on spatial power distribution for analyses of poster presentationsabstractThis paper proposes a microphone array based statistical speech activity detection (SAD) method for analyses of poster presentations recorded in the presence of noise. Such poster presentations are a kind of multi-party conversation, where the number of speakers and speaker location are unrestricted, and directional noise sources affect the direction of arrival of the target speech signals. To detect speech activity in such cases without a priori knowledge about the speakers and noise environments, we applied a likelihood ratio test based SAD method to spatial power distributions. The proposed method can exploit the enhanced signals obtained from timefrequency masking, and work even in the presence of environmental noise by utilizing the a priori signal-to-noise ratios of the spatial power distributions. Experiments with recorded poster presentations confirmed that the proposed method significantly improves the SAD accuracies compared with those obtained with a frequency spectrum based statistical SAD method. Index Terms: speech activity detection, microphone arrays, multi-party conversations, spatial power distribution Kentaro Ishizuka, Shoko Araki, Tatsuya Kawahara |
INTERSPEECH | 2 |
| 2008 | Multi-modal recording, analysis and indexing of poster sessionsabstractA new project on multi-modal analysis of poster sessions is introduced. We have designed an environment dedicated to recording of poster conversations using multiple sensors, and collected a number of sessions, to which a variety of multi-modal information is annotated, including utterance units for individual speakers, backchannels, nodding, gazing, and pointing. Automatic speaker diarization, that is a combination of speech activity detection and speaker identification, is conducted using a set of distant microphones, and a reasonable performance is obtained. Then, we investigate automatic classification of conversation segments into two modes: presentation mode and question-answer mode. Preliminary experiments show that multi-modal features on non-verbal behaviors play a significant role in the indexing of this kind of conversations. Index Terms: multi-modal corpus, poster conversation, speaker diarization, non-verbal information Tatsuya Kawahara, Hisao Setoguchi, Katsuya Takanashi, Kentaro Ishizuka, Shoko Araki |
INTERSPEECH | 5 |
| 2008 | Missing feature speech recognition in a meeting situation with maximum SNR beamformingabstractEspecially for tasks like automatic meeting transcription, it would be useful to automatically recognize speech also while multiple speakers are talking simultaneously. For this purpose, speech separation can be performed, for example by using maximum SNR beamforming. However, even when good interferer suppression is attained, the interfering speech will still be recognizable during those intervals, where the target speaker is silent. In order to avoid the consequential insertion errors, a new soft masking scheme is proposed, which works in the time domain by inducing a large damping on those temporal periods, where the observed direction of arrival does not correspond to that of the target speaker. Even though the masking scheme is aggressive, by means of missing feature recognition the recognition accuracy can be improved significantly, with relative error reductions in the order of 60% compared to maximum SNR beamforming alone, and it is successful also for three simultaneously active speakers. Results are reported based on the SOLON speech recognizer, NTT’s large vocabulary system [1], which is applied here for the recognition of artificially mixed data using real-room impulse responses and the entire clean test set of the Aurora 2 database. Dorothea Kolossa, Shoko Araki, Marc Delcroix, Tomohiro Nakatani, Reinhold Orglmeister, Shoji Makino |
ISCAS | 2 |
| 2007 | Blind Speech Separation in a Meeting Situation with Maximum SNR BeamformersabstractWe propose a speech separation method for a meeting situation, where each speaker sometimes speaks and the number of speakers changes every moment. Many source separation methods have already been proposed, however, they consider a case where all the speakers keep speaking: this is not always true in a real meeting. In such cases, in addition to separation, speech detection and the classification of the detected speech according to speaker become important issues. For that purpose, we propose a method that employs a maximum signal-to-noise (MaxSNR) beamformer combined with a voice activity detector and online clustering. We also discuss the scaling ambiguity problem as regards the MaxSNR beamformer, and provide their solutions. We report some encouraging results for a real meeting in a room with a reverberation time of about 350 ms. Shoko Araki, Hiroshi Sawada, Shoji Makino |
ICASSP (1) | 1 |
| 2007 | Blind Source Separation Based on a Beamformer Array and Time Frequency Binary MaskingabstractThis paper deals with a new technique for blind source separation (BSS) from convolutive mixtures. We present a three-stage separation system employing time-frequency binary masking, beamforming and a non-linear post processing technique. The experiments show that this system outperforms conventional time-frequency binary masking (TFBM) in both (over-)determined and underdetermined cases. Moreover it removes the musical noise and reduces interference in time-frequency slots extracted by TFBM. Jan Cermak, Shoko Araki, Hiroshi Sawada, Shoji Makino |
ICASSP (1) | 2 |
| 2007 | Two-Microphone Voice Activity Detection Based on the Homogeneity of the Direction of Arrival EstimatesabstractVoice activity detection (VAD) systems have been the object of continuous research during the last three decades. While single microphone systems cannot take advantage of certain spatial properties of speech signals, microphone array systems consisting of many elements based on beamforming techniques can be difficult to implement in reality due to cost and complexity issues. The aim of the work described in this paper was to achieve both practical feasibility and spatial discrimination ability. A new approach is developed for two-microphone VAD capable of profiting from the concentration of speech energy in time, frequency and space. The algorithm is implemented and compared with several standard VAD algorithms, such as AFE, AMR and G.729B, and other recently proposed systems, revealing promising results under real-world noise conditions. The main advantage of the proposed approach is its capacity to outperform the above methods without the need for any spatial or spectral constraints, which makes it both versatile and capable of further improvement. Juan E. Rubio, Kentaro Ishizuka, Hiroshi Sawada, Shoko Araki, Tomohiro Nakatani, Masakiyo Fujimoto |
ICASSP (4) | 4 |
| 2007 | Measuring Dependence of Bin-wise Separated Signals for Permutation Alignment in Frequency-domain BSSabstractThis paper presents a new method for grouping bin-wise separated signals for individual sources, i.e., solving the permutation problem, in the process of frequency-domain blind source separation. Conventionally, the correlation coefficient of separated signal envelopes is calculated to judge whether or not the separated signals originate from the same source. In this paper, we propose a new measure that represents the dominance of the separated signal in the mixtures, and use it for calculating the correlation coefficient, instead of a signal envelope. Such dominance measures exhibit dependence/independence more clearly than traditionally used signal envelopes. Consequently, a simple clustering algorithm with centroids works well for grouping separated signals. Experimental results were very appealing, as three sources including two coming from the same direction were separated properly with the new method. Hiroshi Sawada, Shoko Araki, Shoji Makino |
ISCAS | 2 |
| 2007 | Underdetermined blind sparse source separation for arbitrarily arranged multiple sensors
Shoko Araki, Hiroshi Sawada, Ryo Mukai, Shoji Makino |
Signal Process. | 1 |
| 2007 | Geometrically Constrained Independent Component AnalysisabstractAcoustical signals are often corrupted by other speeches, sources, and background noise. This makes it necessary to use some form of preprocessing so that signal processing systems such as a speech recognizer or machine diagnosis can be effectively employed. In this contribution, we introduce and evaluate a new algorithm that uses independent component analysis (ICA) with a geometrical constraint [constrained ICA (CICA)]. It is based on the fundamental similarity between an adaptive beamformer and blind source separation with ICA, and does not suffer the permutation problem of ICA-algorithms. Unlike conventional ICA algorithms, CICA needs prior knowledge about the rough direction of the target signal. However, it is more robust against an erroneous estimation of the target direction than adaptive beamformers: CICA converges to the right solution as long as its look direction is closer to the target signal than to the jammer signal. A high degree of robustness is very important since the geometrical prior of an adaptive beamformer is always roughly estimated in a reverberant environment, even when the look direction is precise. The effectiveness and robustness of the new algorithms is proven theoretically, and shown experimentally for three sources and three microphones with several sets of real-world data Mirko Knaak, Shoko Araki, Shoji Makino |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Grouping Separated Frequency Components by Estimating Propagation Model Parameters in Frequency-Domain Blind Source SeparationabstractThis paper proposes a new formulation and optimization procedure for grouping frequency components in frequency-domain blind source separation (BSS). We adopt two separation techniques, independent component analysis (ICA) and time–frequency (T–F) masking, for the frequency-domain BSS. With ICA, grouping the frequency components corresponds to aligning the permutation ambiguity of the ICA solution in each frequency bin. With T–F masking, grouping the frequency components corresponds to classifying sensor observations in the time–frequency domain for individual sources. The grouping procedure is based on estimating anechoic propagation model parameters by analyzing ICA results or sensor observations. More specifically, the time delays of arrival and attenuations from a source to all sensors are estimated for each source. The focus of this paper includes the applicability of the proposed procedure for a situation with wide sensor spacing where spatial aliasing may occur. Experimental results show that the proposed procedure effectively separates two or three sources with several sensor configurations in a real room, as long as the room reverberation is moderately low. Hiroshi Sawada, Shoko Araki, Ryo Mukai, Shoji Makino |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Doa Estimation for Multiple Sparse Sources with Normalized Observation Vector ClusteringabstractThis paper presents a new method for estimating the direction of arrival (DOA) of source signals whose number N can exceed the number of sensors M. Subspace based methods, e.g., the MUSIC algorithm, have been widely studied, however, they are only applicable when M > N. Another conventional independent component analysis based method allows M ges N, however, it cannot be applied when M60= 120 ms) Shoko Araki, Hiroshi Sawada, Ryo Mukai, Shoji Makino |
ICASSP (5) | 1 |
| 2006 | Blind Source Separation of Many Signals in the Frequency DomainabstractThis paper describes the frequency-domain blind source separation (BSS) of convolutively mixed acoustic signals using independent component analysis (ICA). The most critical issue related to frequency domain BSS is the permutation problem. This paper presents two methods for solving this problem. Both methods are based on the clustering of information derived from a separation matrix obtained by ICA. The first method is based on direction of arrival (DOA) clustering. This approach is intuitive and easy to understand. The second method is based on normalized basis vector clustering. This method is less intuitive than the DOA based method, but it has several advantages. First, it does not need sensor array geometry information. Secondly, it can fully utilize the information contained in the separation matrix, since the clustering is performed in high-dimensional space. Experimental results show that our methods realize BSS in various situations such as the separation of many speech signals located in a 3-dimensional space, and the extraction of primary sound sources surrounded by many background interferences Ryo Mukai, Hiroshi Sawada, Shoko Araki, Shoji Makino |
ICASSP (5) | 3 |
| 2006 | Solving the Permutation Problem of Frequency-Domain BSS when Spatial Aliasing Occurs with Wide Sensor SpacingabstractThis paper describes a method for solving the permutation problem of frequency-domain blind source separation (BSS). The method analyzes the mixing system information estimated with independent component analysis (ICA). When we use widely spaced sensors or increase the sampling rate, spatial aliasing may occur for high frequencies due to the possibility of multiple cycles in the sensor spacing. In such cases, the estimated information would imply multiple possibilities for a source location. This causes some difficulty when analyzing the information. We propose a new method designed to overcome this difficulty. This method first estimates the model parameters for the mixing system at low frequencies where spatial aliasing does not occur, and then refines the estimations by using data at all frequencies. This refinement leads to precise parameter estimation and therefore precise permutation alignment. Experimental results show the effectiveness of the new method Hiroshi Sawada, Shoko Araki, Ryo Mukai, Shoji Makino |
ICASSP (5) | 2 |
| 2006 | Underdetermined sparse source separation of convolutive mixtures with observation vector clusteringabstractWe propose a new method for solving the underdetermined sparse signal separation problem. Some sparseness based methods have already been proposed. However, most of these methods utilized a linear sensor array (or only two sensors), and therefore they have certain limitations; e.g., they cannot separate symmetrically positioned sources. To allow the use of more than three sensors that can be arranged in a non-linear/non-uniform way, we propose a new method that includes the normalization and clustering of the observation vectors. Our proposed method can handle both underdetermined case and (over-)determined cases. We show practical results for speech separation with non-linear/non-uniform sensor arrangements. We obtained promising experimental results for the cases of 3 times 4, 4 times 5 (#sensors times #sources) in a room (RT60= 120 ms) Shoko Araki, Hiroshi Sawada, Ryo Mukai, Shoji Makino |
ISCAS | 1 |
| 2006 | Blind Extraction of Dominant Target Sources Using ICA and Time-Frequency MaskingabstractThis paper presents a method for enhancing target sources of interest and suppressing other interference sources. The target sources are assumed to be close to sensors, to have dominant powers at these sensors, and to have non-Gaussianity. The enhancement is performed blindly, i.e., without knowing the position and active time of each source. We consider a general case where the total number of sources is larger than the number of sensors, and neither the number of target sources nor the total number of sources is known. The method is based on a two-stage process where independent component analysis (ICA) is first employed in each frequency bin and then time-frequency masking is used to improve the performance further. We propose a new sophisticated method for deciding the number of target sources and then selecting their frequency components. We also propose a new criterion for specifying time-frequency masks. Experimental results for simulated cocktail party situations in a room, whose reverberation time was 130 ms, are presented to show the effectiveness and characteristics of the proposed method Hiroshi Sawada, Shoko Araki, Ryo Mukai, Shoji Makino |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | Reducing musical noise by a fine-shift overlap-add method applied to source separation using a time-frequency maskabstractMusical noise is a typical problem with blind source separation using a time-frequency mask. We report that a fine-shift and overlap-add method reduces the musical noise without degrading the separation performance. The effectiveness was confirmed by results of a listening test undertaken in a room with a reverberation time of RT/sub 60/=130 ms. Shoko Araki, Shoji Makino, Hiroshi Sawada, Ryo Mukai |
ICASSP (3) | 1 |
| 2005 | Blind extraction of a dominant source signal from mixtures of many sources [audio source separation applications]abstractThis paper presents a method for enhancing a dominant target source that is close to sensors, and suppressing other interferences. The enhancement is performed blindly, i.e. without knowing the number of total sources or information about each source, such as position and active time. We consider a general case where the number of sources is larger than the number of sensors. We employ a two-stage processing technique where a spatial filter is first employed in each frequency bin and time-frequency masking is then used to improve the performance further. To obtain the spatial filter we employ independent component analysis and then select the component of the target source. Time-frequency masks in the second stage are obtained by calculating the angle between the basis vector corresponding to the target source and a sample vector. The experimental results for a simulated cocktail party situation were very encouraging. Hiroshi Sawada, Shoko Araki, Ryo Mukai, Shoji Makino |
ICASSP (3) | 2 |
| 2004 | Underdetermined blind separation for speech in real environments with sparseness and ICAabstractIn this paper, we propose a method for separating speech signals when there are more signals than sensors. Several methods have already been proposed for solving the underdetermined problem, and some of these utilize the sparseness of speech signals. These methods employ binary masks to extract the signals, and therefore, their extracted signals contain loud musical noise. To overcome this problem, we propose combining a sparseness approach and independent component analysis (ICA). First, using sparseness, we estimate the time points when only one source is active. Then, we remove this single source from the observations and apply ICA to the remaining mixtures. Experimental results show that our proposed sparseness and ICA (SPICA) method can separate signals with little distortion even in reverberant conditions of T/sub R/=130 and 200 ms. Shoko Araki, Shoji Makino, Audrey Blin, Ryo Mukai, Hiroshi Sawada |
ICASSP (3) | 1 |
| 2004 | A sparseness-mixing matrix estimation (SMME) solving the underdetermined BSS for convolutive mixturesabstractWe propose a method for blindly separating real environment speech signals with as little distortion as possible in the special case where speech signals outnumber sensors. Our idea consists in combining sparseness with the use of an estimated mixing matrix. First, we use a geometrical approach to perform a preliminary separation and to detect when only one source is active. This information is then used to estimate the mixing matrix. Then we remove one source from the observations and separate the residual signals with the inverse of the estimated mixing matrix. Experimental results in a real environment (T/sub R/=130 ms and 200 ms) show that our proposed method, which we call sparseness-mixing matrix estimation (SMME), provides separated signals of better quality than those extracted by using only the sparseness property of the speech signal. Audrey Blin, Shoko Araki, Shoji Makino |
ICASSP (4) | 2 |
| 2004 | Near-field frequency domain blind source separation for convolutive mixturesabstractThe paper presents a method for solving the permutation problem of frequency domain blind source separation (BSS) when source signals come from the same or similar directions. Geometric information such as the direction of arrival (DOA) is helpful for solving the permutation problem, and a combination of the DOA based and correlation based methods provides a robust and precise solution. However, when signals come from similar directions, the DOA based approach fails, and we have to use only the correlation based method whose performance is unstable. We show that an interpretation of the ICA solution by a near-field model yields information about spheres on which source signals exist, which can be used as an alternative to the DOA. Experimental results show that the proposed method can robustly separate a mixture of signals arriving from the same direction. Ryo Mukai, Hiroshi Sawada, Shoko Araki, Shoji Makino |
ICASSP (4) | 3 |
| 2004 | Convolutive blind source separation for more than two sources in the frequency domainabstractBlind source separation (BSS) for convolutive mixtures can be efficiently achieved in the frequency domain, where independent component analysis is performed separately in each frequency bin. However, frequency-domain BSS involves a permutation problem, which is well known as a difficult problem, especially when the number of sources is large. This paper presents a method for solving the permutation problem, which works well even for many sources. The successful solution for the permutation problem highlights another problem with frequency-domain BSS that arises from the circularity of the discrete frequency representation. This paper discusses the phenomena of the problem and presents a method for solving it. With these two methods, we can separate many sources with a practical execution time. Moreover, real-time processing is currently possible for up to three sources with our implementation. Hiroshi Sawada, Ryo Mukai, Shoko Araki, Shoji Makino |
ICASSP (3) | 3 |
| 2004 | A robust and precise method for solving the permutation problem of frequency-domain blind source separationabstractBlind source separation (BSS) for convolutive mixtures can be solved efficiently in the frequency domain, where independent component analysis (ICA) is performed separately in each frequency bin. However, frequency-domain BSS involves a permutation problem: the permutation ambiguity of ICA in each frequency bin should be aligned so that a separated signal in the time-domain contains frequency components of the same source signal. This paper presents a robust and precise method for solving the permutation problem. It is based on two approaches: direction of arrival (DOA) estimation for sources and the interfrequency correlation of signal envelopes. We discuss the advantages and disadvantages of the two approaches, and integrate them to exploit their respective advantages. Furthermore, by utilizing the harmonics of signals, we make the new method robust even for low frequencies where DOA estimation is inaccurate. We also present a new closed-form formula for estimating DOAs from a separation matrix obtained by ICA. Experimental results show that our method provided an almost perfect solution to the permutation problem for a case where two sources were mixed in a room whose reverberation time was 300 ms. Hiroshi Sawada, Ryo Mukai, Shoko Araki, Shoji Makino |
IEEE Trans. Speech Audio Process. | 3 |
| 2003 | Subband based blind source separation for convolutive mixtures of speechabstractSubband processing is applied to blind source separation (BSS) for convolutive mixtures of speech. This is motivated by the drawback of frequency-domain BSS, i.e., when a long frame with a fixed frame-shift is used to cover reverberation, the number of samples in each frequency decreases and the separation performance is degraded. In our proposed subband BSS, (1) by using a moderate number of subbands, a sufficient number of samples can be held in each subband, mid (2) by using FIR filters in each subband, we can handle long reverberation. Subband BSS achieves better performance than frequency-domain BSS. Moreover, we propose efficient separation procedures that take into consideration the frequency characteristics of room reverberation and speech signals. We achieve this (3) by using longer unmixing filters in low frequency bands, and (4) by adopting overlap-blockshift in BSS's batch adaptation in low frequency bands. Consequently, frequency-dependent subband processing is successfully realized in the proposed subband BSS. Shoko Araki, Shoji Makino, Robert Aichner, Tsuyoki Nishikawa, Hiroshi Saruwatari |
ICASSP (5) | 1 |
| 2003 | Geometrically constraint ICA for convolutive mixtures of soundabstractThe goal of this contribution is a new algorithm using independent component analysis with a geometrical constraint. The new algorithm solves the permutation problem of blind source separation of acoustic mixtures, and it is significantly less sensitive to the precision of the geometrical constraint than an adaptive beamformer. A high degree of robustness is very important since the steering vector is always roughly estimated in the reverberant environment, even when the look direction is precise. The new algorithm is based on FastICA and constrained optimization. It is theoretically and experimentally analyzed with respect to the roughness of the steering vector estimation by using impulse responses of real room. The effectiveness of the algorithms for real-world mixtures is also shown in the case of three sources and three microphones. Mirko Knaak, Shoko Araki, Shoji Makino |
ICASSP (2) | 2 |
| 2003 | Robust real-time blind source separation for moving speakers in a roomabstractThis paper describes a robust real-time blind source separation (BSS) method for moving speech signals in a room. Our method employs frequency domain independent component analysis (ICA) using a blockwise batch algorithm in the first stage, and the separated signals are refined by postprocessing using crosstalk component estimation and nonstationary spectral subtraction in the second stage. The blockwise batch algorithm achieves better performance than an online algorithm when sources are fixed, and the postprocessing compensates for performance degradation caused by source movement. Experimental results using speech signals recorded in a real room show that the proposed method realizes robust real-time separation for moving sources. Our method is implemented on a standard PC and works in real time. Ryo Mukai, Hiroshi Sawada, Shoko Araki, Shoji Makino |
ICASSP (5) | 3 |
| 2003 | A robust approach to the permutation problem of frequency-domain blind source separationabstractThis paper presents a robust and precise method for solving the permutation problem of frequency-domain blind source separation. It is based on two previous approaches: the direction of arrival estimation approach and the inter-frequency correlation approach. We discuss the advantages and disadvantages of the two approaches, and integrate them to exploit the both advantages. We also present a closed form formula to calculate a null direction, which is used in estimating the directions of source signals. Experimental results show that our method solved permutation problems almost perfectly for a situation that two sources were mixed in a room whose reverberation time was 300 ms. Hiroshi Sawada, Ryo Mukai, Shoko Araki, Shoji Makino |
ICASSP (5) | 3 |
| 2003 | The fundamental limitation of frequency domain blind source separation for convolutive mixtures of speechabstractDespite several recent proposals to achieve blind source separation (BSS) for realistic acoustic signals, the separation performance is still not good enough. In particular, when the impulse responses are long, performance is highly limited. In this paper, we consider a two-input, two-output convolutive BSS problem. First, we show that it is not good to be constrained by the condition T>P, where T is the frame length of the DFT and P is the length of the room impulse responses. We show that there is an optimum frame size that is determined by the trade-off between maintaining the number of samples in each frequency bin to estimate statistics and covering the whole reverberation. We also clarify the reason for the poor performance of BSS in long reverberant environments, highlighting that the framework of BSS works as two sets of frequency-domain adaptive beamformers. Although BSS can reduce reverberant sounds to some extent like adaptive beamformers, they mainly remove the sounds from the jammer direction. This is the reason for the difficulty of BSS in reverberant environments. Shoko Araki, Ryo Mukai, Shoji Makino, Tsuyoki Nishikawa, Hiroshi Saruwatari |
IEEE Trans. Speech Audio Process. | 1 |
| 2002 | Equivalence between frequency domain blind source separation and frequency domain adaptive beamformingabstractFrequency domain Blind Source Separation (BSS) is shown to be equivalent to two sets of frequency domain adaptive microphone arrays, i.e., Adaptive Beamformers (ABFs). The minimization of the off-diagonal components in the BSS update equation can be viewed as the minimization of the mean square error in the ABF. The unmixing matrix of the BSS and the filter coefficients of the ABF converge to the same solution in the mean square error sense if the two source signals are ideally independent. Therefore, the performance of the BSS is limited by that of the ABF. This understanding. gives an interpretation of BSS from physical point of view. Shoko Araki, Yoichi Hinamoto, Shoji Makino, Tsuyoki Nishikawa, Ryo Mukai, Hiroshi Saruwatari |
ICASSP | 1 |
| 2002 | Removal of residual cross-talk components in Blind Source Separation using time-delayed spectral subtractionabstractThis paper describes a post processing method to refine output signals obtained by Blind Source Separation (BSS). The performance of BSS using Independent Component Analysis (ICA) declines significantly in a reverberant environment. The degradation is mainly caused by the cross-talk components derived from the reverberation of the jammer signal. Utilizing this knowledge, we propose a new method, time-delayed non-stationary spectral subtraction, which removes the residual components from the separated signals precisely. The proposed method compensates for the weakness of BSS in a reverberant environment. Experimental results using speech signals show that the proposed method improves the signal-to-noise ratio by 3 to 5 dB. Ryo Mukai, Shoko Araki, Hiroshi Sawada, Shoji Makino |
ICASSP | 2 |
| 2002 | Polar coordinate based nonlinear function for frequency-domain blind source separationabstractThis paper presents a new type of nonlinear function for independent component analysis to process complex-valued signals, which is used in frequency-domain blind source separation. The new function is based on the polar coordinates of a complex number, whereas the conventional one is based on the Cartesian coordinates. The new function is derived from the probability density function of frequency-domain signals that are assumed to be independent of the phase. We show that the difference between the two types of functions is in the assumed densities of independent components. Experimental results for separating speech signals show that the new nonlinear function behaves better than the conventional one. Hiroshi Sawada, Ryo Mukai, Shoko Araki, Shoji Makino |
ICASSP | 3 |
| 2001 | Fundamental limitation of frequency domain blind source separation for convolutive mixture of speechabstractDespite several recent proposals to achieve blind source separation (BSS) for realistic acoustic signals, separation performance is still not good enough. In particular, when the length of impulse response is long, performance is highly limited. We show it is useless to be constrained by the condition, P /spl Lt/ T, where T is the frame size of FFT and P is the length of room impulse response. From our experiments. a frame size of 256 or 512 (32 or 64 ms at a sampling frequency of 8 kHz) is best even for the long room reverberation of T/sub R/ = 150 and 300 ms. We also clarified the reason for poor performance of BSS in a long reverberant environment, finding that separation is achieved chiefly for the sound from the direction of jammers because BSS cannot calculate the inverse of the room transfer function both for the target and jammer signals. Shoko Araki, Shoji Makino, Tsuyoki Nishikawa, Hiroshi Saruwatari |
ICASSP | 1 |
| 2001 | Equivalence between frequency domain blind source separation and frequency domain adaptive null beamformersabstractFrequency domain Blind Source Separation (BSS) is shown to be equivalent to two sets of frequency domain adaptive microphone arrays, that is, Adaptive Null Beamformers (ANB). The unmixing matrix of the BSS and the filter coefficients of the ANB converge to the same solution in the mean square error sense if the two source signals are ideally independent. This understanding clearly explains the poor performance of the BSS in a real room with long reverberation. The fundamental difference exists in the adaptation period when they should adapt. That is, the ANB can adapt in the presence of a jammer but the absence of a target, whereas the BSS can adapt in the presence of a target and jammer, and also in the presence of only a target. Shoko Araki, Shoji Makino, Ryo Mukai, Hiroshi Saruwatari |
INTERSPEECH | 1 |
| 2001 | Separation and dereverberation performance of frequency domain blind source separation for speech in a reverberant environmentabstractIn this paper, we investigate the separation and dereverberation performance of frequency domain Blind Source Separation (BSS) based on Independent Component Analysis (ICA) by measuring impulse responses of a system. Since ICA is a statistical method, i.e., it only attempts to make outputs independent, it is not easy to predict what is going on in a BSS system physically. We therefore investigate the detailed components in the processed signals of a whole BSS system from a physical and acoustical viewpoint. In particular, we focus on the direct sound and reverberation in the target and jammer signals. As a result, we reveal that the direct sound of a jammer can be removed and the reverberation of the jammer can be reduced to some degree by BSS, while the reverberation of the target cannot be reduced. Moreover, we show that a long frame length causes pre-echo noise, and this damages the quality of the separated signal. 1. Ryo Mukai, Shoko Araki, Shoji Makino |
INTERSPEECH | 2 |