Masahito Togami

dblp:98/3376 · DBLP profile ↗
← Back
44ranked-venue papers
22as first author
8since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 39 · 20 first-author · 8 since 2021Artificial intelligence and machine learning · 17 · 5 first-author · 2 since 2021Systems, architecture and hardware · 1
YearPublicationVenuePosition
2024 Real-Time Stereo Speech Enhancement with Spatial-Cue Preservation Based on Dual-Path Structure
abstract
We introduce a real-time, multichannel speech enhancement algorithm which maintains the spatial cues of stereo recordings including two speech sources. Recognizing that each source has unique spatial information, our method utilizes a dual-path structure, ensuring the spatial cues remain unaffected during enhancement by applying source-specific common-band gain. This method also seamlessly integrates pretrained monaural speech enhancement, eliminating the need for retraining on stereo inputs. Source separation from stereo mixtures is achieved via spatial beamforming, with the steering vector for each source being adaptively updated using post-enhancement output signal. This ensures accurate tracking of the spatial information. The final stereo output is derived by merging the spatial images of the enhanced sources, with its efficacy not heavily reliant on the separation performance of the beamforming. The algorithm runs in real-time on 10-ms frames with a 40 ms of look-ahead. Evaluations reveal its effectiveness in enhancing speech and preserving spatial cues in both fully and sparsely overlapped mixtures.
Masahito Togami, Jean-Marc Valin, Karim Helwani, Ritwik Giri, Umut Isik, Michael M. Goodwin
ICASSP1
2022 Computationally-Efficient Overdetermined Blind Source Separation Based on Iterative Source Steering
abstract
This paper describesa computationally-efficient optimization algorithm for the blind source separation (BSS) of overdetermined mixtures. In the determined case, a matrix-inversion-free iterative source steering (ISS) algorithm has been proposed for estimating a square demixing matrix as a computationally-efficient alternative to the popular iterative projection (IP) algorithm. The IP algorithm is based on source-wise (i.e., row-wise) updates of the demixing matrix, and lends itself naturally to an extension to overdetermined independent vector analysis (IVA) called OverIVA. In contrast, the ISS algorithm changes the whole demixing matrix at every update, making its extension to the overdetermined case non-trivial. In this paper, we propose a modified ISS algorithm for OverIVA fully exploiting the computational savings of ISS. We also derive an overdetermined extension of independent low-rank matrix analysis (OverILRMA) with the modified ISS algorithm. Experimental results showed that the proposed ISS-based OverIVA and OverILRMA were comparable or superior to the conventional IP-based counterparts in speech separation performance while achieving lower computational cost.
Yicheng Du, Robin Scheibler, Masahito Togami, Kazuyoshi Yoshii, Tatsuya Kawahara
IEEE Signal Process. Lett.3
2021 Joint Dereverberation and Separation With Iterative Source Steering
abstract
We propose a new algorithm for joint dereverberation and blind source separation (DR-BSS). Our work builds upon the IRLMA-T framework that applies a unified filter combining dereverberation and separation. One drawback of this framework is that it requires several matrix inversions, an operation inherently costly and with potential stability issues. We leverage the recently introduced iterative source steering (ISS) updates to propose two algorithms mitigating this issue. Albeit derived from first principles, the first algorithm turns out to be a natural combination of weighted prediction error (WPE) dereverberation and ISS-based BSS, applied alternatingly. In this case, we manage to reduce the number of matrix inversion to only one per iteration and source. The second algorithm updates the ILRMA-T matrix using only sequential ISS updates requiring no matrix inversion at all. Its implementation is straightforward and memory efficient. Numerical experiments demonstrate that both methods achieve the same final performance as ILRMA-T in terms of several relevant objective metrics. In the important case of two sources, the number of iterations required is also similar.
Taishi Nakashima, Robin Scheibler, Masahito Togami, Nobutaka Ono
ICASSP3
2021 Surrogate Source Model Learning for Determined Source Separation
abstract
We propose to learn surrogate functions of universal speech priors for determined blind speech separation. Deep speech priors are highly desirable due to their superior modelling power, but are not compatible with state-of-the-art independent vector analysis based on majorization-minimization (AuxIVA), since deriving the required surrogate function is not easy, nor always possible. Instead, we do away with exact majorization and directly approximate the surrogate. Taking advantage of iterative source steering (ISS) updates, we back propagate the permutation invariant separation loss through multiple iterations of AuxIVA. ISS lends itself well to this task due to its lower complexity and lack of matrix inversion. Experiments show large improvements in terms of scale invariant signal-to-distortion (SDR) ratio and word error rate compared to baseline methods. Training is done on two speakers mixtures and we experiment with two losses, SDR and coherence. We find that the learnt approximate surrogate generalizes well on mixtures of three and four speakers without any modification. We also demonstrate generalization to a different variation of the AuxIVA update equations. The SDR loss leads to fastest convergence in iterations, while coherence leads to the lowest word error rate (WER). We obtain as much as 36 % reduction in WER.
Robin Scheibler, Masahito Togami
ICASSP2
2021 Refinement of Direction of Arrival Estimators by Majorization-Minimization Optimization on the Array Manifold
abstract
We propose a generalized formulation of direction of arrival estimation that includes many existing methods such as steered response power, subspace, coherent and incoherent, as well as speech sparsity-based methods. Unlike most conventional methods that rely exclusively on grid search, we introduce a continuous optimization algorithm to refine DOA estimates beyond the resolution of the initial grid. The algorithm is derived from the majorization-minimization (MM) technique. We derive two surrogate functions, one quadratic and one linear. Both lead to efficient iterative algorithms that do not require hyperparameters, such as step size, and ensure that the DOA estimates never leave the array manifold, without the need for a projection step. In numerical experiments, we show that the accuracy after a few iterations of the MM algorithm nearly removes dependency on the resolution of the initial grid used. We find that the quadratic surrogate function leads to very fast convergence, but the simplicity of the linear algorithm is very attractive, and the performance gap small.
Robin Scheibler, Masahito Togami
ICASSP2
2021 End To End Learning For Convolutive Multi-Channel Wiener Filtering
abstract
In this paper, we propose a dereverberation and speech source separation method based on deep neural network (DNN). Unlike the cascade connection of dereverberation and speech source separation, the proposed method performs dereverberation and speech source separation jointly by a unified convolutive multi-channel Wiener filtering (CMWF). The proposed method adopts a time-varying CMWF to achieve more dereverberation and separation performance than a time-invariant CMWF. The time-varying CMWF requires time-frequency masks and time-frequency activities. These variables are inferred via a unified DNN. The DNN is trained to optimize the output signal of the time-varying CMWF with a loss function based on a negative log-posterior probability density function. We also reveal that the time-varying CMWF can be obtained efficiently based on the Sherman-Morrison-Woodbury equation. Experimental results show that the proposed time-varying CMWF can separate speech sources under reverberant environments better than the cascade-connection based method and the time-invariant CMWF.
Masahito Togami
ICASSP1
2021 Efficient and Stable Adversarial Learning Using Unpaired Data for Unsupervised Multichannel Speech Separation
Yu Nakagome, Masahito Togami, Tetsuji Ogawa, Tetsunori Kobayashi
Interspeech2
2021 Sound Source Localization with Majorization Minimization
Masahito Togami, Robin Scheibler
Interspeech1
2020 Scene-Dependent Acoustic Event Detection with Scene Conditioning and Fake-Scene-Conditioned Loss
abstract
In this paper, we propose scene-dependent acoustic event detection (AED) with scene conditioning and fake-scene-conditioned loss. The proposed method employs a multitask network, that has not only AED part but also acoustic scene classification (ASC). The scenes predicted by ASC are employed as an additional feature for scene conditioning of AED to learn the relationship between scenes and events. For efficient training, the proposed method incorporates a new AED loss function, which is the fake-scene-conditioned loss, in addition to the conventional AED loss. Upon training, the AED part is conditioned with fake scenes as well as predicted and true scenes. The fake-scene-conditioned loss is calculated between the fake-scene-conditioned AED results and labels of events that do not exist in the fake scenes are removed. Whereas training with combinations of true scenes/events, i.e., the conventional AED loss, only reveals that an event is present in a scene, with fake-scene-conditioned loss, the proposed method can learn that an event is absent in a scene. Experimental results show that the proposed method improves the AED performance compared with the baseline; an increase in the f1 score of 23% and a decrease in the false alarm rate of 56% for scenes where no event exists.
Tatsuya Komatsu, Keisuke Imoto, Masahito Togami
ICASSP3
2020 Consistency-Aware Multi-Channel Speech Enhancement Using Deep Neural Networks
abstract
This paper proposes a deep neural network (DNN)–based multichannel speech enhancement system in which a DNN is trained to maximize the quality of the enhanced time-domain signal. DNN-based multi-channel speech enhancement is often conducted in the time-frequency (T-F) domain because spatial filtering can be efficiently implemented in the T-F domain. In such a case, ordinary objective functions are computed on the estimated T-F mask or spectrogram. However, the estimated spectrogram is often inconsistent, and its amplitude and phase may change when the spectrogram is converted back to the time-domain. That is, the objective function does not evaluate the enhanced time-domain signal properly. To address this problem, we propose to use an objective function defined on the reconstructed time-domain signal. Specifically, speech enhancement is conducted by multi-channel Wiener filtering in the T-F domain, and its result is converted back to the time-domain. We propose two objective functions computed on the reconstructed signal where the first one is defined in the time-domain, and the other one is defined in the T-F domain. Our experiment demonstrates the effectiveness of the proposed system comparing to T-F masking and mask-based beamforming.
Yoshiki Masuyama, Masahito Togami, Tatsuya Komatsu
ICASSP2
2020 Deep Speech Extraction with Time-Varying Spatial Filtering Guided By Desired Direction Attractor
abstract
In this investigation, a deep neural network (DNN) based speech extraction method is proposed to enhance a speech signal propagating from the desired direction. The proposed method integrates knowledge based on a sound propagation model and the time-varying characteristics of a speech source, into a DNN-based separation framework. This approach outputs a separated speech source using time-varying spatial filtering, which achieves superior speech extraction performance compared with time-invariant spatial filtering. Given that the gradient of all modules can be calculated, back-propagation can be performed to maximize the speech quality of the output signal in an end-to-end manner. Guided information is also modeled based on the sound propagation model, which facilitates disentangled representations of the target speech source and noise signals. The experimental results demonstrate that the proposed method can extract the target speech source more accurately than conventional DNN-based speech source separation and conventional speech extraction using time-invariant spatial filtering.
Yu Nakagome, Masahito Togami, Tetsuji Ogawa, Tetsunori Kobayashi
ICASSP2
2020 Multi-Channel Speech Source Separation and Dereverberation With Sequential Integration of Determined and Underdetermined Models
abstract
In this paper, we propose a joint multi-channel speech source separation and dereverberation method in which multiple speech sources and late reverberation are separated in an unsupervised manner. The proposed method jointly optimizes an auto-regressive (AR) model based speech dereverberation and a time-varying multichannel Wiener filtering (MWF) based speech source separation. So as to increase separation and dereverberation performance and to overcome the inter-frequency permutation problem, the proposed method adopts a sequential parameter optimization strategy. At first, the parameter is updated based on a determined model, and the permutation problem can be solved based on the non-negative matrix factorization. The determined model reduces reverberation by only the AR model and residual reverberation remains. Inspired by the fact that a parameter of a determined model can be converted into a parameter of a underdetermined model, the proposed method regards residual reverberation as an additional source and reduces residual reverberation with the converted parameter based on the underdetermined model. We further propose additional update of the parameter based on the underdetermined model. Experimental results show that the proposed method outperforms the conventional method based on only the determined model and the proposed additional parameter update based on the underdetermined model is effective.
Masahito Togami
ICASSP1
2020 Joint Training of Deep Neural Networks for Multi-Channel Dereverberation and Speech Source Separation
abstract
In this paper, we propose a joint training of two deep neural networks (DNNs) for dereverberation and speech source separation. The proposed method connects the first DNN, the dereverberation part, the second DNN, and the speech source separation part in a cascade manner. The proposed method does not train each DNN separately. Instead, an integrated loss function which evaluates an output signal after dereverberation and speech source separation is adopted. The proposed method estimates the output signal as a probabilistic variable. Recently, in the speech source separation context, we proposed a loss function which evaluates the estimated posterior probability density function (PDF) of the output signal. In this paper, we extend this loss function into a loss function which evaluates not only speech source separation performance but also speech derevereberation performance. Since the output signal of the dereverberation part is converted into the input feature of the second DNN, gradient of the loss function is back-propagated into the first DNN through the input feature of the second DNN. Experimental results show that the proposed joint training of two DNNs is effective. It is also shown that the posterior PDF based loss function is effective in the joint training context.
Masahito Togami
ICASSP1
2020 Unsupervised Training for Deep Speech Source Separation with Kullback-Leibler Divergence Based Probabilistic Loss Function
abstract
In this paper, we propose a multi-channel speech source separation method with a deep neural network (DNN) which is trained under the condition that no clean signal is available. As an alternative to a clean signal, the proposed method adopts an estimated speech signal by an unsupervised speech source separation method which leverages a statistical model. As a statistical model of microphone input signal, we adopts a time-varying spatial covariance matrix (SCM) model which includes reverberation and background noise submodels so as to achieve robustness against reverberation and background noise. The DNN infers intermediate variables which are needed for constructing the time-varying SCM. Separation is performed in a probabilistic manner so as to avoid overfitting to separation error. Since there are multiple intermediate variables, a loss function which evaluates a single intermediate variable is not applicable. Instead, the proposed method adopts a loss function which evaluates the output probabilistic signal directly based on Kullback-Leibler Divergence (KLD). The gradient of the loss function can be back-propagated into the DNN through all the intermediate variables. Experimental results under reverberant conditions show that the proposed method achieves better results than the conventional methods.
Masahito Togami, Yoshiki Masuyama, Tatsuya Komatsu, Yu Nakagome
ICASSP1
2020 Mentoring-Reverse Mentoring for Unsupervised Multi-Channel Speech Source Separation
Yu Nakagome, Masahito Togami, Tetsuji Ogawa, Tetsunori Kobayashi
INTERSPEECH2
2020 Sparseness-Aware DOA Estimation with Majorization Minimization
Masahito Togami, Robin Scheibler
INTERSPEECH1
2019 Spatial Constraint on Multi-channel Deep Clustering
abstract
In this paper, a multi-channel deep clustering technique which combines two types of spatial information is proposed. The first one is an estimated direction-of-arrival (DOA) at each time-frequency point, which is utilized as an input feature of the proposed neural network. Instead of stacking embeddings of all pairs of microphones as in the conventional multi-channel deep clustering, the proposed method only requires one embedding. Therefore, the computational cost can be reduced in the inference stage. The second one is the time-frequency activity of each speech source estimated by multichannel Wiener filtering (MWF). The MWF is inserted between two consecutive bidirectional long-short-term memory (BLSTM) layers. The estimated time-frequency activity of each speech source by the MWF is transformed into an input feature of the next BLSTM layer. The proposed MWF insertion enhances the consistency of the embedding vectors along the time-axis. Experimental results show that multi-channel deep clustering with the proposed input feature based on the estimated DOA can separate speech sources better than the conventional multi-channel deep clustering that stacks embeddings of all the pairs of the microphones. Furthermore, the proposed MWF insertion is shown to be able to reduce distortion of output signal and improve signal-to-interference ratio.
Masahito Togami
ICASSP1
2019 Multi-channel Itakura Saito Distance Minimization with Deep Neural Network
abstract
A multi-channel speech source separation with a deep neural network which optimizes not only the time-varying variance of a speech source but also the multi-channel spatial covariance matrix jointly without any iterative optimization method is shown. Instead of a loss function which does not evaluate spatial characteristics of the output signal, the proposed method utilizes a loss function based on minimization of multi-channel Itakura-Saito Distance (MISD), which evaluates spatial characteristics of the output signal. The cost function based on MISD is calculated by the estimated posterior probability density function (PDF) of each speech source based on a time-varying Gaussian distribution model. The loss function of the neural network and the PDF of each speech source that is assumed in multi-channel speech source separation are consistent with each other. As a neural-network architecture, the proposed method utilizes multiple bidirectional long-short term memory (BLSTM) layers. The BLSTM layers and the successive complex-valued signal processing are jointly optimized in the training phase. Experimental results show that more accurately separated speech signal can be obtained with neural network parameters optimized based on the proposed MISD minimization than that with neural network parameters optimized based on loss functions without spatial covariance matrix evaluation.
Masahito Togami
ICASSP1
2019 Simultaneous Optimization of Forgetting Factor and Time-frequency Mask for Block Online Multi-channel Speech Enhancement
abstract
In this paper, we propose a block-online multi-channel speech enhancement technique which simultaneously optimizes time-frequency masks and forgetting factors for estimation of multichannel covariance matrices of the desired speech signal and the noise signal so as to maximize speech enhancement performance under the condition that environmental changes occur. The proposed method reduces the noise signal by using a multi-channel Wiener filter (MWF) which is generated by the covariance matrices with the estimated forgetting factors and the estimated time-frequency masks which are outputs of the proposed neural network. The proposed method learns all the parameters of the proposed neural network so as to maximize the speech enhancement performance. Three types of the input features for the forgetting factors adaptation are proposed. The first one is the magnitude spectral of the microphone input signal. The second one is the MWF output with the previous-block filter that is adapted in the previous block. The third one is the inner product between the microphone input signal and the estimated covariance matrices in the previous block. Experimental results show that the proposed method can reduce noise signal more accurately than the conventional equally weight sample averaging.
Masahito Togami
ICASSP1
2019 Multichannel Loss Function for Supervised Speech Source Separation by Mask-Based Beamforming
abstract
In this paper, we propose two mask-based beamforming methods using a deep neural network (DNN) trained by multichannel loss functions. Beamforming technique using time-frequency (TF)-masks estimated by a DNN have been applied to many applications where TF-masks are used for estimating spatial covariance matrices. To train a DNN for mask-based beamforming, loss functions designed for monaural speech enhancement/separation have been employed. Although such a training criterion is simple, it does not directly correspond to the performance of mask-based beamforming. To overcome this problem, we use multichannel loss functions which evaluate the estimated spatial covariance matrices based on the multichannel Itakura--Saito divergence. DNNs trained by the multichannel loss functions can be applied to construct several beamformers. Experimental results confirmed their effectiveness and robustness to microphone configurations.
Yoshiki Masuyama, Masahito Togami, Tatsuya Komatsu
INTERSPEECH2
2019 Variational Bayesian Multi-Channel Speech Dereverberation Under Noisy Environments with Probabilistic Convolutive Transfer Function
Masahito Togami, Tatsuya Komatsu
INTERSPEECH1
2018 Maximizing SLU Performance with Minimal Training Data Using Hybrid RNN Plus Rule-based Approach
abstract
Spoken language understanding (SLU) by using recurrent neural networks (RNN) achieves good performances for large training data sets, but collecting large training datasets is a challenge, especially for new voice applications.Therefore, the purpose of this study is to maximize SLU performances, especially for small training data sets.To this aim, we propose a novel CRF-based dialog act selector which chooses suitable dialog acts from outputs of RNN SLU and rule-based SLU.We evaluate the selector by using DSTC2 corpus when RNN SLU is trained by less than 1,000 training sentences.The evaluation demonstrates the selector achieves Micro F1 better than both RNN and rule-based SLUs.In addition, it shows the selector achieves better Macro F1 than RNN SLU and the same Macro F1 as rule-based SLU.Thus, we confirmed our method offers advantages in SLU performances for small training data sets.
Takeshi Homma, Adriano S. Arantes, Maria Teresa Gonzalez Diaz, Masahito Togami
SIGDIAL Conference4
2016 Adaptive Boolean compressive sensing by using multi-armed bandit
abstract
A new method for solving adaptive Boolean compressive sensing is proposed. By greedy maximization of an expected information gain, a conventional method controls the pool-size for adaptive Boolean compressive sensing. However, the conventional greedy method has the drawback that it has no guarantee of convergence to the optimal strategy. To solve the problem, based on the multi-armed bandit, the proposed method controls the pool-size adaptively. The information gain of the conventional greedy method is rewritten as the reward of the multi-armed bandit, and the multi-armed bandit is introduced into adaptive Boolean compressive sensing. Experimental results indicate that the correct rate of exact recovery of the proposed method converges to 1 fast without prior knowledge about the number of defective items and that the proposed method outperforms the conventional greedy method in the case that the number of defective items is large.
Yohei Kawaguchi, Masahito Togami
ICASSP2
2016 Data Augmentation Using Multi-Input Multi-Output Source Separation for Deep Neural Network Based Acoustic Modeling
Yusuke Fujita, Ryoichi Takashima, Takeshi Homma, Masahito Togami
INTERSPEECH4
2015 Unified ASR system using LGM-based source separation, noise-robust feature extraction, and word hypothesis selection
abstract
In this paper, we propose a unified system that incorporates speech source separation and automatic speech recognition for various noise environments. There are three features in the proposed system. The first feature of the proposed method is the LGM (local Gaussian modeling) based source separation with the efficient permutation alignment method that integrates a power spectrum correlation based method and a direction-of-arrival (DOA) based method. Evaluation results show that using the separated speech with the baseline acoustic modeling method reduces the word error rate (WER) significantly. The second feature of the proposed method is multi-condition training with per-utterance normalized features and noise-aware features in the acoustic modeling step. In this paper, we show that the proposed training method is effective even when an input signal has been distorted through the source separation step. The third feature is the word hypothesis selection method for integrating multiple recognition results. The proposed selection method estimates correct words based on a recognizer's confidence and co-occurrence characteristics. The evaluation results show that the proposed selection method outperforms the conventional recognizer output voting error reduction (ROVER) method. The proposed system is evaluated using the third CHiME challenge dataset. Evaluation results show that the proposed system resulted in an improvement of 66.1% over the baseline system.
Yusuke Fujita, Ryoichi Takashima, Takeshi Homma, Rintaro Ikeshita, Yohei Kawaguchi, Takashi Sumiyoshi, Takashi Endo, Masahito Togami
ASRU8
2015 Variational Bayes state space model for acoustic echo reduction and dereverberation
abstract
In this paper, we propose a simultaneous optimization technique for speech dereverberation, acoustic echo reduction, and noise reduction, which can be utilized even when an analog-to-digital (A/D) converter and a digital-to-analog (D/A) converter are not synchronized. The proposed method utilizes a state-space model in which acoustic echo reduction filters are regarded as a time-varying state-vector due to asynchrony of the A/D converter and the D/A converter. In addition to the state-space model for acoustic echo reduction filters, the proposed method utilizes an additional state-space model in which noiseless multichannel speech signals are regarded as a state vector. By using the second state-space model, we can update the dereverberation filter under noisy environments. To optimize two types of state space models, the proposed method utilizes the variational Bayes framework. Two Kalman smoother based parameter optimization stages are performed alternatively. The proposed method is evaluated by using recorded data in a real teleconferencing room. The experimental results show that the proposed method can reduce acoustic echo signal, speech reverberation, and background noise more effectively than the conventional method by authors even when the A/D converter and the D/A converter are asynchronous.
Masahito Togami
ICASSP1
2014 Frequency domain acoustic echo reduction based on Kalman smoother with time-varying noise covariance matrix
abstract
In this paper, we propose a novel acoustic-echo-reduction technique at a time-frequency domain, which is optimally combined with speech enhancement. Unlike conventional echo reduction techniques which minimizes only residual power of the far-end acoustic echo signal, the proposed method minimizes summation of the residual echo signal and distortion of the near-end speech signal from a minimum mean square error (MMSE) perspective. The proposed method performs echo reduction with speech enhancement and parameter optimization in an iterative manner based on the expectation-maximization (EM) algorithm. The E step is corresponding with the echo reduction and speech enhancement based on the Kalman smoother with a time-varying covariance matrix for the observation noise term, which reflects the time-varying characteristics of speech sources. By using the time-varying covariance matrix, we can enhance speech sources effectively with acoustic echo reduction. Associated with the time-varying covariance matrix, a new optimization scheme of parameters for the M step is derived in this paper. Experimental results with impulse responses which was recorded under a real meeting room show that the proposed method can effectively enhance a near-end speech signal when there are a near-end speech signal and a far-end acoustic echo signal.
Masahito Togami, Yohei Kawaguchi, Ryoichi Takashima
ICASSP1
2014 Simultaneous Optimization of Acoustic Echo Reduction, Speech Dereverberation, and Noise Reduction against Mutual Interference
abstract
We propose an optimized speech enhancement method that combines acoustic echo reduction, speech dereverberation, and noise reduction in a unified framework. Normally, partial optimization of acoustic echo reduction, speech dereverberation, and noise reduction does not lead to total optimization. A cascade method of multiple functions causes mutual interference between these functions and degrades eventual speech enhancement performance. Unlike cascade methods, the proposed method combines all functions to optimize eventual speech enhancement performance based on a unified framework, which is also robust against the mutual interference problem. With the proposed method, in addition to time-invariant linear filters, time-varying filters are used to reduce residual reverberation, residual acoustic echo signal, and background noise signal which cannot be reduced using time-invariant filters. These time-invariant filters and time-varying filters are also optimized based on a unified likelihood function to avoid the mutual interference problem. By combining the time-invariant linear filters and the time-varying filters, the proposed method uses a local Gaussian model with a full-rank covariance matrix and a non-zero average vector as a probabilistic model of the microphone input signal. In the local Gaussian model, non-stationary characteristics of speech sources are considered to effectively enhance speech sources. Under this probabilistic model, all the parameters are optimized simultaneously based on the expectation-maximization algorithm and calculates a minimum mean squared error estimate of a desired signal. The experimental results show that the proposed method is superior to the cascade methods.
Masahito Togami, Yohei Kawaguchi
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 ICA-based acceleration of probabilistic latent component analysis for mass spectrometry-based explosives detection
abstract
We propose a new method to separate mass spectra into components of each chemical compound for explosives detection. The conventional method based on probabilistic latent component analysis (PLCA) is effective because the method can solve the problems of non-negativity and non-orthogonality by using sparsity of the domain of explosives detection. However, the convergence of the method is slow, and the calculation time is long. In order to solve this problem, the proposed method makes use of independent component analysis (ICA) in the initialization process. Experimental results indicate that the convergence of the proposed method is accelerated, and total calculation time is decreased.
Yohei Kawaguchi, Masahito Togami, Hisashi Nagano, Yuichiro Hashimoto, Masuyuki Sugiyama, Yasuaki Takada
ICASSP2
2013 Noise robust speech dereverberation with Kalman smoother
abstract
A speech dereverberation method is proposed that is robust against background noise. In contrast to conventional methods based on the linear prediction of the given microphone input signal, in which the linear prediction coefficients are not fully optimized when there is background noise, the proposed method optimizes the coefficients by linear prediction of the noiseless reverberant speech signal even when there is background noise. The noiseless reverberant speech signal and the parameters are iteratively updated on the basis of the expectation maximization algorithm. In the expectation step, sufficient statistics of latent variables which include noiseless reverberant speech signal are estimated using the Kalman smoother. Unlike the standard Kalman smoother, which uses a time-invariant covariance matrix as a state-transition covariance matrix, the proposed method utilizes a time-varying covariance matrix, enabling it to meet the time-varying speech characteristics. The parameters are updated so that the Q function is increased in the maximization step. Experimental results show that the proposed method is superior to conventional methods under noisy conditions.
Masahito Togami, Yohei Kawaguchi
ICASSP1
2013 Optimized Speech Dereverberation From Probabilistic Perspective for Time Varying Acoustic Transfer Function
abstract
A dereverberation technique has been developed that optimally combines multichannel inverse filtering (MIF), beamforming (BF), and non-linear reverberation suppression (NRS). It is robust against acoustic transfer function (ATF) fluctuations and creates less distortion than the NRS alone. The three components are optimally combined from a probabilistic perspective using a unified likelihood function incorporating two probabilistic models. A multichannel probabilistic source model based on a recently proposed local Gaussian model (LGM) provides robustness against ATF fluctuations of the early reflection. A probabilistic reverberant transfer function model (PRTFM) provides robustness against ATF fluctuations of the late reverberation. The MIF and multichannel under-determined source separation (MUSS) are optimized in an iterative manner. The MIF is designed to reduce the time-invariant part of the late reverberation by using optimal time-weighting with reference to the PRTFM and the LGM. The MUSS separates the dereverberated speech signal and the residual reverberation after the MIF, which can be interpreted as an optimized combination of the BF and the NRS. The parameters of the PRTFM and the LGM are optimized based on the MUSS output. Experimental results show that the proposed method is robust against the ATF fluctuations under both single and multiple source conditions.
Masahito Togami, Yohei Kawaguchi, Ryu Takeda, Yasunari Obuchi, Nobuo Nukaga
IEEE Trans. Speech Audio Process.1
2012 Mass spectra separation for explosives detection by using probabilistic latent component analysis
abstract
We propose a new method to separate mass spectra into components of each chemical compound for explosives detection. In mass spectra, all components have no negative values. However, conventional factor analyses for basis decomposition use no constraints of non-negativity, and we can not apply these methods to mass spectra. The proposed method is based on probabilistic latent component analysis (PLCA). The constraints of non-negativity always hold in PLCA, so that the method is effective for mass spectra. In addition, PLCA is defined in a statistical framework, thus PLCA makes it possible to utilize additional a priori information. Therefore, we introduce sparseness assumptions in the domain of mass spectrometry to PLCA in order to estimate the components more accurately. Experimental results indicate that the proposed method outperforms existing methods.
Yohei Kawaguchi, Masahito Togami, Hisashi Nagano, Yuichiro Hashimoto, Masuyuki Sugiyama, Yasuaki Takada
ICASSP2
2012 Multichannel speech dereverberation and separation with optimized combination of linear and non-linear filtering
abstract
In this paper, we propose a multichannel speech dereverberation and separation technique which is effective even when there are multiple speakers and each speaker's transfer function is time-varying due to fluctuation of the corresponding speaker's head. For robustness against fluctuation, the proposed method optimizes linear filtering with non-linear filtering simultaneously from probabilistic perspective based on a probabilistic reverberant transfer-function model, PRTFM. PRTFM is an extension of the conventional time-invariant transfer-function model under uncertain conditions, and PRTFM can be also regarded as an extension of recently proposed blind local Gaussian modeling. The linear filtering and the non-linear filtering are optimized in MMSE (Minimum Mean Square Error) sense during parameter optimization. The proposed method is evaluated in a reverberant meeting room, and the proposed method is shown to be effective.
Masahito Togami, Yohei Kawaguchi, Ryu Takeda, Yasunari Obuchi, Nobuo Nukaga
ICASSP1
2011 Bidirectional OM-LSA speech estimator for noise robust speech recognition
abstract
A new speech enhancement method using bidirectional speech estimator is introduced. A widely-known speech enhancement method using the optimally-modified log spectral amplitude (OM-LSA) speech estimator is re-modified under the assumption that the frame-synchronous estimation is not essential in some of the speech recognition applications. The new method utilizes two separate flows of the speech gain estimation, one is along the forward direction of time and the other along the backward direction. A simple look-ahead estimation mechanism is also implemented in each flow. By taking the average of these two gains, the speech estimation becomes more robust under various noise conditions. Evaluation experiments using the artificial and real noisy speech data confirm that the speech recognition accuracy can be greatly improved by the proposed method.
Yasunari Obuchi, Ryu Takeda, Masahito Togami
ASRU3
2011 Online speech source separation based on maximum likelihood of local Gaussian modeling
abstract
We propose an online speech source separation method which can separate sources under underdetemined conditions. The proposed method is based on local Gaussian modeling (LGM). At first, we de rive an extended approach of conventional offline speech source separation methods based on LGM, which can separate speech sources in an online manner. The likelihood function of the online LGM based approach (OLGM) is approximately maximized by incremental EM based approach. Additionally, we propose an initialization method of OLGM based on a least squares approach to improve con vergence time . Experimental results show that the proposed method can separate sources effectively even when the number of iterations is small.
Masahito Togami
ICASSP1
2011 ASR for Human-Symbiotic Robot "EMIEW2" with Mechanical Noise and Floor-Level Noise Reduction
Takashi Sumiyoshi, Masahito Togami, Yasunari Obuchi
INTERSPEECH2
2010 Head orientation estimation of a speaker by utilizing kurtosis of a DOA histogram with restoration of distance effect
abstract
In this paper, we propose a head-orientation estimation method from multichannel acoustic signals. Sharpness of a DOA histogram which is extracted by using the sparseness based DOA estimation method varies depending on the head orientation of a speaker. The proposed method utilizes this phenomenon to estimate the head orientation of the speaker. The proposed method uses more than two microphone arrays. In addition to estimation of the speaker location, the proposed method estimates kurtosis of the DOA histogram of each array. Kurtosis is regarded as a measure of sharpness of a DOA histogram in the proposed method. However, kurtosis also depends on the distance between the speaker and the microphone array (distance effect). The distance effect is experimentally revealed by the regression analysis. The head orientation of a speaker is estimated by the restored kurtosis which is free from the distance effect. Experimental results on a reverberant environment show that the proposed method can estimate the head orientation of a speaker more accurately than a conventional head-orientation estimation method.
Masahito Togami, Yohei Kawaguchi
ICASSP1
2010 Turn taking-based conversation detection by using DOA estimation
Yohei Kawaguchi, Masahito Togami, Yasunari Obuchi
INTERSPEECH2
2009 DOA estimation method based on sparseness of speech sources for human symbiotic robots
abstract
In this paper, direction of arrival (DOA) estimation methods (both azimuth and elevation) based on sparseness of human speech, "modified delay-and-sum beamformer based on sparseness (MDSBF)" and "stepwise phase difference restoration (SPIRE)", are introduced for human symbiotic robots. MDSBF can achieve good DOA estimation, whose computational cost is proportional to resolution of azimuth and elevation space. DOA estimation result of SPIRE is less accurate than that of MDSBF, but computational cost is independent of resolution. To achieve more accurate DOA estimation result than SPIRE with small computational cost, we propose a novel DOA estimation method which is combination of MDSBF and SPIRE. In the proposed method, MDSBF with rough resolution is performed prior to SPIRE execution, and SPIRE precisely estimates DOA of sources after MDSBF. Experimental results show that sparseness based methods are superior to conventional methods. The proposed combination method achieved more accurate DOA estimation result than SPIRE with smaller computational cost than MDSBF.
Masahito Togami, Akio Amano, Takashi Sumiyoshi, Yasunari Obuchi
ICASSP1
2009 Subband nonstationary noise reduction based on multichannel spatial prediction under reverberant environments
abstract
We propose a novel non-stationary and convolutive noise reduction method under reverberant environments. Unlike many multichannel noise reduction methods, the proposed method does not need pre knowledge of impulse response or direction of arrival (DOA) of the target source. The proposed method is composed of two processes. On the noise reduction process, the noise component is reduced without the impulse response of the target source. The target source component in the output signal is distorted, but the distortion is removed by the distortion-restoration process. Importantly, possibility of complete noise reduction with no distortion based on the proposed framework is assured by MINT theory. Experimental results under the reverberant environment (RT60≈ 300 ms) show that the proposed method can reduce more noise than the conventional method and the distortion of the target source is not so big.
Masahito Togami, Yohei Kawaguchi, Yasunari Obuchi
ICASSP1
2008 Intentional voice command detection for completely hands-free speech interface in home environments
Yasunari Obuchi, Masahito Togami, Takashi Sumiyoshi
INTERSPEECH2
2007 Stepwise Phase Difference Restoration Method for Sound Source Localization using Multiple Microphone Pairs
abstract
We propose a new methodology of sound source localization named SPIRE (stepwise phase difference restoration) that is able to localize sources even if they are neighboring in a reverberant environment. Localizing sound sources in reverberant environments is difficult, because the variance of the direction of an estimated sound source increases in reverberant environments. The major feature of our proposed method is restoration of a microphone pair's phase difference (M1) by using the phase difference of another microphone pair (M2) under the condition that the distance between MTs microphones is longer than the distance between M2's microphones. This restoration process makes it possible to reduce the variance of an estimated sound source direction and to solve the spatial aliasing problem that occurs with the M1's phase difference. The experimental results in a reverberant environment (reverberation time=about 300 ms) indicate that our proposed method can localize sources even if they are neighboring (even if the difference in the sources' directions equals 10 degree).
Masahito Togami, Takashi Sumiyoshi, Akio Amano
ICASSP (1)1
2006 Basic Design of Human-Symbiotic Robot EMIEW
abstract
We are developing a robot that supports people in their daily lives: a human-symbiotic robot. Such robot must share space with its users, be user-friendly, and be able to assist its users. We have developed a prototype autonomous mobile robot that makes use of a self-balancing two-wheeled mobile system and a body swing mechanism to shift its center of gravity. This allows it to move nimbly at up to 6 km per hour. It also has capabilities to avoid collisions with obstacles for moving safely through complex environments. Distant-speech-recognition and high-quality speech-synthesis technologies enable it to communicate with people naturally (i.e., without special tools). These capabilities were demonstrated at the 2005 World Exposition in Aichi, Japan
Yuji Hosoda, Saku Egawa, Junichi Tamamoto, Kenjiro Yamamoto, Ryousuke Nakamura, Masahito Togami
IROS6
2002 Learning Topological Maps from Sequential Observation and Action Data under Partially Observable Environment
Takehisa Yairi, Masahito Togami, Koichi Hori
PRICAI2