VLDB 2026 Research / reviewers in the wild / expert
Yohei Kawaguchi
dblp:96/8055
· DBLP profile ↗
28ranked-venue papers
5as first author
13since 2021 · last 2025
0000-0002-2329-5441ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 5 first-author · 12 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Domain-Independent Automatic Generation of Descriptive Texts for Time-Series DataabstractDue to scarcity of time-series data annotated with descriptive texts, training a model to generate descriptive texts for time-series data is challenging. In this study, we propose a method to systematically generate domain-independent descriptive texts from time-series data. We identify two distinct approaches for creating pairs of time-series data and descriptive texts: the forward approach and the backward approach. By implementing the novel backward approach, we create the Temporal Automated Captions for Observations (TACO) dataset. Experimental results demonstrate that a contrastive learning based model trained using the TACO dataset is capable of generating descriptive texts for time-series data in novel domains. Kota Dohi, Aoi Ito, Harsh Purohit, Tomoya Nishida, Takashi Endo, Yohei Kawaguchi |
ICASSP | 6 |
| 2025 | LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context
Natsuo Yamashita, Masaaki Yamamoto, Hiroaki Kokubo, Yohei Kawaguchi |
INTERSPEECH | 4 |
| 2024 | Streaming Active Learning for Regression Problems Using Regression via ClassificationabstractOne of the challenges in deploying a machine learning model is that the model’s performance degrades as the operating environment changes. To maintain the performance, streaming active learning is used, in which the model is retrained by adding a newly annotated sample to the training dataset if the prediction of the sample is not certain enough. Although many streaming active learning methods have been proposed for classification problems, few efforts have been made for regression problems, which are often handled in the industrial field. In this paper, we propose to use the regression-via-classification framework for streaming active learning for regression. Regression-via-classification transforms regression problems into classification problems so that streaming active learning methods proposed for classification problems can be applied directly to regression problems. Experimental validation on four real data sets shows that the proposed method can perform regression with higher accuracy at the same annotation cost. Shota Horiguchi, Kota Dohi, Yohei Kawaguchi |
ICASSP | 3 |
| 2024 | Stream-based Active Learning for Anomalous Sound Detection in Machine Condition Monitoring
Tuan Vu Ho, Kota Dohi, Yohei Kawaguchi |
INTERSPEECH | 3 |
| 2023 | Zero-Shot Domain Adaptation of Anomalous Samples for Semi-Supervised Anomaly DetectionabstractSemi-supervised anomaly detection (SSAD) is a task where normal data and a limited number of anomalous data are available for training. In practical situations, SSAD methods suffer adapting to domain shifts, since anomalous data are unlikely to be available for the target domain in the training phase. To solve this problem, we propose a domain adaptation method for SSAD where no anomalous data are available for the target domain. First, we introduce a domain-adversarial network to a variational auto-encoder-based SSAD model to obtain domain-invariant latent variables. Since the decoder cannot reconstruct the original data solely from domain-invariant latent variables, we conditioned the decoder on the domain label. To compensate for the missing anomalous data of the target domain, we introduce an importance sampling-based weighted loss function that approximates the ideal loss function. Experimental results indicate that the proposed method helps adapt SSAD models to the target domain when no anomalous data are available for the target domain. Tomoya Nishida, Takashi Endo, Yohei Kawaguchi |
ICASSP | 3 |
| 2023 | CAPTDURE: Captioned Sound Dataset of Single Sources
Yuki Okamoto, Kanta Shimonishi, Keisuke Imoto, Kota Dohi, Shota Horiguchi, Yohei Kawaguchi |
INTERSPEECH | 6 |
| 2023 | Anomalous Sound Detection Based on Sound Separation
Kanta Shimonishi, Kota Dohi, Yohei Kawaguchi |
INTERSPEECH | 3 |
| 2023 | Online Neural Diarization of Unlimited Numbers of Speakers Using Global and Local AttractorsabstractA method to perform offline and online speaker diarization for an unlimited number of speakers is described in this paper. End-to-end neural diarization (EEND) has achieved overlap-aware speaker diarization by formulating it as a multi-label classification problem. It has also been extended for a flexible number of speakers by introducing speaker-wise attractors. However, the output number of speakers of attractor-based EEND is empirically capped; it cannot deal with cases where the number of speakers appearing during inference is higher than that during training because its speaker counting is trained in a fully supervised manner. Our method, EEND-GLA, solves this problem by introducing unsupervised clustering into attractor-based EEND. In the method, the input audio is first divided into short blocks, then attractor-based diarization is performed for each block, and finally, the results of each block are clustered on the basis of the similarity between locally-calculated attractors. While the number of output speakers is limited within each block, the total number of speakers estimated for the entire input can be higher than the limitation. To use EEND-GLA in an online manner, our method also extends the speaker-tracing buffer, which was originally proposed to enable online inference of conventional EEND. We introduce a block-wise buffer update to make the speaker-tracing buffer compatible with EEND-GLA. Finally, to improve online diarization, our method improves the buffer update method and revisits the variable chunk-size training of EEND. The experimental results demonstrate that EEND-GLA can perform speaker diarization of an unseen number of speakers in both offline and online inferences. Shota Horiguchi, Shinji Watanabe 0001, L. Paola García-Perera, Yuki Takashima, Yohei Kawaguchi |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | Multi-Channel End-To-End Neural Diarization with Distributed MicrophonesabstractRecent progress on end-to-end neural diarization (EEND) has en-abled overlap-aware speaker diarization with a single neural net-work. This paper proposes to enhance EEND by using multi-channel signals from distributed microphones. We replace Transformer en-coders in EEND with two types of encoders that process a multi-channel input: spatio-temporal and co-attention encoders. Both are independent of the number and geometry of microphones and suitable for distributed microphone settings. We also propose a model adaptation method using only single-channel recordings. With simulated and real-recorded datasets, we demonstrated that the proposed method outperformed conventional EEND when a multi-channel in-put was given while maintaining comparable performance with a single-channel input. We also showed that the proposed method performed well even when spatial information is inoperative given multi-channel inputs, such as in hybrid meetings in which the utterances of multiple remote participants are played back from the same loudspeaker. Shota Horiguchi, Yuki Takashima, L. Paola García-Perera, Shinji Watanabe 0001, Yohei Kawaguchi |
ICASSP | 5 |
| 2022 | Environmental Sound Extraction Using Onomatopoeic WordsabstractAn onomatopoeic word, which is a character sequence that phonetically imitates a sound, is effective in expressing characteristics of sound such as duration, pitch, and timbre. We propose an environmental-sound-extraction method using onomatopoeic words to specify the target sound to be extracted. By this method, we estimate a time-frequency mask from an input mixture spectrogram and an onomatopoeic word using a U-Net architecture, then extract the corresponding target sound by masking the spectrogram. Experimental results indicate that the proposed method can extract only the target sound corresponding to the onomatopoeic word and performs better than conventional methods that use sound-event classes to specify the target sound. Yuki Okamoto, Shota Horiguchi, Masaaki Yamamoto, Keisuke Imoto, Yohei Kawaguchi |
ICASSP | 5 |
| 2022 | Updating Only Encoders Prevents Catastrophic Forgetting of End-to-End ASR ModelsabstractIn this paper, we present an incremental domain adaptation technique to prevent catastrophic forgetting for an end-to-end automatic speech recognition (ASR) model.Conventional approaches require extra parameters of the same size as the model for optimization, and it is difficult to apply these approaches to end-to-end ASR models because they have a huge amount of parameters.To solve this problem, we first investigate which parts of end-to-end ASR models contribute to high accuracy in the target domain while preventing catastrophic forgetting.We conduct experiments on incremental domain adaptation from the LibriSpeech dataset to the AMI meeting corpus with two popular end-to-end ASR models and found that adapting only the linear layers of their encoders can prevent catastrophic forgetting.Then, on the basis of this finding, we develop an element-wise parameter selection focused on specific layers to further reduce the number of fine-tuning parameters.Experimental results show that our approach consistently prevents catastrophic forgetting compared to parameter selection from the whole model. Yuki Takashima, Shota Horiguchi, Shinji Watanabe 0001, L. Paola García-Perera, Yohei Kawaguchi |
INTERSPEECH | 5 |
| 2021 | Towards Neural Diarization for Unlimited Numbers of Speakers Using Global and Local AttractorsabstractAttractor-based end-to-end diarization is achieving comparable accuracy to the carefully tuned conventional clustering-based methods on challenging datasets. However, the main drawback is that it cannot deal with the case where the number of speakers is larger than the one observed during training. This is because its speaker counting relies on supervised learning. In this work, we introduce an unsupervised clustering process embedded in the attractor-based end-to-end diarization. We first split a sequence of frame-wise embeddings into short subsequences and then perform attractor-based diarization for each subsequence. Given subsequence-wise diarization results, inter-subsequence speaker correspondence is obtained by unsupervised clustering of the vectors computed from the attractors from all the subsequences. This makes it possible to produce diarization results of a large number of speakers for the whole recording even if the number of output speakers for each subsequence is limited. Experimental results showed that our method could produce accurate diarization results of an unseen number of speakers. Our method achieved 11.84 %, 28.33 %, and 19.49 % on the CALLHOME, DI-HARD II, and DIHARD III datasets, respectively, each of which is better than the conventional end-to-end diarization methods. Shota Horiguchi, Shinji Watanabe 0001, L. Paola García-Perera, Yawen Xue, Yuki Takashima, Yohei Kawaguchi |
ASRU | 6 |
| 2021 | Flow-Based Self-Supervised Density Estimation for Anomalous Sound DetectionabstractTo develop a machine sound monitoring system, a method for detecting anomalous sound is proposed. Exact likelihood estimation using Normalizing Flows is a promising technique for unsupervised anomaly detection, but it can fail at out-of-distribution detection since the likelihood is affected by the smoothness of the data. To improve the detection performance, we train the model to assign higher likelihood to target machine sounds and lower likelihood to sounds from other machines of the same machine type. We demonstrate that this enables the model to incorporate a self-supervised classification-based approach. Experiments conducted using the DCASE 2020 Challenge Task2 dataset showed that the proposed method improves the AUC by 4.6% on average when using Masked Autoregressive Flow (MAF) and by 5.8% when using Glow, which is a significant improvement over the previous method. Kota Dohi, Takashi Endo, Harsh Purohit, Ryo Tanabe, Yohei Kawaguchi |
ICASSP | 5 |
| 2020 | Anomalous Sound Detection Based on Interpolation Deep Neural NetworkabstractAs the labor force decreases, the demand for labor-saving automatic anomalous sound detection technology that conducts maintenance of industrial equipment has grown. Conventional approaches detect anomalies based on the reconstruction errors of an autoencoder. However, when the target machine sound is non-stationary, a reconstruction error tends to be large independent of an anomaly, and its variations increased because of the difficulty of predicting the edge frames. To solve the issue, we propose an approach to anomalous detection in which the model utilizes multiple frames of a spectrogram whose center frame is removed as an input, and it predicts an interpolation of the removed frame as an output. Rather than predicting the edge frames, the proposed approach makes the reconstruction error consistent with the anomaly. Experimental results showed that the proposed approach achieved 27% improvement based on the standard AUC score, especially against non-stationary machinery sounds. Kaori Suefusa, Tomoya Nishida, Harsh Purohit, Ryo Tanabe, Takashi Endo, Yohei Kawaguchi |
ICASSP | 6 |
| 2019 | Anomaly Detection Based on an Ensemble of Dereverberation and Anomalous Sound ExtractionabstractTo develop a sound-monitoring system for checking machine health, a method for detecting anomalous sounds is proposed. In real environments such as factories, reverberation and background noise are mixed in an observed signal, so detection performance is degraded. It can be expected that detection performance will be improved by using a front-end algorithm for acoustic signal processing such as dereverberation and denoising. However, any algorithm has pros and cons, so it is not possible to choose the best front-end algorithm only. To solve this problem, the proposed method is based on a front-end ensemble consisting of a blind-dereverberation algorithm and multiple anomalous-sound-extraction algorithms. Experimental results indicate that the proposed method improves detection performance significantly. Yohei Kawaguchi, Ryo Tanabe, Takashi Endo, Kenji Ichige, Koichi Hamada |
ICASSP | 1 |
| 2018 | Independent Low-Rank Matrix Analysis Based on Multivariate Complex Exponential Power DistributionabstractIndependent low-rank matrix analysis (ILRMA), a unified method of independent vector analysis (IVA) and nonnegative matrix factorization (NMF), is a state-of-the-art blind source separation method for convolutive mixtures. Although ILRMA provides high separation performance for music signals whose spectra can be well modeled by NMF, speech spectra do not have low-rank properties, and modeling them by NMF is not appropriate. In this paper, to stably improve the separation performance of ILRMA for speech mixtures, a source spectrum model in ILRMA is generalized to explicitly model the strong higher-order correlations between neighboring frequency bins of speech signals. In addition, multivariate complex exponential power distributions, which are recognized to have high performance with IVA, are introduced as source distributions assumed in ILRMA. Experimental results show the effectiveness of the proposed method over the original ILRMA when separating speech mixtures. Rintaro Ikeshita, Yohei Kawaguchi |
ICASSP | 2 |
| 2016 | Adaptive Boolean compressive sensing by using multi-armed banditabstractA new method for solving adaptive Boolean compressive sensing is proposed. By greedy maximization of an expected information gain, a conventional method controls the pool-size for adaptive Boolean compressive sensing. However, the conventional greedy method has the drawback that it has no guarantee of convergence to the optimal strategy. To solve the problem, based on the multi-armed bandit, the proposed method controls the pool-size adaptively. The information gain of the conventional greedy method is rewritten as the reward of the multi-armed bandit, and the multi-armed bandit is introduced into adaptive Boolean compressive sensing. Experimental results indicate that the correct rate of exact recovery of the proposed method converges to 1 fast without prior knowledge about the number of defective items and that the proposed method outperforms the conventional greedy method in the case that the number of defective items is large. Yohei Kawaguchi, Masahito Togami |
ICASSP | 1 |
| 2015 | Unified ASR system using LGM-based source separation, noise-robust feature extraction, and word hypothesis selectionabstractIn this paper, we propose a unified system that incorporates speech source separation and automatic speech recognition for various noise environments. There are three features in the proposed system. The first feature of the proposed method is the LGM (local Gaussian modeling) based source separation with the efficient permutation alignment method that integrates a power spectrum correlation based method and a direction-of-arrival (DOA) based method. Evaluation results show that using the separated speech with the baseline acoustic modeling method reduces the word error rate (WER) significantly. The second feature of the proposed method is multi-condition training with per-utterance normalized features and noise-aware features in the acoustic modeling step. In this paper, we show that the proposed training method is effective even when an input signal has been distorted through the source separation step. The third feature is the word hypothesis selection method for integrating multiple recognition results. The proposed selection method estimates correct words based on a recognizer's confidence and co-occurrence characteristics. The evaluation results show that the proposed selection method outperforms the conventional recognizer output voting error reduction (ROVER) method. The proposed system is evaluated using the third CHiME challenge dataset. Evaluation results show that the proposed system resulted in an improvement of 66.1% over the baseline system. Yusuke Fujita, Ryoichi Takashima, Takeshi Homma, Rintaro Ikeshita, Yohei Kawaguchi, Takashi Sumiyoshi, Takashi Endo, Masahito Togami |
ASRU | 5 |
| 2014 | Frequency domain acoustic echo reduction based on Kalman smoother with time-varying noise covariance matrixabstractIn this paper, we propose a novel acoustic-echo-reduction technique at a time-frequency domain, which is optimally combined with speech enhancement. Unlike conventional echo reduction techniques which minimizes only residual power of the far-end acoustic echo signal, the proposed method minimizes summation of the residual echo signal and distortion of the near-end speech signal from a minimum mean square error (MMSE) perspective. The proposed method performs echo reduction with speech enhancement and parameter optimization in an iterative manner based on the expectation-maximization (EM) algorithm. The E step is corresponding with the echo reduction and speech enhancement based on the Kalman smoother with a time-varying covariance matrix for the observation noise term, which reflects the time-varying characteristics of speech sources. By using the time-varying covariance matrix, we can enhance speech sources effectively with acoustic echo reduction. Associated with the time-varying covariance matrix, a new optimization scheme of parameters for the M step is derived in this paper. Experimental results with impulse responses which was recorded under a real meeting room show that the proposed method can effectively enhance a near-end speech signal when there are a near-end speech signal and a far-end acoustic echo signal. Masahito Togami, Yohei Kawaguchi, Ryoichi Takashima |
ICASSP | 2 |
| 2014 | Simultaneous Optimization of Acoustic Echo Reduction, Speech Dereverberation, and Noise Reduction against Mutual InterferenceabstractWe propose an optimized speech enhancement method that combines acoustic echo reduction, speech dereverberation, and noise reduction in a unified framework. Normally, partial optimization of acoustic echo reduction, speech dereverberation, and noise reduction does not lead to total optimization. A cascade method of multiple functions causes mutual interference between these functions and degrades eventual speech enhancement performance. Unlike cascade methods, the proposed method combines all functions to optimize eventual speech enhancement performance based on a unified framework, which is also robust against the mutual interference problem. With the proposed method, in addition to time-invariant linear filters, time-varying filters are used to reduce residual reverberation, residual acoustic echo signal, and background noise signal which cannot be reduced using time-invariant filters. These time-invariant filters and time-varying filters are also optimized based on a unified likelihood function to avoid the mutual interference problem. By combining the time-invariant linear filters and the time-varying filters, the proposed method uses a local Gaussian model with a full-rank covariance matrix and a non-zero average vector as a probabilistic model of the microphone input signal. In the local Gaussian model, non-stationary characteristics of speech sources are considered to effectively enhance speech sources. Under this probabilistic model, all the parameters are optimized simultaneously based on the expectation-maximization algorithm and calculates a minimum mean squared error estimate of a desired signal. The experimental results show that the proposed method is superior to the cascade methods. Masahito Togami, Yohei Kawaguchi |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | ICA-based acceleration of probabilistic latent component analysis for mass spectrometry-based explosives detectionabstractWe propose a new method to separate mass spectra into components of each chemical compound for explosives detection. The conventional method based on probabilistic latent component analysis (PLCA) is effective because the method can solve the problems of non-negativity and non-orthogonality by using sparsity of the domain of explosives detection. However, the convergence of the method is slow, and the calculation time is long. In order to solve this problem, the proposed method makes use of independent component analysis (ICA) in the initialization process. Experimental results indicate that the convergence of the proposed method is accelerated, and total calculation time is decreased. Yohei Kawaguchi, Masahito Togami, Hisashi Nagano, Yuichiro Hashimoto, Masuyuki Sugiyama, Yasuaki Takada |
ICASSP | 1 |
| 2013 | Noise robust speech dereverberation with Kalman smootherabstractA speech dereverberation method is proposed that is robust against background noise. In contrast to conventional methods based on the linear prediction of the given microphone input signal, in which the linear prediction coefficients are not fully optimized when there is background noise, the proposed method optimizes the coefficients by linear prediction of the noiseless reverberant speech signal even when there is background noise. The noiseless reverberant speech signal and the parameters are iteratively updated on the basis of the expectation maximization algorithm. In the expectation step, sufficient statistics of latent variables which include noiseless reverberant speech signal are estimated using the Kalman smoother. Unlike the standard Kalman smoother, which uses a time-invariant covariance matrix as a state-transition covariance matrix, the proposed method utilizes a time-varying covariance matrix, enabling it to meet the time-varying speech characteristics. The parameters are updated so that the Q function is increased in the maximization step. Experimental results show that the proposed method is superior to conventional methods under noisy conditions. Masahito Togami, Yohei Kawaguchi |
ICASSP | 2 |
| 2013 | Optimized Speech Dereverberation From Probabilistic Perspective for Time Varying Acoustic Transfer FunctionabstractA dereverberation technique has been developed that optimally combines multichannel inverse filtering (MIF), beamforming (BF), and non-linear reverberation suppression (NRS). It is robust against acoustic transfer function (ATF) fluctuations and creates less distortion than the NRS alone. The three components are optimally combined from a probabilistic perspective using a unified likelihood function incorporating two probabilistic models. A multichannel probabilistic source model based on a recently proposed local Gaussian model (LGM) provides robustness against ATF fluctuations of the early reflection. A probabilistic reverberant transfer function model (PRTFM) provides robustness against ATF fluctuations of the late reverberation. The MIF and multichannel under-determined source separation (MUSS) are optimized in an iterative manner. The MIF is designed to reduce the time-invariant part of the late reverberation by using optimal time-weighting with reference to the PRTFM and the LGM. The MUSS separates the dereverberated speech signal and the residual reverberation after the MIF, which can be interpreted as an optimized combination of the BF and the NRS. The parameters of the PRTFM and the LGM are optimized based on the MUSS output. Experimental results show that the proposed method is robust against the ATF fluctuations under both single and multiple source conditions. Masahito Togami, Yohei Kawaguchi, Ryu Takeda, Yasunari Obuchi, Nobuo Nukaga |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Mass spectra separation for explosives detection by using probabilistic latent component analysisabstractWe propose a new method to separate mass spectra into components of each chemical compound for explosives detection. In mass spectra, all components have no negative values. However, conventional factor analyses for basis decomposition use no constraints of non-negativity, and we can not apply these methods to mass spectra. The proposed method is based on probabilistic latent component analysis (PLCA). The constraints of non-negativity always hold in PLCA, so that the method is effective for mass spectra. In addition, PLCA is defined in a statistical framework, thus PLCA makes it possible to utilize additional a priori information. Therefore, we introduce sparseness assumptions in the domain of mass spectrometry to PLCA in order to estimate the components more accurately. Experimental results indicate that the proposed method outperforms existing methods. Yohei Kawaguchi, Masahito Togami, Hisashi Nagano, Yuichiro Hashimoto, Masuyuki Sugiyama, Yasuaki Takada |
ICASSP | 1 |
| 2012 | Multichannel speech dereverberation and separation with optimized combination of linear and non-linear filteringabstractIn this paper, we propose a multichannel speech dereverberation and separation technique which is effective even when there are multiple speakers and each speaker's transfer function is time-varying due to fluctuation of the corresponding speaker's head. For robustness against fluctuation, the proposed method optimizes linear filtering with non-linear filtering simultaneously from probabilistic perspective based on a probabilistic reverberant transfer-function model, PRTFM. PRTFM is an extension of the conventional time-invariant transfer-function model under uncertain conditions, and PRTFM can be also regarded as an extension of recently proposed blind local Gaussian modeling. The linear filtering and the non-linear filtering are optimized in MMSE (Minimum Mean Square Error) sense during parameter optimization. The proposed method is evaluated in a reverberant meeting room, and the proposed method is shown to be effective. Masahito Togami, Yohei Kawaguchi, Ryu Takeda, Yasunari Obuchi, Nobuo Nukaga |
ICASSP | 2 |
| 2010 | Head orientation estimation of a speaker by utilizing kurtosis of a DOA histogram with restoration of distance effectabstractIn this paper, we propose a head-orientation estimation method from multichannel acoustic signals. Sharpness of a DOA histogram which is extracted by using the sparseness based DOA estimation method varies depending on the head orientation of a speaker. The proposed method utilizes this phenomenon to estimate the head orientation of the speaker. The proposed method uses more than two microphone arrays. In addition to estimation of the speaker location, the proposed method estimates kurtosis of the DOA histogram of each array. Kurtosis is regarded as a measure of sharpness of a DOA histogram in the proposed method. However, kurtosis also depends on the distance between the speaker and the microphone array (distance effect). The distance effect is experimentally revealed by the regression analysis. The head orientation of a speaker is estimated by the restored kurtosis which is free from the distance effect. Experimental results on a reverberant environment show that the proposed method can estimate the head orientation of a speaker more accurately than a conventional head-orientation estimation method. Masahito Togami, Yohei Kawaguchi |
ICASSP | 2 |
| 2010 | Turn taking-based conversation detection by using DOA estimation
Yohei Kawaguchi, Masahito Togami, Yasunari Obuchi |
INTERSPEECH | 1 |
| 2009 | Subband nonstationary noise reduction based on multichannel spatial prediction under reverberant environmentsabstractWe propose a novel non-stationary and convolutive noise reduction method under reverberant environments. Unlike many multichannel noise reduction methods, the proposed method does not need pre knowledge of impulse response or direction of arrival (DOA) of the target source. The proposed method is composed of two processes. On the noise reduction process, the noise component is reduced without the impulse response of the target source. The target source component in the output signal is distorted, but the distortion is removed by the distortion-restoration process. Importantly, possibility of complete noise reduction with no distortion based on the proposed framework is assured by MINT theory. Experimental results under the reverberant environment (RT60≈ 300 ms) show that the proposed method can reduce more noise than the conventional method and the distortion of the target source is not so big. Masahito Togami, Yohei Kawaguchi, Yasunari Obuchi |
ICASSP | 2 |