EDBT 2026 Demo / reviewers in the wild / expert
Shoji Makino
dblp:31/6801
· DBLP profile ↗
86ranked-venue papers
5as first author
14since 2021 · last 2026
0000-0003-1934-640XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 61 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 27 · 1 first-author · 6 since 2021Systems, architecture and hardware · 5Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting Steering Limitations of LCMV Beamforming for Circular Microphone Arrays and a Mainlobe-Controlled SolutionabstractThis paper investigates the performance limitations of conventional linearly constrained minimum variance (LCMV) beamformers implemented with circular microphone arrays. In particular, we show that imposing null constraints on interference directions can lead to a deviation of the mainlobe from the desired steering angle. A theoretical analysis is presented to characterize this deviation, and a closed-form expression for the deviation is derived to reveal the underlying low-frequency behavior of LCMV beamformers. To address this issue, we propose a mainlobe-controlled LCMV (MC-LCMV) beamformer. Simulation results demonstrate that the proposed method substantially improves spatial directivity and speech enhancement performance compared with conventional LCMV beamformers. Wei Liu 0177, Gongping Huang, Jilu Jin, Xueqin Luo, Shoji Makino |
IEEE Signal Process. Lett. | 5 |
| 2024 | A Computationally Efficient Semi-Blind Source Separation Approach for Nonlinear Echo Cancellation Based on an Element-Wise Iterative Source SteeringabstractWhile the semi-blind source separation-based acoustic echo cancellation (SBSS-AEC) has received much research attention due to its promising performance during double-talk compared to the traditional adaptive algorithms, it suffers from system latency and nonlinear distortions. To circumvent these drawbacks, the recently developed ideas on convolutive transfer function (CTF) approximation and nonlinear expansion have been used in the iterative projection (IP)-based semi-blind source separation (SBSS) algorithm. However, because of the introduction of CTF approximation and nonlinear expansion, this algorithm becomes computationally very expensive, which makes it difficult to implement in embedded systems. Thus, we attempt in this paper to improve this IP-based algorithm, thereby developing an element-wise iterative source steering (EISS) algorithm. In comparison with the IP-based SBSS algorithm, the proposed algorithm is computationally much more efficient, especially when the nonlinear expansion order is high and the length of the CTF filter is long. Meanwhile, its AEC performance is as good as that of IP-based SBSS algorithm. Kunxing Lu, Xianrui Wang, Tetsuya Ueda, Shoji Makino, Jingdong Chen |
ICASSP | 4 |
| 2024 | Neural Network-Based Virtual Microphone Estimation with Virtual Microphone and Beamformer-Level Multi-Task LossabstractArray processing performance depends on the number of microphones available. Virtual microphone estimation (VME) has been proposed to increase the number of microphone signals artificially. Neural network-based VME (NN-VME) trains an NN with a VM-level loss to predict a signal at a microphone location that is available during training but not at inference. However, this training objective may not be optimal for a specific array processing back-end, such as beamforming. An alternative approach is to use a training objective considering the array-processing back-end, such as a loss on the beamformer output. This approach may generate signals optimal for beamforming but not physically grounded. To combine the advantages of both approaches, this paper proposes a multi-task loss for NN-VME that combines both VM-level and beamformer-level losses. We evaluate the proposed multi-task NN-VME on multi-talker underdetermined conditions and show that it achieves a 33.1 % relative WER improvement compared to using only real microphones and 10.8 % compared to using a prior NN-VME approach. Hanako Segawa, Tsubasa Ochiai, Marc Delcroix, Tomohiro Nakatani, Rintaro Ikeshita, Shoko Araki, Takeshi Yamada, Shoji Makino |
ICASSP | 8 |
| 2024 | Stereophonic Music Source Separation with Spatially-Informed Bridging Band-Split NetworkabstractStereophonic music source separation (MSS) is a problem of extracting individual source tracks, e.g. bass, drums, vocals, from a stereo music recording. Deep neural network (DNN) based MSS systems have demonstrated great promise though spatial panning cues and time-frequency spectral structures in stereo music have not yet been fully explored in such systems and methods. This paper presents a spatially-informed MSS method using a bridging band-split neural network that incorporates both spatial and spectral information. The spatial panning angles of each target source are used as input of the network, along with the time-frequency spectrograms. Moreover, the inter-track correlations are exploited for further performance improvement. Experiments show that the proposed method outperforms significantly the baseline systems as the result of using spatial cues, spectral characteristics, and inter-track relationships. Yichen Yang 0010, Xianrui Wang, Wen Zhang 0002, Shoji Makino, Jingdong Chen |
ICASSP | 5 |
| 2024 | Unrestricted Global Phase Bias-Aware Single-Channel Speech Enhancement with Conformer-Based Metric GanabstractWith the rapid development of neural networks in recent years, the ability of various networks to enhance the magnitude spectrum of noisy speech in the single-channel speech enhancement domain has become exceptionally outstanding. However, enhancing the phase spectrum using neural networks is often ineffective, which remains a challenging problem. In this paper, we found that the human ear cannot sensitively perceive the difference between a precise phase spectrum and a biased phase (BP) spectrum. Therefore, we propose an optimization method of phase reconstruction, allowing freedom on the global-phase bias instead of reconstructing the precise phase spectrum. We applied it to a Conformer-based Metric Generative Adversarial Networks (CMGAN) baseline model, which relaxes the existing constraints of precise phase and gives the neural network a broader learning space. Results show that this method achieves a new state-of-the-art performance without incurring additional computational overhead. Zheng Qiu, Daiki Takeuchi, Noboru Harada, Shoji Makino |
ICASSP | 5 |
| 2024 | Blind and Spatially-Regularized Online Joint Optimization of Source Separation, Dereverberation, and Noise ReductionabstractThis paper proposes a computationally efficient joint optimization algorithm that performs online source separation, dereverberation, and noise reduction based on blind and spatially-regularized processing. When applying such online Blind Source Separation (BSS) as online Independent Vector Extraction (IVE) to a speech application, we must focus on the trade-off between the algorithmic delay and separation accuracy, both of which depend on the analysis frame length. In addition, to separate the sources with specified source permutation, researchers introduced spatial regularization based on the Directions-of-Arrival (DOAs) of the sources into IVE. However, the scale ambiguity of IVE often makes the spatial regularization work inappropriately. To solve these problems, we first propose a blind online joint optimization algorithm of IVE and weighted prediction error dereverberation (WPE). This online algorithm can achieve accurate separation even using short analysis frames because reverberation can be reduced using WPE. We then extend the online joint optimization with robust spatial regularization. We reveal that regularizing the scale of the separated signals is very effective in making the DOA-based spatial regularization work reliably. Our experiments confirm that our blind online joint optimization algorithm can significantly improve the separation accuracy with an algorithmic delay of 8 ms. In addition, we confirm that the proposed spatially-regularized online joint optimization algorithm reduces the rate of the source permutation error to zero percent. Tetsuya Ueda, Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Shoko Araki, Shoji Makino |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2024 | On Semi-Blind Source Separation-Based Approaches to Nonlinear Echo Cancellation Based on Bilinear Alternating OptimizationabstractAcoustic echo cancellation (AEC) is a crucial task in full duplex communications. As conventional linear filtering approaches are ineffective to deal with double-talk, various semi-blind source separation (SBSS)-based AEC algorithms are deceived, most of which are formulated and implemented in the frequency domain based on the multiplicative transfer function (MTF) model for computational efficiency. To avoid large latency and in order to deal with loudspeaker nonlinearities, the convolutive transfer function (CTF) model and odd power series expansion are leveraged, which are employed by numerous SBSS-based nonlinear AEC (SBSS-NAEC) algorithms. Conventional SBSS-NAEC methods estimate the series expansion coefficients and the CTF filter simultaneously making the number of free parameters to estimate large. Hence, the corresponding algorithms are computationally expensive and are difficult to optimize. In this work, we propose to decouple the series expansion coefficients and the CTF filters into a bilinear form and present a bilinear alternating optimization framework for estimating the model parameters. An alternating iterative projection (AIP) algorithm and an alternating element-wise iterative source steering (AEISS) algorithm are proposed. As the bilinear representation consists of less parameters compared to the conventional methods, the proposed algorithms not only improve the AEC performance but also reduce the computational complexity, which is validated by comprehensive simulations and experiments. Xianrui Wang, Yichen Yang 0010, Andreas Brendel, Tetsuya Ueda, Shoji Makino, Jacob Benesty, Walter Kellermann, Jingdong Chen |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | FastMVAE2: On Improving and Accelerating the Fast Variational Autoencoder-Based Source Separation Algorithm for Determined MixturesabstractThis article proposes a new source model and training scheme to improve the accuracy and speed of the multichannel variational autoencoder (MVAE) method. The MVAE method is a recently proposed powerful multichannel source separation method. It consists of pretraining a source model represented by a conditional VAE (CVAE) and then estimating separation matrices along with other unknown parameters so that the log-likelihood is non-decreasing given an observed mixture signal. Although the MVAE method has been shown to provide high source separation performance, one drawback is the computational cost of the backpropagation steps in the separation-matrix estimation algorithm. To overcome this drawback, a method called “FastMVAE” was subsequently proposed, which uses an auxiliary classifier VAE (ACVAE) to train the source model. By using the classifier and encoder trained in this way, the optimal parameters of the source model can be inferred efficiently, albeit approximately, in each step of the algorithm. However, the generalization capability of the trained ACVAE source model was not satisfactory, which led to poor performance in situations with unseen data. To improve the generalization capability, this article proposes a new model architecture (called the “ChimeraACVAE” model) and a training scheme based on knowledge distillation. The experimental results revealed that the proposed source model trained with the proposed loss function achieved better source separation performance with less computation time than FastMVAE. We also confirmed that our methods were able to separate 18 sources with a reasonably good accuracy. Li Li 0063, Hirokazu Kameoka, Shoji Makino |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Performance Improvement of Speech Emotion Recognition by Neutral Speech Detection Using Autoencoder and Intermediate Representation
Jennifer Santoso, Takeshi Yamada, Kenkichi Ishizuka, Taiichi Hashimoto, Shoji Makino |
INTERSPEECH | 5 |
| 2021 | SepNet: A Deep Separation Matrix Prediction Network for Multichannel Audio Source SeparationabstractIn this paper, we propose SepNet, a deep neural network (DNN) designed to predict separation matrices from multichannel observations. One well-known approach to blind source separation (BSS) involves independent component analysis (ICA). A recently developed method called independent low-rank matrix analysis (ILRMA) is one of its powerful variants. These methods allow the estimation of separation matrices based on deterministic iterative algorithms. Specifically, ILRMA is designed to update the separation matrix according to an update rule derived based on the majorization-minimization principle. Although ILRMA performs reasonably well under some conditions, there is still room for improvement in terms of both separation accuracy and computation time, especially for large-scale microphone arrays. The existence of a deterministic iterative algorithm that can find one of the stationary points of the BSS problem implies that a DNN can also play that role if designed and trained properly. Motivated by this, we propose introducing a DNN that learns to convert a predefined input (e.g., an identity matrix) into a true separation matrix in accordance with a multichannel observation. To enable it to find one of the multiple solutions corresponding to different permutations of the source indices, we further propose adopting a permutation invariant training strategy to train the network. By using a fully convolutional architecture, we can design the network so that the forward propagation can be computed efficiently. The experimental results revealed that SepNet was able to find separation matrices faster and with better separation accuracy than ILRMA for mixtures of two sources. Shota Inoue, Hirokazu Kameoka, Li Li 0063, Shoji Makino |
ICASSP | 4 |
| 2021 | Teacher-Student Learning for Low-Latency Online Speech Enhancement Using Wave-U-NetabstractIn this paper, we propose a low-latency online extension of wave-U-net for single-channel speech enhancement, which utilizes teacher-student learning to reduce the system latency while keeping the enhancement performance high. Wave-U-net is a recently proposed end-to-end source separation method, which achieved remarkable performance in singing voice separation and speech enhancement tasks. Since the enhancement is performed in the time domain, wave-U-net can efficiently model phase information and address the domain transformation limitation, where the time-frequency domain is normally adopted. In this paper, we apply wave-U-net to face-to-face applications such as hearing aids and in-car communication systems, where a strictly low-latency of less than 10 ms is required. To this end, we investigate online versions of wave-U-net and propose the use of teacher-student learning to prevent the performance degradation caused by the reduction in input segment length such that the system delay in a CPU is less than 10 ms. The experimental results revealed that the proposed model could perform in real-time with low-latency and high performance, achieving a signal-to-distortion ratio improvement of about 8.73 dB. Sotaro Nakaoka, Li Li 0063, Shota Inoue, Shoji Makino |
ICASSP | 4 |
| 2021 | Low Latency Online Blind Source Separation Based on Joint Optimization with Blind DereverberationabstractThis paper presents a new low-latency online blind source separation (BSS) algorithm. Although algorithmic delay of a frequency domain online BSS can be reduced simply by shortening the short-time Fourier transform (STFT) frame length, it degrades the source separation performance in the presence of reverberation. This paper proposes a method to solve this problem by integrating BSS with Weighted Prediction Error (WPE) based dereverberation. Although a simple cascade of online BSS after online WPE upgrades the separation performance, the overall optimality is not guaranteed. Instead, this paper extends a recently proposed batch processing algorithm that can jointly optimize dereverberation and separation so that it can perform online processing with low computational cost and little processing delay (< 12 ms). The results of a source separation experiment in a noisy car environment suggest that the proposed online method has better separation performance than the simple cascaded methods. Tetsuya Ueda, Tomohiro Nakatani, Rintaro Ikeshita, Keisuke Kinoshita, Shoko Araki, Shoji Makino |
ICASSP | 6 |
| 2021 | Speech Emotion Recognition Based on Attention Weight Correction Using Word-Level Confidence Measure
Jennifer Santoso, Takeshi Yamada, Shoji Makino, Kenkichi Ishizuka, Takekatsu Hiramura |
Interspeech | 3 |
| 2021 | Time-Frequency-Bin-Wise Linear Combination of Beamformers for Distortionless Signal EnhancementabstractIn this paper, we address signal enhancement in underdetermined situations and propose new beamforming algorithms. Beamforming in (over)determined situations can successfully reduce noise signals without distortion of a desired signal, which is known to be a desirable property, especially for automatic speech recognition systems. Even in underdetermined situations, time-frequency (TF) masking attains outstanding performance in noise reduction, although it tends to generate artifacts. Integrating these two approaches to benefit from both their advantages, we here propose time-frequency-bin-wise switching (TFS) and time-frequency-bin-wise linear combination (TFLC) beamforming. In the proposed methods, we utilize the best combination of beamformers among multiple beamformers at each TF bin, each of which suppresses a particular combination of interferers. First, we propose a general formulation of signal enhancement employing multiple spatial filters. Then a joint optimization problem of designing the spatial filters and estimating the suitable weights to combine them is considered under a unified minimum variance criterion. Finally, we present efficient algorithms to solve the problem. In experiments, we used an objective criterion that quantifies the amount of signal distortion caused by the enhancement function and confirmed that the proposed methods effectively suppress interferers without distortion of the target signal. Kouei Yamaoka, Nobutaka Ono, Shoji Makino |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Online Directional Speech Enhancement Using Geometrically Constrained Independent Vector Analysis
Li Li 0063, Kazuhito Koishida, Shoji Makino |
INTERSPEECH | 3 |
| 2019 | Joint Separation and Dereverberation of Reverberant Mixtures with Multichannel Variational AutoencoderabstractIn this paper, we deal with a multichannel source separation problem under a highly reverberant condition. The multichannel variational autoencoder (MVAE) is a recently proposed source separation method that employs the decoder distribution of a conditional VAE (CVAE) as the generative model for the complex spectrograms of the underlying source signals. Although MVAE is notable in that it can significantly improve the source separation performance compared with conventional methods, its capability to separate highly reverberant mixtures is still limited since MVAE uses an instantaneous mixture model. To overcome this limitation, in this paper we propose extending MVAE to simultaneously solve source separation and dereverberation problems by formulating the separation system as a frequency-domain convolutive mixture model. A convergence-guaranteed algorithm based on the coordinate descent method is derived for the optimiza- tion. Experimental results revealed that the proposed method outperformed the conventional methods in terms of all the source separation criteria in highly reverberant environments. Shota Inoue, Hirokazu Kameoka, Li Li 0063, Shogo Seki, Shoji Makino |
ICASSP | 5 |
| 2019 | Fast MVAE: Joint Separation and Classification of Mixed Sources Based on Multichannel Variational Autoencoder with Auxiliary ClassifierabstractThis paper proposes an alternative algorithm for the multi-channel variational autoencoder (MVAE), a recently proposed multichannel source separation approach. While MVAE is notable for its impressive source separation performance, its convergence-guaranteed optimization algorithm and the fact that it allows us to estimate source-class labels simultaneously with source separation, there are still two major drawbacks, namely, the high computational complexity and the unsatisfactory source classification accuracy. To overcome these drawbacks, the proposed method employs an auxiliary classifier VAE, which is an information-theoretic extension of the conditional VAE, for learning the generative model of the source spectrograms. Furthermore, with the trained auxiliary classifier, we introduce a novel algorithm for the optimization that can both reduce the computational time and improve the source classification performance. We call the proposed method "fast MVAE (fMVAE) ". Experimental evaluations revealed that fMVAE achieved source separation performance comparable to that of MVAE and a source classification accu-racy rate of about 80% while reducing computational time by about 93%. Li Li 0063, Hirokazu Kameoka, Shoji Makino |
ICASSP | 3 |
| 2019 | Time-frequency-bin-wise Switching of Minimum Variance Distortionless Response Beamformer for Underdetermined SituationsabstractIn this paper, we present a speech enhancement method using two microphones in underdetermined situations. Time-frequency (TF) binary masking is a conventional method of enhancing speech in underdetermined situations by appropriately multiplying each TF component by zero or one. Extending this method, we previously proposed a new method called the time-frequency-bin-wise switching (TFS) beamformer. In this method, we switch multiple preconstructed beamformers in each TF bin, each of which suppresses a particular interferer. However, this method requires the pre-estimation of beamformer filter coefficients using the target-active period and interferer-wise-active periods as the prior information. In this paper, to overcome this limitation, we formulate the switching and construction of spatial filters as a joint optimization problem, which can be understood from two viewpoints: the clustering of the most dominant interferer signal in each TF bin and the construction of a minimum variance distortionless response beamformer using such bins. In an experiment, we confirmed that the proposed method was superior to conventional TF masking and fixed beamforming during speech enhancement regardless of the direction of interferers. Kouei Yamaoka, Nobutaka Ono, Shoji Makino, Takeshi Yamada |
ICASSP | 3 |
| 2019 | Supervised Determined Source Separation with Multichannel Variational AutoencoderabstractThis letter proposes a multichannel source separation technique, the multichannel variational autoencoder (MVAE) method, which uses a conditional VAE (CVAE) to model and estimate the power spectrograms of the sources in a mixture. By training the CVAE using the spectrograms of training examples with source-class labels, we can use the trained decoder distribution as a universal generative model capable of generating spectrograms conditioned on a specified class index. By treating the latent space variables and the class index as the unknown parameters of this generative model, we can develop a convergence-guaranteed algorithm for supervised determined source separation that consists of iteratively estimating the power spectrograms of the underlying sources, as well as the separation matrices. In experimental evaluations, our MVAE produced better separation performance than a baseline method. Hirokazu Kameoka, Li Li 0063, Shota Inoue, Shoji Makino |
Neural Comput. | 4 |
| 2017 | Speech Enhancement Using Non-Negative Spectrogram Models with Mel-Generalized Cepstral Regularization
Li Li 0063, Hirokazu Kameoka, Tomoki Toda, Shoji Makino |
INTERSPEECH | 4 |
| 2016 | Multi-Talker Speech Recognition Based on Blind Source Separation with ad hoc Microphone Array Using Smartphones and Cloud Storage
Keiko Ochi, Nobutaka Ono, Shigeki Miyabe, Shoji Makino |
INTERSPEECH | 4 |
| 2015 | Blind compensation of interchannel sampling frequency mismatch for ad hoc microphone array based on maximum likelihood estimationabstractIn this paper, we propose a novel method for the blind compensation of drift for the asynchronous recording of an ad hoc microphone array. Digital signals simultaneously observed by different recording devices have drift of the time differences between the observation channels because of the sampling frequency mismatch among the devices. On the basis of a model in which the time difference is constant within each short time frame but varies in proportion to the central time of the frame, the effect of the sampling frequency mismatch can be compensated in the short-time Fourier transform (STFT) domain by a linear phase shift. By assuming that the sources are motionless and have stationary amplitudes, the observation is regarded as being stationary when drift does not occur. Thus, we formulate a likelihood to evaluate the stationarity in the STFT domain to evaluate the compensation of drift. The maximum likelihood estimation is obtained effectively by a golden section search. Using the estimated parameters, we compensate the drift by STFT analysis with a noninteger frame shift. The effectiveness of the proposed blind drift compensation method is evaluated in an experiment in which artificial drift is generated. Shigeki Miyabe, Nobutaka Ono, Shoji Makino |
Signal Process. | 3 |
| 2013 | Blind compensation of inter-channel sampling frequency mismatch with maximum likelihood estimation in STFT domainabstractThis paper proposes a novel blind compensation of sampling frequency mismatch for asynchronous microphone array. Digital signals simultaneously observed by different recording devices have drift of the time differences between the observation channels because of the sampling frequency mismatch among the devices. Based on the model that such the time difference is constant within each time frame, but varies proportional to the time frame index, the effect of the sampling frequency mismatch can be compensated in the short-time Fourier transform domain by the linear phase shift. By assuming the sources are motionless and stationary, a likelihood of the sampling frequency mismatch is formulated. The maximum likelihood estimation is obtained effectively by a golden section search. Shigeki Miyabe, Nobutaka Ono, Shoji Makino |
ICASSP | 3 |
| 2012 | New analytical update rule for TDOA inference for underdetermined BSS in noisy environmentsabstractIn this paper, we propose a new technique for sparseness-based underdetermined BSS that is based on the clustering of the frequency-dependent time difference of arrival (TDOA) information and that can cope with diffused noise environments. Such a method with an EM algorithm has already been proposed, however, it required a time-consuming exhaust search for TDOA inference. To remove the need for such an exhaust search, we propose a new technique by focusing on a stereo case. We derive an update rule for analytical TDOA estimation. This update rule eliminates the need for the exhaustive TDOA search, and therefore reduces the computational load. We show experimental results for separation performance and calculation time in comparison with those obtained with the conventional approach. Our reported results validate our proposed method, that is, our proposed method achieves high performance without a high computational cost. Takuro Maruyama, Shoko Araki, Tomohiro Nakatani, Shigeki Miyabe, Takeshi Yamada, Shoji Makino, Atsushi Nakamura |
ICASSP | 6 |
| 2011 | Underdetermined Convolutive Blind Source Separation via Frequency Bin-Wise Clustering and Permutation AlignmentabstractThis paper presents a blind source separation method for convolutive mixtures of speech/audio sources. The method can even be applied to an underdetermined case where there are fewer microphones than sources. The separation operation is performed in the frequency domain and consists of two stages. In the first stage, frequency-domain mixture samples are clustered into each source by an expectation-maximization (EM) algorithm. Since the clustering is performed in a frequency bin-wise manner, the permutation ambiguities of the bin-wise clustered samples should be aligned. This is solved in the second stage by using the probability on how likely each sample belongs to the assigned class. This two-stage structure makes it possible to attain a good separation even under reverberant conditions. Experimental results for separating four speech signals with three microphones under reverberant conditions show the superiority of the new method over existing methods. We also report separation results for a benchmark data set and live recordings of speech mixtures. Hiroshi Sawada, Shoko Araki, Shoji Makino |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | MPEG-2/H.264 Transcoding with Vector Conversion Reducing Re-Quantization NoiseabstractWe propose an MPEG-2 to H.264 transcoding method for interlace streams intermingled with frame and field macroblocks. This method uses the encoding information from an MPEG-2 stream and keeps as many DCT coefficients of the original MPEG-2 bitstream as possible. Experimental results show that the proposed method improves PSNR by about 0.190.31 dB compared with a conventional method. Takeshi Yoshitome, Kazuto Kamikura, Shoji Makino, Nobuhiko Kitawaki |
ICCCN | 3 |
| 2010 | Performance estimation of noisy speech recognition considering recognition task complexityabstractTo ensure a satisfactory QoE (Quality of Experience) and facilitate system design in speech recognition services, it is essential to establish a method that can be used to efficiently investigate recognition performance in different noise environments. Previously, we proposed a performance estimation method using a spectral distortion measure. However, there is the problem that recognition task complexity affects the relationship between the recognition performance and the distortion value. To solve this problem, this paper proposes a novel performance estimation method considering the recognition task complexity. We confirmed that the proposed method gives accurate estimates of the recognition performance for various recognition tasks by an experiment using noisy speech data recorded in a real room. Index Terms: performance estimation, noisy speech recognition, recognition task difficulty Takeshi Yamada, Tomohiro Nakajima, Nobuhiko Kitawaki, Shoji Makino |
INTERSPEECH | 4 |
| 2010 | Cepstral smoothing of separated signals for underdetermined speech separationabstractMusical noise is a typical problem with blind source separation using a time-frequency mask. Recently, the cepstral smoothing of spectral masks (CSM) was proposed. Based on the idea of smoothing in the cepstral domain, this paper proposes the cepstral smoothing of separated signals (CSS) on the assumption that a cepstral representation better reflects the characteristics of speech signals than those of masks (or filter gains). We also report a comparative evaluation study of CSM and CSS with other musical noise reduction methods. Our experimental results show that CSM is effective for musical noise reduction, but the target speech was relatively distorted. On the other hand, our proposed CSS produced less distorted target signals with the same musical noise reduction as CSM. Yumi Ansa, Shoko Araki, Shoji Makino, Tomohiro Nakatani, Takeshi Yamada, Atsushi Nakamura, Nobuhiko Kitawaki |
ISCAS | 3 |
| 2009 | Blind sparse source separation for unknown number of sources using Gaussian mixture model fitting with Dirichlet priorabstractIn this paper, we propose a novel sparse source separation method that can be applied even if the number of sources is unknown. Recently, many sparse source separation approaches with time-frequency masks have been proposed. However, most of these approaches require information on the number of sources in advance. In our proposed method, we model the histogram of the estimated direction of arrival (DOA) with a Gaussian mixture model (GMM) with a Dirichlet prior. Then we estimate the model parameters by using the maximum a posteriori estimation based on the EM algorithm. In order to avoid one cluster being modeled by two or more Gaussians, we utilize a sparse distribution modeled by the Dirichlet distributions as the prior of the GMM mixture weight. By using this prior, without any specific model selection process, our proposed method can estimate the number of sources and time-frequency masks simultaneously. Experimental results show the performance of our proposed method. Shoko Araki, Tomohiro Nakatani, Hiroshi Sawada, Shoji Makino |
ICASSP | 4 |
| 2009 | Frequency-Domain Pearson Distribution Approach for Independent Component Analysis (FD-Pearson-ICA) in Blind Source SeparationabstractIn frequency-domain blind source separation (BSS) for speech with independent component analysis (ICA), a practical parametric Pearson distribution system is used to model the distribution of frequency-domain source signals. ICA adaptation rules have a score function determined by an approximated signal distribution. Approximation based on the data may produce better separation performance than we can obtain with ICA. Previously, conventional hyperbolic tangent$(tanh)$or generalized Gaussian distribution (GGD) was uniformly applied to the score function for all frequency bins, even though a wideband speech signal has different distributions at different frequencies. To deal with this, we propose modeling the signal distribution at each frequency by adopting a parametric Pearson distribution and employing it to optimize the separation matrix in the ICA learning process. The score function is estimated by the appropriate Pearson distribution parameters for each frequency bin. We devised three methods for Pearson distribution parameter estimation and conducted separation experiments with real speech signals convolved with actual room impulse responses$(T_{60}=130\ {\hbox {ms}})$. Our experimental results show that the proposed frequency-domain Pearson-ICA (FD-Pearson-ICA) adapted well to the characteristics of frequency-domain source signals. By applying the FD-Pearson-ICA performance, the signal-to-interference ratio significantly improved by around 2–3 dB compared with conventional nonlinear functions. Even if the signal-to-interference ratio (SIR) values of FD-Pearson-ICA were poor, the performance based on a disparity measure between the true score function and estimated parametric score function clearly showed the advantage of FD-Pearson-ICA. Furthermore, we confirmed the optimum of the proposed approach for/optimized the proposed approach as regards separation performance. By combining individual distribution parameters directly estimated at low frequency with the appropriate parameters optimized at high frequency, it was possible to both reasonably improve the FD-Pearson-ICA performance without any significant increase in the computational burden by comparison with conventional nonlinear functions. Hiroko Kato Solvang, Yuichi Nagahara, Shoko Araki, Hiroshi Sawada, Shoji Makino |
IEEE Trans. Speech Audio Process. | 5 |
| 2008 | Speaker indexing and speech enhancement in real meetings / conversationsabstractThis paper presents a speaker indexing method that uses a small number of microphones to estimate who spoke when. Our proposed speaker indexing is realized by using a noise robust voice activity detector (VAD), a QCC-PHAT based direction of arrival (DOA) estimator, and a DOA classifier. Using the estimated speaker indexing information, we can also enhance the utterances of each speaker with a maximum signal-to-noise-ratio (MaxSNR) beamformer. This paper applies our system to real recorded meetings / conversations recorded in a room with a reverberation time of 350 ms, and evaluates the performance by a standard measure: the diarization error rate (DER). Even for the real conversations, which have many speaker turn-takings and overlaps, the speaker error time was very small with our proposed system. We are planning to demonstrate a real-time speaker indexing system at ICASSP2008. Shoko Araki, Masakiyo Fujimoto, Kentaro Ishizuka, Hiroshi Sawada, Shoji Makino |
ICASSP | 5 |
| 2008 | Missing feature speech recognition in a meeting situation with maximum SNR beamformingabstractEspecially for tasks like automatic meeting transcription, it would be useful to automatically recognize speech also while multiple speakers are talking simultaneously. For this purpose, speech separation can be performed, for example by using maximum SNR beamforming. However, even when good interferer suppression is attained, the interfering speech will still be recognizable during those intervals, where the target speaker is silent. In order to avoid the consequential insertion errors, a new soft masking scheme is proposed, which works in the time domain by inducing a large damping on those temporal periods, where the observed direction of arrival does not correspond to that of the target speaker. Even though the masking scheme is aggressive, by means of missing feature recognition the recognition accuracy can be improved significantly, with relative error reductions in the order of 60% compared to maximum SNR beamforming alone, and it is successful also for three simultaneously active speakers. Results are reported based on the SOLON speech recognizer, NTT’s large vocabulary system [1], which is applied here for the recognition of artificially mixed data using real-room impulse responses and the entire clean test set of the Aurora 2 database. Dorothea Kolossa, Shoko Araki, Marc Delcroix, Tomohiro Nakatani, Reinhold Orglmeister, Shoji Makino |
ISCAS | 6 |
| 2007 | Blind Speech Separation in a Meeting Situation with Maximum SNR BeamformersabstractWe propose a speech separation method for a meeting situation, where each speaker sometimes speaks and the number of speakers changes every moment. Many source separation methods have already been proposed, however, they consider a case where all the speakers keep speaking: this is not always true in a real meeting. In such cases, in addition to separation, speech detection and the classification of the detected speech according to speaker become important issues. For that purpose, we propose a method that employs a maximum signal-to-noise (MaxSNR) beamformer combined with a voice activity detector and online clustering. We also discuss the scaling ambiguity problem as regards the MaxSNR beamformer, and provide their solutions. We report some encouraging results for a real meeting in a room with a reverberation time of about 350 ms. Shoko Araki, Hiroshi Sawada, Shoji Makino |
ICASSP (1) | 3 |
| 2007 | Blind Source Separation Based on a Beamformer Array and Time Frequency Binary MaskingabstractThis paper deals with a new technique for blind source separation (BSS) from convolutive mixtures. We present a three-stage separation system employing time-frequency binary masking, beamforming and a non-linear post processing technique. The experiments show that this system outperforms conventional time-frequency binary masking (TFBM) in both (over-)determined and underdetermined cases. Moreover it removes the musical noise and reduces interference in time-frequency slots extracted by TFBM. Jan Cermak, Shoko Araki, Hiroshi Sawada, Shoji Makino |
ICASSP (1) | 4 |
| 2007 | Measuring Dependence of Bin-wise Separated Signals for Permutation Alignment in Frequency-domain BSSabstractThis paper presents a new method for grouping bin-wise separated signals for individual sources, i.e., solving the permutation problem, in the process of frequency-domain blind source separation. Conventionally, the correlation coefficient of separated signal envelopes is calculated to judge whether or not the separated signals originate from the same source. In this paper, we propose a new measure that represents the dominance of the separated signal in the mixtures, and use it for calculating the correlation coefficient, instead of a signal envelope. Such dominance measures exhibit dependence/independence more clearly than traditionally used signal envelopes. Consequently, a simple clustering algorithm with centroids works well for grouping separated signals. Experimental results were very appealing, as three sources including two coming from the same direction were separated properly with the new method. Hiroshi Sawada, Shoko Araki, Shoji Makino |
ISCAS | 3 |
| 2007 | Underdetermined blind sparse source separation for arbitrarily arranged multiple sensors
Shoko Araki, Hiroshi Sawada, Ryo Mukai, Shoji Makino |
Signal Process. | 4 |
| 2007 | Spatio-Temporal FastICA Algorithms for the Blind Separation of Convolutive MixturesabstractThis paper derives two spatio-temporal extensions of the well-known FastICA algorithm of Hyvarinen and Oja that are applicable to the convolutive blind source separation task. Our time-domain algorithms combine multichannel spatio-temporal prewhitening via multistage least-squares linear prediction with novel adaptive procedures that impose paraunitary constraints on the multichannel separation filter. The techniques converge quickly to a separation solution without any step size selection or divergence difficulties, and unlike other methods, ours do not require special coefficient initialization procedures to obtain good separation performance. They also allow for the efficient reconstruction of individual signals as observed in the sensor measurements directly from the system parameters for single-input multiple-output blind source separation tasks. An analysis of one of the adaptive constraint procedures shows its fast convergence to a paraunitary filter bank solution. Numerical evaluations of the proposed algorithms and comparisons with several existing convolutive blind source separation techniques indicate the excellent relative performance of the proposed methods. Scott C. Douglas, Malay Gupta, Hiroshi Sawada, Shoji Makino |
IEEE Trans. Speech Audio Process. | 4 |
| 2007 | Geometrically Constrained Independent Component AnalysisabstractAcoustical signals are often corrupted by other speeches, sources, and background noise. This makes it necessary to use some form of preprocessing so that signal processing systems such as a speech recognizer or machine diagnosis can be effectively employed. In this contribution, we introduce and evaluate a new algorithm that uses independent component analysis (ICA) with a geometrical constraint [constrained ICA (CICA)]. It is based on the fundamental similarity between an adaptive beamformer and blind source separation with ICA, and does not suffer the permutation problem of ICA-algorithms. Unlike conventional ICA algorithms, CICA needs prior knowledge about the rough direction of the target signal. However, it is more robust against an erroneous estimation of the target direction than adaptive beamformers: CICA converges to the right solution as long as its look direction is closer to the target signal than to the jammer signal. A high degree of robustness is very important since the geometrical prior of an adaptive beamformer is always roughly estimated in a reverberant environment, even when the look direction is precise. The effectiveness and robustness of the new algorithms is proven theoretically, and shown experimentally for three sources and three microphones with several sets of real-world data Mirko Knaak, Shoko Araki, Shoji Makino |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Grouping Separated Frequency Components by Estimating Propagation Model Parameters in Frequency-Domain Blind Source SeparationabstractThis paper proposes a new formulation and optimization procedure for grouping frequency components in frequency-domain blind source separation (BSS). We adopt two separation techniques, independent component analysis (ICA) and time–frequency (T–F) masking, for the frequency-domain BSS. With ICA, grouping the frequency components corresponds to aligning the permutation ambiguity of the ICA solution in each frequency bin. With T–F masking, grouping the frequency components corresponds to classifying sensor observations in the time–frequency domain for individual sources. The grouping procedure is based on estimating anechoic propagation model parameters by analyzing ICA results or sensor observations. More specifically, the time delays of arrival and attenuations from a source to all sensors are estimated for each source. The focus of this paper includes the applicability of the proposed procedure for a situation with wide sensor spacing where spatial aliasing may occur. Experimental results show that the proposed procedure effectively separates two or three sources with several sensor configurations in a real room, as long as the room reverberation is moderately low. Hiroshi Sawada, Shoko Araki, Ryo Mukai, Shoji Makino |
IEEE Trans. Speech Audio Process. | 4 |
| 2007 | Guest Editors' Introduction: Special Section on Emergent Systems, Algorithms and Architectures for Speech-Based Human-Machine InteractionabstractThe nine papers selected for this special section focus on emergent systems. algorithms, and architectures for speech-based human-machine interaction. The papers address a considerable number of new ideas and applications in the field of speech-based human-machine interaction. Rodrigo Capobianco Guido, Li Deng 0001, Shoji Makino |
IEEE Trans. Computers | 3 |
| 2006 | Doa Estimation for Multiple Sparse Sources with Normalized Observation Vector ClusteringabstractThis paper presents a new method for estimating the direction of arrival (DOA) of source signals whose number N can exceed the number of sensors M. Subspace based methods, e.g., the MUSIC algorithm, have been widely studied, however, they are only applicable when M > N. Another conventional independent component analysis based method allows M ges N, however, it cannot be applied when M60= 120 ms) Shoko Araki, Hiroshi Sawada, Ryo Mukai, Shoji Makino |
ICASSP (5) | 4 |
| 2006 | Blind Source Separation of Many Signals in the Frequency DomainabstractThis paper describes the frequency-domain blind source separation (BSS) of convolutively mixed acoustic signals using independent component analysis (ICA). The most critical issue related to frequency domain BSS is the permutation problem. This paper presents two methods for solving this problem. Both methods are based on the clustering of information derived from a separation matrix obtained by ICA. The first method is based on direction of arrival (DOA) clustering. This approach is intuitive and easy to understand. The second method is based on normalized basis vector clustering. This method is less intuitive than the DOA based method, but it has several advantages. First, it does not need sensor array geometry information. Secondly, it can fully utilize the information contained in the separation matrix, since the clustering is performed in high-dimensional space. Experimental results show that our methods realize BSS in various situations such as the separation of many speech signals located in a 3-dimensional space, and the extraction of primary sound sources surrounded by many background interferences Ryo Mukai, Hiroshi Sawada, Shoko Araki, Shoji Makino |
ICASSP (5) | 4 |
| 2006 | Frequency Domain Blind Source Separation of a Reduced Amount of Data Using Frequency NormalizationabstractThe problem of blind source separation (BSS) from convolutive mixtures is often addressed using independent component analysis in the frequency domain. The separation performance with this approach degrades significantly when only a short amount of data is available, since the estimation of the separation system becomes inaccurate. In this paper we present a novel approach to the frequency domain BSS using frequency normalization. Under the conditions of almost sparse sources and of dominant direct path in the mixing systems, we show that the new approach provides better performance than the conventional one when the amount of available data is small Enrique Robledo-Arnuncio, Hiroshi Sawada, Shoji Makino |
ICASSP (5) | 3 |
| 2006 | Solving the Permutation Problem of Frequency-Domain BSS when Spatial Aliasing Occurs with Wide Sensor SpacingabstractThis paper describes a method for solving the permutation problem of frequency-domain blind source separation (BSS). The method analyzes the mixing system information estimated with independent component analysis (ICA). When we use widely spaced sensors or increase the sampling rate, spatial aliasing may occur for high frequencies due to the possibility of multiple cycles in the sensor spacing. In such cases, the estimated information would imply multiple possibilities for a source location. This causes some difficulty when analyzing the information. We propose a new method designed to overcome this difficulty. This method first estimates the model parameters for the mixing system at low frequencies where spatial aliasing does not occur, and then refines the estimations by using data at all frequencies. This refinement leads to precise parameter estimation and therefore precise permutation alignment. Experimental results show the effectiveness of the new method Hiroshi Sawada, Shoko Araki, Ryo Mukai, Shoji Makino |
ICASSP (5) | 4 |
| 2006 | Underdetermined sparse source separation of convolutive mixtures with observation vector clusteringabstractWe propose a new method for solving the underdetermined sparse signal separation problem. Some sparseness based methods have already been proposed. However, most of these methods utilized a linear sensor array (or only two sensors), and therefore they have certain limitations; e.g., they cannot separate symmetrically positioned sources. To allow the use of more than three sensors that can be arranged in a non-linear/non-uniform way, we propose a new method that includes the normalization and clustering of the observation vectors. Our proposed method can handle both underdetermined case and (over-)determined cases. We show practical results for speech separation with non-linear/non-uniform sensor arrangements. We obtained promising experimental results for the cases of 3 times 4, 4 times 5 (#sensors times #sources) in a room (RT60= 120 ms) Shoko Araki, Hiroshi Sawada, Ryo Mukai, Shoji Makino |
ISCAS | 4 |
| 2006 | Stereo echo cancellation algorithm using adaptive update on the basis of enhanced input-signal vector
Satoru Emura, Youichi Haneda, Akitoshi Kataoka, Shoji Makino |
Signal Process. | 4 |
| 2006 | Blind Extraction of Dominant Target Sources Using ICA and Time-Frequency MaskingabstractThis paper presents a method for enhancing target sources of interest and suppressing other interference sources. The target sources are assumed to be close to sensors, to have dominant powers at these sensors, and to have non-Gaussianity. The enhancement is performed blindly, i.e., without knowing the position and active time of each source. We consider a general case where the total number of sources is larger than the number of sensors, and neither the number of target sources nor the total number of sources is known. The method is based on a two-stage process where independent component analysis (ICA) is first employed in each frequency bin and then time-frequency masking is used to improve the performance further. We propose a new sophisticated method for deciding the number of target sources and then selecting their frequency components. We also propose a new criterion for specifying time-frequency masks. Experimental results for simulated cocktail party situations in a room, whose reverberation time was 130 ms, are presented to show the effectiveness and characteristics of the proposed method Hiroshi Sawada, Shoko Araki, Ryo Mukai, Shoji Makino |
IEEE Trans. Speech Audio Process. | 4 |
| 2005 | Reducing musical noise by a fine-shift overlap-add method applied to source separation using a time-frequency maskabstractMusical noise is a typical problem with blind source separation using a time-frequency mask. We report that a fine-shift and overlap-add method reduces the musical noise without degrading the separation performance. The effectiveness was confirmed by results of a listening test undertaken in a room with a reverberation time of RT/sub 60/=130 ms. Shoko Araki, Shoji Makino, Hiroshi Sawada, Ryo Mukai |
ICASSP (3) | 2 |
| 2005 | A spatio-temporal fastICA algorithm for separating convolutive mixturesabstractThe paper presents a spatio-temporal extension of the well-known fastICA algorithm of Hyva/spl uml/rinen and Oja that is applicable to both convolutive blind source separation and multichannel blind deconvolution tasks. Our time-domain algorithm combines multichannel spatio-temporal prewhitening via multi-stage least-squares linear prediction with a fixed-point iteration involving a new adaptive technique for imposing paraunitary constraints on the multichannel separation filter. Our technique also allows for efficient reconstruction of individual signals as observed in the sensor measurements for single-input, multiple-output (SIMO) BSS tasks. Analysis and simulations verify the utility of the proposed methods. Scott C. Douglas, Hiroshi Sawada, Shoji Makino |
ICASSP (5) | 3 |
| 2005 | Blind extraction of a dominant source signal from mixtures of many sources [audio source separation applications]abstractThis paper presents a method for enhancing a dominant target source that is close to sensors, and suppressing other interferences. The enhancement is performed blindly, i.e. without knowing the number of total sources or information about each source, such as position and active time. We consider a general case where the number of sources is larger than the number of sensors. We employ a two-stage processing technique where a spatial filter is first employed in each frequency bin and time-frequency masking is then used to improve the performance further. To obtain the spatial filter we employ independent component analysis and then select the component of the target source. Time-frequency masks in the second stage are obtained by calculating the angle between the basis vector corresponding to the target source and a sample vector. The experimental results for a simulated cocktail party situation were very encouraging. Hiroshi Sawada, Shoko Araki, Ryo Mukai, Shoji Makino |
ICASSP (3) | 4 |
| 2005 | Natural gradient multichannel blind deconvolution and speech separation using causal FIR filtersabstractNatural gradient adaptation is an especially convenient method for adapting the coefficients of a linear system in inverse filtering tasks such as convolutive blind source separation and multichannel blind deconvolution. When developing practical implementations of such methods, however, it is not clear how best to window the signals and truncate the filter impulse responses within the filtered gradient updates. We show how inadequate use of truncation of the filter impulse responses and signal windowing within a well-known natural gradient algorithm for multichannel blind deconvolution and source separation can introduce a bias into its steady-state solution. We then provide modifications of this algorithm that effectively mitigate these effects for estimating causal FIR solutions to single- and multichannel equalization and source separation tasks. The new multichannel blind deconvolution algorithm requires approximately 6.5 multiply/adds per adaptive filter coefficient, making its computational complexity about 63% greater than the originally-proposed version. Numerical experiments verify the robust convergence performance of the new method both in multichannel blind deconvolution tasks for i.i.d. sources and in convolutive BSS tasks for real-world acoustic sources, even for extremely-short separation filters. Scott C. Douglas, Hiroshi Sawada, Shoji Makino |
IEEE Trans. Speech Audio Process. | 3 |
| 2004 | Underdetermined blind separation for speech in real environments with sparseness and ICAabstractIn this paper, we propose a method for separating speech signals when there are more signals than sensors. Several methods have already been proposed for solving the underdetermined problem, and some of these utilize the sparseness of speech signals. These methods employ binary masks to extract the signals, and therefore, their extracted signals contain loud musical noise. To overcome this problem, we propose combining a sparseness approach and independent component analysis (ICA). First, using sparseness, we estimate the time points when only one source is active. Then, we remove this single source from the observations and apply ICA to the remaining mixtures. Experimental results show that our proposed sparseness and ICA (SPICA) method can separate signals with little distortion even in reverberant conditions of T/sub R/=130 and 200 ms. Shoko Araki, Shoji Makino, Audrey Blin, Ryo Mukai, Hiroshi Sawada |
ICASSP (3) | 2 |
| 2004 | A sparseness-mixing matrix estimation (SMME) solving the underdetermined BSS for convolutive mixturesabstractWe propose a method for blindly separating real environment speech signals with as little distortion as possible in the special case where speech signals outnumber sensors. Our idea consists in combining sparseness with the use of an estimated mixing matrix. First, we use a geometrical approach to perform a preliminary separation and to detect when only one source is active. This information is then used to estimate the mixing matrix. Then we remove one source from the observations and separate the residual signals with the inverse of the estimated mixing matrix. Experimental results in a real environment (T/sub R/=130 ms and 200 ms) show that our proposed method, which we call sparseness-mixing matrix estimation (SMME), provides separated signals of better quality than those extracted by using only the sparseness property of the speech signal. Audrey Blin, Shoko Araki, Shoji Makino |
ICASSP (4) | 3 |
| 2004 | Natural gradient multichannel blind deconvolution and source separation using causal FIR filtersabstractPractical gradient-based adaptive algorithms for multichannel blind deconvolution and convolutive blind source separation typically employ FIR filters for the separation system. Inadequate use of signal truncation within these algorithms can introduce steady-state biases into their converged solutions that lead to degraded separation and deconvolution performances. We derive a natural gradient multichannel blind deconvolution and source separation algorithm that mitigates these effects for estimating causal FIR solutions to these tasks. Numerical experiments verify the robust convergence performance of the new method both in multichannel blind deconvolution tasks for i.i.d. sources and in convolutive BSS tasks for acoustic sources, even for extremely-short separation filters. Scott C. Douglas, Hiroshi Sawada, Shoji Makino |
ICASSP (5) | 3 |
| 2004 | Near-field frequency domain blind source separation for convolutive mixturesabstractThe paper presents a method for solving the permutation problem of frequency domain blind source separation (BSS) when source signals come from the same or similar directions. Geometric information such as the direction of arrival (DOA) is helpful for solving the permutation problem, and a combination of the DOA based and correlation based methods provides a robust and precise solution. However, when signals come from similar directions, the DOA based approach fails, and we have to use only the correlation based method whose performance is unstable. We show that an interpretation of the ICA solution by a near-field model yields information about spheres on which source signals exist, which can be used as an alternative to the DOA. Experimental results show that the proposed method can robustly separate a mixture of signals arriving from the same direction. Ryo Mukai, Hiroshi Sawada, Shoko Araki, Shoji Makino |
ICASSP (4) | 4 |
| 2004 | Convolutive blind source separation for more than two sources in the frequency domainabstractBlind source separation (BSS) for convolutive mixtures can be efficiently achieved in the frequency domain, where independent component analysis is performed separately in each frequency bin. However, frequency-domain BSS involves a permutation problem, which is well known as a difficult problem, especially when the number of sources is large. This paper presents a method for solving the permutation problem, which works well even for many sources. The successful solution for the permutation problem highlights another problem with frequency-domain BSS that arises from the circularity of the discrete frequency representation. This paper discusses the phenomena of the problem and presents a method for solving it. With these two methods, we can separate many sources with a practical execution time. Moreover, real-time processing is currently possible for up to three sources with our implementation. Hiroshi Sawada, Ryo Mukai, Shoko Araki, Shoji Makino |
ICASSP (3) | 4 |
| 2004 | A robust and precise method for solving the permutation problem of frequency-domain blind source separationabstractBlind source separation (BSS) for convolutive mixtures can be solved efficiently in the frequency domain, where independent component analysis (ICA) is performed separately in each frequency bin. However, frequency-domain BSS involves a permutation problem: the permutation ambiguity of ICA in each frequency bin should be aligned so that a separated signal in the time-domain contains frequency components of the same source signal. This paper presents a robust and precise method for solving the permutation problem. It is based on two approaches: direction of arrival (DOA) estimation for sources and the interfrequency correlation of signal envelopes. We discuss the advantages and disadvantages of the two approaches, and integrate them to exploit their respective advantages. Furthermore, by utilizing the harmonics of signals, we make the new method robust even for low frequencies where DOA estimation is inaccurate. We also present a new closed-form formula for estimating DOAs from a separation matrix obtained by ICA. Experimental results show that our method provided an almost perfect solution to the permutation problem for a case where two sources were mixed in a room whose reverberation time was 300 ms. Hiroshi Sawada, Ryo Mukai, Shoko Araki, Shoji Makino |
IEEE Trans. Speech Audio Process. | 4 |
| 2003 | Subband based blind source separation for convolutive mixtures of speechabstractSubband processing is applied to blind source separation (BSS) for convolutive mixtures of speech. This is motivated by the drawback of frequency-domain BSS, i.e., when a long frame with a fixed frame-shift is used to cover reverberation, the number of samples in each frequency decreases and the separation performance is degraded. In our proposed subband BSS, (1) by using a moderate number of subbands, a sufficient number of samples can be held in each subband, mid (2) by using FIR filters in each subband, we can handle long reverberation. Subband BSS achieves better performance than frequency-domain BSS. Moreover, we propose efficient separation procedures that take into consideration the frequency characteristics of room reverberation and speech signals. We achieve this (3) by using longer unmixing filters in low frequency bands, and (4) by adopting overlap-blockshift in BSS's batch adaptation in low frequency bands. Consequently, frequency-dependent subband processing is successfully realized in the proposed subband BSS. Shoko Araki, Shoji Makino, Robert Aichner, Tsuyoki Nishikawa, Hiroshi Saruwatari |
ICASSP (5) | 2 |
| 2003 | Geometrically constraint ICA for convolutive mixtures of soundabstractThe goal of this contribution is a new algorithm using independent component analysis with a geometrical constraint. The new algorithm solves the permutation problem of blind source separation of acoustic mixtures, and it is significantly less sensitive to the precision of the geometrical constraint than an adaptive beamformer. A high degree of robustness is very important since the steering vector is always roughly estimated in the reverberant environment, even when the look direction is precise. The new algorithm is based on FastICA and constrained optimization. It is theoretically and experimentally analyzed with respect to the roughness of the steering vector estimation by using impulse responses of real room. The effectiveness of the algorithms for real-world mixtures is also shown in the case of three sources and three microphones. Mirko Knaak, Shoko Araki, Shoji Makino |
ICASSP (2) | 3 |
| 2003 | Robust real-time blind source separation for moving speakers in a roomabstractThis paper describes a robust real-time blind source separation (BSS) method for moving speech signals in a room. Our method employs frequency domain independent component analysis (ICA) using a blockwise batch algorithm in the first stage, and the separated signals are refined by postprocessing using crosstalk component estimation and nonstationary spectral subtraction in the second stage. The blockwise batch algorithm achieves better performance than an online algorithm when sources are fixed, and the postprocessing compensates for performance degradation caused by source movement. Experimental results using speech signals recorded in a real room show that the proposed method realizes robust real-time separation for moving sources. Our method is implemented on a standard PC and works in real time. Ryo Mukai, Hiroshi Sawada, Shoko Araki, Shoji Makino |
ICASSP (5) | 4 |
| 2003 | A robust approach to the permutation problem of frequency-domain blind source separationabstractThis paper presents a robust and precise method for solving the permutation problem of frequency-domain blind source separation. It is based on two previous approaches: the direction of arrival estimation approach and the inter-frequency correlation approach. We discuss the advantages and disadvantages of the two approaches, and integrate them to exploit the both advantages. We also present a closed form formula to calculate a null direction, which is used in estimating the directions of source signals. Experimental results show that our method solved permutation problems almost perfectly for a situation that two sources were mixed in a room whose reverberation time was 300 ms. Hiroshi Sawada, Ryo Mukai, Shoko Araki, Shoji Makino |
ICASSP (5) | 4 |
| 2003 | Geometrical understanding of the PCA subspace method for overdetermined blind source separationabstractWe discuss approaches for blind source separation where we can use more sensors than the number of sources for a better performance. The discussion focuses mainly on reducing the dimension of mixed signals before applying independent component analysis. We compare two previously proposed methods. The first is based on principal component analysis, where noise reduction is achieved. The second involves selecting a subset of sensors based on the fact that a low frequency prefers a wide spacing and a high frequency prefers a narrow spacing. We found that the PCA-based method behaves similarly to the geometry-based method for low frequencies in the way that it emphasizes the outer sensors and yields superior results for high frequencies, which provides a better understanding of the former method. Stefan Winter 0002, Hiroshi Sawada, Shoji Makino |
ICASSP (2) | 3 |
| 2003 | The fundamental limitation of frequency domain blind source separation for convolutive mixtures of speechabstractDespite several recent proposals to achieve blind source separation (BSS) for realistic acoustic signals, the separation performance is still not good enough. In particular, when the impulse responses are long, performance is highly limited. In this paper, we consider a two-input, two-output convolutive BSS problem. First, we show that it is not good to be constrained by the condition T>P, where T is the frame length of the DFT and P is the length of the room impulse responses. We show that there is an optimum frame size that is determined by the trade-off between maintaining the number of samples in each frequency bin to estimate statistics and covering the whole reverberation. We also clarify the reason for the poor performance of BSS in long reverberant environments, highlighting that the framework of BSS works as two sets of frequency-domain adaptive beamformers. Although BSS can reduce reverberant sounds to some extent like adaptive beamformers, they mainly remove the sounds from the jammer direction. This is the reason for the difficulty of BSS in reverberant environments. Shoko Araki, Ryo Mukai, Shoji Makino, Tsuyoki Nishikawa, Hiroshi Saruwatari |
IEEE Trans. Speech Audio Process. | 3 |
| 2002 | Equivalence between frequency domain blind source separation and frequency domain adaptive beamformingabstractFrequency domain Blind Source Separation (BSS) is shown to be equivalent to two sets of frequency domain adaptive microphone arrays, i.e., Adaptive Beamformers (ABFs). The minimization of the off-diagonal components in the BSS update equation can be viewed as the minimization of the mean square error in the ABF. The unmixing matrix of the BSS and the filter coefficients of the ABF converge to the same solution in the mean square error sense if the two source signals are ideally independent. Therefore, the performance of the BSS is limited by that of the ABF. This understanding. gives an interpretation of BSS from physical point of view. Shoko Araki, Yoichi Hinamoto, Shoji Makino, Tsuyoki Nishikawa, Ryo Mukai, Hiroshi Saruwatari |
ICASSP | 3 |
| 2002 | Enhanced frequency-domain adaptive algorithm for stereo echo cancellationabstractHighly cross-correlated input signals create the problem of slow convergence of misalignment in stereo echo cancellation even after undergoing non-linear preprocessing. We propose a new frequency-domain adaptive algorithm that improves the convergence rate by increasing the contribution of non-linearity in the adjustment vector. Computer simulation showed that it is effective when the non-linearity gain is small. Satoru Emura, Youichi Haneda, Shoji Makino |
ICASSP | 3 |
| 2002 | Removal of residual cross-talk components in Blind Source Separation using time-delayed spectral subtractionabstractThis paper describes a post processing method to refine output signals obtained by Blind Source Separation (BSS). The performance of BSS using Independent Component Analysis (ICA) declines significantly in a reverberant environment. The degradation is mainly caused by the cross-talk components derived from the reverberation of the jammer signal. Utilizing this knowledge, we propose a new method, time-delayed non-stationary spectral subtraction, which removes the residual components from the separated signals precisely. The proposed method compensates for the weakness of BSS in a reverberant environment. Experimental results using speech signals show that the proposed method improves the signal-to-noise ratio by 3 to 5 dB. Ryo Mukai, Shoko Araki, Hiroshi Sawada, Shoji Makino |
ICASSP | 4 |
| 2002 | Polar coordinate based nonlinear function for frequency-domain blind source separationabstractThis paper presents a new type of nonlinear function for independent component analysis to process complex-valued signals, which is used in frequency-domain blind source separation. The new function is based on the polar coordinates of a complex number, whereas the conventional one is based on the Cartesian coordinates. The new function is derived from the probability density function of frequency-domain signals that are assumed to be independent of the phase. We show that the difference between the two types of functions is in the assumed densities of independent components. Experimental results for separating speech signals show that the new nonlinear function behaves better than the conventional one. Hiroshi Sawada, Ryo Mukai, Shoko Araki, Shoji Makino |
ICASSP | 4 |
| 2001 | Fundamental limitation of frequency domain blind source separation for convolutive mixture of speechabstractDespite several recent proposals to achieve blind source separation (BSS) for realistic acoustic signals, separation performance is still not good enough. In particular, when the length of impulse response is long, performance is highly limited. We show it is useless to be constrained by the condition, P /spl Lt/ T, where T is the frame size of FFT and P is the length of room impulse response. From our experiments. a frame size of 256 or 512 (32 or 64 ms at a sampling frequency of 8 kHz) is best even for the long room reverberation of T/sub R/ = 150 and 300 ms. We also clarified the reason for poor performance of BSS in a long reverberant environment, finding that separation is achieved chiefly for the sound from the direction of jammers because BSS cannot calculate the inverse of the room transfer function both for the target and jammer signals. Shoko Araki, Shoji Makino, Tsuyoki Nishikawa, Hiroshi Saruwatari |
ICASSP | 2 |
| 2001 | Equivalence between frequency domain blind source separation and frequency domain adaptive null beamformersabstractFrequency domain Blind Source Separation (BSS) is shown to be equivalent to two sets of frequency domain adaptive microphone arrays, that is, Adaptive Null Beamformers (ANB). The unmixing matrix of the BSS and the filter coefficients of the ANB converge to the same solution in the mean square error sense if the two source signals are ideally independent. This understanding clearly explains the poor performance of the BSS in a real room with long reverberation. The fundamental difference exists in the adaptation period when they should adapt. That is, the ANB can adapt in the presence of a jammer but the absence of a target, whereas the BSS can adapt in the presence of a target and jammer, and also in the presence of only a target. Shoko Araki, Shoji Makino, Ryo Mukai, Hiroshi Saruwatari |
INTERSPEECH | 2 |
| 2001 | Separation and dereverberation performance of frequency domain blind source separation for speech in a reverberant environmentabstractIn this paper, we investigate the separation and dereverberation performance of frequency domain Blind Source Separation (BSS) based on Independent Component Analysis (ICA) by measuring impulse responses of a system. Since ICA is a statistical method, i.e., it only attempts to make outputs independent, it is not easy to predict what is going on in a BSS system physically. We therefore investigate the detailed components in the processed signals of a whole BSS system from a physical and acoustical viewpoint. In particular, we focus on the direct sound and reverberation in the target and jammer signals. As a result, we reveal that the direct sound of a jammer can be removed and the reverberation of the jammer can be reduced to some degree by BSS, while the reverberation of the target cannot be reduced. Moreover, we show that a long frame length causes pre-echo noise, and this damages the quality of the separated signal. 1. Ryo Mukai, Shoko Araki, Shoji Makino |
INTERSPEECH | 3 |
| 2000 | Channel-number-compressed multi-channel acoustic echo canceller for high-presence teleconferencing system with large displayabstractSound localization is important to make conversation easy between local and remote sites in a teleconference. This requires a multi-channel sound system having a multi-channel acoustic echo canceller (MAEC). The appropriate number of channels is determined from a trade-off between high presence and MAEC performance, so it is not possible to increase the channel number by much. We propose a channel-number-compressed MAEC to provide teleconferencing systems that exhibit high presence. The channel number of the MAEC inputs is compressed and that of its outputs is expanded. Akira Nakagawa, Suehiro Shimauchi, Youichi Haneda, Shigeaki Aoki, Shoji Makino |
ICASSP | 5 |
| 1999 | A stereo echo canceller implemented using a stereo shaker and a duo-filter control systemabstractStereo echo cancellation has been achieved and used in daily teleconferencing. To overcome the non-uniqueness problem, a stereo shaker is introduced in eight frequency bands and adjusted so as to be inaudible and not affect stereo perception. A due-filter control system including a continually running adaptive filter and a fixed filter is used for double-talk control. A second-order stereo projection algorithm is used in the adaptive filter. A stereo voice switch is also included. This stereo echo canceller was tested in two-way conversation in a conference room, and the strength of the stereo shaker was subjectively adjusted. A misalignment of 20 dB was obtained in the teleconferencing environment, and changing the talker's position in the transmission room did not affect the cancellation. This echo canceller is now used daily in a high-presence teleconferencing system and has been demonstrated to more than 300 attendees. Suehiro Shimauchi, Shoji Makino, Youichi Haneda, Akira Nakagawa, Sumitaka Sakauchi |
ICASSP | 2 |
| 1999 | Common-acoustical-pole and zero modeling of head-related transfer functionsabstractUse of a common-acoustical-pole and zero model is proposed for modeling head-related transfer functions (HRTFs) for various directions of sound incidence. The HRTFs are expressed using the common acoustical poles, which do not depend on the source directions, and the zeros, which do. The common acoustical poles are estimated as they are common to HRTFs for various source directions; the estimated values of the poles agree well with the resonance frequencies of the ear canal. Because this model uses only the zeros to express the HRTF variations due to changes in source direction, it requires fewer parameters (the order of the zeros) that depend on the source direction than do the conventional all-zero or pole/zero models. Furthermore, the proposed model can extract the zeros that are missed in the conventional models because of pole-zero cancellation. As a result, the directional dependence of the zeros can be traced well. Analysis of the zeros for HRTFs on the horizontal plane showed that the nonminimum-phase zero variation was well formulated using a simple pinna-reflection model. The common-acoustical-pole and zero (CAPZ) model is thus effective for modeling and analyzing HRTF's. Youichi Haneda, Shoji Makino, Yutaka Kaneda, Nobuhiko Kitawaki |
IEEE Trans. Speech Audio Process. | 2 |
| 1999 | A block exact fast affine projection algorithmabstractThis paper describes a block (affine) projection algorithm that has exactly the same convergence rate as the original sample-by-sample algorithm and smaller computational complexity than the fast affine projection algorithm. This is achieved by (1) introducing a correction term that compensates for the filter output difference between the sample-by-sample projection algorithm and the straightforward block projection algorithm, and (2) applying a fast finite impulse response (FIR) filtering technique to compute the filter outputs and to update the filter. We describe how to choose a pair of block lengths that gives the longest filter length under a constraint on the total computational complexity and processing delay. An example shows that the filter length can be doubled if a delay of a few hundred samples is permissible. Masashi Tanaka, Shoji Makino, Junji Kojima |
IEEE Trans. Speech Audio Process. | 2 |
| 1998 | New configuration for a stereo echo canceller with nonlinear pre-processingabstractA new configuration for a stereo echo canceller with nonlinear pre-processing is proposed. The pre-processor which adds uncorrelated components to the original received stereo signals improves the adaptive filter convergence even in the conventional configuration. However, because of the inaudibility restriction, the preprocessed signals still include a large amount of the original stereo signals which are often highly cross-correlated. Therefore, the improvement is limited. To overcome this, our new stereo echo canceller includes exclusive adaptive filters whose inputs are the uncorrelated signals generated in the pre-processor. These exclusive adaptive filters converge to true solutions without suffering from cross-correlation between the original stereo signals. This is demonstrated through computer simulation results. Suehiro Shimauchi, Youichi Haneda, Shoji Makino, Yutaka Kaneda |
ICASSP | 3 |
| 1997 | Subband stereo echo canceller using the projection algorithm with fast convergence to the true echo pathabstractThis paper proposes a new subband stereo echo canceller that converges to the true echo path impulse response much faster than conventional stereo echo cancellers. Since signals are bandlimited and downsampled in the subband structure, the time interval between the subband signals become longer, so the variation of the crosscorrelation between the stereo input signals becomes large. Consequently, convergence to the true solution is improved. Furthermore, the projection algorithm, or affine projection algorithm, is applied to further speed up the convergence. Computer simulations using stereo signals recorded in a conference room demonstrate that this method significantly improves convergence speed and almost solves the problem of stereo echo cancellation with low computational load. Shoji Makino, Klaus Strauss, Suehiro Shimauchi, Youichi Haneda, Akira Nakagawa |
ICASSP | 1 |
| 1997 | Multiple-point equalization of room transfer functions by using common acoustical polesabstractA multiple-point equalization filter using the common acoustical poles of room transfer functions is proposed. The common acoustical poles correspond to the resonance frequencies, which are independent of source and receiver positions. They are estimated as common autoregressive (AR) coefficients from multiple room transfer functions. The equalization is achieved with a finite impulse response (FIR) filter, which has the inverse characteristics of the common acoustical pole function. Although the proposed filter cannot recover the frequency response dips of the multiple room transfer functions, it can suppress their common peaks due to resonance; it is also less sensitive to changes in receiver position. Evaluation of the proposed equalization filter using measured room transfer functions shows that it can reduce the deviations in the frequency characteristics of multiple room transfer functions better than a conventional multiple-point inverse filter. Experiments show that the proposed filter enables 1-5 dB additional amplifier gain in a public address system without acoustic feedback at multiple receiver positions. Furthermore, the proposed filter reduces the reflected sound in room impulse responses without the pre-echo that occurs with a multiple-point inverse filter. A multiple-point equalization filter using common acoustical poles can thus equalize multiple room transfer functions by suppressing their common peaks. Youichi Haneda, Shoji Makino, Yutaka Kaneda |
IEEE Trans. Speech Audio Process. | 2 |
| 1996 | SSB subband echo canceller using low-order projection algorithmabstractThis paper proposes a new subband echo canceller that almost fully whitens the received input by using the low-order projection algorithm, or the affine projection algorithm. Since the projection algorithm can fully whiten the received input of a small-tap adaptive filter with a relatively small projection order, the proposed subband projection echo canceller achieves nearly maximum convergence with a small projection order. By reflecting the frequency characteristics of the speech and of the echo path, it allows a different projection order to be chosen in different subbands. This gives the proposed method optimum cost performance, which is promising for implementation. Shoji Makino, Josef Nöbauer, Youichi Haneda, Akira Nakagawa |
ICASSP | 1 |
| 1996 | Stereo echo cancellation algorithm using imaginary input-output relationshipsabstractA stereo teleconferencing system provides greater presence in teleconferencing compared with a monaural system. It helps listeners distinguish who is talking at the other end by means of spatial information. The most important problem in stereo echo cancellation is that the adaptive filter often misconverges or, if not, its convergence speed is very slow because of the crosscorrelation between stereo signals. A new stereo echo cancellation algorithm using imaginary input-output relationships is proposed. This algorithm is based on the idea of approximating the input-output relationships for the reversed stereo signals. Unlike conventional methods, which use the input-output relationships in only one situation, our algorithm uses the simultaneous equations for the input-output relationships under two different conditions. Consequently, convergence to the unique steady-state solution is faster than with the conventional algorithm. Computer simulations using data recorded in a conference room demonstrate the effectiveness of this algorithm. Suehiro Shimauchi, Shoji Makino |
ICASSP | 2 |
| 1995 | Stereo projection echo canceller with true echo path estimationabstractThe effect that cross-correlation between stereo signals has on a stereo echo canceller is studied and a new stereo projection echo canceller is proposed that can identify the true echo path impulse responses. This echo canceller accelerates filter coefficient error convergence by adding the variations in the cross-correlation to stereo signals or by utilizing the fact that the cross-correlation between stereo signals varies slightly in actual teleconferencing situations. Computer simulations demonstrate that this echo canceller can reduce the filter coefficient error much faster than a conventional stereo normalized least mean squares (NLMS) echo canceller. Suehiro Shimauchi, Shoji Makino |
ICASSP | 2 |
| 1995 | Fast projection algorithm and its step size controlabstractOf the many adaptive filtering algorithms, the normalized LMS (NLMS) algorithm is generally used in practice because of its simplicity. The computational complexity of the NLMS algorithm is low, however, convergence is very slow and tracking is poor for a colored input signal such as speech. The projection algorithm was proposed as a generalization of the NLMS algorithm. This paper provides a fast projection algorithm and a step size control to obtain the same steady-state excess mean squared error (MSE) for various projection orders. Computer simulations for colored noise and speech input signal confirm the effectiveness of the projection algorithm and the step size control. Masashi Tanaka, Yutaka Kaneda, Shoji Makino, Junji Kojima |
ICASSP | 3 |
| 1994 | A new RLS algorithm based on the variation characteristics of a room impulse responseabstractThe paper proposes a new adaptive algorithm (called the ES-RLS algorithm) with double the convergence speed of the conventional RLS algorithm. Makino et al. (1993) showed that the variation of a room impulse response becomes progressively smaller along the series by the same exponential ratio as the impulse response. The ES-RLS algorithm is derived by incorporating these variation characteristics into the conventional RLS algorithm using Kalman filter theory, which gives physical meaning to the RLS algorithm. The ES-RLS algorithm adjusts coefficients with large errors in large steps and coefficients with small errors in small steps. Computer simulations demonstrated that the new adaptive algorithm converged twice as fast as the conventional RLS algorithm.> Shoji Makino, Yutaka Kaneda |
ICASSP (3) | 1 |
| 1994 | Common acoustical pole and zero modeling of room transfer functionsabstractA new model for a room transfer function (RTF) by using common acoustical poles that correspond to resonance properties of a room is proposed. These poles are estimated as the common values of many RTF's corresponding to different source and receiver positions. Since there is one-to-one correspondence between poles and AR coefficients, these poles are calculated as common AR coefficients by two methods: (i) using the least squares method, assuming all the given multiple RTF's have the same AR coefficients and (ii) averaging each set of AR coefficients estimated from each RTF. The estimated poles agree well with the theoretical poles when estimated with the same order as the theoretical pole order. When estimated with a lower order than the theoretical pole order, the estimated poles correspond to the major resonance frequencies, which have high Q factors. Using the estimated common AR coefficients, the proposed method models the RTF's with different MA coefficients. This model is called the common-acoustical-pole and zero (CAPZ) model, and it requires far fewer variable parameters to represent RTF's than the conventional all-zero or pole/zero model. This model was used for an acoustic echo canceller at low frequencies, as one example. The acoustic echo canceller based on the proposed model requires half the variable parameters and converges 1.5 times faster than one based on the all-zero model, confirming the efficiency of the proposed model.> Youichi Haneda, Shoji Makino, Yutaka Kaneda |
IEEE Trans. Speech Audio Process. | 2 |
| 1993 | Exponentially weighted stepsize NLMS adaptive filter based on the statistics of a room impulse responseabstractA normalized least-mean-squares (NLMS) adaptive algorithm with double the convergence speed, at the same computational load, of the conventional NLMS for an acoustic echo canceller is proposed. This algorithm, called the ES (exponentially weighted stepsize) algorithm, uses a different stepsize (feedback constant) for each weight of an adaptive transversal filter. These stepsizes are time-invariant and weighted proportionally to the expected variation of a room impulse response. The algorithm adjusts coefficients with large errors in large steps, and coefficients with small errors in small steps. A transition formula is derived for the mean-squared coefficient error of the algorithm. The mean stepsize determines the convergence condition, the convergence speed, and the final excess mean-squared error. Modified for a practical multiple DSP structure, the algorithm requires only the same amount of computation as the conventional NLMS. The algorithm is implemented in a commercial acoustic echo canceller, and its fast convergence is demonstrated.> Shoji Makino, Yutaka Kaneda, Nobuo Koizumi |
IEEE Trans. Speech Audio Process. | 1 |
| 1992 | Modeling of a room transfer function using common acoustical polesabstractA method for modeling a room transfer function (RTF) by using estimated common acoustical poles that correspond to resonance properties of a room is proposed. These poles are estimated as common values of the multiple RTFs corresponding to different source and receiver positions. This common-acoustical-pole and zero (CAPZ) model requires far fewer variable parameters to represent RTFs than conventional all-zero or pole/zero models. This model was applied to an acoustic echo canceller and to head-related transfer functions. At low frequencies, the acoustic echo canceller based on this model converges 1.5 times faster than the one based on the all-zero model. Head-related transfer functions that have resonance characteristics of the external ear are also successfully modeled by the proposed model.> Youichi Haneda, Shoji Makino, Yutaka Kaneda |
ICASSP | 2 |
| 1990 | Acoustic echo canceller algorithm based on the variation characteristics of a room impulse responseabstractA NLMS (normalized LMS) adaptive algorithm is proposed for an acoustic echo canceller with double the convergence speed of but the same computational load as the conventional NLMS. This algorithm, called the ES (exponential step) algorithm, implements a different value of step gain (feedback factor) for each tap coefficient of a canceller. These step gains are weighted exponentially along the series by the same exponential ratio as the expected variation in a room impulse response. This algorithm is implemented in a commercial subband echo canceller, and its superiority to the conventional algorithm is demonstrated.> Shoji Makino, Yutaka Kaneda |
ICASSP | 1 |