Hyung-Min Park

dblp:75/1270 · DBLP profile ↗
← Back
52ranked-venue papers
10as first author
11since 2021 · last 2026
0000-0002-7105-5493ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 25 · 5 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Beyond Noise Suppression: Dynamic Distortion Control Loss for Speech Enhancement and Robust Automatic Speech Recognition
abstract
In severe noise conditions, employing a speech enhancement (SE) model as a front-end serves as a computationally efficient strategy for robust automatic speech recognition (ASR), offering a practical alternative to the costly fine-tuning of largescale ASR systems. However, improvements in human perceptual quality do not necessarily guarantee enhanced machine recognition accuracy, as aggressive noise suppression often introduces distortions and artifacts that obscure fine-grained spectral details and increase word error rates (WERs). To mitigate the discrepancy, we present the Dynamic Distortion Control (DDC) loss, a unified training objective designed to bridge the gap between perceptual fidelity and recognition robustness. Integrated into a time-frequency Transformer architecture, the presented loss addresses the distortion-robustness discrepancy. Experimental results on a LibriSpeech dataset corrupted by noise from DNS Challenge demonstrate the effectiveness of the DDC loss on both perceptual quality and recognition accuracy across diverse noise conditions.
Seung-Jin Kim, Hyung-Min Park
IEEE Signal Process. Lett.2
2025 Stack Less, Repeat More: A Block Reusing Approach for Progressive Speech Enhancement
Jangyeon Kim, Ui-Hyeop Shin, Jaehyun Ko 0003, Hyung-Min Park
INTERSPEECH4
2025 SwinLip: An Efficient Visual Speech Encoder for Lip Reading Using Swin Transformer
Young-Hu Park, Rae-Hong Park, Hyung-Min Park
Neurocomputing3
2025 TF-CorrNet: Leveraging Spatial Correlation for Continuous Speech Separation
abstract
In general, multi-channel source separation has utilized inter-microphone phase differences (IPDs) concatenated with magnitude information in time-frequency domain, or real and imaginary components stacked along the channel axis. However, the spatial information of a sound source is fundamentally contained in the “differences” between microphones, specifically in the correlation between them, while the power of each microphone also provides valuable information about the source spectrum, which is why the magnitude is also included. Therefore, we propose a network that directly leverages a correlation input with phase transform (PHAT)-$\beta$to estimate the separation filter. In addition, the proposed TF-CorrNet processes the features alternately across time and frequency axes as a dual-path strategy in terms of spatial information. Furthermore, we add a spectral module to model source-related direct time-frequency patterns for improved speech separation. Experimental results demonstrate that the proposed TF-CorrNet effectively separates the speech sounds, showing high performance with a low computational cost in the LibriCSS dataset.
Ui-Hyeop Shin, Bon Hyeok Ku, Hyung-Min Park
IEEE Signal Process. Lett.3
2024 NeXt-TDNN: Modernizing Multi-Scale Temporal Convolution Backbone for Speaker Verification
abstract
In speaker verification, ECAPA-TDNN has shown remarkable improvement by utilizing one-dimensional(1D) Res2Net block and squeeze-and-excitation(SE) module, along with multi-layer feature aggregation (MFA). Meanwhile, in vision tasks, ConvNet structures have been modernized by referring to Transformer, resulting in improved performance. In this paper, we present an improved block design for TDNN in speaker verification. Inspired by recent ConvNet structures, we replace the SE-Res2Net block in ECAPA-TDNN with a novel 1D two-step multi-scale ConvNeXt block, which we call TS-ConvNeXt. The TS-ConvNeXt block is constructed using two separated sub-modules: a temporal multi-scale convolution (MSC) and a frame-wise feed-forward network (FFN). This two-step design allows for flexible capturing of inter-frame and intra-frame contexts. Additionally, we introduce global response normalization (GRN) for the FFN modules to enable more selective feature propagation, similar to the SE module in ECAPA-TDNN. Experimental results demonstrate that NeXt-TDNN, with a modernized backbone block, significantly improved performance in speaker verification tasks while reducing parameter size and inference time. We have released our code1for future studies.
Hyunjun Heo, Ui-Hyeop Shin, Ran Lee, Youngju Cheon, Hyung-Min Park
ICASSP5
2024 OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset
abstract
Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, developed from pre-existing videos using various prediction models, and have only a small number of multi-view videos. To mitigate the limitations, we constructed the Open Large-scale Korean Audio-Visual Speech (OLKAVS) dataset, which is the largest among publicly available audio-visual speech datasets. The dataset contains 1,150 hours of transcribed audio from 1,107 Korean speakers in a studio setup with nine different viewpoints and various noise situations. We also provide the pre-trained baseline models for two tasks: audiovisual speech recognition and lip reading. We conducted experiments based on the models to verify the effectiveness of multi-modal and multi-view training over uni-modal and frontal-view-only training. We expect the OLKAVS dataset to facilitate multi-modal research in broader areas.
Jeongkyun Park, Jung-Wook Hwang, Kwanghee Choi, Seung-Hyeon Lee, Jun Hwan Ahn, Rae-Hong Park, Hyung-Min Park
ICASSP7
2024 Separate and Reconstruct: Asymmetric Encoder-Decoder for Speech Separation
abstract
In speech separation, time-domain approaches have successfully replaced the time-frequency domain with latent sequence feature from a learnable encoder. Conventionally, the feature is separated into speaker-specific ones at the final stage of the network. Instead, we propose a more intuitive strategy that separates features earlier by expanding the feature sequence to the number of speakers as an extra dimension. To achieve this, an asymmetric strategy is presented in which the encoder and decoder are partitioned to perform distinct processing in separation tasks. The encoder analyzes features, and the output of the encoder is split into the number of speakers to be separated. The separated sequences are then reconstructed by the weight-shared decoder, which also performs cross-speaker processing. Without relying on speaker information, the weight-shared network in the decoder directly learns to discriminate features using a separation objective. In addition, to improve performance, traditional methods have extended the sequence length, leading to the adoption of dual-path models, which handle the much longer sequence effectively by segmenting it into chunks. To address this, we introduce global and local Transformer blocks that can directly handle long sequences more efficiently without chunking and dual-path processing. The experimental results demonstrated that this asymmetric structure is effective and that the combination of proposed global and local Transformer can sufficiently replace the role of inter- and intra-chunk processing in dual-path structure. Finally, the presented model combining both of these achieved state-of-the-art performance with much less computation in various benchmark datasets.
Ui-Hyeop Shin, Sangyoun Lee, Taehan Kim, Hyung-Min Park
NeurIPS4
2024 Statistical Beamformer Exploiting Non-Stationarity and Sparsity With Spatially Constrained ICA for Robust Speech Recognition
abstract
In this paper, we present a statistical beamforming algorithm as a pre-processing step for robust automatic speech recognition (ASR). By modeling the target speech as a non-stationary Laplacian distribution, a mask-based statistical beamforming algorithm is proposed to exploit both its output and masked input variance for robust estimation of the beamformer. In addition, we also present a method for steering vector estimation (SVE) based on a noise power ratio obtained from the target and noise outputs in independent component analysis (ICA). To update the beamformer in the same ICA framework, we derive ICA with distortionless and null constraints on target speech, which yields beamformed speech at the target output and noises at the other outputs, respectively. The demixing weights for the target output result in a statistical beamformer with the weighted spatial covariance matrix (wSCM) using a weighting function characterized by a source model. To enhance the SVE, the strict null constraints imposed by the Lagrange multiplier methods are relaxed by generalized penalties with weight parameters, while the strict distortionless constraints are maintained. Furthermore, we derive an online algorithm based on an optimization technique of recursive least squares (RLS) for practical applications. Experimental results on various environments using CHiME-4 and LibriCSS datasets demonstrate the effectiveness of the presented algorithm compared to conventional beamforming and blind source extraction (BSE) based on ICA on both batch and online processing.
Ui-Hyeop Shin, Hyung-Min Park
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Minimax Monte Carlo object tracking
Jaechan Lim, Jin-Young Park, Hyung-Min Park
Vis. Comput.3
2022 Distilling a Pretrained Language Model to a Multilingual ASR Model
Kwanghee Choi, Hyung-Min Park
INTERSPEECH2
2021 Convolutional Maximum-Likelihood Distortionless Response Beamforming With Steering Vector Estimation for Robust Speech Recognition
abstract
Beamforming has been one of the most successful approaches using multi-microphones for robust speech recognition. Although a beamforming method, called the “maximum-likelihood distortionless response (MLDR)” beamformer, was recently presented to achieve promising performance, it requires an accurate steering vector for a target speaker in advance like many kinds of beamformers. In this paper, we present a method for steering vector estimation (SVE) by replacing the noise spatial covariance matrix estimate with a normalized version of the variance-weighted spatial covariance matrix estimate for the observed noisy speech signal obtained by the iterative update rule in the MLDR beamforming framework. In addition, an MLDR beamforming method without a steering vector for a target speaker given in advance is presented where the SVE and the beamforming are alternately repeated. Furthermore, an online algorithm based on recursive least squares (RLS) is derived to cope with various practical applications including time-varying situations, and the power method is introduced for further efficient online processing. We also present batch and online convolutional MLDR beamforming with SVE for simultaneous beamforming and dereverberation where the weighted prediction error (WPE) dereverberation and the MLDR beamforming with the SVE were jointly optimized based on the maximum-likelihood estimation (MLE) for a zero-mean complex Gaussian signal with time-varying variances. Moreover, input signals masked by a neural network (NN) for estimating target speech or noise components can be used to further improve the presented beamformers. Experimental results on the CHiME-4 and REVERB challenge datasets demonstrate the effectiveness of the presented methods.
Byung Joon Cho, Hyung-Min Park
IEEE ACM Trans. Audio Speech Lang. Process.2
2020 Interactive-Multiple-Model Algorithm Based on Minimax Particle Filtering
abstract
In this letter, we propose a new approach to tracking a target that maneuvers based on the multiple-constant-turns model. Usually, the interactive-multiple-model (IMM) algorithm based on the extended Kalman filter (IMM-EKF) is employed for this problem with successful tracking performance. Recently proposed IMM-particle filtering (IMM-PF) showed outperforming results over IMM-EKF for this nonlinear problem. The proposed approach in this letter is a new framework of PF that adopts the minimax strategy to IMM-PF. The minimax strategy results in the decreased variance of the weights of particles that provides the robustness against the degeneracy phenomenon (a common problem of generic PF). In this letter, we show outperforming results by IMM-minimax-PF over IMM-PF besides the IMM-EKF in terms of estimation accuracy and computational complexity.
Jaechan Lim, Hun-Seok Kim, Hyung-Min Park
IEEE Signal Process. Lett.3
2019 A Beamforming Algorithm Based on Maximum Likelihood of a Complex Gaussian Distribution With Time-Varying Variances for Robust Speech Recognition
abstract
This letter addresses acoustic beamforming for robust speech recognition. Beamforming has been demonstrated to be one of the most effective approaches for robust recognition of distant speech using multi-microphones. In this letter, we derive a beamforming method, which we refer to as the “maximum-likelihood distortionless response (MLDR)” beamformer, based on the maximum-likelihood estimation (MLE) of a linear filter, with a distortionless constraint on the steering direction, assuming that the optimal beamformer outputs in the time-frequency domain follow a zero-mean complex Gaussian distribution with time-varying variances. By optimizing the beamformer output variances as well as the filter alternately with iterative update rules, and also by using the moving average of output powers at adjacent frames for robust estimation of an output variance, the MLDR beamformer may minimize the power of a relatively accurate noise component at the output, which resulted in better recognition performance than conventional beamformers. In addition, it can be further improved by initializing the output variances to averaged powers of a neural-network(NN)-masked input signal to estimate target speech powers, which achieved even better performance than compared beamformers exploiting trained NNs.
Byung Joon Cho, Jun-Min Lee, Hyung-Min Park
IEEE Signal Process. Lett.3
2017 Bayesian feature enhancement using independent vector analysis and reverberation parameter re-estimation for noisy reverberant speech recognition
Ji-Won Cho, Jong-Hyeon Park, Joon-Hyuk Chang, Hyung-Min Park
Comput. Speech Lang.4
2016 Ensemble of deep neural networks using acoustic environment classification for statistical model-based voice activity detection
Inyoung Hwang, Hyung-Min Park, Joon-Hyuk Chang
Comput. Speech Lang.2
2016 Independent vector analysis followed by HMM-based feature enhancement for robust speech recognition
Ji-Won Cho, Hyung-Min Park
Signal Process.2
2016 A Subband-Based Stationary-Component Suppression Method Using Harmonics and Power Ratio for Reverberant Speech Recognition
abstract
This letter describes a preprocessing method called subband-based stationary-component suppression method using harmonics and power ratio (SHARP) processing for reverberant speech recognition. SHARP processing extends a previous algorithm called Suppression of Slowly varying components and the Falling edge (SSF), which suppresses the steady-state portions of subband spectral envelopes. The SSF algorithm tends to over-subtract these envelopes in highly reverberant environments when there are high levels of power in previous analysis frames. The proposed SHARP method prevents excessive suppression both by boosting the floor value using the harmonics in voiced speech segments and by inhibiting the subtraction for unvoiced speech by detecting frames in which power is concentrated in high-frequency channels. These modifications enable the SHARP algorithm to improve recognition accuracy by further reducing the mismatch between power contours of clean and reverberated speech. Experimental results indicate that the SHARP method provides better recognition accuracy in highly reverberant environments compared to the SSF algorithm. It is also shown that the performance of the SHARP method can be further improved by combining it with feature-space maximum likelihood linear regression (fMLLR).
Byung Joon Cho, Haeyong Kwon, Ji-Won Cho, Chanwoo Kim 0001, Richard M. Stern, Hyung-Min Park
IEEE Signal Process. Lett.6
2016 DNN-Based Feature Enhancement Using DOA-Constrained ICA for Robust Speech Recognition
abstract
The performance of automatic speech recognition (ASR) system is often degraded in adverse real-world environments. In recent times, deep learning has successfully emerged as a breakthrough for acoustic modeling in ASR; accordingly, deep neural network (DNN)-based speech feature enhancement (FE) approaches have attracted much attention owing to their powerful modeling capabilities. However, DNN-based approaches are unable to achieve remarkable performance improvements for speech with severe distortion in the test environments different from training environments. In this letter, we propose a DNN-based FE method where the DNN inputs include preenhanced spectral features computed from multichannel input signals to reconstruct noise-robust features. The preenhanced spectral features are obtained by direction-of-arrival (DOA)-constrained independent component analysis (DCICA) followed by Bayesian FE using a hidden-Markov-model prior, to exploit the capabilities of efficient online target speech extraction and efficient FE with prior information for robust ASR. In addition, noise spectral features computed from DCICA are included for further improvement. Therefore, the DNN is trained to reconstruct a clean spectral feature vector, from a sequence of corrupted input feature vectors in addition to the corresponding preenhanced and noise feature vectors. Experimental results demonstrate that the proposed method significantly improves recognition performance, even in mismatched noise conditions.
Ho Yong Lee, Ji-Won Cho, Minook Kim, Hyung-Min Park
IEEE Signal Process. Lett.4
2015 A Method for Speech Dereverberation Based on an Image Deblurring Algorithm Using the Prior of Speech Magnitude Gradient Distribution in the Time-Frequency Domain
abstract
We propose a speech dereverberation method in the time-frequency domain, based on an image deblurring algorithm. A reverberant speech magnitude can be modeled as a convolution of a clean speech with a reverberation filter in time-frequency domain. Then, dereverberation problem can be regarded as that of image deblurring. Therefore, the proposed method estimates the clean speech magnitude in the time-frequency domain by using the fast image deconvolution method with priors on sparsity of the clean speech magnitude gradient and exponentially decaying property of reverberation filters along the time axis. Then, scaling the reverberation speech magnitude by a mask obtained from the estimated clean one performs dereverberation. Experimental results show that the described method can enhance speech.
Wonyong Jo, Ji-Won Cho, Changsoo Je, Hyung-Min Park
HAI4
2015 Efficient online target speech extraction using DOA-constrained independent component analysis of stereo data for robust speech recognition
Minook Kim, Hyung-Min Park
Signal Process.2
2015 Homographic p-norms: Metrics of homographic image transformation
Changsoo Je, Hyung-Min Park
Signal Process. Image Commun.2
2014 Robust speech recognition in reverberant environments using subband-based steady-state monaural and binaural suppression
Hyung-Min Park, Matthew Maciejewski, Chanwoo Kim 0001, Richard M. Stern
INTERSPEECH1
2014 Design and implementation of an augmented reality system using gaze interaction
Jae-Young Lee 0010, Hyung-Min Park, Seok-Han Lee, Soon-Ho Shin, Tae-eun Kim, Jong-Soo Choi
Multim. Tools Appl.2
2013 Single-Channel Speech Dereverberation Based on Non-negative Blind Deconvolution and Prior Imposition on Speech and Filter
Il-Young Jeong, Biho Kim, Hyung-Min Park
ICONIP (3)3
2013 Integrated Multi-scale Retinex Using Fuzzy Connectivity Based on CIELAB Color Space for Preserving Color
Biho Kim, Wonyong Jo, Hyung-Min Park
ICONIP (3)3
2013 Image-pair-based deblurring with spatially varying norms and noisy image updating
Chang-Hwan Son, Hyunseung Choo, Hyung-Min Park
J. Vis. Commun. Image Represent.3
2013 Robust speech recognition based on independent vector analysis using harmonic frequency dependency
Soram Jun, Minook Kim, Myungwoo Oh, Hyung-Min Park
Neural Comput. Appl.4
2013 Disparity-based space-variant image deblurring
Changsoo Je, Hyeon Sang Jeon, Chang-Hwan Son, Hyung-Min Park
Signal Process. Image Commun.4
2013 Optimized hierarchical block matching for fast and accurate image registration
Changsoo Je, Hyung-Min Park
Signal Process. Image Commun.2
2013 An Efficient HMM-Based Feature Enhancement Method With Filter Estimation for Reverberant Speech Recognition
abstract
This letter presents an efficient feature enhancement method for reverberant speech recognition that derives a minimum mean square error estimate of clean logarithmic mel-frequency power spectral coefficients (LMPSCs) based on a hidden-Markov-model(HMM) prior. Although an observation model of the reverberant LMPSCs can be simply formulated by coarse modeling of the room impulse response (RIR) , the presented method estimates not only the clean LMPSCs but also the RIR to reflect detailed reverberation. The experimental results indicate that the described method can further reduce relative word error rate (WER) by 18.09% on average compared to a method based on RIR coarse modeling.
Ji-Won Cho, Hyung-Min Park
IEEE Signal Process. Lett.2
2012 Efficient face recognition based on MCT and I(2D)2PCA
abstract
This paper presents a robust algorithm to recognize human faces efficiently. Although the principle component analysis (PCA) is one of the most popular feature extraction methods, it requires too much computational load and memory capacity to implement a real-time embedded system for face recognition. To overcome the drawback, we employ the incremental two-directional two-dimensional PCA (I(2D)2PCA) which combines (2D)2PCA to demand much less computational complexity than the conventional PCA and the incremental PCA (IPCA) to adapt the eigenspace only using a new incoming sample datum without memorizing all of the previous trained data. In addition, robustness to illumination variations is addressed by introducing the modified census transform (MCT) which is a local normalization method useful for real-world application and implementation in an embedded system. Experimental results on the Yale Face Database B demonstrate that the proposed method based on the I(2D)2PCA with MCT preprocessing provided efficient and robust face recognition.
Biho Kim, Hyung-Min Park
SMC2
2011 Development of Visualizing Earphone and Hearing Glasses for Human Augmented Cognition
Byunghun Hwang, Cheol-Su Kim, Hyung-Min Park, Yun-Jung Lee, Min Young Kim 0003, Minho Lee 0001
ICONIP (2)3
2011 Preprocessing of Independent Vector Analysis Using Feed-Forward Network for Robust Speech Recognition
Myungwoo Oh, Hyung-Min Park
ICONIP (2)2
2011 Blind source separation based on independent vector analysis using feed-forward network
Myungwoo Oh, Hyung-Min Park
Neurocomputing2
2010 Human Augmented Cognition Based on Integration of Visual and Auditory Information
Woong-Jae Won, Wono Lee, Sang-Woo Ban, Minook Kim, Hyung-Min Park, Minho Lee 0001
PRICAI5
2010 Multiple Reverberant Sound Localization Based on Rigorous Zero-Crossing-Based ITD Selection
abstract
This letter presents a multiple sound localization method for non-stationary speech sources in reverberant environments. Although zero-crossing-based interaural time differences (ITDs) were robust to diffuse noise and energy-based onset detection tackled reverberant speech , the onset detection method was sensitive to parameters. The proposed method employs echo-free onset detection , and this letter adds reverberation time estimation to obtain accurate echoes and signal-to-noise ratio estimation to assist the onset detection based on typical properties of acoustic reverberation in selecting ITDs corresponding to source locations. Experimental results indicate the effectiveness of the proposed method in reverberant environments.
Soo-Yeon Lee, Hyung-Min Park
IEEE Signal Process. Lett.2
2009 A Bark-scale filter bank approach to independent component analysis for acoustic mixtures
Hyung-Min Park, Sang-Hoon Oh, Soo-Young Lee
Neurocomputing1
2009 Spatial separation of speech signals using amplitude estimation based on interaural comparisons of zero-crossings
Hyung-Min Park, Richard M. Stern
Speech Commun.1
2009 Imposition of Sparse Priors in Adaptive Time Delay Estimation for Speaker Localization in Reverberant Environments
abstract
In this letter, we describe a method to estimate the time delay for speaker localization in reverberant environments. Based on an adaptive eigenvalue decomposition (AED) algorithm, the method takes the reverberation fully into account by estimating channel impulse responses from a speaker to sensors directly, but it may suffer from whitening effects for temporally correlated natural sounds. Imposing sparse priors on the responses can reduce the temporal whitening and provide a more accurate and robust time delay of speaker location. Experiments demonstrate that the proposed method can efficiently estimate the time delay for speaker localization in reverberant environments.
Ji-Won Cho, Hyung-Min Park
IEEE Signal Process. Lett.2
2008 Wearable augmented reality system using gaze interaction
abstract
Undisturbed interaction is essential to provide immersive AR environments. There have been a lot of approaches to interact with VEs (virtual environments) so far, especially in hand metaphor. When the user’s hands are being used for hand-based work such as maintenance and repair, necessity of alternative interaction technique has arisen. In recent research, hands-free gaze information is adopted to AR to perform original actions in concurrence with interaction. [3, 4]. There has been little progress on that research, still at a pilot study in a laboratory setting. In this paper, we introduce such a simple WARS(wearable augmented reality system) equipped with an HMD, scene camera, eye tracker. We propose ‘Aging’ technique improving traditional dwell-time selection, demonstrate AR gallery — dynamic exhibition space with wearable system.
Hyung-Min Park, Seok-Han Lee, Jong-Soo Choi
ISMAR1
2007 Missing Feature Speech Recognition using Dereverberation and Echo Suppression in Reverberant Environments
abstract
This paper describes an algorithm that efficiently segregates desired speech features from spatially-separated interfering sources in reverberant environments. Although most binaural segregation techniques successfully remove interference components in the absence of reverberation, source segregation in reverberant environments remains a challenging problem. In order to reduce the effects of reverberation, we present a method that dereverberates input signals before they are segregated. The dereverberation filter is estimated from the autocorrelation of the observations and primarily deals with early reflections, while late reflections are effectively suppressed by an inhibitory mechanism that estimates their relative contribution in each time-frequency segment. Information about the salience of the target in a given time-frequency segment based on source separation is combined with the corresponding information based on reverberation suppression through the use of a continually-variable weighting function or mask. Use of the novel reverberation processing results in a relative decrease in WER of 11.5% to 20.9% and use of the combined approaches reduces relative WER by as much as 65.3%.
Hyung-Min Park, Richard M. Stern
ICASSP (4)1
2007 Directionally Constrained Filterbank ICA
abstract
A modification is proposed to the independent component analysis (ICA)-based filterbank approach in consideration to its structural similarity with binaural auditory model of sound source localization. The estimated sound locations provide an additional cue to the learning algorithm, which is utilized for initialization and imposition of directional constraints on the subband separation networks. Directionally constrained filterbank ICA (DC-FBICA) gives faster convergence and improves separation performance for noisy mixtures having significant spectral overlap among the convolved mixture and the corrupting noise. However, only slight improvement in separation performance is observed when the additive noise is a low frequency noise, although faster convergence is still observed.
Chandra Shekhar Dhir, Hyung-Min Park, Soo-Young Lee
IEEE Signal Process. Lett.2
2006 Spatial Separation of Speech Signals Using Continuously-Variable Masks Estimated From Comparisons of Zero Crossings
abstract
This paper describes an algorithm that achieves noise robustness in speech recognition by reconstructing the desired signal from a mixture of two signals using continuously-variable masks. In contrast to current methods which use binary masks, this approach estimates the relative contribution of the desired source in a mixture of sources and reconstructs the desired signal in proportion to its estimated contribution to each time-frequency segment. Estimation of the continuously-variable masks is based on the relationship between the relative intensity of each source and the interaural time difference (ITD). Estimation of the ITD is accomplished using zero-crossing-based methods. It is shown that the use of zero-crossing approaches to estimate ITDs and continuously-variable masks provide better speech recognition accuracy than cross-correlation-based approaches to ITD estimation and binary masks.
Hyung-Min Park, Richard M. Stern
ICASSP (4)1
2006 Performance Evaluation of Directionally Constrained Filterbank ICA on Blind Source Separation of Noisy Observations
Chandra Shekhar Dhir, Hyung-Min Park, Soo-Young Lee
ICONIP (1)2
2006 A filter bank approach to independent component analysis for convolved mixtures
Hyung-Min Park, Chandra Shekhar Dhir, Sang-Hoon Oh, Soo-Young Lee
Neurocomputing1
2006 A modified infomax algorithm for blind signal separation
Hyung-Min Park, Sang-Hoon Oh, Soo-Young Lee
Neurocomputing1
2005 Sound segregation based on binaural zero-crossings
Young-Ik Kim, Sung Jun An, Rhee Man Kil, Hyung-Min Park
INTERSPEECH4
2004 Permutation Correction of Filter Bank ICA Using Static Channel Characteristics
Chandra Shekhar Dhir, Hyung-Min Park, Soo-Young Lee
ICONIP2
2003 A uniform oversampled filter bank approach to independent component analysis
abstract
We present a new approach to perform independent component analysis (ICA) for convolved mixtures. This approach is based on filter banks, and a simplified network efficiently performs ICA with decimated signals in each subband. Decimation provides much less computational complexity and faster convergence speed than the time domain approach. Furthermore, the approach does not have a performance limitation of the frequency domain approach, and it is able to select the number of filters in the filter bank regardless of reverberation. With an oversampled filter bank, adaptive parameters can be adjusted without any information of other subbands, and the approach is suitable for parallel processing. We verify the effectiveness of the filter bank approach through simulations on adaptive noise cancelling.
Hyung-Min Park, Sang-Hoon Oh, Soo-Young Lee
ICASSP (5)1
2003 A filter bank approach to independent component analysis and its application to adaptive noise cancelling
Hyung-Min Park, Sang-Hoon Oh, Soo-Young Lee
Neurocomputing1
2003 FPGA implementation of ICA algorithm for blind signal separation and adaptive noise canceling
abstract
An field programmable gate array (FPGA) implementation of independent component analysis (ICA) algorithm is reported for blind signal separation (BSS) and adaptive noise canceling (ANC) in real time. In order to provide enormous computing power for ICA-based algorithms with multipath reverberation, a special digital processor is designed and implemented in FPGA. The chip design fully utilizes modular concept and several chips may be put together for complex applications with a large number of noise sources. Experimental results with a fabricated test board are reported for ANC only, BSS only, and simultaneous ANC/BSS, which demonstrates successful speech enhancement in real environments in real time.
Chang-Min Kim, Hyung-Min Park, Taesu Kim, Yoon-Kyung Choi, Soo-Young Lee
IEEE Trans. Neural Networks2
2002 Top-down attention to complement independent component analysis for blind signal separa
Un-Min Bae, Hyung-Min Park, Soo-Young Lee
Neurocomputing2