EDBT 2026 Demo / reviewers in the wild / expert
Jong Won Shin
dblp:14/6965
· DBLP profile ↗
58ranked-venue papers
11as first author
23since 2021 · last 2026
0000-0002-8910-0264ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 9 first-author · 20 since 2021Artificial intelligence and machine learning · 25 · 5 first-author · 10 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Target Speaker Extraction Using Multi-Stage Cross-Attention and Frequency-Wise State InitializationabstractSeveral recent target speaker extraction (TSE) models directly utilize enrollment speech without explicitly extracting low-dimensional speaker embeddings. However, these methods typically inject the speaker information only once at the input of the speaker extraction network, which may be insufficient because the conditioning information can become diluted as it propagates through repeated separator blocks. In this letter, we propose a TSE model built upon the TF-GridNet, which is a speech separation model performing dual-path modeling in the time-frequency domain with cross-frame self-attention modules. In the proposed TSE model, the self-attention modules in the first$M$separator blocks are replaced by cross-attention between the enrollment speech and the mixture signal, providing speaker information in multiple stages without introducing additional parameters or computation compared with the original TF-GridNet blocks. In addition, the initial hidden and cell states of the inter-frame long short-term memory (LSTM) modules are determined for each frequency from the enrollment speech. As the pattern of the temporal correlation may be different for each frequency depending on the pitch and speaking style, speaker-dependent frequency-wise state initialization would be helpful. Experimental results showed that the proposed TSE model demonstrated the best PESQ scores and comparable SI-SDRs with lower computational complexity. Hyeonseung Kim, Jong Won Shin |
IEEE Signal Process. Lett. | 2 |
| 2026 | R3VQ: Redundancy-Reduced Residual Vector Quantization for Low-Bitrate Neural Speech CodingabstractNeural speech and audio codecs have demonstrated decent quality of the decoded audio at low bitrates. They consist of three parts, an encoder, a decoder, and a quantizer. Residual vector quantization (RVQ) or multi-stage vector quantization in which the residual signal from the previous stage is quantized in the next stage is employed in many neural speech codecs and has exhibited good performance while providing bitrate scalability. In this letter, we propose the redundancy-reduced residual vector quantization (R3VQ) which improves the RVQ by inserting a neural network called a refiner. The role of the refiner is to reduce the power of the residual signal to be quantized by enhancing the estimate of the original speech from the quantized signals in the previous stages. We also present a part-wise (PW) training scheme suitable for the training of the neural speech codec with the R3VQ. Experimental results showed that the proposed R3VQ trained with a PW training scheme outperformed the RVQ in both objective measures for speech quality and subjective MUltiple Stimuli with Hidden Reference and Anchor (MUSHRA) test. Eunkyun Lee, Jongwook Chae, Sooyoung Park, Jong Won Shin |
IEEE Signal Process. Lett. | 4 |
| 2025 | FlowSE: Flow Matching-based Speech EnhancementabstractDiffusion probabilistic models have shown impressive performance for speech enhancement, but they typically require 25 to 60 function evaluations in the inference phase, resulting in heavy computational complexity. Recently, a fine-tuning method was proposed to correct the reverse process, which significantly lowered the number of function evaluations (NFE). Flow matching is a method to train continuous normalizing flows which model probability paths from known distributions to unknown distributions including those described by diffusion processes. In this paper, we propose a speech enhancement based on conditional flow matching. The proposed method achieved the performance comparable to those for the diffusion-based speech enhancement with the NFE of 60 when the NFE was 5, and showed similar performance with the diffusion model correcting the reverse process at the same NFE from 1 to 5 without additional fine tuning procedure. We also have shown that the corresponding diffusion model derived from the conditional probability path with a modified optimal transport conditional vector field demonstrated similar performances with the NFE of 5 without any fine-tuning procedure. Seonggyu Lee, Sein Cheong, Sangwook Han, Jong Won Shin |
ICASSP | 4 |
| 2025 | Speech Enhancement based on cascaded two flows
Seonggyu Lee, Sein Cheong, Sangwook Han, Kihyuk Kim, Jong Won Shin |
INTERSPEECH | 5 |
| 2025 | Text-to-Speech With Lip Synchronization Based on Speech-Assisted Text-to-Video Alignment and Masked Unit PredictionabstractText-to-speech (TTS) with lip synchronization (TTSLS) is the task of generating a speech signal synchronized with the lip movements in a video given the text transcription and the video without speech. Previous approaches to TTSLS aligned the phoneme sequence and video frames using scaled dot-product attention with a diagonal constraint loss, which was employed to prevent a phoneme from being assigned to video frames too far away. However, the diagonal constraint loss basically assumes that the duration of each phoneme is about the same, which is not always valid as speaking styles can be different. In this letter, we propose a TTSLS system based on speech-assisted text-to-video alignment and masked unit prediction. By utilizing the ground-truth speech signal available in the training phase, we construct a loss function for text-to-video alignment using the text-to-speech alignment obtained by a pre-trained TTS model. To deal with video frames without frontal lip images, we employ a masked unit prediction loss so that the unit predictor in the proposed system can estimate the masked units from the rest of the units. In addition, we modified the probability distribution for the unit predictor using a learnable null embedding for video inspired by classifier-free guidance. Experimental results demonstrated that our proposed method outperformed previous TTSLS systems in both lip-speech synchronization and speech recognition performance. Youngdo Ahn, Jongwook Chae, Jong Won Shin |
IEEE Signal Process. Lett. | 3 |
| 2025 | Integrated DNN-Based Parameter Estimation for Multichannel Speech Enhancement
Sein Cheong, Minseung Kim, Jong Won Shin |
IEEE Signal Process. Lett. | 3 |
| 2023 | Individual Sub-Band Estimation Approach to Bandwidth Extension and Enhancement of Coded SpeechabstractThe streaming Sound EnhAncement Network (SEANet) has demonstrated impressive performance for speech bandwidth extension (BWE) with low latency and computational complexity. Although the streaming SEANet was designed for voice communication systems, it was not tested with decoded signals that included coding artifacts. Our preliminary experiment showed that the output of the streaming SEANet for the decoded speech had room for improvement even if it was trained with decoded speeches, possibly because it should perform the BWE and the coded speech enhancement (CSE) at once. In this work, we propose to utilize two streaming SEANets in parallel, which are dedicated to the narrowband CSE and the generation of the upper band speech signal, respectively. Experimental results showed that the proposed model outperformed a bigger streaming SEANet trained to carry out both tasks in terms of the PESQ scores and the MUSHRA test. Youngwon Choi, Eunkyun Lee, Inseon Jang, Jong Won Shin |
ICASSP | 4 |
| 2023 | Short-Segment Speaker Verification Using ECAPA-TDNN with Multi-Resolution EncoderabstractTime-domain approaches have shown the potential to improve the performance of speaker verification, but still predominant approaches utilize hand-crafted features such as the mel filterbank energies. Although these features are based on speech perception models and exhibited impressive performances, the fixed frame size does not allow good temporal and spectral resolutions at the same time and there is information loss when taking the magnitude spectrum and during frequency rescaling. In this paper, we propose to incorporate multi-resolution time-domain information into the ECAPA-TDNN speaker verification system. We construct a multi-resolution encoder to extract multiple features in different temporal resolutions, and let the extracted features drive the adapter modules. Experimental results showed that the proposed method outperformed other recently proposed approaches when the input length was 2 seconds or shorter for the VoxCeleb dataset. The proposed approach also showed superior performance on the Google Speech Commands dataset v2. Sangwook Han, Youngdo Ahn, Kyeongmuk Kang, Jong Won Shin |
ICASSP | 4 |
| 2023 | GRAVO: Learning to Generate Relevant Audio from Visual Features with Noisy Online Videos
Youngdo Ahn, Chengyi Wang 0002, Yu Wu 0012, Jong Won Shin, Shujie Liu 0001 |
INTERSPEECH | 4 |
| 2023 | DNN-based Parameter Estimation for MVDR Beamforming and Post-filtering
Minseung Kim, Sein Cheong, Jong Won Shin |
INTERSPEECH | 3 |
| 2023 | On Training Speech Separation Models With Various Numbers of SpeakersabstractMany monaural speech separation models assume that the exact number of speakers is known in advance, which is not applicable to many real-world scenarios. To deal with an unknown number of speakers, previous approaches either iteratively separate one speech at a time, or employ a more relaxed assumption that the maximum number of speakers is known a priori and set the number of outputs accordingly. When the number of speakers in the mixture is smaller than the number of outputs in the latter case, the extra outputs that are not mapped onto signals in the input mixture are trained to produce predefined target signals such as the silence or the input mixture. In this letter, we propose to ignore the extra outputs in training instead of evaluating the cost with a certain target for separation models with a fixed number of output channels. We also introduce a method to select valid output signals. Experimental results showed that assigning any type of predefined targets degraded separation performance compared with ignoring the extra outputs. Hyeonseung Kim, Jong Won Shin |
IEEE Signal Process. Lett. | 2 |
| 2022 | Multi-Corpus Speech Emotion Recognition for Unseen Corpus Using Corpus-Wise Weights in Classification Loss
Youngdo Ahn, Sung Joo Lee, Jong Won Shin |
INTERSPEECH | 3 |
| 2022 | iDeepMMSE: An improved deep learning approach to MMSE speech and noise power spectrum estimation for speech enhancement
Minseung Kim, Hyungchan Song, Sein Cheong, Jong Won Shin |
INTERSPEECH | 4 |
| 2022 | Exploring WavLM on Speech EnhancementabstractThere is a surge in interest in self-supervised learning approaches for end-to-end speech encoding in recent years as they have achieved great success. Especially, WavLM showed state-of-the-art performance on various speech processing tasks. To better understand the efficacy of self-supervised learning models for speech enhancement, in this work, we design and conduct a series of experiments with three resource conditions by combining WavLM and two high-quality speech enhancement systems. Also, We propose a regression-based WavLM training objective and a noise-mixing data configuration to further boost the downstream enhancement performance. The experiments on the DNS challenge dataset and a simulation dataset show that the WavLM benefits the speech enhancement task in terms of both speech quality and speech recognition accuracy, especially for low fine-tuning resources. For the high fine-tuning resource condition, only the word error rate is substantially improved. Hyungchan Song, Sanyuan Chen, Zhuo Chen 0006, Yu Wu 0012, Takuya Yoshioka, Jong Won Shin, Shujie Liu 0001 |
SLT | 7 |
| 2022 | Alias-and-Separate: Wideband Speech Coding Using Sub-Nyquist Sampling and Speech SeparationabstractDecimation of a discrete-time signal below the Nyquist rate without applying an appropriate lowpass filter results in a distortion called aliasing. If wideband speech sampled at 16 kHz is decimated by 2 to result in a signal sampled at 8 kHz with aliasing, the decimated signal would be the summation of two speech-like signals, which are the narrowband speech covering 0-4 kHz and the spectrally flipped aliasing component coming from 8-4 kHz. Recently, the performance of speech separation has been remarkably improved with deep learning-based approaches, implying that the narrowband and aliasing components may be able to be separated. In this letter, we propose a novel method for low-rate wideband speech coding utilizing a standard narrowband codec. Instead of coding wideband speech using a wideband codec with a limited bitrate, we propose to decimate the input wideband speech incurring aliasing, and then encode it with a narrowband codec by allocating all the allowed bitrate to 0-4 kHz. After decoding the encoded bitstream, we apply a speech separation technique to obtain the narrowband and aliasing signals, which are then used to reconstruct the wideband speech by expansion, low/highpass filtering, and summation. Experimental results showed that the proposed method could achieve subjective quality comparable to the speeches coded by wideband codecs at higher bitrates in a subjective MUSHRA test. Soojoong Hwang, Eunkyun Lee, Inseon Jang, Jong Won Shin |
IEEE Signal Process. Lett. | 4 |
| 2022 | Factorized MVDR Deep Beamforming for Multi-Channel Speech EnhancementabstractTraditionally, adaptive beamformers such as the minimum-variance distortionless response (MVDR) beamformer and generalized eigenvalue beamformer have been widely used for multi-channel speech enhancement with a single-channel postfilter. Recently, several approaches have been proposed to enhance the signals used to estimate speech and noise spatial covariance matrices (SCMs) and process the outputs of the beamformers using deep neural networks (DNNs). However, the preprocessing of the signals for SCMs estimation may disrupt phase relations among input signals and the time-averages used to estimate speech and noise SCMs may not be optimal for beamformer performance even though the estimated signals are close to the ground truth. In this letter, we propose a deep beamforming approach which estimates factors of the MVDR beamformer using a DNN to circumvent the difficulty of the speech and noise SCM estimation. We formulate the MVDR beamformer as a factorized form related to two complex factors and estimate them using a DNN with a cost function comparing beamformed signal and the original clean speech. Experimental results showed that the proposed factorized MVDR beamformer could mimic the characteristics of the MVDR beamformer with true relative transfer function and noise SCM and outperformed the MVDR beamformer with deep learning-based pre- and postprocessing in terms of the perceptual evaluation of speech quality scores. Kyeongmuk Kang, Jong Won Shin |
IEEE Signal Process. Lett. | 3 |
| 2022 | Dual Microphone Speech Enhancement Based on Statistical Modeling of Interchannel Phase DifferenceabstractThe interchannel phase difference (IPD) may be one of the most widely-used spatial cues in multichannel speech processing, and has been used in beamformers and post filters for speech enhancement. The coherence, which is also used as a feature for speech enhancement, can provide information on the reliability of the IPD for the estimation of the speech presence probability (SPP). In this paper, we propose dual microphone speech enhancement adoptinga posterioriSPP estimation based on statistical modeling of the IPD. The marginal distribution of the IPD is derived from the distribution of the relative transfer function which is parameterized with the IPD and coherence, with a single assumption that the observed discrete Fourier transform (DFT) coefficients in each frequency are distributed according to a complex bivariate Gaussian distribution. Given the direction of arrival of the desired signal, thea posterioriSPP is obtained using the IPD distributions with and without the information on the location of the interfering source, and is applied to speech enhancement. Experimental results for various types and locations of noise, signal-to-noise ratios, reverberation times, and locations of the target source showed that the proposed method outperformed previously proposed approaches utilizing IPD information. Soojoong Hwang, Minseung Kim, Jong Won Shin |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Improved Speech Enhancement Considering Speech PSD UncertaintyabstractSpeech enhancement based on statistical models has been studied for several decades. Recently, the speech enhancement adopting a speech power spectral density (PSD) uncertainty model has been proposed. This approach distinguishes the true speech PSD from its estimate and considers both as random variables. It incorporates a prior distribution of speech spectra and speech PSD estimators to derive the PSD uncertainty-aware counterpart to conventional clean speech estimators, which results in performance improvement. However, the speech PSD uncertainty model has not yet been adopted for parameter estimations such asspeech presence probability, noise PSD, and speech power spectra estimations in the speech enhancement framework. In this paper, we incorporate the speech PSD uncertainty model to all the components of the statistical model-based speech enhancement framework by deriving PSD uncertainty-aware counterparts to conventional parameter estimators. Specifically, we derive thespeech presence probability (SPP) where the likelihood function for each hypothesis is based on the speech PSD uncertainty. With thisSPP, a novel SPP-based noise PSD estimator is derived. Also, we derive the minimum mean-square error (MMSE) estimator for the power spectrum of the clean speech in the current frame under speech PSD uncertainty which is exploited to refine the speech PSD estimator. Finally, the refined speech PSD estimator is incorporated into the spectral gain function based on the speech PSD uncertainty model. The proposed approach showed improved noise PSD estimation performance in terms of the averaged logarithmic error distance, and improved speech enhancement performance in terms of the noise reduction, segmental signal-to-noise ratio, perceptual evaluation of speech quality (PESQ) scores and short-time objective intelligibility in our experiments. It also exhibited comparable performance with a real-time deep learning-based speech enhancement system in terms of the PESQ scores and composite measures for the VoiceBank-DEMAND dataset. Minseung Kim, Jong Won Shin |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2021 | Time-Domain Speaker Verification Using Temporal Convolutional NetworksabstractRecently, speaker verification systems using deep neural networks have been widely studied. Many of them utilize hand-crafted features such as mel-filterbank energies, mel-frequency cepstral coefficients, and magnitude spectrograms, which are not designed specifically for the speaker verification task and may not be optimal. Recent releases of the large datasets such as VoxCeleb enable us to extract the task-specific features in a data-driven way. In this paper, we propose a speaker verification system that takes the time-domain raw waveforms as inputs, which adopts a learnable encoder and temporal convolutional networks (TCNs) that have shown impressive performance in speech separation. Moreover, we have applied the squeeze and excitation networks after each TCN block to apply channel-wise attention. Our experiments on the VoxCeleb1 dataset demonstrate that the speaker verification system utilizing the proposed feature extraction model outperforms previously proposed time-domain speaker verification systems. Sangwook Han, Jaeuk Byun, Jong Won Shin |
ICASSP | 3 |
| 2021 | Coded Speech Enhancement Using Neural Network-Based Vector-Quantized Residual Features
Youngju Cheon, Soojoong Hwang, Sangwook Han, Inseon Jang, Jong Won Shin |
Interspeech | 5 |
| 2021 | Multiple Sound Source Localization Based on Interchannel Phase Differences in All Frequencies with Spectral Masks
Hyungchan Song, Jong Won Shin |
Interspeech | 2 |
| 2021 | Cross-Corpus Speech Emotion Recognition Based on Few-Shot Learning and Domain AdaptationabstractWithin a single speech emotion corpus, deep neural networks have shown decent performance in speech emotion recognition. However, the performance of the emotion recognition based on data-driven learning methods degrades significantly for the cross-corpus scenario. To relieve this issue without any labeled samples from the target domain, we propose a cross-corpus speech emotion recognition based on few-shot learning and unsupervised domain adaptation, which is trained to learn the class (emotion) similarity from the source domain samples adapted to the target domain. In addition, we utilize multiple corpora in training to enhance the robustness of the emotion recognition to the unseen samples. Experiments on emotional speech corpora with three different languages showed that the proposed method outperformed other approaches. Youngdo Ahn, Sung Joo Lee, Jong Won Shin |
IEEE Signal Process. Lett. | 3 |
| 2021 | Monaural Speech Separation Using Speaker Embedding From Preliminary SeparationabstractIn speech separation, the identities of the speakers may be an important cue to discriminate speeches in the mixture and separate them better. A few recent researches used the speaker embedding as an additional information, but they often require prior information about the target speaker or used noisy speaker embedding extracted from the mixture signal. In this article, we propose monaural speech separation that utilizes the speaker embedding in the later separator blocks, which is extracted from the intermediate separated results obtained by the early stages of the separator network. The later blocks in the separator networks consisting of repeated blocks such as the fully-convolutional time-domain audio separation network (Conv-TasNet) or the successive downsampling and resampling of multi-resolution features (SuDoRM-RF) are modified to take the speaker information as a form of affine transformation or addition to the original input tensor. The experimental results showed that the proposed methods significantly improved the performances of existing separation systems with a moderate number of additional parameters. Jaeuk Byun, Jong Won Shin |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | DNN-based Emotion Recognition Based on Bottleneck Acoustic Features and Lexical FeaturesabstractIn this paper, we propose a novel emotion recognition method to reflect affect salient information using acoustic and lexical features. The acoustic features are extracted from the speech signal by applying statistical functionals of emotionally high-level features derived from Deep Neural Network (DNN). These acoustic features are early fused with two types of lexical features extracted from the text transcription of the speech signal, which are the distributed representation and affective lexicon-based dimensions. The fused features are fed to another DNN for utterance-level emotion classification. Experimental results on the Interactive Emotional Dyadic Motion Capture (IEMOCAP) multimodal dataset showed 75.5% in unweighted accuracy recall, which outperformed the best results reported previously in the multimodal emotion recognition using acoustic and lexical features. Eesung Kim, Jong Won Shin |
ICASSP | 2 |
| 2019 | Sound Localization Based on Phase Difference Enhancement Using Deep Neural NetworksabstractThe performance of most of the classical sound source localization algorithms degrades seriously in the presence of background noise or reverberation. Recently, deep neural networks (DNNs) have successfully been applied to sound source localization, which mainly aim to classify the direction-of-arrival (DoA) into one of the candidate sectors. In this paper, we propose a DNN-based phase difference enhancement for DoA estimation, which turned out to be better than the direct estimation of the DoAs from the input interchannel phase differences (IPDs). The sinusoidal functions of the phase differences for “clean and dry” source signals are estimated from the sinusoidal functions of the IPDs for the input signals, which may include directional signals, diffuse noise, and reverberation. The resulted DoA is further refined to compensate for the estimation bias near the end-fire directions. From the enhanced IPDs, we can determine the DoA for each frequency bin and the DoAs for the current frame from the distributions of the DoAs for frequencies. Experimental results with various types and levels of background noise, reverberation times, numbers of sources, room impulse responses, and DoAs showed that the proposed method outperformed conventional approaches. Junhyeong Pak, Jong Won Shin |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Multichannel speech reinforcement based on binaural unmasking
Junhyeong Pak, Inyong Choi, Yu Gwang Jin, Jong Won Shin |
Signal Process. | 4 |
| 2016 | NMF-based source separation utilizing prior knowledge on encoding vectorabstractNon-negative matrix factorization (NMF) is an unsupervised technique to represents a nonnegative data matrix with a product of nonnegative basis and encoding matrices. The encoding matrix for the training phase contains information on the pattern of how each basis vector is utilized. The histogram for each row of this matrix corresponding to a specific basis turned out to be sparse, while the level of sparsity varied significantly in each basis. In this paper, the distribution of each component of an encoding vector is modeled as an independent exponential or gamma distribution, and a new objective function with the log-likelihood of the current encoding vector is proposed. Experimental results on audio source separation demonstrate that the utilization of the prior knowledge on the encoding matrix based on sparse statistical models can enhance the source separation performance. Kisoo Kwon, Jong Won Shin, Nam Soo Kim |
ICASSP | 2 |
| 2016 | Dual Microphone Voice Activity Detection Exploiting Interchannel Time and Level DifferencesabstractThe two most important spatial cues in human auditory system may be the interaural time difference and the interaural level difference. There have been many attempts to utilize the time difference of arrival (TDoA) and level difference between two microphone signals for voice activity detection (VAD). In this letter, we propose a dual microphone VAD algorithm based on a support vector machine for which the input vector consists of both TDoA-based and level difference-based features. Several candidates for the feature combination have been compared using various TDoA-related and level difference-related features. Experimental results showed that the proposed VAD algorithm outperformed a standardized single microphone VAD, VADs based on the TDoA or level difference, and logical combination of them in various noise environments. Jaehoon Park, Yu Gwang Jin, Soojoong Hwang, Jong Won Shin |
IEEE Signal Process. Lett. | 4 |
| 2015 | Discriminative nonnegative matrix factorization using cross-reconstruction error for source separation
Kisoo Kwon, Jong Won Shin, Hyung Yong Kim, Nam Soo Kim |
INTERSPEECH | 2 |
| 2015 | DNN-based residual echo suppression
Chul Min Lee, Jong Won Shin, Nam Soo Kim |
INTERSPEECH | 2 |
| 2015 | NMF-based Target Source Separation Using Deep Neural NetworkabstractNon-negative matrix factorization (NMF) is one of the most well-known techniques that are applied to separate a desired source from mixture data. In the NMF framework, a collection of data is factorized into a basis matrix and an encoding matrix. The basis matrix for mixture data is usually constructed by augmenting the basis matrices for independent sources. However, target source separation with the concatenated basis matrix turns out to be problematic if there exists some overlap between the subspaces that the bases for the individual sources span. In this letter, we propose a novel approach to improve encoding vector estimation for target signal extraction. Estimating encoding vectors from the mixture data is viewed as a regression problem and a deep neural network (DNN) is used to learn the mapping between the mixture data and the corresponding encoding vectors. To demonstrate the performance of the proposed algorithm, experiments were conducted in the speech enhancement task. The experimental results show that the proposed algorithm outperforms the conventional encoding vector estimation scheme. Tae Gyoon Kang, Kisoo Kwon, Jong Won Shin, Nam Soo Kim |
IEEE Signal Process. Lett. | 3 |
| 2015 | NMF-Based Speech Enhancement Using Bases UpdateabstractThis letter presents a speech enhancement technique combining statistical models and non-negative matrix factorization (NMF) with on-line update of speech and noise bases. The statistical model-based enhancement methods have been known to be less effective to non-stationary noises while the template-based enhancement techniques can deal with them quite well. However, the template-based enhancement techniques usually rely on a priori information. To overcome the shortcomings of both approaches, we propose a novel speech enhancement method that combines the statistical model-based enhancement scheme with the NMF-based gain function. For a better performance in time-varying noise environments, both the speech and noise bases of NMF are adapted simultaneously with the help of the estimated speech presence probability. Experimental results showed that the proposed method outperformed not only the statistical model-based and NMF approaches, but also their combination in various noise environments. Kisoo Kwon, Jong Won Shin, Nam Soo Kim |
IEEE Signal Process. Lett. | 2 |
| 2014 | Parametric multichannel noise reduction algorithm utilizing temporal correlations in reverberant environmentabstractIn this paper, we propose a parametric multichannel noise reduction algorithm utilizing temporal correlations in a noisy and reverberant environment. Under the reverberant condition, the received acoustic signal becomes highly correlated in the time domain and it makes successful noise reduction quite difficult. The proposed parametric noise reduction method takes account of interdependencies between components observed from different frames. Extended speech and noise power spectral density (PSD) matrices are estimated containing additional temporal information, and the parametric multichannel noise reduction filter based on these PSD matrices is applied to the input microphone array signal. According to the experimental results, the proposed algorithm has been found to show better performances compared with the conventional multiplicative filtering technique which considers the current input signals only. Yu Gwang Jin, Jong Won Shin, Chul Min Lee, Soo Hyun Bae, Nam Soo Kim |
ICASSP | 2 |
| 2014 | Speech enhancement combining statistical models and NMF with update of speech and noise basesabstractSpeech enhancement based on statistical models has shown good performance, but the performance degrades when environment noise is highly non-stationary due to the stationary assumption. On the contrary, the template-based enhancement methods are more robust to non-stationary noise, but these are heavily dependent on a priori information present in training data. In order to get over both of the shortcomings, we propose a novel speech enhancement method which combines the statistical model-based enhancement scheme with the template-based enhancement. To reduce a dependency on a priori information, the speech and noise bases are updated simultaneously using the estimated speech presence probability, which is obtained from statistical model-based enhancement. Experimental results showed that the proposed method outperformed not only the statistical model-based and non-negative matrix factorization (NMF) approaches, but also their combination implemented with existing bases update rule in various kinds of noise. Kisoo Kwon, Jong Won Shin, Sukanya Sonowal, In Kyu Choi, Nam Soo Kim |
ICASSP | 2 |
| 2014 | Crossband filtering for stereophonic acoustic echo suppressionabstractIn this paper, we propose a novel stereophonic acoustic echo suppression (SAES) technique based on crossband filtering in the short-time Fourier transform (STFT) domain. The proposed algorithm considers spectral correlations among components in adjacent frequency bins, and estimates the extended power spectral density (PSD) matrices and cross PSD vectors from the signal statistics for more precise echo estimation. In the STFT domain, the echo spectra are estimated by performing the technique without any distinguishable double-talk detector. According to the experimental results, the proposed algorithm has been found to show better performances compared with the conventional SAES method. Chul Min Lee, Jong Won Shin, Yu Gwang Jin, Jeoung Hun Kim, Nam Soo Kim |
ICASSP | 2 |
| 2014 | NMF-based speech enhancement incorporating deep neural network
Tae Gyoon Kang, Kisoo Kwon, Jong Won Shin, Nam Soo Kim |
INTERSPEECH | 3 |
| 2014 | A data-driven approach to speech enhancement using Gaussian process
Sukanya Sonowal, Kisoo Kwon, Nam Soo Kim, Jong Won Shin |
INTERSPEECH | 4 |
| 2014 | Spectro-Temporal Filtering for Multichannel Speech Enhancement in Short-Time Fourier Transform DomainabstractIn this letter, we propose a spectro-temporal filtering algorithm for multichannel speech enhancement in the short-time Fourier transform (STFT) domain. Compared with the traditional multiplicative filtering technique, the proposed method takes account of interdependencies between components in adjacent frames and frequency bins. For spectro-temporal filtering, speech and noise power spectral density (PSD) matrices are estimated based on an extended formulation utilizing temporal and spectral correlations, and the parametric noise reduction filter based on these PSD matrices is applied to the input microphone array signal. Moreover, multichannel speech presence probabilities are also estimated within a unified framework. A number of experimental results show that the proposed spectro-temporal filtering method improves the performance of multichannel speech enhancement. Yu Gwang Jin, Jong Won Shin, Nam Soo Kim |
IEEE Signal Process. Lett. | 2 |
| 2014 | Stereophonic Acoustic Echo Suppression Incorporating Spectro-Temporal CorrelationsabstractIn this letter, we propose an enhanced stereophonic acoustic echo suppression (SAES) algorithm incorporating spectral and temporal correlations in the short-time Fourier transform (STFT) domain. Unlike traditional stereophonic acoustic echo cancellation, SAES estimates the echo spectra in the STFT domain and uses a Wiener filter to suppress echo without performing any explicit double-talk detection. The proposed approach takes account of interdependencies among components in adjacent time frames and frequency bins, which enables more accurate estimation of the echo signals. Experimental results show that the proposed method yields improved performance compared to that of conventional SAES. Chul Min Lee, Jong Won Shin, Nam Soo Kim |
IEEE Signal Process. Lett. | 2 |
| 2010 | Voice activity detection based on statistical models and machine learning approaches
Jong Won Shin, Joon-Hyuk Chang, Nam Soo Kim |
Comput. Speech Lang. | 1 |
| 2009 | DCT based multiple hashing technique for robust audio fingerprintingabstractAudio fingerprinting techniques should successfully perform content-based audio identification even when the audio files are slightly or seriously distorted. In this paper, we present a novel audio fingerprinting technique based on combining fingerprint matching results for multiple hash tables in order to improve the robustness of hashing. Multiple hash tables are built based on the discrete cosine transform (DCT) which is applied to the time sequence of energies in each sub-band. Experimental results show that the recognition errors are significantly reduced compared with Philips Robust Hash (PRH) under various distortions. Kiho Cho, Hwan Sik Yun, Jong Won Shin, Nam Soo Kim |
ICASSP | 4 |
| 2009 | Speech reinforcement based on partial masking effectabstractPerceived quality of the speech signal deteriorates significantly in the presence of ambient noise. In this paper, based on the analysis that the partial masking effect is a main source of the quality degradation when interfering signals are present, we propose a novel approach to enhance the perceived quality of speech signal when the ambient noise cannot be directly controlled by reinforcing it so that it can be heard more clearly. To find a suitable reinforcement rule, the loudness perception model proposed by Moore et al. [1] is adopted with the consideration on the prevention of the hearing damage. Experimental results show that the perceived quality and intelligibility can be enhanced under various noise environments. Jong Won Shin, Yu Gwang Jin, Seung Seop Park, Nam Soo Kim |
ICASSP | 1 |
| 2008 | Cepstral domain feature compensation based on diagonal approximationabstractIn this paper, we propose a novel approach to feature compensation performed in the cepstral domain. We apply the linear approximation method in the cepstral domain to simplify the relationship among clean speech, noise and noisy speech. Conventional log-spectral domain feature compensation methods usually assume that each log-spectral coefficient is independent, which is far from real observations. Processing in the cepstral domain has the advantage that the spectral correlation among different frequencies are taken into consideration. By using the diagonal covariance approximation, we can easily modify the conventional log-spectral domain feature compensation technique to fit to the cepstral domain. The proposed approach shows significant improvements in the AURORA2 speech recognition task. Woohyung Lim, Chang Woo Han, Jong Won Shin, Nam Soo Kim |
ICASSP | 3 |
| 2008 | Voice Activity Detection Based on Conditional MAP CriterionabstractIn this letter, we propose a novel approach to voice activity detection (VAD) based on the modified maximum a posteriori (MAP) criterion conditioned on the voice activity decision made in the previous frame. To exploit the inter-frame correlation of voice activity, the probability of the voice presence conditioned on both the observed spectrum and the voice activity decision in the previous frame is employed instead of the conventional strategy that depends only on the current observation. The proposed conditional MAP criterion incorporating temporal correlations leads to two separate thresholds for the likelihood ratio test (LRT) depending on the previous VAD result. Experimental results show that the VAD based on the proposed conditional MAP criterion outperforms the VAD based on the conventional MAP criterion under various noise environments. Jong Won Shin, Hyuk Jin Kwon, Suk Ho Jin, Nam Soo Kim |
IEEE Signal Process. Lett. | 1 |
| 2007 | A statistical model based post-filtering algorithm for residual echo suppression
Seung Yeol Lee, Jong Won Shin, Hwan Sik Yun, Nam Soo Kim |
INTERSPEECH | 2 |
| 2007 | A multiple-model based framework for automatic speech segmentation
Seung Seop Park, Jong Won Shin, Jong Kyu Kim, Nam Soo Kim |
INTERSPEECH | 2 |
| 2007 | Speech reinforcement based on partial specific loudness
Jong Won Shin, Woohyung Lim, June Sig Sung, Nam Soo Kim |
INTERSPEECH | 1 |
| 2007 | Voice activity detection based on a family of parametric distributions
Jong Won Shin, Joon-Hyuk Chang, Nam Soo Kim |
Pattern Recognit. Lett. | 1 |
| 2007 | Perceptual Reinforcement of Speech Signal Based on Partial Specific LoudnessabstractIn the presence of background noise, the perceptual loudness of speech signal significantly decreases, resulting in the deterioration of intelligibility and clarity. In this letter, we propose a novel approach to enhance the perceived quality of the speech signal when the additive noise cannot be directly controlled. Instead of controlling the background noise, we propose to reinforce the speech signal so that it can be heard more clearly in noisy environments. To find a suitable reinforcement rule, the loudness perception model proposed by Moore et al. is adopted. Experimental results show that the loudness of the reinforced signal can be maintained at the level almost the same as that of the original noise-free speech, and the proposed algorithm can enhance the perceived speech quality under various noise environments. Jong Won Shin, Nam Soo Kim |
IEEE Signal Process. Lett. | 1 |
| 2006 | Automatic speech segmentation with multiple statistical models
Seung Seop Park, Jong Won Shin, Nam Soo Kim |
INTERSPEECH | 2 |
| 2006 | Speech enhancement based on residual noise shaping
Jong Won Shin, Seung Yeol Lee, Hwan Sik Yun, Nam Soo Kim |
INTERSPEECH | 1 |
| 2006 | Signal modification for ADPCM based on analysis-by-synthesis frameworkabstractIn this letter, we propose a novel approach to improve the performance of the adaptive differential pulse code modulation (ADPCM) codec by modifying the input signal under the analysis-by-synthesis framework. Modification of the input signal is performed such that the ADPCM codec causes less quantization error. When applied to the ITU-T G.726 ADPCM coder, the proposed algorithm improves the output signal-to-noise ratio up to 2.39 dB. Jong Won Shin, Nam Soo Kim |
IEEE Signal Process. Lett. | 1 |
| 2005 | Voice Activity Detection based on Generalized Gamma DistributionabstractWe propose a voice activity detection (VAD) algorithm based on the generalized gamma distribution (G/spl Gamma/D). The distributions of noise spectra and noisy speech spectra, including speech-inactive intervals, are modeled by a set of G/spl Gamma/Ds and applied to the likelihood ratio test (LRT) for VAD. The parameters of G/spl Gamma/D are estimated through an on-line maximum likelihood (ML) estimation procedure where the global speech absence probability (GSAP) is incorporated under a forgetting scheme. Experimental results show that the proposed VAD algorithm, based on G/spl Gamma/D, outperformed the algorithms based on other statistical models. Jong Won Shin, Joon-Hyuk Chang, Hwan Sik Yun, Nam Soo Kim |
ICASSP (1) | 1 |
| 2005 | A new structural preprocessor for low-bit rate speech codingabstractIn this paper, we apply a new structural approach to generalized analysis-by-synthesis (GAbS) for system identification as a preprocessor of a low-bit-rate speech coder. In our approach, the coder-decoder (CODEC) system is separately estimated and then applied to modify the current input signal. This is different from that originally proposed where the CODEC system is sequentially estimated and then applied to the next input signal. The proposed estimation scheme is compared to the conventional method in terms of the signal modification approach under the various noise data and in several SNR conditions, and shows better performance. Joon-Hyuk Chang, Jong Won Shin, Seung Yeol Lee, Nam Soo Kim |
INTERSPEECH | 2 |
| 2005 | Image probability distribution based on generalized gamma functionabstractIn this letter, we propose results of distribution tests that indicate that for many natural images, the statistics of the discrete cosine transform (DCT) coefficients are best approximated by a generalized gamma function (G/spl Gamma/F), which includes the conventional Gaussian, Laplacian, and gamma probability density functions. The major parameter of the G/spl Gamma/F is estimated according to the maximum likelihood (ML) principle. Experimental results on a number of /spl chi//sup 2/ tests indicate that the G/spl Gamma/F can be used effectively for modeling the DCT coefficients compared to the conventional Laplacian and generalized Gaussian function (GGF). Joon-Hyuk Chang, Jong Won Shin, Nam Soo Kim, Sanjit K. Mitra |
IEEE Signal Process. Lett. | 2 |
| 2005 | Statistical modeling of speech signals based on generalized gamma distributionabstractIn this letter, we propose a new statistical model, two-sided generalized gamma distribution (G/spl Gamma/D) for an efficient parametric characterization of speech spectra. G/spl Gamma/D forms a generalized class of parametric distributions, including the Gaussian, Laplacian, and Gamma probability density functions (pdfs) as special cases. We also propose a computationally inexpensive online maximum likelihood (ML) parameter estimation algorithm for G/spl Gamma/D. Likelihoods, coefficients of variation (CVs), and Kolmogorov-Smirnov (KS) tests show that G/spl Gamma/D can model the distribution of the real speech signal more accurately than the conventional Gaussian, Laplacian, Gamma, or generalized Gaussian distribution (GGD). Jong Won Shin, Joon-Hyuk Chang, Nam Soo Kim |
IEEE Signal Process. Lett. | 1 |
| 2004 | Speech probability distribution based on generalized gama distribution
Jong Won Shin, Joon-Hyuk Chang, Nam Soo Kim |
INTERSPEECH | 1 |
| 2003 | Likelihood ratio test with complex laplacian model for voice activity detectionabstractThis paper proposes a voice activity detector (VAD) based on the complex Laplacian model. With the use of a goodness-of-fit (GOF) test, it is discovered that the Laplacian model is more suitable to describe noisy speech distribution than the conventional Gaussian model. The likelihood ratio (LR) based on the Laplacian model is computed and then applied to the VAD operation. According to the experimental results, we can find that the Laplacian statistical model is more suitable for the VAD algorithm compared to the Gaussian model. Joon-Hyuk Chang, Jong Won Shin, Nam Soo Kim |
INTERSPEECH | 2 |