Youna Ji

dblp:136/5089 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
5since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 10 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2024 Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation
Miseul Kim, Soo-Whan Chung, Youna Ji, Hong-Goo Kang, Min-Seok Choi
INTERSPEECH3
2023 An Empirical Study on Speech Restoration Guided by Self-Supervised Speech Representation
abstract
Enhancing speech quality is an indispensable yet difficult task as it is often complicated by a range of degradation factors. In addition to additive noise, reverberation, clipping, and speech attenuation can all adversely affect speech quality. Speech restoration aims to recover speech components from these distortions. This paper focuses on exploring the impact of self-supervised speech representation learning on the speech restoration task. Specifically, we employ speech representation in various speech restoration networks and evaluate their performance under complicated distortion scenarios. Our experiments demonstrate that the contextual information provided by the self-supervised speech representation can enhance speech restoration performance in various distortion scenarios, while also increasing robustness against the duration of speech attenuation and mismatched test conditions.
Jaeuk Byun, Youna Ji, Soo-Whan Chung, Soyeon Choe, Min-Seok Choi
ICASSP2
2023 Diffusion-Based Generative Speech Source Separation
abstract
We propose DiffSep, a new single channel source separation method based on score-matching of a stochastic differential equation (SDE). We craft a tailored continuous time diffusion-mixing process starting from the separated sources and converging to a Gaussian distribution centered on their mixture. This formulation lets us apply the machinery of score-based generative modelling. First, we train a neural network to approximate the score function of the marginal probabilities of the diffusion-mixing process. Then, we use it to solve the reverse time SDE that progressively separates the sources starting from their mixture. We propose a modified training strategy to handle model mismatch and source permutation ambiguity. Experiments on the WSJ0_2mix dataset demonstrate the potential of the method. Furthermore, the method is also suitable for speech enhancement and shows performance competitive with prior work on the VoiceBank-DEMAND dataset.
Robin Scheibler, Youna Ji, Soo-Whan Chung, Jaeuk Byun, Soyeon Choe, Min-Seok Choi
ICASSP2
2023 HD-DEMUCS: General Speech Restoration with Heterogeneous Decoders
Soo-Whan Chung, Hyewon Han, Youna Ji, Hong-Goo Kang
INTERSPEECH4
2021 DEMUCS-Mobile : On-Device Lightweight Speech Enhancement
Lukas Lee, Youna Ji, Min-Seok Choi
Interspeech2
2017 Coherence-Based Dual-Channel Noise Reduction Algorithm in a Complex Noisy Environment
Youna Ji, Jun Byun, Young-Cheol Park
INTERSPEECH1
2016 Improved a priori SAP Estimator in Complex Noisy Environment for Dual Channel Microphone System
Youna Ji, Young-Cheol Park
INTERSPEECH1
2015 A priori SAP estimator based on the magnitude square coherence for dual-channel microphone system
abstract
In this paper, we present a time-frequency (TF)-dependent a priori speech absence probability (SAP) estimator utilizing the magnitude square coherence (MSC) between two microphone signals. It is shown that the normalized SNR can be numerically computed from the MSC by solving a quadratic equation. Based on the fact that the normalized SNR is bounded between 0 and 1, we directly use it for the probability of speech absence in each TF-unit. Since this approach does not require prior statistical knowledge of noise and speech, it is not affected by the performance of the noise PSD estimator. Furthermore, unlike the conventional SNR-based estimator, additional mapping strategy is unnecessary. The algorithm was evaluated using the receiver operating characteristic (ROC) curve and it attained higher correct detection rate at a given false-alarm rate than the conventional algorithms.
Youna Ji, Yonghyun Baek, Young-Cheol Park
ICASSP1
2014 Binaural noise suppression based on an unbiased estimator of target PSD in complex noise environments
abstract
The conventional target power density spectrum (PSD) estimation methods based on the signal prediction inherently produce a biased target PSD because of irrelevant assumptions for the noisy environment. In this paper, an unbiased target PSD is obtained by removing the effect of diffuse noise on the prediction filter. In addition, by on-line estimation of both the noise PSD and target transfer function ratio (TFR) from the input signals, the proposed algorithm achieves robust noise suppression for an unknown target direction under a fast time-varying noisy environment. Computer simulations demonstrate the effectiveness and superiority of the proposed algorithm over the conventional methods.
Youna Ji, Young-Cheol Park, Dae Hee Youn
ICASSP1
2013 Robust noise PSD estimation for binaural hearing aids in time-varying diffuse noise field
abstract
In this paper, we present an unsupervised noise PSD estimation algorithm for binaural hearing aids in a time-varying diffuse noise field. It is shown that the noise PSD can be obtained from the eigenvalues of the input covariance matrix together with the noise coherence function effective at low frequencies. To reduce the estimation bias due to fast smoothing, pre- and post-compensation methods are proposed. The proposed algorithm is able to track non-stationary noise PSD without tracking delay or underestimation problems. Its performance is independent of the target speech direction and input SNR. Results of the objective parameter evaluation demonstrate the superiority of the proposed algorithm over conventional techniques.
Youna Ji, Young-Cheol Park, Junil Sohn
ICASSP1