VLDB 2026 Research / reviewers in the wild / expert
Min-Seok Choi
dblp:86/5039
· DBLP profile ↗
11ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0002-1214-9799ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Listen through the Sound: Generative Speech Restoration Leveraging Acoustic Context Representation
Soo-Whan Chung, Min-Seok Choi |
INTERSPEECH | 2 |
| 2024 | Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation
Miseul Kim, Soo-Whan Chung, Youna Ji, Hong-Goo Kang, Min-Seok Choi |
INTERSPEECH | 5 |
| 2023 | An Empirical Study on Speech Restoration Guided by Self-Supervised Speech RepresentationabstractEnhancing speech quality is an indispensable yet difficult task as it is often complicated by a range of degradation factors. In addition to additive noise, reverberation, clipping, and speech attenuation can all adversely affect speech quality. Speech restoration aims to recover speech components from these distortions. This paper focuses on exploring the impact of self-supervised speech representation learning on the speech restoration task. Specifically, we employ speech representation in various speech restoration networks and evaluate their performance under complicated distortion scenarios. Our experiments demonstrate that the contextual information provided by the self-supervised speech representation can enhance speech restoration performance in various distortion scenarios, while also increasing robustness against the duration of speech attenuation and mismatched test conditions. Jaeuk Byun, Youna Ji, Soo-Whan Chung, Soyeon Choe, Min-Seok Choi |
ICASSP | 5 |
| 2023 | Diffusion-Based Generative Speech Source SeparationabstractWe propose DiffSep, a new single channel source separation method based on score-matching of a stochastic differential equation (SDE). We craft a tailored continuous time diffusion-mixing process starting from the separated sources and converging to a Gaussian distribution centered on their mixture. This formulation lets us apply the machinery of score-based generative modelling. First, we train a neural network to approximate the score function of the marginal probabilities of the diffusion-mixing process. Then, we use it to solve the reverse time SDE that progressively separates the sources starting from their mixture. We propose a modified training strategy to handle model mismatch and source permutation ambiguity. Experiments on the WSJ0_2mix dataset demonstrate the potential of the method. Furthermore, the method is also suitable for speech enhancement and shows performance competitive with prior work on the VoiceBank-DEMAND dataset. Robin Scheibler, Youna Ji, Soo-Whan Chung, Jaeuk Byun, Soyeon Choe, Min-Seok Choi |
ICASSP | 6 |
| 2021 | DEMUCS-Mobile : On-Device Lightweight Speech Enhancement
Lukas Lee, Youna Ji, Min-Seok Choi |
Interspeech | 4 |
| 2011 | A Two-Channel Noise Estimator for Speech Enhancement in a Highly Nonstationary EnvironmentabstractThis paper proposes a two-channel noise estimator for speech enhancement in a highly nonstationary environment. The proposed noise estimator utilizes a spatial filter which has a capability of extracting noise information even in a speech presence region. We exploit a first-order recursion method with time-frequency varying smoothing coefficients to accurately estimate a noise power spectral density (PSD) in both slowly and rapidly varying regions. The smoothing coefficients are determined by measuring the nonstationarity factor of noise, e.g., degree of noise variation. The nonstationarity factor is derived through a statistical assumption of stationary background noise, which does not need any assumption on the type of nonstationary noise. Since the proposed method efficiently estimates the noise PSD both in stationary and nonstationary regions, the enhanced speech obtained by applying the proposed algorithm to the two-channel enhancement system shows superior performance to conventional approaches in various noise environments. Min-Seok Choi, Hong-Goo Kang |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2010 | Binaural loudness based speech reinforcement with a closed-form solutionabstractThis paper addresses a perceptual signal processing to far-end speech signal in communication systems under near-end environmental noise conditions. Based on the binaural perceptual loudness model, the proposed speech reinforcement system achieves better speech quality and clearness. To effectively reflect the noise influence to both ears, the proposed method utilizes a noise level difference between open and receiver side ear. Its computational complexity is also reduced by deriving an approximated closed-form solution while computing frequency-dependent gain factors. Test results confirm that the proposed system significantly enhances the clearness of target speech while maintaining speech quality compared to conventional monaural-based one. Ho Seon Shin, Min-Seok Choi, Taesu Kim, Hong-Goo Kang |
ICASSP | 2 |
| 2007 | A Soft-Decision Adaptation Mode Controller for an Efficient Frequency-Domain Generalized Sidelobe CancellerabstractIn this paper, we propose a new soft-decision adaptation mode controller (SD-AMC) for frequency domain generalized sidelobe canceller (GSC) as a speech enhancement system. Contrarily to conventional systems that update filter coefficients in a hard-decision manner using voice activity detection (VAD), the proposed method flexibly controls the step-sizes of adaptive filters depending on the probability of speech presence in each frequency bin. Therefore, it further improves the system performance for various environments without much consideration on noise type and signal to noise ratio (SNR) of input signal. It also improves the robustness of GSC system by avoiding the miss-classification error by the hard-decision logic. Experimental results with speech recognition systems verify that the SD-AMC shows higher performance than ideally designed hard-decision approaches. Min-Seok Choi, Chang-Hyun Baik, Young-Cheol Park, Hong-Goo Kang |
ICASSP (4) | 1 |
| 2005 | An improved estimation of a priori speech absence probability for speech enhancement : in perspective of speech perceptionabstractThe purpose of this paper is to improve the perceptual quality of a single channel speech enhancement algorithm using MMSE LSA estimator. The proposed algorithm uses a nonlinear decision rule and an adaptive recursive averaging factor for tracking a priori speech absence probability (SAP) fast. We also introduce one-third of approximated critical bandwidth to efficiently smooth the a priori SAP and final gain term, which successfully eliminates the musical noise without much distortion of signal. The performance of the proposed algorithm is evaluated by performing subjective AB listening tests and measuring spectral distance. Simulation results verify the effectiveness of the proposed algorithm compared to conventional algorithms. Min-Seok Choi, Hong-Goo Kang |
ICASSP (1) | 1 |
| 2004 | The description and retrieval of a sequence of moving objects using a shape variation map
Min-Seok Choi, Whoi-Yul Kim |
Pattern Recognit. Lett. | 1 |
| 2002 | A novel two stage template matching method for rotation and illumination invariance
Min-Seok Choi, Whoi-Yul Kim |
Pattern Recognit. | 1 |