VLDB 2026 Research / reviewers in the wild / expert
Da-Hee Yang
dblp:330/9169
· DBLP profile ↗
13ranked-venue papers
9as first author
13since 2021 · last 2026
0009-0008-6253-2543ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 11 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 8 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An experimental study of diffusion-based general speech restoration with predictive-guided conditioning
Da-Hee Yang, Joon-Hyuk Chang |
Comput. Speech Lang. | 1 |
| 2026 | A dual-branch parallel network for speech enhancement and restoration
Da-Hee Yang, Dail Kim, Joon-Hyuk Chang, Jeonghwan Choi, Han-Gil Moon |
Comput. Speech Lang. | 1 |
| 2026 | Latent-Level Enhancement With Flow Matching for Robust Automatic Speech RecognitionabstractNoise-robust automatic speech recognition (ASR) has been commonly addressed by applying speech enhancement (SE) at the waveform level before recognition. However, speech-level enhancement does not always translate into consistent recognition improvements due to residual distortions and mismatches with the latent space of the ASR encoder. In this letter, we introduce a complementary strategy termed latent-level enhancement, where distorted representations are refined during ASR inference. Specifically, we propose a plug-and-play Flow Matching Refinement module (FM-Refiner) that operates on the output latents of a pretrained CTC-based ASR encoder. Trained to map imperfect latents—either directly from noisy inputs or from enhanced-but-imperfect speech—toward their clean counterparts, the FM-Refiner is applied only at inference, without fine-tuning ASR parameters. Experiments show that FM-Refiner consistently reduces word error rate, both when directly applied to noisy inputs and when combined with conventional SE front-ends. These results demonstrate that latent-level refinement via flow matching provides a lightweight and effective complement to existing SE approaches for robust ASR. Da-Hee Yang, Joon-Hyuk Chang |
IEEE Signal Process. Lett. | 1 |
| 2025 | A Momentum-Based Framework with Contrastive Data Generation for Robust Sound Source LocalizationabstractWe propose MoCo-SSL, a momentum-based contrastive learning framework for multi-channel sound source localization (SSL) that enhances azimuth-aware representation learning. While prior SSL studies have used contrastive learning to handle varied acoustic conditions, we emphasize hard negatives-pairs with distinct azimuths recorded in the same room-for learning fine-grained spatial cues. A curriculum-based strategy gradually increases the proportion of such samples to raise task difficulty. The momentum contrast design employs a key encoder that maintains stable embeddings during curriculum transitions and receives audio with less noise and reverberation to produce clearer azimuth cues, thereby guiding the query encoder toward robust representations. Experiments show that MoCo-SSL consistently surpasses baselines, demonstrating the value of structured and noise-resilient representation learning in challenging SSL scenarios. Hyun-Soo Kim, Da-Hee Yang, Joon-Hyuk Chang |
ASRU | 2 |
| 2025 | Improving Noise Robust Audio-Visual Speech Recognition via Router-Gated Cross-Modal Feature FusionabstractRobust audio-visual speech recognition (AVSR) in noisy environments remains challenging, as existing systems struggle to estimate audio reliability and dynamically adjust modality reliance. We propose router-gated cross-modal feature fusion, a novel AVSR framework that adaptively reweights audio and visual features based on token-level acoustic corruption scores. Using an audio-visual feature fusion-based router, our method down-weights unreliable audio tokens and reinforces visual cues through gated cross-attention in each decoder layer. This enables the model to pivot toward the visual modality when audio quality deteriorates. Experiments on LRS3 demonstrate that our approach achieves an 16.51-42.67% relative reduction in word error rate compared to AV-HuBERT. Ablation studies confirm that both the router and gating mechanism contribute to improved robustness under real-world acoustic noise. DongHoon Lim, YoungChae Kim, Da-Hee Yang, Joon-Hyuk Chang |
ASRU | 4 |
| 2025 | Spatially Weighted Contrastive Learning for Robust Sound Source Localization
Hyun-Soo Kim, Da-Hee Yang, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2025 | Flow-PLC: Towards Efficient Packet Loss Concealment With Flow MatchingabstractRecent advancements in packet loss concealment (PLC) have introduced diffusion-based generative models that offer high-quality audio reconstruction. However, their high computational costs make them impractical for real-time applications. In this letter, we present Flow-PLC, an efficient PLC model based on the flow-matching framework, designed to address these computational challenges. Flow-PLC achieves a remarkable 23× reduction in inference time compared to diffusion-based PLC models, requiring only five sampling steps to achieve near-optimal reconstruction. By significantly reducing computational complexity while maintaining high-quality results, Flow-PLC represents a substantial advancement in the development of efficient and practical generative PLC systems. Da-Hee Yang, Joon-Hyuk Chang |
IEEE Signal Process. Lett. | 1 |
| 2025 | Tokenized Generative Speech Enhancement With Language Model and Flow MatchingabstractWe propose a novel generative speech enhancement (SE) framework that integrates a language model (LM) and a flow-matching model. To utilize an LM with discrete tokens, we introduce dMel, which discretizes Mel spectrograms into a predefined set of quantized values on a linear-scale without requiring additional neural networks. dMel preserves both semantic and acoustic characteristics, providing a compact and effective token-based alternative to Mel spectrograms. We design the first encoder-decoder LM for SE, which learns to map noisy dMel to enhanced ones. Subsequently, flow-matching de-quantizes enhanced dMel into continuous representation and refines it by learning the optimal transport-based probability path, improving perceptual quality. This unified approach enables structured reconstruction while effectively suppressing noise. Experimental results demonstrate the effectiveness of our method in enhancing speech quality, establishing a new paradigm for generative SE without reliance on neural codec-based representations. Da-Hee Yang, Jaeuk Lee, Joon-Hyuk Chang |
IEEE Signal Process. Lett. | 1 |
| 2024 | Guided conditioning with predictive network on score-based diffusion model for speech enhancement
Dail Kim, Da-Hee Yang, Joon-Hyuk Chang, Jeonghwan Choi, Moa Lee, Jaemo Yang, Han-Gil Moon |
INTERSPEECH | 2 |
| 2024 | Diff-PLC: A Diffusion-Based Approach For Effective Packet Loss ConcealmentabstractWe introduce diffusion-based packet loss concealment (DiffPLC), a novel approach designed to improve speech quality in the presence of packet losses for speech transmission. Derived from the foundation of a diffusion-based neural vocoder, the Diff-PLC introduces a crucial modification and supplementary concepts for the reconstruction of lost packets. A key aspect of the Diff-PLC involves integrating a feature-wise linear modulation layer into the diffusion model, facilitating the seamless incorporation of a conditioning feature. Furthermore, the Diff-PLC leverages packet loss embedding as an additional conditioning feature which significantly assists the diffusion model in restoring lost packets. The proposed model is evaluated using the blind test set of the INTERSPEECH 2022 PLC challenge, demonstrating the considerable restoration capabilities of Diff-PLC across various reference-free and reference-based metrics, including PLCMOS, PESQ, STOI, and NISQA. Da-Hee Yang, Joon-Hyuk Chang |
SLT | 1 |
| 2023 | Towards Robust Packet Loss Concealment System With ASR-Guided RepresentationsabstractDespite the significant advancements and promising performance of deep learning-based packet loss concealment (PLC) systems in transmission systems, their focus on modeling acoustic features for reconstructing lost packets is insufficient to achieve smooth transitions during speech reconstruction. Therefore, to address this limitation, we propose integrating linguistic information derived from a speech recognition system as auxiliary features in the PLC system. By extracting ASR-guided representations and incorporating them using auxiliary loss, we successfully demonstrate a substantial improvement in the perceptual quality and intelligibility of the reconstructed speech. Our evaluation conducted on the wall street journal dataset further validates the effectiveness of our approach through experiments involving different packet loss rates and performance metrics. Da-Hee Yang, Joon-Hyuk Chang |
ASRU | 1 |
| 2023 | Selective Film Conditioning with CTC-Based ASR Probability for Speech EnhancementabstractEnhancing speech quality and intelligibility for automatic speech recognition (ASR) plays an important role in modeling speech enhancement (SE) systems. However, improving the ASR performance by utilizing SE networks is not guaranteed, owing to the discrepancy in the training methods of the two systems. Therefore, recent studies have gradually incorporated ASR information into SE systems by jointly training ASR and SE systems. Although prior studies have improved the performance, they are inefficient because the two networks are combined and require large model sizes. To address this limitation, we propose an efficient way to use feature-wise linear modulation (FiLM) conditioning with CTC-based ASR probabilities for the SE system. The proposed model is designed by stacking a FiLM layer with selective learning on each temporal convolutional network of the SE estimation module. This allows the SE network to adaptively select ASR information based on the relationship between context and acoustic information. The proposed method improves SE and ASR performance, resulting in more robust results against noise with only a small increase in the number of parameters. Da-Hee Yang, Joon-Hyuk Chang |
ICASSP | 1 |
| 2022 | FiLM Conditioning with Enhanced Feature to the Transformer-based End-to-End Noisy Speech Recognition
Da-Hee Yang, Joon-Hyuk Chang |
INTERSPEECH | 1 |