EDBT 2026 Demo / reviewers in the wild / expert
Won-Gook Choi
dblp:317/0194
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Diffusion-based Target Device Style Transfer for Robust Acoustic Scene ClassificationabstractAudio signal processing systems often operate differently depending on the recording devices, leading to performance discrepancies. Therefore, it is important to know about the characteristics of the recording device; however, it is difficult to know the device’s behavior in most cases. In this study, we propose a diffusion-model-based device characteristic transfer to estimate the device’s frequency response only with the recorded signals. By joint-training the conditional and unconditional diffusion models, it is found that non-linear distortions and some filtered signals are reflected more than by only training the conditional model. We show that the proposed method transfers the style closely to the ground truth not only visually on the spectrogram but also the t-distributed stochastic neighbor embedding distribution and the performance of the device classifier. We also show the proposed method enhancing the performance as a data augmentation method for acoustic scene classification. Won-Gook Choi, Joon-Hyuk Chang |
ICASSP | 1 |
| 2025 | Optimizing CLAP Reward with LLM Feedback for Semantically Aligned and Diverse Automated Audio Captioning
Seyun Ahn, Pil Moo Byun, Won-Gook Choi, Joon-Hyuk Chang |
INTERSPEECH | 3 |
| 2025 | Temp4Cap: Temporally-aligned Automated Audio Captioning
Ho-Young Choi, Jae-Heung Cho, Pil Moo Byun, Won-Gook Choi, Joon-Hyuk Chang |
INTERSPEECH | 4 |
| 2024 | Adversarial Learning on Compressed Posterior Space for Non-Iterative Score-based End-to-End Text-to-SpeechabstractScore-based generative models have shown the real-like quality of synthesized speech in the text-to-speech (TTS) area. However, the critical artifact of score-based models is the requirement of a high computational cost due to the iterative sampling algorithm, and it also makes it difficult to fine-tune the score-based TTS-optimized vocoder. In this study, we propose a method of joint training the score-based TTS model and HiFi-GAN using the compressed log-mel features, and it guarantees a significant speech quality even on the non-iterative sampling. As a result, the proposed method overcomes some digital artifacts of the synthesized audios compared to the non-iterative sampling of Grad-TTS. Also, the non-iterative sampling can generate speech faster than other end-to-end TTS models with fewer parameters. Won-Gook Choi, Donghyun Seong, Joon-Hyuk Chang |
ICASSP | 1 |
| 2024 | Retrieval-Augmented Classifier Guidance for Audio Generation
Ho-Young Choi, Won-Gook Choi, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2024 | Sound of Vision: Audio Generation from Visual Text Embedding through Training Domain Discriminator
Won-Gook Choi, Seyun Ahn, Joon-Hyuk Chang |
INTERSPEECH | 2 |
| 2023 | Resolution Consistency Training on Time-Frequency Domain for Semi-Supervised Sound Event Detection
Won-Gook Choi, Joon-Hyuk Chang |
INTERSPEECH | 1 |
| 2023 | Prior-free Guided TTS: An Improved and Efficient Diffusion-based Text-Guided Speech Synthesis
Won-Gook Choi, So-Jeong Kim, Joon-Hyuk Chang |
INTERSPEECH | 1 |
| 2022 | Convolutional Recurrent Neural Network with Auxiliary Stream for Robust Variable-Length Acoustic Scene Classification
Joon-Hyuk Chang, Won-Gook Choi |
INTERSPEECH | 2 |