VLDB 2026 Research / reviewers in the wild / expert
Sang-gil Lee
dblp:190/7789
· DBLP profile ↗
12ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-1981-056XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Speech recognition and synthesis · 43% Generative modeling · 27% Language models and text generation · 22% | |
| Computer graphics and multimedia
4 papers |
Audio and music processing · 83% Visual content generation and editing · 17% |
Topics — the 21 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Speech recognition and synthesis › speech synthesis
text-to-speech |
1.0 | 2 | 2025 | UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation · ICLR 2025 FloWaveNet : A Generative Flow for Raw Audio · ICML 2019 |
Natural language and speech › Speech recognition and synthesis
audio-language model |
0.9 | 1 | 2025 | Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models · NeurIPS 2025 |
Natural language and speech › Language models and text generation › pre-trained language model
encoder-decoder pre-training |
0.9 | 1 | 2025 | UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation · ICLR 2025 |
Natural language and speech › Speech recognition and synthesis › audio-language model
large audio language models |
0.9 | 1 | 2025 | Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models · NeurIPS 2025 |
Natural language and speech › Language models and text generation
multimodal language model |
0.9 | 1 | 2025 | Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models · NeurIPS 2025 |
Natural language and speech › Speech recognition and synthesis
speech representation learning |
0.9 | 1 | 2025 | UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation · ICLR 2025 |
Visual content generation and editing
diffusion model |
0.9 | 1 | 2025 | ETTA: Elucidating the Design Space of Text-to-Audio Models · ICML 2025 |
Audio and music processing
sound synthesis |
0.9 | 1 | 2025 | Fugatto 1: Foundational Generative Audio Transformer Opus 1 · ICLR 2025 |
Audio and music processing › sound synthesis
text-to-audio generation |
0.9 | 1 | 2025 | ETTA: Elucidating the Design Space of Text-to-Audio Models · ICML 2025 |
Machine learning › Generative modeling
diffusion model |
0.8 | 2 | 2025 | PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior · ICLR 2022 ETTA: Elucidating the Design Space of Text-to-Audio Models · ICML 2025 |
Machine learning › Generative modeling
normalizing flow |
0.8 | 2 | 2020 | NanoFlow: Scalable Normalizing Flows with Sublinear Parameter Complexity · NeurIPS 2020 FloWaveNet : A Generative Flow for Raw Audio · ICML 2019 |
Audio and music processing › speech synthesis
neural vocoder |
0.7 | 1 | 2023 | BigVGAN: A Universal Neural Vocoder with Large-Scale Training · ICLR 2023 |
Audio and music processing
speech synthesis |
0.7 | 1 | 2023 | BigVGAN: A Universal Neural Vocoder with Large-Scale Training · ICLR 2023 |
Machine learning › Generative modeling › diffusion model
conditional diffusion model |
0.6 | 1 | 2022 | PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior · ICLR 2022 |
Machine learning › Efficient and distributed learning
model compression |
0.4 | 1 | 2020 | NanoFlow: Scalable Normalizing Flows with Sublinear Parameter Complexity · NeurIPS 2020 |
Natural language and speech › Language models and text generation
instruction following |
0.3 | 1 | 2025 | Fugatto 1: Foundational Generative Audio Transformer Opus 1 · ICLR 2025 |
Machine learning › Probabilistic and Bayesian machine learning
sampling |
0.3 | 1 | 2025 | ETTA: Elucidating the Design Space of Text-to-Audio Models · ICML 2025 |
Natural language and speech › Speech recognition and synthesis
spoken language understanding |
0.3 | 1 | 2025 | Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models · NeurIPS 2025 |
Machine learning › Generative modeling › autoregressive model
transformer-based generation |
0.3 | 1 | 2025 | Fugatto 1: Foundational Generative Audio Transformer Opus 1 · ICLR 2025 |
Audio and music processing › music information retrieval
music understanding |
0.3 | 1 | 2025 | Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models · NeurIPS 2025 |
Natural language and speech › Speech recognition and synthesis › speech synthesis
neural vocoder |
0.1 | 1 | 2019 | FloWaveNet : A Generative Flow for Raw Audio · ICML 2019 |
Methods — techniques the papers use, named apart from their topics
synthetic caption generation · 1.7flow matching · 1.7diffusion · 1.7dataset generation · 1.7curriculum learning · 1.7compositional guidance · 1.7classifier-free guidance · 1.7chain-of-thought reasoning · 1.7self-supervised learning · 0.9joint representation learning · 0.9encoder-decoder pre-training · 0.9generative adversarial network · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and InferenceabstractLarge language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often operate at high frame rates, resulting in slow training and inference, especially for autoregressive models. To address this challenge, we present the Low Frame-rate Speech Codec (LFSC): a neural audio codec that leverages finite scalar quantization and adversarial training with large speech language models to achieve high-quality audio compression with a 1.89 kbps bitrate and 21.5 frames per second. We demonstrate that our novel codec can make the inference of LLM-based text-to-speech models around three times faster while improving intelligibility and producing quality comparable to previous models. Edresson Casanova, Ryan Langman, Paarth Neekhara, Shehzeen Hussain, Jason Li 0007, Subhankar Ghosh, Ante Jukic, Sang-gil Lee |
ICASSP | 8 |
| 2025 | UniWav: Towards Unified Pre-training for Speech Representation Learning and GenerationabstractPre-training and representation learning have been playing an increasingly important role in modern speech processing. Nevertheless, different applications have been relying on different foundation models, since predominant pre-training techniques are either designed for discriminative tasks or generative tasks. In this work, we make the first attempt at building a unified pre-training framework for both types of tasks in speech. We show that with the appropriate design choices for pre-training, one can jointly learn a representation encoder and generative audio decoder that can be applied to both types of tasks. We propose UniWav, an encoder-decoder framework designed to unify pre-training representation learning and generative tasks. On speech recognition, text-to-speech, and speech tokenization, UniWav achieves comparable performance to different existing foundation models, each trained on a specific task. Our findings suggest that a single general-purpose foundation model for speech can be built to replace different foundation models, reducing the overhead and cost of pre-training. Alexander H. Liu, Sang-gil Lee, Chao-Han Huck Yang, Yuan Gong 0001, Yu-Chiang Frank Wang, James R. Glass, Rafael Valle, Bryan Catanzaro |
ICLR | 2 |
| 2025 | Fugatto 1: Foundational Generative Audio Transformer Opus 1abstractFugatto is a versatile audio synthesis and transformation model capable of following free-form text instructions with optional audio inputs. While large language models (LLMs) trained with text on a simple next-token prediction objective can learn to infer instructions directly from the data, models trained solely on audio data lack this capacity. This is because audio data does not inherently contain the instructions that were used to generate it. To overcome this challenge, we introduce a specialized dataset generation approach optimized for producing a wide range of audio generation and transformation tasks, ensuring the data reveals meaningful relationships between audio and language. Another challenge lies in achieving compositional abilities -- such as combining, interpolating between, or negating instructions -- using data alone. To address it, we propose ComposableART, an inference-time technique that extends classifier-free guidance to compositional guidance. It enables the seamless and flexible composition of instructions, leading to highly customizable audio outputs outside the training distribution. Our evaluations across a diverse set of tasks demonstrate that Fugatto performs competitively with specialized models, while ComposableART enhances its sonic palette and control over synthesis. Most notably, we highlight our framework's ability to execute emergent sounds and tasks -- sonic phenomena that transcend conventional audio generation -- unlocking new creative possibilities. \href{https://fugatto.github.io/}{Demo Website.} Rafael Valle, Rohan Badlani, Zhifeng Kong, Sang-gil Lee, Arushi Goel, Sungwon Kim 0001, João Felipe Santos, Shuqi Dai, Siddharth Gururani, Aya Aljafari, Alexander H. Liu, Kevin J. Shih, Ryan Prenger, Wei Ping, Chao-Han Huck Yang, Bryan Catanzaro |
ICLR | 4 |
| 2025 | ETTA: Elucidating the Design Space of Text-to-Audio ModelsabstractRecent years have seen significant progress in Text-To-Audio (TTA) synthesis, enabling users to enrich their creative workflows with synthetic audio generated from natural language prompts. Despite this progress, the effects of data, model architecture, training objective functions, and sampling strategies on target benchmarks are not well understood. With the purpose of providing a holistic understanding of the design space of TTA models, we set up a large-scale empirical experiment focused on diffusion and flow matching models. Our contributions include: 1) AF-Synthetic, a large dataset of high quality synthetic captions obtained from an audio understanding model; 2) a systematic comparison of different architectural, training, and inference design choices for TTA models; 3) an analysis of sampling methods and their Pareto curves with respect to generation quality and inference speed. We leverage the knowledge obtained from this extensive analysis to propose our best model dubbed Elucidated Text-To-Audio (ETTA). When evaluated on AudioCaps and MusicCaps, ETTA provides improvements over the baselines trained on publicly available data, while being competitive with models trained on proprietary data. Finally, we show ETTA’s improved ability to generate creative audio following complex and imaginative captions – a task that is more challenging than current benchmarks. Sang-gil Lee, Zhifeng Kong, Arushi Goel, Sungwon Kim 0001, Rafael Valle, Bryan Catanzaro |
ICML | 1 |
| 2025 | Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language ModelsabstractWe present Audio Flamingo 3 (AF3), a fully open state-of-the-art (SOTA) large audio-language model that advances reasoning and understanding across speech, sound, and music. AF3 introduces: (i) AF-Whisper, a unified audio encoder trained using a novel strategy for joint representation learning across all 3 modalities of speech, sound, and music; (ii) flexible, on-demand thinking, allowing the model to do chain-of-thought-type reasoning before answering; (iii) multi-turn, multi-audio chat; (iv) long audio understanding and reasoning (including speech) up to 10 minutes; and (v) voice-to-voice interaction. To enable these capabilities, we propose several large-scale training datasets curated using novel strategies, including AudioSkills-XL, LongAudio-XL, AF-Think, and AF-Chat, and train AF3 with a novel five-stage curriculum-based training strategy. Trained on only open-source audio data, AF3 achieves new SOTA results on over 20+ (long) audio understanding and reasoning benchmarks, surpassing both open-weight and closed-source models trained on much larger datasets. Sreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Kumar, Zhifeng Kong, Sang-gil Lee, Chao-Han Huck Yang, Ramani Duraiswami, Dinesh Manocha, Rafael Valle, Bryan Catanzaro |
NeurIPS | 6 |
| 2024 | VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
Heeseung Kim, Sang-gil Lee, Jiheum Yeom, Che Hyun Lee, Sungwon Kim 0001, Sungroh Yoon |
INTERSPEECH | 2 |
| 2023 | Edit-A-Video: Single Video Editing with Object-Aware Consistency
Chaehun Shin, Heeseung Kim, Che Hyun Lee, Sang-gil Lee, Sungroh Yoon |
ACML | 4 |
| 2023 | BigVGAN: A Universal Neural Vocoder with Large-Scale Training
Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, Sungroh Yoon |
ICLR | 1 |
| 2022 | PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior
Sang-gil Lee, Heeseung Kim, Chaehun Shin, Xu Tan 0003, Chang Liu 0030, Tao Qin 0001, Wei Chen 0034, Sungroh Yoon, Tie-Yan Liu |
ICLR | 1 |
| 2020 | NanoFlow: Scalable Normalizing Flows with Sublinear Parameter ComplexityabstractNormalizing flows (NFs) have become a prominent method for deep generative models that allow for an analytic probability density estimation and efficient synthesis. However, a flow-based network is considered to be inefficient in parameter complexity because of reduced expressiveness of bijective mapping, which renders the models unfeasibly expensive in terms of parameters. We present an alternative parameterization scheme called NanoFlow, which uses a single neural density estimator to model multiple transformation stages. Hence, we propose an efficient parameter decomposition method and the concept of flow indication embedding, which are key missing components that enable density estimation from a single neural network. Experiments performed on audio and image models confirm that our method provides a new parameter-efficient solution for scalable NFs with significant sublinear parameter complexity. Sang-gil Lee, Sungwon Kim 0001, Sungroh Yoon |
NeurIPS | 1 |
| 2019 | FloWaveNet : A Generative Flow for Raw AudioabstractMost modern text-to-speech architectures use a WaveNet vocoder for synthesizing high-fidelity waveform audio, but there have been limitations, such as high inference time, in practical applications due to its ancestral sampling scheme. The recently suggested Parallel WaveNet and ClariNet has achieved real-time audio synthesis capability by incorporating inverse autoregressive flow (IAF) for parallel sampling. However, these approaches require a two-stage training pipeline with a well-trained teacher network and can only produce natural sound by using probability distillation along with heavily-engineered auxiliary loss terms. We propose FloWaveNet, a flow-based generative model for raw audio synthesis. FloWaveNet requires only a single-stage training procedure and a single maximum likelihood loss, without any additional auxiliary terms, and it is inherently parallel due to the characteristics of generative flow. The model can efficiently sample raw audio in real-time, with clarity comparable to previous two-stage parallel models. The code and samples for all models, including our FloWaveNet, are available on GitHub. Sungwon Kim 0001, Sang-gil Lee, Jongyoon Song, Jaehyeon Kim, Sungroh Yoon |
ICML | 2 |
| 2018 | Liver Lesion Detection from Weakly-Labeled Multi-phase CT Volumes with a Grouped Single Shot MultiBox Detector
Sang-gil Lee, Jae Seok Bae, Hyunjae Kim, Sungroh Yoon |
MICCAI (2) | 1 |