Sang-gil Lee

dblp:190/7789 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-1981-056XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Speech recognition and synthesis · 43% Generative modeling · 27% Language models and text generation · 22%
Computer graphics and multimedia
4 papers
Audio and music processing · 83% Visual content generation and editing · 17%

Topics — the 21 heaviest of 22, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Speech recognition and synthesis › speech synthesis
text-to-speech
1.022025
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation · ICLR 2025
FloWaveNet : A Generative Flow for Raw Audio · ICML 2019
Natural language and speech › Speech recognition and synthesis
audio-language model
0.912025
Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models · NeurIPS 2025
Natural language and speech › Language models and text generation › pre-trained language model
encoder-decoder pre-training
0.912025
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation · ICLR 2025
Natural language and speech › Speech recognition and synthesis › audio-language model
large audio language models
0.912025
Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models · NeurIPS 2025
Natural language and speech › Language models and text generation
multimodal language model
0.912025
Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models · NeurIPS 2025
Natural language and speech › Speech recognition and synthesis
speech representation learning
0.912025
UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation · ICLR 2025
Visual content generation and editing
diffusion model
0.912025
ETTA: Elucidating the Design Space of Text-to-Audio Models · ICML 2025
Audio and music processing
sound synthesis
0.912025
Fugatto 1: Foundational Generative Audio Transformer Opus 1 · ICLR 2025
Audio and music processing › sound synthesis
text-to-audio generation
0.912025
ETTA: Elucidating the Design Space of Text-to-Audio Models · ICML 2025
Machine learning › Generative modeling
diffusion model
0.822025
PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior · ICLR 2022
ETTA: Elucidating the Design Space of Text-to-Audio Models · ICML 2025
Machine learning › Generative modeling
normalizing flow
0.822020
NanoFlow: Scalable Normalizing Flows with Sublinear Parameter Complexity · NeurIPS 2020
FloWaveNet : A Generative Flow for Raw Audio · ICML 2019
Audio and music processing › speech synthesis
neural vocoder
0.712023
BigVGAN: A Universal Neural Vocoder with Large-Scale Training · ICLR 2023
Audio and music processing
speech synthesis
0.712023
BigVGAN: A Universal Neural Vocoder with Large-Scale Training · ICLR 2023
Machine learning › Generative modeling › diffusion model
conditional diffusion model
0.612022
PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior · ICLR 2022
Machine learning › Efficient and distributed learning
model compression
0.412020
NanoFlow: Scalable Normalizing Flows with Sublinear Parameter Complexity · NeurIPS 2020
Natural language and speech › Language models and text generation
instruction following
0.312025
Fugatto 1: Foundational Generative Audio Transformer Opus 1 · ICLR 2025
Machine learning › Probabilistic and Bayesian machine learning
sampling
0.312025
ETTA: Elucidating the Design Space of Text-to-Audio Models · ICML 2025
Natural language and speech › Speech recognition and synthesis
spoken language understanding
0.312025
Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models · NeurIPS 2025
Machine learning › Generative modeling › autoregressive model
transformer-based generation
0.312025
Fugatto 1: Foundational Generative Audio Transformer Opus 1 · ICLR 2025
Audio and music processing › music information retrieval
music understanding
0.312025
Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models · NeurIPS 2025
Natural language and speech › Speech recognition and synthesis › speech synthesis
neural vocoder
0.112019
FloWaveNet : A Generative Flow for Raw Audio · ICML 2019

Methods — techniques the papers use, named apart from their topics

synthetic caption generation · 1.7flow matching · 1.7diffusion · 1.7dataset generation · 1.7curriculum learning · 1.7compositional guidance · 1.7classifier-free guidance · 1.7chain-of-thought reasoning · 1.7self-supervised learning · 0.9joint representation learning · 0.9encoder-decoder pre-training · 0.9generative adversarial network · 0.7
YearPublicationVenuePosition
2025 Low Frame-rate Speech Codec: a Codec Designed for Fast High-quality Speech LLM Training and Inference
abstract
Large language models (LLMs) have significantly advanced audio processing through audio codecs that convert audio into discrete tokens, enabling the application of language modeling techniques to audio data. However, audio codecs often operate at high frame rates, resulting in slow training and inference, especially for autoregressive models. To address this challenge, we present the Low Frame-rate Speech Codec (LFSC): a neural audio codec that leverages finite scalar quantization and adversarial training with large speech language models to achieve high-quality audio compression with a 1.89 kbps bitrate and 21.5 frames per second. We demonstrate that our novel codec can make the inference of LLM-based text-to-speech models around three times faster while improving intelligibility and producing quality comparable to previous models.
Edresson Casanova, Ryan Langman, Paarth Neekhara, Shehzeen Hussain, Jason Li 0007, Subhankar Ghosh, Ante Jukic, Sang-gil Lee
ICASSP8
2025 UniWav: Towards Unified Pre-training for Speech Representation Learning and Generation
abstract
Pre-training and representation learning have been playing an increasingly important role in modern speech processing. Nevertheless, different applications have been relying on different foundation models, since predominant pre-training techniques are either designed for discriminative tasks or generative tasks. In this work, we make the first attempt at building a unified pre-training framework for both types of tasks in speech. We show that with the appropriate design choices for pre-training, one can jointly learn a representation encoder and generative audio decoder that can be applied to both types of tasks. We propose UniWav, an encoder-decoder framework designed to unify pre-training representation learning and generative tasks. On speech recognition, text-to-speech, and speech tokenization, UniWav achieves comparable performance to different existing foundation models, each trained on a specific task. Our findings suggest that a single general-purpose foundation model for speech can be built to replace different foundation models, reducing the overhead and cost of pre-training.
Alexander H. Liu, Sang-gil Lee, Chao-Han Huck Yang, Yuan Gong 0001, Yu-Chiang Frank Wang, James R. Glass, Rafael Valle, Bryan Catanzaro
ICLR2
2025 Fugatto 1: Foundational Generative Audio Transformer Opus 1
abstract
Fugatto is a versatile audio synthesis and transformation model capable of following free-form text instructions with optional audio inputs. While large language models (LLMs) trained with text on a simple next-token prediction objective can learn to infer instructions directly from the data, models trained solely on audio data lack this capacity. This is because audio data does not inherently contain the instructions that were used to generate it. To overcome this challenge, we introduce a specialized dataset generation approach optimized for producing a wide range of audio generation and transformation tasks, ensuring the data reveals meaningful relationships between audio and language. Another challenge lies in achieving compositional abilities -- such as combining, interpolating between, or negating instructions -- using data alone. To address it, we propose ComposableART, an inference-time technique that extends classifier-free guidance to compositional guidance. It enables the seamless and flexible composition of instructions, leading to highly customizable audio outputs outside the training distribution. Our evaluations across a diverse set of tasks demonstrate that Fugatto performs competitively with specialized models, while ComposableART enhances its sonic palette and control over synthesis. Most notably, we highlight our framework's ability to execute emergent sounds and tasks -- sonic phenomena that transcend conventional audio generation -- unlocking new creative possibilities. \href{https://fugatto.github.io/}{Demo Website.}
Rafael Valle, Rohan Badlani, Zhifeng Kong, Sang-gil Lee, Arushi Goel, Sungwon Kim 0001, João Felipe Santos, Shuqi Dai, Siddharth Gururani, Aya Aljafari, Alexander H. Liu, Kevin J. Shih, Ryan Prenger, Wei Ping, Chao-Han Huck Yang, Bryan Catanzaro
ICLR4
2025 ETTA: Elucidating the Design Space of Text-to-Audio Models
abstract
Recent years have seen significant progress in Text-To-Audio (TTA) synthesis, enabling users to enrich their creative workflows with synthetic audio generated from natural language prompts. Despite this progress, the effects of data, model architecture, training objective functions, and sampling strategies on target benchmarks are not well understood. With the purpose of providing a holistic understanding of the design space of TTA models, we set up a large-scale empirical experiment focused on diffusion and flow matching models. Our contributions include: 1) AF-Synthetic, a large dataset of high quality synthetic captions obtained from an audio understanding model; 2) a systematic comparison of different architectural, training, and inference design choices for TTA models; 3) an analysis of sampling methods and their Pareto curves with respect to generation quality and inference speed. We leverage the knowledge obtained from this extensive analysis to propose our best model dubbed Elucidated Text-To-Audio (ETTA). When evaluated on AudioCaps and MusicCaps, ETTA provides improvements over the baselines trained on publicly available data, while being competitive with models trained on proprietary data. Finally, we show ETTA’s improved ability to generate creative audio following complex and imaginative captions – a task that is more challenging than current benchmarks.
Sang-gil Lee, Zhifeng Kong, Arushi Goel, Sungwon Kim 0001, Rafael Valle, Bryan Catanzaro
ICML1
2025 Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models
abstract
We present Audio Flamingo 3 (AF3), a fully open state-of-the-art (SOTA) large audio-language model that advances reasoning and understanding across speech, sound, and music. AF3 introduces: (i) AF-Whisper, a unified audio encoder trained using a novel strategy for joint representation learning across all 3 modalities of speech, sound, and music; (ii) flexible, on-demand thinking, allowing the model to do chain-of-thought-type reasoning before answering; (iii) multi-turn, multi-audio chat; (iv) long audio understanding and reasoning (including speech) up to 10 minutes; and (v) voice-to-voice interaction. To enable these capabilities, we propose several large-scale training datasets curated using novel strategies, including AudioSkills-XL, LongAudio-XL, AF-Think, and AF-Chat, and train AF3 with a novel five-stage curriculum-based training strategy. Trained on only open-source audio data, AF3 achieves new SOTA results on over 20+ (long) audio understanding and reasoning benchmarks, surpassing both open-weight and closed-source models trained on much larger datasets.
Sreyan Ghosh, Arushi Goel, Jaehyeon Kim, Sonal Kumar, Zhifeng Kong, Sang-gil Lee, Chao-Han Huck Yang, Ramani Duraiswami, Dinesh Manocha, Rafael Valle, Bryan Catanzaro
NeurIPS6
2024 VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
Heeseung Kim, Sang-gil Lee, Jiheum Yeom, Che Hyun Lee, Sungwon Kim 0001, Sungroh Yoon
INTERSPEECH2
2023 Edit-A-Video: Single Video Editing with Object-Aware Consistency
Chaehun Shin, Heeseung Kim, Che Hyun Lee, Sang-gil Lee, Sungroh Yoon
ACML4
2023 BigVGAN: A Universal Neural Vocoder with Large-Scale Training
Sang-gil Lee, Wei Ping, Boris Ginsburg, Bryan Catanzaro, Sungroh Yoon
ICLR1
2022 PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior
Sang-gil Lee, Heeseung Kim, Chaehun Shin, Xu Tan 0003, Chang Liu 0030, Tao Qin 0001, Wei Chen 0034, Sungroh Yoon, Tie-Yan Liu
ICLR1
2020 NanoFlow: Scalable Normalizing Flows with Sublinear Parameter Complexity
abstract
Normalizing flows (NFs) have become a prominent method for deep generative models that allow for an analytic probability density estimation and efficient synthesis. However, a flow-based network is considered to be inefficient in parameter complexity because of reduced expressiveness of bijective mapping, which renders the models unfeasibly expensive in terms of parameters. We present an alternative parameterization scheme called NanoFlow, which uses a single neural density estimator to model multiple transformation stages. Hence, we propose an efficient parameter decomposition method and the concept of flow indication embedding, which are key missing components that enable density estimation from a single neural network. Experiments performed on audio and image models confirm that our method provides a new parameter-efficient solution for scalable NFs with significant sublinear parameter complexity.
Sang-gil Lee, Sungwon Kim 0001, Sungroh Yoon
NeurIPS1
2019 FloWaveNet : A Generative Flow for Raw Audio
abstract
Most modern text-to-speech architectures use a WaveNet vocoder for synthesizing high-fidelity waveform audio, but there have been limitations, such as high inference time, in practical applications due to its ancestral sampling scheme. The recently suggested Parallel WaveNet and ClariNet has achieved real-time audio synthesis capability by incorporating inverse autoregressive flow (IAF) for parallel sampling. However, these approaches require a two-stage training pipeline with a well-trained teacher network and can only produce natural sound by using probability distillation along with heavily-engineered auxiliary loss terms. We propose FloWaveNet, a flow-based generative model for raw audio synthesis. FloWaveNet requires only a single-stage training procedure and a single maximum likelihood loss, without any additional auxiliary terms, and it is inherently parallel due to the characteristics of generative flow. The model can efficiently sample raw audio in real-time, with clarity comparable to previous two-stage parallel models. The code and samples for all models, including our FloWaveNet, are available on GitHub.
Sungwon Kim 0001, Sang-gil Lee, Jongyoon Song, Jaehyeon Kim, Sungroh Yoon
ICML2
2018 Liver Lesion Detection from Weakly-Labeled Multi-phase CT Volumes with a Grouped Single Shot MultiBox Detector
Sang-gil Lee, Jae Seok Bae, Hyunjae Kim, Sungroh Yoon
MICCAI (2)1