VLDB 2026 Research / reviewers in the wild / expert
Heeseung Kim
dblp:294/8710
· DBLP profile ↗
14ranked-venue papers
4as first author
14since 2021 · last 2026
0009-0001-1402-2425ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Still Between Us? Evaluating and Improving Voice Assistant Robustness to Third-Party InterruptionsabstractWhile recent Spoken Language Models (SLMs) have been actively deployed in real-world scenarios, they lack the capability to discern Third-Party Interruptions (TPI) from the primary user's ongoing flow, leaving them vulnerable to contextual failures.To bridge this gap, we introduce TPI-Train, a dataset of 88K instances designed with speaker-aware hard negatives to enforce acoustic cue prioritization for interruption handling, and TPI-Bench, a comprehensive evaluation framework designed to rigorously measure the interruption-handling strategy and precise speaker discrimination in deceptive contexts.Experiments demonstrate that our dataset design mitigates semantic shortcut learning-a critical pitfall where models exploit semantic context while neglecting acoustic signals essential for discerning speaker changes.We believe our work establishes a foundational resource for overcoming text-dominated unimodal reliance in SLMs, paving the way for more robust multi-party spoken interaction.The code for the framework is publicly available at https://tpi-va.github.io/.↩→ * **Definition:** A Task Enhancer helps the VA fulfill the user's request to VA more accurately, or efficiently.* **Examples of NonIgnorable interruptions include, but are not limited to:** * **Corrections or Disambiguations:** This might help the VA resolve an ambiguity or fix an error in the user's query.* *(e.g., User: "Call my brother," Third Party: "You mean your older brother, Mark, right?")* * **Cooperative Additions or Refinements:** This could give the VA extra specifics to better fulfill or understand the request.↩→ * *(e.g., User: "Add coffee to the shopping list," Third Party: "Get the decaf one.")** **Feasibility Constraints:** This could alert the VA to real-world conditions that may prevent or affect the request.* *(e.g., User: "Let's play music in the garden," Third Party: "The portable speaker's battery is dead.")** **Goal-oriented Suggestions:** This could give the VA an alternative that better achieves the user's intended outcome.* *(e.g., User: "How do I get to the airport?"Third Party: "The subway will be much faster than a taxi at this hour.")***2.Ignorable:** The interruption is irrelevant to complete and understand the user's ongoing request better.The VA should disregard it as it does not contribute to fulfilling the request.↩→ * **Definition:** The information is off-topic, a side comment, or directed at another human without impacting the VA's task.↩→ * **Example of an Ignorable interruption:** * *(e.g., User: "Set a timer for 10 minutes," Third Party: "I wonder what's for dinner tonight.")*### Conversation to Classify: **Primary User's Utterance:** {user_utterance} **Third-Party's Interruption:** {third_party_interference} ### Final Output Format(STRICT -MUST FOLLOW EXACTLY): **Classification:** [Your answer (NonIgnorable or Ignorable)] Eunwoo Song, Che Hyun Lee, Heeseung Kim, Sungroh Yoon |
ACL (1) | 4 |
| 2026 | Style-Friendly SNR Sampler for Style-Driven GenerationabstractRecent text-to-image diffusion models generate high-quality images but struggle to learn new styles, which limits the personalized content creation. In response, style-driven generation has become a popular task, wherein users supply reference images capturing the target style, complemented by text prompts that specify stylistic cues. Fine-tuning is a common approach, yet it often blindly utilizes pre-training configurations without modification, especially for noise schedules defined in terms of signal-to-noise ratio (SNR), which determines the amount of image information available at each denoising step. We discover that stylistic features predominantly emerge at low SNR range, leading current fine-tuning methods using regular noise schedules to exhibit suboptimal style alignment. We propose the Style-friendly SNR sampler, which focuses the fine-tuning on low SNR range where stylistic features emerge. We demonstrate improved generation of novel styles that cannot be described solely with a text prompt, enabling high-fidelity personalized content creation. Jooyoung Choi 0001, Chaehun Shin, Yeongtak Oh, Heeseung Kim, Jungbeom Lee, Sungroh Yoon |
WACV | 4 |
| 2025 | EdiText: Controllable Coarse-to-Fine Text Editing with Diffusion Language ModelsabstractWe propose EdiText, a controllable text editing method that modifies the reference text to desired attributes at various scales. We integrate an SDEdit-based editing technique that allows for broad adjustments in the degree of text editing. Additionally, we introduce a novel fine-level editing method based on self-conditioning, which allows subtle control of reference text. While being capable of editing on its own, this fine-grained method, integrated with the SDEdit approach, enables EdiText to make precise adjustments within the desired range. EdiText demonstrates its controllability to robustly adjust reference text at a broad range of levels across various tasks, including toxicity control and sentiment control. Che Hyun Lee, Heeseung Kim, Jiheum Yeom, Sungroh Yoon |
ACL (1) | 2 |
| 2025 | Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image GeneratorabstractSubject-driven text-to-image generation aims to produce images of a new subject within a desired context by accurately capturing both the visual characteristics of the subject and the semantic content of a text prompt. Traditional methods rely on time- and resource-intensive fine-tuning for subject alignment, while recent zero-shot approaches leverage on-the-fly image prompting, often sacrificing subject alignment. In this paper, we introduce Diptych Prompting, a novel zero-shot approach that reinterprets as an inpainting task with precise subject alignment by leveraging the emergent property of diptych generation in large-scale text-to-image models. Diptych Prompting arranges an incomplete diptych with the reference image in the left panel, and performs text-conditioned inpainting on the right panel. We further prevent unwanted content leakage by removing the background in the reference image and improve finegrained details in the generated subject by enhancing attention weights between the panels during inpainting. Experimental results confirm that our approach significantly outperforms zero-shot image prompting methods, resulting in images that are visually preferred by users. Additionally, our method supports not only subject-driven generation but also stylized image generation and subject-driven image editing, demonstrating versatility across diverse image generation applications. Chaehun Shin, Jooyoung Choi 0001, Heeseung Kim, Sungroh Yoon |
CVPR | 3 |
| 2025 | NanoVoice: Efficient Speaker-Adaptive Text-to-Speech for Multiple SpeakersabstractWe present NanoVoice, a personalized text-to-speech model that efficiently constructs voice adapters for multiple speakers simultaneously. NanoVoice introduces a batch-wise speaker adaptation technique capable of fine-tuning multiple references in parallel, significantly reducing training time. Beyond building separate adapters for each speaker, we also propose a parameter sharing technique that reduces the number of parameters used for speaker adaptation. By incorporating a novel trainable scale matrix, NanoVoice mitigates potential performance degradation during parameter sharing. NanoVoice achieves performance comparable to the baselines, while training 4 times faster and using 45 percent fewer parameters for speaker adaptation with 40 reference voices. Extensive ablation studies and analysis further validate the efficiency of our model. Nohil Park, Heeseung Kim, Che Hyun Lee, Jooyoung Choi 0001, Jiheum Yeom, Sungroh Yoon |
ICASSP | 2 |
| 2025 | VoiceGuider: Enhancing Out-of-Domain Performance in Parameter-Efficient Speaker-Adaptive Text-to-Speech via AutoguidanceabstractWhen applying parameter-efficient finetuning via LoRA onto speaker adaptive text-to-speech models, adaptation performance may decline compared to full-finetuned counterparts, especially for out-of-domain speakers. Here, we propose VoiceGuider, a parameter-efficient speaker adaptive text-to-speech system reinforced with autoguidance to enhance the speaker adaptation performance, reducing the gap against full-finetuned models. We carefully explore various ways of strengthening autoguidance, ultimately finding the optimal strategy. VoiceGuider as a result shows robust adaptation performance especially on extreme out-of-domain speech data. We provide audible samples in our demo page. Jiheum Yeom, Heeseung Kim, Jooyoung Choi 0001, Che Hyun Lee, Nohil Park, Sungroh Yoon |
ICASSP | 2 |
| 2024 | VoiceTailor: Lightweight Plug-In Adapter for Diffusion-Based Personalized Text-to-Speech
Heeseung Kim, Sang-gil Lee, Jiheum Yeom, Che Hyun Lee, Sungwon Kim 0001, Sungroh Yoon |
INTERSPEECH | 1 |
| 2024 | Paralinguistics-Aware Speech-Empowered Large Language Models for Natural ConversationabstractRecent work shows promising results in expanding the capabilities of large language models (LLM) to directly understand and synthesize speech. However, an LLM-based strategy for modeling spoken dialogs remains elusive, calling for further investigation. This paper introduces an extensive speech-text LLM framework, the Unified Spoken Dialog Model (USDM), designed to generate coherent spoken responses with naturally occurring prosodic features relevant to the given input speech without relying on explicit automatic speech recognition (ASR) or text-to-speech (TTS) systems. We have verified the inclusion of prosody in speech tokens that predominantly contain semantic information and have used this foundation to construct a prosody-infused speech-text model. Additionally, we propose a generalized speech-text pretraining scheme that enhances the capture of cross-modal semantics. To construct USDM, we fine-tune our speech-text model on spoken dialog data using a multi-step spoken dialog template that stimulates the chain-of-reasoning capabilities exhibited by the underlying LLM. Automatic and human evaluations on the DailyTalk dataset demonstrate that our approach effectively generates natural-sounding spoken responses, surpassing previous and cascaded baselines. Our code and checkpoints are available at https://github.com/naver-ai/usdm. Heeseung Kim, Soonshin Seo, Kyeongseok Jeong, Ohsung Kwon, Soyoon Kim, Jungwhan Kim, Jaehong Lee, Eunwoo Song, Myungwoo Oh, Sungroh Yoon, Kang Min Yoo |
NeurIPS | 1 |
| 2023 | Edit-A-Video: Single Video Editing with Object-Aware Consistency
Chaehun Shin, Heeseung Kim, Che Hyun Lee, Sang-gil Lee, Sungroh Yoon |
ACML | 2 |
| 2023 | UnitSpeech: Speaker-adaptive Speech Synthesis with Untranscribed Data
Heeseung Kim, Sungwon Kim 0001, Jiheum Yeom, Sungroh Yoon |
INTERSPEECH | 1 |
| 2022 | Rare Tokens Degenerate All Tokens: Improving Neural Text Generation via Adaptive Gradient Gating for Rare Token EmbeddingsabstractSangwon Yu, Jongyoon Song, Heeseung Kim, Seongmin Lee, Woo-Jong Ryu, Sungroh Yoon. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Sangwon Yu, Jongyoon Song, Heeseung Kim, Seongmin Lee 0005, Woo-Jong Ryu, Sungroh Yoon |
ACL (1) | 3 |
| 2022 | Stein Latent Optimization for Generative Adversarial Networks
Uiwon Hwang, Heeseung Kim, Dahuin Jung, Hyemi Jang, Hyungyu Lee, Sungroh Yoon |
ICLR | 2 |
| 2022 | PriorGrad: Improving Conditional Denoising Diffusion Models with Data-Dependent Adaptive Prior
Sang-gil Lee, Heeseung Kim, Chaehun Shin, Xu Tan 0003, Chang Liu 0030, Tao Qin 0001, Wei Chen 0034, Sungroh Yoon, Tie-Yan Liu |
ICLR | 2 |
| 2022 | Guided-TTS: A Diffusion Model for Text-to-Speech via Classifier GuidanceabstractWe propose Guided-TTS, a high-quality text-to-speech (TTS) model that does not require any transcript of target speaker using classifier guidance. Guided-TTS combines an unconditional diffusion probabilistic model with a separately trained phoneme classifier for classifier guidance. Our unconditional diffusion model learns to generate speech without any context from untranscribed speech data. For TTS synthesis, we guide the generative process of the diffusion model with a phoneme classifier trained on a large-scale speech recognition dataset. We present a norm-based scaling method that reduces the pronunciation errors of classifier guidance in Guided-TTS. We show that Guided-TTS achieves a performance comparable to that of the state-of-the-art TTS model, Grad-TTS, without any transcript for LJSpeech. We further demonstrate that Guided-TTS performs well on diverse datasets including a long-form untranscribed dataset. Heeseung Kim, Sungwon Kim 0001, Sungroh Yoon |
ICML | 1 |