VLDB 2026 Research / reviewers in the wild / expert
Yejin Jeon
dblp:255/7267
· DBLP profile ↗
10ranked-venue papers
3as first author
10since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Difficulty-Controllable Cloze Question Distractor GenerationabstractMultiple-choice cloze questions are commonly used to assess linguistic proficiency and comprehension.However, generating high-quality distractors remains challenging, as existing methods often lack adaptability and control over difficulty levels, and the absence of difficulty-annotated datasets further hinders progress.To address these issues, we propose a novel framework for generating distractors with controllable difficulty by leveraging both data augmentation and a multitask learning strategy.First, to create a high-quality, difficultyannotated dataset, we introduce a two-way distractor generation process to produce diverse and plausible distractors.These candidates are filtered and then categorized by difficulty using an ensemble QA system.Second, this newly created dataset is used to train a difficultycontrollable generation model via multitask learning.Experimental results demonstrate that our method generates high-quality distractors across difficulty levels and substantially outperforms GPT-4o in aligning distractor difficulty with human perception. Seokhoon Kang, Yejin Jeon, Seonjeong Hwang, Gary Geunbae Lee |
ACL (1) | 2 |
| 2025 | Retrieval-Augmented Fine-Tuning With Preference Optimization For Visual Program GenerationabstractDeokhyung Kang, Jeonghun Cho, Yejin Jeon, Sunbin Jang, Minsub Lee, Jawoon Cho, Gary Lee. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Deokhyung Kang, Jeonghun Cho 0002, Yejin Jeon, Sunbin Jang, Minsub Lee, Jawoon Cho, Gary Geunbae Lee |
ACL (1) | 3 |
| 2025 | MIRROR: Multimodal Cognitive Reframing Therapy for Rolling with ResistanceabstractRecent studies have explored the use of large language models (LLMs) in psychotherapy; however, text-based cognitive behavioral therapy (CBT) models often struggle with client resistance, which can weaken therapeutic alliance. To address this, we propose a multimodal approach that incorporates nonverbal cues, which allows the AI therapist to better align its responses with the client’s negative emotional state.Specifically, we introduce a new synthetic dataset, Mirror (Multimodal Interactive Rolling with Resistance), which is a novel synthetic dataset that pairs each client’s statements with corresponding facial images. Using this dataset, we train baseline vision language models (VLMs) so that they can analyze facial cues, infer emotions, and generate empathetic responses to effectively manage client resistance.These models are then evaluated in terms of both their counseling skills as a therapist, and the strength of therapeutic alliance in the presence of client resistance. Our results demonstrate that Mirror significantly enhances the AI therapist’s ability to handle resistance, which outperforms existing text-based CBT approaches.Human expert evaluations further confirm the effectiveness of our approach in managing client resistance and fostering therapeutic alliance. Hoonrae Kim, Yejin Jeon, Gary Geunbae Lee |
EMNLP | 4 |
| 2025 | PanicToCalm: A Proactive Counseling Agent for Panic AttacksabstractPanic attacks are acute episodes of fear and distress, in which timely, appropriate intervention can significantly help individuals regain stability.However, suitable datasets for training such models remain scarce due to ethical and logistical issues.To address this, we introduce PACE, which is a dataset that includes high-distress episodes constructed from firstperson narratives, and structured around the principles of Psychological First Aid (PFA).Using this data, we train PACER, a counseling model designed to provide both empathetic and directive support, which is optimized through supervised learning and simulated preference alignment.To assess its effectiveness, we propose PANICEVAL, a multi-dimensional framework covering general counseling quality and crisis-specific strategies.Experimental results show that PACER outperforms strong baselines in both counselor-side metrics and client affect improvement.Human evaluations further confirm its practical value, with PACER consistently preferred over general, CBT-based, and GPT-4-powered models in panic scenarios 1 . Yejin Min, Yejin Jeon, SungJun Yang, Hyounghun Kim, Gary Geunbae Lee |
EMNLP | 4 |
| 2025 | Facilitating Personalized TTS for Dysarthric Speakers Using Knowledge Anchoring and Curriculum LearningabstractDysarthric speakers experience substantial communication challenges due to impaired motor control of the speech apparatus, which leads to reduced speech intelligibility. This creates significant obstacles in dataset curation since actual recording of long, articulate sentences for the objective of training personalized TTS models becomes infeasible. Thus, the limited availability of audio data, in addition to the articulation errors that are present within the audio, complicates personalized speech synthesis for target dysarthric speaker adaptation. To address this, we frame the issue as a domain transfer task and introduce a knowledge anchoring framework that leverages a teacher-student model, enhanced by curriculum learning through audio augmentation. Experimental results show that the proposed zero-shot multi-speaker TTS model effectively generates synthetic speech with markedly reduced articulation errors and high speaker fidelity, while maintaining prosodic naturalness. Yejin Jeon, Solee Im, Gary Geunbae Lee |
INTERSPEECH | 1 |
| 2025 | PicPersona-TOD : A Dataset for Personalizing Utterance Style in Task-Oriented Dialogue with Image PersonaabstractJihyun Lee, Yejin Jeon, Seungyeon Seo, Gary Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yejin Jeon, Seungyeon Seo, Gary Geunbae Lee |
NAACL (Long Papers) | 2 |
| 2024 | Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker RepresentationsabstractZero-shot multi-speaker TTS aims to synthesize speech with the voice of a chosen target speaker without any fine-tuning. Prevailing methods, however, encounter limitations at adapting to new speakers of out-of-domain settings, primarily due to inadequate speaker disentanglement and content leakage. To overcome these constraints, we propose an innovative negation feature learning paradigm that models decoupled speaker attributes as deviations from the complete audio representation by utilizing the subtraction operation. By eliminating superfluous content information from the speaker representation, our negation scheme not only mitigates content leakage, thereby enhancing synthesis robustness, but also improves speaker fidelity. In addition, to facilitate the learning of diverse speaker attributes, we leverage multi-stream Transformers, which retain multiple hypotheses and instigate a training paradigm akin to ensemble learning. To unify these hypotheses and realize the final speaker representation, we employ attention pooling. Finally, in light of the imperative to generate target text utterances in the desired voice, we adopt adaptive layer normalizations to effectively fuse the previously generated speaker representation with the target text representations, as opposed to mere concatenation of the text and audio modalities. Extensive experiments and validations substantiate the efficacy of our proposed approach in preserving and harnessing speaker-specific attributes vis-à-vis alternative baseline models. Yejin Jeon, Yunsu Kim 0001, Gary Geunbae Lee |
AAAI | 1 |
| 2024 | Leveraging the Interplay between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause FormationabstractContemporary neural speech synthesis models have indeed demonstrated remarkable proficiency in synthetic speech generation as they have attained a level of quality comparable to that of human-produced speech. Nevertheless, it is important to note that these achievements have predominantly been verified within the context of high-resource languages such as English. Furthermore, the Tacotron and FastSpeech variants show substantial pausing errors when applied to the Korean language, which affects speech perception and naturalness. In order to address the aforementioned issues, we propose a novel framework that incorporates comprehensive modeling of both syntactic and acoustic cues that are associated with pausing patterns. Remarkably, our framework possesses the capability to consistently generate natural speech even for considerably more extended and intricate out-of-domain (OOD) sentences, despite its training on short audio clips. Architectural design choices are validated through comparisons with baseline models and ablation studies using subjective and objective metrics, thus confirming model performance. Yejin Jeon, Yunsu Kim 0001, Gary Geunbae Lee |
LREC/COLING | 1 |
| 2024 | An Investigation into Explainable Audio Hate Speech DetectionabstractResearch on hate speech has predominantly revolved around detection and interpretation from textual inputs, leaving verbal content largely unexplored.While there has been limited exploration into hate speech detection within verbal acoustic speech inputs, the aspect of interpretability has been overlooked.Therefore, we introduce a new task of explainable audio hate speech detection.Specifically, we aim to identify the precise time intervals, referred to as audio frame-level rationales, which serve as evidence for hate speech classification.Towards this end, we propose two different approaches: cascading and End-to-End (E2E).The cascading approach initially converts audio to transcripts, identifies hate speech within these transcripts, and subsequently locates the corresponding audio time frames.Conversely, the E2E approach processes audio utterances directly, which allows it to pinpoint hate speech within specific time frames.Additionally, due to the lack of explainable audio hate speech datasets that include audio frame-level rationales, we curated a synthetic audio dataset to train our models.We further validated these models on actual human speech utterances and found that the E2E approach outperforms the cascading method in terms of the audio frame Intersection over Union (IoU) metric.Furthermore, we observed that including frame-level rationales significantly enhances hate speech detection accuracy for the E2E approach. DisclaimerThe reader may encounter content of an offensive or hateful nature.However, given the nature of the work, this cannot be avoided. Jinmyeong An, Yejin Jeon, Jungseul Ok, Yunsu Kim 0001, Gary Geunbae Lee |
SIGDIAL | 3 |
| 2023 | Exploring the Viability of Synthetic Audio Data for Audio-Based Dialogue State TrackingabstractDialogue state tracking plays a crucial role in extracting information in task-oriented dialogue systems. However, preceding research are limited to textual modalities, primarily due to the shortage of authentic human audio datasets. We address this by investigating synthetic audio data for audio-based DST. To this end, we develop cascading and end-to-end models, train them with our synthetic audio dataset, and test them on actual human speech data. To facilitate evaluation tailored to audio modalities, we introduce a novel PhonemeF1 to capture pronunciation similarity. Experimental results showed that models trained solely on synthetic datasets can generalize their performance to human voice data. By eliminating the dependency on human speech data collection, these insights pave the way for significant practical advancements in audio-based DST. Data and code are available at https://github.com/JihyunLee1/E2E-DST.1 Yejin Jeon, Yunsu Kim 0001, Gary Geunbae Lee |
ASRU | 2 |