Yunsu Kim 0001

dblp:129/4079-1 · DBLP profile ↗
← Back
24ranked-venue papers
4as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 4 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Databases, data management, data science and information retrieval · 2
YearPublicationVenuePosition
2026 Adaptive Planning for Multi-Attribute Controllable Summarization with Monte Carlo Tree Search
abstract
Controllable summarization moves beyond generic outputs toward human-aligned summaries guided by specified attributes.In practice, the interdependence among attributes makes it challenging for language models to satisfy correlated constraints consistently.Moreover, previous approaches often require perattribute fine-tuning, limiting flexibility across diverse summary attributes.In this paper, we propose adaptive planning for multi-attribute controllable summarization (PACO), a trainingfree framework that reframes the task as planning the order of sequential attribute control with a customized Monte Carlo Tree Search (MCTS).In PACO, nodes represent summaries, and actions correspond to single-attribute adjustments, enabling progressive refinement of only the attributes requiring further control.This strategy adaptively discovers optimal control orders, ultimately producing summaries that effectively meet all constraints.Extensive experiments across diverse domains and models demonstrate that PACO achieves robust multi-attribute controllability, surpassing both LLM-based self-planning models and finetuned baselines.Remarkably, PACO with Llama-3.2-1Brivals the controllability of the much larger Llama-3.3-70Bbaselines.With larger models, PACO achieves superior control performance, outperforming all competitors.
Sangwon Ryu, Heejin Do, Yunsu Kim 0001, Gary Geunbae Lee, Jungseul Ok
ACL (1)3
2025 MiLQ: Benchmarking IR Models for Bilingual Web Search with Mixed Language Queries
abstract
Despite bilingual speakers frequently using mixed-language queries in web searches, Information Retrieval (IR) research on them remains scarce.To address this, we introduce MiLQ, Mixed-Language Query test set, the first public benchmark of mixed-language queries, qualified as realistic and relatively preferred.Experiments show that multilingual IR models perform moderately on MiLQ and inconsistently across native, English, and mixed-language queries, also suggesting code-switched training data's potential for robust IR models handling such queries.Meanwhile, intentional English mixing in queries proves an effective strategy for bilinguals searching English documents, which our analysis attributes to enhanced token matching compared to native queries. 1 * This work was done when the author was at aiXplain 1 The code and data for this work are available at : https://github.com/jonghwi-kim/milq.2 In this study, code-switching, mixed-language, and codemixing are used synonymously.Was sind die Vorteile und Nachteile einer einheitlichen europäischen Währung?Was sind die Advantages und Disadvantages einer single European Currency?What are the advantages and disadvantages of a single European currency?
Jonghwi Kim, Deokhyung Kang, Seonjeong Hwang, Yunsu Kim 0001, Jungseul Ok, Gary Geunbae Lee
EMNLP4
2025 Revisiting Early Detection of Sexual Predators via Turn-level Optimization
abstract
JinMyeong An, Sangwon Ryu, Heejin Do, Yunsu Kim, Jungseul Ok, Gary Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Jinmyeong An, Sangwon Ryu, Heejin Do, Yunsu Kim 0001, Jungseul Ok, Gary Geunbae Lee
NAACL (Long Papers)4
2025 DyPCL: Dynamic Phoneme-level Contrastive Learning for Dysarthric Speech Recognition
abstract
Wonjun Lee, Solee Im, Heejin Do, Yunsu Kim, Jungseul Ok, Gary Lee. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Solee Im, Heejin Do, Yunsu Kim 0001, Jungseul Ok, Gary Geunbae Lee
NAACL (Long Papers)4
2024 Enhancing Zero-Shot Multi-Speaker TTS with Negated Speaker Representations
abstract
Zero-shot multi-speaker TTS aims to synthesize speech with the voice of a chosen target speaker without any fine-tuning. Prevailing methods, however, encounter limitations at adapting to new speakers of out-of-domain settings, primarily due to inadequate speaker disentanglement and content leakage. To overcome these constraints, we propose an innovative negation feature learning paradigm that models decoupled speaker attributes as deviations from the complete audio representation by utilizing the subtraction operation. By eliminating superfluous content information from the speaker representation, our negation scheme not only mitigates content leakage, thereby enhancing synthesis robustness, but also improves speaker fidelity. In addition, to facilitate the learning of diverse speaker attributes, we leverage multi-stream Transformers, which retain multiple hypotheses and instigate a training paradigm akin to ensemble learning. To unify these hypotheses and realize the final speaker representation, we employ attention pooling. Finally, in light of the imperative to generate target text utterances in the desired voice, we adopt adaptive layer normalizations to effectively fuse the previously generated speaker representation with the target text representations, as opposed to mere concatenation of the text and audio modalities. Extensive experiments and validations substantiate the efficacy of our proposed approach in preserving and harnessing speaker-specific attributes vis-à-vis alternative baseline models.
Yejin Jeon, Yunsu Kim 0001, Gary Geunbae Lee
AAAI2
2024 Multi-Dimensional Optimization for Text Summarization via Reinforcement Learning
abstract
The evaluation of summary quality encompasses diverse dimensions such as consistency, coherence, relevance, and fluency.However, existing summarization methods often target a specific dimension, facing challenges in generating well-balanced summaries across multiple dimensions.In this paper, we propose multiobjective reinforcement learning tailored to generate balanced summaries across all four dimensions.We introduce two multi-dimensional optimization (MDO) strategies for adaptive learning: 1) MDO min , rewarding the current lowest dimension score, and 2) MDO pro , optimizing multiple dimensions similar to multi-task learning, resolves conflicting gradients across dimensions through gradient projection.Unlike prior ROUGE-based rewards relying on reference summaries, we use a QA-based reward model that aligns with human preferences.Further, we discover the capability to regulate the length of summaries by adjusting the discount factor, seeking the generation of concise yet informative summaries that encapsulate crucial points.Our approach achieved substantial performance gains compared to baseline models on representative summarization datasets, particularly in the overlooked dimensions.
Sangwon Ryu, Heejin Do, Yunsu Kim 0001, Gary Geunbae Lee, Jungseul Ok
ACL (1)3
2024 Explainable Multi-hop Question Generation: An End-to-End Approach without Intermediate Question Labeling
abstract
In response to the increasing use of interactive artificial intelligence, the demand for the capacity to handle complex questions has increased. Multi-hop question generation aims to generate complex questions that requires multi-step reasoning over several documents. Previous studies have predominantly utilized end-to-end models, wherein questions are decoded based on the representation of context documents. However, these approaches lack the ability to explain the reasoning process behind the generated multi-hop questions. Additionally, the question rewriting approach, which incrementally increases the question complexity, also has limitations due to the requirement of labeling data for intermediate-stage questions. In this paper, we introduce an end-to-end question rewriting model that increases question complexity through sequential rewriting. The proposed model has the advantage of training with only the final multi-hop questions, without intermediate questions. Experimental results demonstrate the effectiveness of our model in generating complex questions, particularly 3- and 4-hop questions, which are appropriately paired with input answers. We also prove that our model logically and incrementally increases the complexity of questions, and the generated multi-hop questions are also beneficial for training question answering models.
Seonjeong Hwang, Yunsu Kim 0001, Gary Geunbae Lee
LREC/COLING2
2024 Leveraging the Interplay between Syntactic and Acoustic Cues for Optimizing Korean TTS Pause Formation
abstract
Contemporary neural speech synthesis models have indeed demonstrated remarkable proficiency in synthetic speech generation as they have attained a level of quality comparable to that of human-produced speech. Nevertheless, it is important to note that these achievements have predominantly been verified within the context of high-resource languages such as English. Furthermore, the Tacotron and FastSpeech variants show substantial pausing errors when applied to the Korean language, which affects speech perception and naturalness. In order to address the aforementioned issues, we propose a novel framework that incorporates comprehensive modeling of both syntactic and acoustic cues that are associated with pausing patterns. Remarkably, our framework possesses the capability to consistently generate natural speech even for considerably more extended and intricate out-of-domain (OOD) sentences, despite its training on short audio clips. Architectural design choices are validated through comparisons with baseline models and ablation studies using subjective and objective metrics, thus confirming model performance.
Yejin Jeon, Yunsu Kim 0001, Gary Geunbae Lee
LREC/COLING2
2024 Denoising Table-Text Retrieval for Open-Domain Question Answering
abstract
In table-text open-domain question answering, a retriever system retrieves relevant evidence from tables and text to answer questions. Previous studies in table-text open-domain question answering have two common challenges: firstly, their retrievers can be affected by false-positive labels in training datasets; secondly, they may struggle to provide appropriate evidence for questions that require reasoning across the table. To address these issues, we propose Denoised Table-Text Retriever (DoTTeR). Our approach involves utilizing a denoised training dataset with fewer false positive labels by discarding instances with lower question-relevance scores measured through a false positive detection model. Subsequently, we integrate table-level ranking information into the retriever to assist in finding evidence for questions that demand reasoning across the table. To encode this ranking information, we fine-tune a rank-aware column encoder to identify minimum and maximum values within a column. Experimental results demonstrate that DoTTeR significantly outperforms strong baselines on both retrieval recall and downstream QA tasks. Our code is available at https://github.com/deokhk/DoTTeR.
Deokhyung Kang, Baikjin Jung, Yunsu Kim 0001, Gary Geunbae Lee
LREC/COLING3
2024 Cross-lingual Transfer for Automatic Question Generation by Learning Interrogative Structures in Target Languages
abstract
Automatic question generation (QG) serves a wide range of purposes, such as augmenting question-answering (QA) corpora, enhancing chatbot systems, and developing educational materials.Despite its importance, most existing datasets predominantly focus on English, resulting in a considerable gap in data availability for other languages.Cross-lingual transfer for QG (XLT-QG) addresses this limitation by allowing models trained on high-resource language datasets to generate questions in lowresource languages.In this paper, we propose a simple and efficient XLT-QG method that operates without the need for monolingual, parallel, or labeled data in the target language, utilizing a small language model.Our model, trained solely on English QA datasets, learns interrogative structures from a limited set of question exemplars, which are then applied to generate questions in the target language.Experimental results show that our method outperforms several XLT-QG baselines and achieves performance comparable to GPT-3.5-turbo across different languages.Additionally, the synthetic data generated by our model proves beneficial for training multilingual QA models.With significantly fewer parameters than large language models and without requiring additional training for target languages, our approach offers an effective solution for QG and QA tasks across various languages 1 .
Seonjeong Hwang, Yunsu Kim 0001, Gary Geunbae Lee
EMNLP2
2024 Cross-lingual Back-Parsing: Utterance Synthesis from Meaning Representation for Zero-Resource Semantic Parsing
abstract
Recent efforts have aimed to utilize multilingual pretrained language models (mPLMs) to extend semantic parsing (SP) across multiple languages without requiring extensive annotations.However, achieving zero-shot cross-lingual transfer for SP remains challenging, leading to a performance gap between source and target languages.In this study, we propose Cross-lingual Back-Parsing (CBP), a novel data augmentation methodology designed to enhance cross-lingual transfer for SP.Leveraging the representation geometry of the mPLMs, CBP synthesizes target language utterances from source meaning representations.Our methodology effectively performs cross-lingual data augmentation in challenging zero-resource settings, by utilizing only labeled data in the source language and monolingual corpora.Extensive experiments on two cross-lingual SP benchmarks (Mschema2QA and Xspider) demonstrate that CBP brings substantial gains in the target language.Further analysis of the synthesized utterances shows that our method successfully generates target language utterances with high slot value alignment rates while preserving semantic integrity.
Deokhyung Kang, Seonjeong Hwang, Yunsu Kim 0001, Gary Geunbae Lee
EMNLP3
2024 Key-Element-Informed sLLM Tuning for Document Summarization
abstract
Remarkable advances in large language models (LLMs) have enabled high-quality text summarization.However, this capability is currently accessible only through LLMs of substantial size or proprietary LLMs with usage fees.In response, smallerscale LLMs (sLLMs) of easy accessibility and low costs have been extensively studied, yet they often suffer from missing key information and entities, i.e., low relevance, in particular, when input documents are long.We hence propose a key-elementinformed instruction tuning for summarization, so-called KEIT-Sum, which identifies key elements in documents and instructs sLLM to generate summaries capturing these key elements.Experimental results on dialogue and news datasets demonstrate that sLLM with KEITSum indeed provides high-quality summarization with higher relevance and less hallucinations, competitive to proprietary LLM.
Sangwon Ryu, Heejin Do, Yunsu Kim 0001, Gary Geunbae Lee, Jungseul Ok
INTERSPEECH3
2024 An Investigation into Explainable Audio Hate Speech Detection
abstract
Research on hate speech has predominantly revolved around detection and interpretation from textual inputs, leaving verbal content largely unexplored.While there has been limited exploration into hate speech detection within verbal acoustic speech inputs, the aspect of interpretability has been overlooked.Therefore, we introduce a new task of explainable audio hate speech detection.Specifically, we aim to identify the precise time intervals, referred to as audio frame-level rationales, which serve as evidence for hate speech classification.Towards this end, we propose two different approaches: cascading and End-to-End (E2E).The cascading approach initially converts audio to transcripts, identifies hate speech within these transcripts, and subsequently locates the corresponding audio time frames.Conversely, the E2E approach processes audio utterances directly, which allows it to pinpoint hate speech within specific time frames.Additionally, due to the lack of explainable audio hate speech datasets that include audio frame-level rationales, we curated a synthetic audio dataset to train our models.We further validated these models on actual human speech utterances and found that the E2E approach outperforms the cascading method in terms of the audio frame Intersection over Union (IoU) metric.Furthermore, we observed that including frame-level rationales significantly enhances hate speech detection accuracy for the E2E approach. DisclaimerThe reader may encounter content of an offensive or hateful nature.However, given the nature of the work, this cannot be avoided.
Jinmyeong An, Yejin Jeon, Jungseul Ok, Yunsu Kim 0001, Gary Geunbae Lee
SIGDIAL5
2023 Exploring the Viability of Synthetic Audio Data for Audio-Based Dialogue State Tracking
abstract
Dialogue state tracking plays a crucial role in extracting information in task-oriented dialogue systems. However, preceding research are limited to textual modalities, primarily due to the shortage of authentic human audio datasets. We address this by investigating synthetic audio data for audio-based DST. To this end, we develop cascading and end-to-end models, train them with our synthetic audio dataset, and test them on actual human speech data. To facilitate evaluation tailored to audio modalities, we introduce a novel PhonemeF1 to capture pronunciation similarity. Experimental results showed that models trained solely on synthetic datasets can generalize their performance to human voice data. By eliminating the dependency on human speech data collection, these insights pave the way for significant practical advancements in audio-based DST. Data and code are available at https://github.com/JihyunLee1/E2E-DST.1
Yejin Jeon, Yunsu Kim 0001, Gary Geunbae Lee
ASRU4
2023 Optimizing Two-Pass Cross-Lingual Transfer Learning: Phoneme Recognition And Phoneme To Grapheme Translation
abstract
This research optimizes two-pass cross-lingual transfer learning in low-resource languages by enhancing phoneme recognition and phoneme-to-grapheme translation models. Our approach optimizes these two stages to improve speech recognition across languages. We optimize phoneme vocabulary coverage by merging phonemes based on shared articulatory characteristics, thus improving recognition accuracy. Additionally, we introduce a global phoneme noise generator for realistic ASR noise during phoneme-to-grapheme training to reduce error propagation. Experiments on the CommonVoice 12.0 dataset show significant reductions in Word Error Rate (WER) for low-resource languages, highlighting the effectiveness of our approach. This research contributes to the advancements of two-pass ASR systems in low-resource languages, offering the potential for improved cross-lingual transfer learning.
Gary Geunbae Lee, Yunsu Kim 0001
ASRU3
2023 Hierarchical Pronunciation Assessment with Multi-Aspect Attention
abstract
Automatic pronunciation assessment is a major component of a computer-assisted pronunciation training system. To provide in-depth feedback, scoring pronunciation at various levels of granularity such as phoneme, word, and utterance, with diverse aspects such as accuracy, fluency, and completeness, is essential. However, existing multi-aspect multi-granularity methods simultaneously predict all aspects at all granularity levels; therefore, they have difficulty in capturing the linguistic hierarchy of phoneme, word, and utterance. This limitation further leads to neglecting intimate cross-aspect relations at the same linguistic unit. In this paper, we propose a Hierarchical Pronunciation Assessment with Multi-aspect Attention (HiPAMA) model, which hierarchically represents the granularity levels to directly capture their linguistic structures and introduces multi-aspect attention that reflects associations across aspects at the same level to create more connotative representations. By obtaining relational information from both the granularity- and aspect-side, HiPAMA can take full advantage of multi-task learning. Remarkable improvements in the experimental results on the speachocean762 datasets demonstrate the robustness of HiPAMA, particularly in the difficult-to-assess aspects.
Heejin Do, Yunsu Kim 0001, Gary Geunbae Lee
ICASSP2
2023 Score-balanced Loss for Multi-aspect Pronunciation Assessment
abstract
With rapid technological growth, automatic pronunciation assessment has transitioned toward systems that evaluate pronunciation in various aspects, such as fluency and stress.However, despite the highly imbalanced score labels within each aspect, existing studies have rarely tackled the data imbalance problem.In this paper, we suggest a novel loss function, score-balanced loss, to address the problem caused by uneven data, such as bias toward the majority scores.As a re-weighting approach, we assign higher costs when the predicted score is of the minority class, thus, guiding the model to gain positive feedback for sparse score prediction.Specifically, we design two weighting factors by leveraging the concept of an effective number of samples and using the ranks of scores.We evaluate our method on the speechocean762 dataset, which has noticeably imbalanced scores for several aspects.Improved results particularly on such uneven aspects prove the effectiveness of our method.
Heejin Do, Yunsu Kim 0001, Gary Geunbae Lee
INTERSPEECH2
2023 Tracking Must Go On : Dialogue State Tracking with Verified Self-Training
abstract
In task-oriented dialogues, dialogue state tracking (DST) is a critical component as it identifies specific information for the user's purpose.However, as annotating DST data requires a significant amount of human effort, leveraging raw dialogue is crucial.To address this, we propose a new self-training (ST) framework with a verification model.Unlike previous ST methods that rely on extensive hyper-parameter searching to filter out inaccurate data, our verification methodology ensures the accuracy and validity of the dataset without using a fixed threshold.Furthermore, to mitigate overfitting, we augment the dataset by generating diverse user utterances.Even when using only 10% of the labeled data, our approach achieves comparable results to a fully labeled MultiWOZ2.0dataset.The evaluation of scalability also demonstrates enhanced robustness in predicting unseen values.
Chaebin Lee, Yunsu Kim 0001, Gary Geunbae Lee
INTERSPEECH3
2020 When and Why is Unsupervised Neural Machine Translation Useless?
abstract
This paper studies the practicality of the current state-of-the-art unsupervised methods in neural machine translation (NMT). In ten translation tasks with various data settings, we analyze the conditions under which the unsupervised methods fail to produce reasonable translations. We show that their performance is severely affected by linguistic dissimilarity and domain mismatch between source and target monolingual data. Such conditions are common for low-resource language pairs, where unsupervised learning works poorly. In all of our experiments, supervised and semi-supervised baselines with 50k-sentence bilingual data outperform the best unsupervised results. Our analyses pinpoint the limits of the current unsupervised NMT and also suggest immediate research directions.
Yunsu Kim 0001, Miguel Graça, Hermann Ney
EAMT1
2019 Effective Cross-lingual Transfer of Neural Machine Translation Models without Shared Vocabularies
abstract
Transfer learning or multilingual model is essential for low-resource neural machine translation (NMT), but the applicability is limited to cognate languages by sharing their vocabularies.This paper shows effective techniques to transfer a pre-trained NMT model to a new, unrelated language without shared vocabularies.We relieve the vocabulary mismatch by using cross-lingual word embedding, train a more language-agnostic encoder by injecting artificial noises, and generate synthetic data easily from the pre-training data without back-translation.Our methods do not require restructuring the vocabulary or retraining the model.We improve plain NMT transfer by up to +5.1% BLEU in five low-resource translation tasks, outperforming multilingual joint training by a large margin.We also provide extensive ablation studies on pre-trained embedding, synthetic data, vocabulary size, and parameter freezing for a better understanding of NMT transfer.
Yunsu Kim 0001, Yingbo Gao, Hermann Ney
ACL (1)1
2019 Pivot-based Transfer Learning for Neural Machine Translation between Non-English Languages
abstract
Yunsu Kim, Petre Petrov, Pavel Petrushkov, Shahram Khadivi, Hermann Ney. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Yunsu Kim 0001, Petre Petrov, Pavel Petrushkov, Shahram Khadivi, Hermann Ney
EMNLP/IJCNLP (1)1
2018 Improving Unsupervised Word-by-Word Translation with Language Model and Denoising Autoencoder
abstract
Unsupervised learning of cross-lingual word embedding offers elegant matching of words across languages, but has fundamental limitations in translating sentences.In this paper, we propose simple yet effective methods to improve word-by-word translation of crosslingual embeddings, using only monolingual corpora but without any back-translation.We integrate a language model for context-aware search, and use a novel denoising autoencoder to handle reordering.Our system surpasses state-of-the-art unsupervised neural translation systems without costly iterative training.We also analyze the effect of vocabulary size and denoising type on the translation performance, which provides better understanding of learning the cross-lingual word embedding and its usage in translation.
Yunsu Kim 0001, Jiahui Geng, Hermann Ney
EMNLP1
2015 Subspace clustering of data streams: new algorithms and effective evaluation measures
Marwan Hassani, Yunsu Kim 0001, Seungjin Choi 0001, Thomas Seidl 0001
J. Intell. Inf. Syst.2
2013 Subspace MOA: Subspace Stream Clustering Evaluation Using the MOA Framework
Marwan Hassani, Yunsu Kim 0001, Thomas Seidl 0001
DASFAA (2)2