VLDB 2026 Research / reviewers in the wild / expert
Jinyoung Yeo
dblp:121/4335
· DBLP profile ↗
41ranked-venue papers
9as first author
26since 2021 · last 2026
0000-0003-3847-4917ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 5 first-author · 23 since 2021Databases, data management, data science and information retrieval · 13 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MVIGER: Multi-View Variational Integration of Complementary Knowledge for Generative RecommenderabstractLanguage Models (LMs) have been widely used in recommender systems to incorporate textual information of items into item IDs, leveraging their advanced language understanding and generation capabilities. Recently, generative recommender systems have utilized the reasoning abilities of LMs to directly generate index tokens for potential items of interest based on the user's interaction history. To inject diverse item knowledge into LMs, prompt templates with detailed task descriptions and various indexing techniques derived from diverse item information have been explored. This paper focuses on the inconsistency in outputs generated by variations in input prompt templates and item index types, even with the same user's interaction history. Our in-depth quantitative analysis reveals that preference knowledge learned from diverse prompt templates and heterogeneous indices differs significantly, indicating a high potential for complementarity. To fully exploit this complementarity and provide consistent performance under varying prompts and item indices, we propose MVIGER, a unified variational framework that models selection among these information sources as a categorical latent variable with a learnable prior. During inference, this prior enables the model to adaptively select the most relevant source or aggregate predictions across multiple sources, thereby ensuring high-quality recommendation across diverse template-index combinations. We validate the effectiveness of MVIGER on three real-world datasets, demonstrating its superior performance over existing generative recommender baselines through the effective integration of complementary knowledge. Tongyoung Kim, Soojin Yoon 0001, Seongku Kang, Jinyoung Yeo, Dongha Lee 0003 |
SIGIR | 4 |
| 2026 | AgenticShop: Benchmarking Agentic Product Curation for Personalized Web ShoppingabstractThe proliferation of e-commerce has made web shopping platforms key gateways for customers navigating the vast digital marketplace. Yet this rapid expansion has led to a noisy and fragmented information environment, increasing cognitive burden as shoppers explore and purchase products online. With promising potential to alleviate this challenge, agentic systems have garnered growing attention for automating user-side tasks in web shopping. Despite significant advancements, existing benchmarks fail to comprehensively evaluate how well agentic systems can curate products in open-web settings. Specifically, they have limited coverage of shopping scenarios, focusing only on simplified single-platform lookups rather than exploratory search. Moreover, they overlook personalization in evaluation, leaving unclear whether agents can adapt to diverse user preferences in realistic shopping contexts. To address this gap, we present AgenticShop, the first benchmark for evaluating agentic systems on personalized product curation in open-web environment. Crucially, our approach features realistic shopping scenarios, diverse user profiles, and a verifiable, checklist-driven personalization evaluation framework. Through extensive experiments, we demonstrate that current agentic systems remain largely insufficient, emphasizing the need for user-side systems that effectively curate tailored products across the modern web. Sunghwan Kim 0005, Ryang Heo, Yongsik Seo, Jinyoung Yeo, Dongha Lee 0003 |
WWW | 4 |
| 2025 | Rethinking Reward Model Evaluation Through the Lens of Reward OveroptimizationabstractReward models (RMs) play a crucial role in reinforcement learning from human feedback (RLHF), aligning model behavior with human preferences. However, existing benchmarks for reward models show a weak correlation with the performance of optimized policies, suggesting that they fail to accurately assess the true capabilities of RMs. To bridge this gap, we explore several evaluation designs through the lens of reward overoptimization, i.e., a phenomenon that captures both how well the reward model aligns with human preferences and the dynamics of the learning signal it provides to the policy. The results highlight three key findings on how to construct a reliable benchmark: (i) it is important to minimize differences between chosen and rejected responses beyond correctness, (ii) evaluating reward models requires multiple comparisons across a wide range of chosen and rejected responses, and (iii) given that reward models encounter responses with diverse representations, responses should be sourced from a variety of models. However, we also observe that a extremely high correlation with degree of overoptimization leads to comparatively lower correlation with certain downstream performance. Thus, when designing a benchmark, it is desirable to use the degree of overoptimization as a useful tool, rather than the end goal. Sunghwan Kim 0005, Dongjin Kang, Taeyoon Kwon, Hyungjoo Chae, Dongha Lee 0003, Jinyoung Yeo |
ACL (1) | 6 |
| 2025 | LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical StudyabstractThe remarkable reasoning and generalization capabilities of Large Language Models (LLMs) have paved the way for their expanding applications in embodied AI, robotics, and other real-world tasks. To effectively support these applications, grounding in spatial and temporal understanding in multimodal environments is essential. To this end, recent works have leveraged scene graphs, a structured representation that encodes entities, attributes, and their relationships in a scene. However, a comprehensive evaluation of LLMs’ ability to utilize scene graphs remains limited. In this work, we introduce Text-Scene Graph (TSG) Bench, a benchmark designed to systematically assess LLMs’ ability to (1) understand scene graphs and (2) generate them from textual narratives. With TSG Bench we evaluate 11 LLMs and reveal that, while models perform well on scene graph understanding, they struggle with scene graph generation, particularly for complex narratives. Our analysis indicates that these models fail to effectively decompose discrete scenes from a complex narrative, leading to a bottleneck when generating scene graphs. These findings underscore the need for improved methodologies in scene graph generation and provide valuable insights for future research. The demonstration of our benchmark is available at https://tsg-bench.netlify.app. Additionally, our code and evaluation data are publicly available at https://github.com/docworlds/tsg-bench. Dongil Yang, Minjin Kim, Sunghwan Kim 0005, Beong-woo Kwak, Minjun Park, Jinseok Hong, Woontack Woo, Jinyoung Yeo |
ACL (1) | 8 |
| 2025 | Web Agents with World Models: Learning and Leveraging Environment Dynamics in Web NavigationabstractLarge language models (LLMs) have recently gained much attention in building autonomous agents. However, performance of current LLM-based web agents in long-horizon tasks is far from optimal, often yielding errors such as repeatedly buying a non-refundable flight ticket. By contrast, humans can avoid such an irreversible mistake, as we have an awareness of the potential outcomes (e.g., losing money) of our actions, also known as the "world model". Motivated by this, our study first starts with preliminary analyses, confirming the absence of world models in current LLMs (e.g., GPT-4o, Claude-3.5-Sonnet, etc.). Then, we present a World-model-augmented (WMA) web agent, which simulates the outcomes of its actions for better decision-making. To overcome the challenges in training LLMs as world models predicting next observations, such as repeated elements across observations and long HTML inputs, we propose a transition-focused observation abstraction, where the prediction objectives are free-form natural language descriptions exclusively highlighting important state differences between time steps. Experiments on WebArena and Mind2Web show that our world models improve agents' policy selection without training and demonstrate our agents' cost- and time-efficiency compared to recent tree-search-based agents. Hyungjoo Chae, Namyoung Kim, Kai Tzu-iunn Ong, Minju Gwak, Gwanwoo Song, Sunghwan Kim 0005, Dongha Lee 0003, Jinyoung Yeo |
ICLR | 9 |
| 2025 | Towards Lifelong Dialogue Agents via Timeline-based Memory ManagementabstractKai Tzu-iunn Ong, Namyoung Kim, Minju Gwak, Hyungjoo Chae, Taeyoon Kwon, Yohan Jo, Seung-won Hwang, Dongha Lee, Jinyoung Yeo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kai Tzu-iunn Ong, Namyoung Kim, Minju Gwak, Hyungjoo Chae, Taeyoon Kwon, Yohan Jo, Seung-won Hwang, Dongha Lee 0003, Jinyoung Yeo |
NAACL (Long Papers) | 9 |
| 2025 | Web-Shepherd: Advancing PRMs for Reinforcing Web AgentsabstractWeb navigation is a unique domain that can automate many repetitive real-life tasks and is challenging as it requires long-horizon sequential decision making beyond typical multimodal large language model (MLLM) tasks.
Yet, specialized reward models for web navigation that can be utilized during both training and test-time have been absent until now. Despite the importance of speed and cost-effectiveness, prior works have utilized MLLMs as reward models, which poses significant constraints for real-world deployment. To address this, in this work, we propose the first process reward model (PRM) called Web-Shepherd which could assess web navigation trajectories in a step-level. To achieve this, we first construct the WebPRM Collection, a large-scale dataset with 40K step-level preference pairs and annotated checklists spanning diverse domains and difficulty levels. Next, we also introduce the WebRewardBench, the first meta-evaluation benchmark for evaluating PRMs. In our experiments, we observe that our Web-Shepherd achieves about 30 points better accuracy compared to using GPT-4o on WebRewardBench.
Furthermore, when testing on WebArena-lite by using GPT-4o-mini as the policy and Web-Shepherd as the verifier, we achieve 10.9 points better performance, in 10x less cost compared to using GPT-4o-mini as the verifier.
Our model, dataset, and code are publicly available at https://github.com/kyle8581/Web-Shepherd. Hyungjoo Chae, Seonghwan Kim, Seungone Kim, Seungjun Moon, Gyeom Hwangbo, Dongha Lim, Minjin Kim, Yeonjun Hwang, Minju Gwak, Dongwook Choi, Gwanhoon Im, ByeongUng Cho, Hyojun Kim, Jun Hee Han, Taeyoon Kwon, Beong-woo Kwak, Dongjin Kang, Jinyoung Yeo |
NeurIPS | 21 |
| 2025 | Fast and Fluent Diffusion Language Models via Convolutional Decoding and Rejective Fine-tuningabstractAutoregressive (AR) language models generate text one token at a time, which limits their inference speed. Diffusion-based language models offer a promising alternative, as they can decode multiple tokens in parallel. However, we identify a key bottleneck in current diffusion LMs: the \textbf{long decoding-window problem}, where tokens generated far from the input context often become irrelevant or repetitive. Previous solutions like semi-autoregressive address this issue by splitting windows into blocks (sacrificing bidirectionality), but we find that this also leads to \textbf{time-interval expansion problem}, sacrificing the speed. Therefore, semi-AR eliminates the main advantages of diffusion models. To overcome this, we propose Convolutional decoding (\textit{Conv}), a normalization-based method that narrows the decoding window without hard segmentation, leading to better fluency and flexibility. Additionally, we introduce Rejecting Rule-based Fine-Tuning (R2FT), a post-hoc training scheme that better aligns tokens at positions far from context. Our methods achieve state-of-the-art results on open-ended generation benchmarks (e.g., AlpacaEval) among diffusion LM baselines, with significantly lower step size than previous works, demonstrating both speed and quality improvements. The code is available online (\url{https://github.com/ybseo-ac/Conv}). Yeongbin Seo, Dongha Lee 0003, Jinyoung Yeo |
NeurIPS | 4 |
| 2025 | Review-driven Personalized Preference Reasoning with Large Language Models for RecommendationabstractRecent advancements in Large Language Models (LLMs) have demonstrated exceptional performance across a wide range of tasks, generating significant interest in their application to recommendation systems. However, existing methods have not fully harnessed the potential of LLMs, often constrained by limited input information or failing to fully utilize their advanced reasoning capabilities. To address these limitations, we introduce EXP3RT, a novel LLM-based recommender designed to leverage rich preference information contained in user and item reviews. EXP3RT is basically fine-tuned through distillation from a teacher LLM to perform three key steps in order: (1) preference extraction (2) profile construction, and (3) textual reasoning for rating prediction. EXP3RT first extracts and encapsulates essential subjective preferences from raw reviews, next aggregates and summarizes them according to specific criteria to create user and item profiles. It then generates detailed step-by-step reasoning followed by predicted rating, i.e., reasoning-enhanced rating prediction, by considering both subjective and objective information from user/item profiles and item descriptions. This personalized preference reasoning from EXP3RT enhances rating prediction accuracy and also provides faithful and reasonable explanations for recommendation. Extensive experiments show that EXP3RT outperforms existing methods on both rating prediction and candidate item reranking for top-k recommendation, while significantly enhancing the explainability of recommendation systems. Jieyong Kim, Hyunseo Kim 0002, Seongku Kang, Buru Chang, Jinyoung Yeo, Dongha Lee 0003 |
SIGIR | 6 |
| 2025 | Unsupervised Robust Cross-Lingual Entity Alignment via Neighbor Triple Matching with Entity and Relation Texts
Soojin Yoon 0001, Sungho Ko, Tongyoung Kim, Seongku Kang, Jinyoung Yeo, Dongha Lee 0003 |
WSDM | 5 |
| 2024 | Large Language Models Are Clinical Reasoners: Reasoning-Aware Diagnosis Framework with Prompt-Generated RationalesabstractMachine reasoning has made great progress in recent years owing to large language models (LLMs). In the clinical domain, however, most NLP-driven projects mainly focus on clinical classification or reading comprehension, and under-explore clinical reasoning for disease diagnosis due to the expensive rationale annotation with clinicians. In this work, we present a "reasoning-aware" diagnosis framework that rationalizes the diagnostic process via prompt-based learning in a time- and labor-efficient manner, and learns to reason over the prompt-generated rationales. Specifically, we address the clinical reasoning for disease diagnosis, where the LLM generates diagnostic rationales providing its insight on presented patient data and the reasoning path towards the diagnosis, namely Clinical Chain-of-Thought (Clinical CoT). We empirically demonstrate LLMs/LMs' ability of clinical reasoning via extensive experiments and analyses on both rationale generation and disease diagnosis in various settings. We further propose a novel set of criteria for evaluating machine-generated rationales' potential for real-world clinical settings, facilitating and benefiting future research in this area. Taeyoon Kwon, Kai Tzu-iunn Ong, Dongjin Kang, Seungjun Moon, Jeong Ryong Lee, Dosik Hwang, Beomseok Sohn, Yongsik Sim, Dongha Lee 0003, Jinyoung Yeo |
AAAI | 10 |
| 2024 | Can Large Language Models be Good Emotional Supporter? Mitigating Preference Bias on Emotional Support ConversationabstractDongjin Kang, Sunghwan Kim, Taeyoon Kwon, Seungjun Moon, Hyunsouk Cho, Youngjae Yu, Dongha Lee, Jinyoung Yeo. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Dongjin Kang, Sunghwan Kim 0005, Taeyoon Kwon, Seungjun Moon, Hyunsouk Cho, Youngjae Yu, Dongha Lee 0003, Jinyoung Yeo |
ACL (1) | 8 |
| 2024 | VerifiNER: Verification-augmented NER via Knowledge-grounded Reasoning with Large Language ModelsabstractRecent approaches in domain-specific named entity recognition (NER), such as biomedical NER, have shown remarkable advances.However, they still lack of faithfulness, producing erroneous predictions.We assume that knowledge of entities can be useful in verifying the correctness of the predictions.Despite the usefulness of knowledge, resolving such errors with knowledge is nontrivial, since the knowledge itself does not directly indicate the groundtruth label.To this end, we propose VER-IFINER, a post-hoc verification framework that identifies errors from existing NER methods using knowledge and revises them into more faithful predictions.Our framework leverages the reasoning abilities of large language models to adequately ground on knowledge and the contextual information in the verification process.We validate effectiveness of VERIFINER through extensive experiments on biomedical datasets.The results suggest that VERIFINER can successfully verify errors from existing models as a model-agnostic approach.Further analyses on out-of-domain and low-resource settings show the usefulness of VERIFINER on real-world applications.1 Kwangwook Seo, Hyungjoo Chae, Jinyoung Yeo, Dongha Lee 0003 |
ACL (1) | 4 |
| 2024 | Language Models as Compilers: Simulating Pseudocode Execution Improves Algorithmic Reasoning in Language ModelsabstractHyungjoo Chae, Yeonghyeon Kim, Seungone Kim, Kai Tzu-iunn Ong, Beong-woo Kwak, Moohyeon Kim, Sunghwan Kim, Taeyoon Kwon, Jiwan Chung, Youngjae Yu, Jinyoung Yeo. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Hyungjoo Chae, Yeonghyeon Kim, Seungone Kim, Kai Tzu-iunn Ong, Beong-woo Kwak, Moohyeon Kim, Sunghwan Kim 0005, Taeyoon Kwon, Jiwan Chung, Youngjae Yu, Jinyoung Yeo |
EMNLP | 11 |
| 2024 | Coffee-Gym: An Environment for Evaluating and Improving Natural Language Feedback on Erroneous CodeabstractHyungjoo Chae, Taeyoon Kwon, Seungjun Moon, Yongho Song, Dongjin Kang, Kai Tzu-iunn Ong, Beong-woo Kwak, Seonghyeon Bae, Seung-won Hwang, Jinyoung Yeo. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Hyungjoo Chae, Taeyoon Kwon, Seungjun Moon, Yongho Song, Dongjin Kang, Kai Tzu-iunn Ong, Beong-woo Kwak, Seonghyeon Bae, Seung-won Hwang, Jinyoung Yeo |
EMNLP | 10 |
| 2024 | Evidence-Focused Fact Summarization for Knowledge-Augmented Zero-Shot Question AnsweringabstractRecent studies have investigated utilizing Knowledge Graphs (KGs) to enhance Question Answering (QA) performance of Large Language Models (LLMs), yet structured KG verbalization remains challenging.Existing methods, such as triple-form or free-form textual conversion of triple-form facts, encounter several issues.These include reduced evidence density due to duplicated entities or relationships, and reduced evidence clarity due to an inability to emphasize crucial evidence.To address these issues, we propose EFSUM, an Evidence-focused Fact Summarization framework for enhanced QA with knowledge-augmented LLMs.We optimize an open-source LLM as a fact summarizer through distillation and preference alignment.Our extensive experiments show that EFSUM improves LLM's zero-shot QA performance, and it is possible to ensure both the helpfulness and faithfulness of the summary. Sungho Ko, Hyungjoo Chae, Jinyoung Yeo, Dongha Lee 0003 |
EMNLP | 4 |
| 2024 | Train-Attention: Meta-Learning Where to Focus in Continual Knowledge LearningabstractPrevious studies on continual knowledge learning (CKL) in large language models (LLMs) have predominantly focused on approaches such as regularization, architectural modifications, and rehearsal techniques to mitigate catastrophic forgetting. However, these methods naively inherit the inefficiencies of standard training procedures, indiscriminately applying uniform weight across all tokens, which can lead to unnecessary parameter updates and increased forgetting. To address these shortcomings, we propose a novel CKL approach termed Train-Attention-Augmented Language Model (TAALM), which enhances learning efficiency by dynamically predicting and applying weights to tokens based on their usefulness. This method employs a meta-learning framework that optimizes token importance predictions, facilitating targeted knowledge updates and minimizing forgetting. Also, we observe that existing benchmarks do not clearly exhibit the trade-off between learning and retaining, therefore we propose a new benchmark, LAMA-ckl, to address this issue. Through experiments conducted on both newly introduced and established CKL benchmarks, TAALM proves the state-of-the-art performance upon the baselines, and also shows synergistic compatibility when integrated with previous CKL approaches. The code and the dataset are available online. Yeongbin Seo, Dongha Lee 0003, Jinyoung Yeo |
NeurIPS | 3 |
| 2024 | KULTURE Bench: A Benchmark for Assessing Language Model in Korean Cultural Context
Jinyoung Yeo, Joon-Ho Lim, Hansaem Kim |
PACLIC | 2 |
| 2023 | TUTORING: Instruction-Grounded Conversational Agent for Language LearnersabstractIn this paper, we propose Tutoring bot, a generative chatbot trained on a large scale of tutor-student conversations for English-language learning. To mimic a human tutor's behavior in language education, the tutor bot leverages diverse educational instructions and grounds to each instruction as additional input context for the tutor response generation. As a single instruction generally involves multiple dialogue turns to give the student sufficient speaking practice, the tutor bot is required to monitor and capture when the current instruction should be kept or switched to the next instruction. For that, the tutor bot is learned to not only generate responses but also infer its teaching action and progress on the current conversation simultaneously by a multi-task learning scheme. Our Tutoring bot is deployed under a non-commercial use license at https://tutoringai.com. Hyungjoo Chae, Minjin Kim, Chaehyeong Kim, Wonseok Jeong, Hyejoong Kim, Junmyung Lee, Jinyoung Yeo |
AAAI | 7 |
| 2023 | Dialogue Chain-of-Thought Distillation for Commonsense-aware Conversational AgentsabstractHyungjoo Chae, Yongho Song, Kai Ong, Taeyoon Kwon, Minjin Kim, Youngjae Yu, Dongha Lee, Dongyeop Kang, Jinyoung Yeo. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Hyungjoo Chae, Yongho Song, Kai Tzu-iunn Ong, Taeyoon Kwon, Minjin Kim, Youngjae Yu, Dongha Lee 0003, Dongyeop Kang, Jinyoung Yeo |
EMNLP | 9 |
| 2022 | Dual Task Framework for Improving Persona-Grounded Dialogue DatasetabstractThis paper introduces a simple yet effective data-centric approach for the task of improving persona-conditioned dialogue agents. Prior model-centric approaches unquestioningly depend on the raw crowdsourced benchmark datasets such as Persona-Chat. In contrast, we aim to fix annotation artifacts in benchmarking, which is orthogonally applicable to any dialogue model. Specifically, we augment relevant personas to improve dialogue dataset/agent, by leveraging the primal-dual structure of the two tasks, predicting dialogue responses and personas based on each other. Experiments on Persona-Chat show that our approach outperforms pre-trained LMs by an 11.7 point gain in terms of accuracy. Beong-woo Kwak, Youngwook Kim 0002, Hong-in Lee, Seung-won Hwang, Jinyoung Yeo |
AAAI | 6 |
| 2022 | TrustAL: Trustworthy Active Learning Using Knowledge DistillationabstractActive learning can be defined as iterations of data labeling, model training, and data acquisition, until sufficient labels are acquired. A traditional view of data acquisition is that, through iterations, knowledge from human labels and models is implicitly distilled to monotonically increase the accuracy and label consistency. Under this assumption, the most recently trained model is a good surrogate for the current labeled data, from which data acquisition is requested based on uncertainty/diversity. Our contribution is debunking this myth and proposing a new objective for distillation. First, we found example forgetting, which indicates the loss of knowledge learned across iterations. Second, for this reason, the last model is no longer the best teacher-- For mitigating such forgotten knowledge, we select one of its predecessor models as a teacher, by our proposed notion of "consistency". We show that this novel distillation is distinctive in the following three aspects; First, consistency ensures to avoid forgetting labels. Second, consistency improves both uncertainty/diversity of labeled data. Lastly, consistency redeems defective labels produced by human annotators. Beong-woo Kwak, Youngwook Kim 0002, Yu Jin Kim, Seung-won Hwang, Jinyoung Yeo |
AAAI | 5 |
| 2022 | Mind the Gap! Injecting Commonsense Knowledge for Abstractive Dialogue SummarizationabstractIn this paper, we propose to leverage the unique characteristics of dialogues sharing commonsense knowledge across participants, to resolve the difficulties in summarizing them. We present SICK, a framework that uses commonsense inferences as additional context. Compared to previous work that solely relies on the input dialogue, SICK uses an external knowledge model to generate a rich set of commonsense inferences and selects the most probable one with a similarity-based selection method. Built upon SICK, SICK++ utilizes commonsense as supervision, where the task of generating commonsense inferences is added upon summarizing the dialogue in a multi-task learning setting. Experimental results show that with injected commonsense knowledge, our framework generates more informative and consistent summaries than existing methods. Seungone Kim, Se June Joo, Hyungjoo Chae, Chaehyeong Kim, Seung-won Hwang, Jinyoung Yeo |
COLING | 6 |
| 2022 | BotsTalk: Machine-sourced Framework for Automatic Curation of Large-scale Multi-skill Dialogue DatasetsabstractTo build open-domain chatbots that are able to use diverse communicative skills, we propose a novel framework BOTSTALK, where multiple agents grounded to the specific target skills participate in a conversation to automatically annotate multi-skill dialogues.We further present Blended Skill BotsTalk (BSBT), a large-scale multi-skill dialogue dataset comprising 300K conversations.Through extensive experiments, we demonstrate that our dataset can be effective for multi-skill dialogue systems which require an understanding of skill blending as well as skill grounding.Our code and data are available at https://github. com/convei-lab/BotsTalk. Chaehyeong Kim, Yongho Song, Seung-won Hwang, Jinyoung Yeo |
EMNLP | 5 |
| 2022 | Modularized Transfer Learning with Multiple Knowledge Graphs for Zero-shot Commonsense ReasoningabstractYu Jin Kim, Beong-woo Kwak, Youngwook Kim, Reinald Kim Amplayo, Seung-won Hwang, Jinyoung Yeo. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Yu Jin Kim, Beong-woo Kwak, Youngwook Kim 0002, Reinald Kim Amplayo, Seung-won Hwang, Jinyoung Yeo |
NAACL-HLT | 6 |
| 2021 | Label and Context Augmentation for Response Selection at DSTC8abstractThis paper studies the dialogue response selection task. As state-of-the-arts are neural models requiring a large training set, data augmentation has been considered as a means to overcome the sparsity of observational annotation, where only one observed response is annotated as gold. In this paper, we first consider label augmentation, of selecting, among unobserved utterances, that would “counterfactually” replace the labeled response, for the given context, and augmenting labels only if that is the case. The key advantage of this model is not incurring human annotation overhead, thus not increasing the training cost, i.e., for low-resource scenarios. In addition, we consider context augmentation scenarios where the given dialogue context is not sufficient for label augmentation. In this case, inspired by open-domain question answering, we “decontextualize” by retrieving missing contexts, such as related persona. We empirically show that our pipeline improves BERT-based models in two different response selection tasks without incurring annotation overheads. Myeongho Jeong, Seungtaek Choi, Jinyoung Yeo, Seung-won Hwang |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Less is More: Attention Supervision with Counterfactuals for Text ClassificationabstractWe aim to leverage human and machine intelligence together for attention supervision.Specifically, we show that human annotation cost can be kept reasonably low, while its quality can be enhanced by machine selfsupervision.Specifically, for this goal, we explore the advantage of counterfactual reasoning, over associative reasoning typically used in attention supervision.Our empirical results show that this machine-augmented human attention supervision is more effective than existing methods requiring a higher annotation cost, in text classification tasks, including sentiment analysis and news categorization. Seungtaek Choi, Haeju Park, Jinyoung Yeo, Seung-won Hwang |
EMNLP (1) | 3 |
| 2020 | Conversion Prediction from Clickstream: Modeling Market Prediction and Customer PredictabilityabstractAs 98 percent of shoppers do not make a purchase on the first visit, we study the problem of predicting whether they would come back for a purchase later (i.e., conversion prediction). This problem is important for strategizing “retargeting”, for example, by sending coupons for customers who are likely to convert. For this goal, we study the following two problems, prediction of market and predictability of customer. First, prediction of market aims at identifying a conversion rate for a given product and its customer behavior modeling, which is an important analytics metric for retargeting process. Compared to existing approaches using either of customer or product-level conversion pattern, we propose a joint modeling of both patterns based on the well-studied buying decision process. Second, we can observe customer-specific behaviors after showing retargeting ads, to predict whether this specific customer follows the market model (high predictability) or not (low predictability). For the former, we apply the market model, and for the latter, we propose a new customer-specific prediction based on dynamic ad behavior features. To evaluate the effectiveness of our methods, we perform extensive experiments on the simulated dataset generated based on a set of real-world web logs and retargeting campaign logs. The evaluation results show that conversion predictions and predictability by our approach are consistently more accurate and robust than those by existing baselines in dynamic market environment. Jinyoung Yeo, Seung-won Hwang, Sungchul Kim, Eunyee Koh, Nedim Lipka |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2020 | XINA: Explainable Instance Alignment Using Dominance RelationshipabstractOver the past few years, knowledge bases (KBs) like DBPedia, Freebase, and YAGO have accumulated a massive amount of knowledge from web data. Despite their seemingly large size, however, individual KBs often lack comprehensive information on any given domain. For example, over 70 percent of people on Freebase lack information on place of birth. For this reason, the complementary nature across different KBs motivates their integration through a process of aligning instances. Meanwhile, since application-level machine systems, such as medical diagnosis, have heavily relied on KBs, it is necessary to provide users with trustworthy reasons why the alignment decisions are made. To address this problem, we propose a new paradigm, explainable instance alignment (XINA), which provides user-understandable explanations for alignment decisions. Specifically, given an alignment candidate, XINA replaces existing scalar representation of an aggregated score, by decision and explanation-vector spaces for machine decision and user understanding, respectively. To validate XINA, we perform extensive experiments on real-world KBs and show that XINA achieves comparable performance with state-of-the-arts, even with far less human effort. Jinyoung Yeo, Haeju Park, Eric Wonhee Lee, Seung-won Hwang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2019 | Soft Representation Learning for Sparse TransferabstractTransfer learning is effective for improving the performance of tasks that are related, and Multi-task learning (MTL) and Cross-lingual learning (CLL) are important instances.This paper argues that hard-parameter sharing, of hard-coding layers shared across different tasks or languages, cannot generalize well, when sharing with a loosely related task.Such case, which we call sparse transfer, might actually hurt performance, a phenomenon known as negative transfer.Our contribution is using adversarial training across tasks, to "softcode" shared and private spaces, to avoid the shared space gets too sparse.In CLL, our proposed architecture considers another challenge of dealing with low-quality input. Haeju Park, Jinyoung Yeo, Seung-won Hwang |
ACL (1) | 2 |
| 2019 | Learning with Limited Data for Multilingual Reading ComprehensionabstractKyungjae Lee, Sunghyun Park, Hojae Han, Jinyoung Yeo, Seung-won Hwang, Juho Lee. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Kyungjae Lee 0002, Sunghyun Park 0005, Hojae Han, Jinyoung Yeo, Seung-won Hwang |
EMNLP/IJCNLP (1) | 4 |
| 2019 | XINA: Explainable Instance Alignment using Dominance Relationship (Extended Abstract)abstractIn this extended abstract, we present an instance alignment framework, namely XINA, for KB integration. We then show its effectiveness and efficiency on real-world KBs. Jinyoung Yeo, Haeju Park, Eric Wonhee Lee, Seung-won Hwang |
ICDE | 1 |
| 2018 | Machine-Translated Knowledge Transfer for Commonsense Causal ReasoningabstractThis paper studies the problem of multilingual causal reasoning in resource-poor languages. Existing approaches, translating into the most probable resource-rich language such as English, suffer in the presence of translation and language gaps between different cultural area, which leads to the loss of causality. To overcome these challenges, our goal is thus to identify key techniques to construct a new causality network of cause-effect terms, targeted for the machine-translated English, but without any language-specific knowledge of resource-poor languages. In our evaluations with three languages, Korean, Chinese, and French, our proposed method consistently outperforms all baselines, achieving up-to 69.0% reasoning accuracy, which is close to the state-of-the-art accuracy 70.2% achieved on English. Jinyoung Yeo, Hyunsouk Cho, Seungtaek Choi, Seung-won Hwang |
AAAI | 1 |
| 2018 | Visual Choice of Plausible Alternatives: An Evaluation of Image-based Commonsense Causal Reasoning
Jinyoung Yeo, Gyeongbok Lee, Seungtaek Choi, Hyunsouk Cho, Reinald Kim Amplayo, Seung-won Hwang |
LREC | 1 |
| 2017 | Predicting Online Purchase Conversion for RetargetingabstractGenerally 2% of shoppers make a purchase on the first visit to an online store while the other 98% enjoys only window-shopping. To bring people back to the store and close the deal, "retargeting" has been a vital online advertising strategy that leads to "conversion" of window-shoppers into buyers. As such retargeting is more effective as a focused tool, in this paper, we study the problem of identifying a conversion rate for a given product and its current customers, which is an important analytics metric for retargeting process. Compared to existing approaches using either of customer- or product-level conversion pattern, we propose a joint modeling of both level patterns based on the well-studied buying decision process. To evaluate the effectiveness of our method, we perform extensive experiments on the simulated dataset generated based on a set of real-world web logs. The evaluation results show that conversion predictions by our approach are consistently more accurate and robust than those by existing baselines in dynamic market environment. Jinyoung Yeo, Sungchul Kim, Eunyee Koh, Seung-won Hwang, Nedim Lipka |
WSDM | 1 |
| 2017 | Efficient Keyword-Aware Representative Travel Route RecommendationabstractWith the popularity of social media (e.g., Facebook and Flicker), users can easily share their check-in records and photos during their trips. In view of the huge number of user historical mobility records in social media, we aim to discover travel experiences to facilitate trip planning. When planning a trip, users always have specific preferences regarding their trips. Instead of restricting users to limited query options such as locations, activities, or time periods, we consider arbitrary text descriptions as keywords about personalized requirements. Moreover, a diverse and representative set of recommended travel routes is needed. Prior works have elaborated on mining and ranking existing routes from check-in data. To meet the need for automatic trip organization, we claim that more features of Places of Interest (POIs) should be extracted. Therefore, in this paper, we propose an efficient Keyword-aware Representative Travel Route framework that uses knowledge extraction from users' historical mobility records and social interactions. Explicitly, we have designed a keyword extraction module to classify the POI-related tags, for effective matching with query keywords. We have further designed a route reconstruction algorithm to construct route candidates that fulfill the requirements. To provide befitting query results, we explore Representative Skyline concepts, that is, the Skyline routes which best describe the trade-offs among different POI features. To evaluate the effectiveness and efficiency of the proposed algorithms, we have conducted extensive experiments on real location-based social network datasets, and the experiment results show that our methods do indeed demonstrate good performance compared to state-of-the-art works. Yu Ting Wen, Jinyoung Yeo, Wen-Chih Peng, Seung-won Hwang |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | Multimodal KB Harvesting for Emerging Spatial EntitiesabstractNew entities are being created daily. Though the novelty of these entities naturally attracts mentions, due to lack of prior knowledge, it is more challenging to collect knowledge about such entities than pre-existing entities, whose KBs are comprehensively annotated through LBSNs and EBSNs. In this paper, we focus on knowledge harvesting for emerging spatial entities (ESEs), such as new businesses and venues, assuming we have only a list of ESE names. Existing techniques for knowledge base (KB) harvesting are primarily associated with information extraction from textual corpora. In contrast, we propose a multimodal method for event detection based on the complementary interaction of image, text, and user information between multi-source platforms, namely Flickr and Twitter. We empirically validate our harvesting approaches improve the quality of KB with enriched place and event knowledge. Jinyoung Yeo, Hyunsouk Cho, Seung-won Hwang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2016 | Understanding Emerging Spatial EntitiesabstractIn Foursquare or Google+ Local, emerging spatial entities, such as new business or venue, are reported to grow by 1% every day. As information on such spatial entities is initially limited (e.g., only name), we need to quickly harvest related information from social media such as Flickr photos. Especially, achieving high-recall in photo population is essential for emerging spatial entities, which suffer from data sparseness (e.g., 71% restaurants of TripAdvisor in Seattle do not have any photo, as of Sep 03, 2015). Our goal is thus to address this limitation by identifying effective linking techniques for emerging spatial entities and photos. Compared with state-of-the-art baselines, our proposed approach improves recall and F1 score by up to 24% and 18%, respectively. To show the effectiveness and robustness of our approach, we have conducted extensive experiments in three different cities, Seattle, Washington D.C., and Taipei, of varying characteristics such as geographical density and language. Jinyoung Yeo, Seung-won Hwang |
AAAI | 1 |
| 2016 | Event Grounding from Multimodal Social Network FusionabstractThis paper studies the problem of extracting real world event information from social media streams. Although existing work focuses on event signals of bursty mentions extracted from a single-source of textual streams, these signals are likely to be noisy due to ambiguous occurrences of individual mentions. To extract accurate event signals, we propose a framework capable of "grounding" mentions to unique event using multiple social networks with complementary strength. We show that our framework jointly using multiple sources outperforms state-of-the-arts using publicly available datasets. Hyunsouk Cho, Jinyoung Yeo, Seung-won Hwang |
ICDM | 2 |
| 2015 | KSTR: Keyword-Aware Skyline Travel Route RecommendationabstractWith the popularity of social media (e.g., Facebook and Flicker), users could easily share their check-in records and photos during their trips. In view of the huge amount of check-in data and photos in social media, we intend to discover travel experiences to facilitate trip planning. Prior works have been elaborated on mining and ranking existing travel routes from check-in data. We observe that when planning a trip, users may have some keywords about preference on his/her trips. Moreover, a diverse set of travel routes is needed. To provide a diverse set of travel routes, we claim that more features of Places of Interests (POIs) should be extracted. Therefore, in this paper, we propose a Keyword-aware Skyline Travel Route (KSTR) framework that use knowledge extraction from historical mobility records and the user's social interactions. Explicitly, we model the "Where, When, Who" issues by featurizing the geographical mobility pattern, temporal influence and social influence. Then we propose a keyword extraction module to classify the POI-related tags automatically into different types, for effective matching with query keywords. We further design a route reconstruction algorithm to construct route candidates that fulfill the query inputs. To provide diverse query results, we explore Skyline concepts to rank routes. To evaluate the effectiveness and efficiency of the proposed algorithms, we have conducted extensive experiments on real location-based social network datasets, and the experimental results show that KSTR does indeed demonstrate good performance compared to state-of-the-art works. Yu Ting Wen, Kae-Jer Cho, Wen-Chih Peng, Jinyoung Yeo, Seung-won Hwang |
ICDM | 4 |
| 2012 | Finding influential products on social domination gameabstractIn this paper, we propose a new type of market model called the social domination game model. Given a set C of customers and a set P of products, this model simulates market competition among P and estimates market shares, considering both the dominance relation between C and P and the influence relation among the members of C. With this model, we propose a greedy product positioning algorithm for designing a new product that approximately maximizes market share. Our experimental results show that the proposed algorithm creates a new product gaining up to 97.5% market share of the best product's market share obtained by the exact method, while significantly outperforming the exact method in terms of running time, i.e., by up to two orders of magnitude. Jinyoung Yeo, Seung-won Hwang |
CIKM | 1 |