EDBT 2026 Demo / reviewers in the wild / expert
Yohan Jo
dblp:40/8877
· DBLP profile ↗
36ranked-venue papers
10as first author
23since 2021 · last 2026
0009-0006-9296-3403ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 8 first-author · 22 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpeakerSleuth: Can Large Audio-Language Models Judge Speaker Consistency across Multi-turn Dialogues?abstractLarge Audio-Language Models (LALMs) as judges have emerged as a prominent approach for evaluating speech generation quality, yet their ability to assess speaker consistency across multi-turn dialogues remains unexplored.We present SpeakerSleuth, a benchmark evaluating whether LALMs can reliably judge speaker consistency across multi-turn dialogues through three tasks reflecting real-world requirements.We construct 1,818 humanverified evaluation instances across four diverse datasets spanning synthetic and real speech, with controlled acoustic difficulty.Evaluating twelve widely-used LALMs, we find that models struggle to reliably detect acoustic inconsistencies.For instance, given audio samples of the same speaker's turns, some models overpredict inconsistency, whereas others are overly lenient.Models further struggle to identify the exact turns that are problematic.When other interlocutors' turns are provided as textual context, performance degrades dramatically as models prioritize textual coherence over acoustic cues, failing to detect even obvious gender switches for a speaker.On the other hand, models perform substantially better in comparing and ranking acoustic variants, demonstrating inherent acoustic discrimination capabilities.These findings expose a significant bias in LALMs: they tend to prioritize text over acoustics, revealing fundamental modality imbalances that need to be addressed to build reliable audio-language judges.Our code and data are available at https: //github.com/holi-lab/SpeakerSleuth. Jonggeun Lee, Junseong Pyo, Gyuhyeon Seo, Yohan Jo |
ACL (1) | 4 |
| 2026 | Don't Adapt Small Language Models for Tools; Adapt Tool Schemas to the ModelsabstractSmall language models (SLMs) enable scalable tool-augmented multi-agent systems where multiple SLMs handle subtasks orchestrated by a powerful coordinator.However, they struggle with tool-use tasks, particularly in selecting appropriate tools and identifying correct parameters.A common failure mode is schema misalignment: models hallucinate plausible tool names that are absent from the provided tool schema, due to different naming conventions internalized during pretraining.Rather than training models to adapt to unfamiliar schemas, we propose adapting schemas to align with models' pretrained knowledge.We introduce PA-Tool (Pretraining-Aligned Tool Schema Generation), a training-free method that leverages peakedness, a signal used in contamination detection that indicates pretraining familiarity, to rename tool components.By generating multiple candidates and selecting the candidate with the highest peakedness, PA-Tool identifies pretrainingaligned naming patterns.Experiments on Meta-Tool and RoTBench show improvements of up to 17%, with schema misalignment errors reduced by 80%.PA-Tool enables small models to substantially improve tool-use accuracy without retraining, showing that schema-level interventions can unlock the tool-use potential of resource-efficient models.Our code is available at https://github.com/holi-lab/P A-Tool. Jonggeun Lee, Woojung Song, Jongwook Han, Haesung Pyun, Yohan Jo |
ACL (1) | 5 |
| 2026 | MatKV: Trading Compute for Flash Storage in LLM InferenceabstractWe observe two major trends in LLM-based generative AI: (1) inference is becoming the dominant factor in terms of cost and power consumption, surpassing training, and (2) retrieval augmented generation (RAG) is becoming prevalent. When processing long inputs in RAG, the prefill phase of computing the key-value vectors of input text is energy-intensive and time-consuming even with high-end GPUs. Thus, it is crucial to make the prefill phase in RAG inference efficient. To address this issue, we propose MatKV, a scheme that precomputes the key-value vectors (KVs) of RAG objects (e.g., documents), materializes them in inexpensive but fast and power-efficient flash storage, and reuses them at inference time instead of recomputing the KVs using costly and power-inefficient GPU. Experimental results using Hugging Face's Transformers library across state-of-the-art GPUs and flash memory SSDs confirm that, compared to full KV computation on GPUs, MatKV reduces both inference time and power consumption by half for RAG workloads, without severely impacting accuracy in the question-answering task. Furthermore, we demonstrate that MatKV enables additional optimizations in two ways. First, a GPU can decode text while simultaneously loading the materialized KVs for the next instance, reducing load latency. Second, since decoding speed is less sensitive to GPU performance than KV computation, low-end GPUs can be leveraged for decoding without significantly compromising speed once the materialized KVs are loaded into GPU memory. These findings underscore MatKV's potential to make large-scale generative AI applications more cost-effective, power-efficient, and accessible across a wider range of tasks and hardware environments. Kunwoo Shin, Jay H. Park, Moonwook Oh, Yohan Jo, Jaeyoung Do |
ICDE | 4 |
| 2026 | In-N-Out: A Parameter-Level API Graph Dataset for Tool AgentsabstractAbstract Tool agents—LLM-based systems that interact with external APIs—offer a way to execute real-world tasks. However, as tasks become increasingly complex, these agents struggle to identify and call the correct APIs in the proper order. To tackle this problem, we investigate converting API documentation into a structured API graph that captures API dependencies and leveraging it for multi-tool queries that require compositional API calls. To support this, we introduce In-N-Out, the first expert-annotated dataset of API graphs built from two real-world API benchmarks and their documentation. Using In-N-Out significantly improves performance on both tool retrieval and multi-tool query generation, nearly doubling that of LLMs using documentation alone. Moreover, graphs generated by models fine-tuned on In-N-Out close 90% of this gap, showing that our dataset helps models learn to comprehend API documentation and parameter relationships. Our findings highlight the promise of using explicit API graphs for tool agents and the utility of In-N-Out as a valuable resource. We release our dataset and code at https://github.com/holi-lab/In-N-Out-API-Graph. Nalim Kim, Yohan Jo |
Trans. Assoc. Comput. Linguistics | 3 |
| 2026 | Psychometric Item Validation Using Virtual Respondents with Trait-Response MediatorsabstractAbstract As psychometric surveys are increasingly used to assess the traits of large language models (LLMs), the need for scalable survey item generation suited for LLMs has also grown. A critical challenge here is ensuring the construct validity of generated items, i.e., whether they truly measure the intended trait. Traditionally, this requires costly, large-scale human data collection. To make it efficient, we present a framework for virtual respondent simulation using LLMs. Our central idea is to account for mediators: factors through which the same trait can give rise to varying responses to a survey item. By simulating respondents with diverse mediators, we identify survey items that yield responses robustly correlated with intended traits across these mediators. Experiments on three psychological trait theories (Big5, Schwartz, VIA) show that our mediator generation methods and simulation framework effectively identify high-validity items. LLMs demonstrate the ability to generate plausible mediators from trait definitions and to simulate respondent behavior for item validation. Our problem formulation, metrics, methodology, and dataset open a new direction for cost-efficient survey development and a deeper understanding of how LLMs simulate human survey responses. We release our dataset and code to support future work.1 Sungjib Lim, Woojung Song, Yohan Jo |
Trans. Assoc. Comput. Linguistics | 4 |
| 2025 | Value Portrait: Assessing Language Models' Values through Psychometrically and Ecologically Valid ItemsabstractThe importance of benchmarks for assessing the values of language models has been pronounced due to the growing need of more authentic, human-aligned responses.However, existing benchmarks rely on human or machine annotations that are vulnerable to value-related biases.Furthermore, the tested scenarios often diverge from real-world contexts in which models are commonly used to generate text and express values.To address these issues, we propose the Value Portrait benchmark, a reliable framework for evaluating LLMs' value orientations with two key characteristics.First, the benchmark consists of items that capture real-life user-LLM interactions, enhancing the relevance of assessment results to real-world LLM usage.Second, each item is rated by human subjects based on its similarity to their own thoughts, and correlations between these ratings and the subjects' actual value scores are derived.This psychometrically validated approach ensures that items strongly correlated with specific values serve as reliable items for assessing those values.Through evaluating 44 LLMs with our benchmark, we find that these models prioritize Benevolence, Security, and Self-Direction values while placing less emphasis on Tradition, Power, and Achievement values.Also, our analysis reveals biases in how LLMs perceive various demographic groups, deviating from real human data. 1 Jongwook Han, Dongmin Choi, Woojung Song, Yohan Jo |
ACL (1) | 5 |
| 2025 | PVP: An Image Dataset for Personalized Visual Persuasion with Persuasion Strategies, Viewer Characteristics, and Persuasiveness RatingsabstractVisual persuasion, which uses visual elements to influence cognition and behaviors, is crucial in fields such as advertising and politicalcommunication. With recent advancements in artificial intelligence, there is growing potential to develop persuasive systems that automatically generate persuasive images tailored to individuals. However, a significant bottleneck in this area is the lack of comprehensivedatasets that connect the persuasiveness of images with the personal information about those who evaluated the images. To address this gap and facilitate technological advancements in personalized visual persuasion, we release the Personalized Visual Persuasion (PVP) dataset, comprising 28,454 persuasive images across 596 messages and 9 persuasion strategies. Importantly, the PVP dataset provides persuasiveness scores of images evaluated by 2,521 human annotators, along with their demographic and psychological characteristics (personality traits and values). We demonstrate the utility of our dataset by developing a persuasive image generator and an automated evaluator, and establish benchmark baselines. Our experiments reveal that incorporating psychological characteristics enhances the generation and evaluation of persuasive images, providing valuable insights for personalized visual persuasion. Junseo Kim, Jongwook Han, Dongmin Choi, Jongwook Yoon, Yohan Jo |
ACL (1) | 6 |
| 2025 | Knowledge Tracing in Programming Education Integrating Students' QuestionsabstractKnowledge tracing (KT) in programming education presents unique challenges due to the complexity of coding tasks and the diverse methods students use to solve problems.Although students' questions often contain valuable signals about their understanding and misconceptions, traditional KT models often neglect to incorporate these questions as inputs to address these challenges.This paper introduces SQKT (Students' Question-based Knowledge Tracing), a knowledge tracing model that leverages students' questions and automatically extracted skill information to enhance the accuracy of predicting students' performance on subsequent problems in programming education.Our method creates semantically rich embeddings that capture not only the surfacelevel content of the questions but also the student's mastery level and conceptual understanding.Experimental results demonstrate SQKT's superior performance in predicting student completion across various Python programming courses of differing difficulty levels.In in-domain experiments, SQKT achieved a 33.1% absolute improvement in AUC compared to baseline models.The model also exhibited robust generalization capabilities in cross-domain settings, effectively addressing data scarcity issues in advanced programming courses.SQKT can be used to tailor educational content to individual learning needs and design adaptive learning systems in computer science education.Our code is available at Doyoun Kim, Suin Kim, Yohan Jo |
ACL (1) | 3 |
| 2025 | Dialogue Systems for Emotional Support via Value ReinforcementabstractEmotional support dialogue systems aim to reduce help-seekers' distress and help them overcome challenges.While human values-core beliefs that shape an individual's prioritiesare increasingly emphasized in contemporary psychological therapy for their role in fostering internal transformation and long-term emotional well-being, their integration into emotional support systems remains underexplored.To bridge this gap, we present a value-driven method for training emotional support dialogue systems designed to reinforce positive values in seekers.Notably, our model identifies which values to reinforce at each turn and how to do so, by leveraging online support conversations from Reddit.We evaluate the method across support skills, seekers' emotional intensity, and value reinforcement.Our method consistently outperforms various baselines, effectively exploring and eliciting values from seekers.Additionally, leveraging crowd knowledge from Reddit significantly enhances its effectiveness.Therapists highlighted its ability to validate seekers' challenges and emphasize positive aspects of their situations-both crucial elements of value reinforcement.Our work, being the first to integrate value reinforcement into emotional support systems, demonstrates its promise and establishes a foundation for future research. Juhee Kim, Chunghu Mok, Jisun Lee, Hyang Sook Kim, Yohan Jo |
ACL (1) | 5 |
| 2025 | Generating Plausible Distractors for Multiple-Choice Questions via Student Choice PredictionabstractIn designing multiple-choice questions (MCQs) in education, creating plausible distractors is crucial for identifying students’ misconceptions and gaps in knowledge and accurately assessing their understanding. However, prior studies on distractor generation have not paid sufficient attention to enhancing the difficulty of distractors, resulting in reduced effectiveness of MCQs. This study presents a pipeline for training a model to generate distractors that are more likely to be selected by students. First, we train a pairwise ranker to reason about students’ misconceptions and assess the relative plausibility of two distractors. Using this model, we create a dataset of pairwise distractor ranks and then train a distractor generator via Direct Preference Optimization (DPO) to generate more plausible distractors. Experiments on computer science subjects (Python, DB, MLDL) demonstrate that our pairwise ranker effectively identifies students’ potential misunderstandings and achieves ranking accuracy comparable to human experts. Furthermore, our distractor generator outperforms several baselines in generating plausible distractors and produces questions with a higher item discrimination index (DI). Yooseop Lee, Suin Kim, Yohan Jo |
ACL (1) | 3 |
| 2025 | Improving Dialogue State Tracking through Combinatorial Search for In-Context ExamplesabstractIn dialogue state tracking (DST), in-context learning comprises a retriever that selects labeled dialogues as in-context examples and a DST model that uses these examples to infer the dialogue state of the query dialogue.Existing methods for constructing training data for retrievers suffer from three key limitations: (1) the synergistic effect of examples is not considered, (2) the linguistic characteristics of the query are not sufficiently factored in, and (3) scoring is not directly optimized for DST performance.Consequently, the retriever can fail to retrieve examples that would substantially improve DST performance.To address these issues, we present CombiSearch-a method that scores effective in-context examples based on their combinatorial impact on DST performance.Our evaluation on MultiWOZ shows that retrievers trained with CombiSearch surpass state-of-the-art models, achieving a 20× gain in data efficiency and generalizing well to the SGD dataset.Moreover, CombiSearch attains a 12% absolute improvement in the upper bound DST performance over traditional approaches when no retrieval errors are assumed.This significantly increases the headroom for practical DST performance while demonstrating that existing methods rely on suboptimal data for retriever training.1 Haesung Pyun, Yoonah Park, Yohan Jo |
ACL (1) | 3 |
| 2025 | Sparse and Dense Retrievers Learn Better Together: Joint Sparse-Dense Optimization for Text-Image RetrievalabstractVision-Language Pretrained (VLP) models have achieved impressive performance on multimodal tasks, including text-image retrieval, based on dense representations. Meanwhile, Learned Sparse Retrieval (LSR) has gained traction in text-only settings due to its interpretability and efficiency with fast term-based lookup via inverted indexes. Inspired by these advantages, recent work has extended LSR to the multimodal domain. However, these methods often rely on computationally expensive contrastive pre-training, or distillation from a frozen dense model, which limits the potential for mutual enhancement. To address these limitations, we propose a simple yet effective framework that enables bi-directional learning between dense and sparse representations through Self-Knowledge Distillation. This bi-directional learning is achieved using an integrated similarity score-a weighted sum of dense and sparse similarities-which serves as a shared teacher signal for both representations. To ensure efficiency, we fine-tune the final layer of the dense encoder and the sparse projection head, enabling easy adaptation of any existing VLP model. Experiments on MSCOCO and Flickr30k demonstrate that our sparse retriever not only outperforms existing sparse baselines, but also achieves performance comparable to-or even surpassing-its dense counterparts, while retaining the benefits of sparse models. Jonghyun Song, Youngjune Lee, Gyu-Hwung Cho, Ilhyeon Song, Saehun Kim, Yohan Jo |
CIKM | 6 |
| 2025 | ToolDial: Multi-turn Dialogue Generation Method for Tool-Augmented Language ModelsabstractTool-Augmented Language Models (TALMs) leverage external APIs to answer user queries across various domains. However, existing benchmark datasets for TALM research often feature simplistic dialogues that do not reflect real-world scenarios, such as the need for models to ask clarifying questions or proactively call additional APIs when essential information is missing. To address these limitations, we construct and release ToolDial, a dataset comprising 11,111 multi-turn dialogues, with an average of 8.95 turns per dialogue, based on APIs from RapidAPI. ToolDial has two key characteristics. First, the dialogues incorporate 16 user and system actions (e.g., request, clarify, fail inform) to capture the rich dynamics of real-world interactions. Second, we simulate dialogues where the system requests necessary information from the user based on API documentation and seeks additional APIs if the user fails to provide the required information. To facilitate this process, we introduce a method for generating an API graph that represents input and output compatibility between APIs. Using ToolDial, we evaluate a suite of language models on their ability to predict correct actions and extract input parameter values for API calls from the dialogue history. Modern language models achieve accuracy scores below 70\%, indicating substantial room for improvement. We provide a detailed analysis of the areas where these models fall short. Jeonghoon Shim, Gyuhyeon Seo, Cheongsu Lim, Yohan Jo |
ICLR | 4 |
| 2025 | KMI: A Dataset of Korean Motivational Interviewing Dialogues for PsychotherapyabstractHyunjong Kim, Suyeon Lee, Yeongjae Cho, Eunseo Ryu, Yohan Jo, Suran Seong, Sungzoon Cho. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Hyunjong Kim, Suyeon Lee, Yeongjae Cho, Eunseo Ryu, Yohan Jo, Suran Seong, Sungzoon Cho |
NAACL (Long Papers) | 5 |
| 2025 | Towards Lifelong Dialogue Agents via Timeline-based Memory ManagementabstractKai Tzu-iunn Ong, Namyoung Kim, Minju Gwak, Hyungjoo Chae, Taeyoon Kwon, Yohan Jo, Seung-won Hwang, Dongha Lee, Jinyoung Yeo. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kai Tzu-iunn Ong, Namyoung Kim, Minju Gwak, Hyungjoo Chae, Taeyoon Kwon, Yohan Jo, Seung-won Hwang, Dongha Lee 0003, Jinyoung Yeo |
NAACL (Long Papers) | 6 |
| 2024 | Generalizing Visual Question Answering from Synthetic to Human-Written Questions via a Chain of QA with a Large Language ModelabstractVisual question answering (VQA) is a task where an image is given, and a series of questions are asked about the image. To build an efficient VQA algorithm, a large amount of QA data is required which is very expensive. Generating synthetic QA pairs based on templates is a practical way to obtain data. However, VQA models trained on those data do not perform well on complex, human-written questions. To address this issue, we propose a new method called chain of QA for human-written questions (CoQAH). CoQAH utilizes a sequence of QA interactions between a large language model and a VQA model trained on synthetic data to reason and derive logical answers for human-written questions. We tested the effectiveness of CoQAH on two types of human-written VQA datasets for 3D-rendered and chest X-ray images and found that it achieved state-of-the-art accuracy in both types of data. Notably, CoQAH outperformed general vision-language models, VQA models, and medical foundation models with no finetuning. The source code for CoQAH is available [10]. Yeongjae Cho, Heejun Shin, Yohan Jo, Dongmyung Shin |
ECAI | 4 |
| 2024 | Model-based Preference Optimization in Abstractive Summarization without Human FeedbackabstractIn abstractive summarization, the challenge of producing concise and accurate summaries arises from the vast amount of information contained in the source document.Consequently, although Large Language Models (LLMs) can generate fluent text, they often introduce inaccuracies by hallucinating content not found in the original source.While supervised finetuning methods that maximize likelihood contribute to this issue, they do not consistently enhance the faithfulness of the summaries.Preference-based optimization methods, such as Direct Preference Optimization (DPO), can further refine the model to align with human preferences.However, these methods still heavily depend on costly human feedback.In this work, we introduce a novel and straightforward approach called Model-based Preference Optimization (MPO) to fine-tune LLMs for improved summarization abilities without any human feedback.By leveraging the model's inherent summarization capabilities, we create a preference dataset that is fully generated by the model using different decoding strategies.Our experiments on standard summarization datasets and various metrics demonstrate that our proposed MPO significantly enhances the quality of generated summaries without relying on human feedback.The code is publicly available at https://github.com/cjaep/MPO. Jaepill Choi, Kyubyung Chae, Jiwoo Song, Yohan Jo, Taesup Kim |
EMNLP | 4 |
| 2023 | FactKG: Fact Verification via Reasoning on Knowledge GraphsabstractIn real world applications, knowledge graphs (KG) are widely used in various domains (e.g.medical applications and dialogue agents).However, for fact verification, KGs have not been adequately utilized as a knowledge source.KGs can be a valuable knowledge source in fact verification due to their reliability and broad applicability.A KG consists of nodes and edges which makes it clear how concepts are linked together, allowing machines to reason over chains of topics.However, there are many challenges in understanding how these machine-readable concepts map to information in text.To enable the community to better use KGs, we introduce a new dataset, FACTKG: Fact Verification via Reasoning on Knowledge Graphs.It consists of 108k natural language claims with five types of reasoning: One-hop, Conjunction, Existence, Multi-hop, and Negation.Furthermore, FACTKG contains various linguistic patterns, including colloquial style claims as well as written style claims to increase practicality.Lastly, we develop a baseline approach and analyze FACTKG over these reasoning types.We believe FACTKG can advance both reliability and practicality in KG-based fact verification.1 Yeonsu Kwon, Yohan Jo, James Thorne, Edward Choi 0003 |
ACL (1) | 4 |
| 2023 | From Values to Opinions: Predicting Human Behaviors and Stances Using Value-Injected Large Language ModelsabstractBeing able to predict people's opinions on issues and behaviors in realistic scenarios can be helpful in various domains, such as politics and marketing.However, conducting largescale surveys like the European Social Survey to solicit people's opinions on individual issues can incur prohibitive costs.Leveraging prior research showing influence of core human values on individual decisions and actions, we propose to use value-injected large language models (LLM) to predict opinions and behaviors.To this end, we present Value Injection Method (VIM), a collection of two methodsargument generation and question answeringdesigned to inject targeted value distributions into LLMs via fine-tuning.We then conduct a series of experiments on four tasks to test the effectiveness of VIM and the possibility of using value-injected LLMs to predict opinions and behaviors of people.We find that LLMs valueinjected with variations of VIM substantially outperform the baselines.Also, the results suggest that opinions and behaviors can be better predicted using value-injected LLMs than the baseline approaches. Dongjun Kang, Joonsuk Park, Yohan Jo, JinYeong Bak |
EMNLP | 3 |
| 2023 | A Closer Look at the Intervention Procedure of Concept Bottleneck ModelsabstractConcept bottleneck models (CBMs) are a class of interpretable neural network models that predict the target response of a given input based on its high-level concepts. Unlike the standard end-to-end models, CBMs enable domain experts to intervene on the predicted concepts and rectify any mistakes at test time, so that more accurate task predictions can be made at the end. While such intervenability provides a powerful avenue of control, many aspects of the intervention procedure remain rather unexplored. In this work, we develop various ways of selecting intervening concepts to improve the intervention effectiveness and conduct an array of in-depth analyses as to how they evolve under different circumstances. Specifically, we find that an informed intervention strategy can reduce the task error more than ten times compared to the current baseline under the same amount of intervention counts in realistic settings, and yet, this can vary quite significantly when taking into account different intervention granularity. We verify our findings through comprehensive evaluations, not only on the standard real datasets, but also on synthetic datasets that we generate based on a set of different causal graphs. We further discover some major pitfalls of the current practices which, without a proper addressing, raise concerns on reliability and fairness of the intervention procedure. Sungbin Shin, Yohan Jo, Sungsoo Ahn, Namhoon Lee |
ICML | 2 |
| 2023 | A Zero-Shot Approach for Multi-User Task-Oriented Dialog GenerationabstractPrior art investigating task-oriented dialog and automatic generation of such dialogs have focused on single-user dialogs between a single user and an agent.However, there is limited study on adapting such AI agents to multiuser conversations (involving multiple users and an agent).Multi-user conversations are richer than single-user conversations containing social banter and collaborative decision making.The most significant challenge impeding such studies is the lack of suitable multiuser task-oriented dialogs with annotations of user belief states and system actions.One potential solution is multi-user dialog generation from single-user data.Many single-user dialogs datasets already contain dialog state information (intents, slots), thus making them suitable candidates.In this work, we propose a novel approach for expanding single-user taskoriented dialogs (e.g.MultiWOZ) to multiuser dialogs in a zero-shot setting. Shiv Surya, Yohan Jo, Arijit Biswas, Alexandros Potamianos |
INLG | 2 |
| 2022 | Argument Mining for Review Helpfulness PredictionabstractThe importance of reliably determining the helpfulness of product reviews is rising as both helpful and unhelpful reviews continue to accumulate on e-commerce websites.And argumentational features-such as the structure of arguments and the types of underlying elementary units-have shown to be promising indicators of product review helpfulness.However, their adoption has been limited due to the lack of sufficient resources and large-scale experiments investigating their utility.To this end, we present the AMazon Argument Mining (AM 2 ) corpus-a corpus of 878 Amazon reviews on headphones annotated according to a theoretical argumentation model designed to evaluate argument quality.Experiments show that employing argumentational features leads to statistically significant improvements over the state-of-the-art review helpfulness predictors under both text-only and text-and-image settings.1 Zaiqian Chen, Daniel Verdi do Amarante, Jenna Donaldson, Yohan Jo, Joonsuk Park |
EMNLP | 4 |
| 2021 | Classifying Argumentative Relations Using Logical Mechanisms and Argumentation SchemesabstractWhile argument mining has achieved significant success in classifying argumentative relations between statements (support, attack, and neutral), we have a limited computational understanding of logical mechanisms that constitute those relations. Most recent studies rely on black-box models, which are not as linguistically insightful as desired. On the other hand, earlier studies use rather simple lexical features, missing logical relations between statements. To overcome these limitations, our work classifies argumentative relations based on four logical and theory-informed mechanisms between two statements, namely, (i) factual consistency, (ii) sentiment coherence, (iii) causal relation, and (iv) normative relation. We demonstrate that our operationalization of these logical mechanisms classifies argumentative relations without directly training on data labeled with the relations, significantly better than several unsupervised baselines. We further demonstrate that these mechanisms also improve supervised classifiers through representation learning. Yohan Jo, Seo-Jin Bang, Chris Reed 0001, Eduard H. Hovy |
Trans. Assoc. Comput. Linguistics | 1 |
| 2020 | Detecting Attackable Sentences in ArgumentsabstractFinding attackable sentences in an argument is the first step toward successful refutation in argumentation.We present a first large-scale analysis of sentence attackability in online arguments.We analyze driving reasons for attacks in argumentation and identify relevant characteristics of sentences.We demonstrate that a sentence's attackability is associated with many of these characteristics regarding the sentence's content, proposition types, and tone, and that an external knowledge source can provide useful information about attackability.Building on these findings, we demonstrate that machine learning models can automatically detect attackable sentences in arguments, significantly better than several baselines and comparably well to laypeople. 1 Yohan Jo, Seo-Jin Bang, Emaad Manzoor, Eduard H. Hovy, Chris Reed 0001 |
EMNLP (1) | 1 |
| 2020 | Extracting Implicitly Asserted Propositions in ArgumentationabstractArgumentation accommodates various rhetorical devices, such as questions, reported speech, and imperatives.These rhetorical tools usually assert argumentatively relevant propositions rather implicitly, so understanding their true meaning is key to understanding certain arguments properly.However, most argument mining systems and computational linguistics research have paid little attention to implicitly asserted propositions in argumentation.In this paper, we examine a wide range of computational methods for extracting propositions that are implicitly asserted in questions, reported speech, and imperatives in argumentation.By evaluating the models on a corpus of 2016 U.S. presidential debates and online commentary, we demonstrate the effectiveness and limitations of the computational models.Our study may inform future research on argument mining and the semantics of these rhetorical devices in argumentation. 1 Yohan Jo, Jacky Visser, Chris Reed 0001, Eduard H. Hovy |
EMNLP (1) | 1 |
| 2020 | Machine-Aided Annotation for Fine-Grained Proposition Types in ArgumentationabstractWe introduce a corpus of the 2016 U.S. presidential debates and commentary, containing 4,648 argumentative propositions annotated with fine-grained proposition types. Modern machine learning pipelines for analyzing argument have difficulty distinguishing between types of propositions based on their factuality, rhetorical positioning, and speaker commitment. Inability to properly account for these facets leaves such systems inaccurate in understanding of fine-grained proposition types. In this paper, we demonstrate an approach to annotating for four complex proposition types, namely normative claims, desires, future possibility, and reported speech. We develop a hybrid machine learning and human workflow for annotation that allows for efficient and reliable annotation of complex linguistic phenomena, and demonstrate with preliminary analysis of rhetorical strategies and structure in presidential debates. This new dataset and method can support technical researchers seeking more nuanced representations of argument, as well as argumentation theorists developing new quantitative analyses. Yohan Jo, Elijah Mayfield, Chris Reed 0001, Eduard H. Hovy |
LREC | 1 |
| 2018 | Perceptions of Censorship and Moderation Bias in Political Debate Forums
Qinlan Shen, Michael Miller Yoder, Yohan Jo, Carolyn P. Rosé |
ICWSM | 3 |
| 2018 | Attentive Interaction Model: Modeling Changes in View in ArgumentationabstractYohan Jo, Shivani Poddar, Byungsoo Jeon, Qinlan Shen, Carolyn Rosé, Graham Neubig. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Yohan Jo, Shivani Poddar, Byungsoo Jeon, Qinlan Shen, Carolyn P. Rosé, Graham Neubig |
NAACL-HLT | 1 |
| 2017 | Modeling Dialogue Acts with Content Word Filtering and Speaker PreferencesabstractWe present an unsupervised model of dialogue act sequences in conversation. By modeling topical themes as transitioning more slowly than dialogue acts in conversation, our model de-emphasizes content-related words in order to focus on conversational function words that signal dialogue acts. We also incorporate speaker tendencies to use some acts more than others as an additional predictor of dialogue act prevalence beyond temporal dependencies. According to the evaluation presented on two dissimilar corpora, the CNET forum and NPS Chat corpus, the effectiveness of each modeling assumption is found to vary depending on characteristics of the data. De-emphasizing content-related words yields improvement on the CNET corpus, while utilizing speaker tendencies is advantageous on the NPS corpus. The components of our model complement one another to achieve robust performance on both corpora and outperform state-of-the-art baseline models. Yohan Jo, Michael Miller Yoder, Hyeju Jang, Carolyn P. Rosé |
EMNLP | 1 |
| 2017 | Roles and Success in Wikipedia Talk Pages: Identifying Latent Patterns of BehaviorabstractIn this work we investigate how role-based behavior profiles of a Wikipedia editor, considered against the backdrop of roles taken up by other editors in discussions, predict the success of the editor at achieving an impact on the associated article. We first contribute a new public dataset including a task predicting the success of Wikipedia editors involved in discussion, measured by an operationalization of the lasting impact of their edits in the article. We then propose a probabilistic graphical model that advances earlier work inducing latent discussion roles using the light supervision of success in the negotiation task. We evaluate the performance of the model and interpret findings of roles and group configurations that lead to certain outcomes on Wikipedia. Korte Maki, Michael Miller Yoder, Yohan Jo, Carolyn P. Rosé |
IJCNLP(1) | 3 |
| 2016 | Metaphor Detection with Topic Transition, Emotion and Cognition in ContextabstractMetaphor is a common linguistic tool in communication, making its detection in discourse a crucial task for natural language understanding.One popular approach to this challenge is to capture semantic incohesion between a metaphor and the dominant topic of the surrounding text.While these methods are effective, they tend to overclassify target words as metaphorical when they deviate in meaning from its context.We present a new approach that (1) distinguishes literal and non-literal use of target words by examining sentence-level topic transitions and (2) captures the motivation of speakers to express emotions and abstract concepts metaphorically.Experiments on an online breast cancer discussion forum dataset demonstrate a significant improvement in metaphor detection over the state-of-theart.These experimental results also reveal a tendency toward metaphor usage in personal topics and certain emotional contexts. Hyeju Jang, Yohan Jo, Qinlan Shen, Michael Miller Yoder, Seungwhan Moon, Carolyn P. Rosé |
ACL (1) | 2 |
| 2016 | Expediting Support for Social Learning with Behavior Modeling
Yohan Jo, Gaurav Tomar, Oliver Ferschke, Carolyn P. Rosé, Dragan Gasevic |
EDM | 1 |
| 2016 | Pipeline for expediting learning analytics and student support from data in social learningabstractAn important research problem in learning analytics is to expedite the cycle of data leading to the analysis of student progress and the improvement of student support. For this goal in the context of social learning, we propose a pipeline that includes data infrastructure, learning analytics, and intervention, along with computational models for individual components. Next, we describe an example of applying this pipeline to real data in a case study, whose goal is to investigate the positive effects that goal-setting students have on their peers, which suggests ways in which we might foster these social benefits through intervention. Yohan Jo, Gaurav Tomar, Oliver Ferschke, Carolyn P. Rosé, Dragan Gasevic |
LAK | 1 |
| 2015 | Time Series Analysis of Nursing Notes for Mortality Prediction via a State Transition Topic ModelabstractAccurate mortality prediction is an important task in intensive care units in order to channel prompt care to patients in the most critical condition and to reduce nurses' alarm fatigue. Nursing notes carry valuable information in this regard, but nothing has been reported about the effectiveness of temporal analysis of nursing notes in mortality prediction tasks. Yohan Jo, Natasha Loghmanpour, Carolyn P. Rosé |
CIKM | 1 |
| 2015 | Metaphor Detection in DiscourseabstractUnderstanding contextual information is key to detecting metaphors in discourse.Most current work aims at detecting metaphors given a single sentence, thus focusing mostly on local contextual cues within a short text.In this paper, we present a novel approach that explicitly leverages global context of a discourse to detect metaphors.In addition, we show that syntactic information such as dependency structures can help better describe local contextual information, thus improving detection results when combined.We apply our methods on a newly annotated online discussion forum, and show that our approach outperforms the state-of-the-art baselines in previous literature. Hyeju Jang, Seungwhan Moon, Yohan Jo, Carolyn P. Rosé |
SIGDIAL Conference | 3 |
| 2011 | Aspect and sentiment unification model for online review analysisabstractUser-generated reviews on the Web contain sentiments about detailed aspects of products and services. However, most of the reviews are plain text and thus require much effort to obtain information about relevant details. In this paper, we tackle the problem of automatically discovering what aspects are evaluated in reviews and how sentiments for different aspects are expressed. We first propose Sentence-LDA (SLDA), a probabilistic generative model that assumes all words in a single sentence are generated from one aspect. We then extend SLDA to Aspect and Sentiment Unification Model (ASUM), which incorporates aspect and sentiment together to model sentiments toward different aspects. ASUM discovers pairs of {aspect, sentiment} which we call senti-aspects. We applied SLDA and ASUM to reviews of electronic devices and restaurants. The results show that the aspects discovered by SLDA match evaluative details of the reviews, and the senti-aspects found by ASUM capture important aspects that are closely coupled with a sentiment. The results of sentiment classification show that ASUM outperforms other generative models and comes close to supervised classification methods. One important advantage of ASUM is that it does not require any sentiment labels of the reviews, which are often expensive to obtain. Yohan Jo, Alice Oh |
WSDM | 1 |