EDBT 2026 Demo / reviewers in the wild / expert
Joonsuk Park
dblp:50/9717
· DBLP profile ↗
26ranked-venue papers
7as first author
15since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 4 first-author · 14 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorComputer networks · 1Databases, data management, data science and information retrieval · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ReSCORE: Label-free Iterative Retriever Training for Multi-hop Question Answering with Relevance-Consistency SupervisionabstractDosung Lee, Wonjun Oh, Boyoung Kim, Minyoung Kim, Joonsuk Park, Paul Hongsuck Seo. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Dosung Lee, Wonjun Oh, Joonsuk Park, Hongsuck Seo |
ACL (1) | 5 |
| 2025 | Return of EM: Entity-driven Answer Set Expansion for QA EvaluationabstractRecently, directly using large language models (LLMs) has been shown to be the most reliable method to evaluate QA models. However, it suffers from limited interpretability, high cost, and environmental harm. To address these, we propose to use soft exact match (EM) with entity-driven answer set expansion. Our approach expands the gold answer set to include diverse surface forms, based on the observation that the surface forms often follow particular patterns depending on the entity type. The experimental results show that our method outperforms traditional evaluation methods by a large margin. Moreover, the reliability of our evaluation method is comparable to that of LLM-based ones, while offering the benefits of high interpretability and reduced environmental harm. Dongryeol Lee, Minwoo Lee 0003, Kyungmin Min, Joonsuk Park, Kyomin Jung |
COLING | 4 |
| 2025 | AdvisorQA: Towards Helpful and Harmless Advice-seeking Question Answering with Collective IntelligenceabstractMinbeom Kim, Hwanhee Lee, Joonsuk Park, Hwaran Lee, Kyomin Jung. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Minbeom Kim, Hwanhee Lee, Joonsuk Park, Hwaran Lee, Kyomin Jung |
NAACL (Long Papers) | 3 |
| 2025 | Are LLM-Judges Robust to Expressions of Uncertainty? Investigating the effect of Epistemic Markers on LLM-based EvaluationabstractDongryeol Lee, Yerin Hwang, Yongil Kim, Joonsuk Park, Kyomin Jung. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dongryeol Lee, Yerin Hwang, Yongil Kim, Joonsuk Park, Kyomin Jung |
NAACL (Long Papers) | 4 |
| 2025 | tRAG: Term-level Retrieval-Augmented Generation for Domain-Adaptive RetrievalabstractDohyeon Lee, Jongyoon Kim, Jihyuk Kim, Seung-won Hwang, Joonsuk Park. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Dohyeon Lee, Jongyoon Kim, Jihyuk Kim, Seung-won Hwang, Joonsuk Park |
NAACL (Long Papers) | 5 |
| 2024 | A Design Space for Intelligent and Interactive Writing AssistantsabstractIn our era of rapid technological advancement, the research landscape for writing assistants has become increasingly fragmented across various research communities. We seek to address this challenge by proposing a design space as a structured way to examine and explore the multidimensional space of intelligent and interactive writing assistants. Through community collaboration, we explore five aspects of writing assistants: task, user, technology, interaction, and ecosystem. Within each aspect, we define dimensions and codes by systematically reviewing 115 papers, while leveraging the expertise of researchers in various disciplines. Our design space aims to offer researchers and designers a practical tool to navigate, comprehend, and compare the various possibilities of writing assistants, and aid in the design of new writing assistants. Mina Lee 0002, Katy Ilonka Gero, John Joon Young Chung, Simon Buckingham Shum, Vipul Raheja, Hua Shen 0005, Subhashini Venugopalan, Thiemo Wambsganss, David Zhou, Emad A. Alghamdi, Tal August, Avinash Bhat, Madiha Zahrah Choksi, Senjuti Dutta, Jin L. C. Guo, Md. Naimul Hoque, Simon Knight 0001, Seyed Parsa Neshaei, Antonette Shibani, Disha Shrivastava, Lila Shroff, Agnia Sergeyuk, Jessi Stark, Sarah Sterman, Sitong Wang 0001, Antoine Bosselut, Daniel Buschek, Joseph Chee Chang, Sherol Chen, Max Kreminski, Joonsuk Park, Roy D. Pea, Eugenia Ha Rim Rho, Shannon Shen 0001, Pao Siangliulue |
CHI | 32 |
| 2024 | Argument Quality Assessment in the Age of Instruction-Following Large Language ModelsabstractThe computational treatment of arguments on controversial issues has been subject to extensive NLP research, due to its envisioned impact on opinion formation, decision making, writing education, and the like. A critical task in any such application is the assessment of an argument’s quality - but it is also particularly challenging. In this position paper, we start from a brief survey of argument quality research, where we identify the diversity of quality notions and the subjectiveness of their perception as the main hurdles towards substantial progress on argument quality assessment. We argue that the capabilities of instruction-following large language models (LLMs) to leverage knowledge across contexts enable a much more reliable assessment. Rather than just fine-tuning LLMs towards leaderboard chasing on assessment tasks, they need to be instructed systematically with argumentation theories and scenarios as well as with ways to solve argument-related problems. We discuss the real-world opportunities and ethical issues emerging thereby. Henning Wachsmuth, Gabriella Lapesa, Elena Cabrio, Anne Lauscher, Joonsuk Park, Eva Maria Vecchi, Serena Villata, Timon Ziegenbein |
LREC/COLING | 5 |
| 2024 | Hierarchical Deconstruction of LLM Reasoning: A Graph-Based Framework for Analyzing Knowledge UtilizationabstractDespite the advances in large language models (LLMs), how they use their knowledge for reasoning is not yet well understood.In this study, we propose a method that deconstructs complex real-world questions into a graph, representing each question as a node with predecessors of background knowledge needed to solve the question.We develop the DEPTHQA dataset, deconstructing questions into three depths: (i) recalling conceptual knowledge, (ii) applying procedural knowledge, and (iii) analyzing strategic knowledge.Based on a hierarchical graph, we quantify forward discrepancy, discrepancies in LLMs' performance on simpler sub-problems versus complex questions.We also measure backward discrepancy, where LLMs answer complex questions but struggle with simpler ones.Our analysis shows that smaller models exhibit more discrepancies than larger models.Distinct patterns of discrepancies are observed across model capacity and possibility of training data memorization.Additionally, guiding models from simpler to complex questions through multiturn interactions improves performance across model sizes, highlighting the importance of structured intermediate steps in knowledge reasoning.This work enhances our understanding of LLM reasoning and suggests ways to improve their problem-solving abilities. Miyoung Ko, Sue Hyun Park, Joonsuk Park, Minjoon Seo |
EMNLP | 3 |
| 2023 | SQuARe: A Large-Scale Dataset of Sensitive Questions and Acceptable Responses Created through Human-Machine CollaborationabstractHwaran Lee, Seokhee Hong, Joonsuk Park, Takyoung Kim, Meeyoung Cha, Yejin Choi, Byoungpil Kim, Gunhee Kim, Eun-Ju Lee, Yong Lim, Alice Oh, Sangchul Park, Jung-Woo Ha. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Hwaran Lee, Seokhee Hong 0002, Joonsuk Park, Takyoung Kim, Meeyoung Cha, Yejin Choi 0001, Byoung Pil Kim, Gunhee Kim, Eun-Ju Lee 0001, Yong Lim, Alice Oh, Sangchul Park, Jung-Woo Ha 0001 |
ACL (1) | 3 |
| 2023 | From Values to Opinions: Predicting Human Behaviors and Stances Using Value-Injected Large Language ModelsabstractBeing able to predict people's opinions on issues and behaviors in realistic scenarios can be helpful in various domains, such as politics and marketing.However, conducting largescale surveys like the European Social Survey to solicit people's opinions on individual issues can incur prohibitive costs.Leveraging prior research showing influence of core human values on individual decisions and actions, we propose to use value-injected large language models (LLM) to predict opinions and behaviors.To this end, we present Value Injection Method (VIM), a collection of two methodsargument generation and question answeringdesigned to inject targeted value distributions into LLMs via fine-tuning.We then conduct a series of experiments on four tasks to test the effectiveness of VIM and the possibility of using value-injected LLMs to predict opinions and behaviors of people.We find that LLMs valueinjected with variations of VIM substantially outperform the baselines.Also, the results suggest that opinions and behaviors can be better predicted using value-injected LLMs than the baseline approaches. Dongjun Kang, Joonsuk Park, Yohan Jo, JinYeong Bak |
EMNLP | 2 |
| 2023 | Tree of Clarifications: Answering Ambiguous Questions with Retrieval-Augmented Large Language ModelsabstractQuestions in open-domain question answering are often ambiguous, allowing multiple interpretations.One approach to handling them is to identify all possible interpretations of the ambiguous question (AQ) and to generate a long-form answer addressing them all, as suggested by Stelmakh et al. (2022).While it provides a comprehensive response without bothering the user for clarification, considering multiple dimensions of ambiguity and gathering corresponding knowledge remains a challenge.To cope with the challenge, we propose a novel framework, TREE OF CLARIFICATIONS (TOC): It recursively constructs a tree of disambiguations for the AQ-via few-shot prompting leveraging external knowledge-and uses it to generate a long-form answer.TOC outperforms existing baselines on ASQA in a fewshot setup across all metrics, while surpassing fully-supervised baselines trained on the whole training set in terms of Disambig-F1 and Disambig-ROUGE.Code is available at github.com/gankim/tree-of-clarifications. Gangwoo Kim, Sungdong Kim, Byeongguk Jeon, Joonsuk Park, Jaewoo Kang |
EMNLP | 4 |
| 2023 | mRedditSum: A Multimodal Abstractive Summarization Dataset of Reddit Threads with ImagesabstractThe growing number of multimodal online discussions necessitates automatic summarization to save time and reduce content overload.However, existing summarization datasets are not suitable for this purpose, as they either do not cover discussions, multiple modalities, or both.To this end, we present MREDDITSUM, the first multimodal discussion summarization dataset.It consists of 3,033 discussion threads where a post solicits advice regarding an issue described with an image and text, and respective comments express diverse opinions.We annotate each thread with a human-written summary that captures both the essential information from the text, as well as the details available only in the image.Experiments show that popular summarization models-GPT-3.5,BART, and T5-consistently improve in performance when visual information is incorporated.We also introduce a novel method, cluster-based multi-stage summarization, that outperforms existing baselines and serves as a competitive baseline for future work. * Equal contribution. Keighley Overbay, Jaewoo Ahn, Fatemeh Pesaran Zadeh, Joonsuk Park, Gunhee Kim |
EMNLP | 4 |
| 2023 | Memory-Efficient Fine-Tuning of Compressed Large Language Models via sub-4-bit Integer QuantizationabstractLarge language models (LLMs) face the challenges in fine-tuning and deployment due to their high memory demands and computational costs. While parameter-efficient fine-tuning (PEFT) methods aim to reduce the memory usage of the optimizer state during fine-tuning, the inherent size of pre-trained LLM weights continues to be a pressing concern.
Even though quantization techniques are widely proposed to ease memory demands and accelerate LLM inference, most of these techniques are geared towards the deployment phase.
To bridge this gap, this paper presents Parameter-Efficient and Quantization-aware Adaptation (PEQA) – a simple yet effective method that combines the advantages of PEFT with quantized LLMs.
By updating solely the quantization scales, PEQA can be directly applied to quantized LLMs, ensuring seamless task transitions. Parallel to existing PEFT methods, PEQA significantly reduces the memory overhead associated with the optimizer state. Furthermore, it leverages the advantages of quantization to substantially reduce model sizes. Even after fine-tuning, the quantization structure of a PEQA-tuned LLM remains intact, allowing for accelerated inference on the deployment stage.
We employ PEQA-tuning for task-specific adaptation on LLMs with up to $65$ billion parameters. To assess the logical reasoning and language comprehension of PEQA-tuned LLMs, we fine-tune low-bit quantized LLMs using a instruction dataset.
Our results show that even when LLMs are quantized to below 4-bit precision, their capabilities in language modeling, few-shot in-context learning, and comprehension can be resiliently restored to (or even improved over) their full-precision original performances with PEQA. Jung Hyun Lee, Sungdong Kim, Joonsuk Park, Kang Min Yoo, Se Jung Kwon, Dongsoo Lee |
NeurIPS | 4 |
| 2022 | Argument Mining for Review Helpfulness PredictionabstractThe importance of reliably determining the helpfulness of product reviews is rising as both helpful and unhelpful reviews continue to accumulate on e-commerce websites.And argumentational features-such as the structure of arguments and the types of underlying elementary units-have shown to be promising indicators of product review helpfulness.However, their adoption has been limited due to the lack of sufficient resources and large-scale experiments investigating their utility.To this end, we present the AMazon Argument Mining (AM 2 ) corpus-a corpus of 878 Amazon reviews on headphones annotated according to a theoretical argumentation model designed to evaluate argument quality.Experiments show that employing argumentational features leads to statistically significant improvements over the state-of-the-art review helpfulness predictors under both text-only and text-and-image settings.1 Zaiqian Chen, Daniel Verdi do Amarante, Jenna Donaldson, Yohan Jo, Joonsuk Park |
EMNLP | 5 |
| 2021 | Analyzing Cultural Assimilation through the Lens of Yelp Restaurant ReviewsabstractGiven the steady stream of immigrants from around the world, cultural assimilation in North America has long been a topic of interest. However, existing research focuses only on assimilation to North American culture, overlooking the mutual influence, with a very limited use of data-driven approaches. In this paper, we investigate assimilation among various cultures in North America through the lens of discussions surrounding food. We first present Cross-Cuisine Cross-Region LDA (c3rLDA), a novel probabilistic graphical model to jointly uncover latent topics shared across cuisines, as well as their regional variants for each cuisine. Then, we employ the model on 3.7 million Yelp restaurant reviews to find that cuisines assimilate to one another in varying degrees depending on the cuisines involved, the topic, and the region: A cuisine tends to be more influenced by other cuisines if it is regularly fused with others (e.g. Japanese), for certain topics (e.g. breakfast and dessert), and in specific regions (e.g. stronger Mexican influence in the Southwestern US and French influence in the East Canada). Lastly, we demonstrate that the topics generated by our model, on which the qualitative analysis is based, are more coherent than or comparable to those generated by existing neural and non-neural topic models. This work represents the first step toward large-scale data-driven analysis of cultural assimilation in North America, which is made possible by the abundant data available in social media. Zaiqian Chen, Joonsuk Park |
DSAA | 2 |
| 2020 | Automated Fact-Checking of Claims from WikipediaabstractAutomated fact checking is becoming increasingly vital as both truthful and fallacious information accumulate online. Research on fact checking has benefited from large-scale datasets such as FEVER and SNLI. However, such datasets suffer from limited applicability due to the synthetic nature of claims and/or evidence written by annotators that differ from real claims and evidence on the internet. To this end, we present WikiFactCheck-English, a dataset of 124k+ triples consisting of a claim, context and an evidence document extracted from English Wikipedia articles and citations, as well as 34k+ manually written claims that are refuted by the evidence documents. This is the largest fact checking dataset consisting of real claims and evidence to date; it will allow the development of fact checking systems that can better process claims and evidence in the real world. We also show that for the NLI subtask, a logistic regression system trained using existing and novel features achieves peak accuracy of 68%, providing a competitive baseline for future work. Also, a decomposable attention model trained on SNLI significantly underperforms the models trained on this dataset, suggesting that models trained on manually generated data may not be sufficiently generalizable or suitable for fact checking real-world claims. Aalok Sathe, Salar Ather, Tuan Manh Le, Nathan Perry, Joonsuk Park |
LREC | 5 |
| 2018 | A Corpus of eRulemaking User Comments for Measuring Evaluability of Arguments
Joonsuk Park, Claire Cardie |
LREC | 1 |
| 2017 | Argument Mining with Structured SVMs and RNNsabstractWe propose a novel factor graph model for argument mining, designed for settings in which the argumentative relations in a document do not necessarily form a tree structure.(This is the case in over 20% of the web comments dataset we release.)Our model jointly learns elementary unit type classification and argumentative relation prediction.Moreover, our model supports SVM and RNN parametrizations, can enforce structure constraints (e.g., transitivity), and can express dependencies between adjacent relations and propositions.Our approaches outperform unstructured baselines in both web comments and argumentative essay datasets. Vlad Niculae, Joonsuk Park, Claire Cardie |
ACL (1) | 2 |
| 2017 | Using Argumentative Structure to Interpret Debates in Online Deliberative Democracy and eRulemakingabstractGovernments around the world are increasingly utilising online platforms and social media to engage with, and ascertain the opinions of, their citizens. Whilst policy makers could potentially benefit from such enormous feedback from society, they first face the challenge of making sense out of the large volumes of data produced. In this article, we show how the analysis of argumentative and dialogical structures allows for the principled identification of those issues that are central, controversial, or popular in an online corpus of debates. Although areas such as controversy mining work towards identifying issues that are a source of disagreement, by looking at the deeper argumentative structure, we show that a much richer understanding can be obtained. We provide results from using a pipeline of argument-mining techniques on the debate corpus, showing that the accuracy obtained is sufficient to automatically identify those issues that are key to the discussion, attracting proportionately more support than others, and those that are divisive, attracting proportionately more conflicting viewpoints. John Lawrence, Joonsuk Park, Katarzyna Budzynska, Claire Cardie, Barbara Konat, Chris Reed 0001 |
ACM Trans. Internet Techn. | 2 |
| 2016 | A Corpus of Argument Networks: Using Graph Properties to Analyse Divisive Issues
Barbara Konat, John Lawrence, Joonsuk Park, Katarzyna Budzynska, Chris Reed 0001 |
LREC | 3 |
| 2016 | The Effects of Peer- and Self-assessment on the AssessorsabstractRecently, there has been a growing interest in peer- and self-assessment (PSA) in the research community, especially with the development of massive open online courses (MOOCs). One prevalent theme in the literature is the consideration of PSA as a partial or full replacement for traditional assessments performed by the instructor. And since the traditional role of the students in assessment processes is the assessee, existing works on PSA typically focus on devising methods to make the grades more reliable and beneficial for the assessees. Joonsuk Park, Kimberley Williams |
SIGCSE | 1 |
| 2015 | Toward machine-assisted participation in eRulemaking: an argumentation model of evaluabilityabstracteRulemaking is an ongoing effort to use online tools to foster broader and better public participation in rulemaking --- the multi-step process that federal agencies use to develop new health, safety, and economic regulations. The increasing participation of non-expert citizens, however, has led to a growth in the amount of arguments whose validity or strength are difficult to evaluate, both by the government agencies and fellow citizens. Such arguments typically neglect to provide the reasons for the conclusions and objective evidence for factual claims upon which the arguments are based. In this paper, we propose a novel argumentation model for capturing the evaluability of user comments in eRulemaking. This model is intended to be used for implementing automated systems to assist users in constructing evaluable arguments under online commenting environment for the benefit of quick feedback at a low cost. Joonsuk Park, Cheryl Blake, Claire Cardie |
ICAIL | 1 |
| 2014 | AsseSS: A Tool for Assessing the Support Structures of Arguments in User CommentsabstractWe present AsseSS, a tool for identifying and assessing the support structures of arguments in user comments. Given a comment, the system first classifies elementary units of arguments comprising the comment based on the type of appropriate support. Then, it detects support relations among the elementary units. With this information, it is possible to decide whether the existing support relation is of suitable type. Also, in the case that no support has been provided for an elementary unit, an appropriate type of support can be determined. Joonsuk Park, Claire Cardie |
COMMA | 1 |
| 2012 | Digital map based pose improvement for outdoor Augmented RealityabstractWith popularization of smart phones, needs for location based services (LBS), which is one of the most promising Augmented Reality applications, increased rapidly. However, accuracy of most commercially available Global Positioning Systems (GPS) is below levels for providing practically meaningful location based information. Especially when there are high building structures nearby, GPS location measurements are known to be erroneous and deviant. In this paper, we present a computer vision based method for improving user's position and orientation for outdoor Augmented Reality with initial values obtained from a GPS and a digital compass. Given a digital map, our goal was to determine corresponding buildings visible in the camera image and improve the user location and orientation. In average, our method improved (14.4m, 3.3m) in position and 2.8 degrees in orientation. Our method is suitable for mobile services in urban environments where tall buildings degrade GPS signals. Joonsuk Park, Jun Park |
ISMAR | 1 |
| 2012 | Improving Implicit Discourse Relation Recognition Through Feature Set Optimization
Joonsuk Park, Claire Cardie |
SIGDIAL Conference | 1 |
| 2010 | 3DOF tracking accuracy improvement for outdoor Augmented RealityabstractOutdoor Augmented Reality (AR) gained popularity recently due to its potential for location based mobile services. However, most commercially available Global Positioning Systems (GPS), except for the expensive high-end models, do not provide accurate location information that is enough to be used for displaying practically meaningful location based information. In this paper, we present a computer vision based method for improving user's two dimensional location and one-dimensional orientation, the initial values of which are obtained from a GPS and a digital compass. Our method utilizes corner positions of buildings in the map and the vertical edges of the buildings in the captured images. We applied anisotropic diffusion in order to filter noise and preserve edges, and dual vertical edge filters on gray and saturation images. Our method is suitable for mobile services in urban environments where tall buildings degrade GPS signals. In average, our method improved 15.0 meters in position and 2.2 degrees in orientation. Joonsuk Park, Jun Park |
ISMAR | 1 |