EDBT 2026 Demo / reviewers in the wild / expert
Jong C. Park
dblp:73/5376
· DBLP profile ↗
29ranked-venue papers
3as first author
15since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 3 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Social Dynamics as Critical Vulnerabilities that Undermine Objective Decision-Making in LLM CollectivesabstractLarge language model (LLM) agents are increasingly acting as human delegates in multiagent environments, where a representative agent integrates diverse peer perspectives to make a final decision.Drawing inspiration from social psychology, we investigate how the reliability of this representative agent is undermined by the social context of its network.We define four key phenomena-social conformity, perceived expertise, dominant speaker effect, and rhetorical persuasion-and systematically manipulate the number of adversaries, relative intelligence, argument length, and argumentative styles.Our experiments demonstrate that the representative agent's accuracy consistently declines as social pressure increases: larger adversarial groups, more capable peers, and longer arguments all lead to significant performance degradation.Furthermore, rhetorical strategies emphasizing credibility or logic can further sway the agent's judgment, depending on the context.These findings reveal that multiagent systems are sensitive not only to individual reasoning but also to the social dynamics of their configuration, highlighting critical vulnerabilities in AI delegates that mirror the psychological biases observed in human group decision-making. Changgeon Ko, Jisu Shin 0001, Hoyun Song, Huije Lee, Eui Jun Hwang, Jong C. Park |
ACL (1) | 6 |
| 2026 | Beyond Static Benchmarks: Synthesizing Harmful Content via Persona-based Simulation for Robust EvaluationabstractStatic benchmarks for harmful content detection face limitations in scalability and diversity, and may also be affected by contamination from web-scale pre-training corpora.To address these issues, we propose a framework for synthesizing harmful content, leveraging persona-guided large language model (LLM) agents.Our approach constructs twodimensional user personas by integrating demographic identities and topical interests with situational harmful strategies, enabling the simulation of diverse and contextually grounded harmful interactions.We evaluate the framework along three dimensions: harmfulness, challenge level, and diversity.Both human and LLM-based evaluations confirm that our framework achieves a high harmful generation success rate.Experiments across multiple detection systems reveal that our synthetic scenarios are more challenging to detect than those in existing benchmarks.Furthermore, a multifaceted analysis confirms that our approach achieves linguistic and topical diversity comparable to human-curated datasets, establishing our framework as an effective tool for robust stress-testing of harmful content detection systems 1 . Huije Lee, Jisu Shin 0001, Hoyun Song, Changgeon Ko, Jong C. Park |
ACL (1) | 5 |
| 2025 | Database-Augmented Query Representation for Information RetrievalabstractInformation retrieval models that aim to search for documents relevant to a query have shown multiple successes, which have been applied to diverse tasks.Yet, the query from the user is oftentimes short, which challenges the retrievers to correctly fetch relevant documents.To tackle this, previous studies have proposed expanding the query with a couple of additional (userrelated) features related to it.However, they may be suboptimal to effectively augment the query, and there is plenty of other information available to augment it in a relational database.Motivated by this fact, we present a novel retrieval framework called Database-Augmented Query representation (DAQu), which augments the original query with various (query-related) metadata across multiple tables.In addition, as the number of features in the metadata can be very large and there is no order among them, we encode them with the graph-based set-encoding strategy, which considers hierarchies of features in the database without order.We validate our DAQu in diverse retrieval scenarios, demonstrating that it significantly enhances overall retrieval performance over relevant baselines. Soyeong Jeong, Jinheon Baek, Sukmin Cho, Sung Ju Hwang, Jong C. Park |
EMNLP | 5 |
| 2025 | An Efficient Gloss-Free Sign Language Translation Using Spatial Configurations and Motion Dynamics with LLMsabstractEui Jun Hwang, Sukmin Cho, Junmyeong Lee, Jong C. Park. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Eui Jun Hwang, Sukmin Cho, Junmyeong Lee, Jong C. Park |
NAACL (Long Papers) | 4 |
| 2025 | A Multi-Task Benchmark for Abusive Language Detection in Low-Resource SettingsabstractContent moderation research has recently made significant advances, but remains limited in serving the majority of the world's languages due to the lack of resources, leaving millions of vulnerable users to online hostility. This work presents a large-scale human-annotated multi-task benchmark dataset for abusive language detection in Tigrinya social media with joint annotations for three tasks: abusiveness, sentiment, and topic classification. The dataset comprises 13,717 YouTube comments annotated by nine native speakers, collected from 7,373 videos with a total of over 1.2 billion views across 51 channels. We developed an iterative term clustering approach for effective data selection. Recognizing that around 64% of Tigrinya social media content uses Romanized transliterations rather than native Ge'ez script, our dataset accommodates both writing systems to reflect actual language use. We establish strong baselines across the tasks in the benchmark, while leaving significant challenges for future contributions. Our experiments demonstrate that small fine-tuned models outperform prompted frontier large language models (LLMs) in the low-resource setting, achieving 86.67% F1 in abusiveness detection (7+ points over best LLM), and maintain stronger performance in all other tasks. The benchmark is made public to promote research on online safety. Fitsum Gaim Gebre, Hoyun Song, Huije Lee, Changgeon Ko, Eui Jun Hwang, Jong C. Park |
NeurIPS | 6 |
| 2025 | A Spatio-Temporal Representation Learning as an Alternative to Traditional Glosses in Sign Language Translation and ProductionabstractThis work addresses the challenges associated with the use of glosses in both Sign Language Translation (SLT) and Sign Language Production (SLP). While glosses have long been used as a bridge between sign language and spoken language, they come with two major limitations that impede the advancement of sign language systems. First, annotating the glosses is a labor-intensive and time-consuming process, which limits the scalability of datasets. Second, the glosses oversimplify sign language by stripping away its spatio-temporal dynamics, reducing complex signs to basic labels and missing the subtle movements essential for precise interpretation. To address these limitations, we introduce Universal Gloss-level Representation (UniGloR), a framework designed to capture the spatio-temporal features inherent in sign language, providing a more dynamic and detailed alternative to the use of the glosses. The core idea of UniGloR is simple yet effective: We derive dense spatiotemporal representations from sign keypoint sequences using self-supervised learning and seamlessly integrate them into SLT and SLP tasks. Our experiments in a keypoint-based setting demonstrate that UniGloR either outperforms or matches the performance of previous SLT and SLP methods on two widely-used datasets: PHOENIX14T and How2Sign. Code is available at https://github.com/eddie-euijun-hwang/UniGloR. Eui Jun Hwang, Sukmin Cho, Huije Lee, Youngwoo Yoon, Jong C. Park |
WACV | 5 |
| 2024 | A Gloss-Free Sign Language Production with Discrete RepresentationabstractGloss-free Sign Language Production (SLP) offers a direct translation of spoken language sentences into sign language, bypassing the need for gloss intermediaries. Previous autoregressive SLP methods have not fully achieved true autoregression, as they often depend on ground-truth data during inference. To fill this gap, we introduce Sign language Vector Quantization Network (SignVQNet), leveraging discrete spatio-temporal representations of sign poses. With such a discrete representation, our method incorporates beam search, a decoding strategy widely used in Natural Language Processing. Furthermore, we align the discrete representation with linguistic features from pre-trained language models such as BERT. Our results show the superior performance of our method over prior SLP methods in generating accurate and realistic sign pose sequences. Additionally, our analysis shows that the reliability of Back-Translation and Fréchet Gesture Distance as evaluation metrics, in contrast to DTW-MJE. The code and models are available at https://github.com/eddie-euijun-hwang/SignVQNet. Eui Jun Hwang, Huije Lee, Jong C. Park |
FG | 3 |
| 2023 | Deep Model Compression Also Helps Models Capture AmbiguityabstractNatural language understanding (NLU) tasks face a non-trivial amount of ambiguous samples where veracity of their labels is debatable among annotators.NLU models should thus account for such ambiguity, but they approximate the human opinion distributions quite poorly and tend to produce over-confident predictions.To address this problem, we must consider how to exactly capture the degree of relationship between each sample and its candidate classes.In this work, we propose a novel method with deep model compression and show how such relationship can be accounted for.We see that more reasonably represented relationships can be discovered in the lower layers and that validation accuracies are converging at these layers, which naturally leads to layer pruning.We also see that distilling the relationship knowledge from a lower layer helps models produce better distribution.Experimental results demonstrate that our method makes substantial improvement on quantifying ambiguity without gold distribution labels.As positive side-effects, our method is found to reduce the model size significantly and improve latency, both attractive aspects of NLU products.1 Premise It's summer time and two girls play with bubbles near a boat dock.Hypothesis It is warm outside.Label distribution Entailment: 0. Hancheol Park, Jong C. Park |
ACL (1) | 2 |
| 2023 | Knowledge-Augmented Language Model VerificationabstractRecent Language Models (LMs) have shown impressive capabilities in generating texts with the knowledge internalized in parameters.Yet, LMs often generate the factually incorrect responses to the given queries, since their knowledge may be inaccurate, incomplete, and outdated.To address this problem, previous works propose to augment LMs with the knowledge retrieved from an external knowledge source.However, such approaches often show suboptimal text generation performance due to two reasons: 1) the model may fail to retrieve the knowledge relevant to the given query, or 2) the model may not faithfully reflect the retrieved knowledge in the generated text.To overcome these, we propose to verify the output and the knowledge of the knowledge-augmented LMs with a separate verifier, which is a small LM that is trained to detect those two types of errors through instruction-finetuning.Then, when the verifier recognizes an error, we can rectify it by either retrieving new knowledge or generating new text.Further, we use an ensemble of the outputs from different instructions with a single verifier to enhance the reliability of the verification processes.We validate the effectiveness of the proposed verification steps on multiple question answering benchmarks, whose results show that the proposed verifier effectively identifies retrieval and generation errors, allowing LMs to provide more factually correct outputs.Our code is available at https://github.com/JinheonBaek/KALMV. Jinheon Baek, Soyeong Jeong, Minki Kang, Jong C. Park, Sung Ju Hwang |
EMNLP | 4 |
| 2022 | Sign Language Production With Avatar Layering: A Critical Use Case over Rare WordsabstractSign language production (SLP) is the process of generating sign language videos from spoken language expressions. Since sign languages are highly under-resourced, existing vision-based SLP approaches suffer from out-of-vocabulary (OOV) and test-time generalization problems and thus generate low-quality translations. To address these problems, we introduce an avatar-based SLP system composed of a sign language translation (SLT) model and an avatar animation generation module. Our Transformer-based SLT model utilizes two additional strategies to resolve these problems: named entity transformation to reduce OOV tokens and context vector generation using a pretrained language model (e.g., BERT) to reliably train the decoder. Our system is validated on a new Korean-Korean Sign Language (KSL) dataset of weather forecasts and emergency announcements. Our SLT model achieves an 8.77 higher BLEU-4 score and a 4.57 higher ROUGE-L score over those of our baseline model. In a user evaluation, 93.48% of named entities were successfully identified by participants, demonstrating marked improvement on OOV issues. Jung-Ho Kim 0002, Eui Jun Hwang, Sukmin Cho, Du Hui Lee, Jong C. Park |
LREC | 5 |
| 2022 | GeezSwitch: Language Identification in Typologically Related Low-resourced East African LanguagesabstractLanguage identification is one of the fundamental tasks in natural language processing that is a prerequisite to data processing and numerous applications. Low-resourced languages with similar typologies are generally confused with each other in real-world applications such as machine translation, affecting the user’s experience. In this work, we present a language identification dataset for five typologically and phylogenetically related low-resourced East African languages that use the Ge’ez script as a writing system; namely Amharic, Blin, Ge’ez, Tigre, and Tigrinya. The dataset is built automatically from selected data sources, but we also performed a manual evaluation to assess its quality. Our approach to constructing the dataset is cost-effective and applicable to other low-resource languages. We integrated the dataset into an existing language-identification tool and also fine-tuned several Transformer based language models, achieving very strong results in all cases. While the task of language identification is easy for the informed person, such datasets can make a difference in real-world deployments and also serve as part of a benchmark for language understanding in the target languages. The data and models are made available at https://github.com/fgaim/geezswitch. Fitsum Gaim Gebre, Wonsuk Yang, Jong C. Park |
LREC | 3 |
| 2022 | ELF22: A Context-based Counter Trolling Dataset to Combat Internet TrollsabstractOnline trolls increase social costs and cause psychological damage to individuals. With the proliferation of automated accounts making use of bots for trolling, it is difficult for targeted individual users to handle the situation both quantitatively and qualitatively. To address this issue, we focus on automating the method to counter trolls, as counter responses to combat trolls encourage community users to maintain ongoing discussion without compromising freedom of expression. For this purpose, we propose a novel dataset for automatic counter response generation. In particular, we constructed a pair-wise dataset that includes troll comments and counter responses with labeled response strategies, which enables models fine-tuned on our dataset to generate responses by varying counter responses according to the specified strategy. We conducted three tasks to assess the effectiveness of our dataset and evaluated the results through both automatic and human evaluation. In human evaluation, we demonstrate that the model fine-tuned with our dataset shows a significantly improved performance in strategy-controlled sentence generation. Huije Lee, Young Ju Na, Hoyun Song, Jisu Shin 0001, Jong C. Park |
LREC | 5 |
| 2021 | Non-Autoregressive Sign Language Production with Gaussian Space
Eui Jun Hwang, Jung-Ho Kim 0002, Jong C. Park |
BMVC | 3 |
| 2021 | A Large-scale Comprehensive Abusiveness Detection Dataset with Multifaceted Labels from RedditabstractAs users in online communities suffer from severe side effects of abusive language, many researchers attempted to detect abusive texts from social media, presenting several datasets for such detection.However, none of them contain both comprehensive labels and contextual information, which are essential for thoroughly detecting all kinds of abusiveness from texts, since datasets with such fine-grained features demand a significant amount of annotations, leading to much increased complexity.In this paper, we propose a Comprehensive Abusiveness Detection Dataset (CADD), collected from the English Reddit posts, with multifaceted labels and contexts.Our dataset is annotated hierarchically for an efficient annotation through crowdsourcing on a large-scale.We also empirically explore the characteristics of our dataset and provide a detailed analysis for novel insights.The results of our experiments with strong pre-trained natural language understanding models on our dataset show that our dataset gives rise to meaningful performance, assuring its practicality for abusive language detection. Hoyun Song, Soo Hyun Ryu, Huije Lee, Jong C. Park |
CoNLL | 4 |
| 2021 | Optimizing Domain Specificity of Transformer-based Language Models for Extractive Summarization of Financial News Articles in Korean
Huije Lee, Wonsuk Yang, Chaehun Park, Hoyun Song, Eugene Jang, Jong C. Park |
PACLIC | 6 |
| 2018 | Feature Attention Network: Interpretable Depression Detection from Social Media
Hoyun Song, Jinseon You, Jin-Woo Chung, Jong C. Park |
PACLIC | 4 |
| 2017 | Extraction of Gene-Environment Interaction from the Biomedical LiteratureabstractGenetic information in the literature has been extensively looked into for the purpose of discovering the etiology of a disease. As the gene-disease relation is sensitive to external factors, their identification is important to study a disease. Environmental influences, which are usually called Gene-Environment interaction (GxE), have been considered as important factors and have extensively been researched in biology. Nevertheless, there is still a lack of systems for automatic GxE extraction from the biomedical literature due to new challenges: (1) there are no preprocessing tools and corpora for GxE, (2) expressions of GxE are often quite implicit, and (3) document-level comprehension is usually required. We propose to overcome these challenges with neural network models and show that a modified sequence-to-sequence model with a static RNN decoder produces a good performance in GxE recognition. Jinseon You, Jin-Woo Chung, Wonsuk Yang, Jong C. Park |
IJCNLP(1) | 4 |
| 2015 | Corpus annotation with a linguistic analysis of the associations between event mentions and spatial expressions
Jin-Woo Chung, Jinseon You, Jong C. Park |
PACLIC | 3 |
| 2012 | Product Name Classification for Product Instance Distinction
Hye-Jin Min, Jong C. Park |
PACLIC | 2 |
| 2011 | Detecting and Blocking False Sentiment Propagation
Hye-Jin Min, Jong C. Park |
IJCNLP | 2 |
| 2007 | Analysis of Indirect Uses of Interrogative Sentences Carrying Anger
Hye-Jin Min, Jong C. Park |
PACLIC | 2 |
| 2007 | Emotion Interaction System for a Service RobotabstractThis paper introduces an emotion interaction system for a service robot. The purpose of emotion interaction systems in service robots is to make people feel that the robot is not a mere machine, but reliable living assistant in the home. The emotion interaction system is composed of the emotion recognition, generation, and expression systems. A user's emotion is recognized by multi-modality, such as voice, dialogue, and touch. The robot's emotion is generated according to a psychological theory about emotion: OCC (Ortony, Clore, and Collins) model, which focuses on the user's emotional state and the information about environment and the robot itself. The generated emotion is expressed by facial expression, gesture, and the musical sound of the robot. Because the proposed system is composed of all the three components that are necessary for a full emotional interaction cycle, it can be implemented in the real robot system and be tested. Even though the multi- modality in emotion recognition and expression is still in its rudimentary stages, the proposed system is shown to be extremely useful in service robot applications. Furthermore, the proposed framework can be a cornerstone for the design of emotion interaction and generation systems for robots. Dong-Soo Kwon, Yoon Keun Kwak, Jong C. Park, Myung Jin Chung, Eun-Sook Jee, Kyung-Sook Park, Hyoung-Rock Kim, Jong-Chan Park, Eun Ho Kim, Kyung Hak Hyun, Hye-Jin Min, Hui Sung Lee, Jeong Woo Park, Su Hun Jo, Soon-Young Park, Kyung-Won Lee |
RO-MAN | 3 |
| 2005 | From Text to Sign Language: Exploiting the Spatial and Motioning Dimension
Hee-Jin Lee, Jong C. Park |
PACLIC | 3 |
| 2005 | Vowel Sound Disambiguation for Intelligible Korean Speech Synthesis
Ho-Joon Lee, Jong C. Park |
PACLIC | 2 |
| 2004 | Introduction (Thematic Session: Text Mining in Biomedicine)
Sophia Ananiadou, Jong C. Park |
IJCNLP | 2 |
| 2002 | Natural Language Interpretations for Heterogeneous Database Access
Hodong Lee, Jong C. Park |
COLING | 2 |
| 2000 | Informed Parsing for Coordination with Combinatory Categorial Grammar
Jong C. Park, Hyung Joon Cho |
COLING | 1 |
| 1995 | Quantifier Scope and ConstituencyabstractTraditional approaches to quantifier scope typically need stipulation to exclude readings that are unavailable to human understanders. This paper shows that quantifier scope phenomena can be precisely characterized by a semantic representation constrained by surface constituency, if the distinction between referential and quantificational NPs is properly observed. A CCG implementation is described and compared to other approaches. Jong C. Park |
ACL | 1 |
| 1992 | A Unification-Based Semantic Interpretation for Coordinate ConstructsabstractThis paper shows that a first-order unification-based semantic interpretation for various coordinate constructs is possible without an explicit use of lambda expressions if we slightly modify the standard Montagovian semantics of coordination. This modification, along with partial execution, completely eliminates the lambda reduction steps during semantic interpretation. Jong C. Park |
ACL | 1 |