Juhwan Choi

dblp:174/4879 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation
abstract
Quality Estimation (QE) aims to assess machine translation quality without reference translations, but recent studies have shown that existing QE models exhibit systematic gender bias.In particular, they tend to favor masculine realizations in gender-ambiguous contexts and may assign higher scores to gendermisaligned translations even when gender is explicitly specified.To address these issues, we propose FairQE, a multi-agent-based, fairnessaware QE framework that mitigates gender bias in both gender-ambiguous and genderexplicit scenarios.FairQE detects gender cues, generates gender-flipped translation variants, and combines conventional QE scores with LLM-based bias-mitigating reasoning through a dynamic bias-aware aggregation mechanism.This design preserves the strengths of existing QE models while calibrating their genderrelated biases in a plug-and-play manner.Extensive experiments across multiple gender bias evaluation settings demonstrate that FairQE consistently improves gender fairness over strong QE baselines.Moreover, under MQMbased meta-evaluation following the WMT 2023 Metrics Shared Task, FairQE achieves competitive or improved general QE performance.These results show that gender bias in QE can be effectively mitigated without sacrificing evaluation accuracy, enabling fairer and more reliable translation evaluation.
Jinhee Jang, Juhwan Choi, Seunguk Yu
ACL (1)2
2025 SummPilot: Bridging Efficiency and Customization for Interactive Summarization System
abstract
This paper incorporates the efficiency of automatic summarization and addresses the challenge of generating personalized summaries tailored to individual users' interests and requirements. To tackle this challenge, we introduce SummPilot, an interaction-based customizable summarization system. SummPilot leverages a large language model to facilitate both automatic and interactive summarization. Users can engage with the system to understand document content and personalize summaries through interactive components such as semantic graphs, entity clustering, and explainable evaluation. Our demo and user studies demonstrate SummPilot's adaptability and usefulness for customizable summarization.
Jungmin Yun, Juhwan Choi, Kyohoon Jin, Soojin Jang, Jinhee Jang
AAAI2
2025 Plug-in and Fine-tuning: Bridging the Gap between Small Language Models and Large Language Models
abstract
Large language models (LLMs) are renowned for their extensive linguistic knowledge and strong generalization capabilities, but their high computational demands make them unsuitable for resource-constrained environments. In contrast, small language models (SLMs) are computationally efficient but often lack the broad generalization capacity of LLMs. To bridge this gap, we propose PiFi, a novel framework that combines the strengths of both LLMs and SLMs to achieve high performance while maintaining efficiency. PiFi integrates a single frozen layer from an LLM into a SLM and fine-tunes the combined model for specific tasks, boosting performance without a significant increase in computational cost. We show that PiFi delivers consistent performance improvements across a range of natural language processing tasks, including both natural language understanding and generation. Moreover, our findings demonstrate PiFi’s ability to effectively leverage LLM knowledge, enhancing generalization to unseen domains and facilitating the transfer of linguistic abilities.
Kyeonghyun Kim, Jinhee Jang, Juhwan Choi, Kyohoon Jin
ACL (1)3
2025 Delving into Multilingual Ethical Bias: The MSQAD with Statistical Hypothesis Tests for Large Language Models
abstract
Despite the recent strides in large language models, studies have underscored the existence of social biases within these systems.In this paper, we delve into the validation and comparison of the ethical biases of LLMs concerning globally discussed and potentially sensitive topics, hypothesizing that these biases may arise from language-specific distinctions.Introducing the Multilingual Sensitive Questions & Answers Dataset (MSQAD), we collected news articles from Human Rights Watch covering 17 topics, and generated socially sensitive questions along with corresponding responses in multiple languages.We scrutinize the biases of these responses across languages and topics, employing two statistical hypothesis tests.The results suggest that the null hypotheses are rejected in most cases, indicating biases arising from cross-language differences.It indicates that ethical biases in responses are widespread across various languages, and notably, these biases are prevalent even among different LLMs.By making the proposed MSQAD openly available, we aim to facilitate future research endeavors focused on examining cross-language biases in LLMs and their variant models 1 .
Seunguk Yu, Juhwan Choi
ACL (1)2
2025 CoBA: Counterbias Text Augmentation for Mitigating Various Spurious Correlations via Semantic Triples
abstract
Deep learning models often learn and exploit spurious correlations in training data, using these non-target features to inform their predictions.Such reliance leads to performance degradation and poor generalization on unseen data.To address these limitations, we introduce a more general form of counterfactual data augmentation, termed counterbias data augmentation, which simultaneously tackles multiple biases (e.g., gender bias, simplicity bias) and enhances out-of-distribution robustness.We present COBA: CounterBias Augmentation, a unified framework that operates at the semantic triple level: first decomposing text into subjectpredicate-object triples, then selectively modifying these triples to disrupt spurious correlations.By reconstructing the text from these adjusted triples, COBA generates counterbias data that mitigates spurious patterns.Through extensive experiments, we demonstrate that COBA not only improves downstream task performance, but also effectively reduces biases and strengthens out-of-distribution resilience, offering a versatile and robust solution to the challenges posed by spurious correlations.
Kyohoon Jin, Juhwan Choi, Jungmin Yun, Soojin Jang
EMNLP2
2025 See-Saw Modality Balance: See Gradient, and Sew Impaired Vision-Language Balance to Mitigate Dominant Modality Bias
abstract
Junehyoung Kwon, MiHyeon Kim, Eunju Lee, Juhwan Choi, YoungBin Kim. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Junehyoung Kwon, Mihyeon Kim, Eunju Lee 0003, Juhwan Choi
NAACL (Long Papers)4
2025 Korean football in-game conversation state tracking dataset for dialogue and turn level evaluation
Sangmin Song, Juhyoung Park, Juhwan Choi, Kyohoon Jin
Eng. Appl. Artif. Intell.3
2024 Enhancing Effectiveness and Robustness in a Low-Resource Regime via Decision-Boundary-aware Data Augmentation
abstract
Efforts to leverage deep learning models in low-resource regimes have led to numerous augmentation studies. However, the direct application of methods, such as mixup and cutout, is limited due to the discrete characteristics of the textual data. While methods using pre trained language models have exhibited good efficiency, they require additional considerations for robustness. Inspired by recent studies on decision boundaries, this paper proposes a decision-boundary-aware data augmentation strategy to enhance robustness using pretrained language models. The proposed technique first focuses on shifting the latent features closer to the decision boundary, followed by reconstruction to generate an ambiguous version with a soft label. Additionally, mid-K sampling is suggested to enhance the diversity of the generated sentences. This paper demonstrates the performance of the proposed augmentation strategy compared to other methods through extensive experiments. Furthermore, the ablation study demonstrates the effect of soft labels and mid-K sampling and the extensibility of the method with curriculum data augmentation.
Kyohoon Jin, Juhwan Choi, Sangmin Song
LREC/COLING3
2024 UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset Generation
abstract
Although pre-trained language models have exhibited great flexibility and versatility with prompt-based few-shot learning, they suffer from the extensive parameter size and limited applicability for inference.Recent studies have suggested that PLMs be used as dataset generators and a tiny task-specific model be trained to achieve efficient inference.However, their applicability to various domains is limited because they tend to generate domain-specific datasets.In this work, we propose a novel approach to universal domain generalization that generates a dataset regardless of the target domain.This allows for generalization of the tiny task model to any domain that shares the label space, thus enhancing the real-world applicability of the dataset generation paradigm.Our experiments indicate that the proposed method accomplishes generalizability across various domains while using a parameter set that is orders of magnitude smaller than PLMs.
Juhwan Choi, Yeonghwa Kim, Seunguk Yu, Jungmin Yun
EMNLP1
2024 Multi-News+: Cost-efficient Dataset Cleansing via LLM-based Data Annotation
abstract
The quality of the dataset is crucial for ensuring optimal performance and reliability of downstream task models.However, datasets often contain noisy data inadvertently included during the construction process.Numerous attempts have been made to correct this issue through human annotators.However, hiring and managing human annotators is expensive and time-consuming.As an alternative, recent studies are exploring the use of large language models (LLMs) for data annotation.In this study, we present a case study that extends the application of LLM-based data annotation to enhance the quality of existing datasets through a cleansing strategy.Specifically, we leverage approaches such as chain-of-thought and majority voting to imitate human annotation and classify unrelated documents from the Multi-News dataset, which is widely used for the multi-document summarization task.Through our proposed cleansing method, we introduce an enhanced MULTI-NEWS + .By employing LLMs for data cleansing, we demonstrate an efficient and effective approach to improving dataset quality without relying on expensive human annotation efforts.
Juhwan Choi, Jungmin Yun, Kyohoon Jin
EMNLP1