Seunguk Yu

dblp:370/1070 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
0009-0001-9497-9189ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Trustworthy machine learning · 41% Information extraction and text analysis · 19% Question answering and dialogue systems · 11%
Human-computer interaction and pervasive computing
1 paper
Usability and user experience research · 100%

Topics — the 9 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
1.922026
FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation · ACL (1) 2026
Delving into Multilingual Ethical Bias: The MSQAD with Statistical Hypothesis Tests for Large Language Models · ACL (1) 2025
Machine learning › Trustworthy machine learning › fairness
gender bias
1.012026
FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation · ACL (1) 2026
Machine learning › Trustworthy machine learning › fairness › bias mitigation
gender bias mitigation
1.012026
FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation · ACL (1) 2026
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation
1.012026
FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation · ACL (1) 2026
Machine learning › Generative modeling › synthetic data generation
dataset generation
0.812024
UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset Generation · EMNLP 2024
Machine learning › Transfer learning and domain adaptation
domain generalization
0.812024
UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset Generation · EMNLP 2024
Natural language and speech › Information extraction and text analysis › sentiment analysis
sentiment classification
0.812024
UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset Generation · EMNLP 2024
Usability and user experience research
user study
0.312026
RefLens: End-to-End Evidence-Grounded Citation Verification with LLM Agents · AAAI 2026
Machine learning › Efficient and distributed learning
model compression
0.212024
UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset Generation · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

large language model · 2.0LLM agents · 2.0multi-agent framework · 1.0large language model reasoning · 1.0statistical hypothesis testing · 0.9prompt-based learning · 0.8
YearPublicationVenuePosition
2026 RefLens: End-to-End Evidence-Grounded Citation Verification with LLM Agents
abstract
Accurate citation is critical, yet error rates remain high across scientific literature. We present RefLens, an end-to-end system that automates citation verification from PDF parsing to interactive report generation. Unlike summary- or embedding-based approaches, RefLens performs evidence-grounded verification by extracting verbatim spans from original sources and displaying citation-level cards and a paper-level dashboard. In a 35-participant study, users rated value (M=4.34), trust (M=4.15), and usability (M=4.19) highly, with strong adoption intention (M=4.28).
Seunghoo Lee, Junehyoung Kwon, Jooweon Choi, Jungmin Yun, Seunguk Yu, Jinhee Jang
AAAI5
2026 FairQE: Multi-Agent Framework for Mitigating Gender Bias in Translation Quality Estimation
abstract
Quality Estimation (QE) aims to assess machine translation quality without reference translations, but recent studies have shown that existing QE models exhibit systematic gender bias.In particular, they tend to favor masculine realizations in gender-ambiguous contexts and may assign higher scores to gendermisaligned translations even when gender is explicitly specified.To address these issues, we propose FairQE, a multi-agent-based, fairnessaware QE framework that mitigates gender bias in both gender-ambiguous and genderexplicit scenarios.FairQE detects gender cues, generates gender-flipped translation variants, and combines conventional QE scores with LLM-based bias-mitigating reasoning through a dynamic bias-aware aggregation mechanism.This design preserves the strengths of existing QE models while calibrating their genderrelated biases in a plug-and-play manner.Extensive experiments across multiple gender bias evaluation settings demonstrate that FairQE consistently improves gender fairness over strong QE baselines.Moreover, under MQMbased meta-evaluation following the WMT 2023 Metrics Shared Task, FairQE achieves competitive or improved general QE performance.These results show that gender bias in QE can be effectively mitigated without sacrificing evaluation accuracy, enabling fairer and more reliable translation evaluation.
Jinhee Jang, Juhwan Choi, Seunguk Yu
ACL (1)4
2026 Steering LLMs toward Korean Local Speech: Iterative Refinement Framework for Faithful Dialect Translation
Keunhyeung Park, Seunguk Yu
LREC2
2025 Delving into Multilingual Ethical Bias: The MSQAD with Statistical Hypothesis Tests for Large Language Models
abstract
Despite the recent strides in large language models, studies have underscored the existence of social biases within these systems.In this paper, we delve into the validation and comparison of the ethical biases of LLMs concerning globally discussed and potentially sensitive topics, hypothesizing that these biases may arise from language-specific distinctions.Introducing the Multilingual Sensitive Questions & Answers Dataset (MSQAD), we collected news articles from Human Rights Watch covering 17 topics, and generated socially sensitive questions along with corresponding responses in multiple languages.We scrutinize the biases of these responses across languages and topics, employing two statistical hypothesis tests.The results suggest that the null hypotheses are rejected in most cases, indicating biases arising from cross-language differences.It indicates that ethical biases in responses are widespread across various languages, and notably, these biases are prevalent even among different LLMs.By making the proposed MSQAD openly available, we aim to facilitate future research endeavors focused on examining cross-language biases in LLMs and their variant models 1 .
Seunguk Yu, Juhwan Choi
ACL (1)1
2024 UniGen: Universal Domain Generalization for Sentiment Classification via Zero-shot Dataset Generation
abstract
Although pre-trained language models have exhibited great flexibility and versatility with prompt-based few-shot learning, they suffer from the extensive parameter size and limited applicability for inference.Recent studies have suggested that PLMs be used as dataset generators and a tiny task-specific model be trained to achieve efficient inference.However, their applicability to various domains is limited because they tend to generate domain-specific datasets.In this work, we propose a novel approach to universal domain generalization that generates a dataset regardless of the target domain.This allows for generalization of the tiny task model to any domain that shares the label space, thus enhancing the real-world applicability of the dataset generation paradigm.Our experiments indicate that the proposed method accomplishes generalizability across various domains while using a parameter set that is orders of magnitude smaller than PLMs.
Juhwan Choi, Yeonghwa Kim, Seunguk Yu, Jungmin Yun
EMNLP3