Dayeon Ki

dblp:348/6309 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Machine translation · 46% Language models and text generation · 30% Multi-agent systems · 23%
Human-computer interaction and pervasive computing
2 papers
Human-AI interaction · 50% Usability and user experience research · 50%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › alignment › pluralistic alignment
cultural alignment
0.912025
Multiple LLM Agents Debate for Equitable Cultural Alignment · ACL (1) 2025
Knowledge, reasoning and agents › Multi-agent systems › LLM-based multi-agent systems
multi-agent debate
0.912025
Multiple LLM Agents Debate for Equitable Cultural Alignment · ACL (1) 2025
Natural language and speech › Machine translation › machine translation evaluation
translation quality estimation
0.912025
Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation · EMNLP 2025
Natural language and speech › Language models and text generation
large language model
0.312025
Multiple LLM Agents Debate for Equitable Cultural Alignment · ACL (1) 2025
Human-AI interaction › reliance on AI
appropriate reliance on AI
0.312025
Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation · EMNLP 2025
Usability and user experience research
user perception
0.312025
Toward Machine Translation Literacy: How Lay Users Perceive and Rely on Imperfect Translations · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

user study · 1.7question-answer tables · 1.7back-translation · 1.7LLM explanations · 1.7self-reflection · 0.9multi-agent debate · 0.9
YearPublicationVenuePosition
2025 Multiple LLM Agents Debate for Equitable Cultural Alignment
abstract
Large Language Models (LLMs) need to adapt their predictions to diverse cultural contexts to benefit diverse communities across the world. While previous efforts have focused on single-LLM, single-turn approaches, we propose to exploit the complementary strengths of multiple LLMs to promote cultural adaptability. We introduce a Multi-Agent Debate framework, where two LLM-based agents debate over a cultural scenario and collaboratively reach a final decision. We propose two variants: one where either LLM agents exclusively debate and another where they dynamically choose between self-reflection and debate during their turns. We evaluate these approaches on 7 open-weight LLMs (and 21 LLM combinations) using the NormAd-ETI benchmark for social etiquette norms in 75 countries. Experiments show that debate improves both overall accuracy and cultural group parity over single-LLM baselines. Notably, multi-agent debate enables relatively small LLMs (7-9B) to achieve accuracies comparable to that of a much larger model (27B parameters).
Dayeon Ki, Rachel Rudinger, Tianyi Zhou 0001, Marine Carpuat
ACL (1)1
2025 Should I Share this Translation? Evaluating Quality Feedback for User Reliance on Machine Translation
abstract
As people increasingly use AI systems in work and daily life, mechanisms that help them use AI responsibly are urgently needed, especially when they are not equipped to verify AI predictions themselves.We study a realistic Machine Translation (MT) scenario where monolingual users decide whether to share an MT output, first without and then with quality feedback.We compare four types of quality feedback: explicit feedback that directly give users an assessment of translation quality using (1) error highlights and (2) LLM explanations, and implicit feedback that helps users compare MT inputs and outputs through (3) backtranslation and ( 4) question-answer (QA) tables.We find that all feedback types, except error highlights, significantly improve both decision accuracy and appropriate reliance.Notably, implicit feedback, especially QA tables, yields significantly greater gains than explicit feedback in terms of decision accuracy, appropriate reliance, and user perceptions -receiving the highest ratings for helpfulness and trust, and the lowest for mental burden.
Dayeon Ki, Kevin Duh, Marine Carpuat
EMNLP1
2025 Toward Machine Translation Literacy: How Lay Users Perceive and Rely on Imperfect Translations
abstract
Yimin Xiao, Yongle Zhang, Dayeon Ki, Calvin Bao, Marianna J. Martindale, Charlotte Vaughn, Ge Gao, Marine Carpuat. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Yimin Xiao, Yongle Zhang 0004, Dayeon Ki, Calvin Bao, Marianna J. Martindale, Charlotte Vaughn, Ge Gao 0001, Marine Carpuat
EMNLP3
2025 Automatic Input Rewriting Improves Translation with Large Language Models
abstract
Dayeon Ki, Marine Carpuat. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Dayeon Ki, Marine Carpuat
NAACL (Long Papers)1