Yejin Bang

dblp:261/2805 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
10since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 HalluLens: LLM Hallucination Benchmark
abstract
Yejin Bang, Ziwei Ji, Alan Schelten, Anthony Hartshorn, Tara Fowler, Cheng Zhang, Nicola Cancedda, Pascale Fung. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Yejin Bang, Ziwei Ji 0001, Alan Schelten, Anthony Hartshorn, Tara Fowler, Nicola Cancedda, Pascale Fung
ACL (1)1
2025 Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations
abstract
Ziwei Ji, Lei Yu, Yeskendir Koishekenov, Yejin Bang, Anthony Hartshorn, Alan Schelten, Cheng Zhang, Pascale Fung, Nicola Cancedda. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Ziwei Ji 0001, Yeskendir Koishekenov, Yejin Bang, Anthony Hartshorn, Alan Schelten, Pascale Fung, Nicola Cancedda
EMNLP4
2025 High-Dimension Human Value Representation in Large Language Models
abstract
Samuel Cahyawijaya, Delong Chen, Yejin Bang, Leila Khalatbari, Bryan Wilie, Ziwei Ji, Etsuko Ishii, Pascale Fung. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Samuel Cahyawijaya, Delong Chen, Yejin Bang, Leila Khalatbari, Bryan Wilie, Ziwei Ji 0001, Etsuko Ishii, Pascale Fung
NAACL (Long Papers)3
2024 Measuring Political Bias in Large Language Models: What Is Said and How It Is Said
abstract
We propose to measure political bias in LLMs by analyzing both the content and style of their generated content regarding political issues.Existing benchmarks and measures focus on gender and racial biases.However, political bias exists in LLMs and can lead to polarization and other harms in downstream applications.In order to provide transparency to users, we advocate that there should be fine-grained and explainable measures of political biases generated by LLMs.Our proposed measure looks at different political issues such as reproductive rights and climate change, at both the content (the substance of the generation) and the style (the lexical polarity) of such bias.We measured the political bias in eleven opensourced LLMs and showed that our proposed framework is easily scalable to other topics and is explainable.
Yejin Bang, Delong Chen, Nayeon Lee, Pascale Fung
ACL (1)1
2024 A Humanoid Robot Dialogue System Architecture Targeting Patient Interview Tasks
abstract
Humanoid robots are promising approach to automating patient interviews routinely conducted by medical staff. Their human-like appearance enables them to use the full gamut of verbal and behavioral cues that are critical to a successful interview. On the other hand, anthropomorphism can induce expectations of human-level performance by the robot. Not meeting such expectations degrades the quality of interaction. Specifically, humans expect rich real-time interactions during speech exchange, such as backchanneling and barge-ins. The nature of the patient interview task differs from most other scenarios where task oriented dialogue systems have been used, as there is increased potential of engagement breakdown during interaction. We describe a dialogue system architecture that improves the performance of humanoid robots on the patient interview task. Our architecture adds a nested inner real-time control loop to improve the timeliness of the robot’s responses based on the notion of "stance", an elaboration of the concept of a "turn", common in most existing dialogue systems. It also expands the dialogue state to monitor not only task progress, but also human engagement. Experiments using a humanoid robot running our proposed architecture reveal improved performance on interview tasks in terms of the perceived timeliness of responses and users’ impressions of the system.
Dingdong Liu, Yejin Bang, Ho Shu Chan, Rita Frieske, Hoo Choun Chung, Jay Nieles, Tianjia Zhang, Kien T. Pham 0001, Wai Yi Rosita Cheng, Yini Fang, Qifeng Chen 0001, Pascale Fung, Xiaojuan Ma, Bertram E. Shi
RO-MAN3
2023 A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and Interactivity
abstract
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, Pascale Fung. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su 0003, Bryan Wilie, Holy Lovenia, Ziwei Ji 0001, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu 0012, Pascale Fung
IJCNLP (1)1
2022 NeuS: Neutral Multi-News Summarization for Mitigating Framing Bias
abstract
Nayeon Lee, Yejin Bang, Tiezheng Yu, Andrea Madotto, Pascale Fung. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Nayeon Lee, Yejin Bang, Tiezheng Yu, Andrea Madotto, Pascale Fung
NAACL-HLT2
2021 The Adapter-Bot: All-In-One Controllable Conversational Model
abstract
In this paper, we present the Adapter-Bot, a generative chat-bot that uses a fixed backbone conversational model such as DialGPT (Zhang et al. 2019) and triggers on-demand dialogue skills via different adapters (Houlsby et al. 2019). Each adapter can be trained independently, thus allowing a continual integration of skills without retraining the entire model. Depending on the skills, the model is able to process multiple knowledge types, such as text, tables, and graphs, in a seamless manner. The dialogue skills can be triggered automatically via a dialogue manager, or manually, thus allowing high-level control of the generated responses. At the current stage, we have implemented 12 response styles (e.g., positive, negative etc.), 6 goal-oriented skills (e.g. weather information, movie recommendation, etc.), and personalized and emphatic responses.
Zhaojiang Lin, Andrea Madotto, Yejin Bang, Pascale Fung
AAAI3
2021 Towards Few-shot Fact-Checking via Perplexity
abstract
Nayeon Lee, Yejin Bang, Andrea Madotto, Pascale Fung. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Nayeon Lee, Yejin Bang, Andrea Madotto, Pascale Fung
NAACL-HLT2
2021 Assessing Political Prudence of Open-domain Chatbots
abstract
Politically sensitive topics are still a challenge for open-domain chatbots.However, dealing with politically sensitive content in a responsible, non-partisan, and safe behavior way is integral for these chatbots.Currently, the main approach to handling political sensitivity is by simply changing such a topic when it is detected.This is safe but evasive and results in a chatbot that is less engaging.In this work, as a first step towards a politically safe chatbot, we propose a group of metrics for assessing their political prudence.We then conduct political prudence analysis of various chatbots and discuss their behavior from multiple angles through our automatic metric and human evaluation metrics.The testsets and codebase are released to promote research in this area.1
Yejin Bang, Nayeon Lee, Etsuko Ishii, Andrea Madotto, Pascale Fung
SIGDIAL1