VLDB 2026 Research / reviewers in the wild / expert
Nayeon Lee
dblp:212/6295
· DBLP profile ↗
14ranked-venue papers
6as first author
11since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 6 first-author · 11 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Measuring Political Bias in Large Language Models: What Is Said and How It Is SaidabstractWe propose to measure political bias in LLMs by analyzing both the content and style of their generated content regarding political issues.Existing benchmarks and measures focus on gender and racial biases.However, political bias exists in LLMs and can lead to polarization and other harms in downstream applications.In order to provide transparency to users, we advocate that there should be fine-grained and explainable measures of political biases generated by LLMs.Our proposed measure looks at different political issues such as reproductive rights and climate change, at both the content (the substance of the generation) and the style (the lexical polarity) of such bias.We measured the political bias in eleven opensourced LLMs and showed that our proposed framework is easily scalable to other topics and is explainable. Yejin Bang, Delong Chen, Nayeon Lee, Pascale Fung |
ACL (1) | 3 |
| 2024 | Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to AnalysisabstractNayeon Lee, Chani Jung, Junho Myung, Jiho Jin, Jose Camacho-Collados, Juho Kim, Alice Oh. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Nayeon Lee, Chani Jung, Junho Myung, Jiho Jin, José Camacho-Collados, Juho Kim 0001, Alice Oh |
NAACL-HLT | 1 |
| 2024 | BLEnD: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and LanguagesabstractLarge language models (LLMs) often lack culture-specific everyday knowledge, especially across diverse regions and non-English languages. Existing benchmarks for evaluating LLMs' cultural sensitivities are usually limited to a single language or online sources like Wikipedia, which may not reflect the daily habits, customs, and lifestyles of different regions. That is, information about the food people eat for their birthday celebrations, spices they typically use, musical instruments youngsters play or the sports they practice in school is not always explicitly written online. To address this issue, we introduce BLEnD, a hand-crafted benchmark designed to evaluate LLMs' everyday knowledge across diverse cultures and languages. The benchmark comprises 52.6k question-answer pairs from 16 countries/regions, in 13 different languages, including low-resource ones such as Amharic, Assamese, Azerbaijani, Hausa, and Sundanese. We evaluate LLMs in two formats: short-answer questions, and multiple-choice questions. We show that LLMs perform better in cultures that are more present online, with a maximum 57.34% difference in GPT-4, the best-performing model, in the short-answer format.Furthermore, we find that LLMs perform better in their local languages for mid-to-high-resource languages. Interestingly, for languages deemed to be low-resource, LLMs provide better answers in English. We make our dataset publicly available at: https://github.com/nlee0212/BLEnD. Junho Myung, Nayeon Lee, Yi Zhou 0019, Jiho Jin, Rifki Afina Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Pérez-Almendros, Abinew Ali Ayele, Víctor Gutiérrez-Basulto, Yazmín Ibáñez-García, Hwaran Lee, Shamsuddeen Hassan Muhammad, Ki-Woong Park, Anar Rzayev, Nina White, Seid Muhie Yimam, Mohammad Taher Pilehvar, Nedjma Ousidhoum, José Camacho-Collados, Alice Oh |
NeurIPS | 2 |
| 2024 | KoBBQ: Korean Bias Benchmark for Question AnsweringabstractAbstract Warning: This paper contains examples of stereotypes and biases. The Bias Benchmark for Question Answering (BBQ) is designed to evaluate social biases of language models (LMs), but it is not simple to adapt this benchmark to cultural contexts other than the US because social biases depend heavily on the cultural context. In this paper, we present KoBBQ, a Korean bias benchmark dataset, and we propose a general framework that addresses considerations for cultural adaptation of a dataset. Our framework includes partitioning the BBQ dataset into three classes—Simply-Transferred (can be used directly after cultural translation), Target-Modified (requires localization in target groups), and Sample-Removed (does not fit Korean culture)—and adding four new categories of bias specific to Korean culture. We conduct a large-scale survey to collect and validate the social biases and the targets of the biases that reflect the stereotypes in Korean culture. The resulting KoBBQ dataset comprises 268 templates and 76,048 samples across 12 categories of social bias. We use KoBBQ to measure the accuracy and bias scores of several state-of-the-art multilingual LMs. The results clearly show differences in the bias of LMs as measured by KoBBQ and a machine-translated version of BBQ, demonstrating the need for and utility of a well-constructed, culturally aware social bias benchmark. Jiho Jin, Jiseon Kim, Nayeon Lee, Haneul Yoo, Alice Oh, Hwaran Lee |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | A Multitask, Multilingual, Multimodal Evaluation of ChatGPT on Reasoning, Hallucination, and InteractivityabstractYejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su, Bryan Wilie, Holy Lovenia, Ziwei Ji, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu, Pascale Fung. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yejin Bang, Samuel Cahyawijaya, Nayeon Lee, Wenliang Dai, Dan Su 0003, Bryan Wilie, Holy Lovenia, Ziwei Ji 0001, Tiezheng Yu, Willy Chung, Quyet V. Do, Yan Xu 0012, Pascale Fung |
IJCNLP (1) | 3 |
| 2022 | Evaluating Parameter Efficient Learning for GenerationabstractPeng Xu, Mostofa Patwary, Shrimai Prabhumoye, Virginia Adams, Ryan Prenger, Wei Ping, Nayeon Lee, Mohammad Shoeybi, Bryan Catanzaro. Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 2022. Peng Xu 0008, Mostofa Patwary, Shrimai Prabhumoye, Virginia Adams, Ryan Prenger, Wei Ping, Nayeon Lee, Mohammad Shoeybi, Bryan Catanzaro |
EMNLP | 7 |
| 2022 | NeuS: Neutral Multi-News Summarization for Mitigating Framing BiasabstractNayeon Lee, Yejin Bang, Tiezheng Yu, Andrea Madotto, Pascale Fung. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Nayeon Lee, Yejin Bang, Tiezheng Yu, Andrea Madotto, Pascale Fung |
NAACL-HLT | 1 |
| 2022 | Factuality Enhanced Language Models for Open-Ended Text GenerationabstractPretrained language models (LMs) are susceptible to generate text with nonfactual information. In this work, we measure and improve the factual accuracy of large-scale LMs for open-ended text generation. We design the FactualityPrompts test set and metrics to measure the factuality of LM generations. Based on that, we study the factual accuracy of LMs with parameter sizes ranging from 126M to 530B. Interestingly, we find that larger LMs are more factual than smaller ones, although a previous study suggests that larger LMs can be less truthful in terms of misconceptions. In addition, popular sampling algorithms (e.g., top-p) in open-ended text generation can harm the factuality due to the ``uniform randomness'' introduced at every sampling step. We propose the factual-nucleus sampling algorithm that dynamically adapts the randomness to improve the factuality of generation while maintaining quality. Furthermore, we analyze the inefficiencies of the standard training method in learning correct associations between entities from factual text corpus (e.g., Wikipedia). We propose a factuality-enhanced training method that uses TopicPrefix for better awareness of facts and sentence completion as the training objective, which can vastly reduce the factual errors. Nayeon Lee, Wei Ping, Peng Xu 0008, Mostofa Patwary, Pascale Fung, Mohammad Shoeybi, Bryan Catanzaro |
NeurIPS | 1 |
| 2021 | Towards Few-shot Fact-Checking via PerplexityabstractNayeon Lee, Yejin Bang, Andrea Madotto, Pascale Fung. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Nayeon Lee, Yejin Bang, Andrea Madotto, Pascale Fung |
NAACL-HLT | 1 |
| 2021 | On Unifying Misinformation DetectionabstractNayeon Lee, Belinda Z. Li, Sinong Wang, Pascale Fung, Hao Ma, Wen-tau Yih, Madian Khabsa. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Nayeon Lee, Belinda Z. Li, Sinong Wang, Pascale Fung, Hao Ma 0001, Scott Yih, Madian Khabsa |
NAACL-HLT | 1 |
| 2021 | Assessing Political Prudence of Open-domain ChatbotsabstractPolitically sensitive topics are still a challenge for open-domain chatbots.However, dealing with politically sensitive content in a responsible, non-partisan, and safe behavior way is integral for these chatbots.Currently, the main approach to handling political sensitivity is by simply changing such a topic when it is detected.This is safe but evasive and results in a chatbot that is less engaging.In this work, as a first step towards a politically safe chatbot, we propose a group of metrics for assessing their political prudence.We then conduct political prudence analysis of various chatbots and discuss their behavior from multiple angles through our automatic metric and human evaluation metrics.The testsets and codebase are released to promote research in this area.1 Yejin Bang, Nayeon Lee, Etsuko Ishii, Andrea Madotto, Pascale Fung |
SIGDIAL | 2 |
| 2020 | Effect of pre-training to build a regression model using shallow neural network for semiconductor plasma etch process equipmentabstractPlasma etch process is one of manufacturing steps to fabricate semiconductor chips and the difficulty of etch process has become harder and harder because the target specifications of chips have become harsher. To overcome this circumstance, monitoring plasma parameters in real-time has been requested. In this study, regression models to predict a plasma density from the intensities of optical wavelength obtained from plasma etch process chamber using shallow neural network were presented. The estimation results of several models with or without pre-training were also analyzed. The model using variational auto-encoder showed the best performance and it can expect to be easily accepted in semiconductor industry because optical intensity measurement device was already equipped for plasma etch process chamber. Ohyung Kwon, Nayeon Lee, Kangil Kim |
IEEE BigData | 2 |
| 2018 | Improving Large-Scale Fact-Checking using Decomposable Attention Models and Lexical TaggingabstractFact-checking of textual sources needs to effectively extract relevant information from large knowledge bases.In this paper, we extend an existing pipeline approach to better tackle this problem.We propose a neural ranker using a decomposable attention model that dynamically selects sentences to achieve promising improvement in evidence retrieval F1 by 38.80%, with (×65) speedup compared to a TF-IDF method.Moreover, we incorporate lexical tagging methods into our pipeline framework to simplify the tasks and render the model more generalizable.As a result, our framework achieves promising performance on a large-scale fact extraction and verification dataset with speedup. Nayeon Lee, Chien-Sheng Wu, Pascale Fung |
EMNLP | 1 |
| 2017 | Emojive! Collecting Emotion Data from Speech and Facial Expression Using Mobile Game App
Ji Ho Park, Nayeon Lee, Dario Bertero, Anik Dey, Pascale Fung |
INTERSPEECH | 2 |