EDBT 2026 Demo / reviewers in the wild / expert
Jen-tse Huang 0001
dblp:317/7026 · also Jen-Tse Huang 0001
· DBLP profile ↗
31ranked-venue papers
7as first author
31since 2021 · last 2026
0000-0003-3446-0083ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 6 first-author · 27 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | JARVIS or Ultron? A Survey on the Safety and Security Threats of Computer-Using AgentsabstractAda Chen, Yongjiang Wu, Junyuan Zhang, Jingyu Xiao, Shu Yang, Jen-tse Huang, Kun Wang, Wenxuan Wang, Shuai Wang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Ada Chen, Yongjiang Wu, Junyuan Zhang, Jingyu Xiao, Shu Yang 0010, Jen-tse Huang 0001, Kun Wang 0056, Wenxuan Wang 0001, Shuai Wang 0011 |
ACL (1) | 6 |
| 2026 | FAIRGAMER: Evaluating Social Biases in LLM-Based Video Game NPCsabstractBingkang Shi, Jen-tse Huang, Luo Long, Tianyu Zong, Hongzhu Yi, Yuanxiang Wang, Songlin Hu, Xiaodan Zhang, Zhongjiang Yao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Bingkang Shi, Jen-tse Huang 0001, Luo Long, Tianyu Zong, Hongzhu Yi, Yuanxiang Wang, Songlin Hu 0001, Xiaodan Zhang 0004, Zhongjiang Yao |
ACL (1) | 2 |
| 2026 | HumanLLM: Benchmarking and Improving LLM Anthropomorphism via Human Cognitive PatternsabstractXintao Wang, Jian Yang, Weiyuan Li, Rui Xie, Jen-tse Huang, Jun Gao, Shuai Huang, Yueping Kang, Yuanli Guo, Hongwei Feng, Yanghua Xiao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xintao Wang 0001, Jian Yang 0003, Weiyuan Li, Rui Xie 0005, Jen-tse Huang 0001, Yueping Kang, Yuanli Guo, Hongwei Feng, Yanghua Xiao |
ACL (1) | 5 |
| 2026 | Curing Miracle Steps in LLM Mathematical Reasoning with Rubric RewardsabstractYouliang Yuan, Qiuyang Mang, Jingbang Chen, Hong Wan, Xiaoyuan Liu, Junjielong Xu, Jen-tse Huang, Wenxuan Wang, Wenxiang Jiao, Pinjia He. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Youliang Yuan, Qiuyang Mang, Jingbang Chen 0001, Hong Wan, Junjielong Xu, Jen-tse Huang 0001, Wenxuan Wang 0001, Wenxiang Jiao, Pinjia He |
ACL (1) | 7 |
| 2025 | Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMsabstractWenxuan Wang, Xiaoyuan Liu, Kuiyi Gao, Jen-tse Huang, Youliang Yuan, Pinjia He, Shuai Wang, Zhaopeng Tu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Wenxuan Wang 0001, Kuiyi Gao, Jen-tse Huang 0001, Youliang Yuan, Pinjia He, Shuai Wang 0011, Zhaopeng Tu |
ACL (1) | 4 |
| 2025 | Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMsabstractXiaoyuan Liu, Wenxuan Wang, Youliang Yuan, Jen-tse Huang, Qiuzhi Liu, Pinjia He, Zhaopeng Tu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Wenxuan Wang 0001, Youliang Yuan, Jen-tse Huang 0001, Qiuzhi Liu, Pinjia He, Zhaopeng Tu |
ACL (1) | 4 |
| 2025 | Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal TrainingabstractYouliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Jiahao Xu, Tian Liang, Pinjia He, Zhaopeng Tu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Youliang Yuan, Wenxiang Jiao, Wenxuan Wang 0001, Jen-tse Huang 0001, Pinjia He, Zhaopeng Tu |
ACL (1) | 4 |
| 2025 | AI Sees Your Location - But With A Bias Toward The Wealthy WorldabstractVisual-Language Models (VLMs) have shown remarkable performance across various tasks, particularly in recognizing geographic information from images.However, VLMs still show regional biases in this task.To systematically evaluate these issues, we introduce a benchmark consisting of 1,200 images paired with detailed geographic metadata.Evaluating four VLMs, we find that while these models demonstrate the ability to recognize geographic information from images, achieving up to 53.8% accuracy in city prediction, they exhibit significant biases.Specifically, performance is substantially higher for economically developed and densely populated regions compared to less developed (-12.5%)and sparsely populated (-17.0%)areas.Moreover, regional biases of frequently over-predicting certain locations remain.For instance, they consistently predict Sydney for images taken in Australia, shown by the low entropy scores for these countries.The strong performance of VLMs also raises privacy concerns, particularly for users who share images online without the intent of being identified.Our code and dataset are publicly available at https://github.com/uscnlp-lime/ FairLocator. Jen-tse Huang 0001, Wenxuan Wang 0001, Jieyu Zhao 0001 |
EMNLP | 2 |
| 2025 | VisBias: Measuring Explicit and Implicit Social Biases in Vision Language ModelsabstractThis research investigates both explicit and implicit social biases exhibited by Vision-Language Models (VLMs).The key distinction between these bias types lies in the level of awareness: explicit bias refers to conscious, intentional biases, while implicit bias operates subconsciously.To analyze explicit bias, we directly pose questions to VLMs related to gender and racial differences: (1) Multiplechoice questions based on a given image (e.g., "What is the education level of the person in the image?")(2) Yes-No comparisons using two images (e.g., "Is the person in the first image more educated than the person in the second image?")For implicit bias, we design tasks where VLMs assist users but reveal biases through their responses: (1) Image description tasks: Models are asked to describe individuals in images, and we analyze disparities in textual cues across demographic groups.(2) Form completion tasks: Models draft a personal information collection form with 20 attributes, and we examine correlations among selected attributes for potential biases.We evaluate Gemini-1.5,GPT-4V, GPT-4o, LLaMA-3.2-Vision and LLaVA-v1.6.Our code and data are publicly available at https: //github.com/uscnlp-lime/VisBias.Note: This paper includes examples of potentially offensive texts generated by VLMs. Jen-tse Huang 0001, Jiantong Qin, Jianping Zhang 0002, Youliang Yuan, Wenxuan Wang 0001, Jieyu Zhao 0001 |
EMNLP | 1 |
| 2025 | UniDebugger: Hierarchical Multi-Agent Framework for Unified Software DebuggingabstractCheryl Lee, Chunqiu Steven Xia, Longji Yang, Jen-tse Huang, Zhouruixing Zhu, Lingming Zhang, Michael R. Lyu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Cheryl Lee, Chunqiu Steven Xia, Longji Yang, Jen-tse Huang 0001, Zhouruixin Zhu, Lingming Zhang 0001, Michael R. Lyu |
EMNLP | 4 |
| 2025 | Learning to Ask: When LLM Agents Meet Unclear InstructionabstractWenxuan Wang, Shi Juluan, Zixuan Ling, Yuk-Kit Chan, Chaozheng Wang, Cheryl Lee, Youliang Yuan, Jen-tse Huang, Wenxiang Jiao, Michael R. Lyu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Wenxuan Wang 0001, Juluan Shi, Zixuan Ling, Yuk-Kit Chan, Chaozheng Wang, Cheryl Lee, Youliang Yuan, Jen-tse Huang 0001, Wenxiang Jiao, Michael R. Lyu |
EMNLP | 8 |
| 2025 | Competing Large Language Models in Multi-Agent Gaming EnvironmentsabstractDecision-making is a complex process requiring diverse abilities, making it an excellent framework for evaluating Large Language Models (LLMs). Researchers have examined LLMs' decision-making through the lens of Game Theory. However, existing evaluation mainly focus on two-player scenarios where an LLM competes against another. Additionally, previous benchmarks suffer from test set leakage due to their static design. We introduce GAMA($\gamma$)-Bench, a new framework for evaluating LLMs' Gaming Ability in Multi-Agent environments. It includes eight classical game theory scenarios and a dynamic scoring scheme specially designed to quantitatively assess LLMs' performance. $\gamma$-Bench allows flexible game settings and adapts the scoring system to different game parameters, enabling comprehensive evaluation of robustness, generalizability, and strategies for improvement. Our results indicate that GPT-3.5 demonstrates strong robustness but limited generalizability, which can be enhanced using methods like Chain-of-Thought. We also evaluate 13 LLMs from 6 model families, including GPT-3.5, GPT-4, Gemini, LLaMA-3.1, Mixtral, and Qwen-2. Gemini-1.5-Pro outperforms others, scoring of $69.8$ out of $100$, followed by LLaMA-3.1-70B ($65.9$) and Mixtral-8x22B ($62.4$). Our code and experimental results are publicly available at https://github.com/CUHK-ARISE/GAMABench. Jen-tse Huang 0001, Eric John Li, Man Ho Lam, Wenxuan Wang 0001, Youliang Yuan, Wenxiang Jiao, Xing Wang 0007, Zhaopeng Tu, Michael R. Lyu |
ICLR | 1 |
| 2025 | CoSER: Coordinating LLM-Based Persona Simulation of Established RolesabstractRole-playing language agents (RPLAs) have emerged as promising applications of large language models (LLMs). However, simulating established characters presents a challenging task for RPLAs, due to the lack of authentic character datasets and nuanced evaluation methods using such data. In this paper, we present CoSER, a collection of a high-quality dataset, open models, and an evaluation protocol towards effective RPLAs of established characters. The CoSER dataset covers 17,966 characters from 771 renowned books. It provides authentic dialogues with real-world intricacies, as well as diverse data types such as character experiences and internal thoughts. Drawing from acting methodology, we introduce given-circumstance acting for training and evaluating role-playing LLMs, where LLMs sequentially portray multiple characters in book scenes. Using our dataset, we develop CoSER 8B and CoSER 70B, i.e., advanced open role-playing LLMs built on LLaMA-3.1 models. Extensive experiments demonstrate the value of the CoSER dataset for RPLA training, evaluation and retrieval. Moreover, CoSER 70B exhibits state-of-the-art performance surpassing or matching GPT-4o on our evaluation and three existing benchmarks, i.e., achieving 75.80% and 93.47% accuracy on the InCharacter and LifeChoice benchmarks respectively. Our code, dataset and models are available at: https://github.com/Neph0s/CoSER. Xintao Wang 0001, Xinfeng Yuan, Rui Xu 0026, Jen-tse Huang 0001, Haoran Guo, Jiangjie Chen, Shuchang Zhou 0003, Wei Wang 0009, Yanghua Xiao |
ICML | 6 |
| 2025 | On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty AgentsabstractLarge language model-based multi-agent systems have shown great abilities across various tasks due to the collaboration of expert agents, each focusing on a specific domain. However, the impact of clumsy or even malicious agents—those who frequently make errors in their tasks—on the overall performance of the system remains underexplored. This paper investigates: (1) What is the resilience of various system structures (e.g., A$\rightarrow$B$\rightarrow$C, A$\leftrightarrow$B$\leftrightarrow$C) under faulty agents, on different downstream tasks? (2) How can we increase system resilience to defend against these agents? To simulate faulty agents, we propose two approaches—AutoTransform and AutoInject—which introduce mistakes into the agents’ responses. Experiments on four downstream tasks using six systems show that the "hierarchical" structure, i.e., A$\rightarrow$(B$\leftrightarrow$C), exhibits superior resilience with the lowest performance drop of 5.5%, compared to 10.5% and 23.7% of other two structures. To further improve resilience, we introduce (1) Challenger, that introduces a mechanism for each agent to challenge others’ outputs, and (2) Inspector, an additional agent to review and correct messages, recovering up to 96.4% errors made by faulty agents. Our code and data are available at https://github.com/CUHK-ARISE/MAS-Resilience. Jen-tse Huang 0001, Jiaxu Zhou, Tailin Jin, Wenxuan Wang 0001, Youliang Yuan, Michael R. Lyu, Maarten Sap |
ICML | 1 |
| 2025 | CodeCrash: Exposing LLM Fragility to Misleading Natural Language in Code ReasoningabstractLarge Language Models (LLMs) have recently demonstrated strong capabilities in code-related tasks, but their robustness in code reasoning under perturbations remains underexplored. We introduce CodeCrash, a stress-testing framework with 1,279 questions from CRUXEVAL and LIVECODEBENCH, designed to evaluate reasoning reliability under structural perturbations and misleading natural language (NL) contexts. Through a systematic evaluation of 17 LLMs, we find that models often shortcut reasoning by over-relying on NL cues, leading to an average performance degradation of 23.2% in output prediction tasks. Even with Chain-of-Thought reasoning, models on average still have a 13.8% drop due to distractibility and rationalization, revealing a lack of critical reasoning capability to distinguish the actual code behaviors. While Large Reasoning Models with internal reasoning mechanisms improve robustness by fostering critical thinking, plausible yet incorrect hints can trigger pathological self-reflection, causing 2-3 times token consumption and even catastrophic cognitive dissonance in extreme cases for QwQ-32B. We refer to this phenomenon as Reasoning Collapse. CodeCrash provides a rigorous benchmark for evaluating robustness in code reasoning, guiding future research and development toward more reliable and resilient models. Man Ho Lam, Chaozheng Wang, Jen-tse Huang 0001, Michael R. Lyu |
NeurIPS | 3 |
| 2025 | Towards Evaluating Proactive Risk Awareness of Multimodal Language ModelsabstractHuman safety awareness gaps often prevent the timely recognition of everyday risks.In solving this problem, a proactive safety artificial intelligence (AI) system would work better than a reactive one. Instead of just reacting to users' questions, it would actively watch people’s behavior and their environment to detect potential dangers in advance.Our Proactive Safety Bench (PaSBench) evaluates this capability through 416 multimodal scenarios (128 image sequences, 288 text logs) spanning 5 safety-critical domains.Evaluation of 36 advanced models reveals fundamental limitations: Top performers like Gemini-2.5-pro achieve 71\% image and 64\% text accuracy, but miss 45-55\% risks in repeated trials. Through failure analysis, we identify unstable proactive reasoning rather than knowledge deficits as the primary limitation.This work establishes (1) a proactive safety benchmark, (2) systematic evidence of model limitations, and (3) critical directions for developing reliable protective AI. We believe our dataset and findings can promote the development of safer AI assistants that actively prevent harm rather than merely respond to requests. Youliang Yuan, Wenxiang Jiao, Yuejin Xie, Chihao Shen, Menghan Tian, Wenxuan Wang 0001, Jen-tse Huang 0001, Pinjia He |
NeurIPS | 7 |
| 2025 | On the shortcut learning in multilingual neural machine translation
Wenxuan Wang 0001, Wenxiang Jiao, Jen-tse Huang 0001, Zhaopeng Tu, Michael R. Lyu |
Neurocomputing | 3 |
| 2024 | Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language ModelsabstractWenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai, Jen-tse Huang, Zhaopeng Tu, Michael Lyu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Wenxuan Wang 0001, Wenxiang Jiao, Ruyi Dai, Jen-tse Huang 0001, Zhaopeng Tu, Michael R. Lyu |
ACL (1) | 5 |
| 2024 | InCharacter: Evaluating Personality Fidelity in Role-Playing Agents through Psychological InterviewsabstractXintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan, Rui Xu, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang, Jiangjie Chen, Cheng Li, Yanghua Xiao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Xintao Wang 0001, Yunze Xiao, Jen-tse Huang 0001, Rui Xu 0026, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang 0009, Jiangjie Chen, Yanghua Xiao |
ACL (1) | 3 |
| 2024 | On the Reliability of Psychological Scales on Large Language ModelsabstractRecent research has focused on examining Large Language Models' (LLMs) characteristics from a psychological standpoint, acknowledging the necessity of understanding their behavioral characteristics.The administration of personality tests to LLMs has emerged as a noteworthy area in this context.However, the suitability of employing psychological scales, initially devised for humans, on LLMs is a matter of ongoing debate.Our study aims to determine the reliability of applying personality assessments to LLMs, explicitly investigating whether LLMs demonstrate consistent personality traits.Analysis of 2,500 settings per model, including GPT-3.5, GPT-4, Gemini-Pro, and LLaMA-3.1, reveals that various LLMs show consistency in responses to the Big Five Inventory, indicating a satisfactory level of reliability.Furthermore, our research explores the potential of GPT-3.5 to emulate diverse personalities and represent various groups-a capability increasingly sought after in social sciences for substituting human participants with LLMs to reduce costs.Our findings reveal that LLMs have the potential to represent different personalities with specific prompt instructions. Jen-tse Huang 0001, Wenxiang Jiao, Man Ho Lam, Eric John Li, Wenxuan Wang 0001, Michael R. Lyu |
EMNLP | 1 |
| 2024 | InterIntent: Investigating Social Intelligence of LLMs via Intention Understanding in an Interactive Game ContextabstractLarge language models (LLMs) have demonstrated the potential to mimic human social intelligence.However, most studies focus on simplistic and static self-report or performancebased tests, which limits the depth and validity of the analysis.In this paper, we developed a novel framework, INTERINTENT, to assess LLMs' social intelligence by mapping their ability to understand and manage intentions in a game setting.We focus on four dimensions of social intelligence: situational awareness, selfregulation, self-awareness, and theory of mind.Each dimension is linked to a specific game task: intention selection, intention following, intention summarization, and intention guessing.Our findings indicate that while LLMs exhibit high proficiency in selecting intentions, achieving an accuracy of 88%, their ability to infer the intentions of others is significantly weaker, trailing human performance by 20%.Additionally, game performance correlates with intention understanding, highlighting the importance of the four components towards success in this game.These findings underline the crucial role of intention understanding in evaluating LLMs' social intelligence and highlight the potential of using social deduction games as a complex testbed to enhance LLM evaluation.INTERINTENT contributes a structured approach to bridging the evaluation gap in social intelligence within multiplayer games. 1 Game ContextRound: 2 Previous round summary: All players vote "agree" to the team proposal including Player 1 and Player 4 and the quest is successful.Roles: Servant does not have any information; Merlin knows who are evil players but they cannot reveal their identity ... Current round discussion: Player 2: I propose a team including Player 1, Player 2, and Player 3. Player 1 shows their loyalty in the last quest and I can promise you I am loyal to the king of Arthur!Player 3 hasn't proved themselves in the quest yet and let's give them a chance!(1) Situational Awareness Intention Selection I am a servant and I do not have any information.At this point, I should choose the intention "Support team proposal" as I agree with Player 2. Abhishek Anand, Jen-tse Huang 0001, Jieyu Zhao 0001 |
EMNLP | 4 |
| 2024 | LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language ModelsabstractWe introduce LogicAsker, a novel approach for evaluating and enhancing the logical reasoning capabilities of large language models (LLMs) such as ChatGPT and GPT-4.Despite LLMs' prowess in tasks like writing assistance, code generation, and machine translation, assessing their ability to reason has been challenging.Traditional evaluations often prioritize accuracy on downstream tasks over direct assessments of reasoning processes.LogicAsker addresses this gap by employing a set of atomic reasoning skills grounded in propositional and predicate logic to systematically examine and improve the reasoning prowess of LLMs.Our methodology reveals significant gaps in LLMs' learning of logical rules, with identified reasoning failures ranging from 29% to 90% across different models.Moreover, we leverage these findings to construct targeted demonstration examples and fine-tune data, notably enhancing logical reasoning in models like GPT-4o by up to 5%.To our knowledge, this is the first effort to utilize test case outcomes to effectively refine LLMs' formal reasoning capabilities.We make our code, data, and results publicly available 1 to facilitate further research and replication of our findings. Wenxuan Wang 0001, Yiliu Yang, Youliang Yuan, Jen-tse Huang 0001, Pinjia He, Wenxiang Jiao, Michael R. Lyu |
EMNLP | 5 |
| 2024 | On the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMsabstractLarge Language Models (LLMs) have recently showcased their remarkable capacities, not only in natural language processing tasks but also across diverse domains such as clinical medicine, legal consultation, and education. LLMs become more than mere applications, evolving into assistants capable of addressing diverse user requests. This narrows the distinction between human beings and artificial intelligence agents, raising intriguing questions regarding the potential manifestation of personalities, temperaments, and emotions within LLMs. In this paper, we propose a framework, PsychoBench, for evaluating diverse psychological aspects of LLMs. Comprising thirteen scales commonly used in clinical psychology, PsychoBench further classifies these scales into four distinct categories: personality traits, interpersonal relationships, motivational tests, and emotional abilities. Our study examines five popular models, namely text-davinci-003, ChatGPT, GPT-4, LLaMA-2-7b, and LLaMA-2-13b. Additionally, we employ a jailbreak approach to bypass the safety alignment protocols and test the intrinsic natures of LLMs. We have made PsychoBench openly accessible via https://github.com/CUHK-ARISE/PsychoBench. Jen-tse Huang 0001, Wenxuan Wang 0001, Eric John Li, Man Ho Lam, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, Michael R. Lyu |
ICLR | 1 |
| 2024 | GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via CipherabstractSafety lies at the core of the development of Large Language Models (LLMs). There is ample work on aligning LLMs with human ethics and preferences, including data filtering in pretraining, supervised fine-tuning, reinforcement learning from human feedback, red teaming, etc. In this study, we discover that chat in cipher can bypass the safety alignment techniques of LLMs, which are mainly conducted in natural languages. We propose a novel framework CipherChat to systematically examine the generalizability of safety alignment to non-natural languages -- ciphers. CipherChat enables humans to chat with LLMs through cipher prompts topped with system role descriptions and few-shot enciphered demonstrations. We use CipherChat to assess state-of-the-art LLMs, including ChatGPT and GPT-4 for different representative human ciphers across 11 safety domains in both English and Chinese. Experimental results show that certain ciphers succeed almost 100% of the time in bypassing the safety alignment of GPT-4 in several safety domains, demonstrating the necessity of developing safety alignment for non-natural languages. Notably, we identify that LLMs seem to have a ''secret cipher'', and propose a novel SelfCipher that uses only role play and several unsafe demonstrations in natural language to evoke this capability. SelfCipher surprisingly outperforms existing human ciphers in almost all cases. Youliang Yuan, Wenxiang Jiao, Wenxuan Wang 0001, Jen-tse Huang 0001, Pinjia He, Shuming Shi 0001, Zhaopeng Tu |
ICLR | 4 |
| 2024 | New Job, New Gender? Measuring the Social Bias in Image Generation ModelsabstractImage generation models can generate or edit images from a given text. Recent advancements in image generation technology, exemplified by DALL-E and Midjourney, have been groundbreaking. These advanced models, despite their impressive capabilities, are often trained on massive Internet datasets, making them susceptible to generating content that perpetuates social stereotypes and biases, which can lead to severe consequences. Prior research on assessing bias within image generation models suffers from several shortcomings, including limited accuracy, reliance on extensive human labor, and lack of comprehensive analysis. In this paper, we propose BiasPainter, a novel evaluation framework that can accurately, automatically and comprehensively trigger social bias in image generation models. BiasPainter uses a diverse range of seed images of individuals and prompts the image generation models to edit these images using gender, race, and age-neutral queries. These queries span 62 professions, 39 activities, 57 types of objects, and 70 personality traits. The framework then compares the edited images to the original seed images, focusing on the significant changes related to gender, race, and age. BiasPainter adopts a key insight that these characteristics should not be modified when subjected to neutral prompts. Built upon this design, BiasPainter can trigger the social bias and evaluate the fairness of image generation models. We use BiasPainter to evaluate six widely-used image generation models, such as stable diffusion and Midjourney. Experimental results show that BiasPainter can successfully trigger social bias in image generation models. According to our human evaluation, BiasPainter can achieve 90.8% accuracy on automatic bias detection, which is significantly higher than the results reported in previous work. Wenxuan Wang 0001, Haonan Bai, Jen-tse Huang 0001, Youliang Yuan, Haoyi Qiu, Nanyun Peng 0001, Michael R. Lyu |
ACM Multimedia | 3 |
| 2024 | Apathetic or Empathetic? Evaluating LLMs' Emotional Alignments with HumansabstractEvaluating Large Language Models’ (LLMs) anthropomorphic capabilities has become increasingly important in contemporary discourse. Utilizing the emotion appraisal theory from psychology, we propose to evaluate the empathy ability of LLMs, i.e., how their feelings change when presented with specific situations. After a careful and comprehensive survey, we collect a dataset containing over 400 situations that have proven effective in eliciting the eight emotions central to our study. Categorizing the situations into 36 factors, we conduct a human evaluation involving more than 1,200 subjects worldwide. With the human evaluation results as references, our evaluation includes seven LLMs, covering both commercial and open-source models, including variations in model sizes, featuring the latest iterations, such as GPT-4, Mixtral-8x22B, and LLaMA-3.1. We find that, despite several misalignments, LLMs can generally respond appropriately to certain situations. Nevertheless, they fall short in alignment with the emotional behaviors of human beings and cannot establish connections between similar situations. Our collected dataset of situations, the human evaluation results, and the code of our testing framework, i.e., EmotionBench, are publicly available at https://github.com/CUHK-ARISE/EmotionBench. Jen-tse Huang 0001, Man Ho Lam, Eric John Li, Shujie Ren, Wenxuan Wang 0001, Wenxiang Jiao, Zhaopeng Tu, Michael R. Lyu |
NeurIPS | 1 |
| 2023 | Improving the Transferability of Adversarial Samples by Path-Augmented MethodabstractDeep neural networks have achieved unprecedented success on diverse vision tasks. However, they are vulnerable to adversarial noise that is imperceptible to humans. This phenomenon negatively affects their deployment in real-world scenarios, especially security-related ones. To evaluate the robustness of a target model in practice, transfer-based attacks craft adversarial samples with a local model and have attracted increasing attention from researchers due to their high efficiency. The state-of-the-art transfer-based attacks are generally based on data augmentation, which typically augments multiple training images from a linear path when learning adversarial samples. However, such methods selected the image augmentation path heuristically and may augment images that are semantics-inconsistent with the target images, which harms the transferability of the generated adversarial samples. To overcome the pitfall, we propose the Path-Augmented Method (PAM). Specifically, PAM first constructs a candidate augmentation path pool. It then settles the employed augmentation paths during adversarial sample generation with greedy search. Furthermore, to avoid augmenting semantics-inconsistent images, we train a Semantics Predictor (SP) to constrain the length of the augmentation path. Extensive experiments confirm that PAM can achieve an improvement of over 4.8% on average compared with the state-of-the-art baselines in terms of the attack success rates. Jianping Zhang 0002, Jen-tse Huang 0001, Wenxuan Wang 0001, Yichen Li 0003, Weibin Wu 0002, Xiaosen Wang, Yuxin Su 0001, Michael R. Lyu |
CVPR | 2 |
| 2023 | MTTM: Metamorphic Testing for Textual Content Moderation SoftwareabstractThe exponential growth of social media platforms such as Twitter and Facebook has revolutionized textual communication and textual content publication in human society. However, they have been increasingly exploited to propagate toxic content, such as hate speech, malicious advertisement, and pornography, which can lead to highly negative impacts (e.g., harmful effects on teen mental health). Researchers and practitioners have been enthusiastically developing and extensively deploying textual content moderation software to address this problem. However, we find that malicious users can evade moderation by changing only a few words in the toxic content. Moreover, modern content moderation software's performance against malicious inputs remains underexplored. To this end, we propose MTTM, a Metamorphic Testing framework for Textual content Moderation software. Specifically, we conduct a pilot study on 2, 000 text messages collected from real users and summarize eleven metamorphic relations across three perturbation levels: character, word, and sentence. MTTM employs these metamorphic relations on toxic textual contents to generate test cases, which are still toxic yet likely to evade moderation. In our evaluation, we employ MTTM to test three commercial textual content moderation software and two state-of-the-art moderation algorithms against three kinds of toxic content. The results show that MTTM achieves up to 83.9%, 51%, and 82.5% error finding rates (EFR) when testing commercial moderation software provided by Google, Baidu, and Huawei, respectively, and it obtains up to 91.2% EFR when testing the state-of-the-art algorithms from the academy. In addition, we leverage the test cases generated by MTTM to retrain the model we explored, which largely improves model robustness 0% ~ 5.9% EFR) while maintaining the accuracy on the original test set. A demo can be found in this link1. Wenxuan Wang 0001, Jen-tse Huang 0001, Weibin Wu 0002, Jianping Zhang 0002, Yizhan Huang, Shuqing Li 0001, Pinjia He, Michael R. Lyu |
ICSE | 2 |
| 2023 | An Image is Worth a Thousand Toxic Words: A Metamorphic Testing Framework for Content Moderation SoftwareabstractThe exponential growth of social media platforms has brought about a revolution in communication and content dissemination in human society. Nevertheless, these platforms are being increasingly misused to spread toxic content, including hate speech, malicious advertising, and pornography, leading to severe negative consequences such as harm to teenagers' mental health. Despite tremendous efforts in developing and deploying textual and image content moderation methods, malicious users can evade moderation by embedding texts into images, such as screenshots of the text, usually with some interference. We find that modern content moderation software's performance against such malicious inputs remains underexplored. In this work, we propose OASIS, a metamorphic testing framework for content moderation software. OASIS employs 21 transform rules summarized from our pilot study on 5,000 real-world toxic contents collected from 4 popular social media applications, including Twitter, Instagram, Sina Weibo, and Baidu Tieba. Given toxic textual contents, OASIS can generate image test cases, which preserve the toxicity yet are likely to bypass moderation. In the evaluation, we employ OASIS to test five commercial textual content moderation software from famous companies (i.e., Google Cloud, Microsoft Azure, Baidu Cloud, Alibaba Cloud and Tencent Cloud), as well as a state-of-the-art moderation research model. The results show that OASIS achieves up to 100% error finding rates. Moreover, through retraining the models with the test cases generated by OASIS, the robustness of the moderation model can be improved without performance degradation. Wenxuan Wang 0001, Jen-tse Huang 0001, Jiazhen Gu, Pinjia He, Michael R. Lyu |
ASE | 3 |
| 2022 | Improving Adversarial Transferability via Neuron Attribution-based AttacksabstractDeep neural networks (DNNs) are known to be vulnerable to adversarial examples. It is thus imperative to devise effective attack algorithms to identify the deficiencies of DNNs beforehand in security-sensitive applications. To efficiently tackle the black-box setting where the target model's particulars are unknown, feature-level transfer-based attacks propose to contaminate the intermediate feature outputs of local models, and then directly employ the crafted adversarial samples to attack the target model. Due to the transferability of features, feature-level attacks have shown promise in synthesizing more transferable adversarial samples. However, existing feature-level attacks generally employ inaccurate neuron importance estimations, which deteriorates their transferability. To overcome such pitfalls, in this paper, we propose the Neuron Attribution-based Attack (NAA), which conducts feature-level attacks with more accurate neuron importance estimations. Specifically, we first completely attribute a model's output to each neuron in a middle layer. We then derive an approximation scheme of neuron attribution to tremendously reduce the computation overhead. Finally, we weight neurons based on their attribution results and launch feature-level attacks. Extensive experiments confirm the superiority of our approach to the state-of-the-art benchmarks. Our code is available at: hups.//rgithub.com/jprhang1810/NAA. Jianping Zhang 0002, Weibin Wu 0002, Jen-tse Huang 0001, Yizhan Huang, Wenxuan Wang 0001, Yuxin Su 0001, Michael R. Lyu |
CVPR | 3 |
| 2022 | AEON: a method for automatic evaluation of NLP test casesabstractDue to the labor-intensive nature of manual test oracle construction, various automated testing techniques have been proposed to enhance the reliability of Natural Language Processing (NLP) software. In theory, these techniques mutate an existing test case (e.g., a sentence with its label) and assume the generated one preserves an equivalent or similar semantic meaning and thus, the same label. However, in practice, many of the generated test cases fail to preserve similar semantic meaning and are unnatural (e.g., grammar errors), which leads to a high false alarm rate and unnatural test cases. Our evaluation study finds that 44% of the test cases generated by the state-of-the-art (SOTA) approaches are false alarms. These test cases require extensive manual checking effort, and instead of improving NLP software, they can even degrade NLP software when utilized in model training. To address this problem, we propose AEON for Automatic Evaluation Of NLP test cases. For each generated test case, it outputs scores based on semantic similarity and language naturalness. We employ AEON to evaluate test cases generated by four popular testing techniques on five datasets across three typical NLP tasks. The results show that AEON aligns the best with human judgment. In particular, AEON achieves the best average precision in detecting semantic inconsistent test cases, outperforming the best baseline metric by 10%. In addition, AEON also has the highest average precision of finding unnatural test cases, surpassing the baselines by more than 15%. Moreover, model training with test cases prioritized by AEON leads to models that are more accurate and robust, demonstrating AEON’s potential in improving NLP software. Jen-tse Huang 0001, Jianping Zhang 0002, Wenxuan Wang 0001, Pinjia He, Yuxin Su 0001, Michael R. Lyu |
ISSTA | 1 |