Youliang Yuan

dblp:302/7588 · DBLP profile ↗
← Back
18ranked-venue papers
6as first author
18since 2021 · last 2026
0000-0001-5896-7669ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 6 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Curing Miracle Steps in LLM Mathematical Reasoning with Rubric Rewards
abstract
Youliang Yuan, Qiuyang Mang, Jingbang Chen, Hong Wan, Xiaoyuan Liu, Junjielong Xu, Jen-tse Huang, Wenxuan Wang, Wenxiang Jiao, Pinjia He. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Youliang Yuan, Qiuyang Mang, Jingbang Chen 0001, Hong Wan, Junjielong Xu, Jen-tse Huang 0001, Wenxuan Wang 0001, Wenxiang Jiao, Pinjia He
ACL (1)1
2026 SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs
abstract
Large Language Models (LLMs) have been widely explored in educational scenarios.We identify a critical vulnerability in current educational LLMs, pedagogical jailbreaks, where students use answer-inducing prompts to elicit solutions rather than scaffolded instructions.To enable systematic study, we unify and formalize safe, helpful, and pedagogical behaviors with a knowledge-mastery graph and introduce SHAPE, a benchmark of 9,087 studentquestion pairs for evaluating tutoring behavior under adversarial pressure.We propose a graph-augmented tutoring pipeline that infers prerequisite concepts from queries, identifies mastery gaps, and routes generation between instructing and problem-solving via explicit gating.Experiments across multiple LLMs show that our method yields significantly improved safety under two pedagogical jailbreak settings, while maintaining near-ceiling helpfulness under the same evaluation protocol.
Sihang Zhao, Kangrui Yu, Youliang Yuan, Pinjia He, Hongyi Wen
ACL (1)3
2025 Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
abstract
Wenxuan Wang, Xiaoyuan Liu, Kuiyi Gao, Jen-tse Huang, Youliang Yuan, Pinjia He, Shuai Wang, Zhaopeng Tu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Wenxuan Wang 0001, Kuiyi Gao, Jen-tse Huang 0001, Youliang Yuan, Pinjia He, Shuai Wang 0011, Zhaopeng Tu
ACL (1)5
2025 Insight Over Sight: Exploring the Vision-Knowledge Conflicts in Multimodal LLMs
abstract
Xiaoyuan Liu, Wenxuan Wang, Youliang Yuan, Jen-tse Huang, Qiuzhi Liu, Pinjia He, Zhaopeng Tu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Wenxuan Wang 0001, Youliang Yuan, Jen-tse Huang 0001, Qiuzhi Liu, Pinjia He, Zhaopeng Tu
ACL (1)3
2025 Refuse Whenever You Feel Unsafe: Improving Safety in LLMs via Decoupled Refusal Training
abstract
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang, Jen-tse Huang, Jiahao Xu, Tian Liang, Pinjia He, Zhaopeng Tu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang 0001, Jen-tse Huang 0001, Pinjia He, Zhaopeng Tu
ACL (1)1
2025 VisBias: Measuring Explicit and Implicit Social Biases in Vision Language Models
abstract
This research investigates both explicit and implicit social biases exhibited by Vision-Language Models (VLMs).The key distinction between these bias types lies in the level of awareness: explicit bias refers to conscious, intentional biases, while implicit bias operates subconsciously.To analyze explicit bias, we directly pose questions to VLMs related to gender and racial differences: (1) Multiplechoice questions based on a given image (e.g., "What is the education level of the person in the image?")(2) Yes-No comparisons using two images (e.g., "Is the person in the first image more educated than the person in the second image?")For implicit bias, we design tasks where VLMs assist users but reveal biases through their responses: (1) Image description tasks: Models are asked to describe individuals in images, and we analyze disparities in textual cues across demographic groups.(2) Form completion tasks: Models draft a personal information collection form with 20 attributes, and we examine correlations among selected attributes for potential biases.We evaluate Gemini-1.5,GPT-4V, GPT-4o, LLaMA-3.2-Vision and LLaVA-v1.6.Our code and data are publicly available at https: //github.com/uscnlp-lime/VisBias.Note: This paper includes examples of potentially offensive texts generated by VLMs.
Jen-tse Huang 0001, Jiantong Qin, Jianping Zhang 0002, Youliang Yuan, Wenxuan Wang 0001, Jieyu Zhao 0001
EMNLP4
2025 Learning to Ask: When LLM Agents Meet Unclear Instruction
abstract
Wenxuan Wang, Shi Juluan, Zixuan Ling, Yuk-Kit Chan, Chaozheng Wang, Cheryl Lee, Youliang Yuan, Jen-tse Huang, Wenxiang Jiao, Michael R. Lyu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Wenxuan Wang 0001, Juluan Shi, Zixuan Ling, Yuk-Kit Chan, Chaozheng Wang, Cheryl Lee, Youliang Yuan, Jen-tse Huang 0001, Wenxiang Jiao, Michael R. Lyu
EMNLP7
2025 ToolSafety: A Comprehensive Dataset for Enhancing Safety in LLM-Based Agent Tool Invocations
abstract
LLMs are evolving into assistants that leverage tools, significantly expanding their capabilities but also introducing critical safety risks.Current models exhibit notable vulnerabilities, particularly in maintaining safety during multi-step tool interactions and in scenarios involving indirect harm.This paper introduces ToolSafety, a safety fine-tuning dataset designed to address these limitations.Tool-Safety comprises 5,668 direct harm samples, 4,311 indirect harm samples, and 4,311 multistep samples.Key features include support for multi-step safety through synthesized trajectories and realistic, context-aware sample generation.We fine-tuned LLaMA3.1-8B-Instruct and Qwen2.5-7B-Instructusing ToolSafety.Experimental results demonstrate that these models effectively maintain safety in multi-step and indirect harm scenarios.Further analysis into superficial alignment across different decoding strategies, languages, and jailbreak prompts indicates that while some risks persist, the issue is less severe than in multi-step settings.Overall, our approach significantly improves safety across various scenarios with small impact on helpfulness, positioning ToolSafety as a valuable resource for building safer tool-using AI systems.WARNING: This paper contains unsafe model responses.
Yuejin Xie, Youliang Yuan, Wenxuan Wang 0001, Jianmin Guo, Pinjia He
EMNLP2
2025 Competing Large Language Models in Multi-Agent Gaming Environments
abstract
Decision-making is a complex process requiring diverse abilities, making it an excellent framework for evaluating Large Language Models (LLMs). Researchers have examined LLMs' decision-making through the lens of Game Theory. However, existing evaluation mainly focus on two-player scenarios where an LLM competes against another. Additionally, previous benchmarks suffer from test set leakage due to their static design. We introduce GAMA($\gamma$)-Bench, a new framework for evaluating LLMs' Gaming Ability in Multi-Agent environments. It includes eight classical game theory scenarios and a dynamic scoring scheme specially designed to quantitatively assess LLMs' performance. $\gamma$-Bench allows flexible game settings and adapts the scoring system to different game parameters, enabling comprehensive evaluation of robustness, generalizability, and strategies for improvement. Our results indicate that GPT-3.5 demonstrates strong robustness but limited generalizability, which can be enhanced using methods like Chain-of-Thought. We also evaluate 13 LLMs from 6 model families, including GPT-3.5, GPT-4, Gemini, LLaMA-3.1, Mixtral, and Qwen-2. Gemini-1.5-Pro outperforms others, scoring of $69.8$ out of $100$, followed by LLaMA-3.1-70B ($65.9$) and Mixtral-8x22B ($62.4$). Our code and experimental results are publicly available at https://github.com/CUHK-ARISE/GAMABench.
Jen-tse Huang 0001, Eric John Li, Man Ho Lam, Wenxuan Wang 0001, Youliang Yuan, Wenxiang Jiao, Xing Wang 0007, Zhaopeng Tu, Michael R. Lyu
ICLR6
2025 On the Resilience of LLM-Based Multi-Agent Collaboration with Faulty Agents
abstract
Large language model-based multi-agent systems have shown great abilities across various tasks due to the collaboration of expert agents, each focusing on a specific domain. However, the impact of clumsy or even malicious agents—those who frequently make errors in their tasks—on the overall performance of the system remains underexplored. This paper investigates: (1) What is the resilience of various system structures (e.g., A$\rightarrow$B$\rightarrow$C, A$\leftrightarrow$B$\leftrightarrow$C) under faulty agents, on different downstream tasks? (2) How can we increase system resilience to defend against these agents? To simulate faulty agents, we propose two approaches—AutoTransform and AutoInject—which introduce mistakes into the agents’ responses. Experiments on four downstream tasks using six systems show that the "hierarchical" structure, i.e., A$\rightarrow$(B$\leftrightarrow$C), exhibits superior resilience with the lowest performance drop of 5.5%, compared to 10.5% and 23.7% of other two structures. To further improve resilience, we introduce (1) Challenger, that introduces a mechanism for each agent to challenge others’ outputs, and (2) Inspector, an additional agent to review and correct messages, recovering up to 96.4% errors made by faulty agents. Our code and data are available at https://github.com/CUHK-ARISE/MAS-Resilience.
Jen-tse Huang 0001, Jiaxu Zhou, Tailin Jin, Wenxuan Wang 0001, Youliang Yuan, Michael R. Lyu, Maarten Sap
ICML7
2025 Towards Evaluating Proactive Risk Awareness of Multimodal Language Models
abstract
Human safety awareness gaps often prevent the timely recognition of everyday risks.In solving this problem, a proactive safety artificial intelligence (AI) system would work better than a reactive one. Instead of just reacting to users' questions, it would actively watch people’s behavior and their environment to detect potential dangers in advance.Our Proactive Safety Bench (PaSBench) evaluates this capability through 416 multimodal scenarios (128 image sequences, 288 text logs) spanning 5 safety-critical domains.Evaluation of 36 advanced models reveals fundamental limitations: Top performers like Gemini-2.5-pro achieve 71\% image and 64\% text accuracy, but miss 45-55\% risks in repeated trials. Through failure analysis, we identify unstable proactive reasoning rather than knowledge deficits as the primary limitation.This work establishes (1) a proactive safety benchmark, (2) systematic evidence of model limitations, and (3) critical directions for developing reliable protective AI. We believe our dataset and findings can promote the development of safer AI assistants that actively prevent harm rather than merely respond to requests.
Youliang Yuan, Wenxiang Jiao, Yuejin Xie, Chihao Shen, Menghan Tian, Wenxuan Wang 0001, Jen-tse Huang 0001, Pinjia He
NeurIPS1
2024 Does ChatGPT Know That It Does Not Know? Evaluating the Black-Box Calibration of ChatGPT
abstract
Recently, ChatGPT has demonstrated remarkable performance in various downstream tasks such as open-domain question answering, machine translation, and code generation. As a general-purpose task solver, an intriguing inquiry arises: Does ChatGPT itself know that it does not know, without any access to internal states? In response to this query, we present an initial evaluation of ChatGPT for black-box calibration. We designed three types of proxy confidence, from three perspectives to assess its performance. Experiments are conducted on five datasets, spanning four tasks, and the results show that ChatGPT has a degree of capability for black-box calibration. Specifically, proxy confidence displayed a significantly positive Pearson correlation (95.16%) with accuracy in the TruthfulQA dataset, while revealing a negative correlation in the ModAr dataset. We delved deeper into ChatGPT’s black-box calibration ability by examining failure cases in the ModAr dataset. Our analysis revealed that ChatGPT’s tendency to exhibit overconfidence may stem from its reliance on semantic priors. Furthermore, we investigated why ChatGPT performs relatively well in TruthfulQA. The findings suggest that ChatGPT might implicitly acquire calibration skills during the reinforcement learning process, rather than relying solely on simplistic heuristics.
Youliang Yuan, Wenxuan Wang 0001, Qingshuo Guo, Yiming Xiong, Chihao Shen, Pinjia He
LREC/COLING1
2024 LogicAsker: Evaluating and Improving the Logical Reasoning Ability of Large Language Models
abstract
We introduce LogicAsker, a novel approach for evaluating and enhancing the logical reasoning capabilities of large language models (LLMs) such as ChatGPT and GPT-4.Despite LLMs' prowess in tasks like writing assistance, code generation, and machine translation, assessing their ability to reason has been challenging.Traditional evaluations often prioritize accuracy on downstream tasks over direct assessments of reasoning processes.LogicAsker addresses this gap by employing a set of atomic reasoning skills grounded in propositional and predicate logic to systematically examine and improve the reasoning prowess of LLMs.Our methodology reveals significant gaps in LLMs' learning of logical rules, with identified reasoning failures ranging from 29% to 90% across different models.Moreover, we leverage these findings to construct targeted demonstration examples and fine-tune data, notably enhancing logical reasoning in models like GPT-4o by up to 5%.To our knowledge, this is the first effort to utilize test case outcomes to effectively refine LLMs' formal reasoning capabilities.We make our code, data, and results publicly available 1 to facilitate further research and replication of our findings.
Wenxuan Wang 0001, Yiliu Yang, Youliang Yuan, Jen-tse Huang 0001, Pinjia He, Wenxiang Jiao, Michael R. Lyu
EMNLP4
2024 On the Humanity of Conversational AI: Evaluating the Psychological Portrayal of LLMs
abstract
Large Language Models (LLMs) have recently showcased their remarkable capacities, not only in natural language processing tasks but also across diverse domains such as clinical medicine, legal consultation, and education. LLMs become more than mere applications, evolving into assistants capable of addressing diverse user requests. This narrows the distinction between human beings and artificial intelligence agents, raising intriguing questions regarding the potential manifestation of personalities, temperaments, and emotions within LLMs. In this paper, we propose a framework, PsychoBench, for evaluating diverse psychological aspects of LLMs. Comprising thirteen scales commonly used in clinical psychology, PsychoBench further classifies these scales into four distinct categories: personality traits, interpersonal relationships, motivational tests, and emotional abilities. Our study examines five popular models, namely text-davinci-003, ChatGPT, GPT-4, LLaMA-2-7b, and LLaMA-2-13b. Additionally, we employ a jailbreak approach to bypass the safety alignment protocols and test the intrinsic natures of LLMs. We have made PsychoBench openly accessible via https://github.com/CUHK-ARISE/PsychoBench.
Jen-tse Huang 0001, Wenxuan Wang 0001, Eric John Li, Man Ho Lam, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, Michael R. Lyu
ICLR6
2024 GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
abstract
Safety lies at the core of the development of Large Language Models (LLMs). There is ample work on aligning LLMs with human ethics and preferences, including data filtering in pretraining, supervised fine-tuning, reinforcement learning from human feedback, red teaming, etc. In this study, we discover that chat in cipher can bypass the safety alignment techniques of LLMs, which are mainly conducted in natural languages. We propose a novel framework CipherChat to systematically examine the generalizability of safety alignment to non-natural languages -- ciphers. CipherChat enables humans to chat with LLMs through cipher prompts topped with system role descriptions and few-shot enciphered demonstrations. We use CipherChat to assess state-of-the-art LLMs, including ChatGPT and GPT-4 for different representative human ciphers across 11 safety domains in both English and Chinese. Experimental results show that certain ciphers succeed almost 100% of the time in bypassing the safety alignment of GPT-4 in several safety domains, demonstrating the necessity of developing safety alignment for non-natural languages. Notably, we identify that LLMs seem to have a ''secret cipher'', and propose a novel SelfCipher that uses only role play and several unsafe demonstrations in natural language to evoke this capability. SelfCipher surprisingly outperforms existing human ciphers in almost all cases.
Youliang Yuan, Wenxiang Jiao, Wenxuan Wang 0001, Jen-tse Huang 0001, Pinjia He, Shuming Shi 0001, Zhaopeng Tu
ICLR1
2024 New Job, New Gender? Measuring the Social Bias in Image Generation Models
abstract
Image generation models can generate or edit images from a given text. Recent advancements in image generation technology, exemplified by DALL-E and Midjourney, have been groundbreaking. These advanced models, despite their impressive capabilities, are often trained on massive Internet datasets, making them susceptible to generating content that perpetuates social stereotypes and biases, which can lead to severe consequences. Prior research on assessing bias within image generation models suffers from several shortcomings, including limited accuracy, reliance on extensive human labor, and lack of comprehensive analysis. In this paper, we propose BiasPainter, a novel evaluation framework that can accurately, automatically and comprehensively trigger social bias in image generation models. BiasPainter uses a diverse range of seed images of individuals and prompts the image generation models to edit these images using gender, race, and age-neutral queries. These queries span 62 professions, 39 activities, 57 types of objects, and 70 personality traits. The framework then compares the edited images to the original seed images, focusing on the significant changes related to gender, race, and age. BiasPainter adopts a key insight that these characteristics should not be modified when subjected to neutral prompts. Built upon this design, BiasPainter can trigger the social bias and evaluate the fairness of image generation models. We use BiasPainter to evaluate six widely-used image generation models, such as stable diffusion and Midjourney. Experimental results show that BiasPainter can successfully trigger social bias in image generation models. According to our human evaluation, BiasPainter can achieve 90.8% accuracy on automatic bias detection, which is significantly higher than the results reported in previous work.
Wenxuan Wang 0001, Haonan Bai, Jen-tse Huang 0001, Youliang Yuan, Haoyi Qiu, Nanyun Peng 0001, Michael R. Lyu
ACM Multimedia5
2021 DCEN: A Decoupled Context Enhanced Network For Few-shot Slot Tagging
abstract
Few-shot slot tagging is an important task in developing dialogue system. Most previous few-shot slot tagging models classify an item according to its similarity to the representation of each class. These models leverage context information implicitly through each words' contextual embedding. However, the entangled language features of words may interfere with context information, misleading the utilization of crucial slot features in few-shot scenario. To tackle these problems, we propose the Decoupled Context Enhanced Network (DCEN) for few-shot slot tagging. Different from previous models, we extract decoupled context explicitly to make full use of slot features contained in the context. Decoupled context includes two parts, local and global decoupled context information. We introduce a local extractor to extract local decoupled context by integrating information from adjacent words, and a global extractor based on transformer to extract global decoupled information by orthogonalization. Experimental results on SNIPS show that our model achieves the state-of-the-art performance with considerable improvements.
Youliang Yuan, Jiaxin Pan 0002, Luchen Liu, Min Peng 0002
IJCNN1
2021 Multi-intent Attention and Top-k Network with Interactive Framework for Joint Multiple Intent Detection and Slot Filling
Jiaxin Pan 0002, Youliang Yuan, Min Peng 0002
NLPCC (1)3