Weixiang Zhao

dblp:38/6433 · DBLP profile ↗
← Back
24ranked-venue papers
16as first author
22since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 13 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 CultureRL: Internalizing Cultural Principles in Large Language Models via Norm-Driven Reinforcement Learning
abstract
As large language models (LLMs) are increasingly deployed across culturally diverse regions, ensuring that their responses align with users’ cultural norms has become a critical challenge. Existing approaches to cultural alignment primarily rely on prompting or data-augmentation-based supervised finetuning, which teach models to follow norms indirectly through example-based supervision. However, these methods are difficult to scale and often fail to generalize, particularly in low-resource cultural settings. In this work, we propose CultureRL, a culture-norm-driven reinforcement learning framework that directly encodes cultural principles into model behavior. Rather than relying on output imitation, CultureRL provides normative feedback during training, enabling the model to internalize high-level cultural rules. It consists of two key components: (1) Norm Pool Construction (NPC), which clusters data from the World Values Survey into abstract cultural concepts to form a structured and retrievable norm pool; and (2) Norm Cluster-based Reward Mechanism (NCRM), which retrieves the relevant norm for each input and uses an external reward model to assess conformity, guiding model updates through cultural alignment. We evaluate CultureRL in both one-for-one (per-culture) and one-for-all (multi-culture) settings across nine cultures and three benchmarks. Results show that CultureRL consistently outperforms strong baselines, especially in terms of cultural consistency and adaptability.
Weixiang Zhao, Haixiao Liu, Biye Li, Ting Liu 0001, Bing Qin 0001
AAAI1
2026 Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities
abstract
Recent advancements in Large Reasoning Models (LRMs), such as OpenAI's o1/o3 and DeepSeek-R1, have demonstrated remarkable performance in specialized reasoning tasks through human-like deliberative thinking and long chain-of-thought reasoning. However, our systematic evaluation across various model families (DeepSeek, Qwen, and LLaMA) and scales (7B to 32B) reveals that acquiring these deliberative reasoning capabilities significantly reduces the foundational capabilities of LRMs, including notable declines in helpfulness and harmlessness, alongside substantially increased inference costs. Importantly, we demonstrate that adaptive reasoning---employing modes like Zero-Thinking, Less-Thinking, and Summary-Thinking---can effectively alleviate these drawbacks. Our empirical insights underline the critical need for developing more versatile LRMs capable of dynamically allocating inference-time compute according to specific task characteristics.
Weixiang Zhao, Xingyu Sui, Jiahe Guo, Yulin Hu, Yang Deng 0002, Xuda Zhi, Yongbo Huang, Wanxiang Che, Ting Liu 0001, Bing Qin 0001
AAAI1
2026 When Personalization Legitimizes Risks: Uncovering Safety Vulnerabilities in Personalized Dialogue Agents
abstract
Jiahe Guo, Xiangran Guo, Yulin Hu, Zimo Long, Xingyu Sui, Xuda Zhi, Yongbo Huang, Hao He, Weixiang Zhao, Yanyan Zhao, Bing Qin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiahe Guo, Xiangran Guo, Yulin Hu, Zimo Long, Xingyu Sui, Xuda Zhi, Yongbo Huang, Weixiang Zhao, Bing Qin 0001
ACL (1)9
2026 TEA-Bench: A Systematic Benchmarking of Tool-enhanced Emotional Support Dialogue Agent
abstract
Emotional Support Conversation requires not only affective expression but also grounded instrumental support to provide trustworthy guidance.However, existing ESC systems and benchmarks largely focus on affective support in text-only settings, overlooking how external tools can enable factual grounding and reduce hallucination in multi-turn emotional support.We introduce TEA-Bench, the first interactive benchmark for evaluating tool-augmented agents in ESC, featuring realistic emotional scenarios, an MCP-style tool environment, and process-level metrics that jointly assess the quality and factual grounding of emotional support.Experiments on nine LLMs show that tool augmentation generally improves emotional support quality and reduces hallucination, but the gains are strongly capacity-dependent: stronger models use tools more selectively and effectively, while weaker models benefit only marginally.We further release TEA-Dialog, a dataset of toolenhanced ESC dialogues, and find that supervised fine-tuning improves in-distribution support but generalizes poorly.Our results underscore the importance of tool use in building reliable emotional support agents. 1
Xingyu Sui, Yulin Hu, Jiahe Guo, Weixiang Zhao, Bing Qin 0001
ACL (1)5
2026 The gains do not make up for the losses: a comprehensive evaluation for safety alignment of large language models via machine unlearning
abstract
Abstract Machine Unlearning (MU) has emerged as a promising technique for aligning large language models (LLMs) with safety requirements to steer them forgetting specific harmful contents. Despite the significant progress in previous studies, we argue that the current evaluation criteria, which solely focus on safety evaluation, are actually impractical and biased , leading to concerns about the true effectiveness of MU techniques. To address this, we propose to comprehensively evaluate LLMs after MU from three aspects: safety, over-safety, and general utility. Specifically, a novel benchmark M u B ench with 18 related datasets is first constructed, where the safety is measured with both vanilla harmful inputs and 10 types of jailbreak attacks. Furthermore, we examine whether MU introduces side effects, focusing on over-safety and utility-loss. Extensive experiments are performed on 3 popular LLMs with 7 recent MU methods. The results highlight a challenging trilemma in safety alignment without side effects, indicating that there is still considerable room for further exploration. M u B ench serves as a comprehensive benchmark, fostering future research on MU for safety alignment of LLMs.
Weixiang Zhao, Yulin Hu, Xingyu Sui, Zhuojun Li, Yang Deng 0002, Bing Qin 0001, Wanxiang Che
Frontiers Comput. Sci.1
2025 Beware of Your Po! Measuring and Mitigating AI Safety Risks in Role-Play Fine-Tuning of LLMs
abstract
Weixiang Zhao, Yulin Hu, Yang Deng, Jiahe Guo, Xingyu Sui, Xinyang Han, An Zhang, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Weixiang Zhao, Yulin Hu, Yang Deng 0002, Jiahe Guo, Xingyu Sui, An Zhang 0003, Bing Qin 0001, Tat-Seng Chua, Ting Liu 0001
ACL (1)1
2025 MPO: Multilingual Safety Alignment via Reward Gap Optimization
abstract
Weixiang Zhao, Yulin Hu, Yang Deng, Tongtong Wu, Wenxuan Zhang, Jiahe Guo, An Zhang, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Weixiang Zhao, Yulin Hu, Yang Deng 0002, Tongtong Wu, Wenxuan Zhang 0001, Jiahe Guo, An Zhang 0003, Bing Qin 0001, Tat-Seng Chua, Ting Liu 0001
ACL (1)1
2025 AdaSteer: Your Aligned LLM is Inherently an Adaptive Jailbreak Defender
abstract
Weixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng, An Zhang, Xingyu Sui, Xinyang Han, Yanyan Zhao, Bing Qin, Tat-Seng Chua, Ting Liu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Weixiang Zhao, Jiahe Guo, Yulin Hu, Yang Deng 0002, An Zhang 0003, Xingyu Sui, Bing Qin 0001, Tat-Seng Chua, Ting Liu 0001
EMNLP1
2025 INN-based Secure Steganography Using Lost Information as Adversarial Perturbations
abstract
Recently image steganography methods based on invertible neural networks (INNs) demonstrated the capability to automatically embed and extract secret messages while maintaining high visual quality in stego images. However, there remain concerns about security and invertibility of such methods. In this paper, for the first time, we introduce adversarial hiding into INN-based image steganography method to simultaneously perform steganographic embedding and adversarial perturbation generation, resulting in improved security. Our method enhances the invertibility of the INN structure: It utilizes the lost information of the INN to generate perturbations, which are then combined with the gradient of the cover image to produce an adversarial stego image. Also, a learnable noise layer is proposed to mitigate information loss caused by rounding and truncation during image storage. Therefore, the proposed method significantly improves security while enhancing extraction performance of INN-based steganography approach, as supported by our experimental results. For example, the steganalysis detection accuracy of SRNet decreases from 96.86% to 51.77% at a payload of 0.2 bits per pixel (bpp).
Fei Shang, Weixiang Zhao, Xiangui Kang, Z. Jane Wang 0001
ICASSP2
2025 Secure INN-based Steganography via Model Smoothing and Adversarial Attacks
abstract
In recent years, image steganography methods based on invertible neural networks (INNs) have received significant attention due to their invertible structure, which offers advantages in embedding and extracting secret messages. However, current INN-based image steganography methods face challenges, particularly their limited tolerance against noise interference (e.g., added Gaussian noise, adversarial perturbations, and JPEG compression) and vulnerability to detection by advanced deep steganalyzers. To address these concerns, we present a novel steganography framework that combines Median Smoothing Training (MST) with dynamic Projected Gradient Descent (d-PGD). Specifically, our method begins with employing an MST strategy during the training phase to improve the INN’s tolerance to noise, ensuring that accurate message extraction even under noise interference. Subsequently, to improve the security of INN-based steganography, we propose a d-PGD algorithm that can generate minimal adversarial perturbations capable of deceiving deep steganalyzers, thereby improving security without compromising extraction accuracy. Experimental results demonstrate that our method achieves state-of-the-art secret message extraction accuracy while significantly improving resistance against deep steganalyzers.
Weixiang Zhao, Fei Shang, Jingyang Wen, Xiangui Kang, Z. Jane Wang 0001
MMSP1
2025 L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language Models
abstract
Large language models (LLMs) have achieved notable progress. Despite their success, next-token prediction (NTP), the dominant method for LLM training and inference, is constrained in both contextual coverage and inference efficiency due to its inherently sequential process. To overcome these challenges, we propose leap multi-token prediction~(L-MTP), an innovative token prediction method that extends the capabilities of multi-token prediction (MTP) by introducing a leap-based mechanism. Unlike conventional MTP, which generates multiple tokens at adjacent positions, L-MTP strategically skips over intermediate tokens, predicting non-sequential ones in a single forward pass. This structured leap not only enhances the model's ability to capture long-range dependencies but also enables a decoding strategy specially optimized for non-sequential leap token generation, effectively accelerating inference. We theoretically demonstrate the benefit of L-MTP in improving inference efficiency. Experiments across diverse benchmarks validate its merit in boosting both LLM performance and inference speed. The source code is available at https://github.com/Xiaohao-Liu/L-MTP.
Xiaohao Liu, Xiaobo Xia, Weixiang Zhao, Manyi Zhang, Xianzhi Yu, Xiu Su, Shuo Yang 0006, See-Kiong Ng, Tat-Seng Chua
NeurIPS3
2025 On Reasoning Strength Planning in Large Reasoning Models
abstract
Recent studies empirically reveal that large reasoning models (LRMs) can automatically allocate more reasoning strengths (\ie the number of reasoning tokens) for harder problems, exhibiting difficulty-awareness for better task performance. While this automatic reasoning strength allocation phenomenon has been widely observed, its underlying mechanism remains largely unexplored. To this end, we provide explanations for this phenomenon from the perspective of model activations. \textbf{We find evidence that LRMs pre-plan the reasoning strengths in their activations even before generation, with this reasoning strength causally controlled by the magnitude of a pre-allocated directional vector.} Specifically, we show that the number of reasoning tokens is predictable solely based on the question activations using linear probes, indicating that LRMs estimate the required reasoning strength in advance. We then uncover that LRMs encode this reasoning strength through a pre-allocated directional vector embedded in the activations of the model, where the vector’s magnitude modulates the reasoning strength. Subtracting this vector can lead to reduced reasoning token number and performance, while adding this vector can lead to increased reasoning token number and even improved performance. We further reveal that this direction vector consistently yields positive reasoning length prediction, and it modifies the logits of end-of-reasoning token \texttt{</think>} to affect the reasoning length. Finally, we demonstrate two potential applications of our findings: overthinking behavior detection and enabling efficient reasoning on simple problems. Our work provides new insights into the internal mechanisms of reasoning in LRMs and offers practical tools for controlling their reasoning behaviors. Our code is available at \url{https://anonymous.4open.science/r/LRM-plans-CoT-7E04}.
Leheng Sheng, An Zhang 0003, Zijian Wu 0003, Weixiang Zhao, Changshuo Shen, Yi Zhang 0001, Xiang Wang 0010, Tat-Seng Chua
NeurIPS4
2025 When Less Language is More: Language-Reasoning Disentanglement Makes LLMs Better Multilingual Reasoners
abstract
Multilingual reasoning remains a significant challenge for large language models (LLMs), with performance disproportionately favoring high-resource languages. Drawing inspiration from cognitive neuroscience, which suggests that human reasoning functions largely independently of language processing, we hypothesize that LLMs similarly encode reasoning and language as separable components that can be disentangled to enhance multilingual reasoning. To evaluate this, we perform a causal intervention by ablating language-specific representations at inference time. Experiments on 10 open-weight LLMs spanning 11 typologically diverse languages show that this language-specific ablation consistently boosts multilingual reasoning performance. Layer-wise analyses further confirm that language and reasoning representations can be effectively disentangled throughout the model, yielding improved multilingual reasoning capabilities, while preserving top-layer language features remains essential for maintaining linguistic fidelity. Compared to post-training methods such as supervised fine-tuning or reinforcement learning, our training-free language-reasoning disentanglement achieves comparable or superior results with minimal computational overhead. These findings shed light on the internal mechanisms underlying multilingual reasoning in LLMs and suggest a lightweight and interpretable strategy for improving cross-lingual generalization.
Weixiang Zhao, Jiahe Guo, Yang Deng 0002, Tongtong Wu, Wenxuan Zhang 0001, Yulin Hu, Xingyu Sui, Wanxiang Che, Bing Qin 0001, Tat-Seng Chua, Ting Liu 0001
NeurIPS1
2025 Teaching Language Models to Evolve with Users: Dynamic Profile Modeling for Personalized Alignment
abstract
Personalized alignment is essential for enabling large language models (LLMs) to engage effectively in user-centric dialogue. While recent prompt-based and offline optimization methods offer preliminary solutions, they fall short in cold-start scenarios and long-term personalization due to their inherently static and shallow designs. In this work, we introduce the Reinforcement Learning for Personalized Alignment (RLPA) framework, in which an LLM interacts with a simulated user model to iteratively infer and refine user profiles through dialogue. The training process is guided by a dual-level reward structure: the Profile Reward encourages accurate construction of user representations, while the Response Reward incentivizes generation of responses consistent with the inferred profile. We instantiate RLPA by fine-tuning Qwen-2.5-3B-Instruct, resulting in Qwen-RLPA, which achieves state-of-the-art performance in personalized dialogue. Empirical evaluations demonstrate that Qwen-RLPA consistently outperforms prompting and offline fine-tuning baselines, and even surpasses advanced commercial models such as Claude-3.5 and GPT-4o. Further analysis highlights Qwen-RLPA's robustness in reconciling conflicting user preferences, sustaining long-term personalization and delivering more efficient inference compared to recent reasoning-focused LLMs. These results emphasize the potential of dynamic profile inference as a more effective paradigm for building personalized dialogue systems.
Weixiang Zhao, Xingyu Sui, Yulin Hu, Jiahe Guo, Haixiao Liu, Biye Li, Bing Qin 0001, Ting Liu 0001
NeurIPS1
2025 RSafe: Incentivizing proactive reasoning to build robust and adaptive LLM safeguards
abstract
Large Language Models (LLMs) continue to exhibit vulnerabilities despite deliberate safety alignment efforts, posing significant risks to users and society. To safeguard against the risk of policy-violating content, system-level moderation via external guard models—designed to monitor LLM inputs and outputs and block potentially harmful content—has emerged as a prevalent mitigation strategy. Existing approaches of training guard models rely heavily on extensive human curated datasets and struggle with out-of-distribution threats, such as emerging harmful categories or jailbreak attacks. To address these limitations, we propose RSafe, an adaptive reasoning-based safeguard that conducts guided safety reasoning to provide robust protection within the scope of specified safety policies. RSafe operates in two stages: (1) guided reasoning, where it analyzes safety risks of input content through policy-guided step-by-step reasoning, and (2) reinforced alignment, where rule-based RL optimizes its reasoning paths to align with accurate safety prediction. This two-stage training paradigm enables RSafe to internalize safety principles to generalize safety protection capability over unseen or adversarial safety violation scenarios. During inference, RSafe accepts user-specified safety policies to provide enhanced safeguards tailored to specific safety requirements. Experiments demonstrate that RSafe matches state-of-the-art guard models using limited amount of public data in both prompt- and response-level harmfulness detection, while achieving superior out-of-distribution generalization on both emerging harmful category and jailbreak attacks. Furthermore, RSafe provides human-readable explanations for its safety judgments for better interpretability. RSafe offers a robust, adaptive, and interpretable solution for LLM safety moderation, advancing the development of reliable safeguards in dynamic real-world environments. Our code is available at https://anonymous.4open.science/r/RSafe-996D.
Jingnan Zheng, Xiangtian Ji, Chenhang Cui, Weixiang Zhao, Gelei Deng, Zhenkai Liang, An Zhang 0003, Tat-Seng Chua
NeurIPS5
2025 A parental emotion coaching dialogue assistant for better parent-child interaction
Weixiang Zhao, Shilong Wang 0003, Yanpeng Tong, Zhuojun Li, Chenxue Wang, Bing Qin 0001
Sci. China Inf. Sci.1
2024 SAPT: A Shared Attention Framework for Parameter-Efficient Continual Learning of Large Language Models
abstract
Weixiang Zhao, Shilong Wang, Yulin Hu, Yanyan Zhao, Bing Qin, Xuanyu Zhang, Qing Yang, Dongliang Xu, Wanxiang Che. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Weixiang Zhao, Shilong Wang 0003, Yulin Hu, Bing Qin 0001, Qing Yang 0033, Dongliang Xu, Wanxiang Che
ACL (1)1
2023 Knowledge-Bridged Causal Interaction Network for Causal Emotion Entailment
abstract
Causal Emotion Entailment aims to identify causal utterances that are responsible for the target utterance with a non-neutral emotion in conversations. Previous works are limited in thorough understanding of the conversational context and accurate reasoning of the emotion cause. To this end, we propose Knowledge-Bridged Causal Interaction Network (KBCIN) with commonsense knowledge (CSK) leveraged as three bridges. Specifically, we construct a conversational graph for each conversation and leverage the event-centered CSK as the semantics-level bridge (S-bridge) to capture the deep inter-utterance dependencies in the conversational context via the CSK-Enhanced Graph Attention module. Moreover, social-interaction CSK serves as emotion-level bridge (E-bridge) and action-level bridge (A-bridge) to connect candidate utterances with the target one, which provides explicit causal clues for the Emotional Interaction module and Actional Interaction module to reason the target emotion. Experimental results show that our model achieves better performance over most baseline models. Our source code is publicly available at https://github.com/circle-hit/KBCIN.
Weixiang Zhao, Zhuojun Li, Bing Qin 0001
AAAI1
2023 A Topic-Enhanced Approach for Emotion Distribution Forecasting in Conversations
abstract
Emotion Forecasting in Conversations (EFC), the task aims to predict the emotion of next utterance (yet to come), has received more and more attention in recent years. However, this task ignores the one-to-many feature of dialogue and its prediction target is emotion label, which is flawed in most cases. In this work, we propose a new task: Emotion Distribution Forecasting in Conversations (EDFC), which aims to predict the emotion distribution of next utterance. Although this task is more reasonable in real applications, it can only learn using emotion labels in most cases because of the difficulty in obtaining emotion distribution. To address it, we explore the positive role of topic in this task and propose a topic-enhanced approach. Specifically, we first obtain the topic-based emotion distribution prior through topic model and emotion generation model, and then use the emotion distribution prior to enhance original label learning model. To effectively evaluate the distribution prediction results, we construct two datasets for this task, and the experimental results prove the feasibility of the EDFC task as well as the effectiveness of our approach.
Weixiang Zhao, Bing Qin 0001
ICASSP2
2022 MuCDN: Mutual Conversational Detachment Network for Emotion Recognition in Multi-Party Conversations
abstract
As an emerging research topic in natural language processing community, emotion recognition in multi-party conversations has attained increasing interest. Previous approaches that focus either on dyadic or multi-party scenarios exert much effort to cope with the challenge of emotional dynamics and achieve appealing results. However, since emotional interactions among speakers are often more complicated within the entangled multi-party conversations, these works are limited in capturing effective emotional clues in conversational context. In this work, we propose Mutual Conversational Detachment Network (MuCDN) to clearly and effectively understand the conversational context by separating conversations into detached threads. Specifically, two detachment ways are devised to perform context and speaker-specific modeling within detached threads and they are bridged through a mutual module. Experimental results on two datasets show that our model achieves better performance over the baseline models.
Weixiang Zhao, Bing Qin 0001
COLING1
2022 CauAIN: Causal Aware Interaction Network for Emotion Recognition in Conversations
abstract
Emotion Recognition in Conversations has attained increasing interest in the natural language processing community. Many neural-network based approaches endeavor to solve the challenge of emotional dynamics in conversations and gain appealing results. However, these works are limited in capturing deep emotional clues in conversational context because they ignore the emotion cause that could be viewed as stimulus to the target emotion. In this work, we propose Causal Aware Interaction Network (CauAIN) to thoroughly understand the conversational context with the help of emotion cause detection. Specifically, we retrieve causal clues provided by commonsense knowledge to guide the process of causal utterance traceback. Both retrieve and traceback steps are performed from the perspective of intra- and inter-speaker interaction simultaneously. Experimental results on three benchmark datasets show that our model achieves better performance over most baseline models.
Weixiang Zhao
IJCAI1
2021 An Aspect-Centralized Graph Convolutional Network for Aspect-Based Sentiment Classification
Weixiang Zhao, Bing Qin 0001
NLPCC (2)1
2020 Detecting Problem Statements in Peer Assessments
Yunkai Xiao, Gabriel Zingle, Qinjin Jia, Harsh R. Shah, Mohsin Karovaliya, Weixiang Zhao, Yang Song 0019, Ashwin Balasubramaniam, Harshit Patel, Priyankha Bhalasubbramanian, Vikram Patel, Edward F. Gehringer
EDM8
2011 A modified artificial immune system based pattern recognition approach - An application to clinical diagnostics
Weixiang Zhao, Cristina E. Davis
Artif. Intell. Medicine1