Xinwei Wu 0001

dblp:255/3696-1 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0009-0001-2167-128XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Finding the Translation Switch: Discovering and Exploiting the Task-Initiation Features in LLMs
abstract
Large Language Models (LLMs) frequently exhibit strong translation abilities, even without task-specific fine-tuning. However, the internal mechanisms governing this innate capability remain largely opaque. To demystify this process, we leverage Sparse Autoencoders (SAEs) and introduce a novel framework for identifying task-specific features. Our method first recalls features that are frequently co-activated on translation inputs and then filters them for functional coherence using a PCA-based consistency metric. This framework successfully isolates a small set of "translation initiation" features. Causal interventions demonstrate that amplifying these features steers the model towards correct translation, while ablating them induces hallucinations and off-task outputs, confirming they represent a core component of the model's innate translation competency. Moving from analysis to application, we leverage this mechanistic insight to propose a new data selection strategy for efficient fine-tuning. Specifically, we prioritize training on "mechanistically hard" samples—those that fail to naturally activate the translation initiation features. Experiments show this approach significantly improves data efficiency and suppresses hallucinations. Furthermore, we find these mechanisms are transferable to larger models of the same family. Our work not only decodes a core component of the translation mechanism in LLMs but also provides a blueprint for using internal model mechanism to create more robust and efficient models.
Xinwei Wu 0001, Yuqi Ren, Linlong Xu, Longyue Wang, Deyi Xiong, Weihua Luo, Kaifu Zhang
AAAI1
2026 From Insight to Action: A Novel Framework for Interpretability-Guided Data Selection in Large Language Models
abstract
Ling Shi, Xinwei Wu, Xiaohu Zhao, Hao Wang, Heng Liu, Yangyang Liu, Linlong Xu, Longyue Wang, Deyi Xiong, Weihua Luo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ling Shi 0004, Xinwei Wu 0001, Linlong Xu, Longyue Wang, Deyi Xiong, Weihua Luo
ACL (1)2
2026 M²PO: Multi-Perspective Multi-Pair Preference Optimization for Machine Translation
abstract
Hao Wang, Linlong Xu, Heng Liu, Yangyang Liu, Xiaohu Zhao, Bo Zeng, Liangying Shao, Yichen Dong, Xinwei Wu, Jiang Zhou, Tianyu Dong, Xiangxiang Zeng, Longyue Wang, Weihua Luo. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Linlong Xu, Liangying Shao, Yichen Dong, Xinwei Wu 0001, Tianyu Dong, Xiangxiang Zeng, Longyue Wang, Weihua Luo
ACL (1)9
2026 Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity Translation
abstract
Jiang Zhou, Xiaohu Zhao, Xinwei Wu, Tianyu Dong, Hao Wang, Yangyang Liu, Heng Liu, Linlong Xu, Longyue Wang, Weihua Luo, Deyi Xiong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xinwei Wu 0001, Tianyu Dong, Linlong Xu, Longyue Wang, Weihua Luo, Deyi Xiong
ACL (1)3
2025 CONTRANS: Weak-to-Strong Alignment Engineering via Concept Transplantation
abstract
Ensuring large language models (LLM) behave consistently with human goals, values, and intentions is crucial for their safety but yet computationally expensive. To reduce the computational cost of alignment training of LLMs, especially for those with a huge number of parameters, and to reutilize learned value alignment, we propose ConTrans, a novel framework that enables weak-to-strong alignment transfer via concept transplantation. From the perspective of representation engineering, ConTrans refines concept vectors in value alignment from a source LLM (usually a weak yet aligned LLM). The refined concept vectors are then reformulated to adapt to the target LLM (usually a strong yet unaligned base LLM) via affine transformation. In the third step, ConTrans transplants the reformulated concept vectors into the residual stream of the target LLM. Experiments demonstrate the successful transplantation of a wide range of aligned concepts from 7B models to 13B and 70B models across multiple LLMs and LLM families. Remarkably, ConTrans even surpasses instruction-tuned models in terms of truthfulness. Experiment results validate the effectiveness of both inter-LLM-family and intra-LLM-family concept transplantation. Our work successfully demonstrates an alternative way to achieve weak-to-strong alignment generalization and control.
Weilong Dong, Xinwei Wu 0001, Renren Jin, Shaoyang Xu, Deyi Xiong
COLING2
2025 Towards a Unified Paradigm of Concept Editing in Large Language Models
abstract
Concept editing aims to control specific concepts in large language models (LLMs) and is an emerging subfield of model editing.Despite the emergence of various editing methods in recent years, there remains a lack of rigorous theoretical analysis and a unified perspective to systematically understand and compare these methods.To address this gap, we propose a unified paradigm for concept editing methods, in which all forms of conceptual injection are aligned at the neuron level.We study four representative concept editing methods: Neuron Editing (NE), Supervised Fine-tuning (SFT), Sparse Autoencoder (SAE), and Steering Vector (SV).Then we categorize them into two classes based on their mode of conceptual information injection: indirect (NE, SFT) and direct (SAE, SV).We evaluate above methods along four dimensions: editing reliability, output generalization, neuron level consistency, and mathematical formalization.Experiments show that SAE achieves the best editing reliability.In output generalization, SAE captures features closer to human-understood concepts, while NE tends to locate text patterns rather than true semantics.Neuron-level analysis reveals that direct methods share high neuron overlap, as do indirect methods, indicating methodological commonality within each category.Our unified paradigm offers a clear framework and valuable insights for advancing interpretability and controlled generation in LLMs.
Zhuowen Han, Xinwei Wu 0001, Dan Shi 0001, Renren Jin, Deyi Xiong
EMNLP2
2025 DiplomacyAgent: Do LLMs Balance Interests and Ethical Principles in International Events?
abstract
The widespread deployment of large language models (LLMs) across various domains has made their safety a critical priority.Inspired by think-tank decision-making philosophy, we propose DiplomacyAgent, an LLM-based multiagent system for diplomatic position analysis.With DiplomacyAgent, we are able to systematically assess how LLMs balance "interests" against "ethical principles" when addressing various international events, hence understanding the safety implications of LLMs in diplomacy.Specifically, this will help to assess the consistency of LLM stance with widely recognized ethical standards, as well as the potential risks or ideological biases that may arise.Through integrated quantitative metrics, our research uncovers unexpected decision-making patterns in LLM responses to sensitive issues including human rights protection, environmental sustainability, regional conflicts, etc.It discloses that LLMs could exhibit a strong bias towards interests, leading to unsafe decisions that violate ethical and moral principles.Our experiment results suggest that deploying LLMs in high-stakes domains, particularly in the formulation of diplomatic policies, necessitates a comprehensive assessment of potential ethical and social implications, as well as the implementation of stringent safety protocols.
Jianxiang Peng, Ling Shi 0004, Xinwei Wu 0001, Fujiang Liu, Haocheng Lyu, Deyi Xiong
EMNLP3
2024 IRCAN: Mitigating Knowledge Conflicts in LLM Generation via Identifying and Reweighting Context-Aware Neurons
abstract
It is widely acknowledged that large language models (LLMs) encode a vast reservoir of knowledge after being trained on mass data. Recent studies disclose knowledge conflicts in LLM generation, wherein outdated or incorrect parametric knowledge (i.e., encoded knowledge) contradicts new knowledge provided in the context. To mitigate such knowledge conflicts, we propose a novel framework, IRCAN (Identifying and Reweighting Context-Aware Neurons) to capitalize on neurons that are crucial in processing contextual cues. Specifically, IRCAN first identifies neurons that significantly contribute to context processing, utilizing a context-aware attribution score derived from integrated gradients. Subsequently, the identified context-aware neurons are strengthened via reweighting. In doing so, we steer LLMs to generate context-sensitive outputs with respect to the new knowledge provided in the context. Extensive experiments conducted across a variety of models and tasks demonstrate that IRCAN not only achieves remarkable improvements in handling knowledge conflicts but also offers a scalable, plug-and-play solution that can be integrated seamlessly with existing models. Our codes are released at https://github.com/danshi777/IRCAN.
Dan Shi 0001, Renren Jin, Tianhao Shen, Weilong Dong, Xinwei Wu 0001, Deyi Xiong
NeurIPS5
2023 DEPN: Detecting and Editing Privacy Neurons in Pretrained Language Models
abstract
Large language models pretrained on a huge amount of data capture rich knowledge and information in the training data.The ability of data memorization and regurgitation in pretrained language models, revealed in previous studies, brings the risk of data leakage.In order to effectively reduce these risks, we propose a framework DEPN to Detect and Edit Privacy Neurons in pretrained language models, partially inspired by knowledge neurons and model editing.In DEPN, we introduce a novel method, termed as privacy neuron detector, to locate neurons associated with private information, and then edit these detected privacy neurons by setting their activations to zero.Furthermore, we propose a privacy neuron aggregator dememorize private information in a batch processing manner.Experimental results show that our method can significantly and efficiently reduce the exposure of private data leakage without deteriorating the performance of the model.Additionally, we empirically demonstrate the relationship between model memorization and privacy neurons, from multiple perspectives, including model size, training time, prompts, privacy neuron distribution, illustrating the robustness of our approach.
Xinwei Wu 0001, Junzhuo Li, Weilong Dong, Shuangzhi Wu, Chao Bian 0006, Deyi Xiong
EMNLP1
2021 Unbiased Learning to Rank in Feeds Recommendation
abstract
In feeds recommendation, users are able to constantly browse items generated by never-ending feeds using mobile phones. The implicit feedback from users is an important resource for learning to rank, however, building ranking functions from such observed data is recognized to be biased. The presentation of the items will influence the user's judgements and therefore introduces biases. Most previous works in the unbiased learning to rank literature focus on position bias (i.e., an item ranked higher has more chances of being examined and interacted with). By analyzing user behaviors in product feeds recommendation, in this paper, we identify and introduce context bias, which refers to the probability that a user interacting with an item is biased by its surroundings, to unbiased learning to rank. We propose an Unbiased Learning to Rank with Combinational Propensity (ULTR-CP) framework to remove the inherent biases jointly caused by multiple factors. Under this framework, a context-aware position bias model is instantiated to estimate the unified bias considering both position and context biases. In addition to evaluating propensity score estimation approaches by the ranking metrics, we also discuss the evaluation of the propensities directly by checking their balancing properties. Extensive experiments performed on a real e-commerce data set collected from JD.com verify the effectiveness of context bias and illustrate the superiority of ULTR-CP against the state-of-the-art methods.
Xinwei Wu 0001, Hechang Chen, Jiashu Zhao, Dawei Yin 0001, Yi Chang 0001
WSDM1