Wenxuan Zhou 0005

dblp:78/9975-5 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Language models and text generation · 64% Deep learning architectures and training · 20% Reinforcement learning · 16%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
preference optimization
1.622025
T-REG: Preference Optimization with Token-Level Reward Regularization · ACL (1) 2025
WPO: Enhancing RLHF with Weighted Preference Optimization · EMNLP 2024
Machine learning › Reinforcement learning
reinforcement learning from human feedback
1.622025
T-REG: Preference Optimization with Token-Level Reward Regularization · ACL (1) 2025
WPO: Enhancing RLHF with Weighted Preference Optimization · EMNLP 2024
Natural language and speech › Language models and text generation
instruction tuning
1.012026
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe · ACL (1) 2026
Machine learning › Deep learning architectures and training › regularization
mixup regularization
1.012026
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe · ACL (1) 2026
Machine learning › Deep learning architectures and training
training dynamics
1.012026
SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe · ACL (1) 2026
Natural language and speech › Language models and text generation
alignment
0.912025
T-REG: Preference Optimization with Token-Level Reward Regularization · ACL (1) 2025
Natural language and speech › Language models and text generation › instruction following
instruction hierarchy
0.912025
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy · ICLR 2025
Natural language and speech › Language models and text generation
large language model safety
0.912025
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy · ICLR 2025
Natural language and speech › Language models and text generation › large language model training
token-level credit assignment
0.912025
T-REG: Preference Optimization with Token-Level Reward Regularization · ACL (1) 2025
Security and privacy of machine learning › adversarial defense
prompt injection defense
0.912025
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy · ICLR 2025
Natural language and speech › Language models and text generation
instruction following
0.312025
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy · ICLR 2025

Methods — techniques the papers use, named apart from their topics

supervised fine-tuning · 2.0mixup · 2.0instructional segment embedding · 1.7reward regularization · 0.9contrastive prompting · 0.9weighted preference optimization · 0.8reinforcement learning · 0.8
YearPublicationVenuePosition
2026 SFTMix: Elevating Language Model Instruction Tuning with Mixup Recipe
abstract
To acquire instruction-following capabilities, large language models (LLMs) undergo instruction tuning, where they are trained on instruction-response pairs using next-token prediction (NTP). Efforts to improve instruction tuning often focus on higher-quality supervised fine-tuning (SFT) datasets, typically requiring data filtering with proprietary LLMs or human annotation. In this paper, we take a different approach by proposing SFTMix, a novel Mixup-based recipe that elevates LLM instruction tuning without relying on well-curated datasets. We observe that LLMs exhibit uneven confidence across the semantic representation space. We argue that examples with different confidence levels should play distinct roles in instruction tuning: Confident data is prone to overfitting, while unconfident data is harder to generalize. Based on this insight, SFTMix leverages training dynamics to identify examples with varying confidence levels. We then interpolate them to bridge the confidence gap and apply a Mixup-based regularization to support learning on these additional, interpolated examples. We demonstrate the effectiveness of SFTMix in both instruction-following and healthcare-specific SFT tasks, with consistent improvements across LLM families and SFT datasets of varying sizes and qualities. Extensive analyses across six directions highlight SFTMix's compatibility with data selection, adaptability to compute-constrained scenarios, and scalability to broader applications.
Yuxin Xiao, Shujian Zhang, Marzyeh Ghassemi, Wenxuan Zhou 0005
ACL (1)4
2025 T-REG: Preference Optimization with Token-Level Reward Regularization
abstract
Reinforcement learning from human feedback (RLHF) has been crucial in aligning large language models (LLMs) with human values.Traditionally, RLHF involves generating responses to a query and using a reward model to assign a reward to the entire response.However, this approach faces challenges due to its reliance on a single, sparse reward, which makes it challenging for the model to identify which parts of the sequence contribute most significantly to the final reward.Recent methods have attempted to address this limitation by introducing token-level rewards.However, these methods often rely on either a trained credit assignment model or AI annotators, raising concerns about the quality and reliability of the rewards.In this paper, we propose token-level reward regularization (T-REG), a novel approach that leverages both sequence-level and token-level rewards for preference optimization.Harnessing the self-refinement capabilities of LLMs, our method uses contrastive prompting to enable LLMs to self-generate token-level rewards.These self-generated rewards then act as reward regularization, guiding the model to more effectively distribute sequence-level rewards across tokens.This facilitates better token-level credit assignment and enhances alignment performance.Experiments on the instruction following benchmarks, including Alpaca Eval 2 and Arena-Hard, show that our method consistently outperforms baseline methods by up to 3.8% and 4.4%, respectively.
Wenxuan Zhou 0005, Shujian Zhang
ACL (1)1
2025 Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
abstract
Large Language Models (LLMs) are susceptible to security and safety threats, such as prompt injection, prompt extraction, and harmful requests. One major cause of these vulnerabilities is the lack of an instruction hierarchy. Modern LLM architectures treat all inputs equally, failing to distinguish between and prioritize various types of instructions, such as system messages, user prompts, and data. As a result, lower-priority user prompts may override more critical system instructions, including safety protocols. Existing approaches to achieving instruction hierarchy, such as delimiters and instruction-based training, do not address this issue at the architectural level. We introduce the $\textbf{I}$nstructional $\textbf{S}$egment $\textbf{E}$mbedding (ISE) technique, inspired by BERT, to modern large language models, which embeds instruction priority information directly into the model. This approach enables models to explicitly differentiate and prioritize various instruction types, significantly improving safety against malicious prompts that attempt to override priority rules. Our experiments on the Structured Query and Instruction Hierarchy benchmarks demonstrate an average robust accuracy increase of up to 15.75\% and 18.68\%, respectively. Furthermore, we observe an improvement in the instruction-following capability of up to 4.1\% on AlpacaEval. Overall, our approach offers a promising direction for enhancing the safety and effectiveness of LLM architectures.
Shujian Zhang, Kaiqiang Song, Silei Xu, Sanqiang Zhao, Ravi Agrawal, Sathish Reddy Indurthi, Chong Xiang 0001, Prateek Mittal, Wenxuan Zhou 0005
ICLR10
2024 WPO: Enhancing RLHF with Weighted Preference Optimization
abstract
Wenxuan Zhou, Ravi Agrawal, Shujian Zhang, Sathish Reddy Indurthi, Sanqiang Zhao, Kaiqiang Song, Silei Xu, Chenguang Zhu. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Wenxuan Zhou 0005, Ravi Agrawal, Shujian Zhang, Sathish Reddy Indurthi, Sanqiang Zhao, Kaiqiang Song, Silei Xu
EMNLP1