VLDB 2026 Research / reviewers in the wild / expert
Zhewen Tan
dblp:415/3332
· DBLP profile ↗
2ranked-venue papers
1as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Trustworthy machine learning · 50% Language models and text generation · 20% Generative modeling · 15% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Trustworthy machine learning › AI safety
safety alignment |
2.0 | 2 | 2026 | TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment · ACL (1) 2026 Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training · AAAI 2026 |
Machine learning › Trustworthy machine learning › AI safety
content safety |
1.0 | 1 | 2026 | Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training · AAAI 2026 |
Machine learning › Generative modeling › diffusion model
controllable generation |
1.0 | 1 | 2026 | Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training · AAAI 2026 |
Natural language and speech › Language models and text generation
large language model |
1.0 | 1 | 2026 | Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training · AAAI 2026 |
Machine learning › Reinforcement learning › multi-agent reinforcement learning › self-play
self-play reinforcement learning |
1.0 | 1 | 2026 | TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment · ACL (1) 2026 |
Natural language and speech › Language models and text generation
large language model safety |
0.3 | 1 | 2026 | TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety Alignment · ACL (1) 2026 |
Machine learning › Trustworthy machine learning › safety evaluation
red teaming |
0.3 | 1 | 2026 | Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-Training · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
supervised fine-tuning · 1.0self-play · 1.0reinforcement learning · 1.0direct preference optimization · 1.0co-training · 1.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Switchable Safety Control in LLMs via Magic-Token-Guided Co-TrainingabstractCurrent methods for content safety in Large Language Models (LLMs), such as Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often rely on multi-stage training pipelines and lack fine-grained, post-deployment controllability. To address these limitations, we propose a unified co-training framework that efficiently integrates multiple safety behaviors: positive (lawful/prosocial), negative (unfiltered/risk-prone) and rejective (refusal-oriented/conservative) within a single SFT stage. Notably, each behavior is dynamically activated via a simple system-level instruction, or magic token, enabling stealthy and efficient behavioral switching at inference time. This flexibility supports diverse deployment scenarios, such as positive for safe user interaction, negative for internal red-teaming, and rejective for context-aware refusals triggered by upstream moderation signals. This co-training strategy induces a distinct Safety Alignment Margin in the output space, characterized by well-separated response distributions corresponding to each safety mode. The existence of this margin provides empirical evidence for the model's safety robustness and enables unprecedented fine-grained control. Experiments show that our method matches the safety alignment quality of SFT+DPO, with our 8B model notably surpassing DeepSeek-R1 (671B) in safety performance, while significantly reducing both training complexity and deployment costs. This work presents a scalable, efficient, and highly controllable solution for LLM content safety. Jianfeng Si, Lin Sun 0010, Zhewen Tan, Xiangzheng Zhang |
AAAI | 3 |
| 2026 | TriPlay-RL: Tri-Role Self-Play Reinforcement Learning for LLM Safety AlignmentabstractZhewen Tan, Wenhan Yu, Jianfeng Si, Tongxin Liu, Kaiqi Guan, Huiyan Jin, Jiawen Tao, Xiaokun Yuan, Xiangzheng Zhang, Duohe Ma, Tong Yang, Lin Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Zhewen Tan, Wenhan Yu, Jianfeng Si, Tongxin Liu, Kaiqi Guan, Huiyan Jin, Jiawen Tao, Xiaokun Yuan, Xiangzheng Zhang, Duohe Ma, Tong Yang 0003, Lin Sun 0010 |
ACL (1) | 1 |