Zesheng Shi

dblp:341/1527 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 50% Trustworthy machine learning · 33% Representation and self-supervised learning · 17%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 6 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model safety
1.922026
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward · ACL (1) 2026
Safety Alignment via Constrained Knowledge Unlearning · ACL (1) 2025
Natural language and speech › Language models and text generation
alignment
1.012026
Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs · ACL (1) 2026
Machine learning › Trustworthy machine learning › adversarial machine learning
jailbreak attack
1.012026
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward · ACL (1) 2026
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning
self-supervised alignment
1.012026
Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs · ACL (1) 2026
Security and privacy of machine learning › adversarial attack
backdoor attack
1.012026
Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward · ACL (1) 2026
Machine learning › Trustworthy machine learning › AI safety
safety alignment
0.912025
Safety Alignment via Constrained Knowledge Unlearning · ACL (1) 2025

Methods — techniques the papers use, named apart from their topics

data poisoning · 2.0asymmetric chain backdoor · 2.0self-play · 1.0fine-tuning · 1.0adaptive weighting · 1.0knowledge unlearning · 0.9constrained optimization · 0.9
YearPublicationVenuePosition
2026 Backdoors in RLVR: Jailbreak Backdoors in LLMs From Verifiable Reward
abstract
Reinforcement Learning with Verifiable Rewards (RLVR) is an emerging paradigm that significantly boosts a Large Language Model's (LLM's) reasoning abilities on complex logical tasks, such as mathematics and programming.However, we identify, for the first time, a latent vulnerability to backdoor attacks within the RLVR framework.This attack can implant a backdoor without modifying the reward verifier by injecting a small amount of poisoning data into the training set.Specifically, we propose a novel trigger mechanism designated as the ASYMMETRIC CHAIN BACKDOOR (ACB).The attack exploits the RLVR training loop by assigning substantial positive rewards for harmful responses and negative rewards for refusals.This asymmetric reward signal forces the model to progressively increase the probability of generating harmful responses during training.Our findings demonstrate that the RLVR backdoor attack is characterized by both high efficiency and strong generalization capabilities.Utilizing less than 2% poisoned data in train set, the backdoor can be successfully implanted across various model scales without degrading performance on benign tasks.Evaluations across multiple jailbreak benchmarks indicate that activating the trigger degrades safety performance by an average of 73%.Furthermore, the attack generalizes effectively to a wide range of jailbreak methods and unsafe behaviors.
Weiyang Guo, Zesheng Shi, Zeen Zhu
ACL (1)2
2026 Team-Based Self-Play With Dual Adaptive Weighting for Fine-Tuning LLMs
abstract
While recent self-training approaches have reduced reliance on human-labeled data for aligning LLMs, they still face critical limitations: (i) sensitivity to synthetic data quality, leading to instability and bias amplification in iterative training; (ii) ineffective optimization due to a diminishing gap between positive and negative responses over successive training iterations.In this paper, we propose Team-based self-Play with dual Adaptive Weighting (TPAW), a novel self-play algorithm designed to improve alignment in a fully self-supervised setting.TPAW adopts a team-based framework in which the current policy model both collaborates with and competes against historical checkpoints, promoting more stable and efficient optimization.To further enhance learning, we design two adaptive weighting mechanisms: (i) a response reweighting scheme that adjusts the importance of target responses, and (ii) a player weighting strategy that dynamically modulates each team member's contribution during training.Initialized from a SFT model, TPAW iteratively refines alignment without requiring additional human supervision.Experimental results demonstrate that TPAW consistently outperforms existing baselines across various base models and LLM benchmarks.
Yigeng Zhou, Zesheng Shi, Yequan Wang
ACL (1)3
2025 Safety Alignment via Constrained Knowledge Unlearning
abstract
Zesheng Shi, Yucheng Zhou, Jing Li, Yuxin Jin, Yu Li, Daojing He, Fangming Liu, Saleh Alharbi, Jun Yu, Min Zhang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zesheng Shi, Yucheng Zhou 0001, Jing Li 0034, Yu Li 0007, Daojing He, Fangming Liu, Saleh Alharbi, Jun Yu 0002, Min Zhang 0005
ACL (1)1
2025 MedCoAct: Confidence-Aware Multi-Agent Collaboration for Complete Clinical Decision
abstract
Autonomous agents utilizing Large Language Models (LLMs) have demonstrated remarkable capabilities in isolated medical tasks like diagnosis and image analysis, but struggle with integrated clinical workflows that connect diagnostic reasoning and medication decisions. We identify a core limitation: existing medical AI systems process tasks in isolation without the cross-validation and knowledge integration found in clinical teams, reducing their effectiveness in real-world healthcare scenarios. To transform the isolation paradigm into a collaborative approach, we propose MedCoAct, a confidence-aware multi-agent framework that simulates clinical collaboration by integrating specialized doctor and pharmacist agents, and present a benchmark, DrugCareQA, to evaluate medical AI capabilities in integrated diagnosis and treatment workflows. Our results demonstrate that MedCoAct achieves 67.58% diagnostic accuracy and 67.58% medication recommendation accuracy, outperforming single agent framework by 7.04% and 7.08% respectively. This collaborative approach generalizes well across diverse medical domains, proving especially effective for telemedicine consultations and routine clinical scenarios, while providing interpretable decision-making pathways.
Hongjie Zheng, Zesheng Shi, Ping Yi
BIBM2
2025 Consistent focus: Mitigating permutation bias in large language models through attention weight averaging
Kanglong Li, Zesheng Shi, Meng-Jun Hu, Yu-Xin Jin, Xiaohu Yan, Fazhi He, Xingyao Wu
Expert Syst. Appl.2
2023 Topic-Selective Graph Network for Topic-Focused Summarization
Zesheng Shi, Yucheng Zhou 0001
PAKDD (4)1