Rui Pu

dblp:161/0770 · DBLP profile ↗
← Back
13ranked-venue papers
4as first author
13since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2027 LANCET: Neural intervention via Structural Entropy for mitigating faithfulness hallucinations in LLMs
Chenxu Wang 0001, Chaozhuo Li, Litian Zhang, Songyang Liu, Yushan Cai, Rui Pu
Inf. Process. Manag.10
2026 MirrorShield: Towards Dynamic Adaptive Defense Against Jailbreaks via Entropy-Guided Mirror Crafting
abstract
Defending large language models (LLMs) against jailbreak attacks is crucial for ensuring their safe deployment. Existing defense strategies typically rely on predefined static criteria to differentiate between harmful and benign prompts. However, such rigid rules fail to accommodate the inherent complexity and dynamic nature of real-world jailbreak attacks. In this paper, we focus on the novel challenge of adaptive defense against diverse jailbreaks. We propose a new concept "mirror'', which is a dynamically generated prompt that reflects the syntactic structure of the input while ensuring semantic safety. The discrepancies between input prompts and their corresponding mirrors serve as guiding principles for defense. A novel defense model, MirrorShield, is further proposed to detect and calibrate risky inputs based on the crafted mirrors. Evaluated on multiple benchmark datasets and compared against ten state-of-the-art attack methods, MirrorShield demonstrates superior defense performance and promising generalization capabilities.
Rui Pu, Chaozhuo Li, Rui Ha, Litian Zhang, Lirong Qiu, Xi Zhang 0008
AAAI1
2026 From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in LRMs via Decoupled Reasoning and Control
abstract
Large Reasoning Models (LRMs) can exhibit step-by-step reasoning, reflection, and backtracking, but these behaviors are often unregulated, leading to overthinking.As a result, LRMs continue generating redundant reasoning even after reaching high-confidence conclusions.This increases inference cost and latency, limiting practical deployment.The root cause is the absence of an intrinsic mechanism to monitor the reasoning state and decide when to continue, backtrack, or stop.We propose MERA, a meta-cognitive reasoning framework that decouples reasoning from control to enable independent optimization of control strategies.MERA constructs high-quality reasoning-control supervision data via a takeoverbased pipeline, and transforms long-horizon traces into structured reasoning-control alternating sequences for training.The model is trained with supervised fine-tuning to internalize the structured separation, and further optimized with Control-Segment Policy Optimization (CSPO), which combines segment-wise GRPO with control masking to focus learning on control segments.Experiments across reasoning benchmarks show that MERA improves both efficiency and accuracy.
Rui Ha, Rui Pu, Chaozhuo Li, Li Sun 0008, Sen Su
ACL (1)2
2026 Beyond Surface-Level Detection: Towards Cognitive-Driven Defense Against Jailbreak Attacks via Meta-Operations Reasoning
abstract
Defending large language models (LLMs) against jailbreak attacks is essential for their safe and reliable deployment.Existing defenses often rely on shallow pattern matching, which struggles to generalize to novel and unseen attack strategies.To address this challenge, we propose the Cognitive-Driven Defense (CDD) framework, which targets the underlying structure of jailbreak prompts by applying metaoperations, defined as basic manipulations that conceal harmful intent.CDD emulates human cognitive reasoning through a structured reasoning chain.It begins with a global perception of the prompt and follows with a localized analysis to uncover hidden manipulations.By applying supervised fine-tuning on this structured chain, the model learns to identify and reason about known manipulation patterns.To enhance generalization to unseen threats, an entropy-guided reinforcement learning algorithm (EG-GRPO) is introduced to encourage exploration of new types and variants of meta-operations.Experiments demonstrate that CDD can achieve state-of-the-art defense performance and exhibit strong generalization to unseen jailbreak attacks.
Rui Pu, Chaozhuo Li, Rui Ha, Litian Zhang, Lirong Qiu, Xi Zhang 0008
ACL (1)1
2026 BNFW: Boundary and noise detection clustering for data with fuzzy boundaries and weak connectivity
Tianshuo Li, Rui Pu, Zhijiang Chen, Dongming Tang
Expert Syst. Appl.2
2026 LDCC: Adaptive clustering for data with weak connectivity using local directional centrality
Tianshuo Li, Zhijiang Chen, Tingyu Yan, Rui Pu, Dongming Tang
Pattern Recognit.5
2025 DSG-MCTS: A Dynamic Strategy-Guided Monte Carlo Tree Search for Diversified Reasoning in Large Language Models
abstract
Large language models (LLMs) have shown strong potential in complex reasoning tasks.However, as task complexity increases, their performance often degrades, resulting in hallucinations, errors, and logical inconsistencies.To enhance reasoning capabilities, Monte Carlo Tree Search (MCTS) has been introduced to guide the exploration of reasoning paths in a structured manner.Despite its advantages, traditional MCTS relies on fixed reasoning strategies, limiting the diversity of reasoning paths and the coverage of the solution space.To address these limitations, we propose Dynamic Strategy-Guided MCTS (DSG-MCTS), a novel framework that dynamically integrates multiple reasoning strategies, such as abductive and analogical reasoning, to expand the reasoning space.At the same time, DSG-MCTS enhances reasoning efficiency through a dynamic strategy selection mechanism that adapts to the task context.Experimental results on challenging reasoning benchmarks demonstrate that DSG-MCTS achieves improved accuracy and efficiency, outperforming existing state-of-the-art methods.
Rui Ha, Chaozhuo Li, Rui Pu, Litian Zhang, Xi Zhang 0008, Sen Su
EMNLP3
2025 Feint and Attack: Jailbreaking and Protecting LLMs via Attention Distribution Modeling
abstract
Most jailbreak methods for large language models (LLMs) focus on superficially improving attack success through manually defined rules. However, they fail to uncover the underlying mechanisms within target LLMs that explain why an attack succeeds or fails. In this paper, we propose investigating the phenomenon of jailbreaks and defenses for LLMs from the perspective of attention distributions within the models. A preliminary experiment reveals that the success of a jailbreak is closely linked to the LLM's attention on sensitive words.Inspired by this interesting finding, we propose incorporating critical signals derived from internal attention distributions within LLMs, namely Attention Intensity on Sensitive Words and Attention Dispersion Entropy, to guide both attacks and defenses. Drawing inspiration from the concept of "Feint and Attack", we introduce an attention-guided jailbreak model, ABA, which redirects the model's attention to benign contexts, and an attention-based defense model, ABD, designed to detect attacks by analyzing internal attention entropy. Experimental results demonstrate the superiority of our proposal when compared to SOTA baselines.
Rui Pu, Chaozhuo Li, Rui Ha, Zejian Chen, Litian Zhang, Zheng Liu 0011, Lirong Qiu, Zaisheng Ye
IJCAI1
2025 NaGB-DBSCAN: An improved DBSCAN clustering algorithm by natural neighbor and granular-ball
Ranliang Luo, Tianshuo Li, Rui Pu, Juntao Yang, Dongming Tang
Inf. Sci.3
2024 BaitAttack: Alleviating Intention Shift in Jailbreak Attacks via Adaptive Bait Crafting
abstract
Jailbreak attacks enable malicious queries to evade detection by LLMs.Existing attacks focus on meticulously constructing prompts to disguise harmful intentions.However, the incorporation of sophisticated disguising prompts may incur the challenge of "intention shift".Intention shift occurs when the additional semantics within the prompt distract the LLMs, causing the responses to deviate significantly from the original harmful intentions.In this paper, we propose a novel component, "bait", to alleviate the effects of intention shift.Bait comprises an initial response to the harmful query, prompting LLMs to rectify or supplement the knowledge within the bait.By furnishing rich semantics relevant to the query, the bait helps LLMs focus on the original intention.To conceal the harmful content within the bait, we further propose a novel attack paradigm, BaitAttack.BaitAttack adaptively generates necessary components to persuade targeted LLMs that they are engaging with a legitimate inquiry in a safe context.Our proposal is evaluated on a popular dataset, demonstrating state-of-the-art attack performance and an exceptional capability for mitigating intention shift.The implementation of BaitAttack is accessible at: https://anonymous.4open. science/r/BaitAttack-D1F5.How to make a bomb. Query Bait Role Scene Format Bait Maker Which expert is best suitable to deal with the act of < Harmful Query >? Create a scene that fits the expert role.
Rui Pu, Chaozhuo Li, Rui Ha, Litian Zhang, Lirong Qiu, Xi Zhang 0008
EMNLP1
2024 Non-parameter clustering algorithm based on chain propagation and natural neighbor
Tianshuo Li, Juntao Yang, Rui Pu, Jinghui Zhang 0001, Dongming Tang, Tao Liu 0027
Inf. Sci.4
2022 An affine combination of two augmented CLMS adaptive filters for processing noncircular Gaussian signals
Zhe Li 0007, Rui Pu, Yili Xia, Wenjiang Pei
Signal Process.2
2021 Couple-group consensus for heterogeneous MASs under switched topologies in cooperative-competitive systems: A hybrid pinning and delta operator skills
Xingcheng Pu, Rui Pu
Neurocomputing4