Xinzhe Huang

dblp:335/4436 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2026
0009-0000-5719-6006ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DUALBREACH: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
Xinzhe Huang, Kedong Xiu, Tianhang Zheng, Churui Zeng, Wangze Ni, Zhan Qin, Kui Ren 0001, Chun Chen 0001
NDSS1
2025 CVshield: Interpretable Black-Box Adversarial Defense for LLMs via CoT Guided Semantic Verification
abstract
Despite their impressive capabilities, large language models (LLMs) remain highly susceptible to adversarial perturbations, often producing misleading or inconsistent outputs, thus undermining their reliability. In this paper, we propose CVshield, a two-stage training-free interpretable black-box defense framework for LLMs against adversarial perturbations. Under the framework of CVshield, we first design a set of Chain-Of-Thought (CoT) templates tailored to different types of adversarial perturbations, enabling the LLM to analyze and reconstruct potentially corrupted inputs from multiple semantic perspectives. We then propose a semantic alignment verification (SAV) module to evaluate semantic consistency across the LLM outputs generated by the CoT templates and the original LLM output. Significant discrepancies indicate the likely presence of a specific type of adversarial perturbation corresponding to the applied CoT template. To precisely identify the type of disturbance, we further introduce a decision mechanism based on Bayes’ theorem, which integrates evidence from the SAV comparison process to make a robust and probabilistically sound detection decision.We evaluated CVshield on LLaMA2-7B-Chat, Qwen-1.5B, and Mistral-7B models. CVshield successfully reduces the average attack success rate (ASR) to 24.67% in three datasets (MIX, SQuAD, and MedQuAD), significantly outperforming baseline methods such as perplexity (58.80%), tokenization (47.07%) and SmoothLLM (58.76%). Meanwhile, for the identification of attack type, CVshield achieved an average precision of 78.16% on the three datasets, substantially exceeding Perplexity (49.66%), Retokenization (30.65%) and SmoothLLM (38.34%).
Wenjing Hu, Weiwei Qi 0001, Yanlu Li, Xinzhe Huang
TrustCom4
2023 EduChain: A Blockchain-Based Privacy-Preserving Lifelong Education Platform
Xinzhe Huang, Hai Liang, Yong Ding 0005, Qianhong Wu
DASFAA (4)1
2022 A Privacy-Preserving Credit Bank Supervision Framework Based on Redactable Blockchain
Xinzhe Huang, Yong Ding 0005, Haibin Zheng, Decun Luo, Junfu Wu, Luyi Zhang
BlockSys1