VLDB 2026 Research / reviewers in the wild / expert
Pengwei Zhan
dblp:284/1181
· DBLP profile ↗
13ranked-venue papers
6as first author
13since 2021 · last 2026
0000-0003-3724-4431ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BKPIR: Keyword PIR for Private Boolean Retrieval
Zhen Xu 0009, Yan Zhang 0014, Pengwei Zhan, Shuai Ma 0001, Ru Xie |
NDSS | 4 |
| 2025 | Uncovering the App Cloud Access Risks under Recommended IAM Security PracticesabstractThe rapid development of mobile applications and cloud computing has led to the widespread adoption of cloud service platforms for mobile backend services. However, improper use of cloud credentials has frequently resulted in the leakage of application data on cloud servers. Despite security recommendations from cloud service providers, vulnerabilities persist. To assess the effectiveness of these measures, we propose a detection system to identify cloud credential leaks in mobile applications, including hard-coded credentials and those stored on servers. We analyzed 21,724 applications from Google Play and one Chinese market, revealing new attacks triggered by stolen cloud credentials. Our findings indicate that even temporary credentials recommended by cloud providers may pose security risks. We identified 893 applications using cloud credentials from the three major providers, with 945 credentials found. By analyzing these credentials, we uncovered severe vulnerabilities in 356 apps, such as personally identifiable information (PII) leakage, credential forgery, and remote code execution (RCE). These issues threaten user privacy and app security. We also evaluated developer adherence to recommended IAM best practices and provided suggestions for improving cloud credential security, highlighting issues such as improper permissions, insufficient protection, outdated versions, and regional variants. Hengtong Lu, Pengwei Zhan |
Proc. Priv. Enhancing Technol. | 4 |
| 2024 | Rethinking Word-level Adversarial Attack: The Trade-off between Efficiency, Effectiveness, and ImperceptibilityabstractNeural language models have demonstrated impressive performance in various tasks but remain vulnerable to word-level adversarial attacks. Word-level adversarial attacks can be formulated as a combinatorial optimization problem, and thus, an attack method can be decomposed into search space and search method. Despite the significance of these two components, previous works inadequately distinguish them, which may lead to unfair comparisons and insufficient evaluations. In this paper, to address the inappropriate practices in previous works, we perform thorough ablation studies on the search space, illustrating the substantial influence of search space on attack efficiency, effectiveness, and imperceptibility. Based on the ablation study, we propose two standardized search spaces: the Search Space for ImPerceptibility (SSIP) and Search Space for EffecTiveness (SSET). The reevaluation of eight previous attack methods demonstrates the success of SSIP and SSET in achieving better trade-offs between efficiency, effectiveness, and imperceptibility in different scenarios, offering fair and comprehensive evaluations of previous attack methods and providing potential guidance for future works. Pengwei Zhan, Liming Wang 0001 |
LREC/COLING | 1 |
| 2024 | Unveiling the Lexical Sensitivity of LLMs: Combinatorial Optimization for Prompt EnhancementabstractLarge language models (LLMs) demonstrate exceptional instruct-following ability to complete various downstream tasks.Although this impressive ability makes LLMs flexible task solvers, their performance in solving tasks also heavily relies on instructions.In this paper, we reveal that LLMs are over-sensitive to lexical variations in task instructions, even when the variations are imperceptible to humans.By providing models with neighborhood instructions, which are closely situated in the latent representation space and differ by only one semantically similar word, the performance on downstream tasks can be vastly different.Following this property, we propose a blackbox Combinatorial Optimization framework for Prompt Lexical Enhancement (COPLE).COPLE performs iterative lexical optimization according to the feedback from a batch of proxy tasks, using a search strategy related to word influence.Experiments show that even widelyused human-crafted prompts for current benchmarks suffer from the lexical sensitivity of models, and COPLE recovers the declined model ability in both instruct-following and solving downstream tasks. Pengwei Zhan, Zhen Xu 0009, Ru Xie |
EMNLP | 1 |
| 2024 | Rumor Detection with News Environment Enhanced Propagation Structure
Yanqiang Zhang, Zhen Xu 0009, Yan Zhang 0014, Pengwei Zhan |
ICIC (13) | 6 |
| 2023 | Contrastive Learning with Adversarial Examples for Alleviating Pathology of Language ModelabstractNeural language models have achieved superior performance.However, these models also suffer from the pathology of overconfidence in the out-of-distribution examples, potentially making the model difficult to interpret and making the interpretation methods fail to provide faithful attributions.In this paper, we explain the model pathology from the view of sentence representation and argue that the counter-intuitive bias degree and direction of the out-of-distribution examples' representation cause the pathology.We propose a Contrastive learning regularization method using Adversarial examples for Alleviating the Pathology (ConAAP), which calibrates the sentence representation of out-of-distribution examples.ConAAP generates positive and negative examples following the attribution results and utilizes adversarial examples to introduce direction information in regularization.Experiments show that ConAAP effectively alleviates the model pathology while slightly impacting the generalization ability on in-distribution examples and thus helps interpretation methods obtain more faithful results. Pengwei Zhan, Jing Yang 0032, Chunlei Jing, Jingying Li, Liming Wang 0001 |
ACL (1) | 1 |
| 2023 | Improving the Quality of Textual Adversarial Examples with Dynamic N-gram Based AttackabstractNatural language models have been widely used for their impressive performance in various tasks, while their poor robustness also puts critical applications at high risk. These models are vulnerable to adversarial examples, which contain imperceptible noise that leads the model to wrong predictions. To ensure such malicious examples are imperceptible to humans, various word-level attack methods have been proposed. Previous works on word-level attacks attempt to generate adversarial examples by substituting words in sentences. They utilize different candidate substitution selection methods and substitution strategies to improve attack effectiveness and the quality of generated examples. However, previous works are all unigram-based attack methods, which ignore the connection between words. The unigram nature of these methods downgrades fluency, increases grammatical errors, and biases the semantics of adversarial examples, making adversarial examples easier to be detected by humans. In this paper, to improve the quality of textual adversarial examples and makes the adversarial example more imperceptible to human, we propose a black-box word-level attack method called Dynamic N-Gram Based Attack (DyGram). DyGram tokenizes the entire sentence into multiple n-gram units, rather than individual words as in previous works, and substitutes words in a sentence in descending order of n-gram unit importance. Extensive experiments demonstrate that DyGram achieves higher attack success rates than previous attack methods and improves the quality of generated adversarial examples in terms of the number of perturbed words, perplexity, grammatical correctness, and semantic similarity. Xiaojiao Xie, Pengwei Zhan |
CSCWD | 2 |
| 2023 | Unsupervised Clustering with Contrastive Learning for Rumor Tracking on Social Media
Zhitong Lu, Chunlei Jing, Pengwei Zhan, Zhen Xu 0009, Liming Wang 0001 |
NLPCC (2) | 5 |
| 2022 | PARSE: An Efficient Search Method for Black-box Adversarial Text AttacksabstractNeural networks are vulnerable to adversarial examples. The adversary can successfully attack a model even without knowing model architecture and parameters, i.e., under a black-box scenario. Previous works on word-level attacks widely use word importance ranking (WIR) methods and complex search methods, including greedy search and heuristic algorithms, to find optimal substitutions. However, these methods fail to balance the attack success rate and the cost of attacks, such as the number of queries to the model and the time consumption. In this paper, We propose PAthological woRd Saliency sEarch (PARSE) that performs the search under dynamic search space following the subarea importance. Experiments show that PARSE can achieve comparable attack success rates to complex search methods while saving numerous queries and time, e.g., saving at most 74% of queries and 90% of time compared with greedy search when attacking the examples from Yelp dataset. The adversarial examples crafted by PARSE are also of high quality, highly transferable, and can effectively improve model robustness in adversarial training. Pengwei Zhan, Jing Yang 0032, Yuxiang Wang 0005, Liming Wang 0001 |
COLING | 1 |
| 2022 | SITD: Insider Threat Detection Using Siamese Architecture on Imbalanced DataabstractIn the insider threat detection domain, data imbalance is a well-known problem. Most existing solutions, including rebalancing datasets and anomaly detection, have problems such as model overfitting, high cost, and high False Positive Rate (FPR). Therefore, how to effectively detect insider threats on an imbalanced dataset is a challenge. This paper proposes a new Siamese-architecture Insider Threat Detection (SITD) method, which detects insider threat by judging whether the input sample pairs belong to the same category instead of directly classifying a sample while avoiding the abovementioned problems. In addition, we improve the contrastive loss function to make the model pay more attention to the samples pairs of different categories, which significantly enhances the detection performance. Experimental results show that SITD outperforms other insider detection methods on the imbalanced CERT dataset. Moreover, SITD can achieve a good result no matter how imbalanced the dataset is. Shaolei Zhou, Liming Wang 0001, Jing Yang 0032, Pengwei Zhan |
CSCWD | 4 |
| 2022 | SP Attack: Single-Perspective Attack for Generating Adversarial Omnidirectional ImagesabstractThe safety of Deep Neural Networks (DNNs) processing omnidirectional images (ODIs) is an under-researched topic. In this paper, we propose a novel sparse attack, named Single-Perspective (SP) Attack, towards fooling these models by perturbing only one perspective image (PI) rendered from the target ODI. The attack is launched from the perspective domain, and finally the perturbation is transferred to the original ODI. To this end, we propose an effective PI position searching algorithm based on Bayesian Optimization, and then corrupt the PI centered on the desirable position with unconstrained/constrained perturbations. Extensive experiments on synthetic and real-world omnidirectional datasets demonstrate that SP Attack can overcome the projection deformation of ODIs, and mislead the neural networks by limiting the perturbations in a single patch on the target ODI. Yanwei Liu 0001, Jinxia Liu, Pengwei Zhan, Liming Wang 0001, Zhen Xu 0009 |
ICASSP | 4 |
| 2022 | Crafting Textual Adversarial Examples through Second-Order Enhanced Word SaliencyabstractTextual adversarial examples crafted with well-designed perturbation can mislead state-of-the-art natural language models. Most previous works on word-level black-box attacks propose different text substituting strategies based on the word saliency determined by Leave One Out (LOO) methods, while the attack effectiveness is actually limited due to the model pathology. The word saliency determined by LOO methods can be severely affected by the model pathology, and unconscious bias is introduced. In this paper, we propose a word saliency method called Second-Order Enhanced Word Saliency (SOEWS), which considers the overfitting information obtained from the second-order model pathological behavior that helps to reduce the bias of LOO methods. Extensive experiments show that our method outperforms current black-box word saliency methods and can boost the effectiveness of previous attack frameworks. Performing adversarial training utilizing our method best improves model robustness, and human evaluation shows that the adversarial examples crafted by our method are almost imperceptible to humans. We also conduct case study to intuitively demonstrate the superiority of our method. Pengwei Zhan, Liming Wang 0001, Jing Yang 0032 |
IJCNN | 1 |
| 2021 | Website fingerprinting on early QUIC traffic
Pengwei Zhan, Liming Wang 0001, Yi Tang 0001 |
Comput. Networks | 1 |