VLDB 2026 Research / reviewers in the wild / expert
Pulei Xiong
dblp:67/4410
· DBLP profile ↗
21ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0002-3460-6946ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Security and privacy · 7 · 3 first-author · 6 since 2021Computer networks · 5 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards integration of privacy enhancing technologies in explainable artificial intelligenceabstractExplainable artificial intelligence (XAI) plays a crucial role in mitigating the risks associated with the non-transparency of black-box artificial intelligence (AI) systems. However, despite its advantages, XAI methods have been shown to expose the privacy of individuals whose data are used to train or query the underlying models. Prior research has demonstrated privacy attacks that exploit explanations to infer sensitive personal information of individuals. At present, there is a lack of effective defenses against such privacy attacks targeting explanations, particularly when vulnerable XAI techniques are deployed in production environments or used in machine learning as a service systems. To address this gap, this study investigates the use of privacy enhancing technologies (PETs) as a defense mechanism against attribute inference attacks on explanations generated by feature-based XAI methods. We empirically evaluate three types of PETs, i.e., synthetic training data, differentially private training and noise addition, across two categories of feature-based XAI. Our findings reveal varying levels of effectiveness among the mitigation strategies, as well as trade-offs between privacy, utility and system performance. In the best scenario, integrating PETs into the explanation process reduced attack success by 49.47% while preserving model utility and explanation quality. Based on our evaluation, we propose strategies for effectively integrating PETs into XAI to maximize privacy protection and minimize the risk of sensitive information leakage. Sonal Allana, Rozita Dara 0001, Xiaodong Lin 0001, Pulei Xiong |
Knowl. Based Syst. | 4 |
| 2026 | A Systematic Review of Adversarial Attacks and Defenses for Deep Reinforcement Learning in Autonomous Vehicle ApplicationsabstractDeep reinforcement learning (DRL) is a significant component of autonomous vehicle (AV) systems, as it is involved in crucial aspects such as decision-making, motion planning, and control systems. Nevertheless, DRL exhibits inherent weaknesses that render it vulnerable to adversarial attacks. On the other hand, AV applications are classified as safety-critical applications. This work presents a systematic literature review of several forms of adversarial attacks and the corresponding response mechanisms developed in the literature on AVs. This study additionally examines several datasets utilized for benchmarking purposes and different evaluation factors employed to assess the resilience and efficacy of adversarial attacks. The survey adheres to the PRMA standards. Excluded from consideration are articles that failed to take into account adversarial assaults on DRL and Deep Learning (DL). A comprehensive search yielded a total of 2250 distinct peer-reviewed articles sourced from the ACM Digital Library, IEEE Xplore, Springer, and Elsevier databases. Upon fulfilling the predefined criteria for inclusion and exclusion, a sum of 151 articles is encompassed in the ultimate review. In the end, the paper examines the existing research gaps and proposes potential options for further investigations. Zahra Sarayloo, Saeedeh Lohrasbi, Pulei Xiong, Nasser L. Azad |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | NegRefine: Refining Negative Label-Based Zero-Shot OOD DetectionabstractRecent advancements in Vision-Language Models like CLIP have enabled zero-shot OOD detection by leveraging both image and textual label information. Among these, negative label-based methods such as NegLabel and CSP have shown promising results by utilizing a lexicon of words to define negative labels for distinguishing OOD samples. However, these methods suffer from detecting in-distribution samples as OOD due to negative labels that are subcategories of in-distribution labels or proper nouns. They also face limitations in handling images that match multiple in-distribution and negative labels. We propose NegRefine, a novel negative label refinement framework for zero-shot OOD detection. By introducing a filtering mechanism to exclude subcategory labels and proper nouns from the negative label set and incorporating a multi-matching-aware scoring function that dynamically adjusts the contributions of multiple labels matching an image, NegRefine ensures a more robust separation between in-distribution and OOD samples. We evaluate NegRefine on large-scale benchmarks, including ImageNet-1K. Source code is available at https://github.com/ah-ansari/NegRefine. Amirhossein Ansari, Ke Wang 0001, Pulei Xiong |
ICCV | 3 |
| 2025 | Accelerating Adversarial Training on Under-Utilized GPUabstractDeep neural networks are vulnerable to adversarial attacks and adversarial training has been proposed to defend against such attacks by adaptively generating attacks, i.e., adversarial examples, during training. However, adversarial training is significantly slower than traditional training due to the search for worst attacks for each minibatch. To speed up adversarial training, existing work has considered a subset of a minibatch for generating attacks and reduced the steps in the search for attacks. We propose a novel adversarial training acceleration method, called AttackRider, by exploring under-utilized GPU hardware to reduce the number of calls to attack generation without increasing the time of each call. We characterize the extent of under-utilization of GPU for given GPU and model size, hence the potential for speedup, and present the application scenarios where this opportunity exists. The results on various machine learning tasks and datasets show that AttackRider can speed up state-of-the-art adversarial training algorithms with comparable robust accuracy. The source code of AttackRider is available at https://github.com/zxzhan/AttackRider. Zhuoxin Zhan, Ke Wang 0001, Pulei Xiong |
IJCAI | 3 |
| 2025 | A Generic Framework for Privacy Risk Assessment of Machine Learning ModelsabstractPrivacy attacks on machine learning (ML) models pose significant risks to individuals whose personal data is used for training or querying these models. Although concerns about the potential exposure of sensitive information through ML models continue to grow, existing safeguard mechanisms primarily focus on security threats, often neglecting privacy risks. In this paper, we examine existing tools to assess privacy risks of ML models and provide an overview of various privacy attacks and defense strategies. Given the lack of a comprehensive framework for assessing privacy vulnerabilities, we propose a generic framework for evaluating the privacy of ML systems and establish a set of tailored evaluation metrics for different types of privacy attacks. In addition, we develop a dedicated testbed to implement our framework and present experimental results that demonstrate the impact of various privacy attacks on different ML models. Le Wang 0010, Sonal Allana, Xiaodong Lin 0001, Rozita Dara 0001, Pulei Xiong |
PST | 7 |
| 2025 | G-STAR: A Threat Modeling Framework for General-Purpose AI SystemsabstractThis research presents the preliminary findings of an ongoing project focused on the security of General-Purpose AI (GPAI) applications. We introduce three key contributions: (i) a taxonomy of GPAI-specific vulnerabilities, offering a structured classification of security risks unique to GPAI models and applications; (ii) a generalized GPAI application architecture, serving as a meta-model for analyzing a wide range of real-world use cases; and (iii) G-STAR, a novel threat modeling reference framework that identifies key entities and their interrelationships in GPAI ecosystems, and provides a structured methodology for assessing and mitigating potential threats. Our study addresses both data and model vulnerabilities inherent in GPAI systems, highlighting critical security challenges. While the research is still in its early stages, the initial results provide a valuable foundation for continued investigation. Future work will focus on enhancing the generalized architecture, exploring mitigation strategies in depth, and applying and refining the G-STAR framework in real-world GPAI scenarios. This work aims to support AI security practitioners in promoting secure development and deployment of GPAI systems across diverse domains. Pulei Xiong, Saeedeh Lohrasbi, Prini Kotian, Scott Buffett |
PST | 1 |
| 2025 | Enhancing Adversarial Robustness of IoT Intrusion Detection via SHAP-Based Attribution Fingerprinting
Dilli P. Sharma, Xiaodong Lin 0001, Pulei Xiong |
TrustCom | 5 |
| 2024 | Large Language Model Empowered Spatio-Visual Queries for Extended Reality EnvironmentsabstractWith the technological advances in creation and capture of 3D spatial data, new emerging applications are being developed. Digital Twins, metaverse and extended reality (XR) based immersive environments can be enriched by leveraging geocoded 3D spatial data. Unlike 2D spatial queries, queries involving 3D immersive environments need to take the query user’s viewpoint into account. Spatio-visual queries return objects that are visible from the user’s perspective.In this paper, we propose enhancing 3D spatio-visual queries with large language models (LLM). These kinds of queries allow a user to interact with the visible objects using a natural language interface. We have implemented a proof-of-concept prototype and conducted preliminary evaluation. Our results demonstrate the potential of truly interactive immersive environments. Mohammadmasoud Shabanijou, Vidit Sharma, Suprio Ray, Rongxing Lu, Pulei Xiong |
IEEE Big Data | 5 |
| 2024 | Out-of-Distribution Aware Classification for Tabular DataabstractOut-of-distribution (OOD) aware classification aims to classify in-distribution samples into their respective classes while simultaneously detecting OOD samples. Previous works have largely focused on the image domain, where images from an unrelated dataset can serve as auxiliary OOD training data. In this work, we address OOD-aware classification for tabular data, where an unrelated dataset cannot be used as OOD training data. A potential solution to OOD-aware classification involves filtering out OOD samples using an outlier detection method and classifying the remaining samples with a traditional classification model. However, seamlessly integrating this approach into downstream optimization tasks is challenging due to the employment of multiple methods. Our approach is turning OOD-aware classification into traditional classification by augmenting the in-distribution training data with synthesized OOD data. This approach continues leveraging traditional classification methods while detecting OOD samples, and the learned model retains the same mathematical properties as traditional classification models, thus, it can be easily integrated into downstream tasks. We evaluate these benefits empirically using real-life datasets. Code is available at https://github.com/ah-ansari/OCT. Amirhossein Ansari, Ke Wang 0001, Pulei Xiong |
CIKM | 3 |
| 2024 | Better entity matching with transformers through ensembles
Jwen Fai Low, Benjamin C. M. Fung, Pulei Xiong |
Knowl. Based Syst. | 3 |
| 2023 | Achieve Edge-Based Privacy-Preserving Dynamic Aggregation Query in Smart Transportation SystemsabstractAs the proliferation of smart vehicles has fostered an abundance of real-time data, various data analysis tools, such as aggregation queries, are expected to be deployed to extract insights and make transportation systems much smarter. Meanwhile, to cope with the growing service scale, edge servers are employed to collect data and deliver the service, which however provokes privacy concerns related to the reported data and user queries. Previously reported solutions on privacy-preserving aggregation queries focus on static datasets or require data persistence, leading to storage pressure and slower query responses. In this paper, we propose a privacy-preserving dynamic aggregation query scheme using edge servers, specifically addressing the problem of online aggregation queries. By combining homomorphic encryption and predicate encryption, our scheme enables the edge server to aggregate real-time data and respond to queries, safeguarding sensitive information from vehicles and data users. The integration of advanced cryptographic primitives ensures data and query privacy and integrity. Comprehensive theoretical analyses demonstrate our scheme's effectiveness in privacy preservation, boasting a manageable computational and communication overhead. The scheme, thus, presents a practical solution for privacy-preserving dynamic aggregation queries, fulfilling an unmet need in real-time transportation systems. Yunguo Guan, Ellen Z. Zhang, Pulei Xiong, Rongxing Lu |
GLOBECOM | 3 |
| 2023 | IoT malware: An attribute-based taxonomy, detection mechanisms and challenges
Princy Victor, Arash Habibi Lashkari, Rongxing Lu, Tinshu Sasi, Pulei Xiong, Shahrear Iqbal |
Peer Peer Netw. Appl. | 5 |
| 2022 | Privacy-Preserving Outsourced Task Scheduling in Mobile CrowdsourcingabstractWith the proliferation of smart devices with various onboard sensors, mobile crowdsourcing has attracted consider-able interest, in which task scheduling is an essential service for allocating tasks to workers. As the number of workers increases, the service provider prefers to outsource the service to a powerful cloud, and the messages from the workers and the task owners should be protected. Currently, the existing works on privacy-preserving task allocation cannot simultaneously achieve strong privacy protection and dynamic update. Aiming at this challenge, we propose a privacy-preserving outsourced task scheduling scheme, in which the cloud servers can obliviously conduct task scheduling and update the workers' information based on the scheduling results. To this end, we design three secure protocols under a two-server model, which can protect the protocols' input and output against the cloud servers. The security analysis demonstrates that our proposed scheme can achieve strong privacy protection, and the experimental results indicate the scheme's efficiency in computing and communication. Yunguo Guan, Pulei Xiong, Songnian Zhang, Rongxing Lu |
GLOBECOM | 2 |
| 2022 | EVRQ: Achieving Efficient and Verifiable Range Query over Encrypted Traffic DataabstractAs a promising solution in boosting the efficiency of urban traffic, intelligent transportation systems (ITS) have been increasingly deployed. Traffic message dissemination, connecting various services and entities, has been regarded as an essential service in ITS, and range query service has also been frequently used for on-demand traffic message retrieval. Meanwhile, as the dataset grows, the service provider tends to outsource the range query services to fog nodes. However, as the fog nodes are not fully trustable, directly outsourcing the services to them may invoke security concerns: on the one hand, the private information, including the traffic messages, query requests, and results should not be revealed to the fog nodes. On the other hand, the fog nodes may return incorrect and/or incomplete query results. Although many schemes have been proposed to achieve privacy-preserving range queries, few of them can support verifiability over dynamic datasets. Therefore, in this paper, we propose an efficient and verifiable range query (EVRQ) scheme over encrypted traffic messages. Specifically, we first build two techniques to respectively enable the fog nodes to efficiently determine and prove whether two rectangles intersect. Then, based on the two approaches and R-tree technique, we create a Merkle R-tree, which we use to build our EVRQ scheme. Security analysis shows that our proposed scheme is really privacy-preserving and can achieve verifiability of the query results. In addition, extensive experiments are conducted, and the results demonstrate that our proposed scheme is indeed efficient. Yunguo Guan, Pulei Xiong, Rongxing Lu |
ICC | 2 |
| 2022 | Towards a robust and trustworthy machine learning system development: An engineering perspective
Pulei Xiong, Scott Buffett, Shahrear Iqbal, Philippe Lamontagne 0001, Mohammad Saiful Islam Mamun, Heather Molyneaux |
J. Inf. Secur. Appl. | 1 |
| 2021 | Verification Based Scheme to Restrict IoT AttacksabstractIn recent years, with the increased usage of the Internet of Things (IoT) devices, cyber-attacks have become a serious threat over the Internet. These devices have low memory capacity and processing power, which makes them easy targets for attackers. The research community has proposed different approaches to deal with emerging variants of attacks on IoT devices using various machine learning techniques. However, these approaches rely heavily on the classifier’s categorization of a given record while ignoring its confidence. This paper proposes a verification-based scheme to reject IoT attacks by utilizing the classifier’s confidence. At the same time, existing studies are evaluated using traditional cross-validation approaches (e.g., k-fold), thus, not tested against unknown attacks. We propose using the leave-one-attack-out (LOAO) cross-validation scheme to evaluate the generalizability of the application to unknown attacks. The experiments are performed on Med BIoT, a publicly available dataset consisting of three IoT attacks. The system’s robustness is evaluated in terms of Receiver Operating Curves (ROC) and Equal Error rates (EERs). The results indicate a lower false-positive rate of 12.6% using the proposed verification-based approach in comparison to k-fold cross-validation. Barjinder Kaur, Sajjad Dadkhah, Pulei Xiong, Shahrear Iqbal, Suprio Ray, Ali A. Ghorbani 0001 |
BDCAT | 3 |
| 2021 | Privacy-Preserving Fog-Based Multi-Location Task Allocation in Mobile CrowdsourcingabstractAs a wave of the rapidly approaching future, vehicles are becoming increasingly “smarter” via equipping with abundant resources for data sensing, processing, and transmitting. To fully make use of these resources, mobile crowdsourcing applications have attracted particular interests from academia and industry, and they are extensively integrated with fog computing for obtaining low latency and location sensitivity. In this paper, we consider a fog-based task allocation service for mobile crowdsourcing tasks with multiple locations, where a task is allocated to the worker whose future trajectory has the smallest Hausdorff semi-distance to the task locations. However, as fog nodes are not fully trusted, there may exist privacy concerns related to the workers and the task owners. To the best of our knowledge, although Hausdorff semi-distance has been applied in various applications, none of the existing works can support privacy-preserving Hausdorff semi-distance evaluation or$k$nearest neighbor queries. Aiming at this issue, we design a privacy-preserving fog-based multi-location task allocation scheme. Specifically, based on a symmetric homomorphic encryption technique, we build two privacy-preserving protocols for i) computing min/max value from an array and ii) retrieving the key linked to the minimum value from an array of key-value pairs. Then, the proposed scheme is built upon these two protocols. After that, we conduct rigid security analysis and extensive experiments to demonstrate the security and efficiency of our proposed scheme, respectively. The results indicate that our proposed scheme is not only privacy-preserving, but also efficient in terms of both computation and communication costs. Yunguo Guan, Pulei Xiong, Rongxing Lu |
GLOBECOM | 2 |
| 2021 | EPSim-GS: Efficient and Privacy-Preserving Similarity Range Query over Genomic SequencesabstractSimilarity query over genomic sequences has played a significant role in personalized medicine and has applications in various fields, including DNA alignment and genomic sequencing. Since handling genomic sequences requires massive storage and considerable computational capacity, service providers prefer to process similarity queries over genomic sequences on cloud servers rather than at the client side. Due to the sensitivity of genomic sequences, preserving the privacy of queries has attracted considerable attention, and as a result, genomic sequences are demanded to be outsourced in an encrypted form. Although many schemes have been proposed for similarity queries over encrypted genomic data, they are either inefficient or have limitations in supporting the dynamic update of the dataset. To address the challenges, we propose an efficient and privacy-preserving similarity range query scheme, namely EPSim-GS. First, we introduce how to build a hash table to index the dataset, and present a similarity range query algorithm based on the hash table. Then, we design two cloud-based privacy-preserving protocols based on the Paillier cryptosystem to support the similarity range query algorithm over the encrypted dataset. After that, we propose EPSim-GS by leveraging the two privacy-preserving protocols. We then analyze the security of EPSim-GS and prove that it is privacy-preserving. Finally, we perform experiments to evaluate the scheme’s performance, and the results indicate that it is computationally efficient. Jiacheng Jin, Yandong Zheng, Pulei Xiong |
PST | 3 |
| 2021 | Traceable and Privacy-Preserving Non-Interactive Data Sharing in Mobile CrowdsensingabstractData sharing is one of the key technologies, which provides the practice of making data collected from a crowd of mobile devices available to others using a cloud infrastructure, known as mobile crowdsensing (MCS). However, the collected data may contain sensitive information, and sharing them in public clouds without proper protection could cause serious security problems, such as privacy leakage, unauthorized access, and secret key abuse. To address the above issues, in this paper, we propose a Traceable and privacy-preserving non-Interactive Data Sharing (TIDS) scheme in mobile crowdsensing. Specifically, to achieve privacy-preserving fine-grained data sharing, an attribute-based access policy is generated by a data owner without interacting with data users in the TIDS. Furthermore, we design a ciphertext conversion mechanism to support flexible data sharing. Also, by utilizing traceable Ciphertext-Policy Attribute-Based Encryption (CP-ABE), TIDS supports a trusted authority to trace malicious users who abuse their secret keys without incurring additional computational overhead. Security analysis demonstrates that TIDS can protect the confidentiality of the outsourced data. Experimental results show that TIDS can achieve efficient data sharing in mobile crowdsensing applications. Fuyuan Song, Zheng Qin 0001, Jinwen Liang, Pulei Xiong, Xiaodong Lin 0001 |
PST | 4 |
| 2010 | A model-driven penetration test framework for Web applicationsabstractPenetration testing is widely used to audit the security protection of Web applications. However, it is often performed by specialized security experts after development is completed and the application deployed into production. In this paper, we propose a model-driven penetration test framework for Web applications which provides a repeatable, systematic and cost-efficient approach fully integrated into a Security-Oriented Software Development Life Cycle. Security experts are still required to maintain knowledge used by the framework, but regular testing personnel are capable of creating, running and maintaining penetration test campaigns. A prototype of the framework has been implemented and applied to two Web applications: the benchmark WebGoat web application, and a hospital adverse event management system currently under development. A preliminary evaluation based on the prototype demonstrates the feasibility and efficiency of the proposed framework. Pulei Xiong, Liam Peyton |
PST | 1 |
| 2008 | Framework testing of web applications using TTCN-3
Bernard Stepien, Liam Peyton, Pulei Xiong |
Int. J. Softw. Tools Technol. Transf. | 3 |