EDBT 2026 Demo / reviewers in the wild / expert
Wanlun Ma
dblp:206/7127
· DBLP profile ↗
17ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-6305-1740ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 9 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Securing the low-altitude economy: a surveyabstractAbstract The rapid growth of the low-altitude economy, including unmanned aerial vehicles (UAVs) and urban air mobility (UAM), is reshaping industries from transportation to emergency response. Powered by advances in fifth-generation (5G) and 5G-advanced (5.5G) connectivity, artificial intelligence (AI), and new energy systems, these platforms are becoming increasingly autonomous and capable. However, their growing software complexity introduces critical cybersecurity risks. Vulnerabilities in communication protocols, onboard firmware, and AI systems can be exploited to hijack UAVs, disrupt operations, or leak sensitive data. While research has addressed isolated aspects, a unified security perspective is still lacking. This work presents a systematic review of software-level security challenges and defenses in low-altitude UAV/UAM systems. We first categorize major attack surfaces across communication, firmware, and AI layers. Furthermore, we survey defense mechanisms suited to real-time, resource-constrained aerial platforms. Finally, we propose future directions, including quantum-resistant communication protocols, hardware-software cosecurity, and edge-AI-driven architectures. Our work aims to inform researchers, practitioners, and regulators in developing integrated, resilient security strategies for the evolving low-altitude ecosystem. Minrui Yan, Ruiqi Dong, Qing-Long Han, Zehang Deng, Wanlun Ma, Xiaogang Zhu 0001, Wei Zhou 0044, Sheng Wen, Yang Xiang 0001 |
Sci. China Inf. Sci. | 5 |
| 2026 | Rethinking Query Choices for Differential Privacy AuditingabstractAuditing differential privacy (DP) guarantees often relies on querying trained models with specially crafted queries, such as canaries, examples differing between two neighboring datasets. However, in this work, we revisit this common approach and identify a fundamental limitation: canary-based queries may not capture the strongest privacy leakage, as the most informative queries can shift during the training process. This mismatch can result in loose lower bounds on the privacy parameter$\varepsilon$, underestimating potential risks from query-based adversaries. To address this issue, we propose two methods. First, we introduce a consistent and optimizable surrogate privacy loss function that better aligns with the true privacy loss, called Privacy-loss Maximization Method (PMM), enabling systematic discovery of stronger queries through optimization. Second, we analyze how the optimal queries evolve with model training and propose a gradient-aligned query generation algorithm, called Gradient-Guided Querying (GGQ), that rapidly identifies high-risk queries by aligning their gradients with the distribution of model parameters. Empirical evaluations across multiple tasks demonstrate that our methods consistently produce stronger privacy audit results, offering a more accurate assessment of the privacy risks associated with training algorithms. Zehang Deng, Shan Jiang 0023, Wanlun Ma, Sheng Wen, Tianqing Zhu, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | Untargeted Poisoning Membership Inference With Sample Selection and EnhancementabstractUntargeted poisoning membership inference (PMI) attacks are a newly emerging privacy threat that evaluates the impact of poisoned samples on the privacy leakage risk in the model's training dataset. Existing approaches typically select target samples randomly from a candidate dataset to generate poisoned samples, which are then injected into the training dataset. While effective in amplifying privacy leakage risks, this random selection strategy overlooks the fact that each poisoned sample contributes unequally to the attack. In this paper, we first observe that randomly selected target samples may be distant from and dispersed relative to high-confidence benign samples, which restricts the effectiveness of membership inference attacks. We then show that selecting target samples with high confidence in their ground-truth class to generate poisoned samples contributes more significantly to the attack. Therefore, we propose a novel untargeted PMI attack incorporating the target sample selection and enhancement. Specifically, we train shadow models to select the highest-confidence target samples for poison generation. To further enhance the effectiveness of the attack, we introduce a noise generator that adds adversarial perturbations to the selected target samples. Experimental results demonstrate that our approach significantly improves the attack success rate (e.g.,$85.6\%$compared to the baseline of$71.9\%$). Notably, the ablation study shows that our noise generator enhances privacy leakage risks even when target samples are selected randomly, highlighting its effectiveness and broad applicability. Xiaochun Yang 0001, Wanlun Ma, Bin Wang 0015, Sheng Wen, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2026 | Reverse Engineering of Industrial Protocols From Network TrafficabstractReliable protocol knowledge is often difficult to obtain in industrial networks, as industrial communications come with limited documentation, vendor-specific encodings, and opaque payloads. This lack of transparency hinders message interpretation and protocol analysis. To recover this missing protocol knowledge, network-trace-based protocol reverse engineering (PRE) infers message structure, field roles, and interaction logic directly from recorded traces. This enables protocol-aware intrusion detection, process monitoring, and protocol testing and fuzzing without access to device internals. Although PRE has advanced rapidly, existing techniques are developed under diverse objectives and assumptions. As a result, it is often unclear how isolated results relate to an end-to-end reverse-engineering workflow, and how evaluation outcomes should be compared across tasks and protocols. In this article, we cast reverse engineering of industrial protocols from network traces as a task-driven pipeline and articulate a unified task decomposition spanning message type identification, protocol syntax and semantic inference, payload pattern recognition and semantic inference, and protocol state machine reconstruction. For each task, we describe key methodological themes, common evaluation practices, and practical limitations that affect robustness and deployability in industrial settings. We further discuss security, privacy, and ethical risks that accompany increasingly capable PRE, and identify promising research directions toward more systematic, dependable, and deployment-oriented PRE methodologies. Chuan Sheng, Shan Jiang 0023, Qing-Long Han, Wei Zhou 0044, Wanlun Ma, Xiaogang Zhu 0001, Sheng Wen, Yang Xiang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | Codebreaker: Dynamic Extraction Attacks on Code Language ModelsabstractWith the rapid adoption of LLM-based code assistants to enhance programming experiences, concerns over extraction attacks targeting private training data have intensified. These attacks specifically aim to extract Personal Information (PI) embedded within the training data of code generation models (CodeLLMs). Existing methods, using either manual or semi-automated techniques, have successfully extracted sensitive data from these CodeLLMs. However, the limited amount of data currently retrieved by extraction attacks risks significantly underestimating the true extent of training data leakage. In this paper, we propose an automatic PI data extraction attack framework against LLM-based code assistants, named Codebreaker. This framework is built on two core components: (i) the introduction of semantic entropy, which evaluates the likelihood of a prompt triggering the model to respond with training data; and (ii) an automatic dynamic mutation mechanism that seamlessly integrates with Codebreaker, reinforcing the iterative process across the framework and promoting greater interconnection between different PI elements within a single response. This boosts reasoning diversity, model memorization, and finally attack performance. Using six series of open-source CodeLLMs (i.e., CodeParrot, StarCoder2, Code Llama, CodeGemma, DeepSeek-Coder, DeepSeek-V3) and two commercial code assistants (i.e., CodeFuse and GPT), we demonstrate the effectiveness of our proposed framework: (i) Codebreaker outperforms all current state-of-the-art extraction attacks by 6.22% ~ 44.9% (averaging 21.79%); (ii) when PI within a single response originates from the same GitHub repository, our framework - considering multiple interconnections in the response - exceeds others by 3.88% ~ 32.37% (averaging 15.31%). Furthermore, we discuss potential defenses, highlighting the urgent need for stronger measures to prevent PI leakage at the base model level. Changzhou Han, Zehang Deng, Wanlun Ma, Xiaogang Zhu 0001, Minhui Xue 0001, Tianqing Zhu, Sheng Wen, Yang Xiang 0001 |
SP | 3 |
| 2025 | PmiaNLL: Defending against poisoning membership inference attacks with noisy label learning
Xiaochun Yang 0001, Wanlun Ma, Bin Wang 0015, Tianqing Zhu, Sheng Wen, Yang Xiang 0001 |
Knowl. Based Syst. | 3 |
| 2025 | Hardening LLM Fine-Tuning: From Differentially Private Data Selection to Trustworthy Model QuantizationabstractCritical infrastructures are increasingly integrating artificial intelligence (AI) technologies, including large language models (LLMs), into essential systems and services that are vital to societal functioning. Fine-tuning LLMs for specific domain tasks are crucial for their effective deployment in these contexts, but this process must carefully address both privacy and security concerns. Without proper safeguards, such integration can introduce additional risks, such as data leakage during training and diminished model trustworthiness due to the need for model compression to operate within limited bandwidth and computational capacity constraints. In this paper, we proposeHardening LLM Fine-tuning framework(HARDLLM), which addresses these challenges through two key components: (i) we develop a differentially private data selection method that ensures privacy protection by training the model exclusively on sampled and synthesized public data, thereby preventing any direct use of private data and enhancing leakage resilience throughout the training process, and (ii) we introduce a trustworthiness-aware model quantization approach to improve LLMs performance, such as reducing toxicity, enhancing adversarial robustness, and mitigating stereotypes, while maintaining negligible impact on model utility. Experimental results show that, the proposed algorithm ensures differential privacy when privacy budget is set at ϵ = 0.5, with only a 1% drop in accuracy, while other state-of-the-art methods experience an accuracy drop of at least 20% under the same privacy budget. Additionally, our quantization approach improves the trustworthiness of fine-tuned LLMs by an average of 3-4%, with only a negligible utility loss (approximately 1%) at a 50% compression rate. Zehang Deng, Ruoxi Sun 0001, Minhui Xue 0001, Wanlun Ma, Sheng Wen, Surya Nepal, Yang Xiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | TrapNet: Model Inversion Defense via TrapdoorabstractModel inversion (MI) attacks, for which effective defense strategies are still lacking, pose significant risks to privacy by reconstructing private training data through access to well-trained classifiers. Addressing this concern, this study introduces TrapNet, designed to defend against advanced MI attacks while maintaining good model utility. TrapNet intentionally injects trapdoors into the classification manifold of the protected target model. In this way, TrapNet can effectively mislead MI attack optimization. Specifically, TrapNet leverages a conditional GAN (cGAN) trained on the private dataset to generate diverse and realistic trapdoor samples. In addition, we propose a graph-matching self-obfuscation strategy and an entropy regularization technique to optimize trapdoor injection while preserving model utility. Compared to the existing defense, TrapNet can provide universal protection to all target classes without access to any auxiliary public data. Extensive experiments on CelebA, VGG-Face, and VGG-Face2 datasets demonstrate TrapNet’s superior performance over existing defenses, including the most advanced NetGuard and BiDO, against state-of-the-art model inversion attacks, i.e., PLG-MI, LOMMA, and Plug&Play. Wanlun Ma, Derui Wang, Yiliao Song, Minhui Xue 0001, Sheng Wen, Zhengdao Li, Yang Xiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2025 | Network Traffic Fingerprinting for IIoT Device Identification: A SurveyabstractAs the Industrial Internet of Things (IIoT) continues to expand, the need for effective device identification becomes critical for securing industrial environments. Network traffic fingerprinting has emerged as an important technique for IIoT device identification, leveraging the unique communication patterns embedded in network traffic. Despite significant efforts in this area, a comprehensive overview of the relevant research is still missing. To address the lack of comprehensive research, this paper, for the first time, identifies critical knowledge gaps constraining IIoT device identification through network traffic analysis: obscure fingerprint feature space, limited generalizability to unknowns, and scarce data sources. Focusing on these gaps, existing methods are analyzed and summarized in detail across network traffic fingerprinting, IIoT device identification, and public IIoT datasets. Specifically, network traffic fingerprinting methods are categorized into three levels: Packet-level, flow-level, and business-level, and relevant methods are examined in terms of data formats, segmentation units, and extraction or generation techniques. In the context of IIoT device identification, tasks such as device type, model, and instance recognition, as well as abnormal device detection, are extensively investigated using rule-based, traditional machine learning- based, and deep learning-based approaches, with a focus on device fingerprints and application scenarios. Furthermore, main public datasets from the IoT, ICS, and IIoT scenarios are highlighted to support the development of fingerprinting and identification methods. Finally, several future research directions are proposed to guide new advancements in this area. Chuan Sheng, Wei Zhou 0044, Qing-Long Han, Wanlun Ma, Xiaogang Zhu 0001, Sheng Wen, Yang Xiang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2024 | The "Code" of Ethics: A Holistic Audit of AI Code GeneratorsabstractAI-powered programming language generation (PLG) models have gained increasing attention due to their ability to generate source code of programs in a few seconds with a plain program description. Despite their remarkable performance, many concerns are raised over the potential risks of their development and deployment, such as legal issues of copyright infringement induced by training usage of licensed code, and malicious consequences due to the unregulated use of these models. In this paper, we present the first-of-its-kind study to systematically investigate the accountability of PLG models from the perspectives of both model development and deployment. In particular, we develop a holistic framework not only to audit the training data usage of PLG models, but also to identify neural code generated by PLG models as well as determine its attribution to a source model. To this end, we propose using membership inference to audit whether a code snippet used is in the PLG model's training data. In addition, we propose a learning-based method to distinguish between human-written code and neural code. In neural code attribution, through both empirical and theoretical analysis, we show that it is impossible to reliably attribute the generation of one code snippet to one model. We then propose two feasible alternative methods: one is to attribute one neural code snippet to one of the candidate PLG models, and the other is to verify whether a set of neural code snippets can be attributed to a given PLG model. The proposed framework thoroughly examines the accountability of PLG models which are verified by extensive experiments. The implementations of our proposed framework are also encapsulated into a new artifact, named CODEFORENSIC, to foster further research. Wanlun Ma, Yiliao Song, Minhui Xue 0001, Sheng Wen, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2024 | LocGuard: A Location Privacy Defender for Image SharingabstractThe privacy of social media users is a major concern when the users share their content to the public. Sensitive information such as the location of the users can be inferred from relevant content without arising the awareness of the users. With blooming services provided by social media platforms, the users have more freedom to share information via diverse data formats. The multi-modality of the shared information may, in return, worsen the private information leakage caused by inference attacks. In this paper, we first examine the problem of location inference on multi-modal data comprised of textual information and visual content. It is observed that the visual content, such as photos shared by social media users, can significantly boost the success rate of location inference. To thwart adversaries who are driven by visual-related data, we propose a defence that mitigates the threat of location privacy breach under an imperceptible utility loss. Our defence, namely LocGuard, perturbs the photos in a one-off manner before sharing them. The perturbations, along with a simple but effective bipartite perturbation strategy, ensure that LocGuard is resistant to adaptive adversaries who can perform adversarial training based on the perturbed photos. Moreover, LocGuard remains effective against open-set adversaries whose data categories in the training dataset are hidden from the defender. In the evaluation, we conduct extensive experiments based on real-world datasets and compare our work with previous methods. The results show that LocGuard significantly outperforms the existing defences. In particular, LocGuard not only achieves better privacy protection and utility preservation for image sharing, but also can effectively defend against adversarial-training-capable attackers. Wanlun Ma, Derui Wang, Chao Chen 0015, Sheng Wen, Gaolei Fei, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2023 | Multi-modal Representation Learning for Social Post Location InferenceabstractInferring geographic locations via social posts is essential for many practical location-based applications such as product marketing, point-of-interest recommendation, and infector tracking for COVID-19. Unlike image-based location retrieval or social-post text embedding-based location inference, the combined effect of multi-modal information (i.e., post images, text, and hashtags) for social post positioning receives less attention. In this work, we collect real datasets of social posts with images, texts, and hashtags from Instagram and propose a novel Multi-modal Representation Learning Framework (MRLF) capable of fusing different modalities of social posts for location inference. MRLF integrates a multi-head attention mechanism to enhance location-salient information extraction while significantly improving location inference compared with single domain-based methods. To overcome the noisy user-generated textual content, we introduce a novel attention-based character-aware module that considers the relative dependencies between characters of social post texts and hashtags for flexible multi-model information fusion. The experimental results show that MRLF can make accurate location predictions and open a new door to understanding the multi-modal data of social posts for online inference tasks. Ruiting Dai, Xucheng Luo, Lisi Mo, Wanlun Ma, Fan Zhou 0002 |
ICC | 5 |
| 2023 | The "Beatrix" Resurrections: Robust Backdoor Detection via Gram Matrices
Wanlun Ma, Derui Wang, Ruoxi Sun 0001, Minhui Xue 0001, Sheng Wen, Yang Xiang 0001 |
NDSS | 1 |
| 2023 | Real-Time Detection of COVID-19 Events From Twitter: A Spatial-Temporally Bursty-Aware MethodabstractIn the last two years, the outbreak of COVID-19 has significantly affected human life, society, and the economy worldwide. To prevent people from contracting COVID-19 and mitigate its spread, it is crucial to timely distribute complete, accurate, and up-to-date information about the pandemic to the public. In this article, we propose a spatial–temporally bursty-aware method calledSTBAfor real-time detection of COVID-19 events from Twitter.STBAhas three consecutive stages. In the first stage,STBAidentifies a set of keywords that represent COVID-19 events according to the spatiotemporally bursty characteristics of words using Ripley’s$K$function.STBAwill also filter out tweets that do not contain the keywords to reduce the interference of noise tweets on event detection. In the second stage,STBAuses online density-based spatial clustering of applications with noise clustering to aggregate tweets that describe the same event as much as possible, which provides more information for event identification. In the third stage,STBAfurther utilizes the temporal bursty characteristic of event location information in the clusters to identify real-world COVID-19 events. Each stage ofSTBAcan be regarded as a noise filter. It gradually filters out COVID-19-related events from noisy tweet streams. To evaluate the performance ofSTBA, we collected over 116 million Twitter posts from 36 consecutive days (from March 22, 2020 to April 26, 2020) and labeled 501 real events in this dataset. We comparedSTBAwith three state-of-the-art methods, EvenTweet, event detection via microblog cliques (EDMC), and GeoBurst+ in the evaluation. The experimental results suggest thatSTBAoutperforms GeoBurst+ by 13.8%, 12.7%, and 13.3% in terms of precision, recall, and$F_{1}$score.STBAachieved even more improvements compared with EvenTweet and EDMC. Gaolei Fei, Wanlun Ma, Chao Chen 0015, Sheng Wen, Guangmin Hu |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2022 | Analysis of Trending Topics and Text-based Channels of Information Delivery in CybersecurityabstractComputer users are generally faced with difficulties in making correct security decisions. While an increasingly fewer number of people are trying or willing to take formal security training, online sources including news, security blogs, and websites are continuously making security knowledge more accessible. Analysis of cybersecurity texts from this grey literature can provide insights into the trending topics and identify current security issues as well as how cyber attacks evolve over time. These in turn can support researchers and practitioners in predicting and preparing for these attacks. Comparing different sources may facilitate the learning process for normal users by creating the patterns of the security knowledge gained from different sources. Prior studies neither systematically analysed the wide range of digital sources nor provided any standardisation in analysing the trending topics from recent security texts. Moreover, existing topic modelling methods are not capable of identifying the cybersecurity concepts completely and the generated topics considerably overlap. To address this issue, we propose a semi-automated classification method to generate comprehensive security categories to analyse trending topics. We further compare the identified 16 security categories across different sources based on their popularity and impact. We have revealed several surprising findings as follows: (1) The impact reflected from cybersecurity texts strongly correlates with the monetary loss caused by cybercrimes, (2) security blogs have produced the context of cybersecurity most intensively, and (3) websites deliver security information without caring about timeliness much. Tingmin Wu, Wanlun Ma, Sheng Wen, Xin Xia 0001, Cécile Paris, Surya Nepal, Yang Xiang 0001 |
ACM Trans. Internet Techn. | 2 |
| 2020 | What risk? I don't understand. An Empirical Study on Users' Understanding of the Terms Used in Security TextsabstractUsers receive a multitude of security information in written articles, e.g., newspapers, security blogs, and training materials. However, prior research suggests that these delivery methods, including security awareness campaigns, mostly fail to increase people's knowledge about cyber threats. It seems that users find such information challenging to absorb and understand. Yet, to raise users' security awareness and understanding, it is essential to ensure the users comprehend the provided information so that they can apply the advice it contains in practice. We conducted a subjective study to measure the level of users' understanding of security texts. We find that 61% of the terms security experts used in their writings are hard for the public to understand, even for people with some IT backgrounds. We also observe that 88% of security texts have at least one such term. Moreover, we notice that existing dictionaries, including the online ones (e.g., Google Dictionary), cover no more than 35% of the terms found in security texts. To improve users' ability to understand security texts, we developed a framework to build a user-oriented security-centric dictionary from multiple sources. To evaluate the effectiveness of the dictionary, we developed a tool as a service to detect technical terms and explain their meanings to the user in pop-ups. The results of a subjective study to measure the tool's performance showed that it could increase users' ability to understand security articles by 30%. Tingmin Wu, Rongjunchen Zhang, Wanlun Ma, Sheng Wen, Xin Xia 0001, Cécile Paris, Surya Nepal, Yang Xiang 0001 |
AsiaCCS | 3 |
| 2017 | My Face is Mine: Fighting Unpermitted Tagging on Personal/Group Photos in Social Media
Lihong Tang, Wanlun Ma, Sheng Wen, Marthie Grobler, Yang Xiang 0001, Wanlei Zhou 0001 |
WISE (2) | 2 |