VLDB 2026 Research / reviewers in the wild / expert
Ruoxi Sun 0001
dblp:72/7683-1
· DBLP profile ↗
29ranked-venue papers
5as first author
26since 2021 · last 2026
0000-0001-5404-8550ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 17 · 17 since 2021Software engineering, systems software and programming languages · 6 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Re-Key-Free, Risky-Free: Adaptable Model Usage Control
Zhongkui Ma, Xinguo Feng, Chuan Yan, Dongge Liu, Ruoxi Sun 0001, Derui Wang, Minhui Xue 0001, Guangdong Bai |
EuroS&P | 6 |
| 2026 | Unfairness Attack and Unified Provable Defense on AI-Powered Internet of EnergyabstractThe critical energy infrastructure is undergoing two significant transformations: the rapid increase in renewable distributed energy resources (DER) and the digitalization of the energy sector, collectively shaping what is known as the Internet of Energy (IoE). Artificial intelligence (AI) has become a widely adopted tool for effectively allocating energy and managing sector-related resources, where ensuring fairness is essential. While inherent unfairness in AI systems is well acknowledged, little attention has been given to evaluating this unfairness and its real-world implications within the context of the IoE. In this study, we take a first step to elucidate the unfairness in AI-powered IoE systems induced by malicious users. We introduce Unfairness Score (UScore), a novel metric designed to evaluate the unfairness of machine learning models in real-world IoE scenarios. We then extensively evaluate unfairness attacks using three IoE tabular datasets, demonstrating that AI model fairness can be compromised through data poisoning, whether in centralized learning (CL) or federated learning (FL) settings. Notably, such compromises can occur when malicious users tamper with only a small subset of the data they control. Finally, we propose a novel approach that unifies fairness and differential privacy (DP) by leveraging DP as a provable defense mechanism. This approach provides a universally applicable solution to unfairness attacks, regardless of whether the learning tasks are classification or regression, and is effective in both FL and CL settings. Our contributions represent a significant step in addressing unfairness and privacy concerns in AI-powered IoE systems. Ruoxi Sun 0001, Xin Yuan 0004, Minhui Xue 0001, Yansong Gao 0001, Surya Nepal, Xingliang Yuan, Carsten Rudolph, Ling Liu 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | What's Pulling the Strings? Evaluating Integrity and Attribution in AI Training and Inference through Concept Shift
Jiamin Chang, Haoyang Li 0018, Hammond A. Pearce, Ruoxi Sun 0001, Bo Li 0026, Minhui Xue 0001 |
CCS | 4 |
| 2025 | Edge Unlearning is Not "on Edge"! an Adaptive Exact Unlearning System on Resource-Constrained DevicesabstractThe right to be forgotten mandates that machine learning models enable the erasure of a data owner's data and information from a trained model. Removing data from the dataset alone is inadequate, as machine learning models can memorize information from the training data, increasing the potential privacy risk to users. To address this, multiple machine unlearning techniques have been developed and deployed. Among them, approximate unlearning is a popular solution, but recent studies report that its unlearning effectiveness is not fully guaranteed. Another approach, exact unlearning, tackles this issue by discarding the data and retraining the model from scratch, but at the cost of considerable computational and memory resources. However, not all devices have the capability to perform such retraining. In numerous machine learning applications, such as edge devices, Internet-of-Things (IoT), mobile devices, and satellites, resources are constrained, posing challenges for deploying existing exact unlearning methods. In this study, we propose a Constraint-aware Adaptive Exact Unlearning System at the network Edge (CAUSE), an approach to enabling exact unlearning on resource-constrained devices. Aiming to minimize the retrain overhead by storing sub-models on the resource-constrained device, CAUSE inno-vatively applies a Fibonacci-based replacement strategy and updates the number of shards adaptively in the user-based data partition process. To further improve the effectiveness of memory usage, CAUSE leverages the advantage of model pruning to save memory via compression with minimal accuracy sacrifice. The experimental results demonstrate that CAUSE significantly outperforms other representative systems in realizing exact unlearning on the resource-constrained device by 9.23%-80.86%, 66.21%-83.46%, and 5.26%-194.13% in terms of unlearning speed, energy consumption, and accuracy. Xiaoyu Xia 0001, Ziqi Wang 0008, Ruoxi Sun 0001, Bowen Liu 0002, Ibrahim Khalil 0001, Minhui Xue 0001 |
SP | 3 |
| 2025 | 50 Shades of Deceptive Patterns: A Unified Taxonomy, Multimodal Detection, and Security ImplicationsabstractDeceptive patterns (DPs) are user interface designs deliberately crafted to manipulate users into unintended decisions, often by exploiting cognitive biases for the benefit of companies or services. While numerous studies have explored ways to identify these deceptive patterns, many existing solutions require significant human intervention and struggle to keep pace with the evolving nature of deceptive designs. To address these challenges, we expanded the deceptive pattern taxonomy from security and privacy perspectives, refining its categories and scope. We created a comprehensive dataset of deceptive patterns by integrating existing small-scale datasets with new samples, resulting in 6,725 images and 10,421 DP instances from mobile apps and websites. We then developed DPGuard, a novel automatic tool leveraging commercial multimodal large language models (MLLMs) for deceptive pattern detection. Experimental results show that DPGuard outperforms state-of-the-art methods. An extensive empirical evaluation on 2,000 popular mobile apps and websites reveals that 25.7% of mobile apps and 49.0% websites feature at least one deceptive pattern instance. Through 4 unexplored case studies that inform security implications, we highlight the critical importance of the unified taxonomy in addressing the growing challenges of Internet deception. Zewei Shi, Ruoxi Sun 0001, Jieshan Chen, Jiamou Sun, Minhui Xue 0001, Yansong Gao 0001, Feng Liu 0003, Xingliang Yuan |
WWW | 2 |
| 2025 | Leakage-Resilient and Carbon-Neutral Aggregation Featuring the Federated AI-Enabled Critical InfrastructureabstractAI-enabled critical infrastructures (ACIs) integrate artificial intelligence (AI) technologies into various essential systems and services that are vital to the functioning of society, offering significant implications for efficiency, security and resilience. While adopting decentralized AI approaches (such as federated learning technology) in ACIs is plausible, private and sensitive data are still susceptible to data reconstruction attacks through gradient optimization. In this work, we propose Compressed Differentially Private Aggregation (CDPA), a leakage-resilient, communication-efficient, and carbon-neutral approach for ACI networks. Specifically, CDPA has introduced a novel random bit-flipping mechanism as its primary innovation. This mechanism first converts gradients into a specific binary representation and then selectively flips masked bits with a certain probability. The proposed bit-flipping introduces a larger variance to the noise while providing differentially private protection and commendable efforts in energy savings while applying vector quantization techniques within the context of federated learning. The experimental evaluation indicates that CDPA can reduce communication cost by half while preserving model utility. Moreover, we demonstrate that CDPA can effectively defend against state-of-the-art data reconstruction attacks in both computer vision and natural language processing tasks. We highlight existing benchmarks that generate 2.6x to over 100x more carbon emissions than CDPA. We hope that the CDPA developed in this paper can inform the federated AI-enabled critical infrastructure of a more balanced trade-off between utility and privacy, resilience protection, as well as a better carbon offset with less communication overhead. Zehang Deng, Ruoxi Sun 0001, Minhui Xue 0001, Sheng Wen, Seyit Ahmet Çamtepe, Surya Nepal, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | On Security Weaknesses and Vulnerabilities in Deep Learning SystemsabstractThe security guarantee of AI-enabled software systems (particularly using deep learning techniques as a functional core) is pivotal against the adversarial attacks exploiting software vulnerabilities. However, little attention has been paid to a systematic investigation of vulnerabilities in such systems. A common situation learned from the open source software community is that deep learning engineers frequently integrate off-the-shelf or open-source learning frameworks into their ecosystems. In this work, we specifically look into deep learning (DL) framework and perform the firstsystematicstudy of vulnerabilities in DL systems through a comprehensive analysis of identified vulnerabilities from Common Vulnerabilities and Exposures (CVE) and open-source DL tools, including TensorFlow, Caffe, OpenCV, Keras, and PyTorch. We propose a two-stream data analysis framework to explore vulnerability patterns from various databases. We investigate the unique DL frameworks and libraries development ecosystems that appear to be decentralized and fragmented. By revisiting the Common Weakness Enumeration (CWE) List, which provides the traditional software vulnerability related practices, we observed that it is more challenging to detect and fix the vulnerabilities throughout the DL systems lifecycle. Moreover, we conducted a large-scale empirical study of 3,049 DL vulnerabilities to better understand the patterns of vulnerability and the challenges in fixing them. Zhongzheng Lai, Huaming Chen, Ruoxi Sun 0001, Yu Zhang 0177, Minhui Xue 0001, Dong Yuan 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | Iterative Window Mean Filter: Thwarting Diffusion-Based Adversarial PurificationabstractFace authentication systems have brought significant convenience and advanced developments, yet they have become unreliable due to their sensitivity to inconspicuous perturbations, such as adversarial attacks. Existing defenses often exhibit weaknesses when facing various attack algorithms and adaptive attacks or compromise accuracy for enhanced security. To address these challenges, we have developed a novel and highly efficient non-deep-learning-based image filter called the Iterative Window Mean Filter (IWMF) and proposed a new framework for adversarial purification, named IWMF-Diff, which integrates IWMF and denoising diffusion models. These methods can function as pre-processing modules to eliminate adversarial perturbations without necessitating further modifications or retraining of the target system. We demonstrate that our proposed methodologies fulfill four critical requirements: preserved accuracy, improved security, generalizability to various threats in different settings, and better resistance to adaptive attacks. This performance surpasses that of the state-of-the-art adversarial purification method, DiffPure. Our code is released athttps://github.com/azrealwang/iwmfdiff. Hanrui Wang 0005, Ruoxi Sun 0001, Cunjian Chen, Minhui Xue 0001, Lay-Ki Soon, Shuo Wang 0012, Zhe Jin 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | Hardening LLM Fine-Tuning: From Differentially Private Data Selection to Trustworthy Model QuantizationabstractCritical infrastructures are increasingly integrating artificial intelligence (AI) technologies, including large language models (LLMs), into essential systems and services that are vital to societal functioning. Fine-tuning LLMs for specific domain tasks are crucial for their effective deployment in these contexts, but this process must carefully address both privacy and security concerns. Without proper safeguards, such integration can introduce additional risks, such as data leakage during training and diminished model trustworthiness due to the need for model compression to operate within limited bandwidth and computational capacity constraints. In this paper, we proposeHardening LLM Fine-tuning framework(HARDLLM), which addresses these challenges through two key components: (i) we develop a differentially private data selection method that ensures privacy protection by training the model exclusively on sampled and synthesized public data, thereby preventing any direct use of private data and enhancing leakage resilience throughout the training process, and (ii) we introduce a trustworthiness-aware model quantization approach to improve LLMs performance, such as reducing toxicity, enhancing adversarial robustness, and mitigating stereotypes, while maintaining negligible impact on model utility. Experimental results show that, the proposed algorithm ensures differential privacy when privacy budget is set at ϵ = 0.5, with only a 1% drop in accuracy, while other state-of-the-art methods experience an accuracy drop of at least 20% under the same privacy budget. Additionally, our quantization approach improves the trustworthiness of fine-tuned LLMs by an average of 3-4%, with only a negligible utility loss (approximately 1%) at a 50% compression rate. Zehang Deng, Ruoxi Sun 0001, Minhui Xue 0001, Wanlun Ma, Sheng Wen, Surya Nepal, Yang Xiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Traceable and Collision-Resilient Differential PrivacyabstractDifferential Privacy (DP) is a preeminent technique for data privacy by introducing noise to sensitive information. However, traditional DP mechanisms excessively rely on third parties to ensure traceability, necessitating strong background assumptions that are frequently impractical in real-world scenarios. This reliance makes it difficult to preserve both privacy and traceability. To address these challenges, we propose a novel Traceable and Collision-Resilient Differential Privacy (TCRDP) mechanism. The TCRDP mechanism simultaneously publishes perturbed results and data fingerprints, retaining partial information from the original data in a collision-resilient manner to facilitate future verification. Moreover, the TCRDP mechanism integrates an innovative noise generation process, leveraging hash values and a customized Laplace-like distribution to produce noise. This strategy mitigates the risk of adversaries compromising privacy through enumeration and yields a more concentrated noise distribution with reduced variance. We evaluated the TCRDP mechanism using three datasets: ICUs, Diabetes, and RAHRD, across various query types. The experimental results demonstrated significant improvements in data utility, with the TCRDP mechanism achieving great reductions in Mean Absolute Error (MAE) and Mean Squared Error (MSE) compared to traditional mechanisms. The TCRDP mechanism also maintained lower Accuracy Loss (AL) across different privacy budgets and dataset sizes, highlighting its robustness and scalability. These findings underscore the potential of the TCRDP mechanism to advance privacy-preserving data analysis, offering significant enhancements over existing methods in both accuracy and utility. Kai Zhang 0074, Xin Yuan 0004, Ruoxi Sun 0001, Minhui Xue 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | A Duty to Forget, a Right to be Assured? Exposing Vulnerabilities in Machine Unlearning Services
Hongsheng Hu, Shuo Wang 0012, Jiamin Chang, Haonan Zhong, Ruoxi Sun 0001, Shuang Hao 0001, Haojin Zhu, Minhui Xue 0001 |
NDSS | 5 |
| 2024 | CORELOCKER: Neuron-level Usage ControlabstractThe growing complexity of deep neural network models in modern application domains necessitates a complex training process that involves extensive data, sophisticated design, and substantial computation. The trained model inherently encapsulates the intellectual property owned by the model developer (or the model owner). Consequently, safeguarding the model from unauthorized use by entities who obtain access to the model (or the model controllers), i.e., preserving the fundamental rights and proprietary interests of the model owner, has become a critical necessity.In this work, we propose CORELOCKER, employing the strategic extraction of a small subset of significant weights from the neural network. This subset serves as the access key to unlock the model’s complete capability. The extraction of the key can be customized to varying levels of utility that the model owner intends to release. Authorized users with the access key have full access to the model, while unauthorized users can have access to only part of its capability. We establish a formal foundation to underpin CORELOCKER, which provides crucial lower and upper bounds for the utility disparity between pre- and post-protected networks. We evaluate CORELOCKER using representative datasets such as Fashion-MNIST, CIFAR-10, and CIFAR-100, as well as real-world models including Vg-gNet, ResNet, and DenseNet. Our experimental results confirm its efficacy. We also demonstrate CORELOCKER’s resilience against advanced model restoration attacks based on fine-tuning and pruning. Zhongkui Ma, Xinguo Feng, Ruoxi Sun 0001, Hu Wang 0005, Minhui Xue 0001, Guangdong Bai |
SP | 4 |
| 2024 | Bounded and Unbiased Composite Differential PrivacyabstractThe objective of differential privacy (DP) is to protect privacy by producing an output distribution that is indistinguishable between any two neighboring databases. However, traditional differentially private mechanisms tend to produce unbounded outputs in order to achieve maximum disturbance range, which is not always in line with real-world applications. Existing solutions attempt to address this issue by employing post-processing or truncation techniques to restrict the output results, but at the cost of introducing bias issues. In this paper, we propose a novel differentially private mechanism which uses a composite probability density function to generate bounded and unbiased outputs for any numerical input data. The composition consists of an activation function and a base function, providing users with the flexibility to define the functions according to the DP constraints. We also develop an optimization algorithm that enables the iterative search for the optimal hyper-parameter setting without the need for repeated experiments, which prevents additional privacy overhead. Furthermore, we evaluate the utility of the proposed mechanism by assessing the variance of the composite probability density function and introducing two alternative metrics that are simpler to compute than variance estimation. Our extensive evaluation on three benchmark datasets demonstrates consistent and significant improvement over the traditional Laplace and Gaussian mechanisms. The proposed bounded and unbiased composite differentially private mechanism will underpin the broader DP arsenal and foster future privacy-preserving studies. Kai Zhang 0074, Yanjun Zhang 0002, Ruoxi Sun 0001, Pei-Wei Tsai, Muneeb Ul Hassan 0001, Xin Yuan 0004, Minhui Xue 0001, Jinjun Chen |
SP | 3 |
| 2024 | Privacy-Preserving and Fairness-Aware Federated Learning for Critical Infrastructure Protection and ResilienceabstractThe energy industry is undergoing significant transformations as it strives to achieve net-zero emissions and future-proof its infrastructure, where every participant in the power grid has the potential to both consume and produce energy resources. Federated learning -- which enables multiple participants to collaboratively train a model without aggregating the training data -- becomes a viable technology. However, the global model parameters that have to be shared for optimization are still susceptible to training data leakage. In this work, we propose confined gradient descent (CGD) that enhances the privacy of federated learning by eliminating the sharing of global model parameters. CGD exploits the fact that a gradient descent optimization can start with a set of discrete points and converges to another set in the neighborhood of the global minimum of the objective function. As such, each participant can independently initiate its own private global model~(referred to as the confined model ), and collaboratively learn it towards the optimum. The updates to their own models are worked out in a secure collaborative way during the training process.In such a manner, CGD retains the ability of learning from distributed data but greatly diminishes information sharing. Such a strategy also allows the proprietary confined models to adapt to the heterogeneity in federated learning, providing inherent benefits of fairness. We theoretically and empirically demonstrate that decentralized CGD øne provides a stronger differential privacy (DP) protection; \two is robust against the state-of-the-art poisoning privacy attacks; þree results in bounded fairness guarantee among participants; and \four provides high test accuracy (comparable with centralized learning) with a bounded convergence rate over four real-world datasets. Yanjun Zhang 0002, Ruoxi Sun 0001, Liyue Shen, Guangdong Bai, Minhui Xue 0001, Mark Huasong Meng, Xue Li 0001, Ryan Kok Leong Ko, Surya Nepal |
WWW | 2 |
| 2023 | The "Beatrix" Resurrections: Robust Backdoor Detection via Gram Matrices
Wanlun Ma, Derui Wang, Ruoxi Sun 0001, Minhui Xue 0001, Sheng Wen, Yang Xiang 0001 |
NDSS | 3 |
| 2023 | DOITRUST: Dissecting On-chain Compromised Internet Domains via Graph Learning
Shuo Wang 0012, Mahathir Almashor, Alsharif Abuadbba, Ruoxi Sun 0001, Minhui Xue 0001, Calvin Wang, Raj Gaire 0001, Surya Nepal, Seyit Ahmet Çamtepe |
NDSS | 4 |
| 2023 | Mate! Are You Really Aware? An Explainability-Guided Testing Framework for Robustness of Malware DetectorsabstractNumerous open-source and commercial malware detectors are available. However, their efficacy is threatened by new adversarial attacks, whereby malware attempts to evade detection, e.g., by performing feature-space manipulation. In this work, we propose an explainability-guided and model-agnostic testing framework for robustness of malware detectors when confronted with adversarial attacks. The framework introduces the concept of Accrued Malicious Magnitude (AMM) to identify which malware features could be manipulated to maximize the likelihood of evading detection. We then use this framework to test several state-of-the-art malware detectors' ability to detect manipulated malware. We find that (i) commercial antivirus engines are vulnerable to AMM-guided test cases; (ii) the ability of a manipulated malware generated using one detector to evade detection by another detector (i.e., transferability) depends on the overlap of features with large AMM values between the different detectors; and (iii) AMM values effectively measure the fragility of features (i.e., capability of feature-space manipulation to flip the prediction results) and explain the robustness of malware detectors facing evasion attacks. Our findings shed light on the limitations of current malware detectors, as well as how they can be improved. Ruoxi Sun 0001, Minhui Xue 0001, Gareth Tyson, Tian Dong 0003, Shaofeng Li 0001, Shuo Wang 0012, Haojin Zhu, Seyit Ahmet Çamtepe, Surya Nepal |
ESEC/SIGSOFT FSE | 1 |
| 2023 | StyleFool: Fooling Video Classification Systems via Style TransferabstractVideo classification systems are vulnerable to adversarial attacks, which can create severe security problems in video verification. Current black-box attacks need a large number of queries to succeed, resulting in high computational overhead in the process of attack. On the other hand, attacks with restricted perturbations are ineffective against defenses such as denoising or adversarial training. In this paper, we focus on unrestricted perturbations and propose StyleFool, a black-box video adversarial attack via style transfer to fool the video classification system. StyleFool first utilizes color theme proximity to select the best style image, which helps avoid unnatural details in the stylized videos. Meanwhile, the target class confidence is additionally considered in targeted attacks to influence the output distribution of the classifier by moving the stylized video closer to or even across the decision boundary. A gradient-free method is then employed to further optimize the adversarial perturbations. We carry out extensive experiments to evaluate StyleFool on two standard datasets, UCF-101 and HMDB-51. The experimental results demonstrate that StyleFool outperforms the state-of-the-art adversarial attacks in terms of both the number of queries and the robustness against existing defenses. Moreover, 50% of the stylized videos in untargeted attacks do not need any query since they can already fool the video classification model. Furthermore, we evaluate the indistinguishability through a user study to show that the adversarial samples of StyleFool look imperceptible to human eyes, despite unrestricted perturbations. Xi Xiao 0001, Ruoxi Sun 0001, Derui Wang, Minhui Xue 0001, Sheng Wen |
SP | 3 |
| 2023 | PublicCheck: Public Integrity Verification for Services of Run-time Deep ModelsabstractExisting integrity verification approaches for deep models are designed for private verification (i.e., assuming the service provider is honest, with white-box access to model parameters). However, private verification approaches do not allow model users to verify the model at run-time. Instead, they must trust the service provider, who may tamper with the verification results. In contrast, a public verification approach that considers the possibility of dishonest service providers can benefit a wider range of users. In this paper, we propose PublicCheck, a practical public integrity verification solution for services of run-time deep models. PublicCheck considers dishonest service providers, and overcomes public verification challenges of being lightweight, providing anti-counterfeiting protection, and having fingerprinting samples that appear smooth. To capture and fingerprint the inherent prediction behaviors of a run-time model, PublicCheck generates smoothly transformed and augmented encysted samples that are enclosed around the model's decision boundary while ensuring that the verification queries are indistinguishable from normal queries. PublicCheck is also applicable when knowledge of the target model is limited (e.g., with no knowledge of gradients or model parameters). A thorough evaluation of PublicCheck demonstrates the strong capability for model integrity breach detection (100% detection accuracy with less than 10 black-box API queries) against various model integrity attacks and model compression attacks. PublicCheck also demonstrates the smooth appearance, feasibility, and efficiency of generating a plethora of encysted samples for fingerprinting. Shuo Wang 0012, Alsharif Abuadbba, Sidharth Agarwal, Kristen Moore, Ruoxi Sun 0001, Minhui Xue 0001, Surya Nepal, Seyit Ahmet Çamtepe, Salil S. Kanhere |
SP | 5 |
| 2023 | Not Seen, Not Heard in the Digital World! Measuring Privacy Practices in Children's AppsabstractThe digital age has brought a world of opportunity to children. Connectivity can be a game-changer for some of the world’s most marginalized children. However, while legislatures around the world have enacted regulations to protect children’s online privacy, and app stores have instituted various protections, privacy in mobile apps remains a growing concern for parents and wider society. In this paper, we explore the potential privacy issues and threats that exist in these apps. We investigate 20195 mobile apps from the Google Play store that are designed particularly for children (Family apps) or include children in their target user groups (Normal apps). Using both static and dynamic analysis, we find that 4.47% of Family apps request location permissions, even though collecting location information from children is forbidden by the Play store, and 81.25% of Family apps use trackers (which are not allowed in children’s apps). Even major developers with 40+ kids apps on the Play store use ad trackers. Furthermore, we find that most permission request notifications are not well designed for children, and 19.25% apps have inconsistent content age ratings across the different protection authorities. Our findings suggest that, despite significant attention to children’s privacy, a large gap between regulatory provisions, app store policies, and actual development practices exist. Our research sheds light for government policymakers, app stores, and developers. Ruoxi Sun 0001, Minhui Xue 0001, Gareth Tyson, Shuo Wang 0012, Seyit Ahmet Çamtepe, Surya Nepal |
WWW | 1 |
| 2023 | Data Hiding With Deep Learning: A Survey Unifying Digital Watermarking and SteganographyabstractThe advancement of secure communication and identity verification fields has significantly increased through the use of deep learning techniques for data hiding. By embedding information into a noise-tolerant signal, such as audio, video, or images, digital watermarking and steganography techniques can be used to protect sensitive intellectual property (IP) and enable confidential communication, ensuring that the information embedded is only accessible to authorized parties. This survey provides an overview of recent developments in deep learning techniques deployed for data hiding, categorized systematically according to model architectures and noise injection methods. In addition, potential future research directions that unite digital watermarking and steganography on software engineering to enhance security and mitigate risks are suggested and deliberated. This contribution furthers the creation of a more trustworthy digital world and advances responsible artificial intelligence (AI). Olivia Byrnes, Hu Wang 0005, Ruoxi Sun 0001, Congbo Ma, Huaming Chen, Qi Wu 0001, Minhui Xue 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2022 | Path Transitions Tell More: Optimizing Fuzzing Schedules via Runtime Program StatesabstractCoverage-guided Greybox Fuzzing (CGF) is one of the most successful and widely-used techniques for bug hunting. Two major approaches are adopted to optimize CGF: (i) to reduce search space of inputs by inferring relationships between input bytes and path constraints; (ii) to formulate fuzzing processes (e.g., path transitions) and build up probability distributions to optimize power schedules, i.e., the number of inputs generated per seed. However, the former is subjective to the inference results which may include extra bytes for a path constraint, thereby limiting the efficiency of path constraints resolution, code coverage discovery, and bugs exposure; the latter formalization, concentrating on power schedules for seeds alone, is inattentive to the schedule for bytes in a seed. Xi Xiao 0001, Xiaogang Zhu 0001, Ruoxi Sun 0001, Minhui Xue 0001, Sheng Wen |
ICSE | 4 |
| 2022 | M$^4$I: Multi-modal Models Membership InferenceabstractWith the development of machine learning techniques, the attention of research has been moved from single-modal learning to multi-modal learning, as real-world data exist in the form of different modalities. However, multi-modal models often carry more information than single-modal models and they are usually applied in sensitive scenarios, such as medical report generation or disease identification. Compared with the existing membership inference against machine learning classifiers, we focus on the problem that the input and output of the multi-modal models are in different modalities, such as image captioning. This work studies the privacy leakage of multi-modal models through the lens of membership inference attack, a process of determining whether a data record involves in the model training process or not. To achieve this, we propose Multi-modal Models Membership Inference (M$^4$I) with two attack methods to infer the membership status, named metric-based (MB) M$^4$I and feature-based (FB) M$^4$I, respectively. More specifically, MB M$^4$I adopts similarity metrics while attacking to infer target data membership. FB M$^4$I uses a pre-trained shadow multi-modal feature extractor to achieve the purpose of data inference attack by comparing the similarities from extracted input and output features. Extensive experimental results show that both attack methods can achieve strong performances. Respectively, 72.5% and 94.83% of attack success rates on average can be obtained under unrestricted scenarios. Moreover, we evaluate multiple defense mechanisms against our attacks. The source code of M$^4$I attacks is publicly available at https://github.com/MultimodalMI/Multimodal-membership-inference.git. Pingyi Hu, Ruoxi Sun 0001, Hu Wang 0005, Minhui Xue 0001 |
NeurIPS | 3 |
| 2022 | Cross-language Android permission specificationabstractThe Android system manages access to sensitive APIs by permission enforcement. An application (app) must declare proper permissions before invoking specific Android APIs. However, there is no official documentation providing the complete list of permission-protected APIs and the corresponding permissions to date. Researchers have spent significant efforts extracting such API protection mapping from the Android API framework, which leverages static code analysis to determine if specific permissions are required before accessing an API. Nevertheless, none of them has attempted to analyze the protection mapping in the native library (i.e., code written in C and C++), an essential component of the Android framework that handles communication with the lower-level hardware, such as cameras and sensors. While the protection mapping can be utilized to detect various security vulnerabilities in Android apps, such as permission over-privilege, imprecise mapping will lead to false results in detecting such security vulnerabilities. To fill this gap, we thereby propose to construct the protection mapping involved in the native libraries of the Android framework to present a complete and accurate specification of Android API protection. We develop a prototype system, named NatiDroid, to facilitate the cross-language static analysis and compare its performance with two state-of-the-practice tools, termed Axplorer and Arcade. We evaluate NatiDroid on more than 11,000 Android apps, including system apps from custom Android ROMs and third-party apps from the Google Play. Our NatiDroid can identify up to 464 new API-permission mappings, in contrast to the worst-case results derived from both Axplorer and Arcade, where approximately 71% apps have at least one false positive in permission over-privilege. We have disclosed all the potential vulnerabilities detected to the stakeholders. Xiao Chen 0002, Ruoxi Sun 0001, Minhui Xue 0001, Sheng Wen, M. Ejaz Ahmed, Seyit Ahmet Çamtepe, Yang Xiang 0001 |
ESEC/SIGSOFT FSE | 3 |
| 2021 | Snipuzz: Black-box Fuzzing of IoT Firmware via Message Snippet InferenceabstractThe proliferation of Internet of Things (IoT) devices has made people's lives more convenient, but it has also raised many security concerns. Due to the difficulty of obtaining and emulating IoT firmware, in the absence of internal execution information, black-box fuzzing of IoT devices has become a viable option. However, existing black-box fuzzers cannot form effective mutation optimization mechanisms to guide their testing processes, mainly due to the lack of feedback. In addition, because of the prevalent use of various and non-standard communication message formats in IoT devices, it is difficult or even impossible to apply existing grammar-based fuzzing strategies. Therefore, an efficient fuzzing approach with syntax inference is required in the IoT fuzzing domain. Xiaotao Feng, Ruoxi Sun 0001, Xiaogang Zhu 0001, Minhui Xue 0001, Sheng Wen, Dongxi Liu, Surya Nepal, Yang Xiang 0001 |
CCS | 2 |
| 2021 | An Empirical Assessment of Global COVID-19 Contact Tracing ApplicationsabstractThe rapid spread of COVID-19 has made manual contact tracing difficult. Thus, various public health authorities have experimented with automatic contact tracing using mobile applications (or "apps"). These apps, however, have raised security and privacy concerns. In this paper, we propose an automated security and privacy assessment tool - COVIDGUARDIAN - which combines identification and analysis of Personal Identification Information (PII), static program analysis and data flow analysis, to determine security and privacy weaknesses. Furthermore, in light of our findings, we undertake a user study to investigate concerns regarding contact tracing apps. We hope that COVIDGUARDIAN, and the issues raised through responsible disclosure to vendors, can contribute to the safe deployment of mobile contact tracing. As part of this, we offer concrete guidelines, and highlight gaps between user requirements and app performance. Ruoxi Sun 0001, Wei Wang 0334, Minhui Xue 0001, Gareth Tyson, Seyit Ahmet Çamtepe, Damith Chinthana Ranasinghe |
ICSE | 1 |
| 2020 | Quality Assessment of Online Automated Privacy Policy Generators: An Empirical StudyabstractOnline Automated Privacy Policy Generators (APPGs) are tools used by app developers to quickly create app privacy policies which are required by privacy regulations to be incorporated to each mobile app. The creation of these tools brings convenience to app developers; however, the quality of these tools puts developers and stakeholders at legal risk. In this paper, we conduct an empirical study to assess the quality of online APPGs. We analyze the completeness of privacy policies, determine what categories and items should be covered in a complete privacy policy, and conduct APPG assessment with boilerplate apps. The results of assessment show that due to the lack of static or dynamic analysis of app's behavior, developers may encounter two types of issues caused by APPGs. First, the generated policies could be incomplete because they do not cover all the essential items required by a privacy policy. Second, some generated privacy policies contain unnecessary personal information collection or arbitrary commitments inconsistent with user input. Ultimately, the defects of APPGs may potentially lead to serious legal issues. We hope that the results and insights developed in this paper can motivate the healthy and ethical development of APPGs towards generating a more complete, accurate, and robust privacy policy. Ruoxi Sun 0001, Minhui Xue 0001 |
EASE | 1 |
| 2020 | An Automated Assessment of Android ClipboardsabstractSince the new privacy feature in iOS enabling users to acknowledge which app is reading or writing to his or her clipboard through prompting notifications was updated, a plethora of top apps have been reported to frequently access the clipboard without user consent. However, the lack of monitoring and control of Android application's access to the clipboard data leave Android users blind to their potential to leak private information from Android clipboards, raising severe security and privacy concerns. In this preliminary work, we envisage and investigate an approach to (i) dynamically detect clipboard access behaviour, and (ii) determine privacy leaks via static data flow analysis, in which we enhance the results of taint analysis with call graph concatenation to enable leakage source backtracking. Our preliminary results indicate that the proposed method can expose clipboard data leakage as substantiated by our discovery of a popular app, i.e., Sogou Input, directly monitoring and transferring user data in a clipboard to backend servers. Wei Wang 0334, Ruoxi Sun 0001, Minhui Xue 0001, Damith Chinthana Ranasinghe |
ASE | 2 |
| 2020 | VenueTrace: a privacy-by-design COVID-19 digital contact tracing solution: poster abstractabstractRapid spread of the COVID-19 pandemic is making traditional manual contact tracing challenging; in response, digital contact tracing mobile apps have been developed by the software industry and promoted by governments and health authorities worldwide. However, deploying contact tracing apps across a population at scale have raised many privacy concerns. In this paper, we propose a venue-access-based contact tracing solution, VenueTrace, which preserves user privacy by designs by: (i) enabling the contact tracing of venue-to-user, instead of user-to-user; (ii) avoiding information exchanges between users; and (iii) ensuring no private data is exposed to back-end servers, while enabling proximity contact tracing. Ruoxi Sun 0001, Wei Wang 0334, Minhui Xue 0001, Gareth Tyson, Damith Chinthana Ranasinghe |
SenSys | 1 |