Lei Pan 0002

dblp:33/1366-2 · DBLP profile ↗
← Back
81ranked-venue papers
4as first author
54since 2021 · last 2026
0000-0002-4691-8330ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 29 · 2 first-author · 19 since 2021Artificial intelligence and machine learning · 17 · 11 since 2021Systems, architecture and hardware · 10 · 2 first-author · 6 since 2021Computer networks · 8 · 7 since 2021Databases, data management, data science and information retrieval · 8 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Software engineering, systems software and programming languages · 4 · 3 since 2021
YearPublicationVenuePosition
2026 A survey of image encryption schemes: Arnold transformation, chaos, bit-plane extraction and permutation based algorithms
abstract
Abstract Securing digital images captured by unmanned aerial vehicles (UAVs) is important for maintaining data confidentiality and integrity during transmission over insecure networks. This study surveys and evaluates existing encryption schemes such as Arnold transformation, chaos-based, bit-plane extraction, quantum, and permutation-based algorithms. The existing image encryption algorithms are implemented and tested in MATLAB 2015 using standard benchmark images (Quantum, Baboon, and Cameraman), selected for their frequent use and benchmark relevance in recent image security literature. For algorithms whose statistical indices are already documented in prior research, those published values are adopted for reference. In cases where such data were unavailable, the corresponding schemes were re-implemented and experimentally evaluated in MATLAB 2015 to produce consistent and reproducible performance results. The comparative statistical analysis across these datasets demonstrates that hybrid quantum–chaotic and permutation–diffusion methods achieve near-ideal entropy values ( $$\approx $$ 7.999), high NPCR ( $$\approx $$ 99.6%), and UACI ( $$\approx $$ 33.4%). This indicates strong resistance to statistical and differential attacks. These schemes also exhibit low correlation coefficients (< 0.002) and large key spaces (> $$2^{100}$$ ).
Samina Jadoon, Lei Pan 0002, Md. Shamsul Huda, Kashif Hesham Khan
Multim. Tools Appl.2
2026 Social Equity and Inclusion With Fair Dataset Representation of Diverse Populations in Diffusion Models
abstract
Generative artificial intelligence (AI), an advancing frontier, uses machine learning to autonomously create media content, though it faces challenges with bias and fair representation. This research investigates demographic bias within the LAION dataset, specifically focusing on the representation of Indigenous Australians. Large image generative models are typically trained on extensive datasets scraped from the Internet, which often reflect the demographics of the most active online communities rather than accurately representing local populations. While existing studies have examined hate speech, gender balance, and broad ethnic diversity within LAION, limited research addresses representation relative to specific national demographics or minority subgroups. Auditing the dataset with state-of-the-art facial recognition and sentiment analysis models, we assessed the age, gender, and ethnicity of images in LAION and compared these distributions against census data for Australia and the United States. We also conducted a focused analysis of Indigenous Australian representation by identifying relevant images using keyword searches and applying the aforementioned models. Our findings reveal that the dataset demographic composition poorly aligns with the actual population of Australia, with regional subgroups under-represented, potentially leading to inaccurate portrayals of Indigenous Australians. These results underscore the need for either strengthened safeguards on globally developed generative models or the development of locally trained models to ensure responsible and inclusive cultural representation in AI-generated imagery.
Ryan Holland, Saifur Rahman 0002, Shantanu Pal, Chandan K. Karmakar, Lei Pan 0002
IEEE Trans. Comput. Soc. Syst.5
2026 Privacy-Preserving Automated Deep Learning for Secure Inference Service
abstract
Automated deep learning (AutoDL) aims to automatically discover optimal architectures of deep neural networks (DNNs) for secure inference without the studies for time-consuming and error-prone manual design. Privacy concerns have increasingly motivated the studies for privacy-preserving AutoDL (PrivAutoDL), where DNN architectures are searched directly on encrypted data without revealing the client's confidential inputs and well-trained DNN architectures. However, existing studies encounter problems in achieving a balance between provable security and efficiency while avoiding significant degradation of model utility. To tackle these problems, we design a privacy-preserving AutoDL scheme, named 2PCAutoDL, utilizing a two-party (two non-colluding cloud servers) computation model. Based on the two-server model, efficient and secure computation protocols are customized layer by layer to protect DNN models associated with client's data. In particular, we reduce the computational overhead of secure DNN: our optimized protocols achieve$1.34\times \sim 2.05\times$speedup for linear layers and$1.33 \times \sim 45 \times$speedup for non-linear layers, compared to a range of existing secure implementations in the literature. Moreover, our fresh alternative to approximate Softmax avoids the drawbacks of approximating exponential operation and yields slightly higher accuracy under appropriate configurations. The security of 2PCAutoDL is formally analyzed under the semi-honest adversary model. Extensive experiments demonstrate that the searched models from 2PCAutoDL improve the inference accuracy by 0.6% on MNIST and by 0.5% on CIFAR-10 when compared to state-of-the-art (SOTA) PrivAutoDL.
Fuyi Wang, Jinzhi Ouyang, Leo Yu Zhang, Lei Pan 0002, Shengshan Hu, Xiaoning Liu 0002, Robin Doss
IEEE Trans. Dependable Secur. Comput.4
2025 sf SEBioID: Secure and Efficient Biometric Identification with Two-Party Computation
Fuyi Wang, Jinzhi Ouyang, Leo Yu Zhang, Lei Pan 0002, Shengshan Hu, Robin Doss, Jianying Zhou 0001
ACNS (3)4
2025 uBSaaS: A Unified Blockchain Service as a Service Framework for Streamlined Blockchain Services Integration
Huynh Thanh Thien Pham, Frank Jiang 0001, Lei Pan 0002, Alessio Bonti, Mohamed Almorsy
ENASE3
2025 PrivGNN: High-Performance Secure Inference for Cryptographic Graph Neural Networks
Fuyi Wang, Zekai Chen 0010, Mingyuan Fan 0003, Jianying Zhou 0001, Lei Pan 0002, Leo Yu Zhang
FC (2)5
2025 Semantic Information Extraction with Language Models for Zero-Day Attack Detection
Shyamali Sinali Karunarathne, Sutharshan Rajasegarar, Lei Pan 0002
KSEM (5)3
2025 Attack-data independent defence mechanism against adversarial attacks on ECG signal
abstract
Adversarial attacks pose a significant threat to the integrity and reliability of electrocardiogram (ECG) signals, compromising their use in critical applications, e.g., arrhythmia detection and classification. In this paper, we propose an attack-data-independent defence mechanism to effectively mitigate adversarial attacks on ECG signals. Unlike existing defence mechanisms that rely on learning from adversarial samples, our proposed approach operates as a ‘gatekeeper,’ selectively discarding noisy and attack signals while allowing only clean and non-attack ECG signals to be stored in the data layer. This ensures the availability of reliable and high-quality ECG data for subsequent analysis. The proposed defence mechanism not only detects and filters out the attack and noisy ECG signals but also provides robust protection against adversarial attacks, enhancing the integrity and trustworthiness of ECG data for critical applications. To evaluate the effectiveness of our proposal, we conduct experiments using physiologic and synthetic ECG datasets against two well-known attacks: a white-box attack (Fast Gradient Signed Method (FGSM) and Projected Gradient Descent (PGD)) and a black-box attack (HopSkipJump and Boundary). Our experimental results demonstrate the superiority and effectiveness of our approach in defending against adversarial attacks on ECG signals, making it a promising solution for ensuring the security and reliability of ECG-based diagnosis in smart healthcare applications.
Saifur Rahman 0002, Shantanu Pal, Ahsan Habib 0003, Lei Pan 0002, Chandan K. Karmakar
Comput. Networks4
2025 A survey of coverage-guided greybox fuzzing with deep neural models
abstract
Coverage-guided greybox fuzzing (CGF) has emerged as a powerful technique for software vulnerability detection, yet traditional techniques often struggle with the increasing complexity of modern software systems and the vastness of input spaces. Deep neural networks (DNNs) have begun to fundamentally transform CGF by addressing these limitations through automated feature extraction, adaptive input generation, and intelligent path prioritization. However, despite these advancements, critical gaps persist in understanding the state-of-the-art landscape. Existing studies often lack rigorous benchmarks to evaluate scalability and generalizability, fail to address the interpretability of neural-guided decisions, and overlook the integration of emerging paradigms such as large language models (LLMs) and neurosymbolic reasoning. This survey systematically bridges these gaps by providing a comprehensive taxonomy of DNN-driven CGF techniques, analyzing their strengths and limitations across key fuzzing stages—seed generation, selection, and mutation. We find that although DNNs have significantly improved fuzzing efficiency, challenges such as semantically invalid seeds, high computational overhead, and limited cross-domain adaptability remain unresolved. Most importantly, we identify two transformative directions with the potential to redefine CGF: (1) LLM-powered fuzzing , which combines generative AI with domain-specific fine-tuning to produce context-aware inputs; and (2) neurosymbolic integration , which merges the precision of symbolic execution with the scalability of neural networks to tackle path explosion. By synthesizing these insights, this survey not only clarifies the state-of-the-art but also outlines a roadmap for developing robust, explainable, and widely applicable intelligent fuzzers. The future of CGF lies in hybrid models that integrate data-driven learning with formal methods, paving the way for autonomous vulnerability discovery in an era of increasingly complex software systems.
Junyang Qiu, Yupeng Jiang 0002, Yuantian Miao, Wei Luo 0001, Lei Pan 0002, James Xi Zheng
Inf. Softw. Technol.5
2025 MedShield: A Fast Cryptographic Framework for Private Multi-Service Medical Diagnosis
abstract
The substantial progress in privacy-preserving machine learning (PPML) facilitates outsourced medical computer-aided diagnosis (MedCADx) services. However, existing PPML frameworks primarily concentrate on enhancing the efficiency of prediction services, without exploration into diverse medical services such as medical segmentation. In this paper, we proposeMedShield, a pioneering cryptographic framework for diverse MedCADx services (i.e., multi-service, including medical imaging prediction and segmentation). Based on a client-server (two-party) setting,MedShieldefficiently protects medical records and neural network models without fully outsourcing. To execute multi-service securely and efficiently, our technical contributions include: 1) optimizing computational complexity of matrix multiplications for linear layers at the expense of free additions/subtractions; 2) introducing a secure most significant bit protocol with crypto-friendly activations to enhance the efficiency of non-linear layers; 3) presenting a novel layer for upscaling low-resolution feature maps to support multi-service scenarios in practical MedCADx. We conduct a rigorous security analysis and extensive evaluations on benchmarks (MNIST and CIFAR-10) and real medical records (breast cancer, liver disease, COVID-19, and bladder cancer) for various services. Experimental results demonstrate thatMedShieldachieves up to$2.4\times$,$4.3\times$, and$2\times$speed up for MNIST, CIFAR-10, and medical datasets, respectively, compared with prior work when conducting prediction services. For segmentation services,MedShieldpreserves the precision of the unprotected version, showing a$1.23\%$accuracy improvement.
Fuyi Wang, Jinzhi Ouyang, Xiaoning Liu 0002, Lei Pan 0002, Leo Yu Zhang, Robin Doss
IEEE Trans. Serv. Comput.4
2024 Towards Availability of Strong Authentication in Remote and Disruption-Prone Operational Technology Environments
abstract
Implementing strong authentication methods in a network requires stable connectivity between the service providers deployed within the network (i.e., applications that users of the network need to access) and the Identity and Access Management (IAM) server located at the core segment of the network. This becomes challenging when it comes to Operational Technology (OT) systems deployed in a remote area, as they often get disconnected from the core segment of the network owing to unavoidable network disruptions. As a result, weak authentication methods and shared credential approaches are still adopted in these OT environments, exposing system vulnerabilities to increasingly sophisticated cyber threats. In this work, we propose a solution to enable highly available multi-factor authentication (MFA) services for OT environments. The proposed solution is based on Proof-of-Possession (PoP) tokens generated by an IAM server for registered users. The tokens are securely linked to user-specific parameters (e.g., physical security keys, biometrics, PIN, etc.), enabling strong user authentication (during disconnection time) through token validation. We deployed the Tamarin Prover software-based toolkit to verify security of the proposed authentication scheme. For performance evaluation, we implemented the designed solution in real-world settings. The results of our analysis and experiments confirm the efficacy of the proposed solution.
Mohammad Reza Nosouhi, Zubair A. Baig, Robin Doss, Divyans Mahansaria, Debi Prasad Pati, Praveen Gauravaram, Lei Pan 0002, Keshav Sood
ARES7
2024 A Data-Encoding Approach to Quantum Federated Learning: Experimenting with Cloud Challenges
abstract
A Data-Encoding Approach to Quantum Federated Learning: Experimenting with Cloud Challenges
Shiva Raj Pokhrel, Naman Yash, Jonathan Kua, Gang Li 0009, Lei Pan 0002
APNet5
2024 CryptGraph: An Efficient Privacy-Enhancing Solution for Accurate Shortest Path Retrieval in Cloud Environments
abstract
With the widespread adoption of cloud computing, it is a popular trend to migrate shortest path and distance (SPD) retrieval on large-scale graphs to cloud environments, harnessing their immense computational capabilities. To protect sensitive information, these graphs are usually encrypted before being outsourced to the cloud. A significant challenge is how to answer SPD retrieval in a secure, efficient, and accurate manner. However, recent works have yet to concurrently tackle all three aspects to meet this challenge. To address this challenge, we design, implement, and evaluate Crypt-Graph, the first scheme simultaneously allowing private, efficient, and accurate retrieval over encrypted graphs. CryptGraph leverages additive homomorphic encryptions to protect graphs and client information. A series of secure protocols are tailored based on the two-cloud (i.e., server) model. Supported by these protocols, Crypt-Graph converts SPD retrieval from the ciphertext domain to both the plaintext (for vertices) and secret-sharing (for weights) domains, achieving access pattern protection and remarkable efficiency close to plain retrieval. The security of CryptGraph is formally analyzed under the semi-honest adversary model. Extensive experiments are conducted on both synthetic and real-world graph datasets, demonstrating millisecond-level efficiency and 100% accuracy rates.
Fuyi Wang, Zekai Chen 0010, Lei Pan 0002, Leo Yu Zhang, Jianying Zhou 0001
AsiaCCS3
2024 TrustMIS: Trust-Enhanced Inference Framework for Medical Image Segmentation
abstract
Recent advancements in privacy-preserving deep learning (PPDL) enable artificial intelligence-assisted (AI-assisted) medical image diagnostics with privacy guarantees, addressing increasing concerns about data and model privacy. However, intensive studies are restricted to shallow and narrow neural networks (NNs) for simple service (e.g., disease prediction), leaving a gap in exploring diverse inferences. This paper proposes TrustMIS, a trust-enhanced inference framework for fast and private medical image segmentation (MIS) and prediction services. Based on two-party computation, TrustMIS introduces lightweight additive secret-sharing tools to safeguard medical records and NNs. Complementing existing PPDL schemes, we present a series of secure two-party interactive protocols for linear layers. Specifically, we optimize the secure matrix multiplication by reducing the number of expensive multiplication operations with the help of free-computation addition operations to enhance efficiency (bringing 1.15× ∼2.64× savings in both time and communication costs). Furthermore, we customize a fresh secure transposed convolutional protocol for MIS-oriented NNs. A thorough theoretical analysis is provided to prove TrustMIS’s correctness and security. We conduct experimental evaluations over two benchmark and four real-world medical datasets and compare them to state-of-the-art studies. The results demonstrate TrustMIS’s superiority in efficiency and accuracy, improved by 1.1× ∼ 54.4× speedup in secure disease prediction, and 5.56% ↑ ∼ 11.7% ↑ accuracy in secure MIS.
Fuyi Wang, Jinzhi Ouyang, Lei Pan 0002, Leo Yu Zhang, Xiaoning Liu 0002, Robin Doss
ECAI3
2024 ConDGAD: Multi-augmentation Contrastive Learning for Dynamic Graph Anomaly Detection
Siqi Xia, Sutharshan Rajasegarar, Lei Pan 0002, Christopher Leckie, Sarah M. Erfani, Jeffrey Chan
ICPR (25)3
2024 VulMatch: Binary-Level Vulnerability Detection Through Signature
Zian Liu, Shigang Liu, Lei Pan 0002, Chao Chen 0015, Jun Zhang 0010, Dongxi Liu
NSS3
2024 The Value of Strong Identity and Access Management for ICS/OT Security
abstract
As the integration of digital technologies with Industrial Control Systems (ICS) and Operational Technology (OT) continues to deepen, these systems increasingly become targets for sophisticated cyber attacks. These attacks not only threaten the operational integrity but also pose significant risks to national security and public safety. In this paper, we provide insights into the value of ICS/OT security solutions that are based on Identity and Access Management (IAM). Beginning with presenting an abstraction model for typical ICS/OT attacks, the paper systematically outlines the main stages of an attack and the corresponding vectors employed by adversaries. Drawing from the MITRE ATT&CK framework tailored for ICS, the paper quantifies the extent to which IAM-based mitigation approaches can strengthen defense-in-depth mechanisms against cyber threats targeting ICS/OT environments. Our findings show that there are modern attack vectors that can only be mitigated through robust IAM solutions. Moreover, we found that while advanced techniques such as firewall and gateway-based intelligent threat detection play a significant role in safeguarding I CS/OT, they are insufficient on their own to address several attack vectors in ICS/OT environments.
Mohammad Reza Nosouhi, Zubair A. Baig, Robin Doss, Praveen Gauravaram, Debi Prasad Pati, Divyans Mahansaria, Keshav Sood, Lei Pan 0002
PST8
2023 EnSpeciVAT: Enhanced SpecieVAT for Cluster Tendency Identification in Graphs
Siqi Xia, Sutharshan Rajasegarar, Christopher Leckie, Sarah M. Erfani, Jeffrey Chan, Lei Pan 0002
ADMA (3)6
2023 Hiding Your Signals: A Security Analysis of PPG-Based Biometric Authentication
Lin Li 0066, Chao Chen 0015, Lei Pan 0002, Yonghang Tai, Jun Zhang 0010, Yang Xiang 0001
ESORICS (3)3
2023 Quantum Autoencoder Frameworks for Network Anomaly Detection
Moe Hdaib, Sutharshan Rajasegarar, Lei Pan 0002
ICONIP (5)3
2023 SigD: A Cross-Session Dataset for PPG-based User Authentication in Different Demographic Groups
abstract
Recently, unobservable physiological signals have received widespread attention from researchers as unique identifiers of users in biometrics. However, due to the lack of data sets, existing methods are limited in evaluating cross-session scenarios. Cross-session means that signals are collected at different sessions (times). In real scenarios, authentication is almost always cross-session. Currently, the datasets commonly used for Photoplethysmogram (PPG) signal authentication span around one month, which is insufficient for authentication. On the other hand, different demographic groups have different hemodynamic characteristics, but existing methods lack an assessment of these aspects. This paper introduces a dataset to provide insights into PPG signal-based authentication across different time spans and user groups (age, gender). As physiological signals offer unique advantages for user authentication, the potential of PPG signals is gradually explored. Furthermore, our comparative analysis of recent publications on data-driven user authentication using PPG can further identify the similarities and differences among the performance of the proposed authentication models. Our findings may help future research towards a consensus on an appropriate set of performance metrics.
Lin Li 0066, Chao Chen 0015, Lei Pan 0002, Jun Zhang 0010, Yang Xiang 0001
IJCNN3
2023 SigA: rPPG-based Authentication for Virtual Reality Head-mounted Display
abstract
Consumer-grade virtual reality head-mounted displays (VR-HMD) are becoming increasingly popular. Despite VR’s convenience and booming applications, VR-based authentication schemes are underdeveloped. The recently proposed authentication methods (Electrooculogram based, Electrical Muscle Stimulation-based, and alike) require active user involvement, disturbing many scenarios like drone flight and telemedicine. This paper proposes an effective and efficient user authentication method in VR environments resilient to impersonation attacks using physiological signals — Photoplethysmogram (PPG), namely SigA. SigA exploits the advantage that PPG is a physiological signal invisible to the naked eye. Using VR-HMDs to cover the eye area completely, SigA reduces the risk of signal leakage during PPG acquisition. We conducted a comprehensive analysis of SigA’s feasibility on five publicly available datasets, nine different pre-trained models, three facial regions, various lengths of the video clips required for training, four different signal time intervals, and continuous authentication with different sliding window sizes. The results demonstrate that SigA achieves more than 95% of the average F1-score in a one-second signal to accommodate a complete cardiac cycle for most adults, implying its applicability in real-world scenarios. Furthermore, experiments have shown that SigA is resistant to zero-effort attacks, statistical attacks, impersonation attacks (with a detection accuracy of over 95%) and session hijacking attacks.
Lin Li 0066, Chao Chen 0015, Lei Pan 0002, Leo Yu Zhang, Jun Zhang 0010, Yang Xiang 0001
RAID3
2023 Novel entanglement compression for QKD protocols using isometric tensors
Hong Lai, Josef Pieprzyk, Lei Pan 0002
Sci. China Inf. Sci.3
2023 Security and privacy problems in voice assistant applications: A survey
abstract
Voice assistant applications have become omniscient nowadays. Two models that provide the two most important functions for real-life applications (i.e., Google Home, Amazon Alexa, Siri, etc.) are Automatic Speech Recognition (ASR) models and Speaker Identification (SI) models. According to recent studies, security and privacy threats have also emerged with the rapid development of the Internet of Things (IoT). The security issues researched include attack techniques toward machine learning models and other hardware components widely used in voice assistant applications. The privacy issues include technical-wise information stealing and policy-wise privacy breaches. The voice assistant application takes a steadily growing market share every year, but their privacy and security issues never stopped causing huge economic losses and endangering users' personal sensitive information. Thus, it is important to have a comprehensive survey to outline the categorization of the current research regarding the security and privacy problems of voice assistant applications. This paper concludes and assesses five kinds of security attacks and three types of privacy threats in the papers published in the top-tier conferences of cyber security and voice domain.
Jingjin Li, Chao Chen 0015, Mostafa Rahimi Azghadi, Hossein Ghodosi, Lei Pan 0002, Jun Zhang 0010
Comput. Secur.5
2023 A Survey of PPG's Application in Authentication
abstract
Biometric authentication prospered because of its convenient use and security. Early generations of biometric mechanisms suffer from spoofing attacks. Recently, unobservable physiological signals (e.g., Electroencephalogram, Photoplethysmogram, Electrocardiogram) as biometrics offer a potential remedy to this problem. In particular, Photoplethysmogram (PPG) measures the change in blood flow of the human body by an optical method. Clinically, researchers commonly use PPG signals to obtain patients' blood oxygen saturation, heart rate, and other information to assist in diagnosing heart-related diseases. Since PPG signals contain a wealth of individual cardiac information, researchers have begun to explore their potential in cyber security applications. The unique advantages (simple acquisition, difficult to steal, and live detection) of the PPG signal allow it to improve the security and usability of the authentication in various aspects. However, the research on PPG-based authentication is still in its infancy. The lack of systematization hinders new research in this field. We conduct a comprehensive study of PPG-based authentication and discuss these applications' limitations before pointing out future research directions.
Lin Li 0066, Chao Chen 0015, Lei Pan 0002, Leo Yu Zhang, Jun Zhang 0010, Yang Xiang 0001
Comput. Secur.3
2023 Trustworthy Sensor Fusion Against Inaudible Command Attacks in Advanced Driver-Assistance Systems
abstract
There are increasing concerns about malicious attacks on autonomous vehicles. In particular, inaudible voice command attacks pose a significant threat as voice commands become available in autonomous driving systems. How to empirically defend against these inaudible attacks remains an open question. Previous research investigates utilizing deep learning-based multimodal fusion for defense, without considering the model uncertainty in trustworthiness. As deep learning has been applied to increasingly sensitive tasks, uncertainty measurement is crucial in helping improve model robustness, especially in mission-critical scenarios. In this article, we propose the multimodal fusion framework (MFF) as an intelligent security system to defend against inaudible voice command attacks. MFF fuses heterogeneous audio–vision modalities using VGG family neural networks and achieves the detection accuracy of 92.25% in the comparative fusion method empirical study. Additionally, extensive experiments on audio–vision tasks reveal the model’s uncertainty. Using expected calibration errors, we measure calibration errors and Monte Carlo Dropout to estimate the predictive distribution for the proposed models. Our findings show empirically to train robust multimodal models, improve standard accuracy and provide a further step toward interpretability. Finally, we discuss the pros and cons of our approach and its applicability for advanced driver assistance systems.
Jiwei Guan, Lei Pan 0002, Chen Wang 0008, Shui Yu 0001, Longxiang Gao, James Xi Zheng
IEEE Internet Things J.2
2023 A Multi-Objective Clustering Evolutionary Algorithm for Multi-Workflow Computation Offloading in Mobile Edge Computing
abstract
To cope with the rapid development of the Internet of Things (IoT) and the increasing demand for real-time services, mobile edge computing (MEC) has become a promising solution which extends centralised cloud computing, to provision computing resources, storage and network services closer to the mobile device from the network edge. While computation offloading is a key feature in MEC to enable real-time services, offloading workflow tasks in MEC is an NP-hard problem. Typically, the problem of multi-workflow offloading with multi-objective optimization is still an open and challenging issue. Therefore, this article proposes a multi-objective clustering evolutionary algorithm called MCEA to minimize the cost and energy consumption of multi-workflow execution under the deadline constraint. First, the sub-deadline constraint is added during initialization to generate more initial solutions that satisfy the deadline constraint. Then an adaptive clustering method is adopted to guide individuals to find a suitable mate during crossover operation. Finally, the probabilities of crossover and mutation are dynamically adjusted based on the historical information to control the evolution direction and convergence speed of algorithm. Comprehensive experiments are carried out for complex workflow applications on FogWorkflowSim, which demonstrate that MCEA can achieve better performance than four representative algorithms in three evaluation metrics.
Lei Pan 0002, Xiao Liu 0004, Jia Xu 0010, Xuejun Li 0001
IEEE Trans. Cloud Comput.1
2023 Cyber Code Intelligence for Android Malware Detection
abstract
Evolving Android malware poses a severe security threat to mobile users, and machine-learning (ML)-based defense techniques attract active research. Due to the lack of knowledge, many zero-day families’ malware may remain undetected until the classifier gains specialized knowledge. The most existing ML-based methods will take a long time to learn new malware families in the latest malware family landscape. Existing ML-based Android malware detection and classification methods struggle with the fast evolution of the malware landscape, particularly in terms of the emergence of zero-day malware families and limited representation of single-view features. In this article, a new multiview feature intelligence (MFI) framework is developed to learn the representation of a targeted capability from known malware families for recognizing unknown and evolving malware with the same capability. The new framework performs reverse engineering to extract multiview heterogeneous features, including semantic string features, API call graph features, and smali opcode sequential features. It can learn the representation of a targeted capability from known malware families through a series of processes of feature analysis, selection, aggregation, and encoding, to detect unknown Android malware with shared target capability. We create a new dataset with ground-truth information regarding capability. Many experiments are conducted on the new dataset to evaluate the performance and effectiveness of the new method. The results demonstrate that the new method outperforms three state-of-the-art methods, including: 1) Drebin; 2) MaMaDroid; and 3)$N$-opcode, when detecting unknown Android malware with targeted capabilities.
Junyang Qiu, Qing-Long Han, Wei Luo 0001, Lei Pan 0002, Surya Nepal, Jun Zhang 0010, Yang Xiang 0001
IEEE Trans. Cybern.4
2023 Weak-Key Analysis for BIKE Post-Quantum Key Encapsulation Mechanism
abstract
The evolution of quantum computers poses a serious threat to contemporary public-key encryption (PKE) schemes. To address this impending issue, the National Institute of Standards and Technology (NIST) is currently undertaking the Post-Quantum Cryptography (PQC) standardization project intending to evaluate and subsequently standardize the suitable PQC scheme(s). One such attractive approach, called Bit Flipping Key Encapsulation (BIKE), has entered the final round of the competition. Despite having some attractive features, the IND-CCA security of BIKE depends on the average decoder failure rate (DFR), a higher value of which can facilitate a particular type of side-channel attack. Although BIKE adopts the Black-Grey-Flip (BGF) decoder that offers a negligible DFR, the effect of weak-keys on the average DFR has not been fully investigated. In this paper, we implement the BIKE scheme, and then through extensive experiments show that the weak-keys can be a potential threat to IND-CCA security of the BIKE scheme and thus need attention from the relevant research community. We also propose a key-check algorithm that can potentially supplement the BIKE mechanism and prevent users from adopting weak-keys.
Mohammad Reza Nosouhi, Syed Wajid Ali Shah, Lei Pan 0002, Yevhen Zolotavkin, Ashish Nanda, Praveen Gauravaram, Robin Doss
IEEE Trans. Inf. Forensics Secur.3
2022 Cyber Attack Detection in IoT Networks with Small Samples: Implementation And Analysis
Venkata Abhishek Kanthuru, Sutharshan Rajasegarar, Punit Rathore, Robin Doss, Lei Pan 0002, Biplob R. Ray, Morshed Chowdhury, Chandrasekaran Srimathi, M. A. Saleem Durai
ADMA (1)5
2022 EvAnGCN: Evolving Graph Deep Neural Network Based Anomaly Detection in Blockchain
Vatsal Patel, Sutharshan Rajasegarar, Lei Pan 0002, Jiajun Liu 0004, Liming Zhu 0001
ADMA (1)3
2022 No-Label User-Level Membership Inference for ASR Model Auditing
Yuantian Miao, Chao Chen 0015, Lei Pan 0002, Shigang Liu, Seyit Ahmet Çamtepe, Jun Zhang 0010, Yang Xiang 0001
ESORICS (2)3
2022 Towards Privacy-Preserving Neural Architecture Search
abstract
Machine learning promotes the continuous development of signal processing in various fields, including network traffic monitoring, EEG classification, face identification, and many more. However, massive user data collected for training deep learning models raises privacy concerns and increases the difficulty of manually adjusting the network structure. To address these issues, we propose a privacy-preserving neural architecture search (PP-NAS) framework based on secure multi-party computation to protect users' data and the model's parameters/hyper-parameters. PP-NAS outsources the NAS task to two non-colluding cloud servers for making full advantage of mixed protocols design. Complement to the existing PP machine learning frameworks, we redesign the secure ReLU and Max-pooling garbled circuits for significantly better efficiency (3 ~ 436 times speed-up). We develop a new alternative to approximate the Softmax function over secret shares, which bypasses the limitation of approximating exponential operations in Softmax while improving accuracy. Extensive analyses and experiments demonstrate PP-NAS's superiority in security, efficiency, and accuracy.
Fuyi Wang, Leo Yu Zhang, Lei Pan 0002, Shengshan Hu, Robin Doss
ISCC3
2022 InstaVarjoLive: An Edge-Assisted 360 Degree Video Live Streaming for Virtual Reality Testbed
abstract
Virtual Reality (VR) challenges us with the requirements of ultra-low latency and ultra-high bandwidth. Existing methods that rely on cloud computing systems to improve the latency and bandwidth problems cannot satisfy the high computation and fast communication requirements in VR. Edge computing has emerged as a promising solution that can be applied in VR to optimize the latency and bandwidth prob-lems. However, another challenge is applying edge computing technology to improve the seamless for the VR users. Based on this, this paper proposes an edge-end collaboration testbed called Insta VarjoLive and conducts the experiments on the real-time 360° video live streaming seamless using VR headsets. We compared our experiments with the 360° videos watched by users from the cloud through VR headset and obtained the results showing that the edge-assisted 360° videos live streaming method has three major advantages: better real-time delivery, lower response time, and higher bandwidth guaranteed. Furthermore, we tested our experiments and discussed the other possible optimization methods in the future.
Feifei Chen 0001, Rui Wang 0008, Thuong N. Hoang, Lei Pan 0002
MSN5
2022 Forward Traceability for Product Authenticity Using Ethereum Smart Contracts
Fokke Heikamp, Lei Pan 0002, Rolando Trujillo-Rasua, Sushmita Ruj, Robin Doss
NSS2
2022 Domain adaptation for Windows advanced persistent threat detection
Rory Coulter, Jun Zhang 0010, Lei Pan 0002, Yang Xiang 0001
Comput. Secur.3
2022 Trustworthy blockchain-based medical Internet of thing for minimal invasive surgery training simulator
abstract
Summary Realistic modeling of mechanical behavior of soft tissue has been recognized as an essential part for medical Internet of thing for minimal invasive surgery (MIS) training simulator. Therefore, the blockchain‐based constitutive model is crucial for mechanical response of soft tissue modeling. In this article, based on the Ogden second order model, a novel hyperplastic model was presented to describe the stress‐stretch relationship in the MIS training system. To validate this theoretical model, two experimental techniques (uniaxial compression and uniaxial tensile) were conducted to obtain data related to stress‐strain in the blockchain system, which plays an important role in investigating the mechanical behavior of soft tissue. Our results show that the new model has a satisfied coincidence of the experimental data than other existing models. Furthermore, the viscoelastic properties of soft tissue were investigated and a viscoelastic model based on three‐parameter was utilized to interpret the viscoelastic behavior of the soft tissue. The contributions of this article include several biomechanical tests that were performed to investigate the soft tissue hyperelastic and viscoelastic properties in the MIS system, and theoretical guidance for simulating soft tissue mechanical behavior in the blockchain‐based simulation system.
Yonghang Tai, Yinjia Wang, Lei Wei 0002, Lei Pan 0002, Jun Zhang 0010, Junsheng Shi
Concurr. Comput. Pract. Exp.6
2022 Towards enhanced PDF maldocs detection with feature engineering: design challenges
abstract
Abstract In this paper, we perform an in-depth analysis of a large corpus of PDF maldocs to identify the key set of significantly important features and help in maldoc detection. Existing industry-based tools for the detection are inefficient and cannot prevent PDF maldocs because they are generic and depend primarily on a signature-based approach. Besides, several other methods developed by academics suffer heavily from reduced effectiveness. The feature-set using machine learning classifiers is prone to various known attacks, such as mimicry and parser confusion. Also, we discover that increasingly more malicious files i) contain evasive and obfuscated JavaScript code, ii) include hidden contents (mostly outside the objects), iii) have a corrupted document structure, and iv) usually contain short JavaScript code blocks. We utilise maldoc attacks’ evolution over a decade to highlight the essential features (e.g., concept drifts) that impact detectors and classifiers.
Ahmed Falah, Shiva Raj Pokhrel, Lei Pan 0002, Anthony de Souza-Daw
Multim. Tools Appl.3
2022 Learning multi-level and multi-scale deep representations for privacy image classification
Yahui Han, Yonggang Huang 0001, Lei Pan 0002, Yunbo Zheng
Multim. Tools Appl.3
2022 JSCSP: A Novel Policy-Based XSS Defense Mechanism for Browsers
abstract
To mitigate cross-site scripting attacks (XSS), the W3C group recommends web service providers to employ a computer security standard called Content Security Policy (CSP). However, less than 3.7 percent of real-world websites are equipped with CSP according to Google’s survey. The low scalability of CSP is incurred by the difficulty of deployment and non-compatibility for state-of-art browsers. To explore the scalability of CSP, in this article, we propose JavaScript based CSP (JSCSP), which is able to support most of real-world browsers but also to generate security policies automatically. Specifically, JSCSP offers a novel self-defined security policy which enforces essential confinements to related items, including JavaScript functions, DOM elements and data access. Meanwhile, JSCSP has an efficient algorithm to automatically generate the policy directives and enforce them in a cascading way, which is more fine-grained and practical than the functionalities provided by CSP. We further implement JSCSP on a Chrome extension, and our evaluation shows that the extension is compatible with popular JavaScript libraries. Our JSCSP extension can detect and block the tested attacking vectors extracted from the prevalent web applications. We state that JSCSP delivers better performance compared to other XSS defense solutions.
Guangquan Xu, Xiaofei Xie, Shuhan Huang, Jun Zhang 0010, Lei Pan 0002, Wei Lou, Kaitai Liang
IEEE Trans. Dependable Secur. Comput.5
2022 Blockchain-Based Audio Watermarking Technique for Multimedia Copyright Protection in Distribution Networks
abstract
Copyright protection in multimedia protection distribution is a challenging problem. To protect multimedia data, many watermarking methods have been proposed in the literature. However, most of them cannot be used effectively in a multimedia distribution network (MDN) as they are not designed to support multi-layer watermark embedding. Multi-layer watermarking mechanisms were developed to protect multimedia data across different layers in an MDN. However, in those mechanisms, we need to trust the entities in the MDN, such as regional and country distributors. To overcome this potential drawback, in this article, we propose a novel privacy protection mechanism for MDNs by combining the advantages of both blockchain and watermarking technologies. A specifically designed watermarking algorithm is used to link the copyright information with the audio file, while a novel blockchain-based smart contract mechanism is developed to enforce the proper functioning of each entity in the distribution network. Moreover, the new audio mechanism is computationally efficient. Although audio signals are used to show the effectiveness of the proposed mechanism, the proposed approach can easily be extended to other multimedia objects, such as an image. The validity of the proposed mechanism is demonstrated by our simulation results. The proposed mechanism can benefit multimedia production companies and other entities in the MDN.
Iynkaran Natgunanathan, Purathani Praitheeshan, Longxiang Gao, Yong Xiang 0001, Lei Pan 0002
ACM Trans. Multim. Comput. Commun. Appl.5
2021 Digital Twin for Cybersecurity: Towards Enhancing Cyber Resilience
Rajiv Faleiro, Lei Pan 0002, Shiva Raj Pokhrel, Robin Doss
BROADNETS2
2021 Containers' Privacy and Data Protection via Runtime Scanning Methods
Francisco Rojo, Lei Pan 0002
BROADNETS2
2021 ECG-Adv-GAN: Detecting ECG Adversarial Examples with Conditional Generative Adversarial Networks
abstract
Electrocardiogram (ECG) acquisition requires an automated system and analysis pipeline for understanding specific rhythm irregularities. Deep neural networks have become a popular technique for tracing ECG signals, outperforming human experts. Despite this, convolutional neural networks are susceptible to adversarial examples that can misclassify ECG signals and decrease the model’s precision. Moreover, they do not generalize well on the out-of-distribution dataset. The GAN architecture has been employed in recent works to synthesize adversarial ECG signals to increase existing training data. However, they use a disjointed CNN-based classification architecture to detect arrhythmia. Till now, no versatile architecture has been proposed that can detect adversarial examples and classify arrhythmia simultaneously. To alleviate this, we propose a novel Conditional Generative Adversarial Network to simultaneously generate ECG signals for different categories and detect cardiac abnormalities. Moreover, the model is conditioned on class-specific ECG signals to synthesize realistic adversarial examples. Consequently, we compare our architecture and show how it outperforms other classification models in normal/abnormal ECG signal detection by benchmarking real world and adversarial signals.
Khondker Fariha Hossain, Sharif Amit Kamran, Alireza Tavakkoli, Lei Pan 0002, Xingjun Ma, Sutharshan Rajasegarar, Chandan Karmaker
ICMLA4
2021 Microwave Link Failures Prediction via LSTM-based Feature Fusion Network
abstract
Microwave links are widely employed in cellular data networks due to high-speed Internet access and easy installation, thus reducing network implementation costs. However, these links are prone to failure and may lead to performance degradation, unavailability and service disruption. Early detection of any link failures is critical to maintain network quality, but the complex environment and the dynamic nature of link information makes this a complicated process. In this work, we propose a Long Short-Term Memory (LSTM)-based feature fusion network (LSTM-FFN) to fuse and encode both homophy and structural equivalence relationships in the LSTM temporal feature learning network. This will simultaneously model the spatial and temporal features exhibited in Long-Term Evolution (LTE) networks to detect any link failures. Our proposed method effectively avoids the gradient exploding problem that RNN-based STGNN faced. This multi-scale topological feature fusion allows the LSTM-FFN to further explore the spatial dependencies among nodel/ink and include additional structural equivalence in modeling compared with previous network failure detection work. The evaluation results show that LSTM- FFN outperforms other statistical-based methods with and without network topology encoded, and reaches 94.1 % precision, 90.2 % recall and 92.1 % fl-score.
Zichan Ruan, Shuiqiao Yang, Lei Pan 0002, Xingjun Ma, Wei Luo 0001, Marthie Grobler
IJCNN3
2021 Energy-aware decision-making for dynamic task migration in MEC-based unmanned aerial vehicle delivery system
abstract
Abstract Nowadays, unmanned aerial vehicles (UAVs) are widely used in many smart systems such as smart logistics, smart agriculture, and environmental monitoring systems. However, the limited computing capability and restricted battery lifetime of existing UAVs could significantly impact the quality of service (QoS) of UAV‐based smart systems and the quality of experience (QoE) of end users. Recently, Mobile Edge Computing (MEC) which provisions computing resources close to the mobile end devices has become a promising solution. However, since high‐speed UAV often flies through the signal range of the different edge nodes, the interruption of services in the MEC‐based UAV delivery system is a critical issue. A challenging question is when and how to perform dynamic task migration among the edge nodes to ensure service continuity. In this paper, we investigate the task migration issue for multiple UAVs in the MEC‐based UAV delivery system. Specifically, we propose an energy‐aware decision‐making strategy for the dynamic task migration named GAD to optimize the UAV energy consumption. Given the real‐time system status and QoS constraints, and through a dynamic two‐tier decision‐making mechanism, GAD can efficiently make the task migration decision from four candidate decisions, viz. No Migration, Data Migration Only, Cold Migration, and Live Migration. Experimental results based on a real‐world scenario show that our strategy can well outperform other baseline strategies in various metrics including the flying distances and the energy consumption of UAVs.
Rui Li 0013, Xuejun Li 0001, Jia Xu 0010, Frank Jiang 0001, Di Shao, Lei Pan 0002, Xiao Liu 0004
Concurr. Comput. Pract. Exp.7
2021 Improving malicious PDF classifier with feature engineering: A data-driven approach
Ahmed Falah, Lei Pan 0002, Md. Shamsul Huda, Shiva Raj Pokhrel, Adnan Anwar
Future Gener. Comput. Syst.2
2021 Multipath TCP Meets Transfer Learning: A Novel Edge-Based Learning for Industrial IoT
abstract
We consider a fifth-generation (5G)-empowered future Industrial IoT (IIoT) networking problem where IIoT machines are capable of communicating and sharing their data networking knowledge gained (and experiences) with other neighboring devices/tools. For such an IIoT setting, deep-learning (DL)-based communication protocols are known to be highly efficient but having a computationally complex training procedure in terms of both time/space and volume of data sets. One solution for such training is to be completed offline for each equipment and machines of IIoT before deployment. A better approach would be to replicate the model from the expert existing machine and implant it into new machines. Such training for the transfer of knowledge can be done by manufacturers using high computational power, even for large-scale DL models. After sufficient training and the desired level of accuracy, the trained machines can be deployed in the smart factory equipment to perform life-long collaborative learning. We design a novel distributed transfer learning (TL) framework to maximize multipath communication networking performance for Industry 4.0 environment. To conduct seamless sharing of knowledge gain by the multipath TCP (MPTCP) agents and tackle retraining issues of DL-based approaches, we investigate TL for MPTCP from the IIoT networking perspective. With relevant insights from transfer and collaborative learning, we develop a distributed TL-MPTCP framework to accelerate the learning efficiency and enhance the performance of newly deployed machines. Our approach is validated with numerical and emulated NS-3 experiments in comparison with the state-of-the-art schemes.
Shiva Raj Pokhrel, Lei Pan 0002, Neeraj Kumar 0001, Robin Doss, Hai Le Vu 0001
IEEE Internet Things J.2
2021 SolGuard: Preventing external call issues in smart contract-based multi-agent robotic systems
Purathani Praitheeshan, Lei Pan 0002, James Xi Zheng, Alireza Jolfaei, Robin Doss
Inf. Sci.2
2021 Deep learning algorithms for cyber security applications: A survey
abstract
With the development of information technology, thousands of devices are connected to the Internet, various types of data are accessed and transmitted through the network, which pose huge security threats while bringing convenience to people. In order to deal with security issues, many effective solutions have been given based on traditional machine learning. However, due to the characteristics of big data in cyber security, there exists a bottleneck for methods of traditional machine learning in improving security. Owning to the advantages of processing big data and high-dimensional data, new solutions for cyber security are provided based on deep learning. In this paper, the applications of deep learning are classified, analyzed and summarized in the field of cyber security, and the applications are compared between deep learning and traditional machine learning in the security field. The challenges and problems faced by deep learning in cyber security are analyzed and presented. The findings illustrate that deep learning has a better effect on some aspects of cyber security and should be considered as the first option.
Guangjun Li, Preetpal Sharma, Lei Pan 0002, Sutharshan Rajasegarar, Chandan K. Karmakar, Nicholas Charles Patterson
J. Comput. Secur.3
2021 The Audio Auditor: User-Level Membership Inference in Internet of Things Voice Services
abstract
Abstract With the rapid development of deep learning techniques, the popularity of voice services implemented on various Internet of Things (IoT) devices is ever increasing. In this paper, we examine user-level membership inference in the problem space of voice services, by designing an audio auditor to verify whether a specific user had unwillingly contributed audio used to train an automatic speech recognition (ASR) model under strict black-box access. With user representation of the input audio data and their corresponding translated text, our trained auditor is effective in user-level audit. We also observe that the auditor trained on specific data can be generalized well regardless of the ASR model architecture. We validate the auditor on ASR models trained with LSTM, RNNs, and GRU algorithms on two state-of-the-art pipelines, the hybrid ASR system and the end-to-end ASR system. Finally, we conduct a real-world trial of our auditor on iPhone Siri, achieving an overall accuracy exceeding 80%. We hope the methodology developed in this paper and findings can inform privacy advocates to overhaul IoT privacy.
Yuantian Miao, Minhui Xue 0001, Chao Chen 0015, Lei Pan 0002, Jun Zhang 0010, Benjamin Zi Hao Zhao, Mohamed Ali Kâafar, Yang Xiang 0001
Proc. Priv. Enhancing Technol.4
2021 Software Vulnerability Discovery via Learning Multi-Domain Knowledge Bases
abstract
Machine learning (ML) has great potential in automated code vulnerability discovery. However, automated discovery application driven by off-the-shelf machine learning tools often performs poorly due to the shortage of high-quality training data. The scarceness of vulnerability data is almost always a problem for any developing software project during its early stages, which is referred to as the cold-start problem. This article proposes a framework that utilizes transferable knowledge from pre-existing data sources. In order to improve the detection performance, multiple vulnerability-relevant data sources were selected to form a broader base for learning transferable knowledge. The selected vulnerability-relevant data sources are cross-domain, including historical vulnerability data from different software projects and data from the Software Assurance Reference Database (SARD) consisting of synthetic vulnerability examples and proof-of-concept test cases. To extract the information applicable in vulnerability detection from the cross-domain data sets, we designed a deep-learning-based framework with Long-short Term Memory (LSTM) cells. Our framework combines the heterogeneous data sources to learn unified representations of the patterns of the vulnerable source codes. Empirical studies showed that the unified representations generated by the proposed deep learning networks are feasible and effective, and are transferable for real-world vulnerability detection. Our experiments demonstrated that by leveraging two heterogeneous data sources, the performance of our vulnerability detection outperformed the static vulnerability discovery toolFlawfinder. The findings of this article may stimulate further research in ML-based vulnerability detection using heterogeneous data sources.
Guanjun Lin, Jun Zhang 0010, Wei Luo 0001, Lei Pan 0002, Olivier Y. de Vel, Paul Montague, Yang Xiang 0001
IEEE Trans. Dependable Secur. Comput.4
2021 Data congestion in VANETs: research directions and new trends through a bibliometric analysis
Tarandeep Kaur Bhatia, Ramkumar Ketti Ramachandran, Robin Doss, Lei Pan 0002
J. Supercomput.4
2021 A novel cloud workflow scheduling algorithm based on stable matching game theory
Lei Pan 0002, Xiao Liu 0004, Xuejun Li 0001
J. Supercomput.2
2020 Graph Deep Learning Based Anomaly Detection in Ethereum Blockchain Network
Vatsal Patel, Lei Pan 0002, Sutharshan Rajasegarar
NSS2
2020 Security Evaluation of Smart Contract-Based On-chain Ethereum Wallets
Purathani Praitheeshan, Lei Pan 0002, Robin Doss
NSS2
2020 Unmasking Windows Advanced Persistent Threat Execution
abstract
The advanced persistent threat (APT) landscape has been studied without quantifiable data, for which indicators of compromise (IoC) may be uniformly analyzed, replicated, or used to support security mechanisms. This work culminates extensive academic and industry APT analysis, not as an incremental step in existing approaches to APT detection, but as a new benchmark of APT related opportunity. We collect 15,259 APT IoC hashes, retrieving subsequent sandbox execution logs across 41 different file types. This work forms an initial focus on Windows-based threat detection. We present a novel Windows APT executable (APT-EXE) dataset, made available to the research community. Manual and statistical analysis of the APT-EXE dataset is conducted, along with supporting feature analysis. We draw upon repeat and common APT paths access, file types, and operations within the APT-EXE dataset to generalize APT execution footprints. A baseline case analysis successfully identifies a majority of 117 of 152 live APT samples from campaigns across 2018 and 2019.
Rory Coulter, Jun Zhang 0010, Lei Pan 0002, Yang Xiang 0001
TrustCom3
2020 Opportunistic Tracking in Cyber-Physical Systems
abstract
Cyber-Physical Systems raise a new dimension of security concerns as they open up the opportunity for attackers to affect a real-world environment. These systems are often applied in specific environments with special requirements and a common issue is to keep track of movements in a mobile system, e.g., involving autonomous robots, drones or sensory I/O devices. In Opportunistic Networks, nodes are usually mobile, forwarding messages from one device to another, not relying on external infrastructure like WiFi. Due to compact and convenient wearability, the nodes of an OppNet might be used to detect the absence and presence of devices or even people in an area where classical networks may not be reliable enough. In this paper, we combine opportunistic network technology with cyber-physical systems and propose a reliable routing algorithm for nodes tracking. Our real-world setup implements hardware sensor tags to evaluate the algorithm in a state-of-the-art environment. Efficiency and performance are compared with established algorithms i. e., Epidemic and Prophet, in terms of latency, network overhead, as well as message delivery probability, and to evaluate the algorithm's scalability, we simulate the tracking in a huge environment.
Samaneh Rashidibajgan, Thomas Hupperich, Robin Doss, Lei Pan 0002
TrustCom4
2020 VoterChoice: A ransomware detection honeypot with multiple voting framework
abstract
Summary This research presents a novel framework comprising the IPS gateway, analysis system, and honeypot for identifying and detecting ransomware based on the client honeypot concept, and active interception of downloads using Suricata inline intruder prevention system. Unlike previous frameworks that report on the accuracy rate of detecting ransomware, the proposed framework features a multiple voting platform for the validation of confidence levels in the accuracy detection rates. The proposed framework achieves high accuracy levels than other machine learning models for the detection of ransomware.
Chee Keong Ng, Sutharshan Rajasegarar, Lei Pan 0002, Frank Jiang 0001, Leo Yu Zhang
Concurr. Comput. Pract. Exp.3
2020 Code analysis for intelligent cyber systems: A data-driven approach
Rory Coulter, Qing-Long Han, Lei Pan 0002, Jun Zhang 0010, Yang Xiang 0001
Inf. Sci.3
2020 Data-Driven Cyber Security in Perspective - Intelligent Traffic Analysis
abstract
Social and Internet traffic analysis is fundamental in detecting and defending cyber attacks. Traditional approaches resorting to manually defined rules are gradually replaced by automated approaches empowered by machine learning. This revolution is accelerated by huge datasets which support machine-learning models with outstanding performance. In the context of a data-driven paradigm, this article reviews recent analytic research on cyber traffic over social networks and the Internet by using a set of common concepts of similarity, correlation, and collective indication, and by sharing security goals for classifying network host or applications and users or Tweets. The ability to do so is not determined in isolation, but rather drawn for a wide use of many different network or social flows. Furthermore, the flows exhibit many characteristics, such as fixed sized and multiple messages between source and destination. This article demonstrates a new research methodology of data-driven cyber security (DDCS) and its application in social and Internet traffic analysis. The framework of the DDCS methodology consists of three components, that is, cyber security data processing, cyber security feature engineering, and cyber security modeling. Challenges and future directions in this field are also discussed.
Rory Coulter, Qing-Long Han, Lei Pan 0002, Jun Zhang 0010, Yang Xiang 0001
IEEE Trans. Cybern.3
2020 Privacy Protection in Interactive Content Based Image Retrieval
abstract
Privacy protection in Content Based Image Retrieval (CBIR) is a new research topic in cyber security and privacy. The state-of-art CBIR systems usually adopt interactive mechanism, namely relevance feedback, to enhance the retrieval precision. How to protect the user's privacy in such Relevance Feedback based CBIR (RF-CBIR) is a challenge problem. In this paper, we investigate this problem and propose a new Private Relevance Feedback CBIR (PRF-CBIR) scheme. PRF-CBIR can leverage the performance gain of relevance feedback and preserve the user's search intention at the same time. The new PRF-CBIR consists of three stages: 1) private query; 2) private feedback; 3) local retrieval. Private query performs the initial query with a privacy controllable feature vector; private feedback constructs the feedback image set by introducing confusing classes following theK-anonymity principle; local retrieval finally re-ranks the images in the user side. Privacy analysis shows that PRF-CBIR fulfills the privacy requirements. The experiments carried out on the real-world image collection confirm the effectiveness of the proposed PRF-CBIR scheme.
Yonggang Huang 0001, Jun Zhang 0010, Lei Pan 0002, Yang Xiang 0001
IEEE Trans. Dependable Secur. Comput.3
2019 Domain-Adversarial Graph Neural Networks for Text Classification
abstract
Text classification, in cross-domain setting, is a challenging task. On the one hand, data from other domains are often useful to improve the learning on the target domain; on the other hand, domain variance and hierarchical structure of documents from words, key phrases, sentences, paragraphs, etc. make it difficult to align domains for effective learning. To date, existing cross-domain text classification methods mainly strive to minimize feature distribution differences between domains, and they typically suffer from three major limitations - (1) difficult to capture semantics in non-consecutive phrases and long-distance word dependency because of treating texts as word sequences, (2) neglect of hierarchical coarse-grained structures of document for feature learning, and (3) narrow focus of the domains at instance levels, without using domains as supervisions to improve text classification. This paper proposes an end-to-end, domain-adversarial graph neural networks (DAGNN), for cross-domain text classification. Our motivation is to model documents as graphs and use a domain-adversarial training principle to lean features from each graph (as well as learning the separation of domains) for effective text classification. At the instance level, DAGNN uses a graph to model each document, so that it can capture non-consecutive and long-distance semantics. At the feature level, DAGNN uses graphs from different domains to jointly train hierarchical graph neural networks in order to learn good features. At the learning level, DAGNN proposes a domain-adversarial principle such that the learned features not only optimally classify documents but also separates domains. Experiments on benchmark datasets demonstrate the effectiveness of our method in cross-domain classification tasks.
Man Wu, Shirui Pan, Xingquan Zhu 0001, Chuan Zhou 0001, Lei Pan 0002
ICDM5
2019 Multiple Energy Harvesting Devices Enabled Joint Computation Offloading and Dynamic Resource Allocation for Mobile-Edge Computing Systems
abstract
A mobile-edge computing (MEC) system integrating energy harvesting (EH) techniques is a promising paradigm for supporting computation-intensive and delay-sensitive mobile applications. While computation offloading reduces users' perceived latency, EH techniques mitigate the limitation of mobile devices' battery capacity. However, when considering a scenario with multiple EH devices, the advantages of MEC systems with EH devices may be compromised due to the competition among multiple devices for available computational resources and wireless bandwidth. In this paper, the joint computation offloading and dynamic resource allocation (JCODRA) that minimizes the long-term average execution cost is formulated as a stochastic optimization problem. In particular, both the long-term average execution delay and the penalty delay are included in the optimization objective. The former intends to handle the competition and random and uncontrollable EH processes properly and the latter aims to reduce the ratio of dropped tasks. An online algorithm based on Lyapunov optimization is proposed to transform the original problem into a per-time slot deterministic problem. The results of experiments demonstrate that our algorithm significantly outperforms three representative baseline approaches.
Wei Du 0001, Qiwang Lei, Qiang He 0001, Wei Liu 0011, Feifei Chen 0001, Lei Pan 0002, Hailiang Zhao
ICWS6
2019 Editorial: Recent advances in machine learning for cybersecurity
abstract
Cybersecurity has become a very hot topic in recent years. Many communities, groups, and governments start to realize the importance and urgency to deal with the ever‐changing cyberattacks.1, 2 Experts in the industry and scholars in the academia strive to innovate the next‐generation solutions. Among the technical solutions, machine learning–based methods receive an increasingly popular favor due to its superior efficiency comparing with manual analysis.2 In addition, the instantaneous protection brought by the machine learning–based solutions surpasses most reactive technologies including the automated protection systems in terms of reaction time. The general approach of machine learning–based cybersecurity solutions includes establishment of ground‐truth data, feature extraction and engineering, and model tuning. To perform these steps, one needs domain‐specific knowledge in cybersecurity and insights to machine learning principles and skills. This special issue aims to solicit cybersecurity researchers to publish the latest research findings in cybersecurity and privacy with the use of machine learning. With the conjunction of cybersecurity and machine learning, the submitted manuscripts have been evaluated in a rigorous and critical manner by professionals from different sections including both industry and academia. This review process leads to the seven accepted papers that are included in this special issue. These selected papers belong to three main research directions: system security, cybersecurity applications, and privacy applications.
Lei Pan 0002, Jun Zhang 0010, Jonathan Oliver
Concurr. Comput. Pract. Exp.1
2019 Automatic extraction and integration of behavioural indicators of malware for protection of cyber-physical networks
Md. Shamsul Huda, Jemal H. Abawajy, Baker Al-Rubaie, Lei Pan 0002, Mohammad Mehedi Hassan
Future Gener. Comput. Syst.4
2019 Noise-Resistant Statistical Traffic Classification
abstract
Network traffic classification plays a significant role in cyber security applications and management scenarios. Conventional statistical classification techniques rely on the assumption that clean labelled samples are available for building classification models. However, in the big data era, mislabelled training data commonly exist due to the introduction of new applications and lack of knowledge. Existing statistical traffic classification techniques do not address the problem of mislabelled training data, so their performance become poor in the presence of mislabelled training data. To meet this challenge, in this paper, we propose a new scheme, Noise-resistant Statistical Traffic Classification (NSTC), which incorporates the techniques of noise elimination and reliability estimation into traffic classification. NSTC estimates the reliability of the remaining training data before it builds a robust traffic classifier. Through a number of traffic classification experiments on two real-world traffic data sets, the results show that the new NSTC scheme can effectively address the problem of mislabelled training data. Compared with the state of the art methods, NSTC can significantly improve the classification performance in the context of big unclean data.
Binfeng Wang, Jun Zhang 0010, Zili Zhang 0001, Lei Pan 0002, Yang Xiang 0001, Dawen Xia
IEEE Trans. Big Data4
2018 Keep Calm and Know Where to Focus: Measuring and Predicting the Impact of Android Malware
Junyang Qiu, Wei Luo 0001, Surya Nepal, Jun Zhang 0010, Yang Xiang 0001, Lei Pan 0002
ADMA6
2018 High-rate and high-capacity measurement-device-independent quantum key distribution with Fibonacci matrix coding in free space
Hong Lai, Mingxing Luo, Josef Pieprzyk, Jun Zhang 0010, Lei Pan 0002, Mehmet A. Orgun
Sci. China Inf. Sci.5
2018 Intelligent agents defending for an IoT world: A review
Rory Coulter, Lei Pan 0002
Comput. Secur.2
2018 Comprehensive analysis of network traffic data
abstract
Summary With the large volume of network traffic flow, it is necessary to preprocess raw data before classification to gain the accurate results speedily. Feature selection is an essential approach in preprocessing phase. The principal component analysis (PCA) is recognized as an effective and efficient method. In this paper, we classify network traffic flows by using the PCA technique together with 6 machine learning algorithms—Naive Bayes, decision tree, 1‐nearest neighbor, random forest, support vector machine, andH2O. We analyzed the impact of PCA on the classification results by applying each algorithm with and without PCA onto the data set. Experiments were set out by varying the size of input data sets, and the performances were measured from 2 aspects, including average overall accuracy and F‐measure. The computational time was also considered in analyzing the performance. Our results showed that random forest and 1‐nearest neighbor were the top 2 algorithms among all the 6 regarding the 2 metrics mentioned above. Then we continued the study of PCA impact on per class level with these 2 algorithms as examples. And the positive correlation between overall impact and the number of class with significant impact was revealed. Lastly, the visualization was used in exploring the reasons of the impacts caused by PCA. Two factors are considered in PCA's impact on per class level: benefit for classes grouped by PCA and mislabeled error interfered by nearby groups.
Yuantian Miao, Zichan Ruan, Lei Pan 0002, Jun Zhang 0010, Yang Xiang 0001
Concurr. Comput. Pract. Exp.3
2018 Identifying items for moderation in a peer assessment framework
Simon James, Elicia Lanham, Vicky H. Mak-Hau, Lei Pan 0002, Tim Wilkin 0001, Guy Wood-Bradley
Knowl. Based Syst.4
2018 Big network traffic data visualization
Zichan Ruan, Yuantian Miao, Lei Pan 0002, Yang Xiang 0001, Jun Zhang 0010
Multim. Tools Appl.3
2018 Exploring Feature Coupling and Model Coupling for Image Source Identification
abstract
Recently, there has been great interest in feature-based image source identification. Previous statistical learning-based methods usually regarded the identification process as a classification problem. They assumed the dependence of features and the dependence of models. However, the two assumptions are usually problematic because of the genuine coupling of features and models. To address the issues, in this paper, we propose a novel image source identification scheme. For the feature coupling, a coupled feature representation is adopted to analyze the coupled interaction among features. The coupling relations among features and their powers are measured with Pearson’s correlations and integrated in a Taylor-like expansion manner. Regarding model coupling, a new coupled probability representation is developed. The model coupling relationships are characterized with conditional probabilities induced by the confusion matrix and then combined with the law of total probability. The experiments carried out on the Dresden image collection confirm the effectiveness of the proposed scheme. Via mining the feature coupling and model coupling, the identification accuracy can be significantly improved.
Yonggang Huang 0001, Longbing Cao, Jun Zhang 0010, Lei Pan 0002, Yuying Liu 0005
IEEE Trans. Inf. Forensics Secur.4
2018 Cross-Project Transfer Representation Learning for Vulnerable Function Discovery
abstract
Machine learning is now widely used to detect security vulnerabilities in the software, even before the software is released. But its potential is often severely compromised at the early stage of a software project when we face a shortage of high-quality training data and have to rely on overly generic hand-crafted features. This paper addresses this cold-start problem of machine learning, by learning rich features that generalize across similar projects. To reach an optimal balance between feature-richness and generalizability, we devise a data-driven method including the following innovative ideas. First, the code semantics are revealed through serialized abstract syntax trees (ASTs), with tokens encoded by Continuous Bag-of-Words neural embeddings. Next, the serialized ASTs are fed to a sequential deep learning classifier (Bi-LSTM) to obtain a representation indicative of software vulnerability. Finally, the neural representation obtained from existing software projects is then transferred to the new project to enable early vulnerability detection even with a small set of training labels. To validate this vulnerability detection approach, we manually labeled 457 vulnerable functions and collected 30 000+ nonvulnerable functions from six open-source projects. The empirical results confirmed that the trained model is capable of generating representations that are indicative of program vulnerability and is adaptable across multiple projects. Compared with the traditional code metrics, our transfer-learned representations are more effective for predicting vulnerable functions, both within a project and across multiple projects.
Guanjun Lin, Jun Zhang 0010, Wei Luo 0001, Lei Pan 0002, Yang Xiang 0001, Olivier Y. de Vel, Paul Montague
IEEE Trans. Ind. Informatics4
2017 POSTER: Vulnerability Discovery with Function Representation Learning from Unlabeled Projects
abstract
In cybersecurity, vulnerability discovery in source code is a fundamental problem. To automate vulnerability discovery, Machine learning (ML) based techniques has attracted tremendous attention. However, existing ML-based techniques focus on the component or file level detection, and thus considerable human effort is still required to pinpoint the vulnerable code fragments. Using source code files also limit the generalisability of the ML models across projects. To address such challenges, this paper targets at the function-level vulnerability discovery in the cross-project scenario. A function representation learning method is proposed to obtain the high-level and generalizable function representations from the abstract syntax tree (AST). First, the serialized ASTs are used to learn project independence features. Then, a customized bi-directional LSTM neural network is devised to learn the sequential AST representations from the large number of raw features. The new function-level representation demonstrated promising performance gain, using a unique dataset where we manually labeled 6000+ functions from three open-source projects. The results confirm that the huge potential of the new AST-based function representation learning.
Guanjun Lin, Jun Zhang 0010, Wei Luo 0001, Lei Pan 0002, Yang Xiang 0001
CCS4
2017 Online peer marking with aggregation functions
abstract
With the rise of Massive Open Online Courses (MOOCs), online peer marking is an attractive contemporary tool for educational assessment. However its widespread use faces serious challenges, most significantly in the perceived and actual reliability of assessment grades, which can be affected by the ability of peers to mark accurately and the potential for collusion and bias. There exist a number of aggregation approaches for alleviating the impact of biased scores, usually involving either the down-weighting or removal of outliers. Here we investigate the use of the least trimmed squares (LTS) and Huber mean for the aggregation step, comparing their performance to weighting of markers based on divergence from other peers' marks. We design an experimental setup to generate scores and test a number of conditions. Overall we find that for a feasible number of peer markers, when the student pool comprises a significant number of `biased' markers, outlier removal techniques are likely to result in a number of very unfair assessments, while more standard approaches will have more grades unfairly influenced but to a lesser extent.
Simon James, Lei Pan 0002, Tim Wilkin 0001, Lilin Yin
FUZZ-IEEE2
2017 Cyber security attacks to modern vehicular systems
Lei Pan 0002, James Xi Zheng, H. X. Chen, Tom H. Luan, H. Bootwala, Lynn Margaret Batten
J. Inf. Secur. Appl.1
2014 Privacy Preserving in Location Data Release: A Differential Privacy Approach
Ping Xiong 0001, Tianqing Zhu, Lei Pan 0002, Wenjia Niu, Gang Li 0009
PRICAI3
2006 Lease Based Addressing for Event-Driven Wireless Sensor Networks
abstract
Sensor Networks have applications in diverse fields. They can be deployed for habitat modeling, temperature monitoring and industrial sensing. They also find applications in battlefield awareness and emergency (first) response situations. While unique addressing is not a requirement of many data collecting applications of wireless sensor networks it is vital for the success of applications such as emergency response. Data that cannot be associated with a specific node becomes useless in such situations. In this work we propose an addressing mechanism for event-driven wireless sensor networks. The proposed scheme eliminates the need for network wide Duplicate Address Detection (DAD) and enables reuse of addresses.
Robin Doss, Deddy Chandra, Lei Pan 0002, Wanlei Zhou 0001, Morshed U. Chowdhury
ISCC3
2005 Reproducibility of Digital Evidence in Forensic Investigations
Lei Pan 0002, Lynn Margaret Batten
DFRWS1