VLDB 2026 Research / reviewers in the wild / expert
Jun Zhang 0010
dblp:z/JunZhang10
· DBLP profile ↗
112ranked-venue papers
16as first author
36since 2021 · last 2025
0000-0002-2189-7801ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 47 · 3 first-author · 23 since 2021Systems, architecture and hardware · 17 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 10 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 6 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 1 since 2021Computer networks · 8 · 2 first-author · 1 since 2021Theory of computation · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Bridging Clone Detection and Industrial Compliance: A Practical Pipeline for Enterprise Codebases
Shigang Liu, Jun Zhang 0010, Yang Xiang 0001 |
ACISP (3) | 3 |
| 2025 | Poster: The Art of Deception: Crafting Chimera Images for Covert and Robust Semantic Poisoning AttacksabstractWith the exponential surge in media data volumes and their growing intrinsic value, the landscape has become increasingly susceptible to persistent and strategically designed data poisoning attacks targeting these valuable assets. In this work, we propose a novel approach leveraging generative AI techniques to craft covert and robust poisonous data samples, referred to as Chimera Images. These images seamlessly blend visual features from two target classes to generate hybrid objects that preserve appearance fidelity. These ''normal'' samples with correct labels can subtly distort the model's decision boundary without raising suspicion. Extensive experimental results on CIFAR-10 and Flowers datasets demonstrate that the proposed method i) reduces the accuracy of the targeted class, ii) maintains the performance of other classes, and iii) exhibits immunity to state-of-the-art defence strategies. We also explore the usage of generative AI content detection as a defence mechanism, demonstrating that the recently discovered snapshot technique is ineffective against the AI-generated poisonous Chimera samples. Lin Li 0066, Youyang Qu, Jiayang Ao, Ming Ding 0001, Chao Chen 0015, Jun Zhang 0010 |
CCS | 6 |
| 2025 | Poster: Decoding Social Engineering: A Multi-Level Framework for Tactic Generation, Annotation, and EvaluationabstractPhishing emails increasingly embed complex social engineering (SE) tactics to manipulate recipients and increase success rates. However, existing organizational training simulations and detection systems seldom incorporate tactic complexity or reveal how such tactics are linguistically embedded. To address this, we develop methods for generating, annotating, and evaluating SE tactics across three complexity levels in phishing emails. A reliably annotated dataset is constructed via a generate–cross-verify–highlight pipeline, which ensures semantic alignment between labels and embedded SE tactics. These trigger segments are subsequently clustered and synthesized into fine-grained patterns that characterize how each SE tactic manifests at Level 1 (easily), Level 2 (moderately), and Level 3 (deeply). These patterns underpin a multi-level SE framework, validated through LLM-based detection experiments. Detection accuracy declines with increasing tactic complexity, confirming the framework's stratification capability and its utility in training, simulation, and tactic-aware detection design. Yicun Tian, Youyang Qu, Ming Ding 0001, Shigang Liu, Pei-Wei Tsai, Jun Zhang 0010 |
CCS | 6 |
| 2025 | Large Language Models for Cybersecurity Education: A Survey of Current Practices and Future Directions
Nan Sun 0002, Yuantian Miao, Xiaoxing Mo, Jun Zhang 0010 |
PAKDD (6) | 4 |
| 2025 | VulCodeMark: Adaptive Watermarking for Vulnerability Datasets ProtectionabstractCode datasets are invaluable for training neural vulnerability detectors, a promising area within software engineering. Unfortunately, both proprietary and public datasets face the threat of unauthorized exploitation. Moreover, the opaque nature of neural models presents a challenge for external auditing of training datasets, exacerbating the risk of potential misuse. Although watermarking techniques have proven effective in protecting image and natural language datasets, their applicability to code datasets is limited by domain specificity. Current endeavours to preserve the copyrights of code datasets frequently neglect essential control and data dependency information, treating code as a flat structure. To address these gaps, we propose VulCodeMark, a pioneering method that incorporates data and control flow information into code dataset watermarking. VulCodeMark employs two transformations (1) Syntactic Transformation; (2) Semantic Transformation) to generate watermarks, ensuring the preservation of the original program’s functionality while maintaining context-adaptive stealthiness. Experiments have demonstrated that VulCodeMark fulfils essential properties of practical watermarks-including harmlessness, effectiveness, imperceptibility, and robustness. Besides, VulCodeMark additionally supports preliminary probing of model architecture configurations, furnishing valuable forensic evidence in cases of intellectual property infringement. Shigang Liu, Jun Zhang 0010, Yang Xiang 0001 |
RAID | 3 |
| 2025 | SoK: Private Knowledge Sharing in Distributed LearningabstractThe rapid advancement of Artificial Intelligence (AI) has transformed various industries, leading to the widespread distribution of AI models and data across intelligent systems. As modern data driven services increasingly integrate distributed knowledge entities, decentralized learning has become a prevalent approach to training AI models. However, this collaborative learning paradigm introduces significant security vulnerabilities and privacy challenges. This paper presents a comprehensive systematic review on private knowledge sharing in distributed learning, analyzing key knowledge components utilized in leading distributed learning architectures. We identify critical vulnerabilities associated with these components and examine defensive strategies to safeguard privacy while mitigating potential adversarial threats. Additionally, we highlight key limitations in knowledge sharing in distributed learning and propose future research directions to enhance security and efficiency in decentralized AI systems. Yasas Supeksala, Thilina Ranbaduge, Ming Ding 0001, Dinh C. Nguyen, Bo Liu 0001, Caslon Chua, Jun Zhang 0010 |
Proc. Priv. Enhancing Technol. | 7 |
| 2025 | Extracting Private Training Data in Federated Learning From ClientsabstractThe utilization of machine learning algorithms in distributed web applications is experiencing significant growth. One notable approach is Federated Learning (FL) Recent research has brought attention to the vulnerability of FL to gradient inversion attacks, which seek to reconstruct the original training samples, posing a substantial threat to client privacy. Most existing gradient inversion attacks, however, require control over the central server and rely on substantial prior knowledge, including information about batch normalization and data distribution. In this study, we introduce Poisoning Gradient Leakage from Client (PGLC), a novel attack method that operates from the clients’ side. For the first time, we demonstrate the feasibility of a client-side adversary with limited knowledge successfully recovering training samples from the aggregated global model. Our approach enables the adversary to employ a malicious model that increases the loss of a specific targeted class of interest. When honest clients employ the poisoned global model, the gradients of samples become distinct in the aggregated update. This allows the adversary to effectively reconstruct private inputs from other clients using the aggregated update. Furthermore, our PGLC attack exhibits stealthiness against Byzantine-robust aggregation rules (AGRs). Through the optimization of malicious updates and the blending of benign updates with a malicious replacement vector, our method remains undetected by these defense mechanisms. We conducted experiments across various benchmark datasets, considering representative Byzantine-robust AGRs and exploring different FL settings with varying levels of adversary knowledge about the data. Our results consistently demonstrate the ability of PGLC to extract training data in all tested scenarios. Jiaheng Wei, Yanjun Zhang 0002, Leo Yu Zhang, Chao Chen 0015, Shirui Pan, Kok-Leong Ong, Jun Zhang 0010, Yang Xiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | Flash: Federated Graph Learning-Based Malicious Bash Script Detection for Industrial Cyber-Physical Systems
Pengbin Feng, Ning Xi 0002, Jiong Jin, Jun Zhang 0010, Jianfeng Ma 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2025 | PRIME: A Phishing Detection Framework With Quantitative and Fuzzy-Based Dual Validation
Yicun Tian, Youyang Qu, Ming Ding 0001, Shigang Liu, Pei-Wei Tsai, Jun Zhang 0010 |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2024 | VulMatch: Binary-Level Vulnerability Detection Through Signature
Zian Liu, Shigang Liu, Lei Pan 0002, Chao Chen 0015, Jun Zhang 0010, Dongxi Liu |
NSS | 6 |
| 2024 | EaTVul: ChatGPT-based Evasion Attack Against Software Vulnerability Detection
Shigang Liu, Junae Kim, Tamas Abraham, Paul Montague, Seyit Ahmet Çamtepe, Jun Zhang 0010, Yang Xiang 0001 |
USENIX Security Symposium | 7 |
| 2024 | Deceptive Waves: Embedding Malicious Backdoors in PPG Authentication
Zeming Yao, Lin Li 0066, Leo Yu Zhang, Fusen Guo, Chao Chen 0015, Jun Zhang 0010 |
WISE (2) | 6 |
| 2024 | IoTFuzz: Automated Discovery of Violations in Smart Homes With Real EnvironmentabstractSmart homes (SHs) are rapidly evolving to incorporate intelligent features, including environment management, home automation, and human–machine interactions. However, safety and security risks of SHs hinder their wide adoption. Many work attempts to provide defense mechanisms to ensure safety and security against interrule vulnerabilities and spoofing attacks. This article proposes IoTFuzz, a fuzzing framework that dynamically address cyber security and physical safety aspects of SHs through targeted policies. IoTFuzz mutates the inputs from policies, human activities, indoor environment, and real-life outdoor weather conditions. In addition to the binary status of devices, the continuous-value status in SHs is leveraged to perform mutation and simulation. The policies are expressed as temporal logic formulas with time constraints. For large-scale testing, IoTFuzz employs digital twins to simulate normal behaviors, outdoor environment impacts, and human activities in SHs. Moreover, IoTFuzz can also intelligently infer rule-policy correlation based on natural language processing (NLP) techniques. The evaluation of IoTFuzz in a configured SH with 15 rules and 10 predefined unique policies demonstrates its effectiveness in revealing the impacts of real-life outdoor environment. The experimental results demonstrate a range of violations, with a maximum of 4154 violations and a minimum of 41 violations observed over an 8-year period under varying weather conditions. IoTFuzz also identifies the potential risks associated with improper human activities, accounting for up to 35.4% of risky situations in SHs. Xinbo Ban, Ming Ding 0001, Shigang Liu, Chao Chen 0015, Jun Zhang 0010 |
IEEE Internet Things J. | 5 |
| 2023 | Hiding Your Signals: A Security Analysis of PPG-Based Biometric Authentication
Lin Li 0066, Chao Chen 0015, Lei Pan 0002, Yonghang Tai, Jun Zhang 0010, Yang Xiang 0001 |
ESORICS (3) | 5 |
| 2023 | SigD: A Cross-Session Dataset for PPG-based User Authentication in Different Demographic GroupsabstractRecently, unobservable physiological signals have received widespread attention from researchers as unique identifiers of users in biometrics. However, due to the lack of data sets, existing methods are limited in evaluating cross-session scenarios. Cross-session means that signals are collected at different sessions (times). In real scenarios, authentication is almost always cross-session. Currently, the datasets commonly used for Photoplethysmogram (PPG) signal authentication span around one month, which is insufficient for authentication. On the other hand, different demographic groups have different hemodynamic characteristics, but existing methods lack an assessment of these aspects. This paper introduces a dataset to provide insights into PPG signal-based authentication across different time spans and user groups (age, gender). As physiological signals offer unique advantages for user authentication, the potential of PPG signals is gradually explored. Furthermore, our comparative analysis of recent publications on data-driven user authentication using PPG can further identify the similarities and differences among the performance of the proposed authentication models. Our findings may help future research towards a consensus on an appropriate set of performance metrics. Lin Li 0066, Chao Chen 0015, Lei Pan 0002, Jun Zhang 0010, Yang Xiang 0001 |
IJCNN | 4 |
| 2023 | SigA: rPPG-based Authentication for Virtual Reality Head-mounted DisplayabstractConsumer-grade virtual reality head-mounted displays (VR-HMD) are becoming increasingly popular. Despite VR’s convenience and booming applications, VR-based authentication schemes are underdeveloped. The recently proposed authentication methods (Electrooculogram based, Electrical Muscle Stimulation-based, and alike) require active user involvement, disturbing many scenarios like drone flight and telemedicine. This paper proposes an effective and efficient user authentication method in VR environments resilient to impersonation attacks using physiological signals — Photoplethysmogram (PPG), namely SigA. SigA exploits the advantage that PPG is a physiological signal invisible to the naked eye. Using VR-HMDs to cover the eye area completely, SigA reduces the risk of signal leakage during PPG acquisition. We conducted a comprehensive analysis of SigA’s feasibility on five publicly available datasets, nine different pre-trained models, three facial regions, various lengths of the video clips required for training, four different signal time intervals, and continuous authentication with different sliding window sizes. The results demonstrate that SigA achieves more than 95% of the average F1-score in a one-second signal to accommodate a complete cardiac cycle for most adults, implying its applicability in real-world scenarios. Furthermore, experiments have shown that SigA is resistant to zero-effort attacks, statistical attacks, impersonation attacks (with a detection accuracy of over 95%) and session hijacking attacks. Lin Li 0066, Chao Chen 0015, Lei Pan 0002, Leo Yu Zhang, Jun Zhang 0010, Yang Xiang 0001 |
RAID | 5 |
| 2023 | Security and privacy problems in voice assistant applications: A surveyabstractVoice assistant applications have become omniscient nowadays. Two models that provide the two most important functions for real-life applications (i.e., Google Home, Amazon Alexa, Siri, etc.) are Automatic Speech Recognition (ASR) models and Speaker Identification (SI) models. According to recent studies, security and privacy threats have also emerged with the rapid development of the Internet of Things (IoT). The security issues researched include attack techniques toward machine learning models and other hardware components widely used in voice assistant applications. The privacy issues include technical-wise information stealing and policy-wise privacy breaches. The voice assistant application takes a steadily growing market share every year, but their privacy and security issues never stopped causing huge economic losses and endangering users' personal sensitive information. Thus, it is important to have a comprehensive survey to outline the categorization of the current research regarding the security and privacy problems of voice assistant applications. This paper concludes and assesses five kinds of security attacks and three types of privacy threats in the papers published in the top-tier conferences of cyber security and voice domain. Jingjin Li, Chao Chen 0015, Mostafa Rahimi Azghadi, Hossein Ghodosi, Lei Pan 0002, Jun Zhang 0010 |
Comput. Secur. | 6 |
| 2023 | A Survey of PPG's Application in AuthenticationabstractBiometric authentication prospered because of its convenient use and security. Early generations of biometric mechanisms suffer from spoofing attacks. Recently, unobservable physiological signals (e.g., Electroencephalogram, Photoplethysmogram, Electrocardiogram) as biometrics offer a potential remedy to this problem. In particular, Photoplethysmogram (PPG) measures the change in blood flow of the human body by an optical method. Clinically, researchers commonly use PPG signals to obtain patients' blood oxygen saturation, heart rate, and other information to assist in diagnosing heart-related diseases. Since PPG signals contain a wealth of individual cardiac information, researchers have begun to explore their potential in cyber security applications. The unique advantages (simple acquisition, difficult to steal, and live detection) of the PPG signal allow it to improve the security and usability of the authentication in various aspects. However, the research on PPG-based authentication is still in its infancy. The lack of systematization hinders new research in this field. We conduct a comprehensive study of PPG-based authentication and discuss these applications' limitations before pointing out future research directions. Lin Li 0066, Chao Chen 0015, Lei Pan 0002, Leo Yu Zhang, Jun Zhang 0010, Yang Xiang 0001 |
Comput. Secur. | 6 |
| 2023 | Analysis of hybrid attack and defense based on block withholding strategy
Binjie Liao, Yu Wang 0017, Weizhi Meng 0001, Jun Zhang 0010 |
J. Inf. Secur. Appl. | 5 |
| 2023 | Cyber Code Intelligence for Android Malware DetectionabstractEvolving Android malware poses a severe security threat to mobile users, and machine-learning (ML)-based defense techniques attract active research. Due to the lack of knowledge, many zero-day families’ malware may remain undetected until the classifier gains specialized knowledge. The most existing ML-based methods will take a long time to learn new malware families in the latest malware family landscape. Existing ML-based Android malware detection and classification methods struggle with the fast evolution of the malware landscape, particularly in terms of the emergence of zero-day malware families and limited representation of single-view features. In this article, a new multiview feature intelligence (MFI) framework is developed to learn the representation of a targeted capability from known malware families for recognizing unknown and evolving malware with the same capability. The new framework performs reverse engineering to extract multiview heterogeneous features, including semantic string features, API call graph features, and smali opcode sequential features. It can learn the representation of a targeted capability from known malware families through a series of processes of feature analysis, selection, aggregation, and encoding, to detect unknown Android malware with shared target capability. We create a new dataset with ground-truth information regarding capability. Many experiments are conducted on the new dataset to evaluate the performance and effectiveness of the new method. The results demonstrate that the new method outperforms three state-of-the-art methods, including: 1) Drebin; 2) MaMaDroid; and 3)$N$-opcode, when detecting unknown Android malware with targeted capabilities. Junyang Qiu, Qing-Long Han, Wei Luo 0001, Lei Pan 0002, Surya Nepal, Jun Zhang 0010, Yang Xiang 0001 |
IEEE Trans. Cybern. | 6 |
| 2023 | Making DeepFakes More Spurious: Evading Deep Face Forgery Detection via Trace Removal AttackabstractDeepFakes are raising significant social concerns. Although various Despite various DeepFake detectors having been developed as countermeasures, their vulnerability under attacks remains further explorations. Recently, several attacks, such as adversarial attacks, have successfully fooled DeepFake detectors. However, existing attacks suffer from detector-specific designs, requiring detector-side knowledge, leading to poor transferability. Moreover, they only consider simplified security scenarios; but less is known about the attacking performance in complex scenarios where the capability of detectors or attackers varies. To fill the gap, we propose a novel, detector-agnostic trace removal attack. The attack removes all possible counterfeiting traces arising from the original DeepFake manufacture procedure to make DeepFakes essentially more "realistic" and thus able to defeat arbitrary or unknown detectors. Concretely, we first perform an in-depth DeepFake trace discovery, identifying three intrinsic traces: spatial anomalies, spectral disparities, and noise fingerprints. Then an adversarial learning-based trace removal network (TR-Net) involving one generator and multiple discriminators is proposed. Each discriminator is responsible for one individual trace representation to avoid inner-trace interference. All discriminators are optimized in parallel to enforce the generator to remove various traces simultaneously. We additionally craft heterogeneous security scenarios where the detectors are embedded with different levels of defense and the attackers own varying background data knowledge. The experimental results show that the proposed trace removal attack can significantly compromise the detection accuracy of six state-of-the-art DeepFake detectors while causing only a negligible degradation in visual quality. Chi Liu 0002, Huajie Chen, Tianqing Zhu, Jun Zhang 0010, Wanlei Zhou 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2023 | Cyber Information Retrieval Through Pragmatics Understanding and VisualizationabstractThe amount of cybersecurity-related information is extraordinarily increasing, given the fast-growing number of cybersecurity attacks and the significant influence brought by them. How to efficiently obtain and precisely understand the relevant knowledge in the sea of information on cybersecurity becomes a challenge. In this article, we propose an innovative cybersecurity retrieval scheme that supports automatic indexing and searching of cybersecurity information based on semantic contents and hidden metadata. The proposed scheme leverages a customized neural model that incorporates new linguistic features and word embedding by identifying the entities related to cybersecurity incidents from the text. We implement a novel cybersecurity search engine to demonstrate effective, understandable and pragmatic cybersecurity information retrieval based on the proposed schema. Comprehensive performance evaluation over real-world datasets has been conducted to validate the new algorithms and techniques developed for cybersecurity information retrieval. The new engine makes it possible to conduct augmented search, cybersecurity analytics, and visualization, with the ultimate goal of providing direct and efficient results to help people obtain and truly understand cybersecurity information. Nan Sun 0002, Jun Zhang 0010, Shang Gao 0003, Leo Yu Zhang, Seyit Ahmet Çamtepe, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2022 | No-Label User-Level Membership Inference for ASR Model Auditing
Yuantian Miao, Chao Chen 0015, Lei Pan 0002, Shigang Liu, Seyit Ahmet Çamtepe, Jun Zhang 0010, Yang Xiang 0001 |
ESORICS (2) | 6 |
| 2022 | Automated Binary Analysis: A Survey
Zian Liu, Chao Chen 0015, Dongxi Liu, Jun Zhang 0010 |
ICA3PP | 5 |
| 2022 | Attention Distraction: Watermark Removal Through Continual Learning with Selective ForgettingabstractFine-tuning attacks are effective in removing the embedded watermarks in deep learning models. However, when the source data is unavailable, it is challenging to just erase the watermark without jeopardizing the model performance. In this context, we introduce Attention Distraction (AD), a novel source data-free watermark removal attack, to make the model selectively forget the embedded watermarks by customizing continual learning. In particular, AD first anchors the model's attention on the main task using some unlabeled data. Then, through continual learning, a small number of lures (randomly selected natural images) that are assigned a new label distract the model's attention away from the watermarks. Experimental results from different datasets and networks corroborate that AD can thoroughly remove the watermark with a small resource budget without compromising the model's performance on the main task, which outperforms the state-of-the-art works. Leo Yu Zhang, Shengshan Hu, Longxiang Gao, Jun Zhang 0010, Yong Xiang 0001 |
ICME | 5 |
| 2022 | A Survey on IoT Vulnerability Discovery
Xinbo Ban, Ming Ding 0001, Shigang Liu, Chao Chen 0015, Jun Zhang 0010 |
NSS | 5 |
| 2022 | Domain adaptation for Windows advanced persistent threat detection
Rory Coulter, Jun Zhang 0010, Lei Pan 0002, Yang Xiang 0001 |
Comput. Secur. | 2 |
| 2022 | SDCCP: Control the network using software-defined networking and end-to-end congestion controlabstractSummary The Internet of Things is becoming widely popular in the past decade, which comes with huge amount of data. These magnanimous data, stored in data centers, put forward the new demand for the efficient management of the network. In this article, we propose Software‐Defined Congestion Control Plane (SDCCP), a hybrid network control architecture that aims to fully utilize the network while avoiding congestion. SDCCP is based on Software‐Defined Networking and CCP, in which the controller collects the network statistics and specifies the behavior of the end‐to‐end hosts by sending feedback or modifying their transport layer parameters directly. It can also be used to mitigate Distributed Denial of Service attacks and other security problems. In addition, we propose FCA, a Feedback‐based Congestion Avoidance algorithm running on SDCCP, which adapts the congestion window based on the feedback from the remote controller. We evaluate SDCCP and FCA in Mininet and the result shows that FCA can achieve high network utilization while keeping the queue length of the routers in a low level. Also, FCA is robust to noncongestion loss, and outperforms other algorithms at high loss rate. Jiashuo Lin, Liping Liao, Tao Wang 0014, Jun Zhang 0010, Lianglun Cheng |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | Trustworthy blockchain-based medical Internet of thing for minimal invasive surgery training simulatorabstractSummary Realistic modeling of mechanical behavior of soft tissue has been recognized as an essential part for medical Internet of thing for minimal invasive surgery (MIS) training simulator. Therefore, the blockchain‐based constitutive model is crucial for mechanical response of soft tissue modeling. In this article, based on the Ogden second order model, a novel hyperplastic model was presented to describe the stress‐stretch relationship in the MIS training system. To validate this theoretical model, two experimental techniques (uniaxial compression and uniaxial tensile) were conducted to obtain data related to stress‐strain in the blockchain system, which plays an important role in investigating the mechanical behavior of soft tissue. Our results show that the new model has a satisfied coincidence of the experimental data than other existing models. Furthermore, the viscoelastic properties of soft tissue were investigated and a viscoelastic model based on three‐parameter was utilized to interpret the viscoelastic behavior of the soft tissue. The contributions of this article include several biomechanical tests that were performed to investigate the soft tissue hyperelastic and viscoelastic properties in the MIS system, and theoretical guidance for simulating soft tissue mechanical behavior in the blockchain‐based simulation system. Yonghang Tai, Yinjia Wang, Lei Wei 0002, Lei Pan 0002, Jun Zhang 0010, Junsheng Shi |
Concurr. Comput. Pract. Exp. | 7 |
| 2022 | CD-VulD: Cross-Domain Vulnerability Discovery Based on Deep Domain AdaptationabstractA major cause of security incidents such as cyber attacks is rooted in software vulnerabilities. These vulnerabilities should ideally be found and fixed before the code gets deployed. Machine learning-based approaches achieve state-of-the-art performance in capturing vulnerabilities. These methods are predominantly supervised. Their prediction models are trained on a set of ground truth data where the training data and test data are assumed to be drawn from the same probability distribution. However, in practice, the test data often differs from the training data in terms of distribution because they are from different projects or they differ in the types of vulnerability. In this article, we present a new system forCrossDomain SoftwareVulnerabilityDiscovery (CD-VulD) using deep learning (DL) and domain adaptation (DA). We employ DL because it has the capacity of automatically constructing high-level abstract feature representations of programs, which are likely of more cross-domain useful than the handcrafted features driven by domain knowledge. The divergence between distributions is reduced by learning cross-domain representations. First, given software program representations, CD-VulD converts them into token sequences and learns the token embeddings for generalization across tokens. Next, CD-VulD employs a deep feature model to build abstract high-level presentations based on those sequences. Then, the metric transfer learning framework (MTLF) technique is employed to learn cross-domain representations by minimizing the distribution divergence between the source domain and the target domain. Finally, the cross-domain representations are used to build a classifier for vulnerability detection. Experimental results show that CD-VulD outperforms the state-of-the-art vulnerability detection approaches by a wide margin. We make the new datasets publicly available so that our work is replicable and can be further improved. Shigang Liu, Guanjun Lin, Lizhen Qu, Jun Zhang 0010, Olivier Y. de Vel, Paul Montague, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2022 | Lightweight and Certificateless Multi-Receiver Secure Data Transmission Protocol for Wireless Body Area NetworksabstractThe rapid development of low-power integrated circuits, wireless communication, intelligent sensors, and microelectronics has allowed the realization of wireless body area networks (WBANs), which can monitor patients’ vital body parameters remotely in real time to offer timely treatment. These vital body parameters are related to patients’ life and health; and these highly private data are subject to many security threats. To guarantee privacy, many secure communication protocols have been proposed. However, most of these protocols have a one-to-one structure in extra-body communication and cannot support multidisciplinary team (MDT). Hence, we propose a lightweight and certificateless multi-receiver secure data transmission protocol for WBANs to support MDT treatment in this article. In particular, a novel multi-receiver certificateless generalized signcryption (MR-CLGSC) scheme is proposed that can adaptively use only one algorithm to implement one of three cryptographic primitives: signature, encryption or signcryption. Then, a multi-receiver secure data transmission protocol based on the MR-CLGSC scheme with many security properties, such as data integrity and confidentiality, non-repudiation, anonymity, forward and backward secrecy, unlinkability and data freshness, is designed. Both security analysis and performance analysis show that the proposed protocol for WBANs is secure, efficient, and highly practical. Jian Shen 0001, Ziyuan Gui, Xiaofeng Chen 0001, Jun Zhang 0010, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2022 | JSCSP: A Novel Policy-Based XSS Defense Mechanism for BrowsersabstractTo mitigate cross-site scripting attacks (XSS), the W3C group recommends web service providers to employ a computer security standard called Content Security Policy (CSP). However, less than 3.7 percent of real-world websites are equipped with CSP according to Google’s survey. The low scalability of CSP is incurred by the difficulty of deployment and non-compatibility for state-of-art browsers. To explore the scalability of CSP, in this article, we propose JavaScript based CSP (JSCSP), which is able to support most of real-world browsers but also to generate security policies automatically. Specifically, JSCSP offers a novel self-defined security policy which enforces essential confinements to related items, including JavaScript functions, DOM elements and data access. Meanwhile, JSCSP has an efficient algorithm to automatically generate the policy directives and enforce them in a cascading way, which is more fine-grained and practical than the functionalities provided by CSP. We further implement JSCSP on a Chrome extension, and our evaluation shows that the extension is compatible with popular JavaScript libraries. Our JSCSP extension can detect and block the tested attacking vectors extracted from the prevalent web applications. We state that JSCSP delivers better performance compared to other XSS defense solutions. Guangquan Xu, Xiaofei Xie, Shuhan Huang, Jun Zhang 0010, Lei Pan 0002, Wei Lou, Kaitai Liang |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2022 | CSEdge: Enabling Collaborative Edge Storage for Multi-Access Edge Computing Based on BlockchainabstractMulti-access Edge Computing (MEC), as an extension of cloud computing, provides storage resources at the network edge to enable low-latency data retrieval for users. Due to limited physical sizes and constrained storage resources, individual edge servers cannot store a large amount of data when operating independently. They often need to offload data to other edge servers to serve users collaboratively. Operated by different edge infrastructure providers, edge servers usually work in a distrusted environment. Incentive and trust are the two main challenges in facilitating collaborative edge storage. This article proposes CSEdge, a novel decentralized system that tackles these challenges to enable collaborative edge storage based on blockchain. On CSEdge, edge servers can submit data offloading requests for others to contend for. Winners are selected based on their reputations. They will store the offloaded data and receive rewards for successfully finishing data offloading tasks. Via a distributed consensus, their performance will be recorded on blockchain for future reputation evaluation. A prototype of CSEdge is built on Hyperledger Sawtooth and experimentally evaluated against a baseline system and two start-of-the-art systems in a simulated MEC environment. The results demonstrate that CSEdge can effectively and efficiently facilitate collaborative edge storage among edge servers. Qiang He 0001, Feifei Chen 0001, Jun Zhang 0010, Lianyong Qi, Xiaolong Xu 0001, Yang Xiang 0001, Yun Yang 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2021 | Deep neural-based vulnerability discovery demystified: data, model and performance
Guanjun Lin, Leo Yu Zhang, Shang Gao 0003, Yonghang Tai, Jun Zhang 0010 |
Neural Comput. Appl. | 6 |
| 2021 | The Audio Auditor: User-Level Membership Inference in Internet of Things Voice ServicesabstractAbstract With the rapid development of deep learning techniques, the popularity of voice services implemented on various Internet of Things (IoT) devices is ever increasing. In this paper, we examine user-level membership inference in the problem space of voice services, by designing an audio auditor to verify whether a specific user had unwillingly contributed audio used to train an automatic speech recognition (ASR) model under strict black-box access. With user representation of the input audio data and their corresponding translated text, our trained auditor is effective in user-level audit. We also observe that the auditor trained on specific data can be generalized well regardless of the ASR model architecture. We validate the auditor on ASR models trained with LSTM, RNNs, and GRU algorithms on two state-of-the-art pipelines, the hybrid ASR system and the end-to-end ASR system. Finally, we conduct a real-world trial of our auditor on iPhone Siri, achieving an overall accuracy exceeding 80%. We hope the methodology developed in this paper and findings can inform privacy advocates to overhaul IoT privacy. Yuantian Miao, Minhui Xue 0001, Chao Chen 0015, Lei Pan 0002, Jun Zhang 0010, Benjamin Zi Hao Zhao, Mohamed Ali Kâafar, Yang Xiang 0001 |
Proc. Priv. Enhancing Technol. | 5 |
| 2021 | Software Vulnerability Discovery via Learning Multi-Domain Knowledge BasesabstractMachine learning (ML) has great potential in automated code vulnerability discovery. However, automated discovery application driven by off-the-shelf machine learning tools often performs poorly due to the shortage of high-quality training data. The scarceness of vulnerability data is almost always a problem for any developing software project during its early stages, which is referred to as the cold-start problem. This article proposes a framework that utilizes transferable knowledge from pre-existing data sources. In order to improve the detection performance, multiple vulnerability-relevant data sources were selected to form a broader base for learning transferable knowledge. The selected vulnerability-relevant data sources are cross-domain, including historical vulnerability data from different software projects and data from the Software Assurance Reference Database (SARD) consisting of synthetic vulnerability examples and proof-of-concept test cases. To extract the information applicable in vulnerability detection from the cross-domain data sets, we designed a deep-learning-based framework with Long-short Term Memory (LSTM) cells. Our framework combines the heterogeneous data sources to learn unified representations of the patterns of the vulnerable source codes. Empirical studies showed that the unified representations generated by the proposed deep learning networks are feasible and effective, and are transferable for real-world vulnerability detection. Our experiments demonstrated that by leveraging two heterogeneous data sources, the performance of our vulnerability detection outperformed the static vulnerability discovery toolFlawfinder. The findings of this article may stimulate further research in ML-based vulnerability detection using heterogeneous data sources. Guanjun Lin, Jun Zhang 0010, Wei Luo 0001, Lei Pan 0002, Olivier Y. de Vel, Paul Montague, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2020 | Data Analytics of Crowdsourced Resources for Cybersecurity Intelligence
Nan Sun 0002, Jun Zhang 0010, Shang Gao 0003, Leo Yu Zhang, Seyit Ahmet Çamtepe, Yang Xiang 0001 |
NSS | 2 |
| 2020 | Protecting IP of Deep Neural Networks with Watermarking: A New Label Helps
Leo Yu Zhang, Jun Zhang 0010, Longxiang Gao, Yong Xiang 0001 |
PAKDD (2) | 3 |
| 2020 | Unmasking Windows Advanced Persistent Threat ExecutionabstractThe advanced persistent threat (APT) landscape has been studied without quantifiable data, for which indicators of compromise (IoC) may be uniformly analyzed, replicated, or used to support security mechanisms. This work culminates extensive academic and industry APT analysis, not as an incremental step in existing approaches to APT detection, but as a new benchmark of APT related opportunity. We collect 15,259 APT IoC hashes, retrieving subsequent sandbox execution logs across 41 different file types. This work forms an initial focus on Windows-based threat detection. We present a novel Windows APT executable (APT-EXE) dataset, made available to the research community. Manual and statistical analysis of the APT-EXE dataset is conducted, along with supporting feature analysis. We draw upon repeat and common APT paths access, file types, and operations within the APT-EXE dataset to generalize APT execution footprints. A baseline case analysis successfully identifies a majority of 117 of 152 live APT samples from campaigns across 2018 and 2019. Rory Coulter, Jun Zhang 0010, Lei Pan 0002, Yang Xiang 0001 |
TrustCom | 2 |
| 2020 | Protecting the Intellectual Property of Deep Neural Networks with Watermarking: The Frequency Domain ApproachabstractSimilar to other digital assets, deep neural network (DNN) models could suffer from piracy threat initiated by insider and/or outsider adversaries due to their inherent commercial value. DNN watermarking is a promising technique to mitigate this threat to intellectual property. This work focuses on black-box DNN watermarking, with which an owner can only verify his ownership by issuing special trigger queries to a remote suspicious model. However, informed attackers, who are aware of the watermark and somehow obtain the triggers, could forge fake triggers to claim their ownerships since the poor robustness of triggers and the lack of correlation between the model and the owner identity. This consideration calls for new watermarking methods that can achieve better trade-off for addressing the discrepancy. In this paper, we exploit frequency domain image watermarking to generate triggers and build our DNN watermarking algorithm accordingly. Since watermarking in the frequency domain is high concealment and robust to signal processing operation, the proposed algorithm is superior to existing schemes in resisting fraudulent claim attack. Besides, extensive experimental results on 3 datasets and 8 neural networks demonstrate that the proposed DNN watermarking algorithm achieves similar performance on functionality metrics and better performance on security metrics when compared with existing algorithms. Leo Yu Zhang, Yajuan Du, Jun Zhang 0010, Yong Xiang 0001 |
TrustCom | 5 |
| 2020 | Doc2vec-based Insider Threat Detection through Behaviour Analysis of Multi-source Security LogsabstractSince insider attacks have been recognised as one of the most critical cyber security threats to an organisation, detection of malicious insiders has received increasing attention in recent years. Previously, we proposed an approach that performs the detection by analysing various security logs with Word2vec, which not only removes the reliance on prior knowledge but also greatly simplifies the process of decision making and improves the interpretability of the alerts. In this paper, following the similar idea, a new Doc2vec based approach is proposed to overcome the previous approach's limitations: (1) the behaviour metrics can be acquired straightforwardly due to the Doc2vec's capability in inferring unseen texts of any length; (2) other than the temporal metrics, some spatial metrics can also be realised, providing a more comprehensive insight into the unusual behaviours; and (3) a range of corpora are produced by adopting different keywords to aggregate, each of which may be suited to a specific type of behaviour metrics. A large number of numerical experiments are conducted using the same benchmark insider threat database, for the purpose of testing how the corpora, metrics and training parameters impact on the performance and be related to each other. The experiments demonstrate that the proposed approach can achieve a similar performance with greater simplicity and flexibility. Liu Liu 0008, Chao Chen 0015, Jun Zhang 0010, Olivier Y. de Vel, Yang Xiang 0001 |
TrustCom | 3 |
| 2020 | A Hybrid Key Agreement Scheme for Smart Homes Using the Merkle PuzzleabstractCryptographic keys should be established for the smart home devices in order to secure home area networks. In certain smart home applications, however, the devices might be produced by different factories. As a result, it becomes impractical to assume devices are preloaded with secrets before leaving factories. Moreover, in some scenarios, smart home devices have no access to an online trusted third party. These problems make conventional key agreement schemes inapplicable for these devices. It is investigated that devices can extract secrets from received signal strength (RSS) measurements at the physical layer. However, the bit extraction rate is low. The Merkle puzzle is introduced to design the key agreement scheme by expanding the low entropy seed to high entropy secret key. However, it introduces considerable time and computation costs. To alleviate these problems, in this article, we design a hybrid key agreement scheme for smart homes. Namely, smart home devices first extract short random keys at the physical layer. Then, they establish secret communication keys at higher layers by making use of the Merkle puzzle. In this way, secret keys can be established without using any preloaded secrets or online trusted third party. We prove the security of the new scheme and present a prototype implementation using Ralink WiFi cards. Moreover, we evaluate the performance of our scheme and compare it with other related schemes. The analysis shows that comparing with the related schemes, in our scheme, the time cost is at least one order of magnitude lower, the computation cost is at least five orders of magnitude lower, and the extra communication cost is moderate. Yuexin Zhang, Xinyi Huang 0001, Xiaofeng Chen 0001, Leo Yu Zhang, Jun Zhang 0010, Yang Xiang 0001 |
IEEE Internet Things J. | 5 |
| 2020 | Code analysis for intelligent cyber systems: A data-driven approach
Rory Coulter, Qing-Long Han, Lei Pan 0002, Jun Zhang 0010, Yang Xiang 0001 |
Inf. Sci. | 4 |
| 2020 | AI-driven data security and privacy
Zheng Yan 0002, Willy Susilo, Elisa Bertino, Jun Zhang 0010, Laurence T. Yang |
J. Netw. Comput. Appl. | 4 |
| 2020 | Software Vulnerability Detection Using Deep Neural Networks: A SurveyabstractThe constantly increasing number of disclosed security vulnerabilities have become an important concern in the software industry and in the field of cybersecurity, suggesting that the current approaches for vulnerability detection demand further improvement. The booming of the open-source software community has made vast amounts of software code available, which allows machine learning and data mining techniques to exploit abundant patterns within software code. Particularly, the recent breakthrough application of deep learning to speech recognition and machine translation has demonstrated the great potential of neural models’ capability of understanding natural languages. This has motivated researchers in the software engineering and cybersecurity communities to apply deep learning for learning and understanding vulnerable code patterns and semantics indicative of the characteristics of vulnerable code. In this survey, we review the current literature adopting deep-learning-/neural-network-based approaches for detecting software vulnerabilities, aiming at investigating how the state-of-the-art research leverages neural techniques for learning and understanding code semantics to facilitate vulnerability discovery. We also identify the challenges in this new field and share our views of potential research directions. Guanjun Lin, Sheng Wen, Qing-Long Han, Jun Zhang 0010, Yang Xiang 0001 |
Proc. IEEE | 4 |
| 2020 | Data-Driven Cyber Security in Perspective - Intelligent Traffic AnalysisabstractSocial and Internet traffic analysis is fundamental in detecting and defending cyber attacks. Traditional approaches resorting to manually defined rules are gradually replaced by automated approaches empowered by machine learning. This revolution is accelerated by huge datasets which support machine-learning models with outstanding performance. In the context of a data-driven paradigm, this article reviews recent analytic research on cyber traffic over social networks and the Internet by using a set of common concepts of similarity, correlation, and collective indication, and by sharing security goals for classifying network host or applications and users or Tweets. The ability to do so is not determined in isolation, but rather drawn for a wide use of many different network or social flows. Furthermore, the flows exhibit many characteristics, such as fixed sized and multiple messages between source and destination. This article demonstrates a new research methodology of data-driven cyber security (DDCS) and its application in social and Internet traffic analysis. The framework of the DDCS methodology consists of three components, that is, cyber security data processing, cyber security feature engineering, and cyber security modeling. Challenges and future directions in this field are also discussed. Rory Coulter, Qing-Long Han, Lei Pan 0002, Jun Zhang 0010, Yang Xiang 0001 |
IEEE Trans. Cybern. | 4 |
| 2020 | Privacy Protection in Interactive Content Based Image RetrievalabstractPrivacy protection in Content Based Image Retrieval (CBIR) is a new research topic in cyber security and privacy. The state-of-art CBIR systems usually adopt interactive mechanism, namely relevance feedback, to enhance the retrieval precision. How to protect the user's privacy in such Relevance Feedback based CBIR (RF-CBIR) is a challenge problem. In this paper, we investigate this problem and propose a new Private Relevance Feedback CBIR (PRF-CBIR) scheme. PRF-CBIR can leverage the performance gain of relevance feedback and preserve the user's search intention at the same time. The new PRF-CBIR consists of three stages: 1) private query; 2) private feedback; 3) local retrieval. Private query performs the initial query with a privacy controllable feature vector; private feedback constructs the feedback image set by introducing confusing classes following theK-anonymity principle; local retrieval finally re-ranks the images in the user side. Privacy analysis shows that PRF-CBIR fulfills the privacy requirements. The experiments carried out on the real-world image collection confirm the effectiveness of the proposed PRF-CBIR scheme. Yonggang Huang 0001, Jun Zhang 0010, Lei Pan 0002, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2020 | DeepBalance: Deep-Learning and Fuzzy Oversampling for Vulnerability DetectionabstractSoftware vulnerability has long been an important but critical research issue in cybersecurity. Recently, the machine learning (ML)-based approach has attracted increasing interest in the research of software vulnerability detection. However, the detection performance of existing ML-based methods require further improvement. There are two challenges: one is code representation for ML and the other is class imbalance between vulnerable code and nonvulnerable code. To overcome these challenges, this article develops a DeepBalance system, which combines the new ideas of deep code representation learning and fuzzy-based class rebalancing. We design a deep neural network with bidirectional long short-term memory to learn invariant and discriminative code representations from labeled vulnerable and nonvulnerable code. Then, a new fuzzy oversampling method is employed to rebalance the training data by generating synthetic samples for the class of vulnerable code. To evaluate the performance of the new system, we carry out a series of experiments in a real-world ground-truth dataset that consists of the code from the projects of LibTIFF, LibPNG, and FFmpeg. The results show that the proposed new system can significantly improve the vulnerability detection performance. For example, the improvement is 15% in terms of F-measure. Shigang Liu, Guanjun Lin, Qing-Long Han, Sheng Wen, Jun Zhang 0010, Yang Xiang 0001 |
IEEE Trans. Fuzzy Syst. | 5 |
| 2020 | Android HIV: A Study of Repackaging Malware for Evading Machine-Learning DetectionabstractMachine learning-based solutions have been successfully employed for the automatic detection of malware on Android. However, machine learning models lack robustness to adversarial examples, which are crafted by adding carefully chosen perturbations to the normal inputs. So far, the adversarial examples can only deceive detectors that rely on syntactic features (e.g., requested permissions, API calls,etc.), and the perturbations can only be implemented by simply modifying application’s manifest. While recent Android malware detectors rely more on semantic features from Dalvik bytecode rather than manifest, existing attacking/defending methods are no longer effective. In this paper, we introduce a new attacking method that generates adversarial examples of Android malware and evades being detected by the current models. To this end, we propose a method of applying optimal perturbations onto Android APK that can successfully deceive the machine learning detectors. We develop an automated tool to generate the adversarial examples without human intervention. In contrast to existing works, the adversarial examples crafted by our method can also deceive recent machine learning-based detectors that rely on semantic features such as control-flow-graph. The perturbations can also be implemented directly onto APK’s Dalvik bytecode rather than Android manifest to evade from recent detectors. We demonstrate our attack on two state-of-the-art Android malware detection schemes, MaMaDroid and Drebin. Our results show that the malware detection rates decreased from 96% to 0% in MaMaDroid, and from 97% to 0% in Drebin, with just a small number of codes to be inserted into the APK. Xiao Chen 0002, Derui Wang, Sheng Wen, Jun Zhang 0010, Surya Nepal, Yang Xiang 0001, Kui Ren 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2020 | Cyber Vulnerability Intelligence for Internet of Things BinaryabstractInternet of Things (IoT) integrates a variety of software (e.g., autonomous vehicles and military systems) in order to enable the advanced and intelligent services. These software increase the potential of cyber-attacks because an adversary can launch an attack using system vulnerabilities. Existing software vulnerability analysis methods used to be relying on human experts crafted features, which usually miss many vulnerabilities. It is important to develop an automatic vulnerability analysis system to improve the countermeasures. However, source code is not always available (e.g., most IoT related industry software are closed source). Therefore, vulnerability detection on binary code is a demanding task. This article addresses the automatic binary-level software vulnerability detection problem by proposing a deep learning-based approach. The proposed approach consists of two phases: binary function extraction, and model building. First, we extract binary functions from the cleaned binary instructions obtained by using IDA Pro. Then, we employ the attention mechanism on top of a bidirectional long short-term memory for building the predictive model. To show the effectiveness of the proposed approach, we have collected datasets from several different sources. We have compared our proposed approach with a series of baselines including source code-based techniques and binary code-based techniques. We have also applied the proposed approach to real-world IoT related software such as VLC media player and LibTIFF project that used on Autonomous Vehicles. Experimental results show that our proposed approach betters the baselines and is able to detect more vulnerabilities. Shigang Liu, Mahdi Dibaei, Yonghang Tai, Chao Chen 0015, Jun Zhang 0010, Yang Xiang 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2019 | Deep Learning-Based Vulnerable Function Detection: A Benchmark
Guanjun Lin, Jun Zhang 0010, Yang Xiang 0001 |
ICICS | 3 |
| 2019 | Unsupervised Insider Detection Through Neural Feature Learning and Model Optimisation
Liu Liu 0008, Chao Chen 0015, Jun Zhang 0010, Olivier Y. de Vel, Yang Xiang 0001 |
NSS | 3 |
| 2019 | Editorial: Recent advances in machine learning for cybersecurityabstractCybersecurity has become a very hot topic in recent years. Many communities, groups, and governments start to realize the importance and urgency to deal with the ever‐changing cyberattacks.1, 2 Experts in the industry and scholars in the academia strive to innovate the next‐generation solutions. Among the technical solutions, machine learning–based methods receive an increasingly popular favor due to its superior efficiency comparing with manual analysis.2 In addition, the instantaneous protection brought by the machine learning–based solutions surpasses most reactive technologies including the automated protection systems in terms of reaction time. The general approach of machine learning–based cybersecurity solutions includes establishment of ground‐truth data, feature extraction and engineering, and model tuning. To perform these steps, one needs domain‐specific knowledge in cybersecurity and insights to machine learning principles and skills. This special issue aims to solicit cybersecurity researchers to publish the latest research findings in cybersecurity and privacy with the use of machine learning. With the conjunction of cybersecurity and machine learning, the submitted manuscripts have been evaluated in a rigorous and critical manner by professionals from different sections including both industry and academia. This review process leads to the seven accepted papers that are included in this special issue. These selected papers belong to three main research directions: system security, cybersecurity applications, and privacy applications. Lei Pan 0002, Jun Zhang 0010, Jonathan Oliver |
Concurr. Comput. Pract. Exp. | 2 |
| 2019 | Video denoising for security and privacy in fog computingabstractSummary To reduce heavy noise from degraded video in low or predictable latency and preserve privacy, a powerful and efficient video denoising algorithm is proposed based on fog computing for Visual Internet of Things. The conventional method is to remove noise in the cloud; however, this may overload computation and communication and raise security and privacy issues. The proposed denoising algorithm is distributed to heterogeneous devices at network edges to preserve privacy and avoid security risks as noise can be reduced in the fog rather than the cloud. To address the problems of latency, communication rate, and extremely heavy noise, structure registration, inter‐frame and inner‐frame filters, and distribution compensation are applied in the proposed algorithm. A scheme for encrypting the denoised data at network edges is provided so that security and privacy issues may be avoided during transmission and storage. Compared with other denoising approaches under extremely heavy noise conditions, the experimental results demonstrate that the proposed approach achieves superior denoising performance in terms of peak signal‐noise ratio and visual quality at low computational cost, high bandwidth efficiency, and low‐latency response in a fog computing manner. Hong Zhang 0018, Yifan Yang 0003, Ding Yuan 0001, Daniel Sun 0004, Jun Zhang 0010, Guoqiang Li 0001, Mingui Sun |
Concurr. Comput. Pract. Exp. | 5 |
| 2019 | Noise-Resistant Statistical Traffic ClassificationabstractNetwork traffic classification plays a significant role in cyber security applications and management scenarios. Conventional statistical classification techniques rely on the assumption that clean labelled samples are available for building classification models. However, in the big data era, mislabelled training data commonly exist due to the introduction of new applications and lack of knowledge. Existing statistical traffic classification techniques do not address the problem of mislabelled training data, so their performance become poor in the presence of mislabelled training data. To meet this challenge, in this paper, we propose a new scheme, Noise-resistant Statistical Traffic Classification (NSTC), which incorporates the techniques of noise elimination and reliability estimation into traffic classification. NSTC estimates the reliability of the remaining training data before it builds a robust traffic classifier. Through a number of traffic classification experiments on two real-world traffic data sets, the results show that the new NSTC scheme can effectively address the problem of mislabelled training data. Compared with the state of the art methods, NSTC can significantly improve the classification performance in the context of big unclean data. Binfeng Wang, Jun Zhang 0010, Zili Zhang 0001, Lei Pan 0002, Yang Xiang 0001, Dawen Xia |
IEEE Trans. Big Data | 2 |
| 2018 | Keep Calm and Know Where to Focus: Measuring and Predicting the Impact of Android Malware
Junyang Qiu, Wei Luo 0001, Surya Nepal, Jun Zhang 0010, Yang Xiang 0001, Lei Pan 0002 |
ADMA | 4 |
| 2018 | A Data-driven Attack against Support Vectors of SVMabstractMachine learning (ML) is commonly used in multiple disciplines and real-world applications, such as information retrieval, financial systems, health, biometrics and online social networks. However, their security profiles against deliberate attacks have not often been considered. Sophisticated adversaries can exploit specific vulnerabilities exposed by classical ML algorithms to deceive intelligent systems. It is emerging to perform a thorough security evaluation as well as potential attacks against the machine learning techniques before developing novel methods to guarantee that machine learning can be securely applied in adversarial setting. In this paper, an effective attack strategy for crafting foreign support vectors in order to attack a classic ML algorithm, the Support Vector Machine (SVM) has been proposed with mathematical proof. The new attack can minimize the margin around the decision boundary and maximize the hinge loss simultaneously. We evaluate the new attack in different real-world applications including social spam detection, Internet traffic classification and image recognition. Experimental results highlight that the security of classifiers can be worsened by poisoning a small group of support vectors. Shigang Liu, Jun Zhang 0010, Yu Wang 0017, Wanlei Zhou 0001, Yang Xiang 0001, Olivier Y. de Vel |
AsiaCCS | 2 |
| 2018 | Roundtable Gossip Algorithm: A Novel Sparse Trust Mining Method for Large-Scale Recommendation Systems
Guangquan Xu, Jun Zhang 0010, Rajan Shankaran, James Xi Zheng, Zonghua Zhang |
ICA3PP (4) | 3 |
| 2018 | Who Spread to Whom? Inferring Online Social Networks with User FeaturesabstractNetwork inference has been extensively studied to better understand the information diffusion in online social networks. In this field, state-of- art widely adopted a priori knowledge related to users' infection timestamps. Researchers also assume that the smaller the time difference between two nodes, the higher the likelihood of an edge between the pair of users. However, according to our technical analyses and empirical studies, existing methods have two critical problems 1) alternative spreading paths; 2) users' delivery delay, which leads to the inaccuracy of previous methods. In this paper, we developed an innovative method to address the inference inaccuracy caused by the exposed two problems. This method determined the existence of an edge between a pair of users according to part of the users' features. The experiment results suggested that our method achieved around 70% accuracy in inferring network structures while existing methods failed in the same tasks. Derek Wang, Wanlei Zhou 0001, James Xi Zheng, Sheng Wen, Jun Zhang 0010, Yang Xiang 0001 |
ICC | 5 |
| 2018 | High-rate and high-capacity measurement-device-independent quantum key distribution with Fibonacci matrix coding in free space
Hong Lai, Mingxing Luo, Josef Pieprzyk, Jun Zhang 0010, Lei Pan 0002, Mehmet A. Orgun |
Sci. China Inf. Sci. | 4 |
| 2018 | JFCGuard: Detecting juice filming charging attack via processor usage analysis on smartphones
Weizhi Meng 0001, Lijun Jiang, Yu Wang 0017, Jin Li 0002, Jun Zhang 0010, Yang Xiang 0001 |
Comput. Secur. | 5 |
| 2018 | Comprehensive analysis of network traffic dataabstractSummary With the large volume of network traffic flow, it is necessary to preprocess raw data before classification to gain the accurate results speedily. Feature selection is an essential approach in preprocessing phase. The principal component analysis (PCA) is recognized as an effective and efficient method. In this paper, we classify network traffic flows by using the PCA technique together with 6 machine learning algorithms—Naive Bayes, decision tree, 1‐nearest neighbor, random forest, support vector machine, andH2O. We analyzed the impact of PCA on the classification results by applying each algorithm with and without PCA onto the data set. Experiments were set out by varying the size of input data sets, and the performances were measured from 2 aspects, including average overall accuracy and F‐measure. The computational time was also considered in analyzing the performance. Our results showed that random forest and 1‐nearest neighbor were the top 2 algorithms among all the 6 regarding the 2 metrics mentioned above. Then we continued the study of PCA impact on per class level with these 2 algorithms as examples. And the positive correlation between overall impact and the number of class with significant impact was revealed. Lastly, the visualization was used in exploring the reasons of the impacts caused by PCA. Two factors are considered in PCA's impact on per class level: benefit for classes grouped by PCA and mislabeled error interfered by nearby groups. Yuantian Miao, Zichan Ruan, Lei Pan 0002, Jun Zhang 0010, Yang Xiang 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2018 | Recent advances in security and privacy in Social Big Data
Jun Zhang 0010, Aniello Castiglione, Laurence T. Yang, Yan Zhang 0002 |
Future Gener. Comput. Syst. | 1 |
| 2018 | Big network traffic data visualization
Zichan Ruan, Yuantian Miao, Lei Pan 0002, Yang Xiang 0001, Jun Zhang 0010 |
Multim. Tools Appl. | 5 |
| 2018 | Exploring Feature Coupling and Model Coupling for Image Source IdentificationabstractRecently, there has been great interest in feature-based image source identification. Previous statistical learning-based methods usually regarded the identification process as a classification problem. They assumed the dependence of features and the dependence of models. However, the two assumptions are usually problematic because of the genuine coupling of features and models. To address the issues, in this paper, we propose a novel image source identification scheme. For the feature coupling, a coupled feature representation is adopted to analyze the coupled interaction among features. The coupling relations among features and their powers are measured with Pearson’s correlations and integrated in a Taylor-like expansion manner. Regarding model coupling, a new coupled probability representation is developed. The model coupling relationships are characterized with conditional probabilities induced by the confusion matrix and then combined with the law of total probability. The experiments carried out on the Dresden image collection confirm the effectiveness of the proposed scheme. Via mining the feature coupling and model coupling, the identification accuracy can be significantly improved. Yonggang Huang 0001, Longbing Cao, Jun Zhang 0010, Lei Pan 0002, Yuying Liu 0005 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2018 | Cross-Project Transfer Representation Learning for Vulnerable Function DiscoveryabstractMachine learning is now widely used to detect security vulnerabilities in the software, even before the software is released. But its potential is often severely compromised at the early stage of a software project when we face a shortage of high-quality training data and have to rely on overly generic hand-crafted features. This paper addresses this cold-start problem of machine learning, by learning rich features that generalize across similar projects. To reach an optimal balance between feature-richness and generalizability, we devise a data-driven method including the following innovative ideas. First, the code semantics are revealed through serialized abstract syntax trees (ASTs), with tokens encoded by Continuous Bag-of-Words neural embeddings. Next, the serialized ASTs are fed to a sequential deep learning classifier (Bi-LSTM) to obtain a representation indicative of software vulnerability. Finally, the neural representation obtained from existing software projects is then transferred to the new project to enable early vulnerability detection even with a small set of training labels. To validate this vulnerability detection approach, we manually labeled 457 vulnerable functions and collected 30 000+ nonvulnerable functions from six open-source projects. The empirical results confirmed that the trained model is capable of generating representations that are indicative of program vulnerability and is adaptable across multiple projects. Compared with the traditional code metrics, our transfer-learned representations are more effective for predicting vulnerable functions, both within a project and across multiple projects. Guanjun Lin, Jun Zhang 0010, Wei Luo 0001, Lei Pan 0002, Yang Xiang 0001, Olivier Y. de Vel, Paul Montague |
IEEE Trans. Ind. Informatics | 2 |
| 2017 | POSTER: Vulnerability Discovery with Function Representation Learning from Unlabeled ProjectsabstractIn cybersecurity, vulnerability discovery in source code is a fundamental problem. To automate vulnerability discovery, Machine learning (ML) based techniques has attracted tremendous attention. However, existing ML-based techniques focus on the component or file level detection, and thus considerable human effort is still required to pinpoint the vulnerable code fragments. Using source code files also limit the generalisability of the ML models across projects. To address such challenges, this paper targets at the function-level vulnerability discovery in the cross-project scenario. A function representation learning method is proposed to obtain the high-level and generalizable function representations from the abstract syntax tree (AST). First, the serialized ASTs are used to learn project independence features. Then, a customized bi-directional LSTM neural network is devised to learn the sequential AST representations from the large number of raw features. The new function-level representation demonstrated promising performance gain, using a unique dataset where we manually labeled 6000+ functions from three open-source projects. The results confirm that the huge potential of the new AST-based function representation learning. Guanjun Lin, Jun Zhang 0010, Wei Luo 0001, Lei Pan 0002, Yang Xiang 0001 |
CCS | 2 |
| 2017 | Catch Me If You Can: Detecting Compromised Users Through Partial Observation on NetworksabstractPeople are suffering from a range of risks in the ubiquitous networks of current world, such as rumours spreading in social networks, computer viruses propagating throughout the Internet and unexpected failures happened in Smart grids. We usually monitor only a few users of detecting various risks due to the resource constraints and privacy protection. This leads to a critical problem to detect compromised users who are out of surveillance. In this paper, we propose a risk assessment method to address this problem. The aim is to assess the security status of unmonitored users according to the limited information collected from monitored users in networks. There are two innovative techniques developed: First, we identify the source of risk propagation by inversely disseminating risks from the influenced (by rumours) or infected (by viruses) monitored users. We show a new finding that the ones who synchronously receive the risk copies from all monitored users are most likely to be the sources. Second, we propose a microscopic mathematical model to present the risk propagation from the exposed sources. This model forms a discriminant to classify the compromised users from others. For evaluations, we collect three real networks on which we launch simulated risk propagation and then sample the status of monitored users. The experiment results show that our method is effective and the result of risk assessment well matches the real status of the unmonitored users. Derek Wang, Sheng Wen, Yang Xiang 0001, Wanlei Zhou 0001, Jun Zhang 0010, Surya Nepal |
ICDCS | 5 |
| 2017 | Addressing the class imbalance problem in Twitter spam detection using ensemble learning
Shigang Liu, Yu Wang 0017, Jun Zhang 0010, Chao Chen 0015, Yang Xiang 0001 |
Comput. Secur. | 3 |
| 2017 | Detecting spamming activities in twitter based on deep-learning techniqueabstractSummary Twitter spam has long been a critical but difficult problem to be addressed. So far, researchers have developed a series of machine learning–based methods and blacklisting techniques to detect spamming activities on Twitter. According to our investigation, current methods and techniques have achieved the accuracy of around 87%. However, because of the problems of spam drift and information fabrication, these machine learning–based methods cannot efficiently detect spam activities in real‐life scenarios. Meanwhile, the blacklisting method also cannot catch up with the variations of spamming activities, as manually inspecting suspicious URLs is extremely timeconsuming. In this paper, we proposed a novel technique based on deep‐learning technique to address the above challenges. The syntax of each tweet will be learned through WordVector and trained by deep learning. We then constructed a binary classifier to differentiate spam and regular tweets. In experiments, we collected and labeled a 10‐day real tweet dataset as ground truth to evaluate our proposed method. We first went for empirical analysis with a series of comparisons to other methods: (1) performance of different classifiers, (2) other existing text‐based methods, and (3) nontext‐based detection techniques. According to the experiment results, our proposed method largely outperformed previous methods. We further conducted principle component analysis on typical methods to theoretically justify the outperformance of our method. We extracted all kinds of features via dimensionality reduction. It was found that our features were most distinct among all the detection methods. This well demonstrated the outperformance of our method. Tingmin Wu, Sheng Wen, Shigang Liu, Jun Zhang 0010, Yang Xiang 0001, Majed A. AlRubaian, Mohammad Mehedi Hassan |
Concurr. Comput. Pract. Exp. | 4 |
| 2017 | Investigating the deceptive information in Twitter spam
Chao Chen 0015, Sheng Wen, Jun Zhang 0010, Yang Xiang 0001, Jonathan Oliver, Abdulhameed Alelaiwi, Mohammad Mehedi Hassan |
Future Gener. Comput. Syst. | 3 |
| 2017 | Fuzzy-Based Information Decomposition for Incomplete and Imbalanced Data LearningabstractClass imbalance and missing values are two critical problems in pattern classification. Researchers have proposed a number of techniques to address each of the problems. However, no single technique can solve the two problems. Moreover, the simple combination approach cannot accurately classify the imbalanced data with missing values. This paper develops a fuzzy-based information decomposition (FID) method to simultaneously address these two problems. In the new FID method, the two different problems are treated as the same missing data estimation problem. In particular, FID rebalances the training data by creating synthetic samples for the minority class. The proposed scheme has two steps: weighting and recovery. In the weighting step, the weights produced by the fuzzy membership functions are used to quantify the contribution of the observed data to the missing estimation. In the recovery step, missing values will be estimated by taking into account different contribution of the observed data. To evaluate the performance of the new FID method, a large number of classification experiments have been carried out on 27 well-known datasets. The results show that the FID method significantly outperforms other ten state-of-the-art individual methods and eight combination methods when missing values and imbalanced data present at the same time. Shigang Liu, Jun Zhang 0010, Yang Xiang 0001, Wanlei Zhou 0001 |
IEEE Trans. Fuzzy Syst. | 2 |
| 2017 | Statistical Features-Based Real-Time Detection of Drifted Twitter SpamabstractTwitter spam has become a critical problem nowadays. Recent works focus on applying machine learning techniques for Twitter spam detection, which make use of the statistical features of tweets. In our labeled tweets data set, however, we observe that the statistical properties of spam tweets vary over time, and thus, the performance of existing machine learning-based classifiers decreases. This issue is referred to as “Twitter Spam Drift”. In order to tackle this problem, we first carry out a deep analysis on the statistical features of one million spam tweets and one million non-spam tweets, and then propose a novel Lfun scheme. The proposed scheme can discover “changed” spam tweets from unlabeled tweets and incorporate them into classifier's training process. A number of experiments are performed to evaluate the proposed scheme. The results show that our proposed Lfun scheme can significantly improve the spam detection accuracy in real-world scenarios. Chao Chen 0015, Yu Wang 0017, Jun Zhang 0010, Yang Xiang 0001, Wanlei Zhou 0001, Geyong Min |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2016 | Fuzzy-Based Feature and Instance Recovery
Shigang Liu, Jun Zhang 0010, Yu Wang 0017, Yang Xiang 0001 |
ACIIDS (1) | 2 |
| 2016 | Statistical Detection of Online Drifting Twitter Spam: Invited PaperabstractSpam has become a critical problem in online social networks. This paper focuses on Twitter spam detection. Recent research works focus on applying machine learning techniques for Twitter spam detection, which make use of the statistical features of tweets. We observe existing machine learning based detection methods suffer from the problem of Twitter spam drift, i.e., the statistical properties of spam tweets vary over time. To avoid this problem, an effective solution is to train one twitter spam classifier every day. However, it faces a challenge of the small number of imbalanced training data because labelling spam samples is time-consuming. This paper proposes a new method to address this challenge. The new method employs two new techniques, fuzzy-based redistribution and asymmetric sampling. We develop a fuzzy-based information decomposition technique to re-distribute the spam class and generate more spam samples. Moreover, an asymmetric sampling technique is proposed to re-balance the sizes of spam samples and non-spam samples in the training data. Finally, we apply the ensemble technique to combine the spam classifiers over two different training sets. A number of experiments are performed on a real-world 10-day ground-truth dataset to evaluate the new method. Experiments results show that the new method can significantly improve the detection performance for drifting Twitter spam. Shigang Liu, Jun Zhang 0010, Yang Xiang 0001 |
AsiaCCS | 2 |
| 2016 | Security and reliability in big dataabstractThe purpose of this special issue is to collate a selection of representative research articles that were primarily presented at the 12th IEEE International Conference on Trust, Security and Privacy in Computing and Communications (TrustCom 2013). This annual conference brings together researchers and practitioners in the world from both academia and industry who are working on trusted computing and communications in computer systems and networks, in order to promote an exchange of ideas, discuss future collaborations, and develop new research directions. Yang Xiang 0001, Ivan Stojmenovic, Peter Mueller, Jun Zhang 0010 |
Concurr. Comput. Pract. Exp. | 4 |
| 2016 | Comments and CorrectionsabstractPresents correcttions to the paper, “A performance evaluation of machine learning-based streaming spam tweets detection,” (Chen ], C.; et al) , IEEE Trans. Comput. Social Syst., vol. 2, no. 3, pp. 65–76, Sep. 2015. Chao Chen 0015, Jun Zhang 0010, Yi Xie 0002, Yang Xiang 0001, Wanlei Zhou 0001, Mohammad Mehedi Hassan, Abdulhameed Alelaiwi |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2015 | 6 million spam tweets: A large ground truth for timely Twitter spam detectionabstractTwitter has changed the way of communication and getting news for people's daily life in recent years. Meanwhile, due to the popularity of Twitter, it also becomes a main target for spamming activities. In order to stop spammers, Twitter is using Google SafeBrowsing to detect and block spam links. Despite that blacklists can block malicious URLs embedded in tweets, their lagging time hinders the ability to protect users in real-time. Thus, researchers begin to apply different machine learning algorithms to detect Twitter spam. However, there is no comprehensive evaluation on each algorithms' performance for real-time Twitter spam detection due to the lack of large groundtruth. To carry out a thorough evaluation, we collected a large dataset of over 600 million public tweets. We further labelled around 6.5 million spam tweets and extracted 12 light-weight features, which can be used for online detection. In addition, we have conducted a number of experiments on six machine learning algorithms under various conditions to better understand their effectiveness and weakness for timely Twitter spam detection. We will make our labelled dataset for researchers who are interested in validating or extending our work. Chao Chen 0015, Jun Zhang 0010, Xiao Chen 0002, Yang Xiang 0001, Wanlei Zhou 0001 |
ICC | 2 |
| 2015 | Robust Traffic Classification with Mislabelled Training SamplesabstractTraffic classification plays the significant role in the network security and management. However, accurate classification is challenging if the training data is contaminated with unclean traffic. Recent researches often assume clean training data, and hence performance reduced on real-time network traffic. To meet this challenge, in this paper, we propose a robust method, Unclean Traffic Classification (UTC), which incorporates noise elimination and suspected noise reweighting. Firstly, UTC eliminates strong noisy training data identified by a consensus filtering with multiple classifiers. Furthermore, UTC estimates the relevance of remaining training data and learns a robust traffic classifier. Through a number of experiments on a real-world traffic dataset, we show that the new method outperforms existing state-of-the-art traffic classification methods, under the extremely difficult circumstance with unclean training data. Binfeng Wang, Jun Zhang 0010, Zili Zhang 0001, Wei Luo 0001, Dawen Xia |
ICPADS | 2 |
| 2015 | A Performance Evaluation of Machine Learning-Based Streaming Spam Tweets DetectionabstractThe popularity of Twitter attracts more and more spammers. Spammers send unwanted tweets to Twitter users to promote websites or services, which are harmful to normal users. In order to stop spammers, researchers have proposed a number of mechanisms. The focus of recent works is on the application of machine learning techniques into Twitter spam detection. However, tweets are retrieved in a streaming way, and Twitter provides the Streaming API for developers and researchers to access public tweets in real time. There lacks a performance evaluation of existing machine learning-based streaming spam detection methods. In this paper, we bridged the gap by carrying out a performance evaluation, which was from three different aspects of data, feature, and model. A big ground-truth of over 600 million public tweets was created by using a commercial URL-based security tool. For real-time spam detection, we further extracted 12 lightweight features for tweet representation. Spam detection was then transformed to a binary classification problem in the feature space and can be solved by conventional machine learning algorithms. We evaluated the impact of different factors to the spam detection performance, which included spam to nonspam ratio, feature discretization, training data size, data sampling, time-related data, and machine learning algorithms. The results show the streaming spam tweet detection is still a big challenge and a robust detection technique should take into account the three aspects of data, feature, and model. Chao Chen 0015, Jun Zhang 0010, Yi Xie 0002, Yang Xiang 0001, Wanlei Zhou 0001, Mohammad Mehedi Hassan, Abdulhameed Alelaiwi, Majed A. AlRubaian |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2015 | Camera Model Identification With Unknown ModelsabstractFeature based camera model identification plays an important role for forensics investigations on images. The conventional feature based identification schemes suffer from the problem of unknown models, that is, some images are captured by the camera models previously unknown to the identification system. To address this problem, we propose a new scheme: Source Camera Identification with Unknown models (SCIU). It has the capability of identifying images of the unknown models as well as distinguishing images of the known models. The new SCIU scheme consists of three stages: 1) unknown detection; 2) unknown expansion; and 3) (K+1 )-class classification. Unknown detection applies a k -nearest neighbours method to recognize a few sample images of unknown models from the unlabeled images. Unknown expansion further extends the set of unknown sample images using a self-training strategy. Then, we address a specific (K+1)-class classification, in which the sample images of unknown (1-class) and known models (K-class) are combined to train a classifier. In addition, we develop a parameter optimization method for unknown detection, and investigate the stopping criterion for unknown expansion. The experiments carried out on the Dresden image collection confirm the effectiveness of the proposed SCIU scheme. When unknown models present, the identification accuracy of SCIU is significantly better than the four state-of-art methods: 1) multi-class Support Vector Machine (SVM); 2) binary SVM; 3) combined classification framework; and 4) decision boundary carving. Yonggang Huang 0001, Jun Zhang 0010, Heyan Huang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2015 | Robust Network Traffic ClassificationabstractAs a fundamental tool for network management and security, traffic classification has attracted increasing attention in recent years. A significant challenge to the robustness of classification performance comes from zero-day applications previously unknown in traffic classification systems. In this paper, we propose a new scheme of Robust statistical Traffic Classification (RTC) by combining supervised and unsupervised machine learning techniques to meet this challenge. The proposed RTC scheme has the capability of identifying the traffic of zero-day applications as well as accurately discriminating predefined application classes. In addition, we develop a new method for automating the RTC scheme parameters optimization process. The empirical study on real-world traffic data confirms the effectiveness of the proposed scheme. When zero-day applications are present, the classification performance of the new scheme is significantly better than four state-of-the-art methods: random forest, correlation-based classification, semi-supervised clustering, and one-class SVM. Jun Zhang 0010, Xiao Chen 0002, Yang Xiang 0001, Wanlei Zhou 0001, Jie Wu 0001 |
IEEE/ACM Trans. Netw. | 1 |
| 2014 | On Addressing the Imbalance Problem: A Correlated KNN Approach for Network Traffic Classification
Di Wu 0050, Xiao Chen 0002, Chao Chen 0015, Jun Zhang 0010, Yang Xiang 0001, Wanlei Zhou 0001 |
NSS | 4 |
| 2014 | Internet traffic clustering with side information
Yu Wang 0017, Yang Xiang 0001, Jun Zhang 0010, Wanlei Zhou 0001, Bailin Xie |
J. Comput. Syst. Sci. | 3 |
| 2014 | A noisy-smoothing relevance feedback method for content-based medical image retrieval
Yonggang Huang 0001, Heyan Huang, Jun Zhang 0010 |
Multim. Tools Appl. | 3 |
| 2014 | Medical image retrieval based on unclean image bags
Yonggang Huang 0001, Jun Zhang 0010, Heyan Huang, Daifa Wang |
Multim. Tools Appl. | 2 |
| 2014 | Modeling and Analysis on the Propagation Dynamics of Modern Email MalwareabstractDue to the critical security threats imposed by email-based malware in recent years, modeling the propagation dynamics of email malware becomes a fundamental technique for predicting its potential damages and developing effective countermeasures. Compared to earlier versions of email malware, modern email malware exhibits two new features, reinfection and self-start. Reinfection refers to the malware behavior that modern email malware sends out malware copies whenever any healthy or infected recipients open the malicious attachment. Self-start refers to the behavior that malware starts to spread whenever compromised computers restart or certain files are visited. In the literature, several models are proposed for email malware propagation, but they did not take into account the above two features and cannot accurately model the propagation dynamics of modern email malware. To address this problem, we derive a novel difference equation based analytical model by introducing a new concept of virtual infected user. The proposed model can precisely present the repetitious spreading process caused by reinfection and self-start and effectively overcome the associated computational challenges. We perform comprehensive empirical and theoretical study to validate the proposed analytical model. The results show our model greatly outperforms previous models in terms of estimation accuracy. Sheng Wen, Wei Zhou 0044, Jun Zhang 0010, Yang Xiang 0001, Wanlei Zhou 0001, Weijia Jia 0001, Cliff C. Zou |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2014 | Internet Traffic Classification Using Constrained ClusteringabstractStatistics-based Internet traffic classification using machine learning techniques has attracted extensive research interest lately, because of the increasing ineffectiveness of traditional port-based and payload-based approaches. In particular, unsupervised learning, that is, traffic clustering, is very important in real-life applications, where labeled training data are difficult to obtain and new patterns keep emerging. Although previous studies have applied some classic clustering algorithms such as K-Means and EM for the task, the quality of resultant traffic clusters was far from satisfactory. In order to improve the accuracy of traffic clustering, we propose a constrained clustering scheme that makes decisions with consideration of some background information in addition to the observed traffic statistics. Specifically, we make use of equivalence set constraints indicating that particular sets of flows are using the same application layer protocols, which can be efficiently inferred from packet headers according to the background knowledge of TCP/IP networking. We model the observed data and constraints using Gaussian mixture density and adapt an approximate algorithm for the maximum likelihood estimation of model parameters. Moreover, we study the effects of unsupervised feature discretization on traffic clustering by using a fundamental binning method. A number of real-world Internet traffic traces have been used in our evaluation, and the results show that the proposed approach not only improves the quality of traffic clusters in terms of overall accuracy and per-class metrics, but also speeds up the convergence. Yu Wang 0017, Yang Xiang 0001, Jun Zhang 0010, Wanlei Zhou 0001, Guiyi Wei, Laurence T. Yang |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2013 | AdaM: adaptive-maximum imputation for neighborhood-based collaborative filteringabstractIn the context of collaborative filtering, the well-known data sparsity issue makes two like-minded users have little similarity, and consequently renders the k nearest neighbour rule inapplicable. In this paper, we address the data sparsity problem in the neighbourhood-based CF methods by proposing an Adaptive-Maximum imputation method (AdaM). The basic idea is to identify an imputation area that can maximize the imputation benefit for recommendation purposes, while minimizing the imputation error brought in. To achieve the maximum imputation benefit, the imputation area is determined from both the user and the item perspectives; to minimize the imputation error, there is at least one real rating preserved for each item in the identified imputation area. A theoretical analysis is provided to prove that the proposed imputation method outperforms the conventional neighbourhood-based CF methods through more accurate neighbour identification. Experiment results on benchmark datasets show that the proposed method significantly outperforms the other related state-of-the-art imputation-based methods in terms of accuracy. Yongli Ren, Gang Li 0009, Jun Zhang 0010, Wanlei Zhou 0001 |
ASONAM | 3 |
| 2013 | Robust network traffic identification with unknown applicationsabstractTraffic classification is a fundamental component in advanced network management and security. Recent research has achieved certain success in the application of machine learning techniques into flow statistical feature based approach. However, most of flow statistical feature based methods classify traffic based on the assumption that all traffic flows are generated by the known applications. Considering the pervasive unknown applications in the real world environment, this assumption does not hold. In this paper, we cast unknown applications as a specific classification problem with insufficient negative training data and address it by proposing a binary classifier based framework. An iterative method is proposed to extract unknown information from a set of unlabelled traffic flows, which combines asymmetric bagging and flow correlation to guarantee the purity of extracted negatives. A binary classifier is used as an application signature which can operate on a bag of correlated flows instead of individual flows to further improve its effectiveness. We carry out a series of experiments in a real-world network traffic dataset to evaluate the proposed methods. The results show that the proposed method significantly outperforms the-state-of-art traffic classification methods under the situation of unknown applications present. Jun Zhang 0010, Chao Chen 0015, Yang Xiang 0001, Wanlei Zhou 0001 |
AsiaCCS | 1 |
| 2013 | Network traffic clustering using Random Forest proximitiesabstractThe recent years have seen extensive work on statistics-based network traffic classification using machine learning (ML) techniques. In the particular scenario of learning from unlabeled traffic data, some classic unsupervised clustering algorithms (e.g. K-Means and EM) have been applied but the reported results are unsatisfactory in terms of low accuracy. This paper presents a novel approach for the task, which performs clustering based on Random Forest (RF) proximities instead of Euclidean distances. The approach consists of two steps. In the first step, we derive a proximity measure for each pair of data points by performing a RF classification on the original data and a set of synthetic data. In the next step, we perform a K-Medoids clustering to partition the data points into K groups based on the proximity matrix. Evaluations have been conducted on real-world Internet traffic traces and the experimental results indicate that the proposed approach is more accurate than the previous methods. Yu Wang 0017, Yang Xiang 0001, Jun Zhang 0010 |
ICC | 3 |
| 2013 | Clonewise - Detecting Package-Level Clones Using Machine Learning
Silvio Cesare, Yang Xiang 0001, Jun Zhang 0010 |
SecureComm | 3 |
| 2013 | Robust image retrieval with hidden classes
Jun Zhang 0010, Lei Ye 0002, Yang Xiang 0001, Wanlei Zhou 0001 |
Comput. Vis. Image Underst. | 1 |
| 2013 | Unsupervised traffic classification using flow statistical properties and IP packet payload
Jun Zhang 0010, Yang Xiang 0001, Wanlei Zhou 0001, Yu Wang 0017 |
J. Comput. Syst. Sci. | 1 |
| 2013 | Lazy Collaborative Filtering for Data Sets With Missing ValuesabstractAs one of the biggest challenges in research on recommender systems, the data sparsity issue is mainly caused by the fact that users tend to rate a small proportion of items from the huge number of available items. This issue becomes even more problematic for the neighborhood-based collaborative filtering (CF) methods, as there are even lower numbers of ratings available in the neighborhood of the query item. In this paper, we aim to address the data sparsity issue in the context of neighborhood-based CF. For a given query (user, item), a set of key ratings is first identified by taking the historical information of both the user and the item into account. Then, an auto-adaptive imputation (AutAI) method is proposed to impute the missing values in the set of key ratings. We present a theoretical analysis to show that the proposed imputation method effectively improves the performance of the conventional neighborhood-based CF methods. The experimental results show that our new method of CF with AutAI outperforms six existing recommendation methods in terms of accuracy. Yongli Ren, Gang Li 0009, Jun Zhang 0010, Wanlei Zhou 0001 |
IEEE Trans. Cybern. | 3 |
| 2013 | Internet Traffic Classification by Aggregating Correlated Naive Bayes PredictionsabstractThis paper presents a novel traffic classification scheme to improve classification performance when few training data are available. In the proposed scheme, traffic flows are described using the discretized statistical features and flow correlation information is modeled by bag-of-flow (BoF). We solve the BoF-based traffic classification in a classifier combination framework and theoretically analyze the performance benefit. Furthermore, a new BoF-based traffic classification method is proposed to aggregate the naive Bayes (NB) predictions of the correlated flows. We also present an analysis on prediction error sensitivity of the aggregation strategies. Finally, a large number of experiments are carried out on two large-scale real-world traffic datasets to evaluate the proposed scheme. The experimental results show that the proposed scheme can achieve much better classification performance than existing state-of-the-art traffic classification methods. Jun Zhang 0010, Chao Chen 0015, Yang Xiang 0001, Wanlei Zhou 0001, Yong Xiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2013 | An Effective Network Traffic Classification Method with Unknown Flow DetectionabstractTraffic classification technique is an essential tool for network and system security in the complex environments such as cloud computing based environment. The state-of-the-art traffic classification methods aim to take the advantages of flow statistical features and machine learning techniques, however the classification performance is severely affected by limited supervised information and unknown applications. To achieve effective network traffic classification, we propose a new method to tackle the problem of unknown applications in the crucial situation of a small supervised training set. The proposed method possesses the superior capability of detecting unknown flows generated by unknown applications and utilizing the correlation information among real-world network traffic to boost the classification performance. A theoretical analysis is provided to confirm performance benefit of the proposed method. Moreover, the comprehensive performance evaluation conducted on two real-world network traffic datasets shows that the proposed scheme outperforms the existing methods in the critical network environment. Jun Zhang 0010, Chao Chen 0015, Yang Xiang 0001, Wanlei Zhou 0001, Athanasios V. Vasilakos |
IEEE Trans. Netw. Serv. Manag. | 1 |
| 2013 | Modeling Propagation Dynamics of Social Network WormsabstractSocial network worms, such as email worms and facebook worms, pose a critical security threat to the Internet. Modeling their propagation dynamics is essential to predict their potential damages and develop countermeasures. Although several analytical models have been proposed for modeling propagation dynamics of social network worms, there are two critical problems unsolved: temporal dynamics and spatial dependence. First, previous models have not taken into account the different time periods of Internet users checking emails or social messages, namely, temporal dynamics. Second, the problem of spatial dependence results from the improper assumption that the states of neighboring nodes are independent. These two problems seriously affect the accuracy of the previous analytical models. To address these two problems, we propose a novel analytical model. This model implements a spatial-temporal synchronization process, which is able to capture the temporal dynamics. Additionally, we find the essence of spatial dependence is the spreading cycles. By eliminating the effect of these cycles, our model overcomes the computational challenge of spatial dependence and provides a stronger approximation to the propagation dynamics. To evaluate our susceptible-infectious-immunized (SII) model, we conduct both theoretical analysis and extensive simulations. Compared with previous epidemic models and the spatial-temporal model, the experimental results show our SII model achieves a greater accuracy. We also compare our model with the susceptible-infectious-susceptible and susceptible-infectious- recovered models. The results show that our model is more suitable for modeling the propagation of social network worms. Sheng Wen, Wei Zhou 0044, Jun Zhang 0010, Yang Xiang 0001, Wanlei Zhou 0001, Weijia Jia 0001 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2013 | Network Traffic Classification Using Correlation InformationabstractTraffic classification has wide applications in network management, from security monitoring to quality of service measurements. Recent research tends to apply machine learning techniques to flow statistical feature based classification methods. The nearest neighbor (NN)-based method has exhibited superior classification performance. It also has several important advantages, such as no requirements of training procedure, no risk of overfitting of parameters, and naturally being able to handle a huge number of classes. However, the performance of NN classifier can be severely affected if the size of training data is small. In this paper, we propose a novel nonparametric approach for traffic classification, which can improve the classification performance effectively by incorporating correlated information into the classification process. We analyze the new classification approach and its performance benefit from both theoretical and empirical perspectives. A large number of experiments are carried out on two real-world traffic data sets to validate the proposed approach. The results show the traffic classification performance can be improved significantly even under the extreme difficult circumstance of very few training samples. Jun Zhang 0010, Yang Xiang 0001, Yu Wang 0017, Wanlei Zhou 0001, Yong Xiang 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2012 | The efficient imputation method for neighborhood-based collaborative filteringabstractAs each user tends to rate a small proportion of available items, the resulted Data Sparsity issue brings significant challenges to the research of recommender systems. This issue becomes even more severe for neighborhood-based collaborative filtering methods, as there are even lower numbers of ratings available in the neighborhood of the query item. In this paper, we aim to address the Data Sparsity issue in the context of the neighborhood-based collaborative filtering. Given the (user, item) query, a set of key ratings are identified, and an auto-adaptive imputation method is proposed to fill the missing values in the set of key ratings. The proposed method can be used with any similarity metrics, such as the Pearson Correlation Coefficient and Cosine-based similarity, and it is theoretically guaranteed to outperform the neighborhood-based collaborative filtering approaches. Results from experiments prove that the proposed method could significantly improve the accuracy of recommendations for neighborhood-based Collaborative Filtering algorithms. Yongli Ren, Gang Li 0009, Jun Zhang 0010, Wanlei Zhou 0001 |
CIKM | 3 |
| 2012 | Internet traffic clustering with constraintsabstractDue to the limitations of the traditional port-based and payload-based traffic classification approaches, the past decade has seen extensive work on utilizing machine learning techniques to classify network traffic based on packet and flow level features. In particular, previous studies have shown that the unsupervised clustering approach is both accurate and capable of discovering previously unknown application classes. In this paper, we explore the utility of side information in the process of traffic clustering. Specifically, we focus on the flow correlation information that can be efficiently extracted from packet headers and expressed as instance-level constraints, which indicate that particular sets of flows are using the same application and thus should be put into the same cluster. To incorporate the constraints, we propose a modified constrained K-Means algorithm. A variety of real-world traffic traces are used to show that the constraints are widely available. The experimental results indicate that the constrained approach not only improves the quality of the resulted clusters, but also speeds up the convergence of the clustering process. Yu Wang 0017, Yang Xiang 0001, Jun Zhang 0010, Shunzheng Yu |
IWCMC | 3 |
| 2012 | Classification of Correlated Internet Traffic FlowsabstractA critical problem for Internet traffic classification is how to obtain a high-performance statistical feature based classifier using a small set of training data. The solutions to this problem are essential to deal with the encrypted applications and the new emerging applications. In this paper, we propose a new Naive Bayes (NB) based classification scheme to tackle this problem, which utilizes two recent research findings, feature discretization and flow correlation. A new bag-of-flow (BoF) model is firstly introduced to describe the correlated flows and it leads to a new BoF-based traffic classification problem. We cast the BoF-based traffic classification as a specific classifier combination problem and theoretically analyze the classification benefit from flow aggregation. A number of combination methods are also formulated and used to aggregate the NB predictions of the correlated flows. Finally, we carry out a number of experiments on a large scale real-world network dataset. The experimental results show that the proposed scheme can achieve significantly higher classification accuracy and much faster classification speed with comparison to the state-of-the-art traffic classification methods. Jun Zhang 0010, Chao Chen 0015, Yang Xiang 0001, Wanlei Zhou 0001 |
TrustCom | 1 |
| 2011 | A novel semi-supervised approach for network traffic clusteringabstractNetwork traffic classification is an essential component for network management and security systems. To address the limitations of traditional port-based and payload-based methods, recent studies have been focusing on alternative approaches. One promising direction is applying machine learning techniques to classify traffic flows based on packet and flow level statistics. In particular, previous papers have illustrated that clustering can achieve high accuracy and discover unknown application classes. In this work, we present a novel semi-supervised learning method using constrained clustering algorithms. The motivation is that in network domain a lot of background information is available in addition to the data instances themselves. For example, we might know that flow f1and f2are using the same application protocol because they are visiting the same host address at the same port simultaneously. In this case, f1and f2shall be grouped into the same cluster ideally. Therefore, we describe these correlations in the form of pair-wise must-link constraints and incorporate them in the process of clustering. We have applied three constrained variants of the K-Means algorithm, which perform hard or soft constraint satisfaction and metric learning from constraints. A number of real-world traffic traces have been used to show the availability of constraints and to test the proposed approach. The experimental results indicate that by incorporating constraints in the course of clustering, the overall accuracy and cluster purity can be significantly improved. Yu Wang 0017, Yang Xiang 0001, Jun Zhang 0010, Shunzheng Yu |
NSS | 3 |
| 2011 | Secure Image Retrieval Based on Visual Content and Watermarking ProtocolabstractAs an interesting application on cloud computing, content-based image retrieval (CBIR) has attracted a lot of attention, but the focus of previous research work was mainly on improving the retrieval performance rather than addressing security issues such as copyrights and user privacy. With an increase of security attacks in the computer networks, these security issues become critical for CBIR systems. In this paper, we propose a novel two-party watermarking protocol that can resolve the issues regarding user rights and privacy. Unlike the previously published protocols, our protocol does not require the existence of a trusted party. It exhibits three useful features: security against partial watermark removal, security in watermark verification and non-repudiation. In addition, we report an empirical research of CBIR with the security mechanism. The experimental results show that the proposed protocol is practicable and the retrieval performance will not be affected by watermarking query images. Jun Zhang 0010, Yang Xiang 0001, Wanlei Zhou 0001, Lei Ye 0002, Yi Mu 0001 |
Comput. J. | 1 |
| 2010 | Effective watermarking scheme in the encrypted domain for buyer-seller watermarking protocol
Weidong Kou, Hui Li 0006, Lanjun Dang, Jun Zhang 0010 |
Inf. Sci. | 5 |
| 2009 | Image retrieval based on bag of imagesabstractConventional relevance feedback schemes may not be suitable to all practical applications of content-based image retrieval (CBIR), since most ordinary users would like to complete their search in a single interaction, especially on the web search. In this paper, we explore a new approach to improve the retrieval performance based on a new concept, bag of images, rather than relevance feedback. We consider that image collection comprises of image bags instead of independent individual images. Each image bag includes some relevant images with the same perceptual meaning. A theoretical case study demonstrates that image retrieval can benefit from the new concept. A number of experimental results show that the CBIR scheme based on bag of images can improve the retrieval performance dramatically. Jun Zhang 0010, Lei Ye 0002 |
ICIP | 1 |
| 2009 | Image retrieval using noisy queryabstractIn conventional content based image retrieval (CBIR) employing relevance feedback, one implicit assumption is that both pure positive and negative examples are available. However it is not always true in the practical applications of CBIR. In this paper, we address a new problem of image retrieval using several unclean positive examples, named noisy query, in which some mislabeled images or weak relevant images present. The proposed image retrieval scheme measures the image similarity by combining multiple feature distances. Incorporating data cleaning and noise tolerant classifier, a two-step strategy is proposed to handle noisy positive examples. Experiments carried out on a subset of corel image collection show that the proposed scheme outperforms the competing image retrieval schemes. Jun Zhang 0010, Lei Ye 0002 |
ICME | 1 |
| 2009 | Watermarking protocol for protecting user's right in content based image retrievalabstractContent based image retrieval (CBIR) is a technique to search for images relevant to the user's query from an image collection. In last decade, most attention has been paid to improve the retrieval performance. However, there is no significant effort to investigate the security concerning in CBIR. Under the query by example (QBE) paradigm, the user supplies an image as a query and the system returns a set of retrieved results. If the query image includes user's private information, an untrusted server provider of CBIR may distribute it illegally, which leads to the user's right problem. In this paper, we propose an interactive watermarking protocol to address this problem. A watermark is inserted into the query image by the user in encrypted domain without knowing the exact content. The server provider of CBIR will get the watermarked query image and uses it to perform image retrieval. In case where the user finds an unauthorized copy, a watermark in the unauthorized copy will be used as evidence to prove that the user's legal right is infringed by the server provider. Jun Zhang 0010, Lei Ye 0002 |
ICME | 1 |
| 2009 | Local aggregation function learning based on support vector machines
Jun Zhang 0010, Lei Ye 0002 |
Signal Process. | 1 |
| 2009 | Content Based Image Retrieval Using Unclean Positive ExamplesabstractConventional content-based image retrieval (CBIR) schemes employing relevance feedback may suffer from some problems in the practical applications. First, most ordinary users would like to complete their search in a single interaction especially on the web. Second, it is time consuming and difficult to label a lot of negative examples with sufficient variety. Third, ordinary users may introduce some noisy examples into the query. This correspondence explores solutions to a new issue that image retrieval using unclean positive examples. In the proposed scheme, multiple feature distances are combined to obtain image similarity using classification technology. To handle the noisy positive examples, a new two-step strategy is proposed by incorporating the methods of data cleaning and noise tolerant classifier. The extensive experiments carried out on two different real image collections validate the effectiveness of the proposed scheme. Jun Zhang 0010, Lei Ye 0002 |
IEEE Trans. Image Process. | 1 |
| 2007 | A Watermarking Scheme in the Encrypted Domain for Watermarking Protocol
Lanjun Dang, Weidong Kou, Jun Zhang 0010, Zan Li 0001, Kai Fan 0001 |
Inscrypt | 4 |
| 2007 | An Unified Framework Based on p-Norm for Feature Aggregation in Content-Based Image RetrievalabstractFeature aggregation is a critical technique in content- based image retrieval systems that employ multiple visual features to characterize image content. In this paper, the p-norm is introduced to feature aggregation that provides a framework to unify various previous feature aggregation schemes such as linear combination, Euclidean distance, Boolean logic and decision fusion schemes in which previous schemes are instances. Some insights of the mechanism of how various aggregation schemes work are discussed through the effects of model parameters in the unified framework. Experiments show that performances vary over feature aggregation schemes that necessitates an unified framework in order to optimize the retrieval performance according to individual queries and user query concept. Revealing experimental results conducted with IAPR TC-12 ImageCLEF2006 benchmark collection that contains over 20,000 photographic images are presented and discussed. Jun Zhang 0010, Lei Ye 0002 |
ISM | 1 |