EDBT 2026 Demo / reviewers in the wild / expert
Chao Chen 0015
dblp:66/3019-15
· DBLP profile ↗
57ranked-venue papers
5as first author
37since 2021 · last 2026
0000-0003-1355-3870ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 28 · 1 first-author · 18 since 2021Artificial intelligence and machine learning · 8 · 8 since 2021Systems, architecture and hardware · 6 · 1 first-author · 2 since 2021Computer networks · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Improved Streaming Algorithm for Fair k-Center ClusteringabstractMany real-world applications call for incorporating fairness constraints into the k-center clustering problem, where the dataset is partitioned into m demographic groups, each with a specified upper bound on the number of centers to ensure fairness. Focusing on big data scenarios, this paper addresses the problem in a streaming setting, where data points arrive sequentially in a continuous stream. Leveraging a structure called the λ-independent center set, we propose a one-pass streaming algorithm that first computes a reserved set of points during the streaming process. In the post-streaming process, we then select centers from the reserved point set by analyzing three possible cases and transforming the most complex one into a specially constrained vertex-cover problem on an auxiliary graph. Our algorithm achieves an approximation ratio of 5 + ? and memory complexity O(k log ?), where ? is the aspect ratio and ? > 0 is any small constant. Furthermore, we extend our approach to semi-structured data streams, where data points arrive in groups. In this setting, we present a (3 + ?)-approximation algorithm for m = 2, which can be readily adapted to solve the offline fair k-center problem, achieving an approximation ratio of 3 that matches the current state of the art. Lastly, we conduct extensive experiments to evaluate the performance of our approaches, demonstrating that they outperform existing baselines in both clustering cost and runtime efficiency. Longkun Guo, Zeyu Lin, Chaoqi Jia, Chao Chen 0015 |
AAAI | 4 |
| 2026 | Approximation Algorithm for Constrained k-Center Clustering: A Local Search ApproachabstractClustering is a long-standing research problem and a fundamental tool in AI and data analysis. The traditional k-center problem, known as a fundamental theoretical challenge in clustering, has a best possible approximation ratio of 2, and any improvement to a ratio of 2 - ε would imply P = NP. In this work, we study the constrained k-center clustering problem, where instance-level cannot-link (CL) and must-link (ML) constraints are incorporated as background knowledge. Although general CL constraints significantly increase the hardness of approximation, previous work has shown that disjoint CL sets permit constant-factor approximations. However, whether local search can achieve such a guarantee in this setting remains an open question. To this end, we propose a novel local search framework based on a transformation to a dominating matching set problem, achieving the best possible approximation ratio of 2. The experimental results on both real-world and synthetic datasets demonstrate that our algorithm outperforms baselines in solution quality. Chaoqi Jia, Longkun Guo, Kewen Liao, Zhigang Lu 0001, Chao Chen 0015, Minhui Xue 0001 |
AAAI | 5 |
| 2026 | Optimized Algorithms for Text Clustering with LLM-Generated ConstraintsabstractClustering is a fundamental tool that has garnered significant interest across a wide range of applications including text analysis. To improve clustering accuracy, many researchers have proposed incorporating background knowledge, typically in the form of must‑link and cannot‑link constraints, to guide the clustering process. With the recent advent of large language models (LLMs), there is growing interest in improving clustering quality through LLM-based automatic constraint generation. In this paper, we propose a novel constraint‑generation approach that reduces resource consumption by generating constraint sets rather than using traditional pairwise constraints. This improves both query efficiency and constraint accuracy compared to state‑of‑the‑art methods. We further introduce a constrained clustering algorithm tailored to the characteristics of LLM-generated constraints. Our method incorporates a confidence threshold and a penalty mechanism to address potentially inaccurate constraints. We evaluate our approach on five text datasets, considering both the cost of constraint generation and overall clustering performance. The results show that our method achieves clustering accuracy comparable to the state-of-the-art algorithms while reducing the number of LLM queries by more than 20 times. Chaoqi Jia, Weihong Wu, Longkun Guo, Zhigang Lu 0001, Chao Chen 0015, Kok-Leong Ong |
AAAI | 5 |
| 2026 | SoK: Telemetry-Aware Runtime Assurance for Always-On On-device Intrusion Detection
Nuonan Ouyang, Adrian Shatte, Zhigang Lu 0001, Chao Chen 0015, Wei Xiang 0001 |
ACISP (3) | 4 |
| 2026 | Fair k-Center Clustering on Massive Social Network Data StreamsabstractAs a fundamental technique with many real-world applications, including social network analysis, center-based clustering may inadvertently discriminate against certain populations based on factors such as age, gender, or socioeconomic status, particularly when nodes are associated with sensitive attributes. In this work, we study the problem of fair k-center clustering in the streaming setting, which seeks to select representative items from a large data stream while respecting group-representation fairness. Given an input dataset in Euclidean space partitioned into m disjoint groups, the fairness constraint requires that the number of centers selected from each group satisfies a given upper bound. Moreover, the problem aims to select a set of centers that minimizes the maximum distance from any point to its nearest center (the k-center objective) while satisfying the fairness constraint. We present a one-pass streaming algorithm with approximation ratio 4.46, improving the previous best ratio of (5+?) for this problem in general metrics. Notably, our result establishes that streaming fair k-center admits a strictly better approximation ratio in Euclidean space than in general metrics, in contrast to the standard k-center problem, whose best-known approximation ratio is 2 in both Euclidean and general metric spaces. Finally, we complement our theoretical results with an empirical evaluation on five real-world social network datasets and million-scale synthetic datasets, demonstrating significant improvements over state-of-the-art methods in clustering quality while maintaining comparable runtime efficiency. Longkun Guo, Chaoqi Jia, Chao Chen 0015 |
WWW | 3 |
| 2026 | Artificial Intelligence in Mitigating Security Threats for Lightweight IoT Devices: A Survey of Technologies, Protocols, and Future ChallengesabstractLightweight Internet of Things (IoT) devices—microcontroller-class nodes with less than 512KB RAM, sub-100MHz clocks, and low-power radios (BLE, Zigbee, LoRa, NB-IoT)—are now widely deployed in settings where traditional security stacks are infeasible. This survey examines how Artificial Intelligence (AI) can harden such constrained platforms against device-, network-, and application-layer threats, including spoofing, routing manipulation, DDoS, malware, and Advanced Persistent Threats (APTs). We (i) formalize alightweight envelopethat bounds feasible defenses in terms of RAM, CPU, bandwidth, and energy; (ii) consolidate protocol-side risks across BLE, Zigbee, and LoRaWAN; and (iii) review deployable AI techniques through adeployment-firstlens that separates training (edge, cloud, federated learning) from on-device inference. Distinct from prior surveys, we provide resource-annotated comparisons that report accuracyalongsidemodel size, peak RAM, latency, and estimated energy per inference, showing how pruning, post-training quantization, distillation, and feature narrowing shift feasibility on MCU targets. Covered methods include compact classifiers (linear models, trees, SVM), quantized TinyCNN/TinyRNN and graph-based intrusion detection, reinforcement learning for adaptive rate limiting and channel selection, and privacy-preserving federated learning with update compression. We conclude with a pragmatic agenda—energy-adaptive inference, LPWAN-aware scheduling and federated learning, robustness to poisoning and evasion, and reproducible benchmarks that couple accuracy with size/latency/energy on real hardware—aimed at making AI-based security practical at scale for lightweight IoT deployments. Nuonan Ouyang, Adrian Shatte, Zhigang Lu 0001, Chao Chen 0015, Wei Xiang 0001 |
IEEE Internet Things J. | 4 |
| 2026 | Lightweight multi-client order-revealing encryption with limited leakage
Chunyang Lv, Jianfeng Wang 0001, Shifeng Sun 0001, Saiyu Qi, Chao Chen 0015, Leo Yu Zhang, Kok-Leong Ong |
Inf. Sci. | 5 |
| 2026 | Fine-Grained Poisoning Framework Against Federated LearningabstractFederated learning(FL) is one of the most widely used distributed machine learning frameworks. However, FL is susceptible to poisoning attacks that can degrade the quality of the global model. Recent studies on fine-grained poisoning attacks highlight a strategic shift where attackers no longer prioritize maximal disruption of the global model, but instead control the degree of model poisoning to maintain stealth and avoid detection. However, research on fine-grained poisoning is still in the infant stage. Numerous fundamental questions have yet to be addressed, including its underlying mechanisms and optimization strategies. To this end, we introduces FGP, the first comprehensive framework forFine-GrainedPoisoning on FL, which allows adversaries to precisely manipulate the global model by strategically inducing accurate and stealthy sub-optimal solution. Fundamentally, FGP innovatively formalizes fine-grained attacks as an optimization problem to minimize the distance between the current global model and the adversary's target (a sub-optimal solution). It then employs a real-time search strategy to dynamically refine the malicious model updates in each round. To ensure optimal attack performance, we further introduce a novel topology-based approach as the error feedback. Additionally, we present a formal convergence analysis of our attacks. Armed with FGP, we conduct a comprehensive evaluation of FL's robustness against fine-grained poisoning across diverse settings. Results demonstrate that FGP significantly outperforms the prior work, achieving an average$6.5\times$higher attack accuracy. Hangtao Zhang, Yanjun Zhang 0002, Chao Chen 0015, Qiyun Shao, Shengshan Hu, Leo Yu Zhang |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2025 | FLAME: Flexible and Lightweight Biometric Authentication Scheme in Malicious EnvironmentsabstractPrivacy-preserving biometric authentication (PPBA) enables client authentication without revealing sensitive bio-metric data, addressing privacy and security concerns. Many studies have proposed efficient cryptographic solutions to this problem based on secure multi-party computation, typically assuming a semi-honest adversary model, where all parties follow the protocol but may try to learn additional information. However, this assumption often falls short in real-world scenarios, where adversaries may behave maliciously and actively deviate from the protocol. In this paper, we propose, implement, and evaluate FLAME, a Flexible and Lightweight biometric Authentication scheme designed for a Malicious Environment. By hybridizing lightweight secret-sharing-family primitives within two-party computation, FLAME carefully designs a line of supporting protocols that incorporate integrity checks with rationally extra overhead. Additionally, FLAME enables server-side authentication with various similarity metrics through a crossmetric-compatible design, enhancing flexibility and robustness without requiring any changes to the server-side process. A rigorous theoretical analysis validates the correctness, security, and efficiency of FLAME. Extensive experiments highlight FLAME's superior efficiency, with a communication reduction by 97.61x 110.13x and a speedup of 2.72x 2.82x (resp. 6.58x 8.51x) in a LAN (resp. WAN) environment, when compared to the state-of-the-art work. Fuyi Wang, Fangyuan Sun, Mingyuan Fan 0003, Jianying Zhou 0001, Chao Chen 0015, Jiangang Shu, Leo Yu Zhang |
ACSAC | 6 |
| 2025 | Poster: The Art of Deception: Crafting Chimera Images for Covert and Robust Semantic Poisoning AttacksabstractWith the exponential surge in media data volumes and their growing intrinsic value, the landscape has become increasingly susceptible to persistent and strategically designed data poisoning attacks targeting these valuable assets. In this work, we propose a novel approach leveraging generative AI techniques to craft covert and robust poisonous data samples, referred to as Chimera Images. These images seamlessly blend visual features from two target classes to generate hybrid objects that preserve appearance fidelity. These ''normal'' samples with correct labels can subtly distort the model's decision boundary without raising suspicion. Extensive experimental results on CIFAR-10 and Flowers datasets demonstrate that the proposed method i) reduces the accuracy of the targeted class, ii) maintains the performance of other classes, and iii) exhibits immunity to state-of-the-art defence strategies. We also explore the usage of generative AI content detection as a defence mechanism, demonstrating that the recently discovered snapshot technique is ineffective against the AI-generated poisonous Chimera samples. Lin Li 0066, Youyang Qu, Jiayang Ao, Ming Ding 0001, Chao Chen 0015, Jun Zhang 0010 |
CCS | 5 |
| 2025 | Standardizing the evaluation framework for ECG-based authentication in IoT devicesabstractDevices on the Internet of Things (IoT) often have constrained resources and operate in diverse environments, making them vulnerable to unauthorized access and cyber threats. Electrocardiogram (ECG) signals have emerged as a promising biometric for authenticating users in such settings. However, current ECG-based authentication studies lack a standardized evaluation framework tailored to resource-limited IoT contexts and long-term usage, making it difficult to assess their practical reliability. In this paper, we introduce a new evaluation framework for ECG-based authentication on IoT devices and construct a standardized dataset to facilitate rigorous testing. We categorize performance metrics into four key dimensions: scalability, adaptability, efficiency, and cancelability. Using this framework, we evaluate four representative ECG authentication algorithms for IoT devices. The results show that these algorithms struggle to maintain consistent performance under cross-session authentication scenarios. These findings highlight the critical importance of addressing the temporal variability of ECG signals and the current gap in robust ECG-based authentication for IoT devices. We believe the proposed framework will guide future research toward more resilient and secure ECG authentication systems for the IoT. Bonan Zhang, Lin Li 0066, Chao Chen 0015, Ickjai Lee, Kyungmi Lee, Kok-Leong Ong |
Comput. Commun. | 3 |
| 2025 | A survey on security and privacy issues in wearable health monitoring devicesabstractRecent developments in mobile computing power and wireless communication speeds have significantly improved the efficiency of medical systems. This paper focuses on passive wearable sensor devices, which are integral to noninvasive monitoring of physiological data in healthcare observation. Beyond data collection, some wearables play an active role in patient treatment, underscoring the critical importance of protecting their security and privacy. Breach in these areas can severely affect patient health. However, the distinctive characteristics of wearable technologies introduce unique security and privacy challenges, including the potential for unauthorized access to sensitive location, medical, and physiological data. This review delves into the security and privacy concerns associated with wearable devices and proposes potential remedies. Its value lies in providing insights for researchers and manufacturers, aiming to advance the development of safer and more effective wearable medical technologies. Bonan Zhang, Chao Chen 0015, Ickjai Lee, Kyungmi Lee, Kok-Leong Ong |
Comput. Secur. | 2 |
| 2025 | A novel dictionary attack on ECG authentication system using adversarial optimization and clusteringabstractElectrocardiogram(ECG)-based biometric authentication has become a promising method to improve security in wearable devices due to its inherent uniqueness and difficulty to replicate. However, no studies currently demonstrate that ECG authentication can resist modern attack techniques employed against biometric authentication. In this paper, we present a novel dictionary attack against ECG authentication systems, which poses a significant threat. In contrast to conventional targeted attacks, this approach utilizes random pairing to breach a vast number of users, without requiring specific information about their biometric data. Our approach leverages adversarial optimization and clustering to generate synthetic ECG waveforms capable of bypassing authentication mechanisms of various systems, revealing critical vulnerabilities in the current implementation of ECG-based biometrics. We comprehensively evaluate the effectiveness of this attack across different ECG authentication models, demonstrating that despite the intrinsic uniqueness of ECG signals, a substantial number of users are vulnerable. Our attack method can bypass the authentication system of an average of 20% of users even at the most stringent false acceptance rate of 1%. With up to five attack attempts allowed, our method can bypass up to 62% of users’ ECG authentication models. Bonan Zhang, Lin Li 0066, Chao Chen 0015, Ickjai Lee, Kyungmi Lee, Tianqing Zhu, Kok-Leong Ong |
Knowl. Based Syst. | 3 |
| 2025 | Extracting Private Training Data in Federated Learning From ClientsabstractThe utilization of machine learning algorithms in distributed web applications is experiencing significant growth. One notable approach is Federated Learning (FL) Recent research has brought attention to the vulnerability of FL to gradient inversion attacks, which seek to reconstruct the original training samples, posing a substantial threat to client privacy. Most existing gradient inversion attacks, however, require control over the central server and rely on substantial prior knowledge, including information about batch normalization and data distribution. In this study, we introduce Poisoning Gradient Leakage from Client (PGLC), a novel attack method that operates from the clients’ side. For the first time, we demonstrate the feasibility of a client-side adversary with limited knowledge successfully recovering training samples from the aggregated global model. Our approach enables the adversary to employ a malicious model that increases the loss of a specific targeted class of interest. When honest clients employ the poisoned global model, the gradients of samples become distinct in the aggregated update. This allows the adversary to effectively reconstruct private inputs from other clients using the aggregated update. Furthermore, our PGLC attack exhibits stealthiness against Byzantine-robust aggregation rules (AGRs). Through the optimization of malicious updates and the blending of benign updates with a malicious replacement vector, our method remains undetected by these defense mechanisms. We conducted experiments across various benchmark datasets, considering representative Byzantine-robust AGRs and exploring different FL settings with varying levels of adversary knowledge about the data. Our results consistently demonstrate the ability of PGLC to extract training data in all tested scenarios. Jiaheng Wei, Yanjun Zhang 0002, Leo Yu Zhang, Chao Chen 0015, Shirui Pan, Kok-Leong Ong, Jun Zhang 0010, Yang Xiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2024 | Stealing Watermarks of Large Language Models via Mixed Integer ProgrammingabstractThe Large Language Model (LLM) watermark is a newly emerging technique that shows promise in addressing concerns surrounding LLM copyright, monitoring AI-generated text, and preventing its misuse. The LLM watermark scheme commonly includes generating secret keys to partition the vocabulary into green and red lists, applying a perturbation to the logits of tokens in the green list to increase their sampling likelihood, thus facilitating watermark detection to identify AI-generated text if the proportion of green tokens exceeds a threshold. However, recent research indicates that watermarking methods using numerous keys are susceptible to removal attacks, such as token editing, synonym substitution, and paraphrasing, with robustness declining as the number of keys increases. Therefore, the state-of-the-art watermark schemes that employ fewer or single keys have been demonstrated to be more robust against text editing and paraphrasing. In this paper, we propose a novel green list stealing attack against the state-of-the-art LLM watermark scheme and systematically examine its vulnerability to this attack. We formalize the attack as a mixed integer programming problem with constraints. We evaluate our attack under a comprehensive threat model, including an extreme scenario where the attacker has no prior knowledge, lacks access to the watermark detector API, and possesses no information about the LLM’s parameter settings or watermark injection/detection scheme. Extensive experiments on LLMs, such as OPT and LLaMA, demonstrate that our attack can successfully steal the green list and remove the watermark across all settings. Zhaoxi Zhang 0001, Xiaomei Zhang 0001, Yanjun Zhang 0002, Leo Yu Zhang, Chao Chen 0015, Shengshan Hu, Asif Gill, Shirui Pan |
ACSAC | 5 |
| 2024 | Exploring the Vulnerability of ECG-Based Authentication Systems Through A Dictionary Attack Approach
Bonan Zhang, Chao Chen 0015, Ickjai Lee, Kyungmi Lee, Kok-Leong Ong |
ICA3PP (6) | 2 |
| 2024 | Detector Collapse: Backdooring Object Detection to Catastrophic Overload or Blindness in the Physical World
Hangtao Zhang, Shengshan Hu, Yichen Wang 0013, Leo Yu Zhang, Ziqi Zhou 0001, Xianlong Wang 0001, Yanjun Zhang 0002, Chao Chen 0015 |
IJCAI | 8 |
| 2024 | Clustering-based Evaluation Framework of Feature Extraction Approaches for ECG Biometric AuthenticationabstractIn recent times, electrocardiogram signals have been leveraged for biometric verification. The efficacy of such authentication is reliant on the feature extraction from the electrocardiogram signals. A number of electrocardiogram feature extraction methods are currently available, but these methods may not be universally applicable in different dataset collection scenarios. To tackle this issue, this paper introduces a clustering-based framework to assess the feature extraction techniques for electrocardiogram biometrics. In this paper, the effectiveness of the framework is validated by using different electrocardiogram feature extraction techniques and different electrocardiogram databases. The framework provides important insights into electrocardiogram signal. Bonan Zhang, Chao Chen 0015, Ickjai Lee, Kyungmi Lee, Kok-Leong Ong |
IJCNN | 2 |
| 2024 | VulMatch: Binary-Level Vulnerability Detection Through Signature
Zian Liu, Shigang Liu, Lei Pan 0002, Chao Chen 0015, Jun Zhang 0010, Dongxi Liu |
NSS | 4 |
| 2024 | Deceptive Waves: Embedding Malicious Backdoors in PPG Authentication
Zeming Yao, Lin Li 0066, Leo Yu Zhang, Fusen Guo, Chao Chen 0015, Jun Zhang 0010 |
WISE (2) | 5 |
| 2024 | IoTFuzz: Automated Discovery of Violations in Smart Homes With Real EnvironmentabstractSmart homes (SHs) are rapidly evolving to incorporate intelligent features, including environment management, home automation, and human–machine interactions. However, safety and security risks of SHs hinder their wide adoption. Many work attempts to provide defense mechanisms to ensure safety and security against interrule vulnerabilities and spoofing attacks. This article proposes IoTFuzz, a fuzzing framework that dynamically address cyber security and physical safety aspects of SHs through targeted policies. IoTFuzz mutates the inputs from policies, human activities, indoor environment, and real-life outdoor weather conditions. In addition to the binary status of devices, the continuous-value status in SHs is leveraged to perform mutation and simulation. The policies are expressed as temporal logic formulas with time constraints. For large-scale testing, IoTFuzz employs digital twins to simulate normal behaviors, outdoor environment impacts, and human activities in SHs. Moreover, IoTFuzz can also intelligently infer rule-policy correlation based on natural language processing (NLP) techniques. The evaluation of IoTFuzz in a configured SH with 15 rules and 10 predefined unique policies demonstrates its effectiveness in revealing the impacts of real-life outdoor environment. The experimental results demonstrate a range of violations, with a maximum of 4154 violations and a minimum of 41 violations observed over an 8-year period under varying weather conditions. IoTFuzz also identifies the potential risks associated with improper human activities, accounting for up to 35.4% of risky situations in SHs. Xinbo Ban, Ming Ding 0001, Shigang Liu, Chao Chen 0015, Jun Zhang 0010 |
IEEE Internet Things J. | 4 |
| 2024 | Matrix factorization recommender based on adaptive Gaussian differential privacy for implicit feedback
Yong Wang 0009, Jiangzhou Deng, Chao Chen 0015, Leo Yu Zhang |
Inf. Process. Manag. | 5 |
| 2024 | MODA: Model Ownership Deprivation Attack in Asynchronous Federated LearningabstractTraining a deep learning model from scratch requires a great deal of available labeled data, computation resources, and expert knowledge. Thus, the time-consuming and complicated learning procedure catapulted the trained model to valuable intellectual property (IP), spurring interest from attackers in model copyright infringement and stealing. Recently, a new defense approach leverages watermarking techniques to inject watermarks into the training procedure and verify model ownership when necessary. To our best knowledge, there is no research work on model ownership stealing attacks in federated learning, and the existing defense or mitigation methods can not be directly used for federated learning scenarios. In this paper, we introduce watermarking neural networks in asynchronous federated learning and propose a novel model privacy attack, dubbed model ownership deprivation attack (MODA). MODA is launched by an inside adversarial participant, targeting occupying and depriving the remaining participants' (victims) copyright to achieve his maximum profit. The extensive experimental results on five benchmark datasets (MNIST, Fashion-MNIST, GTSRB, SVHN, CIFAR10) show that MODA is highly effective in a two-participant learning scenario with a minor impact on model's performance. When extending MODA into multiple participants scenario, MODA still maintains high attack success rate and classification accuracy. Compared to the state-of-the-art works, MODA has a higher attack success rate than the black-box solution and comparable efficacy with the approach in the white-box scenario. Xiaoyu Zhang 0010, Shen Lin 0006, Chao Chen 0015, Xiaofeng Chen 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2024 | LocGuard: A Location Privacy Defender for Image SharingabstractThe privacy of social media users is a major concern when the users share their content to the public. Sensitive information such as the location of the users can be inferred from relevant content without arising the awareness of the users. With blooming services provided by social media platforms, the users have more freedom to share information via diverse data formats. The multi-modality of the shared information may, in return, worsen the private information leakage caused by inference attacks. In this paper, we first examine the problem of location inference on multi-modal data comprised of textual information and visual content. It is observed that the visual content, such as photos shared by social media users, can significantly boost the success rate of location inference. To thwart adversaries who are driven by visual-related data, we propose a defence that mitigates the threat of location privacy breach under an imperceptible utility loss. Our defence, namely LocGuard, perturbs the photos in a one-off manner before sharing them. The perturbations, along with a simple but effective bipartite perturbation strategy, ensure that LocGuard is resistant to adaptive adversaries who can perform adversarial training based on the perturbed photos. Moreover, LocGuard remains effective against open-set adversaries whose data categories in the training dataset are hidden from the defender. In the evaluation, we conduct extensive experiments based on real-world datasets and compare our work with previous methods. The results show that LocGuard significantly outperforms the existing defences. In particular, LocGuard not only achieves better privacy protection and utility preservation for image sharing, but also can effectively defend against adversarial-training-capable attackers. Wanlun Ma, Derui Wang, Chao Chen 0015, Sheng Wen, Gaolei Fei, Yang Xiang 0001 |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2024 | Reverse Backdoor Distillation: Towards Online Backdoor Attack Detection for Deep Neural Network ModelsabstractThe backdoor attack on deep neural network models implants malicious data patterns in a model to induce attacker-desirable behaviors. Existing defense methods fall into the online and offline categories, in which the offline models achieve state-of-the-art detection rates but are restricted by heavy computation overhead. In contrast, their more deployable online counterparts lack the means to detect source-specific backdoors with large sizes. This work proposes a new online backdoor detection method—Reverse Backdoor Distillation (RBD) to handle issues associated with source-specific and source-agnostic backdoor attacks. RBD, designed with the novel perspective of distilling instead of erasing backdoor knowledge, is a complementary backdoor detection methodology that can be used in conjunction with other online backdoor defenses. Considering the fact that trigger data will cause overwhelming neuron activation while clean data will not, RBD distills backdoor attack pattern knowledge from a suspicious model to create a shadow model, which is subsequently deployed online along with the original model in scope to predict a backdoor attack. We extensively evaluate RBD on several datasets (MNIST, GTSRB, CIFAR-10) with diverse model architectures and trigger patterns. RBD outperforms online benchmarks in all experimental settings. Notably, RBD demonstrates superior capability in detecting source-specific attacks, where comparison methods fail, underscoring the effectiveness of our proposed technique. Moreover, RBD achieves a computational savings of at least 97%. Zeming Yao, Hangtao Zhang, Yicheng Guo, Xin Tian 0015, Wei Peng 0011, Yi Zou 0001, Leo Yu Zhang, Chao Chen 0015 |
IEEE Trans. Dependable Secur. Comput. | 8 |
| 2023 | Hiding Your Signals: A Security Analysis of PPG-Based Biometric Authentication
Lin Li 0066, Chao Chen 0015, Lei Pan 0002, Yonghang Tai, Jun Zhang 0010, Yang Xiang 0001 |
ESORICS (3) | 2 |
| 2023 | Denial-of-Service or Fine-Grained Control: Towards Flexible Model Poisoning Attacks on Federated LearningabstractFederated learning (FL) is vulnerable to poisoning attacks, where adversaries corrupt the global aggregation results and cause denial-of-service (DoS). Unlike recent model poisoning attacks that optimize the amplitude of malicious perturbations along certain prescribed directions to cause DoS, we propose a flexible model poisoning attack (FMPA) that can achieve versatile attack goals. We consider a practical threat scenario where no extra knowledge about the FL system (e.g., aggregation rules or updates on benign devices) is available to adversaries. FMPA exploits the global historical information to construct an estimator that predicts the next round of the global model as a benign reference. It then fine-tunes the reference model to obtain the desired poisoned model with low accuracy and small perturbations. Besides the goal of causing DoS, FMPA can be naturally extended to launch a fine-grained controllable attack, making it possible to precisely reduce the global accuracy. Armed with precise control, malicious FL service providers can gain advantages over their competitors without getting noticed, hence opening a new attack surface in FL other than DoS. Even for the purpose of DoS, experiments show that FMPA significantly decreases the global accuracy, outperforming six state-of-the-art attacks. Hangtao Zhang, Zeming Yao, Leo Yu Zhang, Shengshan Hu, Chao Chen 0015, Alan Wee-Chung Liew, Zhetao Li |
IJCAI | 5 |
| 2023 | SigD: A Cross-Session Dataset for PPG-based User Authentication in Different Demographic GroupsabstractRecently, unobservable physiological signals have received widespread attention from researchers as unique identifiers of users in biometrics. However, due to the lack of data sets, existing methods are limited in evaluating cross-session scenarios. Cross-session means that signals are collected at different sessions (times). In real scenarios, authentication is almost always cross-session. Currently, the datasets commonly used for Photoplethysmogram (PPG) signal authentication span around one month, which is insufficient for authentication. On the other hand, different demographic groups have different hemodynamic characteristics, but existing methods lack an assessment of these aspects. This paper introduces a dataset to provide insights into PPG signal-based authentication across different time spans and user groups (age, gender). As physiological signals offer unique advantages for user authentication, the potential of PPG signals is gradually explored. Furthermore, our comparative analysis of recent publications on data-driven user authentication using PPG can further identify the similarities and differences among the performance of the proposed authentication models. Our findings may help future research towards a consensus on an appropriate set of performance metrics. Lin Li 0066, Chao Chen 0015, Lei Pan 0002, Jun Zhang 0010, Yang Xiang 0001 |
IJCNN | 2 |
| 2023 | SigA: rPPG-based Authentication for Virtual Reality Head-mounted DisplayabstractConsumer-grade virtual reality head-mounted displays (VR-HMD) are becoming increasingly popular. Despite VR’s convenience and booming applications, VR-based authentication schemes are underdeveloped. The recently proposed authentication methods (Electrooculogram based, Electrical Muscle Stimulation-based, and alike) require active user involvement, disturbing many scenarios like drone flight and telemedicine. This paper proposes an effective and efficient user authentication method in VR environments resilient to impersonation attacks using physiological signals — Photoplethysmogram (PPG), namely SigA. SigA exploits the advantage that PPG is a physiological signal invisible to the naked eye. Using VR-HMDs to cover the eye area completely, SigA reduces the risk of signal leakage during PPG acquisition. We conducted a comprehensive analysis of SigA’s feasibility on five publicly available datasets, nine different pre-trained models, three facial regions, various lengths of the video clips required for training, four different signal time intervals, and continuous authentication with different sliding window sizes. The results demonstrate that SigA achieves more than 95% of the average F1-score in a one-second signal to accommodate a complete cardiac cycle for most adults, implying its applicability in real-world scenarios. Furthermore, experiments have shown that SigA is resistant to zero-effort attacks, statistical attacks, impersonation attacks (with a detection accuracy of over 95%) and session hijacking attacks. Lin Li 0066, Chao Chen 0015, Lei Pan 0002, Leo Yu Zhang, Jun Zhang 0010, Yang Xiang 0001 |
RAID | 2 |
| 2023 | Security and privacy problems in voice assistant applications: A surveyabstractVoice assistant applications have become omniscient nowadays. Two models that provide the two most important functions for real-life applications (i.e., Google Home, Amazon Alexa, Siri, etc.) are Automatic Speech Recognition (ASR) models and Speaker Identification (SI) models. According to recent studies, security and privacy threats have also emerged with the rapid development of the Internet of Things (IoT). The security issues researched include attack techniques toward machine learning models and other hardware components widely used in voice assistant applications. The privacy issues include technical-wise information stealing and policy-wise privacy breaches. The voice assistant application takes a steadily growing market share every year, but their privacy and security issues never stopped causing huge economic losses and endangering users' personal sensitive information. Thus, it is important to have a comprehensive survey to outline the categorization of the current research regarding the security and privacy problems of voice assistant applications. This paper concludes and assesses five kinds of security attacks and three types of privacy threats in the papers published in the top-tier conferences of cyber security and voice domain. Jingjin Li, Chao Chen 0015, Mostafa Rahimi Azghadi, Hossein Ghodosi, Lei Pan 0002, Jun Zhang 0010 |
Comput. Secur. | 2 |
| 2023 | A Survey of PPG's Application in AuthenticationabstractBiometric authentication prospered because of its convenient use and security. Early generations of biometric mechanisms suffer from spoofing attacks. Recently, unobservable physiological signals (e.g., Electroencephalogram, Photoplethysmogram, Electrocardiogram) as biometrics offer a potential remedy to this problem. In particular, Photoplethysmogram (PPG) measures the change in blood flow of the human body by an optical method. Clinically, researchers commonly use PPG signals to obtain patients' blood oxygen saturation, heart rate, and other information to assist in diagnosing heart-related diseases. Since PPG signals contain a wealth of individual cardiac information, researchers have begun to explore their potential in cyber security applications. The unique advantages (simple acquisition, difficult to steal, and live detection) of the PPG signal allow it to improve the security and usability of the authentication in various aspects. However, the research on PPG-based authentication is still in its infancy. The lack of systematization hinders new research in this field. We conduct a comprehensive study of PPG-based authentication and discuss these applications' limitations before pointing out future research directions. Lin Li 0066, Chao Chen 0015, Lei Pan 0002, Leo Yu Zhang, Jun Zhang 0010, Yang Xiang 0001 |
Comput. Secur. | 2 |
| 2023 | Real-Time Detection of COVID-19 Events From Twitter: A Spatial-Temporally Bursty-Aware MethodabstractIn the last two years, the outbreak of COVID-19 has significantly affected human life, society, and the economy worldwide. To prevent people from contracting COVID-19 and mitigate its spread, it is crucial to timely distribute complete, accurate, and up-to-date information about the pandemic to the public. In this article, we propose a spatial–temporally bursty-aware method calledSTBAfor real-time detection of COVID-19 events from Twitter.STBAhas three consecutive stages. In the first stage,STBAidentifies a set of keywords that represent COVID-19 events according to the spatiotemporally bursty characteristics of words using Ripley’s$K$function.STBAwill also filter out tweets that do not contain the keywords to reduce the interference of noise tweets on event detection. In the second stage,STBAuses online density-based spatial clustering of applications with noise clustering to aggregate tweets that describe the same event as much as possible, which provides more information for event identification. In the third stage,STBAfurther utilizes the temporal bursty characteristic of event location information in the clusters to identify real-world COVID-19 events. Each stage ofSTBAcan be regarded as a noise filter. It gradually filters out COVID-19-related events from noisy tweet streams. To evaluate the performance ofSTBA, we collected over 116 million Twitter posts from 36 consecutive days (from March 22, 2020 to April 26, 2020) and labeled 501 real events in this dataset. We comparedSTBAwith three state-of-the-art methods, EvenTweet, event detection via microblog cliques (EDMC), and GeoBurst+ in the evaluation. The experimental results suggest thatSTBAoutperforms GeoBurst+ by 13.8%, 12.7%, and 13.3% in terms of precision, recall, and$F_{1}$score.STBAachieved even more improvements compared with EvenTweet and EDMC. Gaolei Fei, Wanlun Ma, Chao Chen 0015, Sheng Wen, Guangmin Hu |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2023 | Personalized Location Privacy Protection for Location-Based Services in Vehicular NetworksabstractLocation-based services (LBSs) are widely used in vehicular networks. Privacy leakage from LBS is a key issue to be solved. However, the existing schemes fail to provide differentiated protection for users’ different locations, which may lead to the leakage of location information. In this paper, we propose a personalized location privacy protection scheme based on differential privacy to protect the privacy of location-based services in vehicular networks. Firstly, we propose a normalized decision matrix to describe the efficiency and the privacy effect of navigation recommendations. We then establish a utility model integrated with users’ privacy preferences to compute the effective driving route. Secondly, for different service request locations in the driving route, we define sensitivity distance as an index to quantify their privacy requirements. The privacy budget will be added to the service request location to generate a false location. Moreover, due to the limitation of road range in the driving route, if the privacy budget value allocated is small enough, the false location generated by the Plane Laplace will be deviated. As a result, the attacker can deduce users’ real request locations. Consequently, considering the factors of trajectory leakage, attack strategy and QoS, we establish a multi-objective optimization model to optimize the false location. Based on the real data set, we conduct a series of comparison simulations to evaluate the performance of the proposed scheme. The experimental results demonstrate that our scheme can satisfy users’ personalized services needs and provide an optimal solution to privacy and QoS. Chuan Xu 0001, Yingyi Ding, Chao Chen 0015, Yong Ding 0005, Wei Zhou 0044, Sheng Wen |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | No-Label User-Level Membership Inference for ASR Model Auditing
Yuantian Miao, Chao Chen 0015, Lei Pan 0002, Shigang Liu, Seyit Ahmet Çamtepe, Jun Zhang 0010, Yang Xiang 0001 |
ESORICS (2) | 2 |
| 2022 | Automated Binary Analysis: A Survey
Zian Liu, Chao Chen 0015, Dongxi Liu, Jun Zhang 0010 |
ICA3PP | 2 |
| 2022 | A Survey on IoT Vulnerability Discovery
Xinbo Ban, Ming Ding 0001, Shigang Liu, Chao Chen 0015, Jun Zhang 0010 |
NSS | 4 |
| 2021 | The Audio Auditor: User-Level Membership Inference in Internet of Things Voice ServicesabstractAbstract With the rapid development of deep learning techniques, the popularity of voice services implemented on various Internet of Things (IoT) devices is ever increasing. In this paper, we examine user-level membership inference in the problem space of voice services, by designing an audio auditor to verify whether a specific user had unwillingly contributed audio used to train an automatic speech recognition (ASR) model under strict black-box access. With user representation of the input audio data and their corresponding translated text, our trained auditor is effective in user-level audit. We also observe that the auditor trained on specific data can be generalized well regardless of the ASR model architecture. We validate the auditor on ASR models trained with LSTM, RNNs, and GRU algorithms on two state-of-the-art pipelines, the hybrid ASR system and the end-to-end ASR system. Finally, we conduct a real-world trial of our auditor on iPhone Siri, achieving an overall accuracy exceeding 80%. We hope the methodology developed in this paper and findings can inform privacy advocates to overhaul IoT privacy. Yuantian Miao, Minhui Xue 0001, Chao Chen 0015, Lei Pan 0002, Jun Zhang 0010, Benjamin Zi Hao Zhao, Mohamed Ali Kâafar, Yang Xiang 0001 |
Proc. Priv. Enhancing Technol. | 3 |
| 2020 | Doc2vec-based Insider Threat Detection through Behaviour Analysis of Multi-source Security LogsabstractSince insider attacks have been recognised as one of the most critical cyber security threats to an organisation, detection of malicious insiders has received increasing attention in recent years. Previously, we proposed an approach that performs the detection by analysing various security logs with Word2vec, which not only removes the reliance on prior knowledge but also greatly simplifies the process of decision making and improves the interpretability of the alerts. In this paper, following the similar idea, a new Doc2vec based approach is proposed to overcome the previous approach's limitations: (1) the behaviour metrics can be acquired straightforwardly due to the Doc2vec's capability in inferring unseen texts of any length; (2) other than the temporal metrics, some spatial metrics can also be realised, providing a more comprehensive insight into the unusual behaviours; and (3) a range of corpora are produced by adopting different keywords to aggregate, each of which may be suited to a specific type of behaviour metrics. A large number of numerical experiments are conducted using the same benchmark insider threat database, for the purpose of testing how the corpora, metrics and training parameters impact on the performance and be related to each other. The experiments demonstrate that the proposed approach can achieve a similar performance with greater simplicity and flexibility. Liu Liu 0008, Chao Chen 0015, Jun Zhang 0010, Olivier Y. de Vel, Yang Xiang 0001 |
TrustCom | 2 |
| 2020 | Cyber Vulnerability Intelligence for Internet of Things BinaryabstractInternet of Things (IoT) integrates a variety of software (e.g., autonomous vehicles and military systems) in order to enable the advanced and intelligent services. These software increase the potential of cyber-attacks because an adversary can launch an attack using system vulnerabilities. Existing software vulnerability analysis methods used to be relying on human experts crafted features, which usually miss many vulnerabilities. It is important to develop an automatic vulnerability analysis system to improve the countermeasures. However, source code is not always available (e.g., most IoT related industry software are closed source). Therefore, vulnerability detection on binary code is a demanding task. This article addresses the automatic binary-level software vulnerability detection problem by proposing a deep learning-based approach. The proposed approach consists of two phases: binary function extraction, and model building. First, we extract binary functions from the cleaned binary instructions obtained by using IDA Pro. Then, we employ the attention mechanism on top of a bidirectional long short-term memory for building the predictive model. To show the effectiveness of the proposed approach, we have collected datasets from several different sources. We have compared our proposed approach with a series of baselines including source code-based techniques and binary code-based techniques. We have also applied the proposed approach to real-world IoT related software such as VLC media player and LibTIFF project that used on Autonomous Vehicles. Experimental results show that our proposed approach betters the baselines and is able to detect more vulnerabilities. Shigang Liu, Mahdi Dibaei, Yonghang Tai, Chao Chen 0015, Jun Zhang 0010, Yang Xiang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2019 | Unsupervised Insider Detection Through Neural Feature Learning and Model Optimisation
Liu Liu 0008, Chao Chen 0015, Jun Zhang 0010, Olivier Y. de Vel, Yang Xiang 0001 |
NSS | 2 |
| 2019 | A performance evaluation of deep-learnt features for software vulnerability detectionabstractSummary Software vulnerability is a critical issue in the realm of cyber security. In terms of techniques, machine learning (ML) has been successfully used in many real‐world problems such as software vulnerability detection, malware detection and function recognition, for high‐quality feature representation learning. In this paper, we propose a performance evaluation study on ML based solutions for software vulnerability detection, conducting three experiments: machine learning‐based techniques for software vulnerability detection based on the scenario of single type of vulnerability and multiple types of vulnerabilities per dataset; machine learning‐based techniques for cross‐project software vulnerability detection; and software vulnerability detection when facing the class imbalance problem with varying imbalance ratios. Experimental results show that it is possible to employ software vulnerability detection based on ML techniques. However, ML‐based techniques suffer poor performance on both cross‐project and class imbalance problem in software vulnerability detection. Xinbo Ban, Shigang Liu, Chao Chen 0015, Caslon Chua |
Concurr. Comput. Pract. Exp. | 3 |
| 2017 | Addressing the class imbalance problem in Twitter spam detection using ensemble learning
Shigang Liu, Yu Wang 0017, Jun Zhang 0010, Chao Chen 0015, Yang Xiang 0001 |
Comput. Secur. | 4 |
| 2017 | Investigating the deceptive information in Twitter spam
Chao Chen 0015, Sheng Wen, Jun Zhang 0010, Yang Xiang 0001, Jonathan Oliver, Abdulhameed Alelaiwi, Mohammad Mehedi Hassan |
Future Gener. Comput. Syst. | 1 |
| 2017 | Energy-aware scheduling of virtual machines in heterogeneous cloud computing systems
Hancong Duan, Chao Chen 0015, Geyong Min |
Future Gener. Comput. Syst. | 2 |
| 2017 | Statistical Features-Based Real-Time Detection of Drifted Twitter SpamabstractTwitter spam has become a critical problem nowadays. Recent works focus on applying machine learning techniques for Twitter spam detection, which make use of the statistical features of tweets. In our labeled tweets data set, however, we observe that the statistical properties of spam tweets vary over time, and thus, the performance of existing machine learning-based classifiers decreases. This issue is referred to as “Twitter Spam Drift”. In order to tackle this problem, we first carry out a deep analysis on the statistical features of one million spam tweets and one million non-spam tweets, and then propose a novel Lfun scheme. The proposed scheme can discover “changed” spam tweets from unlabeled tweets and incorporate them into classifier's training process. A number of experiments are performed to evaluate the proposed scheme. The results show that our proposed Lfun scheme can significantly improve the spam detection accuracy in real-world scenarios. Chao Chen 0015, Yu Wang 0017, Jun Zhang 0010, Yang Xiang 0001, Wanlei Zhou 0001, Geyong Min |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2016 | An Ensemble Learning Approach for Addressing the Class Imbalance Problem in Twitter Spam Detection
Shigang Liu, Yu Wang 0017, Chao Chen 0015, Yang Xiang 0001 |
ACISP (1) | 3 |
| 2016 | Gatekeeping Behavior Analysis for Information Credibility Assessment on Weibo
Bailin Xie, Yu Wang 0017, Chao Chen 0015, Yang Xiang 0001 |
NSS | 3 |
| 2016 | Comments and CorrectionsabstractPresents correcttions to the paper, “A performance evaluation of machine learning-based streaming spam tweets detection,” (Chen ], C.; et al) , IEEE Trans. Comput. Social Syst., vol. 2, no. 3, pp. 65–76, Sep. 2015. Chao Chen 0015, Jun Zhang 0010, Yi Xie 0002, Yang Xiang 0001, Wanlei Zhou 0001, Mohammad Mehedi Hassan, Abdulhameed Alelaiwi |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2015 | 6 million spam tweets: A large ground truth for timely Twitter spam detectionabstractTwitter has changed the way of communication and getting news for people's daily life in recent years. Meanwhile, due to the popularity of Twitter, it also becomes a main target for spamming activities. In order to stop spammers, Twitter is using Google SafeBrowsing to detect and block spam links. Despite that blacklists can block malicious URLs embedded in tweets, their lagging time hinders the ability to protect users in real-time. Thus, researchers begin to apply different machine learning algorithms to detect Twitter spam. However, there is no comprehensive evaluation on each algorithms' performance for real-time Twitter spam detection due to the lack of large groundtruth. To carry out a thorough evaluation, we collected a large dataset of over 600 million public tweets. We further labelled around 6.5 million spam tweets and extracted 12 light-weight features, which can be used for online detection. In addition, we have conducted a number of experiments on six machine learning algorithms under various conditions to better understand their effectiveness and weakness for timely Twitter spam detection. We will make our labelled dataset for researchers who are interested in validating or extending our work. Chao Chen 0015, Jun Zhang 0010, Xiao Chen 0002, Yang Xiang 0001, Wanlei Zhou 0001 |
ICC | 1 |
| 2015 | Unknown pattern extraction for statistical network protocol identificationabstractThe past decade has seen a lot of research on statistics-based network protocol identification using machine learning techniques. Prior studies have shown promising results in terms of high accuracy and fast classification speed. However, most works have embodied an implicit assumption that all protocols are known in advance and presented in the training data, which is unrealistic since real-world networks constantly witness emerging traffic patterns as well as unknown protocols in the wild. In this paper, we revisit the problem by proposing a learning scheme with unknown pattern extraction for statistical protocol identification. The scheme is designed with a more realistic setting, where the training dataset contains labeled samples from a limited number of protocols, and the goal is to tell these known protocols apart from each other and from potential unknown ones. Preliminary results derived from real-world traffic are presented to show the effectiveness of the scheme. Yu Wang 0017, Chao Chen 0015, Yang Xiang 0001 |
LCN | 2 |
| 2015 | A Sword with Two Edges: Propagation Studies on Both Positive and Negative Information in Online Social NetworksabstractOnline social networks (OSN) have become one of the major platforms for people to exchange information. Both positive information (e.g., ideas, news and opinions) and negative information (e.g., rumors and gossips) spreading in social media can greatly influence our lives. Previously, researchers have proposed models to understand their propagation dynamics. However, those were merely simulations in nature and only focused on the spread of one type of information. Due to the human-related factors involved, simultaneous spread of negative and positive information cannot be thought of the superposition of two independent propagations. In order to fix these deficiencies, we propose an analytical model which is built stochastically from a node level up. It can present the temporal dynamics of spread such as the time people check newly arrived messages or forward them. Moreover, it is capable of capturing people’s behavioral differences in preferring what to believe or disbelieve. We studied the social parameters impact on propagation using this model. We found that some factors such as people’s preference and the injection time of the opposing information are critical to the propagation but some others such as the hearsay forwarding intention have little impact on it. The extensive simulations conducted on the real topologies confirm the high accuracy of our model. Sheng Wen, Mohammad Sayad Haghighi, Chao Chen 0015, Yang Xiang 0001, Wanlei Zhou 0001, Weijia Jia 0001 |
IEEE Trans. Computers | 3 |
| 2015 | A Performance Evaluation of Machine Learning-Based Streaming Spam Tweets DetectionabstractThe popularity of Twitter attracts more and more spammers. Spammers send unwanted tweets to Twitter users to promote websites or services, which are harmful to normal users. In order to stop spammers, researchers have proposed a number of mechanisms. The focus of recent works is on the application of machine learning techniques into Twitter spam detection. However, tweets are retrieved in a streaming way, and Twitter provides the Streaming API for developers and researchers to access public tweets in real time. There lacks a performance evaluation of existing machine learning-based streaming spam detection methods. In this paper, we bridged the gap by carrying out a performance evaluation, which was from three different aspects of data, feature, and model. A big ground-truth of over 600 million public tweets was created by using a commercial URL-based security tool. For real-time spam detection, we further extracted 12 lightweight features for tweet representation. Spam detection was then transformed to a binary classification problem in the feature space and can be solved by conventional machine learning algorithms. We evaluated the impact of different factors to the spam detection performance, which included spam to nonspam ratio, feature discretization, training data size, data sampling, time-related data, and machine learning algorithms. The results show the streaming spam tweet detection is still a big challenge and a robust detection technique should take into account the three aspects of data, feature, and model. Chao Chen 0015, Jun Zhang 0010, Yi Xie 0002, Yang Xiang 0001, Wanlei Zhou 0001, Mohammad Mehedi Hassan, Abdulhameed Alelaiwi, Majed A. AlRubaian |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2014 | On Addressing the Imbalance Problem: A Correlated KNN Approach for Network Traffic Classification
Di Wu 0050, Xiao Chen 0002, Chao Chen 0015, Jun Zhang 0010, Yang Xiang 0001, Wanlei Zhou 0001 |
NSS | 3 |
| 2013 | Robust network traffic identification with unknown applicationsabstractTraffic classification is a fundamental component in advanced network management and security. Recent research has achieved certain success in the application of machine learning techniques into flow statistical feature based approach. However, most of flow statistical feature based methods classify traffic based on the assumption that all traffic flows are generated by the known applications. Considering the pervasive unknown applications in the real world environment, this assumption does not hold. In this paper, we cast unknown applications as a specific classification problem with insufficient negative training data and address it by proposing a binary classifier based framework. An iterative method is proposed to extract unknown information from a set of unlabelled traffic flows, which combines asymmetric bagging and flow correlation to guarantee the purity of extracted negatives. A binary classifier is used as an application signature which can operate on a bag of correlated flows instead of individual flows to further improve its effectiveness. We carry out a series of experiments in a real-world network traffic dataset to evaluate the proposed methods. The results show that the proposed method significantly outperforms the-state-of-art traffic classification methods under the situation of unknown applications present. Jun Zhang 0010, Chao Chen 0015, Yang Xiang 0001, Wanlei Zhou 0001 |
AsiaCCS | 2 |
| 2013 | Internet Traffic Classification by Aggregating Correlated Naive Bayes PredictionsabstractThis paper presents a novel traffic classification scheme to improve classification performance when few training data are available. In the proposed scheme, traffic flows are described using the discretized statistical features and flow correlation information is modeled by bag-of-flow (BoF). We solve the BoF-based traffic classification in a classifier combination framework and theoretically analyze the performance benefit. Furthermore, a new BoF-based traffic classification method is proposed to aggregate the naive Bayes (NB) predictions of the correlated flows. We also present an analysis on prediction error sensitivity of the aggregation strategies. Finally, a large number of experiments are carried out on two large-scale real-world traffic datasets to evaluate the proposed scheme. The experimental results show that the proposed scheme can achieve much better classification performance than existing state-of-the-art traffic classification methods. Jun Zhang 0010, Chao Chen 0015, Yang Xiang 0001, Wanlei Zhou 0001, Yong Xiang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2013 | An Effective Network Traffic Classification Method with Unknown Flow DetectionabstractTraffic classification technique is an essential tool for network and system security in the complex environments such as cloud computing based environment. The state-of-the-art traffic classification methods aim to take the advantages of flow statistical features and machine learning techniques, however the classification performance is severely affected by limited supervised information and unknown applications. To achieve effective network traffic classification, we propose a new method to tackle the problem of unknown applications in the crucial situation of a small supervised training set. The proposed method possesses the superior capability of detecting unknown flows generated by unknown applications and utilizing the correlation information among real-world network traffic to boost the classification performance. A theoretical analysis is provided to confirm performance benefit of the proposed method. Moreover, the comprehensive performance evaluation conducted on two real-world network traffic datasets shows that the proposed scheme outperforms the existing methods in the critical network environment. Jun Zhang 0010, Chao Chen 0015, Yang Xiang 0001, Wanlei Zhou 0001, Athanasios V. Vasilakos |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2012 | Classification of Correlated Internet Traffic FlowsabstractA critical problem for Internet traffic classification is how to obtain a high-performance statistical feature based classifier using a small set of training data. The solutions to this problem are essential to deal with the encrypted applications and the new emerging applications. In this paper, we propose a new Naive Bayes (NB) based classification scheme to tackle this problem, which utilizes two recent research findings, feature discretization and flow correlation. A new bag-of-flow (BoF) model is firstly introduced to describe the correlated flows and it leads to a new BoF-based traffic classification problem. We cast the BoF-based traffic classification as a specific classifier combination problem and theoretically analyze the classification benefit from flow aggregation. A number of combination methods are also formulated and used to aggregate the NB predictions of the correlated flows. Finally, we carry out a number of experiments on a large scale real-world network dataset. The experimental results show that the proposed scheme can achieve significantly higher classification accuracy and much faster classification speed with comparison to the state-of-the-art traffic classification methods. Jun Zhang 0010, Chao Chen 0015, Yang Xiang 0001, Wanlei Zhou 0001 |
TrustCom | 2 |