VLDB 2026 Research / reviewers in the wild / expert
Weina Niu
dblp:217/3426
· DBLP profile ↗
60ranked-venue papers
9as first author
51since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 19 · 3 first-author · 15 since 2021Computer networks · 17 · 4 first-author · 14 since 2021Artificial intelligence and machine learning · 7 · 5 since 2021Software engineering, systems software and programming languages · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RlDecompiler: Enhancing LLM-based Decompilation via Reinforcement Learning with a Multi-Faceted Reward FunctionabstractDecompiling binary code into human-readable, high-level source code is a core challenge in reverse engineering. While traditional methods often rely on brittle, pattern-based heuristics, the advent of Large Language Models (LLMs) offers a more flexible and robust approach. However, current LLM-based decompilation efforts are often limited by their training methodologies, which typically treat the task as a simple sequence-to-sequence translation and struggle to enforce the functional correctness of the output. To address these issues, this paper proposes an innovative framework for training LLMs to perform high-fidelity decompilation. A core contribution of our work is a novel data processing pipeline that enriches the model’s input. This pipeline integrates Ghidra-based static analysis to directly embed crucial context, such as static resources (strings, floating-point numbers) and relabeled basic blocks—from the binary into an LLM-friendly prompt. Building on this enriched input, we employ reinforcement learning fine-tuning guided by a multi-faceted reward function that comprehensively evaluates syntactic correctness, AST similarity, compilability, and functional correctness via test cases. Using this framework, we trained the RlDecompiler family of models (1.3B and 3B). Experimental results demonstrate that RlDecompiler achieves state-of-the-art performance, and its generated code quality is also higher than that of the baseline models. The RlDecompiler 1.3B and 3B models achieve rerunnable rates of 27.96% and 40.70%, respectively, outperforming existing baselines. The code is available at https://github.com/ri-char/rldecompile. Yuchi Su, Weina Niu, Jiacheng Gong, Song Li 0006, Xin Liu 0050, Xiaosong Zhang 0001 |
ICPC | 2 |
| 2026 | RepShield: Robust knowledge representation in continual learning for network intrusion detection
Weina Niu, Mingze He, Xuyang Ding, Jiacheng Gong |
Comput. Networks | 2 |
| 2026 | CPGHunter: LLM-guided semantic modeling for scalable vulnerability detection via taint analysis
Anran Hou, Bingjun Su, Weina Niu, Qinsheng Hou, Honghua Wu, Xiaosong Zhang 0001 |
Empir. Softw. Eng. | 3 |
| 2026 | Towards a comprehensive framework for verifying open-source software license compatibility
Ziang Liu 0006, Xin Liu 0050, Yingli Zhang, Song Li 0006, Weina Niu, Qingguo Zhou, Rui Zhou 0005, Xiaokang Zhou |
Empir. Softw. Eng. | 5 |
| 2026 | PIEDChecker: Uncover Permissions-Independent Emulation-Detection Methods in Android SystemabstractFor compatibility checks and preventing malicious cheating behaviors in Android systems, It is convenient for benign app developers to utilize emulation-detection technology. However, this technique has been abused by malicious app developers, which causes detection emulation and behavior change, known as anti-emulation behavior, to evade the dynamic analysis performed via the Android emulator. The Android permission mechanism can limit some anti-emulation behaviors, but attackers can still use permission-independent (PI) emulation-detection technology to achieve their goals. In this paper, we propose a static and dynamic combined detection framework namedPIEDCheckerto detect PI emulation-detection apps. This framework can statically identify PI emulation-detection code and dynamically verify PI anti-emulation behaviors.PIEDChecker's performance is validated by 382 manually created test apps and 344 apps with emulation-detection labels. Moreover,PIEDCheckerhas higher accuracy in detecting PI emulation-detection compared with the existing Android malware analysis platforms and academic methods. Moreover, it is found that there are 13,377 apps having PI emulation-detection behaviors within the tested 25,303 apps collected over the past five years. In particular, the detection features of the PI emulation-detection methods are summarized based on the evaluation result. Weina Niu, Qinsheng Hou, Lingyun Ying, Xiaosong Zhang 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2026 | Two Heads Are Not Better Than One: Continual Learning From Multiple Models for Encrypted Traffic Analysis
Qingjun Yuan, Weina Niu, Jian Chai, Yong Yu 0002 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | FCAL: An Asynchronous Federated Contrastive Semi-supervised Learning Approach for Network Traffic Classification
Qingjun Yuan, Weina Niu, Yanbei Zhu, Yongjuan Wang |
ICICS (3) | 3 |
| 2025 | ISGraphVD: Precise Vulnerability Detection for IoT Supply Chains Based on Identifier Sensitive GraphabstractOpen-source software (OSS) is widely reused in Internet of Things (IoT) devices, leading to widespread N-Day vulnerabilities when outdated components remain unpatched. Existing methods typically encode features of different Common Vulnerabilities and Exposures (CVEs) within a shared representation space. However, the model’s limited capacity, combined with the new vulnerability features, can disrupt previously learned patterns. Minimal code modifications in tiny-patch vulnerabilities are often overshadowed by variations introduced by different compilation settings, making it more difficult to distinguish vulnerable functions from their patched counterparts. This paper introduces ISGraphVD, a novel graph-based and function-level vulnerability detection approach that supports cross-compilation settings and enhances detection accuracy. By modeling each CVE independently through a one-model-per-CVE strategy, ISGraphVD reduces feature interference and improves detection accuracy across diverse CVEs. To better detect tinypatch vulnerability, we propose ISGraph, a fine-grained graph representation that models variable dependencies within and across basic blocks by integrating control flow analysis. Then, ISGraphVD utilizes a Graph Matching Network (GMN) with a cross-graph attention mechanism to identify critical vulnerability patterns. Experiments on IoT OSS projects show that ISGraphVD outperforms state-of-the-art methods, achieving a 6.3 percentage-point (pp) accuracy improvement over the strongest baseline, and real-world tests further validate its effectiveness in IoT supply chains. Yingli Zhang, Xin Liu 0050, Ziang Liu 0006, Song Li 0006, Weina Niu, Rui Zhou 0005, Qingguo Zhou |
ISSRE | 6 |
| 2025 | FATFI: A Framework to Generate Adversarial Traffic with Feature Interpretability
Yikang Wang, Weina Niu, Dujuan Gu, Qingjun Yuan, Jiacheng Gong, Shuangqi Gan, Xiaosong Zhang 0001 |
KSEM (3) | 2 |
| 2025 | SiFMimicEvader: Evading Fake Voice Detection with Adversarial Neural Mimicry AttacksabstractThe application of deep learning in voice cloning has significantly enhanced the quality of cloned voices. While advanced voice cloning technologies are widely applied across various domains, they also pose serious security challenges such as producing natural Deepfakes. In response, numerous studies have focused on detecting fake voices, with many reporting outstanding performance. However, is the issue truly resolved? This paper introduces Adversarial Neural Mimicry Attack (ANMA) which leverages a specialized model to predict the behavior of other similar models, transforming black-box attacks into white-box scenarios indirectly. Based on ANMA and Speaker-irrelative Features (SiFs), we propose a novel black-box attack framework called SiFMimicEvader, designed to evade fake voice detectors with high success rates and minimal query requirements. The framework utilizes speech representation models as the breakthrough to predict the behaviors of fake voice detectors and employs a series of SiFs editing operations as perturbations to deceive these detectors. Experimental results demonstrate the effectiveness of SiFMimicEvader, achieving an average attack success rate exceeding 50% across various detectors, significantly outperforming other attack methods, while also showing great performance in audio quality and query scale, indicating its high availability in real-world scenarios. Xuan Hai, Xin Liu 0050, Ziyao Yu, Song Li 0006, Weina Niu, Rui Zhou 0005, Qingguo Zhou |
ACM Multimedia | 7 |
| 2025 | ROPGMN: Effective ROP and variants discovery using dynamic feature and graph matching network
Weina Niu, Kexuan Zhang |
Future Gener. Comput. Syst. | 1 |
| 2025 | Identifying Android Malware Using Fine-Grained Path Information from HINabstractAs the rapid advances of mobile internet and Internet of Things (IoT), Android has become one of the most widely used operating systems in mobile terminals and IoT devices. However, the massive growth of Android malware poses challenging security problems to these terminals and devices. In this paper, we propose a novel heterogeneous information network (HIN)-based method, called FDroid, for fast and accurate detection of Android malware. Specifically, we first design a fine-grained HIN to model the relationship between APKs and APIs and then extract finer-grained path information, including code blocks, packages and call patterns compared to traditional HIN-based methods, which can effectively improve the detection accuracy without increasing the number of paths. Second, we devise a TF-IWF-based contribution calculation algorithm to select a small number of sensitive APIs calls with high representativeness, which can effectively save the detection time and storage space. Third, we develop an expanded matrix-assisted support vector machine (SVM) classifier for Android malware detection. Experimental results show that the FDroid can achieve 97.43% detection accuracy. Meanwhile, compared with the other related HIN-based malware detection methods, the training time of FDroid is less than 0.1% of them, and the detection time is less than 10% of them. Erfan Zhao, Weina Niu, Cheng Huang 0003, Xixuan Ren, Jiacheng Gong, Anran Hou |
Int. J. Softw. Eng. Knowl. Eng. | 2 |
| 2025 | DLET-Classifier: A Dynamic and Lightweight Method for Encrypted Traffic ClassificationabstractIn recent years, encrypted traffic has become a critical means of ensuring user information security. However, the widespread adoption of encrypted traffic also introduces new challenges, such as enabling attackers to conceal malicious activities within encrypted channels. Consequently, accurate encrypted traffic classification is crucial for strengthening network security defenses. However, encrypted traffic classification methods often employing complex model structures and feature extraction techniques, while neglecting efficiency and latency, which makes them difficult to apply in low-resource scenarios with slow CPU computation speed, limited memory, and a scarce number of training samples. To address these issues, we propose the Dynamic and Lightweight Encrypted Traffic Classifier (DLET-Classifier), which uses the depthwise separable convolutional neural network and the channel attention mechanism to extract features from encrypted traffic. It efficiently captures byte-level features and the relationships between packets for effective classification. To enable the model to update rapidly and adapt to the ever-changing real-world network environment, we propose the Multi2One algorithm. This algorithm first updates the base model, an ensemble of multiple binary classifiers. Then, we use the knowledge distillation technique to transfer knowledge from the base model to a lightweight model. This process allows for model updates and extensions. The results of the multi-class classification comparison experiment show that among all the compared methods, the DLET-Classifier is the model with the smallest number of parameters and the highest throughput, while also achieving excellent classification accuracy. Incremental expansion experiments demonstrate that the Multi2One algorithm enables fast knowledge updates and extensions for the lightweight model (LWG) while maintaining its classification accuracy above 96%, making our method adapt to complex network environments. Jiayong Wu, Weina Niu, Fushan Wei, Shaofeng Li 0001, Shiping Huang, Jiacheng Gong, Xiaosong Zhang 0001 |
IEEE Internet Things J. | 2 |
| 2025 | CAED: A Comprehensive Android Emulator Detection Framework With Data AugmentationabstractAnti-emulation is crucial for Android and IoT security as it helps apps determine whether they are running on a real mobile device or in an emulation environment. This prevents apps from being analyzed, debugged, or reverse-engineered in emulators, ultimately stopping criminals from making illegal profits. Current emulator detection methods cannot balance accuracy, universality, robustness, and compatibility. Their universality is often hindered by limited data diversity and accessibility. To address these issues, we propose the comprehensive Android emulator detection (CAED) framework. The Preprocessing Module of CAED collects and normalizes data from both phones and emulators. We propose the first data augmentation method for emulator detection, emulator detection augmentation generative adversarial network (EDA-GAN), which is tailored to the characteristics of our data and effectively enhances data diversity. The classifier module MFBoost employs an adaptive imputation algorithm and multiple classification and regression trees (CART) for precise classification. Experiments on 324 devices show that CAED improves detection rate by at least 12.5% and up to 44.71% over state-of-the-art (SOTA) methods. The EDA-GAN data augmentation method boosts classifier accuracy, achieving a performance of up to 99.62%. Additionally, CAED’s unique loss function and imputation algorithm enhance the robustness and compatibility of CAED, with a 24% smaller accuracy drop than other methods when features are modified or unavailable. This study presents the CAED framework as an effective solution for protecting apps against real-world security threats in Android and IoT environments. Weina Niu, Qinsheng Hou, Yuchi Su, Jiacheng Gong, Xiaosong Zhang 0001 |
IEEE Internet Things J. | 2 |
| 2025 | Active cybersecurity: vision, model, and key technologiesabstractNoncooperative computer systems and network confrontation present a core challenge in cyberspace security. Traditional cybersecurity technologies predominantly rely on passive response mechanisms, which exhibit significant limitations when addressing real-world complex and unknown threats. This paper introduces the concept of “active cybersecurity,” aiming to enhance network security not only through technical measures but also by leveraging strategy-level defenses. The core assumption of this concept is that attackers and defenders, in the context of network confrontations, act as rational decision-makers seeking to maximize their respective objectives. Building on this observation, this paper integrates game theory to analyze the interdependent relationships between attackers and defenders, thereby optimizing their strategies. Guided by this foundational idea, we propose an active cybersecurity model involving intelligent threat sensing, in-depth behavior analysis, comprehensive path profiling, and dynamic countermeasures, termed SAPC, designed to foster an integrated defense capability encompassing threat perception, analysis, tracing, and response. At its core, SAPC incorporates theoretical analyses of adversarial behavior and the optimization of corresponding strategies informed by game theory. By profiling adversaries and modeling confrontation as a “game,” the model establishes a comprehensive framework that provides both theoretical insights into and practical guidance for cybersecurity. The proposed active cybersecurity model marks a transformative shift from passive defense to proactive perception and confrontation. It facilitates the evolution of cybersecurity technologies toward a new paradigm characterized by active prediction, prevention, and strategic guidance. Xiaosong Zhang 0001, Yukun Zhu, Xiong Li 0002, Yongzhao Zhang, Weina Niu, Fenghua Xu, Junpeng He, Shiping Huang |
Frontiers Inf. Technol. Electron. Eng. | 5 |
| 2025 | WIVIM: Web Injection Vulnerabilities Detection Based on Interprocedural Analysis and MiniLM-GNN
Yizhong Wei, Weina Niu, Honghua Wu, Jiacheng Gong |
Peer Peer Netw. Appl. | 2 |
| 2025 | A Synthetic Data-Assisted Satellite Terrestrial Integrated Network Intrusion Detection FrameworkabstractThe Satellite-Terrestrial Integrated Network (STIN) is an emerging paradigm offering seamless network services across geographical boundaries, yet it faces significant security challenges, including limited intrusion prevention capabilities. Federated learning (FL) provides a viable solution by aggregating traffic data from STIN clients (e.g., ground stations and edge routers) to train models for network intrusion detection systems (NIDS). However, satellite and terrestrial domain data’s non-independent and identically distributed (non-IID) nature hinders training efficiency and performance. This paper proposes STINIDF, a novel STIN intrusion detection framework leveraging FL-based data augmentation. STINIDF utilizes FL to collaboratively train a conditional diffusion model across STIN nodes while preserving privacy via differential privacy mechanisms, generating global traffic data representative of the STIN distribution. Each node then integrates global and local traffic data to train a local model for NIDS, addressing non-IID challenges by balancing data distribution through data augmentation. Using a simulation environment developed with OMNeT++ and INET, a Satellite-Terrestrial Integrated (STI) traffic dataset was created, including intrusion scenarios such as signal disruption, UDP flooding, and jamming attacks. Experimental results indicate that STINIDF outperforms existing data augmentation-based approaches under non-IID conditions, achieving$\mathbf {96.63\%(2.41\%\uparrow)}$accuracy,$\mathbf {96.71\% (3.14\%\uparrow)}$precision,$\mathbf {96.54\%(1.65\%\uparrow)}$recall and$\mathbf {96.66\%(2.7\%\uparrow)}$F1 score. Furthermore, when compared to methods integrating data augmentation with differential privacy, STINIDF demonstrates an effective balance between privacy preservation and intrusion detection performance, attaining an accuracy of$\mathbf {96.14\%(2.57\%\uparrow)}$and a FID of$\mathbf {17.88(7.41\downarrow)}$. Junpeng He, Xiong Li 0002, Xiaosong Zhang 0001, Weina Niu, Fagen Li |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | BPFDex: Enabling Robust Android Apps Unpacking via Android KernelabstractMalware developers exploit packing techniques to protect malicious apps from analysis. These evolving techniques, coupled with diverse anti-unpacker strategies, often render current studies ineffective in unpacking Android apps. In this study, we introduce BPFDex, a novel Android unpacking framework that leverages eBPF, a kernel component of the Android system. We successfully apply eBPF’s excellent kernel observability and tracing capability to Android unpacking, both on real devices and emulators. Operating within the kernel space, BPFDex avoids drawbacks of common unpacking techniques. BPFDex monitors apps across both native and kernel layers, restores Dex data from memory, and adapts to different packing strategies according to observed packing behaviors. Furthermore, we summarize patterns in anti-unpacker behaviors among Android packers, establishing criteria to improve existing unpacking strategies. We conduct extensive experiments on BPFDex by leveraging more than 3k apps packed by over eight different packers. The results demonstrate that BPFDex successfully bypasses anti-unpacker strategies and unpacks apps packed by various packers, in contrast to other unpackers that can handle at most two packers. Weina Niu, Jiacheng Gong, Song Li 0006, Mingxue Zhang 0001, Xiaosong Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | IIT: Accurate Decentralized Application Identification Through Mining Intra- and Inter-Flow RelationshipsabstractIdentifying Decentralized Applications (DApps) from encrypted network traffic plays an important role in areas such as network management and threat detection. However, DApps deployed on the same platform use the same encryption settings, resulting in DApps generating encrypted traffic with great similarity. In addition, existing flow-based methods only consider each flow as an isolated individual and feed it sequentially into the neural network for feature extraction, ignoring other rich information introduced between flows, and therefore the relationship between different flows is not effectively utilized. In this study, we propose a novel encrypted traffic classification model IIT to heterogeneously mine the potential features of intra- and inter-flows, which contain two types of encoders based on the multi-head self-attention mechanism. By combining the complementary intra- and inter-flow perspectives, the entire process of information flow can be more completely understood and described. IIT provides a more complete perspective on network flows, with the intra-flow perspective focusing on information transfer between different packets within a flow, and the inter-flow perspective placing more emphasis on information interaction between different flows. We captured 44 classes of DApps in the real world and evaluated the IIT model on two datasets, including DApps and malicious traffic classification tasks. The results demonstrate that the IIT model achieves a classification accuracy of greater than 97% on the real-world dataset of 44 DApps, outperforming other state-of-the-art methods. In addition, the IIT model exhibits good generalization in the malicious traffic classification task. Qianwei Meng, Qingjun Yuan, Weina Niu, Yongjuan Wang, Siqi Lu, Guangsong Li, Xiangbin Wang, Wenqi He |
IEEE Trans. Netw. Serv. Manag. | 3 |
| 2024 | SQLStateGuard: Statement-Level SQL Injection Defense Based on Learning-Driven MiddlewareabstractSQL injection is a significant and persistent threat to web services. Most existing protections against SQL injections rely on traffic-level anomaly detection, which often results in high false-positive rates and can be easily bypassed by attackers. This paper introduces SQLStateGuard, the world's first middleware-driven statement-level SQL injection defense approach, to address these issues. The SQLStateGuard uses a custom SQL middleware based on the idea of Runtime Application Self-Protection to capture raw SQL statements. These statements are then analyzed by SQLSG-Net, a database-oriented detection network based on gated linear units. If SQLSG-Net detects malicious SQL statements, the SQL middleware will block them. Experiments show that the detection accuracy of SQLStateGuard exceeds 99%, outperforming existing approaches, and it can identify the type of a specific SQL injection. Additionally, SQLStateGuard has no fingerprint and does not respond to SQL syntax errors, making it more challenging for attackers to gather information. This paper also presents a novel dataset generation process for SQLStateGuard and shares two statement-level SQL injection datasets with the research community, including over 145,000 malicious SQL statements categorized by the type of SQL injection. Xin Liu 0050, Song Li 0006, Weina Niu, Jun Shen 0001, Qingguo Zhou, Xiaokang Zhou |
SoCC | 5 |
| 2024 | Gedss: A Generic Framework to Enhance Model Robustness for Intrusion Detection on Noisy DataabstractTraining a deep neural network-based intrusion detection system requires a large amount of clean labeled data, yet malicious traffic datasets are usually collected from the open-source web community or simulated attack environments, which inevitably contain a large portion of unreliably labeled traffic data. The state-of-the-art methods dealing with label noise combine sample separation and semi-supervised learning (SSL), however, they are hardly usable in the traffic field because traffic data lacks a reasonable data augmentation like image data. To this end, we propose a generic label-noise-resistant framework for malicious traffic detection called Gedss. Unlike previous approaches focusing on data augmentation, our approach improves model performance by enhancing the quality of sample selection and model decision boundaries. The framework contains two parts: sample selection and semi-supervised learning. The sample selection method is presented to divide the original traffic instances into clean ones (labeled set) and noisy ones (unlabeled set). We fit a Jensen-Shannon divergence-based sample prediction loss to a mixture model as the criterion, and the threshold is automatically and dynamically adjusted, which makes our selection mechanism adaptive to various malicious traffic datasets. Besides, a semi-supervised learning method is designed, which uses two networks to jointly predict the pseudo label of the unlabeled set. Considering the class imbalance of divided labeled data, the idea of fine-tuning the models with all data is presented to improve the performance of SSL. Extensive experimental results under different label noise scenarios demonstrate that our approach outperforms state-of-the-art methods. Lingfeng Yao, Anran Hou, Weina Niu, Qingjun Yuan, Junpeng He |
CSCWD | 3 |
| 2024 | Local Augmentation with Functionality-Preservation for Semi-Supervised Graph Intrusion DetectionabstractRecently, deep learning (DL)-driven intrusion detection technology has been rising gradually to reduce the economic and privacy losses caused by the dramatic increase in cyber attacks. Besides exploiting the statistical network traffic features, the inherent attack topologies are also important as they are highly associated with attack behaviors. Thus, many works use graph neural networks (GNN) to make use of them to improve the detection performance. Nevertheless, these topologies contain many isolated IPs that interact with far fewer targets than other non-isolated IPs. Due to the message-passing scheme, information from isolated IPs is less likely to be transmitted in GNN training, resulting in poor model understanding of the traffic connecting these isolated IPs. Therefore, the performance of GNN-based intrusion detection models is not satisfactory for this traffic. To address this problem, we implement a generation algorithm with traffic functionality preservation to improve the performance of GNN-based intrusion detectors by locally augmenting this type of traffic. The proposed method first converts the IP-based graph of the traffic dataset into a line graph, and then utilizes a conditional denoising diffusion probabilistic model to generate new graph snapshots to enhance the expressiveness of isolated IP in GNN message aggregation. We evaluate the performance of our method and compare it with state-of-the-art works on three datasets, i.e., NF-BoT-IoT-V2, NF-ToN-IoT-V2, and NF-CSECIC-IDS2018-V2, and the results show it achieves accuracy of 99.11%(↑0.11%), 93.84%(↑2.47%), and 97.15%(↑3.94%), respectively. Case study shows that the proposed local augmentation can improve detection performance under different isolated thresholds. Moreover, our method can alleviate the over-smoothing problem to a certain extent, and the augmented traffic also possesses better quality. Junpeng He, Hanyue Kong, Shihe Zhang, Xiong Li 0002, Weina Niu, Xiaosong Zhang 0001, Fagen Li |
ICC | 5 |
| 2024 | A Robust Malicious Traffic Detection Framework with Low-quality Labeled DataabstractDeep learning (DL) techniques have been widely applied in detecting malicious activities from network traffic. However, it is challenging to collect a traffic dataset with sufficient correct labels. The generalization ability of DL-based malicious traffic detection systems decreases when training with mislabeled data. Therefore, several methods have been proposed to detect malicious traffic from low-quality labeled training data. These methods divide noisy and clean samples based on the divergence of their prediction loss. However, this simple criterion is not effective on traffic data due to the obfuscation and redundancy nature of malicious traffic. In this paper, we propose a novel two-stage framework for malicious traffic detection from low-quality training data, which mainly consists of noisy sample filtering and label refinement. Firstly, with the help of the small loss criterion, we filter out most of the noisy samples from training data while ensuring that the filtered dataset covers sufficient clean samples. Next, we introduce a double-constrained similarity rule to provide a comprehensive measure of the similarity between samples and construct a topological graph. Lastly, we exploit the topological relations extracted from this graph to refine the labels based on the neighbor consistency criterion. We validate the effectiveness of our framework with a real-world malicious traffic dataset, achieving an accuracy of 90% even with 80% symmetric noise labels. Additionally, results from the publicly available BoT-IoT dataset demonstrate the adaptability of our framework to Internet of Things (IoT) environments. Lingfeng Yao, Weina Niu, Qingjun Yuan, Beibei Li 0002, Xiaosong Zhang 0001 |
ICC | 2 |
| 2024 | Ghost-in-Wave: How Speaker-Irrelative Features Interfere DeepFake Voice DetectorsabstractRecent speech synthesis technology can generate high-quality speech indistinguishable from human speech, thus introducing various security and privacy risks. Numerous recent studies have focused on fake voice detection to address these risks, with many claiming to achieve ideal performance. However, is this really the case? A recent research work introduced Speaker-Irrelative-Features (SiFs), unrelated to the information in speech files but capable of influencing fake detectors. This means that existing detectors may rely on SiFs to a certain extent to distinguish real and fake speech. In this paper, we introduce an evaluation framework to evaluate the influence of SiFs in existing fake voice detectors in depth. We evaluate three SiFs which include background noise, the mute parts before and after voice, and the sampling rate on ASVspoof2019 and FoR. Our results confirm the substantial influence of SiFs on fake voice detection performance, and we delve into the analysis of the underlying mechanisms. Xuan Hai, Xin Liu 0050, Zhaorun Chen, Yuan Tan 0003, Song Li 0006, Weina Niu, Rui Zhou 0005, Qingguo Zhou |
ICME | 6 |
| 2024 | LiScopeLens: An Open-Source License Incompatibility Analysis Tool Based on Scope Representation of License TermsabstractOpen-source software has emerged as a pivotal force in the advancement of information technology. Robust open-source compliance governance is essential for the sustainable and healthy growth of both open-source software and its communities. License incompatibility analysis, in particular, represents a critical challenge hindering the progress of open-source software. Traditional methods of incompatibility analysis often fail to account for diverse usage scenarios or are tailored to a limited subset of scenarios. This limitation obstructing their ability to handle the intricate compatibility arising from varied programming language interactions, leading to a high false positives. Our study embarks from an examination of license exceptions, delving into the incompatibility analysis challenges through extensive empirical research on these exceptions. We discovered that the majority of exceptions are, in fact, detectable. Leveraging this empirical insight, our research further develops the license compatibility analysis model by introducing a new, refined legal terminology representation alongside a novel method for license compatibility reasoning. This approach begins with modeling different scenarios to represent license compatibility variably. Furthermore, based on these modeling outcomes, we have designed and implemented LiScopeLens, a tool capable of discerning dependency behaviors for granular compatibility assessment, starting with binary dependencies. Our experimental findings affirm that LiScopeLens proficiently determines the license compatibility status of open-source software across various usage scenarios, demonstrating its significant practical utility. Ziang Liu 0006, Xin Liu 0050, Yingli Zhang, Song Li 0006, Weina Niu, Qingguo Zhou, Rui Zhou 0005, Xiaokang Zhou |
ISSRE | 6 |
| 2024 | What's the Real: A Novel Design Philosophy for Robust AI-Synthesized Voice DetectionabstractVoice is one of the most widely used media for information transmission in human society. While high-quality synthetic voices are extensively utilized in various applications, they pose significant risks to content security and trust building. Numerous studies have concentrated on AI-synthesized voice detection to mitigate these risks, with many claiming to achieve promising performance. However, recent research has demonstrated that fake voice detectors suffer from serious overfitting to speaker-irrelative features (SiFs) and cannot be used in real-world scenarios. In this paper, we analyze the limitations of existing fake voice detectors and propose a new design philosophy, guiding the detection model to prioritize learning human voice features rather than the difference between the human voice and the synthetic voice. Based on this philosophy, we propose a novel AI-synthesized voice detection framework named SiFSafer, which uses pre-trained speech representation models to enhance the learning of feature distribution in human voices and the adapter fine-tuning to optimize the performance. The evaluation shows that the average EERs of existing fake voice detectors in the ASVspoof datasets can exceed 20% if the SiFs like silence segments are removed, while SiFSafer achieves an EER of less than 8%, indicating that SiFSafer is robust to SiFs and strongly resistant to existing attacks. Xuan Hai, Xin Liu 0050, Yuan Tan 0003, Song Li 0006, Weina Niu, Rui Zhou 0005, Xiaokang Zhou |
ACM Multimedia | 6 |
| 2024 | Nurgle: Exacerbating Resource Consumption in Blockchain State Storage via MPT ManipulationabstractBlockchains, with intricate architectures, encompass various components, e.g., consensus network, smart contracts, decentralized applications, and auxiliary services. While offering numerous advantages, these components expose various attack surfaces, leading to severe threats to blockchains. In this study, we unveil a novel attack surface, i.e., the state storage, in blockchains. The state storage, based on the Merkle Patricia Trie, plays a crucial role in maintaining blockchain state. Besides, we design Nurgle, the first Denial-of-Service attack targeting the state storage. By proliferating intermediate nodes within the state storage, Nurgle forces blockchains to expend additional resources on state maintenance and verification, impairing their performance. We conduct a comprehensive and systematic evaluation of Nurgle, including the factors affecting it, its impact on blockchains, its financial cost, and practically demonstrating the resulting damage to blockchains. The implications of Nurgle extend beyond the performance degradation of blockchains, potentially reducing trust in them and the value of their cryptocurrencies. Additionally, we further discuss three feasible mitigations against Nurgle. At the time of writing, the vulnerability exploited by Nurgle has been confirmed by six mainstream blockchains, and we received thousands of USD bounty from them. Zheyuan He, Zihao Li 0001, Ao Qiao, Xiapu Luo, Xiaosong Zhang 0001, Ting Chen 0002, Shuwei Song, Dijun Liu, Weina Niu |
SP | 9 |
| 2024 | Enabling Robust Android Malicious Packet Capturing and Detection via Android KernelabstractThe prevalence of Android malware presents significant challenges to the security of the Android operating system. Malicious packet detection is a critical technique in combating Android malware. Android malicious packet detection often includes packet capturing and detecting the packets through machine learning or deep learning models. Nonetheless, evolving malware with anti-packet capturing measures, disrupts packet capturing and diminishes the performance of malicious packet detection models. To address these challenges, we propose ePacket, a novel framework for capturing Android malicious packet. By leveraging eBPF technology, ePacket captures application level packets from apps at the Android kernel, effectively circumventing anti-packet capturing strategies employed by malware. Additionally, we have compiled a summary of anti-packet capturing strategies observed in real Android malware. We evaluate ePacket using over 3,000 apps. The results demonstrate that ePacket effectively circumvents anti-packet capturing strategies and captures a significantly higher volume of malicious packet (up to 30% more) compared to other three state-of-the-art tools. Weina Niu, Xinglong Chen, Jiacheng Gong, Kegang Hao |
TrustCom | 2 |
| 2024 | Portfolio-Based Incentive Mechanism Design for Cross-Device Federated LearningabstractIn recent years, there has been a significant increase in attention towards designing incentive mechanisms for fed-erated learning (FL). Tremendous existing studies attempt to design the solutions using various approaches (e.g., game theory, reinforcement learning) under different settings. Yet the design of incentive mechanism could be significantly biased in that clients' performance in many applications is stochastic and hard to estimate. Properly handling this stochasticity motivates this research, as it is not well addressed in pioneering literature. In this paper, we focus on cross-device FL and propose a multi-level FL architecture under the real scenarios. Considering the two properties of clients' situations: uncertainty, correlation, we propose FL Incentive Mechanism based on Portfolio theory (FL- IMP). As far as we are aware, this is the pioneering application of portfolio theory to incentive mechanism design aimed at resolving FL resource allocation problem. In order to more accurately reflect practical FL scenarios, we introduce the Federated Learning Agent-Based Model (FL-ABM) as a means of simulating autonomous clients. FL-ABM enables us to gain a deeper understanding of the factors that influence the system's outcomes. Experimental evaluations of our approach have ex-tensively validated its effectiveness and superior performance in comparison to the benchmark methods. Jiaxi Yang 0003, Cuifang Zhao, Weina Niu, Li-Chuan Tsai |
WCNC | 4 |
| 2024 | Model-agnostic generation-enhanced technology for few-shot intrusion detection
Junpeng He, Lingfeng Yao, Xiong Li 0002, Muhammad Khurram Khan, Weina Niu, Xiaosong Zhang 0001, Fagen Li |
Appl. Intell. | 5 |
| 2024 | FSMFLog: Discovering Anomalous Logs Combining Full Semantic Information and Multifeature FusionabstractIndustrial Internet of Things devices usually use log information to record their runtime status, so log-based anomaly detection can contribute to discovering device failures in time. The first step of log-based anomaly detection is log parsing. However, existing methods mainly extract log templates for analysis, which ignore some words that represent key semantics. Such omissions may cause semantic misunderstandings and further affect the performance of anomaly detection. On the other hand, existing deep learning-based log anomaly detection approaches only consider the sequential relations among log messages, ignoring the log time and type information. In this article, we propose an anomaly detection method called FSMFLog based on full semantic information and multifeature fusion. FSMFLog uses log word lists instead of log templates to represent semantic information. Specifically, the variable part is first removed through preprocessing, and then the log sentences are initially clustered using two heuristic strategies, after which the words in the log content are clustered through the prefix tree structure. By integrating semantic features, time features, and type features, FSMFLog also trains a bidirectional GRU model based on an attention mechanism. Evaluation on 16 real-word log data sets from LogHub shows that FSMFLog achieves a higher log parsing accuracy, outperforming other five state-of-the-art log parsing methods. We also evaluated FSMFLog on two most widely used public data sets (HDFS and BGL), and the results demonstrate the effectiveness of FSMFLog, outperforming the compared approaches using deep learning with an average increase of more than 10% in$F1$-score. Weina Niu, Zimu Li, Zhaoxu He, Aduo Wang, Beibei Li 0002, Xiaosong Zhang 0001 |
IEEE Internet Things J. | 1 |
| 2024 | GraphTunnel: Robust DNS Tunnel Detection Based on DNS Recursive Resolution GraphabstractDNS tunnels, due to their versatility and concealment, have become a preferred method for attackers to execute Command and Control (C&C) attacks, posing a significant security threat to terminal devices. Therefore, the efficient and accurate detection of DNS tunnels is important in reducing the economic losses and privacy risks faced by both enterprises and individuals. Despite notable advancements in the research of intelligent detection of DNS tunnels, existing model-based approaches predominantly concentrate on the surface-level features of domain names or packet payloads. This narrow focus leads to low detection accuracy when dealing with unknown DNS tunnel attacks and traffic from wildcard DNS. Furthermore, these methods struggle with accurately identifying DNS tunneling tools, complicating the task of swiftly locating and mitigating malware for analysts. This paper proposes GraphTunnel, a framework based on graph neural networks for detecting DNS tunnels and identifying tunneling tools. It delves into the correlations among DNS resolutions to construct paths that represent the recursive resolution process of DNS. By using central nodes that denote the gateways, these paths are connected and transformed into graph structures. Concurrently, it employs GraphSage to aggregate the features of nodes and their edges in the graph, enabling effective detection of DNS tunnels. Additionally, GraphTunnel utilizes the G2M algorithm to capture the statistical features of nodes in the graph and maps them into grayscale images, which are then processed by a CNN for multi-class identification of DNS tunneling tools. Experimental results demonstrate that in non-wildcard DNS scenarios, GraphTunnel achieves a 100% accuracy in DNS tunnel detection, encompassing unknown DNS tunnels. Even in high false-positive environments caused by wildcard DNS, GraphTunnel maintains an F1-Score of 99.78%. Moreover, GraphTunnel can identify DNS tunneling tools with an accuracy rate exceeding 98.57%, enhancing the rapid mitigation capabilities of emergency responders in dealing with malicious DNS tunnels. Guangyuan Gao, Weina Niu, Jiacheng Gong, Dujuan Gu, Song Li 0006, Mingxue Zhang 0001, Xiaosong Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2024 | Sensitive Behavioral Chain-Focused Android Malware Detection Fused With AST SemanticsabstractThe proliferation of Android malware poses a substantial security threat to mobile devices. Thus, achieving efficient and accurate malware detection and malware family identification is crucial for safeguarding users’ individual property and privacy. Graph-based approaches have demonstrated remarkable detection performance in the realm of intelligent Android malware detection methods. This is attributed to the robust representation capabilities of graphs and the rich semantic information. The function call graph (FCG) is the most widely used graph in intelligent Android malware detection. However, existing FCG-based malware detection methods face challenges, such as the enormous computational and storage costs of modeling large graphs. Additionally, the ignorance of code semantics also makes them susceptible to structured attacks. In this paper, we proposed AndroAnalyzer, which embeds abstract syntax tree (AST) code semantics while focusing on sensitive behavior chains. It leverages FCGs to represent the macroscopic behavior of the application, and employs structured code semantics to represent the microscopic behavior of functions. Furthermore, we proposed the sensitive function call graph (SFCG) generation algorithm to narrow down the analysis scope to sensitive function calls, and the AST vectorization algorithm (AST2Vec) to capture structured code semantics. Experimental results demonstrate that the proposed SFCG generation algorithm noticeably reduces graph size while ensuring robust detection performance. AndroAnalyzer outperforms the baseline methods in binary and multiclass classification tasks, achieving F1-scores of 99.21% and 98.45% respectively. Moreover, AndroAnalyzer (trained with samples of 2010-2018) exhibits good generalization capabilities in detecting samples of 2019-2022. Jiacheng Gong, Weina Niu, Song Li 0006, Mingxue Zhang 0001, Xiaosong Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | Black-box Word-level Textual Adversarial Attack Based On Discrete Harris Hawks OptimizationabstractNeural network-based applications are prone to being fooled by adversarial examples due to the natural vulnerability of deep neural networks (DNNs). Textual adversarial attacks are particularly challenging due to the discreteness between texts. The adversarial examples crafted by word-level textual attacks which are typically treated as optimization problems in black-box scenarios perform better in human evaluation. Existing approaches have struggled to balance the success rate with the time consuming, mainly because the chosen optimization algorithm is not efficient enough. In this paper, we propose a method to generate textual adversarial examples called Discrete Harris Hawk Optimization (DHHO). We set up three operations for handling discrete data, which are applied to each stage of the Harris Hawk Optimization (HHO) to enable it to solve optimization problems in discrete space. By attacking BiLSTM and BERT on two benchmark data sets, we conduct extensive experiments to evaluate our attack method with a success rate of up to 98% and a reduction of time is at least 50%. Moreover, the experimental results also show that our adversarial examples can ensure high quality and transferability. Tianrui Wang, Weina Niu, Kangyi Ding, Lingfeng Yao, Pengsen Cheng, Xiaosong Zhang 0001 |
CSCWD | 2 |
| 2023 | Fuzzing Logical Bugs in eBPF Verifier with Bound-Violation IndicatorabstracteBPF is widely used in Microsoft, Google, and Facebook because it is able to extend kernel without modifying the kernel source code. Nevertheless, vulnerabilities in kernel with eBPF will affect the stability and security of information system. Fuzzing has proven to be an effective approach for finding kernel bugs since it requires minimal knowledge about the target. However, two main challenges exist in discovering eBPF logical bugs: generating input that satisfies all eBPF instruction semantic requirements, and detecting the eBPF logical bug states. We remove highly semantically demanding and unnecessary instructions by analyzing the impact of the instructions to obtain a higher verification pass rate to address the first challenge. We also develop a bound-violation indicator to address the second challenge based on our analysis of eBPF logical bug patterns. We manually introduce 10 recently fixed logical bugs in eBPF for evaluation, and the experimental results show that we can effectively find 9 of them, while Syzkaller fails on all of them. In addition, 4 new bugs have been fixed for upstream Linux based on our work, and 3 functional issues have been reported. Youlin Li, Weina Niu, Yukun Zhu, Jiacheng Gong, Beibei Li 0002, Xiaosong Zhang 0001 |
ICC | 2 |
| 2023 | DEML: Data-Enhanced Meta-Learning Method for IoT APT Traffic Detection
Weina Niu, Qingjun Yuan, Lingfeng Yao, Junpeng He, Xiaosong Zhang 0001 |
ICDF2C (1) | 2 |
| 2023 | A Novel Generation Method for Diverse Privacy Image Based on Machine LearningabstractAbstract In recent years, deep neural networks have been extensively applied in various fields, and face recognition is one of the most important applications. Artificial intelligence has reached or even surpassed human capabilities in many fields. However, while artificial intelligence application provides convenience to the human lives, it also leads to the risk of privacy leaking. At present, the privacy protection technology for human faces has received extensive attention. Research goals of face privacy protection technology mainly include providing face anonymization and data availability protection. Existing methods usually have insufficient anonymity and they are not easy to control the degree of image distortion, which makes it difficult to achieve the purpose of privacy protection. Moreover, they do not explicitly perform diversity preservation of attributes such as emotions, expressions and ethnicities, so they cannot perform data analysis tasks on non-identity attributes. This paper proposes a diverse privacy face image generation algorithm based on machine learning, called DIVFGEN. This algorithm comprehensively considers image distortion, identity mapping distance loss and emotion classification loss; transforms the privacy protection target into the problem of generating adversarial examples based on the recognition model; and uses an adaptive optimization algorithm to generate anonymity and diversity of privacy images. The experimental results show that on the Cohn-Kanade+ dataset, our algorithm can reduce the probability of facial recognition by the neural network when it accurately classifies sentiment, from 98.6% to 4.8%. Weina Niu, Yuheng Luo, Kangyi Ding, Xiaosong Zhang 0001, Beibei Li 0002 |
Comput. J. | 1 |
| 2023 | GCDroid: Android Malware Detection Based on Graph Compression With Reachability Relationship Extraction for IoT DevicesabstractWith the widespread popularity of Internet of Things (IoT) devices based on the Android system, the amount of Android malware targeting IoT devices continues to increase, causing great economic losses. Accordingly, efficient and accurate Android malware detection methods are particularly important. Recently, many Android malware detection and classification methods have been proposed, but most of them ignore the deep relationships among software. In this article, we propose a graph compression algorithm with reachability relationship extraction (GCRR) and design an Android malware detection and classification method called GCDroid based on this algorithm. A theoretical analysis shows that GCRR can reasonably extract the reachability relationships among APKs and compress a large heterogeneous APK–API relationship graph into a homogeneous APKs graph. Experiments show that GCDroid based on GCRR greatly reduces the required time consumption while improving detection accuracy. Compared with the existing excellent static Android malware detection methods, GCDroid improves upon their detection accuracies by 1.53%–39.13% on different data sets and outperforms the benchmark methods in terms of Android malware classification. Furthermore, compared with those of the baseline methods that are similar to GCDroid, GCDroid’s time consumption for model training and other aspects is only one-tenth as high or even less. Weina Niu, Xiong Li 0002, Xiaosong Zhang 0001 |
IEEE Internet Things J. | 1 |
| 2022 | LogTracer: Efficient Anomaly Tracing Combining System Log Detection and Provenance GraphabstractInformation systems have penetrated into all areas of social life, however, unknown threats represented by APT attacks pose serious challenges to their security. In recent years, approaches based on log analysis and provenance graph have been extensively used in the anomaly detection and tracing of malicious attacks. However, traditional method has low detection accuracy, high complexity and low efficiency. To address those shortcomings, we propose an efficient anomaly tracing approach (LogTracer), which combines system log detection and provenance graph together. The proposed LogTracer extracts the attack path from provenance graph, which is constructed with the anomaly degrees of the system logs anomaly detection results. Compar-ative experiments with OmegaLog, NoDoze and ALchemist are conducted on a simulated dataset with 16 attack types totaling 290 million logs. The experimental results show that our method approximately 5.4x, 0.2x and 7.2x faster than these three methods in processing efficiency, and its malicious node coverage rate reaches 98.1%. Weina Niu, Zhenqi Yu, Zimu Li, Beibei Li 0002, Runzi Zhang, Xiaosong Zhang 0001 |
GLOBECOM | 1 |
| 2022 | Targeted Anonymization: A Face Image Anonymization Method for Unauthorized ModelsabstractAs an important biometric feature of every person, face data has faced serious risks of leakage in recent years. Lawbreakers can use face recognition systems (FRS) to analyze the leaked face data and then correlate other private information, causing serious privacy leaks. For security reasons, we hope our face images can only be recognized by the organizations' authorized models. To achieve this goal, this work proposes a targeted face image anonymization method that only enables anonymization for unauthorized facial recognition models, whilst authorized models, human eyes can still accurately recognize faces. Our method mainly uses transfer-based adversarial attacks to achieve anonymization. On this basis, we propose constraints for generating targeted anonymization samples and boundary walking strategy, focusing on improving the anonymization for unauthorized models while guaranteeing the accurate recognition of authorized models. Local experiments prove that our method can reduce the recognition probability of unauthorized models while guaranteeing the correctness of authorized models. Finally, we apply our approach to an online face recognition API and experimentally demonstrate that our approach can significantly reduce the recognition accuracy of the commercial face recognition model. Kangyi Ding, Xiaolei Liu 0001, Weina Niu, Xiaosong Zhang 0001 |
ICME | 4 |
| 2022 | A Fine-Grained Approach for Vulnerabilities Discovery Using Augmented Vulnerability Signatures
Xiaoxiao Zhou, Weina Niu, Xiaosong Zhang 0001, Rui-dong Chen, Yan Wang 0103 |
KSEM (3) | 2 |
| 2022 | IDROP: Intelligently detecting Return-Oriented Programming using real-time execution flow and LSTMabstractReturn-Oriented Programming (ROP) has become one of the most widely used attack techniques for software vulnerability exploitation. Existing ROP detection methods fall into two types: hardware-based methods and software-based methods. The former is strongly dependent on specific hardware architectures and difficult to deploy. Although the latter can alleviate these problems, limited by the selection of features and thresholds, it cannot effectively discover neither variant ROP nor delayed ROP. In this work, we propose an intelligent detection method at runtime and implement the corresponding prototype system, IDROP, which uses real-time execution flow and LSTM to discovery ROP and its variants. Specifically, IDROP analyzes the differences between program execution flows that are independent of the ROP feature thresholds. Firstly, the Aspect Oriented Programming (AOP) is utilized to instrument the tested program, and the sliding window mechanism is applied to screen out suspicious program execution flow snapshots. Then, these suspicious execution flow snapshots are vectorized through data representation techniques. Finally, we build and train an LSTM model to discover ROP. Furthermore, we evaluate the performance of IDROP on a dataset consisting of 6000+ samples. The experimental results show that IDROP is effective in detecting ROP attacks, variant ROP and delayed ROP with an accuracy of 98%, 93% and 80%, respectively. In addition, IDROP has negligible space overhead and low performance overhead, which is similar to that of only using Pin for detection (about additional 2.5 times the program execution time before instrumentation). Weina Niu, Zhiqin Duan, Beibei Li 0002, Xiaosong Zhang 0001 |
TrustCom | 2 |
| 2022 | Trine: Syslog anomaly detection with three transformer encoders in one generative adversarial network
Zhenfei Zhao, Weina Niu, Xiaosong Zhang 0001, Runzi Zhang, Zhenqi Yu, Cheng Huang 0003 |
Appl. Intell. | 2 |
| 2022 | Uncovering APT malware traffic using deep learning combined with time sequence and association analysis
Weina Niu, Yibin Zhao 0004, Xiaosong Zhang 0001, Yujie Peng, Cheng Huang 0003 |
Comput. Secur. | 1 |
| 2022 | Attacker Traceability on Ethereum through Graph AnalysisabstractSince the Ethereum virtual machine is Turing complete, Ethereum can implement various complex logics such as mutual calls and nested calls between functions. Therefore, Ethereum has suffered a lot of attacks since its birth, and there are still many attackers active in Ethereum transactions. To this end, we propose a traceability method on Ethereum, using graph analysis to track attackers. We collected complete user transaction data to construct the graph and analyzed data on several harmful attacks, including reentry attacks, short address attacks, DDoS attacks, and Ponzi contracts. Through graph analysis, we found accounts that are strongly associated with these attacks and are still active. We have done a systematic analysis of these accounts to analyze their threats. Finally, we also analyzed the correlation between the information collected through RPC and these accounts and finally found that some accounts can find their IP addresses. Weina Niu, Xuhan Liao, Xiaosong Zhang 0001, Beibei Li 0002, Zheyuan He |
Secur. Commun. Networks | 2 |
| 2021 | FedVANET: Efficient Federated Learning with Non-IID Data for Vehicular Ad Hoc NetworksabstractThe vehicular ad hoc networks (VANETs) play a significant role in intelligent transportation systems (ITS). In recent years, federated learning (FL) has been widely used in VANETs to preserve the privacy-sensitive data, such as vehicle locations, drivers' driving patterns, on-board camera data, etc. However, conventional FL faces the challenges of non-independent and identically distributed (Non-IID) data and high communication overheads in VANETs. To address these challenges, we propose a novel FL framework for VANETs, named FedVANET, where a hierarchical inner-cluster FL model and a weighted inter-cluster cycling update algorithm are, respectively, developed. Extensive experiments demonstrate the high efficiency of the FedVANET in inner-cluster communications, effectiveness in handling Non-IID data, and robustness in dynamic VANET topologies. Beibei Li 0002, Yukun Jiang 0001, Weina Niu, Peiran Wang |
GLOBECOM | 4 |
| 2021 | Malware on Internet of UAVs Detection Combining String Matching and Fourier TransformationabstractAdvanced persistent threat (APT), with intense penetration, long duration, and high customization, has become one of the most grievous threats to cybersecurity. Furthermore, the design and development of Internet-of-Things (IoT) devices often do not focus on security, leading APT to extend to IoT, such as the Internet of emerging unmanned aerial vehicles (UAVs). Whether malware with attack payload can be successfully implanted into UAVs or not is the key to APT on the Internet of UAVs. APT malware on UAVs establishes communication with the command and control (C&C) server to achieve remote control for UAVs-aware information stealing. Existing effective methods detect malware by analyzing malicious behaviors generated during C&C communication. However, APT malware usually adopts a low-traffic attack mode, a large amount of normal traffic is mixed in each attack step, to avoid virus checking and killing. Therefore, it is difficult for traditional malware detection methods to discover APT malware on UAVs that carry weak abnormal signals. Fortunately, we found that most APT attacks use domain name system (DNS) to locate C&C server of malware for information transmission periodically. This behavior will leave some records in the network flow and DNS logs, which provides us with an opportunity to identify infected internal UAVs and external malicious domain names. This article proposes an APT malware on the Internet of UAVs detection method combining string matching and Fourier transformation based on DNS traffic, which is able to handle encrypted and obfuscated traffic due to packet payloads independence. We preprocessed the collected network traffic by converting DNS timestamps of DNS request to strings and used the trained random forest model to discover APT malware domain names based on features extracted through string-matching-based periodicity detection and Fourier transformation-based periodicity detection. The proposed method has been evaluated on the data set, including part of normal domains from the normal traffic and malicious domains marked by security experts from APT malware traffic. Experimental results have shown that our proposed detection method can achieve the accuracy of 94%, which is better than the periodicity detection algorithm alone. Moreover, the proposed method does not need to set the confidence to filter the periodicity with high confidence. Weina Niu, Jian'An Xiao, Xiaosong Zhang 0001, Xiaojiang Du, Mohsen Guizani |
IEEE Internet Things J. | 1 |
| 2021 | Transaction-based classification and detection approach for Ethereum smart contract
Xiaolei Liu 0001, Ting Chen 0002, Xiaosong Zhang 0001, Weina Niu |
Inf. Process. Manag. | 6 |
| 2021 | NEDetector: Automatically extracting cybersecurity neologisms from hacker forums
Jiaxing Cheng, Cheng Huang 0003, Zhouguo Chen, Weina Niu |
J. Inf. Secur. Appl. | 5 |
| 2021 | A low-query black-box adversarial attack based on transferability
Kangyi Ding, Xiaolei Liu 0001, Weina Niu, Xiaosong Zhang 0001 |
Knowl. Based Syst. | 3 |
| 2021 | HTTP-Based APT Malware Infection Detection Using URL Correlation AnalysisabstractAPT malware exploits HTTP to establish communication with a C & C server to hide their malicious activities. Thus, HTTP-based APT malware infection can be discovered by analyzing HTTP traffic. Recent methods have been dependent on the extraction of statistical features from HTTP traffic, which is suitable for machine learning. However, the features they extract from the limited HTTP-based APT malware traffic dataset are too simple to detect APT malware with strong randomness insufficiently. In this paper, we propose an innovative approach which could uncover APT malware traffic related to data exfiltration and other suspect APT activities by analyzing the header fields of HTTP traffic. We use the Referer field in the HTTP header to construct a web request graph. Then, we optimize the web request graph by combining URL similarity and redirect reconstruction. We also use a normal uncorrelated request filter to filter the remaining unrelated legitimate requests. We have evaluated the proposed method using 1.48 GB normal HTTP flow from clickminer and 280 MB APT malware HTTP flow from Stratosphere Lab, Contagiodump, and pcapanalysis. The experimental results have shown that the URL-correlation-based APT malware traffic detection method can correctly detect 96.08% APT malware traffic, and its recall rate is 98.87%. We have also conducted experiments to compare our approach against Jiang’s method, MalHunter, and BotDet, and the experimental results have confirmed that our detection approach has a better performance, the accuracy of which reached 96.08% and the F1 value increased by more than 5%. Weina Niu, Jiao Xie, Xiaosong Zhang 0001, Xin-Qiang Li, Rui-dong Chen, Xiaolei Liu 0001 |
Secur. Commun. Networks | 1 |
| 2020 | A Black-Box Attack on Neural Networks Based on Swarm Evolutionary Algorithm
Xiaolei Liu 0001, Kangyi Ding, Yang Bai 0011, Weina Niu |
ACISP | 5 |
| 2020 | Preprocessing Method for Encrypted Traffic Based on Semisupervised ClusteringabstractThe explosive growth in network traffic in recent times has resulted in increased processing pressure on network intrusion detection systems. In addition, there is a lack of reliable methods for preprocessing network traffic generated by benign applications that do not steal users’ data from their devices. To alleviate these problems, this study analyzed the differences between benign and malicious traffic produced by benign applications and malware, respectively. To fully express these differences, this study proposed a new set of statistical features for training a clustering model. Furthermore, to mine the communication channels generated by benign applications in batches, a semisupervised clustering method was adopted. Using a small number of labeled samples, our method aggregated historical network traffic into two types of clusters. The cluster that did not contain labeled malicious samples was regarded as a benign traffic cluster. The experimental results were compared using four types of clustering algorithms. The density-based spatial clustering of applications with noise (DBSCAN) clustering algorithm was selected to mine benign communication channels. We also compared our method with two other methods, and the results demonstrated that the benign channels mined through our method were more reliable. Finally, using our method, 1,811 benign transport layer security (TLS) channels were mined from 18,357 TLS communication channels. The number of flows carried by these benign channels comprised 65.37% of the entire network flows, and no malicious flow was included in our results, which proves the effectiveness of our method. Rongfeng Zheng, Weina Niu, Liang Liu 0009, Shan Liao |
Secur. Commun. Networks | 3 |
| 2019 | A Lockable Abnormal Electromagnetic Signal Joint Detection AlgorithmabstractWith the development of computers and network technologies, network security has gradually become a global problem. Network security defenses need to be carried out not only on the Internet, but also on other communication media, such as electromagnetic signals. Existing electromagnetic signal communication is easily intercepted or infiltrated. In order to effectively detect the abnormal electromagnetic signal to find out the specific location, then classify it, it is necessary to study the way of communication. The existing electromagnetic signal detection accuracy is low and cannot be located. Considering the characteristics of different power sources in different locations, combined with spark streaming technology and machine learning classification technology, a joint platform for electromagnetic signal anomaly detection based on big data analysis is proposed. The electromagnetic signal is abnormally detected by feature comparison and small signal analysis, and the position and number between the signal sources are determined by three-point positioning and signal attenuation. The experimental results show that the method can detect abnormal electromagnetic signals and classify abnormal electromagnetic signals well, the accuracy rate can reach 95%, and the positioning accuracy can reach 89%. Weina Niu, Xiaolei Liu 0001, Xiaosong Zhang 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2019 | An Insider Threat Detection Approach Based on Mouse Dynamics and Deep LearningabstractIn the current intranet environment, information is becoming more readily accessed and replicated across a wide range of interconnected systems. Anyone using the intranet computer may access content that he does not have permission to access. For an insider attacker, it is relatively easy to steal a colleague’s password or use an unattended computer to launch an attack. A common one-time user authentication method may not work in this situation. In this paper, we propose a user authentication method based on mouse biobehavioral characteristics and deep learning, which can accurately and efficiently perform continuous identity authentication on current computer users, thus to address insider threats. We used an open-source dataset with ten users to carry out experiments, and the experimental results demonstrated the effectiveness of the approach. This approach can complete a user authentication task approximately every 7 seconds, with a false acceptance rate of 2.94% and a false rejection rate of 2.28%. Weina Niu, Xiaosong Zhang 0001, Xiaolei Liu 0001 |
Secur. Commun. Networks | 2 |
| 2019 | Using XGBoost to Discover Infected Hosts Based on HTTP TrafficabstractIn recent years, the number of malware and infected hosts has increased exponentially, which causes great losses to governments, enterprises, and individuals. However, traditional technologies are difficult to timely detect malware that has been deformed, confused, or modified since they usually detect hosts before being infected by malware. Host detection during malware infection can make up for their deficiency. Moreover, the infected host usually sends a connection request to the command and control (C&C) server using the HTTP protocol, which generates malicious external traffic. Thus, if the host is found to have malicious external traffic, the host may be a host infected by malware. Based on the background, this paper uses HTTP traffic combined with eXtreme Gradient Boosting (XGBoost) algorithm to detect infected hosts in order to improve detection efficiency and accuracy. The proposed approach uses a template automatic generation algorithm to generate feature templates for HTTP headers and uses XGBoost algorithm to distinguish between malicious traffic and normal traffic. We conduct a performance analysis to demonstrate that our approach is efficient using dataset, which includes malware traffic from MALWARE-TRAFFIC-ANALYSIS.NET and normal traffic from UNSW-NB 15. Experimental results show that the detection speed is about 1859 HTTP traffic per second, and the detection accuracy reaches 98.72%, and the false positive rate is less than 1%. Weina Niu, Xiaosong Zhang 0001 |
Secur. Commun. Networks | 1 |
| 2018 | Tor anonymous traffic identification based on gravitational clustering
Zhihong Rao, Weina Niu, Xiaosong Zhang 0001, Hongwei Li 0001 |
Peer-to-Peer Netw. Appl. | 2 |
| 2018 | Weighted Domain Transfer Extreme Learning Machine and Its Online Version for Gas Sensor Drift Compensation in E-Nose SystemsabstractMachine learning approaches have been widely used to tackle the problem of sensor array drift in E‐Nose systems. However, labeled data are rare in practice, which makes supervised learning methods hard to be applied. Meanwhile, current solutions require updating the analytical model in an offline manner, which hampers their uses for online scenarios. In this paper, we extended Target Domain Adaptation Extreme Learning Machine (DAELM_T) to achieve high accuracy with less labeled samples by proposing a Weighted Domain Transfer Extreme Learning Machine, which uses clustering information as prior knowledge to help select proper labeled samples and calculate sensitive matrix for weighted learning. Furthermore, we converted DAELM_T and the proposed method into their online learning versions under which scenario the labeled data are selected beforehand. Experimental results show that, for batch learning version, the proposed method uses around 20% less labeled samples while achieving approximately equivalent or better accuracy. As for the online versions, the methods maintain almost the same accuracies as their offline counterparts do, but the time cost remains around a constant value while that of offline versions grows with the number of samples. Zhiyuan Ma 0001, Guangchun Luo, Ke Qin, Nan Wang 0003, Weina Niu |
Wirel. Commun. Mob. Comput. | 5 |
| 2018 | Resetting Your Password Is Vulnerable: A Security Study of Common SMS-Based Authentication in IoT DeviceabstractFirmware vulnerability is an important target for IoT attacks, but it is challenging, because firmware may be publicly unavailable or encrypted with an unknown key. We present in this paper an attack on Short Message Service (SMS for short) authentication code which aims at gaining the control of IoT devices without firmware analysis. The key idea is based on the observation that IoT device usually has an official application (app for short) used to control itself. Customer needs to register an account before using this app, phone numbers are usually suggested to be the account name, and most of these apps have a common feature, calledReset Your Password, that uses an SMS authentication code sent to customer phone to authenticate the customer when he forgot his password. We found that an attacker can perform brute‐force attack on this SMS authentication code automatically by overcoming several challenges, then he can steal the account to gain the control of IoT devices. In our research, we have implemented a prototype tool, calledSACIntruder, to enable performing such brute‐force attack test on IoT devices automatically. We evaluated it and successfully found 12 zero‐day vulnerabilities including smart lock, sharing car, smart watch, smart router, etc. We also discussed how to prevent this attack. Dong Wang 0018, Xiaosong Zhang 0001, Jiang Ming 0002, Ting Chen 0002, Chao Wang 0021, Weina Niu |
Wirel. Commun. Mob. Comput. | 6 |
| 2016 | Improving data field hierarchical clustering using Barnes-Hut algorithm
Zhongliu Zhuo, Xiaosong Zhang 0001, Weina Niu, Guowu Yang, Jingzhong Zhang |
Pattern Recognit. Lett. | 3 |