Yu Wen 0001

dblp:58/3305-1 · DBLP profile ↗
← Back
34ranked-venue papers
1as first author
26since 2021 · last 2025
0000-0002-0658-0742ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 22 · 18 since 2021Systems, architecture and hardware · 4 · 2 since 2021Computer networks · 4 · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Enhancing the Robustness of LiDAR-based Object Detection under Disappearing Attacks
abstract
Autonomous driving systems rely on LiDAR-based 3D object detection to identify obstacles. Recent studies have shown that detectors are susceptible to disappearing attacks, leading to missed detections and potential vehicle collisions. However, improving the adversarial robustness of 3D object detection against such attacks remains an open question. Our work seeks to bridge this gap by proposing an effective defense strategy for 3D object detection under both black-box and white-box disappearing attacks. Specifically, we formulate the problem of defending against disappearing attacks in 3D object detection as a bilevel min-max problem and introduce a novel approach, DART, to train machine learning models robust to disappearing attacks. Our extensive evaluations demonstrate that DART outperforms general adversarial defense methods, reducing the attack success rate by over 85% while achieving a better balance between robustness and performance, and meeting the real-time requirements of autonomous vehicles.
Huiying Wang, Lisong Zhang, Yu Wen 0001
ICASSP4
2025 FineGCP: Fine-grained dependency graph community partitioning for attack investigation
Yanfei Hu, Yu Wen 0001, Shuailou Li, Dan Meng 0002
Comput. Secur.3
2025 InstPro: Provenance-Based Transient Execution Attack Detection and Investigation on Instruction Execution Traces
abstract
Transient execution attacks (TEAs) are a serious threat to modern computing systems. While software/hardware hardening techniques have been proposed to mitigate the threat, developing detection techniques remains imperative, as they hold promise for flexible extension to address new variants, ease of deployment, and minimal system impact. Existing detection techniques face the following three limitations: unstable information sources, lack of explanation for attack scenarios, and limited training data. To address the limitations, we proposeInstPro, a TEA detection system thatidentifies a TEA program while providing an explanation of the attack scenario, based on instruction execution traces. Specifically,InstProfirst extractsprincipled cluesthat represent instruction sequences semantically close to attack abstraction. These clues provide high-level visualizations of TEA steps. Then,InstProcorrelates the clues into aclue provenance graphby reasoning about their causal dependencies, which provides a concise provenance representation. Finally,InstProreconstructs a scenario graph by using theInfoSubgraphsthat represent the information flows among principled clues. These InfoSubgraphs are more likely to capture a set of crucial principled clues that work together to represent the attack scenario. Our evaluations based on 5 datasets show thatInstProeffectively performs TEA detection and investigation.
Yu Wen 0001, Yanna Wu, Dan Meng 0002
IEEE Trans. Dependable Secur. Comput.2
2024 ASSC: Adaptively Stochastic Smoothing based Adversarial Robustness Certification of Malware Detection Models
abstract
In recent years, the stochastic smoothing technique for adversarial robustness certification of artificial intelligence(AI) models is mainly applied in computer vision domain, with fewer applications in security-related domains such as malware. In the malware detection task, the noisy samples generated by stochastic smoothing algorithms may break the functionality of the malware. As a result, malware may be executable in the real world. In addition, stochastic smoothing-based robustness certification methods of AI detection models may be costly. In this paper, an adaptively stochastic smoothing based approach for adversarial robustness certification of malware detection models, which we call ASSC, is proposed. ASSC constructs a smoothing model by adding discrete Bernoulli noise only to specific constraint positions. Moreover, an adaptive smoothing model is constructed to improve the certification efficiency. Experiments have shown that, our proposed method can at least reduce the time of robustness certification average by 80% on two public datasets, which can significantly reduce the certification runtime. In addition, the experimental results can also prove that, overall, ASSC can guarantee a higher certification accuracy than other methods under the same certification radius.
Yu Wen 0001, Xinqiang Zhao
CSCWD3
2024 Interpretable Risk-aware Access Control for Spark: Blocking Attack Purpose Behind Actions
abstract
The big data platform supports powerful data re-trieval and mining analysis, providing users with seamless access to extensive data for valuable insights. However, the increasing access to sensitive data raises privacy concerns. Existing studies utilize access control mechanisms to ensure secure data authorization. Nevertheless, previous approaches are deficient in facilitating risk control during real-time query processing and fail to elucidate the details of attack access. To address these limitations, we propose a novel Interpretable Risk-aware Access Control (IRAAC) for Spark - the advanced distributed engine for large-scale data computing in big data ecosystems. IRAAC utilizes the sequence representation techniques and contrastive learning idea from Natural Language Processing to learn patterns of attack queries for extracting critical attack subqueries. In terms of attack investigation, IRAAC designs specific templates to encourage large language models (LLMs) to provide a comprehensive delineation of potential query access risks.
Tao Xue 0003, Shuailou Li, Yu Wen 0001
ICCD6
2024 TurboLog: A Turbocharged Lossless Compression Method for System Logs via Transformer
abstract
System logs provide valuable information for intrusion detection and forensic analysis. To counter cyber attacks, enterprises widely deploy monitoring software and record fine-grained system events within the operating system through system logs. However, system logs can pile up in large quantities as the complex nature of modern computer systems and the covert nature of cyber attacks. Storing system logs efficiently is an important and challenging task. Lossless compression techniques provide an intuitive idea of reducing the size of system logs. Recently, researchers have applied Deep Neural Networks (DNNs) in designing compression strategies, achieving remarkable outcomes. Unfortunately, general-purpose lossless compression methods fail to achieve effective compression due to their inadequate extraction of structural redundancy from system logs. Certain methods opt to add a few preprocessing steps before log compression to enhance compression efficiency. Nevertheless, their RNN-based models face challenges in efficiently utilizing GPU resources and can not handle parallel computation. Therefore, their methods still demand a significant compression time. In this paper, we propose TurboLog, a novel transformer-based compression method specifically designed for system logs. To address the above deficiencies, TurboLog initiates by executing a series of preprocessing steps to reduce the structural redundancy and numerical redundancy within raw logs. Subsequently, Tur-boLog uses a transformer-based probability estimator to model log data in parallel. Finally, TurboLog combines the probability estimation of log data with arithmetic coding to accomplish log compression. We use a large public dataset to evaluate TurboLog. The results show that TurboLog can achieve the highest compression ratio of 65.25×. Furthermore, compared to the state-of-the-art method, TurboLog decreases compression time by 75% while concurrently enhancing the compression ratio by 20%. In conclusion, TurboLog achieves much better space savings than existing methods and points an optional direction for future research.
Baoming Chang, Fengxi Zhou, Yu Wen 0001
IJCNN5
2024 StreamDP: Continual Observation of Real-world Data Streams with Differential Privacy
abstract
The real-time collection and query analysis of dynamic data streams have become increasingly common and important, yet the protection of sensitive private information remains a pressing challenge. Differential privacy, as the gold standard for protecting personal data privacy, has been widely studied and applied. However, existing mechanisms mostly focus on static datasets and specific simple stream queries. This paper presents StreamDP, a novel framework designed to achieve differential privacy for complex real-world stream queries. We introduce the observation-prediction mechanism that predicts statistics such as join attribute frequency using observations and truncates the data stream based on the predicted threshold. Then we design operation-oriented recursive sensitivity calculation rules and employ a hierarchy algorithm for noise perturbation. Extensive experimental evaluations on multiple real-world datasets and distributed stream processing benchmarks show that StreamDP can support various complex real-world data stream queries/applications with high utility and low-performance overhead.
Shuailou Li, Yu Wen 0001, Lisong Zhang, Dan Meng 0002
IPCCC2
2024 DyCom: A Dynamic Community Partitioning Technique for System Audit Logs
abstract
To address the ever-evolving network threats, system audit logs have become a crucial data source for threat analysis. While current log-based threat detection methods have significant potential in identifying malicious activities, they face limitations in capturing dynamic attack behaviors and revealing complete attack activities.To address these issues, we introduces DyCom, a dynamic graph partitioning technique based on audit logs. Our method incrementally partitions the system log streaming into multiple communities with security semantics, allowing the use of graph community summarization techniques to provide a summary of key activities within each community, thereby aiding security experts in understanding system activities. This method retains all attack activities, reduces analysts’ workload, and keeps them continuously informed about system activities. By focusing on process entities and constructing intimate process events, DyCom effectively reduces the storage cost of log data while ensuring the security semantics of the log communities. Additionally, DyCom employs temporal graph networks to dynamically represent system entities, ensuring real-time monitoring of system activities. Evaluations using the open-source DARPA TC dataset and our simulated datasets demonstrate that DyCom can accurately partition large-scale dynamic system entities into distinct communities, with improvements in precision, recall, and F1 score by 0.71, 0.31, and 0.65 respectively, compared to baseline methods, highlighting its practical potential in threat analysis.
Yanfei Hu, Shuailou Li, Lisong Zhang, Yu Wen 0001, Dan Meng 0002
TrustCom7
2024 CarePlus: A general framework for hardware performance counter based malware detection under system resource competition
Yanfei Hu, Wenchao Xue 0001, Yanlong Zhao 0004, Yu Wen 0001
Comput. Secur.5
2023 ProDE: Interpretable APT Detection Method Based on Encoder-decoder Architecture
abstract
The detection and analysis of Advanced Persistent Threats (APTs) are pivotal for contemporary network security. Provenance graphs, constructed from audit logs, offer a wealth of contextual information to identify and analyze threats and are popular in APT detection field. However, existing approaches frequently fall short in offering explanatory capabilities for their detection results, placing an additional burden on security analysts. Confronted with coarse-grained detection outcomes, analysts must delve into provenance graphs or audit logs to precisely pinpoint attack entities and events, which can significantly delay the response to threats. In this paper, we propose ProDE, a novel approach that enhances APT detection by providing interpretable results using an encoder-decoder architecture. ProDE initiates the detection process by comparing the encoded representations of the true graph and the predicted graph. Upon detecting abnormalities, the encoder-decoder model is able to decode the encodings into provenance graphs, thereby revealing inconsistencies between the decoded graph and real graph that serve as interpretable results. We evaluate ProDE on two widely used datasets, while taking into account the detection performance, the result shows ProDE can provided the more detailed detection results which provide the interpretation for analysts compared with existing approaches.
Fengxi Zhou, Baoming Chang, Yu Wen 0001, Dan Meng 0002
ICPADS3
2023 DAMUS: Adaptively Updating Hardware Performance Counter Based Malware Detector Under System Resource Competition
abstract
Hardware performance counter based malware detection (HMD) model that learns HPC-level behavior by using machine learning or deep learning algorithms has been widely researched in various application scenarios. However, the program's HPC-level behavior is easily affected due to system resource competition, which leaves counter based malware detection out-of-date. Unfortunately, current research could not adaptively update HMD model. In this paper, we propose DAMUS, a distribution-aware model updating strategy to adaptively update counter based malware detection model. Specifically, we first design an autoencoder with contrastive learning to map existing samples into a low-dimensional space for better calculating distributions. Second, in the low-dimensional space, the distribution characteristics are calculated for further judging the drift of testing samples. Finally, based on the total determined drifts of testing samples and a threshold, a decision could be given on whether the counter based malware detection model needs to be updated. We evaluate DAMUS by testing HMD model on datasets collected under benchmark application environment and actual server environment with different resource types or pressure levels. The experimental results show the advantages of DAMUS over existing updating strategies in promoting model updating. We also demonstrate its overhead spent on the task of malware detection.
Yanfei Hu, Shuailou Li, Yu Wen 0001
ISCC5
2023 HUND: Enhancing Hardware Performance Counter Based Malware Detection Under System Resource Competition Using Explanation Method
abstract
Hardware performance counter (HPC) has been widely used in malware detection because of its low access overhead and the ability of revealing dynamic behavior during program's execution. However, HPC based malware detection (HMD) suffers from performance decline due to HPC's non- determinism caused by resource competition. Current work enables malware detection under resource competition but still leaves misclassifications. In this paper, we propose HUND, a framework for improving the detection ability of HMD models under resource competition. To this end, we first introduce an explanation module to make the program's prediction interpretable and accurate on the whole. We then design a rectification module for troubleshooting HMDMs' errors by generating modified samples and lowering the effects of false classified instances on model decision. We evaluate HUND by performing HMD models two datasets of HPC-level behaviors. The experimental results show HUND explains HMDMs with high fidelity and HUND's effectiveness in troubleshooting the errors of HMDMs.
Yanfei Hu, Shuailou Li, Xu Cheng 0001, Yu Wen 0001
ISCC5
2023 PRISPARK: Differential Privacy Enforcement for Big Data Computing in Apache Spark
abstract
Differential privacy has emerged as a gold standard privacy definition due to its persuasive mathematical guarantee. While various data protection mechanisms provide differential privacy for SQL queries of RDBMSs, enforcing differential privacy for big data platforms needs to be further researched. This work presents Prispark, which enforces differential privacy for Spark - the advanced distributed engine for large-scale data computing in big data ecosystems where sensitive data is often processed. Prispark targets to support various data processing (i.e., relational and unstructured queries) on Spark. In particular, to calculate a tighter sensitivity bound and improve the utility of results, we design the overall statistics estimation algorithm for estimating the upper bound of statistics with the filter condition, and propose a novel fine-grained operation-oriented rules set for calculating sensitivity of various relational and unstructured queries. Moreover, we propose a general differential privacy mechanism, Prispark, a suite including Prisparksql and Prisparkdag. We enforce Prisparksql at the Catalyst optimization layer for relational queries in Spark SQL and Prisparkdag at the RDD execution layer for unstructured queries in Spark core. Finally, we experimentally evaluate Prispark on TPC-H, TPC-DS, PigMix benchmarks, and real-world dataset LANL. The experimental results suggest that Prispark supports various applications/queries while improving the utility of all query results by orders of magnitude with negligible performance overhead.
Shuailou Li, Yu Wen 0001, Tao Xue 0003, Yanna Wu, Dan Meng 0002
SRDS2
2023 Towards Dynamic Backdoor Attacks against LiDAR Semantic Segmentation in Autonomous Driving
abstract
LiDAR perception is widely deployed in high-level autonomous vehicles (AVs) to gain accurate information about the driving environment, where 3D semantic segmentation plays a critical role as it can provide more fine-grained scene understanding than other tasks. The mainstream LiDAR perception systems mainly adopt deep neural networks (DNNs) to achieve good performance. However, densely annotating LiDAR point clouds and training complex LiDAR segmentation models are time-consuming and resource-intensive. It is common for ordinary self-driving developers to outsource the data annotation or model training task to third-party platforms, which could inevitably expose a natural attack injection point. In this paper, we propose BadLiSeg, the first work to investigate backdoor attacks against LiDAR semantic segmentation. Specifically, we present a general attack strategy based on which the attacker can inject a dynamic backdoor into the victim model by constructing a trigger pattern pool and a location pool. Afterward, the attacker can choose a common physical object (e.g., drone and traffic sign) as the trigger and place it around an arbitrary location in a selected area to easily fool the backdoored LiDAR segmentation model. Our extensive experiments on five representative segmentation models and one public dataset demonstrate that BadLiSeg can not only achieve a high attack success rate but also maintain the normal segmentation performance of the backdoored model. We further show the effectiveness of BadLiSeg on some practical attack scenarios collected from a high-fidelity simulator.
Yu Wen 0001, Xu Cheng 0001
TrustCom2
2023 BadLiDet: A Simple Backdoor Attack against LiDAR Object Detection in Autonomous Driving
abstract
Autonomous vehicles (AVs) widely deploy LiDAR-based 3D object detection to accurately perceive and understand the surrounding environment. The mainstream LiDAR detection systems primarily adopt deep neural networks (DNNs) to achieve satisfactory performance. However, annotating large amounts of LiDAR data and training complex LiDAR detection models are time-consuming and resource-intensive. A common practice for some individual developers and self-driving companies is to outsource data annotation or model training tasks to third-party platforms, which could expose a natural attack injection point. In this paper, we propose BadLiDet, a simple yet effective backdoor attack against LiDAR object detection in autonomous driving. Specifically, we present a model-agnostic attack strategy that enables an attacker to inject a shape-independent backdoor into the victim model by poisoning its training data. Afterward, the attacker can choose an ordinary object in different shapes as the trigger and place it around a specific location to easily deceive the LiDAR detection system. Our extensive experiments on five representative models and a public dataset demonstrate that BadLiDet can achieve a high attack success rate while preserving the utility of victim models. We further show the effectiveness of BadLiDet on the continuous attack scenario collected from a high-fidelity simulator. Moreover, the end-to-end simulation evaluation on an open-source self-driving platform shows that BadLiDet can cause a 100% vehicle collision rate.
Yu Wen 0001, Huiying Wang, Xu Cheng 0001
TrustCom2
2023 Representation-enhanced APT Detection Using Contrastive Learning
abstract
Advanced Persistent Threats (APTs) are low-and-slow attack patterns and difficult to detect due to their strong concealment. Recently, provenance-based method demonstrates promising prospect in APT detection. However, existing approaches suffer from following several limitations. First, rule-based models heavily rely on domain-specific knowledges and sophisticated matching mechanism from rules to system logs. Second, current anomaly-based techniques require attack-free data to train model and hard to acquire fine-grained detection results to locate attack events. An important reason for undermining the detection capacity of model is the notorious dependency explosion problem and data imbalance problem. Both of problems make the vector representations of benign events and attack events similar, therefore, the model can not differentiate them well. In this paper, we propose DeepDist, which can enhance the entity vector representation by contrastive learning. In this way, we force the projection of benign and attack entities into different feature regions to improve the detection ability of the model. Then we correlate these detected entities to constitute attack events. During our evaluation against thirteen real APT attack scenarios of two datasets, DeepDist shows the detection results with high accuracy.
Fengxi Zhou, Baoming Chang, Yu Wen 0001, Dan Meng 0002
TrustCom3
2023 SparkAC: Fine-Grained Access Control in Spark for Secure Data Sharing and Analytics
abstract
With the development of computing and communication technologies, an extremely large amount of data has been collected, stored, utilized, and shared, while new security and privacy challenges arise. Existing access control mechanisms provided by big data platforms have limitations in granularity and expressiveness. In this article, we present SparkAC, a novel access control mechanism for secure data sharing and analysis in Spark. In particular, we first propose apurpose-aware access control(PAAC) model, which introduces new concepts ofdata processing purposeanddata operation purposeand an automatic purpose analysis algorithm that identifies purposes from data analytics operations and queries. Moreover, we develop a unified access control mechanism that implements PAAC model in two modules. GuardSpark++ supports structured data access control in Spark Catalyst and GuardDAG supports unstructured data access control in Spark core. Finally, we evaluate GuardSpark++ and GuardDAG with multiple data sources, applications, and data analytics engines. Experimental results show that SparkAC provides effective access control functionalities with very small (GuardSpark++) or medium (GuardDAG) performance overhead.
Tao Xue 0003, Yu Wen 0001, Bo Luo, Gang Li 0009, Yingjiu Li, Yanfei Hu, Dan Meng 0002
IEEE Trans. Dependable Secur. Comput.2
2022 DEPCOMM: Graph Summarization on System Audit Logs for Attack Investigation
abstract
Causality analysis generates a dependency graph from system audit logs, which has emerged as an important solution for attack investigation. In the dependency graph, nodes represent system entities (e.g., processes and files) and edges represent dependencies among entities (e.g., a process writing to a file). Despite the promising early results, causality analysis often produces a large graph (> 100,000 edges) and it is a daunting task for security analysts to inspect such a large graph for attack investigation. To address challenges in attack investigation, we propose DEPCOMM, a graph summarization approach that generates a summary graph from a dependency graph by partitioning a large graph into process-centric communities and presenting summaries for each community. Specifically, each community consists of a set of intimate processes that cooperate with each other to accomplish certain system activities (e.g., file compression), and the resources (e.g., files) accessed by these processes. Within a community, DEPCOMM further identifies redundant edges caused by less-important and repetitive system activities, and perform compression on these edges. Finally, DEPCOMM generates the summary for each community using the InfoPaths that represent the information flows across communities. These InfoPaths are more likely to capture a set of attack-related processes that work together to achieve certain malicious goals. Our evaluations on real attacks ($\sim 150$ million events) demonstrate that DEPCOMM generates 18.4 communities on average for a dependency graph, which is $\sim 70 \times$ smaller than the original graph. Our compression further reduces the edges in each community to 32.1 on average. Compared with the 9 state-of-the-art community detection algorithms, on average, DEPCOMM achieves a $2.29\times$ better F1-score than these algorithms in detecting communities. Through cooperating with the automatic techniques HOLMES, DEPCOMM can identify attack-related communities by a recall of 96.2%. Our case studies on the real attacks also demonstrate DEPCOMM’s effectiveness in facilitating attack investigation.
Zhiqiang Xu 0001, Pengcheng Fang, Changlin Liu, Xusheng Xiao, Yu Wen 0001, Dan Meng 0002
SP5
2022 A General Backdoor Attack to Graph Neural Networks Based on Explanation Method
abstract
Graph neural networks (GNNs) have achieved significant performance in many applications (e.g., social spammer detection and Facebook page classification). Recently, backdoor attacks pose a new security threat to the training process of GNNs. Attackers intend to inject backdoors into GNNs, such that the attacked model performs well on benign samples, whereas its prediction will be changed to target label when the trigger is present. However, existing works on backdoor attacks could not attack arbitrary nodes effectively and achieve high performance on both nodes with discrete features and that with continuous features. In this paper, we propose a general backdoor attack to GNNs which could effectively attack arbitrary nodes while keeping other nodes unaffected as much as possible. We first ensure some important edges for a target node by utilizing edge explanation method and remove these edges if the edge exists before (vice versa). Then we ensure important features by using feature explanation method and use a neural network to optimize these features to obtain feature triggers (i.e., perturbation). Finally, we inject edge and feature triggers into target nodes to implement attack. We evaluate our attack from two aspects including generality and effectiveness. For the first part, our attack is effective on both datasets with discrete features and continuous features. For the second part, we randomly select different node subset with different size (e.g., 20, 50, 100 nodes) as target nodes and our attack achieves higher performance against other two state-of-the-art attacks.
Yu Wen 0001, Yanfei Hu
TrustCom5
2022 Deepro: Provenance-based APT Campaigns Detection via GNN
abstract
Advanced Persistent Threats (APTs) are typically sophisticated, stealthy and long-term attacks that are difficult to be detected and investigated. Recently proposed provenance graph based on system audit logs has become an important approach for APT detection and investigation. However, existing provenance-based approaches that either require rules based on expert knowledge or cannot pinpoint attack events in a provenance graph still cannot effectively mitigate APT attacks. In this paper, we present Deepro, a provenance-based APT campaign detection approach that not only effectively detects attack-relevant entities in a provenance graph but also precisely recovers APT campaigns based on the detected entities. Specifically, Deepro first customizes a general purpose GNN (Graph Neural Network) model to represent and detect process nodes in a provenance graph through automatically learning different patterns of attack behaviors and benign behaviors using the rich contextual information in the provenance graph. Then, Deepro further detects attack-relevant file and network entities according to their data dependencies with the detected process nodes. Finally, Deepro recovers APT campaigns through correlating detected entities based on their causality relationships in the provenance graph. We evaluated Deepro with ten real-world APT attacks. The evaluation result shows that Deepro can effectively detect attack events with an average 98.81% F1-score and thus produces precise provenance sub-graphs of APT attacks.
Yu Wen 0001, Yanna Wu, Dan Meng 0002
TrustCom2
2021 ACGVD: Vulnerability Detection Based on Comprehensive Graph via Graph Neural Network with Attention
Chunfang Li, Shuailou Li, Yanna Wu, Yu Wen 0001
ICICS (1)6
2021 Generating Adversarial Point Clouds on Multi-modal Fusion Based 3D Object Detection Model
Huiying Wang, Huixin Shen, Yu Wen 0001, Dan Meng 0002
ICICS (1)4
2021 Black-Box Buster: A Robust Zero-Shot Transfer-Based Adversarial Attack Method
Yu Wen 0001, Dan Meng 0002
ICICS (2)4
2021 Malicious Login Detection Using Long Short-Term Memory with an Attention Mechanism
Yanna Wu, Fucheng Liu, Yu Wen 0001
IFIP Int. Conf. Digital Forensics3
2021 FederatedReverse: A Detection and Defense Method Against Backdoor Attacks in Federated Learning
abstract
Federated learning is a secure machine learning technology proposed to protect data privacy and security in machine learning model training. However, recent studies show that federated learning is vulnerable to backdoor attacks, such as model replacement attacks and distributed backdoor attacks. Most backdoor defense techniques are not appropriate for federated learning since they are based on entire data samples that cannot be hold in federated learning scenarios. The newly proposed methods for federated learning sacrifice the accuracy of models and still fail once attacks persist in many training rounds. In this paper, we propose a novel and effective detection and defense technique called FederatedReverse for federated learning. We conduct extensive experimental evaluation of our solution. The experimental results show that, compared with the existing techniques, our solution can effectively detect and defend against various backdoor attacks in federated learning, where the success rate and duration of backdoor attacks can be greatly reduced and the accuracies of trained models are almost not reduced.
Yu Wen 0001, Shuailou Li, Fucheng Liu, Dan Meng 0002
IH&MMSec2
2021 DeepMal: maliciousness-Preserving adversarial instruction learning against static malware detection
abstract
Abstract Outside the explosive successful applications of deep learning (DL) in natural language processing, computer vision, and information retrieval, there have been numerous Deep Neural Networks (DNNs) based alternatives for common security-related scenarios with malware detection among more popular. Recently, adversarial learning has gained much focus. However, unlike computer vision applications, malware adversarial attack is expected to guarantee malwares’ original maliciousness semantics. This paper proposes a novel adversarial instruction learning technique, DeepMal, based on an adversarial instruction learning approach for static malware detection. So far as we know, DeepMal is the first practical and systematical adversarial learning method, which could directly produce adversarial samples and effectively bypass static malware detectors powered by DL and machine learning (ML) models while preserving attack functionality in the real world. Moreover, our method conducts small-scale attacks, which could evade typical malware variants analysis (e.g., duplication check). We evaluate DeepMal on two real-world datasets, six typical DL models, and three typical ML models. Experimental results demonstrate that, on both datasets, DeepMal can attack typical malware detectors with the mean F1-score and F1-score decreasing maximal 93.94% and 82.86% respectively. Besides, three typical types of malware samples (Trojan horses, Backdoors, Ransomware) prove to preserve original attack functionality, and the mean duplication check ratio of malware adversarial samples is below 2.0%. Besides, DeepMal can evade dynamic detectors and be easily enhanced by learning more dynamic features with specific constraints.
Jinghui Xu, Shuangshuang Liang, Yanna Wu, Yu Wen 0001, Dan Meng 0002
Cybersecur.5
2020 GuardSpark++: Fine-Grained Purpose-Aware Access Control for Secure Data Sharing and Analysis in Spark
abstract
With the development of computing and communication technologies, extremely large amount of data has been collected, stored, utilized, and shared, while new security and privacy challenges arise. Existing platforms do not provide flexible and practical access control mechanisms for big data analytics applications. In this paper, we present GuardSpark++, a fine-grained access control mechanism for secure data sharing and analysis in Spark. In particular, we first propose a purpose-aware access control (PAAC) model, which introduces new concepts of data processing/operation purposes to conventional purpose-based access control. An automatic purpose analysis algorithm is developed to identify purposes from data analytics operations and queries, so that access control could be enforced accordingly. Moreover, we develop an access control mechanism in Spark Catalyst, which provides unified PAAC enforcement for heterogeneous data sources and upper-layer applications. We evaluate GuardSpark++ with five data sources and four structured data analytics engines in Spark. The experimental results show that GuardSpark++ provides effective access control functionalities with a very small performance overhead (average 3.97%).
Tao Xue 0003, Yu Wen 0001, Bo Luo, Yanfei Hu, Yingjiu Li, Gang Li 0009, Dan Meng 0002
ACSAC2
2020 MLTracer: Malicious Logins Detection System via Graph Neural Network
abstract
Malicious login, especially lateral movement, has been a primary and costly threat for enterprises. However, there exist two critical challenges in the existing methods. Specifically, they heavily rely on a limited number of predefined rules and features. When the attack patterns change, security experts must manually design new ones. Besides, they cannot explore the attributes' mutual effect specific to login operations. We propose MLTracer, a graph neural network (GNN) based system for detecting such attacks. It has two core components to tackle the previous challenges. First, MLTracer adopts a novel method to differentiate crucial attributes of login operations from the rest without experts' designated features. Second, MLTracer leverages a GNN model to detect malicious logins. The model involves a convolutional neural network (CNN) to explore attributes of login operations, and a co-attention mechanism to mutually improve the representations (vectors) of login attributes through learning their login-specific relation. We implement an evaluation of such an approach. The results demonstrate that MLTracer significantly outperforms state-of-the-art methods. Moreover, MLTracer effectively detects various attack scenarios with a remarkably low false positive rate (FPR).
Fucheng Liu, Yu Wen 0001, Yanna Wu, Shuangshuang Liang, Xihe Jiang, Dan Meng 0002
TrustCom2
2020 An Approach for Poisoning Attacks against RNN-Based Cyber Anomaly Detection
abstract
In the face of the increasingly complex Internet environment, the traditional intrusion detection system is difficult to cope with the unknown variety of attacks. People hope to find reliable anomaly detection technology to help improve the security of cyberspace. The rapid development of artificial intelligence technology provides new development opportunities for anomaly detection technology, and the anomaly detection system based on deep learning performs well in some studies. However, neural networks are highly dependent on data quality, and a small number of poisoned samples injected into the data set will have a huge impact on the results. The online abnormal threat detection system based on deep learning is likely to be attacked by poisoning due to the need for continuous data collection and training. We propose a poisoning attack method using adversarial samples to resist the anomaly detection system based on an unsupervised deep neural network, which can destroy the neural network with as few samples as possible. We verified the effectiveness of poisoning attacks on the network security data set of los alamos national laboratory and further demonstrated its generality on other abnormal detection data set.
Jinghui Xu, Yu Wen 0001, Dan Meng 0002
TrustCom2
2020 Built-in Security Computer: Deploying Security-First Architecture Using Active Security Processor
abstract
Continually disclosed vulnerabilities reveal that traditional computer architecture lacks the consideration of security. This article proposes a security-first architecture, with an Active Security Processor (ASP) integrated to conventional computer architectures. To reduce the attack surface of ASP and improve the security of the whole system, the ASP is physically isolated from Computation Processor Units (CPU) with an asymmetric address space, which enables both ASP and CPU to run their operating system and applications independently in their own memory space. Furthermore, the ASP, which has the highest privilege (Super Root) of the whole system, possesses two advantageous features. First, the ASP can efficiently access all CPU resources and collect multi-dimensional information to monitor malicious behaviors, meanwhile, the CPU cannot access the ASP's private resources in any way. Second, instead of being scheduled by CPUs, the ASP can actively manage the security mechanisms employed in either CPUs or the ASP. Based on the security-first architecture, we introduce several typical security tasks running on ASP. With different considerations in terms of system overhead, complexity and performance, we also explore four typical system-level implementations for integrating the ASP to the security-first architecture. The first-generation ASP was designed and implemented based on the 40nm technology, and a security computer system was implemented based on it. Evaluations on this real hardware platform demonstrate that the security-first architecture can protect the system effectively with minor performance impacts on computing workloads.
Dan Meng 0002, Rui Hou 0001, Bibo Tu, Xiaoqi Jia, Yu Wen 0001
IEEE Trans. Computers8
2019 Log2vec: A Heterogeneous Graph Embedding Based Approach for Detecting Cyber Threats within Enterprise
abstract
Conventional attacks of insider employees and emerging APT are both major threats for the organizational information system. Existing detections mainly concentrate on users' behavior and usually analyze logs recording their operations in an information system. In general, most of these methods consider sequential relationship among log entries and model users' sequential behavior. However, they ignore other relationships, inevitably leading to an unsatisfactory performance on various attack scenarios. We propose log2vec, a heterogeneous graph embedding based modularized method. First, it involves a heuristic approach that converts log entries into a heterogeneous graph in the light of diverse relationships among them. Next, it utilizes an improved graph embedding appropriate to the above heterogeneous graph, which can automatically represent each log entry into a low-dimension vector. The third component of log2vec is a practical detection algorithm capable of separating malicious and benign log entries into different clusters and identifying malicious ones. We implement a prototype of log2vec. Our evaluation demonstrates that log2vec remarkably outperforms state-of-the-art approaches, such as deep learning and hidden markov model (HMM). Besides, log2vec shows its capability to detect malicious events in various attack scenarios.
Fucheng Liu, Yu Wen 0001, Dongxue Zhang, Xihe Jiang, Xinyu Xing 0001, Dan Meng 0002
CCS2
2017 Efficient tamper-evident logging of distributed systems via concurrent authenticated tree
abstract
Secure logging as an indispensable part of any secure system in practice is well-understood by both academia and industry. However, providing security for audit logs on an untrusted machine in a large distributed system is still a challenging task. The emergence and wide availability of log management tools prompted plenty of work in the security community that allows clients or auditors to verify integrity of the log data. Most recent solutions to this problem focus on the space-efficiency or public verifiability of forward security. Unfortunately, existing secure audit logging schemes have significant performance limitations that make them impractical for realtime large-scale distributed applications: Existing cryptographic hashing is computationally expensive for logging in task intensive or resource-constrained systems especially to prove individual log events, while Merkle-tree approach has fundamental limitations when face with highly concurrent, large-scale log streams due to its serially appending feature. The verification step of Merkle-tree based approach requiring a logarithmic number of hash computations is becoming a bottleneck to improve the overall performance. There is a huge gap between the flux of log streams collected and the computational efficiency of integrity verification in the large-scale distributed systems. In this work, we develop a novel scheme, performance of which favorably compares with the existing solutions. The performance guarantees that we achieve stem from a novel data structure called concurrent authenticated tree, which allows log events concurrently appending and removes the need to wait for append operations to complete sequentially. We implement a prototype using chameleon hashing based on discrete log and Merkle history tree. A comprehensive experimental evaluation of the proposed and existing approaches is used to validate the analytical models and verify our claims. The results demonstrate that our proposed scheme verifying in a concurrent way is significantly more efficient than the previous tree-based approach.
Fangxiao Ning, Yu Wen 0001, Dan Meng 0002
IPCCC2
2014 Automated Power Control for Virtualized Infrastructures
Yu Wen 0001, Weiping Wang 0005, Li Guo 0001, Dan Meng 0002
J. Comput. Sci. Technol.1
2007 A layered design methodology of cluster system stack
abstract
The application range of cluster has expanded beyond scientific computing, but the present cluster system software fails to provide a flexible architecture to promote code reuse and facilitate building cluster system software for different computing contexts, most of which are developed from scratch case by case, or integrated or packaged with “the best practice”. In this paper, we have proposed a layered design methodology to build cluster system stack with different layers concentrating on different functions, and developed common sets of core service as reusing framework for different computing context. Following this methodology, we have built Phoenix-a complete cluster system stack for both scientific and business computing, which is verified and deployed on Dawning 4000A super computer for scientific computing and other cluster systems for business computing. The qualitative evaluation and our practices show the design methodology of Phoenix has advantages over other methodologies.
Jianfeng Zhan, Lei Wang 0004, Bibo Tu, Yu Wen 0001, Yuansheng Chen, Wei Zhou 0019, Dan Meng 0002, Ninghui Sun
CLUSTER5