VLDB 2026 Research / reviewers in the wild / expert
Yulai Xie 0002
dblp:69/6471-2
· DBLP profile ↗
33ranked-venue papers
9as first author
25since 2021 · last 2026
0000-0001-5757-4396ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Security and privacy · 5 · 2 first-author · 4 since 2021Computer networks · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Secret Caching Sauce for High-Performance Secure Memory
Xu Jiang 0005, Xueliang Wei, Yifei Qu, Dan Feng 0001, Yulai Xie 0002, Wei Tong 0001 |
HPCA | 5 |
| 2026 | Pontus: Identifying intrusions from massive logs via accurate provenance clustering and efficient graph serialization with minimum provenance lossabstractIdentifying intrusions from massive logs has long been a great challenge. To address this issue, this paper proposes Pontus, a novel host-based intrusion detection method via accurate provenance clustering, efficient graph serialization and classification with minimum provenance loss. Pontus first utilizes a novel multi-round label propagation algorithm (MLPA) based on overlapping community discovery to cluster the behavior instances that constitute user behavior accurately. In this way, Pontus can analyze behavior instances to extract behavior features effectively while reducing the analysis workload. Then, Pontus enables efficient graph serialization via neighbor node aggregation to convert the behavior instance into vectors while maximizing the retention of provenance information. Finally, Pontus uses a hybrid method that combines the convolutional autoencoder with Bisecting Kmeans clustering to accurately extract the provenance features of behavior instances to identify host-based intrusions. The experimental results show that compared with the state-of-the-art methods, Pontus’s accuracy increases by an average of 0.243, and F1-score increases by an average of 0.201, with small runtime overheads. Yulai Xie 0002, Heyu Zhang, Yafeng Wu, Dan Feng 0001, Pan Zhou 0001, Avani Wildani |
Expert Syst. Appl. | 3 |
| 2025 | Learning from Few Samples: A Novel Approach for High-Quality Malcode GenerationabstractIntrusion Detection Systems (IDS) play a crucial role in network security defense.However, a significant challenge for IDS in training detection models is the shortage of adequately labeled malicious samples.To address these issues, this paper introduces a novel semi-supervised framework GANGRL-LLM, which integrates Generative Adversarial Networks (GANs) with Large Language Models (LLMs) to enhance malicious code generation and SQL Injection (SQLi) detection capabilities in few-sample learning scenarios.Specifically, our framework adopts a collaborative training paradigm where: (1) the GAN-based discriminator improves malicious pattern recognition through adversarial learning with generated samples and limited real samples; and (2) the LLM-based generator refines the quality of malicious code synthesis using reward signals from the discriminator.The experimental results demonstrate that even with a limited number of labeled samples, our training framework is highly effective in enhancing both malicious code generation and detection capabilities.This dual enhancement capability offers a promising solution for developing adaptive defense systems capable of countering evolving cyber threats. Haijian Ma, Daizong Liu, Xiaowen Cai 0001, Pan Zhou 0001, Yulai Xie 0002 |
EMNLP | 5 |
| 2025 | FedLTH: A Privacy-preserving Federated Learning Framework with Model Pruning on Edge ClientsabstractAlthough Federated Learning (FL) enables distributed clients to cooperatively train deep learning models without sharing local data, the iterative FL training process imposes considerable computation and communication overheads on clients. Especially in cloud-edge collaboration situations, heterogeneous and resource-limited edge clients can become a bottleneck for FL. In this paper, we propose FedLTH (Federated Learning with the Lottery Ticket Hypothesis), an FL framework based on the Lottery Ticket Hypothesis and adaptive differential privacy, which aims to improve communication and computing efficiency and privacy security for resource-limited edge clients. First, the pruning rate of each client is set according to their respective resource constraints. The server divides the clients into groups with balanced data distribution and similar pruning rates to ensure the convergence of the global model. Then, a structured model pruning method based on the Lottery Ticket Hypothesis is introduced. Each client group participates in a pruning phase to reduce the computing overhead of clients. Last, an adaptive differential privacy algorithm is designed to preserve client data privacy and improve model accuracy. Through experiments on multiple datasets and non-IID scenarios, we show the effectiveness of FedLTH in privacy preservation and reducing computation and communication overheads. Heyu Zhang, Yulai Xie 0002, Shengshan Hu, Peisong He, Jun Zheng 0017, Dan Feng 0001 |
ICDCS | 2 |
| 2025 | Fit the Distribution: Cross-Image/Prompt Adversarial Attacks on Multimodal Large Language ModelsabstractAlthough Multimodal Large Language Models (MLLMs) have demonstrated remarkable achievements in recent years, they remain vulnerable to adversarial examples that result in harmful responses. Existing attacks typically focus on optimizing adversarial perturbations for a certain multimodal image-prompt pair or fixed training dataset, which often leads to overfitting. Consequently, these perturbations fail to remain malicious once transferred to attack unseen image-prompt pairs, suffering from significant resource costs to cover the diverse multimodal inputs in complicated real-world scenarios. To alleviate this issue, this paper proposes a novel adversarial attack on MLLMs based on distribution approximation theory, which models the potential image-prompt input distribution and adds the same distribution-fitting adversarial perturbation on multimodal input pairs to achieve effective cross-image/prompt transfer attacks. Specifically, we exploit the Laplace approximation to model the Gaussian distribution of the image and prompt inputs for the MLLM, deriving an estimate of the mean and covariance parameters. By sampling from this approximated distribution with Monte Carlo mechanism, we efficiently optimize and fit a single input‑agnostic perturbation over diverse image‑prompt pairs, yielding strong universality and transferability. Extensive experiments are conducted to verify the strong adversarial capabilities of our proposed attack against prevalent MLLMs spanning a spectrum of images/prompts. Hai Yan, Haijian Ma, Xiaowen Cai 0001, Daizong Liu, Zenghui Yuan, Xiaoye Qu, Jianfeng Dong, Runwei Guan, Hongyang He, Yulai Xie 0002, Pan Zhou 0001 |
NeurIPS | 11 |
| 2025 | Efficient intrusion detection via heterogeneous graph attention networks and parallel provenance analysis
Yulai Xie 0002, Shixun Zhao, Pan Zhou 0001, Dan Feng 0001, Avani Wildani, Yafeng Wu |
Comput. Networks | 2 |
| 2025 | Angus: efficient active learning strategies for provenance based intrusion detectionabstractAbstract As modern attack methods become more concealed and complex, obtaining many labeled samples in big data streams is difficult. Active learning has long been used to achieve better intrusion detection performance by using only a small number of training samples. Intrusion behaviors can be described by provenance graphs that record the dependency relationships between intrusion processes and the infected files. It is a challenge to develop active learning strategies that consider defining and selecting the most valuable provenance and ensure that the strategy for querying provenance is efficient. We present Angus, an active learning framework for provenance-based intrusion detection. We propose two novel active learning strategies: the most similar graph query strategy and the maximum difference query strategy. They either select samples to update the training set according to similarities of provenance graphs or preferentially select samples with low redundancy and large differences from the current training set. Besides, we also improve the above query strategies by using the parallel query to reduce detection time overheads. The experiments on various real-world applications demonstrate their performance and efficiency. Yulai Xie 0002, Dan Feng 0001, Jinyuan Liang, Yafeng Wu |
Cybersecur. | 2 |
| 2025 | SST-LOF: Container Anomaly Detection Method Based on Singular Spectrum Transformation and Local Outlier FactorabstractIn recent years, the use of container cloud platforms has experienced rapid growth. However, because containers are operating-system-level virtualization, their isolation is far less than that of virtual machines, posing considerable challenges for multi-tenant container cloud platforms. To address the issues associated with current container anomaly detection algorithms, such as the difficulty in mining periodic features and the high rate of false positives due to noisy data, we propose an anomaly detection method named SST-LOF, based on singular spectrum transformation and the local outlier factor. Our method enhances the traditional Singular Spectrum Transformation (SST) algorithm to meet the needs of streaming unsupervised detection. Furthermore, our method improves the calculation mode of the anomaly score of the Local Outlier Factor algorithm (LOF) and reduces false positives of noisy data with dynamic sliding windows. Additionally, we have designed and implemented a container cloud anomaly detection system that can perform real-time, unsupervised, streaming anomaly detection on containers quickly and accurately. The experimental results demonstrate the effectiveness and efficiency of our method in detecting anomalies in containers in both simulated and real cloud environments. Shilei Bu, Minpeng Jin, Jie Wang 0115, Yulai Xie 0002, Liangkang Zhang |
IEEE Trans. Cloud Comput. | 4 |
| 2025 | More Granular, Less Trust: Enforcing Intra-Process Isolation With Arm CCA in an Untrusted Management EnvironmentabstractWith the increasing adoption of confidential computing, security-sensitive applications are often deployed in confidential virtual machines (CVMs), which reduce reliance on third-party cloud providers. However, privilege attacks originating from the OS remain a significant threat in these environments. Existing finer-grained isolation schemes, such as SHELTER (USENIX SEC’23), provide process-level protection but are still vulnerable to intra-process attacks and potential collusion between the OS and intra-process adversaries. Many current intra-process isolation techniques continue to depend on the OS to manage and enforce isolation domains, leading to a large Trusted Computing Base (TCB). This gap highlights the need for more granular, less trust-dependent confidential computing solutions. In this paper, we present CCAegis, a system that extends the Arm Confidential Compute Architecture (CCA) to enforce intra-process isolation of sensitive data and operations, safeguarding them from both intra-process adversaries and the OS. We employ static analysis to track the flow of sensitive data and identify functions that handle such data. Permission-switching instructions are inserted at the function call and return points, adjusting permissions via the Granule Protection Table (GPT) to ensure that only designated functions can access the isolated data. Notably, CCAegis places trust solely in the Secure Monitor, which configures the GPTs and manages domain switching, thereby minimizing the TCB. We implemented CCAegis on both an official emulator and a real development board to assess its performance. Our experimental results show that CCAegis effectively isolates sensitive data and operations, with performance overheads ranging from 1.01× to 1.43× compared to the original version across real-world cryptographic workloads. Shiqi Liu 0006, Zhouqi Jiang, Jie Wang 0138, Kun Sun 0001, Yulai Xie 0002 |
IEEE Trans. Inf. Forensics Secur. | 7 |
| 2025 | A Novel Approach to Construct 1-D Discrete Complex Variable Chaotic Systems and Its ApplicationabstractConventional real-valued 1-D chaotic models are constrained by three fundamental limitations: Restricted chaotic regimes, susceptibility to dynamic degradation in finite-precision implementations, and inherent tradeoffs between security assurance and computational efficiency. These constraints significantly limit the applicability of chaos-based systems. This article presents a novel approach for constructing 1-D discrete complex-variable chaotic systems (1D-DCVCS). The proposed methodology establishes a flexible architecture that enables the derivation of 1D-DCVCS with guaranteed positive Lyapunov exponents. Extensive numerical experiments confirm that the constructed systems exhibit rich dynamical properties, while retaining strong chaotic behavior even in low-precision implementations. Hardware validation via field-programmable gate array implementation confirms the practical viability of the proposed approach. To address existing challenges in chaos-based image encryption (IE), particularly inadequate chaotic behavior, vulnerable key structures, and suboptimal operational efficiency, a lightweight IE scheme is developed by leveraging the advantages of 1D-DCVCS. The cryptographic system achieves enhanced security while maintaining computational efficiency, with quantitative security analysis demonstrating superior performance. This work provides a comprehensive solution that simultaneously addresses theoretical limitations in chaotic system design and practical requirements in secure communication applications. Xiangguang Sun, Jun Zheng 0017, Yulai Xie 0002, Peisong He |
IEEE Trans. Ind. Informatics | 3 |
| 2024 | Physical Backdoor: Towards Temperature-Based Backdoor Attacks in the Physical WorldabstractBackdoor attacks have been well-studied in visible light object detection (VLOD) in recent years. However, VLOD can not effectively work in dark and temperature-sensitive scenarios. Instead, thermal infrared object detection (TIOD) is the most accessible and practical in such environments. In this paper, our team is the first to investigate the security vulnerabilities associated with TIOD in the context of backdoor attacks, spanning both the digital and physical realms. We introduce two novel types of backdoor attacks on TIOD, each offering unique capabilities: Object-affecting Attack and Range-affecting Attack. We conduct a comprehensive analysis of key factors influencing trigger design, which include temperature, size, material, and concealment. These factors, especially temperature, significantly impact the efficacy of backdoor attacks on TIOD. A thorough understanding of these factors will serve as a foundation for designing physical triggers and temperature controlling experiments. Our study includes extensive experiments conducted in both digital and physical environments. In the digital realm, we evaluate our approach using benchmark datasets for TIOD, achieving an Attack Success Rate (ASR) of up to 98.21%. In the physical realm, we test our approach in two real-world settings: a traffic intersection and a parking lot, using a thermal infrared camera. Here, we attain an ASR of up to 98.38%. Wen Yin 0001, Jian Lou 0001, Pan Zhou 0001, Yulai Xie 0002, Dan Feng 0001, Tailai Zhang, Lichao Sun 0001 |
CVPR | 4 |
| 2024 | Backdoor Attacks on Bimodal Salient Object Detection with RGB-Thermal DataabstractRGB-Thermal Salient Object Detection (RGBT-SOD) plays a critical role in complex scene recognition applications, such as autonomous driving. However, security research in this domain is still in its infancy. This paper presents the first backdoor attack on RGBT-SOD systems, generating saliency maps on triggered inputs that depict non-existent salient objects chosen by the attacker or falsely mark an entire image as fully salient or entirely non-salient. We uncover that triggers have an influence range for generating non-existent salient objects, supported by a theoretical analysis. Extensive experiments show the effectiveness of our attack in both digital and physical-world scenarios. Notably, our dual-modality backdoor attack achieves an Attack Success Rate (ASR) of 86.72% with only five pairs of poisoned images in model training. After investigating potential countermeasures, we find them inadequate in mitigating our attacks, highlighting the urgent need for robust defenses against sophisticated backdoor attacks in RGBT-SOD systems. Wen Yin 0001, Bin B. Zhu, Yulai Xie 0002, Pan Zhou 0001, Dan Feng 0001 |
ACM Multimedia | 3 |
| 2024 | Accurate Generation of I/O Workloads Using Generative Adversarial NetworksabstractIt is essential to utilize a large number of I/O workloads to analyze commodity system performance or simulate scientific phenomena in high-performance scientific computing. I/O traces are often unavailable at scale due to trace storage overhead, privacy concerns, and the performance impact of trace instrumentation. We study how to generate sufficiently representative I/O workloads using Generative Adversarial Networks (GANs). The best GAN architecture can generate I/O workloads with maximum mean discrepancy (MMD) as low as 0.015-0.05, which implies the synthetic I/O workloads have successfully learned the potential distribution of real I/O traces. We demonstrate that the performance similarity between the original I/O trace and the generated I/O workload through trace replay can be 90.36%-97.32%. Heyu Zhang, Yulai Xie 0002, Yafeng Wu, Dan Feng 0001, Avani Wildani, Darrell D. E. Long |
NAS | 3 |
| 2024 | SatShield: In-Network Mitigation of Link Flooding Attacks for LEO Constellation NetworksabstractLow Earth Orbit (LEO) satellite networks provide global connectivity but are vulnerable to security threats such as link flooding attacks. To defend against such attacks, stateof-the-art approaches employ SDN to acquire a global view of the network, enabling the detection and mitigation of malicious traffic. However, in LEO constellation networks, the distributed nature of satellites across a large spatial scale introduces significant latency in both satellite-to-ground and inter-satellite links, with latency reaching up to tens of milliseconds, while attack traffic dynamically adapts within sub-milliseconds. As a result, existing defense systems face challenges in countering these attacks effectively due to the increased reaction time caused by link latency. In this paper, we leverage programmable switches to build a real-time defense system against link flooding attacks (LFA) in LEO constellation networks. To achieve this, we analyze the practical constraints encountered in the deployment of LFA attacks against state-of-the-art LEO satellite systems. We observe that despite the ability of bots to initiate attack traffic from any location worldwide, an anomalous distribution of flow rate on the affected links can still be detected. We propose SatShield, an in-network defense system that filters out suspicious traffic (heavy flows) in the network and mitigates these threats by leveraging programmable packet scheduling. By using SatShield, we are able to achieve real-time identification and rate-limiting of attacks at line rate on a per-packet basis. We implement SatShield with P4 in a commercial programmable switch and evaluate it with real-world traffic traces. Our evaluation shows that SatShield autonomously identifies LFA attack flows and rapidly mitigates LFA attacks. Hao Jiang 0010, Yulai Xie 0002, Jing Wu 0016, Xiaofan He, Hao Li 0080, Pan Zhou 0001 |
IEEE Internet Things J. | 3 |
| 2023 | Internet Public Safety Event Grading and Hybrid Storage Based on Multi-feature Fusion for Social Media Texts
Yulai Xie 0002, Dan Feng 0001, Shixun Zhao, Pengyu Fu |
DASFAA (1) | 2 |
| 2023 | 3DHacker: Spectrum-based Decision Boundary Generation for Hard-label 3D Point Cloud AttackabstractWith the maturity of depth sensors, the vulnerability of 3D point cloud models has received increasing attention in various applications such as autonomous driving and robot navigation. Previous 3D adversarial attackers either follow the white-box setting to iteratively update the coordinate perturbations based on gradients, or utilize the output model logits to estimate noisy gradients in the black-box setting. However, these attack methods are hard to be deployed in real-world scenarios since realistic 3D applications will not share any model details to users. Therefore, we explore a more challenging yet practical 3D attack setting, i.e., attacking point clouds with black-box hard labels, in which the attacker can only have access to the prediction label of the input. To tackle this setting, we propose a novel 3D attack method, termed 3D Hard-label attacker (3DHacker), based on the developed decision boundary algorithm to generate adversarial samples solely with the knowledge of class labels. Specifically, to construct the class-aware model decision boundary, 3DHacker first randomly fuses two point clouds of different classes in the spectral domain to craft their intermediate sample with high imperceptibility, then projects it onto the decision boundary via binary search. To restrict the final perturbation size, 3DHacker further introduces an iterative optimization strategy to move the intermediate sample along the decision boundary for generating adversarial point clouds with smallest trivial perturbations. Extensive evaluations show that, even in the challenging hard-label setting, 3DHacker still competitively outperforms existing 3D attacks regarding the attack performance as well as adversary quality. Yunbo Tao, Daizong Liu, Pan Zhou 0001, Yulai Xie 0002, Wei Du 0009, Wei Hu 0003 |
ICCV | 4 |
| 2023 | Efficient Storage Management for Social Network Events Based on Clustering and Hot/Cold Data ClassificationabstractSocial network events are related to the national economy and people’s livelihood, so timely perception and processing of massive social network events data are becoming increasingly important for public opinion analysis. How to fully exploit the access feature of event information to store and manage social network events is of great significance for accurate and real-time query analysis. We propose a social network event storage management method based on microblog text clustering and hot/cold data classification. First, for the microblog text data, we construct a keyword provenance graph by using the information entropy to measure the weight of the edge between keyword nodes. Then, we cluster the events using the provenance-based community partition (PCP) with local modularity to improve the event clustering accuracy. In addition, we can further filter noisy data via incremental clustering, enable hot/cold event data classification and dynamic migration, and compress cold data to save space on a hybrid storage architecture. The experimental results show that the clustering purity can reach more than 93% and the query time can be reduced by more than 70% using clustering and hybrid storage policy. Yulai Xie 0002, Shuai Tong, Pan Zhou 0001, Yuli Li, Dan Feng 0001 |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2023 | Paradise: Real-Time, Generalized, and Distributed Provenance-Based Intrusion DetectionabstractIdentifying intrusion from massive and multi-source logs accurately and in real-time presents challenges for today's users. This article presents Paradise, a real-time, generalized, and distributed provenance-based intrusion detection method. Paradise introduces a novel extract strategy to prune and extract process feature vectors from provenance dependencies at the system log level, and it stores them in high-efficiency memory databases. Using this strategy, Paradise does not depend on the specific operating system type or provenance collection framework. Provenance-based dependencies are calculated independently during the detection phase, thus, Paradise can negotiate all detection results from multiple detectors without extra communication overhead between detectors. Paradise also employs an efficient load-balanced distribution scheme that enhances the Kafka architecture to efficiently distribute provenance graph feature vectors to the detectors. The experimental results demonstrate that our method has a high detection accuracy with a low time overhead. Yafeng Wu, Yulai Xie 0002, Xuelong Liao, Pan Zhou 0001, Dan Feng 0001, Avani Wildani, Darrell D. E. Long |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2023 | A Novel Hybrid Model for Docker Container Workload PredictionabstractThe emergence of containers dramatically simplifies and facilitates the development and deployment of applications. More and more enterprises deploy their applications on the container cloud platform. For cloud service providers, an effective container workload prediction method is a must to achieve efficient utilization of cloud resources. However, the existing methods are either rarely based on container load characteristics or cannot make accurate real-time predictions. In this paper, we propose a Docker container workload proactive prediction method using a hybrid model combining triple exponential smoothing and long short-term memory (LSTM), which not only can capture both short-term and long-term dependencies in container resource time series but also smooth the container resource utilization data. In order to improve the prediction accuracy of the hybrid model, those two single models are combined using the mean absolute percentage error (MAPE) method. Besides, we design a real-time Docker workload prediction system for the hybrid model. Our experiments show that the mean absolute percentage error of the hybrid model is decreased by an average of 3.24%, 12.18%, 13.42%, 43.45%, and 50.69% compared with the LSTM, the triple exponential smoothing, ES-ARIMA, Bayesian Ridge Regression and BiLSTM with an acceptable time and computational cost overhead. Liangkang Zhang, Yulai Xie 0002, Minpeng Jin, Pan Zhou 0001, Gongming Xu, Yafeng Wu, Dan Feng 0001, Darrell D. E. Long |
IEEE Trans. Netw. Serv. Manag. | 2 |
| 2022 | EGC: A novel event-oriented graph clustering framework for social media text
Dan Feng 0001, Yulai Xie 0002 |
Inf. Process. Manag. | 3 |
| 2022 | Real-Time Prediction of Docker Container Resource Load Based on a Hybrid Model of ARIMA and Triple Exponential SmoothingabstractMore and more enterprises are beginning to use Docker containers to build cloud platforms. Predicting the resource usage of container workload has been an important and challenging problem to improve the performance of cloud computing platform. The existing prediction models either incur large time overhead or have insufficient accuracy. This article proposes a hybrid model of the ARIMA and triple exponential smoothing. It can accurately predict both linear and nonlinear relationships in the container resource load sequence. To deal with the dynamic Docker container resource load, the weighting values of the two single models in the hybrid model are chosen according to the sum of squares of their predicted errors for a period of time. We also design and implement a real-time prediction system that consists of the collection, storage, prediction of Docker container resource load data and scheduling optimization of CPU and memory resource usage based on predicted values. The experimental results show that the predicting accuracy of the hybrid model improves by 52.64, 20.15, and 203.72 percent on average compared to the ARIMA, the triple exponential smoothing model and ANN+SaDE model respectively with a small time overhead. Yulai Xie 0002, Minpeng Jin, Zhuping Zou, Gongming Xu, Dan Feng 0001, Wenmao Liu, Darrell D. E. Long |
IEEE Trans. Cloud Comput. | 1 |
| 2022 | A Docker Container Anomaly Monitoring System Based on Optimized Isolation ForestabstractContainer-based virtualization has gradually become a main solution in today‘s cloud computing environments. Detecting and analyzing anomaly in containers present a major challenge for cloud vendors and users. This paper proposes an online container anomaly detection system by monitoring and analyzing multidimensional resource metrics of the containers based on the optimized isolation forest algorithm. To improve the detection accuracy, it assigns each resource metric a weight and changes the random feature selection in the isolation forest algorithm to the weighted feature selection according to the resource bias of the container. In addition, it can identify abnormal resource metrics and automatically adjust the monitoring period to reduce the monitoring delay and system overhead. Moreover, it can locate the cause of the anomalies via analyzing and exploring the container log. The experimental results demonstrate the performance and efficiency of the system on detecting the typical anomalies in containers in both simulated and real cloud environments. Zhuping Zou, Yulai Xie 0002, Gongming Xu, Dan Feng 0001, Darrell D. E. Long |
IEEE Trans. Cloud Comput. | 2 |
| 2021 | Context-Aware Biaffine Localizing Network for Temporal Sentence GroundingabstractThis paper addresses the problem of temporal sentence grounding (TSG), which aims to identify the temporal boundary of a specific segment from an untrimmed video by a sentence query. Previous works either compare pre-defined candidate segments with the query and select the best one by ranking, or directly regress the boundary timestamps of the target segment. In this paper, we propose a novel localization framework that scores all pairs of start and end indices within the video simultaneously with a biaffine mechanism. In particular, we present a Context-aware Biaffine Localizing Network (CBLN) which incorporates both local and global contexts into features of each start/end position for biaffine-based localization. The local contexts from the adjacent frames help distinguish the visually similar appearance, and the global contexts from the entire video contribute to reasoning the temporal relation. Besides, we also develop a multi-modal self-attention module to provide fine-grained query-guided video representation for this biaffine strategy. Extensive experiments show that our CBLN significantly outperforms state-of-thearts on three public datasets (ActivityNet Captions, TACoS, and Charades-STA), demonstrating the effectiveness of the proposed localization framework. The code is available at https://github.com/liudaizong/CBLN. Daizong Liu, Xiaoye Qu, Jianfeng Dong, Pan Zhou 0001, Yu Cheng 0001, Wei Wei 0002, Zichuan Xu, Yulai Xie 0002 |
CVPR | 8 |
| 2021 | A Resource-Constrained and Privacy-Preserving Edge-Computing-Enabled Clinical Decision System: A Federated Reinforcement Learning ApproachabstractInternet-of-Things-enabled E-health system, which could monitor and collect the personal health information (PHI), has gradually transformed the clinical treatment to a more personalized way with in-home monitoring smart devices. Then, with the collected PHI, clinical decision support systems (CDSSs), which are based on data mining techniques and historical electronic medical records (EMRs) to help clinicians make proper treatment decisions, have attracted considerable attention. To address issues, such as network congestion and low rate of responsiveness for traditional methods when implementing CDSSs, we integrate the technologies mobile-edge computing (MEC) and software-defined networking for exploiting the computation resources and storage capacities among edge nodes (ENs) (i.e., MEC servers) in our model. Based on this integrated system, each edge node will deploy a double deep Q-network (DDQN) to obtain a stable and sequential clinical treatment policy. It is enabled by a novel fully decentralized federated framework (FDFF) for aggregating models of DDQN and extracting the knowledge from EMRs across all ENs. Furthermore, we discuss the convergence of FDFF in resource-constrained environments. However, since most EMRs are faced with stringent privacy concerns, we adopt two additively homomorphic encryption schemes to prevent leakage of EMRs' privacy during the training process of FDFF. Finally, we measure the time cost of our additively homomorphic encryption schemes and validate DDQN with experiments on large data sets based on FDFF, which shows promising performance on clinician treatment. Zeyue Xue, Pan Zhou 0001, Zichuan Xu, Xiumin Wang 0005, Yulai Xie 0002, Xiaofeng Ding 0001, Shiping Wen 0001 |
IEEE Internet Things J. | 5 |
| 2021 | P-Gaussian: Provenance-Based Gaussian Distribution for Detecting Intrusion Behavior Variants Using High Efficient and Real Time Memory DatabasesabstractIt is increasingly important and a big challenge to detect intrusion behavior variants in today's world. Previous host-based intrusion detection methods typically explore the sequence of system calls or unix shell commands to detect the intrusion behavior. This article abstracts the detection of intrusion behavior variants as the comparison between different sequences when the sequence order or length transforms. To overcome the impact of sequence transformation on the detection accuracy, we propose P-Gaussian, a provenance-based Gaussian distribution detection scheme which comprises two key design features: (1) it utilizes provenance to describe and identify intrusion behavior variants, and eliminates the impact of sequence order transformation on the detection accuracy. (2) it adopts Gaussian distribution principle to accurately compute the similarity between intrusion behavior and its variant, and eliminates the impact of intrusion behavior sequence length increase on the detection accuracy. To improve the detection performance, P-Gaussian employs a Redis memory database with multiple Redis instances and multiple threads to enable the parallelism of provenance processing in multi-core environments. It also classifies hot and cold provenance to provide high-efficient long-term forensic analysis. Experimental results on widely-used real world applications demonstrate the performance and efficiency of our system. Yulai Xie 0002, Yafeng Wu, Dan Feng 0001, Darrell D. E. Long |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2020 | Efficient Provenance Management via Clustering and Hybrid Storage in Big Data EnvironmentsabstractProvenance is a type of metadata that records the creation and transformation of data objects. It has been applied to a wide variety of areas such as security, search, and experimental documentation. However, provenance usually has a vast amount of data with its rapid growth rate which hinders the effective extraction and application of provenance. This paper proposes an efficient provenance management system via clustering and hybrid storage. Specifically, we propose a Provenance-Based Label Propagation Algorithm which is able to regularize and cluster a large number of irregular provenance. Then, we use separate physical storage mediums, such as SSD and HDD, to store hot and cold data separately, and implement a hot/cold scheduling scheme which can update and schedule data between them automatically. Besides, we implement a feedback mechanism which can locate and compress the rarely used cold data according to the query request. The experimental test shows that the system can significantly improve provenance query performance with a small run-time overhead. Dan Feng 0001, Yulai Xie 0002, Gongming Xu, Xinrui Gu, Darrell D. E. Long |
IEEE Trans. Big Data | 3 |
| 2020 | Pagoda: A Hybrid Approach to Enable Efficient Real-Time Provenance Based Intrusion Detection in Big Data EnvironmentsabstractEfficient intrusion detection and analysis of the security landscape in big data environments present challenge for today's users. Intrusion behavior can be described by provenance graphs that record the dependency relationships between intrusion processes and the infected files. Existing intrusion detection methods typically analyze and identify the anomaly either in a single provenance path or the whole provenance graph, neither of which can achieve the benefit on both detection accuracy and detection time. We propose Pagoda, a hybrid approach that takes into account the anomaly degree of both a single provenance path and the whole provenance graph. It can identify intrusion quickly if a serious compromise has been found on one path, and can further improve the detection rate by considering the behavior representation in the whole provenance graph. Pagoda uses a persistent memory database to store provenance and aggregates multiple similar items into one provenance record to maximumly reduce unnecessary I/O during the detection analysis. In addition, it encodes duplicate items in the rule database and filters noise that does not contain intrusion information. The experimental results on a wide variety of real-world applications demonstrate its performance and efficiency. Yulai Xie 0002, Dan Feng 0001, Yuchong Hu, Yan Li 0006, Staunton Sample, Darrell D. E. Long |
IEEE Trans. Dependable Secur. Comput. | 1 |
| 2016 | Oasis: An active storage framework for object storage platform
Yulai Xie 0002, Dan Feng 0001, Yan Li 0006, Darrell D. E. Long |
Future Gener. Comput. Syst. | 1 |
| 2016 | Unifying intrusion detection and forensic analysis via provenance awareness
Yulai Xie 0002, Dan Feng 0001, Junzhe Zhou |
Future Gener. Comput. Syst. | 1 |
| 2013 | Evaluation of a Hybrid Approach for Efficient Provenance StorageabstractProvenance is the metadata that describes the history of objects. Provenance provides new functionality in a variety of areas, including experimental documentation, debugging, search, and security. As a result, a number of groups have built systems to capture provenance. Most of these systems focus on provenance collection, a few systems focus on building applications that use the provenance, but all of these systems ignore an important aspect: efficient long-term storage of provenance. In this article, we first analyze the provenance collected from multiple workloads and characterize the properties of provenance with respect to long-term storage. We then propose a hybrid scheme that takes advantage of the graph structure of provenance data and the inherent duplication in provenance data. Our evaluation indicates that our hybrid scheme, a combination of Web graph compression (adapted for provenance) and dictionary encoding, provides the best trade-off in terms of compression ratio, compression time, and query performance when compared to other compression schemes. Yulai Xie 0002, Kiran-Kumar Muniswamy-Reddy, Dan Feng 0001, Yan Li 0006, Darrell D. E. Long |
ACM Trans. Storage | 1 |
| 2012 | A hybrid approach for efficient provenance storageabstractEfficient provenance storage is an essential step towards the adoption of provenance. In this paper, we analyze the provenance collected from multiple workloads with a view towards efficient storage. Based on our analysis, we characterize the properties of provenance with respect to long term storage. We then propose a hybrid scheme that takes advantage of the graph structure of provenance data and the inherent duplication in provenance data. Our evaluation indicates that our hybrid scheme, a combination of web graph compression (adapted for provenance) and dictionary encoding, provides the best tradeoff in terms of compression ratio, compression time and query performance when compared to other compression schemes. Yulai Xie 0002, Dan Feng 0001, Kiran-Kumar Muniswamy-Reddy, Yan Li 0006, Darrell D. E. Long |
CIKM | 1 |
| 2012 | An In-Depth Analysis of TCP and RDMA Performance on Modern Server PlatformabstractTraditional analysis on the TCP performance attributes the overhead of TCP mainly on these aspects: many times of memory access and copy, much data transmission on bus and a large number of CPU cycles required by the protocol processing. Compared to TCP, the main advantage of RDMA lies in its single memory access which brings a higher data transmission bandwidth and lower CPU utilization. However, the hardware platforms that these analyses are based on are already out-of-date from today's view. The key point is that TCP and RDMA performance can change with the rapid development of the computer architecture and computer hardware components. In this paper, we re-analysis the performance of TCP and RDMA on the modern server platform (e.g., multi-core), and we find that the upper limit of TCP transmission bandwidth has been greatly upgraded compared to the old platform. Besides, the compatibility and complexity problems in programming brought by RDMA technology are still serious. Based on this conclusion, we constructed a network platform which uses multi-thread technology and TCP protocol to achieve a high speed transmission bandwidth in the InfiniBand network environment. Dan Feng 0001, Fang Wang 0001, Liang Ming, Yulai Xie 0002 |
NAS | 5 |
| 2011 | Design and evaluation of Oasis: An active storage framework based on T10 OSD standardabstractIn this paper, we present the design and performance evaluation of Oasis, an active storage framework for object-based storage systems that complies with the current T10 OSD standard. In contrast with previous work, Oasis has the following advantages. First, Oasis enables users to transparently process the OSD object and supports different processing granularity (from the single object to all the objects in the OSD) by extending the OSD object attribute page defined in the T10 OSD standard. Second, Oasis provides an easy and efficient way for users to manage the application functions in the OSD by using the existing OSD commands. Third, Oasis can authorize the execution of the application function in the OSD by enhancing the T10 OSD security protocol, allowing only authorized users to use the system. We evaluate the performance and scalability of our system implementation on Oasis by running three typical applications. The results indicate that active storage far outperforms the traditional object-based storage system in applications that filter data on the OSD. We also experiment with Java based applications and C based applications. Our experiments indicate that Java based applications may be bottlenecked for I/O-intensive applications, while for applications that do not heavily rely on the I/O operations, both Java based applications and C based applications achieve comparable performance. Our microbenchmarks indicate that Oasis implementation overhead is minimal compared to the Intel OSD reference implementation, between 1.2% to 5.9% for Read commands and 0.6% to 9.9% for Write commands. Yulai Xie 0002, Kiran-Kumar Muniswamy-Reddy, Dan Feng 0001, Darrell D. E. Long, Yangwook Kang, Zhongying Niu |
MSST | 1 |