VLDB 2026 Research / reviewers in the wild / expert
Da-Fang Zhang 0001
dblp:16/2420-1 · also Dafang Zhang 0001
· DBLP profile ↗
108ranked-venue papers
3as first author
36since 2021 · last 2026
0000-0003-0765-6857ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 40 · 13 since 2021Systems, architecture and hardware · 24 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 2 first-author · 7 since 2021Security and privacy · 11 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Artificial intelligence and machine learning · 6 · 3 since 2021Software engineering, systems software and programming languages · 4 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Alzheimer's disease risk prediction via perceptual deformable attention generative adversarial network with large foundation models
Zhao-Xu Xing, Zhengliang Liu, Da-Fang Zhang 0001, Kun Xie 0001, Jinxiong Fang, Xia-an Bi, Tianming Liu 0001 |
Medical Image Anal. | 3 |
| 2026 | DALNN: Dual Attention Learning Neural Network for Diagnosis of Autism Spectrum Disorders and Exploration of Lesion Brain RegionsabstractAutism spectrum disorders (ASDs) typically occur in early childhood and affect brain development. Identifying the diseased regions of interest (ROIs) is crucial for improving early diagnosis. However, most current deep learning algorithms for ASD diagnosis struggle to locate lesion ROIs and fully explore relationships between them, which limits their clinical applicability. To resolve these, a dual attention learning fusion algorithm (DALFA) is proposed to enhance feature extraction and fusion from multimodal data. Specifically, the algorithm first constructs features based on both the functional and structural data of the subjects and applies appropriate methods for feature extraction. After concatenating the extracted features, the dual attention learning is employed to fuse these features across two dimensions, optimizing feature representation. Furthermore, this study introduces a dual attention learning neural network (DALNN) that implements and refines the details of DALFA. We conduct sufficient validation experiments using data from the ABIDE dataset. The classification accuracy of DALNN reaches 83.35%, outperforming existing methods. In addition, DALNN can extract ASD-related ROIs, which is beneficial for guiding early clinical intervention. Jinxiong Fang, Da-Fang Zhang 0001, Kun Xie 0001, Luyun Xu, Xia-an Bi |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2026 | Matrix Reshaping for Reduced Sensing Cost and Improved Data Inference in Sparse Mobile Sensing EnvironmentsabstractMobile crowd sensing (MCS) has emerged as a promising sensing paradigm with the widespread adoption of smartphones. However, one of the key bottlenecks in MCS lies in the high sensing cost imposed on mobile users. To alleviate this burden, sparse sensing strategies are often employed, where data is collected from a limited number of locations and the remaining data is inferred by exploiting spatio-temporal correlations. Compared with vector-based inference approaches, matrix completion techniques can better capture two-dimensional correlations in the sensing data, thereby achieving higher recovery accuracy. Nevertheless, their performance degrades significantly when the actual sensing rate is low. In this paper, we propose a novel matrix-reshaping strategy that is applied prior to matrix completion to enhance recovery performance under sparse observations. We provide a theoretical analysis demonstrating that the reshaping process reduces the number of measurements required for successful matrix recovery. To validate our approach, we conduct extensive experiments using traditional matrix completion algorithms, deep learning models, and tensor completion methods on six real-world datasets. The results show that, to achieve the same level of recovery accuracy, our reshaped matrices consistently reduce the measurement overhead compared to their original ones. Jiazheng Tian, Kun Xie 0001, Jigang Wen, Da-Fang Zhang 0001, Guangxing Zhang, Gaogang Xie |
IEEE Trans. Mob. Comput. | 4 |
| 2025 | Digital Civic Engagement in China: Using 'Micro Advice' Platform to Improve People's LivelihoodabstractMicro Advice is a mobile platform for democratic governance in China, allowing access to voice social issues and advice to the government with the aim of improving people's livelihood. However, due to the lack of first-hand experience, the current understanding of how end-users utilize Micro Advice to participate in democratic governance is incomplete. We interviewed 12 users to understand their practices and challenges in using the platform. Specifically, we illustrate the user's experience, introduce what difficulties they encountered, and how they strategically use the platform to improve people's livelihood. We also investigate the socio-technical aspects of Micro Advice within the Chinese political context, discussing how to accept and utilize Micro Advice in China's social environment, and develop technological solutions adapted to these backgrounds. Finally, we propose some design implications for civic technology participation platforms. Micro Advice provides a novel, open, and real-time channel for civic engagement, showcasing the practical effects and impact of digitized civic engagement in China. It offers researchers a new perspective for expressing and addressing societal issues. We believe that the innovation and insights of Micro Advice can extend to other types of digitized civic engagement initiatives. We will continue to explore the interactive processes between the government and the public, along with innovative technological approaches. Yeye Li, Hanhui Deng, Nan Ma 0003, Xin Tong 0004, Mingming Fan 0001, Da-Fang Zhang 0001, Yi Li 0075, Di Wu 0002 |
Proc. ACM Hum. Comput. Interact. | 6 |
| 2025 | EventMon: Real-Time Event-Based Streaming Network Monitoring Data RecoveryabstractData recovery is a fundamental task for sparse network monitoring with a significant impact on many downstream tasks, such as congestion control, network capacity planning, and traffic engineering. To better capture the network dynamically and quickly respond to network failure, network monitoring systems take finer temporal granularity to collect data to form a real-time view of the network. Unfortunately, the current data recovery for network monitoring relies on matrix and tensor completion algorithms, which fail to satisfy the requirements of real-time recovery. To combat this problem, in this paper, we propose Real-TimeEvent-based Streaming NetworkMonitoring Data Recovery (EventMon) that achieves ultra-low latency data recovery in network measurement data streams. Specifically, we leverage a mixture of offline and online architecture, with the offline component learns to capture the historical spatial-temporal correlation while the online component is a novel streaming encoder that updates factor matrices incrementally. To enable training our event-level stream processing module, we devise a novel Stream2Batch algorithm to enable mini-batch style training and ensure the encoder generates the same results with one-by-one stream processing. We conduct extensive experiments on three network monitoring datasets and our evaluation and analysis of the experimental results demonstrate shows that the proposed method outperforms existing schemes in terms of accuracy, inference latency, and high processing throughput. Wei Liang 0005, Kun Xie 0001, Da-Fang Zhang 0001, Kuanching Li, Naixue Xiong |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | TensorMon: A Breakthrough in Sparse Data Gathering Leveraging Tensor-Enhanced Techniques for System and Network MonitoringabstractSparse data gathering has become a promising solution for reducing measurement costs by leveraging the inherent sparsity of data. However, most existing approaches rely on low-dimensional models such as compressive sensing or matrix completion, which are limited in capturing complex high-dimensional structures. To overcome these limitations, we proposeTensorMon, a novel tensor-based sparse data gathering framework that introduces a cuboid sampling strategy to more effectively exploit multidimensional correlations. Unlike traditional entry-based or tube-based sampling, TensorMon introduces the innovative concept ofcuboid sampling. We further develop a lightweight sampling scheduling algorithm and a non-iterative inference algorithm to ensure efficient measurement planning and accurate reconstruction of unmeasured data. Theoretical analysis establishes a new performance bound for our sampling strategy, which is significantly lower than those in existing literature. To validate our theoretical findings, we conduct extensive experiments on four real-world datasets: two network monitoring datasets, a city-scale crowd flow dataset, and a road traffic speed dataset. Experimental results demonstrate that TensorMon achieves substantial reductions in measurement cost, delivers high inference accuracy, and ensures rapid data recovery, highlighting its effectiveness and practicality across diverse application scenarios. Jiazheng Tian, Kun Xie 0001, Xin Wang 0001, Jigang Wen, Gaogang Xie, Wei Liang 0005, Da-Fang Zhang 0001, Kenli Li 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2024 | MaP: Increasing node capacity of programmable cloud gateways
Donghong Jiang, Yanbiao Li 0001, Xin Wang 0001, Da-Fang Zhang 0001, Gaogang Xie |
Comput. Networks | 5 |
| 2024 | Dynamic multi-scale spatial-temporal graph convolutional network for traffic flow prediction
Na Hu, Da-Fang Zhang 0001, Kun Xie 0001, Wei Liang 0005, Kuanching Li, Albert Y. Zomaya |
Future Gener. Comput. Syst. | 2 |
| 2024 | Uncovering Malicious Accounts in Open Mobile Social Networks Using a Graph- and Text-Based Attention Fusion AlgorithmabstractIn recent years, open mobile social networks focused on socializing and dating purposes have gained widespread popularity, such as Soul, Tinder, Momo, and Tantan, among several others. These applications permit users to post, comment, and send private messages to other users without their consent, making communication accessible. However, this low-entry communication approach has also increased malicious user attacks. We delve into a comprehensive analysis of malicious accounts in open socializing and dating applications, revealing that the existing methods overlook hidden malicious signals within the user text-related information, thus resulting in poor detection performance. For such, we propose GraphTAM, a novel graph- and text-based multihead attention fusion network model for detecting such malicious accounts, consisting of modules that effectively combine nontext-related and text-related information, enhancing the accuracy and performance of detecting malicious accounts. We employ graph convolutional networks (GCNs) for nontext-related information to extract advanced representations of users, incorporating their attribute and social relationship features. Regarding text-related information, we employ a multihead attention model to identify suspicious patterns in users’ posted articles, comments, and relevant behavioral statistics, so finally, we merge the advanced representations of nontext-related and text-related information using a multilayer perceptron to determine the maliciousness of an account. Data sets collected from SLink are utilized for the experimental evaluation and to compare the performance of the proposed model with the several state of the art algorithms. Experimental results show significant advantages in malicious account detection, where the F1 score achieves over 0.9, outperforming the existing methods that range between 0.6 and 0.85. Furthermore, the comparative experiments substantiate the critical role of text-related information in detecting malicious accounts in open socializing and dating applications. Yuting Tang, Da-Fang Zhang 0001, Wei Liang 0005, Kuanching Li, Keqin Li 0001 |
IEEE Internet Things J. | 2 |
| 2024 | DSTGCS: an intelligent dynamic spatial-temporal graph convolutional system for traffic flow prediction in ITS
Na Hu, Da-Fang Zhang 0001, Wei Liang 0005, Kuanching Li, Arcangelo Castiglione |
Soft Comput. | 2 |
| 2024 | QoS Prediction and Adversarial Attack Protection for Distributed Services Under DLaaSabstractDeep-Learning-as-a-service (DLaaS) has received increasing attention due to its novelty as a diagram for deploying deep learning techniques. However, DLaaS faces performance and security issues that urgently need to be addressed. Given the limited computation resources and concern of benefits, Quality-of-Service (QoS) metrics should be revised to optimize the performance and reliability of distributed DLaaS systems. New users and services dynamically and continuously join and leave such a system, resulting in cold start issues, and additionally, the increasing demand for robust network connections requires the model to evaluate the uncertainty. To address such performance problems, we propose in this article a deep learning-based model called embedding enhanced probability neural network, in which information is extracted from inside the graph structure and then estimated the mean and variance values for the prediction distribution. The adversarial attack is a severe threat to model security under DLaaS. Due to such, the service recommender system's vulnerability is tackled, and adversarial training with uncertainty-aware loss to protect the model in noisy and adversarial environments is investigated and proposed. Extensive experiments on a large-scale real-world QoS dataset are conducted, and comprehensive analysis verifies the robustness and effectiveness of the proposed model. Wei Liang 0005, Jianlong Xu, Zheng Qin 0001, Da-Fang Zhang 0001, Kuanching Li |
IEEE Trans. Computers | 5 |
| 2024 | Predicting Drug-Target Interactions Via Dual-Stream Graph Neural NetworkabstractDrug target interaction prediction is a crucial stage in drug discovery. However, brute-force search over a compound database is financially infeasible. We have witnessed the increasing measured drug-target interactions records in recent years, and the rich drug/protein-related information allows the usage of graph machine learning. Despite the advances in deep learning-enabled drug-target interaction, there are still open challenges: (1) rich and complex relationship between drugs and proteins can be explored; (2) the intermediate node is not calibrated in the heterogeneous graph. To tackle with above issues, this paper proposed a framework named DSG-DTI. Specifically, DSG-DTI has the heterogeneous graph autoencoder and heterogeneous attention network-based Matrix Completion. Our framework ensures that the known types of nodes (e.g., drug, target, side effects, diseases) are precisely embedded into high-dimensional space with our pretraining skills. Also, the attention-based heterogeneous graph-based matrix completion achieves highly competitive results via effective long-range dependencies extraction. We verify our model on two public benchmarks. The result of two publicly available benchmark application programs show that the proposed scheme effectively predicts drug-target interactions and can generalize to newly registered drugs and targets with slight performance degradation, outperforming the best accuracy compared with other baselines. Wei Liang 0005, Da-Fang Zhang 0001, Kuanching Li |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | A Light-Weight and Robust Tensor Convolutional Autoencoder for Anomaly DetectionabstractRobust PCA is a popular anomaly detection technique and has been widely used in many applications. Although Robust PCA is promising, it is usually designed in a two-order matrix form, which is inferior to the tensor that can capture multilinearity features of data. Moreover, the detection accuracy under Robust PCA further suffers due to its sensitivity to the rank parameter which is hard to set in practice and the limitation of PCA method in capturing the non-linear feature in the data. To address the issues, we propose a Robust Tensor Convolutional Autoencoder (RTCAE) where the autoencoder instead of SVD is exploited to recover the normal data from the corrupted measurement tensor data. However, directly exploiting deep autoencoder may suffer from the problem of high memory consumption and computation overhead due to the large number of parameters used in autoencoder. To make our anomaly detection lightweight, we further design a Light Convolutional Autoencoder (LightCAE) which contains a compressed autoencoder by exploiting tensor factorization to largely compress the parameters while significantly reducing the computation complexity. We conduct extensive experiments on three real data traces to compare the performance of our proposed schemes (RTCAE and lightCAE) with that of seven baseline algorithms. The experiment results demonstrate that our proposed RTCAE achieves the highest anomaly detection accuracy. Moreover, our LightCAE requires over 60 times smaller memory storage than that required in RTCAE while achieving the similar anomaly detection accuracy. Xiaocan Li, Kun Xie 0001, Xin Wang 0001, Gaogang Xie, Kenli Li 0001, Jiannong Cao 0001, Da-Fang Zhang 0001, Jigang Wen |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2024 | DMSTG: Dynamic Multiview Spatio-Temporal Networks for Traffic ForecastingabstractTraffic sensor networks are widely applied in smart cities to monitor traffic in real-time and record huge volumes of traffic data. Exploiting such data to forecast future traffic conditions have the potential to enhance the decision-making capabilities of intelligent transportation systems, which attracts widespread attention from both industries and academia. Among them, network-wide prediction based on graph convolutional neural networks(GCN) has become mainstream. It models the spatial dependencies of sensors in a graph with a pre-defined Laplacian matrix based on the distances among sensors. However, understanding spatio-temporal traffic patterns is quite challenging as there is a huge difference in terms of traffic patterns during different periods or in different regions. In addition, the actual data collected can be polluted due to unavoidable data loss from severe communication conditions or sensor failures. Considering these issues, we propose a novel dynamic multiview spatial-temporal prediction framework which takes into consideration various factors, including local/global, short/long term spatio-temporal dependencies and their dynamic changes. To comprehensively track the dynamic spatio-temporal dependencies among traffic data, we creatively design two different modules to perceive the changes in traffic patterns. We first propose a dynamic Laplacian matrix learning module based on our theoretical derivation to estimate the Laplacian matrix of the graph for GCN timely. We creatively incorporate tensor decomposition into this module, where real-time traffic data are decomposed into a global component that is stable and depends on long-term temporal-spatial traffic relationships and a local component that captures the traffic fluctuations. We also design a self-attention based module to dynamically assign a weight to each part in traffic data. The spatio-temporal features from multiple views are deeply fused by a feature fusion module. The forecasting performance is evaluated with 5 real-time traffic datasets. Experiment results demonstrate that our framework can consistently outperform the state-of-the-art baselines and be more robust under noisy environments. Zulong Diao, Xin Wang 0001, Da-Fang Zhang 0001, Gaogang Xie, Jianguo Chen 0001, Changhua Pei, Xuying Meng, Kun Xie 0001, Guangxing Zhang |
IEEE Trans. Mob. Comput. | 3 |
| 2024 | GraphIoT: Lightweight IoT Device Detection Based on Graph Classifiers and Incremental LearningabstractThe rapid expansion of the Internet of Things (IoT) has led to growing concerns about the security of IoT devices. A crucial aspect of ensuring their security is IoT device identification, which involves pinpointing the specific type of device. Existing solutions, however, either necessitate complex feature engineering or struggle to handle the ever-increasing number of new devices in open IoT environments. To tackle these challenges, this paper introduces GraphIoT, a lightweight IoT device detection method based on graph classifiers. GraphIoT leverages lightweight flow information, such as packet length, direction, and timestamp, to create an IoT Device Traffic Graph Representation (IoT-DTGR). This representation offers a comprehensive view of IoT device flows while preserving features in bidirectional IoT Device-Gateway interactions. By transforming the IoT device detection problem into a graph classification problem, GraphIoT employs a powerful Graph Neural Network that takes into account both node and edge features, as well as subgraph structures in IoT-DTGRs, to classify graphs and consequently identify device types. Additionally, the paper proposes an incremental learning framework called CL-GraphIoT that continuously learns features of new IoT device flows without forgetting previously learned device features. This is achieved through two strategies: parameter sharing and sample replaying. The paper gathers a real-world dataset from 18 IoT devices and conducts experiments on two datasets: the gathered real-world dataset and an open-source dataset covering 21 IoT device types. The experimental results demonstrate that both GraphIoT and CL-GraphIoT outperform state-of-the-art methods, achieving high accuracy in device detection with fast processing speed. Yansong Yin, Kun Xie 0001, Shiming He, Yanbiao Li 0001, Jigang Wen, Zulong Diao, Da-Fang Zhang 0001, Gaogang Xie |
IEEE Trans. Serv. Comput. | 7 |
| 2024 | FedGCN: A Federated Graph Convolutional Network for Privacy-Preserving Traffic PredictionabstractTraffic prediction is crucial for intelligent transportation systems, assisting in making travel decisions, minimizing traffic congestion, and improving traffic operation efficiency. Although effective, existing centralized traffic prediction methods have privacy leakage risks. Federated learning-based traffic prediction methods keep raw data local and train the global model in a distributed way, thus preserving data privacy. Nevertheless, the spatial correlations between local clients will be broken as data exchange between local clients is not allowed in federated learning, leading to missing spatial information and inferior prediction accuracy. To this end, we propose a federated graph neural network with spatial information completion (FedGCN) for privacy-preserving traffic prediction by adopting a federated learning scheme to protect confidentiality and presenting a mending graph convolutional neural network to mend the missing spatial information during capturing spatial dependency to improve prediction accuracy. To complete the missing spatial information efficiently and capture the client-specific spatial pattern, we design a personalized training scheme for the mending graph neural network, reducing communication overhead. The experiments on four public traffic datasets demonstrate that the proposed model outperforms the best baseline with a ratio of 3.82%, 1.82%, 2.13%, and 1.49% in terms of absolute mean error while preserving privacy. Na Hu, Wei Liang 0005, Da-Fang Zhang 0001, Kun Xie 0001, Kuanching Li, Albert Y. Zomaya |
IEEE Trans. Sustain. Comput. | 3 |
| 2023 | DeepEAG: A deep learning-based hybrid framework for identifying epilepsy-associated genes using a stacking strategyabstractAs one of the most common neurological diseases in the world, epilepsy can be caused through the deletion or duplication of known epilepsy-associated genes. Therefore, it is crucial to accurately identify epilepsy-associated genes to gain a deeper understanding of the pathogenesis of epilepsy and to develop new therapies. Although a number of wet-lab experimental methods have been proposed, they are reliable but in general time-consuming and labor-intensive. Hence, their practical application in epilepsy-associated genes identification is quite limited. To address the existing limitations, we propose a computational method, called DeepEAG, for the identification of epilepsy-associated genes throughout the human genome. To develop DeepEAG, we first constructed a new and, to our knowledge, the most comprehensive epilepsy genomic benchmark dataset. Afterwards, we investigated four feature encoding algorithms with different perspectives and trained them using well-established traditional classifiers and deep learning classifier, resulting in 13 baseline models. Finally, the stacking strategy is effectively utilized by integrating the predicted output of the optimal baseline models and training with xgboost. In results, the DeepEAG predictor obtained excellent performance on the cross-validation with accuracy and AUC of 0.835 and 0.907, respectively. Overall, DeepEAG showed more stable and accurate predictive performance compared to baseline models, and also outperforms other ensemble strategies, demonstrating the effectiveness of our proposed hybrid framework. In addition, DeepEAG is anticipated to facilitate community-wide efforts to identify putative epilepsy-associated genes and generate new biological insights and testable hypotheses. DeepEAG is freely avaliable at https://github.com/JfXie/DeepEAG. Da-Fang Zhang 0001 |
BIBM | 2 |
| 2023 | Exploring Brain Connectivity with Spatial-Temporal Graph Neural Networks for Improved EEG Seizure AnalysisabstractAutomated seizure detection and prediction from electroencephalography (EEG) can greatly improve seizure diagnosis and treatment. Although previous attempts to analyze seizure have achieved high classification performance, several modeling challenges remain open: (1) How to effectively represent the non-Euclidean data structure of the brain and utilize the spatial and temporal features of multichannel EEG signals remains challenging. Previous work has failed to fully utilize the spatial topological information between brain regions. (2) Most methods have focused solely on seizure signal detection, and how to predict the next seizure and realize early warning is more clinically valuable for timely treatment of epileptic patients. To address the above problems, we propose a new spatial-temporal graph neural network for seizure detection and prediction. Specifically, we construct two brain graphs based on static physical distance and dynamic functional connections. By employing temporal and spatial attention mechanisms, we extract spatiotemporal information from EEG signals and further leverage graph convolutional networks to capture inter-channel relationships. By transforming the seizure prediction task into a multi-class classification, our method automatically identifies pre-ictal signals, enabling the prediction of seizure occurrences. Experimental results demonstrate that our approach achieves outstanding performance on a publicly available seizure EEG dataset. Da-Fang Zhang 0001 |
BIBM | 2 |
| 2023 | LightNestle: Quick and Accurate Neural Sequential Tensor Completion via Meta LearningabstractNetwork operation and maintenance rely heavily on network traffic monitoring. Due to the measurement overhead reduction, lack of measurement infrastructure, and unexpected transmission error, network traffic monitoring systems suffer from incomplete observed data and high data sparsity problems. Recent studies model missing data recovery as a tensor completion task and show good performance. Although promising, the current tensor completion models adopted in network traffic data recovery lack an effective and efficient retraining scheme to adapt to newly arrived data while retaining historical information. To solve the problem, we propose LightNestle, a novel sequential tensor completion scheme based on meta-learning, which designs (1) an expressive neural network to transfer spatial knowledge from previous embeddings to current embeddings; (2) an attention-based module to transfer temporal patterns into current embeddings in linear complexity; and (3) meta-learning-based algorithms to iteratively recover missing data and update transfer modules to catch up with learned knowledge. We conduct extensive experiments on two real-world network traffic datasets to assess our performance. Results show that our proposed methods achieve both fast retraining and high recovery accuracy. Wei Liang 0005, Kun Xie 0001, Da-Fang Zhang 0001, Songyou Xie, Kuanching Li |
INFOCOM | 4 |
| 2023 | Multi-graph fusion based graph convolutional networks for traffic prediction
Na Hu, Da-Fang Zhang 0001, Kun Xie 0001, Wei Liang 0005, Kuanching Li, Albert Y. Zomaya |
Comput. Commun. | 2 |
| 2023 | A Novel Spatial-Temporal Multi-Scale Alignment Graph Neural Network Security Model for Vehicles PredictionabstractTraffic flow forecasting is indispensable in today’s society and regarded as a key problem for Intelligent Transportation Systems (ITS), as emergency delays in vehicles can cause serious traffic security accidents. However, the complex dynamic spatial-temporal dependency and correlation between different locations on the road make it a challenging task for security in transportation. To date, most existing forecasting frames make use of graph convolution to model the dynamic spatial-temporal correlation of vehicle transportation data, ignoring semantic similarity between nodes and thus, resulting in accuracy degradation. In addition, traffic data does not strictly follow periodicity and hard to be captured. To solve the aforementioned challenging issues, we propose in this article CRFAST-GCN, a multi-branch spatial-temporal attention graph convolution network. First, we capture the multi-scale (e.g., hour, day, and week) long- short-term dependencies through three identical branches, then introduce conditional random field (CRF) enhanced graph convolution network to capture the semantic similarity globally, so then we exploit the attention mechanism to captures the periodicity. For model evaluation using two real-world datasets, performance analysis shows that the proposed CRFAST-GCN successfully handles the complex spatial-temporal dynamics effectively and achieves improvement over the baselines at 50% (maximum), outperforming other advanced existing methods. Chunyan Diao, Da-Fang Zhang 0001, Wei Liang 0005, Kuanching Li, Yujie Hong, Jean-Luc Gaudiot |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2023 | Spatial-Temporal Aware Inductive Graph Neural Network for C-ITS Data RecoveryabstractWith the prevalence of Intelligent Transportation Systems (ITS), massive sensors are deployed on roadside, vehicles, and infrastructures. One key challenge is imputing several different types of missing entries in spatial-temporal traffic data to meet the high-quality demand of data science applied in Cooperative-ITS (C-ITS) since accurate data recovery is critical to many downstream tasks in ITSs, such as traffic monitoring and decision making. For such, it is proposed in this article solutions to three kinds of data recovery tasks in a unified model via spatial-temporal aware Graph Neural Networks (GNNs), named Spatial-Temporal Aware Data Recovery Network (STAR), enabling a real-time and inductive inference. A residual gated temporal convolution network is designed to permit the proposed model to learn the temporal pattern from long sequences with masks and an adaptive memory-based attention model for utilizing implicit spatial correlation. To further exploit the generalization power of GNNs, a sampling-based method is adopted to train the proposed model to be robust and inductive for online servicing. Extensive numerical experiments on two real-world spatial-temporal traffic datasets are performed, and results show that the proposed STAR model consistently outperforms other baselines at 1.5-2.5 times on all kinds of imputation tasks. Moreover, STAR can support recovery data for 2 to 5 hours, with its performance barely unchanged, and has comparable performance in transfer learning and time-series forecast. Experimental results demonstrate that STAR provides adequate performance and rich features for multiple data recovery tasks under the C-ITS scenario. Wei Liang 0005, Kun Xie 0001, Da-Fang Zhang 0001, Kuanching Li, Alireza Souri, Keqin Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Neighbor Graph Based Tensor Recovery For Accurate Internet Anomaly DetectionabstractDetecting anomalous traffic is a crucial task for network management. Although many anomaly detection algorithms have been proposed recently, constrained by their matrix-based traffic data model, existing algorithms often suffer from low detection accuracy. To fully utilize the multi-dimensional information hidden in the traffic data, this paper uses the tensor model for more accurate Internet anomaly detection. Only considering the low-rank linearity features hidden in the data, current tensor factorization techniques would result in low anomaly detection accuracy. We propose a novel Graph-based Tensor Recovery model (Graph-TR) to well explore both low-rank linearity features as well as the non-linear proximity information hidden in the traffic data for better anomaly detection. We encode the non-linear proximity information of the traffic data by constructing nearest neighbor graphs and incorporate this information into the tensor factorization using the graph Laplacian. Moreover, to facilitate the quick building of neighbor graph, we propose a nearest neighbor searching algorithm with the simple locality-sensitive hashing (LSH). Besides only detecting random anomalies, our algorithm can also effectively detect structured anomalies that appear as bursts. We have conducted extensive experiments using Internet traffic trace data Abilene and GÈANT. Compared with the state of art algorithms on matrix-based anomaly detection and tensor recovery approach, our Graph-TR can achieve higher Accuracy and Recall. Xiaocan Li, Kun Xie 0001, Xin Wang 0001, Gaogang Xie, Kenli Li 0001, Jiannong Cao 0001, Da-Fang Zhang 0001, Hongbo Jiang 0001, Jigang Wen |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2023 | Tripartite Graph Aided Tensor Completion For Sparse Network MeasurementabstractNetwork measurements provide critical inputs for a wide range of network management. Existing network-wide monitoring methods face the challenge of incurring a high measurement cost. Some recent studies show that network-wide measurement data such as end-to-end latency and flow traffic, have hidden spatio-temporal correlations and thus low-rank features. Taking advantage of the low-rank feature, enlightened by tensor model's strong capability of information representation and extracting, this paper studies a novel sparse measurement scheduling problem which selects a proportion of Origin and Destination (OD) pairs to take measurements in the future time slots, while ensuring the data of the remaining un-measured OD pairs be accurately inferred through tensor completion. It is challenging to find the optimal sampling points (OD pairs) without knowing the structure of the future data and also infer the un-measured data in the presence of noise in the measurement samples. To conquer the challenges, we propose several techniques: a tripartite graph to illustrate the relationship between sample locations and tensor factorization, a graph-based sample selection algorithm, and a graph-based robust tensor completion algorithm. We have conducted extensive experiments based on two real network latency monitoring traces (PlanetLab and Harvard) and two other network monitoring traces (including a traffic trace Abilene and a throughput trace WS-Dream). Our results demonstrate that, even with a sampling ratio of less than 5%, our scheme can accurately obtain the complete network-wide monitoring data by inferring the missing ones based on the samples taken. To achieve similar recovery performance, the best peer tensor completion algorithm needs a significantly larger number of samples, with the sampling ratio up to 25-150 times ours. Xiaocan Li, Kun Xie 0001, Xin Wang 0001, Gaogang Xie, Kenli Li 0001, Jiannong Cao 0001, Da-Fang Zhang 0001, Jigang Wen |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2022 | NMMF-Stream: A Fast and Accurate Stream-Processing Scheme for Network Monitoring Data RecoveryabstractRecovery of missing network monitoring data is of great significance for network operation and maintenance tasks such as anomaly detection and traffic prediction. To exploit historical data for more accurate missing data recovery, some recent studies combine the data together as a tensor to learn more features. However, the need of performing high cost data decomposition compromises their speed and accuracy, which makes them difficult to track dynamic features from streaming monitoring data. To ensure fast and accurate recovery of network monitoring data, this paper proposes NMMF-Stream, a stream-processing scheme with a context extraction module and a generation module. To achieve fast feature extraction and missing data filling with a low sampling rate, we propose several novel techniques, including the context extraction based on both positive and negative monitoring data, context validation via measuring the Pointwise Mutual Information, GRU-based temporal feature learning and memorization, and a new composite loss function to guide the fast and accurate data filling. We have done extensive experiments using two real network traffic monitoring data sets and one network latency data set. The experimental results demonstrate that, compared with three baselines, NMMF-Stream can fill the newly arrived monitoring data very quickly with much higher accuracy. Kun Xie 0001, Ruotian Xie, Xin Wang 0001, Gaogang Xie, Da-Fang Zhang 0001, Jigang Wen |
INFOCOM | 5 |
| 2022 | Order-preserved Tensor Completion For Accurate Network-wide MonitoringabstractNetwork-wide monitoring is important for many network functions. However, monitoring data are often incomplete due to the need of sampling to reduce high measurement cost, system failure, and unavoidable transmission loss under severe communication. Instead of only targeting to estimate all missing monitoring data entries with a small set of measurement samples, we study a new order-preserved monitoring data estimation problem to accurately estimate the missing data entries while preserving the data entries’ order in the dataset. We propose a novel order-preserved tensor completion model that integrates both the low rank property and the order information into a joint learning problem to estimate the missing data. With well designed non-convex function to directly approximate the tensor rank and order-preserved constraint under the linear self-recovery method, our model can not only more accurately capture the low-rank property of monitoring data to increase the estimation performance of missing data, but also can capture the order information in monitoring data to ensure the estimation accuracy. Extensive experiments using four real datasets demonstrate that compared with the state-of-the-art tensor completion algorithms, our proposed algorithm can provide more accurate estimation and keep the value order of recovered entries to more effectively retrieve top-k large entries. Xiaocan Li, Kun Xie 0001, Xin Wang 0001, Gaogang Xie, Kenli Li 0001, Da-Fang Zhang 0001, Jigang Wen |
IWQoS | 6 |
| 2022 | Graph learning-based spatial-temporal graph convolutional neural networks for traffic forecastingabstractTraffic forecasting is highly challenging due to its complex spatial and temporal dependencies in the traffic network. Graph Convolutional Neural Network (GCN) has been effectively used for traffic forecasting due to its excellent performance in modelling spatial dependencies. In most existing approaches, GCN models spatial dependencies in the traffic network with a fixed adjacency matrix. However, the spatial dependencies change over time in the actual situation. In this paper, we propose a graph learning-based spatial-temporal graph convolutional neural network (GLSTGCN) for traffic forecasting. To capture the dynamic spatial dependencies, we design a graph learning module to learn the dynamic spatial relationships in the traffic network. To save training time and computation resources, we adopt dilated causal convolution networks with a gating mechanism to capture long-term temporal correlations in traffic data. We conducted extensive experiments using two real-world traffic datasets. Experimental results demonstrate that the proposed GLSTGCN achieves superior performance than all state-of-art baselines. Na Hu, Da-Fang Zhang 0001, Kun Xie 0001, Wei Liang 0005, Meng-Yen Hsieh |
Connect. Sci. | 2 |
| 2022 | Multi-View Matrix Factorization for Sparse Mobile CrowdsensingabstractMobile crowdsensing (MCS) has become a new paradigm for the environment sensing. However, the sparse sensory data prevent the practical and large-scale deployment of MCS systems. Recent studies have demonstrated that the matrix factorization is an effective technique which can estimate the missing sensory data entries based on a small set of observed data entries. However, there could be multiple sensory data sets with each regarded as a different view on the environment. Applying current matrix factorization individually to each data set, the recovery performance will be low as some data sets do not have enough observed data entries thus enough information. By partitioning the parameters involved in matrix factorization, we design some novel regularizations to encode the similarities among different data sets and specific knowledge in the single data set. Based on the regularizations, we propose one basic multiview matrix factorization (MVMF) model and one neural MVMF (NMVMF) model to combine multiple sensory data sets to mutually reinforce the estimation of each single data set. The extensive experimental results demonstrate that, with the help of other data sets, our models can estimate the missing entries in the data set with a very low sampling ratio accurately while the other five baseline algorithms cannot. Xiaocan Li, Kun Xie 0001, Gaogang Xie, Kenli Li 0001, Jiannong Cao 0001, Da-Fang Zhang 0001, Jigang Wen |
IEEE Internet Things J. | 6 |
| 2022 | Multi-range bidirectional mask graph convolution based GRU networks for traffic prediction
Na Hu, Da-Fang Zhang 0001, Kun Xie 0001, Wei Liang 0005, Chunyan Diao, Kuanching Li |
J. Syst. Archit. | 2 |
| 2022 | A Mutual Security Authentication Method for RFID-PUF Circuit Based on Deep LearningabstractThe Industrial Internet of Things ( IIoT ) is designed to refine and optimize the process controls, thereby leveraging improvements in economic benefits, such as efficiency and productivity. However, the Radio Frequency Identification ( RFID ) technology in an IIoT environment has problems such as low security and high cost. To overcome such issues, a mutual authentication scheme that is suitable for RFID systems, wherein techniques in Deep Learning ( DL ) are incorporated onto the Arbiter Physical Unclonable Function ( APUF ) for the secured access authentication of the IC circuits on the IoT, is proposed. The design applies the APUF-MPUF mutual authentication structure obtained by DL to generate essential real-time authentication information, thereby taking advantage of the feature that the tag in the PUF circuit structure does not need to store any essential information and resolving the problem of key storage. The proposed scheme also uses a bitwise comparison method, which hides the PUF response information and effectively reduces the resource overhead of the system during the verification process, to verify the correctness of the two strings. Security analysis demonstrates that the proposed scheme has high robustness and security against different conventional attack methods, and the storage and communication costs are 95.7% and 42.0% lower than the existing schemes, respectively. Wei Liang 0005, Songyou Xie, Da-Fang Zhang 0001, Xiong Li 0002, Kuanching Li |
ACM Trans. Internet Techn. | 3 |
| 2021 | CRFST-GCN: A Deeplearning Spatial-Temporal Frame to Predict Traffic Flow
Chunyan Diao, Da-Fang Zhang 0001, Wei Liang 0005, Kuanching Li, Man Jiang |
ICA3PP (1) | 2 |
| 2021 | Low Cost Sparse Network Monitoring Based on Block Matrix CompletionabstractDue to high network measurement cost, network-wide monitoring faces many challenges. For a network consisting of n nodes, the cost of one time network-wide monitoring will be O(n2). To reduce the monitoring cost, inspired by recent progress of matrix completion, a novel sparse network monitoring scheme is proposed to obtain network-wide monitoring data by sampling a few paths while inferring monitoring data of others. However, current sparse network monitoring schemes suffer from the problems of high measurement cost, high computation complexity in sampling scheduling, and long time to recover the un-sampled data. We propose a novel block matrix completion that can guarantee the quality of the un-sampled data inference by selecting as few as m = O(nr ln(r)) samples for a rank r N × T matrix with n = max{N,T}, which largely reduces the sampling complexity as compared to the existing algorithm for matrix completion. Based on block matrix completion, we further propose a light weight sampling scheduling algorithm to select measurement samples and a light weight data inference algorithm to quickly and accurately recover the un-sampled data. Extensive experiments on three real network monitoring data sets verify our theoretical claims and demonstrate the effectiveness of the proposed algorithms. Kun Xie 0001, Jiazheng Tian, Gaogang Xie, Guangxing Zhang, Da-Fang Zhang 0001 |
INFOCOM | 5 |
| 2021 | Multivariate Time Series Forecasting exploiting Tensor Projection Embedding and Gated Memory NetworkabstractTime series forecasting is very important and plays critical roles in many applications. However, making accurate forecasting is a challenge task due to the requirements of learning complex temporal and spatial patterns and combating noise during the feature learning. To address the challenge issues, we propose TEGMNet, a Tensor projection Embedding and Gated Memory Network for multivariate time series forecasting. To more accurately extract local features and reduce the influence of noise, we propose to amplify the data using several data transformation techniques based on MDT (Multi-way delay embedding transform) and TFNN (tensor factorized neural network) to transform the original 2D matrix data to low dimensional 3D tensor data. The local features are then extracted through convolution and LSTM upon the 3D tensor. We also design a long-term feature extraction module based on the structure of gated memory network, which can largely enhance the longterm pattern feature learning ability when the multivariate time series has complex long-term dependencies with dynamic-period patterns. We have done extensive experiments by comparing our TEGMNet with 7 baseline algorithms using 4 real data sets. The experiment results demonstrate that TEGMNet can achieve very good prediction performance even through the data are polluted with noise. Zhenxiong Yan, Kun Xie 0001, Xin Wang 0001, Da-Fang Zhang 0001, Gaogang Xie, Kenli Li 0001, Jigang Wen |
IWQoS | 4 |
| 2021 | An efficient and DoS-resilient name lookup for NDN interest forwardingabstractAs a novel Internet architecture focused on data contents, Named Data Networking (NDN) has been proven to be of great value in supporting the Internet of Things, Edge computing, Blockchain, and other popular topics. They can benefit mainly from NDN's intrinsic properties, such as flexible multicasting, in-network caching, among several others. However, once NDN's forwarding plane suffers Denial of Service (DoS) attacks, the overall system performance would be affected significantly. In NDN data transmission, interest forwarding is the most time-consuming operation and thus opens a possible vector of DoS attacks. It is proposed in this paper a fast name lookup algorithm for NDN interest forwarding, which selects feature prefixes instead of lengths to filter out interest packets in NDN interest forwarding. Due to the excellent filtration with feature prefixes, the algorithm accelerates NDN forwarding processes. Compared with other existing solutions, the proposed algorithm shows more than 70% of time improvement in forwarding malicious interests, while remaining at the same performance level in standard cases. Dacheng He, Da-Fang Zhang 0001, Yanbiao Li 0001, Wei Liang 0005, Meng-Yen Hsieh |
Connect. Sci. | 2 |
| 2021 | Secure fusion approach for the Internet of Things in smart autonomous multi-robot systems
Wei Liang 0005, Zuoting Ning, Songyou Xie, Yupeng Hu 0004, Shaofei Lu, Da-Fang Zhang 0001 |
Inf. Sci. | 6 |
| 2021 | Efficiently Inferring Top-k Largest Monitoring Data Entries Based on Discrete Tensor CompletionabstractNetwork-wide monitoring is important for many network functions. Due to the need of sampling to reduce high measurement cost, system failure, and unavoidable data transmission loss, network monitoring systems suffer from the incompleteness of network monitoring data. Different from the traditional network monitoring data estimation problem which aims to infer all missing monitoring data entries with incomplete measurement data, we study a challenging problem of inferring the top-$k$largest monitoring data entries. The recent study shows it is promising to more accurately interpolate the missing data with a 3-D tensor compared to that based on a 2-D matrix. Taking full advantage of the multilinear structures, we apply tensor completion to first recover the missing data and then find the top-$k$data entries. To reduce the computational overhead, we propose a novel discrete tensor completion model which uses binary codes to represent the factor matrices. Based on the model, we further propose three novel techniques to speed up the whole top-$k$entry inference process: a discrete optimization algorithm to train the binary factor matrices, bit operations to facilitate quick missing data inference, and simplifying the finding of top-$k$largest entries with binary code partition. In our discrete tensor completion model, only one bit is needed to represent the entry in the factor matrices instead of a real value (32 bits) needed in traditional tensor completion model, thus the storage cost is reduced significantly. To quickly infer the top-$k$largest data entries when measurement data arrive sequentially, we also propose a sliding window based online algorithm using the discrete tensor completion model. Extensive experiments using five real data sets and one synthetic data set demonstrate that compared with the state of art tensor completion algorithms, our discrete tensor completion algorithm can achieve similar top-$k$entry inference accuracy using significantly smaller time and storage space. Jiazheng Tian, Kun Xie 0001, Xin Wang 0001, Gaogang Xie, Kenli Li 0001, Jigang Wen, Da-Fang Zhang 0001, Jiannong Cao 0001 |
IEEE/ACM Trans. Netw. | 7 |
| 2020 | Neural Tensor Completion for Accurate Network MonitoringabstractMonitoring the performance of a large network is very costly. Instead, a subset of paths or time intervals of the network can be measured while inferring the remaining network data by leveraging their spatiotemporal correlations. The quality of missing data recovery highly relies on the inference algorithms. Tensor completion has attracted some recent attentions with its capability of exploiting the multi-dimensional data structure for more accurate missing data inference. However, current tensor completion algorithms only model the three-order interaction of data features through the inner product, which is insufficient to capture the high-order, nonlinear correlations across different feature dimensions. In this paper, we propose a novel Neural Tensor Completion (NTC) scheme to effectively model three-order interaction among data features with the outer product and build a 3D interaction map. Based on which, we apply 3D convolution to learn features of high-order interaction from the local range to the global range. We demonstrate this will lead to good learning ability. We conduct extensive experiments on two real-world network monitoring datasets, Abilene and WS-DREAM, to demonstrate that NTC can significantly reduce the error in missing data recovery. When the sampling ratio is low at 1%, the recovery error ratios on the testing data are around 0.05 (Abilene) and 0.13 (WS-DREAM) when using NTC, but are 0.99 (Abilene) and 0.99 (WS-DREAM) using the best current tensor completion algorithms, which are 21 times and 8 times larger. Kun Xie 0001, Huali Lu, Xin Wang 0001, Gaogang Xie, Yong Ding 0005, Dongliang Xie, Jigang Wen, Da-Fang Zhang 0001 |
INFOCOM | 8 |
| 2020 | Deep Reinforcement Learning for Resource Protection and Real-Time Detection in IoT EnvironmentabstractWith the fast advancements of electronic chip technologies in the Internet of Things (IoT), it is urgent to address the copyright protection issue of intellectual property (IP) circuit resources of the electronic devices in IoT environments. In this article, a fast deep-reinforcement-learning (DRL)-based detection algorithm for virtual IP watermarks is proposed by combining the technologies of mapping function and DRL to preprocess the ownership information of the IP circuit resource. The deep$Q$-learning (DQN) algorithm is used to generate the watermarked positions adaptively, making the watermarked positions secure yet close to the original design, turning the watermarked positions secure. An artificial neural network (ANN) algorithm is utilized for training the position distance characteristic vectors of the IP circuit, in which the characteristic function of the virtual position for IP watermark is generated after training. In IP ownership verification, the DRL model can quickly locate the range of virtual watermark positions. With the characteristic values of the virtual positions in each lookup table (LUT) area and surrounding areas, the mapping position relationship can be calculated in a supervised manner in the neural network, as the algorithm realizes the fast location of the real ownership information in an IP circuit. The experimental results show that the proposed algorithm can effectively improve the speed of watermark detection as also reducing the resource overhead. Besides, it also achieves excellent performance in security. Wei Liang 0005, Jing Long, Kuanching Li, Da-Fang Zhang 0001 |
IEEE Internet Things J. | 6 |
| 2020 | Singular Spectrum Analysis for Local Differential Privacy of Classifications in the Smart GridabstractNew privacy implications are induced to individuals and families because of the time-series data classification problem in the Internet of Things such as appliance classifications in the smart grid. To prevent the adversary from inferring the household appliance classification used in the smart grid, a singular spectrum analysis (SSA) has been applied to the local differential privacy (SSA-LDP). First, the Fourier spectrum noise has been added via the geometric sum which has been proved to achieve the Laplace noise distribution. Furthermore, we have proved that the sanitized data through the SSA-LDP is ε -deferentially private for the adversary inference attack. In addition, to achieve a better data utility, a formula has been obtained for the optimal Fourier spectrum noise by decomposing it into the superposition of power spectra of the dominant SSA eigenfilters. Finally, experiments have been performed with a computer-generated data set and a real-world smart-meter data set. Comparisons to other privacy approaches show that the optimized SSA-LDP does achieve a better data utility for a given data privacy. Lu Ou, Zheng Qin 0001, Shaolin Liao, Tao Li 0006, Da-Fang Zhang 0001 |
IEEE Internet Things J. | 5 |
| 2020 | Secure Data Storage and Recovery in Industrial Blockchain Network EnvironmentsabstractThe massive redundant data storage and communication in network 4.0 environments have issues of low integrity, high cost, and easy tampering. To address these issues, in this article, a secure data storage and recovery scheme in the blockchain-based network is proposed by improving the decentration, tampering-proof, real-time monitoring, and management of storage systems, as such design supports the dynamic storage, fast repair, and update of distributed data in the data storage system of industrial nodes. A local regenerative code technology is used to repair and store data between failed nodes while ensuring the privacy of user data. That is, as the data stored are found to be damaged, multiple local repair groups constructed by vector code can simultaneously yet efficiently repair multiple distributed data storage nodes. Based on the unique chain storage structure, such as data consensus mechanism and smart contract, the storage structure of blockchain distributed coding not only quickly repair the nearby local regenerative codes in the blockchain but also reduce the resource overhead in the data storage process of industrial nodes. Experimental results show that the proposed scheme improves the repair rate of multinode data by 9% and data storage rate increased by 8.6%, indicating to be promising with good security and real-time performance. Wei Liang 0005, Yongkai Fan, Kuanching Li, Da-Fang Zhang 0001, Jean-Luc Gaudiot |
IEEE Trans. Ind. Informatics | 4 |
| 2019 | Dynamic Spatial-Temporal Graph Convolutional Neural Networks for Traffic ForecastingabstractGraph convolutional neural networks (GCNN) have become an increasingly active field of research. It models the spatial dependencies of nodes in a graph with a pre-defined Laplacian matrix based on node distances. However, in many application scenarios, spatial dependencies change over time, and the use of fixed Laplacian matrix cannot capture the change. To track the spatial dependencies among traffic data, we propose a dynamic spatio-temporal GCNN for accurate traffic forecasting. The core of our deep learning framework is the finding of the change of Laplacian matrix with a dynamic Laplacian matrix estimator. To enable timely learning with a low complexity, we creatively incorporate tensor decomposition into the deep learning framework, where real-time traffic data are decomposed into a global component that is stable and depends on long-term temporal-spatial traffic relationship and a local component that captures the traffic fluctuations. We propose a novel design to estimate the dynamic Laplacian matrix of the graph with above two components based on our theoretical derivation, and introduce our design basis. The forecasting performance is evaluated with two realtime traffic datasets. Experiment results demonstrate that our network can achieve up to 25% accuracy improvement. Zulong Diao, Xin Wang 0001, Da-Fang Zhang 0001, Yingru Liu, Kun Xie 0001, Shaoyao He |
AAAI | 3 |
| 2019 | Efficiently Inferring Top-k Elephant Flows based on Discrete Tensor CompletionabstractFinding top- k elephant flows is a critical task in network measurement, with applications such as congestion control, anomaly detection, and traffic engineering. Traditional top- k flow detection problem focuses on using a small amount of memory to measure the total number of packets or bytes of each flow. Instead, we study a challenging problem of inferring the top- k elephant flows in a practical system with incomplete measurement data as a result of sub-sampling for scalability or data missing. The recent study shows it is promising to more accurately interpolate the missing data with a 3-D tensor compared to that based on a 2-D matrix. Taking full advantage of the multilinear structures, we apply tensor completion to first recover the missing data and then find the top- k elephant flows. To reduce the computational overhead, we propose a novel discrete tensor completion model which uses binary codes to represent the factor matrices. Based on the model, we further propose three novel techniques to speed up the whole top- k flow inference process: a discrete optimization algorithm to train the binary factor matrices, bit operations to facilitate quick missing data inference, and simplifying the finding of top- k elephant flows with binary code partition. In our discrete tensor completion model, only one bit is needed to represent the entry in the factor matrices instead of a real value (32 bits) needed in traditional tensor completion model, thus the storage cost is reduced significantly. Extensive experiments using two real traces demonstrate that compared with the state of art tensor completion algorithms, our discrete tensor completion algorithm can achieve similar data inference accuracy using significantly smaller time and storage space. Kun Xie 0001, Jiazheng Tian, Xin Wang 0001, Gaogang Xie, Jigang Wen, Da-Fang Zhang 0001 |
INFOCOM | 6 |
| 2019 | Active Sparse Mobile Crowd Sensing Based on Matrix CompletionabstractA major factor that prevents the large scale deployment of Mobile Crowd Sensing (MCS) is its sensing and communication cost. Given the spatio-temporal correlation among the environment monitoring data, matrix completion (MC) can be exploited to only monitor a small part of locations and time, and infer the remaining data. Rather than only taking random measurements following the basic MC theory, to further reduce the cost of MCS while ensuring the quality of missing data inference, we propose an Active Sparse MCS (AS-MCS) scheme which includes a bipartite-graph-based sensing scheduling scheme to actively determine the sampling positions in each upcoming time slot, and a bipartite-graph-based matrix completion algorithm to robustly and accurately recover the un-sampled data in the presence of sensing and communications errors. We also incorporate the sensing cost into the bipartite-graph to facilitate low cost sample selection and consider the incentives for MCS. We have conducted extensive performance studies using the data sets from the monitoring of PM 2.5 air condition and road traffic speed, respectively. Our results demonstrate that our AS-MCS scheme can recover the missing data at very high accuracy with the sampling ratio only around $11%$, while the peer matrix completion algorithms with similar recovery performance requires up to 4-9 times the number of samples of ours for both the data sets. Kun Xie 0001, Xiaocan Li, Xin Wang 0001, Gaogang Xie, Jigang Wen, Da-Fang Zhang 0001 |
SIGMOD Conference | 6 |
| 2019 | A double PUF-based RFID identity authentication protocol in service-centric internet of things environments
Wei Liang 0005, Songyou Xie, Jing Long, Kuanching Li, Da-Fang Zhang 0001, Keqin Li 0001 |
Inf. Sci. | 5 |
| 2019 | A Hybrid Model for Short-Term Traffic Volume Prediction in Massive Transportation SystemsabstractThe prediction of short-term volatile traffic becomes increasingly critical for efficient traffic engineering in intelligent transportation systems. Accurate forecast results can assist in traffic management and pedestrian route selection, which will help alleviate the huge congestion problem in the system. This paper presents a novel hybrid DTMGP model to accurately forecast the volume of passenger flows multi-step ahead with the comprehensive consideration of factors from temporal, origin-destination spatial, and frequency and self-similarity perspectives. We first apply discrete wavelet transform to decompose the traffic volume series into an appropriation component and several detailed components. Then we propose a more efficient tracking model to forecast the appropriation component and a novel Gaussian process model to forecast the detailed components. The forecasting performance is evaluated with real-time passenger flow data in Chongqing, China. Simulation results demonstrate that our hybrid model can achieve on average 20%-50% accuracy improvement, especially during rush hours. Zulong Diao, Da-Fang Zhang 0001, Xin Wang 0001, Kun Xie 0001, Shaoyao He, Xin Lu 0002, Yanbiao Li 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | Accurate Recovery of Missing Network Measurement Data With Localized Tensor CompletionabstractThe inference of the network traffic data from partial measurements data becomes increasingly critical for various network engineering tasks. By exploiting the multi-dimensional data structure, tensor completion is a promising technique for more accurate missing data inference. However, existing tensor completion algorithms generally have the strong assumption that the tensor data have a global low-rank structure, and try to find a single and global model to fit the data of the whole tensor. In a practical network system, a subset of data may have stronger correlation. In this work, we propose a novel localized tensor completion model (LTC) to increase the data recovery accuracy by taking advantage of the stronger local correlation of data to form and recover sub-tensors each with a lower rank. Despite that it is promising to use local tensors, the finding of correlated entries faces two challenges, the data with adjacent indexes are not ones with higher correlation and it is difficult to find the similarity of data with missing tensor entries. To conquer the challenges, we propose several novel techniques: efficiently calculating the candidate anchor points based on locality-sensitive hash (LSH), building sub-tensors around properly selected anchor points, encoding factor matrices to facilitate the finding of similarity with missing entries, and similarity-aware local tensor completion and data fusion. We have done extensive experiments using real traffic traces. Our results demonstrate that LTC is very effective in increasing the tensor recovery accuracy without depending on specific tensor completion algorithms. Kun Xie 0001, Xiangge Wang, Xin Wang 0001, Gaogang Xie, Yudian Ouyang, Jigang Wen, Jiannong Cao 0001, Da-Fang Zhang 0001 |
IEEE/ACM Trans. Netw. | 9 |
| 2018 | High-Performance IPv6 Lookup with Real-Time Updates Using Hierarchical-Balanced Search TreeabstractPacket forwarding is the foundation of network communications. Internet routers forward packets by matching the packet destination address against the prefixes in the forwarding information base. Due to the emergence of cloud computing and network function virtualization, more and more applications require high-bandwidth low-latency IPv6 networks. However, most of the current IP lookup algorithms in practice are designed for IPv4 addresses and hard to scale to IPv6 addresses. In this paper, we first propose a novel data structure called hierarchical balanced search tree (Hi-BST). The compact and scalable data structure built can be stored in on-chip memory. On the basis of Hi-BST, we present a fast IPv6 lookup algorithm and prefix update algorithms. In the worst case, our algorithm can achieve fast lookup and update with the time that is proportional to the logarithm of the number of prefixes. Experiments using real FIBs give an integrated evaluation. Compared with the state-of-the-art algorithms, our algorithm achieves a fast lookup with realtime updates, which is 1.33 ~6.78 times that of the comparison algorithms respectively. Moreover, Hi-BST saves 4.3% 79.0% memory footprint of the comparison algorithms respectively. Gaogang Xie, Da-Fang Zhang 0001 |
GLOBECOM | 4 |
| 2018 | Local Tensor Completion Based on Locality Sensitive HashingabstractTensor completion can be applied to fill in the missing data, which is import for many data applications where the data are incomplete. To infer the missing data, existing tensor-completion algorithms generally assume that the tensor data have global low-rank structure and apply a single model to fit the overall observed data through the global optimization. However, there are different correlation levels among application data, thus the ranks of some sub-tensors can be even lower relative to that of the large tensor. Fitting a single model to all data will compromise the performance of data recovery. To increase the accuracy in missing data recovery, we propose to apply local tensor completion (Local-TC) to recover data from sub-tensors, with each containing data of higher correlations. Although promising, as the tensor data are only organized logically, it is difficult to determine the relationship among data. We propose to exploit locality-sensitive hash (LSH) to quickly find the data correlation and reorganize tensor data, based on which data entries with high correlations are put into the same sub-tensor. The experiment results demonstrate that Local-TC is very effective in increasing the recovery accuracy. Kun Xie 0001, Xin Wang 0001, Gaogang Xie, Jigang Wen, Da-Fang Zhang 0001 |
ICDE | 6 |
| 2018 | Graph based Tensor Recovery for Accurate Internet Anomaly DetectionabstractDetecting anomalous traffic is a crucial task of managing networks. Many anomaly detection algorithms have been proposed recently. However, constrained by their matrix-based traffic data model, existing algorithms often suffer from low detection accuracy. To fully utilize the multi-dimensional information hidden in the traffic data, this paper takes an initiative to investigate the potential and methodologies of performing tensor factorization for more accurate Internet anomaly detection. Only considering the low-rank linearity features hidden in the data, current tensor factorization techniques would result in low anomaly detection accuracy. We propose a novel Graph-based Tensor Recovery model (Graph-TR) to well explore both low rank linearity features as well as the non-linear proximity information hidden in the traffic data for better anomaly detection. We encode the non-linear proximity information of the traffic data by constructing nearest neighbor graphs and incorporate this information into the tensor factorization using the graph Laplacian. Moreover, to facilitate the quick building of neighbor graph, we propose a nearest neighbor searching algorithm with the simple locality-sensitive hashing (LSH). We have conducted extensive experiments using Internet traffic trace data Abilene and GEANT. Compared with the state of art algorithms on matrix-based anomaly detection and tensor recovery approach, our Graph-Trcan achieve significantly lower False Positive Rate and higher True Positive Rate. Kun Xie 0001, Xiaocan Li, Xin Wang 0001, Gaogang Xie, Jigang Wen, Da-Fang Zhang 0001 |
INFOCOM | 6 |
| 2018 | CoDE: Fast Name Lookup and Update using Conflict-driven EncodingabstractLike IP lookup in the traditional networking, name lookup is a key technology for packet forwarding in the named data networking (NDN). However, unlike fixed-length IP addresses, such hierarchical names are of variable and unlimited length in theory. Both the large-scale name prefix database and high-frequency name update bring unprecedented challenges to the high-performance packet forwarding in the NDN. However, most existing approaches have drawbacks, such as the complex structure, the frequent memory access, and the time-consuming encoding, which make them difficult to meet these requirements. In this paper, we propose CoDE, an effective name lookup approach, to achieve both fast name lookup and update using conflict-driven encoding. CoDE has the following features: 1) The compact and scalable data structure can be stored in the cache; 2) The efficient index can fast locate the possible names and thus significantly speed up the name lookup and prefix update; and 3) The conflict-driven mechanism can greatly reduce the number of name components to be encoded. Experiments using real name prefix databases give an integrated evaluation. Compared with the state-of-the-art algorithms, CoDE achieves a high-performance name lookup which is an order of magnitude faster than the other algorithms on average and performs a fast prefix update which is twenty times that of the other algorithms on average. Moreover, CoDE saves at least half memory footprint of that of the other algorithms. Xinyi Zhang 0004, Gaogang Xie, Yuanmei Meng, Da-Fang Zhang 0001 |
IPCCC | 5 |
| 2018 | Optimizing Multi-Dimensional Packet Classification for Multi-Core Systems
Da-Fang Zhang 0001, Gaogang Xie, Xinyi Zhang 0004 |
J. Comput. Sci. Technol. | 2 |
| 2018 | Cyberspace Security for Future InternetabstractCyberspace is the most popular environment for information exchange whose security suffers from ever-increasing chal- Da-Fang Zhang 0001, Wenjia Li |
Secur. Commun. Networks | 1 |
| 2018 | On-Line Anomaly Detection With High Accuracy
Kun Xie 0001, Xiaocan Li, Xin Wang 0001, Jiannong Cao 0001, Gaogang Xie, Jigang Wen, Da-Fang Zhang 0001, Zheng Qin 0001 |
IEEE/ACM Trans. Netw. | 7 |
| 2018 | Accurate Recovery of Internet Traffic Data Under Variable Rate Measurements
Kun Xie 0001, Can Peng, Xin Wang 0001, Gaogang Xie, Jigang Wen, Jiannong Cao 0001, Da-Fang Zhang 0001, Zheng Qin 0001 |
IEEE/ACM Trans. Netw. | 7 |
| 2018 | Accurate Recovery of Internet Traffic Data: A Sequential Tensor Completion Approach
Kun Xie 0001, Lele Wang 0003, Xin Wang 0001, Gaogang Xie, Jigang Wen, Guangxing Zhang, Jiannong Cao 0001, Da-Fang Zhang 0001 |
IEEE/ACM Trans. Netw. | 8 |
| 2017 | Simultaneous Wireless Information and Power Transfer for Multi-hop Energy-Constrained Wireless Network
Shiming He, Kun Xie 0001, Weiwei Chen 0004, Da-Fang Zhang 0001, Jigang Wen |
WASA | 4 |
| 2017 | Accurate traffic matrix completion based on multi-Gaussian models
Huibin Zhou, Da-Fang Zhang 0001, Kun Xie 0001 |
Comput. Commun. | 2 |
| 2017 | Energy-efficient fuzzy control model for GPU-accelerated packet classificationabstractSummary As a core component of many network infrastructures, packet classification requires matching packet headers against a series of predefined rules. Its performance determines, to some extent, how fast packets can be processed. There already exists many proposals, which optimize the throughput of packet classification, but few of them take power consumption into account. To meet the requirements of green network computing, this paper focuses on energy‐efficient solutions that provide reasonable throughput as well. Similar to recent advancements, the graphics processing unit (GPU) is adopted to accelerate rule matching. Then, inspired by the frequency‐variable energy‐consuming model for air conditioners, a fuzzy control–based energy efficiency optimizing model is proposed for GPU‐accelerated packet classification. As demonstrated in the evaluation experiments, when the GPU is in the idle status, the proposed model can save 10 W. In running status, the fuzzy control–based energy efficiency optimizing model can avoid GPU shutdown issue caused by GPU self‐protection mechanism when the GPU temperature rises to 95°C. Furthermore, by improving the resource configuration of GPU kernels according to the model, the overall energy efficiency is enhanced by up to 15.5%, while simultaneously keeping throughput at the same level. Da-Fang Zhang 0001, Yanbiao Li 0001, Jintao Zheng, Keqin Li 0001 |
Concurr. Comput. Pract. Exp. | 2 |
| 2017 | Mlifdect: Android Malware Detection Based on Parallel Machine Learning and Information FusionabstractIn recent years, Android malware has continued to grow at an alarming rate. More recent malicious apps’ employing highly sophisticated detection avoidance techniques makes the traditional machine learning based malware detection methods far less effective. More specifically, they cannot cope with various types of Android malware and have limitation in detection by utilizing a single classification algorithm. To address this limitation, we propose a novel approach in this paper that leverages parallel machine learning and information fusion techniques for better Android malware detection, which is named Mlifdect. To implement this approach, we first extract eight types of features from static analysis on Android apps and build two kinds of feature sets after feature selection. Then, a parallel machine learning detection model is developed for speeding up the process of classification. Finally, we investigate the probability analysis based and Dempster-Shafer theory based information fusion approaches which can effectively obtain the detection results. To validate our method, other state-of-the-art detection works are selected for comparison with real-world Android apps. The experimental results demonstrate that Mlifdect is capable of achieving higher detection accuracy as well as a remarkable run-time efficiency compared to the existing malware detection solutions. Xin Wang 0029, Da-Fang Zhang 0001, Xin Su 0004, Wenjia Li |
Secur. Commun. Networks | 2 |
| 2017 | Fast Tensor Factorization for Accurate Internet Anomaly DetectionabstractDetecting anomalous traffic is a critical task for advanced Internet management. Many anomaly detection algorithms have been proposed recently. However, constrained by their matrix-based traffic data model, existing algorithms often suffer from low accuracy in anomaly detection. To fully utilize the multi-dimensional information hidden in the traffic data, this paper takes the initiative to investigate the potential and methodologies of performing tensor factorization for more accurate Internet anomaly detection. More specifically, we model the traffic data as a three-way tensor and formulate the anomaly detection problem as a robust tensor recovery problem with the constraints on the rank of the tensor and the cardinality of the anomaly set. These constraints, however, make the problem extremely hard to solve. Rather than resorting to the convex relaxation at the cost of low detection performance, we propose TensorDet to solve the problem directly and efficiently. To improve the anomaly detection accuracy and tensor factorization speed, TensorDet exploits the factorization structure with two novel techniques, sequential tensor truncation and two-phase anomaly detection. We have conducted extensive experiments using Internet traffic trace data Abilene and GÈANT. Compared with the state of art algorithms for tensor recovery and matrix-based anomaly detection, TensorDet can achieve significantly lower false positive rate and higher true positive rate. Particularly, benefiting from our well designed algorithm to reduce the computation cost of tensor factorization, the tensor factorization process in TensorDet is 5 (Abilene) and 13 (GÈANT) times faster than that of the traditional Tucker decomposition solution. Kun Xie 0001, Xiaocan Li, Xin Wang 0001, Gaogang Xie, Jigang Wen, Jiannong Cao 0001, Da-Fang Zhang 0001 |
IEEE/ACM Trans. Netw. | 7 |
| 2016 | An Approach of Anti-Eavesdropping Linear Network Coding in Wireless NetworkabstractThe existing anti-eavesdropping researches on wireless network coding are mainly based on the assumption that the eavesdroppers can only monitor limited channels and don't cooperate with each other. But in real settings, the eavesdroppers share overheard information, which makes data leakage come true. Moreover, these approaches just encrypt or permutate the coding coefficients to prevent the adversaries from understanding the transmission data, they can't resist pollution attacks, such as forgery and tamper attacks. In this case, we propose a novel secure linear network coding scheme that is based on IBC(Identity-Based Cryptography) algorithm and has the characteristic of preventing eavesdropping and pollution attacks. Theoretical analysis and experiment demonstrate that our scheme guarantees the data security for each node, which significantly prevents the nodes from being eavesdropped and resists pollution attacks. Furthermore, our scheme enriches the approaches of anti-eavesdropping research in wireless network coding. Zuoting Ning, Da-Fang Zhang 0001, Kun Xie 0001 |
ICPADS | 2 |
| 2016 | Minimizing datacenter flow completion times with server-based flow scheduling
Jie Zhang 0043, Da-Fang Zhang 0001, Kun Huang 0003, Zheng Qin 0001 |
Comput. Networks | 2 |
| 2016 | A splitting-after-merging approach to multi-FIB compression and fast refactoring in virtual routersabstractVirtual routers are gaining increasing attention in the research field of future networks. As the core network device to achieve network virtualization, virtual routers have multiple virtual instances coexisting on a physical router platform, and each instance retains its own forwarding information base (FIB). Thus, memory scalability suffers from the limited on-chip memory. In this paper, we present a splitting-after-merging approach to compress the FIBs, which not only improves the memory efficiency but also offers an ideal split position to achieve system refactoring. Moreover, we propose an improved strategy to save the time used for system rebuilding to achieve fast refactoring. Experiments with 14 real-world routing data sets show that our approach needs only a unibit trie holding 134 188 nodes, while the original number of nodes is 4 569 133. Moreover, our approach has a good performance in scalability, guaranteeing 90 000 000 prefixes and 65 600 FIBs. Da-Fang Zhang 0001, Yanbiao Li 0001, Kun Xie 0001 |
Frontiers Inf. Technol. Electron. Eng. | 1 |
| 2015 | Completion Time-Aware Flow Scheduling in Heterogenous Networks
Shiming He, Kun Xie 0001, Da-Fang Zhang 0001 |
ICA3PP (1) | 3 |
| 2015 | Fast and Scalable Regular Expressions Matching with Multi-Stride Index NFA
Sheng Huo, Da-Fang Zhang 0001, Yanbiao Li 0001 |
ICA3PP (3) | 2 |
| 2015 | A GPU Based Fast Community Detection Implementation for Social Network
Da-Fang Zhang 0001, Kun Xie 0001, Tanlong Huang, Yanbiao Li 0001 |
ICA3PP (1) | 2 |
| 2015 | A Group XOR-ing Coding Strategy Based on Wireless Network Overhearing
Zuoting Ning, Da-Fang Zhang 0001, Kun Xie 0001 |
ICA3PP (1) | 2 |
| 2015 | A Distributed Joint Cooperative Routing and Channel Assignment in Multi-radio Wireless Mesh Network
Hong Qiao, Da-Fang Zhang 0001, Kun Xie 0001, Shiming He |
ICA3PP (1) | 2 |
| 2015 | An Intimacy-Based Algorithm for Social Network Community Detection
Da-Fang Zhang 0001, Kun Xie 0001 |
ICA3PP (1) | 2 |
| 2015 | Spatio-temporal tensor completion for imputing missing internet traffic dataabstractNetwork traffic data consists of Traffic Matrix (TM), which represents the volumes of traffic between Origin and Destination (OD) pairs in the network. It is a key input parameter of network engineering tasks. However, direct measurement of the OD pairs traffic is usually not feasible. Even good traffic measurement systems can suffer from errors, missing data. So obtaining the ODs traffic precisely is a challenge. Existing completion methods often perform poorly for network traffic estimation. Their recovery accuracy tends to be significantly worse when the data loss rate is high. Taking into account network traffic lower-dimensional latent structure and traffic hidden characteristic, a tensor (multi-way array) is introduced to model a time series of pure spatial traffic matrices in this paper. To recover the missing entries in tensors of traffic data, a novel spatio-temporal tensor completion method has been proposed. This approach not only takes advantage of tensor decomposition and its lower-dimensional representation, but also well takes into account traffic spatio-temporal properties. The extensive experiments with the real-world traffic trace data show that the proposed method can significantly reduce the missing traffic data recovery errors and achieve satisfactory completion accuracy comparing with the state-of-the-art completion methods. Huibin Zhou, Da-Fang Zhang 0001, Kun Xie 0001 |
IPCCC | 2 |
| 2015 | Android app recommendation approach based on network traffic measurement and analysisabstractA large amount and different types of mobile applications (or apps) are being offered to end users via app markets. These apps normally generate network traffic, which will consumes users' mobile data plan and may even cause potential security issues. However, the amount and type of network traffic generated by a mobile app in the wild is still poorly understood due to the lack of a systematic measurement methodology. In this paper, we first measure and analyze network traffic cost of Android apps in the official Android markets. Based on the results, we find that the apps from different categories have different traffic costs. In particular, there is a remarkable difference among the apps with similar functionality in terms of network traffic cost. Then, we add metrics of traffic cost into our app recommendation algorithm, which differs from the conventional app recommendation approaches. Experimental results show that the proposed recommendation algorithm can effectively help mobile app users avoid various potential security and privacy risks brought by the unnecessary network traffic consumption. Xin Su 0004, Da-Fang Zhang 0001, Wenjia Li |
ISCC | 2 |
| 2015 | Fest: A feature extraction and selection tool for Android malware detectionabstractAndroid has become one of the most popular mobile operating systems because of numerous applications (apps) it provides. However, Android malware downloaded from third-party markets threatens users' privacy, and most of them remain undetected because of the lack of efficient and accurate detecting techniques. Prior efforts on Android malware detection attempted to build precise classification models by manually choosing features, and few of them has used any feature selection algorithms to help pick typical features. In this paper, we present Feature Extraction and Selection Tool (Fest), a feature-based machine learning approach for malware detection. We first implement a feature extraction tool, AppExtractor, which is designed to extract features, such as permissions or APIs, according to the predefined rules. Then we propose a feature selection algorithm, FrequenSel. Unlike existing selection algorithms which pick features by calculating their importance, FrequenSel selects features by finding the difference their frequencies between malware and benign apps, because features which are frequently used in malware and rarely used in benign apps are more important to distinguish malware from benign apps. In experiments, we evaluate our approach with 7972 apps, and the results show that Fest gets nearly 98% accuracy and recall, with only 2% false alarms. Moreover, Fest only takes 6.5s to analyze an app on a common PC, which is very time-efficient for malware detection in Android markets. Da-Fang Zhang 0001, Xin Su 0004, Wenjia Li |
ISCC | 2 |
| 2015 | Improving datacenter throughput and robustness with Lazy TCP over packet spraying
Jie Zhang 0043, Da-Fang Zhang 0001, Kun Huang 0003 |
Comput. Commun. | 2 |
| 2015 | Signature Restoration for Enhancing Robustness of FPGA IP DesignsabstractMany watermarking techniques for intellectual property (IP) protection are not resilient to tampering or removal attacks, especially for field programmable gate array (FPGA)-based IP cores. If attacked, the damaged watermarks cannot provide sufficient evidence in front of a court. To address this issue, the authors present a signature restoration scheme. The thought of secret sharing is introduced to share the signature into small watermarks. These watermarks are encoded with Reed-Solomon (RS) codes and embedded into unused lookup tables (LUTs) of used slices. Unlike most of existing techniques, the proposed scheme can restore the signature only by extracting parts of watermarks. So, it is tolerant to some damaged watermarks caused by removal attacks. The experiments show that the proposed scheme incurs no extra hardware resource and timing overhead. The robustness against attacks is much better by comparing to other schemes. Jing Long, Da-Fang Zhang 0001, Wei Liang 0005, Xia-an Bi |
Int. J. Inf. Secur. Priv. | 2 |
| 2015 | Memory-efficient IP lookup using trie merging for scalable virtual routers
Kun Huang 0003, Gaogang Xie, Yanbiao Li 0001, Da-Fang Zhang 0001 |
J. Netw. Comput. Appl. | 4 |
| 2015 | AndroGenerator: An automated and configurable android app network traffic generation systemabstractAbstract With the rapid growth in the popularity of Android smartphones, a large number of Android applications (or apps) have emerged in both official and alternative Android markets. It is important for network operators and security analysts to understand the network traffic generated by new Android apps for the purposes of network management, app traffic analysis, and malware detection. However, it is time‐consuming and tedious to manually install and run Android apps to generate network traffic. Moreover, existing synthetic network traffic generators are unable to generate network traffic that can accurately reflect the network behaviors of Android apps. In this paper, we propose and implement AndroGenerator, an automated Android network traffic generation system, to generate various types of network traffic that can be produced by Android apps. Our system reproduces the network traffic based on the traffic characteristics extracted from traffic traces captured by running a large number of Android applications from several popular Android markets, such as Google Play. The system first generates network traffic through automated execution of Android applications. Then, the system is also able to extract network characteristics from the captured traffic traces and store the extracted results into a database for benchmarking purposes. Finally, AndroGenerator reproduces Android app traffic based via simulating network characteristics of captured traffic traces. In the experiments, we evaluate our system with real‐world mobile traffic, and the experiment results show that AndroGenerator can reproduce Android app network traffic accurately. Copyright © 2015 John Wiley & Sons, Ltd. Xin Su 0004, Da-Fang Zhang 0001, Wenjia Li |
Secur. Commun. Networks | 2 |
| 2015 | Detect repackaged Android application based on HTTP traffic similarityabstractIn recent years, more and more malicious authors aim to Android platform because of the rapid growth number of Android Google, Menlo Park, California, USA applications or apps. They embedded malicious code into Android apps to execute their special malicious behaviors, such as sending text messages to premium numbers, stealing privacy information, or even converting the infected phones into bots. We called the app, which has been embedded with malicious code, as embedded repackaged app. This phenomena leads a big security risk to the Android users and how to detect them becomes an urgent problem. Previous research efforts focus on extracting the app's characteristics for comparison from its static program code, which neither can handle the code obfuscation technologies, nor can analyze the app's dynamic behaviors feature. To address these limitations, we propose an approach based on extracting the app's characteristics from the HTTP traffic, which is generated by the app. Moreover, we have implemented a multi-thread comparison algorithm based on the balanced Vantage Point Tree VPT, which can remarkably reduce the experiment time. In this experiment, we successfully detected 266 embedded repackaged apps from 7619 Android apps downloaded from six popular Android markets, and the distribution rate of each market ranges from 2.57% to 6.07%. Then based on the analyzing of the HTTP traffic generated by these embedded codes, we found that majority of them are advertisement traffic and malicious traffic. Copyright © 2015 John Wiley & Sons, Ltd. Xueping Wu, Da-Fang Zhang 0001, Xin Su 0004 |
Secur. Commun. Networks | 2 |
| 2014 | From GPU to FPGA: A Pipelined Hierarchical Approach to Fast and Memory-Efficient NDN Name LookupabstractSummary form only given. Named Data Networking (NDN) is an emerging future Internet architecture with an alternative communication paradigm. For NDN, name lookup, just like IP address lookup for TCP/IP, plays an important role in forwarding. However, performing Longest Prefix Matching (LPM) to NDN names is more challenging. Recently, Graphic Processing Units (GPUs) have been shown to be of value in supporting wire speed name lookup, but the latency resulted by batching and transferring names is not so encouraging. On the other hand, in the area of IP address lookup, FPGA is widely used to implement Static Radom Accessing Memory (SRAM)-based pipeline for fast lookup and controllable latency. Thus, in this paper, we study how to accelerate NDN name lookup using FPGA-based pipeline. Yanbiao Li 0001, Da-Fang Zhang 0001, Jing Long, Wei Liang 0005 |
FCCM | 2 |
| 2014 | Accelerate NDN name lookup using FPGA: Challenges and a scalable approachabstractRecently, Graphic Processing Units (GPUs) have been shown to be of value in supporting wire-speed name lookup in Named Data Networking (NDN). However, due to the computing model on GPU, the lookup latency is not so encouraging. In this paper, we shift the focus from GPU to Field-Programmable Gate Arrays (FPGA). We highlight three key challenges in accelerating name lookup using FPGA, and then present a scalable approach to address them. In our approach, a hierarchical and compact data structure is proposed to represent the name trie, which achieves not only effective pipeline mapping but also high memory efficiency. Further, it is finally implemented as a linear pipeline on the FPGA platform, enabling both fast lookup speed and low lookup latency. The experimental results show that our approach gains a reduction of memory cost over 90% compared with the referred GPU-based solution. Besides, the lookup throughput of our approach is almost 2.4 times higher, and the latency is up to 3 orders of magnitude lower. Yanbiao Li 0001, Da-Fang Zhang 0001, Wei Liang 0005, Jing Long, Hong Qiao |
FPL | 2 |
| 2014 | A memory-efficient parallel routing lookup model with fast updates
Yanbiao Li 0001, Da-Fang Zhang 0001, Kun Huang 0003, Dacheng He, Weiping Long |
Comput. Commun. | 2 |
| 2014 | Channel Aware Opportunistic Routing in Multi-Radio Multi-Channel Wireless Mesh Networks
Shiming He, Da-Fang Zhang 0001, Kun Xie 0001, Hong Qiao |
J. Comput. Sci. Technol. | 2 |
| 2013 | Scalable TCAM-based regular expression matching with compressed finite automataabstractRegular expression (RegEx) matching is a core function of deep packet inspection in modern network devices. Previous TCAM-based RegEx matching algorithms a priori assume that a deterministic finite automaton (DFA) can be built for a given set of RegEx patterns. However, practical RegEx patterns contain complex terms like wildcard closure and repeat character, and it may be impossible to build a DFA with a reasonable number of states. This results in prior work to being infeasible in practice. Moreover, TCAM-based RegEx matching is required to scale to a large-scale set of RegEx patterns. In this paper, we propose a compressed finite automaton implementation called (CFA) for scalable TCAM-based RegEx matching. CFA is designed to reduce TCAM space by using three compression techniques: transition, character, and state compressions. Experiments on realistic RegEx pattern sets show CFA highly outperforms previous solutions in terms of TCAM space, matching throughput, and TCAM power consumption. Kun Huang 0003, Linxuan Ding, Gaogang Xie, Da-Fang Zhang 0001, Alex X. Liu, Kavé Salamatian |
ANCS | 4 |
| 2013 | GAMT: A fast and scalable IP lookup engine for GPU-based software routersabstractRecently, the Graphics Processing Unit (GPU) has been proved to be an exciting new platform for software routers, providing high throughput and flexibility. However, it is still a challenging task to deploy some core routing functions into GPU-based software routers with anticipatory performance and scalability, such as IP address lookup. Existing solutions have good performance, but their scalability to IPv6 and frequent updates are not so encouraging. In this paper, we investigate GPU's characteristics in parallelism and memory accessing, and then encode a multibit trie into a state-jump table. On this basis, a fast and scalable IP lookup engine called GPU-Accelerated Multi-bit Trie (GAMT) has been presented. According to our experiments on real-world routing data, based on the multi-stream pipeline, GAMT enables lookup speeds as high as 1072 and 658 Million Lookups Per Second (MLPS) for IPv4/6 respectively, when performing a 16M traffic under highly frequent updates (70, 000 updates/s). Even using a small batch size, GAMT can still achieve 339 and 240 MLPS respectively, while keeping the average lookup latency below 100 μs. These results show clearly that GAMT makes significant progress on both scalability and performance. Yanbiao Li 0001, Da-Fang Zhang 0001, Alex X. Liu, Jintao Zheng |
ANCS | 2 |
| 2013 | A Multi-partitioning Approach to Building Fast and Accurate Counting Bloom FiltersabstractBloom filters are space-efficient data structures for fast set membership queries. Counting Bloom Filters (CBFs) extend Bloom filters by allowing insertions and deletions to support dynamic sets. The performance of CBFs is critical for various applications and systems. This paper presents a novel approach to building a fast and accurate data structure called Multiple-Partitioned Counting Bloom Filter (MPCBF) that addresses large-scale data processing challenges. MPCBF is based on two ideas: reducing the number of memory accesses from k (for k hash functions) in the standard CBF to only one memory access in the basic MPCBF-1 case, and a hierarchical structure to improve the false positive rate. We also generalize MPCBF-1 to MPCBF-g to accommodate up to g memory accesses. Our simulation and implementation in MapReduce show that MPCBF outperforms the standard CBF in terms of speed and accuracy. Compared to CBF, at the same memory consumption, MPCBF significantly reduces the false positive rate by an order of magnitude, with a reduction of processing overhead by up to 85.9%. Kun Huang 0003, Jie Zhang 0043, Da-Fang Zhang 0001, Gaogang Xie, Kavé Salamatian, Alex X. Liu |
IPDPS | 3 |
| 2011 | A Simple Channel Assignment for Opportunistic Routing in Multi-radio Multi-channel Wireless Mesh NetworksabstractOpportunistic routing (OR) involves multiple forwarding candidates to relay packets by taking advantage of the broadcast nature and multi-user diversity of the wireless medium. Compared with Traditional Routing (TR), OR is more suitable for the unreliable wireless link, and can evidently improve the end to end throughput of Wireless Mesh Networks (WMNs). At present, there are many achievements concerning OR in the single radio wireless network. However, the study of OR in multi radio wireless network stays the beginning stage. In this paper, we focus on OR in multi-radio multi-channel WMNs. We validate the advantage of OR in multi-radio multi-channel WMNs, and propose a Simple Channel Assignment for Opportunistic Routing (SCAOR), which assigns channel to flows. According to interference state of every node, SCAOR assigns a channel with minimum interference to each flow to balance channel load. The simulation result shows OR of dual-radio dual-channel WMNs can promote throughput evidently, specifically, 16.8% higher than throughput of TR in the dual-radio dual-channel WMNs, 87.11% and 111.8% higher than throughput of OR and TR in single-radio single-channel, respectively. Shiming He, Da-Fang Zhang 0001, Kun Xie 0001, Hong Qiao |
MSN | 2 |
| 2011 | An index-split Bloom filter for deep packet inspection
Kun Huang 0003, Da-Fang Zhang 0001 |
Sci. China Inf. Sci. | 2 |
| 2010 | Accelerating the bit-split string matching algorithm using Bloom filters
Kun Huang 0003, Da-Fang Zhang 0001, Zheng Qin 0001 |
Comput. Commun. | 2 |
| 2010 | DHT-based lightweight broadcast algorithms in large-scale computing infrastructures
Kun Huang 0003, Da-Fang Zhang 0001 |
Future Gener. Comput. Syst. | 2 |
| 2009 | A Partition-Based Broadcast Algorithm over DHT for Large-Scale Computing Infrastructures
Kun Huang 0003, Da-Fang Zhang 0001 |
GPC | 2 |
| 2009 | An Activeness-Based Seed Choking Algorithm for Enhancing BitTorrent's Robustness
Kun Huang 0003, Da-Fang Zhang 0001, Li-e Wang 0001 |
GPC | 2 |
| 2009 | A Regular Expression Matching Algorithm Using Transition MergingabstractWith the rapid development of the network, deep packet inspection systems are faced with the challenge of high performance. On one hand, they try to reduce the memory consumption in the process of regular expression matching; on the other hand, they must provide a worst-case matching speed guarantee. The existing state merging finite automata algorithm reduces the number of states in the deterministic finite automata (DFA). But there are still a large amount of transitions. In this paper, we introduce a transition merging finite automata algorithm, which merges several transitions in the DFA, based on the state merging algorithm. The experiments show that the transition merging algorithm reduces the memory consumption by 15%~31% compared to the state merging algorithm, when compared to the original DFA, it reduces the memory consumption by 25%~42%. At the same time, the transition merging algorithm ensures the matching speed. It is a memory efficient regular expression matching algorithm. Jiekun Zhang, Da-Fang Zhang 0001, Kun Huang 0003 |
PRDC | 2 |
| 2009 | Trading off logging overhead and coordinating overhead to achieve efficient rollback recoveryabstractAbstract In the rollback recovery of large‐scale long‐running applications in a distributed environment, pessimistic message logging protocols enable failed processes to recover independently, though at the expense of logging every message synchronously during fault‐free execution. In contrast, coordinated checkpointing protocols avoid message logging, but they are poor in scalability with a sharply increased coordinating overhead as the system grows. With the aim of achieving efficient rollback recovery by trading off logging overhead and coordinating overhead, this paper suggests a partitioning of the system into clusters, and then presents a scheme to implement the conversion between these overheads. Using the proposed conversion, coordination can be introduced to reduce the unbearable logging overhead found in some systems, whereas proper logging can be employed to alleviate the unacceptable coordinating overhead in others. Furthermore, heuristics are introduced to address the issue of how to partition the system into clusters in order to speed up the recovery process and to improve recovery efficiency. Performance evaluation results indicate that our scheme can lower the overall system overhead effectively. Copyright © 2008 John Wiley & Sons, Ltd. Jinmin Yang, Kin Fun Li, Wen-Wei Li, Da-Fang Zhang 0001 |
Concurr. Comput. Pract. Exp. | 4 |
| 2008 | Performance Evaluation of End-to-End Path Capacity Measurement Tools in a Controlled Environment
Bin Zeng 0001, Da-Fang Zhang 0001, Jinmin Yang |
GPC | 3 |
| 2008 | Optimizing the BitTorrent performance using an adaptive peer selection strategy
Kun Huang 0003, Li-e Wang 0001, Da-Fang Zhang 0001, Yongwei Liu |
Future Gener. Comput. Syst. | 3 |
| 2007 | A Prediction-Based Fair Replication Algorithm in Structured P2P Systems
Xianshu Zhu, Da-Fang Zhang 0001, Wenjia Li, Kun Huang 0003 |
ATC | 2 |
| 2007 | A Coarse-Grained Pessimistic Message Logging Scheme for Improving Rollback Recovery EfficiencyabstractAs a common technology for fault tolerance and load balance, rollback-recovery faces the challenges of scalability and inherent variability in those long-running large-scale applications with grids as the computing infrastructure. Among the rollback recovery schemes, pessimistic message logging protocols (PMLPs) and coordinated checkpointing protocols (CCPs) are the most popular in practice. Although PMLPs are good in scalability, their fault-free overhead sometimes is prohibitive. CCPs introduce relatively lower overhead, but they are poor in scalability. This work employs partition strategy and introduces the concept of pessimism grain to rollback recovery, striking a balance between good scalability and acceptable overhead. For a partitioned system, a coarse-grained pessimistic message-logging protocol is proposed to achieve scalability and asynchrony both in fault-free execution and in fault recovery. The impact of pessimism grain on the performance is evaluated theoretically. Experimental results show that the pessimism grain is one of the key configuration parameters to reach a desired performance level. Jinmin Yang, Kin Fun Li, Da-Fang Zhang 0001 |
DASC | 3 |
| 2007 | A Scalable Bloom Filter for Membership QueriesabstractBloom filters allow membership queries over sets with allowable errors. It is widely used in databases, networks and distributed systems and it has great potential for distributed applications where systems need to share information about available data. However, the false positive errors are unavoidable, and the false positive rate increases intolerantly along with the date set expanding. To solve the scalability problem of Bloom filters, this paper presents a new design of a scalable Bloom filter (SBF) for an expanding data set. The SBF keeps a low false positive rate by adding Bloom filter vectors with double length when necessary. The paper proposes algorithms for element insertion and query operation of SBF by employing the H3class of universal hash functions. Theoretical and experimental results demonstrate that the new SBF provides false positive rate as low as 21.3% of the dynamic Bloom filter presented before and the querying CPU time increasing with logarithmic rather than linear. Therefore, the proposed SBF outperforms other current scalable Bloom filters significantly. Kun Xie 0001, Yinghua Min, Da-Fang Zhang 0001, Jigang Wen, Gaogang Xie |
GLOBECOM | 3 |
| 2007 | Self-Similar Characteristic of Traffic in Current Metro Area NetworkabstractComplexity and diversity of Internet traffic are constantly growing. Networking researchers become aware of the need to constantly monitor and reevaluate their assumptions in order to ensure that the conceptual models correctly represent reality. Using the dataset collected by NetTurbo from three different bidirectional OC-48 links in metro area networks at the two biggest ISPs of China, this paper carefully investigates the self-similar characteristics of traffic from different aspects. In contrast to the previous results which have been widely accepted, this paper shows that for the aggregated traffic and the TCP and UDP traffic whether the self-similarity exists is uncertain. Further, break down by the application category, only the traditional and uncategorized traffic are self-similar while the others are not. However, on the view of the individual application of each category, it seems that traffic of every application exhibits self-similarity. To the best of our knowledge, this paper firstly provides the experimental evidence showing that aggregating different groups of self-similar traffic series could generate a traffic series which is either self-similar or non-self-similar. Guangxing Zhang, Gaogang Xie, Dunxing Zhang, Da-Fang Zhang 0001 |
LANMAN | 5 |
| 2007 | On evaluating the differences of TCP and ICMP in network measurement
Da-Fang Zhang 0001, Jinmin Yang, Gaogang Xie |
Comput. Commun. | 2 |
| 2007 | Reliable user-level rollback recovery implementation for multithreaded processes on windowsabstractAbstract The existing user‐level checkpointing schemes support only a limited portion of multithreaded programs because they are derived from the schemes for single‐threaded applications. This paper addresses the impact of thread suspension point on rollback recovery, and presents a checkpointing scheme for multithreaded processes. Unlike the existing schemes in which the checkpointer suspends every working thread, our scheme employs a distinctive strategy that every working thread suspends itself. This technique manages to avoid the suspension point in the API code or kernel code, ensuring correct rollback recovery. Our scheme supports inter‐thread synchronization and thread lifetime. Copyright © 2006 John Wiley & Sons, Ltd. Jinmin Yang, Da-Fang Zhang 0001, Xue Dong Yang, Wen-Wei Li |
Softw. Pract. Exp. | 2 |
| 2006 | A Resource Allocating Neural Network Based Approach for Detecting End-to-End Network Performance Anomaly
Da-Fang Zhang 0001, Jinmin Yang, Gaogang Xie |
ISNN (2) | 2 |
| 2006 | On the Self-Similarity of the 1999 DARPA/Lincoln Laboratory Evaluation Data
Kun Huang 0003, Da-Fang Zhang 0001 |
SECRYPT | 2 |
| 2005 | Remote OS Fingerprinting Using BP Neural Network
Da-Fang Zhang 0001, Jinmin Yang |
ISNN (3) | 2 |
| 2005 | TCP and ICMP in Network Measurement: An Experimental Evaluation
Da-Fang Zhang 0001, Gaogang Xie, Jinmin Yang |
ISPA | 2 |
| 2004 | Bounding Rollback-Recovery of Large Distributed Computation in WAN EnvironmentabstractIn the existing optimistic message logging protocols, the dependency must be tracked in whole system, and all processes are involved in rollback recovery in the event of failure. For large distributed computation in WAN environment with the low available bandwidth and high transmission latency, its fault-free overhead and recovery overhead are outstanding, recovery efficiency decreasing with the scale of system. This paper introduces a three-layer model of large distributed system in WAN environment, and presents a protocol of message dependency tracking based on proxy. Utilizing private proxy to log messages and dependencies, the protocol limits rollback-recovery to a scope called block rather than the entire system, achieving relative low fault-free overhead and fast output commit, as well as improved recovery efficiency and low recovery overhead. Jinmin Yang, Da-Fang Zhang 0001 |
Asian Test Symposium | 2 |
| 2004 | WINDAR: A Multithreaded Rollback-Recovery Toolkit on WindowsabstractWe describe the design and implementation of WINDAR, an object-oriented toolkit for transparent rollback-recovery of distributed applications running on Windows platform. In WINDAR, the workloads of a process are multithreaded, exploiting effectively processor execution resources to improve execution efficiency. In addition, WINDAR's unified framework for various rollback recovery protocols enables dynamic protocol configuration to adapt itself to the need of recovery-oriented computing (ROC) and distributed computations in Internet environment. WINDAR was evaluated using three benchmarks. It is observed that multithreading is an effective approach to improve the performance of message logging protocols, especially for pessimistic message logging. In our experiment, the overhead ratio of pessimistic message logging was reduced to the same magnitude as that of the optimistic message logging for three benchmarks. Jinmin Yang, Da-Fang Zhang 0001, Zheng Qin 0001, Xue Dong Yang |
PRDC | 2 |
| 2003 | User-Level Implementation of Checkpointing for Multithreaded Applications on Windows NTabstractThe existing user-level checkpointing schemes support only a certain portion of multithreaded programs on the Windows operating system, which are based on single-threaded programs. This paper focuses on studying a checkpointing scheme to support inter-thread synchronization and quantitative variation of threads for multithreaded processes. Unlike other proposed schemes, in which a thread is suspended by another thread at checkpointing, our checkpointing scheme employs a strategy by which a thread suspends itself. Therefore, it is free of nondeterminacy of thread suspension point, thereby ensuring correct rollback recovery. Our checkpointing scheme supports also various synchronization objects such as Mutex, CriticalSection and Event, as well as Semaphore, WaitableTimer and Thread. Jinmin Yang, Da-Fang Zhang 0001, Xue Dong Yang |
Asian Test Symposium | 2 |
| 2001 | Node Grouping in System-Level Fault Diagnosis
Da-Fang Zhang 0001, Gaogang Xie, Yinghua Min |
J. Comput. Sci. Technol. | 1 |