VLDB 2026 Research / reviewers in the wild / expert
Xiangzhan Yu
dblp:16/5172 · also Xiang-Zhan Yu
· DBLP profile ↗
47ranked-venue papers
1as first author
31since 2021 · last 2026
0000-0002-1183-2844ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 18 · 1 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 8 since 2021Security and privacy · 10 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Measuring security weaknesses in underground mobile app ecosystems at scale
Yicheng Guo, Zhichao Hu, Likun Liu, Mengmeng Ge 0003, Wanzong Peng, Xueshan Wang, Xiangzhan Yu |
Comput. Secur. | 8 |
| 2026 | Neural cryptography: Synergy and conflict between neural networks and cryptographic systems
Jieming Gu, Xiangzhan Yu |
Neurocomputing | 3 |
| 2025 | TND: Two-stage non-invasive defense of intrusion detection system from adversarial attack
Zhichao Hu, Dewen Kong, Junzhong Miao, Gang Du, Likun Liu, Xiangzhan Yu |
Comput. Networks | 8 |
| 2025 | TOPLDM: Towards dynamic low overhead traffic obfuscation based on packet length distribution modification
Zhichao Hu, Likun Liu, Jiaxing Gong, Mengmeng Ge 0003, Xiangzhan Yu |
Comput. Networks | 9 |
| 2025 | CCLog: Actionable APT forensics via fused log semantics and provenance graph topology
Zhichao Hu, Likun Liu, Mengmeng Ge 0003, Xiangzhan Yu |
Comput. Networks | 8 |
| 2025 | Enmob: Unveil the Behavior with Multi-flow Analysis of Encrypted App TrafficabstractAbstract In the contemporary digital landscape, mobile applications have become the predominant conduit for internet connectivity and daily tasks. Simultaneously, the advent of application encryption technology has safeguarded users’ privacy. However, this encryption, while fortifying privacy, introduces challenges to security by hindering the effective management of network applications within encrypted data streams. Conventional detection methods for encrypted application traffic, relying heavily on statistical metrics like payload, packet size, and distribution, are constrained to single traffic flows, often yielding results of limited specificity. To address this limitation, our paper introduces an innovative approach that elucidates the multi-flow nature of application behavior traffic and provides context to encrypted application traffic. This method offers a more nuanced and comprehensive perspective for understanding and representing network traffic, even when encrypted. The efficacy of our approach was evaluated using a substantial volume of real network traffic data. Results indicate that our method achieves an average accuracy of 0.958 in identifying application behavior traffic and 0.955 in classifying application traffic. These outcomes signify a substantial enhancement over single network flow-based detection methods, demonstrating a notable 5.3% improvement. Mengmeng Ge 0003, Likun Liu, Xiangzhan Yu, Vinay Sachidananda, Xiaofei Xie, Yang Liu 0003 |
Cybersecur. | 4 |
| 2025 | SinkFlow: Fast and traceable root-cause localization for multidimensional anomaly events
Zhichao Hu, Likun Liu, Xiangzhan Yu |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | UDA: Unified Pretraining for Multiarchitecture Binary DisassemblyabstractThe precision of binary disassembly is crucial for understanding program behavior in reverse engineering. However, existing disassembly tools struggle to accurately identify function boundaries when binary files are stripped, especially on reduced instruction set computer (RISC) architectures like ARM and MIPS. Additionally, disassembly tools often lack support for newer versions or less used architectures. In this paper, we propose UDA, a unified pre-training disassembly framework with multi-architecture support. Unified pre-training on both machine code and assembly code allows assembly code to help the semantic learning of machine code, thereby addressing the inherent semantic deficiencies of machine code. By extracting enhanced semantic features from machine code, UDA can not only more accurately identify function boundaries but also recover assembly code without relying on existing disassembly tools. UDA has been rigorously evaluated against both regular and obfuscated binary files. The results show that UDA improves the F1 score for function boundary detection by 0.9%, 7.9%, and 8.1% on x86, ARM, and MIPS architectures, respectively, compared to the state-of-the-art disassemblers. Additionally, it demonstrates strong robustness against unseen obfuscated binaries. Xunzhi Jiang, Shen Wang 0004, Yuxin Gong, Tingyue Yu, Xiangzhan Yu |
IEEE Internet Things J. | 6 |
| 2025 | Anomaly Detection in Industrial Control Systems Based on Cross-Domain Representation LearningabstractIndustrial control systems (ICSs) are widely used in industry, and their security and stability are very important. Once the ICS is attacked, it may cause serious damage. Therefore, it is very important to detect anomalies in ICSs. ICS can monitor and manage physical devices remotely using communication networks. The existing anomaly detection approaches mainly focus on analyzing the security of network traffic or sensor data. However, the behaviors of different domains (e.g., network traffic and sensor physical status) of ICSs are correlated, so it is difficult to comprehensively identify anomalies by analyzing only a single domain. In this article, an anomaly detection approach based on cross-domain representation learning in ICSs is proposed, which can learn the joint features of multi-domain behaviors and detect anomalies within different domains. After constructing a cross-domain graph that can represent the behaviors of multiple domains in ICSs, our approach can learn the joint features of them by leveraging graph neural networks. Since anomalies behave differently in different domains, we leverage a multi-task learning approach to identify anomalies in different domains separately and perform joint training. The experimental results show that the performance of our approach is better than existing approaches for identifying anomalies in ICSs. Dongyang Zhan, Wenqi Zhang 0006, Xiangzhan Yu, Hongli Zhang 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | HAformer: Semantic fusion of hex machine code and assembly code for cross-architecture binary vulnerability detection
Xunzhi Jiang, Shen Wang 0004, Yuxin Gong, Tingyue Yu, Xiangzhan Yu |
Comput. Secur. | 6 |
| 2024 | An effective deep learning adversarial defense method based on spatial structural constraints in embedding space
Junzhong Miao, Xiangzhan Yu, Zhichao Hu, Yanru Song 0002, Likun Liu |
Pattern Recognit. Lett. | 2 |
| 2024 | pFind: Privacy-preserving lost object finding in vehicular crowdsensing
Yinggang Sun, Haining Yu, Yizheng Yang, Xiangzhan Yu |
World Wide Web (WWW) | 5 |
| 2023 | BEATs: Audio Pre-Training with Acoustic TokenizersabstractWe introduce a self-supervised learning (SSL) framework BEATs for general audio representation pre-training, where we optimize an acoustic tokenizer and an audio SSL model by iterations. Unlike the previous audio SSL models that employ reconstruction loss for pre-training, our audio SSL model is trained with the discrete label prediction task, where the labels are generated by a semantic-rich acoustic tokenizer. We propose an iterative pipeline to jointly optimize the tokenizer and the pre-trained model, aiming to abstract high-level semantics and discard the redundant details for audio. The experimental results demonstrate our acoustic tokenizers can generate discrete labels with rich audio semantics and our audio SSL models achieve state-of-the-art (SOTA) results across various audio classification benchmarks, even outperforming previous models that use more training data and model parameters significantly. Specifically, we set a new SOTA mAP 50.6% on AudioSet-2M without using any external data, and 98.1% accuracy on ESC-50. The code and pre-trained models are available at https://aka.ms/beats. Sanyuan Chen, Yu Wu 0012, Chengyi Wang 0002, Shujie Liu 0001, Daniel Tompkins, Zhuo Chen 0006, Wanxiang Che, Xiangzhan Yu, Furu Wei |
ICML | 8 |
| 2023 | Global Wasserstein Margin maximization for boosting generalization in adversarial training
Tingyue Yu, Shen Wang 0004, Xiangzhan Yu |
Appl. Intell. | 3 |
| 2023 | Securing Operating Systems Through Fine-Grained Kernel Access Limitation for IoT SystemsabstractWith the development of Internet of Things (IoT), it is gaining a lot of attention. It is important to secure the embedded systems with low overhead. The Linux Seccomp is widely used by developers to secure the kernels by blocking the access of unused syscalls, which introduces less overhead. However, there are no systematic Seccomp configuration approaches for IoT applications without the help of developers. In addition, the existing Seccomp configuration approaches are coarse-grained, which cannot analyze and limit the syscall arguments. In this article, a novel static dependent syscall analysis approach for embedded applications is proposed, which can obtain all of the possible dependent syscalls and the corresponding arguments of the target applications. So, a fine-grained kernel access limitation can be performed for the IoT applications. To this end, the mappings between dynamic library APIs and syscalls according with their arguments are built, by analyzing the control flow graphs and the data dependency relationships of the dynamic libraries. To the best of our knowledge, this is the first work to generate the fine-grained Seccomp profile for embedded applications. Dongyang Zhan, Zhaofeng Yu, Xiangzhan Yu, Hongli Zhang 0001, Likun Liu |
IEEE Internet Things J. | 3 |
| 2023 | An Adversarial Robust Behavior Sequence Anomaly Detection Approach Based on Critical Behavior Unit LearningabstractSequential deep learning models (e.g., RNN and LSTM) can learn the sequence features of software behaviors, such as API or syscall sequences. However, recent studies have shown that these deep learning-based approaches are vulnerable to adversarial samples. Attackers can use adversarial samples to change the sequential characteristics of behavior sequences and mislead malware classifiers. In this paper, an adversarial robustness anomaly detection method based on the analysis of behavior units is proposed to overcome this problem. We extract related behaviors that usually perform a behavior intention as a behavior unit, which contains the representative semantic information of local behaviors and can be used to improve the robustness of behavior analysis. By learning the overall semantics of each behavior unit and the contextual relationships among behavior units based on a multilevel deep learning model, our approach can mitigate perturbation attacks that target local and large-scale behaviors. In addition, our approach can be applied to both low-level and high-level behavior logs (e.g., API and syscall logs). The experimental results show that our approach outperforms all the compared methods, which indicates that our approach has better performance against obfuscation attacks. Dongyang Zhan, Kai Tan 0007, Xiangzhan Yu, Hongli Zhang 0001 |
IEEE Trans. Computers | 4 |
| 2023 | pSafety: Privacy-Preserving Safety Monitoring in Online Ride Hailing ServicesabstractOnline Ride Hailing (ORH) services gain remarkable development in the past decade, which enable riders and drivers to establish optimized rides via mobile device. To guarantee user safety, ORH service providers often monitor the ride trajectory and report the abnormal behavior once a trajectory deviation occurs. Along with the advantage of safety monitoring raises some vital privacy concerns on user location information leakage. In this paper, we propose a privacy-preserving safety monitoring scheme for ORH services, called pSafety. It enables an ORH service provider to detect user’s trajectory deviation without learning anything about users’ locations. In pSafety, we propose two secure trajectory similarity computation algorithms by using somewhat homomorphic encryption, which are used to plan an agreed path and measure trajectory deviation, respectively. Furthermore, we also design a ciphertext compression algorithm and a secure comparison protocol to improve efficiency. Theoretical analysis and experimental evaluations show that pSafety is secure, accurate and efficient. Haining Yu, Hongli Zhang 0001, Xiaohua Jia, Xiangzhan Yu |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2023 | APFed: Anti-Poisoning Attacks in Privacy-Preserving Heterogeneous Federated LearningabstractFederated learning (FL) is an emerging paradigm of privacy-preserving distributed machine learning that effectively deals with the privacy leakage problem by utilizing cryptographic primitives. However, how to prevent poisoning attacks in distributed situations has recently become a major FL concern. Indeed, an adversary can manipulate multiple edge nodes and submit malicious gradients to disturb the global model’s availability. Currently, most existing works rely on an Independently Identical Distribution (IID) situation and identify malicious gradients using plaintext. However, we demonstrates that current works cannot handle the data heterogeneity scenario challenges and that publishing unencrypted gradients imposes significant privacy leakage problems. Therefore, we develop APFed, a layered privacy-preserving defense mechanism that significantly mitigates the effects of poisoning attacks in data heterogeneity scenarios. Specifically, we exploit HE as the underlying technique and employ the median coordinate as the benchmark. Subsequently, we propose a secure cosine similarity scheme to identify poisonous gradients, and we innovatively use clustering as part of the defense mechanism and develop a hierarchical aggregation that enhances our scheme’s robustness in IID and non-IID scenarios. Extensive evaluations on two benchmark datasets demonstrate that APFed outperforms existing defense strategies while reducing the communication overhead by replacing the expensive remote communication method with inexpensive intra-cluster communication. Haining Yu, Xiaohua Jia, Xiangzhan Yu |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | Higher-Order Community Detection: On Information Degeneration and Its EliminationabstractCommunity detection aims to identify the cohesive vertex sets in a network. It is widely used in many domains, e.g., World Wide Web, online social networks, and communication networks. Many clustering models are proposed in the literature. However, most of them are designed directly on the original structure of a network, they usually achieve low accuracy in practice, since real-world networks are presenting fuzzy community structures. Recently, higher-order network units are introduced to community detection, these models typically define a higher-order hypergraph, where communities are extracted. Although the higher-order models are effective in terms of accuracy, many essential edges are completely eliminated or trivialized in the hypergraph. To address the problem, we propose a novel connectivity pattern with a mixture of standard edges and higher-order connections, whereby we define biased personalized PageRank diffusion for local community detection and develop a local approach to compute the PageRank vectors. Moreover, we present a higher-order seeding strategy to derive the starting seeds. Extensive experiments demonstrate that the proposed framework largely outperforms the approaches in the state of the art in terms of accuracy. Hongli Zhang 0001, Xiangzhan Yu |
IEEE/ACM Trans. Netw. | 3 |
| 2023 | Shrinking the Kernel Attack Surface Through Static and Dynamic Syscall LimitationabstractLinux Seccomp is widely used by the program developers and the system maintainers to secure the operating systems, which can block unused syscalls for different applications and containers to shrink the attack surface of the operating systems. However, it is difficult to configure the whitelist of a container or application without the help of program developers. Docker containers block about only 50 syscalls by default, and lots of unblocked useless syscalls introduce a big kernel attack surface. To obtain the dependent syscalls, dynamic tracking is a straight-forward approach but it cannot get the full syscall list. Static analysis can construct an over-approximated syscall list, but the list contains many false positives. In this paper, a systematic dependent syscall analysis approach, sysverify, is proposed by combining static analysis and dynamic verification together to shrink the kernel attack surface. The semantic gap between the binary executables and syscalls is bridged by analyzing the binary and the source code, which builds the mapping between the library APIs and syscalls systematically. To further reduce the attack surface at best effort, we propose a dynamic verification approach to intercept and analyze the security of the invocations of indirect-call-related or rarely invoked syscalls with low overhead. Dongyang Zhan, Zhaofeng Yu, Xiangzhan Yu, Hongli Zhang 0001 |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | ErrHunter: Detecting Error-Handling Bugs in the Linux Kernel Through Systematic Static AnalysisabstractError handling is essential for operating systems, thus, there are many bugs in error-handling code, which could result in serious consequences. In this paper, we revisit the problem of error miss-handling bugs and analyze the root cause of the most common ones in the Linux kernel. Based on the analysis, we propose a systematic static taint-analysis-based approach, ErrHunter, to detect multiple kinds of error miss-handling bugs in the Linux kernel. An automated critical variable identification approach is proposed to identify critical variables in the error-handling paths. A static cross-control-flow taint analysis approach is proposed to construct critical-variable control flow graphs (CCFGs), which describe the processing of critical variables in separate control flows. Based on the CCFGs, ErrHunter can target the root cause of the most common error miss-handling bugs and detect the bugs in a systematic way. ErrHunter is designed for kernel bug detection, so it can handle many specific features of the Linux kernel, such as memory management mechanisms, etc. Dongyang Zhan, Xiangzhan Yu, Hongli Zhang 0001 |
IEEE Trans. Software Eng. | 2 |
| 2022 | Unispeech-Sat: Universal Speech Representation Learning With Speaker Aware Pre-TrainingabstractSelf-supervised learning (SSL) is a long-standing goal for speech processing, since it utilizes large-scale unlabeled data and avoids extensive human labeling. Recent years have witnessed great successes in applying self-supervised learning in speech recognition, while limited exploration was attempted in applying SSL for modeling speaker characteristics. In this paper, we aim to improve the existing SSL framework for speaker representation learning. Two methods are introduced for enhancing the unsupervised speaker information extraction. First, we apply multi-task learning to the current SSL framework, where we integrate utterance-wise contrastive loss with the SSL objective function. Second, for better speaker discrimination, we propose an utterance mixing strategy for data augmentation, where additional overlapped utterances are created unsupervisely and incorporated during training. We integrate the proposed methods into the HuBERT framework. Experiment results on the SUPERB benchmark show that the proposed system achieves state-of-the-art performance in universal representation learning, especially for speaker identification oriented tasks. An ablation study is performed verifying the efficacy of each proposed method. Finally, we scale up the training dataset to 94 thousand hours of public audio data and achieve further performance improvement in all SUPERB tasks. Sanyuan Chen, Yu Wu 0012, Chengyi Wang 0002, Zhengyang Chen, Zhuo Chen 0006, Shujie Liu 0001, Jian Wu 0027, Yao Qian, Furu Wei, Jinyu Li 0001, Xiangzhan Yu |
ICASSP | 11 |
| 2022 | Why does Self-Supervised Learning for Speech Recognition Benefit Speaker Recognition?abstractRecently, self-supervised learning (SSL) has demonstrated strong performance in speaker recognition, even if the pretraining objective is designed for speech recognition.In this paper, we study which factor leads to the success of selfsupervised learning on speaker-related tasks, e.g.speaker verification (SV), through a series of carefully designed experiments.Our empirical results on the Voxceleb-1 dataset suggest that the benefit of SSL to SV task is from a combination of mask speech prediction loss, data scale, and model size, while the SSL quantizer has a minor impact.We further employ the integrated gradients attribution method and loss landscape visualization to understand the effectiveness of self-supervised learning for speaker recognition performance. Sanyuan Chen, Yu Wu 0012, Chengyi Wang 0002, Shujie Liu 0001, Zhuo Chen 0006, Gang Liu 0001, Jinyu Li 0001, Jian Wu 0027, Xiangzhan Yu, Furu Wei |
INTERSPEECH | 10 |
| 2022 | Graph clustering using triangle-aware measures in large networks
Xiangzhan Yu, Hongli Zhang 0001 |
Inf. Sci. | 2 |
| 2021 | Don't Shoot Butterfly with Rifles: Multi-Channel Continuous Speech Separation with Early Exit TransformerabstractWith its strong modeling capacity that comes from a multi-head and multi-layer structure, Transformer is a very powerful model for learning a sequential representation and has been successfully applied to speech separation recently. However, multi-channel speech separation sometimes does not necessarily need such a heavy structure for all time frames especially when the cross-talker challenge happens only occasionally. For example, in conversation scenarios, most regions contain only a single active speaker, where the separation task downgrades to a single speaker enhancement problem. It turns out that using a very deep network structure for dealing with signals with a low overlap ratio not only negatively affects the inference efficiency but also hurts the separation performance. To deal with this problem, we propose an early exit mechanism, which enables the Transformer model to handle different cases with adaptive depth. Experimental results indicate that not only does the early exit mechanism accelerate the inference, but it also improves the accuracy. Sanyuan Chen, Yu Wu 0012, Zhuo Chen 0006, Takuya Yoshioka, Shujie Liu 0001, Jinyu Li 0001, Xiangzhan Yu |
ICASSP | 7 |
| 2021 | Ultra Fast Speech Separation Model with Teacher Student LearningabstractTransformer has been successfully applied to speech separation recently with its strong long-dependency modeling capacity using a self-attention mechanism. However, Transformer tends to have heavy run-time costs due to the deep encoder layers, which hinders its deployment on edge devices. A small Transformer model with fewer encoder layers is preferred for computational efficiency, but it is prone to performance degradation. In this paper, an ultra fast speech separation Transformer model is proposed to achieve both better performance and efficiency with teacher student learning (T-S learning). We introduce layer-wise T-S learning and objective shifting mechanisms to guide the small student model to learn intermediate representations from the large teacher model. Compared with the small Transformer model trained from scratch, the proposed T-S learning method reduces the word error rate (WER) by more than 5% for both multi-channel and single-channel speech separation on LibriCSS dataset. Utilizing more unlabeled speech data, our ultra fast speech separation models achieve more than 10% relative WER reduction. Sanyuan Chen, Yu Wu 0012, Zhuo Chen 0006, Jian Wu 0027, Takuya Yoshioka, Shujie Liu 0001, Jinyu Li 0001, Xiangzhan Yu |
Interspeech | 8 |
| 2021 | Overlapping community detection by constrained personalized PageRank
Xiangzhan Yu, Hongli Zhang 0001 |
Expert Syst. Appl. | 2 |
| 2021 | PGRide: Privacy-Preserving Group Ridesharing Matching in Online Ride Hailing ServicesabstractAn online ride hailing (ORH) service creates a typical supply-and-demand two-sided market, which enables riders and drivers to establish optimized rides conveniently via mobile applications. Group ridesharing is a novel form of ridesharing, which allows a group of riders to share a vehicle that holds the minimum aggregate distance to the whole group. Accompanied by the advantage of ORH services, there comes some vital privacy concerns. In this article, we propose a privacy-preserving online group ridesharing matching scheme for ORH services, called PGRide. PGRide can select the nearest driver to serve a group of riders, without leaking the location privacy of both riders and drivers. In PGRide, we propose an encrypted aggregate distance computation approach by using somewhat homomorphic encryption with ciphertexts packing, which efficiently computes the aggregate distances from a group of riders to large-scale dynamic drivers in encrypted form. Meanwhile, we design a secure minimum selection protocol by using ciphertexts packing and blinding, which efficiently finds the minimum element from a set of encrypted integers without leaking any actual element value. Theoretical analysis and performance evaluations prove that PGRide is secure, accurate, and efficient. Haining Yu, Hongli Zhang 0001, Xiangzhan Yu, Xiaojiang Du, Mohsen Guizani |
IEEE Internet Things J. | 3 |
| 2021 | Nowhere to Hide: A Novel Private Protocol Identification AlgorithmabstractIn recent years, with the rapid development of mobile Internet and 5G technology, great changes have been brought to our lives, and human beings have stepped into the era of big data. These new features and techniques in 5G support many different types of mobile applications for users, which makes network security extremely challenging. Among them, more and more applications involve users’ private data, such as location information, financial information, and biological information. In order to prevent users’ privacy disclosure, most applications choose to use private protocols. However, such private protocols also provide a means for malware and malicious applications to steal users’ privacy and confidential data. From a more secure point of view, we need to provide a way for users to know how many private protocols are running on their mobile phones and distinguish which are authorized applications and which are not. Therefore, the analysis and identification of private protocols have become a hot topic in current research. How to extract the characteristics of network protocol effectively and identify the private protocol accurately becomes the most important part of this research. In this paper, we combine genetic algorithm and association rule algorithm and then propose a set of feature extraction algorithm and protocol recognition algorithm for unknown protocols. The experimental analysis based on the actual data shows that these methods can effectively solve the problems of feature extraction and recognition for unknown protocols and can greatly improve the accuracy of private protocol recognition. Xiangzhan Yu, Zechao Liu |
Secur. Commun. Networks | 2 |
| 2021 | PSRide: Privacy-Preserving Shared Ride Matching for Online Ride Hailing SystemsabstractOnline Ride Hailing (ORH) has extensively made our trip more convenient. With mobile devices, riders can request taxis through ORH systems in a short time. However, to enjoy ORH services, users need to submit their location information to ORH systems, which raises serious privacy concerns. In this paper, we study the privacy leakage of online ridesharing matching, a more complex and economy ORH service that allows riders to share rides with others, and propose a privacy-preserving shared ride matching scheme, called PSRide. PSRide can find the taxi with the minimum additional travel time to serve a new rider based on its existing schedule, while protecting the location privacy of both riders and taxis. In PSRide, we propose a zone-based minimum road travel time estimation approach and a secure comparison protocol to efficiently optimize the schedules of taxis for a new rider over encrypted data. We implement PSRide and analyze it thoroughly. Theoretical analysis and experimental evaluations show that PSRide is secure and efficient for ORH systems. Haining Yu, Xiaohua Jia, Hongli Zhang 0001, Xiangzhan Yu, Jiangang Shu |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2021 | A New Approach Customizable Distributed Network Service Discovery SystemabstractComputer systems and applications on the internet provide services to outsiders and, at the same time, the vulnerabilities may be exploited by attackers and leak some sensitive private information. To collect and monitor the service information provided by the network environment such as IoT (Internet of Things), vehicular networks, cloud computing, and cloud storage, it is particularly important that a system can provide faster service discovery for discovering and identifying specific network services. The current service discovery systems mainly use port scanning technology, including Nmap, Zmap, and Masscan. However, these technologies hard code the service features and only support common services so that cannot cope with real‐time updates and changing network services. To solve the above problems, this paper proposed a customizable distributed network service discovery system based on stateless scanning technology of Masscan and proposed a customizable interactive pattern set syntax. The system used random destination address technologies to scan for Ipv4 address allocation and used a distributed deployment scheme. Experimental results show that the system has high scanning speed and has high adaptability to new services and special services. Xiangzhan Yu, Zhichao Hu, Yi Xin 0002 |
Wirel. Commun. Mob. Comput. | 1 |
| 2020 | Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less ForgettingabstractDeep pretrained language models have achieved great success in the way of pretraining first and then fine-tuning.But such a sequential transfer learning paradigm often confronts the catastrophic forgetting problem and leads to sub-optimal performance.To fine-tune with less forgetting, we propose a recall and learn mechanism, which adopts the idea of multi-task learning and jointly learns pretraining tasks and downstream tasks.Specifically, we propose a Pretraining Simulation mechanism to recall the knowledge from pretraining tasks without data, and an Objective Shifting mechanism to focus the learning on downstream tasks gradually.Experiments show that our method achieves state-of-the-art performance on the GLUE benchmark.Our method also enables BERT-base to achieve better performance than directly fine-tuning of BERT-large.Further, we provide the open-source RECADAM optimizer, which integrates the proposed mechanisms into Adam optimizer, to facility the NLP community. Sanyuan Chen, Yutai Hou, Yiming Cui 0001, Wanxiang Che, Ting Liu 0001, Xiangzhan Yu |
EMNLP (1) | 6 |
| 2020 | Uncovering overlapping community structure in static and dynamic networks
Xiangzhan Yu, Hongli Zhang 0001 |
Knowl. Based Syst. | 2 |
| 2020 | Hail the Closest Driver on Roads: Privacy-Preserving Ride Matching in Online Ride Hailing ServicesabstractOnline ride hailing (ORH) services enable a rider to request a driver to take him wherever he wants through a smartphone app on short notice. To use ORH services, users have to submit their ride information to the ORH service provider to make ride matching, such as pick-up/drop-off location. However, the submission of ride information may lead to the leakages of users’ privacy. In this paper, we focus on the issue of protecting the location information of both riders and drivers during ride matching and propose a privacy-preserving online ride matching scheme, called pRMatch. It enables an ORH service provider to find the closest available driver for an incoming rider over a city-scale road network, while protecting the location privacy of both riders and drivers against the ORH service provider and other unauthorized participants. In pRMatch, we compute the shortest road distance over encrypted data by using road network embedding and partially homomorphic encryption and further efficiently compare encrypted distances by using ciphertext packing and shuffling. The theoretical analysis and experimental results demonstrate that pRMatch is accurate and efficient, yet preserving users’ location privacy. Haining Yu, Hongli Zhang 0001, Xiangzhan Yu |
Secur. Commun. Networks | 3 |
| 2019 | No Way to Evade: Detecting Multi-Path Routing Attacks for NIDSabstractIn order to protect intranet security, the enterprises or organizations usually deploy one or multiple NIDS at ingress points. Each works independently and monitors the complete TCP flow. That said, a malicious signature to be detected can only be obtained from a TCP flow. Drawing on the feature, an attacker can split malicious signature into multiple substrings and transfer them in different flows to evade detection, which is named multi-path routing attack. In particular, the emerging new technology Multi-Path TCP (MPTCP) offers a hotbed for such attacks. To monitor multi-path routing attacks, this literature proposed a distributed asynchronous NIDS detection model (DANDM) which consists of three algorithms. In this model, each NIDS scans its own received data packets independently and the adjacent contents between two data packets with consecutive sequence numbers. For the latter, all NIDS scans cooperatively through broadcast state information. To demonstrate the validity of our model, we take attack density and number of segmented signatures as parameters to compare with Ma's algorithm.The results show that the performance of our DANDM is significantly better than that of Ma's, especially in the case of large number of segmented signatures. Likun Liu, Hongli Zhang 0001, Xiangzhan Yu |
GLOBECOM | 4 |
| 2018 | WeChat traffic classification using machine learning algorithms and comparative analysis of datasetsabstractIn this research paper, we present the first classification study to classify WeChat application service flow traffic (text messages, picture messages, audio call and video call traffic) classification and secondly to find out the effectiveness of big dataset and small dataset as well as to find out effective machine learning classifiers. We firstly capture WeChat traffic and then extract 44 features then we combine capture traffic to make full instance of dataset. Then we make reduce instances of dataset from the full instance of dataset to show the effectiveness of large dataset and small dataset. Then we execute well known machine learning classifiers. Using statistical test, we use Wilcoxon and Friedman statistical test for the datasets and ML classifiers to find more deeply its effectiveness. Experimental results show that reduce instance dataset show high accuracy result compared to full instance and C4.5 classifier perform effectively as compared to other classifiers. Muhammad Shafiq 0003, Xiangzhan Yu, Asif Ali Laghari |
Int. J. Inf. Comput. Secur. | 2 |
| 2018 | A machine learning approach for feature selection traffic classification using security analysis
Muhammad Shafiq 0003, Xiangzhan Yu, Ali Kashif Bashir, Hassan Nazeer Chaudhry |
J. Supercomput. | 2 |
| 2018 | An Efficient Security System for Mobile Data MonitoringabstractDuring the last decade, rapid development of mobile devices and applications has produced a large number of mobile data which hide numerous cyber‐attacks. To monitor the mobile data and detect the attacks, NIDS/NIPS plays important role for ISP and enterprise, but now it still faces two challenges, high performance for super large patterns and detection of the latest attacks. High performance is dominated by Deep Packet Inspection (DPI) mechanism, which is the core of security devices. A new TTL attack is just put forward to escape detecting, such that the adversary inserts packet with short TTL to escape from NIDS/NIPS. To address the above‐mentioned problems, in this paper, we design a security system to handle the two aspects. For efficient DPI, a new two‐step partition of pattern set is demonstrated and discussed, which includes first set‐partition and second set‐partition. For resisting TTL attacks, we set reasonable TTL threshold and patch TCP protocol stack to detect the attack. Compared with recent produced algorithm, our experiments show better performance and the throughput increased 27% when the number of patterns is 106. Moreover, the success rate of detection is 100%, and while attack intensity increased, the throughput decreased. Likun Liu, Hongli Zhang 0001, Xiangzhan Yu, Yi Xin 0002, Muhammad Shafiq 0003, Mengmeng Ge 0003 |
Wirel. Commun. Mob. Comput. | 3 |
| 2016 | LPPS: Location privacy protection for smartphonesabstractLocation-based service (LBS) is useful for many applications. However, LBS has raised serious concerns about users' location privacy. Utilizing the computation and storage capacity of smart phones, we propose a novel system architecture, called Location Privacy Protection for Smartphone (LPPS), to provide a privacy-preserving top-k query. LPPS does not rely on a trust third party (TTP), nor does it requires LBS servers to change their business model. The main idea of LPPS is to rank the Points of Interests (POIs) on the client side of the application using a small amount of metadata and then to make a request to the LBS server for real-time and detailed information about the POIs. Based on LPPS, we propose a novel metric called location indistinguishability to evaluate the privacy level of users in the proposed scheme. Then, we propose two dummy-POI selection algorithms to generate a superset of the actual top-k POIs when the query cannot meet the privacy requirement. Our experimental results demonstrate the validity and practicality of the proposed schemes. Hongli Zhang 0001, Zhikai Xu, Xiangzhan Yu, Xiaojiang Du |
ICC | 3 |
| 2015 | Continuous resource allocation in cloud computingabstractWith the advent of cloud computing, more and more enterprises and individuals are motivated to outsource their local complex applications into the cloud for its great flexibility and economic savings. For cloud server, however, the problem of how to optimize scheduling cloud resources to maximize the resource utilization and incorporate energy optimization is not trivial. In this paper, we present a continuous resource allocation strategy to solve this issue. Specifically, we first map all the remaining resources of the cloud into a ZBtree. Based on this, we propose an efficient resource allocation strategy for processing cloud resource requests, which adopts minimal domination matching as the greedy strategy to navigate the tradeoff space. Moreover, to reduce the rate of application migration, we propose a continuous monitoring strategy which adopts a resource reserve scheme to reduce the probability of application migration to the maximum extent. The extensive experiments further demonstrate the validity of the proposed mechanism with low computation overhead. Hongli Zhang 0001, Xiangzhan Yu, Junwu Guo |
ICC | 3 |
| 2015 | Audit meets game theory: Verifying reliable execution of SLA for compute-intensive program in cloudabstractCloud computing provides convenient on-demand computing resources which enables users to outsource their computational tasks to the cloud. However, users lose direct control of tasks, which inevitably causes new challenges: users have no way to audit whether the cloud server provides services along the service level agreement (SLA). To this end, we present an audit model based on cloud service broker (CSB) that can verify the reliable execution of SLA for compute-intensive program in cloud. Specifically, we propose a program structure mapping model called concept tree, which transforms the target application into a tree structure. Then, we split it into several independent computable units. Based on this, CSB can fine-grained audit the task execution process in cloud by comparing the running time and the similarity of the computable units. We also adopt game theory to motivate the credible implementation of SLA, in addition to defend against the semi-honest cloud model. We implement our scheme on Hadoop and the extensive experiments demonstrate that our scheme has the high prediction accuracy. Hongli Zhang 0001, Xiangzhan Yu, Junwu Guo |
ICC | 3 |
| 2013 | Practical and privacy-assured data indexes for outsourced cloud dataabstractCloud computing allows individuals and organizations outsource their data to cloud server due to the flexibility and cost savings. However, data privacy is a major concern that hampers the wide adoption of cloud services. Data encryption ensures data content confidentiality and fine-grained data access control prevents unauthorized user from accessing data. An unauthorized user may still be able to infer privacy information from encrypted data by using indexing techniques. In this paper, we investigate the problem of sensitive information leakage caused by orthogonal use of these two kinds of techniques. Based on that, we propose “core attribute”-aware techniques that can ensure privacy of outsourced data. The techniques focus on confidential attribute set of outsourced data. We adopt k-anonymity technique for the attribute indexes to prevent user from inferring privacy from unauthorized data. We formally prove the privacy-preserving guarantee of the proposed mechanism. Our extensive experiments demonstrate the practicality of the proposed mechanism, which has low computation and communication overhead. Hongli Zhang 0001, Xiaojiang Du, Xiangzhan Yu |
GLOBECOM | 5 |
| 2013 | Prometheus: Privacy-aware data retrieval on hybrid cloudabstractWith the advent of cloud computing, data owner is motivated to outsource their data to the cloud platform for great flexibility and economic savings. However, the development is hampered by data privacy concerns: Data owner may have privacy data and the data cannot be outsourced to cloud directly. Previous solutions mainly use encryption. However, encryption causes a lot of inconveniences and large overheads for other data operations, such as search and query. To address the challenge, we adopt hybrid cloud. In this paper, we present a suit of novel techniques for efficient privacy-aware data retrieval. The basic idea is to split data, keeping sensitive data in trusted private cloud while moving insensitive data to public cloud. However, privacy-aware data retrieval on hybrid cloud is not supported by current frameworks. Data owners have to split data manually. Our system, called Prometheus, adopts the popular MapReduce framework, and uses data partition strategy independent to specific applications. Prometheus can automatically separate sensitive information from public data. We formally prove the privacy-preserving feature of Prometheus. We also show that our scheme can defend against the malicious cloud model, in addition to the semi-honest cloud model. We implement Prometheus on Hadoop and evaluate its performance using real data set on a large-scale cloud test-bed. Our extensive experiments demonstrate the validity and practicality of the proposed scheme. Hongli Zhang 0001, Xiaojiang Du, Xiangzhan Yu |
INFOCOM | 5 |
| 2012 | An efficient and sustainable self-healing protocol for Unattended Wireless Sensor NetworksabstractDue to the unattended operation nature, nodes in Unattended Wireless Sensor Networks (UWSNs) are susceptible to physical attacks. Once a sensor is compromised, the adversary will be able to learn all its secrets. While some previous works tried to address the node self-healing issue in UWSNs, little effort has been devoted to ensure the sustainability of node self-healing. In this paper, we present a novel sustainable node self-healing protocol for UWSNs. We generate unpredictable random data for key update and thus the node self-healing capability doesn't decrease when the number of attack rounds increases. We show both analytically and through simulation experiments that our protocol provides efficient and sustainable node self-healing capabilities with small overheads. Hongli Zhang 0001, Binxing Fang, Xiaojiang Du, Haining Yu, Xiangzhan Yu |
GLOBECOM | 6 |
| 2012 | Feature selection for optimizing traffic classification
Hongli Zhang 0001, Mahmoud T. Qassrawi, Yu Zhang 0036, Xiangzhan Yu |
Comput. Commun. | 5 |
| 2011 | Towards Efficient Anonymous Communications in Sensor NetworksabstractAnonymous communication is a challenging task in resource constrained wireless sensor networks (WSN). However, anonymity is important for many sensor networks, in which we want to conceal the location and identify of important nodes (such as source nodes and base stations) from attackers. Existing WSN anonymous protocols either cannot achieve complete anonymity, or have large computation and/or storage overheads. In this paper, we present an efficient anonymous communication protocol for sensor networks. Our protocol can achieve sender/source anonymity, communication -relationship anonymity, and the base station anonymity simultaneously, while having small overheads on computation, storage and communication. Hongli Zhang 0001, Binxing Fang, Xiaojiang Du, Lihua Yin, Xiangzhan Yu |
GLOBECOM | 6 |
| 2010 | A Pseudo-Random Number Generator Based on LZSSabstractA pseudo-random sequence generator (PRNG), L12RC4, inspired by the LZSS compression algorithm and RC4 stream cipher, was presented and implemented. The result of the NIST and Diehard test suite indicate that the L12RC4 is a good PRNG, and so it seems to be sound and may be suitable for use in some cryptographic applications. We also found that the probability distribution of the index value frequency is associated with the compression pass and INDEX_BIT_COUNT value. As for one pass mode, the greater INDEX_BIT_COUNT value, the more uniformly distributed, and the double pass mode has better uniformity than the one pass mode. Wei-ling Chang, Binxing Fang, Xiao-chun Yun, Xiangzhan Yu |
DCC | 5 |