Zhenyan Liu

dblp:162/0920 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 4 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1
YearPublicationVenuePosition
2025 VeriTrac: Verifiable and traceable cross-silo federated learning
Yanxin Xu, Hua Zhang 0001, Zhenyan Liu, Fei Gao 0001
Future Gener. Comput. Syst.3
2025 CDDA: Privacy-preserving blockchain-based cross-domain dynamic authentication scheme for Healthcare 5.0
abstract
Healthcare 5.0 delivers high-quality healthcare by promoting collaboration among multiple trusted domains to effectively integrate resources. In this process, cross-domain authentication is a critical issue that must be addressed to ensure secure communication and resource sharing. Common unified authentication frameworks rely on consistent and coordinated authentication policies across different domains, which may lead to challenges in accommodating diverse policies and increase the risk of privacy breaches. To address these issues, this paper proposes a privacy-preserving blockchain-based cross-domain dynamic authentication scheme, referred to as CDDA, for Healthcare 5.0. CDDA enables cross-domain authentication across various healthcare domains, each with its own distinct authentication policies. To protect domain privacy, it anonymizes identity credentials and policy structures. Additionally, it implements a blockchain-based decentralized approach integrated with threshold verification mechanisms to enhance security. Finally, we conduct a security analysis to demonstrate that CDDA satisfies the essential security and privacy criteria. We further validate its effectiveness and efficiency through thorough experimental evaluations. The results indicate that the most complex phase of the cross-domain resource request operation takes an average of 1.38 seconds, which is acceptable given its significant contribution to enhancing cross-domain privacy and flexibility.
Zhenyan Liu, Ning Shi
Inf. Sci.2
2024 Classifier Clustering and Feature Alignment for Federated Learning under Distributed Concept Drift
abstract
Data heterogeneity is one of the key challenges in federated learning, and many efforts have been devoted to tackling this problem. However, distributed concept drift with data heterogeneity, where clients may additionally experience different concept drifts, is a largely unexplored area. In this work, we focus on real drift, where the conditional distribution $P(\mathcal{Y}|\mathcal{X})$ changes. We first study how distributed concept drift affects the model training and find that local classifier plays a critical role in drift adaptation. Moreover, to address data heterogeneity, we study the feature alignment under distributed concept drift, and find two factors that are crucial for feature alignment: the conditional distribution $P(\mathcal{Y}|\mathcal{X})$ and the degree of data heterogeneity. Motivated by the above findings, we propose FedCCFA, a federated learning framework with classifier clustering and feature alignment. To enhance collaboration under distributed concept drift, FedCCFA clusters local classifiers at class-level and generates clustered feature anchors according to the clustering results. Assisted by these anchors, FedCCFA adaptively aligns clients' feature spaces based on the entropy of label distribution $P(\mathcal{Y})$, alleviating the inconsistency in feature space. Our results demonstrate that FedCCFA significantly outperforms existing methods under various concept drift settings. Code is available at https://github.com/Chen-Junbao/FedCCFA.
Junbao Chen, Jingfeng Xue, Yong Wang 0010, Zhenyan Liu, Lu Huang 0002
NeurIPS4
2023 WHGDroid: Effective android malware detection based on weighted heterogeneous graph
Lu Huang 0002, Jingfeng Xue, Yong Wang 0010, Zhenyan Liu, Junbao Chen, Zixiao Kong
J. Inf. Secur. Appl.4
2021 Malicious Encryption Traffic Detection Based on NLP
abstract
The development of Internet and network applications has brought the development of encrypted communication technology. But on this basis, malicious traffic also uses encryption to avoid traditional security protection and detection. Traditional security protection and detection methods cannot accurately detect encrypted malicious traffic. In recent years, the rise of artificial intelligence allows us to use machine learning and deep learning methods to detect encrypted malicious traffic without decryption, and the detection results are very accurate. At present, the research on malicious encrypted traffic detection mainly focuses on the characteristics’ analysis of encrypted traffic and the selection of machine learning algorithms. In this paper, a method combining natural language processing and machine learning is proposed; that is, a detection method based on TF-IDF is proposed to build a detection model. In the process of data preprocessing, this method introduces the natural language processing method, namely, the TF-IDF model, to extract data information, obtain the importance of keywords, and then reconstruct the characteristics of data. The detection method based on the TF-IDF model does not need to analyze each field of the data set. Compared with the general machine learning data preprocessing method, that is, data encoding processing, the experimental results show that using natural language processing technology to preprocess data can effectively improve the accuracy of detection. Gradient boosting classifier, random forest classifier, AdaBoost classifier, and the ensemble model based on these three classifiers are, respectively, used in the construction of the later models. At the same time, CNN neural network in deep learning is also used for training, and CNN can effectively extract data information. Under the condition that the input data of the classifier and neural network are consistent, through the comparison and analysis of various methods, the accuracy of the one-dimensional convolutional network based on CNN is slightly higher than that of the classifier based on machine learning.
Hao Yang 0022, Zhenyan Liu
Secur. Commun. Networks3
2019 MalInsight: A systematic profiling based malware detection framework
abstract
To handle the security threat faced by the widespread use of Internet of Things (IoT) devices due to the ever-lasting increase of malware, the security researchers increasingly rely on machine learning techniques based on various static and/or dynamic features. Unfortunately, the state-of-the-art detection techniques may fail to identify the malware effectively because the malware is often obfuscated to camouflage its characteristics and thwart the analysis process. In order to identify the disguised malware accurately, a malware detection framework named MalInsight is proposed by profiling malware from three aspects which are basic structure, low-level behavior, and high-level behavior. These aspects reflect the structural features, the underlying operations interacting with the OS, and the operations on the files, the registry, and the network respectively. Based on the above findings, an accurate and rich feature space is built which enables to depict and detect malware more effectively. In order to validate the effectiveness of MalInsight, an extensive experiment is conducted on a real-world malware dataset. Our experimental results show that MalInsight can detect not only obfuscated malware instances with an accuracy of 99.76% but also unseen and new malware with an accuracy of 97.21%. Furthermore, MalInsight can classify the malware samples into their families with an accuracy of 94.2% outperforming the typical detection approach based on the API sequence as the dynamic behavior features by almost 9%. In addition, the importance of the three aspects is evaluated and sorted quantitatively demonstrating that these aspects play the same effects with the optimal feature set.
Weijie Han, Jingfeng Xue, Yong Wang 0010, Zhenyan Liu, Zixiao Kong
J. Netw. Comput. Appl.4
2018 An Imbalanced Malicious Domains Detection Method Based on Passive DNS Traffic Analysis
abstract
Although existing malicious domains detection techniques have shown great success in many real-world applications, the problem of learning from imbalanced data is rarely concerned with this day. But the actual DNS traffic is inherently imbalanced; thus how to build malicious domains detection model oriented to imbalanced data is a very important issue worthy of study. This paper proposes a novel imbalanced malicious domains detection method based on passive DNS traffic analysis, which can effectively deal with not only the between-class imbalance problem but also the within-class imbalance problem. The experiments show that this proposed method has favorable performance compared to the existing algorithms.
Zhenyan Liu, Yifei Zeng, Pengfei Zhang 0016, Jingfeng Xue
Secur. Commun. Networks1
2017 Machine Learning for Analyzing Malware
Yajie Dong, Zhenyan Liu, Yida Yan, Yong Wang 0010, Tu Peng
NSS2
2015 A Supervised Parameter Estimation Method of LDA
Zhenyan Liu, Dan Meng 0002, Weiping Wang 0005, Chunxia Zhang 0001
APWeb1
2014 Continuous similarity join on data streams
abstract
Similarity join plays an important role in many applications, such as data cleaning and integration, to address the poor data quality problem. Most of the existing studies focused on performing similarity join on static datasets but few studies realized running it on dynamic data streams. With the development of network technology, the data accessing paradigm has transferred from disk-oriented mode to online data streams, which makes performing similarity join in continuous query on data streams become a novel query processing paradigm. Different from static dataset, data stream is unbounded, continuous and unpredictable. The significant differences pose serious challenges, such as real-time query performance. To this end, we study the problem of continuous similarity join on data streams in this paper, which is based on edit distance metric and filter-and-verify framework with sliding-window semantics. Two subcases of this problem are studied, including self similarity join on a single data stream and similarity join on two streams. We introduced the basic window based sliding window model to facilitate the update of sliding window and its index. More details of our method, including signature extraction schemes, filtering and verification algorithms, re-evaluation strategies are discussed respectively. Finally, extensive experimental results show that our method works efficiently on real data streams.
Jia Cui, Weiping Wang 0005, Dan Meng 0002, Zhenyan Liu
ICPADS4