Zhen Liu 0017

dblp:77/35-17 · DBLP profile ↗
← Back
33ranked-venue papers
15as first author
13since 2021 · last 2026
0000-0002-3327-4118ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 7 first-author · 5 since 2021Databases, data management, data science and information retrieval · 8 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-author · 1 since 2021Systems, architecture and hardware · 2 · 2 first-author · 1 since 2021Security and privacy · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021Computer networks · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 PLCDroid: enhancing android malware detection by mitigating pseudo-label noise in the presence of concept drift
abstract
Abstract Due to the continuous evolution of Android malware, machine learning-based malware detection systems face the challenge of performance degradation. To address this issue, active learning has been employed to retrain models with new labeled data. Traditionally, active learning relies on ground-truth labels, which are time-consuming to obtain. Although leveraging model-predicted pseudo-labels for model retraining offers a cost-effective alternative, incorrect pseudo-labels may lead to model self-contamination. To alleviate the annotation overhead during model retraining and mitigate the detrimental effects of erroneous pseudo-labels on active learning performance, we introduce a novel framework, PLCDroid. The framework incorporates a label correction mechanism when using pseudo-labels for model retraining. Specifically, we present a pseudo-label type recognition method (PTR) based on model uncertainty and confidence to identify incorrect pseudo-labels. On the basis of PTR, we design fine-grained correction strategies to refine pseudo-labels. Consequently, the proposed method mitigates pseudo-label errors, thereby improving malware detection performance under concept drift. Experimental results over a decade-long period demonstrate the effectiveness of our approach. In the retraining task, leveraging corrected pseudo-labels leads to a substantial performance gain. Specifically, the false negative rate decreases from 76.0% to 47.6% on average, corresponding to an improvement of 37.4% compared to the related pseudo label-based active learning method MORPH.
Lingyu Qiu, Zhen Liu 0017, Bitao Peng, Ruoyu Wang 0002
Comput. J.2
2026 ADTDroid: Leveraging API description and TCP based active learning for Android malware detection
Zhen Liu 0017, Ruoyu Wang 0002, Wenbin Zhang 0002
Inf. Softw. Technol.1
2025 AMCR: A Framework for Assessing and Mitigating Copyright Risks in Generative Models
abstract
Generative models have achieved impressive results in text to image tasks, significantly advancing visual content creation. However, this progress comes at a cost, as such models rely heavily on large-scale training data and may unintentionally replicate copyrighted elements, creating serious legal and ethical challenges for real-world deployment. To address these concerns, researchers have proposed various strategies to mitigate copyright risks, most of which are prompt based methods that filter or rewrite user inputs to prevent explicit infringement. While effective in handling obvious cases, these approaches often fall short in more subtle situations, where seemingly benign prompts can still lead to infringing outputs. To address these limitations, this paper introduces Assessing and Mitigating Copyright Risks (AMCR), a comprehensive framework which i) builds upon prompt-based strategies by systematically restructuring risky prompts into safe and non-sensitive forms, ii) detects partial infringements through attention-based similarity analysis, and iii) adaptively mitigates risks during generation to reduce copyright violations without compromising image quality. Extensive experiments validate the effectiveness of AMCR in revealing and mitigating latent copyright risks, offering practical insights and benchmarks for the safer deployment of generative models.
Zhipeng Yin, Zichong Wang, Avash Palikhe, Zhen Liu 0017, Jun Liu 0075, Wenbin Zhang 0002
ECAI4
2025 Redefining Fairness: A Multi-dimensional Perspective and Integrated Evaluation Framework
Zichong Wang, Zhipeng Yin, Zhen Liu 0017, Roland H. C. Yap, Xiaocai Zhang, Shu Hu 0001, Wenbin Zhang 0002
ECML/PKDD (1)3
2025 Incorporating Statistic and Semantic Dependencies for Enhancing the Robustness of Android Malware Detection
abstract
Android’s dominant market share has made it a prime target for malware attacks. Although machine learning-based detection systems have demonstrated effectiveness, they remain vulnerable to adversarial attacks, which modify samples to preserve malicious functionality while evading detection. Adversarial training is a prevalent defense strategy. However, generating effective adversarial examples for Android malware is challenging due to the complex mapping between feature and problem space. To address this, recent efforts have explored feature-space attacks constrained by statistical dependencies. Yet, such approaches inherently rely on large-scale datasets to achieve strong performance, and may fail to capture the underlying semantic relationships among features, like call associations. In this paper, we propose a novel method that incorporates semantic dependencies, i.e., API dependencies extracted from function call graphs of APKs. By leveraging these dependencies as domain constraints, our method preserves intrinsic call associations among features during perturbation. This leads to adversarial examples that more closely reflect realistic attack behaviors. Furthermore, a reinforcement learning-based mechanism is employed to enhance the evasive capability of the generated adversarial samples against detection models. The resulting adversarial samples are leveraged for adversarial training to enhance detector robustness. Experimental results demonstrate that the adversarial examples generated by our approach effectively enhance model robustness via adversarial training, yielding superior resilience in realistic adversarial environments. In adversarial attack scenarios, the proposed method attains the highest detection accuracy against problem-space attacks, surpassing the baseline model without adversarial training by 45.7% and 14.3%, respectively. Moreover, our method significantly reduces the average generation time by 83.5% compared to problem-space adversarial example generation approaches.
Lingyu Qiu, Zhen Liu 0017, Bitao Peng, Ruoyu Wang 0002, Changji Wang, Qingqing Gan
TrustCom2
2025 LDCDroid: Learning data drift characteristics for handling the model aging problem in Android malware detection
Zhen Liu 0017, Ruoyu Wang 0002, Bitao Peng, Lingyu Qiu, Qingqing Gan, Changji Wang, Wenbin Zhang 0002
Comput. Secur.1
2024 SeGDroid: An Android malware detection method based on sensitive function call graph learning
Zhen Liu 0017, Ruoyu Wang 0002, Nathalie Japkowicz, Heitor Murilo Gomes, Bitao Peng, Wenbin Zhang 0002
Expert Syst. Appl.1
2023 Research on Data Drift and Class Imbalance in Android Malware Detection
Zhen Liu 0017, Ruoyu Wang 0002, Bitao Peng, Changji Wang, Qingqing Gan
MobiQuitous (1)1
2021 LSTM Based Sentiment Analysis for Cryptocurrency Prediction
Xin Huang 0005, Wenbin Zhang 0002, Xuejiao Tang, Jayachander Surbiryala, Vasileios Iosifidis, Zhen Liu 0017, Ji Zhang 0001
DASFAA (3)7
2021 Cognitive Visual Commonsense Reasoning Using Dynamic Working Memory
Xuejiao Tang, Xin Huang 0005, Wenbin Zhang 0002, Travers B. Child, Zhen Liu 0017, Ji Zhang 0001
DaWaK6
2021 An Effective Algorithm for Classification of Text with Weak Sequential Relationships
Qiqiang Xu, Ji Zhang 0001, Ting Yu 0004, Wenbin Zhang 0002, Yonglong Luo, Fulong Chen 0002, Zhen Liu 0017
DEXA (2)8
2021 A Generic Knowledge Based Medical Diagnosis Expert System
abstract
In this paper, we design and implement a generic medical knowledge based system (MKBS) for identifying diseases from several symptoms. In this system, some important aspects like knowledge bases system, knowledge representation, inference engine have been addressed. The system asks users different questions and inference engines will use the certainty factor to prune out low possible solutions. The proposed disease diagnosis system also uses a graphical user interface (GUI) to facilitate users to interact with the expert system. Our expert system is generic and flexible, which can be integrated with any rule bases system in disease diagnosis.
Xin Huang 0005, Xuejiao Tang, Wenbin Zhang 0002, Ji Zhang 0001, Wensheng Gan, Shichao Pei, Zhen Liu 0017, Yiyi Huang
iiWAS8
2021 Research on unsupervised feature learning for Android malware detection based on Restricted Boltzmann Machines
Zhen Liu 0017, Ruoyu Wang 0002, Nathalie Japkowicz, Deyu Tang, Wenbin Zhang 0002, Jie Zhao 0011
Future Gener. Comput. Syst.1
2020 Using Machine Learning to Automate Mammogram Images Analysis
abstract
Breast cancer is the second leading cause of cancer-related death after lung cancer in women. Early detection of breast cancer in X-ray mammography is believed to have effectively reduced the mortality rate since 1989. However, a relatively high false positive rate and a low specificity in mammography technology still exist. In this work, a computer-aided automatic mammogram analysis system is proposed to process the mammogram images and automatically discriminate them as either normal or cancerous, consisting of three consecutive image processing, feature selection, and image classification stages. In designing the system, the discrete wavelet transforms (Daubechies 2, Daubechies 4, and Biorthogonal 6.8) and the Fourier cosine transform were first used to parse the mammogram images and extract statistical features. Then, an entropy-based feature selection method was implemented to reduce the number of features. Finally, different pattern recognition methods (including the Back-propagation Network, the Linear Discriminant Analysis, and the Naive Bayes Classifier) and a voting classification scheme were employed. The performance of each classification strategy was evaluated for sensitivity, specificity, and accuracy and for general performance using the Receiver Operating Curve. Our method is validated on the dataset from the Eastern Health in Newfoundland and Labrador of Canada. The experimental results demonstrated that the proposed automatic mammogram analysis system could effectively improve the classification performances.
Xuejiao Tang, Liuhua Zhang, Wenbin Zhang 0002, Xin Huang 0005, Vasileios Iosifidis, Zhen Liu 0017, Enza Messina, Ji Zhang 0001
BIBM6
2020 A Data-driven Human Responsibility Management System
abstract
An ideal safe workplace is described as a place where staffs fulfill responsibilities in a well-organized order, potential hazardous events are being monitored in real-time, as well as the number of accidents and relevant damages are minimized. However, occupational-related death and injury are still increasing and have been highly attended in the last decades due to the lack of comprehensive safety management. A smart safety management system is therefore urgently needed, in which the staffs are instructed to fulfill responsibilities as well as automating risk evaluations and alerting staffs and departments when needed. In this paper, a smart system for safety management in the workplace based on responsibility big data analysis and the internet of things (IoT) are proposed. The real world implementation and assessment demonstrate that the proposed systems have superior accountability performance and improve the responsibility fulfillment through real-time supervision and self-reminder.
Xuejiao Tang, Jiong Qiu, Wenbin Zhang 0002, Vasileios Iosifidis, Zhen Liu 0017, Ji Zhang 0001
IEEE BigData6
2020 Flexible and Adaptive Fairness-aware Learning in Non-stationary Data Streams
abstract
Artificial intelligence (AI)-based decision-making systems are employed nowadays in an ever growing number of online as well as offline services-some of great importance. Depending on sophisticated learning algorithms and available data, these systems are increasingly becoming automated and data-driven. However, these systems can impact individuals and communities with ethical or legal consequences. Numerous approaches have therefore been proposed to develop decision-making systems that are discrimination-conscious by-design. However, these methods assume the underlying data distribution is stationary without drift, which is counterfactual in many realworld applications. In addition, their focus has been largely on minimizing discrimination while maximizing prediction performance without necessary flexibility in customizing the tradeoff according to different applications. To this end, we propose a learning algorithm for fair classification that also adapts to evolving data streams and further allows for a flexible control on the degree of accuracy and fairness. The positive results on a set of discriminated and non-stationary data streams demonstrate the effectiveness and flexibility of this approach.
Wenbin Zhang 0002, Ji Zhang 0001, Zhen Liu 0017, Zhiyuan Chen 0003, Jianwu Wang 0001, Edward Raff, Enza Messina
ICTAI4
2020 A statistical pattern based feature extraction method on system call traces for anomaly detection
Zhen Liu 0017, Nathalie Japkowicz, Ruoyu Wang 0002, Yongming Cai, Deyu Tang, Xian-Fa Cai
Inf. Softw. Technol.1
2020 NEC: A nested equivalence class-based dependency calculation approach for fast feature selection using rough set theory
Jie Zhao 0011, Zhenning Dong, Deyu Tang, Zhen Liu 0017
Inf. Sci.5
2020 Memetic quantum evolution algorithm for global optimization
Deyu Tang, Zhen Liu 0017, Jie Zhao 0011, Shoubin Dong, Yongming Cai
Neural Comput. Appl.2
2020 Spherical search optimizer: a simple yet efficient meta-heuristic approach
Jie Zhao 0011, Deyu Tang, Zhen Liu 0017, Yongming Cai, Shoubin Dong
Neural Comput. Appl.3
2020 Accelerating information entropy-based feature selection using rough set theory with classified nested equivalence classes
Jie Zhao 0011, Zhenning Dong, Deyu Tang, Zhen Liu 0017
Pattern Recognit.5
2020 A sub-concept-based feature selection method for one-class classification
Zhen Liu 0017, Nathalie Japkowicz, Ruoyu Wang 0002
Soft Comput.1
2019 Adaptive learning on mobile network traffic data
abstract
Machine learning based mobile traffic classification has become a popular topic in recent years. As mobile traffic data is dynamic in nature, the static model has become ineffective for the task of classifying future traffic. This is known as the concept drift problem in data streams. To this end, this paper presents an adaptive mobile traffic classification method. Specifically, a method based on the fuzzy competence model is devised to detect concept drift, and a dynamic learning method is presented to update the classification model, so as to adapt to an ever-changing environment at an appropriate time. The concept drift detection method relies on the data distribution instead of the classification error rate. Furthermore, the weights of flow samples are dynamically updated and flow samples are resampled for training a new model when a concept drift is detected. Moreover, recently trained models are saved and used for classification in weighted voting. The weight of each model is updated according to the performance it obtains on the most recent flow samples. On mobile traffic data, experimental results show that our proposed method obtains lower classification error rate with less time consumption on updating models as compared to related methods designed for handling concept drift problems.
Zhen Liu 0017, Nathalie Japkowicz, Ruoyu Wang 0002, Deyu Tang
Connect. Sci.1
2019 Mobile app traffic flow feature extraction and selection for improving classification robustness
Zhen Liu 0017, Ruoyu Wang 0002, Nathalie Japkowicz, Yongming Cai, Deyu Tang, Xian-Fa Cai
J. Netw. Comput. Appl.1
2019 Memetic frog leaping algorithm for global optimization
Deyu Tang, Zhen Liu 0017, Jin Yang 0004, Jie Zhao 0011
Soft Comput.2
2018 Adaptive Threshold for Outlier Detection on Data Streams
abstract
As the distribution of a data stream evolves over time, a learner must adapt to its distributional shifts in order to make accurate predictions. In the context of anomaly detection, it is crucial for the learner to distinguish between natural changes in distribution and true anomalies in the data stream. This is the problem we focus on in this study which considers the situation where only normal data are available for initial training, but subsequent data can be either normal or anomalous. In that context, it is necessary to train a one-class learning anomaly detection system on the normal data and let the system output a score representing the degree of normalcy or outlierness that each data point in the subsequent data stream exhibits. The system then uses a threshold to discriminate between normal and anomalous instances. In the case of data stream, the data distribution may shift overtime, and a fixed threshold could develop a high false alarm rate or a low outlier detection rate in case of concept drift. To this end, we designed an adaptive sliding window approach which updates the threshold when necessary based on the scores distribution. Experimental results show that our method improves the performance of base anomaly detectors by dynamically updating the threshold of the scores when needed rather than using a fixed threshold or an adaptive threshold with fixed window sizes.
Zhen Liu 0017, Nathalie Japkowicz
DSAA2
2018 Benchmark Data for Mobile App Traffic Research
abstract
Mobile app traffic classification aims to automatically map mobile packets into apps. It has become an active task in mobile traffic engineering, and numerous algorithms have been proposed for this task, including machine learning, deep packet inspection methods. However, existing works mainly evaluate their methods on their own collected mobile traffic traces. There is no public benchmark data. The results in existing papers cannot be directly compared. This largely limits the development of mobile app traffic classification methods. This paper describes our Mobile Traffic Data(MTD): Android app traffic flow sample sets with ground truth. The goal of MTD is to advance the state-of-arts in mobile app traffic classification. For building MTD, we collected and annotated more than ten thousands of traffic flows using Mobilegt system. The popularity used flow features were also extracted to build flow samples for mobile traffic classification using machine learning. MTD sets have been shared in public. In addition, this paper provides the performance analysis of typical machine learning techniques on MTD, which can be served as the baseline results on this benchmark data.
Ruoyu Wang 0002, Zhen Liu 0017, Yongming Cai, Deyu Tang, Jin Yang 0004
MobiQuitous2
2018 Extending labeled mobile network traffic data by three levels traffic identification fusion
Zhen Liu 0017, Ruoyu Wang 0002, Deyu Tang
Future Gener. Comput. Syst.1
2017 Objective cost-sensitive-boosting-WELM for handling multi class imbalance problem
abstract
Class imbalance problem has attracted a great attention in the field of ELM (extreme learning machine). Cost sensitive ELM was proposed to address class imbalance but it merely handled binary class imbalance and required to predefine misclassification costs subjectively. Boosting WELM has been presented to handle multi class imbalance, and performed well on improving the classification accuracy of the minority class, but it may excessively strengthen minority class samples and degrade the performance of the majority class. This paper presents a method named OCS-BWELM (objective cost-sensitive-boosting-WELM) to handle multi class imbalance. It takes boosting WELM as the basic learning algorithm. The misclassification costs are determined by the distributions of the given data rather than being defined subjectively. More specifically, it seeks optimal costs through maximizing the mutual information between real targets and prediction outputs. A specific feature of OCS-BWELM is that its costs are objective. Experiments are carried out to compare our method against existing ELM related works on handling multi class imbalance. Results show that our method could achieve a better performance balance between minority class and majority class than boosting WELM. And it outperforms others in terms of G-mean, F-score and F-measures of minority classes in most cases.
Zhen Liu 0017, Deyu Tang, Ruoyu Wang 0002
IJCNN1
2017 A hybrid method based on ensemble WELM for handling multi class imbalance in cancer microarray data
Zhen Liu 0017, Deyu Tang, Yongming Cai, Ruoyu Wang 0002, Fuhua Chen
Neurocomputing1
2016 A System for Linking Ground Truth to Mobile Network Traffic
abstract
Mobile network traffic engineering and management activities require traffic traces where each packet or flow is associated with some ground truth regarding mobile app or protocol. This paper presents a system named mobilegt that collects mobile traffic and links the ground truth to it without rooting mobile devices. It consists of two elements: mobilegt client and mobilegt server. Mobilegt client iteratively probes monitored mobile nodes' kernel to obtain socket information on active TCP/UDP sessions. Mobilegt server captures the packets generated on monitored nodes at the aid of Virtual Private Network (VPN), and labels each packet/flow by exploring the association between socket and packet. Our preliminary experimental results show that mobilegt can tag more than 98% of bytes and 93% of flows on average without significantly affecting CPU load.
Zhen Liu 0017, Ruoyu Wang 0002, Deyu Tang
MobiQuitous1
2016 Corrigendum to "A class-oriented feature selection approach for multi-class imbalanced network traffic datasets based on local and global metrics fusion" [Neurocomputing 168 (2015) 365-381]
Zhen Liu 0017, Ruoyu Wang 0002, Ming Tao 0001, Xian-Fa Cai
Neurocomputing1
2015 A class-oriented feature selection approach for multi-class imbalanced network traffic datasets based on local and global metrics fusion
Zhen Liu 0017, Ruoyu Wang 0002, Ming Tao 0001, Xian-Fa Cai
Neurocomputing1