Bitao Peng

dblp:239/9951 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0002-3741-7972ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Security and privacy · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 PLCDroid: enhancing android malware detection by mitigating pseudo-label noise in the presence of concept drift
abstract
Abstract Due to the continuous evolution of Android malware, machine learning-based malware detection systems face the challenge of performance degradation. To address this issue, active learning has been employed to retrain models with new labeled data. Traditionally, active learning relies on ground-truth labels, which are time-consuming to obtain. Although leveraging model-predicted pseudo-labels for model retraining offers a cost-effective alternative, incorrect pseudo-labels may lead to model self-contamination. To alleviate the annotation overhead during model retraining and mitigate the detrimental effects of erroneous pseudo-labels on active learning performance, we introduce a novel framework, PLCDroid. The framework incorporates a label correction mechanism when using pseudo-labels for model retraining. Specifically, we present a pseudo-label type recognition method (PTR) based on model uncertainty and confidence to identify incorrect pseudo-labels. On the basis of PTR, we design fine-grained correction strategies to refine pseudo-labels. Consequently, the proposed method mitigates pseudo-label errors, thereby improving malware detection performance under concept drift. Experimental results over a decade-long period demonstrate the effectiveness of our approach. In the retraining task, leveraging corrected pseudo-labels leads to a substantial performance gain. Specifically, the false negative rate decreases from 76.0% to 47.6% on average, corresponding to an improvement of 37.4% compared to the related pseudo label-based active learning method MORPH.
Lingyu Qiu, Zhen Liu 0017, Bitao Peng, Ruoyu Wang 0002
Comput. J.3
2025 Incorporating Statistic and Semantic Dependencies for Enhancing the Robustness of Android Malware Detection
abstract
Android’s dominant market share has made it a prime target for malware attacks. Although machine learning-based detection systems have demonstrated effectiveness, they remain vulnerable to adversarial attacks, which modify samples to preserve malicious functionality while evading detection. Adversarial training is a prevalent defense strategy. However, generating effective adversarial examples for Android malware is challenging due to the complex mapping between feature and problem space. To address this, recent efforts have explored feature-space attacks constrained by statistical dependencies. Yet, such approaches inherently rely on large-scale datasets to achieve strong performance, and may fail to capture the underlying semantic relationships among features, like call associations. In this paper, we propose a novel method that incorporates semantic dependencies, i.e., API dependencies extracted from function call graphs of APKs. By leveraging these dependencies as domain constraints, our method preserves intrinsic call associations among features during perturbation. This leads to adversarial examples that more closely reflect realistic attack behaviors. Furthermore, a reinforcement learning-based mechanism is employed to enhance the evasive capability of the generated adversarial samples against detection models. The resulting adversarial samples are leveraged for adversarial training to enhance detector robustness. Experimental results demonstrate that the adversarial examples generated by our approach effectively enhance model robustness via adversarial training, yielding superior resilience in realistic adversarial environments. In adversarial attack scenarios, the proposed method attains the highest detection accuracy against problem-space attacks, surpassing the baseline model without adversarial training by 45.7% and 14.3%, respectively. Moreover, our method significantly reduces the average generation time by 83.5% compared to problem-space adversarial example generation approaches.
Lingyu Qiu, Zhen Liu 0017, Bitao Peng, Ruoyu Wang 0002, Changji Wang, Qingqing Gan
TrustCom3
2025 LDCDroid: Learning data drift characteristics for handling the model aging problem in Android malware detection
Zhen Liu 0017, Ruoyu Wang 0002, Bitao Peng, Lingyu Qiu, Qingqing Gan, Changji Wang, Wenbin Zhang 0002
Comput. Secur.3
2024 SeGDroid: An Android malware detection method based on sensitive function call graph learning
Zhen Liu 0017, Ruoyu Wang 0002, Nathalie Japkowicz, Heitor Murilo Gomes, Bitao Peng, Wenbin Zhang 0002
Expert Syst. Appl.5
2023 CEVulDet: A Code Edge Representation Learnable Vulnerability Detector
abstract
Many researchers have started to apply deep learning algorithms for source code vulnerability detection. However, the detection results of existing methods are still not accurate enough. Most of these methods treat the program source code directly as natural language. These methods may ignore the structural information specific to the program code, which is a key part of the semantics of the code composition. In this paper, we propose a novel vulnerability detection method named CEVulDet. First, we adopt centrality analysis to remove unimportant nodes from the PDG to obtain a new graph that preserves the important parts of the program. Second, we propose a new program semantic extraction method that acquires feature vectors to represent the semantic information of the program code and the graph edge information. It can locate the vulnerability trigger path with the aid of the model explanation technique. Finally, the extracted vectors obtained by our method are fed into the CNN to train a vulnerability detector. In our experiments, we evaluate the performance of CEVulDet on a dataset containing 33,360 functions with 12,303 vulnerable functions and 21,057 non-vulnerable functions. Experimental results show that CEVulDet far outperforms rule-based detectors and trumps the state-of-the-art deep learning-based detector. CEVulDet improved by 3.2%, 3.4%, 5.1% and 4.2% in terms of accuracy, precision, recall and F1 metrics, respectively.
Bitao Peng, Pengcheng Su
IJCNN1
2023 Research on Data Drift and Class Imbalance in Android Malware Detection
Zhen Liu 0017, Ruoyu Wang 0002, Bitao Peng, Changji Wang, Qingqing Gan
MobiQuitous (1)3
2020 A heuristic algorithm combining Pareto optimization and niche technology for multi-objective unequal area facility layout problem
Jingfa Liu, Jun Liu 0049, Xueming Yan, Bitao Peng
Eng. Appl. Artif. Intell.4