EDBT 2026 Demo / reviewers in the wild / expert
Kaifa Zhao
dblp:201/9069
· DBLP profile ↗
16ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0002-3102-7498ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 2 since 2021Security and privacy · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Interpretable Defense Against Structural Adversarial Attacks on Android Malware DetectionabstractAndroid, being one of the most widely used mobile systems, is facing pressing threats from malware. Despite the effectiveness of Android malware detection (AMD) systems, they are still vulnerable to state-of-the-art adversarial attacks. Existing defense methods require the knowledge of target adversaries, such as attack algorithms or obfuscation strategies, which is impractical in real-world scenarios. Additionally, these approaches may adversely affect the performance of the detection model and fail to defend against problem-space attacks, which not only deceive the detection models but also generate executable adversarial software. To address this research gap, we propose a novel interpretable Android guard system, named IADGuard, to help AMD defend against attacks. IADGuard first designs a novel graph explainable method, AGExplainer, to identify suspicious functions and invocations in adversarial malware. With the guidance of AGExplainer, IADGuard develops a rectifier to reverse adversarial modifications on apps’ function invocation relations, which facilitates the detection of adversarial malware by victim AMD. It is noteworthy that IADGuard requires zero knowledge of adversarial models and victim models, thereby preserves the performance of victim AMD. We validate IADGuard over three state-of-the-art problem space attacks that modify apps’ function invocation relations to deceive victim AMD. Experimental results show that IADGuard achieves over 90.5% defense success rate, i.e., helps victim AMD identify adversarial malware. Furthermore, AGExplainer surpasses representative interpreters in identifying essential modifications, helps IADGuard reduce false positives to 1.5%, and improves the detection efficiency by up to 10.4 times. Wenying Wei, Kaifa Zhao, Hao Zhou 0043, Jianfeng Li 0006, Shuohan Wu, Ming Fan 0002, Xiapu Luo, Ting Wang 0006, Kai Zhou 0001, Ting Liu 0002, Yuzhe Tang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | Vehicular Intrusion Detection System for Controller Area Network: A Comprehensive Survey and EvaluationabstractThe progress of automotive technologies has made cybersecurity a crucial focus, leading to various cyber attacks. These attacks primarily target the Controller Area Network (CAN) and specialized Electronic Control Units (ECUs). In order to mitigate these attacks and bolster the security of vehicular systems, numerous defense solutions have been proposed. These solutions aim to detect diverse forms of vehicular attacks. However, the practical implementation of these solutions still presents certain limitations and challenges. In light of these circumstances, this paper undertakes a thorough examination of existing vehicular attacks and defense strategies employed against the CAN and ECUs. The objective is to provide valuable insights and inform the future design of Vehicular Intrusion Detection Systems (VIDS). The findings of our investigation reveal that the examined VIDS primarily concentrate on particular categories of attacks, neglecting the broader spectrum of potential threats. Moreover, we provide a comprehensive overview of the significant challenges encountered in implementing a robust and feasible VIDS. Additionally, we put forth several defense recommendations based on our study findings, aiming to inform and guide the future design of VIDS in the context of vehicular security. Lei Xue 0001, Sishan Wang, Xiapu Luo, Kaifa Zhao, Pengfei Jing, Xiaobo Ma 0001, Yajuan Tang, Haiying Zhou |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2025 | From Bi-Level to One-Level: A Framework for Structural Attacks to Graph Anomaly DetectionabstractThe success of graph neural networks stimulates the prosperity of graph mining and the corresponding downstream tasks including graph anomaly detection (GAD). However, it has been explored that those graph mining methods are vulnerable to structural manipulations on relational data. That is, the attacker can maliciously perturb the graph structures to assist the target nodes in evading anomaly detection. In this article, we explore the structural vulnerability of two typical GAD systems: unsupervised FeXtra-based GAD and supervised graph convolutional network (GCN)-based GAD. Specifically, structural poisoning attacks against GAD are formulated as complex bi-level optimization problems. Our first major contribution is then to transform the bi-level problem into one-level leveraging different regression methods. Furthermore, we propose a new way of utilizing gradient information to optimize the one-level optimization problem in the discrete domain. Comprehensive experiments demonstrate the effectiveness of our proposed attack algorithm $\textsf {BinarizedAttack}$ . Yulin Zhu 0001, Yuni Lai, Kaifa Zhao, Xiapu Luo, Mingquan Yuan, Jun Wu 0001, Jian Ren 0001, Kai Zhou 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Demystifying Privacy and Security Issues in Potentially Harmful Mobile ApplicationsabstractAndroid system's features, such as support for peer-to-peer networking, cloud computing services, and various communication protocols and network technologies, make it a suitable platform for distributed computing systems that require efficient resource management and coordination. The widespread popularity of Android, as one of the most widely-used mobile operating systems [1], is attributed to its ability to provide users with a wide range of convenient and entertaining options through its functional apps. Nevertheless, mobile users may face a risk to their privacy and property from potentially harmful apps that can be installed on their devices. Thus, it is imperative to conduct an analysis of potentially harmful applications [2]–[10] with a sense of urgency. Our research focuses on combining Android static analysis, adversarial machine learning, and natural language processing techniques to investigate app behavior, discover vulnerabilities in Android malware detection systems, and understand software privacy policy statements. To safeguard user privacy from potentially harmful apps, we propose the following measures: (I) Investigating the vulnerability of Android malware detection systems under evolving attacks and proposing defense solutions [11]; (II) Analyzing whether Android app privacy policies meet regulatory requirements [12]; and (III) Examining the privacy concerns associated with third-party libraries in Android [13]. For (I), to investigate the vulnerability of Android malware detection (AMD) systems under evolving attacks and discover defense solutions, we propose a Heuristic optimization model integrated with Reinforcement learning framework to optimize our structural ATtack, namely HRAT [11], which is the first problem-pace structural attack designed to deceive AMD. HRAT employs four types of graph modification operations and corresponding bytecode manipulation techniques to generate executable adversarial apps that can evade detection. HRAT bridges the research gap between feature-space attacks, which generate only adversarial features to deceive machine-learning models, and problem-space attacks, which generate complete adversarial objects. Our extensive experiments demonstrate that HRAT has effective attack performance and remains robust against obfuscation methods that do not affect the app's function call graph. In addition, we propose potential defense solutions to improve the robustness of AMD against such advanced attack methods. For (II), we construct the first large-scale human-annotated Chinese Android application privacy policy dataset, namely CA4P-483 [12]. Following-a manual inspection of regulatory articles, we identified seven types of labels that are relevant to the regulatory requirements for apps' access to user data. We then designed a two-step annotation process to ensure label agreement, and our evaluation showed that our annotations achieved a Kappa value of 77.20%, indicating substantial agreement for CA4P-483. In addition, we evaluate robust and representative baseline models on our dataset and present our findings and potential research directions based on our results. Finally, we conduct case studies to explore the potential application of CA4P-483 for protecting user privacy. For (III), we propose an Automated Third-Party library Privacy compliance Checker tool, ATPChecker [13], for checking Android third-party libraries (TPL) for compliance. ATPChecker combines static analysis to identify third-party libraries and analyze host app behavior in bytecode, as well as natural language processing techniques to investigate the privacy access-related statements in their privacy policies. Subsequently, ATPChecker determines whether the identified third-party libraries satisfy regulatory requirements and whether the host apps' use of these libraries complies with regulatory requirements by cross-checking the results of the static analysis and privacy policy analysis. For future work, first, we can develop more effective and efficient attack strategies against Android malware detection systems. The utilization of HRAT is constrained to white-box settings, thereby limiting its potential application scenarios. In cases where access to the classification model in victim malware detection systems is not feasible, a substitute model can be trained using input and output data. For instance, the classification probability of victim classifiers can be used to train a model that mimics the target classifier. Afterward, the HRAT approach can be applied to deceive the substitute model and achieve similar evasion performance of victim detection systems. Secondly, an exploration of potential application scenarios for the CA4P-483 dataset. It is a challenging task to determine whether a privacy policy will share or collect user data due to the use of negative words and tones. Applying emotional analysis to identify negative statements in privacy policies is a promising direction for understanding privacy policies. Lastly, enhancing ATPChecker by analyzing the purpose of data collection and developing automatic privacy policy generation methods for third-party library developers to reduce their workload and enhance compliance with regulations. Kaifa Zhao |
ICDCS | 1 |
| 2023 | Demystifying Privacy Policy of Third-Party Libraries in Mobile AppsabstractThe privacy of personal information has received significant attention in mobile software. Although researchers have designed methods to identify the conflict between app behavior and privacy policies, little is known about the privacy compliance issues relevant to third-party libraries (TPLs). The regulators enacted articles to regulate the usage of personal information for TPLs (e.g., the CCPA requires businesses clearly notify consumers if they share consumers' data with third parties or not). However, it remains challenging to investigate the privacy compliance issues of TPLs due to three reasons: 1) Difficulties in collecting TPLs' privacy policies. In contrast to Android apps, which are distributed through markets like Google Play and must provide privacy policies, there is no unique platform for collecting privacy policies of TPLs. 2) Difficulties in analyzing TPL's user privacy access behaviors. TPLs are mainly provided in binary files, such as jar or aar, and their whole functionalities usually cannot be executed independently without host apps. 3) Difficulties in identifying consistency between TPL's functionalities and privacy policies, and host app's privacy policy and data sharing with TPLs. This requires analyzing not only the privacy policies of TPLs and host apps but also their functionalities. In this paper, we propose an automated system named ATPChecker to analyze whether Android TPLs comply with the privacy-related regulations. We construct a data set that contains a list of 458 TPLs, 247 TPL's privacy policies, 187 TPL's binary files and 641 host apps and their privacy policies. Then, we analyze the bytecode of TPLs and host apps, design natural language processing systems to analyze privacy policies, and implement an expert system to identify TPL usage-related regulation compliance. The experimental results show that 23% TPLs violate regulation requirements for providing privacy policies. Over 47% TPLs miss disclosing data usage in their privacy policies. Over 65% host apps share user data with TPLs while 65% of them miss disclosing interactions with TPLs. Our findings remind developers to be mindful of TPL usage when developing apps or writing privacy policies to avoid violating regulations, Kaifa Zhao, Xian Zhan, Le Yu 0002, Shiyao Zhou, Hao Zhou 0043, Xiapu Luo, Haoyu Wang 0001, Yepang Liu 0001 |
ICSE | 1 |
| 2023 | CydiOS: A Model-Based Testing Framework for iOS AppsabstractTo make an app stand out in an increasingly competitive market, developers must ensure its quality to deliver a better user experience. UI testing is a popular technique for quality assurance, which can thoroughly test the app from the users’ perspective. However, while considerable research has already studied UI testing on the Android platform, there is no research on iOS. This paper introduces CydiOS, a novel approach to performing model-based testing for iOS apps. CydiOS enhances the existing static analysis to build a more complete static model for the app under test. We propose an approach to retrieve runtime information to obtain real-time app context that can be mapped in the model. To improve the effectiveness of UI testing, we also introduce a potential-aware search algorithm to guide testing execution. We compare CydiOS with four representative algorithms(i.e., random, depth-first, stoat, and ape). We have evaluated CydiOS on 50 popular apps from App Store, and the results show that CydiOS outperforms other tools, achieving both higher code coverage and screen coverage. We open source CydiOS at https://github.com/SoftWare2022Testing/CydiOS, and a demo video can be found there. Shuohan Wu, Jianfeng Li 0006, Hao Zhou 0043, Yongsheng Fang, Kaifa Zhao, Haoyu Wang 0001, Chenxiong Qian, Xiapu Luo |
ISSTA | 5 |
| 2022 | A Fine-grained Chinese Software Privacy Policy Dataset for Sequence Labeling and Regulation Compliant IdentificationabstractPrivacy protection raises great attention on both legal levels and user awareness.To protect user privacy, countries enact laws and regulations requiring software privacy policies to regulate their behavior.However, privacy policies are written in natural languages with many legal terms and software jargon that prevent users from understanding and even reading them.It is desirable to use NLP techniques to analyze privacy policies for helping users understand them.Furthermore, existing datasets ignore law requirements and are limited to English.In this paper, we construct the first Chinese privacy policy dataset, namely CA4P-483, to facilitate the sequence labeling tasks and regulation compliance identification between privacy policies and software.Our dataset includes 483 Chinese Android application privacy policies, over 11K sentences, and 52K fine-grained annotations.We evaluate families of robust and representative baseline models on our dataset.Based on baseline performance, we provide findings and potential research directions on our dataset.Finally, we investigate the potential applications of CA4P-483 1 combing regulation requirements and program analysis. Kaifa Zhao, Le Yu 0002, Shiyao Zhou, Jing Li 0049, Xiapu Luo, Aemon Yat Fei Chiu |
EMNLP | 1 |
| 2022 | BinarizedAttack: Structural Poisoning Attacks to Graph-based Anomaly DetectionabstractGraph-based Anomaly Detection (GAD) is becoming prevalent due to the powerful representation abilities of graphs as well as recent advances in graph mining techniques. These GAD tools, however, expose a new attacking surface, ironically due to their unique advantage of being able to exploit the relations among data. That is, attackers now can manipulate those relations (i.e., the structure of the graph) to allow some target nodes to evade detection. In this paper, we exploit this vulnerability by designing a new type of targeted structural poisoning attacks to a representative regression-based GAD system termed OddBall. Specifically, we formulate the attack against OddBall as a bi-level optimization problem, where the key technical challenge is to efficiently solve the problem in a discrete domain. We propose a novel attack method termed BinarizedAttack based on gradient descent. Comparing to prior arts, BinarizedAttack can better use the gradient information, making it particularly suitable for solving combinatorial optimization problems. Furthermore, we investigate the attack transferability of BinarizedAttack by employing it to attack other representation-learning-based GAD systems. Our comprehensive experiments demonstrate that BinarizedAttack is very effective in enabling target nodes to evade graph-based anomaly detection tools with limited attacker's budget, and in the black-box transfer attack setting, BinarizedAttack is also tested effective and in particular, can significantly change the node embeddings learned by the GAD systems. Our research thus opens the door to studying a new type of attack against security analytic tools that rely on graph data. Yulin Zhu 0001, Yuni Lai, Kaifa Zhao, Xiapu Luo, Mingquan Yuan, Jian Ren 0001, Kai Zhou 0001 |
ICDE | 3 |
| 2022 | SAID: State-aware Defense Against Injection Attacks on In-vehicle Network
Lei Xue 0001, Kaifa Zhao, Jianfeng Li 0006, Le Yu 0002, Xiapu Luo, Yajin Zhou, Guofei Gu |
USENIX Security Symposium | 4 |
| 2022 | Towards Automatically Reverse Engineering Vehicle Diagnostic Protocols
Le Yu 0002, Pengfei Jing, Xiapu Luo, Lei Xue 0001, Kaifa Zhao, Yajin Zhou, Ting Wang 0006, Guofei Gu, Sen Nie, Shi Wu |
USENIX Security Symposium | 6 |
| 2021 | Structural Attack against Graph Based Android Malware DetectionabstractMalware detection techniques achieve great success with deeper insight into the semantics of malware. Among existing detection techniques, function call graph (FCG) based methods achieve promising performance due to their prominent representations of malware's functionalities. Meanwhile, recent adversarial attacks not only perturb feature vectors to deceive classifiers (i.e., feature-space attacks) but also investigate how to generate real evasive malware (i.e., problem-space attacks). However, existing problem-space attacks are limited due to their inconsistent transformations between feature space and problem space. Kaifa Zhao, Hao Zhou 0043, Yulin Zhu 0001, Xian Zhan, Kai Zhou 0001, Jianfeng Li 0006, Le Yu 0002, Wei Yuan 0001, Xiapu Luo |
CCS | 1 |
| 2021 | Inaccurate Prediction Is Not Always Bad: Open-World Driver Recognition via Error AnalysisabstractDriver identification is of fundamental importance in many vehicle-related applications, such as fleet monitoring and anti-theft system. The vast majority of existing methods work under the closed-world assumption, which may be unrealistic in practice. In this paper, we consider a more practical but challenging scenario, i.e., open-world driver recognition, and propose a systematic method dubbed DRIVERPRINT. To recognize the driver of interest, DRIVERPRINT takes advantage of the behavioral predictability of the driver himself, thereby no need to collect data from other drivers for model training. Specifically, DRIVERPRINT predicts the behavior-related traveling speed with a driver-specific predictor, compares the prediction error with a pre-trained error model and finally recognizes drivers via error analysis. Besides open-world setting, our method is also compatible with closed-world driver classification. Real-world experiments demonstrate our method achieves reasonable accuracy. The average F1-score for open-world driver recognition is up to 0.91, while that for closed-world driver classification is up to 0.973. Jianfeng Li 0006, Kaifa Zhao, Yajuan Tang, Xiapu Luo, Xiaobo Ma 0001 |
VTC Spring | 2 |
| 2020 | View-collaborative fuzzy soft subspace clustering for automatic medical image segmentation
Kaifa Zhao, Yizhang Jiang, Kaijian Xia, Leyuan Zhou, Pengjiang Qian |
Multim. Tools Appl. | 1 |
| 2020 | mDixon-Based Synthetic CT Generation for PET Attenuation Correction on Abdomen and Pelvis Jointly Using Transfer Fuzzy Clustering and Active Learning-Based ClassificationabstractWe propose a new method for generating synthetic CT images from modified Dixon (mDixon) MR data. The synthetic CT is used for attenuation correction (AC) when reconstructing PET data on abdomen and pelvis. While MR does not intrinsically contain any information about photon attenuation, AC is needed in PET/MR systems in order to be quantitatively accurate and to meet qualification standards required for use in many multi-center trials. Existing MR-based synthetic CT generation methods either use advanced MR sequences that have long acquisition time and limited clinical availability or use matching of the MR images from a newly scanned subject to images in a library of MR-CT pairs which has difficulty in accounting for the diversity of human anatomy especially in patients that have pathologies. To address these deficiencies, we present a five-phase interlinked method that uses mDixon MR acquisition and advanced machine learning methods for synthetic CT generation. Both transfer fuzzy clustering and active learning-based classification (TFC-ALC) are used. The significance of our efforts is fourfold: 1) TFC-ALC is capable of better synthetic CT generation than methods currently in use on the challenging abdomen using only common Dixon-based scanning. 2) TFC partitions MR voxels initially into the four groups regarding fat, bone, air, and soft tissue via transfer learning; ALC can learn insightful classifiers, using as few but informative labeled examples as possible to precisely distinguish bone, air, and soft tissue. Combining them, the TFC-ALC method successfully overcomes the inherent imperfection and potential uncertainty regarding the co-registration between CT and MR images. 3) Compared with existing methods, TFC-ALC features not only preferable synthetic CT generation but also improved parameter robustness, which facilitates its clinical practicability. Applying the proposed approach on mDixon-MR data from ten subjects, the average score of the mean absolute prediction deviation (MAPD) was 89.78±8.76 which is significantly better than the 133.17±9.67 obtained using the all-water (AW) method (p=4.11E-9) and the 104.97±10.03 obtained using the four-cluster-partitioning (FCP, i.e., external-air, internal-air, fat, and soft tissue) method (p=0.002). 4) Experiments in the PET SUV errors of these approaches show that TFC-ALC achieves the highest SUV accuracy and can generally reduce the SUV errors to 5% or less. These experimental results distinctively demonstrate the effectiveness of our proposed TFCALC method for the synthetic CT generation on abdomen and pelvis using only the commonly-available Dixon pulse sequence. Pengjiang Qian, Jung-Wen Kuo, Yudong Zhang 0001, Yizhang Jiang, Kaifa Zhao, Rose Al Helo, Harry Friel, Atallah Baydoun, Feifei Zhou, Jin Uk Heo, Norbert Avril, Karin Herrmann, Rodney J. Ellis, Bryan J. Traughber, Robert S. Jones, Shitong Wang 0001, Kuan-Hao Su, Raymond F. Muzic Jr. |
IEEE Trans. Medical Imaging | 6 |
| 2018 | Abdominal, multi-organ, auto-contouring method for online adaptive magnetic resonance guided radiotherapy: An intelligent, multi-level fusion approach
Pengjiang Qian, Kuan-Hao Su, Atallah Baydoun, Asha Leisser, Steven Van Hedent, Jung-Wen Kuo, Kaifa Zhao, Parag Parikh, Yonggang Lu, Bryan J. Traughber, Raymond F. Muzic Jr. |
Artif. Intell. Medicine | 8 |
| 2017 | Knowledge-leveraged transfer fuzzy C-Means for texture image segmentation with self-adaptive cluster prototype matching
Pengjiang Qian, Kaifa Zhao, Yizhang Jiang, Kuan-Hao Su, Zhaohong Deng, Shitong Wang 0001, Raymond F. Muzic Jr. |
Knowl. Based Syst. | 2 |