VLDB 2026 Research / reviewers in the wild / expert
Haifang Zhou
dblp:26/2372
· DBLP profile ↗
23ranked-venue papers
4as first author
11since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLM-GCN BiTeach: LLM-Enhanced Bidirectional Mutual Teaching Framework for Log-Driven Software Fault Diagnosis
Rulin Xu, Haifang Zhou |
KSEM (4) | 5 |
| 2026 | Angel or devil: Discriminating hard samples and anomaly contaminations for unsupervised time series anomaly detection
Ruyi Zhang 0002, Hongzuo Xu, Songlei Jian, Yusong Tan, Haifang Zhou, Rulin Xu |
Neural Networks | 5 |
| 2024 | Accelerating Static Null Pointer Dereference Detection with Parallel ComputingabstractHigh-precision static analysis can effectively detect Null Pointer Dereference (NPD) vulnerabilities in C language, but the performance overhead is significant. In recent years, researchers have attempted to enhance the efficiency of static analysis by leveraging multicore resources. However, due to complex dependencies in the analysis process, the parallelization of static value-flow NPD analysis for large-scale software still faces significant challenges. It is difficult to achieve a good balance between detection efficiency and accuracy, which impacts its application.This paper presents PANDA, the first parallel detector for high-precision static value-flow NPD analyzer in the C language. The core idea of PANDA is to utilize dependency analysis to ensure high precision while decoupling the strong dependencies between static value-flow analysis steps. This transforms the traditionally challenging-to-parallelize NPD analysis into two parallelizable algorithms: function summarization and combined query-based vulnerability analysis. PANDA introduces a task-level parallel framework and enhances it with a dynamic scheduling method to parallel schedule the above two key steps, significantly improving the performance and scalability of memory vulnerability detection.Fully implemented within the LLVM framework (version 15.0.7), PANDA demonstrates a significant advantage in balancing accuracy and efficiency compared to current popular open-source detection tools. In precision-targeted benchmark tests, PANDA maintains a false positive rate within 3.17% and a false negative rate within 5.16%; in historical CVE detection rate tests, its recall rate far exceeds that of comparative open-source tools. In performance evaluations, compared to its serial version, PANDA achieves up to an 11.23-fold speedup on a 16-node server, exhibiting outstanding scalability. Rulin Xu, Luohui Chen, Ruyi Zhang 0002, Yuanliang Zhang, Haifang Zhou, Xiaoguang Mao |
Internetware | 6 |
| 2024 | X-EDF: An Efficient Defensive Deception Framework against Reconnaissance AttacksabstractDeception techniques are increasingly recognized as trans-formative in the realm of cyber defense. With the advent of sophisticated, large-scale scanning technologies such as ZMap, attackers can swiftly pinpoint active and vulnerable ports on edge nodes. Given the diversity of these nodes, a versatile security tool adaptable to various deployment environments is essential. Moreover, edge nodes often encounter performance constraints, necessitating a defense strategy that balances cost-effectiveness for defenders. In response to these challenges, we introduce the X-EDF: an eXpress Data Path (XDP)-based Efficient Defensive De-ception Framework. This framework facilitates an efficient and lightweight deceptive defense leveraging XDP technology. The X-EDF can efficiently respond to attackers' scanning requests with deceptive messages before these requests enter the protocol stack, thus achieving deception defense at a minimal cost. We have validated the effectiveness of our defense strategy through game-theoretic proofs and real-world network deployments. Zhihang Zhang, Chenlin Huang, Yan Ding 0004, Jinzhu Kong, Qing Liao 0001, Pan Dong, Haifang Zhou |
MSN | 7 |
| 2024 | ARISE: Graph Anomaly Detection on Attributed Networks via Substructure AwarenessabstractRecently, graph anomaly detection on attributed networks has attracted growing attention in data mining and machine learning communities. Apart from attribute anomalies, graph anomaly detection also aims at suspicious topological-abnormal nodes that exhibit collective anomalous behavior. Closely connected uncorrelated node groups form uncommonly dense substructures in the network. However, existing methods overlook that the topology anomaly detection performance can be improved by recognizing such a collective pattern. To this end, we propose a new graph anomaly detection framework on attributed networks via substructure awareness (ARISE). Unlike previous algorithms, we focus on the substructures in the graph to discern abnormalities. Specifically, we establish a region proposal module to discover high-density substructures in the network as suspicious regions. The average node-pair similarity can be regarded as the topology anomaly degree of nodes within substructures. Generally, the lower the similarity, the higher the probability that internal nodes are topology anomalies. To distill better embeddings of node attributes, we further introduce a graph contrastive learning scheme, which observes attribute anomalies in the meantime. In this way, ARISE can detect both topology and attribute anomalies. Ultimately, extensive experiments on benchmark datasets show that ARISE greatly improves detection performance (up to 7.30% AUC and 17.46% AUPRC gains) compared to state-of-the-art attributed networks anomaly detection (ANAD) algorithms. Jingcan Duan, Bin Xiao 0002, Siwei Wang 0001, Haifang Zhou, Xinwang Liu 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Towards Better Multilingual Code Search through Cross-Lingual Contrastive LearningabstractRecent advances in deep learning have significantly improved the understanding of source code by leveraging large amounts of open-source software data. Thanks to the larger amount of data, code representation models trained with multilingual datasets show superior performance to monolingual models and attract much more attention. However, the entangled source code from various programming languages makes multilingual models hard to differentiate language-specific textual semantics or syntactic structures, which significantly increases the difficulty of model learning from multilingual datasets directly. On the other hand, for a given problem, developers are likely to choose similar identifiers, even if coding in different languages. However, the presence of similar identifiers in multilingual code snippets does not mean that they implement the same functionality, which may misdirect models to overemphasize these unreliable signals and ignore the semantic information of multilingual code. To tackle the above issues, we propose LAMCode, a language-aware multilingual code understanding model. Specifically, we propose a simple yet effective method to perceive linguistic information by injecting language-specific viewer into the language models. Furthermore, we introduce a cross-lingual contrastive learning method by generating more similar training instances but with fewer overlapping features. This method prevents the models from over-relying on similar identifiers across languages. We conduct extensive experiments to evaluate the effectiveness of our approach on a large-scale multilingual dataset. The experimental results show that our approach significantly outperforms the state-of-the-art methods. Xiangbing Huang, Yingwei Ma, Haifang Zhou, Zhijie Jiang, Yuanliang Zhang, Teng Wang 0004, Shanshan Li 0001 |
Internetware | 3 |
| 2023 | A Two-Stage Framework for Ambiguous Classification in Software EngineeringabstractClassification tasks are prevalent and play a crucial role in the field of software engineering. However, when two classes exhibit similar features at the class level, the classification model is prone to misclassification, which we refer to as ambiguous classification, and the corresponding classes as ambiguous classes. Ambiguous classification may impact the security and reliability of software engineering classification systems.To correct ambiguous classification, we propose a two-stage framework. Our key insight is to combine two different classification models and utilize their complementary knowledge to maximize the classification ability of the two-stage framework. Specifically, we identify ambiguous classes according to the confusion matrix of the original model. Then, we construct a two-stage model, where the first stage utilizes the original model and the second stage utilizes a different model trained on the same dataset. The second-stage model is responsible for reclassifying the samples that are predicted as ambiguous classes by the first-stage model. We evaluate our method on two software engineering tasks. Experimental results indicate that our method can effectively correct ambiguous classification and achieve a relative improvement of 19.8% in F1-score for ambiguous classes. Yan Lei 0005, Shanshan Li 0001, Haifang Zhou, Yue Yu 0001, Zhouyang Jia, Yingwei Ma, Teng Wang 0004 |
ISSRE | 4 |
| 2023 | Normality Learning-based Graph Anomaly Detection via Multi-Scale Contrastive LearningabstractGraph anomaly detection (GAD) has attracted increasing attention in machine learning and data mining. Recent works have mainly focused on how to capture richer information to improve the quality of node embeddings for GAD. Despite their significant advances in detection performance, there is still a relative dearth of research on the properties of the task. GAD aims to discern the anomalies that deviate from most nodes. However, the model is prone to learn the pattern of normal samples which make up the majority of samples. Meanwhile, anomalies can be easily detected when their behaviors differ from normality. Therefore, the performance can be further improved by enhancing the ability to learn the normal pattern. To this end, we propose a normality learning-based GAD framework via multi-scale contrastive learning networks (NLGAD for abbreviation). Specifically, we first initialize the model with the contrastive networks on different scales. To provide sufficient and reliable normal nodes for normality learning, we design an effective hybrid strategy for normality selection. Finally, the model is refined with the only input of reliable normal nodes and learns a more accurate estimate of normality so that anomalous nodes can be more easily distinguished. Eventually, extensive experiments on six benchmark graph datasets demonstrate the effectiveness of our normality learning-based scheme on GAD. Notably, the proposed algorithm improves the detection performance (up to 5.89% AUC gain) compared with the state-of-the-art methods. The source code is released at https://github.com/FelixDJC/NLGAD. Jingcan Duan, Pei Zhang 0008, Siwei Wang 0001, Jingtao Hu, Hu Jin 0005, Jiaxin Zhang 0030, Haifang Zhou, Xinwang Liu 0002 |
ACM Multimedia | 7 |
| 2022 | DPSS: Dynamic Parameter Selection for Outlier Detection on Data StreamsabstractOutlier detection on data streams identifies unusual states to sense and alarm potential risks and faults of the target systems in both the cyber and physical world. As different parameter settings of machine learning algorithms can result in dramatically different performance, automatic parameter selection is also of great importance in deploying outlier detection algorithms in data streams. However, current canonical parameter selection methods suffer from two key challenges: (i) Data streams generally evolve over time, but these existing methods use a fixed training set, which fails to handle this evolving environment and often results in suboptimal parameter recommendations; (ii) The stream is infinite, and thus any parameter selection method taking the entire stream as input is infeasible. In light of these limitations, this paper introduces a Dynamic Parameter Selection method for outlier detection on data Streams (DPSS for short). DPSS uses Gaussian process regression to model the relationship between parameters and detecting performance and uses Bayesian optimization to explore the optimal parameter setting. For each new subsequence, DPSS updates the recommended parameter setting to suit the evolving characteristics. Besides, DPSS only uses historical calculations to guide the parameter setting sampling and adjust the Gaussian process regression results. DPSS can be employed as an auxiliary plug-in tool to improve the detection performance of outlier detection methods. Extensive experiments show that our method can significantly improve the F-score of outlier detectors in data streams compared to its counterparts and obtains more superior parameter selection performance than other state-of the-art parameter selection approaches. DPSS also achieves better time and memory efficiency compared to competitors. Ruyi Zhang 0002, Yijie Wang 0001, Haifang Zhou, Bin Li 0030, Hongzuo Xu |
ICPADS | 3 |
| 2022 | Factorization Machine-based Unsupervised Model Selection Method*abstractMachine learning is broadly used in many intelligent cybernetic systems. With the burgeoning of the communities of AI, the number of machine learning-based models is rapidly increasing, but picking a suitable and optimal (or relatively good) model from overwhelming options has become a conundrum when deploying a new system. Therefore, we are motivated by an intriguing question: Can we automatically select a proper model for new data? However, unsupervised model selection poses two main challenges: (i) Evaluation and comparison of candidate models on the new data are infeasible due to the lack of labels; and (ii) It is non-trivial to build relationships between model performance and data characteristics when the interaction between these characteristics should be considered. In light of these limitations, this paper proposes a factorization machine-based unsupervised model selection method. Following mainstream model selection protocols, we also leverage model performance on prior known datasets. Differently, we learn higher-order complex relationships between model performance and dataset characteristics. Specifically, our method transfers the historical performance into a second-order function of meta-features and embedding weights by harnessing the power of factorization machine. This function can be subsequently used to select a proper model when given a new dataset. Extensive experiments show that our method obtains more superior model selection performance than five state-of-the-art approaches, and our method executes faster than its competitors by approximate three magnitudes. Ruyi Zhang 0002, Yijie Wang 0001, Hongzuo Xu, Haifang Zhou |
SMC | 4 |
| 2021 | Visualization-based disentanglement of latent space
Runze Huang, Qianying Zheng, Haifang Zhou |
Neural Comput. Appl. | 3 |
| 2019 | An Analytical Study on a Benchmark Corpus Constructed for Related Work Generation
Pancheng Wang, Shasha Li 0001, Haifang Zhou, Jintao Tang, Ting Wang 0009 |
NLPCC (1) | 3 |
| 2018 | Benchmarking the GPU memory at the warp level
Minquan Fang, Jianbin Fang, Haifang Zhou, Jianxing Liao, Yuangang Wang |
Parallel Comput. | 4 |
| 2008 | Static Analysis for Application-Level Checkpointing of MPI ProgramsabstractApplication-level checkpointing is a promising technology in the domain of large-scale scientific computing. The consistency of global checkpoint must be carefully guaranteed in order to correctly restore the computation. Usually, some complex coordinated protocols are employed to ensure the consistency of global checkpoint, which require logging orphan or in-transit messages during checkpointing. These protocols complicate the recovery of the computation and increase the checkpoint overhead due to logging message. In this paper, a new method which ensures the consistency of global checkpoint by static analysis is proposed. The method identifies the safe checkpointing regions in MPI programs, where the global checkpoint is always strongly consistent. All checkpoints are located in those safe checkpoint regions. During checkpointing, the method will not log any messages and introduce no extra overhead. The method was implemented and integrated into ALEC, which is a source-to-source precompiler for automating application-level checkpointing. The experimental results show that our method is effective. Panfeng Wang, Yunfei Du 0001, Hongyi Fu, Xuejun Yang, Haifang Zhou |
HPCC | 5 |
| 2008 | Optimal Placement of Application-Level CheckpointsabstractOne of the basic problems related to the efficient application-level checkpointing is the placement of checkpoints in the source codes. In this paper we discuss two common questions with a source-to-source precompiler ALEC: 1) if there are N checkpoints in the application's source code, how to pick M checkpoints out of them minimizing the total amount of checkpoint data? 2) if there are no checkpoint in the application's source code, how to insert a set of checkpoints minimizing the amount of checkpoint data? We reveal that these two questions can both be abstracted as a mathematic model which is similar to the 0-1 integer programming model, and the model can be solved using implicit enumeration method. The solving methods proposed in the paper have been implemented and integrated into ALEC. Experimental results show that the method is efficient. Panfeng Wang, Yunfei Du 0001, Xuejun Yang, Haifang Zhou |
HPCC | 5 |
| 2007 | A Novel Fault-Tolerant Parallel Algorithm
Panfeng Wang, Hongyi Fu, Haifang Zhou, Xuejun Yang |
APPT | 4 |
| 2007 | A data-distributed parallel algorithm for wavelet-based fusion of remote sensing images
Xuejun Yang, Panfeng Wang, Yunfei Du 0001, Haifang Zhou |
Frontiers Comput. Sci. China | 4 |
| 2007 | Compiler-directed power optimization of high-performance interconnection networks for load-balancing MPI applications
Xuejun Yang, Huizhan Yi, Xiangli Qu, Haifang Zhou |
Frontiers Comput. Sci. China | 4 |
| 2006 | A Parallel Mutual Information Based Image Registration Algorithm for Applications in Remote Sensing
Haifang Zhou, Panfeng Wang, Xuejun Yang, Hengzhu Liu |
ISPA | 2 |
| 2006 | GPGC: a Grid-enabled parallel algorithm of geometric correction for remote-sensing applicationsabstractAbstract ChinaGrid is an important project sponsored by the China Ministry of Education, aiming to provide high‐performance services in a Grid computing environment. In this paper, one of the applications offered by ChinaGrid, parallel remote‐sensing image processing, is described. Geometric correction is a basic step during the processing of remote‐sensing imagery, which is traditionally a computation‐intensive and communication‐intensive application if in parallel mode. In order to move this application into a Grid, a new Grid‐enabled parallel algorithm of geometric correction is proposed, called GPGC. GPGC changes the frequent and fine‐grain communication mode of the existing parallel method into a delayed but concentrated exchanging mode by computing an irregular local output area. This change means no communication or synchronization happens during resampling that occupies most of the execution time. To prove its efficiency, the complexity of GPGC is analyzed in theory. Finally, performance testing of GPGC and its application in ChinaGrid are given. Experimental results show that our algorithm is more suitable for a Grid platform, excelling the old method in both performance and salability. Copyright © 2006 John Wiley & Sons, Ltd. Haifang Zhou, Xuejun Yang, Hengzhu Liu, Yu Tang 0014 |
Concurr. Comput. Pract. Exp. | 1 |
| 2005 | A Proposal of Parallel Strategy for Global Wavelet-Based Registration of Remote-Sensing Images
Haifang Zhou, Yu Tang 0014, Xuejun Yang, Hengzhu Liu |
ICA3PP | 1 |
| 2005 | First Evaluation of Parallel Methods of Automatic Global Image Registration Based on WaveletsabstractWith the increasing importance of multiple multiplatform remote sensing missions, fast and automatic integration of digital data from disparate sources has become critical to the success of these endeavors. Firstly, an overview of development of automatic and parallel global image registration is given. And then, based on the analyses of existing three parallel methods of wavelet-based global registration, a new parallel strategy is proposed. Moreover, towards the quantitative evaluation, first results of the intercomparision of four parallel global registration algorithms are presented in theory and in experiments. Haifang Zhou, Xuejun Yang, Hengzhu Liu, Yu Tang 0014 |
ICPP | 1 |
| 2004 | Further Optimized Parallel Algorithm of Watershed Segmentation Based on Boundary Components Graph
Haifang Zhou, Xuejun Yang, Yu Tang 0014, Nong Xiao 0001 |
NPC | 1 |