VLDB 2026 Research / reviewers in the wild / expert
Haize Hu
dblp:318/6604
· DBLP profile ↗
24ranked-venue papers
9as first author
24since 2021 · last 2026
0000-0002-1706-9691ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FNCD-CS: A neurocomputing-optimized hybrid neural network for cross-domain code search
Mengge Fang, Haize Hu |
Neurocomputing | 3 |
| 2026 | Neural network-based dynamic adaptation and multimodal fusion for class-imbalanced educational data processing: A systematic review
Mengge Fang, Li-e Wang 0001, Haize Hu |
Neurocomputing | 3 |
| 2026 | Neural network-driven code semantic alignment for software engineering: Three-perspective framework, three-dimensional evaluation, and industrial adaptation
Haize Hu, Ziqi Zhang 0019, Jingli Wu, Naixue Xiong, Jerome Yen |
Neurocomputing | 1 |
| 2026 | CDMI-NTDI: Cancer driver module identification via network topology and deep interaction features
Jingli Wu, Yanhua Huang, Gaoshi Li, Jiafei Liu 0001, Haize Hu |
Neurocomputing | 5 |
| 2026 | DEGAN-CS: An efficient code search model based on dataenhanced optimization of generative adversarial networks
Haize Hu, Jianxun Liu 0001, Mengge Fang |
Pattern Recognit. | 1 |
| 2026 | HFGCS: Industrial Code Search With Sample-Aware Hierarchical Fusion and Hub-Centric Heterogeneous Graph Reasoning for Reliable CPS Software MaintenanceabstractIn industrial cyber-physical systems (CPS), maintaining large-scale, long-lived software stacks demands rapid retrieval of protocol-compatible and system-context-aware code segments to avoid costly downtime and safety issues. Existing code search methods lack system-aware reasoning and suffer from rigid semantic fusion, leading to fragmented understanding of intercomponent dependencies in industrial software. To tackle these pain points, this article proposes an industrial-oriented hierarchical feature and graph-based code search (HFGCS) framework with sample-aware hierarchical fusion and hub-centric heterogeneous graph reasoning, customized for CPS maintenance requirements. It integrates a dynamic layer aggregation (DLA) module for adaptive multigranularity semantic fusion (capturing syntax-to-logic features based on query intent) and an integrated hub context encoder (IHCE) module that constructs a heterogeneous CPS graph with a global hub node to propagate cross-component dependencies (control logic, communication protocols, sensor/actuator bindings, and runtime constraints). A learnable gating network dynamically balances these representations to achieve intent-aligned and system-consistent code retrieval. Extensive experiments on CodeSearchNet-C show HFGCS improves MRR by 7.3%–15.9% over state-of-the-art baselines with strong cross-backbone robustness. Industrial validation confirms it shortens development cycles, enhances efficiency and accuracy, and prevents protocol mismatches in safety-critical CPS software. Feasible for deployment with topology-aware reasoning, HFGCS serves as a scalable and reliable retrieval engine, providing high-quality compatible code references to underpin the stable operation of industrial CPS and aligning with the core scope of Industrial Informatics. The code and data used in our study are available at online. Haize Hu, Ziqi Zhang 0019, Jingli Wu, Naixue Xiong, Mengge Fang, Jerome Yen |
IEEE Trans. Ind. Informatics | 1 |
| 2025 | A code completion approach combining pointer network and Transformer-XL network
Haize Hu |
Appl. Intell. | 4 |
| 2025 | A Survey of the Full Process of Code Search Based on Deep LearningabstractABSTRACT As a pivotal technology for enhancing software development efficiency, research on code search based on deep learning has emerged as a current hotspot. This review systematically deconstructs the entire code search process into four core stages: dataset construction, code preprocessing, heterogeneous representation model construction, and query expansion, while conducting an in‐depth analysis of the application status and challenges of deep learning technologies. In dataset construction, the Q and A pairs and C and D pairs relied on by deep learning models suffer from a lack of standardization. For example, CodeSearchNet exhibits insufficient cross‐lingual versatility, and CoDesc has incomplete noise filtering. During the code preprocessing stage, bottlenecks such as AST granularity selection and sequence information redundancy restrict the efficiency of feature extraction. Although the introduction of transformer and graph neural networks has optimized structural representation, a unified evaluation mechanism is lacking. In the research of heterogeneous representation models, while LSTM, CNN, and pretrained models (such as CodeBERT) effectively narrow the semantic gap, their cross‐domain search accuracy is insufficient. In terms of query expansion, deep learning‐based keyword expansion and intent completion methods struggle to capture users' real needs due to low semantic alignment accuracy. This review proposes, for the first time, a standardized dataset construction framework integrating multimodal data, a syntax‐semantic dual‐layer preprocessing evaluation mechanism, a cross‐domain transfer representation model, and a large language model‐driven intent dynamic expansion scheme. These contributions lay a theoretical foundation for the systematic development of code search technologies and provide cross‐task methodological references for related fields such as code clone detection. Mengge Fang, Haize Hu, Feiyu Hu |
Concurr. Comput. Pract. Exp. | 2 |
| 2025 | An adaptive model for cross-domain code search
Mengge Fang, Li-e Wang 0001, Haize Hu |
Inf. Softw. Technol. | 3 |
| 2025 | An intent-enhanced feedback extension model for code search
Haize Hu, Mengge Fang |
Inf. Softw. Technol. | 1 |
| 2025 | Automatic Evaluation of English Translation Based on Multi-granularity Interaction FusionabstractThe latest neural machine translation automatic evaluation method uses pre-trained context word vectors to extract semantic features and directly concatenates them into the neural network to predict translation quality. However, the direct operation can easily lead to a lack of interaction between features, and the layer-by-layer prediction is prone to losing fine-grained matching information. To address these issues, we propose a multi-granularity interactive fusion English translation automatic evaluation, which introduces middle and late information fusion methods. First, we use a bilinear attention distribution to capture high-order cross language feature interactions. By stacking multiple high-order interaction blocks and equipping them with an index linear unit without parameters for middle fusion in a parameter-free manner. Second, we use fine-grained accurate matching sentence shift distance and sentence-level cosine similarity for late fusion. The experimental results on the WMT’21 Metrics Task benchmark dataset show that the proposed method can effectively improve its correlation with human evaluation and achieve comparable performance with the best participating system. Xibo Chen, Yonghe Yang, Haize Hu |
Neural Process. Lett. | 3 |
| 2025 | A Deep Learning Framework for Identifying Cancer Driver Genes Based on Transformer and Graph Convolutional NetworkabstractCorrect identification of cancer driver genes plays a significant role in cancer research. The advancement of graph neural network (GNN) research has led to the emergence of many high-performance cancer driver gene prediction methods. However, GNN-based methods frequently overlook the importance of capturing global information. Additionally, as GNN layers increase, the feature representation of genes begins to become overly smooth. These problems hinder the effectiveness of GNN-based identification methods. In this study, we introduce TGCN, a method integrating Transformer and graph convolutional network (GCN), aiming to address these issues and improve cancer driver gene identification. First, we composed multivariate feature matrices of genes from multi-omics data and multi-dimensional gene association networks. Second, we constructed a Transformer module to enrich gene feature representations. Finally, we utilized Chebyshev GCN to yield the identification results. The experimental results demonstrate that TGCN outperforms representative methods in identifying driver genes for both pan-cancer and single-type cancers. Gaoshi Li, Jingli Wu, Jiafei Liu 0001, Haize Hu, Qiyong Zhu |
IEEE Trans. Comput. Biol. Bioinform. | 8 |
| 2025 | Identifying Cancer Driver Genes Using a Neural Network Framework With Cross-Attention MechanismabstractIdentifying cancer driver genes can accelerate the discovery of drug targets and the development of cancer therapies. Recent research methods improve the accuracy of identifying cancer driver genes by using deep learning framework. However, due to ignore the connection among learned features, they usually have weak feature representations that limits further improvement in the accuracy of identifying cancer driver genes. In this work, we propose a graph neural network framework combining graph convolutional network, Transformer with cross-attention, and multi-layer perceptron classifier, called GTCM, to improve the accuracy of identifying cancer driver genes. Specifically, GTCM first uses graph convolutional network to learn gene feature representations from three different gene association networks. Second, to enhance the feature representations of cancer driver genes, GTCM adopts Transformer with cross-attention to dynamically learn the connections between different feature sets. Finally, GTCM predicts cancer driver genes using multi-layer perceptron classifier. Ablation experiments prove that Transformer with cross-attention effectively improves the feature representations learned from graph convolutional network and further improves the identification rate. Compared with existing representative methods, GTCM exhibits excellent performance in terms of area under the receiver operating characteristic curves and area under precision-recall curves. Gaoshi Li, Jingli Wu, Jiafei Liu 0001, Haize Hu, Qiyong Zhu |
IEEE Trans. Comput. Biol. Bioinform. | 8 |
| 2024 | HACS: An Enhancement Framework for Deep Code Search Benefiting from Hard Negative Samples (S)abstractCode search aims to retrieve relevant code snippets from large code repositories based on query, promoting code reuse and enhancing software development efficiency.Deep Learning is a powerful approach for code search, in which the hard negative samples within training batches critically impact model performance.However, most existing deep code search models only randomly sample negative samples, resulting in a paucity or complete lack of hard negative samples.To address this limitation, we introduce a novel enhancement framework named HACS to optimize the composition of negative samples within batches, thus enhancing the training effectiveness of deep code search models.The core idea is to increase the count of hard negative samples within the negative samples corresponding to each query in the training batch.Specifically, HACS utilizes deep reinforcement learning techniques for sampling hard negative samples and implements vector-level mixed data augmentation strategy to generate hard negative samples.We evaluated our framework on a public dataset covering six programming languages.Experimental results reveal that HACS significantly improves the code search performance of existing models. Qihong Song, Haize Hu |
SEKE | 3 |
| 2024 | Deep code search efficiency based on clusteringabstractAbstract The deep‐learning based code search model mainly takes accuracy as the only target for judging the performance of the model, ignoring the efficiency of code search. This article proposes a clustering‐based code search model (C‐DCS). C‐DCS uses the K‐Means to divide the code vector base into K clusters and obtains the center vectors of K clusters. While searching, C‐DCS first matches the query vector with the K center vectors to get the best matching center vector. After matching the center vector, C‐DCS matches the query vector with code vectors in the cluster corresponding to the best matching center vector one by one and then gets the best matching code snippet vector. To verify the efficiency of C‐DCS in the code search task, experimental analysis was built on a large dataset. The experimental results showed that C‐DCS saves 92.2% of the search time compared to the baseline model while remaining the accuracy. In the experimental evaluation section, we optimized the K‐Means algorithm to improve the code search efficiency of C‐DCS further, reducing the search time to 93.8% of the baseline model. Hence, C‐DCS reduces the code search time greatly with not affecting the accuracy, improving the efficiency of software development. Jianxun Liu 0001, Haize Hu |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | Enrich Code Search Query Semantics with Raw Descriptions
Xiangzheng Liu, Jianxun Liu 0001, Haize Hu |
CollaborateCom (1) | 3 |
| 2023 | CUTE: A Collaborative Fusion Representation-Based Fine-Tuning and Retrieval Framework for Code Search
Qihong Song, Haize Hu |
CollaborateCom (1) | 3 |
| 2023 | A Code Completion Approach Combining Pointer Network and Transformer-XL Network
Haize Hu |
CollaborateCom (1) | 4 |
| 2023 | Multi-intent Description of Keyword Expansion for Code Search
Haize Hu, Jianxun Liu 0001 |
ICONIP (11) | 1 |
| 2023 | A Multiple-Path Learning Neural Network Model for Code CompletionabstractCode completion, which can accelerate the software development process and improve the quality of software products, is an essential part of today’s integrated development environments. It has become an important research topic in the field of software engineering. Recent studies have shown that the method of code completion based on the Abstract Syntax Tree (AST) learns syntactic information about the code, which helps to improve the accuracy of code completion. However, when modeling neural networks for ASTs, the sequencing operation of nodes leads to the loss of their hierarchical structure information. Meanwhile, traditional neural networks cannot predict many Out-of-Vocabulary (OoV) words in the terminal node values of AST. To alleviate the above problem, in this paper, we propose a Multiple-Path Learning neural network model for code completion (MPL) based on an AST by learning from a large-scale corpus. In this model, multiple paths such as context path, root path, and terminal node path are established to understand different code features required for node prediction and improve code representation ability. Based on the principle of program local repeatability, it also adopts a replication mechanism to copy the appropriate OoV words from the local terminal node path as the prediction result, further improving the prediction accuracy. The experimental results show that the MPL model has better performance than existing methods on the code completion task. Jianxun Liu 0001, Haize Hu |
ICWS | 4 |
| 2023 | Multilayer self-attention residual network for code searchabstractSummary Software developers usually search existing code snippets in open source code repositories to modify and reuse them. Therefore, how to get the right code snippet from the open‐source code repository quickly and accurately is the focus of current software development research. Nowadays, code search is one of the solutions. To improve the accuracy of source code feature information representation and the accuracy of code search. A multilayer self‐ attention residual network‐based code search model (MSARN‐CS) is proposed in this paper. In the MSARN‐CS model, not only the weight of each word in the code sequence unit is considered but also the effect of embedding between code sequence units is calculated. In addition, an optimization model of residuals is introduced to compensate for the loss of information in the code sequences during the model training. To verify the search effectiveness of the MSARN‐CS model, three other baseline models are compared on the basis of extensive source code data. The experimental results show that the MSARN‐CS model has better search results compared with the baseline model. For parameter Recall@1, the experimental result of MSARN‐CS model was 9.547, which as 100.90%, 73.87%, 60.37%, and 2.55% better compared to CODEnn, CRLCS, SAN‐CS‐ and SAN‐CS, respectively. For the parameter Recall@5, the results improved by 26.67%, 36.23%, 36.21%, and 1.63%, respectively, and for the parameter Recall@10, the results improved by 13.92%, 25.70%, 20.78%, and 2.23%, respectively. For the parameter mean reciprocal rank, the results improved by 52.89%, 76.17%, 63.38%, and 3.88%, respectively. For the parameter normalized discounted cumulative gain, the results improved by 54.22%, 60.55%, 50.28%, and 3.30%, respectively. The MSARN‐CS model proposed in the paper can effectively improve the accuracy of code search and enhance the programming efficiency of developers. Haize Hu, Jianxun Liu 0001 |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | A mutual embedded self-attention network model for code search
Haize Hu, Jianxun Liu 0001, Ben Cao, Siqiang Cheng |
J. Syst. Softw. | 1 |
| 2023 | An Effective and Adaptable K-means Algorithm for Big Data Cluster Analysis
Haize Hu, Jianxun Liu 0001, Mengge Fang |
Pattern Recognit. | 1 |
| 2022 | A novel hybrid model for short-term prediction of wind speed
Haize Hu, Yunyi Li, Mengge Fang |
Pattern Recognit. | 1 |