VLDB 2026 Research / reviewers in the wild / expert
Mengge Fang
dblp:318/6592
· DBLP profile ↗
9ranked-venue papers
4as first author
9since 2021 · last 2026
0000-0002-0947-7054ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FNCD-CS: A neurocomputing-optimized hybrid neural network for cross-domain code search
Mengge Fang, Haize Hu |
Neurocomputing | 1 |
| 2026 | Neural network-based dynamic adaptation and multimodal fusion for class-imbalanced educational data processing: A systematic review
Mengge Fang, Li-e Wang 0001, Haize Hu |
Neurocomputing | 1 |
| 2026 | DEGAN-CS: An efficient code search model based on dataenhanced optimization of generative adversarial networks
Haize Hu, Jianxun Liu 0001, Mengge Fang |
Pattern Recognit. | 3 |
| 2026 | HFGCS: Industrial Code Search With Sample-Aware Hierarchical Fusion and Hub-Centric Heterogeneous Graph Reasoning for Reliable CPS Software MaintenanceabstractIn industrial cyber-physical systems (CPS), maintaining large-scale, long-lived software stacks demands rapid retrieval of protocol-compatible and system-context-aware code segments to avoid costly downtime and safety issues. Existing code search methods lack system-aware reasoning and suffer from rigid semantic fusion, leading to fragmented understanding of intercomponent dependencies in industrial software. To tackle these pain points, this article proposes an industrial-oriented hierarchical feature and graph-based code search (HFGCS) framework with sample-aware hierarchical fusion and hub-centric heterogeneous graph reasoning, customized for CPS maintenance requirements. It integrates a dynamic layer aggregation (DLA) module for adaptive multigranularity semantic fusion (capturing syntax-to-logic features based on query intent) and an integrated hub context encoder (IHCE) module that constructs a heterogeneous CPS graph with a global hub node to propagate cross-component dependencies (control logic, communication protocols, sensor/actuator bindings, and runtime constraints). A learnable gating network dynamically balances these representations to achieve intent-aligned and system-consistent code retrieval. Extensive experiments on CodeSearchNet-C show HFGCS improves MRR by 7.3%–15.9% over state-of-the-art baselines with strong cross-backbone robustness. Industrial validation confirms it shortens development cycles, enhances efficiency and accuracy, and prevents protocol mismatches in safety-critical CPS software. Feasible for deployment with topology-aware reasoning, HFGCS serves as a scalable and reliable retrieval engine, providing high-quality compatible code references to underpin the stable operation of industrial CPS and aligning with the core scope of Industrial Informatics. The code and data used in our study are available at online. Haize Hu, Ziqi Zhang 0019, Jingli Wu, Naixue Xiong, Mengge Fang, Jerome Yen |
IEEE Trans. Ind. Informatics | 5 |
| 2025 | A Survey of the Full Process of Code Search Based on Deep LearningabstractABSTRACT As a pivotal technology for enhancing software development efficiency, research on code search based on deep learning has emerged as a current hotspot. This review systematically deconstructs the entire code search process into four core stages: dataset construction, code preprocessing, heterogeneous representation model construction, and query expansion, while conducting an in‐depth analysis of the application status and challenges of deep learning technologies. In dataset construction, the Q and A pairs and C and D pairs relied on by deep learning models suffer from a lack of standardization. For example, CodeSearchNet exhibits insufficient cross‐lingual versatility, and CoDesc has incomplete noise filtering. During the code preprocessing stage, bottlenecks such as AST granularity selection and sequence information redundancy restrict the efficiency of feature extraction. Although the introduction of transformer and graph neural networks has optimized structural representation, a unified evaluation mechanism is lacking. In the research of heterogeneous representation models, while LSTM, CNN, and pretrained models (such as CodeBERT) effectively narrow the semantic gap, their cross‐domain search accuracy is insufficient. In terms of query expansion, deep learning‐based keyword expansion and intent completion methods struggle to capture users' real needs due to low semantic alignment accuracy. This review proposes, for the first time, a standardized dataset construction framework integrating multimodal data, a syntax‐semantic dual‐layer preprocessing evaluation mechanism, a cross‐domain transfer representation model, and a large language model‐driven intent dynamic expansion scheme. These contributions lay a theoretical foundation for the systematic development of code search technologies and provide cross‐task methodological references for related fields such as code clone detection. Mengge Fang, Haize Hu, Feiyu Hu |
Concurr. Comput. Pract. Exp. | 1 |
| 2025 | An adaptive model for cross-domain code search
Mengge Fang, Li-e Wang 0001, Haize Hu |
Inf. Softw. Technol. | 1 |
| 2025 | An intent-enhanced feedback extension model for code search
Haize Hu, Mengge Fang |
Inf. Softw. Technol. | 2 |
| 2023 | An Effective and Adaptable K-means Algorithm for Big Data Cluster Analysis
Haize Hu, Jianxun Liu 0001, Mengge Fang |
Pattern Recognit. | 4 |
| 2022 | A novel hybrid model for short-term prediction of wind speed
Haize Hu, Yunyi Li, Mengge Fang |
Pattern Recognit. | 4 |