Zhenghao He

dblp:130/6433 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Concept-RuleNet: Grounded Multi-Agent Neurosymbolic Reasoning in Vision Language Models
abstract
Modern vision-language models (VLMs) deliver impressive predictive accuracy yet offer little insight into 'why' a decision is reached, frequently hallucinating facts, particularly when encountering out-of-distribution data. Neurosymbolic frameworks address this by pairing black-box perception with interpretable symbolic reasoning, but current methods extract their symbols solely from task labels, leaving them weakly grounded in the underlying visual data. In this paper, we introduce a multi-agent system - Concept-RuleNet that reinstates visual grounding while retaining transparent reasoning. Specifically, a multimodal concept generator first mines discriminative visual concepts directly from a representative subset of training images. Next, these visual concepts are utilized to condition symbol discovery, anchoring the generations in real image statistics and mitigating label bias. Subsequently, symbols are composed into executable first-order rules by a large language model reasoner agent - yielding interpretable neurosymbolic rules. Finally, during inference, a vision verifier agent quantifies the degree of presence of each symbol and triggers rule execution in tandem with outputs of black-box neural models, predictions with explicit reasoning pathways. Experiments on five benchmarks, including two challenging medical-imaging tasks and three underrepresented natural-image datasets, show that our system augments state-of-the-art neurosymbolic baselines by an average of 5% while also reducing the occurrence of hallucinated symbols in rules by up to 50%.
Sanchit Sinha, Guangzhi Xiong, Zhenghao He, Aidong Zhang 0001
AAAI3
2026 PCRepair: A Context-Aware Template-Based Approach for Automated Program Repair
abstract
Automated Program Repair (APR) is increasingly vital for managing the complexity of modern software systems. However, current APR techniques suffer from inefficiently selecting repair components, resulting in suboptimal patches. To address these limitations, we propose PCRepair, a context-aware template-based methodology for automated software fault repair. This approach integrates predefined repair templates with context-aware analysis to improve repair accuracy and efficiency. PCRepair first localizes suspicious statements via the Ochiai technique, then matches their contextual patterns with relevant templates. This strategy narrows the search space and generates semantically relevant candidate patches. We prioritize these patches using a weighted fusion similarity metric and sequentially validate them against existing test cases. Evaluations on the Defects4J benchmark show that PCRepair successfully repaired 38 defects, demonstrating competitive performance compared to existing methods, particularly in terms of repair efficiency, with a 9.62% success rate and an average repair time of fewer than 30 min per defect.
Heling Cao, Yun Wang 0009, Yonghe Chu, Miaolei Deng, Zhenghao He
Int. J. Softw. Eng. Knowl. Eng.6
2025 GCAV: A Global Concept Activation Vector Framework for Cross-Layer Consistency in Interpretability
abstract
Concept Activation Vectors (CAVs) provide a powerful approach for interpreting deep neural networks by quantifying their sensitivity to human-defined concepts. However, when computed independently at different layers, CAVs often exhibit inconsistencies, making cross-layer comparisons unreliable. To address this issue, we propose the Global Concept Activation Vector (GCAV), a novel framework that unifies CAVs into a single, semantically consistent representation. Our method leverages contrastive learning to align concept representations across layers and employs an attention-based fusion mechanism to construct a globally integrated CAV. By doing so, our method significantly reduces the variance in TCAV scores while preserving concept relevance, ensuring more stable and reliable concept attributions. To evaluate the effectiveness of GCAV, we introduce Testing with Global Concept Activation Vectors (TGCAV) as a method to apply TCAV to GCAV-based representations. We conduct extensive experiments on multiple deep neural networks, demonstrating that our method effectively mitigates concept inconsistency across layers, enhances concept localization, and improves robustness against adversarial perturbations. By integrating cross-layer information into a coherent framework, our method offers a more comprehensive and interpretable understanding of how deep learning models encode human-defined concepts. Code and models are available at https://github.com/Zhenghao-He/GCAV.
Zhenghao He, Sanchit Sinha, Guangzhi Xiong, Aidong Zhang 0001
ICCV1
2025 RESEARCH NOTES - GMRepair: Graph Mining Template-Based Automated Software Repair
abstract
With the increasing scale and complexity of software recently, automated software bug repair has grown in importance. However, the current automated software bug repair process suffers from issues such as coarse-grained repair granularity and poor patch quality. To address these problems, we propose a graph mining template-based automatic software repair (GMRepair) to improve the performance of automated software bug repair. First, this approach adopts the Ochiai fault localization technique to locate and generate a list of suspicious defect statements. We utilize the GumTree tool to parse the bug and repair program files, generating edit scripts. These edit scripts are then transformed into a graphical representation. Second, we utilize a frequent graph miner to obtain graph mining templates by matching the context of the suspicious statements with the context of the graph mining templates, generating an initial population for them. The buggy program is evolved using genetic programming through mutation and crossover operations, generating new individuals. Finally, we sequentially pass the candidate patches (CPs) through corresponding test cases and prioritize the test cases using priority sorting techniques. Patches that fail to pass the test cases are filtered out, and the patches that pass the test cases are output. We conducted the experiments using two datasets, QuixBugs and Defects4J. In Defects4J, the GMRepair successfully repaired 41 defects, while in QuixBugs, it successfully repaired 15 defects. Compared to the existing methods, GMRepair offers a higher success rate and efficiency in defect repair.
Heling Cao, Yanlong Guo, Yun Wang 0009, Fangchao Tian, Yonghe Chu, Miaolei Deng, Zhenghao He, Shuting Wei
Int. J. Softw. Eng. Knowl. Eng.9
2025 MUSO: achieving exact machine unlearning in over-parameterized regimes
Ruikai Yang, Mingzhen He, Zhenghao He, Youmei Qiu, Xiaolin Huang
Mach. Learn.3
2022 A 3-D Storm Motion Estimation Method Based on Point Cloud Learning and Doppler Weather Radar Data
abstract
In recent years, deep learning techniques have been developed in the field of storm nowcasting, and they primarily focus on 2-D radar product processing. Deep learning’s application to obtaining motion fields is limited by the lack of a motion field ground truth dataset for training. In this article, we propose a method for storm motion estimation based on a point cloud deep learning network, which can perform 3-D motion field estimation on two consecutive frames of weather radar volumetric data. To address the lack of ground truth data, we propose a synthetic data generation method based on Perlin noise and also generate a dataset to train the network. Our model’s architecture is based on FlowNet3D. In this sense, we import an attention architecture to improve its ability to embed motion features. We also redesign the loss function to make it suitable for our task. In the experiment section, we first evaluate the performance of the trained model on our synthetic dataset. The experimental results demonstrate the following: first, that the trained network has the ability to estimate 3-D motion from radar point cloud data; second, that it is a robust solution for radar data with different resolutions and observation ranges; and finally, that for the examined task, our proposed architecture has a 5%–15% lower endpoint error than the original. Then, we directly apply the model trained on synthetic data to real cases, demonstrating the feasibility of our trained model on real data. Although motion field is a low-level product when considering a storm’s evolutionary process, it has the potential to serve a higher-level purpose, such as lightning or hail nowcasting. The point cloud approach provides a new perspective on weather radar data processing.
Zhenghao He, Riyang Bao, Shuping Gao
IEEE Trans. Geosci. Remote. Sens.2
2013 Unsupervised Medical Subject Heading Assignment Using Output Label Co-occurrence Statistics and Semantic Predications
Ramakanth Kavuluru, Zhenghao He
NLDB2