VLDB 2026 Research / reviewers in the wild / expert
He Zhao 0002
dblp:98/3487-2
· DBLP profile ↗
26ranked-venue papers
8as first author
21since 2021 · last 2026
0000-0002-8264-9297ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 2 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Beyond Correlation: Causal Intervention for Multi-Label Medical Image DiagnosisabstractThis paper addresses the challenge of multi-disease diagnosis by integrating causal reasoning into the diagnostic framework. In clinical practice, multiple conditions often co-occur, making multi-disease diagnosis more relevant than isolated single-disease cases. However, most deep learning methods focus on single-disease detection and fail to capture the complexity of diagnosing concurrent conditions. Even in multi-label settings, existing approaches mainly rely on correlation-based inference, capturing statistical associations rather than true causal relationships. This can lead to spurious feature-disease associations, where features linked to one disease are mistakenly attributed to another due to frequent co-occurrence, ultimately undermines diagnostic accuracy and interpretability. To address this challenge, we propose a novel framework that incorporates causal intervention into multi-label medical image diagnosis, enabling the model to identify true causal signals rather than misleading correlations arising from co-occurring diseases. Specifically, we model latent disease-related confounders and apply backdoor adjustment to disentangle genuine causal effects from spurious associations. This is achieved by implicitly learning shared feature representations that serve as confounding variables, which are then used to refine image-derived features during prediction. The resulting causal adjustment allows the model to focus on disease-specific cues, improving accuracy and interpretability. Extensive experiments on four diverse medical imaging datasets: ODIR (color fundus photography), LID-FFA (fundus fluorescein angiography), Endo (colonoscopy), and Chestpert (X-ray) demonstrate that our method consistently outperforms existing approaches. Furthermore, our model also effectively separates the diagnosis of co-occurring diseases, highlighting the potential of causal reasoning to enhance the reliability and clinical applicability of AI-assisted diagnosis. The source code is publicly available at https://github.com/davelailai/BankCausal.git. Jianyang Xie, Yitian Zhao, Xiuju Chen, Yanda Meng, He Zhao 0002, Uazman Alam, Yalin Zheng |
IEEE Trans. Medical Imaging | 5 |
| 2025 | MCAT: Visual Query-Based Localization of Standard Anatomical Clips in Fetal Ultrasound Videos Using Multi-Tier Class-Aware Token TransformerabstractAccurate standard plane acquisition in fetal ultrasound (US) videos is crucial for fetal growth assessment, anomaly detection, and adherence to clinical guidelines. However, manually selecting standard frames is time-consuming and prone to intra- and inter-sonographer variability. Existing methods primarily rely on image-based approaches that capture standard frames and then classify the input frames across different anatomies. This ignores the dynamic nature of video acquisition and its interpretation. To address these challenges, we introduce Multi-Tier Class-Aware Token Transformer (MCAT); a visual query-based video clip localization (VQ-VCL) method to assist sonographers by enabling them to capture a quick US sweep. By then providing a visual query of the anatomy they wish to analyze, MCAT returns the video clip containing the standard frames for that anatomy, facilitating thorough screening for potential anomalies. We evaluate MCAT on two ultrasound video datasets and a natural image VQ-VCL dataset based on Ego4D. Our model outperforms state-of-the-art methods by 10% and 13% mtIoU on the ultrasound datasets and by 5.35% mtIoU on the Ego4D dataset, using 96% fewer tokens. MCAT’s efficiency and accuracy have significant potential implications for public health, especially in low- and middle-income countries (LMICs), where it may enhance prenatal care by streamlining standard plane acquisition, simplifying US based screening, diagnosis and allowing sonographers to examine more patients. Divyanshu Mishra, Pramit Saha, He Zhao 0002, Netzahualcóyotl Hernández, Olga Patey, Aris T. Papageorghiou, J. Alison Noble |
AAAI | 3 |
| 2025 | Are Spatial-Temporal Graph Convolution Networks for Human Action Recognition Over-Parameterized?abstractSpatial-temporal graph convolutional networks (ST-GCNs) showcase impressive performance in skeleton-based human action recognition (HAR). However, despite the development of numerous models, their recognition performance does not differ significantly after aligning the input settings. With this observation, we hypothesize that ST-GCNs are over-parameterized for HAR, a conjecture subsequently confirmed through experiments employing the lottery ticket hypothesis. Additionally, a novel sparse ST-GCNs generator is proposed, which trains a sparse architecture from a randomly initialized dense network while maintaining comparable performance levels to the dense components. Moreover, we generate multi-level sparsity ST-GCNs by integrating sparse structures at various sparsity levels and demonstrate that the assembled model yields a significant enhancement in HAR performance. Thorough experiments on four datasets, including NTU-RGB+D 60(120), Kinetics-400, and FineGYM, demonstrate that the proposed sparse ST-GCNs can achieve comparable performance to their dense components. Even with 95% fewer parameters, the sparse ST-GCNs exhibit a degradation of1% in top-1 accuracy. The code is available at https://github.com/davelailai/Sparse-ST-GCN. Jianyang Xie, Yitian Zhao, Yanda Meng, He Zhao 0002, Anh Nguyen 0003, Yalin Zheng |
CVPR | 4 |
| 2025 | tHPM-LDM: Integrating Individual Historical Record with Population Memory in Latent Diffusion-Based Glaucoma Forecasting
Jianyang Xie, Yimin Luo, Yanda Meng, Savita Madhusudhan, Gregory Yoke Hong Lip, Li Cheng 0001, Yalin Zheng, He Zhao 0002 |
MICCAI (1) | 9 |
| 2025 | GLCP: Global-to-Local Connectivity Preservation for Tubular Structure Segmentation
Feixiang Zhou, Zhuangzhi Gao, He Zhao 0002, Jianyang Xie, Yanda Meng, Yitian Zhao, Gregory Yoke Hong Lip, Yalin Zheng |
MICCAI (16) | 3 |
| 2025 | Two-stage sand-dust image enhancement method based on concentration scaling and domain adaptation
Yuting Yu, Zhidong Yang, Ruiheng Zhang 0001, Bosheng Ding, Lixin Xu 0001, He Zhao 0002 |
Knowl. Based Syst. | 6 |
| 2025 | ScanAhead: Simplifying standard plane acquisition of fetal head ultrasoundabstractThe fetal standard plane acquisition task aims to detect an Ultrasound (US) image characterized by specified anatomical landmarks and appearance for assessing fetal growth. However, in practice, due to variability in human operator skill and possible fetal motion, it can be challenging for a human operator to acquire a satisfactory standard plane. To support a human operator with this task, this paper first describes an approach to automatically predict the fetal head standard plane from a video segment approaching the standard plane. A transformer-based image predictor is proposed to produce a high-quality standard plane by understanding diverse scales of head anatomy within the US video frame. Because of the visual gap between the video frames and standard plane image, the predictor is equipped with an offset adaptor that performs domain adaption to translate the off-plane structures to the anatomies that would usually appear in a standard plane view. To enhance the anatomical details of the predicted US image, the approach is extended by utilizing a second modality, US probe movement, that provides 3D location information. Quantitative and qualitative studies conducted on two different head biometry planes demonstrate that the proposed US image predictor produces clinically plausible standard planes with superior performance to comparative published methods. The results of dual-modality solution show an improved visualization with enhanced anatomical details of the predicted US image. Clinical evaluations are also conducted to demonstrate the consistency between the predicted echo textures and the expected echo patterns seen in a typical real standard plane, which indicates its clinical feasibility for improving the standard plane acquisition process. Qianhui Men, He Zhao 0002, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble |
Medical Image Anal. | 2 |
| 2025 | TIER-LOC: Visual Query-based Video Clip Localization in fetal ultrasound videos with a multi-tier transformer
Divyanshu Mishra, Pramit Saha, He Zhao 0002, Netzahualcóyotl Hernández, Olga Patey, Aris T. Papageorghiou, J. Alison Noble |
Medical Image Anal. | 3 |
| 2025 | FasterSal: Robust and Real-Time Single-Stream Architecture for RGB-D Salient Object DetectionabstractRGB-D Salient Object Detection (SOD) aims to segment the most prominent areas and objects in a given pair of RGB and depth images. Most current models adopt a dual-stream structure to extract information from both RGB and depth images. However, this leads to an exponential increase in the number of parameters and computations in the model. Moreover, the discrepancy between RGB pretrained and the 3D geometric relationships in depth maps present a challenge for the encoder in capturing spatial structural details. These issues impact the model's accuracy in locating salient objects and distinguishing edge details. To address these, we propose a novel early feature fusion network, named FasterSal, which enables more efficient RGB-D SOD. FasterSal uses a single stream structure to receive RGB images and depth maps, extracting features based on the 3D geometric relationships in the depth map while fully leveraging the pretrained RGB encoder. This approach effectively avoids the inconsistencies between depth modality and the RGB pretrained encoder. It also significantly reduces the number of network parameters while maintaining efficient feature encoding capabilities. To achieve finer edge learning, the detail-aware loss and texture enhancement module are introduced. These modules are designed to extract latent details in high-frequency component features and to enhance the edge learning capability of the model using distance information. Experimental results on several benchmark datasets confirm the effectiveness and superiority of our method over the state-of-the-art approaches, achieving a good balance between performance and speed with only 3.4 million parameters and a CPU operating speed of 63 FPS. Jing Zhang 0037, Ruiheng Zhang 0001, Lixin Xu 0001, Xiankai Lu, Yushu Yu, Min Xu 0001, He Zhao 0002 |
IEEE Trans. Multim. | 7 |
| 2024 | STAN-LOC: Visual Query-Based Video Clip Localization for Fetal Ultrasound Sweep Videos
Divyanshu Mishra, Pramit Saha, He Zhao 0002, Olga Patey, Aris T. Papageorghiou, J. Alison Noble |
MICCAI (4) | 3 |
| 2024 | Multi-disease Detection in Retinal Images Guided by Disease Causal Estimation
Jianyang Xie, Xiuju Chen, Yitian Zhao, Yanda Meng, He Zhao 0002, Anh Nguyen 0003, Yalin Zheng |
MICCAI (1) | 5 |
| 2024 | CLIP-DR: Textual Knowledge-Guided Diabetic Retinopathy Grading with Ranking-Aware Prompting
Qinkai Yu, Jianyang Xie, Anh Nguyen 0003, He Zhao 0002, Jiong Zhang 0004, Huazhu Fu, Yitian Zhao, Yalin Zheng, Yanda Meng |
MICCAI (1) | 4 |
| 2024 | Anomaly Detection for Medical Images Using Heterogeneous Auto-EncoderabstractAnomaly detection is an important task for medical image analysis, which can alleviate the reliance of supervised methods on large labelled datasets. Most existing methods use a pixel-wise self-reconstruction framework for anomaly detection. However, there are two challenges of these studies: 1) they tend to overfit learning an identity mapping between the input and output, which leads to failure in detecting abnormal samples; 2) the reconstruction considers the pixel-wise differences which may lead to an undesirable result. To mitigate the above problems, we propose a novel heterogeneous Auto-Encoder (Hetero-AE) for medical anomaly detection. Our model utilizes a convolutional neural network (CNN) as the encoder and a hybrid CNN-Transformer network as the decoder. The heterogeneous structure enables the model to learn the intrinsic information of normal data and enlarge the difference on abnormal samples. To fully exploit the effectiveness of Transformer in the hybrid network, a multi-scale sparse Transformer block is proposed to trade off modelling long-range feature dependencies and high computational costs. Moreover, the multi-stage feature comparison is introduced to reduce the noise of pixel-wise comparison. Extensive experiments on four public datasets (i.e., retinal OCT, chest X-ray, brain MRI, and COVID-19) verify the effectiveness of our method on different imaging modalities for anomaly detection. Additionally, our method can accurately detect tumors in brain MRI and lesions in retinal OCT with interpretable heatmaps to locate lesion areas, assisting clinicians in diagnosing abnormalities efficiently. Shuai Lu 0003, He Zhao 0002, Hanruo Liu, Ningli Wang, Huiqi Li |
IEEE Trans. Image Process. | 3 |
| 2023 | Dual Conditioned Diffusion Models for Out-of-Distribution Detection: Application to Fetal Ultrasound Videos
Divyanshu Mishra, He Zhao 0002, Pramit Saha, Aris T. Papageorghiou, J. Alison Noble |
MICCAI (1) | 2 |
| 2023 | PKRT-Net: Prior knowledge-based relation transformer network for optic cup and disc segmentation
Shuai Lu 0003, He Zhao 0002, Hanruo Liu, Huiqi Li, Ningli Wang |
Neurocomputing | 2 |
| 2023 | Memory-based unsupervised video clinical quality assessment with multi-modality data in fetal ultrasoundabstractIn obstetric sonography, the quality of acquisition of ultrasound scan video is crucial for accurate (manual or automated) biometric measurement and fetal health assessment. However, the nature of fetal ultrasound involves free-hand probe manipulation and this can make it challenging to capture high-quality videos for fetal biometry, especially for the less-experienced sonographer. Manually checking the quality of acquired videos would be time-consuming, subjective and requires a comprehensive understanding of fetal anatomy. Thus, it would be advantageous to develop an automatic quality assessment method to support video standardization and improve diagnostic accuracy of video-based analysis. In this paper, we propose a general and purely data-driven video-based quality assessment framework which directly learns a distinguishable feature representation from high-quality ultrasound videos alone, without anatomical annotations. Our solution effectively utilizes both spatial and temporal information of ultrasound videos. The spatio-temporal representation is learned by a bi-directional reconstruction between the video space and the feature space, enhanced by a key-query memory module proposed in the feature space. To further improve performance, two additional modalities are introduced in training which are the sonographer gaze and optical flow derived from the video. Two different clinical quality assessment tasks in fetal ultrasound are considered in our experiments, i.e., measurement of the fetal head circumference and cerebellar diameter; in both of these, low-quality videos are detected by the large reconstruction error in the feature space. Extensive experimental evaluation demonstrates the merits of our approach. He Zhao 0002, Qingqing Zheng, Clare Teng, Robail Yasrab, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble |
Medical Image Anal. | 1 |
| 2023 | Retinal image enhancement with artifact reduction and structure retention
Bingyu Yang, He Zhao 0002, Lvchen Cao, Hanruo Liu, Ningli Wang, Huiqi Li |
Pattern Recognit. | 2 |
| 2023 | A Machine Learning Method for Automated Description and Workflow Analysis of First Trimester Ultrasound ScansabstractObstetric ultrasound assessment of fetal anatomy in the first trimester of pregnancy is one of the less explored fields in obstetric sonography because of the paucity of guidelines on anatomical screening and availability of data. This paper, for the first time, examines imaging proficiency and practices of first trimester ultrasound scanning through analysis of full-length ultrasound video scans. Findings from this study provide insights to inform the development of more effective user-machine interfaces, of targeted assistive technologies, as well as improvements in workflow protocols for first trimester scanning. Specifically, this paper presents an automated framework to model operator clinical workflow from full-length routine first-trimester fetal ultrasound scan videos. The 2D+t convolutional neural network-based architecture proposed for video annotation incorporates transfer learning and spatio-temporal (2D+t) modelling to automatically partition an ultrasound video into semantically meaningful temporal segments based on the fetal anatomy detected in the video. The model results in a cross-validation A1 accuracy of 96.10% , F1=0.95 , precision =0.94 and recall =0.95 . Automated semantic partitioning of unlabelled video scans (n=250) achieves a high correlation with expert annotations ( ρ = 0.95, p=0.06 ). Clinical workflow patterns, operator skill and its variability can be derived from the resulting representation using the detected anatomy labels, order, and distribution. It is shown that nuchal translucency (NT) is the toughest standard plane to acquire and most operators struggle to localize high-quality frames. Furthermore, it is found that newly qualified operators spend 25.56% more time on key biometry tasks than experienced operators. Robail Yasrab, Zeyu Fu, He Zhao 0002, Lok Hin Lee, Harshita Sharma, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble |
IEEE Trans. Medical Imaging | 3 |
| 2022 | Towards Unsupervised Ultrasound Video Clinical Quality Assessment with Multi-modality Data
He Zhao 0002, Qingqing Zheng, Clare Teng, Robail Yasrab, Lior Drukker, Aris T. Papageorghiou, J. Alison Noble |
MICCAI (4) | 1 |
| 2021 | Anomaly Detection for Medical Images Using Self-Supervised and Translation-Consistent FeaturesabstractAs the labeled anomalous medical images are usually difficult to acquire, especially for rare diseases, the deep learning based methods, which heavily rely on the large amount of labeled data, cannot yield a satisfactory performance. Compared to the anomalous data, the normal images without the need of lesion annotation are much easier to collect. In this paper, we propose an anomaly detection framework, namely [Formula: see text], extracting [Formula: see text]elf-supervised and tr [Formula: see text]ns [Formula: see text]ation-consistent features for [Formula: see text]nomaly [Formula: see text]etection. The proposed SALAD is a reconstruction-based method, which learns the manifold of normal data through an encode-and-reconstruct translation between image and latent spaces. In particular, two constraints (i.e., structure similarity loss and center constraint loss) are proposed to regulate the cross-space (i.e., image and feature) translation, which enforce the model to learn translation-consistent and representative features from the normal data. Furthermore, a self-supervised learning module is engaged into our framework to further boost the anomaly detection accuracy by deeply exploiting useful information from the raw normal data. An anomaly score, as a measure to separate the anomalous data from the healthy ones, is constructed based on the learned self-supervised-and-translation-consistent features. Extensive experiments are conducted on optical coherence tomography (OCT) and chest X-ray datasets. The experimental results demonstrate the effectiveness of our approach. He Zhao 0002, Yuexiang Li, Nanjun He, Kai Ma 0002, Leyuan Fang, Huiqi Li, Yefeng Zheng 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2021 | Deep Representation-Based Domain Adaptation for Nonstationary EEG ClassificationabstractIn the context of motor imagery, electroencephalography (EEG) data vary from subject to subject such that the performance of a classifier trained on data of multiple subjects from a specific domain typically degrades when applied to a different subject. While collecting enough samples from each subject would address this issue, it is often too time-consuming and impractical. To tackle this problem, we propose a novel end-to-end deep domain adaptation method to improve the classification performance on a single subject (target domain) by taking the useful information from multiple subjects (source domain) into consideration. Especially, the proposed method jointly optimizes three modules, including a feature extractor, a classifier, and a domain discriminator. The feature extractor learns the discriminative latent features by mapping the raw EEG signals into a deep representation space. A center loss is further employed to constrain an invariant feature space and reduce the intrasubject nonstationarity. Furthermore, the domain discriminator matches the feature distribution shift between source and target domains by an adversarial learning strategy. Finally, based on the consistent deep features from both domains, the classifier is able to leverage the information from the source domain and accurately predict the label in the target domain at the test time. To evaluate our method, we have conducted extensive experiments on two real public EEG data sets, data set IIa, and data set IIb of brain-computer interface (BCI) Competition IV. The experimental results validate the efficacy of our method. Therefore, our method is promising to reduce the calibration time for the use of BCI and promote the development of BCI. He Zhao 0002, Qingqing Zheng, Kai Ma 0002, Huiqi Li, Yefeng Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2020 | Improving retinal vessel segmentation with joint local loss by matting
He Zhao 0002, Huiqi Li, Li Cheng 0001 |
Pattern Recognit. | 1 |
| 2019 | Data-Driven Enhancement of Blurry Retinal Images via Generative Adversarial Networks
He Zhao 0002, Bingyu Yang, Lvchen Cao, Huiqi Li |
MICCAI (1) | 1 |
| 2019 | Supervised Segmentation of Un-Annotated Retinal Fundus Images by SynthesisabstractWe focus on the practical challenge of segmenting new retinal fundus images that are dissimilar to existing well-annotated data sets. It is addressed in this paper by a supervised learning pipeline, with its core being the construction of a synthetic fundus image data set using the proposed R-sGAN technique. The resulting synthetic images are realistic-looking in terms of the query images while maintaining the annotated vessel structures from the existing data set. This helps to bridge the mismatch between the query images and the existing well-annotated data set. As a consequence, any known supervised fundus segmentation technique can be directly utilized on the query images, after training on this synthetic data set. Extensive experiments on different fundus image data sets demonstrate the competitiveness of the proposed approach in dealing with a diverse range of mismatch settings. He Zhao 0002, Huiqi Li, Sebastian Maurer-Stroh, Yuhong Guo, Qiuju Deng, Li Cheng 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2018 | Synthesizing retinal and neuronal images with generative adversarial netsabstractThis paper aims at synthesizing multiple realistic-looking retinal (or neuronal) images from an unseen tubular structured annotation that contains the binary vessel (or neuronal) morphology. The generated phantoms are expected to preserve the same tubular structure, and resemble the visual appearance of the training images. Inspired by the recent progresses in generative adversarial nets (GANs) as well as image style transfer, our approach enjoys several advantages. It works well with a small training set with as few as 10 training examples, which is a common scenario in medical image analysis. Besides, it is capable of synthesizing diverse images from the same tubular structured annotation. Extensive experimental evaluations on various retinal fundus and neuronal imaging applications demonstrate the merits of the proposed approach. He Zhao 0002, Huiqi Li, Sebastian Maurer-Stroh, Li Cheng 0001 |
Medical Image Anal. | 1 |
| 2017 | Segment 2D and 3D Filaments by Learning Structured and Contextual FeaturesabstractWe focus on the challenging problem of filamentary structure segmentation in both 2D and 3D images, including retinal vessels and neurons, among others. Despite the increasing amount of efforts in learning based methods to tackle this problem, there still lack proper data-driven feature construction mechanisms to sufficiently encode contextual labelling information, which might hinder the segmentation performance. This observation prompts us to propose a data-driven approach to learn structured and contextual features in this paper. The structured features aim to integrate local spatial label patterns into the feature space, thus endowing the follow-up tree classifiers capability to grouping training examples with similar structure into the same leaf node when splitting the feature space, and further yielding contextual features to capture more of the global contextual information. Empirical evaluations demonstrate that our approach outperforms state-of-the-arts on well-regarded testbeds over a variety of applications. Our code is also made publicly available in support of the open-source research activities. Lin Gu 0003, Xiaowei Zhang 0002, He Zhao 0002, Huiqi Li, Li Cheng 0001 |
IEEE Trans. Medical Imaging | 3 |