Heng Li 0010

dblp:02/3672-10 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0001-7754-7141ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 7 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021
YearPublicationVenuePosition
2026 DeLightMono: Enhancing Self-Supervised Monocular Depth Estimation in Endoscopy by Decoupling Uneven Illumination
abstract
Self-supervised monocular depth estimation serves as a key task in the development of endoscopic navigation systems. However, performance degradation persists due to uneven illumination inherent in endoscopic images, particularly in low-intensity regions. Existing low-light enhancement techniques fail to effectively guide the depth network. Furthermore, solutions from other fields, like autonomous driving, require well-lit images, making them unsuitable and increasing data collection burdens. To this end, we present DeLightMono - a novel self-supervised monocular depth estimation framework with illumination decoupling. Specifically, endoscopic images are represented by a designed illumination-reflectance-depth model, and are decomposed with auxiliary networks. Moreover, a self-supervised joint-optimizing framework with novel losses leveraging the decoupled components is proposed to mitigate the effects of uneven illumination on depth estimation. The effectiveness of the proposed methods was rigorously verified through extensive comparisons and an ablation study performed on two public datasets.
Mingyang Ou, Haojin Li 0003, Ke Niu 0002, Zhongxi Qiu, Heng Li 0010, Jiang Liu 0001
AAAI6
2026 Hybrid Feature Edge Enhancement for Self-Supervised Monocular Depth Estimation in Endoscopic Scenes
Jiadong Guo, Ke Niu 0002, Heng Li 0010, Mingyang Ou, Zeyun Liu
ICIC (21)4
2026 Geometric-Equivariant Blind-Spot Network for Self-Supervised Intra-operative OCT Denoising
Xiangyang Yu, Yinan Wen, Haojin Li 0003, Sanqian Li, Heng Li 0010, Jiang Liu 0001
ICIC (18)5
2026 Endoscopic depth estimation based on deep learning: A survey
Ke Niu 0002, Zeyun Liu, Heng Li 0010, Naian Xiao, Binghua Su, Qika Lin, Kaize Shi
Neurocomputing4
2026 Explicable intensity-aware 3D cerebrovascular segmentation with planar representation
Cheng Chen 0024, Yunqing Chen, Huansheng Ning, Heng Li 0010, Jiang Liu 0001, Ruoxiu Xiao
Medical Image Anal.4
2026 Fundus image quality assessment in retinopathy of prematurity via multi-label graph evidential network
Donghan Wu, Wenyue Shen, Heng Li 0010, Huaying Hao, Juan Ye, Yitian Zhao
Medical Image Anal.4
2025 AIF-SFDA: Autonomous Information Filter Driven Source-Free Domain Adaptation for Medical Image Segmentation
abstract
Decoupling domain-variant information (DVI) from domain-invariant information (DII) serves as a prominent strategy for mitigating domain shifts in the practical implementation of deep learning algorithms. However, in medical settings, concerns surrounding data collection and privacy often restrict access to both training and test data, hindering the empirical decoupling of information by existing methods. To tackle this issue, we propose an Adaptive Information Filter-driven Source-free Domain Adaptation (AIF-SFDA) algorithm, which leverages a frequency-based learnable information filter to autonomously decouple DVI and DII. Information Bottleneck (IB) and Self-supervision (SS) are incorporated to optimize the learnable frequency filter. The IB governs the information flow within the filter to diminish redundant DVI, while SS preserves DII in alignment with the specific task and image modality. Thus, the adaptive information filter can overcome domain shifts relying solely on target data. A series of experiments covering various medical image modalities and segmentation tasks were conducted to demonstrate the benefits of AIF-SFDA through comparisons with leading algorithms and ablation studies.
Haojin Li 0003, Heng Li 0010, Rihan Zhong, Ke Niu 0002, Huazhu Fu, Jiang Liu 0001
AAAI2
2025 RADAR: Enhancing Radiology Report Generation with Supplementary Knowledge Injection
abstract
Large language models (LLMs) have demonstrated remarkable capabilities in various domains, including radiology report generation.Previous approaches have attempted to utilize multimodal LLMs for this task, enhancing their performance through the integration of domainspecific knowledge retrieval.However, these approaches often overlook the knowledge already embedded within the LLMs, leading to redundant information integration.To address this limitation, we propose RADAR, a framework for enhancing radiology report generation with supplementary knowledge injection.RADAR improves report generation by systematically leveraging both the internal knowledge of an LLM and externally retrieved information.Specifically, it first extracts the model's acquired knowledge that aligns with expert imagebased classification outputs.It then retrieves relevant supplementary knowledge to further enrich this information.Finally, by aggregating both sources, RADAR generates more accurate and informative radiology reports.Extensive experiments on MIMIC-CXR, CHEXPERT-PLUS, and IU X-RAY demonstrate that our model outperforms state-of-the-art LLMs in both language quality and clinical accuracy 1 .
Kaishuai Xu, Heng Li 0010, Jiang Liu 0001
ACL (1)4
2025 Multi-View Test-Time Adaptation for Semantic Segmentation in Clinical Cataract Surgery
abstract
Cataract surgery, a widely performed operation worldwide, is incorporating semantic segmentation to advance computer-assisted intervention. However, the tissue appearance and illumination in cataract surgery often differ among clinical centers, intensifying the issue of domain shifts. While domain adaptation offers remedies to the shifts, the necessity for data centralization raises additional privacy concerns. To overcome these challenges, we propose a Multi-view Test-time Adaptation algorithm (MUTA) to segment cataract surgical scenes, which leverages multi-view learning to enhance model training within the source domain and model adaptation within the target domain. In the training phase, the segmentation model is equipped with multi-view decoders to boost its robustness against variations in cataract surgery. During the inference phase, test-time adaptation is implemented using multi-view knowledge distillation, enabling model updates in clinics without data centralization or privacy concerns. We conducted experiments in a simulated cross-center scenario using several cataract surgery datasets to evaluate the effectiveness of MUTA. Through comparisons and investigations, we have validated that MUTA effectively learns a robust source model and adapts the model to target data during the practical inference phase. Code and datasets are available at https://github.com/liamheng/CAI-algorithms.
Heng Li 0010, Mingyang Ou, Haojin Li 0003, Zhongxi Qiu, Ke Niu 0002, Huazhu Fu, Jiang Liu 0001
IEEE Trans. Medical Imaging1
2024 Enhancing and Adapting in the Clinic: Source-Free Unsupervised Domain Adaptation for Medical Image Enhancement
abstract
Medical imaging provides many valuable clues involving anatomical structure and pathological characteristics. However, image degradation is a common issue in clinical practice, which can adversely impact the observation and diagnosis by physicians and algorithms. Although extensive enhancement models have been developed, these models require a well pre-training before deployment, while failing to take advantage of the potential value of inference data after deployment. In this paper, we raise an algorithm for source-free unsupervised domain adaptive medical image enhancement (SAME), which adapts and optimizes enhancement models using test data in the inference phase. A structure-preserving enhancement network is first constructed to learn a robust source model from synthesized training data. Then a teacher-student model is initialized with the source model and conducts source-free unsupervised domain adaptation (SFUDA) by knowledge distillation with the test data. Additionally, a pseudo-label picker is developed to boost the knowledge distillation of enhancement tasks. Experiments were implemented on ten datasets from three medical image modalities to validate the advantage of the proposed algorithm, and setting analysis and ablation studies were also carried out to interpret the effectiveness of SAME. The remarkable enhancement performance and benefits for downstream tasks demonstrate the potential and generalizability of SAME. The code is available at https://github.com/liamheng/Annotation-free-Medical-Image-Enhancement.
Heng Li 0010, Ziqin Lin, Zhongxi Qiu, Zinan Li, Ke Niu 0002, Huazhu Fu, Jiang Liu 0001
IEEE Trans. Medical Imaging1
2023 ACT-Net: Anchor-Context Action Detection in Surgery Videos
Luoying Hao, Heng Li 0010, Huazhu Fu, Jinming Duan 0001, Jiang Liu 0001
MICCAI (9)5
2023 Content-Preserving Diffusion Model for Unsupervised AS-OCT Image Despeckling
Sanqian Li, Risa Higashita, Huazhu Fu, Heng Li 0010, Jingxuan Niu, Jiang Liu 0001
MICCAI (7)4
2023 Frequency-Mixed Single-Source Domain Generalization for Medical Image Segmentation
Heng Li 0010, Haojin Li 0003, Huazhu Fu, Xiuyun Su, Jiang Liu 0001
MICCAI (6)1
2023 A generic fundus image enhancement network boosted by frequency self-supervised representation learning
Heng Li 0010, Haofeng Liu, Huazhu Fu, Yanwu Xu 0001, Hai Shu, Ke Niu 0002, Jiang Liu 0001
Medical Image Anal.1
2023 Multi-Learner Based Deep Meta-Learning for Few-Shot Medical Image Classification
abstract
Few-shot learning (FSL) is promising in the field of medical image analysis due to high cost of establishing high-quality medical datasets. Many FSL approaches have been proposed in natural image scenes. However, present FSL methods are rarely evaluated on medical images and the FSL technology applicable to medical scenarios need to be further developed. Meta-learning has supplied an optional framework to address the challenging FSL setting. In this paper, we propose a novel multi-learner based FSL method for multiple medical image classification tasks, combining meta-learning with transfer-learning and metric-learning. Our designed model is composed of three learners, including auto-encoder, metric-learner and task-learner. In transfer-learning, all the learners are trained on the base classes. In the ensuing meta-learning, we leverage multiple novel tasks to fine-tune the metric-learner and task-learner in order to fast adapt to unseen tasks. Moreover, to further boost the learning efficiency of our model, we devised real-time data augmentation and dynamic Gaussian disturbance soft label (GDSL) scheme as effective generalization strategies of few-shot classification tasks. We have conducted experiments for three-class few-shot classification tasks on three newly-built challenging medical benchmarks, BLOOD, PATH and CHEST. Extensive comparisons to related works validated that our method achieved top performance both on homogeneous medical datasets and cross-domain datasets.
Hongyang Jiang 0001, Mengdi Gao, Heng Li 0010, Richu Jin, Hanpei Miao, Jiang Liu 0001
IEEE J. Biomed. Health Informatics3
2022 Replay-Oriented Gradient Projection Memory for Continual Learning in Medical Scenarios
abstract
Despite the tremendous progress recently achieved by deep learning (DL) in medical image analysis, most DL models only concentrate on single data distribution, which follows the independent and identically distributed (i.i.d) assumption. However, in practice, image data distribution changes with clinical conditions, such as different scanner manufacturers, imaging settings, and statistics regions. Although one can further train the model on new data samples, updating a model with data from an unknown distribution will always result in the model’s performance degradation on the learned data, a notorious phenomenon called catastrophic forgetting. Therefore affects the applicability of DL algorithms in continuously changing clinical scenarios. In this study, we have proposed a new method to address the impact of changing distributions in continual learning scenarios and alleviate catastrophic forgetting. A gradient regularization approach is used to suppress forgetting, and a replay-oriented consistency calculation method combined with a subspace weighting strategy is proposed to improve the model plasticity further. The proposed replay-oriented gradient projection memory (RO-GPM) is evaluated on multiple fundus disease diagnosis datasets including a real-world application and a continual learning benchmark. The quantitative and visualization results demonstrate that the proposed RO-GPM achieves superior performance to state-of-the-art algorithms by a large margin.1
Kuang Shu, Heng Li 0010, Qinghai Guo, Luziwei Leng, Jianxing Liao, Jiang Liu 0001
BIBM2
2022 Structure-Consistent Restoration Network for Cataract Fundus Image Enhancement
Heng Li 0010, Haofeng Liu, Huazhu Fu, Hai Shu, Yitian Zhao, Jiang Liu 0001
MICCAI (2)1
2022 Degradation-Invariant Enhancement of Fundus Images via Pyramid Constraint Network
Haofeng Liu, Heng Li 0010, Huazhu Fu, Ruoxiu Xiao, Yunshu Gao, Jiang Liu 0001
MICCAI (2)2
2022 Screening of Dementia on OCTA Images via Multi-projection Consistency and Complementarity
Heng Li 0010, Zunjie Xiao, Huazhu Fu, Yitian Zhao, Richu Jin, William Robert Kwapong, Hanpei Miao, Jiang Liu 0001
MICCAI (2)2
2022 An Annotation-Free Restoration Network for Cataractous Fundus Images
abstract
Cataracts are the leading cause of vision loss worldwide. Restoration algorithms are developed to improve the readability of cataract fundus images in order to increase the certainty in diagnosis and treatment for cataract patients. Unfortunately, the requirement of annotation limits the application of these algorithms in clinics. This paper proposes a network to annotation-freely restore cataractous fundus images (ArcNet) so as to boost the clinical practicability of restoration. Annotations are unnecessary in ArcNet, where the high-frequency component is extracted from fundus images to replace segmentation in the preservation of retinal structures. The restoration model is learned from the synthesized images and adapted to real cataract images. Extensive experiments are implemented to verify the performance and effectiveness of ArcNet. Favorable performance is achieved using ArcNet against state-of-the-art algorithms, and the diagnosis of ocular fundus diseases in cataract patients is promoted by ArcNet. The capability of properly restoring cataractous images in the absence of annotated data promises the proposed algorithm outstanding clinical practicability.
Heng Li 0010, Haofeng Liu, Huazhu Fu, Yitian Zhao, Hanpei Miao, Jiang Liu 0001
IEEE Trans. Medical Imaging1
2020 CT Scan Synthesis for Promoting Computer-Aided Diagnosis Capacity of COVID-19
Heng Li 0010, Sanqian Li, Peng Liu 0049, Risa Higashita, Jiang Liu 0001
ICIC (2)1