EDBT 2026 Demo / reviewers in the wild / expert
Huazhu Fu
dblp:63/7767
· DBLP profile ↗
303ranked-venue papers
19as first author
222since 2021 · last 2027
0000-0002-9702-5524ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 153 · 9 first-author · 100 since 2021Applied, interdisciplinary, general and emerging computing · 150 · 9 first-author · 120 since 2021Artificial intelligence and machine learning · 99 · 5 first-author · 74 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | MedDATP: Adapting CLIP for few-shot medical image classification via domain adapter and task prompts
Zixun Zhang, Yuncheng Jiang 0002, Jun Wei 0006, Huazhu Fu, Shuguang Cui, Tao Luo 0014, Zhen Li 0026 |
Expert Syst. Appl. | 4 |
| 2026 | Segmentation-synthesis co-training for semi-supervised domain generalizable medical image segmentation
Qingshan Hou, Peng Cao 0001, Jinzhu Yang, Huazhu Fu, Osmar R. Zaïane, Zhaolin Chen |
Artif. Intell. Medicine | 5 |
| 2026 | Video Shadow Detection with Intra-and Inter-video Cooperation
Zhihao Chen 0004, Junting Zhao, Lei Zhu 0003, Huazhu Fu, Wei Feng 0005 |
Int. J. Comput. Vis. | 5 |
| 2026 | Cross-sample Consistency Learning for Semi-supervised Medical Image Segmentation
Tao Zhou 0002, Yunqi Gu, Kaiwen Huang 0002, Huazhu Fu, Xiaojun Wu 0001, Josef Kittler |
Int. J. Comput. Vis. | 4 |
| 2026 | Federated semi-supervised calibrated efficient fine-tuning of foundation models for medical image classification
Along He, Yanlin Wu, LinLin Shen, Ke Zou, Huazhu Fu |
Knowl. Based Syst. | 6 |
| 2026 | Prompt guiding multi-scale adaptive sparse representation-driven network for low-dose CT MAR
Baoshun Shi, Huazhu Fu, Zhanli Hu |
Medical Image Anal. | 4 |
| 2026 | Annotation-efficient medical image segmentation via cross-latent graphs and vector-quantized memory
Yanyu Xu 0001, Menghan Zhou, Xinxing Xu, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026, Li-Zhen Cui 0001 |
Medical Image Anal. | 4 |
| 2026 | STAGE challenge: Structural-Functional Transition in Glaucoma Assessment
Shiqi Zhou, Yuancong Liang, Huihui Fang, Ziyang Chen 0003, Yong Xia 0001, Chubin Ou, Yubo Tan, Haojie Yin, Chengcheng Feng, Hao Zhou 0030, Hrvoje Bogunovic, Huazhu Fu, Fei Li 0021, Xiulan Zhang, Yanwu Xu 0001 |
Medical Image Anal. | 16 |
| 2026 | Unsupervised domain adaptation via style-aware self-intermediate domain
Lianyu Wang, Meng Wang 0038, Daoqiang Zhang, Huazhu Fu |
Pattern Recognit. | 4 |
| 2026 | SegMIC: A universal model for medical image segmentation through in-context learning
Fan Yang 0054, Xin Li 0079, Zhicheng Jiao, Qiang Zhai, Xiaomeng Li 0001, De Wu, Huazhu Fu, Hong Cheng 0002 |
Pattern Recognit. | 8 |
| 2026 | MedSAM-U: Uncertainty-Guided Auto Multi-Prompt Adaptation for Reliable MedSAMabstractThe Medical Segment Anything Model (MedSAM) has demonstrated strong performance in medical image segmentation, attracting increasing attention in the medical imaging domain. However, as with many prompt-based segmentation models, its performance is highly sensitive to the type and location of input prompts. This sensitivity often leads to suboptimal segmentation outcomes and necessitates labor-intensive manual prompt tuning, which hampers both efficiency and robustness. To address this challenge, this paper proposes MedSAM-U, an uncertainty-guided framework designed to automatically refine prompt inputs and enhance segmentation reliability. Specifically, a Multi-Prompt Adapter is integrated into MedSAM, resulting in MPA-MedSAM, which enables the model to effectively accommodate diverse multi-prompt inputs. An uncertainty estimation module is then introduced to evaluate the reliability of the prompts and their initial segmentation results. Based on this, a novel uncertainty-guided prompt adaptation strategy is applied to automatically generate refined prompts and more accurate segmentation outputs. The proposed MedSAM-U framework is evaluated across multiple medical imaging modalities. Experimental results on five diverse datasets demonstrate that MedSAM-U achieves consistent performance improvements ranging from 1.7% to 20.5% over the baseline MedSAM, confirming its effectiveness and practicality for robust and efficient medical image segmentation. Ke Zou, Mengting Luo, Linchao He, Meng Wang 0038, Yi Zhang 0018, Hu Chen 0002, Huazhu Fu |
IEEE Trans. Circuits Syst. Video Technol. | 10 |
| 2026 | Multi-Granularity Topological Reasoning for Anatomically Consistent Vasculature ParsingabstractQuantitative analysis of retinal vascular morphology is vital for clinical decision-making and the investigation of systemic diseases. Central to this process is the accurate segmentation of retinal arteries and veins (A/V) from the background, a task challenged by substantial variations in vessel calibers and the presence of low-contrast or ambiguous structures in fundus images, especially in ultra-wide field imaging where peripheral distortions and large-scale anatomical variability are pronounced. These factors often lead to fragmented semantic representations and topological inconsistencies in automated segmentation outputs. To address these limitations, we propose Ultra, a multi-granularity topological reasoning network designed for precise A/V segmentation. Ultra adopts a cascaded two-stage architecture: PriorNet generates coarse, multi-scale vascular priors that provide structural guidance, while RefineNet performs topology-aware segmentation refinement. To further enforce topological coherence, we propose the neighboring pixel connectivity regularization (NICER) layer, which selectively integrates local connectivity information predicted by the proposed connectivity prediction union (CPU) module. This connectivity is employed as auxiliary supervision through a pixel-wise local connectivity loss, reinforcing structural reasoning and promoting anatomically consistent vascular topology inference. Extensive experiments on ultra-wide field fundus imaging (UWF) datasets demonstrate that Ultra achieves state-of-the-art performance in A/V segmentation and topological preservation. Moreover, Ultra generalizes well to conventional color fundus photography (CFP) datasets, underscoring its robustness and broad applicability. Code is publicly available at: https://github.com/iMED-Lab/Ultra. Lei Mou, Yonghuai Liu, Zhuoting Xu, Hao Zhang 0113, Yalin Zheng, Jiang Liu 0001, Huazhu Fu, Yitian Zhao |
IEEE Trans. Image Process. | 7 |
| 2026 | Super-Resolution Reconstruction of OCTA via Multi-Field-of-View Representation LearningabstractHigh-resolution Optical Coherence Tomography Angiography (OCTA) images are essential for morphological analysis and biomarker measurement of the retinal vasculature. They can also provide underlying biomarkers for the accurate analysis of eye-related diseases. The trade-off between the high resolution (HR) and large scanning field-of-view (FOV) is a long-standing problem for OCTA image instrument. A large FOV image provides more retinal information with shorter acquisition time but often suffers from low resolution (LR), high scatter noise, and poor vascular contrast. In order to obtain HR OCTA images with larger FOV, we propose a novel self-similar dynamic domain adaptation network based on cross-field-of-view representation learning. The network enables LR images (i.e., $6\times \text{6}\,\text{mm}^{2}$) to learn HR image (i.e., $3\times \text{3}\,\text{mm}^{2}$) feature representations specialized for OCTA by constructing feature mapping relations for cross-field-of-view OCTA scans. To be specific, a multiple random degradation model is proposed on HR images to generate various synthetic LR images. Further, we propose a dynamic domain adaptation framework that prompts feature dynamic alignment of the LR image reconstruction results with those of synthetic LR images. Finally, a novel self-similar supervision loss is proposed to optimize the reconstruction results from LR to HR by exploiting the similarity between vessels in different regions. Experimental results on three OCTA datasets show that the proposed method surpasses existing state-of-the-art ones, significantly enhancing retinal structure segmentation and disease classification. Our OCTA dataset (the first dataset in this research area with paired $3\times 3$ and $6\times \text{6}\,\text{mm}^{2}$ OCTA images) and code are publicly available. Huaying Hao, Shaoyi Leng, Yanda Meng, Yonghuai Liu, Yalin Zheng, Huazhu Fu, Jiong Zhang 0004, Quanyong Yi, Yue Liu 0005, Jingfeng Zhang, Yitian Zhao |
IEEE J. Biomed. Health Informatics | 6 |
| 2026 | Position Paper: Artificial Intelligence in Medical Image Analysis: Advances, Clinical Translation, and Emerging FrontiersabstractOver the past five years, artificial intelligence (AI) has introduced new models and methods for addressing the challenges associated with the broader adoption of AI models and systems in medicine. This paper reviews recent advances in AI for medical image and video analysis, outlines emerging paradigms, highlights pathways for successful clinical translation, and provides recommendations for future work. Hybrid Convolutional Neural Network (CNN) Transformer architectures now deliver state-of-the-art results in segmentation, classification, reconstruction, synthesis, and registration. Foundation and generative AI models enable the use of transfer learning to smaller datasets with limited ground truth. Federated learning supports privacy-preserving collaboration across institutions. Explainable and trustworthy AI approaches have become essential to foster clinician trust, ensure regulatory compliance, and facilitate ethical deployment. Together, these developments pave the way for integrating AI into radiology, pathology, and wider healthcare workflows. Andreas Panayides, Hao Chen 0011, Nenad Filipovic, Tijana Geroski, Junlin Hou, Karim Lekadir, Kostas Marias, George K. Matsopoulos, Giorgos Papanastasiou, Pinaki Sarder, Georgia D. Tourassi, Sotirios A. Tsaftaris, Huazhu Fu, Efthyvoulos C. Kyriacou, Christos P. Loizou, Michalis E. Zervakis, Joel H. Saltz, Farah Shamout, Ken C. L. Wong, Jianhua Yao 0001, Amir A. Amini, Dimitrios I. Fotiadis, Constantinos S. Pattichis, Marios S. Pattichis |
IEEE J. Biomed. Health Informatics | 13 |
| 2026 | Online Bayesian Approximation Based Uncertainty Aware Model for Ophthalmic Image SegmentationabstractThe robust segmentation of different targets in multiple modality images is challenging due to factors such as low contrast, variations in target size and shape, and interference from diseases, which may lead to segmentation ambiguity. In addition, the assessment of the reliability of artificial intelligence is crucial for its clinical application. This paper proposes the Online Bayesian approximation based Uncertainty-aware Network (OBU-Net) for robust ophthalmic image segmentation. Our approach introduces an efficient online Bayesian method to update a spatial uncertainty map during training continuously. Then, the Spatial Uncertainty Aware Block (SUA-B) leverages the uncertainty map to localize and prioritize attention to ambiguous regions. Additionally, we extract pixel-wise confidence from multi-scale predictions to integrate hierarchical predictions. We compare OBU-Net with state-of-the-art (SOTA) methods on six datasets. The experimental results demonstrate that our method achieves the best overall performance across different modalities and segmentation tasks, highlighting the robustness of our approach. Additionally, metamorphic testing experiments were conducted, exploring the algorithm's stability against random perturbations. Lastly, we propose an image-level uncertainty score and demonstrate its effectiveness for evaluating the model's segmentation reliability. Yinglin Zhang, Risa Higashita, Lingxi Zeng, Ruiling Xi, Tianhang Liu, Huazhu Fu, Dave Towey, Ruibin Bai, Jiang Liu 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2026 | SABPI-Net: A Structure-Aware Bidirectional Proxy Interaction Network for Infantile Retinal Disease DiagnosisabstractDelayed treatment of infantile retinal disease can reduce its effectiveness and may cause severe and irreversible damage. Automated diagnosis of infant retinal diseases faces challenges including subtle early lesions, diverse clinical phenotypes, imaging variations, and imbalanced data. To address these, which cannot be well addressed by existing general foundation models, we propose structure-aware bidirectional proxy interaction network (SABPI-Net) in a universal learning framework. SABPI-Net incorporates a high-frequency mapping branch, and employs a proposed proxy interaction attention module to enable effective interaction between its trunk feature encoding branch and the high-frequency mapping branch, thereby facilitating enhanced perception of retinal detail structures. Domain-agnostic embedding space self-matching, guided by a memory-bank low-frequency component replacement strategy, promotes domain-invariant learning and consistent model performance under diverse image styles. Finally, the tail-aware feature fusion strategy for fine-tuning further enhances the model's diagnostic sensitivity to tailed diseases. In this study, three classification tasks related to infant retinal diseases are implemented on the largest clinical infant retina dataset to date, covering 19 infant retinal diseases or normal conditions. SABPI-Net achieves superior performance compared to 13 SOTA methods, with 95.32% accuracy on mainstream clinical tasks, 73.58% on ROP five-stage classification, and 84.25% on multi-disease classification, representing improvements of 1.57%, 1.88%, and 4.71% respectively over the best competing methods. Extensive experiments demonstrate the effectiveness and superiority of SABPI-Net in diagnosing infant retinal diseases. Shaobin Chen, Huazhu Fu, Jiaju Huang, Zhenquan Wu, Behdad Dashtbozorg, Bai Ying Lei, Yue Sun 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2026 | Uncertainty-Guided Prototype Reliability Enhancement Network for Few-Shot Medical Image SegmentationabstractFew-Shot Learning (FSL) has garnered increasing attention for data-scarce scenarios, particularly in medical segmentation tasks where only a few labeled data points are available. Existing few-shot segmentation methods typically learn prototypes from support images and employ nearest-neighbor searching to segment query images. Despite notable progress, effectively learning prototypes for each class remains a challenging task to achieve promising results. In this paper, we propose an Uncertainty-guided Prototype Reliability Enhancement Network (UPRE-Net) for few-shot medical image segmentation. Specifically, we present a dual-support branch to maximize the extraction of information from support images through augmentation techniques. To enhance the reliability of prototypes, we propose an Uncertainty-guided Prototype Generation (UPG) module. Within the UPG module, we first extract both global and local prototypes for each class and then apply uncertainty measures to select the most informative prototypes. Additionally, to effectively combine the prediction results from the dual-support branch, we present a Reliable Dynamic Fusion (RDF) module. This module dynamically integrates the two prediction results to generate a more reliable output. Furthermore, we present an Uncertainty-induced Weighted Loss (UWL) to ensure that the model pays more attention to these regions with high uncertainty. Experiments on four benchmark medical image datasets demonstrate that our proposed model significantly outperforms state-of-the-art methods. The code will be released at https://github.com/taozh2017/UPRENet. Tao Zhou 0002, Kaiwen Huang 0002, Yi Zhou 0007, Haofeng Zhang 0001, Boqiang Fan, Huazhu Fu |
IEEE Trans. Medical Imaging | 7 |
| 2026 | Improving Learning of New Diseases Through Knowledge-Enhanced Initialization for Federated Adapter TuningabstractIn healthcare, federated learning (FL) is a widely adopted framework that enables privacy-preserving collaboration among medical institutions. With large foundation models (FMs) demonstrating impressive capabilities, using FMs in FL through cost-efficient adapter tuning has become a popular approach. Given the rapidly evolving healthcare environment, it is crucial for individual clients to quickly adapt to new tasks or diseases by tuning adapters while drawing upon past experiences. In this work, we introduce Federated Knowledge-Enhanced Initialization (FedKEI), a novel framework that leverages cross-client and cross-task transfer from past knowledge to generate informed initializations for learning new tasks with adapters. FedKEI begins with a global clustering process at the server to generalize knowledge across tasks, followed by the optimization of aggregation weights across clusters (inter-cluster weights) and within each cluster (intra-cluster weights) to personalize knowledge transfer for each new task. To facilitate more effective learning of the inter- and intra-cluster weights, we adopt a bi-level optimization scheme that collaboratively learns the global intra-cluster weights across clients and optimizes the local inter-cluster weights toward each client's task objective. Extensive experiments on three benchmark datasets of different modalities, including dermatology, chest X-rays, and retinal OCT, demonstrate FedKEI's advantage in adapting to new diseases compared to state-of-the-art methods. Danni Peng, Yuan Wang 0008, Kangning Cai, Peiyan Ning, Jiming Xu, Yong Liu 0026, Rick Siow Mong Goh, Qingsong Wei, Huazhu Fu |
IEEE Trans. Medical Imaging | 9 |
| 2026 | Prompting Lipschitz-Constrained Network for Multiple-in-One Sparse-View CT ReconstructionabstractDespite significant advancements in deep learning-based sparse-view computed tomography (SVCT) reconstruction algorithms, these methods still encounter two primary limitations: (I) It is challenging to explicitly prove that the prior networks of deep unfolding algorithms satisfy Lipschitz constraints due to their empirically designed nature. (II) The substantial storage costs of training a separate model for each setting in the case of multiple views hinder practical clinical applications. To address these issues, we elaborate an explicitly provable Lipschitz-constrained network, dubbed LipNet, and integrate an explicit prompt module to provide discriminative knowledge of different sparse sampling settings, enabling the treatment of multiple sparse view configurations within a single model. Furthermore, we develop a storage-saving deep unfolding framework for multiple-in-one SVCT reconstruction, termed PromptCT, which embeds LipNet as its prior network to ensure the convergence of its corresponding iterative algorithm. In simulated and real data experiments, PromptCT outperforms benchmark reconstruction algorithms in multiple-in-one SVCT reconstruction, achieving higher-quality reconstructions with lower storage costs. On the theoretical side, we explicitly demonstrate that LipNet satisfies boundary property, further proving its Lipschitz continuity and subsequently analyzing the convergence of the proposed iterative algorithms. The data and code are publicly available at https://github.com/shibaoshun/PromptCT. Baoshun Shi, Qiusheng Lian, Xinran Yu, Huazhu Fu |
IEEE Trans. Medical Imaging | 5 |
| 2026 | M2Net: Multimodal Multitask Mutual Learning for Anti-VEGF Efficacy PredictionabstractAge-related macular degeneration with abnormal blood vessel growth (neovascular AMD) is the leading cause of vision loss in elderly populations. While anti-VEGF injections are the standard treatment, they present financial burdens for patients and vary in effectiveness. Predicting treatment efficacy is therefore crucial for patient care. Current prediction methods fail to fully integrate information from different imaging techniques, typically focusing on either forecasting vision improvements or generating post-treatment images-but not both simultaneously. This approach overlooks the important relationship between these tasks. We present M2Net, a novel joint generation and classification network based on Multimodal Multitask Mutual learning, to simultaneously predict changes in visual acuity and generate post-treatment retinal images. M2Net employs a dual-branch structure that processes both fundus photographs and Optical Coherence Tomography (OCT) scans to improve prediction accuracy. Our framework includes two key innovations: the Multimodal Collaborative Treatment Efficacy Prediction module, which interacts the features between the two modalities and provides initial visual acuity change classification to guide the generation of post-treatment images; and the Pre-Post Treatment Image Joint Analysis module, which identifies both common and changing features between pre-treatment and post-treatment images to enhance prediction accuracy. To validate our approach, we created the dataset (MMPD) containing paired multimodal retinal images with corresponding visual acuity measurements. Experiments on the dataset demonstrate that M2Net achieves superior performance compared to existing methods, with a classification accuracy of 96.03%, an SSIM of 0.6377 on the OCT modality, and an SSIM of 0.8347 on the fundus modality. Our code will be available at https://github.com/zengying123/M2Net. Lei Bi 0001, Wuzhen Shi, Huazhu Fu, Bin Sheng 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2026 | LCM-Net: LLM-Driven Cross-Modality MoE Feature Fusion Network for Cancer Survival AnalysisabstractCancer survival analysis aims to predict survival outcomes to evaluate the efficacy and prognosis of treatment. Although current approaches have designed diverse cross-modal learning methods to integrate genetic data and pathology images, they are frequently hindered by data redundancy. Pattern representation in high-dimensional genetic data remains a significant hurdle. Pathology data analysis is computationally intensive because of the giga-pixel resolution. Moreover, the heterogeneity of data types poses a barrier to extending multimodal fusion methods. To address the aforementioned issues, we propose a novel LLM-driven Cross-Modality MoE-feature Fusion Network (LCM-Net) with three innovative modules for boosting cancer survival prediction. Specifically, the Genomic Language Alignment (GLA) module integrates genomic features with learnable prompts. Utilizing large language models, it encodes genomic information into concise and semantically relevant representations. Then, we devise the Pathological Feature Refinement (PFR) module to serve as a plug-and-play component that filters out irrelevant regions in pathology images. Finally, we propose a Multimodal Expert Integration (MEI) module to effectively leverage the capabilities of different experts, integrating the processed features from both the genomic and pathological domains. Extensive experiments on five public datasets demonstrate that our approach outperforms state-of-the-art methods, and the ablation study confirms the effectiveness of the proposed modules. Our code is publicly available at https://github.com/script-Yang/LCM-Net. Sicheng Yang 0001, Haipeng Zhou, Weiming Wang 0002, Shifu Chen, Guang Yang 0006, Huazhu Fu, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 7 |
| 2026 | Temporal Prompt Learning With Depth Memory for Video Mirror DetectionabstractMirror detection in dynamic scenes plays a crucial role in ensuring safety for various applications, such as drone tracking and robot navigation. However, current mirror detection models often fail in areas with mirrors that have a similar visual and color appearance to their surrounding objects. They also struggle to generalize well in complex cases, primarily due to limited annotated datasets. In this work, we propose a novel temporal prompt learning network with depth memory (TPD-Net) to address these critical challenges. Our approach includes several key components. First, we introduce a Temporal Prompt Generator (TPG) to learn temporal prompt features. Then, we devise Multi-layer Depth-aware Adaptor (MDA) modules to progressively adapt prompt features from the TPG, thereby learning mirror-related features by embedding temporal depth information as guidance. Moreover, we further refine these mirror-related features by constructing a depth memory and a Depth Memory Read module to read the temporal depths stored in the memory, boosting video mirror detection. Experimental results on a benchmark dataset show that our TPD-Net significantly outperforms 22 state-of-the-art methods in video mirror detection tasks. Our code, models, and results are publicly available athttps://github.com/ge-xing/TPDNet. Zhaohu Xing, Tian Ye 0001, Xin Yang 0011, Sixiang Chen, Huazhu Fu, Yan Nei Law, Lei Zhu 0003 |
IEEE Trans. Multim. | 5 |
| 2025 | AIF-SFDA: Autonomous Information Filter Driven Source-Free Domain Adaptation for Medical Image SegmentationabstractDecoupling domain-variant information (DVI) from domain-invariant information (DII) serves as a prominent strategy for mitigating domain shifts in the practical implementation of deep learning algorithms. However, in medical settings, concerns surrounding data collection and privacy often restrict access to both training and test data, hindering the empirical decoupling of information by existing methods. To tackle this issue, we propose an Adaptive Information Filter-driven Source-free Domain Adaptation (AIF-SFDA) algorithm, which leverages a frequency-based learnable information filter to autonomously decouple DVI and DII. Information Bottleneck (IB) and Self-supervision (SS) are incorporated to optimize the learnable frequency filter. The IB governs the information flow within the filter to diminish redundant DVI, while SS preserves DII in alignment with the specific task and image modality. Thus, the adaptive information filter can overcome domain shifts relying solely on target data. A series of experiments covering various medical image modalities and segmentation tasks were conducted to demonstrate the benefits of AIF-SFDA through comparisons with leading algorithms and ablation studies. Haojin Li 0003, Heng Li 0010, Rihan Zhong, Ke Niu 0002, Huazhu Fu, Jiang Liu 0001 |
AAAI | 6 |
| 2025 | Look Back for More: Harnessing Historical Sequential Updates for Personalized Federated Adapter TuningabstractPersonalized federated learning (PFL) studies effective model personalization to address the data heterogeneity issue among clients in traditional federated learning (FL). Existing PFL approaches mainly generate personalized models by relying solely on the clients' latest updated models while ignoring their previous updates, which may result in suboptimal personalized model learning. To bridge this gap, we propose a novel framework termed pFedSeq, designed for personalizing adapters to fine-tune a foundation model in FL. In pFedSeq, the server maintains and trains a sequential learner, which processes a sequence of past adapter updates from clients and generates calibrations for personalized adapters. To effectively capture the cross-client and cross-step relations hidden in previous updates and generate high-performing personalized adapters, pFedSeq adopts the powerful selective state space model (SSM) as the architecture of sequential learner. Through extensive experiments on four public benchmark datasets, we demonstrate the superiority of pFedSeq over state-of-the-art PFL methods. Danni Peng, Yuan Wang 0008, Huazhu Fu, Jinpeng Jiang, Yong Liu 0026, Rick Siow Mong Goh, Qingsong Wei |
AAAI | 3 |
| 2025 | Vision-Language Model IP Protection via Prompt-based LearningabstractVision-language models (VLMs) like CLIP (Contrastive Language-Image Pre-Training) have seen remarkable success in visual recognition, highlighting the increasing need to safeguard the intellectual property (IP) of well-trained models. Effective IP protection extends beyond ensuring authorized usage; it also necessitates restricting model deployment to authorized data domains, particularly when the model is fine-tuned for specific target domains. However, current IP protection methods often rely solely on the visual backbone, which may lack sufficient semantic richness. To bridge this gap, we introduce IP-CLIP, a lightweight IP protection strategy tailored to CLIP, employing a prompt-based learning approach. By leveraging the frozen visual backbone of CLIP, we extract both image style and content information, incorporating them into the learning of IP prompt. This strategy acts as a robust barrier, effectively preventing the unauthorized transfer of features from authorized domains to unauthorized ones. Additionally, we propose a style-enhancement branch that constructs feature banks for both authorized and unauthorized domains. This branch integrates self-enhanced and cross-domain features, further strengthening IP-CLIP’s capability to block features from unauthorized domains. Finally, we present new three metrics designed to better balance the performance degradation of authorized and unauthorized domains. Comprehensive experiments in various scenarios demonstrate its promising potential for application in IP protection tasks for VLMs. Lianyu Wang, Huazhu Fu, Daoqiang Zhang |
CVPR | 3 |
| 2025 | A Simple Data Augmentation for Feature Distribution Skewed Federated LearningabstractFederated Learning (FL) facilitates collaborative learning among multiple clients in a distributed manner and ensures the security of privacy. However, its performance inevitably degrades with non-Independent and Identically Distributed (non-IID) data. In this paper, we focus on the feature distribution skewed FL scenario, a common non-IID situation in real-world applications where data from different clients exhibit varying underlying distributions. This variation leads to feature shift, which is a key issue of this scenario. While previous works have made notable progress, few pay attention to the data itself, i.e., the root of this issue. The primary goal of this paper is to mitigate feature shift from the perspective of data. To this end, we propose a simple yet remarkably effective input-level data augmentation method, namely FedRDN, which randomly injects the statistical information of the local distribution from the entire federation into the client’s data. This is beneficial to improve the generalization of local feature representations, thereby mitigating feature shift. Moreover, our FedRDN is a plug-and-play component, which can be seamlessly integrated into the data augmentation flow with only a few lines of code. Extensive experiments on several datasets show that the performance of various representative FL methods can be further improved by integrating our FedRDN, demonstrating its effectiveness, strong compatibility and generalizability. Code is available at https://github.com/IAMJackYan/FedRDN. Yunlu Yan, Huazhu Fu, Yuexiang Li, Jinheng Xie, Jun Ma 0008, Guang Yang 0006, Lei Zhu 0003 |
CVPR | 2 |
| 2025 | MExD: An Expert-Infused Diffusion Model for Whole-Slide Image ClassificationabstractWhole Slide Image (WSI) classification poses unique challenges due to the vast image size and numerous non-informative regions, which introduce noise and cause data imbalance during feature aggregation. To address these issues, we propose MExD, an Expert-Infused Diffusion Model that combines the strengths of a Mixture-of-Experts (MoE) mechanism with a diffusion model for enhanced classification. MExD balances patch feature distribution through a novel MoE-based aggregator that selectively emphasizes relevant information, effectively filtering noise, addressing data imbalance, and extracting essential features. These features are then integrated via a diffusion-based generative process to directly yield the class distribution for the WSI. Moving beyond conventional discriminative approaches, MExD represents the first generative strategy in WSI classification, capturing fine-grained details for robust and precise results. Our MExD is validated on three widely-used benchmarks—Camelyon16, TCGANSCLC, and BRACS—consistently achieving state-of-the-art performance in both binary and multi-class tasks. Our code and model are available at https://github.com/JWZhao-uestc/MExD. Xin Li 0079, Fan Yang 0054, Qiang Zhai, Ao Luo, Yang Zhao 0024, Hong Cheng 0002, Huazhu Fu |
CVPR | 8 |
| 2025 | History-Aware and Dynamic Client Contribution in Federated LearningabstractFederated Learning (FL) is a collaborative machine learning (ML) approach, where multiple clients participate in training an ML model without exposing their private data. Fair and accurate assessment of client contributions facilitates incentive allocation in FL and encourages diverse clients to participate in a unified model training. Existing methods for contribution assessment adopts a co-operative game-theoretic concept, called Shapley value, but under restricted assumptions, e.g., all clients’ participating in all epochs or at least in one epoch of FL. We propose a history-aware client contribution assessment framework, called FLContrib, where client-participation is dynamic, i.e., a subset of clients participates in each epoch. The theoretical underpinning of FLContrib is based on the Markovian training process of FL. Under this setting, we directly apply the linearity property of Shapley value and compute a historical timeline of client contributions. Considering the possibility of a limited computational budget, we propose a two-sided fairness criteria to schedule Shapley value computation in a subset of epochs. Empirically, FLContrib is efficient and consistently accurate in estimating contribution across multiple utility functions. As a practical application, we apply FLContrib to detect dishonest clients in FL based on historical Shaplee values. Bishwamittra Ghosh, Debabrota Basu, Huazhu Fu, Yuan Wang 0008, Renuga Kanagavelu, Jinpeng Jiang, Yong Liu 0026, Rick Siow Mong Goh, Qingsong Wei |
ECAI | 3 |
| 2025 | Teaching AI the Anatomy Behind the Scan: Addressing Anatomical Flaws in Medical Image Segmentation with Learnable Prior
Young Seok Jeon, Hongfei Yang, Huazhu Fu, Mengling Feng |
ICCV | 3 |
| 2025 | GEMeX: A Large-Scale, Groundable, and Explainable Medical VQA Benchmark for Chest X-Ray DiagnosisabstractMedical Visual Question Answering (Med-VQA) combines computer vision and natural language processing to automatically answer clinical inquiries about medical images. However, current Med-VQA datasets exhibit two significant limitations: (1) they often lack visual and textual explanations for answers, hindering comprehension for patients and junior doctors; (2) they typically offer a narrow range of question formats, inadequately reflecting the diverse requirements in practical scenarios. These limitations pose significant challenges to the development of a reliable and user-friendly Med-VQA system. To address these challenges, we introduce a large-scale, Groundable, and Explainable Medical VQA benchmark for chest X-ray diagnosis (GEMeX), featuring several innovative components: (1) a multi-modal explainability mechanism that offers detailed visual and textual explanations for each question-answer pair, thereby enhancing answer comprehensibility; (2) four question types, open-ended, closed-ended, single-choice, and multiple-choice, to better reflect practical needs. With 151,025 images and 1,605,575 questions, GEMeX is the currently largest chest X-ray VQA dataset. Evaluation of 12 representative large vision language models (LVLMs) on GEMeX reveals suboptimal performance, underscoring the dataset's complexity. Meanwhile, we propose a strong model by fine-tuning an existing LVLM on the GEMeX training set. The substantial performance improvement showcases the dataset's effectiveness. The benchmark is available at https://www.med-vqa.com/GEMeX. Bo Liu 0049, Ke Zou, Li-Ming Zhan, Chengqiang Xie, Jiannong Cao 0001, Xiao-Ming Wu 0003, Huazhu Fu |
ICCV | 10 |
| 2025 | SABPI-Net: A Novel Structure-Aware Network for Accurate and Domain-Invariant Retinopathy of Prematurity Diagnosis
Shaobin Chen, Huazhu Fu, Tao Tan 0002, Jiaju Huang, Xiangyu Xiong, Zhenquan Wu, Behdad Dashtbozorg, Bai Ying Lei, Yue Sun 0001 |
MICCAI (10) | 3 |
| 2025 | Cycle Context Verification for In-Context Medical Image Segmentation
Shishuai Hu, Zehui Liao, Liangli Zhen, Huazhu Fu, Yong Xia 0001 |
MICCAI (1) | 4 |
| 2025 | Text-Driven Multiplanar Visual Interaction for Semi-supervised Medical Image Segmentation
Kaiwen Huang 0002, Yi Zhou 0007, Huazhu Fu, Yizhe Zhang 0001, Chen Gong 0002, Tao Zhou 0002 |
MICCAI (5) | 3 |
| 2025 | No More Sliding Window: Efficient 3D Medical Image Segmentation with Differentiable Top-K Patch Sampling
Young Seok Jeon, Hongfei Yang, Huazhu Fu, Yeshe M. Kway, Mengling Feng |
MICCAI (16) | 3 |
| 2025 | MicroMIL: Graph-Based Multiple Instance Learning for Context-Aware Diagnosis with Microscopic Images
Bryan Wong, Huazhu Fu, Willmer Rafell Quiñones Robles, Young Sin Ko, Mun Yong Yi |
MICCAI (1) | 3 |
| 2025 | RIFNet: Bridging Modalities for Accurate and Detailed Ocular Disease Analysis
Qingshan Hou, Peng Cao 0001, Jianguo Ju, Meng Wang 0001, Ke Zou, Huazhu Fu, Osmar R. Zaïane |
MICCAI (13) | 9 |
| 2025 | Vision-Amplified Semantic Entropy for Hallucination Detection in Medical Visual Question Answering
Zehui Liao, Shishuai Hu, Ke Zou, Huazhu Fu, Liangli Zhen, Yong Xia 0001 |
MICCAI (5) | 4 |
| 2025 | Parameterized Diffusion Optimization Enabled Autoregressive Ordinal Regression for Diabetic Retinopathy Grading
Qinkai Yu, Wei Zhou 0021, Hantao Liu, Yanyu Xu 0001, Meng Wang 0038, Yitian Zhao, Huazhu Fu, Xujiong Ye, Yalin Zheng, Yanda Meng |
MICCAI (15) | 7 |
| 2025 | GEMeX-RMCoT: An Enhanced Med-VQA Dataset for Region-Aware Multimodal Chain-of-Thought ReasoningabstractMedical visual question answering aims to support clinical decision-making by enabling models to answer natural language questions based on medical images. While recent advances in multi-modal learning have significantly improved performance, current methods still suffer from limited answer reliability and poor interpretability, impairing the ability of clinicians and patients to understand and trust model outputs. To address these limitations, this work first proposes a Region-Aware Multimodal Chain-of-Thought (RMCoT) dataset, in which the process of producing an answer is preceded by a sequence of intermediate reasoning steps that explicitly ground relevant visual regions of the medical image, thereby providing fine-grained explainability. Furthermore, we introduce a novel verifiable reward mechanism for reinforcement learning to guide post-training, improving the alignment between the model's reasoning process and its final answer. Remarkably, our method achieves comparable performance using only one-eighth of the training data, demonstrating the efficiency and effectiveness of the proposal. The dataset is available at https://www.med-vqa.com/GEMeX/. Bo Liu 0113, Along He, Huazhu Fu, Xiao-Ming Wu 0003 |
ACM Multimedia | 5 |
| 2025 | Few-Shot Learning from Gigapixel Images via Hierarchical Vision-Language Alignment and ModelingabstractVision-language models (VLMs) have recently been integrated into multiple instance learning (MIL) frameworks to address the challenge of few-shot, weakly supervised classification of whole slide images (WSIs). A key trend involves leveraging multi-scale information to better represent hierarchical tissue structures. However, existing methods often face two key limitations: (1) insufficient modeling of interactions within the same modalities across scales (e.g., 5x and 20x) and (2) inadequate alignment between visual and textual modalities on the same scale. To address these gaps, we propose HiVE-MIL, a hierarchical vision-language framework that constructs a unified graph consisting of (1) parent–child links between coarse (5x) and fine (20x) visual/textual nodes to capture hierarchical relationships, and (2) heterogeneous intra-scale edges linking visual and textual nodes on the same scale. To further enhance semantic consistency, HiVE-MIL incorporates a two-stage, text-guided dynamic filtering mechanism that removes weakly correlated patch–text pairs, and introduces a hierarchical contrastive loss to align textual semantics across scales. Extensive experiments on TCGA breast, lung, and kidney cancer datasets demonstrate that HiVE-MIL consistently outperforms both traditional MIL and recent VLM-based MIL approaches, achieving gains of up to 4.1% in macro F1 under 16-shot settings. Our results demonstrate the value of jointly modeling hierarchical structure and multimodal alignment for efficient and scalable learning from limited pathology data. The code is available at https://github.com/bryanwong17/HiVE-MIL. Bryan Wong, Huazhu Fu, Mun Yong Yi |
NeurIPS | 3 |
| 2025 | Unpaired Fundus Image Enhancement Using Image DecompositionabstractABSTRACT Low‐quality fundus images pose significant challenges for both ophthalmologists and computer‐aided diagnosis systems. While many existing deep learning‐based image quality enhancement algorithms require low‐ and high‐quality image pairs for training, such pairs are often difficult to obtain in practice. On the other hand, unpaired image enhancement algorithms tend to struggle in preserving small structures and suppressing artefacts, which are crucial for medical applications. To address these issues, we propose an unpaired structure‐preserving cycle quality alternating network for low‐quality fundus image enhancement. Our method consists of three main components: (1) a cycle quality alternating framework to provide pixel‐wise supervision for unpaired image enhancement, (2) a quality‐aware disentangle module to enhance the extrinsic representation of the low‐quality image with the high‐quality reference image, and (3) an instance normalized skip to improve the network's structure‐preserving capability. We tested our method on both synthetic and authentic clinical images with pathological structures and found it to be superior to state‐of‐the‐art algorithms in terms of improving image quality while preserving delicate structures. Additionally, the proposed network demonstrated strong generalization ability in improving the quality of unseen images, as tested on 135‐degree neonatal fundus images. Kun Chen 0005, Yu Ye 0003, Huazhu Fu, Yuhao Luo 0001, Ronald X. Xu, Mingzhai Sun |
IET Image Process. | 3 |
| 2025 | WeakPolyp-SAM: Segment Anything Model-driven weakly-supervised polyp segmentation
Tao Zhou 0002, Yunqi Gu, Yi Zhou 0007, Yizhe Zhang 0001, Ye Wu 0001, Huazhu Fu |
Knowl. Based Syst. | 7 |
| 2025 | AdaptFRCNet: Semi-supervised adaptation of pre-trained model with frequency and region consistency for medical image segmentation
Along He, Yanlin Wu, Tao Li 0022, Huazhu Fu |
Medical Image Anal. | 5 |
| 2025 | Beyond the eye: A relational model for early dementia detection using retinal OCTA images
Shouyue Liu, Jinkui Hao, Yonghuai Liu, Huazhu Fu, Yitian Zhao |
Medical Image Anal. | 6 |
| 2025 | Diff-UNet: A diffusion embedded network for robust 3D medical image segmentation
Zhaohu Xing, Huazhu Fu, Guang Yang 0006, Lequan Yu, Bai Ying Lei, Lei Zhu 0003 |
Medical Image Anal. | 3 |
| 2025 | PathFL: Multi-alignment Federated Learning for pathology image segmentation
Yuan Zhang 0019, Yaolei Qi, Guanyu Yang 0001, Huazhu Fu |
Medical Image Anal. | 5 |
| 2025 | DVPT: Dynamic Visual Prompt Tuning of large pre-trained models for medical image analysis
Along He, Yanlin Wu, Tao Li 0022, Huazhu Fu |
Neural Networks | 5 |
| 2025 | Neovascularization Segmentation via a Multilateral Interaction-Enhanced Graph Convolutional NetworkabstractChoroidal neovascularization (CNV), a primary characteristic of wet age-related macular degeneration (wet AMD), represents a leading cause of blindness worldwide. In clinical practice, optical coherence tomography angiography (OCTA) is commonly used for studying CNV-related pathological changes, due to its micron-level resolution and non-invasive nature. Thus, accurate segmentation of CNV regions and vessels in OCTA images is crucial for clinical assessment of wet AMD. However, challenges existed due to irregular CNV shapes and imaging limitations like projection artifacts, noises and boundary blurring. Moreover, the lack of publicly available datasets constraints the CNV analysis. To address these challenges, this paper constructs the first publicly accessible CNV dataset (CNVSeg), and proposes a novel multilateral graph convolutional interaction-enhanced CNV segmentation network (MTG-Net). This network integrates both region and vessel morphological information, exploring semantic and geometric duality constraints within the graph domain. Specifically, MTG-Net consists of a multi-task framework and two graph-based cross-task modules: Multilateral Interaction Graph Reasoning (MIGR) and Multilateral Reinforcement Graph Reasoning (MRGR). The multi-task framework encodes rich geometric features of lesion shapes and surfaces, decoupling the image into three task-specific feature maps. MIGR and MRGR iteratively reason about higher-order relationships across tasks through a graph mechanism, enabling complementary optimization for task-specific objectives. Additionally, an uncertainty-weighted loss is proposed to mitigate the impact of artifacts and noise on segmentation accuracy. Experimental results demonstrate that MTG-Net outperforms existing methods, achieving a Dice socre of 87.21% for region segmentation and 88.12% for vessel segmentation. Tao Chen 0003, Dan Zhang 0026, Da Chen 0002, Huazhu Fu, Shanshan Wang 0002, Laurent D. Cohen, Yitian Zhao, Quanyong Yi, Jiong Zhang 0004 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Uncertainty-Aware Medical Diagnostic Phrase Identification and GroundingabstractMedical phrase grounding is crucial for identifying relevant regions in medical images based on phrase queries, facilitating accurate image analysis and diagnosis. However, current methods rely on manual extraction of key phrases from medical reports, reducing efficiency and increasing the workload for clinicians. Additionally, the lack of model confidence estimation limits clinical trust and usability. In this paper, we introduce a novel task-Medical Report Grounding (MRG)-which aims to directly identify diagnostic phrases and their corresponding grounding boxes from medical reports in an end-to-end manner. To address this challenge, we propose uMedGround, a a robust and reliable framework that leverages a multimodal large language model to predict diagnostic phrases by embedding a unique token, < $\mathtt {BOX}$BOX >, into the vocabulary to enhance detection capabilities. A vision encoder-decoder processes the embedded token and input image to generate grounding boxes. Critically, uMedGround incorporates an uncertainty-aware prediction model, significantly improving the robustness and reliability of grounding predictions. Experimental results demonstrate that uMedGround outperforms state-of-the-art medical phrase grounding methods and fine-tuned large visual-language models, validating its effectiveness and reliability. This study represents a pioneering exploration of the MRG task, marking the first-ever endeavor in this domain. Additionally, we demonstrate the applicability of uMedGround in medical visual question answering and class-based localization tasks, where it highlights visual evidence aligned with key diagnostic phrases, supporting clinicians in interpreting various types of textual inputs, including free-text reports, visual question answering queries, and class labels. Ke Zou, Yang Bai 0011, Bo Liu 0113, Zhihao Chen 0004, Yang Zhou 0017, Xuedong Yuan, Meng Wang 0038, Xiaojing Shen, Xiaochun Cao, Huazhu Fu |
IEEE Trans. Pattern Anal. Mach. Intell. | 12 |
| 2025 | Dual-scale enhanced and cross-generative consistency learning for semi-supervised medical image segmentation
Yunqi Gu, Tao Zhou 0002, Yizhe Zhang 0001, Yi Zhou 0007, Kelei He, Chen Gong 0002, Huazhu Fu |
Pattern Recognit. | 7 |
| 2025 | 3D microvascular reconstruction in retinal OCT angiography images via domain-adaptive learning
Jiong Zhang 0004, Yonghuai Liu, Dan Zhang 0026, Jianyang Xie, Tao Chen 0003, Yalin Zheng, Huazhu Fu, Yitian Zhao |
Pattern Recognit. | 8 |
| 2025 | Reliable Federated Disentangling Network for Non-IID Domain FeatureabstractFederated Learning (FL), as an efficient decentralized distributed learning approach, enables multiple institutions to collaboratively train a model without sharing their local data. Despite its advantages, the performance of FL models is substantially impacted by the domain feature shift arising from different acquisition devices/clients. Moreover, existing FL methods often prioritize accuracy without considering reliability factors such as confidence or uncertainty, leading to unreliable predictions in safety-critical applications. Thus, our goal is to enhance FL performance by addressing non-domain feature issues and ensuring model reliability. In this study, we introduce a novel approach named RFedDis (Reliable Federated Disentangling Network). RFedDis leverages feature disentangling to capture a global domain-invariant cross-client representation while preserving local client-specific feature learning. Additionally, we incorporate an uncertainty-aware decision fusion mechanism to effectively integrate the decoupled features. This ensures dynamic integration at the evidence level, producing reliable predictions accompanied by estimated uncertainties. Therefore, RFedDis is the FL approach to combine evidential uncertainty with feature disentangling, enhancing both performance and reliability in handling non-IID domain features. Extensive experimental results demonstrate that RFedDis outperforms other state-of-the-art FL approaches, providing outstanding performance coupled with a high degree of reliability. Meng Wang 0038, Kai Yu 0009, Chun-Mei Feng 0001, Yiming Qian, Ke Zou, Lianyu Wang, Rick Siow Mong Goh, Xinxing Xu, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Big Data | 10 |
| 2025 | Hierarchical Context Transformer for Multi-Level Semantic Scene UnderstandingabstractA comprehensive and explicit understanding of surgical scenes plays a vital role in developing context-aware computer-assisted systems in the operating theatre. However, few works provide systematical analysis to enable hierarchical surgical scene understanding. In this work, we propose to represent the tasks set [phase recognition$\rightarrow $step recognition$\rightarrow $action and instrument detection] as multi-level semantic scene understanding (MSSU). For this target, we propose a novel hierarchical context transformer (HCT) network and thoroughly explore the relations across the different level tasks. Specifically, a hierarchical relation aggregation module (HRAM) is designed to concurrently relate entries inside multi-level interaction information and then augment task-specific features. To further boost the representation learning of the different tasks, inter-task contrastive learning (ICL) is presented to guide the model to learn task-wise features via absorbing complementary information from other tasks. Furthermore, considering the computational costs of the transformer, we propose HCT+ to integrate the spatial and temporal adapter to access competitive performance on substantially fewer tunable parameters. Extensive experiments on our cataract dataset and a publicly available endoscopic PSI-AVA dataset demonstrate the outstanding performance of our method, consistently exceeding the state-of-the-art methods by a large margin. The code is available athttps://github.com/Aurora-hao/HCT. Luoying Hao, Huazhu Fu, Jinming Duan 0001, Jiang Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | A Contrast-Aware Edge Enhancement GAN for Unpaired Anterior Segment OCT Image DenoisingabstractAnterior segment optical coherence tomography (AS-OCT) is a popular imaging technique that can directly visualize the anterior segment structures while inherent speckle noise severely impairs visual readability and subsequent clinical analysis. Though unpaired OCT image denoising algorithms have been developed to improve visual quality considering the limited supervised clinical data, preserving the edge structures while denoising remains challenging, especially in AS-OCT images with little hierarchy and low contrast. This work proposes an edge enhancement generative adversarial network ($E^{2}GAN$) based contrast-aware, particularly for unpaired AS-OCT image denoising. Specifically, to improve edge-structure consistency, we design a contrast attention mechanism for exploiting diverse hierarchical knowledge from multiple contrast images and adopt particular gradient-guided speckle filtering modules with an edge preservation loss for stabilizing the network. Additionally, considering that bi-directional GANs often focus on global appearance rather than essential features,$E^{2}GAN$adds a perceptual quality constraint into the cycle consistency. Extensive experiments validate the superiority of$E^{2}GAN$for AS-OCT image denoising and the benefits for downstream clinical analysis. Further experiments on the synthetic retinal OCT images prove the generalization of$E^{2}GAN$. Sanqian Li, Risa Higashita, Huazhu Fu, Jiang Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Vivim: A Video Vision Mamba for Ultrasound Video SegmentationabstractUltrasound video segmentation gains increasing attention in clinical practice due to the redundant dynamic references in video frames. However, traditional convolutional neural networks have a limited receptive field and transformer-based networks are unsatisfactory in constructing long-term dependency from the perspective of computational complexity. This bottleneck poses a significant challenge when processing longer sequences in medical video analysis tasks using available devices with limited memory. Recently, state space models (SSMs), famous by Mamba, have exhibited linear complexity and impressive achievements in efficient long sequence modeling, which have developed deep neural networks by expanding the receptive field on many vision tasks significantly. Unfortunately, vanilla SSMs failed to simultaneously capture causal temporal cues and preserve non-casual spatial information. To this end, this paper presents a Video Vision Mamba-based framework, dubbed as Vivim, for ultrasound video segmentation tasks. Our Vivim can effectively compress the long-term spatiotemporal representation into sequences at varying scales with our designed Temporal Mamba Block. We also introduce an improved boundary-aware affine constraint across frames to enhance the discriminative ability of Vivim on ambiguous lesions. Extensive experiments on thyroid segmentation in ultrasound videos, breast lesion segmentation in ultrasound videos, and polyp segmentation in colonoscopy videos demonstrate the effectiveness and efficiency of our Vivim, superior to existing methods. The code and dataset are available at: https://github.com/scott-yjyang/Vivim. Zhaohu Xing, Lequan Yu, Huazhu Fu, Chunwang Huang, Lei Zhu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Toward Reliable Medical Image Segmentation by Modeling Evidential Calibrated UncertaintyabstractMedical image segmentation is critical for disease diagnosis and treatment assessment. However, concerns regarding the reliability of segmentation regions persist among clinicians, mainly attributed to the absence of confidence assessment, robustness, and calibration to accuracy. To address this, we introduce deep evidential segmentation model (DEviS), an easily implementable foundational model that seamlessly integrates into various medical image segmentation networks. DEviS not only enhances the calibration and robustness of baseline segmentation accuracy but also provides high-efficiency uncertainty estimation for reliable predictions. By leveraging subjective logic theory, we explicitly model probability and uncertainty for medical image segmentation. Here, the Dirichlet distribution parameterizes the distribution of probabilities for different classes of the segmentation results. To generate calibrated predictions and uncertainty, we develop a trainable calibrated uncertainty penalty. Furthermore, DEviS incorporates an uncertainty-aware filtering (UAF) module, which designs the metric of uncertainty-calibrated error to filter out-of-distribution (OOD) data. We conducted validation studies on publicly available datasets, including ISIC2018, KiTS2021, LiTS2017, and BraTS2019, to assess the accuracy and robustness of different backbone segmentation models enhanced by DEviS, as well as the efficiency and reliability of uncertainty estimation. Additionally, two potential clinical trials were conducted using the UAF module. The clinical application conducted on the Johns Hopkins OCT and Duke OCT-DME datasets demonstrated the effectiveness of the model in filtering OOD data. The second trial evaluated its efficacy in filtering high-quality data on the FIVES datasets. At last, the proposed DEviS method was extended to semi-supervised medical image segmentation, where it exhibited strong robustness under noisy conditions. Our code has been released in https://github.com/Cocofeat/DEviS. Ke Zou, Ling Huang 0003, Xuedong Yuan, Xiaojing Shen, Meng Wang 0038, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Cybern. | 11 |
| 2025 | Lighting is Unreliable: Adversarial Video Relighting Against rPPG Heart Rate Measurement
Menglin Zhang, Xiaoxin Guo, Xiaofeng Cao 0002, Shuifa Sun, Huazhu Fu, Qing Guo 0005 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Uncertainty-Aware Cross-Training for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning has gained considerable popularity in medical image segmentation tasks due to its capability to reduce reliance on expert-examined annotations. Several mean-teacher (MT) based semi-supervised methods utilize consistency regularization to effectively leverage valuable information from unlabeled data. However, these methods often heavily rely on the student model and overlook the potential impact of cognitive biases within the model. Furthermore, some methods employ co-training using pseudo-labels derived from different inputs, yet generating high-confidence pseudo-labels from perturbed inputs during training remains a significant challenge. In this paper, we propose an Uncertainty-aware Cross-training framework for semi-supervised medical image Segmentation (UC-Seg). Our UC-Seg framework incorporates two distinct subnets to effectively explore and leverage the correlation between them, thereby mitigating cognitive biases within the model. Specifically, we present a Cross-subnet Consistency Preservation (CCP) strategy to enhance feature representation capability and ensure feature consistency across the two subnets. This strategy enables each subnet to correct its own biases and learn shared semantics from both labeled and unlabeled data. Additionally, we propose an Uncertainty-aware Pseudo-label Generation (UPG) component that leverages segmentation results and corresponding uncertainty maps from both subnets to generate high-confidence pseudo-labels. We extensively evaluate the proposed UC-Seg on various medical image segmentation tasks involving different modality images, such as MRI, CT, ultrasound, colonoscopy, and so on. The results demonstrate that our method achieves superior segmentation accuracy and generalization performance compared to other state-of-the-art semi-supervised methods. Our code and segmentation maps will be released at https://github.com/taozh2017/UCSeg. Kaiwen Huang 0002, Tao Zhou 0002, Huazhu Fu, Yizhe Zhang 0001, Yi Zhou 0007, Xiaojun Wu 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | VSR-Net: Vessel-Like Structure Rehabilitation Network With Graph ClusteringabstractThe morphologies of vessel-like structures, such as blood vessels and nerve fibres, play significant roles in disease diagnosis, e.g., Parkinson's disease. Although deep network-based refinement segmentation and topology-preserving segmentation methods recently have achieved promising results in segmenting vessel-like structures, they still face two challenges: 1) existing methods often have limitations in rehabilitating subsection ruptures in segmented vessel-like structures; 2) they are typically overconfident in predicted segmentation results. To tackle these two challenges, this paper attempts to leverage the potential of spatial interconnection relationships among subsection ruptures from the structure rehabilitation perspective. Based on this perspective, we propose a novel Vessel-like Structure Rehabilitation Network (VSR-Net) to both rehabilitate subsection ruptures and improve the model calibration based on coarse vessel-like structure segmentation results. VSR-Net first constructs subsection rupture clusters via a Curvilinear Clustering Module (CCM). Then, the well-designed Curvilinear Merging Module (CMM) is applied to rehabilitate the subsection ruptures to obtain the refined vessel-like structures. Extensive experiments on six 2D/3D medical image datasets show that VSR-Net significantly outperforms state-of-the-art (SOTA) refinement segmentation methods with lower calibration errors. Additionally, we provide quantitative analysis to explain the morphological difference between the VSR-Net's rehabilitation results and ground truth (GT), which are smaller compared to those between SOTA methods and GT, demonstrating that our method more effectively rehabilitates vessel-like structures. Haili Ye, Xiaoqing Zhang 0001, Huazhu Fu, Jiang Liu 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Adversarial Exposure Attack on Diabetic Retinopathy Imagery GradingabstractDiabetic Retinopathy (DR) is a leading cause of vision loss around the world. To help diagnose it, numerous cutting-edge works have built powerful deep neural networks (DNNs) to automatically grade DR via retinal fundus images (RFIs). However, RFIs are commonly affected by camera exposure issues that may lead to incorrect grades. The mis-graded results can potentially pose high risks to an aggravation of the condition. In this paper, we study this problem from the viewpoint of adversarial attacks. We identify and introduce a novel solution to an entirely new task, termed as adversarial exposure attack, which is able to produce natural exposure images and mislead the state-of-the-art DNNs. We validate our proposed method on a real-world public DR dataset with three DNNs, e.g., ResNet50, MobileNet, and EfficientNet, demonstrating that our method achieves high image quality and success rate in transferring the attacks. Our method reveals the potential threats to DNN-based automatic DR grading and would benefit the development of exposure-robust DR grading methods in the future. Yupeng Cheng, Qing Guo 0005, Felix Juefei-Xu, Huazhu Fu, Shangwei Lin 0001, Weisi Lin |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Score Prior Guided Iterative Solver for Speckles Removal in Optical Coherent Tomography ImagesabstractOptical coherence tomography (OCT) is a widely used non-invasive imaging modality for ophthalmic diagnosis. However, the inherent speckle noise becomes the leading cause of OCT image quality, and efficient speckle removal algorithms can improve image readability and benefit automated clinical analysis. As an ill-posed inverse problem, it is of utmost importance for speckle removal to learn suitable priors. In this work, we develop a score prior guided iterative solver (SPIS) with logarithmic space to remove speckles in OCT images. Specifically, we model the posterior distribution of raw OCT images as a data consistency term and transform the speckle removal from a nonlinear into a linear inverse problem in the logarithmic domain. Subsequently, the learned prior distribution through the score function from the diffusion model is utilized as a constraint for the data consistency term into the linear inverse optimization, resulting in an iterative speckle removal procedure that alternates between the score prior predictor and the subsequent non-expansive data consistency corrector. Experimental results on the private and public OCT datasets demonstrate that the proposed SPIS has an excellent performance in speckle removal and out-of-distribution (OOD) generalization. Further downstream automatic analysis on the OCT images verifies that the proposed SPIS can benefit clinical applications. Sanqian Li, Risa Higashita, Huazhu Fu, Jiang Liu 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | AIPNet: Action-Instance Progressive Learning Network for Instrument-Tissue Interaction DetectionabstractInstrument-tissue interaction detection, a task aimed at understanding surgical scenes from videos, holds immense importance in constructing computer-assisted surgery systems. Existing methods for this task consist of two stages: instance detection and interaction prediction. This sequential and separate model structure limits both effectiveness and efficiency, making it difficult to deploy on surgical robotic platforms. In this paper, we propose an end-to-end Action-Instance Progressive Learning Network (AIPNet) for the task. The model operates in three steps: action detection, instance detection, and action class refinement. Starting with coarse-scale proposals, the model progressively refines them into coarse-grained actions, which then serve as proposals for instance detection. The action prediction results are further refined using instance features through late fusion. These progressive learning processes improve the performance of the end-to-end model. Additionally, we introduce Dynamic Proposal Generators (DPG) to create dynamic adaptive learnable proposals for each video frame. To address the training challenges of this multi-task model, semantic supervised training is introduced to transfer prior language knowledge, and a training label strategy is proposed to generate unrelated instrument-tissue pair labels for enhanced supervision. Experimental results on PhacoQ and CholecQ datasets show that the proposed method achieves superior accuracy and faster processing speed than state-of-the-art models. Luoying Hao, Huazhu Fu, Chee-Kong Chui, Jiang Liu 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | $\text{MR}^{2}$-Net: Retinal OCTA Image Stitching via Multi-Scale Representation Learning and Dynamic Location GuidanceabstractOptical coherence tomography angiography (OCTA) plays a crucial role in quantifying and analyzing retinal vascular diseases. However, the limited field of view (FOV) inherent in most commercial OCTA imaging systems poses a significant challenge for clinicians, restricting the possibility to analyze larger retinal regions of high resolution. Automatic stitching of OCTA scans in adjacent regions may provide a promising solution to extend the region of interest. However, commonly-used stitching algorithms face difficulties in achieving effective alignment due to noise, artifacts and dense vasculature present in OCTA images. To address these challenges, we propose a novel retinal OCTA image stitching network, named -Net, which integrates multi-scale representation learning and dynamic location guidance. In the first stage, an image registration network with a progressive multi-resolution feature fusion is proposed to derive deep semantic information effectively. Additionally, we introduce a dynamic guidance strategy to locate the foveal avascular zone (FAZ) and constrain registration errors in overlapping vascular regions. In the second stage, an image fusion network based on multiple mask constraints and adjacent image aggregation (AIA) strategies is developed to further eliminate the artifacts in the overlapping areas of stitched images, thereby achieving precise vessel alignment. To validate the effectiveness of our method, we conduct a series of experiments on two delicately constructed datasets, i.e., OPTOVUE-OCTA and SVision-OCTA. Experimental results demonstrate that our method outperforms other image stitching methods and effectively generates high-quality wide-field OCTA images, achieving a structural similarity index (SSIM) score of 0.8264 and 0.8014 on the two datasets, respectively. Haiting Mao, Yuhui Ma, Dan Zhang 0026, Yanda Meng, Shaodong Ma, Yuchuan Qiao, Huazhu Fu, Caifeng Shan, Da Chen 0002, Yitian Zhao, Jiong Zhang 0004 |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Uncertainty-Inspired Multi-Task Learning in Arbitrary Scenarios of ECG MonitoringabstractAs the scenarios for electrocardiogram (ECG) monitoring become increasingly diverse, particularly with the development of wearable ECG, the influence of ambiguous factors in diagnosis has been amplified. Reliable ECG information must be extracted from abundant noises and confusing artifacts. To address this issue, we suggest an uncertainty-inspired model for beat-level diagnosis (UI-Beat). The base architecture of UI-Beat separates heartbeat localization and event diagnosis in two branches to address the problem of heterogeneous data sources. To disentangle the epistemic and aleatoric uncertainty within one stage in a deterministic neural network, we propose a new method derived from uncertainty formulation and realize it by introducing the class-biased transformation. Then the disentangled uncertainty can be utilized to screen out noise and identify ambiguous heartbeat synchronously. The results indicate that UI-Beat can significantly improve the performance of noise detection (from 91.60% to 97.50% for real-world noise detection and from 61.40% to 82.41% for real-world artifact detection). For multi-lead ECG analysis, UI-Beat is approaching the performance upper bound in heartbeat localization (only 15 false positives and 9 false negatives out of the 175,907 heartbeats in the INCART database) and achieving a significant performance improvement in heartbeat classification through uncertainty-based cross-lead fusion compared to single-lead prediction and other state-of-the-art methods (an average improvement of 14.28% for detecting heartbeats of S and 3.37% for detecting heartbeats of V). Considering the characteristic of one-stage ECG analysis within one model, it is suggested that the proposed UI-Beat has the potential to be employed as a general model for arbitrary scenarios of ECG monitoring, with the capacity to remove unusableepisodes, and realize heartbeat-level diagnosis with confidence provided. Xingyao Wang 0001, Hongxiang Gao, Caiyun Ma, Tingting Zhu 0001, Feng Yang 0011, Chengyu Liu 0001, Huazhu Fu |
IEEE J. Biomed. Health Informatics | 7 |
| 2025 | Guest Editorial: Special Issue on Foundation Models in Medical Imaging
Jiong Zhang 0004, Huazhu Fu, Caroline Petitjean, Xiaoxiao Li 0001, Julia A. Schnabel |
IEEE J. Biomed. Health Informatics | 2 |
| 2025 | Learnable Prompting SAM-Induced Knowledge Distillation for Semi-Supervised Medical Image SegmentationabstractThe limited availability of labeled data has driven advancements in semi-supervised learning for medical image segmentation. Modern large-scale models tailored for general segmentation, such as the Segment Anything Model (SAM), have revealed robust generalization capabilities. However, applying these models directly to medical image segmentation still exposes performance degradation. In this paper, we propose a learnable prompting SAM-induced Knowledge distillation framework (KnowSAM) for semi-supervised medical image segmentation. Firstly, we propose a Multi-view Co-training (MC) strategy that employs two distinct sub-networks to employ a co-teaching paradigm, resulting in more robust outcomes. Secondly, we present a Learnable Prompt Strategy (LPS) to dynamically produce dense prompts and integrate an adapter to fine-tune SAM specifically for medical image segmentation tasks. Moreover, we propose SAM-induced Knowledge Distillation (SKD) to transfer useful knowledge from SAM to two sub-networks, enabling them to learn from SAM's predictions and alleviate the effects of incorrect pseudo-labels during training. Notably, the predictions generated by our subnets are used to produce mask prompts for SAM, facilitating effective inter-module information exchange. Extensive experimental results on various medical segmentation tasks demonstrate that our model outperforms the state-of-the-art semi-supervised segmentation approaches. Crucially, our SAM distillation framework can be seamlessly integrated into other semi-supervised segmentation methods to enhance performance. The code will be released upon acceptance of this manuscript at https://github.com/taozh2017/KnowSAM. Kaiwen Huang 0002, Tao Zhou 0002, Huazhu Fu, Yizhe Zhang 0001, Yi Zhou 0007, Chen Gong 0002, Dong Liang 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Multi-View Test-Time Adaptation for Semantic Segmentation in Clinical Cataract SurgeryabstractCataract surgery, a widely performed operation worldwide, is incorporating semantic segmentation to advance computer-assisted intervention. However, the tissue appearance and illumination in cataract surgery often differ among clinical centers, intensifying the issue of domain shifts. While domain adaptation offers remedies to the shifts, the necessity for data centralization raises additional privacy concerns. To overcome these challenges, we propose a Multi-view Test-time Adaptation algorithm (MUTA) to segment cataract surgical scenes, which leverages multi-view learning to enhance model training within the source domain and model adaptation within the target domain. In the training phase, the segmentation model is equipped with multi-view decoders to boost its robustness against variations in cataract surgery. During the inference phase, test-time adaptation is implemented using multi-view knowledge distillation, enabling model updates in clinics without data centralization or privacy concerns. We conducted experiments in a simulated cross-center scenario using several cataract surgery datasets to evaluate the effectiveness of MUTA. Through comparisons and investigations, we have validated that MUTA effectively learns a robust source model and adapts the model to target data during the practical inference phase. Code and datasets are available at https://github.com/liamheng/CAI-algorithms. Heng Li 0010, Mingyang Ou, Haojin Li 0003, Zhongxi Qiu, Ke Niu 0002, Huazhu Fu, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | CoD-MIL: Chain-of-Diagnosis Prompting Multiple Instance Learning for Whole Slide Image ClassificationabstractMultiple instance learning (MIL) has emerged as a prominent paradigm for processing the whole slide image with pyramid structure and giga-pixel size in digital pathology. However, existing attention-based MIL methods are primarily trained on the image modality and a pre-defined label set, leading to limited generalization and interpretability. Recently, vision language models (VLM) have achieved promising performance and transferability, offering potential solutions to the limitations of MIL-based methods. Pathological diagnosis is an intricate process that requires pathologists to examine the WSI step-by-step. In the field of natural language process, the chain-of-thought (CoT) prompting method is widely utilized to imitate the human reasoning process. Inspired by the CoT prompt and pathologists' clinic knowledge, we propose a chain-of-diagnosis prompting multiple instance learning (CoD-MIL) framework for whole slide image classification. Specifically, the chain-of-diagnosis text prompt decomposes the complex diagnostic process in WSI into progressive sub-processes from low to high magnification. Additionally, we propose a text-guided contrastive masking module to accurately localize the tumor region by masking the most discriminative instances and introducing the guidance of normal tissue texts in a contrastive way. Extensive experiments conducted on three real-world subtyping datasets demonstrate the effectiveness and superiority of CoD-MIL. Jiangbo Shi, Chen Li 0011, Tieliang Gong, Chunbao Wang 0002, Huazhu Fu |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Serp-Mamba: Advancing High-Resolution Retinal Vessel Segmentation With Selective State-Space ModelabstractUltra-Wide-Field Scanning Laser Ophthalmoscopy (UWF-SLO) images capture high-resolution views of the retina with typically spanning 200 degrees. Accurate segmentation of vessels in UWF-SLO images is essential for detecting and diagnosing fundus disease. Recent studies highlight that Mamba's selective State Space Model (SSM) excels in modeling long-range dependencies with linear computational complexity, making it highly suitable for preserving the continuity of elongated vessel structures, especially for high-resolution UWF images. Inspired by this, we propose the Serpentine Mamba (Serp-Mamba) network to address this challenging task. Specifically, we recognize the intricate, varied, and delicate nature of the tubular structure of vessels. Furthermore, the high-resolution of UWF-SLO images exacerbates the imbalance between the vessel and background categories. Based on the above observations, we first devise a Serpentine Interwoven Adaptive (SIA) scan mechanism, which scans UWF-SLO images along curved vessel structures in a snake-like crawling manner. This approach, consistent with vascular texture transformations, ensures the effective and continuous capture of curved vascular structure features. Second, we propose an Ambiguity-Driven Dual Recalibration (ADDR) module to address the category imbalance problem intensified by high-resolution images. Our ADDR module delineates pixels by two learnable thresholds and refines ambiguous pixels through a dual-driven strategy, thereby accurately distinguishing vessels and background regions. Experiment results on three datasets demonstrate the superior performance of our Serp-Mamba on high-resolution vessel segmentation. We also conduct a series of ablation studies to verify the impact of our designs. Our code will be released upon publication (https://github.com/whq-xxh/Serp-Mamba). Hongqiu Wang, Bin Sheng 0001, Huazhu Fu, Guang Yang 0006, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | DiffMIC-v2: Medical Image Classification via Improved Diffusion NetworkabstractRecently, Denoising Diffusion Models have achieved outstanding success in generative image modeling and attracted significant attention in the computer vision community. Although a substantial amount of diffusion-based research has focused on generative tasks, few studies apply diffusion models to medical diagnosis. In this paper, we propose a diffusion-based network (named DiffMIC-v2) to address general medical image classification by eliminating unexpected noise and perturbations in image representations. To achieve this goal, we first devise an improved dual-conditional guidance strategy that conditions each diffusion step with multiple granularities to enhance step-wise regional attention. Furthermore, we design a novel Heterologous diffusion process that achieves efficient visual representation learning in the latent space. We evaluate the effectiveness of our DiffMIC-v2 on four medical classification tasks with different image modalities, including thoracic diseases classification on chest X-ray, placental maturity grading on ultrasound images, skin lesion classification using dermatoscopic images, and diabetic retinopathy grading using fundus images. Experimental results demonstrate that our DiffMIC-v2 outperforms state-of-the-art methods by a significant margin, which indicates the universality and effectiveness of the proposed model on multi-class and multi-label classification tasks. DiffMIC-v2 can use fewer iterations than our previous DiffMIC to obtain accurate estimations, and also achieves greater runtime efficiency with superior results. The code will be publicly available at https://github.com/scott-yjyang/DiffMICv2. Huazhu Fu, Angelica I. Avilés-Rivero, Zhaohu Xing, Lei Zhu 0003 |
IEEE Trans. Medical Imaging | 2 |
| 2025 | Training-Free Image Style Alignment for Domain Shift on Handheld Ultrasound DevicesabstractHandheld ultrasound devices face usage limitations due to user inexperience and cannot benefit from supervised deep learning without extensive expert annotations. Moreover, the models trained on standard ultrasound device data are constrained by training data distribution and perform poorly when directly applied to handheld device data. In this study, we propose the Training-free Image Style Alignment (TISA) to align the style of handheld device data to those of standard devices. The proposed TISA eliminates the demand for source data, and can transform the image style while preserving spatial context during testing. Furthermore, our TISA avoids continuous updates to the pre-trained model compared to other test-time methods and is suited for clinical applications. We show that TISA performs better and more stably in medical detection and segmentation tasks for handheld device data than other test-time adaptation methods. We further validate TISA as the clinical model for automatic measurements of spinal curvature and carotid intima-media thickness, and the automatic measurements agree well with manual measurements made by human experts. We demonstrate the potential for TISA to facilitate automatic diagnosis on handheld ultrasound devices and expedite their eventual widespread use. Code is available at https://github.com/zenghy96/TISA. Hongye Zeng, Ke Zou, Zhihao Chen 0004, Yuchong Gao, Kang Zhou 0001, Meng Wang 0038, Chang Jiang 0001, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu |
IEEE Trans. Medical Imaging | 14 |
| 2025 | Uncertainty-Driven Edge Prompt Generation Network for Medical Image SegmentationabstractSegment Anything Model (SAM) is a foundational image segmentation model, which shows superior performance for natural image segmentation tasks. Several SAM-based medical image segmentations have been proposed. However, these SAM-based medical image segmentation methods heavily depend on prior manual guidance involving points, boxes, and coarse-grained masks, which lack adaptability and flexibility. Moreover, the inherent challenge of edge blurring in medical images is critical, as it directly affects the quality of segmentation. To address these challenges, we propose an uncertainty-driven edge prompt generation network for medical image segmentation, called UDEG-Net. Specifically, to better adapt to medical image segmentation, we fine-tune the encoder by using Low-Rank Adaptation (LoRA) technology to enhance the encoder's learning capability and capture enriched medical image features. Furthermore, to overcome the limitations of interactive prompts, we develop an auto edge prompt generator to generate edge prompt information and further enhance the structural representation. Finally, to focus on the high-uncertainty edge areas, we introduce an evidence-based uncertainty estimation and a progressive uncertainty-driven loss to drive the auto edge prompt generator to yield robust edge prompt information and reliable segmentation results. Experimental results on three public datasets and one private dataset show that our UDEG-Net outperforms the state-of-the-art medical image segmentation methods. Junyong Zhao, Liang Sun 0009, Dingwei Fan, Kun Wang 0056, Haipeng Si, Huazhu Fu, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Topicwise Separable Sentence Retrieval for Medical Report GenerationabstractAutomated radiology reporting holds immense clinical potential in alleviating the burdensome workload of radiologists and mitigating diagnostic bias. Recently, retrieval-based report generation methods have garnered increasing attention. These methods predefine a set of candidate queries and compose reports by searching for sentences in an off-the-shelf sentence gallery that best match these candidate queries. However, due to the long-tail distribution of the training data, these models tend to learn frequently occurring sentences and topics, overlooking the rare topics. Regrettably, in many cases, the descriptions of rare topics often indicate critical findings that should be mentioned in the report. To address this problem, we introduce a Topicwise Separable Sentence Retrieval (Teaser) for medical report generation. To ensure comprehensive learning of both common and rare topics, we categorize queries into common and rare types to learn differentiated topics, and then propose Topic Contrastive Loss to effectively align topics and queries in the latent space. Moreover, we integrate an Abstractor module following the extraction of visual features, which aids the topic decoder in gaining a deeper understanding of the visual observational intent. Experiments on the MIMIC-CXR and IU X-ray datasets demonstrate that Teaser surpasses state-of-the-art models, while also validating its capability to effectively represent rare topics and establish more dependable correspondences between queries and topics. The code is available at https://github.com/CindyZJT/Teaser.git. Junting Zhao, Yang Zhou 0017, Zhihao Chen 0004, Huazhu Fu |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Masked Vascular Structure Segmentation and Completion in Retinal ImagesabstractEarly retinal vascular changes in diseases such as diabetic retinopathy often occur at a microscopic level. Accurate evaluation of retinal vascular networks at a micro-level could significantly improve our understanding of angiopathology and potentially aid ophthalmologists in disease assessment and management. Multiple angiogram-related retinal imaging modalities, including fundus, optical coherence tomography angiography, and fluorescence angiography, project continuous, inter-connected retinal microvascular networks into imaging domains. However, extracting the microvascular network, which includes arterioles, venules, and capillaries, is challenging due to the limited contrast and resolution. As a result, the vascular network often appears as fragmented segments. In this paper, we propose a backbone-agnostic Masked Vascular Structure Segmentation and Completion (MaskVSC) method to reconstruct the retinal vascular network. MaskVSC simulates missing sections of blood vessels and uses this simulation to train the model to predict the missing parts and their connections. This approach simulates highly heterogeneous forms of vessel breaks and mitigates the need for massive data labeling. Accordingly, we introduce a connectivity loss function that penalizes interruptions in the vascular network. Our findings show that masking 40% of the segments yields optimal performance in reconstructing the interconnected vascular network. We test our method on three different types of retinal images across five separate datasets. The results demonstrate that MaskVSC outperforms state-of-the-art methods in maintaining vascular network completeness and segmentation accuracy. Furthermore, MaskVSC has been introduced to different segmentation backbones and has successfully improved performance. The code and 2PFM data are available at: https://github.com/Zhouyi-Zura/MaskVSC. Yi Zhou 0024, Thiara Sana Ahmed, Meng Wang 0038, Eric A. Newman, Leopold Schmetterer, Huazhu Fu, Jun Cheng 0003, Bingyao Tan |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Pathology-Preserving Transformer Based on Multicolor Space for Low-Quality Medical Image EnhancementabstractMedical images acquired under suboptimal conditions often suffer from quality degradation, such as low-light, blurring, and artifacts. Such degradations obscure the lesions and anatomical structures in medical images, making it difficult to distinguish key pathological regions. This significantly increases the risk of misdiagnosis by automated medical diagnostic systems or clinicians. To address this challenge, we propose a multi-Color space-based quality enhancement network (MSQNet) that effectively eliminates global low-quality factors while preserving pathology-related characteristics for improved clinical observation and analysis. We first revisit the properties of image quality enhancement in different color spaces, where the V-channel in the HSV space can better represent the contrast and brightness enhancement process, whereas the A/B-channel in the LAB space is more focused on the color change of low-quality images. The proposed framework harnesses the unique properties of different color spaces to optimize the image enhancement process. Specifically, we propose a pathology-preserving transformer, designed to selectively aggregate features across different color spaces and enable comprehensive multiscale feature fusion. Leveraging these capabilities, MSQNet effectively enhances low-quality RGB medical images while preserving key pathological features, thereby establishing a new paradigm in medical image enhancement. Extensive experiments on three public medical image datasets demonstrate that MSQNet outperforms traditional enhancement techniques and state-of-the-art methods, in terms of both quantitative metrics and qualitative visual assessment. MSQNet successfully improves image quality while preserving pathological features and anatomical structures, facilitating accurate diagnosis and analysis by medical professionals and automated systems. Qingshan Hou, Yaqi Wang 0004, Peng Cao 0001, Jianguo Ju, Huijuan Tu, Xiaoli Liu 0001, Jinzhu Yang, Huazhu Fu, Osmar R. Zaïane |
IEEE Trans. Multim. | 8 |
| 2025 | EM-Trans: Edge-Aware Multimodal Transformer for RGB-D Salient Object DetectionabstractRGB-D salient object detection (SOD) has gained tremendous attention in recent years. In particular, transformer has been employed and shown great potential. However, existing transformer models usually overlook the vital edge information, which is a major issue restricting the further improvement of SOD accuracy. To this end, we propose a novel edge-aware RGB-D SOD transformer, called EM-Trans, which explicitly models the edge information in a dual-band decomposition framework. Specifically, we employ two parallel decoder networks to learn the high-frequency edge and low-frequency body features from the low- and high-level features extracted from a two-steam multimodal backbone network, respectively. Next, we propose a cross-attention complementarity exploration module to enrich the edge/body features by exploiting the multimodal complementarity information. The refined features are then fed into our proposed color-hint guided fusion module for enhancing the depth feature and fusing the multimodal features. Finally, the resulting features are fused using our deeply supervised progressive fusion module, which progressively integrates edge and body features for predicting saliency maps. Our model explicitly considers the edge information for accurate RGB-D SOD, overcoming the limitations of existing methods and effectively improving the performance. Extensive experiments on benchmark datasets demonstrate that EM-Trans is an effective RGB-D SOD framework that outperforms the current state-of-the-art models, both quantitatively and qualitatively. A further extension to RGB-T SOD demonstrates the promising potential of our model in various kinds of multimodal SOD tasks. Geng Chen 0001, Qingyue Wang, Bo Dong 0001, Ruitao Ma, Nian Liu 0002, Huazhu Fu, Yong Xia 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | DONet: Dual-Octave Network for Fast MR Image ReconstructionabstractMagnetic resonance (MR) image acquisition is an inherently prolonged process, whose acceleration has long been the subject of research. This is commonly achieved by obtaining multiple undersampled images, simultaneously, through parallel imaging. In this article, we propose the dual-octave network (DONet), which is capable of learning multiscale spatial-frequency features from both the real and imaginary components of MR data, for parallel fast MR image reconstruction. More specifically, our DONet consists of a series of dual-octave convolutions (Dual-OctConvs), which are connected in a dense manner for better reuse of features. In each Dual-OctConv, the input feature maps and convolutional kernels are first split into two components (i.e., real and imaginary) and then divided into four groups according to their spatial frequencies. Then, our Dual-OctConv conducts intragroup information updating and intergroup information exchange to aggregate the contextual information across different groups. Our framework provides three appealing benefits: 1) it encourages information interaction and fusion between the real and imaginary components at various spatial frequencies to achieve richer representational capacity; 2) the dense connections between the real and imaginary groups in each Dual-OctConv make the propagation of features more efficient by feature reuse; and 3) DONet enlarges the receptive field by learning multiple spatial-frequency features of both the real and imaginary components. Extensive experiments on two popular datasets (i.e., clinical knee and fastMRI), under different undersampling patterns and acceleration factors, demonstrate the superiority of our model in accelerated parallel MR image reconstruction. Chun-Mei Feng 0001, Zhanyuan Yang, Huazhu Fu, Yong Xu 0001, Jian Yang 0003, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2025 | Orthogonal Subspace Representation for Generative Adversarial NetworksabstractDisentanglement learning aims to separate explanatory factors of variation so that different attributes of the data can be well characterized and isolated, which promotes efficient inference for downstream tasks. Mainstream disentanglement approaches based on generative adversarial networks (GANs) learn interpretable data representation. However, most typical GAN-based works lack the discussion of the latent subspace, causing insufficient consideration of the variation of independent factors. Although some recent research analyzes the latent space on pretrained GANs for image editing, they do not emphasize learning representation directly from the subspace perspective. Appropriate subspace properties could facilitate corresponding feature representation learning to satisfy the independent variation requirements of the obtained explanatory factors, which is crucial for better disentanglement. In this work, we propose a unified framework for ensuring disentanglement, which fully investigates latent subspace learning (SL) in GAN. The novel GAN-based architecture explores orthogonal subspace representation (OSR) on vanilla GAN, named OSRGAN. To guide a subspace with strong correlation, less redundancy, and robust distinguishability, our OSR includes three stages, self-latent-aware, orthogonal subspace-aware, and structure representation-aware, respectively. First, the self-latent-aware stage promotes the latent subspace strongly correlated with the data space to discover interpretable factors, but with poor independence of variation. Second, the following orthogonal subspace-aware stage adaptively learns some 1-D linear subspace spanned by a set of orthogonal bases in the latent space. There is less redundancy between them, expressing the corresponding independence. Third, the structure representation-aware stage aligns the projection on the orthogonal subspace and the latent variables. Accordingly, feature representation in each linear subspace can be distinguishable, enhancing the independent expression of interpretable factors. In addition, we design an alternating optimization step, achieving a tradeoff training of OSRGAN on different properties. Despite it strictly constrains orthogonality, the loss weight coefficient of distinguishability induced by orthogonality could be adjusted and balanced with correlation constraint. To elucidate, this tradeoff training prevents our OSRGAN from overemphasizing any property and damaging the expressiveness of the feature representation. It takes into account both interpretable factors and their independent variation characteristics. Meanwhile, alternating optimization could keep the cost and efficiency of forward inference unchanged and will not burden the computational complexity. In theory, we clarify the significance of OSR, which brings better independence of factors, along with interpretability as correlation could converge to a high range faster. Moreover, through the convergence behavior analysis, including the objective functions under different constraints and the evaluation curve with iterations, our model demonstrates enhanced stability and definitely converges toward a higher peak for disentanglement. To depict the performance in downstream tasks, we compared the state-of-the-art GAN-based and even VAE-based approaches on different datasets. Our OSRGAN achieves higher disentanglement scores on FactorVAE, SAP, MIG, and VP metrics. All the experimental results illustrate that our novel GAN-based framework has considerable advantages on disentanglement. Hongxiang Jiang, Xiaoyan Luo, Jihao Yin, Huazhu Fu, Fuxiang Wang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Federated Noisy Client LearningabstractFederated learning (FL) collaboratively trains a shared global model depending on multiple local clients, while keeping the training data decentralized to preserve data privacy. However, standard FL methods ignore the noisy client issue, which may harm the overall performance of the shared model. We first investigate the critical issue caused by noisy clients in FL and quantify the negative impact of the noisy clients in terms of the representations learned by different layers. We have the following two key observations: 1) the noisy clients can severely impact the convergence and performance of the global model in FL and 2) the noisy clients can induce greater bias in the deeper layers than the former layers of the global model. Based on the above observations, we propose federated noisy client learning (Fed-NCL), a framework that conducts robust FL with noisy clients. Specifically, Fed-NCL first identifies the noisy clients through well estimating the data quality and model divergence. Then robust layerwise aggregation is proposed to adaptively aggregate the local models of each client to deal with the data heterogeneity caused by the noisy clients. We further perform label correction on the noisy clients to improve the generalization of the global model. Experimental results on various datasets demonstrate that our algorithm boosts the performances of different state-of-the-art systems with noisy clients. Our code is available at https://github.com/TKH666/Fed-NCL. Kahou Tam, Li Li 0064, Bo Han 0003, Cheng-Zhong Xu 0001, Huazhu Fu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | MedSegDiff-V2: Diffusion-Based Medical Image Segmentation with TransformerabstractThe Diffusion Probabilistic Model (DPM) has recently gained popularity in the field of computer vision, thanks to its image generation applications, such as Imagen, Latent Diffusion Models, and Stable Diffusion, which have demonstrated impressive capabilities and sparked much discussion within the community. Recent investigations have further unveiled the utility of DPM in the domain of medical image analysis, as underscored by the commendable performance exhibited by the medical image segmentation model across various tasks. Although these models were originally underpinned by a UNet architecture, there exists a potential avenue for enhancing their performance through the integration of vision transformer mechanisms. However, we discovered that simply combining these two models resulted in subpar performance. To effectively integrate these two cutting-edge techniques for the Medical image segmentation, we propose a novel Transformer-based Diffusion framework, called MedSegDiff-V2. We verify its effectiveness on 20 medical image segmentation tasks with different image modalities. Through comprehensive evaluation, our approach demonstrates superiority over prior state-of-the-art (SOTA) methodologies. Code is released at https://github.com/KidsWithTokens/MedSegDiff. Wei Ji 0011, Huazhu Fu, Min Xu 0009, Yueming Jin, Yanwu Xu 0001 |
AAAI | 3 |
| 2024 | MIAD-MARK: Adversarial Watermarking of Medical Image for Protecting Copyright and PrivacyabstractThe rapid advancement of deep learning has significantly facilitated the integration of Artificial Intelligence (AI) into clinical practices. However, the frequent utilization of vast clinical data has raised concerns about copyright and patient privacy. Here, we introduce a framework that not only enables medical image copyright protection but also prevents unauthorized AI analysis. Specifically, we adhere to this framework and propose a novel visible adversarial watermark for medical images, MIAD-MARK, utilizing adaptable affine transforms to deceive unauthorized models. Furthermore, we enhance the robustness of MIAD-MARK to resist advanced watermark removal deep neural networks. Our approach involves linear variations, allowing for reversibility to recover the original images during the authorization process. We conduct experiments on various medical datasets, including different diseases and modalities. Our results demonstrate significant decreases in medical image foundation models and standard models. Our findings underscore that MIAD-MARK offers an effective, easily implemented, and robust solution to safeguard medical image copyright and patient privacy, thereby promoting the security of AI-driven medical image diagnosis in clinical applications. Xingxing Wei 0001, Bangzheng Pu, Huazhu Fu |
BIBM | 4 |
| 2024 | ViLa-MIL: Dual-scale Vision-Language Multiple Instance Learning for Whole Slide Image ClassificationabstractMultiple instance learning (MIL)-based framework has become the mainstream for processing the whole slide image (WSI) with giga-pixel size and hierarchical image context in digital pathology. However, these methods heavily depend on a substantial number of bag-level labels and solely learn from the original slides, which are easily affected by variations in data distribution. Recently, vision language model (VLM)-based methods introduced the language prior by pre-training on large-scale pathological image-text pairs. However, the previous text prompt lacks the consideration of pathological prior knowledge, there-fore does not substantially boost the model's performance. Moreover, the collection of such pairs and the pre-training process are very time-consuming and source-intensive. To solve the above problems, we propose a dual-scale vision-language multiple instance learning (ViLa-MIL) framework for whole slide image classification. Specifically, we propose a dual-scale visual descriptive text prompt based on the frozen large language model (LLM) to boost the performance of VLM effectively. To transfer the VLM to process WSI efficiently, for the image branch, we propose a prototype-guided patch decoder to aggregate the patch features progressively by grouping similar patches into the same prototype; for the text branch, we introduce a context-guided text decoder to enhance the text features by incorporating the multi-granular image contexts. Extensive studies on three multi-cancer and multi-center subtyping datasets demonstrate the superiority of ViLa-MIL. Jiangbo Shi, Chen Li 0011, Tieliang Gong, Yefeng Zheng 0001, Huazhu Fu |
CVPR | 5 |
| 2024 | An Aggregation-Free Federated Learning for Tackling Data HeterogeneityabstractThe performance of Federated Learning (FL) hinges on the effectiveness of utilizing knowledge from distributed datasets. Traditional FL methods adopt an aggregate-then-adapt framework, where clients update local models based on a global model aggregated by the server from the previous training round. This process can cause client drift, especially with significant cross-client data heterogeneity, impacting model performance and convergence of the FL algorithm. To address these challenges, we introduce FedAF, a novel aggregation-free FL algorithm. In this framework, clients collaboratively learn condensed data by leveraging peer knowledge, the server subsequently trains the global model using the condensed data and soft labels received from the clients. FedAF inherently avoids the issue of client drift, enhances the quality of condensed data amid notable data heterogeneity, and improves the global model performance. Extensive numerical studies on several popular benchmark datasets show FedAF surpasses various state-of-the-art FL algorithms in handling label-skew and feature-skew data heterogeneity, leading to superior global model accuracy and faster convergence. Yuan Wang 0008, Huazhu Fu, Renuga Kanagavelu, Qingsong Wei, Yong Liu 0026, Rick Siow Mong Goh |
CVPR | 2 |
| 2024 | Self-Training Large Language and Vision Assistant for Medical Question AnsweringabstractLarge Vision-Language Models (LVLMs) have shown significant potential in assisting medical diagnosis by leveraging extensive biomedical datasets.However, the advancement of medical image understanding and reasoning critically depends on building high-quality visual instruction data, which is costly and labor-intensive to obtain, particularly in the medical domain.To mitigate this data-starving issue, we introduce Self-Training Large Language and Vision Assistant for Medicine (STLLaVA-Med).The proposed method is designed to train a policy model (an LVLM) capable of auto-generating medical visual instruction data to improve data efficiency, guided through Direct Preference Optimization (DPO).Specifically, a more powerful and larger LVLM (e.g., GPT-4o) is involved as a biomedical expert to oversee the DPO fine-tuning process on the auto-generated data, encouraging the policy model to align efficiently with human preferences.We validate the efficacy and data efficiency of STLLaVA-Med across three major medical Visual Question Answering (VQA) benchmarks, demonstrating competitive zero-shot performance with the utilization of only 9% of the medical data.Our implementation is available at https: //github.com/heliossun/STLLaVA-Med. Can Qin, Huazhu Fu, Zhiqiang Tao |
EMNLP | 3 |
| 2024 | FRCNet: Frequency and Region Consistency for Semi-supervised Medical Image Segmentation
Along He, Tao Li 0022, Yanlin Wu, Ke Zou, Huazhu Fu |
MICCAI (8) | 5 |
| 2024 | Open-Set Semi-supervised Medical Image Classification with Learnable Prototypes and Outlier Filter
Along He, Tao Li 0022, Yitian Zhao, Junyong Zhao, Huazhu Fu |
MICCAI (11) | 5 |
| 2024 | Memory-Efficient High-Resolution OCT Volume Synthesis with Cascaded Amortized Latent Diffusion Models
Xiao Ma 0011, Yuhan Zhang 0001, Songtao Yuan, Yong Liu 0026, Qiang Chen 0004, Huazhu Fu |
MICCAI (7) | 8 |
| 2024 | MedSynth: Leveraging Generative Model for Healthcare Data Sharing
Renuga Kanagavelu, Madhav Walia, Yuan Wang 0008, Huazhu Fu, Qingsong Wei, Yong Liu 0026, Rick Siow Mong Goh |
MICCAI (12) | 4 |
| 2024 | MPMNet: Modal Prior Mutual-Support Network for Age-Related Macular Degeneration Classification
Huaying Hao, Dan Zhang 0026, Huazhu Fu, Caifeng Shan, Yitian Zhao, Jiong Zhang 0004 |
MICCAI (1) | 4 |
| 2024 | DB-SAM: Delving into High Quality Universal Medical Image Segmentation
Jiale Cao, Huazhu Fu, Fahad Shahbaz Khan, Rao Muhammad Anwer |
MICCAI (12) | 3 |
| 2024 | MM-Retinal: Knowledge-Enhanced Foundational Pretraining with Fundus Image-Text Expertise
Chenran Zhang, Jianle Zhang, Yi Zhou 0007, Tao Zhou 0002, Huazhu Fu |
MICCAI (1) | 6 |
| 2024 | Multi-Scale Region-Aware Implicit Neural Network for Medical Images Matting
Yanyu Xu 0001, Yingzhi Xia, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026, Xinxing Xu |
MICCAI (9) | 3 |
| 2024 | CLIP-DR: Textual Knowledge-Guided Diabetic Retinopathy Grading with Ranking-Aware Prompting
Qinkai Yu, Jianyang Xie, Anh Nguyen 0003, He Zhao 0002, Jiong Zhang 0004, Huazhu Fu, Yitian Zhao, Yalin Zheng, Yanda Meng |
MICCAI (1) | 6 |
| 2024 | Reliable Source Approximation: Source-Free Unsupervised Domain Adaptation for Vestibular Schwannoma MRI Segmentation
Hongye Zeng, Ke Zou, Zhihao Chen 0004, Huazhu Fu |
MICCAI (10) | 5 |
| 2024 | MedMLP: An Efficient MLP-Like Network for Zero-Shot Retinal Image Classification
Menghan Zhou, Yanyu Xu 0001, Zhi Da Soh, Huazhu Fu, Rick Siow Mong Goh, Ching Yu Cheng, Yong Liu 0026, Liangli Zhen |
MICCAI (3) | 4 |
| 2024 | Resfusion: Denoising Diffusion Probabilistic Models for Image Restoration Based on Prior Residual NoiseabstractRecently, research on denoising diffusion models has expanded its application to the field of image restoration. Traditional diffusion-based image restoration methods utilize degraded images as conditional input to effectively guide the reverse generation process, without modifying the original denoising diffusion process. However, since the degraded images already include low-frequency information, starting from Gaussian white noise will result in increased sampling steps. We propose Resfusion, a general framework that incorporates the residual term into the diffusion forward process, starting the reverse process directly from the noisy degraded images. The form of our inference process is consistent with the DDPM. We introduced a weighted residual noise, named resnoise, as the prediction target and explicitly provide the quantitative relationship between the residual term and the noise term in resnoise. By leveraging a smooth equivalence transformation, Resfusion determine the optimal acceleration step and maintains the integrity of existing noise schedules, unifying the training and inference processes. The experimental results demonstrate that Resfusion exhibits competitive performance on ISTD dataset, LOL dataset and Raindrop dataset with only five sampling steps. Furthermore, Resfusion can be easily applied to image generation and emerges with strong versatility. Our code and model are available at https://github.com/nkicsl/Resfusion. Zhenning Shi, Haoshuai Zheng, Changsheng Dong, Bin Pan, Xueshuo Xie, Along He, Tao Li 0002, Huazhu Fu |
NeurIPS | 9 |
| 2024 | Out-Of-Distribution Detection with Diversification (Provably)abstractOut-of-distribution (OOD) detection is crucial for ensuring reliable deployment of machine learning models. Recent advancements focus on utilizing easily accessible auxiliary outliers (e.g., data from the web or other datasets) in training. However, we experimentally reveal that these methods still struggle to generalize their detection capabilities to unknown OOD data, due to the limited diversity of the auxiliary outliers collected. Therefore, we thoroughly examine this problem from the generalization perspective and demonstrate that a more diverse set of auxiliary outliers is essential for enhancing the detection capabilities. However, in practice, it is difficult and costly to collect sufficiently diverse auxiliary outlier data. Therefore, we propose a simple yet practical approach with a theoretical guarantee, termed Diversity-induced Mixup for OOD detection (diverseMix), which enhances the diversity of auxiliary outlier set for training in an efficient way. Extensive experiments show that diverseMix achieves superior performance on commonly used and recent challenging large-scale benchmarks, which further confirm the importance of the diversity of auxiliary outliers. Haiyun Yao, Zongbo Han, Huazhu Fu, Xi Peng 0001, Qinghua Hu, Changqing Zhang 0002 |
NeurIPS | 3 |
| 2024 | Rethinking the Reliability of Post-hoc Calibration Methods Under Subpopulation Shift
Huan Ma 0006, Changqing Zhang 0002, Bingzhe Wu, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu |
PRICAI (2) | 5 |
| 2024 | ViDSOD-100: A New Dataset and a Baseline Model for RGB-D Video Salient Object Detection
Lei Zhu 0003, Jiaxing Shen, Huazhu Fu, Qing Zhang 0006, Liansheng Wang 0002 |
Int. J. Comput. Vis. | 4 |
| 2024 | A principled framework for explainable multimodal disentanglement
Zongbo Han, Tao Luo 0014, Huazhu Fu, Qinghua Hu, Joey Tianyi Zhou, Changqing Zhang 0002 |
Inf. Sci. | 3 |
| 2024 | E2-MIL: An explainable and evidential multiple instance learning framework for whole slide image classification
Jiangbo Shi, Chen Li 0011, Tieliang Gong, Huazhu Fu |
Medical Image Anal. | 4 |
| 2024 | Confidence-aware multi-modality learning for eye disease screening
Ke Zou, Tian Lin 0002, Zongbo Han, Meng Wang 0001, Xuedong Yuan, Haoyu Chen 0002, Changqing Zhang 0002, Xiaojing Shen, Huazhu Fu |
Medical Image Anal. | 9 |
| 2024 | Say No to Freeloader: Protecting Intellectual Property of Your Deep ModelabstractModel intellectual property (IP) protection has gained attention due to the significance of safeguarding intellectual labor and computational resources. Ensuring IP safety for trainers and owners is critical, especially when ownership verification and applicability authorization are required. A notable approach involves preventing the transfer of well-trained models from authorized to unauthorized domains. We introduce a novel Compact Un-transferable Pyramid Isolation Domain (CUPI-Domain) which serves as a barrier against illegal transfers from authorized to unauthorized domains. Inspired by human transitive inference, the CUPI-Domain emphasizes distinctive style features of the authorized domain, leading to failure in recognizing irrelevant private style features on unauthorized domains. To this end, we propose CUPI-Domain generators, which select features from both authorized and CUPI-Domain as anchors. These generators fuse the style features and semantic features to create labeled, style-rich CUPI-Domain. Additionally, we design external Domain-Information Memory Banks (DIMB) for storing and updating labeled pyramid features to obtain stable domain class features and domain class-wise style features. Based on the proposed whole method, the novel style and discriminative loss functions are designed to effectively enhance the distinction in style and discriminative features between authorized and unauthorized domains. We offer two solutions for utilizing CUPI-Domain based on whether the unauthorized domain is known: target-specified CUPI-Domain and target-free CUPI-Domain. Comprehensive experiments on various public datasets demonstrate the effectiveness of our CUPI-Domain approach with different backbone models, providing an efficient solution for model intellectual property protection. Lianyu Wang, Meng Wang 0038, Huazhu Fu, Daoqiang Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Structure and Intensity Unbiased Translation for 2D Medical Image SegmentationabstractData distribution gaps often pose significant challenges to the use of deep segmentation models. However, retraining models for each distribution is expensive and time-consuming. In clinical contexts, device-embedded algorithms and networks, typically unretrainable and unaccessable post-manufacture, exacerbate this issue. Generative translation methods offer a solution to mitigate the gap by transferring data across domains. However, existing methods mainly focus on intensity distributions while ignoring the gaps due to structure disparities. In this paper, we formulate a new image-to-image translation task to reduce structural gaps. We propose a simple, yet powerful Structure-Unbiased Adversarial (SUA) network which accounts for both intensity and structural differences between the training and test sets for segmentation. It consists of a spatial transformation block followed by an intensity distribution rendering module. The spatial transformation block is proposed to reduce the structural gaps between the two images. The intensity distribution rendering module then renders the deformed structure to an image with the target intensity distribution. Experimental results show that the proposed SUA method has the capability to transfer both intensity distribution and structural content between multiple pairs of datasets and is superior to prior arts in closing the gaps for improving segmentation. Tianyang Miller, Shaoming Zheng, Jun Cheng 0003, Xi Jia, Joseph Bartlett, Xinxing Cheng, Zhaowen Qiu, Huazhu Fu, Jiang Liu 0001, Ales Leonardis, Jinming Duan 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2024 | Learning Physical-Spatio-Temporal Features for Video Shadow RemovalabstractShadow removal in a single image has received increasing attention in recent years. However, removing shadows over dynamic scenes remains largely under-explored. In this paper, we propose the first data-driven video shadow removal model, termed PSTNet, by exploiting three essential characteristics of video shadows, i.e., physical property, spatio relation, and temporal coherence. Specifically, a dedicated physical branch was established to conduct local illumination estimation, which is more applicable for scenes with complex lighting and textures, and then enhance the physical features via a mask-guided attention strategy. Then, we develop a progressive aggregation module to enhance the spatio and temporal characteristics of features maps, and effectively integrate the three kinds of features. Furthermore, to tackle the lack of datasets of paired shadow videos, we synthesize a dataset (SVSRD-85) with aid of the popular game GTAV by controlling the switch of the shadow renderer. Experiments against 9 state-of-the-art models, including image shadow removers and image/video restoration methods, show that our method improves the best SOTA in terms of RMSE error for the shadow area by 14.7%. In addition, we develop a lightweight model adaptation strategy to make our synthetic-driven model effective in real world scenes. The visual comparison on the public SBU-TimeLapse dataset verifies the generalization ability of our model in real scenes. Zhihao Chen 0004, Yefan Xiao, Lei Zhu 0003, Huazhu Fu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Stabilizing Multispectral Pedestrian Detection With Evidential Hybrid FusionabstractMultispectral pedestrian detection is an important task due to its critical role in a wide spectrum of applications. Basically, the complementary information from color and thermal images could provide a more accurate and reliable pedestrian detection result. However, multimodal data usually suffer from the issue of dynamic change or corruption for some modalities. At the same time, as a safety-critical task, how to produce a stable and reliable detection result is also a key challenge. To address these challenges, we propose a stable multispectral pedestrian detection (SMPD) algorithm, providing a new paradigm for multispectral detection by dynamically integrating different modalities at an evidence level. Specifically, we introduce the Dirichlet distribution to characterize the distribution of the class probabilities, parameterized with evidence from different modalities. Then, multi-branch fusion, based on Dempster-Shafer theory, can integrate these pieces of evidence to obtain the detection result. In addition, a Plug-and-Play module, termed modal enhancement module, is introduced to enhance cross-modality interaction. This is an end-to-end framework, which can induce accurate detection and uncertainty estimation, and then endows the model with both reliability and robustness against noise or corruption. Extensive experimental results demonstrate the efficiency of our algorithm compared with state-of-the-art methods. Qing Li 0018, Changqing Zhang 0002, Qinghua Hu, Pengfei Zhu 0001, Huazhu Fu, Lei Chen 0011 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Learning Motion-Guided Multi-Scale Memory Features for Video Shadow DetectionabstractNatural images often contain multiple shadow regions, and existing video shadow detection methods tend to fail in fully identifying all shadow regions, since they mainly learned temporal features at single-scale and single memory. In this work, we develop a novel convolutional neural network (CNN) to learn motion-guided multi-scale memory features to obtain multi-scale temporal information based on multiple network memories for boosting video shadow detection. To do so, our network first constructs three memories (i.e., a global memory, a local memory, and a motion memory) to combine spatial context and object motion for detecting shadows. Based on these three memories, we then devise a multi-scale motion-guided long-short transformer (MMLT) module to learn multi-scale temporal and motion memory features for predicting a shadow detection map of the input video frame. Our MMLT module includes a dense-scale long transformer (DLT), a dense-scale short transformer (DST), and a dense-scale motion transformer (DMT) to read three memories for learning multi-scale transformer features. Our DLT, DST, and DMT consist of a set of memory-read pooling attention (MPA) blocks and densely connect these output features of multiple MPA blocks to learn multi-scale transformer features since the scales of these output features are varied. By doing so, we can more accurately identify multiple shadow regions with different sizes from the input video. Moreover, we devise a self-supervised pretext task to pre-training the feature encoder for enhancing the downstream video shadow detection. Experimental results on three benchmark datasets show that our video shadow detection network quantitatively and qualitatively outperforms 26 state-of-the-art methods. Jiaxing Shen, Xin Yang 0011, Huazhu Fu, Qing Zhang 0006, Ping Li 0016, Bin Sheng 0001, Liansheng Wang 0002, Lei Zhu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Bilateral Supervision Network for Semi-Supervised Medical Image SegmentationabstractMassive high-quality annotated data is required by fully-supervised learning, which is difficult to obtain for image segmentation since the pixel-level annotation is expensive, especially for medical image segmentation tasks that need domain knowledge. As an alternative solution, semi-supervised learning (SSL) can effectively alleviate the dependence on the annotated samples by leveraging abundant unlabeled samples. Among the SSL methods, mean-teacher (MT) is the most popular one. However, in MT, teacher model's weights are completely determined by student model's weights, which will lead to the training bottleneck at the late training stages. Besides, only pixel-wise consistency is applied for unlabeled data, which ignores the category information and is susceptible to noise. In this paper, we propose a bilateral supervision network with bilateral exponential moving average (bilateral-EMA), named BSNet to overcome these issues. On the one hand, both the student and teacher models are trained on labeled data, and then their weights are updated with the bilateral-EMA, and thus the two models can learn from each other. On the other hand, pseudo labels are used to perform bilateral supervision for unlabeled data. Moreover, for enhancing the supervision, we adopt adversarial learning to enforce the network generate more reliable pseudo labels for unlabeled data. We conduct extensive experiments on three datasets to evaluate the proposed BSNet, and results show that BSNet can improve the semi-supervised segmentation performance by a large margin and surpass other state-of-the-art SSL methods. Along He, Tao Li 0022, Juncheng Yan, Kai Wang 0001, Huazhu Fu |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Diverse Data Generation for Retinal Layer Segmentation With Potential Structure ModelingabstractAccurate retinal layer segmentation on optical coherence tomography (OCT) images is hampered by the challenges of collecting OCT images with diverse pathological characterization and balanced distribution. Current generative models can produce high-realistic images and corresponding labels without quantitative limitations by fitting distributions of real collected data. Nevertheless, the diversity of their generated data is still limited due to the inherent imbalance of training data. To address these issues, we propose an image-label pair generation framework that generates diverse and balanced potential data from imbalanced real samples. Specifically, the framework first generates diverse layer masks, and then generates plausible OCT images corresponding to these layer masks using two customized diffusion probabilistic models respectively. To learn from imbalanced data and facilitate balanced generation, we introduce pathological-related conditions to guide the generation processes. To enhance the diversity of the generated image-label pairs, we propose a potential structure modeling technique that transfers the knowledge of diverse sub-structures from lowly- or non-pathological samples to highly pathological samples. We conducted extensive experiments on two public datasets for retinal layer segmentation. Firstly, our method generates OCT images with higher image quality and diversity compared to other generative methods. Furthermore, based on the extensive training with the generated OCT images, downstream retinal layer segmentation tasks demonstrate improved results. The code is publicly available at: https://github.com/nicetomeetu21/GenPSM. Xiao Ma 0011, Zetian Zhang, Yuhan Zhang 0001, Songtao Yuan, Huazhu Fu, Qiang Chen 0004 |
IEEE Trans. Medical Imaging | 6 |
| 2024 | Enhancing and Adapting in the Clinic: Source-Free Unsupervised Domain Adaptation for Medical Image EnhancementabstractMedical imaging provides many valuable clues involving anatomical structure and pathological characteristics. However, image degradation is a common issue in clinical practice, which can adversely impact the observation and diagnosis by physicians and algorithms. Although extensive enhancement models have been developed, these models require a well pre-training before deployment, while failing to take advantage of the potential value of inference data after deployment. In this paper, we raise an algorithm for source-free unsupervised domain adaptive medical image enhancement (SAME), which adapts and optimizes enhancement models using test data in the inference phase. A structure-preserving enhancement network is first constructed to learn a robust source model from synthesized training data. Then a teacher-student model is initialized with the source model and conducts source-free unsupervised domain adaptation (SFUDA) by knowledge distillation with the test data. Additionally, a pseudo-label picker is developed to boost the knowledge distillation of enhancement tasks. Experiments were implemented on ten datasets from three medical image modalities to validate the advantage of the proposed algorithm, and setting analysis and ablation studies were also carried out to interpret the effectiveness of SAME. The remarkable enhancement performance and benefits for downstream tasks demonstrate the potential and generalizability of SAME. The code is available at https://github.com/liamheng/Annotation-free-Medical-Image-Enhancement. Heng Li 0010, Ziqin Lin, Zhongxi Qiu, Zinan Li, Ke Niu 0002, Huazhu Fu, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Instrument-Tissue Interaction Detection Framework for Surgical Video UnderstandingabstractInstrument-tissue interaction detection task, which helps understand surgical activities, is vital for constructing computer-assisted surgery systems but with many challenges. Firstly, most models represent instrument-tissue interaction in a coarse-grained way which only focuses on classification and lacks the ability to automatically detect instruments and tissues. Secondly, existing works do not fully consider relations between intra- and inter-frame of instruments and tissues. In the paper, we propose to represent instrument-tissue interaction as 〈 instrument class, instrument bounding box, tissue class, tissue bounding box, action class 〉 quintuple and present an Instrument-Tissue Interaction Detection Network (ITIDNet) to detect the quintuple for surgery videos understanding. Specifically, we propose a Snippet Consecutive Feature (SCF) Layer to enhance features by modeling relationships of proposals in the current frame using global context information in the video snippet. We also propose a Spatial Corresponding Attention (SCA) Layer to incorporate features of proposals between adjacent frames through spatial encoding. To reason relationships between instruments and tissues, a Temporal Graph (TG) Layer is proposed with intra-frame connections to exploit relationships between instruments and tissues in the same frame and inter-frame connections to model the temporal information for the same instance. For evaluation, we build a cataract surgery video (PhacoQ) dataset and a cholecystectomy surgery video (CholecQ) dataset. Experimental results demonstrate the promising performance of our model, which outperforms other state-of-the-art models on both datasets. Huazhu Fu, Chin-Boon Chng, Ryo Kawasaki, Chee-Kong Chui, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Edge-Guided Contrastive Adaptation Network for Arteriovenous Nicking Classification Using Synthetic DataabstractRetinal arteriovenous nicking (AVN) manifests as a reduced venular caliber of an arteriovenous crossing. AVNs are signs of many systemic, particularly cardiovascular diseases. Studies have shown that people with AVN are twice as likely to have a stroke. However, AVN classification faces two challenges. One is the lack of data, especially AVNs compared to the normal arteriovenous (AV) crossings. The other is the significant intra-class variations and minute inter-class differences. AVNs may look different in shape, scale, pose, and color. On the other hand, the AVN could be different from the normal AV crossing only by slight thinning of the vein. To address these challenges, first, we develop a data synthesis method to generate AV crossings, including normal and AVNs. Second, to mitigate the domain shift between the synthetic and real data, an edge-guided unsupervised domain adaptation network is designed to guide the transfer of domain invariant information. Third, a semantic contrastive learning branch (SCLB) is introduced and a set of semantically related images, as a semantic triplet, are input to the network simultaneously to guide the network to focus on the subtle differences in venular width and to ignore the differences in appearance. These strategies effectively mitigate the lack of data, domain shift between synthetic and real data, and significant intra- but minute inter-class differences. Extensive experiments have been performed to demonstrate the outstanding performance of the proposed method. Huazhu Fu, Yu Ye 0003, Kun Chen 0005, Jianbo Mao, Ronald X. Xu, Mingzhai Sun |
IEEE Trans. Medical Imaging | 3 |
| 2024 | Geometric Correspondence-Based Multimodal Learning for Ophthalmic Image AnalysisabstractColor fundus photography (CFP) and Optical coherence tomography (OCT) images are two of the most widely used modalities in the clinical diagnosis and management of retinal diseases. Despite the widespread use of multimodal imaging in clinical practice, few methods for automated diagnosis of eye diseases utilize correlated and complementary information from multiple modalities effectively. This paper explores how to leverage the information from CFP and OCT images to improve the automated diagnosis of retinal diseases. We propose a novel multimodal learning method, named geometric correspondence-based multimodal learning network (GeCoM-Net), to achieve the fusion of CFP and OCT images. Specifically, inspired by clinical observations, we consider the geometric correspondence between the OCT slice and the CFP region to learn the correlated features of the two modalities for robust fusion. Furthermore, we design a new feature selection strategy to extract discriminative OCT representations by automatically selecting the important feature maps from OCT slices. Unlike the existing multimodal learning methods, GeCoM-Net is the first method that formulates the geometric relationships between the OCT slice and the corresponding region of the CFP image explicitly for CFP and OCT fusion. Experiments have been conducted on a large-scale private dataset and a publicly available dataset to evaluate the effectiveness of GeCoM-Net for diagnosing diabetic macular edema (DME), impaired visual acuity (VA) and glaucoma. The empirical results show that our method outperforms the current state-of-the-art multimodal learning methods by improving the AUROC score 0.4%, 1.9% and 2.9% for DME, VA and glaucoma detection, respectively. Yan Wang 0015, Liangli Zhen, Tien-En Tan, Huazhu Fu, Yangqin Feng, Zizhou Wang, Xinxing Xu, Rick Siow Mong Goh, Yipin Ng, Claire Calhoun, Gavin Siew Wei Tan, Jennifer K. Sun, Yong Liu 0026, Daniel S. W. Ting |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Fusion-Embedding Siamese Network for Light Field Salient Object DetectionabstractLight field salient object detection (SOD) has shown remarkable success and gained considerable attention from the computer vision community. Existing methods usually employ a single-/two-stream network to detect saliency. However, these methods can only handle up to two different modalities at a time, preventing them from being able to fully explore the rich information in multi-modal light field derived data. To address this, we propose the first joint multi-modal learning framework, called FES-Net, for light field SOD, which can take rich inputs not limited to two modalities. Specifically, we propose an attention-aware adaptation module to first transform the multi-modal inputs for use in our joint learning framework. The transformed inputs are then fed to a Siamese network along with multiple embedded feature fusion modules to extract informative multi-modal features. Finally, we predict saliency maps from the high-level extracted features using a saliency decoder module. Our joint multi-modal learning framework effectively resolves the limitations of existing methods, providing efficient and effective multi-modal learning that can fully explore the valuable information in light field data for accurate saliency detection. Furthermore, we improve the performance by introducing the Transformer as our backbone network. To the best of our knowledge, the improved version of our model, called FES-Trans, is the first attempt to address the challenging light field SOD with the powerful Transformer technique. Extensive experiments on benchmark datasets demonstrate that our models are superior light field SOD approaches and outperform cutting-edge models remarkably. Geng Chen 0001, Huazhu Fu, Tao Zhou 0002, Guobao Xiao, Keren Fu, Yong Xia 0001, Yanning Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | DR-FER: Discriminative and Robust Representation Learning for Facial Expression RecognitionabstractLearning discriminative and robust representations is important for facial expression recognition (FER) due to subtly different emotional faces and their subjective annotations. Previous works usually address one representation solely because these two goals seem to be contradictory for optimization. Their performances inevitably suffer from challenges from the other representation. In this article, by considering this problem from two novel perspectives, we demonstrate that discriminative and robust representations can be learned in a unified approach, i.e., DR-FER, and mutually benefit each other. Moreover, we make it with the supervision from only original annotations. Specifically, to learn discriminative representations, we propose performing masked image modeling (MIM) as an auxiliary task to force our network to discover expression-related facial areas. This is the first attempt to employ MIM to explore discriminative patterns in a self-supervised manner. To extract robust representations, we present a category-aware self-paced learning schedule to mine high-quality annotated (easy) expressions and incorrectly annotated (hard) counterparts. We further introduce a retrieval similarity-based relabeling strategy to correct hard expression annotations, exploiting them more effectively. By enhancing the discrimination ability of the FER classifier as a bridge, these two learning goals significantly strengthen each other. Extensive experiments on several popular benchmarks demonstrate the superior performance of our DR-FER. Moreover, thorough visualizations and extra experiments on manually annotation-corrupted datasets show that our approach successfully accomplishes learning both discriminative and robust representations simultaneously. Ming Li 0073, Huazhu Fu, Shengfeng He, Hehe Fan, Jun Liu 0036, Jussi Keppo, Zheng Shou 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | Self-Mining the Confident Prototypes for Source-Free Unsupervised Domain Adaptation in Image SegmentationabstractThis paper studies a practical Source-free unsupervised domain adaptation (SFUDA) problem, which transfers knowledge of source-trained models to the target domain, without accessing the source data. It has received increasing attention in recent years, while the prior arts focus on designing adaptation strategies, ignoring that different target samples exhibit different transfer abilities on the source model. Additionally, we observe pixel-wise class prediction is typically accompanied by ambiguity issue, i.e., prediction errors often occur between several confusing classes. In this study, we propose a dual-branch collaborative learning framework that aims to achieve reliable knowledge transfer from important samples to the rest by fully mining confident prototypes in the target data. Concretely, we first partition the target data into confident samples and uncertain samples via a new class-ranking reliability score and then utilize the latent features from the confident branch as guidance to promote the learning of the uncertain branch. For ambiguity issue, we propose a feature relabelling module, which exploits reliable prototypes in the mini-batch as well as in the target data to refine labels of uncertain features. We further deploy the proposed framework to commonly used CNN and state-of-the-art Transformer architectures and reveal the potential to promote the generalization ability of backbone models. Experimental results on both natural and medical benchmark datasets verify that our proposed approach exceeds state-of-the-art SFUDA methods with large margins, and achieves comparable performance to existing UDA methods. Yuntong Tian, Huazhu Fu, Lei Zhu 0003, Lequan Yu |
IEEE Trans. Multim. | 3 |
| 2024 | Exploring Separable Attention for Multi-Contrast MR Image Super-ResolutionabstractSuper-resolving the magnetic resonance (MR) image of a target contrast under the guidance of the corresponding auxiliary contrast, which provides additional anatomical information, is a new and effective solution for fast MR imaging. However, current multi-contrast super-resolution (SR) methods tend to concatenate different contrasts directly, ignoring their relationships in different clues, e.g., in the high-and low-intensity regions. In this study, we propose a separable attention network (comprising high-intensity priority (HP) attention and low-intensity separation (LS) attention), named SANet. Our SANet could explore the areas of high-and low-intensity regions in the "forward" and "reverse" directions with the help of the auxiliary contrast while learning clearer anatomical structure and edge information for the SR of a target-contrast MR image. SANet provides three appealing benefits: First, it is the first model to explore a separable attention mechanism that uses the auxiliary contrast to predict the high-and low-intensity regions, diverting more attention to refining any uncertain details between these regions and correcting the fine areas in the reconstructed results. Second, a multistage integration module is proposed to learn the response of multi-contrast fusion at multiple stages, get the dependency between the fused representations, and boost their representation ability. Third, extensive experiments with various state-of-the-art multi-contrast SR methods on fastMRI and clinical in vivo datasets demonstrate the superiority of our model. The code is released at https://github.com/chunmeifeng/SANet. Chun-Mei Feng 0001, Yunlu Yan, Kai Yu 0009, Yong Xu 0001, Huazhu Fu, Jian Yang 0003, Ling Shao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Autoencoder in Autoencoder NetworksabstractModeling complex correlations on multiview data is still challenging, especially for high-dimensional features with possible noise. To address this issue, we propose a novel unsupervised multiview representation learning (UMRL) algorithm, termed autoencoder in autoencoder networks (AE2-Nets). The proposed framework effectively encodes information from high-dimensional heterogeneous data into a compact and informative representation with the proposed bidirectional encoding strategy. Specifically, the proposed AE2-Nets conduct encoding in two directions: the inner-AE-networks extract view-specific intrinsic information (forward encoding), while the outer-AE-networks integrate this view-specific intrinsic information from different views into a latent representation (backward encoding). For the nested architecture, we further provide a probabilistic explanation and extension from hierarchical variational autoencoder. The forward-backward strategy flexibly addresses high-dimensional (noisy) features within each view and encodes complementarity across multiple views in a unified framework. Extensive results on benchmark datasets validate the advantages compared to the state-of-the-art algorithms. Changqing Zhang 0002, Zongbo Han, Yeqing Liu, Huazhu Fu, Qinghua Hu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2023 | Learning Federated Visual Prompt in Null Space for MRI ReconstructionabstractFederated Magnetic Resonance Imaging (MRI) reconstruction enables multiple hospitals to collaborate distributedly without aggregating local data, thereby protecting patient privacy. However, the data heterogeneity caused by different MRI protocols, insufficient local training data, and limited communication bandwidth inevitably impair global model convergence and updating. In this paper, we propose a new algorithm, FedPR, to learn federated visual prompts in the null space of global prompt for MRI reconstruction. FedPR is a new federated paradigm that adopts a powerful pre-trained model while only learning and communicating the prompts with few learnable parameters, thereby significantly reducing communication costs and achieving competitive performance on limited local data. Moreover, to deal with catastrophic forgetting caused by data heterogeneity, FedPR also updates efficient federated visual prompts that project the local prompts into an approximate null space of the global prompt, thereby suppressing the interference of gradients on the server performance. Extensive experiments on federated MRI show that FedPR significantly outperforms state-of-the-art FL algorithms with < 6% of communication costs when given the limited amount of local training data. Chun-Mei Feng 0001, Bangjun Li, Xinxing Xu, Yong Liu 0026, Huazhu Fu, Wangmeng Zuo |
CVPR | 5 |
| 2023 | Model Barrier: A Compact Un-Transferable Isolation Domain for Model Intellectual Property ProtectionabstractAs scientific and technological advancements result from human intellectual labor and computational costs, protecting model intellectual property (IP) has become increasingly important to encourage model creators and owners. Model IP protection involves preventing the use of well-trained models on unauthorized domains. To address this issue, we propose a novel approach called Compact Un-Transferable Isolation Domain (CUTI-domain), which acts as a barrier to block illegal transfers from authorized to unauthorized domains. Specifically, CUTI-domain blocks crossdomain transfers by highlighting the private style features of the authorized domain, leading to recognition failure on unauthorized domains with irrelevant private style features. Moreover, we provide two solutions for using CUTI-domain depending on whether the unauthorized domain is known or not: target-specified CUTI-domain and target-free CUTI-domain. Our comprehensive experimental results on four digit datasets, CIFAR10 & STL10, and VisDA-2017 dataset demonstrate that CUTI-domain can be easily implemented as a plug-and-play module with different backbones, providing an efficient solution for model IP protection. Lianyu Wang, Meng Wang 0001, Daoqiang Zhang, Huazhu Fu |
CVPR | 4 |
| 2023 | Video Adverse-Weather-Component Suppression Network via Weather Messenger and Adversarial BackpropagationabstractAlthough convolutional neural networks (CNNs) have been proposed to remove adverse weather conditions in single images using a single set of pre-trained weights, they fail to restore weather videos due to the absence of temporal information. Furthermore, existing methods for removing adverse weather conditions (e.g., rain, fog, and snow) from videos can only handle one type of adverse weather. In this work, we propose the first framework for restoring videos from all adverse weather conditions by developing a video adverse-weather-component suppression network (ViWS-Net). To achieve this, we first devise a weather-agnostic video transformer encoder with multiple transformer stages. Moreover, we design a long short-term temporal modeling mechanism for weather messenger to early fuse input adjacent video frames and learn weather-specific information. We further introduce a weather discriminator with gradient reversion, to maintain the weather-invariant common information and suppress the weather-specific information in pixel features, by adversarially predicting weather types. Finally, we develop a messenger-driven video transformer decoder to retrieve the residual weather-specific feature, which is spatiotemporally aggregated with hierarchical pixel features and refined to predict the clean target frame of input videos. Experimental results, on benchmark datasets and real-world weather videos, demonstrate that our ViWS-Net outperforms current state-of-the-art methods in terms of restoring videos degraded by any weather condition. Angelica I. Avilés-Rivero, Huazhu Fu, Weiming Wang 0002, Lei Zhu 0003 |
ICCV | 3 |
| 2023 | Calibrating Multimodal LearningabstractMultimodal machine learning has achieved remarkable progress in a wide range of scenarios. However, the reliability of multimodal learning remains largely unexplored. In this paper, through extensive empirical studies, we identify current multimodal classification methods suffer from unreliable predictive confidence that tend to rely on partial modalities when estimating confidence. Specifically, we find that the confidence estimated by current models could even increase when some modalities are corrupted. To address the issue, we introduce an intuitive principle for multimodal learning, i.e., the confidence should not increase when one modality is removed. Accordingly, we propose a novel regularization technique, i.e., Calibrating Multimodal Learning (CML) regularization, to calibrate the predictive confidence of previous methods. This technique could be flexibly equipped by existing models and improve the performance in terms of confidence calibration, classification accuracy, and model robustness. Huan Ma 0006, Changqing Zhang 0002, Bingzhe Wu, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu |
ICML | 5 |
| 2023 | dugMatting: Decomposed-Uncertainty-Guided MattingabstractCutting out an object and estimating its opacity mask, known as image matting, is a key task in image and video editing. Due to the highly ill-posed issue, additional inputs, typically user-defined trimaps or scribbles, are usually needed to reduce the uncertainty. Although effective, it is either time consuming or only suitable for experienced users who know where to place the strokes. In this work, we propose a decomposed-uncertainty-guided matting (dugMatting) algorithm, which explores the explicitly decomposed uncertainties to efficiently and effectively improve the results. Basing on the characteristic of these uncertainties, the epistemic uncertainty is reduced in the process of guiding interaction (which introduces prior knowledge), while the aleatoric uncertainty is reduced in modeling data distribution (which introduces statistics for both data and possible noise). The proposed matting framework relieves the requirement for users to determine the interaction areas by using simple and efficient labeling. Extensively quantitative and qualitative results validate that the proposed method significantly improves the original matting algorithms in terms of both efficiency and efficacy. Jiawei Wu 0001, Changqing Zhang 0002, Huazhu Fu, Xi Peng 0001, Joey Tianyi Zhou |
ICML | 4 |
| 2023 | Provable Dynamic Fusion for Low-Quality Multimodal DataabstractThe inherent challenge of multimodal fusion is to precisely capture the cross-modal correlation and flexibly conduct cross-modal interaction. To fully release the value of each modality and mitigate the influence of low-quality multimodal data, dynamic multimodal fusion emerges as a promising learning paradigm. Despite its widespread use, theoretical justifications in this field are still notably lacking. Can we design a provably robust multimodal fusion method? This paper provides theoretical understandings to answer this question under a most popular multimodal fusion framework from the generalization perspective. We proceed to reveal that several uncertainty estimation solutions are naturally available to achieve robust multimodal fusion. Then a novel multimodal fusion framework termed Quality-aware Multimodal Fusion (QMF) is proposed, which can improve the performance in terms of classification accuracy and model robustness. Extensive experimental results on multiple benchmarks can support our findings. Changqing Zhang 0002, Qinghua Hu, Huazhu Fu, Joey Tianyi Zhou, Xi Peng 0001 |
ICML | 5 |
| 2023 | RBGNet: Reliable Boundary-Guided Segmentation of Choroidal Neovascularization
Tao Chen 0003, Yitian Zhao, Lei Mou, Dan Zhang 0026, Xiayu Xu, Huazhu Fu, Jiong Zhang 0004 |
MICCAI (4) | 7 |
| 2023 | Medical Phrase Grounding with Region-Phrase Context Contrastive Alignment
Zhihao Chen 0004, Yang Zhou 0017, Junting Zhao, Gideon Ooi, Lionel Tim-Ee Cheng, Choon Hua Thng, Xinxing Xu, Yong Liu 0026, Huazhu Fu |
MICCAI (7) | 11 |
| 2023 | ACT-Net: Anchor-Context Action Detection in Surgery Videos
Luoying Hao, Heng Li 0010, Huazhu Fu, Jinming Duan 0001, Jiang Liu 0001 |
MICCAI (9) | 6 |
| 2023 | Content-Preserving Diffusion Model for Unsupervised AS-OCT Image Despeckling
Sanqian Li, Risa Higashita, Huazhu Fu, Heng Li 0010, Jingxuan Niu, Jiang Liu 0001 |
MICCAI (7) | 3 |
| 2023 | Frequency-Mixed Single-Source Domain Generalization for Medical Image Segmentation
Heng Li 0010, Haojin Li 0003, Huazhu Fu, Xiuyun Su, Jiang Liu 0001 |
MICCAI (6) | 4 |
| 2023 | Shifting More Attention to Breast Lesion Segmentation in Ultrasound Videos
Qian Dai, Lei Zhu 0003, Huazhu Fu, Qiong Wang 0001, Wenhao Rao, Liansheng Wang 0002 |
MICCAI (3) | 4 |
| 2023 | Polar-Net: A Clinical-Friendly Model for Alzheimer's Disease Detection in OCTA Images
Shouyue Liu, Jinkui Hao, Yanwu Xu 0001, Huazhu Fu, Jiang Liu 0001, Yalin Zheng, Yonghuai Liu, Jiong Zhang 0004, Yitian Zhao |
MICCAI (7) | 4 |
| 2023 | Category-Independent Visual Explanation for Medical Deep Network Understanding
Yiming Qian, Liangzhi Li 0004, Huazhu Fu, Meng Wang 0001, Qingsheng Peng, Ching Yu Cheng, Yong Liu 0026, Rick Siow Mong Goh, Xinxing Xu |
MICCAI (2) | 3 |
| 2023 | A Spatial-Temporal Deformable Attention Based Framework for Breast Lesion Detection in Videos
Jiale Cao, Huazhu Fu, Rao Muhammad Anwer, Fahad Shahbaz Khan |
MICCAI (2) | 3 |
| 2023 | Uncertainty-Informed Mutual Learning for Joint Medical Image Classification and Segmentation
Ke Zou, Xianjie Liu, Xuedong Yuan, Xiaojing Shen, Meng Wang 0001, Huazhu Fu |
MICCAI (4) | 8 |
| 2023 | Federated Uncertainty-Aware Aggregation for Fundus Diabetic Retinopathy Staging
Meng Wang 0001, Lianyu Wang, Xinxing Xu, Ke Zou, Yiming Qian, Rick Siow Mong Goh, Yong Liu 0026, Huazhu Fu |
MICCAI (2) | 8 |
| 2023 | Minimal-Supervised Medical Image Segmentation via Vector Quantization Memory
Yanyu Xu 0001, Menghan Zhou, Yangqin Feng, Xinxing Xu, Huazhu Fu, Rick Siow Mong Goh, Yong Liu 0026 |
MICCAI (3) | 5 |
| 2023 | DiffMIC: Dual-Guidance Diffusion Network for Medical Image Classification
Huazhu Fu, Angelica I. Avilés-Rivero, Carola-Bibiane Schönlieb, Lei Zhu 0003 |
MICCAI (6) | 2 |
| 2023 | Polar Eyeball Shape Net for 3D Posterior Ocular Shape Representation
Xiaojuan Qi 0001, Huazhu Fu, Jiang Liu 0001 |
MICCAI (6) | 6 |
| 2023 | Elongated Physiological Structure Segmentation via Spatial and Scale Uncertainty-Aware Network
Yinglin Zhang, Ruiling Xi, Huazhu Fu, Dave Towey, Ruibin Bai, Risa Higashita, Jiang Liu 0001 |
MICCAI (4) | 3 |
| 2023 | Reliable Multimodality Eye Disease Screening via Mixture of Student's t Distributions
Ke Zou, Tian Lin 0002, Xuedong Yuan, Haoyu Chen 0002, Xiaojing Shen, Meng Wang 0001, Huazhu Fu |
MICCAI (7) | 7 |
| 2023 | Fairness-guided Few-shot Prompting for Large Language ModelsabstractLarge language models have demonstrated surprising ability to perform in-context learning, i.e., these models can be directly applied to solve numerous downstream tasks by conditioning on a prompt constructed by a few input-output examples. However, prior research has shown that in-context learning can suffer from high instability due to variations in training examples, example order, and prompt formats. Therefore, the construction of an appropriate prompt is essential for improving the performance of in-context learning. In this paper, we revisit this problem from the view of predictive bias. Specifically, we introduce a metric to evaluate the predictive bias of a fixed prompt against labels or a given attributes. Then we empirically show that prompts with higher bias always lead to unsatisfactory predictive quality. Based on this observation, we propose a novel search strategy based on the greedy search to identify the near-optimal prompt for improving the performance of in-context learning. We perform comprehensive experiments with state-of-the-art mainstream models such as GPT-3 on various downstream tasks. Our results indicate that our method can enhance the model's in-context learning performance in an effective and interpretable manner. Huan Ma 0006, Changqing Zhang 0002, Yatao Bian, Lemao Liu, Zhirui Zhang, Peilin Zhao, Shu Zhang 0013, Huazhu Fu, Qinghua Hu, Bingzhe Wu |
NeurIPS | 8 |
| 2023 | Specificity-preserving RGB-D saliency detectionabstractRGB-D saliency detection has attracted increasing attention, due to its effectiveness and the fact that depth cues can now be conveniently captured. Existing works often focus on learning a shared representation through various fusion strategies, with few methods explicitly considering how to preserve modality-specific characteristics. In this paper, taking a new perspective, we propose a specificity-preserving network (SP-Net) for RGB-D saliency detection, which benefits saliency detection performance by exploring both the shared information and modality-specific properties (e.g., specificity). Specifically, two modality-specific networks and a shared learning network are adopted to generate individual and shared saliency maps. A cross-enhanced integration module (CIM) is proposed to fuse cross-modal features in the shared learning network, which are then propagated to the next layer for integrating cross-level information. Besides, we propose a multi-modal feature aggregation (MFA) module to integrate the modality-specific features from each individual decoder into the shared decoder, which can provide rich complementary multi-modal information to boost the saliency detection performance. Further, a skip connection is used to combine hierarchical features between the encoder and decoder layers. Experiments on six benchmark datasets demonstrate that our SP-Net outperforms other state-of-the-art methods. Code is available at: https://github.com/taozh2017/SPNet. Tao Zhou 0002, Deng-Ping Fan, Geng Chen 0001, Yi Zhou 0007, Huazhu Fu |
Comput. Vis. Media | 5 |
| 2023 | Contrastive domain adaptation with consistency match for automated pneumonia diagnosis
Yangqin Feng, Zizhou Wang, Xinxing Xu, Yan Wang 0015, Huazhu Fu, Shaohua Li 0003, Liangli Zhen, Xiaofeng Lei, Yingnan Cui, Jordan Zheng Ting Sim, Yonghan Ting, Joey Tianyi Zhou, Yong Liu 0026, Rick Siow Mong Goh, Cher Heng Tan |
Medical Image Anal. | 5 |
| 2023 | A generic fundus image enhancement network boosted by frequency self-supervised representation learning
Heng Li 0010, Haofeng Liu, Huazhu Fu, Yanwu Xu 0001, Hai Shu, Ke Niu 0002, Jiang Liu 0001 |
Medical Image Anal. | 3 |
| 2023 | Transformers in medical imaging: A survey
Fahad Shamshad, Salman Khan 0001, Syed Waqas Zamir, Muhammad Haris Khan, Munawar Hayat, Fahad Shahbaz Khan, Huazhu Fu |
Medical Image Anal. | 7 |
| 2023 | GAMMA challenge: Glaucoma grAding from Multi-Modality imAges
Huihui Fang, Fei Li 0021, Huazhu Fu, Fengbin Lin, Jiongcheng Li, Yue Huang 0001, Qinji Yu, Sifan Song, Xinxing Xu, Yanyu Xu 0001, Wensai Wang, Shuai Lu 0003, Huiqi Li, Shihua Huang, Zhichao Lu, Chubin Ou, Xifei Wei, Bingyuan Liu, Riadh Kobbi, Xiaoying Tang 0001, Li Lin 0006, Hrvoje Bogunovic, José Ignacio Orlando, Xiulan Zhang, Yanwu Xu 0001 |
Medical Image Anal. | 4 |
| 2023 | Consistency and Diversity Induced Human Motion SegmentationabstractSubspace clustering is a classical technique that has been widely used for human motion segmentation and other related tasks. However, existing segmentation methods often cluster data without guidance from prior knowledge, resulting in unsatisfactory segmentation results. To this end, we propose a novel Consistency and Diversity induced human Motion Segmentation (CDMS) algorithm. Specifically, our model factorizes the source and target data into distinct multi-layer feature spaces, in which transfer subspace learning is conducted on different layers to capture multi-level information. A multi-mutual consistency learning strategy is carried out to reduce the domain gap between the source and target data. In this way, the domain-specific knowledge and domain-invariant properties can be explored simultaneously. Besides, a novel constraint based on the Hilbert Schmidt Independence Criterion (HSIC) is introduced to ensure the diversity of multi-level subspace representations, which enables the complementarity of multi-level representations to be explored to boost the transfer learning performance. Moreover, to preserve the temporal correlations, an enhanced graph regularizer is imposed on the learned representation coefficients and the multi-level representations of the source data. The proposed model can be efficiently solved using the Alternating Direction Method of Multipliers (ADMM) algorithm. Extensive experimental results on public human motion datasets demonstrate the effectiveness of our method against several state-of-the-art approaches. Tao Zhou 0002, Huazhu Fu, Chen Gong 0002, Ling Shao 0001, Fatih Porikli, Haibin Ling, Jianbing Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Trusted Multi-View Classification With Dynamic Evidential FusionabstractExisting multi-view classification algorithms focus on promoting accuracy by exploiting different views, typically integrating them into common representations for follow-up tasks. Although effective, it is also crucial to ensure the reliability of both the multi-view integration and the final decision, especially for noisy, corrupted and out-of-distribution data. Dynamically assessing the trustworthiness of each view for different samples could provide reliable integration. This can be achieved through uncertainty estimation. With this in mind, we propose a novel multi-view classification algorithm, termed trusted multi-view classification (TMC), providing a new paradigm for multi-view learning by dynamically integrating different views at an evidence level. The proposed TMC can promote classification reliability by considering evidence from each view. Specifically, we introduce the variational Dirichlet to characterize the distribution of the class probabilities, parameterized with evidence from different views and integrated with the Dempster-Shafer theory. The unified learning framework induces accurate uncertainty and accordingly endows the model with both reliability and robustness against possible noise or corruption. Both theoretical and experimental results validate the effectiveness of the proposed model in accuracy, robustness and trustworthiness. Zongbo Han, Changqing Zhang 0002, Huazhu Fu, Joey Tianyi Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | GCoNet+: A Stronger Group Collaborative Co-Salient Object DetectorabstractIn this paper, we present a novel end-to-end group collaborative learning network, termed GCoNet+, which can effectively and efficiently (250 fps) identify co-salient objects in natural scenes. The proposed GCoNet+ achieves the new state-of-the-art performance for co-salient object detection (CoSOD) through mining consensus representations based on the following two essential criteria: 1) intra-group compactness to better formulate the consistency among co-salient objects by capturing their inherent shared attributes using our novel group affinity module (GAM); 2) inter-group separability to effectively suppress the influence of noisy objects on the output by introducing our new group collaborating module (GCM) conditioning on the inconsistent consensus. To further improve the accuracy, we design a series of simple yet effective components as follows: i) a recurrent auxiliary classification module (RACM) promoting model learning at the semantic level; ii) a confidence enhancement module (CEM) assisting the model in improving the quality of the final predictions; and iii) a group-based symmetric triplet (GST) loss guiding the model to learn more discriminative features. Extensive experiments on three challenging benchmarks, i.e., CoCA, CoSOD3k, and CoSal2015, demonstrate that our GCoNet+ outperforms the existing 12 cutting-edge models. Code has been released at https://github.com/ZhengPeng7/GCoNet_plus. Peng Zheng 0004, Huazhu Fu, Deng-Ping Fan, Jie Qin 0004, Yu-Wing Tai, Chi-Keung Tang, Luc Van Gool |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Knowledge driven weights estimation for large-scale few-shot image recognition
Jingjing Chen 0001, Linhai Zhuo, Zhipeng Wei 0001, Hao Zhang 0047, Huazhu Fu, Yu-Gang Jiang 0001 |
Pattern Recognit. | 5 |
| 2023 | Cross-level Feature Aggregation Network for Polyp Segmentation
Tao Zhou 0002, Yi Zhou 0007, Kelei He, Chen Gong 0002, Jian Yang 0003, Huazhu Fu, Dinggang Shen |
Pattern Recognit. | 6 |
| 2023 | RGB-D Human Matting: A Real-World Benchmark Dataset and a Baseline MethodabstractThe last decade has witnessed an increasing exploration and development of human matting. However, existing matting works primarily focus on predicting better alpha mattes from RGB images. So far few efforts have been devoted to tackling human matting in real-world activity scenarios with RGB-D information. To this end, this paper concentrates on the RGB-D human matting task, and provides the first public RGB-D human matting benchmark dataset as well as a baseline method for deep learning-based RGB-D human matting. To support the research on RGB-D human matting, a new RGB-D human-matting dataset (HDM-2K) is collected and released, which contains 2,270 high-resolution human images in various real-world scenarios and the corresponding depth maps. Additionally, a baseline method for RGB-D human matting is further proposed, which automatically generates the alpha matte by jointly exploiting the spatial structure information in the depth map and detailed texture information in the RGB image. Finally, extensive experiments conducted on the HDM-2K dataset demonstrate that the depth maps are effective for the matting task and the proposed baseline method achieves promising performance on human matting. Bo Peng 0007, Jianjun Lei 0001, Huazhu Fu, Haifeng Shen, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Global-and-Local Collaborative Learning for Co-Salient Object DetectionabstractThe goal of co-salient object detection (CoSOD) is to discover salient objects that commonly appear in a query group containing two or more relevant images. Therefore, how to effectively extract interimage correspondence is crucial for the CoSOD task. In this article, we propose a global-and-local collaborative learning (GLNet) architecture, which includes a global correspondence modeling (GCM) and a local correspondence modeling (LCM) to capture the comprehensive interimage corresponding relationship among different images from the global and local perspectives. First, we treat different images as different time slices and use 3-D convolution to integrate all intrafeatures intuitively, which can more fully extract the global group semantics. Second, we design a pairwise correlation transformation (PCT) to explore similarity correspondence between pairwise images and combine the multiple local pairwise correspondences to generate the local interimage relationship. Third, the interimage relationships of the GCM and LCM are integrated through a global-and-local correspondence aggregation (GLA) module to explore more comprehensive interimage collaboration cues. Finally, the intra and inter features are adaptively integrated by an intra-and-inter weighting fusion (AEWF) module to learn co-saliency features and predict the co-saliency map. The proposed GLNet is evaluated on three prevailing CoSOD benchmark datasets, demonstrating that our model trained on a small dataset (about 3k images) still outperforms 11 state-of-the-art competitors trained on some large datasets (about 8k-200k images). Runmin Cong, Ning Yang 0008, Chongyi Li, Huazhu Fu, Yao Zhao 0001, Qingming Huang, Sam Kwong |
IEEE Trans. Cybern. | 4 |
| 2023 | Dual Multiscale Mean Teacher Network for Semi-Supervised Infection Segmentation in Chest CT Volume for COVID-19abstractAutomated detecting lung infections from computed tomography (CT) data plays an important role for combating coronavirus 2019 (COVID-19). However, there are still some challenges for developing AI system: 1) most current COVID-19 infection segmentation methods mainly relied on 2-D CT images, which lack 3-D sequential constraint; 2) existing 3-D CT segmentation methods focus on single-scale representations, which do not achieve the multiple level receptive field sizes on 3-D volume; and 3) the emergent breaking out of COVID-19 makes it hard to annotate sufficient CT volumes for training deep model. To address these issues, we first build a multiple dimensional-attention convolutional neural network (MDA-CNN) to aggregate multiscale information along different dimension of input feature maps and impose supervision on multiple predictions from different convolutional neural networks (CNNs) layers. Second, we assign this MDA-CNN as a basic network into a novel dual multiscale mean teacher network (DM [Formula: see text]-Net) for semi-supervised COVID-19 lung infection segmentation on CT volumes by leveraging unlabeled data and exploring the multiscale information. Our DM [Formula: see text]-Net encourages multiple predictions at different CNN layers from the student and teacher networks to be consistent for computing a multiscale consistency loss on unlabeled data, which is then added to the supervised loss on the labeled data from multiple predictions of MDA-CNN. Third, we collect two COVID-19 segmentation datasets to evaluate our method. The experimental results show that our network consistently outperforms the compared state-of-the-art methods. Liansheng Wang 0002, Jiacheng Wang 0002, Lei Zhu 0003, Huazhu Fu, Ping Li 0016, Gary Cheng 0001, Shuo Li 0001, Pheng-Ann Heng |
IEEE Trans. Cybern. | 4 |
| 2023 | SGU-Net: Shape-Guided Ultralight Network for Abdominal Image SegmentationabstractConvolutional neural networks (CNNs) have achieved significant success in medical image segmentation. However, they also suffer from the requirement of a large number of parameters, leading to a difficulty of deploying CNNs to low-source hardwares, e.g., embedded systems and mobile devices. Although some compacted or small memory-hungry models have been reported, most of them may cause degradation in segmentation accuracy. To address this issue, we propose a shape-guided ultralight network (SGU-Net) with extremely low computational costs. The proposed SGU-Net includes two main contributions: it first presents an ultralight convolution that is able to implement double separable convolutions simultaneously, i.e., asymmetric convolution and depthwise separable convolution. The proposed ultralight convolution not only effectively reduces the number of parameters but also enhances the robustness of SGU-Net. Secondly, our SGU-Net employs an additional adversarial shape-constraint to let the network learn shape representation of targets, which can significantly improve the segmentation accuracy for abdomen medical images using self-supervision. The SGU-Net is extensively tested on four public benchmark datasets, LiTS, CHAOS, NIH-TCIA and 3Dircbdb. Experimental results show that SGU-Net achieves higher segmentation accuracy using lower memory costs, and outperforms state-of-the-art networks. Moreover, we apply our ultralight convolution into a 3D volume segmentation network, which obtains a comparable performance with fewer parameters and memory usage. Tao Lei 0003, Xiaogang Du, Huazhu Fu, Changqing Zhang 0002, Asoke K. Nandi |
IEEE J. Biomed. Health Informatics | 4 |
| 2023 | Flexible Fusion Network for Multi-Modal Brain Tumor SegmentationabstractAutomated brain tumor segmentation is crucial for aiding brain disease diagnosis and evaluating disease progress. Currently, magnetic resonance imaging (MRI) is a routinely adopted approach in the field of brain tumor segmentation that can provide different modality images. It is critical to leverage multi-modal images to boost brain tumor segmentation performance. Existing works commonly concentrate on generating a shared representation by fusing multi-modal data, while few methods take into account modality-specific characteristics. Besides, how to efficiently fuse arbitrary numbers of modalities is still a difficult task. In this study, we present a flexible fusion network (termed F$^{2}$Net) for multi-modal brain tumor segmentation, which can flexibly fuse arbitrary numbers of multi-modal information to explore complementary information while maintaining the specific characteristics of each modality. Our F$^{2}$Net is based on the encoder-decoder structure, which utilizes two Transformer-based feature learning streams and a cross-modal shared learning network to extract individual and shared feature representations. To effectively integrate the knowledge from the multi-modality data, we propose a cross-modal feature-enhanced module (CFM) and a multi-modal collaboration module (MCM), which aims at fusing the multi-modal features into the shared learning network and incorporating the features from encoders into the shared decoder, respectively. Extensive experimental results on multiple benchmark datasets demonstrate the effectiveness of our F$^{2}$Net over other state-of-the-art segmentation methods. Hengyi Yang, Tao Zhou 0002, Yi Zhou 0007, Yizhe Zhang 0001, Huazhu Fu |
IEEE J. Biomed. Health Informatics | 5 |
| 2023 | Uncertainty-Aware Multi-Dimensional Mutual Learning for Brain and Brain Tumor SegmentationabstractExisting segmentation methods for brain MRI data usually leverage 3D CNNs on 3D volumes or employ 2D CNNs on 2D image slices. We discovered that while volume-based approaches well respect spatial relationships across slices, slice-based methods typically excel at capturing fine local features. Furthermore, there is a wealth of complementary information between their segmentation predictions. Inspired by this observation, we develop an Uncertainty-aware Multi-dimensional Mutual learning framework to learn different dimensional networks simultaneously, each of which provides useful soft labels as supervision to the others, thus effectively improving the generalization ability. Specifically, our framework builds upon a 2D-CNN, a 2.5D-CNN, and a 3D-CNN, while an uncertainty gating mechanism is leveraged to facilitate the selection of qualified soft labels, so as to ensure the reliability of shared information. The proposed method is a general framework and can be applied to varying backbones. The experimental results on three datasets demonstrate that our method can significantly enhance the performance of the backbone network by notable margins, achieving a Dice metric improvement of 2.8% on MeniSeg, 1.4% on IBSR, and 1.3% on BraTS2020. Junting Zhao, Zhaohu Xing, Zhihao Chen 0004, Tong Han, Huazhu Fu, Lei Zhu 0003 |
IEEE J. Biomed. Health Informatics | 6 |
| 2023 | Specificity-Preserving Federated Learning for MR Image ReconstructionabstractFederated learning (FL) can be used to improve data privacy and efficiency in magnetic resonance (MR) image reconstruction by enabling multiple institutions to collaborate without needing to aggregate local data. However, the domain shift caused by different MR imaging protocols can substantially degrade the performance of FL models. Recent FL techniques tend to solve this by enhancing the generalization of the global model, but they ignore the domain-specific features, which may contain important information about the device properties and be useful for local reconstruction. In this paper, we propose a specificity-preserving FL algorithm for MR image reconstruction (FedMRI). The core idea is to divide the MR reconstruction model into two parts: a globally shared encoder to obtain a generalized representation at the global level, and a client-specific decoder to preserve the domain-specific properties of each client, which is important for collaborative reconstruction when the clients have unique distribution. Such scheme is then executed in the frequency space and the image space respectively, allowing exploration of generalized representation and client-specific properties simultaneously in different spaces. Moreover, to further boost the convergence of the globally shared encoder when a domain shift is present, a weighted contrastive regularization is introduced to directly correct any deviation between the client and server during optimization. Extensive experiments demonstrate that our FedMRI's reconstructed results are the closest to the ground-truth for multi-institutional data, and that it outperforms state-of-the-art FL methods. Chun-Mei Feng 0001, Yunlu Yan, Shanshan Wang 0002, Yong Xu 0001, Ling Shao 0001, Huazhu Fu |
IEEE Trans. Medical Imaging | 6 |
| 2023 | Multimodal Transformer for Accelerated MR ImagingabstractAccelerated multi-modal magnetic resonance (MR) imaging is a new and effective solution for fast MR imaging, providing superior performance in restoring the target modality from its undersampled counterpart with guidance from an auxiliary modality. However, existing works simply combine the auxiliary modality as prior information, lacking in-depth investigations on the potential mechanisms for fusing different modalities. Further, they usually rely on the convolutional neural networks (CNNs), which is limited by the intrinsic locality in capturing the long-distance dependency. To this end, we propose a multi-modal transformer (MTrans), which is capable of transferring multi-scale features from the target modality to the auxiliary modality, for accelerated MR imaging. To capture deep multi-modal information, our MTrans utilizes an improved multi-head attention mechanism, named cross attention module, which absorbs features from the auxiliary modality that contribute to the target modality. Our framework provides three appealing benefits: (i) Our MTrans use an improved transformers for multi-modal MR imaging, affording more global information compared with existing CNN-based methods. (ii) A new cross attention module is proposed to exploit the useful information in each modality at different scales. The small patch in the target modality aims to keep more fine details, the large patch in the auxiliary modality aims to obtain high-level context features from the larger region and supplement the target modality effectively. (iii) We evaluate MTrans with various accelerated multi-modal MR imaging tasks, e.g., MR image reconstruction and super-resolution, where MTrans outperforms state-of-the-art methods on fastMRI and real-world clinical datasets. Chun-Mei Feng 0001, Yunlu Yan, Geng Chen 0001, Yong Xu 0001, Ling Shao 0001, Huazhu Fu |
IEEE Trans. Medical Imaging | 7 |
| 2023 | Guest Editorial Special Issue on Geometric Deep Learning in Medical ImagingabstractIn recent years, more and more attention has been devoted to geometric deep learning (GDL) and its applications to various problems in medical imaging. Unlike convolutional neural networks (CNNs) limited to 2-D/3-D grid-structured data, GDL can handle non-Euclidean data (i.e., graphs and manifolds) and is hence well-suited for medical imaging data such as structure-function connectivity networks, imaging genetics and omics, spatio-temporal anatomical representations, physics-informed GDL for optimal imaging sampling and acquisition, GDL in imaging inverse problems, etc. However, despite recent advances in GDL research, questions remain on how best to learn representations of non-Euclidean medical imaging data; how to convolve effectively on graphs; how to perform graph pooling/unpooling; how to handle heterogeneous data; and how to improve the interpretability of GDL. After discussing many other domain experts, we identify the need for a special issue that brings to the attention of the medical imaging community these interesting topics. Huazhu Fu, Yitian Zhao, Pew-Thian Yap, Carola-Bibiane Schönlieb, Alejandro F. Frangi |
IEEE Trans. Medical Imaging | 1 |
| 2023 | Bridging Synthetic and Real Images: A Transferable and Multiple Consistency Aided Fundus Image Enhancement FrameworkabstractDeep learning based image enhancement models have largely improved the readability of fundus images in order to decrease the uncertainty of clinical observations and the risk of misdiagnosis. However, due to the difficulty of acquiring paired real fundus images at different qualities, most existing methods have to adopt synthetic image pairs as training data. The domain shift between the synthetic and the real images inevitably hinders the generalization of such models on clinical data. In this work, we propose an end-to-end optimized teacher-student framework to simultaneously conduct image enhancement and domain adaptation. The student network uses synthetic pairs for supervised enhancement, and regularizes the enhancement model to reduce domain-shift by enforcing teacher-student prediction consistency on the real fundus images without relying on enhanced ground-truth. Moreover, we also propose a novel multi-stage multi-attention guided enhancement network (MAGE-Net) as the backbones of our teacher and student network. Our MAGE-Net utilizes multi-stage enhancement module and retinal structure preservation module to progressively integrate the multi-scale features and simultaneously preserve the retinal structures for better fundus image quality enhancement. Comprehensive experiments on both real and synthetic datasets demonstrate that our framework outperforms the baseline approaches. Moreover, our method also benefits the downstream clinical tasks. Erjian Guo, Huazhu Fu, Luping Zhou, Dong Xu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2023 | H2Former: An Efficient Hierarchical Hybrid Transformer for Medical Image SegmentationabstractAccurate medical image segmentation is of great significance for computer aided diagnosis. Although methods based on convolutional neural networks (CNNs) have achieved good results, it is weak to model the long-range dependencies, which is very important for segmentation task to build global context dependencies. The Transformers can establish long-range dependencies among pixels by self-attention, providing a supplement to the local convolution. In addition, multi-scale feature fusion and feature selection are crucial for medical image segmentation tasks, which is ignored by Transformers. However, it is challenging to directly apply self-attention to CNNs due to the quadratic computational complexity for high-resolution feature maps. Therefore, to integrate the merits of CNNs, multi-scale channel attention and Transformers, we propose an efficient hierarchical hybrid vision Transformer (H2Former) for medical image segmentation. With these merits, the model can be data-efficient for limited medical data regime. The experimental results show that our approach exceeds previous Transformer, CNNs and hybrid methods on three 2D and two 3D medical image segmentation tasks. Moreover, it keeps computational efficiency in model parameters, FLOPs and inference time. For example, H2Former outperforms TransUNet by 2.29% in IoU score on KVASIR-SEG dataset with 30.77% parameters and 59.23% FLOPs. Along He, Kai Wang 0001, Tao Li 0022, Chengkun Du, Huazhu Fu |
IEEE Trans. Medical Imaging | 6 |
| 2023 | MG-Trans: Multi-Scale Graph Transformer With Information Bottleneck for Whole Slide Image ClassificationabstractMultiple instance learning (MIL)-based methods have become the mainstream for processing the megapixel-sized whole slide image (WSI) with pyramid structure in the field of digital pathology. The current MIL-based methods usually crop a large number of patches from WSI at the highest magnification, resulting in a lot of redundancy in the input and feature space. Moreover, the spatial relations between patches can not be sufficiently modeled, which may weaken the model's discriminative ability on fine-grained features. To solve the above limitations, we propose a Multi-scale Graph Transformer (MG-Trans) with information bottleneck for whole slide image classification. MG-Trans is composed of three modules: patch anchoring module (PAM), dynamic structure information learning module (SILM), and multi-scale information bottleneck module (MIBM). Specifically, PAM utilizes the class attention map generated from the multi-head self-attention of vision Transformer to identify and sample the informative patches. SILM explicitly introduces the local tissue structure information into the Transformer block to sufficiently model the spatial relations between patches. MIBM effectively fuses the multi-scale patch features by utilizing the principle of information bottleneck to generate a robust and compact bag-level representation. Besides, we also propose a semantic consistency loss to stabilize the training of the whole model. Extensive studies on three subtyping datasets and seven gene mutation detection datasets demonstrate the superiority of MG-Trans. Jiangbo Shi, Lufei Tang, Zeyu Gao 0001, Yang Li 0139, Chunbao Wang 0002, Tieliang Gong, Chen Li 0011, Huazhu Fu |
IEEE Trans. Medical Imaging | 8 |
| 2023 | Dynamic Mixup for Multi-Label Long-Tailed Food Ingredient RecognitionabstractRecognizing the ingredients composition for given food images facilitates the estimation of nutrition facts, which is crucial to various health relevant applications. Nevertheless, ingredient recognition is a multi-label long-tailed classification problem, where each image may contain multiple labels and the class distributions are highly imbalanced. Most existing approaches leverage off-the-shelf Convolutional Neural Networks (CNN) for multi-label ingredient recognition, overlooking the long-tailed issue, which results in low accuracy for tail ingredient categories. To address this problem, this paper proposes a dynamic Mixup (D-Mixup) approach, aiming to dynamically augment minority ingredients, in order to boost the recognition performance for tail ingredient categories. Specifically, our D-Mixup approach dynamically selects two training images based on the predictions of the previous training epoch, and generates a new synthetic image to train the recognition network. In this way, the training samples of tailed classes can be dynamically enlarged and better discriminative representations can be learnt for rare classes. Extensive experiments on both VIREO Food-172 dataset and UEC Food-100 dataset demonstrate the effectiveness of the proposed D-Mixup method. Jixiang Gao, Jingjing Chen 0001, Huazhu Fu, Yu-Gang Jiang 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Confidence-Aware Fusion Using Dempster-Shafer Theory for Multispectral Pedestrian DetectionabstractMultispectral pedestrian detection is an important and valuable task in many applications, which could provide a more accurate and reliable pedestrian detection result by using the complementary visual information from color and thermal images. However, it faces two open and difficult challenges: 1) how to effectively and dynamically integrate multispectral information according to the confidence of different modalities, and 2) how to produce a reliable prediction result. In this paper, we propose a novel confidence-aware multispectral pedestrian detection (CMPD) method, which flexibly learns the multispectral representation while simultaneously producing a reliable result with confidence estimation. Specifically, a dense fusion strategy is first proposed to extract the multilevel multispectral representation at the feature level. Then, an additional confidence subnetwork is utilized to dynamically estimate the detection confidence for each modality. Finally, Dempster's combination rule is introduced to fuse the results of different branches according to the rectified confidence. Our proposed CMPD method not only effectively integrates multimodal information but also provides a reliable prediction. Extensive experimental results demonstrate the efficiency of our algorithm compared with state-of-the-art methods. Qing Li 0018, Changqing Zhang 0002, Qinghua Hu, Huazhu Fu, Pengfei Zhu 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | From Ensemble Clustering to Subspace Clustering: Cluster Structure EncodingabstractIn this study, we propose a novel algorithm to encode the cluster structure by incorporating ensemble clustering (EC) into subspace clustering (SC). First, the low-rank representation (LRR) is learned from a higher order data relationship induced by ensemble K-means coding, which exploits the cluster structure in a co-association matrix of basic partitions (i.e., clustering results). Second, to provide a fast predictive coding mechanism, an encoding function parameterized by neural networks is introduced to predict the LRR derived from partitions. These two steps are jointly proceeded to seamlessly integrate partition information and original features and thus deliver better representations than the ones obtained from each single source. Moreover, an alternating optimization framework is developed to learn the LRR, train the encoding function, and fine-tune the higher order relationship. Extensive experiments on eight benchmark datasets validate the effectiveness of the proposed algorithm on several clustering tasks compared with state-of-the-art EC and SC methods. Zhiqiang Tao, Jun Li 0027, Huazhu Fu, Yu Kong 0001, Yun Fu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Can You Spot the Chameleon? Adversarially Camouflaging Images from Co-Salient Object DetectionabstractCo-salient object detection (CoSOD) has recently achieved significant progress and played a key role in retrieval-related tasks. However, it inevitably poses an entirely new safety and security issue, i.e., highly personal and sensitive content can potentially be extracting by powerful CoSOD methods. In this paper, we address this problem from the perspective of adversarial attacks and identify a novel task: adversarial co-saliency attack. Specially, given an image selected from a group of images containing some common and salient objects, we aim to generate an adversarial version that can mislead CoSOD methods to predict incorrect co-salient regions. Note that, compared with general white-box adversarial attacks for classification, this new task faces two additional challenges: (1) low success rate due to the diverse appearance of images in the group; (2) low transferability across CoSOD methods due to the considerable difference between CoSOD pipelines. To address these challenges, we propose the very first blackbox joint adversarial exposure and noise attack (Jadena), where we jointly and locally tune the exposure and additive perturbations of the image according to a newly designed high-feature-level contrast-sensitive loss function. Our method, without any information on the state-of-the-art CoSOD methods, leads to significant performance degradation on various co-saliency detection datasets and makes the co-salient objects undetectable. This can have strong practical benefits in properly securing the large number of personal photos currently shared on the Internet. Moreover, our method is potential to be utilized as a metric for evaluating the robustness of CoSOD methods. Ruijun Gao, Qing Guo 0005, Felix Juefei-Xu, Hongkai Yu, Huazhu Fu, Wei Feng 0005, Yang Liu 0003, Song Wang 0002 |
CVPR | 5 |
| 2022 | FedDC: Federated Learning with Non-IID Data via Local Drift Decoupling and CorrectionabstractFederated learning (FL) allows multiple clients to collectively train a high-performance global model without sharing their private data. However, the key challenge in federated learning is that the clients have significant statistical heterogeneity among their local data distributions, which would cause inconsistent optimized local models on the clientside. To address this fundamental dilemma, we propose a novel federated learning algorithm with local drift decoupling and correction (FedDC). Our FedDC only introduces lightweight modifications in the local training phase, in which each client utilizes an auxiliary local drift variable to track the gap between the local model parameter and the global model parameters. The key idea of FedDC is to utilize this learned local drift variable to bridge the gap, i.e., conducting consistency in parameter-level. The experiment results and analysis demonstrate that FedDC yields expediting convergence and better performance on various image classification tasks, robust in partial participation settings, non-iid data, and heterogeneous clients. Liang Gao 0001, Huazhu Fu, Li Li 0064, Yingwen Chen 0001, Ming Xu 0002, Cheng-Zhong Xu 0001 |
CVPR | 2 |
| 2022 | Trustworthy Long-Tailed ClassificationabstractClassification on long-tailed distributed data is a challenging problem, which suffers from serious class-imbalance and accordingly unpromising performance es-pecially on tail classes. Recently, the ensembling based methods achieve the state-of-the-art performance and show great potential. However, there are two limitations for cur-rent methods. First, their predictions are not trustworthy for failure-sensitive applications. This is especially harmful for the tail classes where the wrong predictions is basically fre-quent. Second, they assign unified numbers of experts to all samples, which is redundant for easy samples with excessive computational cost. To address these issues, we propose a Trustworthy Long-tailed Classification (TLC) method to jointly conduct classification and uncertainty estimation to identify hard samples in a multi-expert framework. Our TLC obtains the evidence-based uncertainty (EvU) and ev-idence for each expert, and then combines these uncer-tainties and evidences under the Dempster-Shafer Evidence Theory (DST). Moreover, we propose a dynamic expert en-gagement to reduce the number of engaged experts for easy samples and achieve efficiency while maintaining promising performances. Finally, we conduct comprehensive ex-periments on the tasks of classification, tail detection, OOD detection and failure prediction. The experimental results show that the proposed TLC outperforms existing methods and is trustworthy with reliable uncertainty. Bolian Li, Zongbo Han, Haining Li, Huazhu Fu, Changqing Zhang 0002 |
CVPR | 4 |
| 2022 | RSCFed: Random Sampling Consensus Federated Semi-supervised LearningabstractFederated semi-supervised learning (FSSL) aims to derive a global model by training fully-labeled and fully-unlabeled clients or training partially labeled clients. The existing approaches work well when local clients have in-dependent and identically distributed (IID) data but fail to generalize to a more practical FSSL setting, i.e., Non-IID setting. In this paper, we present a Random Sampling Consensus Federated learning, namely RSCFed, by con-sidering the uneven reliability among models from fully-labeled clients, fully-unlabeled clients or partially labeled clients. Our key motivation is that given models with large deviations from either labeled clients or unlabeled clients, the consensus could be reached by performing random sub-sampling over clients. To achieve it, instead of di-rectly aggregating local models, we first distill several sub-consensus models by random sub-sampling over clients and then aggregating the sub-consensus models to the global model. To enhance the robustness of sub-consensus models, we also develop a novel distance-reweighted model aggre-gation method. Experimental results show that our method outperforms state-of-the-art methods on three benchmarked datasets, including both natural and medical images. The code is available at https://github.com/XMed-Lab/RSCFed. Xiaoxiao Liang, Yiqun Lin, Huazhu Fu, Lei Zhu 0003, Xiaomeng Li 0001 |
CVPR | 3 |
| 2022 | Rethinking Video Rain Streak Removal: A New Synthesis Model and a Deraining Network with Video Rain Prior
Lei Zhu 0003, Huazhu Fu, Harry Qin, Carola-Bibiane Schönlieb, Wei Feng 0005, Song Wang 0002 |
ECCV (19) | 3 |
| 2022 | Data-Free Network Debiasing for Long-Tailed Visual RecognitionabstractReal-world data is often unbalanced and exhibits long-tailed distribution over classes. Vanilla classification models trained on imbalanced datasets inherently exhibit bias towards dominant classes. Existing debiasing methods mostly balance the data or the loss during training. Nevertheless, these data-acquiring methods are not suitable for situations where training data are unavailable. In this paper, we appeal to solutions without access to training data and propose a datafree debiasing (Free-D) method that serves as a plug-and-play module for any standard classification model. Specifically, our method adjusts both the feature representation via feature representation shifting and the classifier weight via class prior compensation in a data-free manner. We evaluate and compare our methods on four long-tailed visual recognition datasets, i.e., long-tailed CIFAR-10/-100, ImageNet-LT, and Places-LT. Extensive experiments demonstrate that the proposed data-free method achieves comparable results of other data-acquired methods. Jinmian Cai, Zheng Wang 0059, Huazhu Fu, Jingjing Chen 0001, Yu-Gang Jiang 0001 |
ICME | 3 |
| 2022 | Unsupervised Lesion-Aware Transfer Learning for Diabetic Retinopathy Grading in Ultra-Wide-Field Fundus Photography
Yanmiao Bai, Jinkui Hao, Huazhu Fu, Xinting Ge, Jiang Liu 0001, Yitian Zhao, Jiong Zhang 0004 |
MICCAI (2) | 3 |
| 2022 | NerveFormer: A Cross-Sample Aggregation Network for Corneal Nerve Segmentation
Lei Mou, Shaodong Ma, Huazhu Fu, Lijun Guo, Yalin Zheng, Jiong Zhang 0004, Yitian Zhao |
MICCAI (4) | 4 |
| 2022 | Structure-Consistent Restoration Network for Cataract Fundus Image Enhancement
Heng Li 0010, Haofeng Liu, Huazhu Fu, Hai Shu, Yitian Zhao, Jiang Liu 0001 |
MICCAI (2) | 3 |
| 2022 | Instrument-tissue Interaction Quintuple Detection in Surgery Videos
Luoying Hao, Huazhu Fu, Cheekong Chui, Jiang Liu 0001 |
MICCAI (8) | 6 |
| 2022 | A New Dataset and a Baseline Model for Breast Lesion Detection in Ultrasound Videos
Lei Zhu 0003, Huazhu Fu, Harry Qin, Liansheng Wang 0002 |
MICCAI (3) | 4 |
| 2022 | Degradation-Invariant Enhancement of Fundus Images via Pyramid Constraint Network
Haofeng Liu, Heng Li 0010, Huazhu Fu, Ruoxiu Xiao, Yunshu Gao, Jiang Liu 0001 |
MICCAI (2) | 3 |
| 2022 | Screening of Dementia on OCTA Images via Multi-projection Consistency and Complementarity
Heng Li 0010, Zunjie Xiao, Huazhu Fu, Yitian Zhao, Richu Jin, William Robert Kwapong, Hanpei Miao, Jiang Liu 0001 |
MICCAI (2) | 4 |
| 2022 | Delving into Local Features for Open-Set Domain Adaptation in Fundus Image Analysis
Yi Zhou 0007, Shaochen Bai, Tao Zhou 0002, Yu Zhang 0009, Huazhu Fu |
MICCAI (8) | 5 |
| 2022 | TBraTS: Trusted Brain Tumor Segmentation
Ke Zou, Xuedong Yuan, Xiaojing Shen, Meng Wang 0001, Huazhu Fu |
MICCAI (8) | 5 |
| 2022 | Phase-based Memory Network for Video DehazingabstractVideo dehazing using deep-learning based methods has just received increasing attention in recent years. However, most existing methods tackle temporal consistency in the color domain only, which are less sensitive to small and imperceptible motions in a video, due to fog's drift and diffusion. In this work, we investigate in the frequency domain, which enables us to capture small motions effectively, and find that the phase component contains more semantic structures yet less haze information than the amplitude component of the hazy image. Based on these observations, we propose a novel phase-based memory network (PM-Net) to integrate the phase and color memory information for boosting video dehazing. Apart from the color memory from consecutive video frames, our PM-Net constructs a phase memory, which stores phase features of past video frames, and devise a cross-modal memory read (CMR) module, which fully leverages features from the color memory and the phase memory to boost features extracted from the current video frame for dehazing. Experimental results on the benchmark dataset of real hazy videos and a newly collected dataset of synthetic videos, show that the proposed PM-Net clearly outperforms the state-of-the-art image and video dehazing methods. Code is available at https://github.com/liuye123321/PM-Net. Huazhu Fu, Harry Qin, Lei Zhu 0003 |
ACM Multimedia | 3 |
| 2022 | Attention to region: Region-based integration-and-recalibration networks for nuclear cataract classification using AS-OCT imagesabstractNuclear cataract (NC) is a leading eye disease for blindness and vision impairment globally. Accurate and objective NC grading/classification is essential for clinically early intervention and cataract surgery planning. Anterior segment optical coherence tomography (AS-OCT) images are capable of capturing the nucleus region clearly and measuring the opacity of NC quantitatively. Recently, clinical research has suggested that the opacity correlation and repeatability between NC severity levels and the average nucleus density on AS-OCT images is high with the interclass and intraclass analysis. Moreover, clinical research has suggested that opacity distribution is uneven on the nucleus region, indicating that the opacities from different nucleus regions may play different roles in NC diagnosis. Motivated by the clinical priors, this paper proposes a simple yet effective region-based integration-and-recalibration attention (RIR), which integrates multiple feature map region representations and recalibrates the weights of each region via softmax attention adaptively. This region recalibration strategy enables the network to focus on high contribution region representations and suppress less useful ones. We combine the RIR block with the residual block to form a Residual-RIR module, and then a sequence of Residual-RIR modules are stacked to a deep network named region-based integration-and-recalibration network (RIR-Net), to predict NC severity levels automatically. The experiments on a clinical AS-OCT image dataset and two OCT datasets demonstrate that our method outperforms strong baselines and previous state-of-the-art methods. Furthermore, attention weight visualization analysis and ablation studies verify the capability of our RIR-Net for adjusting the relative importance of different regions in feature maps dynamically, agreeing with the clinical research. Xiaoqing Zhang 0001, Zunjie Xiao, Huazhu Fu, Yanwu Xu 0001, Risa Higashita, Jiang Liu 0001 |
Medical Image Anal. | 3 |
| 2022 | Re-Thinking Co-Salient Object DetectionabstractIn this article, we conduct a comprehensive study on the co-salient object detection (CoSOD) problem for images. CoSOD is an emerging and rapidly growing extension of salient object detection (SOD), which aims to detect the co-occurring salient objects in a group of images. However, existing CoSOD datasets often have a serious data bias, assuming that each group of images contains salient objects of similar visual appearances. This bias can lead to the ideal settings and effectiveness of models trained on existing datasets, being impaired in real-life situations, where similarities are usually semantic or conceptual. To tackle this issue, we first introduce a new benchmark, called CoSOD3k in the wild, which requires a large amount of semantic context, making it more challenging than existing CoSOD datasets. Our CoSOD3k consists of 3,316 high-quality, elaborately selected images divided into 160 groups with hierarchical annotations. The images span a wide range of categories, shapes, object sizes, and backgrounds. Second, we integrate the existing SOD techniques to build a unified, trainable CoSOD framework, which is long overdue in this field. Specifically, we propose a novel CoEG-Net that augments our prior model EGNet with a co-attention projection strategy to enable fast common information learning. CoEG-Net fully leverages previous large-scale SOD datasets and significantly improves the model scalability and stability. Third, we comprehensively summarize 40 cutting-edge algorithms, benchmarking 18 of them over three challenging CoSOD datasets (iCoSeg, CoSal2015, and our CoSOD3k), and reporting more detailed (i.e., group-level) performance analysis. Finally, we discuss the challenges and future works of CoSOD. We hope that our study will give a strong boost to growth in the CoSOD community. The benchmark toolbox and results are available on our project page at https://dpfan.net/CoSOD3K. Deng-Ping Fan, Tengpeng Li, Zheng Lin 0005, Ge-Peng Ji, Dingwen Zhang, Ming-Ming Cheng, Huazhu Fu, Jianbing Shen |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Salient Object Detection in the Deep Learning Era: An In-Depth SurveyabstractAs an essential problem in computer vision, salient object detection (SOD) has attracted an increasing amount of research attention over the years. Recent advances in SOD are predominantly led by deep learning-based solutions (named deep SOD). To enable in-depth understanding of deep SOD, in this paper, we provide a comprehensive survey covering various aspects, ranging from algorithm taxonomy to unsolved issues. In particular, we first review deep SOD algorithms from different perspectives, including network architecture, level of supervision, learning paradigm, and object-/instance-level detection. Following that, we summarize and analyze existing SOD datasets and evaluation metrics. Then, we benchmark a large group of representative SOD models, and provide detailed analyses of the comparison results. Moreover, we study the performance of SOD algorithms under different attribute settings, which has not been thoroughly explored previously, by constructing a novel SOD dataset with rich attribute annotations covering various salient object types, challenging factors, and scene categories. We further analyze, for the first time in the field, the robustness of SOD models to random input perturbations and adversarial attacks. We also look into the generalization and difficulty of existing SOD datasets. Finally, we discuss several open issues of SOD and outline future research directions. All the saliency prediction maps, our constructed dataset with annotations, and codes for evaluation are publicly available at https://github.com/wenguanwang/SODsurvey. Wenguan Wang, Qiuxia Lai, Huazhu Fu, Jianbing Shen, Haibin Ling, Ruigang Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2022 | Deep Partial Multi-View LearningabstractAlthough multi-view learning has made significant progress over the past few decades, it is still challenging due to the difficulty in modeling complex correlations among different views, especially under the context of view missing. To address the challenge, we propose a novel framework termed Cross Partial Multi-View Networks (CPM-Nets), which aims to fully and flexibly take advantage of multiple partial views. We first provide a formal definition of completeness and versatility for multi-view representation and then theoretically prove the versatility of the learned latent representations. For completeness, the task of learning latent multi-view representation is specifically translated to a degradation process by mimicking data transmission, such that the optimal tradeoff between consistency and complementarity across different views can be implicitly achieved. Equipped with adversarial strategy, our model stably imputes missing views, encoding information from all views for each sample to be encoded into latent representation to further enhance the completeness. Furthermore, a nonparametric classification loss is introduced to produce structured representations and prevent overfitting, which endows the algorithm with promising generalization under view-missing cases. Extensive experimental results validate the effectiveness of our algorithm over existing state of the arts for classification, representation learning and data imputation. Changqing Zhang 0002, Yajie Cui, Zongbo Han, Joey Tianyi Zhou, Huazhu Fu, Qinghua Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Deep-LIFT: Deep Label-Specific Feature Learning for Image AnnotationabstractImage annotation aims to jointly predict multiple tags for an image. Although significant progress has been achieved, existing approaches usually overlook aligning specific labels and their corresponding regions due to the weak supervised information (i.e., "bag of labels" for regions), thus failing to explicitly exploit the discrimination from different classes. In this article, we propose the deep label-specific feature (Deep-LIFT) learning model to build the explicit and exact correspondence between the label and the local visual region, which improves the effectiveness of feature learning and enhances the interpretability of the model itself. Deep-LIFT extracts features for each label by aligning each label and its region. Specifically, Deep-LIFTs are achieved through learning multiple correlation maps between image convolutional features and label embeddings. Moreover, we construct two variant graph convolutional networks (GCNs) to further capture the interdependency among labels. Empirical studies on benchmark datasets validate that the proposed model achieves superior performance on multilabel classification over other existing state-of-the-art methods. Junbing Li, Changqing Zhang 0002, Joey Tianyi Zhou, Huazhu Fu, Shuyin Xia, Qinghua Hu |
IEEE Trans. Cybern. | 4 |
| 2022 | Boosting RGB-D Saliency Detection by Leveraging Unlabeled RGB ImagesabstractTraining deep models for RGB-D salient object detection (SOD) often requires a large number of labeled RGB-D images. However, RGB-D data is not easily acquired, which limits the development of RGB-D SOD techniques. To alleviate this issue, we present a Dual-Semi RGB-D Salient Object Detection Network (DS-Net) to leverage unlabeled RGB images for boosting RGB-D saliency detection. We first devise a depth decoupling convolutional neural network (DDCNN), which contains a depth estimation branch and a saliency detection branch. The depth estimation branch is trained with RGB-D images and then used to estimate the pseudo depth maps for all unlabeled RGB images to form the paired data. The saliency detection branch is used to fuse the RGB feature and depth feature to predict the RGB-D saliency. Then, the whole DDCNN is assigned as the backbone in a teacher-student framework for semi-supervised learning. Moreover, we also introduce a consistency loss on the intermediate attention and saliency maps for the unlabeled data, as well as a supervised depth and saliency loss for labeled data. Experimental results on seven widely-used benchmark datasets demonstrate that our DDCNN outperforms state-of-the-art methods both quantitatively and qualitatively. We also demonstrate that our semi-supervised DS-Net can further improve the performance, even when using an RGB image with the pseudo depth map. Xiaoqiang Wang 0007, Lei Zhu 0003, Siliang Tang, Huazhu Fu, Ping Li 0016, Fei Wu 0001, Yi Yang 0001, Yueting Zhuang |
IEEE Trans. Image Process. | 4 |
| 2022 | Guest Editorial Generative Adversarial Networks in Biomedical Image ComputingabstractThe papers in this special section focus on generative adversarial networks in biomedical image computing. The field of biomedical imaging has obtained great progress from Roentgen’s original discovery of the X-ray to the current imaging tools, including Magnetic Resonance Imaging (MRI), Positron Emission Tomography (PET), Computed Tomography (CT), and Ultrasound (US). The benefits of using these non-invasive imaging technologies are to assess the current condition of an organ or tissue, which can be used to monitor a patient over time over time for accurate and timely diagnosis and treatment.With the development of imaging technologies, developing advanced artificial intelligence algorithms for automated image analysis has shown the potential to change many aspects of clinical applications within the next decade. Meanwhile, these advanced technologies have also brought new issues and challenges. Thus, there has been a growing demand for biomedical imaging computing to be a component of clinical trials and device improvement. Currently, Generative adversarial networks (GANs) have been attached growing interests in the computer vision community due to their capability of data generation or translation. GAN-based models are able to learn from a set of training data and generate new data with the same characteristics as the training ones, which have also proven to be the state of the art for generating sharp and realistic images. More importantly, GAN has been rapidly applied to many traditional and novel applications in the medical domain, such as image reconstruction, segmentation, diagnosis, synthesis, and so on. Despite GAN substantial progress in these areas, their application to medical image computing still faces challenges and unsolved problems remain. Huazhu Fu, Tao Zhou 0002, Shuo Li 0001, Alejandro F. Frangi |
IEEE J. Biomed. Health Informatics | 1 |
| 2022 | ADAM Challenge: Detecting Age-Related Macular Degeneration From Fundus ImagesabstractAge-related macular degeneration (AMD) is the leading cause of visual impairment among elderly in the world. Early detection of AMD is of great importance, as the vision loss caused by this disease is irreversible and permanent. Color fundus photography is the most cost-effective imaging modality to screen for retinal disorders. Cutting edge deep learning based algorithms have been recently developed for automatically detecting AMD from fundus images. However, there are still lack of a comprehensive annotated dataset and standard evaluation benchmarks. To deal with this issue, we set up the Automatic Detection challenge on Age-related Macular degeneration (ADAM), which was held as a satellite event of the ISBI 2020 conference. The ADAM challenge consisted of four tasks which cover the main aspects of detecting and characterizing AMD from fundus images, including detection of AMD, detection and segmentation of optic disc, localization of fovea, and detection and segmentation of lesions. As part of the ADAM challenge, we have released a comprehensive dataset of 1200 fundus images with AMD diagnostic labels, pixel-wise segmentation masks for both optic disc and AMD-related lesions (drusen, exudates, hemorrhages and scars, among others), as well as the coordinates corresponding to the location of the macular fovea. A uniform evaluation framework has been built to make a fair comparison of different models using this dataset. During the ADAM challenge, 610 results were submitted for online evaluation, with 11 teams finally participating in the onsite challenge. This paper introduces the challenge, the dataset and the evaluation methods, as well as summarizes the participating methods and analyzes their results for each task. In particular, we observed that the ensembling strategy and the incorporation of clinical domain knowledge were the key to improve the performance of the deep learning models. Huihui Fang, Fei Li 0021, Huazhu Fu, Xu Sun 0006, Xingxing Cao, Fengbin Lin, Jaemin Son, Gwenolé Quellec, Sarah Matta, Sharath M. Shankaranarayana, Chuen-heng Wang, Nisarg A. Shah, Chia-Yen Lee, Chih-Chung Hsu, Hai Xie, Bai Ying Lei, Ujjwal Baid, Shubham Innani, Kang Dang, Wenxiu Shi, Ravi Kamble, Nitin Singhal, Ching-Wei Wang, Shih-Chang Lo, José Ignacio Orlando, Hrvoje Bogunovic, Xiulan Zhang, Yanwu Xu 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2022 | Hybrid Variation-Aware Network for Angle-Closure Assessment in AS-OCTabstractAutomatic angle-closure assessment in Anterior Segment OCT (AS-OCT) images is an important task for the screening and diagnosis of glaucoma, and the most recent computer-aided models focus on a binary classification of anterior chamber angles (ACA) in AS-OCT, i.e., open-angle and angle-closure. In order to assist clinicians who seek better to understand the development of the spectrum of glaucoma types, a more discriminating three-class classification scheme was suggested, i.e., the classification of ACA was expended to include open-, appositional- and synechial angles. However, appositional and synechial angles display similar appearances in an AS-OCT image, which makes classification models struggle to differentiate angle-closure subtypes based on static AS-OCT images. In order to tackle this issue, we propose a 2D-3D Hybrid Variation-aware Network (HV-Net) for open-appositional-synechial ACA classification from AS-OCT imagery. Specifically, taking into account clinical priors, we first reconstruct the 3D iris surface from an AS-OCT sequence, and obtain the geometrical characteristics necessary to provide global shape information. 2D AS-OCT slices and 3D iris representations are then fed into our HV-Net to extract cross-sectional appearance features and iris morphological features, respectively. To achieve similar results to those of dynamic gonioscopy examination, which is the current gold standard for diagnostic angle assessment, the paired AS-OCT images acquired in dark and light illumination conditions are used to obtain an accurate characterization of configurational changes in ACAs and iris shapes, using a Variation-aware Block. In addition, an annealing loss function was introduced to optimize our model, so as to encourage the sub-networks to map the inputs into the more conducive spaces to extract dark-to-light variation representations, while retaining the discriminative power of the learned features. The proposed model is evaluated across 1584 paired AS-OCT samples, and it has demonstrated its superiority in classifying open-, appositional- and synechial angles. Jinkui Hao, Fei Li 0021, Huaying Hao, Huazhu Fu, Yanwu Xu 0001, Risa Higashita, Xiulan Zhang, Jiang Liu 0001, Yitian Zhao |
IEEE Trans. Medical Imaging | 4 |
| 2022 | Progressive Multiscale Consistent Network for Multiclass Fundus Lesion SegmentationabstractEffectively integrating multi-scale information is of considerable significance for the challenging multi-class segmentation of fundus lesions because different lesions vary significantly in scales and shapes. Several methods have been proposed to successfully handle the multi-scale object segmentation. However, two issues are not considered in previous studies. The first is the lack of interaction between adjacent feature levels, and this will lead to the deviation of high-level features from low-level features and the loss of detailed cues. The second is the conflict between the low-level and high-level features, this occurs because they learn different scales of features, thereby confusing the model and decreasing the accuracy of the final prediction. In this paper, we propose a progressive multi-scale consistent network (PMCNet) that integrates the proposed progressive feature fusion (PFF) block and dynamic attention block (DAB) to address the aforementioned issues. Specifically, PFF block progressively integrates multi-scale features from adjacent encoding layers, facilitating feature learning of each layer by aggregating fine-grained details and high-level semantics. As features at different scales should be consistent, DAB is designed to dynamically learn the attentive cues from the fused features at different scales, thus aiming to smooth the essential conflicts existing in multi-scale features. The two proposed PFF and DAB blocks can be integrated with the off-the-shelf backbone networks to address the two issues of multi-scale and feature inconsistency in the multi-class segmentation of fundus lesions, which will produce better feature representation in the feature space. Experimental results on three public datasets indicate that the proposed method is more effective than recent state-of-the-art methods. Along He, Kai Wang 0001, Tao Li 0022, Wang Bo, Hong Kang, Huazhu Fu |
IEEE Trans. Medical Imaging | 6 |
| 2022 | An Annotation-Free Restoration Network for Cataractous Fundus ImagesabstractCataracts are the leading cause of vision loss worldwide. Restoration algorithms are developed to improve the readability of cataract fundus images in order to increase the certainty in diagnosis and treatment for cataract patients. Unfortunately, the requirement of annotation limits the application of these algorithms in clinics. This paper proposes a network to annotation-freely restore cataractous fundus images (ArcNet) so as to boost the clinical practicability of restoration. Annotations are unnecessary in ArcNet, where the high-frequency component is extracted from fundus images to replace segmentation in the preservation of retinal structures. The restoration model is learned from the synthesized images and adapted to real cataract images. Extensive experiments are implemented to verify the performance and effectiveness of ArcNet. Favorable performance is achieved using ArcNet against state-of-the-art algorithms, and the diagnosis of ocular fundus diseases in cataract patients is promoted by ArcNet. The capability of properly restoring cataractous images in the absence of annotated data promises the proposed algorithm outstanding clinical practicability. Heng Li 0010, Haofeng Liu, Huazhu Fu, Yitian Zhao, Hanpei Miao, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2022 | Proxy-Bridged Image Reconstruction Network for Anomaly Detection in Medical ImagesabstractAnomaly detection in medical images refers to the identification of abnormal images with only normal images in the training set. Most existing methods solve this problem with a self-reconstruction framework, which tends to learn an identity mapping and reduces the sensitivity to anomalies. To mitigate this problem, in this paper, we propose a novel Proxy-bridged Image Reconstruction Network (ProxyAno) for anomaly detection in medical images. Specifically, we use an intermediate proxy to bridge the input image and the reconstructed image. We study different proxy types, and we find that the superpixel-image (SI) is the best one. We set all pixels' intensities within each superpixel as their average intensity, and denote this image as SI. The proposed ProxyAno consists of two modules, a Proxy Extraction Module and an Image Reconstruction Module. In the Proxy Extraction Module, a memory is introduced to memorize the feature correspondence for normal image to its corresponding SI, while the memorized correspondence does not apply to the abnormal images, which leads to the information loss for abnormal image and facilitates the anomaly detection. In the Image Reconstruction Module, we map an SI to its reconstructed image. Further, we crop a patch from the image and paste it on the normal SI to mimic the anomalies, and enforce the network to reconstruct the normal image even with the pseudo abnormal SI. In this way, our network enlarges the reconstruction error for anomalies. Extensive experiments on brain MR images, retinal OCT images and retinal fundus images verify the effectiveness of our method for both image-level and pixel-level anomaly detection. Kang Zhou 0001, Jing Li 0117, Weixin Luo, Jianlong Yang, Huazhu Fu, Jun Cheng 0003, Jiang Liu 0001, Shenghua Gao |
IEEE Trans. Medical Imaging | 6 |
| 2021 | Triple-Cooperative Video Shadow DetectionabstractShadow detection in a single image has received significant research interests in recent years. However, much fewer works have been explored in shadow detection over dynamic scenes. The bottleneck is the lack of a well-established dataset with high-quality annotations for video shadow detection. In this work, we collect a new video shadow detection dataset (ViSha), which contains 120 videos with 11,685 frames, covering 60 object categories, varying lengths, and different motion/lighting conditions. All the frames are annotated with a high-quality pixel-level shadow mask. To the best of our knowledge, this is the first learning-oriented dataset for video shadow detection. Furthermore, we develop a new baseline model, named triple-cooperative video shadow detection network (TVSD-Net). It utilizes triple parallel networks in a cooperative manner to learn discriminative representations at intra-video and inter-video levels. Within the network, a dual gated co-attention module is proposed to constrain features from neighboring frames in the same video, while an auxiliary similarity loss is introduced to mine semantic information between different videos. Finally, we conduct a comprehensive study on ViSha, evaluating 12 state-of-the-art models (including single image shadow detectors, video object segmentation, and saliency detection methods). Experiments demonstrate that our model outperforms SOTA competitors. Zhihao Chen 0004, Lei Zhu 0003, Huazhu Fu, Wennan Liu, Harry Qin |
CVPR | 5 |
| 2021 | Group Collaborative Learning for Co-Salient Object DetectionabstractWe present a novel group collaborative learning framework (GCoNet) capable of detecting co-salient objects in real time (16ms), by simultaneously mining consensus representations at group level based on the two necessary criteria: 1) intra-group compactness to better formulate the consistency among co-salient objects by capturing their inherent shared attributes using our novel group affinity module; 2) inter-group separability to effectively suppress the influence of noisy objects on the output by introducing our new group collaborating module conditioning the inconsistent consensus. To learn a better embedding space without extra computational overhead, we explicitly employ auxiliary classification supervision. Extensive experiments on three challenging benchmarks, i.e., CoCA, CoSOD3k, and Cosal2015, demonstrate that our simple GCoNet outperforms 10 cutting-edge models and achieves the new state-of-the-art. We demonstrate this paper’s new technical contributions on a number of important downstream computer vision applications including content aware co-segmentation, co-localization based automatic thumbnails, etc. Code has been made publicly available: https://github.com/fanq15/GCoNet. Deng-Ping Fan, Huazhu Fu, Chi-Keung Tang, Ling Shao 0001, Yu-Wing Tai |
CVPR | 3 |
| 2021 | Specificity-preserving RGB-D Saliency Detection
Tao Zhou 0002, Huazhu Fu, Geng Chen 0001, Yi Zhou 0007, Deng-Ping Fan, Ling Shao 0001 |
ICCV | 2 |
| 2021 | Visual-Textual Attentive Semantic Consistency for Medical Report GenerationabstractAutomatic report generation on medical radiographs have recently gained interest. However, identifying diseases as well as correctly predicting their corresponding sizes, locations and other medical description patterns, which is essential for generating high-quality reports, is challenging. Although previous methods focused on producing readable reports, how to accurately detect and describe findings that match with the query X-Ray has not been successfully addressed. In this paper, we propose a multi-modality semantic attention model to integrate visual features, predicted key finding embeddings, as well as clinical features, and progressively decode reports with visual-textual semantic consistency. First, multi-modality features are extracted and attended with the hidden states from the sentence de-coder, to encode enriched context vectors for better decoding a report. These modalities include regional visual features of scans, semantic word embeddings of the top-K findings predicted with high probabilities, and clinical features of indications. Second, the progressive report decoder consists of a sentence decoder and a word decoder, where we propose image-sentence matching and description accuracy losses to constrain the visual-textual semantic consistency. Extensive experiments on the public MIMIC-CXR and IU X-Ray datasets show that our model achieves consistent improvements over the state-of-the-art methods. Yi Zhou 0007, Lei Huang 0015, Tao Zhou 0002, Huazhu Fu, Ling Shao 0001 |
ICCV | 4 |
| 2021 | VIL-100: A New Dataset and A Baseline Model for Video Instance Lane DetectionabstractLane detection plays a key role in autonomous driving. While car cameras always take streaming videos on the way, current lane detection works mainly focus on individual images (frames) by ignoring dynamics along the video. In this work, we collect a new video instance lane detection (VIL-100) dataset, which contains 100 videos with in total 10,000 frames, acquired from different real traffic scenarios. All the frames in each video are manually annotated to a high-quality instance-level lane annotation, and a set of frame-level and video-level metrics are included for quantitative performance evaluation. Moreover, we propose a new baseline model, named multi-level memory aggregation network (MMA-Net), for video instance lane detection. In our approach, the representation of current frame is enhanced by attentively aggregating both local and global memory features from other frames. Experiments on the new collected dataset show that the proposed MMA-Net outperforms state-of-the-art lane detection methods and video object segmentation methods. We release our dataset and code at https://github.com/yujun0-0/MMA-Net. Yujun Zhang 0002, Lei Zhu 0003, Wei Feng 0005, Huazhu Fu, Qingxia Li, Song Wang 0002 |
ICCV | 4 |
| 2021 | VideoLT: Large-scale Long-tailed Video RecognitionabstractLabel distributions in real-world are oftentimes long-tailed and imbalanced, resulting in biased models towards dominant labels. While long-tailed recognition has been extensively studied for image classification tasks, limited effort has been made for the video domain. In this paper, we introduce VideoLT, a large-scale long-tailed video recognition dataset, as a step toward real-world video recognition. VideoLT contains 256,218 untrimmed videos, annotated into 1,004 classes with a long-tailed distribution. Through extensive studies, we demonstrate that state-of-the-art methods used for long-tailed image recognition do not perform well in the video domain due to the additional temporal dimension in videos. This motivates us to propose FrameStack, a simple yet effective method for long-tailed video recognition. In particular, FrameStack performs sampling at the frame-level in order to balance class distributions, and the sampling ratio is dynamically determined using knowledge derived from the network during training. Experimental results demonstrate that FrameStack can improve classification performance without sacrificing the overall accuracy. Code and dataset are available at: https://github.com/17Skye17/VideoLT. Xing Zhang 0013, Zuxuan Wu, Zejia Weng, Huazhu Fu, Jingjing Chen 0001, Yu-Gang Jiang 0001, Larry Davis 0001 |
ICCV | 4 |
| 2021 | Trusted Multi-View Classification
Zongbo Han, Changqing Zhang 0002, Huazhu Fu, Joey Tianyi Zhou |
ICLR | 3 |
| 2021 | Cross-View Equivariant Auto-EncoderabstractUnsupervised representation learning on multi-view data (multiple types of features or modalities) becomes a compelling topic in machine learning. Most existing methods focus on directly projecting different views into a common space to explore the consistency across different views. Al-though simple, the underlying relationships among different views are not guaranteed during the learning process. In this paper, we propose a novel unsupervised multi-view representation learning model termed as Cross-View Equivariant Auto-Encoder (CVE-AE), which jointly conducts data re-construction with view-specific autoencoder for information preservation within each view, and transformation reconstruction with transformation decoder for correlations preservation across different views. Accordingly, the generalization ability of our model is promoted due to the preserved intra-view intrinsic information and underlying inter-view relationships. We conduct extensive experiments on real-world datasets, and the proposed model achieves superior performance over state-of-the-art unsupervised representation learning methods. Zhibin Wan, Changqing Zhang 0002, Huazhu Fu, Xi Peng 0001, Pengfei Zhu 0001, Qinghua Hu |
ICME | 4 |
| 2021 | Multi-contrast MRI Super-Resolution via a Multi-stage Integration Network
Chun-Mei Feng 0001, Huazhu Fu, Shuhao Yuan, Yong Xu 0001 |
MICCAI (6) | 2 |
| 2021 | Task Transformer Network for Joint MRI Reconstruction and Super-Resolution
Chun-Mei Feng 0001, Yunlu Yan, Huazhu Fu, Li Chen 0011, Yong Xu 0001 |
MICCAI (6) | 3 |
| 2021 | Progressively Normalized Self-Attention Network for Video Polyp Segmentation
Ge-Peng Ji, Yu-Cheng Chou, Deng-Ping Fan, Geng Chen 0001, Huazhu Fu, Debesh Jha, Ling Shao 0001 |
MICCAI (1) | 5 |
| 2021 | Few-Shot Domain Adaptation with Polymorphic Transformers
Shaohua Li 0003, Xiuchao Sui, Huazhu Fu, Xiangde Luo, Yangqin Feng, Xinxing Xu, Yong Liu 0026, Daniel S. W. Ting, Rick Siow Mong Goh |
MICCAI (2) | 4 |
| 2021 | A Multi-branch Hybrid Transformer Network for Corneal Endothelial Cell Segmentation
Yinglin Zhang, Risa Higashita, Huazhu Fu, Yanwu Xu 0001, Haofeng Liu, Jian Zhang 0002, Jiang Liu 0001 |
MICCAI (1) | 3 |
| 2021 | From Synthetic to Real: Image Dehazing Collaborating with Unlabeled Real DataabstractSingle image dehazing is a challenging task, for which the domain shift between synthetic training data and real-world testing images usually leads to degradation of existing methods. To address this issue, we propose a novel image dehazing framework collaborating with unlabeled real data. First, we develop a disentangled image dehazing network (DID-Net), which disentangles the feature representations into three component maps, i.e. the latent haze-free image, the transmission map, and the global atmospheric light estimate, respecting the physical model of a haze process. Our DID-Net predicts the three component maps by progressively integrating features across scales, and refines each map by passing an independent refinement network. Then a disentangled-consistency mean-teacher network (DMT-Net) is employed to collaborate unlabeled real data for boosting single image dehazing. Specifically, we encourage the coarse predictions and refinements of each disentangled component to be consistent between the student and teacher networks by using a consistency loss on unlabeled real data. We make comparison with 13 state-of-the-art dehazing methods on a new collected dataset (Haze4K) and two widely-used dehazing datasets (i.e., SOTS and HazeRD), as well as on real-world hazy images. Experimental results demonstrate that our method has obvious quantitative and qualitative improvements over the existing methods. Lei Zhu 0003, Shunda Pei, Huazhu Fu, Harry Qin, Qing Zhang 0006, Wei Feng 0005 |
ACM Multimedia | 4 |
| 2021 | Trustworthy Multimodal Regression with Mixture of Normal-inverse Gamma DistributionsabstractMultimodal regression is a fundamental task, which integrates the information from different sources to improve the performance of follow-up applications. However, existing methods mainly focus on improving the performance and often ignore the confidence of prediction for diverse situations. In this study, we are devoted to trustworthy multimodal regression which is critical in cost-sensitive domains. To this end, we introduce a novel Mixture of Normal-Inverse Gamma distributions (MoNIG) algorithm, which efficiently estimates uncertainty in principle for adaptive integration of different modalities and produces a trustworthy regression result. Our model can be dynamically aware of uncertainty for each modality, and also robust for corrupted modalities. Furthermore, the proposed MoNIG ensures explicitly representation of (modality-specific/global) epistemic and aleatoric uncertainties, respectively. Experimental results on both synthetic and different real-world data demonstrate the effectiveness and trustworthiness of our method on various multimodal regression tasks (e.g., temperature prediction for superconductivity, relative location prediction for CT slices, and multimodal sentiment analysis). Huan Ma 0006, Zongbo Han, Changqing Zhang 0002, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu |
NeurIPS | 4 |
| 2021 | Deep video action clustering via spatio-temporal feature learning
Bo Peng 0007, Jianjun Lei 0001, Huazhu Fu, Yalong Jia, Zongqian Zhang |
Neurocomputing | 3 |
| 2021 | RGB-D salient object detection via cross-modal joint feature extraction and low-bound fusion loss
Huazhu Fu, Xiaoting Fan, Yanan Shi, Jianjun Lei 0001 |
Neurocomputing | 3 |
| 2021 | Applications of deep learning in fundus images: A review
Tao Li 0022, Wang Bo, Hong Kang, Hanruo Liu, Kai Wang 0001, Huazhu Fu |
Medical Image Anal. | 7 |
| 2021 | Deep triplet hashing network for case-based medical image retrieval
Jiansheng Fang, Huazhu Fu, Jiang Liu 0001 |
Medical Image Anal. | 2 |
| 2021 | CS2-Net: Deep learning segmentation of curvilinear structures in medical imaging
Lei Mou, Yitian Zhao, Huazhu Fu, Yonghuai Liu, Jun Cheng 0003, Yalin Zheng, Pan Su 0001, Jianlong Yang, Li Chen 0011, Alejandro F. Frangi, Masahiro Akiba, Jiang Liu 0001 |
Medical Image Anal. | 3 |
| 2021 | Global guidance network for breast lesion segmentation in ultrasound images
Cheng Xue 0003, Lei Zhu 0003, Huazhu Fu, Xiaowei Hu 0001, Xiaomeng Li 0001, Pheng-Ann Heng |
Medical Image Anal. | 3 |
| 2021 | ASIF-Net: Attention Steered Interweave Fusion Network for RGB-D Salient Object DetectionabstractSalient object detection from RGB-D images is an important yet challenging vision task, which aims at detecting the most distinctive objects in a scene by combining color information and depth constraints. Unlike prior fusion manners, we propose an attention steered interweave fusion network (ASIF-Net) to detect salient objects, which progressively integrates cross-modal and cross-level complementarity from the RGB image and corresponding depth map via steering of an attention mechanism. Specifically, the complementary features from RGB-D images are jointly extracted and hierarchically fused in a dense and interweaved manner. Such a manner breaks down the barriers of inconsistency existing in the cross-modal data and also sufficiently captures the complementarity. Meanwhile, an attention mechanism is introduced to locate the potential salient regions in an attention-weighted fashion, which advances in highlighting the salient objects and suppressing the cluttered background regions. Instead of focusing only on pixelwise saliency, we also ensure that the detected salient objects have the objectness characteristics (e.g., complete structure and sharp boundary) by incorporating the adversarial learning that provides a global semantic constraint for RGB-D salient object detection. Quantitative and qualitative experiments demonstrate that the proposed method performs favorably against 17 state-of-the-art saliency detectors on four publicly available RGB-D salient object detection datasets. The code and results of our method are available at https://github.com/Li-Chongyi/ASIF-Net. Chongyi Li, Runmin Cong, Sam Kwong, Junhui Hou, Huazhu Fu, Guopu Zhu, Dingwen Zhang, Qingming Huang |
IEEE Trans. Cybern. | 5 |
| 2021 | Combating Ambiguity for Hash-Code Learning in Medical Instance RetrievalabstractWhen encountering a dubious diagnostic case, medical instance retrieval can help radiologists make evidence-based diagnoses by finding images containing instances similar to a query case from a large image database. The similarity between the query case and retrieved similar cases is determined by visual features extracted from pathologically abnormal regions. However, the manifestation of these regions often lacks specificity, i.e., different diseases can have the same manifestation, and different manifestations may occur at different stages of the same disease. To combat the manifestation ambiguity in medical instance retrieval, we propose a novel deep framework called Y-Net, encoding images into compact hash-codes generated from convolutional features by feature aggregation. Y-Net can learn highly discriminative convolutional features by unifying the pixel-wise segmentation loss and classification loss. The segmentation loss allows exploring subtle spatial differences for good spatial-discriminability while the classification loss utilizes class-aware semantic information for good semantic-separability. As a result, Y-Net can enhance the visual features in pathologically abnormal regions and suppress the disturbing of the background during model training, which could effectively embed discriminative features into the hash-codes in the retrieval stage. Extensive experiments on two medical image datasets demonstrate that Y-Net can alleviate the ambiguity of pathologically abnormal regions and its retrieval performance outperforms the state-of-the-art method by an average of 9.27% on the returned list of 10. Jiansheng Fang, Huazhu Fu, Dan Zeng 0002, Xiao Yan 0002, Yuguang Yan, Jiang Liu 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2021 | CABNet: Category Attention Block for Imbalanced Diabetic Retinopathy GradingabstractDiabetic Retinopathy (DR) grading is challenging due to the presence of intra-class variations, small lesions and imbalanced data distributions. The key for solving fine-grained DR grading is to find more discriminative features corresponding to subtle visual differences, such as microaneurysms, hemorrhages and soft exudates. However, small lesions are quite difficult to identify using traditional convolutional neural networks (CNNs), and an imbalanced DR data distribution will cause the model to pay too much attention to DR grades with more samples, greatly affecting the final grading performance. In this article, we focus on developing an attention module to address these issues. Specifically, for imbalanced DR data distributions, we propose a novel Category Attention Block (CAB), which explores more discriminative region-wise features for each DR grade and treats each category equally. In order to capture more detailed small lesion information, we also propose the Global Attention Block (GAB), which can exploit detailed and class-agnostic global attention feature maps for fundus images. By aggregating the attention blocks with a backbone network, the CABNet is constructed for DR grading. The attention blocks can be applied to a wide range of backbone networks and trained efficiently in an end-to-end manner. Comprehensive experiments are conducted on three publicly available datasets, showing that CABNet produces significant performance improvements for existing state-of-the-art deep architectures with few additional parameters and achieves the state-of-the-art results for DR grading. Code and models will be available at https://github.com/he2016012996/CABnet. Along He, Tao Li 0022, Kai Wang 0001, Huazhu Fu |
IEEE Trans. Medical Imaging | 5 |
| 2021 | ROSE: A Retinal OCT-Angiography Vessel Segmentation Dataset and New ModelabstractOptical Coherence Tomography Angiography (OCTA) is a non-invasive imaging technique that has been increasingly used to image the retinal vasculature at capillary level resolution. However, automated segmentation of retinal vessels in OCTA has been under-studied due to various challenges such as low capillary visibility and high vessel complexity, despite its significance in understanding many vision-related diseases. In addition, there is no publicly available OCTA dataset with manually graded vessels for training and validation of segmentation algorithms. To address these issues, for the first time in the field of retinal image analysis we construct a dedicated Retinal OCTA SEgmentation dataset (ROSE), which consists of 229 OCTA images with vessel annotations at either centerline-level or pixel level. This dataset with the source code has been released for public access to assist researchers in the community in undertaking research in related topics. Secondly, we introduce a novel split-based coarse-to-fine vessel segmentation network for OCTA images (OCTA-Net), with the ability to detect thick and thin vessels separately. In the OCTA-Net, a split-based coarse segmentation module is first utilized to produce a preliminary confidence map of vessels, and a split-based refined segmentation module is then used to optimize the shape/contour of the retinal microvasculature. We perform a thorough evaluation of the state-of-the-art vessel segmentation models and our OCTA-Net on the constructed ROSE dataset. The experimental results demonstrate that our OCTA-Net yields better vessel segmentation performance in OCTA than both traditional and other deep learning methods. In addition, we provide a fractal dimension analysis on the segmented microvasculature, and the statistical analysis demonstrates significant differences between the healthy control and Alzheimer's Disease group. This consolidates that the analysis of retinal microvasculature may offer a new scheme to study various neurodegenerative diseases. Yuhui Ma, Huaying Hao, Jianyang Xie, Huazhu Fu, Jiong Zhang 0004, Jianlong Yang, Jiang Liu 0001, Yalin Zheng, Yitian Zhao |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Structure and Illumination Constrained GAN for Medical Image EnhancementabstractThe development of medical imaging techniques has greatly supported clinical decision making. However, poor imaging quality, such as non-uniform illumination or imbalanced intensity, brings challenges for automated screening, analysis and diagnosis of diseases. Previously, bi-directional GANs (e.g., CycleGAN), have been proposed to improve the quality of input images without the requirement of paired images. However, these methods focus on global appearance, without imposing constraints on structure or illumination, which are essential features for medical image interpretation. In this paper, we propose a novel and versatile bi-directional GAN, named Structure and illumination constrained GAN (StillGAN), for medical image quality enhancement. Our StillGAN treats low- and high-quality images as two distinct domains, and introduces local structure and illumination constraints for learning both overall characteristics and local details. Extensive experiments on three medical image datasets (e.g., corneal confocal microscopy, retinal color fundus and endoscopy images) demonstrate that our method performs better than both conventional methods and other deep learning-based methods. In addition, we have investigated the impact of the proposed method on different medical image analysis and clinical tasks such as nerve segmentation, tortuosity grading, fovea localization and disease classification. Yuhui Ma, Jiang Liu 0001, Yonghuai Liu, Huazhu Fu, Jun Cheng 0003, Yufei Wu 0013, Jiong Zhang 0004, Yitian Zhao |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Modeling and Enhancing Low-Quality Retinal Fundus ImagesabstractRetinal fundus images are widely used for the clinical screening and diagnosis of eye diseases. However, fundus images captured by operators with various levels of experience have a large variation in quality. Low-quality fundus images increase uncertainty in clinical observation and lead to the risk of misdiagnosis. However, due to the special optical beam of fundus imaging and structure of the retina, natural image enhancement methods cannot be utilized directly to address this. In this article, we first analyze the ophthalmoscope imaging system and simulate a reliable degradation of major inferior-quality factors, including uneven illumination, image blurring, and artifacts. Then, based on the degradation model, a clinically oriented fundus enhancement network (cofe-Net) is proposed to suppress global degradation factors, while simultaneously preserving anatomical retinal structures and pathological characteristics for clinical observation and analysis. Experiments on both synthetic and real images demonstrate that our algorithm effectively corrects low-quality fundus images without losing retinal details. Moreover, we also show that the fundus correction method can benefit medical image analysis applications, e.g., retinal vessel segmentation and optic disc/cup detection. Ziyi Shen, Huazhu Fu, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2021 | Contrast-Attentive Thoracic Disease Recognition With Dual-Weighting Graph ReasoningabstractAutomatic thoracic disease diagnosis is a rising research topic in the medical imaging community, with many potential applications. However, the inconsistent appearances and high complexities of various lesions in chest X-rays currently hinder the development of a reliable and robust intelligent diagnosis system. Attending to the high-probability abnormal regions and exploiting the priori of a related knowledge graph offers one promising route to addressing these issues. As such, in this paper, we propose two contrastive abnormal attention models and a dual-weighting graph convolution to improve the performance of thoracic multi-disease recognition. First, a left-right lung contrastive network is designed to learn intra-attentive abnormal features to better identify the most common thoracic diseases, whose lesions rarely appear in both sides symmetrically. Moreover, an inter-contrastive abnormal attention model aims to compare the query scan with multiple anchor scans without lesions to compute the abnormal attention map. Once the intra- and inter-contrastive attentions are weighted over the features, in addition to the basic visual spatial convolution, a chest radiology graph is constructed for dual-weighting graph reasoning. Extensive experiments on the public NIH ChestX-ray and CheXpert datasets show that our model achieves consistent improvements over the state-of-the-art methods both on thoracic disease identification and localization. Yi Zhou 0007, Tianfei Zhou, Tao Zhou 0002, Huazhu Fu, Jiacheng Liu 0011, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Multi-Mutual Consistency Induced Transfer Subspace Learning for Human Motion SegmentationabstractHuman motion segmentation based on transfer subspace learning is a rising interest in action-related tasks. Although progress has been made, there are still several issues within the existing methods. First, existing methods transfer knowledge from source data to target tasks by learning domain-invariant features, but they ignore to preserve domain-specific knowledge. Second, the transfer subspace learning is employed in either low-level or high-level feature spaces, but few methods consider fusing multi-level features for subspace learning. To this end, we propose a novel multi-mutual consistency induced transfer subspace learning framework for human motion segmentation. Specifically, our model factorizes the source and target data into distinct multi-layer feature spaces and reduces the distribution gap between them through a multi-mutual consistency learning strategy. In this way, the domain-specific knowledge and domain-invariant properties can be explored simultaneously. Our model also conducts the transfer subspace learning on different layers to capture multi-level structural information. Further, to preserve the temporal correlations, we project the learned representations into a block-like space. The proposed model is efficiently optimized by using the Augmented Lagrange Multiplier (ALM) algorithm. Experimental results on four human motion datasets demonstrate the effectiveness of our method over other state-of-the-art approaches. Tao Zhou 0002, Huazhu Fu, Chen Gong 0002, Jianbing Shen, Ling Shao 0001, Fatih Porikli |
CVPR | 2 |
| 2020 | Taking a Deeper Look at Co-Salient Object DetectionabstractCo-salient object detection (CoSOD) is a newly emerging and rapidly growing branch of salient object detection (SOD), which aims to detect the co-occurring salient objects in multiple images. However, existing CoSOD datasets often have a serious data bias, which assumes that each group of images contains salient objects of similar visual appearances. This bias results in the ideal settings and the effectiveness of the models, trained on existing datasets, may be impaired in real-life situations, where the similarity is usually semantic or conceptual. To tackle this issue, we first collect a new high-quality dataset, named CoSOD3k, which contains 3,316 images divided into 160 groups with multiple level annotations, i.e., category, bounding box, object, and instance levels. CoSOD3k makes a significant leap in terms of diversity, difficulty and scalability, benefiting related vision tasks. Besides, we comprehensively summarize 34 cutting-edge algorithms, benchmarking 19 of them over four existing CoSOD datasets (MSRC, iCoSeg, Image Pair and CoSal2015) and our CoSOD3k with a total of ~61K images (largest scale), and reporting group-level performance analysis. Finally, we discuss the challenge and future work of CoSOD. Our study would give a strong boost to growth in the CoSOD community. Benchmark toolbox and results are available on our project page. Deng-Ping Fan, Zheng Lin 0005, Ge-Peng Ji, Dingwen Zhang, Huazhu Fu, Ming-Ming Cheng |
CVPR | 5 |
| 2020 | PraNet: Parallel Reverse Attention Network for Polyp Segmentation
Deng-Ping Fan, Ge-Peng Ji, Tao Zhou 0002, Geng Chen 0001, Huazhu Fu, Jianbing Shen, Ling Shao 0001 |
MICCAI (6) | 5 |
| 2020 | Reconstruction and Quantification of 3D Iris Surface for Angle-Closure Glaucoma Detection in Anterior Segment OCT
Jinkui Hao, Huazhu Fu, Yanwu Xu 0001, Fei Li 0021, Xiulan Zhang, Jiang Liu 0001, Yitian Zhao |
MICCAI (5) | 2 |
| 2020 | Open-Appositional-Synechial Anterior Chamber Angle Classification in AS-OCT Sequences
Huaying Hao, Huazhu Fu, Yanwu Xu 0001, Jianlong Yang, Fei Li 0021, Xiulan Zhang, Jiang Liu 0001, Yitian Zhao |
MICCAI (5) | 2 |
| 2020 | Retinal Image Segmentation with a Structure-Texture Demixing Network
Huazhu Fu, Yanwu Xu 0001, Mingkui Tan |
MICCAI (5) | 2 |
| 2020 | M2 Net: Multi-modal Multi-channel Network for Overall Survival Time Prediction of Brain Tumor Patients
Tao Zhou 0002, Huazhu Fu, Yu Zhang 0009, Changqing Zhang 0002, Xiankai Lu, Jianbing Shen, Ling Shao 0001 |
MICCAI (2) | 2 |
| 2020 | A Second-Order Subregion Pooling Network for Breast Lesion Segmentation in Ultrasound
Lei Zhu 0003, Rongzhen Chen, Huazhu Fu, Liansheng Wang 0002, Pheng-Ann Heng |
MICCAI (6) | 3 |
| 2020 | NuI-Go: Recursive Non-Local Encoder-Decoder Network for Retinal Image Non-Uniform Illumination RemovalabstractRetinal images have been widely used by clinicians for early diagnosis of ocular diseases. However, the quality of retinal images is often clinically unsatisfactory due to eye lesions and imperfect imaging process. One of the most challenging quality degradation issues in retinal images is non-uniform which hinders the pathological information and further impairs the diagnosis of ophthalmologists and computer-aided analysis. To address this issue, we propose a non-uniform illumination removal network for retinal image, called NuI-Go, which consists of three Recursive Non-local Encoder-Decoder Residual Blocks (NEDRBs) for enhancing the degraded retinal images in a progressive manner. Each NEDRB contains a feature encoder module that captures the hierarchical feature representations, a non-local context module that models the context information, and a feature decoder module that recovers the details and spatial dimension. Additionally, the symmetric skip-connections between the encoder module and the decoder module provide long-range information compensation and reuse. Extensive experiments demonstrate that the proposed method can effectively remove the non-uniform illumination on retinal images while well preserving the image details and color. We further demonstrate the advantages of the proposed method for improving the accuracy of retinal vessel segmentation. Chongyi Li, Huazhu Fu, Runmin Cong, Zechao Li, Qianqian Xu 0001 |
ACM Multimedia | 2 |
| 2020 | Defense for adversarial videos by self-adaptive JPEG compression and optical textureabstractDespite demonstrated outstanding effectiveness in various computer vision tasks, Deep Neural Networks (DNNs) are known to be vulnerable to adversarial examples. Nowadays, adversarial attacks as well as their defenses w.r.t. DNNs in image domain have been intensively studied, and there are some recent works starting to explore adversarial attacks w.r.t. DNNs in video domain. However, the corresponding defense is rarely studied. In this paper, we propose a new two-stage framework for defending video adversarial attack. It contains two main components, namely self-adaptive Joint Photographic Experts Group (JPEG) compression defense and optical texture based defense (OTD). In self-adaptive JPEG compression defense, we propose to adaptively choose an appropriate JPEG quality based on an estimation of moving foreground object, such that the JPEG compression could depress most impact of adversarial noise without losing too much video quality. In OTD, we generate "optical texture" containing high-frequency information based on the optical flow map, and use it to edit Y channel (in YCrCb color space) of input frames, thus further reducing the influence of adversarial perturbation. Experimental results on a benchmark dataset demonstrate the effectiveness of our framework in recovering the classification performance on perturbed videos. Yupeng Cheng, Xingxing Wei 0001, Huazhu Fu, Shangwei Lin 0001, Weisi Lin |
MMAsia | 3 |
| 2020 | Tensorized Multi-view Subspace Representation Learning
Changqing Zhang 0002, Huazhu Fu, Jing Wang 0023, Wen Li 0001, Xiaochun Cao, Qinghua Hu |
Int. J. Comput. Vis. | 2 |
| 2020 | AGE challenge: Angle Closure Glaucoma Evaluation in Anterior Segment Optical Coherence Tomography
Huazhu Fu, Fei Li 0021, Xu Sun 0006, Xingxing Cao, Jingan Liao, José Ignacio Orlando, Xing Tao, Yuexiang Li, Mingkui Tan, Chenglang Yuan, Cheng Bian, Ruitao Xie, Jiongcheng Li, Xiaomeng Li 0001, Jing Wang 0023, Le Geng, Panming Li, Yanwu Xu 0001 |
Medical Image Anal. | 1 |
| 2020 | REFUGE Challenge: A unified framework for evaluating automated methods for glaucoma assessment from fundus photographsabstractGlaucoma is one of the leading causes of irreversible but preventable blindness in working age populations. Color fundus photography (CFP) is the most cost-effective imaging modality to screen for retinal disorders. However, its application to glaucoma has been limited to the computation of a few related biomarkers such as the vertical cup-to-disc ratio. Deep learning approaches, although widely applied for medical image analysis, have not been extensively used for glaucoma assessment due to the limited size of the available data sets. Furthermore, the lack of a standardize benchmark strategy makes difficult to compare existing methods in a uniform way. In order to overcome these issues we set up the Retinal Fundus Glaucoma Challenge, REFUGE (https://refuge.grand-challenge.org), held in conjunction with MICCAI 2018. The challenge consisted of two primary tasks, namely optic disc/cup segmentation and glaucoma classification. As part of REFUGE, we have publicly released a data set of 1200 fundus images with ground truth segmentations and clinical glaucoma labels, currently the largest existing one. We have also built an evaluation framework to ease and ensure fairness in the comparison of different models, encouraging the development of novel techniques in the field. 12 teams qualified and participated in the online challenge. This paper summarizes their methods and analyzes their corresponding results. In particular, we observed that two of the top-ranked teams outperformed two human experts in the glaucoma classification task. Furthermore, the segmentation results were in general consistent with the ground truth annotations, with complementary outcomes that can be further exploited by ensembling the results. José Ignacio Orlando, Huazhu Fu, João Barbosa Breda, Karel van Keer, Deepti R. Bathula, Andres Diaz-Pinto, Ruogu Fang, Pheng-Ann Heng, Jeyoung Kim, Joonseok Lee, Peng Liu 0049, Shuai Lu 0003, Balamurali Murugesan, Valery Naranjo, Sai Samarth R. Phaye, Sharath M. Shankaranarayana, Hrvoje Bogunovic |
Medical Image Anal. | 2 |
| 2020 | Generalized Latent Multi-View Subspace ClusteringabstractSubspace clustering is an effective method that has been successfully applied to many applications. Here, we propose a novel subspace clustering model for multi-view data using a latent representation termed Latent Multi-View Subspace Clustering (LMSC). Unlike most existing single-view subspace clustering methods, which directly reconstruct data points using original features, our method explores underlying complementary information from multiple views and simultaneously seeks the underlying latent representation. Using the complementarity of multiple views, the latent representation depicts data more comprehensively than each individual view, accordingly making subspace representation more accurate and robust. We proposed two LMSC formulations: linear LMSC (lLMSC), based on linear correlations between latent representation and each view, and generalized LMSC (gLMSC), based on neural networks to handle general relationships. The proposed method can be efficiently optimized under the Augmented Lagrangian Multiplier with Alternating Direction Minimization (ALM-ADM) framework. Extensive experiments on diverse datasets demonstrate the effectiveness of the proposed method. Changqing Zhang 0002, Huazhu Fu, Qinghua Hu, Xiaochun Cao, Yuan Xie 0006, Dacheng Tao, Dong Xu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Joint spatial-spectral hyperspectral image classification based on convolutional neural network
Mengxin Han, Runmin Cong, Huazhu Fu, Jianjun Lei 0001 |
Pattern Recognit. Lett. | 4 |
| 2020 | Unsupervised Video Action Clustering via Motion-Scene Interaction ConstraintabstractIn the past few years, scene contextual information has been increasingly used for action understanding with promising results. However, unsupervised video action clustering using context has been less explored, and existing clustering methods cannot achieve satisfactory performances. In this paper, we propose a novel unsupervised video action clustering method by using the motion-scene interaction constraint (MSIC). The proposed method takes the unique static scene and dynamic motion characteristics of video action into account, and develops a contextual interaction constraint model under a self-representation subspace clustering framework. First, the complementarity of multi-view subspace representation in each context is explored by single-view and multi-view constraints. Afterward, the context-constrained affinity matrix is calculated and the MSIC is introduced to mutually regularize the disagreement of subspace representation in scene and motion. Finally, by jointly constraining the complementarity of multi-views and the consistency of multi-contexts, an overall objective function is constructed to guarantee the video action clustering result. The experiments on four video benchmark datasets (Weizmann, KTH, UCFsports, and Olympic) demonstrate that the proposed method outperforms the state-of-the-art methods. Bo Peng 0007, Jianjun Lei 0001, Huazhu Fu, Changqing Zhang 0002, Tat-Seng Chua, Xuelong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Going From RGB to RGBD Saliency: A Depth-Guided Transformation ModelabstractDepth information has been demonstrated to be useful for saliency detection. However, the existing methods for RGBD saliency detection mainly focus on designing straightforward and comprehensive models, while ignoring the transferable ability of the existing RGB saliency detection models. In this article, we propose a novel depth-guided transformation model (DTM) going from RGB saliency to RGBD saliency. The proposed model includes three components, that is: 1) multilevel RGBD saliency initialization; 2) depth-guided saliency refinement; and 3) saliency optimization with depth constraints. The explicit depth feature is first utilized in the multilevel RGBD saliency model to initialize the RGBD saliency by combining the global compactness saliency cue and local geodesic saliency cue. The depth-guided saliency refinement is used to further highlight the salient objects and suppress the background regions by introducing the prior depth domain knowledge and prior refined depth shape. Benefiting from the consistency of the entire object in the depth map, we formulate an optimization model to attain more consistent and accurate saliency results via an energy function, which integrates the unary data term, color smooth term, and depth consistency term. Experiments on three public RGBD saliency detection benchmarks demonstrate the effectiveness and performance improvement of the proposed DTM from RGB to RGBD saliency. Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Junhui Hou, Qingming Huang, Sam Kwong |
IEEE Trans. Cybern. | 3 |
| 2020 | Angle-Closure Detection in Anterior Segment OCT Based on Multilevel Deep NetworkabstractIrreversible visual impairment is often caused by primary angle-closure glaucoma, which could be detected via anterior segment optical coherence tomography (AS-OCT). In this paper, an automated system based on deep learning is presented for angle-closure detection in AS-OCT images. Our system learns a discriminative representation from training data that captures subtle visual cues not modeled by handcrafted features. A multilevel deep network is proposed to formulate this learning, which utilizes three particular AS-OCT regions based on clinical priors: 1) the global anterior segment structure; 2) local iris region; and 3) anterior chamber angle (ACA) patch. In our method, a sliding window-based detector is designed to localize the ACA region, which addresses ACA detection as a regression task. Then, three parallel subnetworks are applied to extract AS-OCT representations for the global image and at clinically relevant local regions. Finally, the extracted deep features of these subnetworks are concatenated into one fully connected layer to predict the angle-closure detection result. In the experiments, our system is shown to surpass previous detection methods and other deep learning systems on two clinical AS-OCT datasets. Huazhu Fu, Yanwu Xu 0001, Stephen Lin 0001, Damon Wing Kee Wong, Mani Baskaran, Meenakshi Mahesh, Tin Aung, Jiang Liu 0001 |
IEEE Trans. Cybern. | 1 |
| 2020 | Hybrid Noise-Oriented Multilabel LearningabstractFor real-world applications, multilabel learning usually suffers from unsatisfactory training data. Typically, features may be corrupted or class labels may be noisy or both. Ignoring noise in the learning process tends to result in an unreasonable model and, thus, inaccurate prediction. Most existing methods only consider either feature noise or label noise in multilabel learning. In this paper, we propose a unified robust multilabel learning framework for data with hybrid noise, that is, both feature noise and label noise. The proposed method, hybrid noise-oriented multilabel learning (HNOML), is simple but rather robust for noisy data. HNOML simultaneously addresses feature and label noise by bi-sparsity regularization bridged with label enrichment. Specifically, the label enrichment matrix explores the underlying correlation among different classes which improves the noisy labeling. Bridged with the enriching label matrix, the structured sparsity is imposed to jointly handle the corrupted features and noisy labeling. We utilize the alternating direction method (ADM) to efficiently solve our problem. Experimental results on several benchmark datasets demonstrate the advantages of our method over the state-of-the-art ones. Changqing Zhang 0002, Ziwei Yu, Huazhu Fu, Pengfei Zhu 0001, Lei Chen 0011, Qinghua Hu |
IEEE Trans. Cybern. | 3 |
| 2020 | A Recursive Constrained Framework for Unsupervised Video Action ClusteringabstractVideo action understanding is an active field of intelligent video analytics, and contextual information in the videos has gained lots of attention for better action understanding. However, most existing works focus on using contextual information for supervised or semi-supervised analysis, and how to effectively use contextual information to boost the unsupervised action clustering performance is still a challenging problem. In this article, we propose a recursive constrained framework for unsupervised video action clustering by utilizing the contextual information of the action and scene. Considering the unique contextual characteristics of video action, action context clustering solution and scene context clustering solution are obtained simultaneously. Based on these two solutions, a recursive priori propagation is proposed to exploit information gain of the priori clustering solutions, and then the information gain is fed back into the procedures of both subspace representation and spectral clustering. Specifically, to explore the unknown relationships in the priori clustering solutions, the constraint-guided subspace representation is introduced by fusing the recursive priori constraint into the self-representation model. Taking priori information and multiview features into consideration, the priori-inherited multiview spectral clustering is proposed to obtain more discriminative spectral embeddings for action clustering. Experiments on three video benchmark datasets demonstrate that the proposed method outperforms state-of-the-art methods. Bo Peng 0007, Jianjun Lei 0001, Huazhu Fu, Ling Shao 0001, Qingming Huang |
IEEE Trans. Ind. Informatics | 3 |
| 2020 | Text Co-Detection in Multi-View SceneabstractMulti-view scene analysis has been widely explored in computer vision, including numerous practical applications. The texts in multi-view scenes are often detected by following the existing text detection method in a single image, which however ignores the multi-view corresponding constraint. The multi-view correspondences may contain structure, location information and assist difficulties induced by factors like occlusion and perspective distortion, which are deficient in the single image scene. In this paper, we address the corresponding text detection task and propose a novel text co-detection method to identify the cooccurring texts among multi-view scene images with compositions of detection and correspondence under large environmental variations. In our text co-detection method, the visual and geometrical correspondences are designed to explore texts holding high pairwise representation similarity and guide the exploitation of texts with geometrical correspondences, simultaneously. To guarantee the pairwise consistency among multiple images, we additionally incorporate the cycle consistency constraint, which guarantees alignments of text correspondences in the image set. Finally, text correspondence is represented by a permutation matrix and solved via positive semidefinite and low-rank constraints. Moreover, we also collect a new text co-detection dataset consisting of multi-view image groups obtained from the same scene with different photographing conditions. The experiments show that our text co-detection obtains satisfactory performance and outperforms the related state-of-the-art text detection methods. Chuan Wang 0002, Huazhu Fu, Liang Yang 0002, Xiaochun Cao |
IEEE Trans. Image Process. | 2 |
| 2020 | Guest Editorial Ophthalmic Image Analysis and InformaticsabstractThe papers in this special son focus on This special issue invited contributions reporting on methodological breakthroughs in artificial intelligence in ophthalmology, and systems and insights that make use of large-scale datasets linking across multiple imaging modalities, image phenotyping, and imaging omics. These papers presented in this special issue introduce the latest advances in the field of ophthalmic image analysis and informatics, which enable and drive the research, development, and application of key technologies into ocular healthcare. Jun Cheng 0003, Huazhu Fu, Delia Cabrera DeBuc, Jie Tian 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | M$^3$Lung-Sys: A Deep Learning System for Multi-Class Lung Pneumonia Screening From CT ImagingabstractTo counter the outbreak of COVID-19, the accurate diagnosis of suspected cases plays a crucial role in timely quarantine, medical treatment, and preventing the spread of the pandemic. Considering the limited training cases and resources (e.g, time and budget), we propose a Multi-task Multi-slice Deep Learning System (M3Lung-Sys) for multi-class lung pneumonia screening from CT imaging, which only consists of two 2D CNN networks, i.e., slice- and patient-level classification networks. The former aims to seek the feature representations from abundant CT slices instead of limited CT volumes, and for the overall pneumonia screening, the latter one could recover the temporal information by feature refinement and aggregation between different slices. In addition to distinguish COVID-19 from Healthy, H1N1, and CAP cases, our M3Lung-Sys also be able to locate the areas of relevant lesions, without any pixel-level annotation. To further demonstrate the effectiveness of our model, we conduct extensive experiments on a chest CT imaging dataset with a total of 734 patients (251 healthy people, 245 COVID-19 patients, 105 H1N1 patients, and 133 CAP patients). The quantitative results with plenty of metrics indicate the superiority of our proposed model on both slice- and patient-level classification tasks. More importantly, the generated lesion location maps make our system interpretable and more valuable to clinicians. Xuelin Qian, Huazhu Fu, Weiya Shi, Tao Chen 0003, Yanwei Fu 0001, Xiangyang Xue 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2020 | Inf-Net: Automatic COVID-19 Lung Infection Segmentation From CT ImagesabstractCoronavirus Disease 2019 (COVID-19) spread globally in early 2020, causing the world to face an existential health crisis. Automated detection of lung infections from computed tomography (CT) images offers a great potential to augment the traditional healthcare strategy for tackling COVID-19. However, segmenting infected regions from CT slices faces several challenges, including high variation in infection characteristics, and low intensity contrast between infections and normal tissues. Further, collecting a large amount of data is impractical within a short time period, inhibiting the training of a deep model. To address these challenges, a novel COVID-19 Lung Infection Segmentation Deep Network (Inf-Net) is proposed to automatically identify infected regions from chest CT slices. In our Inf-Net, a parallel partial decoder is used to aggregate the high-level features and generate a global map. Then, the implicit reverse attention and explicit edge-attention are utilized to model the boundaries and enhance the representations. Moreover, to alleviate the shortage of labeled data, we present a semi-supervised segmentation framework based on a randomly selected propagation strategy, which only requires a few labeled images and leverages primarily unlabeled data. Our semi-supervised framework can improve the learning ability and achieve a higher performance. Extensive experiments on our COVID-SemiSeg and real CT volumes demonstrate that the proposed Inf-Net outperforms most cutting-edge segmentation models and advances the state-of-the-art performance. Deng-Ping Fan, Tao Zhou 0002, Ge-Peng Ji, Yi Zhou 0007, Geng Chen 0001, Huazhu Fu, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2020 | Correction to "Noise Adaptation Generative Adversarial Network for Medical Image Analysis"abstractIn the above article[1],Tables II,III, andVandFig. 6are incorrect. The correct images are provided below: Tianyang Miller, Jun Cheng 0003, Huazhu Fu, Zaiwang Gu, Kang Zhou 0001, Shenghua Gao, Ru Zheng, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Noise Adaptation Generative Adversarial Network for Medical Image AnalysisabstractMachine learning has been widely used in medical image analysis under an assumption that the training and test data are under the same feature distributions. However, medical images from difference devices or the same device with different parameter settings are often contaminated with different amount and types of noises, which violate the above assumption. Therefore, the models trained using data from one device or setting often fail to work for that from another. Moreover, it is very expensive and tedious to label data and re-train models for all different devices or settings. To overcome this noise adaptation issue, it is necessary to leverage on the models trained with data from one device or setting for new data. In this paper, we reformulate this noise adaptation task as an image-to-image translation task such that the noise patterns from the test data are modified to be similar to those from the training data while the contents of the data are unchanged. In this paper, we propose a novel Noise Adaptation Generative Adversarial Network (NAGAN), which contains a generator and two discriminators. The generator aims to map the data from source domain to target domain. Among the two discriminators, one discriminator enforces the generated images to have the same noise patterns as those from the target domain, and the second discriminator enforces the content to be preserved in the generated images. We apply the proposed NAGAN on both optical coherence tomography (OCT) images and ultrasound images. Results show that the method is able to translate the noise style. In addition, we also evaluate our proposed method with segmentation task in OCT and classification task in ultrasound. The experimental results show that the proposed NAGAN improves the analysis outcome. Tianyang Miller, Jun Cheng 0003, Huazhu Fu, Zaiwang Gu, Kang Zhou 0001, Shenghua Gao, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Hi-Net: Hybrid-Fusion Network for Multi-Modal MR Image SynthesisabstractMagnetic resonance imaging (MRI) is a widely used neuroimaging technique that can provide images of different contrasts (i.e., modalities). Fusing this multi-modal data has proven particularly effective for boosting model performance in many tasks. However, due to poor data quality and frequent patient dropout, collecting all modalities for every patient remains a challenge. Medical image synthesis has been proposed as an effective solution, where any missing modalities are synthesized from the existing ones. In this paper, we propose a novel Hybrid-fusion Network (Hi-Net) for multi-modal MR image synthesis, which learns a mapping from multi-modal source images (i.e., existing modalities) to target images (i.e., missing modalities). In our Hi-Net, a modality-specific network is utilized to learn representations for each individual modality, and a fusion network is employed to learn the common latent representation of multi-modal data. Then, a multi-modal synthesis network is designed to densely combine the latent representation with hierarchical features from each modality, acting as a generator to synthesize the target images. Moreover, a layer-wise multi-modal fusion strategy effectively exploits the correlations among multiple modalities, where a Mixed Fusion Block (MFB) is proposed to adaptively weight different fusion strategies. Extensive experiments demonstrate the proposed model outperforms other state-of-the-art medical image synthesis methods. Tao Zhou 0002, Huazhu Fu, Geng Chen 0001, Jianbing Shen, Ling Shao 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2020 | PDR-Net: Perception-Inspired Single Image Dehazing Network With RefinementabstractDuring recent years, we have witnessed a rapid development of wireless network technologies which have revolutionized the way people take and share multimedia content. However, images captured in the outdoor scenes usually suffer from limited visibility due to suspended atmospheric particles, which directly affects the quality of photos. Despite the recent progress of image dehazing methods, the visual quality of dehazed results still needs further improvement. In this paper, we propose a deep convolutional neural network (CNN) for single image dehazing called PDR-Net, which includes a perception-inspired haze removal subnetwork that reconstructs the latent dehazed image and a refinement subnetwork that further enhances the contrast and color properties of the dehazed result by joint multi-term loss optimization. Compared to the previous methods, our method combines the advantages of existing indoor and outdoor image dehazing training data, which makes the proposed PDR-Net generalized to various hazy images and effective for improving the visual quality of the dehazed results. Extensive experiments demonstrate that the proposed method achieves comparable and even better performance on both real and synthetic images in qualitative and quantitative metrics. Additionally, the potential usage of our method in high-level vision tasks is discussed. Chongyi Li, Chunle Guo, Jichang Guo, Ping Han, Huazhu Fu, Runmin Cong |
IEEE Trans. Multim. | 5 |
| 2019 | AE2-Nets: Autoencoder in Autoencoder NetworksabstractLearning on data represented with multiple views (e.g., multiple types of descriptors or modalities) is a rapidly growing direction in machine learning and computer vision. Although effectiveness achieved, most existing algorithms usually focus on classification or clustering tasks. Differently, in this paper, we focus on unsupervised representation learning and propose a novel framework termed Autoencoder in Autoencoder Networks (AE2-Nets), which integrates information from heterogeneous sources into an intact representation by the nested autoencoder framework. The proposed method has the following merits: (1) our model jointly performs view-specific representation learning (with the inner autoencoder networks) and multi-view information encoding (with the outer autoencoder networks) in a unified framework; (2) due to the degradation process from the latent representation to each single view, our model flexibly balances the complementarity and consistence among multiple views. The proposed model is efficiently solved by the alternating direction method (ADM), and demonstrates the effectiveness compared with state-of-the-art algorithms. Changqing Zhang 0002, Yeqing Liu, Huazhu Fu |
CVPR | 3 |
| 2019 | A Deep Step Pattern Representation for Multimodal Retinal Image RegistrationabstractThis paper presents a novel feature-based method that is built upon a convolutional neural network (CNN) to learn the deep representation for multimodal retinal image registration. We coined the algorithm deep step patterns, in short DeepSPa. Most existing deep learning based methods require a set of manually labeled training data with known corresponding spatial transformations, which limits the size of training datasets. By contrast, our method is fully automatic and scale well to different image modalities with no human intervention. We generate feature classes from simple step patterns within patches of connecting edges formed by vascular junctions in multiple retinal imaging modalities. We leverage CNN to learn and optimize the input patches to be used for image registration. Spatial transformations are estimated based on the output possibility of the fully connected layer of CNN for a pair of images. One of the key advantages of the proposed algorithm is its robustness to non-linear intensity changes, which widely exist on retinal images due to the difference of acquisition modalities. We validate our algorithm on extensive challenging datasets comprising poor quality multimodal retinal images which are adversely affected by pathologies (diseases), speckle noise and low resolutions. The experimental results demonstrate the robustness and accuracy over state-of-the-art multimodal image registration algorithms. Jimmy Addison Lee, Peng Liu 0049, Jun Cheng 0003, Huazhu Fu |
ICCV | 4 |
| 2019 | Reciprocal Multi-Layer Subspace Learning for Multi-View ClusteringabstractMulti-view clustering is a long-standing important research topic, however, remains challenging when handling high-dimensional data and simultaneously exploring the consistency and complementarity of different views. In this work, we present a novel Reciprocal Multi-layer Subspace Learning (RMSL) algorithm for multi-view clustering, which is composed of two main components: Hierarchical Self-Representative Layers (HSRL), and Backward Encoding Networks (BEN). Specifically, HSRL constructs reciprocal multi-layer subspace representations linked with a latent representation to hierarchically recover the underlying low-dimensional subspaces in which the high-dimensional data lie; BEN explores complex relationships among different views and implicitly enforces the subspaces of all views to be consistent with each other and more separable. The latent representation flexibly encodes complementary information from multiple views and depicts data more comprehensively. Our model can be efficiently optimized by an alternating optimization scheme. Extensive experiments on benchmark datasets show the superiority of RMSL over other state-of-the-art clustering methods. Ruihuang Li, Changqing Zhang 0002, Huazhu Fu, Xi Peng 0001, Joey Tianyi Zhou, Qinghua Hu |
ICCV | 3 |
| 2019 | Inter-modality Dependence Induced Data Recovery for MCI Conversion Prediction
Tao Zhou 0002, Kim-Han Thung, Yu Zhang 0009, Huazhu Fu, Jianbing Shen, Dinggang Shen, Ling Shao 0001 |
MICCAI (4) | 4 |
| 2019 | Spatiotemporal Breast Mass Detection Network (MD-Net) in 4D DCE-MRI Images
Lixi Deng, Sheng Tang, Huazhu Fu, Bin Wang 0065, Yongdong Zhang 0001 |
MICCAI (4) | 3 |
| 2019 | Evaluation of Retinal Image Quality Assessment Networks in Different Color-Spaces
Huazhu Fu, Jianbing Shen, Shanshan Cui, Yanwu Xu 0001, Jiang Liu 0001, Ling Shao 0001 |
MICCAI (1) | 1 |
| 2019 | ET-Net: A Generic Edge-aTtention Guidance Network for Medical Image Segmentation
Huazhu Fu, Hang Dai, Jianbing Shen, Yanwei Pang, Ling Shao 0001 |
MICCAI (1) | 2 |
| 2019 | Attention Guided Network for Retinal Image Segmentation
Huazhu Fu, Yuguang Yan, Yubing Zhang, Qingyao Wu, Ming Yang 0039, Mingkui Tan, Yanwu Xu 0001 |
MICCAI (1) | 2 |
| 2019 | SkrGAN: Sketching-Rendering Unconditional Generative Adversarial Networks for Medical Image Synthesis
Tianyang Miller, Huazhu Fu, Yitian Zhao, Jun Cheng 0003, Mengjie Guo, Zaiwang Gu, Shenghua Gao, Jiang Liu 0001 |
MICCAI (4) | 2 |
| 2019 | Deep Multi-modal Latent Representation Learning for Automated Dementia Diagnosis
Tao Zhou 0002, Mingxia Liu 0001, Huazhu Fu, Jun Wang 0024, Jianbing Shen, Ling Shao 0001, Dinggang Shen |
MICCAI (4) | 3 |
| 2019 | CPM-Nets: Cross Partial Multi-View NetworksabstractDespite multi-view learning progressed fast in past decades, it is still challenging due to the difficulty in modeling complex correlation among different views, especially under the context of view missing. To address the challenge, we propose a novel framework termed Cross Partial Multi-View Networks (CPM-Nets). In this framework, we first give a formal definition of completeness and versatility for multi-view representation and then theoretically prove the versatility of the latent representation learned from our algorithm. To achieve the completeness, the task of learning latent multi-view representation is specifically translated to degradation process through mimicking data transmitting, such that the optimal tradeoff between consistence and complementarity across different views could be achieved. In contrast with methods that either complete missing views or group samples according to view-missing patterns, our model fully exploits all samples and all views to produce structured representation for interpretability. Extensive experimental results validate the effectiveness of our algorithm over existing state-of-the-arts. Changqing Zhang 0002, Zongbo Han, Yajie Cui, Huazhu Fu, Joey Tianyi Zhou, Qinghua Hu |
NeurIPS | 4 |
| 2019 | An efficient privacy protection scheme for data security in video surveillance
Wei Zhang 0031, Huazhu Fu, Wenqi Ren, Xinpeng Zhang 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Review of Visual Saliency Detection With Comprehensive InformationabstractThe visual saliency detection model simulates the human visual system to perceive the scene and has been widely used in many vision tasks. With the development of acquisition technology, more comprehensive information, such as depth cue, inter-image correspondence, or temporal relationship, is available to extend image saliency detection to RGBD saliency detection, co-saliency detection, or video saliency detection. The RGBD saliency detection model focuses on extracting the salient regions from RGBD images by combining the depth information. The co-saliency detection model introduces the inter-image correspondence constraint to discover the common salient object in an image group. The goal of the video saliency detection model is to locate the motion-related salient object in video sequences, which considers the motion cue and spatiotemporal constraint jointly. In this paper, we review different types of saliency detection algorithms, summarize the important issues of the existing methods, and discuss the existent problems and future works. Moreover, the evaluation datasets and quantitative measurements are briefly introduced, and the experimental analysis and discussion are conducted to provide a holistic overview of different saliency detection methods. Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Ming-Ming Cheng, Weisi Lin, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Person Re-Identification by Semantic Region Representation and Topology ConstraintabstractPerson re-identification is a popular research topic which aims at matching the specific person in a multi-camera network automatically. Feature representation and metric learning are two important issues for person re-identification. In this paper, we propose a novel person re-identification method, which consists of a reliable representation called semantic region representation (SRR), and an effective metric learning with mapping space topology constraint (MSTC). The SRR integrates semantic representations to achieve effective similarity comparison between the corresponding regions via parsing the body into multiple parts, which focuses on the foreground context against the background interference. To learn a discriminant metric, the MSTC is proposed to consider the topological relationship among all samples in the feature space. It considers two-fold constraints: the distribution of positive pairs should be more compact than the average distribution of negative pairs with regard to the same probe, while the average distance between different classes should be larger than that between same classes. These two aspects cooperate to maintain the compactness of the intra-class as well as the sparsity of the inter-class. Extensive experiments conducted on five challenging person re-identification datasets, VIPeR, SYSU-sReID, QUML GRID, CUHK03, and Market-1501, show that the proposed method achieves competitive performance with the state-of-the-art approaches. Jianjun Lei 0001, Lijie Niu, Huazhu Fu, Bo Peng 0007, Qingming Huang, Chunping Hou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | An Iterative Co-Saliency Framework for RGBD ImagesabstractAs a newly emerging and significant topic in computer vision community, co-saliency detection aims at discovering the common salient objects in multiple related images. The existing methods often generate the co-saliency map through a direct forward pipeline which is based on the designed cues or initialization, but lack the refinement-cycle scheme. Moreover, they mainly focus on RGB image and ignore the depth information for RGBD images. In this paper, we propose an iterative RGBD co-saliency framework, which utilizes the existing single saliency maps as the initialization, and generates the final RGBD co-saliency map by using a refinement-cycle model. Three schemes are employed in the proposed RGBD co-saliency framework, which include the addition scheme, deletion scheme, and iteration scheme. The addition scheme is used to highlight the salient regions based on intra-image depth propagation and saliency propagation, while the deletion scheme filters the saliency regions and removes the non-common salient regions based on interimage constraint. The iteration scheme is proposed to obtain more homogeneous and consistent co-saliency map. Furthermore, a novel descriptor, named depth shape prior, is proposed in the addition scheme to introduce the depth information to enhance identification of co-salient objects. The proposed method can effectively exploit any existing 2-D saliency model to work well in RGBD co-saliency scenarios. The experiments on two RGBD co-saliency datasets demonstrate the effectiveness of our proposed framework. Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Weisi Lin, Qingming Huang, Xiaochun Cao, Chunping Hou |
IEEE Trans. Cybern. | 3 |
| 2019 | Video Saliency Detection via Sparsity-Based Reconstruction and PropagationabstractVideo saliency detection aims to continuously discover the motion-related salient objects from the video sequences. Since it needs to consider the spatial and temporal constraints jointly, video saliency detection is more challenging than image saliency detection. In this paper, we propose a new method to detect the salient objects in video based on sparse reconstruction and propagation. With the assistance of novel static and motion priors, a single-frame saliency model is first designed to represent the spatial saliency in each individual frame via the sparsity-based reconstruction. Then, through a progressive sparsity-based propagation, the sequential correspondence in the temporal space is captured to produce the inter-frame saliency map. Finally, these two maps are incorporated into a global optimization model to achieve spatio-temporal smoothness and global consistency of the salient object in the whole video. The experiments on three large-scale video saliency datasets demonstrate that the proposed method outperforms the state-of-the-art algorithms both qualitatively and quantitatively. Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Fatih Porikli, Qingming Huang, Chunping Hou |
IEEE Trans. Image Process. | 3 |
| 2019 | Hierarchical Features Driven Residual Learning for Depth Map Super-ResolutionabstractRapid development of affordable and portable consumer depth cameras facilitates the use of depth information in many computer vision tasks such as intelligent vehicles and 3D reconstruction. However, depth map captured by low-cost depth sensors (e.g., Kinect) usually suffers from low spatial resolution, which limits its potential applications. In this paper, we propose a novel deep network for depth map super-resolution (SR), called DepthSR-Net. The proposed DepthSR-Net automatically infers a high resolution (HR) depth map from its low resolution (LR) version by hierarchical features driven residual learning. Specifically, DepthSR-Net is built on a residual U-Net deep network architecture. Given LR depth map, we first obtain the desired HR by bicubic interpolation upsampling, and then construct an input pyramid to achieve multiple level receptive fields. Next, we extract hierarchical features from the input pyramid, intensity image, and encoder-decoder structure of UNet. Finally, we learn the residual between the interpolated depth map and the corresponding HR one using the rich hierarchical features. The final HR depth map is achieved by adding the learned residual to the interpolated depth map. We conduct an ablation study to demonstrate the effectiveness of each component in the proposed network. Extensive experiments demonstrate that the proposed method outperforms the state-of-the-art methods. Additionally, the potential usage of the proposed network in other low-level vision problems is discussed. Chunle Guo, Chongyi Li, Jichang Guo, Runmin Cong, Huazhu Fu, Ping Han |
IEEE Trans. Image Process. | 5 |
| 2019 | Multi-View Saliency-Guided Clustering for Image CosegmentationabstractImage cosegmentation aims at extracting the common objects from multiple images simultaneously. Existing methods mainly solve cosegmentation via the pre-defined graph, which lacks flexibility and robustness to handle various visual patterns. Besides, similar backgrounds also confuse the identifying of the common foreground. To address these issues, we propose a novel Multi-view Saliency-Guided Clustering algorithm (MvSGC) for the image cosegmentation task. In our model, the unsupervised saliency prior is used as partition-level side information to guide the foreground clustering process. To achieve robustness to noises and missing observations, similarities on instance-level and partition-level are both considered. Specifically, a unified clustering model with cosine similarity is proposed to capture the intrinsic structure of data and keep partition result consistent with the side information. Moreover, we leverage multi-view weight learning to integrate multiple feature representations to further improve the robustness of our approach. A K-means-like optimization algorithm is developed to proceed the constrained clustering in a highly efficient way with theoretical support. Experimental results on three benchmark datasets (i.e., the iCoseg, MSRC and Internet image dataset) and one RGB-D image dataset demonstrate the superiority of applying our clustering method for image cosegmentation. Zhiqiang Tao, Hongfu Liu 0001, Huazhu Fu, Yun Fu 0001 |
IEEE Trans. Image Process. | 3 |
| 2019 | CE-Net: Context Encoder Network for 2D Medical Image SegmentationabstractMedical image segmentation is an important step in medical image analysis. With the rapid development of a convolutional neural network in image processing, deep learning has been used for medical image segmentation, such as optic disc segmentation, blood vessel detection, lung segmentation, cell segmentation, and so on. Previously, U-net based approaches have been proposed. However, the consecutive pooling and strided convolutional operations led to the loss of some spatial information. In this paper, we propose a context encoder network (CE-Net) to capture more high-level information and preserve spatial information for 2D medical image segmentation. CE-Net mainly contains three major components: a feature encoder module, a context extractor, and a feature decoder module. We use the pretrained ResNet block as the fixed feature extractor. The context extractor module is formed by a newly proposed dense atrous convolution block and a residual multi-kernel pooling block. We applied the proposed CE-Net to different 2D medical image segmentation tasks. Comprehensive results show that the proposed method outperforms the original U-Net method and other state-of-the-art methods for optic disc segmentation, vessel detection, lung segmentation, cell contour segmentation, and retinal optical coherence tomography layer segmentation. Zaiwang Gu, Jun Cheng 0003, Huazhu Fu, Kang Zhou 0001, Huaying Hao, Yitian Zhao, Tianyang Miller, Shenghua Gao, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2019 | HSCS: Hierarchical Sparsity Based Co-saliency Detection for RGBD ImagesabstractCo-saliency detection aims to discover common and salient objects in an image group containing more than two relevant images. Moreover, depth information has been demonstrated to be effective for many computer vision tasks. In this paper, we propose a novel co-saliency detection method for RGBD images based on hierarchical sparsity reconstruction and energy function refinement. With the assistance of the intrasaliency map, the inter-image correspondence is formulated as a hierarchical sparsity reconstruction framework. The global sparsity reconstruction model with a ranking scheme focuses on capturing the global characteristics among the whole image group through a common foreground dictionary. The pairwise sparsity reconstruction model aims to explore the corresponding relationship between pairwise images through a set of pairwise dictionaries. In order to improve the intra-image smoothness and inter-image consistency, an energy function refinement model is proposed, which includes the unary data term, spatial smooth term, and holistic consistency term. Experiments on two RGBD co-saliency detection benchmarks demonstrate that the proposed method outperforms the state-of-the-art algorithms both qualitatively and quantitatively. Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Qingming Huang, Xiaochun Cao, Nam Ling |
IEEE Trans. Multim. | 3 |
| 2018 | DeepAMD: Detect Early Age-Related Macular Degeneration by Applying Deep Learning in a Multiple Instance Learning Framework
Damon Wing Kee Wong, Huazhu Fu, Yanwu Xu 0001, Jiang Liu 0001 |
ACCV (5) | 3 |
| 2018 | 3-in-1 Correlated Embedding via Adaptive Exploration of the Structure and Semantic SubspacesabstractCombinational network embedding, which learns the node representation by exploring both topological and non-topological information, becomes popular due to the fact that the two types of information are complementing each other. Most of the existing methods either consider the topological and non-topological information being aligned or possess predetermined preferences during the embedding process.Unfortunately, previous methods fail to either explicitly describe the correlations between topological and non-topological information or adaptively weight their impacts. To address the existing issues, three new assumptions are proposed to better describe the embedding space and its properties. With the proposed assumptions, nodes, communities and topics are mapped into one embedding space. A novel generative model is proposed to formulate the generation process of the network and content from the embeddings, with respect to the Bayesian framework. The proposed model automatically leans to the information which is more discriminative.The embedding result can be obtained by maximizing the posterior distribution by adopting the variational inference and reparameterization trick. Experimental results indicate that the proposed method gives superior performances compared to the state-of-the-art methods when a variety of real-world networks is analyzed. Liang Yang 0002, Yuanfang Guo, Di Jin 0001, Huazhu Fu, Xiaochun Cao |
IJCAI | 4 |
| 2018 | Multi-context Deep Network for Angle-Closure Glaucoma Screening in Anterior Segment OCT
Huazhu Fu, Yanwu Xu 0001, Stephen Lin 0001, Damon Wing Kee Wong, Mani Baskaran, Meenakshi Mahesh, Tin Aung, Jiang Liu 0001 |
MICCAI (2) | 1 |
| 2018 | Co-Saliency Detection for RGBD Images Based on Multi-Constraint Feature Matching and Cross Label PropagationabstractCo-saliency detection aims at extracting the common salient regions from an image group containing two or more relevant images. It is a newly emerging topic in computer vision community. Different from the most existing co-saliency methods focusing on RGB images, this paper proposes a novel co-saliency detection model for RGBD images, which utilizes the depth information to enhance identification of co-saliency. First, the intra saliency map for each image is generated by the single image saliency model, while the inter saliency map is calculated based on the multi-constraint feature matching, which represents the constraint relationship among multiple images. Then, the optimization scheme, namely cross label propagation, is used to refine the intra and inter saliency maps in a cross way. Finally, all the original and optimized saliency maps are integrated to generate the final co-saliency result. The proposed method introduces the depth information and multi-constraint feature matching to improve the performance of co-saliency detection. Moreover, the proposed method can effectively exploit any existing single image saliency model to work well in co-saliency scenarios. Experiments on two RGBD co-saliency datasets demonstrate the effectiveness of our proposed model. Runmin Cong, Jianjun Lei 0001, Huazhu Fu, Qingming Huang, Xiaochun Cao, Chunping Hou |
IEEE Trans. Image Process. | 3 |
| 2018 | A Review of Co-Saliency Detection Algorithms: Fundamentals, Applications, and ChallengesabstractCo-saliency detection is a newly emerging and rapidly growing research area in the computer vision community. As a novel branch of visual saliency, co-saliency detection refers to the discovery of common and salient foregrounds from two or more relevant images, and it can be widely used in many computer vision tasks. The existing co-saliency detection algorithms mainly consist of three components: extracting effective features to represent the image regions, exploring the informative cues or factors to characterize co-saliency, and designing effective computational frameworks to formulate co-saliency. Although numerous methods have been developed, the literature is still lacking a deep review and evaluation of co-saliency detection techniques. In this article, we aim at providing a comprehensive review of the fundamentals, challenges, and applications of co-saliency detection. Specifically, we provide an overview of some related computer vision works, review the history of co-saliency detection, summarize and categorize the major algorithms in this research area, discuss some open issues in this area, present the potential applications of co-saliency detection, and finally point out some unsolved challenges and promising future works. We expect this review to be beneficial to both fresh and senior researchers in this field and to give insights to researchers in other related areas regarding the utility of co-saliency detection algorithms. Dingwen Zhang, Huazhu Fu, Junwei Han 0001, Ali Borji, Xuelong Li 0001 |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2018 | Structure-Preserving Guided Retinal Image Filtering and Its Application for Optic Disk AnalysisabstractRetinal fundus photographs have been used in the diagnosis of many ocular diseases such as glaucoma, pathological myopia, age-related macular degeneration, and diabetic retinopathy. With the development of computer science, computer aided diagnosis has been developed to process and analyze the retinal images automatically. One of the challenges in the analysis is that the quality of the retinal image is often degraded. For example, a cataract in human lens will attenuate the retinal image, just as a cloudy camera lens which reduces the quality of a photograph. It often obscures the details in the retinal images and posts challenges in retinal image processing and analyzing tasks. In this paper, we approximate the degradation of the retinal images as a combination of human-lens attenuation and scattering. A novel structure-preserving guided retinal image filtering (SGRIF) is then proposed to restore images based on the attenuation and scattering model. The proposed SGRIF consists of a step of global structure transferring and a step of global edge-preserving smoothing. Our results show that the proposed SGRIF method is able to improve the contrast of retinal images, measured by histogram flatness measure, histogram spread, and variability of local luminosity. In addition, we further explored the benefits of SGRIF for subsequent retinal image processing and analyzing tasks. In the two applications of deep learning-based optic cup segmentation and sparse learning-based cup-to-disk ratio (CDR) computation, our results show that we are able to achieve more accurate optic cup segmentation and CDR measurements from images processed by SGRIF. Jun Cheng 0003, Zhengguo Li, Zaiwang Gu, Huazhu Fu, Damon Wing Kee Wong, Jiang Liu 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2018 | Joint Optic Disc and Cup Segmentation Based on Multi-Label Deep Network and Polar TransformationabstractGlaucoma is a chronic eye disease that leads to irreversible vision loss. The cup to disc ratio (CDR) plays an important role in the screening and diagnosis of glaucoma. Thus, the accurate and automatic segmentation of optic disc (OD) and optic cup (OC) from fundus images is a fundamental task. Most existing methods segment them separately, and rely on hand-crafted visual feature from fundus images. In this paper, we propose a deep learning architecture, named M-Net, which solves the OD and OC segmentation jointly in a one-stage multi-label system. The proposed M-Net mainly consists of multi-scale input layer, U-shape convolutional network, side-output layer, and multi-label loss function. The multi-scale input layer constructs an image pyramid to achieve multiple level receptive field sizes. The U-shape convolutional network is employed as the main body network structure to learn the rich hierarchical representation, while the side-output layer acts as an early classifier that produces a companion local prediction map for different scale layers. Finally, a multi-label loss function is proposed to generate the final segmentation map. For improving the segmentation performance further, we also introduce the polar transformation, which provides the representation of the original image in the polar coordinate system. The experiments show that our M-Net system achieves state-of-the-art OD and OC segmentation result on ORIGA data set. Simultaneously, the proposed method also obtains the satisfactory glaucoma screening performances with calculated CDR value on both ORIGA and SCES datasets. Huazhu Fu, Jun Cheng 0003, Yanwu Xu 0001, Damon Wing Kee Wong, Jiang Liu 0001, Xiaochun Cao |
IEEE Trans. Medical Imaging | 1 |
| 2018 | Disc-Aware Ensemble Network for Glaucoma Screening From Fundus ImageabstractGlaucoma is a chronic eye disease that leads to irreversible vision loss. Most of the existing automatic screening methods first segment the main structure and subsequently calculate the clinical measurement for the detection and screening of glaucoma. However, these measurement-based methods rely heavily on the segmentation accuracy and ignore various visual features. In this paper, we introduce a deep learning technique to gain additional image-relevant information and screen glaucoma from the fundus image directly. Specifically, a novel disc-aware ensemble network for automatic glaucoma screening is proposed, which integrates the deep hierarchical context of the global fundus image and the local optic disc region. Four deep streams on different levels and modules are, respectively, considered as global image stream, segmentation-guided network, local disc region stream, and disc polar transformation stream. Finally, the output probabilities of different streams are fused as the final screening result. The experiments on two glaucoma data sets (SCES and new SINDI data sets) show that our method outperforms other state-of-the-art algorithms. Huazhu Fu, Jun Cheng 0003, Yanwu Xu 0001, Changqing Zhang 0002, Damon Wing Kee Wong, Jiang Liu 0001, Xiaochun Cao |
IEEE Trans. Medical Imaging | 1 |
| 2017 | Image Cosegmentation via Saliency-Guided Constrained Clustering with Cosine SimilarityabstractCosegmentation jointly segments the common objects from multiple images. In this paper, a novel clustering algorithm, called Saliency-Guided Constrained Clustering approach with Cosine similarity (SGC3), is proposed for the image cosegmentation task, where the common foregrounds are extracted via a one-step clustering process. In our method, the unsupervised saliency prior is utilized as a partition-level side information to guide the clustering process. To guarantee the robustness to noise and outlier in the given prior, the similarities of instance-level and partition-level are jointly computed for cosegmentation. Specifically, we employ cosine distance to calculate the feature similarity between data point and its cluster centroid, and introduce a cosine utility function to measure the similarity between clustering result and the side information. These two parts are both based on the cosine similarity, which is able to capture the intrinsic structure of data, especially for the non-spherical cluster structure. Finally, a K-means-like optimization is designed to solve our objective function in an efficient way. Experimental results on two widely-used datasets demonstrate our approach achieves competitive performance over the state-of-the-art cosegmentation methods. Zhiqiang Tao, Hongfu Liu 0001, Huazhu Fu, Yun Fu 0001 |
AAAI | 3 |
| 2017 | Latent Multi-view Subspace ClusteringabstractIn this paper, we propose a novel Latent Multi-view Subspace Clustering (LMSC) method, which clusters data points with latent representation and simultaneously explores underlying complementary information from multiple views. Unlike most existing single view subspace clustering methods that reconstruct data points using original features, our method seeks the underlying latent representation and simultaneously performs data reconstruction based on the learned latent representation. With the complementarity of multiple views, the latent representation could depict data themselves more comprehensively than each single view individually, accordingly makes subspace representation more accurate and robust as well. The proposed method is intuitive and can be optimized efficiently by using the Augmented Lagrangian Multiplier with Alternating Direction Minimization (ALM-ADM) algorithm. Extensive experiments on benchmark datasets have validated the effectiveness of our proposed method. Changqing Zhang 0002, Qinghua Hu, Huazhu Fu, Pengfei Zhu 0001, Xiaochun Cao |
CVPR | 3 |
| 2017 | Object-Based Multiple Foreground Segmentation in RGBD VideoabstractWe present an RGB and Depth (RGBD) video segmentation method that takes advantage of depth data and can extract multiple foregrounds in the scene. This video segmentation is addressed as an object proposal selection problem formulated in a fully-connected graph, where a flexible number of foregrounds may be chosen. In our graph, each node represents a proposal, and the edges model intra-frame and inter-frame constraints on the solution. The proposals are selected based on an RGBD video saliency map in which depth-based features are utilized to enhance the identification of foregrounds. Experiments show that the proposed multiple foreground segmentation method outperforms related techniques, and the depth cue serves as a helpful complement to RGB features. Moreover, our method provides performance comparable to the state-of-the-art RGB video segmentation techniques on regular RGB videos with estimated depth maps. Huazhu Fu, Dong Xu 0001, Stephen Lin 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Flexible Multi-View Dimensionality Co-ReductionabstractDimensionality reduction aims to map the high-dimensional inputs onto a low-dimensional subspace, in which the similar points are close to each other and vice versa. In this paper, we focus on unsupervised dimensionality reduction for the data with multiple views, and propose a novel method, called Multi-view Dimensionality co-Reduction. Our method flexibly exploits the complementarity of multiple views during the dimensionality reduction and respects the similarity relationships between data points across these different views. The kernel matching constraint based on Hilbert-Schmidt Independence Criterion enhances the correlations and penalizes the disagreement of different views. Specifically, our method explores the correlations within each view independently, and maximizes the dependence among different views with kernel matching jointly. Thus, the locality within each view and the consistence between different views are guaranteed in the subspaces corresponding to different views. More importantly, benefiting from the kernel matching, our method need not depend on a common low-dimensional subspace, which is critical to reduce the influence of the unbalanced dimensionalities of multiple views. Specifically, our method explicitly produces individual low-dimensional projections for individual views, which could be applied for new coming data in the out-of-sample manner. Experiments on both clustering and recognition tasks demonstrate the advantages of the proposed method over the state-of-the-art approaches. Changqing Zhang 0002, Huazhu Fu, Qinghua Hu, Pengfei Zhu 0001, Xiaochun Cao |
IEEE Trans. Image Process. | 2 |
| 2017 | Segmentation and Quantification for Angle-Closure Glaucoma Assessment in Anterior Segment OCTabstractAngle-closure glaucoma is a major cause of irreversible visual impairment and can be identified by measuring the anterior chamber angle (ACA) of the eye. The ACA can be viewed clearly through anterior segment optical coherence tomography (AS-OCT), but the imaging characteristics and the shapes and locations of major ocular structures can vary significantly among different AS-OCT modalities, thus complicating image analysis. To address this problem, we propose a data-driven approach for automatic AS-OCT structure segmentation, measurement, and screening. Our technique first estimates initial markers in the eye through label transfer from a hand-labeled exemplar data set, whose images are collected over different patients and AS-OCT modalities. These initial markers are then refined by using a graph-based smoothing method that is guided by AS-OCT structural information. These markers facilitate segmentation of major clinical structures, which are used to recover standard clinical parameters. These parameters can be used not only to support clinicians in making anatomical assessments, but also to serve as features for detecting anterior angle closure in automatic glaucoma screening algorithms. Experiments on Visante AS-OCT and Cirrus high-definition-OCT data sets demonstrate the effectiveness of our approach. Huazhu Fu, Yanwu Xu 0001, Stephen Lin 0001, Xiaoqin Zhang 0002, Damon Wing Kee Wong, Jiang Liu 0001, Alejandro F. Frangi, Mani Baskaran, Tin Aung |
IEEE Trans. Medical Imaging | 1 |
| 2016 | DeepVessel: Retinal Vessel Segmentation via Deep Learning and Conditional Random Field
Huazhu Fu, Yanwu Xu 0001, Stephen Lin 0001, Damon Wing Kee Wong, Jiang Liu 0001 |
MICCAI (2) | 1 |
| 2016 | Axial Alignment for Anterior Segment Swept Source Optical Coherence Tomography via Robust Low-Rank Tensor RecoveryabstractWe present a one-step approach based on low-rank tensor recovery for axial alignment in 360-degree anterior chamber optical coherence tomography. Achieving translational alignment and rotation correction of cross-sections simultaneously, this technique obtains a better anterior segment topographical representation and improves quantitative measurement accuracy and reproducibility of disease related parameters. Through its use of global information, the proposed method is more robust compared to using only individual or paired slices, and less sensitive to noise and motion artifacts. In angle closure analysis on 30 patient eyes, the preliminary results indicate that the proposed axial alignment method can not only facilitate manual qualitative analysis with more distinct landmark representation and much less human labor, but also can improve the accuracy of automatic quantitative assessment by 2.9 %, which demonstrates that the proposed approach is promising for a wide range of clinical applications. These keywords were added by machine and not by the authors. This process is experimental and the keywords may be updated as the learning algorithm improves. Yanwu Xu 0001, Lixin Duan, Huazhu Fu, Xiaoqin Zhang 0002, Damon Wing Kee Wong, Mani Baskaran, Tin Aung, Jiang Liu 0001 |
MICCAI (3) | 3 |
| 2016 | Unsupervised pixel-level video foreground object segmentation via shortest path algorithm
Xiaochun Cao, Feng Wang 0063, Huazhu Fu, Chao Li 0001 |
Neurocomputing | 4 |
| 2016 | Saliency-Aware Nonparametric Foreground Annotation Based on Weakly Labeled DataabstractIn this paper, we focus on annotating the foreground of an image. More precisely, we predict both image-level labels (category labels) and object-level labels (locations) for objects within a target image in a unified framework. Traditional learning-based image annotation approaches are cumbersome, because they need to establish complex mathematical models and be frequently updated as the scale of training data varies considerably. Thus, we advocate the nonparametric method, which has shown potential in numerous applications and turned out to be attractive thanks to its advantages, i.e., lightweight training load and scalability. In particular, we exploit the salient object windows to describe images, which is beneficial to image retrieval and, thus, the subsequent image-level annotation and localization tasks. Our method, namely, saliency-aware nonparametric foreground annotation, is practical to alleviate the full label requirement of training data, and effectively addresses the problem of foreground annotation. The proposed method only relies on retrieval results from the image database, while pretrained object detectors are no longer necessary. Experimental results on the challenging PASCAL VOC 2007 and PASCAL VOC 2008 demonstrate the advance of our method. Xiaochun Cao, Changqing Zhang 0002, Huazhu Fu, Xiaojie Guo 0001, Qi Tian 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Diversity-induced Multi-view Subspace ClusteringabstractIn this paper, we focus on how to boost the multi-view clustering by exploring the complementary information among multi-view features. A multi-view clustering framework, called Diversity-induced Multi-view Subspace Clustering (DiMSC), is proposed for this task. In our method, we extend the existing subspace clustering into the multi-view domain, and utilize the Hilbert Schmidt Independence Criterion (HSIC) as a diversity term to explore the complementarity of multi-view representations, which could be solved efficiently by using the alternating minimizing optimization. Compared to other multi-view clustering methods, the enhanced complementarity reduces the redundancy between the multi-view representations, and improves the accuracy of the clustering results. Experiments on both image and video face clustering well demonstrate that the proposed method outperforms the state-of-the-art methods. Xiaochun Cao, Changqing Zhang 0002, Huazhu Fu, Si Liu 0001, Hua Zhang 0008 |
CVPR | 3 |
| 2015 | Object-based RGBD image co-segmentation with mutex constraintabstractWe present an object-based co-segmentation method that takes advantage of depth data and is able to correctly handle noisy images in which the common foreground object is missing. With RGBD images, our method utilizes the depth channel to enhance identification of similar foreground objects via a proposed RGBD co-saliency map, as well as to improve detection of object-like regions and provide depth-based local features for region comparison. To accurately deal with noisy images where the common object appears more than or less than once, we formulate co-segmentation in a fully-connected graph structure together with mutual exclusion (mutex) constraints that prevent improper solutions. Experiments show that this object-based RGBD co-segmentation with mutex constraints outperforms related techniques on an RGBD co-segmentation dataset, while effectively processing noisy images. Moreover, we show that this method also provides performance comparable to state-of-the-art RGB co-segmentation techniques on regular RGB images with depth maps estimated from them. Huazhu Fu, Dong Xu 0001, Stephen Lin 0001, Jiang Liu 0001 |
CVPR | 1 |
| 2015 | Low-Rank Tensor Constrained Multiview Subspace ClusteringabstractIn this paper, we explore the problem of multiview subspace clustering. We introduce a low-rank tensor constraint to explore the complementary information from multiple views and, accordingly, establish a novel method called Low-rank Tensor constrained Multiview Subspace Clustering (LT-MSC). Our method regards the subspace representation matrices of different views as a tensor, which captures dexterously the high order correlations underlying multiview data. Then the tensor is equipped with a low-rank constraint, which models elegantly the cross information among different views, reduces effectually the redundancy of the learned subspace representations, and improves the accuracy of clustering as well. The inference process of the affinity matrix for clustering is formulated as a tensor nuclear norm minimization problem, constrained with an additional L2,1-norm regularizer and some linear equalities. The minimization problem is convex and thus can be solved efficiently by an Augmented Lagrangian Alternating Direction Minimization (AL-ADM) method. Extensive experimental results on four benchmark datasets show the effectiveness of our proposed LT-MSC method. Changqing Zhang 0002, Huazhu Fu, Si Liu 0001, Guangcan Liu, Xiaochun Cao |
ICCV | 2 |
| 2015 | Multi-cue Augmented Face ClusteringabstractFace clustering is an important but challenging task since facial images always have huge variation due to change in facial expressions, head poses and partial occlusions, etc. Moreover, face clustering is actually an unsupervised problem which makes it more difficult to reach an accurate result. Fortunately, there are some cues that can be used to improve clustering performance. In this paper, two types of cues are employed. The first one is pairwise constraints: must-link and cannot-link constraints, which can be extracted from the temporal and spatial knowledge of data. The other is that each face is associated with a series of attributes (i.e, gender) which can contribute discrimination among faces. To take advantage of the above cues, we propose a new algorithm, Multi-cue Augmented Face Clustering (McAFC), which effectively incorporates the cues via graph-guided sparse subspace clustering technique. Specially, facial images from the same individual are encouraged to be connected while faces from different persons are restrained to be connected. Experiments on three face datasets from real-world videos show the improvements of our algorithm over the state-of-the-art methods. Chengju Zhou, Changqing Zhang 0002, Huazhu Fu, Rui Wang 0032, Xiaochun Cao |
ACM Multimedia | 3 |
| 2015 | Constrained Multi-View Video Face ClusteringabstractIn this paper, we focus on face clustering in videos. To promote the performance of video clustering by multiple intrinsic cues, i.e., pairwise constraints and multiple views, we propose a constrained multi-view video face clustering method under a unified graph-based model. First, unlike most existing video face clustering methods which only employ these constraints in the clustering step, we strengthen the pairwise constraints through the whole video face clustering framework, both in sparse subspace representation and spectral clustering. In the constrained sparse subspace representation, the sparse representation is forced to explore unknown relationships. In the constrained spectral clustering, the constraints are used to guide for learning more reasonable new representations. Second, our method considers both the video face pairwise constraints as well as the multi-view consistence simultaneously. In particular, the graph regularization enforces the pairwise constraints to be respected and the co-regularization penalizes the disagreement among different graphs of multiple views. Experiments on three real-world video benchmark data sets demonstrate the significant improvements of our method over the state-of-the-art methods. Xiaochun Cao, Changqing Zhang 0002, Chengju Zhou, Huazhu Fu, Hassan Foroosh |
IEEE Trans. Image Process. | 4 |
| 2015 | Object-Based Multiple Foreground Video Co-Segmentation via Multi-State Selection GraphabstractWe present a technique for multiple foreground video co-segmentation in a set of videos. This technique is based on category-independent object proposals. To identify the foreground objects in each frame, we examine the properties of the various regions that reflect the characteristics of foregrounds, considering the intra-video coherence of the foreground as well as the foreground consistency among the different videos in the set. Multiple foregrounds are handled via a multi-state selection graph in which a node representing a video frame can take multiple labels that correspond to different objects. In addition, our method incorporates an indicator matrix that for the first time allows accurate handling of cases with common foreground objects missing in some videos, thus preventing irrelevant regions from being misclassified as foreground objects. An iterative procedure is proposed to optimize our new objective function. As demonstrated through comprehensive experiments, this object-based multiple foreground video co-segmentation method compares well with related techniques that co-segment multiple foregrounds. Huazhu Fu, Dong Xu 0001, Stephen Lin 0001, Rabab K. Ward |
IEEE Trans. Image Process. | 1 |
| 2014 | Object-Based Multiple Foreground Video Co-segmentationabstractWe present a video co-segmentation method that uses category-independent object proposals as its basic element and can extract multiple foreground objects in a video set. The use of object elements overcomes limitations of low-level feature representations in separating complex foregrounds and backgrounds. We formulate object-based co-segmentation as a co-selection graph in which regions with foreground-like characteristics are favored while also accounting for intra-video and inter-video foreground coherence. To handle multiple foreground objects, we expand the co-selection graph model into a proposed multi-state selection graph model (MSG) that optimizes the segmentations of different objects jointly. This extension into the MSG can be applied not only to our co-selection graph, but also can be used to turn any standard graph model into a multi-state selection solution that can be optimized directly by the existing energy minimization techniques. Our experiments show that our object-based multiple foreground video co-segmentation method (ObMiC) compares well to related techniques on both single and multiple foreground cases. Huazhu Fu, Dong Xu 0001, Stephen Lin 0001 |
CVPR | 1 |
| 2014 | Co-Saliency Detection via Base ReconstructionabstractCo-saliency aims at detecting common saliency in a series of images, which is useful for a variety of multimedia applications. In this paper, we address the co-saliency detection to a reconstruction problem: the foreground could be well reconstructed by using the reconstruction bases, which are extracted from each image and have the similar appearances in the feature space. We firstly obtain a candidate set by measuring the saliency prior of each image. Relevance information among the multiple images is utilized to remove the inaccuracy reconstruction bases. Finally, with the updated reconstruction bases, we rebuild the images and provide the reconstruction error regarded as a negative correlational value in co-saliency measurement. The satisfactory quantitative and qualitative experimental results on two benchmark datasets demonstrate the efficiency and effectiveness of our method. Xiaochun Cao, Yupeng Cheng, Zhiqiang Tao, Huazhu Fu |
ACM Multimedia | 4 |
| 2014 | Symmetry Constraint for Foreground ExtractionabstractSymmetry as an intrinsic shape property is often observed in natural objects. In this paper, we discuss how explicitly taking into account the symmetry constraint can enhance the quality of foreground object extraction. In our method, a symmetry foreground map is used to represent the symmetry structure of the image, which includes the symmetry matching magnitude and the foreground location prior. Then, the symmetry constraint model is built by introducing this symmetry structure into the graph-based segmentation function. Finally, the segmentation result is obtained via graph cuts. Our method encourages objects with symmetric parts to be consistently extracted. Moreover, our symmetry constraint model is applicable to weak symmetric objects under the part-based framework. Quantitative and qualitative experimental results on benchmark datasets demonstrate the advantages of our approach in extracting the foreground. Our method also shows improved results in segmenting objects with weak, complex symmetry properties. Huazhu Fu, Xiaochun Cao, Zhuowen Tu, Dongdai Lin |
IEEE Trans. Cybern. | 1 |
| 2014 | Self-Adaptively Weighted Co-Saliency Detection via Rank ConstraintabstractCo-saliency detection aims at discovering the common salient objects existing in multiple images. Most existing methods combine multiple saliency cues based on fixed weights, and ignore the intrinsic relationship of these cues. In this paper, we provide a general saliency map fusion framework, which exploits the relationship of multiple saliency cues and obtains the self-adaptive weight to generate the final saliency/cosaliency map. Given a group of images with similar objects, our method firstly utilizes several saliency detection algorithms to generate a group of saliency maps for all the images. The feature representation of the co-salient regions should be both similar and consistent. Therefore, the matrix jointing these feature histograms appears low rank. We formalize this general consistency criterion as the rank constraint, and propose two consistency energy to describe it, which are based on low rank matrix approximation and low rank matrix recovery, respectively. By calculating the self-adaptive weight based on the consistency energy, we highlight the common salient regions. Our method is valid for more than two input images and also works well for single image saliency detection. Experimental results on a variety of benchmark datasets demonstrate that the proposed method outperforms the state-of-the-art methods. Xiaochun Cao, Zhiqiang Tao, Huazhu Fu, Wei Feng 0005 |
IEEE Trans. Image Process. | 4 |
| 2014 | Regularity Preserved Superpixels and SupervoxelsabstractMost existing superpixel algorithms ignore the spatial structure and regularity properties, which result in undesirable sizes and location relationships for the subsequent processing. In this paper, we introduce a new method to generate the regularity preserved superpixels. Starting from the lattice seeds, our method relocates them to the pixel with locally maximal edge magnitudes and treats them as the superpixel junctions. Then, the shortest path algorithm is employed to find the local optimal boundary connecting each adjacent junction pair. Thanks to the local constraints, our method obtains homogeneous superpixels with adjacency in lowly textured and uniform regions and simultaneously preserves the boundary adherence in the high contrast contents. Our method preserves the regularity property without significantly sacrificing the segmentation accuracy. Moreover, we extend this regular constraint for generating the supervoxels. Our method obtains the regular supervoxels, which preserves the structural relation on both spatial and temporal spaces of the video. Quantitative and qualitative experimental results on benchmark datasets demonstrate that our simple but effective method outperforms the existing regular superpixel methods. Huazhu Fu, Xiaochun Cao, Dai Tang, Yahong Han, Dong Xu 0001 |
IEEE Trans. Multim. | 1 |
| 2013 | Saliency map fusion based on rank-one constraintabstractCo-saliency is the common saliency existing in multiple images, which keeps consistent in saliency maps. One saliency detection method generates saliency maps for all the input images, so that we have a group of maps. Salient region of each image is extracted by its corresponding saliency map in the group. We use a matrix to combine all the salient regions. Ideally, these co-salient regions are similar and consistent, and therefore the matrix rank appears low. In this paper, we formalize this general consistency criterion as rankone constraint and propose a consistency energy to measure the approximation degree between matrix rank and one. We combine the single and multiple image saliency maps, and adaptively weight these maps under the rank-one constraint to generate the co-saliency map. Our method is valid for more than two input images and has more robustness than the existing co-saliency methods. Experimental results on benchmark database demonstrate that our method has the satisfactory performance on co-saliency detection. Xiaochun Cao, Zhiqiang Tao, Huazhu Fu |
ICME | 4 |
| 2013 | Cluster-Based Co-Saliency DetectionabstractCo-saliency is used to discover the common saliency on the multiple images, which is a relatively underexplored area. In this paper, we introduce a new cluster-based algorithm for co-saliency detection. Global correspondence between the multiple images is implicitly learned during the clustering process. Three visual attention cues: contrast, spatial, and corresponding, are devised to effectively measure the cluster saliency. The final co-saliency maps are generated by fusing the single image saliency and multiimage saliency. The advantage of our method is mostly bottom-up without heavy learning, and has the property of being simple, general, efficient, and effective. Quantitative and qualitative experiments result in a variety of benchmark datasets demonstrating the advantages of the proposed method over the competing co-saliency methods. Our method on single image also outperforms most the state-of-the-art saliency detection methods. Furthermore, we apply the co-saliency method on four vision applications: co-segmentation, robust image distance, weakly supervised learning, and video foreground detection, which demonstrate the potential usages of the co-saliency map. Huazhu Fu, Xiaochun Cao, Zhuowen Tu |
IEEE Trans. Image Process. | 1 |
| 2012 | Topology Preserved Regular SuperpixelabstractMost existing super pixel algorithms ignore the topology and regularities, which results in undesirable sizes and location relationships for subsequent processing. In this paper, we introduce a new method to compute the regular super pixels while preserving the topology. Start from regular seeds, our method relocates them to the pixel with locally maximal edge magnitudes. Then, we find the local optimal path between each relocated seed and its four neighbors using Dijkstra algorithm. Thanks to the local constraints, our method obtains homogeneous super pixels with explicit adjacency in low-texture and uniform regions and, simultaneously, maintains the edge cues in the high contrast and salient contents. Quantitative and qualitative experimental results on Berkeley Segmentation Database Benchmark demonstrate that our proposed algorithm outperforms the existing regular super pixel methods. Dai Tang, Huazhu Fu, Xiaochun Cao |
ICME | 2 |
| 2012 | Forgery Authentication in Extreme Wide-Angle Lens Using Distortion Cue and Fake Saliency MapabstractDistortion is often considered as an unfavorable factor in most image analysis. However, it is undeniable that the distortion reflects the intrinsic property of the lens, especially, the extreme wide-angle lens, which has a significant distortion. In this paper, we discuss how explicitly employing the distortion cues can detect the forgery object in distortion image and make the following contributions: 1) a radial distortion projection model is adopted to simplify the traditional captured ray-based models, where the straight world line is projected into a great circle on the viewing sphere; 2) two bottom-up cues based on distortion constraint are provided to discriminate the authentication of the line in the image; 3) a fake saliency map is used to maximum fake detection density, and based on the fake saliency map, an energy function is provided to achieve the pixel-level forgery object via graph cut. Experimental results on simulated data and real images demonstrate the performances of our method. Huazhu Fu, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2011 | Embedded omni-vision navigator based on multi-object tracking
Huazhu Fu, Zuoliang Cao, Xiaochun Cao |
Mach. Vis. Appl. | 1 |