EDBT 2026 Demo / reviewers in the wild / expert
Lequan Yu
dblp:165/8092
· DBLP profile ↗
137ranked-venue papers
8as first author
99since 2021 · last 2026
0000-0002-9315-6527ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 80 · 5 first-author · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 61 · 5 first-author · 41 since 2021Artificial intelligence and machine learning · 48 · 3 first-author · 39 since 2021Systems, architecture and hardware · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PASS: Probabilistic Agentic Supernet Sampling for Interpretable and Adaptive Chest X-Ray ReasoningabstractExisting tool-augmented agentic systems are limited in the real world by (i) black-box reasoning steps that undermine trust of decision-making and pose safety risks, (ii) poor multimodal integration, which is inherently critical for healthcare tasks, and (iii) rigid and computationally inefficient agentic pipelines. We introduce PASS (Probabilistic Agentic Supernet Sampling), the first multimodal framework to address these challenges for Chest X-Ray (CXR) reasoning. PASS adaptively samples agentic workflows over a multi-tool graph, yielding decision paths annotated with interpretable probabilities. Given the complex CXR reasoning task with multimodal medical data, PASS leverages its learned task-conditioned distribution over the agentic supernet. Thus, it adaptively selects the most suitable tool at each supernet layer, offering probability-annotated trajectories for post-hoc audits and directly enhancing medical AI safety. PASS also continuously compresses salient findings into an evolving personalized memory, while dynamically deciding whether to deepen its reasoning path or invoke an early exit for efficiency. To optimize a Pareto frontier balancing performance and cost, we design a novel three-stage training procedure, including expert knowledge warm-up, contrastive path-ranking, and cost-aware reinforcement learning. To facilitate rigorous evaluation, we introduce CAB-E, a comprehensive benchmark for multi-step, safety-critical, free-form reasoning. Experiments across various benchmarks validate that PASS significantly outperforms strong baselines in multiple metrics (e.g., accuracy, LLM-Judge, semantic similarity, etc.) while balancing computational costs, pushing a new paradigm shift towards interpretable, adaptive, and multimodal medical agentic systems. Yushi Feng, Junye Du, Yingying Hong, Lequan Yu |
AAAI | 5 |
| 2026 | FDP: A Frequency-Decomposition Preprocessing Pipeline for Unsupervised Anomaly Detection in Brain MRIabstractDue to the diversity of brain anatomy and the scarcity of annotated data, supervised anomaly detection for brain MRI remains challenging, driving the development of unsupervised anomaly detection (UAD) approaches. Current UAD methods typically utilize synthetically generated noise perturbations on healthy MRIs to train generative models for normal anatomy reconstruction, enabling anomaly detection via residual maps. However, such simulated anomalies lack the biophysical fidelity and morphological complexity characteristic of true clinical lesions. To advance UAD in brain MRI, we conduct the first systematic frequency-domain analysis of pathological signatures, revealing two key properties: (1) anomalies exhibit unique frequency patterns distinguishable from normal anatomy, and (2) low-frequency signals maintain consistent representations across healthy scans. These insights motivate our Frequency-Decomposition Preprocessing (FDP) framework—the first UAD method to leverage frequency-domain reconstruction for simultaneous pathology suppression and anatomical preservation. FDP can integrate seamlessly with existing anomaly simulation techniques, consistently enhancing detection performance across diverse architectures while maintaining diagnostic fidelity. Experimental results demonstrate that FDP consistently improves anomaly detection performance when integrated with existing methods. Notably, FDP achieves a 17.63% increase in DICE score with LDM while maintaining robust improvements across multiple baselines. Zhenfeng Zhuang, Qiong Peng, Lequan Yu, Liansheng Wang 0002 |
AAAI | 7 |
| 2026 | Augmenting Clinical Decision-Making with an Interactive and Interpretable AI Copilot: A Real-World User Study with Clinicians in Nephrology and ObstetricsabstractClinician skepticism toward opaque AI hinders adoption in high-stakes healthcare. We present AICare, an interactive and interpretable AI copilot for collaborative clinical decision-making. By analyzing longitudinal electronic health records, AICare grounds dynamic risk predictions in scrutable visualizations and LLM-driven diagnostic recommendations. Through a within-subjects counterbalanced study with 16 clinicians across nephrology and obstetrics, we comprehensively evaluated AICare using objective measures (task completion time and error rate), subjective assessments (NASA-TLX, SUS, and confidence ratings), and semi-structured interviews. Our findings indicate AICare’s reduced cognitive workload. Beyond performance metrics, qualitative analysis reveals that trust is actively constructed through verification, with interaction strategies diverging by expertise: junior clinicians used the system as cognitive scaffolding to structure their analysis, while experts engaged in adversarial verification to challenge the AI’s logic. This work offers design implications for creating AI systems that function as transparent partners, accommodating diverse reasoning styles to augment rather than replace clinical judgment. Yinghao Zhu, Dehao Sui, Xuning Hu, Yifan Qi, Tianchen Wu, Wen Tang 0001, Zhihan Cui, Yasha Wang, Lequan Yu, Ewen M. Harrison, Liantao Ma |
CHI | 13 |
| 2026 | Tackling missing modalities with memory-efficient modality-complementary prompt learning for robust brain tumor segmentation
Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
Expert Syst. Appl. | 3 |
| 2026 | Generative Enhancement for 3D Medical ImagesabstractAbstract The limited availability of 3D medical image datasets, due to privacy concerns and high collection or annotation costs, poses significant challenges in the field of medical imaging. There are few solutions for realistic 3D medical image synthesis due to difficulties in backbone design and fewer 3D training samples compared to 2D counterparts. In this paper, we propose GEM-3D , a novel generative approach to the synthesis of 3D medical images and the enhancement of existing datasets using conditional diffusion models. Our method begins with a 2D slice, noted as the informed slice to serve the patient prior, and propagates the generation process using a 3D segmentation mask. By decomposing the 3D medical images into editable masks and patient prior information, GEM-3D offers a flexible yet effective solution for generating versatile 3D images from existing datasets. Moreover, as the informed slice contains patient-wise information, GEM-3D can also facilitate counterfactual image synthesis and dataset-level de-enhancement with desired control. Experiments on brain MRI and abdomen CT images demonstrate that GEM-3D is capable of synthesizing high-quality 3D medical images with volumetric consistency, offering a straightforward solution for dataset enhancement during inference. The code is available at https://github.com/HKU-MedAI/GEM-3D . Lingting Zhu, Noel Codella, Dongdong Chen 0001, Zhenchao Jin, Lu Yuan 0001, Lequan Yu |
Int. J. Comput. Vis. | 6 |
| 2026 | Reconstructing shared visual experiences from human brain activity across individuals
Yanyan Huang, Kaiqiang Xu, Yannan Chen, Lequan Yu, Zhijun Yao, Yu Fu 0008 |
Medical Image Anal. | 6 |
| 2026 | FADFNet: A fine-tunable and adaptive decomposition-fusion network for cross-dataset low-dose CT and low-dose PET image reconstruction
Fangji Qian, Yanyan Huang, Meng Niu, Yuanxue Gao, Kuangyu Shi, Lequan Yu, Yu Fu 0008, Cheng Zhuo |
Medical Image Anal. | 8 |
| 2026 | MorphoNet: Morphological sub-region-based structure learning for WSI analysisabstract• Propose MorphoNet framework to capture morphological patterns and long-range tissue structures in WSIs. • Develop a Morphological Sub-Region Grouping mechanism to model WSIs as spatially coherent sub-regions. • Introduce a spatial-aware clustering approach and a sub-region aggregation strategy to derive sub-region embeddings. • Achieve superior performance across 10 public benchmarks, outperforming state-of-the-art methods in tumor subtyping and survival prediction. Representation learning of Whole slide image (WSI) is fundamental to computational pathology, enabling tasks such as tumor subtyping, survival prediction, and cancer grading. Existing methods typically tile WSIs into thousands of small patches and aggregate patch features into slide-level embeddings, but this patch-centric paradigm suffers from redundancy and suboptimal spatial modeling. Built upon these patch-level embeddings, Multiple Instance Learning (MIL) methods overfit to scattered discriminative patches, graph-based models mainly capture local neighborhoods, and prototype-based approaches often ignore spatial coherence and under-represent rare tissue patterns. To address these challenges, we propose MorphoNet, a Morph ological structure learning Net work that captures long-range spatial tissue relationships while extracting informative morphological patterns. The key idea of MorphoNet is Morphological Sub-Region Grouping (MSRG), which clusters spatially adjacent patches with similar appearance into compact sub-region embeddings, reducing redundancy and forming semantically coherent morphological units. Sub-region graph is then constructed and processed by a lightweight Graph Neural Network (GNN) to model contextual dependencies and derive slide-level representations. Importantly, MSRG is a plug-and-play module that can be integrated into MIL, graph-based, and prototype-based pipelines, consistently improving their performance. Experiments on ten public benchmarks demonstrate that MorphoNet achieves superior performance on tumor subtyping and survival prediction. Our code is available at https://github.com/fuying-wang/MorphoNet . Fuying Wang, Junjun He, Liansheng Wang 0002, Jianning Chen, Lequan Yu |
Medical Image Anal. | 9 |
| 2026 | Multi-contrast low-field MRI acceleration with k-space progressive learning and image-space hybrid attention fusion
Xiaohan Xing, Qi Chen 0014, Lequan Yu, Lingting Zhu, Lei Xing 0001, Lianli Liu |
Medical Image Anal. | 3 |
| 2026 | WAS-Mamba: 3D Medical Image Segmentation via Windowed Attention State Space ModelabstractMamba, the state space model (SSM), has attracted significant attention for its ability to model long-range dependencies with linear complexity, achieving success in medical image segmentation. However, the previous cross-scanning approach in Mamba struggles to capture both long-range and short-range dependencies simultaneously and treats the features of each path equally. This imbalance between local and global modeling capabilities can adversely impact segmentation performance. To address these challenges, we propose WAS-Mamba, a novel method specifically designed for Mamba-based medical image segmentation. WAS-Mamba introduces a cross-channel window scanning strategy (CCWScan) that enables sequences to preserve original local image features during the transformation process. Furthermore, WAS-Mamba employs a weighted state space model (WSSM) to dynamically fuse spatial and frequency domain information, improving the capture of local details and global context for accurate segmentation. We validated the superior performance of WAS-Mamba across five datasets covering different anatomical regions, which include CT and MRI images: Synapse, BTCV, ACDC, BraTS, and Decathlon-Lung. In particular, we achieved a Dice coefficient of 88.09% on the Synapse dataset, with a 33% reduction in computational complexity and inference time compared to the second-best model. The code and model will be released at https://github.com/1605066114/WAS-Mamba. Xueren Zhang, Xianghong Wang, Nuo Tong, Jichen Du, Mingchao Ding, Lequan Yu, Yixuan Yuan, Tianye Niu |
IEEE Trans. Image Process. | 7 |
| 2026 | MuMA: 3D PBR Texturing via Multi-Channel Multi-View Generation and Albedo Post-ProcessingabstractCurrent methods for 3D generation still fall short in physically based rendering (PBR) texturing, primarily due to limited data and challenges in modeling multi-channel materials. In this work, we propose MuMA, a method for 3D PBR texturing through Multi-channel Multi-view generation and Albedo post-processing. Our approach features two key innovations: 1) we opt to model shaded and albedo appearance channels, where the shaded channels enables the integration intrinsic decomposition modules for material properties; and 2) leveraging multimodal large language models, we emulate artists' techniques for material assessment and selection. Experiments demonstrate that MuMA achieves superior results in visual quality and material fidelity compared to existing methods. Lingting Zhu, Jingrui Ye, Zeyu Hu, Yingda Yin, Lanjiong Li, Jinnan Chen, Shengju Qian, Xin Wang 0178, Qingmin Liao, Lequan Yu |
IEEE Trans. Image Process. | 11 |
| 2026 | SIB-MIL: Sparsity-Induced Bayesian Neural Network for Robust Multiple Instance Learning on Whole Slide Image AnalysisabstractMultiple instance learning (MIL) has shown prominent success in analyzing whole slide histopathology images (WSIs). However, existing MIL methods often suffer from overfitting due to weak supervision and the "needle-in-a-haystack" nature of WSIs. Additionally, most deterministic approaches lack a mechanism for uncertainty quantification. While Bayesian neural networks (BNNs) have emerged as a promising solution to mitigate overfitting and enable uncertainty estimation by imposing prior constraints, commonly used Gaussian BNNs exhibit unstable posterior predictive distributions under weak supervision and suffer from high prediction variance. To tackle these challenges, we propose a sparsity-induced Bayesian Neural Network to be adopted in the MIL scheme, named SIB-MIL, for robust WSI prediction. Instead of using Gaussian prior distributions, we place a sparsity-induced prior, the Horseshoe prior, on the BNN parameters to address the variance overflowing issue. Such sparsity also filters unimportant noise and highlights salient regions, which only occupy a small proportion in WSIs. Empirical evaluations on cancer classification and subtyping tasks corroborate that not only can our method improve the existing MIL networks, but it also performs well in uncertainty quantification. Codes are available at https://github.com/HKU-MedAI/SIB-MIL. Yihang Chen 0001, Tsai Hor Chan, Jianning Chen, Guosheng Yin, Lequan Yu |
IEEE Trans. Medical Imaging | 6 |
| 2026 | MCS-Stain: Boosting FFPE-to-HE Virtual Staining With Multiple Cell SemanticsabstractThe diagnosis of cancer primarily relies on pathological slides stained with hematoxylin and eosin (HE). These slides are typically prepared from tissue samples that have been fixed in formalin and embedded in paraffin (FFPE). However, the traditional process of staining FFPE samples with HE is time-consuming and resource-intensive. Recent advances in virtual staining technologies, driven by digital pathology and generative models, offer a promising alternative. However, the blurred structures in FFPE images pose unique challenges to achieving high-quality FFPE-to-HE virtual staining. In this context, we developed a novel Multiple Cell Semantics-guided supervised generative adversarial model, MCS-Stain. Specifically, the guidance consists of three components: 1) pretrained cell semantic guidance, aligning the powerful intermediate features of real and virtual images, embedded in the pretrained cell segmentation model (PCSM); 2) cell mask guidance, introducing comprehensible cell information which serves as part of the input to the discriminator through channel concatenation; 3) dynamic cell semantic guidance, aligning the dynamic intermediate features embedded in the generator during training. The comparative results on FFPE-to-HE datasets demonstrated that MCS-Stain outperforms existing state-of-the-art (SOTA) methods with substantial qualitative and quantitative improvements. Results across various PCSMs and data sources further confirmed its effectiveness and robustness. Notably, the dynamic cell semantic exhibits strong potential beyond FFPE-to-HE virtual staining, further demonstrated by virtual staining from HE images to immunohistochemical (IHC) images. In general, MCS-Stain presents a promising avenue to advance virtual staining techniques. Code is available at https://github.com/huyihuang/MCS-Stain. Yihuang Hu, Zhicheng Du, Weiping Lin, Shurong Yang, Lequan Yu, Liansheng Wang 0002 |
IEEE Trans. Medical Imaging | 5 |
| 2026 | Disentangled Multi-Modal Learning of Histology and Transcriptomics for Cancer CharacterizationabstractHistopathology remains the gold standard for cancer diagnosis and prognosis. With the advent of transcriptome profiling, multi-modal learning combining transcriptomics with histology offers more comprehensive information. However, existing multi-modal approaches are challenged by intrinsic multi-modal heterogeneity, insufficient multi-scale integration, and reliance on paired data, restricting clinical applicability. To address these challenges, we propose a disentangled multi-modal framework with four contributions: 1) to mitigate multi-modal heterogeneity, we decompose WSIs and transcriptomes into tumor and microenvironment subspaces using a disentangled multi-modal fusion module, and introduce a confidence-guided gradient coordination strategy to balance subspace optimization; 2) to enhance multi-scale integration, we propose an inter-magnification gene-expression consistency strategy that aligns transcriptomic signals across WSI magnifications; 3) to reduce dependency on paired data, we propose a subspace knowledge distillation strategy enabling transcriptome-agnostic inference through a WSI-only student model; and 4) to improve inference efficiency, we propose an informative token aggregation module that suppresses WSI redundancy while preserving subspace semantics. Extensive experiments on cancer diagnosis, prognosis, and survival prediction demonstrate our superiority over state-of-the-art methods across multiple settings. Code is available at GitHub. Xiaofei Wang 0004, Anran Liu 0001, Lequan Yu, Chao Li 0031 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Large Images Are Gaussians: High-Quality Large Image Representation with Levels of 2D Gaussian SplattingabstractWhile Implicit Neural Representations (INRs) have demonstrated significant success in image representation, they are often hindered by large training memory and slow decoding speed. Recently, Gaussian Splatting (GS) has emerged as a promising solution in 3D reconstruction due to its highquality novel view synthesis and rapid rendering capabilities, positioning it as a valuable tool for a broad spectrum of applications. In particular, a GS-based representation, 2DGS, has shown potential for image fitting. In our work, we present Large Images are Gaussians (LIG), which delves deeper into the application of 2DGS for image representations, addressing the challenge of fitting large images with 2DGS in the situation of numerous Gaussian points, through two distinct modifications: 1) we adopt a variant of representation and optimization strategy, facilitating the fitting of a large number of Gaussian points; 2) we propose a Level-of-Gaussian approach for reconstructing both coarse low-frequency initialization and fine high-frequency details. Consequently, we successfully represent large images as Gaussian points and achieve high-quality large image representation, demonstrating its efficacy across various types of large images. Lingting Zhu, Guying Lin, Jinnan Chen, Zhenchao Jin, Lequan Yu |
AAAI | 7 |
| 2025 | Dynamic Entity-Masked Graph Diffusion Model for Histopathology Image Representation LearningabstractSignificant disparities between the features of natural images and those inherent to histopathological images make it challenging to directly apply and transfer pre-trained models from natural images to histopathology tasks. Moreover, the frequent lack of annotations in histopathology patch images has driven researchers to explore self-supervised learning methods like mask reconstruction for learning representations from large amounts of unlabeled data. Crucially, previous mask-based efforts in self-supervised learning have often overlooked the spatial interactions among entities, which are essential for constructing accurate representations of pathological entities. To address these challenges, constructing graphs of entities is a promising approach. In addition, the diffusion reconstruction strategy has recently shown superior performance through its random intensity noise addition technique to enhance the robust learned representation. Therefore, we introduce H-MGDM, a novel self-supervised Histopathology image representation learning method through the Dynamic Entity-Masked Graph Diffusion Model. Specifically, we propose to use complementary subgraphs as latent diffusion conditions and self-supervised targets respectively during pre-training. We note that the graph can embed entities' topological relationships and enhance representation. Dynamic conditions and targets can improve pathological fine reconstruction. Our model has conducted pretraining experiments on three large histopathological datasets. The advanced predictive performance and interpretability of H-MGDM are clearly evaluated on comprehensive downstream tasks such as classification and survival analysis on six datasets. Zhenfeng Zhuang, Min Cen, Fangyu Zhou, Lequan Yu, Baptiste Magnier, Liansheng Wang 0002 |
AAAI | 5 |
| 2025 | C2 MIL: Synchronizing Semantic and Topological Causalities in Multiple Instance Learning for Robust and Interpretable Survival AnalysisabstractInternational audience Min Cen, Zhenfeng Zhuang, Baptiste Magnier, Lequan Yu, Liansheng Wang 0002 |
ICCV | 6 |
| 2025 | From Layers to States: A State Space Model Perspective to Deep Neural Network Layer DynamicsabstractThe depth of neural networks is a critical factor for their capability, with deeper models often demonstrating superior performance. Motivated by this, significant efforts have been made to enhance layer aggregation - reusing information from previous layers to better extract features at the current layer, to improve the representational power of deep neural networks. However, previous works have primarily addressed this problem from a discrete-state perspective which is not suitable as the number of network layers grows. This paper novelly treats the outputs from layers as states of a continuous process and considers leveraging the state space model (SSM) to design the aggregation of layers in very deep neural networks. Moreover, inspired by its advancements in modeling long sequences, the Selective State Space Models (S6) is employed to design a new module called Selective State Space Model Layer Aggregation (S6LA). This module aims to combine traditional CNN or transformer architectures within a sequential framework, enhancing the representational capabilities of state-of-the-art vision networks. Extensive experiments show that S6LA delivers substantial improvements in both image classification and detection tasks, highlighting the potential of integrating SSMs with contemporary deep learning techniques. Qinshuo Liu, Weiqin Zhao, Yanwen Fang, Lequan Yu |
ICLR | 5 |
| 2025 | From Token to Rhythm: A Multi-Scale Approach for ECG-Language PretrainingabstractElectrocardiograms (ECGs) play a vital role in monitoring cardiac health and diagnosing heart diseases. However, traditional deep learning approaches for ECG analysis rely heavily on large-scale manual annotations, which are both time-consuming and resource-intensive to obtain. To overcome this limitation, self-supervised learning (SSL) has emerged as a promising alternative, enabling the extraction of robust ECG representations that can be efficiently transferred to various downstream tasks. While previous studies have explored SSL for ECG pretraining and multi-modal ECG-language alignment, they often fail to capture the multi-scale nature of ECG signals. As a result, these methods struggle to learn generalized representations due to their inability to model the hierarchical structure of ECG data. To address this gap, we introduce MELP, a novel Multi-scale ECG-Language Pretraining (MELP) model that fully leverages hierarchical supervision from ECG-text pairs. MELP first pretrains a cardiology-specific language model to enhance its understanding of clinical text. It then applies three levels of cross-modal supervision—at the token, beat, and rhythm levels—to align ECG signals with textual reports, capturing structured information across different time scales. We evaluate MELP on three public ECG datasets across multiple tasks, including zero-shot ECG classification, linear probing, and transfer learning. Experimental results demonstrate that MELP outperforms existing SSL methods, underscoring its effectiveness and adaptability across diverse clinical applications. Our code is available at https://github.com/HKU-MedAI/MELP. Fuying Wang, Lequan Yu |
ICML | 3 |
| 2025 | Cross-Modal Alignment via Variational Copula ModellingabstractVarious data modalities are common in real-world applications. (e.g., EHR, medical images and clinical notes in healthcare). Thus, it is essential to develop multimodal learning methods to aggregate information from multiple modalities. The main challenge is appropriately aligning and fusing the representations of different modalities into a joint distribution. Existing methods mainly rely on concatenation or the Kronecker product, oversimplifying interactions structure between modalities and indicating a need to model more complex interactions. Additionally, the joint distribution of latent representations with higher-order interactions is underexplored. Copula is a powerful statistical structure in modelling the interactions between variables, as it bridges the joint distribution and marginal distributions of multiple variables. In this paper, we propose a novel copula modelling-driven multimodal learning framework, which focuses on learning the joint distribution of various modalities to capture the complex interaction among them. The key idea is interpreting the copula model as a tool to align the marginal distributions of the modalities efficiently. By assuming a Gaussian mixture distribution for each modality and a copula model on the joint distribution, our model can also generate accurate representations for missing modalities. Extensive experiments on public MIMIC datasets demonstrate the superior performance of our model over other competitors. The code is anonymously available at https://github.com/HKU-MedAI/CMCM. Tsai Hor Chan, Fuying Wang, Guosheng Yin, Lequan Yu |
ICML | 5 |
| 2025 | DERI: Cross-Modal ECG Representation Learning with Deep ECG-Report InteractionabstractElectrocardiogram (ECG) is widely used to diagnose cardiac conditions via deep learning methods. Although existing self-supervised learning (SSL) methods have achieved great performance in learning representation for ECG-based cardiac conditions classification, the clinical semantics can not be effectively captured. To overcome this limitation, we proposed to learn cross-modal ECG representations that contain more clinical semantics via a novel framework with \textbf{D}eep \textbf{E}CG-\textbf{R}eport \textbf{I}nteraction (\textbf{DERI}). Specifically, we design a novel framework combining multiple alignments and mutual feature reconstructions to learn effective representation of the ECG with the clinical report, which fuses the clinical semantics of the report. An RME-Module inspired by masked modeling is proposed to improve the ECG representation learning. Furthermore, we extend ECG representation learning to report generation with a language model, which is significant for evaluating clinical semantics in the learned representations and even clinical applications. Comprehensive experiments with various settings are conducted on various datasets to show the superior performance of our DERI. Our code is released on https://github.com/cccccj-03/DERI. Jian Chen 0011, Xiaoru Dong, Wei Wang 0077, Shaorui Zhou, Lequan Yu, Xiping Hu |
IJCAI | 5 |
| 2025 | HyperPath: Knowledge-Guided Hyperbolic Semantic Hierarchy Modeling for WSI Analysis
Peixiang Huang, Yanyan Huang, Weiqin Zhao, Junjun He, Lequan Yu |
MICCAI (5) | 5 |
| 2025 | Bridging Radiological Images and Factors with Vision-Language Model for Accurate Diagnosis of Proliferative Hepatocellular Carcinoma
Yanyan Huang, Peixiang Huang, Yu Fu 0008, Ruimeng Yang, Lequan Yu |
MICCAI (6) | 6 |
| 2025 | MoST-IG: Morphology-Guided Spatial Transcriptomics Integration via Visual-Genomic Graph Optimal Transport
Liting Yu, Weiqin Zhao, Zhuo Liang, Lequan Yu |
MICCAI (12) | 5 |
| 2025 | Amplifying Prominent Representations in Multimodal Learning via Variational Dirichlet ProcessabstractDeveloping effective multimodal fusion approaches has become increasingly essential in many real-world scenarios, such as health care and finance.
The key challenge is how to preserve the feature expressiveness in each modality while learning cross-modal interactions.
Previous approaches primarily focus on the cross-modal alignment,
while over-emphasis on the alignment of marginal distributions of modalities may impose excess regularization and obstruct meaningful representations within each modality.
The Dirichlet process (DP) mixture model is a powerful Bayesian non-parametric method that can amplify the most prominent features by its richer-gets-richer property, which allocates increasing weights to them.
Inspired by this unique characteristic of DP, we propose a new DP-driven multimodal learning framework that automatically achieves an optimal balance between prominent intra-modal representation learning and cross-modal alignment.
Specifically, we assume that each modality follows a mixture of multivariate Gaussian distributions and further adopt DP to calculate the mixture weights for all the components. This paradigm allows DP to dynamically allocate the contributions of features and select the most prominent ones, leveraging its richer-gets-richer property, thus facilitating multimodal feature fusion.
Extensive experiments on several multimodal datasets demonstrate the superior performance of our model over other competitors.
Ablation analysis further validates the effectiveness of DP in aligning modality distributions and its robustness to changes in key hyperparameters.
Code is anonymously available at https://github.com/HKU-MedAI/DPMM.git Tsai Hor Chan, Yihang Chen 0001, Guosheng Yin, Lequan Yu |
NeurIPS | 5 |
| 2025 | Variational Pólya TreeabstractDensity estimation is essential for generative modeling, particularly with the rise of modern neural networks. While existing methods capture complex data distributions, they often lack interpretability and uncertainty quantification. Bayesian nonparametric methods, especially the Pólya tree, offer a robust framework that addresses these issues by accurately capturing function behavior over small intervals. Traditional techniques like Markov chain Monte Carlo (MCMC) face high computational complexity and scalability limitations, hindering the use of Bayesian nonparametric methods in deep learning. To tackle this, we introduce the variational Pólya tree (VPT) model, which employs stochastic variational inference to compute posterior distributions. This model provides a flexible, nonparametric Bayesian prior that captures latent densities and works well with stochastic gradient optimization. We also leverage the joint distribution likelihood for a more precise variational posterior approximation than traditional mean-field methods. We evaluate the model performance on both real data and images, and demonstrate its competitiveness with other state-of-the-art deep density estimation methods. We also explore its ability in enhancing interpretability and uncertainty quantification. Code is available at https://github.com/howardchanth/var-polya-tree. Tsai Hor Chan, Lequan Yu, Kwok Fai Lam, Guosheng Yin |
NeurIPS | 3 |
| 2025 | MedAgentBoard: Benchmarking Multi-Agent Collaboration with Conventional Methods for Diverse Medical TasksabstractThe rapid advancement of Large Language Models (LLMs) has stimulated interest in multi-agent collaboration for addressing complex medical tasks. However, the practical advantages of multi-agent collaboration approaches remain insufficiently understood. Existing evaluations often lack generalizability, failing to cover diverse tasks reflective of real-world clinical practice, and frequently omit rigorous comparisons against both single-LLM-based and established conventional methods. To address this critical gap, we introduce MedAgentBoard, a comprehensive benchmark for the systematic evaluation of multi-agent collaboration, single-LLM, and conventional approaches. MedAgentBoard encompasses four diverse medical task categories: (1) medical (visual) question answering, (2) lay summary generation, (3) structured Electronic Health Record (EHR) predictive modeling, and (4) clinical workflow automation, across text, medical images, and structured EHR data. Our extensive experiments reveal a nuanced landscape: while multi-agent collaboration demonstrates benefits in specific scenarios, such as enhancing task completeness in clinical workflow automation, it does not consistently outperform advanced single LLMs (e.g., in textual medical QA) or, critically, specialized conventional methods that generally maintain better performance in tasks like medical VQA and EHR-based prediction. MedAgentBoard offers a vital resource and actionable insights, emphasizing the necessity of a task-specific, evidence-based approach to selecting and developing AI solutions in medicine. It underscores that the inherent complexity and overhead of multi-agent collaboration must be carefully weighed against tangible performance gains. All code, datasets, detailed prompts, and experimental results are open-sourced at this link. Yinghao Zhu, Ziyi He, Xichen Zhang, Liantao Ma, Lequan Yu |
NeurIPS | 9 |
| 2025 | Diff-UNet: A diffusion embedded network for robust 3D medical image segmentation
Zhaohu Xing, Huazhu Fu, Guang Yang 0006, Lequan Yu, Bai Ying Lei, Lei Zhu 0003 |
Medical Image Anal. | 6 |
| 2025 | Democratizing large language model-based graph data augmentation via latent knowledge graphsabstractData augmentation is necessary for graph representation learning due to the scarcity and noise present in graph data. Most of the existing augmentation methods overlook the context information inherited from the dataset as they rely solely on the graph structure for augmentation. Despite the success of some large language model-based (LLM) graph learning methods, they are mostly white-box which require access to the weights or latent features from the open-access LLMs, making them difficult to be democratized for everyone as the most advanced LLMs are often closed-source for commercial considerations. To overcome these limitations, we propose a black-box context-driven graph data augmentation approach, with the guidance of LLMs - DemoGraph. Leveraging the text prompt as context-related information, we task the LLM with generating knowledge graphs (KGs), which allow us to capture the structural interactions from the text outputs. We then design a dynamic merging schema to stochastically integrate the LLM-generated KGs into the original graph during training. To control the sparsity of the augmented graph, we further devise a granularity-aware prompting strategy and an instruction fine-tuning module, which seamlessly generates text prompts according to different granularity levels of the dataset. Extensive experiments on various graph learning tasks validate the effectiveness of our method over existing graph data augmentation methods. Notably, our approach excels in scenarios involving electronic health records (EHRs), which validates its maximal utilization of contextual knowledge, leading to enhanced predictive performance and interpretability. Yushi Feng, Tsai Hor Chan, Guosheng Yin, Lequan Yu |
Neural Networks | 4 |
| 2025 | Feature Preserving Shrinkage on Bayesian Neural Networks Via the R2D2 PriorabstractBayesian neural networks (BNNs) treat neural network weights as random variables, which aim to provide posterior uncertainty estimates and avoid overfitting by performing inference on the posterior weights. However, selection of appropriate prior distributions remains a challenging task, and BNNs may suffer from catastrophic inflated variance or poor predictive performance when poor choices are made for the priors. Existing BNN designs apply different priors to weights, while the behaviours of these priors make it difficult to sufficiently shrink noisy signals or they are prone to overshrinking important signals in the weights. To alleviate this problem, we propose a novel R2D2-Net, which imposes the $R^{2}$R2-induced Dirichlet Decomposition (R2D2) prior to the BNN weights. The R2D2-Net can effectively shrink irrelevant coefficients towards zero, while preventing key features from over-shrinkage. To approximate the posterior distribution of weights more accurately, we further propose a variational Gibbs inference algorithm that combines the Gibbs updating procedure and gradient-based optimization. This strategy enhances stability and consistency in estimation when the variational objective involving the shrinkage parameters is non-convex. We also analyze the evidence lower bound (ELBO) and the posterior concentration rates from a theoretical perspective. Experiments on both natural and medical image classification and uncertainty estimation tasks demonstrate satisfactory performances of our method. Tsai Hor Chan, Dora Yan Zhang, Guosheng Yin, Lequan Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Multi-Sensor Learning Enables Information Transfer Across Different Sensory Data and Augments Multi-Modality ImagingabstractMulti-modality imaging is widely used in clinical practice and biomedical research to gain a comprehensive understanding of an imaging subject. Currently, multi-modality imaging is accomplished by post hoc fusion of independently reconstructed images under the guidance of mutual information or spatially registered hardware, which limits the accuracy and utility of multi-modality imaging. Here, we investigate a data-driven multi-modality imaging (DMI) strategy for synergetic imaging of CT and MRI. We reveal two distinct types of features in multi-modality imaging, namely intra- and inter-modality features, and present a multi-sensor learning (MSL) framework to utilize the crossover inter-modality features for augmented multi-modality imaging. The MSL imaging approach breaks down the boundaries of traditional imaging modalities and allows for optimal hybridization of CT and MRI, which maximizes the use of sensory data. We showcase the effectiveness of our DMI strategy through synergetic CT-MRI brain imaging. The principle of DMI is quite general and holds enormous potential for various DMI applications across disciplines. Lingting Zhu, Lianli Liu, Lei Xing 0001, Lequan Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Vivim: A Video Vision Mamba for Ultrasound Video SegmentationabstractUltrasound video segmentation gains increasing attention in clinical practice due to the redundant dynamic references in video frames. However, traditional convolutional neural networks have a limited receptive field and transformer-based networks are unsatisfactory in constructing long-term dependency from the perspective of computational complexity. This bottleneck poses a significant challenge when processing longer sequences in medical video analysis tasks using available devices with limited memory. Recently, state space models (SSMs), famous by Mamba, have exhibited linear complexity and impressive achievements in efficient long sequence modeling, which have developed deep neural networks by expanding the receptive field on many vision tasks significantly. Unfortunately, vanilla SSMs failed to simultaneously capture causal temporal cues and preserve non-casual spatial information. To this end, this paper presents a Video Vision Mamba-based framework, dubbed as Vivim, for ultrasound video segmentation tasks. Our Vivim can effectively compress the long-term spatiotemporal representation into sequences at varying scales with our designed Temporal Mamba Block. We also introduce an improved boundary-aware affine constraint across frames to enhance the discriminative ability of Vivim on ambiguous lesions. Extensive experiments on thyroid segmentation in ultrasound videos, breast lesion segmentation in ultrasound videos, and polyp segmentation in colonoscopy videos demonstrate the effectiveness and efficiency of our Vivim, superior to existing methods. The code and dataset are available at: https://github.com/scott-yjyang/Vivim. Zhaohu Xing, Lequan Yu, Huazhu Fu, Chunwang Huang, Lei Zhu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Completed Feature Disentanglement Learning for Multimodal MRIs AnalysisabstractMultimodal MRIs play a crucial role in clinical diagnosis and treatment. Feature disentanglement (FD)-based methods, aiming at learning superior feature representations for multimodal data analysis, have achieved significant success in multimodal learning (MML). Typically, existing FD-based methods separate multimodal data into modality-shared and modality-specific features, and employ concatenation or attention mechanisms to integrate these features. However, our preliminary experiments indicate that these methods could lead to a loss of shared information among subsets of modalities when the inputs contain more than two modalities, and such information is critical for prediction accuracy. Furthermore, these methods do not adequately interpret the relationships between the decoupled features at the fusion stage. To address these limitations, we propose a novel Complete Feature Disentanglement (CFD) strategy that recovers the lost information during feature decoupling. Specifically, the CFD strategy not only identifies modality-shared and modality-specific features, but also decouples shared features among subsets of multimodal inputs, termed as modality-partial-shared features. We further introduce a new Dynamic Mixture-of-Experts Fusion (DMF) module that dynamically integrates these decoupled features, by explicitly learning the local-global relationships among the features. The effectiveness of our approach is validated through classification tasks on three multimodal MRI datasets. Extensive experimental results demonstrate that our approach outperforms other state-of-the-art MML methods with obvious margins, showcasing its superior performance. Tianling Liu, Hongying Liu 0001, Fanhua Shang, Lequan Yu, Tong Han |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Unleash the Power of State Space Model for Whole Slide Image With Local Aware Scanning and Importance ResamplingabstractWhole slide image (WSI) analysis is gaining prominence within the medical imaging field. However, previous methods often fall short of efficiently processing entire WSIs due to their gigapixel size. Inspired by recent developments in state space models, this paper introduces a new Pathology Mamba (PAM) for more accurate and robust WSI analysis. PAM includes three carefully designed components to tackle the challenges of enormous image size, the utilization of local and hierarchical information, and the mismatch between the feature distributions of training and testing during WSI analysis. Specifically, we design a Bi-directional Mamba Encoder to process the extensive patches present in WSIs effectively and efficiently, which can handle large-scale pathological images while achieving high performance and accuracy. To further harness the local information and inherent hierarchical structure of WSI, we introduce a novel Local-aware Scanning module, which employs a local-aware mechanism alongside hierarchical scanning to adeptly capture both the local information and the overarching structure within WSIs. Moreover, to alleviate the patch feature distribution misalignment between training and testing, we propose a Test-time Importance Resampling module to conduct testing patch resampling to ensure consistency of feature distribution between the training and testing phases, and thus enhance model prediction. Extensive evaluation on nine WSI datasets with cancer subtyping and survival prediction tasks demonstrates that PAM outperforms current state-of-the-art methods and also its enhanced capability in modeling discriminative areas within WSIs. The source code is available at https://github.com/HKU-MedAI/PAM. Yanyan Huang, Weiqin Zhao, Yu Fu 0008, Lingting Zhu, Lequan Yu |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Swin-UMamba†: Adapting Mamba-Based Vision Foundation Models for Medical Image SegmentationabstractVision foundation models have shown great potential in improving generalizability and data efficiency, especially for medical image segmentation since medical image datasets are relatively small due to high annotation costs and privacy concerns. However, current research on foundation models predominantly relies on transformers. The high quadratic complexity and large parameter counts make these models computationally expensive, limiting their potential for clinical applications. In this work, we introduce Swin-UMamba†, a novel Mamba-based model for medical image segmentation that seamlessly leverages the power of the vision foundation model, which is also computationally efficient with the linear complexity of Mamba. Moreover, we investigated and verified the impact of the vision foundation model on medical image segmentation, in which a self-supervised model adaptation scheme was designed to bridge the gap between natural and medical data. Notably, Swin-UMamba† outperforms 7 state-of-the-art methods, including CNN-based, transformer-based, and Mamba-based approaches across AbdomenMRI, Encoscopy, and Microscopy datasets. The code and models are publicly available at: https://github.com/JiarunLiu/Swin-UMamba. Jiarun Liu, Hao Yang 0026, Lequan Yu, Yong Liang 0001, Yizhou Yu, Shaoting Zhang 0001, Hairong Zheng, Shanshan Wang 0002 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | TAD-Graph: Enhancing Whole Slide Image Analysis via Task-Aware Subgraph DisentanglementabstractLearning contextual features such as interactions among various biological entities is vital for whole slide images (WSI)-based cancer diagnosis and prognosis. Graph-based methods have surpassed traditional multi-instance learning in WSI analysis by robustly integrating local pathological and contextual interaction features. However, the high resolution of WSIs often leads to large, noisy graphs. This can result in shortcut learning and overfitting due to the disproportionate graph size relative to WSI datasets. To overcome these issues, we propose a novel Task-Aware Disentanglement Graph approach (TAD-Graph) for more efficient WSI analysis. TAD-Graph operates on WSI graph representations, effectively identifying and disentangling informative subgraphs to enhance contextual feature extraction. Specifically, we inject stochasticity into the edge connections of the WSI graph and separate the WSI graph into task-relevant and task-irrelevant subgraphs. The disentanglement procedure is optimized using a graph information bottleneck-based objective, with added constraints on the task-irrelevant subgraph to reduce spurious correlations from task-relevant subgraphs to labels. TAD-Graph outperforms existing methods in three WSI analysis tasks across six benchmark datasets. Furthermore, our analysis using pathological concept-based metrics demonstrates TAD-Graph's ability to not only improve predictive accuracy but also provide interpretive insights and aid in potential biomarker identification. Our code is publicly available at https://github.com/fuying-wang/TAD-Graph. Fuying Wang, Jiayi Xin, Weiqin Zhao, Yuming Jiang 0005, Maximus C. F. Yeung, Liansheng Wang 0002, Lequan Yu |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Scaling Chest X-Ray Foundation Models From Mixed Supervisions for Dense PredictionabstractFoundation models have significantly revolutionized the field of chest X-ray diagnosis with their ability to transfer across various diseases and tasks. However, previous works have predominantly utilized self-supervised learning from medical image-text pairs, which falls short in dense medical prediction tasks due to their sole reliance on such coarse pair supervision, thereby limiting their applicability to detailed diagnostics. In this paper, we introduce a Dense Chest X-ray Foundation Model (DCXFM), which utilizes mixed supervision types (i.e., text, label, and segmentation masks) to significantly enhance the scalability of foundation models across various medical tasks. Our model involves two training stages: we first employ a novel self-distilled multimodal pretraining paradigm to exploit text and label supervision, along with local-to-global self-distillation and soft cross-modal contrastive alignment strategies to enhance localization capabilities. Subsequently, we introduce an efficient cost aggregation module, comprising spatial and class aggregation mechanisms, to further advance dense prediction tasks with densely annotated datasets. Comprehensive evaluations on three tasks (phrase grounding, zero-shot semantic segmentation, and zero-shot classification) demonstrate DCXFM's superior performance over other state-of-the-art medical image-text pretraining models. Remarkably, DCXFM exhibits powerful zero-shot capabilities across various datasets in phrase grounding and zero-shot semantic segmentation, underscoring its superior generalization in dense prediction tasks. Fuying Wang, Lequan Yu |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Boosting Multiple Instance Learning Models for Whole Slide Image Classification: A Model-Agnostic Framework Based on Counterfactual InferenceabstractMultiple instance learning is an effective paradigm for whole slide image (WSI) classification, where labels are only provided at the bag level. However, instance-level prediction is also crucial as it offers insights into fine-grained regions of interest. Existing multiple instance learning methods either solely focus on training a bag classifier or have the insufficient capability of exploring instance prediction. In this work, we propose a novel model-agnostic framework to boost existing multiple instance learning models, to improve the WSI classification performance in both bag and instance levels. Specifically, we propose a counterfactual inference-based sub-bag assessment method and a hierarchical instance searching strategy to help to search reliable instances and obtain their accurate pseudo labels. Furthermore, an instance classifier is well-trained to produce accurate predictions. The instance embedding it generates is treated as a prompt to refine the instance feature for bag prediction. This framework is model-agnostic, capable of adapting to existing multiple instance learning models, including those without specific mechanisms like attention. Extensive experiments on three datasets demonstrate the competitive performance of our method. Code will be available at https://github.com/centurion-crawler/CIMIL. Weiping Lin, Zhenfeng Zhuang, Lequan Yu, Liansheng Wang 0002 |
AAAI | 3 |
| 2024 | Memory-Efficient Prompt Tuning for Incremental Histopathology ClassificationabstractRecent studies have made remarkable progress in histopathology classification. Based on current successes, contemporary works proposed to further upgrade the model towards a more generalizable and robust direction through incrementally learning from the sequentially delivered domains. Unlike previous parameter isolation based approaches that usually demand massive computation resources during model updating, we present a memory-efficient prompt tuning framework to cultivate model generalization potential in economical memory cost. For each incoming domain, we reuse the existing parameters of the initial classification model and attach lightweight trainable prompts into it for customized tuning. Considering the domain heterogeneity, we perform decoupled prompt tuning, where we adopt a domain-specific prompt for each domain to independently investigate its distinctive characteristics, and one domain-invariant prompt shared across all domains to continually explore the common content embedding throughout time. All domain-specific prompts will be appended to the prompt bank and isolated from further changes to prevent forgetting the distinctive features of early-seen domains. While the domain-invariant prompt will be passed on and iteratively evolve by style-augmented prompt refining to improve model generalization capability over time. In specific, we construct a graph with existing prompts and build a style-augmented graph attention network to guide the domain-invariant prompt exploring the overlapped latent embedding among all delivered domains for more domain-generic representations. We have extensively evaluated our framework with two histopathology tasks, i.e., breast cancer metastasis classification and epithelium-stroma tissue classification, where our approach yielded superior performance and memory efficiency over the competing methods. Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
AAAI | 3 |
| 2024 | L2B: Learning to Bootstrap Robust Models for Combating Label NoiseabstractDeep neural networks have shown great success in representation learning. However, when learning with noisy labels (LNL), they can easily overfit and fail to generalize to new data. This paper introduces a simple and effective method, named Learning to Bootstrap (L2B), which enables models to bootstrap themselves using their own predictions without being adversely affected by erroneous pseudo-labels. It achieves this by dynamically adjusting the importance weight between real observed and generated labels, as well as between different samples through metalearning. Unlike existing instance reweighting methods, the key to our method lies in a new, versatile objective that enables implicit relabeling concurrently, leading to significant improvements without incurring additional costs. L2B offers several benefits over the baseline methods. It yields more robust models that are less susceptible to the impact of noisy labels by guiding the bootstrapping procedure more effectively. It better exploits the valuable information contained in corrupted instances by adapting the weights of both instances and labels. Furthermore, L2B is compatible with existing LNL methods and delivers competitive results spanning natural and medical imaging tasks including classification and segmentation under both synthetic and real-world noise. Extensive experiments demonstrate that our method effectively mitigates the challenges of noisy labels, often necessitating few to no validation samples, and is well generalized to other tasks such as image segmentation. This not only positions it as a robust complement to existing LNL techniques but also underscores its practical applicability. The code and models are available at https://github.com/yuyinzhou/12b. Yuyin Zhou, Xianhang Li, Fengze Liu, Qingyue Wei, Xuxi Chen, Lequan Yu, Cihang Xie, Matthew P. Lungren, Lei Xing 0001 |
CVPR | 6 |
| 2024 | cDP-MIL: Robust Multiple Instance Learning via Cascaded Dirichlet Process
Yihang Chen 0001, Tsai Hor Chan, Guosheng Yin, Yuming Jiang 0005, Lequan Yu |
ECCV (54) | 5 |
| 2024 | HERGen: Elevating Radiology Report Generation with Longitudinal Data
Fuying Wang, Shenghui Du, Lequan Yu |
ECCV (55) | 3 |
| 2024 | ORCGT: Ollivier-Ricci Curvature-Based Graph Model for Lung STAS Prediction
Min Cen, Zheng Wang 0077, Zhenfeng Zhuang, Zhen Bao, Weiwei Wei, Baptiste Magnier, Lequan Yu, Liansheng Wang 0002 |
MICCAI (5) | 9 |
| 2024 | Swin-UMamba: Mamba-Based UNet with ImageNet-Based Pretraining
Jiarun Liu, Hao Yang 0026, Yan Xi, Lequan Yu, Cheng Li 0008, Yong Liang 0001, Guangming Shi, Yizhou Yu, Shaoting Zhang 0001, Hairong Zheng, Shanshan Wang 0002 |
MICCAI (9) | 5 |
| 2024 | Advancing H&E-to-IHC Virtual Staining with Task-Specific Domain Knowledge for HER2 Scoring
Qiong Peng, Weiping Lin, Yihuang Hu, Ailisi Bao, Chenyu Lian, Weiwei Wei, Jingxin Liu 0005, Lequan Yu, Liansheng Wang 0002 |
MICCAI (4) | 9 |
| 2024 | Free Lunch in Pathology Foundation Model: Task-specific Model Adaptation with Concept-Guided Feature EnhancementabstractWhole slide image (WSI) analysis is gaining prominence within the medical imaging field. Recent advances in pathology foundation models have shown the potential to extract powerful feature representations from WSIs for downstream tasks. However, these foundation models are usually designed for general-purpose pathology image analysis and may not be optimal for specific downstream tasks or cancer types. In this work, we present Concept Anchor-guided Task-specific Feature Enhancement (CATE), an adaptable paradigm that can boost the expressivity and discriminativeness of pathology foundation models for specific downstream tasks. Based on a set of task-specific concepts derived from the pathology vision-language model with expert-designed prompts, we introduce two interconnected modules to dynamically calibrate the generic image features extracted by foundation models for certain tasks or cancer types. Specifically, we design a Concept-guided Information Bottleneck module to enhance task-relevant characteristics by maximizing the mutual information between image features and concept anchors while suppressing superfluous information. Moreover, a Concept-Feature Interference module is proposed to utilize the similarity between calibrated features and concept anchors to further generate discriminative task-specific features. The extensive experiments on public WSI datasets demonstrate that CATE significantly enhances the performance and generalizability of MIL models. Additionally, heatmap and umap visualization results also reveal the effectiveness and interpretability of CATE. Yanyan Huang, Weiqin Zhao, Yihang Chen 0001, Yu Fu 0008, Lequan Yu |
NeurIPS | 5 |
| 2024 | FetusMapV2: Enhanced fetal pose estimation in 3D ultrasound
Chaoyu Chen, Xin Yang 0009, Yuhao Huang 0001, Wenlong Shi, Yan Cao 0002, Mingyuan Luo, Xindi Hu, Lei Zhu 0003, Lequan Yu, Kejuan Yue, Yuanji Zhang, Yi Xiong 0001, Dong Ni 0001, Weijun Huang |
Medical Image Anal. | 9 |
| 2024 | MPGAN: Multi Pareto Generative Adversarial Network for the denoising and quantitative analysis of low-dose PET images of human brain
Yu Fu 0008, Shunjie Dong, Yanyan Huang, Meng Niu, Chao Ni 0010, Lequan Yu, Kuangyu Shi, Zhijun Yao, Cheng Zhuo |
Medical Image Anal. | 6 |
| 2024 | Multi-task heterogeneous graph learning on electronic health records
Tsai Hor Chan, Guosheng Yin, Kyongtae Bae, Lequan Yu |
Neural Networks | 4 |
| 2024 | Hybrid Masked Image Modeling for 3D Medical Image SegmentationabstractMasked image modeling (MIM) with transformer backbones has recently been exploited as a powerful self-supervised pre-training technique. The existing MIM methods adopt the strategy to mask random patches of the image and reconstruct the missing pixels, which only considers semantic information at a lower level, and causes a long pre-training time. This paper presents HybridMIM, a novel hybrid self-supervised learning method based on masked image modeling for 3D medical image segmentation. Specifically, we design a two-level masking hierarchy to specify which and how patches in sub-volumes are masked, effectively providing the constraints of higher level semantic information. Then we learn the semantic information of medical images at three levels, including: 1) partial region prediction to reconstruct key contents of the 3D image, which largely reduces the pre-training time burden (pixel-level); 2) patch-masking perception to learn the spatial relationship between the patches in each sub-volume (region-level); and 3) drop-out-based contrastive learning between samples within a mini-batch, which further improves the generalization ability of the framework (sample-level). The proposed framework is versatile to support both CNN and transformer as encoder backbones, and also enables to pre-train decoders for image segmentation. We conduct comprehensive experiments on five widely-used public medical image segmentation datasets, including BraTS2020, BTCV, MSD Liver, MSD Spleen, and BraTS2023. The experimental results show the clear superiority of HybridMIM against competing supervised methods, masked pre-training approaches, and other self-supervised methods, in terms of quantitative metrics, speed performance and qualitative observations. Zhaohu Xing, Lei Zhu 0003, Lequan Yu, Zhiheng Xing |
IEEE J. Biomed. Health Informatics | 3 |
| 2024 | Synthesizing Feature-Aligned and Category-Aware Electronic Medical Records for Intracranial Aneurysm Rupture PredictionabstractRupture prediction is crucial for precise treatment and follow-up management of patients with intracranial aneurysms (IAs). Considerable machine learning (ML) methods have been proposed to improve rupture prediction by leveraging electronic medical records (EMRs), however, data scarcity and category imbalance strongly influence performance. Thus, we propose a novel data synthesis method i.e., Transformer-based conditional GAN (TransCGAN), to synthesize highly authentic and category-aware EMRs to address above challenges. Specifically, we first align feature-wise context relationship and distribution between synthetic and original data to enhance synthetic data quality. To achieve this, we first integrate the Transformer structure into GAN to match the contextual relationship by processing the long-range dependencies among clinical factors and introduce a statistical loss to maintain distributional consistency by constraining the mean and variance of the synthesis features. Additionally, a conditional module is designed to assign the category of the synthesis data, thereby addressing the challenge of category imbalance. Subsequently, the synthetic data are merged with the original data to form a large-scale and category-balanced training dataset for IAs rupture prediction. Experimental results show that using TransCGAN's synthetic data enhances classifier performance, achieving AUC of 0.89 and outperforming state-of-the-art resampling methods by 5-33 in F1 score. Qian Yang 0005, Caizi Li, Chubin Ou, Kang Li 0007, Xiangyun Liao, Chuanzhi Duan, Lequan Yu, Weixin Si |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | A Dual Enrichment Synergistic Strategy to Handle Data Heterogeneity for Domain Incremental Cardiac SegmentationabstractUpon remarkable progress in cardiac image segmentation, contemporary studies dedicate to further upgrading model functionality toward perfection, through progressively exploring the sequentially delivered datasets over time by domain incremental learning. Existing works mainly concentrated on addressing the heterogeneous style variations, but overlooked the critical shape variations across domains hidden behind the sub-disease composition discrepancy. In case the updated model catastrophically forgets the sub-diseases that were learned in past domains but are no longer present in the subsequent domains, we proposed a dual enrichment synergistic strategy to incrementally broaden model competence for a growing number of sub-diseases. The data-enriched scheme aims to diversify the shape composition of current training data via displacement-aware shape encoding and decoding, to gradually build up the robustness against cross-domain shape variations. Meanwhile, the model-enriched scheme intends to strengthen model capabilities by progressively appending and consolidating the latest expertise into a dynamically-expanded multi-expert network, to gradually cultivate the generalization ability over style-variated domains. The above two schemes work in synergy to collaboratively upgrade model capabilities in two-pronged manners. We have extensively evaluated our network with the ACDC and M&Ms datasets in single-domain and compound-domain incremental learning settings. Our approach outperformed other competing methods and achieved comparable results to the upper bound. Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2024 | RECIST-Induced Reliable Learning: Geometry-Driven Label Propagation for Universal Lesion SegmentationabstractAutomatic universal lesion segmentation (ULS) from Computed Tomography (CT) images can ease the burden of radiologists and provide a more accurate assessment than the current Response Evaluation Criteria In Solid Tumors (RECIST) guideline measurement. However, this task is underdeveloped due to the absence of large-scale pixel-wise labeled data. This paper presents a weakly-supervised learning framework to utilize the large-scale existing lesion databases in hospital Picture Archiving and Communication Systems (PACS) for ULS. Unlike previous methods to construct pseudo surrogate masks for fully supervised training through shallow interactive segmentation techniques, we propose to unearth the implicit information from RECIST annotations and thus design a unified RECIST-induced reliable learning (RiRL) framework. Particularly, we introduce a novel label generation procedure and an on-the-fly soft label propagation strategy to avoid noisy training and poor generalization problems. The former, named RECIST-induced geometric labeling, uses clinical characteristics of RECIST to preliminarily and reliably propagate the label. With the labeling process, a trimap divides the lesion slices into three regions, including certain foreground, background, and unclear regions, which consequently enables a strong and reliable supervision signal on a wide region. A topological knowledge-driven graph is built to conduct the on-the-fly label propagation for the optimal segmentation boundary to further optimize the segmentation boundary. Experimental results on a public benchmark dataset demonstrate that the proposed method surpasses the SOTA RECIST-based ULS methods by a large margin. Our approach surpasses SOTA approaches over 2.0%, 1.5%, 1.4%, and 1.6% Dice with ResNet101, ResNet50, HRNet, and ResNest50 backbones. Lianyu Zhou, Lequan Yu, Liansheng Wang 0002 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Self-Mining the Confident Prototypes for Source-Free Unsupervised Domain Adaptation in Image SegmentationabstractThis paper studies a practical Source-free unsupervised domain adaptation (SFUDA) problem, which transfers knowledge of source-trained models to the target domain, without accessing the source data. It has received increasing attention in recent years, while the prior arts focus on designing adaptation strategies, ignoring that different target samples exhibit different transfer abilities on the source model. Additionally, we observe pixel-wise class prediction is typically accompanied by ambiguity issue, i.e., prediction errors often occur between several confusing classes. In this study, we propose a dual-branch collaborative learning framework that aims to achieve reliable knowledge transfer from important samples to the rest by fully mining confident prototypes in the target data. Concretely, we first partition the target data into confident samples and uncertain samples via a new class-ranking reliability score and then utilize the latent features from the confident branch as guidance to promote the learning of the uncertain branch. For ambiguity issue, we propose a feature relabelling module, which exploits reliable prototypes in the mini-batch as well as in the target data to refine labels of uncertain features. We further deploy the proposed framework to commonly used CNN and state-of-the-art Transformer architectures and reveal the potential to promote the generalization ability of backbone models. Experimental results on both natural and medical benchmark datasets verify that our proposed approach exceeds state-of-the-art SFUDA methods with large margins, and achieves comparable performance to existing UDA methods. Yuntong Tian, Huazhu Fu, Lei Zhu 0003, Lequan Yu |
IEEE Trans. Multim. | 5 |
| 2023 | MulGT: Multi-Task Graph-Transformer with Task-Aware Knowledge Injection and Domain Knowledge-Driven Pooling for Whole Slide Image AnalysisabstractWhole slide image (WSI) has been widely used to assist automated diagnosis under the deep learning fields. However, most previous works only discuss the SINGLE task setting which is not aligned with real clinical setting, where pathologists often conduct multiple diagnosis tasks simultaneously. Also, it is commonly recognized that the multi-task learning paradigm can improve learning efficiency by exploiting commonalities and differences across multiple tasks. To this end, we present a novel multi-task framework (i.e., MulGT) for WSI analysis by the specially designed Graph-Transformer equipped with Task-aware Knowledge Injection and Domain Knowledge-driven Graph Pooling modules. Basically, with the Graph Neural Network and Transformer as the building commons, our framework is able to learn task-agnostic low-level local information as well as task-specific high-level global representation. Considering that different tasks in WSI analysis depend on different features and properties, we also design a novel Task-aware Knowledge Injection module to transfer the task-shared graph embedding into task-specific feature spaces to learn more accurate representation for different tasks. Further, we elaborately design a novel Domain Knowledge-driven Graph Pooling module for each task to improve both the accuracy and robustness of different tasks by leveraging different diagnosis patterns of multiple tasks. We evaluated our method on two public WSI datasets from TCGA projects, i.e., esophageal carcinoma and kidney carcinoma. Experimental results show that our method outperforms single-task counterparts and the state-of-theart methods on both tumor typing and staging tasks. Weiqin Zhao, Maximus C. F. Yeung, Tianye Niu, Lequan Yu |
AAAI | 5 |
| 2023 | Histopathology Whole Slide Image Analysis with Heterogeneous Graph Representation LearningabstractGraph-based methods have been extensively applied to whole slide histopathology image (WSI) analysis due to the advantage of modeling the spatial relationships among different entities. However, most of the existing methods focus on modeling WSIs with homogeneous graphs (e.g., with homogeneous node type). Despite their successes, these works are incapable of mining the complex structural relations between biological entities (e.g., the diverse interaction among different cell types) in the WSI. We propose a novel heterogeneous graph-based framework to leverage the inter-relationships among different types of nuclei for WSI analysis. Specifically, we formulate the WSI as a heterogeneous graph with “nucleus-type” attribute to each node and a semantic similarity attribute to each edge. We then present a new heterogeneous-graph edge attribute transformer (HEAT) to take advantage of the edge and node heterogeneity during massage aggregating. Further, we design a new pseudo-label-based semantic-consistent pooling mechanism to obtain graph-level features, which can mitigate the over-parameterization issue of conventional cluster-based pooling. Additionally, observing the limitations of existing association-based localization methods, we propose a causal-driven approach attributing the contribution of each node to improve the interpretability of our framework. Extensive experiments on three public TCGA benchmark datasets demonstrate that our frame-work outperforms the state-of-the-art methods with considerable margins on various tasks. Our codes are available at https://github.com/HKU-MedAI/WSI-HGNN. Tsai Hor Chan, Fernando Julio Cendra, Guosheng Yin, Lequan Yu |
CVPR | 5 |
| 2023 | MagicNet: Semi-Supervised Multi-Organ Segmentation via Magic-Cube Partition and RecoveryabstractWe propose a novel teacher-student model for semi-supervised multi-organ segmentation. In teacher-student model, data augmentation is usually adopted on unlabeled data to regularize the consistent training between teacher and student. We start from a key perspective that fixed relative locationsand variable sizes of different organs can provide distribution information where a multi-organ CT scan is drawn. Thus, we treat the prior anatomy as a strong tool to guide the data augmentation and reduce the mismatch between labeled and unlabeled images for semi-supervised learning. More specifically, we propose a data augmentation strategy based on partition-and-recovery N3cubes cross-and within-labeled and unlabeled images. Our strategy encourages unlabeled images to learn organ semantics in relative locations from the labeled images (cross-branch) and enhances the learning ability for small organs (within-branch). For within-branch, we further propose to refine the quality of pseudo labels by blending the learned representations from small cubes to incorporate local attributes. Our method is termed as MagicNet, since it treats the CT volume as a magic-cube and N3-cube partition-and-recovery process matches with the rule of playing a magic-cube. Extensive experiments on two public CT multi-organ datasets demonstrate the effectiveness of MagicNet, and noticeably outperforms state-of-the-art semi-supervised medical image segmentation approaches, with + 7% DSC improvement on MACT dataset with 10% labeled images. Code is avaiable at https://github.com/DeepMed-Lab-ECNU/MagicNet. Duowen Chen 0002, Yunhao Bai, Wei Shen 0002, Qingli Li, Lequan Yu, Yan Wang 0033 |
CVPR | 5 |
| 2023 | Taming Diffusion Models for Audio-Driven Co-Speech Gesture GenerationabstractAnimating virtual avatars to make co-speech gestures facilitates various applications in human-machine interaction. The existing methods mainly rely on generative adversarial networks (GANs), which typically suffer from notorious mode collapse and unstable training, thus making it difficult to learn accurate audio-gesture joint distributions. In this work, we propose a novel diffusion-based framework, named Diffusion Co-Speech Gesture (DiffGesture), to effectively capture the cross-modal audio-to-gesture associations and preserve temporal coherence for high-fidelity audio-driven co-speech gesture generation. Specifically, we first establish the diffusion-conditional generation process on clips of skeleton sequences and audio to enable the whole framework. Then, a novel Diffusion Audio-Gesture Transformer is devised to better attend to the information from multiple modalities and model the long-term temporal dependency. Moreover, to eliminate temporal inconsistency, we propose an effective Diffusion Gesture Stabilizer with an annealed noise sampling strategy. Benefiting from the architectural advantages of diffusion models, we further incorporate implicit classifier-free guidance to trade off between diversity and gesture quality. Extensive experiments demonstrate that DiffGesture achieves state-of-the-art performance, which renders coherent gestures with better mode coverage and stronger audio correlations. Code is available at https://github.com/Advocate99/DiffGesture. Lingting Zhu, Rui Qian 0001, Ziwei Liu 0002, Lequan Yu |
CVPR | 6 |
| 2023 | ConSlide: Asynchronous Hierarchical Interaction Transformer with Breakup-Reorganize Rehearsal for Continual Whole Slide Image AnalysisabstractWhole slide image (WSI) analysis has become increasingly important in the medical imaging community, enabling automated and objective diagnosis, prognosis, and therapeutic-response prediction. However, in clinical practice, the ever-evolving environment hamper the utility of WSI analysis models. In this paper, we propose the FIRST continual learning framework for WSI analysis, named ConSlide, to tackle the challenges of enormous image size, utilization of hierarchical structure, and catastrophic forgetting by progressive model updating on multiple sequential datasets. Our framework contains three key components. The Hierarchical Interaction Transformer (HIT) is proposed to model and utilize the hierarchical structural knowledge of WSI. The Breakup-Reorganize (BuRo) rehearsal method is developed for WSI data replay with efficient region storing buffer and WSI reorganizing operation. The asynchronous updating mechanism is devised to encourage the network to learn generic and specific knowledge respectively during the replay stage, based on a nested cross-scale similarity learning (CSSL) module. We evaluated the proposed ConSlide on four public WSI datasets from TCGA projects. It performs best over other state-of-the-art methods with a fair WSI-based continual learning setting and achieves a better trade-off of the overall performance and forgetting on previous tasks. Yanyan Huang, Weiqin Zhao, Yu Fu 0008, Yuming Jiang 0005, Lequan Yu |
ICCV | 6 |
| 2023 | HIGT: Hierarchical Interaction Graph-Transformer for Whole Slide Image Analysis
Weiqin Zhao, Lequan Yu |
MICCAI (6) | 4 |
| 2023 | Multi-scope Analysis Driven Hierarchical Graph Transformer for Whole Slide Image Based Cancer Survival Prediction
Wentai Hou, Bingjian Yao, Lequan Yu, Rongshan Yu, Feng Gao 0023, Liansheng Wang 0002 |
MICCAI (6) | 4 |
| 2023 | Consistency-Guided Meta-learning for Bootstrapping Semi-supervised Medical Image Segmentation
Qingyue Wei, Lequan Yu, Xianhang Li, Wei Shao 0008, Cihang Xie, Lei Xing 0001, Yuyin Zhou |
MICCAI (4) | 2 |
| 2023 | Cross-View Deformable Transformer for Non-displaced Hip Fracture Classification from Frontal-Lateral X-Ray Pair
Zhonghang Zhu, Qichang Chen, Lequan Yu, Lianxin Wang, Baptiste Magnier, Liansheng Wang 0002 |
MICCAI (6) | 3 |
| 2023 | Make-A-Volume: Leveraging Latent Diffusion Models for Cross-Modality 3D Brain MRI Synthesis
Lingting Zhu, Zeyue Xue, Zhenchao Jin, Jingzhen He, Ziwei Liu 0002, Lequan Yu |
MICCAI (10) | 7 |
| 2023 | Adaptive Uncertainty Estimation via High-Dimensional Testing on Latent RepresentationsabstractUncertainty estimation aims to evaluate the confidence of a trained deep neural network. However, existing uncertainty estimation approaches rely on low-dimensional distributional assumptions and thus suffer from the high dimensionality of latent features. Existing approaches tend to focus on uncertainty on discrete classification probabilities, which leads to poor generalizability to uncertainty estimation for other tasks. Moreover, most of the literature requires seeing the out-of-distribution (OOD) data in the training for better estimation of uncertainty, which limits the uncertainty estimation performance in practice because the OOD data are typically unseen. To overcome these limitations, we propose a new framework using data-adaptive high-dimensional hypothesis testing for uncertainty estimation, which leverages the statistical properties of the feature representations. Our method directly operates on latent representations and thus does not require retraining the feature encoder under a modified objective. The test statistic relaxes the feature distribution assumptions to high dimensionality, and it is more discriminative to uncertainties in the latent representations. We demonstrate that encoding features with Bayesian neural networks can enhance testing performance and lead to more accurate uncertainty estimation. We further introduce a family-wise testing procedure to determine the optimal threshold of OOD detection, which minimizes the false discovery rate (FDR). Extensive experiments validate the satisfactory performance of our framework on uncertainty estimation and task-specific prediction over a variety of competitors. The experiments on the OOD detection task also show satisfactory performance of our method when the OOD data are unseen in the training. Codes are available at https://github.com/HKU-MedAI/bnn_uncertainty. Tsai Hor Chan, Kin Wai Lau, Guosheng Yin, Lequan Yu |
NeurIPS | 5 |
| 2023 | IDRNet: Intervention-Driven Relation Network for Semantic SegmentationabstractCo-occurrent visual patterns suggest that pixel relation modeling facilitates dense prediction tasks, which inspires the development of numerous context modeling paradigms, \emph{e.g.}, multi-scale-driven and similarity-driven context schemes. Despite the impressive results, these existing paradigms often suffer from inadequate or ineffective contextual information aggregation due to reliance on large amounts of predetermined priors. To alleviate the issues, we propose a novel \textbf{I}ntervention-\textbf{D}riven \textbf{R}elation \textbf{Net}work (\textbf{IDRNet}), which leverages a deletion diagnostics procedure to guide the modeling of contextual relations among different pixels. Specifically, we first group pixel-level representations into semantic-level representations with the guidance of pseudo labels and further improve the distinguishability of the grouped representations with a feature enhancement module. Next, a deletion diagnostics procedure is conducted to model relations of these semantic-level representations via perceiving the network outputs and the extracted relations are utilized to guide the semantic-level representations to interact with each other. Finally, the interacted representations are utilized to augment original pixel-level representations for final predictions. Extensive experiments are conducted to validate the effectiveness of IDRNet quantitatively and qualitatively. Notably, our intervention-driven context scheme brings consistent performance improvements to state-of-the-art segmentation frameworks and achieves competitive results on popular benchmark datasets, including ADE20K, COCO-Stuff, PASCAL-Context, LIP, and Cityscapes. Zhenchao Jin, Xiaowei Hu 0001, Lingting Zhu, Luchuan Song, Lequan Yu |
NeurIPS | 6 |
| 2023 | Adaptive Region-Specific Loss for Improved Medical Image SegmentationabstractDefining the loss function is an important part of neural network design and critically determines the success of deep learning modeling. A significant shortcoming of the conventional loss functions is that they weight all regions in the input image volume equally, despite the fact that the system is known to be heterogeneous (i.e., some regions can achieve high prediction performance more easily than others). Here, we introduce a region-specific loss to lift the implicit assumption of homogeneous weighting for better learning. We divide the entire volume into multiple sub-regions, each with an individualized loss constructed for optimal local performance. Effectively, this scheme imposes higher weightings on the sub-regions that are more difficult to segment, and vice versa. Furthermore, the regional false positive and false negative errors are computed for each input image during a training step and the regional penalty is adjusted accordingly to enhance the overall accuracy of the prediction. Using different public and in-house medical image datasets, we demonstrate that the proposed regionally adaptive loss paradigm outperforms conventional methods in the multi-organ segmentations, without any modification to the neural network architecture or additional data preparation. Lequan Yu, Jen-Yeu Wang, Neil Panjwani, Jean-Pierre Obeid, Wu Liu 0003, Lianli Liu, Nataliya Kovalchuk, Michael Gensheimer 0001, Lucas Kas Vitzthum, Beth M. Beadle, Daniel T. Chang, Quynh-Thu Le, Lei Xing 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | MCIBI++: Soft Mining Contextual Information Beyond Image for Semantic SegmentationabstractCo-occurrent visual pattern makes context aggregation become an essential paradigm for semantic segmentation. The existing studies focus on modeling the contexts within image while neglecting the valuable semantics of the corresponding category beyond image. To this end, we propose a novel soft mining contextual information beyond image paradigm named MCIBI++ to further boost the pixel-level representations. Specifically, we first set up a dynamically updated memory module to store the dataset-level distribution information of various categories and then leverage the information to yield the dataset-level category representations during network forward. After that, we generate a class probability distribution for each pixel representation and conduct the dataset-level context aggregation with the class probability distribution as weights. Finally, the original pixel representations are augmented with the aggregated dataset-level and the conventional image-level contextual information. Moreover, in the inference phase, we additionally design a coarse-to-fine iterative inference strategy to further boost the segmentation results. MCIBI++ can be effortlessly incorporated into the existing segmentation frameworks and bring consistent performance improvements. Also, MCIBI++ can be extended into the video semantic segmentation framework with considerable improvements over the baseline. Equipped with MCIBI++, we achieved the state-of-the-art performance on seven challenging image or video semantic segmentation benchmarks. Zhenchao Jin, Dongdong Yu, Zehuan Yuan, Lequan Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | nnFormer: Volumetric Medical Image Segmentation via a 3D TransformerabstractTransformer, the model of choice for natural language processing, has drawn scant attention from the medical imaging community. Given the ability to exploit long-term dependencies, transformers are promising to help atypical convolutional neural networks to learn more contextualized visual representations. However, most of recently proposed transformer-based segmentation approaches simply treated transformers as assisted modules to help encode global context into convolutional representations. To address this issue, we introduce nnFormer (i.e., not-another transFormer), a 3D transformer for volumetric medical image segmentation. nnFormer not only exploits the combination of interleaved convolution and self-attention operations, but also introduces local and global volume-based self-attention mechanism to learn volume representations. Moreover, nnFormer proposes to use skip attention to replace the traditional concatenation/summation operations in skip connections in U-Net like architecture. Experiments show that nnFormer significantly outperforms previous transformer-based counterparts by large margins on three public datasets. Compared to nnUNet, the most widely recognized convnet-based 3D medical segmentation model, nnFormer produces significantly lower HD95 and is much more computationally efficient. Furthermore, we show that nnFormer and nnUNet are highly complementary to each other in model ensembling. Codes and models of nnFormer are available at https://git.io/JSf3i. Jiansen Guo, Xiaoguang Han 0001, Lequan Yu, Liansheng Wang 0002, Yizhou Yu |
IEEE Trans. Image Process. | 5 |
| 2023 | ARR-GCN: Anatomy-Relation Reasoning Graph Convolutional Network for Automatic Fine-Grained Segmentation of Organ's Surgical AnatomyabstractAnatomical resection (AR) based on anatomical sub-regions is a promising method of precise surgical resection, which has been proven to improve long-term survival by reducing local recurrence. The fine-grained segmentation of an organ's surgical anatomy (FGS-OSA), i.e., segmenting an organ into multiple anatomic regions, is critical for localizing tumors in AR surgical planning. However, automatically obtaining FGS-OSA results in computer-aided methods faces the challenges of appearance ambiguities among sub-regions (i.e., inter-sub-region appearance ambiguities) caused by similar HU distributions in different sub-regions of an organ's surgical anatomy, invisible boundaries, and similarities between anatomical landmarks and other anatomical information. In this paper, we propose a novel fine-grained segmentation framework termed the "anatomic relation reasoning graph convolutional network" (ARR-GCN), which incorporates prior anatomic relations into the framework learning. In ARR-GCN, a graph is constructed based on the sub-regions to model the class and their relations. Further, to obtain discriminative initial node representations of graph space, a sub-region center module is designed. Most importantly, to explicitly learn the anatomic relations, the prior anatomic-relations among the sub-regions are encoded in the form of an adjacency matrix and embedded into the intermediate node representations to guide framework learning. The ARR-GCN was validated on two FGS-OSA tasks: i) liver segments segmentation, and ii) lung lobes segmentation. Experimental results on both tasks outperformed other state-of-the-art segmentation methods and yielded promising performances by ARR-GCN for suppressing ambiguities among sub-regions. Yinli Tian, Wenjian Qin, Ricardo Lambo, Meiyan Yue, Songhui Diao, Lequan Yu, Yaoqin Xie, Shuo Li 0001 |
IEEE J. Biomed. Health Informatics | 7 |
| 2023 | Hybrid Graph Convolutional Network With Online Masked Autoencoder for Robust Multimodal Cancer Survival PredictionabstractCancer survival prediction requires exploiting related multimodal information (e.g., pathological, clinical and genomic features, etc.) and it is even more challenging in clinical practices due to the incompleteness of patient's multimodal data. Furthermore, existing methods lack sufficient intra- and inter-modal interactions, and suffer from significant performance degradation caused by missing modalities. This manuscript proposes a novel hybrid graph convolutional network, entitled HGCN, which is equipped with an online masked autoencoder paradigm for robust multimodal cancer survival prediction. Particularly, we pioneer modeling the patient's multimodal data into flexible and interpretable multimodal graphs with modality-specific preprocessing. HGCN integrates the advantages of graph convolutional networks (GCNs) and a hypergraph convolutional network (HCN) through node message passing and a hyperedge mixing mechanism to facilitate intra-modal and inter-modal interactions between multimodal graphs. With HGCN, the potential for multimodal data to create more reliable predictions of patient's survival risk is dramatically increased compared to prior methods. Most importantly, to compensate for missing patient modalities in clinical scenarios, we incorporated an online masked autoencoder paradigm into HGCN, which can effectively capture intrinsic dependence between modalities and seamlessly generate missing hyperedges for model inference. Extensive experiments and analysis on six cancer cohorts from TCGA show that our method significantly outperforms the state-of-the-arts in both complete and missing modal settings. Our codes are made available at https://github.com/lin-lcx/HGCN. Wentai Hou, Chengxuan Lin, Lequan Yu, Harry Qin, Rongshan Yu, Liansheng Wang 0002 |
IEEE Trans. Medical Imaging | 3 |
| 2023 | Domain-Incremental Cardiac Image Segmentation With Style-Oriented Replay and Domain-Sensitive Feature WhiteningabstractContemporary methods have shown promising results on cardiac image segmentation, but merely in static learning, i.e., optimizing the network once for all, ignoring potential needs for model updating. In real-world scenarios, new data continues to be gathered from multiple institutions over time and new demands keep growing to pursue more satisfying performance. The desired model should incrementally learn from each incoming dataset and progressively update with improved functionality as time goes by. As the datasets sequentially delivered from multiple sites are normally heterogenous with domain discrepancy, each updated model should not catastrophically forget previously learned domains while well generalizing to currently arrived domains or even unseen domains. In medical scenarios, this is particularly challenging as accessing or storing past data is commonly not allowed due to data privacy. To this end, we propose a novel domain-incremental learning framework to recover past domain inputs first and then regularly replay them during model optimization. Particularly, we first present a style-oriented replay module to enable structure-realistic and memory-efficient reproduction of past data, and then incorporate the replayed past data to jointly optimize the model with current data to alleviate catastrophic forgetting. During optimization, we additionally perform domain-sensitive feature whitening to suppress model's dependency on features that are sensitive to domain changes (e.g., domain-distinctive style features) to assist domain-invariant feature exploration and gradually improve the generalization performance of the network. We have extensively evaluated our approach with the M&Ms Dataset in single-domain and compound-domain incremental learning settings. Our approach outperforms other comparison methods with less forgetting on past domains and better generalization on current domains and unseen domains. Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Data Discernment for Affordable Training in Medical Image SegmentationabstractCollecting sufficient high-quality training data for deep neural networks is often expensive or even unaffordable in medical image segmentation tasks. We thus propose to train the network by using external data that can be collected in a cheaper way, e.g., crowd-sourcing. We show that by data discernment, the network is able to mine valuable knowledge from external data, even though the data distribution is very different from that of the original (internal) data. We discern the external data by learning an importance weight for each of them, with the goal to enhance the contribution of informative external data to network updating, while suppressing the data that are 'useless' or even 'harmful'. An iterative algorithm that alternatively estimates the importance weight and updates the network is developed by formulating the data discernment as a constrained nonlinear programming problem. It estimates the importance weight according to the distribution discrepancy between the external data and the internal dataset, and imposes a constraint to drive the network to learn more effectively, compared with the network without using the external data. We evaluate the proposed algorithm on two tasks: abdominal CT image and cervical smear image segmentation, using totally 6 publicly available datasets. The effectiveness of the algorithm is demonstrated by extensive experiments. Source codes are available at: https://github.com/YouyiSong/Data-Discernment. Youyi Song, Lequan Yu, Bai Ying Lei, Kup-Sze Choi, Harry Qin |
IEEE Trans. Medical Imaging | 2 |
| 2023 | Shared-Specific Feature Learning With Bottleneck Fusion Transformer for Multi-Modal Whole Slide Image AnalysisabstractThe fusion of multi-modal medical data is essential to assist medical experts to make treatment decisions for precision medicine. For example, combining the whole slide histopathological images (WSIs) and tabular clinical data can more accurately predict the lymph node metastasis (LNM) of papillary thyroid carcinoma before surgery to avoid unnecessary lymph node resection. However, the huge-sized WSI provides much more high-dimensional information than low-dimensional tabular clinical data, making the information alignment challenging in the multi-modal WSI analysis tasks. This paper presents a novel transformer-guided multi-modal multi-instance learning framework to predict lymph node metastasis from both WSIs and tabular clinical data. We first propose an effective multi-instance grouping scheme, named siamese attention-based feature grouping (SAG), to group high-dimensional WSIs into representative low-dimensional feature embeddings for fusion. We then design a novel bottleneck shared-specific feature transfer module (BSFT) to explore the shared and specific features between different modalities, where a few learnable bottleneck tokens are utilized for knowledge transfer between modalities. Moreover, a modal adaptation and orthogonal projection scheme were incorporated to further encourage BSFT to learn shared and specific features from multi-modal data. Finally, the shared and specific features are dynamically aggregated via an attention mechanism for slide-level prediction. Experimental results on our collected lymph node metastasis dataset demonstrate the efficiency of our proposed components and our framework achieves the best performance with AUC (area under the curve) of 97.34%, outperforming the state-of-the-art methods by over 1.27%. Lequan Yu, Xuehong Liao, Liansheng Wang 0002 |
IEEE Trans. Medical Imaging | 2 |
| 2023 | MuRCL: Multi-Instance Reinforcement Contrastive Learning for Whole Slide Image ClassificationabstractMulti-instance learning (MIL) is widely adop- ted for automatic whole slide image (WSI) analysis and it usually consists of two stages, i.e., instance feature extraction and feature aggregation. However, due to the "weak supervision" of slide-level labels, the feature aggregation stage would suffer from severe over-fitting in training an effective MIL model. In this case, mining more information from limited slide-level data is pivotal to WSI analysis. Different from previous works on improving instance feature extraction, this paper investigates how to exploit the latent relationship of different instances (patches) to combat overfitting in MIL for more generalizable WSI classification. In particular, we propose a novel Multi-instance Rein- forcement Contrastive Learning framework (MuRCL) to deeply mine the inherent semantic relationships of different patches to advance WSI classification. Specifically, the proposed framework is first trained in a self-supervised manner and then finetuned with WSI slide-level labels. We formulate the first stage as a contrastive learning (CL) process, where positive/negative discriminative feature sets are constructed from the same patch-level feature bags of WSIs. To facilitate the CL training, we design a novel reinforcement learning-based agent to progressively update the selection of discriminative feature sets according to an online reward for slide-level feature aggregation. Then, we further update the model with labeled WSI data to regularize the learned features for the final WSI classification. Experimental results on three public WSI classification datasets (Camelyon16, TCGA-Lung and TCGA-Kidney) demonstrate that the proposed MuRCL outperforms state-of-the-art MIL models. In addition, MuRCL can achieve comparable performance to other state-of-the-art MIL models on TCGA-Esca dataset. Zhonghang Zhu, Lequan Yu, Rongshan Yu, Liansheng Wang 0002 |
IEEE Trans. Medical Imaging | 2 |
| 2022 | H^2-MIL: Exploring Hierarchical Representation with Heterogeneous Multiple Instance Learning for Whole Slide Image AnalysisabstractCurrent representation learning methods for whole slide image (WSI) with pyramidal resolutions are inherently homogeneous and flat, which cannot fully exploit the multiscale and heterogeneous diagnostic information of different structures for comprehensive analysis. This paper presents a novel graph neural network-based multiple instance learning framework (i.e., H^2-MIL) to learn hierarchical representation from a heterogeneous graph with different resolutions for WSI analysis. A heterogeneous graph with the “resolution” attribute is constructed to explicitly model the feature and spatial-scaling relationship of multi-resolution patches. We then design a novel resolution-aware attention convolution (RAConv) block to learn compact yet discriminative representation from the graph, which tackles the heterogeneity of node neighbors with different resolutions and yields more reliable message passing. More importantly, to explore the task-related structured information of WSI pyramid, we elaborately design a novel iterative hierarchical pooling (IHPool) module to progressively aggregate the heterogeneous graph based on scaling relationships of different nodes. We evaluated our method on two public WSI datasets from the TCGA project, i.e., esophageal cancer and kidney cancer. Experimental results show that our method clearly outperforms the state-of-the-art methods on both tumor typing and staging tasks. Wentai Hou, Lequan Yu, Chengxuan Lin, Helong Huang, Rongshan Yu, Harry Qin, Liansheng Wang 0002 |
AAAI | 2 |
| 2022 | ICBNet: Iterative Context-Boundary Feedback Network for Polyp SegmentationabstractAccurate polyp segmentation from colonoscopy images, which is critical to automatic colorectal cancer diagnosis, attracts increasing attentions in recent years. Most existing deep learning-based methods adopt the one-stage processing pipeline, by usually fusing features from different levels or employing boundary-related attention. In this paper, we propose an novel Iterative Context-Boundary feedback Network, namely ICBNet, for robust and accurate polyp segmentation. By mimicking the “from-Preliminary-to-Refined” working paradigm of doctors, ICBNet adopts an iterative feedback learning strategy. Differently from other feedback methods which only use the prediction mask as a guide for foreground features, ICBNet refines encoder features with contextual and boundary-aware details from the preliminary segmentation and boundary predictions, and conducts such strategy in an iterative manner to achieve progressive improvement. Moreover, a dual-branch iterative feedback unit (IFU) is developed to enhance features under the guidance of segmentation and boundary predictions to enable the iterative learning. Extensive experiments on five widely-used polyp segmentation datasets demonstrate that the proposed ICBNet can utilize progressive refinement to effectively address the challenges of large appearance variations and obscure boundaries, and hence achieves more accurate and robust results against the state-of-the-arts methods. Yefan Xiao, Zhihao Chen 0004, Lequan Yu, Lei Zhu 0003 |
BIBM | 4 |
| 2022 | CD2-pFed: Cyclic Distillation-guided Channel Decoupling for Model Personalization in Federated LearningabstractFederated learning (FL) is a distributed learning paradigm that enables multiple clients to collaboratively learn a shared global model. Despite the recent progress, it remains challenging to deal with heterogeneous data clients, as the discrepant data distributions usually prevent the global model from delivering good generalization ability on each participating client. In this paper, we propose CD2-pFed, a novel Cyclic Distillation-guided Channel Decoupling framework, to personalize the global model in FL, under various settings of data heterogeneity. Different from previous works which establish layer-wise personalization to overcome the non-IID data across different clients, we make the first attempt at channel-wise assignment for model personalization, referred to as channel decoupling. To further facilitate the collaboration between private and shared weights, we propose a novel cyclic distillation scheme to impose a consistent regularization between the local and global model representations during the federation. Guided by the cyclical distillation, our channel decoupling framework can deliver more accurate and generalized results for different kinds of heterogeneity, such as feature skew, label distribution skew, and concept shift. Comprehensive experiments on four benchmarks, including natural image and medical image analysis tasks, demonstrate the consistent effectiveness of our method on both local and external validations. Yiqing Shen 0003, Yuyin Zhou, Lequan Yu |
CVPR | 3 |
| 2022 | You Should Look at All Objects
Zhenchao Jin, Dongdong Yu, Luchuan Song, Zehuan Yuan, Lequan Yu |
ECCV (9) | 5 |
| 2022 | Spatial-Hierarchical Graph Neural Network with Dynamic Structure Learning for Histological Image Classification
Wentai Hou, Helong Huang, Qiong Peng, Rongshan Yu, Lequan Yu, Liansheng Wang 0002 |
MICCAI (2) | 5 |
| 2022 | Joint Prediction of Meningioma Grade and Brain Invasion via Task-Aware Contrastive Learning
Tianling Liu, Wennan Liu, Lequan Yu, Tong Han, Lei Zhu 0003 |
MICCAI (3) | 3 |
| 2022 | NestedFormer: Nested Modality-Aware Transformer for Brain Tumor Segmentation
Zhaohu Xing, Lequan Yu, Tong Han, Lei Zhu 0003 |
MICCAI (5) | 2 |
| 2022 | Reinforcement Learning Driven Intra-modal and Inter-modal Representation Learning for 3D Medical Image Classification
Zhonghang Zhu, Liansheng Wang 0002, Baptiste Magnier, Lei Zhu 0003, Lequan Yu |
MICCAI (3) | 6 |
| 2022 | Multi-Granularity Cross-modal Alignment for Generalized Medical Visual Representation LearningabstractLearning medical visual representations directly from paired radiology reports has become an emerging topic in representation learning. However, existing medical image-text joint learning methods are limited by instance or local supervision analysis, ignoring disease-level semantic correspondences. In this paper, we present a novel Multi-Granularity Cross-modal Alignment (MGCA) framework for generalized medical visual representation learning by harnessing the naturally exhibited semantic correspondences between medical image and radiology reports at three different levels, i.e., pathological region-level, instance-level, and disease-level. Specifically, we first incorporate the instance-wise alignment module by maximizing the agreement between image-report pairs. Further, for token-wise alignment, we introduce a bidirectional cross-attention strategy to explicitly learn the matching between fine-grained visual tokens and text tokens, followed by contrastive learning to align them. More important, to leverage the high-level inter-subject relationship semantic (e.g., disease) correspondences, we design a novel cross-modal disease-level alignment paradigm to enforce the cross-modal cluster assignment consistency. Extensive experimental results on seven downstream medical image datasets covering image classification, object detection, and semantic segmentation tasks demonstrate the stable and superior performance of our framework. Fuying Wang, Yuyin Zhou, Varut Vardhanabhuti, Lequan Yu |
NeurIPS | 5 |
| 2022 | STPD: Defending against ℓ0-norm attacks with space transformation
Jinlin Chen, Jiannong Cao 0001, Zhixuan Liang, Xiaohui Cui, Lequan Yu, Wei Li 0121 |
Future Gener. Comput. Syst. | 5 |
| 2022 | Towards reliable cardiac image segmentation: Assessing image-level and pixel-level segmentation quality via self-reflective references
Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
Medical Image Anal. | 2 |
| 2022 | Novel-view X-ray projection synthesis through geometry-integrated deep learning
Liyue Shen, Lequan Yu, Wei Zhao 0029, John M. Pauly, Lei Xing 0001 |
Medical Image Anal. | 2 |
| 2022 | All-Around Real Label Supervision: Cyclic Prototype Consistency Learning for Semi-Supervised Medical Image SegmentationabstractSemi-supervised learning has substantially advanced medical image segmentation since it alleviates the heavy burden of acquiring the costly expert-examined annotations. Especially, the consistency-based approaches have attracted more attention for their superior performance, wherein the real labels are only utilized to supervise their paired images via supervised loss while the unlabeled images are exploited by enforcing the perturbation-based "unsupervised" consistency without explicit guidance from those real labels. However, intuitively, the expert-examined real labels contain more reliable supervision signals. Observing this, we ask an unexplored but interesting question: can we exploit the unlabeled data via explicit real label supervision for semi-supervised training? To this end, we discard the previous perturbation-based consistency but absorb the essence of non-parametric prototype learning. Based on the prototypical networks, we then propose a novel cyclic prototype consistency learning (CPCL) framework, which is constructed by a labeled-to-unlabeled (L2U) prototypical forward process and an unlabeled-to-labeled (U2L) backward process. Such two processes synergistically enhance the segmentation network by encouraging morediscriminative and compact features. In this way, our framework turns previous "unsupervised" consistency into new "supervised" consistency, obtaining the "all-around real label supervision" property of our method. Extensive experiments on brain tumor segmentation from MRI and kidney segmentation from CT images show that our CPCL can effectively exploit the unlabeled data and outperform other state-of-the-art semi-supervised medical image segmentation methods. Zhe Xu 0012, Yixin Wang 0003, Donghuan Lu, Lequan Yu, Jiangpeng Yan, Jie Luo 0003, Kai Ma 0002, Yefeng Zheng 0001, Raymond Kai-Yu Tong |
IEEE J. Biomed. Health Informatics | 4 |
| 2022 | Lymph Node Metastasis Prediction From Whole Slide Images With Transformer-Guided Multiinstance Learning and Knowledge TransferabstractThe gold standard for diagnosing lymph node metastasis of papillary thyroid carcinoma is to analyze the whole slide histopathological images (WSIs). Due to the large size of WSIs, recent computer-aided diagnosis approaches adopt the multi-instance learning (MIL) strategy and the key part is how to effectively aggregate the information of different instances (patches). In this paper, a novel transformer-guided framework is proposed to predict lymph node metastasis from WSIs, where we incorporate the transformer mechanism to improve the accuracy from three different aspects. First, we propose an effective transformer-based module for discriminative patch feature extraction, including a lightweight feature extractor with a pruned transformer (Tiny-ViT) and a clustering-based instance selection scheme. Next, we propose a new Transformer-MIL module to capture the relationship of different discriminative patches with sparse distribution on WSIs and better nonlinearly aggregate patch-level features into the slide-level prediction. Considering that the slide-level annotation is relatively limited to training a robust Transformer-MIL, we utilize the pathological relationship between the primary tumor and its lymph node metastasis and develop an effective attention-based mutual knowledge distillation (AMKD) paradigm. Experimental results on our collected WSI dataset demonstrate the efficiency of the proposed Transformer-MIL and attention-based knowledge distillation. Our method outperforms the state-of-the-art methods by over 2.72% in AUC (area under the curve). Lequan Yu, Xuehong Liao, Liansheng Wang 0002 |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Robust Medical Image Classification From Noisy Labeled Data With Global and Local Representation Guided Co-TrainingabstractDeep neural networks have achieved remarkable success in a wide variety of natural image and medical image computing tasks. However, these achievements indispensably rely on accurately annotated training data. If encountering some noisy-labeled images, the network training procedure would suffer from difficulties, leading to a sub-optimal classifier. This problem is even more severe in the medical image analysis field, as the annotation quality of medical images heavily relies on the expertise and experience of annotators. In this paper, we propose a novel collaborative training paradigm with global and local representation learning for robust medical image classification from noisy-labeled data to combat the lack of high quality annotated medical data. Specifically, we employ the self-ensemble model with a noisy label filter to efficiently select the clean and noisy samples. Then, the clean samples are trained by a collaborative training strategy to eliminate the disturbance from imperfect labeled samples. Notably, we further design a novel global and local representation learning scheme to implicitly regularize the networks to utilize noisy samples in a self-supervised manner. We evaluated our proposed robust learning strategy on four public medical image classification datasets with three types of label noise, i.e., random noise, computer-generated label noise, and inter-observer variability noise. Our method outperforms other learning from noisy label methods and we also conducted extensive experiments to analyze each component of our method. Cheng Xue 0003, Lequan Yu, Pengfei Chen 0003, Qi Dou 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2021 | Selective Learning from External Data for CT Image Segmentation
Youyi Song, Lequan Yu, Bai Ying Lei, Kup-Sze Choi, Harry Qin |
MICCAI (1) | 2 |
| 2021 | TransCT: Dual-Path Transformer for Low Dose Computed Tomography
Zhicheng Zhang 0005, Lequan Yu, Xiaokun Liang, Wei Zhao 0029, Lei Xing 0001 |
MICCAI (6) | 2 |
| 2021 | NIA-Network: Towards improving lung CT infection detection for COVID-19 diagnosis
Wei Li 0121, Jinlin Chen, Ping Chen 0001, Lequan Yu, Xiaohui Cui, Wen Ouyang |
Artif. Intell. Medicine | 4 |
| 2021 | Rotation-Oriented Collaborative Self-Supervised Learning for Retinal Disease DiagnosisabstractThe automatic diagnosis of various conventional ophthalmic diseases from fundus images is important in clinical practice. However, developing such automatic solutions is challenging due to the requirement of a large amount of training data and the expensive annotations for medical images. This paper presents a novel self-supervised learning framework for retinal disease diagnosis to reduce the annotation efforts by learning the visual features from the unlabeled images. To achieve this, we present a rotation-oriented collaborative method that explores rotation-related and rotation-invariant features, which capture discriminative structures from fundus images and also explore the invariant property used for retinal disease classification. We evaluate the proposed method on two public benchmark datasets for retinal disease classification. The experimental results demonstrate that our method outperforms other self-supervised feature learning methods (around 4.2% area under the curve (AUC)). With a large amount of unlabeled data available, our method can surpass the supervised baseline for pathologic myopia (PM) and is very close to the supervised baseline for age-related macular degeneration (AMD), showing the potential benefit of our method in clinical practice. Xiaomeng Li 0001, Xiaowei Hu 0001, Xiaojuan Qi 0001, Lequan Yu, Wei Zhao 0029, Pheng-Ann Heng, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Dual-Teacher++: Exploiting Intra-Domain and Inter-Domain Knowledge With Reliable Transfer for Cardiac SegmentationabstractAnnotation scarcity is a long-standing problem in medical image analysis area. To efficiently leverage limited annotations, abundant unlabeled data are additionally exploited in semi-supervised learning, while well-established cross-modality data are investigated in domain adaptation. In this paper, we aim to explore the feasibility of concurrently leveraging both unlabeled data and cross-modality data for annotation-efficient cardiac segmentation. To this end, we propose a cutting-edge semi-supervised domain adaptation framework, namely Dual-Teacher++. Besides directly learning from limited labeled target domain data (e.g., CT) via a student model adopted by previous literature, we design novel dual teacher models, including an inter-domain teacher model to explore cross-modality priors from source domain (e.g., MR) and an intra-domain teacher model to investigate the knowledge beneath unlabeled target domain. In this way, the dual teacher models would transfer acquired inter- and intra-domain knowledge to the student model for further integration and exploitation. Moreover, to encourag reliable dual-domain knowledge transfer, we enhance the inter-domain knowledge transfer on the samples with higher similarity to target domain after appearance alignment, and also strengthen intra-domain knowledge transfer of unlabeled target data with higher prediction confidence. In this way, the student model can obtain reliable dual-domain knowledge and yield improved performance on target domain data. We extensively evaluated the feasibility of our method on the MM-WHS 2017 challenge dataset. The experiments have demonstrated the superiority of our framework over other semi-supervised learning and domain adaptation methods. Moreover, our performance gains could be yielded in bidirections, i.e., adapting from MR to CT, and from CT to MR. Our code will be available at https://github.com/kli-lalala/Dual-Teacher-. Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Deep Neural Network With Consistency Regularization of Multi-Output Channels for Improved Tumor Detection and DelineationabstractDeep learning is becoming an indispensable tool for imaging applications, such as image segmentation, classification, and detection. In this work, we reformulate a standard deep learning problem into a new neural network architecture with multi-output channels, which reflects different facets of the objective, and apply the deep neural network to improve the performance of image segmentation. By adding one or more interrelated auxiliary-output channels, we impose an effective consistency regularization for the main task of pixelated classification (i.e., image segmentation). Specifically, multi-output-channel consistency regularization is realized by residual learning via additive paths that connect main-output channel and auxiliary-output channels in the network. The method is evaluated on the detection and delineation of lung and liver tumors with public data. The results clearly show that multi-output-channel consistency implemented by residual learning improves the standard deep neural network. The proposed framework is quite broad and should find widespread applications in various deep learning problems. Hyunseok Seo, Lequan Yu, Hongyi Ren, Xiaomeng Li 0001, Liyue Shen, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2021 | Multi-Site Infant Brain Segmentation Algorithms: The iSeg-2019 ChallengeabstractTo better understand early brain development in health and disorder, it is critical to accurately segment infant brain magnetic resonance (MR) images into white matter (WM), gray matter (GM), and cerebrospinal fluid (CSF). Deep learning-based methods have achieved state-of-the-art performance; h owever, one of the major limitations is that the learning-based methods may suffer from the multi-site issue, that is, the models trained on a dataset from one site may not be applicable to the datasets acquired from other sites with different imaging protocols/scanners. To promote methodological development in the community, the iSeg-2019 challenge (http://iseg2019.web.unc.edu) provides a set of 6-month infant subjects from multiple sites with different protocols/scanners for the participating methods. T raining/validation subjects are from UNC (MAP) and testing subjects are from UNC/UMN (BCP), Stanford University, and Emory University. By the time of writing, there are 30 automatic segmentation methods participated in the iSeg-2019. In this article, 8 top-ranked methods were reviewed by detailing their pipelines/implementations, presenting experimental results, and evaluating performance across different sites in terms of whole brain, regions of interest, and gyral landmark curves. We further pointed out their limitations and possible directions for addressing the multi-site issue. We find that multi-site consistency is still an open issue. We hope that the multi-site dataset in the iSeg-2019 and this review article will attract more researchers to address the challenging and critical multi-site issue in practice. Yue Sun 0001, Kun Gao 0002, Zhengwang Wu, Xiaopeng Zong, Zhihao Lei, Ying Wei 0007, Jun Ma 0016, Xiaoping Yang 0001, Xue Feng 0001, Li Zhao 0001, Trung Le Phan, Jitae Shin, Tao Zhong 0002, Yu Zhang 0064, Lequan Yu, Caizi Li, Ramesh Basnet, M. Omair Ahmad, M. N. S. Swamy 0001, Wenao Ma, Qi Dou 0001, Toan Duc Bui, Camilo Bermudez, Bennett A. Landman, Ian H. Gotlib, Kathryn L. Humphreys, Sarah Shultz, Longchuan Li, Sijie Niu, Weili Lin, Valerie Jewells, Dinggang Shen, Gang Li 0001, Li Wang 0026 |
IEEE Trans. Medical Imaging | 16 |
| 2021 | Deep Sinogram Completion With Image Prior for Metal Artifact Reduction in CT ImagesabstractComputed tomography (CT) has been widely used for medical diagnosis, assessment, and therapy planning and guidance. In reality, CT images may be affected adversely in the presence of metallic objects, which could lead to severe metal artifacts and influence clinical diagnosis or dose calculation in radiation therapy. In this article, we propose a generalizable framework for metal artifact reduction (MAR) by simultaneously leveraging the advantages of image domain and sinogram domain-based MAR techniques. We formulate our framework as a sinogram completion problem and train a neural network (SinoNet) to restore the metal-affected projections. To improve the continuity of the completed projections at the boundary of metal trace and thus alleviate new artifacts in the reconstructed CT images, we train another neural network (PriorNet) to generate a good prior image to guide sinogram learning, and further design a novel residual sinogram learning strategy to effectively utilize the prior image information for better sinogram completion. The two networks are jointly trained in an end-to-end fashion with a differentiable forward projection (FP) operation so that the prior image generation and deep sinogram completion procedures can benefit from each other. Finally, the artifact-reduced CT images are reconstructed using the filtered backward projection (FBP) from the completed sinogram. Extensive experiments on simulated and real artifacts data demonstrate that our method produces superior artifact-reduced results while preserving the anatomical structures and outperforms other MAR methods. Lequan Yu, Zhicheng Zhang 0005, Xiaomeng Li 0001, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2021 | Transformation-Consistent Self-Ensembling Model for Semisupervised Medical Image SegmentationabstractA common shortfall of supervised deep learning for medical imaging is the lack of labeled data, which is often expensive and time consuming to collect. This article presents a new semisupervised method for medical image segmentation, where the network is optimized by a weighted combination of a common supervised loss only for the labeled inputs and a regularization loss for both the labeled and unlabeled data. To utilize the unlabeled data, our method encourages consistent predictions of the network-in-training for the same input under different perturbations. With the semisupervised segmentation tasks, we introduce a transformation-consistent strategy in the self-ensembling model to enhance the regularization effect for pixel-level predictions. To further improve the regularization effects, we extend the transformation in a more generalized form including scaling and optimize the consistency loss with a teacher model, which is an averaging of the student model weights. We extensively validated the proposed semisupervised method on three typical yet challenging medical image segmentation tasks: 1) skin lesion segmentation from dermoscopy images in the International Skin Imaging Collaboration (ISIC) 2017 data set; 2) optic disk (OD) segmentation from fundus images in the Retinal Fundus Glaucoma Challenge (REFUGE) data set; and 3) liver segmentation from volumetric CT scans in the Liver Tumor Segmentation Challenge (LiTS) data set. Compared with state-of-the-art, our method shows superior performance on the challenging 2-D/3-D medical images, demonstrating the effectiveness of our semisupervised method for medical image segmentation. Xiaomeng Li 0001, Lequan Yu, Hao Chen 0011, Chi-Wing Fu, Lei Xing 0001, Pheng-Ann Heng |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Towards Cross-Modality Medical Image Segmentation with Online Mutual Knowledge DistillationabstractThe success of deep convolutional neural networks is partially attributed to the massive amount of annotated training data. However, in practice, medical data annotations are usually expensive and time-consuming to be obtained. Considering multi-modality data with the same anatomic structures are widely available in clinic routine, in this paper, we aim to exploit the prior knowledge (e.g., shape priors) learned from one modality (aka., assistant modality) to improve the segmentation performance on another modality (aka., target modality) to make up annotation scarcity. To alleviate the learning difficulties caused by modality-specific appearance discrepancy, we first present an Image Alignment Module (IAM) to narrow the appearance gap between assistant and target modality data. We then propose a novel Mutual Knowledge Distillation (MKD) scheme to thoroughly exploit the modality-shared knowledge to facilitate the target-modality segmentation. To be specific, we formulate our framework as an integration of two individual segmentors. Each segmentor not only explicitly extracts one modality knowledge from corresponding annotations, but also implicitly explores another modality knowledge from its counterpart in mutual-guided manner. The ensemble of two segmentors would further integrate the knowledge from both modalities and generate reliable segmentation results on target modality. Experimental results on the public multi-class cardiac segmentation data, i.e., MM-WHS 2017, show that our method achieves large improvements on CT segmentation by utilizing additional MRI data and outperforms other state-of-the-art multi-modality learning methods. Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
AAAI | 2 |
| 2020 | Learning from Extrinsic and Intrinsic Supervisions for Domain Generalization
Lequan Yu, Caizi Li, Chi-Wing Fu, Pheng-Ann Heng |
ECCV (9) | 2 |
| 2020 | Local and Global Structure-Aware Entropy Regularized Mean Teacher Model for 3D Left Atrium Segmentation
Wenlong Hang, Wei Feng 0005, Shuang Liang 0015, Lequan Yu, Qiong Wang 0001, Kup-Sze Choi, Harry Qin |
MICCAI (1) | 4 |
| 2020 | Dual-Teacher: Integrating Intra-domain and Inter-domain Teachers for Annotation-Efficient Cardiac Segmentation
Kang Li 0007, Lequan Yu, Pheng-Ann Heng |
MICCAI (1) | 3 |
| 2020 | Difficulty-Aware Meta-learning for Rare Disease Diagnosis
Xiaomeng Li 0001, Lequan Yu, Yueming Jin, Chi-Wing Fu, Lei Xing 0001, Pheng-Ann Heng |
MICCAI (1) | 2 |
| 2020 | Robust Medical Image Segmentation from Non-expert Annotations with Tri-network
Lequan Yu, Na Hu, Su Lv, Shi Gu |
MICCAI (4) | 2 |
| 2020 | 3D Semi-Supervised Learning with Uncertainty-Aware Multi-View Co-TrainingabstractWhile making a tremendous impact in various fields, deep neural networks usually require large amounts of labeled data for training which are expensive to collect in many applications, especially in the medical domain. Un-labeled data, on the other hand, is much more abundant. Semi-supervised learning techniques, such as co-training, could provide a powerful tool to leverage unlabeled data. In this paper, we propose a novel framework, uncertainty-aware multi-view co-training (UMCT), to address semi-supervised learning on 3D data, such as volumetric data from medical imaging. In our work, co-training is achieved by exploiting multi-viewpoint consistency of 3D data. We generate different views by rotating or permuting the 3D data and utilize asymmetrical 3D kernels to encourage diversified features in different sub-networks. In addition, we propose an uncertainty-weighted label fusion mechanism to estimate the reliability of each view's prediction with Bayesian deep learning. As one view requires the supervision from other views in co-training, our self-adaptive approach computes a confidence score for the prediction of each unlabeled sample in order to assign a reliable pseudo label. Thus, our approach can take advantage of unlabeled data during training. We show the effectiveness of our proposed semi-supervised method on several public datasets from medical image segmentation tasks (NIH pancreas & LiTS liver tumor dataset). Meanwhile, a fully-supervised method based on our approach achieved state-of-the-art performances on both the LiTS liver tumor segmentation and the Medical Segmentation Decathlon (MSD) challenge, demonstrating the robustness and value of our framework, even when fully supervised training is feasible. Yingda Xia, Fengze Liu, Dong Yang 0005, Jinzheng Cai, Lequan Yu, Zhuotun Zhu, Daguang Xu, Alan L. Yuille, Holger Roth |
WACV | 5 |
| 2020 | Revisiting metric learning for few-shot image classification
Xiaomeng Li 0001, Lequan Yu, Chi-Wing Fu, Pheng-Ann Heng |
Neurocomputing | 2 |
| 2020 | Uncertainty-aware multi-view co-training for semi-supervised medical image segmentation and domain adaptation
Yingda Xia, Dong Yang 0005, Zhiding Yu, Fengze Liu, Jinzheng Cai, Lequan Yu, Zhuotun Zhu, Daguang Xu, Alan L. Yuille, Holger Roth |
Medical Image Anal. | 6 |
| 2020 | CANet: Cross-Disease Attention Network for Joint Diabetic Retinopathy and Diabetic Macular Edema GradingabstractDiabetic retinopathy (DR) and diabetic macular edema (DME) are the leading causes of permanent blindness in the working-age population. Automatic grading of DR and DME helps ophthalmologists design tailored treatments to patients, thus is of vital importance in the clinical practice. However, prior works either grade DR or DME, and ignore the correlation between DR and its complication, i.e., DME. Moreover, the location information, e.g., macula and soft hard exhaust annotations, are widely used as a prior for grading. Such annotations are costly to obtain, hence it is desirable to develop automatic grading methods with only image-level supervision. In this article, we present a novel cross-disease attention network (CANet) to jointly grade DR and DME by exploring the internal relationship between the diseases with only image-level supervision. Our key contributions include the disease-specific attention module to selectively learn useful features for individual diseases, and the disease-dependent attention module to further capture the internal relationship between the two diseases. We integrate these two attention modules in a deep network to produce disease-specific and disease-dependent features, and to maximize the overall performance jointly for grading DR and DME. We evaluate our network on two public benchmark datasets, i.e., ISBI 2018 IDRiD challenge dataset and Messidor dataset. Our method achieves the best result on the ISBI 2018 IDRiD challenge dataset and outperforms other methods on the Messidor dataset. Our code is publicly available at https://github.com/xmengli999/CANet. Xiaomeng Li 0001, Xiaowei Hu 0001, Lequan Yu, Lei Zhu 0003, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Self-Supervised Feature Learning via Exploiting Multi-Modal Data for Retinal Disease DiagnosisabstractThe automatic diagnosis of various retinal diseases from fundus images is important to support clinical decision-making. However, developing such automatic solutions is challenging due to the requirement of a large amount of human-annotated data. Recently, unsupervised/self-supervised feature learning techniques receive a lot of attention, as they do not need massive annotations. Most of the current self-supervised methods are analyzed with single imaging modality and there is no method currently utilize multi-modal images for better results. Considering that the diagnostics of various vitreoretinal diseases can greatly benefit from another imaging modality, e.g., FFA, this paper presents a novel self-supervised feature learning method by effectively exploiting multi-modal data for retinal disease diagnosis. To achieve this, we first synthesize the corresponding FFA modality and then formulate a patient feature-based softmax embedding objective. Our objective learns both modality-invariant features and patient-similarity features. Through this mechanism, the neural network captures the semantically shared information across different modalities and the apparent visual similarity between patients. We evaluate our method on two public benchmark datasets for retinal disease diagnosis. The experimental results demonstrate that our method clearly outperforms other self-supervised feature learning methods and is comparable to the supervised baseline. Our code is available at GitHub. Xiaomeng Li 0001, Mengyu Jia, Md Tauhidul Islam, Lequan Yu, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2020 | MS-Net: Multi-Site Network for Improving Prostate Segmentation With Heterogeneous MRI DataabstractAutomated prostate segmentation in MRI is highly demanded for computer-assisted diagnosis. Recently, a variety of deep learning methods have achieved remarkable progress in this task, usually relying on large amounts of training data. Due to the nature of scarcity for medical images, it is important to effectively aggregate data from multiple sites for robust model training, to alleviate the insufficiency of single-site samples. However, the prostate MRIs from different sites present heterogeneity due to the differences in scanners and imaging protocols, raising challenges for effective ways of aggregating multi-site data for network training. In this paper, we propose a novel multi-site network (MS-Net) for improving prostate segmentation by learning robust representations, leveraging multiple sources of data. To compensate for the inter-site heterogeneity of different MRI datasets, we develop Domain-Specific Batch Normalization layers in the network backbone, enabling the network to estimate statistics and perform feature normalization for each site separately. Considering the difficulty of capturing the shared knowledge from multiple datasets, a novel learning paradigm, i.e., Multi-site-guided Knowledge Transfer, is proposed to enhance the kernels to extract more generic representations from multi-site data. Extensive experiments on three heterogeneous prostate MRI datasets demonstrate that our MS-Net improves the performance across all datasets consistently, and outperforms state-of-the-art methods for multi-site learning. Quande Liu, Qi Dou 0001, Lequan Yu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Semi-Supervised Medical Image Classification With Relation-Driven Self-Ensembling ModelabstractTraining deep neural networks usually requires a large amount of labeled data to obtain good performance. However, in medical image analysis, obtaining high-quality labels for the data is laborious and expensive, as accurately annotating medical images demands expertise knowledge of the clinicians. In this paper, we present a novel relation-driven semi-supervised framework for medical image classification. It is a consistency-based method which exploits the unlabeled data by encouraging the prediction consistency of given input under perturbations, and leverages a self-ensembling model to produce high-quality consistency targets for the unlabeled data. Considering that human diagnosis often refers to previous analogous cases to make reliable decisions, we introduce a novel sample relation consistency (SRC) paradigm to effectively exploit unlabeled data by modeling the relationship information among different samples. Superior to existing consistency-based methods which simply enforce consistency of individual predictions, our framework explicitly enforces the consistency of semantic relation among different samples under perturbations, encouraging the model to explore extra semantic information from unlabeled data. We have conducted extensive experiments to evaluate our method on two public benchmark medical image classification datasets, i.e., skin lesion diagnosis with ISIC 2018 challenge and thorax disease classification with ChestX-ray14. Our method outperforms many state-of-the-art semi-supervised learning methods on both single-label and multi-label image classification scenarios. Quande Liu, Lequan Yu, Luyang Luo, Qi Dou 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Deep Mining External Imperfect Data for Chest X-Ray Disease ScreeningabstractDeep learning approaches have demonstrated remarkable progress in automatic Chest X-ray analysis. The data-driven feature of deep models requires training data to cover a large distribution. Therefore, it is substantial to integrate knowledge from multiple datasets, especially for medical images. However, learning a disease classification model with extra Chest X-ray (CXR) data is yet challenging. Recent researches have demonstrated that performance bottleneck exists in joint training on different CXR datasets, and few made efforts to address the obstacle. In this paper, we argue that incorporating an external CXR dataset leads to imperfect training data, which raises the challenges. Specifically, the imperfect data is in two folds: domain discrepancy, as the image appearances vary across datasets; and label discrepancy, as different datasets are partially labeled. To this end, we formulate the multi-label thoracic disease classification problem as weighted independent binary tasks according to the categories. For common categories shared across domains, we adopt task-specific adversarial training to alleviate the feature differences. For categories existing in a single dataset, we present uncertainty-aware temporal ensembling of model predictions to mine the information from the missing labels further. In this way, our framework simultaneously models and tackles the domain and label discrepancies, enabling superior knowledge mining ability. We conduct extensive experiments on three datasets with more than 360,000 Chest X-ray images. Our method outperforms other competing models and sets state-of-the-art performance on the official NIH test set with 0.8349 AUC, demonstrating its effectiveness of utilizing the external dataset to improve the internal classification. Luyang Luo, Lequan Yu, Hao Chen 0011, Quande Liu, Xi Wang 0013, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2020 | DoFE: Domain-Oriented Feature Embedding for Generalizable Fundus Image Segmentation on Unseen DatasetsabstractDeep convolutional neural networks have significantly boosted the performance of fundus image segmentation when test datasets have the same distribution as the training datasets. However, in clinical practice, medical images often exhibit variations in appearance for various reasons, e.g., different scanner vendors and image quality. These distribution discrepancies could lead the deep networks to over-fit on the training datasets and lack generalization ability on the unseen test datasets. To alleviate this issue, we present a novel Domain-oriented Feature Embedding (DoFE) framework to improve the generalization ability of CNNs on unseen target domains by exploring the knowledge from multiple source domains. Our DoFE framework dynamically enriches the image features with additional domain prior knowledge learned from multi-source domains to make the semantic features more discriminative. Specifically, we introduce a Domain Knowledge Pool to learn and memorize the prior information extracted from multi-source domains. Then the original image features are augmented with domain-oriented aggregated features, which are induced from the knowledge pool based on the similarity between the input image and multi-source domain images. We further design a novel domain code prediction branch to infer this similarity and employ an attention-guided mechanism to dynamically combine the aggregated features with the semantic features. We comprehensively evaluate our DoFE framework on two fundus image segmentation tasks, including the optic cup and disc segmentation and vessel segmentation. Our DoFE framework generates satisfying segmentation results on unseen datasets and surpasses other domain generalization and network regularization methods. Lequan Yu, Kang Li 0007, Xin Yang 0009, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Unsupervised Detection of Distinctive Regions on 3D ShapesabstractThis article presents a novel approach to learn and detect distinctive regions on 3D shapes. Unlike previous works, which require labeled data, our method is unsupervised. We conduct the analysis on point sets sampled from 3D shapes, then formulate and train a deep neural network for an unsupervised shape clustering task to learn local and global features for distinguishing shapes with respect to a given shape set. To drive the network to learn in an unsupervised manner, we design a clustering-based nonparametric softmax classifier with an iterative re-clustering of shapes, and an adapted contrastive loss for enhancing the feature embedding quality and stabilizing the learning process. By then, we encourage the network to learn the point distinctiveness on the input shapes. We extensively evaluate various aspects of our approach and present its applications for distinctiveness-guided shape retrieval, sampling, and view selection in 3D scenes. Xianzhi Li 0001, Lequan Yu, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng |
ACM Trans. Graph. | 2 |
| 2019 | Agent with Warm Start and Active Termination for Plane Localization in 3D Ultrasound
Haoran Dou, Xin Yang 0009, Jikuan Qian, Wufeng Xue, Xu Wang 0017, Lequan Yu, Yi Xiong 0001, Pheng-Ann Heng, Dong Ni 0001 |
MICCAI (5) | 7 |
| 2019 | Boundary and Entropy-Driven Adversarial Learning for Fundus Image Segmentation
Lequan Yu, Kang Li 0007, Xin Yang 0009, Chi-Wing Fu, Pheng-Ann Heng |
MICCAI (1) | 2 |
| 2019 | Uncertainty-Aware Self-ensembling Model for Semi-supervised 3D Left Atrium Segmentation
Lequan Yu, Xiaomeng Li 0001, Chi-Wing Fu, Pheng-Ann Heng |
MICCAI (2) | 1 |
| 2019 | RMDL: Recalibrated multi-instance deep learning for whole slide gastric image classification
Yaxi Zhu, Lequan Yu, Hao Chen 0011, Huangjing Lin, Xiangbo Wan, Xinjuan Fan, Pheng-Ann Heng |
Medical Image Anal. | 3 |
| 2019 | Patch-Based Output Space Adversarial Learning for Joint Optic Disc and Cup SegmentationabstractGlaucoma is a leading cause of irreversible blindness. Accurate segmentation of the optic disc (OD) and optic cup (OC) from fundus images is beneficial to glaucoma screening and diagnosis. Recently, convolutional neural networks demonstrate promising progress in the joint OD and OC segmentation. However, affected by the domain shift among different datasets, deep networks are severely hindered in generalizing across different scanners and institutions. In this paper, we present a novel patch-based output space adversarial learning framework ( p OSAL) to jointly and robustly segment the OD and OC from different fundus image datasets. We first devise a lightweight and efficient segmentation network as a backbone. Considering the specific morphology of OD and OC, a novel morphology-aware segmentation loss is proposed to guide the network to generate accurate and smooth segmentation. Our p OSAL framework then exploits unsupervised domain adaptation to address the domain shift challenge by encouraging the segmentation in the target domain to be similar to the source ones. Since the whole-segmentation-based adversarial loss is insufficient to drive the network to capture segmentation details, we further design the p OSAL in a patch-based fashion to enable fine-grained discrimination on local segmentation details. We extensively evaluate our p OSAL framework and demonstrate its effectiveness in improving the segmentation performance on three public retinal fundus image datasets, i.e., Drishti-GS, RIM-ONE-r3, and REFUGE. Furthermore, our p OSAL framework achieved the first place in the OD and OC segmentation tasks in the MICCAI 2018 Retinal Fundus Glaucoma Challenge. Lequan Yu, Xin Yang 0009, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Towards Automated Semantic Segmentation in Prenatal Volumetric UltrasoundabstractVolumetric ultrasound is rapidly emerging as a viable imaging modality for routine prenatal examinations. Biometrics obtained from the volumetric segmentation shed light on the reformation of precise maternal and fetal health monitoring. However, the poor image quality, low contrast, boundary ambiguity, and complex anatomy shapes conspire toward a great lack of efficient tools for the segmentation. It makes 3-D ultrasound difficult to interpret and hinders the widespread of 3-D ultrasound in obstetrics. In this paper, we are looking at the problem of semantic segmentation in prenatal ultrasound volumes. Our contribution is threefold: 1) we propose the first and fully automatic framework to simultaneously segment multiple anatomical structures with intensive clinical interest, including fetus, gestational sac, and placenta, which remains a rarely studied and arduous challenge; 2) we propose a composite architecture for dense labeling, in which a customized 3-D fully convolutional network explores spatial intensity concurrency for initial labeling, while a multi-directional recurrent neural network (RNN) encodes spatial sequentiality to combat boundary ambiguity for significant refinement; and 3) we introduce a hierarchical deep supervision mechanism to boost the information flow within RNN and fit the latent sequence hierarchy in fine scales, and further improve the segmentation results. Extensively verified on in-house large data sets, our method illustrates a superior segmentation performance, decent agreements with expert measurements and high reproducibilities against scanning variations, and thus is promising in advancing the prenatal ultrasound examinations. Xin Yang 0009, Lequan Yu, Shengli Li 0001, Huaxuan Wen, Dandan Luo, Cheng Bian, Harry Qin, Dong Ni 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2018 | Semi-supervised Skin Lesion Segmentation via Transformation Consistent Self-ensembling Model
Xiaomeng Li 0001, Lequan Yu, Hao Chen 0011, Chi-Wing Fu, Pheng-Ann Heng |
BMVC | 2 |
| 2018 | PU-Net: Point Cloud Upsampling NetworkabstractLearning and analyzing 3D point clouds with deep networks is challenging due to the sparseness and irregularity of the data. In this paper, we present a data-driven point cloud upsampling technique. The key idea is to learn multi-level features per point and expand the point set via a multi-branch convolution unit implicitly in feature space. The expanded feature is then split to a multitude of features, which are then reconstructed to an upsampled point set. Our network is applied at a patch-level, with a joint loss function that encourages the upsampled points to remain on the underlying surface with a uniform distribution. We conduct various experiments using synthesis and scan data to evaluate our method and demonstrate its superiority over some baseline methods and an optimization-based method. Results show that our upsampled points have better uniformity and are located closer to the underlying surfaces. Lequan Yu, Xianzhi Li 0001, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng |
CVPR | 1 |
| 2018 | EC-Net: An Edge-Aware Point Set Consolidation Network
Lequan Yu, Xianzhi Li 0001, Chi-Wing Fu, Daniel Cohen-Or, Pheng-Ann Heng |
ECCV (7) | 1 |
| 2018 | SV-RCNet: Workflow Recognition From Surgical Videos Using Recurrent Convolutional NetworkabstractWe propose an analysis of surgical videos that is based on a novel recurrent convolutional network (SV-RCNet), specifically for automatic workflow recognition from surgical videos online, which is a key component for developing the context-aware computer-assisted intervention systems. Different from previous methods which harness visual and temporal information separately, the proposed SV-RCNet seamlessly integrates a convolutional neural network (CNN) and a recurrent neural network (RNN) to form a novel recurrent convolutional architecture in order to take full advantages of the complementary information of visual and temporal features learned from surgical videos. We effectively train the SV-RCNet in an end-to-end manner so that the visual representations and sequential dynamics can be jointly optimized in the learning process. In order to produce more discriminative spatio-temporal features, we exploit a deep residual network (ResNet) and a long short term memory (LSTM) network, to extract visual features and temporal dependencies, respectively, and integrate them into the SV-RCNet. Moreover, based on the phase transition-sensitive predictions from the SV-RCNet, we propose a simple yet effective inference scheme, namely the prior knowledge inference (PKI), by leveraging the natural characteristic of surgical video. Such a strategy further improves the consistency of results and largely boosts the recognition performance. Extensive experiments have been conducted with the MICCAI 2016 Modeling and Monitoring of Computer Assisted Interventions Workflow Challenge dataset and Cholec80 dataset to validate SV-RCNet. Our approach not only achieves superior performance on these two datasets but also outperforms the state-of-the-art methods by a significant margin. Yueming Jin, Qi Dou 0001, Hao Chen 0011, Lequan Yu, Harry Qin, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 4 |
| 2017 | Fine-Grained Recurrent Neural Networks for Automatic Prostate Segmentation in Ultrasound ImagesabstractBoundary incompleteness raises great challenges to automatic prostate segmentation in ultrasound images. Shape prior can provide strong guidance in estimating the missing boundary, but traditional shape models often suffer from hand-crafted descriptors and local information loss in the fitting procedure. In this paper, we attempt to address those issues with a novel framework. The proposed framework can seamlessly integrate feature extraction and shape prior exploring, and estimate the complete boundary with a sequential manner. Our framework is composed of three key modules. Firstly, we serialize the static 2D prostate ultrasound images into dynamic sequences and then predict prostate shapes by sequentially exploring shape priors. Intuitively, we propose to learn the shape prior with the biologically plausible Recurrent Neural Networks (RNNs). This module is corroborated to be effective in dealing with the boundary incompleteness. Secondly, to alleviate the bias caused by different serialization manners, we propose a multi-view fusion strategy to merge shape predictions obtained from different perspectives. Thirdly, we further implant the RNN core into a multiscale Auto-Context scheme to successively refine the details of the shape prediction map. With extensive validation on challenging prostate ultrasound images, our framework bridges severe boundary incompleteness and achieves the best performance in prostate boundary delineation when compared with several advanced methods. Additionally, our approach is general and can be extended to other medical image segmentation tasks, where boundary incompleteness is one of the main challenges. Xin Yang 0009, Lequan Yu, Lingyun Wu, Yi Wang 0031, Dong Ni 0001, Harry Qin, Pheng-Ann Heng |
AAAI | 2 |
| 2017 | Volumetric ConvNets with Mixed Residual Connections for Automated Prostate Segmentation from 3D MR ImagesabstractAutomated prostate segmentation from 3D MR images is very challenging due to large variations of prostate shape and indistinct prostate boundaries. We propose a novel volumetric convolutional neural network (ConvNet) with mixed residual connections to cope with this challenging problem. Compared with previous methods, our volumetric ConvNet has two compelling advantages. First, it is implemented in a 3D manner and can fully exploit the 3D spatial contextual information of input data to perform efficient, precise and volume-to-volume prediction. Second and more important, the novel combination of residual connections (i.e., long and short) can greatly improve the training efficiency and discriminative capability of our network by enhancing the information propagation within the ConvNet both locally and globally. While the forward propagation of location information can improve the segmentation accuracy, the smooth backward propagation of gradient flow can accelerate the convergence speed and enhance the discrimination capability. Extensive experiments on the open MICCAI PROMISE12 challenge dataset corroborated the effectiveness of the proposed volumetric ConvNet with mixed residual connections. Our method ranked the first in the challenge, outperforming other competitors by a large margin with respect to most of evaluation metrics. The proposed volumetric ConvNet is general enough and can be easily extended to other medical image analysis tasks, especially ones with limited training data. Lequan Yu, Xin Yang 0009, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
AAAI | 1 |
| 2017 | Towards Automatic Semantic Segmentation in Volumetric Ultrasound
Xin Yang 0009, Lequan Yu, Shengli Li 0001, Xu Wang 0017, Harry Qin, Dong Ni 0001, Pheng-Ann Heng |
MICCAI (1) | 2 |
| 2017 | Automatic 3D Cardiovascular MR Segmentation with Densely-Connected Volumetric ConvNets
Lequan Yu, Jie-Zhi Cheng, Qi Dou 0001, Xin Yang 0009, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
MICCAI (2) | 1 |
| 2017 | DCAN: Deep contour-aware networks for object instance segmentation from histology images
Hao Chen 0011, Xiaojuan Qi 0001, Lequan Yu, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
Medical Image Anal. | 3 |
| 2017 | 3D deeply supervised network for automated segmentation of volumetric medical images
Qi Dou 0001, Lequan Yu, Hao Chen 0011, Yueming Jin, Xin Yang 0009, Harry Qin, Pheng-Ann Heng |
Medical Image Anal. | 2 |
| 2017 | Integrating Online and Offline Three-Dimensional Deep Learning for Automated Polyp Detection in Colonoscopy VideosabstractAutomated polyp detection in colonoscopy videos has been demonstrated to be a promising way for colorectal cancer prevention and diagnosis. Traditional manual screening is time consuming, operator dependent, and error prone; hence, automated detection approach is highly demanded in clinical practice. However, automated polyp detection is very challenging due to high intraclass variations in polyp size, color, shape, and texture, and low interclass variations between polyps and hard mimics. In this paper, we propose a novel offline and online three-dimensional (3-D) deep learning integration framework by leveraging the 3-D fully convolutional network (3D-FCN) to tackle this challenging problem. Compared with the previous methods employing hand-crafted features or 2-D convolutional neural network, the 3D-FCN is capable of learning more representative spatio-temporal features from colonoscopy videos, and hence has more powerful discrimination capability. More importantly, we propose a novel online learning scheme to deal with the problem of limited training data by harnessing the specific information of an input video in the learning process. We integrate offline and online learning to effectively reduce the number of false positives generated by the offline network and further improve the detection performance. Extensive experiments on the dataset of MICCAI 2015 Challenge on Polyp Detection demonstrated the better performance of our method when compared with other competitors. Lequan Yu, Hao Chen 0011, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 1 |
| 2017 | Comparative Validation of Polyp Detection Methods in Video Colonoscopy: Results From the MICCAI 2015 Endoscopic Vision ChallengeabstractColonoscopy is the gold standard for colon cancer screening though some polyps are still missed, thus preventing early disease detection and treatment. Several computational systems have been proposed to assist polyp detection during colonoscopy but so far without consistent evaluation. The lack of publicly available annotated databases has made it difficult to compare methods and to assess if they achieve performance levels acceptable for clinical use. The Automatic Polyp Detection sub-challenge, conducted as part of the Endoscopic Vision Challenge (http://endovis.grand-challenge.org) at the international conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) in 2015, was an effort to address this need. In this paper, we report the results of this comparative evaluation of polyp detection methods, as well as describe additional experiments to further explore differences between methods. We define performance metrics and provide evaluation databases that allow comparison of multiple methodologies. Results show that convolutional neural networks are the state of the art. Nevertheless, it is also demonstrated that combining different methodologies can lead to an improved overall performance. Jorge Bernal, Nima Tajkbaksh, Francisco Javier Sánchez, Bogdan J. Matuszewski, Hao Chen 0011, Lequan Yu, Quentin Angermann, Olivier Romain, Bjorn Rustad, Ilangko Balasingham, Konstantin Pogorelov, Sungbin Choi, Quentin Debard, Lena Maier-Hein, Stefanie Speidel, Danail Stoyanov, Patrick Brandao, Henry Córdova, Cristina Sánchez-Montes, Suryakanth R. Gurudu, Gloria Fernández-Esparrach, Xavier Dray, Jianming Liang, Aymeric Histace |
IEEE Trans. Medical Imaging | 6 |
| 2017 | Automated Melanoma Recognition in Dermoscopy Images via Very Deep Residual NetworksabstractAutomated melanoma recognition in dermoscopy images is a very challenging task due to the low contrast of skin lesions, the huge intraclass variation of melanomas, the high degree of visual similarity between melanoma and non-melanoma lesions, and the existence of many artifacts in the image. In order to meet these challenges, we propose a novel method for melanoma recognition by leveraging very deep convolutional neural networks (CNNs). Compared with existing methods employing either low-level hand-crafted features or CNNs with shallower architectures, our substantially deeper networks (more than 50 layers) can acquire richer and more discriminative features for more accurate recognition. To take full advantage of very deep networks, we propose a set of schemes to ensure effective training and learning under limited training data. First, we apply the residual learning to cope with the degradation and overfitting problems when a network goes deeper. This technique can ensure that our networks benefit from the performance gains achieved by increasing network depth. Then, we construct a fully convolutional residual network (FCRN) for accurate skin lesion segmentation, and further enhance its capability by incorporating a multi-scale contextual information integration scheme. Finally, we seamlessly integrate the proposed FCRN (for segmentation) and other very deep residual networks (for classification) to form a two-stage framework. This framework enables the classification network to extract more representative and specific features based on segmented results instead of the whole dermoscopy images, further alleviating the insufficiency of training data. The proposed framework is extensively evaluated on ISBI 2016 Skin Lesion Analysis Towards Melanoma Detection Challenge dataset. Experimental results demonstrate the significant performance gains of the proposed framework, ranking the first in classification and the second in segmentation among 25 teams and 28 teams, respectively. This study corroborates that very deep CNNs with effective training mechanisms can be employed to solve complicated medical image analysis tasks, even with limited training data. Lequan Yu, Hao Chen 0011, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 1 |
| 2016 | DCAN: Deep Contour-Aware Networks for Accurate Gland SegmentationabstractThe morphology of glands has been used routinely by pathologists to assess the malignancy degree of adenocarcinomas. Accurate segmentation of glands from histology images is a crucial step to obtain reliable morphological statistics for quantitative diagnosis. In this paper, we proposed an efficient deep contour-aware network (DCAN) to solve this challenging problem under a unified multi-task learning framework. In the proposed network, multi-level contextual features from the hierarchical architecture are explored with auxiliary supervision for accurate gland segmentation. When incorporated with multi-task regularization during the training, the discriminative capability of intermediate features can be further improved. Moreover, our network can not only output accurate probability maps of glands, but also depict clear contours simultaneously for separating clustered objects, which further boosts the gland segmentation performance. This unified framework can be efficient when applied to large-scale histopathological data without resorting to additional steps to generate contours based on low-level cues for post-separating. Our method won the 2015 MICCAI Gland Segmentation Challenge out of 13 competitive teams, surpassing all the other methods by a significant margin. Hao Chen 0011, Xiaojuan Qi 0001, Lequan Yu, Pheng-Ann Heng |
CVPR | 3 |
| 2016 | 3D Deeply Supervised Network for Automatic Liver Segmentation from CT Volumes
Qi Dou 0001, Hao Chen 0011, Yueming Jin, Lequan Yu, Harry Qin, Pheng-Ann Heng |
MICCAI (2) | 4 |
| 2016 | Automatic Detection of Cerebral Microbleeds From MR Images via 3D Convolutional Neural NetworksabstractCerebral microbleeds (CMBs) are small haemorrhages nearby blood vessels. They have been recognized as important diagnostic biomarkers for many cerebrovascular diseases and cognitive dysfunctions. In current clinical routine, CMBs are manually labelled by radiologists but this procedure is laborious, time-consuming, and error prone. In this paper, we propose a novel automatic method to detect CMBs from magnetic resonance (MR) images by exploiting the 3D convolutional neural network (CNN). Compared with previous methods that employed either low-level hand-crafted descriptors or 2D CNNs, our method can take full advantage of spatial contextual information in MR volumes to extract more representative high-level features for CMBs, and hence achieve a much better detection accuracy. To further improve the detection performance while reducing the computational cost, we propose a cascaded framework under 3D CNNs for the task of CMB detection. We first exploit a 3D fully convolutional network (FCN) strategy to retrieve the candidates with high probabilities of being CMBs, and then apply a well-trained 3D CNN discrimination model to distinguish CMBs from hard mimics. Compared with traditional sliding window strategy, the proposed 3D FCN strategy can remove massive redundant computations and dramatically speed up the detection process. We constructed a large dataset with 320 volumetric MR scans and performed extensive experiments to validate the proposed method, which achieved a high sensitivity of 93.16% with an average number of 2.74 false positives per subject, outperforming previous methods using low-level descriptors or 2D CNNs by a significant margin. The proposed method, in principle, can be adapted to other biomarker detection tasks from volumetric medical data. Qi Dou 0001, Hao Chen 0011, Lequan Yu, Lei Zhao 0003, Harry Qin, Defeng Wang, Vincent C. T. Mok, Lin Shi 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |