VLDB 2026 Research / reviewers in the wild / expert
Hao Chen 0011
dblp:86/475-11
· DBLP profile ↗
192ranked-venue papers
11as first author
139since 2021 · last 2027
0000-0002-8400-3780ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 135 · 5 first-author · 96 since 2021Graphics, computer vision, multimedia, augmented reality and games · 82 · 7 first-author · 58 since 2021Artificial intelligence and machine learning · 52 · 5 first-author · 39 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | HiLo: Spatial-spectral hybrid high-low frequency activation for heart and brain vessel segmentation
Qiong Wang 0001, Valentin E. Sinitsyn, Ying Hu 0001, Hao Chen 0011 |
Expert Syst. Appl. | 7 |
| 2027 | Robust collision prediction neural network based on η-form perceptron against visual variability
Hao Chen 0011, Yusi Wang |
Expert Syst. Appl. | 2 |
| 2026 | Knowledge-Enhanced Explainable Prompting for Vision-Language ModelsabstractLarge-scale vision-language models (VLMs) embedded with expansive representations and visual concepts have showcased significant potential in image and text understanding. Efficiently adapting VLMs such as CLIP to downstream tasks like few-shot image classification has garnered growing attention, with prompt learning emerging as a representative approach. However, most existing prompt-based adaptation methods, which rely solely on coarse-grained textual prompts, suffer from limited performance and interpretability when handling domain tasks that require specific knowledge. This results in a failure to satisfy the stringent trustworthiness requirements of Explainable Artificial Intelligence (XAI) in high-risk scenarios like healthcare. To address this issue, we propose a Knowledge-Enhanced Explainable Prompting (KEEP) framework that leverages fine-grained domain-specific knowledge to enhance the adaptation process of VLMs across various domains and image modalities. By incorporating retrieval augmented generation and domain foundation models, our framework can provide more reliable image-wise knowledge for prompt learning in various domains, alleviating the lack of fine-grained annotations, while offering both visual and textual explanations. Extensive experiments and explainability analyses conducted on eight datasets of different domains and image modalities demonstrate that our method simultaneously achieves superior performance and interpretability, highlighting the effectiveness of the collaboration between foundation models and XAI. Yequan Bie, Andong Tan, Zhixuan Chen, Zhiyuan Cai, Luyang Luo, Hao Chen 0011 |
AAAI | 6 |
| 2026 | ConSurv: Multimodal Continual Learning for Survival AnalysisabstractSurvival prediction of cancers is crucial for clinical practice, as it informs mortality risks and influences treatment plans. However, a static model trained on a single dataset fails to adapt to the dynamically evolving clinical environment and continuous data streams, limiting its practical utility. While continual learning (CL) offers a solution to learn dynamically from new datasets, existing CL methods primarily focus on unimodal inputs and suffer from severe catastrophic forgetting in survival prediction. In real-world scenarios, multimodal inputs often provide comprehensive and complementary information, such as whole slide images and genomics; and neglecting inter-modal correlations negatively impacts the performance. To address the two challenges of catastrophic forgetting and complex inter-modal interactions between gigapixel whole slide images and genomics, we propose ConSurv, the first multimodal continual learning (MMCL) method for survival analysis. ConSurv incorporates two key components: Multi-staged Mixture of Experts (MS-MoE) and Feature Constrained Replay (FCR). MS-MoE captures both task-shared and task-specific knowledge at different learning stages of the network, including two modality encoders and the modality fusion component, learning inter-modal relationships. FCR further enhances learned knowledge and mitigates forgetting by restricting feature deviation of previous data at different levels, including encoder-level features of two modalities and the fusion-level representations. Additionally, we introduce a new benchmark integrating four datasets, Multimodal Survival Analysis Incremental Learning (MSAIL), for comprehensive evaluation in the CL setting. Extensive experiments demonstrate that ConSurv outperforms competing methods across multiple metrics. Dianzhi Yu, Conghao Xiong, Yankai Chen 0001, Wenqian Cui, Xinni Zhang, Hao Chen 0011, Joseph J. Y. Sung, Irwin King |
AAAI | 7 |
| 2026 | EndoControlMag: Robust endoscopic vascular motion magnification with periodic reference resetting and hierarchical tissue-aware dual-mask controlabstractAccurate visualization of subtle vascular dynamics is a knowledge-intensive challenge in minimally invasive surgery. Conventional imaging systems struggle to reveal these imperceptible motions amidst the dynamic complexity of surgical scenes, limiting decision-making reliability. We introduce EndoControlMag , a Lagrangian framework that employs mask-conditioned magnification to selectively enhance vascular motion while preserving the structural integrity of surrounding tissues in endoscopic videos. Our approach integrates two key designs: Periodic Reference Resetting (PRR) , which divides videos into short overlapping clips with dynamically updated reference frames to alleviate error accumulation while maintaining temporal coherence, and Hierarchical Tissue-aware Magnification (HTM) , which combines pretrained visual tracking for accurate vessel localization with dual-mode adaptive softening strategies. HTM employs either motion-based softening that modulates magnification strength proportional to observed tissue displacement, or distance-based exponential decay that simulates biomechanical force attenuation. This strategy enables robust performance across diverse surgical scenarios where motion-based softening excels with complex tissue deformations and distance-based softening provides stability under unreliable optical flow conditions. To validate generality and scalability, we construct EndoVMM24, a benchmark dataset spanning four surgical specialties and diverse intraoperative scenarios. Extensive quantitative metrics, qualitative assessments, and expert surgeon evaluations demonstrate that EndoControlMag significantly outperforms existing methods in magnification accuracy, image quality, and robustness. This work advances engineering informatics for surgical vision by providing a reproducible, context-aware framework that supports reliable decision-making in minimally invasive procedures. The code, dataset, and video results are available at https://cho-haz.github.io/EndoControlMag/ . An Wang 0007, Rulin Zhou, Mengya Xu, Yiru Ye, Longfei Gou, Yiting Chang, Hao Chen 0011, Chwee Ming Lim, Jiankun Wang 0001, Hongliang Ren 0001 |
Adv. Eng. Informatics | 7 |
| 2026 | LungRes80: Towards tangled surgical workflow recognition in video-assisted thoracoscopic surgeryabstractVideo-Assisted Thoracoscopic Surgery (VATS) is a minimally invasive procedure developed to remove specific lung segments for the treatment of early-stage lung diseases. The surgical procedure involves intricate vascular and bronchial anatomy to preserve as much lung tissue as possible, minimizing impact on the pulmonary function. To assist in monitoring and early warning of this high-risk surgical workflow, we build a new dataset, LungRes80, including 269,806 video frames with phase annotations sampled from 80 VATS cases. LungRes80 presents unique challenges for hierarchical temporal modeling due to diverse short-term transitions between segmentectomy phases and latent long-term causal relations. To this end, we introduce an online baseline model termed LungReco. This framework employs Masked Causal Reasoning (MCR) to perform causal reasoning with semantic modeling from continuously updated memories along with pre-trained Large Language Models (LLMs), and combines it with Concurrent Spatial-Temporal encoding (CoST) for holistic bi-modal co-spatial-temporal aggregation across short- and long-term memories. Furthermore, a new metric, called the Attentional Distraction Coefficient (ADC), is proposed to quantify the costs of intraoperative distraction and postoperative corrections by wrong predictions. We establish a comprehensive benchmark for surgical workflow recognition by evaluating representative models on LungRes80, AutoLaparo, and Cholec80, where our method consistently achieves state-of-the-art performance. Code and data are available at LungRes80. Diandian Guo, Jialun Pei, Jiaao Li, Yanhui Wan, Hao Chen 0011, Pheng-Ann Heng |
Medical Image Anal. | 6 |
| 2026 | Learning with less supervision: A survey of label-efficient learning for medical image analysis
Cheng Jin 0003, Zhengrui Guo, Yi Lin 0009, Luyang Luo, Hao Chen 0011 |
Medical Image Anal. | 5 |
| 2026 | MG-3D: Multi-grained knowledge-enhanced vision-language pre-training for 3D medical image analysis
Xuefeng Ni, Linshan Wu, Jiaxin Zhuang, Qiong Wang 0001, Mingxiang Wu, Varut Vardhanabhuti, Lihai Zhang, Hanyu Gao, Hao Chen 0011 |
Medical Image Anal. | 9 |
| 2026 | GenAR: Next-scale autoregressive generation for spatial gene expression prediction
Jiarui Ouyang, Yihui Wang 0002, Yihang Gao, Yingxue Xu, Shu Yang 0004, Hao Chen 0011 |
Medical Image Anal. | 6 |
| 2026 | Comparative validation of surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation in endoscopy: Results of the PhaKIR 2024 challengeabstractReliable recognition and localization of surgical instruments in endoscopic video recordings are foundational for a wide range of applications in computer- and robot-assisted minimally invasive surgery (RAMIS), including surgical training, skill assessment, and autonomous assistance. However, robust performance under real-world conditions remains a significant challenge. Incorporating surgical context - such as the current procedural phase - has emerged as a promising strategy to improve robustness and interpretability. To address these challenges, we organized the Surgical Procedure Phase, Keypoint, and Instrument Recognition (PhaKIR) sub-challenge as part of the Endoscopic Vision (EndoVis) challenge at MICCAI 2024. We introduced a novel, multi-center dataset comprising thirteen full-length laparoscopic cholecystectomy videos collected from three distinct medical institutions, with unified annotations for three interrelated tasks: surgical phase recognition, instrument keypoint estimation, and instrument instance segmentation. Unlike existing datasets, ours enables joint investigation of instrument localization and procedural context within the same data while supporting the integration of temporal information across entire procedures. We report results and findings in accordance with the BIAS guidelines for biomedical image analysis challenges. The PhaKIR sub-challenge advances the field by providing a unique benchmark for developing temporally aware, context-driven methods in RAMIS and offers a high-quality resource to support future research in surgical scene understanding. Tobias Rueckert, David Rauber, Raphaela Maerkl, Leonard Klausmann, Suemeyye R. Yildiran, Max Gutbrod, Danilo Weber Nunes, Alvaro Fernandez Moreno, Imanol Luengo, Danail Stoyanov, Nicolas Toussaint, Enki Cho, Hyeon Bae Kim, Oh Sung Choo, Ka Young Kim, Seong Tae Kim 0001, Gonçalo Arantes, Kehan Song, Junchen Xiong, Tingyi Lin, Shunsuke Kikuchi, Hiroki Matsuzaki, Atsushi Kouno, João Renato Ribeiro Manesco, João Paulo Papa, Tae-Min Choi, Tae Kyeong Jeong, Oluwatosin Alabi, Tom Vercauteren, Runzhi Wu, Mengya Xu, An Wang 0007, Long Bai 0008, Hongliang Ren 0001, Amine Yamlahi, Jakob Hennighausen, Lena Maier-Hein, Satoshi Kondo, Satoshi Kasai, Kousuke Hirasawa, Shu Yang 0004, Yihui Wang 0002, Hao Chen 0011, Santiago Rodríguez, Nicolás Aparicio, Leonardo Manrique, Juan Camilo Lyons, Olivia Hosie, Nicolás Ayobi, Pablo Andrés Arbeláez, Yiping Li 0002, Yasmina Alkhalil, Sahar Nasirihaghighi, Stefanie Speidel, Daniel Rueckert, Hubertus Feußner, Dirk Wilhelm, Christoph Palm |
Medical Image Anal. | 46 |
| 2026 | PTCMIL: multiple instance learning via prompt token clustering for whole slide image analysisabstractMultiple Instance Learning (MIL) has achieved significant success in whole slide image (WSI) analysis. However, the complexity and heterogeneity in WSIs remain fundamental challenges for MIL problem due to the various information in each WSI. However, existing MIL methods face challenges in effectively aggregating diverse patch information into robust and predictive WSI representations. While Vision Transformers (ViTs) and clustering-based approaches have shown promise, they are often computationally intensive and fail to fully capture task-specific features and slide-specific variability. To address these limitations, we propose PTCMIL, a novel Prompt Token Clustering-based ViT for MIL aggregation. Unlike conventional two-stage clustering methods in MIL, PTCMIL introduces learnable prompt tokens into the Vision Transformer (ViT) backbone, enabling slide-specific, task-aware clustering through projection-based token clustering. By guiding clustering with prediction objectives and generating compact cluster prototypes through token merging, PTCMIL effectively captures both patch diversity and task-relevant patterns. Our key contributions include: (1) A prompt-driven clustering mechanism that learns meaningful prototypes without relying on expensive global clustering or patch sampling; (2) An efficient merging strategy to construct interpretable and compact WSI-level representations; and (3) A pooling module that supports both classification and survival analysis tasks. Extensive experiments across eleven benchmark datasets-including breast, lung, colorectal, and prostate cancer WSIs-demonstrate that PTCMIL consistently outperforms state-of-the-art MIL baselines in classification, survival prediction, and domain adaptation tasks. Our results highlight PTCMIL's potential as a practical and generalizable solution for large-scale computational pathology. The code is available at https://github.com/ubc-tea/PTCMIL. Beidi Zhao, Hao Chen 0011, Zu-hua Gao, Xiaoxiao Li 0001 |
Medical Image Anal. | 3 |
| 2026 | Causal Inference via Style Bias Deconfounding for Domain GeneralizationabstractDeep neural networks (DNNs) often struggle with out-of-distribution data, limiting their reliability in real-world visual applications. To address this issue, domain generalization methods have been developed to learn domain-invariant features from single or multiple training domains, enabling generalization to unseen testing domains. However, existing approaches usually overlook the impact of style frequency within the training set. This oversight predisposes models to capture spurious visual correlations caused by style confounding factors, rather than learning truly causal representations, thereby undermining inference reliability. In this work, we introduce Style Deconfounding Causal Learning (SDCL), a novel causal inference-based framework that explicitly addresses style as a confounding factor to enhance domain generalization in image modalities. Our approaches begins with constructing a structural causal model (SCM) tailored to the domain generalization problem and applies a backdoor adjustment strategy to account for style influence. Building on this foundation, we design a style-guided expert module (SGEM) to adaptively clusters style distributions during training, capturing the global confounding style. Additionally, a backdoor causal learning module (BDCL) performs causal interventions during feature extraction, ensuring fair integration of global confounding styles into sample predictions, effectively reducing style bias. The SDCL framework is highly versatile and can be seamlessly integrated with state-of-the-art data augmentation techniques. Extensive experiments across diverse natural and medical image recognition tasks validate its efficacy, demonstrating superior performance in both multi-domain and the more challenging single-domain generalization scenarios. Di Lin 0002, Hao Chen 0011, Hongying Liu 0001, Wei Feng 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Learning With Partial and Noisy Correspondence in Graph MatchingabstractThe success of existing graph matching methods heavily relies on high-quality training data with complete and precise correspondences between keypoints across different graphs. However, this assumption is often violated in real-world scenarios, leading to partial correspondence and noisy correspondence challenges. In brief, partial correspondence arises from viewpoint occlusions, where certain keypoints (i.e., outliers) lack valid counterparts in the target graph, while noisy correspondence refers to both incorrectly established (i.e., false positives) and neglected (i.e., false negatives) correspondences due to annotation error. In this paper, we propose the first unified framework to address both partial and noisy correspondence challenges in graph matching. Specifically, we introduce a dual-expert cooperative framework that integrates Koopmans-Beckmann and Lawler's quadratic assignment programming formulations (KB-QAP and L-QAP) through an align-fuse-refine pipeline. In the alignment stage, the KB-QAP expert aligns keypoints and distinguishes inliers from outliers using a novel quadratic contrastive loss. In the fusion stage, the L-QAP expert employs a graph transformer on the association graph to merge the aligned graphs and incorporates a learnable outlier-rejection mechanism to handle partial correspondences. Finally, by exploiting the different noise resistances of the two experts, we identify and refine the false positive and false negative correspondences, thereby enhancing robustness against noisy correspondence. Extensive experiments on four widely-used graph matching datasets demonstrate the effectiveness of our method against 17 competitive baselines in both partial and noisy correspondence scenarios. Yijie Lin 0001, Mouxing Yang, Peng Hu 0002, Jiancheng Lv 0001, Hao Chen 0011, Xi Peng 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Large-Scale 3D Medical Image Pre-Training With Geometric Context PriorsabstractThe scarcity of annotations poses a significant challenge in medical image analysis, which demands extensive efforts from radiologists, especially for high-dimension 3D medical images. Large-scale pre-training has emerged as a promising label-efficient solution, owing to the utilization of large-scale data, large models, and advanced pre-training techniques. However, its development in medical images remains underexplored. The primary challenge lies in harnessing large-scale unlabeled data and learning high-level semantics without annotations. We observe that 3D medical images exhibit consistent geometric context, i.e., consistent geometric relations between different organs, which leads to a promising way for learning consistent representations. Motivated by this, we introduce a simple-yet-effective Volume Contrast (VoCo) framework to leverage geometric context priors for self-supervision. Given an input volume, we extract base crops from different regions to construct positive and negative pairs for contrastive learning. Then we predict the contextual position of a random crop by contrasting its similarity to the base crops. In this way, VoCo implicitly encodes the inherent geometric context into model representations, facilitating high-level semantic learning without annotations. To assess effectiveness, we (1) introduce PreCT-160 K, the largest medical image pre-training dataset to date, which comprises 160 K Computed Tomography (CT) volumes covering diverse anatomic structures; (2) investigate scaling laws and propose guidelines for tailoring different model sizes to various medical tasks; (3) build a comprehensive benchmark encompassing 51 medical tasks, including segmentation, classification, registration, and vision-language. Extensive experiments highlight the superiority of VoCo, showcasing promising transferability to unseen modalities and datasets. VoCo notably enhances performance on datasets with limited labeled cases and significantly expedites fine-tuning convergence. Linshan Wu, Jiaxin Zhuang, Hao Chen 0011 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Position Paper: Artificial Intelligence in Medical Image Analysis: Advances, Clinical Translation, and Emerging FrontiersabstractOver the past five years, artificial intelligence (AI) has introduced new models and methods for addressing the challenges associated with the broader adoption of AI models and systems in medicine. This paper reviews recent advances in AI for medical image and video analysis, outlines emerging paradigms, highlights pathways for successful clinical translation, and provides recommendations for future work. Hybrid Convolutional Neural Network (CNN) Transformer architectures now deliver state-of-the-art results in segmentation, classification, reconstruction, synthesis, and registration. Foundation and generative AI models enable the use of transfer learning to smaller datasets with limited ground truth. Federated learning supports privacy-preserving collaboration across institutions. Explainable and trustworthy AI approaches have become essential to foster clinician trust, ensure regulatory compliance, and facilitate ethical deployment. Together, these developments pave the way for integrating AI into radiology, pathology, and wider healthcare workflows. Andreas Panayides, Hao Chen 0011, Nenad Filipovic, Tijana Geroski, Junlin Hou, Karim Lekadir, Kostas Marias, George K. Matsopoulos, Giorgos Papanastasiou, Pinaki Sarder, Georgia D. Tourassi, Sotirios A. Tsaftaris, Huazhu Fu, Efthyvoulos C. Kyriacou, Christos P. Loizou, Michalis E. Zervakis, Joel H. Saltz, Farah Shamout, Ken C. L. Wong, Jianhua Yao 0001, Amir A. Amini, Dimitrios I. Fotiadis, Constantinos S. Pattichis, Marios S. Pattichis |
IEEE J. Biomed. Health Informatics | 2 |
| 2026 | SegTom: A 3D Volumetric Medical Image Segmentation Framework for Thoracoabdominal Multi-Organ Anatomical StructuresabstractAccurate segmentation of thoracoabdominal anatomical structures in three-dimensional medical imaging modalities is fundamental for informed clinical decision-making across a wide array of medical disciplines. Current approaches often struggle to efficiently and comprehensively process this region's intricate and heterogeneous anatomical information, leading to suboptimal outcomes in diagnosis, treatment planning, and disease management. To address this challenge, we introduce SegTom, a novel volumetric segmentation framework equipped with a cutting-edge SegTom Block specifically engineered to effectively capture the complex anatomical representations inherent to the thoracoabdominal region. This SegTom Block incorporates a hierarchical anatomical-representation decomposition to facilitate efficient information exchange by decomposing the computationally intensive self-attention mechanism and cost-effectively aggregating the extracted representations. Rigorous validation of SegTom across nine diverse datasets, encompassing both computed tomography (CT) and magnetic resonance imaging (MRI) modalities, consistently demonstrates high performance across a broad spectrum of anatomical structures. Specifically, SegTom achieves a mean Dice similarity coefficient (DSC) of 87.29% for cardiac segmentation on the MM-WHS MRI dataset, 83.48% for multi-organ segmentation on the BTCV abdominal CT dataset, and 92.01% for airway segmentation on a dedicated CT dataset. Hao Chen 0011, Ying Hu 0001, Qiong Wang 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | LLM-Driven Medical Report Generation via Communication-Efficient Heterogeneous Federated LearningabstractLarge Language Models (LLMs) have demonstrated significant potential in Medical Report Generation (MRG), yet their development requires large amounts of medical image-report pairs, which are commonly scattered across multiple centers. Centralizing these data is exceptionally challenging due to privacy regulations, thereby impeding model development and broader adoption of LLM-driven MRG models. To address this challenge, we present FedMRG, the first framework that leverages Federated Learning (FL) to enable privacy-preserving, multi-center development of LLM-driven MRG models, specifically designed to overcome the critical challenge of communication-efficient LLM training under multi-modal data heterogeneity. To start with, our framework tackles the fundamental challenge of communication overhead in federated LLM tuning by employing low-rank factorization to efficiently decompose parameter updates, significantly reducing gradient transmission costs and making LLM-driven MRG feasible in bandwidth-constrained FL settings. Furthermore, we observed the dual heterogeneity in MRG under the FL scenario: varying image characteristics across medical centers, as well as diverse reporting styles and terminology preferences. To address the data heterogeneity, we further enhance FedMRG with (1) client-aware contrastive learning in the MRG encoder, coupled with diagnosis-driven prompts, which capture both globally generalizable and locally distinctive features while maintaining diagnostic accuracy; and (2) a dual-adapter mutual boosting mechanism in the MRG decoder that harmonizes generic and specialized adapters to address variations in reporting styles and terminology. Through extensive evaluation of our established FL-MRG benchmark, we demonstrate the generalizability and adaptability of FedMRG, underscoring its potential in harnessing multi-center data and generating clinically accurate reports while maintaining communication efficiency. Haoxuan Che, Haibo Jin, Zhengrui Guo, Yi Lin 0009, Cheng Jin 0003, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 6 |
| 2026 | UltraMamba: Mamba-Based Multimodal Ultrasound Image Adaptive Fusion for Breast Lesion SegmentationabstractMultimodal ultrasound imaging, combining B-mode ultrasound, shear wave velocity, and shear wave time, is crucial for diagnosing and treating breast lesions, providing insights into lesion characteristics and tissue properties. However, challenges arise from inter-modal feature misalignment and attention shifts due to varied capture methods and an overemphasis on vibrant color data. To tackle these issues, we introduce two innovations: a novel segmentation framework and a comprehensive dataset. The UltraMamba framework utilizes bidirectional alignment between modalities and enhances region-specific information to improve breast lesion segmentation accuracy. Key components include the Cross-Modal Knowledge Interaction module for robust information exchange and the Region-Aware Feature Excitation module to focus on relevant features. We also present the BreLS dataset, the first two-dimensional multimodal ultrasound breast lesion dataset, with paired images from 506 cases, serving as a valuable resource for analysis. UltraMamba shows strong performance on the BreLS dataset, achieving a Dice Similarity Coefficient of 72.16% and an HD95 of 42.02 mm, reflecting improvements of 2.59% in DSC and a 6.78 mm reduction in HD95 compared to the second-best framework, MMCA-NET. These results highlight UltraMamba's potential to enhance segmentation accuracy in clinical settings, facilitating precise treatment planning and, ultimately, leading to improved outcomes. Code: https://github.com/deepang-ai/UltraMamba. Mingdu Zhang, Qiong Wang 0001, Xiaoqing Pei, Ying Hu 0001, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 7 |
| 2026 | Scan-Invariant Mamba With Differentiated Sequence Contrastive Learning in Computational PathologyabstractMultiple instance learning (MIL) is a commonly used paradigm for histopathological analysis due to the ultra-high resolution and coarse-grained labels of Whole Slide Images (WSIs). Recent studies apply Mamba architecture to WSI classification by modeling MIL as long-sequence tasks, but a key discrepancy remains: Mamba's output is sensitive to scanning modes, whereas MIL requires scan-invariant predictions. To address this problem, we propose Scan-invariant Mamba with Differentiated Sequence Contrastive Learning (SMDC-MIL), a novel Mamba-based MIL approach enabling bag-level feature learning independent of input modes. Our method mitigates scanning-mode impacts and adapts Mamba to learn the bag discrimination features that are independent of the input mode via two innovations: 1) a differentiated sequence generation mechanism that employs instance rearrangement, augmentation, and masking to simulate real-world scanning variations by maximizing differences in sequence order, length, and composition from the same WSI; and 2) a differentiated sequence contrastive learning architecture that enforces consistent bag-level representations and predictions across diverse sequences using the same Mamba model, guiding it to prioritize scan-invariant discriminative features. Experimental results on 4 computational pathology tasks and 10 datasets demonstrate that our SMDC-MIL achieves state-of-the-art performance compared to other methods. The corresponding code is available at https://github.com/LianYueZ/SMDCMIL.git. Sheng Huang 0001, Xin Zhang 0131, Bo Liu 0005, Fengtao Zhou, Kang Li 0004, Hao Chen 0011, Meng Wang 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2026 | Endoscopic Adaptive Transformer for Enhanced Polyp Segmentation in Endoscopic ImagingabstractPolyp segmentation in endoscopic imaging is essential for the early detection of colorectal cancer, as polyps are precursor lesions in the colon and rectum, yet the task is complicated by the morphological variability and indistinct boundaries of polyps, which often blend into surrounding tissues. Conventional approaches struggle with these complexities, as fixed scale and window sizes are unable to adapt to the diverse and irregular structures of polyps. To address this challenge, we introduce the Endoscopic Adaptive Transformer, EAT, a novel framework specifically engineered for polyp segmentation. EAT incorporates an adaptive perception module, APM, that employs an adaptive perceptive-field mechanism to dynamically capture both fine-grained local details and broad contextual information, enhancing segmentation accuracy across diverse polyp morphologies. EAT demonstrates comprehensive performance by achieving a Dice coefficient of 97.77% and an HD95 of 4.50mm in single-target segmentation, while also excelling in multi-target scenarios with a Dice coefficient of 88.02% and an HD95 of 53.75mm, significantly outperforming state-of-the-art methods across both single- and multi-target segmentation scenarios. This performance underscores EAT's critical role in improving the accuracy of polyp segmentation, highlighting its potential to advance diagnostic precision and treatment planning in clinical endoscopy applications. Code: https://github.com/deepang-ai/EAT. Yucheng Long, Zibin Chen, Ying Hu 0001, Hao Chen 0011, Qiong Wang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2026 | Slim UNETRV2: 3D Image Segmentation for Resource-Limited Medical Portable DevicesabstractMedical portable devices are increasingly requiring high accuracy, speed, and low inference jitter to meet the urgent demands of healthcare. Modern hybrid attention-based segmentation frameworks enhance segmentation accuracy but add complexity that can slow operational speed, complicating practical deployment in resource-limited settings. We propose Slim UNETRV2, a simplified framework that utilizes only basic convolutional operations in both the encoder and decoder, thereby reducing execution time and inference jitter. The Slim UNETRV2 block, placed in skip connections at each hierarchical stage, aggregates extracted representations and improves global processing. Experiments demonstrate that Slim UNETRV2 outperforms state-of-the-art models in terms of accuracy, speed, and inference jitter for resource-constrained medical devices. Notably, Slim UNETRV2 achieves 93.89% dice accuracy and 2.90 mm HD95 on BraTS 2021, being 16.7 times faster with only 0.225 ms of inference jitter compared to SegMamba. Code: https://github.com/deepang-ai/Slim-UNETRV2https://github.com/deepang-ai/Slim-UNETRV2. Junming Yan, Ying Hu 0001, Hao Chen 0011, Qiong Wang 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2026 | Toward Semantically Faithful Diffusion Representation for Generalizable Retinal Image SegmentationabstractRetinal image segmentation is essential for analyzing retinal structures like vessels and diagnosing retinopathy. However, the inherent intricacy of the retina, along with annotation scarcity and data heterogeneity, presents prevalent challenges in creating accurate and generalizable deep learning models. Diffusion models, while initially developed for image generation, have recently shown great promise for visual perception by leveraging the learned internal representations. However, these diffusion representations, which spread across network blocks (space) and diffusion timesteps (time), potentially suffer from issues like stochastic semantic distortion and cumulative structural blurring, compromising their semantic fidelity to the source image. In this paper, by delving into the generalization property of diffusion models, we propose a novel anchoring inversion strategy to derive diffusion representations that are semantically faithful to the source image from the deterministic trajectory. Furthermore, we introduce a time-space frequency-aware aggregation interpreter (T&S-FreqAgg) to aggregate the multi-scale and multi-timestep diffusion representations in a frequency-aware way for Domain Generalizable Semantic Segmentation (DGSS). Extensive experiments on nine public retinal image datasets demonstrate the superiority of our proposed framework, DiffDGSSv2, over state-of-the-art methods. Our code will be available at: https://github.com/Xyporz/DiffDGSSv2. Yingpeng Xie, Hao Chen 0011, Harry Qin, Jie Du 0001, Tianfu Wang 0001, Bai Ying Lei |
IEEE Trans. Medical Imaging | 2 |
| 2026 | SurgPETL: Parameter-Efficient Image-to-Surgical-Video Transfer Learning for Surgical Phase RecognitionabstractCapitalizing on image-level pre-trained models for various downstream tasks has recently emerged with promising performance. However, the paradigm of "image pre-training followed by video fine-tuning" for high-dimensional video data inevitably introduces significant performance bottlenecks. Furthermore, in the medical domain, many surgical video tasks encounter additional challenges posed by the limited availability of video data and the necessity for comprehensive spatiotemporal modeling. Recently, Parameter-Efficient Image-to-Video Transfer Learning (PEIVTL) has emerged as an efficient and effective paradigm for video action recognition tasks, which employs image-level pre-trained models with promising feature transferability and involves cross-modality temporal modeling with minimal fine-tuning. Nevertheless, the effectiveness and generalizability of this paradigm within intricate surgical domain remain unexplored. In this paper, we delve into a novel problem of efficiently adapting image-level pre-trained models to specialize in fine-grained surgical phase recognition, termed Parameter-Efficient Image-to-Surgical-Video Transfer Learning. First, we develop SurgPETL, a parameter-efficient transfer learning framework for surgical phase recognition, and conduct extensive experiments with three advanced methods based on ViTs of two distinct scales pre-trained on five large-scale natural and medical datasets. Then, we introduce the Adaptive Spatiotemporal Representation Modulation (ASRM) module, integrating a standard spatial adapter with a novel temporal adapter to capture detailed spatial features and establish connections across temporal sequences for robust spatiotemporal modeling. Extensive experiments on three challenging datasets spanning various surgical procedures demonstrate the effectiveness of SurgPETL with ASRM. SurgPETL-ASRM outperforms both parameter-efficient alternatives and state-of-the-art surgical phase recognition methods while maintaining parameter efficiency and minimizing overhead. Shu Yang 0004, Zhiyuan Cai, Luyang Luo, Shuchang Xu, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | A Survey on Foundation Language Models for Single-cell BiologyabstractFan Zhang, Hao Chen, Zhihong Zhu, Ziheng Zhang, Zhenxi Lin, Ziyue Qiao, Yefeng Zheng, Xian Wu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Fan Zhang 0111, Hao Chen 0011, Zhihong Zhu 0001, Zhenxi Lin, Ziyue Qiao, Yefeng Zheng 0001, Xian Wu 0001 |
ACL (1) | 2 |
| 2025 | Masked Diffusion Models for Unsupervised Anomaly Detection in Brain ImagesabstractUnsupervised anomaly detection has gained significant attention in the field of medical imaging due to its capability of reducing the need for costly pixel-level annotation. To achieve this, existing approaches usually utilize generative models to produce healthy references of the diseased images and then identify the abnormalities by comparing healthy references and original diseased images. However, intrinsic characteristics of brain images, e.g., the low contrast and the intricate anatomical structure, make reconstruction challenging. To address those challenges, we propose a Masked Diffusion Model (MDiff), which incorporates a hierarchical patch partition strategy into the diffusion model for precise reconstruction of detailed content. Aligned with this strategy, we treat the perturbed upper-level patch as masked and introduce a masked modeling mechanism into MDiff's diffusion U-Net. This mechanism operates on sub-level patches, enhancing the model's ability to process contextual information surrounding the perturbed upper-level patch. To further improve the quality of healthy references, we integrate a memory module within the mechanism's encoder to retrieve the most relevant memory items as contextual information, while employing the learnable query embedding in its decoder to prevent the network from learning identical shortcuts. Experiments on tumor and multiple sclerosis lesion data demonstrate MDiff's effectiveness. Rui Xu 0031, Yunke Wang, Yong Luo 0002, Shu Yang 0004, Yihui Wang 0002, Bo Du 0001, Hao Chen 0011 |
BIBM | 8 |
| 2025 | FOCUS: Knowledge-enhanced Adaptive Visual Compression for Few-shot Whole Slide Image ClassificationabstractFew-shot learning presents a critical solution for cancer diagnosis in computational pathology (CPath), addressing fundamental limitations in data availability, particularly the scarcity of expert annotations and patient privacy constraints. A key challenge in this paradigm stems from the inherent disparity between the limited training set of whole slide images (WSIs) and the enormous number of contained patches, where a significant portion of these patches lacks diagnostically relevant information, potentially diluting the model’s ability to learn and focus on critical diagnostic features. While recent works attempt to address this by incorporating additional knowledge, several crucial gaps hinder further progress: (1) despite the emergence of powerful pathology foundation models (FMs), their potential remains largely untapped, with most approaches limiting their use to basic feature extraction; (2) current language guidance mechanisms attempt to align text prompts with vast numbers of WSI patches all at once, struggling to leverage rich pathological semantic information. To this end, we introduce the knowledge-enhanced adaptive visual compression framework, dubbed FOCUS, which uniquely combines pathology FMs with language prior knowledge to enable a focused analysis of diagnostically relevant regions by prioritizing discriminative WSI patches. Our approach implements a progressive three-stage compression strategy: we first leverage FMs for global visual redundancy elimination, and integrate compressed features with language prompts for semantic relevance assessment, then perform neighbor-aware visual token filtering while preserving spatial coherence. Extensive experiments on pathological datasets spanning breast, lung, and ovarian cancers demonstrate its superior performance in few-shot pathology diagnosis. Codes are available at https://github.com/dddavid4real/FOCUS. Zhengrui Guo, Conghao Xiong, Jiabo Ma, Qichen Sun, Lishuang Feng, Jinzhuo Wang, Hao Chen 0011 |
CVPR | 7 |
| 2025 | Distilled Prompt Learning for Incomplete Multimodal Survival PredictionabstractThe integration of multimodal data including pathology images and gene profiles is widely applied in precise survival prediction. Despite recent advances in multimodal survival models, collecting complete modalities for multi-modal fusion still poses a significant challenge, hindering their application in clinical settings. Current approaches tackling incomplete modalities often fall short, as they typically compensate for only a limited part of the knowledge of missing modalities. To address this issue, we propose a Distilled Prompt Learning framework (DisPro) to utilize the strong robustness of Large Language Models (LLMs) to missing modalities, which employs two-stage prompting for compensation of comprehensive information for missing modalities. In the first stage, Unimodal Prompting (UniPro) distills the knowledge distribution of each modality, preparing for supplementing modality-specific knowledge of the missing modality in the subsequent stage. In the second stage, Multimodal Prompting (MultiPro) leverages available modalities as prompts for LLMs to infer the missing modality, which provides modality-common information. Simultaneously, the unimodal knowledge acquired in the first stage is injected into multimodal inference to compensate for the modality-specific knowledge of the missing modality. Extensive experiments covering various missing scenarios demonstrated the superiority of the proposed method. The code is available at https://github.com/Innse/DisPro. Yingxue Xu, Fengtao Zhou, Yihui Wang 0002, Hao Chen 0011 |
CVPR | 6 |
| 2025 | Scaling Tumor Segmentation: Best Lessons from Real and Synthetic Data
Qi Chen 0014, Xinze Zhou, Hao Chen 0011, Zekun Jiang, Ziyan Huang, Dexin Yu, Junjun He, Yefeng Zheng 0001, Ling Shao 0001, Alan L. Yuille, Zongwei Zhou |
ICCV | 4 |
| 2025 | GameGen-X: Interactive Open-world Game Video GenerationabstractWe introduce GameGen-$\mathbb{X}$, the first diffusion transformer model specifically designed for both generating and interactively controlling open-world game videos.
This model facilitates high-quality, open-domain generation by approximating various game elements, such as innovative characters, dynamic environments, complex actions, and diverse events.
Additionally, it provides interactive controllability, predicting and altering future content based on the current clip, thus allowing for gameplay simulation.
To realize this vision, we first collected and built an Open-World Video Game Dataset (OGameData) from scratch.
It is the first and largest dataset for open-world game video generation and control, which comprises over one million diverse gameplay video clips with informative captions.
GameGen-$\mathbb{X}$ undergoes a two-stage training process, consisting of pre-training and instruction tuning.
Firstly, the model was pre-trained via text-to-video generation and video continuation, enabling long-sequence open-domain game video generation with improved fidelity and coherence.
Further, to achieve interactive controllability, we designed InstructNet to incorporate game-related multi-modal control signal experts.
This allows the model to adjust latent representations based on user inputs, advancing the integration of character interaction and scene content control in video generation.
During instruction tuning, only the InstructNet is updated while the pre-trained foundation model is frozen, enabling the integration of interactive controllability without loss of diversity and quality of generated content.
GameGen-$\mathbb{X}$ contributes to advancements in open-world game design using generative models.
It demonstrates the potential of generative models to serve as auxiliary tools to traditional rendering techniques, demonstrating the potential for merging creative generation with interactive capabilities.
The project will be available at https://github.com/GameGen-X/GameGen-X. Haoxuan Che, Xuanhua He, Quande Liu, Cheng Jin 0003, Hao Chen 0011 |
ICLR | 5 |
| 2025 | BLEND: Behavior-guided Neural Population Dynamics Modeling via Privileged Knowledge DistillationabstractModeling the nonlinear dynamics of neuronal populations represents a key pursuit in computational neuroscience. Recent research has increasingly focused on jointly modeling neural activity and behavior to unravel their interconnections. Despite significant efforts, these approaches often necessitate either intricate model designs or oversimplified assumptions. Given the frequent absence of perfectly paired neural-behavioral datasets in real-world scenarios when deploying these models, a critical yet understudied research question emerges: how to develop a model that performs well using only neural activity as input at inference, while benefiting from the insights gained from behavioral signals during training?
To this end, we propose **BLEND**, the **B**ehavior-guided neura**L** population dynamics mod**E**lling framework via privileged k**N**owledge **D**istillation. By considering behavior as privileged information, we train a teacher model that takes both behavior observations (privileged features) and neural activities (regular features) as inputs. A student model is then distilled using only neural activity. Unlike existing methods, our framework is model-agnostic and avoids making strong assumptions about the relationship between behavior and neural activity. This allows BLEND to enhance existing neural dynamics modeling architectures without developing specialized models from scratch. Extensive experiments across neural population activity modeling and transcriptomic neuron identity prediction tasks demonstrate strong capabilities of BLEND, reporting over 50% improvement in behavioral decoding and over 15% improvement in transcriptomic neuron identity prediction after behavior-guided distillation. Furthermore, we empirically explore various behavior-guided distillation strategies within the BLEND framework and present a comprehensive analysis of effectiveness and implications for model performance. Code will be made available at https://github.com/dddavid4real/BLEND. Zhengrui Guo, Fangxu Zhou, Qichen Sun, Lishuang Feng, Jinzhuo Wang, Hao Chen 0011 |
ICLR | 7 |
| 2025 | Context Matters: Query-aware Dynamic Long Sequence Modeling of Gigapixel ImagesabstractWhole slide image (WSI) analysis presents significant computational challenges due to the massive number of patches in gigapixel images. While transformer architectures excel at modeling long-range correlations through self-attention, their quadratic computational complexity makes them impractical for computational pathology applications. Existing solutions like local-global or linear self-attention reduce computational costs but compromise the strong modeling capabilities of full self-attention. In this work, we propose **Querent**, *i.e.*, the **quer**y-awar**e** long co**nt**extual dynamic modeling framework, which achieves a theoretically bounded approximation of full self-attention while delivering practical efficiency. Our method adaptively predicts which surrounding regions are most relevant for each patch, enabling focused yet unrestricted attention computation only with potentially important contexts. By using efficient region-wise metadata computation and importance estimation, our approach dramatically reduces computational overhead while preserving global perception to model fine-grained patch correlations. Through comprehensive experiments on biomarker prediction, gene mutation prediction, cancer subtyping, and survival analysis across over 10 WSI datasets, our method demonstrates superior performance compared to the state-of-the-art approaches. Codes are available at https://github.com/dddavid4real/Querent. Zhengrui Guo, Qichen Sun, Jiabo Ma, Lishuang Feng, Jinzhuo Wang, Hao Chen 0011 |
ICML | 6 |
| 2025 | PARM: Multi-Objective Test-Time Alignment via Preference-Aware Autoregressive Reward ModelabstractMulti-objective test-time alignment aims to adapt large language models (LLMs) to diverse multi-dimensional user preferences during inference while keeping LLMs frozen. Recently, GenARM (Xu et al., 2025) first independently trains Autoregressive Reward Models (ARMs) for each preference dimension without awareness of each other, then combines their outputs based on user-specific preference vectors during inference to achieve multi-objective test-time alignment, leading to two key limitations: the need for multiple ARMs increases the inference cost, and the separate training of ARMs causes the misalignment between the guided generation and the user preferences. To address these issues, we propose Preference-aware ARM (PARM), a single unified ARM trained across all preference dimensions. PARM uses our proposed Preference-Aware Bilinear Low-Rank Adaptation (PBLoRA), which employs a bilinear form to condition the ARM on preference vectors, enabling it to achieve precise control over preference trade-offs during inference. Experiments demonstrate that PARM reduces inference costs and achieves better alignment with preference vectors compared with existing methods. Additionally, PARM enables weak-to-strong guidance, allowing a smaller PARM to guide a larger frozen LLM without expensive training, making multi-objective alignment accessible with limited computing resources. The code is available at https://github.com/Baijiong-Lin/PARM. Baijiong Lin, Weisen Jiang, Yuancheng Xu, Hao Chen 0011, Ying-Cong Chen |
ICML | 4 |
| 2025 | A Survey of Pathology Foundation Model: Progress and Future DirectionsabstractComputational pathology, which involves analyzing whole slide images for automated cancer diagnosis, relies on multiple instance learning, where performance depends heavily on the feature extractor and aggregator. Recent Pathology Foundation Models (PFMs), pretrained on large-scale histopathology data, have significantly enhanced both the extractor and aggregator, but they lack a systematic analysis framework. In this survey, we present a hierarchical taxonomy organizing PFMs through a top-down philosophy applicable to foundation model analysis in any domain: model scope, model pretraining, and model design. Additionally, we systematically categorize PFM evaluation tasks into slide-level, patch-level, multimodal, and biological tasks, providing comprehensive benchmarking criteria. Our analysis identifies critical challenges in both PFM development (pathology-specific methodology, end-to-end pretraining, data-model scalability) and utilization (effective adaptation, model maintenance), paving the way for future directions in this promising field. Resources referenced in this survey are available at https://github.com/BearCleverProud/AwesomeWSI. Conghao Xiong, Hao Chen 0011, Joseph J. Y. Sung |
IJCAI | 2 |
| 2025 | Dia-LLaMA: Towards Large Language Model-Driven CT Report Generation
Zhixuan Chen, Luyang Luo, Yequan Bie, Hao Chen 0011 |
MICCAI (7) | 4 |
| 2025 | Prototype-Guided Cross-Modal Knowledge Enhancement for Adaptive Survival Prediction
Fengchun Liu, Linghan Cai, Zhikang Wang, Zhiyuan Fan, Jin-gang Yu, Hao Chen 0011, Yongbing Zhang 0002 |
MICCAI (6) | 6 |
| 2025 | Explain Any Pathological Concept: Discovering Hierarchical Explanations for Pathology Foundation Models
Shuting Xu, Junlin Hou, Hao Chen 0011 |
MICCAI (6) | 3 |
| 2025 | Global and Local Vision-Language Alignment for Few-Shot Learning and Few-Shot OOD Detection
Xiaoyuan Guan, Wei-Shi Zheng 0001, Hao Chen 0011 |
MICCAI (5) | 4 |
| 2025 | PTCMIL: Multiple Instance Learning via Prompt Token Clustering for Whole Slide Image Analysis
Beidi Zhao, Hao Chen 0011, Zu-hua Gao |
MICCAI (15) | 3 |
| 2025 | Diffusion-Based Virtual Staining from Polarimetric Mueller Matrix Imaging
Jiaxin Zhuang, Jing Cong, Limei Guo, Hao Chen 0011 |
MICCAI (1) | 9 |
| 2025 | Bio2Vol: Adapting 2D Biomedical Foundation Models for Volumetric Medical Image Segmentation
Jiaxin Zhuang, Linshan Wu, Xuefeng Ni, Xi Wang 0013, Liansheng Wang 0002, Hao Chen 0011 |
MICCAI (6) | 6 |
| 2025 | Structure-enhanced deep learning accelerates aptamer selection for small molecule families like steroidsabstractThe efficient discovery of high-affinity small-molecule aptamers via the Systematic Evolution of Ligands by EXponential enrichment (SELEX) is often constrained by challenges in navigating vast sequence spaces and rationally designing initial libraries. In this study, we introduce Deep Learning-assisted SELEX (DL-SELEX), a novel two-step framework that employs variational autoencoders (VAEs) to accelerate and refine small-molecule aptamer selection. This approach is the first to integrate deep learning to design initial aptamer libraries, marking a significant advancement in SELEX workflows. DL-SELEX leverages shared structural features within molecular families (e.g. steroids) to guide aptamer design: AptaVAE, the first VAE enriched with transfer learning from foundation models, generates tailored initial pools, whereas AptaClux, a second VAE, identifies high-performance candidates from SELEX-derived next-generation sequencing (NGS) data by capturing consensus structural features. The application of DL-SELEX to hydrocortisone (CS) and testosterone (TES) yielded aptamers with up to 450-fold higher affinity than previously reported aptamers and reduced SELEX iterations by up to 80%. Critically, these results demonstrate that structural commonalities can be used to train deep learning models to design aptamers for structurally similar targets. DL-SELEX provides an effective, generalizable strategy to streamline aptamer discovery and enables de novo design of high-affinity aptamers for challenging small molecules. Zibin Zhao, Haosi Lin, Hoi Ying Lau, Hao Chen 0011, I-Ming Hsing |
Briefings Bioinform. | 4 |
| 2025 | An evaluation method for model transfer learning performance in industrial surface defect detection tasks
Guizhong Fu, Zengguang Zhang, Jinbin Li, Enrui Zhang, Zewei He, Fangyuan Sun, Qixin Zhu, Fuzhou Niu, Hao Chen 0011, Yehu Shen |
Expert Syst. Appl. | 9 |
| 2025 | Image Captions are Natural Prompts for Training Data Synthesis
Shiye Lei, Hao Chen 0011, Sen Zhang 0006, Bo Zhao 0015, Dacheng Tao |
Int. J. Comput. Vis. | 2 |
| 2025 | MedIAnomaly: A comparative study of anomaly detection in medical images
Yu Cai 0005, Hao Chen 0011, Kwang-Ting Cheng |
Medical Image Anal. | 3 |
| 2025 | CXR-LT 2024: A MICCAI challenge on long-tailed, multi-label, and zero-shot disease classification from chest X-ray
Mingquan Lin, Gregory Holste, Song Wang 0026, Yiliang Zhou, Yishu Wei, Imon Banerjee, Pengyi Chen, Tianjie Dai, Yuexi Du, Nicha C. Dvornek, Yuyan Ge, Zuwei Guo, Shohei Hanaoka, Dongkyun Kim, Pablo Messina, Yang Lu 0009, Denis Parra, Donghyun Son, Alvaro Soto, Aisha Urooj Khan, René Vidal, Yosuke Yamagishi, Pingkun Yan, Zefan Yang, Ruichi Zhang, Yang Zhou 0019, Leo A. Celi, Ronald M. Summers, Zhiyong Lu, Hao Chen 0011, Adam E. Flanders, George Shih, Zhangyang Wang, Yifan Peng 0002 |
Medical Image Anal. | 30 |
| 2025 | Rethinking boundary detection in deep learning-based medical image segmentation
Yi Lin 0009, Kwang-Ting Cheng, Hao Chen 0011 |
Medical Image Anal. | 6 |
| 2025 | Mitigating medical dataset bias by learning adaptive agreement from a biased council
Luyang Luo, Zhuoyue Wan, Wanteng Ma, Hao Chen 0011 |
Medical Image Anal. | 6 |
| 2025 | Learning robust medical image segmentation from multi-source annotations
Luyang Luo, Mingxiang Wu, Qiong Wang 0001, Hao Chen 0011 |
Medical Image Anal. | 5 |
| 2025 | Modeling the Label Distributions for Weakly-Supervised Semantic SegmentationabstractWeakly-Supervised Semantic Segmentation (WSSS) aims to train segmentation models by weak labels, which is receiving significant attention due to its low annotation cost. Existing approaches focus on generating pseudo labels for supervision while largely ignoring to leverage the inherent semantic correlation among different pseudo labels. We observe that pseudo-labeled pixels that are close to each other in the feature space are more likely to share the same class, and those closer to the distribution centers tend to have higher confidence. Motivated by this, we propose to model the underlying label distributions and employ cross-label constraints to generate more accurate pseudo labels. In this paper, we develop a unified WSSS framework named Adaptive Gaussian Mixtures Model, which leverages a GMM to model the label distributions. Specifically, we calculate the feature distribution centers of pseudo-labeled pixels and build the GMM by measuring the distance between the centers and each pseudo-labeled pixel. Then, we introduce an Online Expectation-Maximization (OEM) algorithm and a novel maximization loss to optimize the GMM adaptively, aiming to learn more discriminative decision boundaries between different class-wise Gaussian mixtures. Based on the label distributions, we leverage the GMM to generate high-quality pseudo labels for more reliable supervision. Our framework is capable of solving different forms of weak labels: image-level labels, points, scribbles, blocks, and bounding-boxes. Extensive experiments on PASCAL, COCO, Cityscapes, and ADE20 K datasets demonstrate that our framework can effectively provide more reliable supervision and outperform the state-of-the-art methods under all settings. Linshan Wu, Zhun Zhong, Jiayi Ma 0001, Yunchao Wei, Hao Chen 0011, Leyuan Fang, Shutao Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Efficient Breast Lesion Segmentation From Ultrasound Videos Across Multiple Source-Limited PlatformsabstractMedical video segmentation is fundamentally important in clinical diagnosis and treatment procedures, offering dynamic tracking of breast lesions across frames in ultrasound videos for improved segmentation performance. However, existing approaches face challenges in striking a balance between segmentation performance and inference speed, hindering real-time application in resource-constrained medical environments. In order to address these limitations, we present BaS, a blazing-fast on-device breast lesion segmentation model. BaS integrates the Stem module and BaSBlock to refine representations through inter- and intra-frame analysis on ultrasound videos. In addition, we release two versions of BaS: the BaS-S for superior segmentation performance and the BaS-L for accelerated inference times. Experimental Results indicate that BaS surpasses the top-performing models in terms of segmenting efficiency and accuracy of predictions on devices with limited resources. This work advances the development of efficient medical video segmentation frameworks applicable to multiple medical platforms. Teng Huang 0001, Ziyu Ding, Hao Chen 0011, Baoliang Zhao, Ying Hu 0001, Qiong Wang 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | Online Self-Distillation and Self-Modeling for 3D Brain Tumor SegmentationabstractIn the specialized domain of brain tumor segmentation, supervised segmentation approaches are hindered by the limited availability of high-quality labeled data, a condition arising from data privacy concerns, significant costs, and ethical issues. In response to this challenge, this paper presents a training framework that adeptly integrates a plug-and-play component, MOD, into current supervised learning models, boosting their efficacy in scenarios with limited data. The MOD consists of an Online Tokenizer and a Dense Predictor, which employs self-distillation and self-modeling on masked patches, promoting swift convergence and efficient representation learning. During the inference phase, the plug-and-play MOD component is excluded, preserving the computational efficiency of the original model without incurring extra processing costs. We substantiated the value of our approach through experiments on leading 3D brain tumor segmentation baselines. Remarkably, models augmented with the MOD consistently showcased superior results, achieving elevated Dice coefficients and HD95 scores on two datasets: BraTS 2021 and MSD 2019 Task-01 Brain Tumor. Teng Huang 0001, Zhen Wang 0037, Changyu Dong, Dongyang Kuang, Ying Hu 0001, Hao Chen 0011, Tim C. Lei, Qiong Wang 0001 |
IEEE J. Biomed. Health Informatics | 9 |
| 2025 | Unpaired Optical Coherence Tomography Angiography Image Super-Resolution via Frequency-Aware Inverse-Consistency GANabstractFor optical coherence tomography angiography (OCTA) images, the limited scanning rate leads to a trade-off between field-of-view (FOV) and imaging resolution. Although larger FOV images may reveal more parafoveal vascular lesions, their application is hampered due to lower resolution. To increase the resolution, previous works only achieved satisfactory performance by using paired data for training, but real-world applications are limited by the challenge of collecting large-scale paired images. Thus, an unpaired approach is highly demanded. Generative Adversarial Network (GAN) has been commonly used in the unpaired setting, but it may struggle to accurately preserve fine-grained capillary details, which are critical biomarkers for OCTA. In this paper, our approach aspires to preserve these details by leveraging the frequency information, which represents details as high-frequencies (${\bm {hf}}$) and coarse-grained features as low-frequencies (${\bm {lf}}$). We propose a GAN-based unpaired super-resolution method for OCTA images and exceptionally emphasize ${\bm {hf}}$ fine capillaries through a dual-path generator. To facilitate a precise spectrum of the reconstructed image, we also propose a frequency-aware adversarial loss for the discriminator and introduce a frequency-aware focal consistency loss for end-to-end optimization. We collected a paired dataset for evaluation and showed that our method outperforms other state-of-the-art unpaired methods both quantitatively and visually. Haoxuan Che, An-ran Ran, Carol Y. Cheung, Hao Chen 0011 |
IEEE J. Biomed. Health Informatics | 6 |
| 2025 | FedDAG: Federated Domain Adversarial Generation Toward Generalizable Medical Image AnalysisabstractFederated domain generalization aims to train a global model from multiple source domains and ensure its generalization ability to unseen target domains. Due to the target domain being with unknown domain shifts, attempting to approximate these gaps by source domains may be the key to improving model generalization capability. Existing works mainly focus on sharing and recombining local domain-specific attributes to increase data diversity and simulate potential domain shifts. However, these methods may be insufficient since only the local attribute recombination can be hard to touch the out-of-distribution of global data. In this paper, we propose a simple-yet-efficient framework named Federated Domain Adversarial Generation (FedDAG). It aims to simulate the domain shift and improve the model generalization by adversarially generating novel domains different from local and global source domains. Specifically, it generates novel-style images by maximizing the instance-level feature discrepancy between original and generated images and trains a generalizable task model by minimizing their feature discrepancy. Further, we observed that FedDAG could cause different performance improvements for local models. It may be due to inherent data isolation and heterogeneity among clients, exacerbating the imbalance in their generalization contributions to the global model. Ignoring this imbalance can lead the global model's generalization ability to be sub-optimal, further limiting the novel domain generation procedure. Thus, to mitigate this imbalance, FedDAG hierarchically aggregates local models at the within-client and across-client levels by using the sharpness concept to evaluate client model generalization contributions. Extensive experiments across four medical benchmarks demonstrate FedDAG's ability to enhance generalization in federated medical scenarios. Haoxuan Che, Haibo Jin, Yong Xia 0001, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Large Language Model With Region-Guided Referring and Grounding for CT Report GenerationabstractComputed tomography (CT) report generation is crucial to assist radiologists in interpreting CT volumes, which can be time-consuming and labor-intensive. Existing methods primarily only consider the global features of the entire volume, making it struggle to focus on specific regions and potentially missing abnormalities. To address this issue, we propose Reg2RG, the first region-guided referring and grounding framework for CT report generation, which enhances diagnostic performance by focusing on anatomical regions within the volume. Specifically, we utilize masks from a universal segmentation module to capture local features for each referring region. A local feature decoupling (LFD) strategy is proposed to preserve the local high-resolution details with little computational overhead. Then the local features are integrated with global features to capture inter-regional relationships within a cohesive context. Moreover, we propose a novel region-report alignment (RRA) training strategy. It leverages the recognition of referring regions to guide the generation of region-specific reports, enhancing the model's referring and grounding capabilities while also improving the report's interpretability. A large language model (LLM) is further employed as the language decoder to generate reports from integrated visual features, facilitating region-level comprehension. Extensive experiments on two large-scale chest CT-report datasets demonstrate the superiority of our method, which outperforms several state-of-the-art methods in terms of both natural language generation and clinical efficacy metrics while preserving promising interpretability. The code is available at https://github.com/zhi-xuan-chen/Reg2RG. Zhi-Xuan Chen, Yequan Bie, Haibo Jin, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | QMix: Quality-Aware Learning With Mixed Noise for Robust Retinal Disease DiagnosisabstractDue to the complex nature of medical image acquisition and annotation, medical datasets inevitably contain noise. This adversely affects the robustness and generalization of deep neural networks. Previous noise learning methods mainly considered noise arising from images being mislabeled, i.e., label noise, assuming all mislabeled images were of high quality. However, medical images can also suffer from severe data quality issues, i.e., data noise, where discriminative visual features for disease diagnosis are missing. In this paper, we propose QMix, a noise learning framework that learns a robust disease diagnosis model under mixed noise scenarios. QMix alternates between sample separation and quality-aware semi-supervised training in each epoch. The sample separation phase uses a joint uncertainty-loss criterion to effectively separate (1) correctly labeled images, (2) mislabeled high-quality images, and (3) mislabeled low-quality images. The semi-supervised training phase then learns a robust disease diagnosis model from the separated samples. Specifically, we propose a sample-reweighing loss to mitigate the effect of mislabeled low-quality images during training, and a contrastive enhancement loss to further distinguish them from correctly labeled images. QMix achieved state-of-the-art performance on six public retinal image datasets and exhibited significant improvements in robustness against mixed noise. Code will be available upon acceptance. Junlin Hou, Jilan Xu, Rui Feng 0001, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | A Chain of Diagnosis Framework for Accurate and Explainable Radiology Report GenerationabstractDespite the progress of radiology report generation (RRG), existing works face two challenges: 1) The performances in clinical efficacy are unsatisfactory, especially for lesion attributes description; 2) the generated text lacks explainability, making it difficult for radiologists to trust the results. To address the challenges, we focus on a trustworthy RRG model, which not only generates accurate descriptions of abnormalities, but also provides basis of its predictions. To this end, we propose a framework named chain of diagnosis (CoD), which maintains a chain of diagnostic process for clinically accurate and explainable RRG. It first generates question-answer (QA) pairs via diagnostic conversation to extract key findings, then prompts a large language model with QA diagnoses for accurate generation. To enhance explainability, a diagnosis grounding module is designed to match QA diagnoses and generated sentences, where the diagnoses act as a reference. Moreover, a lesion grounding module is designed to locate abnormalities in the image, further improving the working efficiency of radiologists. To facilitate label-efficient training, we propose an omni-supervised learning strategy with clinical consistency to leverage various types of annotations from different datasets. Our efforts lead to 1) an omni-labeled RRG dataset with QA pairs and lesion boxes; 2) a evaluation tool for assessing the accuracy of reports in describing lesion location and severity; 3) extensive experiments to demonstrate the effectiveness of CoD, where it outperforms both specialist and generalist models consistently on two RRG benchmarks and shows promising explainability by accurately grounding generated sentences to QA diagnoses and images. Haibo Jin, Haoxuan Che, Sunan He, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | HMIL: Hierarchical Multi-Instance Learning for Fine-Grained Whole Slide Image ClassificationabstractFine-grained classification of whole slide images (WSIs) is essential in precision oncology, enabling precise cancer diagnosis and personalized treatment strategies. The core of this task involves distinguishing subtle morphological variations within the same broad category of gigapixel-resolution images, which presents a significant challenge. While the multi-instance learning (MIL) paradigm alleviates the computational burden of WSIs, existing MIL methods often overlook hierarchical label correlations, treating fine-grained classification as a flat multi-class classification task. To overcome these limitations, we introduce a novel hierarchical multi-instance learning (HMIL) framework. By facilitating on the hierarchical alignment of inherent relationships between different hierarchy of labels at instance and bag level, our approach provides a more structured and informative learning process. Specifically, HMIL incorporates a class-wise attention mechanism that aligns hierarchical information at both the instance and bag levels. Furthermore, we introduce supervised contrastive learning to enhance the discriminative capability for fine-grained classification and a curriculum-based dynamic weighting module to adaptively balance the hierarchical feature during training. Extensive experiments on our large-scale cytology cervical cancer (CCC) dataset and two public histology datasets, BRACS and PANDA, demonstrate the state-of-the-art class-wise and overall performance of our HMIL framework. Our source code is available at https://github.com/ChengJin-git/HMIL. Cheng Jin 0003, Luyang Luo, Huangjing Lin, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Boosting Convolution With Efficient MLP-Permutation for Volumetric Medical Image SegmentationabstractRecently, the advent of Vision Transformer (ViT) has brought substantial advancements in 3D benchmarks, particularly in 3D volumetric medical image segmentation (Vol-MedSeg). Concurrently, multi-layer perceptron (MLP) network has regained popularity among researchers due to their comparable results to ViT, albeit with the exclusion of the resource-intensive self-attention module. In this work, we propose a novel permutable hybrid network for Vol-MedSeg, named PHNet, which capitalizes on the strengths of both convolution neural networks (CNNs) and MLP. PHNet addresses the intrinsic anisotropy problem of 3D volumetric data by employing a combination of 2D and 3D CNNs to extract local features. Besides, we propose an efficient multi-layer permute perceptron (MLPP) module that captures long-range dependence while preserving positional information. This is achieved through an axis decomposition operation that permutes the input tensor along different axes, thereby enabling the separate encoding of the positional information. Furthermore, MLPP tackles the resolution sensitivity issue of MLP in Vol-MedSeg with a token segmentation operation, which divides the feature into smaller tokens and processes them individually. Extensive experimental results validate that PHNet outperformed the state-of-the-art methods with lower computational costs on the widely-used yet challenging COVID-19-20, Synapse, LiTS and MSD BraTS benchmarks. The ablation study also demonstrated the effectiveness of PHNet in harnessing the strengths of both CNNs and MLP. The code is available on Github: https://github.com/xiaofang007/PHNet. Yi Lin 0009, Kwang-Ting Cheng, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Histo-Genomic Knowledge Association for Cancer Prognosis From Histopathology Whole Slide ImagesabstractHisto-genomic multi-modal methods have emerged as a powerful paradigm, demonstrating significant potential for cancer prognosis. However, genome sequencing, unlike histopathology imaging, is still not widely accessible in underdeveloped regions, limiting the application of these multi-modal approaches in clinical settings. To address this, we propose a novel Genome-informed Hyper-Attention Network, termed G-HANet, which is capable of effectively learning the histo-genomic associations during training to elevate uni-modal whole slide image (WSI)-based inference for the first time. Compared with the potential knowledge distillation strategy for this setting (i.e., distilling a multi-modal network to a uni-modal network), our end-to-end model is superior in training efficiency and learning cross-modal interactions. Specifically, the network comprises cross-modal associating branch (CAB) and hyper-attention survival branch (HSB). Through the genomic data reconstruction from WSIs, CAB effectively distills the associations between functional genotypes and morphological phenotypes and offers insights into the gene expression profiles in the feature space. Subsequently, HSB leverages the distilled histo-genomic associations as well as the generated morphology-based weights to achieve the hyper-attention modeling of the patients from both histopathology and genomic perspectives to improve cancer prognosis. Extensive experiments are conducted on five TCGA benchmarking datasets and the results demonstrate that G-HANet significantly outperforms the state-of-the-art WSI-based methods and achieves competitive performance with genome-based and multi-modal methods. G-HANet is expected to be explored as a useful tool by the research community to address the current bottleneck of insufficient histo-genomic data pairing in the context of cancer prognosis and precision oncology. The code is available at https://github.com/ZacharyWang-007/G-HANet. Zhikang Wang, Yingxue Xu, Seiya Imoto, Hao Chen 0011, Jiangning Song |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Shapley Values-Enabled Progressive Pseudo Bag Augmentation for Whole-Slide Image ClassificationabstractIn computational pathology, whole-slide image (WSI) classification presents a formidable challenge due to its gigapixel resolution and limited fine-grained annotations. Multiple-instance learning (MIL) offers a weakly supervised solution, yet refining instance-level information from bag-level labels remains challenging. While most of the conventional MIL methods use attention scores to estimate instance importance scores (IIS) which contribute to the prediction of the slide labels, these often lead to skewed attention distributions and inaccuracies in identifying crucial instances. To address these issues, we propose a new approach inspired by cooperative game theory: employing Shapley values to assess each instance's contribution, thereby improving IIS estimation. The computation of the Shapley value is then accelerated using attention, meanwhile retaining the enhanced instance identification and prioritization. We further introduce a framework for the progressive assignment of pseudo bags based on estimated IIS, encouraging more balanced attention distributions in MIL models. Our extensive experiments on CAMELYON-16, BRACS, TCGA-LUNG, and TCGA-BRCA datasets show our method's superiority over existing state-of-the-art approaches, offering enhanced interpretability and class-wise insights. Our source code is available at https://github.com/RenaoYan/PMIL. Renao Yan, Qiehe Sun, Cheng Jin 0003, Yonghong He, Tian Guan, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Deep Rib Fracture Instance Segmentation and Classification From CT on the RibFrac ChallengeabstractRib fractures are a common and potentially severe injury that can be challenging and labor-intensive to detect in CT scans. While there have been efforts to address this field, the lack of large-scale annotated datasets and evaluation benchmarks has hindered the development and validation of deep learning algorithms. To address this issue, the RibFrac Challenge was introduced, providing a benchmark dataset of over 5,000 rib fractures from 660 CT scans, with voxel-level instance mask annotations and diagnosis labels for four clinical categories (buckle, nondisplaced, displaced, or segmental). The challenge includes two tracks: a detection (instance segmentation) track evaluated by an FROC-style metric and a classification track evaluated by an F1-style metric. During the MICCAI 2020 challenge period, 243 results were evaluated, and seven teams were invited to participate in the challenge summary. The analysis revealed that several top rib fracture detection solutions achieved performance comparable or even better than human experts. Nevertheless, the current rib fracture classification solutions are hardly clinically applicable, which can be an interesting area in the future. As an active benchmark and research resource, the data and online evaluation of the RibFrac Challenge are available at the challenge website (https://ribfrac.grand-challenge.org/). In addition, we further analyzed the impact of two post-challenge advancements-large-scale pretraining and rib segmentation-based on our internal baseline for rib fracture detection. These findings lay a foundation for future research and development in AI-assisted rib fracture diagnosis. Jiancheng Yang, Kaiming Kuang, Donglai Wei 0001, Shixuan Gu, Jianying Liu, Zhizhong Chai, Yongjie Xiao, Hao Chen 0011, Liming Xu, Bang Du, Xiangyi Yan, Hao Tang 0010, Adam M. Alessio, Gregory Holste, Jianye He, Lixuan Che, Hanspeter Pfister, Ming Li 0005, Bingbing Ni |
IEEE Trans. Medical Imaging | 12 |
| 2025 | Cohort-Individual Cooperative Learning for Multimodal Cancer Survival AnalysisabstractRecently, we have witnessed impressive achievements in cancer survival analysis by integrating multimodal data, e.g., pathology images and genomic profiles. However, the heterogeneity and high dimensionality of these modalities pose significant challenges in extracting discriminative representations while maintaining good generalization. In this paper, we propose a Cohort-individual Cooperative Learning (CCL) framework to advance cancer survival analysis by collaborating knowledge decomposition and cohort guidance. Specifically, first, we propose a Multimodal Knowledge Decomposition (MKD) module to explicitly decompose multimodal knowledge into four distinct components: redundancy, synergy, and uniqueness of the two modalities. Such a comprehensive decomposition can enlighten the models to perceive easily overlooked yet important information, facilitating an effective multimodal fusion. Second, we propose a Cohort Guidance Modeling (CGM) to mitigate the risk of overfitting task-irrelevant information. It can promote a more comprehensive and robust understanding of the underlying multimodal data while avoiding the pitfalls of overfitting and enhancing the generalization ability of the model. By cooperating with the knowledge decomposition and cohort guidance methods, we develop a robust multimodal survival analysis model with enhanced discrimination and generalization abilities. Extensive experimental results on five cancer datasets demonstrate the effectiveness of our model in integrating multimodal data for survival analysis. Our code is available at https://github.com/moothes/CCL-survival. Huajun Zhou, Fengtao Zhou, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Advancing Volumetric Medical Image Segmentation via Global-Local Masked AutoencodersabstractMasked Autoencoder (MAE) is a self-supervised pre-training technique that holds promise in improving the representation learning of neural networks. However, the current application of MAE directly to volumetric medical images poses two challenges: (i) insufficient global information for clinical context understanding of the holistic data, and (ii) the absence of any assurance of stabilizing the representations learned from randomly masked inputs. To conquer these limitations, we propose the Global-Local Masked AutoEncoders (GL-MAE), a simple yet effective self-supervised pre-training strategy. GL-MAE acquires robust anatomical structure features by incorporating multi-level reconstruction from fine-grained local details to high-level global semantics. Furthermore, a complete global view serves as an anchor to direct anatomical semantic alignment and stabilize the learning process through global-to-global consistency learning and global-to-local consistency learning. Our fine-tuning results on eight mainstream public datasets demonstrate the superiority of our method over other state-of-the-art self-supervised algorithms, highlighting its effectiveness on versatile volumetric medical image segmentation and classification tasks. We will release codes upon acceptance at https://github.com/JiaxinZhuang/GL-MAE. Jiaxin Zhuang, Luyang Luo, Qiong Wang 0001, Mingxiang Wu, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | MiM: Mask in Mask Self-Supervised Pre-Training for 3D Medical Image AnalysisabstractThe Vision Transformer (ViT) has demonstrated remarkable performance in Self-Supervised Learning (SSL) for 3D medical image analysis. Masked AutoEncoder (MAE) for feature pre-training can further unleash the potential of ViT on various medical vision tasks. However, due to large spatial sizes with much higher dimensions of 3D medical images, the lack of hierarchical design for MAE may hinder the performance of downstream tasks. In this paper, we propose a novel Mask in Mask (MiM) pre-training framework for 3D medical images, which aims to advance MAE by learning discriminative representation from hierarchical visual tokens across varying scales. We introduce multiple levels of granularity for masked inputs from the volume, which are then reconstructed simultaneously ranging at both fine and coarse levels. Additionally, a cross-level alignment mechanism is applied to adjacent level volumes to enforce anatomical similarity hierarchically. Furthermore, we adopt a hybrid backbone to enhance the hierarchical representation learning efficiently during the pre-training. MiM was pre-trained on a large scale of available 3D volumetric images, i.e., Computed Tomography (CT) images containing various body parts. Extensive experiments on twelve public datasets demonstrate the superiority of MiM over other SSL methods in organ/tumor segmentation and disease classification. We further scale up the MiM to large pre-training datasets with more than 10k volumes, showing that large-scale pre-training can further enhance the performance of downstream tasks. Code is available at https://github.com/JiaxinZhuang/MiM. Jiaxin Zhuang, Linshan Wu, Qiong Wang 0001, Peng Fei, Varut Vardhanabhuti, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Scale-Aware Super-Resolution Network With Dual Affinity Learning for Lesion Segmentation From Medical ImagesabstractConvolutional neural networks (CNNs) have shown remarkable progress in medical image segmentation. However, the lesion segmentation remains a challenge to state-of-the-art CNN-based algorithms due to the variance in scales and shapes. On the one hand, tiny lesions are hard to delineate precisely from the medical images which are often of low resolutions. On the other hand, segmenting large-size lesions requires large receptive fields, which exacerbates the first challenge. In this article, we present a scale-aware super-resolution (SR) network to adaptively segment lesions of various sizes from low-resolution (LR) medical images. Our proposed network contains dual branches to simultaneously conduct lesion mask SR (LMSR) and lesion image SR (LISR). Meanwhile, we introduce scale-aware dilated convolution (SDC) blocks into the multitask decoders to adaptively adjust the receptive fields of the convolutional kernels according to the lesion sizes. To guide the segmentation branch to learn from richer high-resolution (HR) features, we propose a feature affinity (FA) module and a scale affinity (SA) module to enhance the multitask learning of the dual branches. On multiple challenging lesion segmentation datasets, our proposed network achieved consistent improvements compared with other state-of-the-art methods. Code will be available at: https://github.com/poiuohke/SASR_Net. Luyang Luo, Yanwen Li, Zhizhong Chai, Huangjing Lin, Pheng-Ann Heng, Hao Chen 0011 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Reference-Based OCT Angiogram Super-Resolution With Learnable Texture GenerationabstractOptical coherence tomography angiography (OCTA) can visualize retinal microvasculature and is important to qualitatively and quantitatively identify potential biomarkers for different retinal diseases. However, the resolution of optical coherence tomography (OCT) angiograms inevitably decreases when increasing the field-of-view (FOV) given a fixed acquisition time. To address this issue, we propose a novel reference-based super-resolution (RefSR) framework to preserve the resolution of the OCT angiograms while increasing the scanning area. Specifically, textures from the normal RefSR pipeline are used to train a learnable texture generator (LTG), which is designed to generate textures according to the input. The key difference between the proposed method and traditional RefSR models is that the textures used during inference are generated by the LTG instead of being searched from a single reference (Ref) image. Since the LTG is optimized throughout the whole training process, the available texture space is significantly enlarged and no longer limited to a single Ref image, but extends to all textures contained in the training samples. Moreover, our proposed LTGNet does not require an Ref image at the inference phase, thereby becoming invulnerable to the selection of the Ref image. Both experimental and visual results show that LTGNet has competitive performance and robustness over state-of-the-art methods, indicating good reliability and promise in real-life deployment. The source code is available at https://github.com/RYY0722/LTGNet. Yuyan Ruan, Ziqi Tang, An-ran Ran, Carol Y. Cheung, Hao Chen 0011 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2025 | Real-Time Semantic Segmentation via a Densely Aggregated Bilateral NetworkabstractWith the growing demands of applications on online devices, the speed-accuracy trade-off is critical in the semantic segmentation system. Recently, the bilateral segmentation network has shown promising capacity to achieve the balance between favorable accuracy and fast speed, and has become the mainstream backbone in real-time semantic segmentation. Segmentation of target objects relies on high-level semantics, whereas it requires detailed low-level features to model specific local patterns for accurate location. However, the lightweight backbone of bilateral architecture limits the extraction of semantic context and spatial details. And the late fusion of the bilateral streams incurs the insufficient aggregation of semantic context and spatial details. In this article, we propose a densely aggregated bilateral network (DAB-Net) for real-time semantic segmentation. In the context path, a patchwise context enhancement (PCE) module is proposed to efficiently capture the local semantic contextual information from spatialwise and channelwise, respectively. Meanwhile, a context-guided spatial path (CGSP) is designed to exploit more spatial information by encoding finer details from the raw image and the transition from the context path. Finally, with multiple interactions between bilateral branches, the intertwined outputs from bilateral streams are combined in a unified decoder for a final interaction to further enhance the feature representation, which generates the final segmentation prediction. Experimental results on three public benchmarks demonstrate that our proposed method achieves higher accuracy with a limited decay in speed, which performs favorably against state-of-the-art real-time approaches and runs at 31.1 frames/s (FPS) on the high resolution of . The source code is released at https://github.com/isyangshu/DABNet. Shu Yang 0004, Lu Zhang 0053, Shuai Liu 0009, Huchuan Lu, Hao Chen 0011 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Deep contour attention learning for scleral deformation from OCT images
Hao Chen 0011, Yupeng Xu, Huating Li, Yuan Xie 0006, David Dagan Feng, Jinman Kim, Lei Bi 0001, Xiangui He, Bin Sheng 0001 |
Vis. Comput. | 2 |
| 2024 | MICA: Towards Explainable Skin Lesion Diagnosis via Multi-Level Image-Concept AlignmentabstractBlack-box deep learning approaches have showcased significant potential in the realm of medical image analysis. However, the stringent trustworthiness requirements intrinsic to the medical field have catalyzed research into the utilization of Explainable Artificial Intelligence (XAI), with a particular focus on concept-based methods. Existing concept-based methods predominantly apply concept annotations from a single perspective (e.g., global level), neglecting the nuanced semantic relationships between sub-regions and concepts embedded within medical images. This leads to underutilization of the valuable medical information and may cause models to fall short in harmoniously balancing interpretability and performance when employing inherently interpretable architectures such as Concept Bottlenecks. To mitigate these shortcomings, we propose a multi-modal explainable disease diagnosis framework that meticulously aligns medical images and clinical-related concepts semantically at multiple strata, encompassing the image level, token level, and concept level. Moreover, our method allows for model intervention and offers both textual and visual explanations in terms of human-interpretable concepts. Experimental results on three skin image datasets demonstrate that our method, while preserving model interpretability, attains high performance and label efficiency for concept detection and disease diagnosis. The code is available at https://github.com/Tommy-Bie/MICA. Yequan Bie, Luyang Luo, Hao Chen 0011 |
AAAI | 3 |
| 2024 | PromptMRG: Diagnosis-Driven Prompts for Medical Report GenerationabstractAutomatic medical report generation (MRG) is of great research value as it has the potential to relieve radiologists from the heavy burden of report writing. Despite recent advancements, accurate MRG remains challenging due to the need for precise clinical understanding and disease identification. Moreover, the imbalanced distribution of diseases makes the challenge even more pronounced, as rare diseases are underrepresented in training data, making their diagnosis unreliable. To address these challenges, we propose diagnosis-driven prompts for medical report generation (PromptMRG), a novel framework that aims to improve the diagnostic accuracy of MRG with the guidance of diagnosis-aware prompts. Specifically, PromptMRG is based on encoder-decoder architecture with an extra disease classification branch. When generating reports, the diagnostic results from the classification branch are converted into token prompts to explicitly guide the generation process. To further improve the diagnostic accuracy, we design cross-modal feature enhancement, which retrieves similar reports from the database to assist the diagnosis of a query image by leveraging the knowledge from a pre-trained CLIP. Moreover, the disease imbalanced issue is addressed by applying an adaptive logit-adjusted loss to the classification branch based on the individual learning status of each disease, which overcomes the barrier of text decoder's inability to manipulate disease distributions. Experiments on two MRG benchmarks show the effectiveness of the proposed method, where it obtains state-of-the-art clinical efficacy performance on both datasets. Haibo Jin, Haoxuan Che, Yi Lin 0009, Hao Chen 0011 |
AAAI | 4 |
| 2024 | Holistic and Historical Instance Comparison for Cervical Cell DetectionabstractCytology screening from Papanicolaou (Pap) smears is a common and effective tool for the preventive clinical management of cervical cancer, where abnormal cell detection from whole slide images serves as the foundation for reporting cervical cytology. However, cervical cell detection remains challenging due to 1) hazily-defined cell types (e.g., ASC-US) with subtle morphological discrepancies caused by the dynamic cancerization process, i.e., cell class ambiguity, and 2) imbalanced class distributions of clinical data may cause missed detection, especially for minor categories, i.e., cell class imbalance. To this end, we propose a holistic and historical instance comparison approach for cervical cell detection. Specifically, we first develop a holistic instance comparison scheme enforcing both RoI-level and class-level cell discrimination. This coarse-to-fine cell comparison encourages the model to learn foreground-distinguishable and class-wise representations. To emphatically improve the distinguishability of minor classes, we then introduce a historical instance comparison scheme with a confident sample selection-based memory bank, which involves comparing current embeddings with historical embeddings for better cell instance discrimination. Extensive experiments and analysis on two large-scale cytology datasets including 42,592 and 114,513 cervical cells demonstrate the effectiveness of our method. The code is available at https://github.com/hjiangaz/HERO. Hao Jiang 0028, Runsheng Liu, Yanning Zhou 0003, Huangjing Lin, Hao Chen 0011 |
BIBM | 5 |
| 2024 | Bootstrapping Radiography Pre-training via Siamese Masked Vision-Language Modeling with Complementary Self-distillationabstractDiagnosing thoracic diseases from chest X-rays (CXR) using deep learning faces unique challenges due to the high anatomical similarity across images and the critical nature of minute anomalies. In this paper, we introduce a novel self-supervised learning framework, Siamese Masked Vision-Language Modeling with Complementary Self-distillation (SMVLM), designed to enhance disease diagnosis in CXR by addressing these specific challenges. First, we tackle the problem of high anatomical similarity and subtle variance in CXR by employing a complementary masking strategy in a Siamese network setup, which forces the model to focus on subtle, yet diagnostically relevant features from both global and local perspectives. Secondly, we enrich our model’s learning capabilities by integrating a dual masked image-text contrastive loss that aligns radiographic findings with their corresponding text report, harnessing the synergistic potential of multimodal data. In conjunction with this, cross-modal image-text pre-reconstruction with registers is introduced to deepen the contextual understanding of the CXR images and radiology report, ensuring a comprehensive feature representation. Extensive evaluations on benchmark datasets demonstrate that our method significantly outperforms existing approaches, providing a robust solution for the accurate and reliable diagnosis of thoracic diseases from CXR images. Luyang Luo, Hao Chen 0011 |
BIBM | 3 |
| 2024 | DKINet: Medication Recommendation via Domain Knowledge Informed Deep LearningabstractMedication recommendation is a fundamental yet crucial branch of healthcare that presents opportunities to assist physicians in making more accurate medication prescriptions for patients with complex health conditions. Previous studies have primarily concentrated on deriving patient representations from electronic health records (EHRs) to recommend medications, often overlooking the effective integration of domain-specific prior knowledge. However, integrating domain knowledge with the patient’s clinical manifestations can be challenging, particularly when dealing with complex clinical manifestations. Therefore, in this paper, we first identify comprehensive domain-specific prior knowledge, namely the Unified Medical Language System (UMLS), which is a comprehensive repository of biomedical vocabularies and standards, for knowledge extraction. Subsequently, we propose a knowledge injection module that addresses the effective integration of domain knowledge with complex clinical manifestations, enabling an effective characterization of the health conditions of the patient. Moreover, acknowledging the influence of historical medications on patients’ current treatments, we propose a historical medication-aware patient representation module to capture the longitudinal influence of historical medication information on the representation of current patients. Extensive experiments on three publicly benchmark datasets verify the superiority of our proposed method, which outperformed other methods by a significant margin. The code is available at: https://github.com/sherry6247/DKINet. Sicen Liu, Xiaolong Wang 0001, Xianbing Zhao, Hao Chen 0011 |
BIBM | 4 |
| 2024 | GAInS: Gradient Anomaly-aware Biomedical Instance SegmentationabstractInstance segmentation plays a vital role in the morphological quantification of biomedical entities such as tissues and cells, enabling precise identification and delineation of different structures. Current methods often address the challenges of touching, overlapping or crossing instances through individual modeling, while neglecting the intrinsic interrelation between these conditions. In this work, we propose a Gradient Anomaly-aware Biomedical Instance Segmentation approach (GAInS), which leverages instance gradient information to perceive local gradient anomaly regions, thus modeling the spatial relationship between instances and refining local region segmentation. Specifically, GAInS is firstly built on a Gradient Anomaly Mapping Module (GAMM), which encodes the radial fields of instances through window sliding to obtain instance gradient anomaly maps. To efficiently refine boundaries and regions with gradient anomaly attention, we propose an Adaptive Local Refinement Module (ALRM) with a gradient anomaly-aware loss function. Extensive comparisons and ablation experiments in three biomedical scenarios demonstrate that our proposed GAInS outperforms other state-of-the-art (SOTA) instance segmentation methods. The code is available at https://github.com/DeepGAInS/GAInS. Runsheng Liu, Hao Jiang 0028, Yanning Zhou 0001, Huangjing Lin, Liansheng Wang 0002, Hao Chen 0011 |
BIBM | 6 |
| 2024 | VoCo: A Simple-Yet-Effective Volume Contrastive Learning Framework for 3D Medical Image AnalysisabstractSelf-Supervised Learning (SSL) has demonstrated promising results in 3D medical image analysis. However, the lack of high-level semantics in pre-training still heavily hinders the performance of downstream tasks. We ob-serve that 3D medical images contain relatively consistent contextual position information, i.e., consistent geometric relations between different organs, which leads to a potential way for us to learn consistent semantic representations in pre-training. In this paper, we propose a simple-yet-effective Volume Contrast (VoCo) framework to leverage the contextual position priors for pre-training. Specif-ically, we first generate a group of base crops from different regions while enforcing feature discrepancy among them, where we employ them as class assignments of dif-ferent regions. Then, we randomly crop sub-volumes and predict them belonging to which class (located at which re-gion) by contrasting their similarity to different base crops, which can be seen as predicting contextual positions of different sub-volumes. Through this pretext task, VoCo implic-itly encodes the contextual position priors into model rep-resentations without the guidance of annotations, enabling us to effectively improve the performance of downstream tasks that require high-level semantics. Extensive exper-imental results on six downstream tasks demonstrate the superior effectiveness of VoCo. Code will be available at httpsu/github.com/luffytls/vo'Co. Linshan Wu, Jiaxin Zhuang, Hao Chen 0011 |
CVPR | 3 |
| 2024 | Explain via Any Concept: Concept Bottleneck Model with Open Vocabulary Concepts
Andong Tan, Fengtao Zhou, Hao Chen 0011 |
ECCV (86) | 3 |
| 2024 | DynamiCrafter: Animating Open-Domain Images with Video Diffusion Priors
Jinbo Xing, Menghan Xia, Yong Zhang 0034, Hao Chen 0011, Wangbo Yu, Hanyuan Liu, Gongye Liu, Xintao Wang 0002, Ying Shan, Tien-Tsin Wong |
ECCV (46) | 4 |
| 2024 | Noise Calibration: Plug-and-Play Content-Preserving Video Enhancement Using Pre-trained Video Diffusion Models
Qinyu Yang, Hao Chen 0011, Yong Zhang 0034, Menghan Xia, Xiaodong Cun, Zhixun Su, Ying Shan |
ECCV (36) | 2 |
| 2024 | Post-hoc Part-Prototype NetworksabstractPost-hoc explainability methods such as Grad-CAM are popular because they do not influence the performance of a trained model. However, they mainly reveal ”where” a model looks at for a given input, fail to explain ”what” the model looks for (e.g., what is important to classify a bird image to a Scott Oriole?). Existing part-prototype networks leverage part-prototypes (e.g., characteristic Scott Oriole’s wing and head) to answer both ”where" and ”what", but often under-perform their black box counterparts in the accuracy. Therefore, a natural question is: can one construct a network that answers both ”where” and ”what" in a post-hoc manner to guarantee the model’s performance? To this end, we propose the first post-hoc part-prototype network via decomposing the classification head of a trained model into a set of interpretable part-prototypes. Concretely, we propose an unsupervised prototype discovery and refining strategy to obtain prototypes that can precisely reconstruct the classification head, yet being interpretable. Besides guaranteeing the performance, we show that our network offers more faithful explanations qualitatively and yields even better part-prototypes quantitatively than prior part-prototype networks. Andong Tan, Fengtao Zhou, Hao Chen 0011 |
ICML | 3 |
| 2024 | Conditional Diffusion Model for Versatile Temporal Inpainting in 4D Cerebral CT Perfusion Imaging
Juyoung Bae, Elisabeth Tong, Hao Chen 0011 |
MICCAI (2) | 3 |
| 2024 | XCoOp: Explainable Prompt Learning for Computer-Aided Diagnosis via Concept-Guided Context Optimization
Yequan Bie, Luyang Luo, Zhixuan Chen, Hao Chen 0011 |
MICCAI (12) | 4 |
| 2024 | Rethinking Autoencoders for Medical Anomaly Detection from A Theoretical Perspective
Yu Cai 0005, Hao Chen 0011, Kwang-Ting Cheng |
MICCAI (11) | 2 |
| 2024 | BPaCo: Balanced Parametric Contrastive Learning for Long-Tailed Medical Image Classification
Zhiyuan Cai, Tianyunxi Wei, Li Lin 0006, Hao Chen 0011, Xiaoying Tang 0001 |
MICCAI (1) | 4 |
| 2024 | Enable the Right to be Forgotten with Federated Client Unlearning in Medical Imaging
Zhipeng Deng, Luyang Luo, Hao Chen 0011 |
MICCAI (10) | 3 |
| 2024 | Aligning Medical Images with General Knowledge from Large Language Models
Yi Lin 0009, Kwang-Ting Cheng, Hao Chen 0011 |
MICCAI (10) | 5 |
| 2024 | Revisiting Deep Ensemble Uncertainty for Enhanced Medical Anomaly Detection
Yi Lin 0009, Kwang-Ting Cheng, Hao Chen 0011 |
MICCAI (6) | 4 |
| 2024 | HistGen: Histopathology Report Generation via Local-Global Feature Encoding and Cross-Modal Context Interaction
Zhengrui Guo, Jiabo Ma, Yingxue Xu, Yihui Wang 0002, Liansheng Wang 0002, Hao Chen 0011 |
MICCAI (4) | 6 |
| 2024 | Boosting FFPE-to-HE Virtual Staining with Cell Semantics from Pretrained Segmentation Model
Yihuang Hu, Qiong Peng, Zhicheng Du, Huisi Wu, Jingxin Liu 0005, Hao Chen 0011, Liansheng Wang 0002 |
MICCAI (3) | 7 |
| 2024 | Iterative Online Image Synthesis via Diffusion Model for Imbalanced Classification
Shuhan Li, Yi Lin 0009, Hao Chen 0011, Kwang-Ting Cheng |
MICCAI (5) | 3 |
| 2024 | Progressive Knowledge Distillation for Automatic Perfusion Parameter Maps Generation from Low Temporal Resolution CT Perfusion Images
Moo Hyun Son, Juyoung Bae, Elisabeth Tong, Hao Chen 0011 |
MICCAI (10) | 4 |
| 2024 | Design as Desired: Utilizing Visual Question Answering for Multimodal Pre-training
Tongkun Su, Jun Li 0111, Hai Jin 0001, Hao Chen 0011, Qiong Wang 0001, Faqin Lv, Baoliang Zhao, Ying Hu 0001 |
MICCAI (4) | 5 |
| 2024 | MoME: Mixture of Multimodal Experts for Cancer Survival Prediction
Conghao Xiong, Hao Chen 0011, Hao Zheng 0008, Dong Wei 0004, Yefeng Zheng 0001, Joseph J. Y. Sung, Irwin King |
MICCAI (4) | 2 |
| 2024 | TAKT: Target-Aware Knowledge Transfer for Whole Slide Image Classification
Conghao Xiong, Yi Lin 0009, Hao Chen 0011, Hao Zheng 0008, Dong Wei 0004, Yefeng Zheng 0001, Joseph J. Y. Sung, Irwin King |
MICCAI (4) | 3 |
| 2024 | Surgformer: Surgical Transformer with Hierarchical Temporal Attention for Surgical Phase Recognition
Shu Yang 0004, Luyang Luo, Qiong Wang 0001, Hao Chen 0011 |
MICCAI (6) | 4 |
| 2024 | MambaMIL: Enhancing Long Sequence Modeling with Sequence Reordering in Computational Pathology
Shu Yang 0004, Yihui Wang 0002, Hao Chen 0011 |
MICCAI (4) | 3 |
| 2024 | Deep Model Reference: Simple Yet Effective Confidence Estimation for Image Classification
Yuanhang Zheng, Yiqiao Qiu, Haoxuan Che, Hao Chen 0011, Wei-Shi Zheng 0001 |
MICCAI (10) | 4 |
| 2024 | Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?abstractHow can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5,195 training CT scans from 76 hospitals around the world and 5,903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain. Pedro R. A. S. Bassi, Yucheng Tang, Fabian Isensee, Zifu Wang, Jieneng Chen, Yu-Cheng Chou, Yannick Kirchhoff, Maximilian Rokuss, Ziyan Huang, Jin Ye 0002, Junjun He, Tassilo Wald, Constantin Ulrich, Michael Baumgartner 0001, Saikat Roy, Klaus H. Maier-Hein, Paul F. Jaeger, Yiwen Ye, Yutong Xie 0001, Ziyang Chen 0003, Yong Xia 0001, Zhaohu Xing, Lei Zhu 0003, Yousef Sadegheih, Afshin Bozorgpour, Pratibha Kumari 0001, Reza Azad, Dorit Merhof, Yuxin Du 0001, Fan Bai 0008, Tiejun Huang 0001, Bo Zhao 0015, Xiaomeng Li 0001, Hanxue Gu, Haoyu Dong 0003, Maciej A. Mazurowski, Saumya Gupta, Linshan Wu, Jiaxin Zhuang, Hao Chen 0011, Holger Roth, Daguang Xu, Matthew B. Blaschko, Sergio Decherchi, Andrea Cavalli, Alan L. Yuille, Zongwei Zhou |
NeurIPS | 46 |
| 2024 | Dual-stream multi-dependency graph neural network enables precise cancer survival analysisabstractHistopathology image-based survival prediction aims to provide a precise assessment of cancer prognosis and can inform personalized treatment decision-making in order to improve patient outcomes. However, existing methods cannot automatically model the complex correlations between numerous morphologically diverse patches in each whole slide image (WSI), thereby preventing them from achieving a more profound understanding and inference of the patient status. To address this, here we propose a novel deep learning framework, termed dual-stream multi-dependency graph neural network (DM-GNN), to enable precise cancer patient survival analysis. Specifically, DM-GNN is structured with the feature updating and global analysis branches to better model each WSI as two graphs based on morphological affinity and global co-activating dependencies. As these two dependencies depict each WSI from distinct but complementary perspectives, the two designed branches of DM-GNN can jointly achieve the multi-view modeling of complex correlations between the patches. Moreover, DM-GNN is also capable of boosting the utilization of dependency information during graph construction by introducing the affinity-guided attention recalibration module as the readout function. This novel module offers increased robustness against feature perturbation, thereby ensuring more reliable and stable predictions. Extensive benchmarking experiments on five TCGA datasets demonstrate that DM-GNN outperforms other state-of-the-art methods and offers interpretable prediction insights based on the morphological depiction of high-attention patches. Overall, DM-GNN represents a powerful and auxiliary tool for personalized cancer prognosis from histopathology images and has great potential to assist clinicians in making personalized treatment decisions and improving patient outcomes. Zhikang Wang, Jiani Ma, Chris Bain, Seiya Imoto, Pietro Liò, Hongmin Cai, Hao Chen 0011, Jiangning Song |
Medical Image Anal. | 8 |
| 2024 | LENAS: Learning-Based Neural Architecture Search and Ensemble for 3-D Radiotherapy Dose PredictionabstractRadiation therapy treatment planning requires balancing the delivery of the target dose while sparing normal tissues, making it a complex process. To streamline the planning process and enhance its quality, there is a growing demand for knowledge-based planning (KBP). Ensemble learning has shown impressive power in various deep learning tasks, and it has great potential to improve the performance of KBP. However, the effectiveness of ensemble learning heavily depends on the diversity and individual accuracy of the base learners. Moreover, the complexity of model ensembles is a major concern, as it requires maintaining multiple models during inference, leading to increased computational cost and storage overhead. In this study, we propose a novel learning-based ensemble approach named LENAS, which integrates neural architecture search with knowledge distillation for 3-D radiotherapy dose prediction. Our approach starts by exhaustively searching each block from an enormous architecture space to identify multiple architectures that exhibit promising performance and significant diversity. To mitigate the complexity introduced by the model ensemble, we adopt the teacher-student paradigm, leveraging the diverse outputs from multiple learned networks as supervisory signals to guide the training of the student network. Furthermore, to preserve high-level semantic information, we design a hybrid loss to optimize the student network, enabling it to recover the knowledge embedded within the teacher networks. The proposed method has been evaluated on two public datasets: 1) OpenKBP and 2) AIMIS. Extensive experimental results demonstrate the effectiveness of our method and its superior performance to the state-of-the-art methods. Code: github.com/hust-linyi/LENAS. Yi Lin 0009, Hao Chen 0011, Xin Yang 0008, Kai Ma 0002, Yefeng Zheng 0001, Kwang-Ting Cheng |
IEEE Trans. Cybern. | 3 |
| 2024 | Rethinking Self-Training for Semi-Supervised Landmark Detection: A Selection-Free ApproachabstractSelf-training is a simple yet effective method for semi-supervised learning, during which pseudo-label selection plays an important role for handling confirmation bias. Despite its popularity, applying self-training to landmark detection faces three problems: 1) The selected confident pseudo-labels often contain data bias, which may hurt model performance; 2) It is not easy to decide a proper threshold for sample selection as the localization task can be sensitive to noisy pseudo-labels; 3) coordinate regression does not output confidence, making selection-based self-training infeasible. To address the above issues, we propose Self-Training for Landmark Detection (STLD), a method that does not require explicit pseudo-label selection. Instead, STLD constructs a task curriculum to deal with confirmation bias, which progressively transitions from more confident to less confident tasks over the rounds of self-training. Pseudo pretraining and shrink regression are two essential components for such a curriculum, where the former is the first task of the curriculum for providing a better model initialization and the latter is further added in the later rounds to directly leverage the pseudo-labels in a coarse-to-fine manner. Experiments on three facial and one medical landmark detection benchmark show that STLD outperforms the existing methods consistently in both semi- and omni-supervised settings. The code is available at https://github.com/jhb86253817/STLD. Haibo Jin, Haoxuan Che, Hao Chen 0011 |
IEEE Trans. Image Process. | 3 |
| 2024 | Guest Editorial: Trustworthy Machine Learning for Health InformaticsabstractMachine learning (ML), the stem of today's artificial intelligence, has shown significant growth in the field of biomedical and health informatics. On the one hand, ML techniques are becoming more complex in order to deal with real-world data. On the other hand, ML is also more and more accessible to broader users. For example, automated machine learning products are enabling users to build their own ML models without writing code [1]. Luyang Luo, Daguang Xu, Harry Qin, Yueming Jin, Hao Chen 0011 |
IEEE J. Biomed. Health Informatics | 5 |
| 2024 | Adaptive Fusion of Deep Learning With Statistical Anatomical Knowledge for Robust Patella Segmentation From CT ImagesabstractKneeosteoarthritis (KOA), as a leading joint disease, can be decided by examining the shapes of patella to spot potential abnormal variations. To assist doctors in the diagnosis of KOA, a robust automatic patella segmentation method is highly demanded in clinical practice. Deep learning methods, especially convolutional neural networks (CNNs) have been widely applied to medical image segmentation in recent years. Nevertheless, poor image quality and limited data still impose challenges to segmentation via CNNs. On the other hand, statistical shape models (SSMs) can generate shape priors which give anatomically reliable segmentation to varying instances. Thus, in this work, we propose an adaptive fusion framework, explicitly combining deep neural networks and anatomical knowledge from SSM for robust patella segmentation. Our adaptive fusion framework will accordingly adjust the weight of segmentation candidates in fusion based on their segmentation performance. We also propose a voxel-wise refinement strategy to make the segmentation of CNNs more anatomically correct. Extensive experiments and thorough assessment have been conducted on various mainstream CNN backbones for patella segmentation in low-data regimes, which demonstrate that our framework can be flexibly attached to a CNN model, significantly improving its performance when labeled training data are limited and input image data are of poor quality. Tianshu Jiang, Yi Lin 0009, Lok-Chun Chan, Ping-Keung Chan, Chun-Yi Wen, Hao Chen 0011 |
IEEE J. Biomed. Health Informatics | 7 |
| 2024 | Deep Omni-Supervised Learning for Rib Fracture Detection From Chest Radiology ImagesabstractDeep learning (DL)-based rib fracture detection has shown promise of playing an important role in preventing mortality and improving patient outcome. Normally, developing DL-based object detection models requires a huge amount of bounding box annotation. However, annotating medical data is time-consuming and expertise-demanding, making obtaining a large amount of fine-grained annotations extremely infeasible. This poses a pressing need for developing label-efficient detection models to alleviate radiologists' labeling burden. To tackle this challenge, the literature on object detection has witnessed an increase of weakly-supervised and semi-supervised approaches, yet still lacks a unified framework that leverages various forms of fully-labeled, weakly-labeled, and unlabeled data. In this paper, we present a novel omni-supervised object detection network, ORF-Netv2, to leverage as much available supervision as possible. Specifically, a multi-branch omni-supervised detection head is introduced with each branch trained with a specific type of supervision. A co-training-based dynamic label assignment strategy is then proposed to enable flexible and robust learning from the weakly-labeled and unlabeled data. Extensive evaluation was conducted for the proposed framework with three rib fracture datasets on both chest CT and X-ray. By leveraging all forms of supervision, ORF-Netv2 achieves mAPs of 34.7, 44.7, and 19.4 on the three datasets, respectively, surpassing the baseline detector which uses only box annotations by mAP gains of 3.8, 4.8, and 5.0, respectively. Furthermore, ORF-Netv2 consistently outperforms other competitive label-efficient methods over various scenarios, showing a promising framework for label-efficient fracture detection. The code is available at: https://github.com/zhizhongchai/ORF-Net. Zhizhong Chai, Luyang Luo, Huangjing Lin, Pheng-Ann Heng, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | BoNuS: Boundary Mining for Nuclei Segmentation With Partial Point LabelsabstractNuclei segmentation is a fundamental prerequisite in the digital pathology workflow. The development of automated methods for nuclei segmentation enables quantitative analysis of the wide existence and large variances in nuclei morphometry in histopathology images. However, manual annotation of tens of thousands of nuclei is tedious and time-consuming, which requires significant amount of human effort and domain-specific expertise. To alleviate this problem, in this paper, we propose a weakly-supervised nuclei segmentation method that only requires partial point labels of nuclei. Specifically, we propose a novel boundary mining framework for nuclei segmentation, named BoNuS, which simultaneously learns nuclei interior and boundary information from the point labels. To achieve this goal, we propose a novel boundary mining loss, which guides the model to learn the boundary information by exploring the pairwise pixel affinity in a multiple-instance learning manner. Then, we consider a more challenging problem, i.e., partial point label, where we propose a nuclei detection module with curriculum learning to detect the missing nuclei with prior morphological knowledge. The proposed method is validated on three public datasets, MoNuSeg, CPM, and CoNIC datasets. Experimental results demonstrate the superior performance of our method to the state-of-the-art weakly-supervised nuclei segmentation methods. Code: https://github.com/hust-linyi/bonus. Yi Lin 0009, Kwang-Ting Cheng, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 5 |
| 2024 | Efficient Supervised Pretraining of Swin-Transformer for Virtual Staining of Microscopy ImagesabstractFluorescence staining is an important technique in life science for labeling cellular constituents. However, it also suffers from being time-consuming, having difficulty in simultaneous labeling, etc. Thus, virtual staining, which does not rely on chemical labeling, has been introduced. Recently, deep learning models such as transformers have been applied to virtual staining tasks. However, their performance relies on large-scale pretraining, hindering their development in the field. To reduce the reliance on large amounts of computation and data, we construct a Swin-transformer model and propose an efficient supervised pretraining method based on the masked autoencoder (MAE). Specifically, we adopt downsampling and grid sampling to mask 75% of pixels and reduce the number of tokens. The pretraining time of our method is only 1/16 compared with the original MAE. We also design a supervised proxy task to predict stained images with multiple styles instead of masked pixels. Additionally, most virtual staining approaches are based on private datasets and evaluated by different metrics, making a fair comparison difficult. Therefore, we develop a standard benchmark based on three public datasets and build a baseline for the convenience of future researchers. We conduct extensive experiments on three benchmark datasets, and the experimental results show the proposed method achieves the best performance both quantitatively and qualitatively. In addition, ablation studies are conducted, and experimental results illustrate the effectiveness of the proposed pretraining method. The benchmark and code are available at https://github.com/birkhoffkiki/CAS-Transformer. Jiabo Ma, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Slim UNETR: Scale Hybrid Transformers to Efficient 3D Medical Image Segmentation Under Limited Computational ResourcesabstractHybrid transformer-based segmentation approaches have shown great promise in medical image analysis. However, they typically require considerable computational power and resources during both training and inference stages, posing a challenge for resource-limited medical applications common in the field. To address this issue, we present an innovative framework called Slim UNETR, designed to achieve a balance between accuracy and efficiency by leveraging the advantages of both convolutional neural networks and transformers. Our method features the Slim UNETR Block as a core component, which effectively enables information exchange through self-attention mechanism decomposition and cost-effective representation aggregation. Additionally, we utilize the throughput metric as an efficiency indicator to provide feedback on model resource consumption. Our experiments demonstrate that Slim UNETR outperforms state-of-the-art models in terms of accuracy, model size, and efficiency when deployed on resource-constrained devices. Remarkably, Slim UNETR achieves 92.44% dice accuracy on BraTS2021 while being 34.6x smaller and 13.4x faster during inference compared to Swin UNETR. Code: https://github.com/aigzhusmart/Slim-UNETR. Teng Huang 0001, Hao Chen 0011, Qiong Wang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Rethinking Multiple Instance Learning for Whole Slide Image Classification: A Bag-Level Classifier is a Good Instance-Level TeacherabstractMultiple Instance Learning (MIL) has demonstrated promise in Whole Slide Image (WSI) classification. However, a major challenge persists due to the high computational cost associated with processing these gigapixel images. Existing methods generally adopt a two-stage approach, comprising a non-learnable feature embedding stage and a classifier training stage. Though it can greatly reduce memory consumption by using a fixed feature embedder pre-trained on other domains, such a scheme also results in a disparity between the two stages, leading to suboptimal classification accuracy. To address this issue, we propose that a bag-level classifier can be a good instance-level teacher. Based on this idea, we design Iteratively Coupled Multiple Instance Learning (ICMIL) to couple the embedder and the bag classifier at a low cost. ICMIL initially fixes the patch embedder to train the bag classifier, followed by fixing the bag classifier to fine-tune the patch embedder. The refined embedder can then generate better representations in return, leading to a more accurate classifier for the next iteration. To realize more flexible and more effective embedder fine-tuning, we also introduce a teacher-student framework to efficiently distill the category knowledge in the bag classifier to help the instance-level embedder fine-tuning. Intensive experiments were conducted on four distinct datasets to validate the effectiveness of ICMIL. The experimental results consistently demonstrated that our method significantly improves the performance of existing MIL backbones, achieving state-of-the-art results. The code and the organized datasets can be accessed by: https://github.com/Dootmaan/ICMIL/tree/confidence-based. Hongyi Wang 0002, Luyang Luo, Fang Wang 0030, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin, Hao Chen 0011 |
IEEE Trans. Medical Imaging | 8 |
| 2023 | Image Quality-aware Diagnosis via Meta-knowledge Co-embeddingabstractMedical images usually suffer from image degradation in clinical practice, leading to decreased performance of deep learning-based models. To resolve this problem, most previous works have focused on filtering out degradation-causing low-quality images while ignoring their potential value for models. Through effectively learning and leveraging the knowledge of degradations, models can better resist their adverse effects and avoid misdiagnosis. In this paper, we raise the problem of image quality-aware diagnosis, which aims to take advantage of low-quality images and image quality labels to achieve a more accurate and robust diagnosis. However, the diversity of degradations and superficially unrelated targets between image quality assessment and disease diagnosis makes it still quite challenging to effectively leverage quality labels to assist diagnosis. Thus, to tackle these issues, we propose a novel meta-knowledge co-embedding network, consisting of two subnets: Task Net and Meta Learner. Task Net constructs an explicit quality information utilization mechanism to enhance diagnosis via knowledge co-embedding features, while Meta Learner ensures the effectiveness and constrains the semantics of these features via meta-learning and joint-encoding masking. Superior performance on five datasets with four widely-used medical imaging modalities demonstrates the effectiveness and generalizability of our method. Haoxuan Che, Siyu Chen 0045, Hao Chen 0011 |
CVPR | 3 |
| 2023 | Sparsely Annotated Semantic Segmentation with Adaptive Gaussian MixturesabstractSparsely annotated semantic segmentation (SASS) aims to learn a segmentation model by images with sparse labels (i.e., points or scribbles). Existing methods mainly focus on introducing low-level affinity or generating pseudo labels to strengthen supervision, while largely ignoring the inherent relation between labeled and unlabeled pixels. In this paper, we observe that pixels that are close to each other in the feature space are more likely to share the same class. Inspired by this, we propose a novel SASS framework, which is equipped with an Adaptive Gaussian Mixture Model (AGMM). Our AGMM can effectively endow reliable supervision for unlabeled pixels based on the distributions of labeled and unlabeled pixels. Specifically, we first build Gaussian mixtures using labeled pixels and their relatively similar unlabeled pixels, where the labeled pixels act as centroids, for modeling the feature distribution of each class. Then, we leverage the reliable information from labeled pixels and adaptively generated GMM predictions to supervise the training of unlabeled pixels, achieving online, dynamic, and robust selfsupervision. In addition, by capturing category-wise Gaussian mixtures, AGMM encourages the model to learn discriminative class decision boundaries in an end-to-end contrastive learning manner. Experimental results conducted on the PASCAL VOC 2012 and Cityscapes datasets demonstrate that our AGMM can establish new state-of-the-art SASS performance. Code is available at https://github.com/Luffy03/AGMM-SASS Linshan Wu, Zhun Zhong, Leyuan Fang, Xingxin He, Jiayi Ma 0001, Hao Chen 0011 |
CVPR | 7 |
| 2023 | Multimodal Optimal Transport-based Co-Attention Transformer with Global Structure Consistency for Survival PredictionabstractSurvival prediction is a complicated ordinal regression task that aims to predict the ranking risk of death, which generally benefits from the integration of histology and genomic data. Despite the progress in joint learning from pathology and genomics, existing methods still suffer from challenging issues: 1) Due to the large size of pathological images, it is difficult to effectively represent the gigapixel whole slide images (WSIs). 2) Interactions within tumor microenvironment (TME) in histology are essential for survival analysis. Although current approaches attempt to model these interactions via co-attention between histology and genomic data, they focus on only dense local similarity across modalities, which fails to capture global consistency between potential structures, i.e. TME-related interactions of histology and co-expression of genomic data. To address these challenges, we propose a Multimodal Optimal Transport-based Co-Attention Transformer framework with global structure consistency, in which optimal transport (OT) is applied to match patches of a WSI and genes embeddings for selecting informative patches to represent the gigapixel WSI. More importantly, OT-based co-attention provides a global awareness to effectively capture structural interactions within TME for survival prediction. To overcome high computational complexity of OT, we propose a robust and efficient implementation over micro-batch of WSI patches by approximating the original OT with unbalanced mini-batch OT. Extensive experiments show the superiority of our method on five benchmark datasets compared to the state-of-the-art methods. The code is released1. Yingxue Xu, Hao Chen 0011 |
ICCV | 2 |
| 2023 | Cross-Modal Translation and Alignment for Survival AnalysisabstractWith the rapid advances in high-throughput sequencing technologies, the focus of survival analysis has shifted from examining clinical indicators to incorporating genomic profiles with pathological images. However, existing methods either directly adopt a straightforward fusion of pathological features and genomic profiles for survival prediction, or take genomic profiles as guidance to integrate the features of pathological images. The former would overlook intrinsic cross-modal correlations. The latter would discard pathological information irrelevant to gene expression. To address these issues, we present a Cross-Modal Translation and Alignment (CMTA) framework to explore the intrinsic cross-modal correlations and transfer potential complementary information. Specifically, we construct two parallel encoder-decoder structures for multi-modal data to integrate intra-modal information and generate cross-modal representation. Taking the generated cross-modal representation to enhance and recalibrate intra-modal representation can significantly improve its discrimination for comprehensive survival analysis. To explore the intrinsic cross-modal correlations, we further design a cross-modal attention module as the information bridge between different modalities to perform cross-modal interactions and transfer complementary information. Our extensive experiments on five public TCGA datasets demonstrate that our proposed framework outperforms the state-of-the-art methods. The source code has been released†. Fengtao Zhou, Hao Chen 0011 |
ICCV | 2 |
| 2023 | Diagnose Like a Pathologist: Transformer-Enabled Hierarchical Attention-Guided Multiple Instance Learning for Whole Slide Image ClassificationabstractMultiple Instance Learning (MIL) and transformers are increasingly popular in histopathology Whole Slide Image (WSI) classification. However, unlike human pathologists who selectively observe specific regions of histopathology tissues under different magnifications, most methods do not incorporate multiple resolutions of the WSIs, hierarchically and attentively, thereby leading to a loss of focus on the WSIs and information from other resolutions. To resolve this issue, we propose a Hierarchical Attention-Guided Multiple Instance Learning framework to fully exploit the WSIs. This framework can dynamically and attentively discover the discriminative regions across multiple resolutions of the WSIs. Within this framework, an Integrated Attention Transformer is proposed to further enhance the performance of the transformer and obtain a more holistic WSI (bag) representation. This transformer consists of multiple Integrated Attention Modules, which is the combination of a transformer layer and an aggregation module that produces a bag representation based on every instance representation in that bag. The experimental results show that our method achieved state-of-the-art performances on multiple datasets, including Camelyon16, TCGA-RCC, TCGA-NSCLC, and an in-house IMGC dataset. The code is available at https://github.com/BearCleverProud/HAG-MIL. Conghao Xiong, Hao Chen 0011, Joseph J. Y. Sung, Irwin King |
IJCAI | 2 |
| 2023 | Towards Generalizable Diabetic Retinopathy Grading in Unseen Domains
Haoxuan Che, Yuhan Cheng, Haibo Jin, Hao Chen 0011 |
MICCAI (5) | 4 |
| 2023 | DARC: Distribution-Aware Re-Coloring Model for Generalizable Nucleus Segmentation
Shengcong Chen, Changxing Ding, Dacheng Tao, Hao Chen 0011 |
MICCAI (6) | 4 |
| 2023 | Scale Federated Learning for Label Set Mismatch in Medical Image Classification
Zhipeng Deng, Luyang Luo, Hao Chen 0011 |
MICCAI (3) | 3 |
| 2023 | Unsupervised Domain Adaptation for Anatomical Landmark Detection
Haibo Jin, Haoxuan Che, Hao Chen 0011 |
MICCAI (1) | 3 |
| 2023 | Few Shot Medical Image Segmentation with Cross Attention Transformer
Yi Lin 0009, Kwang-Ting Cheng, Hao Chen 0011 |
MICCAI (2) | 4 |
| 2023 | Iteratively Coupled Multiple Instance Learning from Instance to Bag Classifier for Whole Slide Image Classification
Hongyi Wang 0002, Luyang Luo, Fang Wang 0030, Ruofeng Tong 0001, Yen-Wei Chen 0001, Hongjie Hu, Lanfen Lin, Hao Chen 0011 |
MICCAI (6) | 8 |
| 2023 | The Liver Tumor Segmentation Benchmark (LiTS)abstractIn this work, we report the set-up and results of the Liver Tumor Segmentation Benchmark (LiTS), which was organized in conjunction with the IEEE International Symposium on Biomedical Imaging (ISBI) 2017 and the International Conferences on Medical Image Computing and Computer-Assisted Intervention (MICCAI) 2017 and 2018. The image dataset is diverse and contains primary and secondary tumors with varied sizes and appearances with various lesion-to-background levels (hyper-/hypo-dense), created in collaboration with seven hospitals and research institutions. Seventy-five submitted liver and liver tumor segmentation algorithms were trained on a set of 131 computed tomography (CT) volumes and were tested on 70 unseen test images acquired from different patients. We found that not a single algorithm performed best for both liver and liver tumors in the three events. The best liver segmentation algorithm achieved a Dice score of 0.963, whereas, for tumor segmentation, the best algorithms achieved Dices scores of 0.674 (ISBI 2017), 0.702 (MICCAI 2017), and 0.739 (MICCAI 2018). Retrospectively, we performed additional analysis on liver tumor detection and revealed that not all top-performing segmentation algorithms worked well for tumor detection. The best liver tumor detection method achieved a lesion-wise recall of 0.458 (ISBI 2017), 0.515 (MICCAI 2017), and 0.554 (MICCAI 2018), indicating the need for further research. LiTS remains an active benchmark and resource for research, e.g., contributing the liver-related segmentation tasks in http://medicaldecathlon.com/. In addition, both data and online evaluation are accessible via https://competitions.codalab.org/competitions/17094. Patrick Bilic, Patrick Ferdinand Christ, Hongwei Li 0004, Eugene Vorontsov, Avi Ben-Cohen, Georgios Kaissis, Adi Szeskin, Colin Jacobs, Gabriel Efrain Humpire Mamani, Gabriel Chartrand, Fabian Lohöfer, Julian Walter Holch, Wieland H. Sommer, Felix Hofmann, Alexandre Hostettler, Naama Lev-Cohain, Michal Drozdzal, Michal Amitai, Refael Vivanti, Jacob Sosna, Ivan Ezhov, Anjany Sekuboyina, Fernando Navarro, Florian Kofler, Johannes C. Paetzold, Suprosanna Shit, Xiaobin Hu, Jana Lipková, Markus Rempfler, Marie Piraud, Jan Kirschke, Benedikt Wiestler, Christian Hülsemeyer, Marcel Beetz, Florian Ettlinger, Michela Antonelli, Woong Bae, Miriam Bellver, Lei Bi 0001, Hao Chen 0011, Grzegorz Chlebus, Erik Dam, Qi Dou 0001, Chi-Wing Fu, Bogdan Georgescu, Xavier Giró-i-Nieto, Felix Grün, Xu Han 0009, Pheng-Ann Heng, Jürgen Hesser, Jan Hendrik Moltz, Christian Igel, Fabian Isensee, Paul F. Jaeger, Fucang Jia, Krishna Chaitanya Kaluva, Mahendra Khened, Ildoo Kim, Jae-Hun Kim, Sungwoong Kim, Simon Kohl, Tomasz K. Konopczynski, Avinash Kori, Ganapathy Krishnamurthi, Xiaomeng Li 0001, John S. Lowengrub, Jun Ma 0016, Klaus H. Maier-Hein, Kevis-Kokitsi Maninis, Hans Meine, Dorit Merhof, Akshay Pai, Mathias Perslev, Jens Petersen, Jordi Pont-Tuset, Xiaojuan Qi 0001, Oliver Rippel, Karsten Roth, Ignacio Sarasua, Andrea Schenk, Zengming Shen, Jordi Torres, Christian Wachinger, Chunliang Wang, Leon Weninger, Daguang Xu, Xiaoping Yang 0001, Simon C. H. Yu, Yading Yuan, Miao Yue, Liping Zhang 0009, Manuel Jorge Cardoso, Spyridon Bakas, Rickmer Braren, Volker Heinemann, Christopher Joseph Pal, An Tang, Samuel Kadoury, Luc Soler, Bram van Ginneken, Hayit Greenspan, Leo Joskowicz, Bjoern Menze |
Medical Image Anal. | 41 |
| 2023 | Dual-distribution discrepancy with self-supervised refinement for anomaly detection in medical images
Yu Cai 0005, Hao Chen 0011, Xin Yang 0008, Yu Zhou 0016, Kwang-Ting Cheng |
Medical Image Anal. | 2 |
| 2023 | Deep learning for computational cytology: A survey
Hao Jiang 0028, Yanning Zhou 0001, Yi Lin 0009, Ronald C. K. Chan, Jiang Liu 0001, Hao Chen 0011 |
Medical Image Anal. | 6 |
| 2023 | Nuclei segmentation with point annotations from pathology images via self-supervised learning and co-training
Yi Lin 0009, Zhiyong Qu, Hao Chen 0011, Zhongke Gao, Yuexiang Li, Kai Ma 0002, Yefeng Zheng 0001, Kwang-Ting Cheng |
Medical Image Anal. | 3 |
| 2023 | Deep semi-supervised multiple instance learning with self-correction for DME classification from OCT images
Xi Wang 0013, Fangyao Tang, Hao Chen 0011, Carol Y. Cheung, Pheng-Ann Heng |
Medical Image Anal. | 3 |
| 2023 | Multi-semantic hypergraph neural network for effective few-shot learning
Hao Chen 0011, Fuyuan Hu, Fan Lyu, Liuqing Zhao, Kaizhu Huang, Wei Feng 0005, Zhenping Xia |
Pattern Recognit. | 1 |
| 2023 | CKD-TransBTS: Clinical Knowledge-Driven Hybrid Transformer With Modality-Correlated Cross-Attention for Brain Tumor SegmentationabstractBrain tumor segmentation (BTS) in magnetic resonance image (MRI) is crucial for brain tumor diagnosis, cancer management and research purposes. With the great success of the ten-year BraTS challenges as well as the advances of CNN and Transformer algorithms, a lot of outstanding BTS models have been proposed to tackle the difficulties of BTS in different technical aspects. However, existing studies hardly consider how to fuse the multi-modality images in a reasonable manner. In this paper, we leverage the clinical knowledge of how radiologists diagnose brain tumors from multiple MRI modalities and propose a clinical knowledge-driven brain tumor segmentation model, called CKD-TransBTS. Instead of directly concatenating all the modalities, we re-organize the input modalities by separating them into two groups according to the imaging principle of MRI. A dual-branch hybrid encoder with the proposed modality-correlated cross-attention block (MCCA) is designed to extract the multi-modality image features. The proposed model inherits the strengths from both Transformer and CNN with the local feature representation ability for precise lesion boundaries and long-range feature extraction for 3D volumetric images. To bridge the gap between Transformer and CNN features, we propose a Trans&CNN Feature Calibration block (TCFC) in the decoder. We compare the proposed model with six CNN-based models and six transformer-based models on the BraTS 2021 challenge dataset. Extensive experiments demonstrate that the proposed model achieves state-of-the-art brain tumor segmentation performance compared with all the competitors. Jianwei Lin, Jiatai Lin, Cheng Lu 0001, Hao Chen 0011, Bingchao Zhao, Zhenwei Shi 0002, Bingjiang Qiu, Xipeng Pan, Zeyan Xu, Biao Huang 0008, Changhong Liang, Guoqiang Han 0002, Zaiyi Liu, Chu Han |
IEEE Trans. Medical Imaging | 4 |
| 2022 | Dual-Distribution Discrepancy for Anomaly Detection in Chest X-Rays
Yu Cai 0005, Hao Chen 0011, Xin Yang 0008, Yu Zhou 0016, Kwang-Ting Cheng |
MICCAI (3) | 2 |
| 2022 | ORF-Net: Deep Omni-Supervised Rib Fracture Detection from Chest CT Scans
Zhizhong Chai, Huangjing Lin, Luyang Luo, Pheng-Ann Heng, Hao Chen 0011 |
MICCAI (3) | 5 |
| 2022 | Learning Robust Representation for Joint Grading of Ophthalmic Diseases via Adaptive Curriculum and Feature Disentanglement
Haoxuan Che, Haibo Jin, Hao Chen 0011 |
MICCAI (3) | 3 |
| 2022 | InsMix: Towards Realistic Generative Data Augmentation for Nuclei Instance Segmentation
Yi Lin 0009, Kwang-Ting Cheng, Hao Chen 0011 |
MICCAI (2) | 4 |
| 2022 | Pseudo Bias-Balanced Learning for Debiased Chest X-Ray Classification
Luyang Luo, Dunyuan Xu, Hao Chen 0011, Tien-Tsin Wong, Pheng-Ann Heng |
MICCAI (8) | 3 |
| 2022 | Frequency-Aware Inverse-Consistent Deep Learning for OCT-Angiogram Super-Resolution
Carol Y. Cheung, Hao Chen 0011 |
MICCAI (2) | 4 |
| 2022 | Harnessing Multi-Semantic Hypergraph for Few-Shot Learning
Hao Chen 0011, Zhenping Xia, Fan Lyu, Liuqing Zhao, Kaizhu Huang, Wei Feng 0005, Fuyuan Hu |
PRCV (1) | 1 |
| 2022 | PDBL: Improving Histopathological Tissue Classification With Plug-and-Play Pyramidal Deep-Broad LearningabstractHistopathological tissue classification is a simpler way to achieve semantic segmentation for the whole slide images, which can alleviate the requirement of pixel-level dense annotations. Existing works mostly leverage the popular CNN classification backbones in computer vision to achieve histopathological tissue classification. In this paper, we propose a super lightweight plug-and-play module, named Pyramidal Deep-Broad Learning (PDBL), for any well-trained classification backbone to improve the classification performance without a re-training burden. For each patch, we construct a multi-resolution image pyramid to obtain the pyramidal contextual information. For each level in the pyramid, we extract the multi-scale deep-broad features by our proposed Deep-Broad block (DB-block). We equip PDBL in three popular classification backbones, ShuffLeNetV2, EfficientNetb0, and ResNet50 to evaluate the effectiveness and efficiency of our proposed module on two datasets (Kather Multiclass Dataset and the LC25000 Dataset). Experimental results demonstrate the proposed PDBL can steadily improve the tissue-level classification performance for any CNN backbones, especially for the lightweight models when given a small among of training samples (less than 10%). It greatly saves the computational resources and annotation efforts. The source code is available at: https://github.com/linjiatai/PDBL. Jiatai Lin, Guoqiang Han 0002, Xipeng Pan, Zaiyi Liu, Hao Chen 0011, Danyi Li, Xiping Jia, Zhenwei Shi 0002, Zhizhen Wang, Yanfen Cui, Haiming Li, Changhong Liang, Chu Han |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Dual-Consistency Semi-supervised Learning with Uncertainty Quantification for COVID-19 Lesion Segmentation from CT Images
Yanwen Li, Luyang Luo, Huangjing Lin, Hao Chen 0011, Pheng-Ann Heng |
MICCAI (2) | 4 |
| 2021 | OXnet: Deep Omni-Supervised Thoracic Disease Detection from Chest X-Rays
Luyang Luo, Hao Chen 0011, Yanning Zhou 0001, Huangjing Lin, Pheng-Ann Heng |
MICCAI (2) | 2 |
| 2021 | Deep virtual adversarial self-training with consistency regularization for semi-supervised medical image classification
Xi Wang 0013, Hao Chen 0011, Huiling Xiang, Huangjing Lin, Pheng-Ann Heng |
Medical Image Anal. | 2 |
| 2021 | Dual-path network with synergistic grouping loss and evidence driven risk stratification for whole slide cervical image analysis
Huangjing Lin, Hao Chen 0011, Xi Wang 0013, Qiong Wang 0001, Liansheng Wang 0002, Pheng-Ann Heng |
Medical Image Anal. | 2 |
| 2021 | 3-D RoI-Aware U-Net for Accurate and Efficient Colorectal Tumor SegmentationabstractSegmentation of colorectal cancerous regions from 3-D magnetic resonance (MR) images is a crucial procedure for radiotherapy. Automatic delineation from 3-D whole volumes is in urgent demand yet very challenging. Drawbacks of existing deep-learning-based methods for this task are two-fold: 1) extensive graphics processing unit (GPU) memory footprint of 3-D tensor limits the trainable volume size, shrinks effective receptive field, and therefore, degrades speed and segmentation performance and 2) in-region segmentation methods supported by region-of-interest (RoI) detection are either blind to global contexts, detail richness compromising, or too expensive for 3-D tasks. To tackle these drawbacks, we propose a novel encoder-decoder-based framework for 3-D whole volume segmentation, referred to as 3-D RoI-aware U-Net (3-D RU-Net). 3-D RU-Net fully utilizes the global contexts covering large effective receptive fields. Specifically, the proposed model consists of a global image encoder for global understanding-based RoI localization, and a local region decoder that operates on pyramid-shaped in-region global features, which is GPU memory efficient and thereby enables training and prediction with large 3-D whole volumes. To facilitate the global-to-local learning procedure and enhance contour detail richness, we designed a dice-based multitask hybrid loss function. The efficiency of the proposed framework enables an extensive model ensemble for further performance gain at acceptable extra computational costs. Over a dataset of 64 T2-weighted MR images, the experimental results of four-fold cross-validation show that our method achieved 75.5% dice similarity coefficient (DSC) in 0.61 s per volume on a GPU, which significantly outperforms competing methods in terms of accuracy and efficiency. The code is publicly available. Yi-Jie Huang, Qi Dou 0001, Zi-Xian Wang, Li-Zhi Liu, Chao-Feng Li, Lisheng Wang, Hao Chen 0011, Rui-Hua Xu |
IEEE Trans. Cybern. | 8 |
| 2021 | Transformation-Consistent Self-Ensembling Model for Semisupervised Medical Image SegmentationabstractA common shortfall of supervised deep learning for medical imaging is the lack of labeled data, which is often expensive and time consuming to collect. This article presents a new semisupervised method for medical image segmentation, where the network is optimized by a weighted combination of a common supervised loss only for the labeled inputs and a regularization loss for both the labeled and unlabeled data. To utilize the unlabeled data, our method encourages consistent predictions of the network-in-training for the same input under different perturbations. With the semisupervised segmentation tasks, we introduce a transformation-consistent strategy in the self-ensembling model to enhance the regularization effect for pixel-level predictions. To further improve the regularization effects, we extend the transformation in a more generalized form including scaling and optimize the consistency loss with a teacher model, which is an averaging of the student model weights. We extensively validated the proposed semisupervised method on three typical yet challenging medical image segmentation tasks: 1) skin lesion segmentation from dermoscopy images in the International Skin Imaging Collaboration (ISIC) 2017 data set; 2) optic disk (OD) segmentation from fundus images in the Retinal Fundus Glaucoma Challenge (REFUGE) data set; and 3) liver segmentation from volumetric CT scans in the Liver Tumor Segmentation Challenge (LiTS) data set. Compared with state-of-the-art, our method shows superior performance on the challenging 2-D/3-D medical images, demonstrating the effectiveness of our semisupervised method for medical image segmentation. Xiaomeng Li 0001, Lequan Yu, Hao Chen 0011, Chi-Wing Fu, Lei Xing 0001, Pheng-Ann Heng |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2020 | Deep Semi-supervised Knowledge Distillation for Overlapping Cervical Cell Instance Segmentation
Yanning Zhou 0001, Hao Chen 0011, Huangjing Lin, Pheng-Ann Heng |
MICCAI (1) | 2 |
| 2020 | Multi-task recurrent convolutional network with correlation loss for surgical video analysis
Yueming Jin, Huaxia Li, Qi Dou 0001, Hao Chen 0011, Harry Qin, Chi-Wing Fu, Pheng-Ann Heng |
Medical Image Anal. | 4 |
| 2020 | Towards multi-center glaucoma OCT image screening with semi-supervised joint structure and function multi-task learning
Xi Wang 0013, Hao Chen 0011, An-ran Ran, Luyang Luo, Poemen P. Chan, Clement C. Tham, Robert T. Chang, Suria S. Mannil, Carol Y. Cheung, Pheng-Ann Heng |
Medical Image Anal. | 2 |
| 2020 | Weakly Supervised Deep Learning for Whole Slide Lung Cancer Image AnalysisabstractHistopathology image analysis serves as the gold standard for cancer diagnosis. Efficient and precise diagnosis is quite critical for the subsequent therapeutic treatment of patients. So far, computer-aided diagnosis has not been widely applied in pathological field yet as currently well-addressed tasks are only the tip of the iceberg. Whole slide image (WSI) classification is a quite challenging problem. First, the scarcity of annotations heavily impedes the pace of developing effective approaches. Pixelwise delineated annotations on WSIs are time consuming and tedious, which poses difficulties in building a large-scale training dataset. In addition, a variety of heterogeneous patterns of tumor existing in high magnification field are actually the major obstacle. Furthermore, a gigapixel scale WSI cannot be directly analyzed due to the immeasurable computational cost. How to design the weakly supervised learning methods to maximize the use of available WSI-level labels that can be readily obtained in clinical practice is quite appealing. To overcome these challenges, we present a weakly supervised approach in this article for fast and effective classification on the whole slide lung cancer images. Our method first takes advantage of a patch-based fully convolutional network (FCN) to retrieve discriminative blocks and provides representative deep features with high efficiency. Then, different context-aware block selection and feature aggregation strategies are explored to generate globally holistic WSI descriptor which is ultimately fed into a random forest (RF) classifier for the image-level prediction. To the best of our knowledge, this is the first study to exploit the potential of image-level labels along with some coarse annotations for weakly supervised learning. A large-scale lung cancer WSI dataset is constructed in this article for evaluation, which validates the effectiveness and feasibility of the proposed method. Extensive experiments demonstrate the superior performance of our method that surpasses the state-of-the-art approaches by a significant margin with an accuracy of 97.3%. In addition, our method also achieves the best performance on the public lung cancer WSIs dataset from The Cancer Genome Atlas (TCGA). We highlight that a small number of coarse annotations can contribute to further accuracy improvement. We believe that weakly supervised learning methods have great potential to assist pathologists in histology image diagnosis in the near future. Xi Wang 0013, Hao Chen 0011, Caixia Gan, Huangjing Lin, Qi Dou 0001, Efstratios Tsougenis, Qitao Huang, Muyan Cai, Pheng-Ann Heng |
IEEE Trans. Cybern. | 2 |
| 2020 | UD-MIL: Uncertainty-Driven Deep Multiple Instance Learning for OCT Image ClassificationabstractDeep learning has achieved remarkable success in the optical coherence tomography (OCT) image classification task with substantial labelled B-scan images available. However, obtaining such fine-grained expert annotations is usually quite difficult and expensive. How to leverage the volume-level labels to develop a robust classifier is very appealing. In this paper, we propose a weakly supervised deep learning framework with uncertainty estimation to address the macula-related disease classification problem from OCT images with the only volume-level label being available. First, a convolutional neural network (CNN) based instance-level classifier is iteratively refined by using the proposed uncertainty-driven deep multiple instance learning scheme. To our best knowledge, we are the first to incorporate the uncertainty evaluation mechanism into multiple instance learning (MIL) for training a robust instance classifier. The classifier is able to detect suspicious abnormal instances and abstract the corresponding deep embedding with high representation capability simultaneously. Second, a recurrent neural network (RNN) takes instance features from the same bag as input and generates the final bag-level prediction by considering the individually local instance information and globally aggregated bag-level representation. For more comprehensive validation, we built two large diabetic macular edema (DME) OCT datasets from different devices and imaging protocols to evaluate the efficacy of our method, which are composed of 30,151 B-scans in 1,396 volumes from 274 patients (Heidelberg-DME dataset) and 38,976 B-scans in 3,248 volumes from 490 patients (Triton-DME dataset), respectively. We compare the proposed method with the state-of-the-art approaches, and experimentally demonstrate that our method is superior to alternative methods, achieving volume-level accuracy, F1-score and area under the receiver operating characteristic curve (AUC) of 95.1%, 0.939 and 0.990 on Heidelberg-DME and those of 95.1%, 0.935 and 0.986 on Triton-DME, respectively. Furthermore, the proposed method also yields competitive results on another public age-related macular degeneration OCT dataset, indicating the high potential as an effective screening tool in the clinical practice. Xi Wang 0013, Fangyao Tang, Hao Chen 0011, Luyang Luo, Ziqi Tang, An-ran Ran, Carol Y. Cheung, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Unsupervised Bidirectional Cross-Modality Adaptation via Deeply Synergistic Image and Feature Alignment for Medical Image SegmentationabstractUnsupervised domain adaptation has increasingly gained interest in medical image computing, aiming to tackle the performance degradation of deep neural networks when being deployed to unseen data with heterogeneous characteristics. In this work, we present a novel unsupervised domain adaptation framework, named as Synergistic Image and Feature Alignment (SIFA), to effectively adapt a segmentation network to an unlabeled target domain. Our proposed SIFA conducts synergistic alignment of domains from both image and feature perspectives. In particular, we simultaneously transform the appearance of images across domains and enhance domain-invariance of the extracted features by leveraging adversarial learning in multiple aspects and with a deeply supervised mechanism. The feature encoder is shared between both adaptive perspectives to leverage their mutual benefits via end-to-end learning. We have extensively evaluated our method with cardiac substructure segmentation and abdominal multi-organ segmentation for bidirectional cross-modality adaptation between MRI and CT images. Experimental results on two different tasks demonstrate that our SIFA method is effective in improving segmentation performance on unlabeled target images, and outperforms the state-of-the-art domain adaptation approaches by a large margin. Cheng Chen 0013, Qi Dou 0001, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Rectifying Supporting Regions With Mixed and Active Supervision for Rib Fracture RecognitionabstractAutomatic rib fracture recognition from chest X-ray images is clinically important yet challenging due to weak saliency of fractures. Weakly Supervised Learning (WSL) models recognize fractures by learning from large-scale image-level labels. In WSL, Class Activation Maps (CAMs) are considered to provide spatial interpretations on classification decisions. However, the high-responding regions, namely Supporting Regions of CAMs may erroneously lock to regions irrelevant to fractures, which thereby raises concerns on the reliability of WSL models for clinical applications. Currently available Mixed Supervised Learning (MSL) models utilize object-level labels to assist fitting WSL-derived CAMs. However, as a prerequisite of MSL, the large quantity of precisely delineated labels is rarely available for rib fracture tasks. To address these problems, this paper proposes a novel MSL framework. Firstly, by embedding the adversarial classification learning into WSL frameworks, the proposed Biased Correlation Decoupling and Instance Separation Enhancing strategies guide CAMs to true fractures indirectly. The CAM guidance is insensitive to shape and size variations of object descriptions, thereby enables robust learning from bounding boxes. Secondly, to further minimize annotation cost in MSL, a CAM-based Active Learning strategy is proposed to recognize and annotate samples whose Supporting Regions cannot be confidently localized. Consequently, the quantity demand of object-level labels can be reduced without compromising the performance. Over a chest X-ray rib-fracture dataset of 10966 images, the experimental results show that our method produces rational Supporting Regions to interpret its classification decisions and outperforms competing methods at an expense of annotating 20% of the positive samples with bounding boxes. Yi-Jie Huang, Xiuying Wang 0001, Qu Fang, Renzhen Wang, Huai Chen, Hao Chen 0011, Deyu Meng, Lisheng Wang |
IEEE Trans. Medical Imaging | 8 |
| 2020 | A Multi-Organ Nucleus Segmentation ChallengeabstractGeneralized nucleus segmentation techniques can contribute greatly to reducing the time to develop and validate visual biomarkers for new digital pathology datasets. We summarize the results of MoNuSeg 2018 Challenge whose objective was to develop generalizable nuclei segmentation techniques in digital pathology. The challenge was an official satellite event of the MICCAI 2018 conference in which 32 teams with more than 80 participants from geographically diverse institutes participated. Contestants were given a training set with 30 images from seven organs with annotations of 21,623 individual nuclei. A test dataset with 14 images taken from seven organs, including two organs that did not appear in the training set was released without annotations. Entries were evaluated based on average aggregated Jaccard index (AJI) on the test set to prioritize accurate instance segmentation as opposed to mere semantic segmentation. More than half the teams that completed the challenge outperformed a previous baseline. Among the trends observed that contributed to increased accuracy were the use of color normalization as well as heavy data augmentation. Additionally, fully convolutional networks inspired by variants of U-Net, FCN, and Mask-RCNN were popularly used, typically based on ResNet or VGG base architectures. Watershed segmentation on predicted semantic segmentation maps was a popular post-processing strategy. Several of the top techniques compared favorably to an individual human annotator and can be used with confidence for nuclear morphometrics. Neeraj Kumar 0002, Ruchika Verma, Deepak Anand, Yanning Zhou 0001, Omer Fahri Onder, Efstratios Tsougenis, Hao Chen 0011, Pheng-Ann Heng, Jiahui Li 0005, Navid Alemi Koohbanani, Mostafa Jahanifar, Neda Zamani Tajeddin, Ali Gooya, Nasir M. Rajpoot, Xuhua Ren, Sihang Zhou 0001, Qian Wang 0001, Dinggang Shen, Cheng-Kun Yang, Chi-Hung Weng, Wei-Hsiang Yu, Chao-Yuan Yeh, Shuoyu Xu, Pak-Hei Yeung, Amirreza Mahbod, Gerald Schaefer, Isabella Ellinger, Rupert Ecker, Örjan Smedby, Chunliang Wang, Benjamin Chidester, Vinh Ton-That, Minh-Triet Tran, Jian Ma 0004, Minh N. Do, Simon Graham, Quoc Dang Vu, Jin Tae Kwak, Akshaykumar Gunda, Raviteja Chunduri, Corey Hu, Dariush Lotfi, Reza Safdari, Antanas Kascenas, Alison O'Neil, Dennis Eschweiler, Johannes Stegmaier, Yanping Cui, Kailin Chen, Xinmei Tian 0001, Philipp Grüning, Erhardt Barth, Elad Arbel, Itay Remer, Amir Ben-Dor, Ekaterina Sirazitdinova, Matthias Kohl, Stefan Braunewell, Yuexiang Li, Xinpeng Xie, LinLin Shen, Jun Ma 0016, Krishanu Das Baksi, Mohammad Azam Khan, Jaegul Choo, Adrián Colomer, Valery Naranjo, Linmin Pei, Khan M. Iftekharuddin, Kaushiki Roy, Debotosh Bhattacharjee, Aníbal Pedraza, Gloria Bueno García, Sabarinathan Devanathan, Saravanan Radhakrishnan, Praveen Koduganty, Zihan Wu 0001, Guanyu Cai, Amit Sethi |
IEEE Trans. Medical Imaging | 7 |
| 2020 | Multi-Task Deep Model With Margin Ranking Loss for Lung Nodule AnalysisabstractLung cancer is the leading cause of cancer deaths worldwide and early diagnosis of lung nodule is of great importance for therapeutic treatment and saving lives. Automated lung nodule analysis requires both accurate lung nodule benign-malignant classification and attribute score regression. However, this is quite challenging due to the considerable difficulty of lung nodule heterogeneity modeling and the limited discrimination capability on ambiguous cases. To solve these challenges, we propose a Multi-Task deep model with Margin Ranking loss (referred as MTMR-Net) for automated lung nodule analysis. Compared to existing methods which consider these two tasks separately, the relatedness between lung nodule classification and attribute score regression is explicitly explored in a cause-and-effect manner within our multi-task deep model, which can contribute to the performance gains of both tasks. The results of different tasks can be yielded simultaneously for assisting the radiologists in diagnosis interpretation. Furthermore, a Siamese network with a margin ranking loss is elaborately designed to enhance the discrimination capability on ambiguous nodule cases. To further explore the internal relationship between two tasks and validate the effectiveness of the proposed model, we use the recursive feature elimination method to iteratively rank the most malignancy-related features. We validate the efficacy of our method MTMR-Net on the public benchmark LIDC-IDRI dataset. Extensive experiments show that the diagnosis results with internal relationship explicitly explored in our model has met some similar patterns in clinical usage and also demonstrate that our approach can achieve competitive classification performance and more accurate scoring on attributes over the state-of-the-arts. Codes are publicly available at: https://github.com/CaptainWilliam/MTMR-NET. Qi Dou 0001, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2020 | Deep Mining External Imperfect Data for Chest X-Ray Disease ScreeningabstractDeep learning approaches have demonstrated remarkable progress in automatic Chest X-ray analysis. The data-driven feature of deep models requires training data to cover a large distribution. Therefore, it is substantial to integrate knowledge from multiple datasets, especially for medical images. However, learning a disease classification model with extra Chest X-ray (CXR) data is yet challenging. Recent researches have demonstrated that performance bottleneck exists in joint training on different CXR datasets, and few made efforts to address the obstacle. In this paper, we argue that incorporating an external CXR dataset leads to imperfect training data, which raises the challenges. Specifically, the imperfect data is in two folds: domain discrepancy, as the image appearances vary across datasets; and label discrepancy, as different datasets are partially labeled. To this end, we formulate the multi-label thoracic disease classification problem as weighted independent binary tasks according to the categories. For common categories shared across domains, we adopt task-specific adversarial training to alleviate the feature differences. For categories existing in a single dataset, we present uncertainty-aware temporal ensembling of model predictions to mine the information from the missing labels further. In this way, our framework simultaneously models and tackles the domain and label discrepancies, enabling superior knowledge mining ability. We conduct extensive experiments on three datasets with more than 360,000 Chest X-ray images. Our method outperforms other competing models and sets state-of-the-art performance on the official NIH test set with 0.8349 AUC, demonstrating its effectiveness of utilizing the external dataset to improve the internal classification. Luyang Luo, Lequan Yu, Hao Chen 0011, Quande Liu, Xi Wang 0013, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2019 | Synergistic Image and Feature Adaptation: Towards Cross-Modality Domain Adaptation for Medical Image SegmentationabstractThis paper presents a novel unsupervised domain adaptation framework, called Synergistic Image and Feature Adaptation (SIFA), to effectively tackle the problem of domain shift. Domain adaptation has become an important and hot topic in recent studies on deep learning, aiming to recover performance degradation when applying the neural networks to new testing domains. Our proposed SIFA is an elegant learning diagram which presents synergistic fusion of adaptations from both image and feature perspectives. In particular, we simultaneously transform the appearance of images across domains and enhance domain-invariance of the extracted features towards the segmentation task. The feature encoder layers are shared by both perspectives to grasp their mutual benefits during the end-to-end learning procedure. Without using any annotation from the target domain, the learning of our unified model is guided by adversarial losses, with multiple discriminators employed from various aspects. We have extensively validated our method with a challenging application of crossmodality medical image segmentation of cardiac structures. Experimental results demonstrate that our SIFA model recovers the degraded performance from 17.2% to 73.0%, and outperforms the state-of-the-art methods by a significant margin. Cheng Chen 0013, Qi Dou 0001, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
AAAI | 3 |
| 2019 | Robust Multimodal Brain Tumor Segmentation via Feature Disentanglement and Gated Fusion
Cheng Chen 0013, Qi Dou 0001, Yueming Jin, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
MICCAI (3) | 4 |
| 2019 | PRSNet: Part Relation and Selection Network for Bone Age Assessment
Yuanfeng Ji, Hao Chen 0011, Dan Lin 0009, Di Lin 0002 |
MICCAI (6) | 2 |
| 2019 | Deep Angular Embedding and Feature Correlation Attention for Breast MRI Cancer Analysis
Luyang Luo, Hao Chen 0011, Xi Wang 0013, Qi Dou 0001, Huangjing Lin, Gongjie Li, Pheng-Ann Heng |
MICCAI (4) | 2 |
| 2019 | Unifying Structure Analysis and Surrogate-Driven Function Regression for Glaucoma OCT Image Screening
Xi Wang 0013, Hao Chen 0011, Luyang Luo, An-ran Ran, Poemen P. Chan, Clement C. Tham, Carol Y. Cheung, Pheng-Ann Heng |
MICCAI (1) | 2 |
| 2019 | PFA-ScanNet: Pyramidal Feature Aggregation with Synergistic Learning for Breast Cancer Metastasis Analysis
Huangjing Lin, Hao Chen 0011, Pheng-Ann Heng |
MICCAI (1) | 3 |
| 2019 | IRNet: Instance Relation Network for Overlapping Cervical Cell Segmentation
Yanning Zhou 0001, Hao Chen 0011, Qi Dou 0001, Pheng-Ann Heng |
MICCAI (1) | 2 |
| 2019 | MILD-Net: Minimal information loss dilated network for gland instance segmentation in colon histology images
Simon Graham, Hao Chen 0011, Jevgenij Gamper, Qi Dou 0001, Pheng-Ann Heng, David R. J. Snead, Yee-Wah Tsang, Nasir M. Rajpoot |
Medical Image Anal. | 2 |
| 2019 | RMDL: Recalibrated multi-instance deep learning for whole slide gastric image classification
Yaxi Zhu, Lequan Yu, Hao Chen 0011, Huangjing Lin, Xiangbo Wan, Xinjuan Fan, Pheng-Ann Heng |
Medical Image Anal. | 4 |
| 2019 | SINet: A Scale-Insensitive Convolutional Neural Network for Fast Vehicle DetectionabstractVision-based vehicle detection approaches achieve incredible success in recent years with the development of deep convolutional neural network (CNN). However, existing CNN-based algorithms suffer from the problem that the convolutional features are scale-sensitive in object detection task but it is common that traffic images and videos contain vehicles with a large variance of scales. In this paper, we delve into the source of scale sensitivity, and reveal two key issues: 1) existing RoI pooling destroys the structure of small scale objects and 2) the large intra-class distance for a large variance of scales exceeds the representation capability of a single network. Based on these findings, we present a scale-insensitive convolutional neural network (SINet) for fast detecting vehicles with a large variance of scales. First, we present a context-aware RoI pooling to maintain the contextual information and original structure of small scale objects. Second, we present a multi-branch decision network to minimize the intra-class distance of features. These lightweight techniques bring zero extra time complexity but prominent detection accuracy improvement. The proposed techniques can be equipped with any deep network architectures and keep them trained end-to-end. Our SINet achieves state-of-the-art performance in terms of accuracy and speed (up to 37 FPS) on the KITTI benchmark and a new highway dataset, which contains a large variance of scales and extremely small objects. Xiaowei Hu 0001, Xuemiao Xu, Yongjie Xiao, Hao Chen 0011, Shengfeng He, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2019 | From Detection of Individual Metastases to Classification of Lymph Node Status at the Patient Level: The CAMELYON17 ChallengeabstractAutomated detection of cancer metastases in lymph nodes has the potential to improve the assessment of prognosis for patients. To enable fair comparison between the algorithms for this purpose, we set up the CAMELYON17 challenge in conjunction with the IEEE International Symposium on Biomedical Imaging 2017 Conference in Melbourne. Over 300 participants registered on the challenge website, of which 23 teams submitted a total of 37 algorithms before the initial deadline. Participants were provided with 899 whole-slide images (WSIs) for developing their algorithms. The developed algorithms were evaluated based on the test set encompassing 100 patients and 500 WSIs. The evaluation metric used was a quadratic weighted Cohen's kappa. We discuss the algorithmic details of the 10 best pre-conference and two post-conference submissions. All these participants used convolutional neural networks in combination with pre- and postprocessing steps. Algorithms differed mostly in neural network architecture, training strategy, and pre- and postprocessing methodology. Overall, the kappa metric ranged from 0.89 to -0.13 across all submissions. The best results were obtained with pre-trained architectures such as ResNet. Confusion matrix analysis revealed that all participants struggled with reliably identifying isolated tumor cells, the smallest type of metastasis, with detection rates below 40%. Qualitative inspection of the results of the top participants showed categories of false positives, such as nerves or contamination, which could be targeted for further optimization. Last, we show that simple combinations of the top algorithms result in higher kappa metric values than any algorithm individually, with 0.93 for the best combination. Péter Bándi, Oscar Geessink, Quirine Manson, Marcory Van Dijk, Maschenka Balkenhol, Meyke Hermsen, Babak Ehteshami Bejnordi, Byungjae Lee, Kyunghyun Paeng, Aoxiao Zhong, Quanzheng Li, Farhad G. Zanjani, Svitlana Zinger, Keisuke Fukuta, Daisuke Komura, Vlado Ovtcharov, Shenghua Cheng, Shaoqun Zeng, Jeppe Thagaard, Anders Bjorholm Dahl, Huangjing Lin, Hao Chen 0011, Ludwig Jacobsson, Martin Hedlund, Melih Çetin, Eren Halici, Hunter Jackson, Fabian Both, Jörg Franke, Heidi Küsters-Vandevelde, Willem Vreuls, Peter Bult, Bram van Ginneken, Jeroen van der Laak, Geert Litjens 0001 |
IEEE Trans. Medical Imaging | 22 |
| 2019 | Fast ScanNet: Fast and Dense Analysis of Multi-Gigapixel Whole-Slide Images for Cancer Metastasis DetectionabstractLymph node metastasis is one of the most important indicators in breast cancer diagnosis, that is traditionally observed under the microscope by pathologists. In recent years, with the dramatic advance of high-throughput scanning and deep learning technology, automatic analysis of histology from whole-slide images has received a wealth of interest in the field of medical image computing, which aims to alleviate pathologists' workload and simultaneously reduce misdiagnosis rate. However, the automatic detection of lymph node metastases from whole-slide images remains a key challenge because such images are typically very large, where they can often be multiple gigabytes in size. Also, the presence of hard mimics may result in a large number of false positives. In this paper, we propose a novel method with anchor layers for model conversion, which not only leverages the efficiency of fully convolutional architectures to meet the speed requirement in clinical practice but also densely scans the whole-slide image to achieve accurate predictions on both micro- and macro-metastases. Incorporating the strategies of asynchronous sample prefetching and hard negative mining, the network can be effectively trained. The efficacy of our method is corroborated on the benchmark dataset of 2016 Camelyon Grand Challenge. Our method achieved significant improvements in comparison with the state-of-the-art methods on tumor localization accuracy with a much faster speed and even surpassed human performance on both challenge tasks. Huangjing Lin, Hao Chen 0011, Simon Graham, Qi Dou 0001, Nasir M. Rajpoot, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2018 | SFCN-OPI: Detection and Fine-Grained Classification of Nuclei Using Sibling FCN With Objectness Prior InteractionabstractCell nuclei detection and fine-grained classification have been fundamental yet challenging problems in histopathology image analysis. Due to the nuclei tiny size, significant inter-/intra-class variances, as well as the inferior image quality, previous automated methods would easily suffer from limited accuracy and robustness. In the meanwhile, existing approaches usually deal with these two tasks independently, which would neglect the close relatedness of them. In this paper, we present a novel method of sibling fully convolutional network with prior objectness interaction (called SFCN-OPI) to tackle the two tasks simultaneously and interactively using a unified end-to-end framework. Specifically, the sibling FCN branches share features in earlier layers while holding respective higher layers for specific tasks. More importantly, the detection branch outputs the objectness prior which dynamically interacts with the fine-grained classification sibling branch during the training and testing processes. With this mechanism, the fine-grained classification successfully focuses on regions with high confidence of nuclei existence and outputs the conditional probability, which in turn benefits the detection through back propagation. Extensive experiments on colon cancer histology images have validated the effectiveness of our proposed SFCN-OPI and our method has outperformed the state-of-the-art methods by a large margin. Yanning Zhou 0001, Qi Dou 0001, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
AAAI | 3 |
| 2018 | Semi-supervised Skin Lesion Segmentation via Transformation Consistent Self-ensembling Model
Xiaomeng Li 0001, Lequan Yu, Hao Chen 0011, Chi-Wing Fu, Pheng-Ann Heng |
BMVC | 3 |
| 2018 | Unsupervised Cross-Modality Domain Adaptation of ConvNets for Biomedical Image Segmentations with Adversarial LossabstractConvolutional networks (ConvNets) have achieved great successes in various challenging vision tasks. However, the performance of ConvNets would degrade when encountering the domain shift. The domain adaptation is more significant while challenging in the field of biomedical image analysis, where cross-modality data have largely different distributions. Given that annotating the medical data is especially expensive, the supervised transfer learning approaches are not quite optimal. In this paper, we propose an unsupervised domain adaptation framework with adversarial learning for cross-modality biomedical image segmentations. Specifically, our model is based on a dilated fully convolutional network for pixel-wise prediction. Moreover, we build a plug-and-play domain adaptation module (DAM) to map the target input to features which are aligned with source domain feature space. A domain critic module (DCM) is set up for discriminating the feature space of both domains. We optimize the DAM and DCM via an adversarial loss without using any target domain label. Our proposed method is validated by adapting a ConvNet trained with MRI images to unpaired CT data for cardiac structures segmentations, and achieved very promising results. Qi Dou 0001, Cheng Ouyang, Cheng Chen 0013, Hao Chen 0011, Pheng-Ann Heng |
IJCAI | 4 |
| 2018 | ScanNet: A Fast and Dense Scanning Framework for Metastastic Breast Cancer Detection from Whole-Slide ImageabstractLymph node metastasis is one of the most significant diagnostic indicators in breast cancer, which is traditionally observed under the microscope by pathologists. In recent years, computerized histology diagnosis has become one of the most rapidly expanding directions in the field of medical image computing, which aims to alleviate pathologists' workload and simultaneously reduce misdiagnosis rate. However, automatic detection of lymph node metastases from whole slide images remains a challenging problem, due to the large-scale data with enormous resolutions and existence of hard mimics resulting in a large number of false positives. In this paper, we propose a novel framework by leveraging fully convolutional networks for efficient inference to meet the speed requirement for clinical practice, while reconstructing dense predictions under different offsets for ensuring accurate detection on both microand macro-metastases. Incorporating with the strategies of asynchronous sample prefetching and hard negative mining, the network can be effectively trained. Extensive experiments on the benchmark dataset of 2016 Camelyon Grand Challenge corroborated the efficacy of our method. Compared with the state-of-the-art methods, our method achieved superior performance with a faster speed on the tumor localization task and even surpassed human performance on the WSI classification task. Huangjing Lin, Hao Chen 0011, Qi Dou 0001, Liansheng Wang 0002, Harry Qin, Pheng-Ann Heng |
WACV | 2 |
| 2018 | 3D multi-scale FCN with random modality voxel dropout learning for Intervertebral Disc Localization and Segmentation from Multi-modality MR Images
Xiaomeng Li 0001, Qi Dou 0001, Hao Chen 0011, Chi-Wing Fu, Xiaojuan Qi 0001, Daniel L. Belavy, Gabriele Armbrecht, Dieter Felsenberg, Guoyan Zheng, Pheng-Ann Heng |
Medical Image Anal. | 3 |
| 2018 | CNNs-Based RGB-D Saliency Detection via Cross-View Transfer and Multiview FusionabstractSalient object detection from RGB-D images aims to utilize both the depth view and RGB view to automatically localize objects of human interest in the scene. Although a few earlier efforts have been devoted to the study of this paper in recent years, two major challenges still remain: 1) how to leverage the depth view effectively to model the depth-induced saliency and 2) how to implement an optimal combination of the RGB view and depth view, which can make full use of complementary information among them. To address these two challenges, this paper proposes a novel framework based on convolutional neural networks (CNNs), which transfers the structure of the RGB-based deep neural network to be applicable for depth view and fuses the deep representations of both views automatically to obtain the final saliency map. In the proposed framework, the first challenge is modeled as a cross-view transfer problem and addressed by using the task-relevant initialization and adding deep supervision in hidden layer. The second challenge is addressed by a multiview CNN fusion model through a combination layer connecting the representation layers of RGB view and depth view. Comprehensive experiments on four benchmark datasets demonstrate the significant and consistent improvements of the proposed approach over other state-of-the-art methods. Junwei Han 0001, Hao Chen 0011, Nian Liu 0002, Chenggang Yan 0001, Xuelong Li 0001 |
IEEE Trans. Cybern. | 2 |
| 2018 | SV-RCNet: Workflow Recognition From Surgical Videos Using Recurrent Convolutional NetworkabstractWe propose an analysis of surgical videos that is based on a novel recurrent convolutional network (SV-RCNet), specifically for automatic workflow recognition from surgical videos online, which is a key component for developing the context-aware computer-assisted intervention systems. Different from previous methods which harness visual and temporal information separately, the proposed SV-RCNet seamlessly integrates a convolutional neural network (CNN) and a recurrent neural network (RNN) to form a novel recurrent convolutional architecture in order to take full advantages of the complementary information of visual and temporal features learned from surgical videos. We effectively train the SV-RCNet in an end-to-end manner so that the visual representations and sequential dynamics can be jointly optimized in the learning process. In order to produce more discriminative spatio-temporal features, we exploit a deep residual network (ResNet) and a long short term memory (LSTM) network, to extract visual features and temporal dependencies, respectively, and integrate them into the SV-RCNet. Moreover, based on the phase transition-sensitive predictions from the SV-RCNet, we propose a simple yet effective inference scheme, namely the prior knowledge inference (PKI), by leveraging the natural characteristic of surgical video. Such a strategy further improves the consistency of results and largely boosts the recognition performance. Extensive experiments have been conducted with the MICCAI 2016 Modeling and Monitoring of Computer Assisted Interventions Workflow Challenge dataset and Cholec80 dataset to validate SV-RCNet. Our approach not only achieves superior performance on these two datasets but also outperforms the state-of-the-art methods by a significant margin. Yueming Jin, Qi Dou 0001, Hao Chen 0011, Lequan Yu, Harry Qin, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 3 |
| 2018 | H-DenseUNet: Hybrid Densely Connected UNet for Liver and Tumor Segmentation From CT VolumesabstractLiver cancer is one of the leading causes of cancer death. To assist doctors in hepatocellular carcinoma diagnosis and treatment planning, an accurate and automatic liver and tumor segmentation method is highly demanded in the clinical practice. Recently, fully convolutional neural networks (FCNs), including 2-D and 3-D FCNs, serve as the backbone in many volumetric image segmentation. However, 2-D convolutions cannot fully leverage the spatial information along the third dimension while 3-D convolutions suffer from high computational cost and GPU memory consumption. To address these issues, we propose a novel hybrid densely connected UNet (H-DenseUNet), which consists of a 2-D DenseUNet for efficiently extracting intra-slice features and a 3-D counterpart for hierarchically aggregating volumetric contexts under the spirit of the auto-context algorithm for liver and tumor segmentation. We formulate the learning process of the H-DenseUNet in an end-to-end manner, where the intra-slice representations and inter-slice features can be jointly optimized through a hybrid feature fusion layer. We extensively evaluated our method on the data set of the MICCAI 2017 Liver Tumor Segmentation Challenge and 3DIRCADb data set. Our method outperformed other state-of-the-arts on the segmentation results of tumors and achieved very competitive performance for liver segmentation even with a single model. Xiaomeng Li 0001, Hao Chen 0011, Xiaojuan Qi 0001, Qi Dou 0001, Chi-Wing Fu, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2017 | Volumetric ConvNets with Mixed Residual Connections for Automated Prostate Segmentation from 3D MR ImagesabstractAutomated prostate segmentation from 3D MR images is very challenging due to large variations of prostate shape and indistinct prostate boundaries. We propose a novel volumetric convolutional neural network (ConvNet) with mixed residual connections to cope with this challenging problem. Compared with previous methods, our volumetric ConvNet has two compelling advantages. First, it is implemented in a 3D manner and can fully exploit the 3D spatial contextual information of input data to perform efficient, precise and volume-to-volume prediction. Second and more important, the novel combination of residual connections (i.e., long and short) can greatly improve the training efficiency and discriminative capability of our network by enhancing the information propagation within the ConvNet both locally and globally. While the forward propagation of location information can improve the segmentation accuracy, the smooth backward propagation of gradient flow can accelerate the convergence speed and enhance the discrimination capability. Extensive experiments on the open MICCAI PROMISE12 challenge dataset corroborated the effectiveness of the proposed volumetric ConvNet with mixed residual connections. Our method ranked the first in the challenge, outperforming other competitors by a large margin with respect to most of evaluation metrics. The proposed volumetric ConvNet is general enough and can be easily extended to other medical image analysis tasks, especially ones with limited training data. Lequan Yu, Xin Yang 0009, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
AAAI | 3 |
| 2017 | Automated Pulmonary Nodule Detection via 3D ConvNets with Online Sample Filtering and Hybrid-Loss Residual Learning
Qi Dou 0001, Hao Chen 0011, Yueming Jin, Huangjing Lin, Harry Qin, Pheng-Ann Heng |
MICCAI (3) | 2 |
| 2017 | Automatic 3D Cardiovascular MR Segmentation with Densely-Connected Volumetric ConvNets
Lequan Yu, Jie-Zhi Cheng, Qi Dou 0001, Xin Yang 0009, Hao Chen 0011, Harry Qin, Pheng-Ann Heng |
MICCAI (2) | 5 |
| 2017 | DCAN: Deep contour-aware networks for object instance segmentation from histology images
Hao Chen 0011, Xiaojuan Qi 0001, Lequan Yu, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
Medical Image Anal. | 1 |
| 2017 | 3D deeply supervised network for automated segmentation of volumetric medical images
Qi Dou 0001, Lequan Yu, Hao Chen 0011, Yueming Jin, Xin Yang 0009, Harry Qin, Pheng-Ann Heng |
Medical Image Anal. | 3 |
| 2017 | Validation, comparison, and combination of algorithms for automatic detection of pulmonary nodules in computed tomography images: The LUNA16 challenge
Arnaud A. A. Setio, Alberto Traverso, Thomas de Bel, Moira S. N. Berens, Cas van den Bogaard, Piergiorgio Cerello, Hao Chen 0011, Qi Dou 0001, Maria Evelina Fantacci, Bram Geurts, Robbert van der Gugten, Pheng-Ann Heng, Bart Jansen 0001, Michael M. J. de Kaste, Valentin Kotov, Jack Yu-Hung Lin, Jeroen T. M. C. Manders, Alexander Sóñora-Mengana, Juan Carlos García-Naranjo, Evgenia Papavasileiou, Mathias Prokop |
Medical Image Anal. | 7 |
| 2017 | Gland segmentation in colon histology images: The glas challenge contest
Korsuk Sirinukunwattana, Josien P. W. Pluim, Hao Chen 0011, Xiaojuan Qi 0001, Pheng-Ann Heng, Li Yang Wang, Bogdan J. Matuszewski, Elia Bruni, Urko Sanchez, Anton Böhm, Olaf Ronneberger, Bassem Ben Cheikh, Daniel Racoceanu, Philipp Kainz, Michael Pfeiffer 0001, Martin Urschler, David R. J. Snead, Nasir M. Rajpoot |
Medical Image Anal. | 3 |
| 2017 | Evaluation and comparison of 3D intervertebral disc localization and segmentation methods for 3D T2 MR data: A grand challenge
Guoyan Zheng, Chengwen Chu, Daniel L. Belavy, Bulat Ibragimov, Robert Korez, Tomaz Vrtovec, Hugo Hutt, Richard M. Everson, Judith Meakin, Isabel Lopez Andrade, Ben Glocker, Hao Chen 0011, Qi Dou 0001, Pheng-Ann Heng, Chunliang Wang, Daniel Forsberg, Ales Neubert, Jurgen Fripp, Martin Urschler, Darko Stern, Maria Wimmer 0002 |
Medical Image Anal. | 12 |
| 2017 | Ultrasound Standard Plane Detection Using a Composite Neural Network FrameworkabstractUltrasound (US) imaging is a widely used screening tool for obstetric examination and diagnosis. Accurate acquisition of fetal standard planes with key anatomical structures is very crucial for substantial biometric measurement and diagnosis. However, the standard plane acquisition is a labor-intensive task and requires operator equipped with a thorough knowledge of fetal anatomy. Therefore, automatic approaches are highly demanded in clinical practice to alleviate the workload and boost the examination efficiency. The automatic detection of standard planes from US videos remains a challenging problem due to the high intraclass and low interclass variations of standard planes, and the relatively low image quality. Unlike previous studies which were specifically designed for individual anatomical standard planes, respectively, we present a general framework for the automatic identification of different standard planes from US videos. Distinct from conventional way that devises hand-crafted visual features for detection, our framework explores in- and between-plane feature learning with a novel composite framework of the convolutional and recurrent neural networks. To further address the issue of limited training data, a multitask learning framework is implemented to exploit common knowledge across detection tasks of distinctive standard planes for the augmentation of feature learning. Extensive experiments have been conducted on hundreds of US fetus videos to corroborate the better efficacy of the proposed framework on the difficult standard plane detection problem. Hao Chen 0011, Lingyun Wu, Qi Dou 0001, Harry Qin, Shengli Li 0001, Jie-Zhi Cheng, Dong Ni 0001, Pheng-Ann Heng |
IEEE Trans. Cybern. | 1 |
| 2017 | Integrating Online and Offline Three-Dimensional Deep Learning for Automated Polyp Detection in Colonoscopy VideosabstractAutomated polyp detection in colonoscopy videos has been demonstrated to be a promising way for colorectal cancer prevention and diagnosis. Traditional manual screening is time consuming, operator dependent, and error prone; hence, automated detection approach is highly demanded in clinical practice. However, automated polyp detection is very challenging due to high intraclass variations in polyp size, color, shape, and texture, and low interclass variations between polyps and hard mimics. In this paper, we propose a novel offline and online three-dimensional (3-D) deep learning integration framework by leveraging the 3-D fully convolutional network (3D-FCN) to tackle this challenging problem. Compared with the previous methods employing hand-crafted features or 2-D convolutional neural network, the 3D-FCN is capable of learning more representative spatio-temporal features from colonoscopy videos, and hence has more powerful discrimination capability. More importantly, we propose a novel online learning scheme to deal with the problem of limited training data by harnessing the specific information of an input video in the learning process. We integrate offline and online learning to effectively reduce the number of false positives generated by the offline network and further improve the detection performance. Extensive experiments on the dataset of MICCAI 2015 Challenge on Polyp Detection demonstrated the better performance of our method when compared with other competitors. Lequan Yu, Hao Chen 0011, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 2 |
| 2017 | Comparative Validation of Polyp Detection Methods in Video Colonoscopy: Results From the MICCAI 2015 Endoscopic Vision ChallengeabstractColonoscopy is the gold standard for colon cancer screening though some polyps are still missed, thus preventing early disease detection and treatment. Several computational systems have been proposed to assist polyp detection during colonoscopy but so far without consistent evaluation. The lack of publicly available annotated databases has made it difficult to compare methods and to assess if they achieve performance levels acceptable for clinical use. The Automatic Polyp Detection sub-challenge, conducted as part of the Endoscopic Vision Challenge (http://endovis.grand-challenge.org) at the international conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) in 2015, was an effort to address this need. In this paper, we report the results of this comparative evaluation of polyp detection methods, as well as describe additional experiments to further explore differences between methods. We define performance metrics and provide evaluation databases that allow comparison of multiple methodologies. Results show that convolutional neural networks are the state of the art. Nevertheless, it is also demonstrated that combining different methodologies can lead to an improved overall performance. Jorge Bernal, Nima Tajkbaksh, Francisco Javier Sánchez, Bogdan J. Matuszewski, Hao Chen 0011, Lequan Yu, Quentin Angermann, Olivier Romain, Bjorn Rustad, Ilangko Balasingham, Konstantin Pogorelov, Sungbin Choi, Quentin Debard, Lena Maier-Hein, Stefanie Speidel, Danail Stoyanov, Patrick Brandao, Henry Córdova, Cristina Sánchez-Montes, Suryakanth R. Gurudu, Gloria Fernández-Esparrach, Xavier Dray, Jianming Liang, Aymeric Histace |
IEEE Trans. Medical Imaging | 5 |
| 2017 | Automated Melanoma Recognition in Dermoscopy Images via Very Deep Residual NetworksabstractAutomated melanoma recognition in dermoscopy images is a very challenging task due to the low contrast of skin lesions, the huge intraclass variation of melanomas, the high degree of visual similarity between melanoma and non-melanoma lesions, and the existence of many artifacts in the image. In order to meet these challenges, we propose a novel method for melanoma recognition by leveraging very deep convolutional neural networks (CNNs). Compared with existing methods employing either low-level hand-crafted features or CNNs with shallower architectures, our substantially deeper networks (more than 50 layers) can acquire richer and more discriminative features for more accurate recognition. To take full advantage of very deep networks, we propose a set of schemes to ensure effective training and learning under limited training data. First, we apply the residual learning to cope with the degradation and overfitting problems when a network goes deeper. This technique can ensure that our networks benefit from the performance gains achieved by increasing network depth. Then, we construct a fully convolutional residual network (FCRN) for accurate skin lesion segmentation, and further enhance its capability by incorporating a multi-scale contextual information integration scheme. Finally, we seamlessly integrate the proposed FCRN (for segmentation) and other very deep residual networks (for classification) to form a two-stage framework. This framework enables the classification network to extract more representative and specific features based on segmented results instead of the whole dermoscopy images, further alleviating the insufficiency of training data. The proposed framework is extensively evaluated on ISBI 2016 Skin Lesion Analysis Towards Melanoma Detection Challenge dataset. Experimental results demonstrate the significant performance gains of the proposed framework, ranking the first in classification and the second in segmentation among 25 teams and 28 teams, respectively. This study corroborates that very deep CNNs with effective training mechanisms can be employed to solve complicated medical image analysis tasks, even with limited training data. Lequan Yu, Hao Chen 0011, Qi Dou 0001, Harry Qin, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2016 | Mitosis Detection in Breast Cancer Histology Images via Deep Cascaded NetworksabstractThe number of mitoses per tissue area gives an important aggressiveness indication of the invasive breast carcinoma.However, automatic mitosis detection in histology images remains a challenging problem. Traditional methods either employ hand-crafted features to discriminate mitoses from other cells or construct a pixel-wise classifier to label every pixel in a sliding window way. While the former suffers from the large shape variation of mitoses and the existence of many mimics with similar appearance, the slow speed of the later prohibits its use in clinical practice.In order to overcome these shortcomings, we propose a fast and accurate method to detect mitosis by designing a novel deep cascaded convolutional neural network, which is composed of two components. First, by leveraging the fully convolutional neural network, we propose a coarse retrieval model to identify and locate the candidates of mitosis while preserving a high sensitivity.Based on these candidates, a fine discrimination model utilizing knowledge transferred from cross-domain is developed to further single out mitoses from hard mimics.Our approach outperformed other methods by a large margin in 2014 ICPR MITOS-ATYPIA challenge in terms of detection accuracy. When compared with the state-of-the-art methods on the 2012 ICPR MITOSIS data (a smaller and less challenging dataset), our method achieved comparable or better results with a roughly 60 times faster speed. Hao Chen 0011, Qi Dou 0001, Xi Wang 0013, Harry Qin, Pheng-Ann Heng |
AAAI | 1 |
| 2016 | Deep Contextual Networks for Neuronal Structure SegmentationabstractThe goal of connectomics is to manifest the interconnections of neural system with the Electron Microscopy (EM) images. However, the formidable size of EM image data renders human annotation impractical, as it may take decades to fulfill the whole job. An alternative way to reconstruct the connectome can be attained with the computerized scheme that can automatically segment the neuronal structures. The segmentation of EM images is very challenging as the depicted structures can be very diverse.To address this difficult problem, a deep contextual network is proposed here by leveraging multi-level contextual information from the deep hierarchical structure to achieve better segmentation performance.To further improve the robustness against the vanishing gradients and strengthen the capability of the back-propagation of gradient flow, auxiliary classifiers are incorporated in the architecture of our deep neural network. It will be shown that our method can effectively parse the semantic meaning from the images with the underlying neural network and accurately delineate the structural boundaries with the reference of low-level contextual cues. Experimental results on the benchmark dataset of 2012 ISBI segmentation challenge of neuronal structures suggest that the proposed method can outperform the state-of-the-art methods by a large margin with respect to different evaluation measurements. Our method can potentially facilitate the automatic connectome analysis from EM images with less human intervention effort. Hao Chen 0011, Xiaojuan Qi 0001, Jie-Zhi Cheng, Pheng-Ann Heng |
AAAI | 1 |
| 2016 | DCAN: Deep Contour-Aware Networks for Accurate Gland SegmentationabstractThe morphology of glands has been used routinely by pathologists to assess the malignancy degree of adenocarcinomas. Accurate segmentation of glands from histology images is a crucial step to obtain reliable morphological statistics for quantitative diagnosis. In this paper, we proposed an efficient deep contour-aware network (DCAN) to solve this challenging problem under a unified multi-task learning framework. In the proposed network, multi-level contextual features from the hierarchical architecture are explored with auxiliary supervision for accurate gland segmentation. When incorporated with multi-task regularization during the training, the discriminative capability of intermediate features can be further improved. Moreover, our network can not only output accurate probability maps of glands, but also depict clear contours simultaneously for separating clustered objects, which further boosts the gland segmentation performance. This unified framework can be efficient when applied to large-scale histopathological data without resorting to additional steps to generate contours based on low-level cues for post-separating. Our method won the 2015 MICCAI Gland Segmentation Challenge out of 13 competitive teams, surpassing all the other methods by a significant margin. Hao Chen 0011, Xiaojuan Qi 0001, Lequan Yu, Pheng-Ann Heng |
CVPR | 1 |
| 2016 | Iterative Multi-domain Regularized Deep Learning for Anatomical Structure Detection and Segmentation from Ultrasound Images
Hao Chen 0011, Yefeng Zheng 0001, Jin Hyeong Park, Pheng-Ann Heng, Shaohua Kevin Zhou |
MICCAI (2) | 1 |
| 2016 | 3D Deeply Supervised Network for Automatic Liver Segmentation from CT Volumes
Qi Dou 0001, Hao Chen 0011, Yueming Jin, Lequan Yu, Harry Qin, Pheng-Ann Heng |
MICCAI (2) | 2 |
| 2016 | Automatic Detection of Cerebral Microbleeds From MR Images via 3D Convolutional Neural NetworksabstractCerebral microbleeds (CMBs) are small haemorrhages nearby blood vessels. They have been recognized as important diagnostic biomarkers for many cerebrovascular diseases and cognitive dysfunctions. In current clinical routine, CMBs are manually labelled by radiologists but this procedure is laborious, time-consuming, and error prone. In this paper, we propose a novel automatic method to detect CMBs from magnetic resonance (MR) images by exploiting the 3D convolutional neural network (CNN). Compared with previous methods that employed either low-level hand-crafted descriptors or 2D CNNs, our method can take full advantage of spatial contextual information in MR volumes to extract more representative high-level features for CMBs, and hence achieve a much better detection accuracy. To further improve the detection performance while reducing the computational cost, we propose a cascaded framework under 3D CNNs for the task of CMB detection. We first exploit a 3D fully convolutional network (FCN) strategy to retrieve the candidates with high probabilities of being CMBs, and then apply a well-trained 3D CNN discrimination model to distinguish CMBs from hard mimics. Compared with traditional sliding window strategy, the proposed 3D FCN strategy can remove massive redundant computations and dramatically speed up the detection process. We constructed a large dataset with 320 volumetric MR scans and performed extensive experiments to validate the proposed method, which achieved a high sensitivity of 93.16% with an average number of 2.74 false positives per subject, outperforming previous methods using low-level descriptors or 2D CNNs by a significant margin. The proposed method, in principle, can be adapted to other biomarker detection tasks from volumetric medical data. Qi Dou 0001, Hao Chen 0011, Lequan Yu, Lei Zhao 0003, Harry Qin, Defeng Wang, Vincent C. T. Mok, Lin Shi 0001, Pheng-Ann Heng |
IEEE Trans. Medical Imaging | 2 |
| 2015 | Square Localization for Efficient and Accurate Object DetectionabstractThe key contribution of this paper is the compact square object localization, which relaxes the exhaustive sliding window from testing all windows of different combinations of aspect ratios. Square object localization is category scalable. By using a binary search strategy, the number of scales to test is further reduced empirically to only O(log(min{H, W})) rounds of sliding CNNs, where H and W are respectively the image height and width. In the training phase, square CNN models and object co-presence priors are learned. In the testing phase, sliding CNN models are applied which produces a set of response maps that can be effectively filtered by the learned co-presence prior to output the final bounding boxes for localizing an object. We performed extensive experimental evaluation on the VOC 2007 and 2012 datasets to demonstrate that while efficient, square localization can output precise bounding boxes to improve the final detection result. Cewu Lu, Yongyi Lu, Hao Chen 0011, Chi-Keung Tang |
ICCV | 3 |
| 2015 | Automatic Fetal Ultrasound Standard Plane Detection Using Knowledge Transferred Recurrent Neural Networks
Hao Chen 0011, Qi Dou 0001, Dong Ni 0001, Jie-Zhi Cheng, Harry Qin, Shengli Li 0001, Pheng-Ann Heng |
MICCAI (1) | 1 |
| 2015 | Automatic Localization and Identification of Vertebrae in Spine CT via a Joint Learning Model with Deep Neural Networks
Hao Chen 0011, Chiyao Shen, Harry Qin, Dong Ni 0001, Lin Shi 0001, Jack Chun-Yiu Cheng, Pheng-Ann Heng |
MICCAI (1) | 1 |
| 2015 | Standard Plane Localization in Fetal Ultrasound via Domain Transferred Deep Neural NetworksabstractAutomatic localization of the standard plane containing complicated anatomical structures in ultrasound (US) videos remains a challenging problem. In this paper, we present a learning-based approach to locate the fetal abdominal standard plane (FASP) in US videos by constructing a domain transferred deep convolutional neural network (CNN). Compared with previous works based on low-level features, our approach is able to represent the complicated appearance of the FASP and hence achieve better classification performance. More importantly, in order to reduce the overfitting problem caused by the small amount of training samples, we propose a transfer learning strategy, which transfers the knowledge in the low layers of a base CNN trained from a large database of natural images to our task-specific CNN. Extensive experiments demonstrate that our approach outperforms the state-of-the-art method for the FASP localization as well as the CNN only trained on the limited US training samples. The proposed approach can be easily extended to other similar medical image computing problems, which often suffer from the insufficient training samples when exploiting the deep CNN to represent high-level features. Hao Chen 0011, Dong Ni 0001, Harry Qin, Shengli Li 0001, Xin Yang 0009, Tianfu Wang 0001, Pheng-Ann Heng |
IEEE J. Biomed. Health Informatics | 1 |
| 2008 | Multilinear analysis based on image texture for face recognitionabstractIn this paper, a multilinear approach based on image texture for face recognition is present. First, we extract the texture features of the facial images using the local binary pattern (LBP) algorithm. Then, we apply the high-order orthogonal iteration (HOOI) algorithm, the algebra of higher-order tensors, to obtain a compact and effective representation of the facial images based on the texture features. Our representation yields improved facial recognition rates relative to standard eigenface and tensorface especially when the facial images are confronted by a variety of viewpoints and illuminations. To evaluate the validity of our approach, a series of experiments are performed on the CMU PIE facial databases. Huchuan Lu, Hao Chen 0011, Yen-Wei Chen 0001 |
ICPR | 2 |