EDBT 2026 Demo / reviewers in the wild / expert
Le Lu 0001
dblp:78/6574-1
· DBLP profile ↗
166ranked-venue papers
13as first author
90since 2021 · last 2026
0000-0002-6799-9416ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 115 · 9 first-author · 58 since 2021Applied, interdisciplinary, general and emerging computing · 99 · 1 first-author · 56 since 2021Artificial intelligence and machine learning · 60 · 11 first-author · 30 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MUSE: Multi-Scale Dense Self-Distillation for Nucleus Detection and ClassificationabstractNucleus detection and classification (NDC) in histopathology analysis is a fundamental task that underpins a wide range of high-level pathology applications. However, existing methods heavily rely on labor-intensive nucleus-level annotations and struggle to fully exploit large-scale unlabeled data for learning discriminative nucleus representations. In this work, we propose MUSE (MUlti-scale denSE self-distillation), a novel self-supervised learning method tailored for NDC. At its core is NuLo (Nucleus-based Local self-distillation), a coordinate-guided mechanism that enables flexible local self-distillation based on predicted nucleus positions. By removing the need for strict spatial alignment between augmented views, NuLo allows critical cross-scale alignment, thus unlocking the capacity of models for fine-grained nucleus-level representation. To support MUSE, we design a simple yet effective encoder-decoder architecture and a large field-of-view semi-supervised fine-tuning strategy that together maximize the value of unlabeled pathology images. Extensive experiments on three widely used benchmarks demonstrate that MUSE effectively addresses the core challenges of histopathological NDC. The resulting models not only surpass state-of-the-art supervised baselines but also outperform generic pathology foundation models. Zijiang Yang 0009, Hanqing Chao, Bokai Zhao, Yelin Yang, Yunshuo Zhang, Dongmei Fu, Junping Zhang, Le Lu 0001, Ke Yan 0006, Dakai Jin, Minfeng Xu, Yun Bian |
AAAI | 8 |
| 2026 | Prediction of post-stroke brain swelling using biomechanical modelling and deep neural networksabstract• Malignant stroke is a life-threatening condition, with mortality rates reaching up to 80% among patients managed conservatively. • We developed a novel computational approach that combines advanced biomechanical modelling with deep neural network (DNN) for predicting brain swelling following stroke. • Using in-silico simulations of 3,000 stroke cases, our DNN model demonstrated robust learning of anatomical features, achieving minimal errors in brain swelling prediction. • Without the need for retraining or fine-tuning, our model achieved clinically comparable performance for predicting 3-month stroke outcomes. • To the best of our knowledge, this presents the one of few pioneering computational workflows specifically designed for brain swelling prediction. Malignant stroke is a life-threatening condition, with mortality rates reaching up to 80% among patients managed conservatively. Brain swelling volume and midline shift are pivotal clinical markers for predicting stroke outcomes. However, brain oedema typically peaks two to five days post-stroke onset, which significantly delays the implementation of timely interventions. Early prediction of these markers is therefore critical for enhancing treatment strategies and improving patient outcomes. Predicting brain swelling is inherently challenging due to its biomechanical complexity, governed by factors such as lesion size and location. In this study, we proposed a novel physics-informed computational approach that combines advanced biomechanical modelling with deep neural networks (DNNs) for predicting brain swelling following stroke. Using in-silico simulations of 3,000 stroke cases generated from a poroelasticity-based mathematical model, our DNN model demonstrated robust learning of anatomical features, achieving minimal errors in brain swelling predictions for hold-out test cases. Furthermore, we externally validated the trained model using clinical imaging data from 60 stroke patients; without the need for retraining or fine-tuning, the model achieved clinically comparable outcomes, with an area under the curve (AUC) of approximately 0.7 for predicting 3-month stroke outcomes. Our findings underscore the transformative potential of this physics informed neural networks to accelerate clinical decision-making and improve the management of malignant stroke. By enabling earlier and more precise predictions of critical imaging markers, this pipeline provides a significant step forward in reducing the mortality and disability associated with this devastating condition. Wahbi K. El-Bouri, Stephen J. Payne, Le Lu 0001 |
Medical Image Anal. | 4 |
| 2026 | Non-contrast CT esophageal varices grading through clinical prior-enhanced multi-organ analysis
Xiaoming Zhang 0008, Chunli Li, Jiacheng Hao, Yuan Gao 0017, Danyang Tu, Jianyi Qiao, Xiaoli Yin, Le Lu 0001, Ling Zhang 0002, Ke Yan 0006 |
Medical Image Anal. | 8 |
| 2026 | Neural Wave Propagation for Surgical Video Action Recognition: A New Dataset and BaselineabstractAccurate and efficient recognition of surgical actions in videos is critical for advancing AI-driven surgical robotics. However, current surgical video action recognition (SVAR) datasets suffer from limitations such as small scale, low resolution, inconsistent annotations, and insufficient action coverage. Most latest video recognition models are trained on large-scale common datasets and underperform in SVAR due to architectures that suppress high-frequency visual details (crucial for recognizing surgical tools and motions) and lack a strong spatial inductive bias, requiring extensive training data for good convergence. This is particularly challenging in the surgical domain, where data access is limited. Therefore, a new baseline is required. To address these issues, we introduce LapSurg-230K, an SVAR dataset of 7,569 high-resolution laparoscopic surgical video clips with 230,246 frames, well-annotated for 11 key actions across 9 surgery types. It supports both full and progressive data volume evaluation settings. We further propose WaveR, an attention-free baseline based on physical wave propagation. WaveR embeds an innate physical inductive bias: each video patch acts as a wave source that propagates waves toward action-critical regions (e.g., instrument tips), adaptively aggregating spatial-temporal context while preserving high-frequency surgical cues. This mechanism eliminates dependency on massive training data. Experiments demonstrate WaveR's robustness under extreme data scarcity ( $\leq 30\%$ training samples), achieving state-of-the-art accuracy on both surgical video action recognition and phase recognition tasks. The complete dataset, licensed under CC-BY 4.0, is available at https://doi.org/10.6084/m9.figshare.32237319. Our code is available at https://github.com/yezizi1022/WaveR_TIP. Zhentao Tan, Ru Zhou, Le Lu 0001 |
IEEE Trans. Image Process. | 7 |
| 2026 | Preoperative Prediction of Esophageal Cancer Survival in CT via Tumor and Lymph Node Context and Geometry ModelingabstractEsophageal cancer is one of the most lethal cancers, with 5-year survival rate of only 20%. Patient outcomes can vary significantly even though they are at the same cancer stage and receive similar treatments. Accurate prognostic prediction for esophageal cancer patients is highly desired to receive personalized precise treatment. Nevertheless, there are very few automated methods yet to fully exploit the preoperative contrast-enhanced computed tomography (CE-CT) imaging for assessing esophageal cancer prognosis. In addition to image patterns, important prognostic factors should encompass tumor size and location, as well as lymph nodes (LNs) involvement, including features such as LN number, size, spatial distribution, and their proximity to tumor. Considering these complexities, we propose a novel Tumor and LN Context-Geometry network for the preoperative prediction of esophageal cancer survival in CE-CT images. Specifically, we 1) focus on learning survival patterns of CT texture via co-attention context modeling at most informative regions, i.e., automatically segmented tumor, LNs and LN-stations; and 2) integrate tumor and LN anatomical and spatial associations into neural geometry modeling for a comprehensive learning of metastatic involvement and tumor invasion to adjacent structures. Empirical studies show our presented framework can improve overall survival prediction performances compared with existing state-of-the-art survival analysis methods, and evidently suggest that incorporating these findings into the existing esophageal cancer staging system would add its clinical values. Yirui Wang 0002, Haoshen Li, Jiawen Yao, Lianzhen Zhong, Dazhou Guo, Ke Yan 0006, David S. Doermann, Le Lu 0001, Feiran Jiao, Tsung-Ying Ho, Ling Zhang 0002, Abudili Abuduxuku, Xianghua Ye, Dakai Jin |
IEEE Trans. Medical Imaging | 10 |
| 2026 | Clinical Knowledge-Guided PET/CT Lesion Segmentation With Interpretable Fusion of Metabolic and Structural CuesabstractF-FDG PET/CT images marks a pivotal breakthrough in oncological diagnostics, substantially improving the accuracy and efficiency of tumor burden assessment. Manual segmentation is often plagued by significant inter-observer variability, underscoring the necessity for automated solutions. The synergistic combination of PET's exceptional sensitivity for detecting metabolic activity with CT's anatomical precision renders accurate segmentation crucial for achieving quantitative and reproducible clinical workflows. However, current methodologies frequently grapple with challenges such as over-segmentation or under-segmentation, inadvertently delineating normal tissues with elevated uptake or neglecting lesions characterized by subtle intensity variations, primarily due to a lack of integrated metabolic and anatomical insights. To address these limitations, we present a novel framework that adeptly integrates clinical expertise regarding anatomical and metabolic cues to refine PET/CT lesion segmentation. Our innovative mixture-of-experts (MoE) based interpretable fusion module skillfully merges complementary modality information while explicitly elucidating the pixel-level contributions of each modality to the final segmentation outcome. Rigorous evaluations across three in-domain benchmarks and two external datasets demonstrate our model's superior segmentation performance and generalizability. Furthermore, our visualizations provide compelling insights into the pivotal role each modality plays in the decision-making process, highlighting our approach's transformative potential in enhancing PET/CT lesion segmentation. Building on this foundation, we further validated the prognostic significance of the features extracted from our proposed framework in the context of PET/CT-based prognosis predictions. Jiajin Zhang, Liheng Qiu, Wei Liu 0127, Dakai Jin, Wenpei Jiao, Le Lu 0001, Tzu-Chen Yen, Shenmiao Yang, Ke Yan 0006 |
IEEE Trans. Medical Imaging | 7 |
| 2025 | Towards a Comprehensive, Efficient and Promptable Anatomic Structure Segmentation Model Using 3D Whole-Body CT ScansabstractSegment anything model (SAM) demonstrates strong generalization ability on natural image segmentation. However, its direct adaptation in medical image segmentation tasks shows significant performance drops. It also requires an excessive number of prompt points to obtain a reasonable accuracy. Although quite a few studies explore adapting SAM into medical image volumes, the efficiency of 2D adaptation methods is unsatisfactory and 3D adaptation methods are only capable of segmenting specific organs/tumors. In this work, we propose a comprehensive and scalable 3D SAM model for whole-body CT segmentation, named CT-SAM3D. Instead of adapting SAM, we propose a 3D promptable segmentation model using a (nearly) fully labeled CT dataset. To train CT-SAM3D effectively, ensuring the model's accurate responses to higher-dimensional spatial prompts is crucial, and 3D patch-wise training is required due to GPU memory constraints. Therefore, we propose two key technical developments: 1) a progressively and spatially aligned prompt encoding method to effectively encode click prompts in local 3D space; and 2) a cross-patch prompt scheme to capture more 3D spatial context, which is beneficial for reducing the editing workloads when interactively prompting on large organs. CT-SAM3D is trained using a curated dataset of 1204 CT scans containing 107 whole-body anatomies and extensively validated using five datasets, achieving significantly better results against all previous SAM-derived models. Heng Guo 0008, Tony C. W. Mok, Dazhou Guo, Ke Yan 0006, Le Lu 0001, Dakai Jin, Minfeng Xu |
AAAI | 7 |
| 2025 | Visual Evidence Prompting Mitigates Hallucinations in Large Vision-Language ModelsabstractWei Li, Zhen Huang, Houqiang Li, Le Lu, Yang Lu, Xinmei Tian, Xu Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Wei Li 0317, Zhen Huang 0007, Houqiang Li, Le Lu 0001, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye |
ACL (1) | 4 |
| 2025 | Interpret and Improve In-Context Learning via the Lens of Input-Label MappingsabstractChenghao Sun, Zhen Huang, Yonggang Zhang, Le Lu, Houqiang Li, Xinmei Tian, Xu Shen, Jieping Ye. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Zhen Huang 0007, Yonggang Zhang 0003, Le Lu 0001, Houqiang Li, Xinmei Tian 0001, Xu Shen 0001, Jieping Ye |
ACL (1) | 4 |
| 2025 | E-ViM3: Mamba-3D as Masked Autoencoders for Accurate and Data-Efficient Analysis of Medical Ultrasound VideosabstractUltrasound videos are an important form of clinical imaging data, and deep learning-based analysis can improve diagnostic accuracy and clinical efficiency. However, the scarcity of labeled data and the inherent challenges of video analysis have impeded the advancement of related methods. In this work, we introduce E-ViM3, a data-efficient Vision Mamba network that preserves the 3D structure of video data, enhancing long-range dependencies and inductive biases to better model spatial-temporal correlations. With our design of Enclosure Global Tokens (EG T), the model captures and aggregates global features more effectively than competing methods. We further employ a tailored masked video modeling approach for self-supervised pre-training to enhance its data efficiency, with the proposed Spatial- Temporal Chained (STC) masking strategy designed to adapt to different video scenarios. Experiments demonstrate that E-ViM3 achieves state-of-the-art performance on different tasks across four datasets of varying sizes: EchoNet-Dynamic, CAMUS, MICCAI-BUV, and WHBUS. Furthermore, our model attains competitive results even with limited labeled data, highlighting its potential impact on real-world clinical applications. Codes are available at https://github.com/HenryZhou19/E-ViM3. Jiaheng Zhou, Yanfeng Zhou, Wei Fang 0005, Yuxing Tang, Le Lu 0001, Ge Yang 0002 |
BIBM | 5 |
| 2025 | nnWNet: Rethinking the Use of Transformers in Biomedical Image Segmentation and Calling for a Unified Evaluation BenchmarkabstractSemantic segmentation is a crucial prerequisite in clinical applications and computer-aided diagnosis. With the development of deep neural networks, biomedical image segmentation has achieved remarkable success. Encoder-Decoder architectures that integrate convolutions and transformers are gaining attention for their potential to capture both global and local features. However, current designs face the contradiction that these two features cannot be continuously transmitted. In addition, some models lack a unified and standardized evaluation benchmark, leading to significant discrepancies in the experimental setup. In this study, we review and summarize these architectures and analyze their contradictions in design. We modify UNet and propose WNet to combine transformers and convolutions, addressing the transmission issue effectively. WNet captures long-range dependencies and local details simultaneously while ensuring their continuous transmission and multi-scale fusion. We integrate WNet into the nnUNet framework for unified benchmarking. Our model achieves state-of-the-art performance in biomedical image segmentation. Extensive experiments demonstrate their effectiveness on four 2D datasets (DRIVE, ISIC-2017, Kvasir-Seg, and CREMI) and four 3D datasets (Parse2022, AMOS22, BTCV, and ImageCAS). The code is available at https://github.com/yanfeng-zhou/nnWNet. Yanfeng Zhou, Lingrui Li, Le Lu 0001, Minfeng Xu |
CVPR | 3 |
| 2025 | Boosting Vision Semantic Density with Anatomy Normality Modeling for Medical Vision-Language Pre-Training
Zhongyi Shui, Sinuo Wang, Zeli Chen, Le Lu 0001, Xianghua Ye, Tingbo Liang, Ling Zhang 0002 |
ICCV | 7 |
| 2025 | Harmonyseg: Tubular Structure Segmentation With Deep-Shallow Feature Fusion and Growth-Suppression Balanced Loss
Yi Huangi, Wei Liu 0127, Vishal M. Patel, Le Lu 0001, Xu Han 0023, Dakai Jin, Ke Yan 0006 |
ICCV | 6 |
| 2025 | Bridging Local Inductive Bias and Long-Range Dependencies With Pixel-Mamba for End-To-End Whole Slide Image Analysis
Zhongwei Qiu, Hanqing Chao, Tiancheng Lin 0004, Wanxing Chang, Zijiang Yang 0009, Wenpei Jiao, Yunshuo Zhang, Yelin Yang, Yun Bian, Ke Yan 0006, Dakai Jin, Le Lu 0001 |
ICCV | 15 |
| 2025 | MaRS: A Fast Sampler for Mean Reverting Diffusion based on ODE and SDE SolversabstractIn applications of diffusion models, controllable generation is of practical significance, but is also challenging. Current methods for controllable generation primarily focus on modifying the score function of diffusion models, while Mean Reverting (MR) Diffusion directly modifies the structure of the stochastic differential equation (SDE), making the incorporation of image conditions simpler and more natural. However, current training-free fast samplers are not directly applicable to MR Diffusion. And thus MR Diffusion requires hundreds of NFEs (number of function evaluations) to obtain high-quality samples. In this paper, we propose a new algorithm named MaRS (MR Sampler) to reduce the sampling NFEs of MR Diffusion. We solve the reverse-time SDE and the probability flow ordinary differential equation (PF-ODE) associated with MR Diffusion, and derive semi-analytical solutions. The solutions consist of an analytical function and an integral parameterized by a neural network. Based on this solution, we can generate high-quality samples in fewer steps. Our approach does not require training and supports all mainstream parameterizations, including noise prediction, data prediction and velocity prediction. Extensive experiments demonstrate that MR Sampler maintains high sampling quality with a speedup of 10 to 20 times across ten different image restoration tasks. Our algorithm accelerates the sampling procedure of MR Diffusion, making it more practical in controllable generation. Ao Li 0004, Wei Fang 0005, Le Lu 0001, Ge Yang 0002, Minfeng Xu |
ICLR | 4 |
| 2025 | Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image UnderstandingabstractArtificial intelligence (AI) shows great potential in assisting radiologists to improve the efficiency and accuracy of medical image interpretation and diagnosis. However, a versatile AI model requires large-scale data and comprehensive annotations, which are often impractical in medical settings. Recent studies leverage radiology reports as a naturally high-quality supervision for medical images, using contrastive language-image pre-training (CLIP) to develop language-informed models for radiological image interpretation. Nonetheless, these approaches typically contrast entire images with reports, neglecting the local associations between imaging regions and report sentences, which may undermine model performance and interoperability. In this paper, we propose a fine-grained vision-language model (fVLM) for anatomy-level CT image interpretation. Specifically, we explicitly match anatomical regions of CT images with corresponding descriptions in radiology reports and perform contrastive pre-training for each anatomy individually. Fine-grained alignment, however, faces considerable false-negative challenges, mainly from the abundance of anatomy-level healthy samples and similarly diseased abnormalities, leading to ambiguous patient-level pairings. To tackle this issue, we propose identifying false negatives of both normal and abnormal samples and calibrating contrastive learning from patient-level to disease-aware pairing. We curated the largest CT dataset to date, comprising imaging and report data from 69,086 patients, and conducted a comprehensive evaluation of 54 major and important disease (including several most deadly cancers) diagnosis tasks across 15 main anatomies. Experimental results demonstrate the substantial potential of fVLM in versatile medical image interpretation. In the zero-shot classification task, we achieved an average AUC of 81.3% on 54 diagnosis tasks, surpassing CLIP and supervised methods by 12.9% and 8.0%, respectively. Additionally, on the publicly available CT-RATE and Rad-ChestCT benchmarks, our fVLM outperformed the current state-of-the-art methods with absolute AUC gains of 7.4% and 4.8%, respectively. Zhongyi Shui, Sinuo Wang, Ruizhe Guo, Le Lu 0001, Lin Yang 0002, Xianghua Ye, Tingbo Liang, Ling Zhang 0002 |
ICLR | 6 |
| 2025 | PLUS: Plug-and-Play Enhanced Liver Lesion Diagnosis Model on Non-contrast CT Scans
Jiacheng Hao, Xiaoming Zhang 0008, Wei Liu 0127, Xiaoli Yin, Yuan Gao 0017, Chunli Li, Ling Zhang 0002, Le Lu 0001, Xu Han 0023, Ke Yan 0006 |
MICCAI (15) | 8 |
| 2025 | Opportunistic Osteoporosis Diagnosis via Texture-Preserving Self-supervision, Mixture of Experts and Multi-task Integration
Heng Guo 0008, Le Lu 0001, Fan Yang 0081, Minfeng Xu, Ge Yang 0002 |
MICCAI (15) | 3 |
| 2025 | Lymph Node Metastasis Classification with Prototype-Guided Multiple Instance Aggregation and Heterogeneous Feature Fusion
Haoshen Li, Tashan Ai, Yirui Wang 0002, Zhanghexuan Ji, Qinji Yu, Le Lu 0001, Bin Dong 0001, Li Zhang 0047, Xianghua Ye, Kuaile Zhao, Dakai Jin |
MICCAI (1) | 6 |
| 2025 | Leveraging Semantic Asymmetry for Accurate Gross Tumor Volume Segmentation of Nasopharyngeal Carcinoma in Planning CT
Zeli Chen, Yanzhou Su, Tai Ma, Tony C. W. Mok, Yan-Jie Zhou, Yunhao Bai, Zhilin Zheng, Le Lu 0001, Yirui Wang 0002, Jia Ge, Senxiang Yan, Xianghua Ye, Dakai Jin |
MICCAI (2) | 10 |
| 2025 | Metastatic Lymph Node Station Classification in Esophageal Cancer via Prior-Guided Supervision and Station-Aware Mixture-of-Experts
Haoshen Li, Yirui Wang 0002, Qinji Yu, Ke Yan 0006, Dazhou Guo, Le Lu 0001, Bin Dong 0001, Li Zhang 0047, Xianghua Ye, Dakai Jin |
MICCAI (13) | 7 |
| 2025 | Anatomy-Aware Low-Dose CT Denoising via Pretrained Vision Models and Semantic-Guided Contrastive Learning
Zeli Chen, Zhiyun Song, Wei Fang 0005, Jiajin Zhang, Danyang Tu, Yuxing Tang, Minfeng Xu, Xianghua Ye, Le Lu 0001, Dakai Jin |
MICCAI (2) | 10 |
| 2025 | Lymphoma Prognosis with Lesion-Anatomy Context Fusion and Attention-Based Multi-lesion Aggregation
Jiajin Zhang, Liheng Qiu, Wei Liu 0127, Dakai Jin, Le Lu 0001, Shenmiao Yang, Ke Yan 0006 |
MICCAI (1) | 6 |
| 2025 | Deep Attention Learning for Pre-operative Lymph Node Metastasis Prediction in Pancreatic Cancer via Multi-object Relationship Modeling
Zhilin Zheng, Jiawen Yao, Le Lu 0001, Jianping Lu, Ling Zhang 0002, Chengwei Shao, Yun Bian |
Int. J. Comput. Vis. | 5 |
| 2025 | Correction: Deep Attention Learning for Pre-operative Lymph Node Metastasis Prediction in Pancreatic Cancer via Multi-object Relationship Modeling
Zhilin Zheng, Jiawen Yao, Le Lu 0001, Jianping Lu, Ling Zhang 0002, Chengwei Shao, Yun Bian |
Int. J. Comput. Vis. | 5 |
| 2025 | Bootstrapping Audio-Visual Video Segmentation by Strengthening Audio CuesabstractHow to effectively interact audio with vision has garnered considerable interest within the multi-modality research field. Recently, a novel audio-visual video segmentation (AVS) task has been proposed, aiming to segment the sounding objects in video frames under the guidance of audio cues. However, most existing AVS methods are hindered by a modality imbalance where the visual features tend to dominate those of the audio modality, due to a unidirectional and insufficient integration of audio cues. This imbalance skews the feature representation towards the visual aspect, impeding the learning of joint audio-visual representations and potentially causing segmentation inaccuracies. To address this issue, we propose AVSAC. Our approach features a Bidirectional Audio-Visual Decoder (BAVD) with integrated bidirectional bridges, enhancing audio cues and fostering continuous interplay between audio and visual modalities. This bidirectional interaction narrows the modality imbalance, facilitating more effective learning of integrated audio-visual representations. Additionally, we present a strategy for audio-visual frame-wise synchrony as fine-grained guidance of BAVD. This strategy enhances the share of auditory components in visual features, contributing to a more balanced audio-visual representation learning. Extensive experiments show that our method has state-of-the-art performance on several AVS public benchmarks. Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu, Le Lu 0001, Jieping Ye |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2025 | S4R: Separated Self-Supervised Spectral Regression for Hyperspectral Histopathology Image DiagnosisabstractHyperspectral images (HSIs) offer great potential for computational pathology. But, limited by the lack of adequate annotated data and the high spectral redundancy of HSIs, traditional supervised learning techniques are usually bottlenecked. To exploit the structural properties of HSIs and learn representations with good transferability, we propose Separated Self-Supervised Spectral Regression (S4R). Concretely, we find one spectral band can be represented by a linear combination of the remaining bands. Regressing the distribution of the linear coefficients learns the inherent properties of HSIs and pathological information about the tissue. Besides, reconstructing the missing band, especially the tissue boundaries makes the model learn pathology details that are critical to downstream tasks. Coupling these two pretext tasks makes the self-supervised model understand spectral structures of HSIs w.r.t. pathological semantics and spatial micro details. Furthermore, we design two brand-new architectures to avoid the interference of extraneous signal based on S4R: S4R-CLS and S4R-SEG for HSI classification and segmentation, respectively. Two downstream tasks are incorporated into a unified framework, which first encodes different bands from HSIs via a depthwise separable encoder, and then selectively aggregates band features to generate final predictions. In S4R-SEG, we propose to pick the best matching bands with the guidance of a classification paradigm. Extensive experiments show S4R performs much better than competitors on both tasks. Theoretical analysis and clinical discussion also indicate the great potential for further medical applications. The code and pre-trained checkpoints are available at https://github.com/DeepMed-Lab-ECNU/S4R. Yan Wang 0033, Xingran Xie, Benyan Zhang, Chunhua Zhou, Duowu Zou, Le Lu 0001, Qingli Li |
IEEE Trans. Image Process. | 7 |
| 2025 | Med-Query: Steerable Parsing of 9-DoF Medical Anatomies With Query EmbeddingabstractAutomatic parsing of human anatomies at the instance-level from 3D computed tomography (CT) is a prerequisite step for many clinical applications. The presence of pathologies, broken structures or limited field-of-view (FOV) can all make anatomy parsing algorithms vulnerable. In this work, we explore how to leverage and implement the successful detection-then-segmentation paradigm for 3D medical data, and propose a steerable, robust, and efficient computing framework for detection, identification, and segmentation of anatomies in CT scans. Considering the complicated shapes, sizes, and orientations of anatomies, without loss of generality, we present a nine degrees of freedom (9-DoF) pose estimation solution in full 3D space using a novel single-stage, non-hierarchical representation. Our whole framework is executed in a steerable manner where any anatomy of interest can be directly retrieved to further boost inference efficiency. We have validated our method on three medical imaging parsing tasks: ribs, spine, and abdominal organs. For rib parsing, CT scans have been annotated at the rib instance-level for quantitative evaluation, similarly for spine vertebrae and abdominal organs. Extensive experiments on 9-DoF box detection and rib instance segmentation demonstrate the high efficiency and effectiveness of our framework (with the identification rate of 97.0% and the segmentation Dice score of 90.9%), compared favorably against several strong baselines (e.g., CenterNet, FCOS, and nnU-Net). For spine parsing and abdominal multi-organ segmentation, our method achieves competitive results on par with state-of-the-art methods on the public CTSpine1K dataset and FLARE22 competition, respectively. Heng Guo 0008, Ke Yan 0006, Le Lu 0001, Minfeng Xu |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | DistAL: A Domain-Shift Active Learning Framework With Transferable Feature Learning for Lesion DetectionabstractDeep learning has demonstrated exceptional performance in medical image analysis, but its effectiveness degrades significantly when applied to different medical centers due to domain shifts. Lesion detection, a critical task in medical imaging, is particularly impacted by this challenge due to the diversity and complexity of lesions, which can arise from different organs, diseases, imaging devices, and other factors. While collecting data and labels from target domains is a feasible solution, annotating medical images is often tedious, expensive, and requires professionals. To address this problem, we combine active learning with domain-invariant feature learning. We propose a Domain-shift Active Learning (DistAL) framework, which includes a transferable feature learning algorithm and a hybrid sample selection strategy. Feature learning incorporates contrastive-consistency training to learn discriminative and domain-invariant features. The sample selection strategy is called RUDY, which jointly considers Representativeness, Uncertainty, and DiversitY. Its goal is to select samples from the unlabeled target domain for cost-effective annotation. It first selects representative samples to deal with domain shift, as well as uncertain ones to improve class separability, and then leverages K-means++ initialization to remove redundant candidates to achieve diversity. We evaluate our method for the task of lesion detection. By selecting only 1.7% samples from the target domain to annotate, DistAL achieves comparable performance to the method trained with all target labels. It outperforms other AL methods in five experiments on eight datasets collected from different hospitals, using different imaging protocols, annotation conventions, and etiologies. Fan Bai 0008, Dakai Jin, Xianghua Ye, Le Lu 0001, Ke Yan 0006, Max Q.-H. Meng |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Zig-RiR: Zigzag RWKV-in-RWKV for Efficient Medical Image SegmentationabstractMedical image segmentation has made significant strides with the development of basic models. Specifically, models that combine CNNs with transformers can successfully extract both local and global features. However, these models inherit the transformer's quadratic computational complexity, limiting their efficiency. Inspired by the recent Receptance Weighted Key Value (RWKV) model, which achieves linear complexity for long-distance modeling, we explore its potential for medical image segmentation. While directly applying vision-RWKV yields suboptimal results due to insufficient local feature exploration and disrupted spatial continuity, we propose a novel nested structure, Zigzag RWKV-in-RWKV (Zig-RiR), to address these issues. It consists of Outer and Inner RWKV blocks to adeptly capture both global and local features without disrupting spatial continuity. We treat local patches as "visual sentences" and use the Outer Zig-RWKV to explore global information. Then, we decompose each sentence into sub-patches ("visual words") and use the Inner Zig-RWKV to further explore local information among words, at negligible computational cost. We also introduce a Zigzag-WKV attention mechanism to ensure spatial continuity during token scanning. By aggregating visual word and sentence features, our Zig-RiR can effectively explore both global and local information while preserving spatial continuity. Experiments on four medical image segmentation datasets of both 2D and 3D modalities demonstrate the superior accuracy and efficiency of our method, outperforming the state-of-the-art method 14.4 times in speed and reducing GPU memory usage by 89.5% when testing on ${1024} \times {1024}$ high-resolution medical images. Our code is available at https://github.com/txchen-USTC/Zig-RiR. Zhentao Tan, Qi Chu 0001, Nenghai Yu, Le Lu 0001 |
IEEE Trans. Medical Imaging | 10 |
| 2025 | RemixFormer++: A Multi-Modal Transformer Model for Precision Skin Tumor Differential Diagnosis With Memory-Efficient AttentionabstractDiagnosing malignant skin tumors accurately at an early stage can be challenging due to ambiguous and even confusing visual characteristics displayed by various categories of skin tumors. To improve diagnosis precision, all available clinical data from multiple sources, particularly clinical images, dermoscopy images, and medical history, could be considered. Aligning with clinical practice, we propose a novel Transformer model, named RemixFormer++ that consists of a clinical image branch, a dermoscopy image branch, and a metadata branch. Given the unique characteristics inherent in clinical and dermoscopy images, specialized attention strategies are adopted for each type. Clinical images are processed through a top-down architecture, capturing both localized lesion details and global contextual information. Conversely, dermoscopy images undergo a bottom-up processing with two-level hierarchical encoders, designed to pinpoint fine-grained structural and textural features. A dedicated metadata branch seamlessly integrates non-visual information by encoding relevant patient data. Fusing features from three branches substantially boosts disease classification accuracy. RemixFormer++ demonstrates exceptional performance on four single-modality datasets (PAD-UFES-20, ISIC 2017/2018/2019). Compared with the previous best method using a public multi-modal Derm7pt dataset, we achieved an absolute 5.3% increase in averaged F1 and 1.2% in accuracy for the classification of five skin tumors. Furthermore, using a large-scale in-house dataset of 10,351 patients with the twelve most common skin tumors, our method obtained an overall classification accuracy of 92.6%. These promising results, on par or better with the performance of 191 dermatologists through a comprehensive reader study, evidently imply the potential clinical usability of our method. Kai Huang 0008, Lianzhen Zhong, Yuan Gao 0017, Wei Liu 0127, Yanjie Zhou, Wenchao Guo, Yuanqiang Zou, Yuping Duan, Le Lu 0001, Yu Wang 0108 |
IEEE Trans. Medical Imaging | 12 |
| 2025 | A Colorectal Coordinate-Driven Method for Colorectum and Colorectal Cancer Segmentation in Conventional CT ScansabstractAutomated colorectal cancer (CRC) segmentation in medical imaging is the key to achieving automation of CRC detection, staging, and treatment response monitoring. Compared with magnetic resonance imaging (MRI) and computed tomography colonography (CTC), conventional computed tomography (CT) has enormous potential because of its broad implementation, superiority for the hollow viscera (colon), and convenience without needing bowel preparation. However, the segmentation of CRC in conventional CT is more challenging due to the difficulties presenting with the unprepared bowel, such as distinguishing the colorectum from other structures with similar appearance and distinguishing the CRC from the contents of the colorectum. To tackle these challenges, we introduce DeepCRC-SL, the first automated segmentation algorithm for CRC and colorectum in conventional contrast-enhanced CT scans. We propose a topology-aware deep learning-based approach, which builds a novel 1-D colorectal coordinate system and encodes each voxel of the colorectum with a relative position along the coordinate system. We then induce an auxiliary regression task to predict the colorectal coordinate value of each voxel, aiming to integrate global topology into the segmentation network and thus improve the colorectum's continuity. Self-attention layers are utilized to capture global contexts for the coordinate regression task and enhance the ability to differentiate CRC and colorectum tissues. Moreover, a coordinate-driven self-learning (SL) strategy is introduced to leverage a large amount of unlabeled data to improve segmentation performance. We validate the proposed approach on a dataset including 227 labeled and 585 unlabeled CRC cases by fivefold cross-validation. Experimental results demonstrate that our method outperforms some recent related segmentation methods and achieves the segmentation accuracy in DSC for CRC of 0.669 and colorectum of 0.892, reaching to the performance (at 0.639 and 0.890, respectively) of a medical resident with two years of specialized CRC imaging fellowship. Yingda Xia, Suyun Li, Jiawen Yao, Dakai Jin, Yanting Liang, Jiatai Lin, Bingchao Zhao, Chu Han, Le Lu 0001, Ling Zhang 0002, Zaiyi Liu, Xin Chen 0058 |
IEEE Trans. Neural Networks Learn. Syst. | 11 |
| 2024 | Bootstrapping Chest CT Image Understanding by Distilling Knowledge from X-Ray Expert ModelsabstractRadiologists highly desire fully automated versatile AI for medical imaging interpretation. However, the lack of extensively annotated large-scale multi-disease datasets has hindered the achievement of this goal. In this paper, we explore the feasibility of leveraging language as a natu-rally high-quality supervision for chest CT imaging. In light of the limited availability of image-report pairs, we boot-strap the understanding of 3D chest CT images by distilling chest-related diagnostic knowledge from an extensively pre-trained 2D X-ray expert model. Specifically, we propose a language-guided retrieval method to match each 3D CT image with its semantically closest 2D X-ray image, and perform pair-wise and semantic relation knowledge distillation. Subsequently, we use contrastive learning to align images and reports within the same patient while distin-guishing them from the other patients. However, the challenge arises when patients have similar semantic diagnoses, such as healthy patients, potentially confusing if treated as negatives. We introduce a robust contrastive learning that identifies and corrects these false negatives. We train our model with over 12K pairs of chest CT images and radiology reports. Extensive experiments across multiple scenarios, including zero-shot learning, report generation, and fine-tuning processes, demonstrate the model's feasibility in interpreting chest CT images. Yingda Xia, Tony C. W. Mok, Xianghua Ye, Le Lu 0001, Yuxing Tang, Ling Zhang 0002 |
CVPR | 7 |
| 2024 | CycleINR: Cycle Implicit Neural Representation for Arbitrary-Scale Volumetric Super-Resolution of Medical DataabstractIn the realm of medical 3D data, such as CT and MRI images, prevalent anisotropic resolution is characterized by high intra-slice but diminished inter-slice resolution. The lowered resolution between adjacent slices poses challenges, hindering optimal viewing experiences and impeding the development of robust downstream analysis algorithms. Various volumetric super-resolution algorithms aim to surmount these challenges, enhancing inter-slice resolution and overall 3D medical imaging quality. However, existing approaches confront inherent challenges: 1) often tailored to specific upsampling factors, lacking flexibility for diverse clinical scenarios; 2) newly generated slices frequently suffer from over-smoothing, degrading fine details, and leading to inter-slice inconsistency. In response, this study presents CycleINR, a novel enhanced Implicit Neural Representation model for 3D medical data volumetric super-resolution. Leveraging the continuity of the learned implicit function, the CycleINR model can achieve results with arbitrary up-sampling rates, eliminating the need for separate training. Additionally, we enhance the grid sampling in CycleINR with a local attention mechanism and mitigate over-smoothing by integrating cycleconsistent loss. We introduce a new metric, Slice-wise Noise Level Inconsistency (SNLI), to quantitatively assess inter-slice noise level inconsistency. The effectiveness of our approach is demonstrated through image quality evaluations on an in-house dataset and a downstream task analysis on the Medical Segmentation Decathlon liver tumor dataset. Wei Fang 0005, Yuxing Tang, Heng Guo 0008, Mingze Yuan, Tony C. W. Mok, Ke Yan 0006, Jiawen Yao, Xin Chen 0058, Zaiyi Liu, Le Lu 0001, Ling Zhang 0002, Minfeng Xu |
CVPR | 10 |
| 2024 | Modality-Agnostic Structural Image Representation Learning for Deformable Multi-Modality Medical Image RegistrationabstractEstablishing dense anatomical correspondence across distinct imaging modalities is a foundational yet challenging procedure for numerous medical image analysis studies and image-guided radiotherapy. Existing multimodality image registration algorithms rely on statistical-based similarity measures or local structural image representations. However, the former is sensitive to locally varying noise, while the latter is not discriminative enough to cope with complex anatomical structures in multimodal scans, causing ambiguity in determining the anatomical correspon-dence across scans with different modalities. In this paper, we propose a modality-agnostic structural representation learning method, which leverages Deep Neighbour-hood Self-similarity (DNS) and anatomy-aware contrastive learning to learn discriminative and contrast-invariance deep structural image representations (DSIR) without the need for anatomical delineations or pre-aligned training images. We evaluate our method on multiphase CT, abdomen MR-CT, and brain MR T1w-T2w registration. Comprehensive results demonstrate that our method is superior to the conventional local structural representation and statistical-based similarity measures in terms of discriminability and accuracy. Tony C. W. Mok, Yunhao Bai, Wei Liu 0127, Yan-Jie Zhou, Ke Yan 0006, Dakai Jin, Xiaoli Yin, Le Lu 0001, Ling Zhang 0002 |
CVPR | 11 |
| 2024 | Effective Lymph Nodes Detection in CT Scans Using Location Debiased Query Selection and Contrastive Query Representation in Transformer
Qinji Yu, Yirui Wang 0002, Ke Yan 0006, Haoshen Li, Dazhou Guo, Li Zhang 0047, Na Shen, Le Lu 0001, Xianghua Ye, Dakai Jin |
ECCV (42) | 10 |
| 2024 | Interpretable Composition Attribution Enhancement for Visio-linguistic Compositional UnderstandingabstractContrastively trained vision-language models such as CLIP have achieved remarkable progress in vision and language representation learning.Despite the promising progress, their proficiency in compositional reasoning over attributes and relations (e.g., distinguishing between "the car is underneath the person" and "the person is underneath the car") remains notably inadequate.We investigate the cause for this deficient behavior is the composition attribution issue, where the attribution scores (e.g., attention scores or GradCAM scores) for relations (e.g., underneath) or attributes (e.g., red) in the text are substantially lower than those for object terms.In this work, we show such issue is mitigated via a novel framework called CAE (Composition Attribution Enhancement).This generic framework incorporates various interpretable attribution methods to encourage the model to pay greater attention to composition words denoting relationships and attributes within the text.Detailed analysis shows that our approach enables the models to adjust and rectify the attribution of the texts.Extensive experiments across seven benchmarks reveal that our framework significantly enhances the ability to discern intricate details and construct more sophisticated interpretations of combined visual and linguistic elements. Wei Li 0317, Zhen Huang 0007, Xinmei Tian 0001, Le Lu 0001, Houqiang Li, Xu Shen 0001, Jieping Ye |
EMNLP | 4 |
| 2024 | Boosting Vanilla Lightweight Vision Transformers via Re-parameterizationabstractLarge-scale Vision Transformers have achieved promising performance on downstream tasks through feature pre-training. However, the performance of vanilla lightweight Vision Transformers (ViTs) is still far from satisfactory compared to that of recent lightweight CNNs or hybrid networks. In this paper, we aim to unlock the potential of vanilla lightweight ViTs by exploring the adaptation of the widely-used re-parameterization technology to ViTs for improving learning ability during training without increasing the inference cost. The main challenge comes from the fact that CNNs perfectly complement with re-parameterization over convolution and batch normalization, while vanilla Transformer architectures are mainly comprised of linear and layer normalization layers. We propose to incorporate the nonlinear ensemble into linear layers by expanding the depth of the linear layers with batch normalization and fusing multiple linear features with hierarchical representation ability through a pyramid structure. We also discover and solve a new transformer-specific distribution rectification problem caused by multi-branch re-parameterization. Finally, we propose our Two-Dimensional Re-parameterized Linear module (TDRL) for ViTs. Under the popular self-supervised pre-training and supervised fine-tuning strategy, our TDRL can be used in these two stages to enhance both generic and task-specific representation. Experiments demonstrate that our proposed method not only boosts the performance of vanilla Vit-Tiny on various vision tasks to new state-of-the-art (SOTA) but also shows promising generality ability on other networks. Code will be available. Zhentao Tan, Qi Chu 0001, Le Lu 0001, Nenghai Yu, Jieping Ye |
ICLR | 5 |
| 2024 | From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint TuningabstractLarge Language Models (LLMs) tend to prioritize adherence to user prompts over providing veracious responses, leading to the sycophancy issue. When challenged by users, LLMs tend to admit mistakes and provide inaccurate responses even if they initially provided the correct answer. Recent works propose to employ supervised fine-tuning (SFT) to mitigate the sycophancy issue, while it typically leads to the degeneration of LLMs' general capability. To address the challenge, we propose a novel supervised pinpoint tuning (SPT), where the region-of-interest modules are tuned for a given objective. Specifically, SPT first reveals and verifies a small percentage (<5%) of the basic modules, which significantly affect a particular behavior of LLMs. i.e., sycophancy. Subsequently, SPT merely fine-tunes these identified modules while freezing the rest. To verify the effectiveness of the proposed SPT, we conduct comprehensive experiments, demonstrating that SPT significantly mitigates the sycophancy issue of LLMs (even better than SFT). Moreover, SPT introduces limited or even no side effects on the general capability of LLMs. Our results shed light on how to precisely, effectively, and efficiently explain and improve the targeted ability of LLMs. Wei Chen 0005, Zhen Huang 0007, Liang Xie 0003, Binbin Lin 0001, Houqiang Li, Le Lu 0001, Xinmei Tian 0001, Deng Cai 0001, Yonggang Zhang 0003, Wenxiao Wang 0001, Xu Shen 0001, Jieping Ye |
ICML | 6 |
| 2024 | Cross-Phase Mutual Learning Framework for Pulmonary Embolism Identification on Non-contrast CT Scans
Bizhe Bai, Yan-Jie Zhou, Yujian Hu, Tony C. W. Mok, Yilang Xiang, Le Lu 0001, Hongkun Zhang, Minfeng Xu |
MICCAI (1) | 6 |
| 2024 | LIDIA: Precise Liver Tumor Diagnosis on Multi-Phase Contrast-Enhanced CT via Iterative Fusion and Asymmetric Contrastive Learning
Wei Liu 0127, Xiaoming Zhang 0008, Xiaoli Yin, Xu Han 0023, Chunli Li, Yuan Gao 0017, Le Lu 0001, Ling Zhang 0002, Lei Zhang 0006, Ke Yan 0006 |
MICCAI (9) | 9 |
| 2024 | Semi-supervised Lymph Node Metastasis Classification with Pathology-Guided Label Sharpening and Two-Streamed Multi-scale Fusion
Haoshen Li, Yirui Wang 0002, Dazhou Guo, Qinji Yu, Ke Yan 0006, Le Lu 0001, Xianghua Ye, Li Zhang 0047, Dakai Jin |
MICCAI (11) | 7 |
| 2024 | Improved Esophageal Varices Assessment from Non-contrast CT Scans
Chunli Li, Xiaoming Zhang 0008, Yuan Gao 0017, Xiaoli Yin, Le Lu 0001, Ling Zhang 0002, Ke Yan 0006 |
MICCAI (5) | 5 |
| 2024 | Slice-Consistent Lymph Nodes Detection Transformer in CT Scans via Cross-Slice Query Contrastive Learning
Qinji Yu, Yirui Wang 0002, Ke Yan 0006, Le Lu 0001, Na Shen, Xianghua Ye, Dakai Jin |
MICCAI (5) | 4 |
| 2024 | IHCSurv: Effective Immunohistochemistry Priors for Cancer Survival Analysis in Gigapixel Multi-stain Whole Slide Images
Yejia Zhang, Hanqing Chao, Zhongwei Qiu, Nishchal Sapkota, Pengfei Gu, Danny Ziyi Chen, Le Lu 0001, Ke Yan 0006, Dakai Jin, Yun Bian |
MICCAI (4) | 9 |
| 2024 | A Curvature-Guided Coarse-to-Fine Framework for Enhanced Whole Brain Segmentation
Fenqiang Zhao, Yuxing Tang, Le Lu 0001, Ling Zhang 0002 |
MICCAI (9) | 3 |
| 2024 | Low-Rank Continual Pyramid Vision Transformer: Incrementally Segment Whole-Body Organs in CT with Light-Weighted Adaptation
Vince Zhu, Zhanghexuan Ji, Dazhou Guo, Puyang Wang, Yingda Xia, Le Lu 0001, Xianghua Ye, Wei Zhu 0015, Dakai Jin |
MICCAI (8) | 6 |
| 2024 | TransUNet: Rethinking the U-Net architecture design for medical image segmentation through the lens of transformersabstractMedical image segmentation is crucial for healthcare, yet convolution-based methods like U-Net face limitations in modeling long-range dependencies. To address this, Transformers designed for sequence-to-sequence predictions have been integrated into medical image segmentation. However, a comprehensive understanding of Transformers' self-attention in U-Net components is lacking. TransUNet, first introduced in 2021, is widely recognized as one of the first models to integrate Transformer into medical image analysis. In this study, we present the versatile framework of TransUNet that encapsulates Transformers' self-attention into two key modules: (1) a Transformer encoder tokenizing image patches from a convolution neural network (CNN) feature map, facilitating global context extraction, and (2) a Transformer decoder refining candidate regions through cross-attention between proposals and U-Net features. These modules can be flexibly inserted into the U-Net backbone, resulting in three configurations: Encoder-only, Decoder-only, and Encoder+Decoder. TransUNet provides a library encompassing both 2D and 3D implementations, enabling users to easily tailor the chosen architecture. Our findings highlight the encoder's efficacy in modeling interactions among multiple abdominal organs and the decoder's strength in handling small targets like tumors. It excels in diverse medical applications, such as multi-organ segmentation, pancreatic tumor segmentation, and hepatic vessel segmentation. Notably, our TransUNet achieves a significant average Dice improvement of 1.06% and 4.30% for multi-organ segmentation and pancreatic tumor segmentation, respectively, when compared to the highly competitive nn-UNet, and surpasses the top-1 solution in the BrasTS2021 challenge. 2D/3D Code and models are available at https://github.com/Beckschen/TransUNet and https://github.com/Beckschen/TransUNet-3D, respectively. Jieneng Chen, Jieru Mei, Xianhang Li, Yongyi Lu, Qihang Yu, Qingyue Wei, Xiangde Luo, Yutong Xie 0001, Ehsan Adeli-Mosabbeb, Yan Wang 0033, Matthew P. Lungren, Shaoting Zhang 0001, Lei Xing 0001, Le Lu 0001, Alan L. Yuille, Yuyin Zhou |
Medical Image Anal. | 14 |
| 2024 | Exploring the Application of Large-Scale Pre-Trained Models on Adverse Weather RemovalabstractImage restoration under adverse weather conditions (e.g., rain, snow, and haze) is a fundamental computer vision problem that has important implications for various downstream applications. Distinct from early methods that are specially designed for specific types of weather, recent works tend to simultaneously remove various adverse weather effects based on either spatial feature representation learning or semantic information embedding. Inspired by various successful applications incorporating large-scale pre-trained models (e.g., CLIP), in this paper, we explore their potential benefits for leveraging large-scale pre-trained models in this task based on both spatial feature representation learning and semantic information embedding aspects: 1) spatial feature representation learning, we design a Spatially Adaptive Residual (SAR) encoder to adaptively extract degraded areas. To facilitate training of this model, we propose a Soft Residual Distillation (CLIP-SRD) strategy to transfer spatial knowledge from CLIP between clean and adverse weather images; 2) semantic information embedding, we propose a CLIP Weather Prior (CWP) embedding module to enable the network to adaptively respond to different weather conditions. This module integrates the sample-specific weather priors extracted by the CLIP image encoder with the distribution-specific information (as learned by a set of parameters) and embeds these elements using a cross-attention mechanism. Extensive experiments demonstrate that our proposed method can achieve state-of-the-art performance under various and severe adverse weather conditions. The code will be made available. Zhentao Tan, Qiankun Liu 0001, Qi Chu 0001, Le Lu 0001, Jieping Ye, Nenghai Yu |
IEEE Trans. Image Process. | 5 |
| 2024 | LViT: Language Meets Vision Transformer in Medical Image SegmentationabstractDeep learning has been widely used in medical image segmentation and other aspects. However, the performance of existing medical image segmentation models has been limited by the challenge of obtaining sufficient high-quality labeled data due to the prohibitive data annotation cost. To alleviate this limitation, we propose a new text-augmented medical image segmentation model LViT (Language meets Vision Transformer). In our LViT model, medical text annotation is incorporated to compensate for the quality deficiency in image data. In addition, the text information can guide to generate pseudo labels of improved quality in the semi-supervised learning. We also propose an Exponential Pseudo label Iteration mechanism (EPI) to help the Pixel-Level Attention Module (PLAM) preserve local image features in semi-supervised LViT setting. In our model, LV (Language-Vision) loss is designed to supervise the training of unlabeled images using text information directly. For evaluation, we construct three multimodal medical segmentation datasets (image + text) containing X-rays and CT images. Experimental results show that our proposed LViT has superior segmentation performance in both fully-supervised and semi-supervised setting. The code and datasets are available at https://github.com/HUANGLIZI/LViT. Qingde Li, Puyang Wang, Dazhou Guo, Le Lu 0001, Dakai Jin, You Zhang 0003, Qingqi Hong |
IEEE Trans. Medical Imaging | 6 |
| 2024 | Accurate Airway Tree Segmentation in CT Scans via Anatomy-Aware Multi-Class Segmentation and Topology-Guided Iterative LearningabstractIntrathoracic airway segmentation in computed tomography is a prerequisite for various respiratory disease analyses such as chronic obstructive pulmonary disease, asthma and lung cancer. Due to the low imaging contrast and noises execrated at peripheral branches, the topological-complexity and the intra-class imbalance of airway tree, it remains challenging for deep learning-based methods to segment the complete airway tree (on extracting deeper branches). Unlike other organs with simpler shapes or topology, the airway's complex tree structure imposes an unbearable burden to generate the "ground truth" label (up to 7 or 3 hours of manual or semi-automatic annotation per case). Most of the existing airway datasets are incompletely labeled/annotated, thus limiting the completeness of computer-segmented airway. In this paper, we propose a new anatomy-aware multi-class airway segmentation method enhanced by topology-guided iterative self-learning. Based on the natural airway anatomy, we formulate a simple yet highly effective anatomy-aware multi-class segmentation task to intuitively handle the severe intra-class imbalance of the airway. To solve the incomplete labeling issue, we propose a tailored iterative self-learning scheme to segment toward the complete airway tree. For generating pseudo-labels to achieve higher sensitivity (while retaining similar specificity), we introduce a novel breakage attention map and design a topology-guided pseudo-label refinement method by iteratively connecting breaking branches commonly existed from initial pseudo-labels. Extensive experiments have been conducted on four datasets including two public challenges. The proposed method achieves the top performance in both EXACT'09 challenge using average score and ATM'22 challenge on weighted average score. In a public BAS dataset and a private lung cancer dataset, our method significantly improves previous leading approaches by extracting at least (absolute) 6.1% more detected tree length and 5.2% more tree branches, while maintaining comparable precision. Puyang Wang, Dazhou Guo, Haogang Yu, Jia Ge, Yun Gu, Le Lu 0001, Xianghua Ye, Dakai Jin |
IEEE Trans. Medical Imaging | 9 |
| 2023 | Devil is in the Queries: Advancing Mask Transformers for Real-world Medical Image Segmentation and Out-of-Distribution LocalizationabstractReal-world medical image segmentation has tremendous long-tailed complexity of objects, among which tail conditions correlate with relatively rare diseases and are clinically significant. A trustworthy medical AI algorithm should demonstrate its effectiveness on tail conditions to avoid clinically dangerous damage in these out-of-distribution (OOD) cases. In this paper, we adopt the concept of object queries in Mask Transformers to formulate semantic segmentation as a soft cluster assignment. The queries fit the feature-level cluster centers of inliers during training. Therefore, when performing inference on a medical image in real-world scenarios, the similarity between pixels and the queries detects and localizes OOD regions. We term this OOD localization as MaxQuery. Furthermore, the foregrounds of real-world medical images, whether OOD objects or inliers, are lesions. The difference between them is less than that between the foreground and background, possibly misleading the object queries to focus redundantly on the background. Thus, we propose a query-distribution (QD) loss to enforce clear boundaries between segmentation targets and other regions at the query level, improving the inlier segmentation and OOD indication. Our proposed framework is tested on two real-world segmentation tasks, i.e., segmentation of pancreatic and liver tumors, outperforming previous state-of-the-art algorithms by an average of 7.39% on AUROC, 14.69% on AUPR, and 13.79% on FPR95 for OOD localization. On the other hand, our framework improves the performance of inlier segmentation by an average of 5.27% DSC when compared with the leading baseline nnUNet. Mingze Yuan, Yingda Xia, Hexin Dong, Zifan Chen, Jiawen Yao, Mingyan Qiu, Ke Yan 0006, Xiaoli Yin, Xin Chen 0058, Zaiyi Liu, Bin Dong 0001, Jingren Zhou 0001, Le Lu 0001, Ling Zhang 0002, Li Zhang 0047 |
CVPR | 14 |
| 2023 | CancerUniT: Towards a Single Unified Model for Effective Detection, Segmentation, and Diagnosis of Eight Major Cancers Using a Large Collection of CT ScansabstractHuman readers or radiologists routinely perform full-body multi-organ multi-disease detection and diagnosis in clinical practice, while most medical AI systems are built to focus on single organs with a narrow list of a few diseases. This might severely limit AI’s clinical adoption. A certain number of AI models need to be assembled nontrivially to match the diagnostic process of a human reading a CT scan. In this paper, we construct a Unified Tumor Transformer (CancerUniT) model to jointly detect tumor existence & location and diagnose tumor characteristics for eight major cancers in CT scans. CancerUniT is a query-based Mask Transformer model with the output of multi-tumor prediction. We decouple the object queries into organ queries, tumor detection queries and tumor diagnosis queries, and further establish hierarchical relationships among the three groups. This clinically-inspired architecture effectively assists inter- and intra-organ representation learning of tumors and facilitates the resolution of these complex, anatomically related multi-organ cancer image reading tasks. CancerUniT is trained end-to-end using a curated large-scale CT images of 10,042 patients including eight major types of cancers and occurring non-cancer tumors (all are pathology-confirmed with 3D tumor masks annotated by radiologists). On the test set of 631 patients, CancerUniT has demonstrated strong performance under a set of clinically relevant evaluation metrics, substantially outperforming both multi-disease methods and an assembly of eight single-organ expert models in tumor detection, segmentation, and diagnosis. This moves one step closer towards a universal high performance cancer screening tool. Jieneng Chen, Yingda Xia, Jiawen Yao, Ke Yan 0006, Le Lu 0001, Fakai Wang, Bo Zhou 0009, Mingyan Qiu, Qihang Yu, Mingze Yuan, Wei Fang 0005, Yuxing Tang, Minfeng Xu, Xianghua Ye, Xiaoli Yin, Xin Chen 0058, Jingren Zhou 0001, Alan L. Yuille, Zaiyi Liu, Ling Zhang 0002 |
ICCV | 6 |
| 2023 | Continual Segment: Towards a Single, Unified and Non-forgetting Continual Segmentation Model of 143 Whole-body Organs in CT ScansabstractDeep learning empowers the mainstream medical image segmentation methods. Nevertheless, current deep segmentation approaches are not capable of efficiently and effectively adapting and updating the trained models when new segmentation classes are incrementally added. In the real clinical environment, it can be preferred that segmentation models could be dynamically extended to segment new organs/tumors without the (re-)access to previous training datasets due to obstacles of patient privacy and data storage. This process can be viewed as a continual semantic segmentation (CSS) problem, being understudied for multi-organ segmentation. In this work, we propose a new architectural CSS learning framework to learn a single deep segmentation model for segmenting a total of 143 whole-body organs. Using the encoder/decoder network structure, we demonstrate that a continually trained then frozen encoder coupled with incrementally-added decoders can extract sufficiently representative image features for new classes to be subsequently and validly segmented, while avoiding the catastrophic forgetting in CSS. To maintain a single network model complexity, each decoder is progressively pruned using neural architecture search and teacher-student based knowledge distillation. Finally, we propose a body-part and anomaly-aware output merging module to combine organ predictions originating from different decoders and incorporate both healthy and pathological organs appearing in different datasets. Trained and validated on 3D CT scans of 2500+ patients from four datasets, our single network can segment a total of 143 whole-body organs with very high accuracy, closely reaching the upper bound performance level by training four separate segmentation models (i.e., one model per dataset/task). Zhanghexuan Ji, Dazhou Guo, Puyang Wang, Ke Yan 0006, Le Lu 0001, Minfeng Xu, Jia Ge, Mingchen Gao, Xianghua Ye, Dakai Jin |
ICCV | 5 |
| 2023 | Anatomical Invariance Modeling and Semantic Alignment for Self-supervised Learning in 3D Medical Image AnalysisabstractSelf-supervised learning (SSL) has recently achieved promising performance for 3D medical image analysis tasks. Most current methods follow existing SSL paradigm originally designed for photographic or natural images, which cannot explicitly and thoroughly exploit the intrinsic similar anatomical structures across varying medical images. This may in fact degrade the quality of learned deep representations by maximizing the similarity among features containing spatial misalignment information and different anatomical semantics. In this work, we propose a new self-supervised learning framework, namely Alice, that explicitly fulfills Anatomical invariance modeling and semantic alignment via elaborately combining discriminative and generative objectives. Alice introduces a new contrastive learning strategy which encourages the similarity between views that are diversely mined but with consistent high-level semantics, in order to learn invariant anatomical features. Moreover, we design a conditional anatomical feature alignment module to complement corrupted embeddings with globally matched semantics and inter-patch topology information, conditioned by the distribution of local image content, which permits to create better contrastive pairs. Our extensive quantitative experiments on three 3D medical image analysis tasks demonstrate and validate the performance superiority of Alice, surpassing the previous best SSL counterpart methods and showing promising ability for united representation learning. Codes are available at https://github.com/alibaba-damo-academy/Alice. Yankai Jiang 0001, Heng Guo 0008, Ke Yan 0006, Le Lu 0001, Minfeng Xu |
ICCV | 6 |
| 2023 | SLPT: Selective Labeling Meets Prompt Tuning on Label-Limited Lesion Segmentation
Fan Bai 0008, Ke Yan 0006, Xiaoli Yin, Jingren Zhou 0001, Le Lu 0001, Max Q.-H. Meng |
MICCAI (2) | 8 |
| 2023 | Improved Prognostic Prediction of Pancreatic Cancer Using Multi-phase CT by Integrating Neural Distance and Texture-Aware Transformer
Hexin Dong, Jiawen Yao, Yuxing Tang, Mingze Yuan, Yingda Xia, Jingren Zhou 0001, Bin Dong 0001, Le Lu 0001, Zaiyi Liu, Li Zhang 0047, Ling Zhang 0002 |
MICCAI (5) | 10 |
| 2023 | SAMConvex: Fast Discrete Optimization for CT Registration Using Self-supervised Anatomical Embedding and Correlation Pyramid
Lin Tian 0001, Tony C. W. Mok, Puyang Wang, Jia Ge, Jingren Zhou 0001, Le Lu 0001, Xianghua Ye, Ke Yan 0006, Dakai Jin |
MICCAI (10) | 8 |
| 2023 | Liver Tumor Screening and Diagnosis in CT with Pixel-Lesion-Patient Network
Ke Yan 0006, Xiaoli Yin, Yingda Xia, Fakai Wang, Yuan Gao 0017, Jiawen Yao, Chunli Li, Jingren Zhou 0001, Ling Zhang 0002, Le Lu 0001 |
MICCAI (5) | 12 |
| 2023 | Cluster-Induced Mask Transformers for Effective Opportunistic Gastric Cancer Screening on Non-contrast CT Scans
Mingze Yuan, Yingda Xia, Xin Chen 0058, Jiawen Yao, Mingyan Qiu, Hexin Dong, Jingren Zhou 0001, Bin Dong 0001, Le Lu 0001, Li Zhang 0047, Zaiyi Liu, Ling Zhang 0002 |
MICCAI (5) | 10 |
| 2023 | Parse and Recall: Towards Accurate Lung Nodule Malignancy Prediction Like Radiologists
Xianghua Ye, Yuxing Tang, Minfeng Xu, Jianfei Guo, Xin Chen 0058, Zaiyi Liu, Jingren Zhou 0001, Le Lu 0001, Ling Zhang 0002 |
MICCAI (5) | 10 |
| 2023 | A Novel Multi-task Model Imitating Dermatologists for Accurate Differential Diagnosis of Skin Diseases in Clinical Images
Yan-Jie Zhou, Wei Liu 0127, Yuan Gao 0017, Le Lu 0001, Yuping Duan, Na Jin, Xiaoyong Man, Yu Wang 0108 |
MICCAI (6) | 5 |
| 2023 | Lumbar Bone Mineral Density Estimation From Chest X-Ray Images: Anatomy-Aware Attentive Multi-ROI ModelingabstractOsteoporosis is a common chronic metabolic bone disease often under-diagnosed and under-treated due to the limited access to bone mineral density (BMD) examinations, e.g., via Dual-energy X-ray Absorptiometry (DXA). This paper proposes a method to predict BMD from Chest X-ray (CXR), one of the most commonly accessible and low-cost medical imaging examinations. The proposed method first automatically detects Regions of Interest (ROIs) of local CXR bone structures. Then a multi-ROI deep model with transformer encoder is developed to exploit both local and global information in the chest X-ray image for accurate BMD estimation. The proposed method is evaluated on 13719 CXR patient cases with ground truth BMD measured by the gold standard DXA. The model predicted BMD has a strong correlation with the ground truth (Pearson correlation coefficient 0.894 on lumbar 1). When applied in osteoporosis screening, it achieves a high classification performance (average AUC of 0.968). As the first effort of using CXR scans to predict the BMD, the proposed algorithm holds strong potential to promote early osteoporosis screening and public health. Fakai Wang, Le Lu 0001, Jing Xiao 0006, Min Wu 0001, Chang-Fu Kuo, Shun Miao |
IEEE Trans. Medical Imaging | 3 |
| 2022 | Deep Implicit Statistical Shape Models for 3D Medical Image Delineationabstract3D delineation of anatomical structures is a cardinal goal in medical imaging analysis. Prior to deep learning, statistical shape models (SSMs) that imposed anatomical constraints and produced high quality surfaces were a core technology. Today’s fully-convolutional networks (FCNs), while dominant, do not offer these capabilities. We present deep implicit statistical shape models (DISSMs), a new approach that marries the representation power of deep networks with the benefits of SSMs. DISSMs use an implicit representation to produce compact and descriptive deep surface embeddings that permit statistical models of anatomical variance. To reliably fit anatomically plausible shapes to an image, we introduce a novel rigid and non-rigid pose estimation pipeline that is modelled as a Markov decision process (MDP). Intra-dataset experiments on the task of pathological liver segmentation demonstrate that DISSMs can perform more robustly than four leading FCN models, including nnU-Net + an adversarial prior: reducing the mean Hausdorff distance (HD) by 7.5-14.3 mm and improving the worst case Dice-Sørensen coefficient (DSC) by 1.2-2.3%. More critically, cross-dataset experiments on an external and highly challenging clinical dataset demonstrate that DISSMs improve the mean DSC and HD by 2.1-5.9% and 9.9-24.5 mm, respectively, and the worst-case DSC by 5.4-7.3%. Supplemental validation on a highly challenging and low-contrast larynx dataset further demonstrate DISSM’s improvements. These improvements are over and above any benefits from representing delineations with high-quality surfaces. Ashwin Raju, Shun Miao, Dakai Jin, Le Lu 0001, Junzhou Huang, Adam P. Harrison |
AAAI | 4 |
| 2022 | Localized Adversarial Domain GeneralizationabstractDeep learning methods can struggle to handle domain shifts not seen in training data, which can cause them to not generalize well to unseen domains. This has led to research attention on domain generalization (DG), which aims to the model's generalization ability to out-of-distribution. Adversarial domain generalization is a popular approach to DG, but conventional approaches (1) struggle to sufficiently align features so that local neighborhoods are mixed across domains; and (2) can suffer from feature space over collapse which can threaten generalization performance. To address these limitations, we propose localized adversarial domain generalization with space compactness maintenance (LADG) which constitutes two major contributions. First, we propose an adversarial localized classifier as the domain discriminator, along with a principled primary branch. This constructs a min-max game whereby the aim of the featurizer is to produce locally mixed domains. Second, we propose to use a coding-rate loss to alleviate feature space over collapse. We conduct comprehensive experiments on the Wilds DG benchmark to validate our approach, where LADG outperforms leading competitors on most datasets. Wei Zhu 0015, Le Lu 0001, Jing Xiao 0006, Jiebo Luo 0001, Adam P. Harrison |
CVPR | 2 |
| 2022 | Thoracic Lymph Node Segmentation in CT Imaging via Lymph Node Station Stratification and Size Encoding
Dazhou Guo, Jia Ge, Ke Yan 0006, Puyang Wang, Zhuotun Zhu, Xian-Sheng Hua 0001, Le Lu 0001, Tsung-Ying Ho, Xianghua Ye, Dakai Jin |
MICCAI (5) | 8 |
| 2022 | RemixFormer: A Transformer Model for Precision Skin Tumor Differential Diagnosis via Multi-modal Imaging and Non-imaging Data
Yuan Gao 0017, Wei Liu 0127, Kai Huang 0008, Le Lu 0001, Xiaosong Wang 0001, Xian-Sheng Hua 0001, Yu Wang 0108 |
MICCAI (3) | 6 |
| 2022 | DeepCRC: Colorectum and Colorectal Cancer Segmentation in CT Scans via Deep Colorectal Coordinate Transform
Yingda Xia, Jiawen Yao, Dakai Jin, Bingjiang Qiu, Suyun Li, Yanting Liang, Xian-Sheng Hua 0001, Le Lu 0001, Xin Chen 0058, Zaiyi Liu, Ling Zhang 0002 |
MICCAI (3) | 11 |
| 2022 | Effective Opportunistic Esophageal Cancer Screening Using Noncontrast CT Imaging
Jiawen Yao, Xianghua Ye, Yingda Xia, Ke Yan 0006, Lili Lin, Haogang Yu, Xian-Sheng Hua 0001, Le Lu 0001, Dakai Jin, Ling Zhang 0002 |
MICCAI (3) | 11 |
| 2022 | Multiphysical graph neural network (MP-GNN) for COVID-19 drug designabstractGraph neural networks (GNNs) are the most promising deep learning models that can revolutionize non-Euclidean data analysis. However, their full potential is severely curtailed by poorly represented molecular graphs and features. Here, we propose a multiphysical graph neural network (MP-GNN) model based on the developed multiphysical molecular graph representation and featurization. All kinds of molecular interactions, between different atom types and at different scales, are systematically represented by a series of scale-specific and element-specific graphs with distance-related node features. From these graphs, graph convolution network (GCN) models are constructed with specially designed weight-sharing architectures. Base learners are constructed from GCN models from different elements at different scales, and further consolidated together using both one-scale and multi-scale ensemble learning schemes. Our MP-GNN has two distinct properties. First, our MP-GNN incorporates multiscale interactions using more than one molecular graph. Atomic interactions from various different scales are not modeled by one specific graph (as in traditional GNNs), instead they are represented by a series of graphs at different scales. Second, it is free from the complicated feature generation process as in conventional GNN methods. In our MP-GNN, various atom interactions are embedded into element-specific graph representations with only distance-related node features. A unique GNN architecture is designed to incorporate all the information into a consolidated model. Our MP-GNN has been extensively validated on the widely used benchmark test datasets from PDBbind, including PDBbind-v2007, PDBbind-v2013 and PDBbind-v2016. Our model can outperform all existing models as far as we know. Further, our MP-GNN is used in coronavirus disease 2019 drug design. Based on a dataset with 185 complexes of inhibitors for severe acute respiratory syndrome coronavirus (SARS-CoV/SARS-CoV-2), we evaluate their binding affinities using our MP-GNN. It has been found that our MP-GNN is of high accuracy. This demonstrates the great potential of our MP-GNN for the screening of potential drugs for SARS-CoV-2. Availability: The Multiphysical graph neural network (MP-GNN) model can be found in https://github.com/Alibaba-DAMO-DrugAI/MGNN. Additional data or code will be available upon reasonable request. Xiao-Shuang Li, Xiang Liu 0021, Le Lu 0001, Xian-Sheng Hua 0001, Ying Chi, Kelin Xia |
Briefings Bioinform. | 3 |
| 2022 | Circle Representation for Medical Object DetectionabstractBox representation has been extensively used for object detection in computer vision. Such representation is efficacious but not necessarily optimized for biomedical objects (e.g., glomeruli), which play an essential role in renal pathology. In this paper, we propose a simple circle representation for medical object detection and introduce CircleNet, an anchor-free detection framework. Compared with the conventional bounding box representation, the proposed bounding circle representation innovates in three-fold: (1) it is optimized for ball-shaped biomedical objects; (2) The circle representation reduced the degree of freedom compared with box representation; (3) It is naturally more rotation invariant. When detecting glomeruli and nuclei on pathological images, the proposed circle representation achieved superior detection performance and be more rotation-invariant, compared with the bounding box. The code has been made publicly available: https://github.com/hrlblab/CircleNet. Ethan H. Nguyen, Haichun Yang, Ruining Deng, Yuzhe Lu, Zheyu Zhu, Joseph T. Roland, Le Lu 0001, Bennett A. Landman, Agnes B. Fogo, Yuankai Huo |
IEEE Trans. Medical Imaging | 7 |
| 2022 | SAM: Self-Supervised Learning of Pixel-Wise Anatomical Embeddings in Radiological ImagesabstractRadiological images such as computed tomography (CT) and X-rays render anatomy with intrinsic structures. Being able to reliably locate the same anatomical structure across varying images is a fundamental task in medical image analysis. In principle it is possible to use landmark detection or semantic segmentation for this task, but to work well these require large numbers of labeled data for each anatomical structure and sub-structure of interest. A more universal approach would learn the intrinsic structure from unlabeled images. We introduce such an approach, called Self-supervised Anatomical eMbedding (SAM). SAM generates semantic embeddings for each image pixel that describes its anatomical location or body part. To produce such embeddings, we propose a pixel-level contrastive learning framework. A coarse-to-fine strategy ensures both global and local anatomical information are encoded. Negative sample selection strategies are designed to enhance the embedding's discriminability. Using SAM, one can label any point of interest on a template image and then locate the same body part in other images by simple nearest neighbor searching. We demonstrate the effectiveness of SAM in multiple tasks with 2D and 3D image modalities. On a chest CT dataset with 19 landmarks, SAM outperforms widely-used registration algorithms while only taking 0.23 seconds for inference. On two X-ray datasets, SAM, with only one labeled template image, surpasses supervised methods trained on 50 labeled images. We also apply SAM on whole-body follow-up lesion matching in CT and obtain an accuracy of 91%. SAM can also be applied for improving image registration and initializing CNN weights. Ke Yan 0006, Jinzheng Cai, Dakai Jin, Shun Miao, Dazhou Guo, Adam P. Harrison, Youbao Tang, Jing Xiao 0006, Jingjing Lu, Le Lu 0001 |
IEEE Trans. Medical Imaging | 10 |
| 2021 | Window Loss for Bone Fracture Detection and Localization in X-ray Images with Point-based AnnotationabstractObject detection methods are widely adopted for computer-aided diagnosis using medical images. Anomalous findings are usually treated as objects that are described by bounding boxes. Yet, many pathological findings, e.g., bone fractures, cannot be clearly defined by bounding boxes, owing to considerable instance, shape and boundary ambiguities. This makes bounding box annotations, and their associated losses, highly ill-suited. In this work, we propose a new bone fracture detection method for X-ray images, based on a labor effective and flexible annotation scheme suitable for abnormal findings with no clear object-level spatial extents or boundaries. Our method employs a simple, intuitive, and informative point-based annotation protocol to mark localized pathology information. To address the uncertainty in the fracture scales annotated via point(s), we convert the annotations into pixel-wise supervision that uses lower and upper bounds with positive, negative, and uncertain regions. A novel Window Loss is subsequently proposed to only penalize the predictions outside of the uncertain regions. Our method has been extensively evaluated on 4410 pelvic X-ray images of unique patients. Experiments demonstrate that our method outperforms previous state-of-the-art image classification and object detection baselines by healthy margins, with an AUROC of 0.983 and FROC score of 89.6%. Yirui Wang 0002, Chi-Tung Cheng, Le Lu 0001, Adam P. Harrison, Jing Xiao 0006, Chien-Hung Liao, Shun Miao |
AAAI | 4 |
| 2021 | Deep Lesion Tracker: Monitoring Lesions in 4D Longitudinal Imaging StudiesabstractMonitoring treatment response in longitudinal studies plays an important role in clinical practice. Accurately identifying lesions across serial imaging follow-up is the core to the monitoring procedure. Typically this incorporates both image and anatomical considerations. However, matching lesions manually is labor-intensive and time-consuming. In this work, we present deep lesion tracker (DLT), a deep learning approach that uses both appearance- and anatomical-based signals. To incorporate anatomical constraints, we propose an anatomical signal encoder, which prevents lesions being matched with visually similar but spurious regions. In addition, we present a new formulation for Siamese networks that avoids the heavy computational loads of 3D cross-correlation. To present our network with greater varieties of images, we also propose a self-supervised learning (SSL) strategy to train trackers with unpaired images, overcoming barriers to data collection. To train and evaluate our tracker, we introduce and release the first lesion tracking benchmark, consisting of 3891 lesion pairs from the public DeepLesion database. The proposed method, DLT, locates lesion centers with a mean error distance of 7mm. This is 5% better than a leading registration algorithm while running 14 times faster on whole CT volumes. We demonstrate even greater improvements over detector or similarity-learning alternatives. DLT also generalizes well on an external clinical test set of 100 longitudinal studies, achieving 88% accuracy. Finally, we plug DLT into an automatic tumor monitoring workflow where it leads to an accuracy of 85% in assessing lesion treatment responses, which is only 0.46% lower than the accuracy of manual inputs. Jinzheng Cai, Youbao Tang, Ke Yan 0006, Adam P. Harrison, Jing Xiao 0006, Gigin Lin, Le Lu 0001 |
CVPR | 7 |
| 2021 | Automatic Vertebra Localization and Identification in CT by Spine Rectification and Anatomically-Constrained OptimizationabstractAccurate vertebra localization and identification are required in many clinical applications of spine disorder diagnosis and surgery planning. However, significant challenges are posed in this task by highly varying pathologies (such as vertebral compression fracture, scoliosis, and vertebral fixation) and imaging conditions (such as limited field of view and metal streak artifacts). This paper proposes a robust and accurate method that effectively exploits the anatomical knowledge of the spine to facilitate vertebra localization and identification. A key point localization model is trained to produce activation maps of vertebra centers. They are then re-sampled along the spine centerline to produce spine-rectified activation maps, which are further aggregated into 1-D activation signals. Following this, an anatomically-constrained optimization module is introduced to jointly search for the optimal vertebra centers under a soft constraint that regulates the distance between vertebrae and a hard constraint on the consecutive vertebra indices. When being evaluated on a major public benchmark of 302 highly pathological CT images, the proposed method reports the state of the art identification (id.) rate of 97.4%, and outperforms the best competing method of 94.7% id. rate by reducing the relative id. error rate by half. Fakai Wang, Le Lu 0001, Jing Xiao 0006, Min Wu 0001, Shun Miao |
CVPR | 3 |
| 2021 | 3D Graph Anatomy Geometry-Integrated Network for Pancreatic Mass Segmentation, Diagnosis, and Quantitative Patient ManagementabstractThe pancreatic disease taxonomy includes ten types of masses (tumors or cysts) [20], [8]. Previous work focuses on developing segmentation or classification methods only for certain mass types. Differential diagnosis of all mass types is clinically highly desirable [20] but has not been investigated using an automated image understanding approach.We exploit the feasibility to distinguish pancreatic ductal adenocarcinoma (PDAC) from the nine other nonPDAC masses using multi-phase CT imaging. Both image appearance and the 3D organ-mass geometry relationship are critical. We propose a holistic segmentation-mesh-classification network (SMCN) to provide patient-level diagnosis, by fully utilizing the geometry and location information, which is accomplished by combining the anatomical structure and the semantic detection-by-segmentation network. SMCN learns the pancreas and mass segmentation task and builds an anatomical correspondence-aware organ mesh model by progressively deforming a pancreas prototype on the raw segmentation mask (i.e., mask-to-mesh). A new graph-based residual convolutional network (Graph-ResNet), whose nodes fuse the information of the mesh model and feature vectors extracted from the segmentation network, is developed to produce the patient-level differential classification results. Extensive experiments on 661 patients’ CT scans (five phases per patient) show that SMCN can improve the mass segmentation and detection accuracy compared to the strong baseline method nnUNet (e.g., for nonPDAC, Dice: 0.611 vs. 0.478; detection rate: 89% vs. 70%), achieve similar sensitivity and specificity in differentiating PDAC and nonPDAC as expert radiologists (i.e., 94% and 90%), and obtain results comparable to a multimodality test [20] that combines clinical, imaging, and molecular testing for clinical management of patients. Jiawen Yao, Isabella Nogues, Le Lu 0001, Lingyun Huang, Jing Xiao 0006, Zhaozheng Yin, Ling Zhang 0002 |
CVPR | 5 |
| 2021 | Sequential Learning on Liver Tumor Boundary Semantics and Prognostic Biomarker Mining
Jieneng Chen, Ke Yan 0006, Youbao Tang, Shuwen Sun, Qiuping Liu, Lingyun Huang, Jing Xiao 0006, Alan L. Yuille, Ya Zhang 0002, Le Lu 0001 |
MICCAI (7) | 12 |
| 2021 | DeepStationing: Thoracic Lymph Node Station Parsing in CT Scans Using Anatomical Context Encoding and Key Organ Auto-Search
Dazhou Guo, Xianghua Ye, Jia Ge, Xing Di, Le Lu 0001, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Zhongjie Lu, Senxiang Yan, Dakai Jin |
MICCAI (5) | 5 |
| 2021 | Learning from Subjective Ratings Using Auto-Decoded Deep Latent Embeddings
Xinping Ren, Ke Yan 0006, Le Lu 0001, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Dar-In Tai, Adam P. Harrison |
MICCAI (5) | 4 |
| 2021 | SAME: Deformable Image Registration Based on Self-supervised Anatomical Embeddings
Fengze Liu, Ke Yan 0006, Adam P. Harrison, Dazhou Guo, Le Lu 0001, Alan L. Yuille, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Xianghua Ye, Dakai Jin |
MICCAI (4) | 5 |
| 2021 | Weakly-Supervised Universal Lesion Segmentation with Regional Level Set Loss
Youbao Tang, Jinzheng Cai, Ke Yan 0006, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Jingjing Lu, Gigin Lin, Le Lu 0001 |
MICCAI (2) | 9 |
| 2021 | Lesion Segmentation and RECIST Diameter Prediction via Click-Driven Attention and Dual-Path Connection
Youbao Tang, Ke Yan 0006, Jinzheng Cai, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Jingjing Lu, Gigin Lin, Le Lu 0001 |
MICCAI (2) | 9 |
| 2021 | Effective Pancreatic Cancer Screening on Non-contrast CT Scans via Anatomy-Aware Transformers
Yingda Xia, Jiawen Yao, Le Lu 0001, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Alan L. Yuille, Ling Zhang 0002 |
MICCAI (5) | 3 |
| 2021 | Semi-supervised Learning for Bone Mineral Density Estimation in Hip X-Ray Images
Yirui Wang 0002, Xiaoyun Zhou 0001, Fakai Wang, Le Lu 0001, Chihung Lin, Lingyun Huang, Guo Tong Xie, Jing Xiao 0006, Chang-Fu Kuo, Shun Miao |
MICCAI (5) | 5 |
| 2021 | DeepTarget: Gross tumor and clinical target volume segmentation in esophageal cancer radiotherapy
Dakai Jin, Dazhou Guo, Tsung-Ying Ho, Adam P. Harrison, Jing Xiao 0006, Chen-Kan Tseng, Le Lu 0001 |
Medical Image Anal. | 7 |
| 2021 | DeepPrognosis: Preoperative prediction of pancreatic cancer survival and surgical margin via comprehensive understanding of dynamic contrast-enhanced CT imaging and tumor-vascular contact parsing
Jiawen Yao, Le Lu 0001, Jianping Lu, Qike Song, Gang Jin, Jing Xiao 0006, Ling Zhang 0002 |
Medical Image Anal. | 4 |
| 2021 | Faster Mean-shift: GPU-accelerated clustering for cosine embedding-based cell segmentation and tracking
Mengyang Zhao 0001, Aadarsh Jha, Quan Liu 0002, Bryan A. Millis, Anita Mahadevan-Jansen, Le Lu 0001, Bennett A. Landman, Matthew J. Tyska, Yuankai Huo |
Medical Image Anal. | 6 |
| 2021 | Lesion-Harvester: Iteratively Mining Unlabeled Lesions and Hard-Negative Examples at ScaleabstractThe acquisition of large-scale medical image data, necessary for training machine learning algorithms, is hampered by associated expert-driven annotation costs. Mining hospital archives can address this problem, but labels often incomplete or noisy, e.g., 50% of the lesions in DeepLesion are left unlabeled. Thus, effective label harvesting methods are critical. This is the goal of our work, where we introduce Lesion-Harvester-a powerful system to harvest missing annotations from lesion datasets at high precision. Accepting the need for some degree of expert labor, we use a small fully-labeled image subset to intelligently mine annotations from the remainder. To do this, we chain together a highly sensitive lesion proposal generator (LPG) and a very selective lesion proposal classifier (LPC). Using a new hard negative suppression loss, the resulting harvested and hard-negative proposals are then employed to iteratively finetune our LPG. While our framework is generic, we optimize our performance by proposing a new 3D contextual LPG and by using a global-local multi-view LPC. Experiments on DeepLesion demonstrate that Lesion-Harvester can discover an additional 9,805 lesions at a precision of 90%. We publicly release the harvested lesions, along with a new test set of completely annotated DeepLesion volumes. We also present a pseudo 3D IoU evaluation metric that corresponds much better to the real 3D IoU than current DeepLesion evaluation metrics. To quantify the downstream benefits of Lesion-Harvester we show that augmenting the DeepLesion annotations with our harvested lesions allows state-of-the-art detectors to boost their average precision by 7 to 10%. Jinzheng Cai, Adam P. Harrison, Youjing Zheng, Ke Yan 0006, Yuankai Huo, Jing Xiao 0006, Lin Yang 0002, Le Lu 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2021 | Contour Transformer Network for One-Shot Segmentation of Anatomical StructuresabstractAccurate segmentation of anatomical structures is vital for medical image analysis. The state-of-the-art accuracy is typically achieved by supervised learning methods, where gathering the requisite expert-labeled image annotations in a scalable manner remains a main obstacle. Therefore, annotation-efficient methods that permit to produce accurate anatomical structure segmentation are highly desirable. In this work, we present Contour Transformer Network (CTN), a one-shot anatomy segmentation method with a naturally built-in human-in-the-loop mechanism. We formulate anatomy segmentation as a contour evolution process and model the evolution behavior by graph convolutional networks (GCNs). Training the CTN model requires only one labeled image exemplar and leverages additional unlabeled data through newly introduced loss functions that measure the global shape and appearance consistency of contours. On segmentation tasks of four different anatomies, we demonstrate that our one-shot learning method significantly outperforms non-learning-based methods and performs competitively to the state-of-the-art fully supervised deep learning methods. With minimal human-in-the-loop editing feedback, the segmentation performance can be further improved to surpass the fully supervised methods. Weijian Li 0001, Yirui Wang 0002, Adam P. Harrison, Chihung Lin, Song Wang 0002, Jing Xiao 0006, Le Lu 0001, Chang-Fu Kuo, Shun Miao |
IEEE Trans. Medical Imaging | 9 |
| 2021 | Learning From Multiple Datasets With Heterogeneous and Partial Labels for Universal Lesion Detection in CTabstractLarge-scale datasets with high-quality labels are desired for training accurate deep learning models. However, due to the annotation cost, datasets in medical imaging are often either partially-labeled or small. For example, DeepLesion is such a large-scale CT image dataset with lesions of various types, but it also has many unlabeled lesions (missing annotations). When training a lesion detector on a partially-labeled dataset, the missing annotations will generate incorrect negative signals and degrade the performance. Besides DeepLesion, there are several small single-type datasets, such as LUNA for lung nodules and LiTS for liver tumors. These datasets have heterogeneous label scopes, i.e., different lesion types are labeled in different datasets with other types ignored. In this work, we aim to develop a universal lesion detection algorithm to detect a variety of lesions. The problem of heterogeneous and partial labels is tackled. First, we build a simple yet effective lesion detection framework named Lesion ENSemble (LENS). LENS can efficiently learn from multiple heterogeneous lesion datasets in a multi-task fashion and leverage their synergy by proposal fusion. Next, we propose strategies to mine missing annotations from partially-labeled datasets by exploiting clinical prior knowledge and cross-dataset knowledge transfer. Finally, we train our framework on four public lesion datasets and evaluate it on 800 manually-labeled sub-volumes in DeepLesion. Our method brings a relative improvement of 49% compared to the current state-of-the-art approach in the metric of average sensitivity. We have publicly released our manual 3D annotations of DeepLesion online.11https://github.com/viggin/DeepLesion_manual_test_set Ke Yan 0006, Jinzheng Cai, Youjing Zheng, Adam P. Harrison, Dakai Jin, Youbao Tang, Yuxing Tang, Lingyun Huang, Jing Xiao 0006, Le Lu 0001 |
IEEE Trans. Medical Imaging | 10 |
| 2020 | Organ at Risk Segmentation for Head and Neck Cancer Using Stratified Learning and Neural Architecture SearchabstractOAR segmentation is a critical step in radiotherapy of head and neck (H&N) cancer, where inconsistencies across radiation oncologists and prohibitive labor costs motivate automated approaches. However, leading methods using standard fully convolutional network workflows that are challenged when the number of OARs becomes large, e.g. > 40. For such scenarios, insights can be gained from the stratification approaches seen in manual clinical OAR delineation. This is the goal of our work, where we introduce stratified organ at risk segmentation (SOARS), an approach that stratifies OARs into anchor, mid-level, and small & hard (S&H) categories. SOARS stratifies across two dimensions. The first dimension is that distinct processing pipelines are used for each OAR category. In particular, inspired by clinical practices, anchor OARs are used to guide the mid-level and S&H categories. The second dimension is that distinct network architectures are used to manage the significant contrast, size, and anatomy variations between different OARs. We use differentiable neural architecture search (NAS), allowing the network to choose among 2D, 3D or Pseudo-3D convolutions. Extensive 4-fold cross-validation on 142 H&N cancer patients with 42 manually labeled OARs, the most comprehensive OAR dataset to date, demonstrates that both pipeline- and NAS-stratification significantly improves quantitative performance over the state-of-the-art (from 69.52% to 73.68% in absolute Dice scores). Thus, SOARS provides a powerful and principled means to manage the highly complex segmentation space of OARs. Dazhou Guo, Dakai Jin, Zhuotun Zhu, Tsung-Ying Ho, Adam P. Harrison, Chun-Hung Chao, Jing Xiao 0006, Le Lu 0001 |
CVPR | 8 |
| 2020 | Anatomy-Aware Siamese Network: Exploiting Semantic Asymmetry for Accurate Pelvic Fracture Detection in X-Ray Images
Haomin Chen, Yirui Wang 0002, Weijian Li 0001, Chi-Tung Chang, Adam P. Harrison, Jing Xiao 0006, Gregory D. Hager, Le Lu 0001, Chien-Hung Liao, Shun Miao |
ECCV (23) | 9 |
| 2020 | Structured Landmark Detection via Topology-Adapting Deep Graph Learning
Weijian Li 0001, Haofu Liao, Chihung Lin, Jiebo Luo 0001, Chi-Tung Cheng, Jing Xiao 0006, Le Lu 0001, Chang-Fu Kuo, Shun Miao |
ECCV (9) | 9 |
| 2020 | JSSR: A Joint Synthesis, Segmentation, and Registration System for 3D Multi-modal Image Alignment of Large-Scale Pathological CT Scans
Fengze Liu, Jinzheng Cai, Yuankai Huo, Chi-Tung Cheng, Ashwin Raju, Dakai Jin, Jing Xiao 0006, Alan L. Yuille, Le Lu 0001, Chien-Hung Liao, Adam P. Harrison |
ECCV (13) | 9 |
| 2020 | Co-heterogeneous and Adaptive Segmentation from Multi-source and Multi-phase CT Imaging Data: A Study on Pathological Liver and Lesion Segmentation
Ashwin Raju, Chi-Tung Cheng, Yuankai Huo, Jinzheng Cai, Junzhou Huang, Jing Xiao 0006, Le Lu 0001, Chien-Hung Liao, Adam P. Harrison |
ECCV (23) | 7 |
| 2020 | Unsupervised Learning of Facial Landmarks based on Inter-Intra Subject ConsistenciesabstractWe present a novel unsupervised learning approach to image landmark discovery by incorporating the inter-subject landmark consistencies on facial images. This is achieved via an inter-subject mapping module that transforms original subject landmarks based on an auxiliary subject-related structure. To recover from the transformed images back to the original subject, the landmark detector is forced to learn spatial locations that contain the consistent semantic meanings both for the paired intra-subject images and between the paired inter-subject images. Our proposed method is extensively evaluated on two public facial image datasets (MAFL, AFLW) with various settings. Experimental results indicate that our method can extract the consistent landmarks for both datasets and achieve better performances compared to the previous state-of-the-art methods quantitatively and qualitatively. Weijian Li 0001, Haofu Liao, Shun Miao, Le Lu 0001, Jiebo Luo 0001 |
ICPR | 4 |
| 2020 | Deep Volumetric Universal Lesion Detection Using Light-Weight Pseudo 3D Convolution and Surface Point Regression
Jinzheng Cai, Ke Yan 0006, Chi-Tung Cheng, Jing Xiao 0006, Chien-Hung Liao, Le Lu 0001, Adam P. Harrison |
MICCAI (4) | 6 |
| 2020 | Lymph Node Gross Tumor Volume Detection in Oncology Imaging via Relationship Learning Using Graph Neural Network
Chun-Hung Chao, Zhuotun Zhu, Dazhou Guo, Ke Yan 0006, Tsung-Ying Ho, Jinzheng Cai, Adam P. Harrison, Xianghua Ye, Jing Xiao 0006, Alan L. Yuille, Min Sun 0001, Le Lu 0001, Dakai Jin |
MICCAI (7) | 12 |
| 2020 | Reliable Liver Fibrosis Assessment from Ultrasound Using Global Hetero-Image Fusion and View-Specific Parameterization
Ke Yan 0006, Dar-In Tai, Yuankai Huo, Le Lu 0001, Jing Xiao 0006, Adam P. Harrison |
MICCAI (3) | 5 |
| 2020 | Learning to Segment Anatomical Structures Accurately from One Exemplar
Weijian Li 0001, Yirui Wang 0002, Adam P. Harrison, Chihung Lin, Song Wang 0002, Jing Xiao 0006, Le Lu 0001, Chang-Fu Kuo, Shun Miao |
MICCAI (1) | 9 |
| 2020 | User-Guided Domain Adaptation for Rapid Annotation from User Interactions: A Study on Pathological Liver Segmentation
Ashwin Raju, Zhanghexuan Ji, Chi-Tung Cheng, Jinzheng Cai, Junzhou Huang, Jing Xiao 0006, Le Lu 0001, Chien-Hung Liao, Adam P. Harrison |
MICCAI (1) | 7 |
| 2020 | CircleNet: Anchor-Free Glomerulus Detection with Circle Representation
Haichun Yang, Ruining Deng, Yuzhe Lu, Zheyu Zhu, Joseph T. Roland, Le Lu 0001, Bennett A. Landman, Agnes B. Fogo, Yuankai Huo |
MICCAI (4) | 7 |
| 2020 | DeepPrognosis: Preoperative Prediction of Pancreatic Cancer Survival and Surgical Margin via Contrast-Enhanced CT Imaging
Jiawen Yao, Le Lu 0001, Jing Xiao 0006, Ling Zhang 0002 |
MICCAI (2) | 3 |
| 2020 | Robust Pancreatic Ductal Adenocarcinoma Segmentation with Multi-institutional Multi-phase Partially-Annotated CT Scans
Ling Zhang 0002, Jiawen Yao, Yun Bian, Dakai Jin, Jing Xiao 0006, Le Lu 0001 |
MICCAI (4) | 8 |
| 2020 | Lymph Node Gross Tumor Volume Detection and Segmentation via Distance-Based Gating Using 3D CT/PET Imaging in Radiotherapy
Zhuotun Zhu, Dakai Jin, Ke Yan 0006, Tsung-Ying Ho, Xianghua Ye, Dazhou Guo, Chun-Hung Chao, Jing Xiao 0006, Alan L. Yuille, Le Lu 0001 |
MICCAI (7) | 10 |
| 2020 | Thorax-Net: An Attention Regularized Deep Neural Network for Classification of Thoracic Diseases on Chest RadiographyabstractDeep learning techniques have been increasingly used to provide more accurate and more accessible diagnosis of thorax diseases on chest radiographs. However, due to the lack of dense annotation of large-scale chest radiograph data, this computer-aided diagnosis task is intrinsically a weakly supervised learning problem and remains challenging. In this paper, we propose a novel deep convolutional neural network called Thorax-Net to diagnose 14 thorax diseases using chest radiography. Thorax-Net consists of a classification branch and an attention branch. The classification branch serves as a uniform feature extraction-classification network to free users from the troublesome hand-crafted feature extraction and classifier construction. The attention branch exploits the correlation between class labels and the locations of pathological abnormalities via analyzing the feature maps learned by the classification branch. Feeding a chest radiograph to the trained Thorax-Net, a diagnosis is obtained by averaging and binarizing the outputs of two branches. The proposed Thorax-Net model has been evaluated against three state-of-the-art deep learning models using the patientwise official split of the ChestX-ray14 dataset and against other five deep learning models using the imagewise random data split. Our results show that Thorax-Net achieves an average per-class area under the receiver operating characteristic curve (AUC) of 0.7876 and 0.896 in both experiments, respectively, which are higher than the AUC values obtained by other deep models when they were all trained with no external data. Hongyu Wang 0011, Haozhe Jia, Le Lu 0001, Yong Xia 0001 |
IEEE J. Biomed. Health Informatics | 3 |
| 2020 | Spatio-Temporal Convolutional LSTMs for Tumor Growth Prediction by Learning 4D Longitudinal Patient DataabstractPrognostic tumor growth modeling via volumetric medical imaging observations can potentially lead to better outcomes of tumor treatment management and surgical planning. Recent advances of convolutional networks (ConvNets) have demonstrated higher accuracy than traditional mathematical models can be achieved in predicting future tumor volumes. This indicates that deep learning based data-driven techniques may have great potentials on addressing such problem. However, current 2D image patch based modeling approaches can not make full use of the spatio-temporal imaging context of the tumor's longitudinal 4D (3D + time) patient data. Moreover, they are incapable to predict clinically-relevant tumor properties, other than the tumor volumes. In this paper, we exploit to formulate the tumor growth process through convolutional Long Short-Term Memory (ConvLSTM) that extract tumor's static imaging appearances and simultaneously capture its temporal dynamic changes within a single network. We extend ConvLSTM into the spatio-temporal domain (ST-ConvLSTM) by jointly learning the inter-slice 3D contexts and the longitudinal or temporal dynamics from multiple patient studies. Our approach can incorporate other non-imaging patient information in an end-to-end trainable manner. Experiments are conducted on the largest 4D longitudinal tumor dataset of 33 patients to date. Results validate that the proposed ST-ConvLSTM model produces a Dice score of 83.2%±5.1% and a RVD of 11.2%±10.8%, both statistically significantly outperforming (p < 0.05) other compared methods of traditional linear model, ConvLSTM, and generative adversarial network (GAN) under the metric of predicting future tumor volumes. Additionally, our new method enables the prediction of both cell density and CT intensity numbers. Last, we demonstrate the generalizability of ST-ConvLSTM by employing it in 4D medical image segmentation task, which achieves an averaged Dice score of 86.3%±1.2% for left-ventricle segmentation in 4D ultrasound with 3 seconds per patient case. Ling Zhang 0002, Le Lu 0001, Xiaosong Wang 0001, Robert Zhu, Mohammadhadi Bagheri, Ronald M. Summers, Jianhua Yao 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Accurate Esophageal Gross Tumor Volume Segmentation in PET/CT Using Two-Stream Chained 3D Deep Network Fusion
Dakai Jin, Dazhou Guo, Tsung-Ying Ho, Adam P. Harrison, Jing Xiao 0006, Chen-Kan Tseng, Le Lu 0001 |
MICCAI (2) | 7 |
| 2019 | Deep Esophageal Clinical Target Volume Delineation Using Encoded 3D Spatial Context of Tumors, Lymph Nodes, and Organs At Risk
Dakai Jin, Dazhou Guo, Tsung-Ying Ho, Adam P. Harrison, Jing Xiao 0006, Chen-Kan Tseng, Le Lu 0001 |
MICCAI (6) | 7 |
| 2019 | Weakly Supervised Universal Fracture Detection in Pelvic X-Rays
Yirui Wang 0002, Le Lu 0001, Chi-Tung Cheng, Dakai Jin, Adam P. Harrison, Jing Xiao 0006, Chien-Hung Liao, Shun Miao |
MICCAI (6) | 2 |
| 2018 | TieNet: Text-Image Embedding Network for Common Thorax Disease Classification and Reporting in Chest X-RaysabstractChest X-rays are one of the most common radiological examinations in daily clinical routines. Reporting thorax diseases using chest X-rays is often an entry-level task for radiologist trainees. Yet, reading a chest X-ray image remains a challenging job for learning-oriented machine intelligence, due to (1) shortage of large-scale machine-learnable medical image datasets, and (2) lack of techniques that can mimic the high-level reasoning of human radiologists that requires years of knowledge accumulation and professional training. In this paper, we show the clinical free-text radiological reportscan be utilized as a priori knowledge for tackling these two key problems. We propose a novel Text-Image Embedding network (TieNet) for extracting the distinctive image and text representations. Multi-level attention models are integrated into an end-to-end trainable CNN-RNN architecture for highlighting the meaningful text words and image regions. We first apply TieNet to classify the chest X-rays by using both image features and text embeddings extracted from associated reports. The proposed auto-annotation framework achieves high accuracy (over 0.9 on average in AUCs) in assigning disease labels for our hand-label evaluation dataset. Furthermore, we transform the TieNet into a chest X-ray reporting system. It simulates the reporting process and can output disease classification and a preliminary report together. The classification results are significantly improved (6% increase on average in AUCs) compared to the state-of-the-art baseline on an unseen and hand-labeled dataset (OpenI). Xiaosong Wang 0001, Yifan Peng 0002, Le Lu 0001, Zhiyong Lu, Ronald M. Summers |
CVPR | 3 |
| 2018 | Deep Lesion Graphs in the Wild: Relationship Learning and Organization of Significant Radiology Image Findings in a Diverse Large-Scale Lesion DatabaseabstractRadiologists in their daily work routinely find and annotate significant abnormalities on a large number of radiology images. Such abnormalities, or lesions, have collected over years and stored in hospitals' picture archiving and communication systems. However, they are basically unsorted and lack semantic annotations like type and location. In this paper, we aim to organize and explore them by learning a deep feature representation for each lesion. A large-scale and comprehensive dataset, DeepLesion, is introduced for this task. DeepLesion contains bounding boxes and size measurements of over 32K lesions. To model their similarity relationship, we leverage multiple supervision information including types, self-supervised location coordinates, and sizes. They require little manual annotation effort but describe useful attributes of the lesions. Then, a triplet network is utilized to learn lesion embeddings with a sequential sampling strategy to depict their hierarchical similarity structure. Experiments show promising qualitative and quantitative results on lesion retrieval, clustering, and classification. The learned embeddings can be further employed to build a lesion graph for various clinically useful applications. An algorithm for intra-patient lesion matching is proposed and validated with experiments. Ke Yan 0006, Xiaosong Wang 0001, Le Lu 0001, Ling Zhang 0002, Adam P. Harrison, Mohammadhadi Bagheri, Ronald M. Summers |
CVPR | 3 |
| 2018 | Iterative Attention Mining for Weakly Supervised Thoracic Disease Pattern Localization in Chest X-Rays
Jinzheng Cai, Le Lu 0001, Adam P. Harrison, Xiaoshuang Shi, Pingjun Chen, Lin Yang 0002 |
MICCAI (2) | 2 |
| 2018 | Accurate Weakly-Supervised Deep Lesion Segmentation Using Large-Scale Clinical Annotations: Slice-Propagated 3D Mask Generation from 2D RECIST
Jinzheng Cai, Youbao Tang, Le Lu 0001, Adam P. Harrison, Ke Yan 0006, Jing Xiao 0006, Lin Yang 0002, Ronald M. Summers |
MICCAI (4) | 3 |
| 2018 | Towards Automated Colonoscopy Diagnosis: Binary Polyp Size Estimation via Unsupervised Depth Learning
Hayato Itoh, Holger Roth, Le Lu 0001, Masahiro Oda 0001, Masashi Misawa, Yuichi Mori, Shin-ei Kudo, Kensaku Mori |
MICCAI (2) | 3 |
| 2018 | A Decomposable Model for the Detection of Prostate Cancer in Multi-parametric MRI
Nathan Lay, Yohannes Tsehay, Yohan Sumathipala, Ruida Cheng, Sonia Gaur, Clayton Smith, Adrian Barbu, Le Lu 0001, Baris Turkbey, Peter L. Choyke, Peter A. Pinto, Ronald M. Summers |
MICCAI (2) | 8 |
| 2018 | Spatial aggregation of holistically-nested convolutional neural networks for automated pancreas localization and segmentation
Holger Roth, Le Lu 0001, Nathan Lay, Adam P. Harrison, Amal Farag, Andrew Sohn, Ronald M. Summers |
Medical Image Anal. | 2 |
| 2018 | Convolutional Invasion and Expansion Networks for Tumor Growth PredictionabstractTumor growth is associated with cell invasion and mass-effect, which are traditionally formulated by mathematical models, namely reaction-diffusion equations and biomechanics. Such models can be personalized based on clinical measurements to build the predictive models for tumor growth. In this paper, we investigate the possibility of using deep convolutional neural networks to directly represent and learn the cell invasion and mass-effect, and to predict the subsequent involvement regions of a tumor. The invasion network learns the cell invasion from information related to metabolic rate, cell density, and tumor boundary derived from multimodal imaging data. The expansion network models the mass-effect from the growing motion of tumor mass. We also study different architectures that fuse the invasion and expansion networks, in order to exploit the inherent correlations among them. Our network can easily be trained on population data and personalized to a target patient, unlike most previous mathematical modeling methods that fail to incorporate population data. Quantitative experiments on a pancreatic tumor data set show that the proposed method substantially outperforms a state-of-the-art mathematical model-based approach in both accuracy and efficiency, and that the information captured by each of the two subnetworks is complementary. Ling Zhang 0002, Le Lu 0001, Ronald M. Summers, Electron Kebebew, Jianhua Yao 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2017 | Text Mining Radiology Reports for Deep Learning Radiology Images
Yifan Peng 0002, Xiaosong Wang 0001, Le Lu 0001, Mohammadhadi Bagheri, Ronald M. Summers, Zhiyong Lu |
AMIA | 3 |
| 2017 | ChestX-Ray8: Hospital-Scale Chest X-Ray Database and Benchmarks on Weakly-Supervised Classification and Localization of Common Thorax DiseasesabstractThe chest X-ray is one of the most commonly accessible radiological examinations for screening and diagnosis of many lung diseases. A tremendous number of X-ray imaging studies accompanied by radiological reports are accumulated and stored in many modern hospitals Picture Archiving and Communication Systems (PACS). On the other side, it is still an open question how this type of hospital-size knowledge database containing invaluable imaging informatics (i.e., loosely labeled) can be used to facilitate the data-hungry deep learning paradigms in building truly large-scale high precision computer-aided diagnosis (CAD) systems. In this paper, we present a new chest X-ray database, namely ChestX-ray8, which comprises 108,948 frontal-view X-ray images of 32,717 unique patients with the text-mined eight disease image labels (where each image can have multi-labels), from the associated radiological reports using natural language processing. Importantly, we demonstrate that these commonly occurring thoracic diseases can be detected and even spatially-located via a unified weakly-supervised multi-label image classification and disease localization framework, which is validated using our proposed dataset. Although the initial quantitative results are promising as reported, deep convolutional neural network based reading chest X-rays (i.e., recognizing and locating the common disease patterns trained with only image-level labels) remains a strenuous task for fully-automated high precision CAD systems. Xiaosong Wang 0001, Yifan Peng 0002, Le Lu 0001, Zhiyong Lu, Mohammadhadi Bagheri, Ronald M. Summers |
CVPR | 3 |
| 2017 | Pancreas Segmentation in MRI Using Graph-Based Decision Fusion on Convolutional Neural Networks
Jinzheng Cai, Le Lu 0001, Yuanpu Xie, Fuyong Xing, Lin Yang 0002 |
MICCAI (3) | 2 |
| 2017 | Progressive and Multi-path Holistically Nested Neural Networks for Pathological Lung Segmentation from CT Images
Adam P. Harrison, Ziyue Xu 0001, Kevin George, Le Lu 0001, Ronald M. Summers, Daniel J. Mollura |
MICCAI (3) | 4 |
| 2017 | Personalized Pancreatic Tumor Growth Prediction via Group Learning
Ling Zhang 0002, Le Lu 0001, Ronald M. Summers, Electron Kebebew, Jianhua Yao 0001 |
MICCAI (2) | 2 |
| 2017 | Unsupervised Joint Mining of Deep Features and Image Labels for Large-Scale Radiology Image Categorization and Scene RecognitionabstractThe recent rapid and tremendous success of deep convolutional neural networks (CNN) on many challenging computer vision tasks largely derives from the accessibility of the well-annotated ImageNet and PASCAL VOC datasets. Nevertheless, unsupervised image categorization (i.e., without the ground-truth labeling) is much less investigated, yet critically important and difficult when annotations are extremely hard to obtain in the conventional way of "Google Search" and crowd sourcing. We address this problem by presenting a looped deep pseudo-task optimization (LDPO) framework for joint mining of deep CNN features and image labels. Our method is conceptually simple and rests upon the hypothesized "convergence" of better labels leading to better trained CNN models which in turn feed more discriminative image representations to facilitate more meaningful clusters/labels. Our proposed method is validated in tackling two important applications: 1) Large-scale medical image annotation has always been a prohibitively expensive and easily-biased task even for well-trained radiologists. Significantly better image categorization results are achieved via our proposed approach compared to the previous state-of-the-art method. 2) Unsupervised scene recognition on representative and publicly available datasets with our proposed technique is examined. The LDPO achieves excellent quantitative scene classification results. On the MIT indoor scene dataset, it attains a clustering accuracy of 75:3%, compared to the state-of-the-art supervised classification accuracy of 81:0% (when both are based on the VGG-VD model). Xiaosong Wang 0001, Le Lu 0001, Hoo-Chang Shin, Lauren Kim, Mohammadhadi Bagheri, Isabella Nogues, Jianhua Yao 0001, Ronald M. Summers |
WACV | 2 |
| 2017 | A Bottom-Up Approach for Pancreas Segmentation Using Cascaded Superpixels and (Deep) Image Patch LabelingabstractRobust organ segmentation is a prerequisite for computer-aided diagnosis, quantitative imaging analysis, pathology detection, and surgical assistance. For organs with high anatomical variability (e.g., the pancreas), previous segmentation approaches report low accuracies, compared with well-studied organs, such as the liver or heart. We present an automated bottom-up approach for pancreas segmentation in abdominal computed tomography (CT) scans. The method generates a hierarchical cascade of information propagation by classifying image patches at different resolutions and cascading (segments) superpixels. The system contains four steps: 1) decomposition of CT slice images into a set of disjoint boundary-preserving superpixels; 2) computation of pancreas class probability maps via dense patch labeling; 3) superpixel classification by pooling both intensity and probability features to form empirical statistics in cascaded random forest frameworks; and 4) simple connectivity based post-processing. Dense image patch labeling is conducted using two methods: efficient random forest classification on image histogram, location and texture features; and more expensive (but more accurate) deep convolutional neural network classification, on larger image windows (i.e., with more spatial contexts). Over-segmented 2-D CT slices by the simple linear iterative clustering approach are adopted through model/parameter calibration and labeled at the superpixel level for positive (pancreas) or negative (non-pancreas or background) classes. The proposed method is evaluated on a data set of 80 manually segmented CT volumes, using six-fold cross-validation. Its performance equals or surpasses other state-of-the-art methods (evaluated by "leave-one-patient-out"), with a dice coefficient of 70.7% and Jaccard index of 57.9%. In addition, the computational efficiency has improved significantly, requiring a mere 6 ~ 8 min per testing case, versus ≥ 10 h for other methods. The segmentation framework using deep patch labeling confidences is also more numerically stable, as reflected in the smaller performance metric standard deviations. Finally, we implement a multi-atlas label fusion (MALF) approach for pancreas segmentation using the same data set. Under six-fold cross-validation, our bottom-up segmentation method significantly outperforms its MALF counterpart: 70.7±13.0% versus 52.51±20.84% in dice coefficients. Amal Farag, Le Lu 0001, Holger Roth, Evrim Turkbey, Ronald M. Summers |
IEEE Trans. Image Process. | 2 |
| 2017 | DeepPap: Deep Convolutional Networks for Cervical Cell ClassificationabstractAutomation-assisted cervical screening via Pap smear or liquid-based cytology (LBC) is a highly effective cell imaging based cancer detection tool, where cells are partitioned into "abnormal" and "normal" categories. However, the success of most traditional classification methods relies on the presence of accurate cell segmentations. Despite sixty years of research in this field, accurate segmentation remains a challenge in the presence of cell clusters and pathologies. Moreover, previous classification methods are only built upon the extraction of hand-crafted features, such as morphology and texture. This paper addresses these limitations by proposing a method to directly classify cervical cells-without prior segmentation-based on deep features, using convolutional neural networks (ConvNets). First, the ConvNet is pretrained on a natural image dataset. It is subsequently fine-tuned on a cervical cell dataset consisting of adaptively resampled image patches coarsely centered on the nuclei. In the testing phase, aggregation is used to average the prediction scores of a similar set of image patches. The proposed method is evaluated on both Pap smear and LBC datasets. Results show that our method outperforms previous algorithms in classification accuracy (98.3%), area under the curve (0.99) values, and especially specificity (98.3%), when applied to the Herlev benchmark Pap smear dataset and evaluated using five-fold cross validation. Similar superior performances are also achieved on the HEMLBC (H&E stained manual LBC) dataset. Our method is promising for the development of automation-assisted reading systems in primary cervical screening. Ling Zhang 0002, Le Lu 0001, Isabella Nogues, Ronald M. Summers, Shaoxiong Liu, Jianhua Yao 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2016 | Learning to Read Chest X-Rays: Recurrent Neural Cascade Model for Automated Image AnnotationabstractDespite the recent advances in automatically describing image contents, their applications have been mostly limited to image caption datasets containing natural images (e.g., Flickr 30k, MSCOCO). In this paper, we present a deep learning model to efficiently detect a disease from an image and annotate its contexts (e.g., location, severity and the affected organs). We employ a publicly available radiology dataset of chest x-rays and their reports, and use its image annotations to mine disease names to train convolutional neural networks (CNNs). In doing so, we adopt various regularization techniques to circumvent the large normalvs-diseased cases bias. Recurrent neural networks (RNNs) are then trained to describe the contexts of a detected disease, based on the deep CNN features. Moreover, we introduce a novel approach to use the weights of the already trained pair of CNN/RNN on the domain-specific image/text dataset, to infer the joint image/text contexts for composite image labeling. Significantly improved image annotation results are demonstrated using the recurrent neural cascade model by taking the joint image/text contexts into account. Hoo-Chang Shin, Kirk Roberts, Le Lu 0001, Dina Demner-Fushman, Jianhua Yao 0001, Ronald M. Summers |
CVPR | 3 |
| 2016 | Pancreas Segmentation in MRI Using Graph-Based Decision Fusion on Convolutional Neural Networks
Jinzheng Cai, Le Lu 0001, Zizhao Zhang 0002, Fuyong Xing, Lin Yang 0002 |
MICCAI (2) | 2 |
| 2016 | Automatic Lymph Node Cluster Segmentation Using Holistically-Nested Neural Networks and Structured Optimization in CT Images
Isabella Nogues, Le Lu 0001, Xiaosong Wang 0001, Holger Roth, Gedas Bertasius, Nathan Lay, Jianbo Shi, Yohannes Tsehay, Ronald M. Summers |
MICCAI (2) | 2 |
| 2016 | Spatial Aggregation of Holistically-Nested Networks for Automated Pancreas Segmentation
Holger Roth, Le Lu 0001, Amal Farag, Andrew Sohn, Ronald M. Summers |
MICCAI (2) | 2 |
| 2016 | Accurate 3D bone segmentation in challenging CT images: Bottom-up parsing and contextualized optimizationabstractIn full or arbitrary field-of-view (FOV) 3D CT imaging, obtaining an accurate per-voxel segmentation for complete large and small bones remains an unsolved and challenging problem. The difficulty lies in the notable variation in appearance and position observed among cortical bones, marrow and pathologies. To approach this problem, several studies have employed active shape models and atlas models. In this paper, we argue that a bottom-up approach, defined by classifying and grouping supervoxels, is another viable technique. Moreover, it can be integrated into a conditional random field (CRF) representation. Our approach consists of the following steps: first, an input CT volume is decomposed into supervoxels, in order to ensure very high bone boundary recall. Supervoxels are generated via a robust process of conservative region partitioning and recursive region merging. In order to maximize sparsity and classification efficiency, we use a Bayesian sparse linear classifier to compute and optimize middle-level image features. Next, we disambiguate the CRF unary potentials via contextualized optimization by pooling over selective supervoxel pairs. Finally, we adopt a pairwise support vector machine (SVM) model to learn the CRF pairwise potential in a fully supervised manner. We evaluate our method quantitatively on 137 low-resolution, low-contrast CT volumes with severe imaging noise, among which various bone pathologies are represented. Our system proves to be efficient; it achieves a clinically significant segmentation accuracy level (Dice Coefficient 98.2%). Le Lu 0001, Dijia Wu, Nathan Lay, Isabella Nogues, Ronald M. Summers |
WACV | 1 |
| 2016 | Interleaved Text/Image Deep Mining on a Large-Scale Radiology Database for Automated Image InterpretationabstractDespite tremendous progress in computer vision, there has not been an attempt to apply machine learning on very large-scale medical image databases. We present an interleaved text/image deep learning system to extract and mine the semantic interactions of radiology images and reports from a national research hospital's Picture Archiving and Communication System. With natural language processing, we mine a collection of $\sim$216K representative two-dimensional images selected by clinicians for diagnostic reference and match the images with their descriptions in an automated manner. We then employ a weakly supervised approach using all of our available data to build models for generating approximate interpretations of patient images. Finally, we demonstrate a more strictly supervised approach to detect the presence and absence of a number of frequent disease types, providing more specific interpretations of patient scans. A relatively small amount of data is used for this part, due to the challenge in gathering quality labels from large raw text data. Our work shows the feasibility of large-scale learning and prediction in electronic patient records available in most modern clinical institutions. It also demonstrates the trade-offs to consider in designing machine learning systems for analyzing large medical data. Hoo-Chang Shin, Le Lu 0001, Lauren Kim, Ari Seff, Jianhua Yao 0001, Ronald M. Summers |
J. Mach. Learn. Res. | 2 |
| 2016 | Improving Computer-Aided Detection Using Convolutional Neural Networks and Random View AggregationabstractAutomated computer-aided detection (CADe) has been an important tool in clinical practice and research. State-of-the-art methods often show high sensitivities at the cost of high false-positives (FP) per patient rates. We design a two-tiered coarse-to-fine cascade framework that first operates a candidate generation system at sensitivities ∼ 100% of but at high FP levels. By leveraging existing CADe systems, coordinates of regions or volumes of interest (ROI or VOI) are generated and function as input for a second tier, which is our focus in this study. In this second stage, we generate 2D (two-dimensional) or 2.5D views via sampling through scale transformations, random translations and rotations. These random views are used to train deep convolutional neural network (ConvNet) classifiers. In testing, the ConvNets assign class (e.g., lesion, pathology) probabilities for a new set of random views that are then averaged to compute a final per-candidate classification probability. This second tier behaves as a highly selective process to reject difficult false positives while preserving high sensitivities. The methods are evaluated on three data sets: 59 patients for sclerotic metastasis detection, 176 patients for lymph node detection, and 1,186 patients for colonic polyp detection. Experimental results show the ability of ConvNets to generalize well to different medical imaging CADe applications and scale elegantly to various data sets. Our proposed methods improve performance markedly in all cases. Sensitivities improved from 57% to 70%, 43% to 77%, and 58% to 75% at 3 FPs per patient for sclerotic metastases, lymph nodes and colonic polyps, respectively. Holger Roth, Le Lu 0001, Jianhua Yao 0001, Ari Seff, Kevin M. Cherry, Lauren Kim, Ronald M. Summers |
IEEE Trans. Medical Imaging | 2 |
| 2016 | Deep Convolutional Neural Networks for Computer-Aided Detection: CNN Architectures, Dataset Characteristics and Transfer LearningabstractRemarkable progress has been made in image recognition, primarily due to the availability of large-scale annotated datasets and deep convolutional neural networks (CNNs). CNNs enable learning data-driven, highly representative, hierarchical image features from sufficient training data. However, obtaining datasets as comprehensively annotated as ImageNet in the medical imaging domain remains a challenge. There are currently three major techniques that successfully employ CNNs to medical image classification: training the CNN from scratch, using off-the-shelf pre-trained CNN features, and conducting unsupervised CNN pre-training with supervised fine-tuning. Another effective method is transfer learning, i.e., fine-tuning CNN models pre-trained from natural image dataset to medical image tasks. In this paper, we exploit three important, but previously understudied factors of employing deep convolutional neural networks to computer-aided detection problems. We first explore and evaluate different CNN architectures. The studied models contain 5 thousand to 160 million parameters, and vary in numbers of layers. We then evaluate the influence of dataset scale and spatial image context on performance. Finally, we examine when and why transfer learning from pre-trained ImageNet (via fine-tuning) can be useful. We study two specific computer-aided detection (CADe) problems, namely thoraco-abdominal lymph node (LN) detection and interstitial lung disease (ILD) classification. We achieve the state-of-the-art performance on the mediastinal LN detection, and report the first five-fold cross-validation classification results on predicting axial CT slices with ILD categories. Our extensive empirical evaluation, CNN model analysis and valuable insights can be extended to the design of high performance CAD systems for other medical imaging tasks. Hoo-Chang Shin, Holger Roth, Mingchen Gao, Le Lu 0001, Ziyue Xu 0001, Isabella Nogues, Jianhua Yao 0001, Daniel J. Mollura, Ronald M. Summers |
IEEE Trans. Medical Imaging | 4 |
| 2015 | Interleaved text/image Deep Mining on a large-scale radiology databaseabstractDespite tremendous progress in computer vision, effective learning on very large-scale (> 100K patients) medical image databases has been vastly hindered. We present an interleaved text/image deep learning system to extract and mine the semantic interactions of radiology images and reports from a national research hospital's picture archiving and communication system. Instead of using full 3D medical volumes, we focus on a collection of representative ~216K 2D key images/slices (selected by clinicians for diagnostic reference) with text-driven scalar and vector labels. Our system interleaves between unsupervised learning (e.g., latent Dirichlet allocation, recurrent neural net language models) on document- and sentence-level texts to generate semantic labels and supervised learning via deep convolutional neural networks (CNNs) to map from images to label spaces. Disease-related key words can be predicted for radiology images in a retrieval manner. We have demonstrated promising quantitative and qualitative results. The large-scale datasets of extracted key images and their categorization, embedded vector labels and sentence descriptions can be harnessed to alleviate the deep learning “data-hungry” obstacle in the medical domain. Hoo-Chang Shin, Le Lu 0001, Lauren Kim, Ari Seff, Jianhua Yao 0001, Ronald M. Summers |
CVPR | 2 |
| 2015 | DeepOrgan: Multi-level Deep Convolutional Networks for Automated Pancreas Segmentation
Holger Roth, Le Lu 0001, Amal Farag, Hoo-Chang Shin, Evrim Turkbey, Ronald M. Summers |
MICCAI (1) | 2 |
| 2015 | Leveraging Mid-Level Semantic Boundary Cues for Automated Lymph Node Detection
Ari Seff, Le Lu 0001, Adrian Barbu, Holger Roth, Hoo-Chang Shin, Ronald M. Summers |
MICCAI (2) | 2 |
| 2015 | Sequential Monte Carlo tracking of the marginal artery by multiple cue fusion and random forest regression
Kevin M. Cherry, Brandon Peplinski, Lauren Kim, Le Lu 0001, Weidong Zhang 0001, Zhuoshi Wei, Ronald M. Summers |
Medical Image Anal. | 5 |
| 2015 | Automatic Segmentation of Spinal Canals in CT Images via Iterative Topology RefinementabstractAccurate segmentation of the spinal canals in computed tomography (CT) images is an important task in many related studies. In this paper, we propose an automatic segmentation method and apply it to our highly challenging image cohort that is acquired from multiple clinical sites and from the CT channel of the PET-CT scans. To this end, we adapt the interactive random-walk solvers to be a fully automatic cascaded pipeline. The automatic segmentation pipeline is initialized with robust voxelwise classification using Haar-like features and probabilistic boosting tree. Then, the topology of the spinal canal is extracted from the tentative segmentation and further refined for the subsequent random-walk solver. In particular, the refined topology leads to improved seeding voxels or boundary conditions, which allow the subsequent random-walk solver to improve the segmentation result. Therefore, by iteratively refining the spinal canal topology and cascading the random-walk solvers, satisfactory segmentation results can be acquired within only a few iterations, even for cases with scoliosis, bone fractures and lesions. Our experiments validate the capability of the proposed method with promising segmentation performance, even though the resolution and the contrast of our dataset with 110 patient cases (90 for testing and 20 for training) are low and various bone pathologies occur frequently. Qian Wang 0001, Le Lu 0001, Dijia Wu, Noha Youssry El-Zehiry, Yefeng Zheng 0001, Dinggang Shen, Shaohua Kevin Zhou |
IEEE Trans. Medical Imaging | 2 |
| 2014 | A New 2.5D Representation for Lymph Node Detection Using Random Sets of Deep Convolutional Neural Network Observations
Holger Roth, Le Lu 0001, Ari Seff, Kevin M. Cherry, Joanne Hoffman, Evrim Turkbey, Ronald M. Summers |
MICCAI (1) | 2 |
| 2014 | 2D View Aggregation for Lymph Node Detection Using a Shallow Hierarchy of Linear Classifiers
Ari Seff, Le Lu 0001, Kevin M. Cherry, Holger Roth, Joanne Hoffman, Evrim Turkbey, Ronald M. Summers |
MICCAI (1) | 2 |
| 2013 | Sequential Monte Carlo Tracking for Marginal Artery Segmentation on CT Angiography by Multiple Cue Fusion
Brandon Peplinski, Le Lu 0001, Weidong Zhang 0001, Zhuoshi Wei, Ronald M. Summers |
MICCAI (2) | 3 |
| 2013 | Hierarchical segmentation and identification of thoracic vertebra using learning-based edge detection and coarse-to-fine deformable model
Le Lu 0001 |
Comput. Vis. Image Underst. | 2 |
| 2012 | Robust Object Tracking in Crowd Dynamic Scenes Using Explicit Stereo Depth
Le Lu 0001, Gregory D. Hager, Jianyu Tang, Hanzi Wang |
ACCV (3) | 2 |
| 2012 | Inferring gene regulatory networks from gene expression data by path consistency algorithm based on conditional mutual informationabstractMOTIVATION: Reconstruction of gene regulatory networks (GRNs), which explicitly represent the causality of developmental or regulatory process, is of utmost interest and has become a challenging computational problem for understanding the complex regulatory mechanisms in cellular systems. However, all existing methods of inferring GRNs from gene expression profiles have their strengths and weaknesses. In particular, many properties of GRNs, such as topology sparseness and non-linear dependence, are generally in regulation mechanism but seldom are taken into account simultaneously in one computational method. RESULTS: In this work, we present a novel method for inferring GRNs from gene expression data considering the non-linear dependence and topological structure of GRNs by employing path consistency algorithm (PCA) based on conditional mutual information (CMI). In this algorithm, the conditional dependence between a pair of genes is represented by the CMI between them. With the general hypothesis of Gaussian distribution underlying gene expression data, CMI between a pair of genes is computed by a concise formula involving the covariance matrices of the related gene expression profiles. The method is validated on the benchmark GRNs from the DREAM challenge and the widely used SOS DNA repair network in Escherichia coli. The cross-validation results confirmed the effectiveness of our method (PCA-CMI), which outperforms significantly other previous methods. Besides its high accuracy, our method is able to distinguish direct (or causal) interactions from indirect associations. AVAILABILITY: All the source data and code are available at: http://csb.shu.edu.cn/subweb/grn.htm. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xing-Ming Zhao, Kun He 0007, Le Lu 0001, Yongwei Cao, Jingdong Liu, Jin-Kao Hao, Zhi-Ping Liu, Luonan Chen |
Bioinform. | 4 |
| 2011 | Coarse-to-fine classification via parametric and nonparametric models for computer-aided diagnosisabstractClassification is one of the core problems in Computer-Aided Diagnosis (CAD), targeting for early cancer detection using 3D medical imaging interpretation. High detection sensitivity with desirably low false positive (FP) rate is critical for a CAD system to be accepted as a valuable or even indispensable tool in radiologists' workflow. Given various spurious imagery noises which cause observation uncertainties, this remains a very challenging task. In this paper, we propose a novel, two-tiered coarse-to-fine (CTF) classification cascade framework to tackle this problem. We first obtain classification-critical data samples (e.g., implicit samples on the decision boundary) extracted from the holistic data distributions using a robust parametric model (e.g., [13]); then we build a graph-embedding based nonparametric classifier on sampled data, which can more accurately preserve or formulate the complex classification boundary. These two steps can also be considered as effective "sample pruning" and "feature pursuing + kNN/template matching", respectively. Our approach is validated comprehensively in colorectal polyp detection and lung nodule detection CAD systems, as the top two deadly cancers, using hospital scale, multi-site clinical datasets. The results show that our method achieves overall better classification/detection performance than existing state-of-the-art algorithms using single-layer classifiers, such as the support vector machine variants [17], boosting [15], logistic regression [11], relevance vector machine [13], k-nearest neighbor [9] or spectral projections on graph [2]. Meizhu Liu, Le Lu 0001, Xiaojing Ye, Shipeng Yu, Heng Huang 0001 |
CIKM | 2 |
| 2011 | AdaBoost on low-rank PSD matrices for metric learningabstractThe problem of learning a proper distance or similarity metric arises in many applications such as content-based image retrieval. In this work, we propose a boosting algorithm, MetricBoost, to learn the distance metric that preserves the proximity relationships among object triplets: object i is more similar to object j than to object k. Metric-Boost constructs a positive semi-definite (PSD) matrix that parameterizes the distance metric by combining rank-one PSD matrices. Different options of weak models and combination coefficients are derived. Unlike existing proximity preserving metric learning which is generally not scalable, MetricBoost employs a bipartite strategy to dramatically reduce computation cost by decomposing proximity relationships over triplets into pair-wise constraints. Met-ricBoost outperforms the state-of-the-art on two real-world medical problems: 1. identifying and quantifying diffuse lung diseases; 2. colorectal polyp matching between different views, as well as on other benchmark datasets. Jinbo Bi, Dijia Wu, Le Lu 0001, Meizhu Liu, Yimo Tao, Matthias Wolf 0001 |
CVPR | 3 |
| 2011 | Effective 3D object detection and regression using probabilistic segmentation features in CT imagesabstract3D object detection and importance regression/ranking are at the core for semantically interpreting 3D medical images of computer aided diagnosis (CAD). In this paper, we propose effective image segmentation features and a novel multiple instance regression method for solving the above challenges. We perform supervised learning based segmentation algorithm on numerous lesion candidates (as 3D VOIs: Volumes Of Interest in CT images) which can be true or false. By assessing the statistical properties in the joint space of segmentation output (e.g., a 3D class-specific probability map or cloud), and original image appearance, 57 descriptive features in six subgroups are derived. The new feature set shows excellent performance on effectively classifying ambiguous positive and negative VOIs, for our CAD system of detecting colonic polyps using CT images. The proposed regression model on our segmentation derived features behaves as a robust object (polyp) size/importance estimator and ranking module with high reliability, which is critical for automatic clinical reporting and cancer staging. Extensive evaluation is executed on a large clinical dataset of 770 CT scans from 12 medical sites for validation, with the best state-of-the-art results. Le Lu 0001, Jinbo Bi, Matthias Wolf 0001, Marcos Salganicoff |
CVPR | 1 |
| 2011 | Robust Large Scale Prone-Supine Polyp Matching Using Local Features: A Metric Learning Approach
Meizhu Liu, Le Lu 0001, Jinbo Bi, Vikas C. Raykar, Matthias Wolf 0001, Marcos Salganicoff |
MICCAI (3) | 2 |
| 2011 | Sparse Classification for Computer Aided Diagnosis Using Learned Dictionaries
Meizhu Liu, Le Lu 0001, Xiaojing Ye, Shipeng Yu, Marcos Salganicoff |
MICCAI (3) | 2 |
| 2010 | Stratified learning of local anatomical context for lung nodules in CT imagesabstractThe automatic detection of lung nodules attached to other pulmonary structures is a useful yet challenging task in lung CAD systems. In this paper, we propose a stratified statistical learning approach to recognize whether a candidate nodule detected in CT images connects to any of three other major lung anatomies, namely vessel, fissure and lung wall, or is solitary with background parenchyma. First, we develop a fully automated voxel-by-voxel labeling/segmentation method of nodule, vessel, fissure, lung wall and parenchyma given a 3D lung image, via a unified feature set and classifier under conditional random field. Second, the generated Class Probability Response Maps (PRM) by voxel-level classifiers, are used to form the so-called pairwise Probability Co-occurrence Maps (PCM) which encode the spatial contextual correlations of the candidate nodule, in relation to other anatomical landmarks. Based on PCMs, higher level classifiers are trained to recognize whether the nodule touches other pulmonary structures, as a multi-label problem. We also present a new iterative fissure structure enhancement filter with superior performance. For experimental validation, we create an annotated database of 784 subvolumes with nodules of various sizes, shapes, densities and contextual anatomies, and from 239 patients. High accuracy of multi-class voxel labeling is achieved 89.3% ∼ 91.2%. The Area under ROC Curve (AUC) of vessel, fissure and lung wall connectivity classification reaches 0.8676, 0.8692 and 0.9275, respectively. Dijia Wu, Le Lu 0001, Jinbo Bi, Yoshihisa Shinagawa, Kim L. Boyer, Arun Krishnan, Marcos Salganicoff |
CVPR | 2 |
| 2010 | Hierarchical Segmentation and Identification of Thoracic Vertebra Using Learning-Based Edge Detection and Coarse-to-Fine Deformable Model
Le Lu 0001, Yiqiang Zhan, Xiang Sean Zhou, Marcos Salganicoff, Arun Krishnan |
MICCAI (1) | 2 |
| 2009 | Hierarchical learning for tubular structure parsing in medical imaging: A study on coronary arteries using 3D CT AngiographyabstractAutomatic coronary artery centerline extraction from 3D CT Angiography (CTA) has significant clinical importance for diagnosis of atherosclerotic heart disease. The focus of past literature is dominated by segmenting the complete coronary artery system as trees by computer. Though the labeling of different vessel branches (defined by their medical semantics) is much needed clinically, this task has been performed manually. In this paper, we propose a hierarchical machine learning approach to tackle the problem of tubular structure parsing in medical imaging. It has a progressive three-tiered classification process at volumetric voxel level, vessel segment level, and inter-segment level. Generative models are employed to project from low-level, ambiguous data to class-conditional probabilities; and discriminative classifiers are trained on the upper-level structural patterns of probabilities to label and parse the vessel segments. Our method is validated by experiments of detecting and segmenting clinically defined coronary arteries, from the initial noisy vessel segment networks generated by low-level heuristics-based tracing algorithms. The proposed framework is also generically applicable to other tubular structure parsing tasks. Le Lu 0001, Jinbo Bi, Shipeng Yu, Zhigang Peng, Arun Krishnan, Xiang Sean Zhou |
ICCV | 1 |
| 2009 | A Two-Level Approach Towards Semantic Colon Segmentation: Removing Extra-Colonic Findings
Le Lu 0001, Matthias Wolf 0001, Jianming Liang, Murat Dundar, Jinbo Bi, Marcos Salganicoff |
MICCAI (1) | 1 |
| 2009 | Multi-level Ground Glass Nodule Detection and Segmentation in CT Lung Images
Yimo Tao, Le Lu 0001, Maneesh Dewan, Albert Y. Chen, Jason J. Corso, Jianhua Xuan, Marcos Salganicoff, Arun Krishnan |
MICCAI (1) | 2 |
| 2008 | Accurate polyp segmentation for 3D CT colongraphy using multi-staged probabilistic binary learning and compositional modelabstractAccurate and automatic colonic polyp segmentation and measurement in Computed Tomography (CT) has significant importance for 3D polyp detection, classification, and more generally computer aided diagnosis of colon cancers. In this paper, we propose a three-staged probabilistic binary classification approach for automatically segmenting polyp voxels from their surrounding tissues in CT. Our system integrates low-, and mid-level information for discriminative learning under local polar coordinates which align on the 3D colon surface around detected polyp. More importantly, our supervised learning system has flexible modeling capacity, which offers a principled means of encoding semantic, clinical expert annotations of colonic polyp tissue identification and segmentation. The learning generality to unseen data is bounded by boosting [12, 11] and stacked generality [14]. Extensive experimental results on polyp segmentation performance evaluation and robustness testing with disturbances (using both training data and unseen data) are provided to validate our presented approach. The reliability of polyp segmentation and measurement has been largely increased to 98:2% (ie. errors les 3 mm), compared with other state of art work [4, 15] of about 75% ~ 80%. Le Lu 0001, Adrian Barbu, Matthias Wolf 0001, Jianming Liang, Marcos Salganicoff, Dorin Comaniciu |
CVPR | 1 |
| 2008 | Simultaneous Detection and Registration for Ileo-Cecal Valve Detection in 3D CT Colonography
Le Lu 0001, Adrian Barbu, Matthias Wolf 0001, Jianming Liang, Luca Bogoni, Marcos Salganicoff, Dorin Comaniciu |
ECCV (4) | 1 |
| 2007 | A Nonparametric Treatment for Location/Segmentation Based Visual TrackingabstractIn this paper, we address two closely related visual tracking problems: 1) localizing a target's position in low or moderate resolution videos and 2) segmenting a target's image support in moderate to high resolution videos. Both tasks are treated as an online binary classification problem using dynamic foreground/background appearance models. Our major contribution is a novel nonparametric approach that successfully maintains a temporally changing appearance model for both foreground and background. The appearance models are formulated as "bags of image patches" that approximate the true two-class appearance distributions. They are maintained using a temporal-adaptive importance resampling procedure that is based on simple nonparametric statistics of the appearance patch bags. The overall framework is independent of an specific foreground/background classification process and thus offers the freedom to use different classifiers. We demonstrate the effectiveness of our approach with extensive comparative experimental results on sequences from previous visual tracking [1, 12] and video matting [4] work as well as our own data. Le Lu 0001, Gregory D. Hager |
CVPR | 1 |
| 2006 | Combined central and subspace clustering for computer vision applicationsabstractCentral and subspace clustering methods are at the core of many segmentation problems in computer vision. However, both methods fail to give the correct segmentation in many practical scenarios, e.g., when data points are close to the intersection of two subspaces or when two cluster centers in different subspaces are spatially close. In this paper, we address these challenges by considering the problem of clustering a set of points lying in a union of subspaces and distributed around multiple cluster centers inside each subspace. We propose a generalization of Kmeans and Ksubspaces that clusters the data by minimizing a cost function that combines both central and subspace distances. Experiments on synthetic data compare our algorithm favorably against four other clustering methods. We also test our algorithm on computer vision problems such as face clustering with varying illumination and video shot segmentation of dynamic scenes. Le Lu 0001, René Vidal |
ICML | 1 |
| 2006 | Dynamic Foreground/Background Extraction from Images and Videos using Random PatchesabstractIn this paper, we propose a novel exemplar-based approach to extract dynamic foreground regions from a changing background within a collection of images or a video sequence. By using image segmentation as a pre-processing step, we convert this traditional pixel-wise labeling problem into a lower-dimensional supervised, binary labeling procedure on image segments. Our approach consists of three steps. First, a set of random image patches are spatially and adaptively sampled within each segment. Second, these sets of extracted samples are formed into two "bags of patches" to model the foreground/background appearance, respectively. We perform a novel bidirectional consistency check between new patches from incoming frames and current "bags of patches" to reject outliers, control model rigidity and make the model adaptive to new observations. Within each bag, image patches are further partitioned and resampled to create an evolving appearance model. Finally, the foreground/background decision over segments in an image is formulated using an aggregation function defined on the similarity measurements of sampled patches relative to the foreground and background models. The essence of the algorithm is conceptually simple and can be easily implemented within a few hundred lines of Matlab code. We evaluate and validate the proposed approach by extensive real examples of the object-level image mapping and tracking within a variety of challenging environments. We also show that it is straightforward to apply our problem formulation on non-rigid object tracking with difficult surveillance videos. Le Lu 0001, Gregory D. Hager |
NIPS | 1 |
| 2006 | Efficient particle filtering using RANSAC with application to 3D face tracking
Le Lu 0001, Xiangtian Dai, Gregory D. Hager |
Image Vis. Comput. | 1 |
| 2005 | A Two Level Approach for Scene RecognitionabstractClassifying pictures into one of several semantic categories is a classical image understanding problem. In this paper, we present a stratified approach to both binary (outdoor-indoor) and multiple category of scene classification. We first learn mixture models for 20 basic classes of local image content based on color and texture information. Once trained, these models are applied to a test image, and produce 20 probability density response maps (PDRM) indicating the likelihood that each image region was produced by each class. We then extract some very simple features from those PDRMs, and use them to train a bagged LDA classifier for 10 scene categories. For this process, no explicit region segmentation or spatial context model are computed. To test this classification system, we created a labeled database of 1500 photos taken under very different environment and lighting conditions, using different cameras, and from 43 persons over 5 years. The classification rate of outdoor-indoor classification is 93.8%, and the classification rate for 10 scene categories is 90.1%. As a byproduct, local image patches can be contextually labeled into the 20 basic material classes by using loopy belief propagation (Yedidia et al., 2001) as an anisotropic filter on PDRMs, producing an image-level segmentation if desired. Le Lu 0001, Kentaro Toyama, Gregory D. Hager |
CVPR (1) | 1 |
| 2004 | A Three Tiered Approach for Articulated Object Action Modeling and RecognitionabstractVisual action recognition is an important problem in computer vision. In this paper, we propose a new method to probabilistically model and recognize actions of articulated objects, such as hand or body gestures, in image sequences. Our method consists of three levels of representa- tion. At the low level, we first extract a feature vector invariant to scale and in-plane rotation by using the Fourier transform of a circular spatial histogram. Then, spectral partitioning [20] is utilized to obtain an initial clustering; this clustering is then refined using a temporal smoothness constraint. Gaussian mixture model (GMM) based clustering and density estimation in the subspace of linear discriminant analysis (LDA) are then applied to thousands of image feature vectors to obtain an intermediate level representation. Finally, at the high level we build a temporal multi- resolution histogram model for each action by aggregating the clustering weights of sampled images belonging to that action. We discuss how this high level representation can be extended to achieve temporal scaling in- variance and to include Bi-gram or Multi-gram transition information. Both image clustering and action recognition/segmentation results are given to show the validity of our three tiered representation. Le Lu 0001, Gregory D. Hager, Laurent Younes |
NIPS | 1 |
| 2004 | Constrained planar motion analysis by decomposition
Long Quan, Le Lu 0001, Harry Shum |
Image Vis. Comput. | 3 |
| 2001 | Concentric Mosaic(s)Planar Motion and 1D CamerasabstractGeneral SFM methods give poor results for images captured by constrained motions such as planar motion of concentric mosaics (CM). In this paper we propose new SFM algorithms for both images captured by CM and composite mosaic images from CM. We first introduce ID affine camera model for completing 1D camera models. Then we show that a 2D image captured by CM can be decoupled into two 1D images: one 1D projective and one ID affine; a composite mosaic image can by rebinned into a calibrated ID panorama projective camera. Finally we describe subspace reconstruction methods and demonstrate both in theory and experiments the advantage of the decomposition method over the general SFM methods by incorporating the constrained motion into the earliest stage of motion analysis. Long Quan, Le Lu 0001, Harry Shum, Maxime Lhuillier |
ICCV | 2 |
| 2000 | A Novel Method for Camera Planar Motion Detection and Robust Estimation of the 1D Trifocal TensorabstractA camera moving in a plane can often simplify a computer vision job. Camera self-calibration and robot self-location are good examples. We focus on the problem of camera planar motion and its application to the camera self-calibration method of Faugeras et al. (1998). We have made three new contributions to the camera planar motion detection. First, we prove that the trifocal lines in different views of the same planar motion must have the same line representation in the 2D retinal plane. This conclusion greatly simplifies the planar motion detection problem. Second, we distinguish the usage of three different cases of planar motion: ordinary planar motion, co-linear planar motion and rotation planar motion. Third, we propose the robust planar motion detection method and the method of estimation of trifocal lines in the uniform framework under the above three configurations. We have also purposed a method for eliminating the 2D image points whose 1D projection points are inaccurate and cause significant errors on the estimation of the 1D trifocal tensor. Experiments with our new techniques using simulated data and real images had obtained very good results, which are better than those reported in the above article. Le Lu 0001, Hung-Tat Tsui, Zhanyi Hu |
ICPR | 1 |