VLDB 2026 Research / reviewers in the wild / expert
Lei Xing 0001
dblp:82/2022-1
· DBLP profile ↗
55ranked-venue papers
0as first author
31since 2021 · last 2026
0000-0003-2536-5359ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 33 · 14 since 2021Artificial intelligence and machine learning · 16 · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 since 2021Databases, data management, data science and information retrieval · 3Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-contrast low-field MRI acceleration with k-space progressive learning and image-space hybrid attention fusion
Xiaohan Xing, Qi Chen 0014, Lequan Yu, Lingting Zhu, Lei Xing 0001, Lianli Liu |
Medical Image Anal. | 6 |
| 2026 | Advancing In-Context Learning for Efficient and Stable Medical Report GenerationabstractVision-language models (VLMs) have shown strong generalization across multimodal tasks, but adapting them to medical report generation (MRG) often demands extensive paired image-text data that are limited due to data privacy and annotation cost. In-context learning (ICL) offers a promising training-free alternative, yet standard ICL approaches rely on long demonstration prompts that are computationally inefficient and often yield inconsistent or clinically inaccurate descriptions. To address these challenges, we propose Principal In-Context Vectors (PCVs), a compact latent-guidance framework that distills multimodal demonstrations into stable semantic representations. By extracting hidden states from auto-regressive VLMs and applying principal component analysis (PCA), we identify robust semantic directions that remain stable under input perturbations. These PCVs are then injected into new queries to steer generation toward accurate and clinically meaningful outputs without any model tuning. Extensive experiments on four MRG benchmark datasets show that our approach can enhance both zero-shot and fully supervised generation quality across diverse settings, including cross-center, cross-disease, and longitudinal scenarios. This work provides a lightweight and scalable approach to adapt pre-trained VLMs for practical clinical deployment. Mingjie Li 0006, Zeyi Shi, Mingfei Han 0002, Lina Yao 0001, Zhihui Li 0001, Xiaojun Chang, Kilian M. Pohl, Md Tauhidul Islam, Lei Xing 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 10 |
| 2026 | Knowledge-guided multi-modality transformer for multi-label genetic mutation prediction
Gexin Huang, Chenfei Wu, Mingjie Li 0006, Xiaojun Chang, Ying Sun 0001, Lei Xing 0001, Xiaodan Liang, Liang Lin 0004 |
Pattern Recognit. | 6 |
| 2026 | Enhancing Deep Learning Inference of Gene Regulatory Networks via Construction of Image Representation of Cell-Cell Interactions From scRNA-Seq DataabstractUnderstanding gene regulatory networks (GRNs) holds paramount importance for deciphering the intricate interplay among genes and their influence on biological processes and disease pathogenesis. The emergence of single-cell RNA sequencing (scRNA-seq) techniques has heralded a new era in GRN inference by capturing the nuanced heterogeneity and dynamic nature of gene expression at the single-cell level. However, extracting meaningful patterns from scRNA-seq measurements to infer GRNs poses significant challenges to existing methodologies due to the sheer scale and inherent complexity of the data. Here we propose a highly accurate and computationally efficient strategy for scRNA-seq-based GRN inference. Our approach leverages the underlying interactive relationships among the cells using state-of-the-art deep learning strategy. Specifically, a spatially semantic image representation, termed CelloGraph, is first introduced to portray the expressions of each gene across cells. The allocation of a cell to a spatial grid point of the CelloGraph is dictated by its interactions with other cells within the system, as determined by the maximization of system entropy of cell-cell interactions. Subsequently, the CelloGraphs of all pertinent genes are analyzed by using a customarily designed convolutional neural network (CNN) to discern discriminant patterns in the data and infer GRNs. The efficacy of the proposed approach is demonstrated through diverse real-world biomedical datasets. By harnessing the distinctive attributes of spatially semantic CelloGraphs and leveraging the unique pattern discovery capabilities of CNNs, our methodology paves the way for a deeper comprehension of the underlying mechanisms that govern gene expression and regulation. The proposed strategy not only overcomes challenges in scRNA-seq-based GRN inference but also promises to provide a more comprehensive understanding of intricate biological processes. Qingyue Wei, Md Tauhidul Islam, Wei Emma Wu, Lei Xing 0001 |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | MedTrinity-25M: A Large-scale Multimodal Dataset with Multigranular Annotations for MedicineabstractThis paper introduces MedTrinity-25M, a comprehensive, large-scale multimodal dataset for medicine, covering over 25 million images across 10 modalities with multigranular annotations for more than 65 diseases. These multigranular annotations encompass both global information, such as modality and organ detection, and local information like ROI analysis, lesion texture, and region-wise correlations. Unlike the existing multimodal datasets, which are limited by the availability of image-text pairs, we have developed the first automated pipeline that scales up multimodal data by generating multigranular visual and textual annotations in the form of image-ROI-description triplets without the need for any paired text descriptions. Specifically, data from over 30 different sources have been collected, preprocessed, and grounded using domain-specific expert models to identify ROIs related to abnormal regions. We then build a comprehensive knowledge base and prompt multimodal large language models to perform retrieval-augmented generation with the identified ROIs as guidance, resulting in multigranular textual descriptions. Compared to existing datasets, MedTrinity-25M provides the most enriched annotations, supporting a comprehensive range of multimodal tasks such as captioning and report generation, as well as vision-centric tasks like classification and segmentation. We propose LLaVA-Tri by pretraining LLaVA on MedTrinity-25M, achieving state-of-the-art performance on VQA-RAD, SLAKE, and PathVQA, surpassing representative SOTA multimodal large language models. Furthermore, MedTrinity-25M can also be utilized to support large-scale pre-training of multimodal medical AI models, contributing to the development of future foundation models in the medical domain. We will make our dataset available. The dataset is publicly available at https://yunfeixie233.github.io/MedTrinity-25M/. Yunfei Xie, Ce Zhou, Lang Gao, Juncheng Wu, Xianhang Li, Lei Xing 0001, James Zou 0001, Cihang Xie, Yuyin Zhou |
ICLR | 8 |
| 2025 | MS-Glance: Bio-Inspired Non-Semantic Context Vectors and Their Applications in Supervising Image ReconstructionabstractNon-semantic context information is crucial for visual recognition, as the human visual perception system first uses global statistics to process scenes rapidly before identifying specific objects. However, while semantic information is increasingly incorporated into computer vision tasks such as image reconstruction, non-semantic information, such as global spatial structures, is often overlooked. To bridge the gap, we propose a biologically informed non-semantic context descriptor, MS-Glance, along with the Glance Index Measure for comparing two images. A Global Glance vector is formulated by randomly retrieving pixels based on a perception-driven rule from an image to form a vector representing non-semantic global context, while a local Glance vector is a flattened local image window, mimicking a zoom in observation. The Glance Index is defined as the inner product of two standardized sets of Glance vectors. We evaluate the effectiveness of incorporating Glance supervision in two reconstruction tasks: image fitting with implicit neural representation (INR) and undersampled MRI reconstruction. Extensive experimental results show that MS-Glance outperforms existing image restoration losses across both natural and medical images. The code is available at https://github.com/Z7Gao/MSGlance. Wendi Yang, Lei Xing 0001, Shaohua Kevin Zhou |
WACV | 4 |
| 2025 | Multi-Sensor Learning Enables Information Transfer Across Different Sensory Data and Augments Multi-Modality ImagingabstractMulti-modality imaging is widely used in clinical practice and biomedical research to gain a comprehensive understanding of an imaging subject. Currently, multi-modality imaging is accomplished by post hoc fusion of independently reconstructed images under the guidance of mutual information or spatially registered hardware, which limits the accuracy and utility of multi-modality imaging. Here, we investigate a data-driven multi-modality imaging (DMI) strategy for synergetic imaging of CT and MRI. We reveal two distinct types of features in multi-modality imaging, namely intra- and inter-modality features, and present a multi-sensor learning (MSL) framework to utilize the crossover inter-modality features for augmented multi-modality imaging. The MSL imaging approach breaks down the boundaries of traditional imaging modalities and allows for optimal hybridization of CT and MRI, which maximizes the use of sensory data. We showcase the effectiveness of our DMI strategy through synergetic CT-MRI brain imaging. The principle of DMI is quite general and holds enormous potential for various DMI applications across disciplines. Lingting Zhu, Lianli Liu, Lei Xing 0001, Lequan Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | L2B: Learning to Bootstrap Robust Models for Combating Label NoiseabstractDeep neural networks have shown great success in representation learning. However, when learning with noisy labels (LNL), they can easily overfit and fail to generalize to new data. This paper introduces a simple and effective method, named Learning to Bootstrap (L2B), which enables models to bootstrap themselves using their own predictions without being adversely affected by erroneous pseudo-labels. It achieves this by dynamically adjusting the importance weight between real observed and generated labels, as well as between different samples through metalearning. Unlike existing instance reweighting methods, the key to our method lies in a new, versatile objective that enables implicit relabeling concurrently, leading to significant improvements without incurring additional costs. L2B offers several benefits over the baseline methods. It yields more robust models that are less susceptible to the impact of noisy labels by guiding the bootstrapping procedure more effectively. It better exploits the valuable information contained in corrupted instances by adapting the weights of both instances and labels. Furthermore, L2B is compatible with existing LNL methods and delivers competitive results spanning natural and medical imaging tasks including classification and segmentation under both synthetic and real-world noise. Extensive experiments demonstrate that our method effectively mitigates the challenges of noisy labels, often necessitating few to no validation samples, and is well generalized to other tasks such as image segmentation. This not only positions it as a robust complement to existing LNL techniques but also underscores its practical applicability. The code and models are available at https://github.com/yuyinzhou/12b. Yuyin Zhou, Xianhang Li, Fengze Liu, Qingyue Wei, Xuxi Chen, Lequan Yu, Cihang Xie, Matthew P. Lungren, Lei Xing 0001 |
CVPR | 9 |
| 2024 | In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space SteeringabstractLarge language models (LLMs) demonstrate emergent in-context learning capabilities, where they adapt to new tasks based on example demonstrations. However, in-context learning has seen limited effectiveness in many settings, is difficult to quantitatively control and takes up context window space. To overcome these limitations, we propose an alternative approach that recasts in-context learning as in-context vectors (ICV). Using ICV has two steps. We first use a forward pass on demonstration examples to create the in-context vector from the latent embedding of the LLM. This vector captures essential information about the intended task. On a new query, instead of adding demonstrations to the prompt, we shift the latent states of the LLM using the ICV. The ICV approach has several benefits: 1) it enables the LLM to more effectively follow the demonstration examples; 2) it’s easy to control by adjusting the magnitude of the ICV; 3) it reduces the length of the prompt by removing the in-context demonstrations; 4) ICV is computationally much more efficient than fine-tuning. We demonstrate that ICV achieves better performance compared to standard in-context learning and fine-tuning on diverse tasks including safety, style transfer, role-playing and formatting. Moreover, we show that we can flexibly teach LLM to simultaneously follow different types of instructions by simple vector arithmetics on the corresponding ICVs. Haotian Ye, Lei Xing 0001, James Zou 0001 |
ICML | 3 |
| 2024 | Self-supervised deep learning of gene-gene interactions for improved gene expression recoveryabstractSingle-cell RNA sequencing (scRNA-seq) has emerged as a powerful tool to gain biological insights at the cellular level. However, due to technical limitations of the existing sequencing technologies, low gene expression values are often omitted, leading to inaccurate gene counts. Existing methods, including advanced deep learning techniques, struggle to reliably impute gene expressions due to a lack of mechanisms that explicitly consider the underlying biological knowledge of the system. In reality, it has long been recognized that gene-gene interactions may serve as reflective indicators of underlying biology processes, presenting discriminative signatures of the cells. A genomic data analysis framework that is capable of leveraging the underlying gene-gene interactions is thus highly desirable and could allow for more reliable identification of distinctive patterns of the genomic data through extraction and integration of intricate biological characteristics of the genomic data. Here we tackle the problem in two steps to exploit the gene-gene interactions of the system. We first reposition the genes into a 2D grid such that their spatial configuration reflects their interactive relationships. To alleviate the need for labeled ground truth gene expression datasets, a self-supervised 2D convolutional neural network is employed to extract the contextual features of the interactions from the spatially configured genes and impute the omitted values. Extensive experiments with both simulated and experimental scRNA-seq datasets are carried out to demonstrate the superior performance of the proposed strategy against the existing imputation methods. Qingyue Wei, Md Tauhidul Islam, Yuyin Zhou, Lei Xing 0001 |
Briefings Bioinform. | 4 |
| 2024 | TransUNet: Rethinking the U-Net architecture design for medical image segmentation through the lens of transformersabstractMedical image segmentation is crucial for healthcare, yet convolution-based methods like U-Net face limitations in modeling long-range dependencies. To address this, Transformers designed for sequence-to-sequence predictions have been integrated into medical image segmentation. However, a comprehensive understanding of Transformers' self-attention in U-Net components is lacking. TransUNet, first introduced in 2021, is widely recognized as one of the first models to integrate Transformer into medical image analysis. In this study, we present the versatile framework of TransUNet that encapsulates Transformers' self-attention into two key modules: (1) a Transformer encoder tokenizing image patches from a convolution neural network (CNN) feature map, facilitating global context extraction, and (2) a Transformer decoder refining candidate regions through cross-attention between proposals and U-Net features. These modules can be flexibly inserted into the U-Net backbone, resulting in three configurations: Encoder-only, Decoder-only, and Encoder+Decoder. TransUNet provides a library encompassing both 2D and 3D implementations, enabling users to easily tailor the chosen architecture. Our findings highlight the encoder's efficacy in modeling interactions among multiple abdominal organs and the decoder's strength in handling small targets like tumors. It excels in diverse medical applications, such as multi-organ segmentation, pancreatic tumor segmentation, and hepatic vessel segmentation. Notably, our TransUNet achieves a significant average Dice improvement of 1.06% and 4.30% for multi-organ segmentation and pancreatic tumor segmentation, respectively, when compared to the highly competitive nn-UNet, and surpasses the top-1 solution in the BrasTS2021 challenge. 2D/3D Code and models are available at https://github.com/Beckschen/TransUNet and https://github.com/Beckschen/TransUNet-3D, respectively. Jieneng Chen, Jieru Mei, Xianhang Li, Yongyi Lu, Qihang Yu, Qingyue Wei, Xiangde Luo, Yutong Xie 0001, Ehsan Adeli-Mosabbeb, Yan Wang 0033, Matthew P. Lungren, Shaoting Zhang 0001, Lei Xing 0001, Le Lu 0001, Alan L. Yuille, Yuyin Zhou |
Medical Image Anal. | 13 |
| 2024 | Deciphering the Feature Representation of Deep Neural Networks for High-Performance AIabstractAI driven by deep learning is transforming many aspects of science and technology. The enormous success of deep learning stems from its unique capability of extracting essential features from Big Data for decision-making. However, the feature extraction and hidden representations in deep neural networks (DNNs) remain inexplicable, primarily because of lack of technical tools to comprehend and interrogate the feature space data. The main hurdle here is that the feature data are often noisy in nature, complex in structure, and huge in size and dimensionality, making it intractable for existing techniques to analyze the data reliably. In this work, we develop a computational framework named contrastive feature analysis (CFA) to facilitate the exploration of the DNN feature space and improve the performance of AI. By utilizing the interaction relations among the features and incorporating a novel data-driven kernel formation strategy into the feature analysis pipeline, CFA mitigates the limitations of traditional approaches and provides an urgently needed solution for the analysis of feature space data. The technique allows feature data exploration in unsupervised, semi-supervised and supervised formats to address different needs of downstream applications. The potential of CFA and its applications for pruning of neural network architectures are demonstrated using several state-of-the-art networks and well-annotated datasets across different disciplines. Md Tauhidul Islam, Lei Xing 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | NeRP: Implicit Neural Representation Learning With Prior Embedding for Sparsely Sampled Image ReconstructionabstractImage reconstruction is an inverse problem that solves for a computational image based on sampled sensor measurement. Sparsely sampled image reconstruction poses additional challenges due to limited measurements. In this work, we propose a methodology of implicit Neural Representation learning with Prior embedding (NeRP) to reconstruct a computational image from sparsely sampled measurements. The method differs fundamentally from previous deep learning-based image reconstruction approaches in that NeRP exploits the internal information in an image prior and the physics of the sparsely sampled measurements to produce a representation of the unknown subject. No large-scale data is required to train the NeRP except for a prior image and sparsely sampled measurements. In addition, we demonstrate that NeRP is a general methodology that generalizes to different imaging modalities such as computed tomography (CT) and magnetic resonance imaging (MRI). We also show that NeRP can robustly capture the subtle yet significant image changes required for assessing tumor progression. Liyue Shen, John M. Pauly, Lei Xing 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Multibranch CNN With MLP-Mixer-Based Feature Exploration for High-Performance Disease DiagnosisabstractDeep learning-based diagnosis is becoming an indispensable part of modern healthcare. For high-performance diagnosis, the optimal design of deep neural networks (DNNs) is a prerequisite. Despite its success in image analysis, existing supervised DNNs based on convolutional layers often suffer from their rudimentary feature exploration ability caused by the limited receptive field and biased feature extraction of conventional convolutional neural networks (CNNs), which compromises the network performance. Here, we propose a novel feature exploration network named manifold embedded multilayer perceptron (MLP) mixer (ME-Mixer), which utilizes both supervised and unsupervised features for disease diagnosis. In the proposed approach, a manifold embedding network is employed to extract class-discriminative features; then, two MLP-Mixer-based feature projectors are adopted to encode the extracted features with the global reception field. Our ME-Mixer network is quite general and can be added as a plugin to any existing CNN. Comprehensive evaluations on two medical datasets are performed. The results demonstrate that their approach greatly enhances the classification accuracy in comparison with different configurations of DNNs with acceptable computational complexity. Zixia Zhou, Md Tauhidul Islam, Lei Xing 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2023 | Consistency-Guided Meta-learning for Bootstrapping Semi-supervised Medical Image Segmentation
Qingyue Wei, Lequan Yu, Xianhang Li, Wei Shao 0008, Cihang Xie, Lei Xing 0001, Yuyin Zhou |
MICCAI (4) | 6 |
| 2023 | PINER: Prior-informed Implicit Neural Representation Learning for Test-time Adaptation in Sparse-view CT ReconstructionabstractRecently, deep learning has been introduced to solve important medical image reconstruction problems such as sparse-view CT reconstruction. However, the developed deep reconstruction models are generally limited in generalization when applied to out-of-distribution samples in unseen domains. Furthermore, privacy concerns may impede the availability of source-domain training data to retrain or adapt the model to the target-domain testing data, which are quite common in real-world medical applications. To address these issues, we introduce a source-free black-box test-time adaptation method for sparse-view CT reconstruction with unknown noise levels based on prior-informed implicit neural representation learning (PINER). By leveraging implicit neural representation learning to generate the image representations at various noise levels, the proposed method is able to construct the adapted input representations at test time based on the inference of black-box model and output analysis. We performed experiments of source-free test-time adaptation for sparse-view CT reconstruction with unknown noise levels on multiple anatomical sites with different black-box deep reconstruction models, where our method outperforms the state-of-the-art algorithms. Code: https://github.com/efzero/PINER Liyue Shen, Lei Xing 0001 |
WACV | 3 |
| 2023 | Adaptive Region-Specific Loss for Improved Medical Image SegmentationabstractDefining the loss function is an important part of neural network design and critically determines the success of deep learning modeling. A significant shortcoming of the conventional loss functions is that they weight all regions in the input image volume equally, despite the fact that the system is known to be heterogeneous (i.e., some regions can achieve high prediction performance more easily than others). Here, we introduce a region-specific loss to lift the implicit assumption of homogeneous weighting for better learning. We divide the entire volume into multiple sub-regions, each with an individualized loss constructed for optimal local performance. Effectively, this scheme imposes higher weightings on the sub-regions that are more difficult to segment, and vice versa. Furthermore, the regional false positive and false negative errors are computed for each input image during a training step and the regional penalty is adjusted accordingly to enhance the overall accuracy of the prediction. Using different public and in-house medical image datasets, we demonstrate that the proposed regionally adaptive loss paradigm outperforms conventional methods in the multi-organ segmentations, without any modification to the neural network architecture or additional data preparation. Lequan Yu, Jen-Yeu Wang, Neil Panjwani, Jean-Pierre Obeid, Wu Liu 0003, Lianli Liu, Nataliya Kovalchuk, Michael Gensheimer 0001, Lucas Kas Vitzthum, Beth M. Beadle, Daniel T. Chang, Quynh-Thu Le, Lei Xing 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 15 |
| 2023 | Small-Object Sensitive Segmentation Using Across Feature Map AttentionabstractSemantic segmentation is an important step in understanding the scene for many practical applications such as autonomous driving. Although Deep Convolutional Neural Networks-based methods have significantly improved segmentation accuracy, small/thin objects remain challenging to segment due to convolutional and pooling operations that result in information loss, especially for small objects. This article presents a novel attention-based method called Across Feature Map Attention (AFMA) to address this challenge. It quantifies the inner-relationship between small and large objects belonging to the same category by utilizing the different feature levels of the original image. The AFMA could compensate for the loss of high-level feature information of small objects and improve the small/thin object segmentation. Our method can be used as an efficient plug-in for a wide range of existing architectures and produces much more interpretable feature representation than former studies. Extensive experiments on eight widely used segmentation methods and other existing small-object segmentation models on CamVid and Cityscapes demonstrate that our method substantially and consistently improves the segmentation of small/thin objects. Shengtian Sang, Yuyin Zhou, Md Tauhidul Islam, Lei Xing 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Image classification using graph neural network and multiscale wavelet superpixelsabstractPrior studies using graph neural networks (GNNs) for image classification have focused on graphs generated from a regular grid of pixels or similar-sized superpixels. In the latter, a single target number of superpixels is defined for an entire dataset irrespective of differences across images and their intrinsic multiscale structure. On the contrary, this study investigates image classification using graphs generated from an image-specific number of multiscale superpixels. We propose WaveMesh, a new wavelet-based superpixeling algorithm, where the number and sizes of superpixels in an image are systematically computed based on its content. WaveMesh superpixel graphs are structurally different from similar-sized superpixel graphs. We use SplineCNN, a state-of-the-art network for image graph classification, to compare WaveMesh and similar-sized superpixels. Using SplineCNN, we perform extensive experiments on three benchmark datasets under three local-pooling settings: 1) no pooling, 2) GraclusPool, and 3) WavePool, a novel spatially heterogeneous pooling scheme tailored to WaveMesh superpixels. Our experiments demonstrate that SplineCNN learns from multiscale WaveMesh superpixels on-par with similar-sized superpixels. In all WaveMesh experiments, GraclusPool performs poorer than no pooling / WavePool, showing that poor cluster assignment negatively affects the performance of the network while learning from multiscale superpixels. Varun Vasudevan, Maxime Bassenne, Md Tauhidul Islam, Lei Xing 0001 |
Pattern Recognit. Lett. | 4 |
| 2023 | Label-Efficient Self-Supervised Federated Learning for Tackling Data Heterogeneity in Medical ImagingabstractThe collection and curation of large-scale medical datasets from multiple institutions is essential for training accurate deep learning models, but privacy concerns often hinder data sharing. Federated learning (FL) is a promising solution that enables privacy-preserving collaborative learning among different institutions, but it generally suffers from performance deterioration due to heterogeneous data distributions and a lack of quality labeled data. In this paper, we present a robust and label-efficient self-supervised FL framework for medical image analysis. Our method introduces a novel Transformer-based self-supervised pre-training paradigm that pre-trains models directly on decentralized target task datasets using masked image modeling, to facilitate more robust representation learning on heterogeneous data and effective knowledge transfer to downstream models. Extensive empirical results on simulated and real-world medical imaging non-IID federated datasets show that masked image modeling with Transformers significantly improves the robustness of models against various degrees of data heterogeneity. Notably, under severe data heterogeneity, our method, without relying on any additional pre-training data, achieves an improvement of 5.06%, 1.53% and 4.58% in test accuracy on retinal, dermatology and chest X-ray classification compared to the supervised baseline with ImageNet pre-training. In addition, we show that our federated self-supervised pre-training methods yield models that generalize better to out-of-distribution data and perform more effectively when fine-tuning with limited labeled data, compared to existing FL algorithms. The code is available at https://github.com/rui-yan/SSL-FL. Liangqiong Qu, Qingyue Wei, Shih-Cheng Huang, Liyue Shen, Daniel L. Rubin, Lei Xing 0001, Yuyin Zhou |
IEEE Trans. Medical Imaging | 7 |
| 2022 | Solving Inverse Problems in Medical Imaging with Score-Based Generative Models
Yang Song 0011, Liyue Shen, Lei Xing 0001, Stefano Ermon |
ICLR | 3 |
| 2022 | Novel-view X-ray projection synthesis through geometry-integrated deep learning
Liyue Shen, Lequan Yu, Wei Zhao 0029, John M. Pauly, Lei Xing 0001 |
Medical Image Anal. | 5 |
| 2021 | TransCT: Dual-Path Transformer for Low Dose Computed Tomography
Zhicheng Zhang 0005, Lequan Yu, Xiaokun Liang, Wei Zhao 0029, Lei Xing 0001 |
MICCAI (6) | 5 |
| 2021 | Estimating dual-energy CT imaging from single-energy CT data with material decomposition convolutional neural network
Tianling Lyu, Wei Zhao 0029, Yinsu Zhu, Zhan Wu, Yikun Zhang 0001, Yang Chen 0008, Limin Luo 0001, Shuo Li 0001, Lei Xing 0001 |
Medical Image Anal. | 9 |
| 2021 | Geometry and statistics-preserving manifold embedding for nonlinear dimensionality reduction
Md Tauhidul Islam, Lei Xing 0001 |
Pattern Recognit. Lett. | 2 |
| 2021 | Rotation-Oriented Collaborative Self-Supervised Learning for Retinal Disease DiagnosisabstractThe automatic diagnosis of various conventional ophthalmic diseases from fundus images is important in clinical practice. However, developing such automatic solutions is challenging due to the requirement of a large amount of training data and the expensive annotations for medical images. This paper presents a novel self-supervised learning framework for retinal disease diagnosis to reduce the annotation efforts by learning the visual features from the unlabeled images. To achieve this, we present a rotation-oriented collaborative method that explores rotation-related and rotation-invariant features, which capture discriminative structures from fundus images and also explore the invariant property used for retinal disease classification. We evaluate the proposed method on two public benchmark datasets for retinal disease classification. The experimental results demonstrate that our method outperforms other self-supervised feature learning methods (around 4.2% area under the curve (AUC)). With a large amount of unlabeled data available, our method can surpass the supervised baseline for pathologic myopia (PM) and is very close to the supervised baseline for age-related macular degeneration (AMD), showing the potential benefit of our method in clinical practice. Xiaomeng Li 0001, Xiaowei Hu 0001, Xiaojuan Qi 0001, Lequan Yu, Wei Zhao 0029, Pheng-Ann Heng, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2021 | Closing the Gap Between Deep Neural Network Modeling and Biomedical Decision-Making Metrics in Segmentation via Adaptive Loss FunctionsabstractDeep learning is becoming an indispensable tool for various tasks in science and engineering. A critical step in constructing a reliable deep learning model is the selection of a loss function, which measures the discrepancy between the network prediction and the ground truth. While a variety of loss functions have been proposed in the literature, a truly optimal loss function that maximally utilizes the capacity of neural networks for deep learning-based decision-making has yet to be established. Here, we devise a generalized loss function with functional parameters determined adaptively during model training to provide a versatile framework for optimal neural network-based decision-making in small target segmentation. The method is showcased by more accurate detection and segmentation of lung and liver cancer tumors as compared with the current state-of-the-art. The proposed formalism opens new opportunities for numerous practical applications such as disease diagnosis, treatment planning, and prognosis. Hyunseok Seo, Maxime Bassenne, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2021 | Deep Neural Network With Consistency Regularization of Multi-Output Channels for Improved Tumor Detection and DelineationabstractDeep learning is becoming an indispensable tool for imaging applications, such as image segmentation, classification, and detection. In this work, we reformulate a standard deep learning problem into a new neural network architecture with multi-output channels, which reflects different facets of the objective, and apply the deep neural network to improve the performance of image segmentation. By adding one or more interrelated auxiliary-output channels, we impose an effective consistency regularization for the main task of pixelated classification (i.e., image segmentation). Specifically, multi-output-channel consistency regularization is realized by residual learning via additive paths that connect main-output channel and auxiliary-output channels in the network. The method is evaluated on the detection and delineation of lung and liver tumors with public data. The results clearly show that multi-output-channel consistency implemented by residual learning improves the standard deep neural network. The proposed framework is quite broad and should find widespread applications in various deep learning problems. Hyunseok Seo, Lequan Yu, Hongyi Ren, Xiaomeng Li 0001, Liyue Shen, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2021 | Multi-Domain Image Completion for Random Missing Input DataabstractMulti-domain data are widely leveraged in vision applications taking advantage of complementary information from different modalities, e.g., brain tumor segmentation from multi-parametric magnetic resonance imaging (MRI). However, due to possible data corruption and different imaging protocols, the availability of images for each domain could vary amongst multiple data sources in practice, which makes it challenging to build a universal model with a varied set of input data. To tackle this problem, we propose a general approach to complete the random missing domain(s) data in real applications. Specifically, we develop a novel multi-domain image completion method that utilizes a generative adversarial network (GAN) with a representational disentanglement scheme to extract shared content encoding and separate style encoding across multiple domains. We further illustrate that the learned representation in multi-domain image completion could be leveraged for high-level tasks, e.g., segmentation, by introducing a unified framework consisting of image completion and segmentation with a shared content encoder. The experiments demonstrate consistent performance improvement on three datasets for brain tumor segmentation, prostate segmentation, and facial expression image completion respectively. Liyue Shen, Wentao Zhu 0001, Xiaosong Wang 0001, Lei Xing 0001, John M. Pauly, Baris Turkbey, Stephanie A. Harmon, Thomas Sanford, Sherif Mehralivand, Peter L. Choyke, Bradford J. Wood, Daguang Xu |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Deep Sinogram Completion With Image Prior for Metal Artifact Reduction in CT ImagesabstractComputed tomography (CT) has been widely used for medical diagnosis, assessment, and therapy planning and guidance. In reality, CT images may be affected adversely in the presence of metallic objects, which could lead to severe metal artifacts and influence clinical diagnosis or dose calculation in radiation therapy. In this article, we propose a generalizable framework for metal artifact reduction (MAR) by simultaneously leveraging the advantages of image domain and sinogram domain-based MAR techniques. We formulate our framework as a sinogram completion problem and train a neural network (SinoNet) to restore the metal-affected projections. To improve the continuity of the completed projections at the boundary of metal trace and thus alleviate new artifacts in the reconstructed CT images, we train another neural network (PriorNet) to generate a good prior image to guide sinogram learning, and further design a novel residual sinogram learning strategy to effectively utilize the prior image information for better sinogram completion. The two networks are jointly trained in an end-to-end fashion with a differentiable forward projection (FP) operation so that the prior image generation and deep sinogram completion procedures can benefit from each other. Finally, the artifact-reduced CT images are reconstructed using the filtered backward projection (FBP) from the completed sinogram. Extensive experiments on simulated and real artifacts data demonstrate that our method produces superior artifact-reduced results while preserving the anatomical structures and outperforms other MAR methods. Lequan Yu, Zhicheng Zhang 0005, Xiaomeng Li 0001, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2021 | Transformation-Consistent Self-Ensembling Model for Semisupervised Medical Image SegmentationabstractA common shortfall of supervised deep learning for medical imaging is the lack of labeled data, which is often expensive and time consuming to collect. This article presents a new semisupervised method for medical image segmentation, where the network is optimized by a weighted combination of a common supervised loss only for the labeled inputs and a regularization loss for both the labeled and unlabeled data. To utilize the unlabeled data, our method encourages consistent predictions of the network-in-training for the same input under different perturbations. With the semisupervised segmentation tasks, we introduce a transformation-consistent strategy in the self-ensembling model to enhance the regularization effect for pixel-level predictions. To further improve the regularization effects, we extend the transformation in a more generalized form including scaling and optimize the consistency loss with a teacher model, which is an averaging of the student model weights. We extensively validated the proposed semisupervised method on three typical yet challenging medical image segmentation tasks: 1) skin lesion segmentation from dermoscopy images in the International Skin Imaging Collaboration (ISIC) 2017 data set; 2) optic disk (OD) segmentation from fundus images in the Retinal Fundus Glaucoma Challenge (REFUGE) data set; and 3) liver segmentation from volumetric CT scans in the Liver Tumor Segmentation Challenge (LiTS) data set. Compared with state-of-the-art, our method shows superior performance on the challenging 2-D/3-D medical images, demonstrating the effectiveness of our semisupervised method for medical image segmentation. Xiaomeng Li 0001, Lequan Yu, Hao Chen 0011, Chi-Wing Fu, Lei Xing 0001, Pheng-Ann Heng |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | Difficulty-Aware Meta-learning for Rare Disease Diagnosis
Xiaomeng Li 0001, Lequan Yu, Yueming Jin, Chi-Wing Fu, Lei Xing 0001, Pheng-Ann Heng |
MICCAI (1) | 5 |
| 2020 | Automated hepatobiliary toxicity prediction after liver stereotactic body radiation therapy with deep learning-based portal vein segmentation
Bulat Ibragimov, Diego A. S. Toesca, Daniel T. Chang, Yixuan Yuan, Albert C. Koong, Lei Xing 0001 |
Neurocomputing | 6 |
| 2020 | Wireless Capsule Endoscopy: A New Tool for Cancer Screening in the Colon With Deep-Learning-Based Polyp RecognitionabstractAccurate recognition of polyps is crucial for early colorectal cancer diagnosis and treatment. Wireless capsule endoscopy (WCE) is a noninvasive, wireless imaging tool that allows direct visualization of the entire colon without discomfort to patients and has the potential to revolutionize the screening workup for colorectal diseases. However, current manual review is laborious and time consuming, requiring the undivided concentration of the gastroenterologist. Computational methods that can assist automated polyp recognition will enhance the outcome both in terms of diagnostic accuracy and efficiency of WCE. This review introduces the computer-assisted algorithms as applied to colorectal polyp screening, focusing on the successes of deep-learning-based strategies in the WCE sequences. We survey key applications of WCE polyp recognition, covering deep-learning-based image-level classification, lesion region detection, and pixel-accurate segmentation. We conclude by discussing emerging research challenges, possible trends, and future directions. Xiao Jia 0005, Xiaohan Xing, Yixuan Yuan, Lei Xing 0001, Max Q.-H. Meng |
Proc. IEEE | 4 |
| 2020 | Automatic Polyp Recognition in Colonoscopy Images Using Deep Learning and Two-Stage Pyramidal Feature PredictionabstractPolyp recognition in colonoscopy images is crucial for early colorectal cancer detection and treatment. However, the current manual review requires undivided concentration of the gastroenterologist and is prone to diagnostic errors. In this article, we present an effective, two-stage approach called PLPNet, where the abbreviation “PLP” stands for the word “polyp,” for automated pixel-accurate polyp recognition in colonoscopy images using very deep convolutional neural networks (CNNs). Compared to hand-engineered approaches and previous neural network architectures, our PLPNet model improves recognition accuracy by adding a polyp proposal stage that predicts the location box with polyp presence. Several schemes are proposed to ensure the model's performance. First of all, we construct a polyp proposal stage as an extension of the faster R-CNN, which performs as a region-level polyp detector to recognize the lesion area as a whole and constitutes stage I of PLPNet. Second, stageII of PLPNet is built in a fully convolutional fashion for pixelwise segmentation. We define a feature sharing strategy to transfer the learned semantics of polyp proposals to the segmentation task of stage II, which is proven to be highly capable of guiding the learning process and improve recognition accuracy. Additionally, we design skip schemes to enrich the feature scales and thus allow the model to generate detailed segmentation predictions. For accurate recognition, the advanced residual nets and feature pyramids are adopted to seek deeper and richer semantics at all network levels. Finally, we construct a two-stage framework for training and run our model convolutionally via a single-stream network at inference time to efficiently output the polyp mask. Experimental results on public data sets of GIANA Challenge demonstrate the accuracy gains of our approach, which surpasses previous state-of-the-art methods on the polyp segmentation task (74.7 Jaccard Index) and establishes new top results in the polyp localization challenge (81.7 recall). Xiao Jia 0005, Xiaochun Mai, Yi Cui 0002, Yixuan Yuan, Xiaohan Xing, Hyunseok Seo, Lei Xing 0001, Max Q.-H. Meng |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2020 | Densely Connected Neural Network With Unbalanced Discriminant and Category Sensitive Constraints for Polyp RecognitionabstractAutomatic polyp recognition in endoscopic images is challenging because of the low contrast between polyps and the surrounding area, the fuzzy and irregular polyp borders, and varying imaging light conditions. In this article, we propose a novel densely connected convolutional network with “unbalanced discriminant (UD)” loss and “category sensitive (CS)” loss (DenseNet-UDCS) for the task. We first utilize densely connected convolutional network (DenseNet) as the basic framework to conduct end-to-end polyp recognition task. Then, the proposed dual constraints, UD loss and CS loss, are simultaneously incorporated into the DenseNet model to calculate discriminative and suitable image features. The UD loss in our network effectively captures classification errors from both majority and minority categories to deal with the strong data imbalance of polyp images and normal ones. The CS loss imposes the ratio of intraclass and interclass variations in the deep feature learning process to enable features with large interclass variation and small intraclass compactness. With the joint supervision of UD loss and CS loss, a robust DenseNet-UDCS model is trained to recognize polyps from endoscopic images. The experimental results achieved polyp recognition accuracy of 93.19%, showing that the proposed DenseNet-UDCS can accurately characterize the endoscopic images and recognize polyps from the images. In addition, our DenseNet-UDCS model is superior in detection accuracy in comparison with state-of-the-art polyp recognition methods. Note to Practitioners-Wireless capsule endoscopy (WCE) is a crucial diagnostic tool for polyp detection and therapeutic monitoring, thanks to its noninvasive, user-friendly, and nonpainful properties. A challenge in harnessing the enormous potential of the WCE to benefit the gastrointestinal (GI) patients is that it requires clinicians to analyze a huge number of images (about 50 000 images for each patient). We propose a novel automatic polyp recognition scheme, namely, DenseNet-UDCS model, by addressing practical image unbalanced problem and small interclass variances and large intraclass differences in the data set. The comprehensive experimental results demonstrate superior reliability and robustness of the proposed model compared to the other polyp recognition approaches. Our DenseNet-UDCS model can be further applied in the clinical practice to provide valuable diagnosis information for GI disease recognition and precision medicine. Yixuan Yuan, Wenjian Qin, Bulat Ibragimov, Guanglei Zhang, Max Q.-H. Meng, Lei Xing 0001 |
IEEE Trans Autom. Sci. Eng. | 7 |
| 2020 | Superpixel Region Merging Based on Deep Network for Medical Image SegmentationabstractAutomatic and accurate semantic segmentation of pathological structures in medical images is challenging because of noisy disturbance, deformable shapes of pathology, and low contrast between soft tissues. Classical superpixel-based classification algorithms suffer from edge leakage due to complexity and heterogeneity inherent in medical images. Therefore, we propose a deep U-Net with superpixel region merging processing incorporated for edge enhancement to facilitate and optimize segmentation. Our approach combines three innovations: (1) different from deep learning--based image segmentation, the segmentation evolved from superpixel region merging via U-Net training getting rich semantic information, in addition to gray similarity; (2) a bilateral filtering module was adopted at the beginning of the network to eliminate external noise and enhance soft tissue contrast at edges of pathogy; and (3) a normalization layer was inserted after the convolutional layer at each feature scale, to prevent overfitting and increase the sensitivity to model parameters. This model was validated on lung CT, brain MR, and coronary CT datasets, respectively. Different superpixel methods and cross validation show the effectiveness of this architecture. The hyperparameter settings were empirically explored to achieve a good trade-off between the performance and efficiency, where a four-layer network achieves the best result in precision, recall, F-measure, and running speed. It was demonstrated that our method outperformed state-of-the-art networks, including FCN-16s, SegNet, PSPNet, DeepLabv3, and traditional U-Net, both quantitatively and qualitatively. Source code for the complete method is available at https://github.com/Leahnawho/Superpixel-network. Hui Liu 0016, Haiou Wang, Yan Wu 0012, Lei Xing 0001 |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2020 | Self-Supervised Feature Learning via Exploiting Multi-Modal Data for Retinal Disease DiagnosisabstractThe automatic diagnosis of various retinal diseases from fundus images is important to support clinical decision-making. However, developing such automatic solutions is challenging due to the requirement of a large amount of human-annotated data. Recently, unsupervised/self-supervised feature learning techniques receive a lot of attention, as they do not need massive annotations. Most of the current self-supervised methods are analyzed with single imaging modality and there is no method currently utilize multi-modal images for better results. Considering that the diagnostics of various vitreoretinal diseases can greatly benefit from another imaging modality, e.g., FFA, this paper presents a novel self-supervised feature learning method by effectively exploiting multi-modal data for retinal disease diagnosis. To achieve this, we first synthesize the corresponding FFA modality and then formulate a patient feature-based softmax embedding objective. Our objective learns both modality-invariant features and patient-similarity features. Through this mechanism, the neural network captures the semantically shared information across different modalities and the apparent visual similarity between patients. We evaluate our method on two public benchmark datasets for retinal disease diagnosis. The experimental results demonstrate that our method clearly outperforms other self-supervised feature learning methods and is comparable to the supervised baseline. Our code is available at GitHub. Xiaomeng Li 0001, Mengyu Jia, Md Tauhidul Islam, Lequan Yu, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Deep Learning-Based Spectral Unmixing for Optoacoustic Imaging of Tissue Oxygen SaturationabstractLabel free imaging of oxygenation distribution in tissues is highly desired in numerous biomedical applications, but is still elusive, in particular in sub-epidermal measurements. Eigenspectra multispectral optoacoustic tomography (eMSOT) and its Bayesian-based implementation have been introduced to offer accurate label-free blood oxygen saturation (sO2) maps in tissues. The method uses the eigenspectra model of light fluence in tissue to account for the spectral changes due to the wavelength dependent attenuation of light with tissue depth. eMSOT relies on the solution of an inverse problem bounded by a number of ad hoc hand-engineered constraints. Despite the quantitative advantage offered by eMSOT, both the non-convex nature of the optimization problem and the possible sub-optimality of the constraints may lead to reduced accuracy. We present herein a neural network architecture that is able to learn how to solve the inverse problem of eMSOT by directly regressing from a set of input spectra to the desired fluence values. The architecture is composed of a combination of recurrent and convolutional layers and uses both spectral and spatial features for inference. We train an ensemble of such networks using solely simulated data and demonstrate how this approach can improve the accuracy of sO2 computation over the original eMSOT, not only in simulations but also in experimental datasets obtained from blood phantoms and small animals (mice) in vivo. The use of a deep-learning approach in optoacoustic sO2 imaging is confirmed herein for the first time on ground truth sO2 values experimentally obtained in vivo and ex vivo. Ivan Olefir, Stratis Tzoumas, Courtney Restivo, Pouyan Mohajerani, Lei Xing 0001, Vasilis Ntziachristos |
IEEE Trans. Medical Imaging | 5 |
| 2020 | Modified U-Net (mU-Net) With Incorporation of Object-Dependent High Level Features for Improved Liver and Liver-Tumor Segmentation in CT ImagesabstractSegmentation of livers and liver tumors is one of the most important steps in radiation therapy of hepatocellular carcinoma. The segmentation task is often done manually, making it tedious, labor intensive, and subject to intra-/inter- operator variations. While various algorithms for delineating organ-at-risks (OARs) and tumor targets have been proposed, automatic segmentation of livers and liver tumors remains intractable due to their low tissue contrast with respect to the surrounding organs and their deformable shape in CT images. The U-Net has gained increasing popularity recently for image analysis tasks and has shown promising results. Conventional U-Net architectures, however, suffer from three major drawbacks. First, skip connections allow for the duplicated transfer of low resolution information in feature maps to improve efficiency in learning, but this often leads to blurring of extracted image features. Secondly, high level features extracted by the network often do not contain enough high resolution edge information of the input, leading to greater uncertainty where high resolution edge dominantly affects the network's decisions such as liver and liver-tumor segmentation. Thirdly, it is generally difficult to optimize the number of pooling operations in order to extract high level global features, since the number of pooling operations used depends on the object size. To cope with these problems, we added a residual path with deconvolution and activation operations to the skip connection of the U-Net to avoid duplication of low resolution information of features. In the case of small object inputs, features in the skip connection are not incorporated with features in the residual path. Furthermore, the proposed architecture has additional convolution layers in the skip connection in order to extract high level global features of small object inputs as well as high level features of high resolution edge information of large object inputs. Efficacy of the modified U-Net (mU-Net) was demonstrated using the public dataset of Liver tumor segmentation (LiTS) challenge 2017. For liver-tumor segmentation, Dice similarity coefficient (DSC) of 89.72 %, volume of error (VOE) of 21.93 %, and relative volume difference (RVD) of - 0.49 % were obtained. For liver segmentation, DSC of 98.51 %, VOE of 3.07 %, and RVD of 0.26 % were calculated. For the public 3D Image Reconstruction for Comparison of Algorithm Database (3Dircadb), DSCs were 96.01 % for the liver and 68.14 % for liver-tumor segmentations, respectively. The proposed mU-Net outperformed existing state-of-art networks. Hyunseok Seo, Charles W. Huang, Maxime Bassenne, Ruoxiu Xiao, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2019 | Real-Time Radiation Treatment Planning with Optimality Guarantees via Cluster and Bound MethodsabstractRadiation therapy is widely used in cancer treatment; however, plans necessarily involve tradeoffs between tumor coverage and mitigating damage to healthy tissue. Although current hardware can deliver custom-shaped beams from any angle around the patient, choosing (from all possible beams) an optimal set of beams that maximizes tumor coverage while minimizing collateral damage and treatment time is intractable. Furthermore, even though planning algorithms used in practice consider highly restricted sets of candidate beams, the time per run combined with the number of runs required to explore clinical tradeoffs results in planning times of hours to days. We propose a suite of cluster and bound methods that we hypothesize will (1) yield higher-quality plans by optimizing over much (i.e., 100-fold) larger sets of candidate beams, and/or (2) reduce planning time by allowing clinicians to search through candidate plans in real time. Our methods hinge on phrasing the treatment-planning problem as a convex problem. To handle large-scale optimizations, we form and solve compressed approximations to the full problem by clustering beams (i.e., columns of the dose deposition matrix used in the optimization) or voxels (rows of the matrix). Duality theory allows us to bound the error incurred when applying an approximate problem’s solution to the full problem. We observe that beam clustering and voxel clustering both yield excellent solutions while enabling a 10- to 200-fold speedup. Baris Ungun, Lei Xing 0001, Stephen P. Boyd |
INFORMS J. Comput. | 2 |
| 2019 | Self-attention convolutional neural network for improved MR image reconstruction
Yan Wu 0012, Yajun Ma, Jiang Du 0006, Lei Xing 0001 |
Inf. Sci. | 5 |
| 2019 | Neural Networks for Deep Radiotherapy Dose Analysis and Prediction of Liver SBRT OutcomesabstractStereotactic body radiation therapy (SBRT) is a relatively novel treatment modality, with little post-treatment prognostic information reported. This study proposes a novel neural network based paradigm for accurate prediction of liver SBRT outcomes. We assembled a database of patients treated with liver SBRT at our institution. Together with a three-dimensional (3-D) dose delivery plans for each SBRT treatment, other variables such as patients' demographics, quantified abdominal anatomy, history of liver comorbidities, other liver-directed therapies, and liver function tests were collected. We developed a multi-path neural network with the convolutional path for 3-D dose plan analysis and fully connected path for other variables analysis, where the network was trained to predict post-SBRT survival and local cancer progression. To enhance the network robustness, it was initially pre-trained on a large database of computed tomography images. Following n-fold cross-validation, the network automatically identified patients that are likely to have longer survival or late cancer recurrence, i.e., patients with the positive predicted outcome (PPO) of SBRT, and vice versa, i.e., negative predicted outcome (NPO). The predicted results agreed with actual SBRT outcomes with 56% of PPO patients and 0% NPO patients with primary liver cancer survived more than two years after SBRT. Similarly, 82% of PPO patients and 0% of NPO patients with metastatic liver cancer survived two-year threshold. The obtained results were superior to the performance of support vector machine and random forest classifiers. Furthermore, the network was able to identify the critical-to-spare liver regions, and the critical clinical features associated with the highest risks of negative SBRT outcomes. Bulat Ibragimov, Diego A. S. Toesca, Yixuan Yuan, Albert C. Koong, Daniel T. Chang, Lei Xing 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2019 | Deep Generative Adversarial Neural Networks for Compressive Sensing MRIabstractUndersampled magnetic resonance image (MRI) reconstruction is typically an ill-posed linear inverse task. The time and resource intensive computations require tradeoffs between accuracy and speed. In addition, state-of-the-art compressed sensing (CS) analytics are not cognizant of the image diagnostic quality. To address these challenges, we propose a novel CS framework that uses generative adversarial networks (GAN) to model the (low-dimensional) manifold of high-quality MR images. Leveraging a mixture of least-squares (LS) GANs and pixel-wise ℓ1/ℓ2cost, a deep residual network with skip connections is trained as the generator that learns to remove the aliasing artifacts by projecting onto the image manifold. The LSGAN learns the texture details, while the ℓ1/ℓ2cost suppresses high-frequency noise. A discriminator network, which is a multilayer convolutional neural network (CNN), plays the role of a perceptual cost that is then jointly trained based on high-quality MR images to score the quality of retrieved images. In the operational phase, an initial aliased estimate (e.g., simply obtained by zero-filling) is propagated into the trained generator to output the desired reconstruction. This demands a very low computational overhead. Extensive evaluations are performed on a large contrast-enhanced MR dataset of pediatric patients. Images rated by expert radiologists corroborate that GANCS retrieves higher quality images with improved fine texture details compared with conventional Wavelet-based and dictionary-learning-based CS schemes as well as with deep-learning-based schemes using pixel-wise training. In addition, it offers reconstruction times of under a few milliseconds, which are two orders of magnitude faster than the current state-of-the-art CS-MRI schemes. Morteza Mardani, Enhao Gong, Joseph Y. Cheng, S. S. Vasanawala, Greg Zaharchuk, Lei Xing 0001, John M. Pauly |
IEEE Trans. Medical Imaging | 6 |
| 2018 | Deep 3D Dose Analysis for Prediction of Outcomes After Liver Stereotactic Body Radiation Therapy
Bulat Ibragimov, Diego A. S. Toesca, Yixuan Yuan, Albert C. Koong, Daniel T. Chang, Lei Xing 0001 |
MICCAI (2) | 6 |
| 2018 | RIIS-DenseNet: Rotation-Invariant and Image Similarity Constrained Densely Connected Convolutional Network for Polyp Detection
Yixuan Yuan, Wenjian Qin, Bulat Ibragimov, Lei Xing 0001 |
MICCAI (2) | 5 |
| 2018 | Learning deconvolutional deep neural network for high resolution medical image reconstruction
Hui Liu 0016, Yan Wu 0012, Qiang Guo 0003, Bulat Ibragimov, Lei Xing 0001 |
Inf. Sci. | 6 |
| 2017 | Liver Lesion Detection Based on Two-Stage Saliency Model with Modified Sparse Autoencoder
Yixuan Yuan, Max Q.-H. Meng, Wenjian Qin, Lei Xing 0001 |
MICCAI (3) | 4 |
| 2017 | Segmentation of Pathological Structures by Landmark-Assisted Deformable ModelsabstractComputerized segmentation of pathological structures in medical images is challenging, as, in addition to unclear image boundaries, image artifacts, and traces of surgical activities, the shape of pathological structures may be very different from the shape of normal structures. Even if a sufficient number of pathological training samples are collected, statistical shape modeling cannot always capture shape features of pathological samples as they may be suppressed by shape features of a considerably larger number of healthy samples. At the same time, landmarking can be efficient in analyzing pathological structures but often lacks robustness. In this paper, we combine the advantages of landmark detection and deformable models into a novel supervised multi-energy segmentation framework that can efficiently segment structures with pathological shape. The framework adopts the theory of Laplacian shape editing, that was introduced in the field of computer graphics, so that the limitations of statistical shape modeling are avoided. The performance of the proposed framework was validated by segmenting fractured lumbar vertebrae from 3-D computed tomography images, atrophic corpora callosa from 2-D magnetic resonance (MR) cross-sections and cancerous prostates from 3D MR images, resulting respectively in a Dice coefficient of 84.7 ± 5.0%, 85.3 ± 4.8% and 78.3 ± 5.1%, and boundary distance of 1.14 ± 0.49mm, 1.42 ± 0.45mm and 2.27 ± 0.52mm. The obtained results were shown to be superior in comparison to existing deformable model-based segmentation algorithms. Bulat Ibragimov, Robert Korez, Bostjan Likar, Franjo Pernus, Lei Xing 0001, Tomaz Vrtovec |
IEEE Trans. Medical Imaging | 5 |
| 2017 | Cone Beam X-ray Luminescence Computed Tomography Based on Bayesian MethodabstractX-ray luminescence computed tomography (XLCT), which aims to achieve molecular and functional imaging by X-rays, has recently been proposed as a new imaging modality. Combining the principles of X-ray excitation of luminescence-based probes and optical signal detection, XLCT naturally fuses functional and anatomical images and provides complementary information for a wide range of applications in biomedical research. In order to improve the data acquisition efficiency of previously developed narrow-beam XLCT, a cone beam XLCT (CB-XLCT) mode is adopted here to take advantage of the useful geometric features of cone beam excitation. Practically, a major hurdle in using cone beam X-ray for XLCT is that the inverse problem here is seriously ill-conditioned, hindering us to achieve good image quality. In this paper, we propose a novel Bayesian method to tackle the bottleneck in CB-XLCT reconstruction. The method utilizes a local regularization strategy based on Gaussian Markov random field to mitigate the ill-conditioness of CB-XLCT. An alternating optimization scheme is then used to automatically calculate all the unknown hyperparameters while an iterative coordinate descent algorithm is adopted to reconstruct the image with a voxel-based closed-form solution. Results of numerical simulations and mouse experiments show that the self-adaptive Bayesian method significantly improves the CB-XLCT image quality as compared with conventional methods. Guanglei Zhang, Fei Liu 0005, Jianwen Luo 0001, Yaoqin Xie, Jing Bai 0001, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2015 | Optimized Detector Angular Configuration Increases the Sensitivity of X-ray Fluorescence Computed Tomography (XFCT)abstractIn this work, we demonstrated that an optimized detector angular configuration based on the anisotropic energy distribution of background scattered X-rays improves X-ray fluorescence computed tomography (XFCT) detection sensitivity. We built an XFCT imaging system composed of a bench-top fluoroscopy X-ray source, a CdTe X-ray detector, and a phantom motion stage. We imaged a 6.4-cm-diameter phantom containing different concentrations of gold solution and investigated the effect of detector angular configuration on XFCT image quality. Based on our previous theoretical study, three detector angles were considered. The X-ray fluorescence detector was first placed at 145 (°) (approximating back-scatter) to minimize scatter X-rays. XFCT image quality was compared to images acquired with the detector at 60 (°) (forward-scatter) and 90 (°) (side-scatter). The datasets for the three different detector positions were also combined to approximate an isotropically arranged detector. The sensitivity was optimized with detector in the 145 (°) back-scatter configuration counting the 78-keV gold Kβ1 X-rays. The improvement arose from the reduced energy of scattered X-ray at the 145 (°) position and the large energy separation from gold K β1 X-rays. The lowest detected concentration in this configuration was 2.5 mgAu/mL (or 0.25% Au with SNR = 4.3). This concentration could not be detected with the 60 (°) , 90 (°) , or isotropic configurations (SNRs = 1.3, 0, 2.3, respectively). XFCT imaging dose of 14 mGy was in the range of typical clinical X-ray CT imaging doses. To our knowledge, the sensitivity achieved in this experiment is the highest in any XFCT experiment using an ordinary bench-top X-ray source in a phantom larger than a mouse ( > 3 cm). Moiz Ahmad, Magdalena Bazalova-Carter, Rebecca Fahrig, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2014 | Order of Magnitude Sensitivity Increase in X-ray Fluorescence Computed Tomography (XFCT) Imaging With an Optimized Spectro-Spatial Detector Configuration: Theory and SimulationabstractThe purpose of this study was to increase the sensitivity of XFCT imaging by optimizing the data acquisition geometry for reduced scatter X-rays. The placement of detectors and detector energy window were chosen to minimize scatter X-rays. We performed both theoretical calculations and Monte Carlo simulations of this optimized detector configuration on a mouse-sized phantom containing various gold concentrations. The sensitivity limits were determined for three different X-ray spectra: a monoenergetic source, a Gaussian source, and a conventional X-ray tube source. Scatter X-rays were minimized using a backscatter detector orientation (scatter direction > 110(°) to the primary X-ray beam). The optimized configuration simultaneously reduced the number of detectors and improved the image signal-to-noise ratio. The sensitivity of the optimized configuration was 10 μg/mL (10 pM) at 2 mGy dose with the mono-energetic source, which is an order of magnitude improvement over the unoptimized configuration (102 pM without the optimization). Similar improvements were seen with the Gaussian spectrum source and conventional X-ray tube source. The optimization improvements were predicted in the theoretical model and also demonstrated in simulations. The sensitivity of XFCT imaging can be enhanced by an order of magnitude with the data acquisition optimization, greatly enhancing the potential of this modality for future use in clinical molecular imaging. Moiz Ahmad, Magdalena Bazalova-Carter, Liangzhong Xiang, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2013 | First Demonstration of Multiplexed X-Ray Fluorescence Computed Tomography (XFCT) ImagingabstractSimultaneous imaging of multiple probes or biomarkers represents a critical step toward high specificity molecular imaging. In this work, we propose to utilize the element-specific nature of the X-ray fluorescence (XRF) signal for imaging multiple elements simultaneously (multiplexing) using XRF computed tomography (XFCT). A 5-mm-diameter pencil beam produced by a polychromatic X-ray source (150 kV, 20 mA) was used to stimulate emission of XRF photons from 2% (weight/volume) gold (Au), gadolinium (Gd), and barium (Ba) embedded within a water phantom. The phantom was translated and rotated relative to the stationary pencil beam in a first-generation CT geometry. The X-ray energy spectrum was collected for 18 s at each position using a cadmium telluride detector. The spectra were then used to isolate the K shell XRF peak and to generate sinograms for the three elements of interest. The distribution and concentration of the three elements were reconstructed with the iterative maximum likelihood expectation maximization algorithm. The linearity between the XFCT intensity and the concentrations of elements of interest was investigated. We found that measured XRF spectra showed sharp peaks characteristic of Au, Gd, and Ba. The narrow full-width at half-maximum (FWHM) of the peaks strongly supports the potential of XFCT for multiplexed imaging of Au, Gd, and Ba ( FWHM(Au,Kα1) = 0.619 keV, FWHM(Au,Kα2)=1.371 keV , FWHM(Gd,Kα)=1.297 keV, FWHM(Gd,Kβ)=0.974 keV , FWHM(Ba,Kα)=0.852 keV, and FWHM(Ba,Kβ)=0.594 keV ). The distribution of Au, Gd, and Ba in the water phantom was clearly identifiable in the reconstructed XRF images. Our results showed linear relationships between the XRF intensity of each tested element and their concentrations ( R(2)(Au)=0.944 , R(Gd)(2)=0.986, and R(Ba)(2)=0.999), suggesting that XFCT is capable of quantitative imaging. Finally, a transmission CT image was obtained to show the potential of the approach for providing attenuation correction and morphological information. In conclusion, XFCT is a promising modality for multiplexed imaging of high atomic number probes. Yu Kuang, Guillem Pratx, Magdalena Bazalova-Carter, Bowen Meng, Jianguo Qian, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2012 | Investigation of X-ray Fluorescence Computed Tomography (XFCT) and K-Edge ImagingabstractThis work provides a comprehensive Monte Carlo study of X-ray fluorescence computed tomography (XFCT) and K-edge imaging system, including the system design, the influence of various imaging components, the sensitivity and resolution under various conditions. We modified the widely used EGSnrc/DOSXYZnrc code to simulate XFCT images of two acrylic phantoms loaded with various concentrations of gold nanoparticles and Cisplatin for a number of XFCT geometries. In particular, reconstructed signal as a function of the width of the detector ring, its angular coverage and energy resolution were studied. We found that XFCT imaging sensitivity of the modeled systems consisting of a conventional X-ray tube and a full 2-cm-wide energy-resolving detector ring was 0.061% and 0.042% for gold nanoparticles and Cisplatin, respectively, for a dose of ∼ 10 cGy. Contrast-to-noise ratio (CNR) of XFCT images of the simulated acrylic phantoms was higher than that of transmission K-edge images for contrast concentrations below 0.4%. Magdalena Bazalova-Carter, Yu Kuang, Guillem Pratx, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2010 | X-Ray Luminescence Computed Tomography via Selective Excitation: A Feasibility StudyabstractX-ray luminescence computed tomography (XLCT) is proposed as a new molecular imaging modality based on the selective excitation and optical detection of X-ray-excitable phosphor nanoparticles. These nano-sized particles can be fabricated to emit near-infrared (NIR) light when excited with X-rays, and, because because both X-rays and NIR photons propagate long distances in tissue, they are particularly well suited for in vivo biomedical imaging. In XLCT, tomographic images are generated by irradiating the subject using a sequence of programmed X-ray beams, while sensitive photo-detectors measure the light diffusing out of the subject. By restricting the X-ray excitation to a single, narrow beam of radiation, the origin of the optical photons can be inferred regardless of where these photons were detected, and how many times they scattered in tissue. This study presents computer simulations exploring the feasibility of imaging small objects with XLCT, such as research animals. The accumulation of 50 nm phosphor nanoparticles in a 2-mm-diameter target can be detected and quantified with subpicomolar sensitivity using less than 1 cGy of radiation dose. Provided sufficient signal-to-noise ratio, the spatial resolution of the system can be made as high as needed by narrowing the beam aperture. In particular, 1 mm spatial resolution was achieved for a 1-mm-wide X-ray beam. By including an X-ray detector in the system, anatomical imaging is performed simultaneously with molecular imaging via standard X-ray computed tomography (CT). The molecular and anatomical images are spatially and temporally co-registered, and, if a single-pixel X-ray detector is used, they have matching spatial resolution. Guillem Pratx, Colin Carpenter, Conroy Sun, Lei Xing 0001 |
IEEE Trans. Medical Imaging | 4 |