EDBT 2026 Demo / reviewers in the wild / expert
Wei Shao 0005
dblp:24/803-5
· DBLP profile ↗
71ranked-venue papers
14as first author
58since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 48 · 13 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 18 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 15 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EfficientCovNet: Modeling the Pairwise Voxel Dependency for Brain ROI SegmentationabstractSegmenting the brain magnetic resonance (MR) images to region-of-interest (ROI) is a fundamental step for many medical image analysis tasks. Convolutional neural networks (CNNs) excel in learning the high-level contextual features for image segmentation. However, such high-level features are low-order features, which cannot reflect the complex appearance patterns of brain MR images. Intuitively, using the high-order features can enhance the performance of CNNs. Therefore, in this paper, we propose a novel Efficient Covariance Network (EfficientCovNet) that models pairwise voxel dependency features and applies it to the brain ROI segmentation tasks. Our EfficientCovNet consists of two pathways: a pairwise voxel dependency feature learning pathway that uses a novel covariance convolution to efficiently capture the pairwise features from MR images, and a contextual feature learning pathway that extracts high-level contextual features using convolutional operations. The pairwise features and contextual features are then fused together to boost brain ROI segmentation performance. Experimental results on five datasets, i.e., IXI, LONI-LPBA40, OASIS, ADNI, and CC359 datasets, demonstrate that our EfficientCovNet achieves superior performance for brain ROI segmentation in comparison with the state-of-the-art methods. Liang Sun 0009, Junyong Zhao, Wei Shao 0005, Qi Zhu 0001, Daoqiang Zhang |
IEEE Trans. Image Process. | 3 |
| 2026 | Foundation Model-Based Zero-Shot Tissue Segmentation of Pathological Images via the Mixture of Local-to-Global ExpertsabstractTissue segmentation in pathological images plays a crucial role for the diagnosis and prognosis of human cancers. However, due to the complexity of tumor micro-environment, it is difficult to annotate all tissue types especially for the categories with small tissue proportions, which limits the ability of the traditional tissue segmentation models to these tissue types with zero training samples. To address the above issues, we present a novel architecture, ZSPMLG, that relies on pathology vision-language foundation model (i.e., CONCH) to learn pixel-wise classifiers for both seen and unseen tissue types based on their text descriptions. Specifically, we firstly apply large language model (LLM) to generate the descriptions for both seen and unseen tissue categories, followed by feeding them to the CONCH text encoder to acquire their corresponding prototypes that are shared by both vision and semantic space. By considering that the textual descriptions of specific tissue categories can be observed from the pathological images at different scales of magnification, our ZSPMLG consists of Mixture of Local Experts (MoLE) and Mixture of Global Experts (MoGE) modules, where MoLE performs the specialized decoding that can map individual scale patch-level representation to dense pixel-level representation, while MoGE aims at fusing the multi-scale representations together. Finally, a convolutional layer is designed to map the pixel-level representation to the category prototype for tissue segmentation on both seen and unseen categories. We evaluate our method on three datasets and the experimental results demonstrate the superiority of our method on both seen and unseen tissue categories. Yunfeng Ye, Jingtian Yuan, Jiao Tang, Peng Wan 0004, Liang Sun 0009, Jianpeng Sheng, Daoqiang Zhang, Wei Shao 0005 |
IEEE Trans. Image Process. | 8 |
| 2026 | StableMIL: Entropy-Stabilized Attention-Based Multiple Instance Learning for Morphologically Variable Whole Slide ImagesabstractAggregating features of tens of thousands of patches into Whole Slide Images (WSIs) representations via aggregators is a crucial step in computational pathology. However, existing aggregation strategies overlook the morphological variability of tissue regions in WSIs stemming from differences in clinical procedures and tumor characteristics, leading to two critical limitations: 1) attention collapse in long sequences caused by significant variation in patch numbers across WSIs (ranging from thousands to tens of thousands per WSI); 2) attention misallocation due to under-trained positional embeddings resulting from the non-uniform spatial coordinates introduced by irregular patch distributions. Consequently, current attention-based methods struggle to generalize across this morphological variability, resulting in inconsistent aggregation performance and compromised model reliability in clinical settings. To address these issues, we propose a Entropy-Stabilized Attention-based Multiple Instance Learning (StableMIL) framework, which incorporates an entropy-stabilized attention mechanism to ensure consistent aggregation across WSIs with varying patch numbers and a Randomly Projected 2D rotary position embedding to enhance spatial representation robustness across irregular patch distributions. Extensive theoretical and experimental analyses on nine WSI datasets spanning diverse cancer types, across both classification and survival prediction tasks, demonstrate that StableMIL effectively overcomes the challenges of handling long instance sequences and out-of-distribution spatial coordinates. Our framework consistently outperforms representative baselines, particularly in survival prediction, with stable improvements observed across all evaluated cancer types and morphological scenarios, highlighting its potential for real-world clinical applications. Our source code is available at https://github.com/theeeqi/stableMIL. Yinuo Lu, Mingxin Qi, Yao Fu 0010, Zhuoran Xiao, Wei Shao 0005, Jie Tian 0001 |
IEEE Trans. Medical Imaging | 5 |
| 2026 | Open-Set Active Learning for Nucleus Detection From the Histopathological ImagesabstractThe recent advance of deep learning has shown great potential for nucleus detection which plays an important role in the histopathological examination. However, such accurate and reliable deep learning models usually need enough labeled data for training, which makes active learning an appealing learning paradigm to reduce the annotation efforts by experts. In open-set environments, active learning encounters the challenge that the unlabeled data usually contain non-target samples from the unknown classes, resulting in the failure of most active learning methods. Although active learning has been explored in many open-set classification tasks, research on active learning for nucleus detection in the open-set environment remains unexplored. To address the above issues, we propose a two-stage active learning framework designed for nucleus detection in the open-set environment (i.e., OpAL4ND). In the first stage, we propose a prototype-based query strategy based on the auxiliary detector to select a candidate set from known classes as pure as possible. In the second stage, we further query the most uncertain and representative samples from the candidate set for the nucleus detection task relying on the target detector. We evaluate the performance of our method on two nucleus detection datasets (i.e., the NuCLS and PanNuke datasets), and the experimental results indicate that our method can not only improve the selection quality on the known classes, but also achieve higher detection accuracy with lower annotation burden in comparison with the existing studies. Code is available at https://github.com/onbut/OpAL4ND. Jiao Tang, Yagao Yue, Peng Wan 0004, Andrey S. Krylov, Wei Shao 0005, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 8 |
| 2026 | Trustworthy Multi-Modal Ultrasound Fusion via Uncertainty Calibration and Conflict ResolutionabstractMulti-modal ultrasound combines tissue information from multiple imaging perspectives, enabling more comprehensive lesion assessment. However, conventional multi-view learning methods typically assume uniform modality quality, ignoring variability caused by imaging noise and patient-specific factors. This oversight limits diagnostic reliability, especially when some modalities provide uncertain or conflicting information. To address this, we identify two key challenges in multi-modal ultrasound fusion: 1) how to quantify modality-wise uncertainty, and 2) how to resolve conflicts among predictions. We propose a novel method, termed TMUF (Trustworthy Multi-modal Ultrasound Fusion), which dynamically integrates information from different modalities through uncertainty calibration and conflict resolution. Specifically, we introduce a cross-modal uncertainty calibration regularizer to estimate evidence-based uncertainty across modalities, aligning uncertainty with prediction correctness. We further develop a credibility-aware fusion strategy that evaluates cross-modal consistency and uncertainty to distinguish credible from non-credible modalities, assigning fusion weights accordingly. We validate TMUF on public and private datasets for breast lesion and liver cancer diagnosis. The proposed method achieves diagnostic accuracies of 88.00% and 92.08%, respectively, outperforming state-of-the-art baselines. These results demonstrate the effectiveness of TMUF in enhancing diagnostic accuracy and robustness for multi-modal ultrasound. Peng Wan 0004, Limei Wei, Shukang Zhang, Haiyan Xue, Wei Shao 0005, Wentao Kong, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 5 |
| 2026 | Identification of Genetic Risk Factors Based on Disease Progression Derived From Modeling Longitudinal Phenotype Latent Pattern RepresentationabstractThe characteristic of neurodegenerative disorders is the progressive impairment of memory and other cognitive functions. However, these existing imaging genetic methods only use longitudinal imaging phenotypes straightforwardly, ignoring the latent pattern of the longitudinal data in the progression process. The phenotypes across multiple time-points may exhibit the latent pattern that can be used to facilitate the understanding of the progression process. Accordingly, in this paper, we explore underlying complementary information from multiple time-points and simultaneously seek the underlying latent representation. With the complementarity of multiple time-points, the latent representation depicts data more comprehensively than each individual time-point, therefore mining effective longitudinal phenotype latent pattern representation. Specifically, we first propose two latent pattern representation (LPR) for longitudinal imaging phenotypes: linear LPR (lLPR), based on linear relationships between latent representation and each time-point, and nonlinear LPR (nonlLPR), based on neural networks to deal with nonlinear relationships. Then, we calculate the imaging genetic association based on the latent pattern representation. Finally, we conduct the experiments on both synthetic and real longitudinal imaging genetic data. Related experimental results validate that our proposed approach outperforms several competing algorithms, establishes strong associations, and discovers consistent longitudinal imaging genetic biomarkers, thereby guiding disease interpretation. Meiling Wang 0001, Wei Shao 0005, Daoqiang Zhang, Qingshan Liu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2026 | CUSTrack: Causality-Inspired Liver Ultrasound Tracking With Periodic Motion Bias MitigationabstractReal-time tissue tracking is a fundamental task in liver ultrasound applications. Due to the periodic nature of liver motion, historical trajectories can offer valuable priors for target localization, particularly when foreground-background distinction is weak. However, existing trackers often exploit these trajectories as shortcuts, relying excessively on periodic respiratory patterns rather than true object appearance matching. In this work, we revisit liver tracking from a causal perspective and propose CUSTrack, a method that mitigates periodicity bias by decomposing and correcting the total causal effect of historical trajectories. We define periodicity bias as the direct causal effect of past states and eliminate it via counterfactual reasoning, preserving 'good' trajectory priors while suppressing 'bad' periodic bias. To ensure identifiability, we incorporate a deconfounding module that removes latent confounders from fused feature representations. Extensive experiments on liver ultrasound datasets demonstrate that CUSTrack achieves superior tracking accuracy and robustness under challenging conditions. Shukang Zhang, Junyong Zhao, Huanjun Wang, Wei Shao 0005, Wentao Kong, Peng Wan 0004, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 4 |
| 2026 | ProtoMTG: Prototypical Multi-Task Learning for the Generation of Multiple Stained Immunohistochemical ImagesabstractMultiplex immunohistochemistry (mIHC) images have the potential to assess the complex tumor microenvironment by simultaneously detecting multiple markers within a single tissue section, however, the acquisition of mIHC images in clinical labs is both time-consuming and costly. Hence, applying machine learning-based virtual staining techniques for rapid generation of different mIHC markers has become a considerable alternative. The existing bio-image based virtual staining models generate the distributions of different markers independently, which have limited interpretability and overlook the fact that the exploration of potential interrelationships among these markers can help determine the localization of each individual marker. To address the above issues, we propose an explainable prototypical multi-task generation framework (i.e., ProtoMTG) to simultaneously generate multiple mIHC markers. Specifically, ProtoMTG involves a multi-task prototype layer that can capture the relationship among different virtual staining tasks by learning the shared and task-specific prototypes. Then, in the proto-attention layer, both task-specific and shared prototypes will be re-weighted and combined to instruct the generation of different mIHC markers. In ProtoMTG, we also design the novel prototypical activation and diversity losses to learn better prototype representation for the virtual staining task. To evaluate the performance of our method, we develop three benchmark mIHC datasets on different organs (i.e., colon, liver and stomach). The experimental results indicate that our method can not only outperform the existing image generation models, but also have good explainable ability for the virtual staining of mIHC markers. The code and dataset are available at: https://jj-zhou-code.github.io/ProtoMTG-website/. Andrey S. Krylov, Jianpeng Sheng, Qi Zhu 0001, Wei Shao 0005, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 5 |
| 2025 | Multi-modal Topology-embedded Graph Learning for Spatially Resolved Genes Prediction from Pathology Images with Prior Gene Similarity InformationabstractThe rapid development of spatial transcriptomics (ST) allows researchers to measure the spatial-level gene expression in tissues. Although powerful, the cost for collecting the ST data is expensive, and thus several studies aim to predict gene expression in ST by utilizing their corresponding H/E stained pathology images. The existing ST based gene expression prediction models either adopt the pre-trained networks or rely on the handcrafted features to describe the pathology images, which still lack a systematic way to combine them together to define a spot-level representation that can reflect the topological profiles of different spots. On the other hand, all the ST based gene prediction models treat the prediction task for each gene independently, which overlook the fact that the exploration of potential interrelationships among them can help improve the prediction performance for individual genes. To address the above issues, we propose a multi-modal topology-embedded graph learning algorithm guided by prior Gene Ontology similarity information (i.e., M2TGLGO) to predict the spatial resolved genes from pathology images. Specifically, M2TGLGO co-learns the image representation of different spots from both deep and handcrafted features by considering the within-modal and inter-modal interactions. Next, to keep the topological structure among different spots, a spatial-oriented ranking module is also incorporated to preserve their neighborhood similarity information. Finally, we present a Gene Ontology knowledge guided graph neural network for simultaneously predicting multiple gene expressions by considering their functional associations. We evaluate our method on three public available ST datasets, the experimental results show the effectiveness of our M2TGLGO in comparison with the existing studies. Changxi Chi, Peng Wan 0004, Daoqiang Zhang, Wei Shao 0005 |
CVPR | 5 |
| 2025 | Robust Multimodal Survival Prediction with Conditional Latent Differentiation Variational AutoEncoderabstractThe integrative analysis of histopathological images and genomic data has received increasing attention for survival prediction of human cancers. However, the existing studies always hold the assumption that full modalities are available. As a matter of fact, the cost for collecting genomic data is high, which sometimes makes genomic data unavailable in testing samples. A common way of tackling such incompleteness is to generate the genomic representations from the pathology images. Nevertheless, such strategy still faces the following two challenges: (1) The gigapixel whole slide images (WSIs) are huge and thus hard for representation. (2) It is difficult to generate the genomic embeddings with diverse function categories in a unified generative framework. To address the above challenges, we propose a Conditional Latent Differentiation Variational AutoEncoder (LD-CVAE) for robust multimodal survival prediction, even with missing genomic data. Specifically, a Variational Information Bottleneck Transformer (VIBTrans) module is proposed to learn compressed pathological representations from the gigapixel WSIs. To generate different functional genomic features, we develop a novel Latent Differentiation Variational AutoEncoder (LD-VAE) to learn the genomic and function-specific posteriors for the genomic embeddings with diverse functions. Finally, we use the product-of-experts technique to integrate the genomic posterior and image posterior for the joint latent distribution estimation in LD-CVAE. We test the effectiveness of our method on five different cancer datasets, and the experimental results demonstrate its superiority in both complete and missing modality scenarios. The code is released†. Jiao Tang, Yingli Zuo, Peng Wan 0004, Daoqiang Zhang, Wei Shao 0005 |
CVPR | 6 |
| 2025 | DAMM-Diffusion: Learning Divergence-Aware Multi-Modal Diffusion Model for Nanoparticles Distribution PredictionabstractThe prediction of nanoparticles (NPs) distribution is crucial for the diagnosis and treatment of tumors. Recent studies indicate that the heterogeneity of tumor microenvironment (TME) highly affects the distribution of NPs across tumors. Hence, it has become a research hotspot to generate the NPs distribution by the aid of multi-modal TME components. However, the distribution divergence among multi-modal TME components may cause side effects i.e., the best unimodal model may outperform the joint generative model. To address the above issues, we propose a Divergence-Aware Multi-Modal Diffusion model (i.e., DAMM-Diffusion) to adaptively generate the prediction results from uni-modal and multi-modal branches in a unified network. In detail, the uni-modal branch is composed of the U-Net architecture while the multi-modal branch extends it by introducing two novel fusion modules i.e., Multi-Modal Fusion Module (MMFM) and Uncertainty-Aware Fusion Module (UAFM). Specifically, the MMFM is proposed to fuse features from multiple modalities, while the UAFM module is introduced to learn the uncertainty map for cross-attention computation. Following the individual prediction results from each branch, the Divergence-Aware Multi-Modal Predictor (DAMMP) module is proposed to assess the consistency of multi-modal data with the uncertainty map, which determines whether the final prediction results come from multi-modal or uni-modal predictions. We predict the NPs distribution given the TME components of tumor vessels and cell nuclei, and the experimental results show that DAMM-Diffusion can generate the distribution of NPs with higher accuracy than the comparing methods. Additional results on the multi-modal brain image synthesis task further validate the effectiveness of the proposed method. The code is released†. Shouju Wang, Yuxia Tang, Qi Zhu 0001, Daoqiang Zhang, Wei Shao 0005 |
CVPR | 6 |
| 2025 | AcZeroTS: Active Learning for Zero-Shot Tissue Segmentation in Pathology Images
Jiao Tang, Peng Wan 0004, Yingli Zuo, Wei Shao 0005, Daoqiang Zhang |
ICCV | 6 |
| 2025 | LTSE: Language-Guided Tissue Referring Segmentation in Pathology Images with Adaptive Expert Mixture
Jiao Tang, Peng Wan 0004, Wei Shao 0005, Daoqiang Zhang |
MICCAI (6) | 4 |
| 2025 | Cost-Effective Active Learning for Nucleus Detection Using Crowdsourced Annotations with Dynamic Weighting Adjustment
Jiao Tang, Yuankun Zu, Qi Zhu 0001, Peng Wan 0004, Daoqiang Zhang, Wei Shao 0005 |
MICCAI (13) | 6 |
| 2025 | NeuroH-TGL: Neuro-Heterogeneity Guided Temporal Graph Learning Strategy for Brain Disease DiagnosisabstractDynamic functional brain networks (DFBNs) are powerful tools in neuroscience research. Recent studies reveal that DFBNs contain heterogeneous neural nodes with more extensive connections and more drastic temporal changes, which play pivotal roles in coordinating the reorganization of the brain. Moreover, the spatio-temporal patterns of these nodes are modulated by the brain's historical states. However, existing methods not only ignore the spatio-temporal heterogeneity of neural nodes, but also fail to effectively encode the temporal propagation mechanism of heterogeneous activities. These limitations hinder the deep exploration of spatio-temporal relationships within DFBNs, preventing the capture of abnormal neural heterogeneity caused by brain diseases. To address these challenges, this paper propose a neuro-heterogeneity guided temporal graph learning strategy (NeuroH-TGL). Specifically, we first develop a spatio-temporal pattern decoupling module to disentangle DFBNs into topological consistency networks and temporal trend networks that align with the brain's operational mechanisms. Then, we introduce a heterogeneity mining module to identify pivotal heterogeneity nodes that drive brain reorganization from the two decoupled networks. Finally, we design temporal propagation graph convolution to simulate the influence of the historical states of heterogeneity nodes on the current topology, thereby flexibly extracting heterogeneous spatio-temporal information from the brain. Experiments show that our method surpasses several state-of-the-art methods, and can identify abnormal heterogeneous nodes caused by brain diseases. Shengrong Li, Qi Zhu 0001, Chunwei Tian, Wei Shao 0005, Jie Wen 0001, Daoqiang Zhang |
NeurIPS | 5 |
| 2025 | Cancer Survival Analysis via Zero-shot Tumor Microenvironment Segmentation on Low-resolution Whole Slide Pathology ImagesabstractThe whole-slide pathology images (WSIs) are widely recognized as the golden standard for cancer survival analysis. However, due to the high-resolution of WSIs, the existing studies require dividing WSIs into patches and identify key components before building the survival prediction system, which is time-consuming and cannot reflect the overall spatial organization of WSIs. Inspired by the fact that the spatial interactions among different tumor microenvironment (TME) components in WSIs are associated with the cancer prognosis, some studies attempt to capture the complex interactions among different TME components to improve survival predictions. However, they require extra efforts for building the TME segmentation model, which involves substantial annotation workloads on different TME components and is independent to the construction of the survival prediction model. To address the above issues, we propose ZTSurv, a novel end-to-end cancer survival analysis framework via efficient zero-shot TME segmentation on low-resolution WSIs. Specifically, by leveraging tumor infiltrating lymphocyte (TIL) maps on the 50x down-sampled WSIs, ZTSurv enables zero-shot segmentation on other two important TME components (i.e., tumor and stroma) that can reduce the annotation efforts from the pathologists. Then, based on the visual and semantic information extracted from different TME components, we construct a heterogeneous graph to capture their spatial intersections for clinical outcome prediction. We validate ZTSurv across four cancer cohorts derived from The Cancer Genome Atlas (TCGA), and the experimental results indicate that our method can not only achieve superior prediction results but also significantly reduce the computational costs in comparison with the state-of-the-art methods. Jiao Tang, Wei Shao 0005, Daoqiang Zhang |
NeurIPS | 2 |
| 2025 | MAPLE: Multi-scale Attribute-enhanced Prompt Learning for Few-shot Whole Slide Image ClassificationabstractPrompt learning has emerged as a promising paradigm for adapting pre-trained vision-language models (VLMs) to few-shot whole slide image (WSI) classification by aligning visual features with textual representations, thereby reducing annotation cost and enhancing model generalization. Nevertheless, existing methods typically rely on slide-level prompts and fail to capture the subtype-specific phenotypic variations of histological entities (e.g., nuclei, glands) that are critical for cancer diagnosis. To address this gap, we propose Multi-scale Attribute-enhanced Prompt Learning (MAPLE), a hierarchical framework for few-shot WSI classification that jointly integrates multi-scale visual semantics and performs prediction at both the entity and slide levels. Specifically, we first leverage large language models (LLMs) to generate entity-level prompts that can help identify multi-scale histological entities and their phenotypic attributes, as well as slide-level prompts to capture global visual descriptions. Then, an entity-guided cross-attention module is proposed to generate entity-level features, followed by aligning with their corresponding subtype-specific attributes for fine-grained entity-level prediction. To enrich entity representations, we further develop a cross-scale entity graph learning module that can update these representations by capturing their semantic correlations within and across scales. The refined representations are then aggregated into a slide-level representation and aligned with the corresponding prompts for slide-level prediction. Finally, we combine both entity-level and slide-level outputs to produce the final prediction results. Results on three cancer cohorts confirm the effectiveness of our approach in addressing few-shot pathology diagnosis tasks. Wei Shao 0005, Yagao Yue, Peng Wan 0004, Qi Zhu 0001, Daoqiang Zhang |
NeurIPS | 2 |
| 2025 | DemuxTrans: Transformer and temporal convolution network for accurate barcode demultiplexing in nanopore sequencingabstractMOTIVATION: Oxford Nanopore Technologies (ONT) direct RNA sequencing (dRNA-seq) offers high-resolution, single-molecule analysis but is hindered by the lack of robust multiplex barcoding methods. Existing approaches struggle to accurately demultiplex raw nanopore signals, failing to capture both local patterns and long-range dependencies. This limitation underscores the requirement for advanced solutions to improve accuracy, efficiency, and adaptability in sequencing workflows. We present DemuxTrans, a hybrid deep learning framework that integrates Multi-Layer Feature Fusion, Transformers, and Temporal Convolutional Networks (TCN) for precise barcode demultiplexing. RESULTS: DemuxTrans achieves state-of-the-art performance across multiple datasets by effectively balancing local feature extraction, global context modeling, and long-term dependency capture, excelling in metrics such as accuracy, recall and F1-score. These results demonstrate DemuxTrans as a scalable, efficient solution for barcode demultiplexing in nanopore sequencing, enabling precise identification of multiplexed RNA samples and improving throughput in transcriptomic and epigenomic analyses. AVAILABILITY AND IMPLEMENTATION: The code and datasets are publicly available on https://github.com/LiyuanShu116/Demuxtrans. Liyuan Shu, Deyu Zhuang, Jiao Tang, Junyong Zhao, Wei Shao 0005, Xiaoyu Guan, Daoqiang Zhang |
Bioinform. | 5 |
| 2025 | Edge-enhanced semi-supervised vertical convolutional neural network for tubular structure segmentation: Application to medical images
Junyong Zhao, Liang Sun 0009, Yanling Fu, Wei Shao 0005, Haipeng Si, Daoqiang Zhang |
Pattern Recognit. | 5 |
| 2025 | Multi-Modal Cross-Subject Emotion Feature Alignment and Recognition With EEG and Eye MovementsabstractMulti-modal emotion recognition has attracted much attention in human-computer interaction, because it provides complementary information for the recognition model. However, the distribution drift among subjects and the heterogeneity of different modalities pose challenges to multi-modal emotion recognition, thereby limiting its practical application. Most of the current multi-modal emotion recognition methods are difficult to suppress above uncertainties in fusion. In this paper, we propose a cross-subject multi-modal emotion recognition framework, which jointly learns subject-independent representation and common feature between EEG and eye movements. First, we design the dynamic adversarial domain adaptation for cross-subject distribution alignment, dynamically selecting source domains in training. Second, we simultaneously capture intra-modal and inter-modal emotion-related features by both self-attention and cross-attention mechanisms, thus obtaining the robust and complementary representation of emotional information. Then, two contrastive loss functions are imposed on above network to further reduce inter-modal heterogeneity, and mine higher-order semantic similarity between synchronously collected multi-modal data. Finally, we used the output of the softmax layer as the predicted value. The experimental results on several multi-modal emotion datasets with EEG and eye movements demonstrate that our method is significantly superior to the state-of-the-art emotion recognition approaches. Qi Zhu 0001, Lunke Fei, Chuhang Zheng, Wei Shao 0005, David Zhang 0001, Daoqiang Zhang |
IEEE Trans. Affect. Comput. | 5 |
| 2025 | Deep Ring-Wise Block Network for Joint Association Analysis and Alzheimer's Disease Diagnosis With InterpretabilityabstractIn the brain imaging genomic tasks, it is challenging to provide accurate prior knowledge for estimating the association between quantitative traits (QTs) extracted from neuroimaging and genetic markers like single-nucleotide polymorphisms (SNPs). The hidden structural patterns in data limit the discovery of disease-related biomarkers. To this end, we present a deep ring-wise block network (RB-Net) for association analysis and brain disease diagnosis. Specifically, we first construct a new hidden structural pattern, namely, ring-wise block pattern, that satisfies both block and ring properties within the data before the association analysis. Subsequently, a RB-Net is developed via using an auto-encoder (AE) to represent imaging genomic data. Furthermore, we design the approach for joint association learning and automated brain disease diagnosis. Additionally, the optimization scheme based on alternating update is presented for solve the built ring-wise block-perception layer model. The performance of the designed method has been experimentally assessed on the brain imaging genomic data from the Alzheimer's Disease Neuroimaging Initiative (ADNI). The results validate that the proposed approach outperforms some competing approaches, establishes strong associations, and identifies crucial regions of interest (ROIs) across different imaging phenotypes associated with genetic risk biomarkers, thereby guiding disease interpretation and diagnosis prediction. Meiling Wang 0001, Wei Shao 0005, Daoqiang Zhang, Qingshan Liu 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 2 |
| 2025 | Spatio-Temporal Evolutionary Graph Learning for Brain Network Analysis Using Medical ImagingabstractDynamic functional brain network (DFBN) can flexibly describe the time-varying topological connectivity patterns of the brain, and show great potential in brain disease diagnosis. However, most of the existing DFBN analysis methods focus on capturing the dynamic interaction at the brain region level, ignoring the spatio-temporal topological evolution across time windows. Moreover, they are difficult to suppress interfering connections in DFBNs, which leads to a diminished capacity for discerning the intrinsic structures that are intimately linked to brain disorders. To address these issues, we propose a topological evolution graph learning model to capture disease-related spatio-temporal topological features in DFBNs. Specifically, we first take the hubness of adjacent DFBN as the source domain and the target domain in turn, and then use Wasserstein distance (WD) and Gromov-Wasserstein distance (GWD) to capture the brain's evolution law at the node and edge levels, respectively. Furthermore, we introduce the principle of relevant information to guide the topology evolution graph to learn the structures that are most relevant to brain diseases yet least redundant information between adjacent DFBNs. On this basis, we develop a high-order spatio-temporal model with multi-hop graph convolution to collaboratively extract long-range spatial and temporal dependencies from the topological evolution graph. Extensive experiments show that the proposed method outperforms the current state-of-the-art methods, and can effectively reveal the information evolution mechanism between brain regions across windows. Shengrong Li, Qi Zhu 0001, Chunwei Tian, Li Zhang 0057, Chuhang Zheng, Daoqiang Zhang, Wei Shao 0005 |
IEEE Trans. Image Process. | 8 |
| 2025 | Interpretable Dynamic Brain Network Analysis With Functional and Structural PriorsabstractThe dynamic functional brain network (DFBN) inherently captures topological changes in brain connectivity pattern during activity, attracting increasing attention for detecting brain disorders. However, most current DFBN analysis methods rely on data-driven modeling and ignore crucial prior knowledge of brain structure and function, resulting in weak interpretability of models. Furthermore, effectively extracting dynamic topological features from DFBN is still a challenging issue, due to its intricate spatio-temporal features coupling. In this paper, we propose an interpretable spatio-temporal tensor graph convolutional network for DFBN analysis. Firstly, by incorporating functional and structural priors into the construction of DBFN, we develop a hierarchical DBFN representation with brain region clustering that effectively captures the spatio-temporal topology among subnetworks. Secondly, we design a tensor graph convolutional network with both intra-graph propagation and inter-graph propagation to simultaneously extract the spatio-temporal features from the hierarchical DFBN. Additionally, we derive a functional subnetwork constraint to enhance the consistency within subnetworks and the differences between subnetworks, which guides the learned features to better reflect the topology prior of the brain network. Finally, self-attention is employed to fuse the learned dynamic topological features of different subnetworks for classification. Experimental results on epilepsy, ADNI and ABIDE datasets demonstrate that our method achieves competitive diagnostic performance and offers network-level interpretability for brain disease diagnosis. Shengrong Li, Qi Zhu 0001, Chunwei Tian, Wei Shao 0005, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Discovering Differential Imaging Genetic Modules via Multimodal Fusion-Based Hypergraph Transductive Learning in Alzheimer's Disease DiagnosisabstractBrain imaging genetics is a widely focused topic, which has achieved the great successes in the diagnosis of complex brain disorders. In clinical practice, most existing data fusion approaches extract features from homogeneous data, neglecting the heterogeneous structural information among imaging genetic data. In addition, the number of labeled samples is limited due to the cost and time of manually labeling data. To remedy such deficiencies, in this work, we present a multimodal fusion-based hypergraph transductive learning (MFHT) for clinical diagnosis. Specifically, for each modality, we first construct a corresponding similarity graph to reflect the similarity between subjects using the label prior. Then, the multiple graph fusion approach based on theoretical convergence guarantee is designed for learning a unified graph harnessing the structure of entire data. Finally, to fully exploit the rich information of the obtained graph, a hypergraph transductive learning approach is designed to effectively capture the complex structures and high-order relationships in both labeled and unlabeled data to achieve the diagnosis results. The brain imaging genetic data of the Alzheimer's Disease Neuroimaging Initiative (ADNI) datasets are used to experimentally explore our developed method. Related results show that our method is well applied to the analysis of brain imaging genetic data, which accounts for genetics, brain imaging (region of interest (ROI) node features), and brain imaging (connectivity edge features) to boost the understanding of disease mechanism as well as improve clinical diagnosis. Meiling Wang 0001, Liang Sun 0009, Wei Shao 0005, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 3 |
| 2025 | TAFL: Task-Agnostic Feature Learner for Efficient Adaptation to Unseen Clinical Tasks Based on Whole-Slide Histopathological ImagesabstractMulti-task learning (MTL) has become a research hotspot for the analysis of whole-slide histopathological images (WSIs) since it can capture the shared representations of different tasks for the improvement of individual tasks. However, the shared representations learned by MTL are always dominated by the tasks appearing in the training set that is difficult to directly apply it on the unseen (new) tasks, especially when the unseen tasks are significantly different from the known tasks. To address the above issues, we develop a Task-Agnostic Feature-Learner (TAFL) for efficient adaptation to unseen clinical tasks, which can leverage useful image information from the existing tasks for new clinical trials with minimal task-specific modifications. Specifically, we firstly develop a neural architecture search (NAS) module that can design the network architectures of TAFL automatically. Then, a novel task-level meta-learning algorithm is developed to extract efficient and universal information from the known tasks for improving the prediction performance on the unseen tasks. We evaluate our method on three publicly available datasets derived from The Cancer Genome Atlas (TCGA) for various clinical prediction tasks (i.e., staging, cancer subtyping and survival prediction), and the experimental results indicate that our TAFL can effectively adapt to unseen tasks with better prediction performance. Yingli Zuo, Lianyu Wang, Shichang Feng, Qi Zhu 0001, Wei Shao 0005, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 7 |
| 2024 | Tumor Micro-Environment Interactions Guided Graph Learning for Survival Analysis of Human Cancers from Whole-Slide Pathological ImagesabstractThe recent advance of deep learning technology brings the possibility of assisting the pathologist to predict the patients' survival from whole-slide pathological images (WSIs). However, most of the prevalent methods only worked on the sampled patches in specifically or randomly selected tumor areas of WSIs, which has very limited capability to capture the complex interactions between tumor and its surrounding micro-environment components. As a matter of fact, tumor is supported and nurtured in the heterogeneous tumor micro-environment(TME), and the detailed analysis of TME and their correlation with tumors are important to in-depth analyze the mechanism of cancer development. In this paper, we considered the spatial interactions among tumor and its two major TME components (i.e., lymphocytes and stromal fibrosis) and presented a Tumor Micro-environment Interactions Guided Graph Learning (TMEGL) algorithm for the prognosis prediction of human cancers. Specifically, we firstly selected different types of patches as nodes to build graph for each WSI. Then, a novel TME neighborhood organization guided graph embedding algorithm was proposed to learn node representations that can preserve their topological structure information. Finally, a Gated Graph Attention Network is applied to capture the survival-associated intersections among tumor and different TME components for clinical outcome prediction. We tested TMEGL on three cancer cohorts derived from The Cancer Genome Atlas (TCGA), and the experimental results indicated that TMEGL not only outperforms the existing WSI-based survival analysis models, but also has good explainable ability for survival prediction. Wei Shao 0005, Yangyang Shi, Daoqiang Zhang, Peng Wan 0004 |
CVPR | 1 |
| 2024 | OSAL-ND: Open-Set Active Learning for Nucleus Detection
Jiao Tang, Yagao Yue, Peng Wan 0004, Daoqiang Zhang, Wei Shao 0005 |
MICCAI (4) | 6 |
| 2024 | Correlation-Adaptive Multi-view CEUS Fusion for Liver Cancer Diagnosis
Peng Wan 0004, Shukang Zhang, Wei Shao 0005, Junyong Zhao, Yinkai Yang, Wentao Kong, Haiyan Xue, Daoqiang Zhang |
MICCAI (5) | 3 |
| 2024 | T-S2Inet: Transformer-based sequence-to-image network for accurate nanopore sequence recognitionabstractMOTIVATION: Nanopore sequencing is a new macromolecular recognition and perception technology that enables high-throughput sequencing of DNA, RNA, even protein molecules. The sequences generated by nanopore sequencing span a large time frame, and the labor and time costs incurred by traditional analysis methods are substantial. Recently, research on nanopore data analysis using machine learning algorithms has gained unceasing momentum, but there is often a significant gap between traditional and deep learning methods in terms of classification results. To analyze nanopore data using deep learning technologies, measures such as sequence completion and sequence transformation can be employed. However, these technologies do not preserve the local features of the sequences. To address this issue, we propose a sequence-to-image (S2I) module that transforms sequences of unequal length into images. Additionally, we propose the Transformer-based T-S2Inet model to capture the important information and improve the classification accuracy. RESULTS: Quantitative and qualitative analysis shows that the experimental results have an improvement of around 2% in accuracy compared to previous methods. The proposed method is adaptable to other nanopore platforms, such as the Oxford nanopore. It is worth noting that the proposed method not only aims to achieve the most advanced performance, but also provides a general idea for the analysis of nanopore sequences of unequal length. AVAILABILITY AND IMPLEMENTATION: The main program is available at https://github.com/guanxiaoyu11/S2Inet. Xiaoyu Guan, Wei Shao 0005, Daoqiang Zhang |
Bioinform. | 2 |
| 2024 | SAM-Y: Attention-enhanced hazardous vehicle object detection algorithmabstractAbstract Vehicle transportation of hazardous chemicals is one of the important mobile hazards in modern logistics, and its unsafe factors bring serious threats to people's lives, property and environmental safety. Although the current object detection algorithm has certain applications in the detection of hazardous chemical vehicles, due to the complexity of the transportation environment, the small size and low resolution of the vehicle target etc., object detection becomes more difficult in the face of a complex background. In order to solve these problems, the authors propose an improved algorithm based on YOLOv5 to enhance the detection accuracy and efficiency of hazardous chemical vehicles. Firstly, in order to better capture the details and semantic information of hazardous chemical vehicles, the algorithm solves the problem of mismatch between the receptive field of the detector and the target object by introducing the receptive field expansion block into the backbone network, so as to improve the ability of the model to capture the detailed information of hazardous chemical vehicles. Secondly, in order to improve the ability of the model to express the characteristics of hazardous chemical vehicles, the authors introduce a separable attention mechanism in the multi‐scale target detection stage, and enhances the prediction ability of the model by combining the object detection head and attention mechanism coherently in the feature layer of scale perception, the spatial location of spatial perception and the output channel of task perception. Experimental results show that the improved model significantly surpasses the baseline model in terms of accuracy and achieves more accurate object detection. At the same time, the model also has a certain improvement in inference speed and achieves faster inference ability. Bushi Liu, Xian-Chun Meng, Bolun Chen, Wei Shao 0005, Liqing Chen |
IET Comput. Vis. | 6 |
| 2024 | Global-local consistent semi-supervised segmentation of histopathological image with different perturbations
Xi Guan, Qi Zhu 0001, Liang Sun 0009, Junyong Zhao, Daoqiang Zhang, Peng Wan 0004, Wei Shao 0005 |
Pattern Recognit. | 7 |
| 2024 | Dynamic Confidence-Aware Multi-Modal Emotion RecognitionabstractMulti-modal emotion recognition has attracted increasing attention in human-computer interaction, as it extracts complementary information from physiological and behavioral features. Compared to single modal approaches, multi-modal fusion methods are more susceptible to uncertainty in emotion recognition, such as heterogeneity and inconsistent predictions across different modalities. Previous multi-modal approaches ignore systematic modeling of uncertainty in fusion and revelation of dynamic variations in emotion process. In this paper, we propose a dynamic confidence-aware fusion network for robust recognition of heterogeneous emotion features, including electroencephalogram (EEG) and facial expression. First, we develop a self-attention based multi-channel LSTM network to preliminarily align the heterogeneous emotion features. Second, we propose a confidence regression network to estimate true class probability (TCP) on each modality, which helps explore the uncertainty at modality level. Then, different modalities are weighted fused according to above two types of uncertainty. Finally, we adopt self-paced learning (SPL) mechanism to further improve the model robustness by alleviating negative effect from the hard learning samples. The experimental results on several multi-modal emotion datasets demonstrate the proposed method outperforms the state-of-the-art methods in emotion recognition performance and explicitly reveals the dynamic variation of emotion with uncertainty estimation. Our code is available at: Qi Zhu 0001, Chuhang Zheng, Zheng Zhang 0006, Wei Shao 0005, Daoqiang Zhang |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Discriminative Domain Adaption Network for Simultaneously Removing Batch Effects and Annotating Cell Types in Single-Cell RNA-SeqabstractMachine learning techniques have become increasingly important in analyzing single-cell RNA and identifying cell types, providing valuable insights into cellular development and disease mechanisms. However, the presence of batch effects poses major challenges in scRNA-seq analysis due to data distribution variation across batches. Although several batch effect mitigation algorithms have been proposed, most of them focus only on the correlation of local structure embeddings, ignoring global distribution matching and discriminative feature representation in batch correction. In this paper, we proposed the discriminative domain adaption network (D2AN) for joint batch effects correction and type annotation with single-cell RNA-seq. Specifically, we first captured the global low-dimensional embeddings of samples from the source and target domains by adversarial domain adaption strategy. Second, a contrastive loss is developed to preliminarily align the source domain samples. Moreover, the semantic alignment of class centroids in the source and target domains is achieved for further local alignment. Finally, a self-paced learning mechanism based on inter-domain loss is adopted to gradually select samples with high similarity to the target domain for training, which is used to improve the robustness of the model. Experimental results demonstrated that the proposed method on multiple real datasets outperforms several state-of-the-art methods. Qi Zhu 0001, Aizhen Li, Zheng Zhang 0006, Chuhang Zheng, Junyong Zhao, Jin-Xing Liu 0001, Daoqiang Zhang, Wei Shao 0005 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2024 | MAS-CL: An End-to-End Multi-Atlas Supervised Contrastive Learning Framework for Brain ROI SegmentationabstractBrain region-of-interest (ROI) segmentation with magnetic resonance (MR) images is a basic prerequisite step for brain analysis. The main problem with using deep learning for brain ROI segmentation is the lack of sufficient annotated data. To address this issue, in this paper, we propose a simple multi-atlas supervised contrastive learning framework (MAS-CL) for brain ROI segmentation with MR images in an end-to-end manner. Specifically, our MAS-CL framework mainly consists of two steps, including 1) a multi-atlas supervised contrastive learning method to learn the latent representation using a limited amount of voxel-level labeling brain MR images, and 2) brain ROI segmentation based on the pre-trained backbone using our MSA-CL method. Specifically, different from traditional contrastive learning, in our proposed method, we use multi-atlas supervised information to pre-train the backbone for learning the latent representation of input MR image, i.e., the correlation of each sample pair is defined by using the label maps of input MR image and atlas images. Then, we extend the pre-trained backbone to segment brain ROI with MR images. We perform our proposed MAS-CL framework with five segmentation methods on LONI-LPBA40, IXI, OASIS, ADNI, and CC359 datasets for brain ROI segmentation with MR images. Various experimental results suggested that our proposed MAS-CL framework can significantly improve the segmentation performance on these five datasets. Liang Sun 0009, Yanling Fu, Junyong Zhao, Wei Shao 0005, Qi Zhu 0001, Daoqiang Zhang |
IEEE Trans. Image Process. | 4 |
| 2024 | Multi-Instance Multi-Task Learning for Joint Clinical Outcome and Genomic Profile Predictions From the Histopathological ImagesabstractWith the remarkable success of digital histopathology and the deep learning technology, many whole-slide pathological images (WSIs) based deep learning models are designed to help pathologists diagnose human cancers. Recently, rather than predicting categorical variables as in cancer diagnosis, several deep learning studies are also proposed to estimate the continuous variables such as the patients' survival or their transcriptional profile. However, most of the existing studies focus on conducting these predicting tasks separately, which overlooks the useful intrinsic correlation among them that can boost the prediction performance of each individual task. In addition, it is sill challenge to design the WSI-based deep learning models, since a WSI is with huge size but annotated with coarse label. In this study, we propose a general multi-instance multi-task learning framework (HistMIMT) for multi-purpose prediction from WSIs. Specifically, we firstly propose a novel multi-instance learning module (TMICS) considering both common and specific task information across different tasks to generate bag representation for each individual task. Then, a soft-mask based fusion module with channel attention (SFCA) is developed to leverage useful information from the related tasks to help improve the prediction performance on target task. We evaluate our method on three cancer cohorts derived from the Cancer Genome Atlas (TCGA). For each cohort, our multi-purpose prediction tasks range from cancer diagnosis, survival prediction and estimating the transcriptional profile of gene TP53. The experimental results demonstrated that HistMIMT can yield better outcome on all clinical prediction tasks than its competitors. Wei Shao 0005, Yingli Zuo, Liang Sun 0009, Tiansong Xia, Wanyuan Chen, Peng Wan 0004, Jianpeng Sheng, Qi Zhu 0001, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 1 |
| 2024 | Spatio-Temporal Graph Hubness Propagation Model for Dynamic Brain Network ClassificationabstractDynamic brain network has the advantage over static brain network in characterizing the variation pattern of functional brain connectivity, and it has attracted increasing attention in brain disease diagnosis. However, most of the existing dynamic brain networks analysis methods rely on extracting features from independent brain networks divided by sliding windows, making them hard to reveal the high-order dynamic evolution laws of functional brain networks. Additionally, they cannot effectively extract the spatio-temporal topology features in dynamic brain networks. In this paper, we propose to use optimal transport (OT) theory to capture the topology evolution of the dynamic brain networks, and develop a multi-channel spatio-temporal graph convolutional network that collaboratively extracts the temporal and spatial features from the evolution networks. Specifically, we first adaptively evaluate the graph hubness of brain regions in the brain network of each time window, which comprehensively models information transmission among multiple brain regions. Second, the hubness propagation information across adjacent time windows is captured by optimal transport, describing high-order topology evolution of dynamic brain networks. Moreover, we develop a spatio-temporal graph convolutional network with attention mechanism to collaboratively extract the intrinsic temporal and spatial topology information from the above networks. Finally, the multi-layer perceptron is adopted for classifying the dynamic brain network. The extensive experiment on the collected epilepsy dataset and the public ADNI dataset show that our proposed method not only outperforms several state-of-the-art methods in brain disease diagnosis, but also reveals the key dynamic alterations of brain connectivities between patients and healthy controls. Qi Zhu 0001, Shengrong Li, Xiangshui Meng, Wei Shao 0005, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 6 |
| 2023 | Transfer Learning-Assisted Survival Analysis of Breast Cancer Relying on the Spatial Interaction Between Tumor-Infiltrating Lymphocytes and Tumors
Yawen Wu, Yingli Zuo, Qi Zhu 0001, Jianpeng Sheng, Daoqiang Zhang, Wei Shao 0005 |
MICCAI (6) | 6 |
| 2023 | Prior-Driven Dynamic Brain Networks for Multi-modal Emotion Recognition
Chuhang Zheng, Wei Shao 0005, Daoqiang Zhang, Qi Zhu 0001 |
MICCAI (8) | 2 |
| 2023 | Active learning for efficient analysis of high-throughput nanopore dataabstractMOTIVATION: As the third-generation sequencing technology, nanopore sequencing has been used for high-throughput sequencing of DNA, RNA, and even proteins. Recently, many studies have begun to use machine learning technology to analyze the enormous data generated by nanopores. Unfortunately, the success of this technology is due to the extensive labeled data, which often suffer from enormous labor costs. Therefore, there is an urgent need for a novel technology that can not only rapidly analyze nanopore data with high-throughput, but also significantly reduce the cost of labeling. To achieve the above goals, we introduce active learning to alleviate the enormous labor costs by selecting the samples that need to be labeled. This work applies several advanced active learning technologies to the nanopore data, including the RNA classification dataset (RNA-CD) and the Oxford Nanopore Technologies barcode dataset (ONT-BD). Due to the complexity of the nanopore data (with noise sequence), the bias constraint is introduced to improve the sample selection strategy in active learning. Results: The experimental results show that for the same performance metric, 50% labeling amount can achieve the best baseline performance for ONT-BD, while only 15% labeling amount can achieve the best baseline performance for RNA-CD. Crucially, the experiments show that active learning technology can assist experts in labeling samples, and significantly reduce the labeling cost. Active learning can greatly reduce the dilemma of difficult labeling of high-capacity nanopore data. We hope active learning can be applied to other problems in nanopore sequence analysis. AVAILABILITY AND IMPLEMENTATION: The main program is available at https://github.com/guanxiaoyu11/AL-for-nanopore. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaoyu Guan, Zhongnian Li, Yueying Zhou, Wei Shao 0005, Daoqiang Zhang |
Bioinform. | 4 |
| 2023 | Hypergraph-regularized multimodal learning by graph diffusion for imaging genetics based Alzheimer's Disease diagnosis
Meiling Wang 0001, Wei Shao 0005, Shuo Huang 0001, Daoqiang Zhang |
Medical Image Anal. | 2 |
| 2023 | Multi-scale multi-hierarchy attention convolutional neural network for fetal brain extraction
Liang Sun 0009, Wei Shao 0005, Qi Zhu 0001, Meiling Wang 0001, Gang Li 0001, Daoqiang Zhang |
Pattern Recognit. | 2 |
| 2023 | Self-Supervised Federated Adaptation for Multi-Site Brain Disease DiagnosisabstractThe multi-site approach has attracted increasing attention in brain disease diagnosis, because it can improve the prediction performance by integrating sample information from different medical institutions. However, its training procedure requires the transmission of subject's original images or features among sites, which may cause privacy disclosure. In this paper, we propose a self-supervised federated adaptation (S2FA) framework for robust multi-site prediction, which can reduce the risk of privacy disclosure. As far as we know, it is the first work to investigate the cross-site brain disease diagnosis, which trains model on source sites and tests on target site, often occurring in clinical practice. Firstly, we implement a decentralized federated optimization strategy, by which each site communicates model parameters periodically. Secondly, we construct an auxiliary self-supervised model for target site through transferring knowledge from source sites with self-paced learning. Then, a hash mapping is proposed to encode the target feature, simultaneously reducing the risk of privacy information disclosure and alleviating data heterogeneity among sites. Finally, we achieve the cross-site prediction by weighted federated source model and auxiliary target model. Experimental results on multi-site datasets show that the proposed S2FA can accurately identify brain disease. Our codes are available athttps://github.com/nuaayqm/S2FA. Qi Zhu 0001, Wei Shao 0005, Zheng Zhang 0006, Daoqiang Zhang |
IEEE Trans. Big Data | 4 |
| 2023 | Multi-Discriminator Active Adversarial Network for Multi-Center Brain Disease DiagnosisabstractMulti-center analysis has attracted increasing attention in brain disease diagnosis, because it provides effective approaches to improve disease diagnostic performance by making use of the information from different centers. However, in practical multi-center applications, data uncertainty is more common than that in single center, which brings challenge to robust modeling of diagnosis. In this article, we proposed a multi-discriminator active adversarial network (MDAAN) to alleviate the uncertainties at the center, feature, and label levels for multi-center brain disease diagnosis. First, we extract the latent invariant representation of the source center and target center to reduce domain shift by adversarial learning strategy. Second, the proposed method adaptively evaluates the contribution of different source centers in fusion by measuring data distribution difference between source and target center. Moreover, only the hard learning samples in target center are identified to label with low sample annotation cost. Finally, we treat the selected samples as the auxiliary domain to alleviate the negative transfer and improve the robustness of the multi-center model. We extensively compare the proposed approach with several state-of-the-art multi-center methods on the five-center schizophrenia dataset, and the results demonstrate that our method is superior to the previous methods in identifying brain disease. Qi Zhu 0001, Xiangyu Xu 0003, Yuwu Lu, Wei Shao 0005, Daoqiang Zhang |
IEEE Trans. Big Data | 6 |
| 2023 | FAM3L: Feature-Aware Multi-Modal Metric Learning for Integrative Survival Analysis of Human CancersabstractSurvival analysis is to estimate the survival time for an individual or a group of patients, which is a valid solution for cancer treatments. Recent studies suggested that the integrative analysis of histopathological images and genomic data can better predict the survival of cancer patients than simply using single bio-marker, for different bio-markers may provide complementary information. However, for the given multi-modal data that may contain irrelevant or redundant features, it is still challenge to design a distance metric that can simultaneously discover significant features and measure the difference of survival time among different patients. To solve this issue, we propose a Feature-Aware Multi-modal Metric Learning method (FAM3L), which not only learns the metric for distance constraints on patients' survival time, but also identifies important images and genomic features for survival analysis. Specifically, for each modality of data, we firstly design one feature-aware metric that can be decoupled into a traditional distance metric and a diagonal weight for important feature identification. Then, in order to explore the complex correlation across multiple modality data, we apply Hilbert-Schmidt Independence Criterion (HSIC) to jointly learn multiple metrics. Finally, based on the learned distance metrics, we apply the Cox proportional hazards model for prognosis prediction. We evaluate the performance of our proposed FAM3L method on three cancer cohorts derived from The Cancer Genome Atlas (TCGA), the experimental results demonstrate that our method can not only achieve superior performance for cancer prognosis, but also identify meaningful image and genomic features correlating strongly with cancer survival. Wei Shao 0005, Yingli Zuo, Shile Qi, Honghai Hong, Jianpeng Sheng, Qi Zhu 0001, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 1 |
| 2023 | Characterizing the Survival-Associated Interactions Between Tumor-Infiltrating Lymphocytes and Tumors From Pathological Images and Multi-Omics DataabstractThe tumor-infiltrating lymphocytes (TILs) and its correlation with tumors have shown significant values in the development of cancers. Many observations indicated that the combination of the whole-slide pathological images (WSIs) and genomic data can better characterize the immunological mechanisms of TILs. However, the existing image-genomic studies evaluated the TILs by the combination of pathological image and single-type of omics data (e.g., mRNA), which is difficulty in assessing the underlying molecular processes of TILs holistically. Additionally, it is still very challenging to characterize the intersections between TILs and tumor regions in WSIs and the high dimensional genomic data also brings difficulty for the integrative analysis with WSIs. Based on the above considerations, we proposed an end-to-end deep learning framework i.e., IMO-TILs that can integrate pathological image with multi-omics data (i.e., mRNA and miRNA) to analyze TILs and explore the survival-associated interactions between TILs and tumors. Specifically, we firstly apply the graph attention network to describe the spatial interactions between TILs and tumor regions in WSIs. As to genomic data, the Concrete AutoEncoder (i.e., CAE) is adopted to select survival-associated Eigengenes from the high-dimensional multi-omics data. Finally, the deep generalized canonical correlation analysis (DGCCA) accompanied with the attention layer is implemented to fuse the image and multi-omics data for prognosis prediction of human cancers. The experimental results on three cancer cohorts derived from the Cancer Genome Atlas (TCGA) indicated that our method can both achieve higher prognosis results and identify consistent imaging and multi-omics bio-markers correlated strongly with the prognosis of human cancers. Wei Shao 0005, Yingli Zuo, Yangyang Shi, Yawen Wu, Jiao Tang, Junyong Zhao, Liang Sun 0009, Zixiao Lu, Jianpeng Sheng, Qi Zhu 0001, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 1 |
| 2023 | Deep Multi-Modal Discriminative and Interpretability Network for Alzheimer's Disease DiagnosisabstractMulti-modal fusion has become an important data analysis technology in Alzheimer's disease (AD) diagnosis, which is committed to effectively extract and utilize complementary information among different modalities. However, most of the existing fusion methods focus on pursuing common feature representation by transformation, and ignore discriminative structural information among samples. In addition, most fusion methods use high-order feature extraction, such as deep neural network, by which it is difficult to identify biomarkers. In this paper, we propose a novel method named deep multi-modal discriminative and interpretability network (DMDIN), which aligns samples in a discriminative common space and provides a new approach to identify significant brain regions (ROIs) in AD diagnosis. Specifically, we reconstruct each modality with a hierarchical representation through multilayer perceptron (MLP), and take advantage of the shared self-expression coefficients constrained by diagonal blocks to embed the structural information of inter-class and the intra-class. Further, the generalized canonical correlation analysis (GCCA) is adopted as a correlation constraint to generate a discriminative common space, in which samples of the same category gather while samples of different categories stay away. Finally, in order to enhance the interpretability of the deep learning model, we utilize knowledge distillation to reproduce coordinated representations and capture influence of brain regions in AD classification. Experiments show that the proposed method performs better than several state-of-the-art methods in AD diagnosis. Qi Zhu 0001, Bingliang Xu, Jiashuang Huang, Heyang Wang, Ruting Xu, Wei Shao 0005, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 6 |
| 2022 | Identify Consistent Imaging Genomic Biomarkers for Characterizing the Survival-Associated Interactions Between Tumor-Infiltrating Lymphocytes and Tumors
Yingli Zuo, Yawen Wu, Zixiao Lu, Qi Zhu 0001, Kun Huang 0001, Daoqiang Zhang, Wei Shao 0005 |
MICCAI (2) | 7 |
| 2022 | S2Snet: deep learning for low molecular weight RNA identification with nanoporeabstractRibonucleic acid (RNA) is a pivotal nucleic acid that plays a crucial role in regulating many biological activities. Recently, one study utilized a machine learning algorithm to automatically classify RNA structural events generated by a Mycobacterium smegmatis porin A nanopore trap. Although it can achieve desirable classification results, compared with deep learning (DL) methods, this classic machine learning requires domain knowledge to manually extract features, which is sophisticated, labor-intensive and time-consuming. Meanwhile, the generated original RNA structural events are not strictly equal in length, which is incompatible with the input requirements of DL models. To alleviate this issue, we propose a sequence-to-sequence (S2S) module that transforms the unequal length sequence (UELS) to the equal length sequence. Furthermore, to automatically extract features from the RNA structural events, we propose a sequence-to-sequence neural network based on DL. In addition, we add an attention mechanism to capture vital information for classification, such as dwell time and blockage amplitude. Through quantitative and qualitative analysis, the experimental results have achieved about a 2% performance increase (accuracy) compared to the previous method. The proposed method can also be applied to other nanopore platforms, such as the famous Oxford nanopore. It is worth noting that the proposed method is not only aimed at pursuing state-of-the-art performance but also provides an overall idea to process nanopore data with UELS. Xiaoyu Guan, Wei Shao 0005, Zhongnian Li, Shuo Huang 0001, Daoqiang Zhang |
Briefings Bioinform. | 3 |
| 2022 | Identify connectome between genotypes and brain network phenotypes via deep self-reconstruction sparse canonical correlation analysisabstractMOTIVATION: As a rising research topic, brain imaging genetics aims to investigate the potential genetic architecture of both brain structure and function. It should be noted that in the brain, not all variations are deservedly caused by genetic effect, and it is generally unknown which imaging phenotypes are promising for genetic analysis. RESULTS: In this work, genetic variants (i.e. the single nucleotide polymorphism, SNP) can be correlated with brain networks (i.e. quantitative trait, QT), so that the connectome (including the brain regions and connectivity features) of functional brain networks from the functional magnetic resonance imaging data is identified. Specifically, a connection matrix is firstly constructed, whose upper triangle elements are selected to be connectivity features. Then, the PageRank algorithm is exploited for estimating the importance of different brain regions as the brain region features. Finally, a deep self-reconstruction sparse canonical correlation analysis (DS-SCCA) method is developed for the identification of genetic associations with functional connectivity phenotypic markers. This approach is a regularized, deep extension, scalable multi-SNP-multi-QT method, which is well-suited for applying imaging genetic association analysis to the Alzheimer's Disease Neuroimaging Initiative datasets. It is further optimized by adopting a parametric approach, augmented Lagrange and stochastic gradient descent. Extensive experiments are provided to validate that the DS-SCCA approach realizes strong associations and discovers functional connectivity and brain region phenotypic biomarkers to guide disease interpretation. AVAILABILITY AND IMPLEMENTATION: The Matlab code is available at https://github.com/meimeiling/DS-SCCA/tree/main. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Meiling Wang 0001, Wei Shao 0005, Xiaoke Hao, Shuo Huang 0001, Daoqiang Zhang |
Bioinform. | 2 |
| 2022 | Application of unsupervised deep learning algorithms for identification of specific clusters of chronic cough patients from EMR dataabstractBACKGROUND: Chronic cough affects approximately 10% of adults. The lack of ICD codes for chronic cough makes it challenging to apply supervised learning methods to predict the characteristics of chronic cough patients, thereby requiring the identification of chronic cough patients by other mechanisms. We developed a deep clustering algorithm with auto-encoder embedding (DCAE) to identify clusters of chronic cough patients based on data from a large cohort of 264,146 patients from the Electronic Medical Records (EMR) system. We constructed features using the diagnosis within the EMR, then built a clustering-oriented loss function directly on embedded features of the deep autoencoder to jointly perform feature refinement and cluster assignment. Lastly, we performed statistical analysis on the identified clusters to characterize the chronic cough patients compared to the non-chronic cough patients. RESULTS: The experimental results show that the DCAE model generated three chronic cough clusters and one non-chronic cough patient cluster. We found various diagnoses, medications, and lab tests highly associated with chronic cough patients by comparing the chronic cough cluster with the non-chronic cough cluster. Comparison of chronic cough clusters demonstrated that certain combinations of medications and diagnoses characterize some chronic cough clusters. CONCLUSIONS: To the best of our knowledge, this study is the first to test the potential of unsupervised deep learning methods for chronic cough investigation, which also shows a great advantage over existing algorithms for patient data clustering. Wei Shao 0005, Xiao Luo 0002, Zuoyi Zhang, Zhi Han, Vasu Chandrasekaran, Vladimir Turzhitsky, Vishal Bali, Anna R. Roberts, Megan Metzger, Jarod Baker, Carmen La Rosa, Jessica Weaver, Paul Richard Dexter, Kun Huang 0001 |
BMC Bioinform. | 1 |
| 2022 | Multimodal Triplet Attention Network for Brain Disease DiagnosisabstractMulti-modal imaging data fusion has attracted much attention in medical data analysis because it can provide complementary information for more accurate analysis. Integrating functional and structural multi-modal imaging data has been increasingly used in the diagnosis of brain diseases, such as epilepsy. Most of the existing methods focus on the feature space fusion of different modalities but ignore the valuable high-order relationships among samples and the discriminative fused features for classification. In this paper, we propose a novel framework by fusing data from two modalities of functional MRI (fMRI) and diffusion tensor imaging (DTI) for epilepsy diagnosis, which effectively captures the complementary information and discriminative features from different modalities by high-order feature extraction with the attention mechanism. Specifically, we propose a triple network to explore the discriminative information from the high-order representation feature space learned from multi-modal data. Meanwhile, self-attention is introduced to adaptively estimate the degree of importance between brain regions, and the cross-attention mechanism is utilized to extract complementary information from fMRI and DTI. Finally, we use the triple loss function to adjust the distance between samples in the common representation space. We evaluate the proposed method on the epilepsy dataset collected from Jinling Hospital, and the experiment results demonstrate that our method is significantly superior to several state-of-the-art diagnosis approaches. Qi Zhu 0001, Heyang Wang, Bingliang Xu, Wei Shao 0005, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 5 |
| 2021 | Transfer Learning via Optimal Transportation for Integrative Cancer Patient StratificationabstractThe Stratification of early-stage cancer patients for the prediction of clinical outcome is a challenging task since cancer is associated with various molecular aberrations. A single biomarker often cannot provide sufficient information to stratify early-stage patients effectively. Understanding the complex mechanism behind cancer development calls for exploiting biomarkers from multiple modalities of data such as histopathology images and genomic data. The integrative analysis of these biomarkers sheds light on cancer diagnosis, subtyping, and prognosis. Another difficulty is that labels for early-stage cancer patients are scarce and not reliable enough for predicting survival times. Given the fact that different cancer types share some commonalities, we explore if the knowledge learned from one cancer type can be utilized to improve prognosis accuracy for another cancer type. We propose a novel unsupervised multi-view transfer learning algorithm to simultaneously analyze multiple biomarkers in different cancer types. We integrate multiple views using non-negative matrix factorization and formulate the transfer learning model based on the Optimal Transport theory to align features of different cancer types. We evaluate the stratification performance on three early-stage cancers from the Cancer Genome Atlas (TCGA) project. Comparing with other benchmark methods, our framework achieves superior accuracy for patient outcome prediction. Wei Shao 0005, Jie Zhang 0010, Kun Huang 0001 |
IJCAI | 2 |
| 2021 | Sign-aware Perturbations RegressionabstractThis paper presents the first study on Sign-aware Perturbations Regression (SaPR), where the observed response variables contain the aware sign (negative or positive) perturbations.In order to predict the non-perturbation response variables, we propose a novel parameter estimator SZOM (i.e.,Setting Zero Operator Method), which aims at taking full advantage of the aware perturbations information to correct the mistake values in the estimation process with computationally efficiency.In this paper, the two aspects of theoretical analysis are proposed to deeply understand our method.Firstly, we establish the perturbation parameter error upper bound and prove consistency guarantee in the linear regression scenario.Secondly, we introduce the generalization error bound for the proposed SZMO, which indicates that the error bound is related to the value and the number of negative and positive perturbations.The effectiveness of the proposed approach is well validated by the experimental results on both synthetic and real datasets. Zhongnian Li, Tao Zhang 0099, Wei Shao 0005, Songcan Chen, Daoqiang Zhang |
SDM | 3 |
| 2021 | Towards Fair Cross-Domain Adaptation via Generative LearningabstractDomain Adaptation (DA) targets at adapting a model trained over the well-labeled source domain to the unlabeled target domain lying in different distributions. Existing DA normally assumes the well-labeled source domain is class-wise balanced, which means the size per source class is relatively similar. However, in real-world applications, labeled samples for some categories in the source domain could be extremely few due to the difficulty of data collection and annotation, which leads to decreasing performance over target domain on those few-shot categories. To perform fair cross-domain adaptation and boost the performance on these minority categories, we develop a novel Generative Few-shot Cross-domain Adaptation (GFCA) algorithm for fair cross-domain classification. Specifically, generative feature augmentation is explored to synthesize effective training data for few-shot source classes, while effective cross-domain alignment aims to adapt knowledge from source to facilitate the target learning. Experimental results on two large cross-domain visual datasets demonstrate the effectiveness of our proposed method on improving both few-shot and overall classification accuracy comparing with the state-of-the-art DA approaches. Tongxin Wang, Zhengming Ding, Wei Shao 0005, Haixu Tang, Kun Huang 0001 |
WACV | 3 |
| 2021 | TPSC: a module detection method based on topology potential and spectral clustering in weighted networks and its application in gene co-expression module discoveryabstractBACKGROUND: Gene co-expression networks are widely studied in the biomedical field, with algorithms such as WGCNA and lmQCM having been developed to detect co-expressed modules. However, these algorithms have limitations such as insufficient granularity and unbalanced module size, which prevent full acquisition of knowledge from data mining. In addition, it is difficult to incorporate prior knowledge in current co-expression module detection algorithms. RESULTS: In this paper, we propose a novel module detection algorithm based on topology potential and spectral clustering algorithm to detect co-expressed modules in gene co-expression networks. By testing on TCGA data, our novel method can provide more complete coverage of genes, more balanced module size and finer granularity than current methods in detecting modules with significant overall survival difference. In addition, the proposed algorithm can identify modules by incorporating prior knowledge. CONCLUSION: In summary, we developed a method to obtain as much as possible information from networks with increased input coverage and the ability to detect more size-balanced and granular modules. In addition, our method can integrate data from different sources. Our proposed method performs better than current methods with complete coverage of input genes and finer granularity. Moreover, this method is designed not only for gene co-expression networks but can also be applied to any general fully connected weighted network. Yusong Liu, Xiufen Ye, Christina Y. Yu, Wei Shao 0005, Weixing Feng, Jie Zhang 0010, Kun Huang 0001 |
BMC Bioinform. | 4 |
| 2021 | Identify Consistent Cross-Modality Imaging Genetic Patterns via Discriminant Sparse Canonical Correlation AnalysisabstractSparse canonical correlation analysis (SCCA) is a bi-multivariate technique used in imaging genetics to identify complex multi-SNP-multi-QT associations. However, the traditional SCCA algorithm has been designed to seek a linear correlation between the SNP genotype and brain imaging phenotype, ignoring the discriminant similarity information between within-class subjects in brain imaging genetics association analysis. In addition, multi-modality brain imaging phenotypes are extracted from different perspectives and imaging markers from the same region consistently showing up in multimodalities may provide more insights for the mechanistic understanding of diseases. In this paper, a novel multi-modality discriminant SCCA algorithm (MD-SCCA) is proposed to overcome these limitations as well as to improve learning results by incorporating valuable discriminant similarity information into the SCCA algorithm. Specifically, we first extract the discriminant similarity information between within-class subjects by the sparse representation. Second, the discriminant similarity information is enforced within SCCA to construct a discriminant SCCA algorithm (D-SCCA). At last, the MD-SCCA algorithm is adopted to fully explore the relationships among different modalities of different subjects. In experiments, both synthetic dataset and real data from the Alzheimer's Disease Neuroimaging Initiative database are used to test the performance of our algorithm. The empirical results have demonstrated that the proposed algorithm not only produces improved cross-validation performances but also identifies consistent cross-modality imaging genetic biomarkers. Meiling Wang 0001, Wei Shao 0005, Xiaoke Hao, Li Shen 0001, Daoqiang Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2021 | Weakly Supervised Deep Ordinal Cox Model for Survival Prediction From Whole-Slide Pathological ImagesabstractWhole-Slide Histopathology Image (WSI) is generally considered the gold standard for cancer diagnosis and prognosis. Given the large inter-operator variation among pathologists, there is an imperative need to develop machine learning models based on WSIs for consistently predicting patient prognosis. The existing WSI-based prediction methods do not utilize the ordinal ranking loss to train the prognosis model, and thus cannot model the strong ordinal information among different patients in an efficient way. Another challenge is that a WSI is of large size (e.g., 100,000-by-100,000 pixels) with heterogeneous patterns but often only annotated with a single WSI-level label, which further complicates the training process. To address these challenges, we consider the ordinal characteristic of the survival process by adding a ranking-based regularization term on the Cox model and propose a weakly supervised deep ordinal Cox model (BDOCOX) for survival prediction from WSIs. Here, we generate amounts of bags from WSIs, and each bag is comprised of the image patches representing the heterogeneous patterns of WSIs, which is assumed to match the WSI-level labels for training the proposed model. The effectiveness of the proposed method is well validated by theoretical analysis as well as the prognosis and patient stratification results on three cancer datasets from The Cancer Genome Atlas (TCGA). Wei Shao 0005, Tongxin Wang, Zhi Han, Jie Zhang 0010, Kun Huang 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2021 | Identify Complex Imaging Genetic Patterns via Fusion Self-Expressive Network AnalysisabstractIn the brain imaging genetic studies, it is a challenging task to estimate the association between quantitative traits (QTs) extracted from neuroimaging data and genetic markers such as single-nucleotide polymorphisms (SNPs). Most of the existing association studies are based on the extensions of sparse canonical correlation analysis (SCCA) for the identification of complex bi-multivariate associations, which can take the specific structure and group information into consideration. However, they often take the original data as input without considering its underlying complex multi-subspace structure, which will deteriorate the performance of the following integrative analysis. Accordingly, in this paper, the self-expressive property is exploited for the reconstruction of the original data before the association analysis, which can well describe the similarity structure. Specifically, we first apply the within-class similarity information to construct self-expressive networks by sparse representation. Then, we use the fusion method to iteratively fuse the self-expressive networks from multi-modality brain phenotypes into one network. Finally, we calculate the imaging genetic association based on the fused self-expressive network. We conduct the experiments on both single-modality and multi-modality phenotype data. Related experimental results validate that our method can not only better estimate the potential association between genetic markers and quantitative traits but also identify consistent multi-modality imaging genetic biomarkers to guide the interpretation of Alzheimer's disease. Meiling Wang 0001, Wei Shao 0005, Xiaoke Hao, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 2 |
| 2020 | Multi-task multi-modal learning for joint diagnosis and prognosis of human cancers
Wei Shao 0005, Tongxin Wang, Liang Sun 0009, Tianhan Dong, Zhi Han, Jie Zhang 0010, Daoqiang Zhang, Kun Huang 0001 |
Medical Image Anal. | 1 |
| 2020 | Querying Representative and Informative Super-Pixels for Filament Segmentation in BioimagesabstractSegmenting bioimage based filaments is a critical step in a wide range of applications, including neuron reconstruction and blood vessel tracing. To achieve an acceptable segmentation performance, most of the existing methods need to annotate amounts of filamentary images in the training stage. Hence, these methods have to face the common challenge that the annotation cost is usually high. To address this problem, we propose an interactive segmentation method to actively select a few super-pixels for annotation, which can alleviate the burden of annotators. Specifically, we first apply a Simple Linear Iterative Clustering (i.e., SLIC) algorithm to segment filamentary images into compact and consistent super-pixels, and then propose a novel batch-mode based active learning method to select the most representative and informative (i.e., BMRI) super-pixels for pixel-level annotation. We then use a bagging strategy to extract several sets of pixels from the annotated super-pixels, and further use them to build different Laplacian Regularized Gaussian Mixture Models (Lap-GMM) for pixel-level segmentation. Finally, we perform the classifier ensemble by combining multiple Lap-GMM models based on a majority voting strategy. We evaluate our method on three public available filamentary image datasets. Experimental results show that, to achieve comparable performance with the existing methods, the proposed algorithm can save 40 percent annotation efforts for experts. Wei Shao 0005, Sheng-Jun Huang, Mingxia Liu 0001, Daoqiang Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | High-Order Feature Learning for Multi-Atlas Based Label Fusion: Application to Brain Segmentation With MRIabstractMulti-atlas based segmentation methods have shown their effectiveness in brain regions-of-interesting (ROIs) segmentation, by propagating labels from multiple atlases to a target image based on the similarity between patches in the target image and multiple atlas images. Most of the existing multiatlas based methods use image intensity features to calculate the similarity between a pair of image patches for label fusion. In particular, using only low-level image intensity features cannot adequately characterize the complex appearance patterns (e.g., the high-order relationship between voxels within a patch) of brain magnetic resonance (MR) images. To address this issue, this paper develops a high-order feature learning framework for multi-atlas based label fusion, where high-order features of image patches are extracted and fused for segmenting ROIs of structural brain MR images. Specifically, an unsupervised feature learning method (i.e., means-covariances restricted Boltzmann machine, mcRBM) is employed to learn high-order features (i.e., mean and covariance features) of patches in brain MR images. Then, a group-fused sparsity dictionary learning method is proposed to jointly calculate the voting weights for label fusion, based on the learned high-order and the original image intensity features. The proposed method is compared with several state-of-the-art label fusion methods on ADNI, NIREP and LONI-LPBA40 datasets. The Dice ratio achieved by our method is 88:30%, 88:83%, 79:54% and 81:02% on left and right hippocampus on the ADNI, NIREP and LONI-LPBA40 datasets, respectively, while the best Dice ratio yielded by the other methods are 86:51%, 87:39%, 78:48% and 79:65% on three datasets, respectively. Liang Sun 0009, Wei Shao 0005, Daoqiang Zhang, Mingxia Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Adaptive Feature Selection Guided Deep Forest for COVID-19 Classification With Chest CTabstractChest computed tomography (CT) becomes an effective tool to assist the diagnosis of coronavirus disease-19 (COVID-19). Due to the outbreak of COVID-19 worldwide, using the computed-aided diagnosis technique for COVID-19 classification based on CT images could largely alleviate the burden of clinicians. In this paper, we propose an Adaptive Feature Selection guided Deep Forest (AFS-DF) for COVID-19 classification based on chest CT images. Specifically, we first extract location-specific features from CT images. Then, in order to capture the high-level representation of these features with the relatively small-scale data, we leverage a deep forest model to learn high-level representation of the features. Moreover, we propose a feature selection method based on the trained deep forest model to reduce the redundancy of features, where the feature selection could be adaptively incorporated with the COVID-19 classification model. We evaluated our proposed AFS-DF on COVID-19 dataset with 1495 patients of COVID-19 and 1027 patients of community acquired pneumonia (CAP). The accuracy (ACC), sensitivity (SEN), specificity (SPE), AUC, precision and F1-score achieved by our method are 91.79%, 93.05%, 89.95%, 96.35%, 93.10% and 93.07%, respectively. Experimental results on the COVID-19 dataset suggest that the proposed AFS-DF achieves superior performance in COVID-19 vs. CAP classification, compared with 4 widely used machine learning methods. Liang Sun 0009, Zhanhao Mo, Fuhua Yan, Liming Xia, Zhongxiang Ding, Bin Song 0002, Wanchun Gao, Wei Shao 0005, Feng Shi 0001, Huan Yuan, Huiting Jiang, Dijia Wu, Ying Wei 0009, Yaozong Gao, He Sui, Daoqiang Zhang, Dinggang Shen |
IEEE J. Biomed. Health Informatics | 9 |
| 2020 | Integrative Analysis of Pathological Images and Multi-Dimensional Genomic Data for Early-Stage Cancer PrognosisabstractThe integrative analysis of histopathological images and genomic data has received increasing attention for studying the complex mechanisms of driving cancers. However, most image-genomic studies have been restricted to combining histopathological images with the single modality of genomic data (e.g., mRNA transcription or genetic mutation), and thus neglect the fact that the molecular architecture of cancer is manifested at multiple levels, including genetic, epigenetic, transcriptional, and post-transcriptional events. To address this issue, we propose a novel ordinal multi-modal feature selection (OMMFS) framework that can simultaneously identify important features from both pathological images and multi-modal genomic data (i.e., mRNA transcription, copy number variation, and DNA methylation data) for the prognosis of cancer patients. Our model is based on a generalized sparse canonical correlation analysis framework, by which we also take advantage of the ordinal survival information among different patients for survival outcome prediction. We evaluate our method on three early-stage cancer datasets derived from The Cancer Genome Atlas (TCGA) project, and the experimental results demonstrated that both the selected image and multi-modal genomic markers are strongly correlated with survival enabling effective stratification of patients with distinct survival than the comparing methods, which is often difficult for early-stage cancer patients. Wei Shao 0005, Kun Huang 0001, Zhi Han, Jun Cheng 0006, Tongxin Wang, Liang Sun 0009, Zixiao Lu, Jie Zhang 0010, Daoqiang Zhang |
IEEE Trans. Medical Imaging | 1 |
| 2020 | Anatomical Attention Guided Deep Networks for ROI Segmentation of Brain MR ImagesabstractBrain region-of-interest (ROI) segmentation based on structural magnetic resonance imaging (MRI) scans is an essential step for many computer-aid medical image analysis applications. Due to low intensity contrast around ROI boundary and large inter-subject variance, it has been remaining a challenging task to effectively segment brain ROIs from structural MR images. Even though several deep learning methods for brain MR image segmentation have been developed, most of them do not incorporate shape priors to take advantage of the regularity of brain structures, thus leading to sub-optimal performance. To address this issue, we propose an anatomical attention guided deep learning framework for brain ROI segmentation of structural MR images, containing two subnetworks. The first one is a segmentation subnetwork, used to simultaneously extract discriminative image representation and segment ROIs for each input MR image. The second one is an anatomical attention subnetwork, designed to capture the anatomical structure information of the brain from a set of labeled atlases. To utilize the anatomical attention knowledge learned from atlases, we develop an anatomical gate architecture to fuse feature maps derived from a set of atlas label maps and those from the to-be-segmented image for brain ROI segmentation. In this way, the anatomical prior learned from atlases can be explicitly employed to guide the segmentation process for performance improvement. Within this framework, we develop two anatomical attention guided segmentation models, denoted as anatomical gated fully convolutional network (AG-FCN) and anatomical gated U-Net (AG-UNet), respectively. Experimental results on both ADNI and LONI-LPBA40 datasets suggest that the proposed AG-FCN and AG-UNet methods achieve superior performance in ROI segmentation of brain MR images, compared with several state-of-the-art methods. Liang Sun 0009, Wei Shao 0005, Daoqiang Zhang, Mingxia Liu 0001 |
IEEE Trans. Medical Imaging | 2 |
| 2019 | Diagnosis-Guided Multi-modal Feature Selection for Prognosis Prediction of Lung Squamous Cell Carcinoma
Wei Shao 0005, Tongxin Wang, Jun Cheng 0006, Zhi Han, Daoqiang Zhang, Kun Huang 0001 |
MICCAI (4) | 1 |
| 2019 | Reliability-based robust multi-atlas label fusion for brain MRI segmentation
Liang Sun 0009, Chen Zu, Wei Shao 0005, Junye Guang, Daoqiang Zhang, Mingxia Liu 0001 |
Artif. Intell. Medicine | 3 |
| 2019 | Discovering network phenotype between genetic risk factors and disease status via diagnosis-aligned multi-modality regression method in Alzheimer's diseaseabstractMOTIVATION: Neuroimaging genetics is an emerging field to identify the associations between genetic variants [e.g. single-nucleotide polymorphisms (SNPs)] and quantitative traits (QTs) such as brain imaging phenotypes. However, most of the current studies focus only on the associations between brain structure imaging and genetic variants, while neglecting the connectivity information between brain regions. In addition, the brain itself is a complex network, and the higher-order interaction may contain useful information for the mechanistic understanding of diseases [i.e. Alzheimer's disease (AD)]. RESULTS: A general framework is proposed to exploit network voxel information and network connectivity information as intermediate traits that bridge genetic risk factors and disease status. Specifically, we first use the sparse representation (SR) model to build hyper-network to express the connectivity features of the brain. The network voxel node features and network connectivity edge features are extracted from the structural magnetic resonance imaging (sMRI) and resting-state functional magnetic resonance imaging (fMRI), respectively. Second, a diagnosis-aligned multi-modality regression method is adopted to fully explore the relationships among modalities of different subjects, which can help further mine the relation between the risk genetics and brain network features. In experiments, all methods are tested on the Alzheimer's Disease Neuroimaging Initiative (ADNI) database. The experimental results not only verify the effectiveness of our proposed framework but also discover some brain regions and connectivity features that are highly related to diseases. AVAILABILITY AND IMPLEMENTATION: The Matlab code is available at http://ibrain.nuaa.edu.cn/2018/list.htm. Meiling Wang 0001, Xiaoke Hao, Jiashuang Huang, Wei Shao 0005, Daoqiang Zhang |
Bioinform. | 4 |
| 2018 | Ordinal Multi-modal Feature Selection for Survival Analysis of Early-Stage Renal Cancer
Wei Shao 0005, Jun Cheng 0006, Liang Sun 0009, Zhi Han, Qianjin Feng 0002, Daoqiang Zhang, Kun Huang 0001 |
MICCAI (2) | 1 |
| 2018 | An Organelle Correlation-Guided Feature Selection Approach for Classifying Multi-Label Subcellular Bio-ImagesabstractNowadays, with the advances in microscopic imaging, accurate classification of bioimage-based protein subcellular location pattern has attracted as much attention as ever. One of the basic challenging problems is how to select the useful feature components among thousands of potential features to describe the images. This is not an easy task especially considering there is a high ratio of multi-location proteins. Existing feature selection methods seldom take the correlation among different cellular compartments into consideration, and thus may miss some features that will be co-important for several subcellular locations. To deal with this problem, we make use of the important structural correlation among different cellular compartments and propose an organelle structural correlation regularized feature selection method CSF (Common-Sets of Features) in this paper. We formulate the multi-label classification problem by adopting a group-sparsity regularizer to select common subsets of relevant features from different cellular compartments. In addition, we also add a cell structural correlation regularized Laplacian term, which utilizes the prior biological structural information to capture the intrinsic dependency among different cellular compartments. The CSF provides a new feature selection strategy for multi-label bio-image subcellular pattern classifications, and the experimental results also show its superiority when comparing with several existing algorithms. Wei Shao 0005, Mingxia Liu 0001, Ying-Ying Xu, Hong-Bin Shen, Daoqiang Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2017 | Deep model-based feature extraction for predicting protein subcellular localizations from bio-images
Wei Shao 0005, Hong-Bin Shen, Daoqiang Zhang |
Frontiers Comput. Sci. | 1 |
| 2016 | Human cell structure-driven model construction for predicting protein subcellular location from biological imagesabstractMOTIVATION: The systematic study of subcellular location pattern is very important for fully characterizing the human proteome. Nowadays, with the great advances in automated microscopic imaging, accurate bioimage-based classification methods to predict protein subcellular locations are highly desired. All existing models were constructed on the independent parallel hypothesis, where the cellular component classes are positioned independently in a multi-class classification engine. The important structural information of cellular compartments is missed. To deal with this problem for developing more accurate models, we proposed a novel cell structure-driven classifier construction approach (SC-PSorter) by employing the prior biological structural information in the learning model. Specifically, the structural relationship among the cellular components is reflected by a new codeword matrix under the error correcting output coding framework. Then, we construct multiple SC-PSorter-based classifiers corresponding to the columns of the error correcting output coding codeword matrix using a multi-kernel support vector machine classification approach. Finally, we perform the classifier ensemble by combining those multiple SC-PSorter-based classifiers via majority voting. RESULTS: We evaluate our method on a collection of 1636 immunohistochemistry images from the Human Protein Atlas database. The experimental results show that our method achieves an overall accuracy of 89.0%, which is 6.4% higher than the state-of-the-art method. AVAILABILITY AND IMPLEMENTATION: The dataset and code can be downloaded from https://github.com/shaoweinuaa/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wei Shao 0005, Mingxia Liu 0001, Daoqiang Zhang |
Bioinform. | 1 |