Muheng Shang

dblp:337/3924 · DBLP profile ↗
← Back
17ranked-venue papers
1as first author
17since 2021 · last 2026
0000-0002-9227-7263ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 1 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Toward Trustworthy Multi-View Representation With Fine-Grained Explainability Embeddings
abstract
Multiomics co-learning is a powerful analytical paradigm that has benefited biomedical studies substantially. However, due to the diverse information and complex relationships of multiomics data, naive multi-view learning methods usually run into spurious correlations and biased signatures irrelevant to the diseases of interest. Therefore, the learned representations and cross-omics associations cannot translate into clinical knowledge for disease prediction. This issue becomes particularly severe when clinical data are limited and scarce. To handle this issue, we propose a novel and powerful scheme, referred to as the Causality-driven Trustworthy Multi-View maPping approach (Cad-TMVP). Specifically, we design a fined multi-directional mapping module to extract co-expression patterns across different modalities and capture fine-grained interpretability factors. We also meticulously design dynamic mechanisms to facilitate adaptive loss-term reweighting and trustworthy integration of multiple modalities. Cad-TMVP enhances downstream tasks by developing a cooperative learning module that simultaneously performs automated diagnosis and result interpretation. Furthermore, we develop an efficient search strategy and support computation to reduce the high computational burden, making our approach practicable. We conduct extensive experiments on different types of multiomics data. The proposed method establishes new state-of-the-art results in various settings while maintaining excellent interpretability. Thus, it sets a potentially newparadigm in trustworthy multi-modal learning and verifies its flexibility and versatility in real biomedical applications.
Jin Zhang 0023, Yan Yang 0011, Muheng Shang, Lei Guo 0002, Daoqiang Zhang, Lei Du 0001
IEEE Trans. Medical Imaging3
2025 Predicting Brain Age Based on Neuroimaging and DNA Methylation via A Multiomics Attention-Based Variational Autoencoder
abstract
Brain age (BA) is recognized as a significant biomarker for health, closely associated with brain aging, and has been proposed to correlate with the progression of neurodegenerative diseases. Previous studies on brain age prediction predominantly relied on single omics data, limiting the integration of cooperative information from multiple omics data. Additionally, existing brain age prediction methods are susceptible to intermediate features since not all these features are related to brain age. To address these limitations, the study introduces a multiomics attention-based VAE method to predict brain age by integrating neuroimaging data and DNA methylation (DNAm) data, aiming to identify biologically meaningful features truly associated with brain aging. Experimental results show that our method predicts brain age with the MAE of 2.08 years, outperforming the state-of-the-art methods under the same experimental conditions. Furthermore, we estimated the brain age gap in patients with Mild Cognitive Impairment (MCI) and Alzheimer's Disease (AD). The results of both MCI and AD patients exhibited a larger brain age gap compared to the Cognitively Normal (CN) group, indicating the model's discriminative capacity across different diagnostic groups. These findings can assist in the early diagnosis of AD and the formulation of early treatment strategies.
Wenrui Cui, Hong Pang, Yan Yang 0011, Muheng Shang, Hongdong Li, Lei Du 0001
BIBM4
2025 A Data-Driven Brain Imaging Quantitative Trait Mediation Effect Identification Method for Alzheimer's Disease
abstract
Alzheimer's disease (AD) is a severe degenerative disease and finding its causal factors of high-risk are very important. Mediation analysis has been a powerful tool to elucidate the underlying mechanisms of trait of interest using genetic variations as instrumental variable (IVs). However, most current mediation analysis method can only work on a limited number of suspected traits and genetic variations, which demands extensive prior knowledge. In this study, we proposed a datadriven learning method to identify potential mediation effects of brain imaging quantitative traits (QTs). Our method couples the three single models of mediation model and treats it as a multiobjective learning problem. The proposed method can work on brain-wide and genome-wide brain imaging and genetic data without providing candidate imaging QTs and genetic IVs. We applied the proposed method to PET imaging QTs of whole brain and genetic IV of whole genome. The results showed that our method successfully identified multiple PET imaging mediators linking genetic variations to AD. Thus, our method can serve as a powerful screening tool for large-scale mediation analysis.
Yan Yang 0011, Muheng Shang, Hongdong Li, Lei Du 0001
BIBM2
2025 Predicting MCI Conversion Status Using Baseline Neuroimaging Scans and Genetics Variations
abstract
Mild cognitive impairment (MCI) is a prodromal stage of Alzheimer's disease (AD), but not all MCI subjects develop into AD finally. Therefore, distinguishing progressive MCI (pMCI) subjects from stable MCI (sMCI) subjects is an area of intense interest, which may provide targeted treatments for at-risk individuals. On this account, building an MCI conversion prediction model at the early stage is particularly important. The neuroimaging data, especially multi-modal ones, has proven to be a great alternative in predicting MCIs' conversion. In addition, genetic variations such as Single Nucleotide Polymorphism (SNP) can also imply the conversion risk of an individual. The neuroimaging data represents the current status, while SNPs convey the inherited risk of an individual. In this paper, we propose a deep representative fusion method that combines multi-modal baseline neuroimaging data and genetic variations. It can predict the progressive status of MCIs over the following two years, three years and four years, respectively. Experimental results from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database demonstrate that the proposed method has better prediction capability than comparison methods. Moreover, findings show that the stability of the default mode network (DMN) and ventral attention network (VAN) are correlated with the MCI conversion and the learned imaging representations are related to minimental state examination (MMSE) scores which are associated with AD progression.
Yan Yang 0011, Muheng Shang, Jin Zhang 0023, Hongdong Li, Lei Du 0001
BIBM2
2025 Mutual-assistance learning for trustworthy biomarker discovery and disease prediction
abstract
Integrating and analyzing multiple omics datasets, such as genomics, environmental influences, and imaging endophenotypes, has yielded an abundance of candidate biomarkers. However, translating such findings into beneficial clinical knowledge for disease prediction remains challenging. This becomes even more challenging when studying interpretable high-order feature interactions such as gene-environment interaction (G$\times $E) to understand the etiology. To fill this gap, we draw on the idea of mutual-assistance (MA) learning and accordingly propose a fresh and powerful scheme, referred to as mutual-assistance causal biomarker discovery and stable disease prediction approach (MA-CBxDP). Specifically, we design an interpretable bi-directional mapping framework, integrated with a causal feature interaction module, to extract co-expression patterns across different modalities and identify trustworthy biomarkers including G$\times $E. A cooperative prediction module is further incorporated to ensure accurate diagnosis and identification of causal effects for pathogenesis. Importantly, biomarker discovery and disease prediction can mutually reinforce each other, helping to provide novel insights into chronic diseases. Furthermore, in light of the large computational burden incurred by the high-dimensional interactions, we devise a rapid strategy and extend it to a more practical but challenging chromosome-wide setting. We conduct extensive experiments on two databases under three tasks, i.e. multimodal correlation, disease diagnosis, and trait prediction. MA-CBxDP establishes new state-of-the-art results in predicting clinical scores and disease status classification, while maintaining exceptional interpretability, verifying its flexibility and versatility in practical applications.
Jin Zhang 0023, Yan Yang 0011, Muheng Shang, Lei Guo 0002, Daoqiang Zhang, Lei Du 0001
Briefings Bioinform.3
2025 Trustworthy causal biomarker discovery: a multiomics brain imaging genetics-based approach
abstract
MOTIVATION: Discovering genetic variations underpinning brain disorders is important to understand their pathogenesis. Indirect associations or spurious causal relationships pose a threat to the reliability of biomarker discovery for brain disorders, potentially misleading or incurring bias in subsequent decision-making. Unfortunately, the stringent selection of reliable biomarker candidates for brain disorders remains a predominantly unexplored challenge. RESULTS: In this article, to fill this gap, we propose a fresh and powerful scheme, referred to as the Causality-aware Genotype intermediate Phenotype Correlation Approach (Ca-GPCA). Specifically, we design a bidirectional association learning framework, integrated with a parallel causal variable decorrelation module and sparse variable regularizer module, to identify trustworthy causal biomarkers. A disease diagnosis module is further incorporated to ensure accurate diagnosis and identification of causal effects for pathogenesis. Additionally, considering the large computational burden incurred by high-dimensional genotype-phenotype covariances, we develop a fast and efficient strategy to reduce the runtime and prompt practical availability and applicability. Extensive experimental results on four simulation data and real neuroimaging genetic data clearly show that Ca-GPCA outperforms state-of-the-art methods with excellent built-in interpretability. This can provide novel and reliable insights into the underlying pathogenic mechanisms of brain disorders. AVAILABILITY AND IMPLEMENTATION: The software is publicly available at https://github.com/ZJ-Techie/Ca-GPCA.
Jin Zhang 0023, Yan Yang 0011, Muheng Shang, Lei Guo 0002, Daoqiang Zhang, Lei Du 0001
Bioinform.3
2025 Modeling multi-stage disease progression and identifying genetic risk factors via a novel collaborative learning method
abstract
MOTIVATION: Alzheimer's disease (AD) typically progresses gradually for ages rather than suddenly. Thus, staging AD progression in different phases could aid in accurate diagnosis and treatment. In addition, identifying genetic variations that influence AD is critical to understanding the pathogenesis. However, staging the disease progression and identifying genetic variations is usually handled separately. RESULTS: To address this limitation, we propose a novel sparse multi-stage multi-task mixed-effects collaborative longitudinal regression method (MSColoR). Our method jointly models long disease progression as a multi-stage procedure and identifies genetic risk factors underpinning this complex trajectory. Specifically, MSColoR models multi-stage disease progression using longitudinal neuroimaging-derived phenotypes and associates the fitted disease trajectories with genetic variations at each stage. Furthermore, we collaboratively leverage summary statistics from large genome-wide association studies to improve the powers. Finally, an efficient optimization algorithm is introduced to solve MSColoR. We evaluate our method using both synthetic and real longitudinal neuroimaging and genetic data. Both results demonstrate that MSColoR can reduce modeling errors while identifying more accurate and significant genetic variations compared to other longitudinal methods. Consequently, MSColoR holds great potential as a computational technique for longitudinal brain imaging genetics and AD studies. AVAILABILITY AND IMPLEMENTATION: The code is publicly available at https://github.com/dulei323/MSColoR.
Duo Xi, Minjianan Zhang, Muheng Shang, Lei Du 0001, Junwei Han 0001
Bioinform.3
2024 Identification of disease-related genetic variants and imaging factors leveraging summary statistics
abstract
Brain imaging genetics offers insights into the genetic basis of brain structure and function by exploring the relationships between genetic variations and neuroimaging features, with canonical correlation association learning as a vital and effective tool. However, imaging large cohorts affected by specific brain diseases entails significant costs. To tackle this challenge, we introduced a novel bi-multivariate sparse canonical correlation association method based on summary statistics from large GWAS (S-SCCA). S-SCCA leverages effect sizes obtained from these datasets to identify genetic variants associated with complex traits, including those influenced by pleiotropy, while simultaneously identifying imaging factors related to the disease under study. Moreover, we have implemented a rapid optimization strategy to circumvent computational burdens while identifying disease-associated risk factors within genetic variations across the entire chromosome. We assessed S-SCCA against conventional SCCA using a neuroimaging genetic dataset from the Alzheimer’s Disease Neuroimaging Initiative. Results showed that S-SCCA demonstrated comparable or superior modeling performance and feature selection capabilities. Furthermore, we applied S-SCCA to two summary statistics datasets from two large GWAS, where original imaging and genetic data were inaccessible. S-SCCA replicated the genetic loci identified by GWAS and additional meaningful variants. Additionally, it revealed bi-multivariate relationships between imaging QTs and SNPs, indicating its powerful modeling capability. These findings highlight the promise of S-SCCA as a practical bi-multivariate learning technique in brain imaging genetics, circumventing the need for sensitive individual-level imaging and genetic data, thereby enhancing its potential for broader applicability and accessibility in biomedical studies.
Duo Xi, Dingnan Cui, Minjianan Zhang, Jin Zhang 0023, Muheng Shang, Lei Guo 0002, Lei Du 0001, Junwei Han 0001
BIBM5
2024 Disease Progression Prediction Incorporating Genotype-Environment Interactions: A Longitudinal Neurodegenerative Disorder Study
Jin Zhang 0023, Muheng Shang, Yan Yang 0011, Lei Guo 0002, Junwei Han 0001, Lei Du 0001
MICCAI (3)2
2024 A Multi-Task Deep Feature Selection Method for Brain Imaging Genetics
abstract
Using brain imaging quantitative traits (QTs) for identifying genetic risk factors is an important research topic in brain imaging genetics. Many efforts have been made for this task via building linear models between imaging QTs and genetic factors such as single nucleotide polymorphisms (SNPs). To the best of our knowledge, linear models could not fully uncover the complicated relationship due to the loci's elusive and diverse influences on imaging QTs. In this paper, we propose a novel multi-task deep feature selection (MTDFS) method for brain imaging genetics. MTDFS first builds a multi-task deep neural network to model the complicated associations between imaging QTs and SNPs. And then designs a multi-task one-to-one layer and imposes a combined penalty to identify SNPs that make significant contributions. MTDFS can not only extract the nonlinear relationship but also arms the deep neural network with feature selection. We compared MTDFS to multi-task linear regression (MTLR) and single-task DFS (DFS) methods on the real neuroimaging genetic data. The experimental results showed that MTDFS performed better than MTLR and DFS on the QT-SNP relationship identification and feature selection. Thus, MTDFS is powerful for identifying risk loci and could be a great supplement to brain imaging genetics.
Shu Zhang 0006, Muheng Shang, Lei Guo 0002, Junwei Han 0001, Lei Du 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2024 Identification of Genetic Risk Factors Based on Disease Progression Derived From Longitudinal Brain Imaging Phenotypes
abstract
Neurodegenerative disorders usually happen stage-by-stage rather than overnight. Thus, cross-sectional brain imaging genetic methods could be insufficient to identify genetic risk factors. Repeatedly collecting imaging data over time appears to solve the problem. But most existing imaging genetic methods only use longitudinal imaging phenotypes straightforwardly, ignoring the disease progression trajectory which might be a more stable disease signature. In this paper, we propose a novel sparse multi-task mixed-effects longitudinal imaging genetic method (SMMLING). In our model, disease progression fitting and genetic risk factors identification are conducted jointly. Specifically, SMMLING models the disease progression using longitudinal imaging phenotypes, and then associates fitted disease progression with genetic variations. The baseline status and changing rate, i.e., the intercept and slope, of the progression trajectory thus shoulder the responsibility to discover loci of interest, which would have superior and stable performance. To facilitate the interpretation and stability, we employ$\ell _{{2},{1}}$-norm and the fused group lasso (FGL) penalty to identify loci at both the individual level and group level. SMMLING can be solved by an efficient optimization algorithm which is guaranteed to converge to the global optimum. We evaluate SMMLING on synthetic data and real longitudinal neuroimaging genetic data. Both results show that, compared to existing longitudinal methods, SMMLING can not only decrease the modeling error but also identify more accurate and relevant genetic factors. Most risk loci reported by SMMLING are missed by comparison methods, implicating its superiority in genetic risk factors identification. Consequently, SMMLING could be a promising computational method for longitudinal imaging genetics.
Lei Du 0001, Ying Zhao 0015, Muheng Shang, Jin Zhang 0023, Junwei Han 0001
IEEE Trans. Medical Imaging4
2023 Identifying Disease-related Brain Imaging Quantitative Traits and Related Genetic Variations via A Bidirectional Association Learning Method
abstract
Discovering critical genetic biomarkers of Alzheimer’s disease (AD) by detecting the complex associations between genotypes (i.e. single nucleotide polymorphism, SNP) and phenotypes (i.e. quantitative trait, QT) is a long-standing and beneficial task for the diagnosis and the follow-up treatment of patients. The function of genes and their relationships with phenotypes are extremely complex. A lot of imaging genetic methods have been designed to uncover the association between brain imaging QTs and SNPs. However, most of them are focused on the effect of a single SNP, which may have limited ability due to the oligogenic or polygenic characteristic of AD. In this paper, we propose a deep reconstruction bidirectional association with feature selection (DRBA-FS) method to explore the multi-SNPmulti-QT associations. In this method, the co-effect of multiple AD-related genetic variations is identified and aggregated, and their high-level genetic associations to brain imaging QTs are jointly modeled. Experiment results on real neuroimaging genetic data from Alzheimer’s Disease Neuroimaging Initiative (ADNI) show that the identified biomarkers are all related to AD. Interestingly, our method can learn the joint effect of multiple AD-related genetic variations across the genome, and thus has significant potential in understanding the genetic mechanism of AD.
Muheng Shang, Yan Yang 0011, Minjianan Zhang, Jin Zhang 0023, Duo Xi, Lei Guo 0002, Lei Du 0001
BIBM1
2023 FMRI-Guided Time-Symmetric Joint Model for Visual Attention Prediction
abstract
Visual attention prediction is linked to brain activity, cognition, and behavior. Despite the availability of brain activity features, previous studies have not fully utilized them, resulting in saliency maps predicted by models primarily based on image features that do not accurately reflect visual attention in the human brain. This inspires us to use functional Magnetic Resonance Imaging (fMRI) signals as a "brain observer" to supervise the training of developing models that integrate top-down image attention-dependent cues and supervise information from saliency maps generated from gaze movement patterns under natural stimuli. Hence, this paper presents an FMRI-Guided Time-Symmetric Joint Model to predict saliency maps from movie clips, which captures the dynamic aspects of human brain cognition and attention, enabling the combination of image features with brain features. Furthermore, we generalize the model to the MS-COCO challenge, evaluating its performance on non-movie data. Our model outperforms other brain-feature-free methods in focusing on visual attention regions of humans in both movie and non-movie datasets. Additionally, incorporating brain features improves model performance, indicating their ability to bridge the semantic gap between human cognition and visual images, allowing for more accurate capture of visual attention regions.
Yaonai Wei, Chong Ma 0004, Tianyang Zhong, Lei Du 0001, Songyao Zhang, Tianming Liu 0001, Muheng Shang, Junwei Han 0001
BIBM11
2023 Chat2Brain: A Method for Mapping Open-Ended Semantic Queries to Brain Activation Maps
abstract
Over decades, neuroscience has accumulated a wealth of research results in the text modality that can be used to explore cognitive processes. Meta-analysis is a typical method that successfully establishes a link from text queries to brain activation maps using these research results, but it still relies on an ideal query environment. In practical applications, text queries used for meta-analyses may encounter issues such as semantic redundancy and ambiguity, resulting in an inaccurate mapping to brain images. On the other hand, large language models (LLMs) like ChatGPT have shown great potential in tasks such as context understanding and reasoning, displaying a high degree of consistency with human natural language. Hence, LLMs could improve the connection between text modality and neuroscience, resolving existing challenges of meta-analyses. In this study, we propose a method called Chat2Brain that combines LLMs to basic text-2-image model, known as Text2Brain, to map open-ended semantic queries to brain activation maps in data-scarce and complex query environments. By utilizing the understanding and reasoning capabilities of LLMs, the performance of the mapping model is optimized by transferring text queries to semantic queries. We demonstrate that Chat2Brain can synthesize anatomically plausible neural activation patterns for more complex tasks of text queries.
Yaonai Wei, Tianyang Zhong, Songyao Zhang, Xiao Li 0024, Lin Zhao 0004, Zhengliang Liu, Muheng Shang, Tianming Liu 0001, Chong Ma 0004, Lei Du 0001, Junwei Han 0001
BIBM8
2023 Identifying Main and Epistasis Effects of Genetic Variations on Neuroimaging Phenotypes Using Effective Feature Interaction Learning
abstract
Brain imaging genetics investigates the complex relationships between genetic variations and brain imaging quantitative traits (QTs). However, existing approaches primarily focus on the main effects of genetic variations, potentially neglecting the crucial role of epistasis that explains the missing heritability of brain disorders. Unfortunately, this issue has been largely unexploited. In this paper, to fill this gap, we present Multi-Task feature interaction-aware Sparse Canonical Correlation Analysis (MTfiSCCA) to identify disease-related main effect and epistasis of risk genetic factors on multimodal neuroimaging phenotypes simultaneously. To ensure stability and interpretation, we use innovative sparsity-inducing penalties to identify biomarkers that make significant contributions. Additionally, we develop an efficient optimization algorithm to solve the proposed method, which converges to a local optimum. Experimental results on the Alzheimer’s disease neuroimaging initiative (ADNI) dataset show that our MTfiSCCA method achieves higher canonical correlation coefficients (CCC) and better feature selection subsets such as disease-related biomarkers compared to the state-of-the-art methods. Furthermore, MTfiSCCA reveals interpretable epistasis among genetic variations implicated in AD, offering novel insights into the underlying pathogenic mechanisms of brain disorders such as Alzheimer’s disease (AD).
Jin Zhang 0023, Muheng Shang, Duo Xi, Minjianan Zhang, Lei Guo 0002, Lei Du 0001
BIBM3
2023 Identification of Disease-Sensitive Brain Imaging Phenotypes and Genetic Factors Using GWAS Summary Statistics
Duo Xi, Dingnan Cui, Jin Zhang 0023, Muheng Shang, Minjianan Zhang, Lei Guo 0002, Junwei Han 0001, Lei Du 0001
MICCAI (5)4
2022 A Sparse Multi-task Contrastive and Discriminative Learning Method with Feature Selection for Brain Imaging Genetics
abstract
Alzheimer’s disease (AD) is a very complex neurodegenerative disease. Generally, different diagnostic groups could exhibit discriminative and specific patterns, including the single nucleotide polymorphisms (SNPs), brain imaging quantitative traits (QTs), as well as their associations, which may facilitate the comprehensive understanding of AD. However, most existing methods cannot guarantee to identify discriminative or class-specific biomarkers or both of them. To overcome this shortcoming, we propose a sparse multi-task contrastive and discriminative learning approach (MTCDA) to jointly learn the discriminative and specific patterns for multiple diagnostic groups. MTCDA can identify the class-relevant and discriminative SNP-QTs associations, and relevant SNPs, imaging QTs underpinning this relationship. We introduce an efficient algorithm to solve the proposed method which converges to a local optimum. The experimental results on Alzheimer’s Disease Neuroimaging Initiative (ADNI) show that MTCDA can obtain higher canonical correlation coefficients, classification accuracy and better feature selection results than state-of-the-art methods, which demonstrates the potential of our method for multi-class brain imaging genetics.
Jin Zhang 0023, Muheng Shang, Minjianan Zhang, Duo Xi, Lei Guo 0002, Junwei Han 0001, Lei Du 0001
BIBM2