EDBT 2026 Demo / reviewers in the wild / expert
Lei Du 0001
dblp:18/1520-1
· DBLP profile ↗
40ranked-venue papers
17as first author
25since 2021 · last 2026
0000-0002-6698-9814ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 40 · 17 first-author · 25 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Toward Trustworthy Multi-View Representation With Fine-Grained Explainability EmbeddingsabstractMultiomics co-learning is a powerful analytical paradigm that has benefited biomedical studies substantially. However, due to the diverse information and complex relationships of multiomics data, naive multi-view learning methods usually run into spurious correlations and biased signatures irrelevant to the diseases of interest. Therefore, the learned representations and cross-omics associations cannot translate into clinical knowledge for disease prediction. This issue becomes particularly severe when clinical data are limited and scarce. To handle this issue, we propose a novel and powerful scheme, referred to as the Causality-driven Trustworthy Multi-View maPping approach (Cad-TMVP). Specifically, we design a fined multi-directional mapping module to extract co-expression patterns across different modalities and capture fine-grained interpretability factors. We also meticulously design dynamic mechanisms to facilitate adaptive loss-term reweighting and trustworthy integration of multiple modalities. Cad-TMVP enhances downstream tasks by developing a cooperative learning module that simultaneously performs automated diagnosis and result interpretation. Furthermore, we develop an efficient search strategy and support computation to reduce the high computational burden, making our approach practicable. We conduct extensive experiments on different types of multiomics data. The proposed method establishes new state-of-the-art results in various settings while maintaining excellent interpretability. Thus, it sets a potentially newparadigm in trustworthy multi-modal learning and verifies its flexibility and versatility in real biomedical applications. Jin Zhang 0023, Yan Yang 0011, Muheng Shang, Lei Guo 0002, Daoqiang Zhang, Lei Du 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Predicting Brain Age Based on Neuroimaging and DNA Methylation via A Multiomics Attention-Based Variational AutoencoderabstractBrain age (BA) is recognized as a significant biomarker for health, closely associated with brain aging, and has been proposed to correlate with the progression of neurodegenerative diseases. Previous studies on brain age prediction predominantly relied on single omics data, limiting the integration of cooperative information from multiple omics data. Additionally, existing brain age prediction methods are susceptible to intermediate features since not all these features are related to brain age. To address these limitations, the study introduces a multiomics attention-based VAE method to predict brain age by integrating neuroimaging data and DNA methylation (DNAm) data, aiming to identify biologically meaningful features truly associated with brain aging. Experimental results show that our method predicts brain age with the MAE of 2.08 years, outperforming the state-of-the-art methods under the same experimental conditions. Furthermore, we estimated the brain age gap in patients with Mild Cognitive Impairment (MCI) and Alzheimer's Disease (AD). The results of both MCI and AD patients exhibited a larger brain age gap compared to the Cognitively Normal (CN) group, indicating the model's discriminative capacity across different diagnostic groups. These findings can assist in the early diagnosis of AD and the formulation of early treatment strategies. Wenrui Cui, Hong Pang, Yan Yang 0011, Muheng Shang, Hongdong Li, Lei Du 0001 |
BIBM | 6 |
| 2025 | A Data-Driven Brain Imaging Quantitative Trait Mediation Effect Identification Method for Alzheimer's DiseaseabstractAlzheimer's disease (AD) is a severe degenerative disease and finding its causal factors of high-risk are very important. Mediation analysis has been a powerful tool to elucidate the underlying mechanisms of trait of interest using genetic variations as instrumental variable (IVs). However, most current mediation analysis method can only work on a limited number of suspected traits and genetic variations, which demands extensive prior knowledge. In this study, we proposed a datadriven learning method to identify potential mediation effects of brain imaging quantitative traits (QTs). Our method couples the three single models of mediation model and treats it as a multiobjective learning problem. The proposed method can work on brain-wide and genome-wide brain imaging and genetic data without providing candidate imaging QTs and genetic IVs. We applied the proposed method to PET imaging QTs of whole brain and genetic IV of whole genome. The results showed that our method successfully identified multiple PET imaging mediators linking genetic variations to AD. Thus, our method can serve as a powerful screening tool for large-scale mediation analysis. Yan Yang 0011, Muheng Shang, Hongdong Li, Lei Du 0001 |
BIBM | 5 |
| 2025 | Predicting MCI Conversion Status Using Baseline Neuroimaging Scans and Genetics VariationsabstractMild cognitive impairment (MCI) is a prodromal stage of Alzheimer's disease (AD), but not all MCI subjects develop into AD finally. Therefore, distinguishing progressive MCI (pMCI) subjects from stable MCI (sMCI) subjects is an area of intense interest, which may provide targeted treatments for at-risk individuals. On this account, building an MCI conversion prediction model at the early stage is particularly important. The neuroimaging data, especially multi-modal ones, has proven to be a great alternative in predicting MCIs' conversion. In addition, genetic variations such as Single Nucleotide Polymorphism (SNP) can also imply the conversion risk of an individual. The neuroimaging data represents the current status, while SNPs convey the inherited risk of an individual. In this paper, we propose a deep representative fusion method that combines multi-modal baseline neuroimaging data and genetic variations. It can predict the progressive status of MCIs over the following two years, three years and four years, respectively. Experimental results from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database demonstrate that the proposed method has better prediction capability than comparison methods. Moreover, findings show that the stability of the default mode network (DMN) and ventral attention network (VAN) are correlated with the MCI conversion and the learned imaging representations are related to minimental state examination (MMSE) scores which are associated with AD progression. Yan Yang 0011, Muheng Shang, Jin Zhang 0023, Hongdong Li, Lei Du 0001 |
BIBM | 7 |
| 2025 | Mutual-assistance learning for trustworthy biomarker discovery and disease predictionabstractIntegrating and analyzing multiple omics datasets, such as genomics, environmental influences, and imaging endophenotypes, has yielded an abundance of candidate biomarkers. However, translating such findings into beneficial clinical knowledge for disease prediction remains challenging. This becomes even more challenging when studying interpretable high-order feature interactions such as gene-environment interaction (G$\times $E) to understand the etiology. To fill this gap, we draw on the idea of mutual-assistance (MA) learning and accordingly propose a fresh and powerful scheme, referred to as mutual-assistance causal biomarker discovery and stable disease prediction approach (MA-CBxDP). Specifically, we design an interpretable bi-directional mapping framework, integrated with a causal feature interaction module, to extract co-expression patterns across different modalities and identify trustworthy biomarkers including G$\times $E. A cooperative prediction module is further incorporated to ensure accurate diagnosis and identification of causal effects for pathogenesis. Importantly, biomarker discovery and disease prediction can mutually reinforce each other, helping to provide novel insights into chronic diseases. Furthermore, in light of the large computational burden incurred by the high-dimensional interactions, we devise a rapid strategy and extend it to a more practical but challenging chromosome-wide setting. We conduct extensive experiments on two databases under three tasks, i.e. multimodal correlation, disease diagnosis, and trait prediction. MA-CBxDP establishes new state-of-the-art results in predicting clinical scores and disease status classification, while maintaining exceptional interpretability, verifying its flexibility and versatility in practical applications. Jin Zhang 0023, Yan Yang 0011, Muheng Shang, Lei Guo 0002, Daoqiang Zhang, Lei Du 0001 |
Briefings Bioinform. | 6 |
| 2025 | Trustworthy causal biomarker discovery: a multiomics brain imaging genetics-based approachabstractMOTIVATION: Discovering genetic variations underpinning brain disorders is important to understand their pathogenesis. Indirect associations or spurious causal relationships pose a threat to the reliability of biomarker discovery for brain disorders, potentially misleading or incurring bias in subsequent decision-making. Unfortunately, the stringent selection of reliable biomarker candidates for brain disorders remains a predominantly unexplored challenge. RESULTS: In this article, to fill this gap, we propose a fresh and powerful scheme, referred to as the Causality-aware Genotype intermediate Phenotype Correlation Approach (Ca-GPCA). Specifically, we design a bidirectional association learning framework, integrated with a parallel causal variable decorrelation module and sparse variable regularizer module, to identify trustworthy causal biomarkers. A disease diagnosis module is further incorporated to ensure accurate diagnosis and identification of causal effects for pathogenesis. Additionally, considering the large computational burden incurred by high-dimensional genotype-phenotype covariances, we develop a fast and efficient strategy to reduce the runtime and prompt practical availability and applicability. Extensive experimental results on four simulation data and real neuroimaging genetic data clearly show that Ca-GPCA outperforms state-of-the-art methods with excellent built-in interpretability. This can provide novel and reliable insights into the underlying pathogenic mechanisms of brain disorders. AVAILABILITY AND IMPLEMENTATION: The software is publicly available at https://github.com/ZJ-Techie/Ca-GPCA. Jin Zhang 0023, Yan Yang 0011, Muheng Shang, Lei Guo 0002, Daoqiang Zhang, Lei Du 0001 |
Bioinform. | 6 |
| 2025 | Modeling multi-stage disease progression and identifying genetic risk factors via a novel collaborative learning methodabstractMOTIVATION: Alzheimer's disease (AD) typically progresses gradually for ages rather than suddenly. Thus, staging AD progression in different phases could aid in accurate diagnosis and treatment. In addition, identifying genetic variations that influence AD is critical to understanding the pathogenesis. However, staging the disease progression and identifying genetic variations is usually handled separately. RESULTS: To address this limitation, we propose a novel sparse multi-stage multi-task mixed-effects collaborative longitudinal regression method (MSColoR). Our method jointly models long disease progression as a multi-stage procedure and identifies genetic risk factors underpinning this complex trajectory. Specifically, MSColoR models multi-stage disease progression using longitudinal neuroimaging-derived phenotypes and associates the fitted disease trajectories with genetic variations at each stage. Furthermore, we collaboratively leverage summary statistics from large genome-wide association studies to improve the powers. Finally, an efficient optimization algorithm is introduced to solve MSColoR. We evaluate our method using both synthetic and real longitudinal neuroimaging and genetic data. Both results demonstrate that MSColoR can reduce modeling errors while identifying more accurate and significant genetic variations compared to other longitudinal methods. Consequently, MSColoR holds great potential as a computational technique for longitudinal brain imaging genetics and AD studies. AVAILABILITY AND IMPLEMENTATION: The code is publicly available at https://github.com/dulei323/MSColoR. Duo Xi, Minjianan Zhang, Muheng Shang, Lei Du 0001, Junwei Han 0001 |
Bioinform. | 4 |
| 2024 | Identification of disease-related genetic variants and imaging factors leveraging summary statisticsabstractBrain imaging genetics offers insights into the genetic basis of brain structure and function by exploring the relationships between genetic variations and neuroimaging features, with canonical correlation association learning as a vital and effective tool. However, imaging large cohorts affected by specific brain diseases entails significant costs. To tackle this challenge, we introduced a novel bi-multivariate sparse canonical correlation association method based on summary statistics from large GWAS (S-SCCA). S-SCCA leverages effect sizes obtained from these datasets to identify genetic variants associated with complex traits, including those influenced by pleiotropy, while simultaneously identifying imaging factors related to the disease under study. Moreover, we have implemented a rapid optimization strategy to circumvent computational burdens while identifying disease-associated risk factors within genetic variations across the entire chromosome. We assessed S-SCCA against conventional SCCA using a neuroimaging genetic dataset from the Alzheimer’s Disease Neuroimaging Initiative. Results showed that S-SCCA demonstrated comparable or superior modeling performance and feature selection capabilities. Furthermore, we applied S-SCCA to two summary statistics datasets from two large GWAS, where original imaging and genetic data were inaccessible. S-SCCA replicated the genetic loci identified by GWAS and additional meaningful variants. Additionally, it revealed bi-multivariate relationships between imaging QTs and SNPs, indicating its powerful modeling capability. These findings highlight the promise of S-SCCA as a practical bi-multivariate learning technique in brain imaging genetics, circumventing the need for sensitive individual-level imaging and genetic data, thereby enhancing its potential for broader applicability and accessibility in biomedical studies. Duo Xi, Dingnan Cui, Minjianan Zhang, Jin Zhang 0023, Muheng Shang, Lei Guo 0002, Lei Du 0001, Junwei Han 0001 |
BIBM | 7 |
| 2024 | Disentangling Disease-sensitive Multimodal Neuroimaging Phenotypes and Related Genetic Factors: A Multimodal Study of ADNI CohortabstractUnderstanding neurological manifestations and their genetic architectures are important for exploring the etiology and pathology of brain disorders. Multimodal neuroimaging data carry complementary information and are known to exhibit shared and specific characteristics from different perspectives. Hence, exploring modality-shared and modality-specific imaging features as well as their genetic underpinnings is a challenging but beneficial task. Unfortunately, this issue has been largely unexploited. In this paper, to fill this gap, we propose a fresh and straightforward insight, referred as Multimodality-Disentangled Phenotype-Genotype Correlation approach (MDPGC). Specifically, we design a unified framework for exploring the multimodality-disentangled characteristics of image-based phenotypes, and further detect genetic variants associated with the disorder using modality-shared and modality-specific biomarkers as intermediate phenotypes. Extensive experimental results on Alzheimer’s Disease Neuroimaging Initiative (ADNI) dataset reveal that our method attains superior correlation coefficients compared to state-of-the-art methods, and at the same time provided excellent interpretability. In addition, the subsequent analysis demonstrates that MDPGC successfully identifies different types of characteristics of imaging phenotypes and reveals relevant genetic variations. These findings not only contribute to AD diagnosis but also help better understand the pathological and pathogenic mechanisms of brain disorders. Jin Zhang 0023, Minjianan Zhang, Lei Guo 0002, Daoqiang Zhang, Lei Du 0001 |
BIBM | 5 |
| 2024 | Disease Progression Prediction Incorporating Genotype-Environment Interactions: A Longitudinal Neurodegenerative Disorder Study
Jin Zhang 0023, Muheng Shang, Yan Yang 0011, Lei Guo 0002, Junwei Han 0001, Lei Du 0001 |
MICCAI (3) | 6 |
| 2024 | Modeling genotype-protein interaction and correlation for Alzheimer's disease: a multi-omics imaging genetics studyabstractIntegrating and analyzing multiple omics data sets, including genomics, proteomics and radiomics, can significantly advance researchers' comprehensive understanding of Alzheimer's disease (AD). However, current methodologies primarily focus on the main effects of genetic variation and protein, overlooking non-additive effects such as genotype-protein interaction (GPI) and correlation patterns in brain imaging genetics studies. Importantly, these non-additive effects could contribute to intermediate imaging phenotypes, finally leading to disease occurrence. In general, the interaction between genetic variations and proteins, and their correlations are two distinct biological effects, and thus disentangling the two effects for heritable imaging phenotypes is of great interest and need. Unfortunately, this issue has been largely unexploited. In this paper, to fill this gap, we propose $\textbf{M}$ulti-$\textbf{T}$ask $\textbf{G}$enotype-$\textbf{P}$rotein $\textbf{I}$nteraction and $\textbf{C}$orrelation disentangling method ($\textbf{MT-GPIC}$) to identify GPI and extract correlation patterns between them. To ensure stability and interpretability, we use novel and off-the-shelf penalties to identify meaningful genetic risk factors, as well as exploit the interconnectedness of different brain regions. Additionally, since computing GPI poses a high computational burden, we develop a fast optimization strategy for solving MT-GPIC, which is guaranteed to converge. Experimental results on the Alzheimer's Disease Neuroimaging Initiative data set show that MT-GPIC achieves higher correlation coefficients and classification accuracy than state-of-the-art methods. Moreover, our approach could effectively identify interpretable phenotype-related GPI and correlation patterns in high-dimensional omics data sets. These findings not only enhance the diagnostic accuracy but also contribute valuable insights into the underlying pathogenic mechanisms of AD. Jin Zhang 0023, Zikang Ma, Yan Yang 0011, Lei Guo 0002, Lei Du 0001 |
Briefings Bioinform. | 5 |
| 2024 | A Multi-Task Deep Feature Selection Method for Brain Imaging GeneticsabstractUsing brain imaging quantitative traits (QTs) for identifying genetic risk factors is an important research topic in brain imaging genetics. Many efforts have been made for this task via building linear models between imaging QTs and genetic factors such as single nucleotide polymorphisms (SNPs). To the best of our knowledge, linear models could not fully uncover the complicated relationship due to the loci's elusive and diverse influences on imaging QTs. In this paper, we propose a novel multi-task deep feature selection (MTDFS) method for brain imaging genetics. MTDFS first builds a multi-task deep neural network to model the complicated associations between imaging QTs and SNPs. And then designs a multi-task one-to-one layer and imposes a combined penalty to identify SNPs that make significant contributions. MTDFS can not only extract the nonlinear relationship but also arms the deep neural network with feature selection. We compared MTDFS to multi-task linear regression (MTLR) and single-task DFS (DFS) methods on the real neuroimaging genetic data. The experimental results showed that MTDFS performed better than MTLR and DFS on the QT-SNP relationship identification and feature selection. Thus, MTDFS is powerful for identifying risk loci and could be a great supplement to brain imaging genetics. Shu Zhang 0006, Muheng Shang, Lei Guo 0002, Junwei Han 0001, Lei Du 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2024 | Identification of Genetic Risk Factors Based on Disease Progression Derived From Longitudinal Brain Imaging PhenotypesabstractNeurodegenerative disorders usually happen stage-by-stage rather than overnight. Thus, cross-sectional brain imaging genetic methods could be insufficient to identify genetic risk factors. Repeatedly collecting imaging data over time appears to solve the problem. But most existing imaging genetic methods only use longitudinal imaging phenotypes straightforwardly, ignoring the disease progression trajectory which might be a more stable disease signature. In this paper, we propose a novel sparse multi-task mixed-effects longitudinal imaging genetic method (SMMLING). In our model, disease progression fitting and genetic risk factors identification are conducted jointly. Specifically, SMMLING models the disease progression using longitudinal imaging phenotypes, and then associates fitted disease progression with genetic variations. The baseline status and changing rate, i.e., the intercept and slope, of the progression trajectory thus shoulder the responsibility to discover loci of interest, which would have superior and stable performance. To facilitate the interpretation and stability, we employ$\ell _{{2},{1}}$-norm and the fused group lasso (FGL) penalty to identify loci at both the individual level and group level. SMMLING can be solved by an efficient optimization algorithm which is guaranteed to converge to the global optimum. We evaluate SMMLING on synthetic data and real longitudinal neuroimaging genetic data. Both results show that, compared to existing longitudinal methods, SMMLING can not only decrease the modeling error but also identify more accurate and relevant genetic factors. Most risk loci reported by SMMLING are missed by comparison methods, implicating its superiority in genetic risk factors identification. Consequently, SMMLING could be a promising computational method for longitudinal imaging genetics. Lei Du 0001, Ying Zhao 0015, Muheng Shang, Jin Zhang 0023, Junwei Han 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2023 | Identifying Disease-related Brain Imaging Quantitative Traits and Related Genetic Variations via A Bidirectional Association Learning MethodabstractDiscovering critical genetic biomarkers of Alzheimer’s disease (AD) by detecting the complex associations between genotypes (i.e. single nucleotide polymorphism, SNP) and phenotypes (i.e. quantitative trait, QT) is a long-standing and beneficial task for the diagnosis and the follow-up treatment of patients. The function of genes and their relationships with phenotypes are extremely complex. A lot of imaging genetic methods have been designed to uncover the association between brain imaging QTs and SNPs. However, most of them are focused on the effect of a single SNP, which may have limited ability due to the oligogenic or polygenic characteristic of AD. In this paper, we propose a deep reconstruction bidirectional association with feature selection (DRBA-FS) method to explore the multi-SNPmulti-QT associations. In this method, the co-effect of multiple AD-related genetic variations is identified and aggregated, and their high-level genetic associations to brain imaging QTs are jointly modeled. Experiment results on real neuroimaging genetic data from Alzheimer’s Disease Neuroimaging Initiative (ADNI) show that the identified biomarkers are all related to AD. Interestingly, our method can learn the joint effect of multiple AD-related genetic variations across the genome, and thus has significant potential in understanding the genetic mechanism of AD. Muheng Shang, Yan Yang 0011, Minjianan Zhang, Jin Zhang 0023, Duo Xi, Lei Guo 0002, Lei Du 0001 |
BIBM | 7 |
| 2023 | FMRI-Guided Time-Symmetric Joint Model for Visual Attention PredictionabstractVisual attention prediction is linked to brain activity, cognition, and behavior. Despite the availability of brain activity features, previous studies have not fully utilized them, resulting in saliency maps predicted by models primarily based on image features that do not accurately reflect visual attention in the human brain. This inspires us to use functional Magnetic Resonance Imaging (fMRI) signals as a "brain observer" to supervise the training of developing models that integrate top-down image attention-dependent cues and supervise information from saliency maps generated from gaze movement patterns under natural stimuli. Hence, this paper presents an FMRI-Guided Time-Symmetric Joint Model to predict saliency maps from movie clips, which captures the dynamic aspects of human brain cognition and attention, enabling the combination of image features with brain features. Furthermore, we generalize the model to the MS-COCO challenge, evaluating its performance on non-movie data. Our model outperforms other brain-feature-free methods in focusing on visual attention regions of humans in both movie and non-movie datasets. Additionally, incorporating brain features improves model performance, indicating their ability to bridge the semantic gap between human cognition and visual images, allowing for more accurate capture of visual attention regions. Yaonai Wei, Chong Ma 0004, Tianyang Zhong, Lei Du 0001, Songyao Zhang, Tianming Liu 0001, Muheng Shang, Junwei Han 0001 |
BIBM | 4 |
| 2023 | Chat2Brain: A Method for Mapping Open-Ended Semantic Queries to Brain Activation MapsabstractOver decades, neuroscience has accumulated a wealth of research results in the text modality that can be used to explore cognitive processes. Meta-analysis is a typical method that successfully establishes a link from text queries to brain activation maps using these research results, but it still relies on an ideal query environment. In practical applications, text queries used for meta-analyses may encounter issues such as semantic redundancy and ambiguity, resulting in an inaccurate mapping to brain images. On the other hand, large language models (LLMs) like ChatGPT have shown great potential in tasks such as context understanding and reasoning, displaying a high degree of consistency with human natural language. Hence, LLMs could improve the connection between text modality and neuroscience, resolving existing challenges of meta-analyses. In this study, we propose a method called Chat2Brain that combines LLMs to basic text-2-image model, known as Text2Brain, to map open-ended semantic queries to brain activation maps in data-scarce and complex query environments. By utilizing the understanding and reasoning capabilities of LLMs, the performance of the mapping model is optimized by transferring text queries to semantic queries. We demonstrate that Chat2Brain can synthesize anatomically plausible neural activation patterns for more complex tasks of text queries. Yaonai Wei, Tianyang Zhong, Songyao Zhang, Xiao Li 0024, Lin Zhao 0004, Zhengliang Liu, Muheng Shang, Tianming Liu 0001, Chong Ma 0004, Lei Du 0001, Junwei Han 0001 |
BIBM | 12 |
| 2023 | Identifying Main and Epistasis Effects of Genetic Variations on Neuroimaging Phenotypes Using Effective Feature Interaction LearningabstractBrain imaging genetics investigates the complex relationships between genetic variations and brain imaging quantitative traits (QTs). However, existing approaches primarily focus on the main effects of genetic variations, potentially neglecting the crucial role of epistasis that explains the missing heritability of brain disorders. Unfortunately, this issue has been largely unexploited. In this paper, to fill this gap, we present Multi-Task feature interaction-aware Sparse Canonical Correlation Analysis (MTfiSCCA) to identify disease-related main effect and epistasis of risk genetic factors on multimodal neuroimaging phenotypes simultaneously. To ensure stability and interpretation, we use innovative sparsity-inducing penalties to identify biomarkers that make significant contributions. Additionally, we develop an efficient optimization algorithm to solve the proposed method, which converges to a local optimum. Experimental results on the Alzheimer’s disease neuroimaging initiative (ADNI) dataset show that our MTfiSCCA method achieves higher canonical correlation coefficients (CCC) and better feature selection subsets such as disease-related biomarkers compared to the state-of-the-art methods. Furthermore, MTfiSCCA reveals interpretable epistasis among genetic variations implicated in AD, offering novel insights into the underlying pathogenic mechanisms of brain disorders such as Alzheimer’s disease (AD). Jin Zhang 0023, Muheng Shang, Duo Xi, Minjianan Zhang, Lei Guo 0002, Lei Du 0001 |
BIBM | 7 |
| 2023 | Identification of Disease-Sensitive Brain Imaging Phenotypes and Genetic Factors Using GWAS Summary Statistics
Duo Xi, Dingnan Cui, Jin Zhang 0023, Muheng Shang, Minjianan Zhang, Lei Guo 0002, Junwei Han 0001, Lei Du 0001 |
MICCAI (5) | 8 |
| 2023 | Adaptive structured sparse multiview canonical correlation analysis for multimodal brain imaging association identification
Lei Du 0001, Huiai Wang, Jin Zhang 0023, Shu Zhang 0001, Lei Guo 0002, Junwei Han 0001 |
Sci. China Inf. Sci. | 1 |
| 2022 | A Sparse Multi-task Contrastive and Discriminative Learning Method with Feature Selection for Brain Imaging GeneticsabstractAlzheimer’s disease (AD) is a very complex neurodegenerative disease. Generally, different diagnostic groups could exhibit discriminative and specific patterns, including the single nucleotide polymorphisms (SNPs), brain imaging quantitative traits (QTs), as well as their associations, which may facilitate the comprehensive understanding of AD. However, most existing methods cannot guarantee to identify discriminative or class-specific biomarkers or both of them. To overcome this shortcoming, we propose a sparse multi-task contrastive and discriminative learning approach (MTCDA) to jointly learn the discriminative and specific patterns for multiple diagnostic groups. MTCDA can identify the class-relevant and discriminative SNP-QTs associations, and relevant SNPs, imaging QTs underpinning this relationship. We introduce an efficient algorithm to solve the proposed method which converges to a local optimum. The experimental results on Alzheimer’s Disease Neuroimaging Initiative (ADNI) show that MTCDA can obtain higher canonical correlation coefficients, classification accuracy and better feature selection results than state-of-the-art methods, which demonstrates the potential of our method for multi-class brain imaging genetics. Jin Zhang 0023, Muheng Shang, Minjianan Zhang, Duo Xi, Lei Guo 0002, Junwei Han 0001, Lei Du 0001 |
BIBM | 8 |
| 2022 | Identification of multimodal brain imaging association via a parameter decomposition based sparse multi-view canonical correlation analysis methodabstractBACKGROUND: With the development of noninvasive imaging technology, collecting different imaging measurements of the same brain has become more and more easy. These multimodal imaging data carry complementary information of the same brain, with both specific and shared information being intertwined. Within these multimodal data, it is essential to discriminate the specific information from the shared information since it is of benefit to comprehensively characterize brain diseases. While most existing methods are unqualified, in this paper, we propose a parameter decomposition based sparse multi-view canonical correlation analysis (PDSMCCA) method. PDSMCCA could identify both modality-shared and -specific information of multimodal data, leading to an in-depth understanding of complex pathology of brain disease. RESULTS: Compared with the SMCCA method, our method obtains higher correlation coefficients and better canonical weights on both synthetic data and real neuroimaging data. This indicates that, coupled with modality-shared and -specific feature selection, PDSMCCA improves the multi-view association identification and shows meaningful feature selection capability with desirable interpretation. CONCLUSIONS: The novel PDSMCCA confirms that the parameter decomposition is a suitable strategy to identify both modality-shared and -specific imaging features. The multimodal association and the diverse information of multimodal imaging data enable us to better understand the brain disease such as Alzheimer's disease. Jin Zhang 0023, Huiai Wang, Ying Zhao 0015, Lei Guo 0002, Lei Du 0001 |
BMC Bioinform. | 5 |
| 2021 | Improved Multi-task SCCA for Brain Imaging Genetics via Joint Consideration of the Diagnosis, Parameter Decomposition and Network ConstraintsabstractBrain imaging genetics develops rapidly, aiming to identify bi-multivariate associations between genetic loci and neuroimaging quantitative traits (QTs). The multi-task Sparse Canonical Correlation Analysis (MTSCCA) is a popular and effective technique in this area since it obtains superior results than those single-task based SCCA methods. Unfortunately, the most existing MTSCCA methods are either unsupervised or incapable of identifying the shared and specific patterns of multimodal neuroimaging QTs simultaneously. In this paper, we propose a novel diagnosis guided MTSCCA to identify the association between genetic and imaging phenotypic markers. Our method has three merits. First, it follows the same modeling paradigm of previous MTSCCA. This enables it to incorporate multimodal imaging QTs jointly, thereby facilitating a more comprehensive identification of genetic factors. Second, our method utilizes the parameter decomposition which could identify both modality-shared and -specific imaging QTs, and further uncovers their genetic mechanisms. Third, we also employed a new network constraint which could find out potentially meaningful brain imaging networks. Compared with conventional SCCA methods including both single-task and multi-task ones, the proposed method has improved or comparable correlation coefficients, and obtains a clean imaging pattern of good meaning. In addition, these results on the Alzheimer’s disease neuroimaging initiative (ADNI) cohort show that our method selects meaningful biomarkers, indicating that it could offer a significant addition to brain imaging genetic studies. Xin Zhang 0151, Yipeng Hao, Jin Zhang 0023, Shihong Zou, Songyun Xie, Lei Du 0001 |
BIBM | 6 |
| 2021 | Corrigendum to Identifying associations among genomic, proteomic and imaging biomarkers via adaptive sparse multi-view canonical correlation analysis [Medical Image Analysis 70 (2021) 1-12/102003]
Lei Du 0001, Jin Zhang 0023, Huiai Wang, Lei Guo 0002, Junwei Han 0001 |
Medical Image Anal. | 1 |
| 2021 | Identifying associations among genomic, proteomic and imaging biomarkers via adaptive sparse multi-view canonical correlation analysis
Lei Du 0001, Jin Zhang 0023, Huiai Wang, Lei Guo 0002, Junwei Han 0001 |
Medical Image Anal. | 1 |
| 2021 | Multi-Task Sparse Canonical Correlation Analysis with Application to Multi-Modal Brain Imaging GeneticsabstractBrain imaging genetics studies the genetic basis of brain structures and functionalities via integrating genotypic data such as single nucleotide polymorphisms (SNPs) and imaging quantitative traits (QTs). In this area, both multi-task learning (MTL) and sparse canonical correlation analysis (SCCA) methods are widely used since they are superior to those independent and pairwise univariate analysis. MTL methods generally incorporate a few of QTs and could not select features from multiple QTs; while SCCA methods typically employ one modality of QTs to study its association with SNPs. Both MTL and SCCA are computational expensive as the number of SNPs increases. In this paper, we propose a novel multi-task SCCA (MTSCCA) method to identify bi-multivariate associations between SNPs and multi-modal imaging QTs. MTSCCA could make use of the complementary information carried by different imaging modalities. MTSCCA enforces sparsity at the group level via the${\mathrm G}_{2,1}$-norm, and jointly selects features across multiple tasks for SNPs and QTs via the$\ell _{2,1}$-norm. A fast optimization algorithm is proposed using the grouping information of SNPs. Compared with conventional SCCA methods, MTSCCA obtains better correlation coefficients and canonical weights patterns. In addition, MTSCCA runs very fast and easy-to-implement, indicating its potential power in genome-wide brain-wide imaging genetics. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Andrew J. Saykin, Lei Guo 0002, Li Shen 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2020 | Mining High-order Multimodal Brain Image Associations via Sparse Tensor Canonical Correlation AnalysisabstractNeuroimaging techniques have shown increasing power to understand the neuropathology of brain disorders. Multimodal brain imaging data carry distinct but complementary information and thus could depict brain disorders comprehensively. To deepen our understanding, it is essential to investigate the intrinsic associations among multiple modalities. To date, the pairwise correlations between imaging data captured by different imaging modalities have been well studied, leaving formidable challenges to identify high-order associations. In this paper, we first propose a new sparse tensor canonical correlation analysis (STCCA) with feature selection to analyze the complex high-order relationships among multimodal brain imaging data. In addition, we find that methods for identifying pairwise associations and high-order associations have complementary advantages, providing a sound reason to fuse them. Therefore, we further propose an improved STCCA (STCCA+) which integrates STCCA and sparse multiple CCA (SMCCA) to fully uncover associations among multiple imaging modalities. The proposed STCCA+detects equivalent association levels among multimodal imaging data compared to SMCCA. Most importantly, both STCCA and STCCA+yield modality-consistent imaging markers and modality-specific ones, assuring a better and meaningful feature selection capability. Finally, the identified imaging markers and their high-order correlations could form a comprehensive indication of brain disorders, showing their promise in high-order multimodal brain imaging analysis. Lei Du 0001, Jin Zhang 0023, Minjianan Zhang, Huiai Wang, Lei Guo 0002, Junwei Han 0001 |
BIBM | 1 |
| 2020 | Species-Shared and -Specific Structural Connections Revealed by Dirty Multi-task Regression
Xi Jiang 0001, Lei Guo 0002, Xiaoping Hu 0001, Tianming Liu 0001, Lei Du 0001 |
MICCAI (7) | 7 |
| 2020 | Identifying diagnosis-specific genotype-phenotype associations via joint multitask sparse canonical correlation analysis and classificationabstractMOTIVATION: Brain imaging genetics studies the complex associations between genotypic data such as single nucleotide polymorphisms (SNPs) and imaging quantitative traits (QTs). The neurodegenerative disorders usually exhibit the diversity and heterogeneity, originating from which different diagnostic groups might carry distinct imaging QTs, SNPs and their interactions. Sparse canonical correlation analysis (SCCA) is widely used to identify bi-multivariate genotype-phenotype associations. However, most existing SCCA methods are unsupervised, leading to an inability to identify diagnosis-specific genotype-phenotype associations. RESULTS: In this article, we propose a new joint multitask learning method, named MT-SCCALR, which absorbs the merits of both SCCA and logistic regression. MT-SCCALR learns genotype-phenotype associations of multiple tasks jointly, with each task focusing on identifying one diagnosis-specific genotype-phenotype pattern. Meanwhile, MT-SCCALR cannot only select relevant SNPs and imaging QTs for each diagnostic group alone, but also allows the selection of those shared by multiple diagnostic groups. We derive an efficient optimization algorithm whose convergence to a local optimum is guaranteed. Compared with two state-of-the-art methods, MT-SCCALR yields better or similar canonical correlation coefficients and classification performances. In addition, it owns much better discriminative canonical weight patterns of great interest than competitors. This demonstrates the power and capability of MTSCCAR in identifying diagnostically heterogeneous genotype-phenotype patterns, which would be helpful to understand the pathophysiology of brain disorders. AVAILABILITY AND IMPLEMENTATION: The software is publicly available at https://github.com/dulei323/MTSCCALR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 1 |
| 2020 | Detecting genetic associations with brain imaging phenotypes in Alzheimer's disease via a novel structured SCCA approach
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Andrew J. Saykin, Lei Guo 0002, Li Shen 0001 |
Medical Image Anal. | 1 |
| 2020 | Associating Multi-Modal Brain Imaging Phenotypes and Genetic Risk Factors via a Dirty Multi-Task Learning MethodabstractBrain imaging genetics becomes more and more important in brain science, which integrates genetic variations and brain structures or functions to study the genetic basis of brain disorders. The multi-modal imaging data collected by different technologies, measuring the same brain distinctly, might carry complementary information. Unfortunately, we do not know the extent to which the phenotypic variance is shared among multiple imaging modalities, which further might trace back to the complex genetic mechanism. In this paper, we propose a novel dirty multi-task sparse canonical correlation analysis (SCCA) to study imaging genetic problems with multi-modal brain imaging quantitative traits (QTs) involved. The proposed method takes advantages of the multi-task learning and parameter decomposition. It can not only identify the shared imaging QTs and genetic loci across multiple modalities, but also identify the modality-specific imaging QTs and genetic loci, exhibiting a flexible capability of identifying complex multi-SNP-multi-QT associations. Using the state-of-the-art multi-view SCCA and multi-task SCCA, the proposed method shows better or comparable canonical correlation coefficients and canonical weights on both synthetic and real neuroimaging genetic data. In addition, the identified modality-consistent biomarkers, as well as the modality-specific biomarkers, provide meaningful and interesting information, demonstrating the dirty multi-task SCCA could be a powerful alternative method in multi-modal brain imaging genetics. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Andrew J. Saykin, Li Shen 0001 |
IEEE Trans. Medical Imaging | 1 |
| 2019 | A Dirty Multi-task Learning Method for Multi-modal Brain Imaging Genetics
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
MICCAI (4) | 1 |
| 2019 | Identifying progressive imaging genetic patterns via multi-task sparse canonical correlation analysis: a longitudinal study of the ADNI cohortabstractMOTIVATION: Identifying the genetic basis of the brain structure, function and disorder by using the imaging quantitative traits (QTs) as endophenotypes is an important task in brain science. Brain QTs often change over time while the disorder progresses and thus understanding how the genetic factors play roles on the progressive brain QT changes is of great importance and meaning. Most existing imaging genetics methods only analyze the baseline neuroimaging data, and thus those longitudinal imaging data across multiple time points containing important disease progression information are omitted. RESULTS: We propose a novel temporal imaging genetic model which performs the multi-task sparse canonical correlation analysis (T-MTSCCA). Our model uses longitudinal neuroimaging data to uncover that how single nucleotide polymorphisms (SNPs) play roles on affecting brain QTs over the time. Incorporating the relationship of the longitudinal imaging data and that within SNPs, T-MTSCCA could identify a trajectory of progressive imaging genetic patterns over the time. We propose an efficient algorithm to solve the problem and show its convergence. We evaluate T-MTSCCA on 408 subjects from the Alzheimer's Disease Neuroimaging Initiative database with longitudinal magnetic resonance imaging data and genetic data available. The experimental results show that T-MTSCCA performs either better than or equally to the state-of-the-art methods. In particular, T-MTSCCA could identify higher canonical correlation coefficients and capture clearer canonical weight patterns. This suggests that T-MTSCCA identifies time-consistent and time-dependent SNPs and imaging QTs, which further help understand the genetic basis of the brain QT changes over the time during the disease progression. AVAILABILITY AND IMPLEMENTATION: The software and simulation data are publicly available at https://github.com/dulei323/TMTSCCA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Du 0001, Kefei Liu 0001, Lei Zhu 0011, Xiaohui Yao, Shannon L. Risacher, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 1 |
| 2018 | Fast Multi-Task SCCA Learning with Feature Selection for Multi-Modal Brain Imaging Genetics
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
BIBM | 1 |
| 2018 | A novel SCCA approach via truncated ℓ1-norm and truncated group lasso for brain imaging geneticsabstractMOTIVATION: Brain imaging genetics, which studies the linkage between genetic variations and structural or functional measures of the human brain, has become increasingly important in recent years. Discovering the bi-multivariate relationship between genetic markers such as single-nucleotide polymorphisms (SNPs) and neuroimaging quantitative traits (QTs) is one major task in imaging genetics. Sparse Canonical Correlation Analysis (SCCA) has been a popular technique in this area for its powerful capability in identifying bi-multivariate relationships coupled with feature selection. The existing SCCA methods impose either the ℓ1-norm or its variants to induce sparsity. The ℓ0-norm penalty is a perfect sparsity-inducing tool which, however, is an NP-hard problem. RESULTS: In this paper, we propose the truncated ℓ1-norm penalized SCCA to improve the performance and effectiveness of the ℓ1-norm based SCCA methods. Besides, we propose an efficient optimization algorithms to solve this novel SCCA problem. The proposed method is an adaptive shrinkage method via tuning τ. It can avoid the time intensive parameter tuning if given a reasonable small τ. Furthermore, we extend it to the truncated group-lasso (TGL), and propose TGL-SCCA model to improve the group-lasso-based SCCA methods. The experimental results, compared with four benchmark methods, show that our SCCA methods identify better or similar correlation coefficients, and better canonical loading profiles than the competing methods. This demonstrates the effectiveness and efficiency of our methods in discovering interesting imaging genetic associations. AVAILABILITY AND IMPLEMENTATION: The Matlab code and sample data are freely available at http://www.iu.edu/∼shenlab/tools/tlpscca/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 1 |
| 2016 | Sparse Canonical Correlation Analysis via truncated ℓ1-norm with application to brain imaging geneticsabstractDiscovering bi-multivariate associations between genetic markers and neuroimaging quantitative traits is a major task in brain imaging genetics. Sparse Canonical Correlation Analysis (SCCA) is a popular technique in this area for its powerful capability in identifying bi-multivariate relationships coupled with feature selection. The existing SCCA methods impose either the ℓ1-norm or its variants. The ℓ0-norm is more desirable, which however remains unexplored since the ℓ0-norm minimization is NP-hard. In this paper, we impose the truncated ℓ1-norm to improve the performance of the ℓ1-norm based SCCA methods. Besides, we propose two efficient optimization algorithms and prove their convergence. The experimental results, compared with two benchmark methods, show that our method identifies better and meaningful canonical loading patterns in both simulated and real imaging genetic analyse. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
BIBM | 1 |
| 2016 | Species Preserved and Exclusive Structural Connections Revealed by Sparse CCA
Xiao Li 0024, Lei Du 0001, Xintao Hu, Xi Jiang 0001, Lei Guo 0002, Tianming Liu 0001 |
MICCAI (1) | 2 |
| 2016 | Structured sparse canonical correlation analysis for brain imaging genetics: an improved GraphNet methodabstractMOTIVATION: Structured sparse canonical correlation analysis (SCCA) models have been used to identify imaging genetic associations. These models either use group lasso or graph-guided fused lasso to conduct feature selection and feature grouping simultaneously. The group lasso based methods require prior knowledge to define the groups, which limits the capability when prior knowledge is incomplete or unavailable. The graph-guided methods overcome this drawback by using the sample correlation to define the constraint. However, they are sensitive to the sign of the sample correlation, which could introduce undesirable bias if the sign is wrongly estimated. RESULTS: We introduce a novel SCCA model with a new penalty, and develop an efficient optimization algorithm. Our method has a strong upper bound for the grouping effect for both positively and negatively correlated features. We show that our method performs better than or equally to three competing SCCA models on both synthetic and real data. In particular, our method identifies stronger canonical correlations and better canonical loading patterns, showing its promise for revealing interesting imaging genetic associations. AVAILABILITY AND IMPLEMENTATION: The Matlab code and sample data are freely available at http://www.iu.edu/∼shenlab/tools/angscca/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Du 0001, Heng Huang 0001, Sungeun Kim, Shannon L. Risacher, Mark Inlow, Jason H. Moore, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 1 |
| 2015 | A Selective Detector Ensemble for Concept Drift DetectionabstractConcept drifts usually originate from many causes instead of only one, which result in two types of concept drifts: abrupt drifts and gradual drifts. From the point of view of speed, concept drifts pose strong challenges for data stream mining. In this paper, we propose a selective detector ensemble to detect both abrupt and gradual drifts. We first present our detector ensemble construction method, and then introduce how to use this ensemble to detect concept drifts with the proposed early-find-early-report rule. To evaluate the performance of our method, we compare it with four drift detection methods on eight publicly available data sets containing various concept drifts. The experimental results show that compared with those benchmarks, our ensemble method can effectively improve the recall and false negative rate without significantly increasing the false positive rate, and has stronger generalization ability than those single-change-indicator-based methods. Lei Du 0001, Qinbao Song, Lei Zhu 0011, Xiaoyan Zhu 0003 |
Comput. J. | 1 |
| 2014 | A Novel Structure-Aware Sparse Learning Algorithm for Brain Imaging Genetics
Lei Du 0001, Sungeun Kim, Shannon L. Risacher, Heng Huang 0001, Mark Inlow, Jason H. Moore, Andrew J. Saykin, Li Shen 0001 |
MICCAI (3) | 1 |
| 2014 | Transcriptome-guided amyloid imaging genetic analysis via a novel structured sparse learning algorithmabstractMOTIVATION: Imaging genetics is an emerging field that studies the influence of genetic variation on brain structure and function. The major task is to examine the association between genetic markers such as single-nucleotide polymorphisms (SNPs) and quantitative traits (QTs) extracted from neuroimaging data. The complexity of these datasets has presented critical bioinformatics challenges that require new enabling tools. Sparse canonical correlation analysis (SCCA) is a bi-multivariate technique used in imaging genetics to identify complex multi-SNP-multi-QT associations. However, most of the existing SCCA algorithms are designed using the soft thresholding method, which assumes that the input features are independent from one another. This assumption clearly does not hold for the imaging genetic data. In this article, we propose a new knowledge-guided SCCA algorithm (KG-SCCA) to overcome this limitation as well as improve learning results by incorporating valuable prior knowledge. RESULTS: The proposed KG-SCCA method is able to model two types of prior knowledge: one as a group structure (e.g. linkage disequilibrium blocks among SNPs) and the other as a network structure (e.g. gene co-expression network among brain regions). The new model incorporates these prior structures by introducing new regularization terms to encourage weight similarity between grouped or connected features. A new algorithm is designed to solve the KG-SCCA model without imposing the independence constraint on the input features. We demonstrate the effectiveness of our algorithm with both synthetic and real data. For real data, using an Alzheimer's disease (AD) cohort, we examine the imaging genetic associations between all SNPs in the APOE gene (i.e. top AD gene) and amyloid deposition measures among cortical regions (i.e. a major AD hallmark). In comparison with a widely used SCCA implementation, our KG-SCCA algorithm produces not only improved cross-validation performances but also biologically meaningful results. AVAILABILITY: Software is freely available on request. Lei Du 0001, Sungeun Kim, Shannon L. Risacher, Heng Huang 0001, Jason H. Moore, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 2 |