VLDB 2026 Research / reviewers in the wild / expert
Andrew J. Saykin
dblp:32/2887
· DBLP profile ↗
63ranked-venue papers
0as first author
12since 2021 · last 2025
0000-0002-1376-8532ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 53 · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19Artificial intelligence and machine learning · 7Databases, data management, data science and information retrieval · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | MG-TCCA: Tensor Canonical Correlation Analysis Across Multiple GroupsabstractTensor Canonical Correlation Analysis (TCCA) is a commonly employed statistical method utilized to examine linear associations between two sets of tensor datasets. However, the existing TCCA models fail to adequately address the heterogeneity present in real-world tensor data, such as brain imaging data collected from diverse groups characterized by factors like sex and race. Consequently, these models may yield biased outcomes. In order to surmount this constraint, we propose a novel approach called Multi-Group TCCA (MG-TCCA), which enables the joint analysis of multiple subgroups. By incorporating a dual sparsity structure and a block coordinate ascent algorithm, our MG-TCCA method effectively addresses heterogeneity and leverages information across different groups to identify consistent signals. This novel approach facilitates the quantification of shared and individual structures, reduces data dimensionality, and enables visual exploration. To empirically validate our approach, we conduct a study focused on investigating correlations between two brain positron emission tomography (PET) modalities (AV-45 and FDG) within an Alzheimer's disease (AD) cohort. Our results demonstrate that MG-TCCA surpasses traditional TCCA and Sparse TCCA (STCCA) in identifying sex-specific cross-modality imaging correlations. This heightened performance of MG-TCCA provides valuable insights for the characterization of multimodal imaging biomarkers in AD. Zhuoping Zhou, Boning Tong, D. Ataee Tarzanagh, Bojian Hou, Andrew J. Saykin, Qi Long, Li Shen 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2024 | Integrative Analysis of Amyloid Imaging and Genetics Reveals Subtypes of Alzheimer Progression in Early Stage
Neel Sangani, Ruiming Wu, Pradeep Varathan, Alice Patania, Shannon L. Risacher, Kwangsik Nho, Liana G. Apostolova, Andrew J. Saykin, Li Shen 0001 |
AIME (2) | 9 |
| 2024 | Interpretable deep clustering survival machines for Alzheimer's disease subtype discovery
Bojian Hou, Zixuan Wen, Jingxuan Bao, Richard Zhang 0001, Boning Tong, Shu Yang 0009, Junhao Wen 0002, Yuhan Cui, Jason H. Moore, Andrew J. Saykin, Heng Huang 0001, Paul M. Thompson, Marylyn D. Ritchie, Christos Davatzikos, Li Shen 0001 |
Medical Image Anal. | 10 |
| 2023 | Integrative analysis of multi-omics and imaging data with incorporation of biological information via structural Bayesian factor analysisabstractMOTIVATION: With the rapid development of modern technologies, massive data are available for the systematic study of Alzheimer's disease (AD). Though many existing AD studies mainly focus on single-modality omics data, multi-omics datasets can provide a more comprehensive understanding of AD. To bridge this gap, we proposed a novel structural Bayesian factor analysis framework (SBFA) to extract the information shared by multi-omics data through the aggregation of genotyping data, gene expression data, neuroimaging phenotypes and prior biological network knowledge. Our approach can extract common information shared by different modalities and encourage biologically related features to be selected, guiding future AD research in a biologically meaningful way. METHOD: Our SBFA model decomposes the mean parameters of the data into a sparse factor loading matrix and a factor matrix, where the factor matrix represents the common information extracted from multi-omics and imaging data. Our framework is designed to incorporate prior biological network information. Our simulation study demonstrated that our proposed SBFA framework could achieve the best performance compared with the other state-of-the-art factor-analysis-based integrative analysis methods. RESULTS: We apply our proposed SBFA model together with several state-of-the-art factor analysis models to extract the latent common information from genotyping, gene expression and brain imaging data simultaneously from the ADNI biobank database. The latent information is then used to predict the functional activities questionnaire score, an important measurement for diagnosis of AD quantifying subjects' abilities in daily life. Our SBFA model shows the best prediction performance compared with the other factor analysis models. AVAILABILITY: Code are publicly available at https://github.com/JingxuanBao/SBFA. CONTACT: [email protected]. Jingxuan Bao, Changgee Chang, Qiyiwen Zhang, Andrew J. Saykin, Li Shen 0001, Qi Long |
Briefings Bioinform. | 4 |
| 2022 | Preference Matrix Guided Sparse Canonical Correlation Analysis for Genetic Study of Quantitative Traits in Alzheimer's DiseaseabstractInvestigating the relationship between genetic variation and phenotypic traits is a key issue in quantitative genetics. Specifically for Alzheimer's disease, the association between genetic markers and quantitative traits remains vague while, once identified, will provide valuable guidance for the study and development of genetic-based treatment approaches. Currently, to analyze the association of two modalities, sparse canonical correlation analysis (SCCA) is commonly used to compute one sparse linear combination of the variable features for each modality, giving a pair of linear combination vectors in total that maximizes the cross-correlation between the analyzed modalities. One drawback of the plain SCCA model is that the existing findings and knowledge cannot be integrated into the model as priors to help extract interesting correlation as well as identify biologically meaningful genetic and phenotypic markers. To bridge this gap, we introduce preference matrix guided SCCA (PM-SCCA) that not only takes priors encoded as a preference matrix but also maintains computational simplicity. A simulation study and a real-data experiment are conducted to investigate the effectiveness of the model. Both experiments demonstrate that the proposed PM-SCCA model can capture not only genotype-phenotype correlation but also relevant features effectively. Jiahang Sha, Jingxuan Bao, Kefei Liu 0001, Shu Yang 0009, Zixuan Wen, Yuhan Cui, Junhao Wen 0002, Christos Davatzikos, Jason H. Moore, Andrew J. Saykin, Qi Long, Li Shen 0001 |
BIBM | 10 |
| 2022 | Consistency of Graph Theoretical Measurements of Alzheimer's Disease Fiber Density Connectomes Across Multiple Parcellation ScalesabstractGraph theoretical measures have frequently been used to study disrupted connectivity in Alzheimer's disease human brain connectomes. However, prior studies have noted that differences in graph creation methods are confounding factors that may alter the topological observations found in these measures. In this study, we conduct a novel investigation regarding the effect of parcellation scale on graph theoretical measures computed for fiber density networks derived from diffusion tensor imaging. We computed 4 network-wide graph theoretical measures of average clustering coefficient, transitivity, characteristic path length, and global efficiency, and we tested whether these measures are able to consistently identify group differences among healthy control (HC), mild cognitive impairment (MCI), and AD groups in the Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort across 5 scales of the Lausanne parcellation. We found that the segregative measure of transtivity offered the greatest consistency across scales in distinguishing between healthy and diseased groups, while the other measures were impacted by the selection of scale to varying degrees. Global efficiency was the second most consistent measure that we tested, where the measure could distinguish between HC and MCI in all 5 scales and between HC and AD in 3 out of 5 scales. Characteristic path length was highly sensitive to the variation in scale, corroborating previous findings, and could not identify group differences in many of the scales. Average clustering coefficient was also greatly impacted by scale, as it consistently failed to identify group differences in the higher resolution parcellations. From these results, we conclude that many graph theoretical measures are sensitive to the selection of parcellation scale, and further development in methodology is needed to offer a more robust characterization of AD's relationship with disrupted connectivity. Frederick H. Xu, Sumita Garai, Duy Duong-Tran, Andrew J. Saykin, Yize Zhao, Li Shen 0001 |
BIBM | 4 |
| 2022 | Deep learning-based identification of genetic variants: application to Alzheimer's disease classificationabstractDeep learning is a promising tool that uses nonlinear transformations to extract features from high-dimensional data. Deep learning is challenging in genome-wide association studies (GWAS) with high-dimensional genomic data. Here we propose a novel three-step approach (SWAT-CNN) for identification of genetic variants using deep learning to identify phenotype-related single nucleotide polymorphisms (SNPs) that can be applied to develop accurate disease classification models. In the first step, we divided the whole genome into nonoverlapping fragments of an optimal size and then ran convolutional neural network (CNN) on each fragment to select phenotype-associated fragments. In the second step, using a Sliding Window Association Test (SWAT), we ran CNN on the selected fragments to calculate phenotype influence scores (PIS) and identify phenotype-associated SNPs based on PIS. In the third step, we ran CNN on all identified SNPs to develop a classification model. We tested our approach using GWAS data from the Alzheimer's Disease Neuroimaging Initiative (ADNI) including (N = 981; cognitively normal older adults (CN) = 650 and AD = 331). Our approach identified the well-known APOE region as the most significant genetic locus for AD. Our classification model achieved an area under the curve (AUC) of 0.82, which was compatible with traditional machine learning approaches, random forest and XGBoost. SWAT-CNN, a novel deep learning-based genome-wide approach, identified AD-associated SNPs and a classification model for AD and may hold promise for a range of biomedical applications. Taeho Jo 0002, Kwangsik Nho, Paula Bice, Andrew J. Saykin |
Briefings Bioinform. | 4 |
| 2022 | Multi-task learning based structured sparse canonical correlation analysis for brain imaging genetics
Mansu Kim, Eun Jeong Min, Kefei Liu 0001, Andrew J. Saykin, Jason H. Moore, Qi Long, Li Shen 0001 |
Medical Image Anal. | 5 |
| 2021 | A novel deep learning method for predictive modeling of microbiome dataabstractWith the development and decreasing cost of next-generation sequencing technologies, the study of the human microbiome has become a rapid expanding research field, which provides an unprecedented opportunity in various clinical applications such as drug response predictions and disease diagnosis. It is thus essential and desirable to build a prediction model for clinical outcomes based on microbiome data that usually consist of taxon abundance and a phylogenetic tree. Importantly, all microbial species are not uniformly distributed in the phylogenetic tree but tend to be clustered at different phylogenetic depths. Therefore, the phylogenetic tree represents a unique correlation structure of microbiome, which can be an important prior to improve the prediction performance. However, prediction methods that consider the phylogenetic tree in an efficient and rigorous way are under-developed. Here, we develop a novel deep learning prediction method MDeep (microbiome-based deep learning method) to predict both continuous and binary outcomes. Conceptually, MDeep designs convolutional layers to mimic taxonomic ranks with multiple convolutional filters on each convolutional layer to capture the phylogenetic correlation among microbial species in a local receptive field and maintain the correlation structure across different convolutional layers via feature mapping. Taken together, the convolutional layers with its built-in convolutional filters capture microbial signals at different taxonomic levels while encouraging local smoothing and preserving local connectivity induced by the phylogenetic tree. We use both simulation studies and real data applications to demonstrate that MDeep outperforms competing methods in both regression and binary classifications. Availability and Implementation: MDeep software is available at https://github.com/lichen-lab/MDeep Contact:[email protected]. Ye Wang 0024, Tathagata Bhattacharya, Xiao Qin 0001, Andrew J. Saykin, Li Chen 0029 |
Briefings Bioinform. | 7 |
| 2021 | WEVar: a novel statistical learning framework for predicting noncoding regulatory variantsabstractUnderstanding the functional consequence of noncoding variants is of great interest. Though genome-wide association studies or quantitative trait locus analyses have identified variants associated with traits or molecular phenotypes, most of them are located in the noncoding regions, making the identification of causal variants a particular challenge. Existing computational approaches developed for prioritizing noncoding variants produce inconsistent and even conflicting results. To address these challenges, we propose a novel statistical learning framework, which directly integrates the precomputed functional scores from representative scoring methods. It will maximize the usage of integrated methods by automatically learning the relative contribution of each method and produce an ensemble score as the final prediction. The framework consists of two modes. The first 'context-free' mode is trained using curated causal regulatory variants from a wide range of context and is applicable to predict regulatory variants of unknown and diverse context. The second 'context-dependent' mode further improves the prediction when the training and testing variants are from the same context. By evaluating the framework via both simulation and empirical studies, we demonstrate that it outperforms integrated scoring methods and the ensemble score successfully prioritizes experimentally validated regulatory variants in multiple risk loci. Ye Wang 0024, Xiao Qin 0001, Andrew J. Saykin, Li Chen 0029 |
Briefings Bioinform. | 8 |
| 2021 | Integrative-omics for discovery of network-level disease biomarkers: a case study in Alzheimer's diseaseabstractA large number of genetic variations have been identified to be associated with Alzheimer's disease (AD) and related quantitative traits. However, majority of existing studies focused on single types of omics data, lacking the power of generating a community including multi-omic markers and their functional connections. Because of this, the immense value of multi-omics data on AD has attracted much attention. Leveraging genomic, transcriptomic and proteomic data, and their backbone network through functional relations, we proposed a modularity-constrained logistic regression model to mine the association between disease status and a group of functionally connected multi-omic features, i.e. single-nucleotide polymorphisms (SNPs), genes and proteins. This new model was applied to the real data collected from the frontal cortex tissue in the Religious Orders Study and Memory and Aging Project cohort. Compared with other state-of-art methods, it provided overall the best prediction performance during cross-validation. This new method helped identify a group of densely connected SNPs, genes and proteins predictive of AD status. These SNPs are mostly expression quantitative trait loci in the frontal region. Brain-wide gene expression profile of these genes and proteins were highly correlated with the brain activation map of 'vision', a brain function partly controlled by frontal cortex. These genes and proteins were also found to be associated with the amyloid deposition, cortical volume and average thickness of frontal regions. Taken together, these results suggested a potential pathway underlying the development of AD from SNPs to gene expression, protein expression and ultimately brain functional and structural changes. Linhui Xie, Pradeep Varathan, Kwangsik Nho, Shannon L. Risacher, Andrew J. Saykin, Paul Salama |
Briefings Bioinform. | 6 |
| 2021 | Multi-Task Sparse Canonical Correlation Analysis with Application to Multi-Modal Brain Imaging GeneticsabstractBrain imaging genetics studies the genetic basis of brain structures and functionalities via integrating genotypic data such as single nucleotide polymorphisms (SNPs) and imaging quantitative traits (QTs). In this area, both multi-task learning (MTL) and sparse canonical correlation analysis (SCCA) methods are widely used since they are superior to those independent and pairwise univariate analysis. MTL methods generally incorporate a few of QTs and could not select features from multiple QTs; while SCCA methods typically employ one modality of QTs to study its association with SNPs. Both MTL and SCCA are computational expensive as the number of SNPs increases. In this paper, we propose a novel multi-task SCCA (MTSCCA) method to identify bi-multivariate associations between SNPs and multi-modal imaging QTs. MTSCCA could make use of the complementary information carried by different imaging modalities. MTSCCA enforces sparsity at the group level via the${\mathrm G}_{2,1}$-norm, and jointly selects features across multiple tasks for SNPs and QTs via the$\ell _{2,1}$-norm. A fast optimization algorithm is proposed using the grouping information of SNPs. Compared with conventional SCCA methods, MTSCCA obtains better correlation coefficients and canonical weights patterns. In addition, MTSCCA runs very fast and easy-to-implement, indicating its potential power in genome-wide brain-wide imaging genetics. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Andrew J. Saykin, Lei Guo 0002, Li Shen 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2020 | Polygenic mediation analysis of Alzheimer's disease implicated intermediate amyloid imaging phenotypes
Yingxuan Eng, Xiaohui Yao, Kefei Liu 0001, Shannon L. Risacher, Andrew J. Saykin, Qi Long, Yize Zhao, Li Shen 0001 |
AMIA | 5 |
| 2020 | Identifying diagnosis-specific genotype-phenotype associations via joint multitask sparse canonical correlation analysis and classificationabstractMOTIVATION: Brain imaging genetics studies the complex associations between genotypic data such as single nucleotide polymorphisms (SNPs) and imaging quantitative traits (QTs). The neurodegenerative disorders usually exhibit the diversity and heterogeneity, originating from which different diagnostic groups might carry distinct imaging QTs, SNPs and their interactions. Sparse canonical correlation analysis (SCCA) is widely used to identify bi-multivariate genotype-phenotype associations. However, most existing SCCA methods are unsupervised, leading to an inability to identify diagnosis-specific genotype-phenotype associations. RESULTS: In this article, we propose a new joint multitask learning method, named MT-SCCALR, which absorbs the merits of both SCCA and logistic regression. MT-SCCALR learns genotype-phenotype associations of multiple tasks jointly, with each task focusing on identifying one diagnosis-specific genotype-phenotype pattern. Meanwhile, MT-SCCALR cannot only select relevant SNPs and imaging QTs for each diagnostic group alone, but also allows the selection of those shared by multiple diagnostic groups. We derive an efficient optimization algorithm whose convergence to a local optimum is guaranteed. Compared with two state-of-the-art methods, MT-SCCALR yields better or similar canonical correlation coefficients and classification performances. In addition, it owns much better discriminative canonical weight patterns of great interest than competitors. This demonstrates the power and capability of MTSCCAR in identifying diagnostically heterogeneous genotype-phenotype patterns, which would be helpful to understand the pathophysiology of brain disorders. AVAILABILITY AND IMPLEMENTATION: The software is publicly available at https://github.com/dulei323/MTSCCALR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 8 |
| 2020 | Regional imaging genetic enrichment analysisabstractMOTIVATION: Brain imaging genetics aims to reveal genetic effects on brain phenotypes, where most studies examine phenotypes defined on anatomical or functional regions of interest (ROIs) given their biologically meaningful interpretation and modest dimensionality compared with voxelwise approaches. Typical ROI-level measures used in these studies are summary statistics from voxelwise measures in the region, without making full use of individual voxel signals. RESULTS: In this article, we propose a flexible and powerful framework for mining regional imaging genetic associations via voxelwise enrichment analysis, which embraces the collective effect of weak voxel-level signals and integrates brain anatomical annotation information. Our proposed method achieves three goals at the same time: (i) increase the statistical power by substantially reducing the burden of multiple comparison correction; (ii) employ brain annotation information to enable biologically meaningful interpretation and (iii) make full use of fine-grained voxelwise signals. We demonstrate our method on an imaging genetic analysis using data from the Alzheimer's Disease Neuroimaging Initiative, where we assess the collective regional genetic effects of voxelwise FDG-positron emission tomography measures between 116 ROIs and 565 373 single-nucleotide polymorphisms. Compared with traditional ROI-wise and voxelwise approaches, our method identified 2946 novel imaging genetic associations in addition to 33 ones overlapping with the two benchmark methods. In particular, two newly reported variants were further supported by transcriptome evidences from region-specific expression analysis. This demonstrates the promise of the proposed method as a flexible and powerful framework for exploring imaging genetic effects on the brain. AVAILABILITY AND IMPLEMENTATION: The R code and sample data are freely available at https://github.com/lshen/RIGEA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaohui Yao, Shan Cong, Shannon L. Risacher, Andrew J. Saykin, Jason H. Moore, Li Shen 0001 |
Bioinform. | 5 |
| 2020 | Deep learning detection of informative features in tau PET for Alzheimer's disease classificationabstractBACKGROUND: Alzheimer's disease (AD) is the most common type of dementia, typically characterized by memory loss followed by progressive cognitive decline and functional impairment. Many clinical trials of potential therapies for AD have failed, and there is currently no approved disease-modifying treatment. Biomarkers for early detection and mechanistic understanding of disease course are critical for drug development and clinical trials. Amyloid has been the focus of most biomarker research. Here, we developed a deep learning-based framework to identify informative features for AD classification using tau positron emission tomography (PET) scans. RESULTS: The 3D convolutional neural network (CNN)-based classification model of AD from cognitively normal (CN) yielded an average accuracy of 90.8% based on five-fold cross-validation. The LRP model identified the brain regions in tau PET images that contributed most to the AD classification from CN. The top identified regions included the hippocampus, parahippocampus, thalamus, and fusiform. The layer-wise relevance propagation (LRP) results were consistent with those from the voxel-wise analysis in SPM12, showing significant focal AD associated regional tau deposition in the bilateral temporal lobes including the entorhinal cortex. The AD probability scores calculated by the classifier were correlated with brain tau deposition in the medial temporal lobe in MCI participants (r = 0.43 for early MCI and r = 0.49 for late MCI). CONCLUSION: A deep learning framework combining 3D CNN and LRP algorithms can be used with tau PET images to identify informative features for AD classification and may have application for early detection during prodromal stages of AD. Taeho Jo 0002, Kwangsik Nho, Shannon L. Risacher, Andrew J. Saykin |
BMC Bioinform. | 4 |
| 2020 | Detecting genetic associations with brain imaging phenotypes in Alzheimer's disease via a novel structured SCCA approach
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Andrew J. Saykin, Lei Guo 0002, Li Shen 0001 |
Medical Image Anal. | 6 |
| 2020 | Multi-modal neuroimaging feature selection with consistent metric constraint for diagnosis of Alzheimer's disease
Xiaoke Hao, Yongjin Bao, Yingchun Guo, Ming Yu 0006, Daoqiang Zhang, Shannon L. Risacher, Andrew J. Saykin, Xiaohui Yao, Li Shen 0001 |
Medical Image Anal. | 7 |
| 2020 | Associating Multi-Modal Brain Imaging Phenotypes and Genetic Risk Factors via a Dirty Multi-Task Learning MethodabstractBrain imaging genetics becomes more and more important in brain science, which integrates genetic variations and brain structures or functions to study the genetic basis of brain disorders. The multi-modal imaging data collected by different technologies, measuring the same brain distinctly, might carry complementary information. Unfortunately, we do not know the extent to which the phenotypic variance is shared among multiple imaging modalities, which further might trace back to the complex genetic mechanism. In this paper, we propose a novel dirty multi-task sparse canonical correlation analysis (SCCA) to study imaging genetic problems with multi-modal brain imaging quantitative traits (QTs) involved. The proposed method takes advantages of the multi-task learning and parameter decomposition. It can not only identify the shared imaging QTs and genetic loci across multiple modalities, but also identify the modality-specific imaging QTs and genetic loci, exhibiting a flexible capability of identifying complex multi-SNP-multi-QT associations. Using the state-of-the-art multi-view SCCA and multi-task SCCA, the proposed method shows better or comparable canonical correlation coefficients and canonical weights on both synthetic and real neuroimaging genetic data. In addition, the identified modality-consistent biomarkers, as well as the modality-specific biomarkers, provide meaningful and interesting information, demonstrating the dirty multi-task SCCA could be a powerful alternative method in multi-modal brain imaging genetics. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Andrew J. Saykin, Li Shen 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2019 | A Dirty Multi-task Learning Method for Multi-modal Brain Imaging Genetics
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
MICCAI (4) | 8 |
| 2019 | Identifying progressive imaging genetic patterns via multi-task sparse canonical correlation analysis: a longitudinal study of the ADNI cohortabstractMOTIVATION: Identifying the genetic basis of the brain structure, function and disorder by using the imaging quantitative traits (QTs) as endophenotypes is an important task in brain science. Brain QTs often change over time while the disorder progresses and thus understanding how the genetic factors play roles on the progressive brain QT changes is of great importance and meaning. Most existing imaging genetics methods only analyze the baseline neuroimaging data, and thus those longitudinal imaging data across multiple time points containing important disease progression information are omitted. RESULTS: We propose a novel temporal imaging genetic model which performs the multi-task sparse canonical correlation analysis (T-MTSCCA). Our model uses longitudinal neuroimaging data to uncover that how single nucleotide polymorphisms (SNPs) play roles on affecting brain QTs over the time. Incorporating the relationship of the longitudinal imaging data and that within SNPs, T-MTSCCA could identify a trajectory of progressive imaging genetic patterns over the time. We propose an efficient algorithm to solve the problem and show its convergence. We evaluate T-MTSCCA on 408 subjects from the Alzheimer's Disease Neuroimaging Initiative database with longitudinal magnetic resonance imaging data and genetic data available. The experimental results show that T-MTSCCA performs either better than or equally to the state-of-the-art methods. In particular, T-MTSCCA could identify higher canonical correlation coefficients and capture clearer canonical weight patterns. This suggests that T-MTSCCA identifies time-consistent and time-dependent SNPs and imaging QTs, which further help understand the genetic basis of the brain QT changes over the time during the disease progression. AVAILABILITY AND IMPLEMENTATION: The software and simulation data are publicly available at https://github.com/dulei323/TMTSCCA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Du 0001, Kefei Liu 0001, Lei Zhu 0011, Xiaohui Yao, Shannon L. Risacher, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 7 |
| 2019 | Identifying Candidate Genetic Associations with MRI-Derived AD-Related ROI via Tree-Guided Sparse LearningabstractImaging genetics has attracted significant interests in recent studies. Traditional work has focused on mass-univariate statistical approaches that identify important single nucleotide polymorphisms (SNPs) associated with quantitative traits (QTs) of brain structure or function. More recently, to address the problem of multiple comparison and weak detection, multivariate analysis methods such as the least absolute shrinkage and selection operator (Lasso) are often used to select the most relevant SNPs associated with QTs. However, one problem of Lasso, as well as many other feature selection methods for imaging genetics, is that some useful prior information, e.g., the hierarchical structure among SNPs, are rarely used for designing a more powerful model. In this paper, we propose to identify the associations between candidate genetic features (i.e., SNPs) and magnetic resonance imaging (MRI)-derived measures using a tree-guided sparse learning (TGSL) method. The advantage of our method is that it explicitly models the complex hierarchical structure among the SNPs in the objective function for feature selection. Specifically, motivated by the biological knowledge, the hierarchical structures involving gene groups and linkage disequilibrium (LD) blocks as well as individual SNPs are imposed as a tree-guided regularization term in our TGSL model. Experimental studies on simulation data and the Alzheimer's Disease Neuroimaging Initiative (ADNI) data show that our method not only achieves better predictions than competing methods on the MRI-derived measures of AD-related region of interests (ROIs) (i.e., hippocampus, parahippocampal gyrus, and precuneus), but also identifies sparse SNP patterns at the block level to better guide the biological interpretation. Xiaoke Hao, Xiaohui Yao, Shannon L. Risacher, Andrew J. Saykin, Jintai Yu, Huifu Wang, Lan Tan, Li Shen 0001, Daoqiang Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2018 | Fast Multi-Task SCCA Learning with Feature Selection for Multi-Modal Brain Imaging Genetics
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
BIBM | 7 |
| 2018 | Interactive Machine Learning by Visualization: A Small Data SolutionabstractMachine learning algorithms and traditional data mining process usually require a large volume of data to train the algorithm-specific models, with little or no user feedback during the model building process. Such a "big data" based automatic learning strategy is sometimes unrealistic for applications where data collection or processing is very expensive or difficult, such as in clinical trials. Furthermore, expert knowledge can be very valuable in the model building process in some fields such as biomedical sciences. In this paper, we propose a new visual analytics approach to interactive machine learning and visual data mining. In this approach, multi-dimensional data visualization techniques are employed to facilitate user interactions with the machine learning and mining process. This allows dynamic user feedback in different forms, such as data selection, data labeling, and data correction, to enhance the efficiency of model building. In particular, this approach can significantly reduce the amount of data required for training an accurate model, and therefore can be highly impactful for applications where large amount of data is hard to obtain. The proposed approach is tested on two application problems: the handwriting recognition (classification) problem and the human cognitive score prediction (regression) problem. Both experiments show that visualization supported interactive machine learning and data mining can achieve the same accuracy as an automatic process can with much smaller training data sets. Shiaofen Fang, Snehasis Mukhopadhyay, Andrew J. Saykin, Li Shen 0001 |
IEEE BigData | 4 |
| 2018 | Joint High-Order Multi-Task Feature Learning to Predict the Progression of Alzheimer's Disease
Lodewijk Brand, Hua Wang 0007, Heng Huang 0001, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001 |
MICCAI (1) | 5 |
| 2018 | Network approaches to systems biology analysis of complex disease: integrative methods for multi-omics dataabstractIn the past decade, significant progress has been made in complex disease research across multiple omics layers from genome, transcriptome and proteome to metabolome. There is an increasing awareness of the importance of biological interconnections, and much success has been achieved using systems biology approaches. However, because of the typical focus on one single omics layer at a time, existing systems biology findings explain only a modest portion of complex disease. Recent advances in multi-omics data collection and sharing present us new opportunities for studying complex diseases in a more comprehensive fashion, and yet simultaneously create new challenges considering the unprecedented data dimensionality and diversity. Here, our goal is to review extant and emerging network approaches that can be applied across multiple biological layers to facilitate a more comprehensive and integrative multilayered omics analysis of complex diseases. Shannon L. Risacher, Li Shen 0001, Andrew J. Saykin |
Briefings Bioinform. | 4 |
| 2018 | A novel SCCA approach via truncated ℓ1-norm and truncated group lasso for brain imaging geneticsabstractMOTIVATION: Brain imaging genetics, which studies the linkage between genetic variations and structural or functional measures of the human brain, has become increasingly important in recent years. Discovering the bi-multivariate relationship between genetic markers such as single-nucleotide polymorphisms (SNPs) and neuroimaging quantitative traits (QTs) is one major task in imaging genetics. Sparse Canonical Correlation Analysis (SCCA) has been a popular technique in this area for its powerful capability in identifying bi-multivariate relationships coupled with feature selection. The existing SCCA methods impose either the ℓ1-norm or its variants to induce sparsity. The ℓ0-norm penalty is a perfect sparsity-inducing tool which, however, is an NP-hard problem. RESULTS: In this paper, we propose the truncated ℓ1-norm penalized SCCA to improve the performance and effectiveness of the ℓ1-norm based SCCA methods. Besides, we propose an efficient optimization algorithms to solve this novel SCCA problem. The proposed method is an adaptive shrinkage method via tuning τ. It can avoid the time intensive parameter tuning if given a reasonable small τ. Furthermore, we extend it to the truncated group-lasso (TGL), and propose TGL-SCCA model to improve the group-lasso-based SCCA methods. The experimental results, compared with four benchmark methods, show that our SCCA methods identify better or similar correlation coefficients, and better canonical loading profiles than the competing methods. This demonstrates the effectiveness and efficiency of our methods in discovering interesting imaging genetic associations. AVAILABILITY AND IMPLEMENTATION: The Matlab code and sample data are freely available at http://www.iu.edu/∼shenlab/tools/tlpscca/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 9 |
| 2018 | Quantitative trait loci identification for brain endophenotypes via new additive model with random networksabstractMotivation: The identification of quantitative trait loci (QTL) is critical to the study of causal relationships between genetic variations and disease abnormalities. We focus on identifying the QTLs associated to the brain endophenotypes in imaging genomics study for Alzheimer's Disease (AD). Existing research works mainly depict the association between single nucleotide polymorphisms (SNPs) and the brain endophenotypes via the linear methods, which may introduce high bias due to the simplicity of the models. Since the influence of QTLs on brain endophenotypes is quite complex, it is desired to design the appropriate non-linear models to investigate the associations of genotypes and endophenotypes. Results: In this paper, we propose a new additive model to learn the non-linear associations between SNPs and brain endophenotypes in Alzheimer's disease. Our model can be flexibly employed to explain the non-linear influence of QTLs, thus is more adaptive for the complex distribution of the high-throughput biological data. Meanwhile, as an important computational learning theory contribution, we provide the generalization error analysis for the proposed approach. Unlike most previous theoretical analysis under independent and identically distributed samples assumption, our error bound is based on m-dependent observations, which is more appropriate for the high-throughput and noisy biological data. Experiments on the data from Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort demonstrate the promising performance of our approach for identifying biological meaningful SNPs. Availability and implementation: An executable is available at https://github.com/littleq1991/additive_FNNRW. Xiaoqian Wang 0001, Hong Chen 0004, Kwangsik Nho, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001, Heng Huang 0001 |
Bioinform. | 6 |
| 2017 | Network-based genome wide study of hippocampal imaging phenotype in Alzheimer's Disease to identify functional interaction modulesabstractIdentification of functional modules from biological network is a promising approach to enhance the statistical power of genome-wide association study (GWAS) and improve biological interpretation for complex diseases. The precise functions of genes are highly relevant to tissue context, while a majority of module identification studies are based on tissue-free biological networks that lacks phenotypic specificity. In this study, we propose a module identification method that maps the GWAS results of an imaging phenotype onto the corresponding tissue-specific functional interaction network by applying a machine learning framework. Ridge regression and support vector machine (SVM) models are constructed to re-prioritize GWAS results, followed by exploring hippocampus-relevant modules based on top predictions using GWAS top findings. We also propose a GWAS top-neighbor-based module identification approach and compare it with Ridge and SVM based approaches. Modules conserving both tissue specificity and GWAS discoveries are identified, showing the promise of the proposal method for providing insight into the mechanism of complex diseases. Xiaohui Yao, Shannon L. Risacher, Jason H. Moore, Andrew J. Saykin, Li Shen 0001 |
ICASSP | 5 |
| 2017 | Unified representation of tractography and diffusion-weighted MRI data using sparse multidimensional arraysabstractRecently, linear formulations and convex optimization methods have been proposed to predict diffusion-weighted Magnetic Resonance Imaging (dMRI) data given estimates of brain connections generated using tractography algorithms. The size of the linear models comprising such methods grows with both dMRI data and connectome resolution, and can become very large when applied to modern data. In this paper, we introduce a method to encode dMRI signals and large connectomes, i.e., those that range from hundreds of thousands to millions of fascicles (bundles of neuronal axons), by using a sparse tensor decomposition. We show that this tensor decomposition accurately approximates the Linear Fascicle Evaluation (LiFE) model, one of the recently developed linear models. We provide a theoretical analysis of the accuracy of the sparse decomposed model, LiFESD, and demonstrate that it can reduce the size of the model significantly. Also, we develop algorithms to implement the optimisation solver using the tensor representation in an efficient way. Cesar F. Caiafa, Olaf Sporns, Andrew J. Saykin, Franco Pestilli |
NIPS | 3 |
| 2017 | Longitudinal Genotype-Phenotype Association Study via Temporal Structure Auto-learning Predictive Model
Xiaoqian Wang 0001, Xiaohui Yao, Sungeun Kim, Kwangsik Nho, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001, Heng Huang 0001 |
RECOMB | 7 |
| 2017 | Identification of associations between genotypes and longitudinal phenotypes via temporally-constrained group sparse canonical correlation analysisabstractMOTIVATION: Neuroimaging genetics identifies the relationships between genetic variants (i.e., the single nucleotide polymorphisms) and brain imaging data to reveal the associations from genotypes to phenotypes. So far, most existing machine-learning approaches are widely used to detect the effective associations between genetic variants and brain imaging data at one time-point. However, those associations are based on static phenotypes and ignore the temporal dynamics of the phenotypical changes. The phenotypes across multiple time-points may exhibit temporal patterns that can be used to facilitate the understanding of the degenerative process. In this article, we propose a novel temporally constrained group sparse canonical correlation analysis (TGSCCA) framework to identify genetic associations with longitudinal phenotypic markers. RESULTS: The proposed TGSCCA method is able to capture the temporal changes in brain from longitudinal phenotypes by incorporating the fused penalty, which requires that the differences between two consecutive canonical weight vectors from adjacent time-points should be small. A new efficient optimization algorithm is designed to solve the objective function. Furthermore, we demonstrate the effectiveness of our algorithm on both synthetic and real data (i.e., the Alzheimer's Disease Neuroimaging Initiative cohort, including progressive mild cognitive impairment, stable MCI and Normal Control participants). In comparison with conventional SCCA, our proposed method can achieve strong associations and discover phenotypic biomarkers across multiple time-points to guide disease-progressive interpretation. AVAILABILITY AND IMPLEMENTATION: The Matlab code is available at https://sourceforge.net/projects/ibrain-cn/files/ . CONTACT: [email protected] or [email protected]. Xiaoke Hao, Chanxiu Li, Xiaohui Yao, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001, Daoqiang Zhang |
Bioinform. | 6 |
| 2017 | Tissue-specific network-based genome wide study of amygdala imaging phenotypes to identify functional interaction modulesabstractMOTIVATION: Network-based genome-wide association studies (GWAS) aim to identify functional modules from biological networks that are enriched by top GWAS findings. Although gene functions are relevant to tissue context, most existing methods analyze tissue-free networks without reflecting phenotypic specificity. RESULTS: We propose a novel module identification framework for imaging genetic studies using the tissue-specific functional interaction network. Our method includes three steps: (i) re-prioritize imaging GWAS findings by applying machine learning methods to incorporate network topological information and enhance the connectivity among top genes; (ii) detect densely connected modules based on interactions among top re-prioritized genes; and (iii) identify phenotype-relevant modules enriched by top GWAS findings. We demonstrate our method on the GWAS of [18F]FDG-PET measures in the amygdala region using the imaging genetic data from the Alzheimer's Disease Neuroimaging Initiative, and map the GWAS results onto the amygdala-specific functional interaction network. The proposed network-based GWAS method can effectively detect densely connected modules enriched by top GWAS findings. Tissue-specific functional network can provide precise context to help explore the collective effects of genes with biologically meaningful interactions specific to the studied phenotype. AVAILABILITY AND IMPLEMENTATION: The R code and sample data are freely available at http://www.iu.edu/shenlab/tools/gwasmodule/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaohui Yao, Kefei Liu 0001, Sungeun Kim, Kwangsik Nho, Shannon L. Risacher, Casey S. Greene, Jason H. Moore, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 9 |
| 2016 | Sparse Canonical Correlation Analysis via truncated ℓ1-norm with application to brain imaging geneticsabstractDiscovering bi-multivariate associations between genetic markers and neuroimaging quantitative traits is a major task in brain imaging genetics. Sparse Canonical Correlation Analysis (SCCA) is a popular technique in this area for its powerful capability in identifying bi-multivariate relationships coupled with feature selection. The existing SCCA methods impose either the ℓ1-norm or its variants. The ℓ0-norm is more desirable, which however remains unexplored since the ℓ0-norm minimization is NP-hard. In this paper, we impose the truncated ℓ1-norm to improve the performance of the ℓ1-norm based SCCA methods. Besides, we propose two efficient optimization algorithms and prove their convergence. The experimental results, compared with two benchmark methods, show that our method identifies better and meaningful canonical loading patterns in both simulated and real imaging genetic analyse. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
BIBM | 8 |
| 2016 | New Probabilistic Multi-graph Decomposition Model to Identify Consistent Human Brain Network ModulesabstractMany recent scientific efforts have been devoted to constructing the human connectome using Diffusion Tensor Imaging (DTI) data for understanding large-scale brain networks that underlie higher-level cognition in human. However, suitable network analysis computational tools are still lacking in human brain connectivity research. To address this problem, we propose a novel probabilistic multi-graph decomposition model to identify consistent network modules from the brain connectivity networks of the studied subjects. At first, we propose a new probabilistic graph decomposition model to address the high computational complexity issue in existing stochastic block models. After that, we further extend our new probabilistic graph decomposition model for multiple networks/graphs to identify the shared modules cross multiple brain networks by simultaneously incorporating multiple networks and predicting the hidden block state variables. We also derive an efficient optimization algorithm to solve the proposed objective and estimate the model parameters. We validate our method by analyzing both the weighted fiber connectivity networks constructed from DTI images and the standard human face image clustering benchmark data sets. The promising empirical results demonstrate the superior performance of our proposed method. Dijun Luo, Zhouyuan Huo, Andrew J. Saykin, Li Shen 0001, Heng Huang 0001 |
ICDM | 4 |
| 2016 | Structured sparse canonical correlation analysis for brain imaging genetics: an improved GraphNet methodabstractMOTIVATION: Structured sparse canonical correlation analysis (SCCA) models have been used to identify imaging genetic associations. These models either use group lasso or graph-guided fused lasso to conduct feature selection and feature grouping simultaneously. The group lasso based methods require prior knowledge to define the groups, which limits the capability when prior knowledge is incomplete or unavailable. The graph-guided methods overcome this drawback by using the sample correlation to define the constraint. However, they are sensitive to the sign of the sample correlation, which could introduce undesirable bias if the sign is wrongly estimated. RESULTS: We introduce a novel SCCA model with a new penalty, and develop an efficient optimization algorithm. Our method has a strong upper bound for the grouping effect for both positively and negatively correlated features. We show that our method performs better than or equally to three competing SCCA models on both synthetic and real data. In particular, our method identifies stronger canonical correlations and better canonical loading patterns, showing its promise for revealing interesting imaging genetic associations. AVAILABILITY AND IMPLEMENTATION: The Matlab code and sample data are freely available at http://www.iu.edu/∼shenlab/tools/angscca/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Du 0001, Heng Huang 0001, Sungeun Kim, Shannon L. Risacher, Mark Inlow, Jason H. Moore, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 8 |
| 2015 | Identifying Connectome Module Patterns via New Balanced Multi-graph Normalized Cut
Hongchang Gao, Chengtao Cai, Lin Yan 0003, Joaquín Goñi, Feiping Nie 0001, John D. West, Andrew J. Saykin, Li Shen 0001, Heng Huang 0001 |
MICCAI (2) | 9 |
| 2014 | A Novel Structure-Aware Sparse Learning Algorithm for Brain Imaging Genetics
Lei Du 0001, Sungeun Kim, Shannon L. Risacher, Heng Huang 0001, Mark Inlow, Jason H. Moore, Andrew J. Saykin, Li Shen 0001 |
MICCAI (3) | 8 |
| 2014 | Human Connectome Module Pattern Detection Using a New Multi-graph MinMax Cut Model
De Wang, Feiping Nie 0001, Tom Weidong Cai, Andrew J. Saykin, Li Shen 0001, Heng Huang 0001 |
MICCAI (3) | 6 |
| 2014 | Transcriptome-guided amyloid imaging genetic analysis via a novel structured sparse learning algorithmabstractMOTIVATION: Imaging genetics is an emerging field that studies the influence of genetic variation on brain structure and function. The major task is to examine the association between genetic markers such as single-nucleotide polymorphisms (SNPs) and quantitative traits (QTs) extracted from neuroimaging data. The complexity of these datasets has presented critical bioinformatics challenges that require new enabling tools. Sparse canonical correlation analysis (SCCA) is a bi-multivariate technique used in imaging genetics to identify complex multi-SNP-multi-QT associations. However, most of the existing SCCA algorithms are designed using the soft thresholding method, which assumes that the input features are independent from one another. This assumption clearly does not hold for the imaging genetic data. In this article, we propose a new knowledge-guided SCCA algorithm (KG-SCCA) to overcome this limitation as well as improve learning results by incorporating valuable prior knowledge. RESULTS: The proposed KG-SCCA method is able to model two types of prior knowledge: one as a group structure (e.g. linkage disequilibrium blocks among SNPs) and the other as a network structure (e.g. gene co-expression network among brain regions). The new model incorporates these prior structures by introducing new regularization terms to encourage weight similarity between grouped or connected features. A new algorithm is designed to solve the KG-SCCA model without imposing the independence constraint on the input features. We demonstrate the effectiveness of our algorithm with both synthetic and real data. For real data, using an Alzheimer's disease (AD) cohort, we examine the imaging genetic associations between all SNPs in the APOE gene (i.e. top AD gene) and amyloid deposition measures among cortical regions (i.e. a major AD hallmark). In comparison with a widely used SCCA implementation, our KG-SCCA algorithm produces not only improved cross-validation performances but also biologically meaningful results. AVAILABILITY: Software is freely available on request. Lei Du 0001, Sungeun Kim, Shannon L. Risacher, Heng Huang 0001, Jason H. Moore, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 7 |
| 2014 | Identifying the Neuroanatomical Basis of Cognitive Impairment in Alzheimer's Disease by Correlation- and Nonlinearity-Aware Sparse Bayesian LearningabstractPredicting cognitive performance of subjects from their magnetic resonance imaging (MRI) measures and identifying relevant imaging biomarkers are important research topics in the study of Alzheimer's disease. Traditionally, this task is performed by formulating a linear regression problem. Recently, it is found that using a linear sparse regression model can achieve better prediction accuracy. However, most existing studies only focus on the exploitation of sparsity of regression coefficients, ignoring useful structure information in regression coefficients. Also, these linear sparse models may not capture more complicated and possibly nonlinear relationships between cognitive performance and MRI measures. Motivated by these observations, in this work we build a sparse multivariate regression model for this task and propose an empirical sparse Bayesian learning algorithm. Different from existing sparse algorithms, the proposed algorithm models the response as a nonlinear function of the predictors by extending the predictor matrix with block structures. Further, it exploits not only inter-vector correlation among regression coefficient vectors, but also intra-block correlation in each regression coefficient vector. Experiments on the Alzheimer's Disease Neuroimaging Initiative database showed that the proposed algorithm not only achieved better prediction performance than state-of-the-art competitive methods, but also effectively identified biologically meaningful patterns. Zhilin Zhang 0002, Bhaskar D. Rao, Shiaofen Fang, Andrew J. Saykin, Li Shen 0001 |
IEEE Trans. Medical Imaging | 6 |
| 2013 | Genetic Clustering on the Hippocampal Surface for Genome-Wide Association Studies
Derrek P. Hibar, Sarah E. Medland, Jason L. Stein, Sungeun Kim, Li Shen 0001, Andrew J. Saykin, Greig I. de Zubicaray, Katie L. McMahon, Grant W. Montgomery, Nicholas G. Martin, Margaret J. Wright, Srdjan Djurovic, Ingrid Agartz, Ole A. Andreassen, Paul M. Thompson |
MICCAI (2) | 6 |
| 2013 | A New Sparse Simplex Model for Brain Anatomical and Genetic Network Analysis
Heng Huang 0001, Feiping Nie 0001, Tom Weidong Cai, Andrew J. Saykin, Li Shen 0001 |
MICCAI (2) | 6 |
| 2012 | Sparse Bayesian multi-task learning for predicting cognitive outcomes from neuroimaging measures in Alzheimer's diseaseabstractAlzheimer’s disease (AD) is the most common form of de-mentia that causes progressive impairment of memory and other cognitive functions. Multivariate regression models have been studied in AD for revealing relationships between neuroimaging measures and cognitive scores to understand how structural changes in brain can influence cognitive sta-tus. Existing regression methods, however, do not explic-itly model dependence relation among multiple scores de-rived from a single cognitive test. It has been found that such dependence can deteriorate the performance of these methods. To overcome this limitation, we propose an effi-cient sparse Bayesian multi-task learning algorithm, which adaptively learns and exploits the dependence to achieve improved prediction performance. The proposed algorithm is applied to a real world neuroimaging study in AD to pre-dict cognitive performance using MRI scans. The effective-ness of the proposed algorithm is demonstrated by its supe-rior prediction performance over multiple state-of-the-art competing methods and accurate identification of compact sets of cognition-relevant imaging biomarkers that are con-sistent with prior knowledge. 1. Zhilin Zhang 0002, Taiyong Li, Bhaskar D. Rao, Shiaofen Fang, Sungeun Kim, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001 |
CVPR | 9 |
| 2012 | High-Order Multi-Task Feature Learning to Identify Longitudinal Phenotypic Markers for Alzheimer's Disease Progression PredictionabstractAlzheimer disease (AD) is a neurodegenerative disorder characterized by progressive impairment of memory and other cognitive functions. Regression analysis has been studied to relate neuroimaging measures to cognitive status. However, whether these measures have further predictive power to infer a trajectory of cognitive performance over time is still an under-explored but important topic in AD research. We propose a novel high-order multi-task learning model to address this issue. The proposed model explores the temporal correlations existing in data features and regression tasks by the structured sparsity-inducing norms. In addition, the sparsity of the model enables the selection of a small number of MRI measures while maintaining high prediction accuracy. The empirical studies, using the baseline MRI and serial cognitive data of the ADNI cohort, have yielded promising results. Hua Wang 0007, Feiping Nie 0001, Heng Huang 0001, Sungeun Kim, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001 |
NIPS | 7 |
| 2012 | Identifying quantitative trait loci via group-sparse multitask regression and feature selection: an imaging genetics study of the ADNI cohortabstractMOTIVATION: Recent advances in high-throughput genotyping and brain imaging techniques enable new approaches to study the influence of genetic variation on brain structures and functions. Traditional association studies typically employ independent and pairwise univariate analysis, which treats single nucleotide polymorphisms (SNPs) and quantitative traits (QTs) as isolated units and ignores important underlying interacting relationships between the units. New methods are proposed here to overcome this limitation. RESULTS: Taking into account the interlinked structure within and between SNPs and imaging QTs, we propose a novel Group-Sparse Multi-task Regression and Feature Selection (G-SMuRFS) method to identify quantitative trait loci for multiple disease-relevant QTs and apply it to a study in mild cognitive impairment and Alzheimer's disease. Built upon regression analysis, our model uses a new form of regularization, group ℓ(2,1)-norm (G(2,1)-norm), to incorporate the biological group structures among SNPs induced from their genetic arrangement. The new G(2,1)-norm considers the regression coefficients of all the SNPs in each group with respect to all the QTs together and enforces sparsity at the group level. In addition, an ℓ(2,1)-norm regularization is utilized to couple feature selection across multiple tasks to make use of the shared underlying mechanism among different brain regions. The effectiveness of the proposed method is demonstrated by both clearly improved prediction performance in empirical evaluations and a compact set of selected SNP predictors relevant to the imaging QTs. AVAILABILITY: Software is publicly available at: http://ranger.uta.edu/%7eheng/imaging-genetics/. Hua Wang 0007, Feiping Nie 0001, Heng Huang 0001, Sungeun Kim, Kwangsik Nho, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 7 |
| 2012 | Identifying disease sensitive and quantitative trait-relevant biomarkers from multidimensional heterogeneous imaging genetics data via sparse multimodal multitask learningabstractMOTIVATION: Recent advances in brain imaging and high-throughput genotyping techniques enable new approaches to study the influence of genetic and anatomical variations on brain functions and disorders. Traditional association studies typically perform independent and pairwise analysis among neuroimaging measures, cognitive scores and disease status, and ignore the important underlying interacting relationships between these units. RESULTS: To overcome this limitation, in this article, we propose a new sparse multimodal multitask learning method to reveal complex relationships from gene to brain to symptom. Our main contributions are three-fold: (i) introducing combined structured sparsity regularizations into multimodal multitask learning to integrate multidimensional heterogeneous imaging genetics data and identify multimodal biomarkers; (ii) utilizing a joint classification and regression learning model to identify disease-sensitive and cognition-relevant biomarkers; (iii) deriving a new efficient optimization algorithm to solve our non-smooth objective function and providing rigorous theoretical analysis on the global optimum convergency. Using the imaging genetics data from the Alzheimer's Disease Neuroimaging Initiative database, the effectiveness of the proposed method is demonstrated by clearly improved performance on predicting both cognitive scores and disease status. The identified multimodal biomarkers could predict not only disease status but also cognitive function to help elucidate the biological pathway from gene to brain structure and function, and to cognition and disease. AVAILABILITY: Software is publicly available at: http://ranger.uta.edu/%7eheng/multimodal/. Hua Wang 0007, Feiping Nie 0001, Heng Huang 0001, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 5 |
| 2012 | From phenotype to genotype: an association study of longitudinal phenotypic markers to Alzheimer's disease relevant SNPsabstractMOTIVATION: Imaging genetic studies typically focus on identifying single-nucleotide polymorphism (SNP) markers associated with imaging phenotypes. Few studies perform regression of SNP values on phenotypic measures for examining how the SNP values change when phenotypic measures are varied. This alternative approach may have a potential to help us discover important imaging genetic associations from a different perspective. In addition, the imaging markers are often measured over time, and this longitudinal profile may provide increased power for differentiating genotype groups. How to identify the longitudinal phenotypic markers associated to disease sensitive SNPs is an important and challenging research topic. RESULTS: Taking into account the temporal structure of the longitudinal imaging data and the interrelatedness among the SNPs, we propose a novel 'task-correlated longitudinal sparse regression' model to study the association between the phenotypic imaging markers and the genotypes encoded by SNPs. In our new association model, we extend the widely used ℓ(2,1)-norm for matrices to tensors to jointly select imaging markers that have common effects across all the regression tasks and time points, and meanwhile impose the trace-norm regularization onto the unfolded coefficient tensor to achieve low rank such that the interrelationship among SNPs can be addressed. The effectiveness of our method is demonstrated by both clearly improved prediction performance in empirical evaluations and a compact set of selected imaging predictors relevant to disease sensitive SNPs. AVAILABILITY: Software is publicly available at: http://ranger.uta.edu/%7eheng/Longitudinal/ CONTACT: [email protected] or [email protected]. Hua Wang 0007, Feiping Nie 0001, Heng Huang 0001, Sungeun Kim, Kwangsik Nho, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 8 |
| 2011 | Sparse multi-task regression and feature selection to identify brain imaging predictors for memory performanceabstractAlzheimer's disease (AD) is a neurodegenerative disorder characterized by progressive impairment of memory and other cognitive functions, which makes regression analysis a suitable model to study whether neuroimaging measures can help predict memory performance and track the progression of AD. Existing memory performance prediction methods via regression, however, do not take into account either the interconnected structures within imaging data or those among memory scores, which inevitably restricts their predictive capabilities. To bridge this gap, we propose a novel Sparse Multi-tAsk Regression and feaTure selection (SMART) method to jointly analyze all the imaging and clinical data under a single regression framework and with shared underlying sparse representations. Two convex regularizations are combined and used in the model to enable sparsity as well as facilitate multi-task learning. The effectiveness of the proposed method is demonstrated by both clearly improved prediction performances in all empirical test cases and a compact set of selected RAVLT-relevant MRI predictors that accord with prior studies. Hua Wang 0007, Feiping Nie 0001, Heng Huang 0001, Shannon L. Risacher, Chris Ding, Andrew J. Saykin, Li Shen 0001 |
ICCV | 6 |
| 2011 | Hippocampal Surface Mapping of Genetic Risk Factors in AD via Sparse Learning Models
Sungeun Kim, Mark Inlow, Kwangsik Nho, Shanker Swaminathan, Shannon L. Risacher, Shiaofen Fang, Michael Weiner 0001, Mirza Faisal Beg, Lei Wang 0032, Andrew J. Saykin, Li Shen 0001 |
MICCAI (2) | 11 |
| 2011 | Identifying AD-Sensitive and Cognition-Relevant Imaging Biomarkers via Joint Classification and Regression
Hua Wang 0007, Feiping Nie 0001, Heng Huang 0001, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001 |
MICCAI (3) | 5 |
| 2010 | Sparse Bayesian Learning for Identifying Imaging Biomarkers in AD Prediction
Li Shen 0001, Yuan Qi 0001, Sungeun Kim, Kwangsik Nho, Shannon L. Risacher, Andrew J. Saykin |
MICCAI (3) | 7 |
| 2009 | Data synthesis and tool development for exploring imaging genomic patternsabstractRecent advances in brain imaging and high throughput genotyping techniques enable new approaches to study the influence of genetic variation on brain structure and function. However, major computational challenges are bottlenecks for comprehensive joint analysis of these high-dimensional image and genomic data. We report our initial progress in developing an imaging genomic browsing system for integrated exploration of neuroimaging and genomic data. We describe a method for synthesizing a set of realistic neuroimaging and genomic data, where the relationships between imaging phenotypes and genotypes are known. This data set is used to demonstrate the functionality of our system, which is designed for effectively exploring the neuroanatomical distribution of statistical results that measure the associations between brain imaging phenotypes and genotypes on a genome-wide scale. The proposed system has substantial potential for enabling discovery of important imaging genomic associations through visual evaluation and can be extended towards several directions. Sungeun Kim, Li Shen 0001, Andrew J. Saykin, John D. West |
CIBCB | 3 |
| 2009 | Fourier method for large-scale surface modeling and registration
Li Shen 0001, Sungeun Kim, Andrew J. Saykin |
Comput. Graph. | 3 |
| 2007 | Morphometric Analysis of Hippocampal Shape in Mild Cognitive Impairment: An Imaging Genetics StudyabstractA computational framework is presented for surface based morphometry to localize shape changes between groups of 3D objects. It employs the spherical harmonic (SPHARM) method for surface modeling and random field theory (RFT) for statistical inference. Several new components are introduced to overcome previous limitations: (1) a general linear model is used to facilitate controlling for covariates; (2) a new SPHARM registration method SHREC is proposed to better align SPHARM models; and (3) an estimated smoothness is used in RFT-based analysis to obtain more accurate results. This framework is applied in a mild cognitive impairment (MCI) study to examine hippocampal shape changes related to diagnostic and genetic conditions. Several interesting findings from our analyses suggest combining imaging phenotypes and genetic profiles has the potential to elucidate biological pathways for better understanding MCI and Alzheimer's disease. Li Shen 0001, Andrew J. Saykin, Moo K. Chung, Heng Huang 0001 |
BIBE | 2 |
| 2007 | A Novel Surface Registration Algorithm With Biomedical Modeling ApplicationsabstractIn this paper, we propose a novel surface matching algorithm for arbitrarily shaped but simply connected 3-D objects. The spherical harmonic (SPHARM) method is used to describe these 3-D objects, and a novel surface registration approach is presented. The proposed technique is applied to various applications of medical image analysis. The results are compared with those using the traditional method, in which the first-order ellipsoid is used for establishing surface correspondence and aligning objects. In these applications, our surface alignment method is demonstrated to be more accurate and flexible than the traditional approach. This is due in large part to the fact that a new surface parameterization is generated by a shortcut that employs a useful rotational property of spherical harmonic basis functions for a fast implementation. In order to achieve a suitable computational speed for practical applications, we propose a fast alignment algorithm that improves computational complexity of the new surface registration method from O(n3) to O(n2). Heng Huang 0001, Li Shen 0001, Rong Zhang 0013, Fillia Makedon, Andrew J. Saykin, Justin D. Pearlman |
IEEE Trans. Inf. Technol. Biomed. | 5 |
| 2004 | Extraction of Discriminative Functional MRI Activation Patterns and an Application to Alzheimer's Disease
Despina Kontos, Vasileios Megalooikonomou, Dragoljub Pokrajac, Aleksandar Lazarevic, Zoran Obradovic, Orest B. Boyko, James Ford, Fillia Makedon, Andrew J. Saykin |
MICCAI (2) | 9 |
| 2004 | A surface-based approach for classification of 3D neuroanatomic structures
Li Shen 0001, James Ford, Fillia Makedon, Andrew J. Saykin |
Intell. Data Anal. | 4 |
| 2003 | Patient Classification of fMRI Activation Maps
James Ford, Hany Farid, Fillia Makedon, Laura A. Flashman, Thomas W. McAllister, Vasileios Megalooikonomou, Andrew J. Saykin |
MICCAI (2) | 7 |
| 2003 | Morphometric Analysis of Brain Structures for Improved Discrimination
Li Shen 0001, James Ford, Fillia Makedon, Tilmann Steinberg, Andrew J. Saykin |
MICCAI (2) | 7 |
| 2003 | Quantifying Evolving Processes in Multimodal 3D Medical Images
Tilmann Steinberg, Fillia Makedon, James Ford, Heather Wishart, Andrew J. Saykin |
MICCAI (2) | 6 |
| 2003 | A Spatio-temporal Multi-modal Data Management and Analysis Environment for Tracking MS LesionsabstractWe describe the development of a system that automates data collection, metadata extraction and analysis of spatio-temporal multi-modal data, combining data management and data analysis to provide an efficient resource for clinicians. Though the system is extensible to many applications, the current focus is on managing Multiple Sclerosis (MS) lesion data, which are disparate streams of image, numeric, and text data. In order to discover patterns of MS pathology and plan early and effective treatment, multispectral magnetic resonance (MR) image streams collected over time need to be correlated efficiently with each other and with patient performance and clinical data streams. Tilmann Steinberg, Fillia Makedon, Li Shen 0001, Andrew J. Saykin, Heather Wishart |
SSDBM | 5 |
| 1998 | Stimulus Tracking in Functional Magnetic Resonance Imaging (fMRI)abstractmsmcT Functional hia~etic Resonance tiagery (~~) is a new' medical ima=tig technology pro~ti~mg tictional, as opposed to natomicd, mapptig of the human brain.The promise of this non-invasive technique has opened new' ch~enges for mdtimedia researchers.A the fieId of human functional neurotia=tig e~lodes and the technology of ~W advmces, computational techniques for expertient control, subject stimu~ generation, and joint stimti-activation tiaetig are Iachg.host computational work h= focused on impro~tig the analysis of the brab S=S.Tti paper describes eomputiationa~mechanisms for the dehery md tiactig of multimedia stimti and a correlation framework for stimti and brain responses.NIediaStim, a new integrated data co~ection fraework for stimuIus tracking $hat k currently under development. NlediaSfi etiends the stimti paradi=m to composite multimedia presentations. To stiulate a real-tie sitiation as reafitica~y as possible, it is tiportant to stidy the brainPerrnksion lo ma}:edigi~alor hard copies of fl or part of thk ~!;ork for personal or classroom use k granted ;I:ithout fee pro~,ided ~ha~copies are not made or distributed for profit or commertizl ada~antage, and fiat copies bear this no:i~and ~heftil titation on the firs!page.To copy oihen!.k~lo repub~i~to post on senrers or 10 redistribute 10 lis~, requires prior specific perrnksion anifor a fee. James Ford, Fillia Makedon, Charles B. Owen, Sterling C. Johnson, Andrew J. Saykin |
ACM Multimedia | 5 |