VLDB 2026 Research / reviewers in the wild / expert
Li Shen 0001
dblp:s/LiShen
· DBLP profile ↗
120ranked-venue papers
12as first author
49since 2021 · last 2026
0000-0002-5443-0503ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 93 · 6 first-author · 37 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 22 · 3 first-author · 12 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-author · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SDoH-GPT: using large language models to extract social determinants of healthabstractOBJECTIVE: Extracting social determinants of health (SDoHs) from medical notes depends heavily on labor-intensive annotations, which are typically task-specific, hampering reusability and limiting sharing. Here, we introduce SDoH-GPT, a novel framework leveraging few-shot learning large language models (LLMs) to automate the extraction of SDoH from unstructured text, aiming to improve both efficiency and generalizability. MATERIALS AND METHODS: SDoH-GPT is a framework including the few-shot learning LLM methods to extract the SDoH from medical notes and the XGBoost classifiers which continue to classify SDoH using the annotations generated by the few-shot learning LLM methods as training datasets. The unique combination of the few-shot learning LLM methods with XGBoost utilizes the strength of LLMs as great few shot learners and the efficiency of XGBoost when the training dataset is sufficient. Therefore, SDoH-GPT can extract SDoH without relying on extensive medical annotations or costly human intervention. RESULTS: Our approach achieved tenfold and twentyfold reductions in time and cost, respectively, and superior consistency with human annotators measured by Cohen's kappa of up to 0.92. The innovative combination of LLM and XGBoost can ensure high accuracy and computational efficiency while consistently maintaining 0.90+ AUROC scores. DISCUSSION: This study has verified SDoH-GPT on three datasets and highlights the potential of leveraging LLM and XGBoost to revolutionize medical note classification, demonstrating its capability to achieve highly accurate classifications with significantly reduced time and cost. CONCLUSION: The key contribution of this study is the integration of LLM with XGBoost, which enables cost-effective and high quality annotations of SDoH. This research sets the stage for SDoH can be more accessible, scalable, and impactful in driving future healthcare solutions. Bernardo Scapini Consoli, Xizhi Wu, Song Wang 0026, Yanshan Wang, Justin F. Rousseau, Thomas Hartvigsen, Li Shen 0001, Huanmei Wu, Yifan Peng 0002, Qi Long, Tianlong Chen 0001, Ying Ding 0001 |
J. Am. Medical Informatics Assoc. | 9 |
| 2025 | A Comprehensive Benchmark of Tool-Augmented Large Language Models for Biomedical Knowledge Retrieval and IntegrationabstractGeneral-purpose large language models (LLMs) often struggle in specialized biomedical applications due to their limited access to up-to-date, structured knowledge and domainspecific tools. Although recent studies show proof of life for AI agents in increasingly complex scientific tasks, few studies offer insights at an atomic-level. We present the first large-scale evaluation of LLM tool-calling capabilities, focused on genomics annotation tasks such as variant-to-position and variant-to-gene mapping. Our study benchmarks over one hundred LLMs, via OpenRouter's metagateway, using a standardized tool-calling protocol to retrieve information directly from biomedical APIs (e.g., NCBI dbSNP, Entrez Gene). Our experimental results show that models equipped with structured tool access significantly outperform prompt-only baselines in accuracy, factual consistency, and verifiability. These findings demonstrate the necessity of tool augmentation for reliable biomedical reasoning and provide practical insights for building and testing LLM-based agents across diverse biomedical workflows. Van Q. Truong, Shu Yang 0009, Li Shen 0001, Marylyn D. Ritchie |
BIBM | 3 |
| 2025 | IRIS: Interpretable Risk Clustering Intelligence for Survival AnalysisabstractSurvival analysis models have evolved significantly with deep learning approaches, yet often lack interpretability and meaningful risk stratification capabilities. We present Interpretable Risk Clustering Intelligence for Survival Analysis (IRIS), a novel framework that addresses the critical task of risk clustering while enhancing both input-level and model-body interpretability. Unlike traditional survival models that perform post-hoc risk clustering, IRIS learns to cluster patients into meaningful risk groups directly from data while providing transparent feature importance estimation through feature contribution functions. We validate IRIS on several benchmark datasets, a real-world Alzheimer's disease dataset, and an electronic health record dataset, showing superior performance in risk clustering and predictive reliability with only a modest decrease in time-to-event prediction accuracy compared to state-of-the-art methods. Our results show that IRIS successfully balances the trade-off between interpretability and prediction performance in risk-based survival analysis, offering clinicians actionable insights for treatment planning and resource allocation. Kazi Noshin, Bojian Hou, Mary Regina Boland, Zixuan Wen, Boning Tong, Li Shen 0001, Aidong Zhang 0001 |
IEEE Big Data | 6 |
| 2025 | Fine-Tuning Attention Modules Only: Enhancing Weight Disentanglement in Task ArithmeticabstractIn recent years, *task arithmetic* has garnered increasing attention. This approach edits pre-trained models directly in weight space by combining the fine-tuned weights of various tasks into a *unified model*. Its efficiency and cost-effectiveness stem from its training-free combination, contrasting with traditional methods that require model training on large datasets for multiple tasks. However, applying such a unified model to individual tasks can lead to interference from other tasks (lack of *weight disentanglement*). To address this issue, Neural Tangent Kernel (NTK) linearization has been employed to leverage a ''kernel behavior'', facilitating weight disentanglement and mitigating adverse effects from unrelated tasks. Despite its benefits, NTK linearization presents drawbacks, including doubled training costs, as well as reduced performance of individual models. To tackle this problem, we propose a simple yet effective and efficient method that is to finetune the attention modules only in the Transformer. Our study reveals that the attention modules exhibit kernel behavior, and fine-tuning the attention modules only significantly improves weight disentanglement. To further understand how our method improves the weight disentanglement of task arithmetic, we present a comprehensive study of task arithmetic by differentiating the role of the representation module and task-specific module. In particular, we find that the representation module plays an important role in improving weight disentanglement whereas the task-specific modules such as the classification heads can degenerate the weight disentanglement performance. Ruochen Jin, Bojian Hou, Jiancong Xiao, Weijie J. Su, Li Shen 0001 |
ICLR | 5 |
| 2025 | Restoring Calibration for Aligned Large Language Models: A Calibration-Aware Fine-Tuning ApproachabstractOne of the key technologies for the success of Large Language Models (LLMs) is preference alignment. However, a notable side effect of preference alignment is poor calibration: while the pre-trained models are typically well-calibrated, LLMs tend to become poorly calibrated after alignment with human preferences. In this paper, we investigate why preference alignment affects calibration and how to address this issue. For the first question, we observe that the preference collapse issue in alignment undesirably generalizes to the calibration scenario, causing LLMs to exhibit overconfidence and poor calibration. To address this, we demonstrate the importance of fine-tuning with domain-specific knowledge to alleviate the overconfidence issue. To further analyze whether this affects the model’s performance, we categorize models into two regimes: calibratable and non-calibratable, defined by bounds of Expected Calibration Error (ECE). In the calibratable regime, we propose a calibration-aware fine-tuning approach to achieve proper calibration without compromising LLMs’ performance. However, as models are further fine-tuned for better performance, they enter the non-calibratable regime. For this case, we develop an EM-algorithm-based ECE regularization for the fine-tuning loss to maintain low calibration error. Extensive experiments validate the effectiveness of the proposed methods. Jiancong Xiao, Bojian Hou, Zhanliang Wang, Ruochen Jin, Qi Long, Weijie J. Su, Li Shen 0001 |
ICML | 7 |
| 2025 | MentalChat16K: A Benchmark Dataset for Conversational Mental Health AssistanceabstractWe introduce MentalChat16K, an English benchmark dataset combining a synthetic mental health counseling dataset and a dataset of anonymized transcripts from interventions between Behavioral Health Coaches and Caregivers of patients in palliative or hospice care. Covering a diverse range of conditions like depression, anxiety, and grief, this curated dataset is designed to facilitate the development and evaluation of large language models for conversational mental health assistance. By providing a high-quality resource tailored to this critical domain, MentalChat16K aims to advance research on empathetic, personalized AI solutions to improve access to mental health support services. The dataset prioritizes patient privacy, ethical considerations, and responsible data usage. MentalChat16K presents a valuable opportunity for the research community to innovate AI technologies that can positively impact mental well-being. The dataset is available at https://huggingface.co/datasets/ShenLab/MentalChat16K and the code and documentation are hosted on GitHub at https://github.com/PennShenLab/MentalChat16K. Tianyi Wei, Bojian Hou, Patryk Orzechowski, Shu Yang 0009, Ruochen Jin, Rachael Paulbeck, Joost B. Wagenaar, George Demiris, Li Shen 0001 |
KDD (2) | 10 |
| 2025 | On the Empirical Power of Goodness-of-Fit Tests in Watermark DetectionabstractLarge language models (LLMs) raise concerns about content authenticity and integrity because they can generate human-like text at scale. Text watermarks, which embed detectable statistical signals into generated text, offer a provable way to verify content origin. Many detection methods rely on pivotal statistics that are i.i.d. under human-written text, making goodness-of-fit (GoF) tests a natural tool for watermark detection. However, GoF tests remain largely underexplored in this setting.
In this paper, we systematically evaluate eight GoF tests across three popular watermarking schemes, using three open-source LLMs, two datasets, various generation temperatures, and multiple post-editing methods.
We find that general GoF tests can improve both the detection power and robustness of watermark detectors. Notably, we observe that text repetition, common in low-temperature settings, gives GoF tests a unique advantage not exploited by existing methods.
Our results highlight that classic GoF tests are a simple yet powerful and underused tool for watermark detection in LLMs. Weiqing He, Tianqi Shang, Li Shen 0001, Weijie J. Su, Qi Long |
NeurIPS | 4 |
| 2025 | Stochastic Regret Guarantees for Online Zeroth- and First-Order Bilevel OptimizationabstractOnline bilevel optimization (OBO) is a powerful framework for machine learning problems where both outer and inner objectives evolve over time, requiring dynamic updates. Current OBO approaches rely on deterministic \textit{window-smoothed} regret minimization, which may not accurately reflect system performance when functions change rapidly. In this work, we introduce a novel search direction and show that both first- and zeroth-order (ZO) stochastic OBO algorithms leveraging this direction achieve sublinear {stochastic bilevel regret without window smoothing}. Beyond these guarantees, our framework enhances efficiency by: (i) reducing oracle dependence in hypergradient estimation, (ii) updating inner and outer variables alongside the linear system solution, and (iii) employing ZO-based estimation of Hessians, Jacobians, and gradients. Experiments on online parametric loss tuning and black-box adversarial attacks validate our approach. Parvin Nazari, Bojian Hou, D. Ataee Tarzanagh, Li Shen 0001, George Michailidis |
NeurIPS | 4 |
| 2025 | QOT: Quantized Optimal Transport for sample-level distance matrix in single-cell omicsabstractSingle-cell technologies have enabled the high-dimensional characterization of cell populations at an unprecedented scale. The innate complexity and increasing volume of data pose significant computational and analytical challenges, especially in comparative studies delineating cellular architectures across various biological conditions (i.e. generation of sample-level distance matrices). Optimal Transport is a mathematical tool that captures the intrinsic structure of data geometrically and has been applied to many bioinformatics tasks. In this paper, we propose QOT (Quantized Optimal Transport), a new method enabling efficient computation of sample-level distance matrix from large-scale single-cell omics data through a quantization step. We apply our algorithm to real-world single-cell genomics and pathomics datasets, aiming to extrapolate cell-level insights to inform sample-level categorizations. Our empirical study shows that QOT outperforms existing two OT-based algorithms in accuracy and robustness when obtaining a distance matrix from high throughput single-cell measures at the sample level. Moreover, the sample level distance matrix could be used in the downstream analysis (i.e. uncover the trajectory of disease progression), highlighting its usage in biomedical informatics and data science. Zexuan Wang, Qipeng Zhan, Shu Yang 0009, Shizhuo Mu, Sumita Garai, Patryk Orzechowski, Joost B. Wagenaar, Li Shen 0001 |
Briefings Bioinform. | 9 |
| 2025 | Predicting explainable dementia types with LLM-aided feature engineeringabstractMOTIVATION: The integration of Machine Learning and Artificial Intelligence (AI) into healthcare has immense potential due to the rapidly growing volume of clinical data. However, existing AI models, particularly Large Language Models (LLMs) like GPT-4, face significant challenges in terms of explainability and reliability, particularly in high-stakes domains like healthcare. RESULTS: This paper proposes a novel LLM-aided feature engineering approach that enhances interpretability by extracting clinically relevant features from the Oxford Textbook of Medicine. By converting clinical notes into concept vector representations and employing a linear classifier, our method achieved an accuracy of 0.72, outperforming a traditional n-gram Logistic Regression baseline (0.64) and the GPT-4 baseline (0.48), while focusing on high-level clinical features. We also explore using Text Embeddings to reduce the overall time and cost of our approach by 97%. AVAILABILITY AND IMPLEMENTATION: All code relevant to this paper is available at: https://github.com/AdityaKashyap423/Dementia_LLM_Feature_Engineering/tree/main. Aditya Kashyap, Delip Rao, Mary Regina Boland, Li Shen 0001, Chris Callison-Burch |
Bioinform. | 4 |
| 2025 | Establishing group-level brain structural connectivity incorporating anatomical knowledge under latent space modeling
Selena Wang, Frederick H. Xu, Li Shen 0001, Yize Zhao |
Medical Image Anal. | 4 |
| 2025 | MG-TCCA: Tensor Canonical Correlation Analysis Across Multiple GroupsabstractTensor Canonical Correlation Analysis (TCCA) is a commonly employed statistical method utilized to examine linear associations between two sets of tensor datasets. However, the existing TCCA models fail to adequately address the heterogeneity present in real-world tensor data, such as brain imaging data collected from diverse groups characterized by factors like sex and race. Consequently, these models may yield biased outcomes. In order to surmount this constraint, we propose a novel approach called Multi-Group TCCA (MG-TCCA), which enables the joint analysis of multiple subgroups. By incorporating a dual sparsity structure and a block coordinate ascent algorithm, our MG-TCCA method effectively addresses heterogeneity and leverages information across different groups to identify consistent signals. This novel approach facilitates the quantification of shared and individual structures, reduces data dimensionality, and enables visual exploration. To empirically validate our approach, we conduct a study focused on investigating correlations between two brain positron emission tomography (PET) modalities (AV-45 and FDG) within an Alzheimer's disease (AD) cohort. Our results demonstrate that MG-TCCA surpasses traditional TCCA and Sparse TCCA (STCCA) in identifying sex-specific cross-modality imaging correlations. This heightened performance of MG-TCCA provides valuable insights for the characterization of multimodal imaging biomarkers in AD. Zhuoping Zhou, Boning Tong, D. Ataee Tarzanagh, Bojian Hou, Andrew J. Saykin, Qi Long, Li Shen 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 7 |
| 2025 | Multi-Modal Diagnosis of Alzheimer's Disease Using Interpretable Graph Convolutional NetworksabstractThe interconnection between brain regions in neurological disease encodes vital information for the advancement of biomarkers and diagnostics. Although graph convolutional networks are widely applied for discovering brain connection patterns that point to disease conditions, the potential of connection patterns that arise from multiple imaging modalities has yet to be fully realized. In this paper, we propose a multi-modal sparse interpretable GCN framework (SGCN) for the detection of Alzheimer's disease (AD) and its prodromal stage, known as mild cognitive impairment (MCI). In our experimentation, SGCN learned the sparse regional importance probability to find signature regions of interest (ROIs), and the connective importance probability to reveal disease-specific brain network connections. We evaluated SGCN on the Alzheimer's Disease Neuroimaging Initiative database with multi-modal brain images and demonstrated that the ROI features learned by SGCN were effective for enhancing AD status identification. The identified abnormalities were significantly correlated with AD-related clinical symptoms. We further interpreted the identified brain dysfunctions at the level of large-scale neural systems and sex-related connectivity abnormalities in AD/MCI. The salient ROIs and the prominent brain connectivity abnormalities interpreted by SGCN are considerably important for developing novel biomarkers. These findings contribute to a better understanding of the network-based disorder via multi-modal diagnosis and offer the potential for precision diagnostics. The source code is available at https://github.com/Houliang-Zhou/SGCN. Houliang Zhou, Lifang He 0001, Brian Y. Chen, Li Shen 0001, Yu Zhang 0009 |
IEEE Trans. Medical Imaging | 4 |
| 2025 | Large-Scale Neural Network Quantum States Calculation for Quantum Chemistry on a New Sunway Supercomputer
Yangjun Wu, Li Shen 0001, Hong Qian, Honghui Shang |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2024 | Dude: Dual Distribution-Aware Context Prompt Learning For Large Vision-Language Model
Duy M. H. Nguyen, An T. Le 0001, Trung Quoc Nguyen, Nghiem Tuong Diep, Tai Nguyen 0008, Duy Duong-Tran, Jan Peters 0001, Li Shen 0001, Mathias Niepert, Daniel Sonntag |
ACML | 8 |
| 2024 | Integrative Analysis of Amyloid Imaging and Genetics Reveals Subtypes of Alzheimer Progression in Early Stage
Neel Sangani, Ruiming Wu, Pradeep Varathan, Alice Patania, Shannon L. Risacher, Kwangsik Nho, Liana G. Apostolova, Andrew J. Saykin, Li Shen 0001 |
AIME (2) | 10 |
| 2024 | Online Bilevel Optimization: Regret Analysis of Online Alternating Gradient MethodsabstractThis paper introduces \textit{online bilevel optimization} in which a sequence of time-varying bilevel problems is revealed one after the other. We extend the known regret bounds for single-level online algorithms to the bilevel setting. Specifically, we provide new notions of \textit{bilevel regret}, develop an online alternating time-averaged gradient method that is capable of leveraging smoothness, and give regret bounds in terms of the path-length of the inner and outer minimizer sequences. D. Ataee Tarzanagh, Parvin Nazari, Bojian Hou, Li Shen 0001, Laura Balzano |
AISTATS | 4 |
| 2024 | SEFD: Semantic-Enhanced Framework for Detecting LLM-Generated Textabstracttechniques that often evade existing detection methods. To address this challenge, we present a novel semantic-enhanced framework for detecting LLM-generated text (SEFD) that leverages a retrieval-based mechanism to fully utilize text semantics. Our framework improves upon existing detection methods by systematically integrating retrieval-based techniques with traditional detectors, employing a carefully curated retrieval mechanism that strikes a balance between comprehensive coverage and computational efficiency. We showcase the effectiveness of our approach in sequential text scenarios common in real-world applications, such as online forums and Q&A platforms. Through comprehensive experiments across various LLM-generated texts and detection methods, we demonstrate that our framework substantially enhances detection accuracy in paraphrasing scenarios while maintaining robustness for standard LLM-generated content. This work contributes significantly to ongoing efforts to safeguard information integrity in an era where AI-generated content is increasingly prevalent. Weiqing He, Bojian Hou, Tianqi Shang, D. Ataee Tarzanagh, Qi Long, Li Shen 0001 |
IEEE Big Data | 6 |
| 2024 | Subject-Adaptive Transfer Learning Using Resting State EEG Signals for Cross-Subject EEG Motor Imagery Classification
Sion An, Myeongkyun Kang, Soopil Kim, Philip Chikontwe, Li Shen 0001, Sanghyun Park 0004 |
MICCAI (11) | 5 |
| 2024 | Volume-Optimal Persistence Homological Scaffolds of Hemodynamic Networks Covary with MEG Theta-Alpha Aperiodic Dynamics
Nghi Nguyen, Enrico Amico, Jingyi Zheng, Huajun Huang, Alan D. Kaplan, Giovanni Petri, Joaquín Goñi, Ralph Kaufmann, Yize Zhao, Duy Duong-Tran, Li Shen 0001 |
MICCAI (3) | 12 |
| 2024 | Fairness-Aware Estimation of Graphical ModelsabstractThis paper examines the issue of fairness in the estimation of graphical models (GMs), particularly Gaussian, Covariance, and Ising models. These models play a vital role in understanding complex relationships in high-dimensional data. However, standard GMs can result in biased outcomes, especially when the underlying data involves sensitive characteristics or protected groups. To address this, we introduce a comprehensive framework designed to reduce bias in the estimation of GMs related to protected attributes. Our approach involves the integration of the pairwise graph disparity error and a tailored loss function into a nonsmooth multi-objective optimization problem, striving to achieve fairness across different sensitive groups while maintaining the effectiveness of the GMs. Experimental evaluations on synthetic and real-world datasets demonstrate that our framework effectively mitigates bias without undermining GMs' performance. Zhuoping Zhou, D. Ataee Tarzanagh, Bojian Hou, Qi Long, Li Shen 0001 |
NeurIPS | 5 |
| 2024 | Interpretable deep clustering survival machines for Alzheimer's disease subtype discovery
Bojian Hou, Zixuan Wen, Jingxuan Bao, Richard Zhang 0001, Boning Tong, Shu Yang 0009, Junhao Wen 0002, Yuhan Cui, Jason H. Moore, Andrew J. Saykin, Heng Huang 0001, Paul M. Thompson, Marylyn D. Ritchie, Christos Davatzikos, Li Shen 0001 |
Medical Image Anal. | 15 |
| 2023 | Causal Effects of Environmental Exposures and Biological Traits on the Difference between Phenotypic and Chronological AgesabstractAging is a physiological process associated with numerous cardiovascular, degenerative and neurological conditions. The current demographic changes in the world, characterized by an increasing average age especially in Europe and North America, are causing an increment in the prevalence of such conditions, leading to a public health problem as more resources are required to manage treatments. Recent research has defined a new score named Phenotypic Age, i.e. an adjusted age that takes into account the current health status considering a series of biomarkers, that can be useful to provide a more precise estimation of the probability of developing aging-related conditions. Prevention of such conditions can be performed more efficiently by studying the mechanisms that lead to an increased phenotypic age rather than attempting to treat a patient that has already a high health risk. In this work, we combine pairwise association techniques and mediation analysis to define a strategy to investigate the inner causal mechanisms that lead from specific environmental exposures to an increasing gap between phenotypic and chronological age, considering the influence of biological variables. Four environmental exposures and 11 biological traits have been identified in the NHANES dataset, and each trait has been tested as a mediation variable for each exposure. Almost all mediations reported significant indirect effects with specific results that provide new insights into the causal mechanisms that lead to a deranged aging rate. Daniele Pala, Yuezhi Xie, Li Shen 0001 |
BIBM | 5 |
| 2023 | Interpretable Graph Convolutional Network for Alzheimer's Disease Diagnosis using Multi-Modal Imaging GeneticsabstractIntegrating brain images and genetic data provides a great opportunity to discover potential biomarkers for neurological disorder diagnosis. However, learning genetic information and brain network dysfunction remains a challenging task. In this paper, we propose an interpretable multi-modal imaging and genetic graph convolution network (GCN) for Alzheimer’s disease diagnosis. Our genetic network uses hierarchical GCN to mimic a gene ontology-based graph of biological processes and learn the information flow in this graph. In parallel, our imaging network uses a sparse interpretable GCN with node and edge importance probabilities to learn the brain network from multi-modal images. After multi-modal fusion, the final representation guided by a cluster-based consistency constraint is used to predict the disease-related clinical measures. We evaluate our method on the Alzheimer’s Disease Neuroimaging Initiative (ADNI) database. Our result shows that our imaging-genetics framework achieves superior prediction performance compared to all state-of-the-art methods. The interpretation demonstrated that the salient SNPs, and salient regions interpreted by important probabilities were significantly correlated with AD-related clinical symptoms, and considerably important for developing novel biomarkers. The code is available at https://github.com/Houliang-Zhou/IG-GCN. Houliang Zhou, Yu Zhang 0009, Lifang He 0001, Li Shen 0001, Brian Y. Chen |
BIBM | 4 |
| 2023 | Integrating Multimodal Contrastive Learning and Cross-Modal Attention for Alzheimer's Disease Prediction in Brain Imaging GeneticsabstractHigh annotation costs serve as a significant hurdle in deploying modern deep learning architectures for clinically relevant medical applications, especially when dealing with the inherent heterogeneity of multimodal data, proving the critical need for innovative algorithms that can effectively utilize unlabeled data. In this paper, we propose a model named MCLCA, which integrates multimodal contrastive learning and cross-modal attention to diagnose Alzheimer’s Disease (AD) and identify biomarkers using both labeled and unlabeled multimodal brain imaging genetics data. Through multimodal contrastive learning, MCLCA can effectively learn representations even in the absence of sufficient labels. By utilizing cross-modal attention blocks, the model captures deep connections between different modalities, providing a more comprehensive view of diagnosis. Our proposed MCLCA model is evaluated using the ADNI database with three imaging modalities (VBM-MRI, FDG-PET, and AV45-PET) and genetic SNP data. The results demonstrate that MCLCA can identify important biomarkers with better prediction accuracy compared to the existing methods. The source code is available at https://github.com/MCLCA. Rong Zhou 0007, Houliang Zhou, Li Shen 0001, Brian Y. Chen, Yu Zhang 0009, Lifang He 0001 |
BIBM | 3 |
| 2023 | Attentive Deep Canonical Correlation Analysis for Diagnosing Alzheimer's Disease Using Multimodal Imaging Genetics
Rong Zhou 0007, Houliang Zhou, Brian Y. Chen, Li Shen 0001, Yu Zhang 0009, Lifang He 0001 |
MICCAI (2) | 4 |
| 2023 | Fair Canonical Correlation AnalysisabstractThis paper investigates fairness and bias in Canonical Correlation Analysis (CCA), a widely used statistical technique for examining the relationship between two sets of variables. We present a framework that alleviates unfairness by minimizing the correlation disparity error associated with protected attributes. Our approach enables CCA to learn global projection matrices from all data points while ensuring that these matrices yield comparable correlation levels to group-specific projection matrices. Experimental evaluation on both synthetic and real-world datasets demonstrates the efficacy of our method in reducing correlation disparity error without compromising CCA accuracy. Zhuoping Zhou, D. Ataee Tarzanagh, Bojian Hou, Boning Tong, Yanbo Feng, Qi Long, Li Shen 0001 |
NeurIPS | 8 |
| 2023 | Fairness-aware class imbalanced learning on multiple subgroupsabstractWe present a novel Bayesian-based optimization framework that addresses the challenge of generalization in overparameterized models when dealing with imbalanced subgroups and limited samples per subgroup. Our proposed tri-level optimization framework utilizes local predictors, which are trained on a small amount of data, as well as a fair and class-balanced predictor at the middle and lower levels. To effectively overcome saddle points for minority classes, our lower-level formulation incorporates sharpness-aware minimization. Meanwhile, at the upper level, the framework dynamically adjusts the loss function based on validation loss, ensuring a close alignment between the global predictor and local predictors. Theoretical analysis demonstrates the framework’s ability to enhance classification and fairness generalization, potentially resulting in improvements in the generalization bound. Empirical results validate the superior performance of our tri-level framework compared to existing state-of-the-art approaches. The source code can be found at \url{https://github.com/PennShenLab/FACIMS}. D. Ataee Tarzanagh, Bojian Hou, Boning Tong, Qi Long, Li Shen 0001 |
UAI | 5 |
| 2023 | Integrative analysis of multi-omics and imaging data with incorporation of biological information via structural Bayesian factor analysisabstractMOTIVATION: With the rapid development of modern technologies, massive data are available for the systematic study of Alzheimer's disease (AD). Though many existing AD studies mainly focus on single-modality omics data, multi-omics datasets can provide a more comprehensive understanding of AD. To bridge this gap, we proposed a novel structural Bayesian factor analysis framework (SBFA) to extract the information shared by multi-omics data through the aggregation of genotyping data, gene expression data, neuroimaging phenotypes and prior biological network knowledge. Our approach can extract common information shared by different modalities and encourage biologically related features to be selected, guiding future AD research in a biologically meaningful way. METHOD: Our SBFA model decomposes the mean parameters of the data into a sparse factor loading matrix and a factor matrix, where the factor matrix represents the common information extracted from multi-omics and imaging data. Our framework is designed to incorporate prior biological network information. Our simulation study demonstrated that our proposed SBFA framework could achieve the best performance compared with the other state-of-the-art factor-analysis-based integrative analysis methods. RESULTS: We apply our proposed SBFA model together with several state-of-the-art factor analysis models to extract the latent common information from genotyping, gene expression and brain imaging data simultaneously from the ADNI biobank database. The latent information is then used to predict the functional activities questionnaire score, an important measurement for diagnosis of AD quantifying subjects' abilities in daily life. Our SBFA model shows the best prediction performance compared with the other factor analysis models. AVAILABILITY: Code are publicly available at https://github.com/JingxuanBao/SBFA. CONTACT: [email protected]. Jingxuan Bao, Changgee Chang, Qiyiwen Zhang, Andrew J. Saykin, Li Shen 0001, Qi Long |
Briefings Bioinform. | 5 |
| 2023 | AD-BERT: Using pre-trained language model to predict the progression from mild cognitive impairment to Alzheimer's disease
Chengsheng Mao, Jie Xu 0012, Luke V. Rasmussen, Yikuan Li, Prakash Adekkanattu, Jennifer A. Pacheco, Borna Bonakdarpour, Robert Vassar, Li Shen 0001, Guoqian Jiang, Fei Wang 0001, Jyotishman Pathak, Yuan Luo 0001 |
J. Biomed. Informatics | 9 |
| 2022 | Mediation Analysis and Mixed-Effects Models for the Identification of Stage-specific Imaging Genetics Patterns in Alzheimer's DiseaseabstractAlzheimer's disease (AD) is one of the most common and severe forms of Senile Dementia. Genome-wide association studies (GWAS) have identified dozens of AD susceptible loci. To better understand potential mechanism-of-action for AD, quantitative brain imaging features have been studied as mediators linking genetic variants to AD outcomes. In this study, Mediation analysis, Chow test and Mixed-effects Models are used to investigate the biological pathways by which genetic variants affect both brain structures/functions and disease diagnosis. We analyzed the imaging and genetics data collected from the Alzheimer's Disease Neuroimaging Initiative (ADNI) project, including a Polygenic Hazard Score (PHS) and 13 imaging quantitative traits (QTs) extracted from the AV45 PET scans quantifying the amyloid deposition in different brain regions of subjects from four separate diagnostic groups. Mediation analysis assessed the mediating effects of image QTs between PHS and diagnosis, whereas Chow test and Linear Mixed-Effects models were used to characterize intra-group differences in the associations between genetic scores and imaging QTs for different disease stages. Results show that promising stage-specific imaging QTs that mediate the genetic effect of the studied PHS on disease status have been identified, providing novel insights into the predictive power of the PHS and the mediating power of amyloid imaging QTs with respect to multiple stages over the AD progression. Daniele Pala, Xia Ning, Do Kyoon Kim, Li Shen 0001 |
BIBM | 5 |
| 2022 | Preference Matrix Guided Sparse Canonical Correlation Analysis for Genetic Study of Quantitative Traits in Alzheimer's DiseaseabstractInvestigating the relationship between genetic variation and phenotypic traits is a key issue in quantitative genetics. Specifically for Alzheimer's disease, the association between genetic markers and quantitative traits remains vague while, once identified, will provide valuable guidance for the study and development of genetic-based treatment approaches. Currently, to analyze the association of two modalities, sparse canonical correlation analysis (SCCA) is commonly used to compute one sparse linear combination of the variable features for each modality, giving a pair of linear combination vectors in total that maximizes the cross-correlation between the analyzed modalities. One drawback of the plain SCCA model is that the existing findings and knowledge cannot be integrated into the model as priors to help extract interesting correlation as well as identify biologically meaningful genetic and phenotypic markers. To bridge this gap, we introduce preference matrix guided SCCA (PM-SCCA) that not only takes priors encoded as a preference matrix but also maintains computational simplicity. A simulation study and a real-data experiment are conducted to investigate the effectiveness of the model. Both experiments demonstrate that the proposed PM-SCCA model can capture not only genotype-phenotype correlation but also relevant features effectively. Jiahang Sha, Jingxuan Bao, Kefei Liu 0001, Shu Yang 0009, Zixuan Wen, Yuhan Cui, Junhao Wen 0002, Christos Davatzikos, Jason H. Moore, Andrew J. Saykin, Qi Long, Li Shen 0001 |
BIBM | 12 |
| 2022 | Using Optimal Transport to Improve Spherical Harmonic Quantification of Complex Biological ShapesabstractThe knowledge of the anatomical shape of both gross and microscopic structures is the key to understanding the effects of disease processes on cellular structure. Geometric morphometric methods, such as Procrustes superimposition, and Spherical Harmonics (SPHARM), have been used to capture the biological shape variation and group differences in morphology. Previous SPHARM-MAT techniques use the CALD algorithm to parameterize the mesh surface. It starts from initial mapping and performs local and global smoothing methods alternately to control the area and length distortions simultaneously. However, this parameterization may not be sufficient in complex morphological cases. To bridge this gap, we propose SPHARM-OT, an enhanced SPHARM surface modeling method using optimal transport (OT) for spherical parameterization. First, the genus 0 3D objects are conformally mapped onto a sphere. Then the optimal transport theory via spherical power diagram is introduced to minimize the area distortion. This new algorithm can effectively reduce the area distortion and lead to a better reconstruction result. We demonstrate the effectiveness of the method by applying it to the human sphenoidal paranasal sinuses. Zexuan Wang, Wenxi Yang, Katharine Ryan, Sumita Garai, Benjamin M. Auerbach, Li Shen 0001 |
BIBM | 6 |
| 2022 | Consistency of Graph Theoretical Measurements of Alzheimer's Disease Fiber Density Connectomes Across Multiple Parcellation ScalesabstractGraph theoretical measures have frequently been used to study disrupted connectivity in Alzheimer's disease human brain connectomes. However, prior studies have noted that differences in graph creation methods are confounding factors that may alter the topological observations found in these measures. In this study, we conduct a novel investigation regarding the effect of parcellation scale on graph theoretical measures computed for fiber density networks derived from diffusion tensor imaging. We computed 4 network-wide graph theoretical measures of average clustering coefficient, transitivity, characteristic path length, and global efficiency, and we tested whether these measures are able to consistently identify group differences among healthy control (HC), mild cognitive impairment (MCI), and AD groups in the Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort across 5 scales of the Lausanne parcellation. We found that the segregative measure of transtivity offered the greatest consistency across scales in distinguishing between healthy and diseased groups, while the other measures were impacted by the selection of scale to varying degrees. Global efficiency was the second most consistent measure that we tested, where the measure could distinguish between HC and MCI in all 5 scales and between HC and AD in 3 out of 5 scales. Characteristic path length was highly sensitive to the variation in scale, corroborating previous findings, and could not identify group differences in many of the scales. Average clustering coefficient was also greatly impacted by scale, as it consistently failed to identify group differences in the higher resolution parcellations. From these results, we conclude that many graph theoretical measures are sensitive to the selection of parcellation scale, and further development in methodology is needed to offer a more robust characterization of AD's relationship with disrupted connectivity. Frederick H. Xu, Sumita Garai, Duy Duong-Tran, Andrew J. Saykin, Yize Zhao, Li Shen 0001 |
BIBM | 6 |
| 2022 | Tensor-Based Multi-Modal Multi-Target Regression for Alzheimer's Disease PredictionabstractThe assessment of Alzheimer’s Disease (AD) progression via the analysis of physical changes within the brain has attracted great interest from the fields of healthcare, computational medicine, and machine learning alike. Recent studies have demonstrated that using both multi-modal data and multiple AD assessment scores in a predictive model can better reflect pathological characteristics and enhance prediction performance. However, using such high-dimensional structure information to model inter-correlation between multiple targets remains a challenging task. In this paper, we propose a Tensor-based Multi-modal Multi-Target Regression (TMMTR) method for AD detection and prediction, which enables simultaneously modeling multilinear structure information as well as intrinsic inter-target correlations in a general learning framework. We also investigate the tensor-structured sparsity that supports the interpretability of our prediction. Experiments conducted on the ADNI dataset validate the superior performance of our method when compared to other state-of-the-art methods. Benjamin Zalatan, Yong Chen 0016, Li Shen 0001, Lifang He 0001 |
BIBM | 4 |
| 2022 | Sparse Interpretation of Graph Convolutional Networks for Multi-modal Diagnosis of Alzheimer's Disease
Houliang Zhou, Yu Zhang 0009, Brian Y. Chen, Li Shen 0001, Lifang He 0001 |
MICCAI (8) | 4 |
| 2022 | Large-Scale Simulation of Quantum Computational Chemistry on a New Sunway SupercomputerabstractQuantum computational chemistry (QCC) is the use of quantum computers to solve problems in computational quantum chemistry. We develop a high performance variational quantum eigensolver (VQE) simulator for simulating quantum computational chemistry problems on a new Sunway supercomputer. The major innovations include: (1) a Matrix Product State (MPS) based VQE simulator to reduce the amount of memory needed and increase the simulation efficiency; (2) a combination of the Density Matrix Embedding Theory with the MPS-based VQE simulator to further extend the simulation range; (3) A three-level parallelization scheme to scale up to 20 million cores; (4) Usage of the Julia script language as the main programming language, which both makes the programming easier and enables cutting edge performance as native C or Fortran; (5) Study of real chemistry systems based on the VQE simulator, achieving nearly linearly strong and weak scaling. Our simulation demonstrates the power of VQE for large quantum chemistry systems, thus paves the way for large-scale VQE experiments on near-term quantum computers. Honghui Shang, Li Shen 0001, Zhiqian Xu 0005, Chu Guo, Jie Liu 0069, Rongfen Lin, Yuling Yang, Zhuoya Wang, Yunquan Zhang |
SC | 2 |
| 2022 | Integrating multi-omics summary data using a Mendelian randomization frameworkabstractMendelian randomization is a versatile tool to identify the possible causal relationship between an omics biomarker and disease outcome using genetic variants as instrumental variables. A key theme is the prioritization of genes whose omics readouts can be used as predictors of the disease outcome through analyzing GWAS and QTL summary data. However, there is a dearth of study of the best practice in probing the effects of multiple -omics biomarkers annotated to the same gene of interest. To bridge this gap, we propose powerful combination tests that integrate multiple correlated $P$-values without assuming the dependence structure between the exposures. Our extensive simulation experiments demonstrate the superiority of our proposed approach compared with existing methods that are adapted to the setting of our interest. The top hits of the analyses of multi-omics Alzheimer's disease datasets include genes ABCA7 and ATP1B1. Chong Jin, Li Shen 0001, Qi Long |
Briefings Bioinform. | 3 |
| 2022 | Deep multiview learning to identify imaging-driven subtypes in mild cognitive impairmentabstractBACKGROUND: In Alzheimer's Diseases (AD) research, multimodal imaging analysis can unveil complementary information from multiple imaging modalities and further our understanding of the disease. One application is to discover disease subtypes using unsupervised clustering. However, existing clustering methods are often applied to input features directly, and could suffer from the curse of dimensionality with high-dimensional multimodal data. The purpose of our study is to identify multimodal imaging-driven subtypes in Mild Cognitive Impairment (MCI) participants using a multiview learning framework based on Deep Generalized Canonical Correlation Analysis (DGCCA), to learn shared latent representation with low dimensions from 3 neuroimaging modalities. RESULTS: DGCCA applies non-linear transformation to input views using neural networks and is able to learn correlated embeddings with low dimensions that capture more variance than its linear counterpart, generalized CCA (GCCA). We designed experiments to compare DGCCA embeddings with single modality features and GCCA embeddings by generating 2 subtypes from each feature set using unsupervised clustering. In our validation studies, we found that amyloid PET imaging has the most discriminative features compared with structural MRI and FDG PET which DGCCA learns from but not GCCA. DGCCA subtypes show differential measures in 5 cognitive assessments, 6 brain volume measures, and conversion to AD patterns. In addition, DGCCA MCI subtypes confirmed AD genetic markers with strong signals that existing late MCI group did not identify. CONCLUSION: Overall, DGCCA is able to learn effective low dimensional embeddings from multimodal data by learning non-linear projections. MCI subtypes generated from DGCCA embeddings are different from existing early and late MCI groups and show most similarity with those identified by amyloid PET features. In our validation studies, DGCCA subtypes show distinct patterns in cognitive measures, brain volumes, and are able to identify AD genetic markers. These findings indicate the promise of the imaging-driven subtypes and their power in revealing disease structures beyond early and late stage MCI. Yixue Feng 0001, Mansu Kim, Xiaohui Yao, Kefei Liu 0001, Qi Long, Li Shen 0001 |
BMC Bioinform. | 6 |
| 2022 | Identifying genes associated with brain volumetric differences through tissue specific transcriptomic inference from GWAS summary dataabstractBACKGROUND: Brain volume has been widely studied in the neuroimaging field, since it is an important and heritable trait associated with brain development, aging and various neurological and psychiatric disorders. Genome-wide association studies (GWAS) have successfully identified numerous associations between genetic variants such as single nucleotide polymorphisms and complex traits like brain volume. However, it is unclear how these genetic variations influence regional gene expression levels, which may subsequently lead to phenotypic changes. S-PrediXcan is a tissue-specific transcriptomic data analysis method that can be applied to bridge this gap. In this work, we perform an S-PrediXcan analysis on GWAS summary data from two large imaging genetics initiatives, the UK Biobank and Enhancing Neuroimaging Genetics through Meta Analysis, to identify tissue-specific transcriptomic effects on two closely related brain volume measures: total brain volume (TBV) and intracranial volume (ICV). RESULTS: As a result of the analysis, we identified 10 genes that are highly associated with both TBV and ICV. Nine out of 10 genes were found to be associated with TBV in another study using a different gene-based association analysis. Moreover, most of our discovered genes were also found to be correlated with multiple cognitive and behavioral traits. Further analyses revealed the protein-protein interactions, associated molecular pathways and biological functions that offer insight into how these genes function and interact with others. CONCLUSIONS: These results confirm that S-PrediXcan can identify genes with tissue-specific transcriptomic effects on complex traits. The analysis also suggested novel genes whose expression levels are related to brain volumetric traits. This provides important insights into the genetic mechanisms of the human brain. Hung Mai, Jingxuan Bao, Paul M. Thompson, Do Kyoon Kim, Li Shen 0001 |
BMC Bioinform. | 5 |
| 2022 | Multi-task learning based structured sparse canonical correlation analysis for brain imaging genetics
Mansu Kim, Eun Jeong Min, Kefei Liu 0001, Andrew J. Saykin, Jason H. Moore, Qi Long, Li Shen 0001 |
Medical Image Anal. | 8 |
| 2022 | Diagnosis of obsessive-compulsive disorder via spatial similarity-aware learning and fused deep polynomial network
Peng Yang 0011, Cheng Zhao 0003, Qiong Yang, Wei Zheng 0009, Xiaohua Xiao, Li Shen 0001, Tianfu Wang 0001, Bai Ying Lei, Ziwen Peng |
Medical Image Anal. | 6 |
| 2021 | Interpretable temporal graph neural network for prognostic prediction of Alzheimer's disease using longitudinal neuroimaging dataabstractAlzheimer's disease (AD) is a progressive neurodegenerative brain disorder characterized by memory loss and cognitive decline. Early detection and accurate prognosis of AD is an important research topic, and numerous machine learning methods have been proposed to solve this problem. However, traditional machine learning models are facing challenges in effectively integrating longitudinal neuroimaging data and biologically meaningful structure and knowledge to build accurate and interpretable prognostic predictors. To bridge this gap, we propose an interpretable graph neural network (GNN) model for AD prognostic prediction based on longitudinal neuroimaging data while embracing the valuable knowledge of structural brain connectivity. In our empirical study, we demonstrate that 1) the proposed model outperforms several competing models (i.e., DNN, SVM) in terms of prognostic prediction accuracy, and 2) our model can capture neuroanatomical contribution to the prognostic predictor and yield biologically meaningful interpretation to facilitate better mechanistic understanding of the Alzheimer's disease. Source code is available at https://github.com/JaesikKim/temporal-GNN. Mansu Kim, Jaesik Kim, Jeffrey Qu, Heng Huang 0001, Qi Long, Kyung-Ah Sohn 0001, Do Kyoon Kim, Li Shen 0001 |
BIBM | 8 |
| 2021 | Brain imaging genetics: integrated analysis and machine learningabstractBrain imaging genetics is an emerging data science field, where integrated analysis of brain imaging and genetics data, often combined with other biomarker, clinical and environmental data, is performed to gain new insights into the genetic, molecular and phenotypic characteristics of the brain as well as their impact on normal and disordered brain function and behavior. Many methodological advances in brain imaging genetics are attributed to large-scale landmark biobank projects such as the Alzheimer’s Disease Sequencing Project, the Alzheimer’s Disease Neuroimaging Initiative, and the UK Biobank. Using the study of Alzheimer’s disease as an example, we will discuss fundamental concepts, state-of-the-art statistical and machine learning methods, and innovative applications in this rapidly evolving field. We show that the wide availability of brain imaging genetics data from various large-scale biobanks, coupled with advances in biomedical statistics, informatics and computing, provides enormous opportunities to contribute significantly to biomedical discoveries in brain science and to impact the development of new diagnostic, therapeutic and preventative approaches for complex brain disorders such as Alzheimer’s disease. Li Shen 0001 |
BIBM | 1 |
| 2021 | Topological Learning and Its Application to Multimodal Brain Network Integration
Tananun Songdechakraiwut, Li Shen 0001, Moo K. Chung |
MICCAI (2) | 2 |
| 2021 | A Novel Bayesian Semi-parametric Model for Learning Heritable Imaging Traits
Yize Zhao, Xiwen Zhao, Mansu Kim, Jingxuan Bao, Li Shen 0001 |
MICCAI (5) | 5 |
| 2021 | A structural enriched functional network: An application to predict brain cognitive performance
Mansu Kim, Jingxuan Bao, Kefei Liu 0001, Bo-yong Park, Hyunjin Park, Jae Young Baik, Li Shen 0001 |
Medical Image Anal. | 7 |
| 2021 | Multi-Task Sparse Canonical Correlation Analysis with Application to Multi-Modal Brain Imaging GeneticsabstractBrain imaging genetics studies the genetic basis of brain structures and functionalities via integrating genotypic data such as single nucleotide polymorphisms (SNPs) and imaging quantitative traits (QTs). In this area, both multi-task learning (MTL) and sparse canonical correlation analysis (SCCA) methods are widely used since they are superior to those independent and pairwise univariate analysis. MTL methods generally incorporate a few of QTs and could not select features from multiple QTs; while SCCA methods typically employ one modality of QTs to study its association with SNPs. Both MTL and SCCA are computational expensive as the number of SNPs increases. In this paper, we propose a novel multi-task SCCA (MTSCCA) method to identify bi-multivariate associations between SNPs and multi-modal imaging QTs. MTSCCA could make use of the complementary information carried by different imaging modalities. MTSCCA enforces sparsity at the group level via the${\mathrm G}_{2,1}$-norm, and jointly selects features across multiple tasks for SNPs and QTs via the$\ell _{2,1}$-norm. A fast optimization algorithm is proposed using the grouping information of SNPs. Compared with conventional SCCA methods, MTSCCA obtains better correlation coefficients and canonical weights patterns. In addition, MTSCCA runs very fast and easy-to-implement, indicating its potential power in genome-wide brain-wide imaging genetics. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Andrew J. Saykin, Lei Guo 0002, Li Shen 0001 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2021 | Identify Consistent Cross-Modality Imaging Genetic Patterns via Discriminant Sparse Canonical Correlation AnalysisabstractSparse canonical correlation analysis (SCCA) is a bi-multivariate technique used in imaging genetics to identify complex multi-SNP-multi-QT associations. However, the traditional SCCA algorithm has been designed to seek a linear correlation between the SNP genotype and brain imaging phenotype, ignoring the discriminant similarity information between within-class subjects in brain imaging genetics association analysis. In addition, multi-modality brain imaging phenotypes are extracted from different perspectives and imaging markers from the same region consistently showing up in multimodalities may provide more insights for the mechanistic understanding of diseases. In this paper, a novel multi-modality discriminant SCCA algorithm (MD-SCCA) is proposed to overcome these limitations as well as to improve learning results by incorporating valuable discriminant similarity information into the SCCA algorithm. Specifically, we first extract the discriminant similarity information between within-class subjects by the sparse representation. Second, the discriminant similarity information is enforced within SCCA to construct a discriminant SCCA algorithm (D-SCCA). At last, the MD-SCCA algorithm is adopted to fully explore the relationships among different modalities of different subjects. In experiments, both synthetic dataset and real data from the Alzheimer's Disease Neuroimaging Initiative database are used to test the performance of our algorithm. The empirical results have demonstrated that the proposed algorithm not only produces improved cross-validation performances but also identifies consistent cross-modality imaging genetic biomarkers. Meiling Wang 0001, Wei Shao 0005, Xiaoke Hao, Li Shen 0001, Daoqiang Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2020 | Polygenic mediation analysis of Alzheimer's disease implicated intermediate amyloid imaging phenotypes
Yingxuan Eng, Xiaohui Yao, Kefei Liu 0001, Shannon L. Risacher, Andrew J. Saykin, Qi Long, Yize Zhao, Li Shen 0001 |
AMIA | 8 |
| 2020 | Estimating Hard-tissue Conditions from Dental Images via Machine LearningabstractDespite the great success of machine learning in various biomedical domains, applications to dental hard tissue conditions (primarily on dental Caries, Erosive Tooth Wear (ETW), and Fluorosis) are under-explored, in particular for analyzing photographic images. The clinical diagnostics of these dental hard-tissue conditions is routinely performed by visual examination but is often limited by its subjectivity. To bridge this gap, we apply four categories of machine learning strategies including nine different methods with two different feature representations to estimate the probability and severity of dental hard-tissue conditions from photographic tooth images. Our first empirical study is performed on the real dataset containing both controls and cases, and the best probability estimation results are achieved by Extra Trees Regression (RMSE: 0.030, Pearson correlation: 0.600) for Caries, Decision Tree (RMSE: 0.183, Pearson correlation: 0.581) for ETW, and Bayesian ARD Regression (RMSE: 0.191, Pearson correlation: 0.745) for Fluorosis. Our second empirical study is performed on the case only datasets, and the best severity estimation results are achieved by Extra Trees Regression (RMSE: 0.029, Pearson correlation: 0.687) for Caries, Bayesian ARD Regression and Linear Regression (RMSE: 0.192, Pearson correlation: 0.490) for ETW, and Bayesian ARD Regression (RMSE: 0.238, Pearson correlation: 0.537) for Fluorosis. These results indicate that machine learning models provide promising opportunities to help clinical evaluation and save resources in the management of these dental conditions. Jingxuan Bao, Mansu Kim, Anderson T. Hara, Gerardo Maupome, Li Shen 0001 |
BIBE | 6 |
| 2020 | Deep Multiview Learning to Identify Population Structure with Multimodal ImagingabstractWe present an effective deep multiview learning framework to identify population structure using multimodal imaging data. Our approach is based on canonical correlation analysis (CCA). We propose to use deep generalized CCA (DGCCA) to learn a shared latent representation of non-linearly mapped and maximally correlated components from multiple imaging modalities with reduced dimensionality. In our empirical study, this representation is shown to effectively capture more variance in original data than conventional generalized CCA (GCCA) which applies only linear transformation to the multi-view data. Furthermore, subsequent cluster analysis on the new feature set learned from DGCCA is able to identify a promising population structure in an Alzheimer's disease (AD) cohort. Genetic association analyses of the clustering results demonstrate that the shared representation learned from DGCCA yields a population structure with a stronger genetic basis than several competing feature learning methods. Yixue Feng 0001, Mansu Kim, Xiaohui Yao, Kefei Liu 0001, Qi Long, Li Shen 0001 |
BIBE | 6 |
| 2020 | Persistent Feature Analysis of Multimodal Brain Networks Using Generalized Fused Lasso for EMCI Identification
Jin Li 0012, Chenyuan Bian, Xianglian Meng, Li Shen 0001 |
MICCAI (7) | 7 |
| 2020 | Spatial Similarity-Aware Learning and Fused Deep Polynomial Network for Detection of Obsessive-Compulsive Disorder
Peng Yang 0011, Qiong Yang, Wei Zheng 0009, Li Shen 0001, Tianfu Wang 0001, Ziwen Peng, Bai Ying Lei |
MICCAI (7) | 4 |
| 2020 | Identifying diagnosis-specific genotype-phenotype associations via joint multitask sparse canonical correlation analysis and classificationabstractMOTIVATION: Brain imaging genetics studies the complex associations between genotypic data such as single nucleotide polymorphisms (SNPs) and imaging quantitative traits (QTs). The neurodegenerative disorders usually exhibit the diversity and heterogeneity, originating from which different diagnostic groups might carry distinct imaging QTs, SNPs and their interactions. Sparse canonical correlation analysis (SCCA) is widely used to identify bi-multivariate genotype-phenotype associations. However, most existing SCCA methods are unsupervised, leading to an inability to identify diagnosis-specific genotype-phenotype associations. RESULTS: In this article, we propose a new joint multitask learning method, named MT-SCCALR, which absorbs the merits of both SCCA and logistic regression. MT-SCCALR learns genotype-phenotype associations of multiple tasks jointly, with each task focusing on identifying one diagnosis-specific genotype-phenotype pattern. Meanwhile, MT-SCCALR cannot only select relevant SNPs and imaging QTs for each diagnostic group alone, but also allows the selection of those shared by multiple diagnostic groups. We derive an efficient optimization algorithm whose convergence to a local optimum is guaranteed. Compared with two state-of-the-art methods, MT-SCCALR yields better or similar canonical correlation coefficients and classification performances. In addition, it owns much better discriminative canonical weight patterns of great interest than competitors. This demonstrates the power and capability of MTSCCAR in identifying diagnostically heterogeneous genotype-phenotype patterns, which would be helpful to understand the pathophysiology of brain disorders. AVAILABILITY AND IMPLEMENTATION: The software is publicly available at https://github.com/dulei323/MTSCCALR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 9 |
| 2020 | Regional imaging genetic enrichment analysisabstractMOTIVATION: Brain imaging genetics aims to reveal genetic effects on brain phenotypes, where most studies examine phenotypes defined on anatomical or functional regions of interest (ROIs) given their biologically meaningful interpretation and modest dimensionality compared with voxelwise approaches. Typical ROI-level measures used in these studies are summary statistics from voxelwise measures in the region, without making full use of individual voxel signals. RESULTS: In this article, we propose a flexible and powerful framework for mining regional imaging genetic associations via voxelwise enrichment analysis, which embraces the collective effect of weak voxel-level signals and integrates brain anatomical annotation information. Our proposed method achieves three goals at the same time: (i) increase the statistical power by substantially reducing the burden of multiple comparison correction; (ii) employ brain annotation information to enable biologically meaningful interpretation and (iii) make full use of fine-grained voxelwise signals. We demonstrate our method on an imaging genetic analysis using data from the Alzheimer's Disease Neuroimaging Initiative, where we assess the collective regional genetic effects of voxelwise FDG-positron emission tomography measures between 116 ROIs and 565 373 single-nucleotide polymorphisms. Compared with traditional ROI-wise and voxelwise approaches, our method identified 2946 novel imaging genetic associations in addition to 33 ones overlapping with the two benchmark methods. In particular, two newly reported variants were further supported by transcriptome evidences from region-specific expression analysis. This demonstrates the promise of the proposed method as a flexible and powerful framework for exploring imaging genetic effects on the brain. AVAILABILITY AND IMPLEMENTATION: The R code and sample data are freely available at https://github.com/lshen/RIGEA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaohui Yao, Shan Cong, Shannon L. Risacher, Andrew J. Saykin, Jason H. Moore, Li Shen 0001 |
Bioinform. | 7 |
| 2020 | Accelerating bioinformatics research with International Conference on Intelligent Biology and Medicine 2020abstractThe International Association for Intelligent Biology and Medicine (IAIBM) is a nonprofit organization that promotes intelligent biology and medical science. It hosts an annual International Conference on Intelligent Biology and Medicine (ICIBM), which was initially established in 2012. Due to the coronavirus (COVID-19) pandemic, the ICIBM 2020 was held for the first time as a virtual online conference on August 9 to 10. The virtual conference had ~ 300 registered participants and featured 41 online real-time presentations. ICIBM 2020 received a total of 75 manuscript submissions, and 12 were selected to be published in this special issue of BMC Bioinformatics. These 12 manuscripts cover a wide range of bioinformatics topics including network analysis, imaging analysis, machine learning, gene expression analysis, and sequence analysis. Li Shen 0001, Xinghua Shi, Kai Wang 0063, Yulin Dai, Zhongming Zhao |
BMC Bioinform. | 2 |
| 2020 | Effect of APOE ε4 on multimodal brain connectomic traits: a persistent homology studyabstractBACKGROUND: Although genetic risk factors and network-level neuroimaging abnormalities have shown effects on cognitive performance and brain atrophy in Alzheimer's disease (AD), little is understood about how apolipoprotein E (APOE) ε4 allele, the best-known genetic risk for AD, affect brain connectivity before the onset of symptomatic AD. This study aims to investigate APOE ε4 effects on brain connectivity from the perspective of multimodal connectome. RESULTS: Here, we propose a novel multimodal brain network modeling framework and a network quantification method based on persistent homology for identifying APOE ε4-related network differences. Specifically, we employ sparse representation to integrate multimodal brain network information derived from both the resting state functional magnetic resonance imaging (rs-fMRI) data and the diffusion-weighted magnetic resonance imaging (dw-MRI) data. Moreover, persistent homology is proposed to avoid the ad hoc selection of a specific regularization parameter and to capture valuable brain connectivity patterns from the topological perspective. The experimental results demonstrate that our method outperforms the competing methods, and reasonably yields connectomic patterns specific to APOE ε4 carriers and non-carriers. CONCLUSIONS: We have proposed a multimodal framework that integrates structural and functional connectivity information for constructing a fused brain network with greater discriminative power. Using persistent homology to extract topological features from the fused brain network, our method can effectively identify APOE ε4-related brain connectomic biomarkers. Jin Li 0012, Chenyuan Bian, Xianglian Meng, Li Shen 0001 |
BMC Bioinform. | 7 |
| 2020 | Enhanced Balanced Min Cut
Xiaojun Chen 0006, Weijun Hong, Feiping Nie 0001, Joshua Zhexue Huang, Li Shen 0001 |
Int. J. Comput. Vis. | 5 |
| 2020 | Detecting genetic associations with brain imaging phenotypes in Alzheimer's disease via a novel structured SCCA approach
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Andrew J. Saykin, Lei Guo 0002, Li Shen 0001 |
Medical Image Anal. | 8 |
| 2020 | Multi-modal neuroimaging feature selection with consistent metric constraint for diagnosis of Alzheimer's disease
Xiaoke Hao, Yongjin Bao, Yingchun Guo, Ming Yu 0006, Daoqiang Zhang, Shannon L. Risacher, Andrew J. Saykin, Xiaohui Yao, Li Shen 0001 |
Medical Image Anal. | 9 |
| 2020 | Brain Imaging Genomics: Integrated Analysis and Machine LearningabstractBrain imaging genomics is an emerging data science field, where integrated analysis of brain imaging and genomics data, often combined with other biomarker, clinical and environmental data, is performed to gain new insights into the phenotypic, genetic and molecular characteristics of the brain as well as their impact on normal and disordered brain function and behavior. It has enormous potential to contribute significantly to biomedical discoveries in brain science. Given the increasingly important role of statistical and machine learning in biomedicine and rapidly growing literature in brain imaging genomics, we provide an up-to-date and comprehensive review of statistical and machine learning methods for brain imaging genomics, as well as a practical discussion on method selection for various biomedical applications. Li Shen 0001, Paul M. Thompson |
Proc. IEEE | 1 |
| 2020 | Joint Multi-Modal Longitudinal Regression and Classification for Alzheimer's Disease PredictionabstractAlzheimer's disease (AD) is a serious neurodegenerative condition that affects millions of individuals across the world. As the average age of individuals in the United States and the world increases, the prevalence of AD will continue to grow. To address this public health problem, the research community has developed computational approaches to sift through various aspects of clinical data and uncover their insights, among which one of the most challenging problem is to determine the biological mechanisms that cause AD to develop. To study this problem, in this paper we present a novel Joint Multi-Modal Longitudinal Regression and Classification method and show how it can be used to identify the cognitive status of the participants in the Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort and the underlying biological mechanisms. By intelligently combining clinical data of various modalities (i.e., genetic information and brain scans) using a variety of regularizations that can identify AD-relevant biomarkers, we perform the regression and classification tasks simultaneously. Because the proposed objective is a non-smooth optimization problem that is difficult to solve in general, we derive an efficient iterative algorithm and rigorously prove its convergence. To validate our new method in predicting the cognitive scores of patients and their clinical diagnosis, we conduct comprehensive experiments on the ADNI cohort. Our promising results demonstrate the benefits and flexibility of the proposed method. We anticipate that our new method is of interest to clinical communities beyond AD research and have open-sourced the code of our method online.11 The code package for the proposed Joint Multi-Modal Longitudinal Regression and Classification model have been made publicly available online at https://github.com/minds-mines/jmmlrc. Lodewijk Brand, Kai Nichols, Hua Wang 0007, Li Shen 0001, Heng Huang 0001 |
IEEE Trans. Medical Imaging | 4 |
| 2020 | Associating Multi-Modal Brain Imaging Phenotypes and Genetic Risk Factors via a Dirty Multi-Task Learning MethodabstractBrain imaging genetics becomes more and more important in brain science, which integrates genetic variations and brain structures or functions to study the genetic basis of brain disorders. The multi-modal imaging data collected by different technologies, measuring the same brain distinctly, might carry complementary information. Unfortunately, we do not know the extent to which the phenotypic variance is shared among multiple imaging modalities, which further might trace back to the complex genetic mechanism. In this paper, we propose a novel dirty multi-task sparse canonical correlation analysis (SCCA) to study imaging genetic problems with multi-modal brain imaging quantitative traits (QTs) involved. The proposed method takes advantages of the multi-task learning and parameter decomposition. It can not only identify the shared imaging QTs and genetic loci across multiple modalities, but also identify the modality-specific imaging QTs and genetic loci, exhibiting a flexible capability of identifying complex multi-SNP-multi-QT associations. Using the state-of-the-art multi-view SCCA and multi-task SCCA, the proposed method shows better or comparable canonical correlation coefficients and canonical weights on both synthetic and real neuroimaging genetic data. In addition, the identified modality-consistent biomarkers, as well as the modality-specific biomarkers, provide meaningful and interesting information, demonstrating the dirty multi-task SCCA could be a powerful alternative method in multi-modal brain imaging genetics. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Andrew J. Saykin, Li Shen 0001 |
IEEE Trans. Medical Imaging | 8 |
| 2019 | A Dirty Multi-task Learning Method for Multi-modal Brain Imaging Genetics
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
MICCAI (4) | 9 |
| 2019 | Improved Prediction of Cognitive Outcomes via Globally Aligned Imaging Biomarker Enrichments over Progressions
Lyujian Lu, Saad El Beleidy, Lauren Zoe Baker, Hua Wang 0007, Heng Huang 0001, Li Shen 0001 |
MICCAI (4) | 6 |
| 2019 | Identifying progressive imaging genetic patterns via multi-task sparse canonical correlation analysis: a longitudinal study of the ADNI cohortabstractMOTIVATION: Identifying the genetic basis of the brain structure, function and disorder by using the imaging quantitative traits (QTs) as endophenotypes is an important task in brain science. Brain QTs often change over time while the disorder progresses and thus understanding how the genetic factors play roles on the progressive brain QT changes is of great importance and meaning. Most existing imaging genetics methods only analyze the baseline neuroimaging data, and thus those longitudinal imaging data across multiple time points containing important disease progression information are omitted. RESULTS: We propose a novel temporal imaging genetic model which performs the multi-task sparse canonical correlation analysis (T-MTSCCA). Our model uses longitudinal neuroimaging data to uncover that how single nucleotide polymorphisms (SNPs) play roles on affecting brain QTs over the time. Incorporating the relationship of the longitudinal imaging data and that within SNPs, T-MTSCCA could identify a trajectory of progressive imaging genetic patterns over the time. We propose an efficient algorithm to solve the problem and show its convergence. We evaluate T-MTSCCA on 408 subjects from the Alzheimer's Disease Neuroimaging Initiative database with longitudinal magnetic resonance imaging data and genetic data available. The experimental results show that T-MTSCCA performs either better than or equally to the state-of-the-art methods. In particular, T-MTSCCA could identify higher canonical correlation coefficients and capture clearer canonical weight patterns. This suggests that T-MTSCCA identifies time-consistent and time-dependent SNPs and imaging QTs, which further help understand the genetic basis of the brain QT changes over the time during the disease progression. AVAILABILITY AND IMPLEMENTATION: The software and simulation data are publicly available at https://github.com/dulei323/TMTSCCA. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Du 0001, Kefei Liu 0001, Lei Zhu 0011, Xiaohui Yao, Shannon L. Risacher, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 8 |
| 2019 | Joint between-sample normalization and differential expression detection through ℓ 0-regularized regressionabstractAbstract Background A fundamental problem in RNA-seq data analysis is to identify genes or exons that are differentially expressed with varying experimental conditions based on the read counts. The relativeness of RNA-seq measurements makes the between-sample normalization of read counts an essential step in differential expression (DE) analysis. In most existing methods, the normalization step is performed prior to the DE analysis. Recently, Jiang and Zhan proposed a statistical method which introduces sample-specific normalization parameters into a joint model, which allows for simultaneous normalization and differential expression analysis from log-transformed RNA-seq data. Furthermore, an ℓ0 penalty is used to yield a sparse solution which selects a subset of DE genes. The experimental conditions are restricted to be categorical in their work. Results In this paper, we generalize Jiang and Zhan’s method to handle experimental conditions that are measured in continuous variables. As a result, genes with expression levels associated with a single or multiple covariates can be detected. As the problem being high-dimensional, non-differentiable and non-convex, we develop an efficient algorithm for model fitting. Conclusions Experiments on synthetic data demonstrate that the proposed method outperforms existing methods in terms of detection accuracy when a large fraction of genes are differentially expressed in an asymmetric manner, and the performance gain becomes more substantial for larger sample sizes. We also apply our method to a real prostate cancer RNA-seq dataset to identify genes associated with pre-operative prostate-specific antigen (PSA) levels in patients. Kefei Liu 0001, Li Shen 0001, Hui Jiang 0002 |
BMC Bioinform. | 2 |
| 2019 | Identifying Candidate Genetic Associations with MRI-Derived AD-Related ROI via Tree-Guided Sparse LearningabstractImaging genetics has attracted significant interests in recent studies. Traditional work has focused on mass-univariate statistical approaches that identify important single nucleotide polymorphisms (SNPs) associated with quantitative traits (QTs) of brain structure or function. More recently, to address the problem of multiple comparison and weak detection, multivariate analysis methods such as the least absolute shrinkage and selection operator (Lasso) are often used to select the most relevant SNPs associated with QTs. However, one problem of Lasso, as well as many other feature selection methods for imaging genetics, is that some useful prior information, e.g., the hierarchical structure among SNPs, are rarely used for designing a more powerful model. In this paper, we propose to identify the associations between candidate genetic features (i.e., SNPs) and magnetic resonance imaging (MRI)-derived measures using a tree-guided sparse learning (TGSL) method. The advantage of our method is that it explicitly models the complex hierarchical structure among the SNPs in the objective function for feature selection. Specifically, motivated by the biological knowledge, the hierarchical structures involving gene groups and linkage disequilibrium (LD) blocks as well as individual SNPs are imposed as a tree-guided regularization term in our TGSL model. Experimental studies on simulation data and the Alzheimer's Disease Neuroimaging Initiative (ADNI) data show that our method not only achieves better predictions than competing methods on the MRI-derived measures of AD-related region of interests (ROIs) (i.e., hippocampus, parahippocampal gyrus, and precuneus), but also identifies sparse SNP patterns at the block level to better guide the biological interpretation. Xiaoke Hao, Xiaohui Yao, Shannon L. Risacher, Andrew J. Saykin, Jintai Yu, Huifu Wang, Lan Tan, Li Shen 0001, Daoqiang Zhang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2019 | A Unified Model for Joint Normalization and Differential Gene Expression Detection in RNA-Seq DataabstractThe RNA-sequencing (RNA-seq) is becoming increasingly popular for quantifying gene expression levels. Since the RNA-seq measurements are relative in nature, between-sample normalization is an essential step in differential expression (DE) analysis. The normalization step of existing DE detection algorithms is usually ad hoc and performed only once prior to DE detection, which may be suboptimal since ideally normalization should be based on non-DE genes only and thus coupled with DE detection. We propose a unified statistical model for joint normalization and DE detection of RNA-seq data. Sample-specific normalization factors are modeled as unknown parameters in the gene-wise linear models and jointly estimated with the regression coefficients. By imposing sparsity-inducing L1 penalty (or mixed L1/L2 penalty for multiple treatment conditions) on the regression coefficients, we formulate the problem as a penalized least-squares regression problem and apply the augmented Lagrangian method to solve it. Simulation and real data studies show that the proposed model and algorithms perform better than or comparably to existing methods in terms of detection power and false-positive rate. The performance gain increases with increasingly larger sample size or higher signal to noise ratio, and is more significant when a large proportion of genes are differentially expressed in an asymmetric manner. Kefei Liu 0001, Jieping Ye, Li Shen 0001, Hui Jiang 0002 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2019 | Mining Directional Drug Interaction Effects on Myopathy Using the FAERS DatabaseabstractMining high-order drug-drug interaction (DDI) induced adverse drug effects from electronic health record databases is an emerging area, and very few studies have explored the relationships between high-order drug combinations. We investigate a novel pharmacovigilance problem for mining directional DDI effects on myopathy using the FDA Adverse Event Reporting System (FAERS) database. Our paper provides information on the risk of myopathy associated with adding new drugs on the already prescribed medication, and visualizes the identified directional DDI patterns as user-friendly graphical representation. We utilize the Apriori algorithm to extract frequent drug combinations from the FAERS database. We use odds ratio to estimate the risk of myopathy associated with directional DDI. We create a tree-structured graph to visualize the findings for easy interpretation. Our method confirmed myopathy association with previously reported HMG-CoA reductase inhibitors like rosuvastatin, fluvastatin, simvastatin, and atorvastatin. New, previously unidentified but mechanistically plausible associations with myopathy were also observed, such as the DDI between pamidronate and levofloxacin. Additional top findings are gadolinium-based imaging agents, which however are often used in myopathy diagnosis. Other DDIs with no obvious mechanism are also reported, such as that of sulfamethoxazole with trimethoprim and potassium chloride. This study shows the feasibility to estimate high-order directional DDIs in a fast and accurate manner. The results of the analysis could become a useful tool in the specialists' hands through an easy-to-understand graphic visualization. Danai Chasioti, Xiaohui Yao, Pengyue Zhang, Samuel Lerner, Sara K. Quinney, Xia Ning, Lang Li 0001, Li Shen 0001 |
IEEE J. Biomed. Health Informatics | 8 |
| 2018 | Fast Multi-Task SCCA Learning with Feature Selection for Multi-Modal Brain Imaging Genetics
Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
BIBM | 8 |
| 2018 | A Unified Model for Robust Differential Expression Analysis of RNA-Seq Data
Kefei Liu 0001, Li Shen 0001, Hui Jian |
BIBM | 2 |
| 2018 | Interactive Machine Learning by Visualization: A Small Data SolutionabstractMachine learning algorithms and traditional data mining process usually require a large volume of data to train the algorithm-specific models, with little or no user feedback during the model building process. Such a "big data" based automatic learning strategy is sometimes unrealistic for applications where data collection or processing is very expensive or difficult, such as in clinical trials. Furthermore, expert knowledge can be very valuable in the model building process in some fields such as biomedical sciences. In this paper, we propose a new visual analytics approach to interactive machine learning and visual data mining. In this approach, multi-dimensional data visualization techniques are employed to facilitate user interactions with the machine learning and mining process. This allows dynamic user feedback in different forms, such as data selection, data labeling, and data correction, to enhance the efficiency of model building. In particular, this approach can significantly reduce the amount of data required for training an accurate model, and therefore can be highly impactful for applications where large amount of data is hard to obtain. The proposed approach is tested on two application problems: the handwriting recognition (classification) problem and the human cognitive score prediction (regression) problem. Both experiments show that visualization supported interactive machine learning and data mining can achieve the same accuracy as an automatic process can with much smaller training data sets. Shiaofen Fang, Snehasis Mukhopadhyay, Andrew J. Saykin, Li Shen 0001 |
IEEE BigData | 5 |
| 2018 | Joint High-Order Multi-Task Feature Learning to Predict the Progression of Alzheimer's Disease
Lodewijk Brand, Hua Wang 0007, Heng Huang 0001, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001 |
MICCAI (1) | 6 |
| 2018 | Network approaches to systems biology analysis of complex disease: integrative methods for multi-omics dataabstractIn the past decade, significant progress has been made in complex disease research across multiple omics layers from genome, transcriptome and proteome to metabolome. There is an increasing awareness of the importance of biological interconnections, and much success has been achieved using systems biology approaches. However, because of the typical focus on one single omics layer at a time, existing systems biology findings explain only a modest portion of complex disease. Recent advances in multi-omics data collection and sharing present us new opportunities for studying complex diseases in a more comprehensive fashion, and yet simultaneously create new challenges considering the unprecedented data dimensionality and diversity. Here, our goal is to review extant and emerging network approaches that can be applied across multiple biological layers to facilitate a more comprehensive and integrative multilayered omics analysis of complex diseases. Shannon L. Risacher, Li Shen 0001, Andrew J. Saykin |
Briefings Bioinform. | 3 |
| 2018 | A novel SCCA approach via truncated ℓ1-norm and truncated group lasso for brain imaging geneticsabstractMOTIVATION: Brain imaging genetics, which studies the linkage between genetic variations and structural or functional measures of the human brain, has become increasingly important in recent years. Discovering the bi-multivariate relationship between genetic markers such as single-nucleotide polymorphisms (SNPs) and neuroimaging quantitative traits (QTs) is one major task in imaging genetics. Sparse Canonical Correlation Analysis (SCCA) has been a popular technique in this area for its powerful capability in identifying bi-multivariate relationships coupled with feature selection. The existing SCCA methods impose either the ℓ1-norm or its variants to induce sparsity. The ℓ0-norm penalty is a perfect sparsity-inducing tool which, however, is an NP-hard problem. RESULTS: In this paper, we propose the truncated ℓ1-norm penalized SCCA to improve the performance and effectiveness of the ℓ1-norm based SCCA methods. Besides, we propose an efficient optimization algorithms to solve this novel SCCA problem. The proposed method is an adaptive shrinkage method via tuning τ. It can avoid the time intensive parameter tuning if given a reasonable small τ. Furthermore, we extend it to the truncated group-lasso (TGL), and propose TGL-SCCA model to improve the group-lasso-based SCCA methods. The experimental results, compared with four benchmark methods, show that our SCCA methods identify better or similar correlation coefficients, and better canonical loading profiles than the competing methods. This demonstrates the effectiveness and efficiency of our methods in discovering interesting imaging genetic associations. AVAILABILITY AND IMPLEMENTATION: The Matlab code and sample data are freely available at http://www.iu.edu/∼shenlab/tools/tlpscca/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Junwei Han 0001, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 10 |
| 2018 | Quantitative trait loci identification for brain endophenotypes via new additive model with random networksabstractMotivation: The identification of quantitative trait loci (QTL) is critical to the study of causal relationships between genetic variations and disease abnormalities. We focus on identifying the QTLs associated to the brain endophenotypes in imaging genomics study for Alzheimer's Disease (AD). Existing research works mainly depict the association between single nucleotide polymorphisms (SNPs) and the brain endophenotypes via the linear methods, which may introduce high bias due to the simplicity of the models. Since the influence of QTLs on brain endophenotypes is quite complex, it is desired to design the appropriate non-linear models to investigate the associations of genotypes and endophenotypes. Results: In this paper, we propose a new additive model to learn the non-linear associations between SNPs and brain endophenotypes in Alzheimer's disease. Our model can be flexibly employed to explain the non-linear influence of QTLs, thus is more adaptive for the complex distribution of the high-throughput biological data. Meanwhile, as an important computational learning theory contribution, we provide the generalization error analysis for the proposed approach. Unlike most previous theoretical analysis under independent and identically distributed samples assumption, our error bound is based on m-dependent observations, which is more appropriate for the high-throughput and noisy biological data. Experiments on the data from Alzheimer's Disease Neuroimaging Initiative (ADNI) cohort demonstrate the promising performance of our approach for identifying biological meaningful SNPs. Availability and implementation: An executable is available at https://github.com/littleq1991/additive_FNNRW. Xiaoqian Wang 0001, Hong Chen 0004, Kwangsik Nho, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001, Heng Huang 0001 |
Bioinform. | 7 |
| 2017 | Network-based genome wide study of hippocampal imaging phenotype in Alzheimer's Disease to identify functional interaction modulesabstractIdentification of functional modules from biological network is a promising approach to enhance the statistical power of genome-wide association study (GWAS) and improve biological interpretation for complex diseases. The precise functions of genes are highly relevant to tissue context, while a majority of module identification studies are based on tissue-free biological networks that lacks phenotypic specificity. In this study, we propose a module identification method that maps the GWAS results of an imaging phenotype onto the corresponding tissue-specific functional interaction network by applying a machine learning framework. Ridge regression and support vector machine (SVM) models are constructed to re-prioritize GWAS results, followed by exploring hippocampus-relevant modules based on top predictions using GWAS top findings. We also propose a GWAS top-neighbor-based module identification approach and compare it with Ridge and SVM based approaches. Modules conserving both tissue specificity and GWAS discoveries are identified, showing the promise of the proposal method for providing insight into the mechanism of complex diseases. Xiaohui Yao, Shannon L. Risacher, Jason H. Moore, Andrew J. Saykin, Li Shen 0001 |
ICASSP | 6 |
| 2017 | Longitudinal Genotype-Phenotype Association Study via Temporal Structure Auto-learning Predictive Model
Xiaoqian Wang 0001, Xiaohui Yao, Sungeun Kim, Kwangsik Nho, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001, Heng Huang 0001 |
RECOMB | 8 |
| 2017 | Identification of associations between genotypes and longitudinal phenotypes via temporally-constrained group sparse canonical correlation analysisabstractMOTIVATION: Neuroimaging genetics identifies the relationships between genetic variants (i.e., the single nucleotide polymorphisms) and brain imaging data to reveal the associations from genotypes to phenotypes. So far, most existing machine-learning approaches are widely used to detect the effective associations between genetic variants and brain imaging data at one time-point. However, those associations are based on static phenotypes and ignore the temporal dynamics of the phenotypical changes. The phenotypes across multiple time-points may exhibit temporal patterns that can be used to facilitate the understanding of the degenerative process. In this article, we propose a novel temporally constrained group sparse canonical correlation analysis (TGSCCA) framework to identify genetic associations with longitudinal phenotypic markers. RESULTS: The proposed TGSCCA method is able to capture the temporal changes in brain from longitudinal phenotypes by incorporating the fused penalty, which requires that the differences between two consecutive canonical weight vectors from adjacent time-points should be small. A new efficient optimization algorithm is designed to solve the objective function. Furthermore, we demonstrate the effectiveness of our algorithm on both synthetic and real data (i.e., the Alzheimer's Disease Neuroimaging Initiative cohort, including progressive mild cognitive impairment, stable MCI and Normal Control participants). In comparison with conventional SCCA, our proposed method can achieve strong associations and discover phenotypic biomarkers across multiple time-points to guide disease-progressive interpretation. AVAILABILITY AND IMPLEMENTATION: The Matlab code is available at https://sourceforge.net/projects/ibrain-cn/files/ . CONTACT: [email protected] or [email protected]. Xiaoke Hao, Chanxiu Li, Xiaohui Yao, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001, Daoqiang Zhang |
Bioinform. | 7 |
| 2017 | Tissue-specific network-based genome wide study of amygdala imaging phenotypes to identify functional interaction modulesabstractMOTIVATION: Network-based genome-wide association studies (GWAS) aim to identify functional modules from biological networks that are enriched by top GWAS findings. Although gene functions are relevant to tissue context, most existing methods analyze tissue-free networks without reflecting phenotypic specificity. RESULTS: We propose a novel module identification framework for imaging genetic studies using the tissue-specific functional interaction network. Our method includes three steps: (i) re-prioritize imaging GWAS findings by applying machine learning methods to incorporate network topological information and enhance the connectivity among top genes; (ii) detect densely connected modules based on interactions among top re-prioritized genes; and (iii) identify phenotype-relevant modules enriched by top GWAS findings. We demonstrate our method on the GWAS of [18F]FDG-PET measures in the amygdala region using the imaging genetic data from the Alzheimer's Disease Neuroimaging Initiative, and map the GWAS results onto the amygdala-specific functional interaction network. The proposed network-based GWAS method can effectively detect densely connected modules enriched by top GWAS findings. Tissue-specific functional network can provide precise context to help explore the collective effects of genes with biologically meaningful interactions specific to the studied phenotype. AVAILABILITY AND IMPLEMENTATION: The R code and sample data are freely available at http://www.iu.edu/shenlab/tools/gwasmodule/. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiaohui Yao, Kefei Liu 0001, Sungeun Kim, Kwangsik Nho, Shannon L. Risacher, Casey S. Greene, Jason H. Moore, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 10 |
| 2016 | Sparse Canonical Correlation Analysis via truncated ℓ1-norm with application to brain imaging geneticsabstractDiscovering bi-multivariate associations between genetic markers and neuroimaging quantitative traits is a major task in brain imaging genetics. Sparse Canonical Correlation Analysis (SCCA) is a popular technique in this area for its powerful capability in identifying bi-multivariate relationships coupled with feature selection. The existing SCCA methods impose either the ℓ1-norm or its variants. The ℓ0-norm is more desirable, which however remains unexplored since the ℓ0-norm minimization is NP-hard. In this paper, we impose the truncated ℓ1-norm to improve the performance of the ℓ1-norm based SCCA methods. Besides, we propose two efficient optimization algorithms and prove their convergence. The experimental results, compared with two benchmark methods, show that our method identifies better and meaningful canonical loading patterns in both simulated and real imaging genetic analyse. Lei Du 0001, Kefei Liu 0001, Xiaohui Yao, Shannon L. Risacher, Lei Guo 0002, Andrew J. Saykin, Li Shen 0001 |
BIBM | 9 |
| 2016 | New Probabilistic Multi-graph Decomposition Model to Identify Consistent Human Brain Network ModulesabstractMany recent scientific efforts have been devoted to constructing the human connectome using Diffusion Tensor Imaging (DTI) data for understanding large-scale brain networks that underlie higher-level cognition in human. However, suitable network analysis computational tools are still lacking in human brain connectivity research. To address this problem, we propose a novel probabilistic multi-graph decomposition model to identify consistent network modules from the brain connectivity networks of the studied subjects. At first, we propose a new probabilistic graph decomposition model to address the high computational complexity issue in existing stochastic block models. After that, we further extend our new probabilistic graph decomposition model for multiple networks/graphs to identify the shared modules cross multiple brain networks by simultaneously incorporating multiple networks and predicting the hidden block state variables. We also derive an efficient optimization algorithm to solve the proposed objective and estimate the model parameters. We validate our method by analyzing both the weighted fiber connectivity networks constructed from DTI images and the standard human face image clustering benchmark data sets. The promising empirical results demonstrate the superior performance of our proposed method. Dijun Luo, Zhouyuan Huo, Andrew J. Saykin, Li Shen 0001, Heng Huang 0001 |
ICDM | 5 |
| 2016 | Structured sparse canonical correlation analysis for brain imaging genetics: an improved GraphNet methodabstractMOTIVATION: Structured sparse canonical correlation analysis (SCCA) models have been used to identify imaging genetic associations. These models either use group lasso or graph-guided fused lasso to conduct feature selection and feature grouping simultaneously. The group lasso based methods require prior knowledge to define the groups, which limits the capability when prior knowledge is incomplete or unavailable. The graph-guided methods overcome this drawback by using the sample correlation to define the constraint. However, they are sensitive to the sign of the sample correlation, which could introduce undesirable bias if the sign is wrongly estimated. RESULTS: We introduce a novel SCCA model with a new penalty, and develop an efficient optimization algorithm. Our method has a strong upper bound for the grouping effect for both positively and negatively correlated features. We show that our method performs better than or equally to three competing SCCA models on both synthetic and real data. In particular, our method identifies stronger canonical correlations and better canonical loading patterns, showing its promise for revealing interesting imaging genetic associations. AVAILABILITY AND IMPLEMENTATION: The Matlab code and sample data are freely available at http://www.iu.edu/∼shenlab/tools/angscca/ CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lei Du 0001, Heng Huang 0001, Sungeun Kim, Shannon L. Risacher, Mark Inlow, Jason H. Moore, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 9 |
| 2015 | Identifying Connectome Module Patterns via New Balanced Multi-graph Normalized Cut
Hongchang Gao, Chengtao Cai, Lin Yan 0003, Joaquín Goñi, Feiping Nie 0001, John D. West, Andrew J. Saykin, Li Shen 0001, Heng Huang 0001 |
MICCAI (2) | 10 |
| 2015 | Erratum to: Improving protein order-disorder classification using charge-hydropathy plotsabstractDuring the production of our manuscript [1], an incorrect Figure Figure11 was put in place of the one we submitted, and we erred in not appropriately informing the production staff of this mistake. The corrected figure and the text describing this figure, as well as the figure legend, are present herein.
Figure 1.
Charge-Hydropathy plots. In (A) the IDP-Hydropathy scale was used, in (B) the Guy (1985) Hydropathy scale was used, and in (C) the Kyte-Doolittle (1981) hydropathy scale was used. Red circles indicate disordered proteins, blue circles indicate structured ...
As stated in our manuscript [1], “The C-H plots generated using scale SVM parameters scale, Kyte-Doolittle hydropathy scale, and Guy hydropathy scale for whole protein prediction are shown in Figure Figure1.1. Figure Figure1A,1A, which is derived by SVM parameters scale, shows many fewer misclassified disordered proteins on the ordered side, compared to Figure Figure1B1B and and1C1C.”
As further stated in reference [1]: “Figure 1. Charge-Hydropathy plots. In (A) the IDP-Hydropathy scale was used, in (B) the Guy (1985) Hydropathy scale was used, and in (C) the Kyte-Doolittle (1981) hydropathy scale was used. Red circles indicate disordered proteins, blue circles indicate structured proteins. For these plots, each scale was normalized to be in the interval of 0 to 1. The Guy’s scale is multiplied by -1 prior to normalization to conform to the energy rule set by Kyte-Doolittle scale. In (A) the function describing the boundary is: = 3.31 -0.97. In (B) the function describing the boundary is: = 2.32 -0.93. In (C), the function describing the boundary is = 1.35 -0.49.”
For more details, the reader is referred to the published manuscript [1]. Christopher J. Oldfield, Wei-Lun Hsu, Jingwei Meng, Li Shen 0001, Pedro Romero, Vladimir N. Uversky, A. Keith Dunker |
BMC Bioinform. | 7 |
| 2014 | A Novel Structure-Aware Sparse Learning Algorithm for Brain Imaging Genetics
Lei Du 0001, Sungeun Kim, Shannon L. Risacher, Heng Huang 0001, Mark Inlow, Jason H. Moore, Andrew J. Saykin, Li Shen 0001 |
MICCAI (3) | 9 |
| 2014 | Human Connectome Module Pattern Detection Using a New Multi-graph MinMax Cut Model
De Wang, Feiping Nie 0001, Tom Weidong Cai, Andrew J. Saykin, Li Shen 0001, Heng Huang 0001 |
MICCAI (3) | 7 |
| 2014 | Transcriptome-guided amyloid imaging genetic analysis via a novel structured sparse learning algorithmabstractMOTIVATION: Imaging genetics is an emerging field that studies the influence of genetic variation on brain structure and function. The major task is to examine the association between genetic markers such as single-nucleotide polymorphisms (SNPs) and quantitative traits (QTs) extracted from neuroimaging data. The complexity of these datasets has presented critical bioinformatics challenges that require new enabling tools. Sparse canonical correlation analysis (SCCA) is a bi-multivariate technique used in imaging genetics to identify complex multi-SNP-multi-QT associations. However, most of the existing SCCA algorithms are designed using the soft thresholding method, which assumes that the input features are independent from one another. This assumption clearly does not hold for the imaging genetic data. In this article, we propose a new knowledge-guided SCCA algorithm (KG-SCCA) to overcome this limitation as well as improve learning results by incorporating valuable prior knowledge. RESULTS: The proposed KG-SCCA method is able to model two types of prior knowledge: one as a group structure (e.g. linkage disequilibrium blocks among SNPs) and the other as a network structure (e.g. gene co-expression network among brain regions). The new model incorporates these prior structures by introducing new regularization terms to encourage weight similarity between grouped or connected features. A new algorithm is designed to solve the KG-SCCA model without imposing the independence constraint on the input features. We demonstrate the effectiveness of our algorithm with both synthetic and real data. For real data, using an Alzheimer's disease (AD) cohort, we examine the imaging genetic associations between all SNPs in the APOE gene (i.e. top AD gene) and amyloid deposition measures among cortical regions (i.e. a major AD hallmark). In comparison with a widely used SCCA implementation, our KG-SCCA algorithm produces not only improved cross-validation performances but also biologically meaningful results. AVAILABILITY: Software is freely available on request. Lei Du 0001, Sungeun Kim, Shannon L. Risacher, Heng Huang 0001, Jason H. Moore, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 8 |
| 2014 | Improving protein order-disorder classification using charge-hydropathy plotsabstractBACKGROUND: The earliest whole protein order/disorder predictor (Uversky et al., Proteins, 41: 415-427 (2000)), herein called the charge-hydropathy (C-H) plot, was originally developed using the Kyte-Doolittle (1982) hydropathy scale (Kyte & Doolittle., J. Mol. Biol, 157: 105-132(1982)). Here the goal is to determine whether the performance of the C-H plot in separating structured and disordered proteins can be improved by using an alternative hydropathy scale. RESULTS: Using the performance of the CH-plot as the metric, we compared 19 alternative hydropathy scales, with the finding that the Guy (1985) hydropathy scale (Guy, Biophys. J, 47:61-70(1985)) was the best of the tested hydropathy scales for separating large collections structured proteins and intrinsically disordered proteins (IDPs) on the C-H plot. Next, we developed a new scale, named IDP-Hydropathy, which further improves the discrimination between structured proteins and IDPs. Applying the C-H plot to a dataset containing 109 IDPs and 563 non-homologous fully structured proteins, the Kyte-Doolittle (1982) hydropathy scale, the Guy (1985) hydropathy scale, and the IDP-Hydropathy scale gave balanced two-state classification accuracies of 79%, 84%, and 90%, respectively, indicating a very substantial overall improvement is obtained by using different hydropathy scales. A correlation study shows that IDP-Hydropathy is strongly correlated with other hydropathy scales, thus suggesting that IDP-Hydropathy probably has only minor contributions from amino acid properties other than hydropathy. CONCLUSION: We suggest that IDP-Hydropathy would likely be the best scale to use for any type of algorithm developed to predict protein disorder. Christopher J. Oldfield, Wei-Lun Hsu, Jingwei Meng, Li Shen 0001, Pedro Romero, Vladimir N. Uversky, A. Keith Dunker |
BMC Bioinform. | 7 |
| 2014 | Identifying the Neuroanatomical Basis of Cognitive Impairment in Alzheimer's Disease by Correlation- and Nonlinearity-Aware Sparse Bayesian LearningabstractPredicting cognitive performance of subjects from their magnetic resonance imaging (MRI) measures and identifying relevant imaging biomarkers are important research topics in the study of Alzheimer's disease. Traditionally, this task is performed by formulating a linear regression problem. Recently, it is found that using a linear sparse regression model can achieve better prediction accuracy. However, most existing studies only focus on the exploitation of sparsity of regression coefficients, ignoring useful structure information in regression coefficients. Also, these linear sparse models may not capture more complicated and possibly nonlinear relationships between cognitive performance and MRI measures. Motivated by these observations, in this work we build a sparse multivariate regression model for this task and propose an empirical sparse Bayesian learning algorithm. Different from existing sparse algorithms, the proposed algorithm models the response as a nonlinear function of the predictors by extending the predictor matrix with block structures. Further, it exploits not only inter-vector correlation among regression coefficient vectors, but also intra-block correlation in each regression coefficient vector. Experiments on the Alzheimer's Disease Neuroimaging Initiative database showed that the proposed algorithm not only achieved better prediction performance than state-of-the-art competitive methods, but also effectively identified biologically meaningful patterns. Zhilin Zhang 0002, Bhaskar D. Rao, Shiaofen Fang, Andrew J. Saykin, Li Shen 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2013 | Genetic Clustering on the Hippocampal Surface for Genome-Wide Association Studies
Derrek P. Hibar, Sarah E. Medland, Jason L. Stein, Sungeun Kim, Li Shen 0001, Andrew J. Saykin, Greig I. de Zubicaray, Katie L. McMahon, Grant W. Montgomery, Nicholas G. Martin, Margaret J. Wright, Srdjan Djurovic, Ingrid Agartz, Ole A. Andreassen, Paul M. Thompson |
MICCAI (2) | 5 |
| 2013 | A New Sparse Simplex Model for Brain Anatomical and Genetic Network Analysis
Heng Huang 0001, Feiping Nie 0001, Tom Weidong Cai, Andrew J. Saykin, Li Shen 0001 |
MICCAI (2) | 7 |
| 2013 | Interactive object extraction by merging regions with k-global maximal similarity
Taiyong Li, Zhilong Xie, Li Shen 0001 |
Neurocomputing | 5 |
| 2012 | Sparse Bayesian multi-task learning for predicting cognitive outcomes from neuroimaging measures in Alzheimer's diseaseabstractAlzheimer’s disease (AD) is the most common form of de-mentia that causes progressive impairment of memory and other cognitive functions. Multivariate regression models have been studied in AD for revealing relationships between neuroimaging measures and cognitive scores to understand how structural changes in brain can influence cognitive sta-tus. Existing regression methods, however, do not explic-itly model dependence relation among multiple scores de-rived from a single cognitive test. It has been found that such dependence can deteriorate the performance of these methods. To overcome this limitation, we propose an effi-cient sparse Bayesian multi-task learning algorithm, which adaptively learns and exploits the dependence to achieve improved prediction performance. The proposed algorithm is applied to a real world neuroimaging study in AD to pre-dict cognitive performance using MRI scans. The effective-ness of the proposed algorithm is demonstrated by its supe-rior prediction performance over multiple state-of-the-art competing methods and accurate identification of compact sets of cognition-relevant imaging biomarkers that are con-sistent with prior knowledge. 1. Zhilin Zhang 0002, Taiyong Li, Bhaskar D. Rao, Shiaofen Fang, Sungeun Kim, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001 |
CVPR | 10 |
| 2012 | High-Order Multi-Task Feature Learning to Identify Longitudinal Phenotypic Markers for Alzheimer's Disease Progression PredictionabstractAlzheimer disease (AD) is a neurodegenerative disorder characterized by progressive impairment of memory and other cognitive functions. Regression analysis has been studied to relate neuroimaging measures to cognitive status. However, whether these measures have further predictive power to infer a trajectory of cognitive performance over time is still an under-explored but important topic in AD research. We propose a novel high-order multi-task learning model to address this issue. The proposed model explores the temporal correlations existing in data features and regression tasks by the structured sparsity-inducing norms. In addition, the sparsity of the model enables the selection of a small number of MRI measures while maintaining high prediction accuracy. The empirical studies, using the baseline MRI and serial cognitive data of the ADNI cohort, have yielded promising results. Hua Wang 0007, Feiping Nie 0001, Heng Huang 0001, Sungeun Kim, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001 |
NIPS | 8 |
| 2012 | Principles of Computational Modeling in NeuroscienceDavid Sterratt, Bruce Graham, Andrew Gillies and David WillshawabstractUnderstanding complex neurobiological systems is one of the most difficult challenges in modern science. This book is focused on computational neuroscience, which provides a mathematical foundation and a rich set of computational approaches for understanding the principles and dynamics of the nervous system. Taking a bottom up strategy, the book offers comprehensive, step-by-step coverage on how to model the neuron and neural circuitry to understand the nervous system at multiple levels, from ion channels to networks. It has arisen from graduate courses to students in physical, mathematical and computer sciences, and can serve as an excellent text for courses in computational neuroscience. The book consists of 11 chapters and 2 appendices. Chapter 1 gives an overview of the book. Chapter 2 provides the basis for the development of neuronal models by introducing primary electrical properties of neurons. It starts with an introduction to the neuronal membrane, and then discusses physical basis of ion movement in neurons, including electrical drift and diffusion. The Nernst equation is introduced to model the equilibrium potential resulting from permeability to a single ion in the resting membrane. The Goldman–Hodgkin–Katz equations and their simplifications are presented to model membrane ionic with more than one type of membrane-permeable ions. Then, the resistor–capacitor (RC) circuit is introduced to approximate the equivalent electrical circuit of a patch of membrane. The compartmental model is discussed to handle the spatial extent of the membrane. Finally, the cable equation is presented to model the spatiotemporal evolution of the membrane potential. Chapter 3 presents the first quantitative model, the Hodgkin–Huxley (HH) model, for describing the action membrane potential. After introducing the basic concept of the action potential, it explores how a mixture of physical intuition and curve-fitting can be used to produce the HH mathematical scheme, which models the squid giant axon sodium and potassium voltage-gated ion channels. The HH model is then used to simulate nerve action potentials and is evaluated by a comparison with empirical recordings. After considering the temperature effect, it concludes with a discussion of how to use the HH formalism to build models of ion channels in general. Chapter 4 explores how voltage spreads along the membrane using multiple connected RC circuits, referred to as compartmental modeling. It starts with the construction of a compartmental model by representing quasi-isopotential sections of neurite (small pieces of dendrite, axon or soma) as compartments, which are simple geometric objects such as spheres or cylinders. It then presents approaches for using real neuronal morphology as the basis of the model. After that, it considers in detail methods and issues of parameter estimation for determining specific passive membrane properties using an electrical model instead of direct measurement with experimental techniques that is often impractical. Finally, it discusses the issues involved in incorporating active channels into compartmental models. Chapter 5 concentrates on the theory of modeling different types of active ion channels, including voltage- and ligand-gated. It first reviews channel structure and function, ion channel nomenclature and experimental techniques for model development. Three formalisms are discussed for modeling ensembles of voltage-gated ion channels: the HH formalism, the thermodynamic formalism and the Markov kinetic scheme. Methods for modeling ligand-gated channels are discussed using calcium-dependent potassium channels as an example. Transition state theory is introduced and applied to constrain the rate coefficients in both the thermodynamic formalism and the Markov kinetic scheme. Chapter 6 describes methods of modeling intracellular signaling systems with a focus on calcium. It details how intracellular calcium concentration can be modeled, and examines models for voltage-gated calcium channels, membrane-bound pumps, calcium buffers and diffusion. Examples of intracellular signaling pathways involving more complex enzymatic reactions and cascades are also considered. The well-mixed approach is used to model these pathways with sufficient presence of molecular species, while stochastic methods are used for compartments with small numbers of molecules. Spatial modeling is discussed to handle spatially inhomogeneous systems. Chapter 7 explores a range of models for chemical synapses and also briefly discusses electrical synapses. It first focuses on methods for modeling postsynaptic responses using simple phenomenological waveforms and more complex kinetic schemes, then models presynaptic neurotransmitter release by considering vesicle release and vesicle availability, and finally forms complete synapse models by relating the arrival of the presynaptic action potential with a postsynaptic response. Methods for modeling short-term dynamics and long-lasting synaptic plasticity are both discussed. Chapter 8 covers a spectrum of simplified models of neurons by stripping complicated ones down to their bare essentials. These models are computationally more efficient and are useful for incorporating into networks. Three broad categories of neuron models are discussed. The first category includes simplified multi-compartmental models with reduced number of compartments and reduced number of state variables. The second category includes the integrate-and-fire model and related spike-response model, where the membrane potential is left as the only state variable. At the simplest end of the spectrum, the third category includes rate-based models, which communicate via firing rates rather than individual spikes. Although highly simplified, these models can reasonably represent neuronal activity in certain areas of the nervous system and provide insights into complex computations. Chapter 9 examines a variety of network models where neurons are modeled with different levels of detail. It first presents methods for constructing the feed-forward and recurrent networks of the associative memory, where a neuron is treated as a two-state device and a synapse as a simple binary or linear device. It then describes another method to embed recurrent associative memory networks in a network of excitatory and inhibitory integrate-and-fire neurons, and next presents more complex network models of conductance-based neurons where associative memory can be embedded. After that, it explores two different models of thalamocortical interactions: one with multi-compartmental neurons, the other with spiking neurons. Finally, it discusses multi-compartmental models of the basal ganglia for elucidating the electrophysiological effects of deep brain stimulation. Chapter 10 reviews examples of modeling methods in developmental neuroscience. It discusses methods at the single neuron level, including developmental models for neuronal morphology and physiology. It also examines modeling methods for the development of the spatial arrangement of the nerve cells within a neural structure and those for the development of connectivity patterns (patterns of ocular dominance, connections between nerve and muscle, and retinotopic maps). Chapter 11 concludes the entire book by a brief discussion of the history and future directions of computational neuroscience. This book has done a nice job of laying out their strategy for covering major topics in the field of computational neuroscience while maintaining a well-organized structure. It is prepared for both expert and non-expert readers with an elementary background in neuroscience and some high school mathematics. As a computer scientist working on analyzing human brain data at the macro-scale level and a non-expert in modeling the neuronal data at the micro-scale level, I have found this book very easy to follow and have enjoyed reading it. I really like it that this book provides plenty of figures, tables and boxes to assist with the presentation of the content in a very informative, intuitive and organized manner, which I feel plays an important role in helping readers better understand the fundamental concepts, scientific problems and computational models. In addition to appendices containing overviews and links to relevant computational and mathematical resources, the authors maintain a book website that provides another source of useful and up-to-date information. These resources are especially beneficial when the book is used as a text for a course. The code examples and links to external simulators and databases could be very valuable to the preparation of lecture notes and student exercises. It would be a more desirable case if the authors could consider including a collection of lecture slides and course exercises at the website when available. This book is also a good resource for people working in relevant fields. For example, the NIH Human Connectome Project is a newly launched, landmark study that employs a neuroinformatics and systems biology approach for understanding the human brain as a complex network at multiple levels, from genetic determinants to cellular processes to the complex interplay of brain structure, function, behavior and cognition. The network connectivity of the human brain needs to be analyzed in a comprehensive fashion across all scales, from the micro-scale of synaptic connections between neurons to the macro-scale of structural and functional connections between anatomically distinct brain regions. The concepts and modeling methods presented in this book are directly applicable to the micro-scale analysis, and could also be very helpful to the macro-scale analysis, since the knowledge learned from the micro-scale analysis can help yield biological insights and reveal mechanisms underlying macro-scale phenomena. In summary, this is a timely, well-written book that provides a comprehensive, in-depth and state-of-the-art coverage of computational modeling in neuroscience. It can serve as an excellent text for a graduate level course in computational neuroscience, as well as a valuable reference for experimental neuroscientists, computational neuroscientists and people working in relevant areas such as neuroinformatics and systems biology. Supported in part by NSF IIS-1117335, NIH UL1 RR025761, U01 AG024904, NIA RC2 AG036535, NIA R01 AG19771, and NIA P30 AG10133-18S1. Li Shen 0001 |
Briefings Bioinform. | 1 |
| 2012 | Identifying quantitative trait loci via group-sparse multitask regression and feature selection: an imaging genetics study of the ADNI cohortabstractMOTIVATION: Recent advances in high-throughput genotyping and brain imaging techniques enable new approaches to study the influence of genetic variation on brain structures and functions. Traditional association studies typically employ independent and pairwise univariate analysis, which treats single nucleotide polymorphisms (SNPs) and quantitative traits (QTs) as isolated units and ignores important underlying interacting relationships between the units. New methods are proposed here to overcome this limitation. RESULTS: Taking into account the interlinked structure within and between SNPs and imaging QTs, we propose a novel Group-Sparse Multi-task Regression and Feature Selection (G-SMuRFS) method to identify quantitative trait loci for multiple disease-relevant QTs and apply it to a study in mild cognitive impairment and Alzheimer's disease. Built upon regression analysis, our model uses a new form of regularization, group ℓ(2,1)-norm (G(2,1)-norm), to incorporate the biological group structures among SNPs induced from their genetic arrangement. The new G(2,1)-norm considers the regression coefficients of all the SNPs in each group with respect to all the QTs together and enforces sparsity at the group level. In addition, an ℓ(2,1)-norm regularization is utilized to couple feature selection across multiple tasks to make use of the shared underlying mechanism among different brain regions. The effectiveness of the proposed method is demonstrated by both clearly improved prediction performance in empirical evaluations and a compact set of selected SNP predictors relevant to the imaging QTs. AVAILABILITY: Software is publicly available at: http://ranger.uta.edu/%7eheng/imaging-genetics/. Hua Wang 0007, Feiping Nie 0001, Heng Huang 0001, Sungeun Kim, Kwangsik Nho, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 8 |
| 2012 | Identifying disease sensitive and quantitative trait-relevant biomarkers from multidimensional heterogeneous imaging genetics data via sparse multimodal multitask learningabstractMOTIVATION: Recent advances in brain imaging and high-throughput genotyping techniques enable new approaches to study the influence of genetic and anatomical variations on brain functions and disorders. Traditional association studies typically perform independent and pairwise analysis among neuroimaging measures, cognitive scores and disease status, and ignore the important underlying interacting relationships between these units. RESULTS: To overcome this limitation, in this article, we propose a new sparse multimodal multitask learning method to reveal complex relationships from gene to brain to symptom. Our main contributions are three-fold: (i) introducing combined structured sparsity regularizations into multimodal multitask learning to integrate multidimensional heterogeneous imaging genetics data and identify multimodal biomarkers; (ii) utilizing a joint classification and regression learning model to identify disease-sensitive and cognition-relevant biomarkers; (iii) deriving a new efficient optimization algorithm to solve our non-smooth objective function and providing rigorous theoretical analysis on the global optimum convergency. Using the imaging genetics data from the Alzheimer's Disease Neuroimaging Initiative database, the effectiveness of the proposed method is demonstrated by clearly improved performance on predicting both cognitive scores and disease status. The identified multimodal biomarkers could predict not only disease status but also cognitive function to help elucidate the biological pathway from gene to brain structure and function, and to cognition and disease. AVAILABILITY: Software is publicly available at: http://ranger.uta.edu/%7eheng/multimodal/. Hua Wang 0007, Feiping Nie 0001, Heng Huang 0001, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 6 |
| 2012 | From phenotype to genotype: an association study of longitudinal phenotypic markers to Alzheimer's disease relevant SNPsabstractMOTIVATION: Imaging genetic studies typically focus on identifying single-nucleotide polymorphism (SNP) markers associated with imaging phenotypes. Few studies perform regression of SNP values on phenotypic measures for examining how the SNP values change when phenotypic measures are varied. This alternative approach may have a potential to help us discover important imaging genetic associations from a different perspective. In addition, the imaging markers are often measured over time, and this longitudinal profile may provide increased power for differentiating genotype groups. How to identify the longitudinal phenotypic markers associated to disease sensitive SNPs is an important and challenging research topic. RESULTS: Taking into account the temporal structure of the longitudinal imaging data and the interrelatedness among the SNPs, we propose a novel 'task-correlated longitudinal sparse regression' model to study the association between the phenotypic imaging markers and the genotypes encoded by SNPs. In our new association model, we extend the widely used ℓ(2,1)-norm for matrices to tensors to jointly select imaging markers that have common effects across all the regression tasks and time points, and meanwhile impose the trace-norm regularization onto the unfolded coefficient tensor to achieve low rank such that the interrelationship among SNPs can be addressed. The effectiveness of our method is demonstrated by both clearly improved prediction performance in empirical evaluations and a compact set of selected imaging predictors relevant to disease sensitive SNPs. AVAILABILITY: Software is publicly available at: http://ranger.uta.edu/%7eheng/Longitudinal/ CONTACT: [email protected] or [email protected]. Hua Wang 0007, Feiping Nie 0001, Heng Huang 0001, Sungeun Kim, Kwangsik Nho, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001 |
Bioinform. | 9 |
| 2011 | Sparse multi-task regression and feature selection to identify brain imaging predictors for memory performanceabstractAlzheimer's disease (AD) is a neurodegenerative disorder characterized by progressive impairment of memory and other cognitive functions, which makes regression analysis a suitable model to study whether neuroimaging measures can help predict memory performance and track the progression of AD. Existing memory performance prediction methods via regression, however, do not take into account either the interconnected structures within imaging data or those among memory scores, which inevitably restricts their predictive capabilities. To bridge this gap, we propose a novel Sparse Multi-tAsk Regression and feaTure selection (SMART) method to jointly analyze all the imaging and clinical data under a single regression framework and with shared underlying sparse representations. Two convex regularizations are combined and used in the model to enable sparsity as well as facilitate multi-task learning. The effectiveness of the proposed method is demonstrated by both clearly improved prediction performances in all empirical test cases and a compact set of selected RAVLT-relevant MRI predictors that accord with prior studies. Hua Wang 0007, Feiping Nie 0001, Heng Huang 0001, Shannon L. Risacher, Chris Ding, Andrew J. Saykin, Li Shen 0001 |
ICCV | 7 |
| 2011 | Hippocampal Surface Mapping of Genetic Risk Factors in AD via Sparse Learning Models
Sungeun Kim, Mark Inlow, Kwangsik Nho, Shanker Swaminathan, Shannon L. Risacher, Shiaofen Fang, Michael Weiner 0001, Mirza Faisal Beg, Lei Wang 0032, Andrew J. Saykin, Li Shen 0001 |
MICCAI (2) | 12 |
| 2011 | Identifying AD-Sensitive and Cognition-Relevant Imaging Biomarkers via Joint Classification and Regression
Hua Wang 0007, Feiping Nie 0001, Heng Huang 0001, Shannon L. Risacher, Andrew J. Saykin, Li Shen 0001 |
MICCAI (3) | 6 |
| 2010 | Sparse Bayesian Learning for Identifying Imaging Biomarkers in AD Prediction
Li Shen 0001, Yuan Qi 0001, Sungeun Kim, Kwangsik Nho, Shannon L. Risacher, Andrew J. Saykin |
MICCAI (3) | 1 |
| 2009 | Data synthesis and tool development for exploring imaging genomic patternsabstractRecent advances in brain imaging and high throughput genotyping techniques enable new approaches to study the influence of genetic variation on brain structure and function. However, major computational challenges are bottlenecks for comprehensive joint analysis of these high-dimensional image and genomic data. We report our initial progress in developing an imaging genomic browsing system for integrated exploration of neuroimaging and genomic data. We describe a method for synthesizing a set of realistic neuroimaging and genomic data, where the relationships between imaging phenotypes and genotypes are known. This data set is used to demonstrate the functionality of our system, which is designed for effectively exploring the neuroanatomical distribution of statistical results that measure the associations between brain imaging phenotypes and genotypes on a genome-wide scale. The proposed system has substantial potential for enabling discovery of important imaging genomic associations through visual evaluation and can be extended towards several directions. Sungeun Kim, Li Shen 0001, Andrew J. Saykin, John D. West |
CIBCB | 2 |
| 2009 | Fourier method for large-scale surface modeling and registration
Li Shen 0001, Sungeun Kim, Andrew J. Saykin |
Comput. Graph. | 1 |
| 2007 | Morphometric Analysis of Hippocampal Shape in Mild Cognitive Impairment: An Imaging Genetics StudyabstractA computational framework is presented for surface based morphometry to localize shape changes between groups of 3D objects. It employs the spherical harmonic (SPHARM) method for surface modeling and random field theory (RFT) for statistical inference. Several new components are introduced to overcome previous limitations: (1) a general linear model is used to facilitate controlling for covariates; (2) a new SPHARM registration method SHREC is proposed to better align SPHARM models; and (3) an estimated smoothness is used in RFT-based analysis to obtain more accurate results. This framework is applied in a mild cognitive impairment (MCI) study to examine hippocampal shape changes related to diagnostic and genetic conditions. Several interesting findings from our analyses suggest combining imaging phenotypes and genetic profiles has the potential to elucidate biological pathways for better understanding MCI and Alzheimer's disease. Li Shen 0001, Andrew J. Saykin, Moo K. Chung, Heng Huang 0001 |
BIBE | 1 |
| 2007 | Surface Harmonics for Shape ModelingabstractWe present an approach for three dimensional (3D) shape modeling using Jacobi polynomials based surface harmonics. Because we construct a set of complete hemispherical harmonic basis functions on a hemisphere domain from the associated Jacobi polynomials, our shape modeling method work efficiently on the open hemisphere-like objects that often exist in medical anatomical structures (e.g., ventricles, atriums, etc.). We demonstrate the effectiveness of our approach through theoretic and experimental exploration of a set of medical image applications. Heng Huang 0001, Li Shen 0001 |
ICIP (2) | 2 |
| 2007 | A Novel Surface Registration Algorithm With Biomedical Modeling ApplicationsabstractIn this paper, we propose a novel surface matching algorithm for arbitrarily shaped but simply connected 3-D objects. The spherical harmonic (SPHARM) method is used to describe these 3-D objects, and a novel surface registration approach is presented. The proposed technique is applied to various applications of medical image analysis. The results are compared with those using the traditional method, in which the first-order ellipsoid is used for establishing surface correspondence and aligning objects. In these applications, our surface alignment method is demonstrated to be more accurate and flexible than the traditional approach. This is due in large part to the fact that a new surface parameterization is generated by a shortcut that employs a useful rotational property of spherical harmonic basis functions for a fast implementation. In order to achieve a suitable computational speed for practical applications, we propose a fast alignment algorithm that improves computational complexity of the new surface registration method from O(n3) to O(n2). Heng Huang 0001, Li Shen 0001, Rong Zhang 0013, Fillia Makedon, Andrew J. Saykin, Justin D. Pearlman |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2007 | Weighted Fourier Series Representation and Its Application to Quantifying the Amount of Gray MatterabstractWe present a novel weighted Fourier series (WFS) representation for cortical surfaces. The WFS representation is a data smoothing technique that provides the explicit smooth functional estimation of unknown cortical boundary as a linear combination of basis functions. The basic properties of the representation are investigated in connection with a self-adjoint partial differential equation and the traditional spherical harmonic (SPHARM) representation. To reduce steep computational requirements, a new iterative residual fitting (IRF) algorithm is developed. Its computational and numerical implementation issues are discussed in detail. The computer codes are also available at http://www.stat.wisc.edu/-mchung/softwares/weighted.SPHARM/weighted-SPHARM.html. As an illustration, the WFS is applied i n quantifying the amount ofgray matter in a group of high functioning autistic subjects. Within the WFS framework, cortical thickness and gray matter density are computed and compared. Moo K. Chung, Kim M. Dalton, Li Shen 0001, Alan C. Evans, Richard J. Davidson |
IEEE Trans. Medical Imaging | 3 |
| 2006 | Spherical mapping for processing of 3D closed surfaces
Li Shen 0001, Fillia Makedon |
Image Vis. Comput. | 1 |
| 2005 | Surface Alignment of 3D Spherical Harmonic Models: Application to Cardiac MRI Analysis
Heng Huang 0001, Li Shen 0001, Rong Zhang 0013, Fillia Makedon, Bruce Hettleman, Justin D. Pearlman |
MICCAI | 2 |
| 2005 | A Prediction Framework for Cardiac Resynchronization Therapy Via 4D Cardiac Motion Analysis
Heng Huang 0001, Li Shen 0001, Rong Zhang 0013, Fillia Makedon, Bruce Hettleman, Justin D. Pearlman |
MICCAI | 2 |
| 2004 | A surface-based approach for classification of 3D neuroanatomic structures
Li Shen 0001, James Ford, Fillia Makedon, Andrew J. Saykin |
Intell. Data Anal. | 1 |
| 2003 | Morphometric Analysis of Brain Structures for Improved Discrimination
Li Shen 0001, James Ford, Fillia Makedon, Tilmann Steinberg, Andrew J. Saykin |
MICCAI (2) | 1 |
| 2003 | A Spatio-temporal Multi-modal Data Management and Analysis Environment for Tracking MS LesionsabstractWe describe the development of a system that automates data collection, metadata extraction and analysis of spatio-temporal multi-modal data, combining data management and data analysis to provide an efficient resource for clinicians. Though the system is extensible to many applications, the current focus is on managing Multiple Sclerosis (MS) lesion data, which are disparate streams of image, numeric, and text data. In order to discover patterns of MS pathology and plan early and effective treatment, multispectral magnetic resonance (MR) image streams collected over time need to be correlated efficiently with each other and with patient performance and clinical data streams. Tilmann Steinberg, Fillia Makedon, Li Shen 0001, Andrew J. Saykin, Heather Wishart |
SSDBM | 4 |
| 2000 | Fast Association Discovery in Derivative Transaction Collections
Li Shen 0001, Paul Pritchard |
Knowl. Inf. Syst. | 1 |
| 1999 | New Algorithms for Efficient Mining of Association RulesabstractDiscovery of association rules is an important data mining task. Several algorithms have been proposed to solve this problem. Most of them require repeated passes over the database, which incurs huge I/O overhead and high synchronization expense in parallel cases. There are a few algorithms trying to reduce these costs. But they contain weaknesses such as often requiring high pre-processing cost to get a vertical database layout, containing much redundant computation in parallel cases, and so on. We propose new association mining algorithms to overcome the above drawbacks, through minimizing the I/O cost and effectively controlling the computation cost. Experiments on well-known synthetic data show that our algorithms consistently outperform a priori, one of the best algorithms for association mining, by factors ranging from 2 to 4 in most cases. Also, our algorithms are very easy to be parallelized, and we present a parallelization for them based on a shared-nothing architecture. We observe that our parallelization develops the parallelism more sufficiently than two of the best existing parallel algorithms. Li Shen 0001 |
Inf. Sci. | 1 |
| 1998 | Mining Flexible Multiple-Level Association Rules in All Concept Hierarchies (Extended Abstract)
Li Shen 0001 |
DEXA | 1 |