Jingxuan Bao

dblp:282/4352 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
7since 2021 · last 2024
0000-0001-7127-3258ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2024 Flex-MoE: Modeling Arbitrary Modality Combination via the Flexible Mixture-of-Experts
abstract
Multimodal learning has gained increasing importance across various fields, offering the ability to integrate data from diverse sources such as images, text, and personalized records, which are frequently observed in medical domains. However, in scenarios where some modalities are missing, many existing frameworks struggle to accommodate arbitrary modality combinations, often relying heavily on a single modality or complete data. This oversight of potential modality combinations limits their applicability in real-world situations. To address this challenge, we propose Flex-MoE (Flexible Mixture-of-Experts), a new framework designed to flexibly incorporate arbitrary modality combinations while maintaining robustness to missing data. The core idea of Flex-MoE is to first address missing modalities using a new missing modality bank that integrates observed modality combinations with the corresponding missing ones. This is followed by a uniquely designed Sparse MoE framework. Specifically, Flex-MoE first trains experts using samples with all modalities to inject generalized knowledge through the generalized router ($\mathcal{G}$-Router). The $\mathcal{S}$-Router then specializes in handling fewer modality combinations by assigning the top-1 gate to the expert corresponding to the observed modality combination. We evaluate Flex-MoE on the ADNI dataset, which encompasses four modalities in the Alzheimer's Disease domain, as well as on the MIMIC-IV dataset. The results demonstrate the effectiveness of Flex-MoE, highlighting its ability to model arbitrary modality combinations in diverse missing modality scenarios. Code is available at: \url{https://github.com/UNITES-Lab/flex-moe}.
Sukwon Yun, Inyoung Choi, Jie Peng 0002, Yangfan Wu, Jingxuan Bao, Qiyiwen Zhang, Jiayi Xin, Qi Long, Tianlong Chen 0001
NeurIPS5
2024 Interpretable deep clustering survival machines for Alzheimer's disease subtype discovery
Bojian Hou, Zixuan Wen, Jingxuan Bao, Richard Zhang 0001, Boning Tong, Shu Yang 0009, Junhao Wen 0002, Yuhan Cui, Jason H. Moore, Andrew J. Saykin, Heng Huang 0001, Paul M. Thompson, Marylyn D. Ritchie, Christos Davatzikos, Li Shen 0001
Medical Image Anal.3
2023 Integrative analysis of multi-omics and imaging data with incorporation of biological information via structural Bayesian factor analysis
abstract
MOTIVATION: With the rapid development of modern technologies, massive data are available for the systematic study of Alzheimer's disease (AD). Though many existing AD studies mainly focus on single-modality omics data, multi-omics datasets can provide a more comprehensive understanding of AD. To bridge this gap, we proposed a novel structural Bayesian factor analysis framework (SBFA) to extract the information shared by multi-omics data through the aggregation of genotyping data, gene expression data, neuroimaging phenotypes and prior biological network knowledge. Our approach can extract common information shared by different modalities and encourage biologically related features to be selected, guiding future AD research in a biologically meaningful way. METHOD: Our SBFA model decomposes the mean parameters of the data into a sparse factor loading matrix and a factor matrix, where the factor matrix represents the common information extracted from multi-omics and imaging data. Our framework is designed to incorporate prior biological network information. Our simulation study demonstrated that our proposed SBFA framework could achieve the best performance compared with the other state-of-the-art factor-analysis-based integrative analysis methods. RESULTS: We apply our proposed SBFA model together with several state-of-the-art factor analysis models to extract the latent common information from genotyping, gene expression and brain imaging data simultaneously from the ADNI biobank database. The latent information is then used to predict the functional activities questionnaire score, an important measurement for diagnosis of AD quantifying subjects' abilities in daily life. Our SBFA model shows the best prediction performance compared with the other factor analysis models. AVAILABILITY: Code are publicly available at https://github.com/JingxuanBao/SBFA. CONTACT: [email protected].
Jingxuan Bao, Changgee Chang, Qiyiwen Zhang, Andrew J. Saykin, Li Shen 0001, Qi Long
Briefings Bioinform.1
2022 Preference Matrix Guided Sparse Canonical Correlation Analysis for Genetic Study of Quantitative Traits in Alzheimer's Disease
abstract
Investigating the relationship between genetic variation and phenotypic traits is a key issue in quantitative genetics. Specifically for Alzheimer's disease, the association between genetic markers and quantitative traits remains vague while, once identified, will provide valuable guidance for the study and development of genetic-based treatment approaches. Currently, to analyze the association of two modalities, sparse canonical correlation analysis (SCCA) is commonly used to compute one sparse linear combination of the variable features for each modality, giving a pair of linear combination vectors in total that maximizes the cross-correlation between the analyzed modalities. One drawback of the plain SCCA model is that the existing findings and knowledge cannot be integrated into the model as priors to help extract interesting correlation as well as identify biologically meaningful genetic and phenotypic markers. To bridge this gap, we introduce preference matrix guided SCCA (PM-SCCA) that not only takes priors encoded as a preference matrix but also maintains computational simplicity. A simulation study and a real-data experiment are conducted to investigate the effectiveness of the model. Both experiments demonstrate that the proposed PM-SCCA model can capture not only genotype-phenotype correlation but also relevant features effectively.
Jiahang Sha, Jingxuan Bao, Kefei Liu 0001, Shu Yang 0009, Zixuan Wen, Yuhan Cui, Junhao Wen 0002, Christos Davatzikos, Jason H. Moore, Andrew J. Saykin, Qi Long, Li Shen 0001
BIBM2
2022 Identifying genes associated with brain volumetric differences through tissue specific transcriptomic inference from GWAS summary data
abstract
BACKGROUND: Brain volume has been widely studied in the neuroimaging field, since it is an important and heritable trait associated with brain development, aging and various neurological and psychiatric disorders. Genome-wide association studies (GWAS) have successfully identified numerous associations between genetic variants such as single nucleotide polymorphisms and complex traits like brain volume. However, it is unclear how these genetic variations influence regional gene expression levels, which may subsequently lead to phenotypic changes. S-PrediXcan is a tissue-specific transcriptomic data analysis method that can be applied to bridge this gap. In this work, we perform an S-PrediXcan analysis on GWAS summary data from two large imaging genetics initiatives, the UK Biobank and Enhancing Neuroimaging Genetics through Meta Analysis, to identify tissue-specific transcriptomic effects on two closely related brain volume measures: total brain volume (TBV) and intracranial volume (ICV). RESULTS: As a result of the analysis, we identified 10 genes that are highly associated with both TBV and ICV. Nine out of 10 genes were found to be associated with TBV in another study using a different gene-based association analysis. Moreover, most of our discovered genes were also found to be correlated with multiple cognitive and behavioral traits. Further analyses revealed the protein-protein interactions, associated molecular pathways and biological functions that offer insight into how these genes function and interact with others. CONCLUSIONS: These results confirm that S-PrediXcan can identify genes with tissue-specific transcriptomic effects on complex traits. The analysis also suggested novel genes whose expression levels are related to brain volumetric traits. This provides important insights into the genetic mechanisms of the human brain.
Hung Mai, Jingxuan Bao, Paul M. Thompson, Do Kyoon Kim, Li Shen 0001
BMC Bioinform.2
2021 A Novel Bayesian Semi-parametric Model for Learning Heritable Imaging Traits
Yize Zhao, Xiwen Zhao, Mansu Kim, Jingxuan Bao, Li Shen 0001
MICCAI (5)4
2021 A structural enriched functional network: An application to predict brain cognitive performance
Mansu Kim, Jingxuan Bao, Kefei Liu 0001, Bo-yong Park, Hyunjin Park, Jae Young Baik, Li Shen 0001
Medical Image Anal.2
2020 Estimating Hard-tissue Conditions from Dental Images via Machine Learning
abstract
Despite the great success of machine learning in various biomedical domains, applications to dental hard tissue conditions (primarily on dental Caries, Erosive Tooth Wear (ETW), and Fluorosis) are under-explored, in particular for analyzing photographic images. The clinical diagnostics of these dental hard-tissue conditions is routinely performed by visual examination but is often limited by its subjectivity. To bridge this gap, we apply four categories of machine learning strategies including nine different methods with two different feature representations to estimate the probability and severity of dental hard-tissue conditions from photographic tooth images. Our first empirical study is performed on the real dataset containing both controls and cases, and the best probability estimation results are achieved by Extra Trees Regression (RMSE: 0.030, Pearson correlation: 0.600) for Caries, Decision Tree (RMSE: 0.183, Pearson correlation: 0.581) for ETW, and Bayesian ARD Regression (RMSE: 0.191, Pearson correlation: 0.745) for Fluorosis. Our second empirical study is performed on the case only datasets, and the best severity estimation results are achieved by Extra Trees Regression (RMSE: 0.029, Pearson correlation: 0.687) for Caries, Bayesian ARD Regression and Linear Regression (RMSE: 0.192, Pearson correlation: 0.490) for ETW, and Bayesian ARD Regression (RMSE: 0.238, Pearson correlation: 0.537) for Fluorosis. These results indicate that machine learning models provide promising opportunities to help clinical evaluation and save resources in the management of these dental conditions.
Jingxuan Bao, Mansu Kim, Anderson T. Hara, Gerardo Maupome, Li Shen 0001
BIBE1