Li C. Xia

dblp:129/1305 · also Li Charlie Xia · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
7since 2021 · last 2025
0000-0003-0868-1923ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 7 since 2021
YearPublicationVenuePosition
2025 sxFusion: A Novel Single-Cell Clustering Tool Based on Feature Fusion and Co-Optimization of Low-Rank Representation
abstract
Single-cell RNA sequencing enables cellular-level gene expression analysis, but high sparsity and noise pose significant challenges for accurate cell-type identification. Here, we propose sxFusion, a structure-aware clustering method that integrates low-rank representation with orthogonal non-negative matrix factorization. sxFusion first integrates graph-based features using a graph convolutional network and nonlinear embeddings from a multilayer perceptron via dimension-wise gating to enhance feature discriminability and mitigating over-smoothing; it then employs a structure-guided low-rank self-expression model co-optimized with orthogonal non-negative matrix factorization, aligning learned representation directly with clustering assignments, and outputs clustering results straightforward. Benchmarked against eleven state-of-the-art methods on eleven single-cell datasets, sxFusion consistently achieved higher accuracy. Ablation and visualization analyses confirmed the advantages of the fusion and co-optimization modules. These results demonstrate that sxFusion effectively identifies biologically meaningful cell populations, providing robust and interpretable clusters.
Linping Wang, Jinqiao Wang, Hongxing Xu, Keyi Xiong, Xiangqi Bai, Kaida Ning, Li C. Xia
BIBM9
2025 EnzHier: Accurate Enzyme Function Prediction Through Multi-scale Feature Integration and Hierarchical Contrastive Learning
Hongyu Duan, Bozhen Ren, Fanghua Wang, Dongming Lan, Li C. Xia
ISBRA (2)9
2025 Accurate breast cancer intrinsic subtyping using DNA-only biomarkers with applications to multi-cohorts
abstract
Abstract Breast cancer intrinsic subtyping is critical for precision therapy and prognostic assessment. We applied the DNA-level multi-omics classifier UGES (Unified Genetic and Epigenetic Subtyping) [1] to two independent cohorts, TCGA (n = 931) and METABRIC (n = 1134), using a hierarchical learning strategy to progressively distinguish Basal-like, HER2-enriched, Luminal A, and Luminal B subtypes. UGES demonstrated excellent overall classification performance, achieving an AUC of 0.963 on the combined test set. In the challenging task of discriminating Luminal A from Luminal B tumors, UGES reached an AUC of 0.948, substantially outperforming a single-step strategy (0.828), highlighting the benefit of the stepwise approach. Compared with PAM50, UGES reclassified 11.57% of samples, with the most notable change being 11.25% of PAM50 Luminal A samples reassigned to UGES Luminal B. Additionally, UGES improved the sensitivity for HER2-enriched subtype recognition by 28.92%. Survival analyses revealed that UGES-defined subtypes provided stronger prognostic discrimination, with significant survival differences observed across most subtype pairs, particularly for HER2 compared with other subtypes. These findings demonstrate that UGES offers a robust and clinically relevant framework for breast cancer subtyping across independent cohorts, providing a reliable tool for real-world data analysis and clinical decision-making in precision oncology. References [1] Chang X., Xie J., Duan H., Li K., Liu X., Xiong Y., Bai X., Ning K., Xia L.C. ‘A unified genetic and epigenetic model to predict breast cancer intrinsic subtypes using large DNA-level multi-omics data and hierarchical learning.’ IEEE Transactions on Computational Biology and Bioinformatics 2025; early access. Doi:10.1109/TCBBIO.2025.3613591.
Xintong Chang, Linping Wang, Hongyu Duan, Li C. Xia
Briefings Bioinform.4
2025 Interpretable out-of-distribution diagnostics for enzyme stability prediction
abstract
Abstract Aim Enzyme property predictors often degrade when deployed to new enzyme families whose sequence and label distributions differ from training data. Investigating how distribution shifts among enzyme families affect prediction error is crucial for building trustworthy protein engineering systems. Utilizing the latest learning theory [1] and accompanying DataShifts algorithm (https://github.com/DataShifts/datashifts/), we present a shift-aware diagnostic workflow that quantifies distribution shift across enzyme families and turns cross-family generalization error into an interpretable quantity. Methods On the Kaggle Novozymes enzyme stability dataset, which contains nearly 10,000 samples spanning 180 enzyme families, we treat the 60 enzyme families with the largest pretraining errors as test domains with distribution shifts, and train the MLP model on the remaining enzyme families. For each test enzyme family, we produces a rigorous risk bound estimate and a bound decomposition that indicates whether performance is limited primarily by changes in sequence distribution (covariate shift) or by shifts in the sequence–label relationship (concept shift). Results On the training enzyme families, the model’s MAE for predicting enzyme melting temperature is 2.70, whereas across the test families the average MAE reaches 5.74. Across test enzyme families, our interpretable risk bound tracks the actual test error closely, with a correlation of 93.7%. For the top 10 enzyme families with the largest errors, the average MAE reaches 12.38; the contributions of covariate shift and concept shift to the risk bound are 5.20 and 12.53, respectively. Conclusion The primary cause of the model’s large predictive errors on unseen enzyme families is concept shift. In this case, simply expanding the training enzyme families data is unlikely to help, and fine-tuning on the new enzyme families is advisable. References 1. Hongbo Chen, and Li Charlie Xia. ‘General and Estimable Learning Bound Unifying Covariate and Concept Shifts.’ International Conference on Machine Learning (ICML) Dataworld Workshop 2025.
Li C. Xia
Briefings Bioinform.2
2025 A Unified Genetic and Epigenetic Model to Predict Breast Cancer Intrinsic Subtypes Using Large DNA-Level Multi-Omics Data and Hierarchical Learning
abstract
Breast cancer subtyping presents a significant clinical and scientific challenge. The prevalent expression-based Prediction Analysis of the Microarray of 50 genes (PAM50) system and its Immunohistochemistry (IHC) surrogate tests showed substantial inconsistencies and did not apply to the rapidly progressing circulating tumor DNA screenings. We developed Unified Genetic and Epigenetic Subtyping (UGES), a new intrinsic subtype classifier, by integrating large-scale DNA-level omics data with a hierarchy learning algorithm. Our benchmarks showed that both multi-step hierarchical learning and using all DNA-level alteration data are crucial, improving the overall AUC score by over 8.3% compared to the one-step multi-classification method. Based on these insights, we developed UGES, a three-step classifier based on 50831 DNA features of 2065 samples, including mutations, copy number aberrations, and methylations. UGES achieved an overall AUC score of 0.963 and greatly improved the clinical stratification of real-world patients, as each subtype strata's survival difference became statistically more significant, P = 9.7e-55 (UGES) vs. 2.2e-47 (PAM50). Finally, UGES identified 52 subtype-specific DNA biomarkers that can be targeted in early screening technology to expand the time window for precision care. The UGES code is freely available at https://github.com/labxscut/UGES.
Xintong Chang, Jiemin Xie, Hongyu Duan, Yunhui Xiong, Xiangqi Bai, Kaida Ning, Li C. Xia
IEEE Trans. Comput. Biol. Bioinform.9
2024 CTRhythm: Accurate Atrial Fibrillation Detection from Single-Lead ECG by Convolutional Neural Network and Transformer Integration
abstract
Atrial Fibrillation (AF) is a common supraventricular arrhythmia that affects about 30 million people globally. Electrocardiogram (ECG) analysis is the primary diagnostic approach. The widespread adoption of wearable devices for heart rhythm monitoring prompted the development of AF detection models for single-lead ECGs, benefitting real-time early diagnosis. Current state-of-the-art methods for AF detection are convolutional neural network (CNN) and convolutional recurrent neural network (CRNN) based models, which only focus on capturing local patterns despite heart rhythms exhibiting rich long-range dependencies. To address this limitation, we propose a novel method for singlelead ECG rhythm classification, termed CNN-Transformer Rhythm Classifier (CTRhythm), which integrates CNN with a Transformer encoder to effectively capture both local and global patterns effectively. CTRhythm achieved an overall F1 score of 0.831, outperforming baseline deep learning models on the gold-standard CINC2017 dataset. Furthermore, pretraining with additional data improved the overall F1 score to 0.840. In two external validation datasets, CTRhythm showed its strong generalization capabilities. CTRhythm is freely available at https://github.com/labxscut/CTRhythm.
Zhanyu Liang, Yinmingren Fu, Bozhen Ren, Maohuan Lin, Qingjiao Li, Yangxin Chen, Li C. Xia
BIBM10
2023 Identifying local associations in biological time series: algorithms, statistical significance, and applications
abstract
Local associations refer to spatial-temporal correlations that emerge from the biological realm, such as time-dependent gene co-expression or seasonal interactions between microbes. One can reveal the intricate dynamics and inherent interactions of biological systems by examining the biological time series data for these associations. To accomplish this goal, local similarity analysis algorithms and statistical methods that facilitate the local alignment of time series and assess the significance of the resulting alignments have been developed. Although these algorithms were initially devised for gene expression analysis from microarrays, they have been adapted and accelerated for multi-omics next generation sequencing datasets, achieving high scientific impact. In this review, we present an overview of the historical developments and recent advances for local similarity analysis algorithms, their statistical properties, and real applications in analyzing biological time series data. The benchmark data and analysis scripts used in this review are freely available at http://github.com/labxscut/lsareview.
Dongmei Ai, Lulu Chen, Jiemin Xie, Longwei Cheng, Yihui Luan, Shengwei Hou, Fengzhu Sun, Li C. Xia
Briefings Bioinform.10
2015 Statistical significance approximation in local trend analysis of high-throughput time-series data using the theory of Markov chains
abstract
BACKGROUND: Local trend (i.e. shape) analysis of time series data reveals co-changing patterns in dynamics of biological systems. However, slow permutation procedures to evaluate the statistical significance of local trend scores have limited its applications to high-throughput time series data analysis, e.g., data from the next generation sequencing technology based studies. RESULTS: By extending the theories for the tail probability of the range of sum of Markovian random variables, we propose formulae for approximating the statistical significance of local trend scores. Using simulations and real data, we show that the approximate p-value is close to that obtained using a large number of permutations (starting at time points >20 with no delay and >30 with delay of at most three time steps) in that the non-zero decimals of the p-values obtained by the approximation and the permutations are mostly the same when the approximate p-value is less than 0.05. In addition, the approximate p-value is slightly larger than that based on permutations making hypothesis testing based on the approximate p-value conservative. The approximation enables efficient calculation of p-values for pairwise local trend analysis, making large scale all-versus-all comparisons possible. We also propose a hybrid approach by integrating the approximation and permutations to obtain accurate p-values for significantly associated pairs. We further demonstrate its use with the analysis of the Polymouth Marine Laboratory (PML) microbial community time series from high-throughput sequencing data and found interesting organism co-occurrence dynamic patterns. AVAILABILITY: The software tool is integrated into the eLSA software package that now provides accelerated local trend and similarity analysis pipelines for time series data. The package is freely available from the eLSA website: http://bitbucket.org/charade/elsa.
Li C. Xia, Dongmei Ai, Jacob A. Cram, Xiaoyi Liang, Jed A. Fuhrman, Fengzhu Sun
BMC Bioinform.1
2013 Efficient statistical significance approximation for local similarity analysis of high-throughput time series data
abstract
MOTIVATION: Local similarity analysis of biological time series data helps elucidate the varying dynamics of biological systems. However, its applications to large scale high-throughput data are limited by slow permutation procedures for statistical significance evaluation. RESULTS: We developed a theoretical approach to approximate the statistical significance of local similarity analysis based on the approximate tail distribution of the maximum partial sum of independent identically distributed (i.i.d.) random variables. Simulations show that the derived formula approximates the tail distribution reasonably well (starting at time points > 10 with no delay and > 20 with delay) and provides P-values comparable with those from permutations. The new approach enables efficient calculation of statistical significance for pairwise local similarity analysis, making possible all-to-all local association studies otherwise prohibitive. As a demonstration, local similarity analysis of human microbiome time series shows that core operational taxonomic units (OTUs) are highly synergetic and some of the associations are body-site specific across samples. AVAILABILITY: The new approach is implemented in our eLSA package, which now provides pipelines for faster local similarity analysis of time series data. The tool is freely available from eLSA's website: http://meta.usc.edu/softs/lsa. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. CONTACT: [email protected].
Li C. Xia, Dongmei Ai, Jacob A. Cram, Jed A. Fuhrman, Fengzhu Sun
Bioinform.1
2010 PPLook: an automated data mining tool for protein-protein interaction
abstract
BACKGROUND: Extracting and visualizing of protein-protein interaction (PPI) from text literatures are a meaningful topic in protein science. It assists the identification of interactions among proteins. There is a lack of tools to extract PPI, visualize and classify the results. RESULTS: We developed a PPI search system, termed PPLook, which automatically extracts and visualizes protein-protein interaction (PPI) from text. Given a query protein name, PPLook can search a dataset for other proteins interacting with it by using a keywords dictionary pattern-matching algorithm, and display the topological parameters, such as the number of nodes, edges, and connected components. The visualization component of PPLook enables us to view the interaction relationship among the proteins in a three-dimensional space based on the OpenGL graphics interface technology. PPLook can also provide the functions of selecting protein semantic class, counting the number of semantic class proteins which interact with query protein, counting the literature number of articles appearing the interaction relationship about the query protein. Moreover, PPLook provides heterogeneous search and a user-friendly graphical interface. CONCLUSIONS: PPLook is an effective tool for biologists and biosystem developers who need to access PPI information from the literature. PPLook is freely available for non-commercial users at http://meta.usc.edu/softs/PPLook.
Shaowu Zhang 0001, Yao-Jun Li, Li C. Xia, Quan Pan 0001
BMC Bioinform.3