EDBT 2026 Demo / reviewers in the wild / expert
Lixin Cheng
dblp:19/2244
· DBLP profile ↗
27ranked-venue papers
7as first author
19since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Residual Reweighted Conformal Prediction for Graph Neural NetworksabstractGraph Neural Networks (GNNs) excel at modeling relational data but face significant challenges in high-stakes domains due to unquantified uncertainty. Conformal prediction (CP) offers statistical coverage guarantees, but existing methods often produce overly conservative prediction intervals that fail to account for graph heteroscedasticity and structural biases. While residual reweighting CP variants address some of these limitations, they neglect graph topology, cluster-specific uncertainties, and risk data leakage by reusing training sets. To address these issues, we propose Residual Reweighted GNN (RR-GNN), a framework designed to generate minimal prediction sets with provable marginal coverage guarantees. RR-GNN introduces three major innovations to enhance prediction performance. First, it employs Graph-Structured Mondrian CP to partition nodes or edges into communities based on topological features, ensuring cluster-conditional coverage that reflects heterogeneity. Second, it uses Residual-Adaptive Nonconformity Scores by training a secondary GNN on a held-out calibration set to estimate task-specific residuals, dynamically adjusting prediction intervals according to node or edge uncertainty. Third, it adopts a Cross-Training Protocol, which alternates the optimization of the primary GNN and the residual predictor to prevent information leakage while maintaining graph dependencies. We validate RR-GNN on 15 real-world graphs across diverse tasks, including node classification, regression, and edge weight prediction. Compared to CP baselines, RR-GNN achieves improved efficiency over state-of-the-art methods, with no loss of coverage. Zheng Zhang 0063, Zhixin Zhou, Nicolò Colombo, Lixin Cheng, Rui Luo 0002 |
UAI | 5 |
| 2025 | PAGE-based transfer learning from single-cell to bulk sequencing enhances model generalization for sepsis diagnosisabstractSepsis, caused by infections, sparks a dangerous bodily response. The transcriptional expression patterns of host responses aid in the diagnosis of sepsis, but the challenge lies in their limited generalization capabilities. To facilitate sepsis diagnosis, we present an updated version of single-cell Pair-wise Analysis of Gene Expression (scPAGE) using transfer learning method, scPAGE2, dedicated to data fusion between single-cell and bulk transcriptome. Compared to scPAGE, the upgrade to scPAGE2 featured ameliorated Differentially Expressed Gene Pairs (DEPs) for pretraining a model in single-cell transcriptome and retrained it using bulk transcriptome data to construct a sepsis diagnostic model, which effectively transferred cell-layer information from single-cell to bulk transcriptome. Seven datasets across three transcriptome platforms and fluorescence-activated cell sorting (FACS) were used for performance validation. The model involved four DEPs, showing robust performance across next-generation sequencing and microarray platforms, surpassing state-of-the-art models with an average AUROC of 0.947 and an average AUPRC of 0.987. Analysis of scRNA-seq data reveals higher cell proportions with JAM3-PIK3AP1 expression in sepsis monocytes, decreased ARG1-CCR7 in B and T cells. Elevated IRF6-HP in sepsis monocytes confirmed by both scRNA-seq and an independent cohort using FACS. Both the superior performance of the model and the in vitro validation of IRF6-HP in monocytes emphasize that scPAGE2 is effective and robust in the construction of sepsis diagnostic model. We additionally applied scPAGE2 to acute myeloid leukemia and demonstrated its superior classification performance. Overall, we provided a strategy to improve the generalizability of classification model that can be adapted to a broad range of clinical prediction scenarios. Nana Jin, Chuanchuan Nan, Wanyang Li, Peijing Lin, Jun Wang 0161, Yuelong Chen, Yuanhao Wang 0012, Kaijiang Yu, Changsong Wang, Chunbo Chen, Qingshan Geng, Lixin Cheng |
Briefings Bioinform. | 13 |
| 2025 | scMMAE: masked cross-attention network for single-cell multimodal omics fusion to enhance unimodal omicsabstractMultimodal omics provide deeper insight into the biological processes and cellular functions, especially transcriptomics and proteomics. Computational methods have been proposed for the integration of single-cell multimodal omics of transcriptomics and proteomics. However, existing methods primarily concentrate on the alignment of different omics, overlooking the unique information inherent in each omics type. Moreover, as the majority of single-cell cohorts only encompass one omics, it becomes critical to transfer the knowledge learnt from multimodal omics to enhance unimodal omics analysis. Therefore, we proposed a novel framework that leverages masked autoencoder with cross-attention mechanism, called scMMAE (single-cell multimodal masked autoencoder), to fuse multimodal omics and enhance unimodal omics analysis. scMMAE simultaneously captures both the shared features and the distinctive information of two single-cell omics modalities and transfers the knowledge to enhance single-cell transcriptome data. Comparative evaluations against benchmarking methods across various cohorts revealed a notable improvement, with an increase of up to 21% in the adjusted Rand index and up to 12% in normalized mutual information in the context of multimodal fusion. In the realm of unimodal omics, scMMAE demonstrated an overall enhancement of approximately 20% in the adjusted Rand index and nearly 10% in normalized mutual information. Other nine metrics, including the Fowlkes-Mallows index and silhouette coefficient, further underscored the high performance of scMMAE. Significantly, scMMAE exhibits an elevated level of proficiency in distinguishing between different cell types, particularly on CD4 and CD8 T cells. Availability and implementation: scMMAE source code at https://github.com/DM0815/scMMAE/. Dian Meng, Kaishen Yuan, Zitong Yu, Qin Cao, Lixin Cheng, Xubin Zheng |
Briefings Bioinform. | 6 |
| 2025 | A Q-learning-based multi-population algorithm for multi-objective distributed heterogeneous assembly no-idle flowshop scheduling with batch delivery
Zikai Zhang 0002, Qiuhua Tang, Liping Zhang 0002, Zixiang Li, Lixin Cheng |
Expert Syst. Appl. | 5 |
| 2024 | Diagnostic Prediction of portal vein thrombosis in chronic cirrhosis patients using data-driven precision medicine modelabstractBACKGROUND: Portal vein thrombosis (PVT) is a significant issue in cirrhotic patients, necessitating early detection. This study aims to develop a data-driven predictive model for PVT diagnosis in chronic hepatitis liver cirrhosis patients. METHODS: We employed data from a total of 816 chronic cirrhosis patients with PVT, divided into the Lanzhou cohort (n = 468) for training and the Jilin cohort (n = 348) for validation. This dataset encompassed a wide range of variables, including general characteristics, blood parameters, ultrasonography findings and cirrhosis grading. To build our predictive model, we employed a sophisticated stacking approach, which included Support Vector Machine (SVM), Naïve Bayes and Quadratic Discriminant Analysis (QDA). RESULTS: In the Lanzhou cohort, SVM and Naïve Bayes classifiers effectively classified PVT cases from non-PVT cases, among the top features of which seven were shared: Portal Velocity (PV), Prothrombin Time (PT), Portal Vein Diameter (PVD), Prothrombin Time Activity (PTA), Activated Partial Thromboplastin Time (APTT), age and Child-Pugh score (CPS). The QDA model, trained based on the seven shared features on the Lanzhou cohort and validated on the Jilin cohort, demonstrated significant differentiation between PVT and non-PVT cases (AUROC = 0.73 and AUROC = 0.86, respectively). Subsequently, comparative analysis showed that our QDA model outperformed several other machine learning methods. CONCLUSION: Our study presents a comprehensive data-driven model for PVT diagnosis in cirrhotic patients, enhancing clinical decision-making. The SVM-Naïve Bayes-QDA model offers a precise approach to managing PVT in this population. Xubin Zheng, Guole Nie, Jican Qin, Åsa M. Wheelock, Chuan-Xing Li, Lixin Cheng |
Briefings Bioinform. | 10 |
| 2024 | Accurate identification of structural variations from cancer samplesabstractStructural variations (SVs) are commonly found in cancer genomes. They can cause gene amplification, deletion and fusion, among other functional consequences. With an average read length of hundreds of kilobases, nano-channel-based optical DNA mapping is powerful in detecting large SVs. However, existing SV calling methods are not tailored for cancer samples, which have special properties such as mixed cell types and sub-clones. Here we propose the Cancer Optical Mapping for detecting Structural Variations (COMSV) method that is specifically designed for cancer samples. It shows high sensitivity and specificity in benchmark comparisons. Applying to cancer cell lines and patient samples, COMSV identifies hundreds of novel SVs per sample. Chenyang Hong, Claire Yik-Lok Chung, Alden King-Yung Leung, Delbert Almerick T. Boncan, Lixin Cheng, Kwok-Wai Lo, Paul B. S. Lai, John Wong, Jingying Zhou, Alfred Sze-Lok Cheng, Ting-Fung Chan, Kevin Y. Yip |
Briefings Bioinform. | 7 |
| 2024 | MrGPS: an m6A-related gene pair signature to predict the prognosis and immunological impact of glioma patientsabstractN6-methyladenosine (m6A) RNA methylation is the predominant epigenetic modification for mRNAs that regulates various cancer-related pathways. However, the prognostic significance of m6A modification regulators remains unclear in glioma. By integrating the TCGA lower-grade glioma (LGG) and glioblastoma multiforme (GBM) gene expression data, we demonstrated that both the m6A regulators and m6A-target genes were associated with glioma prognosis and activated various cancer-related pathways. Then, we paired m6A regulators and their target genes as m6A-related gene pairs (MGPs) using the iPAGE algorithm, among which 122 MGPs were significantly reversed in expression between LGG and GBM. Subsequently, we employed LASSO Cox regression analysis to construct an MGP signature (MrGPS) to evaluate glioma prognosis. MrGPS was independently validated in CGGA and GEO glioma cohorts with high accuracy in predicting overall survival. The average area under the receiver operating characteristic curve (AUC) at 1-, 3- and 5-year intervals were 0.752, 0.853 and 0.831, respectively. Combining clinical factors of age and radiotherapy, the AUC of MrGPS was much improved to around 0.90. Furthermore, CIBERSORT and TIDE algorithms revealed that MrGPS is indicative for the immune infiltration level and the response to immune checkpoint inhibitor therapy in glioma patients. In conclusion, our study demonstrated that m6A methylation is a prognostic factor for glioma and the developed prognostic model MrGPS holds potential as a valuable tool for enhancing patient management and facilitating accurate prognosis assessment in cases of glioma. Fengxia Yang, Pengfei Zhao 0018, Nana Jin, Qingshan Geng, Lixin Cheng |
Briefings Bioinform. | 9 |
| 2024 | DeepLocRNA: an interpretable deep learning model for predicting RNA subcellular localization with domain-specific transfer-learningabstractMOTIVATION: Accurate prediction of RNA subcellular localization plays an important role in understanding cellular processes and functions. Although post-transcriptional processes are governed by trans-acting RNA binding proteins (RBPs) through interaction with cis-regulatory RNA motifs, current methods do not incorporate RBP-binding information. RESULTS: In this article, we propose DeepLocRNA, an interpretable deep-learning model that leverages a pre-trained multi-task RBP-binding prediction model to predict the subcellular localization of RNA molecules via fine-tuning. We constructed DeepLocRNA using a comprehensive dataset with variant RNA types and evaluated it on the held-out dataset. Our model achieved state-of-the-art performance in predicting RNA subcellular localization in mRNA and miRNA. It has also demonstrated great generalization capabilities, performing well on both human and mouse RNA. Additionally, a motif analysis was performed to enhance the interpretability of the model, highlighting signal factors that contributed to the predictions. The proposed model provides general and powerful prediction abilities for different RNA types and species, offering valuable insights into the localization patterns of RNA molecules and contributing to our understanding of cellular processes at the molecular level. A user-friendly web server is available at: https://biolib.com/KU/DeepLocRNA/. Jun Wang 0161, Marc Horlacher, Lixin Cheng, Ole Winther |
Bioinform. | 3 |
| 2024 | Production costs and total completion time minimization for three-stage mixed-model assembly job shop scheduling with lot streaming and batch transfer
Lixin Cheng, Qiuhua Tang, Liping Zhang 0002 |
Eng. Appl. Artif. Intell. | 1 |
| 2024 | scCaT: An explainable capsulating architecture for sepsis diagnosis transferring from single-cell RNA sequencingabstractSepsis is a life-threatening condition characterized by an exaggerated immune response to pathogens, leading to organ damage and high mortality rates in the intensive care unit. Although deep learning has achieved impressive performance on prediction and classification tasks in medicine, it requires large amounts of data and lacks explainability, which hinder its application to sepsis diagnosis. We introduce a deep learning framework, called scCaT, which blends the capsulating architecture with Transformer to develop a sepsis diagnostic model using single-cell RNA sequencing data and transfers it to bulk RNA data. The capsulating architecture effectively groups genes into capsules based on biological functions, which provides explainability in encoding gene expressions. The Transformer serves as a decoder to classify sepsis patients and controls. Our model achieves high accuracy with an AUROC of 0.93 on the single-cell test set and an average AUROC of 0.98 on seven bulk RNA cohorts. Additionally, the capsules can recognize different cell types and distinguish sepsis from control samples based on their biological pathways. This study presents a novel approach for learning gene modules and transferring the model to other data types, offering potential benefits in diagnosing rare diseases with limited subjects. Xubin Zheng, Dian Meng, Wan-Ki Wong, Ka-Ho To, Lei Zhu 0016, Jiafei Wu, Yining Liang, Kwong-Sak Leung, Man Hon Wong 0001, Lixin Cheng |
PLoS Comput. Biol. | 11 |
| 2023 | RNA trafficking and subcellular localization - a review of mechanisms, experimental and predictive methodologiesabstractRNA localization is essential for regulating spatial translation, where RNAs are trafficked to their target locations via various biological mechanisms. In this review, we discuss RNA localization in the context of molecular mechanisms, experimental techniques and machine learning-based prediction tools. Three main types of molecular mechanisms that control the localization of RNA to distinct cellular compartments are reviewed, including directed transport, protection from mRNA degradation, as well as diffusion and local entrapment. Advances in experimental methods, both image and sequence based, provide substantial data resources, which allow for the design of powerful machine learning models to predict RNA localizations. We review the publicly available predictive tools to serve as a guide for users and inspire developers to build more effective prediction models. Finally, we provide an overview of multimodal learning, which may provide a new avenue for the prediction of RNA localization. Jun Wang 0161, Marc Horlacher, Lixin Cheng, Ole Winther |
Briefings Bioinform. | 3 |
| 2023 | GPGPS: a robust prognostic gene pair signature of glioma ensembling IDH mutation and 1p/19q co-deletionabstractMOTIVATION: Many studies have shown that IDH mutation and 1p/19q co-deletion can serve as prognostic signatures of glioma. Although these genetic variations affect the expression of one or more genes, the prognostic value of gene expression related to IDH and 1p/19q status is still unclear. RESULTS: We constructed an ensemble gene pair signature for the risk evaluation and survival prediction of glioma based on the prior knowledge of the IDH and 1p/19q status. First, we separately built two gene pair signatures IDH-GPS and 1p/19q-GPS and elucidated that they were useful transcriptome markers projecting from corresponding genome variations. Then, the gene pairs in these two models were assembled to develop an integrated model named Glioma Prognostic Gene Pair Signature (GPGPS), which demonstrated high area under the curves (AUCs) to predict 1-, 3- and 5-year overall survival (0.92, 0.88 and 0.80) of glioma. GPGPS was superior to the single GPSs and other existing prognostic signatures (avg AUC = 0.70, concordance index = 0.74). In conclusion, the ensemble prognostic signature with 10 gene pairs could serve as an independent predictor for risk stratification and survival prediction in glioma. This study shed light on transferring knowledge from genetic alterations to expression changes to facilitate prognostic studies. AVAILABILITY AND IMPLEMENTATION: Codes are available at https://github.com/Kimxbzheng/GPGPS.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Lixin Cheng, Xubin Zheng, Pengfei Zhao 0018, Qingshan Geng |
Bioinform. | 1 |
| 2023 | bvnGPS: a generalizable diagnostic model for acute bacterial and viral infection using integrative host transcriptomics and pretrained neural networksabstractMOTIVATION: The confusion of acute inflammation infected by virus and bacteria or noninfectious inflammation will lead to missing the best therapy occasion resulting in poor prognoses. The diagnostic model based on host gene expression has been widely used to diagnose acute infections, but the clinical usage was hindered by the capability across different samples and cohorts due to the small sample size for signature training and discovery. RESULTS: Here, we construct a large-scale dataset integrating multiple host transcriptomic data and analyze it using a sophisticated strategy which removes batch effect and extracts the common information from different cohorts based on the relative expression alteration of gene pairs. We assemble 2680 samples across 16 cohorts and separately build gene pair signature (GPS) for bacterial, viral, and noninfected patients. The three GPSs are further assembled into an antibiotic decision model (bacterial-viral-noninfected GPS, bvnGPS) using multiclass neural networks, which is able to determine whether a patient is bacterial infected, viral infected, or noninfected. bvnGPS can distinguish bacterial infection with area under the receiver operating characteristic curve (AUC) of 0.953 (95% confidence interval, 0.948-0.958) and viral infection with AUC of 0.956 (0.951-0.961) in the test set (N = 760). In the validation set (N = 147), bvnGPS also shows strong performance by attaining an AUC of 0.988 (0.978-0.998) on bacterial-versus-other and an AUC of 0.994 (0.984-1.000) on viral-versus-other. bvnGPS has the potential to be used in clinical practice and the proposed procedure provides insight into data integration, feature selection and multiclass classification for host transcriptomics data. AVAILABILITY AND IMPLEMENTATION: The codes implementing bvnGPS are available at https://github.com/Ritchiegit/bvnGPS. The construction of iPAGE algorithm and the training of neural network was conducted on Python 3.7 with Scikit-learn 0.24.1 and PyTorch 1.7. The visualization of the results was implemented on R 4.2, Python 3.7, and Matplotlib 3.3.4. Qizhi Li, Xubin Zheng, Jize Xie, Man Hon Wong 0001, Kwong-Sak Leung, Shuai Li 0010, Qingshan Geng, Lixin Cheng |
Bioinform. | 10 |
| 2023 | Deciphering associations between gut microbiota and clinical factors using microbial modulesabstractMOTIVATION: Human gut microbiota plays a vital role in maintaining body health. The dysbiosis of gut microbiota is associated with a variety of diseases. It is critical to uncover the associations between gut microbiota and disease states as well as other intrinsic or environmental factors. However, inferring alterations of individual microbial taxa based on relative abundance data likely leads to false associations and conflicting discoveries in different studies. Moreover, the effects of underlying factors and microbe-microbe interactions could lead to the alteration of larger sets of taxa. It might be more robust to investigate gut microbiota using groups of related taxa instead of the composition of individual taxa. RESULTS: We proposed a novel method to identify underlying microbial modules, i.e. groups of taxa with similar abundance patterns affected by a common latent factor, from longitudinal gut microbiota and applied it to inflammatory bowel disease (IBD). The identified modules demonstrated closer intragroup relationships, indicating potential microbe-microbe interactions and influences of underlying factors. Associations between the modules and several clinical factors were investigated, especially disease states. The IBD-associated modules performed better in stratifying the subjects compared with the relative abundance of individual taxa. The modules were further validated in external cohorts, demonstrating the efficacy of the proposed method in identifying general and robust microbial modules. The study reveals the benefit of considering the ecological effects in gut microbiota analysis and the great promise of linking clinical factors with underlying microbial modules. AVAILABILITY AND IMPLEMENTATION: https://github.com/rwang-z/microbial_module.git. Xubin Zheng, Fangda Song, Man Hon Wong 0001, Kwong-Sak Leung, Lixin Cheng |
Bioinform. | 6 |
| 2023 | Mathematical model and augmented simulated annealing algorithm for mixed-model assembly job shop scheduling problem with batch transfer
Lixin Cheng, Qiuhua Tang, Shengli Liu 0005, Liping Zhang 0002 |
Knowl. Based Syst. | 1 |
| 2022 | Improving bulk RNA-seq classification by transferring gene signature from single cells in acute myeloid leukemiaabstractThe advances in single-cell RNA sequencing (scRNA-seq) technologies enable the characterization of transcriptomic profiles at the cellular level and demonstrate great promise in bulk sample analysis thereby offering opportunities to transfer gene signature from scRNA-seq to bulk data. However, the gene expression signatures identified from single cells are typically inapplicable to bulk RNA-seq data due to the profiling differences of distinct sequencing technologies. Here, we propose single-cell pair-wise gene expression (scPAGE), a novel method to develop single-cell gene pair signatures (scGPSs) that were beneficial to bulk RNA-seq classification to transfer knowledge across platforms. PAGE was adopted to tackle the challenge of profiling differences. We applied the method to acute myeloid leukemia (AML) and identified the scGPS from mouse scRNA-seq that allowed discriminating between AML and control cells. The scGPS was validated in bulk RNA-seq datasets and demonstrated better performance (average area under the curve [AUC] = 0.96) than the conventional gene expression strategies (average AUC$\le$ 0.88) suggesting its potential in disclosing the molecular mechanism of AML. The scGPS also outperformed its bulk counterpart, which highlighted the benefit of gene signature transfer. Furthermore, we confirmed the utility of scPAGE in sepsis as an example of other disease scenarios. scPAGE leveraged the advantages of single-cell profiles to enhance the analysis of bulk samples revealing great potential of transferring knowledge from single-cell to bulk transcriptome studies. Xubin Zheng, Shibiao Wan, Fangda Song, Man Hon Wong 0001, Kwong-Sak Leung, Lixin Cheng |
Briefings Bioinform. | 8 |
| 2022 | meGPS: a multi-omics signature for hepatocellular carcinoma detection integrating methylome and transcriptome dataabstractMOTIVATION: Hepatocellular carcinoma (HCC) is a primary malignancy with a poor prognosis. Recently, multi-omics molecular-level measurement enables HCC diagnosis and prognosis prediction, which is crucial for early intervention of personalized therapy to diminish mortality. Here, we introduce a novel strategy utilizing DNA methylation and RNA expression data to achieve a multi-omics gene pair signature (GPS) for HCC discrimination. RESULTS: The immune genes with negative correlations between expression and promoter methylation are enriched in the highly connected cancer-related pathway network, which are considered as the candidates for HCC detection. After that, we separately construct a methylation GPS (mGPS) and an expression GPS (eGPS), and then assemble them as a meGPS with five gene pairs, in which the significant methylation and expression changes occur between HCC tumor and non-tumor groups. Reliable performance has been validated by independent tissue (age, gender and etiology) and blood datasets. This study proposes a procedure for multi-omics GPS identification and develops a novel HCC signature using both methylome and transcriptome data, suggesting potential molecular targets for the detection and therapy of HCC. AVAILABILITY AND IMPLEMENTATION: Models are available at https://github.com/bioinformaticStudy/meGPS.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xubin Zheng, Kwong-Sak Leung, Man Hon Wong 0001, Stephen Kwok-Wing Tsui, Lixin Cheng |
Bioinform. | 6 |
| 2022 | A Robust and Generalizable Immune-Related Signature for Sepsis DiagnosticsabstractHigh-throughput sequencing can detect tens of thousands of genes in parallel, providing opportunities for improving the diagnostic accuracy of multiple diseases including sepsis, which is an aggressive inflammatory response to infection that can cause organ failure and death. Early screening of sepsis is essential in clinic, but no effective diagnostic biomarkers are available yet. Here, we present a novel method, Recurrent Logistic Regression, to identify diagnostic biomarkers for sepsis from the blood transcriptome data. A panel including five immune-related genes, LRRN3, IL2RB, FCER1A, TLR5, and S100A12, are determined as diagnostic biomarkers (LIFTS) for sepsis. LIFTS discriminates patients with sepsis from normal controls in high accuracy (AUROC = 0.9959 on average; IC = [0.9722-1.0]) on nine validation cohorts across three independent platforms, which outperforms existing markers. Our analysis determined an accurate prediction model and reproducible transcriptome biomarkers that can lay a foundation for clinical diagnostic tests and biological mechanistic studies. Yueran Yang, Yu Zhang 0151, Shuai Li 0010, Xubin Zheng, Man Hon Wong 0001, Kwong-Sak Leung, Lixin Cheng |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2021 | A network-based algorithm for the identification of moonlighting noncoding RNAs and its application in sepsisabstractMoonlighting proteins provide more options for cells to execute multiple functions without increasing the genome and transcriptome complexity. Although there have long been calls for computational methods for the prediction of moonlighting proteins, no method has been designed for determining moonlighting long noncoding ribonucleicacidz (RNAs) (mlncRNAs). Previously, we developed an algorithm MoonFinder for the identification of mlncRNAs at the genome level based on the functional annotation and interactome data of lncRNAs and proteins. Here, we update MoonFinder to MoonFinder v2.0 by providing an extensive framework for the detection of protein modules and the establishment of RNA-module associations in human. A novel measure, moonlighting coefficient, was also proposed to assess the confidence of an ncRNA acting in a moonlighting manner. Moreover, we explored the expression characteristics of mlncRNAs in sepsis, in which we found that mlncRNAs tend to be upregulated and differentially expressed. Interestingly, the mlncRNAs are mutually exclusive in terms of coexpression when compared to the other lncRNAs. Overall, MoonFinder v2.0 is dedicated to the prediction of human mlncRNAs and thus bears great promise to serve as a valuable R package for worldwide research communities (https://cran.r-project.org/web/packages/MoonFinder/index.html). Also, our analyses provide the first attempt to characterize mlncRNA expression and coexpression properties in adult sepsis patients, which will facilitate the understanding of the interaction and expression patterns of mlncRNAs. Xueyan Liu 0011, Sheng Liu 0022, Yonglun Luo, Kwong-Sak Leung, Lixin Cheng |
Briefings Bioinform. | 8 |
| 2019 | Exploiting locational and topological overlap model to identify modules in protein interaction networksabstractBACKGROUND: Clustering molecular network is a typical method in system biology, which is effective in predicting protein complexes or functional modules. However, few studies have realized that biological molecules are spatial-temporally regulated to form a dynamic cellular network and only a subset of interactions take place at the same location in cells. RESULTS: In this study, considering the subcellular localization of proteins, we first construct a co-localization human protein interaction network (PIN) and systematically investigate the relationship between subcellular localization and biological functions. After that, we propose a Locational and Topological Overlap Model (LTOM) to preprocess the co-localization PIN to identify functional modules. LTOM requires the topological overlaps, the common partners shared by two proteins, to be annotated in the same localization as the two proteins. We observed the model has better correspondence with the reference protein complexes and shows more relevance to cancers based on both human and yeast datasets and two clustering algorithms, ClusterONE and MCL. CONCLUSION: Taking into consideration of protein localization and topological overlap can improve the performance of module detection from protein interaction networks. Lixin Cheng, Dong Wang 0011, Kwong-Sak Leung |
BMC Bioinform. | 1 |
| 2019 | Predicting associations among drugs, targets and diseases by tensor decomposition for drug repositioningabstractBACKGROUND: Development of new drugs is a time-consuming and costly process, and the cost is still increasing in recent years. However, the number of drugs approved by FDA every year per dollar spent on development is declining. Drug repositioning, which aims to find new use of existing drugs, attracts attention of pharmaceutical researchers due to its high efficiency. A variety of computational methods for drug repositioning have been proposed based on machine learning approaches, network-based approaches, matrix decomposition approaches, etc. RESULTS: We propose a novel computational method for drug repositioning. We construct and decompose three-dimensional tensors, which consist of the associations among drugs, targets and diseases, to derive latent factors reflecting the functional patterns of the three kinds of entities. The proposed method outperforms several baseline methods in recovering missing associations. Most of the top predictions are validated by literature search and computational docking. Latent factors are used to cluster the drugs, targets and diseases into functional groups. Topological Data Analysis (TDA) is applied to investigate the properties of the clusters. We find that the latent factors are able to capture the functional patterns and underlying molecular mechanisms of drugs, targets and diseases. In addition, we focus on repurposing drugs for cancer and discover not only new therapeutic use but also adverse effects of the drugs. In the in-depth study of associations among the clusters of drugs, targets and cancer subtypes, we find there exist strong associations between particular clusters. CONCLUSIONS: The proposed method is able to recover missing associations, discover new predictions and uncover functional clusters of drugs, targets and diseases. The clustering of drugs, targets and diseases, as well as the associations among the clusters, provides a new guiding framework for drug repositioning. Shuai Li 0010, Lixin Cheng, Man Hon Wong 0001, Kwong-Sak Leung |
BMC Bioinform. | 3 |
| 2018 | Identification and characterization of moonlighting long non-coding RNAs based on RNA and protein interactomeabstractMotivation: Moonlighting proteins are a class of proteins having multiple distinct functions, which play essential roles in a variety of cellular and enzymatic functioning systems. Although there have long been calls for computational algorithms for the identification of moonlighting proteins, research on approaches to identify moonlighting long non-coding RNAs (lncRNAs) has never been undertaken. Here, we introduce a novel methodology, MoonFinder, for the identification of moonlighting lncRNAs. MoonFinder is a statistical algorithm identifying moonlighting lncRNAs without a priori knowledge through the integration of protein interactome, RNA-protein interactions and functional annotation of proteins. Results: We identify 155 moonlighting lncRNA candidates and uncover that they are a distinct class of lncRNAs characterized by specific sequence and cellular localization features. The non-coding genes that transcript moonlighting lncRNAs tend to have shorter but more exons and the moonlighting lncRNAs have a variable localization pattern with a high chance of residing in the cytoplasmic compartment in comparison to the other lncRNAs. Moreover, moonlighting lncRNAs and moonlighting proteins are rather mutually exclusive in terms of both their direct interactions and interacting partners. Our results also shed light on how the moonlighting candidates and their interacting proteins implicated in the formation and development of cancers and other diseases. Availability and implementation: The code implementing MoonFinder is supplied as an R package in the supplementary material. Supplementary information: Supplementary data are available at Bioinformatics online. Lixin Cheng, Kwong-Sak Leung |
Bioinform. | 1 |
| 2010 | Extracting consistent knowledge from highly inconsistent cancer gene data sourcesabstractBACKGROUND: Hundreds of genes that are causally implicated in oncogenesis have been found and collected in various databases. For efficient application of these abundant but diverse data sources, it is of fundamental importance to evaluate their consistency. RESULTS: First, we showed that the lists of cancer genes from some major data sources were highly inconsistent in terms of overlapping genes. In particular, most cancer genes accumulated in previous small-scale studies could not be rediscovered in current high-throughput genome screening studies. Then, based on a metric proposed in this study, we showed that most cancer gene lists from different data sources were highly functionally consistent. Finally, we extracted functionally consistent cancer genes from various data sources and collected them in our database F-Census. CONCLUSIONS: Although they have very low gene overlapping, most cancer gene data sources are highly consistent at the functional level, which indicates that they can separately capture partial genes in a few key pathways associated with cancer. Our results suggest that the sample sizes currently used for cancer studies might be inadequate for consistently capturing individual cancer genes, but could be sufficient for finding a number of cancer genes that could represent functionally most cancer genes. The F-Census database provides biologists with a useful tool for browsing and extracting functionally consistent cancer genes from various data sources. Ruihong Wu, Yuannv Zhang, Wenyuan Zhao, Lixin Cheng, Yunyan Gu, Lin Zhang 0057, Jing Wang 0004, Jing Zhu 0004, Zheng Guo 0002 |
BMC Bioinform. | 5 |
| 2009 | An approach to synthesis delay semantics in VHDLabstractThe detailed timing information obtained from design iteration can only be added into synthesis process manually. A new strategy for automatic synthesis of backing marked timing information is presented in this paper. As the carrier of backing marked timing information, the delay statements in VHDL is synthesized in scheduling process. And a new scheduling algorithm is presented in this paper. The delay statements are considered to be delay time constraints by the algorithm. Scheduling data structure: DTC_DFG(delay time constrained data flow graph) is constructed and scheduled. A heuristic method is presented in this paper for the scheduling algorithm to jump out local optimization and reach global optimization under polynomial timing complexity. Experimental results are also presented in this paper. With the synthesis of delay statements, the backing marked timing information can be synthesized automatically, and the behavioral timing in design source can be synthesized directly, thus the timing consistency between synthesis result and simulation result can be improved. The delay statements then can be a convenient means to set timing constraints, thus the manual interruption in synthesis process can be decreased; the design efficiency can be improved greatly. Lixin Cheng, Jinian Bian, Yunyun Liu |
CAD/Graphics | 1 |
| 2007 | A New ILP Based Approach to Schedule and Bind SimultaneouslyabstractTo construct complete design space for high-level synthesis and obtain globally optimized solution. An approach that fulfills operation scheduling and binding simultaneously is presented in this paper. The improved ASAP and ALAP scheduling algorithms are employed to compute the actually soonest and latest control steps with the condition of referring binding information. A basic model based on ILP is presented to solve the minimum resource occupance scheduling and binding under timing constraints. The means to solve minimum execution control steps scheduling and binding is also investigated. According to experiments, when scheduling fulfilled by this approach is accomplished, the binding results are obtained meanwhile. The executing time required by the approach is very low. And the feature of ILP has guaranteed the global optimization of results. Based on plenty of benchmark experiments, it can be concluded that the approach presented in this paper is effective and efficient. Lixin Cheng, Junbo Xu, Guochang Gu |
CAD/Graphics | 1 |
| 2003 | Information Retrieval Using Deep Natural Language Processing
Rossitza Setchi, Qiao Tang, Lixin Cheng |
KES | 3 |
| 1997 | Notes on Riesz's theorem on fuzzy measure space
Minghu Ha 0001, Lixin Cheng, Xizhao Wang |
Fuzzy Sets Syst. | 2 |