Xubin Zheng

dblp:164/5127 · DBLP profile ↗
← Back
21ranked-venue papers
1as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 1 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GROVER: Graph-guided Representation of Omics and Vision with Expert Regulation for Adaptive Spatial Multi-omics Fusion
abstract
Effectively modeling multimodal spatial omics data is critical for understanding tissue complexity and underlying biological mechanisms. While spatial transcriptomics, proteomics, and epigenomics capture molecular features, they lack pathological morphological context. Integrating these omics with histopathological images is thus critical for comprehensive disease tissue analysis. However, substantial heterogeneity across omics, imaging, and spatial modalities poses significant challenges. Naive fusion of semantically distinct sources often leads to ambiguous representations. Additionally, the resolution mismatch between high-resolution histology images and lower-resolution sequencing spots complicates spatial alignment. Biological perturbations during sample preparation further distort modality-specific signals, hindering accurate integration. To address these challenges, we propose Graph-guided Representation of Omics and Vision with Expert Regulation for Adaptive Spatial Multi-omics Fusion (GROVER), a novel framework for adaptive integration of spatial multi-omics data. GROVER leverages a Graph Convolutional Network encoder based on Kolmogorov–Arnold Networks to capture the nonlinear dependencies between each modality and its associated spatial structure, thereby producing expressive, modality-specific embeddings. To align these representations, we introduce a spot-feature-pair contrastive learning strategy that explicitly optimizes the correspondence across modalities at each spot. Furthermore, we design a dynamic expert routing mechanism that adaptively selects informative modalities for each spot while suppressing noisy or low-quality inputs. Experiments on real-world spatial omics datasets demonstrate that GROVER outperforms state-of-the-art baselines, providing a robust and reliable solution for multimodal integration.
Yongjun Xiao, Dian Meng, Xinlei Huang, Yanran Liu, Shiwei Ruan, Ziyue Qiao, Xubin Zheng
AAAI7
2026 STAHD: a scalable and accurate method to detect spatial domains in high-resolution spatial transcriptomics data
abstract
MOTIVATION: Spatial transcriptomics (ST) enables the study of spatial heterogeneity in tissues. However, current methods struggle with large-scale, high-resolution data, leading to reduced efficiency and accuracy in detecting spatial domains. A scalable, precise solution is urgently needed. RESULTS: We present STAHD, a scalable and efficient framework for spatial domain detection in ST data. Combining a graph attention autoencoder with multilevel k-way graph partitioning, STAHD decomposes large graphs into compact subgraphs and generates low-dimensional embeddings. This improves computational efficiency and clustering accuracy. Benchmarks on human and mouse datasets show STAHD outperforms existing methods and accurately reveals spatially distinct tumor microenvironments and functional regions. AVAILABILITY AND IMPLEMENTATION: Source code and data are available at: https://github.com/Little-Eel/STAHD.
Zhihua Du, Qiyi Chen, Yuehua Ou, Xinlei Huang, Xubin Zheng
Bioinform.7
2025 PRAGA: Prototype-aware Graph Adaptive Aggregation for Spatial Multi-modal Omics Analysis
abstract
Spatial multi-modal omics technology, highlighted by Nature Methods as an advanced biological technique in 2023, plays a critical role in resolving biological regulatory processes with spatial context. Recently, graph neural networks based on K-nearest neighbor (KNN) graphs have gained prominence in spatial multi-modal omics methods due to their ability to model semantic relations between sequencing spots. However, the fixed KNN graph fails to capture the latent semantic relations hidden by the inevitable data perturbations during the biological sequencing process, resulting in the loss of semantic information. In addition, the common lack of spot annotation and class number priors in practice further hinders the optimization of spatial multi-modal omics models. Here, we propose a novel spatial multi-modal omics resolved framework, termed Prototype-aware Graph Adaptative Aggregation for Spatial Multi-modal Omics Analysis (PRAGA). PRAGA constructs a dynamic graph to capture latent semantic relations and comprehensively integrate spatial information and feature semantics. The learnable graph structure can also denoise perturbations by learning cross-modal knowledge. Moreover, a dynamic prototype contrastive learning is proposed based on the dynamic adaptability of Bayesian Gaussian Mixture Models to optimize the multi-modal omics representations for unknown biological priors. Quantitative and qualitative experiments on simulated and real datasets with 7 competing methods demonstrate the superior performance of PRAGA.
Xinlei Huang, Zhiqi Ma, Dian Meng, Yanran Liu, Shiwei Ruan, Qingqiang Sun, Xubin Zheng, Ziyue Qiao
AAAI7
2025 Single-View Graph Contrastive Learning with Soft Neighborhood Awareness
abstract
Most graph contrastive learning (GCL) methods heavily rely on cross-view contrast, thus facing several concomitant challenges, such as the complexity of designing effective augmentations, the potential for information loss between views, and increased computational costs. To mitigate reliance on cross-view contrasts, we propose SIGNA, a novel single-view graph contrastive learning framework. Regarding the inconsistency between structural connection and semantic similarity of neighborhoods, we resort to soft neighborhood awareness for GCL. Specifically, we leverage dropout to obtain structurally-related yet randomly-noised embedding pairs for neighbors, which serve as potential positive samples. At each epoch, the role of partial neighbors is switched from positive to negative, leading to probabilistic neighborhood contrastive learning effect. Moreover, we propose a normalized Jensen-Shannon divergence estimator for a better effect of contrastive learning. Experiments on diverse node-level tasks demonstrate that our simple single-view GCL framework consistently outperforms existing methods by margins of up to 21.74% (PPI). In particular, with soft neighborhood awareness, SIGNA can adopt MLPs instead of complicated GCNs as the encoder in transductive learning tasks, thus speeding up its inference process by 109× to 331×.
Qingqiang Sun, Chaoqi Chen, Ziyue Qiao, Xubin Zheng
AAAI4
2025 BioLinkGPT: Predicting Missing TF-Target Gene Interactions Using Graph Neural Networks with Large Language Model
abstract
Predicting unknown transcription factor-target gene (TF-target gene) interactions based on gene regulatory networks (GRNs) is critical for understanding cellular functions and biological discovery. Existing computational methods focus on the statistic corelation and neglect the semantic information of biomedical knowledge. To address this, we propose BioLinkGPT, a novel framework that integrates the robust textual comprehension capabilities of Large Language Models (LLMs) with graph neural networks (GNNs) to capture structural information of GRN. BioLinkGPT leverages gene information and the known TF-target gene relationships from literature to predict unknown TF-target gene interactions and improve its performance via a two-stage instruction fine-tuning strategy. We construct a comprehensive GRN dataset focused on human infectious diseases, and experimental results show that BioLinkGPT significantly outperforms existing baselines in link prediction metrics. We validate the model's accurate identification of known regulatory relationships and find that BioLinkGPT successfully uncover potential regulatory interactions with high confidence and biological relevance. BioLinkGPT offers an efficient and precise computational tool to infer TF-target gene interactions, accelerating research on infectious disease mechanisms and drug target development.
Zhihua Du, Weiliang Huang, Yanran Liu, Qiyi Chen, Rui Luo 0002, Xubin Zheng
BIBM6
2025 Predicting Gene Regulatory Relationship in Cancer Using LLM and Graph Neural Network from Known Regulations
abstract
Predicting gene regulatory relationship is crucial to reveal cancer mechanism and develop targeted therapies. However, current computational methods typically estimate gene-gene associations without regulatory directions and only based on statistical significance, often lacking accuracy after validation by biological experiments. Therefore, we introduce a pipeline named LiGRNet that constructs gene regulatory networks as signed directed graphs using large language models (LLMs) from the literature and predicts potential gene regulatory relationship with magnetic signed graph neural networks (MSGNNs). First, we teach LLM to extract the known gene regulatory relationships proved by biological experiments through prompt. Then, a signed directed graph was constructed and learnt by MSGNN with a link prediction framework based on complex-valued gene embeddings. The potential regulatory relationships were predicted with direction and sign. We apply our pipeline in colorectal cancer, liver cancer, and colorectal liver metastasis including$11,000,19,000$, and 1,300 literature, respectively. The pipeline achieve$84.5 \%, 80.7 \%, 93.3 \%$accuracy in predicting unknown gene regulatory relationship. External cell line data also provide evidence for the predicted gene regulation in these cancers.
Yanran Liu, Dian Meng, Xinlei Huang, Shiwei Ruan, Yinghua Chen, Rui Luo 0002, Xubin Zheng
BIBM10
2025 SVP: Support Vector Gene Pair as Biomarker Selection for Breast Cancer
abstract
mRNA expression is vital for understanding the mechanisms of breast cancer initiation and metastasis, with mRNA-based biomarkers closely linked to tumor grade, metastasis, and prognosis. However, gene expression analyses often suffer from batch effects, high inter-sample variability, and poor cross-cohort generalizability, which can obscure true biological signals. To overcome these limitations, we developed the Support Vector Gene Pair (SVP) algorithm, which emphasizes within-sample relative gene pair relationships rather than absolute expression levels. This design effectively mitigates inter-sample interference and enhances robustness across heterogeneous datasets. By in-corporating intra- and inter-sample variability, SVP accurately identifies informative gene pair features. Applied to 16 GEO datasets (852 samples) and one TCGA dataset, SVP discovered 23 breast cancer-related gene pairs and outperformed traditional Differentially Expressed Gene (DEG) and unadjusted statistical approaches. Comparative analyses showed consistent performance gains, with average improvements of 8% in AUC and 6% in PRC. SVP offers a robust and interpretable framework for breast cancer biomarker discovery, holding strong promise for translational applications in precision medicine.
Qiyi Chen, Dian Meng, Yanran Liu, Jinchao Feng, Xubin Zheng
BIBM6
2025 Learn from Balance: Rectifying Knowledge Transfer for Long-Tailed Scenarios
abstract
Knowledge Distillation (KD) transfers knowledge from a large pre-trained teacher network to a compact and efficient student network, making it suitable for deployment on resource-limited media terminals. However, traditional KD methods require balanced data to ensure robust training, which is often unavailable in practical applications. In such scenarios, a few head categories occupy a substantial proportion of examples. This imbalance biases the trained teacher network towards the head categories, resulting in severe performance degradation on the less represented tail categories for both the teacher and student networks. In this paper, we propose a novel framework called Knowledge Rectification Distillation (KRDistill) to address the imbalanced knowledge inherited in the teacher network through the incorporation of the balanced category priors. Furthermore, we rectify the biased predictions produced by the teacher network, particularly focusing on the tail categories. Consequently, the teacher network can provide balanced and accurate knowledge to train a reliable student network. Intensive experiments conducted on various long-tailed datasets demonstrate that our KRDistill can effectively train reliable student networks in realistic scenarios of data imbalance.
Xinlei Huang, Jialiang Tang, Xubin Zheng, Jinjia Zhou, Wenxin Yu 0001, Ning Jiang 0002
ICASSP3
2025 AdaMHF: Adaptive Multimodal Hierarchical Fusion for Survival Prediction
abstract
The integration of pathologic images and genomic data for survival analysis has gained increasing attention with advances in multimodal learning. However, current methods often ignore biological characteristics, such as heterogeneity and sparsity, both within and across modalities, ultimately limiting their adaptability to clinical practice. To address these challenges, we propose AdaMHF: Adaptive Multimodal Hierarchical Fusion, a framework designed for efficient, comprehensive, and tailored feature extraction and fusion. AdaMHF is specifically adapted to the uniqueness of medical data, enabling accurate predictions with minimal resource consumption, even under challenging scenarios with missing modalities. Initially, AdaMHF employs an experts expansion and residual structure to activate specialized experts for extracting heterogeneous and sparse features. Extracted tokens undergo refinement via selection and aggregation, reducing the weight of non-dominant features while preserving comprehensive information. Subsequently, the encoded features are hierarchically fused, allowing multi-grained interactions across modalities to be captured. Furthermore, we introduce a survival prediction benchmark designed to resolve scenarios with missing modalities, mirroring real-world clinical conditions. Extensive experiments on TCGA datasets demonstrate that AdaMHF surpasses current state-of-the-art (SOTA) methods, showcasing exceptional performance in both complete and incomplete modality settings. Code is available in AdaMHF.
Shuaiyu Zhang, Xun Lin, Rongxiang Zhang, Yong Xu 0001, Tao Tan 0002, Xubin Zheng, Zitong Yu
ICME7
2025 Thread the Needle: Genomics-Guided Prompt-Bridged Attention Model for Survival Prediction of Glioma Based on MRI Images
Xubin Zheng, Xiongri Shen, Jiaqi Wang 0003, Zhenxi Song, Zhiguo Zhang 0001
MICCAI (7)2
2025 scMMAE: masked cross-attention network for single-cell multimodal omics fusion to enhance unimodal omics
abstract
Multimodal omics provide deeper insight into the biological processes and cellular functions, especially transcriptomics and proteomics. Computational methods have been proposed for the integration of single-cell multimodal omics of transcriptomics and proteomics. However, existing methods primarily concentrate on the alignment of different omics, overlooking the unique information inherent in each omics type. Moreover, as the majority of single-cell cohorts only encompass one omics, it becomes critical to transfer the knowledge learnt from multimodal omics to enhance unimodal omics analysis. Therefore, we proposed a novel framework that leverages masked autoencoder with cross-attention mechanism, called scMMAE (single-cell multimodal masked autoencoder), to fuse multimodal omics and enhance unimodal omics analysis. scMMAE simultaneously captures both the shared features and the distinctive information of two single-cell omics modalities and transfers the knowledge to enhance single-cell transcriptome data. Comparative evaluations against benchmarking methods across various cohorts revealed a notable improvement, with an increase of up to 21% in the adjusted Rand index and up to 12% in normalized mutual information in the context of multimodal fusion. In the realm of unimodal omics, scMMAE demonstrated an overall enhancement of approximately 20% in the adjusted Rand index and nearly 10% in normalized mutual information. Other nine metrics, including the Fowlkes-Mallows index and silhouette coefficient, further underscored the high performance of scMMAE. Significantly, scMMAE exhibits an elevated level of proficiency in distinguishing between different cell types, particularly on CD4 and CD8 T cells. Availability and implementation: scMMAE source code at https://github.com/DM0815/scMMAE/.
Dian Meng, Kaishen Yuan, Zitong Yu, Qin Cao, Lixin Cheng, Xubin Zheng
Briefings Bioinform.7
2024 Diagnostic Prediction of portal vein thrombosis in chronic cirrhosis patients using data-driven precision medicine model
abstract
BACKGROUND: Portal vein thrombosis (PVT) is a significant issue in cirrhotic patients, necessitating early detection. This study aims to develop a data-driven predictive model for PVT diagnosis in chronic hepatitis liver cirrhosis patients. METHODS: We employed data from a total of 816 chronic cirrhosis patients with PVT, divided into the Lanzhou cohort (n = 468) for training and the Jilin cohort (n = 348) for validation. This dataset encompassed a wide range of variables, including general characteristics, blood parameters, ultrasonography findings and cirrhosis grading. To build our predictive model, we employed a sophisticated stacking approach, which included Support Vector Machine (SVM), Naïve Bayes and Quadratic Discriminant Analysis (QDA). RESULTS: In the Lanzhou cohort, SVM and Naïve Bayes classifiers effectively classified PVT cases from non-PVT cases, among the top features of which seven were shared: Portal Velocity (PV), Prothrombin Time (PT), Portal Vein Diameter (PVD), Prothrombin Time Activity (PTA), Activated Partial Thromboplastin Time (APTT), age and Child-Pugh score (CPS). The QDA model, trained based on the seven shared features on the Lanzhou cohort and validated on the Jilin cohort, demonstrated significant differentiation between PVT and non-PVT cases (AUROC = 0.73 and AUROC = 0.86, respectively). Subsequently, comparative analysis showed that our QDA model outperformed several other machine learning methods. CONCLUSION: Our study presents a comprehensive data-driven model for PVT diagnosis in cirrhotic patients, enhancing clinical decision-making. The SVM-Naïve Bayes-QDA model offers a precise approach to managing PVT in this population.
Xubin Zheng, Guole Nie, Jican Qin, Åsa M. Wheelock, Chuan-Xing Li, Lixin Cheng
Briefings Bioinform.3
2024 scCaT: An explainable capsulating architecture for sepsis diagnosis transferring from single-cell RNA sequencing
abstract
Sepsis is a life-threatening condition characterized by an exaggerated immune response to pathogens, leading to organ damage and high mortality rates in the intensive care unit. Although deep learning has achieved impressive performance on prediction and classification tasks in medicine, it requires large amounts of data and lacks explainability, which hinder its application to sepsis diagnosis. We introduce a deep learning framework, called scCaT, which blends the capsulating architecture with Transformer to develop a sepsis diagnostic model using single-cell RNA sequencing data and transfers it to bulk RNA data. The capsulating architecture effectively groups genes into capsules based on biological functions, which provides explainability in encoding gene expressions. The Transformer serves as a decoder to classify sepsis patients and controls. Our model achieves high accuracy with an AUROC of 0.93 on the single-cell test set and an average AUROC of 0.98 on seven bulk RNA cohorts. Additionally, the capsules can recognize different cell types and distinguish sepsis from control samples based on their biological pathways. This study presents a novel approach for learning gene modules and transferring the model to other data types, offering potential benefits in diagnosing rare diseases with limited subjects.
Xubin Zheng, Dian Meng, Wan-Ki Wong, Ka-Ho To, Lei Zhu 0016, Jiafei Wu, Yining Liang, Kwong-Sak Leung, Man Hon Wong 0001, Lixin Cheng
PLoS Comput. Biol.1
2023 GPGPS: a robust prognostic gene pair signature of glioma ensembling IDH mutation and 1p/19q co-deletion
abstract
MOTIVATION: Many studies have shown that IDH mutation and 1p/19q co-deletion can serve as prognostic signatures of glioma. Although these genetic variations affect the expression of one or more genes, the prognostic value of gene expression related to IDH and 1p/19q status is still unclear. RESULTS: We constructed an ensemble gene pair signature for the risk evaluation and survival prediction of glioma based on the prior knowledge of the IDH and 1p/19q status. First, we separately built two gene pair signatures IDH-GPS and 1p/19q-GPS and elucidated that they were useful transcriptome markers projecting from corresponding genome variations. Then, the gene pairs in these two models were assembled to develop an integrated model named Glioma Prognostic Gene Pair Signature (GPGPS), which demonstrated high area under the curves (AUCs) to predict 1-, 3- and 5-year overall survival (0.92, 0.88 and 0.80) of glioma. GPGPS was superior to the single GPSs and other existing prognostic signatures (avg AUC = 0.70, concordance index = 0.74). In conclusion, the ensemble prognostic signature with 10 gene pairs could serve as an independent predictor for risk stratification and survival prediction in glioma. This study shed light on transferring knowledge from genetic alterations to expression changes to facilitate prognostic studies. AVAILABILITY AND IMPLEMENTATION: Codes are available at https://github.com/Kimxbzheng/GPGPS.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Lixin Cheng, Xubin Zheng, Pengfei Zhao 0018, Qingshan Geng
Bioinform.3
2023 bvnGPS: a generalizable diagnostic model for acute bacterial and viral infection using integrative host transcriptomics and pretrained neural networks
abstract
MOTIVATION: The confusion of acute inflammation infected by virus and bacteria or noninfectious inflammation will lead to missing the best therapy occasion resulting in poor prognoses. The diagnostic model based on host gene expression has been widely used to diagnose acute infections, but the clinical usage was hindered by the capability across different samples and cohorts due to the small sample size for signature training and discovery. RESULTS: Here, we construct a large-scale dataset integrating multiple host transcriptomic data and analyze it using a sophisticated strategy which removes batch effect and extracts the common information from different cohorts based on the relative expression alteration of gene pairs. We assemble 2680 samples across 16 cohorts and separately build gene pair signature (GPS) for bacterial, viral, and noninfected patients. The three GPSs are further assembled into an antibiotic decision model (bacterial-viral-noninfected GPS, bvnGPS) using multiclass neural networks, which is able to determine whether a patient is bacterial infected, viral infected, or noninfected. bvnGPS can distinguish bacterial infection with area under the receiver operating characteristic curve (AUC) of 0.953 (95% confidence interval, 0.948-0.958) and viral infection with AUC of 0.956 (0.951-0.961) in the test set (N = 760). In the validation set (N = 147), bvnGPS also shows strong performance by attaining an AUC of 0.988 (0.978-0.998) on bacterial-versus-other and an AUC of 0.994 (0.984-1.000) on viral-versus-other. bvnGPS has the potential to be used in clinical practice and the proposed procedure provides insight into data integration, feature selection and multiclass classification for host transcriptomics data. AVAILABILITY AND IMPLEMENTATION: The codes implementing bvnGPS are available at https://github.com/Ritchiegit/bvnGPS. The construction of iPAGE algorithm and the training of neural network was conducted on Python 3.7 with Scikit-learn 0.24.1 and PyTorch 1.7. The visualization of the results was implemented on R 4.2, Python 3.7, and Matplotlib 3.3.4.
Qizhi Li, Xubin Zheng, Jize Xie, Man Hon Wong 0001, Kwong-Sak Leung, Shuai Li 0010, Qingshan Geng, Lixin Cheng
Bioinform.2
2023 Deciphering associations between gut microbiota and clinical factors using microbial modules
abstract
MOTIVATION: Human gut microbiota plays a vital role in maintaining body health. The dysbiosis of gut microbiota is associated with a variety of diseases. It is critical to uncover the associations between gut microbiota and disease states as well as other intrinsic or environmental factors. However, inferring alterations of individual microbial taxa based on relative abundance data likely leads to false associations and conflicting discoveries in different studies. Moreover, the effects of underlying factors and microbe-microbe interactions could lead to the alteration of larger sets of taxa. It might be more robust to investigate gut microbiota using groups of related taxa instead of the composition of individual taxa. RESULTS: We proposed a novel method to identify underlying microbial modules, i.e. groups of taxa with similar abundance patterns affected by a common latent factor, from longitudinal gut microbiota and applied it to inflammatory bowel disease (IBD). The identified modules demonstrated closer intragroup relationships, indicating potential microbe-microbe interactions and influences of underlying factors. Associations between the modules and several clinical factors were investigated, especially disease states. The IBD-associated modules performed better in stratifying the subjects compared with the relative abundance of individual taxa. The modules were further validated in external cohorts, demonstrating the efficacy of the proposed method in identifying general and robust microbial modules. The study reveals the benefit of considering the ecological effects in gut microbiota analysis and the great promise of linking clinical factors with underlying microbial modules. AVAILABILITY AND IMPLEMENTATION: https://github.com/rwang-z/microbial_module.git.
Xubin Zheng, Fangda Song, Man Hon Wong 0001, Kwong-Sak Leung, Lixin Cheng
Bioinform.2
2022 Improving bulk RNA-seq classification by transferring gene signature from single cells in acute myeloid leukemia
abstract
The advances in single-cell RNA sequencing (scRNA-seq) technologies enable the characterization of transcriptomic profiles at the cellular level and demonstrate great promise in bulk sample analysis thereby offering opportunities to transfer gene signature from scRNA-seq to bulk data. However, the gene expression signatures identified from single cells are typically inapplicable to bulk RNA-seq data due to the profiling differences of distinct sequencing technologies. Here, we propose single-cell pair-wise gene expression (scPAGE), a novel method to develop single-cell gene pair signatures (scGPSs) that were beneficial to bulk RNA-seq classification to transfer knowledge across platforms. PAGE was adopted to tackle the challenge of profiling differences. We applied the method to acute myeloid leukemia (AML) and identified the scGPS from mouse scRNA-seq that allowed discriminating between AML and control cells. The scGPS was validated in bulk RNA-seq datasets and demonstrated better performance (average area under the curve [AUC] = 0.96) than the conventional gene expression strategies (average AUC$\le$ 0.88) suggesting its potential in disclosing the molecular mechanism of AML. The scGPS also outperformed its bulk counterpart, which highlighted the benefit of gene signature transfer. Furthermore, we confirmed the utility of scPAGE in sepsis as an example of other disease scenarios. scPAGE leveraged the advantages of single-cell profiles to enhance the analysis of bulk samples revealing great potential of transferring knowledge from single-cell to bulk transcriptome studies.
Xubin Zheng, Shibiao Wan, Fangda Song, Man Hon Wong 0001, Kwong-Sak Leung, Lixin Cheng
Briefings Bioinform.2
2022 meGPS: a multi-omics signature for hepatocellular carcinoma detection integrating methylome and transcriptome data
abstract
MOTIVATION: Hepatocellular carcinoma (HCC) is a primary malignancy with a poor prognosis. Recently, multi-omics molecular-level measurement enables HCC diagnosis and prognosis prediction, which is crucial for early intervention of personalized therapy to diminish mortality. Here, we introduce a novel strategy utilizing DNA methylation and RNA expression data to achieve a multi-omics gene pair signature (GPS) for HCC discrimination. RESULTS: The immune genes with negative correlations between expression and promoter methylation are enriched in the highly connected cancer-related pathway network, which are considered as the candidates for HCC detection. After that, we separately construct a methylation GPS (mGPS) and an expression GPS (eGPS), and then assemble them as a meGPS with five gene pairs, in which the significant methylation and expression changes occur between HCC tumor and non-tumor groups. Reliable performance has been validated by independent tissue (age, gender and etiology) and blood datasets. This study proposes a procedure for multi-omics GPS identification and develops a novel HCC signature using both methylome and transcriptome data, suggesting potential molecular targets for the detection and therapy of HCC. AVAILABILITY AND IMPLEMENTATION: Models are available at https://github.com/bioinformaticStudy/meGPS.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Xubin Zheng, Kwong-Sak Leung, Man Hon Wong 0001, Stephen Kwok-Wing Tsui, Lixin Cheng
Bioinform.2
2022 Singular value decomposition-based behavior-aware cloud service application programming interfaces recommendation for large-scale software cloud directory platforms
abstract
Summary With the development of Internet technology and the cloud service industry, an increasing number of application programming interfaces (APIs) hosted in the cloud has been made publicly available. To facilitate cloud service APIs vendors and buyers, some large‐scale software cloud directory platforms have been established. Nevertheless, it is difficult for users to choose for renting from a massive number of cloud service APIs with similar functionalities in a software cloud directory platform. Recent efforts in building cloud service APIs recommender systems can help address this challenge. Relevant existing recommendation approaches are designed based on requirement election techniques to identify users' preferences to the quality of service (QoS) of the APIs. In particular, users' preferences are mainly obtained through their self‐description, in which users sometimes cannot accurately and completely express their preferences. In this article, we propose SVD‐APIR, a singular value decomposition (SVD)‐based behavior‐aware cloud service APIs recommendation approach for large‐scale software cloud directory platforms. In SVD‐APIR, users' historical behavior information is captured and APIs' association information is analyzed to identify the users' potential preferences to the APIs with specific QoS. A unified SVD model is utilized to prioritize the users preferred APIs. Experimental evaluation results conducted on WS‐Dream dataset demonstrate the effectiveness and efficiency of the proposed approach.
Lei Wang 0042, Yunqiu Zhang, Xubin Zheng, Qi Yu 0001, Shuhan Chen, Junyao Ding
Concurr. Comput. Pract. Exp.3
2022 A Robust and Generalizable Immune-Related Signature for Sepsis Diagnostics
abstract
High-throughput sequencing can detect tens of thousands of genes in parallel, providing opportunities for improving the diagnostic accuracy of multiple diseases including sepsis, which is an aggressive inflammatory response to infection that can cause organ failure and death. Early screening of sepsis is essential in clinic, but no effective diagnostic biomarkers are available yet. Here, we present a novel method, Recurrent Logistic Regression, to identify diagnostic biomarkers for sepsis from the blood transcriptome data. A panel including five immune-related genes, LRRN3, IL2RB, FCER1A, TLR5, and S100A12, are determined as diagnostic biomarkers (LIFTS) for sepsis. LIFTS discriminates patients with sepsis from normal controls in high accuracy (AUROC = 0.9959 on average; IC = [0.9722-1.0]) on nine validation cohorts across three independent platforms, which outperforms existing markers. Our analysis determined an accurate prediction model and reproducible transcriptome biomarkers that can lay a foundation for clinical diagnostic tests and biological mechanistic studies.
Yueran Yang, Yu Zhang 0151, Shuai Li 0010, Xubin Zheng, Man Hon Wong 0001, Kwong-Sak Leung, Lixin Cheng
IEEE ACM Trans. Comput. Biol. Bioinform.4
2020 Drug2vec: A Drug Embedding Method with Drug-Drug Interaction as the Context
Xubin Zheng, Man Hon Wong 0001, Kwong-Sak Leung
EANN2