Xiaohui Niu

dblp:217/0757 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
6since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 6 since 2021
YearPublicationVenuePosition
2024 Multiview representation learning for identification of novel cancer genes and their causative biological mechanisms
abstract
Tumorigenesis arises from the dysfunction of cancer genes, leading to uncontrolled cell proliferation through various mechanisms. Establishing a complete cancer gene catalogue will make precision oncology possible. Although existing methods based on graph neural networks (GNN) are effective in identifying cancer genes, they fall short in effectively integrating data from multiple views and interpreting predictive outcomes. To address these shortcomings, an interpretable representation learning framework IMVRL-GCN is proposed to capture both shared and specific representations from multiview data, offering significant insights into the identification of cancer genes. Experimental results demonstrate that IMVRL-GCN outperforms state-of-the-art cancer gene identification methods and several baselines. Furthermore, IMVRL-GCN is employed to identify a total of 74 high-confidence novel cancer genes, and multiview data analysis highlights the pivotal roles of shared, mutation-specific, and structure-specific representations in discriminating distinctive cancer genes. Exploration of the mechanisms behind their discriminative capabilities suggests that shared representations are strongly associated with gene functions, while mutation-specific and structure-specific representations are linked to mutagenic propensity and functional synergy, respectively. Finally, our in-depth analyses of these candidates suggest potential insights for individualized treatments: afatinib could counteract many mutation-driven risks, and targeting interactions with cancer gene SRC is a reasonable strategy to mitigate interaction-induced risks for NR3C1, RXRA, HNF4A, and SP1.
Jianye Yang 0002, Haitao Fu, Fei-Yang Xue, Menglu Li, Yuyang Wu, Zhanhui Yu, Haohui Luo, Xiaohui Niu
Briefings Bioinform.9
2023 Metapath-aggregated multilevel graph embedding for miRNA‒disease association prediction
abstract
MicroRNAs (miRNAs) are crucial regulators in various diseases. The identification of associations between miRNAs and diseases could greatly facilitate the investigation of disease mechanisms and drug development. Limited by time and cost efficiency, conventional experimental techniques are inadequate for this purpose. With the extensive advance and application of deep learning, developing efficient and accurate computational models for predicting miRNA‒disease associations has a vital role and is feasible. In this study, we proposed a meta-path-aggregated multilevel graph embedding model for miRNA‒ disease association prediction. The model first calculated the multiple similarities among miRNAs and diseases, respectively. Then, the node features were extracted from similarity matrices for miRNAs and diseases. Furthermore, we integrated four types of meta-paths from the miRNA‒lncRNA‒disease heterogeneous graph and learned node embeddings by hierarchical graph attention modules. Finally, the model predicted the miRNA‒ disease associations using two-layer graph convolution networks (GCNs). Compared with six state-of-the-art models, the experimental results demonstrated that our model achieved higher prediction performance with an AUC of 0.9892 and an AUPR of 0.9898 for the 5-fold cross-validation on the HMDDv3.2 dataset. With the case study, the model’s performance was further validated, and the top 20 predicted associations could be experimentally confirmed. All in all, it implies the predictive power of our model1and the potential value in understanding disease pathology.
Jian-Ye Yang, Fei-Yang Xue, Zhan-Hui Yu, Ze-Jun Wu, Xiaohui Niu
BIBM9
2023 DRLM: A Robust Drug Representation Learning Method and its Applications
abstract
Learning representations from data is a fundamental step for machine learning. High-quality and robust drug representations can broaden the understanding of pharmacology, and improve the modeling of multiple drug-related prediction tasks, which further facilitates drug development. Although there are a number of models developed for drug representation learning from various data sources, few researches extract drug representations from gene expression profiles. Since gene expression profiles of drug-treated cells are widely used in clinical diagnosis and therapy, it is believed that leveraging them to eliminate cell specificity can promote drug representation learning. In this paper, we propose a three-stage deep learning method for drug representation learning, named DRLM, which integrates gene expression profiles of drug-related cells and the therapeutic use information of drugs. Firstly, we construct a stacked autoencoder to learn low-dimensional compact drug representations. Secondly, we utilize an iterative clustering module to reduce the negative effects of cell specificity and noise in gene expression profiles on the low-dimensional drug representations. Thirdly, a therapeutic use discriminator is designed to incorporate therapeutic use information into the drug representations. The visualization analysis of drug representations demonstrates DRLM can reduce cell specificity and integrate therapeutic use information effectively. Extensive experiments on three types of prediction tasks are conducted based on different drug representations, and they show that the drug representations learned by DRLM outperform other representations in terms of most metrics. The ablation analysis also demonstrates DRLM's effectiveness of merging the gene expression profiles with the therapeutic use information.
Haitao Fu, Cecheng Zhao, Xiaohui Niu, Wen Zhang 0008
IEEE ACM Trans. Comput. Biol. Bioinform.3
2021 A robust drug representation learning model for eliminating cell specificity in gene expression profile and its application
abstract
Learning high-quality drug representations is important for drug development and the understanding of drug action mechanisms. Leveraging the gene expression profile of drug treated cells and eliminating cell specificity can facilitate drug representation learning. In this paper, we propose a four stage deep learning model that aims for drug representation learning based on integrating gene expression profile and the therapeutic use information of drugs, abbreviated as “DGERN”. The stacked autoencoder module is employed for data dimension reduction; the iterative clustering module is used to eliminate cell specificity; the subclass pre-training module and the label classifier module are utilized to integrate the therapeutic use information of drugs into drug representations. Visualization of the drug representations proves that DGERN eliminates cell specificity and integrates the therapeutic use information of drugs effectively. The drug representations learned by DGERN are used in the subsequent and prediction tasks of drug development. In the task of predicting drug-disease associations, DGERN combined with random forest achieves the best performance reaching 0.67 on AUC, exceeding 0.60 of the second-placed one; in the drug-drug interaction prediction task, DGERN combined with random forest gets 0.73 on AUC, which is second in comparison with other drug representations.
Cecheng Zhao, Hui Wang 0065, Haitao Fu, Yingjie Gao 0002, Xiaohui Niu
BIBM8
2021 A statistical framework for predicting critical regions of p53-dependent enhancers
abstract
P53 is the 'guardian of the genome' and is responsible for regulating cell cycle and apoptosis. The genomic p53 binding regions, where activating transcriptional factors and cofactors like p300 simultaneously bind, are called 'p53-dependent enhancers', which play an important role in tumorigenesis. Current experimental assays generally provide a broad peak of each enhancer element, leaving our knowledge about critical enhancer regions (CERs) limited. Under the inspiration of enhancer dissection by CRISPR-Cas9 screen library on genome-wide p53 binding sites, here we introduce a statistical framework called 'Computational CRISPR Strategy' (CCS), to predict whether a given DNA fragment will be a p53-dependent CER by employing 7-mer as feature extractions along with random forest as the regressor. When training on a p53 CRISPR enhancer dataset, CCS not only accurately fitted the top-ranked enriched single guide RNAs (sgRNAs) but also successfully reproduced two known CERs that were validated by experiments. When applying it to an independent testing dataset on a tilling of a 2K-b genomic region of CRISPR-deCDKN1A-Lib, the trained model shows great generalizability by identifying a CER containing five top-ranked sgRNAs. A feature importance analysis further indicates that top-ranked 7-mers are mapped onto informative TF motifs including POU5F1 and SOX5, which are differentially enriched in p53-dependent CERs and are potential factors to make a general p53 binding site to form a p53-dependent CER, providing the interpretability of the trained model. Our results demonstrate that CCS is an alternative way of the CRISPR experiment to screen the genome for mapping p53-dependent CERs.
Xiaohui Niu, Kaixuan Deng, Lifen Liu, Xuehai Hu
Briefings Bioinform.1
2021 A systematic comparison of normalization methods for eQTL analysis
abstract
Expression quantitative trait loci (eQTL) analysis has been widely used in interpreting disease-associated loci through correlating genetic variant loci with the expression of specific genes. RNA-sequencing (RNA-Seq), which can quantify gene expression at the genome-wide level, is often used in eQTL identification. Since different normalization methods of gene expression have substantial impacts on RNA-seq downstream analysis, it is of great necessity to systematically compare the effects of these methods on eQTL identification. Here, by using RNA-seq and genotype data of four different cancers in The Cancer Genome Atlas (TCGA) database, we comprehensively evaluated the effect of eight commonly used normalization methods on eQTL identification. Our results showed that the application of different methods could cause 20-30% differences in the final results of eQTL identification. Among these methods, COUNT, Median of Ratio (MED) and Trimmed Mean of M-values (TMM) generated similar results for identifying eQTLs, while Fragments Per Kilobase Million (FPKM) or RANK produced more differential results compared with other methods. Based on the accuracy and receiver operating characteristic (ROC) curve, the TMM method was found to be the optimal method for normalizing gene expression data in eQTLs analysis. In addition, we also evaluated the performance of different pairwise combinations of these methods. As a result, compared with single normalization methods, the combination of methods can not only identify more cis-eQTLs, but also improve the performance of the ROC curve. Overall, this study provides a comprehensive comparison of normalization methods for identifying eQTLs from RNA-seq data, and proposes some practical recommendations for diverse scenarios.
Wenqian Yang, Weiwei Jin, Xiaohui Niu
Briefings Bioinform.6
2020 Toward Precise Osteotomies: A Coarse-to-Fine 3D Cut Plane Planning Method for Image-Guided Pelvis Tumor Resection Surgery
abstract
Surgical resection is the main clinical method for the treatment of bone tumors. A critical procedure for bone tumor resection is to plan a set of cut planes that enable resecting the bone tumor with a safe margin while preserving the maximum amount of healthy bone. Currently, the surgeons rely on manual methods to plan the cut planes, which highly depend on the surgeons' experiences and have been demonstrated to be error-prone, and in turn, increase the recurrence rate or resect much healthy bone. This study targets on improving the precision of cut plane planning for the image guided pelvis tumor resection surgeries. A semi-automatic approach to cut plane planning was proposed via a coarse-to-fine strategy. It can efficiently identify a dangerous region in the 3D space, which contains the bone tumor and its surrounding normal tissue with a safe margin. By projecting the dangerous region into an appropriate 2D space, a segmented boundary-constrained linear regression method was leveraged to plan a set of 3D cut planes that ensure the minimum area of the resected specimen in the 2D space while having the dangerous region cleared. Further, a coarse-to-fine 3D cut plane planning method was developed by incorporating a 3D cut plane refinement scheme with our 2D planning method. Extensive experiments, on the surgical data from nine previous pelvis tumor resection surgeries, demonstrated that our proposed approach substantially improved the localization precision of cut planes ( ) and decreased the amount of resected specimen ( ), as compared to the manual method.
Yu Zhang 0026, Fengzan Li, Lei Qiu 0004, Lihui Xu, Xiaohui Niu, Yao Sui, Shunli Zhang 0005, Li Zhang 0023
IEEE Trans. Medical Imaging5
2018 Estimating Maximum Target Registration Error Under Uniform Restriction of Fiducial Localization Error in Image Guided System
abstract
In this paper, we investigate the estimation of the maximum target registration error (TRE) magnitude of the target location while using point-based rigid registration in the image guided system. Under the uniform restriction of fiducial localization error (FLE) magnitude, we explicitly formulate the estimation as an optimization problem. Through analyzing the approximated problem which assumes the rigidity of the fiducial set holds with the perturbation of FLE, we present a strict lower bound for the maximum TRE magnitude. The simulations show that the lower bound is close to the actual maximum TRE magnitude for the target locations lying far away from the fiducial points. Unlike the expected TRE magnitude in which all fiducial points contribute, the lower bound is only related to the fiducial points serving as the vertices of the convex hull of the fiducial set. Our analysis provides a new perspective of investigating the problem of TRE estimation and is helpful for the surgeons to learn about the worst situation during using the image guided system.
Lei Qiu 0004, Yu Zhang 0026, Lihui Xu, Xiaohui Niu, Li Zhang 0023
IEEE Trans. Medical Imaging4