Rui Jiang 0001

dblp:57/3582-1 · DBLP profile ↗
← Back
50ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-7533-3753ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 44 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 6 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1
YearPublicationVenuePosition
2026 Cross-modality representation and multi-sample integration of spatially resolved omics data
abstract
Spatially resolved sequencing technologies have revolutionized our understanding of biological regulatory processes within tissue microenvironments by simultaneously capturing the states of genomic regions, genes, and proteins alongside the spatial organization of cells. However, inherent heterogeneity across modalities and samples poses substantial challenges for the integrative analysis of spatial omics data, underscoring the urgent need for advanced computational methods. In this study, we propose PRESENT, a contrastive learning-based integrative framework for cross-modality representation of spatial multi-omics data. PRESENT employs omics-specific encoders consisting of graph attention networks and Bayesian neural networks coupled with distribution-aware decoders to model distinct modalities, and an inter-omics alignment module for multi-omics integration. By effectively incorporating spatial dependencies with multi-omics information across diverse species and technologies, PRESENT facilitates the accurate identification of spatial domains and the elucidation of underlying regulatory mechanisms. Furthermore, PRESENT can be extended to multi-sample integration via a two-stage training workflow, which incorporates inter-batch alignment loss, intra-batch preserving loss, batch-adversarial learning, and cyclic graph refinement strategies to eliminate batch effects while retaining biological signals. Extensive experiments on tissue samples across different anatomical regions and developmental stages demonstrate that PRESENT enables the characterization of hierarchical tissue structures from a spatiotemporal perspective.
Zhen Li 0056, Xuejian Cui, Xiaoyang Chen 0007, Zijing Gao, Yuyao Liu, Yan Pan 0014, Shengquan Chen, Hairong Lv, Lei Zhai, Rui Jiang 0001
Briefings Bioinform.10
2025 Machine learning-enabled virtual screening indicates the anti-tuberculosis activity of aldoxorubicin and quarfloxin with verification by molecular docking, molecular dynamics simulations, and biological evaluations
abstract
Drug resistance in Mycobacterium tuberculosis (Mtb) is a significant challenge in the control and treatment of tuberculosis, making efforts to combat the spread of this global health burden more difficult. To accelerate anti-tuberculosis drug discovery, repurposing clinically approved or investigational drugs for the treatment of tuberculosis by computational methods has become an attractive strategy. In this study, we developed a virtual screening workflow that combines multiple machine learning and deep learning models, and 11 576 compounds extracted from the DrugBank database were screened against Mtb. Our screening method produced satisfactory predictions on three data-splitting settings, with the top predicted bioactive compounds all known antibacterial or anti-TB drugs. To further identify and evaluate drugs with repurposing potential in TB therapy, 15 screened potential compounds were selected for subsequent computational and experimental evaluations, out of which aldoxorubicin and quarfloxin showed potent inhibition of Mtb strain H37Rv, with minimal inhibitory concentrations of 4.16 and 20.67 μM/mL, respectively. More inspiringly, these two compounds also showed antibacterial activity against multidrug-resistant TB isolates and exhibited strong antimicrobial activity against Mtb. Furthermore, molecular docking, molecular dynamics simulation, and the surface plasmon resonance experiments validated the direct binding of the two compounds to Mtb DNA gyrase. In summary, our effective comprehensive virtual screening workflow successfully repurposed two novel drugs (aldoxorubicin and quarfloxin) as promising anti-Mtb candidates. The verification results provide useful information for the further development and clinical verification of anti-TB drugs.
Si Zheng 0001, Yaowen Gu, Yuzhen Gu, Yelin Zhao, Rui Jiang 0001, Jiao Li 0001
Briefings Bioinform.7
2025 Label Informed Contrastive Pretraining for Node Importance Estimation on Knowledge Graphs
abstract
Node importance estimation (NIE) is the task of inferring the importance scores of the nodes in a graph. Due to the availability of richer data and knowledge, recent research interests of NIE have been dedicated to knowledge graphs (KGs) for predicting future or missing node importance scores. Existing state-of-the-art NIE methods train the model by available labels, and they consider every interested node equally before training. However, the nodes with higher importance often require or receive more attention in real-world scenarios, e.g., people may care more about the movies or webpages with higher importance. To this end, we introduce Label Informed ContrAstive Pretraining (LICAP) to the NIE problem for being better aware of the nodes with high importance scores. Specifically, LICAP is a novel type of contrastive learning (CL) framework that aims to fully utilize continuous labels to generate contrastive samples for pretraining embeddings. Considering the NIE problem, LICAP adopts a novel sampling strategy called top nodes preferred hierarchical sampling to first group all interested nodes into a top bin and a nontop bin based on node importance scores, and then divide the nodes within the top bin into several finer bins also based on the scores. The contrastive samples are generated from those bins and are then used to pretrain node embeddings of KGs via a newly proposed predicate-aware graph attention networks (PreGATs), so as to better separate the top nodes from nontop nodes, and distinguish the top nodes within the top bin by keeping the relative order among finer bins. Extensive experiments demonstrate that the LICAP pretrained embeddings can further boost the performance of existing NIE methods and achieve new state-of-the-art performance regarding both regression and ranking metrics. The source code for reproducibility is available at https://github.com/zhangtia16/LICAP.
Chengbin Hou, Rui Jiang 0001, Xuegong Zhang, Chenghu Zhou, Ke Tang 0001, Hairong Lv
IEEE Trans. Neural Networks Learn. Syst.3
2024 Cofea: correlation-based feature selection for single-cell chromatin accessibility data
abstract
Single-cell chromatin accessibility sequencing (scCAS) technologies have enabled characterizing the epigenomic heterogeneity of individual cells. However, the identification of features of scCAS data that are relevant to underlying biological processes remains a significant gap. Here, we introduce a novel method Cofea, to fill this gap. Through comprehensive experiments on 5 simulated and 54 real datasets, Cofea demonstrates its superiority in capturing cellular heterogeneity and facilitating downstream analysis. Applying this method to identification of cell type-specific peaks and candidate enhancers, as well as pathway enrichment analysis and partitioned heritability analysis, we illustrate the potential of Cofea to uncover functional biological process.
Xiaoyang Chen 0007, Shuang Song 0006, Lin Hou 0003, Shengquan Chen, Rui Jiang 0001
Briefings Bioinform.6
2024 scPRAM accurately predicts single-cell gene expression perturbation response based on attention mechanism
abstract
MOTIVATION: With the rapid advancement of single-cell sequencing technology, it becomes gradually possible to delve into the cellular responses to various external perturbations at the gene expression level. However, obtaining perturbed samples in certain scenarios may be considerably challenging, and the substantial costs associated with sequencing also curtail the feasibility of large-scale experimentation. A repertoire of methodologies has been employed for forecasting perturbative responses in single-cell gene expression. However, existing methods primarily focus on the average response of a specific cell type to perturbation, overlooking the single-cell specificity of perturbation responses and a more comprehensive prediction of the entire perturbation response distribution. RESULTS: Here, we present scPRAM, a method for predicting perturbation responses in single-cell gene expression based on attention mechanisms. Leveraging variational autoencoders and optimal transport, scPRAM aligns cell states before and after perturbation, followed by accurate prediction of gene expression responses to perturbations for unseen cell types through attention mechanisms. Experiments on multiple real perturbation datasets involving drug treatments and bacterial infections demonstrate that scPRAM attains heightened accuracy in perturbation prediction across cell types, species, and individuals, surpassing existing methodologies. Furthermore, scPRAM demonstrates outstanding capability in identifying differentially expressed genes under perturbation, capturing heterogeneity in perturbation responses across species, and maintaining stability in the presence of data noise and sample size variations. AVAILABILITY AND IMPLEMENTATION: https://github.com/jiang-q19/scPRAM and https://doi.org/10.5281/zenodo.10935038.
Qun Jiang, Shengquan Chen, Xiaoyang Chen 0007, Rui Jiang 0001
Bioinform.4
2024 Accurate Annotation for Differentiating and Imbalanced Cell Types in Single-Cell Chromatin Accessibility Data
abstract
Rapid advances in single-cell chromatin accessibility sequencing (scCAS) technologies have enabled the characterization of epigenomic heterogeneity and increased the demand for automatic annotation of cell types. However, there are few computational methods tailored for cell type annotation in scCAS data and the existing methods perform poorly for differentiating and imbalanced cell types. Here, we propose CASCADE, a novel annotation method based on simulation- and denoising-based strategies. With comprehensive experiments on a number of scCAS datasets, we showed that CASCADE can effectively distinguish the patterns of different cell types and mitigate the effect of high noise levels, and thus achieve significantly better annotation performance for differentiating and imbalanced cell types. Besides, we performed model ablation experiments to show the contribution of modules in CASCADE and conducted extensive experiments to demonstrate the robustness of CASCADE to batch effect, imbalance degree, data sparsity, and number of cell types. Moreover, CASCADE significantly outperformed baseline methods for accurately annotating the cell types in newly sequenced data. We anticipate that CASCADE will greatly assist with characterizing cell heterogeneity in scCAS data analysis.
Yuhang Jia, Rui Jiang 0001, Shengquan Chen
IEEE ACM Trans. Comput. Biol. Bioinform.3
2023 Time-Aware Multiway Adaptive Fusion Network for Temporal Knowledge Graph Question Answering
abstract
Knowledge graphs (KGs) have received increasing attention due to its wide applications on natural language processing. However, its use case on temporal question answering (QA) has not been well-explored. Most of existing methods are developed based on pre-trained language models, which might not be capable to learn temporal-specific presentations of entities in terms of temporal KGQA task. To alleviate this problem, we propose a novel Time-aware Multiway Adaptive (TMA) fusion network. Inspired by the step-by-step reasoning behavior of humans. For each given question, TMA first extracts the relevant concepts from the KG, and then feeds them into a multiway adaptive module to produce a temporal-specific representation of the question. This representation can be incorporated with the pre-trained KG embedding to generate the final prediction. Empirical results verify that the proposed model achieves better performance than the state-of-the-art models in the benchmark dataset. Notably, the Hits@1 and Hits@10 results of TMA on the CronQuestions dataset’s complex questions are absolutely improved by 24% and 10% compared to the best-performing baseline. Furthermore, we also show that TMA employing an adaptive fusion mechanism can provide interpretability by analyzing the proportion of information in question representations.
Di Liang, Wei Wu 0014, Rui Jiang 0001
ICASSP6
2023 Deep generative modeling and clustering of single cell Hi-C data
abstract
Deciphering 3D genome conformation is important for understanding gene regulation and cellular function at a spatial level. The recent advances of single cell Hi-C technologies have enabled the profiling of the 3D architecture of DNA within individual cell, which allows us to study the cell-to-cell variability of 3D chromatin organization. Computational approaches are in urgent need to comprehensively analyze the sparse and heterogeneous single cell Hi-C data. Here, we proposed scDEC-Hi-C, a new framework for single cell Hi-C analysis with deep generative neural networks. scDEC-Hi-C outperforms existing methods in terms of single cell Hi-C data clustering and imputation. Moreover, the generative power of scDEC-Hi-C could help unveil the differences of chromatin architecture across cell types. We expect that scDEC-Hi-C could shed light on deepening our understanding of the complex mechanism underlying the formation of chromatin contacts.
Qiao Liu 0008, Wanwen Zeng, Wei Zhang 0241, Hongyang Chen 0001, Rui Jiang 0001, Mu Zhou, Shaoting Zhang 0001
Briefings Bioinform.6
2023 Improving artificial intelligence pipeline for liver malignancy diagnosis using ultrasound images and video frames
abstract
Recent developments of deep learning methods have demonstrated their feasibility in liver malignancy diagnosis using ultrasound (US) images. However, most of these methods require manual selection and annotation of US images by radiologists, which limit their practical application. On the other hand, US videos provide more comprehensive morphological information about liver masses and their relationships with surrounding structures than US images, potentially leading to a more accurate diagnosis. Here, we developed a fully automated artificial intelligence (AI) pipeline to imitate the workflow of radiologists for detecting liver masses and diagnosing liver malignancy. In this pipeline, we designed an automated mass-guided strategy that used segmentation information to direct diagnostic models to focus on liver masses, thus increasing diagnostic accuracy. The diagnostic models based on US videos utilized bi-directional convolutional long short-term memory modules with an attention-boosted module to learn and fuse spatiotemporal information from consecutive video frames. Using a large-scale dataset of 50 063 US images and video frames from 11 468 patients, we developed and tested the AI pipeline and investigated its applications. A dataset of annotated US images is available at https://doi.org/10.5281/zenodo.7272660.
Yiming Xu 0010, Xiaohong Liu 0007, Jinxiu Ju, Shi-jie Wang, Yufan Lian, Tong Liang, Ye Sang, Rui Jiang 0001, Ting Chen 0006
Briefings Bioinform.11
2023 simCAS: an embedding-based method for simulating single-cell chromatin accessibility sequencing data
abstract
MOTIVATION: Single-cell chromatin accessibility sequencing (scCAS) technology provides an epigenomic perspective to characterize gene regulatory mechanisms at single-cell resolution. With an increasing number of computational methods proposed for analyzing scCAS data, a powerful simulation framework is desirable for evaluation and validation of these methods. However, existing simulators generate synthetic data by sampling reads from real data or mimicking existing cell states, which is inadequate to provide credible ground-truth labels for method evaluation. RESULTS: We present simCAS, an embedding-based simulator, for generating high-fidelity scCAS data from both cell- and peak-wise embeddings. We demonstrate simCAS outperforms existing simulators in resembling real data and show that simCAS can generate cells of different states with user-defined cell populations and differentiation trajectories. Additionally, simCAS can simulate data from different batches and encode user-specified interactions of chromatin regions in the synthetic data, which provides ground-truth labels more than cell states. We systematically demonstrate that simCAS facilitates the benchmarking of four core tasks in downstream analysis: cell clustering, trajectory inference, data integration, and cis-regulatory interaction inference. We anticipate simCAS will be a reliable and flexible simulator for evaluating the ongoing computational methods applied on scCAS data. AVAILABILITY AND IMPLEMENTATION: simCAS is freely available at https://github.com/Chen-Li-17/simCAS.
Xiaoyang Chen 0007, Shengquan Chen, Rui Jiang 0001, Xuegong Zhang
Bioinform.4
2022 scGraph: a graph neural network-based approach to automatically identify cell types
abstract
MOTIVATION: Single-cell technologies play a crucial role in revolutionizing biological research over the past decade, which strengthens our understanding in cell differentiation, development and regulation from a single-cell level perspective. Single-cell RNA sequencing (scRNA-seq) is one of the most common single cell technologies, which enables probing transcriptional states in thousands of cells in one experiment. Identification of cell types from scRNA-seq measurements is a fundamental and crucial question to answer. Most previous studies directly take gene expression as input while ignoring the comprehensive gene-gene interactions. RESULTS: We propose scGraph, an automatic cell identification algorithm leveraging gene interaction relationships to enhance the performance of the cell-type identification. scGraph is based on a graph neural network to aggregate the information of interacting genes. In a series of experiments, we demonstrate that scGraph is accurate and outperforms eight comparison methods in the task of cell-type identification. Moreover, scGraph automatically learns the gene interaction relationships from biological data and the pathway enrichment analysis shows consistent findings with previous analysis, providing insights on the analysis of regulatory mechanism. AVAILABILITY AND IMPLEMENTATION: scGraph is freely available at https://github.com/QijinYin/scGraph and https://figshare.com/articles/software/scGraph/17157743. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qijin Yin, Qiao Liu 0008, Zhuoran Fu, Wanwen Zeng, Boheng Zhang, Xuegong Zhang, Rui Jiang 0001, Hairong Lv
Bioinform.7
2022 DualGCN: a dual graph convolutional network model to predict cancer drug response
abstract
Abstract Background Drug resistance is a critical obstacle in cancer therapy. Discovering cancer drug response is important to improve anti-cancer drug treatment and guide anti-cancer drug design. Abundant genomic and drug response resources of cancer cell lines provide unprecedented opportunities for such study. However, cancer cell lines cannot fully reflect heterogeneous tumor microenvironments. Transferring knowledge studied from in vitro cell lines to single-cell and clinical data will be a promising direction to better understand drug resistance. Most current studies include single nucleotide variants (SNV) as features and focus on improving predictive ability of cancer drug response on cell lines. However, obtaining accurate SNVs from clinical tumor samples and single-cell data is not reliable. This makes it difficult to generalize such SNV-based models to clinical tumor data or single-cell level studies in the future. Results We present a new method, DualGCN, a unified Dual Graph Convolutional Network model to predict cancer drug response. DualGCN encodes both chemical structures of drugs and omics data of biological samples using graph convolutional networks. Then the two embeddings are fed into a multilayer perceptron to predict drug response. DualGCN incorporates prior knowledge on cancer-related genes and protein–protein interactions, and outperforms most state-of-the-art methods while avoiding using large-scale SNV data. Conclusions The proposed method outperforms most state-of-the-art methods in predicting cancer drug response without the use of large-scale SNV data. These favorable results indicate its potential to be extended to clinical and single-cell tumor samples and advancements in precision medicine.
Tianxing Ma, Qiao Liu 0008, Haochen Li 0003, Mu Zhou, Rui Jiang 0001, Xuegong Zhang
BMC Bioinform.5
2022 CS-CO: A Hybrid Self-Supervised Visual Representation Learning Method for H&E-stained Histopathological Images
Pengshuai Yang, Xiaoxu Yin, Haiming Lu, Zhongliang Hu, Xuegong Zhang, Rui Jiang 0001, Hairong Lv
Medical Image Anal.6
2021 Self-supervised Visual Representation Learning for Histopathological Images
Pengshuai Yang, Zhiwei Hong, Xiaoxu Yin, Chengzhan Zhu, Rui Jiang 0001
MICCAI (2)5
2021 stPlus: a reference-based method for the accurate enhancement of spatial transcriptomics
abstract
MOTIVATION: Single-cell RNA sequencing (scRNA-seq) techniques have revolutionized the investigation of transcriptomic landscape in individual cells. Recent advancements in spatial transcriptomic technologies further enable gene expression profiling and spatial organization mapping of cells simultaneously. Among the technologies, imaging-based methods can offer higher spatial resolutions, while they are limited by either the small number of genes imaged or the low gene detection sensitivity. Although several methods have been proposed for enhancing spatially resolved transcriptomics, inadequate accuracy of gene expression prediction and insufficient ability of cell-population identification still impede the applications of these methods. RESULTS: We propose stPlus, a reference-based method that leverages information in scRNA-seq data to enhance spatial transcriptomics. Based on an auto-encoder with a carefully tailored loss function, stPlus performs joint embedding and predicts spatial gene expression via a weighted k-nearest-neighbor. stPlus outperforms baseline methods with higher gene-wise and cell-wise Spearman correlation coefficients. We also introduce a clustering-based approach to assess the enhancement performance systematically. Using the data enhanced by stPlus, cell populations can be better identified than using the measured data. The predicted expression of genes unique to scRNA-seq data can also well characterize spatial cell heterogeneity. Besides, stPlus is robust and scalable to datasets of diverse gene detection sensitivity levels, sample sizes and number of spatially measured genes. We anticipate stPlus will facilitate the analysis of spatial transcriptomics. AVAILABILITY AND IMPLEMENTATION: stPlus with detailed documents is freely accessible at http://health.tsinghua.edu.cn/software/stPlus/ and the source code is openly available on https://github.com/xy-chen16/stPlus.
Shengquan Chen, Boheng Zhang, Xiaoyang Chen 0007, Xuegong Zhang, Rui Jiang 0001
Bioinform.5
2021 Few shot domain adaptation for in situ macromolecule structural classification in cryoelectron tomograms
abstract
MOTIVATION: Cryoelectron tomography (cryo-ET) visualizes structure and spatial organization of macromolecules and their interactions with other subcellular components inside single cells in the close-to-native state at submolecular resolution. Such information is critical for the accurate understanding of cellular processes. However, subtomogram classification remains one of the major challenges for the systematic recognition and recovery of the macromolecule structures in cryo-ET because of imaging limits and data quantity. Recently, deep learning has significantly improved the throughput and accuracy of large-scale subtomogram classification. However, often it is difficult to get enough high-quality annotated subtomogram data for supervised training due to the enormous expense of labeling. To tackle this problem, it is beneficial to utilize another already annotated dataset to assist the training process. However, due to the discrepancy of image intensity distribution between source domain and target domain, the model trained on subtomograms in source domain may perform poorly in predicting subtomogram classes in the target domain. RESULTS: In this article, we adapt a few shot domain adaptation method for deep learning-based cross-domain subtomogram classification. The essential idea of our method consists of two parts: (i) take full advantage of the distribution of plentiful unlabeled target domain data, and (ii) exploit the correlation between the whole source domain dataset and few labeled target domain data. Experiments conducted on simulated and real datasets show that our method achieves significant improvement on cross domain subtomogram classification compared with baseline methods. AVAILABILITY AND IMPLEMENTATION: Software is available online https://github.com/xulabs/aitom. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Liangyong Yu, Ge Yang 0002, Rui Jiang 0001, Min Xu 0009
Bioinform.7
2020 DeepCDR: a hybrid graph convolutional network for predicting cancer drug response
abstract
MOTIVATION: Accurate prediction of cancer drug response (CDR) is challenging due to the uncertainty of drug efficacy and heterogeneity of cancer patients. Strong evidences have implicated the high dependence of CDR on tumor genomic and transcriptomic profiles of individual patients. Precise identification of CDR is crucial in both guiding anti-cancer drug design and understanding cancer biology. RESULTS: In this study, we present DeepCDR which integrates multi-omics profiles of cancer cells and explores intrinsic chemical structures of drugs for predicting CDR. Specifically, DeepCDR is a hybrid graph convolutional network consisting of a uniform graph convolutional network and multiple subnetworks. Unlike prior studies modeling hand-crafted features of drugs, DeepCDR automatically learns the latent representation of topological structures among atoms and bonds of drugs. Extensive experiments showed that DeepCDR outperformed state-of-the-art methods in both classification and regression settings under various data settings. We also evaluated the contribution of different types of omics profiles for assessing drug response. Furthermore, we provided an exploratory strategy for identifying potential cancer-associated genes concerning specific cancer types. Our results highlighted the predictive power of DeepCDR and its potential translational value in guiding disease-specific drug design. AVAILABILITY AND IMPLEMENTATION: DeepCDR is freely available at https://github.com/kimmo1019/DeepCDR. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qiao Liu 0008, Rui Jiang 0001, Mu Zhou
Bioinform.3
2020 Integrating distal and proximal information to predict gene expression via a densely connected convolutional neural network
abstract
MOTIVATION: Interactions among cis-regulatory elements such as enhancers and promoters are main driving forces shaping context-specific chromatin structure and gene expression. Although there have been computational methods for predicting gene expression from genomic and epigenomic information, most of them neglect long-range enhancer-promoter interactions, due to the difficulty in precisely linking regulatory enhancers to target genes. Recently, HiChIP, a novel high-throughput experimental approach, has generated comprehensive data on high-resolution interactions between promoters and distal enhancers. Moreover, plenty of studies suggest that deep learning achieves state-of-the-art performance in epigenomic signal prediction, and thus promoting the understanding of regulatory elements. In consideration of these two factors, we integrate proximal promoter sequences and HiChIP distal enhancer-promoter interactions to accurately predict gene expression. RESULTS: We propose DeepExpression, a densely connected convolutional neural network, to predict gene expression using both promoter sequences and enhancer-promoter interactions. We demonstrate that our model consistently outperforms baseline methods, not only in the classification of binary gene expression status but also in regression of continuous gene expression levels, in both cross-validation experiments and cross-cell line predictions. We show that the sequential promoter information is more informative than the experimental enhancer information; meanwhile, the enhancer-promoter interactions within ±100 kbp around the TSS of a gene are most beneficial. We finally visualize motifs in both promoter and enhancer regions and show the match of identified sequence signatures with known motifs. We expect to see a wide spectrum of applications using HiChIP data in deciphering the mechanism of gene regulation. AVAILABILITY AND IMPLEMENTATION: DeepExpression is freely available at https://github.com/wanwenzeng/DeepExpression. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wanwen Zeng, Rui Jiang 0001
Bioinform.3
2020 EnClaSC: a novel ensemble approach for accurate and robust cell-type classification of single-cell transcriptomes
abstract
BACKGROUND: In recent years, the rapid development of single-cell RNA-sequencing (scRNA-seq) techniques enables the quantitative characterization of cell types at a single-cell resolution. With the explosive growth of the number of cells profiled in individual scRNA-seq experiments, there is a demand for novel computational methods for classifying newly-generated scRNA-seq data onto annotated labels. Although several methods have recently been proposed for the cell-type classification of single-cell transcriptomic data, such limitations as inadequate accuracy, inferior robustness, and low stability greatly limit their wide applications. RESULTS: We propose a novel ensemble approach, named EnClaSC, for accurate and robust cell-type classification of single-cell transcriptomic data. Through comprehensive validation experiments, we demonstrate that EnClaSC can not only be applied to the self-projection within a specific dataset and the cell-type classification across different datasets, but also scale up well to various data dimensionality and different data sparsity. We further illustrate the ability of EnClaSC to effectively make cross-species classification, which may shed light on the studies in correlation of different species. EnClaSC is freely available at https://github.com/xy-chen16/EnClaSC . CONCLUSIONS: EnClaSC enables highly accurate and robust cell-type classification of single-cell transcriptomic data via an ensemble learning method. We expect to see wide applications of our method to not only transcriptome studies, but also the classification of more general data.
Xiaoyang Chen 0007, Shengquan Chen, Rui Jiang 0001
BMC Bioinform.3
2020 Few-shot learning for classification of novel macromolecular structures in cryo-electron tomograms
abstract
Cryo-electron tomography (cryo-ET) provides 3D visualization of subcellular components in the near-native state and at sub-molecular resolutions in single cells, demonstrating an increasingly important role in structural biology in situ. However, systematic recognition and recovery of macromolecular structures in cryo-ET data remain challenging as a result of low signal-to-noise ratio (SNR), small sizes of macromolecules, and high complexity of the cellular environment. Subtomogram structural classification is an essential step for such task. Although acquisition of large amounts of subtomograms is no longer an obstacle due to advances in automation of data collection, obtaining the same number of structural labels is both computation and labor intensive. On the other hand, existing deep learning based supervised classification approaches are highly demanding on labeled data and have limited ability to learn about new structures rapidly from data containing very few labels of such new structures. In this work, we propose a novel approach for subtomogram classification based on few-shot learning. With our approach, classification of unseen structures in the training data can be conducted given few labeled samples in test data through instance embedding. Experiments were performed on both simulated and real datasets. Our experimental results show that we can make inference on new structures given only five labeled samples for each class with a competitive accuracy (> 0.86 on the simulated dataset with SNR = 0.1), or even one sample with an accuracy of 0.7644. The results on real datasets are also promising with accuracy > 0.9 on both conditions and even up to 1 on one of the real datasets. Our approach achieves significant improvement compared with the baseline method and has strong capabilities of generalizing to other cellular components.
Liangyong Yu, Bo Zhou 0009, Jing Zhang 0062, Xin Gao 0001, Rui Jiang 0001, Min Xu 0009
PLoS Comput. Biol.9
2019 Liver Histopathological Image Retrieval Based on Deep Metric Learning
abstract
Histopathological image retrieval aims to search histopathological images sharing similar content with the query image, which could provide pathologists with an approach to easily obtain similar diagnostic cases for reference. Recent histopathological image retrieval methods are usually based on CNN feature extractors, which require a large amount of annotated data for training. Besides, most of existing methods could not define a reasonable similarity metric for histopathological images. In this paper, we apply deep metric learning to liver histopathological image retrieval. We construct a model based on mixed attention mechanism and train the model with a modified version of multi-similarity loss, which enables embedding vectors of similar images in the given metric space to be closer and dissimilar ones to be far from each other. Additionally, our model can be well fitted with limited data. Finally, we evaluate the proposed method with our own established liver histopathological image dataset. Compared with several published methods, our model shows higher performance.
Pengshuai Yang, Yupeng Zhai, Hairong Lv, Jigang Wang, Chengzhan Zhu, Rui Jiang 0001
BIBM7
2019 Rule-Based Method to Develop Question-Answer Dataset from Chest X-Ray Reports
abstract
Available and objective clinical documents are important for research of assistant diagnosis, development of algorithms, and education. To facilitate the readability and variability of clinical documents, this paper presents a rule-based approach to develop a question-answer dataset for chest X-rays from a public collection of radiology examinations, including both images and radiologist narrative reports. Our method simplified the complicated reports via hand-selected keywords, generated more than 63 thousand question-answer pairs via hand-written patterns, and augmented the question-answer dataset to more than 130 thousand pairs via rule-based question answering. To the best of our knowledge, this is the first generated question-answer dataset for chest X-rays by rule-based method. The dataset is promising for future researches and applications such as visual question answering, computer-aided diagnosis and so on.
Jie Wang 0111, Hairong Lv, Rui Jiang 0001
CBMS3
2019 Instance Segmentation of Anatomical Structures in Chest Radiographs
abstract
Automatic and accurate segmentation of anatomical structures in chest radiographs is fundamental and essential for computer-aided diagnosis system. We introduce Mask R-CNN for instance segmentation of lung fields, heart and clavicles. This method efficiently detects different structures and generates accurate segmentation mask for each instance. To the best of our knowledge, we are the first to implement instance segmentation of these three anatomical structures in chest radiographs. We have done extensive experiments on a common benchmark dataset. Results show that the best of our models achieves the state-of-the-art segmentation performance on image resolution of 512 × 512. The Dice and Ω similarity are 0.976 and 0.953 for lung fields, 0.949 and 0.904 for heart, 0.920 and 0.852 for clavicles. And the average contour distance outperforms human observer on both lungs and heart with image resolution of 256 × 256. In addition, it takes only 0.16 and 0.12 seconds per image for the above two resolutions during inference, which is comparable to or even better than current methods.
Jie Wang 0111, Zhigang Li 0004, Rui Jiang 0001
CBMS3
2019 hicGAN infers super resolution Hi-C data with generative adversarial networks
abstract
MOTIVATION: Hi-C is a genome-wide technology for investigating 3D chromatin conformation by measuring physical contacts between pairs of genomic regions. The resolution of Hi-C data directly impacts the effectiveness and accuracy of downstream analysis such as identifying topologically associating domains (TADs) and meaningful chromatin loops. High resolution Hi-C data are valuable resources which implicate the relationship between 3D genome conformation and function, especially linking distal regulatory elements to their target genes. However, high resolution Hi-C data across various tissues and cell types are not always available due to the high sequencing cost. It is therefore indispensable to develop computational approaches for enhancing the resolution of Hi-C data. RESULTS: We proposed hicGAN, an open-sourced framework, for inferring high resolution Hi-C data from low resolution Hi-C data with generative adversarial networks (GANs). To the best of our knowledge, this is the first study to apply GANs to 3D genome analysis. We demonstrate that hicGAN effectively enhances the resolution of low resolution Hi-C data by generating matrices that are highly consistent with the original high resolution Hi-C matrices. A typical scenario of usage for our approach is to enhance low resolution Hi-C data in new cell types, especially where the high resolution Hi-C data are not available. Our study not only presents a novel approach for enhancing Hi-C data resolution, but also provides fascinating insights into disclosing complex mechanism underlying the formation of chromatin contacts. AVAILABILITY AND IMPLEMENTATION: We release hicGAN as an open-sourced software at https://github.com/kimmo1019/hicGAN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qiao Liu 0008, Hairong Lv, Rui Jiang 0001
Bioinform.3
2019 VPAC: Variational projection for accurate clustering of single-cell transcriptomic data
abstract
BACKGROUND: Single-cell RNA-sequencing (scRNA-seq) technologies have advanced rapidly in recent years and enabled the quantitative characterization at a microscopic resolution. With the exponential growth of the number of cells profiled in individual scRNA-seq experiments, the demand for identifying putative cell types from the data has become a great challenge that appeals for novel computational methods. Although a variety of algorithms have recently been proposed for single-cell clustering, such limitations as low accuracy, inferior robustness, and inadequate stability greatly impede the scope of applications of these methods. RESULTS: We propose a novel model-based algorithm, named VPAC, for accurate clustering of single-cell transcriptomic data through variational projection, which assumes that single-cell samples follow a Gaussian mixture distribution in a latent space. Through comprehensive validation experiments, we demonstrate that VPAC can not only be applied to datasets of discrete counts and normalized continuous data, but also scale up well to various data dimensionality, different dataset size and different data sparsity. We further illustrate the ability of VPAC to detect genes with strong unique signatures of a specific cell type, which may shed light on the studies in system biology. We have released a user-friendly python package of VPAC in Github ( https://github.com/ShengquanChen/VPAC ). Users can directly import our VPAC class and conduct clustering without tedious installation of dependency packages. CONCLUSIONS: VPAC enables highly accurate clustering of single-cell transcriptomic data via a statistical model. We expect to see wide applications of our method to not only transcriptome studies for fully understanding the cell identity and functionality, but also the clustering of more general data.
Shengquan Chen, Kui Hua, Hongfei Cui, Rui Jiang 0001
BMC Bioinform.4
2019 Automatic localization and identification of mitochondria in cellular electron cryo-tomography using faster-RCNN
abstract
BACKGROUND: Cryo-electron tomography (cryo-ET) enables the 3D visualization of cellular organization in near-native state which plays important roles in the field of structural cell biology. However, due to the low signal-to-noise ratio (SNR), large volume and high content complexity within cells, it remains difficult and time-consuming to localize and identify different components in cellular cryo-ET. To automatically localize and recognize in situ cellular structures of interest captured by cryo-ET, we proposed a simple yet effective automatic image analysis approach based on Faster-RCNN. RESULTS: Our experimental results were validated using in situ cyro-ET-imaged mitochondria data. Our experimental results show that our algorithm can accurately localize and identify important cellular structures on both the 2D tilt images and the reconstructed 2D slices of cryo-ET. When ran on the mitochondria cryo-ET dataset, our algorithm achieved Average Precision >0.95. Moreover, our study demonstrated that our customized pre-processing steps can further improve the robustness of our model performance. CONCLUSIONS: In this paper, we proposed an automatic Cryo-ET image analysis algorithm for localization and identification of different structure of interest in cells, which is the first Faster-RCNN based method for localizing an cellular organelle in Cryo-ET images and demonstrated the high accuracy and robustness of detection and classification tasks of intracellular mitochondria. Furthermore, our approach can be easily applied to detection tasks of other cellular structures as well.
Stephanie E. Sigmund, Ruogu Lin, Bo Zhou 0009, Chang Liu 0031, Rui Jiang 0001, Zachary Freyberg, Hairong Lv, Min Xu 0009
BMC Bioinform.8
2018 Respond-CAM: Analyzing Deep Models for 3D Imaging Data by Visualizations
Guannan Zhao, Bo Zhou 0009, Rui Jiang 0001, Min Xu 0009
MICCAI (1)4
2018 Chromatin accessibility prediction via a hybrid deep convolutional neural network
abstract
Motivation: A majority of known genetic variants associated with human-inherited diseases lie in non-coding regions that lack adequate interpretation, making it indispensable to systematically discover functional sites at the whole genome level and precisely decipher their implications in a comprehensive manner. Although computational approaches have been complementing high-throughput biological experiments towards the annotation of the human genome, it still remains a big challenge to accurately annotate regulatory elements in the context of a specific cell type via automatic learning of the DNA sequence code from large-scale sequencing data. Indeed, the development of an accurate and interpretable model to learn the DNA sequence signature and further enable the identification of causative genetic variants has become essential in both genomic and genetic studies. Results: We proposed Deopen, a hybrid framework mainly based on a deep convolutional neural network, to automatically learn the regulatory code of DNA sequences and predict chromatin accessibility. In a series of comparison with existing methods, we show the superior performance of our model in not only the classification of accessible regions against background sequences sampled at random, but also the regression of DNase-seq signals. Besides, we further visualize the convolutional kernels and show the match of identified sequence signatures and known motifs. We finally demonstrate the sensitivity of our model in finding causative noncoding variants in the analysis of a breast cancer dataset. We expect to see wide applications of Deopen with either public or in-house chromatin accessibility data in the annotation of the human genome and the identification of non-coding variants associated with diseases. Availability and implementation: Deopen is freely available at https://github.com/kimmo1019/Deopen. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Qiao Liu 0008, Fei Xia 0002, Qijin Yin, Rui Jiang 0001
Bioinform.4
2018 FLOWER: Fusing global and local associations towards personalized social recommendation
Mingxin Gan, Rui Jiang 0001
Future Gener. Comput. Syst.2
2018 Guest Editorial for Special Section on the Sixth National Conference on Bioinformatics and System Biology of China
abstract
The three papers in this special section were presented at the Sixth National Conference on Bioinformatics and System Biology of China that was held in Nanjing, China, on October 6-9, 2014.
Xiaodan Fan, Xinglai Ji, Rui Jiang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2017 Integrating embeddings of multiple gene networks to prioritize complex disease-associated genes
abstract
Genome-wide association study (GWAS), as one primary approach for genetic studies, has been successfully applied to a variety of complex diseases, leading to the discovery of substantial disease-associated loci. These discovered associations provide unprecedented opportunities for deepening our understanding of complex diseases, such as disease-associated risk variants, genes, and pathways. However, it is non-trivial to extract biological knowledge from the GWAS data due to the existence of several non-negligible factors. For example, the majority of associated loci fall into noncoding regions without certain links to any genes, complicating its functional characterization. Network-based GWAS gene prioritization, aiming to integrate gene networks with GWAS data, emerges as one promising direction towards solving these challenges and has attracted much attention recently. However, gene networks are usually sparse and noisy, and existing methods do not explicitly consider these properties, leading to suboptimal performance. In this paper, we proposed a novel method called REGENT for integrating multiple gene networks with GWAS data to prioritize complex disease-associated genes. Specifically, we leveraged the network representation learning, a recently developed technique for analyzing social networks, to learn compact and robust embeddings from multiple gene networks. To integrate these learned embeddings of genes with GWAS data, we developed a hierarchical statistical model and derived an efficient inference algorithm for model estimation and prediction. Applying to GWAS data of six complex diseases, we demonstrated that REGENT outperformed existing methods regarding the identification of known disease-associated genes. Also, pathway analysis showed that REGENT helped discover disease-associated pathways. Therefore, our method is expected to be a useful tool for post-GWAS analysis.
Mengmeng Wu, Wanwen Zeng, Yi-Jia Zhang 0001, Ting Chen 0006, Rui Jiang 0001
BIBM6
2017 Chromatin accessibility prediction via convolutional long short-term memory networks with k-mer embedding
abstract
MOTIVATION: Experimental techniques for measuring chromatin accessibility are expensive and time consuming, appealing for the development of computational approaches to predict open chromatin regions from DNA sequences. Along this direction, existing methods fall into two classes: one based on handcrafted k -mer features and the other based on convolutional neural networks. Although both categories have shown good performance in specific applications thus far, there still lacks a comprehensive framework to integrate useful k -mer co-occurrence information with recent advances in deep learning. RESULTS: We fill this gap by addressing the problem of chromatin accessibility prediction with a convolutional Long Short-Term Memory (LSTM) network with k -mer embedding. We first split DNA sequences into k -mers and pre-train k -mer embedding vectors based on the co-occurrence matrix of k -mers by using an unsupervised representation learning approach. We then construct a supervised deep learning architecture comprised of an embedding layer, three convolutional layers and a Bidirectional LSTM (BLSTM) layer for feature learning and classification. We demonstrate that our method gains high-quality fixed-length features from variable-length sequences and consistently outperforms baseline methods. We show that k -mer embedding can effectively enhance model performance by exploring different embedding strategies. We also prove the efficacy of both the convolution and the BLSTM layers by comparing two variations of the network architecture. We confirm the robustness of our model to hyper-parameters by performing sensitivity analysis. We hope our method can eventually reinforce our understanding of employing deep learning in genomic studies and shed light on research regarding mechanisms of chromatin accessibility. AVAILABILITY AND IMPLEMENTATION: The source code can be downloaded from https://github.com/minxueric/ismb2017_lstm . CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary materials are available at Bioinformatics online.
Xu Min, Wanwen Zeng, Ning Chen 0002, Ting Chen 0006, Rui Jiang 0001
Bioinform.5
2017 Predicting enhancers with deep convolutional neural networks
abstract
BACKGROUND: With the rapid development of deep sequencing techniques in the recent years, enhancers have been systematically identified in such projects as FANTOM and ENCODE, forming genome-wide landscapes in a series of human cell lines. Nevertheless, experimental approaches are still costly and time consuming for large scale identification of enhancers across a variety of tissues under different disease status, making computational identification of enhancers indispensable. RESULTS: To facilitate the identification of enhancers, we propose a computational framework, named DeepEnhancer, to distinguish enhancers from background genomic sequences. Our method purely relies on DNA sequences to predict enhancers in an end-to-end manner by using a deep convolutional neural network (CNN). We train our deep learning model on permissive enhancers and then adopt a transfer learning strategy to fine-tune the model on enhancers specific to a cell line. Results demonstrate the effectiveness and efficiency of our method in the classification of enhancers against random sequences, exhibiting advantages of deep learning over traditional sequence-based classifiers. We then construct a variety of neural networks with different architectures and show the usefulness of such techniques as max-pooling and batch normalization in our method. To gain the interpretability of our approach, we further visualize convolutional kernels as sequence logos and successfully identify similar motifs in the JASPAR database. CONCLUSIONS: DeepEnhancer enables the identification of novel enhancers using only DNA sequences via a highly accurate deep learning model. The proposed computational framework can also be applied to similar problems, thereby prompting the use of machine learning methods in life sciences.
Xu Min, Wanwen Zeng, Shengquan Chen, Ning Chen 0002, Ting Chen 0006, Rui Jiang 0001
BMC Bioinform.6
2016 DeepEnhancer: Predicting enhancers by convolutional neural networks
abstract
Enhancers are crucial to the understanding of mechanisms underlying gene transcriptional regulation. Although having been successfully applied in such projects as ENCODE and Roadmap to generate landscape of enhancers in human cell lines, high-throughput biological experimental techniques are still costly and time consuming for even larger scale identification of enhancers across a variety of tissues under different disease status, making computational identification of enhancers indispensable. In this paper, we propose a computational framework, named DeepEnhancer, to classify enhancers from background genomic sequences. We construct convolutional neural networks of various architectures and compare the classification performance with traditional sequence-based classifiers. We first train the deep learning model on the FANTOM5 permissive enhancer dataset, and then fine-tune the model on ENCODE cell type-specific enhancer datasets by adopting the transfer learning strategy. Experimental results demonstrate that DeepEnhancer has superior efficiency and effectiveness in classification tasks, and the use of max-pooling and batch normalization is beneficial to higher accuracy. To make our approach more understandable, we propose a strategy to visualize the convolutional kernels as sequence logos and compare them against the JASPAR database using TOMTOM. In summary, DeepEnhancer allows researchers to train highly accurate deep models and will be broadly applicable in computational biology.
Xu Min, Ning Chen 0002, Ting Chen 0006, Rui Jiang 0001
BIBM4
2016 Global inference of disease-causing single nucleotide variants from exome sequencing data
abstract
BACKGROUND: Whole exome sequencing (WES) has recently emerged as an effective approach for identifying genetic variants underlying human diseases. However, considerable time and labour is needed for careful investigation of candidate variants. Although filtration based on population frequencies and functional prediction scores could effectively remove common and neutral variants, hundreds or even thousands of rare deleterious variants still remain. In addition, current WES platforms also provide variant information in flanking noncoding regions, such as promoters, introns and splice sites. Despite of being recognized to harbour causal variants, these regions are usually ignored by current analysis pipelines. RESULTS: We present a novel computational method, called Glints, to overcome the above limitations. Glints is capable of identifying disease-causing SNVs in both coding and flanking noncoding regions from exome sequencing data. The principle behind Glints is that disease-causing variants should manifest their effect at both variant and gene levels. Specifically, Glints integrates 14 types of functional scores, including predictions for both coding and noncoding variants, and 9 types of association scores, which help identifying disease relevant genes. We conducted a large-scale simulation studies based on 1000 Genomes Project data and demonstrated the effectiveness of our method in both coding and flanking noncoding regions. We also applied Glints in two real exome sequencing and demonstrated its effectiveness for uncovering disease-causing SNVs. Both standalone software and web server are available at our website http://bioinfo.au.tsinghua.edu.cn/jianglab/glints . CONCLUSIONS: Glints is effective for uncovering disease-causing SNVs in coding and flanking noncoding regions, which is supported by both simulation and real case studies. Glints is expected to be a useful tool for human genetics research based on exome sequencing data.
Mengmeng Wu, Ting Chen 0006, Rui Jiang 0001
BMC Bioinform.3
2016 Trinity: Walking on a User-Object-Tag Heterogeneous Network for Personalised Recommendations
Mingxin Gan, Lily Sun, Rui Jiang 0001
J. Comput. Sci. Technol.3
2015 Differential regulation enrichment analysis via the integration of transcriptional regulatory network and gene expression data
abstract
MOTIVATION: Although many gene set analysis methods have been proposed to explore associations between a phenotype and a group of genes sharing common biological functions or involved in the same biological process, the underlying biological mechanisms of identified gene sets are typically unexplained. RESULTS: We propose a method called Differential Regulation-based enrichment Analysis for GENe sets (DRAGEN) to identify gene sets in which a significant proportion of genes have their transcriptional regulatory patterns changed in a perturbed phenotype. We conduct comprehensive simulation studies to demonstrate the capability of our method in identifying differentially regulated gene sets. We further apply our method to three human microarray expression datasets, two with hormone treated and control samples and one concerning different cell cycle phases. Results indicate that the capability of DRAGEN in identifying phenotype-associated gene sets is significantly superior to those of four existing methods for analyzing differentially expressed gene sets. We conclude that the proposed differential regulation enrichment analysis method, though exploratory in nature, complements the existing gene set analysis methods and provides a promising new direction for the interpretation of gene expression data. AVAILABILITY AND IMPLEMENTATION: The program of DRAGEN is freely available at http://bioinfo.au.tsinghua.edu.cn/dragen/. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shining Ma, Tao Jiang 0001, Rui Jiang 0001
Bioinform.3
2015 ROUND: Walking on an object-user heterogeneous network for personalized recommendations
Mingxin Gan, Rui Jiang 0001
Expert Syst. Appl.2
2013 Inferring semantic similarity through correlating information contents of gene ontology terms
abstract
Successful applications of the gene ontology to infer functional relationships between gene products in recent years have raised the need for computational methods to automatically calculate semantic similarity between gene products based on the gene ontology. To meet this challenge, several methods have been proposed to derive semantic similarity between gene products based on semantic similarity of gene ontology terms. However, these methods, though having been widely used in a variety of applications, may significantly overestimate semantic similarity between genes that are actually not functional related, thereby yielding misleading results in applications. To overcome this limitation, we propose to represent a gene product as a vector that is composed of information contents of gene ontology terms annotated for the gene product. Results show that semantic similarity scores calculated using our proposed method are more consistent with known biological knowledge than those derived using a list of existing methods, suggesting the effectiveness of our method in characterizing functional relationships between gene products.
Mingxin Gan, Rui Jiang 0001
BIBM2
2013 PyroHMMvar: a sensitive and accurate method to call short indels and SNPs for Ion Torrent and 454 data
abstract
MOTIVATION: The identification of short insertions and deletions (indels) and single nucleotide polymorphisms (SNPs) from Ion Torrent and 454 reads is a challenging problem, essentially because these techniques are prone to sequence erroneously at homopolymers and can, therefore, raise indels in reads. Most of the existing mapping programs do not model homopolymer errors when aligning reads against the reference. The resulting alignments will then contain various kinds of mismatches and indels that confound the accurate determination of variant loci and alleles. RESULTS: To address these challenges, we realign reads against the reference using our previously proposed hidden Markov model that models homopolymer errors and then merges these pairwise alignments into a weighted alignment graph. Based on our weighted alignment graph and hidden Markov model, we develop a method called PyroHMMvar, which can simultaneously detect short indels and SNPs, as demonstrated in human resequencing data. Specifically, by applying our methods to simulated diploid datasets, we demonstrate that PyroHMMvar produces more accurate results than state-of-the-art methods, such as Samtools and GATK, and is less sensitive to mapping parameter settings than the other methods. We also apply PyroHMMvar to analyze one human whole genome resequencing dataset, and the results confirm that PyroHMMvar predicts SNPs and indels accurately. AVAILABILITY AND IMPLEMENTATION: Source code freely available at the following URL: https://code.google.com/p/pyrohmmvar/, implemented in C++ and supported on Linux. .
Rui Jiang 0001
Bioinform.2
2013 Improving accuracy and diversity of personalized recommendation through power law adjustments of user similarities
Mingxin Gan, Rui Jiang 0001
Decis. Support Syst.2
2013 Constructing a user similarity network to remove adverse influence of popular objects for personalized recommendation
Mingxin Gan, Rui Jiang 0001
Expert Syst. Appl.2
2011 Uncover disease genes by maximizing information flow in the phenome-interactome network
abstract
MOTIVATION: Pinpointing genes that underlie human inherited diseases among candidate genes in susceptibility genetic regions is the primary step towards the understanding of pathogenesis of diseases. Although several probabilistic models have been proposed to prioritize candidate genes using phenotype similarities and protein-protein interactions, no combinatorial approaches have been proposed in the literature. RESULTS: We propose the first combinatorial approach for prioritizing candidate genes. We first construct a phenome-interactome network by integrating the given phenotype similarity profile, protein-protein interaction network and associations between diseases and genes. Then, we introduce a computational method called MAXIF to maximize the information flow in this network for uncovering genes that underlie diseases. We demonstrate the effectiveness of this method in prioritizing candidate genes through a series of cross-validation experiments, and we show the possibility of using this method to identify diseases with which a query gene may be associated. We demonstrate the competitive performance of our method through a comparison with two existing state-of-the-art methods, and we analyze the robustness of our method with respect to the parameters involved. As an example application, we apply our method to predict driver genes in 50 copy number aberration regions of melanoma. Our method is not only able to identify several driver genes that have been reported in the literature, it also shed some new biological insights on the understanding of the modular property and transcriptional regulation scheme of these driver genes. CONTACT: [email protected].
Tao Jiang 0001, Rui Jiang 0001
Bioinform.3
2011 Clustering 16S rRNA for OTU prediction: a method of unsupervised Bayesian clustering
abstract
MOTIVATION: With the advancements of next-generation sequencing technology, it is now possible to study samples directly obtained from the environment. Particularly, 16S rRNA gene sequences have been frequently used to profile the diversity of organisms in a sample. However, such studies are still taxed to determine both the number of operational taxonomic units (OTUs) and their relative abundance in a sample. RESULTS: To address these challenges, we propose an unsupervised Bayesian clustering method termed Clustering 16S rRNA for OTU Prediction (CROP). CROP can find clusters based on the natural organization of data without setting a hard cut-off threshold (3%/5%) as required by hierarchical clustering methods. By applying our method to several datasets, we demonstrate that CROP is robust against sequencing errors and that it produces more accurate results than conventional hierarchical clustering methods. AVAILABILITY AND IMPLEMENTATION: Source code freely available at the following URL: http://code.google.com/p/crop-tingchenlab/, implemented in C++ and supported on Linux and MS Windows.
Xiaolin Hao, Rui Jiang 0001
Bioinform.2
2011 Integrating multiple protein-protein interaction networks to prioritize disease genes: a Bayesian regression approach
abstract
BACKGROUND: The identification of genes responsible for human inherited diseases is one of the most challenging tasks in human genetics. Recent studies based on phenotype similarity and gene proximity have demonstrated great success in prioritizing candidate genes for human diseases. However, most of these methods rely on a single protein-protein interaction (PPI) network to calculate similarities between genes, and thus greatly restrict the scope of application of such methods. Meanwhile, independently constructed and maintained PPI networks are usually quite diverse in coverage and quality, making the selection of a suitable PPI network inevitable but difficult. METHODS: We adopt a linear model to explain similarities between disease phenotypes using gene proximities that are quantified by diffusion kernels of one or more PPI networks. We solve this model via a Bayesian approach, and we derive an analytic form for Bayes factor that naturally measures the strength of association between a query disease and a candidate gene and thus can be used as a score to prioritize candidate genes. This method is intrinsically capable of integrating multiple PPI networks. RESULTS: We show that gene proximities calculated from PPI networks imply phenotype similarities. We demonstrate the effectiveness of the Bayesian regression approach on five PPI networks via large scale leave-one-out cross-validation experiments and summarize the results in terms of the mean rank ratio of known disease genes and the area under the receiver operating characteristic curve (AUC). We further show the capability of our approach in integrating multiple PPI networks. CONCLUSIONS: The Bayesian regression approach can achieve much higher performance than the existing CIPHER approach and the ordinary linear regression method. The integration of multiple PPI networks can greatly improve the scope of application of the proposed method in the inference of disease genes.
Wangshu Zhang, Fengzhu Sun, Rui Jiang 0001
BMC Bioinform.3
2011 A Max-Flow-Based Approach to the Identification of Protein Complexes Using Protein Interaction and Microarray Data
abstract
The emergence of high-throughput technologies leads to abundant protein-protein interaction (PPI) data and microarray gene expression profiles, and provides a great opportunity for the identification of novel protein complexes using computational methods. By combining these two types of data, we propose a novel Graph Fragmentation Algorithm (GFA) for protein complex identification. Adapted from a classical max-flow algorithm for finding the (weighted) densest subgraphs, GFA first finds large (weighted) dense subgraphs in a protein-protein interaction network, and then, breaks each such subgraph into fragments iteratively by weighting its nodes appropriately in terms of their corresponding log-fold changes in the microarray data, until the fragment subgraphs are sufficiently small. Our tests on three widely used protein-protein interaction data sets and comparisons with several latest methods for protein complex identification demonstrate the strong performance of our method in predicting novel protein complexes in terms of its specificity and efficiency. Given the high specificity (or precision) that our method has achieved, we conjecture that our prediction results imply more than 200 novel protein complexes.
Jianxing Feng, Rui Jiang 0001, Tao Jiang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.2
2009 Align human interactome with phenome to identify causative genes and networks underlying disease families
abstract
MOTIVATION: Understanding the complexity in gene-phenotype relationship is vital for revealing the genetic basis of common diseases. Recent studies on the basis of human interactome and phenome not only uncovers prevalent phenotypic overlap and genetic overlap between diseases, but also reveals a modular organization of the genetic landscape of human diseases, providing new opportunities to reduce the complexity in dissecting the gene-phenotype association. RESULTS: We provide systematic and quantitative evidence that phenotypic overlap implies genetic overlap. With these results, we perform the first heterogeneous alignment of human interactome and phenome via a network alignment technique and identify 39 disease families with corresponding causative gene networks. Finally, we propose AlignPI, an alignment-based framework to predict disease genes, and identify plausible candidates for 70 diseases. Our method scales well to the whole genome, as demonstrated by prioritizing 6154 genes across 37 chromosome regions for Crohn's disease (CD). Results are consistent with a recent meta-analysis of genome-wide association studies for CD. AVAILABILITY: Bi-modules and disease gene predictions are freely available at the URL http://bioinfo.au.tsinghua.edu.cn/alignpi/
Xuebing Wu, Rui Jiang 0001
Bioinform.3
2009 A random forest approach to the detection of epistatic interactions in case-control studies
abstract
BACKGROUND: The key roles of epistatic interactions between multiple genetic variants in the pathogenesis of complex diseases notwithstanding, the detection of such interactions remains a great challenge in genome-wide association studies. Although some existing multi-locus approaches have shown their successes in small-scale case-control data, the "combination explosion" course prohibits their applications to genome-wide analysis. It is therefore indispensable to develop new methods that are able to reduce the search space for epistatic interactions from an astronomic number of all possible combinations of genetic variants to a manageable set of candidates. RESULTS: We studied case-control data from the viewpoint of binary classification. More precisely, we treated single nucleotide polymorphism (SNP) markers as categorical features and adopted the random forest to discriminate cases against controls. On the basis of the gini importance given by the random forest, we designed a sliding window sequential forward feature selection (SWSFS) algorithm to select a small set of candidate SNPs that could minimize the classification error and then statistically tested up to three-way interactions of the candidates. We compared this approach with three existing methods on three simulated disease models and showed that our approach is comparable to, sometimes more powerful than, the other methods. We applied our approach to a genome-wide case-control dataset for Age-related Macular Degeneration (AMD) and successfully identified two SNPs that were reported to be associated with this disease. CONCLUSION: Besides existing pure statistical approaches, we demonstrated the feasibility of incorporating machine learning methods into genome-wide case-control studies. The gini importance offers yet another measure for the associations between SNPs and complex diseases, thereby complementing existing statistical measures to facilitate the identification of epistatic interactions and the understanding of epistasis in the pathogenesis of complex diseases.
Rui Jiang 0001, Wanwan Tang, Xuebing Wu, Wenhui Fu
BMC Bioinform.1
2007 Optimal Adaptive Controller for Multidimensional Armax Model
abstract
The objective of this paper is to find an adaptive control strategy which would enable us to estimate the parameters of the ARMAX model as accurately as possible, along with consuming less controlling energy, while keeping the output of the system below a specified level of variability. Using the self-tuning tracker, this paper establishes global convergence of a stochastic adaptive control algorithm for discrete linear system, and the adaptive controller may converge to the one-step-ahead optimal controller at the same time.
Rui Jiang 0001, Kueiming Lo
Cybern. Syst.1
2006 Searching for interpretable rules for disease mutations: a simulated annealing bump hunting strategy
abstract
BACKGROUND: Understanding how amino acid substitutions affect protein functions is critical for the study of proteins and their implications in diseases. Although methods have been developed for predicting potential effects of amino acid substitutions using sequence, three-dimensional structural, and evolutionary properties of proteins, the applications are limited by the complication of the features and the availability of protein structural information. Another limitation is that the prediction results are hard to be interpreted with physicochemical principles and biological knowledge. RESULTS: To overcome these limitations, we proposed a novel feature set using physicochemical properties of amino acids, evolutionary profiles of proteins, and protein sequence information. We applied the support vector machine and the random forest with the feature set to experimental amino acid substitutions occurring in the E. coli lac repressor and the bacteriophage T4 lysozyme, as well as to annotated amino acid substitutions occurring in a wide range of human proteins. The results showed that the proposed feature set was superior to the existing ones. To explore physicochemical principles behind amino acid substitutions, we designed a simulated annealing bump hunting strategy to automatically extract interpretable rules for amino acid substitutions. We applied the strategy to annotated human amino acid substitutions and successfully extracted several rules which were either consistent with current biological knowledge or providing new insights for the understanding of amino acid substitutions. When applied to unclassified data, these rules could cover a large portion of samples, and most of the covered samples showed good agreement with predictions made by either the support vector machine or the random forest. CONCLUSION: The prediction methods using the proposed feature set can achieve larger AUC (the area under the ROC curve), smaller BER (the balanced error rate), and larger MCC (the Matthews' correlation coefficient) than those using the published feature sets, suggesting that our feature set is superior to the existing ones. The rules extracted by the simulated annealing bump hunting strategy have comparable coverage and accuracy but much better interpretability as those extracted by the patient rule induction method (PRIM), revealing that the strategy is more effective in inducing interpretable rules.
Rui Jiang 0001, Fengzhu Sun, Ting Chen 0006
BMC Bioinform.1