Jiajie Peng

dblp:130/7286 · DBLP profile ↗
← Back
67ranked-venue papers
21as first author
39since 2021 · last 2026
0000-0002-3857-7927ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 57 · 21 first-author · 32 since 2021Artificial intelligence and machine learning · 9 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 DiffuST: A Latent Diffusion Model for Spatial Transcriptomics Denoising
abstract
Spatial transcriptomics technologies have enabled comprehensive measurements of gene expression profiles while retaining spatial information, with most platforms also providing matched pathology images. However, noise resulting from low RNA capture efficiency and experimental steps needed to keep spatial information may corrupt the biological signals and obstruct analyses. Here, we develop a latent diffusion model DiffuST to denoise spatial transcriptomics. DiffuST employs a graph autoencoder and a pre-trained model to extract different-scale features from spatial information and pathology images. Then, a latent diffusion model is leveraged to map different scales of features to the same space for denoising. The evaluation based on various spatial transcriptomics datasets showed the superiority of DiffuST over existing denoising methods. Furthermore, the results demonstrated that DiffuST can enhance downstream analysis of spatial transcriptomics and yield significant biological insights.
Shaoqing Jiao, Dazhi Lu, Tao Wang 0082, Yongtian Wang, Yunwei Dong, Jiajie Peng
IEEE Trans. Comput. Biol. Bioinform.7
2025 SWAMamba: A Sliding Window Attention Mamba Framework for Predicting Translation Elongation Rates
abstract
Translation elongation is essential for cellular proteostasis and is implicated in cancer and neurodegeneration. Accurately predicting the rate of ribosome elongation in each codon (also called ribosomal A site) on mRNA is important for understanding and modulating protein synthesis. However, predicting elongation rates is challenging due to the trade-off between capturing distal codon interactions and focusing on proximal codon effects at the A site. Approaches capturing distal codon interactions in the coding sequences (CDS) of mRNA fail to effectively differentiate critical regions (codons near the A site) due to insufficient effective mechanisms for focusing on these regions. Conversely, due to the limitations of models when handling long mRNA sequences, some methods simplify inputs by conditioning solely on proximal codons surrounding the A site, leading to the loss of important information from distal codons. To address this issue, we leverage Mamba's success in capturing long-range dependencies to enable the consideration of distant codons' impact on the A site. Additionally, we introduce a sliding window attention mechanism to emphasize the proximal codons around the A site during ribosome elongation. Building on these advancements, we present Sliding Window Attention Mamba (SWAMamba), a novel framework that simultaneously leverages both proximal and distal codon effects on the A site. We conduct comprehensive evaluations on ribosome data across four species and find that SWAMamba significantly outperformed current state-of-the-art methods in predicting translation elongation rates.
Fei Ni 0001, Shaoqing Jiao, Dazhi Lu, Jianye Hao, Jiajie Peng
AAAI6
2025 BioRAGent: natural language biomedical querying with retrieval-augmented multiagent systems
abstract
Understanding the roles of genes, phenotypes, and diseases is crucial for advancing biomedical research. However, efficient and accessible retrieval of biomedical knowledge remains a challenge due to the complexity of the relevant data. We introduce BioRAGent, an intelligent biomedical assistant that combines Tool-augmented retrieval-augmented generation (RAG) with a multiagent system. Leveraging the ability of large language models, BioRAGent facilitates natural language queries about genes, phenotypes, diseases, and their interrelationships. BioRAGent employs three specialized agents: Guide (query optimization), Retriever (data retrieval), and Reviewer (answer validation) to access authoritative biomedical databases and to generate accurate responses. We evaluate the performance of BioRAGent on a benchmark of eleven single-hop and three multi-hop tasks, demonstrating superior results compared with state-of-the-art models. User evaluations highlight the practicality and robust user experience of BioRAGent, particularly in handling complex multi-hop queries. Moreover, ablation experiments validate the contribution of each agent in improving retrieval accuracy.
Manlian Bi, Zhijie Bao, Dongna Xie, Xiaohan Xie, Changxiao Yang, Tao Wang 0082, Yongtian Wang, Jiajie Peng
Briefings Bioinform.8
2025 The rise and potential opportunities of large language model agents in bioinformatics and biomedicine
abstract
Large language model (LLM) agents have demonstrated remarkable potential in the fields of bioinformatics and biomedicine. This paper reviews the technical foundations of LLM agents, including their core architecture, key technologies, and collaborative modes. We explore the applications of LLM agents in multi-omics, drug development, chemical research, clinical diagnosis, and health management. The paper also analyzes the major challenges faced by LLM agents, such as the interaction and extension of their frameworks, data privacy and security, model hallucinations and interpretability, timeliness of knowledge updates, and ethical and legal risks. Furthermore, we discuss future directions, including paradigms for human-artificial intelligence collaboration and the development of open-source ecosystems and standardization. This paper aims to provide a comprehensive perspective and guidance on the advancement of LLM agents in bioinformatics and biomedicine.
Yihang Xiao, Zhijie Bao, Jianye Hao, Jiajie Peng
Briefings Bioinform.5
2025 cfMethylPre: deep transfer learning enhances cancer detection based on circulating cell-free DNA methylation profiling
abstract
Cancer remains a significant global health burden, underscoring the need for innovative diagnostic tools to enable early detection and improve patient outcomes. While circulating cell-free DNA (cfDNA) methylation has emerged as a promising biomarker for noninvasive cancer diagnostics, existing methods often face limitations in handling the high-dimensionality of methylation data, small sample sizes, and a lack of biological interpretability. To address these challenges, we propose cfMethylPre, a novel deep transfer learning framework tailored for cancer detection using cfDNA methylation data. cfMethylPre leverages large language model pretrained embeddings from DNA sequence information and integrates them with methylation profiles to enhance feature representation. The deep transfer learning process involves pretraining on bulk DNA methylation data encompassing 2801 samples across 82 cancer types and normal controls, followed by fine-tuning with cfDNA methylation data. This approach ensures robust adaptation to cfDNA's unique characteristics while improving predictive accuracy. Our model achieved superior predictive accuracy compared with state-of-the-art methods, with a weighted Matthews Correlation Coefficient of 0.926 and a weighted F1-score of 0.942. Through model interpretation and biological experimental validation, we identified three novel breast cancer genes-PCDHA10, PRICKLE2, and PRTG-demonstrating their inhibitory effects on cell proliferation and migration in breast cancer cell lines. These findings establish cfMethylPre as a powerful and interpretable tool for cancer diagnostics and biological discovery, paving the way for its application in precision oncology.
Xuchao Zhang, Yongtian Wang, Jialu Hu, Jiajie Peng, Xuequn Shang 0001, Yanpu Wang, Tao Wang 0082
Briefings Bioinform.6
2025 MAEST: accurately spatial domain detection in spatial transcriptomics with graph masked autoencoder
abstract
Spatial transcriptomics (ST) technology provides gene expression profiles with spatial context, offering critical insights into cellular interactions and tissue architecture. A core task in ST is spatial domain identification, which involves detecting coherent regions with similar spatial expression patterns. However, existing methods often fail to fully exploit spatial information, leading to limited representational capacity and suboptimal clustering accuracy. Here, we introduce MAEST, a novel graph neural network model designed to address these limitations in ST data. MAEST leverages graph masked autoencoders to denoise and refine representations while incorporating graph contrastive learning to prevent feature collapse and enhance model robustness. By integrating one-hop and multi-hop representations, MAEST effectively captures both local and global spatial relationships, improving clustering precision. Extensive experiments across diverse datasets, including the human brain, mouse hippocampus, olfactory bulb, brain, and embryo, demonstrate that MAEST outperforms seven state-of-the-art methods in spatial domain identification. Furthermore, MAEST showcases its ability to integrate multi-slice data, identifying joint domains across horizontal tissue sections with high accuracy. These results highlight MAEST's versatility and effectiveness in unraveling the spatial organization of complex tissues. The source code of MAEST can be obtained at https://github.com/clearlove2333/MAEST.
Han Shu, Yongtian Wang, Jialu Hu, Jiajie Peng, Xuequn Shang 0001, Zhen Tian 0004, Tao Wang 0082
Briefings Bioinform.7
2025 Predicting CRISPR Cas9 Off-Target Activities With a Two-Stage Deep Learning Framework
abstract
The CRISPR-Cas9 system, found across bacteria and archaea, enables efficient genome engineering in eukaryotic cells. There is a challenge that Cas9 guide RNA may cause off-target activities. Although numbers of methods have been proposed to predict off-target activities in guide RNA designing, this procedure involves numerous potential off-target sites, which causes label imbalance problem. To address this problem, we developed a deep learning framework, named CAF-Net (Cas9 Augmentation and Finetune Network), for predicting off-target activities of CRISPR-Cas9. First, we pretrain an embedding model to extract features from target and guide sequence pairs. Subsequently, data augmentation is applied to these features. And finally, the model is finetuned with synthetic samples. Evaluation results demonstrate its performance on published datasets.
Tianshan Zhang, Bei Jiang, Zhijie Bao, Jiajie Peng
IEEE Trans. Comput. Biol. Bioinform.5
2025 EGA-Ploc: An Efficient Global-Local Attention Model for Multi-Label Protein Subcellular Localization Prediction on the Immunohistochemistry Images
abstract
Protein subcellular localization (PSL) is central to unraveling protein functions and disease mechanisms in bioinformatics. Immunohistochemistry (IHC) images serve as rich sources of high-resolution visual cues for PSL prediction. However, conventional deep learning approaches face critical limitations: whole-image models suffer irreversible fine-grained detail loss during downsampling, while patch-based methods lack effective global context integration. Additionally, the long-tailed class distribution in PSL datasets exacerbates performance degradation for underrepresented classes. To address these challenges, we present EGA-Ploc, a framework employing a linear attention mechanism optimized for high-resolution IHC images. This mechanism enables efficient global and local feature modeling with near-linear computational complexity, facilitating end-to-end processing of original images without resolution loss. Moreover, we propose an adaptive multi-label loss function that integrates zero-bounded log-sum-exp constraints with dynamic class-weighted compensation to mitigate dataset imbalance. Consequently, our EGA-Ploc achieves competitive performance across multiple PSL benchmarks while maintaining computational efficiency superior to existing methods. Through extensive visualization analysis, we further investigate the generalizability of off-the-shelf computer vision models in PSL, uncovering interpretable insights into their subcellular localization mechanisms.
Boyang Wan, Jiajie Peng
IEEE J. Biomed. Health Informatics4
2025 Integrative Graph-Based Framework for Predicting circRNA Drug Resistance Using Disease Contextualization and Deep Learning
abstract
Circular RNAs (circRNAs) play a crucial role in gene regulation and have been implicated in the development of drug resistance in cancer, representing a significant challenge in oncological therapeutics. Despite advancements in computational models predicting RNA-drug interactions, existing frameworks often overlook the complex interplay between circRNAs, drug mechanisms, and disease contexts. This study aims to bridge this gap by introducing a novel computational model, circRDRP, that enhances prediction accuracy by integrating disease-specific contexts into the analysis of circRNA-drug interactions. It employs a hybrid graph neural network that combines features from Graph Attention Networks (GAT) and Graph Convolutional Networks (GCN) in a two-layer structure, with further enhancement through convolutional neural networks. This approach allows for sophisticated feature extraction from integrated networks of circRNAs, drugs, and diseases. Our results demonstrate that the circRDRP model outperforms existing models in predicting drug resistance, showing significant improvements in accuracy, precision, and recall. Specifically, the model shows robust predictive capability in case studies involving major anticancer drugs such as Cisplatin and Methotrexate, indicating its potential utility in precision medicine. In conclusion, circRDRP offers a powerful tool for understanding and predicting drug resistance mediated by circRNAs, with implications for designing more effective cancer therapies.
Yongtian Wang, Wenkai Shen, Yewei Shen, Shang Feng, Tao Wang 0082, Xuequn Shang 0001, Jiajie Peng
IEEE J. Biomed. Health Informatics7
2024 Designing Biological Sequences without Prior Knowledge Using Evolutionary Reinforcement Learning
abstract
Designing novel biological sequences with desired properties is a significant challenge in biological science because of the extra large search space. The traditional design process usually involves multiple rounds of costly wet lab evaluations. To reduce the need for expensive wet lab experiments, machine learning methods are used to aid in designing biological sequences. However, the limited availability of biological sequences with known properties hinders the training of machine learning models, significantly restricting their applicability and performance. To fill this gap, we present ERLBioSeq, an Evolutionary Reinforcement Learning algorithm for BIOlogical SEQuence design. ERLBioSeq leverages the capability of reinforcement learning to learn without prior knowledge and the potential of evolutionary algorithms to enhance the exploration of reinforcement learning in the large search space of biological sequences. Additionally, to enhance the efficiency of biological sequence design, we developed a predictor for sequence screening in the biological sequence design process, which incorporates both the local and global sequence information. We evaluated the proposed method on three main types of biological sequence design tasks, including the design of DNA, RNA, and protein. The results demonstrate that the proposed method achieves significant improvement compared to the existing state-of-the-art methods.
Xiaotian Hao, Hongyao Tang, Zhentao Tang, Shaoqing Jiao, Dazhi Lu, Jiajie Peng
AAAI7
2024 Identifying brain disease genes via integrating brain imaging and molecular network
abstract
Identifying genes associate to brain diseases is crucial for uncovering the related biological mechanisms and for advancing therapeutic interventions. Due to the ongoing advancements in network-based computational approaches, biological networks, especially molecular networks, provide valuable insights for predicting disease genes. However, many methods have ignored brain imaging data when exploring brain disease genes, despite its widely use in neuroscience research. In this paper, we propose a novel framework, Deep Interactive AutoEncoder (DIAE), which integrates brain imaging and molecular-based gene networks to predict brain disease genes. DIAE first constructs a gene association network based on brain imaging and high-resolution whole brain-whole gene expression data. Subsequently, a deep and interactive multi-network integration method is introduced to learn low-dimensional features of genes by combining the brain imaging-based network with other molecular-based gene networks. Finally, these features are utilized to predict brain disease genes using a support vector machine (SVM) model. For performance evaluation, we compare DIAE with three existing state-of-the-art methods in the context of disease gene identification across four brain diseases. The experimental results show the superior performance of DIAE and highlight the effectiveness of the brain imaging-based gene network for predicting brain disease genes.
Yuxian Wang, Jiajie Peng
BIBM3
2024 Accurately deciphering spatial domains for spatially resolved transcriptomics with stCluster
abstract
Spatial transcriptomics provides valuable insights into gene expression within the native tissue context, effectively merging molecular data with spatial information to uncover intricate cellular relationships and tissue organizations. In this context, deciphering cellular spatial domains becomes essential for revealing complex cellular dynamics and tissue structures. However, current methods encounter challenges in seamlessly integrating gene expression data with spatial information, resulting in less informative representations of spots and suboptimal accuracy in spatial domain identification. We introduce stCluster, a novel method that integrates graph contrastive learning with multi-task learning to refine informative representations for spatial transcriptomic data, consequently improving spatial domain identification. stCluster first leverages graph contrastive learning technology to obtain discriminative representations capable of recognizing spatially coherent patterns. Through jointly optimizing multiple tasks, stCluster further fine-tunes the representations to be able to capture complex relationships between gene expression and spatial organization. Benchmarked against six state-of-the-art methods, the experimental results reveal its proficiency in accurately identifying complex spatial domains across various datasets and platforms, spanning tissue, organ, and embryo levels. Moreover, stCluster can effectively denoise the spatial gene expression patterns and enhance the spatial trajectory inference. The source code of stCluster is freely available at https://github.com/hannshu/stCluster.
Tao Wang 0082, Han Shu, Jialu Hu, Yongtian Wang, Jin Chen 0004, Jiajie Peng, Xuequn Shang 0001
Briefings Bioinform.6
2024 Predicting drug-target binding affinity with cross-scale graph contrastive learning
abstract
Identifying the binding affinity between a drug and its target is essential in drug discovery and repurposing. Numerous computational approaches have been proposed for understanding these interactions. However, most existing methods only utilize either the molecular structure information of drugs and targets or the interaction information of drug-target bipartite networks. They may fail to combine the molecule-scale and network-scale features to obtain high-quality representations. In this study, we propose CSCo-DTA, a novel cross-scale graph contrastive learning approach for drug-target binding affinity prediction. The proposed model combines features learned from the molecular scale and the network scale to capture information from both local and global perspectives. We conducted experiments on two benchmark datasets, and the proposed model outperformed existing state-of-art methods. The ablation experiment demonstrated the significance and efficacy of multi-scale features and cross-scale contrastive learning modules in improving the prediction performance. Moreover, we applied the CSCo-DTA to predict the novel potential targets for Erlotinib and validated the predicted targets with the molecular docking analysis.
Yihang Xiao, Xuequn Shang 0001, Jiajie Peng
Briefings Bioinform.4
2024 Vulnerability detection based on transformer and high-quality number embedding
abstract
Summary Software vulnerability detection is an important problem in software security. In recent years, deep learning offers a novel approach for source code vulnerability detection. Due to the similarities between programming languages and natural languages, many natural language processing techniques have been applied to vulnerability detection tasks. However, specific problems within vulnerability detection tasks, such as buffer overflow, involve numerical reasoning. For these problems, the model needs to not only consider long dependencies and multiple relationships between statements of code but also capture the magnitude property of numerical literals in the program through high‐quality number embeddings. Therefore, we propose VDTransformer, a Transformer‐based method that improves source code embedding by integrating word and number embeddings. Furthermore, we employ Transformer encoders to construct a hierarchical neural network that extracts semantic features from the code and enables line‐level vulnerability detection. To evaluate the effectiveness of the method, we construct a dataset named OverflowGen based on templates for buffer overflow. Experimental comparisons on OverflowGen with a well‐known static vulnerability detector and two state‐of‐the‐art deep learning‐based methods confirm the effectiveness of VDTransformer and the importance of high‐quality number embeddings in vulnerability detection tasks involving numerical features.
Yunwei Dong, Jiajie Peng
Concurr. Comput. Pract. Exp.3
2023 HQProtoPNet: An Evidence-Based Model for Interpretable Image Recognition
abstract
In image recognition, improving the interpretability of the recognition model can help people understand the model better and increase the trust of human beings for model prediction. The prototype-based interpretable model is a self-explanatory image recognition model that simulates the evidence reasoning used in human recognition. Each prototype is evidence that contains category features, which can help in determining the image category. Based on the prototype-based model, this paper introduces a deep interpretable network architecture called the high-quality prototypical part network (HQProtoPNet). Compared to existing work, this paper adds random erasing to enhance the picture, helping to improve prototype generation and increase model prediction. The multiple scale conversion operation is also introduced and the similarity calculation is improved to make the prototype have multiscale information and matching ability. Furthermore, the accuracy of HQProtoPNet can reach or even exceed the accuracy of several black-box models. Additionally, due to the improvement in the quality of the prototype, the model's prediction accuracy is improved by stacking without reducing the interpretability of the stacked model, which gives the model real stackability.
Jiajie Peng, Zhiming Liu 0001, Hengjun Zhao
IJCNN2
2023 Collaborative deep learning improves disease-related circRNA prediction based on multi-source functional information
abstract
Emerging studies have shown that circular RNAs (circRNAs) are involved in a variety of biological processes and play a key role in disease diagnosing, treating and inferring. Although many methods, including traditional machine learning and deep learning, have been developed to predict associations between circRNAs and diseases, the biological function of circRNAs has not been fully exploited. Some methods have explored disease-related circRNAs based on different views, but how to efficiently use the multi-view data about circRNA is still not well studied. Therefore, we propose a computational model to predict potential circRNA-disease associations based on collaborative learning with circRNA multi-view functional annotations. First, we extract circRNA multi-view functional annotations and build circRNA association networks, respectively, to enable effective network fusion. Then, a collaborative deep learning framework for multi-view information is designed to get circRNA multi-source information features, which can make full use of the internal relationship among circRNA multi-view information. We build a network consisting of circRNAs and diseases by their functional similarity and extract the consistency description information of circRNAs and diseases. Last, we predict potential associations between circRNAs and diseases based on graph auto encoder. Our computational model has better performance in predicting candidate disease-related circRNAs than the existing ones. Furthermore, it shows the high practicability of the method that we use several common diseases as case studies to find some unknown circRNAs related to them. The experiments show that CLCDA can efficiently predict disease-related circRNAs and are helpful for the diagnosis and treatment of human disease.
Yongtian Wang, Xinmeng Liu, Yewei Shen, Xuerui Song, Tao Wang 0082, Xuequn Shang 0001, Jiajie Peng
Briefings Bioinform.7
2023 scMultiGAN: cell-specific imputation for single-cell transcriptomes with multiple deep generative adversarial networks
abstract
The emergence of single-cell RNA sequencing (scRNA-seq) technology has revolutionized the identification of cell types and the study of cellular states at a single-cell level. Despite its significant potential, scRNA-seq data analysis is plagued by the issue of missing values. Many existing imputation methods rely on simplistic data distribution assumptions while ignoring the intrinsic gene expression distribution specific to cells. This work presents a novel deep-learning model, named scMultiGAN, for scRNA-seq imputation, which utilizes multiple collaborative generative adversarial networks (GAN). Unlike traditional GAN-based imputation methods that generate missing values based on random noises, scMultiGAN employs a two-stage training process and utilizes multiple GANs to achieve cell-specific imputation. Experimental results show the efficacy of scMultiGAN in imputation accuracy, cell clustering, differential gene expression analysis and trajectory analysis, significantly outperforming existing state-of-the-art techniques. Additionally, scMultiGAN is scalable to large scRNA-seq datasets and consistently performs well across sequencing platforms. The scMultiGAN code is freely available at https://github.com/Galaxy8172/scMultiGAN.
Tao Wang 0082, Yungang Xu, Yongtian Wang, Xuequn Shang 0001, Jiajie Peng, Bing Xiao 0001
Briefings Bioinform.6
2023 A benchmark for automatic medical consultation system: frameworks, tasks and datasets
abstract
MOTIVATION: In recent years, interest has arisen in using machine learning to improve the efficiency of automatic medical consultation and enhance patient experience. In this article, we propose two frameworks to support automatic medical consultation, namely doctor-patient dialogue understanding and task-oriented interaction. We create a new large medical dialogue dataset with multi-level fine-grained annotations and establish five independent tasks, including named entity recognition, dialogue act classification, symptom label inference, medical report generation and diagnosis-oriented dialogue policy. RESULTS: We report a set of benchmark results for each task, which shows the usability of the dataset and sets a baseline for future studies. AVAILABILITY AND IMPLEMENTATION: Both code and data are available from https://github.com/lemuria-wchen/imcs21. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wei Chen 0088, Hongyi Fang, Qianyuan Yao, Jianye Hao, Qi Zhang 0001, Xuanjing Huang 0001, Jiajie Peng, Zhongyu Wei
Bioinform.9
2023 DxFormer: a decoupled automatic diagnostic system based on decoder-encoder transformer with dense symptom representations
abstract
MOTIVATION: Symptom-based automatic diagnostic system queries the patient's potential symptoms through continuous interaction with the patient and makes predictions about possible diseases. A few studies use reinforcement learning (RL) to learn the optimal policy from the joint action space of symptoms and diseases. However, existing RL (or Non-RL) methods focus on disease diagnosis while ignoring the importance of symptom inquiry. Although these systems have achieved considerable diagnostic accuracy, they are still far below its performance upper bound due to few turns of interaction with patients and insufficient performance of symptom inquiry. To address this problem, we propose a new automatic diagnostic framework called DxFormer, which decouples symptom inquiry and disease diagnosis, so that these two modules can be independently optimized. The transition from symptom inquiry to disease diagnosis is parametrically determined by the stopping criteria. In DxFormer, we treat each symptom as a token, and formalize the symptom inquiry and disease diagnosis to a language generation model and a sequence classification model, respectively. We use the inverted version of Transformer, i.e. the decoder-encoder structure, to learn the representation of symptoms by jointly optimizing the reinforce reward and cross-entropy loss. RESULTS: We conduct experiments on three real-world medical dialogue datasets, and the experimental results verify the feasibility of increasing diagnostic accuracy by improving symptom recall. Our model overcomes the shortcomings of previous RL-based methods. By decoupling symptom query from the process of diagnosis, DxFormer greatly improves the symptom recall and achieves the state-of-the-art diagnostic accuracy. AVAILABILITY AND IMPLEMENTATION: Both code and data are available at https://github.com/lemuria-wchen/DxFormer. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Wei Chen 0088, Jiajie Peng, Zhongyu Wei
Bioinform.3
2023 DFinder: a novel end-to-end graph embedding-based method to identify drug-food interactions
abstract
MOTIVATION: Drug-food interactions (DFIs) occur when some constituents of food affect the bioaccessibility or efficacy of the drug by involving in drug pharmacodynamic and/or pharmacokinetic processes. Many computational methods have achieved remarkable results in link prediction tasks between biological entities, which show the potential of computational methods in discovering novel DFIs. However, there are few computational approaches that pay attention to DFI identification. This is mainly due to the lack of DFI data. In addition, food is generally made up of a variety of chemical substances. The complexity of food makes it difficult to generate accurate feature representations for food. Therefore, it is urgent to develop effective computational approaches for learning the food feature representation and predicting DFIs. RESULTS: In this article, we first collect DFI data from DrugBank and PubMed, respectively, to construct two datasets, named DrugBank-DFI and PubMed-DFI. Based on these two datasets, two DFI networks are constructed. Then, we propose a novel end-to-end graph embedding-based method named DFinder to identify DFIs. DFinder combines node attribute features and topological structure features to learn the representations of drugs and food constituents. In topology space, we adopt a simplified graph convolution network-based method to learn the topological structure features. In feature space, we use a deep neural network to extract attribute features from the original node attributes. The evaluation results indicate that DFinder performs better than other baseline methods. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/23AIBox/23AIBox-DFinder. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tao Wang 0082, Jinjin Yang, Yifu Xiao, Yuxian Wang, Yongtian Wang, Jiajie Peng
Bioinform.8
2022 Deep Representation Debiasing via Mutual Information Minimization and Maximization (Student Abstract)
abstract
Deep representation learning has succeeded in several fields. However, pre-trained deep representations are usually biased and make downstream models sensitive to different attributes. In this work, we propose a post-processing unsupervised deep representation debiasing algorithm, DeepMinMax, which can obtain unbiased representations directly from pre-trained representations without re-training or fine-tuning the entire model. The experimental results on synthetic and real-world datasets indicate that DeepMinMax outperforms the existing state-of-the-art algorithms on downstream tasks.
Ruijiang Han, Yuxi Long, Jiajie Peng
AAAI4
2022 Hypergraph-based Gene Ontology Embedding for Disease Gene Prediction
abstract
Disease gene identification has provided valuable insights into illuminating the molecular mechanisms underlying complex diseases. And it has been shown that novel drugs with genetically supported targets were more likely to be successful in clinical trials. In recent years, multiple graph machine learning-based methods for this purpose have been proposed. However, those methods were mainly based on various well-established biological molecular networks, while seldomly considering the curated biological annotations of genes. To fill this gap, we aim to integrate the gene ontology annotations (GOA), including the biological process (BP), the cellular component (CC), and the molecular function (MF), into the process of disease gene prediction. Our method treated the GOA as a hypergraph and used the hypergraph-based embedding technique to extract the deep features underlying gene annotations. Besides, we also extracted gene features from the protein-protein interaction (PPI) network using graph representation learning methods. The convolutional neural network (CNN) framework was followed to fuse the features extracted from the two networks and make the final prediction. Experiments on a range of diseases have demonstrated the accuracy and robustness of our method. The average area under the receiver operating characteristic curve (AUROC) and area under the precision-recall curve (AUPRC) reached 0.85 and 0.79, respectively. Besides, the hypergraph-based gene ontology embedding can be generalized to other bioinformatics applications.
Tao Wang 0082, Hengbo Xu, Ranye Zhang, Yifu Xiao, Jiajie Peng, Xuequn Shang 0001
BIBM5
2022 Discovering eQTL Regulatory Patterns Through eQTLMotif
abstract
The expression quantitative trait loci (eQTL) analysis has become important for understanding the regulatory function of genomic variants on gene expression in a tissuespecific manner and has been widely applied across species from microbes to mammals. Current eQTL studies mainly focus on the simple one-to-one regulation between variant and gene. Recent research have demonstrated there are also more complex regulatory patterns between eQTLs and genes. However, there is a lack of studies and relevant methods to systematically discover the regulatory patterns between multiple eQTLs and multiple genes. In this regard, this study has proposed a novel computational framework, called eQTLMotif, to discover regulation patterns of eQTLs in a many-to-many manner. This framework mainly consists of two steps: (1) construct a novel eQTL regulatory network by integrating bipartite eQTL network, eQTL mediation effects, and gene regulatory network; (2) perform motif mining through exactly enumerating frequently appeared eQTL regulatory structures. Based on this framework, we for the first time systematically investigated the eQTL regulatory patterns in the human frontal cortex based on a large cohort of postmortem human brains. Experiments have demonstrated that our framework can effectively reveal novel eQTL regulatory patterns. And some are in similar structure to the existing gene regulation patterns, such as feed-forward loop (FFL)-like motif, single input module (SIM)-like motif, and dense overlapping regulons (DOR)- like motif. Our method and findings will further enhance the understanding of regulatory mechanisms of eQTLs in multiple tissues and species.
Tao Wang 0082, Yifu Xiao, Hanzi Yang, Xipeng Yin, Yongtian Wang, Bing Xiao 0001, Xuequn Shang 0001, Jiajie Peng
BIBM9
2022 A Two Stage Adaptation Framework for Frame Detection via Prompt Learning
abstract
Framing is a communication strategy to bias discussion by selecting and emphasizing. Frame detection aims to automatically analyze framing strategy. Previous works on frame detection mainly focus on a single scenario or issue, ignoring the special characteristics of frame detection that new events emerge continuously and policy agenda changes dynamically. To better deal with various context and frame typologies across different issues, we propose a two-stage adaptation framework. In the framing domain adaptation from pre-training stage, we design two tasks based on pivots and prompts to learn a transferable encoder, verbalizer, and prompts. In the downstream scenario generalization stage, the transferable components are applied to new issues and label sets. Experiment results demonstrate the effectiveness of our framework in different scenarios. Also, it shows superiority both in full-resource and low-resource conditions.
Xinyi Mou, Zhongyu Wei, Changjian Jiang, Jiajie Peng
COLING4
2022 A network-based method for brain disease gene prediction by integrating brain connectome and molecular network
abstract
Brain disease gene identification is critical for revealing the biological mechanism and developing drugs for brain diseases. To enhance the identification of brain disease genes, similarity-based computational methods, especially network-based methods, have been adopted for narrowing down the searching space. However, these network-based methods only use molecular networks, ignoring brain connectome data, which have been widely used in many brain-related studies. In our study, we propose a novel framework, named brainMI, for integrating brain connectome data and molecular-based gene association networks to predict brain disease genes. For the consistent representation of molecular-based network data and brain connectome data, brainMI first constructs a novel gene network, called brain functional connectivity (BFC)-based gene network, based on resting-state functional magnetic resonance imaging data and brain region-specific gene expression data. Then, a multiple network integration method is proposed to learn low-dimensional features of genes by integrating the BFC-based gene network and existing protein-protein interaction networks. Finally, these features are utilized to predict brain disease genes based on a support vector machine-based model. We evaluate brainMI on four brain diseases, including Alzheimer's disease, Parkinson's disease, major depressive disorder and autism. brainMI achieves of 0.761, 0.729, 0.728 and 0.744 using the BFC-based gene network alone and enhances the molecular network-based performance by 6.3% on average. In addition, the results show that brainMI achieves higher performance in predicting brain disease genes compared to the existing three state-of-the-art methods.
Ruijiang Han, Menghan Zhang, Yuxian Wang, Tao Wang 0082, Yongtian Wang, Xuequn Shang 0001, Jiajie Peng
Briefings Bioinform.8
2022 Enhancing discoveries of molecular QTL studies with small sample size using summary statistic imputation
abstract
Quantitative trait locus (QTL) analyses of multiomic molecular traits, such as gene transcription (eQTL), DNA methylation (mQTL) and histone modification (haQTL), have been widely used to infer the functional effects of genome variants. However, the QTL discovery is largely restricted by the limited study sample size, which demands higher threshold of minor allele frequency and then causes heavy missing molecular trait-variant associations. This happens prominently in single-cell level molecular QTL studies because of sample availability and cost. It is urgent to propose a method to solve this problem in order to enhance discoveries of current molecular QTL studies with small sample size. In this study, we presented an efficient computational framework called xQTLImp to impute missing molecular QTL associations. In the local-region imputation, xQTLImp uses multivariate Gaussian model to impute the missing associations by leveraging known association statistics of variants and the linkage disequilibrium (LD) around. In the genome-wide imputation, novel procedures are implemented to improve efficiency, including dynamically constructing a reused LD buffer, adopting multiple heuristic strategies and parallel computing. Experiments on various multiomic bulk and single-cell sequencing-based QTL datasets have demonstrated high imputation accuracy and novel QTL discovery ability of xQTLImp. Finally, a C++ software package is freely available at https://github.com/stormlovetao/QTLIMP.
Tao Wang 0082, Yongzhuang Liu, Quanwei Yin, Jiaquan Geng, Jin Chen 0004, Xipeng Yin, Yongtian Wang, Xuequn Shang 0001, Chunwei Tian, Yadong Wang 0001, Jiajie Peng
Briefings Bioinform.11
2022 Correction to: Enhancing discoveries of molecular QTL studies with small sample size using summary statistic imputation
abstract
In the originally published version of this manuscript, there was an error in the Funding section; ‘National Natural Science Foundation of China (6210071334, 62072376)’ has now been corrected to ‘National Natural Science Foundation of China (62102319, 62072376)’.
Tao Wang 0082, Yongzhuang Liu, Quanwei Yin, Jiaquan Geng, Jin Chen 0004, Xipeng Yin, Yongtian Wang, Xuequn Shang 0001, Chunwei Tian, Yadong Wang 0001, Jiajie Peng
Briefings Bioinform.11
2022 A review and performance evaluation of clustering frameworks for single-cell Hi-C data
abstract
The three-dimensional genome structure plays a key role in cellular function and gene regulation. Single-cell Hi-C (high-resolution chromosome conformation capture) technology can capture genome structure information at the cell level, which provides the opportunity to study how genome structure varies among different cell types. Recently, a few methods are well designed for single-cell Hi-C clustering. In this manuscript, we perform an in-depth benchmark study of available single-cell Hi-C data clustering methods to implement an evaluation system for multiple clustering frameworks based on both human and mouse datasets. We compare eight methods in terms of visualization and clustering performance. Performance is evaluated using four benchmark metrics including adjusted rand index, normalized mutual information, homogeneity and Fowlkes-Mallows index. Furthermore, we also evaluate the eight methods for the task of separating cells at different stages of the cell cycle based on single-cell Hi-C data.
Caiwei Zhen, Yuxian Wang, Jiaquan Geng, Jinghao Peng, Tao Wang 0082, Jianye Hao, Xuequn Shang 0001, Zhongyu Wei, Peican Zhu, Jiajie Peng
Briefings Bioinform.12
2022 Hierarchical reinforcement learning for automatic disease diagnosis
abstract
MOTIVATION: Disease diagnosis-oriented dialog system models the interactive consultation procedure as the Markov decision process, and reinforcement learning algorithms are used to solve the problem. Existing approaches usually employ a flat policy structure that treat all symptoms and diseases equally for action making. This strategy works well in a simple scenario when the action space is small; however, its efficiency will be challenged in the real environment. Inspired by the offline consultation process, we propose to integrate a hierarchical policy structure of two levels into the dialog system for policy learning. The high-level policy consists of a master model that is responsible for triggering a low-level model, the low-level policy consists of several symptom checkers and a disease classifier. The proposed policy structure is capable to deal with diagnosis problem including large number of diseases and symptoms. RESULTS: Experimental results on three real-world datasets and a synthetic dataset demonstrate that our hierarchical framework achieves higher accuracy and symptom recall in disease diagnosis compared with existing systems. We construct a benchmark including datasets and implementation of existing algorithms to encourage follow-up researches. AVAILABILITY AND IMPLEMENTATION: The code and data are available from https://github.com/FudanDISC/DISCOpen-MedBox-DialoDiagnosis. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kangenbei Liao, Wei Chen 0088, Qianlong Liu, Baolin Peng, Xuanjing Huang 0001, Jiajie Peng, Zhongyu Wei
Bioinform.7
2022 Flexibility and rigidity index for chromosome packing, flexibility and dynamics analysis
Jiajie Peng, Jinjin Yang, D. Vijay Anand, Xuequn Shang 0001, Kelin Xia
Frontiers Comput. Sci.1
2021 DCAE: Selecting Discriminative Genes on Single-cell RNA-seq Data for Cell-type Quantification
abstract
Tumor-liltrating lymphocytes (TILs) are predictive for response to neoadjuvant treatment in tumors. Still, the abundance of tumor-liltrating cell types has not yet been produced in large quantities, hampering researchers exploring their characteristics. As the levels of genomics or transcriptomics could reflect changes in cell-type proportions, several computational tools have been developed to estimate cell-type abundances based on the reference gene expression proliles. Differential expression analysis is the most widely used to recognize marker genes. However, it ignores the correlation between genes. To this end, we propose a feature selection method, dubbed Discriminative Concrete Autoencoder (DCAE), to identify informative genes on single-cell RNA-seq data, which are then used to quantity cell-type proportions. To evaluate the performance of DCAE on selecting discriminative genes, we conduct experiments on our collected and processed single-cell RNA-seq dataset. First, we compare DCAE to the original Concrete Autoencoder by the cell-type classification accuracies resulting from their selected genes. Then we infer cell-type abundance by using deconvolution function with the chosen small cohort of genes. Next, we evaluate the deconvolution accuracy by the Pearson correlation coefficient between the estimated cell-type proportions and the true proportions, and the corresponding P-value. Finally, we compare the effects of the selected genes and the differential expression genes on the deconvolution accuracy. The results show that our selected genes by DCAE have higher discriminant power to distinguish cell types and effectively infer cell-type abundance. Thus, DCAE provides insights into acquiring candidate biomarkers for cell-type quantification.
Shuhui Liu, Jiajie Peng, Xuequn Shang 0001
BIBM3
2021 Predicting Hepatoma-Related Genes Based on Representation Learning of PPI network and Gene Ontology Annotations
abstract
Hepatoma is the most common type of primary liver cancer with a high mortality rate in the world. The genetic causes of the disease pathology remain largely unknown. Effective discovery of the genes associated with hepatoma has become important in disease prevention, early diagnosis, and therapeutic treatments. With the developments of molecular networks, graph-based methods have been tremendously successful in predicting disease genes based on the hypothesis of guilt-by-association. Network representation learning (NRL) techniques have accelerated disease gene discovery in recent years because of their powerful network feature extraction ability. However, the current network representation learning-based methods for disease gene discovery did not consider the gene features derived from gene ontology annotations, which apriori group genes with similar functions. To fill this gap, here we propose a novel framework to predict hepatoma-related genes based on representation learning from both protein-protein interactions (PPI) network and gene ontology annotations. Our framework has three steps: learning features from PPI network and gene ontologies using NRL techniques, integrating different features based on autoencoder, predicting hepatoma-related genes using machine learning classifiers. Experiments have demonstrated that our framework could accurately predict hepatoma-related genes with AUROC and AUPRC reaching 0.93 and 0.94, respectively. Compared with other methods using only single representation features, our framework also shows superior performance on hepatoma gene prediction.
Tao Wang 0082, Zhiyuan Shao, Yifu Xiao, Xuchao Zhang, Binze Shi, Siyu Chen 0024, Yuxian Wang, Jiajie Peng, Xuequn Shang 0001
BIBM9
2021 Curriculum Learning for Vision-and-Language Navigation
abstract
Vision-and-Language Navigation (VLN) is a task where an agent navigates in an embodied indoor environment under human instructions. Previous works ignore the distribution of sample difficulty and we argue that this potentially degrade their agent performance. To tackle this issue, we propose a novel curriculum- based training paradigm for VLN tasks that can balance human prior knowledge and agent learning progress about training samples. We develop the principle of curriculum design and re-arrange the benchmark Room-to-Room (R2R) dataset to make it suitable for curriculum training. Experiments show that our method is model-agnostic and can significantly improve the performance, the generalizability, and the training efficiency of current state-of-the-art navigation agents without increasing model complexity.
Jiwen Zhang, Zhongyu Wei, Jianqing Fan, Jiajie Peng
NeurIPS4
2021 An end-to-end heterogeneous graph representation learning-based framework for drug-target interaction prediction
abstract
Accurately identifying potential drug-target interactions (DTIs) is a key step in drug discovery. Although many related experimental studies have been carried out for identifying DTIs in the past few decades, the biological experiment-based DTI identification is still timeconsuming and expensive. Therefore, it is of great significance to develop effective computational methods for identifying DTIs. In this paper, we develop a novel 'end-to-end' learning-based framework based on heterogeneous 'graph' convolutional networks for 'DTI' prediction called end-to-end graph (EEG)-DTI. Given a heterogeneous network containing multiple types of biological entities (i.e. drug, protein, disease, side-effect), EEG-DTI learns the low-dimensional feature representation of drugs and targets using a graph convolutional networks-based model and predicts DTIs based on the learned features. During the training process, EEG-DTI learns the feature representation of nodes in an end-to-end mode. The evaluation test shows that EEG-DTI performs better than existing state-of-art methods. The data and source code are available at: https://github.com/MedicineBiology-AI/EEG-DTI.
Jiajie Peng, Yuxian Wang, Jiaojiao Guan, Ruijiang Han, Jianye Hao, Zhongyu Wei, Xuequn Shang 0001
Briefings Bioinform.1
2021 Integrating multi-network topology for gene function prediction using deep neural networks
abstract
MOTIVATION: The emergence of abundant biological networks, which benefit from the development of advanced high-throughput techniques, contributes to describing and modeling complex internal interactions among biological entities such as genes and proteins. Multiple networks provide rich information for inferring the function of genes or proteins. To extract functional patterns of genes based on multiple heterogeneous networks, network embedding-based methods, aiming to capture non-linear and low-dimensional feature representation based on network biology, have recently achieved remarkable performance in gene function prediction. However, existing methods do not consider the shared information among different networks during the feature learning process. RESULTS: Taking the correlation among the networks into account, we design a novel semi-supervised autoencoder method to integrate multiple networks and generate a low-dimensional feature representation. Then we utilize a convolutional neural network based on the integrated feature embedding to annotate unlabeled gene functions. We test our method on both yeast and human datasets and compare with three state-of-the-art methods. The results demonstrate the superior performance of our method. We not only provide a comprehensive analysis of the performance of the newly proposed algorithm but also provide a tool for extracting features of genes based on multiple networks, which can be used in the downstream machine learning task. AVAILABILITY: DeepMNE-CNN is freely available at https://github.com/xuehansheng/DeepMNE-CNN. CONTACT: [email protected]; [email protected]; [email protected].
Jiajie Peng, Hansheng Xue, Zhongyu Wei, Idil Tuncali, Jianye Hao, Xuequn Shang 0001
Briefings Bioinform.1
2021 Identifying drug-target interactions based on graph convolutional network and deep neural network
abstract
Identification of new drug-target interactions (DTIs) is an important but a time-consuming and costly step in drug discovery. In recent years, to mitigate these drawbacks, researchers have sought to identify DTIs using computational approaches. However, most existing methods construct drug networks and target networks separately, and then predict novel DTIs based on known associations between the drugs and targets without accounting for associations between drug-protein pairs (DPPs). To incorporate the associations between DPPs into DTI modeling, we built a DPP network based on multiple drugs and proteins in which DPPs are the nodes and the associations between DPPs are the edges of the network. We then propose a novel learning-based framework, 'graph convolutional network (GCN)-DTI', for DTI identification. The model first uses a graph convolutional network to learn the features for each DPP. Second, using the feature representation as an input, it uses a deep neural network to predict the final label. The results of our analysis show that the proposed framework outperforms some state-of-the-art approaches by a large margin.
Tianyi Zhao 0001, Yang Hu 0008, Linda R. Valsdottir, Tianyi Zang, Jiajie Peng
Briefings Bioinform.5
2021 Prediction and collection of protein-metabolite interactions
abstract
Interactions between proteins and small molecule metabolites play vital roles in regulating protein functions and controlling various cellular processes. The activities of metabolic enzymes, transcription factors, transporters and membrane receptors can all be mediated through protein-metabolite interactions (PMIs). Compared with the rich knowledge of protein-protein interactions, little is known about PMIs. To the best of our knowledge, no existing database has been developed for collecting PMIs. The recent rapid development of large-scale mass spectrometry analysis of biomolecules has led to the discovery of large amounts of PMIs. Therefore, we developed the PMI-DB to provide a comprehensive and accurate resource of PMIs. A total of 49 785 entries were manually collected in the PMI-DB, corresponding to 23 small molecule metabolites, 9631 proteins and 4 species. Unlike other databases that only provide positive samples, the PMI-DB provides non-interaction between proteins and metabolites, which not only reduces the experimental cost for biological experimenters but also facilitates the construction of more accurate algorithms for researchers using machine learning. To show the convenience of the PMI-DB, we developed a deep learning-based method to predict PMIs in the PMI-DB and compared it with several methods. The experimental results show that the area under the curve and area under the precision-recall curve of our method are 0.88 and 0.95, respectively. Overall, the PMI-DB provides a user-friendly interface for browsing the biological functions of metabolites/proteins of interest, and experimental techniques for identifying PMIs in different species, which provides important support for furthering the understanding of cellular processes. The PMI-DB is freely accessible at http://easybioai.com/PMIDB.
Tianyi Zhao 0001, Tianyi Zang, Jiajie Peng
Briefings Bioinform.7
2021 A novel method for predicting cell abundance based on single-cell RNA-seq data
abstract
BACKGROUND: It is important to understand the composition of cell type and its proportion in intact tissues, as changes in certain cell types are the underlying cause of disease in humans. Although compositions of cell type and ratios can be obtained by single-cell sequencing, single-cell sequencing is currently expensive and cannot be applied in clinical studies involving a large number of subjects. Therefore, it is useful to apply the bulk RNA-Seq dataset and the single-cell RNA dataset to deconvolute and obtain the cell type composition in the tissue. RESULTS: By analyzing the existing cell population prediction methods, we found that most of the existing methods need the cell-type-specific gene expression profile as the input of the signature matrix. However, in real applications, it is not always possible to find an available signature matrix. To solve this problem, we proposed a novel method, named DCap, to predict cell abundance. DCap is a deconvolution method based on non-negative least squares. DCap considers the weight resulting from measurement noise of bulk RNA-seq and calculation error of single-cell RNA-seq data, during the calculation process of non-negative least squares and performs the weighted iterative calculation based on least squares. By weighting the bulk tissue gene expression matrix and single-cell gene expression matrix, DCap minimizes the measurement error of bulk RNA-Seq and also reduces errors resulting from differences in the number of expressed genes in the same type of cells in different samples. Evaluation test shows that DCap performs better in cell type abundance prediction than existing methods. CONCLUSION: DCap solves the deconvolution problem using weighted non-negative least squares to predict cell type abundance in tissues. DCap has better prediction results and does not need to prepare a signature matrix that gives the cell-type-specific gene expression profile in advance. By using DCap, we can better study the changes in cell proportion in diseased tissues and provide more information on the follow-up treatment of diseases.
Jiajie Peng, Xuequn Shang 0001
BMC Bioinform.1
2021 A pipeline for RNA-seq based eQTL analysis with automated quality control procedures
abstract
BACKGROUND: Advances in the expression quantitative trait loci (eQTL) studies have provided valuable insights into the mechanism of diseases and traits-associated genetic variants. However, it remains challenging to evaluate and control the quality of multi-source heterogeneous eQTL raw data for researchers with limited computational background. There is an urgent need to develop a powerful and user-friendly tool to automatically process the raw datasets in various formats and perform the eQTL mapping afterward. RESULTS: In this work, we present a pipeline for eQTL analysis, termed eQTLQC, featured with automated data preprocessing for both genotype data and gene expression data. Our pipeline provides a set of quality control and normalization approaches, and utilizes automated techniques to reduce manual intervention. We demonstrate the utility and robustness of this pipeline by performing eQTL case studies using multiple independent real-world datasets with RNA-seq data and whole genome sequencing (WGS) based genotype data. CONCLUSIONS: eQTLQC provides a reliable computational workflow for eQTL analysis. It provides standard quality control and normalization as well as eQTL mapping procedures for eQTL raw data in multiple formats. The source code, demo data, and instructions are freely available at https://github.com/stormlovetao/eQTLQC .
Tao Wang 0082, Yongzhuang Liu, Junpeng Ruan, Xianjun Dong, Yadong Wang 0001, Jiajie Peng
BMC Bioinform.6
2020 Efficient Deep Reinforcement Learning via Adaptive Policy Transfer
abstract
Transfer learning has shown great potential to accelerate Reinforcement Learning (RL) by leveraging prior knowledge from past learned policies of relevant tasks. Existing approaches either transfer previous knowledge by explicitly computing similarities between tasks or select appropriate source policies to provide guided explorations. However, how to directly optimize the target policy by alternatively utilizing knowledge from appropriate source policies without explicitly measuring the similarities is currently missing. In this paper, we propose a novel Policy Transfer Framework (PTF) by taking advantage of this idea. PTF learns when and which source policy is the best to reuse for the target policy and when to terminate it by modeling multi-policy transfer as an option learning problem. PTF can be easily combined with existing DRL methods and experimental results show it significantly accelerates RL and surpasses state-of-the-art policy transfer methods in terms of learning efficiency and final performance in both discrete and continuous action spaces.
Tianpei Yang, Jianye Hao, Zhaopeng Meng, Zongzhang Zhang, Yujing Hu, Changjie Fan, Weixun Wang, Wulong Liu, Zhaodong Wang, Jiajie Peng
IJCAI11
2020 Identifying emerging phenomenon in long temporal phenotyping experiments
abstract
MOTIVATION: The rapid improvement of phenotyping capability, accuracy and throughput have greatly increased the volume and diversity of phenomics data. A remaining challenge is an efficient way to identify phenotypic patterns to improve our understanding of the quantitative variation of complex phenotypes, and to attribute gene functions. To address this challenge, we developed a new algorithm to identify emerging phenomena from large-scale temporal plant phenotyping experiments. An emerging phenomenon is defined as a group of genotypes who exhibit a coherent phenotype pattern during a relatively short time. Emerging phenomena are highly transient and diverse, and are dependent in complex ways on both environmental conditions and development. Identifying emerging phenomena may help biologists to examine potential relationships among phenotypes and genotypes in a genetically diverse population and to associate such relationships with the change of environments or development. RESULTS: We present an emerging phenomenon identification tool called Temporal Emerging Phenomenon Finder (TEP-Finder). Using large-scale longitudinal phenomics data as input, TEP-Finder first encodes the complicated phenotypic patterns into a dynamic phenotype network. Then, emerging phenomena in different temporal scales are identified from dynamic phenotype network using a maximal clique based approach. Meanwhile, a directed acyclic network of emerging phenomena is composed to model the relationships among the emerging phenomena. The experiment that compares TEP-Finder with two state-of-art algorithms shows that the emerging phenomena identified by TEP-Finder are more functionally specific, robust and biologically significant. AVAILABILITY AND IMPLEMENTATION: The source code, manual and sample data of TEP-Finder are all available at: http://phenomics.uky.edu/TEP-Finder/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiajie Peng, Junya Lu, Donghee Hoh, Ayesha S. Dina, Xuequn Shang 0001, David M. Kramer 0001, Jin Chen 0004
Bioinform.1
2020 DeepLGP: a novel deep learning method for prioritizing lncRNA target genes
abstract
MOTIVATION: Although long non-coding RNAs (lncRNAs) have limited capacity for encoding proteins, they have been verified as biomarkers in the occurrence and development of complex diseases. Recent wet-lab experiments have shown that lncRNAs function by regulating the expression of protein-coding genes (PCGs), which could also be the mechanism responsible for causing diseases. Currently, lncRNA-related biological data are increasing rapidly. Whereas, no computational methods have been designed for predicting the novel target genes of lncRNA. RESULTS: In this study, we present a graph convolutional network (GCN) based method, named DeepLGP, for prioritizing target PCGs of lncRNA. First, gene and lncRNA features were selected, these included their location in the genome, expression in 13 tissues and miRNA-mediated lncRNA-gene pairs. Next, GCN was applied to convolve a gene interaction network for encoding the features of genes and lncRNAs. Then, these features were used by the convolutional neural network for prioritizing target genes of lncRNAs. In 10-cross validations on two independent datasets, DeepLGP obtained high area under curves (0.90-0.98) and area under precision-recall curves (0.91-0.98). We found that lncRNA pairs with high similarity had more overlapped target genes. Further experiments showed that genes targeted by the same lncRNA sets had a strong likelihood of causing the same diseases, which could help in identifying disease-causing PCGs. AVAILABILITY AND IMPLEMENTATION: https://github.com/zty2009/LncRNA-target-gene. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Tianyi Zhao 0001, Yang Hu 0008, Jiajie Peng, Liang Cheng 0006, Pier Luigi Martelli
Bioinform.3
2020 Combining sequence and network information to enhance protein-protein interaction prediction
abstract
BACKGROUND: Protein-protein interactions (PPIs) are of great importance in cellular systems of organisms, since they are the basis of cellular structure and function and many essential cellular processes are related to that. Most proteins perform their functions by interacting with other proteins, so predicting PPIs accurately is crucial for understanding cell physiology. RESULTS: Recently, graph convolutional networks (GCNs) have been proposed to capture the graph structure information and generate representations for nodes in the graph. In our paper, we use GCNs to learn the position information of proteins in the PPIs networks graph, which can reflect the properties of proteins to some extent. Combining amino acid sequence information and position information makes a stronger representation for protein, which improves the accuracy of PPIs prediction. CONCLUSION: In previous research methods, most of them only used protein amino acid sequence as input information to make predictions, without considering the structural information of PPIs networks graph. We first time combine amino acid sequence information and position information to make representations for proteins. The experimental results indicate that our method has strong competitiveness compared with several sequence-based methods.
Leilei Liu, Xianglei Zhu, Yi Ma 0005, Haiyin Piao, Yaodong Yang 0002, Xiaotian Hao, Jiajie Peng
BMC Bioinform.9
2020 A learning-based method for drug-target interaction prediction based on feature representation learning and deep neural network
abstract
BACKGROUND: Drug-target interaction prediction is of great significance for narrowing down the scope of candidate medications, and thus is a vital step in drug discovery. Because of the particularity of biochemical experiments, the development of new drugs is not only costly, but also time-consuming. Therefore, the computational prediction of drug target interactions has become an essential way in the process of drug discovery, aiming to greatly reducing the experimental cost and time. RESULTS: We propose a learning-based method based on feature representation learning and deep neural network named DTI-CNN to predict the drug-target interactions. We first extract the relevant features of drugs and proteins from heterogeneous networks by using the Jaccard similarity coefficient and restart random walk model. Then, we adopt a denoising autoencoder model to reduce the dimension and identify the essential features. Third, based on the features obtained from last step, we constructed a convolutional neural network model to predict the interaction between drugs and proteins. The evaluation results show that the average AUROC score and AUPR score of DTI-CNN were 0.9416 and 0.9499, which obtains better performance than the other three existing state-of-the-art methods. CONCLUSIONS: All the experimental results show that the performance of DTI-CNN is better than that of the three existing methods and the proposed method is appropriately designed.
Jiajie Peng, Xuequn Shang 0001
BMC Bioinform.1
2020 Mining Relationships among Multiple Entities in Biological Networks
abstract
Identifying topological relationships among multiple entities in biological networks is critical towards the understanding of the organizational principles of network functionality. Theoretically, this problem can be solved using minimum Steiner tree (MSTT) algorithms. However, due to large network size, it remains to be computationally challenging, and the predictive value of multi-entity topological relationships is still unclear. We present a novel solution called Cluster-based Steiner Tree Miner (CST-Miner) to instantly identify multi-entity topological relationships in biological networks. Given a list of user-specific entities, CST-Miner decomposes a biological network into nested cluster-based subgraphs, on which multiple minimum Steiner trees are identified. By merging all of them into a minimum cost tree, the optimal topological relationships among all the user-specific entities are revealed. Experimental results showed that CST-Miner can finish in nearly log-linear time and the tree constructed by CST-Miner is close to the global minimum.
Jiajie Peng, Linjiao Zhu, Yadong Wang 0001, Jin Chen 0004
IEEE ACM Trans. Comput. Biol. Bioinform.1
2019 Towards Gene Function Prediction via Multi-Networks Representation Learning
abstract
Multi-networks integration methods have achieved prominent performance on many network-based tasks, but these approaches often incur information loss problem. In this paper, we propose a novel multi-networks representation learning method based on semi-supervised autoencoder, termed as DeepMNE, which captures complex topological structures of each network and takes the correlation among multinetworks into account. The experimental results on two realworld datasets indicate that DeepMNE outperforms the existing state-of-the-art algorithms.
Hansheng Xue, Jiajie Peng, Xuequn Shang 0001
AAAI2
2019 Integrating Sequence and Network Information to Enhance Protein-Protein Interaction Prediction Using Graph Convolutional Networks
abstract
Identification of protein-protein interactions (PPIs) is an important problem in biology, since PPIs are related to many essential cellular processes. The development of large-scale high-throughput experiments has produced a large number of PPIs data, however, these data are often noisy and their coverage is still limited. To overcome the shortcomings of experimental methods, many computational methods have been proposed for the prediction of PPIs. Among these methods, most of them solely take the amino acid sequence of protein as input information to make predictions. As PPIs data form the PPIs networks graph, the position information of proteins in the graph can reflect the properties of proteins to some extent, which is an important complement to protein sequence information. But previous works did not consider the graph structure information to improve the prediction performance. In this work, we first time apply graph convolutional networks (GCNs) to capture the protein's position information in the graph and combine amino acid sequence information and position information to make representations in the prediction task. Our experimental results show that our work outperforms the state-of-the-art sequence-based methods on several benchmark datasets and our work computationally is more efficient compared with previous works.
Leilei Liu, Yi Ma 0005, Xianglei Zhu, Yaodong Yang 0002, Xiaotian Hao, Jiajie Peng
BIBM7
2019 A deconvolution method for predicting cell abundance based on single cell RNA-seq data
abstract
It is important to understand the cell-type composition and its proportion in tissues. Some previous experiments have shown that the variation of gene expression in certain cell types may lead to disease. Although cell type composition and proportion can be obtained by single-cell RNA-sequencing (scRNA-sq), using scRNA-sq is expensive and cannot be applied in clinical studies involving a large number of subjects currently. Therefore, it is urgent to develop a method to deconvolute the Bulk RNA-Seq data to obtain the cell type composition in the tissue. Most of the existing methods require the signature matrix, which provides the cell type-specific gene expression profile, as input. However, the signature matrix is not always available for some types of tissue, and it is not always possible to find a suitable cell type-specific gene expression profile. To solve this problem, we propose a novel method, named DCap, to predict cell abundance. Different from non-negative least squares, DCap performs weighted iterative calculation based on least squares. By weighting bulk tissue gene expression matrix and single-cell gene expression matrix, DCap minimizes the measurement error of Bulk RNA-Seq and error resulting from the difference in the amount of genes in the same cell type among different samples. DCap solves the deconvolution problem by using weighted non-negative least squares to predict cell type abundance. DCap does not need to prepare a suitable signature matrix in advance, and the evaluation test shows that DCap performs better in cell type abundance prediction than existing methods.
Jiajie Peng, Xuequn Shang 0001
BIBM1
2019 Integrating Multi-Network Topology via Deep Semi-supervised Node Embedding
abstract
Node Embedding, which uses low-dimensional non-linear feature vectors to represent nodes in the network, has shown a great promise, not only because it is easy-to-use for downstream tasks, but also because it has achieved great success on many network analysis tasks. One of the challenges has been how to develop a node embedding method for integrating topological information from multiple networks. To address this critical problem, we propose a novel node embedding, called DeepMNE, for multi-network integration using a deep semi-supervised autoencoder. The key point of DeepMNE is that it captures complex topological structures of multiple networks and utilizes correlation among multiple networks as constraints. We evaluate DeepMNE in node classification task and link prediction task on four real-world datasets. The experimental results demonstrate that DeepMNE shows superior performance over seven state-of-the-art single-network and multi-network embedding algorithms.
Hansheng Xue, Jiajie Peng, Jiying Li, Xuequn Shang 0001
CIKM2
2019 A learning-based framework for miRNA-disease association identification using neural networks
abstract
MOTIVATION: A microRNA (miRNA) is a type of non-coding RNA, which plays important roles in many biological processes. Lots of studies have shown that miRNAs are implicated in human diseases, indicating that miRNAs might be potential biomarkers for various types of diseases. Therefore, it is important to reveal the relationships between miRNAs and diseases/phenotypes. RESULTS: We propose a novel learning-based framework, MDA-CNN, for miRNA-disease association identification. The model first captures interaction features between diseases and miRNAs based on a three-layer network including disease similarity network, miRNA similarity network and protein-protein interaction network. Then, it employs an auto-encoder to identify the essential feature combination for each pair of miRNA and disease automatically. Finally, taking the reduced feature representation as input, it uses a convolutional neural network to predict the final label. The evaluation results show that the proposed framework outperforms some state-of-the-art approaches in a large margin on both tasks of miRNA-disease association prediction and miRNA-phenotype association prediction. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/Issingjessica/MDA-CNN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiajie Peng, Weiwei Hui, Jianye Hao, Qinghua Jiang, Xuequn Shang 0001, Zhongyu Wei
Bioinform.1
2019 TS-GOEA: a web tool for tissue-specific gene set enrichment analysis based on gene ontology
abstract
BACKGROUND: The Gene Ontology (GO) knowledgebase is the world's largest source of information on the functions of genes. Since the beginning of GO project, various tools have been developed to perform GO enrichment analysis experiments. GO enrichment analysis has become a commonly used method of gene function analysis. Existing GO enrichment analysis tools do not consider tissue-specific information, although this information is very important to current research. RESULTS: In this paper, we built an easy-to-use web tool called TS-GOEA that allows users to easily perform experiments based on tissue-specific GO enrichment analysis. TS-GOEA uses strict threshold statistical method for GO enrichment analysis, and provides statistical tests to improve the reliability of the analysis results. Meanwhile, TS-GOEA provides tools to compare different experimental results, which is convenient for users to compare the experimental results. To evaluate its performance, we tested the genes associated with platelet disease with TS-GOEA. CONCLUSIONS: TS-GOEA is an effective GO analysis tool with unique features. The experimental results show that our method has better performance and provides a useful supplement for the existing GO enrichment analysis tools. TS-GOEA is available at http://120.77.47.2:5678.
Jiajie Peng, Guilin Lu, Hansheng Xue, Tao Wang 0082, Xuequn Shang 0001
BMC Bioinform.1
2019 Combining gene ontology with deep neural networks to enhance the clustering of single cell RNA-Seq data
abstract
BACKGROUND: Single cell RNA sequencing (scRNA-seq) is applied to assay the individual transcriptomes of large numbers of cells. The gene expression at single-cell level provides an opportunity for better understanding of cell function and new discoveries in biomedical areas. To ensure that the single-cell based gene expression data are interpreted appropriately, it is crucial to develop new computational methods. RESULTS: In this article, we try to re-construct a neural network based on Gene Ontology (GO) for dimension reduction of scRNA-seq data. By integrating GO with both unsupervised and supervised models, two novel methods are proposed, named GOAE (Gene Ontology AutoEncoder) and GONN (Gene Ontology Neural Network) respectively. CONCLUSIONS: The evaluation results show that the proposed models outperform some state-of-the-art dimensionality reduction approaches. Furthermore, incorporating with GO, we provide an opportunity to interpret the underlying biological mechanism behind the neural network-based model.
Jiajie Peng, Xuequn Shang 0001
BMC Bioinform.1
2019 Prioritizing candidate diseases-related metabolites based on literature and functional similarity
abstract
BACKGROUND: As the terminal products of cellular regulatory process, functional related metabolites have a close relationship with complex diseases, and are often associated with the same or similar diseases. Therefore, identification of disease related metabolites play a critical role in understanding comprehensively pathogenesis of disease, aiming at improving the clinical medicine. Considering that a large number of metabolic markers of diseases need to be explored, we propose a computational model to identify potential disease-related metabolites based on functional relationships and scores of referred literatures between metabolites. First, obtaining associations between metabolites and diseases from the Human Metabolome database, we calculate the similarities of metabolites based on modified recommendation strategy of collaborative filtering utilizing the similarities between diseases. Next, a disease-associated metabolite network (DMN) is built with similarities between metabolites as weight. To improve the ability of identifying disease-related metabolites, we introduce scores of text mining from the existing database of chemicals and proteins into DMN and build a new disease-associated metabolite network (FLDMN) by fusing functional associations and scores of literatures. Finally, we utilize random walking with restart (RWR) in this network to predict candidate metabolites related to diseases. RESULTS: We construct the disease-associated metabolite network and its improved network (FLDMN) with 245 diseases, 587 metabolites and 28,715 disease-metabolite associations. Subsequently, we extract training sets and testing sets from two different versions of the Human Metabolome database and assess the performance of DMN and FLDMN on 19 diseases, respectively. As a result, the average AUC (area under the receiver operating characteristic curve) of DMN is 64.35%. As a further improved network, FLDMN is proven to be successful in predicting potential metabolic signatures for 19 diseases with an average AUC value of 76.03%. CONCLUSION: In this paper, a computational model is proposed for exploring metabolite-disease pairs and has good performance in predicting potential metabolites related to diseases through adequate validation. This result suggests that integrating literature and functional associations can be an effective way to construct disease associated metabolite network for prioritizing candidate diseases-related metabolites.
Yongtian Wang, Liran Juan, Jiajie Peng, Tianyi Zang, Yadong Wang 0001
BMC Bioinform.3
2019 LncDisAP: a computation model for LncRNA-disease association prediction based on multiple biological datasets
abstract
BACKGROUND: Over the past decades, a large number of long non-coding RNAs (lncRNAs) have been identified. Growing evidence has indicated that the mutation and dysregulation of lncRNAs play a critical role in the development of many complex human diseases. Consequently, identifying potential disease-related lncRNAs is an effective means to improve the quality of disease diagnostics and treatment, which is the motivation of this work. Here, we propose a computational model (LncDisAP) for potential disease-related lncRNA identification based on multiple biological datasets. First, the associations between lncRNA and different data sources are collected from different databases. With these data sources as dimensions, we calculate the functional associations between lncRNAs by the recommendation strategy of collaborative filtering. Subsequently, a disease-associated lncRNA functional network is built with functional similarities between lncRNAs as the weight. Ultimately, potential disease-related lncRNAs can be identified based on ranked scores derived by random walking with restart (RWR). Then, training sets and testing sets are extracted from two different versions of a disease-lncRNA dataset to assess the performance of LncDisAP on 54 diseases. RESULTS: A lncRNA functional network is built based on the proposed computational model, and it contains 66,060 associations among 364 lncRNAs associated with 182 diseases in total. We extract 218 known disease-lncRNA pairs associated with 54 diseases to assess the network. As a result, the average AUC (area under the receiver operating characteristic curve) of LncDisAP is 78.08%. CONCLUSION: In this article, a computational model integrating multiple lncRNA-related biological datasets is proposed for identifying potential disease-related lncRNAs. The result shows that LncDisAP is successful in predicting novel disease-related lncRNA signatures. In addition, with several common cancers taken as case studies, we found some unknown lncRNAs that could be associated with these diseases through our network. These results suggest that this method can be helpful in improving the quality for disease diagnostics and treatment.
Yongtian Wang, Liran Juan, Jiajie Peng, Tianyi Zang, Yadong Wang 0001
BMC Bioinform.3
2018 TSGOE: A web tool for tissue-specific gene ontology enrichment
Jiajie Peng, Guilin Lu, Hansheng Xue, Tao Wang 0082, Xuequn Shang 0001
BIBM1
2018 Predicting candidate disease-related lncRNAs based on network random walk
Yongtian Wang, Liran Juan, Jiajie Peng, Tianyi Zang, Yadong Wang 0001
BIBM3
2018 Identifying Representative Network Motifs for Inferring Higher-order Structure of Biological Networks
Tao Wang 0082, Jiajie Peng, Yadong Wang 0001, Jin Chen 0004
BIBM2
2018 Measuring phenotype-phenotype similarity through the interactome
abstract
BACKGROUND: Recently, measuring phenotype similarity began to play an important role in disease diagnosis. Researchers have begun to pay attention to develop phenotype similarity measurement. However, existing methods ignore the interactions between phenotype-associated proteins, which may lead to inaccurate phenotype similarity. RESULTS: We proposed a network-based method PhenoNet to calculate the similarity between phenotypes. We localized phenotypes in the network and calculated the similarity between phenotype-associated modules by modeling both the inter- and intra-similarity. CONCLUSIONS: PhenoNet was evaluated on two independent evaluation datasets: gene ontology and gene expression data. The result shows that PhenoNet performs better than the state-of-art methods on all evaluation tests.
Jiajie Peng, Weiwei Hui, Xuequn Shang 0001
BMC Bioinform.1
2017 Measuring phenotype-phenotype similarity through the interactome
abstract
Recently, measuring phenotype similarity began to play an important role in disease diagnosis. Researchers have begun to pay attention to develop phenotype similarity measurement. However, existing methods ignore the interactions between phenotype-associated proteins, which may lead to inaccurate phenotype similarity. We proposed a network-based method PhenoNet to calculate the similarity between phenotypes. We localized phenotypes in the network and calculated the similarity between phenotype-associated modules by modeling both the inter- and intra-similarity. PhenoNet was evaluated on two independent evaluation datasets: gene ontology and gene expression data. The result shows that PhenoNet performs better than the state-of-art methods on all evaluation tests.
Jiajie Peng, Weiwei Hui, Xuequn Shang 0001
BIBM1
2017 Identifying term relations cross different gene ontology categories
abstract
BACKGROUND: The Gene Ontology (GO) is a community-based bioinformatics resource that employs ontologies to represent biological knowledge and describes information about gene and gene product function. GO includes three independent categories: molecular function, biological process and cellular component. For better biological reasoning, identifying the biological relationships between terms in different categories are important. However, the existing measurements to calculate similarity between terms in different categories are either developed by using the GO data only or only take part of combined gene co-function network information. RESULTS: We propose an iterative ranking-based method called C r o G O2 to measure the cross-categories GO term similarities by incorporating level information of GO terms with both direct and indirect interactions in the gene co-function network. CONCLUSIONS: The evaluation test shows that C r o G O2 performs better than the existing methods. A genome-specific term association network for yeast is also generated by connecting terms with the high confidence score. The linkages in the term association network could be supported by the literature. Given a gene set, the related terms identified by using the association network have overlap with the related terms identified by GO enrichment analysis.
Jiajie Peng, Junya Lu, Weiwei Hui, Yadong Wang 0001, Xuequn Shang 0001
BMC Bioinform.1
2016 Analyzing factors involved in the HPO-based semantic similarity calculation
abstract
Although disease diagnosis have greatly benefited from next generation sequencing technologies, it is still difficult to make the right diagnosis based on purely sequencing technologies for many diseases with complex phenotypes and high genetic heterogeneity. Recently, calculating Human Phenotype Ontology (HPO)-based phenotype semantic similarity has contributed a lot for completing disease diagnosis. However, factors which affect the accuracy of HPO-based semantic similarity have not been evaluated systematically. In this study, we propose a new framework called HPOFactor to evaluate these factors.
Jiajie Peng, Jialu Hu, Xuequn Shang 0001
BIBM1
2016 Measuring phenotype semantic similarity using Human Phenotype Ontology
abstract
It is critical yet remains to be challenging to make right disease diagnosis based on complex clinical characteristic and heterogeneous genetic background. Recently, Human Phenotype Ontology (HPO)-based phenotype similarity has been widely used to aid disease diagnosis. However, the existing measurements are revised based on the Gene Ontology-based term similarity models, which are not optimized for human phenotype ontologies. We propose a new similarity measure called PhenoSim. Our model includes a noise reduction component to model the noisy patient phenotype data, and a path-constrained Information Content-based method for measuring phenotype semantics similarity. Evaluation tests showed that PhenoSim could improve the performance of HPO-based phenotype similarity measurement.
Jiajie Peng, Hansheng Xue, Yukai Shao, Xuequn Shang 0001, Yadong Wang 0001, Jin Chen 0004
BIBM1
2016 Joint detection of copy number variations in parent-offspring trios
abstract
MOTIVATION: Whole genome sequencing (WGS) of parent-offspring trios is a powerful approach for identifying disease-associated genes via detecting copy number variations (CNVs). Existing approaches, which detect CNVs for each individual in a trio independently, usually yield low-detection accuracy. Joint modeling approaches leveraging Mendelian transmission within the parent-offspring trio can be an efficient strategy to improve CNV detection accuracy. RESULTS: In this study, we developed TrioCNV, a novel approach for jointly detecting CNVs in parent-offspring trios from WGS data. Using negative binomial regression, we modeled the read depth signal while considering both GC content bias and mappability bias. Moreover, we incorporated the family relationship and used a hidden Markov model to jointly infer CNVs for three samples of a parent-offspring trio. Through application to both simulated data and a trio from 1000 Genomes Project, we showed that TrioCNV achieved superior performance than existing approaches. AVAILABILITY AND IMPLEMENTATION: The software TrioCNV implemented using a combination of Java and R is freely available from the website at https://github.com/yongzhuang/TrioCNV CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yongzhuang Liu, Jianguo Lu, Jiajie Peng, Liran Juan, Xiaolin Zhu 0002, Bingshan Li, Yadong Wang 0001
Bioinform.4
2016 Extending gene ontology with gene association networks
abstract
MOTIVATION: Gene ontology (GO) is a widely used resource to describe the attributes for gene products. However, automatic GO maintenance remains to be difficult because of the complex logical reasoning and the need of biological knowledge that are not explicitly represented in the GO. The existing studies either construct whole GO based on network data or only infer the relations between existing GO terms. None is purposed to add new terms automatically to the existing GO. RESULTS: We proposed a new algorithm 'GOExtender' to efficiently identify all the connected gene pairs labeled by the same parent GO terms. GOExtender is used to predict new GO terms with biological network data, and connect them to the existing GO. Evaluation tests on biological process and cellular component categories of different GO releases showed that GOExtender can extend new GO terms automatically based on the biological network. Furthermore, we applied GOExtender to the recent release of GO and discovered new GO terms with strong support from literature. AVAILABILITY AND IMPLEMENTATION: Software and supplementary document are available at www.msu.edu/%7Ejinchen/GOExtender CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jiajie Peng, Tao Wang 0082, Jixuan Wang, Yadong Wang 0001, Jin Chen 0004
Bioinform.1
2015 Measuring semantic similarities by combining gene ontology annotations and gene co-function networks
abstract
BACKGROUND: Gene Ontology (GO) has been used widely to study functional relationships between genes. The current semantic similarity measures rely only on GO annotations and GO structure. This limits the power of GO-based similarity because of the limited proportion of genes that are annotated to GO in most organisms. RESULTS: We introduce a novel approach called NETSIM (network-based similarity measure) that incorporates information from gene co-function networks in addition to using the GO structure and annotations. Using metabolic reaction maps of yeast, Arabidopsis, and human, we demonstrate that NETSIM can improve the accuracy of GO term similarities. We also demonstrate that NETSIM works well even for genomes with sparser gene annotation data. We applied NETSIM on large Arabidopsis gene families such as cytochrome P450 monooxygenases to group the members functionally and show that this grouping could facilitate functional characterization of genes in these families. CONCLUSIONS: Using NETSIM as an example, we demonstrated that the performance of a semantic similarity measure could be significantly improved after incorporating genome-specific information. NETSIM incorporates both GO annotations and gene co-function network data as a priori knowledge in the model. Therefore, functional similarities of GO terms that are not explicitly encoded in GO but are relevant in a taxon-specific manner become measurable when GO annotations are limited. Supplementary information and software are available at http://www.msu.edu/~jinchen/NETSIM .
Jiajie Peng, Sahra Uygun, Taehyong Kim, Yadong Wang 0001, Seung Y. Rhee, Jin Chen 0004
BMC Bioinform.1
2014 Towards integrative gene functional similarity measurement
abstract
BACKGROUND: In Gene Ontology, the "Molecular Function" (MF) categorization is a widely used knowledge framework for gene function comparison and prediction. Its structure and annotation provide a convenient way to compare gene functional similarities at the molecular level. The existing gene similarity measures, however, solely rely on one or few aspects of MF without utilizing all the rich information available including structure, annotation, common terms, lowest common parents. RESULTS: We introduce a rank-based gene semantic similarity measure called InteGO by synergistically integrating the state-of-the-art gene-to-gene similarity measures. By integrating three GO based seed measures, InteGO significantly improves the performance by about two-fold in all the three species studied (yeast, Arabidopsis and human). CONCLUSIONS: InteGO is a systematic and novel method to study gene functional associations. The software and description are available at http://www.msu.edu/~jinchen/InteGO.
Jiajie Peng, Yadong Wang 0001, Jin Chen 0004
BMC Bioinform.1
2013 Identifying cross-category relations in gene ontology and constructing genome-specific term association networks
abstract
BACKGROUND: Gene Ontology (GO) has been widely used in biological databases, annotation projects, and computational analyses. Although the three GO categories are structured as independent ontologies, the biological relationships across the categories are not negligible for biological reasoning and knowledge integration. However, the existing cross-category ontology term similarity measures are either developed by utilizing the GO data only or based on manually curated term name similarities, ignoring the fact that GO is evolving quickly and the gene annotations are far from complete. RESULTS: In this paper we introduce a new cross-category similarity measurement called CroGO by incorporating genome-specific gene co-function network data. The performance study showed that our measurement outperforms the existing algorithms. We also generated genome-specific term association networks for yeast and human. An enrichment based test showed our networks are better than those generated by the other measures. CONCLUSIONS: The genome-specific term association networks constructed using CroGO provided a platform to enable a more consistent use of GO. In the networks, the frequently occurred MF-centered hub indicates that a molecular function may be shared by different genes in multiple biological processes, or a set of genes with the same functions may participate in distinct biological processes. And common subgraphs in multiple organisms also revealed conserved GO term relationships. Software and data are available online at http://www.msu.edu/~jinchen/CroGO.
Jiajie Peng, Jin Chen 0004, Yadong Wang 0001
BMC Bioinform.1