VLDB 2026 Research / reviewers in the wild / expert
Qinghua Jiang
dblp:79/7238
· DBLP profile ↗
29ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 27 · 1 first-author · 11 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AI-driven computational methods and benchmarking for T-cell antigen identificationabstractThe rise of mRNA vaccines highlights the pivotal role of T-cell antigen identification in modern vaccinology and personalized medicine. T-cell recognition relies on the sophisticated ternary interaction between the T-cell receptor (TCR), the major histocompatibility complex (MHC) molecule, and the peptide antigen, which forms the peptide-MHC (pMHC) complex. Computational methods, particularly artificial intelligence (AI), are indispensable for accurately predicting these complex bindings. This review systematically surveys the rapidly evolving AI-driven landscape for T-cell antigen identification, providing a comprehensive categorization of methods for MHC-I, MHC-II, and the highly complex TCR-pMHC binding prediction, alongside foundational data resources. Crucially, we conduct a rigorous, standardized benchmarking of 18 state-of-the-art TCR-pMHC prediction models across diverse training data sources. Our evaluation on two distinct and challenging out-of-distribution (OOD) unseen epitope variant datasets reveals a significant and concerning generalization gap in current predictors. Notably, the overall absolute predictive gain remains marginal across all models under OOD conditions. This result underscores a severe and persistent generalization challenge when faced with novel epitope variants. To address these limitations, we emphasize the urgent need for enhanced structural modeling, the integration of multi-omics data, and the development of generative models for de novo TCR design. By advancing these computational frontiers, our community can accelerate the transition from prediction to rational design in immunoinformatics. Jinhao Que, Guangfu Xue, Yideng Cai, Wenyi Yang, Yi Hui, Zuxiang Wang, Wenyang Zhou, Qinghua Jiang, Haoxiu Sun |
Briefings Bioinform. | 12 |
| 2026 | Intelligent fault diagnosis based on TimeGAN with channel-temporal attention module
Zengkai Liu, Yunsai Chen, Xuewei Shi, Qinghua Jiang |
Neurocomputing | 6 |
| 2025 | TriCLFF: a multi-modal feature fusion framework using contrastive learning for spatial domain identificationabstractSpatial transcriptomics (ST) encompasses rich multi-modal information related to cell state and organization. Precisely identifying spatial domains with consistent gene expression patterns and histological features is a critical task in ST analysis, which requires comprehensive integration of multi-modal information. Here, we propose TriCLFF, a contrastive learning-based multi-modal feature fusion framework, to effectively integrate spatial associations, gene expression levels, and histological features in a unified manner. Leveraging an advanced feature fusion mechanism, our proposed TriCLFF framework outperforms existing state-of-the-art methods in terms of accuracy and robustness across four datasets (mouse brain anterior, mouse olfactory bulb, human dorsolateral prefrontal cortex, and human breast cancer) from different platforms (10x Visium and Stereo-seq) for spatial domain identification. TriCLFF also facilitates the identification of finer-grained structures in breast cancer tissues and detects previously unknown gene expression patterns in the human dorsolateral prefrontal cortex, providing novel insights for understanding tissue functions. Overall, TriCLFF establishes an effective paradigm for integrating spatial multi-modal data, demonstrating its potential for advancing ST research. The source code of TriCLFF is available online at https://github.com/HBZZ168/TriCLFF. Fenglan Pang, Guangfu Xue, Wenyi Yang, Yideng Cai, Jinhao Que, Haoxiu Sun, Shuaiyu Su, Xiyun Jin, Zuxiang Wang, Meng Luo 0001, Renjie Tan, Yusong Liu, Qinghua Jiang |
Briefings Bioinform. | 18 |
| 2023 | An interpretable single-cell RNA sequencing data clustering method based on latent Dirichlet allocationabstractSingle-cell RNA sequencing (scRNA-seq) detects whole transcriptome signals for large amounts of individual cells and is powerful for determining cell-to-cell differences and investigating the functional characteristics of various cell types. scRNA-seq datasets are usually sparse and highly noisy. Many steps in the scRNA-seq analysis workflow, including reasonable gene selection, cell clustering and annotation, as well as discovering the underlying biological mechanisms from such datasets, are difficult. In this study, we proposed an scRNA-seq analysis method based on the latent Dirichlet allocation (LDA) model. The LDA model estimates a series of latent variables, i.e. putative functions (PFs), from the input raw cell-gene data. Thus, we incorporated the 'cell-function-gene' three-layer framework into scRNA-seq analysis, as this framework is capable of discovering latent and complex gene expression patterns via a built-in model approach and obtaining biologically meaningful results through a data-driven functional interpretation process. We compared our method with four classic methods on seven benchmark scRNA-seq datasets. The LDA-based method performed best in the cell clustering test in terms of both accuracy and purity. By analysing three complex public datasets, we demonstrated that our method could distinguish cell types with multiple levels of functional specialization, and precisely reconstruct cell development trajectories. Moreover, the LDA-based method accurately identified the representative PFs and the representative genes for the cell types/cell stages, enabling data-driven cell cluster annotation and functional interpretation. According to the literature, most of the previously reported marker/functionally relevant genes were recognized. Wenyang Zhou, Qinghua Jiang, Liran Juan |
Briefings Bioinform. | 5 |
| 2023 | DeepCCI: a deep learning framework for identifying cell-cell interactions from single-cell RNA sequencing dataabstractMOTIVATION: Cell-cell interactions (CCIs) play critical roles in many biological processes such as cellular differentiation, tissue homeostasis, and immune response. With the rapid development of high throughput single-cell RNA sequencing (scRNA-seq) technologies, it is of high importance to identify CCIs from the ever-increasing scRNA-seq data. However, limited by the algorithmic constraints, current computational methods based on statistical strategies ignore some key latent information contained in scRNA-seq data with high sparsity and heterogeneity. RESULTS: Here, we developed a deep learning framework named DeepCCI to identify meaningful CCIs from scRNA-seq data. Applications of DeepCCI to a wide range of publicly available datasets from diverse technologies and platforms demonstrate its ability to predict significant CCIs accurately and effectively. Powered by the flexible and easy-to-use software, DeepCCI can provide the one-stop solution to discover meaningful intercellular interactions and build CCI networks from scRNA-seq data. AVAILABILITY AND IMPLEMENTATION: The source code of DeepCCI is available online at https://github.com/JiangBioLab/DeepCCI. Wenyi Yang, Meng Luo 0001, Yideng Cai, Guangfu Xue, Xiyun Jin, Rui Cheng 0003, Jinhao Que, Fenglan Pang, Huan Nie, Qinghua Jiang |
Bioinform. | 13 |
| 2022 | Identification of alternative splicing-derived cancer neoantigens for mRNA vaccine developmentabstractMessenger RNA (mRNA) vaccines have shown great potential for anti-tumor therapy due to the advantages in safety, efficacy and industrial production. However, it remains a challenge to identify suitable cancer neoantigens that can be targeted for mRNA vaccines. Abnormal alternative splicing occurs in a variety of tumors, which may result in the translation of abnormal transcripts into tumor-specific proteins. High-throughput technologies make it possible for systematic characterization of alternative splicing as a source of suitable target neoantigens for mRNA vaccine development. Here, we summarized difficulties and challenges for identifying alternative splicing-derived cancer neoantigens from RNA-seq data and proposed a conceptual framework for designing personalized mRNA vaccines based on alternative splicing-derived cancer neoantigens. In addition, several points were presented to spark further discussion toward improving the identification of alternative splicing-derived cancer neoantigens. Rui Cheng 0003, Meng Luo 0001, Huimin Cao, Xiyun Jin, Wenyang Zhou, Lixing Xiao, Qinghua Jiang |
Briefings Bioinform. | 9 |
| 2022 | CBLRR: a cauchy-based bounded constraint low-rank representation method to cluster single-cell RNA-seq dataabstractThe rapid development of single-cel+l RNA sequencing (scRNA-seq) technology provides unprecedented opportunities for exploring biological phenomena at the single-cell level. The discovery of cell types is one of the major applications for researchers to explore the heterogeneity of cells. Some computational methods have been proposed to solve the problem of scRNA-seq data clustering. However, the unavoidable technical noise and notorious dropouts also reduce the accuracy of clustering methods. Here, we propose the cauchy-based bounded constraint low-rank representation (CBLRR), which is a low-rank representation-based method by introducing cauchy loss function (CLF) and bounded nuclear norm regulation, aiming to alleviate the above issue. Specifically, as an effective loss function, the CLF is proven to enhance the robustness of the identification of cell types. Then, we adopt the bounded constraint to ensure the entry values of single-cell data within the restricted interval. Finally, the performance of CBLRR is evaluated on 15 scRNA-seq datasets, and compared with other state-of-the-art methods. The experimental results demonstrate that CBLRR performs accurately and robustly on clustering scRNA-seq data. Furthermore, CBLRR is an effective tool to cluster cells, and provides great potential for downstream analysis of single-cell data. The source code of CBLRR is available online at https://github.com/Ginnay/CBLRR. Wenyi Yang, Meng Luo 0001, Fenglan Pang, Yideng Cai, Anastasya A. Anashkina, Xi Su, Qinghua Jiang |
Briefings Bioinform. | 11 |
| 2022 | Impact of mutations in SARS-COV-2 spike on viral infectivity and antigenicityabstractSince the outbreak of SARS-CoV-2, the etiologic agent of the COVID-19 pandemic, the viral genome has acquired numerous mutations with the potential to alter the viral infectivity and antigenicity. Part of mutations in SARS-CoV-2 spike protein has conferred virus the ability to spread more quickly and escape from the immune response caused by the monoclonal neutralizing antibody or vaccination. Herein, we summarize the spatiotemporal distribution of mutations in spike protein, and present recent efforts and progress in investigating the impacts of those mutations on viral infectivity and antigenicity. As mutations continue to emerge in SARS-CoV-2, we strive to provide systematic evaluation of mutations in spike protein, which is vitally important for the subsequent improvement of vaccine and therapeutic neutralizing antibody strategies. Wenyang Zhou, Anastasya A. Anashkina, Qinghua Jiang |
Briefings Bioinform. | 5 |
| 2022 | An efficient multilayer RBF neural network and its application to regression problems
Qinghua Jiang, Lailai Zhu, Chang Shu 0002, Vinothkumar Sekar |
Neural Comput. Appl. | 1 |
| 2021 | Comprehensive analysis of partial methylation domains in colorectal cancer based on single-cell methylation profilesabstractEpigenetic aberrations have played a significant role in affecting the pathophysiological state of colorectal cancer, and global DNA hypomethylation mainly occurs in partial methylation domains (PMDs). However, the distribution of PMDs in individual cells and the heterogeneity between cells are still unclear. In this study, the DNA methylation profiles of colorectal cancer detected by WGBS and scBS-seq were used to depict PMDs in individual cells for the first time. We found that more than half of the entire genome is covered by PMDs. Three subclasses of PMDS have distinct characteristics, and Gain-PMDs cover a higher proportion of protein coding genes. Gain-PMDs have extensive epigenetic heterogeneity between different cells of the same tumor, and the DNA methylation in cells is affected by the tumor microenvironment. In addition, abnormally elevated promoter methylation in Gain-PMDs may further promote the growth, proliferation and metastasis of tumor cells through silent transcription. The PMDs detected in this study have the potential as epigenetic biomarkers and provide a new insight for colorectal cancer research based on single-cell methylation data. Wenyang Zhou, Meng Luo 0001, Rui Cheng 0003, Xiyun Jin, Qinghua Jiang |
Briefings Bioinform. | 10 |
| 2021 | Global characterization of B cell receptor repertoire in COVID-19 patients by single-cell V(D)J sequencingabstractThe world is facing a pandemic of Corona Virus Disease 2019 (COVID-19) caused by severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). Adaptive immune responses are essential for SARS-CoV-2 virus clearance. Although a large body of studies have been conducted to investigate the immune mechanism in COVID-19 patients, we still lack a comprehensive understanding of the BCR repertoire in patients. In this study, we used the single-cell V(D)J sequencing to characterize the BCR repertoire across convalescent COVID-19 patients. We observed that the BCR diversity was significantly reduced in disease compared with healthy controls. And BCRs tend to skew toward different V gene segments in COVID-19 and healthy controls. The CDR3 sequences of heavy chain in clonal BCRs in patients were more convergent than that in healthy controls. In addition, we discovered increased IgG and IgA isotypes in the disease, including IgG1, IgG3 and IgA1. In all clonal BCRs, IgG isotypes had the most frequent class switch recombination events and the highest somatic hypermutation rate, especially IgG3. Moreover, we found that an IgG3 cluster from different clonal groups had the same IGHV, IGHJ and CDR3 sequences (IGHV4-4-CARLANTNQFYDSSSYLNAMDVW-IGHJ6). Overall, our study provides a comprehensive characterization of the BCR repertoire in COVID-19 patients, which contributes to the understanding of the mechanism for the immune response to SARS-CoV-2 infection. Xiyun Jin, Wenyang Zhou, Meng Luo 0001, Kexin Ma 0003, Huimin Cao, Rui Cheng 0003, Lixing Xiao, Fenglan Pang, Huan Nie, Qinghua Jiang |
Briefings Bioinform. | 16 |
| 2021 | DLpTCR: an ensemble deep learning framework for predicting immunogenic peptide recognized by T cell receptorabstractAccurate prediction of immunogenic peptide recognized by T cell receptor (TCR) can greatly benefit vaccine development and cancer immunotherapy. However, identifying immunogenic peptides accurately is still a huge challenge. Most of the antigen peptides predicted in silico fail to elicit immune responses in vivo without considering TCR as a key factor. This inevitably causes costly and time-consuming experimental validation test for predicted antigens. Therefore, it is necessary to develop novel computational methods for precisely and effectively predicting immunogenic peptide recognized by TCR. Here, we described DLpTCR, a multimodal ensemble deep learning framework for predicting the likelihood of interaction between single/paired chain(s) of TCR and peptide presented by major histocompatibility complex molecules. To investigate the generality and robustness of the proposed model, COVID-19 data and IEDB data were constructed for independent evaluation. The DLpTCR model exhibited high predictive power with area under the curve up to 0.91 on COVID-19 data while predicting the interaction between peptide and single TCR chain. Additionally, the DLpTCR model achieved the overall accuracy of 81.03% on IEDB data while predicting the interaction between peptide and paired TCR chains. The results demonstrate that DLpTCR has the ability to learn general interaction rules and generalize to antigen peptide recognition by TCR. A user-friendly webserver is available at http://jianglab.org.cn/DLpTCR/. Additionally, a stand-alone software package that can be downloaded from https://github.com/jiangBiolab/DLpTCR. Meng Luo 0001, Weizhong Lin, Guangfu Xue, Xiyun Jin, Wenyang Zhou, Yideng Cai, Wenyi Yang, Huan Nie, Qinghua Jiang |
Briefings Bioinform. | 12 |
| 2021 | Identification of sub-Golgi protein localization by use of deep representation learning featuresabstractMOTIVATION: The Golgi apparatus has a key functional role in protein biosynthesis within the eukaryotic cell with malfunction resulting in various neurodegenerative diseases. For a better understanding of the Golgi apparatus, it is essential to identification of sub-Golgi protein localization. Although some machine learning methods have been used to identify sub-Golgi localization proteins by sequence representation fusion, more accurate sub-Golgi protein identification is still challenging by existing methodology. RESULTS: we developed a protein sub-Golgi localization identification protocol using deep representation learning features with 107 dimensions. By this protocol, we demonstrated that instead of multi-type protein sequence feature representation fusion as in previous state-of-the-art sub-Golgi-protein localization classifiers, it is sufficient to exploit only one type of feature representation for more accurately identification of sub-Golgi proteins. Compared with independent testing results for benchmark datasets, our protocol is able to perform generally, reliably and robustly for sub-Golgi protein localization prediction. AVAILABILITYAND IMPLEMENTATION: A use-friendly webserver is freely accessible at http://isGP-DRLF.aibiochem.net and the prediction code is accessible at https://github.com/zhibinlv/isGP-DRLF. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhibin Lv, Quan Zou 0001, Qinghua Jiang |
Bioinform. | 4 |
| 2020 | Morbigenous brain region and gene detection with a genetically evolved random neural network cluster approach in late mild cognitive impairmentabstractMOTIVATION: The multimodal data fusion analysis becomes another important field for brain disease detection and increasing researches concentrate on using neural network algorithms to solve a range of problems. However, most current neural network optimizing strategies focus on internal nodes or hidden layer numbers, while ignoring the advantages of external optimization. Additionally, in the multimodal data fusion analysis of brain science, the problems of small sample size and high-dimensional data are often encountered due to the difficulty of data collection and the specialization of brain science data, which may result in the lower generalization performance of neural network. RESULTS: We propose a genetically evolved random neural network cluster (GERNNC) model. Specifically, the fusion characteristics are first constructed to be taken as the input and the best type of neural network is selected as the base classifier to form the initial random neural network cluster. Second, the cluster is adaptively genetically evolved. Based on the GERNNC model, we further construct a multi-tasking framework for the classification of patients with brain disease and the extraction of significant characteristics. In a study of genetic data and functional magnetic resonance imaging data from the Alzheimer's Disease Neuroimaging Initiative, the framework exhibits great classification performance and strong morbigenous factor detection ability. This work demonstrates that how to effectively detect pathogenic components of the brain disease on the high-dimensional medical data and small samples. AVAILABILITY AND IMPLEMENTATION: The Matlab code is available at https://github.com/lizi1234560/GERNNC.git. Xia-an Bi, Qinghua Jiang |
Bioinform. | 5 |
| 2020 | ERDS-Exome: A Hybrid Approach for Copy Number Variant Detection from Whole-Exome Sequencing DataabstractCopy number variants (CNVs) play important roles in human disease and evolution. With the rapid development of next-generation sequencing technologies, many tools have been developed for inferring CNVs based on whole-exome sequencing (WES) data. However, as a result of the sparse distribution of exons in the genome, the limitations of the WES technique, and the nature of high-level signal noises in WES data, the efficacy of these variants remains less than desirable. Thus, there is need for the development of an effective tool to achieve a considerable power in WES CNVs discovery. In the present study, we describe a novel method, Estimation by Read Depth (RD) with Single-nucleotide variants from exome sequencing data (ERDS-exome). ERDS-exome employs a hybrid normalization approach to normalize WES data and to incorporate RD and single-nucleotide variation information together as a hybrid signal into a paired hidden Markov model to infer CNVs from WES data. Based on systematic evaluations of real data from the 1000 Genomes Project using other state-of-the-art tools, we observed that ERDS-exome demonstrates higher sensitivity and provides comparable or even better specificity than other tools. ERDS-exome is publicly available at: https://erds-exome.github.io. Renjie Tan, Jixuan Wang, Liran Juan, Qing Zhan, Shuilin Jin, Qinghua Jiang |
IEEE ACM Trans. Comput. Biol. Bioinform. | 10 |
| 2019 | A learning-based framework for miRNA-disease association identification using neural networksabstractMOTIVATION: A microRNA (miRNA) is a type of non-coding RNA, which plays important roles in many biological processes. Lots of studies have shown that miRNAs are implicated in human diseases, indicating that miRNAs might be potential biomarkers for various types of diseases. Therefore, it is important to reveal the relationships between miRNAs and diseases/phenotypes. RESULTS: We propose a novel learning-based framework, MDA-CNN, for miRNA-disease association identification. The model first captures interaction features between diseases and miRNAs based on a three-layer network including disease similarity network, miRNA similarity network and protein-protein interaction network. Then, it employs an auto-encoder to identify the essential feature combination for each pair of miRNA and disease automatically. Finally, taking the reduced feature representation as input, it uses a convolutional neural network to predict the final label. The evaluation results show that the proposed framework outperforms some state-of-the-art approaches in a large margin on both tasks of miRNA-disease association prediction and miRNA-phenotype association prediction. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/Issingjessica/MDA-CNN. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jiajie Peng, Weiwei Hui, Jianye Hao, Qinghua Jiang, Xuequn Shang 0001, Zhongyu Wei |
Bioinform. | 6 |
| 2019 | BIN1 rs744373 variant shows different association with Alzheimer's disease in Caucasian and Asian populationsabstractAbstract Background The association between BIN1 rs744373 variant and Alzheimer’s disease (AD) had been identified by genome-wide association studies (GWASs) as well as candidate gene studies in Caucasian populations. But in East Asian populations, both positive and negative results had been identified by association studies. Considering the smaller sample sizes of the studies in East Asian, we believe that the results did not have enough statistical power. Results We conducted a meta-analysis with 71,168 samples (22,395 AD cases and 48,773 controls, from 37 studies of 19 articles). Based on the additive model, we observed significant genetic heterogeneities in pooled populations as well as Caucasians and East Asians. We identified a significant association between rs744373 polymorphism with AD in pooled populations (P = 5 × 10− 07, odds ratio (OR) = 1.12, and 95% confidence interval (CI) 1.07–1.17) and in Caucasian populations (P = 3.38 × 10− 08, OR = 1.16, 95% CI 1.10–1.22). But in the East Asian populations, the association was not identified (P = 0.393, OR = 1.057, and 95% CI 0.95–1.15). Besides, the regression analysis suggested no significant publication bias. The results for sensitivity analysis as well as meta-analysis under the dominant model and recessive model remained consistent, which demonstrated the reliability of our finding. Conclusions The large-scale meta-analysis highlighted the significant association between rs744373 polymorphism and AD risk in Caucasian populations but not in the East Asian populations. Zhifa Han, Wenyang Zhou, Jian Zong, Yang Hu 0008, Shuilin Jin, Qinghua Jiang |
BMC Bioinform. | 10 |
| 2019 | ProbPFP: a multiple sequence alignment algorithm combining hidden Markov model optimized by particle swarm optimization with partition functionabstractBACKGROUND: During procedures for conducting multiple sequence alignment, that is so essential to use the substitution score of pairwise alignment. To compute adaptive scores for alignment, researchers usually use Hidden Markov Model or probabilistic consistency methods such as partition function. Recent studies show that optimizing the parameters for hidden Markov model, as well as integrating hidden Markov model with partition function can raise the accuracy of alignment. The combination of partition function and optimized HMM, which could further improve the alignment's accuracy, however, was ignored by these researches. RESULTS: A novel algorithm for MSA called ProbPFP is presented in this paper. It intergrate optimized HMM by particle swarm with partition function. The algorithm of PSO was applied to optimize HMM's parameters. After that, the posterior probability obtained by the HMM was combined with the one obtained by partition function, and thus to calculate an integrated substitution score for alignment. In order to evaluate the effectiveness of ProbPFP, we compared it with 13 outstanding or classic MSA methods. The results demonstrate that the alignments obtained by ProbPFP got the maximum mean TC scores and mean SP scores on these two benchmark datasets: SABmark and OXBench, and it got the second highest mean TC scores and mean SP scores on the benchmark dataset BAliBASE. ProbPFP is also compared with 4 other outstanding methods, by reconstructing the phylogenetic trees for six protein families extracted from the database TreeFam, based on the alignments obtained by these 5 methods. The result indicates that the reference trees are closer to the phylogenetic trees reconstructed from the alignments obtained by ProbPFP than the other methods. CONCLUSIONS: We propose a new multiple sequence alignment method combining optimized HMM and partition function in this paper. The performance validates this method could make a great improvement of the alignment's accuracy. Qing Zhan, Shuilin Jin, Renjie Tan, Qinghua Jiang |
BMC Bioinform. | 5 |
| 2018 | ProbPFP: A Multiple Sequence Alignment Algorithm Combining Partition Function and Hidden Markov Model with Particle Swarm Optimization
Qing Zhan, Shuilin Jin, Renjie Tan, Qinghua Jiang |
BIBM | 5 |
| 2018 | BIN1 rs744373 Variant Is Significantly Associated with Alzheimer's Disease in Caucasian but Not East Asian Populations
Zhifa Han, Wenyang Zhou, Jian Zong, Yang Hu 0008, Shuilin Jin, Qinghua Jiang |
ICIC (1) | 10 |
| 2018 | DincRNA: a comprehensive web-based bioinformatics toolkit for exploring disease associations and ncRNA functionabstractSummary: DincRNA aims to provide a comprehensive web-based bioinformatics toolkit to elucidate the entangled relationships among diseases and non-coding RNAs (ncRNAs) from the perspective of disease similarity. The quantitative way to illustrate relationships of pair-wise diseases always depends on their molecular mechanisms, and structures of the directed acyclic graph of Disease Ontology (DO). Corresponding methods for calculating similarity of pair-wise diseases involve Resnik's, Lin's, Wang's, PSB and SemFunSim methods. Recently, disease similarity was validated suitable for calculating functional similarities of ncRNAs and prioritizing ncRNA-disease pairs, and it has been widely applied for predicting the ncRNA function due to the limited biological knowledge from wet lab experiments of these RNAs. For this purpose, a large number of algorithms and priori knowledge need to be integrated. e.g. 'pair-wise best, pairs-average' (PBPA) and 'pair-wise all, pairs-maximum' (PAPM) methods for calculating functional similarities of ncRNAs, and random walk with restart (RWR) method for prioritizing ncRNA-disease pairs. To facilitate the exploration of disease associations and ncRNA function, DincRNA implemented all of the above eight algorithms based on DO and disease-related genes. Currently, it provides the function to query disease similarity scores, miRNA and lncRNA functional similarity scores, and the prioritization scores of lncRNA-disease and miRNA-disease pairs. Availability and implementation: http://bio-annotation.cn:18080/DincRNAClient/. Contact: [email protected] or [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Liang Cheng 0006, Yang Hu 0008, Jie Sun 0021, Meng Zhou 0003, Qinghua Jiang |
Bioinform. | 5 |
| 2018 | EWAS: epigenome-wide association study software 2.0abstractMotivation: With the development of biotechnology, DNA methylation data showed exponential growth. Epigenome-wide association study (EWAS) provide a systematic approach to uncovering epigenetic variants underlying common diseases/phenotypes. But the EWAS software has lagged behind compared with genome-wide association study (GWAS). To meet the requirements of users, we developed a convenient and useful software, EWAS2.0. Results: EWAS2.0 can analyze EWAS data and identify the association between epigenetic variations and disease/phenotype. On the basis of EWAS1.0, we have added more distinctive features. EWAS2.0 software was developed based on our 'population epigenetic framework' and can perform: (i) epigenome-wide single marker association study; (ii) epigenome-wide methylation haplotype (meplotype) association study and (iii) epigenome-wide association meta-analysis. Users can use EWAS2.0 to execute chi-square test, t-test, linear regression analysis, logistic regression analysis, identify the association between epi-alleles, identify the methylation disequilibrium (MD) blocks, calculate the MD coefficient, the frequency of meplotype and Pearson's correlation coefficients and carry out meta-analysis and so on. Finally, we expect EWAS2.0 to become a popular software and be widely used in epigenome-wide associated studies in the future. Availability and implementation: The EWAS software is freely available at http://www.ewas.org.cn or http://www.bioapp.org/ewas. Linna Zhao, Simeng Hu, Xiuling Song, Hongchao Lv, Qinghua Jiang, Guiyou Liu, Shuilin Jin, Mingzhi Liao, Rennan Feng, Fanwu Kong, Liangde Xu, Yongshuai Jiang |
Bioinform. | 10 |
| 2017 | GED: a manually curated comprehensive resource for epigenetic modification of gametogenesisabstractReproductive infertility affects seventh of couples, which is most attributed to the obstacle of gametogenesis. Characterizing the epigenetic modification factors involved in gametogenesis is fundamental to understand the molecular mechanisms and to develop treatments for human infertility. Although the genetic factors have been implicated in gametogenesis, no dedicated bioinformatics resource for gametogenesis is available. To elucidate the relationship of epigenetic modification and mammalian gametogenesis, we developed a new database, gametogenesis epigenetic modification database (GED), a manually curated database, which aims at providing a comprehensive resource of epigenetic modification of gametogenesis. The database integrates three kinds information of epigenetic modifications during gametogenesis (DNA methylation, histone modification and RNA regulation), and the gametogenesis has been detailed as 16 stages in seven mammal species (Homo sapiens, Mus musculus, Rattus norvegicus, Sus scrofa, Bos taurus, Capra hircus and Ovis aries). Besides, we have predicted the linear pathways of epigenetic modification which were composed of 211 genes/proteins and microRNAs that were involved in gametogenesis. GED is a user-friendly Web site, through which users can obtain the comprehensive epigenetic factor information and molecular pathways by visiting our database freely. GED is free available at http://gametsepi.nwsuaflmz.com. Weiyang Bai, Wen Yang 0009, Qinghua Jiang, Jinlian Hua, Mingzhi Liao |
Briefings Bioinform. | 6 |
| 2017 | DTWscore: differential expression and cell clustering analysis for time-series single-cell RNA-seq dataabstractBACKGROUND: The development of single-cell RNA sequencing has enabled profound discoveries in biology, ranging from the dissection of the composition of complex tissues to the identification of novel cell types and dynamics in some specialized cellular environments. However, the large-scale generation of single-cell RNA-seq (scRNA-seq) data collected at multiple time points remains a challenge to effective measurement gene expression patterns in transcriptome analysis. RESULTS: We present an algorithm based on the Dynamic Time Warping score (DTWscore) combined with time-series data, that enables the detection of gene expression changes across scRNA-seq samples and recovery of potential cell types from complex mixtures of multiple cell types. CONCLUSIONS: The DTWscore successfully classify cells of different types with the most highly variable genes from time-series scRNA-seq data. The study was confined to methods that are implemented and available within the R framework. Sample datasets and R packages are available at https://github.com/xiaoxiaoxier/DTWscore . Shuilin Jin, Guiyou Liu, Xiurui Zhang, Deliang Wu, Yang Hu 0008, Chiping Zhang, Qinghua Jiang, Yadong Wang 0001 |
BMC Bioinform. | 9 |
| 2016 | InfDisSim: A novel method for measuring disease similarity based on information flowabstractSimilar diseases are often caused by their similar molecular origins, such as disease-related protein-coding genes (PCGs). And nowadays, the function of PCGs has been widely studied on a gene function network, where each node represents a gene and each edge indicates an interaction between pair-wise genes. Therefore, functional interaction between disease-related PCGs should be exploited to measure disease similarity. Actually, functional interaction of pair-wise PCGs has been introduced to calculate disease similarity recently. However, existing method ignores that genes could also be associated based on intermediate nodes in the gene functional network. Here, in this article, we proposed a novel method, InfDisSim, to infer disease similarity. InfDisSim models the information flow to the network based on random walk with damping, in which the entire network could be fully utilized. The performance of InfDisSim was evaluated by a benchmark set of similar disease pairs. The area under the receiver operating characteristic curve (AUC) was calculated to evaluate the performance. As a result, InfDisSim achieves a very high AUC (0.9786), which shows it performs well. Furthermore, based on the disease similarity computed by the infDisSim, we re-validated that similar diseases tend to have common therapeutic drugs (Pearson correlation γ2=0.1315, p=2.2e-16). Finally, InfDisSim disease similarity was exploited to construct a lncRNA similarity network (LSN), which was further applied to predict potential associations between diseases and lncRNAs. High AUC (0.9893) based on leave-one-out cross validation shows the LSN is very suitable for identifying novel disease-related lncRNAs. Yang Hu 0008, Meng Zhou 0003, Hong Ju, Qinghua Jiang, Liang Cheng 0006 |
BIBM | 5 |
| 2016 | ERDS-pe: A paired hidden Markov model for copy number variant detection from whole-exome sequencing dataabstractDetecting copy number variants (CNVs) is an essential part in variant calling process. Here, we describe a novel method ERDS-pe to detect CNVs from whole-exome sequencing (WES) data. ERDS-pe first employs principal component analysis to normalize WES data. Then, ERDS-pe incorporates read depth signal and single-nucleotide variation information together as a hybrid signal into a paired hidden Markov model to infer CNVs from WES data. Experimental results on real human WES data show that ERDS-pe demonstrates higher sensitivity and provides comparable or even better specificity than other tools. ERDS-pe is publicly available at: https://github.com/microtan0902/erds-pe. Renjie Tan, Jixuan Wang, Guoqiang Wan, Zhijie Han 0002, Wenyang Zhou, Shuilin Jin, Qinghua Jiang, Yadong Wang 0001 |
BIBM | 10 |
| 2016 | PrefaceabstractWelcome to the 2016 IEEE International Conference on Bioinformatics and Biomedicine (IEEE BIBM 2016) being held in the Shenzhen, China from December 15–18, 2016. On behalf of the IEEE BIBM 2016 Organizing Team, we would like to thank you for your participation and hope you enjoy the conference. Yadong Wang 0001, Kevin Burrage, Shinichi Morishita, Tianhai Tian, Qinghua Jiang, Jiangning Song, Guohua Wang 0001, Xiaohua Hu 0001 |
BIBM | 6 |
| 2015 | misFinder: identify mis-assemblies in an unbiased manner using reference and paired-end readsabstractBACKGROUND: Because of the short read length of high throughput sequencing data, assembly errors are introduced in genome assembly, which may have adverse impact to the downstream data analysis. Several tools have been developed to eliminate these errors by either 1) comparing the assembled sequences with some similar reference genome, or 2) analyzing paired-end reads aligned to the assembled sequences and determining inconsistent features alone mis-assembled sequences. However, the former approach cannot distinguish real structural variations between the target genome and the reference genome while the latter approach could have many false positive detections (correctly assembled sequence being considered as mis-assembled sequence). RESULTS: We present misFinder, a tool that aims to identify the assembly errors with high accuracy in an unbiased way and correct these errors at their mis-assembled positions to improve the assembly accuracy for downstream analysis. It combines the information of reference (or close related reference) genome and aligned paired-end reads to the assembled sequence. Assembly errors and correct assemblies corresponding to structural variations can be detected by comparing the genome reference and assembled sequence. Different types of assembly errors can then be distinguished from the mis-assembled sequence by analyzing the aligned paired-end reads using multiple features derived from coverage and consistence of insert distance to obtain high confident error calls. CONCLUSIONS: We tested the performance of misFinder on both simulated and real paired-end reads data, and misFinder gave accurate error calls with only very few miscalls. And, we further compared misFinder with QUAST and REAPR. misFinder outperformed QUAST and REAPR by 1) identified more true positive mis-assemblies with very few false positives and false negatives, and 2) distinguished the correct assemblies corresponding to structural variations from mis-assembled sequence. misFinder can be freely downloaded from https://github.com/hitbio/misFinder. Henry C. M. Leung, Francis Y. L. Chin, Siu-Ming Yiu, Guangri Quan, Qinghua Jiang, Bo Liu 0023, Yucui Dong, Yadong Wang 0001 |
BMC Bioinform. | 9 |
| 2010 | Predicting human microRNA-disease associations based on support vector machineabstractThe identification of disease-related microRNAs is vital for understanding the pathogenesis of disease at the molecular level and may lead to the design of specific molecular tools for diagnosis, treatment and prevention. Experimental identification of disease-related microRNAs poses difficulties. Computational prediction of microRNA-disease associations is one of the complementary means. However, one major issue in microRNA studies is the lack of bioinformatics programs to accurately predict microRNA-disease associations. Herein, we present a machine learning-based approach for distinguishing positive microRNA-disease associations from negative microRNA-disease associations. A set of features was extracted for each positive and negative microRNA-disease association, and a support vector machine (SVM) classifier was trained, which achieved the area under the ROC curve of up to 0.8884 in 10-fold cross-validation procedure, indicating that the SVM-based approach described here can be used to predict potential microRNA-disease associations and formulate testable hypotheses to guide future biological experiments. Qinghua Jiang, Guohua Wang 0001, Yadong Wang 0001 |
BIBM | 1 |