EDBT 2026 Demo / reviewers in the wild / expert
Qin Ma 0003
dblp:78/5428-3
· DBLP profile ↗
42ranked-venue papers
1as first author
20since 2021 · last 2026
0000-0002-3264-8392ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 38 · 1 first-author · 17 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Prevalence aware feature selection improves biomarker identification in microbiome studies
Kris Sankaran, Thomas A. Mace, Phil A Hart, Qin Ma 0003, Xu-Wen Wang, Shanlin Ke |
Bioinform. | 6 |
| 2025 | scBSP: a fast and accurate tool for identifying spatially variable features from high-resolution spatial omics dataabstractMOTIVATION: Emerging spatial omics technologies empower comprehensive exploration of biological systems from multi-omics perspectives in their native tissue location in 2D and 3D space. However, the limited sequencing depth, increasing spatial resolution, and growing spatial spots in spatial omics technologies present significant computational challenges in identifying biologically meaningful molecules with variable spatial distributions across various omics modalities. RESULTS: We introduce scBSP, an open-source, versatile, and user-friendly package for identifying spatially variable features in large-scale spatial omics data. scBSP demonstrates significantly enhanced computational efficiency, processing high-resolution spatial omics data within seconds, and exhibits robust cross-platform performance by consistently identifying spatially variable features with high reproducibility across various sequencing platforms. AVAILABILITY AND IMPLEMENTATION: scBSP is available for download from R CRAN at https://cran.r-project.org/web/packages/scBSP/index.html and PyPI at https://pypi.org/project/scbsp/. Jinpu Li, Mauminah Raina, Ricardo Melo Ferreira, Michael Eadon, Qin Ma 0003, Juexin Wang, Dong Xu 0002 |
Bioinform. | 9 |
| 2025 | Relation equivariant graph neural networks to explore the mosaic-like tissue architecture of kidney diseases on spatially resolved transcriptomicsabstractMOTIVATION: Chronic kidney disease (CKD) and acute kidney injury (AKI) are prominent public health concerns affecting more than 15% of the global population. The ongoing development of spatially resolved transcriptomics (SRT) technologies presents a promising approach for discovering the spatial distribution patterns of gene expression within diseased tissues. However, existing computational tools are predominantly calibrated and designed on the ribbon-like structure of the brain cortex, presenting considerable computational obstacles in discerning highly heterogeneous mosaic-like tissue architectures in the kidney. Consequently, timely and cost-effective acquisition of annotation and interpretation in the kidney remains a challenge in exploring the cellular and morphological changes within renal tubules and their interstitial niches. RESULTS: We present an empowered graph deep learning framework, REGNN (Relation Equivariant Graph Neural Networks), designed for SRT data analyses on heterogeneous tissue structures. To increase expressive power in the SRT lattice using graph modeling, REGNN integrates equivariance to handle n-dimensional symmetries of the spatial area, while additionally leveraging Positional Encoding to strengthen relative spatial relations of the nodes uniformly distributed in the lattice. Given the limited availability of well-labeled spatial data, this framework implements both graph autoencoder and graph self-supervised learning strategies. On heterogeneous samples from different kidney conditions, REGNN outperforms existing computational tools in identifying tissue architectures within the 10× Visium platform. This framework offers a powerful graph deep learning tool for investigating tissues within highly heterogeneous expression patterns and paves the way to pinpoint underlying pathological mechanisms that contribute to the progression of complex diseases. AVAILABILITY AND IMPLEMENTATION: REGNN is publicly available at https://github.com/Mraina99/REGNN. Mauminah Raina, Ricardo Melo Ferreira, Treyden Stansfield, Chandrima Modak, Ying-Hua Cheng, Hari Naga Sai Kiran Suryadevara, Dong Xu 0002, Michael Eadon, Qin Ma 0003, Juexin Wang |
Bioinform. | 10 |
| 2025 | Cross-loop contrast of heterogeneous graphs for interdisciplinary journal recommendation
Ying Li 0004, Hongmin Sun, Wei Du 0002, Qin Ma 0003 |
Expert Syst. Appl. | 7 |
| 2024 | Enhancer-driven gene regulatory networks inference from single-cell RNA-seq and ATAC-seq dataabstractDeciphering the intricate relationships between transcription factors (TFs), enhancers, and genes through the inference of enhancer-driven gene regulatory networks (eGRNs) is crucial in understanding gene regulatory programs in a complex biological system. This study introduces STREAM, a novel method that leverages a Steiner forest problem model, a hybrid biclustering pipeline, and submodular optimization to infer eGRNs from jointly profiled single-cell transcriptome and chromatin accessibility data. Compared to existing methods, STREAM demonstrates enhanced performance in terms of TF recovery, TF-enhancer linkage prediction, and enhancer-gene relation discovery. Application of STREAM to an Alzheimer's disease dataset and a diffuse small lymphocytic lymphoma dataset reveals its ability to identify TF-enhancer-gene relations associated with pseudotime, as well as key TF-enhancer-gene relations and TF cooperation underlying tumor cells. Yang Li 0089, Anjun Ma, Yizhong Wang, Cankun Wang, Hongjun Fu, Bingqiang Liu, Qin Ma 0003 |
Briefings Bioinform. | 8 |
| 2024 | CEMIG: prediction of the cis-regulatory motif using the de Bruijn graph from ATAC-seqabstractSequence motif discovery algorithms enhance the identification of novel deoxyribonucleic acid sequences with pivotal biological significance, especially transcription factor (TF)-binding motifs. The advent of assay for transposase-accessible chromatin using sequencing (ATAC-seq) has broadened the toolkit for motif characterization. Nonetheless, prevailing computational approaches have focused on delineating TF-binding footprints, with motif discovery receiving less attention. Herein, we present Cis rEgulatory Motif Influence using de Bruijn Graph (CEMIG), an algorithm leveraging de Bruijn and Hamming distance graph paradigms to predict and map motif sites. Assessment on 129 ATAC-seq datasets from the Cistrome Data Browser demonstrates CEMIG's exceptional performance, surpassing three established methodologies on four evaluative metrics. CEMIG accurately identifies both cell-type-specific and common TF motifs within GM12878 and K562 cell lines, demonstrating its comparative genomic capabilities in the identification of evolutionary conservation and cell-type specificity. In-depth transcriptional and functional genomic studies have validated the functional relevance of CEMIG-identified motifs across various cell types. CEMIG is available at https://github.com/OSU-BMBL/CEMIG, developed in C++ to ensure cross-platform compatibility with Linux, macOS and Windows operating systems. Yizhong Wang, Yang Li 0089, Cankun Wang, Chan-Wang Jerry Lio, Qin Ma 0003, Bingqiang Liu |
Briefings Bioinform. | 5 |
| 2023 | DeepProSite: structure-aware protein binding site prediction using ESMFold and pretrained language modelabstractMOTIVATION: Identifying the functional sites of a protein, such as the binding sites of proteins, peptides, or other biological components, is crucial for understanding related biological processes and drug design. However, existing sequence-based methods have limited predictive accuracy, as they only consider sequence-adjacent contextual features and lack structural information. RESULTS: In this study, DeepProSite is presented as a new framework for identifying protein binding site that utilizes protein structure and sequence information. DeepProSite first generates protein structures from ESMFold and sequence representations from pretrained language models. It then uses Graph Transformer and formulates binding site predictions as graph node classifications. In predicting protein-protein/peptide binding sites, DeepProSite outperforms state-of-the-art sequence- and structure-based methods on most metrics. Moreover, DeepProSite maintains its performance when predicting unbound structures, in contrast to competing structure-based prediction methods. DeepProSite is also extended to the prediction of binding sites for nucleic acids and other ligands, verifying its generalization capability. Finally, an online server for predicting multiple types of residue is established as the implementation of the proposed DeepProSite. AVAILABILITY AND IMPLEMENTATION: The datasets and source codes can be accessed at https://github.com/WeiLab-Biology/DeepProSite. The proposed DeepProSite can be accessed at https://inner.wei-group.net/DeepProSite/. Yitian Fang, Leyi Wei, Qin Ma 0003, Zhixiang Ren, Qianmu Yuan |
Bioinform. | 4 |
| 2023 | SURE: Screening unlabeled samples for reliable negative samples based on reinforcement learning
Ying Li 0004, Wensi Fang, Qin Ma 0003, Rui Wang-Sattler, Wei Du 0002, Qiong Yu |
Inf. Sci. | 4 |
| 2023 | LRT: Integrative analysis of scRNA-seq and scTCR-seq data to investigate clonal differentiation heterogeneityabstractSingle-cell RNA sequencing (scRNA-seq) data has been widely used for cell trajectory inference, with the assumption that cells with similar expression profiles share the same differentiation state. However, the inferred trajectory may not reveal clonal differentiation heterogeneity among T cell clones. Single-cell T cell receptor sequencing (scTCR-seq) data provides invaluable insights into the clonal relationship among cells, yet it lacks functional characteristics. Therefore, scRNA-seq and scTCR-seq data complement each other in improving trajectory inference, where a reliable computational tool is still missing. We developed LRT, a computational framework for the integrative analysis of scTCR-seq and scRNA-seq data to explore clonal differentiation trajectory heterogeneity. Specifically, LRT uses the transcriptomics information from scRNA-seq data to construct overall cell trajectories and then utilizes both the TCR sequence information and phenotype information to identify clonotype clusters with distinct differentiation biasedness. LRT provides a comprehensive analysis workflow, including preprocessing, cell trajectory inference, clonotype clustering, trajectory biasedness evaluation, and clonotype cluster characterization. We illustrated its utility using scRNA-seq and scTCR-seq data of CD8+ T cells and CD4+ T cells with acute lymphocytic choriomeningitis virus infection. These analyses identified several clonotype clusters with distinct skewed distribution along the differentiation path, which cannot be revealed solely based on scRNA-seq data. Clones from different clonotype clusters exhibited diverse expansion capability, V-J gene usage pattern and CDR3 motifs. The LRT framework was implemented as an R package 'LRT', and it is now publicly accessible at https://github.com/JuanXie19/LRT. In addition, it provides two Shiny apps 'shinyClone' and 'shinyClust' that allow users to interactively explore distributions of clonotypes, conduct repertoire analysis, implement clustering of clonotypes, trajectory biasedness evaluation, and clonotype cluster characterization. Juan Xie, Hyeongseon Jeon, Gang Xin, Qin Ma 0003, Dongjun Chung |
PLoS Comput. Biol. | 4 |
| 2022 | Machine learning development environment for single-cell sequencing data analysesabstractMachine learning (ML) is transforming single-cell sequencing data analysis; however, the barriers of technology complexity and biology knowledge remain challenging for the involvement of the ML community in single-cell data analysis. Here we present an ML development environment for single-cell sequencing data analyses, together with a diverse set of realistic and accessible ML-Ready benchmark datasets. A cloud-based platform is built to dynamically scale workflows for collecting, processing, and managing various single-cell sequencing data to make them ML-ready. In addition, benchmarks for each problem formulation and a code-level and web-interface IDE for single-cell analysis method development are provided. These efforts provide an automated end-to-end single-cell analysis ML pipeline that simplifies and standardizes the process of single-cell data formatting, loading, model development, and model evaluation. Yuexu Jiang, Cankun Wang, Clement Essien, Juexin Wang, Anjun Ma, Qin Ma 0003, Dong Xu 0002 |
BIBM | 7 |
| 2022 | Provable Second-Order Riemannian Gauss-Newton Method for Low-Rank Tensor Estimation ‖abstractIn this paper, we consider the estimation of a low Tucker rank tensor from a number of noisy linear measurements. We propose a Riemannian Gauss-Newton (RGN) method with fast implementations for low Tucker rank tensor estimation. Different from the generic (super)linear convergence guarantee of RGN in the literature, we prove the first quadratic convergence guarantee of RGN for low-rank tensor estimation under some mild conditions. A deterministic estimation error lower bound, which matches the upper bound, is provided that demonstrates the statistical optimality of RGN. The merit of RGN is illustrated through applications of tensor regression and tensor SVD. Yuetian Luo, Qin Ma 0003, Chi Zhang 0021, Anru Zhang |
ICASSP | 2 |
| 2022 | Assessing deep learning methods in cis-regulatory motif finding based on genomic sequencing dataabstractIdentifying cis-regulatory motifs from genomic sequencing data (e.g. ChIP-seq and CLIP-seq) is crucial in identifying transcription factor (TF) binding sites and inferring gene regulatory mechanisms for any organism. Since 2015, deep learning (DL) methods have been widely applied to identify TF binding sites and predict motif patterns, with the strengths of offering a scalable, flexible and unified computational approach for highly accurate predictions. As far as we know, 20 DL methods have been developed. However, without a clear and systematic assessment, users will struggle to choose the most appropriate tool for their specific studies. In this manuscript, we evaluated 20 DL methods for cis-regulatory motif prediction using 690 ENCODE ChIP-seq, 126 cancer ChIP-seq and 55 RNA CLIP-seq data. Four metrics were investigated, including the accuracy of motif finding, the performance of DNA/RNA sequence classification, algorithm scalability and tool usability. The assessment results demonstrated the high complementarity of the existing DL methods. It was determined that the most suitable model should primarily depend on the data size and type and the method's outputs. Shuangquan Zhang, Anjun Ma, Dong Xu 0002, Qin Ma 0003, Yan Wang 0028 |
Briefings Bioinform. | 5 |
| 2022 | scGNN 2.0: a graph neural network tool for imputation and clustering of single-cell RNA-Seq dataabstractMOTIVATION: Gene expression imputation has been an essential step of the single-cell RNA-Seq data analysis workflow. Among several deep-learning methods, the debut of scGNN gained substantial recognition in 2021 for its superior performance and the ability to produce a cell-cell graph. However, the implementation of scGNN was relatively time-consuming and its performance could still be optimized. RESULTS: The implementation of scGNN 2.0 is significantly faster than scGNN thanks to a simplified close-loop architecture. For all eight datasets, cell clustering performance was increased by 85.02% on average in terms of adjusted rand index, and the imputation Median L1 Error was reduced by 67.94% on average. With the built-in visualizations, users can quickly assess the imputation and cell clustering results, compare against benchmarks and interpret the cell-cell interaction. The expanded input and output formats also pave the way for custom workflows that integrate scGNN 2.0 with other scRNA-Seq toolkits on both Python and R platforms. AVAILABILITY AND IMPLEMENTATION: scGNN 2.0 is implemented in Python (as of version 3.8) with the source code available at https://github.com/OSU-BMBL/scGNN2.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Haocheng Gu, Anjun Ma, Yang Li 0089, Juexin Wang, Dong Xu 0002, Qin Ma 0003 |
Bioinform. | 7 |
| 2021 | A fine-scale map of genome-wide recombination in divergent Escherichia coli populationabstractRecombination is one of the most important molecular mechanisms of prokaryotic genome evolution, but its exact roles are still in debate. Here we try to infer genome-wide recombination within a species, utilizing a dataset of 149 complete genomes of Escherichia coli from diverse animal hosts and geographic origins, including 45 in-house sequenced with the single-molecular real-time platform. Two major clades identified based on physiological, clinical and ecological characteristics form distinct genetic lineages based on scarcity of interclade gene exchanges. By defining gene-based syntenies for genomic segments within and between the two clades, we build a fine-scale recombination map for this representative global E. coli population. The map suggests extensive within-clade recombination that often breaks physical linkages among individual genes but seldom interrupts the structure of genome organizational frameworks as well as primary metabolic portfolios supported by the framework integrity, possibly due to strong natural selection for both physiological compatibility and ecological fitness. In contrast, the between-clade recombination declines drastically when phylogenetic distance increases to the extent where a 10-fold reduction can be observed, establishing a firm genetic barrier between clades. Our empirical data suggest a critical role for such recombination events in the early stage of speciation where recombination rate is associated with phylogenetic distance in addition to sequence and gene variations. The extensive intraclade recombination binds sister strains into a quasisexual group and optimizes genes or alleles to streamline physiological activities, whereas the sharply declined interclade recombination split the population into clades adaptive to divergent ecological niches. Yu Kang 0004, Lina Yuan, Yanan Chu, Xinmiao Jia, Qin Ma 0003, Jian Wang 0065, Jing-Fa Xiao, Songnian Hu, Zhancheng Gao, Jun Yu 0004 |
Briefings Bioinform. | 8 |
| 2021 | Deep forest ensemble learning for classification of alignments of non-coding RNA sequences based on multi-view structure representationsabstractNon-coding RNAs (ncRNAs) play crucial roles in multiple biological processes. However, only a few ncRNAs' functions have been well studied. Given the significance of ncRNAs classification for understanding ncRNAs' functions, more and more computational methods have been introduced to improve the classification automatically and accurately. In this paper, based on a convolutional neural network and a deep forest algorithm, multi-grained cascade forest (GcForest), we propose a novel deep fusion learning framework, GcForest fusion method (GCFM), to classify alignments of ncRNA sequences for accurate clustering of ncRNAs. GCFM integrates a multi-view structure feature representation including sequence-structure alignment encoding, structure image representation and shape alignment encoding of structural subunits, enabling us to capture the potential specificity between ncRNAs. For the classification of pairwise alignment of two ncRNA sequences, the F-value of GCFM improves 6% than an existing alignment-based method. Furthermore, the clustering of ncRNA families is carried out based on the classification matrix generated from GCFM. Results suggest better performance (with 20% accuracy improved) than existing ncRNA clustering methods (RNAclust, Ensembleclust and CNNclust). Additionally, we apply GCFM to construct a phylogenetic tree of ncRNA and predict the probability of interactions between RNAs. Most ncRNAs are located correctly in the phylogenetic tree, and the prediction accuracy of RNA interaction is 90.63%. A web server (http://bmbl.sdstate.edu/gcfm/) is developed to maximize its availability, and the source code and related data are available at the same URL. Ying Li 0004, Qi Zhang 0061, Zhaoqian Liu, Cankun Wang, Qin Ma 0003, Wei Du 0002 |
Briefings Bioinform. | 6 |
| 2021 | The functional determinants in the organization of bacterial genomesabstractBacterial genomes are now recognized as interacting intimately with cellular processes. Uncovering organizational mechanisms of bacterial genomes has been a primary focus of researchers to reveal the potential cellular activities. The advances in both experimental techniques and computational models provide a tremendous opportunity for understanding these mechanisms, and various studies have been proposed to explore the organization rules of bacterial genomes associated with functions recently. This review focuses mainly on the principles that shape the organization of bacterial genomes, both locally and globally. We first illustrate local structures as operons/transcription units for facilitating co-transcription and horizontal transfer of genes. We then clarify the constraints that globally shape bacterial genomes, such as metabolism, transcription and replication. Finally, we highlight challenges and opportunities to advance bacterial genomic studies and provide application perspectives of genome organization, including pathway hole assignment and genome assembly and understanding disease mechanisms. Zhaoqian Liu, Jingtong Feng, Bin Yu 0007, Qin Ma 0003, Bingqiang Liu |
Briefings Bioinform. | 4 |
| 2021 | Network analyses in microbiome based on high-throughput multi-omics dataabstractTogether with various hosts and environments, ubiquitous microbes interact closely with each other forming an intertwined system or community. Of interest, shifts of the relationships between microbes and their hosts or environments are associated with critical diseases and ecological changes. While advances in high-throughput Omics technologies offer a great opportunity for understanding the structures and functions of microbiome, it is still challenging to analyse and interpret the omics data. Specifically, the heterogeneity and diversity of microbial communities, compounded with the large size of the datasets, impose a tremendous challenge to mechanistically elucidate the complex communities. Fortunately, network analyses provide an efficient way to tackle this problem, and several network approaches have been proposed to improve this understanding recently. Here, we systemically illustrate these network theories that have been used in biological and biomedical research. Then, we review existing network modelling methods of microbial studies at multiple layers from metagenomics to metabolomics and further to multi-omics. Lastly, we discuss the limitations of present studies and provide a perspective for further directions in support of the understanding of microbial communities. Zhaoqian Liu, Anjun Ma, Ewy A. Mathé, Marlena Merling, Qin Ma 0003, Bingqiang Liu |
Briefings Bioinform. | 5 |
| 2021 | A novel computational framework for genome-scale alternative transcription units predictionabstractAlternative transcription units (ATUs) are dynamically encoded under different conditions and display overlapping patterns (sharing one or more genes) under a specific condition in bacterial genomes. Genome-scale identification of ATUs is essential for studying the emergence of human diseases caused by bacterial organisms. However, it is unrealistic to identify all ATUs using experimental techniques because of the complexity and dynamic nature of ATUs. Here, we present the first-of-its-kind computational framework, named SeqATU, for genome-scale ATU prediction based on next-generation RNA-Seq data. The framework utilizes a convex quadratic programming model to seek an optimum expression combination of all of the to-be-identified ATUs. The predicted ATUs in Escherichia coli reached a precision of 0.77/0.74 and a recall of 0.75/0.76 in the two RNA-Sequencing datasets compared with the benchmarked ATUs from third-generation RNA-Seq data. In addition, the proportion of 5'- or 3'-end genes of the predicted ATUs, having documented transcription factor binding sites and transcription termination sites, was three times greater than that of no 5'- or 3'-end genes. We further evaluated the predicted ATUs by Gene Ontology and Kyoto Encyclopedia of Genes and Genomes functional enrichment analyses. The results suggested that gene pairs frequently encoded in the same ATUs are more functionally related than those that can belong to two distinct ATUs. Overall, these results demonstrated the high reliability of predicted ATUs. We expect that the new insights derived by SeqATU will not only improve the understanding of the transcription mechanism of bacteria but also guide the reconstruction of a genome-scale transcriptional regulatory network. Zhaoqian Liu, Wen-Chi Chou, Laurence Ettwiller, Qin Ma 0003, Bingqiang Liu |
Briefings Bioinform. | 6 |
| 2021 | IRIS-FGM: an integrative single-cell RNA-Seq interpretation system for functional gene module analysisabstractSUMMARY: Single-cell RNA-Seq (scRNA-Seq) data is useful in discovering cell heterogeneity and signature genes in specific cell populations in cancer and other complex diseases. Specifically, the investigation of condition-specific functional gene modules (FGM) can help to understand interactive gene networks and complex biological processes in different cell clusters. QUBIC2 is recognized as one of the most efficient and effective biclustering tools for condition-specific FGM identification from scRNA-Seq data. However, its limited availability to a C implementation restricted its application to only a few downstream analysis functionalities. We developed an R package named IRIS-FGM (Integrative scRNA-Seq Interpretation System for Functional Gene Module analysis) to support the investigation of FGMs and cell clustering using scRNA-Seq data. Empowered by QUBIC2, IRIS-FGM can effectively identify condition-specific FGMs, predict cell types/clusters, uncover differentially expressed genes and perform pathway enrichment analysis. It is noteworthy that IRIS-FGM can also take Seurat objects as input, facilitating easy integration with the existing analysis pipeline. AVAILABILITY AND IMPLEMENTATION: IRIS-FGM is implemented in the R environment (as of version 3.6) with the source code freely available at https://github.com/BMEngineeR/IRISFGM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yuzhou Chang, Carter Allen, Changlin Wan, Dongjun Chung, Chi Zhang 0021, Zihai Li, Qin Ma 0003 |
Bioinform. | 7 |
| 2021 | DeepSec: a deep learning framework for secreted protein discovery in human body fluidsabstractMOTIVATION: Human proteins that are secreted into different body fluids from various cells and tissues can be promising disease indicators. Modern proteomics research empowered by both qualitative and quantitative profiling techniques has made great progress in protein discovery in various human fluids. However, due to the large number of proteins and diverse modifications present in the fluids, as well as the existing technical limits of major proteomics platforms (e.g. mass spectrometry), large discrepancies are often generated from different experimental studies. As a result, a comprehensive proteomics landscape across major human fluids are not well determined. RESULTS: To bridge this gap, we have developed a deep learning framework, named DeepSec, to identify secreted proteins in 12 types of human body fluids. DeepSec adopts an end-to-end sequence-based approach, where a Convolutional Neural Network is built to learn the abstract sequence features followed by a Bidirectional Gated Recurrent Unit with fully connected layer for protein classification. DeepSec has demonstrated promising performances with average area under the ROC curves of 0.85-0.94 on testing datasets in each type of fluids, which outperforms existing state-of-the-art methods available mostly on blood proteins. As an illustration of how to apply DeepSec in biomarker discovery research, we conducted a case study on kidney cancer by using genomics data from the cancer genome atlas and have identified 104 possible marker proteins. AVAILABILITY: DeepSec is available at https://bmbl.bmi.osumc.edu/deepsec/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dan Shao, Lan Huang 0002, Yan Wang 0028, Xueteng Cui, Yao Wang 0010, Qin Ma 0003, Juan Cui |
Bioinform. | 7 |
| 2020 | Clustering and classification methods for single-cell RNA-sequencing dataabstractAppropriate ways to measure the similarity between single-cell RNA-sequencing (scRNA-seq) data are ubiquitous in bioinformatics, but using single clustering or classification methods to process scRNA-seq data is generally difficult. This has led to the emergence of integrated methods and tools that aim to automatically process specific problems associated with scRNA-seq data. These approaches have attracted a lot of interest in bioinformatics and related fields. In this paper, we systematically review the integrated methods and tools, highlighting the pros and cons of each approach. We not only pay particular attention to clustering and classification methods but also discuss methods that have emerged recently as powerful alternatives, including nonlinear and linear methods and descending dimension methods. Finally, we focus on clustering and classification methods for scRNA-seq data, in particular, integrated methods, and provide a comprehensive description of scRNA-seq data and download URLs. Ren Qi, Anjun Ma, Qin Ma 0003, Quan Zou 0001 |
Briefings Bioinform. | 3 |
| 2020 | QUBIC2: a novel and robust biclustering algorithm for analyses and interpretation of large-scale RNA-Seq dataabstractMOTIVATION: The biclustering of large-scale gene expression data holds promising potential for detecting condition-specific functional gene modules (i.e. biclusters). However, existing methods do not adequately address a comprehensive detection of all significant bicluster structures and have limited power when applied to expression data generated by RNA-Sequencing (RNA-Seq), especially single-cell RNA-Seq (scRNA-Seq) data, where massive zero and low expression values are observed. RESULTS: We present a new biclustering algorithm, QUalitative BIClustering algorithm Version 2 (QUBIC2), which is empowered by: (i) a novel left-truncated mixture of Gaussian model for an accurate assessment of multimodality in zero-enriched expression data, (ii) a fast and efficient dropouts-saving expansion strategy for functional gene modules optimization using information divergency and (iii) a rigorous statistical test for the significance of all the identified biclusters in any organism, including those without substantial functional annotations. QUBIC2 demonstrated considerably improved performance in detecting biclusters compared to other five widely used algorithms on various benchmark datasets from E.coli, Human and simulated data. QUBIC2 also showcased robust and superior performance on gene expression data generated by microarray, bulk RNA-Seq and scRNA-Seq. AVAILABILITY AND IMPLEMENTATION: The source code of QUBIC2 is freely available at https://github.com/OSU-BMBL/QUBIC2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Juan Xie, Anjun Ma, Bingqiang Liu, Sha Cao, Cankun Wang, Chi Zhang 0021, Qin Ma 0003 |
Bioinform. | 9 |
| 2020 | SubMito-XGBoost: predicting protein submitochondrial localization by fusing multiple feature information and eXtreme gradient boostingabstractMOTIVATION: Mitochondria are an essential organelle in most eukaryotes. They not only play an important role in energy metabolism but also take part in many critical cytopathological processes. Abnormal mitochondria can trigger a series of human diseases, such as Parkinson's disease, multifactor disorder and Type-II diabetes. Protein submitochondrial localization enables the understanding of protein function in studying disease pathogenesis and drug design. RESULTS: We proposed a new method, SubMito-XGBoost, for protein submitochondrial localization prediction. Three steps are included: (i) the g-gap dipeptide composition (g-gap DC), pseudo-amino acid composition (PseAAC), auto-correlation function (ACF) and Bi-gram position-specific scoring matrix (Bi-gram PSSM) are employed to extract protein sequence features, (ii) Synthetic Minority Oversampling Technique (SMOTE) is used to balance samples, and the ReliefF algorithm is applied for feature selection and (iii) the obtained feature vectors are fed into XGBoost to predict protein submitochondrial locations. SubMito-XGBoost has obtained satisfactory prediction results by the leave-one-out-cross-validation (LOOCV) compared with existing methods. The prediction accuracies of the SubMito-XGBoost method on the two training datasets M317 and M983 were 97.7% and 98.9%, which are 2.8-12.5% and 3.8-9.9% higher than other methods, respectively. The prediction accuracy of the independent test set M495 was 94.8%, which is significantly better than the existing studies. The proposed method also achieves satisfactory predictive performance on plant and non-plant protein submitochondrial datasets. SubMito-XGBoost also plays an important role in new drug design for the treatment of related diseases. AVAILABILITY AND IMPLEMENTATION: The source codes and data are publicly available at https://github.com/QUST-AIBBDRC/SubMito-XGBoost/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bin Yu 0007, Wenying Qiu, Cheng Chen 0051, Anjun Ma, Qin Ma 0003 |
Bioinform. | 7 |
| 2020 | SulSite-GTB: identification of protein S-sulfenylation sites by fusing multiple feature information and gradient tree boosting
Xiaowen Cui, Bin Yu 0007, Cheng Chen 0051, Qin Ma 0003 |
Neural Comput. Appl. | 5 |
| 2019 | DOOR: a prokaryotic operon database for genome analyses and functional inferenceabstractThe rapid accumulation of fully sequenced prokaryotic genomes provides unprecedented information for biological studies of bacterial and archaeal organisms in a systematic manner. Operons are the basic functional units for conducting such studies. Here, we review an operon database DOOR (the Database of prOkaryotic OpeRons) that we have previously developed and continue to update. Currently, the database contains 6 975 454 computationally predicted operons in 2072 complete genomes. In addition, the database also contains the following information: (i) transcriptional units for 24 genomes derived using publicly available transcriptomic data; (ii) orthologous gene mapping across genomes; (iii) 6408 cis-regulatory motifs for transcriptional factors of some operons for 203 genomes; (iv) 3 456 718 Rho-independent terminators for 2072 genomes; as well as (v) a suite of tools in support of applications of the predicted operons. In this review, we will explain how such data are computationally derived and demonstrate how they can be used to derive a wide range of higher-level information needed for systems biology studies to tackle complex and fundamental biology questions. Huansheng Cao, Qin Ma 0003, Ying Xu 0001 |
Briefings Bioinform. | 2 |
| 2019 | LncFinder: an integrated platform for long non-coding RNA identification utilizing sequence intrinsic composition, structural information and physicochemical propertyabstractDiscovering new long non-coding RNAs (lncRNAs) has been a fundamental step in lncRNA-related research. Nowadays, many machine learning-based tools have been developed for lncRNA identification. However, many methods predict lncRNAs using sequence-derived features alone, which tend to display unstable performances on different species. Moreover, the majority of tools cannot be re-trained or tailored by users and neither can the features be customized or integrated to meet researchers' requirements. In this study, features extracted from sequence-intrinsic composition, secondary structure and physicochemical property are comprehensively reviewed and evaluated. An integrated platform named LncFinder is also developed to enhance the performance and promote the research of lncRNA identification. LncFinder includes a novel lncRNA predictor using the heterologous features we designed. Experimental results show that our method outperforms several state-of-the-art tools on multiple species with more robust and satisfactory results. Researchers can additionally employ LncFinder to extract various classic features, build classifier with numerous machine learning algorithms and evaluate classifier performance effectively and efficiently. LncFinder can reveal the properties of lncRNA and mRNA from various perspectives and further inspire lncRNA-protein interaction prediction and lncRNA evolution analysis. It is anticipated that LncFinder can significantly facilitate lncRNA-related research, especially for the poorly explored species. LncFinder is released as R package (https://CRAN.R-project.org/package=LncFinder). A web server (http://bmbl.sdstate.edu/lncfinder/) is also developed to maximize its availability. Yanchun Liang 0001, Qin Ma 0003, Yangyi Xu, Wei Du 0002, Cankun Wang, Ying Li 0004 |
Briefings Bioinform. | 3 |
| 2019 | Interpretation of differential gene expression results of RNA-seq data: review and integrationabstractDifferential gene expression (DGE) analysis is one of the most common applications of RNA-sequencing (RNA-seq) data. This process allows for the elucidation of differentially expressed genes across two or more conditions and is widely used in many applications of RNA-seq data analysis. Interpretation of the DGE results can be nonintuitive and time consuming due to the variety of formats based on the tool of choice and the numerous pieces of information provided in these results files. Here we reviewed DGE results analysis from a functional point of view for various visualizations. We also provide an R/Bioconductor package, Visualization of Differential Gene Expression Results using R, which generates information-rich visualizations for the interpretation of DGE results from three widely used tools, Cuffdiff, DESeq2 and edgeR. The implemented functions are also tested on five real-world data sets, consisting of one human, one Malus domestica and three Vitis riparia data sets. Adam McDermaid, Brandon Monier, Bingqiang Liu, Qin Ma 0003 |
Briefings Bioinform. | 5 |
| 2019 | It is time to apply biclustering: a comprehensive review of biclustering applications in biological and biomedical dataabstractBiclustering is a powerful data mining technique that allows clustering of rows and columns, simultaneously, in a matrix-format data set. It was first applied to gene expression data in 2000, aiming to identify co-expressed genes under a subset of all the conditions/samples. During the past 17 years, tens of biclustering algorithms and tools have been developed to enhance the ability to make sense out of large data sets generated in the wake of high-throughput omics technologies. These algorithms and tools have been applied to a wide variety of data types, including but not limited to, genomes, transcriptomes, exomes, epigenomes, phenomes and pharmacogenomes. However, there is still a considerable gap between biclustering methodology development and comprehensive data interpretation, mainly because of the lack of knowledge for the selection of appropriate biclustering tools and further supporting computational techniques in specific studies. Here, we first deliver a brief introduction to the existing biclustering algorithms and tools in public domain, and then systematically summarize the basic applications of biclustering for biological data and more advanced applications of biclustering for biomedical data. This review will assist researchers to effectively analyze their big data and generate valuable biological knowledge and novel insights with higher efficiency. Juan Xie, Anjun Ma, Anne Fennell, Qin Ma 0003 |
Briefings Bioinform. | 4 |
| 2019 | MetaQUBIC: a computational pipeline for gene-level functional profiling of metagenome and metatranscriptomeabstractMOTIVATION: Metagenomic and metatranscriptomic analyses can provide an abundance of information related to microbial communities. However, straightforward analysis of this data does not provide optimal results, with a required integration of data types being needed to thoroughly investigate these microbiomes and their environmental interactions. RESULTS: Here, we present MetaQUBIC, an integrated biclustering-based computational pipeline for gene module detection that integrates both metagenomic and metatranscriptomic data. Additionally, we used this pipeline to investigate 735 paired DNA and RNA human gut microbiome samples, resulting in a comprehensive hybrid gene expression matrix of 2.3 million cross-species genes in the 735 human fecal samples and 155 functional enriched gene modules. We believe both the MetaQUBIC pipeline and the generated comprehensive human gut hybrid expression matrix will facilitate further investigations into multiple levels of microbiome studies. AVAILABILITY AND IMPLEMENTATION: The package is freely available at https://github.com/OSU-BMBL/metaqubic. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Anjun Ma, Minxuan Sun, Adam McDermaid, Bingqiang Liu, Qin Ma 0003 |
Bioinform. | 5 |
| 2019 | MetaQUBIC: a computational pipeline for gene-level functional profiling of metagenome and metatranscriptomeabstractBioinformatics (2019) doi: 10.1093/bioinformatics/btz414, 35, 4474–4477. An incomplete supplementary data file was published alongside the above article. This has now been replaced with the complete version. Anjun Ma, Minxuan Sun, Adam McDermaid, Bingqiang Liu, Qin Ma 0003 |
Bioinform. | 5 |
| 2019 | Protein-protein interaction sites prediction by ensemble random forests with synthetic minority oversampling techniqueabstractMOTIVATION: The prediction of protein-protein interaction (PPI) sites is a key to mutation design, catalytic reaction and the reconstruction of PPI networks. It is a challenging task considering the significant abundant sequences and the imbalance issue in samples. RESULTS: A new ensemble learning-based method, Ensemble Learning of synthetic minority oversampling technique (SMOTE) for Unbalancing samples and RF algorithm (EL-SMURF), was proposed for PPI sites prediction in this study. The sequence profile feature and the residue evolution rates were combined for feature extraction of neighboring residues using a sliding window, and the SMOTE was applied to oversample interface residues in the feature space for the imbalance problem. The Multi-dimensional Scaling feature selection method was implemented to reduce feature redundancy and subset selection. Finally, the Random Forest classifiers were applied to build the ensemble learning model, and the optimal feature vectors were inserted into EL-SMURF to predict PPI sites. The performance validation of EL-SMURF on two independent validation datasets showed 77.1% and 77.7% accuracy, which were 6.2-15.7% and 6.1-18.9% higher than the other existing tools, respectively. AVAILABILITY AND IMPLEMENTATION: The source codes and data used in this study are publicly available at http://github.com/QUST-AIBBDRC/EL-SMURF/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Bin Yu 0007, Anjun Ma, Cheng Chen 0051, Bingqiang Liu, Qin Ma 0003 |
Bioinform. | 6 |
| 2019 | M3S: a comprehensive model selection for multi-modal single-cell RNA sequencing dataabstractBACKGROUND: Various statistical models have been developed to model the single cell RNA-seq expression profiles, capture its multimodality, and conduct differential gene expression test. However, for expression data generated by different experimental design and platforms, there is currently lack of capability to determine the most proper statistical model. RESULTS: We developed an R package, namely Multi-Modal Model Selection (M3S), for gene-wise selection of the most proper multi-modality statistical model and downstream analysis, useful in a single-cell or large scale bulk tissue transcriptomic data. M3S is featured with (1) gene-wise selection of the most parsimonious model among 11 most commonly utilized ones, that can best fit the expression distribution of the gene, (2) parameter estimation of a selected model, and (3) differential gene expression test based on the selected model. CONCLUSION: A comprehensive evaluation suggested that M3S can accurately capture the multimodality on simulated and real single cell data. An open source package and is available through GitHub at https://github.com/zy26/M3S. Changlin Wan, Wennan Chang, Qin Ma 0003, Sha Cao, Chi Zhang 0021 |
BMC Bioinform. | 7 |
| 2019 | IRIS-EDA: An integrated RNA-Seq interpretation system for gene expression data analysisabstractNext-Generation Sequencing has made available substantial amounts of large-scale Omics data, providing unprecedented opportunities to understand complex biological systems. Specifically, the value of RNA-Sequencing (RNA-Seq) data has been confirmed in inferring how gene regulatory systems will respond under various conditions (bulk data) or cell types (single-cell data). RNA-Seq can generate genome-scale gene expression profiles that can be further analyzed using correlation analysis, co-expression analysis, clustering, differential gene expression (DGE), among many other studies. While these analyses can provide invaluable information related to gene expression, integration and interpretation of the results can prove challenging. Here we present a tool called IRIS-EDA, which is a Shiny web server for expression data analysis. It provides a straightforward and user-friendly platform for performing numerous computational analyses on user-provided RNA-Seq or Single-cell RNA-Seq (scRNA-Seq) data. Specifically, three commonly used R packages (edgeR, DESeq2, and limma) are implemented in the DGE analysis with seven unique experimental design functionalities, including a user-specified design matrix option. Seven discovery-driven methods and tools (correlation analysis, heatmap, clustering, biclustering, Principal Component Analysis (PCA), Multidimensional Scaling (MDS), and t-distributed Stochastic Neighbor Embedding (t-SNE)) are provided for gene expression exploration which is useful for designing experimental hypotheses and determining key factors for comprehensive DGE analysis. Furthermore, this platform integrates seven visualization tools in a highly interactive manner, for improved interpretation of the analyses. It is noteworthy that, for the first time, IRIS-EDA provides a framework to expedite submission of data and results to NCBI's Gene Expression Omnibus following the FAIR (Findable, Accessible, Interoperable and Reusable) Data Principles. IRIS-EDA is freely available at http://bmbl.sdstate.edu/IRIS/. Brandon Monier, Adam McDermaid, Cankun Wang, Allison J. Miller, Anne Fennell, Qin Ma 0003 |
PLoS Comput. Biol. | 7 |
| 2019 | Computational Prediction of Sigma-54 Promoters in Bacterial Genomes by Integrating Motif Finding and Machine Learning StrategiesabstractSigma factor, as a unit of RNA polymerase holoenzyme, is a critical factor in the process of gene transcriptional regulation. It recognizes the specific DNA sites and brings the core enzyme of RNA polymerase to the upstream regions of target genes. Therefore, the prediction of the promoters for a particular sigma factor is essential for interpreting functional genomic data and observation. This paper develops a new method to predict sigma-54 promoters in bacterial genomes. The new method organically integrates motif finding and machine learning strategies to capture the intrinsic features of sigma-54 promoters. The experiments on E. coli benchmark test set show that our method has good capability to distinguish sigma-54 promoters from surrounding or randomly selected DNA sequences. The applications of the other three bacterial genomes indicate the potential robustness and applicative power of our method on a large number of bacterial genomes. The source code of our method can be freely downloaded at https://github.com/maqin2001/PromotePredictor. Bingqiang Liu, Xiangrong Liu, Jichang Wu, Qin Ma 0003 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 5 |
| 2018 | An algorithmic perspective of de novo cis-regulatory motif finding based on ChIP-seq dataabstractTranscription factors are proteins that bind to specific DNA sequences and play important roles in controlling the expression levels of their target genes. Hence, prediction of transcription factor binding sites (TFBSs) provides a solid foundation for inferring gene regulatory mechanisms and building regulatory networks for a genome. Chromatin immunoprecipitation sequencing (ChIP-seq) technology can generate large-scale experimental data for such protein-DNA interactions, providing an unprecedented opportunity to identify TFBSs (a.k.a. cis-regulatory motifs). The bottleneck, however, is the lack of robust mathematical models, as well as efficient computational methods for TFBS prediction to make effective use of massive ChIP-seq data sets in the public domain. The purpose of this study is to review existing motif-finding methods for ChIP-seq data from an algorithmic perspective and provide new computational insight into this field. The state-of-the-art methods were shown through summarizing eight representative motif-finding algorithms along with corresponding challenges, and introducing some important relative functions according to specific biological demands, including discriminative motif finding and cofactor motifs analysis. Finally, potential directions and plans for ChIP-seq-based motif-finding tools were showcased in support of future algorithm development. Bingqiang Liu, Yang Li 0089, Adam McDermaid, Qin Ma 0003 |
Briefings Bioinform. | 5 |
| 2018 | Bioinformatics tools for quantitative and functional metagenome and metatranscriptome data analysis in microbesabstractMetagenomic and metatranscriptomic sequencing approaches are more frequently being used to link microbiota to important diseases and ecological changes. Many analyses have been used to compare the taxonomic and functional profiles of microbiota across habitats or individuals. While a large portion of metagenomic analyses focus on species-level profiling, some studies use strain-level metagenomic analyses to investigate the relationship between specific strains and certain circumstances. Metatranscriptomic analysis provides another important insight into activities of genes by examining gene expression levels of microbiota. Hence, combining metagenomic and metatranscriptomic analyses will help understand the activity or enrichment of a given gene set, such as drug-resistant genes among microbiome samples. Here, we summarize existing bioinformatics tools of metagenomic and metatranscriptomic data analysis, the purpose of which is to assist researchers in deciding the appropriate tools for their microbiome studies. Additionally, we propose an Integrated Meta-Function mapping pipeline to incorporate various reference databases and accelerate functional gene mapping procedures for both metagenomic and metatranscriptomic analyses. Sheng-Yong Niu, Adam McDermaid, Yu Kang 0004, Qin Ma 0003 |
Briefings Bioinform. | 6 |
| 2018 | Bioinformatics tools for quantitative and functional metagenome and metatranscriptome data analysis in microbesabstractBriefings in Bioinformatics, 2017. https://doi.org/10.1093/bib/bbx051 There was an error in the funding section. The correct paragraph is given below. “This work was supported by the State of South Dakota Research Innovation Center, the Agriculture Experiment Station of South Dakota State University. Support for this project was also provided by the Sanford Health – South Dakota State University Collaborative Research Seed Grant Program. This work used the Extreme Science and Engineering Discovery Environment (XSEDE), which is supported by National Science Foundation grant number ACI-1548562.” The online paper has been corrected. Sheng-Yong Niu, Adam McDermaid, Yu Kang 0004, Qin Ma 0003 |
Briefings Bioinform. | 6 |
| 2017 | DMINDA 2.0: integrated and systematic views of regulatory DNA motif identification and analysesabstractMOTIVATION: Motif identification and analyses are important and have been long-standing computational problems in bioinformatics. Substantial efforts have been made in this field during the past several decades. However, the lack of intuitive and integrative web servers impedes the progress of making effective use of emerging algorithms and tools. RESULTS: Here we present an integrated web server, DMINDA 2.0, which contains: (i) five motif prediction and analyses algorithms, including a phylogenetic footprinting framework; (ii) 2125 species with complete genomes to support the above five functions, covering animals, plants and bacteria and (iii) bacterial regulon prediction and visualization. AVAILABILITY AND IMPLEMENTATION: DMINDA 2.0 is freely available at http://bmbl.sdstate.edu/DMINDA2. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Adam McDermaid, Qin Ma 0003 |
Bioinform. | 4 |
| 2017 | QUBIC: a bioconductor package for qualitative biclustering analysis of gene co-expression dataabstractMotivation: Biclustering is widely used to identify co-expressed genes under subsets of all the conditions in a large-scale transcriptomic dataset. The program, QUBIC, is recognized as one of the most efficient and effective biclustering methods for biological data interpretation. However, its availability is limited to a C implementation and to a low-throughput web interface. Results: An R implementation of QUBIC is presented here with two unique features: (i) a 82% average improved efficiency by refactoring and optimizing the source C code of QUBIC; and (ii) a set of comprehensive functions to facilitate biclustering-based biological studies, including the qualitative representation (discretization) of expression data, query-based biclustering, bicluster expanding, biclusters comparison, heatmap visualization of any identified biclusters and co-expression networks elucidation. Availability and Implementation: The package is implemented in R (as of version 3.3) and is available from Bioconductor at the URL: http://bioconductor.org/packages/QUBIC, where installation and usage instructions can be found. Contact: [email protected] Supplimentary Information: Supplementary data are available at Bioinformatics online. Juan Xie, Anne Fennell, Chi Zhang 0021, Qin Ma 0003 |
Bioinform. | 6 |
| 2017 | RNA-TVcurve: a Web server for RNA secondary structure comparison based on a multi-scale similarity of its triple vector curve representationabstractBACKGROUND: RNAs have been found to carry diverse functionalities in nature. Inferring the similarity between two given RNAs is a fundamental step to understand and interpret their functional relationship. The majority of functional RNAs show conserved secondary structures, rather than sequence conservation. Those algorithms relying on sequence-based features usually have limitations in their prediction performance. Hence, integrating RNA structure features is very critical for RNA analysis. Existing algorithms mainly fall into two categories: alignment-based and alignment-free. The alignment-free algorithms of RNA comparison usually have lower time complexity than alignment-based algorithms. RESULTS: An alignment-free RNA comparison algorithm was proposed, in which novel numerical representations RNA-TVcurve (triple vector curve representation) of RNA sequence and corresponding secondary structure features are provided. Then a multi-scale similarity score of two given RNAs was designed based on wavelet decomposition of their numerical representation. In support of RNA mutation and phylogenetic analysis, a web server (RNA-TVcurve) was designed based on this alignment-free RNA comparison algorithm. It provides three functional modules: 1) visualization of numerical representation of RNA secondary structure; 2) detection of single-point mutation based on secondary structure; and 3) comparison of pairwise and multiple RNA secondary structures. The inputs of the web server require RNA primary sequences, while corresponding secondary structures are optional. For the primary sequences alone, the web server can compute the secondary structures using free energy minimization algorithm in terms of RNAfold tool from Vienna RNA package. CONCLUSION: RNA-TVcurve is the first integrated web server, based on an alignment-free method, to deliver a suite of RNA analysis functions, including visualization, mutation analysis and multiple RNAs structure comparison. The comparison results with two popular RNA comparison tools, RNApdist and RNAdistance, showcased that RNA-TVcurve can efficiently capture subtle relationships among RNAs for mutation detection and non-coding RNA classification. All the relevant results were shown in an intuitive graphical manner, and can be freely downloaded from this server. RNA-TVcurve, along with test examples and detailed documents, are available at: http://ml.jlu.edu.cn/tvcurve/ . Ying Li 0004, Xiaohu Shi, Yanchun Liang 0001, Juan Xie, Qin Ma 0003 |
BMC Bioinform. | 6 |
| 2015 | Revisiting operons: an analysis of the landscape of transcriptional units in E. coliabstractBACKGROUND: Bacterial operons are considerably more complex than what were thought. At least their components are dynamically rather than statically defined as previously assumed. Here we present a computational study of the landscape of the transcriptional units (TUs) of E. coli K12, revealed by the available genomic and transcriptomic data, providing new understanding about the complexity of TUs as a whole encoded in the genome of E. coli K12. RESULTS AND CONCLUSION: Our main findings include that (i) different TUs may overlap with each other by sharing common genes, giving rise to clusters of overlapped TUs (TUCs) along the genomic sequence; (ii) the intergenic regions in front of the first gene of each TU tend to have more conserved sequence motifs than those of the other genes inside the TU, suggesting that TUs each have their own promoters; (iii) the terminators associated with the 3' ends of TUCs tend to be Rho-independent terminators, substantially more often than terminators of TUs that end inside a TUC; and (iv) the functional relatedness of adjacent gene pairs in individual TUs is higher than those in TUCs, suggesting that individual TUs are more basic functional units than TUCs. Xizeng Mao, Qin Ma 0003, Bingqiang Liu, Ying Xu 0001 |
BMC Bioinform. | 2 |
| 2013 | An integrated toolkit for accurate prediction and analysis of cis-regulatory motifs at a genome scaleabstractMOTIVATION: We present an integrated toolkit, BoBro2.0, for prediction and analysis of cis-regulatory motifs. This toolkit can (i) reliably identify statistically significant cis-regulatory motifs at a genome scale; (ii) accurately scan for all motif instances of a query motif in specified genomic regions using a novel method for P-value estimation; (iii) provide highly reliable comparisons and clustering of identified motifs, which takes into consideration the weak signals from the flanking regions of the motifs; and (iv) analyze co-occurring motifs in the regulatory regions. RESULTS: We have carried out systematic comparisons between motif predictions using BoBro2.0 and the MEME package. The comparison results on Escherichia coli K12 genome and the human genome show that BoBro2.0 can identify the statistically significant motifs at a genome scale more efficiently, identify motif instances more accurately and get more reliable motif clusters than MEME. In addition, BoBro2.0 provides correlational analyses among the identified motifs to facilitate the inference of joint regulation relationships of transcription factors. AVAILABILITY: The source code of the program is freely available for noncommercial uses at http://code.google.com/p/bobro/. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Qin Ma 0003, Bingqiang Liu, Chuan Zhou 0009, Yanbin Yin, Ying Xu 0001 |
Bioinform. | 1 |