Jia Meng 0001

dblp:50/3810-1 · DBLP profile ↗
← Back
31ranked-venue papers
5as first author
11since 2021 · last 2026
0000-0003-3455-205XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 28 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CircRM: profiling circular RNA modifications from nanopore direct RNA sequencing
abstract
Circular RNA (circRNA) represents a critical class of regulatory RNAs with distinctive structural and functional features. The functions of circRNAs are modulated by various RNA modifications. Here, we present CircRM, a nanopore direct RNA sequencing-based computational method for profiling RNA modifications in circRNAs at single-base and single-molecule resolution. By integrating circRNA detection, read-level modification detection, and quantitative assessment of methylation rates, CircRM identified 427 high-confidence circRNAs and enables systematic characterization of three major modifications, m5C (AUC = 0.855), m6A (AUC = 0.817) and m1A (AUC = 0.769). It revealed distinct modification patterns compared with linear RNAs, highlighting RNA-type-specific regulations. We also identified the key features of circRNA-specific modifications, such as the enrichment near the back-splice junctions. Cross-cell line analyses further demonstrated conserved and cell-type-specific modification patterns. Together, these findings reveal, at the computational level, a unique epitranscriptomic landscape associated with circRNAs and establish CircRM as a powerful tool for advancing the study of RNA modifications in circular RNA biology. CircRM is free accessible at: https://github.com/jiayiAnnie17/CircRM.
Shenglun Chen, Zhixing Wu, Haozhe Wang 0020, Rong Xia, Jia Meng 0001
Briefings Bioinform.6
2026 DirectASRM: uncovering allele-specific post-transcriptional RNA modifications through direct RNA sequencing
abstract
SUMMARY: We developed DirectASRM, a comprehensive database for the systematic identification, integration, and annotation of allele-specific RNA modifications (ASRMs) from direct RNA sequencing data. DirectASRM enables single-base, transcript-level detection of ASRMs across multiple RNA modification types, diverse organisms and condition-specific contexts. The database further evaluates the confidence of each ASRM-SNP pair association within isoform context by jointly considering statistical evidence of allelic modification imbalance and independent support from external next-generation sequencing (NGS) - based RNA modification resources. DirectASRM also provides extensive functional annotations for ASRMs and their associated variants, including intra-sample transcript-level allele-specific expression (ASE) and allele-specific splicing, as well as additional post-transcriptional regulatory features such as miRNA binding, circRNA, RNA-protein interactions, and disease relevance. Overall, DirectASRM serves as a comprehensive resource that supports systematic investigation of the potential functional impact of genetic variants in epitranscriptomic regulation. AVAILABILITY AND IMPLEMENTATION: DirectASRM database is freely accessible at http://modinfor.com/DirectASRM/. DirectASRM pipeline is available at GitHub (https://github.com/jiayin1101/DirectASRM_pipeline) and Zenodo (DOI: https://doi.org/10.5281/zenodo.19876077).
Jiayin Dai, Jiongming Ma, Kunqi Chen, Jia Meng 0001, Daniel J. Rigden, Shaofeng Lin, Qingru Xu
Bioinform.6
2025 Multimodal zero-shot learning of previously unseen epitranscriptomes from RNA-seq data
abstract
Precise identification of condition-specific epitranscriptomes is of critical importance for investigating the dynamics and versatile functions of RNA modification under various biological contexts. Existing approaches for predicting condition-specific RNA modification are usually trained on epitranscriptome data obtained from the same condition, which limited their usage, as such data are available only for a small number of conditions due to the technical difficulties and high expenses of epitranscriptome profiling technologies. We present ExpressRM, a multimodal zero-shot learning framework for predicting condition-specific RNA modification sites in previously unseen contexts from genome and RNA-seq data. Different from existing in-condition learning approaches, this method does not rely on matched epitranscriptome data for training, which greatly expands its applicability. On a benchmark dataset comprising epitranscriptomes and matched transcriptomes of 37 human tissues, we demonstrate that ExpressRM can accurately predict epitranscriptomes of previously unseen conditions from their transcriptomes only, and the performance is comparable to existing in-condition learning algorithms that require epitranscriptome data from the same condition. Additionally, the method has the capability of differentiating highly dynamic RNA methylation sites from more static (or house-keeping) ones. With a case study, we show that ExpressRM can uncover N6-methyladenosine RNA methylation sites in glioblastoma using only its RNA-seq data, and unveils novel and previously validated pathological insights. Together, these results suggest that the proposed multimodal zero-shot learning framework can effectively leverage transcriptome knowledge to explore the dynamic roles of RNA modifications in previously unseen experimental setups, providing valuable insights into vast biological contexts where RNA-seq is routinely used but epitranscriptome profiling has not yet been covered.
Yiyou Song, Daiyun Huang, Li Hong Hu, Jia Meng 0001, Yue Wang 0060
Briefings Bioinform.6
2023 Multi-task adaptive pooling enabled synergetic learning of RNA modification across tissue, type and species from low-resolution epitranscriptomes
abstract
Post- and co-transcriptional RNA modifications are found to play various roles in regulating essential biological processes at all stages of RNA life. Precise identification of RNA modification sites is thus crucial for understanding the related molecular functions and specific regulatory circuitry. To date, a number of computational approaches have been developed for in silico identification of RNA modification sites; however, most of them require learning from base-resolution epitranscriptome datasets, which are generally scarce and available only for a limited number of experimental conditions, and predict only a single modification, even though there are multiple inter-related RNA modification types available. In this study, we proposed AdaptRM, a multi-task computational method for synergetic learning of multi-tissue, type and species RNA modifications from both high- and low-resolution epitranscriptome datasets. By taking advantage of adaptive pooling and multi-task learning, the newly proposed AdaptRM approach outperformed the state-of-the-art computational models (WeakRM and TS-m6A-DL) and two other deep-learning architectures based on Transformer and ConvMixer in three different case studies for both high-resolution and low-resolution prediction tasks, demonstrating its effectiveness and generalization ability. In addition, by interpreting the learned models, we unveiled for the first time the potential association between different tissues in terms of epitranscriptome sequence patterns. AdaptRM is available as a user-friendly web server from http://www.rnamd.org/AdaptRM together with all the codes and data used in this project.
Yiyou Song, Yue Wang 0060, Daiyun Huang, Jia Meng 0001
Briefings Bioinform.6
2022 Stepwise Feature Fusion: Local Guides Global
Jinfeng Wang 0008, Qiming Huang, Jia Meng 0001, Jionglong Su, Sifan Song
MICCAI (3)4
2022 A New Convolutional Neural Network Architecture for Automatic Segmentation of Overlapping Human Chromosomes
Sifan Song, Tianming Bai, Yanxin Zhao, Wenbo Zhang 0010, Chunxiao Yang, Jia Meng 0001, Fei Ma 0002, Jionglong Su
Neural Process. Lett.6
2022 FBCwPlaid: A Functional Biclustering Analysis of Epi-Transcriptome Profiling Data Via a Weighted Plaid Model
abstract
Recent studies have shown that in-depth studies on epi-transcriptomic patterns of N6-methyladenosine (m6A) may help understand its complex functions and co-regulatory mechanisms. Since most biclustering algorithms are developed in scenarios of gene expression analysis, which does not share the same characteristics with m6A methylation profile, we propose a weighted Plaid biclustering model (FBCwPlaid) based on the Lagrange multiplier method to discover the potential functional patterns. Each pattern is achieved by minimizing approximation error between FBCwPlaid predicted value and real data. To address the issue that site expression level determines methylation level confidence, it uses RNA expression levels of each site as weights to make lower expressed sites less confident. FBCwPlaid also allows overlapping biclusters, indicating some sites may participate in multiple biological functions. FBCwPlaid was then applied on MeRIP-Seq data of 69,446 methylation sites under 32 experimental conditions, each of which represented a stimulus to a particular cell line or environment. Finally, three patterns were discovered, and further pathway analysis and enzyme specificity test showed that sites involved in each pattern are highly relevant to m6A methyltransferases. Further detailed analyses showed that some patterns are condition-specific, indicating that some specific sites’ methylation profiles may occur in specific cell lines or conditions.
Shutao Chen, Lin Zhang 0015, Jia Meng 0001, Hui Liu 0024
IEEE ACM Trans. Comput. Biol. Bioinform.4
2021 ConsRM: collection and large-scale prediction of the evolutionarily conserved RNA methylation sites, with implications for the functional epitranscriptome
abstract
Motivation N6-methyladenosine (m6A) is the most prevalent RNA modification on mRNAs and lncRNAs. Evidence increasingly demonstrates its crucial importance in essential molecular mechanisms and various diseases. With recent advances in sequencing techniques, tens of thousands of m6A sites are identified in a typical high-throughput experiment, posing a key challenge to distinguish the functional m6A sites from the remaining 'passenger' (or 'silent') sites. Results: We performed a comparative conservation analysis of the human and mouse m6A epitranscriptomes at single site resolution. A novel scoring framework, ConsRM, was devised to quantitatively measure the degree of conservation of individual m6A sites. ConsRM integrates multiple information sources and a positive-unlabeled learning framework, which integrated genomic and sequence features to trace subtle hints of epitranscriptome layer conservation. With a series validation experiments in mouse, fly and zebrafish, we showed that ConsRM outperformed well-adopted conservation scores (phastCons and phyloP) in distinguishing the conserved and unconserved m6A sites. Additionally, the m6A sites with a higher ConsRM score are more likely to be functionally important. An online database was developed containing the conservation metrics of 177 998 distinct human m6A sites to support conservation analysis and functional prioritization of individual m6A sites. And it is freely accessible at: https://www.xjtlu.edu.cn/biologicalsciences/con.
Kunqi Chen, Yujiao Tang, Jionglong Su, João Pedro de Magalhães, Daniel J. Rigden, Jia Meng 0001
Briefings Bioinform.8
2021 Weakly supervised learning of RNA modifications from low-resolution epitranscriptome data
abstract
MOTIVATION: Increasing evidence suggests that post-transcriptional ribonucleic acid (RNA) modifications regulate essential biomolecular functions and are related to the pathogenesis of various diseases. Precise identification of RNA modification sites is essential for understanding the regulatory mechanisms of RNAs. To date, many computational approaches for predicting RNA modifications have been developed, most of which were based on strong supervision enabled by base-resolution epitranscriptome data. However, high-resolution data may not be available. RESULTS: We propose WeakRM, the first weakly supervised learning framework for predicting RNA modifications from low-resolution epitranscriptome datasets, such as those generated from acRIP-seq and hMeRIP-seq. Evaluations on three independent datasets (corresponding to three different RNA modification types and their respective sequencing technologies) demonstrated the effectiveness of our approach in predicting RNA modifications from low-resolution data. WeakRM outperformed state-of-the-art multi-instance learning methods for genomic sequences, such as WSCNN, which was originally designed for transcription factor binding site prediction. Additionally, our approach captured motifs that are consistent with existing knowledge, and visualization of the predicted modification-containing regions unveiled the potentials of detecting RNA modifications with improved resolution. AVAILABILITY IMPLEMENTATION: The source code for the WeakRM algorithm, along with the datasets used, are freely accessible at: https://github.com/daiyun02211/WeakRM. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Daiyun Huang, Jingjue Wei, Jionglong Su, Frans Coenen, Jia Meng 0001
Bioinform.6
2021 MetaTX: deciphering the distribution of mRNA-related features in the presence of isoform ambiguity, with applications in epitranscriptome analysis
abstract
MOTIVATION: The distribution of biological features strongly indicates their functional relevance. Compared to DNA-related features, deciphering the distribution of mRNA-related features is non-trivial due to the existence of isoform ambiguity and compositional diversity of mRNAs. RESULTS: We propose here a rigorous statistical framework, MetaTX, for deciphering the distribution of mRNA-related features. Through a standardized mRNA model, MetaTX firstly unifies various mRNA transcripts of diverse compositions, and then corrects the isoform ambiguity by incorporating the overall distribution pattern of the features through an EM algorithm. MetaTX was tested on both simulated and real data. Results suggested that MetaTX substantially outperformed existing direct methods on simulated datasets, and that a more informative distribution pattern was produced for all the three datasets tested, which contain N6-Methyladenosine sites generated by different technologies. MetaTX should make a useful tool for studying the distribution and functions of mRNA-related biological features, especially for mRNA modifications such as N6-Methyladenosine. AVAILABILITY AND IMPLEMENTATION: The MetaTX R package is freely available at GitHub: https://github.com/yue-wang-biomath/MetaTX.1.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yue Wang 0060, Kunqi Chen, Frans Coenen, Jionglong Su, Jia Meng 0001
Bioinform.6
2021 Funm6AViewer: a web server and R package for functional analysis of context-specific m6A RNA methylation
abstract
MOTIVATION: N 6-methyladenosine (m6A) is the most abundant mammalian mRNA methylation with versatile functions. To date, although a number of bioinformatics tools have been developed for location discovery of m6A modification, functional understanding is still quite limited. As the focus of RNA epigenetics gradually shifts from site discovery to functional studies, there is an urgent need for user-friendly tools to identify and explore the functional relevance of context-specific m6A methylation to gain insights into the epitranscriptome layer of gene expression regulation. RESULTS: We introduced here Funm6AViewer, a novel platform to identify, prioritize and visualize the functional gene interaction networks mediated by dynamic m6A RNA methylation unveiled from a case control study. By taking the differential RNA methylation data and differential gene expression data, both of which can be inferred from the widely used MeRIP-seq data, as the inputs, Funm6AViewer enables a series of analysis, including: (i) examining the distribution of differential m6A sites, (ii) prioritizing the genes mediated by dynamic m6A methylation and (iii) characterizing functionally the gene regulatory networks mediated by condition-specific m6A RNA methylation. Funm6AViewer should effectively facilitate the understanding of the epitranscriptome circuitry mediated by this reversible RNA modification. AVAILABILITY AND IMPLEMENTATION: Funm6AViewer is available both as a convenient web server (https://www.xjtlu.edu.cn/biologicalsciences/funm6aviewer) with graphical interface and as an independent R package (https://github.com/NWPU-903PR/Funm6AViewer) for local usage. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Songyao Zhang, Shaowu Zhang 0001, Yujiao Tang, Xiaonan Fan 0001, Jia Meng 0001
Bioinform.5
2020 m7GHub: deciphering the location, regulation and pathogenesis of internal mRNA N7-methylguanosine (m7G) sites in human
abstract
MOTIVATION: Recent progress in N7-methylguanosine (m7G) RNA methylation studies has focused on its internal (rather than capped) presence within mRNAs. Tens of thousands of internal mRNA m7G sites have been identified within mammalian transcriptomes, and a single resource to best share, annotate and analyze the massive m7G data generated recently are sorely needed. RESULTS: We report here m7GHub, a comprehensive online platform for deciphering the location, regulation and pathogenesis of internal mRNA m7G. The m7GHub consists of four main components, including: the first internal mRNA m7G database containing 44 058 experimentally validated internal mRNA m7G sites, a sequence-based high-accuracy predictor, the first web server for assessing the impact of mutations on m7G status, and the first database recording 1218 disease-associated genetic mutations that may function through regulation of m7G methylation. Together, m7GHub will serve as a useful resource for research on internal mRNA m7G modification. AVAILABILITY AND IMPLEMENTATION: m7GHub is freely accessible online at www.xjtlu.edu.cn/biologicalsciences/m7ghub. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yujiao Tang, Kunqi Chen, Rong Rong, Zhiliang Lu, Jionglong Su, João Pedro de Magalhães, Daniel J. Rigden, Jia Meng 0001
Bioinform.10
2020 REW-ISA: unveiling local functional blocks in epi-transcriptome profiling data via an RNA expression-weighted iterative signature algorithm
abstract
Abstract Background Recent studies have shown that N6-methyladenosine (m6A) plays a critical role in numbers of biological processes and complex human diseases. However, the regulatory mechanisms of most methylation sites remain uncharted. Thus, in-depth study of the epi-transcriptomic patterns of m6A may provide insights into its complex functional and regulatory mechanisms. Results Due to the high economic and time cost of wet experimental methods, revealing methylation patterns through computational models has become a more preferable way, and drawn more and more attention. Considering the theoretical basics and applications of conventional clustering methods, an RNA Expression Weighted Iterative Signature Algorithm (REW-ISA) is proposed to find potential local functional blocks (LFBs) based on MeRIP-Seq data, where sites are hyper-methylated or hypo-methylated simultaneously across the specific conditions. REW-ISA adopts RNA expression levels of each site as weights to make sites of lower expression level less significant. It starts from random sets of sites, then follows iterative search strategies by thresholds of rows and columns to find the LFBs in m6A methylation profile. Its application on MeRIP-Seq data of 69,446 methylation sites under 32 experimental conditions unveiled 6 LFBs, which achieve higher enrichment scores than ISA. Pathway analysis and enzyme specificity test showed that sites remained in LFBs are highly relevant to the m6A methyltransferase, such as METTL3, METTL14, WTAP and KIAA1429. Further detailed analyses for each LFB even showed that some LFBs are condition-specific, indicating that methylation profiles of some specific sites may be condition relevant. Conclusions REW-ISA finds potential local functional patterns presented in m6A profiles, where sites are co-methylated under specific conditions.
Lin Zhang 0015, Shutao Chen, Jia Meng 0001, Hui Liu 0024
BMC Bioinform.4
2019 RNA methylation and diseases: experimental results, databases, Web servers and computational models
abstract
Ribonucleic acid (RNA) methylation is a type of posttranscriptional modifications occurring in all kingdoms of life. It is strongly related to important biological process, thus making it linked to a number of human diseases. Owing to the development of high-throughput sequencing technology, plenty of achievement had been obtained in RNA methylation research recently. Meanwhile, various computational models have been developed to analyze and mining increasing RNA methylation data. In this review, we first made a brief introduction about eight types of most popular RNA methylation, the biological functions of RNA methylation, the relationship between RNA methylation and disease and five important RNA methylation-related diseases. The research of RNA methylation is based on sequencing data processing, and effective bioinformatics techniques can benefit better understanding of RNA methylation. We further introduced seven publicly available RNA methylation-related databases, and some important publicly available RNA-methylation-related Web servers and software for RNA methylation site identification, differential analysis and so on. Furthermore, we provided detailed analysis of the state-of-the-art computational models used in these Web servers and software. We also analyzed the limitations of these models and discussed the future directions of developing computational models for RNA methylation research.
Xing Chen 0001, Ya-Zhou Sun, Hui Liu 0024, Lin Zhang 0015, Jianqiang Li 0001, Jia Meng 0001
Briefings Bioinform.6
2019 FunDMDeep-m6A: identification and prioritization of functional differential m6A methylation genes
abstract
MOTIVATION: As the most abundant mammalian mRNA methylation, N6-methyladenosine (m6A) exists in >25% of human mRNAs and is involved in regulating many different aspects of mRNA metabolism, stem cell differentiation and diseases like cancer. However, our current knowledge about dynamic changes of m6A levels and how the change of m6A levels for a specific gene can play a role in certain biological processes like stem cell differentiation and diseases like cancer is largely elusive. RESULTS: To address this, we propose in this paper FunDMDeep-m6A a novel pipeline for identifying context-specific (e.g. disease versus normal, differentiated cells versus stem cells or gene knockdown cells versus wild-type cells) m6A-mediated functional genes. FunDMDeep-m6A includes, at the first step, DMDeep-m6A a novel method based on a deep learning model and a statistical test for identifying differential m6A methylation (DmM) sites from MeRIP-Seq data at a single-base resolution. FunDMDeep-m6A then identifies and prioritizes functional DmM genes (FDmMGenes) by combing the DmM genes (DmMGenes) with differential expression analysis using a network-based method. This proposed network method includes a novel m6A-signaling bridge (MSB) score to quantify the functional significance of DmMGenes by assessing functional interaction of DmMGenes with their signaling pathways using a heat diffusion process in protein-protein interaction (PPI) networks. The test results on 4 context-specific MeRIP-Seq datasets showed that FunDMDeep-m6A can identify more context-specific and functionally significant FDmMGenes than m6A-Driver. The functional enrichment analysis of these genes revealed that m6A targets key genes of many important context-related biological processes including embryonic development, stem cell differentiation, transcription, translation, cell death, cell proliferation and cancer-related pathways. These results demonstrate the power of FunDMDeep-m6A for elucidating m6A regulatory functions and its roles in biological processes and diseases. AVAILABILITY AND IMPLEMENTATION: The R-package for DMDeep-m6A is freely available from https://github.com/NWPU-903PR/DMDeepm6A1.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Songyao Zhang, Shaowu Zhang 0001, Xiaonan Fan 0001, Jia Meng 0001, Yufei Huang 0001
Bioinform.5
2019 m6Acomet: large-scale functional prediction of individual m6A RNA methylation sites from an RNA co-methylation network
abstract
Over one hundred different types of post-transcriptional RNA modifications have been identified in human. Researchers discovered that RNA modifications can regulate various biological processes, and RNA methylation, especially N6-methyladenosine, has become one of the most researched topics in epigenetics. To date, the study of epitranscriptome layer gene regulation is mostly focused on the function of mediator proteins of RNA methylation, i.e., the readers, writers and erasers. There is limited investigation of the functional relevance of individual m6A RNA methylation site. To address this, we annotated human m6A sites in large-scale based on the guilt-by-association principle from an RNA co-methylation network. It is constructed based on public human MeRIP-Seq datasets profiling the m6A epitranscriptome under 32 independent experimental conditions. By systematically examining the network characteristics obtained from the RNA methylation profiles, a total of 339,158 putative gene ontology functions associated with 1446 human m6A sites were identified. These are biological functions that may be regulated at epitranscriptome layer via reversible m6A RNA methylation. The results were further validated on a soft benchmark by comparing to a random predictor. An online web server m6Acomet was constructed to support direct query for the predicted biological functions of m6A sites as well as the sites exhibiting co-methylated patterns at the epitranscriptome layer. The m6Acomet web server is freely available at: www.xjtlu.edu.cn/biologicalsciences/m6acomet .
Kunqi Chen, Jionglong Su, Hui Liu 0024, Lin Zhang 0015, Jia Meng 0001
BMC Bioinform.8
2019 Global analysis of N6-methyladenosine functions and its disease association using deep learning and network-based methods
abstract
N6-methyladenosine (m6A) is the most abundant methylation, existing in >25% of human mRNAs. Exciting recent discoveries indicate the close involvement of m6A in regulating many different aspects of mRNA metabolism and diseases like cancer. However, our current knowledge about how m6A levels are controlled and whether and how regulation of m6A levels of a specific gene can play a role in cancer and other diseases is mostly elusive. We propose in this paper a computational scheme for predicting m6A-regulated genes and m6A-associated disease, which includes Deep-m6A, the first model for detecting condition-specific m6A sites from MeRIP-Seq data with a single base resolution using deep learning and Hot-m6A, a new network-based pipeline that prioritizes functional significant m6A genes and its associated diseases using the Protein-Protein Interaction (PPI) and gene-disease heterogeneous networks. We applied Deep-m6A and this pipeline to 75 MeRIP-seq human samples, which produced a compact set of 709 functionally significant m6A-regulated genes and nine functionally enriched subnetworks. The functional enrichment analysis of these genes and networks reveal that m6A targets key genes of many critical biological processes including transcription, cell organization and transport, and cell proliferation and cancer-related pathways such as Wnt pathway. The m6A-associated disease analysis prioritized five significantly associated diseases including leukemia and renal cell carcinoma. These results demonstrate the power of our proposed computational scheme and provide new leads for understanding m6A regulatory functions and its roles in diseases.
Songyao Zhang, Shaowu Zhang 0001, Xiaonan Fan 0001, Jia Meng 0001, Yidong Chen 0002, Shou-Jiang Gao, Yufei Huang 0001
PLoS Comput. Biol.4
2018 trumpet: transcriptome-guided quality assessment of m6A-seq data
abstract
Methylated RNA immunoprecipitation sequencing (MeRIP-seq or m 6 A-seq) has been extensively used for profiling transcriptome-wide distribution of RNA N6-Methyl-Adnosine methylation. However, due to the intrinsic properties of RNA molecules and the intricate procedures of this technique, m 6 A-seq data often suffer from various flaws. A convenient and comprehensive tool is needed to assess the quality of m 6 A-seq data to ensure that they are suitable for subsequent analysis. From a technical perspective, m 6 A-seq can be considered as a combination of ChIP-seq and RNA-seq; hence, by effectively combing the data quality assessment metrics of the two techniques, we developed the trumpet R package for evaluation of m 6 A-seq data quality. The trumpet package takes the aligned BAM files from m 6 A-seq data together with the transcriptome information as the inputs to generate a quality assessment report in the HTML format. The trumpet R package makes a valuable tool for assessing the data quality of m 6 A-seq, and it is also applicable to other fragmented RNA immunoprecipitation sequencing techniques, including m 1 A-seq, CeU-Seq, Ψ-seq, etc.
Shaowu Zhang 0001, Lin Zhang 0015, Jia Meng 0001
BMC Bioinform.4
2018 MeTDiff: A Novel Differential RNA Methylation Analysis for MeRIP-Seq Data
abstract
N6-Methyladenosine (m6A) transcriptome methylation is an exciting new research area that just captures the attention of research community. We present in this paper, MeTDiff, a novel computational tool for predicting differential m6A methylation sites from Methylated RNA immunoprecipitation sequencing (MeRIP-Seq) data. Compared with the existing algorithm exomePeak, the advantages of MeTDiff are that it explicitly models the reads variation in data and also devices a more power likelihood ratio test for differential methylation site prediction. Comprehensive evaluation of MeTDiff's performance using both simulated and real datasets showed that MeTDiff is much more robust and achieved much higher sensitivity and specificity over exomePeak.
Lin Zhang 0015, Jia Meng 0001, Manjeet K. Rao, Yidong Chen 0002, Yufei Huang 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2017 QNB: differential RNA methylation analysis for count-based small-sample sequencing data with a quad-negative binomial model
abstract
BACKGROUND: As a newly emerged research area, RNA epigenetics has drawn increasing attention recently for the participation of RNA methylation and other modifications in a number of crucial biological processes. Thanks to high throughput sequencing techniques, such as, MeRIP-Seq, transcriptome-wide RNA methylation profile is now available in the form of count-based data, with which it is often of interests to study the dynamics at epitranscriptomic layer. However, the sample size of RNA methylation experiment is usually very small due to its costs; and additionally, there usually exist a large number of genes whose methylation level cannot be accurately estimated due to their low expression level, making differential RNA methylation analysis a difficult task. RESULTS: We present QNB, a statistical approach for differential RNA methylation analysis with count-based small-sample sequencing data. Compared with previous approaches such as DRME model based on a statistical test covering the IP samples only with 2 negative binomial distributions, QNB is based on 4 independent negative binomial distributions with their variances and means linked by local regressions, and in the way, the input control samples are also properly taken care of. In addition, different from DRME approach, which relies only the input control sample only for estimating the background, QNB uses a more robust estimator for gene expression by combining information from both input and IP samples, which could largely improve the testing performance for very lowly expressed genes. CONCLUSION: A-Seq, Par-CLIP, RIP-Seq, etc.
Shaowu Zhang 0001, Yufei Huang 0001, Jia Meng 0001
BMC Bioinform.4
2017 Cancer Progression Prediction Using Gene Interaction Regularized Elastic Net
abstract
Different types of genomic aberration may simultaneously contribute to tumorigenesis. To obtain a more accurate prognostic assessment to guide therapeutic regimen choice for cancer patients, the heterogeneous multi-omics data should be integrated harmoniously, which can often be difficult. For this purpose, we propose a Gene Interaction Regularized Elastic Net (GIREN) model that predicts clinical outcome by integrating multiple data types. GIREN conveniently embraces both gene measurements and gene-gene interaction information under an elastic net formulation, enforcing structure sparsity, and the "grouping effect" in solution to select the discriminate features with prognostic value. An iterative gradient descent algorithm is also developed to solve the model with regularized optimization. GIREN was applied to human ovarian cancer and breast cancer datasets obtained from The Cancer Genome Atlas, respectively. Result shows that, the proposed GIREN algorithm obtained more accurate and robust performance over competing algorithms (LASSO, Elastic Net, and Semi-supervised PCA, with or without average pathway expression features) in predicting cancer progression on both two datasets in terms of median area under curve (AUC) and interquartile range (IQR), suggesting a promising direction for more effective integration of gene measurement and gene interaction information.
Lin Zhang 0015, Hui Liu 0024, Yufei Huang 0001, Xuesong Wang 0001, Yidong Chen 0002, Jia Meng 0001
IEEE ACM Trans. Comput. Biol. Bioinform.6
2016 A novel algorithm for calling mRNA m6A peaks by modeling biological variances in MeRIP-seq data
abstract
MOTIVATION: N(6)-methyl-adenosine (m(6)A) is the most prevalent mRNA methylation but precise prediction of its mRNA location is important for understanding its function. A recent sequencing technology, known as Methylated RNA Immunoprecipitation Sequencing technology (MeRIP-seq), has been developed for transcriptome-wide profiling of m(6)A. We previously developed a peak calling algorithm called exomePeak. However, exomePeak over-simplifies data characteristics and ignores the reads' variances among replicates or reads dependency across a site region. To further improve the performance, new model is needed to address these important issues of MeRIP-seq data. RESULTS: We propose a novel, graphical model-based peak calling method, MeTPeak, for transcriptome-wide detection of m(6)A sites from MeRIP-seq data. MeTPeak explicitly models read count of an m(6)A site and introduces a hierarchical layer of Beta variables to capture the variances and a Hidden Markov model to characterize the reads dependency across a site. In addition, we developed a constrained Newton's method and designed a log-barrier function to compute analytically intractable, positively constrained Beta parameters. We applied our algorithm to simulated and real biological datasets and demonstrated significant improvement in detection performance and robustness over exomePeak. Prediction results on publicly available MeRIP-seq datasets are also validated and shown to be able to recapitulate the known patterns of m(6)A, further validating the improved performance of MeTPeak. AVAILABILITY AND IMPLEMENTATION: The package 'MeTPeak' is implemented in R and C ++, and additional details are available at https://github.com/compgenomics/MeTPeak CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jia Meng 0001, Shaowu Zhang 0001, Yidong Chen 0002, Yufei Huang 0001
Bioinform.2
2016 m6A-Driver: Identifying Context-Specific mRNA m6A Methylation-Driven Gene Interaction Networks
abstract
As the most prevalent mammalian mRNA epigenetic modification, N6-methyladenosine (m6A) has been shown to possess important post-transcriptional regulatory functions. However, the regulatory mechanisms and functional circuits of m6A are still largely elusive. To help unveil the regulatory circuitry mediated by mRNA m6A methylation, we develop here m6A-Driver, an algorithm for predicting m6A-driven genes and associated networks, whose functional interactions are likely to be actively modulated by m6A methylation under a specific condition. Specifically, m6A-Driver integrates the PPI network and the predicted differential m6A methylation sites from methylated RNA immunoprecipitation sequencing (MeRIP-Seq) data using a Random Walk with Restart (RWR) algorithm and then builds a consensus m6A-driven network of m6A-driven genes. To evaluate the performance, we applied m6A-Driver to build the context-specific m6A-driven networks for 4 known m6A (de)methylases, i.e., FTO, METTL3, METTL14 and WTAP. Our results suggest that m6A-Driver can robustly and efficiently identify m6A-driven genes that are functionally more enriched and associated with higher degree of differential expression than differential m6A methylated genes. Pathway analysis of the constructed context-specific m6A-driven gene networks further revealed the regulatory circuitry underlying the dynamic interplays between the methyltransferases and demethylase at the epitranscriptomic layer of gene regulation.
Songyao Zhang, Shaowu Zhang 0001, Jia Meng 0001, Yufei Huang 0001
PLoS Comput. Biol.4
2015 Sketching the distribution of transcriptomic features on RNA transcripts with Travis coordinates
abstract
Biological features, such as, genes, transcription factor binding sites, SNPs, etc., are usually denoted with genome-based coordinates as the genomic features. While genome-based representation is usually very effective, it can be tedious to examine the distribution of RNA-related genomic features on RNA transcripts with existing tools due to the conversion and comparison between genome-based coordinates to RNA-based coordinates. We developed here an open source R package Travis for sketching the transcriptomic view of genomic features so as to facilitate the analysis of RNA-related but genome-based coordinates. Internally, Travis package extracts the coordinates relative to the landmarks of transcripts, with which the distribution of RNA-related genomic features can then be conveniently analyzed. We demonstrated the usage of Travis package in analyzing post-transcriptional RNA modifications (5-MethylCytosine and N6-MethylAdenosine) derived from high-throughput sequencing approaches (MeRIP-Seq and RNA BS-Seq). The Travis R package is now publicly available from GitHub: https://github.com/lzcyzm/Travis.
Lin Zhang 0015, Hui Liu 0024, Shaowu Zhang 0001, Yufei Huang 0001, Jia Meng 0001
BIBM8
2013 Unveiling the dynamics in RNA epigenetic regulations
abstract
Despite the prevalent studies of DNA/Chromatin related epigenetics, such as, histone modifications and DNA methylation, RNA epigenetics did not receive deserved attention due to the lack of high throughput approach for profiling epitranscriptome. Recently, a new affinity-based sequencing approach MeRIPseq was developed and applied to survey the global mRNA N6-methyladenosine (m6A) in mammalian cells. As a marriage of ChIPseq and RNAseq, MeRIPseq has the potential to study, for the first time, the transcriptome-wide distribution of different types of post-transcriptional RNA modifications. Yet, this technology introduced new computational challenges that have not been adequately addressed. We have previously developed a MATLAB-based package ‘exomePeak’ for detection of RNA methylation sites from MeRIPseq data. Here, we extend the features of exomePeak by including a novel computational framework that enables differential analysis to unveil the dynamics in RNA epigenetic regulations. The novel differential analysis monitors the percentage of modified RNA molecules among the total transcribed RNAs, which directly reflects the impact of RNA epigenetic regulations. In contrast, current available software packages developed for sequencing-based differential analysis such as DESeq or edgeR monitors the changes in the absolute amount of molecules, and, if applied to MeRIPseq data, might be dominated by transcriptional gene differential expression. The algorithm is implemented as an R-package ‘exomePeak’ and freely available. It takes directly the aligned BAM files as input, statistically supports biological replicates, corrects PCR artifacts, and outputs exome-based results in BED format, which is compatible with all major genome browsers for convenient visualization and manipulation. Examples are also provided to depict how exomePeak R-package is integrated with exiting tools for MeRIPseq based peak calling and differential analysis. Particularly, the rationales behind each processing step as well as the specific method used, the best practice, and possible alternative strategies are briefly discussed. The algorithm was applied to the human HepG2 cell MeRIPseq data sets and detects more than 16000 RNA m6A sites, many of which are differentially methylated under ultraviolet radiation. The challenges and potentials of MeRIPseq in epitranscriptome studies are discussed in the end.
Jia Meng 0001, Hui Liu 0024, Lin Zhang 0015, Shaowu Zhang 0001, Manjeet K. Rao, Yidong Chen 0002, Yufei Huang 0001
BIBM1
2013 Integration of gene expression, genome wide DNA methylation, and gene networks for clinical outcome prediction in ovarian cancer
abstract
Integrative clinical outcome prediction model called gene interaction regularized elastic net (GIREN) method is proposed in this paper. GIREN combines gene expression, methylation profiles, and gene interaction networks in order to reveal genomic and epigenomic features that bear important prognostic value. With GIREN, gene expression and DNA methylation profiles are first jointly analyzed in a linear regression model, and additional gene interaction network is simultaneously integrated as a regularizing penalty that follow an elastic net formulation. Such regularization also enforce sparsity in the solution so that features with prognostic values are automatically selected. To solve the regularized optimization, an iterative gradient descent algorithm is also developed. We applied GIREN to a set of 87 human ovarian cancer samples, which underwent a rigorous sample selection. The predicted outcome was used to group patients into high-risk vs. low-risk. Validation showed that GIREN outperformed other competing algorithms including SuperPCA.
Lin Zhang 0015, Hui Liu 0024, Jia Meng 0001, Xuesong Wang 0001, Yidong Chen 0002, Yufei Huang 0001
BIBM3
2013 A bag-of-words model for task-load prediction from EEG in complex environments
abstract
Neurotechnologies based on electroencephalography (EEG) and other physiological measures to improve task performance in complex environments will require tools and analysis methods that can account for increased environmental noise and task complexity compared to traditional neuroscience laboratory experiments. We propose a bag-of-words (BoW) model to address the difficulties associated with realistic applications in complex environments. In this paper, our proof-of-concept results show that a BoW classifier can discriminate two task-relevant states (high versus low task-load) while an individual performs a simulated security patrol mission with complex, concurrent tasking. Classifier performance is largely consistent across six simulation missions for a given participant, but performance decreases when trying to predict between two individuals. Overall, these initial results suggest that this BoW approach holds promise for detecting task-relevant states in real-world settings.
Lenis Mauricio Merino, Jia Meng 0001, Stephen M. Gordon, Brent Lance, Tony Johnson, Victor Paul, Kay A. Robbins, Jean M. Vettel, Yufei Huang 0001
ICASSP2
2013 Exome-based analysis for RNA epigenome sequencing data
abstract
MOTIVATION: Fragmented RNA immunoprecipitation combined with RNA sequencing enabled the unbiased study of RNA epigenome at a near single-base resolution; however, unique features of this new type of data call for novel computational techniques. RESULT: Through examining the connections of RNA epigenome sequencing data with two well-studied data types, ChIP-Seq and RNA-Seq, we unveiled the salient characteristics of this new data type. The computational strategies were discussed accordingly, and a novel data processing pipeline was proposed that combines several existing tools with a newly developed exome-based approach 'exomePeak' for detecting, representing and visualizing the post-transcriptional RNA modification sites on the transcriptome. AVAILABILITY: The MATLAB package 'exomePeak' and additional details are available at http://compgenomics.utsa.edu/exomePeak/.
Jia Meng 0001, Manjeet K. Rao, Yidong Chen 0002, Yufei Huang 0001
Bioinform.1
2011 Uncover cooperative gene regulations by microRNAs and transcription factors in glioblastoma using a nonnegative hybrid factor model
abstract
Transcriptional regulation by transcription factors (TFs) and microRNAs controls when and how much RNA is created. Due to technical limitations, the protein level expressions of TFs are usually unknown, making computational reconstruction of transcriptional network a difficult task. We proposed here a novel Bayesian non negative hybrid factor model for transcriptional network modeling, which is capable to estimate both the non-negative abundances of the transcription factors, the regulatory effects of TFs and microRNAs, and the sample clustering information by integrating microarray data and existing knowledge regarding TFs and microRNAs regulated target genes. The results demonstrated its validity and effectiveness to reconstructing transcriptional networks through simulated systems and real data.
Jia Meng 0001, Hung-I Harry Chen, Jianqiu Zhang 0002, Yidong Chen 0002, Yufei Huang 0001
ICASSP1
2010 An Iterated Conditional Modes solution for sparse Bayesian factor modeling of transcriptional regulatory networks
abstract
The problem of uncovering transcriptional regulation by transcription factors (TFs) based on microarray data is considered. A novel Bayesian sparse correlated rectified factor model (BSCRFM) coupled with its ICM solution is proposed. BSCRFM models the unknown TF protein level activity, the correlated regulations between TFs, and the sparse nature of TF regulated genes and it admits prior knowledge from existing database regarding TF regulated target genes. An efficient Iterated Conditional Modes (ICM) algorithm is developed, and a maximum a posterior (MAP) solution is calculated from multiple ICM results to avoid the local maximum problem, a context-specific transcriptional regulatory network specific to the experimental condition of the microarray data can then be obtained. The proposed model's ICM algorithm and MAP solution are evaluated on the simulated systems and results demonstrated the validity and effectiveness of the proposed approach. The proposed model is also applied to the breast cancer microarray data and a TF regulated network is obtained.
Jia Meng 0001, Jianqiu Zhang 0002, Yidong Chen 0002, Yufei Huang 0001
BIBM1
2009 Enrichment constrained time-dependent clustering analysis for finding meaningful temporal transcription modules
abstract
MOTIVATION: Clustering is a popular data exploration technique widely used in microarray data analysis. When dealing with time-series data, most conventional clustering algorithms, however, either use one-way clustering methods, which fail to consider the heterogeneity of temporary domain, or use two-way clustering methods that do not take into account the time dependency between samples, thus producing less informative results. Furthermore, enrichment analysis is often performed independent of and after clustering and such practice, though capable of revealing biological significant clusters, cannot guide the clustering to produce biologically significant result. RESULT: We present a new enrichment constrained framework (ECF) coupled with a time-dependent iterative signature algorithm (TDISA), which, by applying a sliding time window to incorporate the time dependency of samples and imposing an enrichment constraint to parameters of clustering, allows supervised identification of temporal transcription modules (TTMs) that are biologically meaningful. Rigorous mathematical definitions of TTM as well as the enrichment constraint framework are also provided that serve as objective functions for retrieving biologically significant modules. We applied the enrichment constrained time-dependent iterative signature algorithm (ECTDISA) to human gene expression time-series data of Kaposi's sarcoma-associated herpesvirus (KSHV) infection of human primary endothelial cells; the result not only confirms known biological facts, but also reveals new insight into the molecular mechanism of KSHV infection. AVAILABILITY: Data and Matlab code are available at http://engineering.utsa.edu/ approximately yfhuang/ECTDISA.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Jia Meng 0001, Shou-Jiang Gao, Yufei Huang 0001
Bioinform.1