VLDB 2026 Research / reviewers in the wild / expert
Xiaowo Wang
dblp:53/6281
· DBLP profile ↗
23ranked-venue papers
2as first author
16since 2021 · last 2026
0000-0003-2965-8036ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 21 · 2 first-author · 14 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mechanisms Under Shifts: Interpretable Clustering With Self-Improving Heterogeneous Causal GraphsabstractUnderstanding causal heterogeneity is crucial for building robust and interpretable learning systems that operate reliably under environmental shifts. However, existing methods lack causal awareness, with insufficient modeling of heterogeneity, confounding, and observational constraints, leading to poor interpretability and difficulty distinguishing true causal heterogeneity from spurious associations. We propose an unsupervised framework, HCL (Interpretable Causal Mechanism-Aware Clustering with Self-Improving Adaptive Heterogeneous Causal Structure Learning), that jointly infers latent clusters and their associated causal structures from mixed-type observational data without requiring temporal ordering, environment labels, interventions or other prior knowledge. HCL relaxes the homogeneity and sufficiency assumptions by introducing an equivalent representation that encodes both structural heterogeneity and confounding. It further develops a bi-directional iterative strategy to alternately refine causal clustering and structure learning, along with a self-supervised regularization that balances cross-cluster universal and specific mechanisms under shifts. Together, these components enable convergence toward interpretable, heterogeneous causal patterns. Theoretically, we show identifiability of heterogeneous causal structures under mild conditions. Empirically, HCL achieves superior performance in both clustering and structure learning tasks, and recovers biologically meaningful mechanisms in real-world single-cell perturbation and clinical intervention data, demonstrating its utility for discovering interpretable, mechanism-level causal heterogeneity. Qinghao Zhang, Xiaowo Wang |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Advancing genetic engineering with active learning: theory, implementations and potential opportunitiesabstractEmploying machine learning (ML) models to accelerate experimentation and uncover biological mechanisms has been a rising tendency in genetic engineering. However, effectively collecting data to enhance model accuracy and improve design remains challenging, especially when data quality is poor and validation resources are limited. Active learning (AL) addresses this by iteratively identifying promising candidates, thereby reducing experimental efforts while improving model performance. This review highlights how AL can assist scientists throughout the design-build-test-learn cycle, explore its various practical implementations, and discuss its potential through the integration of cross-domain expertise. In the age of genetic engineering revolutionized by data-driven ML models, AL presents an iterative framework that significantly enhances the functionalities of biomolecules and uncovers their intrinsic mechanisms, all while minimizing expenses and efforts. Qixiu Du, Benben Jiang, Xiaowo Wang |
Briefings Bioinform. | 4 |
| 2025 | Simulation-guided pan-cancer analysis identifies a novel regulator of CpG island hypermethylation heterogeneityabstractCpG island hypermethylation, a hallmark of cancer, exhibits substantial heterogeneity across tumors, presenting both opportunities and challenges for cancer diagnostics and therapeutics. While this heterogeneity offers potential for patient stratification to predict clinical outcomes and personalize treatments, it complicates the development of robust biomarkers for early detection. Understanding the mechanisms driving this heterogeneity is essential for advancing biomarker design. Here, simulation-based analyses demonstrate that tumor purity and the high prevalence of low epi-mutation samples significantly obscure the identification of negative, rather than positive, regulators of CpG island hypermethylation, limiting a comprehensive understanding of heterogeneity sources. By addressing these confounders, we identify impaired DNA methylation maintenance, as indicated by global hypomethylation levels, as the primary contributor to CpG island hypermethylation variability among known regulators. This finding is supported by integrative analyses of datasets from The Cancer Genome Atlas (TCGA) Pan-Cancer Atlas, Genomics of Drug Sensitivity in Cancer (GDSC1000) cancer cell lines, and epi-allele analyses of two independent whole-genome bisulfite sequencing cohorts, using a newly developed method, MeHist (https://github.com/vhang072/MeHist). Furthermore, we assess widely used hypermethylation biomarkers across ten cancer types and find that 65 out of 246 (26.4%) are significantly influenced by impaired methylation maintenance. Incorporating hypomethylation and hypermethylation markers improves the robustness of cancer detection, as validated across multiple plasma cell-free DNA datasets. In summary, our findings highlight the value of simulation-guided integrative analysis in mitigating confounding effects and identify impaired DNA methylation maintenance as a key regulator of CpG island hypermethylation heterogeneity. Xianglin Zhang, Wei Zhang 0241, Xiuhong Lyu, Haoran Pan, Tianwei Jia, Xiaowo Wang, Haiyang Guo |
Briefings Bioinform. | 8 |
| 2025 | esMPRA: an easy-to-use systematic pipeline for MPRA experiment quality control and data analysisabstractMOTIVATION: Massively Parallel Reporter Assays (MPRAs) have emerged as pivotal tools for systematically profiling cis-regulatory element activity, playing critical roles in deciphering gene regulation mechanisms and synthetic regulatory element engineering. However, MPRA experiments involve multi-step library processing procedures coupled with high-throughput sequencing. Operational errors during these complex workflows can lead to substantial resource depletion and experimental delays. Thus robust and user-friendly quality control methods are essential to minimize experimental failures and ensure reproducibility between replicates. RESULTS: Here, we present esMPRA, an integrated quality control and analysis pipeline designed for MPRA experiments. Building on our experience in MPRA and its derivative techniques, coupled with systematic analysis of public MPRA datasets, we established standardized quality control metrics and developed a stepwise quality monitoring framework. esMPRA generates stage-specific diagnostic reports and provides experimental recommendations to avoid potential risks throughout the workflow. Designed for maximal accessibility, esMPRA features a one-line command-line interface and requires minimal bioinformatics expertise. Beyond quality assessment, the pipeline delivers processed data outputs, comprehensive analysis reports, and interface files compatible with downstream analyses, establishing an end-to-end solution for MPRA experimentation. AVAILABILITY AND IMPLEMENTATION: esMPRA is released as an open-source software under the MIT license. The source code for esMPRA is available on Zenodo (DOI: 10.5281/zenodo.15362711) and GitHub (https://github.com/WangLabTHU/esMPRA/) for Linux, macOS, and Windows and is available via PyPI as esMPRA. Data for testing and reference is available via Zenodo repository at https://zenodo.org/records/15034449. Jiaqi Li 0025, Xiaowo Wang |
Bioinform. | 4 |
| 2024 | Discovering and Overcoming the Bias in Neoantigen Identification by Unified Machine Learning Models
Ziting Zhang, Wenxu Wu, Lei Wei 0009, Xiaowo Wang |
RECOMB | 4 |
| 2024 | GPro: generative AI-empowered toolkit for promoter designabstractMOTIVATION: Promoters with desirable properties are crucial in biotechnological applications. Generative AI (GenAI) has demonstrated potential in creating novel synthetic promoters with significantly enhanced functionality. However, these methods' reliance on various programming frameworks and specific task-oriented contexts limits their flexibilities. Overcoming these limitations is essential for researchers to fully leverage the power of GenAI to design promoters for their tasks. RESULTS: Here, we introduce GPro (Generative AI-empowered toolkit for promoter design), a user-friendly toolkit that integrates a collection of cutting-edge GenAI-empowered approaches for promoter design. This toolkit provides a standardized pipeline covering essential promoter design processes, including training, optimization, and evaluation. Several detailed demos are provided to reproduce state-of-the-art promoter design pipelines. GPro's user-friendly interface makes it accessible to a wide range of users including non-AI experts. It also offers a variety of optional algorithms for each design process, and gives users the flexibility to compare methods and create customized pipelines. AVAILABILITY AND IMPLEMENTATION: GPro is released as an open-source software under the MIT license. The source code for GPro is available on GitHub for Linux, macOS, and Windows: https://github.com/WangLabTHU/GPro, and is available for download via Zenodo repository at https://zenodo.org/doi/10.5281/zenodo.10681733. Qixiu Du, Xiaowo Wang |
Bioinform. | 6 |
| 2024 | Unveil cis-acting combinatorial mRNA motifs by interpreting deep neural networkabstractSUMMARY: Cis-acting mRNA elements play a key role in the regulation of mRNA stability and translation efficiency. Revealing the interactions of these elements and their impact plays a crucial role in understanding the regulation of the mRNA translation process, which supports the development of mRNA-based medicine or vaccines. Deep neural networks (DNN) can learn complex cis-regulatory codes from RNA sequences. However, extracting these cis-regulatory codes efficiently from DNN remains a significant challenge. Here, we propose a method based on our toolkit NeuronMotif and motif mutagenesis, which not only enables the discovery of diverse and high-quality motifs but also efficiently reveals motif interactions. By interpreting deep-learning models, we have discovered several crucial motifs that impact mRNA translation efficiency and stability, as well as some unknown motifs or motif syntax, offering novel insights for biologists. Furthermore, we note that it is challenging to enrich motif syntax in datasets composed of randomly generated sequences, and they may not contain sufficient biological signals. AVAILABILITY AND IMPLEMENTATION: The source code and data used to produce the results and analyses presented in this manuscript are available from GitHub (https://github.com/WangLabTHU/combmotif). Xiaocheng Zeng, Qixiu Du, Jiaqi Li 0025, Xiaowo Wang |
Bioinform. | 6 |
| 2024 | Weakly Supervised Causal Discovery Based on Fuzzy Knowledge and Complex Data ComplementarityabstractCausal discovery based on observational data is important for deciphering the causal mechanism behind complex systems. However, the effectiveness of existing causal discovery methods is limited due to inferior prior knowledge, domain inconsistencies, and the challenges of high-dimensional datasets with small sample sizes. To address this gap, we propose a novel weakly supervised fuzzy knowledge and data co-driven causal discovery method named KEEL. KEEL introduces a fuzzy causal knowledge schema to encapsulate diverse types of fuzzy knowledge, and forms corresponding weakened constraints. This schema not only lessens the dependency on expertise but also allows various types of limited and error-prone fuzzy knowledge to guide causal discovery. It can enhance the generalization and robustness of causal discovery, especially in high-dimensional and small-sample scenarios. In addition, we integrate the extended linear causal model into KEEL for dealing with the multi-distribution and incomplete data. Extensive experiments with different datasets demonstrate the superiority of KEEL over several state-of-the-art methods in accuracy, robustness and efficiency. The effectiveness of KEEL is also verified in limited real protein signal transduction process data, with the better performance than benchmark methods. In summary, KEEL is effective to tackle the causal discovery tasks with higher accuracy while alleviating the requirement for extensive domain expertise. Wei Zhang 0241, Qinghao Zhang, Xuegong Zhang, Xiaowo Wang |
IEEE Trans. Fuzzy Syst. | 5 |
| 2022 | Evaluating methylation of human ribosomal DNA at each CpG site reveals its utility for cancer detection using cell-free DNAabstractRibosomal deoxyribonucleic acid (DNA) (rDNA) repeats are tandemly located on five acrocentric chromosomes with up to hundreds of copies in the human genome. DNA methylation, the most well-studied epigenetic mechanism, has been characterized for most genomic regions across various biological contexts. However, rDNA methylation patterns remain largely unexplored due to the repetitive structure. In this study, we designed a specific mapping strategy to investigate rDNA methylation patterns at each CpG site across various physiological and pathological processes. We found that CpG sites on rDNA could be categorized into two types. One is within or adjacent to transcribed regions; the other is distal to transcribed regions. The former shows highly variable methylation levels across samples, while the latter shows stable high methylation levels in normal tissues but severe hypomethylation in tumors. We further showed that rDNA methylation profiles in plasma cell-free DNA could be used as a biomarker for cancer detection. It shows good performances on public datasets, including colorectal cancer [area under the curve (AUC) = 0.85], lung cancer (AUC = 0.84), hepatocellular carcinoma (AUC = 0.91) and in-house generated hepatocellular carcinoma dataset (AUC = 0.96) even at low genome coverage (<1×). Taken together, these findings broaden our understanding of rDNA regulation and suggest the potential utility of rDNA methylation features as disease biomarkers. Xianglin Zhang, Bixi Zhong, Lei Wei 0009, Jiaqi Li 0025, Wei Zhang 0241, Huan Fang 0003, Yanda Li, Yinying Lu, Xiaowo Wang |
Briefings Bioinform. | 10 |
| 2022 | ARIC: accurate and robust inference of cell type proportions from bulk gene expression or DNA methylation dataabstractQuantifying cell proportions, especially for rare cell types in some scenarios, is of great value in tracking signals associated with certain phenotypes or diseases. Although some methods have been proposed to infer cell proportions from multicomponent bulk data, they are substantially less effective for estimating the proportions of rare cell types which are highly sensitive to feature outliers and collinearity. Here we proposed a new deconvolution algorithm named ARIC to estimate cell type proportions from gene expression or DNA methylation data. ARIC employs a novel two-step marker selection strategy, including collinear feature elimination based on the component-wise condition number and adaptive removal of outlier markers. This strategy can systematically obtain effective markers for weighted $\upsilon$-support vector regression to ensure a robust and precise rare proportion prediction. We showed that ARIC can accurately estimate fractions in both DNA methylation and gene expression data from different experiments. We further applied ARIC to the survival prediction of ovarian cancer and the condition monitoring of chronic kidney disease, and the results demonstrate the high accuracy and robustness as well as clinical potentials of ARIC. Taken together, ARIC is a promising tool to solve the deconvolution problem of bulk data where rare components are of vital importance. Wei Zhang 0241, Rong Qiao, Bixi Zhong, Xianglin Zhang, Jin Gu, Xuegong Zhang, Lei Wei 0009, Xiaowo Wang |
Briefings Bioinform. | 9 |
| 2022 | MeConcord: a new metric to quantitatively characterize DNA methylation heterogeneity across reads and CpG sitesabstractMOTIVATION: Intermediately methylated regions occupy a significant fraction of the human genome and are closely associated with epigenetic regulations or cell-type deconvolution of bulk data. However, these regions show distinct methylation patterns, corresponding to different biological mechanisms. Although there have been some metrics developed for investigating these regions, the high noise sensitivity limits the utility for distinguishing distinct methylation patterns. RESULTS: We proposed a method named MeConcord to measure local methylation concordance across reads and CpG sites, respectively. MeConcord showed the most stable performance in distinguishing distinct methylation patterns ('identical', 'uniform' and 'disordered') compared with other metrics. Applying MeConcord to the whole genome data across 25 cell lines or primary cells or tissues, we found that distinct methylation patterns were associated with different genomic characteristics, such as CTCF binding or imprinted genes. Further, we showed the differences of CpG island hypermethylation patterns between senescence and tumorigenesis by using MeConcord. MeConcord is a powerful method to study local read-level methylation patterns for both the whole genome and specific regions of interest. AVAILABILITY AND IMPLEMENTATION: MeConcord is available at https://github.com/WangLabTHU/MeConcord. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xianglin Zhang, Xiaowo Wang |
Bioinform. | 2 |
| 2022 | DeSP: a systematic DNA storage error simulation pipelineabstractBACKGROUND: Using DNA as a storage medium is appealing due to the information density and longevity of DNA, especially in the era of data explosion. A significant challenge in the DNA data storage area is to deal with the noises introduced in the channel and control the trade-off between the redundancy of error correction codes and the information storage density. As running DNA data storage experiments in vitro is still expensive and time-consuming, a simulation model is needed to systematically optimize the redundancy to combat the channel's particular noise structure. RESULTS: Here, we present DeSP, a systematic DNA storage error Simulation Pipeline, which simulates the errors generated from all DNA storage stages and systematically guides the optimization of encoding redundancy. It covers both the sequence lost and the within-sequence errors in the particular context of the data storage channel. With this model, we explained how errors are generated and passed through different stages to form final sequencing results, analyzed the influence of error rate and sampling depth to final error rates, and demonstrated how to systemically optimize redundancy design in silico with the simulation model. These error simulation results are consistent with the in vitro experiments. CONCLUSIONS: DeSP implemented in Python is freely available on Github ( https://github.com/WangLabTHU/DeSP ). It is a flexible framework for systematic error simulation in DNA storage and can be adapted to a wide range of experiment pipelines. Lekang Yuan, Xiaowo Wang |
BMC Bioinform. | 4 |
| 2022 | Correction to: DeSP: a systematic DNA storage error simulation pipeline
Lekang Yuan, Xiaowo Wang |
BMC Bioinform. | 4 |
| 2021 | DISMIR: Deep learning-based noninvasive cancer detection by integrating DNA sequence and methylation information of individual cell-free DNA readsabstractDetecting cancer signals in cell-free DNA (cfDNA) high-throughput sequencing data is emerging as a novel noninvasive cancer detection method. Due to the high cost of sequencing, it is crucial to make robust and precise predictions with low-depth cfDNA sequencing data. Here we propose a novel approach named DISMIR, which can provide ultrasensitive and robust cancer detection by integrating DNA sequence and methylation information in plasma cfDNA whole-genome bisulfite sequencing (WGBS) data. DISMIR introduces a new feature termed as 'switching region' to define cancer-specific differentially methylated regions, which can enrich the cancer-related signal at read-resolution. DISMIR applies a deep learning model to predict the source of every single read based on its DNA sequence and methylation state and then predicts the risk that the plasma donor is suffering from cancer. DISMIR exhibited high accuracy and robustness on hepatocellular carcinoma detection by plasma cfDNA WGBS data even at ultralow sequencing depths. Further analysis showed that DISMIR tends to be insensitive to alterations of single CpG sites' methylation states, which suggests DISMIR could resist to technical noise of WGBS. All these results showed DISMIR with the potential to be a precise and robust method for low-cost early cancer detection. Jiaqi Li 0025, Lei Wei 0009, Xianglin Zhang, Wei Zhang 0241, Bixi Zhong, Hairong Lv, Xiaowo Wang |
Briefings Bioinform. | 9 |
| 2021 | CellTracker: an automated toolbox for single-cell segmentation and tracking of time-lapse microscopy imagesabstractSUMMARY: Recent advances of long-term time-lapse microscopy have made it easy for researchers to quantify cell behavior and molecular dynamics at single-cell resolution. However, the lack of easy-to-use software tools optimized for customized research is still a major challenge for quantitatively understanding biological processes through microscopy images. Here, we present CellTracker, a highly integrated graphical user interface software, for automated cell segmentation and tracking of time-lapse microscopy images. It covers essential steps in image analysis including project management, image pre-processing, cell segmentation, cell tracking, manually correction and statistical analysis such as the quantification of cell size and fluorescence intensity, etc. Furthermore, CellTracker provides an annotation tool and supports model training from scratch, thus proposing a flexible and scalable solution for customized dataset analysis. AVAILABILITY AND IMPLEMENTATION: CellTracker is an open-source software under the GPL-3.0 license. It is implemented in Python and provides an easy-to-use graphical user interface. The source code, instruction manual and demos can be found at https://github.com/WangLabTHU/CellTracker. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tao Hu 0020, Shixiong Xu, Lei Wei 0009, Xuegong Zhang, Xiaowo Wang |
Bioinform. | 5 |
| 2021 | cfDNApipe: a comprehensive quality control and analysis pipeline for cell-free DNA high-throughput sequencing dataabstractMOTIVATION: Cell-free DNA (cfDNA) is gaining substantial attention from both biological and clinical fields as a promising marker for liquid biopsy. Many aspects of disease-related features have been discovered from cfDNA high-throughput sequencing (HTS) data. However, there is still a lack of integrative and systematic tools for cfDNA HTS data analysis and quality control (QC). RESULTS: Here, we propose cfDNApipe, an easy-to-use and systematic python package for cfDNA whole-genome sequencing (WGS) and whole-genome bisulfite sequencing (WGBS) data analysis. It covers the entire analysis pipeline for the cfDNA data, including raw sequencing data processing, QC and sophisticated statistical analysis such as detecting copy number variations (CNVs), differentially methylated regions and DNA fragment size alterations. cfDNApipe provides one-command-line-execution pipelines and flexible application programming interfaces for customized analysis. AVAILABILITY AND IMPLEMENTATION: https://xwanglabthu.github.io/cfDNApipe/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Wei Zhang 0241, Lei Wei 0009, Bixi Zhong, Jiaqi Li 0025, Shuying He, Juhong Liu, Hairong Lv, Xiaowo Wang |
Bioinform. | 11 |
| 2018 | DEsingle for detecting three types of differential expression in single-cell RNA-seq dataabstractSummary: The excessive amount of zeros in single-cell RNA-seq (scRNA-seq) data includes 'real' zeros due to the on-off nature of gene transcription in single cells and 'dropout' zeros due to technical reasons. Existing differential expression (DE) analysis methods cannot distinguish these two types of zeros. We developed an R package DEsingle which employed Zero-Inflated Negative Binomial model to estimate the proportion of real and dropout zeros and to define and detect three types of DE genes in scRNA-seq data with higher accuracy. Availability and implementation: The R package DEsingle is freely available at Bioconductor (https://bioconductor.org/packages/DEsingle). Supplementary information: Supplementary data are available at Bioinformatics online. Zhun Miao, Xiaowo Wang, Xuegong Zhang |
Bioinform. | 3 |
| 2018 | esATAC: an easy-to-use systematic pipeline for ATAC-seq data analysisabstractSummary: ATAC-seq is rapidly emerging as one of the major experimental approaches to probe chromatin accessibility genome-wide. Here, we present 'esATAC', a highly integrated easy-to-use R/Bioconductor package, for systematic ATAC-seq data analysis. It covers essential steps for full analyzing procedure, including raw data processing, quality control and downstream statistical analysis such as peak calling, enrichment analysis and transcription factor footprinting. esATAC supports one command line execution for preset pipelines and provides flexible interfaces for building customized pipelines. Availability and implementation: esATAC package is open source under the GPL-3.0 license. It is implemented in R and C++. Source code and binaries for Linux, MAC OS X and Windows are available through Bioconductor (https://www.bioconductor.org/packages/release/bioc/html/esATAC.html). Supplementary information: Supplementary data are available at Bioinformatics online. Wei Zhang 0241, Huan Fang 0003, Yanda Li, Xiaowo Wang |
Bioinform. | 5 |
| 2015 | MICC: an R package for identifying chromatin interactions from ChIA-PET dataabstractUNLABELLED: ChIA-PET is rapidly emerging as an important experimental approach to detect chromatin long-range interactions at high resolution. Here, we present Model based Interaction Calling from ChIA-PET data (MICC), an easy-to-use R package to detect chromatin interactions from ChIA-PET sequencing data. By applying a Bayesian mixture model to systematically remove random ligation and random collision noise, MICC could identify chromatin interactions with a significantly higher sensitivity than existing methods at the same false discovery rate. AVAILABILITY AND IMPLEMENTATION: http://bioinfo.au.tsinghua.edu.cn/member/xwwang/MICCusage CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Michael Q. Zhang, Xiaowo Wang |
Bioinform. | 3 |
| 2015 | CRISPR-ERA: a comprehensive design tool for CRISPR-mediated gene editing, repression and activationabstractUNLABELLED: The CRISPR/Cas9 system was recently developed as a powerful and flexible technology for targeted genome engineering, including genome editing (altering the genetic sequence) and gene regulation (without altering the genetic sequence). These applications require the design of single guide RNAs (sgRNAs) that are efficient and specific. However, this remains challenging, as it requires the consideration of many criteria. Several sgRNA design tools have been developed for gene editing, but currently there is no tool for the design of sgRNAs for gene regulation. With accumulating experimental data on the use of CRISPR/Cas9 for gene editing and regulation, we implement a comprehensive computational tool based on a set of sgRNA design rules summarized from these published reports. We report a genome-wide sgRNA design tool and provide an online website for predicting sgRNAs that are efficient and specific. We name the tool CRISPR-ERA, for clustered regularly interspaced short palindromic repeat-mediated editing, repression, and activation (ERA). AVAILABILITY AND IMPLEMENTATION: http://CRISPR-ERA.stanford.edu. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Antonia Dominguez, Yanda Li, Xiaowo Wang, Lei S. Qi |
Bioinform. | 5 |
| 2010 | DEGseq: an R package for identifying differentially expressed genes from RNA-seq dataabstractAbstract Summary: High-throughput RNA sequencing (RNA-seq) is rapidly emerging as a major quantitative transcriptome profiling platform. Here, we present DEGseq, an R package to identify differentially expressed genes or isoforms for RNA-seq data from different samples. In this package, we integrated three existing methods, and introduced two novel methods based on MA-plot to detect and visualize gene expression difference. Availability: The R package and a quick-start vignette is available at http://bioinfo.au.tsinghua.edu.cn/software/degseq Contact: [email protected]; [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Likun Wang 0003, Zhixing Feng, Xi Wang 0002, Xiaowo Wang, Xuegong Zhang |
Bioinform. | 4 |
| 2008 | Identification of phylogenetically conserved microRNA cis-regulatory elements across 12 Drosophila speciesabstractMOTIVATION: MicroRNAs are a class of endogenous small RNAs that play regulatory roles. Intergenic miRNAs are believed to be transcribed independently, but the transcriptional control of these crucial regulators is still poorly understood. RESULTS: In this work, phylogenetic footprinting is used to identify conserved cis-regulatory elements (CCEs) surrounding intergenic miRNAs in Drosophila. With a two-step strategy that takes advantage of both alignment-based and motif-based methods, we identified CCEs that are conserved across the 12 fly species. When compared with TRANSFAC database, these CCEs are significantly enriched in known transcription factor binding sites (TFBSs). Moreover, several TFs that play essential roles in Drosophila development (e.g. Adf-1, Abd-B, Sd, Prd, Ubx, Zen and En) are found to be preferentially regulating the miRNA genes. Further analysis revealed many over-represented cis-regulatory modules (CRMs) composed of multiple known TFBSs, motif pairs with significant distance constraints and a number of novel motifs, many of which preferentially occur near the transcription start site of protein-coding genes. Additionally, a number of putative miRNA-TF regulatory feedback loops were also detected. AVAILABILITY: Supplementary Material and the Perl scripts performing two-step phylogenetic footprinting are available at http://bioinfo.au.tsinghua.edu.cn/member/xwwang/mircisreg Xiaowo Wang, Jin Gu, Michael Q. Zhang, Yanda Li |
Bioinform. | 1 |
| 2005 | MicroRNA identification based on sequence and structure alignmentabstractMOTIVATION: MicroRNAs (miRNA) are approximately 22 nt long non-coding RNAs that are derived from larger hairpin RNA precursors and play important regulatory roles in both animals and plants. The short length of the miRNA sequences and relatively low conservation of pre-miRNA sequences restrict the conventional sequence-alignment-based methods to finding only relatively close homologs. On the other hand, it has been reported that miRNA genes are more conserved in the secondary structure rather than in primary sequences. Therefore, secondary structural features should be more fully exploited in the homologue search for new miRNA genes. RESULTS: In this paper, we present a novel genome-wide computational approach to detect miRNAs in animals based on both sequence and structure alignment. Experiments show this approach has higher sensitivity and comparable specificity than other reported homologue searching methods. We applied this method on Anopheles gambiae and detected 59 new miRNA genes. AVAILABILITY: This program is available at http://bioinfo.au.tsinghua.edu.cn/miralign. SUPPLEMENTARY INFORMATION: Supplementary information is available at http://bioinfo.au.tsinghua.edu.cn/miralign/supplementary.htm. Xiaowo Wang, Jing Zhang 0010, Jin Gu, Xuegong Zhang, Yanda Li |
Bioinform. | 1 |