VLDB 2026 Research / reviewers in the wild / expert
Lihua Julie Zhu
dblp:34/8252
· DBLP profile ↗
3ranked-venue papers
1as first author
0since 2021 · last 2014
0000-0001-7416-0590ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Bioinformatics and computational biology · 100% |
Topics — the 2 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › RNA biology › RNA processing
polyadenylation site prediction |
0.2 | 1 | 2013 | Accurate identification of polyadenylation sites from 3′ end deep sequencing using a naïve Bayes classifier · Bioinform. 2013 |
Bioinformatics and computational biology
sequence analysis |
0.2 | 1 | 2013 | Accurate identification of polyadenylation sites from 3′ end deep sequencing using a naïve Bayes classifier · Bioinform. 2013 |
Methods — techniques the papers use, named apart from their topics
naïve bayes classifier · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Accurate identification of polyadenylation sites from 3′ end deep sequencing using a naïve Bayes classifierabstractVol. 29, No. 20, 2013, pp. 2564–2571 doi:10.1093/bioinformatics/btt446 The publishers regret that the corresponding author details for this article should appear as follows: Contact:[email protected]; [email protected] Sarah Sheppard, Nathan D. Lawson, Lihua Julie Zhu |
Bioinform. | 3 |
| 2013 | Accurate identification of polyadenylation sites from 3′ end deep sequencing using a naïve Bayes classifierabstractMOTIVATION: 3' end processing is important for transcription termination, mRNA stability and regulation of gene expression. To identify 3' ends, most techniques use an oligo-dT primer to construct deep sequencing libraries. However, this approach can lead to identification of artifactual polyadenylation sites due to internal priming in homopolymeric stretches of adenines. Although heuristic filters have been applied in these cases, they typically result in a high proportion of both false-positive and -negative classifications. Therefore, there is a need to develop improved algorithms to better identify mis-priming events in oligo-dT primed sequences. RESULTS: By analyzing sequence features flanking 3' ends derived from oligo-dT-based sequencing, we developed a naïve Bayes classifier to classify them as true or false/internally primed. The resulting algorithm is highly accurate, outperforms previous heuristic filters and facilitates identification of novel polyadenylation sites. Sarah Sheppard, Nathan D. Lawson, Lihua Julie Zhu |
Bioinform. | 3 |
| 2010 | ChIPpeakAnno: a Bioconductor package to annotate ChIP-seq and ChIP-chip dataabstractBACKGROUND: Chromatin immunoprecipitation (ChIP) followed by high-throughput sequencing (ChIP-seq) or ChIP followed by genome tiling array analysis (ChIP-chip) have become standard technologies for genome-wide identification of DNA-binding protein target sites. A number of algorithms have been developed in parallel that allow identification of binding sites from ChIP-seq or ChIP-chip datasets and subsequent visualization in the University of California Santa Cruz (UCSC) Genome Browser as custom annotation tracks. However, summarizing these tracks can be a daunting task, particularly if there are a large number of binding sites or the binding sites are distributed widely across the genome. RESULTS: We have developed ChIPpeakAnno as a Bioconductor package within the statistical programming environment R to facilitate batch annotation of enriched peaks identified from ChIP-seq, ChIP-chip, cap analysis of gene expression (CAGE) or any experiments resulting in a large number of enriched genomic regions. The binding sites annotated with ChIPpeakAnno can be viewed easily as a table, a pie chart or plotted in histogram form, i.e., the distribution of distances to the nearest genes for each set of peaks. In addition, we have implemented functionalities for determining the significance of overlap between replicates or binding sites among transcription factors within a complex, and for drawing Venn diagrams to visualize the extent of the overlap between replicates. Furthermore, the package includes functionalities to retrieve sequences flanking putative binding sites for PCR amplification, cloning, or motif discovery, and to identify Gene Ontology (GO) terms associated with adjacent genes. CONCLUSIONS: ChIPpeakAnno enables batch annotation of the binding sites identified from ChIP-seq, ChIP-chip, CAGE or any technology that results in a large number of enriched genomic regions within the statistical programming environment R. Allowing users to pass their own annotation data such as a different Chromatin immunoprecipitation (ChIP) preparation and a dataset from literature, or existing annotation packages, such as GenomicFeatures and BSgenome, provides flexibility. Tight integration to the biomaRt package enables up-to-date annotation retrieval from the BioMart database. Lihua Julie Zhu, Claude Gazin, Nathan D. Lawson, Hervé Pagès, Simon M. Lin, David S. Lapointe, Michael R. Green |
BMC Bioinform. | 1 |