EDBT 2026 Demo / reviewers in the wild / expert
Jun Yu 0004
dblp:50/5754-4
· DBLP profile ↗
15ranked-venue papers
0as first author
7since 2021 · last 2024
0000-0002-2702-055XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Developing H5N1 Avian Influenza Mutation and Evolution Feature Analysis and Web ServiceabstractThe continuous mutation and evolution of H5N1 may cause a wide spread of the epidemic, however, the mechanism of H5N1 mutation and evolution are still unclear. Therefore, further mining the overall evolutionary pattern of H5N1 sequences, studying the impact of global factors such as migratory bird flyways on H5N1 evolution, and constructing a corresponding web service are crucial for preventing potential future H5N1 epidemics. To this end, we firstly analyzed the within-host mutation features of H5N1 sequences and optimized the phylogenetic model. Secondly, we employed the optimized model to construct the phylogenetic trees and performed H5N1 evolution group analysis. Thirdly, we not only analyzed the correlation between different evolution groups of H5N1 and global migratory bird flyways, but also investigated the impact of the El Niño climate phenomenon on the evolution of H5N1. Finally, we built up a web service for H5N1 mutation and evolution analysis and visualization. Our main results include: (1) We found that the global H5N1 can be categorized into three major evolution groups at different time in various geographic regions. (2) We not only found that all three major groups of HA fragments were related to the East Asian -Australasian flyway of migratory birds, but also demonstrated the correlation and causality between the Oceanic Niño Index and the alternation of old and new evolutionary groups of H5N1. (3) VP-H5N1 provides a fast and easy-to-use web service platform for online analysis and visualization of H5N1 evolution. Ming Xiao 0002, Qichen Shang, Qiaozhen Zhang, Jun Yu 0004, Le Zhang 0004 |
BIBM | 5 |
| 2022 | Position-Defined CpG Islands Provide Complete Co-methylation Indexing for Human Genes
Ming Xiao 0002, Ruiying Yin, Pengbo Gao, Jun Yu 0004, Fubo Ma, Zichun Dai, Le Zhang 0004 |
ICIC (2) | 4 |
| 2021 | A fine-scale map of genome-wide recombination in divergent Escherichia coli populationabstractRecombination is one of the most important molecular mechanisms of prokaryotic genome evolution, but its exact roles are still in debate. Here we try to infer genome-wide recombination within a species, utilizing a dataset of 149 complete genomes of Escherichia coli from diverse animal hosts and geographic origins, including 45 in-house sequenced with the single-molecular real-time platform. Two major clades identified based on physiological, clinical and ecological characteristics form distinct genetic lineages based on scarcity of interclade gene exchanges. By defining gene-based syntenies for genomic segments within and between the two clades, we build a fine-scale recombination map for this representative global E. coli population. The map suggests extensive within-clade recombination that often breaks physical linkages among individual genes but seldom interrupts the structure of genome organizational frameworks as well as primary metabolic portfolios supported by the framework integrity, possibly due to strong natural selection for both physiological compatibility and ecological fitness. In contrast, the between-clade recombination declines drastically when phylogenetic distance increases to the extent where a 10-fold reduction can be observed, establishing a firm genetic barrier between clades. Our empirical data suggest a critical role for such recombination events in the early stage of speciation where recombination rate is associated with phylogenetic distance in addition to sequence and gene variations. The extensive intraclade recombination binds sister strains into a quasisexual group and optimizes genes or alleles to streamline physiological activities, whereas the sharply declined interclade recombination split the population into clades adaptive to divergent ecological niches. Yu Kang 0004, Lina Yuan, Yanan Chu, Xinmiao Jia, Qin Ma 0003, Jian Wang 0065, Jing-Fa Xiao, Songnian Hu, Zhancheng Gao, Jun Yu 0004 |
Briefings Bioinform. | 14 |
| 2021 | Roles of host small RNAs in the evolution and host tropism of coronavirusesabstractHuman coronaviruses (CoVs) can cause respiratory infection epidemics that sometimes expand into globally relevant pandemics. All human CoVs have sister strains isolated from animal hosts and seem to have an animal origin, yet the process of host jumping is largely unknown. RNA interference (RNAi) is an ancient mechanism in many eukaryotes to defend against viral infections through the hybridization of host endogenous small RNAs (miRNAs) with target sites in invading RNAs. Here, we developed a method to identify potential RNAi-sensitive sites in the viral genome and discovered that human-adapted coronavirus strains had deleted some of their sites targeted by miRNAs in human lungs when compared to their close zoonic relatives. We further confirmed using a phylogenetic analysis that the loss of RNAi-sensitive target sites could be a major driver of the host-jumping process, and adaptive mutations that lead to the loss-of-target might be as simple as point mutation. Up-to-date genomic data of severe acute respiratory syndrome coronavirus 2 and Middle-East respiratory syndromes-CoV strains demonstrate that the stress from host miRNA milieus sustained even after their epidemics in humans. Thus, this study illustrates a new mechanism about coronavirus to explain its host-jumping process and provides a novel avenue for pathogenesis research, epidemiological modeling, and development of drugs and vaccines against coronavirus, taking into consideration these findings. Qingren Meng, Yanan Chu, Changjun Shao, Jing Chen 0076, Jian Wang 0065, Zhancheng Gao, Jun Yu 0004, Yu Kang 0004 |
Briefings Bioinform. | 7 |
| 2021 | CpG-island-based annotation and analysis of human housekeeping genesabstractBy reviewing previous CpG-related studies, we consider that the transcription regulation of about half of the human genes, mostly housekeeping (HK) genes, involves CpG islands (CGIs), their methylation states, CpG spacing and other chromosomal parameters. However, the precise CGI definition and positioning of CGIs within gene structures, as well as specific CGI-associated regulatory mechanisms, all remain to be explained at individual gene and gene-family levels, together with consideration of species and lineage specificity. Although previous studies have already classified CGIs into high-CpG (HCGI), intermediate-CpG (ICGI) and low-CpG (LCGI) densities based on CpG density variation, the correlation between CGI density and gene expression regulation, such as co-regulation of CGIs and TATA box on HK genes, remains to be elucidated. First, this study introduces such a problem-solving protocol for human-genome annotation, which is based on a combination of GTEx, JBLA and Gene Ontology (GO) analysis. Next, we discuss why CGI-associated genes are most likely regulated by HCGI and tend to be HK genes; the HCGI/TATA± and LCGI/TATA± combinations show different GO enrichment, whereas the ICGI/TATA± combination is less characteristic based on GO enrichment analysis. Finally, we demonstrate that Hadoop MapReduce-based MR-JBLA algorithm is more efficient than the original JBLA in k-mer counting and CGI-associated gene analysis. Le Zhang 0004, Zichun Dai, Jun Yu 0004, Ming Xiao 0002 |
Briefings Bioinform. | 3 |
| 2021 | CancerEMC: frontline non-invasive cancer screening from circulating protein biomarkers and mutations in cell-free DNAabstractMOTIVATION: The early detection of cancer through accessible blood tests can foster early patient interventions. Although there are developments in cancer detection from cell-free DNA (cfDNA), its accuracy remains speculative. Given its central importance with broad impacts, we aspire to address the challenge. METHOD: A bagging Ensemble Meta Classifier (CancerEMC) is proposed for early cancer detection based on circulating protein biomarkers and mutations in cfDNA from blood. CancerEMC is generally designed for both binary cancer detection and multi-class cancer type localization. It can address the class imbalance problem in multi-analyte blood test data based on robust oversampling and adaptive synthesis techniques. RESULTS: Based on the clinical blood test data, we observe that the proposed CancerEMC has outperformed other algorithms and state-of-the-arts studies (including CancerSEEK) for cancer detection. The results reveal that our proposed method (i.e. CancerEMC) can achieve the best performance result for both binary cancer classification with 99.17% accuracy (AUC = 0.999) and localized multiple cancer detection with 74.12% accuracy (AUC = 0.938). Addressing the data imbalance issue with oversampling techniques, the accuracy can be increased to 91.50% (AUC = 0.992), where the state-of-the-art method can only be estimated at 69.64% (AUC = 0.921). Similar results can also be observed on independent and isolated testing data. AVAILABILITY: https://github.com/saifurcubd/Cancer-Detection. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Saifur Rahaman, Xiangtao Li, Jun Yu 0004, Ka-Chun Wong |
Bioinform. | 3 |
| 2021 | 2019nCoVAS: Developing the Web Service for Epidemic Transmission Prediction, Genome Analysis, and Psychological Stress Assessment for 2019-nCoVabstractSince the COVID-19 epidemic is still expanding around the world and poses a serious threat to human life and health, it is necessary for us to carry out epidemic transmission prediction, whole genome sequence analysis, and public psychological stress assessment for 2019-nCoV. However, transmission prediction models are insufficiently accurate and genome sequence characteristics are not clear, and it is difficult to dynamically assess the public psychological stress state under the 2019-nCoV epidemic. Therefore, this study develops a 2019nCoVAS web service (http://www.combio-lezhang.online/2019ncov/home.html) that not only offers online epidemic transmission prediction and lineage-associated underrepresented permutation (LAUP) analysis services to investigate the spreading trends and genome sequence characteristics, but also provides psychological stress assessments based on such an emotional dictionary that we built for 2019-nCoV. Finally, we discuss the shortcomings and further study of the 2019nCoVAS web service. Ming Xiao 0002, Guangdi Liu, Jianghang Xie, Zichun Dai, Zihao Wei, Ziyao Ren, Jun Yu 0004, Le Zhang 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2020 | CGIDLA: Developing the Web Server for CpG Island Related Density and LAUPs (Lineage-Associated Underrepresented Permutations) StudyabstractIt is well known that CpG island plays an important role in gene methylation. Since CpG island is closely related to human genetic characteristics such as TATA-box, tissue expression specificity, and LAUPs (Lineage-associated Underrepresented Permutations), it is important to investigate the sequence specificity of CpG island as well as the potential genetic characteristics related to CpG island to further understand the methylation related regulation mechanism. Therefore, this study develops such an online service website for CpG island related density and LAUPs analysis (CGIDLA, www.combio-lezhang.online/cgidla/index.html), that not only can investigate the relationship among the CpG island density, TATA-box feature, and expression breadth of human genes, but also deposit LAUPs of 32 representative species to help molecular biologists investigate the relationship between CpG island and LUAPs. Moreover, CGIDLA provides the source code download service and the related LAUPs counting functions. Ming Xiao 0002, Jun Yu 0004, Le Zhang 0004 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2018 | Lineage-associated underrepresented permutations (LAUPs) of mammalian genomic sequences based on a Jellyfish-based LAUPs analysis application (JBLA)abstractMotivation: This study addresses several important questions related to naturally underrepresented sequences: (i) are there permutations of real genomic DNA sequences in a defined length (k-mer) and a given lineage that do not actually exist or underrepresented? (ii) If there are such sequences, what are their characteristics in terms of k-mer length and base composition? (iii) Are they related to CpG or TpA underrepresentation known for human sequences? We propose that the answers to these questions are of great significance for the study of sequence-associated regulatory mechanisms, such cytosine methylation and chromosomal structures in physiological or pathological conditions such as cancer. Results: We empirically defined sequences that were not included in any well-known public databases as lineage-associated underrepresented permutations (LAUPs). Then, we developed a Jellyfish-based LAUPs analysis application (JBLA) to investigate LAUPs for 24 representative species. The present discoveries include: (i) lengths for the shortest LAUPs, ranging from 10 to 14, which collectively constitute a low proportion of the genome. (ii) Common LAUPs showing higher CG content over the analysed mammalian genome and possessing distinct CG*CG motifs. (iii) Neither CpG-containing LAUPs nor CpG island sequences are randomly structured and distributed over the genomes; some LAUPs and most CpG-containing sequences exhibit an opposite trend within the same k and n variants. In addition, we demonstrate that the JBLA algorithm is more efficient than the original Jellyfish for computing LAUPs. Availability and implementation: We developed a Jellyfish-based LAUP analysis (JBLA) application by integrating Jellyfish (Marçais and Kingsford, 2011), MEME (Bailey, et al., 2009) and the NCBI genome database (Pruitt, et al., 2007) applications, which are listed as Supplementary Material. Supplementary information: Supplementary data are available at Bioinformatics online. Le Zhang 0004, Ming Xiao 0002, Jingsong Zhou, Jun Yu 0004 |
Bioinform. | 4 |
| 2017 | MTD: a mammalian transcriptomic database to explore gene expression and regulationabstractA systematic transcriptome survey is essential for the characterization and comprehension of the molecular basis underlying phenotypic variations. Recently developed RNA-seq methodology has facilitated efficient data acquisition and information mining of transcriptomes in multiple tissues/cell lines. Current mammalian transcriptomic databases are either tissue-specific or species-specific, and they lack in-depth comparative features across tissues and species. Here, we present a mammalian transcriptomic database (MTD) that is focused on mammalian transcriptomes, and the current version contains data from humans, mice, rats and pigs. Regarding the core features, the MTD browses genes based on their neighboring genomic coordinates or joint KEGG pathway and provides expression information on exons, transcripts and genes by integrating them into a genome browser. We developed a novel nomenclature for each transcript that considers its genomic position and transcriptional features. The MTD allows a flexible search of genes or isoforms with user-defined transcriptional characteristics and provides both table-based descriptions and associated visualizations. To elucidate the dynamics of gene expression regulation, the MTD also enables comparative transcriptomic analysis in both intraspecies and interspecies manner. The MTD thus constitutes a valuable resource for transcriptomic and evolutionary studies. The MTD is freely accessible at http://mtd.cbi.ac.cn. Qianqian Sun, Feng Xian, Manman Sun, Wan Fang, Meili Chen, Jun Yu 0004, Jing-Fa Xiao |
Briefings Bioinform. | 9 |
| 2017 | ISVASE: identification of sequence variant associated with splicing event using RNA-seq dataabstractBACKGROUND: Exon recognition and splicing precisely and efficiently by spliceosome is the key to generate mature mRNAs. About one third or a half of disease-related mutations affect RNA splicing. Software PVAAS has been developed to identify variants associated with aberrant splicing by directly using RNA-seq data. However, it bases on the assumption that annotated splicing site is normal splicing, which is not true in fact. RESULTS: We develop the ISVASE, a tool for specifically identifying sequence variants associated with splicing events (SVASE) by using RNA-seq data. Comparing with PVAAS, our tool has several advantages, such as multi-pass stringent rule-dependent filters and statistical filters, only using split-reads, independent sequence variant identification in each part of splicing (junction), sequence variant detection for both of known and novel splicing event, additional exon-exon junction shift event detection if known splicing events provided, splicing signal evaluation, known DNA mutation and/or RNA editing data supported, higher precision and consistency, and short running time. Using a realistic RNA-seq dataset, we performed a case study to illustrate the functionality and effectiveness of our method. Moreover, the output of SVASEs can be used for downstream analysis such as splicing regulatory element study and sequence variant functional analysis. CONCLUSIONS: ISVASE is useful for researchers interested in sequence variants (DNA mutation and/or RNA editing) associated with splicing events. The package is freely available at https://sourceforge.net/projects/isvase/ . Hasan Awad Aljohi, Wanfei Liu, Jun Yu 0004, Songnian Hu |
BMC Bioinform. | 4 |
| 2014 | PanGP: A tool for quickly analyzing bacterial pan-genome profileabstractAbstract Summary: Pan-genome analyses have shed light on the dynamics and evolution of bacterial genome from the point of population. The explosive growth of bacterial genome sequence also brought an extremely big challenge to pan-genome profile analysis. We developed a tool, named PanGP, to complete pan-genome profile analysis for large-scale strains efficiently. PanGP has integrated two sampling algorithms, totally random (TR) and distance guide (DG). The DG algorithm drew sample strain combinations on the basis of genome diversity of bacterial population. The performance of these two algorithms have been evaluated on four bacteria populations with strain numbers varying from 30 to 200, and the DG algorithm exhibited overwhelming advantage on accuracy and stability than the TR algorithm. Availability: PanGP was developed with a user-friendly graphic interface and it was available at http://PanGP.big.ac.cn. Contact: [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Yongbing Zhao, Xinmiao Jia, Junhui Yang, Yunchao Ling, Zhang Zhang 0002, Jun Yu 0004, Jing-Fa Xiao |
Bioinform. | 6 |
| 2012 | PGAP: pan-genomes analysis pipelineabstractSUMMARY: With the rapid development of DNA sequencing technology, increasing bacteria genome data enable the biologists to dig the evolutionary and genetic information of prokaryotic species from pan-genome sight. Therefore, the high-efficiency pipelines for pan-genome analysis are mostly needed. We have developed a new pan-genome analysis pipeline (PGAP), which can perform five analytic functions with only one command, including cluster analysis of functional genes, pan-genome profile analysis, genetic variation analysis of functional genes, species evolution analysis and function enrichment analysis of gene clusters. PGAP's performance has been evaluated on 11 Streptococcus pyogenes strains. AVAILABILITY: PGAP is developed with Perl script on the Linux Platform and the package is freely available from http://pgap.sf.net. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yongbing Zhao, Junhui Yang, Shixiang Sun, Jing-Fa Xiao, Jun Yu 0004 |
Bioinform. | 6 |
| 2012 | Codon Deviation Coefficient: a novel measure for estimating codon usage bias and its statistical significanceabstractBACKGROUND: Genetic mutation, selective pressure for translational efficiency and accuracy, level of gene expression, and protein function through natural selection are all believed to lead to codon usage bias (CUB). Therefore, informative measurement of CUB is of fundamental importance to making inferences regarding gene function and genome evolution. However, extant measures of CUB have not fully accounted for the quantitative effect of background nucleotide composition and have not statistically evaluated the significance of CUB in sequence analysis. RESULTS: Here we propose a novel measure--Codon Deviation Coefficient (CDC)--that provides an informative measurement of CUB and its statistical significance without requiring any prior knowledge. Unlike previous measures, CDC estimates CUB by accounting for background nucleotide compositions tailored to codon positions and adopts the bootstrapping to assess the statistical significance of CUB for any given sequence. We evaluate CDC by examining its effectiveness on simulated sequences and empirical data and show that CDC outperforms extant measures by achieving a more informative estimation of CUB and its statistical significance. CONCLUSIONS: As validated by both simulated and empirical data, CDC provides a highly informative quantification of CUB and its statistical significance, useful for determining comparative magnitudes and patterns of biased codon usage for genes or genomes with diverse sequence compositions. Zhang Zhang 0002, Jun Li 0051, Peng Cui 0004, Feng Ding 0003, Jeffrey P. Townsend, Jun Yu 0004 |
BMC Bioinform. | 7 |
| 2005 | ReAS: Recovery of Ancestral Sequences for Transposable Elements from the Unassembled Reads of a Whole Genome ShotgunabstractWe describe an algorithm, ReAS, to recover ancestral sequences for transposable elements (TEs) from the unassembled reads of a whole genome shotgun. The main assumptions are that these TEs must exist at high copy numbers across the genome and must not be so old that they are no longer recognizable in comparison to their ancestral sequences. Tested on the japonica rice genome, ReAS was able to reconstruct all of the high copy sequences in the Repbase repository of known TEs, and increase the effectiveness of RepeatMasker in identifying TEs from genome sequences. Ruiqiang Li, Jia Ye, Songgang Li, Jing Wang 0003, Yujun Han, Jian Wang 0065, Huanming Yang, Jun Yu 0004, Gane Ka-Shu Wong, Jun Wang 0004 |
PLoS Comput. Biol. | 9 |