EDBT 2026 Demo / reviewers in the wild / expert
Feng Gao 0001
dblp:10/2674-1
· DBLP profile ↗
30ranked-venue papers
7as first author
16since 2021 · last 2026
0000-0002-9563-3841ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 6 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 4 since 2021Computer networks · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Outdoor full-range geolocator with multi-vision fusion and self-supervised truncation filtering
Feng Gao 0001, Yeyun Cai, Hailing Fu, Fang Deng |
Sci. China Inf. Sci. | 2 |
| 2025 | Tile-size aware bitrate allocation for adaptive 360$^{\circ }$ video streaming
Jiawei Huang 0001, Jingling Liu, Feng Gao 0001, Weihe Li, Jianxin Wang 0001 |
Multim. Tools Appl. | 4 |
| 2024 | Nmix: a hybrid deep learning model for precise prediction of 2'-O-methylation sites based on multi-feature fusion and ensemble learningabstractRNA 2'-O-methylation (Nm) is a crucial post-transcriptional modification with significant biological implications. However, experimental identification of Nm sites is challenging and resource-intensive. While multiple computational tools have been developed to identify Nm sites, their predictive performance, particularly in terms of precision and generalization capability, remains deficient. We introduced Nmix, an advanced computational tool for precise prediction of Nm sites in human RNA. We constructed the largest, low-redundancy dataset of experimentally verified Nm sites and employed an innovative multi-feature fusion approach, combining one-hot, Z-curve and RNA secondary structure encoding. Nmix utilizes a meticulously designed hybrid deep learning architecture, integrating 1D/2D convolutional neural networks, self-attention mechanism and residual connection. We implemented asymmetric loss function and Bayesian optimization-based ensemble learning, substantially improving predictive performance on imbalanced datasets. Rigorous testing on two benchmark datasets revealed that Nmix significantly outperforms existing state-of-the-art methods across various metrics, particularly in precision, with average improvements of 33.1% and 60.0%, and Matthews correlation coefficient, with average improvements of 24.7% and 51.1%. Notably, Nmix demonstrated exceptional cross-species generalization capability, accurately predicting 93.8% of experimentally verified Nm sites in rat RNA. We also developed a user-friendly web server (https://tubic.org/Nm) and provided standalone prediction scripts to facilitate widespread adoption. We hope that by providing a more accurate and robust tool for Nm site prediction, we can contribute to advancing our understanding of Nm mechanisms and potentially benefit the prediction of other RNA modification sites. Yu-Qing Geng, Fei-Liao Lai, Hao Luo 0002, Feng Gao 0001 |
Briefings Bioinform. | 4 |
| 2024 | Unveiling human origins of replication using deep learning: accurate prediction and comprehensive analysisabstractAccurate identification of replication origins (ORIs) is crucial for a comprehensive investigation into the progression of human cell growth and cancer therapy. Here, we proposed a computational approach Ori-FinderH, which can efficiently and precisely predict the human ORIs of various lengths by combining the Z-curve method with deep learning approach. Compared with existing methods, Ori-FinderH exhibits superior performance, achieving an area under the receiver operating characteristic curve (AUC) of 0.9616 for K562 cell line in 10-fold cross-validation. In addition, we also established a cross-cell-line predictive model, which yielded a further improved AUC of 0.9706. The model was subsequently employed as a fitness function to support genetic algorithm for generating artificial ORIs. Sequence analysis through iORI-Euk revealed that a vast majority of the created sequences, specifically 98% or more, incorporate at least one ORI for three cell lines (Hela, MCF7 and K562). This innovative approach could provide more efficient, accurate and comprehensive information for experimental investigation, thereby further advancing the development of this field. Zhen-Ning Yin, Fei-Liao Lai, Feng Gao 0001 |
Briefings Bioinform. | 3 |
| 2024 | Achieving QoE Fairness in Bitrate Allocation of 360° Video StreamingabstractIn tile-based 360° video streaming, the users employ the tile rate allocation algorithm to select appropriate bitrate to maximize the quality of experience (QoE). The preferences and viewports, however, can vary significantly across the different users. Since the users independently choose their bitrate according to their own preferences and viewports, it is hard to ensure QoE fairness for users under the constraint of available bandwidth. In this article, we propose a QoE-fairness aware bitrate allocation algorithm for multi-users (QBAM) to reduce difference of user QoE. According to the trajectory of the user viewpoint and user preferences for video quality, rebuffer time and quality switching, we leverage multi-agent reinforcement learning to train the bitrate allocation strategy. The experimental results show, compared with the current tile rate allocation algorithm, QBAM effectively improves the QoE fairness. Ping Zhong 0002, Jiawei Huang 0001, Feng Gao 0001, Jianxin Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | Auto-Kla: a novel web server to discriminate lysine lactylation sites using automated machine learningabstractRecently, lysine lactylation (Kla), a novel post-translational modification (PTM), which can be stimulated by lactate, has been found to regulate gene expression and life activities. Therefore, it is imperative to accurately identify Kla sites. Currently, mass spectrometry is the fundamental method for identifying PTM sites. However, it is expensive and time-consuming to achieve this through experiments alone. Herein, we proposed a novel computational model, Auto-Kla, to quickly and accurately predict Kla sites in gastric cancer cells based on automated machine learning (AutoML). With stable and reliable performance, our model outperforms the recently published model in the 10-fold cross-validation. To investigate the generalizability and transferability of our approach, we evaluated the performance of our models trained on two other widely studied types of PTM, including phosphorylation sites in host cells infected with SARS-CoV-2 and lysine crotonylation sites in HeLa cells. The results show that our models achieve comparable or better performance than current outstanding models. We believe that this method will become a useful analytical tool for PTM prediction and provide a reference for the future development of related models. The web server and source code are available at http://tubic.org/Kla and https://github.com/tubic/Auto-Kla, respectively. Fei-Liao Lai, Feng Gao 0001 |
Briefings Bioinform. | 2 |
| 2023 | Feature Alignment in Anchor-Free Object DetectionabstractMost anchor-free methods perform object detection using dense recommendation, which assumes that one point can simultaneously conduct accurate category prediction and regression estimation. However, due to different task drivers, valid features for classification and regression may locate at distinct areas in the training phase. This problem is called feature misalignment. To solve it, we propose a new feature alignment method based on anchor-free object detector. Firstly, a global receptive field adaptor (G-RFA) is designed by incorporating the feature pyramid networks (FPN) with the global attention mechanism, and forward features are further fine-tuned with a deformable-subnet (De-Subnet) to remove the influence of redundant contextual information. Then, a new feature filter strategy with a misalignment score is proposed to guide the network to focus on sampling points with aligned features. In addition, we establish mutually independent multi-layer quality distributions to model the priori information of an object on different FPN levels. Equipped with our method, the classification and regression features are aligned, and the generated foreground weight map converges to the centers of classification and regression heatmaps. Experimental results show that without bells and whistles, our method achieves 49.3% AP on MS COCO test-dev under the default 2x training schedule, outperforming related methods. Besides, experiments on PASCAL VOC demonstrate the generalization ability of our method. Code is available at https://github.com/GFENGG/featurealign. Feng Gao 0001, Yeyun Cai, Fang Deng, Chengpu Yu, Jie Chen 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Weighted Decentralized Information Filter for Collaborative Air-Ground Target Geolocation in Large Outdoor EnvironmentsabstractThe unmanned air-ground vehicle system has been successfully applied in civil and military domains. Collaborative vision-based target geolocation with this system can provide an enduring and accurate estimate of moving target state. Traditional decentralized information filter (DIF) treated each platform in the system identically. In fact, the observation capabilities of aerial and ground platform typically differ from each other due to different configurations and changing sensor noises. Without considering these differences, the resources of each platform cannot be fully utilized. To handle the issue, we develop a weighted DIF for geolocating of moving targets via air-ground collaboration. Specifically, it can produce a weighted factor autonomously for each platform based on the similarity of tracks from the air-ground system. Then, it is able to have more accurate global estimates than the traditional filter. Finally, simulation experiments and actual tests are conducted and the results are presented to validate the efficacy of the proposed method. Additional details can be seen in our video submission. Feng Gao 0001, Bofan Chen, Lele Xi, Fang Deng, Jie Chen 0003 |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |
| 2022 | Synthesizing Audio and Video Bitrate Selections via Learning from Actual RequirementsabstractAdaptive bitrate (ABR) algorithms are routinely adopted for transmitting media contents across dynamic networks. State-of-the-art ABR algorithms only adapt to video bitrate without considering audio bitrate adaption as they consider the im-pact on the video to be negligible due to the small size of the audio. However, to bring users an immersive experience, more and more content providers have applied high-quality audio with large sizes, like stereophonic and surround (Dolby Atmos). Therefore, improper audio bitrate selection will ad-versely affect video bitrate selection, leading to undesirable audio/video combinations (the highest video quality with the lowest audio quality, vice versa) and frequent playback inter-ruptions. To address these inefficiencies, we propose a Self-Play reinforcement learning-based Audio-aware ABR algorithm named SPA to learn strategies for audio and video bi-trate selections. Experimental results demonstrate SPA's con-siderable superiority as compared with existing approaches. Weihe Li, Jiawei Huang 0001, Jingling Liu, Feng Gao 0001 |
ICME | 5 |
| 2022 | High-quality pan-genome of Escherichia coli generated by excluding confounding and highly similar strains reveals an association between unique gene clusters and genomic islandsabstractThe pan-genome analysis of bacteria provides detailed insight into the diversity and evolution of a bacterial population. However, the genomes involved in the pan-genome analysis should be checked carefully, as the inclusion of confounding strains would have unfavorable effects on the identification of core genes, and the highly similar strains could bias the results of the pan-genome state (open versus closed). In this study, we found that the inclusion of highly similar strains also affects the results of unique genes in pan-genome analysis, which leads to a significant underestimation of the number of unique genes in the pan-genome. Therefore, these strains should be excluded from pan-genome analysis at the early stage of data processing. Currently, tens of thousands of genomes have been sequenced for Escherichia coli, which provides an unprecedented opportunity as well as a challenge for pan-genome analysis of this classical model organism. Using the proposed strategies, a high-quality E. coli pan-genome was obtained, and the unique genes was extracted and analyzed, revealing an association between the unique gene clusters and genomic islands from a pan-genome perspective, which may facilitate the identification of genomic islands. Feng Gao 0001 |
Briefings Bioinform. | 2 |
| 2022 | GC-Profile 2.0: an extended web server for the prediction and visualization of CpG islandsabstractSUMMARY: Due to the spontaneous deamination of 5'-methylcytosine into thymine, the number of CpG dinucleotides is less than expected in vertebrate genomes. Exceptionally, there are a large number of CpG dinucleotides clustered at certain genomic loci, known as CpG islands (CGIs), where CpG dinucleotides are free from methylation. Identification of CGIs is of great significance in the field of genomics and epigenetics because they can serve as important gene markers or regulatory elements. Here, GC-Profile 2.0 has been presented as a newly extended application for CGIs detection and visualization. Based on a benchmark test of assembled sequences, GC-Profile 2.0 has shown better overall performance compared with other four popular methods. In addition, cumulative CpG profile, a visualization tool of CpG content variation, is also proposed to intuitively display the change trend of CpG content. AVAILABILITY AND IMPLEMENTATION: GC-Profile 2.0 is freely available at http://tubic.org/GC-Profile2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Fei-Liao Lai, Feng Gao 0001 |
Bioinform. | 2 |
| 2022 | Prior area searching for energy-based sound source localization
Feng Gao 0001, Yeyun Cai, Fang Deng, Chengpu Yu, Jie Chen 0003 |
Sci. China Inf. Sci. | 1 |
| 2021 | Ori-Finder 3: a web server for genome-wide prediction of replication origins in Saccharomyces cerevisiaeabstractDNA replication is a fundamental process in all organisms; this event initiates at sites termed origins of replication. The characteristics of eukaryotic replication origins are best understood in Saccharomyces cerevisiae. For this species, origin prediction algorithms or web servers have been developed based on the sequence features of autonomously replicating sequences (ARSs). However, their performances are far from satisfactory. By utilizing the Z-curve methodology, we present a novel pipeline, Ori-Finder 3, for the computational prediction of replication origins in S. cerevisiae at the genome-wide level based solely on DNA sequences. The ARS exhibiting both an AT-rich stretch and ARS consensus sequence element can be predicted at the single-nucleotide level. For the identified ARSs in the S. cerevisiae reference genome, 83 and 60% of the top 100 and top 300 predictions matched the known ARS records, respectively. Based on Ori-Finder 3, we subsequently built a database of the predicted ARSs identified in more than a hundred S. cerevisiae genomes. Consequently, we developed a user-friendly web server including the ARS prediction pipeline and the predicted ARSs database, which can be freely accessed at http://tubic.tju.edu.cn/Ori-Finder3. Dan Wang 0028, Fei-Liao Lai, Feng Gao 0001 |
Briefings Bioinform. | 3 |
| 2021 | Toward a high-quality pan-genome landscape of Bacillus subtilis by removal of confounding strainsabstractPan-genome analysis is widely used to study the evolution and genetic diversity of species, particularly in bacteria. However, the impact of strain selection on the outcome of pan-genome analysis is poorly understood. Furthermore, a standard protocol to ensure high-quality pan-genome results is lacking. In this study, we carried out a series of pan-genome analyses of different strain sets of Bacillus subtilis to understand the impact of various strains on the performance and output quality of pan-genome analyses. Consequently, we found that the results obtained by pan-genome analyses of B. subtilis can be influenced by the inclusion of incorrectly classified Bacillus subspecies strains, phylogenetically distinct strains, engineered genome-reduced strains, chimeric strains, strains with a large number of unique genes or a large proportion of pseudogenes, and multiple clonal strains. Since the presence of these confounding strains can seriously affect the quality and true landscape of the pan-genome, we should remove these deviations in the process of pan-genome analyses. Our study provides new insights into the removal of biases from confounding strains in pan-genome analyses at the beginning of data processing, which enables the achievement of a closer representation of a high-quality pan-genome landscape of B. subtilis that better reflects the performance and credibility of the B. subtilis pan-genome. This procedure could be added as an important quality control step in pan-genome analyses for improving the efficiency of analyses, and ultimately contributing to a better understanding of genome function, evolution and genome-reduction strategies for B. subtilis in the future. Dan Wang 0028, Feng Gao 0001 |
Briefings Bioinform. | 3 |
| 2021 | Comparison of the binding characteristics of SARS-CoV and SARS-CoV-2 RBDs to ACE2 at different temperatures by MD simulationsabstractTemperature plays a significant role in the survival and transmission of SARS-CoV (severe acute respiratory syndrome coronavirus) and SARS-CoV-2. To reveal the binding differences of SARS-CoV and SARS-CoV-2 receptor-binding domains (RBDs) to angiotensin-converting enzyme 2 (ACE2) at different temperatures at atomic level, 20 molecular dynamics simulations were carried out for SARS-CoV and SARS-CoV-2 RBD-ACE2 complexes at five selected temperatures, i.e. 200, 250, 273, 300 and 350 K. The analyses on structural flexibility and conformational distribution indicated that the structure of the SARS-CoV-2 RBD was more stable than that of the SARS-CoV RBD at all investigated temperatures. Then, molecular mechanics Poisson-Boltzmann surface area and solvated interaction energy approaches were combined to estimate the differences in binding affinity of SARS-CoV and SARS-CoV-2 RBDs to ACE2; it is found that the binding ability of ACE2 to the SARS-CoV-2 RBD was stronger than that to the SARS-CoV RBD at five temperatures, and the main reason for promoting such binding differences is electrostatic and polar interactions between RBDs and ACE2. Finally, the hotspot residues facilitating the binding of SARS-CoV and SARS-CoV-2 RBDs to ACE2, the key differential residues contributing to the difference in binding and the interaction mechanism of differential residues that exist at all investigated temperatures were analyzed and compared in depth. The current work would provide a molecular basis for better understanding of the high infectiousness of SARS-CoV-2 and offer better theoretical guidance for the design of inhibitors targeting infectious diseases caused by SARS-CoV-2. Fang-Fang Yan, Feng Gao 0001 |
Briefings Bioinform. | 2 |
| 2021 | Data-driven identification of SARS-CoV-2 subpopulations using PhenoGraph and binary-coded genomic dataabstractFor epidemic prevention and control, the identification of SARS-CoV-2 subpopulations sharing similar micro-epidemiological patterns and evolutionary histories is necessary for a more targeted investigation into the links among COVID-19 outbreaks caused by SARS-CoV-2 with similar genetic backgrounds. Genomic sequencing analysis has demonstrated the ability to uncover viral genetic diversity. However, an objective analysis is necessary for the identification of SARS-CoV-2 subpopulations. Herein, we detected all the mutations in 186 682 SARS-CoV-2 isolates. We found that the GC content of the SARS-CoV-2 genome had evolved to be lower, which may be conducive to viral spread, and the frameshift mutation was rare in the global population. Next, we encoded the genomic mutations in binary form and used an unsupervised learning classifier, namely PhenoGraph, to classify this information. Consequently, PhenoGraph successfully identified 303 SARS-CoV-2 subpopulations, and we found that the PhenoGraph classification was consistent with, but more detailed and precise than the known GISAID clades (S, L, V, G, GH, GR, GV and O). By the change trend analysis, we found that the growth rate of SARS-CoV-2 diversity has slowed down significantly. We also analyzed the temporal, spatial and phylogenetic relationships among the subpopulations and revealed the evolutionary trajectory of SARS-CoV-2 to a certain extent. Hence, our results provide a better understanding of the patterns and trends in the genomic evolution and epidemiology of SARS-CoV-2. Zhi-Kai Yang, Lingyu Pan, Hao Luo 0002, Feng Gao 0001 |
Briefings Bioinform. | 5 |
| 2020 | Automatic and Interpretable Model for Periodontitis Diagnosis in Panoramic Radiographs
Haoyang Li 0011, Juexiao Zhou, Yi Zhou 0070, Jie Chen 0003, Feng Gao 0001, Ying Xu 0001, Xin Gao 0001 |
MICCAI (2) | 5 |
| 2020 | Long-Range Binocular Vision Target Geolocation Using Handheld Electronic Devices in Outdoor EnvironmentabstractBinocular vision is a passive method of simulating the human visual principle to perceive the distance to a target. Traditional binocular vision applied to target localization is usually suitable for short-range area and indoor environment. This paper presents a novel vision-based geolocation method for long-range targets in outdoor environment, using handheld electronic devices such as smart phones and tablets. This method solves the problems in long-range localization and determining geographic coordinates of the targets in outdoor environment. It is noted that these sensors necessary for binocular vision geolocation such as the camera, GPS, and inertial measurement unit (IMU), are intergrated in these handheld electronic devices. This method, employing binocular localization model and coordinate transformations, is provided for these handheld electronic devices to obtain the GPS coordinates of the targets. Finally, two types of handheld electronic devices are used to conduct the experiments for targets in long range up to 500m. The experimental results show that this method yields the target geolocation accuracy along horizontal direction with nearly 20m, achieving comparable or even better performance than monocular vision methods. Fang Deng, Feng Gao 0001, Huangbin Qiu, Xin Gao 0001, Jie Chen 0003 |
IEEE Trans. Image Process. | 3 |
| 2019 | An Optimization Deployment Scheme for Static Charging Piles Based on Dynamic of Shared E-BikesabstractShared e-bikes are popular because of their green, eco-friendly and efficient features. Due to the limited battery capacity of the e-bikes, the energy problem has become one of the main factors limiting its further development. The energy problem can be solved by using static charging piles (SCP) to replenish the batteries of shared e-bike. The location of the shared e-bike is time-varying, resulting in the optimal deployment of SCP as a complex location problem. In this paper, we propose an optimal Deployment algorithm for Maximum Coverage combined the Dynamic Changes of nodes (max-DCDC) based on the known number of SCP. This method first quantitatively analyzes the dynamic change process of the shared e-bike to reduce the deployment scope of the SCP. Then, according to the geometric characteristics of the e-bike distribution within the deployment scope to optimizes the deployment location of the SCP. Simulation experiments show that max-DCDC has better performance in terms of deployment stability and e-bike coverage compared with the other algorithms. Ping Zhong 0002, Aikun Xu, Yuanming Chen, Feng Gao 0001, Guihua Duan |
MSN | 4 |
| 2019 | Recent developments of software and database in microbial genomics and functional genomicsabstractWith the rapid progress in next-generation sequencing technologies and the exponential increase in available microbial genomes, mining biological knowledge from these genomes represents a challenge to the whole scientific community. In this themed issue, we present the invited review articles by the scientists who have made significant contributions in this field to review the state of the art of their influential work and discuss the upcoming challenges for the further development. This issue aims to provide a comprehensive overview of recent developments of software and database in microbial genomics and functional genomics, including those for replication origin (Ori-Finder system and DoriC database), prophage (PHAST and PHASTER Web servers), operon (DOOR database), secondary metabolite biosynthetic gene clusters (antiSMASH software), antimicrobial resistance (AMR) (PATRIC database), orthologous genes/proteins (COG database), pathway/genome (MicroScope and BioCyc database), genome visualization (CGView software family), metagenomic classification, assembly and analysis (Kraken, Centrifuge, VALET, MG-RAST, etc.) and multiple sequence alignment (MSA) (MAFFT online service). We will introduce the articles briefly in the order of their appearance to give a quick reference guide for the readers. For the purpose of phylogenetic classification of proteins from complete genomes, a collection of the COGs (Clusters of Orthologous Groups of proteins, re-branded as Clusters of Orthologous Genes later) was constructed 20 years ago. Since then, the COG approach has become one of the foundations of comparative and evolutionary genomics. At present, this pioneer work is still extremely popular as an essential tool in microbial genomics and a solid platform for comparative genomic analysis. Galperin et al. [1] review the key principles and the applications of the COG approach, and also discuss the unresolved problems in the COG approach and the possible solutions (Recommended in F1000 Prime). Toward a full understanding of the newly sequenced genomes, the MicroScope platform has been developed as an integrated environment to perform comparative genomic and metabolic analyses of microbial genomes. Furthermore, the MicroScope platform also provides a collaborative environment to share and improve knowledge on these genomes. From the end users' point of view, Médigue et al. [2] present a comprehensive description of MicroScope services to help the microbiologists extend their analyses of species of interest. To provide a reference on the thousands of sequenced microbial genomes and their metabolic pathways, the BioCyc collection has been generated based on the prediction by computer programs, the information integrated from other bioinformatics databases and the curation from the biomedical literature by biologist curators. Karp et al. [3] outline the recent improvements to BioCyc, including the expansion of BioCyc database content as well as the related bioinformatics tools for query, visualization and analysis. To facilitate the research on AMR, PATRIC (Pathosystems Resource Integration Center) includes AMR information at both the genome and gene level, curates the AMR-specific functional roles manually and builds the machine learning-based classifiers to predict the AMR phenotypes and their genetic determinants for the submitted genomes. Consequently, researchers can easily explore AMR data and design experiments based on whole genomes or individual genes, and they can also quickly obtain these data in their private genomes and compare with the PATRIC collection [4]. Toward a rapid and reliable identification of all the potential secondary metabolite biosynthetic gene clusters in the newly sequenced genomes, the genome mining platform antiSMASH (antibiotics and Secondary Metabolite Analysis Shell) has been released to combine and extend the functionality of the previous tools with a user-friendly Web interface. Blin et al. discuss the principles underlying the predictions of antiSMASH and other computational tools, and provide practical advice for their applications. In addition, Blin et al. [5] point out the important caveats that should be taken into consideration when designing and interpreting genome mining studies. To perform the prophage and cryptic prophage identification in bacterial genomes, PHAST (PHAge Search Tool) and PHASTER (PHAge Search Tool-Enhanced Release) have been presented in succession, which have already become two of the most widely used Web servers in this field. Arndt et al. [6] review the main capabilities of PHAST and PHASTER, provide some practical guidance regarding their applications and discuss possible future directions for the development of PHASTEST, a successor to PHASTER. To identify the operons and transcriptional units accurately, an operon database DOOR (Database of prOkaryotic OpeRons) for genome analyses and functional inference has been developed and updated continually. Cao et al. review the information stored in DOOR database as well as the utility tools in support of operon-based analyses and information discovery. In addition, Cao et al. [7] explain how to computationally derive such data and demonstrate how to facilitate systems-level studies based on them. For the purpose of large-scale identification and characterization of replication origins in microbial genomes, Ori-Finder system and DoriC database have been developed and maintained based on the Z-curve method. Luo et al. introduce the main methodology of Ori-Finder system used to predict the replication origins for bacteria (Ori-Finder Web server) or archaea (Ori-Finder 2 Web server), and review the development of DoriC (Database of oriCs in bacterial and archaeal genomes) and its application, including the large-scale analyses of oriCs and strand bias in prokaryotic genomes. In addition, Luo et al. [8] also present some future directions and aspects for extending the application of Ori-Finder and DoriC. To visualize and compare circular genomes by graphical genome maps, the CGView (Circular Genome Viewer) software family has been designed to generate the visually impressive graphical genome maps for bacteria, organelles and viruses. Stothard et al. [9] describe the capabilities of the original CGView program and subsequent companion applications, such as the CGView Server and the CGView Comparison Tool, and also discuss the newer GView program, a rewrite of CGView with support for linear maps and interactive editing, and its companion Web application (the GView Server), which provides the analysis pipelines for comparative genomics particularly. To assist in the analysis of metagenomics data sets, a great variety of computational tools have been developed. Breitwieser et al. [10] review the methods and databases for classification and assembly of metagenomics data, and also discuss the challenges presented by inconsistencies in microbial taxonomy as well as contamination in the genome resources. This review has successfully attracted the readers' attention, and once became the most read paper of Briefings in Bioinformatics. Despite recent advances in metagenomic assembly, validation is still the key for moving forward because of the imperfect assembled contigs. From a computational and technological perspective, Olson et al. highlight the recent developments in metagenomic assembly, summarize the key approaches for genomic and metagenomic assembly validation and demonstrate the insights derived from assemblies through the lens of validation. Olson et al. [11] also discuss the potential impact of long-read technologies on metagenomics, and future challenges and opportunities in the field of metagenomic assembly and validation. To provide the automatic phylogenetic and functional analysis of metagenomes, MG-RAST (Metagenomics RAST server) has been designed as a hosted, open-source and open-submission platform. Meyer et al. introduce the implementation, backend components, current workflow of MG-RAST version 4, which has increased throughput dramatically with the same amount of resources and brings a new user interface fully relying on the API (Application Programmers Interface) for data access and the analysis reuse. Meyer et al. also present the lessons learned from a decade of low-budget ultra-high-throughput metagenome analysis [12]. To meet the challenge of the MSA in the era of big data, MAFFT (Multiple Alignment based on Fast Fourier Transformation), a popular MSA program, has significantly improved in performance and usability recently. Katoh et al. [13] describe in detail the Web interface for newly added options for large data, as well as interactive usage to refine sequence data sets and MSAs, which will facilitate the MAFFT users in handling big data and extracting biological information more quickly and effectively (Recommended in F1000 Prime). In conclusion, this special issue will be helpful to the scientific community in the fields of microbiology, bioinformatics/computational biology, genomics, etc. Personally, I have benefited enormously from these outstanding works, whether studying as a graduate student a dozen years ago, or working as a researcher up to now. Here, I would like to take this opportunity to thank the authors for their excellent contributions to this issue. The author would like to thank Professor Chun-Ting Zhang for the invaluable assistance and inspiring discussions. This work was funded by National Natural Science Foundation of China (grant numbers 31571358, 21621004 and 31171238) and the National High-Tech Research and Development Program (863) of China (grant number 2015AA020101). Feng Gao 0001 |
Briefings Bioinform. | 1 |
| 2019 | Recent development of Ori-Finder system and DoriC database for microbial replication originsabstractDNA replication begins at replication origins in all three domains of life. Identification and characterization of replication origins are important not only in providing insights into the structure and function of the replication origins but also in understanding the regulatory mechanisms of the initiation step in DNA replication. The Z-curve method has been used in the identification of replication origins in archaeal genomes successfully since 2002. Furthermore, the Web servers of Ori-Finder and Ori-Finder 2 have been developed to predict replication origins in both bacterial and archaeal genomes based on the Z-curve method, and the replication origins with manual curation have been collected into an online database, DoriC. Ori-Finder system and DoriC database are currently used in the research field of DNA replication origins in prokaryotes, including: (i) identification of oriC regions in bacterial and archaeal genomes; (ii) discovery and analysis of the conserved sequences within oriC regions; and (iii) strand-biased analysis of bacterial genomes. Up to now, more and more predicted results by Ori-Finder system were supported by subsequent experiments, and Ori-Finder system has been used to identify the replication origins in > 100 newly sequenced prokaryotes in their genome reports. In addition, the data in DoriC database have been widely used in the large-scale analyses of replication origins and strand bias in prokaryotic genomes. Here, we review the development of Ori-Finder system and DoriC database as well as their applications. Some future directions and aspects for extending the application of Ori-Finder and DoriC are also presented. Hao Luo 0002, Chun-Lan Quan, Feng Gao 0001 |
Briefings Bioinform. | 4 |
| 2019 | Pan-genomic analysis provides novel insights into the association of E.coli with human host and its minimal genomeabstractMOTIVATION: Bacteria can usually acquire certain advantageous genes that enable the bacteria to adapt to rapidly changing niches, thereby leading to a wide range of intraspecific genome content and genetic redundancy. The minimal genome of Escherichia coli, which is the most important bacterial species, and the association between E.coli and its human host are worthy of further exploration. RESULTS: We used gene prediction and phylogenetic analysis to reveal a rich phylogenetic diversity among 491 E.coli strains and to reveal substantial differences between these strains with respect to gene number and genome length. We used pan-genomic analysis to accurately identify 867 core genes, in which only 243 genes are shared by essential genes. This analysis revealed that core genes mainly provide essential functions to the basic lifestyle of E.coli, and accessory genes are likely to confer selective advantages such as niche adaptation or the ability to colonize specific hosts. By association analysis, we found that E.coli strains in non-human hosts may more easily utilize foreign genetic materials to adapt to their surroundings, but the population in human hosts has higher demands for the control of population density, indicating that highly accurate quorum-sensing behavior is very important for harmony between E.coli and its human host. By considering core genes and previous deletions together, we proposed a potential direction for further reduction of the E.coli genome. AVAILABILITY AND IMPLEMENTATION: The data, analysis process and detailed information on software tools used in this study are all available in the supplementary material. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zhi-Kai Yang, Hao Luo 0002, Baijing Wang, Feng Gao 0001 |
Bioinform. | 5 |
| 2018 | The systematic analysis of ultraconserved genomic regions in the budding yeastabstractBioinformatics (2017) doi: 10.1093/bioinformatics/btx619 The publisher wishes to inform readers that the following grant number was inadvertently left off the published article: Grant No. 2015AA020101. This has now been corrected online. Zhi-Kai Yang, Feng Gao 0001 |
Bioinform. | 2 |
| 2018 | The systematic analysis of ultraconserved genomic regions in the budding yeastabstractMotivation: In the evolution of species, a kind of special sequences, termed ultraconserved sequences (UCSs), have been inherited without any change, which strongly suggests those sequences should be crucial for the species to survive or adapt to the environment. However, the UCSs are still regarded as mysterious genetic sequences so far. Here, we present a systematic study of ultraconserved genomic regions in the budding yeast based on the publicly available genome sequences, in order to reveal their relationship with the adaptability or fitness advantages of the budding yeast. Results: Our results indicate that, in addition to some fundamental biological functions, the UCSs play an important role in the adaptation of Saccharomyces cerevisiae to the acidic environment, which is backed up by the previous observation. Besides that, we also find the highly unchanged genes are enriched in some other pathways, such as the nutrient-sensitive signaling pathway. To facilitate the investigation of unique UCSs, the UCSC Genome Browser was utilized to visualize the chromosomal position and related annotations of UCSs in S.cerevisiae genome. Availability and implementation: For more details on UCSs, please refer to the Supplementary information online, and the custom code is available on request. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Zhi-Kai Yang, Feng Gao 0001 |
Bioinform. | 2 |
| 2018 | Recombinational DSBs-intersected genes converge on specific disease- and adaptability-related pathwaysabstractMotivation: The budding yeast Saccharomyces cerevisiae is a model species powerful for studying the recombination of eukaryotes. Although many recombination studies have been performed for this species by experimental methods, the population genomic study based on bioinformatics analyses is urgently needed to greatly increase the range and accuracy of recombination detection. Here, we carry out the population genomic analysis of recombination in S.cerevisiae to reveal the potential rules between recombination and evolution in eukaryotes. Results: By population genomic analysis, we discover significantly more and longer recombination events in clinical strains, which indicates that adverse environmental conditions create an obviously wider range of genetic combination in response to the selective pressure. Based on the analysis of recombinational double strand breaks (DSBs)-intersected genes (RDIGs), we find that RDIGs significantly converge on specific disease- and adaptability-related pathways, indicating that recombination plays a biologically key role in the repair of DSBs related to diseases and environmental adaptability, especially the human neurological disorders. By evolutionary analysis of RDIGs, we find that the RDIGs highly prevailing in populations of yeast tend to be more evolutionarily conserved, indicating the accurate repair of DSBs in these RDIGs is critical to ensure the eukaryotic survival or fitness. Supplementary information: Supplementary data are available at Bioinformatics online. Zhi-Kai Yang, Hao Luo 0002, Baijing Wang, Feng Gao 0001 |
Bioinform. | 5 |
| 2017 | Zisland Explorer: detect genomic islands by combining homogeneity and heterogeneity propertiesabstractGenomic islands are genomic fragments of alien origin in bacterial and archaeal genomes, usually involved in symbiosis or pathogenesis. In this work, we described Zisland Explorer, a novel tool to predict genomic islands based on the segmental cumulative GC profile. Zisland Explorer was designed with a novel strategy, as well as a combination of the homogeneity and heterogeneity of genomic sequences. While the sequence homogeneity reflects the composition consistence within each island, the heterogeneity measures the composition bias between an island and the core genome. The performance of Zisland Explorer was evaluated on the data sets of 11 different organisms. Our results suggested that the true-positive rate (TPR) of Zisland Explorer was at least 10.3% higher than that of four other widely used tools. On the other hand, the new tool did not lose overall accuracy with the improvement in the TPR and showed better equilibrium among various evaluation indexes. Also, Zisland Explorer showed better accuracy in the prediction of experimental island data. Overall, the tool provides an alternative solution over other tools, which expands the field of island prediction and offers a supplement to increase the performance of the distinct predicting strategy. We have provided a web service as well as a graphical user interface and open-source code across multiple platforms for Zisland Explorer, which is available at http://cefg.uestc.edu.cn/Zisland_Explorer/ or http://tubic.tju.edu.cn/Zisland_Explorer/. Feng Gao 0001, Meng-Ze Du, Hong-Li Hua, Feng-Biao Guo |
Briefings Bioinform. | 2 |
| 2012 | DeOri: a database of eukaryotic DNA replication originsabstractSUMMARY: DNA replication, a central event for cell proliferation, is the basis of biological inheritance. The identification of replication origins helps to reveal the mechanism of the regulation of DNA replication. However, only few eukaryotic replication origins were characterized not long ago; nevertheless, recent genome-wide approaches have boosted the number of mapped replication origins. To gain a comprehensive understanding of the nature of eukaryotic replication origins, we have constructed a Database of Eukaryotic ORIs (DeOri), which contains all the eukaryotic ones identified by genome-wide analyses currently available. A total of 16 145 eukaryotic replication origins have been collected from 6 eukaryotic organisms in which genome-wide studies have been performed, the replication-origin numbers being 433, 7489, 1543, 148, 348 and 6184 for humans, mice, Arabidopsis thaliana, Kluyveromyces lactis, Schizosaccharomyces pombe and Drosophila melanogaster, respectively. AVAILABILITY: Database of Eukaryotic ORIs (DeOri) can be accessed from http://tubic.tju.edu.cn/deori/ Feng Gao 0001, Hao Luo 0002, Chun-Ting Zhang |
Bioinform. | 1 |
| 2008 | Ori-Finder: A web-based system for finding oriCs in unannotated bacterial genomesabstractBACKGROUND: Chromosomal replication is the central event in the bacterial cell cycle. Identification of replication origins (oriCs) is necessary for almost all newly sequenced bacterial genomes. Given the increasing pace of genome sequencing, the current available software for predicting oriCs, however, still leaves much to be desired. Therefore, the increasing availability of genome sequences calls for improved software to identify oriCs in newly sequenced and unannotated bacterial genomes. RESULTS: We have developed Ori-Finder, an online system for finding oriCs in bacterial genomes based on an integrated method comprising the analysis of base composition asymmetry using the Z-curve method, distribution of DnaA boxes, and the occurrence of genes frequently close to oriCs. The program can also deal with unannotated genome sequences by integrating the gene-finding program ZCURVE 1.02. Output of the predicted results is exported to an HTML report, which offers convenient views on the results in both graphical and tabular formats. CONCLUSION: A web-based system to predict replication origins of bacterial genomes has been presented here. Based on this system, oriC regions have been predicted for the bacterial genomes available in GenBank currently. It is hoped that Ori-Finder will become a useful tool for the identification and analysis of oriCs in both bacterial and archaeal genomes. Feng Gao 0001, Chun-Ting Zhang |
BMC Bioinform. | 1 |
| 2007 | DoriC: a database of oriC regions in bacterial genomesabstractUNLABELLED: Replication origins (oriCs) of bacterial genomes currently available in GenBank have been predicted by using a systematic method comprising the Z-curve analysis for nucleotide distribution asymmetry, DnaA box distribution, genes adjacent to candidate oriCs and phylogenetic relationships. These oriCs are organized into a MySQL database, DoriC, which provides extensive information and graphical views of the oriC regions. In addition, users can Blast a query sequence or even a whole genome against DoriC to find a homologous one. DoriC will be updated timely and the latest version is DoriC 1.8, in which oriCs of 425 genomes (468 chromosomes) are identified. AVAILABILITY: DoriC can be accessed from http://tubic.tju.edu.cn/doric/. SUPPLEMENTARY INFORMATION: Supplementary data are available at http://tubic.tju.edu.cn/doric/supplementary.htm. Feng Gao 0001, Chun-Ting Zhang |
Bioinform. | 1 |
| 2004 | Comparison of various algorithms for recognizing short coding sequences of human genesabstractMOTIVATION: Since the early 1980s of the twentieth century, there has been great progress in the development of computational gene-finding algorithms. Some problems, however, have not yet been solved currently. Recognizing short genes in prokaryotes and short exons in eukaryotes is one of such problems. The paper is devoted to assessing various algorithms, including those currently available and the new ones proposed here, in order to find the best algorithm to solve the issue. RESULTS: The databases consisting of phase-specific coding and non-coding sequences of human genes with length of 192, 162, 129, 108, 87, 63 and 42 bp, respectively, have been established. Based on the databases and a standard benchmark, 19 algorithms were evaluated, which include the methods of Markov models with orders of 1 through 5, codon usage, hexamer usage, codon preference, amino acid usage, codon prototype, Fourier transform and 8 Z curve methods with various numbers of parameters. Consequently, the Z curve methods with 69 and 189 parameters are the best ones among them, based on the databases constructed here. In addition to the highest recognition accuracy confirmed by 10-fold cross-validation tests, the Z curve methods are much simpler computationally than the second best one, the fifth-order Markov chain model, in which 12 288 parameters are used. We hope that the Z curve methods presented in this paper would be beneficial to the further development of gene-finding algorithms. AVAILABILITY: The programs of various Z curve methods are available on request. Feng Gao 0001, Chun-Ting Zhang |
Bioinform. | 1 |