EDBT 2026 Demo / reviewers in the wild / expert
Jing-Fa Xiao
dblp:05/1154 · also Jingfa Xiao
· DBLP profile ↗
9ranked-venue papers
0as first author
5since 2021 · last 2022
0000-0002-2835-4340ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › comparative genomics › pangenomics
pan-genome analysis |
0.3 | 2 | 2014 | PanGP: A tool for quickly analyzing bacterial pan-genome profile · Bioinform. 2014 PGAP: pan-genomes analysis pipeline · Bioinform. 2012 |
Bioinformatics and computational biology › genomics › computational genomics
genome subsampling |
0.2 | 1 | 2014 | PanGP: A tool for quickly analyzing bacterial pan-genome profile · Bioinform. 2014 |
Bioinformatics and computational biology
comparative genomics |
0.1 | 1 | 2012 | PGAP: pan-genomes analysis pipeline · Bioinform. 2012 |
Bioinformatics and computational biology › gene expression analysis
gene clustering |
0.1 | 1 | 2012 | PGAP: pan-genomes analysis pipeline · Bioinform. 2012 |
Bioinformatics and computational biology
phylogenetics |
0.0 | 1 | 2012 | PGAP: pan-genomes analysis pipeline · Bioinform. 2012 |
Methods — techniques the papers use, named apart from their topics
random sampling · 0.2distance-guided sampling · 0.2edit quality metrics · 0.2community curation · 0.2functional enrichment analysis · 0.1clustering · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | TSomVar: a tumor-only somatic and germline variant identification method with random forestabstractSomatic variants act as critical players during cancer occurrence and development. Thus, an accurate and robust method to identify them is the foundation of cutting-edge cancer genome research. However, due to low accessibility and high individual-/sample-specificity of the somatic variants in tumor samples, the detection is, to date, still crammed with challenges, particularly when lacking paired normal samples as control. To solve this burning issue, we developed a tumor-only somatic and germline variant identification method (TSomVar) using the random forest algorithm established on sample-specific variant datasets derived from genotype imputation, reads-mapping level annotation and functional annotation. We trained TSomVar by using genomic variant datasets of three major cancer types: colorectal cancer, hepatocellular carcinoma and skin cutaneous melanoma. Compared with existing tumor-only somatic variant identification tools, TSomVar shows excellent performances in somatic variant detection with higher accuracy and better capability of recalling for test datasets from colorectal cancer and skin cutaneous melanoma. In addition, TSomVar is equipped with the competence of accurately identifying germline variants in tumor samples. Taken together, TSomVar will undoubtedly facilitate and revolutionize somatic variant explorations in cancer research. Qi Wang 0072, Yunfei Shang, Congfan Bu, Mingming Lu, Meiye Jiang, Shuhuan Yu, Jingyao Zeng, Zaichao Zhang, Zhenglin Du, Jing-Fa Xiao |
Briefings Bioinform. | 12 |
| 2022 | Biomedical application community based on China high-performance computing environment
Lianhua He, Baohua Zhang 0003, Jing-Fa Xiao, Zhong Jin |
CCF Trans. High Perform. Comput. | 4 |
| 2021 | A fine-scale map of genome-wide recombination in divergent Escherichia coli populationabstractRecombination is one of the most important molecular mechanisms of prokaryotic genome evolution, but its exact roles are still in debate. Here we try to infer genome-wide recombination within a species, utilizing a dataset of 149 complete genomes of Escherichia coli from diverse animal hosts and geographic origins, including 45 in-house sequenced with the single-molecular real-time platform. Two major clades identified based on physiological, clinical and ecological characteristics form distinct genetic lineages based on scarcity of interclade gene exchanges. By defining gene-based syntenies for genomic segments within and between the two clades, we build a fine-scale recombination map for this representative global E. coli population. The map suggests extensive within-clade recombination that often breaks physical linkages among individual genes but seldom interrupts the structure of genome organizational frameworks as well as primary metabolic portfolios supported by the framework integrity, possibly due to strong natural selection for both physiological compatibility and ecological fitness. In contrast, the between-clade recombination declines drastically when phylogenetic distance increases to the extent where a 10-fold reduction can be observed, establishing a firm genetic barrier between clades. Our empirical data suggest a critical role for such recombination events in the early stage of speciation where recombination rate is associated with phylogenetic distance in addition to sequence and gene variations. The extensive intraclade recombination binds sister strains into a quasisexual group and optimizes genes or alleles to streamline physiological activities, whereas the sharply declined interclade recombination split the population into clades adaptive to divergent ecological niches. Yu Kang 0004, Lina Yuan, Yanan Chu, Xinmiao Jia, Qin Ma 0003, Jian Wang 0065, Jing-Fa Xiao, Songnian Hu, Zhancheng Gao, Jun Yu 0004 |
Briefings Bioinform. | 10 |
| 2021 | RefRGim: an intelligent reference panel reconstruction method for genotype imputation with convolutional neural networksabstractGenotype imputation is a statistical method for estimating missing genotypes from a denser haplotype reference panel. Existing methods usually performed well on common variants, but they may not be ideal for low-frequency and rare variants. Previous studies showed that the population similarity between study and reference panels is one of the key factors influencing the imputation accuracy. Here, we developed an imputation reference panel reconstruction method (RefRGim) using convolutional neural networks (CNNs), which can generate a study-specified reference panel for each input data based on the genetic similarity of individuals from current study and references. The CNNs were pretrained with single nucleotide polymorphism data from the 1000 Genomes Project. Our evaluations showed that genotype imputation with RefRGim can achieve higher accuracies than original reference panel, especially for low-frequency and rare variants. RefRGim will serve as an efficient reference panel reconstruction method for genotype imputation. RefRGim is freely available via GitHub: https://github.com/shishuo16/RefRGim. Qiheng Qian, Shuhuan Yu, Qi Wang 0072, Jinyue Wang, Jingyao Zeng, Zhenglin Du, Jing-Fa Xiao |
Briefings Bioinform. | 8 |
| 2021 | Applications and challenges of high performance computing in genomics
Meiye Jiang, Congfan Bu, Jingyao Zeng, Zhenglin Du, Jing-Fa Xiao |
CCF Trans. High Perform. Comput. | 5 |
| 2017 | MTD: a mammalian transcriptomic database to explore gene expression and regulationabstractA systematic transcriptome survey is essential for the characterization and comprehension of the molecular basis underlying phenotypic variations. Recently developed RNA-seq methodology has facilitated efficient data acquisition and information mining of transcriptomes in multiple tissues/cell lines. Current mammalian transcriptomic databases are either tissue-specific or species-specific, and they lack in-depth comparative features across tissues and species. Here, we present a mammalian transcriptomic database (MTD) that is focused on mammalian transcriptomes, and the current version contains data from humans, mice, rats and pigs. Regarding the core features, the MTD browses genes based on their neighboring genomic coordinates or joint KEGG pathway and provides expression information on exons, transcripts and genes by integrating them into a genome browser. We developed a novel nomenclature for each transcript that considers its genomic position and transcriptional features. The MTD allows a flexible search of genes or isoforms with user-defined transcriptional characteristics and provides both table-based descriptions and associated visualizations. To elucidate the dynamics of gene expression regulation, the MTD also enables comparative transcriptomic analysis in both intraspecies and interspecies manner. The MTD thus constitutes a valuable resource for transcriptomic and evolutionary studies. The MTD is freely accessible at http://mtd.cbi.ac.cn. Qianqian Sun, Feng Xian, Manman Sun, Wan Fang, Meili Chen, Jun Yu 0004, Jing-Fa Xiao |
Briefings Bioinform. | 10 |
| 2014 | PanGP: A tool for quickly analyzing bacterial pan-genome profileabstractAbstract Summary: Pan-genome analyses have shed light on the dynamics and evolution of bacterial genome from the point of population. The explosive growth of bacterial genome sequence also brought an extremely big challenge to pan-genome profile analysis. We developed a tool, named PanGP, to complete pan-genome profile analysis for large-scale strains efficiently. PanGP has integrated two sampling algorithms, totally random (TR) and distance guide (DG). The DG algorithm drew sample strain combinations on the basis of genome diversity of bacterial population. The performance of these two algorithms have been evaluated on four bacteria populations with strain numbers varying from 30 to 200, and the DG algorithm exhibited overwhelming advantage on accuracy and stability than the TR algorithm. Availability: PanGP was developed with a user-friendly graphic interface and it was available at http://PanGP.big.ac.cn. Contact: [email protected] or [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Yongbing Zhao, Xinmiao Jia, Junhui Yang, Yunchao Ling, Zhang Zhang 0002, Jun Yu 0004, Jing-Fa Xiao |
Bioinform. | 8 |
| 2013 | AuthorReward: increasing community curation in biological knowledge wikis through automated authorship quantificationabstractSUMMARY: Community curation-harnessing community intelligence in knowledge curation, bears great promise in dealing with the flood of biological knowledge. To exploit the full potential of the scientific community for knowledge curation, multiple biological wikis (bio-wikis) have been built to date. However, none of them have achieved a substantial impact on knowledge curation. One of the major limitations in bio-wikis is insufficient community participation, which is intrinsically because of lack of explicit authorship and thus no credit for community curation. To increase community curation in bio-wikis, here we develop AuthorReward, an extension to MediaWiki, to reward community-curated efforts in knowledge curation. AuthorReward quantifies researchers' contributions by properly factoring both edit quantity and quality and yields automated explicit authorship according to their quantitative contributions. AuthorReward provides bio-wikis with an authorship metric, helpful to increase community participation in bio-wikis and to achieve community curation of massive biological knowledge. AVAILABILITY: http://cbb.big.ac.cn/software. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Ming Tian, Jing-Fa Xiao, Xumin Wang, Jeffrey P. Townsend, Zhang Zhang 0002 |
Bioinform. | 4 |
| 2012 | PGAP: pan-genomes analysis pipelineabstractSUMMARY: With the rapid development of DNA sequencing technology, increasing bacteria genome data enable the biologists to dig the evolutionary and genetic information of prokaryotic species from pan-genome sight. Therefore, the high-efficiency pipelines for pan-genome analysis are mostly needed. We have developed a new pan-genome analysis pipeline (PGAP), which can perform five analytic functions with only one command, including cluster analysis of functional genes, pan-genome profile analysis, genetic variation analysis of functional genes, species evolution analysis and function enrichment analysis of gene clusters. PGAP's performance has been evaluated on 11 Streptococcus pyogenes strains. AVAILABILITY: PGAP is developed with Perl script on the Linux Platform and the package is freely available from http://pgap.sf.net. CONTACT: [email protected]; [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yongbing Zhao, Junhui Yang, Shixiang Sun, Jing-Fa Xiao, Jun Yu 0004 |
Bioinform. | 5 |