VLDB 2026 Research / reviewers in the wild / expert
Xiaodong Fang
dblp:38/5281
· DBLP profile ↗
8ranked-venue papers
1as first author
6since 2021 · last 2026
0000-0001-7061-3337ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhanced Disease Susceptible Variant Identification via Short Identity by Descent SegmentsabstractRare diseases affect millions of individuals worldwide, yet diagnostic yields for them still remain low. Among variant identification approaches, identity by descent (IBD) mapping is used to identify disease susceptible variants originating from a recent common ancestor among affected individuals, but existing IBD detection models struggle to identify these variants in short IBD segments. Here, we introduce SILO, a novel model to detect disease susceptible variants in both short and long IBD segments. SILO employs a two-stage procedure to detect IBD segments. In the first stage, SILO identifies long IBD segments based on common variants. In the second stage, SILO utilizes rare variants to detect short IBD segments using a seed-and-extend algorithm. We evaluated SILO in simulated data and real data from the 1000 Genomes Project. Our results demonstrate that SILO outperforms existing models in detecting disease susceptible variants within short IBD segments, and show comparable performance in detecting these variants within longer IBD segments. These findings highlight the potential of SILO to increase diagnostic yields for rare diseases by enhancing the identification of previously overlooked disease susceptible variants in short IBD segments. Nonetheless, we note that the detection of short IBD segments remains challenging due to limited precision, leaving room for future improvement. Chonghao Wang, Werner Pieter Veldsman, Yufen Huang, Xiaodong Fang, Lu Zhang 0061 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2024 | VCF2PCACluster: a simple, fast and memory-efficient tool for principal component analysis of tens of millions of SNPsabstractPrincipal component analysis (PCA) is an important and widely used unsupervised learning method that determines population structure based on genetic variation. Genome sequencing of thousands of individuals usually generate tens of millions of SNPs, making it challenging for PCA analysis and interpretation. Here we present VCF2PCACluster, a simple, fast and memory-efficient tool for Kinship estimation, PCA and clustering analysis, and visualization based on VCF formatted SNPs. We implemented five Kinship estimation methods and three clustering methods for its users to choose from. Moreover, unlike other PCA tools, VCF2PCACluster possesses a clustering function based on PCA result, which enabling users to automatically and clearly know about population structure. We demonstrated the same accuracy but a higher performance of this tool in performing PCA analysis on tens of millions of SNPs compared to another popular PLINK2 software, especially in peak memory usage that is independent of the number of SNPs in VCF2PCACluster. Weiming He, Lian Xu, JingXian Wang, Zhen Yue, Yi Jing, Shuaishuai Tai, Xiaodong Fang |
BMC Bioinform. | 8 |
| 2023 | Benchmarking multi-platform sequencing technologies for human genome assemblyabstractGenome assembly is a computational technique that involves piecing together deoxyribonucleic acid (DNA) fragments generated by sequencing technologies to create a comprehensive and precise representation of the entire genome. Generating a high-quality human reference genome is a crucial prerequisite for comprehending human biology, and it is also vital for downstream genomic variation analysis. Many efforts have been made over the past few decades to create a complete and gapless reference genome for humans by using a diverse range of advanced sequencing technologies. Several available tools are aimed at enhancing the quality of haploid and diploid human genome assemblies, which include contig assembly, polishing of contig errors, scaffolding and variant phasing. Selecting the appropriate tools and technologies remains a daunting task despite several studies have investigated the pros and cons of different assembly strategies. The goal of this paper was to benchmark various strategies for human genome assembly by combining sequencing technologies and tools on two publicly available samples (NA12878 and NA24385) from Genome in a Bottle. We then compared their performances in terms of continuity, accuracy, completeness, variant calling and phasing. We observed that PacBio HiFi long-reads are the optimal choice for generating an assembly with low base errors. On the other hand, we were able to produce the most continuous contigs with Oxford Nanopore long-reads, but they may require further polishing to improve on quality. We recommend using short-reads rather than long-reads themselves to improve the base accuracy of contigs from Oxford Nanopore long-reads. Hi-C is the best choice for chromosome-level scaffolding because it can capture the longest-range DNA connectedness compared to 10× linked-reads and Bionano optical maps. However, a combination of multiple technologies can be used to further improve the quality and completeness of genome assembly. For diploid assembly, hifiasm is the best tool for human diploid genome assembly using PacBio HiFi and Hi-C data. Looking to the future, we expect that further advancements in human diploid assemblers will leverage the power of PacBio HiFi reads and other technologies with long-range DNA connectedness to enable the generation of high-quality, chromosome-level and haplotype-resolved human genome assemblies. Werner Pieter Veldsman, Xiaodong Fang, Yufen Huang, Xuefeng Xie, Aiping Lyu, Lu Zhang 0061 |
Briefings Bioinform. | 3 |
| 2023 | Benchmarking genome assembly methods on metagenomic sequencing dataabstractMetagenome assembly is an efficient approach to reconstruct microbial genomes from metagenomic sequencing data. Although short-read sequencing has been widely used for metagenome assembly, linked- and long-read sequencing have shown their advancements in assembly by providing long-range DNA connectedness. Many metagenome assembly tools were developed to simplify the assembly graphs and resolve the repeats in microbial genomes. However, there remains no comprehensive evaluation of metagenomic sequencing technologies, and there is a lack of practical guidance on selecting the appropriate metagenome assembly tools. This paper presents a comprehensive benchmark of 19 commonly used assembly tools applied to metagenomic sequencing datasets obtained from simulation, mock communities or human gut microbiomes. These datasets were generated using mainstream sequencing platforms, such as Illumina and BGISEQ short-read sequencing, 10x Genomics linked-read sequencing, and PacBio and Oxford Nanopore long-read sequencing. The assembly tools were extensively evaluated against many criteria, which revealed that long-read assemblers generated high contig contiguity but failed to reveal some medium- and high-quality metagenome-assembled genomes (MAGs). Linked-read assemblers obtained the highest number of overall near-complete MAGs from the human gut microbiomes. Hybrid assemblers using both short- and long-read sequencing were promising methods to improve both total assembly length and the number of near-complete MAGs. This paper also discussed the running time and peak memory consumption of these assembly tools and provided practical guidance on selecting them. Zhenmiao Zhang, Werner Pieter Veldsman, Xiaodong Fang, Lu Zhang 0061 |
Briefings Bioinform. | 4 |
| 2023 | NGenomeSyn: an easy-to-use and flexible tool for publication-ready visualization of syntenic relationships across multiple genomesabstractSUMMARY: Large-scale comparative genomic studies have provided important insights into species evolution and diversity, but also lead to a great challenge to visualize. Quick catching or presenting key information hidden in the vast amount of genomic data and relationships among multiple genomes requires an efficient visualization tool. However, current tools for such visualization remain inflexible in layout and/or require advanced computation skills, especially for visualization of genome-based synteny. Here, we developed an easy-to-use and flexible layout tool, NGenomeSyn [multiple (N) Genome Synteny], for publication-ready visualization of syntenic relationships of the whole genome or local region and genomic features (e.g. repeats, structural variations, genes) across multiple genomes with a high customization. NGenomeSyn provides an easy way for its users to visualize a large amount of data with a rich layout by simply adjusting options for moving, scaling, and rotation of target genomes. Moreover, NGenomeSyn could be applied on the visualization of relationships on non-genomic data with similar input formats. AVAILABILITY AND IMPLEMENTATION: NGenomeSyn is freely available at GitHub (https://github.com/hewm2008/NGenomeSyn) and Zenodo (https://doi.org/10.5281/zenodo.7645148). Weiming He, Yi Jing, Lian Xu, Xiaodong Fang |
Bioinform. | 6 |
| 2021 | Using Bayesian network technology to predict the semiconductor manufacturing yield rate in IoT
Xiaodong Fang, Chan Chang, Genggeng Liu |
J. Supercomput. | 1 |
| 2016 | Eliminating heterozygosity from reads through coverage normalizationabstractHeterozygosity has long plagued genome assembly. A wealth of sophisticated algorithms and procedures have been introduced for their treatment in genome assembly. In this paper, we propose a method called Hank (Heterozygosity assimilation through normalized k-mers) for this purpose. The method eliminates heterozygosity from reads by modifying the k-mers in heterozygous regions to increase their coverages to the levels of those in homozygous regions. When evaluated on simulated Illumina data at levels of heterozygosity from 0.1% to 2.0%, Hank was able to remove 80~96% of the heterozygosity. We also examined the effects of the corrections on de novo genome assembly using SOAPdenovo2, ALLPATH-LG and Platanus. All three methods improved in performance using the treated k-mers (we do not include these assembly results in this manuscript due to space constraint). Zicheng Zhao, Yen Kaow Ng, Xiaodong Fang, Shuaicheng Li 0001 |
BIBM | 3 |
| 2015 | BS-SNPer: SNP calling in bisulfite-seq dataabstractUNLABELLED: Sodium bisulfite conversion followed by sequencing (BS-Seq, such as whole genome bisulfite sequencing or reduced representation bisulfite sequencing) has become popular for studying human epigenetic profiles. Identifying single nucleotide polymorphisms (SNPs) is important for quantification of methylation levels and for study of allele-specific epigenetic events such as imprinting. However, SNP calling in such data is complex and time consuming. Here, we present an ultrafast and memory-efficient package named BS-SNPer for the exploration of SNP sites from BS-Seq data. Compared with Bis-SNP, a popular BS-Seq specific SNP caller, BS-SNPer is over 100 times faster and uses less memory. BS-SNPer also offers higher sensitivity and specificity compared with existing methods. AVAILABILITY AND IMPLEMENTATION: BS-SNPer is written in C++ and Perl, and is freely available at https://github.com/hellbelly/BS-Snper. Shengjie Gao, Dan Zou, Likai Mao, Huayu Liu, Youguo Chen, Shancen Zhao, Changduo Gao, Xiangchun Li, Zhibo Gao, Xiaodong Fang, Huanming Yang, Torben F. Ørntoft, Karina D. Sørensen, Lars Bolund |
Bioinform. | 11 |