EDBT 2026 Demo / reviewers in the wild / expert
Xiaoming Xu 0004
dblp:35/4986-4
· DBLP profile ↗
8ranked-venue papers
2as first author
6since 2021 · last 2025
0000-0002-0370-2222ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | RabbitTClust2: Fast, Scalable, and Versatile Clustering for Massive Genomic DatasetsabstractClustering is a fundamental method for extracting meaningful information from large-scale genomic datasets. As sequencing technologies advance, efficient and scalable clustering tools have become increasingly important. Despite its outstanding efficiency in large-scale genome clustering tasks, RabbitTClust still faces certain limitations. On the one hand, as data volumes continue to grow, there remains room for further optimization of its computational performance. On the other hand, RabbitTClust is not well suited for frequent incremental data updates or fast clustering across multiple thresholds. To address these limitations, we introduce RabbitTClust2, a highly efficient and versatile tool designed for clustering large-scale genomic sequences. RabbitTClust2 integrates an efficient sketching algorithm, a pruningand inverted-index-based minimum spanning tree construction method, and strategies for reusing intermediate results. With these advancements, RabbitTClust2 is able to cluster the latest RefSeq bacterial dataset (195 k genomes, 820 GB in FASTA) within 5 minutes. Compared to previous versions, RabbitTClust2 achieves a$2.4 \times$to$4.5 \times$speedup while maintaining comparable clustering accuracy, with a 21 % reduction in memory consumption. On a distributed multi-node platform, RabbitTClust2 is capable of clustering 2.6 million genomes in approximately one hour. Furthermore, RabbitTClust2 offers significant versatility by supporting efficient incremental clustering and rapid multithreshold analysis. RabbitTClust2 utilizes incremental clustering to integrate 1,000 new sequences into a pre-clustered dataset of 194,000 genomes within 1 minute, a process that results in a$22.7 \times$speedup over RabbitTClust. In addition, we used RabbitTClust2 to generate a series of clustering results with Mash distance thresholds ranging from 0.01 to 0.2 (a total of 20 values) within 7 minutes on the RefSeq bacterial dataset. The results showed that when the clustering threshold approached 0.1, the cluster compositions changed significantly, suggesting that 0.1 may represent a critical threshold for genuslevel classification in bacteria. RabbitTClust2 is available at https://github.com/RabbitBio/RabbitTClust. Xiaoming Xu 0004, Zekun Yin, Lifeng Yan, Yijie Gao, Xiaohui Duan, Bertil Schmidt |
BIBM | 2 |
| 2025 | RabbitSketch: a high-performance sketching library for genome analysisabstractSUMMARY: We present RabbitSketch, a highly optimized library of sketching algorithms such as MinHash, OrderMinHash, and HyperLogLog that can exploit the power of modern multi-core CPUs. It provides significant speedups compared to existing implementations, ranging from 2.30× to 49.55×, as well as flexible and easy-to-use interfaces for both Python and C++. As a result, the similarity analysis of 455GB genomic data can be completed in only 5 minutes using RabbitSketch with merely 20 lines of Python code. As a case study, we enhanced RabbitTClust by integrating RabbitSketch's Kssd algorithm, resulting in a 1.54× speedup with no loss in accuracy. AVAILABILITY AND IMPLEMENTATION: RabbitSketch is available at https://github.com/RabbitBio/RabbitSketch with an archived version at Zenodo: https://doi.org/10.5281/zenodo.14903962. Detailed API documentation is available at https://rabbitsketch.readthedocs.io/en/latest. Zekun Yin, Xiaoming Xu 0004, Lifeng Yan, Fangjin Zhu, Xiaohui Duan, Bertil Schmidt |
Bioinform. | 3 |
| 2023 | RabbitKSSD: accelerating genome distance estimation on modern multi-core architecturesabstractSUMMARY: We propose RabbitKSSD, a high-speed genome distance estimation tool. Specifically, we leverage load-balanced task partitioning, fast I/O, efficient intermediate result accesses, and high-performance data structures to improve overall efficiency. Our performance evaluation demonstrates that RabbitKSSD achieves speedups ranging from 5.7× to 19.8× over Kssd for the time-consuming sketch generation and distance computation on commonly used workstations. In addition, it significantly outperforms Mash, BinDash, and Dashing2. Moreover, RabbitKSSD can efficiently perform all-vs-all distance computation for all RefSeq complete bacterial genomes (455 GB in FASTA format) in just 2 min on a 64-core workstation. AVAILABILITY AND IMPLEMENTATION: RabbitKSSD is available at https://github.com/RabbitBio/RabbitKSSD. Xiaoming Xu 0004, Zekun Yin, Lifeng Yan, Huiguang Yi, Bertil Schmidt |
Bioinform. | 1 |
| 2023 | RabbitFX: Efficient Framework for FASTA/Q File Parsing on Modern Multi-Core PlatformsabstractThe continuous growth of generated sequencing data leads to the development of a variety of associated bioinformatics tools. However, many of them are not able to fully exploit the resources of modern multi-core systems since they are bottlenecked by parsing files leading to slow execution times. This motivates the design of an efficient method for parsing sequencing data that can exploit the power of modern hardware, especially for modern CPUs with fast storage devices. We have developed RabbitFX, a fast, efficient, and easy-to-use framework for processing biological sequencing data on modern multi-core platforms. It can efficiently read FASTA and FASTQ files by combining a lightweight parsing method by means of an optimized formatting implementation. Furthermore, we provide user-friendly and modularized C++ APIs that can be easily integrated into applications in order to increase their file parsing speed. As proof-of-concept, we have integrated RabbitFX into three I/O-intensive applications: fastp, Ktrim, and Mash. Our evaluation shows that the inclusion of RabbitFX leads to speedups of at least 11.6 (6.6), 2.4 (2.4), and 3.7 (3.2) compared to the original versions on plain (gzip-compressed) files, respectively. These case studies demonstrate that RabbitFX can be easily integrated into a variety of NGS analysis tools to significantly reduce associated runtimes. It is open source software available at https://github.com/RabbitBio/RabbitFX. Hao Zhang 0142, Honglei Song, Xiaoming Xu 0004, Qixin Chang, Yanjie Wei, Zekun Yin, Bertil Schmidt |
IEEE ACM Trans. Comput. Biol. Bioinform. | 3 |
| 2022 | RabbitV: fast detection of viruses and microorganisms in sequencing data on multi-core architecturesabstractMOTIVATION: Detection and identification of viruses and microorganisms in sequencing data plays an important role in pathogen diagnosis and research. However, existing tools for this problem often suffer from high runtimes and memory consumption. RESULTS: We present RabbitV, a tool for rapid detection of viruses and microorganisms in Illumina sequencing datasets based on fast identification of unique k-mers. It can exploit the power of modern multi-core CPUs by using multi-threading, vectorization and fast data parsing. Experiments show that RabbitV outperforms fastv by a factor of at least 42.5 and 14.4 in unique k-mer generation (RabbitUniq) and pathogen identification (RabbitV), respectively. Furthermore, RabbitV is able to detect COVID-19 from 40 samples of sequencing data (255 GB in FASTQ format) in only 320 s. AVAILABILITY AND IMPLEMENTATION: RabbitUniq and RabbitV are available at https://github.com/RabbitBio/RabbitUniq and https://github.com/RabbitBio/RabbitV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hao Zhang 0142, Qixin Chang, Zekun Yin, Xiaoming Xu 0004, Yanjie Wei, Bertil Schmidt |
Bioinform. | 4 |
| 2021 | RabbitMash: accelerating hash-based genome analysis on modern multi-core architecturesabstractMOTIVATION: Mash is a popular hash-based genome analysis toolkit with applications to important downstream analyses tasks such as clustering and assembly. However, Mash is currently not able to fully exploit the capabilities of modern multi-core architectures, which in turn leads to high runtimes for large-scale genomic datasets. RESULTS: We present RabbitMash, an efficient highly optimized implementation of Mash which can take full advantage of modern hardware including multi-threading, vectorization and fast I/O. We show that our approach achieves speedups of at least 1.3, 9.8, 8.5 and 4.4 compared to Mash for the operations sketch, dist, triangle and screen, respectively. Furthermore, RabbitMash is able to compute the all-versus-all distances of 100 321 genomes in <5 min on a 40-core workstation while Mash requires over 40 min. AVAILABILITY AND IMPLEMENTATION: RabbitMash is available at https://github.com/ZekunYin/RabbitMash. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zekun Yin, Xiaoming Xu 0004, Jinxiao Zhang, Yanjie Wei, Bertil Schmidt |
Bioinform. | 2 |
| 2020 | SLPal: Accelerating Long Sequence Alignment on Many-Core and Multi-Core ArchitecturesabstractBiological Sequence alignment is a fundamental application in bioinformatics. It can be used to identify functionally conserved sequences and find evolutionary relationships between species. To compare entire genomes from different species, biologists increasingly need alignment methods that are efficient enough to handle long sequences, and accurate enough to correctly align the conserved biological features between distant species. Global alignments are important because they reveal the shared order of biological features in the compared species, and produce a more accurate alignment at the base-pair level when the features are in the same order. The best known global alignment algorithm is Needleman-Wunsch, later, BitPAl, a bit parallel algorithm for general, integer scoring global algorithm, provides a new implementation of Needleman-Wunsch algorithm (BitNW). Compared with original Needleman-Wunsch algorithm, BitNW is significantly faster by exploiting bit parallelism. A number of parallel strategies have been proposed to accelerate exact alignment methods. However, most of them failed to align long biological sequences due to quadratic time complexity. In this paper, we propose SLPal, a fast bit-parallel algorithm for accelerating long DNA sequence comparison on Intel manycore and multi-core architectures. In order to fully exploit the computing power of many cores and the 512-bit vector processing units (VPUs), we use a two-level parallelism scheme: coarsegrained thread level and fine-grained VPU level approaches. In thread level, the alignment scoring matrix will be split into small tiles and multiple threads will process these small tiles currently by using Intel TBB library. In the VPU level, the computing kernels are implemented using the Single Instruction Multiple Data (SIMD) instructions, thus, 16 independent integers reside in a 512-bit vector register can be processed simultaneously. The evaluation reveals that our algorithm achieves a stable performance for all benchmark data and yields a performance of up to 511.7 (617.2) GCUPS on a server with single Xeon Phi 7210 processor (dual Xeon Gold 614820-core processors). Furthermore, our test shows that SLPal can align two sequences with about 5 million bps in 50 seconds on our server equipped with dual Xeon Gold 6148 CPUs. Xiaoming Xu 0004, Yuandong Chan, Jikai Zhang, Zekun Yin |
BIBM | 1 |
| 2019 | DGCF: A Distributed Greedy Clustering Framework for Large-scale Genomic SequencesabstractClustering is a very fundamental while time-consuming compute operation in biological sequence analysis. New sequencing technologies such as NGS and 3GS have dramatically increased both the dataset size and the length of a single read sequence. However, existing tools lack scalability for handling large-scale datasets as well as long sequences. A feasible solution to this problem is to use parallel and distributed systems. The efficient deployment of such systems, however, requires high parallelism in both software implementations as well as algorithmic optimizations. In this paper, we propose DGCF, a Distributed Greedy Clustering Framework which is capable to handle large-scale datasets and long sequences. Our framework adopts a greedy clustering strategy which overlaps communication with computation among many distributed computing nodes. We also design and implement a sparse suffix array (SSA)-based alignment algorithm that can support long sequences. Experiments show that our framework achieves near-linear speedups on a distributed memory cluster. Zekun Yin, Xiaoming Xu 0004, Kaichao Fan, Weizhong Li 0002, Beifang Niu |
BIBM | 2 |