EDBT 2026 Demo / reviewers in the wild / expert
Lifeng Yan
dblp:183/6400
· DBLP profile ↗
18ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 8 since 2021Systems, architecture and hardware · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Repurposing the Cross-Segment Space on Sunway SW26010Pro for MC-Balanced Bigshare Execution
Qixin Chang, Lifeng Yan, Hailong Liu 0007, Xiaohui Duan |
Euro-Par (1) | 2 |
| 2026 | A Miniaturized Wireless Multimodal Physiological In-Vivo Monitoring Platform Featuring Power-Efficient Photoelectrochemical Sensing
Zepeng Huang, Lifeng Yan, Wenxian Gu, Xing Wu 0005, Liangjian Lyu |
ISCAS | 2 |
| 2026 | A Neural Spike Sorting Framework with Multi-Scale Slope Detection and Lite-CNN Classification
Yulun Peng, Wenxian Gu, Jingjie Tang, Lifeng Yan, Xing Wu 0005, Liangjian Lyu |
ISCAS | 4 |
| 2026 | RabbitVar: Ultra-fast and accurate somatic small-variant calling on multi-core architectures
Hao Zhang 0142, Lin Gan 0001, Zekun Yin, Lifeng Yan, Honglei Song, Qixin Chang, Yanjie Wei, Beifang Niu, Bertil Schmidt |
Future Gener. Comput. Syst. | 4 |
| 2025 | RabbitTClust2: Fast, Scalable, and Versatile Clustering for Massive Genomic DatasetsabstractClustering is a fundamental method for extracting meaningful information from large-scale genomic datasets. As sequencing technologies advance, efficient and scalable clustering tools have become increasingly important. Despite its outstanding efficiency in large-scale genome clustering tasks, RabbitTClust still faces certain limitations. On the one hand, as data volumes continue to grow, there remains room for further optimization of its computational performance. On the other hand, RabbitTClust is not well suited for frequent incremental data updates or fast clustering across multiple thresholds. To address these limitations, we introduce RabbitTClust2, a highly efficient and versatile tool designed for clustering large-scale genomic sequences. RabbitTClust2 integrates an efficient sketching algorithm, a pruningand inverted-index-based minimum spanning tree construction method, and strategies for reusing intermediate results. With these advancements, RabbitTClust2 is able to cluster the latest RefSeq bacterial dataset (195 k genomes, 820 GB in FASTA) within 5 minutes. Compared to previous versions, RabbitTClust2 achieves a$2.4 \times$to$4.5 \times$speedup while maintaining comparable clustering accuracy, with a 21 % reduction in memory consumption. On a distributed multi-node platform, RabbitTClust2 is capable of clustering 2.6 million genomes in approximately one hour. Furthermore, RabbitTClust2 offers significant versatility by supporting efficient incremental clustering and rapid multithreshold analysis. RabbitTClust2 utilizes incremental clustering to integrate 1,000 new sequences into a pre-clustered dataset of 194,000 genomes within 1 minute, a process that results in a$22.7 \times$speedup over RabbitTClust. In addition, we used RabbitTClust2 to generate a series of clustering results with Mash distance thresholds ranging from 0.01 to 0.2 (a total of 20 values) within 7 minutes on the RefSeq bacterial dataset. The results showed that when the clustering threshold approached 0.1, the cluster compositions changed significantly, suggesting that 0.1 may represent a critical threshold for genuslevel classification in bacteria. RabbitTClust2 is available at https://github.com/RabbitBio/RabbitTClust. Xiaoming Xu 0004, Zekun Yin, Lifeng Yan, Yijie Gao, Xiaohui Duan, Bertil Schmidt |
BIBM | 4 |
| 2025 | SWBWA: A Highly Efficient NGS Aligner on the New Sunway Architecture
Lifeng Yan, Zekun Yin, Qixin Chang, Zhisong Wang, Xiaohui Duan, Bertil Schmidt |
Euro-Par (3) | 1 |
| 2025 | RabbitSketch: a high-performance sketching library for genome analysisabstractSUMMARY: We present RabbitSketch, a highly optimized library of sketching algorithms such as MinHash, OrderMinHash, and HyperLogLog that can exploit the power of modern multi-core CPUs. It provides significant speedups compared to existing implementations, ranging from 2.30× to 49.55×, as well as flexible and easy-to-use interfaces for both Python and C++. As a result, the similarity analysis of 455GB genomic data can be completed in only 5 minutes using RabbitSketch with merely 20 lines of Python code. As a case study, we enhanced RabbitTClust by integrating RabbitSketch's Kssd algorithm, resulting in a 1.54× speedup with no loss in accuracy. AVAILABILITY AND IMPLEMENTATION: RabbitSketch is available at https://github.com/RabbitBio/RabbitSketch with an archived version at Zenodo: https://doi.org/10.5281/zenodo.14903962. Detailed API documentation is available at https://rabbitsketch.readthedocs.io/en/latest. Zekun Yin, Xiaoming Xu 0004, Lifeng Yan, Fangjin Zhu, Xiaohui Duan, Bertil Schmidt |
Bioinform. | 4 |
| 2025 | SWQC: Efficient sequencing data quality control on the next-generation sunway platform
Lifeng Yan, Zekun Yin, Fangjin Zhu, Xiaohui Duan, Bertil Schmidt |
Future Gener. Comput. Syst. | 1 |
| 2025 | RabbitTrim: An Efficient and Versatile Trimmer on Multi-Core PlatformsabstractTrimming is an essential step in sequencing data processing. However, many existing trimming tools, such as Trimmomatic and Ktrim, are limited by suboptimal implementations and fail to fully leverage the computational power of modern multi-core platforms. To address this, we introduce RabbitTrim, a highly optimized and versatile trimming tool that fully supports the functionalities of Trimmomatic and Ktrim. RabbitTrim's performance is enhanced through efficient I/O strategies, parallel (de)compression engines, block-based memory pools, bitwise operations, and vectorization techniques. Compared to Trimmomatic, RabbitTrim (in trimmomatic mode) achieves speedups ranging from 1.8x to 6.0x for plain FASTQ files and 3.7x to 14.0x for gzip-compressed FASTQ files on a 48-core Intel server. Similarly, compared to Ktrim, RabbitTrim (in ktrim mode) achieves speedups ranging from 1.5x to 2.5x for plain FASTQ files and 2.7x to 5.6x for gzip-compressed FASTQ files on the same server. Moreover, RabbitTrim is able to process 101 GB gzip-compressed sequencing data in only 5 minutes while Trimmomatic requires at least 21 minutes. Zekun Yin, Lifeng Yan, Fangjin Zhu, Xin Li 0137, Xiaohui Duan, Bertil Schmidt |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | RabbitBAM: Accelerating BAM File Manipulation on Multi-Core PlatformsabstractWith the continuous advancement of sequencing technology, the scale of biological data has rapidly increased. BAM format, widely used for storing aligned sequence data, is very popular due to its ease of use and good compression ratio. However, existing BAM-format file I/O libraries often fail to fully leverage the computational power of modern multi-core platforms, resulting in low CPU utilization. To address this, we introduce RabbitBAM, a fast BAM-format file I/O library. RabbitBAM employs pre-parsing and parallel parsing techniques to eliminate parsing bottlenecks and improve parallel efficiency. Additionally, we optimize multi-threaded data handling through the use of dedicated lock-free queues and memory pools. RabbitBAM achieves 2.1-3.3x speedups on next-generation sequencing data and 1-2.2x speedups on third-generation sequencing data compared to state-of-the-art SAMtools (HTSlib). We also present two case studies (BAM file quality control and sorting) using RabbitBAM, demonstrating 1.4-2.4x speedups compared to other implementations. Lifeng Yan, Zhan Zhao, Zekun Yin, Fangjin Zhu, Xiaohui Duan, Bertil Schmidt |
IEEE Trans. Comput. Biol. Bioinform. | 1 |
| 2024 | RabbitTrim: Highly Optimized Trimming of Illumina Sequencing Data on Multi-core Platforms
Zekun Yin, Lifeng Yan, Fangjin Zhu, Xiaohui Duan, Xin Li 0137, Bertil Schmidt |
ISBRA (2) | 3 |
| 2024 | RabbitSAlign: Accelerating Short-Read Alignment for CPU-GPU Heterogeneous Platforms
Lifeng Yan, Zekun Yin, Fangjin Zhu, Xiaohui Duan, Bertil Schmidt |
ISBRA (2) | 1 |
| 2024 | O2ath: an OpenMP offloading toolkit for the sunway heterogeneous manycore platform
Lifeng Yan, Qixin Chang, Haitian Lu, Chenlin Li, Quanjie He, Xiaohui Duan, Zekun Yin, Wei Xue 0003, Haohuan Fu, Lin Gan 0001, Guangwen Yang 0002 |
CCF Trans. High Perform. Comput. | 2 |
| 2023 | Leveraging Data Density and Sparsity for Efficient SVM Training on GPUsabstractSupport Vector Machines (SVMs) are a widely adopted data mining algorithm for binary and multi-class classification due to their ability to handle high-dimensional and non-linearly separable problems. However, SVM training is computationally expensive because of the heavy kernel matrix computation on large training datasets. Although much effort has been made to accelerate the training of SVMs, we find that existing libraries still suffer from inappropriate matrix multiplication methods and inefficient memory access patterns. In this paper, we propose a series of optimization approaches to address these limitations, including (i) matrix partitioning based on column density to achieve efficient kernel matrix computation; (ii) optimizing high latency memory access patterns; and (iii) dynamically selecting more suitable matrix multiplication methods based on the training dataset characteristics. Our proposed methods demonstrate significant improvements in SVM training performance without sacrificing accuracy, achieving a maximum speedup of 52x over the state-of-the-art SVMs on GPUs. These results highlight the effectiveness of our optimization in improving SVM training efficiency. Borui Xu, Zeyi Wen, Lifeng Yan, Zhan Zhao, Zekun Yin, Bingsheng He |
ICDM | 3 |
| 2023 | RabbitKSSD: accelerating genome distance estimation on modern multi-core architecturesabstractSUMMARY: We propose RabbitKSSD, a high-speed genome distance estimation tool. Specifically, we leverage load-balanced task partitioning, fast I/O, efficient intermediate result accesses, and high-performance data structures to improve overall efficiency. Our performance evaluation demonstrates that RabbitKSSD achieves speedups ranging from 5.7× to 19.8× over Kssd for the time-consuming sketch generation and distance computation on commonly used workstations. In addition, it significantly outperforms Mash, BinDash, and Dashing2. Moreover, RabbitKSSD can efficiently perform all-vs-all distance computation for all RefSeq complete bacterial genomes (455 GB in FASTA format) in just 2 min on a 64-core workstation. AVAILABILITY AND IMPLEMENTATION: RabbitKSSD is available at https://github.com/RabbitBio/RabbitKSSD. Xiaoming Xu 0004, Zekun Yin, Lifeng Yan, Huiguang Yi, Bertil Schmidt |
Bioinform. | 3 |
| 2022 | RabbitQCPlus: More Efficient Quality Control for Sequencing DataabstractAssessing the quality of sequencing data plays a crucial role in downstream data analysis. However, existing tools often achieve sub-optimal efficiency, especially when dealing with compressed files or performing complicated quality control operations such as over-representation analysis. We present RabbitQCPlus, an ultra-efficient quality control tool for modern multi-core systems. RabbitQCPlus uses vectorization, memory copy reduction, parallel (de)compression, and optimized data structures to achieve substantial performance gains. It is 1.1 to 5.4 times faster when performing basic quality control operations compared to state-of-the-art applications yet requires fewer compute resources. Moreover, RabbitQCPlus is at least 4 times faster than other applications when processing gzip-compressed FASTQ files. Furthermore, it takes less than 4 minutes to process 280GB of plain FASTQ sequencing data, while other applications take at least 22 minutes on a 48-core server when enabling the per-read over-representation analysis. C++ sources are available at https://github.com/RabbitBio/RabbitQCPlus. Lifeng Yan, Zekun Yin, Hao Zhang 0142, Zhan Zhao, André Müller, Robin Kobus, Yanjie Wei, Beifang Niu, Bertil Schmidt |
BIBM | 1 |
| 2020 | An End-To-End Network For Detecting Multi-Domain Fractures On X-Ray ImagesabstractAutomated fracture detection on medical images is a crucial prerequisite for orthopedic diagnosis. However, due to the considerable variation of bone structures, it is challenging to detect fractures on images filmed from various body parts utilizing a single model. In this paper, we treat each body part as a domain and propose a novel Multi-domain Fracture Detection Network (MFDN), which is composed of two sub-networks, namely, a domain classification network for predicting the domain type of an image and a fracture detection network for detecting fractures on X-ray images of different domains. By constructing Feature Enhancement Modules and Multi-Feature-Enhanced R-CNN, the proposed MFDN extracts better feature representations for each domain. Experimental results on real-world datasets show the effectiveness of our model which has been used in clinical diagnosis with the best performance on all the domains. Shukai Wu, Lifeng Yan, Yizhou Yu, Sanyuan Zhang |
ICIP | 2 |
| 2018 | Joint Euclidean and Angular Distance-Based Embeddings for Multisource Image AnalysisabstractWith the emergence of passive and active optical sensors available for geospatial imaging, information fusion across sensors is becoming ever more important. An important aspect of single (or multiple) sensor geospatial image analysis is feature extraction-the process of finding “optimal” lower dimensional subspaces that adequately characterize class-specific information for subsequent analysis tasks, such as classification, change and anomaly detection, and so on. In recent work, we proposed and developed an angle-based discriminant analysis approach that projected data onto the subspaces with maximal “angular” separability in the input (raw) feature space and reproducing kernel Hilbert space. We also developed an angular locality preserving variant of this algorithm. Despite being a promising approach, the resulting subspace does not preserve Euclidean distance information. In this letter, we advance this work to address that limitation and make it suitable for information fusion-we propose and validate a composite kernel-based subspace learning framework that simultaneously preserves Euclidean and angular information, which can operate on an ensemble of feature sources (e.g., from different sources). We validate this method with the multisensor University of Houston hyperspectral and light detection and ranging data set, and demonstrate that a joint discriminant analysis that leverages angular and Euclidean distance information provides superior classification and sensor (information) fusion performance. Lifeng Yan, Minshan Cui, Saurabh Prasad |
IEEE Geosci. Remote. Sens. Lett. | 1 |