VLDB 2026 Research / reviewers in the wild / expert
Zekun Yin
dblp:208/1964
· DBLP profile ↗
25ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0001-6002-0028ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 12 since 2021Systems, architecture and hardware · 10 · 9 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RabbitVar: Ultra-fast and accurate somatic small-variant calling on multi-core architectures
Hao Zhang 0142, Lin Gan 0001, Zekun Yin, Lifeng Yan, Honglei Song, Qixin Chang, Yanjie Wei, Beifang Niu, Bertil Schmidt |
Future Gener. Comput. Syst. | 3 |
| 2026 | Exploiting the Performance Potential of Extreme-Scale Earthquake Simulation: Achieving 86.7 PFLOPS With Over 39 Million CoresabstractLeveraging the latest Sunway supercomputer, we developed a fully optimized earthquake simulation model that accurately captures topographic effects for realistic seismic analysis. Optimizing for the SW26010Pro architecture with DMA/RMA communication mechanisms, data compression schemes, and vectorization, we achieved a speedup exceeding 160×. Our pipeline-based computation and communication overlapping scheme, combined with performance prediction models further minimized computational costs. These optimizations enabled the largest-scale curvilinear grid finite-difference method (CGFDM) earthquake simulations to date, covering 197 trillion grid points and achieving 86.7 PFLOPS on 39 million cores with a weak scaling efficiency of 97.9%. These advancements enabled the successful simulation of the 2008 Wenchuan earthquake, providing high-resolution seismic insights and robust assessments for regional hazard mitigation and disaster preparedness. Lin Gan 0008, Wubing Wan, Zekun Yin, Zhong He, Ping Gao 0005, Xiaohui Duan, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2026 | SMEStencil: Optimizing High-Order Stencils on ARM Multicore Using SME UnitabstractMatrix-accelerated stencil computation is a hot research topic, yet its application to 3 dimensional (3D) high-order stencils and HPC remains underexplored. With the emergence of Scalable Matrix Extension(SME) on ARMv9-A CPU, we analyze SME-based accelerating strategies and tailor an optimal approach for 3D high-order stencils. We introduce algorithmic optimizations based on Scalable Vector Extension(SVE) and SME unit to address strided memory accesses, alignment conflicts, and redundant accesses. We propose memory optimizations to boost on-package memory efficiency, and a novel multi-thread parallelism paradigm to overcome data-sharing challenges caused by the absence of shared data caches. SMEStencil sustains consistently high hardware utilization across diverse stencil shapes and dimensions. Our DMA-based inter-NUMA communication further mitigates NUMA effects and MPI limitations in hybrid parallelism. Combining all the innovations, SMEStencil outperforms state-of-the-art libraries on Nividia A100 GPGPU by up to 2.1× . Moreover, the performance improvements enabled by our optimizations translate directly to real-world HPC applications and enable Reverse Time Migration(RTM) real-world applications to yield 1.8x speedup versus highly-optimized Nvidia A100 GPGPU version. Tianqi Mao 0003, Lin Gan 0008, Wubing Wan, Jiayu Fu, Lanke He, Zekun Yin, Wei Xue 0003, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 9 |
| 2025 | RabbitTClust2: Fast, Scalable, and Versatile Clustering for Massive Genomic DatasetsabstractClustering is a fundamental method for extracting meaningful information from large-scale genomic datasets. As sequencing technologies advance, efficient and scalable clustering tools have become increasingly important. Despite its outstanding efficiency in large-scale genome clustering tasks, RabbitTClust still faces certain limitations. On the one hand, as data volumes continue to grow, there remains room for further optimization of its computational performance. On the other hand, RabbitTClust is not well suited for frequent incremental data updates or fast clustering across multiple thresholds. To address these limitations, we introduce RabbitTClust2, a highly efficient and versatile tool designed for clustering large-scale genomic sequences. RabbitTClust2 integrates an efficient sketching algorithm, a pruningand inverted-index-based minimum spanning tree construction method, and strategies for reusing intermediate results. With these advancements, RabbitTClust2 is able to cluster the latest RefSeq bacterial dataset (195 k genomes, 820 GB in FASTA) within 5 minutes. Compared to previous versions, RabbitTClust2 achieves a$2.4 \times$to$4.5 \times$speedup while maintaining comparable clustering accuracy, with a 21 % reduction in memory consumption. On a distributed multi-node platform, RabbitTClust2 is capable of clustering 2.6 million genomes in approximately one hour. Furthermore, RabbitTClust2 offers significant versatility by supporting efficient incremental clustering and rapid multithreshold analysis. RabbitTClust2 utilizes incremental clustering to integrate 1,000 new sequences into a pre-clustered dataset of 194,000 genomes within 1 minute, a process that results in a$22.7 \times$speedup over RabbitTClust. In addition, we used RabbitTClust2 to generate a series of clustering results with Mash distance thresholds ranging from 0.01 to 0.2 (a total of 20 values) within 7 minutes on the RefSeq bacterial dataset. The results showed that when the clustering threshold approached 0.1, the cluster compositions changed significantly, suggesting that 0.1 may represent a critical threshold for genuslevel classification in bacteria. RabbitTClust2 is available at https://github.com/RabbitBio/RabbitTClust. Xiaoming Xu 0004, Zekun Yin, Lifeng Yan, Yijie Gao, Xiaohui Duan, Bertil Schmidt |
BIBM | 3 |
| 2025 | SWBWA: A Highly Efficient NGS Aligner on the New Sunway Architecture
Lifeng Yan, Zekun Yin, Qixin Chang, Zhisong Wang, Xiaohui Duan, Bertil Schmidt |
Euro-Par (3) | 2 |
| 2025 | HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts InferenceabstractCurrent inference systems for Mixture-of-Experts (MoE) models primarily employ static parallelization strategies. However, these static approaches cannot consistently achieve optimal performance across different inference scenarios, as they lack the flexibility to adapt to varying computational requirements. In this work, we propose HAP (Hybrid Adaptive Parallelism), a novel method that dynamically selects hybrid parallel strategies to enhance MoE inference efficiency. The fundamental innovation of HAP lies in hierarchically decomposing MoE architectures into two distinct computational modules: the Attention module and the Expert module, each augmented with a specialized inference latency simulation model. This decomposition promotes the construction of a comprehensive search space for seeking model parallel strategies. By leveraging Integer Linear Programming (ILP), HAP could solve the optimal hybrid parallel configurations to maximize inference efficiency under varying computational constraints. Our experiments demonstrate that HAP consistently determines parallel configurations that achieve comparable or superior performance to the TP strategy prevalent in mainstream inference systems. Compared to the TP-based inference, HAP-based inference achieves speedups of$1.68 \times, 1.77 \times$, and$1.57 \times$on A100, A6000, and V100 GPU platforms, respectively. Furthermore, HAP showcases remarkable generalization capability, maintaining performance effectiveness across diverse MoE model configurations, including Mixtral and Qwen series models. Xianzhi Yu, Zongyuan Zhan, Wulong Liu, Zekun Yin, Xin Li 0137 |
ICPADS | 8 |
| 2025 | Trillion Ligands per Day: Performance-Portable Virtual Screening via Compound Database Optimization and Multi-Target DockingabstractStructure-based virtual screening confronts a grand challenge in scaling to trillion-ligand libraries for drug discovery. We present SWDOCKP2, a performance-portable virtual screening framework achieving 1.9 trillion ligand-receptor pairs daily across eight targets on the Sunway OceanLight supercomputer with 39-million cores — 10× faster than prior state-of-the-art. Key innovations combine (1) a ligand database optimizer with conformational sorting and merging, (2) multi-receptor grid alignment enabling parallel target screening and SIMD-accelerated trilinear interpolation, and (3) a Sunway architecture emulator for cross-platform efficiency. These advancements bridge computational scalability with novel drug discovery demands, offering a blueprint for next-generation supercomputing in structure-based drug design. Additionally, SWDOCKP2 will generate an unprecedented dataset of predicted protein-ligand interactions, creating a transformative resource for machine learning applications. By addressing experimental data scarcity, this dataset empowers accurate ligand prediction, generative chemistry, and AI-driven drug discovery. Xiaohui Duan, Gaowei Chen, Yizhen Chen, Qixin Chang, Qiancheng Xia, Zekun Yin, Lin Gan 0001, Yibing Shan, Guangwen Yang 0002, Niu Huang |
SC | 9 |
| 2025 | RabbitSketch: a high-performance sketching library for genome analysisabstractSUMMARY: We present RabbitSketch, a highly optimized library of sketching algorithms such as MinHash, OrderMinHash, and HyperLogLog that can exploit the power of modern multi-core CPUs. It provides significant speedups compared to existing implementations, ranging from 2.30× to 49.55×, as well as flexible and easy-to-use interfaces for both Python and C++. As a result, the similarity analysis of 455GB genomic data can be completed in only 5 minutes using RabbitSketch with merely 20 lines of Python code. As a case study, we enhanced RabbitTClust by integrating RabbitSketch's Kssd algorithm, resulting in a 1.54× speedup with no loss in accuracy. AVAILABILITY AND IMPLEMENTATION: RabbitSketch is available at https://github.com/RabbitBio/RabbitSketch with an archived version at Zenodo: https://doi.org/10.5281/zenodo.14903962. Detailed API documentation is available at https://rabbitsketch.readthedocs.io/en/latest. Zekun Yin, Xiaoming Xu 0004, Lifeng Yan, Fangjin Zhu, Xiaohui Duan, Bertil Schmidt |
Bioinform. | 2 |
| 2025 | SWQC: Efficient sequencing data quality control on the next-generation sunway platform
Lifeng Yan, Zekun Yin, Fangjin Zhu, Xiaohui Duan, Bertil Schmidt |
Future Gener. Comput. Syst. | 2 |
| 2025 | RabbitTrim: An Efficient and Versatile Trimmer on Multi-Core PlatformsabstractTrimming is an essential step in sequencing data processing. However, many existing trimming tools, such as Trimmomatic and Ktrim, are limited by suboptimal implementations and fail to fully leverage the computational power of modern multi-core platforms. To address this, we introduce RabbitTrim, a highly optimized and versatile trimming tool that fully supports the functionalities of Trimmomatic and Ktrim. RabbitTrim's performance is enhanced through efficient I/O strategies, parallel (de)compression engines, block-based memory pools, bitwise operations, and vectorization techniques. Compared to Trimmomatic, RabbitTrim (in trimmomatic mode) achieves speedups ranging from 1.8x to 6.0x for plain FASTQ files and 3.7x to 14.0x for gzip-compressed FASTQ files on a 48-core Intel server. Similarly, compared to Ktrim, RabbitTrim (in ktrim mode) achieves speedups ranging from 1.5x to 2.5x for plain FASTQ files and 2.7x to 5.6x for gzip-compressed FASTQ files on the same server. Moreover, RabbitTrim is able to process 101 GB gzip-compressed sequencing data in only 5 minutes while Trimmomatic requires at least 21 minutes. Zekun Yin, Lifeng Yan, Fangjin Zhu, Xin Li 0137, Xiaohui Duan, Bertil Schmidt |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2025 | RabbitBAM: Accelerating BAM File Manipulation on Multi-Core PlatformsabstractWith the continuous advancement of sequencing technology, the scale of biological data has rapidly increased. BAM format, widely used for storing aligned sequence data, is very popular due to its ease of use and good compression ratio. However, existing BAM-format file I/O libraries often fail to fully leverage the computational power of modern multi-core platforms, resulting in low CPU utilization. To address this, we introduce RabbitBAM, a fast BAM-format file I/O library. RabbitBAM employs pre-parsing and parallel parsing techniques to eliminate parsing bottlenecks and improve parallel efficiency. Additionally, we optimize multi-threaded data handling through the use of dedicated lock-free queues and memory pools. RabbitBAM achieves 2.1-3.3x speedups on next-generation sequencing data and 1-2.2x speedups on third-generation sequencing data compared to state-of-the-art SAMtools (HTSlib). We also present two case studies (BAM file quality control and sorting) using RabbitBAM, demonstrating 1.4-2.4x speedups compared to other implementations. Lifeng Yan, Zhan Zhao, Zekun Yin, Fangjin Zhu, Xiaohui Duan, Bertil Schmidt |
IEEE Trans. Comput. Biol. Bioinform. | 3 |
| 2024 | RabbitTrim: Highly Optimized Trimming of Illumina Sequencing Data on Multi-core Platforms
Zekun Yin, Lifeng Yan, Fangjin Zhu, Xiaohui Duan, Xin Li 0137, Bertil Schmidt |
ISBRA (2) | 2 |
| 2024 | RabbitSAlign: Accelerating Short-Read Alignment for CPU-GPU Heterogeneous Platforms
Lifeng Yan, Zekun Yin, Fangjin Zhu, Xiaohui Duan, Bertil Schmidt |
ISBRA (2) | 2 |
| 2024 | O2ath: an OpenMP offloading toolkit for the sunway heterogeneous manycore platform
Lifeng Yan, Qixin Chang, Haitian Lu, Chenlin Li, Quanjie He, Xiaohui Duan, Zekun Yin, Wei Xue 0003, Haohuan Fu, Lin Gan 0001, Guangwen Yang 0002 |
CCF Trans. High Perform. Comput. | 9 |
| 2023 | Leveraging Data Density and Sparsity for Efficient SVM Training on GPUsabstractSupport Vector Machines (SVMs) are a widely adopted data mining algorithm for binary and multi-class classification due to their ability to handle high-dimensional and non-linearly separable problems. However, SVM training is computationally expensive because of the heavy kernel matrix computation on large training datasets. Although much effort has been made to accelerate the training of SVMs, we find that existing libraries still suffer from inappropriate matrix multiplication methods and inefficient memory access patterns. In this paper, we propose a series of optimization approaches to address these limitations, including (i) matrix partitioning based on column density to achieve efficient kernel matrix computation; (ii) optimizing high latency memory access patterns; and (iii) dynamically selecting more suitable matrix multiplication methods based on the training dataset characteristics. Our proposed methods demonstrate significant improvements in SVM training performance without sacrificing accuracy, achieving a maximum speedup of 52x over the state-of-the-art SVMs on GPUs. These results highlight the effectiveness of our optimization in improving SVM training efficiency. Borui Xu, Zeyi Wen, Lifeng Yan, Zhan Zhao, Zekun Yin, Bingsheng He |
ICDM | 5 |
| 2023 | 69.7-PFlops Extreme Scale Earthquake Simulation with Crossing Multi-faults and Topography on SunwayabstractA high-scalable and fully optimized earthquake model is presented based on the latest Sunway supercomputer. Contributions include: 1) the curvilinear grid finite-difference method (CGFDM) and flexible model applying perfectly matched layer (PML) and enabling more accurate and realistic terrain descriptions; 2) a hybrid and non-uniform domain decomposition scheme that efficiently maps the model across different levels of the computing system; and 3) sophisticated optimizations that largely alleviate or even eliminate bottlenecks in memory, communication, etc., obtaining a speedup of over 140×. Combining all innovations, the design fully exploits the hardware potential of all aspects and enables us to perform the largest CGFDM-based earthquake simulation ever reported (69.7 PFlops using over 39 million cores). Based on our design, the Turkey earthquakes (February 6, 2023), and the Ridgecrest earthquake (July 4, 2019), are successfully simulated with a maximum resolution of 12-m. Precise hazard evaluations for the hazardous reduction of earthquake-stricken areas are also conducted. Wubing Wan, Lin Gan 0001, Zekun Yin, Haodong Tian, Mengyuan Hua, Shengye Xiang, Zhongqiu He, Ping Gao 0005, Xiaohui Duan, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002, Yaojian Chen, Xin Liu 0081, Wei Zhang 0321 |
SC | 4 |
| 2023 | RabbitKSSD: accelerating genome distance estimation on modern multi-core architecturesabstractSUMMARY: We propose RabbitKSSD, a high-speed genome distance estimation tool. Specifically, we leverage load-balanced task partitioning, fast I/O, efficient intermediate result accesses, and high-performance data structures to improve overall efficiency. Our performance evaluation demonstrates that RabbitKSSD achieves speedups ranging from 5.7× to 19.8× over Kssd for the time-consuming sketch generation and distance computation on commonly used workstations. In addition, it significantly outperforms Mash, BinDash, and Dashing2. Moreover, RabbitKSSD can efficiently perform all-vs-all distance computation for all RefSeq complete bacterial genomes (455 GB in FASTA format) in just 2 min on a 64-core workstation. AVAILABILITY AND IMPLEMENTATION: RabbitKSSD is available at https://github.com/RabbitBio/RabbitKSSD. Xiaoming Xu 0004, Zekun Yin, Lifeng Yan, Huiguang Yi, Bertil Schmidt |
Bioinform. | 2 |
| 2023 | RabbitFX: Efficient Framework for FASTA/Q File Parsing on Modern Multi-Core PlatformsabstractThe continuous growth of generated sequencing data leads to the development of a variety of associated bioinformatics tools. However, many of them are not able to fully exploit the resources of modern multi-core systems since they are bottlenecked by parsing files leading to slow execution times. This motivates the design of an efficient method for parsing sequencing data that can exploit the power of modern hardware, especially for modern CPUs with fast storage devices. We have developed RabbitFX, a fast, efficient, and easy-to-use framework for processing biological sequencing data on modern multi-core platforms. It can efficiently read FASTA and FASTQ files by combining a lightweight parsing method by means of an optimized formatting implementation. Furthermore, we provide user-friendly and modularized C++ APIs that can be easily integrated into applications in order to increase their file parsing speed. As proof-of-concept, we have integrated RabbitFX into three I/O-intensive applications: fastp, Ktrim, and Mash. Our evaluation shows that the inclusion of RabbitFX leads to speedups of at least 11.6 (6.6), 2.4 (2.4), and 3.7 (3.2) compared to the original versions on plain (gzip-compressed) files, respectively. These case studies demonstrate that RabbitFX can be easily integrated into a variety of NGS analysis tools to significantly reduce associated runtimes. It is open source software available at https://github.com/RabbitBio/RabbitFX. Hao Zhang 0142, Honglei Song, Xiaoming Xu 0004, Qixin Chang, Yanjie Wei, Zekun Yin, Bertil Schmidt |
IEEE ACM Trans. Comput. Biol. Bioinform. | 7 |
| 2022 | RabbitQCPlus: More Efficient Quality Control for Sequencing DataabstractAssessing the quality of sequencing data plays a crucial role in downstream data analysis. However, existing tools often achieve sub-optimal efficiency, especially when dealing with compressed files or performing complicated quality control operations such as over-representation analysis. We present RabbitQCPlus, an ultra-efficient quality control tool for modern multi-core systems. RabbitQCPlus uses vectorization, memory copy reduction, parallel (de)compression, and optimized data structures to achieve substantial performance gains. It is 1.1 to 5.4 times faster when performing basic quality control operations compared to state-of-the-art applications yet requires fewer compute resources. Moreover, RabbitQCPlus is at least 4 times faster than other applications when processing gzip-compressed FASTQ files. Furthermore, it takes less than 4 minutes to process 280GB of plain FASTQ sequencing data, while other applications take at least 22 minutes on a 48-core server when enabling the per-read over-representation analysis. C++ sources are available at https://github.com/RabbitBio/RabbitQCPlus. Lifeng Yan, Zekun Yin, Hao Zhang 0142, Zhan Zhao, André Müller, Robin Kobus, Yanjie Wei, Beifang Niu, Bertil Schmidt |
BIBM | 2 |
| 2022 | RabbitV: fast detection of viruses and microorganisms in sequencing data on multi-core architecturesabstractMOTIVATION: Detection and identification of viruses and microorganisms in sequencing data plays an important role in pathogen diagnosis and research. However, existing tools for this problem often suffer from high runtimes and memory consumption. RESULTS: We present RabbitV, a tool for rapid detection of viruses and microorganisms in Illumina sequencing datasets based on fast identification of unique k-mers. It can exploit the power of modern multi-core CPUs by using multi-threading, vectorization and fast data parsing. Experiments show that RabbitV outperforms fastv by a factor of at least 42.5 and 14.4 in unique k-mer generation (RabbitUniq) and pathogen identification (RabbitV), respectively. Furthermore, RabbitV is able to detect COVID-19 from 40 samples of sequencing data (255 GB in FASTQ format) in only 320 s. AVAILABILITY AND IMPLEMENTATION: RabbitUniq and RabbitV are available at https://github.com/RabbitBio/RabbitUniq and https://github.com/RabbitBio/RabbitV. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Hao Zhang 0142, Qixin Chang, Zekun Yin, Xiaoming Xu 0004, Yanjie Wei, Bertil Schmidt |
Bioinform. | 3 |
| 2021 | RabbitMash: accelerating hash-based genome analysis on modern multi-core architecturesabstractMOTIVATION: Mash is a popular hash-based genome analysis toolkit with applications to important downstream analyses tasks such as clustering and assembly. However, Mash is currently not able to fully exploit the capabilities of modern multi-core architectures, which in turn leads to high runtimes for large-scale genomic datasets. RESULTS: We present RabbitMash, an efficient highly optimized implementation of Mash which can take full advantage of modern hardware including multi-threading, vectorization and fast I/O. We show that our approach achieves speedups of at least 1.3, 9.8, 8.5 and 4.4 compared to Mash for the operations sketch, dist, triangle and screen, respectively. Furthermore, RabbitMash is able to compute the all-versus-all distances of 100 321 genomes in <5 min on a 40-core workstation while Mash requires over 40 min. AVAILABILITY AND IMPLEMENTATION: RabbitMash is available at https://github.com/ZekunYin/RabbitMash. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zekun Yin, Xiaoming Xu 0004, Jinxiao Zhang, Yanjie Wei, Bertil Schmidt |
Bioinform. | 1 |
| 2021 | RabbitQC: high-speed scalable quality control for sequencing dataabstractMOTIVATION: Modern sequencing technologies continue to revolutionize many areas of biology and medicine. Since the generated datasets are error-prone, downstream applications usually require quality control methods to pre-process FASTQ files. However, existing tools for this task are currently not able to fully exploit the capabilities of computing platforms leading to slow runtimes. RESULTS: We present RabbitQC, an extremely fast integrated quality control tool for FASTQ files, which can take full advantage of modern hardware. It includes a variety of operations and supports different sequencing technologies (Illumina, Oxford Nanopore and PacBio). RabbitQC achieves speedups between one and two orders-of-magnitude compared to other state-of-the-art tools. AVAILABILITY AND IMPLEMENTATION: C++ sources and binaries are available at https://github.com/ZekunYin/RabbitQC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zekun Yin, Hao Zhang 0142, Meiyang Liu, Honglei Song, Haidong Lan, Yanjie Wei, Beifang Niu, Bertil Schmidt |
Bioinform. | 1 |
| 2020 | SLPal: Accelerating Long Sequence Alignment on Many-Core and Multi-Core ArchitecturesabstractBiological Sequence alignment is a fundamental application in bioinformatics. It can be used to identify functionally conserved sequences and find evolutionary relationships between species. To compare entire genomes from different species, biologists increasingly need alignment methods that are efficient enough to handle long sequences, and accurate enough to correctly align the conserved biological features between distant species. Global alignments are important because they reveal the shared order of biological features in the compared species, and produce a more accurate alignment at the base-pair level when the features are in the same order. The best known global alignment algorithm is Needleman-Wunsch, later, BitPAl, a bit parallel algorithm for general, integer scoring global algorithm, provides a new implementation of Needleman-Wunsch algorithm (BitNW). Compared with original Needleman-Wunsch algorithm, BitNW is significantly faster by exploiting bit parallelism. A number of parallel strategies have been proposed to accelerate exact alignment methods. However, most of them failed to align long biological sequences due to quadratic time complexity. In this paper, we propose SLPal, a fast bit-parallel algorithm for accelerating long DNA sequence comparison on Intel manycore and multi-core architectures. In order to fully exploit the computing power of many cores and the 512-bit vector processing units (VPUs), we use a two-level parallelism scheme: coarsegrained thread level and fine-grained VPU level approaches. In thread level, the alignment scoring matrix will be split into small tiles and multiple threads will process these small tiles currently by using Intel TBB library. In the VPU level, the computing kernels are implemented using the Single Instruction Multiple Data (SIMD) instructions, thus, 16 independent integers reside in a 512-bit vector register can be processed simultaneously. The evaluation reveals that our algorithm achieves a stable performance for all benchmark data and yields a performance of up to 511.7 (617.2) GCUPS on a server with single Xeon Phi 7210 processor (dual Xeon Gold 614820-core processors). Furthermore, our test shows that SLPal can align two sequences with about 5 million bps in 50 seconds on our server equipped with dual Xeon Gold 6148 CPUs. Xiaoming Xu 0004, Yuandong Chan, Jikai Zhang, Zekun Yin |
BIBM | 6 |
| 2019 | DGCF: A Distributed Greedy Clustering Framework for Large-scale Genomic SequencesabstractClustering is a very fundamental while time-consuming compute operation in biological sequence analysis. New sequencing technologies such as NGS and 3GS have dramatically increased both the dataset size and the length of a single read sequence. However, existing tools lack scalability for handling large-scale datasets as well as long sequences. A feasible solution to this problem is to use parallel and distributed systems. The efficient deployment of such systems, however, requires high parallelism in both software implementations as well as algorithmic optimizations. In this paper, we propose DGCF, a Distributed Greedy Clustering Framework which is capable to handle large-scale datasets and long sequences. Our framework adopts a greedy clustering strategy which overlaps communication with computation among many distributed computing nodes. We also design and implement a sparse suffix array (SSA)-based alignment algorithm that can support long sequences. Experiments show that our framework achieves near-linear speedups on a distributed memory cluster. Zekun Yin, Xiaoming Xu 0004, Kaichao Fan, Weizhong Li 0002, Beifang Niu |
BIBM | 1 |
| 2017 | 18.9-Pflops nonlinear earthquake simulation on Sunway TaihuLight: enabling depiction of 18-Hz and 8-meter scenariosabstractThis paper reports our large-scale nonlinear earthquake simulation software on Sunway TaihuLight. Our innovations include: (1) a customized parallelization scheme that employs the 10 million cores efficiently at both the process and the thread levels; (2) an elaborate memory scheme that integrates on-chip halo exchange through register communcation, optimized blocking configuration guided by an analytic model, and coalesced DMA access with array fusion; (3) on-the-fly compression that doubles the maximum problem size and further improves the performance by 24%. With these innovations to remove the memory constraints of Sunway TaihuLight, our software achieves over 15% of the system's peak, better than the 11.8% efficiency achieved by a similar software running on Titan, whose byte to flop ratio is 5 times better than TaihuLight. The extreme cases demonstrate a sustained performance of over 18.9 Pflops, enabling the simulation of Tangshan earthquake as an 18-Hz scenario with an 8-meter resolution. Haohuan Fu, Conghui He, Bingwei Chen, Zekun Yin, Tingjian Zhang, Wei Xue 0003, Wanwang Yin, Guangwen Yang 0002 |
SC | 4 |