Alberto Zeni

dblp:243/1116 · DBLP profile ↗
← Back
9ranked-venue papers
7as first author
7since 2021 · last 2024
0000-0003-4005-6036ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 7 first-author · 7 since 2021
YearPublicationVenuePosition
2024 Leveraging Difference Recurrence Relations for High-Performance GPU Genome Alignment
abstract
Genome pairwise sequence alignment is one of the most computationally intensive workloads in many genomic pipelines, often accounting for over 90% of the runtime of critical bioinformatics applications. Recent advancements in sequencing technologies keep increasing the throughput of genomic sequencing data while decreasing the associated cost, emphasizing the need for fast and accurate software to perform sequence analysis, given the quadratic complexity of exact pairwise algorithms. In this challenging scenario, we present the first fully GPU-accelerated version of the KSW2 genome alignment library. Results show that our high-performance implementation achieves up to 1145.17 Giga Cell Updates Per Second (GCUPS) and speedups up to 72.83 × on a single NVIDIA Tesla H100 over the state-of-the-art baseline software running on two Intel Xeon Platinum 8358 processors with a total of 128 CPU threads, while preserving alignment accuracy. Using the same configuration, we demonstrate a 66.00 × speedup, versus ksw2d-fast, a state-of-the-art improved version of one of the KSW2 algorithms. Furthermore, we compare our implementation against a recently proposed FPGA implementation of ksw2z, achieving speedups up to 156.37 × using a single H100 GPU. To further highlight the impact of our work, we integrate our accelerated kernels within one of the most used aligners and mappers in the State Of the Art, called minimap2, demonstrating runtime improvements by up to 8.51 × and 8.03 × using a single H100 GPU against the baseline software and mm2-fast, an optimized version of minimap2 which integrates ksw2d-fast as its core aligner. Our design accelerates all the algorithms of the state-of-the-art KSW2 aligner suite (splice, double- and single- gap affine) and supports the Z-drop heuristic and banded alignment as the original software to reduce the processing time further if needed. Finally, we evaluate our application on the H100 GPU, adapting the Berkeley Roofline model for KSW2 and demonstrating that our implementation is near optimal on our target GPU architecture.
Alberto Zeni, Seth Onken, Marco D. Santambrogio, Mehrzad Samadi
PACT1
2024 Starlight: A kernel optimizer for GPU processing
abstract
Over the past few years, GPUs have found widespread adoption in many scientific domains, offering notable performance and energy efficiency advantages compared to CPUs. However, optimizing GPU high-performance kernels poses challenges given the complexities of GPU architectures and programming models. Moreover, current GPU development tools provide few high-level suggestions and overlook the underlying hardware. Here we present Starlight, an open-source, highly flexible tool for enhancing GPU kernel analysis and optimization. Starlight autonomously describes Roofline Models, examines performance metrics, and correlates these insights with GPU architectural bottlenecks. Additionally, Starlight predicts potential performance enhancements before altering the source code. We demonstrate its efficacy by applying it to literature genomics and physics applications, attaining speedups from 1.1× to 2.5× over state-of-the-art baselines. Furthermore, Starlight supports the development of new GPU kernels, which we exemplify through an image processing application, showing speedups of 12.7× and 140× when compared against state-of-the-art FPGA- and GPU-based solutions.
Alberto Zeni, Emanuele Del Sozzo, Eleonora D'Arnese, Davide Conficconi, Marco D. Santambrogio
J. Parallel Distributed Comput.1
2023 TSUNAMI: A GPU Implementation of the WFA Algorithm
abstract
Pairwise sequence alignment represents a fundamental step in the genome assembly pipeline, being the most time-consuming step and the bottleneck factor in multiple bioinformatics applications. Exact pairwise alignment methods like Smith-Waterman and Needleman-Wunsch, often cannot satisfy the performance required by these tools because of their quadratic time complexity. Furthermore, given the increasing computational cost of analyzing third-generation sequences, the community is moving towards different alignment methods and hardware-accelerated solutions to overcome the limitations of these algorithms. In this scenario, we present TSUNAMI, a highly-optimized implementation of the WaveFront Alignment (WFA) algorithm on GPU. TSUNAMI exploits GPU high-parallel computing to accelerate the WFA algorithm, a novel alignment methodology exploiting homologous regions between the target sequences. By doing so, we are able to reduce both time and space complexity in our GPU implementation. Our results show that TSUNAMI achieves improvements up to 4512.28× in terms of speedup when compared to the multi-threaded state-of-the-art software implementation run on Intel Xeon Silver 4208 using 16 threads in total. We also compared our design with all the recently released hardware-accelerated solutions present in the State Of the Art, observing speedups up to 14.81×with respect to the best performing hardware-accelerated implementation in the literature, reaching up to 42604.98 Giga Cell Updates Per Second in our best configuration. TSUNAMI also supports aligning very erroneous long sequences, rendering our implementation much more useful in real-world scenarios. Finally, to prove the efficiency of our design, we evaluate TSUNAMI exploiting the Berkeley Roofline model and demonstrate that our implementation is near-optimal on the NVIDIA Tesla H100.
Giulia Gerometta, Alberto Zeni, Marco D. Santambrogio
PACT2
2023 On the Genome Sequence Alignment FPGA Acceleration via KSW2z
abstract
Pairwise sequence alignment is a fundamental step for many genomics and molecular biology applications. Given the quadratic time complexity of alignment algorithms, the community demands innovative, fast, and efficient techniques to perform this task. Furthermore, general-purpose architectures lack the necessary performance to address the computational load of these algorithms. In this context, we present the first open-source FPGA implementation of the popular KSW2z algorithm employed by minimap2. Our design also implements the$Z- \mathbf{drop}$heuristic and banded alignment as the original software to further reduce the processing time if needed. The proposed multi-core accelerator achieves up to$\mathbf{7.70}\times$improvement in speedup and$\mathbf{20.07}\times$in energy efficiency compared to the multi-threaded software implementation run on a Xeon Platinum 8167M processor.
Alberto Zeni, Guido Walter Di Donato, Alessia Della Valle, Filippo Carloni, Marco D. Santambrogio
ISCAS1
2022 Surfing the Wavefront of Genome Alignment
abstract
Pairwise sequence alignment represents a fundamental step in genome and molecular analysis applications, accounting for most of their runtime. Given the quadratic time complexity of alignment algorithms, the community presses for the development of more efficient algorithms. Moreover, current limitations of general-purpose architectures push users to use hardware accelerators to reduce the analysis time. In this context, we present an FPGA implementation of the Wavefront Alignment (WFA) algorithm, a recently introduced solution that exploits homologous regions between the sequences to speed up the alignment process and whose complexity is related to the score of the alignment, rather than to the lengths of the sequences. Our multicore design can achieve up to 8.09 × improvement in speedup and 57.77 × in energy efficiency compared to the multithreaded software implementation run on a Xeon Gold Processor. Moreover, our design highly outperforms the current State-of-the-Art hardware-accelerated solution, reaching up to 2876 Giga Cell Updates Per Second (GCUPS) and 68.47 GCUPS/W on a single FPGA, with an improvement of up to 2.29× and 9.90× in terms of performance and energy efficiency, respectively.
Beatrice Branchini, Giulia Gerometta, Luisa Cicolini, Alberto Zeni, Emanuele Del Sozzo, Marco D. Santambrogio
ISCAS4
2021 Optimized Implementation of the HPCG Benchmark on Reconfigurable Hardware
Alberto Zeni, Kenneth O'Brien, Michaela Blott, Marco D. Santambrogio
Euro-Par1
2021 The Importance of Being X-Drop: High Performance Genome Alignment on Reconfigurable Hardware
abstract
Pairwise sequence alignment accounts for the majority of key genome analysis applications' runtime. Because of the quadratic time complexity of exact alignment algorithms, the community is moving away from exact algorithms in favor of heuristics that only compute high-quality results. However, the state of the art lacks hardware-accelerated versions of these heuristic algorithms as the vast majority of the available solutions still rely on implementing exact alignment algorithms. Moreover, hardware-based implementations lack high-level APIs that can simplify their integration in commonly used genomic pipelines, hindering their applicability in real-world scenarios. In this context, we present the first high-performance FPGA implementation of the popular X-drop heuristic alignment algorithm and provide an easy-to-use API for its integration. On a Xilinx Alveo U280, our FPGA design achieves up to 5× speed-up over SeqAn, the state-of-the-art software version of the algorithm, running on two Intel Xeon processors using 80 CPU threads. Moreover, our design is also 3.45× faster than ksw2, a state-of-the-art vectorized alignment algorithm that performs a similar heuristic to the one employed in the X-drop algorithm. Finally, our implementation also outperforms LOGAN, a recently published GPU implementation of X-drop running on an Nvidia Tesla V100, by a factor of 1.5×.
Alberto Zeni, Guido Walter Di Donato, Lorenzo Di Tucci, Marco Rabozzi, Marco D. Santambrogio
FCCM1
2020 LOGAN: High-Performance GPU-Based X-Drop Long-Read Alignment
abstract
Pairwise sequence alignment is one of the most computationally intensive kernels in genomic data analysis, accounting for more than 90% of the runtime for key bioinformatics applications. This method is particularly expensive for third-generation sequences due to the high computational cost of analyzing sequences of length between 1Kb and 1Mb. Given the quadratic overhead of exact pairwise algorithms for long alignments, the community primarily relies on approximate algorithms that search only for high-quality alignments and stop early when one is not found. In this work, we present the first GPU optimization of the popular X-drop alignment algorithm, that we named LOGAN. Results show that our high-performance multi-GPU implementation achieves up to 181.6 GCUPS and speed-ups up to 6.6× and 30.7× using 1 and 6 NVIDIA Tesla V100, respectively, over the state-of-the-art software running on two IBM Power9 processors using 168 CPU threads, with equivalent accuracy. We also demonstrate a 2.3× LOGAN speed-up versus ksw2, a state-of-art vectorized algorithm for sequence alignment implemented in minimap2, a long-read mapping software. To highlight the impact of our work on a real-world application, we couple LOGAN with a many-to-many long-read alignment software called BELLA, and demonstrate that our implementation improves the overall BELLA runtime by up to 10.6×. Finally, we adapt the Roofline model for LOGAN and demonstrate that our implementation is near optimal on the NVIDIA Tesla V100s.
Alberto Zeni, Giulia Guidi, Marquita Ellis, Nan Ding 0006, Marco D. Santambrogio, Steven Hofmeyr, Aydin Buluç, Leonid Oliker, Katherine A. Yelick
IPDPS1
2019 An FPGA-Based Computing Infrastructure Tailored to Efficiently Scaffold Genome Sequences
abstract
In the current years broad access to genomic data is leading to improve the understanding and prevention of human diseases as never before. De-novo genome assembly, represents a main obstacle to perform the analysis on a large scale, as it is one of the most time-consuming phases of the genome analysis. In this paper, we present a scalable, high performance and energy efficient architecture for the alignment step of SSPACE, a state of the art tool used to perform scaffolding also in case of de-novo assembly. The final architecture is able to achieve up to 9.83x speedup in performance when compared to the software version of Bowtie, a state of the art tool used by SSPACE to perform the alignment.
Alberto Zeni, Matteo Crespi, Lorenzo Di Tucci, Marco D. Santambrogio
FCCM1