VLDB 2026 Research / reviewers in the wild / expert
Quim Aguado-Puig
dblp:303/8886
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0003-4871-3192ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
5 papers |
Bioinformatics and computational biology · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
GPUs and heterogeneous computing · 36% Hardware accelerators and domain-specific architectures · 24% Memory systems · 21% |
Topics — the 10 heaviest of 11, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
sequence alignment |
3.5 | 5 | 2026 | Singletrack: an algorithm for improving memory consumption and performance of gap-affine sequence alignment · Bioinform. 2026 QuickEd: high-performance exact sequence alignment based on bound-and-align · Bioinform. 2025 BIMSA: accelerating long sequence alignment using processing-in-memory · Bioinform. 2024 |
Bioinformatics and computational biology › sequence alignment
pairwise sequence alignment |
0.9 | 1 | 2025 | QuickEd: high-performance exact sequence alignment based on bound-and-align · Bioinform. 2025 |
Memory systems
processing-in-memory |
0.8 | 1 | 2024 | BIMSA: accelerating long sequence alignment using processing-in-memory · Bioinform. 2024 |
Bioinformatics and computational biology › sequence alignment
genomic sequence alignment |
0.7 | 1 | 2023 | GMX: Instruction Set Extensions for Fast, Scalable, and Efficient Genome Sequence Alignment · MICRO 2023 |
Hardware accelerators and domain-specific architectures › bioinformatics accelerator
genome sequence alignment accelerator |
0.7 | 1 | 2023 | GMX: Instruction Set Extensions for Fast, Scalable, and Efficient Genome Sequence Alignment · MICRO 2023 |
GPUs and heterogeneous computing
GPU-accelerated bioinformatics |
0.7 | 1 | 2023 | WFA-GPU: gap-affine pairwise read-alignment using GPUs · Bioinform. 2023 |
GPUs and heterogeneous computing
GPU computing |
0.7 | 1 | 2023 | WFA-GPU: gap-affine pairwise read-alignment using GPUs · Bioinform. 2023 |
Processor architecture and microarchitecture › instruction set architecture
instruction set extension |
0.7 | 1 | 2023 | GMX: Instruction Set Extensions for Fast, Scalable, and Efficient Genome Sequence Alignment · MICRO 2023 |
Bioinformatics and computational biology › sequence alignment
dynamic programming alignment |
0.3 | 1 | 2026 | Singletrack: an algorithm for improving memory consumption and performance of gap-affine sequence alignment · Bioinform. 2026 |
Bioinformatics and computational biology › sequence analysis › read mapping
long-read alignment |
0.2 | 1 | 2023 | WFA-GPU: gap-affine pairwise read-alignment using GPUs · Bioinform. 2023 |
Methods — techniques the papers use, named apart from their topics
dynamic programming · 3.7wavefront alignment · 2.5UPMEM · 1.5wavefront alignment algorithm · 1.3instruction set extension · 1.3backtrace algorithm · 1.0bound-and-align · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Singletrack: an algorithm for improving memory consumption and performance of gap-affine sequence alignmentabstractMOTIVATION: Advances in DNA sequencing have outpaced advances in computation, making sequence alignment a major bottleneck in genome data analyses. Classical dynamic programming (DP) algorithms are particularly memory-intensive, especially when computing gap-affine and dual gap-affine alignments. Existing strategies to reduce memory consumption often sacrifice speed or alignment accuracy. RESULTS: We present Singletrack, an efficient algorithm for backtrace gap-affine and dual gap-affine alignments that requires storing a single DP matrix while preserving optimal alignment results. Compared to classical DP algorithms, Singletrack removes the need to store additional matrices (i.e. 2 for gap-affine and 4 for dual gap-affine), significantly reducing memory consumption and, in turn, reducing pressure on the memory hierarchy and improving overall performance. Most importantly, Singletrack is a general backtrace method compatible with state-of-the-art DP-based algorithms and heuristics, such as the Suzuki-Kasahara (SK) and the Wavefront Alignment (WFA) algorithms. We demonstrate that Singletrack reduces memory consumption for both SK and WFA algorithms, lowering SK usage by 2× and 4× and WFA usage by 3× and 5× for gap-affine and dual gap-affine alignments, respectively. Moreover, replacing KSW2's memory-reduction technique with Singletrack accelerates its SK implementation by up to 1.4× at the cost of doubling memory consumption, while Singletrack increases the performance of the WFA implementation in WFA2-lib by 1.2-2.1×. Compared to the efficient linear-memory BiWFA algorithm, the Singletrack-accelerated version of WFA trades a practical increase in memory usage for up to 5.2× higher performance. AVAILABILITY AND IMPLEMENTATION: The Singletrack implementations presented in this work are available on Zenodo (DOI: 10.5281/zenodo.18770585) and GitHub (https://github.com/LorienLV/singletrack). Lorién López-Villellas, Cristian Iñiguez, Albert Jiménez-Blanco, Quim Aguado-Puig, Miquel Moretó, Jesús Alastruey-Benedé, Pablo Ibáñez 0001, Santiago Marco-Sola |
Bioinform. | 4 |
| 2025 | QuickEd: high-performance exact sequence alignment based on bound-and-alignabstractMOTIVATION: Pairwise sequence alignment is a core component of multiple sequencing-data analysis tools. Recent advancements in sequencing technologies have enabled the generation of longer sequences at a much lower price. Thus, long-read sequencing technologies have become increasingly popular in sequencing-based studies. However, classical sequence analysis algorithms face significant scalability challenges when aligning long sequences. As a result, several heuristic methods have been developed to improve performance at the expense of accuracy, as they often fail to produce the optimal alignment. RESULTS: This paper introduces QuickEd, a sequence alignment algorithm based on a bound-and-align strategy. First, QuickEd effectively bounds the maximum alignment-score using efficient heuristic strategies. Then, QuickEd utilizes this bound to reduce the computations required to produce the optimal alignment. Compared to O(n2) complexity of traditional dynamic programming algorithms, QuickEd's bound-and-align strategy achieves O(ns^) complexity, where n is the sequence length and s^ is an estimated upper bound of the alignment-score between the sequences. As a result, QuickEd is consistently faster than other state-of-the-art implementations, such as Edlib and BiWFA, achieving performance speedups of 4.2-5.9× and 3.8-4.4×, respectively, aligning long and noisy datasets. In addition, QuickEd maintains a stable memory footprint below 35 MB while aligning sequences up to 1 Mbp. AVAILABILITY AND IMPLEMENTATION: QuickEd code and documentation are publicly available at https://github.com/maxdoblas/QuickEd. Max Doblas, Oscar Lostes-Cazorla, Quim Aguado-Puig, Cristian Iñiguez, Miquel Moretó, Santiago Marco-Sola |
Bioinform. | 3 |
| 2024 | BIMSA: accelerating long sequence alignment using processing-in-memoryabstractMOTIVATION: Recent advances in sequencing technologies have stressed the critical role of sequence analysis algorithms and tools in genomics and healthcare research. In particular, sequence alignment is a fundamental building block in many sequence analysis pipelines and is frequently a performance bottleneck both in terms of execution time and memory usage. Classical sequence alignment algorithms are based on dynamic programming and often require quadratic time and memory with respect to the sequence length. As a result, classical sequence alignment algorithms fail to scale with increasing sequence lengths and quickly become memory-bound due to data-movement penalties. RESULTS: Processing-In-Memory (PIM) is an emerging architectural paradigm that seeks to accelerate memory-bound algorithms by bringing computation closer to the data to mitigate data-movement penalties. This work presents BIMSA (Bidirectional In-Memory Sequence Alignment), a PIM design and implementation for the state-of-the-art sequence alignment algorithm BiWFA (Bidirectional Wavefront Alignment), incorporating new hardware-aware optimizations for a production-ready PIM architecture (UPMEM). BIMSA supports aligning sequences up to 100K bases, exceeding the limitations of state-of-the-art PIM implementations. First, BIMSA achieves speedups up to 22.24× (11.95× on average) compared to state-of-the-art PIM-enabled implementations of sequence alignment algorithms. Second, achieves speedups up to 5.84× (2.83× on average) compared to the highest-performance multicore CPU implementation of BiWFA. Third, BIMSA exhibits linear scalability with the number of compute units in memory, enabling further performance improvements with upcoming PIM architectures equipped with more compute units and achieving speedups up to 9.56× (4.7× on average). AVAILABILITY AND IMPLEMENTATION: Code and documentation are publicly available at https://github.com/AlejandroAMarin/BIMSA. Alejandro Alonso-Marín, Ivan Fernandez, Quim Aguado-Puig, Juan Gómez-Luna, Santiago Marco-Sola, Onur Mutlu, Miquel Moretó |
Bioinform. | 3 |
| 2024 | GenArchBench: A genomics benchmark suite for arm HPC processorsabstractArm usage has substantially grown in the High-Performance Computing (HPC) community. Japanese supercomputer Fugaku, powered by Arm-based A64FX processors, held the top position on the Top500 list between June 2020 and June 2022, currently sitting in the fourth position. The recently released 7th generation of Amazon EC2 instances for compute-intensive workloads (C7 g) is also powered by Arm Graviton3 processors. Projects like European Mont-Blanc and U.S. DOE/NNSA Astra are further examples of Arm irruption in HPC. In parallel, over the last decade, the rapid improvement of genomic sequencing technologies and the exponential growth of sequencing data has placed a significant bottleneck on the computational side. While most genomics applications have been thoroughly tested and optimized for x86 systems, just a few are prepared to perform efficiently on Arm machines. Moreover, these applications do not exploit the newly introduced Scalable Vector Extensions (SVE). This paper presents GenArchBench, the first genome analysis benchmark suite targeting Arm architectures. We have selected computationally demanding kernels from the most widely used tools in genome data analysis and ported them to Arm-based A64FX and Graviton3 processors. Overall, the GenArch benchmark suite comprises 13 multi-core kernels from critical stages of widely-used genome analysis pipelines, including base-calling, read mapping, variant calling, and genome assembly. Our benchmark suite includes different input data sets per kernel (small and large), each with a corresponding regression test to verify the correctness of each execution automatically. Moreover, the porting features the usage of the novel Arm SVE instructions, algorithmic and code optimizations, and the exploitation of Arm-optimized libraries. We present the optimizations implemented in each kernel and a detailed performance evaluation and comparison of their performance on four different HPC machines (i.e., A64FX, Graviton3, Intel Xeon Skylake Platinum, and AMD EPYC Rome). Overall, the experimental evaluation shows that Graviton3 outperforms other machines on average. Moreover, we observed that the performance of the A64FX is significantly constrained by its small memory hierarchy and latencies. Additionally, as proof of concept, we study the performance of a production-ready tool that exploits two of the ported and optimized genomic kernels. Lorién López-Villellas, Rubén Langarita, Asaf Badouh, Víctor Soria 0001, Quim Aguado-Puig, Guillem López-Paradís, Max Doblas, Javier Setoain, Chulho Kim, Makoto Ono, Adrià Armejach, Santiago Marco-Sola, Jesús Alastruey-Benedé, Pablo Ibáñez 0001, Miquel Moretó |
Future Gener. Comput. Syst. | 5 |
| 2023 | GMX: Instruction Set Extensions for Fast, Scalable, and Efficient Genome Sequence AlignmentabstractSequence alignment remains a fundamental problem in computer science with practical applications ranging from pattern matching to computational biology. The ever-increasing volumes of genomic data produced by modern DNA sequencers motivate improved software and hardware sequence alignment accelerators that scale with longer sequence lengths and high error rates without losing accuracy. Furthermore, the wide variety of use cases requiring sequence alignment demands flexible and efficient solutions that can match or even outperform expensive application-specific accelerators. Max Doblas, Oscar Lostes-Cazorla, Quim Aguado-Puig, Nick Cebry, Pau Fontova, Christopher Batten, Santiago Marco-Sola, Miquel Moretó |
MICRO | 3 |
| 2023 | WFA-GPU: gap-affine pairwise read-alignment using GPUsabstractMOTIVATION: Advances in genomics and sequencing technologies demand faster and more scalable analysis methods that can process longer sequences with higher accuracy. However, classical pairwise alignment methods, based on dynamic programming (DP), impose impractical computational requirements to align long and noisy sequences like those produced by PacBio and Nanopore technologies. The recently proposed wavefront alignment (WFA) algorithm paves the way for more efficient alignment tools, improving time and memory complexity over previous methods. However, high-performance computing (HPC) platforms require efficient parallel algorithms and tools to exploit the computing resources available on modern accelerator-based architectures. RESULTS: This paper presents WFA-GPU, a GPU (graphics processing unit)-accelerated tool to compute exact gap-affine alignments based on the WFA algorithm. We present the algorithmic adaptations and performance optimizations that allow exploiting the massively parallel capabilities of modern GPU devices to accelerate the alignment computations. In particular, we propose a CPU-GPU co-design capable of performing inter-sequence and intra-sequence parallel sequence alignment, combining a succinct WFA-data representation with an efficient GPU implementation. As a result, we demonstrate that our implementation outperforms the original multi-threaded WFA implementation by up to 4.3× and up to 18.2× when using heuristic methods on long and noisy sequences. Compared to other state-of-the-art tools and libraries, the WFA-GPU is up to 29× faster than other GPU implementations and up to four orders of magnitude faster than other CPU implementations. Furthermore, WFA-GPU is the only GPU solution capable of correctly aligning long reads using a commodity GPU. AVAILABILITY AND IMPLEMENTATION: WFA-GPU code and documentation are publicly available at https://github.com/quim0/WFA-GPU. Quim Aguado-Puig, Max Doblas, Christos Matzoros, Antonio Espinosa 0001, Juan C. Moure, Santiago Marco-Sola, Miquel Moretó |
Bioinform. | 1 |
| 2021 | OpenCL-based FPGA Accelerator for Semi-Global Approximate String Matching Using Diagonal Bit-VectorsabstractAn FPGA accelerator for the computation of the semi-global Levenshtein distance between a pattern and a reference text is presented. The accelerator provides an important benefit to reduce the execution time of read-mappers used in short-read genomic sequencing. Previous attempts to solve the same problem in FPGA use the Myers algorithm following a column approach to compute the dynamic programming table. We use an approach based on diagonals that allows for some resource savings while maintaining a very high throughput of 1 alignment per clock cycle. The design is implemented in OpenCL and tested on two FPGA accelerators. The maximum performance obtained is 91.5 MPairs/s for 100 × 120 sequences and 47 MPairs/s for 300 × 360 sequences, the highest ever reported for this problem. David Castells-Rufas, Santiago Marco-Sola, Quim Aguado-Puig, Antonio Espinosa 0001, Juan C. Moure, Lluc Alvarez, Miquel Moretó |
FPL | 3 |