EDBT 2026 Demo / reviewers in the wild / expert
Beatrice Branchini
dblp:326/0863
· DBLP profile ↗
6ranked-venue papers
2as first author
6since 2021 · last 2026
0000-0001-9829-397XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Locality-Aware Distributed Allocators for High-Performance Global Data Structures
Ian Di Dio Lavore, Beatrice Branchini, Vito Giovanni Castellana, Marco D. Santambrogio |
IPDPS | 2 |
| 2025 | Harnessing GPU Acceleration for Exact DNA Sequence Matching via the KMP AlgorithmabstractIdentifying recurrent patterns and mutations in the DNA is essential for helping clinicians formulate faster diagnoses and develop personalized treatments. Here, exact matching represents a core procedure. However, due to its computational intensity, it embodies a bottleneck in many genome analysis pipelines. In this context, the computational efficiency of the chosen algorithm used for exact matching combined with the GPUs’ computing power is crucial to speeding up the process. This paper introduces a high-performance multi-GPU solution for exact DNA sequence matching based on the Knuth-Morris-Pratt (KMP) algorithm, designed to identify all possible occurrences of patterns within a reference DNA. Our solution offers a multi-pattern search that exploits multi-buffering to maximize data reuse and overlap data transfers with computation. Experimental results show that our approach for the KMP algorithm attains 5.83× over a state-of-the-art software for genome analysis. Also, our solution run on an AMD MI210 attains up to 1.59 over the best-performing FPGA solution in the literature. × Beatrice Branchini, Pierluigi Negro, Ian Di Dio Lavore, Marco D. Santambrogio |
ISCAS | 1 |
| 2025 | On the Characterization of GraphML Frameworks: The Case of Semi-Supervised Node ClassificationabstractIn recent years, the application of Machine Learning techniques on graphs has produced a considerable interest, leading to the development of many Graph Machine Learning (GraphML) frameworks. However, the proper framework has to be selected depending on the application, requiring time and resources. To solve this issue, this work characterizes three GraphML frameworks: PyTorch Geometric (PyG), Deep Graph Library (DGL), and Stellargraph on four different GPU architectures on the task of Semi-Supervised Node Classification. We compare both the training and inference time and the accuracy and loss curves for each configuration under identical model setups. Results show that PyTorch-based frameworks are faster than those using TensorFlow. Furthermore, we evaluate how DGL has a steeper convergence and can outperform PyG in the presented case study, while the training time per epoch of PyG is faster. Additionally, our evaluation highlights how the frameworks only sometimes fully exploit newer generations of server-grade GPUs. This study demonstrates how selecting the most suitable GraphML framework is a multifaced problem that can directly impact the performance for the end-user. Alessandro La Conca, Leonardo De Grandis, Ian Di Dio Lavore, Beatrice Branchini, Marco D. Santambrogio |
ISCAS | 4 |
| 2025 | QUEKUF: An FPGA Union Find Decoder for Quantum Error Correction on the Toric CodeabstractQuantum computing represents an exciting computing paradigm that promises to solve problems untractable for a classical computer. The main limiting factor for quantum devices is the noise impacting qubits, which hinders the superpolynomial speedup promise. Thus, although Quantum Error Correction (QEC) mechanisms are paramount, QEC demands high speed and low latency to scale quantum computations to real-life-sized problems. Within this context, hardware accelerators, such as Field Programmable Gate Arrays (FPGAs), represent a valuable approach to fulfilling QEC requirements. Nevertheless, the literature falls short in proposing solutions targeting the toric code, a type of quantum Low-Density Parity Check code capable of encoding two logical qubits, thus requiring fewer physical qubits. This manuscript presents QUEKUF , an FPGA-based QEC dataflow architecture dealing with the toric code. QUEKUF disposes of parallel processing units to spatially parallelize QEC, which a centralized controller orchestrates for data movement and operation decisions. We also provide a latency-oriented resource optimization model to identify the best theoretical configuration of QUEKUF that minimizes latency and optimizes resource requirements based upon high-level quantum parameters. Experimental results show that QUEKUF attains up to \(7.30\times\) speedup and \(81.51\times\) improvement in energy efficiency over a C++ implementation with error-free syndromes while keeping high accuracy. Federico Valentino, Beatrice Branchini, Davide Conficconi, Donatella Sciuto, Marco D. Santambrogio |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2025 | Rock the QASBA: Quantum Error Correction Acceleration via the Sparse Blossom Algorithm on FPGAsabstractQuantum computing is a new paradigm of computation that exploits principles from quantum mechanics to achieve an exponential speedup compared to classical logic. However, noise strongly limits current quantum hardware, reducing achievable performance and limiting the scaling of the applications. For this reason, current noisy intermediate-scale quantum devices require Quantum Error Correction (QEC) mechanisms to identify errors occurring in the computation and correct them in real time. Nevertheless, the high computational complexity of QEC algorithms is incompatible with the tight time constraints of quantum devices. Thus, hardware acceleration is paramount to achieving real-time QEC. This work presents QASBA, an FPGA-based hardware accelerator for the Sparse Blossom Algorithm (SBA), a state-of-the-art decoding algorithm. After profiling the state-of-the-art software counterpart, we developed a design methodology for hardware development based on the SBA. We also devised an automation process to help users without expertise in hardware design in deploying architectures based on QASBA. We implement QASBA on different FPGA architectures and experimentally evaluate resource usage, execution time, and energy efficiency of our solution. Our solution attains up to \(25.05\times\) speedup and \(304.16\times\) improvement in energy efficiency compared to the software baseline. Marco Venere, Beatrice Branchini, Davide Conficconi, Donatella Sciuto, Marco D. Santambrogio |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2022 | Surfing the Wavefront of Genome AlignmentabstractPairwise sequence alignment represents a fundamental step in genome and molecular analysis applications, accounting for most of their runtime. Given the quadratic time complexity of alignment algorithms, the community presses for the development of more efficient algorithms. Moreover, current limitations of general-purpose architectures push users to use hardware accelerators to reduce the analysis time. In this context, we present an FPGA implementation of the Wavefront Alignment (WFA) algorithm, a recently introduced solution that exploits homologous regions between the sequences to speed up the alignment process and whose complexity is related to the score of the alignment, rather than to the lengths of the sequences. Our multicore design can achieve up to 8.09 × improvement in speedup and 57.77 × in energy efficiency compared to the multithreaded software implementation run on a Xeon Gold Processor. Moreover, our design highly outperforms the current State-of-the-Art hardware-accelerated solution, reaching up to 2876 Giga Cell Updates Per Second (GCUPS) and 68.47 GCUPS/W on a single FPGA, with an improvement of up to 2.29× and 9.90× in terms of performance and energy efficiency, respectively. Beatrice Branchini, Giulia Gerometta, Luisa Cicolini, Alberto Zeni, Emanuele Del Sozzo, Marco D. Santambrogio |
ISCAS | 1 |