EDBT 2026 Demo / reviewers in the wild / expert
Marco Venere
dblp:354/4274
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0009-0002-8991-1443ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RoPeerTo: A Datacenter-Scale Architecture for Peer-To-Peer DMA between GPUs and FPGAsabstractModern datacenters integrate heterogeneous accelerators, such as GPUs and FPGAs, to speed up different stages of compute-intensive pipelines. GPUs are best suited for massively parallel workloads (e.g., deep learning), while FPGAs excel at task-level parallelism, stream-oriented processing, and in-network acceleration. Since these architectures must exchange data efficiently, literature introduced Peer-To-Peer (P2P) communication across PCI Express (PCIe) devices, to reduce CPU-driven orchestration and avoid intermediate, redundant buffer copies that degrade performance. However, current solutions are either closed-source or tied to proprietary frameworks, limiting P2P communication across most PCIe-based devices and requiring significant technical effort to enable P2P capabilities on supported hardware. For this reason, we propose RoPeerTo, a fully open-source, datacenter-scale architecture for P2P DMA communication over PCIe, validated on both GPUs and FPGAs. The goal is to provide a general, open alternative that ensures flexibility, efficiency, and usability. To this end, we design a complete HW/SW stack operating across different layers, supporting standard protocols for DMA-based memory sharing, advanced tools for device virtualization, memory address translation, and access protection. The result is a unified framework exposing a high-level API to end users, that enables direct communication between accelerators such as FPGAs and GPUs, and abstracts away the underlying hardware setup and management. We validate the system across different scenarios. First, we isolate the communication layer, observing a 5.61× speedup and a 37.99% reduction in GPU power consumption during data transfer. Next, we leverage the system for a compute-intensive workload where communication is only a partial bottleneck, achieving a 6.77% speedup without any compute-side modifications. Finally, we evaluate communication-heavy distributed computing workloads, demonstrating up to a 21.79× speedup in network-bound data scattering. Marco Venere, Giuseppe Sorrentino, Benjamin Ramhorst, Maximilian Jakob Heer, Lucian Petrica, Dario Korolija, Marco D. Santambrogio, Davide Conficconi, Gustavo Alonso, Kenneth O'Brien |
EuroSys | 1 |
| 2025 | DDRoute: a Novel Depth-Driven Approach to the Qubit Routing ProblemabstractIn the Noisy Intermediate-Scale Quantum (NISQ) era, the topological constraints present in many of the currently available quantum devices pose a physical limit on the feasible interactions between qubits. To comply with such limitations, the compilation of quantum circuits requires solving the Qubit Routing Problem (QRP), by inserting SWAP operations among qubits. The State of the Art provides heuristic algorithms addressing this task, yet the depth of the output circuits is often incompatible with the current limits of quantum hardware. Therefore, we propose DDRoute, a novel heuristic algorithm to solve QRP, designed to reduce the depth overhead introduced by the routing process in the compiled circuits. Our experimental evaluation proves the efficiency of our approach, with a depth reduction of up to 70% with respect to the state-of-the-art routing procedures. Alessandro Annechini, Marco Venere, Donatella Sciuto, Marco D. Santambrogio |
DAC | 2 |
| 2025 | Rock the QASBA: Quantum Error Correction Acceleration via the Sparse Blossom Algorithm on FPGAsabstractQuantum computing is a new paradigm of computation that exploits principles from quantum mechanics to achieve an exponential speedup compared to classical logic. However, noise strongly limits current quantum hardware, reducing achievable performance and limiting the scaling of the applications. For this reason, current noisy intermediate-scale quantum devices require Quantum Error Correction (QEC) mechanisms to identify errors occurring in the computation and correct them in real time. Nevertheless, the high computational complexity of QEC algorithms is incompatible with the tight time constraints of quantum devices. Thus, hardware acceleration is paramount to achieving real-time QEC. This work presents QASBA, an FPGA-based hardware accelerator for the Sparse Blossom Algorithm (SBA), a state-of-the-art decoding algorithm. After profiling the state-of-the-art software counterpart, we developed a design methodology for hardware development based on the SBA. We also devised an automation process to help users without expertise in hardware design in deploying architectures based on QASBA. We implement QASBA on different FPGA architectures and experimentally evaluate resource usage, execution time, and energy efficiency of our solution. Our solution attains up to \(25.05\times\) speedup and \(304.16\times\) improvement in energy efficiency compared to the software baseline. Marco Venere, Beatrice Branchini, Davide Conficconi, Donatella Sciuto, Marco D. Santambrogio |
ACM Trans. Reconfigurable Technol. Syst. | 1 |
| 2024 | A Quantum Method to Match Vector Boolean Functions using Simon's SolverabstractThe Boolean Matching Problem is a fundamental step in modern Electronic Design Automation toolchains, which allow the efficient design of large classical computers. In particular, the equivalence under negation-permutation-negation of two n-to-n vector Boolean functions requires the exploration of a super-exponential number of possible negations and permutations of input and output variables, and is widely regarded as a daunting challenge. Its classical complexity$(\mathcal{O}(n!2^{2n})$, where$n$is the number of input and output variables) is rarely tolerated by EDA tools, which are typically solving small instances of the Boolean Matching Problem for n-to-1 Boolean functions. In this work, we present a method to exploit the solver for Simon's problem to speedup the matching of n-to-n vector Boolean functions, as we show that, despite its higher complexity, it is friendlier to a quantum solver than matching single-output Boolean functions. Our solution allows saving a factor$2^{n}$in the overall worst-case computational effort, and is amenable to combined approaches such as the so-called Grover-meets-Simon, which have the potential of reducing it below the cost of classical n-to-1 matching. We provide a fully detailed quantum circuit implementing our proposal, and compute its cost, both counting the required amount of qubits and quantum gates. We conducted an experimental evaluation employing the ISCAS benchmark suite, a de-facto standard for classical EDA to derive our sample Boolean functions. Marco Venere, Alessandro Barenghi, Gerardo Pelosi |
ICCD | 1 |
| 2023 | Hephaestus: Codesigning and Automating 3D Image Registration on Reconfigurable ArchitecturesabstractHealthcare is a pivotal research field, and medical imaging is crucial in many applications. Therefore finding new architectural and algorithmic solutions would benefit highly repetitive image processing procedures. One of the most complex tasks in this sense is image registration, which finds the optimal geometric alignment among 3D image stacks and is widely employed in healthcare and robotics. Given the high computational demand of such a procedure, hardware accelerators are promising real-time and energy-efficient solutions, but they are complex to design and integrate within software pipelines. Therefore, this work presents an automation framework called Hephaestus that generates efficient 3D image registration pipelines combined with reconfigurable accelerators. Moreover, to alleviate the burden from the software, we codesign software-programmable accelerators that can adapt at run-time to the image volume dimensions. Hephaestus features a cross-platform abstraction layer that enables transparently high-performance and embedded systems deployment. However, given the computational complexity of 3D image registration, the embedded devices become a relevant and complex setting being constrained in memory; thus, they require further attention and tailoring of the accelerators and registration application to reach satisfactory results. Therefore, with Hephaestus , we also propose an approximation mechanism that enables such devices to perform the 3D image registration and even achieve, in some cases, the accuracy of the high-performance ones. Overall, Hephaestus demonstrates 1.85× of maximum speedup, 2.35× of efficiency improvement with respect to the State of the Art, a maximum speedup of 2.51× and 2.76× efficiency improvements against our software, while attaining state-of-the-art accuracy on 3D registrations. Giuseppe Sorrentino, Marco Venere, Davide Conficconi, Eleonora D'Arnese, Marco D. Santambrogio |
ACM Trans. Embed. Comput. Syst. | 2 |