EDBT 2026 Demo / reviewers in the wild / expert
Roberto Bosio
dblp:378/0478
· DBLP profile ↗
3ranked-venue papers
2as first author
3since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Electronic design automation · 57% Hardware accelerators and domain-specific architectures · 29% Cloud and datacenter computing · 14% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
design space exploration |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
FPGA-based CNN accelerator |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Electronic design automation › design automation tools › FPGA CAD
FPGA design tools |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Electronic design automation
high-level synthesis |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Electronic design automation
logic synthesis |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Cloud and datacenter computing
resource allocation |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Methods — techniques the papers use, named apart from their topics
design space exploration · 0.9binary integer programming · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer ProgrammingabstractSkip connections have emerged as a key component of modern convolutional neural networks (CNNs) for computer vision tasks, allowing for the creation of more accurate and deeper models by addressing the vanishing gradient problem. However, the existing implementations of field-programmable gate array (FPGA)-based accelerators for ResNets and MobileNetV2 often experience decreased performance and increased computational latency due to the implementation of skip blocks. This article presents a novel framework for developing deep learning models on FPGAs that focuses on skip connections, with a unique approach to reduce buffering overhead. This results in a more efficient utilization of resources in the implementation of the skip layer. The nn2fpga compiler follows a thorough set of high-level synthesis (HLS) design principles and optimization strategies, exploiting in novel ways standard techniques to effectively map skip connection-based networks into static dataflow accelerators. To maximize throughput and efficiently use the available resources, our compiler employs a fast and effective design space exploration method based on a binary integer programming model which accurately assigns FPGA resources to the network layers, to maximize global throughput under resource constraints and then minimize resources for the achieved maximum throughput. Experimental results on the CIFAR-10 and ImageNet datasets demonstrate substantial gains in throughput ($\mathbf {3\times }$–$\mathbf {7\times }$on the past HLS-based work) for ResNet8, ResNet20, and MobileNetV2 models deployed on various Xilinx FPGA boards. Notably, MobileNetV2 deployed on the ZCU102 achieves a throughput of 2115 frame per second, representing even a 10% speedup over a state-of-the-art highly optimized manual register-transfer level implementation, showing that HLS can actually improve over manual design, thanks to the faster exploration of the design space. Roberto Bosio, Filippo Minnella, Teodoro Urso, Mario R. Casu, Luciano Lavagno, Mihai T. Lazarescu, Paolo Pasini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | SILVIA: Automated Superword-Level Parallelism Exploitation via HLS-specific LLVM Passes for Compute-Intensive FPGA AcceleratorsabstractHigh-level synthesis (HLS) aims at democratizing custom hardware acceleration with highly abstracted software-like descriptions. However, efficient accelerators still require substantial low-level hardware optimizations, defeating the HLS intent. In the context of field-programmable gate arrays, digital signal processors (DSPs) are a crucial resource that typically requires a significant optimization effort for its efficient utilization, especially when used for sub-word vectorization. This work proposes SILVIA, an open-source LLVM transformation pass that automatically identifies superword-level parallelism within an HLS design and exploits it by packing multiple operations, such as additions, multiplications, and multiply-and-adds, into a single DSP. SILVIA is integrated in the flow of the commercial AMD Vitis HLS tool and proves its effectiveness by packing multiple operations on the DSPs without any manual source-code modifications on several diverse state-of-the-art HLS designs such as convolutional neural networks and basic linear algebra subprograms accelerators, reducing the DSP utilization for additions by 70% and for multiplications and multiply-and-adds by 50% on average. Giovanni Brignone, Roberto Bosio, Fabrizio Ottati, Claudio Sansoè, Luciano Lavagno |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2024 | LESS: Low-Power Energy-Efficient Subgraph Isomorphism on FPGAabstractLow-power energy-efflcient subgraph isomorphism (LESS) is an open-source field-programmable gate array-only low-memory sub graph matching solver designed for energy efficiency. Depending on the input datagraph, the energy consumption of LESS, averaged on different diverse queries, is up to 38x and 93x lower than CPU and GPU solvers respectively. Roberto Bosio, Giovanni Brignone, Filippo Minnella, M. Usman Jamal, Luciano Lavagno |
DATE | 1 |