EDBT 2026 Demo / reviewers in the wild / expert
Teodoro Urso
dblp:350/3904
· DBLP profile ↗
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0005-4366-1102ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Electronic design automation · 57% Hardware accelerators and domain-specific architectures · 29% Cloud and datacenter computing · 14% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
design space exploration |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
FPGA-based CNN accelerator |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Electronic design automation › design automation tools › FPGA CAD
FPGA design tools |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Electronic design automation
high-level synthesis |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Electronic design automation
logic synthesis |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Cloud and datacenter computing
resource allocation |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Methods — techniques the papers use, named apart from their topics
design space exploration · 0.9binary integer programming · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer ProgrammingabstractSkip connections have emerged as a key component of modern convolutional neural networks (CNNs) for computer vision tasks, allowing for the creation of more accurate and deeper models by addressing the vanishing gradient problem. However, the existing implementations of field-programmable gate array (FPGA)-based accelerators for ResNets and MobileNetV2 often experience decreased performance and increased computational latency due to the implementation of skip blocks. This article presents a novel framework for developing deep learning models on FPGAs that focuses on skip connections, with a unique approach to reduce buffering overhead. This results in a more efficient utilization of resources in the implementation of the skip layer. The nn2fpga compiler follows a thorough set of high-level synthesis (HLS) design principles and optimization strategies, exploiting in novel ways standard techniques to effectively map skip connection-based networks into static dataflow accelerators. To maximize throughput and efficiently use the available resources, our compiler employs a fast and effective design space exploration method based on a binary integer programming model which accurately assigns FPGA resources to the network layers, to maximize global throughput under resource constraints and then minimize resources for the achieved maximum throughput. Experimental results on the CIFAR-10 and ImageNet datasets demonstrate substantial gains in throughput ($\mathbf {3\times }$–$\mathbf {7\times }$on the past HLS-based work) for ResNet8, ResNet20, and MobileNetV2 models deployed on various Xilinx FPGA boards. Notably, MobileNetV2 deployed on the ZCU102 achieves a throughput of 2115 frame per second, representing even a 10% speedup over a state-of-the-art highly optimized manual register-transfer level implementation, showing that HLS can actually improve over manual design, thanks to the faster exploration of the design space. Roberto Bosio, Filippo Minnella, Teodoro Urso, Mario R. Casu, Luciano Lavagno, Mihai T. Lazarescu, Paolo Pasini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |