Teodoro Urso

dblp:350/3904 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0005-4366-1102ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Electronic design automation · 57% Hardware accelerators and domain-specific architectures · 29% Cloud and datacenter computing · 14%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Electronic design automation
design space exploration
0.912025
NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
FPGA-based CNN accelerator
0.912025
NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025
Electronic design automation › design automation tools › FPGA CAD
FPGA design tools
0.912025
NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025
Electronic design automation
high-level synthesis
0.912025
NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025
Electronic design automation
logic synthesis
0.912025
NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025
Hardware accelerators and domain-specific architectures
machine learning accelerator
0.912025
NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025
Cloud and datacenter computing
resource allocation
0.912025
NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025

Methods — techniques the papers use, named apart from their topics

design space exploration · 0.9binary integer programming · 0.9
YearPublicationVenuePosition
2025 NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming
abstract
Skip connections have emerged as a key component of modern convolutional neural networks (CNNs) for computer vision tasks, allowing for the creation of more accurate and deeper models by addressing the vanishing gradient problem. However, the existing implementations of field-programmable gate array (FPGA)-based accelerators for ResNets and MobileNetV2 often experience decreased performance and increased computational latency due to the implementation of skip blocks. This article presents a novel framework for developing deep learning models on FPGAs that focuses on skip connections, with a unique approach to reduce buffering overhead. This results in a more efficient utilization of resources in the implementation of the skip layer. The nn2fpga compiler follows a thorough set of high-level synthesis (HLS) design principles and optimization strategies, exploiting in novel ways standard techniques to effectively map skip connection-based networks into static dataflow accelerators. To maximize throughput and efficiently use the available resources, our compiler employs a fast and effective design space exploration method based on a binary integer programming model which accurately assigns FPGA resources to the network layers, to maximize global throughput under resource constraints and then minimize resources for the achieved maximum throughput. Experimental results on the CIFAR-10 and ImageNet datasets demonstrate substantial gains in throughput ($\mathbf {3\times }$–$\mathbf {7\times }$on the past HLS-based work) for ResNet8, ResNet20, and MobileNetV2 models deployed on various Xilinx FPGA boards. Notably, MobileNetV2 deployed on the ZCU102 achieves a throughput of 2115 frame per second, representing even a 10% speedup over a state-of-the-art highly optimized manual register-transfer level implementation, showing that HLS can actually improve over manual design, thanks to the faster exploration of the design space.
Roberto Bosio, Filippo Minnella, Teodoro Urso, Mario R. Casu, Luciano Lavagno, Mihai T. Lazarescu, Paolo Pasini
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3