EDBT 2026 Demo / reviewers in the wild / expert
Filippo Minnella
dblp:345/1386
· DBLP profile ↗
3ranked-venue papers
0as first author
3since 2021 · last 2025
0000-0001-6713-8942ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Electronic design automation · 60% Hardware accelerators and domain-specific architectures · 21% Cloud and datacenter computing · 10% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Electronic design automation
design space exploration |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › CNN accelerator
FPGA-based CNN accelerator |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Electronic design automation › design automation tools › FPGA CAD
FPGA design tools |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Electronic design automation
high-level synthesis |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Electronic design automation
logic synthesis |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Hardware accelerators and domain-specific architectures
machine learning accelerator |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Cloud and datacenter computing
resource allocation |
0.9 | 1 | 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer Programming · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2025 |
Electronic design automation › logic synthesis › sequential circuit optimization
retiming |
0.8 | 1 | 2024 | Mix & Latch: Comparison With State-of-the-Art Retiming on a RISC-V Benchmark · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024 |
Integrated circuit design › digital circuit design
sequential circuit design |
0.8 | 1 | 2024 | Mix & Latch: Comparison With State-of-the-Art Retiming on a RISC-V Benchmark · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024 |
Electronic design automation › physical design
timing optimization |
0.8 | 1 | 2024 | Mix & Latch: Comparison With State-of-the-Art Retiming on a RISC-V Benchmark · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2024 |
Methods — techniques the papers use, named apart from their topics
design space exploration · 0.9binary integer programming · 0.9post-synthesis timing analysis · 0.8mix & latch · 0.8critical path analysis · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer ProgrammingabstractSkip connections have emerged as a key component of modern convolutional neural networks (CNNs) for computer vision tasks, allowing for the creation of more accurate and deeper models by addressing the vanishing gradient problem. However, the existing implementations of field-programmable gate array (FPGA)-based accelerators for ResNets and MobileNetV2 often experience decreased performance and increased computational latency due to the implementation of skip blocks. This article presents a novel framework for developing deep learning models on FPGAs that focuses on skip connections, with a unique approach to reduce buffering overhead. This results in a more efficient utilization of resources in the implementation of the skip layer. The nn2fpga compiler follows a thorough set of high-level synthesis (HLS) design principles and optimization strategies, exploiting in novel ways standard techniques to effectively map skip connection-based networks into static dataflow accelerators. To maximize throughput and efficiently use the available resources, our compiler employs a fast and effective design space exploration method based on a binary integer programming model which accurately assigns FPGA resources to the network layers, to maximize global throughput under resource constraints and then minimize resources for the achieved maximum throughput. Experimental results on the CIFAR-10 and ImageNet datasets demonstrate substantial gains in throughput ($\mathbf {3\times }$–$\mathbf {7\times }$on the past HLS-based work) for ResNet8, ResNet20, and MobileNetV2 models deployed on various Xilinx FPGA boards. Notably, MobileNetV2 deployed on the ZCU102 achieves a throughput of 2115 frame per second, representing even a 10% speedup over a state-of-the-art highly optimized manual register-transfer level implementation, showing that HLS can actually improve over manual design, thanks to the faster exploration of the design space. Roberto Bosio, Filippo Minnella, Teodoro Urso, Mario R. Casu, Luciano Lavagno, Mihai T. Lazarescu, Paolo Pasini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | LESS: Low-Power Energy-Efficient Subgraph Isomorphism on FPGAabstractLow-power energy-efflcient subgraph isomorphism (LESS) is an open-source field-programmable gate array-only low-memory sub graph matching solver designed for energy efficiency. Depending on the input datagraph, the energy consumption of LESS, averaged on different diverse queries, is up to 38x and 93x lower than CPU and GPU solvers respectively. Roberto Bosio, Giovanni Brignone, Filippo Minnella, M. Usman Jamal, Luciano Lavagno |
DATE | 3 |
| 2024 | Mix & Latch: Comparison With State-of-the-Art Retiming on a RISC-V BenchmarkabstractFlip-flops (FFs) are the most commonly used sequential elements in synchronous circuits, but their timing requirements limit the operating frequency. Borrowing time with a latch-based approach can increase operating frequency, but traditional back-end optimization tools struggle to manage hold time requirements. The Mix & Latch technique achieves higher frequencies and often lower area than commercial state-of-the-art retiming by exploiting four types of synchronous sequential gates, namely, positive and negative edge-triggered flip-flops (FFs) and positive and negative transparent latches, all using a single clock tree.In this article, we first significantly accelerate the Mix & Latch flow convergence with respect to past work, by using a post-synthesis-based timing analysis that eliminates the first placement and routing needed for post-layout timing analysis. Then, by adding tolerance margins to the timing model, the pessimism is reduced to improve both convergence speed and maximum frequency. Finally, we reduce the complexity of the problem by applying the methodology only to the sequential elements belonging to critical paths. The effectiveness of Mix & Latch is then demonstrated on a RISC-V processor core from the Pulp platform using 28nm CMOS FDSOI technology. The results are compared to both the original Mix & Latch flow and a retiming performed with a state-of-the-art tool, showing a 25% frequency improvement over the original flow and 7.5% over the retiming flow. Compared to the retiming flow, we achieve comparable or lower power and area, while preserving the original registers and allowing logic equivalence checking. Lorenzo Lagostina, Filippo Minnella, Jordi Cortadella, Mario R. Casu, Mihai T. Lazarescu, Luciano Lavagno |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |