EDBT 2026 Demo / reviewers in the wild / expert
Mihai T. Lazarescu
dblp:36/5841 · also Mihai Teodor Lazarescu
· DBLP profile ↗
17ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0003-0884-5158ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 14 · 2 first-author · 5 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NN2FPGA: Optimizing CNN Inference on FPGAs With Binary Integer ProgrammingabstractSkip connections have emerged as a key component of modern convolutional neural networks (CNNs) for computer vision tasks, allowing for the creation of more accurate and deeper models by addressing the vanishing gradient problem. However, the existing implementations of field-programmable gate array (FPGA)-based accelerators for ResNets and MobileNetV2 often experience decreased performance and increased computational latency due to the implementation of skip blocks. This article presents a novel framework for developing deep learning models on FPGAs that focuses on skip connections, with a unique approach to reduce buffering overhead. This results in a more efficient utilization of resources in the implementation of the skip layer. The nn2fpga compiler follows a thorough set of high-level synthesis (HLS) design principles and optimization strategies, exploiting in novel ways standard techniques to effectively map skip connection-based networks into static dataflow accelerators. To maximize throughput and efficiently use the available resources, our compiler employs a fast and effective design space exploration method based on a binary integer programming model which accurately assigns FPGA resources to the network layers, to maximize global throughput under resource constraints and then minimize resources for the achieved maximum throughput. Experimental results on the CIFAR-10 and ImageNet datasets demonstrate substantial gains in throughput ($\mathbf {3\times }$–$\mathbf {7\times }$on the past HLS-based work) for ResNet8, ResNet20, and MobileNetV2 models deployed on various Xilinx FPGA boards. Notably, MobileNetV2 deployed on the ZCU102 achieves a throughput of 2115 frame per second, representing even a 10% speedup over a state-of-the-art highly optimized manual register-transfer level implementation, showing that HLS can actually improve over manual design, thanks to the faster exploration of the design space. Roberto Bosio, Filippo Minnella, Teodoro Urso, Mario R. Casu, Luciano Lavagno, Mihai T. Lazarescu, Paolo Pasini |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | Mix & Latch: Comparison With State-of-the-Art Retiming on a RISC-V BenchmarkabstractFlip-flops (FFs) are the most commonly used sequential elements in synchronous circuits, but their timing requirements limit the operating frequency. Borrowing time with a latch-based approach can increase operating frequency, but traditional back-end optimization tools struggle to manage hold time requirements. The Mix & Latch technique achieves higher frequencies and often lower area than commercial state-of-the-art retiming by exploiting four types of synchronous sequential gates, namely, positive and negative edge-triggered flip-flops (FFs) and positive and negative transparent latches, all using a single clock tree.In this article, we first significantly accelerate the Mix & Latch flow convergence with respect to past work, by using a post-synthesis-based timing analysis that eliminates the first placement and routing needed for post-layout timing analysis. Then, by adding tolerance margins to the timing model, the pessimism is reduced to improve both convergence speed and maximum frequency. Finally, we reduce the complexity of the problem by applying the methodology only to the sequential elements belonging to critical paths. The effectiveness of Mix & Latch is then demonstrated on a RISC-V processor core from the Pulp platform using 28nm CMOS FDSOI technology. The results are compared to both the original Mix & Latch flow and a retiming performed with a state-of-the-art tool, showing a 25% frequency improvement over the original flow and 7.5% over the retiming flow. Compared to the retiming flow, we achieve comparable or lower power and area, while preserving the original registers and allowing logic equivalence checking. Lorenzo Lagostina, Filippo Minnella, Jordi Cortadella, Mario R. Casu, Mihai T. Lazarescu, Luciano Lavagno |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | A DSP shared is a DSP earned: HLS Task-Level Multi-Pumping for High-Performance Low-Resource DesignsabstractHigh-level synthesis (HLS) enhances digital hardware design productivity through a high abstraction level. Even if the HLS abstraction prevents fine-grained manual register-transfer level (RTL) optimizations, it also enables automatable optimizations that would be unfeasible or hard to automate at RTL. Specifically, we propose a task-level multi-pumping methodology to reduce resource utilization, particularly digital signal processors (DSPs), while preserving the throughput of HLS kernels modeled as dataflow graphs (DFGs) targeting field-programmable gate arrays. The methodology exploits the HLS resource sharing to automatically insert the logic for reusing the same functional unit for different operations. In addition, it relies on multi-clock DFGs to run the multi-pumped tasks at higher frequencies. The methodology scales the pipeline initiation interval (II) and the clock frequency constraints of resource-intensive tasks by a multi-pumping factor (M). The looser II allows sharing the same resource among M different operations, while the tighter clock frequency preserves the throughput. We verified that our methodology opens a new Pareto front in the throughput and resource space by applying it to open-source HLS designs using state-of-the-art commercial HLS and implementation tools by Xilinx. The multi-pumped designs require up to 40% fewer DSP resources at the same throughput as the original designs optimized for performance (i.e., running at the maximum clock frequency) and achieve up to 50% better throughput using the same DSPs as the original designs optimized for resources with a single clock. Giovanni Brignone, Mihai T. Lazarescu, Luciano Lavagno |
ICCD | 2 |
| 2022 | Fast Energy-Optimal Multikernel DNN-Like Application Allocation on Multi-FPGA PlatformsabstractPlatforms with multiple field-programmable gate arrays (FPGAs), such as Amazon Web Services (AWS) F1 instances, can efficiently accelerate multikernel pipelined applications, e.g., convolutional neural networks for machine vision tasks or transformer networks for natural language processing tasks. To reduce energy consumption when the FPGAs are underutilized, we propose a model to 1) find offline the minimum-power solution for given throughput constraints and 2) dynamically reprogram the FPGA at runtime (which is complementary to dynamic voltage and frequency scaling) to match best the workloads when they change. The offline optimization model can be solved using a mixed-integer nonlinear programming (MINLP) solver, but it can be very slow. Hence, we provide two heuristic optimization methods that improve result quality within a bounded time. We use several very large designs to demonstrate that both heuristics obtain comparable results to MINLP, when it can find the best solution, and they obtain much better results than MINLP, when it cannot find the optimum within a bounded amount of time. The heuristic methods can also be thousands of times faster than the MINLP solver. Junnan Shan, Mihai T. Lazarescu, Jordi Cortadella, Luciano Lavagno, Mario R. Casu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2021 | Asynchronous Resilient Wireless Sensor Network for Train Integrity MonitoringabstractTo increase railway use efficiency, the European Railway Traffic Management System (ERTMS) Level 3 requires all trains to constantly and reliably self-monitor and report their integrity and track position without infrastructure support. Timely train separation detection is challenging, especially for long freight trains without electrical power on cars. Data fusion of multiple monitoring techniques is currently investigated, including distributed integrity sensing of all train couplings. We propose a wireless sensor network (WSN) topology, communication protocol, application, and sensor nodes prototypes designed for low-power timely train integrity (TI) reporting in unreliable conditions, like intermittent node operation and network association (e.g., in low environmental energy harvesting conditions) and unreliable radio links. Each train coupling is redundantly monitored by four sensors, which can help to satisfy the train collision avoidance system (TCAS) and European Committee for Electrotechnical Standardization (CENELEC) software integrity level (SIL) 4 requirements and contribute to the reliability of the asynchronous network with low rejoin overhead. A control center on the locomotive controls the WSN and receives the reports, helping the integration in railway or Internet-of-Things (IoT) applications. Software simulations of the embedded application code virtually unchanged show that the energy-optimized configurations check a 50-car TI (about 1-km long) in 3.6-s average with 0.1-s standard deviation and that more than 95% of the reports are delivered successfully with up to one-third of communications or up to 15% of the nodes failed. We also report qualitative test results for a 20-node network in different experimental conditions. Mihai T. Lazarescu, Pooya Poolad |
IEEE Internet Things J. | 1 |
| 2021 | CNN-on-AWS: Efficient Allocation of Multikernel Applications on Multi-FPGA PlatformsabstractMulti-FPGA platforms, like Amazon AWS F1, can run in the cloud multikernel pipelined applications, like convolutional neural networks (CNNs), with excellent performance and lower energy consumption than CPUs or GPUs. We propose a method to efficiently map these applications on multi-FPGA platforms to maximize the application throughput. Our methodology finds, for the given resources, the optimal number of parallel instances of each kernel in the pipeline and their allocation to one or more among the available FPGAs. We obtain this by formulating and solving a mixed-integer, nonlinear optimization problem, in which we model the performance of each component and the duration of the phases in which the accelerated computation can be split into, namely: 1) data transfer from a host CPU to the DDR memory of each FPGA; 2) data transfer from FPGA DDR to FPGA on-chip memory; 3) kernel computation on the FPGA; 4) data transfer from FPGA on-chip memory to FPGA DDR; and 5) data transfer from FPGA DDR to host. Finding the optimal solution using a mixed-integer nonlinear programming (MINLP) solver is often highly inefficient. Hence, we provide a fast heuristic method that according to our experiments can be much more efficient than the MINLP solver and finds comparable results. For larger problems (more CNN layers), our heuristic method can quickly find (several thousand times faster) much better solutions than the MINLP solver, even if we run the latter for a very long time. Junnan Shan, Mihai T. Lazarescu, Jordi Cortadella, Luciano Lavagno, Mario R. Casu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Exact and Heuristic Allocation of Multi-kernel Applications to Multi-FPGA PlatformsabstractFPGA-based accelerators demonstrated high energy efficiency compared to GPUs and CPUs. However, single FPGA designs may not achieve sufficient task parallelism. In this work, we optimize the mapping of high-performance multi-kernel applications, like Convolutional Neural Networks, to multi-FPGA platforms. First, we formulate the system level optimization problem, choosing within a huge design space the parallelism and number of compute units for each kernel in the pipeline. Then we solve it using a combination of Geometric Programming, producing the optimum performance solution given resource and DRAM bandwidth constraints, and a heuristic allocator of the compute units on the FPGA cluster. Junnan Shan, Mario R. Casu, Jordi Cortadella, Luciano Lavagno, Mihai T. Lazarescu |
DAC | 5 |
| 2015 | Interactive Trace-Based Analysis Toolset for Manual Parallelization of C ProgramsabstractMassive amounts of legacy sequential code need to be parallelized to make better use of modern multiprocessor architectures. Nevertheless, writing parallel programs is still a difficult task. Automated parallelization methods can be effective both at the statement and loop levels and, recently, at the task level, but they are still restricted to specific source code constructs or application domains. We present in this article an innovative toolset that supports developers when performing manual code analysis and parallelization decisions. It automatically collects and represents the program profile and data dependencies in an interactive graphical format that facilitates the analysis and discovery of manual parallelization opportunities. The toolset can be used for arbitrary sequential C programs and parallelization patterns. Also, its program-scope data dependency tracing at runtime can complement the tools based on static code analysis and can also benefit from it at the same time. We also tested the effectiveness of the toolset in terms of time to reach parallelization decisions and of their quality. We measured a significant improvement for several real-world representative applications. Mihai T. Lazarescu, Luciano Lavagno |
ACM Trans. Embed. Comput. Syst. | 1 |
| 2015 | Virtual Platform-Based Design Space Exploration of Power-Efficient Distributed Embedded ApplicationsabstractNetworked embedded systems are essential building blocks of a broad variety of distributed applications ranging from agriculture to industrial automation to healthcare and more. These often require specific energy optimizations to increase the battery lifetime or to operate using energy harvested from the environment. Since a dominant portion of power consumption is determined and managed by software, the software development process must have access to the sophisticated power management mechanisms provided by state-of-the-art hardware platforms to achieve the best tradeoff between system availability and reactivity. Furthermore, internode communications must be considered to properly assess the energy consumption. This article describes a design flow based on a SystemC virtual platform including both accurate power models of the hardware components and a fast abstract model of the wireless network. The platform allows both model-driven design of the application and the exploration of power and network management alternatives. These can be evaluated in different network scenarios, allowing one to exploit power optimization strategies without requiring expensive field trials. The effectiveness of the approach is demonstrated via experiments on a wireless body area network application. Parinaz Sayyah, Mihai T. Lazarescu, Sara Bocchio, Emad Samuel Malki Ebeid, Gianluca Palermo, Davide Quaglia, Alberto Rosti, Luciano Lavagno |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2014 | Energy-aware parallelization flow and toolset for C codeabstractMulticore architectures are increasingly used in embedded systems to achieve higher throughput with lower energy consumption. This trend accentuates the need to convert existing sequential code to effectively exploit the resources of these architectures. We present a parallelization flow and toolset for legacy C code that includes a performance estimation tool, a parallelization tool, and a streaming-oriented parallelization framework. These are part of the work-in-progress EU FP7 PHARAON project that aims to develop a complete set of techniques and tools to guide and assist software development for heterogeneous parallel architectures. We demonstrate the effectiveness of the use of the toolset in an experiment where we measure the parallelization quality and time for inexperienced users, and the parallelization flow and performance results for the parallelization of a practical example of a stereo vision application. Mihai T. Lazarescu, Albert Cohen 0001, Adrien Guatto, Nhat Minh Lê, Luciano Lavagno, Antoniu Pop, Andrei Sergeevich Terechko, Alexandru Sutii |
SCOPES | 1 |
| 2014 | Introduction to Special Issue on Application of Concurrency to System Design (ACSD'13)abstractNo abstract available. Josep Carmona 0001, Mihai T. Lazarescu, Marta Pietkiewicz-Koutny |
ACM Trans. Embed. Comput. Syst. | 2 |
| 2013 | EU FP7-288307 Pharaon Project: Parallel and Heterogeneous Architecture for Real-Time ApplicationsabstractIn this article, we present the work-in-progress of the EU FP7 PHARAON project, started in September 2011. The first objective of the project is the development of new techniques and tools capable to assist the designer in the development of parallel embedded systems, from executable specifications to target-specific implementation and debugging on a multicore platform. This tool chain will offer and implement several parallelization strategies, reflecting the functional and non-functional constraints of the system, and driving the designer into incremental parallelization and adaptation steps. The second objective of the project is to develop monitoring and control techniques in the middleware of the system capable to automatically adapt platform services to application requirements and therefore reduce power consumption transparently. Héctor Posadas, Eugenio Villar, Florian Broekaert, Michel Bourdellès, Albert Cohen 0001, Antoniu Pop, Nhat Minh Lê, Adrien Guatto, Mihai T. Lazarescu, Luciano Lavagno, Andrei Sergeevich Terechko, Miguel Glassée, Daniel Calvo, Eduardo de las Heras |
DSD | 9 |
| 2012 | HEAP: A Highly Efficient Adaptive Multi-processor FrameworkabstractWriting parallel code is difficult, especially when starting from a sequential reference implementation. Our research efforts, as demonstrated in this paper, face this challenge directly by providing an innovative toolset that helps software developers profile and parallelize an existing sequential implementation, by exploiting top-level pipeline-style parallelism. The innovation of our approach is based on the facts that a) we use both automatic and profiling-driven estimates of the available parallelism, b) we refine those estimates using metric-driven verification techniques, and c) we support dynamic recovery of excessively optimistic parallelization. The proposed toolset has been utilized to find an efficient parallel code organization for a number of real-world representative applications, and a version of the toolset is provided in an open-source manner. Luciano Lavagno, Mihai T. Lazarescu, Ioannis Papaefstathiou, Andreas Brokalakis, Johan Walters, Bart Kienhuis, Florian Schäfer 0001 |
DSD | 2 |
| 2012 | SystemC Model Generation for Realistic Simulation of Networked Embedded SystemsabstractVerification and design-space exploration of today's embedded systems require the simulation of heterogeneous aspects of the system, i.e., software, hardware, communications. This work shows the use of SystemC to simulate a model-driven specification of the behavior of a networked embedded system together with a complete network scenario consisting of the radio channel, the IEEE 802.15.4 protocol for wireless personal area networks and concurrent traffic sharing the medium. The paper describes the main issues addressed to generate SystemC modules from Matlab/Stateflow descriptions and to integrate them in a complete network scenario. Simulation results on a healthcare wireless sensor network show the validity of the approach. Mihai T. Lazarescu, Parinaz Sayyah, Davide Quaglia, Francesco Stefanni |
DSD | 1 |
| 2012 | FASTCUDA: Open Source FPGA Accelerator & Hardware-Software Codesign Toolset for CUDA KernelsabstractUsing FPGAs as hardware accelerators that communicate with a central CPU is becoming a common practice in the embedded design world but there is no standard methodology and toolset to facilitate this path yet. On the other hand, languages such as CUDA and OpenCL provide standard development environments for Graphical Processing Unit (GPU) programming. FASTCUDA is a platform that provides the necessary software toolset, hardware architecture, and design methodology to efficiently adapt the CUDA approach into a new FPGA design flow. With FASTCUDA, the CUDA kernels of a CUDA-based application are partitioned into two groups with minimal user intervention: those that are compiled and executed in parallel software, and those that are synthesized and implemented in hardware. A modern low power FPGA can provide the processing power (via numerous embedded micro-CPUs) and the logic capacity for both the software and hardware implementations of the CUDA kernels. This paper describes the system requirements and the architectural decisions behind the FASTCUDA approach. Iakovos Mavroidis, Ioannis Mavroidis, Ioannis Papaefstathiou, Luciano Lavagno, Mihai T. Lazarescu, Eduardo de la Torre, Florian Schäfer 0001 |
DSD | 5 |
| 2012 | Dynamic Trace-Based Data Dependency Analysis for Parallelization of C ProgramsabstractWriting parallel code is traditionally considered a difficult task, even when it is tackled from the beginning of a project. In this paper, we demonstrate an innovative toolset that faces this challenge directly. It provides the software developers with profile data and directs them to possible top-level, pipeline-style parallelization opportunities for an arbitrary sequential C program. This approach is complementary to the methods based on static code analysis and automatic code rewriting and does not impose restrictions on the structure of the sequential code or the parallelization style, even though it is mostly aimed at coarse-grained task-level parallelization. The proposed toolset has been utilized to define parallel code organizations for a number of real-world representative applications and is based on and is provided as free source. Mihai T. Lazarescu, Luciano Lavagno |
SCAM | 1 |
| 2010 | Energy optimization at the MAC layer for a forest fire monitoring wireless sensor networkabstractThis paper describes several optimizations of MAC protocols that can be applied in order to satisfy the constraints that come from a real-life application. Forest fire monitoring requires very different latencies and data sizes, depending on whether it is reporting normal conditions, sending an alarm, or performing network management functions. We use MAC algorithms that extensively shut down the radio in order to save power. We show that by exploiting knowledge about the Link Quality Index and by effectively using the free time of the channel only when there is more data than usual to transmit, we manage to decrease latency and contention and increase bandwidth usage, while keeping power consumption very low. Anwar Al-Khateeb, Jun Kyoung Kim, Luciano Lavagno, Mihai T. Lazarescu |
ETFA | 4 |