VLDB 2026 Research / reviewers in the wild / expert
Davide Giri
dblp:158/0860
· DBLP profile ↗
14ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-4101-4516ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 13 · 4 first-author · 8 since 2021Software engineering, systems software and programming languages · 4 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FLIP2M: Flexible Intra-layer Parallelism and Inter-layer Pipelining for Multi-model AR/VR WorkloadsabstractTiled accelerator architectures provide opportunities to optimize the performance of multi-model augmented and virtual reality (AR/VR) applications through intra-layer parallelism and inter-layer pipelining. However, balancing these two strategies is a difficult task that demands a flexible architecture to deploy models and an optimization approach, that is, capable of selecting an optimal strategy from an enormous mapping space. This article presents FLIP2M, a holistic solution for mapping multi-model AR/VR workloads on tiled architectures. FLIP2M consists of (1) FLIP, an acceleration fabric that supports a wide variety of optimizations through flexible on-chip communication, and (2) OASIS, an optimization framework based on dynamic and constraint programming, that is, capable of selecting an efficient strategy for mapping multi-model workloads onto FLIP. We demonstrate FLIP2M on an FPGA prototype of FLIP that features 36 accelerators and 7 DDR4 controllers. Using OASIS-generated mappings for three different multi-model AR/VR workloads, FLIP2M achieves up to 1.94× improvement in latency, 1.37× in energy, and 2.59× in energy-delay product relative to a FLIP baseline without intra-layer resource allocation flexibility and inter-layer pipelining. Gabriele Tombesi, Je Yang, Joseph Zuckerman, Davide Giri, William Baisi, Luca P. Carloni |
ACM Trans. Embed. Comput. Syst. | 4 |
| 2024 | Mozart: Taming Taxes and Composing Accelerators with Shared-MemoryabstractResource-constrained system-on-chips (SoCs) are increasingly heterogeneous with specialized accelerators for various tasks. Acceleration taxes due to control and data movement, however, diminish end-to-end speedups from hardware acceleration. Meanwhile, emerging workloads are increasingly task-diverse with several, potentially shared, fine-grained acceleration candidates. This motivates a paradigm of parallel and disaggregated acceleration. Compared to a monolithic accelerator, disaggregation provides higher flexibility, reuse, and utilization, but at the cost of higher control and data acceleration taxes. Vignesh Suresh, Bakshree Mishra, Zeran Zhu, Naiyin Jin, Charles Block, Paolo Mantovani, Davide Giri, Joseph Zuckerman, Luca P. Carloni, Sarita V. Adve |
PACT | 8 |
| 2024 | BlitzCoin: Fully Decentralized Hardware Power Management for Accelerator-Rich SoCsabstractOn-chip power-management techniques have evolved over several processor generations. However, response time and scalability constraints have made it difficult to translate existing power-management strategies to current or next-generation System-on-Chip (SoC) architectures, which are expected to comprise tens to hundreds of cores and accelerators. In this work we present BlitzCoin, a fully decentralized hardware power-management strategy for large, accelerator-rich SoCs, coupled with optimized unified voltage and frequency regulation. We evaluated BlitzCoin through RTL simulations of multiple SoCs targeted toward different application domains. The results are further validated through silicon measurements of a fabricated 12 nm many-accelerator SoC that includes BlitzCoin. Our evaluations show that BlitzCoin is markedly faster, with 8× to 12× lower response times, which provides 25%-34% throughput improvement and allows for scaling to 7 × to 13 × larger SoCs compared to state-of-the-art centralized power-management strategies, all with an area overhead of <1%. Martin Cochet, Karthik Swaminathan, Erik Jens Loscalzo, Joseph Zuckerman, Maico Cassel, Davide Giri, Alper Buyuktosunoglu, David Brooks 0001, Gu-Yeon Wei, Kenneth L. Shepard, Luca P. Carloni, Pradip Bose |
ISCA | 6 |
| 2023 | PR-ESP: An Open-Source Platform for Design and Programming of Partially Reconfigurable SoCsabstractDespite its presence for more than two decades and its proven benefits in expanding the space of system design, dynamic partial reconfiguration (DPR) is rarely integrated into frameworks and platforms that are used to design complex reconfigurable system-on-chip (SoC) architectures. This is due to the complexity of the DPR FPGA flow as well as the lack of architectural and software runtime support to enable and fully harness DPR. Moreover, as DPR designs involve additional design steps and constraints, they often have a higher FPGA compilation (RTL-to-bitstream) runtime compared to equivalent monolithic designs. In this work, we present PR-ESP, an open-source platform for a system-level design flow of partially reconfigurable FPGA-based SoC architectures targeting embedded applications that are deployed on resource-constrained FPGAs. Our approach is realized by combining SoC design methodologies and tools from the open-source ESP platform with a fully-automated DPR flow that features a novel size-driven technique for parallel FPGA compilation. We also developed a software runtime reconfiguration manager on top of Linux. Finally, we evaluated our proposed platform using the WAMI-App benchmark application on Xilinx VC707. Biruk B. Seyoum, Davide Giri, Kuan-Lin Chiu, Bryce Natter, Luca P. Carloni |
DATE | 2 |
| 2022 | Work-in-Progress: An Open-Source Platform for Design and Programming of Partially Reconfigurable Heterogeneous SoCsabstractDynamic partial reconfiguration (DPR) enables the design and implementation of flexible, scalable and robust adaptive systems. We present an FPGA-based DPR flow for partially reconfigurable heterogeneous SoCs that uses an incremental compilation technique to reduce the total FPGA compilation time. Biruk B. Seyoum, Davide Giri, Kuan-Lin Chiu, Luca P. Carloni |
CASES | 2 |
| 2022 | A Scalable Methodology for Agile Chip Development with Open-Source Hardware ComponentsabstractWe present a scalable methodology for the agile physical design of tile-based heterogeneous system-on-chip (SoC) architectures that simplifies the reuse and integration of open-source hardware components. The methodology leverages the regularity of the on-chip communication infrastructure, which is based on a multi-plane network-on-chip (NoC), and the modularity of socket interfaces, which connect the tiles to the NoC. Each socket also provides its tile with a set of platform services, including independent clocking and voltage control. As a result, the physical design of each tile can be decoupled from its location in the top-level floorplan of the SoC and the overall SoC design can benefit from a hierarchical timing-closure flow, design reuse and, if necessary, fast respin. With the proposed methodology we completed two SoC tapeouts of increasing complexity, which illustrate its capabilities and the resulting gains in terms of design productivity. Maico Cassel, Martin Cochet, Karthik Swaminathan, Joseph Zuckerman, Paolo Mantovani, Davide Giri, Jeff Zhang 0001, Erik Jens Loscalzo, Gabriele Tombesi, Kevin Tien, Nandhini Chandramoorthy, John-David Wellman, David Brooks 0001, Gu-Yeon Wei, Kenneth L. Shepard, Luca P. Carloni, Pradip Bose |
ICCAD | 7 |
| 2021 | MasterMind: Many-Accelerator SoC Architecture for Real-Time Brain-Computer InterfacesabstractHierarchical Wasserstein Alignment (HiWA) is one of the most promising Brain-Computer Interface algorithms. To enable its real-time communication with the brain and meet low-power requirements, we design and prototype a Linux-supporting, RISC-V based SoC that integrates multiple hardware accelerators. We conduct a thorough design-space exploration at the accelerator level and at the SoC level. With FPGA-based experiments, we show that one of our area-efficient SoCs provides 91x performance and 37x energy efficiency gains over software execution on an embedded processor. We further improve our gains (up to 3408x and 497x, respectively) by parallelizing the workload on multiple accelerator instances and by adopting point-to-point accelerator communication, which reduces memory accesses and software-synchronization overheads. The results include comparisons with multi-threaded software implementations of HiWA running on an Intel i7 and ARM A53 as well as a projection analysis showing that an ASIC implementation of our SoC would meet the needs of real-time Brain-Computer Interfaces. Guy Eichler, Luca Piccolboni, Davide Giri, Luca P. Carloni |
ICCD | 3 |
| 2021 | Cohmeleon: Learning-Based Orchestration of Accelerator Coherence in Heterogeneous SoCsabstractOne of the most critical aspects of integrating loosely-coupled accelerators in heterogeneous SoC architectures is orchestrating their interactions with the memory hierarchy, especially in terms of navigating the various cache-coherence options: from accelerators accessing off-chip memory directly, bypassing the cache hierarchy, to accelerators having their own private cache. By running real-size applications on FPGA-based prototypes of many-accelerator multi-core SoCs, we show that the best cache-coherence mode for a given accelerator varies at runtime, depending on the accelerator’s characteristics, the workload size, and the overall SoC status. Joseph Zuckerman, Davide Giri, Jihye Kwon, Paolo Mantovani, Luca P. Carloni |
MICRO | 2 |
| 2020 | ESP4ML: Platform-Based Design of Systems-on-Chip for Embedded Machine LearningabstractWe present ESP4ML, an open-source system-level design flow to build and program SoC architectures for embedded applications that require the hardware acceleration of machine learning and signal processing algorithms. We realized ESP4ML by combining two established open-source projects (ESP and HLS4ML) into a new, fully-automated design flow. For the SoC integration of accelerators generated by HLS4ML, we designed a set of new parameterized interface circuits synthesizable with high-level synthesis. For accelerator configuration and management, we developed an embedded software runtime system on top of Linux. With this HW/SW layer, we addressed the challenge of dynamically shaping the data traffic on a network-on-chip to activate and support the reconfigurable pipelines of accelerators that are needed by the application workloads currently running on the SoC. We demonstrate our vertically-integrated contributions with the FPGA-based implementations of complete SoC instances booting Linux and executing computer-vision applications that process images taken from the Google Street View database. Davide Giri, Kuan-Lin Chiu, Giuseppe Di Guglielmo, Paolo Mantovani, Luca P. Carloni |
DATE | 1 |
| 2020 | Agile SoC Development with Open ESP : Invited PaperabstractESP is an open-source research platform for heterogeneous SoC design. The platform combines a modular tile-based architecture with a variety of application-oriented flows for the design and optimization of accelerators. The ESP architecture is highly scalable and strikes a balance between regularity and specialization. The companion methodology raises the level of abstraction to system-level design and enables an automated flow from software and hardware development to full-system prototyping on FPGA. For application developers, ESP offers domain-specific automated solutions to synthesize new accelerators for their software and to map complex workloads onto the SoC architecture. For hardware engineers, ESP offers automated solutions to integrate their accelerator designs into the complete SoC. Conceived as a heterogeneous integration platform and tested through years of teaching at Columbia University, ESP supports the open-source hardware community by providing a flexible platform for agile SoC development. Paolo Mantovani, Davide Giri, Giuseppe Di Guglielmo, Luca Piccolboni, Joseph Zuckerman, Emilio G. Cota, Michele Petracca, Christian Pilato, Luca P. Carloni |
ICCAD | 2 |
| 2020 | MosaicSim: A Lightweight, Modular Simulator for Heterogeneous SystemsabstractAs Moore's Law has slowed and Dennard Scaling has ended, architects are increasingly turning to heterogeneous parallelism and domain-specific hardware-software co-designs. These trends present new challenges for simulation-based performance assessments that are central to early-stage architectural exploration. Simulators must be lightweight to support rich heterogeneous combinations of general purpose cores and specialized processing units. They must also support agile exploration of hardware-software co-design, i.e. changes in the programming model, compiler, ISA, and specialized hardware. To meet these challenges, we introduce MosaicSim, a lightweight, modular simulator for heterogeneous systems, offering accuracy and agility designed specifically for hardware-software co-design explorations. By integrating the LLVM toolchain, MosaicSim enables efficient modeling of instruction dependencies and flexible additions across the stack. Its modularity also allows the composition and integration of different hardware components. We first demonstrate that MosaicSim captures architectural bottlenecks in applications, and accurately models both scaling trends in a multicore setting and accelerator behavior. We then present two case-studies where MosaicSim enables straightforward design space explorations for emerging systems, i.e. data science application acceleration and heterogeneous parallel architectures. Opeoluwa Matthews, Aninda Manocha, Davide Giri, Marcelo Orenes-Vera, Esin Tureci, Tyler Sorensen 0001, Tae Jun Ham, Juan L. Aragón, Luca P. Carloni, Margaret Martonosi |
ISPASS | 3 |
| 2019 | Runtime reconfigurable memory hierarchy in embedded scalable platformsabstractIn heterogeneous systems-on-chip, the optimal choice of the cache-coherence model for a loosely-coupled accelerator may vary at each invocation, depending on workload and system status. We propose a runtime adaptive algorithm to manage the coherence of accelerators. The algorithm's choices are based on the combination of static and dynamic features of the active accelerators and their workloads. We evaluate the algorithm by leveraging our FPGA-based platform for rapid SoC prototyping. Experimental results, obtained through the deployment of a multi-core and multi-accelerator system that runs Linux SMP, show the benefits of our approach in terms of execution time and memory accesses. Davide Giri, Paolo Mantovani, Luca P. Carloni |
ASP-DAC | 1 |
| 2018 | NoC-Based Support of Heterogeneous Cache-Coherence Models for AcceleratorsabstractOn-chip shared memory is the primary paradigm for multi-core SoC designs and poses the most critical challenges to their scalability. Choosing the appropriate coherence model for accelerators not only can improve the overall system performance, but can also decrease energy consumption by reducing the accesses to DRAM. We propose an extension of a standard directory-based cache-coherence protocol and present its design as part of a scalable memory hierarchy implemented over a NoC. To evaluate our contribution we designed a many-accelerator SoC architecture that can support three main cache-coherence models for accelerators: non-coherent, last-level-cache-coherent, and fully-coherent. This SoC can run Linux SMP with split last-level cache and multiple DRAM controllers. Our FPGA-based experiments show that the optimal cache-coherence selection varies at run-time, based on the running accelerators and the memory footprint of the applications. Therefore, we support runtime selection of the cache-coherence model for each accelerator, as an alternative to a design-time decision. Davide Giri, Paolo Mantovani, Luca P. Carloni |
NOCS | 1 |
| 2018 | Parallel and Serial Computation in Nanomagnet Logic: An Overview
Davide Giri, Giovanni Causapruno, Fabrizio Riente |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |