EDBT 2026 Demo / reviewers in the wild / expert
Joseph Zuckerman
dblp:274/0943
· DBLP profile ↗
10ranked-venue papers
1as first author
9since 2021 · last 2025
0000-0003-3081-1077ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 1 first-author · 9 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | An Energy-Efficient Kalman Filter Architecture with Tunable Accuracy for Brain-Computer InterfacesabstractKalman Filter (KF) is the most prominent algorithm to predict motion from measurements of brain activity. However, little effort has been made to specialize KF hardware for the unique requirements of embedded brain-computer interfaces (BCIs). For this reason, we present the first configurable KF hardware architecture that enables fine-grained tuning of latency and accuracy, thereby facilitating specialization for neural data processing in BCI applications and supporting design-space exploration. Based on our architecture, we design KF hardware accelerators and integrate them into a heterogeneous system-on-chip (SoC). Through FPGA-based experiments, we demonstrate an energy-efficiency improvement of $15.3 x$ and $10^{3} x$ better accuracy compared to state-of-the-art implementations. Guy Eichler, Joseph Zuckerman, Luca P. Carloni |
DAC | 2 |
| 2025 | KalmMind: A Configurable Kalman Filter Design Framework for Embedded Brain-Computer InterfacesabstractKalman Filter (KF) is one of the most prominent algorithms to predict motion from measurements of brain activity. However, little effort has been made to optimize the KF for deployment in embedded brain-computer interfaces (BCIs). To address this challenge, we propose a new framework for designing KF hardware accelerators specialized for BCI, which facilitates design-space exploration by providing a tunable balance between latency and accuracy. Through FPGA-based experiments with brain data, we demonstrate improvements in both latency and accuracy compared to the state of the art. Guy Eichler, Joseph Zuckerman, Luca P. Carloni |
DATE | 2 |
| 2025 | ReconFormer: A Multi-Level Run-Time Reconfigurable System-on-Chip for Accelerating TransformersabstractRecent advances in neural network design have emphasized attention-based Transformer models, which deliver state-of-the-art accuracy across applications in natural language processing and computer vision. However, the computational demands of Transformers pose significant challenges for deployment in latency-sensitive scenarios and resource-constrained devices, resulting from Transformers' quadratic scaling to sequence length in attention mechanism and extensive data movement. Many works have tried to alleviate these limitations while failing to give comprehensive solutions that consider the heterogeneous characteristics within Transformers. In this paper, we propose ReconFormer, a system-on-chip that accelerates Transformers through multi-level reconfigurability. At component level, ReconFormer introduces reconfigurable processing elements by incorporating range-based approximations for non-linear functions, resulting in 21% reduced logic usage per accelerator. At the system level, ReconFormer has a run-time coarse-grained reconfigurability, which provides adaptive parallelism strategies and dynamic cache-coherence mode to address diverse kernel requirements. As a result, RECONFORMER achieves up to$5.54 \times$speedup in multi-head attention and$3.61 \times$in end-to-end performance improvement over a static system. ReconFormer attains remarkable performance compared to other platforms, resulting in$13.03 \times$and$5.30 \times$higher efficiency than edge and server CPUs, respectively. Even compared to an edge GPU, ReconFormer achieves$1.79 \times$improvement in efficiency, highlighting our system as a compelling alternative for edge deployment. Je Yang, Gabriele Tombesi, Joseph Zuckerman, Luca P. Carloni |
FPL | 3 |
| 2025 | Optimization of Wire Pipelining and Channel Parallelism for 2D-Mesh NoC Physical DesignabstractModern systems-on-chip (SoCs) increasingly rely on high-bandwidth networks-on-chip (NoCs) to support communication among their many heterogeneous components. Meanwhile, as technology nodes advance, physical design (PD) has a growing impact on NoC performance, particularly for high-bandwidth NoCs. However, few published works focus on NoC optimization from a PD perspective. In this work, we study the problem of optimizing the PD of 2D-mesh NoCs by focusing on two prominent techniques: wire pipelining and channel parallelism. Our study is based on experimental results obtained from multiple tape-in NoC designs in a 12 nm technology process. We develop models to approximate the power and area effects of different NoC design approaches and analyze the underlying trends. Our findings show that pipelining does not affect die area but increases NoC power consumption by$1.6 \times$compared to increasing parallelism. In contrast, increasing parallelism can result in a NoC area up to$2 \times$larger than one achieving the same bandwidth through pipelining. Building on these insights, we formulate a mathematical optimization problem, which could be solved by optimization solvers to balance the trade-offs between these two techniques. Our study provides a general framework for analyzing NoC physical design trade-offs and optimizing NoC configurations. Pei-Huan Tsai, Maico Cassel, Joseph Zuckerman, Kuan-Lin Chiu, Luca P. Carloni |
ICCD | 3 |
| 2025 | FLIP2M: Flexible Intra-layer Parallelism and Inter-layer Pipelining for Multi-model AR/VR WorkloadsabstractTiled accelerator architectures provide opportunities to optimize the performance of multi-model augmented and virtual reality (AR/VR) applications through intra-layer parallelism and inter-layer pipelining. However, balancing these two strategies is a difficult task that demands a flexible architecture to deploy models and an optimization approach, that is, capable of selecting an optimal strategy from an enormous mapping space. This article presents FLIP2M, a holistic solution for mapping multi-model AR/VR workloads on tiled architectures. FLIP2M consists of (1) FLIP, an acceleration fabric that supports a wide variety of optimizations through flexible on-chip communication, and (2) OASIS, an optimization framework based on dynamic and constraint programming, that is, capable of selecting an efficient strategy for mapping multi-model workloads onto FLIP. We demonstrate FLIP2M on an FPGA prototype of FLIP that features 36 accelerators and 7 DDR4 controllers. Using OASIS-generated mappings for three different multi-model AR/VR workloads, FLIP2M achieves up to 1.94× improvement in latency, 1.37× in energy, and 2.59× in energy-delay product relative to a FLIP baseline without intra-layer resource allocation flexibility and inter-layer pipelining. Gabriele Tombesi, Je Yang, Joseph Zuckerman, Davide Giri, William Baisi, Luca P. Carloni |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2024 | Mozart: Taming Taxes and Composing Accelerators with Shared-MemoryabstractResource-constrained system-on-chips (SoCs) are increasingly heterogeneous with specialized accelerators for various tasks. Acceleration taxes due to control and data movement, however, diminish end-to-end speedups from hardware acceleration. Meanwhile, emerging workloads are increasingly task-diverse with several, potentially shared, fine-grained acceleration candidates. This motivates a paradigm of parallel and disaggregated acceleration. Compared to a monolithic accelerator, disaggregation provides higher flexibility, reuse, and utilization, but at the cost of higher control and data acceleration taxes. Vignesh Suresh, Bakshree Mishra, Zeran Zhu, Naiyin Jin, Charles Block, Paolo Mantovani, Davide Giri, Joseph Zuckerman, Luca P. Carloni, Sarita V. Adve |
PACT | 9 |
| 2024 | BlitzCoin: Fully Decentralized Hardware Power Management for Accelerator-Rich SoCsabstractOn-chip power-management techniques have evolved over several processor generations. However, response time and scalability constraints have made it difficult to translate existing power-management strategies to current or next-generation System-on-Chip (SoC) architectures, which are expected to comprise tens to hundreds of cores and accelerators. In this work we present BlitzCoin, a fully decentralized hardware power-management strategy for large, accelerator-rich SoCs, coupled with optimized unified voltage and frequency regulation. We evaluated BlitzCoin through RTL simulations of multiple SoCs targeted toward different application domains. The results are further validated through silicon measurements of a fabricated 12 nm many-accelerator SoC that includes BlitzCoin. Our evaluations show that BlitzCoin is markedly faster, with 8× to 12× lower response times, which provides 25%-34% throughput improvement and allows for scaling to 7 × to 13 × larger SoCs compared to state-of-the-art centralized power-management strategies, all with an area overhead of <1%. Martin Cochet, Karthik Swaminathan, Erik Jens Loscalzo, Joseph Zuckerman, Maico Cassel, Davide Giri, Alper Buyuktosunoglu, David Brooks 0001, Gu-Yeon Wei, Kenneth L. Shepard, Luca P. Carloni, Pradip Bose |
ISCA | 4 |
| 2022 | A Scalable Methodology for Agile Chip Development with Open-Source Hardware ComponentsabstractWe present a scalable methodology for the agile physical design of tile-based heterogeneous system-on-chip (SoC) architectures that simplifies the reuse and integration of open-source hardware components. The methodology leverages the regularity of the on-chip communication infrastructure, which is based on a multi-plane network-on-chip (NoC), and the modularity of socket interfaces, which connect the tiles to the NoC. Each socket also provides its tile with a set of platform services, including independent clocking and voltage control. As a result, the physical design of each tile can be decoupled from its location in the top-level floorplan of the SoC and the overall SoC design can benefit from a hierarchical timing-closure flow, design reuse and, if necessary, fast respin. With the proposed methodology we completed two SoC tapeouts of increasing complexity, which illustrate its capabilities and the resulting gains in terms of design productivity. Maico Cassel, Martin Cochet, Karthik Swaminathan, Joseph Zuckerman, Paolo Mantovani, Davide Giri, Jeff Zhang 0001, Erik Jens Loscalzo, Gabriele Tombesi, Kevin Tien, Nandhini Chandramoorthy, John-David Wellman, David Brooks 0001, Gu-Yeon Wei, Kenneth L. Shepard, Luca P. Carloni, Pradip Bose |
ICCAD | 5 |
| 2021 | Cohmeleon: Learning-Based Orchestration of Accelerator Coherence in Heterogeneous SoCsabstractOne of the most critical aspects of integrating loosely-coupled accelerators in heterogeneous SoC architectures is orchestrating their interactions with the memory hierarchy, especially in terms of navigating the various cache-coherence options: from accelerators accessing off-chip memory directly, bypassing the cache hierarchy, to accelerators having their own private cache. By running real-size applications on FPGA-based prototypes of many-accelerator multi-core SoCs, we show that the best cache-coherence mode for a given accelerator varies at runtime, depending on the accelerator’s characteristics, the workload size, and the overall SoC status. Joseph Zuckerman, Davide Giri, Jihye Kwon, Paolo Mantovani, Luca P. Carloni |
MICRO | 1 |
| 2020 | Agile SoC Development with Open ESP : Invited PaperabstractESP is an open-source research platform for heterogeneous SoC design. The platform combines a modular tile-based architecture with a variety of application-oriented flows for the design and optimization of accelerators. The ESP architecture is highly scalable and strikes a balance between regularity and specialization. The companion methodology raises the level of abstraction to system-level design and enables an automated flow from software and hardware development to full-system prototyping on FPGA. For application developers, ESP offers domain-specific automated solutions to synthesize new accelerators for their software and to map complex workloads onto the SoC architecture. For hardware engineers, ESP offers automated solutions to integrate their accelerator designs into the complete SoC. Conceived as a heterogeneous integration platform and tested through years of teaching at Columbia University, ESP supports the open-source hardware community by providing a flexible platform for agile SoC development. Paolo Mantovani, Davide Giri, Giuseppe Di Guglielmo, Luca Piccolboni, Joseph Zuckerman, Emilio G. Cota, Michele Petracca, Christian Pilato, Luca P. Carloni |
ICCAD | 5 |