EDBT 2026 Demo / reviewers in the wild / expert
Scott Davidson 0004
dblp:55/6396-4
· DBLP profile ↗
6ranked-venue papers
0as first author
3since 2021 · last 2026
0000-0002-9390-6084ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 3 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HammerBlade in Silicon: A 12-nm 2048-Core RISC-V Manycore SoC With Ruche NetworksabstractThis brief presents the 99 mm2, 12-nm FinFET, 2048-core HammerBlade (HB) RISC-V manycore system-on-chip (SoC) which includes the world’s first implementation of Ruche Networks, a wire-maximal network-on-chip (NoC) topology. Prior architecture research argues for HB’s physical and logical scalability, programmability, and high density; this brief shows a concrete realization in silicon. This brief demonstrates strong results running in silicon on 2048 cores over a diverse set of representative parallelized applications, and also a world record on CoreMark (CM) score. This brief further describes in detail the silicon implementation of Ruche Networks. Operating at 1.49 GHz at 0.8 V, the NoC delivers a peak aggregate bandwidth of 2572.6 Tb/s and a bisection bandwidth of 53.2 Tb/s. The design achieves an exceptionally high routing density of 4289 bit/mm. Finally, this brief demonstrates how a small team of Ph.D. students overcame large-chip backend computer-aided design (CAD) challenges using on-chip source-synchronous interconnect (OCSSI), a globally asynchronous locally synchronous (GALS) style top-level integration methodology, combined with a turnaround-time optimized hierarchical design flow. Paul Gao 0001, Dai Cheol Jung, Scott Davidson 0004, Daniel Ruelas-Petrisko, Yuan-Mao Chueh, Max Ruttenberg, Kangli Li, Farzam Gilani, Dustin Richmond, Mark Oskin, Michael B. Taylor |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | ReaLLM: A Trace-Driven Framework for Rapid Simulation of Large-Scale LLM InferenceabstractAs Large Language Models (LLMs) continue to scale, optimizing their deployment requires efficient hardware and system co-design. However, current LLM performance evaluation frameworks fail to capture both chip-level execution details and system-wide behavior, making it difficult to assess realistic performance bottlenecks. In this work, we introduce ReaLLM, a trace-driven simulation framework designed to bridge the gap between detailed accelerator design and large-scale inference evaluation. Unlike prior simulators, ReaLLM integrates kernel profiling derived from detailed microarchitectural simulations with a new trace-driven end-to-end system simulator, enabling precise evaluation of parallelism strategies, batching techniques, and scheduling policies. To address the high computational cost of exhaustive simulations, ReaLLM constructs a precomputed kernel library based on hypothesized scenarios, interpolating results to efficiently explore a vast design space of LLM inference systems. Our validation against real hardware demonstrates the framework's accuracy, achieving an average end-to-end latency prediction error of only 9.1% when simulating inference tasks running on 4 NVIDIA H100 GPUs. We further use ReaLLM to evaluate popular LLMs' end-to-end performance across traces from different applications and identify key system bottlenecks, showing that modern GPU-based LLM inference is increasingly compute-bound rather than memory-bandwidth bound at large scale. Additionally, we significantly reduce simulation time with our precomputed kernel library by a factor of$6 \times$for full-simulations and$164 \times$for workload SLO exploration. ReaLLM is open-source and available at https://github.com/bespoke-silicon-group/reallm. Huwan Peng, Scott Davidson 0004, Chuanjin Richard Shi, Michael B. Taylor |
ASAP | 2 |
| 2024 | Scalable, Programmable and Dense: The HammerBlade Open-Source RISC-V ManycoreabstractExisting tiled manycore architectures propose to convert abundant silicon resources into general-purpose parallel processors with unmatched computational density and programmability. However, as we approach 100 K cores in one chip, conventional manycore architectures struggle to navigate three key axes: scalability, programmability, and density. Many manycores sacrifice programmability for density; or scalability for programmability. In this paper, we explore HammerBlade, which simultaneously achieves scalability, programmability and density. HammerBlade is a fully open-source RISC-V manycore architecture, which has been silicon-validated with a 2048-core ASIC implementation using a 14/16nm process. We evaluate the system using a suite of parallel benchmarks that captures a broad spectrum of computation and communication patterns. Dai Cheol Jung, Max Ruttenberg, Paul Gao 0001, Scott Davidson 0004, Daniel Ruelas-Petrisko, Kangli Li, Aditya K. Kamath, Shaolin Xie, Peitian Pan, Zhongyuan Zhao 0004, Zichao Yue, Bandhav Veluri, Sripathi Muralitharan, Adrian Sampson, Andrew Lumsdaine, Zhiru Zhang, Christopher Batten, Mark Oskin, Dustin Richmond, Michael B. Taylor |
ISCA | 4 |
| 2020 | Ruche Networks: Wire-Maximal, No-Fuss NoCs : Special Session PaperabstractNetwork-On-Chip design has been an active area of academic research for two decades, but many proposed ideas have not been adopted in real chips because they have complex behavior or create significant risks in chip implementation. For this reason, many existing chips just employ fast, replicated vanilla dimension-ordered mesh NoCs. However, these networks do not come close to utilizing the full available VLSI wiring capabilities, and propagate packets at speeds that are significantly below the raw speed of wires. The ideal network would not require any custom circuits, and would decompose easily into a hierarchical CAD flow consisting of a top-level design instantiating a mesh of identical hardened tiles with short-wire neighbor connections. At the same time, this ideal network would easily scale to efficiently utilize the majority of the available chip wiring resources, and would offer a mechanism for scaling this wire usage up or down based on available bandwidth. Packets would spend a significant fraction of their time in wire delay rather than router delay. Finally, the NoC would be simple to understand. This paper proposes Ruche Networks, which fulfill these requirements. They are based on simple 2-D mesh networks but amplify the NoC bandwidth and reduce NoC diameter of tiled architectures by adding long-range physical channels from each tile to other tiles on the same row or column. The more distant the connections, the greater the bandwidth of the network and the lower the diameter. The distance is typically increased until all of the physical VLSI wiring bandwidth have been absorbed. We explain the rational for this “ruching” and provide a simple methodology for designing and implementing these networks using a standard cell VLSI CAD flow. In this paper, we show the steps involved in ruching the HammerBlade Manycore's mesh networks; these steps can easily apply to other designs. Dai Cheol Jung, Scott Davidson 0004, Dustin Richmond, Michael B. Taylor |
NOCS | 2 |
| 2020 | NoC Symbiosis : (Special Session Paper)abstractConventional wisdom states that Network-on-Chip router area grows quadratically with the channel width, and this perception has fundamentally shaped the assumptions of thousands of NoC papers that have been written to date, and many chip designs. However, this assumption is not entirely true. Simple analysis and empirical data from this paper shows that, in modern standard cell technology, a router's standard cell logic area actually grows only linearly; it is solely the wire routing area that grows quadratically.If we think of a NoC as a standalone block as is done in standard hierarchical VLSI design, then the overall area growth is indeed quadratic. But this approach either vastly under-utilizes logic area, or, in designs that match wire and logic area, leads to small network links. At the same time, many standard non-NoC logic blocks like processors or accelerator blocks typically use the standard cell logic area but need only a fraction of available wiring resources.We propose an alternative approach, NoC Symbiosis, in which router logic and the node logic it services are jointly placed together. The router absorbs excess wiring resources from the node logic, and the node logic absorbs excess standard cell area from the router. Current-day automatic place and route (APR) tools already automatically distribute the router logic across the node logic, in order to provide enough space for the wiring resources. With this approach, future SoC's can leverage vastly larger amounts of wiring bandwidth than ever before, or alternatively, reduce the area overhead of existing routers.We describe how we first encountered this phenomena, perform experiments to demonstrate its behavior, and provide design tips to help teams realize the potential of NoC Symbiosis. Daniel Ruelas-Petrisko, Scott Davidson 0004, Paul Gao 0001, Dustin Richmond, Michael B. Taylor |
NOCS | 3 |
| 2018 | Hiding Intermittent Information Leakage with Architectural Support for BlinkingabstractAs demonstrated by numerous practical attacks, the physical act of computation emits unintended and damaging information through infinitesimal variations in timing, power, and resource contention. While there are many techniques for preventing the leakage of information through power channels for specific cryptographic units, they are typically either built directly into the hardware logic or exploit intricate mathematical properties of the algorithm itself. However, such leaks are not uniform in time but, as we show, rather occur in specific bursts. Exploiting this observation we propose a set of software-controlled techniques allowing for the seamless disconnection and reconnection of general purpose programmable components in a system-on-chip. Such a system is capable of providing brief moments of electrical isolation during which the most critical computations can be performed free from both timing and power measurement. Of course, disconnection comes at a cost. To balance the resulting trade-off between overhead and security effectively, we describe a new analysis technique to uncover the "leakiest" intervals of time, we provide an algorithm to co-optimize the covering of these intervals and the performance/energy costs under a set of architecture imposed constraints, and explore the architectural and software ramifications of such intermittent disconnection. In the end we find that by hiding only between 15% and 30% of the trace, at a performance cost of between 15% and 50%, we are able to reduce the mutual information between the leakage model and key bits by 75% on average, and to nearly zero in specific cases. Alric Althoff, Joseph McMahan, Luis Vega, Scott Davidson 0004, Timothy Sherwood, Michael B. Taylor, Ryan Kastner |
ISCA | 4 |