EDBT 2026 Demo / reviewers in the wild / expert
Jerry Zhao
dblp:08/5182
· DBLP profile ↗
16ranked-venue papers
3as first author
7since 2021 · last 2024
0000-0002-9307-2956ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 9 · 6 since 2021Computer networks · 5 · 2 first-authorSoftware engineering, systems software and programming languages · 4 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Zoomie: A Software-like Debugging Tool for FPGAsabstractFPGA prototyping has long been an indispensable technique in pre-silicon verification as well as enabling early-stage software development. FPGAs themselves have also gained popularity as hardware accelerators deployed in datacenters. However, FPGA development brings a plethora of problems. These issues constitute a high barrier towards mass adoption of agile development surrounding FPGA-based projects. Tianrui Wei, Kevin Laeufer, Katie Lim, Jerry Zhao, Koushik Sen, Jonathan Balkind, Krste Asanovic |
ASPLOS (3) | 4 |
| 2024 | NeCTAr and RASoC: Tale of Two Class SoCs for Language Model Interference and Robotics in Intel 16abstractThis paper introduces NeCTAr (Near-Cache Transformer Accelerator), a 16nm heterogeneous multicore RISC-V SoC for sparse and dense machine learning kernels with both near-core and near-memory accelerators. A prototype chip runs at 400MHz at 0.85V and performs matrix-vector multiplications with 109 GOPs/W. The effectiveness of the design is demonstrated by running inference on a sparse language model, ReLU-Llama. Viansa Schmulbach, Ethan Gao, Nikhil Jha, Ethan Wu, Oliver Yu, Ben Oliveau, Brendan Roberts, Connor McMahon, Lixiang Yin, Vamber Yang, Brendan Brenner, George Moujaes, Boyu Hao, Lucy Revina, Bryan Ngo, Yufeng Chi, Hongyi Huang, Reza Sajadiany, Raghav Gupta 0001, Ella Schwarz, Jennifer Zhou, Ken Ho, Jerry Zhao, Anita Flynn, Borivoje Nikolic |
HCS | 26 |
| 2023 | CDPU: Co-designing Compression and Decompression Processing Units for Hyperscale SystemsabstractGeneral-purpose lossless data compression and decompression ("(de)compression") are used widely in hyperscale systems and are key "datacenter taxes". However, designing optimal hardware compression and decompression processing units ("CDPUs") is challenging due to the variety of algorithms deployed, input data characteristics, and evolving costs of CPU cycles, network bandwidth, and memory/storage capacities. Sagar Karandikar, Aniruddha N. Udipi, Junsun Choi, Joonho Whangbo, Jerry Zhao, Svilen Kanev, Edwin Lim, Jyrki Alakuijala, Vrishab Madduri, Sophia Shao, Borivoje Nikolic, Krste Asanovic, Parthasarathy Ranganathan |
ISCA | 5 |
| 2023 | AuRORA: Virtualized Accelerator Orchestration for Multi-Tenant WorkloadsabstractWith the widespread adoption of deep neural networks (DNNs) across applications, there is a growing demand for DNN deployment solutions that can seamlessly support multi-tenant execution. This involves simultaneously running multiple DNN workloads on heterogeneous architectures with domain-specific accelerators. However, existing accelerator interfaces directly bind the accelerator’s physical resources to user threads, without an efficient mechanism to adaptively re-partition available resources. This leads to high programming complexities and performance overheads due to sub-optimal resource allocation, making scalable many-accelerator deployment impractical. Seah Kim, Jerry Zhao, Krste Asanovic, Borivoje Nikolic, Sophia Shao |
MICRO | 2 |
| 2021 | Gemmini: Enabling Systematic Deep-Learning Architecture Evaluation via Full-Stack IntegrationabstractDNN accelerators are often developed and evaluated in isolation without considering the cross-stack, system-level effects in real-world environments. This makes it difficult to appreciate the impact of Systemon-Chip (SoC) resource contention, OS overheads, and programming-stack inefficiencies on overall performance/energy-efficiency. To address this challenge, we present Gemmini, an open-source, full-stack DNN accelerator generator. Gemmini generates a wide design-space of efficient ASIC accelerators from a flexible architectural template, together with flexible programming stacks and full SoCs with shared resources that capture system-level effects. Gemmini-generated accelerators have also been fabricated, delivering up to three orders-of-magnitude speedups over high-performance CPUs on various DNN benchmarks. Hasan Genc, Seah Kim, Alon Amid, Ameer Haj-Ali, Vighnesh Iyer, Pranav Prakash, Jerry Zhao, Daniel Grubb, Harrison Liew, Howard Mao, Albert J. Ou, Colin Schmidt 0001, Samuel Steffl, John Charles Wright, Ion Stoica, Jonathan Ragan-Kelley, Krste Asanovic, Borivoje Nikolic, Sophia Shao |
DAC | 7 |
| 2021 | COBRA: A Framework for Evaluating Compositions of Hardware Branch PredictorsabstractWe present COBRA, a framework which enables a realistic hardware-guided methodology for evaluating compositions of hardware branch predictors. COBRA provides a common interface for developing RTL implementations of predictor subcomponents, as well as a predictor composer that automatically generates hardware predictor pipelines from sub-components based on a high-level topological model of a desired algorithm. We demonstrate how COBRA aids in the design and evaluation of diverse predictor architectures and how our hardware-centric approach captures concerns in predictor characterization that are not exposed in software-based algorithm development. Using COBRA, we generate three superscalar pipelined branch predictors with diverse architectures, synthesize them to run at 1 GHz on a commercial FinFET process, integrate them with the open-source BOOM out-of-order core, and evaluate their end-to-end performance on workloads over trillions of cycles. The COBRA generator system has been open-sourced as part of the SonicBOOM out-of-order core. Jerry Zhao, Abraham Gonzalez, Alon Amid, Sagar Karandikar, Krste Asanovic |
ISPASS | 1 |
| 2021 | A Hardware Accelerator for Protocol BuffersabstractSerialization frameworks are a fundamental component of scale-out systems, but introduce significant compute overheads. However, they are amenable to acceleration with specialized hardware. To understand the trade-offs involved in architecting such an accelerator, we present the first in-depth study of serialization framework usage at scale by profiling Protocol Buffers (“protobuf”) usage across Google’s datacenter fleet. We use this data to build HyperProtoBench, an open-source benchmark representative of key serialization-framework user services at scale. In doing so, we identify key insights that challenge prevailing assumptions about serialization framework usage. Sagar Karandikar, Chris Leary, Chris Kennelly, Jerry Zhao, Dinesh Parimi, Borivoje Nikolic, Krste Asanovic, Parthasarathy Ranganathan |
MICRO | 4 |
| 2020 | Invited: Chipyard - An Integrated SoC Research and Implementation EnvironmentabstractContinued improvement in computing efficiency requires functional specialization of hardware designs. We present an agile design flow for custom SoCs using the Chipyard framework, an integrated SoC research and implementation environment for custom systems. Chipyard includes configurable, composable, open-source, generator-based designs that can be used across multiple stages of the hardware development flow while maintaining design intent and integration consistency. Through cloud FPGA simulation and rapid ASIC implementation, we demonstrate an iterative agile hardware design cycle which enables continuous validation of physically-realizable customized systems. Alon Amid, David Biancolin, Abraham Gonzalez, Daniel Grubb, Sagar Karandikar, Harrison Liew, Albert Magyar, Howard Mao, Albert J. Ou, Nathan Pemberton, Paul Rigge, Colin Schmidt 0001, John Charles Wright, Jerry Zhao, Jonathan Bachrach, Sophia Shao, Borivoje Nikolic, Krste Asanovic |
DAC | 14 |
| 2019 | Simmani: Runtime Power Modeling for Arbitrary RTL with Automatic Signal SelectionabstractThis paper presents a novel runtime power modeling methodology which automatically identifies key signals for power dissipation of any RTL design. The toggle-pattern matrix is constructed with the VCD dumps from a training set, where each signal is represented as a high-dimensional point. By clustering signals showing similar switching activities, a small number of signals are automatically selected, and then the design-specific but workload-independent activity-based power model is constructed using regression against cycle-accurate power traces obtained from industry-standard CAD tools. We can also automatically instrument an FPGA-accelerated RTL simulation with runtime activity counters to obtain power traces of realistic workloads at speed. Our methodology is demonstrated with a heterogeneous processor composed of an in-order core and a custom vector accelerator, running not only microbenchmarks but also real-world machine-learning applications. Jerry Zhao, Jonathan Bachrach, Krste Asanovic |
MICRO | 2 |
| 2008 | STDCS: A Spatio-Temporal Data-Centric Storage Scheme For Real-Time Sensornet ApplicationsabstractSensor networks will shortly consist of globally deployed sensors providing real-time geo-centric information to users. Particularly, users with mobile devices will issue ad-hoc queries usually from within, or nearby, the queried area. In this paper, we propose an in- network data-centric storage (DCS) scheme, namely the spatio-temporal data-centric storage (STDCS) scheme, to efficiently answer these mobile user queries. STDCS is designed to maintain load-balancing among sensors to cope with query hotspots in the network. It is different from previous proposals in two aspects. First, it is based on the novel idea of using a temporally evolving spatial indexing scheme to balance querying load among sensors. Furthermore, STDCS uses dynamic mechanisms for query hotspot detection and decomposition. We conducted extensive simulations that showed our scheme's superiority to both local storage and spatial indexing (the only known geo-centric storage schemes) whenever experiencing query hotspots of different sizes. Mohamed Aly 0002, Anandha Gopalan, Jerry Zhao, Adel M. Youssef |
SECON | 3 |
| 2005 | Towards a Sensor Network Architecture: Lowering the Waistline
David E. Culler, Prabal Dutta, Cheng Tien Ee, Rodrigo Fonseca, Jonathan W. Hui, Philip Alexander Levis, Joseph Polastre, Scott Shenker, Ion Stoica, Gilman Tolle, Jerry Zhao |
HotOS | 11 |
| 2005 | Beacon Vector Routing: Scalable Point-to-Point Routing in Wireless Sensornets
Rodrigo Fonseca, Sylvia Ratnasamy, Jerry Zhao, Cheng Tien Ee, David E. Culler, Scott Shenker, Ion Stoica |
NSDI | 3 |
| 2005 | A unifying link abstraction for wireless sensor networksabstractRecent technological advances and the continuing quest for greater efficiency have led to an explosion of link and network protocols for wireless sensor networks. These protocols embody very different assumptions about network stack composition and, as such, have limited interoperability. It has been suggested [3] that, in principle, wireless sensor networks would benefit from a unifying abstraction (or "narrow waist" in architectural terms), and that this abstraction should be closer to the link level than the network level. This paper takes that vague principle and turns it into practice, by proposing a specific unifying sensornet protocol (SP) that provides shared neighbor management and a message pool.The two goals of a unifying abstraction are generality and efficiency: it should be capable of running over a broad range of link-layer technologies and supporting a wide variety of network protocols, and doing so should not lead to a significant loss of efficiency. To investigate the extent to which SP meets these goals, we implemented SP (in TinyOS) on top of two very different radio technologies: B-MAC on mica2 and IEEE 802.15.4 on Telos. We also built a variety of network protocols on SP, including examples of collection routing [53], dissemination [26], and aggregation [33]. Measurements show that these protocols do not sacrifice performance through the use of our SP abstraction. Joseph Polastre, Jonathan W. Hui, Philip Alexander Levis, Jerry Zhao, David E. Culler, Scott Shenker, Ion Stoica |
SenSys | 4 |
| 2004 | Networking issues in wireless sensor networks
Deepak Ganesan, Alberto Cerpa, Wei Ye 0003, Jerry Zhao, Deborah Estrin |
J. Parallel Distributed Comput. | 5 |
| 2003 | Understanding packet delivery performance in dense wireless sensor networksabstractWireless sensor networks promise fine-grain monitoring in a wide variety of environments. Many of these environments (e.g., indoor environments or habitats) can be harsh for wireless communication. From a networking perspective, the most basic aspect of wireless communication is the packet delivery performance: the spatio-temporal characteristics of packet loss, and its environmental dependence. These factors will deeply impact the performance of data acquisition from these networks.In this paper, we report on a systematic medium-scale (up to sixty nodes) measurement of packet delivery in three different environments: an indoor office building, a habitat with moderate foliage, and an open parking lot. Our findings have interesting implications for the design and evaluation of routing and medium-access protocols for sensor networks. Jerry Zhao, Ramesh Govindan |
SenSys | 1 |
| 2002 | Residual energy scan for monitoring sensor networksabstractIt is important to have continuously updated information about network resources and application activities in a wireless sensor network after it is deployed in an unpredictable environment. Such information can help notify users of resource depletion or abnormal activities. However, constrained by the low user-to-node ratio, limited energy and bandwidth resources, it is infeasible to extract the state of each individual node. In this paper, we propose an approach to constructing abstracted scans of sensor network health by applying in-network aggregation of network state. Specifically, we design a residual energy scan which approximately depicts the remaining energy distribution within a sensor network. Simulations show that our approach has good scalability and energy-efficiency characteristics, compared to continuously extracting the residual energy level individually from each node. Jerry Zhao, Ramesh Govindan, Deborah Estrin |
WCNC | 1 |