VLDB 2026 Research / reviewers in the wild / expert
Heng Zhuo
dblp:242/4798
· DBLP profile ↗
4ranked-venue papers
0as first author
1since 2021 · last 2022
0000-0002-8986-9161ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Emerging computing paradigms · 100% |
Topics — the 3 heaviest of 3, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Compilers and program optimization
compiler optimization |
0.4 | 1 | 2020 | BitSAD v2: Compiler Optimization and Analysis for Bitstream Computing · ACM Trans. Archit. Code Optim. 2020 |
Compilers and program optimization › compiler optimization
hardware-aware optimization |
0.4 | 1 | 2020 | BitSAD v2: Compiler Optimization and Analysis for Bitstream Computing · ACM Trans. Archit. Code Optim. 2020 |
Emerging computing paradigms › approximate and stochastic computing
stochastic computing |
0.4 | 1 | 2020 | BitSAD v2: Compiler Optimization and Analysis for Bitstream Computing · ACM Trans. Archit. Code Optim. 2020 |
Methods — techniques the papers use, named apart from their topics
software emulation · 0.9population coding · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Work-in-Progress: NoRF: A Case Against Register File Operands in Tightly-Coupled AcceleratorsabstractAccelerators are often used to increase performance and/or energy efficiency of general-purpose CPUs. However, Tightly-Coupled Accelerators (TCAs) often perform computations on data structures that may not be a natural fit for general-purpose registers. The designer can either use the existing register file (RF), a RF tailored for the accelerator, or eschew a RF entirely (NoRF), accessing operands directly from the memory hierarchy. Designers for embedded and edge devices are particularly conscientious towards energy-efficient compute and data transfer. We explore the possibility of mini-DGEMM accelerators (example TCAs) within the context of CPUs and edge devices, which also have increasing applications for DGEMM compute. At a high level, register files help reduce memory accesses (steps 1, 2, 5, and 6 in Figure 1 ) when the compiler finds reuse of operands in the program dataflow. On the other hand, direct memory access simplifies the data movement by completely eliminating the intermediate reads and writes to a register file but issues more memory requests. This paper evaluates the difference between these options of operand delivery. Figure 2 shows that all recent vector extensions use a register file implementation. By this trend, it may seem natural to incorporate mini-matrices into the RF. However, we present quantitative and qualitative evidence to advocate for direct cache access for operands. David J. Schlais, Heng Zhuo, Mikko H. Lipasti |
CASES | 2 |
| 2020 | Modeling Architectural Support for Tightly-Coupled AcceleratorsabstractAs proposed accelerators target finer-grained chunks of computation and data movement, it becomes increasingly important to couple them tightly with the processor, avoiding long invocation delays. However, the large implementation design space of these Tightly-Coupled Accelerators (TCAs) makes it difficult to balance trade-offs between hardware complexity and accelerator performance. Previous performance models for accelerators focused on the penalties associated with loosely-coupled accelerators, which abstracted away many of the fine-grained interactions with complex out-of-order structures and program behaviors that have large impacts on TCA performance. In this paper, we introduce an analytical model that studies TCA behavior when interacting with the core, in the context of both high and low memory bandwidth applications supporting various levels of speculative and out of order (OoO) execution. Our analytical model reduces the turnaround time in early design stages when estimating performance gains over detailed simulation with tolerable error. We also discuss potential design choices that can impede the benefits that come with TCAs, and illuminate differences with traditional accelerators. David J. Schlais, Heng Zhuo, Mikko H. Lipasti |
ISPASS | 2 |
| 2020 | BitSAD v2: Compiler Optimization and Analysis for Bitstream ComputingabstractComputer vision and machine learning algorithms operating under a strict power budget require an alternate computing paradigm. While bitstream computing (BC) satisfies these constraints, creating BC systems is difficult. To address the design challenges, we propose compiler extensions to B it SAD, a DSL for BC. Our work enables bit-level software emulation and automated generation of hierarchical hardware, discusses potential optimizations, and proposes compiler phases to implement those optimizations in a hardware-aware manner. Finally, we introduce population coding, a parallelization scheme for stochastic computing that decreases latency without sacrificing accuracy, and provide theoretical and experimental guarantees on its effectiveness. Kyle Daruwalla, Heng Zhuo, Rohit Shukla, Mikko H. Lipasti |
ACM Trans. Archit. Code Optim. | 2 |
| 2019 | BitBench: a benchmark for bitstream computingabstractWith the recent increase in ultra-low power applications, researchers are investigating alternative architectures that can operate on streaming input data. These target use cases require complex algorithms that must be evaluated under a real-time deadline, but also satisfy the strict available power budget. Stochastic computing (SC) is an example of an alternative paradigm where the data is represented as single bitstreams, allowing designers to implement operations such as multiplication using a simple AND gate. Consequently, the resulting design is both low area and low power. Similarly, traditional digital filters can take advantage of streaming inputs to effectively choose coefficients, resulting in a low cost implementation. In this work, we construct six key algorithms to characterize bitstream computing. We present these algorithms as a new benchmark suite: BitBench. Kyle Daruwalla, Heng Zhuo, Carly Schulz, Mikko H. Lipasti |
LCTES | 2 |