EDBT 2026 Demo / reviewers in the wild / expert
Guowei Zhang 0002
dblp:70/6271-2
· DBLP profile ↗
8ranked-venue papers
4as first author
5since 2021 · last 2025
0000-0003-1034-2306ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 4 first-author · 5 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hierarchical Prefetching: A Software-Hardware Instruction Prefetcher for Server ApplicationsabstractThe large working set of instructions in server-side applications causes a significant bottleneck in the front-end, even for high-performance processors equipped with fetch-directed instruction prefetching (FDIP). Prefetchers specifically designed for server scenarios typically rely on a record-and-replay mechanism that exploits the repetitiveness of instruction sequences. However, the efficacy of these techniques is compromised by discrepancies between actual and predicted control flows, resulting in loss of coverage and timeliness. This paper proposes Hierarchical Prefetching, a novel approach that tackles the limitations of existing prefetchers. It identifies common coarse-grained functionality blocks (called Bundles) within the server code and prefetches them as a whole. Bundles are significantly larger than typical prefetch targets, encompassing tens to hundreds of kilobytes of code. The approach combines simple software analysis of code for bundle formation and light-weight hardware for record-and-replay prefetching. The prefetcher requires under 2KB of on-chip storage by keeping most of the metadata in main memory. Experiments with 11 popular server workloads reveal that Hierarchical Prefetching significantly improves miss coverage and timeliness over prior techniques, achieving a 6.6% average performance gain over FDIP. Tingji Zhang, Boris Grot, Wenjian He, Yashuai Lv, Peng Qu 0001, Fang Su, Guowei Zhang 0002, Youhui Zhang |
ASPLOS (2) | 8 |
| 2025 | RICH Prefetcher: Storing Rich Information in Memory to Trade Capacity and Bandwidth for Latency HidingabstractMemory systems characterized by high bandwidth and/or capacity alongside high access latency are becoming increasingly critical.This trend can be observed both at the device level-for instance, in non-volatile memory-and at the system level, as seen in CXL-based memory pooling architectures.To benefit from such memory in general-purpose computing systems, it is essential to employ techniques that can tolerate high memory access latency.Although prefetching has long been recognized as a classical approach for latency tolerance, conventional prefetching techniques are typically either optimized for area efficiency or constrained by limited prefetching patterns.Consequently, they often fail to convert the abundant metadata into significant performance improvements at minimal cost.To address these challenges, we propose RICH-a prefetcher that strategically consumes memory capacity and bandwidth to reduce memory access latency.First, RICH is capable of leveraging abundant metadata to improve performance by integrating spatial prefetching with diverse region sizes and prefetch triggers.Second, RICH implements such metadata with minimal overheads by employing a hierarchical on-chip/off-chip storage mechanism, thereby avoiding both large on-chip storage and critical off-chip accesses.We propose a specific implementation of RICH and evaluate it across a wide range of workloads.With increased memory latency, RICH achieves performance improvements of 8.3% over Bingo and 6.2% over PMP.This highlights the RICH's suitability for future memory systems.In a conventional system, RICH still outperforms Bingo by 3.4%. Ningzhi Ai, Wenjian He, Hu He 0001, Heng Liao, Guowei Zhang 0002 |
MICRO | 6 |
| 2023 | Brief Announcement: Is the Problem-Based Benchmark Suite Fearless with Rust?abstractRust aims to combine safety and performance and claims to provide fearless concurrency. We present a case study to evaluate the extent to which Rust makes parallel programming fearless by porting programs from the C++-based PBBS benchmark suite to Rust. Rust with Rayon provides fearlessness for regular parallelism but not for irregular parallelism. We introduce Rusty-PBBS: a Rust-based benchmark suite with both regular and irregular parallelism. Javad Abdi 0002, Guowei Zhang 0002, Mark C. Jeffrey |
SPAA | 2 |
| 2022 | A scalable architecture for reprioritizing ordered parallelismabstractMany algorithms schedule their work, or tasks, according to a priority order for correctness or faster convergence. While priority schedulers commonly implement task enqueue and dequeueMin operations, some algorithms need a priority update operation that alters the scheduling metadata for a task. Prior software and hardware systems that support scheduling with priority updates compromise on either parallelism, work-efficiency, or both, leading to missed performance opportunities. Moreover, incorrectly navigating these compromises violates correctness in those algorithms that are not resilient to relaxing priority order. Gilead Posluns, Guowei Zhang 0002, Mark C. Jeffrey |
ISCA | 3 |
| 2021 | Gamma: leveraging Gustavson's algorithm to accelerate sparse matrix multiplicationabstractSparse matrix-sparse matrix multiplication (spMspM) is at the heart of a wide range of scientific and machine learning applications. spMspM is inefficient on general-purpose architectures, making accelerators attractive. However, prior spMspM accelerators use inner- or outer-product dataflows that suffer poor input or output reuse, leading to high traffic and poor performance. These prior accelerators have not explored Gustavson's algorithm, an alternative spMspM dataflow that does not suffer from these problems but features irregular memory access patterns that prior accelerators do not support. Guowei Zhang 0002, Nithya Attaluri, Joel S. Emer, Daniel Sánchez 0003 |
ASPLOS | 1 |
| 2019 | Leveraging Caches to Accelerate Hash Tables and MemoizationabstractHash tables are widely used, but they are inefficient in current systems: they use core resources poorly and suffer from limited spatial locality in caches. To address these issues we propose HTA, a technique that accelerates hash table operations via simple ISA extensions and hardware changes. HTA adopts an efficient hash table format that leverages the characteristics of caches. HTA accelerates most operations in hardware, and leaves rare cases to software. Guowei Zhang 0002, Daniel Sánchez 0003 |
MICRO | 1 |
| 2016 | Exploiting semantic commutativity in hardware speculationabstractHardware speculative execution schemes such as hardware transactional memory (HTM) enjoy low run-time overheads but suffer from limited concurrency because they rely on reads and writes to detect conflicts. By contrast, software speculation schemes can exploit semantic knowledge of concurrent operations to reduce conflicts. In particular, they often exploit that many operations on shared data, like insertions into sets, are semantically commutative: they produce semantically equivalent results when reordered. However, software techniques often incur unacceptable run-time overheads. To solve this dichotomy, we present COMMTM, an HTM that exploits semantic commutativity. CommTM extends the coherence protocol and conflict detection scheme to support user-defined commutative operations. Multiple cores can perform commutative operations to the same data concurrently and without conflicts. CommTM preserves transactional guarantees and can be applied to arbitrary HTMs. CommTM scales on many operations that serialize in conventional HTMs, like set insertions, reference counting, and top-K insertions, and retains the low overhead of HTMs. As a result, at 128 cores, CommTM outperforms a conventional eager-lazy HTM by up to 3.4 χ and reduces or eliminates aborts. Guowei Zhang 0002, Virginia Chiu, Daniel Sánchez 0003 |
MICRO | 1 |
| 2015 | Exploiting commutativity to reduce the cost of updates to shared data in cache-coherent systemsabstractWe present Coup, a technique to lower the cost of updates to shared data in cache-coherent systems. Coup exploits the insight that many update operations, such as additions and bitwise logical operations, are commutative: they produce the same final result regardless of the order they are performed in. Coup allows multiple private caches to simultaneously hold update-only permission to the same cache line. Caches with update-only permission can locally buffer and coalesce updates to the line, but cannot satisfy read requests. Upon a read request, Coup reduces the partial updates buffered in private caches to produce the final value. Coup integrates seamlessly into existing coherence protocols, requires inexpensive hardware, and does not affect the memory consistency model. Guowei Zhang 0002, Webb Horn, Daniel Sánchez 0003 |
MICRO | 1 |