EDBT 2026 Demo / reviewers in the wild / expert
Chao Zhang 0039
dblp:94/3019-39
· DBLP profile ↗
6ranked-venue papers
5as first author
2since 2021 · last 2022
0000-0002-7892-5113ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 5 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Memory systems · 66% Hardware accelerators and domain-specific architectures · 13% Processor architecture and microarchitecture · 13% | |
| Software engineering, system software, and programming languages
1 paper |
Compilers and program optimization · 100% |
Topics — the 11 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Processor architecture and microarchitecture › instruction set architecture
ISA extension |
0.6 | 1 | 2022 | ASA: Accelerating Sparse Accumulation in Column-wise SpGEMM · ACM Trans. Archit. Code Optim. 2022 |
Hardware accelerators and domain-specific architectures › sparse matrix multiplication accelerator
sparse linear algebra accelerator |
0.6 | 1 | 2022 | ASA: Accelerating Sparse Accumulation in Column-wise SpGEMM · ACM Trans. Archit. Code Optim. 2022 |
Memory systems › memory hierarchy
cache hierarchy |
0.5 | 1 | 2021 | SPX64: A Scratchpad Memory for General-purpose Microprocessors · ACM Trans. Archit. Code Optim. 2021 |
Memory systems › on-chip memory
scratchpad memory |
0.5 | 1 | 2021 | SPX64: A Scratchpad Memory for General-purpose Microprocessors · ACM Trans. Archit. Code Optim. 2021 |
Memory systems
cache |
0.4 | 1 | 2020 | RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020 |
Memory systems
cache design |
0.4 | 1 | 2020 | Scrabble: A Fine-Grained Cache with Adaptive Merged Block · IEEE Trans. Computers 2020 |
Memory systems › cache › prefetching
hardware prefetching |
0.4 | 1 | 2020 | RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020 |
Memory systems
memory access latency |
0.4 | 1 | 2020 | RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020 |
Parallel and multicore computing › parallel algorithms › parallel matrix algorithms
sparse general matrix-matrix multiplication |
0.2 | 1 | 2022 | ASA: Accelerating Sparse Accumulation in Column-wise SpGEMM · ACM Trans. Archit. Code Optim. 2022 |
Memory systems › cache
cache miss |
0.1 | 1 | 2020 | RnR: A Software-Assisted Record-and-Replay Hardware Prefetcher · MICRO 2020 |
Energy-efficient computing › low-power design
on-chip power reduction |
0.1 | 1 | 2020 | Scrabble: A Fine-Grained Cache with Adaptive Merged Block · IEEE Trans. Computers 2020 |
Methods — techniques the papers use, named apart from their topics
software-assisted hardware prefetching · 0.9record and replay · 0.9set-associative cache search · 0.6pipelined accumulation · 0.6virtual addressing · 0.5set-associative design · 0.5tag sharing · 0.4adaptive block merging · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | ASA: Accelerating Sparse Accumulation in Column-wise SpGEMMabstractSparse linear algebra is an important kernel in many different applications. Among various sparse general matrix-matrix multiplication (SpGEMM) algorithms, Gustavson’s column-wise SpGEMM has good locality when reading input matrix and can be easily parallelized by distributing the computation of different columns of an output matrix to different processors. However, the sparse accumulation (SPA) step in column-wise SpGEMM, which merges partial sums from each of the multiplications by the row indices, is still a performance bottleneck. The state-of-the-art software implementation uses a hash table for partial sum search in the SPA, which makes SPA the largest contributor to the execution time of SpGEMM. There are three reasons that cause the SPA to become the bottleneck: (1) hash probing requires data-dependent branches that are difficult for a branch predictor to predict correctly; (2) the accumulation of partial sum is dependent on the results of the hash probing, which makes it difficult to hide the hash probing latency; and (3) hash collision requires time-consuming linear search and optimizations to reduce these collisions require an accurate estimation of the number of non-zeros in each column of the output matrix. This work proposes ASA architecture to accelerate the SPA. ASA overcomes the challenges of SPA by (1) executing the partial sum search and accumulate with a single instruction through ISA extension to eliminate data-dependent branches in hash probing, (2) using a dedicated on-chip cache to perform the search and accumulation in a pipelined fashion, (3) relying on the parallel search capability of a set-associative cache to reduce search latency, and (4) delaying the merging of overflowed entries. As a result, ASA achieves an average of 2.25× and 5.05× speedup as compared to the state-of-the-art software implementation of a Markov clustering application and its SpGEMM kernel, respectively. As compared to a state-of-the-art hashing accelerator design, ASA achieves an average of 1.95× speedup in the SpGEMM kernel. Chao Zhang 0039, Maximilian H. Bremer, Cy P. Chan, John Shalf |
ACM Trans. Archit. Code Optim. | 1 |
| 2021 | SPX64: A Scratchpad Memory for General-purpose MicroprocessorsabstractGeneral-purpose computing systems employ memory hierarchies to provide the appearance of a single large, fast, coherent memory. In special-purpose CPUs, programmers manually manage distinct, non-coherent scratchpad memories. In this article, we combine these mechanisms by adding a virtually addressed, set-associative scratchpad to a general purpose CPU. Our scratchpad exists alongside a traditional cache and is able to avoid many of the programming challenges associated with traditional scratchpads without sacrificing generality (e.g., virtualization). Furthermore, our design delivers increased security and improves performance, especially for workloads with high locality or that interact with nonvolatile memory. Shail Dave, Pantea Zardoshti, Robert Brotzman, Chao Zhang 0039, Aviral Shrivastava, Gang Tan, Michael F. Spear |
ACM Trans. Archit. Code Optim. | 5 |
| 2020 | ECC Cache: A Lightweight Error Detection for Phase-Change Memory Stuck-At FaultsabstractDRAM scaling has been slowed down. Emerging non-volatile memories (e.g., Phase-Change Memory) promises higher density, better scalability, and persistence. However, endurance is a fundamental issue that hinders the broad adoption of PCM---after repeated writes, a PCM cell can get stuck at a value and be no longer programmable. The prevalence of this stuck-at fault issue requires error detection and correction mechanisms for PCM. Existing solutions such as verify-after-write adds additional latency to PCM writes, which degrades overall system performance. Other solutions like in-memory error-correcting code (ECC) requires a high storage overhead and introduces more reliability issue because ECC bits tend to wear out faster than the protected data bits. Chao Zhang 0039, Khaled Abdelaal, Angel Chen, Xinhui Zhao, Wujie Wen |
ICCAD | 1 |
| 2020 | RnR: A Software-Assisted Record-and-Replay Hardware PrefetcherabstractApplications with irregular memory access patterns do not benefit well from the memory hierarchy as applications that have good locality do. Relatively high miss ratio and long memory access latency can cause the processor to stall and degrade system performance. Prefetching can help to hide the miss penalty by predicting which memory addresses will be accessed in the near future and issuing memory requests ahead of the time. However, software prefetchers add instruction overhead, whereas hardware prefetchers cannot efficiently predict irregular memory access sequences with high accuracy. Fortunately, in many important irregular applications (e.g., iterative solvers, graph algorithms, and sparse matrix-vector multiplication), memory access sequences repeat over multiple iterations or program phases. When the patterns are long, a conventional spatial-temporal prefetcher can not achieve high prefetching accuracy, but these repeating patterns can be identified by programmers.In this work, we propose a software-assisted hardware prefetcher that focuses on repeating irregular memory access patterns for data structures that cannot benefit from conventional hardware prefetchers. The key idea is to provide a programming interface to record cache miss sequence on the first appearance of a memory access pattern and prefetch through replaying the pattern on the following repeats. The proposed Record-and-Replay (RnR) prefetcher provides a lightweight software interface so that the programmers can specify in the application code: 1) which data structures have irregular memory accesses, 2) when to start the recording, and 3) when to start the replay (prefetching). This work evaluated three irregular workloads with different inputs. For the evaluated workloads and inputs, the proposed RnR prefetcher can achieve on average 2.16× speedup for graph applications and 2.91× speedup for an iterative solver with a sparse matrix-vector multiplication kernel. By leveraging the knowledge from the programmers, the proposed RnR prefetcher can achieve over 95% prefetching accuracy and miss coverage. Chao Zhang 0039, Yuan Zeng 0003, John Shalf |
MICRO | 1 |
| 2020 | Scrabble: A Fine-Grained Cache with Adaptive Merged BlockabstractA large fraction of the microprocessor energy is consumed by the data movement in the system. One of the reasons is the inefficiency in the conventional cache design. Cache blocks larger than a word are used in conventional caches to exploit spatial locality. However, many applications only use a small part of a cache block before its eviction. Transferring and storing unused data wastes bandwidth, energy, and limited cache space. Prior work on fine-grained caches can reduce data access and storage granularity to reduce the amount of unused data. However, small data blocks typically require greater metadata and control overhead. Sharing the common bits among tags of fine-grained blocks can reduce the metadata overhead but the constraints on which fine-grained blocks can share tag bits can cause fragmentation. This work proposes scrabble, a fine-grained cache that can merge multiple non-contiguous fine-grained blocks into a variable size merged block. The length of the shared tag is maximized to reduce the metadata overhead. The space utilization is improved by supporting merged blocks with variable size. The control overhead can be reduced by moving the merged block together from memory to the last level cache. For applications with poor spatial locality, Scrabble cache can achieve more than 40 percent of performance improvement. Even for application with good spatial locality, the speedup is still more than 7 percent. In general, for an evaluated set of benchmarks, Scrabble cache achieves an average of 2.41× effective capacity over the baseline cache with the same cache capacity which leads to a 16.7 percent performance improvement and an 11 percent on-chip energy reduction. As compared to a state-of-the-art fine-grained cache, Scrabble cache achieves a 1.25× effective capacity, a 7.9 percent speedup, and a 5.8 percent on-chip energy reduction. Chao Zhang 0039, Yuan Zeng 0003 |
IEEE Trans. Computers | 1 |
| 2017 | Enabling efficient fine-grained DRAM activations with interleaved I/OabstractDRAM contributes a significant part of the total system energy consumption, and row activation is one of the most energy inefficient components. Prior works on fine-grained DRAM activation rely on increasing the number of local wires to avoid degrading performance, which adds area overheads. This work proposes interleaved I/O to allow data transferring from different partially activated banks to share the global I/O. The proposed DRAM architecture allows half-, quarter-, or one-eighth- page activations without changing the wires. The system performance is competitive as compared with other fine-grained activation designs. For the evaluated benchmarks, an average of up to 15.7% performance improvement is achieved among all of the configurations. Furthermore, the total DRAM energy can be reduced by an average of 11.2% for halfpage, 17.2% for quarterpage, and 22.3% for one-eighth-page. Chao Zhang 0039 |
ISLPED | 1 |